<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>R-bloggers</title>
	<atom:link href="https://www.r-bloggers.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://www.r-bloggers.com</link>
	<description>R news and tutorials contributed by hundreds of R bloggers</description>
	<lastBuildDate>Wed, 30 Sep 2026 00:00:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=5.5.22</generator>

<image>
	<url>https://i0.wp.com/www.r-bloggers.com/wp-content/uploads/2016/08/cropped-R_single_01-200.png?fit=32%2C32&#038;ssl=1</url>
	<title>R-bloggers</title>
	<link>https://www.r-bloggers.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">11524731</site>	<item>
		<title>BioC2026 conference recap</title>
		<link>https://www.r-bloggers.com/2026/09/bioc2026-conference-recap/</link>
		
		<dc:creator><![CDATA[Laurah Ondari]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; ">
<p>The North American Bioconductor Conference 2026 (([BioC2026])(https://bioc2026.bioconductor.org/)) took place from August 10-12, 2026, at the Fred Hutchinson Cancer Center in Seattle, Washington. The conference brought together researchers, dev...</p></div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/bioc2026-conference-recap/">BioC2026 conference recap</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/"> Bioconductor community blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
 





<p><a href="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/bioc2026_groupshot.jpg?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-1" rel="nofollow" target="_blank"><img src="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/bioc2026_groupshot.jpg?w=578&#038;ssl=1" class="zoomable img-fluid" style="width:100.0%" data-recalc-dims="1"></a></p>
<p>The North American Bioconductor Conference 2026 (([BioC2026])(https://bioc2026.bioconductor.org/)) took place from August 10-12, 2026, at the Fred Hutchinson Cancer Center in Seattle, Washington. The conference brought together researchers, developers, educators, and community members to explore current developments within and beyond the Bioconductor project. BioC2026 welcomed 108 attendees, with 70 participating in person and 38 joining virtually. Across the three days, the programme featured three keynote speakers, 20 contributed speakers, and 17 sessions. The programme highlighted the breadth of biological research supported by Bioconductor, as well as the tools, infrastructure, and communities that enable open and reproducible computational biology. The conference was then followed by two community events on August 13-14: a Carpentries workshop on orchestrating large-scale single-cell analysis with Bioconductor and the Bioconductor North America 2026 Hackathon.</p>
<p><a href="https://i1.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/Closing%20Remarks%202026.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-2" rel="nofollow" target="_blank"><img src="https://i1.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/Closing%20Remarks%202026.png?w=578&#038;ssl=1" class="zoomable img-fluid" style="width:100.0%" data-recalc-dims="1"></a></p>
<section id="participants-by-country" class="level2">
<h2 class="anchored" data-anchor-id="participants-by-country">Participants by country</h2>
<p>Participants travelled to Seattle from across North America and beyond, reflecting the international reach of the Bioconductor community. Attendees joined from the United States, Canada, Germany, Ireland, the United Kingdom, Australia, and Nigeria. The audience included researchers at different stages of their careers, with staff representing 42% of attendees, followed by students and postdoctoral researchers at 28%, faculty at 13.8%, industry at 13.8%, and government at 2.4%.</p>
<iframe src="https://blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/eurobioc2026-participants-map.html" width="450" height="600" frameborder="0">
</iframe>
</section>
<section id="programme-overview" class="level2">
<h2 class="anchored" data-anchor-id="programme-overview">Programme overview</h2>
<section id="keynotes" class="level3">
<h3 class="anchored" data-anchor-id="keynotes">Keynotes</h3>
<p>The conference featured three keynote speakers: &#8211; <a href="https://www.youtube.com/watch?v=FuFSePQXVhM&#038;t=187s" rel="nofollow" target="_blank">Jeff Leek speaking on “Building Open Infrastructure for Translational AI in a Cancer Center”</a> &#8211; <a href="https://www.youtube.com/watch?v=ZG0weEXGvlA&#038;t=21s" rel="nofollow" target="_blank">Ting Ye speaking on “From Association to Causation: Genetic-Anchored Causal Inference in Human Biology”</a> &#8211; <a href="https://www.youtube.com/watch?v=_-sitxM0pNM" rel="nofollow" target="_blank">Michael Lawrence speaking on “S7 for Bioconductor”</a></p>
<div style="display:flex; gap:3px;">
<p><img src="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/leek_keynote.jpeg?w=578&#038;ssl=1" style="width:33.33%; height:200px; object-fit:cover;" data-recalc-dims="1"></p>
<p><img src="https://i2.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/lawrence_keynote.jpeg?w=578&#038;ssl=1" style="width:33.33%; height:200px; object-fit:cover;" data-recalc-dims="1"></p>
<p><img src="https://i2.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/ting_keynote.jpg?w=578&#038;ssl=1" style="width:33.33%; height:200px; object-fit:cover;" data-recalc-dims="1"></p>
</div>
<div style="text-align:center; font-style:italic; font-size:0.9em; margin-top:5px;">
<p>(Left to Right: Jeff Leek, Michael Lawrence and Ting Ye giving their keynotes at BioC2026)</p>
</div>
</section>
<section id="short-talks" class="level3">
<h3 class="anchored" data-anchor-id="short-talks">Short talks</h3>
<p>Twenty contributed speakers shared their work across the conference programme. The short talks at BioC2026 reflected the breadth of research and development across the Bioconductor community. Topics ranged from community training and developer engagement to multi-omics exploration, reproducible workflows, and AI infrastructure. The programme also covered emerging areas of biological data analysis, including spatial transcriptomics, antibody repertoire analysis, rare variant testing, DNA methylation, structural variation, and Perturb-seq.</p>
</section>
<section id="poster-sessions-and-package-demo-sessions" class="level3">
<h3 class="anchored" data-anchor-id="poster-sessions-and-package-demo-sessions">Poster sessions and package demo sessions</h3>
<p>The poster session showcased a diverse range of research applications and software development projects from across the Bioconductor community. Topics included single-cell and spatial transcriptomics, multi-omics integration, microbiome analysis, proteomics, antibody repertoire analysis, functional enrichment, ancestry inference, genomic data processing, and AI-guided drug discovery. The posters provided an opportunity for attendees to discuss new methods and tools in a more informal setting, while connecting researchers, package developers, and community members working across different areas of computational biology. The sessions also complemented the package demonstrations, creating space for participants to explore new software and engage directly with the people developing and applying Bioconductor tools.</p>
<p>Feedback from the conference showed that these interactions were particularly valuable, with attendees highlighting conversations around posters and package demonstrations as some of the most useful networking opportunities.</p>
<img src="https://i2.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/bioc2026_poster_session.jpg?w=578&#038;ssl=1" class="zoomable img-fluid" style="width:100.0%" data-recalc-dims="1">
<center>
<p><em>Participants during the poster session.</em></p>
</center>
</section>
<section id="celebrating-25-years-of-bioconductor" class="level3">
<h3 class="anchored" data-anchor-id="celebrating-25-years-of-bioconductor">Celebrating 25 years of Bioconductor</h3>
<p>The celebrations marking Bioconductor’s 25th anniversary continued at BioC2026 following the celebrations that began at EuroBioC2026 earlier in the year. The milestone was marked with a specially branded anniversary cake, thanking participants and community members for being part of the Bioconductor community and contributing to its growth over the past 25 years.</p>
<div style="display:flex; gap:3px;">
<p><img src="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/bioc_25_cake.jpeg?w=578&#038;ssl=1" style="width:calc(50% - 1.5px); height:250px; object-fit:cover;" data-recalc-dims="1"></p>
<p><img src="https://i2.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/Bioc2026_25years.png?w=578&#038;ssl=1" style="width:calc(50% - 1.5px); height:250px; object-fit:cover;" data-recalc-dims="1"></p>
</div>
<center>
<p><em>Bioconductor at 25 celebrations.</em></p>
</center>
</section>
</section>
<section id="community-and-networking" class="level2">
<h2 class="anchored" data-anchor-id="community-and-networking">Community and networking</h2>
<p>One of the strongest themes to emerge from the BioC2026 survey was the value of the Bioconductor community itself. Participants described the conference as welcoming, collaborative, and interactive, and appreciated the opportunity to reconnect with colleagues and meet new people. All respondents who answered the question said they would recommend BioC2026 to their colleagues. The smaller size of the conference was also seen as a strength. Several respondents noted that the more intimate format made it easier to engage in conversations and build connections across different roles and career stages. Networking was consistently identified as one of the most valuable aspects of the event. Participants appreciated both the formal scientific programme and the informal conversations that happened around talks, posters, and demonstrations. Looking ahead, many attendees expressed interest in even more dedicated time for networking and informal interaction.</p>
</section>
<section id="infrastructure-and-tools" class="level2">
<h2 class="anchored" data-anchor-id="infrastructure-and-tools">Infrastructure and tools</h2>
<section id="zulip" class="level3">
<h3 class="anchored" data-anchor-id="zulip">Zulip</h3>
<p>During BioC2026, Zulip provided a dedicated space for participants to connect, interact and stay engaged throughout the conference. The conference space included topic-based discussions covering: &#8211; <strong>Workshops</strong>: Dedicated topics for individual workshops, including Single Cell, Bioconductor in Workflows, microbiome, WebR, and other training sessions. &#8211; <strong>Training &#038; education</strong>: A space for broader conversations around training activities and educational sessions. &#8211; <strong>Social</strong>: Topics for informal conversations, social activities and connecting with other participants. &#8211; <strong>Introductions</strong>: Participants could introduce themselves and connect with others attending the conference. &#8211; <strong>Conference logistics channels</strong>: For travel, rooms, schedules, Day 1 information, and general conference support. &#8211; <strong>Community engagement</strong>: Dedicated spaces for the developer survey, student and early-career researcher activities, and other community initiatives. &#8211; <strong>Help and support</strong>: Dedicated topics such as the helpdesk, lost and found, and issues accessing virtual sessions.</p>
</section>
<section id="sticker-contest-winner" class="level3">
<h3 class="anchored" data-anchor-id="sticker-contest-winner">Sticker contest winner</h3>
<p><img src="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/BioC2026_sticker.png?w=578&#038;ssl=1" style="float:right; width:40%; margin-left:20px; margin-bottom:10px;" data-recalc-dims="1"></p>
<p>
Kayva Vaddadi was BioC2026’s sticker design contest winner. Kayva’s design celebrated Seattle as the host city of BioC2026, featuring well-known landmarks such as the Space Needle, the Seattle Great Wheel, Mount Rainier, and the surrounding evergreen forests. Puget Sound and a Washington State ferry reflect Seattle’s maritime landscape and the culture of connection across the region. The BioC2026 hot air balloon on the sticker was inspired by the scenic balloon views popular in the Seattle area and represents the celebratory and exploratory spirit of the Bioconductor community.
</p>
<p>
See the announcement here: <a href="https://lnkd.in/p/emAWMGX" rel="nofollow" target="_blank">https://lnkd.in/p/emAWMGX</a>
</p>
<div style="clear:both;">

</div>
</section>
</section>
<section id="conference-materials" class="level2">
<h2 class="anchored" data-anchor-id="conference-materials">Conference materials</h2>
<p>All conference recordings are available on the <a href="https://www.youtube.com/watch?v=OzpgIa2W414&#038;list=PLMYePXTcNLQo" rel="nofollow" target="_blank">Bioconductor YouTube channel</a>. BioC2026 workshops will also be available in the “Archived” section of <a href="http://workshop.bioconductor.org/" rel="nofollow" target="_blank">workshop.bioconductor.org</a>, alongside workshops from previous BioC, EuroBioC, and BiocAsia events. If you’d like to add additional materials or use the platform for your own workshop or course, contact us at <a href="http://support.bioconductor.org/" rel="nofollow" target="_blank">support.bioconductor.org</a>.</p>
</section>
<section id="post-conference-events" class="level2">
<h2 class="anchored" data-anchor-id="post-conference-events">Post-conference events</h2>
<section id="orchestrating-large-scale-single-cell-analysis-with-bioconductor" class="level3">
<h3 class="anchored" data-anchor-id="orchestrating-large-scale-single-cell-analysis-with-bioconductor">Orchestrating Large-Scale Single-Cell Analysis with Bioconductor</h3>
<p>On August 13–14, participants took part in a Carpentries workshop focused on orchestrating large-scale single-cell analysis with Bioconductor led by Ted Laderas and Wes Wilson. The workshop was attended by 11 participants and provided a hands-on environment for participants to work through approaches for analysing large-scale single-cell datasets using Bioconductor tools and workflows. Feedback highlighted the value of the live coding, small-group format and complementary teaching approaches.</p>
<div style="display:flex; gap:3px;">
<p><img src="https://i1.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/carpentry-workshop1.jpeg?w=578&#038;ssl=1" style="width:calc(50% - 1.5px); height:250px; object-fit:cover;" data-recalc-dims="1"></p>
<p><img src="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/scrna_bioc2026_workshop_1.jpeg?w=578&#038;ssl=1" style="width:calc(50% - 1.5px); height:250px; object-fit:cover;" data-recalc-dims="1"></p>
</div>
<center>
<p><em>Participants during the scRNA-seq workshop.</em></p>
</center>
</section>
<section id="bioconductor-north-america-2026-hackathon" class="level3">
<h3 class="anchored" data-anchor-id="bioconductor-north-america-2026-hackathon">Bioconductor North America 2026 Hackathon</h3>
<p>The Bioconductor North America 2026 Hackathon took place on August 13–14, bringing community members together to collaborate on software and infrastructure projects that strengthen interoperability, software development, and the wider Bioconductor ecosystem. These projects included: AI harness tools for standardizing how users within organizations would interact with agents on shared projects and on shared institutional resources. Link: <a href="https://github.com/eisenefaust/harness_agent" rel="nofollow" target="_blank">https://github.com/eisenefaust/harness_agent</a></p>
<p>Bioconductor tooling for including CLI executable components in Bioconductor packages. Link: <a href="https://blog.bioconductor.org/posts/2026-08-21-biocjobs/" rel="nofollow" target="_blank">https://blog.bioconductor.org/posts/2026-08-21-biocjobs/</a></p>
<p>An exploration on benchmarking LLM generated workflows, building out infrastructure and paradigms for how we can evaluate what LLMs return as work product to scientific questions. Link: <a href="https://github.com/Amisor/LLM_BioBenchmarking/tree/main" rel="nofollow" target="_blank">https://github.com/Amisor/LLM_BioBenchmarking/tree/main</a></p>
<img src="https://i1.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/bioc2026_hackathon.jpeg?w=578&#038;ssl=1" class="zoomable img-fluid" style="width:100.0%" data-recalc-dims="1">
<center>
<p><em>Participants during the Hackathon.</em></p>
</center>
</section>
</section>
<section id="up-next" class="level2">
<h2 class="anchored" data-anchor-id="up-next">Up next…</h2>
<p>Later this year, the Bioconductor community will head to Melbourne, Australia, for BioCAsia2026, taking place on November 19-20, 2026. <a href="https://biocasia2026.bioconductor.org/" rel="nofollow" target="_blank">BioCAsia2026</a> brings together researchers, developers, and data scientists from across the Asia-Pacific region and beyond to share advances in bioinformatics, computational biology, and open-source software. The meeting will highlight practical training, community building, and the latest developments in R and Bioconductor. BioC2027 will then take place in early summer in Boston, Massachusetts, continuing the North American Bioconductor conference series and bringing the community together once again for scientific exchange, collaboration, and community building. The community will thereafter gather in Basel, Switzerland, for EuroBioC2027, taking place from September 8–10, 2027. The European Bioconductor Conference will continue to bring together researchers, developers, and educators from across Europe and beyond to share new software, methods, and applications in computational biology.</p>
</section>
<section id="acknowledgements" class="level2">
<h2 class="anchored" data-anchor-id="acknowledgements">Acknowledgements</h2>
<p>BioC2026 gratefully acknowledges the support of all sponsors and partners whose contributions and support made the conference possible.</p>
<section id="sponsors" class="level3">
<h3 class="anchored" data-anchor-id="sponsors">Sponsors</h3>
<p>BioC2026 was proudly supported by Cortex, R consortium and Posit</p>
<p><a href="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/Bioc2026_sponsors.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-3" rel="nofollow" target="_blank"><img src="https://i0.wp.com/blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/Bioc2026_sponsors.png?w=578&#038;ssl=1" class="zoomable img-fluid" style="width:100.0%" data-recalc-dims="1"></a></p>
</section>
<section id="organising-committee" class="level3">
<h3 class="anchored" data-anchor-id="organising-committee">Organising committee</h3>
<p>We thank the local organisers, programme committee, workshop instructors, keynote speakers, volunteers, sponsors, and all participants whose contributions made BioC2026 a success.</p>
<ul>
<li>Marc Carlson, Seattle Children’s Hospital</li>
<li>Nicholas Cooley, University of Limerick</li>
<li>Sean Davis, University of Colorado Cancer Center</li>
<li>Maria Doyle, University of Limerick</li>
<li>Mikhail Dozmorov, Virginia Commonwealth University</li>
<li>Jenny Drnevich, University of Illinois at Urbana-Champaign</li>
<li>Russell Eberts, Fred Hutchinson Cancer Center</li>
<li>Erica Feick, Dana-Farber Cancer Institute</li>
<li>Sana Hirata, Fred Hutchinson Cancer Center</li>
<li>Ted Laderas, Fred Hutchinson Cancer Center</li>
<li>Alexandru Mahmoud, University of Limerick</li>
<li>Matthew McCall, University of Rochester Medical Center</li>
<li>Lori (Shepherd) Kern, Roswell Park Comprehensive Cancer Center</li>
<li>Laurah Nyasita Ondari, International Institute of Tropical Agriculture (IITA)</li>
<li>Charlotte Soneson, Friedrich Miescher Institute for Biomedical Research</li>
<li>Tim Triche, Van Andel Institute</li>
<li>Levi Waldron, CUNY Graduate School of Public Health and Health Policy</li>
<li>Wes Wilson, Princess Margaret Cancer Centre</li>
</ul>


</section>
</section>

<p>
© 2026 Bioconductor. Content is published under <a href="https://creativecommons.org/licenses/by/4.0/" rel="nofollow" target="_blank">Creative Commons CC-BY-4.0 License</a> for the text and <a href="https://opensource.org/licenses/BSD-3-Clause" rel="nofollow" target="_blank">BSD 3-Clause License</a> for any code. | <a href="https://www.r-bloggers.com/" rel="nofollow" target="_blank">R-Bloggers</a>
</p> 
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/"> Bioconductor community blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/bioc2026-conference-recap/">BioC2026 conference recap</a>]]></content:encoded>
					
		
		<enclosure url="https://blog.bioconductor.org/posts/2026-09-30-BioC2026-recap/media/bioc2026_groupshot.jpg" length="0" type="image/jpeg" />

		<post-id xmlns="com-wordpress:feed-additions:1">404015</post-id>	</item>
		<item>
		<title>My first attempt at porting a Shiny app failed, and not for a technical reason</title>
		<link>https://www.r-bloggers.com/2026/09/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/</link>
		
		<dc:creator><![CDATA[Arthur Bréant]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 16:14:32 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://rtask.thinkr.fr/?p=30097</guid>

					<description><![CDATA[<p>You can read the original post in its original format on Rtask website by ThinkR here: My first attempt at porting a Shiny app failed, and not for a technical reason<br />
Diary of a rewrite, episode 2 of 4. Translating your code line by line is the surest way to get it ...</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/">My first attempt at porting a Shiny app failed, and not for a technical reason</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://rtask.thinkr.fr/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/"> Rtask</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>You can read the original post in its original format on <a rel="nofollow" href="https://rtask.thinkr.fr/" target="_blank">Rtask</a> website by ThinkR here: <a rel="nofollow" href="https://rtask.thinkr.fr/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/" target="_blank">My first attempt at porting a Shiny app failed, and not for a technical reason</a></p>
<p><em>Diary of a rewrite, episode 2 of 4.</em></p>
<p>Translating your code line by line is the surest way to get it all wrong…</p>
<p>In <a href="https://rtask.thinkr.fr/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/" rel="nofollow" target="_blank">the previous episode</a> we told the story of why we cancelled our own stack decision halfway through, and why we went back to R. If you have not read it, no worries: this one stands on its own.</p>
<p>Today we are talking about the port itself. And above all about the way I got it wrong the first time.</p>
<p>A quick bit of context. StaffPuzzle is our internal staffing tool: it says who works on what, half-day by half-day. Its Shiny version is 1,200 lines, reads the team’s Google calendars, and has been running for years. The goal was to rewrite it in <a href="https://github.com/hyperverse-r/htmxr" rel="nofollow" target="_blank">htmxr</a>, meaning HTML rendered server-side, hydrated by htmx, with plumber2 behind it.</p>
<p>Here is how I went about it.</p>
<p>I opened <code>app_server.R</code>, I listed every <code>observeEvent</code>, and I started turning them into routes. One route per observer. It is methodical, it is exhaustive… and it produces an application full of endpoints that serve no purpose.</p>
<p>A quick reminder: what is a route? In a classic web application, the browser asks for an address and the server sends back a response. A route is that pair: an address, and the function that builds what we send back. Where Shiny holds a continuous conversation with the browser for the whole visit, a route-based application answers requests that are independent of one another, with no memory of the previous one.</p>
<p>Concretely, here is the kind of translation I was doing. On the Shiny side, the observer that rebuilds the main plot whenever a filter moves (I have shortened and renamed it for readability, the original is bushier):</p>
<pre>observeEvent(c(input$refresh, rv$slots, rv$data_ready), {
  req(rv$slots, rv$consultants, rv$start, rv$end, rv$sort_by)
  rv$filtered_slots &lt;- rv$slots |&gt;
    filter(member %in% rv$consultants) |&gt;
    mark_focus(types = rv$slot_types) |&gt;
    add_weekdays()
  rv$plot &lt;- rv$filtered_slots |&gt;
    team_availability_plot(
      start   = rv$start,
      end     = rv$end,
      sort_by = rv$sort_by,
      order   = rv$consultants
    )
})</pre>
<p>And, a hundred lines further down in the same file, the output that displays it:</p>
<pre>output$plot &lt;- renderPlotly({
  rv$plot |&gt; ggplotly(tooltip = &quot;text&quot;)
})</pre>
<p>On the routes side, it becomes an address and a function:</p>
<pre>#* @get /plot
function(consultants, start, end, sort_by, slot_types) {
  slots |&gt;
    filter(member %in% consultants) |&gt;
    mark_focus(types = slot_types) |&gt;
    add_weekdays() |&gt;
    team_availability_plot(
      start   = start,
      end     = end,
      sort_by = sort_by,
      order   = consultants
    )
}</pre>
<p>The browser will ask for <code>/plot?consultants=arthur,vincent&start=2026-10-05&end=2026-11-30&sort_by=alphabetical</code>, and the server will answer with the corresponding piece of page.</p>
<p>Notice in passing what has melted away on the Shiny side: I need an observer, a slot in <code>reactiveValues</code> to store the result, and a <code>render*</code> elsewhere in the file to display it. On the route side, all that is left is a function that takes arguments and returns something. That one translates well: it produces something visible, and that something can state its address.</p>
<p>The problem is all the others.</p>
<p>My mistake did not come from not knowing the destination tool. I was translating code instead of changing paradigm. And it took me several days to notice!</p>
<p>It is a bit like moving house. You do not pack by photographing each room so you can rebuild it identically in the new flat. You look at what you actually use, and you leave the rest at the car boot sale.</p>
<p>So here is what I wish someone had told me before I started.</p>
<h2>Why inventorying the observers is the wrong entry point</h2>
<p>The reason is structural. Shiny is driven by a reactive graph, and a good share of the observers have no visible output at all. They copy a value from one place to another, they disable a button, they keep two inputs in sync. In short, they maintain the consistency of an in-memory state.</p>
<p>But that state does not exist in htmxr. Neither does the work of those observers. They do not translate: they disappear.</p>
<p>I remember the sentence I blurted out in a meeting, on 30 June, right in the middle of the port:</p>
<blockquote><p>
I feel like saying “I’ve got this or that observeEvent, so I’ve got this or that endpoint” isn’t ok. Whereas I thought starting from the data to serve is a good entry point.
</p></blockquote>
<p>It has been my working rule ever since.</p>
<h2>Start from the data being served</h2>
<p>An htmxr application is a set of addressable blocks of data. Every block on screen has a URL that renders it. The controls that drive it (filters, dates, sorting) are the parameters of that URL. We are in a hypermedia frame, no longer in a dependency graph.</p>
<p>The method comes down to 5 moves, and the order matters:</p>
<ol style="list-style-type: decimal">
<li><strong>Inventory the screens, screenshot by screenshot.</strong> Not what the code does: what the user sees.</li>
<li><strong>Cut each screen into blocks of data</strong>: header, filters, grid, detail panel.</li>
<li><strong>Give every block a URL.</strong> A block that cannot state its address is a badly cut block.</li>
<li><strong>Turn the controls into parameters.</strong> <code>?start=2026-10-01&n_weeks=8</code>, not observers.</li>
<li><strong>For each interaction, ask yourself a single question</strong>: is this a mutation of persistent business state?</li>
</ol>
<p>The fifth one is the one I kept missing. Clicking a filter, changing a sort, picking a period: those are <code>GET</code>s, even if “something changes” on screen. Nothing is modified server-side, we are simply asking for another view of the same data.</p>
<p>Tip: if your inventory produces a lot of <code>POST</code>s, chances are some <code>GET</code>s have slipped in. Go back over them.</p>
<h2>60 lines with nothing left to do</h2>
<p>Here is the example that tipped my understanding.</p>
<p>My Shiny version devoted 60 lines to keeping the URL and the inputs in sync, in both directions. Good code, by the way: that application had shareable URLs, which is not that common in Shiny.</p>
<p>On the Shiny side, 60 lines. Here are the two ends of it, abridged and renamed:</p>
<pre>## Direction 1: the URL arrives, we put the inputs in the right state
observeEvent(session$clientData$url_search, {
  params &lt;- parseQueryString(session$clientData$url_search)
  url_params$consultants &lt;- params$consultants
  url_params$slot_types  &lt;- params$types
  url_params$services    &lt;- params$services
  url_params$period      &lt;- params$period
  if (!is.null(url_params$consultants)) {
    updateSelectInput(
      session  = session,
      inputId  = &quot;consultants&quot;,
      selected = str_to_title(
        str_split_1(str_remove_all(url_params$consultants, &quot; &quot;), pattern = &quot;,&quot;)
      )
    )
  }
  ## … and the same treatment for period, slot_types and services
})
## Direction 2: an input moves, we rewrite the URL
observeEvent(
  c(input$consultants, input$slot_types, input$services,
    input$period, input$sort_by),
  ignoreInit = TRUE, {
    consultants &lt;- str_to_lower(glue_collapse(input$consultants, sep = &quot;,&quot;))
    slot_types  &lt;- str_to_lower(glue_collapse(input$slot_types, sep = &quot;,&quot;))
    period      &lt;- glue_collapse(input$period, sep = &quot;,&quot;)
    url &lt;- glue(&quot;?consultants={consultants}&#038;types={slot_types}&#038;period={period}&quot;)
    updateQueryString(session = session, url, mode = &quot;replace&quot;)
  })</pre>
<p>On the htmxr side, one argument:</p>
<pre>hx_set(
  form,
  get      = &quot;.&quot;,
  target   = &quot;#trainings-view&quot;,
  trigger  = &quot;change&quot;,
  push_url = &quot;true&quot;
)</pre>
<p>Careful, this is not a matter of concision. Those 60 lines have not been shortened: they have nothing left to do. They were maintaining a mirror between two representations of the same state, the URL on one side, the inputs on the other. When the URL <em>is</em> the state, there is simply no mirror left to hold.</p>
<p>Back to our house move: these are not furniture you carry up to the new flat. There is no room for them.</p>
<p>Want a simple test to see where you stand? If your port produces code that synchronises two things, you have kept the original paradigm. A good port deletes whole categories of code, it does not rewrite them shorter.</p>
<h2>What happens to each building block</h2>
<p>Once the frame is set, the port becomes mechanical. Let us look at what becomes of the blocks you have in front of you in your <code>server.R</code>.</p>
<p><strong>A module becomes a route and a rendering function.</strong> My biggest module was 267 lines: 4 <code>selectInput</code>, one <code>dateRangeInput</code>, 6 observers, and the URL synchronisation. All that is left is an HTML form and some parameters. The <code>ns()</code> disappears entirely, since it only existed to avoid identifier collisions between instances. With no input binding, there is nothing left to name.</p>
<p><strong><code>input$x</code> becomes an argument of your function.</strong> There is no <code>input</code> object any more. The value arrives in the request: as a URL parameter for a <code>GET</code>, in the body for a <code>POST</code>. It is the most disorienting change in the first few days, and the most liberating one afterwards. A view function receives its inputs as arguments, so it can be tested!</p>
<p><strong><code>output$x &lt;- render*()</code> becomes a function that returns HTML.</strong> The output/render pair disappears. All that is left is a function and its return value, called by a route. A <code>renderPlot</code> becomes a URL that serves the SVG, a <code>renderUI</code> a URL that serves a fragment.</p>
<p><strong><code>updateSelectInput()</code> becomes: nothing.</strong> We no longer update a control. The server renders the <code>&lt;select&gt;</code> in the state it should be in, and htmx swaps the fragment that contains it. The whole <code>update*Input()</code> family evaporates with it.</p>
<p><strong><code>invalidateLater()</code> becomes an attribute.</strong> A periodic refresh is written <code>trigger = &quot;every 15s&quot;</code>. The browser asks for the fragment again, the server recomputes it. Nothing to invalidate, since there is no state.</p>
<p><strong><code>showNotification()</code> and <code>withProgress()</code> become a fragment, or CSS.</strong> A message is a piece of HTML returned with the response. A waiting indicator is an element that htmx shows during the request and hides afterwards. A CSS class, not R code.</p>
<p><strong><code>reactiveValues()</code> and <code>session$userData</code> become the URL, or the database.</strong> In the URL if it describes what we are looking at, in the database if it has to survive closing the tab. And if nothing is left after that sorting? Then it was plumbing state, and it goes with the rest <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f609.png" alt="😉" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>
<h3>Seeing it run</h3>
<p>All of this stays abstract until you have it in front of you. So let us take Shiny’s canonical example, the one you get with <code>shiny::runExample(&quot;01_hello&quot;)</code>: the Old Faithful geyser histogram, with a slider for the number of bins.</p>
<ul>
<li><a href="https://hello.demo.hyperverse.world/" rel="nofollow" target="_blank">The application, written in htmxr</a>. One slider, one plot, and that is all.</li>
<li><a href="https://hello.demo.hyperverse.world/plot?bins=30" rel="nofollow" target="_blank">The URL that serves the plot</a>. Open it directly in your browser: you land on the SVG itself, the drawing spelled out. That is exactly what “give every block a URL” means.</li>
</ul>
<p>Have fun changing <code>bins=30</code> to <code>bins=5</code>, then to <code>bins=50</code>. The server recomputes and sends you back the drawing. No session, no state, no open conversation.</p>
<h3>And the code, on both sides</h3>
<p>You already know the Shiny version: it is the one from the tutorial. Here it is, stripped to the essentials.</p>
<pre>ui &lt;- page_sidebar(
  title = &quot;Hello Shiny!&quot;,
  sidebar = sidebar(
    sliderInput(&quot;bins&quot;, &quot;Number of bins:&quot;, min = 1, max = 50, value = 30)
  ),
  plotOutput(&quot;distPlot&quot;)
)
server &lt;- function(input, output) {
  output$distPlot &lt;- renderPlot({
    x    &lt;- faithful$waiting
    bins &lt;- seq(min(x), max(x), length.out = input$bins + 1)
    hist(x, breaks = bins, col = &quot;darkgray&quot;, border = &quot;white&quot;)
  })
}</pre>
<p>And the htmxr version, the one running behind the link above. I have removed the Bootstrap classes, they teach you nothing.</p>
<pre>#* @get /
#* @serializer html
function() {
  hx_page(
    hx_head(title = &quot;Old Faithful Geyser Data&quot;),
    hx_slider_input(
      id      = &quot;bins&quot;,
      label   = &quot;Number of bins:&quot;,
      value   = 30, min = 1, max = 50,
      get     = &quot;plot&quot;,
      trigger = &quot;input changed delay:300ms&quot;,
      target  = &quot;#plot&quot;
    ),
    tags$div(id = &quot;plot&quot;) |&gt;
      hx_set(get = &quot;plot&quot;, trigger = &quot;load&quot;, target = &quot;#plot&quot;)
  )
}
#* @get /plot
#* @serializer none
function(query) {
  xmlSVG({
    x    &lt;- faithful$waiting
    bins &lt;- seq(min(x), max(x), length.out = as.numeric(query$bins %||% 30) + 1)
    hist(x, breaks = bins, col = &quot;darkgray&quot;, border = &quot;white&quot;)
  }, width = 7, height = 5)
}</pre>
<p>Three things to look at, and they sum up the whole article.</p>
<p><strong>The slider is the same object.</strong> <code>sliderInput(id, label, min, max, value)</code> becomes <code>hx_slider_input(id, label, min, max, value)</code>. There is nothing to relearn.</p>
<p><strong>The plot computation has not moved by a single line.</strong> The four lines of <code>hist()</code> are identical on both sides. That is what I had not understood when I started: it is not your business code you are porting, it is the plumbing around it.</p>
<p><strong>What changes is the wiring, and it becomes visible.</strong> In Shiny, nothing says that moving the slider redraws the plot: the reactive graph works it out by itself from the presence of <code>input$bins</code> inside <code>renderPlot</code>. In htmxr, it is written on the slider itself: <code>get = &quot;plot&quot;</code> (go fetch this address), <code>target = &quot;#plot&quot;</code> (put the result there). You read the wiring instead of guessing it.</p>
<p>And <code>server &lt;- function(input, output)</code> is gone. All that is left is two functions that take arguments and return something.</p>
<h3>Try it yourself</h3>
<p>No need to copy anything: the example ships with the package.</p>
<pre># install.packages(&quot;pak&quot;)
pak::pak(&quot;hyperverse-r/htmxr&quot;)
htmxr::hx_run_example(&quot;hello&quot;)</pre>
<p>Your browser is waiting at <code>http://127.0.0.1:8080</code>. Move the slider: every move fires a request to <code>/plot</code>, and the server returns a new SVG. Open your browser’s network tab to watch them go by, it beats any explanation.</p>
<p>And if you call <code>hx_run_example()</code> with no argument, it lists the others: <code>select-input</code>, <code>multi-select</code>, <code>delete-row</code>, <code>toast-notification</code>, <code>infinity-scroll</code>, <code>page-or-fragment</code>, <code>json-endpoint</code> and <code>reactive-values</code>. That last one is worth a look if you are wondering what becomes of a <code>reactiveValues()</code> once the reactive graph is gone.</p>
<p>And since the plot is a URL, anyone can consume it. <a href="https://helloreact-mu.vercel.app/" rel="nofollow" target="_blank">Here is the same application written in React</a>, calling exactly the same R endpoint. The backend does not know who is asking, and it does not care.</p>
<p>This is the most underrated property of cutting things into addressable blocks: it does not lock you in. The day you want another frontend, or a mobile app, or simply for a colleague to pull your plot into their own tool, your R logic is already an API. It became one without you doing anything extra.</p>
<h2>What costs, and what does not</h2>
<p>Everything above is mechanical. On my port, only 3 patterns called for a real decision, and they are the ones that make an estimate vary.</p>
<table>
<colgroup>
<col width="33%" />
<col width="33%" />
<col width="33%" />
</colgroup>
<thead>
<tr class="header">
<th>Shiny pattern</th>
<th>Cost</th>
<th>Equivalent</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><code>selectInput</code>, <code>actionButton</code>, <code>sliderInput</code></td>
<td>low</td>
<td><code>hx_select_input()</code>, <code>hx_button()</code>, <code>hx_slider_input()</code></td>
</tr>
<tr class="even">
<td><code>dateRangeInput</code></td>
<td>low</td>
<td>no dedicated constructor: two <code>&lt;input type=&quot;date&quot;&gt;</code></td>
</tr>
<tr class="odd">
<td>URL <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/2194.png" alt="↔" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎ inputs sync</td>
<td>low</td>
<td>the <code>push_url</code> argument</td>
</tr>
<tr class="even">
<td>static <code>ggplot</code></td>
<td>low</td>
<td><code>svglite</code> + htmx swap</td>
</tr>
<tr class="odd">
<td><code>DT</code> (display)</td>
<td>low</td>
<td><code>hx_table()</code>, <code>hx_table_rows()</code></td>
</tr>
<tr class="even">
<td>long task + notifications</td>
<td>medium</td>
<td>external job + periodic polling</td>
</tr>
<tr class="odd">
<td>interactive <code>plotly</code></td>
<td><strong>high</strong></td>
<td>plotly.js, or drop the interactivity</td>
</tr>
<tr class="even">
<td><code>rhandsontable</code></td>
<td><strong>high</strong></td>
<td>JS library, or one form per cell</td>
</tr>
</tbody>
</table>
<p><strong>The editable table has no native equivalent.</strong> Either you embed a JavaScript library, or you rebuild it with HTML fields and one save per cell. The criterion is dimensional: under twenty rows or so HTML wins, at a thousand rows with multi-cell copy-paste the library becomes the right choice again.</p>
<p><strong>The interactive plot deserves a prior question</strong>, and it is precisely the one I forgot to ask myself: is it really a plot? When it is in fact a data grid, a schedule, a calendar or a heatmap, porting it as a plot is a mistake. Mine was one because Shiny made tables hard, not because the table was the right object. A real plot, on the other hand, stays an SVG served by a URL, and that is an easy case.</p>
<p><strong>An opaque internal dependency is not worked around, it is isolated.</strong> You definitely have one: the homegrown package that fetches the data from somewhere, that nobody has really read in years. The temptation is to start with it, to “understand before porting”. Do not. You will spend three days on it and you will have ported nothing.</p>
<p>Develop everything else against a demo dataset, hard-coded, and only plug the real package in at the very end, in one go. By then you will know exactly what you want from it: a function, some arguments, a table back. And if it is still incomprehensible, call it as it is without trying to rewrite it. That is not cowardice, it is what stops it from blocking the whole port.</p>
<h2>The most useful side effect</h2>
<p>Shiny encourages mixing business logic and reactive orchestration: they live in the same file, often in the same function. In my v1, the rule that decides whether a slot is highlighted was computed twice, with different rules, 15 lines apart. That was not carelessness on my part, it is what a file where business and plumbing cohabit produces.</p>
<p>The port forces the separation. A function called by a route can no longer lean on an ambient reactive context. What is left is ordinary R, which takes data and returns data or HTML. So it can be tested.</p>
<p>The tests-to-code ratio went from 0.26 to 0.77: 544 tests, no browser, a few milliseconds. With the interface having become a pure function from data to HTML, it falls into the same test harness as the rest. The Shiny equivalent needs <code>shinytest2</code>, a headless Chrome and snapshots. It is slow, it is fragile, and it breaks when a margin moves.</p>
<p>Two pieces of honesty are in order here. My new version does more than the old one (writing to calendars, draft mode, historised indicators), and part of the difference in volume comes from that, not from the paradigm. And nothing stops you from testing the business logic of a Shiny application: v1 did it over 300 lines. What changes is that the interface stops being the area you never test.</p>
<h2>What you do not get back</h2>
<p>Let us be clear: the reactive graph does render a real service when everything depends on everything. Think of a dashboard where 6 linked plots recompute together, with cross dependencies and expensive derived values. Rebuilding that with fragments and URL parameters is possible, but you are hand-writing what Shiny gives you for free.</p>
<p>Hence the only decision criterion that is worth anything, to my mind: does my application’s state fit in a URL?</p>
<p>For mine, yes: a period, some filters, a sort. The port was a good bet. For an exploration tool where the user builds a complex working session, the answer is no, and Shiny remains the right tool. That is not a polite concession, it is the same reasoning applied to a different problem.</p>
<h2>Conclusion</h2>
<p>I still write Shiny, for our clients and for ourselves. This port did not prove that one technology beats another.</p>
<p>What it did show me is that for years I had been solving a page-rendering problem with a tool designed to hold state in memory. And that the cost of that gap stayed invisible until I had paid it once.</p>
<p>So before you set off, ask yourself the URL question. It will save you the few days I lost!</p>
<p><strong>Do you have a Shiny application in this situation?</strong> We read it and give you back a note that says whether the port is worth it, including when the answer is no. <a href="https://rtask.thinkr.fr/contact/" rel="nofollow" target="_blank">Write to us</a> in two lines: the size of your application, and how long it has been running.</p>
<p>Next week, episode 3: what leaving the managed platform cost us, bill in hand.</p>
<hr />
<p><strong>Diary of a rewrite</strong>, four episodes:</p>
<ol style="list-style-type: decimal">
<li><strong>We chose React to learn it. The AI wrote it for us.</strong></li>
<li><strong>My first attempt at porting a Shiny app failed, and not for a technical reason</strong> <em>(you are here)</em></li>
<li><strong>We left the managed platform. Here is the real bill.</strong> <em>(6 October)</em></li>
<li><strong>Five signs your Shiny app has outgrown Shiny</strong> <em>(13 October)</em></li>
</ol>
<p><em>The application cited is StaffPuzzle, ThinkR’s internal staffing tool, written with htmxr, alpiner and plumber2. The figures are measured on the repository’s two branches, not estimated. Disclosure of interest: I am the author of htmxr and alpiner.</em></p>
<p>This post is better presented on its original ThinkR website here: <a rel="nofollow" href="https://rtask.thinkr.fr/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/" target="_blank">My first attempt at porting a Shiny app failed, and not for a technical reason</a></p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://rtask.thinkr.fr/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/"> Rtask</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/my-first-attempt-at-porting-a-shiny-app-failed-and-not-for-a-technical-reason/">My first attempt at porting a Shiny app failed, and not for a technical reason</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403991</post-id>	</item>
		<item>
		<title>August 2026 Top 40 New CRAN Packages</title>
		<link>https://www.r-bloggers.com/2026/09/august-2026-top-40-new-cran-packages/</link>
		
		<dc:creator><![CDATA[Joseph Rickert]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://rworks.dev/posts/august-2026-top-40-new-cran-packages/</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; ">
<p>Three hundred twenty-nine new packages were submitted to CRAN in July. Here are my Top 40 picks in nineteen categories: Agriculture, Artificial Intelligence, Climate Studies, Computational Methods, Ecology, Epidemiology, Genomics, Machine Learni...</p></div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/august-2026-top-40-new-cran-packages/">August 2026 Top 40 New CRAN Packages</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://rworks.dev/posts/august-2026-top-40-new-cran-packages/"> R Works</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
 





<p>Three hundred twenty-nine new packages were submitted to CRAN in July. Here are my Top 40 picks in nineteen categories: Agriculture, Artificial Intelligence, Climate Studies, Computational Methods, Ecology, Epidemiology, Genomics, Machine Learning, Medical Statistics, Meta-Analysis, Networks, Process control, Programming, Public Transit, Statistics, Surveys, Time Series, Utilities, and Visualization.</p>
<div class="columns">
<div class="column" style="width:45%;">
<section id="agriculture" class="level3">
<h3 class="anchored" data-anchor-id="agriculture">Agriculture</h3>
<p><a href="https://cran.r-project.org/package=agridatasets" rel="nofollow" target="_blank">agridatasets</a> v0.1.1: Offers a rich and diverse collection of datasets focused on agriculture, agronomy, animal science, and related fields. The package includes experimental, observational, and field-trial data on crops such as rice, wheat, corn, soybean, cotton, coffee, avocado, and orange, as well as forestry species including bamboo, eucalyptus, and timber. Datasets cover plant breeding and genetics, factorial and randomized block experiments, herbicide and insecticide efficacy trials, pest and disease infestation, soil characteristics and land suitability, plant growth regulators, seed germination, and crop yield modeling. See the <a href="https://cran.r-project.org/web/packages/agridatasets/vignettes/agridatasets.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/agridatasets.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-1" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/agridatasets.png?w=578&#038;ssl=1" class="img-fluid" alt="Rice vs Wheat Time Series" data-recalc-dims="1"></a></p>
</section>
<section id="artificial-intelligence" class="level2">
<h2 class="anchored" data-anchor-id="artificial-intelligence">Artificial Intelligence</h2>
<p><a href="https://cran.r-project.org/package=commons" rel="nofollow" target="_blank">commons</a> v0.1.0: Implements trustworthy large language model agents. Connect raw data sources, provides a pool of trusted calculations, and a searchable context layer that demonstrates how to interpret them. Then, deploy data agents that answer questions, log interactions, and can be evaluated and improved over time. See the vignettes <a href="https://cran.r-project.org/web/packages/commons/vignettes/commons.html" rel="nofollow" target="_blank">Introduction</a> and <a href="https://cran.r-project.org/web/packages/commons/vignettes/governance.html" rel="nofollow" target="_blank">Security and governance</a>.</p>
<p><a href="https://cran.r-project.org/package=diffuseR" rel="nofollow" target="_blank">diffuseR</a> v0.2.2: A native <code>R</code> implementation of diffusion models providing a functional interface to state-of-the-art generative AI. Inspired by the <code>Python</code> library <code>diffusers</code> from <a href="https://huggingface.co/" rel="nofollow" target="_blank">Hugging Face</a>, functions generate and manipulate images from text prompts using models such as <code>Stable Diffusion</code>, with no <code>Python</code> dependency. Supports multiple diffusion schedulers and device acceleration. See the <a href="https://cran.r-project.org/web/packages/diffuseR/vignettes/performance-levers.html" rel="nofollow" target="_blank">vignette</a> and <a href="https://cran.r-project.org/web/packages/diffuseR/readme/README.html" rel="nofollow" target="_blank">README</a> for more information.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/diffuseR.jpg?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-2" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/diffuseR.jpg?w=578&#038;ssl=1" class="img-fluid" alt="Sample Output" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=Rhobots" rel="nofollow" target="_blank">Rhobots</a> v0.1.10: Implements the <code>BERTopic</code> topic modeling pipeline directly in <code>R</code>: Provides transformer-based sentence embedding, uniform manifold approximation and projection dimensionality reduction, hierarchical density-based spatial clustering of applications with noise clustering, and class-based term frequency-inverse document frequency topic extraction without any dependency on <code>Python</code>, <code>conda</code>, or <code>reticulate</code>. Every stage runs in <code>R</code> through <code>torch</code>, <code>safetensors</code>, <code>tok</code>, <code>uwot</code>, and <code>dbscan</code>. The package mirrors the accessor API of the original <code>Python</code> package, adds integrated quality metrics and hyperparameter search tools, and introduces part-of-speech filtered and C-value-ranked representation models. See <a href="https://cran.r-project.org/web/packages/Rhobots/readme/README.html" rel="nofollow" target="_blank">README</a> for more information.</p>
<section id="climate-studies" class="level3">
<h3 class="anchored" data-anchor-id="climate-studies">Climate Studies</h3>
<p><a href="https://cran.r-project.org/package=topocast" rel="nofollow" target="_blank">topocast</a> v0.0.5: Downscales coarse-resolution raster data to a finer grid by fitting local linear regressions of a response, such as a climate variable, on one or more fine-resolution predictors, such as elevation and other terrain indices, within a moving window. Multiplicative and additive anomaly application downscale time series relative to a baseline climatology. Follows the regression-on-elevation approach used for high-resolution climate surfaces described in <a href="doi:10.1038/sdata.2017.122" rel="nofollow" target="_blank">Karger et al. (2017)</a>. See the vignettes <a href="https://cran.r-project.org/web/packages/topocast/vignettes/getting-started.html" rel="nofollow" target="_blank">Getting started</a> and <a href="https://cran.r-project.org/web/packages/topocast/vignettes/how-it-works.html" rel="nofollow" target="_blank">How moving-window downscaling works</a>.</p>
<p><a href="https://rworks.dev/posts/august-2026-top-40-new-cran-packages/topocast.svg" class="lightbox" data-gallery="quarto-lightbox-gallery-3" rel="nofollow" target="_blank"><img src="https://rworks.dev/posts/august-2026-top-40-new-cran-packages/topocast.svg" class="img-fluid" alt="Heat map showing the result of downscaling"></a></p>
</section>
<section id="computational-methods" class="level3">
<h3 class="anchored" data-anchor-id="computational-methods">Computational Methods</h3>
<p><a href="https://cran.r-project.org/package=openfhe.R" rel="nofollow" target="_blank">openfhe.R</a> v1.5.1.1: Provides an <code>R</code> interface to <code>penFHE</code>, the open-source <code>C++</code> library for fully homomorphic encryption <a href="https://eprint.iacr.org/2022/915" rel="nofollow" target="_blank">Badawi et al. (2022)</a>, which allows computation directly on encrypted data without access to the secret key. Supports the <a href="https://eprint.iacr.org/2012/144" rel="nofollow" target="_blank">Brakerski-Fan-Vercauteren (2012)</a>, <a href="doi:10.1145/2633600" rel="nofollow" target="_blank">Brakerski-Gentry-Vaikuntanathan (2014)</a>, and <a href="https://eprint.iacr.org/2016/421" rel="nofollow" target="_blank">Cheon-Kim-Kim-Song (2017)</a> schemes for arithmetic on encrypted numbers. There are three vignettes including <a href="https://cran.r-project.org/web/packages/openfhe.R/vignettes/introduction.html" rel="nofollow" target="_blank">Introduction</a> and <a href="https://cran.r-project.org/web/packages/openfhe.R/vignettes/ckks-bootstrapping.html" rel="nofollow" target="_blank">CKKS Bootstrapping</a>.</p>
</section>
<section id="ecology" class="level3">
<h3 class="anchored" data-anchor-id="ecology">Ecology</h3>
<p><a href="https://cran.r-project.org/package=FINN" rel="nofollow" target="_blank">FINN</a> v0.1.0: Implements a hybrid dynamic forest model that can be configured as a fully mechanistic, process-based model, like classic forest gap models, or with its demographic processes (growth, mortality, regeneration) replaced by deep neural networks, or any combination of the two. Provides functions to define a model and its mechanistic or empirical components, calibrate it to forest inventory data, and interpret the calibrated processes. FINN is implemented with the <code>torch</code> package but no knowledge of <code>torch</code> is required. The hybrid modeling approach is described in <a href="doi:10.1111/2041-210x.70347" rel="nofollow" target="_blank">Pichler and Käber (2026)</a>. There are five vignettes including <a href="https://cran.r-project.org/web/packages/FINN/vignettes/A-Introduction_to_FINN.html" rel="nofollow" target="_blank">Introduction</a> and <a href="https://cran.r-project.org/web/packages/FINN/vignettes/E-Mortality.html" rel="nofollow" target="_blank">Mortality: a binomial response and a neural-network process</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/FINN.jpeg?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-4" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/FINN.jpeg?w=578&#038;ssl=1" class="img-fluid" alt="Illustration of how FINN works" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=gbif.range" rel="nofollow" target="_blank">gbif.range</a> v1.9.2: Implements an end-to-end workflow to generate ecologically informed species range maps from sparse observations using environmental clustering and convex hulls. Serves as a standalone framework or complementary approach to species distribution models. By constraining estimated ranges within authoritative or custom ecoregion boundaries, the approach prevents spurious range over-prediction common in geometric hull methods. There are five vignettes including <a href="https://cran.r-project.org/web/packages/gbif.range/vignettes/getting-started.html" rel="nofollow" target="_blank">Getting Started</a> and <a href="https://cran.r-project.org/web/packages/gbif.range/vignettes/ecoregion-constrained-range-inference.html" rel="nofollow" target="_blank">Part 2: Ecoregion-Based Range Inference</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/gbif.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-5" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/gbif.png?w=578&#038;ssl=1" class="img-fluid" alt="Map showing 10 species classes over the European Alps, using two CHELSA bioclimatic layers" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=PhysMove" rel="nofollow" target="_blank">PhysMove</a> v1.2.5: Provides tools to analyse animal movement and space-use patterns from telemetry data using methods derived from statistical physics. Methods span displacement-based approaches, distribution fitting, space-use metrics, including the influence of correlations on space-use, network-based community detection, and measures of entropy and predictability. The package enables characterization of these patterns across spatial and temporal scales. For applications of these methods in ecological studies see <a href="doi:10.1038/s41598-017-00165-0" rel="nofollow" target="_blank">Rodríguez et al. (2017)</a> and <a href="doi:10.1073/pnas.1716137115" rel="nofollow" target="_blank">Sequeira et al. (2018)</a>. There are four vignettes including <a href="https://cran.r-project.org/web/packages/PhysMove/vignettes/pt1_introduction.html" rel="nofollow" target="_blank">Introduction</a> and <a href="https://cran.r-project.org/web/packages/PhysMove/vignettes/pt2_movement_patterns.html" rel="nofollow" target="_blank">Movement Patterns</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/PhysMove.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-6" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/PhysMove.png?w=578&#038;ssl=1" class="img-fluid" alt="Map showing simulated movement patterns" data-recalc-dims="1"></a></p>
</section>
<section id="epidemiology" class="level3">
<h3 class="anchored" data-anchor-id="epidemiology">Epidemiology</h3>
<p><a href="https://cran.r-project.org/package=RtForecastR" rel="nofollow" target="_blank">RtForecastR</a> v0.1.1: Provides functions for filtered (real-time/causal) and smoothed (retrospective) estimation of the time-varying effective reproduction number from case-count time series, using the EpiFilter algorithm of <a href="doi:10.1371/journal.pcbi.1009347" rel="nofollow" target="_blank">Parag (2021)</a>, together with a one-step-ahead in-sample prediction check, a genuine out-of-sample one-step forecast with predictive intervals, elimination probability <img src="https://latex.codecogs.com/png.latex?P(R_t%20%3C%201)">, and forecast calibration metrics. See the <a href="https://cran.r-project.org/web/packages/RtForecastR/vignettes/rtforecastr-walkthrough.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/RtForecastR.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-7" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/RtForecastR.png?w=578&#038;ssl=1" class="img-fluid" alt="Filtered realtime vs smoothed R_t" data-recalc-dims="1"></a></p>
</section>
<section id="genomics" class="level3">
<h3 class="anchored" data-anchor-id="genomics">Genomics</h3>
<p><a href="https://cran.r-project.org/package=diffwrap" rel="nofollow" target="_blank">diffwrap</a> v0.6-3: Provides functions for differential expression analysis of read counts from messenger RNA (mRNA) sequencing (RNA-Seq) data or micro RNA (miRNA) expression values generated by the Comprehensive Analysis Pipeline for microRNA Sequencing <code>expression_reports.sh</code> script. The workflow follows the Bioconductor <code>edgeR</code>&#8211;<code>limma</code> expression data analysis pipeline providing options for different approaches, such as pure <code>edgeR</code>, <code>voom</code> or paired samples. Methods are described in <a href="doi:10.1093/bioinformatics/btp616" rel="nofollow" target="_blank">Robinson, McCarthy and Smyth (2010)</a>, <a href="doi:10.1093/nar/gkv007" rel="nofollow" target="_blank">Ritchie et al. (2015)</a>, <a href="doi:10.1186/gb-2014-15-2-r29" rel="nofollow" target="_blank">Law et al. (2014</a> and <a href="doi:10.1186/1471-2164-15-423" rel="nofollow" target="_blank">Sun et al. (2014)</a>. See the <a href="https://rworks.dev/posts/august-2026-top-40-new-cran-packages/" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/diffwrap.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-8" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/diffwrap.png?w=578&#038;ssl=1" class="img-fluid" alt="Heat map of genes over FDR-filtered samples" data-recalc-dims="1"></a></p>
</section>
<section id="machine-learning" class="level3">
<h3 class="anchored" data-anchor-id="machine-learning">Machine Learning</h3>
<p><a href="https://cran.r-project.org/package=figsr" rel="nofollow" target="_blank">figsr</a> v0.1.1: Implements a flexible, interpretable machine learning algorithm for additive tree sums. Fits a sum of shallow classification and regression trees (CART) by greedily minimizing residual impurity, growing a new tree or deepening an existing one at each step, whichever reduces the residuals most. Supports regression and two-class classification, variable importance, bootstrap ensembling and seamless integration with <code>parsnip</code> and <code>tidymodels</code> workflows. The method is described in <a href="doi:10.1073/pnas.2310151122" rel="nofollow" target="_blank">Tan et al. (2023)</a>. See the <a href="https://cran.r-project.org/web/packages/figsr/vignettes/figsr-intro.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/figsr.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-9" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/figsr.png?w=578&#038;ssl=1" class="img-fluid" alt="Visualization of the tree sum" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=neuralsbi" rel="nofollow" target="_blank">neuralsbi</a> v0.3.2: Provides a native <code>R</code> implementation of a neural simulation-based estimator that runs on the <code>torch</code> back end and is focused on Neural Posterior Estimation. Given a prior over parameters and a simulator, functions train a conditional neural density estimator to approximate the Bayesian posterior, enabling amortized, likelihood-free inference. It targets applied researchers who want an approachable interface with sensible defaults and built-in posterior diagnostics. There are four vignettes including <a href="https://cran.r-project.org/web/packages/neuralsbi/vignettes/neuralsbi.html" rel="nofollow" target="_blank">Getting started</a> and <a href="https://cran.r-project.org/web/packages/neuralsbi/vignettes/sir-epidemic.html" rel="nofollow" target="_blank">Case study: inferring epidemic parameters (SIR)</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/neuralsbi.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-10" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/neuralsbi.png?w=578&#038;ssl=1" class="img-fluid" alt="Plots showing fit of multi-modal estimator" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=NBvarsel" rel="nofollow" target="_blank">NBvarsel</a> v0.1.1: Performs exhaustive or groupwise (backward elimination) variable selection for binary outcome prediction models using cross-validated net benefit as the optimization criterion. It supports predictor costs, restricted cubic splines, interaction terms, permutation importance, and parallel computation. It includes visualizations for model comparison and variable importance. References include <a href="doi:10.1177/0272989X06295361" rel="nofollow" target="_blank">Vickers &#038; Elkin (2006)</a> and <a href="doi:10.1016/j.eururo.2018.08.038" rel="nofollow" target="_blank">Van Calster et al. (2018)</a>. See the <a href="https://cran.r-project.org/web/packages/NBvarsel/vignettes/nb-varsel-tutorial.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/NBvarsel.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-11" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/NBvarsel.png?w=578&#038;ssl=1" class="img-fluid" alt="Adjusted net benefit for predictors" data-recalc-dims="1"></a></p>
</section>
<section id="medical-statistics" class="level3">
<h3 class="anchored" data-anchor-id="medical-statistics">Medical Statistics</h3>
<p><a href="https://cran.r-project.org/package=CausalState" rel="nofollow" target="_blank">CausalState</a> v0.10.2: Implements sequential doubly robust and infinite-dimensional targeted maximum likelihood estimators for longitudinal modified treatment policies in settings with transitioning states, such as ICU, ward, or emergency department care episodes. Treatment is permitted in active states and becomes structurally inapplicable after a state transition (e.g. discharge or death). Methods based on <a href="doi:10.1080/01621459.2021.1955691" rel="nofollow" target="_blank">Diaz et al. (2021)</a> and <a href="doi:10.48550/arXiv.1705.02459" rel="nofollow" target="_blank">Luedtke et al. (2017)</a>. There are three vignettes including <a href="https://cran.r-project.org/web/packages/CausalState/vignettes/getting-started.html" rel="nofollow" target="_blank">Getting started</a> and <a href="https://cran.r-project.org/web/packages/CausalState/vignettes/wb-metalearner.html" rel="nofollow" target="_blank">Wu-Benkeser density-ratio metalearner</a>.</p>
<p><a href="https://cran.r-project.org/package=orthoMTL" rel="nofollow" target="_blank">orthoMTL</a> v0.1.0: Fits regularized multi-task learning models where relationships between tasks are controlled via orthogonality or disjoint-support constraints. Supports regression, binary classification, and censored survival data. In survival mode, time-to-event outcomes are converted into binary labels at user-defined thresholds, enabling the discovery of features with time-varying effects that standard proportional-hazards models cannot detect. Implements the penalty described in <a href="https://hal.science/hal-00985654" rel="nofollow" target="_blank">Vervier et al. (2014)</a>. See the <a href="https://cran.r-project.org/package=orthoMTL" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/orthoMTL.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-12" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/orthoMTL.png?w=578&#038;ssl=1" class="img-fluid" alt="Heatmap showing coefficient weights across time thresholds" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=expoquimR" rel="nofollow" target="_blank">expoquimR</a> v0.1.0: Provides a unified toolkit for occupational chemical exposure risk assessment, implementing three internationally recognized methods: the qualitative control-banding methods COSHH Essentials (UK Health and Safety Executive) and the method of the French National Research and Safety Institute (INRS), together with the quantitative statistical procedure of the UNE-EN 689 standard for comparing measured exposure levels against occupational exposure limits. Every step of each method is implemented as a small, independently callable, and unit-tested function, so assessments are reproducible and auditable. Optional <code>shiny</code> applications provide a guided, interactive workflow. References: <a href="https://www.hse.gov.uk/coshh/essentials/index.htm" rel="nofollow" target="_blank">UK Health and Safety Executive (2003)</a> and Mallet, Pilorget and Berne (2013), ISBN:978-2-7389-2166-2. There are three vignettes including <a href="https://cran.r-project.org/web/packages/expoquimR/vignettes/coshh-essentials.html" rel="nofollow" target="_blank">COSHH Essentials: Qualitative Chemical Risk Assessment</a> and <a href="https://cran.r-project.org/web/packages/expoquimR/vignettes/inrs-method.html" rel="nofollow" target="_blank">INRS Method: Qualitative Inhalation Risk Assessment</a></p>
<p><a href="https://cran.r-project.org/package=rdborrow" rel="nofollow" target="_blank">rdborrow</a> v0.0.4.0: Implements causal inference methods for incorporating external control data into randomized controlled trials with longitudinal outcomes. Provides an analysis module supporting weighting-based methods such as inverse probability weighting and augmented inverse probability weighting, difference-in-differences. Methods are based on <a href="doi:10.1093/biostatistics/kxae012" rel="nofollow" target="_blank">Zhou et al. (2024)</a> and <a href="doi:10.1080/01621459.2024.2395586" rel="nofollow" target="_blank">Zhou et al. (2024)</a>. There are five vignettes including <a href="https://cran.r-project.org/web/packages/rdborrow/index.html" rel="nofollow" target="_blank">Introduction</a> and <a href="https://cran.r-project.org/web/packages/rdborrow/vignettes/primary_analysis_workflow.html" rel="nofollow" target="_blank">Primary Analysis Workflow</a>.</p>
</section>
<section id="meta-analysis" class="level3">
<h3 class="anchored" data-anchor-id="meta-analysis">Meta Analysis</h3>
<p><a href="https://cran.r-project.org/package=metaselection" rel="nofollow" target="_blank">metaselection</a> v0.3.0: Fits a flexible class of p-value selection models for meta-analysis and meta-regression models, providing standard errors and confidence intervals based on either cluster-robust variance estimators (i.e., sandwich estimators) or cluster-level bootstrapping to handle dependent effect size estimates, as described in <a href="doi:10.31222/osf.io/qg5x6_v1" rel="nofollow" target="_blank">Pustejovsky, Citkowicz, and Joshi (2025)</a> and <a href="doi:10.31222/osf.io/wjpxk_v1" rel="nofollow" target="_blank">Citkowicz, Pustejovsky, and Joshi (2026)</a>. Supported models include generalizations of the step-function selection model as proposed by <a href="doi:10.1007/BF02294384" rel="nofollow" target="_blank">Vevea and Hedges (1995)</a> and the beta-function selection model as proposed by <a href="doi:10.1037/met0000119" rel="nofollow" target="_blank">Citkowicz and Vevea (2017)</a>. See the <a href="https://cran.r-project.org/web/packages/metaselection/vignettes/selection-models.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/metaselection.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-13" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/metaselection.png?w=578&#038;ssl=1" class="img-fluid" alt="Distribution of one sided p-value" data-recalc-dims="1"></a></p>
</section>
</section>
</div><div class="column" style="width:10%;">

</div><div class="column" style="width:45%;">
<section id="networks" class="level3">
<h3 class="anchored" data-anchor-id="networks">Networks</h3>
<p><a href="https://cran.r-project.org/package=idiographic" rel="nofollow" target="_blank">idiographic</a> v0.3.4: Provides functions to make person-specific and within-person network estimation from intensive longitudinal and panel data. Estimators include ordinary vector autoregression (VAR), graphical vector autoregression, multilevel vector autoregression, rolling ordinary and graphical VAR, native Bayesian VAR, multilevel Bayesian VAR, unified Structural Equation Modeling, and Group Iterative Multiple Model Estimation. Methods are described in <a href="doi:10.1007/978-3-031-95365-1_20" rel="nofollow" target="_blank">Saqr et al. 2025</a> and <a href="doi:10.1080/00273171.2018.1454823" rel="nofollow" target="_blank">Epskamp et al. (2018)</a>. There are seven vignettes including an <a href="https://cran.r-project.org/web/packages/idiographic/vignettes/idiographic.html" rel="nofollow" target="_blank">introduction</a> and <a href="https://cran.r-project.org/web/packages/idiographic/vignettes/mlvar.html" rel="nofollow" target="_blank">Multilevel VAR</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/idiographic.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-14" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/idiographic.png?w=578&#038;ssl=1" class="img-fluid" alt="Example of within person network" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=lame" rel="nofollow" target="_blank">lame</a> v1.3.4: Implements additive and multiplicative effects models for both cross-sectional and longitudinal network analysis and supports square and rectangular network structures. Key features include: (1) Cross-sectional network analysis with support for binary, continuous, ordinal, and count data; (2) Longitudinal network analysis with additive sender/receiver and multiplicative latent-factor effects that can evolve over time through AR(1) processes’ (<a href="doi:10.1080/01621459.2014.988214" rel="nofollow" target="_blank">Sewell and Chen (2015)</a> and <a href="doi:10.1093/biomet/asu040" rel="nofollow" target="_blank">Durante and Dunson (2014)</a>); (3) Handling of changing actor compositions across time periods in longitudinal models; and (4) Performance improvements. There are seven vignettes including <a href="https://cran.r-project.org/web/packages/lame/vignettes/lame-overview.html" rel="nofollow" target="_blank">Overview</a> and <a href="https://cran.r-project.org/web/packages/lame/vignettes/lame.html" rel="nofollow" target="_blank">Getting Started</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/lame.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-15" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/lame.png?w=578&#038;ssl=1" class="img-fluid" alt="Circular plot showing multiplicative effects" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=tabulergm" rel="nofollow" target="_blank">tabulergm</a> v0.1.0: Creates publication-ready tables documenting exponential-family random graph models (ERGMs), a class of statistical models for social networks (<a href="doi:10.1016/j.socnet.2006.08.002" rel="nofollow" target="_blank">Robins et al. (2007)</a>). Tables describe model terms through their definitions, mathematical representations, and graphical representations, and can be generated from ERGM formulas or from models fitted with the <code>ergm</code> package (<a href="doi:10.18637/jss.v024.i03" rel="nofollow" target="_blank">Hunter et al. (2008)</a>. Resulting tables can be integrated into <code>quarto</code> and <code>rmarkdown</code> documents.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/tabulergm.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-16" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/tabulergm.png?w=578&#038;ssl=1" class="img-fluid" alt="Table with glyphs and equations" data-recalc-dims="1"></a></p>
</section>
<section id="process-control" class="level3">
<h3 class="anchored" data-anchor-id="process-control">Process Control</h3>
<p><a href="https://cran.r-project.org/package=gpci" rel="nofollow" target="_blank">gpci</a> v0.1.0: Implements a comprehensive, generalized framework for computing, estimating, and validating generalized process capability indices. Supports user-supplied probability density functions, cumulative distribution functions, survival functions, and quantile functions with uncensored data parameter estimation via Maximum Likelihood Estimation. Provides several classical and non-normal capability indices, including Cpy (<a href="doi:10.1080/16843703.2010.11673233" rel="nofollow" target="_blank">Maiti, Saha and Nanda, (2010)</a>, Spmk (<a href="doi:10.1007/s41872-019-00081-4" rel="nofollow" target="_blank">Dey and Saha (2019)</a>, CpTk (<a href="doi:10.1007/s13198-019-00789-7" rel="nofollow" target="_blank">Saha, Dey and Maiti (2019)</a> and others. Functions also compute parametric and non-parametric bootstrap confidence intervals, confidence levels using percentile, highest posterior density intervals and Heidelberger-Welch convergence diagnostics. See <a href="doi:10.1080/16843703.2010.11673233" rel="nofollow" target="_blank">Maiti, Saha and Nanda (2010)</a> and <a href="doi:10.1080/21681015.2018.1437793" rel="nofollow" target="_blank">Saha, Dey and Maiti (2018)</a> for background. See the vignettes <a href="https://cran.r-project.org/web/packages/gpci/vignettes/getting-started.html" rel="nofollow" target="_blank">Getting Started</a> and <a href="https://cran.r-project.org/web/packages/gpci/vignettes/custom-distributions.html" rel="nofollow" target="_blank">Using Custom Distributions</a>.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/gpci.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-17" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/gpci.png?w=578&#038;ssl=1" class="img-fluid" alt="Process run chart" data-recalc-dims="1"></a></p>
</section>
<section id="programming" class="level3">
<h3 class="anchored" data-anchor-id="programming">Programming</h3>
<p><a href="https://cran.r-project.org/package=polyglotSQL" rel="nofollow" target="_blank">polyglotSQL</a> v0.1.0: Provides functions to parse, tokenize, validate, format, analyze and translate SQL between more than 30 dialects (<code>PostgreSQL</code>, <code>MySQL</code>, <code>BigQuery</code>, <code>Snowflake</code>, <code>DuckDB</code>, <code>T-SQL</code>, and others) using <code>polyglot-sql</code> <a href="https://github.com/tobilg/polyglot" rel="nofollow" target="_blank">Rust crate</a>. All processing happens locally in the <code>R</code> session. Includes column-level lineage, structural query analysis, query optimization, <code>AST</code> diffing and <code>OpenLineage</code> facet generation. There are four vignettes including <a href="https://cran.r-project.org/web/packages/polyglotSQL/vignettes/getting-started.html" rel="nofollow" target="_blank">Getting Started</a> and <a href="https://cran.r-project.org/web/packages/polyglotSQL/vignettes/parsing-validation-lineage.html" rel="nofollow" target="_blank">Parsing, validation and lineage</a>.</p>
</section>
<section id="public-transit" class="level3">
<h3 class="anchored" data-anchor-id="public-transit">Public Transit</h3>
<p><a href="https://cran.r-project.org/package=transittraj" rel="nofollow" target="_blank">transittraj</a> v1.1.0: Today’s public transit vehicles produce a large amount of automatic vehicle location (AVL) data which is very useful for planning and performance studies, but can be noisy, error-prone, and sparse. This package provides tools for cleaning AVL point data and turning it into continuous, differentiable, monotonic, and invertible vehicle trajectory functions, based on the work of <a href="doi:10.48550/arXiv.2509.00119" rel="nofollow" target="_blank">Robbennolt et al. (2025)</a> and <a href="doi:10.1109/ITSC57777.2023.10422524" rel="nofollow" target="_blank">Huang et al. (2023)</a>. See the vignettes <a href="https://cran.r-project.org/web/packages/transittraj/vignettes/intro-trajectories.html" rel="nofollow" target="_blank">Introduction</a> and <a href="https://cran.r-project.org/web/packages/transittraj/vignettes/intro-trajectories.html" rel="nofollow" target="_blank">The AVL Cleaning Workflow</a>.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/transittraj.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-18" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/transittraj.png?w=578&#038;ssl=1" class="img-fluid" alt="Plot of LA E line vehicle trajectories" data-recalc-dims="1"></a></p>
</section>
<section id="statistics" class="level3">
<h3 class="anchored" data-anchor-id="statistics">Statistics</h3>
<p><a href="https://cran.r-project.org/package=chaidr" rel="nofollow" target="_blank">chaidr</a> v0.1.0: Implements the CHAID (Chi-squared Automatic Interaction Detection) decision tree algorithm of <a href="doi:10.2307/2986296" rel="nofollow" target="_blank">Kass (1980)</a> and the Exhaustive CHAID variant of <a href="doi:10.1080/02664769100000005" rel="nofollow" target="_blank">Biggs, de Ville, and Suen (1991)</a>, as specified in the <code>IBM SPSS</code> Statistics Algorithms documentation. Supports nominal, ordinal with floating missing category, and continuous predictors, and nominal, ordinal, and continuous response variables using Pearson chi-squared, Goodman row-effects, and one-way ANOVA F tests respectively. Includes prediction, rule extraction, gains and lift analysis, validation on holdout data, and visualization via base graphics, <code>Graphviz</code> DOT, <code>plotly</code>, and conversion to <code>partykit</code> objects. See the <a href="https://cran.r-project.org/web/packages/chaidr/vignettes/chaidr.html" rel="nofollow" target="_blank">Introduction</a> and the <a href="https://cran.r-project.org/web/packages/chaidr/vignettes/chaidr-ja.html" rel="nofollow" target="_blank">Japanese tutorial</a>.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/chaidr.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-19" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/chaidr.png?w=578&#038;ssl=1" class="img-fluid" alt="CHAID decision tree" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=citcdf" rel="nofollow" target="_blank">citcdf</a> v1.1.0: Enables complex hypothesis testing through conditional cumulative distribution function estimation. Method is detailed in: <a href="doi:10.1101/2021.05.21.445165" rel="nofollow" target="_blank">Gauthier et al. (2021)</a>. See the <a href="https://cran.r-project.org/web/packages/citcdf/vignettes/citcdf-userguide.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/citcdf.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-20" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/citcdf.png?w=578&#038;ssl=1" class="img-fluid" alt="Plot of p-values sorted against BH threshold" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=falsifyr" rel="nofollow" target="_blank">falsifyr</a> v1.0.0: Attacks fitted <code>R</code> model claims by searching for small plausible perturbations that make a target result disappear. The package focuses on claim-level fragility, smallest-kill reporting, and reproducible caveated robustness checks for ordinary fitted model objects. The methods draw on the fragility-index concept of <a href="doi:10.1016/j.jclinepi.2013.10.019" rel="nofollow" target="_blank">Walsh et al. (2014)</a>, multiverse analysis of <a href="doi:10.1177/1745691616658637" rel="nofollow" target="_blank">Steegen et al. (2016)</a>, specification-curve analysis of <a href="doi:10.1038/s41562-020-0912-z" rel="nofollow" target="_blank">Simonsohn et al. (2020)</a>, and robust covariance estimation of <a href="doi:10.18637/jss.v011.i10" rel="nofollow" target="_blank">Zeileis (2004)</a>. See the vignettes <a href="https://cran.r-project.org/web/packages/falsifyr/vignettes/attacking-a-regression-claim.html" rel="nofollow" target="_blank">Attacking a Regression Claim</a> and <a href="https://cran.r-project.org/web/packages/falsifyr/vignettes/interpreting-survival-scores.html" rel="nofollow" target="_blank">Interpreting Survival Scores</a>.</p>
<p><a href="https://cran.r-project.org/package=FPScausal" rel="nofollow" target="_blank">FPScausal</a> v0.1.1: Implements functional propensity score weighting for causal inference with functional treatments. The method estimates weights that balance observed confounders by removing their dependence on the functional treatment and uses a dual formulation of the weighting problem for efficient unconstrained optimization. The framework supports scalar, binary, and functional outcomes, as well as functional covariates, and can be used to estimate marginal causal effects in settings with time-varying exposures. The methodology follows <a href="doi:10.48550/arXiv.2608.03200" rel="nofollow" target="_blank">Ciardulli et al. (2026)</a>. See the <a href="https://cran.r-project.org/web/packages/FPScausal/vignettes/FPScausal.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/FPScausal.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-21" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/FPScausal.png?w=578&#038;ssl=1" class="img-fluid" alt="Plot weighted vs unweighted causal effects" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=LRErdd" rel="nofollow" target="_blank">LRErdd</a> v0.1.0: Provides functions for the design and analysis of Regression Discontinuity Designs as local randomized experiments within the potential outcome approach as formalized in <a href="doi:10.1214/15-AOAS809" rel="nofollow" target="_blank">Li, Mattei and Mealli (2015)</a> including functions to implement the design phase of the study, where the focus is on the selection of suitable subpopulations for which valid causal inference can be drawn. These functions provide summary statistics of pre- and post-treatment variables by treatment status and select suitable subpopulations around the threshold where pre-treatment variables are well balanced between treatment groups. See the <a href="https://cran.r-project.org/web/packages/LRErdd/vignettes/LRErdd.html" rel="nofollow" target="_blank">vignette</a>. There is also a <code>Shiny</code> application.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/LRErdd.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-22" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/LRErdd.png?w=578&#038;ssl=1" class="img-fluid" alt="Distribution plots for Fisher test." data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=mvboxcox" rel="nofollow" target="_blank">mvboxcox</a> v0.1.4: Fits bivariate logistic Box-Cox regression models for binary outcomes and positive continuous predictors. Transformation parameters are selected by cross-validated grid search with adaptive refinement and thin-plate spline smoothing. The package also provides prediction, empirical and sampling-weighted median effects, simulation tools, and sampling-weighted model fitting. The methodology extends the logistic Box-Cox approach of <a href="doi:10.1002/cjs.11587" rel="nofollow" target="_blank">Xing et al. (2021)</a>. See the <a href="https://cran.r-project.org/web/packages/mvboxcox/vignettes/introduction.html" rel="nofollow" target="_blank">vignette</a>.</p>
</section>
<section id="surveys" class="level3">
<h3 class="anchored" data-anchor-id="surveys">Surveys</h3>
<p><a href="https://cran.r-project.org/package=sondage" rel="nofollow" target="_blank">sondage</a> v0.9.1: Implements survey sampling algorithms for single-stage probability sampling from finite populations, written in <code>C</code>. Provides equal probability methods (simple random sampling, systematic, Bernoulli), unequal probability methods (conditional Poisson / maximum entropy, Sampford, Brewer, systematic PPS, Pareto, sequential Poisson, Poisson, Chromy’s minimum replacement, multinomial), balanced sampling via the cube method, and spatially balanced sampling via the local pivotal method and spatially correlated Poisson sampling. Functions compute joint inclusion probabilities, pairwise expectations, and sampling covariances for variance estimation. See <a href="doi:10.1007/0-387-34240-0" rel="nofollow" target="_blank">Tillé (2006)</a> for background. There are two vignettes <a href="https://cran.r-project.org/web/packages/sondage/vignettes/sondage.html" rel="nofollow" target="_blank">Getting Started</a> and <a href="https://cran.r-project.org/web/packages/sondage/vignettes/custom-methods.html" rel="nofollow" target="_blank">Extending sondage with Custom Methods</a>.</p>
</section>
<section id="time-series" class="level3">
<h3 class="anchored" data-anchor-id="time-series">Time Series</h3>
<p><a href="https://cran.r-project.org/package=scanr" rel="nofollow" target="_blank">scanr</a> v0.1.1: Detects change points in long univariate time series using the <a href="https://arxiv.org/html/2608.28110v1" rel="nofollow" target="_blank">SCAN framework</a>. The implementation uses a native <code>Rust</code> backend exposed to <code>R</code> via <code>extendr</code>. See the <a href="https://cran.r-project.org/web/packages/scanr/vignettes/scanr-introduction.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/scanr.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-23" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/scanr.png?w=578&#038;ssl=1" class="img-fluid" alt="Plot of time series with detected change points" data-recalc-dims="1"></a></p>
</section>
<section id="utilities" class="level3">
<h3 class="anchored" data-anchor-id="utilities">Utilities</h3>
<p><a href="https://cran.r-project.org/package=rmoriebricklayer" rel="nofollow" target="_blank">rmoriebricklayer</a> v0.5.1: Tools for building brick-proof, reproducible, self-contained data capsules. Resolves open-data sources through the <a href="https://ckan.org/" rel="nofollow" target="_blank">Comprehensive Knowledge Archive Network</a> <code>package_show</code>and <code>package_search</code> endpoints, records and verifies provenance with Secure Hash Algorithm 256 (SHA-256) digests and Internet Archive <a href="https://web.archive.org/" rel="nofollow" target="_blank">Wayback Machine</a> snapshots, validates downloaded data against a pinned schema. Run records are captured in a manifest plus a plain-language summary so any result can be traced back to its inputs. Tests distributional drift between a pinned capsule and a fresh fetch because a re-released extract can be statistically identical yet differ byte-for-byte. Manifests can be authenticated with keyed digests (HMAC-SHA-256, RFC 2104) or post-quantum hash-based signatures. There are eight vignettes including <a href="https://cran.r-project.org/web/packages/rmoriebricklayer/vignettes/capsules.html" rel="nofollow" target="_blank">Building reproducible data capsules</a> and <a href="https://cran.r-project.org/web/packages/rmoriebricklayer/vignettes/provenance.html" rel="nofollow" target="_blank">Provenance you can verify</a>.</p>
<p><a href="https://cran.r-project.org/package=scimesh" rel="nofollow" target="_blank">scimesh</a> v0.4.0: Implements a fast, GPU-free 3D software renderer written in modern <code>C++17</code> with native <code>R</code> bindings. Renders triangle meshes to publication-quality images entirely on the CPU, requiring no display server or graphics hardware. Features multi-light Blinn-Phong shading, screen-space ambient occlusion, anti-aliasing, depth fog, transparency, wireframe rendering, texture mapping, screen-space lines and text labels, and procedural geometry generation. Supports standard mesh file formats with PNG and PPM output. Works on high-performance computing clusters, headless servers, containers, and continuous integration pipelines, making it suitable for scientific visualization across neuro-imaging, molecular structures, and general 3D graphics. See the <a href="https://cran.r-project.org/web/packages/scimesh/vignettes/scimesh.html" rel="nofollow" target="_blank">vignette</a>.</p>
</section>
<section id="visualization" class="level3">
<h3 class="anchored" data-anchor-id="visualization">Visualization</h3>
<p><a href="https://cran.r-project.org/web/packages/dgraphs/vignettes/data-derived-graph-workflow.html" rel="nofollow" target="_blank">dgraphs</a> v0.2.0: Constructs data-derived graphs from numerical observations using mutual, shared-neighbor, intersection, geodesic, radius, adaptive-radius, and minimum-spanning-tree completion methods. Provides graph conversion, pruning, diagnostics, spectral embedding, endpoint detection, and path utilities. The implemented graph constructions include methods described by <a href="doi:10.1109/T-C.1973.223640" rel="nofollow" target="_blank">Jarvis and Patrick (1973)</a>, <a href="doi:10.1016/S0167-7152(96)00213-1" rel="nofollow" target="_blank">Brito et al. (1997)</a>, <a href="doi:10.3934/fods.2019001" rel="nofollow" target="_blank">Berry and Sauer (2019)</a>, and <a href="doi:10.2307/2346439" rel="nofollow" target="_blank">Gower and Ross (1969)</a>. See the <a href="https://cran.r-project.org/web/packages/dgraphs/vignettes/data-derived-graph-workflow.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/dgraphs.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-24" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/dgraphs.png?w=578&#038;ssl=1" class="img-fluid" alt="Continuous-kNN graph on the variable-density circular point cloud" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=gghotelling" rel="nofollow" target="_blank">gghotelling</a> v0.2.1: Calculate Hotelling’s T² ellipses and detect multivariate outliers both for base <code>R</code> plots and <code>ggplot2</code> plots. Optionally, uses robust covariance estimation to reduce the influence of outliers on the ellipses. Also included: bagplots, kernel density plots and outlier diagnostic plots. See the <a href="https://cran.r-project.org/web/packages/gghotelling/vignettes/gghotelling.html" rel="nofollow" target="_blank">vignette</a>.</p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/gghotelling.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-25" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/gghotelling.png?w=578&#038;ssl=1" class="img-fluid" alt="Plot of Hotelling ellipses with minimum covariance determinant estimator" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=ggmultiglyph" rel="nofollow" target="_blank">ggmultiglyph</a> v0.1.0: Provides <code>ggplot2</code> geoms for visualizing multivariate data using glyphs. The package implements several established glyph designs described in the information visualization literature, including the review by <a href="doi:10.2312/conf/EG2013/stars/039-063" rel="nofollow" target="_blank">Borgo et al. (2013)</a>.</p>
<p><a href="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/ggmultiglyph.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-26" rel="nofollow" target="_blank"><img src="https://i2.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/ggmultiglyph.png?w=578&#038;ssl=1" class="img-fluid" alt="Glyphs" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=grip" rel="nofollow" target="_blank">grip</a> v0.2.0: Implements GRIP multiscale graph layout with a unified choice between hop-count and geometry-aware edge-length graph metrics in 2D and 3D. Provides layout scoring, candidate comparison, multiscale trace diagnostics, synthetic graph families, and advanced experimental geodesic-KK utilities for weighted-layout evaluation and polish. Based on <a href="doi:10.7155/jgaa.00052" rel="nofollow" target="_blank">Gajer and Kobourov (2002)</a> and <a href="doi:10.1016/j.comgeo.2004.03.014" rel="nofollow" target="_blank">Gajer, Goodrich and Kobourov (2004)</a>. There are four vignettes including <a href="https://cran.r-project.org/web/packages/grip/vignettes/grip-examples.html" rel="nofollow" target="_blank">Getting Started</a> and <a href="https://cran.r-project.org/web/packages/grip/vignettes/weighted-grip-intro.html" rel="nofollow" target="_blank">Weighted Graph Layouts</a>.</p>
<p><a href="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/grip.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-27" rel="nofollow" target="_blank"><img src="https://i0.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/grip.png?w=578&#038;ssl=1" class="img-fluid" alt="Plots of large graph weighted patterns" data-recalc-dims="1"></a></p>
<p><a href="https://cran.r-project.org/package=orbis" rel="nofollow" target="_blank">orbis</a> v0.1.0: Implements a layered grammar of graphics that compiles plots to a resolution-independent scene description and renders it through two back-ends: a self-contained SVG writer with embedded <code>JavaScript</code> for interactive figures (tooltips, hover highlighting, zoom, pan and legend toggling) and <code>R</code>’s own graphics devices for publication-quality output at any resolution. Geographic layers are first class. The layered grammar follows <a href="doi:10.1198/jcgs.2009.07098" rel="nofollow" target="_blank">Wickham (2010)</a>; projections follow <a href="doi:10.3133/pp1395" rel="nofollow" target="_blank">Snyder (1987)</a> and, for Equal Earth, <a href="doi:10.1080/13658816.2018.1504949" rel="nofollow" target="_blank">Savric, Patterson and Jenny (2019)</a>; line simplification uses <a href="doi:10.3138/FM57-6770-U75U-7727" rel="nofollow" target="_blank">Douglas and Peucker (1973)</a>; the default colour scales follow the guidance on perceptually uniform palettes of <a href="doi:10.1038/s41467-020-19160-7" rel="nofollow" target="_blank">Crameri, Shephard and Heron (2020)</a>. See the <a href="https://cran.r-project.org/web/packages/orbis/vignettes/orbis.html" rel="nofollow" target="_blank">vignette</a></p>
<p><a href="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/orbis.png?ssl=1" class="lightbox" data-gallery="quarto-lightbox-gallery-28" rel="nofollow" target="_blank"><img src="https://i1.wp.com/rworks.dev/posts/august-2026-top-40-new-cran-packages/orbis.png?w=578&#038;ssl=1" class="img-fluid" alt="World choropleth: Robinson projection" data-recalc-dims="1"></a></p>
</section>
</div>
</div>



 
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://rworks.dev/posts/august-2026-top-40-new-cran-packages/"> R Works</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/august-2026-top-40-new-cran-packages/">August 2026 Top 40 New CRAN Packages</a>]]></content:encoded>
					
		
		<enclosure url="https://rworks.dev/posts/august-2026-top-40-new-cran-packages/gbif.png" length="0" type="image/png" />

		<post-id xmlns="com-wordpress:feed-additions:1">403988</post-id>	</item>
		<item>
		<title>Celebrating the rOpenSci Community</title>
		<link>https://www.r-bloggers.com/2026/09/celebrating-the-ropensci-community/</link>
		
		<dc:creator><![CDATA[rOpenSci]]></dc:creator>
		<pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://ropensci.org/blog/2026/09/29/celebrating-ropensci/</guid>

					<description><![CDATA[<p>Read it in: Español. rOpenSci turns 15 and I’ve been a part of this community for a good chunk of that time. As I reflected on those years, a lot of great memories came to mind, and I realized some of them taught me some lessons that I would ...</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/celebrating-the-ropensci-community/">Celebrating the rOpenSci Community</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://ropensci.org/blog/2026/09/29/celebrating-ropensci/"> rOpenSci - open tools for open science</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<p><a href='https://ropensci.org/es/blog/2026/09/29/celebrando-ropensci/' rel="nofollow" target="_blank">Read it in: Español</a>.</p> <p>rOpenSci turns 15 and I’ve been a part of this community for a good chunk of that time. As I reflected on those years, a lot of great memories came to mind, and I realized some of them taught me some lessons that I would love to share.</p>
<h2>
Stretch your comfort zone
</h2><p>Early in 2018 I came across <a href="https://ropensci.org/blog/2018/02/08/unconf2018/" rel="nofollow" target="_blank">this post</a>:</p>
<blockquote>
<p><strong>Apply to attend rOpenSci unconf 2018!</strong><br>
We’re organizing unconf18 to bring together scientists, developers, and open data enthusiasts from academia, industry, government, and non-profits to get together for a couple of days to hack on various projects and generally enrich our community.</p>
</blockquote>
<p>In hindsight I was a good fit. I had recently transitioned from academia to my first job as a research software engineer. And yet, did I think I was a solid candidate? Oh no, I had lots of doubts.</p>
<p>But I know the greatest things often happen a little stretch outside your comfort zone. So I applied anyway, got invited, and spent two days collaborating with <a href="https://unconf18.ropensci.org/#participants" rel="nofollow" target="_blank">an incredible bunch of people</a>. Some of them were already rockstars, many others eventually became influential in their own niche, and all of them had this great attitude that at the time I lacked words to describe. I now do: they were respectful and kind.</p>
<h2>
Be respectful and kind
</h2><p>In 2022 <a href="https://ropensci.org/blog/2022/09/22/launch-champions-program/" rel="nofollow" target="_blank">rOpenSci launched the first cohort of the rOpenSci Champions Program</a>. By then I already had a few years of experience as an associate editor of <a href="https://ropensci.org/software-review/" rel="nofollow" target="_blank">rOpenSci software peer-review</a>, and <a href="https://ropensci.org/author/yanina-bellini-saibene/" rel="nofollow" target="_blank">Yani</a> invited me to talk to our champions about the process.</p>
<p>As I prepared <a href="https://ropensci-training.github.io/software-review/en/" rel="nofollow" target="_blank">that talk</a> and discussed some ideas with Yani, she made me realize how much the rOpenSci community cares about the way we communicate with one another<sup id="fnref:1"><a href="https://ropensci.org/blog/2026/09/29/celebrating-ropensci/#fn:1" class="footnote-ref" role="doc-noteref" rel="nofollow" target="_blank">1</a></sup>. So much so that our guides for <a href="https://devguide.ropensci.org/softwarereview_reviewer.html" rel="nofollow" target="_blank">reviewers</a> and <a href="https://devguide.ropensci.org/softwarereview_editor.html" rel="nofollow" target="_blank">editors</a> open with this message:</p>
<blockquote>
<p>rOpenSci’s community is our best asset. We aim for reviews to be open, non-adversarial, and focused on improving software quality. Be respectful and kind! See our reviewers’ guide and <a href="https://ropensci.org/code-of-conduct/" rel="nofollow" target="_blank">code of conduct</a> for more.</p>
</blockquote>
<p>Coming from such a deeply technical guide, that opening might surprise you. But it makes sense to me now; broken code is way easier to fix than broken human relationships.</p>
<h2>
Contribute in your own way
</h2><p>At rOpenSci you can contribute in so many ways. We have a <a href="https://contributing.ropensci.org/" rel="nofollow" target="_blank">community contributing guide</a> but here’s a list of some of my own contributions:</p>
<ul>
<li>Attend <a href="https://ropensci.org/community/" rel="nofollow" target="_blank">events</a>.</li>
<li>Help to welcome and onboard new members.</li>
<li>Ask or answer questions on Slack, or share or discuss ideas or jobs.</li>
<li>Improve our documentation, e.g. fix a typo.</li>
<li>Write a <a href="https://ropensci.org/blog/" rel="nofollow" target="_blank">blog</a> post.</li>
<li><a href="https://ropensci.org/multilingual-publishing/" rel="nofollow" target="_blank">Translate</a> or review some work in your native language.</li>
<li>Pilot a new process.</li>
<li>Author, review, or edit an R package.</li>
<li>Lead a workshop.</li>
<li>Build a tool to enhance some process.</li>
<li>Give and get support, e.g. mentor, review grant applications or talks, advice or recommendations for jobs, help unavailable or overwhelmed people.</li>
</ul>
<p>As you can see, most of my contributions did not involve any code, and only a few required an invitation. Some of them are the kind of thing you could add to your CV, and the most important ones only belong in your heart.</p>
<h2>
Attract great people
</h2><p>Great communities are made of great people, and you can play an active role in growing it in the direction you want. If you know someone that would fit in and enjoy the rOpenSci community, you can try to <a href="https://contributing.ropensci.org/" rel="nofollow" target="_blank">find a way to bring them in</a>. For example, I’ve encouraged and helped people to submit and review packages. Also as an editor I have the occasional privilege to nominate other editors. Using this super-power I’ve attracted two <a href="https://ropensci.org/software-review/" rel="nofollow" target="_blank">editors</a> and I now get to enjoy interacting with them quite regularly.</p>
<h2>
Keep the human in your human connections
</h2><p>AI is changing the way we contribute to open source software. But we don’t know much about its effect on the open communities around that software. Earlier this year rOpenSci published a <a href="https://ropensci.org/blog/2026/02/26/ropensci-ai-policy/" rel="nofollow" target="_blank">preliminary set of policies</a>, and you may want to watch rOpenSci’s <a href="https://ropensci.org/news/" rel="nofollow" target="_blank">newsletter</a> and this <a href="https://openscapes.org/events/2026-10-01-community-call-open-communities-ai/" rel="nofollow" target="_blank">Openscapes community call</a>:</p>
<blockquote>
<p><strong>Open Communities in the Age of AI</strong><br>
{Open communities} help us connect on a human level around science and data and the things that we are passionate about. They change careers, they change lives. (…) But now, people are not necessarily finding their communities and getting the benefit of those communities in the way that they used to, because they’re more easily able to get help and answers by using AI. (…) These connections and communities are more important than ever, as we need to find ways to support each other and ourselves as AI rapidly changes the landscape right under our feet.</p>
</blockquote>
<p>Personally, over the past year I’ve used AI heavily and learned several lessons. The most relevant here is that I don’t want AI to impersonate me or the other person in a <a href="https://contributing.ropensci.org/motivations.html#connect" rel="nofollow" target="_blank">human-to-human connection</a>. The <a href="https://contributing.ropensci.org/intro.html#humans" rel="nofollow" target="_blank">humans of rOpenSci</a> are real and wonderful people. Here are some that I’ve recently met in person:</p>
<div class="row">
<div class="col-md-6">
<figure><img src="https://i0.wp.com/ropensci.org/blog/2026/09/29/celebrating-ropensci/2024_boston_zci_yani-noam-mauro.JPG?w=578&#038;ssl=1"
alt="Noam Ross, Yanina Bellini Saibene, and Mauro Lepore at the 2024 CZI meeting in Boston, USA" data-recalc-dims="1"><figcaption>
<p><a href='https://www.linkedin.com/in/noamross/' rel="nofollow" target="_blank">Noam Ross</a>, <a href='https://www.linkedin.com/in/yabellini/' rel="nofollow" target="_blank">Yanina Bellini Saibene</a> and <a href='https://www.linkedin.com/in/mauro-lepore/' rel="nofollow" target="_blank">me</a> at the 2024 <a href='https://chanzuckerberg.com/' rel="nofollow" target="_blank">CZI</a> meeting in Boston, USA.</p>
</figcaption>
</figure>
</div>
<div class="col-md-6">
<figure><img src="https://i0.wp.com/ropensci.org/blog/2026/09/29/celebrating-ropensci/2024_seattle_posit-conf_monica-stefanie-sean-julia-mauro-kelly.png?w=578&#038;ssl=1"
alt="Monica Gerber, Stefanie Butland, Sean Kross, Julia Stewart Lowndes, Kelly O&#39;Briant, and Mauro Lepore at the 2024 Posit conference in Seattle, USA" data-recalc-dims="1"><figcaption>
<p><a href='https://www.linkedin.com/in/monica-gerber/' rel="nofollow" target="_blank">Monica Gerber</a>, <a href='https://www.linkedin.com/in/stefaniebutland/' rel="nofollow" target="_blank">Stefanie Butland</a>, <a href='https://www.linkedin.com/in/seankross/' rel="nofollow" target="_blank">Sean Kross</a>, <a href='https://www.linkedin.com/in/julia-stewart-lowndes/' rel="nofollow" target="_blank">Julia Stewart Lowndes</a>, <a href='https://www.linkedin.com/in/kellyobriant/' rel="nofollow" target="_blank">Kelly O’Briant</a> and me at the 2024 Posit conference in Seattle, USA.</p>
</figcaption>
</figure>
</div>
</div>
<div class="row">
<div class="col-md-6">
<figure><img src="https://i0.wp.com/ropensci.org/blog/2026/09/29/celebrating-ropensci/2025_atlanta_posit-conf_luis-diana-mauro.jpg?w=578&#038;ssl=1"
alt="Luis D. Verde Arregoitia, Diana Garcia Cortes, and Mauro Lepore at the Georgia Aquarium after the 2025 Posit conference in Atlanta, USA" data-recalc-dims="1"><figcaption>
<p><a href='https://www.linkedin.com/in/luis-d-verde-arregoitia-a20339209/' rel="nofollow" target="_blank">Luis D. Verde Arregoitia</a>, <a href='https://www.linkedin.com/in/ddiannae/' rel="nofollow" target="_blank">Diana Garcia Cortes</a> and me at the Georgia Aquarium after the 2025 Posit conference in Atlanta, USA.</p>
</figcaption>
</figure>
</div>
<div class="col-md-6">
<figure><img src="https://i1.wp.com/ropensci.org/blog/2026/09/29/celebrating-ropensci/2025_san-jose_costa-rica_ronny.jpg?w=578&#038;ssl=1"
alt="Ronny A. Hernández Mora and Mauro Lepore at Ronny&#39;s family gathering near San José, Costa Rica in 2025" data-recalc-dims="1"><figcaption>
<p><a href='https://www.linkedin.com/in/ronny-hernandez-mora/' rel="nofollow" target="_blank">Ronny A. Hernández Mora</a> and me at Ronny’s family gathering near San José, Costa Rica in 2025.</p>
</figcaption>
</figure>
</div>
</div>
<div class="row">
<div class="col-md-6">
<figure><img src="https://i1.wp.com/ropensci.org/blog/2026/09/29/celebrating-ropensci/2025_sao-pablo_brasil_bea-mauro.jpg?w=578&#038;ssl=1"
alt="Beatriz Milz and Mauro Lepore after coffee and pastries in São Paulo, Brazil in 2025" data-recalc-dims="1"><figcaption>
<p><a href='https://www.linkedin.com/in/beatrizmilz/' rel="nofollow" target="_blank">Beatriz Milz</a> and me after coffee and pastries in São Paulo, Brazil in 2025.</p>
</figcaption>
</figure>
</div>
</div>
<p>Thanks for joining me in celebrating rOpenSci. And if, like me, you <a href="https://contributing.ropensci.org/intro.html#community" rel="nofollow" target="_blank">self-identify with our community</a>, then happy birthday to you!</p>
<div class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1">
<p>Two books I like are <a href="https://openlibrary.org/books/OL31981288M/How_to_Win_Friends_and_Influence_People" rel="nofollow" target="_blank">How to Win Friends and Influence People</a> and <a href="https://openlibrary.org/works/OL282391W/Crucial_Conversations" rel="nofollow" target="_blank">Crucial Conversations (Third Edition): Tools for Talking When Stakes Are High</a>. <a href="https://ropensci.org/blog/2026/09/29/celebrating-ropensci/#fnref:1" class="footnote-backref" role="doc-backlink" rel="nofollow" target="_blank"><img src="https://s.w.org/images/core/emoji/13.0.0/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></p>
</li>
</ol>
</div>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://ropensci.org/blog/2026/09/29/celebrating-ropensci/"> rOpenSci - open tools for open science</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/celebrating-the-ropensci-community/">Celebrating the rOpenSci Community</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403978</post-id>	</item>
		<item>
		<title>X Reasons for policy professionals to get into programming in 202X</title>
		<link>https://www.r-bloggers.com/2026/09/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/</link>
		
		<dc:creator><![CDATA[Giles]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 10:06:06 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://www.gilesd-j.com/?p=4701</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> In 2019 I wrote a post with seven reasons policy professionals should learn to code. Now that AI can write reasonable code, I’ve had a number of people ask me if it’s still worth learning to program. Based on the evidence I’ve seen, AI appears to be a ...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/">X Reasons for policy professionals to get into programming in 202X</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/"> Data Analytics and AI Archives - Giles</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<p class="wp-block-paragraph"><strong>TLDR:</strong> In 2019 I wrote a post with seven reasons policy professionals should learn to code. Now that AI can write <em>reasonable</em> code<em>,</em> I’ve had a number of people ask me if it’s still worth learning to program. Based on the evidence I’ve seen, AI appears to be a an amplifier, rather than a substitute for human intelligence in the policy space: meaning the people that get the most out of AI are likely to be those that know how to instruct AI to tackle a problem and tell good output from bad. So my answer is yes, learning to code still makes sense. Not necessarily so you can write the code yourself, but so you can make AI a force multiplier for your work.</p>



<p class="wp-block-paragraph"><strong>Note: </strong>I’m intending to make updates and refinements to this post over time: both to correct the typos that I haven’t picked up and to incorporate feedback from other coders in the policy community. Feel free to <a href="https://www.gilesd-j.com/contact/" rel="nofollow" target="_blank">reach out if you have suggestions</a>.  </p>



<p class="wp-block-paragraph"><strong>How AI was used</strong>: I wrote the first draft of this post. AI was used (minimally) to refine how some ideas were communicated. </p>



<h3 class="wp-block-heading">Background</h3>



<p class="wp-block-paragraph">In 2019 I wrote <a href="https://www.gilesd-j.com/2019/01/07/7-reasons-for-policy-professionals-to-get-pumped-about-r-programming-in-2019/" rel="nofollow" target="_blank">a listicle outlining why people working on public policy should learn how to program</a>.</p>



<p class="wp-block-paragraph">Putting aside that <s>R</s> <em>python</em> was probably the answer to which language to learn, the point of the post was to share some of the practical benefits I’d seen as an applied economist and policy advisor from learning to code.</p>



<p class="wp-block-paragraph">Despite the listicle being one of my lower effort posts, it seemed to strike a chord: my inbox was quickly flooded with messages thanking me for articulating the benefits others in the public policy community had seen from leveraging code in their work. To my surprise, the article was also used as a reference for several university programming courses, resulting in my inbox still receiving messages from people starting their coding journey.</p>



<p class="wp-block-paragraph">But a lot has changed since I first wrote the article: the UK voted to leave the European Union; the World Health Organization declared a global pandemic; and 25 movies in the Marvel Cinematic universe were released.</p>



<p class="wp-block-paragraph">Oh! And a quaint lil’ startup called OpenAI launched something called ChatGPT.</p>



<p class="wp-block-paragraph">And while there’s certainly a lot that could be said about the unfolding Marvel Universe’s implication for learning to code, this will have to be a subject of a future post. As today, I’d like to attempt to answer a question I’ve been increasingly asked since Artificial Intelligence (AI) tools and/or Large Language Models (LLMs) have made producing code trivial.</p>



<h3 class="wp-block-heading">It’s sunk costs all the way down</h3>



<p class="wp-block-paragraph">At the outset, I’m clearly not an impartial observer. I was motivated enough to teach myself R, to write the original post and create a dedicated MOOC on using R for policy analysis. I also <em>enjoy</em> coding, so I have a vested interest in answering <em>yes</em> to the question. But, in my defence, these reasons also mean that I’ve thought about it a lot. I’m also a frequent user of AI in my work and find it has greatly expanded the scope and scale of what I can achieve.</p>



<p class="wp-block-paragraph">This post is therefore my attempt to provide an answer to the question <em>Is it still worth learning to code in 202X?</em> My answer is, yes, but for different reasons to the 2019 listicle.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i1.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-11.png?w=450&#038;ssl=1" alt="" class="wp-image-4704" srcset_temp="https://i1.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-11.png?w=450&#038;ssl=1 634w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-11-300x256.png 300w" sizes="auto, (max-width: 634px) 100vw, 634px" data-recalc-dims="1" /><figcaption class="wp-element-caption"><strong>Source: </strong>rogierK @ Twitter (post no longer available)</figcaption></figure>



<h3 class="wp-block-heading">Doing smart stuff quickly</h3>



<p class="wp-block-paragraph">What got me excited about learning to program when I first started, was its ability to expand the scope and scale of analysis I could do and questions I could answer. In the world of public policy, there’s an almost unlimited number of interesting questions that could be asked, but rarely the time and data to answer them. Sometimes this is simply because the data didn’t exist, which makes learning to code not particularly helpful. However, often there are cases where data <em>does</em> exist, just not in a format that can be readily analysed with the commercial tools made available to you:</p>



<ul class="wp-block-list">
<li class="">Want to import thousands of individual excel files and merge them into a single dataframe?</li>



<li class="">Have a repetitive data cleaning task that would take hours to do using a spreadsheet?</li>



<li class="">Want to apply a statistical technique that hasn’t been added to your stats software?</li>
</ul>



<p class="wp-block-paragraph">Programming languages makes solving these problems trivial, but with the right prompt, so can AI.</p>



<h3 class="wp-block-heading">Probability-based slop machines</h3>



<p class="wp-block-paragraph">But let’s be clear: at their core, LLMs are probabilistic slop machines. I don’t mean <em>slop</em> in the <a href="https://en.wikipedia.org/wiki/AI_slop" rel="nofollow" target="_blank">pejorative sense</a>, but as a palpable description of what LLMs do: slap together <em>intelligent looking</em> responses based on <a href="https://ai.stackexchange.com/questions/50233/is-next-token-prediction-sufficient-to-explain-emergent-capabilities-like-comple" rel="nofollow" target="_blank">next-token prediction</a>.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn1" id="_ftnref1" rel="nofollow" target="_blank">[1]</a> I say <em>intelligent looking<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn2" id="_ftnref2" rel="nofollow" target="_blank"><strong>[2]</strong></a></em> as LLMs have been trained to produce outputs that look intelligent, rather than to exhibit intelligence in a human sense:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>“…AI systems have demonstrated their capability to solve cognitive ability test problems, primarily through guided training (e.g., Zhuo &#038; Kankanhalli, 2020) or programmed approaches to transform problems into algorithmically solvable formats (e.g., Schmidhuber, 2004). While remarkable, it is debatable whether these accomplishments signify intelligence, given that the capabilities of most current AI systems are limited to specific programming and/or training data, without the necessary demonstration of novel problem-solving ability characteristic of human intelligence (Davidson &#038; Downing, 2000; Raaheim &#038; Brun, 1985). Consequently, many AI systems might be more aptly recognised as having the capacity to exhibit artificial achievement or artificial expertise…”</em></p>
<cite><em>Gignac, G.E. and Szodorai, E.T., 2024. Defining intelligence: Bridging the gap between human and artificial perspectives. Intelligence, 104, p.101832 (</em><a href="https://www.sciencedirect.com/science/article/pii/S0160289624000266#article" rel="nofollow" target="_blank"><em>link</em></a><em>).</em></cite></blockquote>



<p class="wp-block-paragraph">But sometimes a little bit of AI slop can be helpful. When producing code, I’ll often get AI to produce the first draft before making tweaks to the approach and style.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn3" id="_ftnref3" rel="nofollow" target="_blank">[3]</a> I use AI to help with writing and research. For this post I asked AI for good search terms to find research exploring the cognitive and economic benefits of learning to code, so I could check my assumptions and think clearly about what has changed since I first wrote the post. I then asked AI for ideas to better communicate key ideas presented in this post.  </p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>“The majority of human production has always been slop. Mediocrity is not a bug of technology; it is the baseline of culture. The canon of art we revere today – those few thousands works in museums and textbooks – is the surviving tip of an immense iceberg of forgotten, derivative, or simply boring creations…”</em></p>
<cite>D’Isa, Francesco (1 December 2025). “The Idea of ‘AI Slop’ Is Slop”. The Philosophical Salon (<a href="https://thephilosophicalsalon.com/the-idea-of-ai-slop-is-slop/" rel="nofollow" target="_blank">link</a>)</cite></blockquote>



<p class="wp-block-paragraph">If you’re wondering why I’m waxing lyrical about intelligence and the mechanics of LLMs, it’s because I think they help decide two things: which problems AI is likely to be good at solving and whether (or when) it’s worth learning to code. After all, if LLMs can already code better and faster than you, why bother learning how to do it yourself? Particularly if the type of <em>intelligence</em> AI outputs can mimic have sufficient crossover with our own to make it practically indistinguishable in the realm of applied economics and public policy.</p>



<h3 class="wp-block-heading">Widening the scope, scale and quality(?) of analysis</h3>



<p class="wp-block-paragraph">Looking back at <a href="https://www.gilesd-j.com/2019/01/07/7-reasons-for-policy-professionals-to-get-pumped-about-r-programming-in-2019/" rel="nofollow" target="_blank">my original post</a>, almost all the cited benefits come from programming languages changing the scope, scale and quality of what you can produce as an applied policy analyst. And if I had to put a number on it: of the seven reasons cited in the listicle, six can be achieved without having to write the code yourself.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn4" id="_ftnref4" rel="nofollow" target="_blank">[4]</a> Have some messy data that needs to be cleaned and reshaped before analysis? Share a sample with AI and let it write the code to take care of the data cleaning for you. Have research question you want answered? LLMs can suggest alternative statistical approaches you might not have thought of and write the code to implement them. Need to produce a set of publication ready plots? AI can write the code to do this too in a format that tells the story of your choice.</p>



<p class="wp-block-paragraph">It would seem to be game over for learning to code and my 2019 post.</p>



<p class="wp-block-paragraph">Of course, you already know that I don’t believe this but let me tell you why…</p>



<h3 class="wp-block-heading">The jagged frontier of AI-assisted public policy analysis</h3>



<p class="wp-block-paragraph">Firstly, whether AI is <em>intelligent</em> or merely produces outputs that make it appear to be, early evidence suggests its ability to substitute for <em>human</em> intelligence varies across tasks. Researchers that tested the productivity benefits of AI termed this as a <em>jagged frontier,</em> that the users had to carefully navigate to benefit from using AI in their work.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn5" id="_ftnref5" rel="nofollow" target="_blank">[5]</a> It’s also not clear whether improving model performance is likely to expand the frontier equally across tasks, with many of the benchmark scores not necessarily generalizing outside test questions and being fragile to irrelevant and/or incorrect context.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn6" id="_ftnref6" rel="nofollow" target="_blank">[6]</a> Finally, at the time of writing this post, the types of tasks that frontier AI models had the hardest time completing successfully, looked strikingly similar to the problems encountered in public policy: messy, interdependent and heavily reliant on context.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn7" id="_ftnref7" rel="nofollow" target="_blank">[7]</a></p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-13.png?w=450&#038;ssl=1" alt="" class="wp-image-4709" srcset_temp="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-13.png?w=450&#038;ssl=1 501w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-13-295x300.png 295w" sizes="auto, (max-width: 501px) 100vw, 501px" data-recalc-dims="1" /><figcaption class="wp-element-caption"><strong>Source: </strong>Ethan Mollick, Centaurs and Cyborgs on the Jagged Frontier, One Useful Thing, <a href="https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the-jagged" rel="nofollow" target="_blank">link</a></figcaption></figure>



<p class="wp-block-paragraph">Secondly, while there’s solid evidence suggesting that AI <em>can</em> lead to productivity gains for knowledge workers, this isn’t automatic and appears to lean on the taste, judgement and experience of the user. <a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn8" id="_ftnref8" rel="nofollow" target="_blank">[8]</a> This was demonstrated by a field experiment in Kenya,<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn9" id="_ftnref9" rel="nofollow" target="_blank">[9]</a> where a little over 600 entrepreneurs were randomly provided access to an AI-powered business assistant. The performance of those with the assistant was then tracked over time and compared to those without the AI assistant. The result: top performers gained around 15 percent from having access to AI, while weaker ones lost around 10. Entrepreneurs that benefited the most from AI appeared to be better at identifying and acting on contextually appropriate advice, while poor performers tended to adopt generic strategies proposed by the tool.</p>



<p class="wp-block-paragraph">A much smaller experimental study closer to the world of coding pointed to a similar idea,  with both the 2025 (n=16) and 2026 (n=57) replication suggesting the productivity returns to using AI are ambiguous <em>and</em> vary based on the experience of the user.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn10" id="_ftnref10" rel="nofollow" target="_blank">[10]</a> Both versions of the study took the same basic form: experienced developers were randomly assigned to work with AI or not on their own repositories and the performance benefits were compared between groups. In the 2025 study, the authors found the performance benefits were <em>on average</em> negative: experienced developers took longer to complete the same tasks and overestimated the time saved.  In the 2026 replication,<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn11" id="_ftnref11" rel="nofollow" target="_blank">[11]</a> the authors found performance increased on average, but not for all developers.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn12" id="_ftnref12" rel="nofollow" target="_blank">[12]</a> Although the authors note their estimates are likely biased downwards due to difficulties in recruiting developers willing not to use AI, they paint the same basic picture: the benefits of using AI are not unambiguously positive and depend on the skills and experience of the user.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn13" id="_ftnref13" rel="nofollow" target="_blank">[13]</a>  </p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“…The right analogy for AI is not humans, but an alien intelligence with a distinct set of capabilities and limitations. Just because it exceeds human ability at one task doesn’t mean it can do all related work at human level. Although AIs and humans can perform some similar tasks, the underlying “cognitive” processes are fundamentally different…”</p>
<cite>Mollick, E. (2024, May 12). Superhuman? What does it mean for AI to be better than a human? And how can we tell? One Useful Thing (<a href="https://www.oneusefulthing.org/p/superhuman" rel="nofollow" target="_blank">link</a>).</cite></blockquote>



<h3 class="wp-block-heading">There’s no accounting for taste</h3>



<p class="wp-block-paragraph">To be clear, my point here isn’t to suggest AI tools aren’t helpful, just that the size of the benefit depends on the judgement, expertise and taste of the user. AI certainly can suggest a variety of statistical techniques for answering your research question, but are they the right ones?<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn14" id="_ftnref14" rel="nofollow" target="_blank">[14]</a> Sure, AI can quickly produce code for twenty reasonable looking plots, but do you understand the precise changes needed to accurately present your results and apply your corporate style guidelines and colours? Certainly, AI might simplify data cleaning, but does the approach make the right assumptions given the context, your research question and how the data was collected? And, while it might be true that research built on AI-generated code is more reproducible than a spreadsheet with a tangled collection of formulas, it may just be encoding hidden errors, the wrong methodology and creating a new source of technical debt you and your team will have to pay in the future.</p>



<p class="wp-block-paragraph">In my original post I was excited at programming expanding the scale, scope and quality of analysis I could do. AI can help with the first two, but its influence on the latter seems to depend on the user. AI may be able to slap together a solid draft that gets you 90 percent of the way, which for some tasks is all you need, but deciding what good looks like <em>and</em> instructing AI how to help relies on building sufficient judgement, taste and expertise. </p>



<h3 class="wp-block-heading">Vibing fast and slow</h3>



<p class="wp-block-paragraph">The need to develop judgement, expertise and professional taste is also not unique to coding. But like musicians with sheet music, a mathematician with an equation and an economist with a set of national account statistics, anyone working with code is likely to perform better with some understanding of how it works, even with the help of AI. In <a href="https://www.linkedin.com/posts/andrewyng_some-people-today-are-discouraging-others-share-7305984835037118464-RPIX/" rel="nofollow" target="_blank">the words of Andrew Ng</a>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph"><em>“…people who understand the language of software through their knowledge of coding can tell an LLM or an AI-enabled IDE what they want much more precisely, and get much better results.”</em></p>
</blockquote>



<p class="wp-block-paragraph">And once again, there’s research that backs this up. For instance, pre-AI a meta-analysis of the effects of learning to code points to benefits to creativity and metacognition i.e. <em>the processes underlying the monitoring, adaptation, evaluation, and planning of thinking and behaviour,</em><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn15" id="_ftnref15" rel="nofollow" target="_blank">[15]</a> both which are plausibly linked to the ability to properly prompt and manage AI to perform tasks. A small observational study<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn16" id="_ftnref16" rel="nofollow" target="_blank">[16]</a>  (n=100) seems to support this idea, with their analysis finding a positive association between vibe-coding efficacy and computer science (CS) domain expertise. Although the exact questions used to evaluate CS expertise aren’t explicitly outlined, their cited source points to them covering the same thinking and skills that learning to program helps to build,<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn17" id="_ftnref17" rel="nofollow" target="_blank">[17]</a> with the authors concluding that effectively <em>vibe-coding likely draws on problem decomposition and algorithmic thinking skills</em>.</p>



<p class="wp-block-paragraph">The book <em>Messy Jobs</em> also makes this point by noting that <em>by cheapening the cost of cognition, AI simultaneously raises the value of judgement and taste.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn18" id="_ftnref18" rel="nofollow" target="_blank"><strong>[18]</strong></a></em> The consequence of this for vibe-coding is that both instructing AI how to solve a problem and being able to judge the quality of outputs relies on knowing how code works, otherwise you might just be doing stupid things faster, in the words of Hadley Wickham:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">“…LLMs and AI are, I think, a massive amplifier. You can certainly use them to amplify the worst aspects of yourself, but you can also use them to amplify the best aspects…”</p>
<cite>Wickham, H. (2026, July 25). y code when ai? A talk on AI and coding. Tidy Design. <a href="https://tidydesign.substack.com/p/y-code-when-ai" rel="nofollow" target="_blank">https://tidydesign.substack.com/p/y-code-when-ai</a></cite></blockquote>



<h3 class="wp-block-heading">Summing up</h3>



<p class="wp-block-paragraph">In my 2019 listicle, I argued that learning to code was worth it so you can answer a wider variety of questions more quickly and at a higher standard to traditional point-and-click software. Seven years later and AI tools can produce code for a wide variety of tasks more quickly than an experienced developer. But that doesn’t make what it produces <em>good</em>. </p>



<p class="wp-block-paragraph">None of this means that I think you should avoid using AI before having mastered how to code, just that to be a force-multiplier it helps to know what good code looks like. Treat it as a brainy intern: motivated, well read, and in need of supervision. Give it clear instructions and it might come back with something useful. Don’t and its code might set the office microwave on fire. Learning to code can help you tell the difference between the two. </p>



<p class="wp-block-paragraph">That matters because you’re accountable for the result as the human given the the instructions. Use AI to decide how to prioritize which grants to cancel and you might cause unnecessary harm,<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn19" id="_ftnref19" rel="nofollow" target="_blank">[19]</a> trust its recommendations without validating them and you might just erroneously reinforce cultural stereotypes<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn20" id="_ftnref20" rel="nofollow" target="_blank">[20]</a>, and have it author a report for you, and it might just cite hallucinated sources.<a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftn21" id="_ftnref21" rel="nofollow" target="_blank">[21]</a> If used well, AI can help you reach your destination quicker and expand your professional frontier. If used poorly, it might just having you producing more errors at scale.</p>



<p class="wp-block-paragraph">I don’t know how long this holds. The studies here are small and not focused directly on the tasks we do in public policy. My best guess is that the jagged frontier will continue to expand, but unevenly, which will continue to place a premium on professional expertise and taste for tasks AI continues to struggle with. </p>



<p class="wp-block-paragraph">So: should you still learn to code in 202X? Yes, but not for the reason I gave in 2019. Learn to code so you can tell good code from bad and steer the tool in the right direction. </p>



<hr class="wp-block-separator has-alpha-channel-opacity"/>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref1" id="_ftn1" rel="nofollow" target="_blank">[1]</a>The <a href="https://gmcgoldr.github.io/2026/09/04/llm-next-token-predictors.html" rel="nofollow" target="_blank">push-back I’ve seen to this description</a> of how LLMs work doesn’t change the core point. LLMs are trained to probabilistically predict the next token to produce intelligent-looking outputs at scale.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref2" id="_ftn2" rel="nofollow" target="_blank">[2]</a> LLMs are trained on the finished product of human intelligence, such as published research, books and social media posts, but not the thinking that went into producing them. Newer models that ‘think’ before responding to a prompt are meant as a partial workaround to this, but don’t change the basic point that LLMs are trained to mimic <em>articulated</em> <em>intelligence</em>, rather than the process of producing. </p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref3" id="_ftn3" rel="nofollow" target="_blank">[3]</a> A practice that might have the opposite result if you’re not careful, see: Goldsmith-Pinkham, P. (2026, September 10). How do we know if writing is AI? Some initial analyses of AI’s impact on economics writing. A Causal Affair (<a href="https://paulgp.substack.com/p/how-do-we-know-if-writing-is-ai" rel="nofollow" target="_blank">link</a>) </p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref4" id="_ftn4" rel="nofollow" target="_blank">[4]</a> The one that may no longer be true is the trend towards open-source statistics. Certainly, I’ve been experimenting with ways to use AI to better extract government data (<a href="https://policyanalysislab.com/shared/general/250312_mvp/media/260625_MVP_Demo.mp4" rel="nofollow" target="_blank">see example</a>), but it’s entirely possible LLMs will make the internet a more difficult resource to navigate as text can be generated at larger scales.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref5" id="_ftn5" rel="nofollow" target="_blank">[5]</a> Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K.C., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K.R., 2026. Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37(2), pp.403-423.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref6" id="_ftn6" rel="nofollow" target="_blank">[6]</a> Akimitsu, P., 2026. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam. arXiv preprint arXiv:2607.23424.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref7" id="_ftn7" rel="nofollow" target="_blank">[7]</a> METR. (2026, May). Task-completion time horizons of frontier AI models. <a href="https://metr.org/time-horizons/" rel="nofollow" target="_blank">https://metr.org/time-horizons/</a></p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref8" id="_ftn8" rel="nofollow" target="_blank">[8]</a> Ethan Mollick uses the term <em>Centaurs</em> and <em>Cyborgs</em> todescribe the two most common approaches used by  workers that were able to navigate this frontier. In both cases, using AI effectively requires intelligently and selectively integrating the tools into their work without ‘falling asleep at the wheel’ (<a href="https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the-jagged" rel="nofollow" target="_blank">link</a>).</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref9" id="_ftn9" rel="nofollow" target="_blank">[9]</a> Otis, N.G., Clarke, R., Delecourt, S., Holtz, D. and Koning, R., 2023. The Uneven Impact of Generative AI on Entrepreneurial Performance: Evidence from a Field Experiment in Kenya (No. 24-042). Harvard Business School Working Paper.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref10" id="_ftn10" rel="nofollow" target="_blank">[10]</a> Becker, J., Rush, N., Barnes, B., &#038; Rein, D. (2025, July 10). Measuring the impact of early-2025 AI on experienced open-source developer productivity. METR. <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" rel="nofollow" target="_blank">https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/</a></p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref11" id="_ftn11" rel="nofollow" target="_blank">[11]</a> Becker, J., Rush, N., Cunningham, T., Rein, D., &#038; Mahamud, K. (2026, February 24). We are changing our developer productivity experiment design. METR. <a href="https://metr.org/blog/2026-02-24-uplift-update/" rel="nofollow" target="_blank">https://metr.org/blog/2026-02-24-uplift-update/</a></p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref12" id="_ftn12" rel="nofollow" target="_blank">[12]</a> The authors note it had become increasingly in their 2026 follow-up study to recruit developers willing to work without AI, which likely resulted in estimates of AI-assisted speedup being biased downward.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref13" id="_ftn13" rel="nofollow" target="_blank">[13]</a> Although the authors recommend their results are interpreted with caution, the 2025 study not only found that developers using AI were slower, but that more experienced developers tended to overestimate the time they would save from using AI.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref14" id="_ftn14" rel="nofollow" target="_blank">[14]</a>Like humans, LLMs exhibit systematic biases in the way they work and evaluate problems. This might come in subtle forms that influence outputs subtly, such as assuming doctors are more likely to be men than women. Or explicit, such as producing outputs that reflect how prominent a solution is in its training set, rather than necessarily being appropriate for the task at hand (<a href="https://arxiv.org/html/2411.10915v1#S3" rel="nofollow" target="_blank">link</a>).  </p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref15" id="_ftn15" rel="nofollow" target="_blank">[15]</a> Scherer, R., Siddiq, F. and Sánchez Viveros, B., 2019. The cognitive benefits of learning computer programming: A meta-analysis of transfer effects. Journal of Educational Psychology, 111(5), p.764.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref16" id="_ftn16" rel="nofollow" target="_blank">[16]</a> Thorgeirsson, S., Weidmann, T.B. and Su, Z., 2026, April. Computer science achievement and writing skills predict vibe coding proficiency. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1-17).</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref17" id="_ftn17" rel="nofollow" target="_blank">[17]</a> Specifically, code tracing, code completion, conditionals, loops and logical operators, see: Parker, M.C., Solomon, A., Pritchett, B., Illingworth, D.A., Marguilieux, L.E. and Guzdial, M., 2018, August. Socioeconomic status and computer science achievement: Spatial ability as a mediating variable in a novel model of understanding. In Proceedings of the 2018 ACM Conference on International Computing Education Research (pp. 97-105).</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref18" id="_ftn18" rel="nofollow" target="_blank">[18]</a> Garicano, L, Li, J &#038; Wu, Y 2026, Messy Jobs: The Work That AI Cannot Reach, Upriver Press</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref19" id="_ftn19" rel="nofollow" target="_blank">[19]</a> Alibašić, H., 2026. Introduction: The Imperative of Hybrid Intelligence in Digital Governance. In Hybrid Intelligence for Effective Digital Governance: AI in Administration (pp. 3-48). Cham: Springer Nature Switzerland.</p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref20" id="_ftn20" rel="nofollow" target="_blank">[20]</a>Saar Alon-Barkat, Madalina Busuioc, Human–AI Interactions in Public Sector Decision Making: “Automation Bias” and “Selective Adherence” to Algorithmic Advice, Journal of Public Administration Research and Theory, Volume 33, Issue 1, January 2023, Pages 153–169, <a href="https://doi.org/10.1093/jopart/muac007" rel="nofollow" target="_blank">https://doi.org/10.1093/jopart/muac007</a> </p>



<p class="wp-block-paragraph"><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/#_ftnref21" id="_ftn21" rel="nofollow" target="_blank">[21]</a> Karp, P. (2025, November 6). AI-tainted Deloitte report was worse than previously thought. Australian Financial Review. <a href="https://www.afr.com/politics/ai-tainted-deloitte-report-was-worse-than-previously-thought-20251106-p5n863" rel="nofollow" target="_blank">https://www.afr.com/politics/ai-tainted-deloitte-report-was-worse-than-previously-thought-20251106-p5n863</a></p>
<p>The post <a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/" rel="nofollow" target="_blank">X Reasons for policy professionals to get into programming in 202X</a> appeared first on <a href="https://www.gilesd-j.com/" rel="nofollow" target="_blank">Giles</a>.</p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://www.gilesd-j.com/2026/09/28/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/"> Data Analytics and AI Archives - Giles</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/x-reasons-for-policy-professionals-to-get-into-programming-in-202x/">X Reasons for policy professionals to get into programming in 202X</a>]]></content:encoded>
					
		
		<enclosure url="https://policyanalysislab.com/shared/general/250312_mvp/media/260625_MVP_Demo.mp4" length="64065940" type="video/mp4" />

		<post-id xmlns="com-wordpress:feed-additions:1">403955</post-id>	</item>
		<item>
		<title>Adding a free AI chatbot to a blogdown website</title>
		<link>https://www.r-bloggers.com/2026/09/adding-a-free-ai-chatbot-to-a-blogdown-website/</link>
		
		<dc:creator><![CDATA[R-bloggers on Almost Random]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://www.malte-grosser.com/post/adding-a-free-ai-chatbot/</guid>

					<description><![CDATA[<p>I wanted to add a small “Ask AI” window to this website. Visitors should be able to ask about my projects, publications or background and get a short answer with a useful link. Running it should cost nothing.<br />
I described my website setup to ChatGPT, fo...</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/adding-a-free-ai-chatbot-to-a-blogdown-website/">Adding a free AI chatbot to a blogdown website</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://www.malte-grosser.com/post/adding-a-free-ai-chatbot/"> R-bloggers on Almost Random</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
I wanted to add a small “Ask AI” window to this website. Visitors should be able to ask about my projects, publications or background and get a short answer with a useful link. Running it should cost nothing.
I described my website setup to ChatGPT, followed its suggestions and discussed the parts I wanted to change. That covered choosing a provider, writing the code and adjusting the window for desktop and mobile.
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://www.malte-grosser.com/post/adding-a-free-ai-chatbot/"> R-bloggers on Almost Random</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/adding-a-free-ai-chatbot-to-a-blogdown-website/">Adding a free AI chatbot to a blogdown website</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403962</post-id>	</item>
		<item>
		<title>rOpenSci News Digest, September 2026</title>
		<link>https://www.r-bloggers.com/2026/09/ropensci-news-digest-september-2026/</link>
		
		<dc:creator><![CDATA[rOpenSci]]></dc:creator>
		<pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://ropensci.org/blog/2026/09/28/news-september-2026/</guid>

					<description><![CDATA[<p>Dear rOpenSci friends, it’s time for our monthly news roundup!  You can read this post on our blog. Now let’s dive into the activity at and around rOpenSci!</p>
<p>rOpenSci HQ</p>
<p>Yani and Mark at “Open Communities in the age of AI”<br />
rO...</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/ropensci-news-digest-september-2026/">rOpenSci News Digest, September 2026</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://ropensci.org/blog/2026/09/28/news-september-2026/"> rOpenSci - open tools for open science</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<!-- Before sending DELETE THE INDEX_CACHE and re-knit! -->
<p>Dear rOpenSci friends, it’s time for our monthly news roundup! <!-- blabla --> You can read this post <a href="https://ropensci.org/blog/2026/09/28/news-september-2026" rel="nofollow" target="_blank">on our blog</a>. Now let’s dive into the activity at and around rOpenSci!</p>
<h2>
rOpenSci HQ
</h2><h3>
Yani and Mark at “Open Communities in the age of AI”
</h3><p>rOpenSci community manager <a href="https://ropensci.org/author/yanina-bellini-saibene" rel="nofollow" target="_blank">Yani</a> and software-review lead <a href="https://ropensci.org/author/mark-padgham" rel="nofollow" target="_blank">Mark</a> will participate in the upcoming Openscapes community call on October 1st (this Thursday!) at 9:30AM PT (16:30 UTC) entitled <em>“Open Communities in the Age of AI”</em>, together with Mara Averick, Senior Developer Advocate at Quansight and Hadley Wickham, Chief Scientist at Posit.</p>
<p><a href="https://openscapes.org/events/2026-10-01-community-call-open-communities-ai/" rel="nofollow" target="_blank">Event page</a>, including link for free registration.</p>
<h3>
Champions Program update
</h3><p>Our current 2026–2027 cohort has completed the training phase and is now focused on developing their individual projects with the support of their mentors, as well as on their outreach activities.</p>
<p>Meanwhile, the 2025–2026 cohort will wrap up their journey with a closing community call, <a href="https://ropensci.org/commcalls/mas-alla-del-codigo-2026/" rel="nofollow" target="_blank"><strong>Más Allá del Código: muestra abierta de los proyectos de nuestros campeon(a|e)s (Beyond the Code: an open showcase of our Champions’ Projects)</strong></a>. At the call, the Champions will present their final projects, from creating, improving, and reviewing R packages to the outreach work that brought those projects to their communities. They’ll also share what their projects led to, including scholarships, new collaborations, and professional growth. The event is in Spanish. The call will take place on <a href="https://ropensci.org/commcalls/mas-alla-del-codigo-2026/" rel="nofollow" target="_blank">Monday, October 12 at 15:00 UTC</a>. Come get inspired, ask your questions live, and help us celebrate our growing open-source community!</p>
<h2>
We’re still celebrating our 15th anniversary! <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f389.png" alt="🎉" class="wp-smiley" style="height: 1em; max-height: 1em;" />
</h2><p>In July, we started to share stories from members of our community about their experiences with rOpenSci. Our second story features <a href="https://ropensci.org/author/yi-chin-sunny-tseng/" rel="nofollow" target="_blank">Yi-Chin Sunny Tseng</a> and her connection with rOpenSci. Read it on our blog: <a href="https://ropensci.org/blog/2026/09/15/birthday-post-sunny/" rel="nofollow" target="_blank">Happy Birthday rOpenSci — My Journey from First-Time Developer to Current Opportunities</a> Stay tuned for more stories from our community as we continue celebrating 15 years of rOpenSci!</p>
<h2>
rOpenSci at LatinR 2026 in Medellín, Colombia. See you there!
</h2><p><a href="https://www.eventbrite.com.ar/e/1998690018649" rel="nofollow" target="_blank">Registration for LatinR 2026 (November 11–13, Universidad de Antioquia, Medellín) is now open</a>, and rOpenSci will have a strong presence.</p>
<p><a href="https://ropensci.org/author/jeroen-ooms" rel="nofollow" target="_blank">Jeroen</a> will give a talk on R-Universe. <a href="https://ropensci.org/author/yanina-bellini-saibene" rel="nofollow" target="_blank">Yani</a> will lead a workshop on R package development and give a talk on the Champions Program’s open curriculum. <a href="https://ropensci.org/author/nic-crane" rel="nofollow" target="_blank">Nic Crane</a> is one of the conference’s keynote speakers. <a href="https://ropensci.org/author/evelia-lorena-coss-navarrete/" rel="nofollow" target="_blank">Evelia Lorena Coss Navarrete</a>, one of our Champions, will talk about the package she developed during the program. Several other rOpenSci folks — <a href="https://ropensci.org/author/evelia-lorena-coss-navarrete/" rel="nofollow" target="_blank">Natalia Da Silva</a>, <a href="https://ropensci.org/author/luis-d.-verde-arregoitia/" rel="nofollow" target="_blank">Luis Verde</a>, and <a href="https://ropensci.org/author/francisco-cardozo/" rel="nofollow" target="_blank">Francisco Cardozo</a>, among others — will also be there. Jeroen, Francisco, and Yani will take part in the hackathon during the conference.</p>
<p>Join us in Medellín!</p>
<h3>
Coworking
</h3><p>Read <a href="https://ropensci.org/blog/2023/06/21/coworking/" rel="nofollow" target="_blank">all about coworking</a>!</p>
<ul>
<li>Tuesday October 6th, 09:00 Americas Pacific (16:00 UTC) <a href="https://ropensci.org/events/coworking-2026-10/" rel="nofollow" target="_blank">“Writing Tests &#038; Testing in R”</a>, with <a href="https://ropensci.org/author/yanina-bellini-saibene" rel="nofollow" target="_blank">Yanina Bellini Saibene</a> and co-host <a href="https://ropensci.org/author/olivier-leroy/" rel="nofollow" target="_blank">Olivier Leroy</a>.
<ul>
<li>Explore how to write tests for R and add some tests to your work or packages</li>
<li>Meet co-host, Olivier Leroy, and chat about testing</li>
</ul>
</li>
<li>Tuesday November 3rd, 09:00 Australia Western (01:00 UTC) <a href="https://ropensci.org/events/coworking-2026-11/" rel="nofollow" target="_blank">“Climate Science in R”</a>, with <a href="https://ropensci.org/author/steffi-lazerte" rel="nofollow" target="_blank">Steffi LaZerte</a> and co-host <a href="https://ropensci.org/author/elio-campitelli/" rel="nofollow" target="_blank">Elio Campitelli</a>.
<ul>
<li>Explore how R is used to study the climate</li>
<li>Meet co-host, Elio Campitelli, and discuss Climate Science in R</li>
</ul>
</li>
<li>Tuesday December 8th<sup>*</sup>, 14:00 Europe Central (12:00 UTC) <a href="https://ropensci.org/events/coworking-2026-12/" rel="nofollow" target="_blank">“Code Linting in R”</a>, with <a href="https://ropensci.org/author/steffi-lazerte" rel="nofollow" target="_blank">Steffi LaZerte</a> and co-host <a href="https://ropensci.org/author/etienne-bacher/" rel="nofollow" target="_blank">Etienne Bacher</a>.
<ul>
<li>Read up on Code Linting and apply some linters to your R code</li>
<li>Meet co-host, Etienne Bacher, and discuss code linting in general, or flir and Jarl in particular<br>
* Note that December coworking is a week later than usual</li>
</ul>
</li>
</ul>
<p>And remember, you can always cowork independently on work related to R, work on packages that tend to be neglected, or work on what ever you need to get done!</p>
<h2>
Software <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f4e6.png" alt="📦" class="wp-smiley" style="height: 1em; max-height: 1em;" />
</h2><p>The following two packages recently became a part of our software suite:</p>
<ul>
<li>
<p><a href="https://docs.ropensci.org/ciecl" rel="nofollow" target="_blank">ciecl</a>, developed by Rodolfo Tasso Suazo: Tools for working with the International Classification of Diseases (ICD-10 Chile official MINSAL/DEIS v2018). Includes optimized SQL search with SQLite, fuzzy matching of medical terms (Jaro-Winkler), Charlson and Elixhauser comorbidity calculation, WHO ICD-11 API integration, and hierarchical code validation. Data from Centro FIC Chile DEIS <a href="https://deis.minsal.cl/centrofic/" rel="nofollow" target="_blank">https://deis.minsal.cl/centrofic/</a>. It has been <a href="https://github.com/ropensci/software-review/issues/765" rel="nofollow" target="_blank">reviewed</a> by Maëlle Salmon and Yanina Bellini.</p>
</li>
<li>
<p><a href="https://docs.ropensci.org/brapiR2" rel="nofollow" target="_blank">brapiR2</a>, developed by Joash Joshua Ayo: Provides pipe-friendly, stateless read access to the Breeding API (BrAPI) v2.1 specification, an open community standard for plant breeding data interchange maintained by the BrAPI project <a href="https://brapi.org/" rel="nofollow" target="_blank">https://brapi.org</a>. Wraps 32 of the 37 BrAPI v2.1 entities across all four modules, Core, Germplasm, Phenotyping, and Genotyping, covering 49 of the specifications 138 retrieval (GET and search) endpoints and returning tidy tibbles ready for analysis. Write and update endpoints are out of scope by design. Features include automatic pagination, async search handling, response caching, parallel batch fetching, and convenience functions for genomic selection workflows (e.g. dosage matrix extraction). Designed for plant breeders and bioinformaticians who need programmatic access to plant breeding databases that implement the BrAPI’ v2 specification. It has been <a href="https://github.com/ropensci/software-review/issues/792" rel="nofollow" target="_blank">reviewed</a> by David Waring and Jenna Hershberger.</p>
</li>
</ul>
<p>Discover <a href="https://ropensci.org/packages" rel="nofollow" target="_blank">more packages</a>, read more about <a href="https://ropensci.org/software-review" rel="nofollow" target="_blank">Software Peer Review</a>.</p>
<h3>
New versions
</h3><p>The following twenty-two packages have had an update since the last newsletter: <a href="https://docs.ropensci.org/visdat" title="Preliminary Visualisation of Data" rel="nofollow" target="_blank">visdat</a> (<a href="https://github.com/ropensci/visdat/releases/tag/v0.6.1" rel="nofollow" target="_blank"><code>v0.6.1</code></a>), <a href="https://docs.ropensci.org/nycOpenData" title="A Lightweight Interface to NYC Open Data APIs" rel="nofollow" target="_blank">nycOpenData</a> (<a href="https://github.com/ropensci/nycOpenData/releases/tag/v0.2.3" rel="nofollow" target="_blank"><code>v0.2.3</code></a>), <a href="https://docs.ropensci.org/ruODK" title="An R Client for the ODK Central API" rel="nofollow" target="_blank">ruODK</a> (<a href="https://github.com/ropensci/ruODK/releases/tag/v1.6.0" rel="nofollow" target="_blank"><code>v1.6.0</code></a>), <a href="https://docs.ropensci.org/stats19" title="Work with Open Road Traffic Casualty Data from Great Britain" rel="nofollow" target="_blank">stats19</a> (<a href="https://github.com/ropensci/stats19/releases/tag/v4.1.0" rel="nofollow" target="_blank"><code>v4.1.0</code></a>), <a href="https://docs.ropensci.org/npi" title="Access the U.S. National Provider Identifier Registry API" rel="nofollow" target="_blank">npi</a> (<a href="https://github.com/ropensci/npi/releases/tag/v0.3.1" rel="nofollow" target="_blank"><code>v0.3.1</code></a>), <a href="https://docs.ropensci.org/promoutils" title="Utilities for Promoting rOpenSci" rel="nofollow" target="_blank">promoutils</a> (<a href="https://github.com/ropensci-org/promoutils/releases/tag/v0.7.0" rel="nofollow" target="_blank"><code>v0.7.0</code></a>), <a href="https://docs.ropensci.org/brapiR2" title="A Tidyverse-Native Client for the BrAPI v2 (Breeding API) Specification" rel="nofollow" target="_blank">brapiR2</a> (<a href="https://github.com/ropensci/brapiR2/releases/tag/v0.2.0" rel="nofollow" target="_blank"><code>v0.2.0</code></a>), <a href="https://docs.ropensci.org/ciecl" title="International Classification of Diseases ICD-10/ICD-11 for Chile" rel="nofollow" target="_blank">ciecl</a> (<a href="https://github.com/ropensci/ciecl/releases/tag/v1.0.0" rel="nofollow" target="_blank"><code>v1.0.0</code></a>), <a href="https://docs.ropensci.org/bibtex" title="Bibtex Parser" rel="nofollow" target="_blank">bibtex</a> (<a href="https://github.com/ropensci/bibtex/releases/tag/v0.5.3" rel="nofollow" target="_blank"><code>v0.5.3</code></a>), <a href="https://docs.ropensci.org/ckanr" title="Client for the Comprehensive Knowledge Archive Network (CKAN) API" rel="nofollow" target="_blank">ckanr</a> (<a href="https://github.com/ropensci/ckanr/releases/tag/v0.9.0" rel="nofollow" target="_blank"><code>v0.9.0</code></a>), <a href="https://docs.ropensci.org/osmapiR" title="OpenStreetMap API" rel="nofollow" target="_blank">osmapiR</a> (<a href="https://github.com/ropensci/osmapiR/releases/tag/v0.2.6" rel="nofollow" target="_blank"><code>v0.2.6</code></a>), <a href="https://docs.ropensci.org/distionary" title="Create and Evaluate Probability Distributions" rel="nofollow" target="_blank">distionary</a> (<a href="https://github.com/probaverse/distionary/releases/tag/v0.2.0" rel="nofollow" target="_blank"><code>v0.2.0</code></a>), <a href="https://docs.ropensci.org/GLMMcosinor" title="Fit a Cosinor Model Using a Generalized Mixed Modeling Framework" rel="nofollow" target="_blank">GLMMcosinor</a> (<a href="https://github.com/ropensci/GLMMcosinor/releases/tag/v0.2.2" rel="nofollow" target="_blank"><code>v0.2.2</code></a>), <a href="https://docs.ropensci.org/reviser" title="Analyzing Revisions in Real-Time Time Series Vintages" rel="nofollow" target="_blank">reviser</a> (<a href="https://github.com/ropensci/reviser/releases/tag/v0.3.1" rel="nofollow" target="_blank"><code>v0.3.1</code></a>), <a href="https://docs.ropensci.org/autotest" title="Automatic Package Testing" rel="nofollow" target="_blank">autotest</a> (<a href="https://github.com/ropensci-review-tools/autotest/releases/tag/v0.2" rel="nofollow" target="_blank"><code>v0.2</code></a>), <a href="https://docs.ropensci.org/occCite" title="Querying and Managing Large Biodiversity Occurrence Datasets" rel="nofollow" target="_blank">occCite</a> (<a href="https://github.com/ropensci/occCite/releases/tag/v0.6.3" rel="nofollow" target="_blank"><code>v0.6.3</code></a>), <a href="https://docs.ropensci.org/readODS" title="Read and Write ODS Files" rel="nofollow" target="_blank">readODS</a> (<a href="https://github.com/ropensci/readODS/releases/tag/v2.3.6" rel="nofollow" target="_blank"><code>v2.3.6</code></a>), <a href="https://docs.ropensci.org/osmdata" title="Import OpenStreetMap Data as Simple Features or Spatial Objects" rel="nofollow" target="_blank">osmdata</a> (<a href="https://github.com/ropensci/osmdata/releases/tag/v0.4.1" rel="nofollow" target="_blank"><code>v0.4.1</code></a>), <a href="https://docs.ropensci.org/goodpractice" title="Advice on R Package Building" rel="nofollow" target="_blank">goodpractice</a> (<a href="https://github.com/ropensci-review-tools/goodpractice/releases/tag/v1.2.0" rel="nofollow" target="_blank"><code>v1.2.0</code></a>), <a href="https://docs.ropensci.org/galamm" title="Generalized Additive Latent and Mixed Models" rel="nofollow" target="_blank">galamm</a> (<a href="https://github.com/ropensci/galamm/releases/tag/v0.4.1" rel="nofollow" target="_blank"><code>v0.4.1</code></a>), <a href="https://docs.ropensci.org/RAMEN" title="Regional Association of Methylome variability with the Exposome and geNome" rel="nofollow" target="_blank">RAMEN</a> (<a href="https://github.com/ropensci/RAMEN/releases/tag/v2.1.2" rel="nofollow" target="_blank"><code>v2.1.2</code></a>), and <a href="https://docs.ropensci.org/EDIutils" title="An API Client for the Environmental Data Initiative Repository" rel="nofollow" target="_blank">EDIutils</a> (<a href="https://github.com/ropensci/EDIutils/releases/tag/v3.0.1" rel="nofollow" target="_blank"><code>v3.0.1</code></a>).</p>
<h2>
Software Peer Review
</h2><p>There are seventeen recently closed and active submissions and 4 submissions on hold. Issues are at different stages:</p>
<ul>
<li>
<p>Two at <a href="https://github.com/ropensci/software-review/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3A%226/approved%22" rel="nofollow" target="_blank">‘6/approved’</a>:</p>
<ul>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/792" rel="nofollow" target="_blank">brapiR2</a>, A Tidyverse-Native Client for the BrAPI v2 (Breeding API) Specification. Submitted by <a href="https://orcid.org/0009-0007-1642-0172" rel="nofollow" target="_blank">Joash Joshua Ayo</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/765" rel="nofollow" target="_blank">ciecl</a>, International Classification of Diseases ICD-10/ICD-11 for Chile. Submitted by <a href="https://github.com/Rodotasso" rel="nofollow" target="_blank">Rodolfo Tasso</a>.</p>
</li>
</ul>
</li>
<li>
<p>Three at <a href="https://github.com/ropensci/software-review/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3A%225/awaiting-reviewer(s)-response%22" rel="nofollow" target="_blank">‘5/awaiting-reviewer(s)-response’</a>:</p>
<ul>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/787" rel="nofollow" target="_blank">ibger</a>, Access the IBGE Aggregate Data API from R. Submitted by <a href="https://castlab.org/" rel="nofollow" target="_blank">Andre Leite Wanderley</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/744" rel="nofollow" target="_blank">RAQSAPI</a>, A Simple Interface to the US EPA Air Quality System Data Mart API. Submitted by <a href="https://github.com/mccroweyclinton-EPA" rel="nofollow" target="_blank">mccroweyclinton-EPA</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/717" rel="nofollow" target="_blank">coevolve</a>, Fit Bayesian Generalized Dynamic Phylogenetic Models using Stan. Submitted by <a href="https://scottclaessens.github.io/" rel="nofollow" target="_blank">Scott Claessens</a>. (Stats).</p>
</li>
</ul>
</li>
<li>
<p>One at <a href="https://github.com/ropensci/software-review/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3A%224/review(s)-in-awaiting-changes%22" rel="nofollow" target="_blank">‘4/review(s)-in-awaiting-changes’</a>:</p>
<ul>
<li><a href="https://github.com/ropensci/software-review/issues/718" rel="nofollow" target="_blank">rcrisp</a>, Automate the Delineation of Urban River Spaces. Submitted by <a href="https://github.com/cforgaci" rel="nofollow" target="_blank">Claudiu Forgaci</a>. (Stats).</li>
</ul>
</li>
<li>
<p>Four at <a href="https://github.com/ropensci/software-review/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3A%223/reviewer(s)-assigned%22" rel="nofollow" target="_blank">‘3/reviewer(s)-assigned’</a>:</p>
<ul>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/799" rel="nofollow" target="_blank">camtrapReport</a>, Camera-Trap Report Generator. Submitted by <a href="https://www.wur.nl/en/persons/e-elham-ebrahimi" rel="nofollow" target="_blank">Elham Ebrahimi</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/785" rel="nofollow" target="_blank">nert</a>, Curated Access to TERN Environmental Raster Data. Submitted by <a href="https://scholar.google.com.au/citations?user=zG1uKrcAAAAJ&#038;hl=en" rel="nofollow" target="_blank">Max Moldovan</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/769" rel="nofollow" target="_blank">rfastlowess</a>, High-Performance LOWESS Smoothing for R. Submitted by <a href="https://github.com/thisisamirv" rel="nofollow" target="_blank">Amir Valizadeh</a>. (Stats).</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/799" rel="nofollow" target="_blank">camtrapReport</a>, Camera-Trap Report Generator. Submitted by <a href="https://www.wur.nl/en/persons/e-elham-ebrahimi" rel="nofollow" target="_blank">Elham Ebrahimi</a>.</p>
</li>
</ul>
</li>
<li>
<p>Two at <a href="https://github.com/ropensci/software-review/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3A%222/seeking-reviewer(s)%22" rel="nofollow" target="_blank">‘2/seeking-reviewer(s)’</a>:</p>
<ul>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/775" rel="nofollow" target="_blank">grumpy</a>, Read NumPy .npy and .npz Files. Submitted by <a href="https://hugogruson.fr/" rel="nofollow" target="_blank">Hugo Gruson</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/774" rel="nofollow" target="_blank">tezr</a>, Access Thesis Metadata from Turkiye’s National Thesis Center. Submitted by <a href="https://emraher.com/" rel="nofollow" target="_blank">Emrah Er</a>.</p>
</li>
</ul>
</li>
<li>
<p>Five at <a href="https://github.com/ropensci/software-review/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3A%221/editor-checks%22" rel="nofollow" target="_blank">‘1/editor-checks’</a>:</p>
<ul>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/811" rel="nofollow" target="_blank">ocean3d</a>, Three-Dimensional Marine Spatial Analysis. Submitted by <a href="https://jmatsushiba.com/" rel="nofollow" target="_blank">Jay Matsushiba</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/808" rel="nofollow" target="_blank">corila</a>, Sparse modelling with grouped and correlated features allowing for privileged information. Submitted by <a href="https://rauschenberger.github.io/" rel="nofollow" target="_blank">Armin Rauschenberger</a>. (Stats).</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/805" rel="nofollow" target="_blank">tarpolyglot</a>, Run Python, Julia, Rust, and C++ Inside targets Pipeline Steps. Submitted by <a href="https://github.com/Pierre9344" rel="nofollow" target="_blank">Pierre9344</a>.</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/777" rel="nofollow" target="_blank">OptSurvCutR</a>, Optimal Survival Cut-Point Discovery for Time-to-Event Analysis with OptSurvCutR. Submitted by <a href="https://github.com/paytonyau" rel="nofollow" target="_blank">Payton Yau</a>. (Stats).</p>
</li>
<li>
<p><a href="https://github.com/ropensci/software-review/issues/766" rel="nofollow" target="_blank">HydraR</a>, Stateful Agentic Orchestration for Scientific Reproducibility. Submitted by <a href="https://www.mq.edu.au/research/research-centres-groups-and-facilities/facilities/australian-proteome-analysis-facility" rel="nofollow" target="_blank">Ignatius Pang</a>.</p>
</li>
</ul>
</li>
</ul>
<p>Find out more about <a href="https://ropensci.org/software-review" rel="nofollow" target="_blank">Software Peer Review</a> and how to get involved.</p>
<h2>
On the blog
</h2><!-- Do not forget to rebase your branch! -->
<ul>
<li><a href="https://ropensci.org/blog/2026/09/15/birthday-post-sunny" rel="nofollow" target="_blank">Happy Birthday rOpenSci — My Journey from First-Time Developer to Current Opportunities</a> by Yi-Chin Sunny Tseng. Yi-Chin Sunny Tseng shares how the rOpenSci Champions Program helped her grow from a first-time R package developer into an open science contributor, creating tools for biodiversity research and discovering new opportunities along the way.</li>
</ul>
<h2>
Calls for contributions
</h2><h3>
Calls for maintainers
</h3><p>If you’re interested in maintaining any of the R packages below, you might enjoy reading our blog post <a href="https://ropensci.org/blog/2023/02/07/what-does-it-mean-to-maintain-a-package/" rel="nofollow" target="_blank">What Does It Mean to Maintain a Package?</a>.</p>
<ul>
<li><a href="https://docs.ropensci.org/charlatan" rel="nofollow" target="_blank">charlatan</a>, create fake data in R. <a href="https://github.com/ropensci/charlatan/issues/150" rel="nofollow" target="_blank">Issue for volunteering</a>.</li>
</ul>
<h3>
Calls for contributions
</h3><p>Refer to our <a href="https://ropensci.org/help-wanted/" rel="nofollow" target="_blank">help wanted page</a> – before opening a PR, we recommend asking in the issue whether help is still needed.</p>
<h2>
Package development corner
</h2><p>Some useful information for R package developers. <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f440.png" alt="👀" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>
<h3>
{covr2gh}: Coverage summary as GitHub comment
</h3><p>The covr2gh package by Dragoș Moldovan-Grünfeld provides an automated way to summarise the impact of a pull request (PR) on test coverage directly within GitHub. See it in action in a PR in the <a href="https://github.com/GSK-Biostatistics/tfrmt/pull/844#issuecomment-5424645453" rel="nofollow" target="_blank">tfrmt R package</a>.</p>
<h3>
Deprecation messages for package data
</h3><p>Hugo Gruson wrote an <a href="https://hugogruson.fr/posts/deprecation-pkg-data/" rel="nofollow" target="_blank">exhaustive post</a> about the deprecation of <em>data</em> in an R package. The post features the <a href="https://rdrr.io/r/base/delayedAssign.html" rel="nofollow" target="_blank"><code>delayedAssign()</code></a> function.</p>
<h3>
A refactoring story featuring people
</h3><p>Athanasia Mo Mowinckel published <a href="https://drmowinckels.io/blog/2026/atlases-as-functions/" rel="nofollow" target="_blank">“Why ggseg Atlases Became Function Calls”</a>, where she explains how she made data into objects into a package in order to allow re-exporting them. She furthermore tells how she got to that conclusion by trying out different solutions and discussing with other package developers in the rOpenSci Slack workspace.</p>
<h3>
usethis 3.2.2
</h3><p>The usethis package was updated on CRAN. The <a href="https://usethis.r-lib.org/news/index.html#usethis-322" rel="nofollow" target="_blank">changelog</a> features many quality of life improvements around Git and GitHub, and a new <code>use_readme_qmd()</code> function for drafting a README in the Quarto format.</p>
<h2>
Last words
</h2><p>Thanks for reading! If you want to get involved with rOpenSci, check out our <a href="https://contributing.ropensci.org/" rel="nofollow" target="_blank">Contributing Guide</a>. This guide will help direct you to the right place, whether you want to make code contributions, non-code contributions, or contribute in other ways such as through sharing use cases. You can also support our work through <a href="https://ropensci.org/donate" rel="nofollow" target="_blank">donations</a>.</p>
<p>If you haven’t subscribed to our newsletter yet, you can <a href="https://ropensci.org/news/" rel="nofollow" target="_blank">do so though our signup form</a>. Until it’s time for our next newsletter, you can keep in touch with us through our <a href="https://ropensci.org/" rel="nofollow" target="_blank">website</a>, <a href="https://hachyderm.io/@rOpenSci" rel="nofollow" target="_blank">Mastodon</a>, or <a href="https://www.linkedin.com/company/ropensci/" rel="nofollow" target="_blank">LinkedIn</a>. See you soon!</p>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://ropensci.org/blog/2026/09/28/news-september-2026/"> rOpenSci - open tools for open science</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/ropensci-news-digest-september-2026/">rOpenSci News Digest, September 2026</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403951</post-id>	</item>
		<item>
		<title>JAGS 5.0.0 is now available for macOS</title>
		<link>https://www.r-bloggers.com/2026/09/jags-5-0-0-is-now-available-for-macos/</link>
		
		<dc:creator><![CDATA[Martyn]]></dc:creator>
		<pubDate>Sun, 27 Sep 2026 17:41:49 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">http://martynplummer.wordpress.com/?p=2098</guid>

					<description><![CDATA[<p>This is a guest post by Matt Denwood. After a longer-than-anticipated delay, the macOS installer for JAGS 5.0.0 is now available via SourceForge. There are some big changes for how the official binaries of JAGS 5.x are installed on macOS, … Continue reading →</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/jags-5-0-0-is-now-available-for-macos/">JAGS 5.0.0 is now available for macOS</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://martynplummer.wordpress.com/2026/09/27/jags-5-0-0-is-now-available-for-macos/"> R – JAGS News</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<p class="wp-block-paragraph"><em>This is a guest post by Matt Denwood</em>.</p>



<p class="wp-block-paragraph">After a longer-than-anticipated delay, the macOS installer for JAGS 5.0.0 is now available via <a href="https://sourceforge.net/projects/mcmc-jags/files/JAGS/5.x/macOS/" rel="nofollow" target="_blank">SourceForge</a>. There are some big changes for how the official binaries of JAGS 5.x are installed on macOS, which is the topic covered by this news post.</p>



<h2 class="wp-block-heading"><strong>Multi-threading and BLAS linkage</strong></h2>



<p class="wp-block-paragraph">JAGS now has multi-threading support to allow parallel execution of independent chains, which uses Apple’s Grand Central Dispatch for the macOS binaries. This is the default build, as most users will want to benefit from the parallel computation. However, a single-threaded build is also provided as part of the installer, which has the same behaviour as JAGS 4.3.2. A third build is also provided within the installer, which uses single threading along with Netlib’s reference implementation of BLAS and LAPACK rather than Apple’s vecLib (Accelerate framework). By default, all three of these builds are installed so that the user can later switch between them;  however, the installer can be customised to deselect builds that are not needed.  For any users wishing to avoid the larger download and installation footprint (e.g. build/test machines), we have also provided single-build installers via SourceForge. For discussion on the difference between BLAS implementations, see the relevant section of the <a href="https://cran.r-project.org/bin/macosx/RMacOSX-FAQ.html" rel="nofollow" target="_blank">R for macOS FAQ</a>. It is also possible to build JAGS using OpenMP on macOS – this is not provided as part of the official macOS installer, but the necessary compilation instructions will be added to the macOS section of the JAGS installation manual in due course.</p>



<h2 class="wp-block-heading"><strong>Installation location</strong></h2>



<ol class="wp-block-list">
<li>Although there is no standardised location for installing system-level software on macOS, the <code>/opt</code> directory is now more frequently used (by e.g. homebrew and some elements of R) than <code>/usr/local</code>. Part of the reason for this is that installation under <code>/opt/jags</code> ensures a clean separation between software relating to JAGS and other software that you may have installed under other subdirectories of <code>/opt</code>.</li>



<li>The installer now contains multiple builds of JAGS as well as utilities that are specific to macOS (see below) – the versioned directory structure allows the different builds and utilities to coexist.</li>



<li>Having a versioned directory structure also allows different major versions of JAGS to be installed on the same system – we hope that this will ease the transition from JAGS 4.3.2 to JAGS 5.0.0 on CRAN’s build systems.</li>
</ol>



<p class="wp-block-paragraph">The installer also contains a build of JAGS 4.3.2 that installs under the same <code>/opt/jags</code> directory, with symbolic links to the older /usr/local directory that allow the CRAN build of rjags-4 to continue to function as expected. These symlinks also allow JAGS 5.x to be accessed from the Terminal (and by pkg-config) without updating the PATH variable, although we recommend adding the new directory to the PATH explicitly using (for example): <code>echo 'export PATH=&quot;/opt/jags/bin:$PATH&quot;' &gt;&gt; ~/.zshrc</code></p>



<h2 class="wp-block-heading"><strong>JAGS utilities</strong></h2>



<p class="wp-block-paragraph">As well as JAGS itself, the macOS installer provides the following three utilities:</p>



<ul class="wp-block-list">
<li><strong>jags-version</strong> – this allows the active build of JAGS to be switched, both for the command-line version and for rjags.  Switching between major versions of JAGS requires re-installation of the rjags package with matching version, as well as re-compilation of the runjags package. Switching between builds of JAGS with the same major version does not require re-installation of any R package (but may require R to be restarted, if the required packages are loaded).</li>



<li><strong>jags-uninstall</strong> – this allows for complete removal of any JAGS installations detected either under the new /opt/jags directory or the legacy /usr/local directory.</li>



<li><strong>pkgconf-lite</strong> – as macOS does not include pkg-config by default, this small utility (provided by pkgconf) is installed in order to facilitate compilation of rjags on vanilla macOS installations.</li>
</ul>



<p class="wp-block-paragraph">See the provided man pages for jags-version and jags-uninstall for more information on these utilities.  Note that these utilities are <strong>not</strong> installed if using the single-build installers.</p>



<h2 class="wp-block-heading"><strong>Getting help</strong></h2>



<p class="wp-block-paragraph">If you encounter any issues with the macOS installer or utilities, please either post to the macOS-specific JAGS installation thread on the <a href="https://sourceforge.net/p/mcmc-jags/discussion/610037/thread/115edca29d" rel="nofollow" target="_blank">JAGS discussion forum</a>. If you have comments or suggestions regarding the macOS utilities or build process for JAGS then please feel free to open an issue on the <a href="https://github.com/mdenwood/JAGS-build-macOS" rel="nofollow" target="_blank">GitHub repo for the macOS build scripts</a>.</p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://martynplummer.wordpress.com/2026/09/27/jags-5-0-0-is-now-available-for-macos/"> R – JAGS News</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/jags-5-0-0-is-now-available-for-macos/">JAGS 5.0.0 is now available for macOS</a>]]></content:encoded>
					
		
		<enclosure url="https://0.gravatar.com/avatar/fdc509bd31ae635d89cccbdc64ef09464ea1c20d7858c4089a07ea3bea91b8e3?s=96&#038;d=identicon&#038;r=G" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403941</post-id>	</item>
		<item>
		<title>Analytics Pipeline for Dashboards, with Python, R and Javascript</title>
		<link>https://www.r-bloggers.com/2026/09/analytics-pipeline-for-dashboards-with-python-r-and-javascript/</link>
		
		<dc:creator><![CDATA[T. Moudiki]]></dc:creator>
		<pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://thierrymoudiki.github.io//blog/2026/09/27/python/r/javascript/analytics-pipeline</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> I had a lot of fun this morning, brainstorming and assembling this _Analytics Pipeline_ for Dashboards with Claude (and yes it takes much, much more than only 10 prompts).</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/analytics-pipeline-for-dashboards-with-python-r-and-javascript/">Analytics Pipeline for Dashboards, with Python, R and Javascript</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://thierrymoudiki.github.io//blog/2026/09/27/python/r/javascript/analytics-pipeline"> T. Moudiki's Webpage - R</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>I had a lot of fun this morning, brainstorming and assembling this <em>Analytics Pipeline</em> for Dashboards with Claude (and yes it takes much, much more than only 10 prompts).</p>

<p>The philosophy is borrowed from Observable Framework’s data loader, but implemented from scratch with voluntarily opinionated choices of libraries for Python, R and Javascript (and no Markdown): pre-compute data with <strong>Python + Polars</strong> (Polars is my new <em>crush</em>) or <strong>R + dplyr</strong>, publish it as static files, and explore them with <strong>JavaScript</strong> in the browser (<a href="https://github.com/uwdata/arquero" rel="nofollow" target="_blank">Arquero</a> for data wrangling, <a href="https://observablehq.github.io/plot/" rel="nofollow" target="_blank">Observable Plot</a> or <a href="https://www.highcharts.com/" rel="nofollow" target="_blank">Highcharts</a> for charts)</p>

<p>The repository is available on <a href="https://github.com/thierrymoudiki/analytics-pipeline" rel="nofollow" target="_blank">GitHub</a> and here’s the quick start guide to run it locally (you need to have <a href="https://www.python.org/downloads/" rel="nofollow" target="_blank">Python</a> and <a href="https://www.r-project.org/" rel="nofollow" target="_blank">R</a> installed on your machine):</p>

<pre>make                        # list all targets
uv venv venv                # or: make venv (python -m venv venv without uv)
source venv/bin/activate    # Windows: venv\Scripts\activate
make install                # Polars into venv/, dplyr into .uvr/library/
make dev                    # starts http://localhost:3000 and opens your browser
</pre>

<p>The repository’s README file is also informative and contains a few more details about the pipeline.</p>

<p><img src="https://i1.wp.com/thierrymoudiki.github.io/images/2026-09-27/2026-09-27-image1.png?w=578&#038;ssl=1" alt="image-title-here" class="img-responsive" data-recalc-dims="1" /></p>


<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://thierrymoudiki.github.io//blog/2026/09/27/python/r/javascript/analytics-pipeline"> T. Moudiki's Webpage - R</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/analytics-pipeline-for-dashboards-with-python-r-and-javascript/">Analytics Pipeline for Dashboards, with Python, R and Javascript</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403939</post-id>	</item>
		<item>
		<title>Autumn 2026 Data Science Training Courses</title>
		<link>https://www.r-bloggers.com/2026/09/autumn-2026-data-science-training-courses/</link>
		
		<dc:creator><![CDATA[The Jumping Rivers Blog]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 23:59:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://www.jumpingrivers.com/blog/autumn-2026-training-courses/</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; ">
<p>There are ten public training courses left in 2026, running through October and November. If you’ve been meaning to learn Python, build your first Shiny app, or get to grips with machine learning in R, there’s still time to do it befo...</p></div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/autumn-2026-data-science-training-courses/">Autumn 2026 Data Science Training Courses</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://www.jumpingrivers.com/blog/autumn-2026-training-courses/"> The Jumping Rivers Blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<p>
<a href = "https://www.jumpingrivers.com/blog/autumn-2026-training-courses/">
<img src="https://i1.wp.com/www.jumpingrivers.com/blog/autumn-2026-training-courses/featured.png?w=400&#038;ssl=1" style="width:400px" class="image-center" style="display: block; margin: auto;" data-recalc-dims="1" />
</a>
</p>
<p>There are ten public training courses left in 2026, running through October and November. If you’ve been meaning to learn Python, build your first Shiny app, or get to grips with machine learning in R, there’s still time to do it before the end of the year.</p>
<p>All courses run online and are delivered live by data scientists and engineers who work on client projects. Every participant gets course notes and scripts, live demonstrations, and hands-on exercises throughout.</p>
<h2 id="whats-on-this-autumn">What’s On This Autumn</h2>
<h3 id="python-in-october">Python in October</h3>
<p>October is a good month to pick up Python. <strong>Introduction to Python</strong> runs on 12th–13th October, <strong>Programming with Python</strong> follows a week later, and <strong>Data Visualisation with Python</strong> runs straight after that. You can take all three in turn to go from your first lines of Python to building your own functions and charts in about two weeks.</p>
<h3 id="shiny">Shiny</h3>
<p><strong>Introduction to Shiny</strong> runs on 5th–6th October and takes you from an empty script to a working interactive web app in R. If you already build Shiny apps, <strong>Advanced Concepts in Shiny</strong> on 23rd–24th November covers how to write maintainable code, build robust apps that handle unexpected user input, and get more out of reactive programming.</p>
<h3 id="visualisation-and-statistics-with-r">Visualisation and Statistics with R</h3>
<p><strong>Data Visualisation with ggplot2</strong> runs on 2nd–3rd November. It’s a good fit for anyone who uses R already and wants to produce clear, publication-quality charts. <strong>Statistical Modelling with R</strong> runs on 16th–17th November and covers hypothesis testing, regression, clustering and principal components analysis.</p>
<h3 id="machine-learning-with-tidymodels">Machine Learning with Tidymodels</h3>
<p><strong>Machine Learning with Tidymodels</strong> runs on 4th–5th November, and <strong>Advanced Machine Learning with Tidymodels</strong> follows on 23rd–24th November. Taking both gives you a complete machine learning workflow in R, from pre-processing and model fitting through to tuning.</p>
<h3 id="git">Git</h3>
<p><strong>Introduction to Git</strong> runs on the mornings of 16th–17th November. It covers why version control matters, how to use it properly, and how to start collaborating with others using Git.</p>
<p>Git and Advanced Machine Learning with Tidymodels run in the mornings, so you can take them alongside an afternoon course in the same week.</p>
<aside class="advert">
<p>
Whether you want to start from scratch, or improve your skills, <a href="https://www.jumpingrivers.com/training/?utm_source=blog&#038;utm_medium=banner&#038;utm_campaign=2026-autumn-training-courses" rel="nofollow" target="_blank">Jumping Rivers has a training course for you</a>.
</p>
</aside>
<h2 id="autumn-2026-schedule">Autumn 2026 Schedule</h2>
<p>Each course runs over two sessions of three and a half hours. Afternoon sessions run from 13:30 to 17:00 and morning sessions from 09:00 to 12:30 (UK time). Click a course to see the full outline and book.</p>
<table>
<thead>
<tr>
<th>Date</th>
<th>Course</th>
<th>Level</th>
</tr>
</thead>
<tbody>
<tr>
<td>5th–6th October</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-introduction-shiny-application-web/#event-2026-10-05-13-30-online" rel="nofollow" target="_blank">Introduction to Shiny</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>12th–13th October</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-introduction-visualisation-manipulation/#event-2026-10-12-13-30-online" rel="nofollow" target="_blank">Introduction to Python</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>19th–20th October</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-programming-control-flow-functions/#event-2026-10-19-13-30-online" rel="nofollow" target="_blank">Programming with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>21st–22nd October</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-matplotlib-seaborn-visualisation/#event-2026-10-21-13-30-online" rel="nofollow" target="_blank">Data Visualisation with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>2nd–3rd November</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-advanced-graphics-ggplot2-plotly-themes-scaling-faceting/#event-2026-11-02-13-30-online" rel="nofollow" target="_blank">Data Visualisation with ggplot2</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>4th–5th November</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-prediction-inference-analytics-machine-learning-tidymodels/#event-2026-11-04-13-30-online" rel="nofollow" target="_blank">Machine Learning with Tidymodels</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>16th–17th November (mornings)</td>
<td><a href="https://www.jumpingrivers.com/training/course/intro-to-git/#event-2026-11-16-09-00-online" rel="nofollow" target="_blank">Introduction to Git</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>16th–17th November</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-statistics-modelling-linear-regression-clustering/#event-2026-11-16-13-30-online" rel="nofollow" target="_blank">Statistical Modelling with R</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>23rd–24th November (mornings)</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-prediction-inference-tidymodels-lda-pre-processing-tree-based-models/#event-2026-11-23-09-00-online" rel="nofollow" target="_blank">Advanced Machine Learning with Tidymodels</a></td>
<td>Advanced</td>
</tr>
<tr>
<td>23rd–24th November</td>
<td><a href="https://www.jumpingrivers.com/training/course/advanced-shiny/#event-2026-11-23-13-30-online" rel="nofollow" target="_blank">Advanced Concepts in Shiny</a></td>
<td>Advanced</td>
</tr>
</tbody>
</table>
<p>Booking closes one week before each course starts, so Introduction to Shiny closes on 28th September.</p>
<p>Early bird pricing is still available on some November courses: until 4th October for Introduction to Git and Statistical Modelling with R, and until 11th October for both Advanced courses.</p>
<h2 id="planning-ahead-book-now-for-2027">Planning Ahead? Book Now for 2027</h2>
<p>If the autumn is too busy, you can already book for next year. Our public schedule for January to June 2027 is live, and every course on it is currently at the early bird price.</p>
<p>It follows the same pattern as this year. Introduction to R and Data Wrangling in the Tidyverse run in January and again in April, and the Python courses run in February and May. Machine learning with Tidymodels runs in March.</p>
<p>It also includes <strong>Introduction to Bayesian Inference using RStan</strong> in January. It runs over four afternoons and covers writing Stan programs for a range of statistical models and checking MCMC diagnostics in R.</p>
<p>Early bird pricing ends about six weeks before each course, so for the January courses that means booking by late November or early December.</p>
<table>
<thead>
<tr>
<th>Date</th>
<th>Course</th>
<th>Level</th>
</tr>
</thead>
<tbody>
<tr>
<td>11th–12th January</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-introduction-tidyverse-readr-ggplot2-dplyr/#event-2027-01-11-13-30-online" rel="nofollow" target="_blank">Introduction to R</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>18th–21st January</td>
<td><a href="https://www.jumpingrivers.com/training/course/introduction-bayesian-inference-rstan-monte-carlo/#event-2027-01-18-13-30-online" rel="nofollow" target="_blank">Introduction to Bayesian Inference using RStan</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>25th–26th January</td>
<td><a href="https://www.jumpingrivers.com/training/course/data-tidyverse-dplyr-tidyr-lubridate-forcats/#event-2027-01-25-13-30-online" rel="nofollow" target="_blank">Data Wrangling in the Tidyverse</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>1st–2nd February</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-advanced-graphics-ggplot2-plotly-themes-scaling-faceting/#event-2027-02-01-13-30-online" rel="nofollow" target="_blank">Data Visualisation with ggplot2</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>8th–9th February</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-programming-functions-looping-conditionals/#event-2027-02-08-13-30-online" rel="nofollow" target="_blank">Programming with R</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>15th–16th February</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-introduction-visualisation-manipulation/#event-2027-02-15-13-30-online" rel="nofollow" target="_blank">Introduction to Python</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>22nd–23rd February</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-programming-control-flow-functions/#event-2027-02-22-13-30-online" rel="nofollow" target="_blank">Programming with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>1st–2nd March</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-matplotlib-seaborn-visualisation/#event-2027-03-01-13-30-online" rel="nofollow" target="_blank">Data Visualisation with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>8th–9th March</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-statistics-modelling-linear-regression-clustering/#event-2027-03-08-13-30-online" rel="nofollow" target="_blank">Statistical Modelling with R</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>15th–16th March</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-prediction-inference-analytics-machine-learning-tidymodels/#event-2027-03-15-13-30-online" rel="nofollow" target="_blank">Machine Learning with Tidymodels</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>22nd–23rd March</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-prediction-inference-tidymodels-lda-pre-processing-tree-based-models/#event-2027-03-22-13-30-online" rel="nofollow" target="_blank">Advanced Machine Learning with Tidymodels</a></td>
<td>Advanced</td>
</tr>
<tr>
<td>12th–13th April</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-introduction-tidyverse-readr-ggplot2-dplyr/#event-2027-04-12-13-30-online" rel="nofollow" target="_blank">Introduction to R</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>19th–20th April</td>
<td><a href="https://www.jumpingrivers.com/training/course/data-tidyverse-dplyr-tidyr-lubridate-forcats/#event-2027-04-19-13-30-online" rel="nofollow" target="_blank">Data Wrangling in the Tidyverse</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>26th–27th April</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-advanced-graphics-ggplot2-plotly-themes-scaling-faceting/#event-2027-04-26-13-30-online" rel="nofollow" target="_blank">Data Visualisation with ggplot2</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>10th–11th May</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-programming-functions-looping-conditionals/#event-2027-05-10-13-30-online" rel="nofollow" target="_blank">Programming with R</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>17th–18th May</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-introduction-visualisation-manipulation/#event-2027-05-17-13-30-online" rel="nofollow" target="_blank">Introduction to Python</a></td>
<td>Foundation</td>
</tr>
<tr>
<td>24th–25th May</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-programming-control-flow-functions/#event-2027-05-24-13-30-online" rel="nofollow" target="_blank">Programming with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>14th–15th June</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-matplotlib-seaborn-visualisation/#event-2027-06-14-13-30-online" rel="nofollow" target="_blank">Data Visualisation with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>16th–17th June</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-statistics-modelling-linear-regression-clustering/#event-2027-06-16-13-30-online" rel="nofollow" target="_blank">Statistical Modelling with R</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>21st–22nd June</td>
<td><a href="https://www.jumpingrivers.com/training/course/reporting-with-quarto/#event-2027-06-21-13-30-online" rel="nofollow" target="_blank">Reporting with Quarto</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>23rd–24th June</td>
<td><a href="https://www.jumpingrivers.com/training/course/python-programming-control-flow-functions/#event-2027-06-23-13-30-online" rel="nofollow" target="_blank">Programming with Python</a></td>
<td>Intermediate</td>
</tr>
<tr>
<td>28th–29th June</td>
<td><a href="https://www.jumpingrivers.com/training/course/r-prediction-inference-analytics-machine-learning-tidymodels/#event-2027-06-28-13-30-online" rel="nofollow" target="_blank">Machine Learning with Tidymodels</a></td>
<td>Intermediate</td>
</tr>
</tbody>
</table>
<p>We’ll add dates for the second half of 2027 later in the year.</p>
<h2 id="why-train-with-jumping-rivers">Why Train with Jumping Rivers</h2>
<p>Our trainers are practising data scientists and engineers. The examples and exercises in every course come from the work we do with clients, so you learn how these tools are used in real projects.</p>
<p>Every participant receives:</p>
<ul>
<li>Comprehensive PDF notes and scripts to keep after the course</li>
<li>Live demonstrations and hands-on exercises throughout</li>
<li>Direct access to a trainer for questions</li>
</ul>
<p>We have delivered over 1,000 courses to organisations including NHS Scotland, Shell, Wessex Water and the Royal Statistical Society.</p>
<h2 id="additional-perks">Additional Perks</h2>
<p>We run free webinars throughout the year. Attending a Jumping Rivers webinar gives you:</p>
<ul>
<li>Early exposure to new topics in data science and analytics</li>
<li>Up to 20% off training courses</li>
<li>Up to 20% off <a href="https://ai-in-production.jumpingrivers.com/" rel="nofollow" target="_blank">Jumping Rivers conferences</a></li>
</ul>
<p><a href="https://jumpingrivers.typeform.com/to/UmdyNbAs" rel="nofollow" target="_blank">Register for upcoming webinars here.</a></p>
<h2 id="training-for-teams">Training for Teams</h2>
<p>We also run in-house training for organisations that want to develop their teams. We tailor each course to your workflows, tools and experience levels, and we can deliver any course in our <a href="https://www.jumpingrivers.com/training/all-courses/" rel="nofollow" target="_blank">full catalogue</a>, not just the ones on the public schedule. Group bookings and returning clients get a discount.</p>
<p>To discuss options, email <a href="mailto:training@jumpingrivers.com" rel="nofollow" target="_blank">training@jumpingrivers.com</a>.</p>
<h2 id="book-your-place">Book Your Place</h2>
<p>View the full schedule and book at <a href="https://www.jumpingrivers.com/training/public/" rel="nofollow" target="_blank">jumpingrivers.com/training/public</a>. Places on public courses are limited, so book early if you have a particular date in mind.</p>
<p>
For updates and revisions to this article, see the <a href = "https://www.jumpingrivers.com/blog/autumn-2026-training-courses/">original post</a>
</p>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://www.jumpingrivers.com/blog/autumn-2026-training-courses/"> The Jumping Rivers Blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/autumn-2026-data-science-training-courses/">Autumn 2026 Data Science Training Courses</a>]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">403880</post-id>	</item>
		<item>
		<title>On Exact Hat Algebras</title>
		<link>https://www.r-bloggers.com/2026/09/on-exact-hat-algebras/</link>
		
		<dc:creator><![CDATA[https://pacha.dev/blog]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 23:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://pacha.dev/blog/2026/09/26/index.html</guid>

					<description><![CDATA[<p>What is a hat algebra and why it is useful to solve maximization problems</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/on-exact-hat-algebras/">On Exact Hat Algebras</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://pacha.dev/blog/2026/09/26/index.html"> https://pacha.dev/blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p><em>Before the main content:</em></p>
<ul>
<li><em>I am creating an R Community on Google Groups. You can join the group using this <a href="https://docs.google.com/forms/d/e/1FAIpQLSdMAj4adRAT4Gyuwt_9dPvxRvOUPml9AD59vuI7qS7XDlp48g/viewform?usp=dialog" rel="nofollow" target="_blank">form</a>.</em></li>
<li><em>With all the funding cuts to PhD studies, I’d appreciate if you can donate <a href="https://buymeacoffee.com/pacha" rel="nofollow" target="_blank">here</a>.</em></li>
</ul>
<p>In Economics we often use reduced-form models, which are simplified ways to describe a process. I tend to see the criticism to this as unfair, as in physics they do the same, and so many models in physics assume “spherical objects in a vacuum” as a simplification to make real-world problems mathematically solvable.</p>
<p>With reduced form models we are not interested in how faithful those ressemble reality but how useful these are. As an example in support of reduced form models, consider the two extreme cases of models to answer “how to get from 37 Old Queen St to St. James Park underground station?”:</p>
<ol>
<li>Person 1 draws a quick sketch like this on an old receipt</li>
</ol>
<pre>      _ Embassy
     |   
  ---
 |
Underground</pre>
<ol>
<li>Person 2 built a very detailed 1:10000 architectural model of St. James neighbourhood and shows you the tiny streets, cars, etc.</li>
</ol>
<p>For the question, person 1 sketch is sufficient.</p>
<p>The Armington model of trade sits between these two extremes. It considers a world of \(N\) countries, each country \(i\) has a fixed labour endowment \(w_i\) and there is an endogenous labour wage. Income is given by \(Y_i = w_i L_i\). Trade flows are given by \(X_{ij}\) (and so the sales share is \(\gamma_{ij} = X_{ij} / Y_i\)). There is a productivity shift \(\x_i\), a trade cost $_{ij}, and an elasticity of substitution \(\varepsilon\).</p>
<p>The strong assumption in this model is a known and homogenous elasticity of substitution. In this model countries can substitute between Apple and Android phones even when many would argue that those phones compete in different markets.</p>
<p>The model imposes a market clearing condition \(w_i L_i = \sum_{j=1}^N \lambda_{ij} w_j L_j\).</p>
<p>There are many values \(\lambda_{12}, \lambda_{13}, \ldots, \lambda_{1N}\) that balance the market clearing condition but in this model the solution is a gravity-type equation that leads to $<em>{ij} = (<em>i / (</em>{ij} w_i)^{}) / sum</em>{l=1}^N (<em>l / (</em>{lj} w_l)^{}). More about gravity-type equations can be read in <a href="https://yotoyotov.com/Gravity_Undergrads.html" rel="nofollow" target="_blank">Gravity for Undergrads</a> by Prof. Dr. Yoto V. Yotov.</p>
<p>Consider, for example, the EU increase on low value imports from the UK described [https://hboltd.co.uk/uk-eu-shipping-changes-2026-customs-overhaul/]. In this model, instead of a product-by-product specific tariff have an initial UK-EU average trade cost of \(\tau_{ij}\) that increased to \(\tau_{ij}&#8217;\). The hat form of this is \(\hat{\tau_{ij}} = \tau_{ij}&#8217; / \tau_{ij}\).</p>
<p>After the trade cost change the market clearing would be \(w_i&#8217; L_i&#8217; = \sum_{j=1}^N \lambda_{ij}&#8217; w_j&#8217; L_j&#8217;\), and then</p>
<p>\[
\hat{w_i} \hat{L_i} = \frac{\sum_{j=1}^N \lambda_{ij}&#8217; w_j&#8217; L_j&#8217;}{\sum_{j=1}^N \lambda_{ij} w_j L_j}
\]</p>
<p>\[
\hat{w_i} \hat{L_i} = \frac{\sum_{j=1}^N X_{ij}&#8217;}{w_i L_i}
\]</p>
<p>\[
\hat{w_i} \hat{L_i} = \sum_{j=1}^N \gamma_{ij} \hat{X_{ij}}
\]</p>
<p>Similarly, $ = .</p>
<p>The new hat equations are useful, for example, to determine the effect of the change to the trade cost over wages \(\hat{w_i}\). The hat market clearing means that \(\hat{w_i} = \sum_{j=1}^N \gamma_{ij} \hat{X_{ij}}\) (labour is exogenous), and therefore \(\hat{w_i} = \sum_{j=1}^N \gamma_{ij} \hat{\lambda_{ij}} \hat{w_j}\).</p>
<p>Merging the hat market clearing with the hat gravity equation, we get</p>
<p>\[
\hat{w_i} = \sum_{j=1}^N \frac{\gamma_{ij} \hat{w_j} / (\hat{\tau_{ij}} \hat{w_i})^{\varepsilon}}{\sum_{l=1}^N \gamma_{lj} \hat{w_l} / (\hat{\tau_{lj}} \hat{w_l})^{\varepsilon}}
\]</p>
<p>From the initial equilibrium shares \(\lambda_{ij}\) and \(\gamma_{ij}\) and a fixed \(\varepsilon\) (for example, estimated using USITC data), we have a system of \(N \times N\) equations that we can solve to get \(\hat{w_j}\) and then recover \(\hat{\lambda_{ij}}\). In other words, we simplified an identification problem by solving the differences by resorting on a sufficient statistic instead of estimating all trade costs.</p>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://pacha.dev/blog/2026/09/26/index.html"> https://pacha.dev/blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/on-exact-hat-algebras/">On Exact Hat Algebras</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403915</post-id>	</item>
		<item>
		<title>Quantitative Event-Driven Modeling: BTC Balance Sheet Shocks with SEC EDGAR Filings in R</title>
		<link>https://www.r-bloggers.com/2026/09/quantitative-event-driven-modeling-btc-balance-sheet-shocks-with-sec-edgar-filings-in-r/</link>
		
		<dc:creator><![CDATA[Selcuk Disci]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 10:50:06 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">http://datageeek.com/?p=12790</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> Executive Summary MicroStrategy (MSTR) has fundamentally transitioned from a traditional enterprise software firm into an equity-based Bitcoin (BTC) holding vehicle and treasury operation. Evaluating MSTR using conventional corporate finance metrics (such as Price-to-Earnings or EBITDA multiples) fails to capture the core driver of its equity valuation: the dynamic Net Asset ...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/quantitative-event-driven-modeling-btc-balance-sheet-shocks-with-sec-edgar-filings-in-r/">Quantitative Event-Driven Modeling: BTC Balance Sheet Shocks with SEC EDGAR Filings in R</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://datageeek.com/2026/09/25/quantitative-event-driven-modeling-btc-balance-sheet-shocks-with-sec-edgar-filings-in-r/"> DataGeeek</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<h2 class="wp-block-heading">Executive Summary</h2>



<p class="wp-block-paragraph">MicroStrategy (<strong>MSTR</strong>) has fundamentally transitioned from a traditional enterprise software firm into an equity-based Bitcoin (<strong>BTC</strong>) holding vehicle and treasury operation. Evaluating MSTR using conventional corporate finance metrics (such as Price-to-Earnings or <strong>EBITDA </strong>multiples) fails to capture the core driver of its equity valuation: the dynamic Net Asset Value (<strong>NAV</strong>) premium driven by programmatic Bitcoin treasury expansion.</p>



<p class="wp-block-paragraph">This technical article explores an end-to-end, reproducible quantitative pipeline built in <strong>R</strong>. Inspired by recent advancements in automated <strong>SEC </strong>parsing tools—specifically the secfile architecture highlighted in <a href="https://www.interactivebrokers.com/campus/ibkr-quant-news/secfile-sec-edgar-filings-in-r-and-python/" rel="nofollow" target="_blank"><strong>Interactive Brokers Quant Blog</strong></a>—this pipeline ingests real-time <strong>SEC EDGAR</strong> filings, dynamically extracts digital asset balance sheet facts, aligns mixed-frequency financial and market data, and models <strong>MSTR </strong>daily equity log returns using an out-of-sample <strong>82.08% R-Squared</strong> structural tidymodels framework.</p>



<h2 class="wp-block-heading">1. The SEC EDGAR Challenge &#038; The secfile Paradigm</h2>



<h3 class="wp-block-heading">Unstructured Financial Data vs. Machine-Readable XBRL</h3>



<p class="wp-block-paragraph">Historically, extraction of balance sheet metrics directly from SEC filings (<strong>Forms 10-K and 10-Q</strong>) required brittle HTML regex scraping or costly third-party commercial APIs. Corporate disclosures often vary across reporting periods, particularly when handling emerging asset classes like digital currencies. MicroStrategy’s <strong>XBRL </strong>taxonomy has evolved, tagging Bitcoin holdings under varying terms such as <strong>DigitalAssets</strong>, <strong>CryptocurrencyHoldings</strong>, or within broader <strong>Assets </strong>line items.</p>



<p class="wp-block-paragraph">As demonstrated in the <strong>Interactive Brokers Quant Blog</strong> overview of secfile, direct ingestion of <strong>SEC EDGAR</strong> financial facts via compliant <strong>XBRL </strong>parsing solves three primary institutional hurdles:</p>



<ol class="wp-block-list">
<li><strong>Regulatory Compliance &#038; Transparency</strong>: Enforcing strict HTTP User-Agent headers compliant with SEC EDGAR access policies ensures uninterrupted data pipeline ingestion.</li>



<li><strong>Dynamic Taxonomy Mapping</strong>: By searching the taxonomy dictionary programmatically, the pipeline automatically handles changes in reporting terminology without hardcoded column dependencies.</li>



<li><strong>Point-in-Time Alignment</strong>: Raw SEC filings record historical submission dates, preventing look-ahead bias when backtesting event-driven trading strategies against historical market prices.</li>
</ol>



<p class="wp-block-paragraph">Using <strong>tidy evaluation</strong> via the <strong>rlang </strong>package, the pipeline dynamically isolates balance sheet events safely without breaking execution when <strong>SEC XBRL</strong> tags undergo structural revisions.</p>



<h2 class="wp-block-heading">2. Mathematical Structure &#038; Structural Economics</h2>



<p class="wp-block-paragraph">The core hypothesis of this quantitative framework is that MSTR daily return dynamics are governed by two orthogonal components:</p>



<ol class="wp-block-list">
<li><strong>Systematic Asset Return Co-movement</strong>: Direct market return exposure to spot Bitcoin returns.</li>



<li><strong>Balance Sheet Shock Factors</strong>: Discontinuous structural changes in treasury balance sheet holdings resulting from capital raises or debt-funded BTC acquisitions.</li>
</ol>



<h3 class="wp-block-heading">Daily Log Returns</h3>



<p class="wp-block-paragraph">Continuous log returns for asset prices are calculated by taking the <strong>natural logarithm</strong> of the <strong>current price</strong> divided by <strong>the previous day’s price</strong>. Taking natural logarithms guarantees <strong>additivity </strong>over time horizons and prevents <strong>non-negative</strong> boundary issues inherent to simple percentage returns.</p>



<h3 class="wp-block-heading">Balance Sheet Shock Factor</h3>



<p class="wp-block-paragraph">To quantify <strong>treasury expansion</strong> independently of market price <strong>fluctuation</strong>, the Balance Sheet Shock metric measures the <strong>logarithmic growth</strong> rate of total Bitcoins held on MicroStrategy’s balance sheet.</p>



<p class="wp-block-paragraph">Because corporate filings occur at <strong>discrete quarterly intervals</strong>, daily balance sheet values are carried forward using <strong>Last Observation Carried Forward</strong> via the <strong>zoo </strong>package. Consequently, the balance sheet shock value equals zero on non-reporting days, triggering discrete structural impulses only on filings or event reporting dates.</p>



<h3 class="wp-block-heading">Econometric Specification</h3>



<p class="wp-block-paragraph">The <strong>structural linear model</strong> specification expresses <strong>the daily log return</strong> of MSTR as a function of three main components:</p>



<ul class="wp-block-list">
<li>The <strong>intercept term</strong>, representing <strong>base drift</strong>.</li>



<li>The spot <strong>Bitcoin log return</strong>, weighted by the <strong>Equity-to-BTC Beta Elasticity </strong>coefficient<strong> </strong>(measuring leverage and <strong>NAV premium</strong> response).</li>



<li><strong>The Balance Sheet Shock factor</strong>, weighted by the <strong>Treasury Expansion</strong> Impact coefficient.</li>



<li>A <strong>residual </strong>error term capturing unexplained market <strong>noise</strong>.</li>
</ul>



<h2 class="wp-block-heading">3. Pipeline Architecture &#038; Package Integration</h2>



<p class="wp-block-paragraph">The pipeline follows a modern, production-grade functional programming architecture using the tidyverse ecosystem:</p>



<ul class="wp-block-list">
<li><strong>secfile</strong>: High-throughput scraper and <strong>XBRL </strong>parser for <strong>SEC EDGAR</strong> datasets.</li>



<li><strong>tidyquant</strong>: Financial wrapper connecting <strong>quantmod</strong> and <strong>PerformanceAnalytics</strong> packages into tidy data frames.</li>



<li><strong>timetk</strong>: Provides time-series splitting for chronological out-of-sample data partitioning without temporal leakage.</li>



<li><strong>tidymodels</strong>: Modular modeling framework combining <strong>recipe</strong> creation, <strong>parsnip </strong>engine specification, and <strong>workflow </strong>execution.</li>



<li><strong>plotly &#038; ggtext</strong>: Interactive visualization engine supporting <strong>HTML/CSS</strong> styled markdown <strong>tooltips </strong>and dynamic price projection bands.</li>



<li><strong>zoo</strong>: Handles non-homogeneous time-series alignment via Last Observation Carried Forward (<strong>na.locf</strong>).</li>
</ul>



<h2 class="wp-block-heading">4. Chronological Splitting &#038; Model Training</h2>



<p class="wp-block-paragraph">Financial time series violate the Independent and Identically Distributed (<strong>i.i.d.</strong>) assumption of standard k-fold cross-validation due to <strong>autocorrelation</strong>. To preserve time order, we employ strict time-based windowing via <strong>timetk</strong>.</p>



<p class="wp-block-paragraph">Using recipes and workflows, we define the feature roles and couple them with a native Ordinary Least Squares (<strong>OLS</strong>) estimation engine. This guarantees clean execution without <strong>data leakage</strong> across the training split and testing horizon.</p>



<h2 class="wp-block-heading">5. Empirical Results &#038; Performance Evaluation</h2>



<p class="wp-block-paragraph">Out-of-sample validation was conducted over a <strong>15-day forward horizon</strong>, evaluating <strong>the structural model</strong> against <strong>a Naive Baseline Benchmark</strong> (1-day lagged return persistence model where predicted return equals yesterday’s actual return).</p>



<h3 class="wp-block-heading">Performance Metrics Output</h3>



<ul class="wp-block-list">
<li>Structural Model RMSE: <strong>0.0241</strong></li>



<li>Naive Baseline RMSE: <strong>0.0583</strong></li>



<li>Structural Model R-Squared: <strong>82.08%</strong></li>



<li>Naive Baseline R-Squared: <strong>4.12%</strong></li>
</ul>



<p class="wp-block-paragraph">The Structural Model achieved an out-of-sample R-Squared of 82.08%, <strong>significantly outperforming</strong> the Naive Baseline. This confirms that equity return variance in MSTR is <strong>predominantly explained</strong> by spot Bitcoin fluctuations and treasury balance sheet updates rather than simple price momentum.</p>


<pre>
# ==============================================================================
# SEC XBRL DRIVEN QUANT PIPELINE: MICROSTRATEGY BALANCE SHEET SHOCKS VS BTC
# Author: Selcuk Disci (datageeek.com)
# ==============================================================================

# 1. LOAD REQUIRED LIBRARIES (AUTOMATED PACMAN ENTIRE PIPELINE INGESTION)
# ------------------------------------------------------------------------------
# Check and install pacman package manager if not already available
if (!require(&quot;pacman&quot;)) install.packages(&quot;pacman&quot;)
# Ingest core financial, data manipulation, SEC scraping, and modeling libraries
pacman::p_load(secfile, tidyquant, tidyverse, zoo, timetk, tidymodels)

# 2. DEFINE SEC COMPLIANT USER AGENT AND FETCH DATA
# ------------------------------------------------------------------------------
# Set contact email required by SEC EDGAR fair access policy header guidelines
user_agent &lt;- &quot;&lt;your_email_address&gt;&quot;

# Fetch SEC Central Index Key (CIK) identifier for MicroStrategy Inc.
mstr_cik &lt;- get_ciks(&quot;MSTR&quot;, user_agent = user_agent)

# Retrieve corporate submissions metadata and extract structured XBRL financial facts
mstr_subs &lt;- get_submissions(mstr_cik, user_agent = user_agent)
mstr_facts &lt;- get_data(mstr_subs, user_agent = user_agent)

# 3. EXTRACTION OF BITCOIN HOLDINGS DYNAMICALLY FROM WIDE FORMAT
# ------------------------------------------------------------------------------
# Extract column names from SEC dataset to locate dynamic crypto asset tags
available_columns &lt;- names(mstr_facts)
# Detect target columns matching SEC XBRL taxonomy for digital holdings
target_btc_column &lt;- available_columns[str_detect(available_columns, &quot;DigitalAsset|CryptocurrencyHoldings&quot;)]

# Fallback safety net to ensure execution if specific digital asset tags are absent
if(length(target_btc_column) == 0) {
  target_btc_column &lt;- &quot;Assets&quot; 
} else {
  target_btc_column &lt;- target_btc_column[1]
}

# Clean and transform raw Bitcoin holdings time series data
mstr_btc_holdings &lt;- mstr_facts %&gt;%
  select(report_date, !!sym(target_btc_column)) %&gt;% # Dynamically unquote &#039;target_btc_column&#039; using !!sym() to evaluate the string as a column name
  rename(date = report_date, BTC_Held = !!sym(target_btc_column)) %&gt;%
  mutate(
    date = as.Date(date),
    BTC_Held = as.numeric(BTC_Held) 
  ) %&gt;%
  filter(!is.na(BTC_Held)) %&gt;%
  distinct(date, .keep_all = TRUE) %&gt;%
  arrange(date)

# 4. FETCH AND COMPUTE LOG RETURNS USING TIDYQUANT
# ------------------------------------------------------------------------------
# Define equity and cryptocurrency symbols for market data retrieval
tickers &lt;- c(&quot;MSTR&quot;, &quot;BTC-USD&quot;)

# Download adjusted close daily historical price series and calculate log returns
market_returns &lt;- tq_get(tickers, from = &quot;2020-01-01&quot;, get = &quot;stock.prices&quot;) %&gt;%
  group_by(symbol) %&gt;%
  tq_mutate(select = adjusted,
            mutate_fun = periodReturn,
            period = &quot;daily&quot;,
            type = &quot;log&quot;,
            col_rename = &quot;log_return&quot;) %&gt;%
  select(date, symbol, log_return, adjusted) %&gt;%
  ungroup()

# Pivot market data into a wide format and standardize asset-specific return/price columns
market_pivoted &lt;- market_returns %&gt;%
  pivot_wider(names_from = symbol, values_from = c(log_return, adjusted)) %&gt;%
  rename(
    MSTR_Log_Return = log_return_MSTR,
    BTC_Log_Return = `log_return_BTC-USD`,
    MSTR_Close = adjusted_MSTR,
    BTC_Close = `adjusted_BTC-USD`
  ) %&gt;%
  mutate(date = as.Date(date))

# 5. ALIGN MIXED FREQUENCY DATA AND COMPUTE SHOCKS
# ------------------------------------------------------------------------------
# Join balance sheet data with market returns and impute missing daily holdings values via forward fill
processed_model_data &lt;- market_pivoted %&gt;%
  left_join(mstr_btc_holdings, by = &quot;date&quot;) %&gt;%
  mutate(BTC_Held_Daily = na.locf(BTC_Held, na.rm = FALSE)) %&gt;%
  filter(!is.na(BTC_Held_Daily) & !is.na(MSTR_Log_Return) & !is.na(BTC_Log_Return)) %&gt;%
  mutate(Balance_Sheet_Shock = log(BTC_Held_Daily / lag(BTC_Held_Daily))) %&gt;%
  filter(!is.na(Balance_Sheet_Shock) & is.finite(Balance_Sheet_Shock))

# 6. TIME-BASED DATA SPLITTING VIA TIMETK (STRICT TIME WINDOWS)
# ------------------------------------------------------------------------------
# Partition dataset chronologically to prevent future data leakage during evaluation
data_splits &lt;- time_series_split(
  data       = processed_model_data,
  date_var   = date,
  initial    = &quot;1 year&quot;,   
  assess     = &quot;15 days&quot;,  
  cumulative = FALSE       
)

# Extract historical training split and testing horizon
train_data &lt;- training(data_splits)
test_data  &lt;- testing(data_splits)

# 7. ESTIMATE LINEAR REGRESSION VIA NATIVE TIDYMODELS COMPONENT
# ------------------------------------------------------------------------------
# Define features profile inside the standard recipe framework
mstr_recipe &lt;- recipe(MSTR_Log_Return ~ date + BTC_Log_Return + Balance_Sheet_Shock, data = train_data) %&gt;%
  update_role(date, new_role = &quot;id&quot;)

# Specify parsnip linear regression engine spec
lm_spec &lt;- linear_reg() %&gt;%
  set_engine(&quot;lm&quot;) %&gt;%
  set_mode(&quot;regression&quot;)

# Bind graph components into a clean workflow architecture
mstr_workflow &lt;- workflow() %&gt;%
  add_recipe(mstr_recipe) %&gt;%
  add_model(lm_spec)

# Fit model natively on the chronological training split partition
fitted_lm_workflow &lt;- fit(mstr_workflow, data = train_data)

# 8. OUT-OF-SAMPLE PERFORMANCE TESTING WITH NAIVE BENCHMARK USING YARDSTICK
# ------------------------------------------------------------------------------
# Generate clean out-of-sample predictions via standard tidymodels syntax
test_data_predictions &lt;- predict(fitted_lm_workflow, new_data = test_data)

# Construct evaluation frame with actual returns and naive lag benchmark
evaluation_df &lt;- test_data %&gt;%
  select(date, MSTR_Log_Return, MSTR_Close) %&gt;%
  bind_cols(test_data_predictions) %&gt;%
  rename(Predicted_Return = .pred) %&gt;%
  mutate(Naive_Predicted_Return = lag(MSTR_Log_Return, default = first(MSTR_Log_Return)))

# Reshape predictions to comparative long format for unified performance calculation
evaluation_long &lt;- evaluation_df %&gt;%
  select(date, MSTR_Log_Return, Predicted_Return, Naive_Predicted_Return) %&gt;%
  pivot_longer(
    cols = c(Predicted_Return, Naive_Predicted_Return),
    names_to = &quot;model_type&quot;,
    values_to = &quot;estimate&quot;
  ) %&gt;%
  rename(truth = MSTR_Log_Return) %&gt;%
  mutate(model_type = if_else(model_type == &quot;Predicted_Return&quot;, &quot;Structural_Model&quot;, &quot;Naive_Baseline&quot;))

# Compute RMSE and R-Squared accuracy metrics across structural vs naive models
my_financial_metrics   &lt;- metric_set(rmse, rsq)
accuracy_report_tibble &lt;- evaluation_long %&gt;%
  group_by(model_type) %&gt;%
  my_financial_metrics(truth = truth, estimate = estimate) %&gt;%
  ungroup() %&gt;%
  arrange(.metric, model_type)

# Print execution metric table to console
print(accuracy_report_tibble)

# 9. MODERN INTERACTIVE PLOTLY VISUALIZATION (PRICE-BASED DYNAMIC RSI)
# ------------------------------------------------------------------------------
# Load interactive graphics packages required for dynamic reporting
if (!require(&quot;pacman&quot;)) install.packages(&quot;pacman&quot;)
pacman::p_load(plotly, scales, glue, ggtext)

# Extract out-of-sample RMSE metric value for confidence band projection
trusted_rmse &lt;- accuracy_report_tibble %&gt;%
  filter(model_type == &quot;Structural_Model&quot; & .metric == &quot;rmse&quot;) %&gt;%
  pull(.estimate)

# Extract out-of-sample R-Squared metric value for header annotation
rsq_val &lt;- accuracy_report_tibble %&gt;%
  filter(model_type == &quot;Structural_Model&quot; & .metric == &quot;rsq&quot;) %&gt;%
  pull(.estimate)

# Convert predicted log returns back into absolute dollar prices with confidence intervals
df_eval &lt;- evaluation_df %&gt;%
  mutate(MSTR_Yesterday_Close = lag(MSTR_Close, default = first(MSTR_Close))) %&gt;%
  mutate(
    actual  = MSTR_Close,
    pred    = MSTR_Yesterday_Close * exp(Predicted_Return),
    conf_hi = MSTR_Yesterday_Close * exp(Predicted_Return + (2 * trusted_rmse)),
    conf_lo = MSTR_Yesterday_Close * exp(Predicted_Return - (2 * trusted_rmse))
  ) %&gt;%
  filter(date &gt; min(date))

# Build customized hover text data frames for interactive Plotly tooltips
df_plot_actual &lt;- df_eval %&gt;% select(date, actual)  %&gt;% mutate(text_actual = glue(&quot;&lt;b&gt;Actual MSTR Price:&lt;/b&gt; ${round(actual, 2)}&lt;br&gt;&lt;b&gt;Date:&lt;/b&gt; {format(date, &#039;%b %d, %Y&#039;)}&quot;))
df_plot_pred   &lt;- df_eval %&gt;% select(date, pred)    %&gt;% mutate(text_pred   = glue(&quot;&lt;b&gt;Linear AI Pred:&lt;/b&gt; ${round(pred, 2)}&lt;br&gt;&lt;b&gt;Date:&lt;/b&gt; {format(date, &#039;%b %d, %Y&#039;)}&quot;))
df_plot_hi     &lt;- df_eval %&gt;% select(date, conf_hi) %&gt;% mutate(text_hi     = glue(&quot;&lt;b&gt;Overbought Ceiling:&lt;/b&gt; ${round(conf_hi, 2)}&lt;br&gt;&lt;b&gt;Date:&lt;/b&gt; {format(date, &#039;%b %d, %Y&#039;)}&quot;))
df_plot_lo     &lt;- df_eval %&gt;% select(date, conf_lo) %&gt;% mutate(text_lo     = glue(&quot;&lt;b&gt;Oversold Floor:&lt;/b&gt; ${round(conf_lo, 2)}&lt;br&gt;&lt;b&gt;Date:&lt;/b&gt; {format(date, &#039;%b %d, %Y&#039;)}&quot;))

# Assemble base ggplot layer with confidence bands, actuals, and model projections
p &lt;- ggplot() +
  geom_ribbon(data = df_eval, aes(x = date, ymin = conf_lo, ymax = conf_hi), fill = &quot;#808080&quot;, alpha = 0.18) +
  geom_point(data = df_plot_hi, aes(x = date, y = conf_hi, text = text_hi), color = &quot;transparent&quot;, alpha = 0, size = 3) +
  geom_point(data = df_plot_lo, aes(x = date, y = conf_lo, text = text_lo), color = &quot;transparent&quot;, alpha = 0, size = 3) +
  geom_line(data = df_plot_actual, aes(x = date, y = actual), color = &quot;#2c3e50&quot;, linewidth = 1.2) +
  geom_point(data = df_plot_actual, aes(x = date, y = actual, text = text_actual), color = &quot;#2c3e50&quot;, size = 2) +
  geom_line(data = df_plot_pred, aes(x = date, y = pred), color = &quot;#e74c3c&quot;, linetype = &quot;dashed&quot;, linewidth = 1.2) +
  geom_point(data = df_plot_pred, aes(x = date, y = pred, text = text_pred), color = &quot;#e74c3c&quot;, size = 2) +
  scale_y_continuous(labels = dollar_format(accuracy = 1)) +
  labs(
    x = &quot;&quot;, y = &quot;&quot;,
    title = paste0(
      &quot;MicroStrategy (MSTR) &lt;span style = &#039;color:#2c3e50&#039;&gt;Actual Prices&lt;/span&gt; vs &quot;,
      &quot;&lt;span style = &#039;color:#e74c3c&#039;&gt;Tidymodels Linear AI Forecast&lt;/span&gt;&lt;br&gt;&quot;,
      &quot;&lt;span style=&#039;font-size:12px; color:#555555;&#039;&gt;15-Day Trading Horizon | Out-of-Sample R-Squared: &quot;, round(rsq_val * 100, 2), &quot;%&lt;/span&gt;&quot;
    )
  ) +
  theme_minimal() +
  theme(plot.title = element_markdown(hjust = 0.5, face = &quot;bold&quot;),
        plot.background = element_rect(fill = &quot;#ffffff&quot;, color = NA),
        panel.background = element_rect(fill = &quot;#ffffff&quot;, color = NA),
        panel.grid.minor = element_blank())

# Set typography styles for HTML dashboard output
font_family &lt;- list(family = &quot;Roboto Slab, Sans-Serif&quot;, size = 16)
label_font  &lt;- list(font = list(family = &quot;Roboto Slab, Sans-Serif&quot;, size = 13))

# Convert static ggplot to fully interactive HTML Plotly widget
ggplotly(p, tooltip = &quot;text&quot;) %&gt;% 
  style(hoverlabel = label_font) %&gt;% 
  layout(font = font_family) %&gt;% 
  config(displayModeBar = FALSE)
</pre>


<h2 class="wp-block-heading">6. Price Transformation &#038; Dynamic Volatility Bands</h2>



<p class="wp-block-paragraph">Log return forecasts are <strong>converted </strong>back into actionable nominal price predictions using <strong>exponential </strong>compounding anchored to the previous trading day’s close price. The predicted price equals yesterday’s closing price <strong>multiplied by</strong> the <strong>exponential </strong>of the predicted log return.</p>



<h3 class="wp-block-heading">Empirical Confidence Envelopes (Overbought / Oversold Zones)</h3>



<p class="wp-block-paragraph">Using the empirical Root Mean Squared Error (<strong>RMSE</strong>) derived from out-of-sample evaluation, <strong>dynamic 2-sigma volatility</strong> boundaries are constructed around the predicted price path:</p>



<ul class="wp-block-list">
<li><strong>Overbought Ceiling</strong>: Calculated as yesterday’s close multiplied by the exponential of the predicted return plus twice the RMSE.</li>



<li><strong>Oversold Floor</strong>: Calculated as yesterday’s close multiplied by the exponential of the predicted return minus twice the RMSE.</li>
</ul>



<figure data-wp-context="{"imageId":"6ab6532adce04"}" data-wp-interactive="core/image" data-wp-key="6ab6532adce04" class="wp-block-image size-large wp-lightbox-container"><img loading="lazy" data-attachment-id="12794" data-permalink="https://datageeek.com/2026/09/25/quantitative-event-driven-modeling-btc-balance-sheet-shocks-with-sec-edgar-filings-in-r/secfile/" data-orig-file="https://datageeek.com/wp-content/uploads/2026/09/secfile.png" data-orig-size="1125,756" data-comments-opened="1" data-image-title="secfile" data-image-description="" data-image-caption="" data-large-file="https://i1.wp.com/datageeek.com/wp-content/uploads/2026/09/secfile.png?w=450&#038;ssl=1" data-wp-class--hide="state.isContentHidden" data-wp-class--show="state.isContentVisible" data-wp-init="callbacks.setButtonStyles" data-wp-on--click="actions.showLightbox" data-wp-on--load="callbacks.setButtonStyles" data-wp-on--pointerdown="actions.preloadImage" data-wp-on--pointerenter="actions.preloadImageWithDelay" data-wp-on--pointerleave="actions.cancelPreload" data-wp-on-window--resize="callbacks.setButtonStyles" src="https://i1.wp.com/datageeek.com/wp-content/uploads/2026/09/secfile.png?w=450&#038;ssl=1" alt="" class="wp-image-12794" srcset_temp="https://i1.wp.com/datageeek.com/wp-content/uploads/2026/09/secfile.png?w=450&#038;ssl=1 1024w, https://datageeek.com/wp-content/uploads/2026/09/secfile.png?w=150 150w, https://datageeek.com/wp-content/uploads/2026/09/secfile.png?w=300 300w, https://datageeek.com/wp-content/uploads/2026/09/secfile.png?w=768 768w, https://datageeek.com/wp-content/uploads/2026/09/secfile.png 1125w" sizes="(max-width: 1024px) 100vw, 1024px" data-recalc-dims="1" /><button
			class="lightbox-trigger"
			type="button"
			aria-haspopup="dialog"
			data-wp-bind--aria-label="state.thisImage.triggerButtonAriaLabel"
			data-wp-init="callbacks.initTriggerButton"
			data-wp-on--click="actions.showLightbox"
			data-wp-style--right="state.thisImage.buttonRight"
			data-wp-style--top="state.thisImage.buttonTop"
		>
			<svg xmlns="http://www.w3.org/2000/svg" width="12" height="12" fill="none" viewBox="0 0 12 12">
				<path fill="#fff" d="M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z" />
			</svg>
		</button></figure>



<h3 class="wp-block-heading">Visual Output Analysis</h3>



<ol class="wp-block-list">
<li><strong>Directional Accuracy</strong>: The predicted price path (<strong>dashed red line</strong>) tracks the actual price trajectory (<strong>dark solid line</strong>) with high fidelity across the 15-day evaluation window.</li>



<li><strong>Volatility Envelope Containment</strong>: Actual spot equity prices remain fully bounded within the 2-sigma gray confidence envelope, validating the structural model’s error boundary calibration.</li>



<li><strong>Price Acceleration Tracking</strong>: The sharp upward price repricing between September 14 and September 21 is effectively captured by the structural model due to the immediate integration of spot BTC log returns.</li>
</ol>



<h2 class="wp-block-heading">Conclusion</h2>



<p class="wp-block-paragraph">By combining automated <strong>SEC EDGAR parsing via secfile</strong> package with functional machine learning pipelines in tidymodels, quantitative analysts can build scalable, production-ready asset pricing engines. In the case of MicroStrategy (MSTR), incorporating SEC XBRL balance sheet tracking alongside high-frequency spot crypto returns yields an empirical out-of-sample explanatory power (R-Squared) exceeding 82%.</p>



<p class="wp-block-paragraph"></p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://datageeek.com/2026/09/25/quantitative-event-driven-modeling-btc-balance-sheet-shocks-with-sec-edgar-filings-in-r/"> DataGeeek</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/quantitative-event-driven-modeling-btc-balance-sheet-shocks-with-sec-edgar-filings-in-r/">Quantitative Event-Driven Modeling: BTC Balance Sheet Shocks with SEC EDGAR Filings in R</a>]]></content:encoded>
					
		
		<enclosure url="https://datageeek.com/wp-content/uploads/2026/09/secfile.png" length="0" type="" />
<enclosure url="https://1.gravatar.com/avatar/db5e3f9ef188ea98fe38ab05c5a3fad9fb52fe3472715a8fc02f7ea41731f77c?s=96&#038;d=identicon&#038;r=G" length="0" type="" />
<enclosure url="https://datageeek.com/wp-content/uploads/2026/09/secfile.png?w=1024" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403882</post-id>	</item>
		<item>
		<title>Data dictionaries for humans, agents and jumping frogs</title>
		<link>https://www.r-bloggers.com/2026/09/data-dictionaries-for-humans-agents-and-jumping-frogs/</link>
		
		<dc:creator><![CDATA[R-bloggers on Almost Random]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://www.malte-grosser.com/post/data-dict-jumping-frogs/</guid>

					<description><![CDATA[<p>Apparently, you can rent a frog and enter a jumping competition. You might find yourself competing against teams with years or even decades of experience choosing and preparing theirs. Welcome to the Calaveras County Jumping Frog Jubilee. In Astley et ...</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/data-dictionaries-for-humans-agents-and-jumping-frogs/">Data dictionaries for humans, agents and jumping frogs</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://www.malte-grosser.com/post/data-dict-jumping-frogs/"> R-bloggers on Almost Random</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
Apparently, you can rent a frog and enter a jumping competition. You might find yourself competing against teams with years or even decades of experience choosing and preparing theirs. Welcome to the Calaveras County Jumping Frog Jubilee. In Astley et al. (2013), researchers studied bullfrog jumps at the event, comparing frogs rented by fairgoers with those entered by experienced teams. Their measurements now provide the example for the Data Dict quickstart.
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://www.malte-grosser.com/post/data-dict-jumping-frogs/"> R-bloggers on Almost Random</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/data-dictionaries-for-humans-agents-and-jumping-frogs/">Data dictionaries for humans, agents and jumping frogs</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403821</post-id>	</item>
		<item>
		<title>We chose React to learn it. The AI wrote it for us.</title>
		<link>https://www.r-bloggers.com/2026/09/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/</link>
		
		<dc:creator><![CDATA[Arthur Bréant]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 13:26:30 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://rtask.thinkr.fr/?p=30054</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> You can read the original post in its original format on Rtask website by ThinkR here: We chose React to learn it. The AI wrote it for us.<br />
Diary of a rewrite, episode 1 of 4. In February 2026 we chose React to rebuild an internal tool, knowing full well it would take ...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/">We chose React to learn it. The AI wrote it for us.</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://rtask.thinkr.fr/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/"> Rtask</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>You can read the original post in its original format on <a rel="nofollow" href="https://rtask.thinkr.fr/" target="_blank">Rtask</a> website by ThinkR here: <a rel="nofollow" href="https://rtask.thinkr.fr/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/" target="_blank">We chose React to learn it. The AI wrote it for us.</a></p>
<p><em>Diary of a rewrite, episode 1 of 4.</em></p>
<p>In February 2026 we chose React to rebuild an internal tool, knowing full well it would take <strong>twice as long as doing it in R</strong>. It was a deliberate trade: we were not buying an application, <strong>we were buying skills</strong>.</p>
<p>In June, halfway through, with two milestones already delivered, we stopped everything and went back to R.</p>
<p>Here are the three reasons. Only one of them is “we were wrong”. The one that settled it is more uncomfortable than that: <strong>the learning we were paying a premium for had simply never happened</strong>, because the code was arriving faster than we could read it.</p>
<p>StaffPuzzle is our internal staffing tool: who works on what, which half-day. Version 1 is a read-only Shiny application, 1,200 lines, that has been running for years.</p>
<p><img loading="lazy" fetchpriority="high" decoding="async" src="https://i0.wp.com/rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v1-1024x626.png?w=450&#038;ssl=1" alt="staffpuzzle v1" class="aligncenter size-large wp-image-30049" srcset_temp="https://i0.wp.com/rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v1-1024x626.png?w=450&#038;ssl=1 1024w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v1-300x183.png 300w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v1-768x470.png 768w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v1-1536x939.png 1536w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v1-2048x1252.png 2048w" sizes="(max-width: 1024px) 100vw, 1024px" data-recalc-dims="1" /></p>
<p>Version 2 was meant to add writing to Google Calendar, a simulation mode and historised indicators. We made the stack decision twice, four months apart, and the second one undid the first. <strong>Both are written down, with their criteria. That is what makes the story tellable, and it is also, as we will see, what made it possible.</strong></p>
<p>This series covers the four months that followed, in four episodes. It ends on the real numbers: what the rewrite cost, what the hosting costs, and what the licence we avoided was worth.</p>
<h2>19 February: choosing React</h2>
<p>Three options on the table: all JavaScript (React, Supabase), all R (raw HTMX on raw plumber2), or a hybrid. The criteria table was explicit, and it is worth reading exactly as it was written:</p>
<table>
<thead>
<tr>
<th>Criterion</th>
<th>Option A: React</th>
<th>Option B: R</th>
<th>Four months on</th>
</tr>
</thead>
<tbody>
<tr>
<td>Learning modern frontend</td>
<td>Strong</td>
<td>Weak</td>
<td>never happened</td>
</tr>
<tr>
<td>Time to market</td>
<td>~19-24 half-days</td>
<td>~12 half-days</td>
<td>held</td>
</tr>
<tr>
<td>Reuse of existing R</td>
<td>None</td>
<td>Total</td>
<td>held</td>
</tr>
<tr>
<td>Transferability to other web projects</td>
<td>High</td>
<td>Low</td>
<td>wrong criterion</td>
</tr>
<tr>
<td>Rich interactivity</td>
<td>Natural</td>
<td>Difficult</td>
<td>factually false</td>
</tr>
</tbody>
</table>
<p>Look at the second row. We knew React would cost twice as much, and we chose it anyway. This is not naivety, it is a deliberate trade. The team is six people, it knows R and has almost no exposure to modern web stacks; version 1 works, there is no urgency. The Architecture Decision Record (ADR) says so plainly: the team’s comfort with React is “junior”, and that is “a deliberate investment, not an asset we already hold”.</p>
<p>In other words, the February decision was not buying an application. It was buying skills, and accepting to pay for them in delay.</p>
<blockquote><p>
  <strong>Architecture Decision Record</strong>: a document that records a significant software architecture choice and the reasoning behind it.
</p></blockquote>
<h3>The option that was not in the table</h3>
<p>You may have noticed: “stay on Shiny” is not there. The three options examined were all-JavaScript, all-R on HTMX, and a hybrid. None of them meant keeping the tool as it was and adding what it lacked.</p>
<p>This was not an oversight. The question had been settled upstream, in the project brief written a month earlier, under the heading “Technical constraints”:</p>
<blockquote><p>
  Target stack: exit Shiny, REST architecture, classic web site.
</p></blockquote>
<p>Written there, at the same level as Google API quotas and OAuth token handling. But Google’s quotas are a constraint: they are imposed on us. “Exit Shiny” is a decision. Filing a decision under constraints takes it out of the debate. You do not argue with a constraint, you work around it.</p>
<p>Nothing forced us off Shiny. Writing to a calendar, simulating, historising indicators: Shiny can do all of that. So the February ADR diligently compared three options inside a perimeter nobody had justified. And in June, when we reopened the file, we went back over every criterion one by one without ever questioning the perimeter.</p>
<blockquote><p>
  <strong>What matters for what follows.</strong> A decision justified by learning is only valid if the learning happens. That is a hypothesis, not an asset, and it can be checked.
</p></blockquote>
<h2>3 June: the U-turn</h2>
<p>Four months later, two milestones delivered in React, the team reopens the file. Out of it comes a second decision that cancels the first. Three things had moved, and they are not of the same nature.</p>
<h3>One: the option we had refused no longer existed</h3>
<p>In February, the R option on the table was raw HTMX on raw plumber2. Writing the HTML by hand, setting the attributes by hand. That was a fair criticism, and it had weighed.</p>
<p>In the meantime, two R packages had been released, <a href="https://github.com/hyperverse-r/htmxr" rel="nofollow" target="_blank"><code>{htmxr}</code></a> and <a href="https://github.com/hyperverse-r/alpiner" rel="nofollow" target="_blank"><code>{alpiner}</code></a>, which give that approach the high-level R abstractions it was missing. <strong>June’s option B is not February’s option B. That is not the same thing as having been wrong: the world had changed, and the decision had no mechanism for noticing that on its own.</strong></p>
<h3>Two: two of the arguments were simply false</h3>
<p>February’s table rated rich interactivity as “difficult” on the HTMX side, with drag and drop as the example. In June, someone went and read the React code that had actually been written: there is no drag and drop. It is a lasso selection, which is precisely what <a href="https://github.com/hyperverse-r/alpiner" rel="nofollow" target="_blank"><code>{alpiner}</code></a> does natively. The argument was defending a difficulty that did not exist.</p>
<p>The second one needs more care. The criterion said: React skills are “highly transferable to other web projects”, R skills much less so. That is true, and it is the wrong criterion. It measures what a developer carries away on their CV, not what a company accumulates. For a firm whose business is R, going deeper into R is not low transferability: it is the main asset getting thicker. Maintaining packages headed for CRAN makes us more visible and more capable than a surface knowledge of React ever would have.</p>
<h3>Three: the learning had not happened</h3>
<p><strong>This is the one that hurts, and it is the one that settled the decision.</strong> June’s ADR puts it in a single sentence:</p>
<blockquote><p>
  Learning React, the central motivation for February’s decision, is not happening in practice: the developer does not have time to read the code generated by the AI, so the skills objective is not being met.</p>
<p>  <em>ADR-003, 3 June 2026</em>
</p></blockquote>
<p>It is worth weighing what that says. The sole justification for four months of extra cost was skills. The code had been written, the application worked, the milestones were delivered. And the benefit that paid for the extra cost had not materialised.</p>
<p><strong>With an AI assisting, a stack decision justified by learning can sabotage itself.</strong> The code arrives faster than you understand it, delivery moves forward, and the illusion holds until somebody asks to see what was learned. The subject goes well beyond this article and we will come back to it; here it is enough to note that it changed a technical decision halfway through a project.</p>
<h2>What it cost: the price of the U-turn</h2>
<p>A reversal is not free, and June’s ADR lists its own downsides before concluding.</p>
<p>The first one is not technical:</p>
<ul>
<li>Redoing the communication with our internal users. <strong>Changing technical direction mid-project, in front of people waiting for a tool, requires an explanation you would rather not have to give.</strong></li>
<li>Stabilising still-experimental packages in order to put them in production.</li>
<li>The bus factor: the ecosystem we chose has one main maintainer, and it is the same person as the project’s developer.</li>
<li>Salvaging the documentary value of the React code before deleting it: the types and the tests described a data model that was still valid.</li>
</ul>
<p>The arithmetic, though, was favourable: six to eight half-days to redo everything, against eight to ten to finish in React. A proof of concept written straight afterwards, 280 lines of R reproducing the main screen at 70 milliseconds per request, turned that estimate into a commitment.</p>
<p>The numbers are what convinced us. They are not what decided it.</p>
<h2>Three months on: where we are</h2>
<p>The application runs in production, writes to the calendars, and historises its indicators. The tests-to-code ratio went from <strong>0.26 to 0.77</strong>: 544 tests, no browser, a single language from the calendar all the way to the screen.</p>
<p>That is the gain we had not anticipated. The interface having become a pure function from data to HTML, it is tested like the rest of the code, with no browser and no screenshots. We justified the return by consistency of language; what we got that was most valuable is testability.</p>
<p>And one benefit that was in no table at all: the application became the first real production case for our own packages, the ones in the <a href="https://hyperverse.world/" rel="nofollow" target="_blank">hyperverse</a>, which no demo example would ever have given us.</p>
<p><img loading="lazy" decoding="async" src="https://i2.wp.com/rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v2-1024x644.png?w=450&#038;ssl=1" alt="staffpuzzle v2" class="aligncenter size-large wp-image-30051" srcset_temp="https://i2.wp.com/rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v2-1024x644.png?w=450&#038;ssl=1 1024w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v2-300x189.png 300w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v2-768x483.png 768w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v2-1536x966.png 1536w, https://rtask.thinkr.fr/wp-content/uploads/staffpuzzle-v2-2048x1288.png 2048w" sizes="(max-width: 1024px) 100vw, 1024px" data-recalc-dims="1" /></p>
<h2>What we take away</h2>
<p><strong>Writing decisions down is what lets you undo them.</strong> That is the main lesson, and it sounds banal right up until you need it. In June, nobody had to reconstruct from memory why React had been chosen: the criteria were written, dated, weighted. It was enough to go back over them one by one and see which no longer held. Without that document, the discussion would have been about people, about who had been right, instead of about criteria.</p>
<p><strong>Separate “the world changed” from “we were wrong”.</strong> Of the three reasons, only one is an error of judgement: the drag-and-drop argument, which reading the code would have disproved back in February. The other two are of a different kind: a tool appeared, and a hypothesis about ourselves turned out to be false. Confusing the three produces either pointless guilt or an inability to turn back.</p>
<p><strong>A decision justified by a non-technical benefit must schedule its own verification.</strong> February was buying learning. Nobody had planned to check that it was arriving. We found out by accident, four months later, reopening the file for another reason. That is the one thing we would do differently: if a criterion decides, it needs a review date.</p>
<blockquote><p>
  <strong>Would we do it again.</strong> We would probably choose React again, in February 2026, with what we knew then. That is what makes the story interesting: the U-turn is not the correction of a blunder, it is how a decision behaves when you have taken the trouble to write it down.
</p></blockquote>
<hr />
<p><strong>Has your Shiny application outgrown itself, and you are wondering what to do with it?</strong> We read it and give you back a note that says so, including when the answer is “leave it alone”. <a href="https://rtask.thinkr.fr/contact/" rel="nofollow" target="_blank">Write to us</a> in two lines: the size of your application, and how long it has been running.</p>
<hr />
<p><strong>Diary of a rewrite</strong>, four episodes:</p>
<ol>
<li><strong>We chose React to learn it. The AI wrote it for us.</strong> <em>(you are here)</em></li>
<li><strong>My first attempt at porting a Shiny app failed, and not for a technical reason</strong> <em>(29 September 2026)</em></li>
<li><strong>We left the managed platform. Here is the real bill.</strong> <em>(6 October 2026)</em></li>
<li><strong>Five signs your Shiny app has outgrown Shiny</strong> <em>(13 October 2026)</em></li>
</ol>
<p><em>The two decisions cited are Architecture Decision Records from the StaffPuzzle repository, dated 19 February and 3 June 2026. The figures are measured on the repository’s branches, not estimated. Disclosure of interest: htmxr and alpiner are written by Arthur Bréant, who led this project and maintains those packages.</em></p>
<p>This post is better presented on its original ThinkR website here: <a rel="nofollow" href="https://rtask.thinkr.fr/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/" target="_blank">We chose React to learn it. The AI wrote it for us.</a></p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://rtask.thinkr.fr/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/"> Rtask</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/we-chose-react-to-learn-it-the-ai-wrote-it-for-us/">We chose React to learn it. The AI wrote it for us.</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403787</post-id>	</item>
		<item>
		<title>Más Allá del Código: muestra abierta de los proyectos de nuestros campeon(a&#124;e)s</title>
		<link>https://www.r-bloggers.com/2026/09/mas-alla-del-codigo-muestra-abierta-de-los-proyectos-de-nuestros-campeonaes/</link>
		
		<dc:creator><![CDATA[rOpenSci]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://ropensci.org/commcalls/mas-alla-del-codigo-2026/</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> How to join this free online event with Ana Carolina Moreno, Diana Garcia Cortes, Erick Navarro Delgado, Guadalupe Pascal, Valentina Clavijo Mesa, Monika Avila Marquez, Soledad Araya Orrego and Yanina Bellini Saibene.<br />
Más Allá del Código es un encuentro abierto para conocer los proyectos finales de nuestra cohorte: ...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/mas-alla-del-codigo-muestra-abierta-de-los-proyectos-de-nuestros-campeonaes/">Más Allá del Código: muestra abierta de los proyectos de nuestros campeon(a|e)s</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://ropensci.org/commcalls/mas-alla-del-codigo-2026/"> rOpenSci - open tools for open science</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<p><a href="https://ropensci.org/commcalls/mas-alla-del-codigo-2026/" rel="nofollow" target="_blank">How to join this free online event with Ana Carolina Moreno, Diana Garcia Cortes, Erick Navarro Delgado, Guadalupe Pascal, Valentina Clavijo Mesa, Monika Avila Marquez, Soledad Araya Orrego and Yanina Bellini Saibene.</a></p>
<p>Más Allá del Código es un encuentro abierto para conocer los proyectos finales de nuestra cohorte: desde la creación, mejora y revisión de paquetes de software hasta las iniciativas de divulgación que los llevaron a la comunidad. Acompañanos a descubrir el impacto técnico detrás de cada herramienta y las oportunidades que abrieron estos proyectos, incluyendo becas, nuevas colaboraciones y el crecimiento profesional de sus creadores.</p>
<p>¡Sumate a la sesión para inspirarte con sus proyectos, hacer tus preguntas en vivo y celebrar el crecimiento de nuestra comunidad de código abierto!</p>
<p>Consulta a continuación las biografías de las personas que exponen y los recursos del evento.</p>
<h2 class="title">
Speakers
</h2>
<h3>Diana Garcia</h3>
<img src="https://i2.wp.com/ropensci.org/img/team/diana-garcia.jpg?w=578&#038;ssl=1" alt="Portrait of Diana Garcia" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Bióloga Computacional del Breast Oncology Program en el Dana Farber Cancer Institute, donde investiga las alteraciones del genoma asociadas a resistencia a terapias en cáncer de mama mestastático. Tiene un doctorado en Ciencias Biomédicas y una maestría en Ciencias de la Computación y experiencia desarrollando software tanto en la academia como en la industria. Disfruta mucho enseñar programación, fue profesora en CETYS Universidad, Campus Tijuana, y tallerista en la Escuela de Código Pilares en CDMX. Forma parte de R-Ladies Boston y del programa de campeones de rOpenSci, previamente participó en PyLadies CDMX y en Women Who Code CDMX.
</p>
<h3>Erick Navarro</h3>
<img src="https://i0.wp.com/ropensci.org/img/team/erick-navarro-delgado.jpg?w=578&#038;ssl=1" alt="Portrait of Erick Navarro" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Licenciado en biología por la Universidad Nacional Autónoma de México, y candidato a Doctor en Bioinformática por The University of British Columbia. Actual campeón del Programa de rOpenSci.
</p>
<h3>Ana Carolina Moreno</h3>
<img src="https://i1.wp.com/ropensci.org/img/team/ana-carolina-moreno.jpeg?w=578&#038;ssl=1" alt="Portrait of Ana Carolina Moreno" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Periodista especializada en datos e inteligencia artificial. Cuenta con casi 20 años de experiencia en redacciones de Brasil y España, habiendo trabajado en TV Globo, G1, Folha de S. Paulo, La Voz de Galicia, Jornal da Tarde y Terra Magazine. Fundadora del capítulo de RLadies+ São Paulo
</p>
<h3>Soledad Andrea Araya Orrego</h3>
<img src="https://i2.wp.com/ropensci.org/img/team/soledad-araya.jpeg?w=578&#038;ssl=1" alt="Portrait of Soledad Andrea Araya Orrego" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Cientista política especializada en métodos cuantitativos, análisis de datos y ciencia de datos reproducible aplicada a investigación social. Trabajo con R para procesamiento, análisis y visualización de datos, automatización de flujos de trabajo y desarrollo de herramientas abiertas.
</p>
<h3>Maria Valentina Clavijo Mesa</h3>
<img src="https://i0.wp.com/ropensci.org/img/team/maria-valentina-clavijo-mesa.png?w=578&#038;ssl=1" alt="Portrait of Maria Valentina Clavijo Mesa" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Estudiante de doctorado en el Politecnico di Milano | Investigo la resiliencia de las infraestructuras críticas expuestas al cambio climático | Cofundadora de la sección de Medellín de RLadies+
</p>
<h3>Guadalupe Pascal</h3>
<img src="https://i0.wp.com/ropensci.org/img/team/guadalupe-pascal.jpg?w=578&#038;ssl=1" alt="Portrait of Guadalupe Pascal" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Soy investigadora en el ámbito de la optimización basada en datos para los procesos de toma de decisiones en empresas y sistemas sociales, todo ello desde una perspectiva regional centrada en el Sur Global. Trabajo desde una perspectiva de género e interseccional basada en los principios de la ciencia abierta y la justicia epistémica.
</p>
<h3>Monika Avila Marquez</h3>
<img src="https://i1.wp.com/ropensci.org/img/team/monika-avila-marquez.jpeg?w=578&#038;ssl=1" alt="Portrait of Monika Avila Marquez" style=" object-fit: cover; object-position: center; height: 250px; width: 200; margin-right: 15px" data-recalc-dims="1" />
<p>
Soy econometrista y estadística y me dedico a la inferencia causal a partir de datos observacionales, con especial atención a los entornos con interferencia, así como al uso de métodos de aprendizaje automático en la econometría de datos de panel. También trabajo en la selección de modelos de efectos aleatorios cruzados para datos experimentales.
</p>
<h2 class="title">Resources</h2>
<ul class="resources-list">
<li>
<a href="https://github.com/snaraya/votosCL" rel="nofollow" target="_blank">paquete votosCL que tiene como objetivo facilitar el almacenamiento y manejo de los datos electorales de Chile.</a>
</li>
<li>
<a href="https://github.com/ropensci/software-review/issues/743" rel="nofollow" target="_blank">Issue con la revision por pares del paquete RAMEN</a>
</li>
<li>
<a href="https://github.com/ropensci/RAMEN" rel="nofollow" target="_blank">Regional Association of DNA Methylome variability with the Exposome and geNome (RAMEN)</a>
</li>
<li>
<a href="https://docs.google.com/document/d/16FA9ywfg7jtJWq325mnCAPT5ejFVqcsEYi-q1yHBEsI/edit?usp=sharing" rel="nofollow" target="_blank">Collaborative notes</a>
</li>
</ul>
</div>
<h2>Join Us!</h2>
<ul>
<li>
<strong>Who</strong>
Everyone is welcome. No RSVP needed, simply connect and/or dial in at the time of the event.
</li>
<li>
<strong>When</strong>
Monday, 12 October 2026 08:00 PDT (Monday, 12 October 2026 15:00 UTC)
</li>
<li>
<a href="https://www.timeanddate.com/worldclock/fixedtime.html?iso=2026-10-12T15%3A00%3A00&#038;ah=1&#038;msg=M%C3%A1s%20All%C3%A1%20del%20C%C3%B3digo%3A%20muestra%20abierta%20de%20los%20proyectos%20de%20nuestros%20campeon%28a%7Ce%29s%20(rOpenSci%20Comm%20Call)" rel="nofollow" target="_blank">Find your timezone</a>
</li>
<li>
<a href="https://ropensci.org/commcalls/mas-alla-del-codigo-2026//index.ics" rel="nofollow" target="_blank">Add to Calendar.</a>
</li>
<li>
<strong>How</strong>
<p>Everyone is welcome. No RSVP needed.</p>
<p>Test your Zoom setup <a href ="https://zoom.us/test">https://zoom.us/test</a>.</p>
<p>Join the meeting: <a href="https://numfocus-org.zoom.us/j/86115160074?pwd=5aGfse7U7xxrABA5lmgdYRRD8hdEtx.1" rel="nofollow" target="_blank">https://numfocus-org.zoom.us/j/86115160074?pwd=5aGfse7U7xxrABA5lmgdYRRD8hdEtx.1</a></p>
ID de la reunión: 86115160074
Código de acceso: 689296
<p>Find your local number to join by phone: <a href="https://zoom.us/u/adAyZGMYrE" rel="nofollow" target="_blank">https://zoom.us/u/adAyZGMYrE</a></p>
</li>
</ul>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://ropensci.org/commcalls/mas-alla-del-codigo-2026/"> rOpenSci - open tools for open science</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/mas-alla-del-codigo-muestra-abierta-de-los-proyectos-de-nuestros-campeonaes/">Más Allá del Código: muestra abierta de los proyectos de nuestros campeon(a|e)s</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403953</post-id>	</item>
		<item>
		<title>Friedman test in R, or the nonparametric version of the repeated measures ANOVA</title>
		<link>https://www.r-bloggers.com/2026/09/friedman-test-in-r-or-the-nonparametric-version-of-the-repeated-measures-anova/</link>
		
		<dc:creator><![CDATA[R on Stats and R]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; ">
<p>Introduction<br />
Data<br />
Friedman test</p>
<p>Aim and hypotheses<br />
Assumptions<br />
In R</p>
<p>With base R<br />
With the {rstatix} package</p>
<p>Interpretations</p>
<p>Post-hoc tests</p>
<p>Pairwise Wilcoxon signed-rank tests<br />
Nemenyi test<br />
Conover test</p>
<p>Combination of statistical results and plo...</p></div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/friedman-test-in-r-or-the-nonparametric-version-of-the-repeated-measures-anova/">Friedman test in R, or the nonparametric version of the repeated measures ANOVA</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/"> R on Stats and R</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>


<div id="TOC">
<ul>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#introduction" id="toc-introduction" rel="nofollow" target="_blank">Introduction</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#data" id="toc-data" rel="nofollow" target="_blank">Data</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#friedman-test" id="toc-friedman-test" rel="nofollow" target="_blank">Friedman test</a>
<ul>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#aim-and-hypotheses" id="toc-aim-and-hypotheses" rel="nofollow" target="_blank">Aim and hypotheses</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#assumptions" id="toc-assumptions" rel="nofollow" target="_blank">Assumptions</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#in-r" id="toc-in-r" rel="nofollow" target="_blank">In R</a>
<ul>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#with-base-r" id="toc-with-base-r" rel="nofollow" target="_blank">With base R</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#with-the-rstatix-package" id="toc-with-the-rstatix-package" rel="nofollow" target="_blank">With the {rstatix} package</a></li>
</ul></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#interpretations" id="toc-interpretations" rel="nofollow" target="_blank">Interpretations</a></li>
</ul></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#post-hoc-tests" id="toc-post-hoc-tests" rel="nofollow" target="_blank">Post-hoc tests</a>
<ul>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#pairwise-wilcoxon-signed-rank-tests" id="toc-pairwise-wilcoxon-signed-rank-tests" rel="nofollow" target="_blank">Pairwise Wilcoxon signed-rank tests</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#nemenyi-test" id="toc-nemenyi-test" rel="nofollow" target="_blank">Nemenyi test</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#conover-test" id="toc-conover-test" rel="nofollow" target="_blank">Conover test</a></li>
</ul></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#combination-of-statistical-results-and-plot" id="toc-combination-of-statistical-results-and-plot" rel="nofollow" target="_blank">Combination of statistical results and plot</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#summary" id="toc-summary" rel="nofollow" target="_blank">Summary</a></li>
<li><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#references" id="toc-references" rel="nofollow" target="_blank">References</a></li>
</ul>
</div>

<p><img src="https://i2.wp.com/statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/images/friedman-test-nonparametric-version-repeated-measures-anova.jpg?w=578&#038;ssl=1" style="width:100.0%" data-recalc-dims="1" /></p>
<div id="introduction" class="section level1">
<h1>Introduction</h1>
<p>In a previous article, we showed how to perform a <a href="https://statsandr.com/blog/repeated-measures-anova-in-r/" rel="nofollow" target="_blank">repeated measures ANOVA in R</a> to compare a <a href="https://statsandr.com/blog/variable-types-and-examples/#quantitative" rel="nofollow" target="_blank">quantitative variable</a> measured on the same subjects under three or more related conditions, or at three or more points in time.</p>
<p>As for many <a href="https://statsandr.com/blog/what-statistical-test-should-i-do/" rel="nofollow" target="_blank">statistical tests</a>, its results can only be trusted if some assumptions are met, in particular the <a href="https://statsandr.com/blog/do-my-data-follow-a-normal-distribution-a-note-on-the-most-widely-used-distribution-and-how-to-test-for-normality-in-r/" rel="nofollow" target="_blank">normality</a> of the residuals (at least for small samples) and the <strong>sphericity</strong> of the data. When they are not, or when the dependent variable is only ordinal, its nonparametric version can be used: the <strong>Friedman test</strong>, proposed by the economist Milton Friedman in 1937 in a paper whose title sums up its purpose: “The use of ranks to avoid the assumption of normality implicit in the analysis of variance” <span class="citation">(<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#ref-friedman1937use" rel="nofollow" target="_blank">Friedman 1937</a>)</span>.</p>
<p>The article about the <a href="https://statsandr.com/blog/kruskal-wallis-test-nonparametric-version-anova/" rel="nofollow" target="_blank">Kruskal-Wallis test</a>, the nonparametric version of the <a href="https://statsandr.com/blog/anova-in-r/" rel="nofollow" target="_blank">one-way ANOVA</a>, already pointed in this direction: if the observations between samples are dependent (for instance, the same individuals measured before, during and after a treatment), the Friedman test should be preferred to take this dependency into account. The present article picks up that thread.</p>
<p>These four tests, all used to compare three groups or more, form two pairs of parametric and nonparametric tests:</p>
<table>
<colgroup>
<col width="11%" />
<col width="48%" />
<col width="39%" />
</colgroup>
<thead>
<tr class="header">
<th></th>
<th><strong>Independent samples</strong></th>
<th><strong>Related samples</strong></th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><strong>Parametric</strong></td>
<td><a href="https://statsandr.com/blog/anova-in-r/" rel="nofollow" target="_blank">One-way ANOVA</a></td>
<td><a href="https://statsandr.com/blog/repeated-measures-anova-in-r/" rel="nofollow" target="_blank">Repeated measures ANOVA</a></td>
</tr>
<tr class="even">
<td><strong>Nonparametric</strong></td>
<td><a href="https://statsandr.com/blog/kruskal-wallis-test-nonparametric-version-anova/" rel="nofollow" target="_blank">Kruskal-Wallis test</a></td>
<td>Friedman test</td>
</tr>
</tbody>
</table>
<p>In other words, the Friedman test is the “related-samples” counterpart of the Kruskal-Wallis test. Both work on ranks instead of raw values, but the Kruskal-Wallis test ranks all observations together, whereas the Friedman test ranks the measurements <strong>within each subject</strong>, which is how it takes the dependency between the samples into account.</p>
<p>In the rest of the article, we show how to perform and interpret the Friedman test in R, how to follow it up with post-hoc tests, and how to present all the results on a single plot.</p>
</div>
<div id="data" class="section level1">
<h1>Data</h1>
<p>We use the same scenario as in the <a href="https://statsandr.com/blog/repeated-measures-anova-in-r/#data" rel="nofollow" target="_blank">repeated measures ANOVA article</a>: a treatment against chronic pain, with pain measured from 0 (no pain at all) to 100 (unbearable pain) on the same patients (i) before, (ii) during and (iii) one month after the treatment. This time, however, the data come from a small pilot study with 16 patients, and some scores are affected by flare-ups:</p>
<pre># number of patients
n &lt;- 16

# each patient has its own baseline level of pain
patient_effect &lt;- rnorm(n, mean = 0, sd = 6)

# pain score at the 3 moments, with right-skewed fluctuations (flare-ups)
before &lt;- 60 + patient_effect + rexp(n, rate = 1 / 6)
during &lt;- 48 + patient_effect + rexp(n, rate = 1 / 6)
after &lt;- 46 + patient_effect + rexp(n, rate = 1 / 6)

# dataset in the long format
dat &lt;- data.frame(
  patient = factor(rep(1:n, times = 3)),
  time = factor(rep(c(&quot;before&quot;, &quot;during&quot;, &quot;after&quot;), each = n),
    levels = c(&quot;before&quot;, &quot;during&quot;, &quot;after&quot;)
  ),
  pain = round(c(before, during, after), 1)
)

head(dat)
##   patient   time pain
## 1       1 before 75.7
## 2       2 before 58.8
## 3       3 before 91.4
## 4       4 before 67.8
## 5       5 before 92.4
## 6       6 before 60.7</pre>
<p>(A seed has been set in the background with <code>set.seed(42)</code>, so the data and all the results below are reproducible.)</p>
<p>As before, the term <code>patient_effect</code> creates the dependency between the three scores of a given patient, and the data are in the <strong>long format</strong> (one row per patient and per moment). The difference lies in the fluctuations around the usual level of each patient, drawn from a strongly right-skewed exponential distribution (<code>rexp()</code>) instead of a normal one. This mimics flare-ups: most patients report a score close to their usual level, but a few of them report a much higher one. With so few patients, normality cannot be taken for granted.</p>
<p>Since we are going to use a test based on ranks, the <a href="https://statsandr.com/blog/descriptive-statistics-in-r/#median" rel="nofollow" target="_blank">median</a> and the <a href="https://statsandr.com/blog/descriptive-statistics-in-r/#interquartile-range" rel="nofollow" target="_blank">interquartile range</a> are the most relevant <a href="https://statsandr.com/blog/descriptive-statistics-in-r/" rel="nofollow" target="_blank">descriptive statistics</a>, but the <a href="https://statsandr.com/blog/descriptive-statistics-in-r/#mean" rel="nofollow" target="_blank">mean</a> is shown for comparison:</p>
<pre># install.packages(&quot;dplyr&quot;)
library(dplyr)

dat %&gt;%
  group_by(time) %&gt;%
  summarise(
    n = n(),
    median = median(pain),
    IQR = IQR(pain),
    mean = round(mean(pain), 2)
  ) %&gt;%
  as.data.frame() # to print all decimals
##     time  n median    IQR  mean
## 1 before 16  68.50 15.325 71.14
## 2 during 16  57.45 12.025 59.45
## 3  after 16  52.10  7.475 55.54</pre>
<p>The <a href="https://statsandr.com/blog/descriptive-statistics-in-r/#boxplot" rel="nofollow" target="_blank">boxplots</a> compare the three moments, and the second plot joins the successive scores of each patient:</p>
<pre># install.packages(&quot;ggplot2&quot;)
library(ggplot2)

ggplot(dat) +
  aes(x = time, y = pain, fill = time) +
  geom_boxplot() +
  theme(legend.position = &quot;none&quot;) +
  labs(
    x = &quot;Moment of the measurement&quot;,
    y = &quot;Pain score&quot;
  )</pre>
<p><img src="https://i1.wp.com/statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/index_files/figure-html/unnamed-chunk-3-1.png?w=450&#038;ssl=1" alt="" style="display: block; margin: auto;" data-recalc-dims="1" /></p>
<pre>ggplot(dat) +
  aes(x = time, y = pain, group = patient) +
  geom_line(alpha = 0.4) +
  geom_point(alpha = 0.4) +
  labs(
    x = &quot;Moment of the measurement&quot;,
    y = &quot;Pain score&quot;
  )</pre>
<p><img src="https://i2.wp.com/statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/index_files/figure-html/unnamed-chunk-3-2.png?w=450&#038;ssl=1" alt="" style="display: block; margin: auto;" data-recalc-dims="1" /></p>
<p>In our <a href="https://statsandr.com/blog/what-is-the-difference-between-population-and-sample/" rel="nofollow" target="_blank">sample</a>, the median pain score decreased from 68.5 before the treatment to 57.45 during the treatment and 52.1 one month after it, and most patients follow this downward trend. Both plots also reveal two isolated high scores (one during and one after the treatment) due to flare-ups. (The low score flagged after the treatment is less extreme, and simply reflects a patient who responded particularly well.) These flare-ups pull the mean upwards (55.5 versus a median of 52.1 one month after the treatment) and would weigh heavily on a test based on means, whereas in a test based on ranks they simply count as the highest score of the patients concerned.</p>
<p>Only a sound statistical test will tell us whether these differences can be generalized to the <a href="https://statsandr.com/blog/what-is-the-difference-between-population-and-sample/" rel="nofollow" target="_blank">population</a>.</p>
</div>
<div id="friedman-test" class="section level1">
<h1>Friedman test</h1>
<div id="aim-and-hypotheses" class="section level2">
<h2>Aim and hypotheses</h2>
<p>The Friedman test is used to compare <span class="math inline">\(k \geq 3\)</span> related conditions (or points in time) in terms of a quantitative or <a href="https://statsandr.com/blog/variable-types-and-examples/#ordinal" rel="nofollow" target="_blank">ordinal</a> variable measured on the same subjects. In our example, it helps us to answer the question: “Is the pain score different before, during and after the treatment?”.</p>
<p>Its principle is simple: the <span class="math inline">\(k\)</span> measurements of each subject are ranked from 1 (the smallest) to <span class="math inline">\(k\)</span> (the largest), with tied values receiving the average of their ranks, and the ranks are then summed for each condition. If the conditions do not differ, these sums, denoted <span class="math inline">\(R_1, \dots, R_k\)</span>, should all be close to <span class="math inline">\(n(k+1)/2\)</span>, where <span class="math inline">\(n\)</span> is the number of subjects. The test statistic measures how far they are from this value:</p>
<p><span class="math display">\[Q = \frac{12}{n k (k+1)} \sum_{j=1}^k R_j^2 - 3n(k+1)\]</span></p>
<p>Under the null hypothesis, and if <span class="math inline">\(n\)</span> is not too small, <span class="math inline">\(Q\)</span> follows approximately a chi-square distribution with <span class="math inline">\(k - 1\)</span> degrees of freedom (with a correction in case of ties).<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#fn1" class="footnote-ref" id="fnref1" rel="nofollow" target="_blank"><sup>1</sup></a> Since the ranking is done within each subject, the general level of each patient plays no role, just as the repeated measures ANOVA removes the variability between subjects from its error term.</p>
<p>The null and alternative hypotheses of the Friedman test are:</p>
<ul>
<li><span class="math inline">\(H_0\)</span>: the <span class="math inline">\(k\)</span> related conditions have the same distribution, so no condition tends to receive higher or lower ranks than the others</li>
<li><span class="math inline">\(H_1\)</span>: at least one condition is different from the others</li>
</ul>
<p>In our example, <span class="math inline">\(H_0\)</span> means that the pain score has the same distribution before, during and after the treatment, and <span class="math inline">\(H_1\)</span> that at least one of the three moments differs from the other two.</p>
<p>Be careful that, as for the ANOVA and the Kruskal-Wallis test, the alternative hypothesis is <strong><em>not</em></strong> that all conditions are different from each other. If the null hypothesis is rejected, we only know that <em>at least</em> one moment differs from the others; post-hoc tests, covered later, tell us which ones.</p>
<p>Note also that the Friedman test is often presented as a comparison of medians. This shortcut is only valid under additional assumptions (in particular, distributions with the same shape that only differ by a shift), which are not needed as long as the test is used to compare conditions in general, as we do here.</p>
</div>
<div id="assumptions" class="section level2">
<h2>Assumptions</h2>
<p>First, the Friedman test requires one dependent variable, at least ordinal, measured on the same subjects across <span class="math inline">\(k \geq 3\)</span> related conditions or points in time: a within-subjects design in which each subject is measured once in each condition (subjects with a missing measurement are removed by <code>friedman.test()</code>). Here, the pain score is measured at 3 moments on the same 16 patients, so this assumption is met.<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#fn2" class="footnote-ref" id="fnref2" rel="nofollow" target="_blank"><sup>2</sup></a></p>
<p>Second, the subjects must be <strong>independent from each other</strong>. Independence is required between subjects, but not within them: the <span class="math inline">\(k\)</span> measurements of a given patient are of course dependent (this is the whole point of the design, and exactly what the ranking within subjects accounts for), but the scores of one patient must not influence those of another. This is verified based on the design of the study: here, patients have been selected at random and treated individually.</p>
<p>Third, as a nonparametric test, the Friedman test does <strong>not require normality</strong>, and since it compares neither means nor variances of differences, it does not require sphericity either. This is precisely why it is preferred over the repeated measures ANOVA when the <a href="https://statsandr.com/blog/repeated-measures-anova-in-r/#assumptions" rel="nofollow" target="_blank">normality or sphericity assumptions</a> of the latter are not satisfied.</p>
<p>In our example, normality can be assessed on the residuals of a model including the moment and the patient, as in the <a href="https://statsandr.com/blog/repeated-measures-anova-in-r/#normality" rel="nofollow" target="_blank">repeated measures ANOVA article</a>:</p>
<pre># residuals of the model, taking the patient effect into account
res_lm &lt;- lm(pain ~ time + patient, data = dat)

# install.packages(&quot;car&quot;)
library(car)

qqPlot(residuals(res_lm),
  id = FALSE # remove point identification
)</pre>
<p><img src="https://i0.wp.com/statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/index_files/figure-html/unnamed-chunk-4-1.png?w=450&#038;ssl=1" alt="" style="display: block; margin: auto;" data-recalc-dims="1" /></p>
<pre>shapiro.test(residuals(res_lm))
## 
## 	Shapiro-Wilk normality test
## 
## data:  residuals(res_lm)
## W = 0.89685, p-value = 0.0004984</pre>
<p>Several points of the <a href="https://statsandr.com/blog/descriptive-statistics-in-r/#qq-plot" rel="nofollow" target="_blank">QQ-plot</a> lie far from the straight line and outside the confidence bands, and the Shapiro-Wilk test rejects the normality of the residuals (<em>p</em>-value < 0.05). With only 16 patients, the central limit theorem cannot be relied upon to make this deviation harmless, so the Friedman test is the appropriate choice.</p>
<p>This flexibility comes at a price, though: since only the order of the measurements within each subject is used, the Friedman test is generally less powerful than the repeated measures ANOVA when the assumptions of the latter are met. It is an alternative, not a default choice.</p>
</div>
<div id="in-r" class="section level2">
<h2>In R</h2>
<div id="with-base-r" class="section level3">
<h3>With base R</h3>
<p>In R, the Friedman test is done with the <code>friedman.test()</code> function. With data in the long format, the easiest is to use its formula interface, <code>dependent variable ~ conditions | subjects</code>:</p>
<pre>friedman.test(pain ~ time | patient,
  data = dat
)
## 
## 	Friedman rank sum test
## 
## data:  pain and time and patient
## Friedman chi-squared = 17.375, df = 2, p-value = 0.0001687</pre>
<p>It also accepts a matrix with one row per subject and one column per condition (the wide format), which can be obtained with the <code>pivot_wider()</code> function of the <code>{tidyr}</code> package:</p>
<pre># install.packages(&quot;tidyr&quot;)
library(tidyr)

# from the long format to the wide format
dat_wide &lt;- dat %&gt;%
  pivot_wider(names_from = time, values_from = pain)

head(dat_wide)
## # A tibble: 6 × 4
##   patient before during after
##   &lt;fct&gt;    &lt;dbl&gt;  &lt;dbl&gt; &lt;dbl&gt;
## 1 1         75.7   66.8  55.2
## 2 2         58.8   49    52.5
## 3 3         91.4   50.4  55  
## 4 4         67.8   53.9  51.5
## 5 5         92.4   52.8  51.7
## 6 6         60.7   58.8  46.1
# Friedman test on the 3 columns of scores (without the patient column)
friedman.test(as.matrix(dat_wide[, -1]))
## 
## 	Friedman rank sum test
## 
## data:  as.matrix(dat_wide[, -1])
## Friedman chi-squared = 17.375, df = 2, p-value = 0.0001687</pre>
<p>Both give the same results: the test statistic <span class="math inline">\(Q\)</span> (<code>Friedman chi-squared</code>), the number of degrees of freedom (<code>df</code>, equal to <span class="math inline">\(k - 1 = 2\)</span>) and the <a href="https://statsandr.com/blog/student-s-t-test-in-r-and-by-hand-how-to-compare-two-groups-under-different-scenarios/#a-note-on-p-value-and-significance-level-alpha" rel="nofollow" target="_blank"><em>p</em>-value</a>, which we interpret in the next section.</p>
<p>For the curious reader, the test statistic can easily be recomputed from the ranks within patients:</p>
<pre># rank the 3 scores within each patient
ranks &lt;- t(apply(dat_wide[, -1], 1, rank))

# sum of the ranks for each moment
R &lt;- colSums(ranks)
R
## before during  after 
##     45     29     22
# test statistic
k &lt;- 3
12 / (n * k * (k + 1)) * sum(R^2) - 3 * n * (k + 1)
## [1] 17.375</pre>
<p>Under the null hypothesis, each sum of ranks would be close to <span class="math inline">\(n(k+1)/2 = 32\)</span>. Here, the moment before the treatment collects much higher ranks, and we find the same statistic as <code>friedman.test()</code>.</p>
</div>
<div id="with-the-rstatix-package" class="section level3">
<h3>With the {rstatix} package</h3>
<p>For an output consistent with the tidyverse, the <code>friedman_test()</code> function of the <code>{rstatix}</code> package uses the same formula and returns a data frame:</p>
<pre># install.packages(&quot;rstatix&quot;)
library(rstatix)

dat %&gt;%
  friedman_test(pain ~ time | patient)
## # A tibble: 1 × 6
##   .y.       n statistic    df        p method       
## * &lt;chr&gt; &lt;int&gt;     &lt;dbl&gt; &lt;dbl&gt;    &lt;dbl&gt; &lt;chr&gt;        
## 1 pain     16      17.4     2 0.000169 Friedman test</pre>
<p>The same package provides an effect size, Kendall’s <span class="math inline">\(W\)</span>:</p>
<pre>dat %&gt;%
  friedman_effsize(pain ~ time | patient)
## # A tibble: 1 × 5
##   .y.       n effsize method    magnitude
## * &lt;chr&gt; &lt;int&gt;   &lt;dbl&gt; &lt;chr&gt;     &lt;ord&gt;    
## 1 pain     16   0.543 Kendall W large</pre>
<p>Kendall’s <span class="math inline">\(W\)</span>, computed as <span class="math inline">\(Q / (n(k - 1))\)</span>, ranges from 0 (no consistent ordering of the conditions across subjects) to 1 (all subjects rank the conditions in the same order). <code>{rstatix}</code> interprets it with the usual guidelines: 0.1 to < 0.3 for a small effect, 0.3 to < 0.5 for a moderate effect, and 0.5 or more for a large effect. Here, <span class="math inline">\(W =\)</span> 0.54, so a large effect.</p>
</div>
</div>
<div id="interpretations" class="section level2">
<h2>Interpretations</h2>
<p>The <em>p</em>-value is smaller than the significance level <span class="math inline">\(\alpha = 0.05\)</span>, so we reject the null hypothesis and we conclude that the pain score is not the same at the three moments (<span class="math inline">\(\chi^2(2) = 17.38\)</span>, <em>p</em>-value < 0.001, Kendall’s <span class="math inline">\(W = 0.54\)</span>).</p>
<p>(<em>For the sake of illustration</em>, with a <em>p</em>-value larger than 0.05, we could not have rejected the null hypothesis, and thus could not have concluded that the pain score changed over time.)</p>
<p>As any omnibus test, however, the Friedman test does not tell us which moments differ.</p>
</div>
</div>
<div id="post-hoc-tests" class="section level1">
<h1>Post-hoc tests</h1>
<p>To find out, we need post-hoc tests (in Latin, “after this”, so after a significant Friedman test), which compare the conditions two by two while controlling for multiple comparisons (see more details in the <a href="https://statsandr.com/blog/anova-in-r/#post-hoc-test" rel="nofollow" target="_blank">one-way ANOVA</a> article). Since the data are related, these comparisons must take the within-subject structure into account: whereas the Dunn test follows a Kruskal-Wallis test, the most common post-hoc tests after a Friedman test are:</p>
<ul>
<li>pairwise Wilcoxon signed-rank tests</li>
<li>the Nemenyi test</li>
<li>the Conover test</li>
</ul>
<p>Unlike pairwise Wilcoxon tests, which only use the two conditions being compared, the Nemenyi and Conover tests compare the rank sums of the Friedman test itself, computed on all conditions at once. This makes them more consistent with the omnibus test, but it also has a drawback: the conclusion for a given pair of conditions can change depending on which other conditions are included in the study. For this reason, some authors, such as <span class="citation">Benavoli et al. (<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#ref-benavoli2016should" rel="nofollow" target="_blank">2016</a>)</span>, recommend pairwise Wilcoxon signed-rank tests instead. In practice, both approaches are widely used, and they often lead to the same conclusions, as in our example below.</p>
<div id="pairwise-wilcoxon-signed-rank-tests" class="section level2">
<h2>Pairwise Wilcoxon signed-rank tests</h2>
<p>The <a href="https://statsandr.com/blog/wilcoxon-test-in-r-how-to-compare-2-groups-under-the-non-normality-assumption/#paired-samples" rel="nofollow" target="_blank">Wilcoxon signed-rank test</a> compares two paired samples. It belongs to the same family of rank-based tests as the Friedman test, and it is actually a <a href="https://statsandr.com/blog/one-sample-wilcoxon-test-in-r/" rel="nofollow" target="_blank">one-sample Wilcoxon test</a> applied to the differences within subjects. Pairwise tests with the Holm adjustment are obtained with <code>pairwise_wilcox_test()</code> and the argument <code>paired = TRUE</code>:<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#fn3" class="footnote-ref" id="fnref3" rel="nofollow" target="_blank"><sup>3</sup></a></p>
<pre>dat %&gt;%
  pairwise_wilcox_test(pain ~ time,
    paired = TRUE,
    p.adjust.method = &quot;holm&quot;
  )
## # A tibble: 3 × 9
##   .y.   group1 group2    n1    n2 statistic       p  p.adj p.adj.signif
## * &lt;chr&gt; &lt;chr&gt;  &lt;chr&gt;  &lt;int&gt; &lt;int&gt;     &lt;dbl&gt;   &lt;dbl&gt;  &lt;dbl&gt; &lt;chr&gt;       
## 1 pain  before during    16    16       117 0.00919 0.0184 *           
## 2 pain  before after     16    16       122 0.00336 0.0101 *           
## 3 pain  during after     16    16        95 0.175   0.175  ns</pre>
<p>(With <code>paired = TRUE</code>, observations are paired according to their order in the dataset, so the patients must appear in the same order at each moment, which is the case here.)</p>
<p>The base R function <code>pairwise.wilcox.test()</code> gives the same <em>p</em>-values:</p>
<pre>pairwise.wilcox.test(dat$pain, dat$time,
  paired = TRUE,
  p.adjust.method = &quot;holm&quot;
)
## 
## 	Pairwise comparisons using Wilcoxon signed rank exact test 
## 
## data:  dat$pain and dat$time 
## 
##        before during
## during 0.018  -     
## after  0.010  0.175 
## 
## P value adjustment method: holm</pre>
<p>Comparing the adjusted <em>p</em>-values (the <code>p.adj</code> column of the first output) to the 5% significance level, we conclude that:</p>
<ul>
<li>the pain score differs significantly between before and during the treatment (adjusted <em>p</em>-value = 0.018),</li>
<li>it differs significantly between before the treatment and one month after it (adjusted <em>p</em>-value = 0.010), and</li>
<li>it does not differ significantly between during the treatment and one month after it (adjusted <em>p</em>-value = 0.175).</li>
</ul>
<p>In other words, the treatment is associated with a significant decrease of the pain score (median from 68.5 to 57.45), which persists one month after the treatment (median of 52.1), while the further decrease after the treatment is not significant.</p>
</div>
<div id="nemenyi-test" class="section level2">
<h2>Nemenyi test</h2>
<p>The Nemenyi test is available in the <code>{PMCMRplus}</code> package:</p>
<pre># install.packages(&quot;PMCMRplus&quot;)
library(PMCMRplus)

frdAllPairsNemenyiTest(pain ~ time | patient,
  data = dat
)
##        before  during 
## during 0.01299 -      
## after  0.00014 0.43104</pre>
<p>These <em>p</em>-values (based on the studentized range distribution, which already accounts for multiple comparisons) lead to the same conclusions: the moment before the treatment differs significantly from both other moments, while during and after do not differ significantly.</p>
</div>
<div id="conover-test" class="section level2">
<h2>Conover test</h2>
<p>Like the Nemenyi test, the Conover test compares the rank sums of the Friedman test, but it relies on a Student’s <em>t</em> distribution with an error term estimated from the ranks, which generally makes it more powerful than the Nemenyi test. It is available in the same package, with the <code>frdAllPairsConoverTest()</code> function. By default, this function uses a single-step adjustment based on the studentized range distribution (as for the Nemenyi test), but other adjustments can be chosen with the <code>p.adjust.method</code> argument, in which case the <em>p</em>-values are computed from the <em>t</em> distribution and then adjusted. For consistency with the pairwise Wilcoxon tests, we use the Holm method. Note that, unlike <code>frdAllPairsNemenyiTest()</code>, this function does not accept a formula, so the response, the conditions and the subjects are passed as separate vectors:</p>
<pre>frdAllPairsConoverTest(
  y = dat$pain,
  groups = dat$time,
  blocks = dat$patient,
  p.adjust.method = &quot;holm&quot;
)
##        before  during 
## during 0.00066 -      
## after  6.9e-06 0.08650</pre>
<p>These adjusted <em>p</em>-values lead to the same conclusions: the pain score before the treatment differs significantly from the pain score during and one month after the treatment, while the pain scores during and after the treatment do not differ significantly.</p>
<p>Note that you may also come across this test under the name <strong>Durbin-Conover test</strong>, for instance in the <code>{ggstatsplot}</code> package used in the next section. The Durbin test is a generalization of the Friedman test to designs in which each subject is measured under only some of the conditions (the so-called balanced incomplete block designs), and the Durbin-Conover test is its post-hoc test. When all subjects are measured under all conditions, as in our example, the Durbin test reduces to the Friedman test and the Durbin-Conover test reduces to the Conover test: despite the different names, the computations are exactly the same (the <code>durbinAllPairsTest()</code> function of the <code>{PMCMRplus}</code> package would give the same results as above).</p>
</div>
</div>
<div id="combination-of-statistical-results-and-plot" class="section level1">
<h1>Combination of statistical results and plot</h1>
<p>As for the Kruskal-Wallis test, the <code>{ggstatsplot}</code> package can display the data and all the statistical results on a single plot, here with <code>ggwithinstats()</code>, designed for within-subjects designs:</p>
<pre># install.packages(&quot;ggstatsplot&quot;)
library(ggstatsplot)

ggwithinstats(
  data = dat,
  x = time,
  y = pain,
  subject.id = patient,
  type = &quot;nonparametric&quot;, # repeated measures ANOVA or Friedman test
  pairwise.display = &quot;all&quot;, # display all pairwise comparisons
  bf.message = FALSE
)</pre>
<p><img src="https://i1.wp.com/statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/index_files/figure-html/unnamed-chunk-14-1.png?w=450&#038;ssl=1" alt="" style="display: block; margin: auto;" data-recalc-dims="1" /></p>
<p>The results of the Friedman test are shown in the subtitle (the <em>p</em>-value is after <code>p =</code>), the boxplots and violin plots show the distribution at each moment, and the post-hoc tests are displayed between each pair of moments. For these post-hoc tests, <code>ggwithinstats()</code> uses the Durbin-Conover test with the Holm adjustment which, as explained above, is the Conover test under another name. The adjusted <em>p</em>-values displayed on the plot are therefore exactly those obtained with <code>frdAllPairsConoverTest()</code> in the previous section.</p>
</div>
<div id="summary" class="section level1">
<h1>Summary</h1>
<p>In this article, we presented the Friedman test, the nonparametric version of the repeated measures ANOVA, used to compare three or more related conditions in terms of a quantitative or ordinal variable. Since it ranks the measurements within each subject, it requires independence between subjects but neither normality nor sphericity. In R, it is performed with <code>friedman.test()</code> or <code>friedman_test()</code> from <code>{rstatix}</code>, and a significant result, which only indicates that at least one condition differs, is followed by post-hoc tests such as pairwise Wilcoxon signed-rank tests with adjusted <em>p</em>-values. If you hesitate between several methods, see this <a href="https://statsandr.com/blog/what-statistical-test-should-i-do/" rel="nofollow" target="_blank">overview of the most common statistical tests</a>.</p>
<p>Thanks for reading.</p>
<p>As always, if you have a question or a suggestion related to the topic covered in this article, please add it as a comment so other readers can benefit from the discussion.</p>
</div>
<div id="references" class="section level1 unnumbered">
<h1>References</h1>
<div id="refs" class="references csl-bib-body hanging-indent">
<div id="ref-benavoli2016should" class="csl-entry">
Benavoli, Alessio, Giorgio Corani, and Francesca Mangili. 2016. <span>“Should We Really Use Post-Hoc Tests Based on Mean-Ranks?”</span> <em>The Journal of Machine Learning Research</em> 17 (1): 152–61.
</div>
<div id="ref-friedman1937use" class="csl-entry">
Friedman, Milton. 1937. <span>“The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance.”</span> <em>Journal of the American Statistical Association</em> 32 (200): 675–701.
</div>
</div>
</div>
<div class="footnotes footnotes-end-of-document">
<hr />
<ol>
<li id="fn1"><p>A common rule of thumb is <span class="math inline">\(n > 15\)</span> or <span class="math inline">\(k > 4\)</span>. For smaller samples, exact tables or permutation <em>p</em>-values are more appropriate.<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#fnref1" class="footnote-back" rel="nofollow" target="_blank"><img src="https://s.w.org/images/core/emoji/13.0.0/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></p></li>
<li id="fn2"><p>In theory, the Friedman test can also be used with two related conditions, but it then only uses the sign of the difference within each subject (it is equivalent to the sign test). In practice, the <a href="https://statsandr.com/blog/wilcoxon-test-in-r-how-to-compare-2-groups-under-the-non-normality-assumption/#paired-samples" rel="nofollow" target="_blank">Wilcoxon signed-rank test</a>, which also uses the size of the differences, is preferred for two related samples.<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#fnref2" class="footnote-back" rel="nofollow" target="_blank"><img src="https://s.w.org/images/core/emoji/13.0.0/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></p></li>
<li id="fn3"><p>The Holm method is less conservative than the Bonferroni correction, while controlling the same global error rate. See <code>?p.adjust</code> for the other available methods.<a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/#fnref3" class="footnote-back" rel="nofollow" target="_blank"><img src="https://s.w.org/images/core/emoji/13.0.0/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" />︎</a></p></li>
</ol>
</div>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://statsandr.com/blog/friedman-test-nonparametric-version-repeated-measures-anova/"> R on Stats and R</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/friedman-test-in-r-or-the-nonparametric-version-of-the-repeated-measures-anova/">Friedman test in R, or the nonparametric version of the repeated measures ANOVA</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403766</post-id>	</item>
		<item>
		<title>O*NET ratings, job descriptions and classification codes as measures of occupational similarity</title>
		<link>https://www.r-bloggers.com/2026/09/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/</link>
		
		<dc:creator><![CDATA[Giles Dickenson-Jones]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 10:30:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://www.gilesd-j.com/?p=4561</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> The last post identified potential transition pathways for truck drivers based on how closely other occupations matched them on skills, abilities and knowledge. This post follows the same basic approach across a wider range of occupations and asks whether two proxies could stand in for the ratings where no O*...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/">O*NET ratings, job descriptions and classification codes as measures of occupational similarity</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/"> Data Analytics and AI Archives - Giles</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>

<p class="wp-block-paragraph"><em><strong>TLDR:</strong></em> <em>The last post identified potential transition pathways for truck drivers based on how closely other occupations matched them on skills, abilities and knowledge. This post follows the same basic approach across a wider range of occupations and asks whether two proxies could stand in for the ratings where no O*NET exists: a small language model comparing the meaning of job descriptions and the distance between two jobs’ US Standard Occupational Classification (SOC)</em> <em>codes. Both agree with the skill-based scores on average and appear to complement each other. The SOC code helps separate jobs in different classes of work and the descriptions are better at ranking occupations within similar groups. Together they explain about a third of the variation in the skill-based score overall, and 10 to 15 percent within a major group, which is where a real job seeker is likely to look. Neither is a substitute for the O*NET’s worker characteristic data, but both provide a reasonable sense check</em> <em>of</em> <em>transition pathway estimates which can be useful when it’s not possible to use data from the O*NET.</em></p>



<h2 class="wp-block-heading">Background</h2>



<p class="wp-block-paragraph">In the <a href="https://www.gilesd-j.com/2026/08/28/identifying-skills-based-occupational-transition-pathways-with-the-onet/" rel="nofollow" target="_blank">last post</a> I reproduced the similarity scores estimated by <a href="https://doi.org/10.1016/j.tranpol.2025.103907" rel="nofollow" target="_blank">Bratanova et al. (2026)</a>. They used worker characteristic data from the O*NET to do this, with the basic idea being that two jobs that demand similar things of their workers should <em>on average</em> be easier to transition between<em>.</em> There’s of course more to their paper than this, but this is the crux of the approach.</p>



<p class="wp-block-paragraph">They’re certainly not the first to use the O*NET data in this way.<sup data-fn="c857b964-35bf-4c85-8b8f-31555c2650e1" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#c857b964-35bf-4c85-8b8f-31555c2650e1" id="c857b964-35bf-4c85-8b8f-31555c2650e1-link" rel="nofollow" target="_blank">1</a></sup> There are also some great interactive tools<sup data-fn="33ae5a5a-24da-480e-8785-5b3ac56a9c2f" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#33ae5a5a-24da-480e-8785-5b3ac56a9c2f" id="33ae5a5a-24da-480e-8785-5b3ac56a9c2f-link" rel="nofollow" target="_blank">2</a></sup> that intuitively illustrate how the approach works. Unfortunately, only found these after I’d pored over the literature as an adviser on a similar project in India. Because the SOC didn’t map cleanly to India’s labour market<sup data-fn="3799ac63-09ad-45f4-aa5d-a000497f84ab" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#3799ac63-09ad-45f4-aa5d-a000497f84ab" id="3799ac63-09ad-45f4-aa5d-a000497f84ab-link" rel="nofollow" target="_blank">3</a></sup>, we decided to dropped the O*NET in any case. But, the project left me with three questions that this post tries to answer:</p>



<ul class="wp-block-list">
<li class=""><em>How much value does detailed worker characteristics data add when compared to similarity measures built from descriptive text, such as job descriptions?</em></li>



<li class=""><em>If the two approaches differ, are the differences significant enough to matter in practice?</em></li>



<li class=""><em>Given occupational classification systems are usually designed to group comparable jobs together, can they serve as proxy for skill gaps?</em></li>
</ul>



<h3 class="wp-block-heading">Occupational Similarity: Three Ways</h3>



<p class="wp-block-paragraph">Had we used the O*NET, answering these questions would have been relatively simple: calculate similarity scores using occupation characteristic from the O*NET and compare these with scores calculated from text-based measures and a numeric indicator of classification similarity, which is more or less what I’m going to do here.</p>



<h3 class="wp-block-heading">A primer on skill similarity and the O*NET</h3>



<p class="wp-block-paragraph">If you’re unfamiliar with the territory I’d recommend taking a look at <a href="https://www.gilesd-j.com/2026/08/31/identifying-skills-based-occupational-transition-pathways-with-the-onet/" rel="nofollow" target="_blank">my previous post</a> where I explain why skill similarity measures are used for identifying transition pathways <em>and</em> the basic structure of the O*NET data that is drawn on. But, for those unwilling to leave this page:</p>



<ul class="wp-block-list">
<li class=""><strong>The O*NET:</strong> rates 800+ occupations across a standard set of characteristics that describe common abilities, skills and knowledge required to perform an occupation. Each characteristic is then given a rating based on the <em>level</em> required and its <em>importance.</em> Importance is meant to measure how critical a characteristic is to a job, whereas the level signifies the level of proficiency required / complexity of the task. For instance, the skill of <em>speaking</em> is important for both lawyers and paralegals, but lawyers are expected to have a higher <em>level</em> of speaking skills compared to paralegals (<a href="https://www.onetonline.org/help/online/scales" rel="nofollow" target="_blank">see here)</a>.</li>



<li class=""><strong>Skill similarity as a proxy for viable transition pathways:</strong> When two occupations demand a similar set of abilities, skills and knowledge it’s assumed a worker would find it easier to transition between the two, <em><a href="https://en.wikipedia.org/wiki/Ceteris_paribus" rel="nofollow" target="_blank">on average</a></em>. For instance, because the skills, abilities and knowledge required of truck drivers resembles bus drivers, it might be easier (and require less training) for them to transition between these jobs compared to becoming a data scientist. Both transitions are possible, but one involves overcoming more friction than the other, making the pathway less <em>viable</em>.<sup data-fn="1173c6b3-4b61-451e-bf9b-866d20c4ebce" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#1173c6b3-4b61-451e-bf9b-866d20c4ebce" id="1173c6b3-4b61-451e-bf9b-866d20c4ebce-link" rel="nofollow" target="_blank">4</a></sup> </li>



<li class=""><strong>The O*NET and US Standard Occupational Classification (SOC):</strong> The SOC uses a tiered classification system to group occupations, with jobs organized into 23 major groups, 98 minor groups and 459 broad occupations. Each SOC code specifies where a job has been assigned, for instance the SOC code for Chief Executives SOC is 11-1011.00, which places it in:</li>
</ul>



<h3 class="wp-block-heading">Three ways to test whether two jobs are alike</h3>



<p class="wp-block-paragraph">Because I’m interested in <em>testing</em> the ideas rather than describing them, I’m not going to cover each approach in too much detail, but the TLDR summary of each approach is below:</p>



<ul class="wp-block-list">
<li class=""><strong>Ratings similarity:</strong> occupations are compared based on the <em>level</em> and <em>importance of</em> skills, abilities and knowledge required to complete the job.</li>



<li class=""><strong>Semantic similarity:</strong> Occupations are compared to one-another based on the similarity of job descriptions. Measurement is based on <em>semantic</em> similarity of job descriptions, which attempts to judge text based on its meaning rather than contents.</li>



<li class=""><strong>SOC code distance:</strong> Differences in SOC code assignment is used to provide a <em>naive</em> measure of how similar two occupations are likely to be. Because occupations in the SOC are classified based on work performed and, in some cases, on the skills, education and/or training needed to perform the work<sup data-fn="74b5b1c4-493f-46fb-ba4d-a2a6ae82cfb0" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#74b5b1c4-493f-46fb-ba4d-a2a6ae82cfb0" id="74b5b1c4-493f-46fb-ba4d-a2a6ae82cfb0-link" rel="nofollow" target="_blank">5</a></sup>, it’s expected that similarly classified jobs will <em>on average</em> share similar characteristics (although this won’t necessarily be linear).</li>
</ul>



<h2 class="wp-block-heading">Project set-up</h2>



<h3 class="wp-block-heading">Data:</h3>



<p class="wp-block-paragraph">Like <a href="https://www.gilesd-j.com/2026/08/31/identifying-skills-based-occupational-transition-pathways-with-the-onet/" rel="nofollow" target="_blank">the previous O*NET post</a>, analysis uses version 31.0 of <a href="https://www.onetcenter.org/database.html" rel="nofollow" target="_blank">the O*NET database</a>. You can download the data used in this post <a href="http://gilesd-j.com/shared_resources/blogs/260921_onet_proxies/260921_onet_proxy_data.zip" rel="nofollow" target="_blank">here</a>.</p>



<p class="wp-block-paragraph">Everything a reader might want to change sits in this one chunk: packages, paths, the rating scale bounds, the embedding model, the labels and the chart theme.</p>



<pre>library(tidyverse)
library(janitor)
library(readxl)
library(text)
library(scales)

#where the data lives and where the figures get written
ref_dir_data   &lt;- file.path(&quot;.&quot;, &quot;Data&quot;)
ref_dir_images &lt;- file.path(&quot;.&quot;, &quot;images&quot;)
dir.create(ref_dir_images, showWarnings = FALSE)

#the O*NET rating files. The NAME of each entry becomes the element set label
#in the combined table, which is how essential and transferable skills stay apart.
ref_files_onet &lt;- c(
  abilities           = &quot;Abilities.xlsx&quot;,
  skills_essential    = &quot;Essential Skills.xlsx&quot;,
  skills_transferable = &quot;Transferable Skills.xlsx&quot;,
  knowledge           = &quot;Knowledge.xlsx&quot;
)

#the O*NET file holding one plain-English description per occupation
ref_file_desc &lt;- &quot;Occupation_Data.xlsx&quot;

#O*NET rating scale bounds, used to put both ratings on a 0-100 scale.
#Importance is collected on a 1-5 scale, level on a 0-7 scale.
ref_scale_level_max &lt;- 7
ref_scale_imp_min   &lt;- 1
ref_scale_imp_max   &lt;- 5

#For semantic similarity, all-MiniLM-L6-v2 has been used. 
ref_model_embed &lt;- &quot;sentence-transformers/all-MiniLM-L6-v2&quot;

#readable labels for the five ratings we're comparing 
ref_labels_set &lt;- c(
  abilities           = &quot;Abilities&quot;,
  skills_essential    = &quot;Essential skills&quot;,
  skills_transferable = &quot;Transferable skills&quot;,
  knowledge           = &quot;Knowledge&quot;,
  all                 = &quot;Average Across Worker Characteristics&quot;
)

#for chart labels
ref_labels_measure &lt;- c(ref_labels_set, semantic = &quot;Semantic (descriptions)&quot;)


#import SOC classification structure
ref_path_soc &lt;- file.path(&quot;Data&quot;, &quot;soc_structure_2018.xlsx&quot;)

# Import data
dta_soc_raw &lt;- read_excel(
  ref_path_soc,
  sheet = &quot;2018 Structure&quot;,
  skip = 8,
  col_names = c(&quot;major&quot;, &quot;minor&quot;, &quot;broad&quot;, &quot;detailed&quot;, &quot;soc_title&quot;),
  col_types = &quot;text&quot;
)

# Data wrangling
# One row per code, with its level
dta_soc_long &lt;- dta_soc_raw |&gt;
  pivot_longer(major:detailed, names_to = &quot;soc_level&quot;, values_to = &quot;soc_code&quot;, values_drop_na = TRUE)

# One row per detailed occupation with parent codes and titles
lkp_soc_title &lt;- setNames(dta_soc_long$soc_title, dta_soc_long$soc_code)

dta_soc_hier &lt;- dta_soc_raw |&gt;
  fill(major, minor, broad) |&gt;
  filter(!is.na(detailed)) |&gt;
  mutate(
    major_title = lkp_soc_title[major],
    minor_title = lkp_soc_title[minor],
    broad_title = lkp_soc_title[broad]
  ) |&gt;
  select(major, major_title, minor, minor_title, broad, broad_title, detailed, detailed_title = soc_title)


#convert the SOC classification to format for plots
lkp_soc_major &lt;- dta_soc_long |&gt;
  filter(soc_level == &quot;major&quot;) |&gt;
  transmute(soc_major       = str_sub(soc_code, 1, 2),
            soc_major_label = str_c(soc_major, &quot; &quot;, str_remove(soc_title, &quot; Occupations$&quot;)))

#colours for figures.
ref_col_desc        &lt;- &quot;#D4A017&quot;
ref_col_soc_rank    &lt;- &quot;#1F3A5F&quot;
ref_col_soc_raw     &lt;- &quot;#7A94B8&quot;
ref_col_combo       &lt;- &quot;#2F7E6E&quot;
ref_col_combo_light &lt;- &quot;#86B8AB&quot;
ref_col_primary &lt;- &quot;#1E298D&quot;  # dark blue
ref_col_accent  &lt;- &quot;#26CDDA&quot;  # cyan
ref_col_grid    &lt;- &quot;#E5E7EB&quot;
ref_col_text    &lt;- &quot;#121212&quot;

#a shared theme so every chart in the post looks the same
ref_theme_post &lt;- theme_minimal(base_size = 11) +
  theme(
    panel.grid.minor    = element_blank(),
    panel.grid.major    = element_line(colour = ref_col_grid),
    axis.text           = element_text(colour = ref_col_text),
    axis.title          = element_text(colour = ref_col_primary, face = &quot;bold&quot;),
    strip.text          = element_text(colour = ref_col_text, face = &quot;bold&quot;, hjust = 0),
    legend.position     = &quot;bottom&quot;,
    plot.title.position = &quot;plot&quot;
  )

#the extra theme settings the three heatmaps share: no gridlines, small
#rotated axis text and a wide legend bar
ref_theme_heatmap &lt;- ref_theme_post +
  theme(
    panel.grid       = element_blank(),
    axis.text.x      = element_text(angle = 90, vjust = 0.5, hjust = 1, size = 7),
    axis.text.y      = element_text(size = 7),
    legend.key.width = unit(1.5, &quot;cm&quot;)
  )</pre>



<h3 class="wp-block-heading">Import the data</h3>



<p class="wp-block-paragraph">The code below reads O*NET data adding source labels.</p>



<pre>#read each O*NET rating file, tag it with its element set, and stack them
dta_onet_raw &lt;- imap(ref_files_onet, \(ref_file, ref_set) {
  read_excel(file.path(ref_dir_data, ref_file)) |&gt;
    clean_names() |&gt;
    mutate(element_set = ref_set)
}) |&gt;
  list_rbind() |&gt;
  rename(soc_code  = o_net_soc_code,
         soc_label = title)

#one row per occupation: code, title and description
dta_desc_raw &lt;- read_excel(file.path(ref_dir_data, ref_file_desc)) |&gt;
  clean_names() |&gt;
  rename(soc_code  = o_net_soc_code,
         soc_label = title)

#a lookup of code to title, taken from the ratings files
lkp_soc_labels &lt;- dta_onet_raw |&gt;
  distinct(soc_code, soc_label)</pre>



<h2 class="wp-block-heading">Exploratory analysis</h2>



<h3 class="wp-block-heading">Checking coverage</h3>



<p class="wp-block-paragraph">The code below does a simple sense check of the imported data by presenting the number of occupations covered and the number of characteristics measured across O*NET abilities, knowledge and skills. Each characteristic (or element set) should have the same number of values for importance and level.</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th>Element set</th><th>Importance</th><th>Level</th></tr></thead><tbody><tr><td>Abilities</td><td>52</td><td>52</td></tr><tr><td>Knowledge</td><td>33</td><td>33</td></tr><tr><td>Skills (essential)</td><td>10</td><td>10</td></tr><tr><td>Skills (transferable)</td><td>25</td><td>25</td></tr></tbody><tfoot><tr><td>Total</td><td>120</td><td>120</td></tr></tfoot></table></figure>



<pre>#characteristic and description counts
sum_onet_coverage &lt;- tibble(
  occupations_rated      = n_distinct(dta_onet_raw$soc_code),
  occupations_described  = n_distinct(dta_desc_raw$soc_code),
  rated_with_description = n_distinct(intersect(dta_onet_raw$soc_code,
                                                dta_desc_raw$soc_code))
)
#
sum_onet_coverage


#characteristic counts 
sum_onet_element_counts &lt;- dta_onet_raw |&gt;
  distinct(element_set, element_name, scale_name) |&gt;
  count(element_set, scale_name, name = &quot;nmb_elements&quot;) |&gt;
  pivot_wider(names_from = scale_name, values_from = nmb_elements)

sum_onet_element_counts</pre>



<h2 class="wp-block-heading">Calculating skill-based similarity</h2>



<p class="wp-block-paragraph">We’ll follow a similar approach to <a href="https://www.gilesd-j.com/2026/08/31/identifying-skills-based-occupational-transition-pathways-with-the-onet/" rel="nofollow" target="_blank">the previous post</a> when calculating skill-based similarity scores, with two main differences:</p>



<ol class="wp-block-list">
<li class="">We use <em>Manhattan distance</em> to calculate skill gaps via R’s <code>dist()</code> function. This is essentially the same calculation, but more memory efficient.</li>



<li class="">Average skill gaps are calculated for each worker characteristic and <em>then</em> averaged, so abilities, knowledge, essential and transferable skills are weighted equally.</li>
</ol>



<h3 class="wp-block-heading">Put level and importance on a common scale</h3>



<p class="wp-block-paragraph">The O*NET rates level from 0 to 7 and importance from 1 to 5, so the code below scales both ratings so they are on a common scale:</p>



<pre>dta_onet_scaled &lt;- dta_onet_raw |&gt;
  mutate(scale_name = str_to_lower(scale_name)) |&gt;
  mutate(value_0_100 = case_when(
    scale_name == &quot;level&quot;      ~ 100 * data_value / ref_scale_level_max,
    scale_name == &quot;importance&quot; ~ 100 * (data_value - ref_scale_imp_min) /
                                       (ref_scale_imp_max - ref_scale_imp_min)
  )) |&gt;
  select(soc_code, soc_label, element_set, element_name, scale_name, value_0_100)</pre>



<h3 class="wp-block-heading">Calculating skill gaps</h3>



<p class="wp-block-paragraph">In the last post, each origin occupation was paired with each potential destination job using a join before skill gaps were calculated. This worked, but required creating a large table with a lot of redundant information before making the calculation. The code below uses <code>dist()</code>, which is a more memory efficient way to make the same calculation by:</p>



<ul class="wp-block-list">
<li class="">Laying out each characteristic and scale across columns, with one job per row.</li>



<li class="">Calculating the total skill gaps across all worker characteristics for each job pair.</li>



<li class="">Dividing the total skill gap for each job pair by the number of columns to get the average.</li>



<li class="">Averaging skill gaps across worker characteristic categories.</li>
</ul>



<pre>#split the rescaled ratings into the four sets, then add a fifth entry holding
#all of them together
dta_onet_by_set &lt;- split(dta_onet_scaled, dta_onet_scaled$element_set)


# Stage 1: widen a set of ratings to one row per occupation, one column per
# element and scale (with columns names based on the element set, name and scale ).
fnc_widen_ratings &lt;- function(dta_ratings) {
  dta_ratings |&gt;
    select(soc_code, element_set, element_name, scale_name, value_0_100) |&gt;
    pivot_wider(names_from  = c(element_set, element_name, scale_name),
                values_from = value_0_100)
}

#create wide dataframe with one row per occupation 
dta_onet_wide &lt;- dta_onet_by_set |&gt;
  map(fnc_widen_ratings)

#to see an example of the input data:
# tmp_dta_onet_wide_example&lt;-dta_onet_wide$skills_transferable |&gt; head()
# rm(tmp_dta_onet_wide_example)

#check that each worker requirement has the expected number of columns and rows
# (894 occupations and 1 column for each element_name and level e.g 52 abilities and two ratings for each (level and importance)  (52*2=104)) 
chk_wide_dimensions &lt;- dta_onet_wide |&gt;
  map(\(dta_wide) tibble(nmb_occupations = nrow(dta_wide),
                         nmb_columns     = ncol(dta_wide) - 1)) |&gt;
  list_rbind(names_to = &quot;element_set&quot;)

chk_wide_dimensions

# Stage 2: Calculate the Manhattan distance between every pair of rows, divided by the
# number of columns to make it an average gap. 
fnc_pair_gaps &lt;- function(dta_wide) {
  
  tmp_matrix &lt;- dta_wide |&gt;
    #move the occupation code from a column to the row labels
    column_to_rownames(&quot;soc_code&quot;) |&gt;
    #dist() needs a matrix of numbers, not a data frame
    as.matrix()
  
  #dist() calculates the total absolute skill gap between occupations across all characteristics and then takes the average.  
  #as.matrix() lays the input data out as a square with one row and one column
  #per occupation, presenting all job pair transiton paths 
  tmp_gaps &lt;- as.matrix(dist(tmp_matrix, method = &quot;manhattan&quot;)) / ncol(tmp_matrix)
  
  tmp_gaps |&gt;
    #back to a data frame, with the row labels as a column of origin codes
    as_tibble(rownames = &quot;soc_code_a&quot;) |&gt;
    #stack the destination columns into rows: one row per pair
    pivot_longer(-soc_code_a, names_to = &quot;soc_code_b&quot;, values_to = &quot;gap_mean&quot;)
}

#average gap within each of the four sets
dta_gaps_by_set &lt;- dta_onet_wide |&gt;
  map(fnc_pair_gaps) |&gt;
  list_rbind(names_to = &quot;element_set&quot;)

#the combined score: average the four set-level gaps, so each set counts
#equally whether it has 33 characteristics or 52
dta_gaps_all &lt;- dta_gaps_by_set |&gt;
  summarise(gap_mean = mean(gap_mean), .by = c(soc_code_a, soc_code_b)) |&gt;
  mutate(element_set = &quot;all&quot;)

dta_gaps &lt;- bind_rows(dta_gaps_by_set, dta_gaps_all)

#every set should produce the same number of pairs: occupations squared
chk_pairs_per_set &lt;- dta_gaps |&gt;
  count(element_set, name = &quot;nmb_pairs&quot;)

chk_pairs_per_set</pre>



<h3 class="wp-block-heading">Turn gaps into scores</h3>



<p class="wp-block-paragraph">The code below rescales skill gaps so each job pair gets a similarity score between 0 and 100. Job pairs that require similar worker characteristics (and have smaller skill gaps) receive scores closer to 100, while job pairs with larger skill gaps receive similarity scores closer to zero. Because job pairs include moving to the same occupation, the minimum skill gap is zero.</p>



<pre>#| label: similarity-scores

rlt_sim_ratings &lt;- dta_gaps |&gt;
  mutate(
    sim_score = 100 * (1 - (gap_mean - min(gap_mean)) / (max(gap_mean) - min(gap_mean))),
    .by = element_set
  )

#every occupation against itself, which should be 100 in every set
chk_self_pairs &lt;- rlt_sim_ratings |&gt;
  filter(soc_code_a == soc_code_b) |&gt;
  summarise(sim_min = min(sim_score), sim_max = max(sim_score), .by = element_set)

chk_self_pairs</pre>



<h3 class="wp-block-heading">Scores seem to cluster by broad work classes</h3>



<p class="wp-block-paragraph">The heatmap above presents the average similarity score by each SOC major group excluding identical job pairs.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i1.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/skill_heatmap.png?w=450&#038;ssl=1" alt="" class="wp-image-4618" srcset_temp="https://i1.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/skill_heatmap.png?w=450&#038;ssl=1 682w, https://www.gilesd-j.com/wp-content/uploads/2026/09/skill_heatmap-300x263.png 300w" sizes="auto, (max-width: 682px) 100vw, 682px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">Because self-comparisons have been removed, the downward sloping diagonal line shows the average similarity score for jobs pairs in the same major group. As one might expect, similarity scores tend to be higher for occupation pairs that are both in the same major group. There is also <em>some</em> visual support for similarity scores being lower for larger differences in SOC codes, with lower average similarity scores in the bottom left and upper right third of the heatmap (where the SOC classification differences are highest).</p>



<p class="wp-block-paragraph">However, neither pattern is unambiguously dominant, with a number of clusters being apparent in the heatmap. For instance, major groups 41 and 42 share unusually high average similarity scores across most other major groups up to 45 to 53. Correspondingly, the average similarity scores appear to be higher for occupation pairs within groups 45 to 53 than outside them. Major groups 11 to 19 exhibit a similar pattern.</p>



<p class="wp-block-paragraph">Although exploring the source of this is worthy of another post, one explanation for the observed clustering might be that some worker requirements aren’t as widely shared as others, which might point to some inter-group transitions being more difficult than others. This clustering also isn’t all that surprising given the SOC has been designed to group similar worker groups together.</p>



<pre>#average score between every pair of major groups, self-pairs excluded
sum_sim_ratings_major &lt;- rlt_sim_ratings |&gt;
  filter(soc_code_a != soc_code_b,
         element_set==&quot;all&quot;) |&gt;
  mutate(major_a = str_sub(soc_code_a, 1, 2),
         major_b = str_sub(soc_code_b, 1, 2)) |&gt;
  summarise(value = mean(sim_score), .by = c(element_set, major_a, major_b)) |&gt;
  mutate(panel = factor(ref_labels_set[element_set], levels = ref_labels_set))

plt_sim_ratings_major &lt;- sum_sim_ratings_major |&gt;
  #swap the two-digit codes for the readable major group names
  mutate(major_a = factor(major_a, levels = lkp_soc_major$soc_major,
                          labels = lkp_soc_major$soc_major_label),
         major_b = factor(major_b, levels = lkp_soc_major$soc_major)) |&gt;
  ggplot(aes(x = major_b, y = major_a, fill = value)) +
  geom_tile(colour = &quot;white&quot;, linewidth = 0.3) +
  #darker always means more alike
  scale_fill_gradient(low = ref_col_accent, high = ref_col_primary,
                      name = &quot;Average similarity score (0-100)&quot;) +
  scale_y_discrete(limits = rev) +
  coord_equal() +
  facet_wrap(vars(panel)) +
  ref_theme_heatmap +
  labs(title =&quot;Average similarity score by SOC major group&quot;, 
       x = &quot;SOC major group&quot;, y = &quot;SOC Major Group&quot;)

plt_sim_ratings_major</pre>



<h2 class="wp-block-heading">Estimate similarity from job descriptions</h2>



<p class="wp-block-paragraph">To estimate similarity scores from job descriptions, the O*NET’s occupation description data is used. For instance, the description for Emergency Management Directors (11-9161.00) is:<br><em>Plan and direct disaster response or crisis management activities, provide disaster preparedness training, and prepare emergency plans and procedures for natural (e.g., hurricanes, floods, earthquakes), wartime, or technological (e.g., nuclear power plant emergencies or hazardous materials spills) disasters or hostage situations.</em></p>



<p class="wp-block-paragraph">These descriptions aren’t really designed to describe what a job requires in detail and is better thought of as narrative text meant to support the SOC taxonomy. Although this makes them a poor substitute for real job posting data, they <em>do</em> resemble the narrative descriptions used by other occupational classification systems, which is a good <em>worst case</em> scenario to test.</p>



<p class="wp-block-paragraph">A sentence-embedding model is used to compare how similarly the meaning of two job descriptions are. These are neural network models trained on a large amount of text with the aim of turning a paragraph into a list of numbers designed to encode its underlying meaning. Because it isn’t simply comparing whether each paragraph share similar words, phrases like “drive a truck” and “operate a heavy vehicle” will receive similar scores that indicate they mean something similar. We’ll use cosine similarity, which gives a score of 1 when the meaning of seem to be pointing the same direction, 0 when they’re unrelated and <0 when they have opposite meanings.</p>



<p class="wp-block-paragraph">I’ve arbitrarily chosen the <a href="https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2" rel="nofollow" target="_blank">all-MiniLM-L6-v2 transformer model</a> to do compare job descriptions via the text package. I’m not going to claim a particularly rigorous process was used when selecting the model<sup data-fn="e1bb8487-9a3f-4168-8b6d-5fc0c5068ee5" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#e1bb8487-9a3f-4168-8b6d-5fc0c5068ee5" id="e1bb8487-9a3f-4168-8b6d-5fc0c5068ee5-link" rel="nofollow" target="_blank">6</a></sup>, but it holds two advantages for this post:</p>



<ol class="wp-block-list">
<li class="">It’s small enough to be run locally on a laptop; and</li>



<li class="">The model has been successfully used enough by researchers in adjacent fields<sup data-fn="4797dff9-9b8f-4fe4-83a0-8850cea9b3ee" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#4797dff9-9b8f-4fe4-83a0-8850cea9b3ee" id="4797dff9-9b8f-4fe4-83a0-8850cea9b3ee-link" rel="nofollow" target="_blank">7</a></sup>.</li>
</ol>



<p class="wp-block-paragraph">The code below has the model process the text of each O*NET job description that has a similarity score. Be warned, this can take time to run as it requires having the model process each of the 894 occupations</p>



<p class="wp-block-paragraph"><strong>Note:</strong> The first run also downloads the model, which is about 100MB (excluding additional library and Python-related requirements).</p>



<pre>textrpp_initialize()

#keep the descriptions of rated occupations only, in code order
dta_desc &lt;- dta_desc_raw |&gt;
  semi_join(lkp_soc_labels, by = join_by(soc_code)) |&gt;
  arrange(soc_code)

#turn each description into a 384-number embedding. The last layer, averaged
#over tokens, is how this model's own authors produce sentence embeddings.
dta_desc_embed &lt;- textEmbed(
  texts  = dta_desc |&gt; select(description),
  model  = ref_model_embed,
  layers = -1,
  aggregation_from_layers_to_tokens = &quot;concatenate&quot;,
  aggregation_from_tokens_to_texts  = &quot;mean&quot;,
  keep_token_embeddings = FALSE)

#one row per occupation, one column per embedding dimension
dta_desc_vectors &lt;- dta_desc_embed$texts$description</pre>



<h3 class="wp-block-heading">Compare every description with every other</h3>



<p class="wp-block-paragraph"><code>textSimilarityMatrix()</code> computes the cosine similarity between every job description pair to provide a standardized similarity measure across job description pairs. The argument <code>center=TRUE</code> tells the function to subtract values by their corresponding column means. This is the function’s default behaviour and is meant to reduce the chance of scores reflecting paragraphs <em>looking</em> similar as a result of sharing a common format and/or style e.g. all being job descriptions with a comparable style and format.</p>



<pre>#cosine similarity between every pair of descriptions, after centring
mat_desc_sim &lt;- as.matrix(
  textSimilarityMatrix(dta_desc_vectors, method = &quot;cosine&quot;, center = TRUE)
)
dimnames(mat_desc_sim) &lt;- list(dta_desc$soc_code, dta_desc$soc_code)

#from a matrix to one row per pair, same shape as the ratings table
rlt_sim_semantic &lt;- mat_desc_sim |&gt;
  as_tibble(rownames = &quot;soc_code_a&quot;) |&gt;
  pivot_longer(-soc_code_a, names_to = &quot;soc_code_b&quot;, values_to = &quot;sim_semantic&quot;)

#a description is identical to itself, so the diagonal should be 1
chk_semantic_self &lt;- rlt_sim_semantic |&gt;
  filter(soc_code_a == soc_code_b)</pre>



<h3 class="wp-block-heading">Descriptions show higher average similarity within the same group</h3>



<p class="wp-block-paragraph">The code below produces a heatmap of average semantic similarity scores for each job description pair once identical job pairs are dropped.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i0.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/description_heatmap.png?w=450&#038;ssl=1" alt="" class="wp-image-4615" srcset_temp="https://i0.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/description_heatmap.png?w=450&#038;ssl=1 682w, https://www.gilesd-j.com/wp-content/uploads/2026/09/description_heatmap-300x263.png 300w" sizes="auto, (max-width: 682px) 100vw, 682px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">Once again, the downward sloping diagonal follows a similar pattern to the skill-based similarity scores: average similarity scores are higher for job pairs within the same SOC major group. The skill-based clustering of groups into wider families of work like the skill-based heatmap isn’t as apparent. Exploring why would make a good excuse for another post, but if I had to guess it’s probably a result of some combination of the following:</p>



<ul class="wp-block-list">
<li class=""><strong>Occupation descriptions from the same group use common terms and formats:</strong> Marketing Managers, Sales Managers and Public Relations Managers have descriptions that start with <em>“Plan, direct, or coordinate…”.</em></li>



<li class=""><strong>Terms and format don’t appear to carry across groups:</strong> a cursory glance at the descriptions used across group clusters in the skill-based score heatmap suggests the format and terms used differ a lot between groups. For instance, occupations in the observed cluster between groups 45 to 53 tend to be described by tools, equipment and outputs that are specific to their work (repairing wind turbines, installing roof support bolts, harvesting vegetables etc), with the result being that similar will often look very different to the model.</li>



<li class=""><strong>Descriptions vary in length and specificity:</strong> the length of descriptions vary from single short sentences to lengthy paragraphs. Unlike the O*NET’s worker characteristic data they also haven’t been standardized to allow comparison across all occupations in the SOC.</li>
</ul>



<p class="wp-block-paragraph"><strong>Note:</strong> It’s also likely that the descriptions of jobs in the same group leaned more on one another than jobs from other groups when they were drafted.</p>



<pre>sum_sim_semantic_major &lt;- rlt_sim_semantic |&gt;
  filter(soc_code_a != soc_code_b) |&gt;
  mutate(major_a = str_sub(soc_code_a, 1, 2),
         major_b = str_sub(soc_code_b, 1, 2)) |&gt;
  summarise(value = mean(sim_semantic), .by = c(major_a, major_b))

plt_sim_semantic_major &lt;- sum_sim_semantic_major |&gt;
  mutate(major_a = factor(major_a, levels = lkp_soc_major$soc_major,
                          labels = lkp_soc_major$soc_major_label),
         major_b = factor(major_b, levels = lkp_soc_major$soc_major)) |&gt;
  ggplot(aes(x = major_b, y = major_a, fill = value)) +
  geom_tile(colour = &quot;white&quot;, linewidth = 0.3) +
  scale_fill_gradient(low = ref_col_accent, high = ref_col_primary,
                      name = &quot;Average cosine similarity&quot;) +
  scale_y_discrete(limits = rev) +
  coord_equal() +
  ref_theme_heatmap +
  labs(x = &quot;SOC major group&quot;, y = NULL)


plt_sim_semantic_major</pre>



<h2 class="wp-block-heading">Use the SOC code distance as a “measure” of similarity</h2>



<p class="wp-block-paragraph">I’ve used double quotes around <em>measure</em> to signify I’m making <em>air quotes</em> here, but to be clear: by <em>measure</em> I mean <em>proxy.</em> As <em>even</em> if similar jobs have been placed in the same group <em>and</em> groups are designed to correspond with wider groupings, such as white-collar, service, blue-collar and members of the military, code differences are not designed to precisely quantify where a job sits on some comparable continuum.</p>



<p class="wp-block-paragraph">To provide an example of what I mean by this: Surgical Assistants (29-9093) and Healthcare Practitioners and Technical Workers (29-9099) are next to each other on the SOC, with codes that are six apart from one another. Yet, Home Health Aides (31-1121) has a code that’s a little over 12 thousand apart from both jobs, despite the three jobs being broadly comparable.</p>



<p class="wp-block-paragraph">Although this is an extreme example, it does illustrate one of the core problems with relying on raw SOC codes as a proxy: small differences in assignments can result in large code differences that don’t correspond to how similar (or different) two jobs are. Added to this, even when these differences do provide a reasonable proxy for how similar (or different) two jobs are, this is unlikely to be constant across job pairs.</p>



<p class="wp-block-paragraph">Knowing all this, a reasonable person might ask <em>why look at SOC codes at all</em>. Four reasons:</p>



<ol class="wp-block-list">
<li class=""><strong>Data availability:</strong> Data as rich as the O*NET sometimes isn’t available, can’t be drawn on, or isn’t appropriate to use for the labour market being examined.</li>



<li class=""><strong>A bad proxy can still have useful information</strong>: Occupation codes might get the ordering of similar jobs right, even when it gets the size of differences wrong.</li>



<li class=""><strong>To complement other similarity measures:</strong> When data is scarce, classification code differences might provide useful information that other data sources lack, making it useful as part of a wider portfolio of measures, such as part of a composite index.</li>



<li class=""><strong>As an intuitive sense-check:</strong> When adjacent codes tend to be used for similar jobs, score differences can provide a simple sanity check of other similarity scores.</li>
</ol>



<p class="wp-block-paragraph">To limit the effect of adjacent occupations being assigned to different groups, the code below creates a difference measure based on an occupation’s ranking in the SOC, rather than raw code differences. This assumes that adjacent occupations are equally alike across the SOC, regardless of which groups they fall in. This approach was chosen as it feels easier to justify than relying on the raw codes, which sometimes exhibit large differences across classifications that are unlikely to correspond to how similar two jobs are.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/soc_heatmap.png?w=450&#038;ssl=1" alt="" class="wp-image-4621" srcset_temp="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/soc_heatmap.png?w=450&#038;ssl=1 682w, https://www.gilesd-j.com/wp-content/uploads/2026/09/soc_heatmap-300x263.png 300w" sizes="auto, (max-width: 682px) 100vw, 682px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">Given differences in SOC code rankings are linear, it’s not surprising that the heatmap presents limited variation. However, it <em>does</em> present two patterns that you might expect real similarity scores <em>should</em> follow* when mapped on a classification system like the SOC:</p>



<ol class="wp-block-list">
<li class="">Occupations pairs from the same group <em>should</em> exhibit higher similarity with one another; and</li>



<li class="">Similarity scores should decrease as the distance between groups increase.</li>
</ol>



<p class="wp-block-paragraph">*(Provided the classification system groups jobs in a way that aligns with the nature of the work.)</p>



<p class="wp-block-paragraph"><strong>Note:</strong> The code also estimates raw SOC distance to allow testing whether within-group code differences have predictive power within the same minor group.</p>



<pre>#position of each occupation in the SOC list: 1 = first code, 894 = last
lkp_soc_rank &lt;- lkp_soc_labels |&gt;
  arrange(soc_code) |&gt;
  mutate(soc_rank = row_number()) |&gt;
  select(soc_code, soc_rank)

#two readings of the SOC for every pair:
#1.) differences in code rankings
#2.) differences in raw codes
sum_soc_distance &lt;- rlt_sim_ratings |&gt;
  distinct(soc_code_a, soc_code_b) |&gt;
  left_join(lkp_soc_rank, by = join_by(soc_code_a == soc_code)) |&gt;
  left_join(lkp_soc_rank, by = join_by(soc_code_b == soc_code),
            suffix = c(&quot;_a&quot;, &quot;_b&quot;)) |&gt;
  mutate(
    soc_rank_distance = abs(soc_rank_a - soc_rank_b),
    #the full code as one integer, suffix included: &quot;11-9199.11&quot; -&gt; 11919911.
    #Removing the point as well as the hyphen keeps the suffix as whole digits,
    #so two specialisations of one code are 1 apart rather than 0.01, which
    #rounding and floating point would otherwise erase.
    soc_numeric_a     = as.numeric(str_remove_all(soc_code_a, &quot;[-.]&quot;)),
    soc_numeric_b     = as.numeric(str_remove_all(soc_code_b, &quot;[-.]&quot;)),
    soc_code_distance = abs(soc_numeric_a - soc_numeric_b),
    major_a           = str_sub(soc_code_a, 1, 2),
    major_b           = str_sub(soc_code_b, 1, 2)
  )
#average list-position distance between every pair of major groups, self-pairs
#excluded. The raw code is kept for later; drawn up here it would show the same
#bands with different numbers.
sum_soc_distance_major &lt;- sum_soc_distance |&gt;
  filter(soc_code_a != soc_code_b) |&gt;
  summarise(value = mean(soc_rank_distance), .by = c(major_a, major_b))

plt_soc_distance_major &lt;- sum_soc_distance_major |&gt;
  mutate(major_a = factor(major_a, levels = lkp_soc_major$soc_major,
                          labels = lkp_soc_major$soc_major_label),
         major_b = factor(major_b, levels = lkp_soc_major$soc_major)) |&gt;
  ggplot(aes(x = major_b, y = major_a, fill = value)) +
  geom_tile(colour = &quot;white&quot;, linewidth = 0.3) +
  #the ramp is flipped here, so darker still means closer
  scale_fill_gradient(low = ref_col_primary, high = ref_col_accent,
                      name = &quot;Average difference in list position (darker = closer)&quot;) +
  scale_y_discrete(limits = rev) +
  coord_equal() +
  ref_theme_heatmap +
  labs(x = &quot;SOC major group&quot;, y = NULL,
       title = &quot;Average difference in SOC code ranking by major group&quot;)

plt_soc_distance_major</pre>



<h2 class="wp-block-heading">Do the measures agree?</h2>



<p class="wp-block-paragraph">But, do the three measures <del>blend</del> agree with one another?</p>



<p class="wp-block-paragraph">The code below takes all three measures and combines them into a single dataframe so individual similarity measures can be compared against one another. Groupings are specified to help communicate insights from exploratory analysis (and arguments with Claude) that aren’t shown here for the sake of brevity. In short, the pairs cover:</p>



<ul class="wp-block-list">
<li class=""><strong>Groupings based on where occupation codes differ:</strong> Such as job pairs that differ by major, minor, or detailed SOC groupings. This allows us to ask whether the explanatory power of compared measures change when job pairs are drawn from different classifications.</li>



<li class=""><strong>Groupings where jobs are in the same SOC groups:</strong> Such as where compared jobs are in the same major, minor or SOC group. This is to see if the predictive power of measures change when job pairs are drawn from closer occupational groups.</li>
</ul>



<p class="wp-block-paragraph">The second group comes from the exploratory analysis left out of the post, which pointed to the SOC distance measure’s explanatory power varying based on how differently two jobs were classified. Because the ranking ignores the size of classification differences <em>within</em> an occupational group, raw SOC code differences might be better at separating jobs in the same group than the ranking, which I wanted to test.</p>



<pre>#each pair is labelled by the first level at which its two codes differ. A pair
#that differs at the minor group shares the major group above it; a pair that
#differs only in the suffix is two specialisations of one detailed occupation.
ref_labels_differ &lt;- c(
  &quot;0&quot; = &quot;Differ at major group&quot;,
  &quot;1&quot; = &quot;Differ at minor group&quot;,
  &quot;2&quot; = &quot;Differ at broad occupation&quot;,
  &quot;3&quot; = &quot;Differ at detailed occupation&quot;,
  &quot;4&quot; = &quot;Differ in suffix only&quot;
)

#labels for the second cut: the group both codes share
ref_labels_share &lt;- c(
  &quot;All pairs&quot;,
  &quot;Same major group&quot;,
  &quot;Same minor group&quot;,
  &quot;Same broad occupation&quot;,
  &quot;Same detailed occupation&quot;
)

#combine ratings in a single dataframe
rlt_pairs &lt;- rlt_sim_ratings |&gt;
  #the combined ratings score only; the four sets had their turn above
  filter(element_set == &quot;all&quot;) |&gt;
  select(soc_code_a, soc_code_b, sim_ratings = sim_score) |&gt;
  inner_join(rlt_sim_semantic, by = join_by(soc_code_a, soc_code_b)) |&gt;
  inner_join(sum_soc_distance |&gt;
               select(soc_code_a, soc_code_b, soc_rank_distance, soc_code_distance),
             by = join_by(soc_code_a, soc_code_b)) |&gt;
  #each pair once, lower code first, self-pairs dropped
  filter(soc_code_a &lt; soc_code_b) |&gt;
  mutate(
    #distances flipped to closeness, so every measure reads higher = more alike
    soc_closeness      = -soc_rank_distance,
    soc_code_closeness = -soc_code_distance,
    #&quot;53-3032.00&quot;: major &quot;53&quot;, minor &quot;53-3&quot;, broad &quot;53-303&quot;, detailed &quot;53-3032&quot;
    shared_depth = case_when(
      str_sub(soc_code_a, 1, 7) == str_sub(soc_code_b, 1, 7) ~ 4,
      str_sub(soc_code_a, 1, 6) == str_sub(soc_code_b, 1, 6) ~ 3,
      str_sub(soc_code_a, 1, 4) == str_sub(soc_code_b, 1, 4) ~ 2,
      str_sub(soc_code_a, 1, 2) == str_sub(soc_code_b, 1, 2) ~ 1,
      TRUE                                                   ~ 0
    ),
    differ_at = factor(ref_labels_differ[as.character(shared_depth)],
                       levels = ref_labels_differ)
  )</pre>



<h3 class="wp-block-heading">Examining the number in each group</h3>



<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th>Subset</th><th>Pairs</th></tr></thead><tbody><tr><th colspan="2">By the group the codes share</th></tr><tr><td>All pairs</td><td>399,171</td></tr><tr><td>Same major group</td><td>24,526</td></tr><tr><td>Same minor group</td><td>7,851</td></tr><tr><td>Same broad occupation</td><td>1,097</td></tr><tr><td>Same detailed occupation</td><td>233</td></tr></tbody></table></figure>



<p class="wp-block-paragraph">The code below outputs a frequency table for each group. An important characteristic of the groups shown in the table above is that <em>most</em> job pairs span different major groups.</p>



<pre>#one row per cut and level, holding the pairs that belong to it
dta_pairs_subsets &lt;- bind_rows(
  tibble(cut_by = &quot;By where the codes first differ&quot;, depth = 0:4, subset = ref_labels_differ),
  tibble(cut_by = &quot;By the group the codes share&quot;,    depth = 0:4, subset = ref_labels_share)
) |&gt;
  mutate(
    cut_by    = fct_inorder(cut_by),
    subset = factor(subset, levels = c(ref_labels_differ, ref_labels_share)),
    pairs  = map2(cut_by, depth, \(ref_cut, ref_depth) {
      if (ref_cut == &quot;By where the codes first differ&quot;) {
        filter(rlt_pairs, shared_depth == ref_depth)
      } else {
        filter(rlt_pairs, shared_depth &gt;= ref_depth)
      }
    }),
    nmb_pairs = map_int(pairs, nrow)
  )

dta_pairs_subsets |&gt;
  select(cut_by, subset, nmb_pairs)</pre>



<h3 class="wp-block-heading">Job description vs. skill-based similarity</h3>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/semantic.png?w=450&#038;ssl=1" alt="" class="wp-image-4612" srcset_temp="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/semantic.png?w=450&#038;ssl=1 679w, https://www.gilesd-j.com/wp-content/uploads/2026/09/semantic-300x233.png 300w" sizes="auto, (max-width: 679px) 100vw, 679px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">The plot above compares similarity scores estimated from job descriptions and the skill-based scores. The relationship is positive, indicating that <em>on average</em> the same job pairs tend to have higher similarity scores using either measure. But, the fit isn’t great, particularly for job pairs with low similarity scores (which account for a large share of job pairs). The glib conclusion is that both measures tend to agree when the similarity of two jobs can be picked up by both measures, but not so much otherwise.</p>



<pre>#20,000 random pairs, fixed by the seed so the figure is the same every render.
#The same sample feeds the SOC figure below.
set.seed(20260918)
dta_sample_pairs &lt;- rlt_pairs |&gt;
  slice_sample(n = 20000)

plt_scatter_descriptions &lt;- dta_sample_pairs |&gt;
  ggplot(aes(x = sim_semantic, y = sim_ratings)) +
  geom_point(colour = ref_col_desc, alpha = 0.1, size = 0.6) +
  geom_smooth(method = &quot;loess&quot;, se = FALSE, colour = ref_col_text, linewidth = 1.1) +
  ref_theme_post +
  labs(x = &quot;Semantic similarity of descriptions (cosine; more alike to the right)&quot;,
       y = &quot;Ratings similarity (0-100)&quot;,
       title = &quot;Ratings similarity against semantic similarity of descriptions, 20,000 random pairs&quot;)

plt_scatter_descriptions</pre>



<h3 class="wp-block-heading">SOC distance vs. skill-based similarity</h3>



<p class="wp-block-paragraph">The code below uses a scatter plot to compare SOC code distance measures for job pairs with the similarity scores estimated from worker characteristics.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/soc_distance.png?w=450&#038;ssl=1" alt="" class="wp-image-4609" srcset_temp="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/soc_distance.png?w=450&#038;ssl=1 1006w, https://www.gilesd-j.com/wp-content/uploads/2026/09/soc_distance-300x157.png 300w, https://www.gilesd-j.com/wp-content/uploads/2026/09/soc_distance-768x403.png 768w" sizes="auto, (max-width: 1006px) 100vw, 1006px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">Although I spent quite a bit of time digging into how the two measures correspond with one another (or don’t), I’ve dropped quite a lot of this analysis to keep the post short. Most of the analysis told the same basic story: both distance measures provide <em>some*</em> explanatory power across occupations, but most of their power comes from separating distinct classes of work from one another. The loess curve flattening at greater distances says the rest: once two jobs have SOC list differences of 300 or more, it has little to say about how alike they might be.</p>



<p class="wp-block-paragraph">*(I was actually surprised how much explanatory power the measure had.)</p>



<pre>#ggplot2 apparently can't give two facets different axis transformations, so the two
#panels are drawn separately and placed side by side with patchwork
library(patchwork)

plt_soc_position &lt;- dta_sample_pairs |&gt;
  ggplot(aes(x = soc_rank_distance, y = sim_ratings)) +
  geom_point(colour = ref_col_soc_rank, alpha = 0.1, size = 0.6) +
  geom_smooth(method = &quot;loess&quot;, se = FALSE, colour = ref_col_text, linewidth = 1.1) +
  #reversed so closer pairs sit on the right
  scale_x_reverse(labels = label_comma()) +
  ref_theme_post +
  labs(x = &quot;Difference in list position (closer pairs to the right)&quot;,
       y = &quot;Ratings similarity (0-100)&quot;,
       subtitle = &quot;SOC list position&quot;)

plt_soc_raw &lt;- dta_sample_pairs |&gt;
  ggplot(aes(x = soc_code_distance, y = sim_ratings)) +
  geom_point(colour = ref_col_soc_raw, alpha = 0.1, size = 0.6) +
  geom_smooth(method = &quot;loess&quot;, se = FALSE, colour = ref_col_text, linewidth = 1.1) +
 
  scale_x_continuous(trans = compose_trans( &quot;reverse&quot;),
                     labels = label_comma()) +
  ref_theme_post +
  labs(x = &quot;Difference in raw code (closer pairs to the right)&quot;,
       y = NULL,
       subtitle = &quot;SOC raw code&quot;)

plt_scatter_soc &lt;- plt_soc_position + plt_soc_raw +
  plot_annotation(title = &quot;Ratings similarity against the SOC code, (20,000 random pairs)&quot;)

plt_scatter_soc</pre>



<h3 class="wp-block-heading">Correlation plot</h3>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-5.png?w=450&#038;ssl=1" alt="" class="wp-image-4580" srcset_temp="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-5.png?w=450&#038;ssl=1 729w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-5-300x197.png 300w" sizes="auto, (max-width: 729px) 100vw, 729px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">The code above summarizes the basic message as the scatter plots using a correlation matrix. All measures agree with one another, but the SOC distance measures agrees more closely with the skill-based scores than the descriptions do. This doesn’t necessarily make distance measures a better proxy, but might just point to a similar set of worker characteristics being embodied by both the structure of the SOC and the O*NET’s worker characteristic ratings (more on this below).</p>



<pre>#readable names for the four measures, in the order they were introduced
ref_labels_measure &lt;- c(
  sim_ratings        = &quot;Skill-based Ratings&quot;,
  sim_semantic       = &quot;Descriptions&quot;,
  soc_closeness      = &quot;SOC list position&quot;,
  soc_code_closeness = &quot;SOC raw code&quot;
)

#rank correlation between every pair of measures, all pairs
sum_cor &lt;- rlt_pairs |&gt;
  select(all_of(names(ref_labels_measure))) |&gt;
  cor(method = &quot;spearman&quot;) |&gt;
  as_tibble(rownames = &quot;measure_a&quot;) |&gt;
  pivot_longer(-measure_a, names_to = &quot;measure_b&quot;, values_to = &quot;correlation&quot;) |&gt;
  mutate(
    measure_a = factor(ref_labels_measure[measure_a], levels = ref_labels_measure),
    measure_b = factor(ref_labels_measure[measure_b], levels = ref_labels_measure)
  ) |&gt;
  #the matrix is symmetric, so show each pair once, below the diagonal
  filter(as.integer(measure_a) &gt; as.integer(measure_b))

plt_cor &lt;- sum_cor |&gt;
  ggplot(aes(x = measure_b, y = measure_a, fill = correlation)) +
  geom_tile(colour = &quot;white&quot;, linewidth = 1) +
  geom_text(aes(label = round_half_up(correlation, 2),
                #dark text on light tiles, light text on dark ones
                colour = correlation &gt; 0.5),
            size = 4.5) +
  scale_fill_gradient(low = &quot;white&quot;, high = ref_col_soc_rank,
                      limits = c(0, 1), name = &quot;Rank correlation&quot;) +
  scale_colour_manual(values = c(`TRUE` = &quot;white&quot;, `FALSE` = ref_col_text),
                      guide = &quot;none&quot;) +
  scale_y_discrete(limits = rev) +
  coord_equal() +
  ref_theme_post +
  theme(panel.grid = element_blank(),
        axis.text.x = element_text(angle = 30, hjust = 1)) +
  labs(x = NULL, y = NULL,
       title = &quot;Rank correlation between the four measures, all pairs&quot;)

plt_cor</pre>



<h3 class="wp-block-heading">Three measures of the same hierarchy(?)</h3>



<p class="wp-block-paragraph">The code below produces a table of average similarity scores by each measure based on job pair grouping differences.</p>



<p class="wp-block-paragraph">The table points to the SOC hierarchy doing a reasonable job of grouping similar jobs together. It also shows that most job pairs sit in different major groups and score a little above 50. It’s tempting to interpret this as indicating that the scores aren’t particularly good at differentiating jobs. But, this is actually what you’d expect of a database that’s designed to cover everything, rather than provide a representative snapshot of the labour market.</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><thead><tr><th>Where the codes first differ</th><th>Pairs</th><th>Ratings similarity</th><th>Semantic similarity</th><th>SOC rank distance</th></tr></thead><tbody><tr><td>Major group</td><td>374,645</td><td>55</td><td>0.20</td><td>316</td></tr><tr><td>Minor group</td><td>16,675</td><td>70</td><td>0.34</td><td>29</td></tr><tr><td>Broad occupation</td><td>6,754</td><td>74</td><td>0.39</td><td>11</td></tr><tr><td>Detailed occupation</td><td>864</td><td>77</td><td>0.48</td><td>3</td></tr><tr><td>Suffix only</td><td>233</td><td>76</td><td>0.44</td><td>2</td></tr></tbody></table></figure>



<pre>#average of each measure by where the codes first differ
sum_by_level &lt;- rlt_pairs |&gt;
  summarise(
    nmb_pairs         = n(),
    sim_ratings       = mean(sim_ratings),
    sim_semantic      = mean(sim_semantic),
    soc_rank_distance = mean(soc_rank_distance),
    .by = differ_at
  ) |&gt;
  arrange(differ_at) |&gt;
  mutate(sim_ratings       = round_half_up(sim_ratings, 0),
         sim_semantic      = round_half_up(sim_semantic, 2),
         soc_rank_distance = round_half_up(soc_rank_distance, 0))

sum_by_level</pre>



<p class="wp-block-paragraph">But the table also points to an uncomfortable truth about all three measures: it becomes increasingly difficult to use similarity scores to differentiate jobs once they share the same major group. Whether a job pair is in the same minor group, broad occupation, or have differences in the last two digits of the SOC code, the average similarity score shows limited movement. The plot below illustrates this by presenting the entire distribution of each measure by each group. Although the average score rises as the classification narrows, the boxes overlap so much that the classification tells you little about how alike two jobs are once they share a major group. The score still varies inside each grouping, which points to similarity scores being best used as a means to rank alternative pathways, rather than to put a precise number on how similar two jobs are.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i1.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-6.png?w=450&#038;ssl=1" alt="" class="wp-image-4584" srcset_temp="https://i1.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-6.png?w=450&#038;ssl=1 798w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-6-300x153.png 300w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-6-768x393.png 768w" sizes="auto, (max-width: 798px) 100vw, 798px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph"><strong>Note:</strong> the association between similarity measures and SOC groups is not new, it’s a design principle of the system. The SOC classifies occupations <em>based upon work performed, skills, education, training, and credentials</em><sup data-fn="fed2e599-bca3-4d9a-99a9-26064a7d49cf" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#fed2e599-bca3-4d9a-99a9-26064a7d49cf" id="fed2e599-bca3-4d9a-99a9-26064a7d49cf-link" rel="nofollow" target="_blank">8</a></sup> <em>,</em> which are many of the same characteristics the skill based scores were based on. In addition, because managers are intentionally grouped with workers they manage and/or supervise, skill similarity is likely to be deeply nested in the hierarchy. In short, a skills-based score agreeing with a system designed around skills is exactly what you’d expect. It’s also why I felt a SOC code distance measure might provide a plausible proxy for similarity scores in the first place.</p>



<pre>dta_plt_level_dist &lt;- rlt_pairs |&gt;
  transmute(differ_at,
            `Skill-based ratings`   = percent_rank(sim_ratings),
            `Job descriptions`      = percent_rank(sim_semantic),
            `SOC list position`     = percent_rank(soc_closeness)) |&gt;
  pivot_longer(-differ_at, names_to = &quot;measure&quot;, values_to = &quot;percentile&quot;) |&gt;
  mutate(measure = fct_inorder(measure))

plt_level_dist &lt;- dta_plt_level_dist |&gt;
  ggplot(aes(x = differ_at, y = percentile)) +
  geom_boxplot(fill = ref_col_grid, colour = ref_col_text, outlier.alpha = 0.03,
               outlier.size = 0.3, width = 0.6) +
  #the level means, which are the numbers in the table
  stat_summary(fun = mean, geom = &quot;point&quot;, colour = ref_col_desc, size = 2.5) +
  scale_x_discrete(labels = \(x) str_remove(x, &quot;^Differ &quot;)) +
  facet_wrap(vars(measure), nrow = 1) +
  ref_theme_post +
  theme(axis.text.x = element_text(angle = 25, hjust = 1)) +
  labs(x = &quot;Where the two codes first differ&quot;, y = &quot;Percentile of the measure&quot;,
       subtitle = &quot;All three measures rise with the hierarchy on average (gold points), and all three overlap heavily below the major group&quot;)

 
plt_level_dist</pre>



<h3 class="wp-block-heading">The explanatory power of descriptions appears to be steady across the hierarchy</h3>



<p class="wp-block-paragraph">To prove that I’m not siding with SOC distance measures as a proxy for similarity paths, the plot below presents insights resulting from a long series of arguments with Claude.</p>



<p class="wp-block-paragraph">Presented in the plot are rank correlation of each measure with skill-based scores based on how job pairs are grouped within the SOC. Notice that across the largest set of pairs, SOC distance measures agree more closely with the skill-based rankings. But, in every other group job descriptions do a better job of predicting skill-based rankings. Notice also that the raw SOC code doesn’t beat the list measure either, which points to the information lost by using the ranking being minimal, even when jobs are in the same narrow classification.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i0.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/tracks_ratings_by_subset.png?w=450&#038;ssl=1" alt="" class="wp-image-4606" srcset_temp="https://i0.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/tracks_ratings_by_subset.png?w=450&#038;ssl=1 768w, https://www.gilesd-j.com/wp-content/uploads/2026/09/tracks_ratings_by_subset-300x155.png 300w" sizes="auto, (max-width: 768px) 100vw, 768px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">Aside from this supporting the idea that the distance measure is proxying characteristics embodied by the SOC, it also points to job descriptions providing a more reliable proxy for skill similarity the deeper in the hierarchy the jobs being compared sit. And since that is where most workers are likely to look when they think about changing jobs, descriptions are likely to be the more useful measure for sense checking potential transition pathways.</p>



<pre># Rank correlation of the ratings score with each measure, on one set of
# pairs. Written as a function because it runs once per subset.
fnc_cor_with_ratings &lt;- function(dta) {
  summarise(
    dta,
    `Descriptions`  = cor(sim_ratings, sim_semantic,       method = &quot;spearman&quot;),
    `List position` = cor(sim_ratings, soc_closeness,      method = &quot;spearman&quot;),
    `Raw code`      = cor(sim_ratings, soc_code_closeness, method = &quot;spearman&quot;)
  )
}

sum_cor_by_subset &lt;- dta_pairs_subsets |&gt;
  mutate(result = map(pairs, fnc_cor_with_ratings)) |&gt;
  select(cut_by, subset, nmb_pairs, result) |&gt;
  unnest(result)

sum_cor_by_subset |&gt;
  mutate(across(where(is.numeric), \(x) round_half_up(x, 2)))

plt_cor_by_subset &lt;- sum_cor_by_subset |&gt;
  ggplot(aes(y = subset)) +
  geom_vline(xintercept = 0, colour = ref_col_grid) +
  #the gap between the two measures in each row is the story
  geom_segment(aes(x = `List position`, xend = Descriptions, yend = subset),
               colour = ref_col_grid, linewidth = 2) +
  geom_point(aes(x = `Raw code`, colour = &quot;Raw code&quot;),
             shape = 21, fill = &quot;white&quot;, size = 3.2, stroke = 1.1) +
  geom_point(aes(x = `List position`, colour = &quot;List position&quot;), size = 4) +
  geom_point(aes(x = Descriptions, colour = &quot;Descriptions&quot;), size = 4) +
  geom_text(aes(x = Descriptions, label = round_half_up(Descriptions, 2)),
            vjust = -1.2, size = 3, colour = ref_col_desc) +
  geom_text(aes(x = `List position`, label = round_half_up(`List position`, 2)),
            vjust = 2.1, size = 3, colour = ref_col_soc_rank) +
  #the pair count, at the left edge of each row
  geom_text(aes(x = -Inf, label = paste0(format(nmb_pairs, big.mark = &quot;,&quot;, trim = TRUE),
                                         &quot; pairs&quot;)),
            hjust = -0.1, size = 2.8, colour = &quot;grey50&quot;) +
  scale_colour_manual(
    values = c(Descriptions = ref_col_desc, `List position` = ref_col_soc_rank,
               `Raw code` = ref_col_soc_raw),
    name   = NULL,
    guide  = guide_legend(override.aes = list(shape = c(16, 16, 21), fill = &quot;white&quot;))
  ) +
  scale_y_discrete(limits = rev) +
  scale_x_continuous(expand = expansion(mult = c(0.25, 0.08))) +
  #one panel per cut_by, each with its own five rows
  facet_wrap(vars(cut_by), scales = &quot;free_y&quot;) +
  ref_theme_post +
  theme(panel.grid.major.y = element_blank()) +
  labs(x = &quot;Rank correlation with the ratings score&quot;, y = NULL,
       title = &quot;How each measure tracks the ratings, within each subset of pairs&quot;)

plt_cor_by_subset</pre>



<h3 class="wp-block-heading">Does the agreement hold across occupation groups?</h3>



<p class="wp-block-paragraph">The plot below presents the rank correlation of each measure with skill-based scores by occupational group, where job pairs are in the same major group.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i0.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/major_group_corr.png?w=450&#038;ssl=1" alt="" class="wp-image-4603" srcset_temp="https://i0.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/major_group_corr.png?w=450&#038;ssl=1 754w, https://www.gilesd-j.com/wp-content/uploads/2026/09/major_group_corr-300x283.png 300w" sizes="auto, (max-width: 754px) 100vw, 754px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">Once again, description-based scores outperform the SOC distance measure in most groups, and where they don’t the gap is small. The clear exception is Architecture and Engineering, which plausibly stems from the SOC hierarchy separating engineers from the technicians and drafters who support them, a technical grouping that the description-based similarity scores largely miss.</p>



<pre>#pairs where both occupations sit in the same major group. Each pair belongs to
#exactly one group, so no need to count both directions.
sum_cor_by_group &lt;- rlt_pairs |&gt;
  filter(shared_depth &gt;= 1) |&gt;
  mutate(major = str_sub(soc_code_a, 1, 2)) |&gt;
  summarise(
    nmb_pairs        = n(),
    `Descriptions`   = cor(sim_ratings, sim_semantic,  method = &quot;spearman&quot;),
    `List position`  = cor(sim_ratings, soc_closeness, method = &quot;spearman&quot;),
    .by = major
  ) |&gt;
  pivot_longer(c(Descriptions, `List position`),
               names_to = &quot;measure&quot;, values_to = &quot;correlation&quot;) |&gt;
  mutate(major = factor(major, levels = lkp_soc_major$soc_major,
                        labels = lkp_soc_major$soc_major_label))

plt_cor_by_group &lt;- sum_cor_by_group |&gt;
  ggplot(aes(y = major)) +
  geom_vline(xintercept = 0, colour = ref_col_grid) +
  geom_point(aes(x = correlation, colour = measure), size = 2.8) +
  #pair count at the left edge of each row, so the thin groups announce themselves
  geom_text(data = \(d) distinct(d, major, nmb_pairs),
            aes(x = -Inf, label = paste0(format(nmb_pairs, big.mark = &quot;,&quot;, trim = TRUE),
                                         &quot; pairs&quot;)),
            hjust = -0.1, size = 2.6, colour = &quot;grey50&quot;) +
  scale_colour_manual(values = c(Descriptions    = ref_col_desc,
                                 `List position` = ref_col_soc_rank),
                      name = NULL) +
  scale_y_discrete(limits = rev) +
  scale_x_continuous(expand = expansion(mult = c(0.3, 0.05))) +
  ref_theme_post +
  labs(x = &quot;Rank correlation with the skill-based score&quot;, y = NULL,
       title = &quot;How well each proxy measure tracks the skill-based score for pairs within a SOC major group&quot;)

plt_cor_by_group</pre>



<h3 class="wp-block-heading">Comparing the predictive power of similarity measures</h3>



<p class="wp-block-paragraph">The code below attempts to answer the question that inspired this post in the first place: if you don’t have detailed job characteristic data like the O*NET, how far might you get using national classification structures and job descriptions?</p>



<p class="wp-block-paragraph">In an attempt to answer this, the explanatory power of each measure is compared across the same five job pair groupings and a group with all job pairs. For each group, a linear regression model is used to estimate the explanatory power of measures individually and when combined together. A model with a higher adjusted R squared score explains a larger share of differences in skill-based score. The gap between the best measure and the combined model can be thought of as the marginal predictive value of adding the other measure.</p>



<figure class="wp-block-image aligncenter size-full"><img loading="lazy" decoding="async" loading="lazy" src="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-9.png?w=450&#038;ssl=1" alt="" class="wp-image-4593" srcset_temp="https://i2.wp.com/www.gilesd-j.com/wp-content/uploads/2026/09/image-9.png?w=450&#038;ssl=1 875w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-9-300x145.png 300w, https://www.gilesd-j.com/wp-content/uploads/2026/09/image-9-768x371.png 768w" sizes="auto, (max-width: 875px) 100vw, 875px" data-recalc-dims="1" /></figure>



<p class="wp-block-paragraph">The results support the same stories that were presented before:</p>



<ul class="wp-block-list">
<li class=""><strong>Neither measure has much explanatory power for skill-based scores</strong>: A model that combines both measures at once can explain about a third of variations in skill-based scores across all pairs, but only 10 to 15 percent once two jobs are from the same major group.</li>



<li class=""><strong>Both proxies appear to work as complements to one another:</strong></li>
</ul>



<p class="wp-block-paragraph"><strong>Note:</strong> p-values and standard errors have not been reported as the observations used in the model aren’t independent e.g. as the same job is included multiple times across job pairs. This isn’t expected to impact the r-squared scores, but <em>will</em> influence anything that relies on the number of independent observations.</p>



<pre>#the models, defined once: names, formulas and colours in the same order
ref_models &lt;- tribble(
  ~model,                             ~formula,
  &quot;Descriptions only&quot;,                sim_ratings ~ sim_semantic,
  &quot;SOC list position only&quot;,           sim_ratings ~ soc_closeness,
  &quot;Descriptions + SOC list position&quot;, sim_ratings ~ sim_semantic + soc_closeness
)

ref_col_models &lt;- c(
  `Descriptions only`                = ref_col_desc,
  `SOC list position only`           = ref_col_soc_rank,
  `Descriptions + SOC list position` = ref_col_combo
)

#the adjusted R-squared of each model on one set of pairs. Written as a function
#because it runs once per subset.
fnc_r_squared &lt;- function(dta) {
  ref_models |&gt;
    mutate(r_squared = map_dbl(formula,
                               \(ref_formula) summary(lm(ref_formula, data = dta))$adj.r.squared)) |&gt;
    select(model, r_squared)
}

#all pairs first, as the whole-dataset benchmark, then the pairs cut by where
#their codes first differ
sum_models &lt;- dta_pairs_subsets |&gt;
  filter(subset == &quot;All pairs&quot; | cut_by == &quot;By where the codes first differ&quot;) |&gt;
  mutate(result = map(pairs, fnc_r_squared)) |&gt;
  select(subset, nmb_pairs, result) |&gt;
  unnest(result) |&gt;
  mutate(subset = fct_relevel(subset, &quot;All pairs&quot;),
         model  = factor(model, levels = ref_models$model))

sum_models |&gt;
  mutate(r_squared = round_half_up(r_squared, 3)) |&gt;
  select(-nmb_pairs) |&gt;
  pivot_wider(names_from = model, values_from = r_squared)

plt_models &lt;- sum_models |&gt;
  ggplot(aes(y = subset)) +
  geom_vline(xintercept = 0, colour = ref_col_grid) +
  #the gap between the best single measure and the combined model is what the
  #second measure adds
  geom_segment(data = \(d) d |&gt; summarise(x = max(r_squared[model != &quot;Descriptions + SOC list position&quot;]),
                                          xend = r_squared[model == &quot;Descriptions + SOC list position&quot;],
                                          .by = subset),
               aes(x = x, xend = xend, yend = subset), colour = ref_col_grid, linewidth = 2) +
  geom_point(aes(x = r_squared, colour = model), size = 4) +
  geom_text(data = \(d) distinct(d, subset, nmb_pairs),
            aes(x = -Inf, label = paste0(format(nmb_pairs, big.mark = &quot;,&quot;, trim = TRUE), &quot; pairs&quot;)),
            hjust = -0.1, size = 2.8, colour = &quot;grey50&quot;) +
  scale_colour_manual(values = ref_col_models, name = NULL) +
  scale_y_discrete(limits = rev) +
  scale_x_continuous(expand = expansion(mult = c(0.4, 0.05))) +
  ref_theme_post +
  theme(panel.grid.major.y = element_blank()) +
  labs(x = &quot;Share of variation in the skill-based score explained (adjusted R²)&quot;, y = NULL,
       title = &quot;How much of the skill-based score each proxy recovers, for all pairs and by where the two codes first differ&quot;)

plt_models</pre>



<h3 class="wp-block-heading">What each measure brings</h3>



<p class="wp-block-paragraph">The upshot of the regression results is that neither one of the measures can stand in as a substitute for skill-based similarity scores, but <em>both</em> measures appear to hold value in the absence of occupational data as detailed as the O*NET. A classification distance measure might say little about transition pathways for jobs that resemble each other, but might provide a useful metric for splitting occupations into distinct groups between which transition is less likely. From there, similarity scores based on job descriptions might be useful for ranking potential transition paths between similarly grouped job pairs.</p>



<p class="wp-block-paragraph">The analysis also points to the measures being practically useful for sense-checking similarity scores, which was what I wanted to test in the first place. For instance, where a country’s occupational classification standards align with the ILO’s International Standard Classification of Occupations, the broad structure of occupation codes <em>should</em> behave in a similar way to the SOC. And classification systems that include text-based information of each job and occupational group provide a means for producing semantic similarity scores. Neither measure is likely to be as rich as the O*NET. But, when the O*NET can’t (or shouldn’t be) used for analysis, both provide accessible metrics for validating transition pathways estimated from non-traditional and unstructured data sources, such as online job postings, vocational curricula and survey data.<sup data-fn="4d1c3aa7-88ba-4368-af85-b87dcc1ce9ed" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#4d1c3aa7-88ba-4368-af85-b87dcc1ce9ed" id="4d1c3aa7-88ba-4368-af85-b87dcc1ce9ed-link" rel="nofollow" target="_blank">9</a></sup></p>



<h2 class="wp-block-heading">Summing up</h2>



<p class="wp-block-paragraph">The post has its origins in a project to identify potential occupational pathways for a country with limited data and a labour force that looked nothing like most of the OECD. Neither problem is unique, but the comparability of local occupations to their US counterparts is a critically important consideration when deciding whether to use the O*NET, particularly when the pathways are intended to inform public policy.<sup data-fn="f7b5fac9-e34f-451d-82a3-c7972879e493" class="fn"><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#f7b5fac9-e34f-451d-82a3-c7972879e493" id="f7b5fac9-e34f-451d-82a3-c7972879e493-link" rel="nofollow" target="_blank">10</a></sup> </p>



<p class="wp-block-paragraph">One solution to this is to leverage local data sources to develop a local database of occupational characteristics as a substitute to the O*NET. But, even <em>if</em> sufficient data and money exist to make this possible, it can be hard to know whether the identified pathways make sense. This post was meant to test whether description-based similarity measures and classification distance scores might serve as a basic sense check, which the analysis indicates they can. The classification distance measure might point to whether a pathway crosses a major occupational boundary that doesn’t make sense, the descriptions should help check if the rankings of transition pathways within a group look sane. Neither is likely to tell you that a particular pathway is correct, but together they will hopefully point to job transitions that are implausible.</p>



<p class="wp-block-paragraph">One of the things I argued with Claude about while finalizing this post, was its use of the word “cheap” to describe the two proxy measures. I didn’t like the phrasing (and I still don’t), but Claude is right that both measures are <em>cheap</em>. One is produced by subtracting one classification code from the another. The other comes from a transformer model that runs on my laptop, takes <100MB of space and was deployed across a set of job descriptions that were never meant to be comprehensive outlines of a job. The measures are <em>cheap</em>, which makes it rather extraordinary that they can explain so much of a far richer dataset.</p>



<p class="wp-block-paragraph">Another point that came to mind while writing this is that by making the skill-based score the thing to predict, I’ve implicitly assumed it’s the standard other measures should be judged against. However, it’s also possible that all three measures carry valuable information about an occupation that isn’t mutually shared. This is testable with the right data, but it’s a point worth keeping in mind as the skill-based score not aligning with either measure might also reflect it lacking important information. If so, the richest and most expensive dataset in the room comes with its own blind spots.</p>



<p class="wp-block-paragraph"><strong>How AI was used to write this post:</strong> AI produced the first draft of the code and based on the code used in my last set of analysis of the O*NET. I then proceeded to heavily edit this until it answered the questions a human being (me) might be interested in. The bulk of the writing is my own, with AI only used when I needed inspiration for improving how some points were communicated.</p>



<p class="wp-block-paragraph"></p>


<ol class="wp-block-footnotes"><li id="c857b964-35bf-4c85-8b8f-31555c2650e1">Raimi, D. and Greenspon, J., 2025. Finding the Right Fit: What Jobs Offer a Good Match for Fossil Fuel Workers’ Skills? (No. 25-06). Resources for the Future. <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#c857b964-35bf-4c85-8b8f-31555c2650e1-link" aria-label="Jump to footnote reference 1" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="33ae5a5a-24da-480e-8785-5b3ac56a9c2f">Resources for the Future, Skills Matching Explorer, <a href="https://www.rff.org/publications/data-tools/skills-matching-explorer/" rel="nofollow" target="_blank">https://www.rff.org/publications/data-tools/skills-matching-explorer/</a> <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#33ae5a5a-24da-480e-8785-5b3ac56a9c2f-link" aria-label="Jump to footnote reference 2" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="3799ac63-09ad-45f4-aa5d-a000497f84ab">Nor was it necessarily appropriate: Lo Bello, S., Sanchez Puerta, M.L. and Winkler, H., 2019. From Ghana to America: The skill content of jobs and economic development (No. 12259). IZA Discussion Papers. <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3390249" rel="nofollow" target="_blank">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3390249</a> <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#3799ac63-09ad-45f4-aa5d-a000497f84ab-link" aria-label="Jump to footnote reference 3" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="1173c6b3-4b61-451e-bf9b-866d20c4ebce">As noted in the previous post, other factors matter too, such as wage differentials, the availability (and proximity) of jobs and the intrinsic benefits of occupations being compared. <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#1173c6b3-4b61-451e-bf9b-866d20c4ebce-link" aria-label="Jump to footnote reference 4" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="74b5b1c4-493f-46fb-ba4d-a2a6ae82cfb0">U.S. Bureau of Labor Statistics (2018), 2018 SOC User Guide: Classification Principles and Coding Guidelines, <a href="https://www.bls.gov/soc/2018/soc_2018_class_prin_cod_guide.pdf" rel="nofollow" target="_blank">https://www.bls.gov/soc/2018/soc_2018_class_prin_cod_guide.pdf</a> <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#74b5b1c4-493f-46fb-ba4d-a2a6ae82cfb0-link" aria-label="Jump to footnote reference 5" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="e1bb8487-9a3f-4168-8b6d-5fc0c5068ee5">Claude suggested it and I checked if it made sense <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#e1bb8487-9a3f-4168-8b6d-5fc0c5068ee5-link" aria-label="Jump to footnote reference 6" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="4797dff9-9b8f-4fe4-83a0-8850cea9b3ee">See: Saroglou, S., Diamantaras, K., Preta, F., Delianidi, M., Benisis, A. and Meyer, C.J., 2025. Enhancing job matching: occupation, skill and qualification linking with the ESCO and EQF taxonomies. arXiv preprint arXiv:2512.03195. <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#4797dff9-9b8f-4fe4-83a0-8850cea9b3ee-link" aria-label="Jump to footnote reference 7" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="fed2e599-bca3-4d9a-99a9-26064a7d49cf">U.S. Bureau of Labor Statistics, <em>Standard Occupational Classification: User Guide</em>, “Classification Principles”. <a href="https://www.bls.gov/soc/soc-user-guide.htm" rel="nofollow" target="_blank">https://www.bls.gov/soc/soc-user-guide.htm</a>. Accessed 19 September 2026. <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#fed2e599-bca3-4d9a-99a9-26064a7d49cf-link" aria-label="Jump to footnote reference 8" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="4d1c3aa7-88ba-4368-af85-b87dcc1ce9ed">For instance, see: World Economic Forum, 2018. Towards a reskilling revolution: A future of jobs for all. Report, <a href="https://www3.weforum.org/docs/WEF_FOW_Reskilling_Revolution.pdf" rel="nofollow" target="_blank">(link)</a>; Lassébie, J., Marcolin, L., Vandeweyer, M. and Vignal, B., 2021. Speaking the same language: A machine learning approach to classify skills in Burning Glass Technologies data. OECD Social, Employment and Migration Working Papers, <a href="https://one.oecd.org/document/DELSA/ELSA/WD/SEM(2021)10/en/pdf" rel="nofollow" target="_blank">(link)</a>; and Granata, J., Posadas, J. and Testaverde, M., 2021. Indonesia’s Online Vacancy Outlook: From Online Job Postings to Labor Market Intelligence 2020. World Bank: Washington, DC, USA. (<a href="https://documents.worldbank.org/en/publication/documents-reports/documentdetail/936031636696719107/indonesia-s-online-vacancy-outlook-from-online-job-postings-to-labor-market-intelligence-2020" rel="nofollow" target="_blank">link</a>). <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#4d1c3aa7-88ba-4368-af85-b87dcc1ce9ed-link" aria-label="Jump to footnote reference 9" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li><li id="f7b5fac9-e34f-451d-82a3-c7972879e493">For instance, Lo Bello, S., Sanchez Puerta, M.L. and Winkler find large differences between non-routine and manual tasks when comparing developed and developing countries. See: Lo Bello, S., Sanchez Puerta, M.L. and Winkler, H., 2019. From Ghana to America: The skill content of jobs and economic development (No. 12259). IZA Discussion Papers. <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3390249" rel="nofollow" target="_blank">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3390249</a> <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/#f7b5fac9-e34f-451d-82a3-c7972879e493-link" aria-label="Jump to footnote reference 10" rel="nofollow" target="_blank"><img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/21a9.png?w=578&#038;ssl=1" alt="&#x21a9;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" />︎</a></li></ol><p>The post <a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/" rel="nofollow" target="_blank">O*NET ratings, job descriptions and classification codes as measures of occupational similarity</a> appeared first on <a href="https://www.gilesd-j.com/" rel="nofollow" target="_blank">Giles</a>.</p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://www.gilesd-j.com/2026/09/21/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/"> Data Analytics and AI Archives - Giles</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/onet-ratings-job-descriptions-and-classification-codes-as-measures-of-occupational-similarity/">O*NET ratings, job descriptions and classification codes as measures of occupational similarity</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403749</post-id>	</item>
		<item>
		<title>When Chain-Ladder Factors Misbehave: Smoothing Reserves with Polars and Whittaker-Henderson</title>
		<link>https://www.r-bloggers.com/2026/09/when-chain-ladder-factors-misbehave-smoothing-reserves-with-polars-and-whittaker-henderson/</link>
		
		<dc:creator><![CDATA[Christian Lorentzen]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 03:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://lorentzen.ch/?p=2096</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> Chain-Ladder reserving in polars, validated against R and Whittaker-Henderson smoothing for a difficult development pattern.</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/when-chain-ladder-factors-misbehave-smoothing-reserves-with-polars-and-whittaker-henderson/">When Chain-Ladder Factors Misbehave: Smoothing Reserves with Polars and Whittaker-Henderson</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://lorentzen.ch/index.php/2026/09/21/when-chain-ladder-factors-misbehave-smoothing-reserves-with-polars-and-whittaker-henderson/"> R – Michael&#039;s and Christian&#039;s Blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<div style="background-color: #d9edf7; color: #31708f; border-left-color: #31708f; " class="ub-styled-box ub-notification-box wp-block-ub-styled-box" id="ub-styled-box-bef7e861-d08d-4714-9cba-9fab98f5f233">
<p id="ub-styled-box-notification-content-"><strong>TD;DR</strong></p>



<ul class="wp-block-list">
<li>We run a full Chain-Ladder reserving calculation for 233 insurers in a handful of <code>polars</code> expressions, cross-checked against R’s <code>ChainLadder</code> package.</li>



<li>One insurer’s incurred development pattern turns out non-monotonic — the kind of pattern that trips up standard parametric curve fits.</li>



<li><strong>Whittaker-Henderson</strong> smoothing, newly available as <code>scipy.signal.whittaker_henderson</code>, handles it gracefully and shifts the estimated reserve by about 3%.<br></li>
</ul>


</div>


<h2 class="wp-block-heading" id="0-a-reserving-problem-solved-without-excel-"><strong>A Reserving Problem, Solved Without Excel</strong></h2>



<p>Ask a reserving actuary how they run a Chain-Ladder and you’ll usually hear “Excel” or the name of a pricey specialized tool. It turns out a modern dataframe library handles it just as well — in a few lines, for hundreds of companies at once.</p>



<p>We use the <a href="https://www.casact.org/publications-research/research/research-resources/loss-reserving-data-pulled-naic-schedule-p" rel="nofollow" target="_blank">CAS loss reserving data</a>, specifically the <em>other liability</em> line of business (LoB): 233 US insurers (“GRNAME”), 10 accident years (1998–2007), paid and incurred losses at every development lag (1-10). We treat 2007 as our reporting year, i.e. we simulate a year-end closing.</p>



<p>After reading in the data, we prepare it a bit</p>



<pre>import polars as pl

# Other Liability Data Set, December 2025
df = pl.read_csv(
    &quot;https://www.casact.org/sites/default/files/2026-03/othliab_pos_98-07.csv&quot;
)
df_triangle = (
    df
    .group_by(&quot;GRNAME&quot;, &quot;AccidentYear&quot;, &quot;DevelopmentLag&quot;)
    .agg(pl.sum(&quot;IncurredLosses&quot;).alias(&quot;Incurred&quot;), pl.sum(&quot;CumPaidLoss&quot;).alias(&quot;Paid&quot;))
    .sort(&quot;GRNAME&quot;, &quot;AccidentYear&quot;, &quot;DevelopmentLag&quot;)
)</pre>



<p>This is still the “full triangle”, i.e. all 10 development lags for all 10 accident years. Note that <code>df_triangle</code> is not in a triangle format but in a long data format, see <a href="https://doi.org/10.18637/jss.v059.i10" rel="nofollow" target="_blank">Tidy Data by Hadley Wickham</a>.</p>



<h2 class="wp-block-heading" id="1-chain-ladder-in-a-few-lines-of-polars-"><strong>Chain-Ladder in a Few Lines of Polars</strong></h2>



<p>The mechanics are the textbook ones: aggregate claims by accident year i and development lag j (starts at 1, not 0), keep only the upper-left (already observed) triangle, and compute CL development factors</p>



<pre>f^{CL}_j = \frac{\sum_{i=1998}^{2008-j} C_{i,j}}{\sum_{i=1998}^{2008-j} C_{i,j-1}}</pre>



<p>separately for every company. In polars this is just a <code>group_by</code> plus a couple of window (<code>.over(...)</code>) expressions — no manual loops over triangles required:</p>



<pre>def df_2_cl_factors(df, group_by=None):
    g = [] if group_by is None else group_by

    cl_factors = (
        df
        .group_by(g + [&quot;AccidentYear&quot;, &quot;DevelopmentLag&quot;])
        .agg(pl.sum(&quot;Paid&quot;), pl.sum(&quot;Incurred&quot;))
        # Filter upper left triangle
        .filter(REPORTING_YEAR &gt;= pl.col(&quot;AccidentYear&quot;) + pl.col(&quot;DevelopmentLag&quot;) - 1)
        .sort(g + [&quot;AccidentYear&quot;, &quot;DevelopmentLag&quot;])
        .with_columns(
            PreviousPaid=pl.col(&quot;Paid&quot;).shift(1).over(g + [&quot;AccidentYear&quot;]),
            PreviousIncurred=pl.col(&quot;Incurred&quot;).shift(1).over(g + [&quot;AccidentYear&quot;]),
        )
        # Calculate volume-weighted factors per lag
        .group_by(g + [&quot;DevelopmentLag&quot;])
        .agg(
            number_of_years=pl.len(),
            Paid=pl.col(&quot;Paid&quot;).sum(),
            PreviousPaid=pl.col(&quot;PreviousPaid&quot;).sum(),
            Incurred=pl.col(&quot;Incurred&quot;).sum(),
            PreviousIncurred=pl.col(&quot;PreviousIncurred&quot;).sum(),
        )
        .with_columns(
            f_CL_paid=pl.when(pl.col(&quot;PreviousPaid&quot;) == 0).then(1).otherwise(pl.col(&quot;Paid&quot;) / pl.col(&quot;PreviousPaid&quot;)),
            f_CL_inc=pl.when(pl.col(&quot;PreviousIncurred&quot;) == 0).then(1).otherwise(pl.col(&quot;Incurred&quot;) / pl.col(&quot;PreviousIncurred&quot;)),
        )
        .sort(g + [&quot;DevelopmentLag&quot;])
    )
    return cl_factors

cl_factors = df_2_cl_factors(df_triangle, group_by=[&quot;GRNAME&quot;])</pre>



<p>We validated the result against R’s <code>ChainLadder</code> package for one insurer, Grinnell Mut Grp, and the incurred factors matched.</p>



<p>Next, we calculate the Chain-Ladder development factors, again for both paid and incurred, this time separately for each company. Note that we take care to only account for the <em>upper left</em> triangle, which is what’s usually available in practice.</p>



<pre>def df_2_cl_factors(df, group_by=None):
    g = [] if group_by is None else group_by

    cl_factors = (
        df
        .group_by(g + [&quot;AccidentYear&quot;, &quot;DevelopmentLag&quot;])
        .agg(pl.sum(&quot;Paid&quot;), pl.sum(&quot;Incurred&quot;))
        # Filter upper left triangle
        .filter(REPORTING_YEAR &gt;= pl.col(&quot;AccidentYear&quot;) + pl.col(&quot;DevelopmentLag&quot;) - 1)
        .sort(g + [&quot;AccidentYear&quot;, &quot;DevelopmentLag&quot;])
        .with_columns(
            PreviousPaid=pl.col(&quot;Paid&quot;).shift(1).over(g + [&quot;AccidentYear&quot;]),
            PreviousIncurred=pl.col(&quot;Incurred&quot;).shift(1).over(g + [&quot;AccidentYear&quot;]),
        )
        # Calculate volume-weighted factors per lag
        .group_by(g + [&quot;DevelopmentLag&quot;])
        .agg(
            number_of_years=pl.len(),
            Paid=pl.col(&quot;Paid&quot;).sum(),
            PreviousPaid=pl.col(&quot;PreviousPaid&quot;).sum(),
            Incurred=pl.col(&quot;Incurred&quot;).sum(),
            PreviousIncurred=pl.col(&quot;PreviousIncurred&quot;).sum(),
        )
        .with_columns(
            f_CL_paid=pl.when(pl.col(&quot;PreviousPaid&quot;) == 0).then(1).otherwise(pl.col(&quot;Paid&quot;) / pl.col(&quot;PreviousPaid&quot;)),
            f_CL_inc=pl.when(pl.col(&quot;PreviousIncurred&quot;) == 0).then(1).otherwise(pl.col(&quot;Incurred&quot;) / pl.col(&quot;PreviousIncurred&quot;)),
        )
        .sort(g + [&quot;DevelopmentLag&quot;])
    )
    return cl_factors

cl_factors = df_2_cl_factors(df_triangle, group_by=[&quot;GRNAME&quot;])
cl_factors.filter(pl.col(&quot;GRNAME&quot;) == &quot;Grinnell Mut Grp&quot;)
┌────────────┬────────────┬────────────┬────────┬───┬──────────┬────────────┬───────────┬──────────┐
│ GRNAME     ┆ Developmen ┆ number_of_ ┆ Paid   ┆ … ┆ Incurred ┆ PreviousIn ┆ f_CL_paid ┆ f_CL_inc │
│ ---        ┆ tLag       ┆ years      ┆ ---    ┆   ┆ ---      ┆ curred     ┆ ---       ┆ ---      │
│ str        ┆ ---        ┆ ---        ┆ i64    ┆   ┆ i64      ┆ ---        ┆ f64       ┆ f64      │
│            ┆ i64        ┆ u32        ┆        ┆   ┆          ┆ i64        ┆           ┆          │
╞════════════╪════════════╪════════════╪════════╪═══╪══════════╪════════════╪═══════════╪══════════╡
│ Grinnell   ┆ 1          ┆ 10         ┆ 59563  ┆ … ┆ 191044   ┆ 0          ┆ 1.0       ┆ 1.0      │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 2          ┆ 9          ┆ 89554  ┆ … ┆ 172605   ┆ 166116     ┆ 1.727141  ┆ 1.039063 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 3          ┆ 8          ┆ 107323 ┆ … ┆ 151543   ┆ 151052     ┆ 1.38351   ┆ 1.003251 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 4          ┆ 7          ┆ 106828 ┆ … ┆ 129664   ┆ 131101     ┆ 1.135997  ┆ 0.989039 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 5          ┆ 6          ┆ 96179  ┆ … ┆ 105967   ┆ 106851     ┆ 1.088687  ┆ 0.991727 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 6          ┆ 5          ┆ 82541  ┆ … ┆ 86134    ┆ 86800      ┆ 1.05192   ┆ 0.992327 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 7          ┆ 4          ┆ 64703  ┆ … ┆ 66151    ┆ 66159      ┆ 1.019475  ┆ 0.999879 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 8          ┆ 3          ┆ 48924  ┆ … ┆ 49409    ┆ 49457      ┆ 1.008015  ┆ 0.999029 │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 9          ┆ 2          ┆ 32840  ┆ … ┆ 33029    ┆ 32969      ┆ 1.006343  ┆ 1.00182  │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
│ Grinnell   ┆ 10         ┆ 1          ┆ 15785  ┆ … ┆ 15915    ┆ 15908      ┆ 1.002413  ┆ 1.00044  │
│ Mut Grp    ┆            ┆            ┆        ┆   ┆          ┆            ┆           ┆          │
└────────────┴────────────┴────────────┴────────┴───┴──────────┴────────────┴───────────┴──────────┘</pre>



<p>We validated the result against R’s <code>ChainLadder</code> package for one insurer, Grinnell Mut Grp, and the incurred factors matched, see the linked notebook.</p>



<figure class="wp-block-image size-large"><img loading="lazy" fetchpriority="high" decoding="async" src="https://i1.wp.com/lorentzen.ch/wp-content/uploads/2026/09/image-1-1024x442.png?w=450&#038;ssl=1" alt="" class="wp-image-2111" srcset_temp="https://i1.wp.com/lorentzen.ch/wp-content/uploads/2026/09/image-1-1024x442.png?w=450&#038;ssl=1 1024w, https://lorentzen.ch/wp-content/uploads/2026/09/image-1-300x129.png 300w, https://lorentzen.ch/wp-content/uploads/2026/09/image-1-768x331.png 768w, https://lorentzen.ch/wp-content/uploads/2026/09/image-1.png 1050w" sizes="(max-width: 1024px) 100vw, 1024px" data-recalc-dims="1" /><figcaption class="wp-element-caption">Chain-Ladder factors for the 5 largest companies of LoB other liability, plus Grinnell Mut Grp and Virginia Mut Ins Co.</figcaption></figure>



<h2 class="wp-block-heading" id="2-a-pattern-that-doesnt-play-nice-"><strong>A Pattern That Doesn’t Play Nice</strong></h2>



<p>Zooming into the CL factors of individual companies, two stood out: the paid pattern of Virginia Mut Ins Co and the incurred pattern of Grinnell Mut Grp. Neither is monotonic — a red flag, since most commercial reserving tools only offer parametric curve fits that <em>are</em> strictly monotonic (to be fair, parametric curve fits are more for tail factor estimation).</p>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" src="https://i1.wp.com/lorentzen.ch/wp-content/uploads/2026/09/image-3.png?w=450&#038;ssl=1" alt="" class="wp-image-2113" srcset_temp="https://i1.wp.com/lorentzen.ch/wp-content/uploads/2026/09/image-3.png?w=450&#038;ssl=1 576w, https://lorentzen.ch/wp-content/uploads/2026/09/image-3-300x236.png 300w" sizes="(max-width: 576px) 100vw, 576px" data-recalc-dims="1" /></figure>



<p>Grinnell’s incurred CL factors rise above 1 early on, dip below 1 (case reserves getting released, maybe subrogation), then climb back above 1 with some zig-zag in later years. A parametric curve simply can’t represent that saddle-shaped pattern.</p>



<h2 class="wp-block-heading" id="3-enter-whittaker-henderson-"><strong>Enter Whittaker-Henderson</strong></h2>



<p><a href="https://en.wikipedia.org/wiki/Whittaker%E2%80%93Henderson_smoothing" rel="nofollow" target="_blank">Whittaker-Henderson (WH) smoothing</a> is a non-parametric, penalized smoother — no functional form assumed, just a trade-off between fitting the data and penalizing roughness. That trade-off is controlled by two knobs: the penalty <code>order</code> (we use the standard <code>order=2</code>, i.e. penalizing curvature) and the penalty strength <code>lamb</code>, which we set by eye.</p>



<p>It happens to be a perfect match here: development factors are a discrete-time signal with equal time steps, exactly what WH smoothing was designed for (it dates back to Georg Bohlmann in 1899 — arguably it should be called Bohlmann-Whittaker-Henderson). As of scipy 1.18, it ships out of the box as <code>scipy.signal.whittaker_henderson</code> — full disclosure, I contributed that implementation, so I might be a little biased towards finding excuses to use it <img src="https://i2.wp.com/s.w.org/images/core/emoji/17.0.2/72x72/1f604.png?w=578&#038;ssl=1" alt="&#x1f604;" class="wp-smiley" style="height: 1em; max-height: 1em;" data-recalc-dims="1" /></p>



<pre>from scipy.signal import whittaker_henderson

f_smooth = whittaker_henderson(
  signal=f_cl_inc, weights=previous_incurred, lamb=1e4
).x</pre>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" src="https://i2.wp.com/lorentzen.ch/wp-content/uploads/2026/09/image.png?w=450&#038;ssl=1" alt="" class="wp-image-2125" srcset_temp="https://i2.wp.com/lorentzen.ch/wp-content/uploads/2026/09/image.png?w=450&#038;ssl=1 576w, https://lorentzen.ch/wp-content/uploads/2026/09/image-300x236.png 300w" sizes="(max-width: 576px) 100vw, 576px" data-recalc-dims="1" /></figure>



<p>One neat property of WH smoothing: it preserves the weighted in-sample sum of the signal, so the smoothed factors reproduce the observed incurred losses exactly on the fitted range. Out-of-sample — i.e. for the not-yet-observed lower-right triangle — the smoothed and raw factors diverge, which is exactly where it matters for reserving.</p>



<h2 class="wp-block-heading" id="4-does-it-change-the-reserve-"><strong>Does It Change the Reserve?</strong></h2>



<p>Yes, a bit:</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><tbody><tr><td>accident year</td><td class="has-text-align-right" data-align="right">reserve CL incurred</td><td class="has-text-align-right" data-align="right">reserve smoothed</td></tr><tr><td>1998</td><td class="has-text-align-right" data-align="right">0.0</td><td class="has-text-align-right" data-align="right">0.0</td></tr><tr><td>1999</td><td class="has-text-align-right" data-align="right">7.5</td><td class="has-text-align-right" data-align="right">20.6</td></tr><tr><td>2000</td><td class="has-text-align-right" data-align="right">37.2</td><td class="has-text-align-right" data-align="right">39.0</td></tr><tr><td>2001</td><td class="has-text-align-right" data-align="right">21.5</td><td class="has-text-align-right" data-align="right">38.3</td></tr><tr><td>2002</td><td class="has-text-align-right" data-align="right">23.3</td><td class="has-text-align-right" data-align="right">13.7</td></tr><tr><td>2003</td><td class="has-text-align-right" data-align="right">-124.9</td><td class="has-text-align-right" data-align="right">-118.9</td></tr><tr><td>2004</td><td class="has-text-align-right" data-align="right">-336.1</td><td class="has-text-align-right" data-align="right">-358.5</td></tr><tr><td>2005</td><td class="has-text-align-right" data-align="right">-522.0</td><td class="has-text-align-right" data-align="right">-525.9</td></tr><tr><td>2006</td><td class="has-text-align-right" data-align="right">-482.1</td><td class="has-text-align-right" data-align="right">-456.4</td></tr><tr><td>2007</td><td class="has-text-align-right" data-align="right">394.4</td><td class="has-text-align-right" data-align="right">398.2</td></tr><tr><td><strong>TOTAL</strong></td><td class="has-text-align-right" data-align="right"><strong>-981.1</strong></td><td class="has-text-align-right" data-align="right"><strong>-950.0</strong></td></tr></tbody></table></figure>



<p>The total reserve shifts by about 3% (+31.1) — modest for the total, though individual accident years move more (some by double digits in percentage terms), since smoothing lets a factor’s neighbors pull it away from its own noisy ratio.</p>



<p>If we compare against the incurred losses after all 10 years of development — the closest proxy we have to the true ultimate loss — both CL and the WH-smoothed CL turn out to overestimate it. The un-smoothed CL is only marginally closer.</p>



<figure class="wp-block-table"><table class="has-fixed-layout"><tbody><tr><td>incurred</td><td class="has-text-align-right" data-align="right">ultimate CL</td><td class="has-text-align-right" data-align="right">ultimate smoothed</td></tr><tr><td>189,901</td><td class="has-text-align-right" data-align="right">194,067</td><td class="has-text-align-right" data-align="right">194,098</td></tr></tbody></table></figure>



<p>So for this one company, smoothing made the development pattern easier to reason about, but it didn’t make the forecast more accurate — the difference (about 31) is small compared to the 4,582 standard error Mack’s method reports for this triangle, so neither method is clearly better here. A good reminder to check against a holdout whenever one is available.</p>



<h2 class="wp-block-heading" id="5-takeaways-"><strong>Takeaways</strong></h2>



<ul class="wp-block-list">
<li>polars makes Chain-Ladder wrangling compact and fast, even across hundreds of companies at once.</li>



<li>Volume-weighted CL factors computed in polars match R’s <code>ChainLadder</code> package exactly — always reassuring when switching tools.</li>



<li>Whittaker-Henderson smoothing is a flexible alternative to parametric curve fitting whenever development factors are noisy or non-monotonic, and it’s now built into scipy. A smoother pattern isn’t automatically a more accurate one, though — always check against a holdout when you can.</li>
</ul>



<p>Natural next steps could be to add a tail factor, estimate reserve uncertainty (à la Mack’s method), or let REML pick <code>lamb</code> automatically instead of choosing it by eye.</p>



<p>The full notebook — all the polars code, the charts, and the R comparison — is on <a href="https://github.com/lorentzenchr/notebooks/blob/master/blogposts/2026-09-21%20Chain%20Ladder.ipynb" rel="nofollow" target="_blank">GitHub</a>. This post as well as the notebook was AI reviewed.</p>



<p>Spotted a bug, or have a favorite way to smooth development factors? Let me know in the comments!<br></p>



<p></p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://lorentzen.ch/index.php/2026/09/21/when-chain-ladder-factors-misbehave-smoothing-reserves-with-polars-and-whittaker-henderson/"> R – Michael&#039;s and Christian&#039;s Blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/when-chain-ladder-factors-misbehave-smoothing-reserves-with-polars-and-whittaker-henderson/">When Chain-Ladder Factors Misbehave: Smoothing Reserves with Polars and Whittaker-Henderson</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403737</post-id>	</item>
		<item>
		<title>Semi-parametric option pricing based on underlying&#8217;s historical data (accepted at the osQF 2026 (ex R/Finance) conference)</title>
		<link>https://www.r-bloggers.com/2026/09/semi-parametric-option-pricing-based-on-underlyings-historical-data-accepted-at-the-osqf-2026-ex-r-finance-conference/</link>
		
		<dc:creator><![CDATA[T. Moudiki]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://thierrymoudiki.github.io//blog/2026/09/20/r/semi-parametric-pricing-osqf</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> This post is a follow-up to my previous posts on semi-parametric option pricing. A link to the study (accepted for presentation at the osQF 2026 conference) is provided at the end of this post.</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/semi-parametric-option-pricing-based-on-underlyings-historical-data-accepted-at-the-osqf-2026-ex-r-finance-conference/">Semi-parametric option pricing based on underlying’s historical data (accepted at the osQF 2026 (ex R/Finance) conference)</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://thierrymoudiki.github.io//blog/2026/09/20/r/semi-parametric-pricing-osqf"> T. Moudiki's Webpage - R</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>This post is a follow-up to my previous posts on semi-parametric option pricing. A link to the study (accepted for presentation at the osQF 2026 conference) is provided at the end of this post.</p>

<p>In this study (R code provided), we build an empirical pricing measure for options directly from their underlying’s historical dynamics, requiring no market of option prices to calibrate against. Starting from a single observed discounted price series, we filter out linear predictability via an AR(1) fit and preserve the remaining dependence structure through a stationary block bootstrap of the residuals. The resulting empirical distribution is then given a minimal adjustment consistent with no-arbitrage: a single scalar correction, following Duan and Simonato [1998], that enforces the martingale condition required by the Fundamental Theorem of Asset Pricing (FTAP). We are explicit that this construction is not a recovery of “the” risk-neutral measure itself in the usual economic sense, but rather a pricing measure obtained by disturbing the historical dynamics as little as possible, to verify the FTAP. We validate the methodology empirically against the implied volatility surface of DAX index European options quoted July 5, 2002, comparing prices directly rather than implied volatilities. Because the construction requires no option-market input at any stage, it extends naturally to path-dependent payoffs; we illustrate this on arithmetic Asian options. 1 Motivation Standard option pricing practice requires the assumption of a parametric family for the option underlying’s dynamics (geometric Brownian motion [Black and Scholes, 1973], stochastic volatility [Heston, 1993], jump-diffusion [Merton, 1976]), and calibrating the parameters of that parametric family to observed option prices. This works well when a liquid market of option prices exists to calibrate against. It does not help when no such market exists, for examples for path-dependent or structured payoffs written on an underlying with no quoted option market (an Asian option on a private index, an embedded option inside an insurance product, a participation certificate). This note develops an alternative construction for option pricing that requires no market prices at all. It builds an empirical pricing measure directly from the historical dynamics of the underlying asset, and imposes no parametric distributional assumption beyond what the data itself exhibits. The construction is then made consistent with the Fundamental Theorem of Asset Pricing (FTAP). We first state the FTAP theorem, then show how it motivates the step of our pricing estimator, before turning to empirical validation.</p>

<p><a href="https://www.researchgate.net/publication/414513095_Semi-parametric_option_pricing_based_on_underlying's_historical_data" rel="nofollow" target="_blank">https://www.researchgate.net/publication/414513095_Semi-parametric_option_pricing_based_on_underlying’s_historical_data</a></p>

<p>I’m now interested in constructive remarks and feedback on the study, that will allow to enrich, improve and robustify the methodology.</p>

<p><img src="https://i2.wp.com/thierrymoudiki.github.io/images/2026-09-20/2026-09-20-image1.png?w=578&#038;ssl=1" alt="image-title-here" class="img-responsive" data-recalc-dims="1" /></p>


<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://thierrymoudiki.github.io//blog/2026/09/20/r/semi-parametric-pricing-osqf"> T. Moudiki's Webpage - R</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/semi-parametric-option-pricing-based-on-underlyings-historical-data-accepted-at-the-osqf-2026-ex-r-finance-conference/">Semi-parametric option pricing based on underlying’s historical data (accepted at the osQF 2026 (ex R/Finance) conference)</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403730</post-id>	</item>
		<item>
		<title>The AI Hammer</title>
		<link>https://www.r-bloggers.com/2026/09/the-ai-hammer/</link>
		
		<dc:creator><![CDATA[https://pacha.dev/blog]]></dc:creator>
		<pubDate>Sat, 19 Sep 2026 23:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://pacha.dev/blog/2026/09/20/ai-hammer/index.html</guid>

					<description><![CDATA[<p>A short reflection on the use of tools without consideration</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/the-ai-hammer/">The AI Hammer</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://pacha.dev/blog/2026/09/20/ai-hammer/index.html"> https://pacha.dev/blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p><em>Before the main content: I am creating an R Community on Google Groups. You can join the group using this <a href="https://docs.google.com/forms/d/e/1FAIpQLSdMAj4adRAT4Gyuwt_9dPvxRvOUPml9AD59vuI7qS7XDlp48g/viewform?usp=dialog" rel="nofollow" target="_blank">form</a>.</em></p>
<p>With Obama’s recent <a href="https://www.youtube.com/watch?v=Ask5wzUo9kQ" rel="nofollow" target="_blank">comments on AI</a>, I was thinking about what once was the Open Government initiative that, while remarkable in many senses, used to confuse data and information. That was back in 2009, and a large fraction of my work now and then consisted in obtaining data, this is values might pull out of a database such as region, city and variables of interest provided in formats meant to provide interpretations and drawing conclusions. In other words, a part of my work has the complication of dealing with data provided in the wrong format for such goal, as it is the case of releasing data in PDF format.</p>
<p>When an economist like me goes to a government website such as the <a href="https://www.usitc.gov/data/gravity/gravity_portal" rel="nofollow" target="_blank">USITC</a>, it is often the case that he/she is looking for raw data so that it can be analyzed and shaped in different ways, usually detecting inconsistencies, and interpreted into the form ot summary tables and plots inside a document. USITC uses the correct formats for that goal but other public offices tend to insist on the wrong formats even now that AI has become a salient issue and it is often perceived as a multi-purpose tool.</p>
<p>Sharing data in PDF format will continue to be a common practise, and often once that overly complicates scientific work and investigative journalism as it was the case of <a href="https://www.propublica.org/article/updated-dollars-for-docs-heres-whats-new" rel="nofollow" target="_blank">Dollars for Docs</a>. The popular saying is that if you have a hammer, then everything looks like a nail. AI-based workflow may increase the speed at which we can create a software prototype and still work within an organizational culture that confuses data and information.</p>
<p>With the rise of AI, government agencies can streamline the workflows that result in a PDF for effective information distribution. The information in a PDF document can be read in its electronic form, printed and reshared easily. However, adding AI to government workflows can add a vicious instead of a virtuous element. For the economist who wants raw data, PDFs made faster as in more frequent updates is the incorrect choice. The risk is that AI can become a new hammer, a tool to try to hit everything with in the same way as we continue to use PDFs without thinking of the suitability behind the technological choice.</p>
<p>I am quite sure that many data issues, not just in government but also in businesses, do not require AI at all. Around five years ago, a much less salient hype was about Machine Learning (ML), and many comments gave me the impression that ML was becoming a new hammer. In particular, many data issues such as bottlenecks and organization can be solved by moving data from multiple files into a proper Structured Query Language (SQL) database engine such as PostgreSQL or MariaDB, both with a proven track record.</p>
<p>Raw data also requires metadata describing the time range, source, and descriptions for the individual variables, and multiple important properties. A PDF seems to be useful to provide a data dictionary, and so could be a Word document or even a plain TXT files, which had led me to think about the lack of debate on the suitability of AI and many other tools. For the SQL case, it is justified for large raw datasets where Excel spreadsheets fall short, and there will be cases where Excel spreadsheet will be just right.</p>
<p>There are many other raw data formats that might be more suited to particular needs such as plain spreadsheet files (CSV). There are specific needs, as it is the case of geographical data, for which we have shapefiles (SHP) for holding spatial feature information with their respective database version known as PostGIS. Like different data format serve different goals, it is up to us to see the AI hammer in the toolbox and know when to use it.</p>
<p>How is it relevant to the R community? I have received an increasing number of low effort Pull Requests (PRs) to the different R packages that I maintain and I see the IA hammer in action. Most of those PRs failed to answer simple questions from my side such as “is this data free to be re-distributed?”. I usually get a positive feeling when I receive an email or see an open issue on GitHub from a user asking me about a particular dataset. A few times, fortunately a few ones, I have received a PR with a simple change because I documented a function with a bad wording or I had typos, and so I am thankful for those corrections that sometimes came without much debate.</p>
<p>I think that I will continue to receive AI-based PRs, which it not bad by itself, and so it is not using it to add more examples, add checks or rewrite parts of a code. My same critique will apply if, instead of AI-based changes I would be receiving changes consisting in copying and pasting without much consideration about the relevance of doing it.</p>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://pacha.dev/blog/2026/09/20/ai-hammer/index.html"> https://pacha.dev/blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/the-ai-hammer/">The AI Hammer</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403735</post-id>	</item>
		<item>
		<title>Bootstrap v traditional asymptotic normal assumptions by @ellis2013nz</title>
		<link>https://www.r-bloggers.com/2026/09/bootstrap-v-traditional-asymptotic-normal-assumptions-by-ellis2013nz/</link>
		
		<dc:creator><![CDATA[free range statistics - R]]></dc:creator>
		<pubDate>Sat, 19 Sep 2026 13:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://freerangestats.info/blog/2026/09/20/clt-boot-comparison</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> Today’s just a very short sequel to last week’s post, where I had a look at some very skewed distributions to test the idea that sample sizes sometimes need to be in the tens of thousands for the sample mean to have a normal distribution. Turns out the...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/bootstrap-v-traditional-asymptotic-normal-assumptions-by-ellis2013nz/">Bootstrap v traditional asymptotic normal assumptions by @ellis2013nz</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://freerangestats.info/blog/2026/09/20/clt-boot-comparison"> free range statistics - R</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>Today’s just a very short sequel to <a href="https://freerangestats.info/blog/2026/09/16/clt-simulations" rel="nofollow" target="_blank">last week’s post</a>, where I had a look at some very skewed distributions to test the idea that sample sizes sometimes need to be in the tens of thousands for the sample mean to have a normal distribution. Turns out they do.</p>

<p>I had a bit of unfinished business at the back of my mind, which was “would a <a href="https://en.wikipedia.org/wiki/Bootstrapping_(statistics)" rel="nofollow" target="_blank">bootstrap</a> confidence interval do any better?”. Hence today’s new set of simulations.</p>

<p>I compared the coverage of a 95% confidence interval for the mean constructed the traditional way—like they teach it in basic stats courses—from a few heavily skewed distributions. I also constructed a 95% confidence interval using the bias-corrected and adjusted bootstrap method, which I believe is the best candidate to work in a wide variety of bias and skew situations.</p>

<p>To skip to the chase, here’s the results. Turns out that a) the bootstrap does indeed do considerably better than just relying on the central limit theorem, particularly with smaller sample sizes; and b) it’s still got coverage a lot less than the 95% we wanted:</p>

<object type="image/svg+xml" data="https://freerangestats.info/img/0333-sims-results.svg" width="450"><img src="https://i1.wp.com/freerangestats.info/img/0333-sims-results.png?w=450&#038;ssl=1" data-recalc-dims="1" /></object>

<p>No surprise here; from what I understand of the history, this is pretty much exactly what the BCa bootstrap was developed for. So we’re on it’s home ground, and it does (relatively) well. But those actual coverage numbers are still well below 95%, for both methods.</p>

<p>Here’s the code that did that. It’s very similar to the one from a few days ago.</p>

<figure class="highlight"><pre>library(tidyverse)
library(actuar)
library(glue)
library(scales)
library(boot)

# set the below to TRUE if running for the first time
run_sims &lt;- FALSE

set.seed(123)

# Number of repeats for each combination of sample size and population:
today_reps &lt;- 1000

# Population size:
N &lt;- 1e6

# Modified version of the function we used last week, this time just looking at
# coverage of confidence intervals and using a BCa bootstrap to compare to the
# traditional asuymptotic CLT/normal assumed one:
sim_clt2 &lt;- function(
  x,
  n = 30,
  reps = today_reps,
  replace = TRUE,
  conf = 0.95,
  boot_R = 3001,
  ...
) {
  true_mean &lt;- mean(x)

  samples &lt;- replicate(
    reps,
    sample(x = x, size = n, replace = replace),
    simplify = FALSE
  )
  means &lt;- sapply(samples, mean)

  covered_clt &lt;- sapply(samples, function(s) {
    # rely on asymptotic normality to estimate a confidence interval and check
    # for coverage for each sample
    se &lt;- stats::sd(s) / sqrt(length(s))
    ci &lt;- mean(s) + c(-1, 1) * qnorm((1 - conf) / 2 + conf) * se
    ci[1] &lt;= true_mean &#038; true_mean &lt;= ci[2]
  })

  covered_boot &lt;- sapply(samples, function(s) {
    b &lt;- boot::boot(
      data = s,
      statistic = function(x, w) {
        mean(x[w])
      },
      R = boot_R
    )
    ci_boot_res &lt;- boot::boot.ci(b, conf = conf, type = &quot;bca&quot;)
    ci_boot &lt;- ci_boot_res$bca[4:5]
    ci_boot[1] &lt;= true_mean &#038; true_mean &lt;= ci_boot[2]
  })

  return(list(
    coverage_clt = mean(covered_clt),
    coverage_boot = mean(covered_boot)
  ))
}

# Populations we're going to use
pops &lt;- list(
  exp(rnorm(N)),
  exp(rnorm(N, sd = 2)),
  exp(rexp(N, rate = 2))
)

# Sample sizes we're going to use
ns &lt;- c(10, 30, 200, 1000)

# Run simulations:
if (run_sims) {
  results &lt;- expand_grid(pop = 1:3, n = ns) |&gt;
    mutate(coverage_clt = NA, coverage_boot = NA)

  # this - obviously when you think about what it's doing - will take a long time
  # (~2 hours) to run. It's embarassingly parallel so could consider parallelising
  # it easily enough, but there is a lot of demands on memory so for my laptop is
  # probably not going to be worth trying this as the machine wouldn't be able to
  # do multiple goes of the 3000 rep bootstrap, 1000 rep simulation from a 1e6
  # population at once.
  for (i in 1:nrow(results)) {
    cat(i)
    param &lt;- results[i, ]
    tmp &lt;- sim_clt2(pops[[param$pop]], n = param$n)
    results[i, ]$coverage_clt &lt;- tmp$coverage_clt
    results[i, ]$coverage_boot &lt;- tmp$coverage_boot
  }

  save(results, file = glue(&quot;0333-boot-results-{Sys.Date()}.rda&quot;))
} else {
  lf &lt;- sort(
    list.files(pattern = &quot;0333-boot-results.*\\.rda$&quot;),
    decreasing = TRUE
  )
  load(lf[1])
}

# labels for the populations:
pop_labs &lt;- c(&quot;log normal(0,1)&quot;, &quot;log normal(0,2)&quot;, &quot;exponential(2)&quot;)

# Draw plot:
p &lt;- results |&gt;
  mutate(lab = pop_labs[pop]) |&gt;
  mutate(lab = fct_reorder(lab, coverage_boot)) |&gt;
  ggplot(aes(x = coverage_clt, y = coverage_boot, colour = lab)) +
  geom_abline(slope = 1, intercept = 0, colour = &quot;grey50&quot;) +
  geom_point(size = 2) +
  geom_text_repel(aes(label = comma(n)), seed = 123, alpha = 0.5) +
  coord_equal() +
  scale_x_continuous(label = percent) +
  scale_y_continuous(label = percent) +
  labs(
    x = &quot;Confidence interval from asymptotic normality includes the mean&quot;,
    y = &quot;Confidence interval from BCa bootstrap includes the mean&quot;,
    title = &quot;Bootstrap outperforms asymptotic normality assumption with smaller n.&quot;,
    subtitle = &quot;Proportion of time the 95% confidence interval actually contains the true value.
Labelled numbers indicate sample sizes. Diagonal line shows equal performance.&quot;,
    colour = &quot;Population distribution:&quot;
  )

 print(p)</pre></figure>

<p>I still haven’t looked at the point—raised by Professor Harrell himself after my last post—of the assymetry of these confidence intervals, which causes a whole new set of problems. I think I’ve run out of oomph for looking at that, but it is actually an important point to remember. Maybe some time later.</p>

<p>That’s it for today really. I still think the bootstrap is a close to magic as you get in frequentist statistics, and I thoroughly recommend it. It’s good stuff. But when you’ve got a sample size of 10, 30, 200—sometimes even when you’ve got 1,000, 10,000 or 50,000—there’s just limits to what you can do.</p>


<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://freerangestats.info/blog/2026/09/20/clt-boot-comparison"> free range statistics - R</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/bootstrap-v-traditional-asymptotic-normal-assumptions-by-ellis2013nz/">Bootstrap v traditional asymptotic normal assumptions by @ellis2013nz</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403725</post-id>	</item>
		<item>
		<title>From URL to Theme: How {css2r} Steals a Website’s Colours, and Where It Gives Up</title>
		<link>https://www.r-bloggers.com/2026/09/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/</link>
		
		<dc:creator><![CDATA[Arthur Bréant]]></dc:creator>
		<pubDate>Fri, 18 Sep 2026 16:53:26 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://rtask.thinkr.fr/?p=30036</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> You can read the original post in its original format on Rtask website by ThinkR here: From URL to Theme: How {css2r} Steals a Website’s Colours, and Where It Gives Up<br />
You have thirty minutes before the demo, and your Shiny app still looks like Bootstrap 5 out of ...</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/">From URL to Theme: How {css2r} Steals a Website’s Colours, and Where It Gives Up</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://rtask.thinkr.fr/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/"> Rtask</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>You can read the original post in its original format on <a rel="nofollow" href="https://rtask.thinkr.fr/" target="_blank">Rtask</a> website by ThinkR here: <a rel="nofollow" href="https://rtask.thinkr.fr/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/" target="_blank">From URL to Theme: How {css2r} Steals a Website’s Colours, and Where It Gives Up</a></p>
<p><em>You have thirty minutes before the demo, and your Shiny app still looks like Bootstrap 5 out of the box.</em></p>
<p>We have all been there. The analysis is solid, the model is good, the tables are clean. And then someone from the communication team walks past your screen and says, very politely, that “it doesn’t really look like us”. They are right. Your app is blue-ish, your company is orange. Somewhere there is a 60-page brand book you have never opened, and you have no intention of reading it before 2pm.</p>
<p>So here is the naive question we asked ourselves at ThinkR: the company website already carries the brand. The colours are in its CSS, right there, publicly served over HTTP. Why not just go and take them?</p>
<p>That is <code>{css2r}</code>, and the small Shiny app that wraps it, <strong>Shiny Copy</strong>. You give it a URL, it gives you a <code>bslib::bs_theme()</code> call. It is live at <strong><a href="https://connect.thinkr.fr/css2r/" rel="nofollow" target="_blank">connect.thinkr.fr/css2r</a></strong>. Open it in another tab, paste your company’s address, and come back. This article will still be here.</p>
<p>It works. And the interesting part of this article is not that it works. It is <em>everything the web refuses to give you along the way</em>. Buckle up, we’re talking about CSS. <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f3a8.png" alt="🎨" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>
<h2>Thirty seconds: a URL in, a theme out</h2>
<pre>remotes::install_github(&quot;ThinkR-open/css2r&quot;)
library(css2r)
thinkr &lt;- css2r$new(url = &quot;https://thinkr.fr&quot;)
#&gt; &#x2714; Internet ok
#&gt; &#x2714; html page downloaded
#&gt; &#x2714; CSS links extracted
#&gt; &#x2714; CSS links filtered
#&gt; &#x2714; CSS downloaded
#&gt; &#x2714; Colors extracted successfully
#&gt; &#x2714; Colors analyzed successfully.
#&gt; &#x2139; No Google Fonts detected.
#&gt; &#x2714; Shiny theme code generated.</pre>
<p>Nine steps, a few seconds, and the object holds everything it found:</p>
<pre>thinkr$top_colors
#&gt; $white_black
#&gt;      Color Count
#&gt; 32 #FFFFFF     1
#&gt;
#&gt; $top_colors
#&gt;     Color Count
#&gt; 1 #38404C   133
#&gt; 2 #0046C8    72
#&gt; 3 #F05622    40
#&gt; 4 #20B8D6    23
cat(thinkr$shiny_code)
#&gt; fluidPage(
#&gt;   theme = bslib::bs_theme(
#&gt;     bg = &quot;#FFFFFF&quot;,
#&gt;     fg = &quot;#38404C&quot;,
#&gt;     primary = &quot;#38404C&quot;,
#&gt;     secondary = &quot;#0046C8&quot;
#&gt;   ),
#&gt;   h1(&quot;Hello World primary&quot;, class = &quot;text-center text-secondary&quot;),
#&gt;   h1(&quot;Hello World secondary&quot;, class = &quot;text-center text-primary&quot;)
#&gt; )</pre>
<p>Copy, paste into your <code>app_ui.R</code>, done.</p>
<p>And if you prefer clicking to typing, that is exactly what the hosted version does: swatches, a copy button, and a live Bootstrap mockup showing what your app would look like before you write a single line. Locally, <code>css2r::run_app()</code> launches the same thing.</p>
<p>Now look at that theme again. Look at it properly.</p>
<p><code>#F05622</code> is the ThinkR orange. The one on the logo, on the slides, on the mugs. It came <strong>third</strong>, with 40 occurrences. It is nowhere in the generated theme. What <code>{css2r}</code> picked as <code>primary</code> is <code>#38404C</code>, a dark slate grey: the body-text colour, which by construction appears everywhere.</p>
<p>We ran our own tool on our own website and it did not find our own brand colour. That deserves an explanation, and the explanation is the rest of this article.</p>
<h2>Under the hood: nine steps and one regular expression</h2>
<p>The <code>css2r</code> R6 class is deliberately small. The pipeline is linear and every step can be called by hand with <code>on_initialize = FALSE</code>:</p>
<ol style="list-style-type: decimal">
<li><code>check_internet()</code>, via <code>{curl}</code></li>
<li><code>download_html()</code>, via <code>{rvest}</code></li>
<li><code>extract_css_links()</code>, every <code>&lt;link rel=&quot;stylesheet&quot;&gt;</code></li>
<li><code>filter_css_links()</code>, keeping only the ones on the same domain</li>
<li><code>download_css_files()</code>, via <code>{httr}</code></li>
<li><code>extract_colors()</code>, the heart of the machine</li>
<li><code>analyze_colors()</code>, splitting neutrals from brand colours</li>
<li><code>detect_google_fonts()</code>, reading the Google Fonts URL parameters</li>
<li><code>generate_shiny_code()</code>, assembling the <code>bs_theme()</code> call</li>
</ol>
<p>Step 6, the one that does all the real work, is five lines:</p>
<pre>pattern &lt;- &quot;#[0-9A-Fa-f]{6}&quot;
matches &lt;- gregexpr(pattern = pattern, text = self$css_content, perl = TRUE)
colors_hex &lt;- regmatches(x = self$css_content, m = matches) |&gt;
  unlist() |&gt;
  toupper()
table_colors_hex &lt;- colors_hex |&gt; table() |&gt; as.data.frame(stringsAsFactors = FALSE)</pre>
<p>Find every six-digit hex code, count how many times each one appears, sort. Then <code>analyze_colors()</code> puts <code>#FFFFFF</code> and <code>#000000</code> aside as neutrals, keeps the top four, and <code>generate_shiny_code()</code> assigns background and foreground by relative luminance:</p>
<pre>hex_luminance = function(hex) {
  hex &lt;- sub(&quot;^#&quot;, &quot;&quot;, hex)
  r &lt;- strtoi(substr(hex, 1, 2), 16L) / 255
  g &lt;- strtoi(substr(hex, 3, 4), 16L) / 255
  b &lt;- strtoi(substr(hex, 5, 6), 16L) / 255
  0.2126 * r + 0.7152 * g + 0.0722 * b
}</pre>
<p>Lightest colour becomes <code>bg</code>, darkest becomes <code>fg</code>, most frequent becomes <code>primary</code>, second becomes <code>secondary</code>. That is the entire theory.</p>
<p>It is a heuristic, and a good one: it produces a usable theme on a large share of ordinary websites. But it rests on three assumptions about the web, and the modern web breaks all three.</p>
<h2>Why this is harder than it looks</h2>
<p>To make this concrete rather than theoretical, we pointed the extraction at a handful of real sites in August 2026 and counted what was actually reachable. The results are more instructive than any diagram.</p>
<table style="width:100%;">
<colgroup>
<col width="11%" />
<col width="11%" />
<col width="11%" />
<col width="11%" />
<col width="11%" />
<col width="11%" />
<col width="11%" />
<col width="11%" />
<col width="11%" />
</colgroup>
<thead>
<tr class="header">
<th>Site</th>
<th>CSS links</th>
<th>Same-domain</th>
<th>Inline <code>&lt;style&gt;</code></th>
<th>6-digit hex</th>
<th>3-digit hex</th>
<th><code>rgb()</code></th>
<th><code>var(--…)</code></th>
<th>Verdict</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>thinkr.fr</td>
<td>4</td>
<td>3</td>
<td>3 blocks (8.3 kB)</td>
<td>309</td>
<td>234</td>
<td>70</td>
<td>37</td>
<td>works, wrong primary</td>
</tr>
<tr class="even">
<td>lemonde.fr</td>
<td>2</td>
<td>2</td>
<td>0</td>
<td>805</td>
<td>140</td>
<td>82</td>
<td>707</td>
<td>runs, all-grey theme</td>
</tr>
<tr class="odd">
<td>posit.co</td>
<td>5</td>
<td>4</td>
<td>1</td>
<td>137</td>
<td>42</td>
<td>72</td>
<td><strong>1 545</strong></td>
<td>works, misses the palette</td>
</tr>
<tr class="even">
<td>tailwindcss.com</td>
<td>2</td>
<td>2</td>
<td>0</td>
<td>973</td>
<td>46</td>
<td>7</td>
<td><strong>8 371</strong></td>
<td>300 unique colours, no signal</td>
</tr>
<tr class="odd">
<td>stripe.com</td>
<td>5</td>
<td><strong>0</strong></td>
<td>0</td>
<td>n/a</td>
<td>n/a</td>
<td>n/a</td>
<td>n/a</td>
<td><strong>total failure</strong></td>
</tr>
</tbody>
</table>
<p>Three distinct failure modes hide in that table.</p>
<h3>1. The CSS you cannot reach</h3>
<p><code>filter_css_links()</code> keeps only stylesheets served from the same domain as the page:</p>
<pre>keep(.p = ~ domain(.x) == self$domain)</pre>
<p>That rule exists for a good reason. On thinkr.fr it correctly throws away a <code>highlight.js</code> theme hosted on cdnjs, a syntax-highlighting palette that has nothing to do with our brand and would have polluted the count with random purples.</p>
<p>On stripe.com, the same rule throws away <strong>everything</strong>. All five stylesheets live on <code>b.stripecdn.com</code>. Same company, different hostname, and <code>{css2r}</code> returns empty-handed:</p>
<blockquote><p>
<strong>External stylesheets only.</strong> This website loads 5 stylesheet(s), but all of them are served from another domain, so ShinyCopy cannot tell brand CSS from third-party CSS.
</p></blockquote>
<p>Any organisation that serves its assets from a dedicated CDN, which is to say most large organisations, is invisible to us.</p>
<p>There is a subtlety here that cost us a bug, and it is worth naming because it is the kind of mistake that hides behind a plausible error message. <code>urltools::domain()</code> does not return the registrable domain despite its name: it returns the full host, subdomain included. So for a site that redirects its apex to <code>www</code>, the page we downloaded and the stylesheets it references ended up on <code>www.example.com</code> while the filter was still comparing against the <code>example.com</code> the user had typed. Every stylesheet was discarded, and the app confidently blamed inline CSS. The filter now compares eTLD+1, and resolves everything against the URL actually reached after redirects, which keeps the useful behaviour (cdnjs still goes) without the self-inflicted wound.</p>
<p>Then there is the CSS that never appears in a <code>&lt;link&gt;</code> at all. thinkr.fr ships three inline <code>&lt;style&gt;</code> blocks, 8.3 kB of critical CSS injected straight into the HTML for faster first paint. It is a completely standard performance practice, and <code>extract_css_links()</code> walks right past it because it only looks at <code>link[rel='stylesheet']</code>. Shiny Copy at least has the decency to tell you:</p>
<blockquote><p>
<strong>No stylesheet found.</strong> No <code>&lt;link rel=&quot;stylesheet&quot;&gt;</code> tag was found on thinkr.fr. This usually means the site ships its CSS inline in the HTML, or renders the page with JavaScript. ShinyCopy supports neither yet.
</p></blockquote>
<p>And finally, the pages that have no meaningful HTML at load time because everything is rendered client-side by JavaScript. <code>{rvest}</code> fetches the document as the server sent it, not as the browser eventually paints it. For a single-page app, there is often nothing to read.</p>
<h3>2. The colours you cannot parse</h3>
<p>This one is my favourite, because the tool fails <em>silently</em> and <em>partially</em>.</p>
<p><code>#[0-9A-Fa-f]{6}</code> matches six-digit hex codes. Nothing else. On thinkr.fr that means we see 309 colour declarations and quietly ignore <strong>234 three-digit hex codes</strong> (<code>#fff</code>, <code>#333</code>) and <strong>70 <code>rgb()</code> / <code>rgba()</code> calls</strong>. Roughly forty percent of the colour information on our own site is dropped on the floor, and nothing in the output tells you so. Add <code>hsl()</code>, which appears 60 times on lemonde.fr, plus CSS named colours like <code>rebeccapurple</code>, and the blind spot widens further.</p>
<p>The fix is not conceptually hard: normalise every colour notation to hex before counting. It is just work that has not been done yet.</p>
<p>The deeper problem is CSS custom properties. Look at that <code>var(--…)</code> column again. On tailwindcss.com there are <strong>8 371</strong> references to CSS variables and on posit.co <strong>1 545</strong>. A modern design system declares its brand colour exactly once:</p>
<pre>:root { --brand-primary: #F05622; }</pre>
<p>…and then uses it a thousand times through <code>var(--brand-primary)</code>. Our frequency counter sees that colour <strong>once</strong>. Meanwhile a grey declared inline in fifteen legacy components scores fifteen. We are not measuring importance, we are measuring how badly the stylesheet has aged.</p>
<h3>3. Frequency is not identity</h3>
<p>Which brings us back to the ThinkR orange, and to the assumption underneath the whole approach: <em>the most frequent colour is the most important colour</em>.</p>
<p>It simply is not true. The most frequent colour in a stylesheet is almost always the interface plumbing: text grey, border grey, disabled grey. Brand colours are used sparingly, and that is the entire point of a brand colour. A logo orange that appears on three buttons carries far more identity than a <code>#38404C</code> used on every paragraph of the site.</p>
<p><img decoding="async" src="https://i1.wp.com/thinkr.fr/wp-content/uploads/frequency-en.png?w=578&#038;ssl=1" data-recalc-dims="1" /></p>
<p>lemonde.fr is the purest illustration, and a worse case than our own site. Every step of the pipeline succeeds, no error is raised, and this comes out:</p>
<pre>$top_colors
    Color Count
1 #2A303C    77
2 #E8EAEE    62
3 #F5F6F8    40
4 #464F5F    25
bslib::bs_theme(bg = &quot;#FFFFFF&quot;, fg = &quot;#2A303C&quot;,
                primary = &quot;#2A303C&quot;, secondary = &quot;#E8EAEE&quot;)</pre>
<p>All four are greys. Le Monde’s blue, <code>#015AAD</code>, sits at <strong>rank 10</strong> with 17 occurrences: ranks one through nine are neutrals, every single one. And <code>secondary</code> lands on <code>#E8EAEE</code>, a contrast ratio of <strong>1.09</strong> against the white background. A <code>.btn-secondary</code> would be, quite literally, invisible.</p>
<p>On tailwindcss.com the top two are white (214) and black (129), and since <code>analyze_colors()</code> deliberately sets those aside, we fall through to <code>#030712</code>, which is… very slightly less black. Below that, 300 unique colours with no meaningful frequency gap between them, because a design system publishes its entire palette in one file.</p>
<p>There is one more consequence worth naming. On thinkr.fr, <code>#38404C</code> is both the most frequent colour <em>and</em> the darkest one, so it becomes <code>fg</code> <strong>and</strong> <code>primary</code> at the same time. A <code>.text-primary</code> element is then exactly the same colour as body text: technically valid, visually pointless. <code>{css2r}</code> performs no contrast check and no de-duplication between roles.</p>
<h2>So we asked a human</h2>
<p>Here is the pragmatic conclusion we reached while building Shiny Copy: if the machine cannot reliably tell a brand colour from interface plumbing, do not let it decide alone.</p>
<p>The app extracts the top four candidates, generates its best guess, and then displays clickable swatches for the <code>primary</code> and <code>secondary</code> roles. One click re-themes the whole preview live:</p>
<pre>observeEvent(input$preview_primary, {
  rv$preview_primary &lt;- input$preview_primary
  session$setCurrentTheme(
    bslib::bs_theme_update(
      rv$site$shiny_theme,
      primary   = rv$preview_primary,
      secondary = rv$preview_secondary
    )
  )
})</pre>
<p>The frequency ranking stops being an answer and becomes a shortlist. You still get your ThinkR orange. You just have to click on it, which takes about a second, and you are the one who knows what your brand looks like.</p>
<p>Except that a shortlist is only ever as good as the ranking that produced it, and lemonde.fr shows exactly where that breaks: the four swatches on offer are four greys, and the brand blue at rank 10 is not among them. There is nothing left for the human to rescue. Deferring to the user fixes a ranking that is merely imperfect; it does nothing for one that is wrong from the first row.</p>
<h2>What it would take to go further</h2>
<p>The limits above are not mysteries, they are a roadmap. In rough order of payoff:</p>
<ul>
<li><strong>Rank by chroma, not just by count.</strong> The cheapest fix with the largest visible effect. Discard near-neutrals before choosing <code>primary</code>, using nothing more than the spread between the RGB channels:
<pre>chroma &lt;- function(hex) {
  v &lt;- c(strtoi(substr(hex, 2, 3), 16L),
         strtoi(substr(hex, 4, 5), 16L),
         strtoi(substr(hex, 6, 7), 16L))
  max(v) - min(v)
}</pre>
<p><code>#38404C</code> scores 18, <code>#F05622</code> scores 206. The separation is not subtle. Applied to our two examples, <code>primary</code> / <code>secondary</code> would go from <code>#38404C</code> / <code>#0046C8</code> to <code>#0046C8</code> / <code>#F05622</code> on thinkr.fr, and from two greys to <code>#015AAD</code> on lemonde.fr. It also happens to solve the <code>primary == fg</code> collision for free, since a dark neutral can no longer win.</p>
</li>
<li>
<p><strong>Normalise colour notations.</strong> A <code>parse_css_color()</code> helper turning <code>#fff</code>, <code>rgb()</code>, <code>rgba()</code>, <code>hsl()</code> and named colours into a canonical hex would immediately recover the ~40% we currently drop.</p>
</li>
<li>
<p><strong>Read inline <code>&lt;style&gt;</code> blocks.</strong> One extra <code>html_nodes(&quot;style&quot;) |&gt; html_text()</code> and a lot of critical CSS stops being invisible.</p>
</li>
<li>
<p><strong>Handle asset CDNs.</strong> Comparing registrable domains fixed the apex/www case, but <code>b.stripecdn.com</code> is genuinely a different domain from <code>stripe.com</code> and no amount of string comparison will bridge that. It needs a different rule: an allowlist, or scoring stylesheets rather than excluding them outright.</p>
</li>
<li>
<p><strong>Resolve CSS custom properties.</strong> Parse <code>--var: #hex</code> declarations, then count <code>var(--var)</code> usages against them. This is the single biggest fidelity gain available, and the one that would make the tool work on modern design systems at all.</p>
</li>
<li>
<p><strong>Weight by context, not by count.</strong> A colour on a <code>.btn</code> background is not worth the same as a gradient stop in a hero image. That means a real CSS parser instead of a regular expression. A much bigger project, and the point where “small useful tool” becomes “product”.</p>
</li>
<li>
<p><strong>Check contrast before emitting.</strong> A quick WCAG ratio on the <code>bg</code> / <code>fg</code> pair, and a rule preventing <code>primary</code> from collapsing onto <code>fg</code>.</p>
</li>
<li>
<p><strong>Headless rendering</strong> for JavaScript-heavy pages, via <code>{chromote}</code>. Powerful, heavy, and a whole other class of deployment problem.</p>
</li>
</ul>
<p>We shipped without any of these on purpose. A tool that solves the common case in five seconds and tells you honestly when it cannot is more useful than one that never leaves the lab. Every one of those error modals in Shiny Copy exists so that you find out immediately, rather than shipping a grey app and wondering why.</p>
<h2>Try it, break it, tell us</h2>
<p>The app is live and free, no installation required:</p>
<p><strong><img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f449.png" alt="👉" class="wp-smiley" style="height: 1em; max-height: 1em;" /> <a href="https://connect.thinkr.fr/css2r/" rel="nofollow" target="_blank">connect.thinkr.fr/css2r</a></strong></p>
<p>And if you would rather have it in your own scripts:</p>
<pre>remotes::install_github(&quot;ThinkR-open/css2r&quot;)
css2r::run_app()</pre>
<p>Point it at your company website. If it works, you have saved yourself an afternoon in the brand book. If it fails, we would genuinely like to know which of the three failure modes above caught you. The source is on <a href="https://github.com/ThinkR-open/css2r" rel="nofollow" target="_blank">GitHub</a>, issues and pull requests are very welcome, and the colour-notation parser is a lovely first contribution. <img src="https://s.w.org/images/core/emoji/13.0.0/72x72/1f609.png" alt="😉" class="wp-smiley" style="height: 1em; max-height: 1em;" /></p>
<p>And if what you actually need is a Shiny application that is properly branded, properly tested and properly deployed, <a href="https://thinkr.fr/" rel="nofollow" target="_blank">get in touch</a>. That part we do by hand.</p>
<hr />
<p><em>Figures measured in August 2026; websites change, so your numbers may differ.</em></p>
<p>This post is better presented on its original ThinkR website here: <a rel="nofollow" href="https://rtask.thinkr.fr/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/" target="_blank">From URL to Theme: How {css2r} Steals a Website’s Colours, and Where It Gives Up</a></p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://rtask.thinkr.fr/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/"> Rtask</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/from-url-to-theme-how-css2r-steals-a-websites-colours-and-where-it-gives-up/">From URL to Theme: How {css2r} Steals a Website’s Colours, and Where It Gives Up</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403704</post-id>	</item>
		<item>
		<title>R Batteries Included (R 4.6.1 with OpenBLAS and Intel MKL) is now publicly accessible</title>
		<link>https://www.r-bloggers.com/2026/09/r-batteries-included-r-4-6-1-with-openblas-and-intel-mkl-is-now-publicly-accessible/</link>
		
		<dc:creator><![CDATA[https://pacha.dev/blog]]></dc:creator>
		<pubDate>Thu, 17 Sep 2026 23:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://pacha.dev/blog/2026/09/18/r-batteries-included-open/index.html</guid>

					<description><![CDATA[<div style = "width:60%; display: inline-block; float:left; "> R on Windows made faster</div>
<div style = "width: 40%; display: inline-block; float:right;"></div>
<div style="clear: both;"></div>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/r-batteries-included-r-4-6-1-with-openblas-and-intel-mkl-is-now-publicly-accessible/">R Batteries Included (R 4.6.1 with OpenBLAS and Intel MKL) is now publicly accessible</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://pacha.dev/blog/2026/09/18/r-batteries-included-open/index.html"> https://pacha.dev/blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p><em>Before the main content: I am creating an R Community on Google Groups. You can join the group using this <a href="https://docs.google.com/forms/d/e/1FAIpQLSdMAj4adRAT4Gyuwt_9dPvxRvOUPml9AD59vuI7qS7XDlp48g/viewform?usp=dialog" rel="nofollow" target="_blank">form</a>.</em></p>
<p>Have you ever heard someone claim that R is “too slow”?</p>
<p>As a Linux user who does almost everything on the command line, I used to find this idea hard to believe. R has always been quite efficient for my work. But recently, everything clicked.</p>
<p>While attending an R Sprint, I had to run R inside a Windows virtual machine to test that my suggested improvements to R worked there too. Linear models and some mathematical functions ran much slower compared to another Linux virtual machine setup.</p>
<p>Here is why: When you install Linux, something naive such as “pacman -S r” manages high-performance numerical libraries (like OpenBLAS) and makes sure everything is neat. On Windows, base R doesn’t ship with these optimizations out of the box, and setting them up requires an intricate separate installation.</p>
<p>A while back, Microsoft R Open solved this exact bottleneck by shipping with Intel MKL built-in. MRO is now retired and it closed a gap.</p>
<p>To bridge this gap again, I created R Batteries Included. It is a custom R distribution for Windows that ships pre-configured with OpenBLAS and Intel MKL. No messy setups and no extra steps. Just high-performance R on Windows, right out of the box.</p>
<p><a href="https://www.linkedin.com/in/g-nono-gueye-ph-d-4106a445/" rel="nofollow" target="_blank">G. N. Gueye</a> suggested that my R distribution that makes R fast on Windows should be openly available with sponsorships. I followed that advice and R Batteries Included is here:</p>
<ul>
<li><a href="https://github.com/pachadotdev/r-batteries-included" rel="nofollow" target="_blank">Code</a></li>
<li><a href="https://github.com/pachadotdev/r-batteries-included/releases/download/4.6.1/R-4.6.1-Batteries-Included.exe" rel="nofollow" target="_blank">Windows installer</a></li>
</ul>
<p>See the plots showing the speed gains. I used the famous AT&#038;T Benchmark.</p>
<p><img src="https://i0.wp.com/pacha.dev/blog/2026/09/18/r-batteries-included-open/bench1.jpeg?w=50%25&#038;ssl=1"  data-recalc-dims="1"> <img src="https://i1.wp.com/pacha.dev/blog/2026/09/18/r-batteries-included-open/bench2.jpeg?w=50%25&#038;ssl=1"  data-recalc-dims="1"> <img src="https://i0.wp.com/pacha.dev/blog/2026/09/18/r-batteries-included-open/bench3.jpeg?w=50%25&#038;ssl=1"  data-recalc-dims="1"></p>
<p>Please re-post, try it, and add a star to the repository. If you can, consider sponsoring this work on:</p>
<ul>
<li><a href="https://github.com/sponsors/pachadotdev" rel="nofollow" target="_blank">GitHub</a></li>
<li><a href="https://buymeacoffee.com/pacha" rel="nofollow" target="_blank">Buy me a Coffee</a></li>
</ul>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://pacha.dev/blog/2026/09/18/r-batteries-included-open/index.html"> https://pacha.dev/blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/r-batteries-included-r-4-6-1-with-openblas-and-intel-mkl-is-now-publicly-accessible/">R Batteries Included (R 4.6.1 with OpenBLAS and Intel MKL) is now publicly accessible</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403702</post-id>	</item>
		<item>
		<title>Join the R-Community on Google Groups and the next London R Users Group Social Event</title>
		<link>https://www.r-bloggers.com/2026/09/join-the-r-community-on-google-groups-and-the-next-london-r-users-group-social-event/</link>
		
		<dc:creator><![CDATA[https://pacha.dev/blog]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 23:00:00 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://pacha.dev/blog/2026/09/17/r-community/index.html</guid>

					<description><![CDATA[<p>Think about Stackoverflow/Reddit minus the negativity.</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/join-the-r-community-on-google-groups-and-the-next-london-r-users-group-social-event/">Join the R-Community on Google Groups and the next London R Users Group Social Event</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://pacha.dev/blog/2026/09/17/r-community/index.html"> https://pacha.dev/blog</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<h2>R-Community on Google Groups</h2>
<p>I am in the process of creating a welcoming, collaborative space where anyone can ask and answer R questions without judgment and share R packages or their work using R. Think of it as the helpful, supportive sides of StackOverflow or Reddit, minus the negativity and elitism.</p>
<p>The idea of using Google Groups is to avoid making people creating accounts on different platforms, as these days the R community is spreaded between LinkedIn, X, Bluesky, Mastodon, and others I have not explored. Google Groups works with any email regardless if it is hosted by Google and you can interact with it from groups.google.com in the browser or from your email client.</p>
<p>No matter your programming style, all R preferences are welcome here. Whether you are working with base R, the Tidyverse, data.table, or exploring niche packages in the wider R ecosystem, your questions and insights are valuable. The idea is are all here to learn from one another.</p>
<p>Code of Conduct: Strictly friendly and welcoming environment. Zero-tolerance policy for bullying or hate.</p>
<p>Click here to join the community form: <a href="https://docs.google.com/forms/d/e/1FAIpQLSdMAj4adRAT4Gyuwt_9dPvxRvOUPml9AD59vuI7qS7XDlp48g/viewform" rel="nofollow" target="_blank">https://docs.google.com/forms/d/e/1FAIpQLSdMAj4adRAT4Gyuwt_9dPvxRvOUPml9AD59vuI7qS7XDlp48g/viewform</a></p>

<h2>London R Users Group: Upcoming Social</h2>
<p>How does meeting up during the second week of October sound?</p>
<p>I am planning an informal social event at a local Wetherspoons. It will be a completely relaxed evening to grab a pint (or a soft drink!), catch up in person, and chat about all things R. Whether you are a seasoned package developer or just starting your data journey, I would love to have you there.</p>
<p>If you are interested in coming along, please let me know:</p>
<ul>
<li><a href="https://pacha.dev/blog/2026/09/17/r-community/m.vargas.sepulveda@gmail.com" rel="nofollow" target="_blank">Email me directly</a></li>
<li>Or join the discussion over on the R-Community Google Group</li>
</ul>
<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://pacha.dev/blog/2026/09/17/r-community/index.html"> https://pacha.dev/blog</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/join-the-r-community-on-google-groups-and-the-next-london-r-users-group-social-event/">Join the R-Community on Google Groups and the next London R Users Group Social Event</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403685</post-id>	</item>
		<item>
		<title>coinclp: the COIN-OR Clp linear programming solver is back on CRAN</title>
		<link>https://www.r-bloggers.com/2026/09/coinclp-the-coin-or-clp-linear-programming-solver-is-back-on-cran/</link>
		
		<dc:creator><![CDATA[Sam Lovick]]></dc:creator>
		<pubDate>Wed, 16 Sep 2026 07:05:34 +0000</pubDate>
				<category><![CDATA[R bloggers]]></category>
		<guid isPermaLink="false">https://lovickconsulting.com/2026/09/16/coinclp-coin-or-clp-back-on-cran/</guid>

					<description><![CDATA[<p>Two small R packages of mine have gone to CRAN. coinclp binds the COIN-OR Clp linear programming solver and was accepted on […]<br />
The post coinclp: the COIN-OR Clp linear programming solver is back on CRAN appeared first on Sam Lovick Consulting.</p>
<strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/coinclp-the-coin-or-clp-linear-programming-solver-is-back-on-cran/">coinclp: the COIN-OR Clp linear programming solver is back on CRAN</a>]]></description>
										<content:encoded><![CDATA[<!-- 
<div style="min-height: 30px;">
[social4i size="small" align="align-left"]
</div>
-->

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 12px;">
[This article was first published on  <strong><a href="https://lovickconsulting.com/2026/09/16/coinclp-coin-or-clp-back-on-cran/"> R Archives - Sam Lovick Consulting</a></strong>, and kindly contributed to <a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers</a>].  (You can report issue about the content on this page <a href="https://www.r-bloggers.com/contact-us/">here</a>)
<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div>
<p>Two small R packages of mine have gone to CRAN. <a href="https://cran.r-project.org/package=coinclp" rel="nofollow" target="_blank">coinclp</a> binds the <a href="https://github.com/coin-or/Clp" rel="nofollow" target="_blank">COIN-OR Clp</a> linear programming solver and was accepted on 15 September 2026; <a href="https://github.com/SamLovick/ROI.plugin.coinclp" rel="nofollow" target="_blank">ROI.plugin.coinclp</a> registers it with the R Optimization Infrastructure and is in the submission queue behind it. Between them they put back something R lost at the end of 2021, and this post is a short account of what that was, what the new packages do, and — because it would be dishonest to leave it out — when you should reach for a different solver altogether.</p>
<h2>Where this comes from</h2>
<p>Clp is one of the original COIN-OR codes. COIN-OR started inside IBM Research in 2000 as an open-source home for operations research software, and Clp — John Forrest’s simplex code, with a barrier method alongside — has been its linear programming workhorse ever since. It is mature, free under the Eclipse Public License, and good at the thing simplex codes are for: solving a model, changing it a little, and solving it again from the basis you already have.</p>
<p>R had a binding to it for a decade. Gabriel Gelius-Dietrich’s <a href="https://cran.r-project.org/package=clpAPI" rel="nofollow" target="_blank">clpAPI</a>, written at Heinrich Heine University Düsseldorf for the <em>sybil</em> metabolic-modelling toolkit, was on CRAN from 2011. In 2017 Benoit Thieurmel built <a href="https://cran.r-project.org/package=ROI.plugin.clp" rel="nofollow" target="_blank">ROI.plugin.clp</a> on top of it, so that anyone using ROI could pick Clp with one argument. Then clpAPI’s CRAN checks started failing, nobody fixed them, and it was archived on 30 November 2021. ROI.plugin.clp had done nothing wrong, but it depended on an archived package, so it followed six weeks later.</p>
<p>I had been using clpAPI in my own models since 2018 and noticed the way everyone does: a fresh R install, a script that would not load. The particular annoyance is that on Windows the Clp library is <em>already on the machine</em> — Rtools has shipped it since version 4.3 — so the only thing missing was a few hundred lines of C++ to get at it.</p>
<h2>What coinclp does</h2>
<p>The bindings are written from scratch against the current Clp callable library, with registered entry points and external pointers that clean up after themselves, and they build on R 4.5 and 4.6. There are three ways in.</p>
<p>The first is one call. Give it an objective, a constraint matrix, directions and a right-hand side, and get back the solution, the objective, the shadow prices and the reduced costs together:</p>
<pre>library(coinclp)

A &lt;- rbind(material = c(120, 210),
           labour   = c(110,  30),
           capacity = c(  1,   1))

fit &lt;- clp_solve(c(143, 60), A, &quot;&lt;=&quot;, c(15000, 4000, 75), max = TRUE)

fit$objval    # 6315.625
fit$solution  # 21.875 53.125
fit$duals     # 0.0000 1.0375 28.8750</pre>
<p>The constraint matrix can be an ordinary dense matrix, a <code>Matrix</code> sparse matrix, a <code>slam</code> triplet matrix or plain <code>i</code>/<code>j</code>/<code>v</code> triplets. A sparse matrix goes to Clp as sparse: the package passes only the non-zero entries, and never expands the matrix into a full grid of mostly zeros first, which for a large model is the difference between fitting in memory and not.</p>
<p>The second is the whole callable library: build a model, keep it, change bounds or coefficients, hand back the basis and re-solve. That is what Clp is for, and it is where the one-call interfaces of most R solver packages let you down — a parametric study or a column-generation loop that rebuilds the model each iteration throws away exactly the information that makes simplex fast. In the vignette a tightened re-solve from a saved basis takes zero iterations.</p>
<p>The third is for old code. Every function clpAPI exported — <code>initProbCLP()</code>, <code>loadProblemCLP()</code>, <code>solveInitialCLP()</code> and the rest — is reproduced with the same names and arguments, so a script written against clpAPI needs only a new <code>library()</code> line. None of clpAPI’s code is reused; the layer is an independent implementation of its interface.</p>
<p>ROI.plugin.coinclp is the ROI side. The solver is called <code>&quot;coinclp&quot;</code> rather than <code>&quot;clp&quot;</code>, because ROI takes the name from the package, and unlike the 2017 plugin it returns duals, reduced costs and row activities alongside the primal solution:</p>
<pre>library(ROI)
library(ROI.plugin.coinclp)

res &lt;- ROI_solve(op, solver = &quot;coinclp&quot;)
solution(res, &quot;dual&quot;)</pre>
<h2>When not to use it</h2>
<p>Clp is not the fastest open-source LP solver any more, and I would rather say so than have you find out. The obvious comparison is <a href="https://highs.dev/" rel="nofollow" target="_blank">HiGHS</a>, from Julian Hall’s group at Edinburgh, which is now the default LP solver in SciPy and in MATLAB. On Hans Mittelmann’s <a href="https://plato.asu.edu/ftp/lpopt.html" rel="nofollow" target="_blank">LPopt benchmark</a> as of September 2026, HiGHS solves 54 of the 65 test problems within the time limit and Clp solves 40, and HiGHS is about twice as fast on the scaled geometric mean. The commercial codes are further ahead again: COPT solves all 65 and is roughly 27 times faster than Clp on the same measure.</p>
<p>The benchmark also shows why the honest answer is “it depends on the model”. On a handful of instances Clp is the quicker of the two — <code>Linf_520c</code> takes Clp 36 seconds and HiGHS 872, and Clp solves <code>datt256</code> where HiGHS times out — but on more of them the reverse holds, sometimes by a wide margin. If you have one large LP to solve once, try <a href="https://cran.r-project.org/package=highs" rel="nofollow" target="_blank">highs</a> or <a href="https://cran.r-project.org/package=ROI.plugin.highs" rel="nofollow" target="_blank">ROI.plugin.highs</a> first; with ROI the switch is one argument, so trying both costs nothing.</p>
<p>And Clp solves linear programs only. If your variables are integer you need a MIP solver: HiGHS again, or <a href="https://cran.r-project.org/package=Rglpk" rel="nofollow" target="_blank">Rglpk</a>, <a href="https://cran.r-project.org/package=lpSolve" rel="nofollow" target="_blank">lpSolve</a> or COIN-OR’s own <a href="https://cran.r-project.org/package=Rsymphony" rel="nofollow" target="_blank">Rsymphony</a>. The ROI plugin will refuse an integer problem rather than quietly relax it, and <code>clp_solve()</code> has no notion of integer variables at all.</p>
<p>Where Clp still earns its place is the kind of work it was built for: models that are solved many times with small changes, where a warm start from the previous basis matters more than raw speed on a cold solve; anything that already speaks Clp’s API, which is more code than you might think; and Windows machines, where it is the one LP solver you get for free with Rtools and no further installation.</p>
<h2>Installing it</h2>
<p>On Windows, <code>install.packages(&quot;coinclp&quot;)</code> is the whole job — from source with Rtools installed until CRAN’s Windows binaries appear, which usually takes a few days. Elsewhere Clp is a system library and has to be there first: <code>coinor-libclp-dev</code> on Debian and Ubuntu, <code>coin-or-Clp-devel</code> on Fedora, <code>brew install clp</code> on macOS, <code>coin-or-clp</code> from conda-forge. Until the ROI plugin clears the CRAN queue it installs from GitHub:</p>
<pre>install.packages(&quot;coinclp&quot;)
remotes::install_github(&quot;SamLovick/ROI.plugin.coinclp&quot;)</pre>
<p>Source, issues and the vignette are at <a href="https://github.com/SamLovick/coinclp" rel="nofollow" target="_blank">github.com/SamLovick/coinclp</a> and <a href="https://github.com/SamLovick/ROI.plugin.coinclp" rel="nofollow" target="_blank">github.com/SamLovick/ROI.plugin.coinclp</a>. Both packages are under the Eclipse Public License, matching COIN-OR. If you were a clpAPI user and something in the compatibility layer does not behave as it used to, an issue with the script that broke would be very welcome.</p>
<p>The post <a href="https://lovickconsulting.com/2026/09/16/coinclp-coin-or-clp-back-on-cran/" rel="nofollow" target="_blank">coinclp: the COIN-OR Clp linear programming solver is back on CRAN</a> appeared first on <a href="https://lovickconsulting.com/" rel="nofollow" target="_blank">Sam Lovick Consulting</a>.</p>

<div style="border: 1px solid; background: none repeat scroll 0 0 #EDEDED; margin: 1px; font-size: 13px;">
<div style="text-align: center;">To <strong>leave a comment</strong> for the author, please follow the link and comment on their blog: <strong><a href="https://lovickconsulting.com/2026/09/16/coinclp-coin-or-clp-back-on-cran/"> R Archives - Sam Lovick Consulting</a></strong>.</div>
<hr />
<a href="https://www.r-bloggers.com/" rel="nofollow">R-bloggers.com</a> offers <strong><a href="https://feedburner.google.com/fb/a/mailverify?uri=RBloggers" rel="nofollow">daily e-mail updates</a></strong> about <a title="The R Project for Statistical Computing" href="https://www.r-project.org/" rel="nofollow">R</a> news and tutorials about <a title="R tutorials" href="https://www.r-bloggers.com/how-to-learn-r-2/" rel="nofollow">learning R</a> and many other topics. <a title="Data science jobs" href="https://www.r-users.com/" rel="nofollow">Click here if you're looking to post or find an R/data-science job</a>.

<hr>Want to share your content on R-bloggers?<a href="https://www.r-bloggers.com/add-your-blog/" rel="nofollow"> click here</a> if you have a blog, or <a href="http://r-posts.com/" rel="nofollow"> here</a> if you don't.
</div><strong>Continue reading</strong>: <a href="https://www.r-bloggers.com/2026/09/coinclp-the-coin-or-clp-linear-programming-solver-is-back-on-cran/">coinclp: the COIN-OR Clp linear programming solver is back on CRAN</a>]]></content:encoded>
					
		
		<enclosure url="" length="0" type="" />

		<post-id xmlns="com-wordpress:feed-additions:1">403657</post-id>	</item>
	</channel>
</rss>
