<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	xmlns:georss="http://www.georss.org/georss" xmlns:geo="http://www.w3.org/2003/01/geo/wgs84_pos#" xmlns:media="http://search.yahoo.com/mrss/"
	>

<channel>
	<title>Online Journalism Blog</title>
	<atom:link href="https://onlinejournalismblog.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://onlinejournalismblog.com</link>
	<description>Comment, analysis and links covering online journalism and online news, citizen journalism, blogging, vlogging, photoblogging, podcasts, vodcasts, interactive storytelling, publishing, Computer Assisted Reporting, User Generated Content, searching and all things internet.</description>
	<lastBuildDate>Tue, 18 Aug 2026 15:05:05 +0000</lastBuildDate>
	<language>en</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>http://wordpress.com/</generator>
<site xmlns="com-wordpress:feed-additions:1">722736</site><cloud domain='onlinejournalismblog.com' port='80' path='/?rsscloud=notify' registerProcedure='' protocol='http-post' />
<image>
		<url>https://s2.wp.com/i/webclip.png</url>
		<title>Online Journalism Blog</title>
		<link>https://onlinejournalismblog.com</link>
	</image>
	<atom:link rel="search" type="application/opensearchdescription+xml" href="https://onlinejournalismblog.com/osd.xml" title="Online Journalism Blog" />
	<atom:link rel='hub' href='https://onlinejournalismblog.com/?pushpress=hub'/>
	<item>
		<title>How to judge (and minimise) the risk of using sensitive information with an AI chatbot</title>
		<link>https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/</link>
					<comments>https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Tue, 18 Aug 2026 15:05:04 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[ethics]]></category>
		<category><![CDATA[law]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[anonymisation]]></category>
		<category><![CDATA[artificial intelligence]]></category>
		<category><![CDATA[ChatGPT]]></category>
		<category><![CDATA[Claude]]></category>
		<category><![CDATA[Copilot]]></category>
		<category><![CDATA[data retention]]></category>
		<category><![CDATA[extraction attacks]]></category>
		<category><![CDATA[gemini]]></category>
		<category><![CDATA[Henk Van Ess]]></category>
		<category><![CDATA[knowledge cutoff]]></category>
		<category><![CDATA[legal orders]]></category>
		<category><![CDATA[pretraining]]></category>
		<category><![CDATA[privacy]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reverse search warrants]]></category>
		<category><![CDATA[RLHF]]></category>
		<category><![CDATA[security]]></category>
		<category><![CDATA[threat modelling]]></category>
		<category><![CDATA[threat models]]></category>
		<category><![CDATA[tokens]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31647</guid>

					<description><![CDATA[What are the risks of information you put into an AI chat becoming public? Evaluating those risks is slightly different to other information security challenges because of the way large language models work, so here&#8217;s a guide to assessing and managing information security with AI. Understanding risk with AI is about understanding probabilities One of [&#8230;]]]></description>
										<content:encoded><![CDATA[
<h4 class="wp-block-heading"><strong><em>What are the risks of information you put into an AI chat becoming public? Evaluating those risks is slightly different to other information security challenges because of the way large language models work, so here&#8217;s a guide to assessing and managing information security with AI.</em></strong></h4>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png"><img width="1024" height="569" data-attachment-id="31726" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/4-ai-security-considerations-1/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png" data-orig-size="1209,672" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="4 AI security considerations (1)" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png?w=1024" alt="
4 security questions for genAI

Uniqueness
How much does the LLM already know about this?

Sensitivity
How bad would disclosure be?

Exposure
Who or what could access it besides the model?

Alternatives
Are there safer ways to complete the task?

Paul Bradshaw 2026 | Icons generated by Google Gemini
" class="wp-image-31726" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/4-ai-security-considerations-1.png 1209w" sizes="(max-width: 1024px) 100vw, 1024px" /></a></figure>



<span id="more-31647"></span>



<h2 class="wp-block-heading">Understanding risk with AI is about understanding probabilities</h2>



<p class="wp-block-paragraph">One of the most common misunderstandings about large language models (LLMs) is the idea that if you upload a document such as an interview transcript, confidential leak or sensitive dataset to an AI tool like ChatGPT, it might reveal that document to another user.</p>



<p class="wp-block-paragraph">There are no publicly documented cases of this happening. </p>



<p class="wp-block-paragraph">The reason for this is the same reason why AI shouldn&#8217;t be treated as a search engine: large language models like GPT-5 don&#8217;t work by retrieving documents for users — they work by separating documents into parts of language (&#8216;tokens&#8217; in the jargon), adding those to a vast ocean of other tokens, and <em>training</em> on a database of <em>relationships</em> between those.</p>



<p class="wp-block-paragraph">When you ask a question like &#8220;Who is the CEO of Apple?&#8221; a genAI tool doesn&#8217;t retrieve information from a document. Instead it breaks down your question into tokens like &#8220;CEO&#8221; and &#8220;Apple&#8221; and &#8220;Who&#8221;, calculates the most statistically likely meaning of those things (e.g. is that apple probably a fruit or a company), and then the most statistically likely sequence of words which might form a meaningful answer to that question. </p>



<p class="wp-block-paragraph">(In this example, &#8220;Tim&#8221; and &#8220;Cook&#8221; are likely to be more strongly associated with &#8220;CEO&#8221; and &#8220;Apple&#8221; in its training data than other tokens, and so they will be formed into a response: &#8220;Tim Cook&#8221;.)</p>



<p class="wp-block-paragraph">If information you provide in a prompt is used in an LLM&#8217;s training (see below), then, it will be a) broken down into tokens and b) those tokens&#8217; relationships with other tokens will be competing with lots of other token relationships to be the &#8216;most likely&#8217; answer to any question. </p>



<p class="wp-block-paragraph">So when do those <strong>tokens and relationships</strong> become a risk?</p>



<h2 class="wp-block-heading">Probabilities are shaped by training</h2>



<p class="wp-block-paragraph">There are two main steps in a large language model&#8217;s training: the most important is <strong>pretraining</strong>: this is when the model is trained using an enormous collection of language. If your information ends up being included in that collection, it is a drop in the ocean of those trillions of words.</p>



<p class="wp-block-paragraph">There will also be a delay before it is included in responses, because each model is trained a number of months before it is release — this is called the <strong>knowledge cutoff</strong>. </p>



<p class="wp-block-paragraph">ChatGPT&#8217;s latest model, GPT-5.6, has a knowledge cutoff of five months ago, which means it doesn&#8217;t contain any information after mid-February. If Apple&#8217;s CEO has changed since then, that won&#8217;t be reflected in its core training data. You can <a href="https://aiknowledgecutoff.com/">see each model&#8217;s cutoff dates here</a>.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg"><img width="1024" height="557" data-attachment-id="31746" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/dla2/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg" data-orig-size="5739,3122" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="dla2" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=1024" alt="Group photo of DLA members" class="wp-image-31746" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=2048 2048w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=1440 1440w" sizes="(max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">The Data Labelers Association (DLA) represents people paid by AI companies as part of models&#8217; training process. <em>Image: <a href="https://datalabelers.org/">Data Labelers Association</a></em></figcaption></figure>



<p class="wp-block-paragraph">The other key step in a model&#8217;s training is called <strong>reinforcement learning</strong> from human feedback (<a href="https://huggingface.co/blog/rlhf">RLHF</a>).</p>



<p class="wp-block-paragraph">Whenever ChatGPT presents you with two possible responses and asks to choose which one you prefer, that&#8217;s you helping to train the model by reinforcing a preference. </p>



<p class="wp-block-paragraph">AI companies pay humans to train the model too (there are <a href="https://www.amnesty.org/en/documents/pol40/0996/2026/en/">human rights</a> and <a href="https://africauncensored.online/blog/2025/08/26/fuelling-the-agi-hype-the-recruitment-playbook-to-land-big-tech-contracts/">exploitation</a> concerns around this). Those people train the model by looking at a small sample of conversations and selecting the better responses.</p>



<p class="wp-block-paragraph">Your information <em>might</em> be selected for this sample (out of millions of conversations every day). If that happens it will be incorporated into a model&#8217;s training with a shorter time lag, and weighted more heavily. For example, even where the majority of pretraining data says the Apple CEO is Tim Cook, if a human has ranked a different, more up-to-date, response as more or less accurate, then it can outweigh that other training.</p>



<p class="wp-block-paragraph">The chances of your information either being incorporated into pretraining data, or selected for reinforcement training, will depend to some extent on policy.</p>



<h2 class="wp-block-heading">Check the policies</h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png"><img width="1024" height="184" data-attachment-id="31749" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/ai-privacy-policy-table/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png" data-orig-size="1600,288" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="ai-privacy-policy-table" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=1024" alt="Table showing privacy policies for " class="wp-image-31749" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=1440 1440w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png 1600w" sizes="(max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">The privacy policies of major AI companies in 2025. <em>Image: <a href="https://arxiv.org/pdf/2509.05382">User Privacy and Large Language Models: An Analysis of Frontier Developers’ Privacy Policies</a></em></figcaption></figure>



<p class="wp-block-paragraph">Free consumer accounts with ChatGPT, Gemini, Copilot and Claude <strong>opt users in by default</strong> to agreeing to allow their conversations to be used for training. The chances of a conversation being used is still tiny, but you can reduce it further by <a href="https://builtin.com/articles/ai-training-data-opt-out">opting out</a> (you might also consider opting out of AI training on <a href="https://www.wired.com/story/how-to-stop-your-data-from-being-used-to-train-ai/">platforms like Adobe and AWS</a> where you might use sensitive data). </p>



<p class="wp-block-paragraph">Enterprise and business accounts and APIs tend to opt users out by default. </p>



<p class="wp-block-paragraph"><strong>Data retention</strong> is another factor to consider. Some developers <a href="https://hai.stanford.edu/news/be-careful-what-you-tell-your-ai-chatbot">retain chats indefinitely</a>, and <a href="https://www.linkedin.com/pulse/illusion-incognito-what-ai-chatbots-really-keep-when-you-ciappelli-ozrzc/">even &#8216;incognito&#8217; or &#8216;temporary&#8217; chats are likely to be retained</a> for at least 30 days. Some accounts will have a <strong>zero data retention</strong> policy — PC Tech Magazine <a href="https://pctechmag.com/2026/07/zero-data-retention-how-to-enforce-it-across-ai-providers/">has a breakdown of different providers&#8217; policies and processes</a> for this.</p>



<p class="wp-block-paragraph">Your chats are more <em>likely</em> to be selected for reinforcement training — and retained for longer — if they are &#8216;flagged&#8217; for training for some reason. The most obvious reason that a chat might be flagged is for <a href="https://www.anthropic.com/legal/aup">safety reasons</a>, or because you have given feedback on a response in some way.</p>



<p class="wp-block-paragraph">Chats can be flagged for training <strong>even</strong> <strong>if you have opted out</strong> of training on a free account, as a recent update to Anthropic&#8217;s privacy policy <a href="https://techcoffeehouse.com/2026/06/09/claude-training-data-opt-out-carve-out/">revealed</a>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">&#8220;Conversations that Anthropic’s systems flag for safety review can still be used to train its models, regardless of a user’s stated preference. Anthropic’s policy does not define what triggers a safety flag, nor does it commit to notifying users when one occurs.&#8221;</p>
</blockquote>



<h2 class="wp-block-heading">The less information about an entity, the higher the risk</h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png"><img loading="lazy" width="1024" height="561" data-attachment-id="31759" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/ai-network-privacy/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png" data-orig-size="1693,929" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="ai-network-privacy" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=1024" alt="Side-by-side digital illustration comparing how new information is represented in two AI knowledge networks. On the left, a dense web of interconnected nodes surrounds a large central node labelled “Microsoft”. The only surrounding labels are HIRING, POLICY, EMPLOYMENT, WORKFORCE, EMPLOYEES, HR, PAY, CONDITIONS, HOLIDAY, RECRUITMENT and BONUSES. A document icon points towards the network, where its new connections appear as many thin, faint lines that blend into the existing structure. On the right, a sparse network contains only six nodes labelled PREDATOR NAME and VICTIM NAME 1 through VICTIM NAME 5. A document icon points towards this network, where bright, thick connections between the names form a prominent, isolated cluster. The illustration contrasts how new information can become diluted within a large, well-connected network but stand out clearly in a sparse one.
" class="wp-image-31759" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png?w=1440 1440w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-network-privacy.png 1693w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Image generated by GPT-5</figcaption></figure>



<p class="wp-block-paragraph">If your information was used to train a large language model — either through pretraining or reinforcement training — there is still no guarantee that it will find its way into a response to a user&#8217;s prompt. </p>



<p class="wp-block-paragraph">What will make this more likely, however, is if your information relates to an entity or a question that the LLM &#8216;knows&#8217; less about (i.e. has less training data).</p>



<p class="wp-block-paragraph">Here are two scenarios to illustrate:</p>



<p class="wp-block-paragraph">In Scenario A, a journalist uploads a sensitive document about Microsoft&#8217;s hiring policies to Gemini and it ends up being part of Gemini&#8217;s training data. All the &#8216;tokens&#8217; (words and parts of words) from that document are already in its training, and this document changes the strengths of their relationships slightly so some have a stronger relationship with tokens like &#8216;Microsoft&#8217; or concepts like &#8216;bias&#8217;.</p>



<p class="wp-block-paragraph">At various points Gemini users will ask a question about Microsoft&#8217;s hiring policies. Because the sensitive document is one of many documents about Microsoft and hiring that Gemini has been trained on, its information is very unlikely to make a meaningful difference to most responses. </p>



<p class="wp-block-paragraph">In Scenario B, on the other hand, a journalist uploads a list of the names of victims of a known predator. In this case, some of the names are quite unique and Gemini has either very little or no training data related to those. As a result the information in the document (the relationships between words) will be much more heavily weighted. Instead of being a drop in this particular ocean, the document constitutes a much larger part of it. </p>



<p class="wp-block-paragraph">Now when users ask a question about one of those people, Gemini has much less training to draw on — and is therefore more likely to replicate the particular patterns of words (the &#8216;fact pattern&#8217;) in the sensitive document than other patterns it might predict. </p>



<p class="wp-block-paragraph">Remember, however: it is still not <em>retrieving</em> a sentence from a document — it is only <em>predicting</em> one based on mathematical relationships.</p>



<p class="wp-block-paragraph">Notably, if they asked about the known predator in general the chances of those names being generated by the AI tool is lower, because there are more, and stronger, connections to other facts, which might be more heavily weighted (depending on the question). </p>



<p class="wp-block-paragraph">The complicating factor, of course, is that the user has no way of knowing whether the &#8216;fact&#8217; in either scenario is true. When asked questions where very little information exists, AI tools are much more likely to hallucinate (hence <a href="https://www.damiencharlotin.com/hallucinations/?sort_by=-date&amp;states=USA&amp;period_idx=0">court cases where AI tools have defamed individuals</a>). And because the sensitive material is not online in either scenario, the chatbot cannot point to it as a source for the user to check with. The user has the suggestion of a fact, but not the proof.</p>



<h2 class="wp-block-heading">When the risk calculation is different: sharing chats, internal access, hacks and law</h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png"><img loading="lazy" width="1024" height="784" data-attachment-id="31752" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/chatgpt-tweet-share2/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png" data-orig-size="1154,884" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="chatgpt-tweet-share2" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png?w=1024" alt="Tweet:         We just removed a feature from @ChatGPTapp that allowed users to make their conversations discoverable by search engines, such as Google. This was a short-lived experiment to help people discover useful conversations. This feature required users to opt-in, first by picking a chat to share, then by clicking a checkbox for it to be shared with search engines (see below).

Ultimately we think this feature introduced too many opportunities for folks to accidentally share things they didn't intend to, so we're removing the option. We're also working to remove indexed content from the relevant search engines. This change is rolling out to all users through tomorrow morning.

Security and privacy are paramount for us, and we'll keep working to maximally reflect that in our products and features." class="wp-image-31752" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/chatgpt-tweet-share2.png 1154w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">OpenAI <a href="https://threadreaderapp.com/thread/1951041845938499669.html">removed</a> the ability for shared chats to be found via Google after Henk Van Ess <a href="https://www.digitaldigging.org/p/chatgpt-confessions-gone-they-are">revealed</a> that they were public &#8211; but they are still on Archive.org.</figcaption></figure>



<p class="wp-block-paragraph">Outside of user interactions with the AI tool, there are a number of other risks to consider. </p>



<p class="wp-block-paragraph">The first is a simple one: if you choose to <strong>share</strong> the chat at any point this creates a risk that it becomes public. </p>



<p class="wp-block-paragraph">Last year an investigation by <strong>Henk Van Ess</strong>&#8216;s Digital Digging newsletter <a href="https://www.digitaldigging.org/p/chatgpt-confessions-gone-they-are">uncovered</a> over 100,000 ChatGPT conversations which were findable with advanced search techniques. Last month he <a href="https://www.digitaldigging.org/p/96477-private-chats-went-public-in">found another 90,000 shared conversations</a> from multiple chatbots. The problem isn&#8217;t the AI tool — it&#8217;s the user sharing their chat. As Henk writes: &#8220;It looks like people thought &#8220;share&#8221; means <em>send</em>. But it means <em>publish</em>.&#8221;</p>



<p class="wp-block-paragraph">It&#8217;s also worth mentioning that some <strong>engineering, support, and safety/abuse teams</strong> at AI companies may have authorisation to access AI conversations. As with reinforcement training, this is more likely if content is flagged for specific reasons, to provide account support, or for legal reasons, and in the context of millions of prompts each day. As with any online service, there is also the potential for an employee to misuse their access. </p>



<p class="wp-block-paragraph">The same security risks that apply to online behaviour in general apply here, too: if you are using unsecured or public wifi to access an AI tool, it could be intercepted. The provider&#8217;s servers could be hacked or breached, as could your account with the provider if, for example, your password is weak. Your employer&#8217;s network and your own computer may also be hacked or accessed.</p>



<p class="wp-block-paragraph">A more specific security risk is <strong>extraction attacks</strong>: these are specific attempts to extract training data from large language models (<a href="https://github.com/kzhao5/ModelExtractionPapers">documented in literature listed here</a>).</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png"><img loading="lazy" width="1024" height="707" data-attachment-id="31708" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/openai_govtrequests2025/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png" data-orig-size="1286,888" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="openAI_govtRequests2025" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=1024" alt="Pie chart and table showing number of government requests received by OpenAI: 75 content requests received, 62 where data was disclosed." class="wp-image-31708" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png 1286w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">OpenAI received 75 Government requests for content between July and December 2025. <a href="https://cdn.openai.com/trust-and-transparency/report-2025h2-government-requests-for-user-data.pdf">Source: OpenAI transparency report</a></figcaption></figure>



<p class="wp-block-paragraph">Finally, there are <strong>legal orders</strong>, which may be more important to consider than technical risks. The most famous of these is <a href="https://arstechnica.com/tech-policy/2025/11/openai-fights-order-to-hand-over-20-million-private-chatgpt-conversations/">the court ruling that forced OpenAI to give millions of user chats to news organisations</a> in a copyright case, but chats are increasingly showing up as evidence in criminal and civil cases, with one legal website <a href="https://www.cslawreport.com/21343316/genai-chats-becoming-evidence-law-enforcement-warrants-and-subpoenas.thtml">advising</a> &#8220;to prepare for a steady increase of government and litigation requests for Gen AI data&#8221;. In the US, court rulings have <a href="https://www.whitecase.com/insight-alert/attorney-client-privilege-and-work-product-age-generative-ai">indicated</a> that chats with AI tools &#8220;were not protected by attorney-client privilege or the work product doctrine&#8221; (see <a href="https://caselaw.nationalarchives.gov.uk/ukut/iac/2026/81?"><em>Munir v Secretary of State for the Home Department</em></a> for a <a href="https://www.osborneclarke.com/insights/ai-tools-and-privilege-uk-what-are-risks">similar</a> point in the UK).</p>



<p class="wp-block-paragraph">Some of these requests will be specific to a particular user, but a warrant in October 2025 <a href="https://www.forbes.com/sites/thomasbrewster/2025/10/20/openai-ordered-to-unmask-writer-of-prompts/">revealed</a> that they can also relate to a particular prompt: what is <a href="https://www.eff.org/deeplinks/2025/12/ai-chatbot-companies-should-protect-your-conversations-bulk-surveillance">known</a> as <strong>&#8220;reverse&#8221; search warrants</strong>. </p>



<h2 class="wp-block-heading">What can you do about it? Anonymisation and switching to local models</h2>



<p class="wp-block-paragraph">One way to reduce the risk involved in uploading sensitive information to an AI chatbot is to <strong>anonymise</strong> it. This involves either removing information that identifies individuals, or replacing it with unique IDs that you can use to re-identify them outside of the chat. </p>



<p class="wp-block-paragraph">It&#8217;s not just names that might identify individuals: any collection of details might be connected with an individual (<a href="https://www.taylorhampton.co.uk/jigsaw-identification-highlights-the-dangers-that-publications-can-pose-in-libel-cases/">jigsaw identification</a>) — remember that it&#8217;s these <em>relationships</em> that large language models are designed to calculate. Some <a href="https://eleks.com/research/data-anonymization-working-solution/">technical solutions and libraries</a> already exist for this.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png"><img loading="lazy" width="1000" height="644" data-attachment-id="31755" data-permalink="https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/data-privacy-venn-diagram/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png" data-orig-size="1000,644" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="data-privacy-venn-diagram" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png?w=1000" alt="Venn diagram showing two circles: 
sensitive data - not already public, may cause harm if disclosed
personal data - defined and governed by data protection and regulation. 
Where they overlap: 'private data'" class="wp-image-31755" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png 1000w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/data-privacy-venn-diagram.png?w=768 768w" sizes="auto, (max-width: 1000px) 100vw, 1000px" /></a><figcaption class="wp-element-caption">The ODI&#8217;s <a href="https://theodi.org/news-and-events/blog/anonymisation-and-synthetic-data-towards-trustworthy-data/">guide on anonymisation and synthetic data</a> covers both personal and sensitive data</figcaption></figure>



<p class="wp-block-paragraph">It&#8217;s also worth considering whether uploading the information is necessary at all: just providing a description of the document or data may be enough to get a useful response, especially if you are focusing on a process (<a href="https://onlinejournalismblog.com/2025/12/02/journey-prompts-and-destination-prompts-how-to-avoid-becoming-deskilled-when-using-ai/">journey prompting</a>) rather than an end result.</p>



<p class="wp-block-paragraph">If the risk of using an AI service is too high, another option is to use a large language model locally — i.e. on your own computer rather than a server belonging to OpenAI, Google or another provider.</p>



<p class="wp-block-paragraph">A number of tools make it relatively easy to run AI models locally, including <a href="https://lmstudio.ai/">LM Studio</a>, <a href="https://anythingllm.com/">AnythingLLM</a> and <a href="https://ollama.com/">Ollama</a>. These allow you to choose from a range of models, download them to your computer, (make sure you use the local models as some also offer &#8216;cloud&#8217; models), and then run prompts against them. </p>



<p class="wp-block-paragraph">You will need a computer with enough processing power and RAM (<a href="https://www.canirun.ai/">CanIRun.ai</a> provides a breakdown of which models you may be able to run based on detecting your computer&#8217;s specs), and the larger the model, the more power you will need. Prompts will typically run more slowly than on mainstream AI chatbots, too. </p>



<h2 class="wp-block-heading">Make a threat model</h2>



<p class="wp-block-paragraph">A useful process to map all this out is to create a <strong>threat model</strong> that <a href="https://onlinejournalismblog.com/2014/07/16/why-every-journalist-should-have-a-threat-model-with-cats/">asks four simple questions</a>:</p>



<ol class="wp-block-list">
<li><strong>What information do you not want other people to know?</strong> (This can be anything from passwords to contacts’ details, data and documents)</li>



<li><strong>Why </strong>might someone want that information? <strong>Who?</strong></li>



<li><strong>What can they do</strong> to get it?</li>



<li><strong>What might happen</strong> if they do?</li>
</ol>



<p class="wp-block-paragraph">Using the information in this post to map out <em>realistic</em> threats will make any subsequent steps much more manageable — and less anxiety-provoking.</p>



<p class="wp-block-paragraph"><strong><em>This is a work in progress. Contributions, suggestions and updates are <a href="https://www.linkedin.com/in/paulbradshawuk/">welcome</a>.</em></strong></p>



<p class="wp-block-paragraph"><em>Thanks to Laura Isotalo who raised a question around security and suggested writing about it after my talk on <a href="https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/">FOI management with AI</a> at Dataharvest this year.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/08/18/how-to-judge-and-minimise-the-risk-of-using-sensitive-information-with-an-ai-chatbot/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31647</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/08/dla2.jpg?w=1024" medium="image">
			<media:title type="html">Group photo of DLA members</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/08/ai-privacy-policy-table.png?w=1024" medium="image">
			<media:title type="html">Table showing privacy policies for </media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/07/openai_govtrequests2025.png?w=1024" medium="image">
			<media:title type="html">Pie chart and table showing number of government requests received by OpenAI: 75 content requests received, 62 where data was disclosed.</media:title>
		</media:content>
	</item>
		<item>
		<title>How data journalism is done on TikTok: six key takeaways from watching 50 videos</title>
		<link>https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/</link>
					<comments>https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/#respond</comments>
		
		<dc:creator><![CDATA[James Hickman]]></dc:creator>
		<pubDate>Mon, 10 Aug 2026 12:06:00 +0000</pubDate>
				<category><![CDATA[data journalism]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[online video]]></category>
		<category><![CDATA[bar charts]]></category>
		<category><![CDATA[BBC]]></category>
		<category><![CDATA[Choropleth maps]]></category>
		<category><![CDATA[FT]]></category>
		<category><![CDATA[Guardian]]></category>
		<category><![CDATA[ITV News]]></category>
		<category><![CDATA[James Hickman]]></category>
		<category><![CDATA[John Burn-Murdock]]></category>
		<category><![CDATA[length]]></category>
		<category><![CDATA[Line charts]]></category>
		<category><![CDATA[location]]></category>
		<category><![CDATA[TikTok]]></category>
		<category><![CDATA[Times]]></category>
		<category><![CDATA[vertical video]]></category>
		<category><![CDATA[visualisation]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31768</guid>

					<description><![CDATA[In a guest post for OJB, James Hickman pulls together a set of basic best-practice guidelines for creating data journalism videos on TikTok, from length and location to animation and style. Short-form vertical video is quickly becoming a key format for telling all types of news stories — and data journalism is no exception. As a data journalist, I wanted [&#8230;]]]></description>
										<content:encoded><![CDATA[
<h3 class="wp-block-heading"><em>In a guest post for OJB, James Hickman pulls together a set of basic best-practice guidelines for creating data journalism videos on TikTok, from length and location to animation and style.</em></h3>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png"><img loading="lazy" width="1024" height="384" data-attachment-id="31811" data-permalink="https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/video-best-practices/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png" data-orig-size="2170,814" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="video best practices" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=1024" alt="6 best practice guidelines for video data shorts:
1. Keep your data shorts under 90 seconds
2. Maintain a human connection
3. Locations reinforce authority
4. Show don't tell
5. Animate your graphs
6. Don't be afraid to innovate and adapt" class="wp-image-31811" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=2048 2048w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/video-best-practices.png?w=1440 1440w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">Short-form vertical video is quickly <a href="https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2026" target="_blank" rel="noopener">becoming a key format for telling</a> all types of news stories — and data journalism is no exception. </p>



<p class="wp-block-paragraph">As a data journalist, I wanted to know how these outlets use this new and evolving digital video format to create “<strong>data shorts</strong>”: short-form social media videos in a vertical, reel-style format, that analyses or reports on specific datasets or newly released data.</p>



<p class="wp-block-paragraph"><a href="https://docs.google.com/spreadsheets/d/1Am_0HAx_kY59FTKs7ZMfzFbfB82QC7sS0eqmUpsCpxY/edit?usp=sharing" target="_blank" rel="noopener">I analysed 50 data shorts</a> from the<em> </em><a href="http://tiktok.com/" target="_blank" rel="noopener"><em>TikTok </em></a>accounts of five major UK news publications —<em> </em>the <a href="https://www.tiktok.com/@bbcnews" target="_blank" rel="noopener"><em>BBC,</em></a><em> </em><a href="https://www.tiktok.com/@itvnews" target="_blank" rel="noopener"><em>ITV News</em></a><em>,</em> <a href="https://www.tiktok.com/@guardian" target="_blank" rel="noopener"><em>Guardian,</em></a> <a href="https://www.tiktok.com/@thetimes" target="_blank" rel="noopener"><em>Times</em></a><em> and</em> <a href="https://www.tiktok.com/@financialtimes" target="_blank" rel="noopener"><em>Financial Times</em></a><em> — </em>to examine how these digital news outlets approach the format. I looked at the genre of story they covered, the <a href="https://onlinejournalismblog.com/2020/08/11/here-are-the-7-types-of-stories-most-often-found-in-data/" target="_blank" rel="noopener">angles they focused on</a>, their length, and their use of animation, to help inform my guidelines.</p>



<span id="more-31768"></span>



<h2 class="wp-block-heading"><strong>Tip 1: Keep your data shorts under 90 seconds</strong></h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png"><img loading="lazy" width="1024" height="828" data-attachment-id="31813" data-permalink="https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/mean-tiktok-video-lengths-vary/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png" data-orig-size="1348,1090" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="Mean TikTok video lengths vary" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=1024" alt="Bar chart: Mean video lengths vary across publications, but none produce data shorts averaging over two minutes. The BBC has the shortest bar at 58 seconds, the FT the longest at 105 seconds." class="wp-image-31813" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png 1348w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">My analysed data shorts averaged a runtime of <strong>just under 84 seconds</strong>.</p>



<p class="wp-block-paragraph">The <strong>BBC</strong> <strong>ranked lowest</strong> for average video length, at just under a minute, while <strong>the Financial Times had the highest</strong> at 105 seconds.</p>



<p class="wp-block-paragraph">Notably, <strong>none of the publications had an average video length over two minutes</strong>, despite<em> </em>TikTok<em> </em>allowing longer uploads up to 60 minutes. </p>



<p class="wp-block-paragraph">By keeping their runtimes short and using an average speaking rate of <a href="https://virtualspeech.com/blog/average-speaking-rate-words-per-minute" target="_blank" rel="noopener">100–150 words per minute</a>, <strong>data shorts act as video equivalents to the traditional News in Brief</strong> (NIB) format — <a href="https://roughhousemedia.co.uk/our-simple-guide-to-journalistic-jargon/" target="_blank" rel="noopener">short news articles usually summarised in a single paragraph</a> — rather than conventional 300–500-word news stories.</p>



<p class="wp-block-paragraph">Based on the average runtime, <strong>a script for a data short should be around 200 words</strong>, presenting  a condensed overview rather than deep-dives into causes and effects which would require a more time-demanding feature-length format.</p>



<p class="wp-block-paragraph">By condensing the runtime of a data short, journalists can also <strong>get to the story’s point more quickly</strong>, maintaining audience engagement with the video by removing unnecessary details, pauses or words that may increase the chances of viewers scrolling to the next video in their timeline.</p>



<h2 class="wp-block-heading">Tip 2: <strong>Maintain a human connection and get in front of the camera</strong></h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png"><img loading="lazy" width="1024" height="827" data-attachment-id="31814" data-permalink="https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/shorts-favour-the-visual-presence-of-journalists/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png" data-orig-size="1344,1086" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="shorts favour the visual presence of journalists" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png?w=1024" alt="Donut chart: My analysed shorts favour the visual presence of their journalists/narrators, rather than just the inclusion of their voice in data shorts
86% of the donut is piece to camera, the other 14% is audio only." class="wp-image-31814" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/shorts-favour-the-visual-presence-of-journalists.png 1344w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">Journalists had a presence in all 50 shorts.</p>



<p class="wp-block-paragraph">43 out of the 50 shorts were shot in the <strong>piece-to-camera</strong> format, physically establishing a human relationship between a journalist and the data story — and all the videos featured a journalist&#8217;s voice-over.</p>



<p class="wp-block-paragraph">Featuring the voice and face of a presenter in a data short is important, as vertical video platforms like TikTok are personality-driven and built on relationships with audiences.</p>



<p class="wp-block-paragraph">By being present, journalists and presenters act as a point of engagement, recognition, and credibility with their audience, building trust faster and establishing themselves as a reliable news source and industry presence.</p>



<h2 class="wp-block-heading">Tip 3: <strong>Locations reinforce your authority</strong></h2>



<p class="wp-block-paragraph">Where you film your piece-to-camera is also important.</p>



<p class="wp-block-paragraph">All the data shorts with pieces-to-camera I analysed were filmed in a recording studio or a newsroom setting, except one, which was recorded on location.</p>



<p class="wp-block-paragraph">By recording in professional environments, you further reinforce your reliability as a news source, creating an innate sense of trust with your audience by visually communicating your professionalism.</p>



<h2 class="wp-block-heading">Tip 4: <strong>Show don’t tell</strong></h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png"><img loading="lazy" width="1024" height="825" data-attachment-id="31815" data-permalink="https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/line-charts-were-the-most-commonly-used-visualisation/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png" data-orig-size="1344,1084" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="Line charts were the most commonly used visualisation" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png?w=1024" alt="Bar chart: Line charts were the most commonly used visualisation in the data shorts I analysed
The longest bar is for line chart (15 uses), followed by bar chart (11)" class="wp-image-31815" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/line-charts-were-the-most-commonly-used-visualisation.png 1344w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">Data visualisations are commonly used in data journalism — and data shorts are no exception, with 33 of the 50 TikTok videos analysed including some form of data visualisation.</p>



<p class="wp-block-paragraph"><strong>Line charts</strong> (used 15 times) were the most common form of visualisation, suggesting that change stories such as <a href="https://www.tiktok.com/@financialtimes/video/7548754660601908502">this video by the FT&#8217;s John Burn-Murdoch</a> were one of the most common angles chosen.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@financialtimes/video/7548754660601908502" data-video-id="7548754660601908502" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@financialtimes" href="https://www.tiktok.com/@financialtimes?refer=embed">@financialtimes</a> <p>One section of society that continues to steer clear of the conversation over concern about declining birth rates is the left. FT chief data reporter John Burn-Murdoch explains why progressives should care about falling birth rates. Tap the link to find out more. <a title="ft" target="_blank" href="https://www.tiktok.com/tag/ft?refer=embed">#FT</a> <a title="financialtimes" target="_blank" href="https://www.tiktok.com/tag/financialtimes?refer=embed">#FinancialTimes</a></p> <a target="_blank" title="♬ original sound - FinancialTimes" href="https://www.tiktok.com/music/original-sound-7548754667640195842?refer=embed">♬ original sound &#8211; FinancialTimes</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph"><strong>Bar charts</strong> —<a href="https://www.tiktok.com/@itvpolitics/video/7652400134977752342" target="_blank" rel="noopener"> used primarily to show ranking comparisons</a> — were the second-most popular form of visualisation.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@itvpolitics/video/7652400134977752342" data-video-id="7652400134977752342" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@itvpolitics" href="https://www.tiktok.com/@itvpolitics?refer=embed">@itvpolitics</a> <p>Reporter Lewis Denison takes a look at the impact of Restore Britain in the Makerfield by-election, which is being held on June 18 <a title="politics" target="_blank" href="https://www.tiktok.com/tag/politics?refer=embed">#politics</a> @itvnews</p> <a target="_blank" title="♬ original sound - ITV Politics - ITV Politics" href="https://www.tiktok.com/music/original-sound-ITV-Politics-7652400172386880278?refer=embed">♬ original sound &#8211; ITV Politics &#8211; ITV Politics</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph"><strong>Choropleth maps</strong> (appearing four times) were the third most prevalent visualisation, used in political and economic stories to visualise statistical differences between geographical areas.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@bbcnews/video/7567771000696294678" data-video-id="7567771000696294678" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@bbcnews" href="https://www.tiktok.com/@bbcnews?refer=embed">@bbcnews</a> <p>A minister said the figures were an indictment of previous policies and the government was &#8220;tackling the root causes of deprivation head on&#8221;. <a title="england" target="_blank" href="https://www.tiktok.com/tag/england?refer=embed">#England</a> <a title="data" target="_blank" href="https://www.tiktok.com/tag/data?refer=embed">#Data</a> <a title="neighbourhood" target="_blank" href="https://www.tiktok.com/tag/neighbourhood?refer=embed">#Neighbourhood</a> <a title="ukgovernment" target="_blank" href="https://www.tiktok.com/tag/ukgovernment?refer=embed">#UKGovernment</a> <a title="bbcnews" target="_blank" href="https://www.tiktok.com/tag/bbcnews?refer=embed">#BBCNews</a></p> <a target="_blank" title="♬ original sound - BBC News - BBC News" href="https://www.tiktok.com/music/original-sound-BBC-News-7567770992240511766?refer=embed">♬ original sound &#8211; BBC News &#8211; BBC News</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph">More complex visualisations, like dot plots and stacked area charts, appeared only once each throughout my analysis, suggesting that <strong>intricate graphs may be too complicated to use effectively </strong>in short-form vertical-video formats.</p>



<p class="wp-block-paragraph">Only eight of the 33 data shorts that included visualisations used more than one type of chart, suggesting that the short time frame may further restrict the amount of analysis achievable through visualisations.</p>



<h2 class="wp-block-heading">Tip 5: <strong>Animate your graphs to maintain momentum</strong></h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png"><img loading="lazy" width="1024" height="824" data-attachment-id="31817" data-permalink="https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/most-tiktok-data-visualisations-featured-some-form-of-animation/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png" data-orig-size="1362,1096" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="Most TikTok data visualisations featured some form of animation" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png?w=1024" alt="Donut chart: Most of the data visualisations seen within my analysis featured some form of animation
Number of animated graphs: 29
Static graphs: 4" class="wp-image-31817" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/08/most-tiktok-data-visualisations-featured-some-form-of-animation.png 1362w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">Digitally animated graphics are a visual storytelling technique frequently used in short-form vertical video to engage viewers.</p>



<p class="wp-block-paragraph"><strong>29 of the 33 data visualisations in my analysis featured animation</strong>, either zooming in on specific data, revealing data over time or highlighting specific elements through using colour.</p>



<p class="wp-block-paragraph">This proved effective in adding further storytelling to the graphs, while allowing the video to elaborate on specific datapoints within the visualisation.</p>



<p class="wp-block-paragraph">Animated text was equally as common, primarily used to emphasise statistical information or to draw attention to particular words or names that were important to the story. An example can be seen in<a href="https://www.tiktok.com/@thetimes/video/7657589204440009987" target="_blank" rel="noopener"> <em>the Times</em>’ 2026 data short on the increase in UK expatriates</a>, which uses text animations to highlight the rising figures, as well as the factors driving the increase in people leaving the UK.</p>



<h2 class="wp-block-heading">Tip 6: <strong>Don’t be afraid to innovate and adapt</strong> </h2>



<p class="wp-block-paragraph">Although the analysis documented the most common practices used by five major news brands, there are no set-in-stone methods for creating a data short — and my analysis also highlighted some newly emerging storytelling techniques.</p>



<p class="wp-block-paragraph">Two of the videos, for example, used <strong>satellite imagery as data visualisations</strong>. <a href="https://nightingaledvs.com/from-space-to-story-in-data-journalism/" target="_blank" rel="noopener">Robert Simmon</a> and <a href="https://nielsdehoog.medium.com/the-extent-of-flooding-in-western-europe-67d09668de67">Niels de Hoog</a> have used this approach in the past to show how geographical areas have changed over time, and on TikTok both <a href="http://medium.com/r?url=https%3A%2F%2Fwww.tiktok.com%2F%40guardian%2Fvideo%2F7619367712904465686">The Guardian</a> and <a href="https://www.tiktok.com/@bbcnews/video/7656510851813248278" target="_blank" rel="noopener">the BBC</a> experimented with satellite data in a similar way, creating effective, visually-led data shorts that show real geographical change rather than simply describing it. Both videos keep the runtime of these data shorts under the two-minute mark, and also allow time to give context on how and why these changes are occurring.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@guardian/video/7619367712904465686" data-video-id="7619367712904465686" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@guardian" href="https://www.tiktok.com/@guardian?refer=embed">@guardian</a> <p>“Three years into this war, the biggest threat to people isn’t just the fighting, it’s what these pictures seem to reveal: the starvation being created by the RSF on purpose.” As it prepared its 18-month siege of El Fasher in Sudan, the Rapid Support Forces, a paramilitary group, started by attacking the farming villages around the city. It burned down 41 villages over three months, attacking them several times, destroying everything they needed to produce food.  Research by the Yale Humanitarian Research Lab using satellite imagery and remote sensing suggests a “starvation strategy” was enforced on these rural communities. Legal experts say the RSF used mass starvation as a weapon against civilians, which amounts to a war crime.  We’ve worked with Yale’s Humanitarian Research Lab to reveal the damage and the effects it had on people who were displaced from these villages. Watch this video with Guardian reporter Kaamil Ahmed to learn more – and for our in-depth visual investigation tap the link in bio.</p> <a target="_blank" title="♬ original sound - The Guardian - The Guardian" href="https://www.tiktok.com/music/original-sound-The-Guardian-7619367770538609431?refer=embed">♬ original sound &#8211; The Guardian &#8211; The Guardian</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph">The Guardian <a href="http://medium.com/r?url=https%3A%2F%2Fwww.tiktok.com%2F%40guardian%2Fvideo%2F7388201464289643809">frequently implemented more elaborate animations and graphics within their work</a>, giving many of their data shorts greater visual spectacle and flair than those produced by other publications.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@guardian/video/7388201464289643809" data-video-id="7388201464289643809" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@guardian" href="https://www.tiktok.com/@guardian?refer=embed">@guardian</a> <p>What exactly is a super-majority? And why has it been brought up in British politics so much lately? While the term doesn’t actually apply in the UK really &#8211; it&#8217;s an American expression referring to a majority in the US Congress big enough to withstand a filibuster (a device used by a smaller group of senators to block legislation) – it was has been used repeatedly by the Conservatives and Reform UK throughout the six weeks of the general election campaigning. Now, the results are in, and despite the lowest voter turnout since 2001, Labour has claimed its largest majority government in 25 years. Keir Starmer’s party has secured 412 seats in parliament – well above the 326 required for a majority … but is that as good as it sounds? And what don’t the numbers tell us? Our political correspondent, Kiran Stacey explains more. <a title="ukelection" target="_blank" href="https://www.tiktok.com/tag/ukelection?refer=embed">#ukelection</a> <a title="labour" target="_blank" href="https://www.tiktok.com/tag/labour?refer=embed">#labour</a> <a title="labourparty" target="_blank" href="https://www.tiktok.com/tag/labourparty?refer=embed">#labourparty</a> <a title="keirstarmer" target="_blank" href="https://www.tiktok.com/tag/keirstarmer?refer=embed">#keirstarmer</a> <a title="supermajority" target="_blank" href="https://www.tiktok.com/tag/supermajority?refer=embed">#supermajority</a> <a title="ukpolitics" target="_blank" href="https://www.tiktok.com/tag/ukpolitics?refer=embed">#ukpolitics</a> <a title="politics" target="_blank" href="https://www.tiktok.com/tag/politics?refer=embed">#politics</a> <a title="uknews" target="_blank" href="https://www.tiktok.com/tag/uknews?refer=embed">#uknews</a></p> <a target="_blank" title="♬ original sound - The Guardian" href="https://www.tiktok.com/music/original-sound-7388201609739668257?refer=embed">♬ original sound &#8211; The Guardian</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph">Don’t be afraid to experiment and innovate with your data shorts to push the medium further.</p>



<h4 class="wp-block-heading"><em>If you have produced or work in the field of data shorts or vertical video journalism, I would love to get in contact with you in the comments of this post, through </em><a href="http://jameshickman2004@gmail.com/" target="_blank" rel="noopener"><em>email </em></a><em>or on </em><a href="http://linkedin.com/in/james-henry-hickman-058218242/?skipRedirect=true" target="_blank" rel="noopener"><em>LinkedIn</em></a><em>.</em></h4>



<p class="wp-block-paragraph"><em>A version of this post with further analysis and more detail about methodology was <a href="https://medium.com/@jameshickman2004/telling-data-stories-with-vertical-video-a-content-analysis-ddf0ec35eaeb">first published on Medium</a>. </em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/08/10/how-data-journalism-is-done-on-tiktok-six-key-takeaways-from-watching-50-videos/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31768</post-id>
		<media:content url="https://1.gravatar.com/avatar/4033760335cbd08e77e1a5c841ac775082a5d047b3e8aed6238514a6182e561f?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">jameshickman2004</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/08/mean-tiktok-video-lengths-vary.png?w=1024" medium="image">
			<media:title type="html">Bar chart: Mean video lengths vary across publications, but none produce data shorts averaging over two minutes. The BBC has the shortest bar at 58 seconds, the FT the longest at 105 seconds.</media:title>
		</media:content>
	</item>
		<item>
		<title>This one story can be used to discuss seven different types of bias</title>
		<link>https://onlinejournalismblog.com/2026/07/03/this-one-story-can-be-used-to-discuss-seven-different-types-of-bias/</link>
					<comments>https://onlinejournalismblog.com/2026/07/03/this-one-story-can-be-used-to-discuss-seven-different-types-of-bias/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Fri, 03 Jul 2026 09:51:26 +0000</pubDate>
				<category><![CDATA[online journalism]]></category>
		<category><![CDATA[SEO]]></category>
		<category><![CDATA[A/B testing]]></category>
		<category><![CDATA[Amal Clooney]]></category>
		<category><![CDATA[analytics]]></category>
		<category><![CDATA[availability heuristic]]></category>
		<category><![CDATA[bias]]></category>
		<category><![CDATA[Chicago Tribune]]></category>
		<category><![CDATA[cognitive bias]]></category>
		<category><![CDATA[confirmation bias]]></category>
		<category><![CDATA[diversity]]></category>
		<category><![CDATA[ITV News]]></category>
		<category><![CDATA[John Torode]]></category>
		<category><![CDATA[Lisa Faulkner]]></category>
		<category><![CDATA[Meghan Markle]]></category>
		<category><![CDATA[Nicole Kidman]]></category>
		<category><![CDATA[Taylor Swift]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31650</guid>

					<description><![CDATA[The latest &#8220;wife of&#8221; headline — ITV News&#8217;s report on the actor Lisa Faulkner revealing that she has undergone surgery after a cancer diagnosis — is an opportunity to get journalism students exploring how different forms of bias might shape news reporting — and not just the obvious ones. An opening question might be why a [&#8230;]]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png"><img loading="lazy" width="1024" height="733" data-attachment-id="31651" data-permalink="https://onlinejournalismblog.com/2026/07/03/this-one-story-can-be-used-to-discuss-seven-different-types-of-bias/torodeswife/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png" data-orig-size="1204,862" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="torodeswife" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=1024" alt="ITV News headline: John Torode’s wife Lisa Faulkner reveals breast cancer diagnosis" class="wp-image-31651" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png 1204w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">The latest &#8220;wife of&#8221; headline — ITV News&#8217;s <a href="https://www.itv.com/news/2026-07-02/john-torodes-wife-lisa-faulkner-reveals-breast-cancer-diagnosis">report</a> on the actor Lisa Faulkner revealing that she has undergone surgery after a cancer diagnosis — is an opportunity to get journalism students exploring how different forms of bias might shape news reporting — and not just the obvious ones.</p>



<span id="more-31650"></span>



<p class="wp-block-paragraph">An opening question might be why a reporter or editor would put another person&#8217;s name at the front of the headline &#8220;John Torode’s wife Lisa Faulkner reveals breast cancer diagnosis&#8221;. </p>



<p class="wp-block-paragraph">A male-dominated newsroom might be one answer. Another might be a newsroom that is more likely to watch a middle-class cooking programme (<em>Masterchef</em>) than a working-class soap opera (<em>Eastenders</em>). It might also be driven by <strong>search analytics</strong>: four times as many searches were <a href="https://trends.google.com/explore?q=%2Fm%2F0bv9zc%2C%2Fm%2F068fkb&amp;date=today%201-y&amp;geo=GB">made for Torode in the last 12 months</a> than Faulkner.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png"><img loading="lazy" width="752" height="643" data-attachment-id="31658" data-permalink="https://onlinejournalismblog.com/2026/07/03/this-one-story-can-be-used-to-discuss-seven-different-types-of-bias/trendstorodevfaulkner/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png" data-orig-size="752,643" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="TrendsTorodeVFaulkner" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png?w=752" alt="Line chart comparing search volumes for John Torode (peaks 12 months ago and generally higher) and Lisa Faulkner (peaks in the last week): In the last 12 months, four times as many searches were made for John Torode than Lisa Faulkner" class="wp-image-31658" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png 752w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png?w=300 300w" sizes="auto, (max-width: 752px) 100vw, 752px" /></a></figure>



<p class="wp-block-paragraph">Search analytics are closely related to bias around <strong>news values</strong>: why, for example, might reporting <a href="https://www.hellomagazine.com/celebrities/813918/taylor-swifts-boyfriend-travis-kelce-opens-up-over-loss-im-sorry-how-it-ended/">lead on &#8220;Taylor Swift&#8217;s boyfriend&#8221;</a> or &#8220;<a href="https://www.yahoo.com/entertainment/celebrity/articles/nicole-kidmans-husbands-flirty-move-213015535.html?guccounter=1&amp;guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&amp;guce_referrer_sig=AQAAADfQ8xHwbVxPQH0qLj3SPwitw0XOqEELezWSbADl5-xLfXWb-fOJiwpv1ueYrdLhCS--P79iHWjLRUsbaPrNlWeM_gCYdop34m5LJtfssQ_S8ELk5ykOh1I8XHQY3CBJzsrC24VDlF26rs2ehqlKm3hLsoJmxGMMjb9wrMjMhrAF">Nicole Kidman’s husband</a>&#8221; instead? How do news values come into play with <a href="https://fortune.com/2015/09/01/just-how-sexist-was-the-ap-tweet-calling-amal-clooney-wife-of-actor/">Amal Clooney</a> or <a href="https://www.tandfonline.com/doi/full/10.1080/14680777.2021.1928258">Meghan Markle</a> or The Chicago Tribune leading on the bronze medal-winning &#8220;<a href="https://www.poynter.org/reporting-editing/2016/whats-behind-sexist-reporting-at-the-olympics-lack-of-newsroom-diversity-and-experience/">Wife of a Bears’ lineman</a>&#8220;? And how do they intersect with other forces?</p>



<p class="wp-block-paragraph">Increasingly, there will also be the potential for <strong>algorithmic bias</strong>: might the headline have been AI-generated or -selected? What training data might have informed that? Could <a href="https://www.niemanlab.org/2021/10/how-a-b-testing-can-and-cant-improve-your-headline-writing/">A/B testing</a> have shaped it?</p>



<p class="wp-block-paragraph">A role may be played by <strong>cognitive bias</strong> too: the <a href="https://onlinejournalismblog.com/2023/02/14/availability-bias-a-guide-for-journalists/">availability heuristic</a> means people are more likely to connect new information with what happened most recently and was most widely reported (e.g. John Torode&#8217;s sacking, responsible for the spike in searches at the start of that chart above).</p>



<p class="wp-block-paragraph">It&#8217;s important to emphasise that no one escapes the trap of cognitive biases, either: an opportunity for reflection around <a href="https://onlinejournalismblog.com/2020/04/07/how-to-prevent-confirmation-bias-affecting-your-journalism/">confirmation bias</a> can do this. Compare the initial reaction to the headline (did it &#8216;confirm&#8217; existing assumptions?) with the potential for a more complex picture with multiple social and cultural forces at work.</p>



<p class="wp-block-paragraph">The point here is not what the &#8216;real&#8217; cause of the headline was: it&#8217;s that confirmation bias prevents us from exploring counter hypotheses or less simple explanations before drawing that conclusion.</p>



<p class="wp-block-paragraph">Each of those biases is a force to name and address, in both critical and practical terms. We can ask what forces act to introduce bias into reporting, and at what levels (individual, team, culture, technology, audience, institution, business model)? What is considered best practice in this area? How should journalists balance chasing search traffic with their <a href="https://www.ipso.co.uk/news-analysis/ipso-analysis-how-clause-12-discrimination-works/">duty to avoid prejudice</a>, to be objective and accurate? How do we <a href="https://onlinejournalismblog.com/2025/05/07/why-im-no-longer-saying-ai-is-biased/">design prompts to avoid</a> (or <a href="https://onlinejournalismblog.com/2024/10/17/identifying-bias-in-your-writing-with-generative-ai/">identify</a>) bias? How do we recruit or source in ways that reduce the opportunities for bias? And how do we slow down to prevent cognitive biases driving our behaviour?</p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/07/03/this-one-story-can-be-used-to-discuss-seven-different-types-of-bias/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31650</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/07/torodeswife.png?w=1024" medium="image">
			<media:title type="html">ITV News headline: John Torode’s wife Lisa Faulkner reveals breast cancer diagnosis</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/07/trendstorodevfaulkner.png?w=752" medium="image">
			<media:title type="html">Line chart comparing search volumes for John Torode (peaks 12 months ago and generally higher) and Lisa Faulkner (peaks in the last week): In the last 12 months, four times as many searches were made for John Torode than Lisa Faulkner</media:title>
		</media:content>
	</item>
		<item>
		<title>Managing a mass FOI project? Here&#8217;s an AI-assisted methodology for that</title>
		<link>https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/</link>
					<comments>https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Tue, 30 Jun 2026 07:28:55 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[data journalism]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[DataHarvest]]></category>
		<category><![CDATA[foi]]></category>
		<category><![CDATA[NotebookLM]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31491</guid>

					<description><![CDATA[Sending FOIs to multiple bodies across the country to get the big picture on an issue sounds like a great idea — until the responses start to trickle in. Differences between responses often make mass FOI projects extremely time-consuming as you try to get everything into a format that allows you to ask journalistic questions [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><strong><em>Sending FOIs to multiple bodies across the country to get the big picture on an issue sounds like a great idea — until the responses start to trickle in. Differences between responses often make mass FOI projects extremely time-consuming as you try to get everything into a format that allows you to ask journalistic questions and compare different authorities. Can AI help?</em></strong></p>



<p class="wp-block-paragraph">On one recent project I decided to put together a methodology that made the process less stressful, faster and more accurate. Here&#8217;s how it works.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png"><img loading="lazy" width="960" height="540" data-attachment-id="31578" data-permalink="https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/foiai-vibe-coding-mass-foi-projects/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png" data-orig-size="960,540" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="FOI+AI vibe coding mass FOI projects" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png?w=960" alt="Data structure

Extract &amp; reshape

Check &amp; verify

Combine

Audit &amp; prioritise

Audit responses to identify the level of detail in each response and identify edge cases. Include a caveats column.
Augment manual audit with NotebookLM audit.
Identify a priority order for data, e.g. totals by outcome, hospital, category or year where these are provided separately


Design a data structure that can accommodate all responses
Structure should follow ‘tidy’ data principles, i.e. one row per combination of features (force, category, hospital, outcome, year)
Structure should include source details, e.g. filename, sheet name, name of person entering data


PDFs: use Tabula or 
vibe coding (design a prompt template to generate code to attempt to extract data). Multi-sheet XLS files: use Open Refine to import and combine sheets
Design a prompt template for generating code to reshape CSV responses


Manual checks (e.g. compare entries, check page-ending rows)
Analysis-based checks (e.g pivots, totals)
AI-based checks using a prompt template (e.g. compare files)


Use OpenRefine or: Design a prompt template for generating code to combine the resulting CSV files.
" class="wp-image-31578" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/foiai-vibe-coding-mass-foi-projects.png?w=768 768w" sizes="auto, (max-width: 960px) 100vw, 960px" /></a></figure>



<span id="more-31491"></span>



<h2 class="wp-block-heading">Problem 1: Responses are inconsistent</h2>



<p class="wp-block-paragraph">The main problem journalists face with responses to a bulk FOI request is their inconsistency. This can be reduced by specifying the year used (financial or calendar), the format of the response (spreadsheet) and even providing a template, but inevitably some authorities will ignore either or both.</p>



<p class="wp-block-paragraph">On top of that, authorities will often use slightly different language to refer to the same thing, or different levels of categories. And some </p>



<p class="wp-block-paragraph">With so many different inconsistencies, it helps to break those down and make a list that you can turn into a series of steps based on priority. For example:</p>



<ul class="wp-block-list">
<li>Inconsistent <strong>filetype</strong>: we first need to extract data into one format (CSV)</li>



<li>Inconsistent<strong> shape</strong>: &#8230;then we need to reshape data (wide to long, all fields)</li>



<li>Inconsistent <strong>naming</strong>: &#8230;then we need to standardise categories or other fields</li>



<li>Inconsistent <strong>timescale</strong>: &#8230;then we need to make an editorial decision whether to use all responses and, if needed, how to report the use of different timescales</li>
</ul>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png"><img loading="lazy" width="960" height="540" data-attachment-id="31581" data-permalink="https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/foi-responses/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png" data-orig-size="960,540" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="foi responses" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png?w=960" alt="Screenshots of spreadsheets with different structures" class="wp-image-31581" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png?w=768 768w" sizes="auto, (max-width: 960px) 100vw, 960px" /></a><figcaption class="wp-element-caption">Responses to the same question will often come in different shapes and use different terms</figcaption></figure>



<h2 class="wp-block-heading">Problem 2: You need to choose a focus</h2>



<p class="wp-block-paragraph">That inconsistency will apply across every question that you asked in the FOI — so before you spend any time making the responses consistent, it makes sense to decide which <em>parts</em> of the responses you&#8217;re actually going to need.</p>



<p class="wp-block-paragraph">This is often shaped by how many questions were answered by authorities, and to what level of detail. In the FOI project I was working on, for example, the request asked for details of crimes in hospitals, but not all forces provided totals that combined crime category and hospital names and outcomes. Our story would have to focus on either categories, or hospitals, or outcomes.</p>



<p class="wp-block-paragraph">To choose a focus you need to know how many responses answer each question — you need an <strong>audit</strong> that does the following:</p>



<ul class="wp-block-list">
<li>Browse the FOIs, or a representative sample</li>



<li>Identify the range of information covered</li>



<li>Are there any questions which some bodies refused to answer?</li>



<li>Or levels of detail they couldn’t provide?</li>



<li>Identify any differences (e.g. financial year vs calendar year)</li>



<li>Decide what takes priority where multiple tables are given (e.g. offence category, not hospital or outcome)</li>
</ul>



<p class="wp-block-paragraph">You can do this manually, but this is also a task for which AI is well suited: a classification and summarisation task to provide an overview of a large document set. Using the <a href="https://www.youtube.com/watch?v=u3w5SegQPMw">Pulitzer Center&#8217;s risk assessment</a> this task also sits in AI&#8217;s &#8220;sweet spot&#8221; because it&#8217;s not audience-facing and you don&#8217;t need high accuracy.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/u3w5SegQPMw?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">Specifically, it&#8217;s well suited to <a href="https://notebooklm.google.com/">NotebookLM</a>, an AI tool for working with documents using the Gemini LLM. </p>



<p class="wp-block-paragraph">Those documents will need to be in PDF or Word format, or Google Sheets, and there&#8217;s a limit of 50 documents (you can combined PDFs with Adobe Acrobat and other online tools). An alternative is <a href="https://journaliststudio.google.com/pinpoint/">Google Pinpoint</a>.</p>



<p class="wp-block-paragraph">A prompt template to audit your responses might look like this:</p>



<p class="wp-block-paragraph"><code>OBJECTIVE: You are a journalist looking to combine data from multiple FOI responses into a single table that can be analysed to identify trends over time and compare categories or bodies.<br>TASK: Audit these responses and produce a table identifying the range of information covered for each authority. </code><br><code>Each column should identify Y/N if they provided that information.<br>Include a column indicating what type of year was used (e.g. financial vs calendar) </code><br><code>and a column indicating whether data is provided for each incident, or per year, month, quarter, another period.<br>If the response includes any warnings or caveats, quote these and the page number in a caveats column.</code></p>



<p class="wp-block-paragraph">The resulting table should give you an idea of the coverage of the FOI responses, including which questions might have the most comprehensive responses, and which questions will have patchier data.</p>



<p class="wp-block-paragraph">You can now decide on a focus which isn&#8217;t going to lead to frustration and having to start over.</p>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(625 / 352)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="351" alt="A table classifying what's in each FOI request" class="wp-block-jetpack-slideshow_image wp-image-31631" data-id="31631" data-aspect-ratio="625 / 352" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png 960w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">A manual audit of FOI responses</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="351" alt="A table classifying what's in each FOI request" class="wp-block-jetpack-slideshow_image wp-image-31630" data-id="31630" data-aspect-ratio="625 / 352" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png 960w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">This audit of FOI responses was compiled by NotebookLM</figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<h2 class="wp-block-heading">Problem 3:&nbsp;You need to extract the data</h2>



<p class="wp-block-paragraph">Now you can start extracting that <em>target</em> data from the FOI requests. For data in Word documents and email tables, or single-sheet XLSX files, copy and paste the data into a spreadsheet and save it as a CSV. </p>



<p class="wp-block-paragraph">It&#8217;s tempting to reach straight for AI tools to tackle trickier formats, but there two tools to try first that are faster, more accurate, more explainable, less energy intensive, and avoid deskilling.</p>



<ul class="wp-block-list">
<li><strong>Try <a href="https://openrefine.org/">Open Refine</a></strong> to combine sheets from XLSX — or search DuckDuckGo to find other solutions <a href="https://www.exceldemy.com/excel-combine-data-from-multiple-sheets/">like this</a>&nbsp;</li>



<li><strong>Try <a href="https://tabula.technology/">Tabula</a></strong> first for PDF extraction</li>
</ul>



<p class="wp-block-paragraph">You should also conduct a <strong>security and privacy assessment</strong> before reaching for AI: on the whole information provided under FOI laws should not raise issues around security or privacy, but the organisation may have made a mistake, or there could be issues with the particular nature of the information provided that mean you would not want a large language model to be learning from it. </p>



<p class="wp-block-paragraph">If you need to use AI to extract data from an XLSX file where it has been provided in different sheets (for example one for each year), here&#8217;s a template prompt you can adapt:</p>


<div class="wp-block-code">
	<div class="cm-editor">
		<div class="cm-scroller">
			
<pre>
<code><div class="cm-line">This XLS file contains a sheet for each financial year. </div><div class="cm-line">Each sheet has three tables.</div><div class="cm-line">Extract the second table, in columns D-E, from each sheet</div><div class="cm-line">Create a CSV containing the combined tables with the following columns:</div><div class="cm-line">Offence Code | Total | Year from | Year to</div><div class="cm-line">Ignore the first two sheets which do not relate to financial years</div><div class="cm-line"></div><div class="cm-line"></div><div class="cm-line"></div></code></pre>
		</div>
	</div>
</div>


<p class="wp-block-paragraph">The prompt should provide <strong>context</strong> for what is in the sheets, name <strong>specific columns or rows</strong>, specify the <strong>structure</strong> and the <strong>output</strong>, and tell it what <em>not</em> to do.</p>



<p class="wp-block-paragraph">PDF extraction is more problematic, because there&#8217;s an explainability challenge (it&#8217;s hard to explain how it extracted the data). I&#8217;m going to suggest a way to address this in a moment, but if you have to use AI for PDF extraction, here&#8217;s a template prompt for that:</p>


<div class="wp-block-code">
	<div class="cm-editor">
		<div class="cm-scroller">
			
<pre>
<code><div class="cm-line">Attached is a PDF. </div><div class="cm-line">Extract ONLY the table on pages 2 and 3 which shows crime totals by category and year. </div><div class="cm-line">If any line breaks or hyphenations cause split labels, reconstruct using nearest-neighbor/line-merge logic and re-validate.</div><div class="cm-line">Export as a CSV.</div><div class="cm-line">Check you have grabbed rows at the ends and beginnings of pages. </div><div class="cm-line">Check that figures in the extracted CSV, when totalled match totals in the PDF</div><div class="cm-line">If any uncertainty remains (e.g., an OCR-confused label), </div><div class="cm-line">include a “flags” list with the exact text span and your best-guess normalization.</div><div class="cm-line"></div><div class="cm-line"></div><div class="cm-line"></div></code></pre>
		</div>
	</div>
</div>


<p class="wp-block-paragraph">Again, this names specific columns and rows, and specifies output. It addresses particular problems with PDF extraction: line breaks where column or row titles are split, provides space for uncertainty (AI is ultimately a machine for guessing) and includes a validation step.</p>



<p class="wp-block-paragraph">That reduces some of the most common risks in using AI for this — but explainability remains a problem: how do we know what it did? And how can we use that to help us <strong>verify</strong> the results? </p>



<p class="wp-block-paragraph">A better approach than is to ask an AI tool help <em>us</em> to extract the data — by giving us some code that we can run ourselves in <a href="https://colab.research.google.com/">Google Colab</a> (an app in Google Drive, so it requires no installation).</p>



<p class="wp-block-paragraph">The template prompt for this is long, so I&#8217;m not going to paste it here. Instead you can <a href="https://github.com/paulbradshaw/mass_foi_projects/blob/main/prompts/code_extract_pdf.md">find it on the GitHub repo</a> for my Dataharvest talk on this. But it has a few key elements to highlight:</p>



<ul class="wp-block-list">
<li>Set a clear objective including any prioritisation of data to extract</li>



<li>Break down the task step by step </li>



<li>Describe the PDF and the specific pages and tables you want to target, including the structure. Say whether the PDF has gridlines and any text that appears before or after it. </li>



<li>Provide any extra details </li>
</ul>



<p class="wp-block-paragraph">One thing to highlight in the prompt: it should specify that code &#8220;be commented so that a non-coder can understand what is happening&#8221; (explainability again).</p>



<p class="wp-block-paragraph">Once the AI tool has provided you with some code, <a href="http://colab.research.google.com/">create a Colab notebook</a> and copy and paste that code into the empty code block, then run it by pressing the &#8216;play&#8217; button to the left. </p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-4-3 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/UaFDfWH7Zf4?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">Once a code block has run, look for an ‘Upload’ button appearing underneath the code block, where you can upload the PDF</p>



<p class="wp-block-paragraph">Create extra code blocks for further code by clicking the button marked &#8216;+ Code&#8217; at the top of the notebook (the same button will also appear above or below a block when you hover). </p>



<p class="wp-block-paragraph">Once you&#8217;ve done this for one response, you can adapt the same template prompt for other responses, adjusting any details on the PDF description etc.</p>



<p class="wp-block-paragraph">For extra confidence (thanks to someone at Dataharvest for suggesting this), you can ask an AI tool to generate code again, but in a different language (<a href="https://www.geeksforgeeks.org/r-language/how-to-use-r-with-google-colaboratory/">you can run R in Colab</a>). This prompt <a href="https://github.com/paulbradshaw/mass_foi_projects/blob/main/prompts/code_extract_pdf.md#a-prompt-for-generating-a-second-script-in-r">would be the same with one change</a>:</p>



<p class="wp-block-paragraph"><code>Look at the attached PDF and suggest code in R that will work in Colab notebook (explain how I run R in Colab).</code></p>



<h2 class="wp-block-heading">Problem 4:&nbsp;You need to check it’s accurate</h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png"><img loading="lazy" width="960" height="540" data-attachment-id="31634" data-permalink="https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/checklist-for-extraction/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png" data-orig-size="960,540" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="checklist for extraction" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png?w=960" alt="Data validation checklist
Check
Blanks/merged cells in original: check values haven’t shifted
Multi‑line cells: check for no truncation
First/last rows on PDF pages: check values extracted
Headers &amp; labels: check extracted and match data type
Totals: check they are extracted/omitted as needed
Non-numeric data: (“–”, “&lt;5”): extracted accurately
Spot‑checks: for individual values at random
Values count: actual numbers vs expected (e.g. 5 cols x 8 rows)
Non-numeric data: (“–”, “&lt;5”): impacts on totals/counts 
Row totals*: calculate from values &amp; match to PDF totals
Column totals*: calculate &amp; match to PDF
Pivot table**: construct &amp; match to PDF
Sequence/ID continuity***: check for gaps using formulae
AI check: value matching
AI check: row matching
AI check: total checks
AI check: label checks
AI check: count checks

" class="wp-image-31634" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/checklist-for-extraction.png?w=768 768w" sizes="auto, (max-width: 960px) 100vw, 960px" /></a><figcaption class="wp-element-caption">Create a data validation checklist to ensure you cover all the checks for each extracted dataset</figcaption></figure>



<p class="wp-block-paragraph">After you&#8217;ve extracted the data from a response, you need to check that it&#8217;s extracted the data correctly. I combined three methods to do this:</p>



<ol class="wp-block-list">
<li><strong>“Eyeball” tests</strong>: manual checks for known risks, vulnerabilities and issues</li>



<li><strong>Analysis tests</strong>: calculations such as totals and value counts to cross-check</li>



<li><strong>AI tests</strong>: document comparison and repeat of above tests</li>
</ol>



<p class="wp-block-paragraph">Particular things to check manually (&#8220;eyeball&#8221; tests) include where the PDF has blank or merged cells (values sometimes get shifted here), where text runs across multiple lines (sometimes only the first line is extracted), and the first and last rows on each page. </p>



<p class="wp-block-paragraph">Analysis tests are where you don&#8217;t compare values directly, but add up all your numbers to see if it matches a total in the original. Any discrepancy is extraction should result in a different figure. </p>



<p class="wp-block-paragraph">I&#8217;ve created a <a href="https://docs.google.com/document/d/1uQZ6dpAIpPwABLkLotBiQucXXFye7X48Q-woA3QBBLU/edit?usp=sharing">template checklist</a> which lists all the tests I could identify. Using a checklist makes it possible to systematically go through and tick each test (some will not be applicable, e.g. where you don&#8217;t have totals in the original PDF). </p>



<p class="wp-block-paragraph">The AI tests just repeat the above, adding an extra layer on top of the manual checks, which can be useful in identifying human error in your own checks. A <a href="https://github.com/paulbradshaw/mass_foi_projects/blob/main/prompts/checking.md">prompt template for AI validation</a> is also on the repo (again, it&#8217;s too long to paste here). </p>



<h2 class="wp-block-heading">Problem 5: You need to get the responses into one dataset</h2>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(625 / 352)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="351" alt="A table with a column for year from and year to, and separate category and code" class="wp-block-jetpack-slideshow_image wp-image-31608" data-id="31608" data-aspect-ratio="625 / 352" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png 960w" sizes="(max-width: 625px) 100vw, 625px" /></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="351" alt="Table with a column for each year and descriptions mixing codes and categories" class="wp-block-jetpack-slideshow_image wp-image-31609" data-id="31609" data-aspect-ratio="625 / 352" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png 960w" sizes="(max-width: 625px) 100vw, 625px" /></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">Now you have a collection of CSVs with data extracted from FOI responses, and checked against the originals. The next step is to combine those into a single dataset you can question.</p>



<p class="wp-block-paragraph">That means deciding on a single structure. It will need to be a structure that can accommodate the variety of responses, and that allows you to ask different questions (e.g. pivot tables). Some tips for that structure include:</p>



<ul class="wp-block-list">
<li><strong>Have columns for filename and authority</strong> for traceability/combination</li>



<li><strong>Have ‘year from’/‘year to’ </strong>and store the year as a value, not a column name (the data should be long, not wide)</li>



<li>Have a column that measures the <strong>number</strong> of events (e.g. crimes) — or two columns for &#8216;number from&#8217; and &#8216;number to&#8217; if organisations are likely to give a range (&lt;5 can be encoded as from 1 to 5, for example)*</li>



<li><strong>Have a category column</strong> that classifies the type of event (e.g. ‘theft’)</li>



<li>You might have sub-category column as well, or columns for category code, outcome, location, etc.</li>



<li><strong>Allow columns to have ‘NOT SPECIFIED’</strong> where bodies didn’t provide that level of detail (or combined detail)</li>
</ul>



<p class="wp-block-paragraph">Once you&#8217;ve identified a structure, you&#8217;re going to need to reshape some of your data to fit into that. This is another task that can be done well by AI (there are manual alternatives if you are able to code), but again it is better to ask AI to generate and comment code that you can run, rather than performing the reshaping itself.</p>



<p class="wp-block-paragraph"><a href="https://github.com/paulbradshaw/mass_foi_projects/blob/main/prompts/code_reshape.md">A prompt template for reshaping data in this way is available here</a>. It sets an objective, describes the target structure, and describes the task step by step. </p>



<p class="wp-block-paragraph">You can validate reshaped data against the original using the same checklist as before (or you may decide to only validate at this point). In particular, where data has been reshaped from a pivot table-type response, you can validate it against the original data by generating pivot tables to replicate the original shape for comparison. </p>



<h2 class="wp-block-heading">Problem 6:&nbsp;You need to clean inconsistent categories or entities</h2>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/-aa02-9lf8o?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">It&#8217;s likely that different authorities will have used different language to refer to the same entities or similar categories. This is a problem if you want to tell a story about how many events there were in a particular category (or connected to an entity of interest), or which category or entity ranked top or bottom. </p>



<p class="wp-block-paragraph">Fixing this problem is best left until you have all the responses in a single dataset, as this makes it easier to identify clusters, and means you only have to clean one dataset, rather than each one. </p>



<p class="wp-block-paragraph">Once you have a single dataset, generate a list of all the unique categories or entities, and have a quick look to get a feel for the problems it presents. The key here isn&#8217;t to see this as a single cleaning job —&nbsp;you&#8217;ll need to break it up into multiple cleaning stages, each of which tackles a different problem. </p>



<p class="wp-block-paragraph">Here are six likely stages you might need to clean for, in the following likely order:</p>



<ol class="wp-block-list">
<li><strong>Mixed data</strong> (for example, category and code, or crime category and Act/Section): this might be fixed by <a href="https://www.exceldemy.com/convert-text-to-columns-excel/">using the spreadsheet&#8217;s <em>Text to columns </em>tool</a>. This is probably the first stage of cleaning to perform, as it will isolate strings of text for further cleaning.</li>



<li><strong>Acronyms</strong> (for example some authorities using GBH while others use “Grievous bodily harm”): use an <code>IF</code> or <a href="https://www.extendoffice.com/documents/excel/1055-excel-identify-case.html"><code>EXACT</code> function combined with <code>UPPER</code></a> or <code><a href="https://spreadsheetpoint.com/formulas/regexmatch-function-google-sheets/">REGEXMATCH</a></code> (Google Sheets only) to identify all-caps entries.</li>



<li><strong>Upper/lower case inconsistency</strong>: this can be fixed by using a text formatting function like <a href="https://www.excelmojo.com/proper-excel-function/"><code>PROPER</code></a>. However, this will also turn acronyms like GBH to &#8220;Gbh&#8221; so you will need to fix acronyms first, or filter them out while cleaning everything to proper or lower case.</li>



<li><strong>Extra/analogous characters</strong>: some responses may use quotation marks, spaces, apostrophes, commas, and currency symbols where others don&#8217;t. Or some may be “and” while others use “&amp;”. Using <em>Edit &gt; Find and replace</em> can deal with some of this, as can the <code>SUBSTITUTE</code> function. The <code>TRIM</code> function can get rid of spaces at the start and end, and <code><a href="https://support.google.com/docs/answer/3098245?hl=en">REGEXREPLACE</a></code> can replace certain patterns of text.</li>



<li><strong>Slight variations</strong> (e.g. one authority uses &#8220;Criminal damage and arson&#8221; while another users &#8220;Criminal damage and arson offences&#8221; or &#8220;Starbucks&#8221; and &#8220;Starbucks Ltd&#8221;): Open Refine&#8217;s <a href="https://www.youtube.com/watch?v=-aa02-9lf8o">cluster and edit tool</a> is very powerful for tackling this, and you can also sort the categories to bring similar ones together to edit them manually</li>



<li><strong>Different category level</strong> (e.g. some use the top level categories, others sub-category, and others lower level categories): this is the trickiest to deal with, and best left until you&#8217;ve cleaned for all the problems above. You will need a lookup dataset with categories at different levels, and then a lookup function to make matches and establish category levels. </li>
</ol>



<p class="wp-block-paragraph">Obtaining reference data on categories is extremely useful in general: it helps give you a target for your own data. If there are 12 official top level categories, for example, then you know that you need to get your list into those 12 categories (by assigning sub categories to their parent category, for example)</p>



<p class="wp-block-paragraph">For crime the&nbsp;<a href="https://www.gov.uk/government/publications/counting-rules-for-recorded-crime">Home Office Crime Recording Rules for frontline officers, staff</a>&nbsp;provides a very useful lookup table listing four different levels of category for crime (from specific offences to offence categories, sub classes and classes) as well as their Home Office code, any of which might be used in responses.</p>



<p class="wp-block-paragraph">You can use AI tools to assist with &#8216;auditing&#8217; your data along these lines too. <a href="https://github.com/paulbradshaw/mass_foi_projects/blob/main/prompts/auditcategorylist.md">A template prompt is suggested here</a>. This also provides code that you can use to run the audit yourself, and role prompting to highlight any potential problems.</p>



<p class="wp-block-paragraph">The output from that prompt can be used to sort the data for each stage of cleaning. Make sure that you retain the original list of &#8216;dirty&#8217; terms alongside the progressively cleaned versions. This will give you a lookup table you can use to translate from inconsistent categories to a consistent list. </p>



<p class="wp-block-paragraph">A major benefit of this approach is that the <strong>results are valuable for any future FOI project</strong>, too. That lookup from dirty to clean categories can be used and reused.</p>



<h2 class="wp-block-heading">Finally, the project is ready to answer questions accurately</h2>



<p class="wp-block-paragraph">At this point you&#8217;ve broken down and tackled each of the distinctive problems that mass FOI projects involve:</p>



<ul class="wp-block-list">
<li>You&#8217;ve extracted information from responses in different formats, and put them in a single format (CSV)</li>



<li>You&#8217;ve converted information in different shapes (wide, long) and put them in a single shape</li>



<li>You&#8217;ve validated that extracted information to check that data is consistent with the source, and is not missing</li>



<li>You&#8217;ve combined data from different responses into a single file</li>



<li>You&#8217;ve standardised the language used in different responses so that answers to questions will be accurate and consistent</li>
</ul>



<p class="wp-block-paragraph">Now you should finally have a dataset that can answer questions.</p>



<h2 class="wp-block-heading">Giving the whole problem to AI agents doesn&#8217;t work</h2>



<p class="wp-block-paragraph">One final note: I tried agentic AI tools — ChatGPT&#8217;s Codex and Claude Code — on two of these problems, to see how their performance compared with the more hybrid human+AI approach. Both produced <em>plausible</em> outputs that could easily have been mistaken for the data trapped in the FOI responses, but&#8230;</p>



<p class="wp-block-paragraph">When given the task of getting a folder of FOI responses from police forces into a consistent shape, the resulting CSV was <strong>missing half of incidents and one third of forces</strong>. Most forces were missing incidents. </p>



<p class="wp-block-paragraph">The same problems were seen when Codex and Claude Code were given the challenge of cleaning up categories with a target table of official categories. Codex only managed to classify 35% of the categories correctly, while Claude Code managed 53%.</p>



<p class="wp-block-paragraph">More detailed prompting (perhaps drawing on the experiences above) might produce better results, but the results aren&#8217;t really the challenge here: the key challenge is how to <em>know</em> that you have all the results, or the correct matches. For that, there still needs to be some manual validation step, especially focused on typical problems with PDF extraction.</p>



<p class="wp-block-paragraph"><strong><em>This is a work in progress. If you have other tips, or a collection of FOI responses to test with this, let me know.</em></strong></p>



<p class="wp-block-paragraph"><em>*Thanks to Martin Rosenbaum for <a href="https://www.linkedin.com/feed/update/urn:li:activity:7478749956690964480?commentUrn=urn%3Ali%3Acomment%3A%28activity%3A7478749956690964480%2C7478763922133659649%29&amp;dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287478763922133659649%2Curn%3Ali%3Aactivity%3A7478749956690964480%29">suggesting</a> using two columns for &#8216;value from&#8217; and &#8216;value to&#8217; where organisations do not provide a single number.</em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/06/30/managing-a-mass-foi-project-heres-a-methodology-for-that/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31491</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/foi-responses.png?w=960" medium="image">
			<media:title type="html">Screenshots of spreadsheets with different structures</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditmanual.png?w=625" medium="image">
			<media:title type="html">A table classifying what&#039;s in each FOI request</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/auditai.png?w=625" medium="image">
			<media:title type="html">A table classifying what&#039;s in each FOI request</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/tidystructure.png?w=625" medium="image">
			<media:title type="html">A table with a column for year from and year to, and separate category and code</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/badstructure.png?w=625" medium="image">
			<media:title type="html">Table with a column for each year and descriptions mixing codes and categories</media:title>
		</media:content>
	</item>
		<item>
		<title>Showing charts on video? Here are two essential techniques to make them effective</title>
		<link>https://onlinejournalismblog.com/2026/06/23/showing-charts-on-video-here-are-two-essential-techniques-to-make-them-effective/</link>
					<comments>https://onlinejournalismblog.com/2026/06/23/showing-charts-on-video-here-are-two-essential-techniques-to-make-them-effective/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Tue, 23 Jun 2026 13:31:49 +0000</pubDate>
				<category><![CDATA[data journalism]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[online video]]></category>
		<category><![CDATA[television]]></category>
		<category><![CDATA[Animation]]></category>
		<category><![CDATA[dataviz]]></category>
		<category><![CDATA[design]]></category>
		<category><![CDATA[motion]]></category>
		<category><![CDATA[remove to improve]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31361</guid>

					<description><![CDATA[Using visualisation on TV and video is very different to using charts and maps online. In video, the audience has very little time to absorb the information contained in the chart — so you need to get them to that information as quickly as possible. Every bad example of charts in videos forgets this. And [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Using visualisation on TV and video is very different to using charts and maps online. In video, the audience has very little time to absorb the information contained in the chart — so you need to get them to that information as quickly as possible.</p>



<p class="wp-block-paragraph">Every bad example of charts in videos forgets this. And every good example uses two essential techniques: <strong>keeping things simple</strong>, and<strong> adding motion</strong>.</p>



<span id="more-31361"></span>



<h2 class="wp-block-heading">Remove to improve: edit your chart ruthlessly</h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif"><img loading="lazy" width="640" height="460" data-attachment-id="31469" data-permalink="https://onlinejournalismblog.com/2026/06/23/showing-charts-on-video-here-are-two-essential-techniques-to-make-them-effective/removetoimprove/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif" data-orig-size="640,460" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="removetoimprove" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif?w=640" alt="Remove to improve: animated gif showing a chart having elements steadily removed (gridlines, legend, etc), improving clarity" class="wp-image-31469" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif 640w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif?w=300 300w" sizes="auto, (max-width: 640px) 100vw, 640px" /></a><figcaption class="wp-element-caption">Remove to improve is a key principle of good datavis — but it&#8217;s especially important in video</figcaption></figure>



<p class="wp-block-paragraph">Viewers looking at an on-screen chart have to take in at least three types of information in a short space of time: <strong>shapes</strong> (bars, slices, lines, etc.), <strong>colours</strong>, and <strong>text</strong> (titles, subtitles, labels, annotations and legends).</p>



<p class="wp-block-paragraph">The more you can reduce the time needed to interpret each of those, the quicker the viewer can understand the chart.</p>



<p class="wp-block-paragraph">The best place to start is <strong>reducing the number of data points</strong> to those that are absolutely essential. This will reduce the number of shapes the viewer has to absorb. </p>



<p class="wp-block-paragraph">If your story is about ranking the worst or best areas for a problem, for example, a top ten is too much for video. Cut back to just the top three or five. </p>



<p class="wp-block-paragraph">If you are comparing change over time with a multiple line chart, experiment to find out what the smallest number of lines is that you can use and it still work.</p>



<p class="wp-block-paragraph">If your story is about parts of a whole then, again, you don&#8217;t need to break down each part: your pie chart might simply show two slices: the part you are focusing on, and the rest of the whole.</p>



<p class="wp-block-paragraph">Next, <strong>reduce the number of colours</strong>: this is good practice generally, but on video it&#8217;s especially important: make the element that is the focus of your story (the bar, line or slice that the story is about) a strong colour. Then, make all other elements a neutral or non-colour such as grey or pale blue. This tells the viewer immediately which bar, line or slice they should be looking at.</p>



<h3 class="wp-block-heading">Good titles remove the need to read labels or axes</h3>



<p class="wp-block-paragraph">Now, <strong>reinforce the story with text</strong>: edit the title of the chart so that it tells the story, succinctly. For example, a bar chart showing the worst places for pollution that is titled &#8220;Worst cities for pollution&#8221; or &#8220;Pollution in each city&#8221; creates unnecessary work for your viewer. A title that says <em>&#8220;Scunthorpe is the most polluted city in the UK&#8221;</em> tells them precisely what the chart is doing — and what city the highlighted bar relates to, saving them the work of trying to find the label for that bar.</p>



<p class="wp-block-paragraph">Similarly a title that says &#8220;Crime has dropped 10% in the last three years&#8221; or &#8220;40% of donations to the party came from Company X&#8221; allow a viewer to understand a chart without having to check labels or axes compare slices or lines.</p>



<p class="wp-block-paragraph">You should also <strong>remove unnecessary text</strong>. For example there&#8217;s no need for both a legend and direct labelling of bars: direct labelling is quicker and easier to understand. Likewise, it might be easier to directly label values rather than place them on an axis. Subheadings are likely to be a luxury you can do without. Use trial and error to see how what can be removed or moved — make sure you test drafts on someone unfamiliar with the chart.</p>



<p class="wp-block-paragraph">You may be able to remove further elements from the chart if they are being narrated in the video, or if animation will add them later.</p>



<h2 class="wp-block-heading">Direct attention with motion</h2>



<figure class="wp-block-image size-large is-resized"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif"><img loading="lazy" width="320" height="180" data-attachment-id="31485" data-permalink="https://onlinejournalismblog.com/2026/06/23/showing-charts-on-video-here-are-two-essential-techniques-to-make-them-effective/linechartwipe/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif" data-orig-size="320,180" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="linechartwipe" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif?w=320" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif?w=320" alt="Line chart with a wipe effect" class="wp-image-31485" style="width:494px;height:auto" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif 320w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif?w=300 300w" sizes="auto, (max-width: 320px) 100vw, 320px" /></a></figure>



<p class="wp-block-paragraph">Colour and text are just two techniques for directing attention in data visualisation — but in video there&#8217;s a third: <strong>motion</strong>.</p>



<p class="wp-block-paragraph">Here are just some of the methods you can use to add motion to a chart:</p>



<p class="wp-block-paragraph"><strong>Zoom</strong>: zoom in to a focal data point to direct attention, or to zoom out to put a data point into the bigger picture. Zooming might involve <a href="https://x.com/cherdarchuk/status/1525178697568747520">expanding or shrinking the scale</a> to illustrate how a particular data point affects it.</p>



<p class="wp-block-paragraph"><strong>Reveal</strong>: instead of the chart appearing fully-drawn, it can be revealed in stages. For example, we might start with the whole amount, and then reveal specific parts (slices) of that amount, one by one. Or we might start with a scatterplot, then <a href="https://www.youtube.com/watch?v=jV49bDzkQhY&amp;t=64s">add a trend line, then add a highlighted area</a>, and so on. Typically this is accompanied by narration that relates to each reveal. </p>



<p class="wp-block-paragraph"><strong>Wipe</strong>: often used with a line chart. The wipe <a href="https://www.youtube.com/watch?v=S1m-KgEpoow&amp;t=364s">reveals the line from left to right</a>. Because a line chart shows change over time this is particularly effective as we see that line <em>change over time</em>. A vertical wipe can also be used with bar charts and histograms to create the effect of the bar(s) <a href="https://www.tiktok.com/@theeconomist/video/7606417592877714710">growing upwards or outwards</a>. Wipes can even be used to reveal titles or labels to &#8216;read&#8217; those out.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/S1m-KgEpoow?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;start=364&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph"><strong>Highlight or fade</strong>: an element can be highlighted — or the rest of the chart quickly faded — to direct attention to that particular element or series of elements. Conversely, a focal element can be faded back or de-highlighted to indicate that the focus is about to shift to the whole, or to another part (which might then be highlighted in turn).</p>



<p class="wp-block-paragraph"><strong>Removal</strong>: instead of revealing elements, you can hide them. So instead of starting with simplicity and gradually adding more detail, we start with complexity and strip it back to focus on something in particular.</p>



<p class="wp-block-paragraph"><strong>Scroll</strong> <strong>or pan</strong>: a more complex chart might need to be revealed vertically, as we scroll down to specific elements of interest. The Economist&#8217;s <a href="https://www.tiktok.com/@theeconomist/video/7606417592877714710">Epstein Files video on TikTok</a> provides one example. Or you might <a href="https://x.com/cherdarchuk/status/1525178675225710592">pan across a wide chart horizontally</a> to emphasise a scale.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@theeconomist/video/7606417592877714710" data-video-id="7606417592877714710" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@theeconomist" href="https://www.tiktok.com/@theeconomist?refer=embed">@theeconomist</a> <p>Who had the most contact with Jeffrey Epstein in the final decade of his life? In the years after Epstein pled guilty to soliciting sex from a minor, he maintained consistent contact with a vast network of rich and powerful figures. The Economist’s data team analysed 1.4m of Epstein’s emails to map their relationships. <a title="epstein" target="_blank" href="https://www.tiktok.com/tag/epstein?refer=embed">#Epstein</a> <a title="epsteinfiles" target="_blank" href="https://www.tiktok.com/tag/epsteinfiles?refer=embed">#Epsteinfiles</a> <a title="data" target="_blank" href="https://www.tiktok.com/tag/data?refer=embed">#data</a> <a title="news" target="_blank" href="https://www.tiktok.com/tag/news?refer=embed">#news</a> <a title="analysis" target="_blank" href="https://www.tiktok.com/tag/analysis?refer=embed">#analysis</a></p> <a target="_blank" title="♬ original sound - The Economist - The Economist" href="https://www.tiktok.com/music/original-sound-The-Economist-7606417673916631830?refer=embed">♬ original sound &#8211; The Economist &#8211; The Economist</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph"><strong>Counter</strong>: often used for scale stories that focus on a single number, this technique uses a counter to draw attention to the final figure that it arrives at.</p>



<p class="wp-block-paragraph"><strong>Annotation</strong>: extra text is added on screen to direct attention to a specific element, often with an arrow or circle to connect the text with that element. This is especially useful with scatterplots where the story might need to focus on different data points (you can <a href="https://www.youtube.com/watch?v=jV49bDzkQhY&amp;t=63s">see an example in this Pudding video at 1&#8217;03</a>). </p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/jV49bDzkQhY?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;start=56&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph"><strong>Timelapse</strong>: bar chart races are a particularly popular example of animating the transition between multiple &#8216;snapshots&#8217; of a chart at different points in time, and the same principle can be applied to maps and other charts (Joey Chedarchuk <a href="https://x.com/cherdarchuk/status/1525160488589414403?s=21&amp;t=MxDxMWULNYaUSO8_okcXug">collects a number of examples in this Twitter thread</a>).</p>



<p class="wp-block-paragraph"><strong>Presenter interaction</strong>: a common way to add motion with charts on TV is for a presenter to interact with it: <a href="https://www.youtube.com/watch?v=lCeea3jOvs8">pointing</a> to different parts of it, or &#8216;operating&#8217; the chart via touchscreen, clicker or <a href="https://www.youtube.com/watch?v=rTbMHBkdw2U&amp;t=23s">tablet</a>. A classic example is Hans Rosling&#8217;s <em><a href="https://www.youtube.com/watch?v=Z8t4k0Q8e8Y&amp;time_continue=31&amp;source_ve_path=NzY3NTg&amp;embeds_referring_euri=https%3A%2F%2Fdocs.google.com%2F&amp;embeds_referring_origin=https%3A%2F%2Fdocs.google.com">200 years in 4 minutes</a></em>. </p>



<p class="wp-block-paragraph">A newer form of this on TikTok is for a presenter to <a href="https://www.tiktok.com/@theeuropeancorrespondent/video/7637181529218731296">appear on top of a particular part the chart and then move to a different point</a> to match the focus of the narration.</p>



<figure class="wp-block-embed is-type-video is-provider-tiktok wp-block-embed-tiktok"><div class="wp-block-embed__wrapper">
<div class="embed-tiktok"><blockquote class="tiktok-embed" cite="https://www.tiktok.com/@theeuropeancorrespondent/video/7637181529218731296" data-video-id="7637181529218731296" data-embed-from="oembed" style="max-width:605px; min-width:325px;"> <section> <a target="_blank" title="@theeuropeancorrespondent" href="https://www.tiktok.com/@theeuropeancorrespondent?refer=embed">@theeuropeancorrespondent</a> <p>Your internet might be running faster or slower than that of your fellow Europeans, depending on what part of the Continent you live in.</p> <a target="_blank" title="♬ Originalton - The European Correspondent - The European Correspondent" href="https://www.tiktok.com/music/Originalton-The-European-Correspondent-7637181541692574497?refer=embed">♬ Originalton &#8211; The European Correspondent &#8211; The European Correspondent</a> </section> </blockquote> <script async src="https://www.tiktok.com/embed.js"></script></div>
</div></figure>



<p class="wp-block-paragraph">Do you know of any other key techniques for using charts in video storytelling? Let me know in the comments or at <a href="https://www.linkedin.com/in/paulbradshawuk/">linkedin.com/in/paulbradshawuk</a></p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/06/23/showing-charts-on-video-here-are-two-essential-techniques-to-make-them-effective/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31361</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/removetoimprove.gif?w=640" medium="image">
			<media:title type="html">Remove to improve: animated gif showing a chart having elements steadily removed (gridlines, legend, etc), improving clarity</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/linechartwipe.gif?w=320" medium="image">
			<media:title type="html">Line chart with a wipe effect</media:title>
		</media:content>
	</item>
		<item>
		<title>&#8220;Many newsrooms are not optimised for what humans do best&#8221; (but we have an opportunity to change that)</title>
		<link>https://onlinejournalismblog.com/2026/06/18/many-newsrooms-are-not-optimised-for-what-humans-do-best-but-we-have-an-opportunity-to-change-that/</link>
					<comments>https://onlinejournalismblog.com/2026/06/18/many-newsrooms-are-not-optimised-for-what-humans-do-best-but-we-have-an-opportunity-to-change-that/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Thu, 18 Jun 2026 14:55:23 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[Agnes Stenbom Swedling]]></category>
		<category><![CDATA[automation]]></category>
		<category><![CDATA[cory doctorow]]></category>
		<category><![CDATA[human in the loop]]></category>
		<category><![CDATA[reverse centaurs]]></category>
		<category><![CDATA[The AI Shift]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31459</guid>

					<description><![CDATA[Some essential reading by Agnes Stenbom Swedling explores how news organisations integrate AI into their workflows and the idea of the &#8220;human in the loop&#8220;. Many newsrooms, she points out, &#8220;are not optimised for what humans do best&#8221;, and so far the introduction of AI hasn&#8217;t involved a critical consideration of whether we want to [&#8230;]]]></description>
										<content:encoded><![CDATA[
<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png"><img loading="lazy" width="1024" height="563" data-attachment-id="31572" data-permalink="https://onlinejournalismblog.com/2026/06/18/many-newsrooms-are-not-optimised-for-what-humans-do-best-but-we-have-an-opportunity-to-change-that/reversecentaurdoctorow/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png" data-orig-size="1658,912" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="reversecentaurDoctorow" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=1024" alt="Amazon worker with horse head" class="wp-image-31572" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=1440 1440w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png 1658w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">Image by Cory Doctorow, from <a href="https://onezero.medium.com/revenge-of-the-chickenized-reverse-centaurs-b2e8d5cda826">Revenge of the Chickenized Reverse-Centaurs</a></figcaption></figure>



<p class="wp-block-paragraph">Some <a href="https://reutersinstitute.politics.ox.ac.uk/news/are-human-journalists-truly-irreplaceable-how-safeguard-public-interest-journalism-age-ai">essential reading by Agnes Stenbom Swedling</a> explores how news organisations integrate AI into their workflows and the idea of the &#8220;<strong>human in the loop</strong>&#8220;. Many newsrooms, she points out, &#8220;are not optimised for what humans do best&#8221;, and so far the introduction of AI hasn&#8217;t involved a critical consideration of whether we want to embed those features in new systems, or rethink them:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">&#8220;What is being built – incrementally, often unintentionally – is a form of machine-centric hybridisation. Workflows are optimised for what machines do well: speed, scale, pattern recognition, cost efficiency. Humans are then positioned around those systems, adapting their tasks, roles, and decision-making to fit the logics of machines. </p>



<p class="wp-block-paragraph">&#8220;The consequence is a subtle but significant inversion: rather than engaging in uniquely human activities, work is reorganised to fit machine-driven processes. And once that inversion is embedded at the infrastructural level, it becomes increasingly difficult to reverse.&#8221;</p>
</blockquote>



<span id="more-31459"></span>



<p class="wp-block-paragraph">The writer Cory Doctorow has a term for those humans: <strong><a href="https://boingboing.net/2026/05/14/cory-doctorows-reverse-centaur-book-why-ai-is-hell-on-workers.html">reverse centaurs</a></strong>, &#8220;a person who has been conscripted to assist a machine.&#8221;</p>



<p class="wp-block-paragraph">A recent issue of the FT newsletter The AI Shift <a href="https://ep.ft.com/permalink/emails/eyJlbWFpbCI6ImRmNTAxNzhlMmNlYzVlZjZiZDY2ZmM3Nzg0NGEyMjRkZWNjNjU2ODk0YjhkNTQiLCAidHJhbnNhY3Rpb25JZCI6ImNlOTcxNjY5LTUwZTYtNDQ3Ny04Nzg4LWZhOTczZDgyNTM1YiIsICJiYXRjaElkIjoiODk4ZjM5MWQtNjhjZC00NzRlLThmZWYtM2FhMDNkYmQ4ODBiIn0=">documents</a> this among software developers: &#8220;the most creative activities of the job are being replaced by the tedium of checking machine output,&#8221; writes one. </p>



<p class="wp-block-paragraph">In contrast, a plain old &#8220;centaur&#8221; in tech terms is someone who enlists technology to assist <em>them</em>. <strong>Florent Daudens</strong> <a href="https://fdaudens.substack.com/p/meet-your-new-colleagues">highlights</a> one example of this recently in a Toronto Star <a href="https://www.thestar.com/gift-redeem?t=28e68667-609a-441d-9e9a-2c216e34c5d6">investigation</a> into Ontario&#8217;s strong-mayor powers. &#8220;The method,&#8221; he writes, &#8220;deserves just as much attention and shows what happens when you apply editorial judgment at scale. The team first made their editorial judgment explicit [and] a custom AI tool then applied those criteria across thousands of documents. </p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">&#8220;The reporter hadn’t been replaced. They’d been <em>amplified</em>. By an agent built around their judgment, not against it.&#8221;</p>
</blockquote>



<p class="wp-block-paragraph">What these examples point to is that machine-centric hybridisation is not inevitable. At the same time as that dark future looms (if many of us aren&#8217;t already living it), AI also opens up a counter-intuitive opportunity to make work <em>more</em> human. </p>



<p class="wp-block-paragraph">After all, one of the qualities that often makes work feel dehumanising is when it becomes overly routine. Many newsroom workflows are built on such habits: check the news diary and forward planning, check the emails, check the wires, rewrite a press release. </p>



<p class="wp-block-paragraph">A part-automated workflow introduces opportunities to step back from that routine because:</p>



<ul class="wp-block-list">
<li>We have to <strong>design</strong> prompts and strategies, and <strong>engage</strong> with agents in ways that force us to make more explicit <strong>judgements</strong> about values</li>



<li>We spend less time on (re)producing, monitoring and chasing, and more time on <strong>imagining, choosing, checking, reviewing and editing. </strong></li>
</ul>



<p class="wp-block-paragraph">The AI Shift newsletter documents this too. One developer explains that their work now focuses more on “architectural decisions (I.e. how do we best assemble all of those small chunks of code to do something complicated?) </p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">&#8220;This aspect of development is often overlooked by non-technical leaders because, yes, in the short term, you *can* mash a load of LLM-generated code together and be somewhat confident it’ll do the thing you’ve asked it to do… but you run out of runway quickly as your product becomes more complex. It’s also a non-starter for anything that could end with a lawsuit.&#8221;</p>
</blockquote>



<p class="wp-block-paragraph">And journalism involves enough legal risk and complexity for all of the above to apply to newsrooms too.  </p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/06/18/many-newsrooms-are-not-optimised-for-what-humans-do-best-but-we-have-an-opportunity-to-change-that/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31459</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/reversecentaurdoctorow.png?w=1024" medium="image">
			<media:title type="html">Amazon worker with horse head</media:title>
		</media:content>
	</item>
		<item>
		<title>Caught in a trap: what journalists can learn from systems thinking</title>
		<link>https://onlinejournalismblog.com/2026/06/16/caught-in-a-trap-what-journalists-can-learn-from-systems-thinking/</link>
					<comments>https://onlinejournalismblog.com/2026/06/16/caught-in-a-trap-what-journalists-can-learn-from-systems-thinking/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Tue, 16 Jun 2026 13:03:52 +0000</pubDate>
				<category><![CDATA[data journalism]]></category>
		<category><![CDATA[investigative journalism]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[donella meadows]]></category>
		<category><![CDATA[drift to low performance]]></category>
		<category><![CDATA[escalation]]></category>
		<category><![CDATA[ideas]]></category>
		<category><![CDATA[policy resistance]]></category>
		<category><![CDATA[rule beating]]></category>
		<category><![CDATA[seeking the wrong goal]]></category>
		<category><![CDATA[Shifting the burden to the intervenor]]></category>
		<category><![CDATA[success to the successful]]></category>
		<category><![CDATA[system traps]]></category>
		<category><![CDATA[Thinking in Systems]]></category>
		<category><![CDATA[tragedy of the commons]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=29858</guid>

					<description><![CDATA[One of the most powerful ways to generate original journalism is to look at the systems behind stories — particularly the points where those systems fail. For investigative work, those points are central. Surface-level scandals often stem from deeper systemic problems. So what tools do we have for recognising those patterns? Donella Meadows&#8217;s classic book [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">One of the most powerful ways to generate original journalism is to look at the <strong>systems</strong> behind stories — particularly the points where those systems fail.</p>



<p class="wp-block-paragraph">For investigative work, those points are central. Surface-level scandals often stem from deeper systemic problems. So what tools do we have for recognising those patterns?</p>



<p class="wp-block-paragraph">Donella Meadows&#8217;s classic book <a href="https://www.chelseagreen.com/product/thinking-in-systems/"><em>Thinking in Systems</em></a> offers one: &#8220;<strong>system traps</strong>&#8221; — patterns that explain how systems get stuck, break down, or behave in ways nobody intends. They are &#8220;traps&#8221; because attempts to escape them often backfire.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png"><img loading="lazy" width="1024" height="714" data-attachment-id="31436" data-permalink="https://onlinejournalismblog.com/2026/06/16/caught-in-a-trap-what-journalists-can-learn-from-systems-thinking/system-traps/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png" data-orig-size="1206,842" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="System traps" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png?w=1024" alt="System trap

Journalism examples

Policy resistance

The war on drugs; reforms that fail; missed targets

Overuse leading to shortages; climate change impacts; AI

Tragedy of the commons

Drift to low performance

Normalisation of poor performance or low productivity

Escalation

Arms races; races to the bottom

Success to the successful

Increasing concentration of wealth or resources

Shifting the burden to the intervenor

Subsidies, price fixes and delaying the impact/cost of a policy

Rule beating

Tax avoidance, loopholes

Seeking the wrong goal

Schools focusing on targets over pupil welfare; 
" class="wp-image-31436" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/system-traps.png 1206w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a></figure>



<p class="wp-block-paragraph">In this post I’ll explain each trap, what it looks like in the wild, and how to use it as a lens for story ideas.</p>



<span id="more-29858"></span>



<h2 class="wp-block-heading">1. &#8216;Fixes that fail&#8217;: policy resistance </h2>



<p class="wp-block-paragraph">Why wasn’t the war on drugs ever won? Why do healthcare reforms so often backfire? Why are targets to reduce emissions missed? These are classic <strong>policy resistance</strong> stories, where different parts of a system pull in opposite directions.</p>



<p class="wp-block-paragraph">Policy resistance is perhaps the most visible &#8216;system failure&#8217; — but recognising it means we know what questions to ask, and what to look for.</p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: Reforms that don’t work, or make things worse. Outcomes that don&#8217;t improve. Missed targets, broken promises.</li>



<li><strong>Questions to ask</strong>: Who has a different goal within the system? Is one part of the system cancelling out another? Are there places where the problem is being tackled successfully?</li>
</ul>



<p class="wp-block-paragraph">In environmental reporting, for example, stories regularly emerge from conflicts between policies aimed at reducing emissions and policies aimed at increasing production, consumption or employment.</p>



<p class="wp-block-paragraph">One of the simplest ways to identify and reveal &#8220;<em>fixes that fail</em>&#8221; is the &#8216;<strong>lack of change</strong>&#8216; story: once a policy has been in place for a long enough period, investigate if it seems to be working. Use data to spot where a system is resisting its own policy.</p>



<p class="wp-block-paragraph">Look for <strong>resistance or opposition to change</strong>: The Washington Post, for example, &#8220;<a href="https://web.archive.org/web/20210610141033/https://www.washingtonpost.com/investigations/interactive/2021/police-reform-failure/">examined</a> three historic firsts in policing reforms&#8221; to reveal reforms &#8220;stifled by entrenched cultures, systemic dysfunction, shifts in leadership and swings in public mood&#8221;. </p>



<p class="wp-block-paragraph">That <strong>resistance can be external</strong>, too, such as <a href="https://www.ft.com/content/c4cedc55-f654-4abf-8f1f-5231b3abef20">local communities resisting green energy plans</a> due to concerns over property prices or local wildlife</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/BYSFvq8sr7A?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph"><strong>Anniversaries</strong> can provide useful hooks to review attempts at reform, as in <a href="https://www.propublica.org/article/why-police-reform-stalled-elizabeth-glazer">ProPublica&#8217;s feature</a> on the stalling of policing reform &#8220;More Than Two Years After George Floyd’s Murder Sparked a Movement&#8221;.</p>



<p class="wp-block-paragraph">Another approach is to look for <strong>unintended consequences</strong>. The increase in orphans and <a href="https://en.wikipedia.org/wiki/1980s%E2%80%931990s_Romanian_orphans_phenomenon">orphanages in Romania in the 1980s and 90s</a> is one example (the unintended consequence of a policy whose intention was to increase birth rates); communities <a href="https://www.npr.org/2021/06/17/1006495476/after-50-years-of-the-war-on-drugs-what-good-is-it-doing-for-us">scarred</a> by the war on drugs is another.</p>



<h2 class="wp-block-heading">2. <strong>Tragedy of the commons</strong> </h2>



<p class="wp-block-paragraph">This trap appears when shared resources — <a href="https://gijn.org/stories/investigation-reveals-global-fisheries-already-collapsed/">fisheries</a>, roads, emergency services — are used by individuals in ways that degrade the whole. </p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: Unregulated or poorly monitored access to a shared resource. Growing use without growing responsibility.</li>



<li><strong>Questions to ask</strong>: Who uses the resource? Who monitors it? Are there rules — and do they work?</li>
</ul>



<figure class="wp-block-embed is-type-rich is-provider-embed-handler wp-block-embed-embed-handler"><div class="wp-block-embed__wrapper">
<a href="https://www.solvingforpattern.org/wp-content/uploads/2013/07/Ostrom-design-principles.png"><img src="https://www.solvingforpattern.org/wp-content/uploads/2013/07/Ostrom-design-principles.png" style="max-width:100%;" /></a>
</div><figcaption class="wp-element-caption">Political economist Elinor Ostrom won the Nobel Prize in Economics for her research into how communities successfully manage commons. She identifies eight &#8216;pressure points&#8217;, <a href="https://www.solvingforpattern.org/2013/07/17/making-the-practice-turn-the-active-voice/">adapted here by Howard Silverman</a></figcaption></figure>



<p class="wp-block-paragraph">A potential story on this system failure might draw on political scientist <strong>Elinor Ostrom</strong>&#8216;s eight pressure points to look in the following areas: </p>



<ul class="wp-block-list">
<li>What are the <strong>boundaries</strong>? Who counts as a “user” or stakeholder?</li>



<li>Is the system being <strong>monitored</strong>? If so, by whom?</li>



<li>Who has <strong>decision-making power</strong> — and is it legitimate?</li>



<li>Are rules being <strong>broken or gamed</strong>, and if so, how are they enforced?</li>



<li>What does the <strong>conflict resolution process</strong> look like — or is it missing?</li>



<li>Are <strong>local voices</strong> recognised — or excluded?</li>
</ul>



<p class="wp-block-paragraph"><em>Science</em>&#8216;s <a href="https://gijn.org/stories/investigation-reveals-global-fisheries-already-collapsed/">investigation into fisheries</a>, for example, &#8220;tested how accurate estimates of fish stocks actually are&#8221; (monitoring). Daniel Wizenberg&#8217;s reporting on cruise ship pollution <a href="https://www.theguardian.com/travel/2023/oct/19/europe-ports-bear-brunt-of-cruise-ship-pollution">reports on ports considering bans and restrictions</a> (punishments) and <a href="https://pulitzercenter.org/stories/small-arctic-village-where-tourists-are-replacing-narwhals-spanish">&#8220;a divide among residents about the benefits and drawbacks of tourism&#8221;</a> (the balance of costs and benefits).</p>



<p class="wp-block-paragraph"><a href="https://greentogrey.eu/">Green to Grey: How Europe Is Squandering the Little Nature It Has Left</a> fills a gap in monitoring, identifies the ineffectiveness of non-binding EU goals, and quotes local voices that were ignored. </p>



<p class="wp-block-paragraph">But this trap doesn&#8217;t just apply to environmental resources: a story about the increasing demands being placed on A&amp;E services might similarly focus on boundaries (e.g. patients using A&amp;E instead of other services) and relationships. Similar areas could be looked at in a story about housing stock being rented out to tourists. <a href="https://www.thebureauinvestigates.com/stories/2025-03-13/how-decades-of-factory-farming-paved-the-way-for-todays-superbugs-crisis">Antibiotic overuse</a> is another example. </p>



<p class="wp-block-paragraph">Jason Del Rey&#8217;s <a href="https://www.vox.com/recode/23170900/leaked-amazon-memo-warehouses-hiring-shortage">story</a> on internal Amazon research warning that the company was &#8220;running out of people to hire&#8221; is a tragedy of the commons, and the &#8220;attention economy&#8221; is one way to describe another commons that is being exhausted. </p>



<h2 class="wp-block-heading">3. <strong>Drift to low performance (aka eroding goals/boiling frog syndrome)</strong></h2>



<figure class="wp-block-embed alignright is-type-rich is-provider-embed-handler wp-block-embed-embed-handler"><div class="wp-block-embed__wrapper">
<a href="https://thesystemsthinker.com/wp-content/uploads/2016/02/resolved-by-taking-corrective-action.jpg"><img src="https://thesystemsthinker.com/wp-content/uploads/2016/02/resolved-by-taking-corrective-action.jpg" style="max-width:100%;" /></a>
</div><figcaption class="wp-element-caption"><a href="https://thesystemsthinker.com/drifting-goals-the-boiled-frog-syndrome/">Image from The Systems Thinker</a></figcaption></figure>



<p class="wp-block-paragraph">A typical &#8220;drift to low performance&#8221; trap is a negative feedback loop, where systems slowly get worse, people adjust — and low expectations (&#8216;<a href="https://systemsandus.com/archetypes/eroding-goals/">eroding goals</a>&#8216;) become the new normal.</p>



<p class="wp-block-paragraph">A good example of this being made explicit is the BBC story <em><a href="https://www.bbc.co.uk/news/articles/ckgvl8l5q0xo">Harm at risk of being normalised in maternity care</a></em>. </p>



<p class="wp-block-paragraph">In education, the Bureau of Investigative Journalism&#8217;s <a href="https://schoolsweek.co.uk/investigation-the-broken-special-needs-system/">investigation</a> into the &#8220;broken&#8221; special needs system identified how thresholds had been raised to &#8220;make it more difficult for children to receive support&#8221;, while <a href="https://www.independent.co.uk/news/education/education-news/universities-investigation-increase-grade-inflation-b2157709.html">grade inflation</a> at universities is another symptom of standards being lowered. </p>



<p class="wp-block-paragraph">In environmental journalism, you might look for <a href="https://www.theguardian.com/environment/2022/dec/09/government-to-weaken-water-pollution-goals-river-health">targets being weakened</a>, <a href="https://www.theguardian.com/environment/2025/mar/12/changes-to-bathing-water-status-test-will-deny-rivers-protection-say-critics">barriers to protection</a>, or a <a href="https://www.theguardian.com/business/2025/mar/27/raw-sewage-hours-dumped-water-companies-england-last-year-data?utm_source=chatgpt.com">lack of enforcement</a>. </p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: Declining standards, redefined targets, morale collapse.</li>



<li><strong>Questions to ask</strong>: Have goals or benchmarks been quietly lowered? Are people adapting to failure instead of fixing it? Are regulators turning a blind eye?</li>
</ul>



<h2 class="wp-block-heading">4. <strong>Escalation</strong> (arms races and races to the bottom)</h2>



<p class="wp-block-paragraph"><strong>Escalation traps</strong> form when organisations <a href="https://thesystemsthinker.com/using-escalation-to-change-the-competitive-game/">try to outdo each other</a>, creating a feedback loop as each responds to the other. Examples include arms races, military conflicts, political polarisation, and cat-and-mouse relationships (e.g. spammers versus email filters).</p>



<p class="wp-block-paragraph">&#8216;<strong>Races to the bottom</strong>&#8216; are an escalation trap. Examples include negative <a href="https://hardestquestions.substack.com/p/bogeyman-politics-and-the-race-to">political campaigns</a>, increasing <a href="https://www.ft.com/content/bd7d488a-f40c-40a4-86e2-4eaef2346088">deregulation</a> and <a href="https://www.bbc.co.uk/news/business-56500673">lowering taxes</a>, and <strong>price wars</strong> — where competing companies make products increasingly cheaper to undercut rivals.</p>



<p class="wp-block-paragraph">Escalation of even apparently positive behaviours can also become an escalation trap: Meadows gives the example of morality leading to sanctimoniousness, Puritanism and <strong>intolerance</strong>. Other examples might include healthy eating becoming orthorexia, and healthy scepticism turning into conspiracism.</p>



<p class="wp-block-paragraph">Many investigations into social media algorithms are stories about escalation traps: the Wall Street Journal&#8217;s<em> <a href="https://www.wsj.com/articles/the-facebook-files-11631713039">Facebook Files</a></em>, for example, shows how the company&#8217;s desire for growth led it to ignore warnings that algorithm changes were making &#8220;Facebook, and those who used it, angrier&#8221;, while NRK&#8217;s <a href="https://www.nrk.no/ostfold/tiktok_s-muscle-power_-how-children-are-drawn-into-a-world-of-extreme-exercise-1.16212259">investigation into TikTok</a> showed &#8220;how TikTok&#8217;s algorithms drag children into a world of extreme exercise&#8221;.  Algorithm Watch&#8217;s <a href="https://algorithmwatch.org/en/instagram-algorithm-nudity/">investigation</a> revealed the Instagram algorithms that encourage &#8220;showing skin&#8221;: users themselves are often caught in an arms race too.</p>



<p class="wp-block-paragraph">Political arms races include parties <a href="https://www.fightingknifecrime.london/news-posts/when-fear-meets-facts-the-politics-of-crime-prevention">trying to outdo each other on being &#8220;tough on crime&#8221;</a> or <a href="https://iea.org.uk/blog/this-anti-immigration-arms-race-exemplifies-everything-that%E2%80%99s-wrong-with-politicians">anti-immigration</a>. Stories on the escalation trap might also focus on how <strong>inflationary forces</strong> are playing out, such as tech companies or sports clubs trying to outspend each other on talent.</p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: reactions and counter-reactions. Unhealthy or extreme behaviours being reinforced. Cat-and-mouse games between rule-breakers and enforcers.</li>



<li><strong>Questions to ask</strong>: What concerns are being raised? What’s driving the escalation? Who benefits? Is there a way out? How might &#8216;unilateral disarmament&#8217; be facilitated (e.g. through regulation, agreement or redesign)? </li>
</ul>



<h2 class="wp-block-heading">5. <strong>Success to the successful/competitive exclusion</strong></h2>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/oovsyzB9L-s?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">The rich getting richer and the poor getting poorer are well known examples of the <strong>competitive exclusion</strong> system trap. This is another feedback loop, where &#8216;winners&#8217; in a system have an unfair advantage when it comes to competing for new opportunities or resources, leading to a vicious circle. </p>



<p class="wp-block-paragraph">Cory Doctorow&#8217;s book <a href="https://www.versobooks.com/en-gb/products/3341-enshittification">Enshittification</a> provides an endless list of examples of this system trap from multiple fields, along with the playbook that companies follow (including <strong><a href="https://doctorow.medium.com/https-pluralistic-net-2025-10-23-traveling-salesman-solution-pee-bottles-9f276109d8c1">chickenisation</a></strong> and the <a href="https://doctorow.medium.com/tiktoks-enshittification-bb3f5df91979">giant teddy bear tactic</a>), and the forces that lead to it.</p>



<p class="wp-block-paragraph">Education is an area where this trap shows itself regularly: look for evidence that <a href="https://www.independent.co.uk/news/uk/gap-england-department-for-education-government-data-b2634966.html">the gap between private and state school pupils going to university is widening</a>, for example, how the use of private tutoring <a href="https://inews.co.uk/news/education/gcse-results-face-national-scandal-of-inequality-as-post-covid-boom-in-private-tuition-deepens-division-2567907">&#8220;deepens inequality&#8221;</a>, or the &#8220;<a href="https://www.theguardian.com/education/2023/mar/01/england-poorer-pupils-face-exclusion-from-top-state-schools-study">geographic exclusion</a>&#8221; of poorer children from higher performing schools.</p>



<p class="wp-block-paragraph">Tech is another: in 2015 Reuters <a href="https://www.reuters.com/legal/litigation/amazon-copied-products-rigged-search-results-promote-its-own-brands-documents-2021-10-13/">revealed</a> that Amazon used its access to seller data to design its own private‑label products and &#8220;manipulat[ed] search results to boost its own product lines in India&#8221;, while in 2015 The <em>Wall Street Journal</em> <a href="https://www.wsj.com/articles/inside-the-u-s-antitrust-probe-of-google-1426793274">used FOI to reveal</a> &#8220;evidence that Google’s algorithm was demoting the search results of competing services while placing its own higher on the search results page&#8221;.</p>



<p class="wp-block-paragraph">Focusing on <strong>supplier vulnerability</strong> can help expose where dominant companies may be exploiting their power. <em>The Guardian</em>&#8216;s <a href="https://www.theguardian.com/small-business-network/2015/jun/25/tesco-supermarkets-behaving-badly-suppliers?utm_source=chatgpt.com">coverage</a> of the UK supermarket sector describes multiple examples, while its <a href="https://www.theguardian.com/business/2015/feb/05/supermarket-watchdog-inquiry-tesco-treatment-suppliers">reporting</a> notably focuses on the rules designed to prevent such exploitation, and the regulator&#8217;s (lack of) power to impose effective punishment. </p>



<p class="wp-block-paragraph">Broader scrutiny of <strong>regulation</strong> (complaints, investigations, reports, data), <strong>tax laws</strong> (such as taxes on inheritance but also <a href="https://www.icij.org/investigations/uber-files/uber-tax-havens-dodge-drivers/">the use of &#8220;sweetheart tax deals for companies like Uber</a>&#8220;), and <strong>anti-monopoly laws</strong>— how often they are enforced, the effectiveness (or not) of punishments, how companies avoid those, and the need for new rules — provides another potential basis of reporting. </p>



<p class="wp-block-paragraph">In sport, for example, journalists have <a href="https://www.thebureauinvestigates.com/stories/2023-11-15/roman-abramovichs-hidden-football-deals-during-chelseas-time-at-the-top">investigated</a> how &#8220;Abramovich’s Chelsea may have secretly bypassed Financial Fair Play (FFP) rules governing world football, helping propel the club to the pinnacle of the game, at the expense of its rivals&#8221;. </p>



<figure class="wp-block-embed is-type-rich is-provider-embed-handler wp-block-embed-embed-handler"><div class="wp-block-embed__wrapper">
<a href="https://thesystemsthinker.com/wp-content/uploads/images/volume-4/originally-designed-to-slow-typists.jpg"><img src="https://thesystemsthinker.com/wp-content/uploads/images/volume-4/originally-designed-to-slow-typists.jpg" style="max-width:100%;" /></a>
</div><figcaption class="wp-element-caption">The QWERTY keyboard benefited from first-mover advantage and the costs of learning a new system (the competency trap). <a href="https://thesystemsthinker.com/using-success-to-the-successful-to-avoid-competency-traps/">Image: The Systems Thinker</a></figcaption></figure>



<p class="wp-block-paragraph">Other forces to consider include <a href="https://en.wikipedia.org/wiki/First-mover_advantage">first-mover advantage</a>, the <a href="https://thesystemsthinker.com/using-success-to-the-successful-to-avoid-competency-traps/">“competency trap”</a> (when people don&#8217;t switch to alternatives because they are used to a particular product), and &#8216;<strong>network effects</strong>&#8216;. </p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: Self-reinforcing inequality (e.g. funding that rewards institutions with better resources, <a href="https://parliamentnews.co.uk/nhs-review-based-funding-plan-raises-concerns">reviews</a> or results), whistleblowers, network effects, regulatory effectiveness, lobbying for regulatory change, tax deals and subsidies, &#8216;bundling&#8217; of products</li>



<li><strong>Questions to ask</strong>: Who controls key resources? Are newer players being excluded? What is the experience of those at the bottom of the supply chain? What complaints are being raised with regulators? What reports are produced by regulators internally?</li>
</ul>



<h2 class="wp-block-heading">6. Treating symptoms not causes: <strong>shifting the burden to the intervenor</strong></h2>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png"><img loading="lazy" width="1024" height="541" data-attachment-id="30158" data-permalink="https://onlinejournalismblog.com/2026/06/16/caught-in-a-trap-what-journalists-can-learn-from-systems-thinking/teach-a-man-to-fish/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png" data-orig-size="1894,1002" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;}" data-image-title="teach a man to fish" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=1024" alt="&quot;If you give a man a fish, he is hungry again in an hour. If you teach him to catch a fish, you do him a good turn.&quot; Anne Isabella Thackeray Ritchie" class="wp-image-30158" srcset="https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=1440 1440w, https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png 1894w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption"><a href="https://www.quoteslyfe.com/quote/If-you-give-a-man-a-fish-376894">Proverbs about teaching a man to fish</a> teach the &#8216;trap&#8217; in treating symptoms instead of causes</figcaption></figure>



<p class="wp-block-paragraph">Examples of &#8220;shifting the burden to the intervenor&#8221; include the use of <strong>subsidies</strong> or <strong>price fixes</strong> to prop up failing industries or maintain access to cheap fuel or credit, or <strong>practices that will be damaging in the long term</strong> (e.g. <a href="https://www.theguardian.com/environment/2023/mar/12/scientists-warn-of-phosphogeddon-fertiliser-shortages-loom?_hsenc=p2ANqtz--mM2FgIG-tZzdMjs4DCCrv0EU4E-hmN8D-8Fxmw1WgYsw_AD32jOdUUy_SsYtswoVga1wT">overuse of fertiliser</a>, drilling for more oil) instead of more sustainable options.</p>



<p class="wp-block-paragraph">The <strong>Bureau of Investigative Journalism</strong>&#8216;s <a href="https://www.thebureauinvestigates.com/projects/the-housing-crisis/">investigation into the housing crisis</a>, for example, highlights the problems <a href="https://neweconomics.org/2023/02/government-to-spend-over-46bn-more-subsidising-private-landlords-than-on-its-programme-to-build-affordable-homes">caused by subsidies</a> that ultimately flow to landlords, instead of regulating rents or building affordable homes. And the <strong>Marshall Project</strong> is just one of a number of organisations to have investigated how <a href="https://www.themarshallproject.org/impact/prosecuting-drug-crimes">jails are used as a quick &#8216;fix&#8217;</a> for drug and mental health problems instead of treatment, or <a href="https://www.bbc.co.uk/news/uk-39655259">police being used to deal with public health failures</a>.</p>



<p class="wp-block-paragraph"><strong>Outsourcing</strong> is a <a href="https://www.wearethepractitioners.com/index.php/topics/art-analysis/systems-thinking/Shifting-the-Burden-to-the-Intervenor">common example</a> of shifting the burden, as it often results in public bodies losing the ability to provide services themselves and becoming dependent on private agencies (who may prioritise costs over outcomes). When there are <a href="https://www.theguardian.com/society/2023/apr/21/regulator-to-review-safety-concerns-over-medicines-courier-sciensus-nhs">safety concerns</a> over those private agencies, or they <a href="https://lowdownnhs.info/explainers/50-failures-in-nhs-outsourcing-2013-2019/">fail to deliver</a>, it creates further problems. </p>



<p class="wp-block-paragraph">This system trap is often also called <strong>&#8220;addiction&#8221;</strong>, because the &#8216;trap&#8217; is that these are short term &#8216;fixes&#8217; offering immediate relief but creating a cycle that’s increasingly hard to break without pain or disruption (addiction itself is just <a href="https://durmonski.com/self-improvement/shifting-the-burden/">one example of this trap</a>). </p>



<p class="wp-block-paragraph">The first episode of the FACTA.eu investigation <em><a href="https://facta.eu/it/non-e-unagricoltura-per-piccoli/">Non è un’agricoltura per piccoli</a></em>, for example, quotes owners of small farms who have difficulty accessing subsidies and report &#8220;a widespread feeling of being faced with a trap rather than a support&#8221;. </p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: Band-aid policies, outsourcing, bailouts, pressures in the system spreading to other parts (e.g. from health to police), recurring crises, companies promoting &#8216;quick fixes&#8217; to deeper problems (e.g. AI), &#8220;addiction&#8221; to or dependence on subsidies or emergency measures</li>



<li><strong>Questions to ask</strong>: What deeper problem is being avoided? Who&#8217;s intervening? Are outcomes improving—or simply being masked? What skills or oversight has been lost as a result of outsourcing? What happens when the &#8216;contractor &#8216;intervenor&#8217; fails? Who profits, and who pays when it goes wrong? How effective is regulation of outsourcing, privatisation or services &#8216;filling the gap&#8217; for others? What is being done to tackle the root causes?</li>
</ul>



<h2 class="wp-block-heading">7.<strong> Gaming the system: rule beating</strong></h2>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(2082 / 1330)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="2082" height="1330" alt="How an obscure legal doctrine called qualified immunity protects police accused of excessive force Shielded A REUTERS INVESTIGATION" class="wp-block-jetpack-slideshow_image wp-image-31443" data-id="31443" data-aspect-ratio="2082 / 1330" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png 2082w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png?w=150&amp;h=96 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png?w=300&amp;h=192 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png?w=768&amp;h=491 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png?w=1024&amp;h=654 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png?w=1440&amp;h=920 1440w" sizes="(max-width: 2082px) 100vw, 2082px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Reuters&#8217; investigation focuses on how laws are used to avoid scrutiny</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1230" height="1164" alt="Histogram showing contract spending dropping after a new law is introduced, but then returning to previous levels" class="wp-block-jetpack-slideshow_image wp-image-31442" data-id="31442" data-aspect-ratio="1230 / 1164" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png 1230w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png?w=150&amp;h=142 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png?w=300&amp;h=284 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png?w=768&amp;h=727 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png?w=1024&amp;h=969 1024w" sizes="(max-width: 1230px) 100vw, 1230px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Il Post&#8217;s investigation shows how spending on temporary staff returned to previous levels despite a law intended to reduce it</figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">In this trap people and organisations follow the rules but violate the spirit of the law. The most widely reported example of this is <strong>tax avoidance</strong>: strategies for avoiding tax which are legal but widely perceived to be unethical (illegal strategies are <a href="https://www.dbtandpartners.co.uk/tax-evasion/tax-avoidance-vs-tax-evasion/">called</a> <strong>tax evasion</strong>).</p>



<p class="wp-block-paragraph">Rule beating can be a good subject for <strong>factchecks</strong> or <strong>explainers</strong>, such as Channel 4 News&#8217;s <a href="https://www.channel4.com/news/factcheck/factcheck-why-does-amazon-pay-so-little-tax">factcheck of why Amazon UK pays so little tax</a> and Ethical Consumer&#8217;s <a href="https://www.ethicalconsumer.org/ethical-campaigns-boycotts/amazon-uks-substantial-tax-avoidance">explainer on how much that costs citizens in lost taxes</a>. Reporting can also focus on <strong>reactions</strong> to legal but unethical behaviour from <a href="https://www.theguardian.com/technology/2020/sep/08/amazon-uk-pays-3-more-in-tax-despite-35-rise-in-profits">experts and campaigners</a>, including <a href="https://www.standard.co.uk/business/business-news/calls-for-tax-loopholes-to-be-closed-as-shell-makes-ps1-4bn-more-than-expected-b1078838.html">calls</a> for loopholes to be fixed. </p>



<p class="wp-block-paragraph">Deeper investigations can focus on: </p>



<ul class="wp-block-list">
<li>Uncovering the use of <strong>loopholes</strong> as in Sky News&#8217;s <a href="https://news.sky.com/story/major-uk-recruiters-linked-to-tax-avoidance-schemes-after-workers-hit-with-crippling-hmrc-demands-13326107">investigation</a> into recruitment companies paying workers &#8220;by third-party umbrella companies &#8230; what were technically loans&#8221;.</li>



<li><strong>Using laws to avoid scrutiny or accountability</strong>, such as the doctrine of qualified immunity investigated in Reuters&#8217; series <em><a href="https://www.reuters.com/investigates/section/usa-police-immunity/">Shielded</a></em>, or <a href="https://democracyforsale.substack.com/p/how-lawyers-to-the-super-rich-strangled?utm_source=substack&amp;utm_medium=email">SLAPPs</a>.</li>



<li>Revealing <strong>unusual behaviour</strong>, such as iWatchAfrica <a href="https://iwatchafrica.org/2020/03/how-multinational-tech-companies-exploit-tax-laws-and-shift-profit-a-focus-on-ghana-and-nigeria/">revealing</a> that &#8220;Facebook had not paid any direct taxes in Ghana and Nigeria since it began operations over a decade ago&#8221;</li>



<li>The <strong>scale</strong> of such behaviour (&#8220;<em><a href="https://www.theguardian.com/politics/2022/sep/24/one-in-six-uk-public-procurement-contracts-had-tax-haven-link-study-finds">One in six UK public procurement contracts had tax haven link</a></em>&#8221; is one example; <a href="https://en.wikipedia.org/wiki/LuxLeaks">LuxLeaks</a> is another. Il Post&#8217;s <a href="https://www.ilpost.it/2026/05/13/gettonisti-ospedali/">story on the use of agency staff in hospitals</a> identifies its scale through the use of a different procurement code, but also identifies other tactics used to get around the law)</li>



<li>Unfair <strong>variation</strong>, as in <a href="https://www.nrk.no/norge/store-summer-fra-vindkraft-til-skatteparadis_-_-for-jaevlig-1.17346628">NRK&#8217;s report</a> on foreign-owned wind power companies paying less corporate tax than Norwegian-owned ones. </li>
</ul>



<p class="wp-block-paragraph">These approaches can be adapted to other systems where rule beating can be found: in 2023 the Freedom of the Press Foundation <a href="https://freedom.press/issues/data-broker-loophole-threatens-journalists-and-whistleblowers">reported</a> that &#8220;intelligence and law enforcement agencies have been using data brokers as a loophole to get around the warrant requirement and other legal restrictions on accessing information for their investigations&#8221;, while in online gambling, The Investigative Post <a href="https://www.investigativepost.org/2024/10/20/the-legal-loopholes-in-online-gambling-regulations-political-will-or-political-gridlock/">revealed</a> that &#8220;Some betting apps promote fantasy sports formats to avoid classification as gambling. In states where online gambling is banned—like California—offshore casinos step in.&#8221;</p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: Stagnant or reversed progress, exploited loopholes, odd behaviour at the end of the financial year, unexpected numbers, rules failing to keep up with new technologies, calls for rules to be changed</li>



<li><strong>Questions to ask</strong>: Who benefits from the current rules? Who loses out? Are rules being circumvented? Why does it matter if rules are circumvented? Have rules kept up to date? </li>
</ul>



<h2 class="wp-block-heading">8. S<strong>eeking the wrong goal</strong></h2>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/O4qOE1U9I8o?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;start=63&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div><figcaption class="wp-element-caption">The goal of high GDP is often criticised as creating a perverse incentive for governments</figcaption></figure>



<p class="wp-block-paragraph">The final system trap is where a person or organisation is motivated by the wrong goal, often because of the metrics that are used for rewards or punishment. As a result, what should be the real goal is undermined (&#8220;What gets measured gets managed&#8221;) and <a href="https://en.wikipedia.org/wiki/Goodhart%27s_law">the measure ceases to become meaningful</a>.</p>



<p class="wp-block-paragraph">For example, treatment targets may drive hospitals to discharge patients early, when it is not in the best interests of those patients. Police may misclassify crimes to improve their crime figures, and schools might teach students how to pass exams, rather than educating them. </p>



<p class="wp-block-paragraph">The <a href="https://en.wikipedia.org/wiki/Stafford_Hospital_scandal">Stafford Hospital scandal</a> is one high-profile example: &#8220;appalling&#8221; conditions at the hospital, <a href="https://pressgazette.co.uk/publishers/regional-newspapers/reporters-role-exposing-nhs-scandal-i-came-industry-make-difference/">exposed by local reporter Shaun Lintern</a>, led to an inquiry which <a href="https://web.archive.org/web/20090325080101/http://www.telegraph.co.uk/health/healthnews/5030012/Staffordshire-hospital-scandal-the-hidden-story.html">noted</a> discussions at the hospital&#8217;s board &#8220;were dominated by finance, target and achieving foundation trust status. There is little evidence that poor standards of nursing care were identified and discussed.&#8221; </p>



<p class="wp-block-paragraph">In education journalists have <a href="https://www.theguardian.com/education/2019/oct/11/one-in-10-pupils-removed-from-school-rolls-to-boost-gcse-results">exposed</a> the practice of &#8220;off-rolling&#8221; poor-performing pupils to boost results, while in crime the Tampa Bay Times <a href="https://web.archive.org/web/20200105132725/https://www.tampabay.com/florida-politics/buzz/2020/01/05/how-the-pinellas-sheriffs-office-boosts-its-rape-stats-without-solving-cases/">exposed</a> how police boosted statistics for rape cases despite not having &#8220;identified a suspect, assigned a detective or even confirmed that a crime had occurred&#8221;. </p>



<p class="wp-block-paragraph">Sometimes the measure doesn’t even have to be a target: Tim Harford’s book <a href="https://timharford.com/books/messy/">Messy</a> dedicates a chapter to various examples of incentives and measures skewing behaviour, <a href="https://jimbodonahue.github.io/blog/data/targets/">such as</a> “how the APGAR measure went from a metric for newborn babies to the target that drove (at least in part) the rise of C-sections in the United States.”</p>



<p class="wp-block-paragraph">Identifying how an organisation&#8217;s performance is measured, particularly where it determines funding, can be a key step in identifying the &#8216;wrong goal&#8217; trap in action. </p>



<ul class="wp-block-list">
<li><strong>What to look for</strong>: How organisations&#8217; or individuals&#8217; performance is measured (particularly where it determines funding or promotion). Perverse incentives, metric-driven behaviour, goal distortion, signs of &#8216;gaming&#8217; the system. In data: special/vague categories (which might be used for misclassification), failing to record data, or <a href="https://onlinejournalismblog.com/2020/05/04/how-should-journalists-report-fiddling-the-figures-on-coronavirus-tests/">changing classifications</a>. Board reports that indicate &#8216;wrong&#8217; priorities.</li>



<li><strong>Questions to ask</strong>: Are the chosen goals actually improving outcomes? Is effort being rewarded instead of results? How should goals be changed? </li>
</ul>



<h4 class="wp-block-heading"><em>I&#8217;m on the lookout for more examples of investigations based on system traps &#8211; if you know of one please let me know in the comments or <a href="https://www.linkedin.com/in/paulbradshawuk/">on LinkedIn</a>.</em></h4>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/06/16/caught-in-a-trap-what-journalists-can-learn-from-systems-thinking/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">29858</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2025/07/teach-a-man-to-fish.png?w=1024" medium="image">
			<media:title type="html">&#034;If you give a man a fish, he is hungry again in an hour. If you teach him to catch a fish, you do him a good turn.&#034; Anne Isabella Thackeray Ritchie</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/shielded.png" medium="image">
			<media:title type="html">How an obscure legal doctrine called qualified immunity protects police accused of excessive force Shielded A REUTERS INVESTIGATION</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/la-spesa-per-i-gettonisti.png" medium="image">
			<media:title type="html">Histogram showing contract spending dropping after a new law is introduced, but then returning to previous levels</media:title>
		</media:content>
	</item>
		<item>
		<title>How to: generate hundreds of maps by combining QGIS with Python (code included!)</title>
		<link>https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/</link>
					<comments>https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Mon, 08 Jun 2026 10:44:37 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[data journalism]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[automation]]></category>
		<category><![CDATA[mapping]]></category>
		<category><![CDATA[parameterisation]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[QGIS]]></category>
		<category><![CDATA[vibe coding]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31502</guid>

					<description><![CDATA[At this year&#8217;s Dataharvest I delivered a workshop on using Python in QGIS to automate the process of exporting maps for multiple locations. Here&#8217;s how to do it (you can find a GitHub repository with materials and links here). Making a map for a story is cool —&#160;but what if you could make a map [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><strong><em>At this year&#8217;s Dataharvest I delivered a workshop on using Python in QGIS to automate the process of exporting maps for multiple locations. Here&#8217;s how to do it (you can <a href="https://github.com/paulbradshaw/QGIS_param/tree/main">find a GitHub repository with materials and links here</a>).</em></strong></p>



<p class="wp-block-paragraph">Making a map for a story is cool —&nbsp;but what if you could make a map for every reader? Or if you&#8217;re working on a project involving teams in different regions or countries, what if you could give each one of those teams a map centred on their own patch?</p>



<p class="wp-block-paragraph">Normally you would have to manually move the map to centre it on a key city, and then export an image. Then do it again and again and again for every area. </p>



<p class="wp-block-paragraph">Luckily, QGIS has the ability to run code. And this is a great excuse to start using it.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png"><img loading="lazy" width="1024" height="574" data-attachment-id="31538" data-permalink="https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/qgismap/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png" data-orig-size="2298,1290" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="qgismap" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=1024" alt="" class="wp-image-31538" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=2048 2048w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=1440 1440w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">By organising the layers on the left you can put shapes such as flood defences over a base OpenStreetMap layer. You can also change the scale in the box underneath the map</figcaption></figure>



<span id="more-31502"></span>



<h2 class="wp-block-heading">Start by creating your base map</h2>



<p class="wp-block-paragraph">Before getting into the code, however, you need a map for the code to move around</p>



<p class="wp-block-paragraph">Create a new project in QGIS and import any shape file(s) that you want to show on your map, along with a base map underneath that. I&#8217;ve <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/01_make_a_map.md">created a tutorial here</a> which imports flood defences and then adds an OpenStreetMap base layer underneath those.</p>



<p class="wp-block-paragraph">Tweak the design of the map until you&#8217;re happy, and then zoom into one of the locations on your map (for example, a key city). We need to use this to create a test &#8216;area map&#8217; so we have an idea of the scale and proportions we can use when we automate things.</p>



<p class="wp-block-paragraph">Once zoomed in, create an image layout by selecting <strong><em>Project &gt; New Print Layout…</em></strong>. I&#8217;ve <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/02_layout.md">created a tutorial for that, too, which you can follow here</a>. This allows you to identify what scale to use when you start to add code. </p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png"><img loading="lazy" width="1024" height="610" data-attachment-id="31537" data-permalink="https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/layoutqgis/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png" data-orig-size="2258,1346" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="layoutqgis" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=1024" alt="" class="wp-image-31537" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=2048 2048w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=1440 1440w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">When zoomed in to Cheltenham, the Layout view allows us to round down to a scale of 25000:1</figcaption></figure>



<p class="wp-block-paragraph"><em>PS: you can make the map entirely in code itself without doing any of this, but I find it quicker to work with the map on screen first so the results can be checked and tweaked more quickly</em>.</p>



<h2 class="wp-block-heading">Download the template Python script and open in QGIS</h2>



<p class="wp-block-paragraph">I&#8217;ve written a Python script that you can import into QGIS to generate multiple images from our map. You can <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/qgis_take_pix.py">see the code here</a>. To download it, click the three dots in the upper right corner and select <strong>Download</strong>. <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/03_qgis_python.md">I&#8217;ve created a tutorial here which walks through the steps</a>.</p>



<p class="wp-block-paragraph">The code does the following:</p>



<ol class="wp-block-list">
<li>Specifies a folder to store all the images</li>



<li>Stores a list of lat-longs</li>



<li>Creates a transformer to convert those to the projection used by your shape files (the coordinate reference system, or CRS)</li>



<li>Sets a map size and scale</li>



<li>Loops through each location in the list, zooms to that scale, and adds a label with the location and scale</li>



<li>Sets a filename that includes the name of the location</li>



<li>Exports an image with that name (for each location in the list)</li>
</ol>



<p class="wp-block-paragraph">Download it to the same folder as your QGIS project and open it in QGIS by going to&nbsp;<strong>Plugins &gt; Python console</strong> as outlined in <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/03_qgis_python.md">the tutorial</a>.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png"><img loading="lazy" width="960" height="540" data-attachment-id="31509" data-permalink="https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/dataharvest-qgis-parameterisation/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png" data-orig-size="960,540" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="Dataharvest QGIS parameterisation" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png?w=960" alt="" class="wp-image-31509" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png?w=768 768w" sizes="auto, (max-width: 960px) 100vw, 960px" /></a><figcaption class="wp-element-caption">To open the Python file you&#8217;ll need to click the File Editor button in the Python console and then the &#8216;open file&#8217; button to open the script</figcaption></figure>



<h2 class="wp-block-heading">Run the script — and customise it</h2>



<p class="wp-block-paragraph">Once you&#8217;ve downloaded the script and opened it in QGIS, you can check if it works. </p>



<p class="wp-block-paragraph">Click the green &#8216;Run&#8217; button to run the script. As it runs the Console area on the left of the Python Console will display any <code>print</code> messages that are in the script to make it easier to track where the script is (like <code>running script</code> and <code>imported libraries</code>).</p>



<p class="wp-block-paragraph">The two key things to keep an eye on are whether it exports to the right location, and whether the list of locations works.</p>



<p class="wp-block-paragraph">The code includes a line which sets a target directory on your computer:</p>



<p class="wp-block-paragraph"><code>out_dir = os.path.join(os.path.expanduser("~"), "Downloads", "qgis_images")</code></p>



<p class="wp-block-paragraph">If it works, you should find the images in your Downloads folder, in a folder called qgis_images. </p>



<p class="wp-block-paragraph">But it might not work on your computer. If it doesn&#8217;t &#8211; if you get an error &#8211; or if you just want to store the images somewhere else, you can do that by <strong>uncommenting</strong> the following line of code (i.e. removing the hash character) and changing it to describe a path on <em>your</em> computer:</p>



<p class="wp-block-paragraph"># <code>out_prefix = "/Users/paul/Downloads/testqgis/images/"</code></p>



<p class="wp-block-paragraph">You can edit the code directly in the editor window in QGIS (the part of the Python Console where you can see the code) &#8211; make sure that you save it once you&#8217;ve made any changes. </p>



<p class="wp-block-paragraph">The list of locations is in this block of code:</p>


<div class="wp-block-code">
	<div class="cm-editor">
		<div class="cm-scroller">
			
<pre>
<code><div class="cm-line">listofdicstocsv = [</div><div class="cm-line">    {&apos;lat&apos;: 50.844441271809465, &apos;long&apos;: -0.2985307978506628, &apos;officialname&apos;: &apos;Adur District Council&apos;}, </div><div class="cm-line">    {&apos;lat&apos;: 54.71108952940033, &apos;long&apos;: -3.2472347297148736, &apos;officialname&apos;: &apos;Allerdale Borough Council&apos;}</div><div class="cm-line">]</div></code></pre>
		</div>
	</div>
</div>


<p class="wp-block-paragraph">This should work with the flood defences data, but if you are using a different map you will need different lat-longs and place names. You can edit these manually for now with your own test lat-longs and names if needed.</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png"><img loading="lazy" width="1024" height="623" data-attachment-id="31540" data-permalink="https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/qgisconsole/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png" data-orig-size="1958,1192" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="qgisconsole" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=1024" alt="" class="wp-image-31540" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=1024 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=1440 1440w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png 1958w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></a><figcaption class="wp-element-caption">The console on the left will show any print messages generate by the script on the right as it runs, as well as any errors</figcaption></figure>



<h2 class="wp-block-heading">Expand the list</h2>



<p class="wp-block-paragraph">Once you&#8217;ve got it working with two locations, it&#8217;s time to update the code so that it uses a <em>full</em> list of locations. Automating a process like this is called <strong>parameterisation</strong>, because it involves generating multiple outputs based on one or more parameters (in this case the lat and long and name).</p>



<p class="wp-block-paragraph">In the tutorial we are using a list of major cities in the UK. You might need a similar list for another country, or it might be centre points of different regions, or the locations of key events.</p>



<p class="wp-block-paragraph">Whatever the locations are, you&#8217;ll need to convert that list into a Python list of <strong>dictionaries</strong>.</p>



<p class="wp-block-paragraph">A dictionary looks like this:</p>



<p class="wp-block-paragraph"><code>{'lat': 50.844441271809465, 'long': -0.2985307978506628, 'officialname': 'Adur District Council'}</code></p>



<p class="wp-block-paragraph">The curly brackets are the giveaway: those mark the start and end of a dictionary. Inside them is a series of pairs, which are equivalent to a label (lat, long, officialname), or <strong>key</strong>, followed by a <strong>value</strong> (50.8, -0.29, &#8216;Adur District Council). Each of these pairs (called <strong>key-value pairs</strong>) is separated by a comma.</p>



<p class="wp-block-paragraph">A dictionary is like a row in a table: the keys are the column headings and stay the same each time, while the values are different in each row. So when you have a list of dictionaries, it&#8217;s basically a list of rows. In other words, it is just another way of storing a table. </p>



<p class="wp-block-paragraph">So we are going to need a table of locations (the ones we want to create maps for). We can then convert that table to a list of dictionaries (a list of table rows). </p>



<p class="wp-block-paragraph">You might be tempted to paste your table or attach a CSV in a prompt and ask an AI tool to convert that into a list of Python dictionaries. But <a href="https://onlinejournalismblog.com/2025/06/19/how-to-reduce-the-environmental-impact-of-using-ai/">it&#8217;s important when using AI to assess whether the task justifies the environmental impact</a> and consider non-AI alternatives that achieve the same results. </p>



<p class="wp-block-paragraph">For example, the website&nbsp;<a href="https://csvjson.com/csv2json">CSV2JSON</a>&nbsp;will perform this task (JSON is just a list of dictionaries). It will take a CSV file and provide it as a list of dictionaries that you can copy and paste into your Python script. I&#8217;ve also <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/optionalfiles/turn_table_into_list.md">written a tutorial on other options</a> to achieve the same results with a spreadsheet formula.</p>



<p class="wp-block-paragraph">Once you&#8217;ve got your list of dictionaries, you can replace the list in the template Python script that starts and ends with square brackets here:</p>


<div class="wp-block-code">
	<div class="cm-editor">
		<div class="cm-scroller">
			
<pre>
<code><div class="cm-line">[</div><div class="cm-line">    {&apos;lat&apos;: 50.844441271809465, &apos;long&apos;: -0.2985307978506628, &apos;officialname&apos;: &apos;Adur District Council&apos;}, </div><div class="cm-line">    {&apos;lat&apos;: 54.71108952940033, &apos;long&apos;: -3.2472347297148736, &apos;officialname&apos;: &apos;Allerdale Borough Council&apos;}</div><div class="cm-line">]</div></code></pre>
		</div>
	</div>
</div>


<p class="wp-block-paragraph">Note that the <strong>list</strong> starts and ends with square brackets, and each <strong>dictionary</strong> in that list is separated by a comma.</p>



<p class="wp-block-paragraph">Once your list is pasted, you&#8217;ll probably have a big chunk of code (or a very long line of code) with all those locations.</p>



<h2 class="wp-block-heading">Testing the expanded list</h2>



<p class="wp-block-paragraph">You can now test the code works with more locations by saving the updated code and running it again. </p>



<p class="wp-block-paragraph">But you might not want to export hundreds of images straight away, as this will take a while. </p>



<p class="wp-block-paragraph">Instead, you can try a <strong>slice</strong> of the list by adding square brackets when you run your loop to specify a limited range of items, like this:</p>



<p class="wp-block-paragraph"><code>for rec in listofdicstocsv[10:20]:</code></p>



<p class="wp-block-paragraph">That <code>[10:20]</code> means just loop through items at positions 10 to 19 (it stops just before the 20th item) in the list <code>listofdictstocsv</code>. You can obviously change those numbers to specify a different range.</p>



<p class="wp-block-paragraph">Once it runs fine with the range specified, you can then look at the results and decide whether you want to export all the images, or tweak your code further to a different scale — or multiple scales&#8230;</p>



<h2 class="wp-block-heading">Exporting more than one scale </h2>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(625 / 442)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="441" alt="Map of Westminster's flood defences" class="wp-block-jetpack-slideshow_image wp-image-31550" data-id="31550" data-aspect-ratio="625 / 442" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">For much closer maps a 25000:1 scale might be best</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="441" alt="Map of Belfast's flood defences" class="wp-block-jetpack-slideshow_image wp-image-31545" data-id="31545" data-aspect-ratio="625 / 442" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Some cities might work best at 75000:1</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="441" alt="Map of Derby's flood defences" class="wp-block-jetpack-slideshow_image wp-image-31547" data-id="31547" data-aspect-ratio="625 / 442" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Some cities might work best at 120000:1</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="441" alt="Map of Wolverhampton's flood defences" class="wp-block-jetpack-slideshow_image wp-image-31548" data-id="31548" data-aspect-ratio="625 / 442" src="https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Larger conurbations might work best at 175000:1 or higher</figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">As you export multiple images you might notice that not all locations are the same. Some locations are self-contained and suit a smaller scale (e.g. more rural cities or a London borough), while others are large conurbations (e.g. Birmingham) that need a larger scale to show all the different parts. </p>



<p class="wp-block-paragraph">Now we need to adapt the code to export at more than one scale. </p>



<p class="wp-block-paragraph">You can <a href="https://github.com/paulbradshaw/QGIS_param/blob/main/qgis_take_pix_twoSize.py">see an example of this adaptation in a second Python script</a>. This adds code that does the following:</p>



<ol class="wp-block-list">
<li>Creates a list of scales to use</li>



<li>Adds an extra loop inside the one that loops through each location, which loops through each scale</li>



<li>Zooms to that location at each scale, and exports an image for that scale (with an appropriate label and filename)</li>
</ol>



<p class="wp-block-paragraph">The key block of code to adapt is this one:</p>


<div class="wp-block-code">
	<div class="cm-editor">
		<div class="cm-scroller">
			
<pre>
<code><div class="cm-line">MAP_SCALES = [</div><div class="cm-line">    175000,   # original scale</div><div class="cm-line">    120000    # zoomed-in version</div><div class="cm-line">]</div></code></pre>
		</div>
	</div>
</div>


<p class="wp-block-paragraph">You can replace those values with your own preferred scales, and it should work. </p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-4-3 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/HfOlnNp3v1E?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">But what about other adaptations — and trouble-shooting? This might be where you turn to AI&#8230;</p>



<h2 class="wp-block-heading">Using AI to adapt QGIS Python code further</h2>



<p class="wp-block-paragraph">Large language models (LLMs) are particularly well suited to helping with coding and coding problems — but that comes with the risk of <strong>deskilling</strong>, or denying you the opportunity to learn. To manage that risk, try a prompt design which begins with this template:</p>


<div class="wp-block-code">
	<div class="cm-editor">
		<div class="cm-scroller">
			
<pre>
<code><div class="cm-line">You are my mentor, a data journalist with over a decade&apos;s experience in the field. </div><div class="cm-line">You have advanced statistical knowledge as well as a healthy scepticism when dealing with both data and human sources. </div><div class="cm-line">I will ask you for help with some coding - you are happy to guide me, but you don&apos;t want me to become deskilled and too reliant on you. </div><div class="cm-line">Your advice will always be designed to force me to think for myself, learn new skills and concepts, and practise those.</div><div class="cm-line"></div><div class="cm-line">When you provide code add comments explaining each step as if to a person with no coding experience. </div><div class="cm-line">I prefer simple code that requires multiple lines to complex code in one or two lines.</div><div class="cm-line"></div><div class="cm-line">Flag any assumptions you are making about the data or my question. </div><div class="cm-line">Warn me if the question does not contain enough information to answer accurately.</div><div class="cm-line">Warn me if the result could be misleading without additional context.</div></code></pre>
		</div>
	</div>
</div>


<p class="wp-block-paragraph">Responses using this template are likely to help you better understand the code that is being produced, as well as identifying things to consider (such as the risks of mixing layers with different projections) and pushing back if you&#8217;re leaving important information out. </p>



<p class="wp-block-paragraph">To adapt the Python code, you might then attach the script with a prompt like this:</p>



<p class="wp-block-paragraph"><code>Attached is some Python that generates images for maps centred at hundreds of locations. Adapt this so that it now generates those images at two different scales. Comment the code to highlight the section of code where it does this, so that I can tweak it and try different zoom levels.</code> </p>



<p class="wp-block-paragraph">Other ways to address deskilling include&nbsp;<a href="https://onlinejournalismblog.com/2026/02/02/parallel-prompting-another-way-to-avoiding-deskilling-with-ai/">parallel prompting</a>,&nbsp;<a href="https://onlinejournalismblog.com/2025/12/02/journey-prompts-and-destination-prompts-how-to-avoid-becoming-deskilled-when-using-ai/">journey prompts</a>&nbsp;and&nbsp;<a href="https://onlinejournalismblog.com/2026/01/13/how-to-stop-ai-making-you-stupid-hybrid-destination-journey-prompting/">destination-journey prompting</a>.</p>



<p class="wp-block-paragraph"><em>You can find the slides from my Dataharvest workshop, including Spanish, Italian, Portuguese, French and Dutch versions, along with walkthroughs, materials and useful links, in <a href="https://github.com/paulbradshaw/QGIS_param/tree/main">this GitHub repo</a>.<br></em></p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/06/08/how-to-generate-hundreds-of-maps-by-combining-qgis-with-python-code-included/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31502</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgismap.png?w=1024" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/layoutqgis.png?w=1024" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/dataharvest-qgis-parameterisation.png?w=960" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/qgisconsole.png?w=1024" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/westminster_25000.png?w=625" medium="image">
			<media:title type="html">Map of Westminster&#039;s flood defences</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/bradford_75000.png?w=625" medium="image">
			<media:title type="html">Map of Belfast&#039;s flood defences</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/derby_120000.png?w=625" medium="image">
			<media:title type="html">Map of Derby&#039;s flood defences</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/06/wolverhampton_175000.png?w=625" medium="image">
			<media:title type="html">Map of Wolverhampton&#039;s flood defences</media:title>
		</media:content>
	</item>
		<item>
		<title>FAQ: AI, misinformation and journalism</title>
		<link>https://onlinejournalismblog.com/2026/05/23/faq-ai-misinformation-and-journalism/</link>
					<comments>https://onlinejournalismblog.com/2026/05/23/faq-ai-misinformation-and-journalism/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Sat, 23 May 2026 08:03:59 +0000</pubDate>
				<category><![CDATA[online journalism]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31415</guid>

					<description><![CDATA[In this latest post in the FAQ series, I am sharing some responses to a radio interview about AI&#8217;s impact on journalism. Q: Is the continuous growth of AI-generated content online a danger for journalism? It is certainly a problem yes, in three ways: it makes reporting harder, it makes it harder to support journalism financially, [&#8230;]]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph"><em>In this latest post in <a href="https://onlinejournalismblog.com/category/faq/">the FAQ series</a>, I am sharing some responses to a radio interview about AI&#8217;s impact on journalism.</em></p>



<h2 class="wp-block-heading">Q: Is the continuous growth of AI-generated content online a danger for journalism? </h2>



<p class="wp-block-paragraph">It is certainly a <em>problem</em> yes, in three ways: it makes reporting harder, it makes it harder to support journalism financially, and it makes it harder for audiences to trust your reporting. </p>



<span id="more-31415"></span>



<ol class="wp-block-list"></ol>



<p class="wp-block-paragraph">It makes reporting harder because a journalist has to deal not only with a lot more misinformation but also a lot more information full stop.</p>



<p class="wp-block-paragraph">It makes the business model more difficult because there&#8217;s more competition from AI-generated slop.</p>



<p class="wp-block-paragraph">And that&#8217;s compounded by the fact that audiences are less likely to trust your reporting — either because they might dismiss it as AI generated, or because they might have already begun to believe AI misinformation that your reporting contradicts.</p>



<h2 class="wp-block-heading">Q: Is the industry equipped to handle it?</h2>



<p class="wp-block-paragraph">I don&#8217;t think any industry is equipped to handle how AI is changing the information environment, and I think it will take some time for society to adjust to a world of unlimited cheap and unreliable information.</p>



<p class="wp-block-paragraph">But journalism is probably in a better position than most industries to handle that change because it has rules and workflows that a lot of other industries do not. </p>



<p class="wp-block-paragraph">For example, journalists are trained to research both sides of the story and seek a right of reply, and to find more than one source for key facts. A high value is put on speaking to real people and capturing raw footage of events. Reporters are encouraged to be sceptical and apply verification and factchecking skills, and their stories go through editors who are supposed to look at the story dispassionately. These are not qualities that all industries have.</p>



<p class="wp-block-paragraph">Also, the industry has been exploring applications of artificial intelligence for over a decade, so there&#8217;s technical knowledge there too which many organisations will lack.</p>



<p class="wp-block-paragraph">The biggest challenge is the collapse of the financial basis for journalism and I think there will have to come a point where more public or charitable funding is used to support it. </p>



<h2 class="wp-block-heading">Q: When AI can fabricate convincing news content at scale, who is ultimately responsible for the integrity of what we read and watch online? </h2>



<p class="wp-block-paragraph">Primarily it&#8217;s whoever created that content. When the printing press came along someone might have asked who is responsible for these books that are being published? And of course we know the answer now is the authors and the publishers. </p>



<p class="wp-block-paragraph">But technology companies are also responsible, because AI is not a neutral tool like a printing press. </p>



<p class="wp-block-paragraph">ChatGPT is a collection of recipes for predicting text, and so OpenAI, which owns ChatGPT, has responsibility for the way it has written those recipes because it can change them. It has changed them, and does change them, regularly. </p>



<p class="wp-block-paragraph">If you try to create racist or other offensive content in ChatGPT or Gemini or Claude, for example, it will generally refuse to do so. If you create images in Gemini a watermark is added, so it can be identified as AI generated. If you ask Claude for medical advice it will recommend consulting a health professional. These are editorial choices by the technology company.</p>



<p class="wp-block-paragraph">So while individual crimes involving fabrication might be tackled by identifying and charging the creator, the ultimate responsibility for any systemic social harm will lie with technology companies who haven&#8217;t considered key risks and taken steps to prevent that. And there will need to be regulation and enforcement to ensure that happens.</p>



<h2 class="wp-block-heading">Q: What does this mean for the public&#8217;s trust in journalism?</h2>



<p class="wp-block-paragraph">We have always had to earn that trust, and that has become more and more important as competition for the public&#8217;s attention has increased. </p>



<p class="wp-block-paragraph">Part of that trust is about building understanding: not only explaining how we actually do journalism, and why we do it that way, but also listening to our audiences to understand what they need and why. </p>



<p class="wp-block-paragraph">In order for the public to trust journalism they have to see how it is distinct from other sources of power, and holds that power to account. They have to see how it is on their side, rather than looking down on them.</p>



<h2 class="wp-block-heading">Q: When fabricated content goes viral, what happens to people&#8217;s ability to make informed decisions?</h2>



<p class="wp-block-paragraph">I think we already had a problem with people being able to make informed decisions, as it became easier to find information that confirmed a person&#8217;s existing beliefs rather than information that challenged those.</p>



<p class="wp-block-paragraph">Ultimately the more misinformation there is, the more likely it is that people are either making decisions based on flawed information, or unable to make decisions because they can&#8217;t be 100% sure a certain piece of information is true.</p>



<p class="wp-block-paragraph">I think we are going to learn some quite painful lessons in the next few years because most people think they won&#8217;t fall for this stuff, and yet most people do. I am an intelligent person who trains people in verification and factchecking, and I have retweeted information that is not true.</p>



<p class="wp-block-paragraph">There are two key things that we need to remember: firstly, that you will fall for misinformation at some point, and probably already have, that is certain. </p>



<p class="wp-block-paragraph">Secondly, slow down and look for counter-evidence.</p>



<h2 class="wp-block-heading">Q: For an ordinary person watching the news or scrolling their feed, what practical steps can they take right now to check the veracity of what they are seeing?</h2>



<p class="wp-block-paragraph">There&#8217;s a great piece of research in psychology that finds when we look at an image of a crowd of faces, we pay attention to the angry faces more: our eyes spend more time on those.</p>



<p class="wp-block-paragraph">Our brains are designed to pay more attention to threats, and social media algorithms have learned this, so they prioritise material that triggers anger and fear because we spend more time looking at it, and are more likely to react to it. </p>



<p class="wp-block-paragraph">So the first practical step people can take is to slow down. When we come across new information two parts of our brain processes it: the first part is fast and instinctive. It makes a decision whether to pay attention and what to do with the information. That&#8217;s the part that shares or likes or comments on a chat message or video. </p>



<p class="wp-block-paragraph">Don&#8217;t share. Don&#8217;t like. Don&#8217;t react. </p>



<p class="wp-block-paragraph">The key thing is to let the information move through to a second phase, the part of the brain that actually processes that information rationally.</p>



<p class="wp-block-paragraph">Now, ask yourself if that update is triggering some core emotion. Is it confirming something you believed? Is it making you angry or afraid? Why? The short answer is: because it keeps you on that site or app for longer. But is it true?</p>



<p class="wp-block-paragraph">The main way to check veracity isn&#8217;t to look for clues that something is fake. An image or video might be real but from a different time or place. Something might be factually true but misleading because it&#8217;s missing context. Trusted friends and authority figures will share this stuff, so don&#8217;t use that as a primary signal.</p>



<p class="wp-block-paragraph">The main way to check veracity is to look for evidence that challenges it. </p>



<p class="wp-block-paragraph">For example you can use a website called TinEye to see where an image has appeared before (you can also freeze frame videos and take screenshots). You can put the information into Google with the word &#8216;hoax&#8217; or &#8216;factcheck&#8217; to see if it&#8217;s already been debunked. </p>



<p class="wp-block-paragraph">You can try to identify the original source, rather than whoever passed it on to you. That might give you clues about their independence, expertise or agenda.</p>



<h2 class="wp-block-heading">Q: And people can use AI tools to verify as well?</h2>



<p class="wp-block-paragraph">Using AI to verify material is very dangerous. We&#8217;ve seen a big increase in the last few years of people using AI tools to check if an image or a video is real, and very often they get it wrong.</p>



<p class="wp-block-paragraph">Part of the reason is that people ask the wrong question. If you ask AI &#8220;is this fake?&#8221; then you have to remember tools like ChatGPT and Grok are sycophantic, they are designed to please, so they will say &#8220;yes&#8221; even if they&#8217;re not certain. They are also overconfident.</p>



<p class="wp-block-paragraph">The key thing with AI in general is to always ask a neutral question. Ask it what evidence there might be to help you identify whether something is true or false. Ask it what methods you can use. </p>



<p class="wp-block-paragraph">AI is best used to challenge you and open up other ideas and techniques. Don&#8217;t use it to confirm things because it will just tell you what you want to hear.</p>



<p class="wp-block-paragraph">Never rely on AI&#8217;s answer as the end of a process &#8211; it can only ever point you to next steps. Remember that it is not a factual tool, it is a language prediction tool. So it can predict the language of a piece of advice around verifying something, but it cannot <em>know</em> if something is real or fake, because that&#8217;s not what AI is designed to do.  </p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/05/23/faq-ai-misinformation-and-journalism/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31415</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>
	</item>
		<item>
		<title>Words as data: how data journalists tell stories about documents and text</title>
		<link>https://onlinejournalismblog.com/2026/05/14/how-to-tell-stories-about-documents-and-text-using-data-journalism/</link>
					<comments>https://onlinejournalismblog.com/2026/05/14/how-to-tell-stories-about-documents-and-text-using-data-journalism/#respond</comments>
		
		<dc:creator><![CDATA[Paul Bradshaw]]></dc:creator>
		<pubDate>Thu, 14 May 2026 12:29:00 +0000</pubDate>
				<category><![CDATA[data journalism]]></category>
		<category><![CDATA[online journalism]]></category>
		<category><![CDATA[7 angles]]></category>
		<category><![CDATA[7 angles of data journalism]]></category>
		<category><![CDATA[change]]></category>
		<category><![CDATA[documents]]></category>
		<category><![CDATA[economist]]></category>
		<category><![CDATA[exploratory]]></category>
		<category><![CDATA[FT]]></category>
		<category><![CDATA[Guardian]]></category>
		<category><![CDATA[John Burn-Murdoch]]></category>
		<category><![CDATA[leads]]></category>
		<category><![CDATA[press association]]></category>
		<category><![CDATA[Quartz]]></category>
		<category><![CDATA[Ranking]]></category>
		<category><![CDATA[relationships]]></category>
		<category><![CDATA[scale]]></category>
		<category><![CDATA[speeches]]></category>
		<category><![CDATA[TEXT]]></category>
		<category><![CDATA[The Outlier]]></category>
		<category><![CDATA[The Pudding]]></category>
		<category><![CDATA[topic modelling]]></category>
		<category><![CDATA[USA Today]]></category>
		<category><![CDATA[variation]]></category>
		<category><![CDATA[Washington Post]]></category>
		<guid isPermaLink="false">http://onlinejournalismblog.com/?p=31215</guid>

					<description><![CDATA[Documents and other collections of text can be goldmines for data journalism — if you know how to approach them as data. Here are some techniques and inspiration for your next data project. From stories about political speech and song lyrics, to street names and social media chatter, data journalists now have a wide range [&#8230;]]]></description>
										<content:encoded><![CDATA[
<h3 class="wp-block-heading"><em>Documents and other collections of text can be goldmines for data journalism — if you know how to approach them as data. Here are some techniques and inspiration for your next data project.</em></h3>



<p class="wp-block-paragraph">From stories about <a href="https://www.theguardian.com/politics/ng-interactive/2026/feb/25/how-rightwing-rhetoric-has-risen-sharply-in-the-uk-parliament-an-exclusive-visual-analysis">political speech</a> and <a href="https://www.lexicodosamba.com.br/en/">song lyrics</a>, to <a href="https://www.rspb.org.uk/whats-happening/news/street-names-evoke-nature-as-the-real-thing-vanishes">street names</a> and <a href="https://www.theguardian.com/uk/interactive/2011/dec/07/london-riots-twitter">social media chatter</a>, data journalists now have a wide range of examples of text-as-data to draw inspiration and guidance from, while tools such as <a href="https://journaliststudio.google.com/pinpoint/">Pinpoint</a> and <a href="https://notebooklm.google.com/">NotebookLM</a> are making text analysis easier than ever.</p>



<p class="wp-block-paragraph">I <a href="https://github.com/paulbradshaw/dealingwithdocuments/blob/master/textdatasets/djstories_usingtext.csv">compiled a list of over 200 pieces of data journalism</a> where text or documents were used as sources. Quantification techniques ranged from <a href="https://www.shropshirestar.com/news/uk-news/2020/01/30/from-nine-mentions-a-year-to-9000-how-mps-caught-the-brexit-bug/">counting the frequency of a single word</a> and <a href="https://colinmorris.github.io/blog/size-of-things">using Google&#8217;s ngram viewer</a>, to <a href="https://www.turing.ac.uk/news/seven-ten-premier-league-footballers-face-twitter-abuse">machine learning</a> and <a href="https://www.spiegel.de/politik/bundestagswahl-2025-datenanalyse-der-wahlkampfreden-der-kanzlerkandidaten-a-92e1b537-5f39-44ba-8f39-1f2433a5c647">topic modelling</a>. </p>



<p class="wp-block-paragraph">Looking at those articles it&#8217;s clear that, once quantified, journalists tell the same stories about text as any other piece of data: using the <a href="https://onlinejournalismblog.com/2020/08/11/here-are-the-7-types-of-stories-most-often-found-in-data/">seven most common angles</a>. </p>



<p class="wp-block-paragraph">But <em>how</em> those angles are used — and how <em>often</em> — is where it gets interesting&#8230;</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png"><img loading="lazy" width="625" height="373" data-attachment-id="31413" data-permalink="https://onlinejournalismblog.com/2026/05/14/how-to-tell-stories-about-documents-and-text-using-data-journalism/7-angles-for-data-stories-text_docs-1/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png" data-orig-size="959,573" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="7 angles for data stories &amp;#8211; TEXT_DOCS (1)" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png?w=625" alt="7 common angles for data stories: text and documents 
Scale: how often words/phrases are used
Change: how language has changed
Ranking: the most/least common words/phrases
Variation: e.g. in relation to gender, ethnicity, ideology etc.
Exploration: journeys through multiple angles; interactives
Relationships: correlations, similarities and connections
Meta: ‘how we quantified text’
Leads: clusters, patterns or themes for further digging
" class="wp-image-31413" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/7-angles-for-data-stories-text_docs-1.png 959w" sizes="auto, (max-width: 625px) 100vw, 625px" /></a></figure>



<span id="more-31215"></span>



<h2 class="wp-block-heading">Text-based data journalism most often uses an exploratory feature format</h2>



<p class="wp-block-paragraph">While <a href="https://onlinejournalismblog.com/2020/08/11/here-are-the-7-types-of-stories-most-often-found-in-data/">most data stories about <em>numerical</em> data focused on scale or change</a>, the data stories about text I looked at were overwhelmingly dominated by <strong>exploratory</strong> formats.</p>



<p class="wp-block-paragraph">There are two obvious possible explanations for this. First, analysing text normally requires more time and skill than working with numerical data. That would be hard to justify for a simpler news article revealing scale or change.</p>



<p class="wp-block-paragraph">Second, text is often rich and complex, lending itself more to exploration. The analysis itself will often require explanation too, especially if it involves classification.</p>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(1166 / 710)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1166" height="710" alt="What 1.2 million parliamentary speeches can teach us about gender representation" class="wp-block-jetpack-slideshow_image wp-image-31222" data-id="31222" data-aspect-ratio="1166 / 710" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png 1166w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png?w=150&amp;h=91 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png?w=300&amp;h=183 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png?w=768&amp;h=468 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png?w=1024&amp;h=624 1024w" sizes="(max-width: 1166px) 100vw, 1166px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Pudding <a href="https://pudding.cool/2018/07/women-in-parliament/">explore political speech</a> through a series of categories.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="960" height="540" alt="The hidden structure of the Apple keynote" class="wp-block-jetpack-slideshow_image wp-image-31227" data-id="31227" data-aspect-ratio="960 / 540" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/applekeynote.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/applekeynote.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/applekeynote.png?w=150&amp;h=84 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/applekeynote.png?w=300&amp;h=169 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/applekeynote.png?w=768&amp;h=432 768w" sizes="(max-width: 960px) 100vw, 960px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption"><a href="https://qz.com/261181/the-hidden-structure-of-the-apple-keynote">This Quartz feature</a> on Apple keynotes quantifies text in a range of ways</figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph"><strong>The Pudding</strong>&#8216;s <a href="https://pudding.cool/2018/07/women-in-parliament/">scrollytell data feature on speeches in Parliament</a> is a good example of an exploratory approach: text from over a million speeches is classified in two ways: by the gender of the speaker, and by the topic it relates to. It is quantified in terms of the percentage of time that speakers spend on each topic (code is shared in a <a href="https://github.com/dldx/women-in-parliament/tree/master">GitHub repo</a>). </p>



<p class="wp-block-paragraph">The story explores the data through a number of themes: the economy, welfare, education, and so on. Although the focus is largely on <em>variation</em> between men and women, this is secondary to the primary exploratory angle.</p>



<p class="wp-block-paragraph">The same thematic approach is adopted by South African data journalism website <strong>The Outlier</strong> in <a href="https://web.archive.org/web/20220211154430/https://theoutlier.co.za/news/82289/sona2022-how-does-it-compare-to-the-last-5-speeches-by-ramaphosa"><em>#SONA2022: How does it compare to the last 5 speeches by Ramaphosa?</em></a> and by <strong>Sueddeutsche Zeitung</strong> in <a href="https://www.sueddeutsche.de/projekte/artikel/politik/russland-propaganda-ria-novosti-e261162/">their feature on why Russians perceive the West as a threat</a>.</p>



<p class="wp-block-paragraph">Another way to structure an exploratory feature about text is by breaking the feature up around different <strong>questions</strong>. This is the approach taken by Quartz in <a href="https://qz.com/261181/the-hidden-structure-of-the-apple-keynote">The hidden structure of the Apple keynote</a> (&#8220;Who&#8217;s on stage?&#8221;, &#8220;Who’s the funniest?&#8221; and &#8220;When is the unveil?&#8221;)</p>



<h2 class="wp-block-heading">Stories about changes in language</h2>



<p class="wp-block-paragraph">Text can act as a proxy for cultural attitudes and fashions, or political focus, so it&#8217;s not surprising that many stories based on data analysis focus on <strong>changes</strong> in society or power that language can reveal.</p>



<p class="wp-block-paragraph">Text from social media and forums can be analysed to answer questions about whether the mood of populations is changing (<em><a href="https://www.economist.com/graphic-detail/2022/03/12/the-war-in-ukraine-has-made-russian-social-media-users-glum">The war in Ukraine has made Russian social-media users glum</a></em>) or if people are getting ruder (<em><a href="https://www.economist.com/britain/2017/09/16/foul-mouthed-mothers-are-causing-problems-for-mumsnet">Foul-mouthed mothers are causing problems for Mumsnet</a></em>). </p>



<p class="wp-block-paragraph">Song lyrics can be analysed to identify <a href="https://www.statsignificant.com/p/how-have-song-lyrics-changed-since">trends in negativity and complexity</a> while, over a longer timescale, book context can be <a href="https://www.ft.com/content/e577411e-3bf2-4fb4-872a-8b7d5e9139d3">analysed</a> to conclude, as <strong>Jon Burn-Murdoch</strong> <a href="https://x.com/jburnmurdoch/status/1743238493248037167?s=12&amp;t=oEwMkjIJwOFpyXJc2zZyxw">one Twitter thread</a>, that &#8220;western society is shifting away from a culture of progress, and towards one of caution, worry and risk-aversion&#8221;.</p>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(960 / 540)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="960" height="540" alt="" class="wp-block-jetpack-slideshow_image wp-image-31231" data-id="31231" data-aspect-ratio="960 / 540" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/these-graphs-show-how-much-words-of-the-year-actually-get-used-after-they-join-the-dictionary.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/these-graphs-show-how-much-words-of-the-year-actually-get-used-after-they-join-the-dictionary.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/these-graphs-show-how-much-words-of-the-year-actually-get-used-after-they-join-the-dictionary.png?w=150&amp;h=84 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/these-graphs-show-how-much-words-of-the-year-actually-get-used-after-they-join-the-dictionary.png?w=300&amp;h=169 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/these-graphs-show-how-much-words-of-the-year-actually-get-used-after-they-join-the-dictionary.png?w=768&amp;h=432 768w" sizes="(max-width: 960px) 100vw, 960px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">This Vice story (which no longer hosts the charts) uses trend charts to anchor text commentary</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1796" height="1046" alt="" class="wp-block-jetpack-slideshow_image wp-image-31242" data-id="31242" data-aspect-ratio="1796 / 1046" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png 1796w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png?w=150&amp;h=87 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png?w=300&amp;h=175 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png?w=768&amp;h=447 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png?w=1024&amp;h=596 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png?w=1440&amp;h=839 1440w" sizes="(max-width: 1796px) 100vw, 1796px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Economist analysed language on Mumsnet</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1604" height="736" alt="" class="wp-block-jetpack-slideshow_image wp-image-31241" data-id="31241" data-aspect-ratio="1604 / 736" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png 1604w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png?w=150&amp;h=69 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png?w=300&amp;h=138 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png?w=768&amp;h=352 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png?w=1024&amp;h=470 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png?w=1440&amp;h=661 1440w" sizes="(max-width: 1604px) 100vw, 1604px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Russian sentiment on social media was a novel angle on the war in Ukraine</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1188" height="940" alt="" class="wp-block-jetpack-slideshow_image wp-image-31245" data-id="31245" data-aspect-ratio="1188 / 940" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png 1188w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png?w=150&amp;h=119 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png?w=300&amp;h=237 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png?w=768&amp;h=608 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png?w=1024&amp;h=810 1024w" sizes="(max-width: 1188px) 100vw, 1188px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Change stories about text can focus on what has changed in a single document</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1314" height="970" alt="&quot;Laïcité&quot;, mot rare chez le président Nous avons comparé les occurrences du mot «laïcité» dans les discours du président Macron en 2018, les discours du chef de l'Etat en 2017 et les discours du candidat Macron." class="wp-block-jetpack-slideshow_image wp-image-31396" data-id="31396" data-aspect-ratio="1314 / 970" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png 1314w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png?w=150&amp;h=111 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png?w=300&amp;h=221 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png?w=768&amp;h=567 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png?w=1024&amp;h=756 1024w" sizes="(max-width: 1314px) 100vw, 1314px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Paris Match&#8217;s <a href="https://www.parismatch.com/Actu/Politique/Immigration-et-laicite-comment-la-rhetorique-de-Macron-evolue-1599491">data analysis</a> identified how Emmanuel Macron&#8217;s speeches had mentioned secularism much more often as a candidate than as president.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="2088" height="1190" alt="How TMC's case for deep-sea mining changed over the years: line chart
This graph tracks how often TMC uses terms linked to either a “green frame” or a “defense frame” for deep-sea mining in its press releases over time.
" class="wp-block-jetpack-slideshow_image wp-image-31487" data-id="31487" data-aspect-ratio="2088 / 1190" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png 2088w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png?w=150&amp;h=85 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png?w=300&amp;h=171 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png?w=768&amp;h=438 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png?w=1024&amp;h=584 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/mining-language-change.png?w=1440&amp;h=821 1440w" sizes="(max-width: 2088px) 100vw, 2088px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">From <a href="https://pulitzercenter.org/resource/how-we-analyzed-metals-companys-public-messaging-deep-sea-mining">How We Analyzed The Metals Company’s Public Messaging on Deep-Sea Mining</a></figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">Changes in <strong>political speech</strong> are an important indicator not only of the priorities of those in power, but also their relationships with each other and their role in setting the tone of national conversation. </p>



<p class="wp-block-paragraph">One Guardian analysis, for example, leads on <em><a href="https://www.theguardian.com/politics/ng-interactive/2026/feb/25/how-rightwing-rhetoric-has-risen-sharply-in-the-uk-parliament-an-exclusive-visual-analysis">How rightwing rhetoric has risen sharply in the UK parliament</a></em>, while the award-winning USA Today feature <em>Hope’ is out, ‘fight’ is in: Does tweeting divide Congress, or simply echo its division?”</em>&nbsp;reports that &#8220;Language has become more divided and emotional&#8221; based on an analysis of more than 2.8 million tweets posted by members of Congress (the original article is no longer online but <a href="https://nationalpress.org/award-story/usa-today-data-and-visual-journalists-win-innovative-storytelling-award/">parts can be seen here</a>). </p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/Vexb0dx7zQ0?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">More simply, the Press Association used the official record of Parliament to calculate that mentions of &#8220;Brexit&#8221; had gone <em><a href="https://www.shropshirestar.com/news/uk-news/2020/01/30/from-nine-mentions-a-year-to-9000-how-mps-caught-the-brexit-bug/">From nine mentions a year to 9,000</a></em>.</p>



<p class="wp-block-paragraph">A more investigative use of change is provided by the Washington Post&#8217;s <a href="https://www.washingtonpost.com/investigations/whistleblowers-say-usaids-ig-removed-critical-details-from-public-reports/2014/10/22/68fbc1a0-4031-11e4-b03f-de718edeb92f_story.html">investigation into claims that &#8220;USAID’s IG removed critical details from public reports&#8221;</a>. By obtaining draft versions of audits reporters were able to identify what changes were made before publication: &#8220;more than 400 negative references were removed from the audits between the draft and final versions&#8221;</p>



<h2 class="wp-block-heading">Ranking the most used words and phrases<a href="https://github.com/paulbradshaw/dealingwithdocuments/edit/master/textstories.md#ranking-stories"></a></h2>



<p class="wp-block-paragraph">&#8220;Doncaster is the least sexist place in Britain when it comes to street names,&#8221; the <strong>Doncaster Free Press</strong> <a href="https://www.doncasterfreepress.co.uk/news/people/doncaster-has-the-least-sexist-street-names-in-britain-according-to-new-survey-2843062">reported</a> in 2020. The newspaper hadn&#8217;t done the analysis: a PR firm had scraped the names of over 200,000 streets and classified them along gender lines. They had also ranked the places with the biggest gap between male and female proportions, and the most common names (Victoria and John).</p>



<p class="wp-block-paragraph">Ranking can be a quick way to get a story out of a corpus of text — if you can identify a pattern to extract data from. <strong>ABC News</strong> in Australia, for example, analysed over 1,000 job descriptions to <a href="https://www.abc.net.au/news/2018-05-03/what-job-ads-reveal-about-the-rising-internship-culture/9713918">reveal</a> &#8220;some of the most common words and phrases advertisers used to attract interns&#8221;, while <em><a href="https://colinmorris.github.io/blog/size-of-things">The size of things: an ngram experiment</a></em> &#8220;used Google Books’ Ngram dataset to find the most popular size analogies in English books&#8221;. </p>



<p class="wp-block-paragraph">It turns out that &#8220;size of a pea&#8221; was the most common. </p>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(1260 / 758)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1260" height="758" alt="Alright, so we already know that “Baby” is the most popular name by far, but check out the other most popular names. Apart from Jesus, these names have a long history of popularity in the US. While many of these names hit their peak popularity in the 1950s, many are still popular today, with Michael and John still ranking as the 8th and 26th and most common boy’s names in the US in 2016, respectively." class="wp-block-jetpack-slideshow_image wp-image-31313" data-id="31313" data-aspect-ratio="1260 / 758" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png 1260w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png?w=150&amp;h=90 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png?w=300&amp;h=180 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png?w=768&amp;h=462 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png?w=1024&amp;h=616 1024w" sizes="(max-width: 1260px) 100vw, 1260px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Pudding&#8217;s <a href="https://pudding.cool/2019/05/names-in-songs/">exploratory feature on song names</a> includes lots of ranking.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1598" height="1102" alt="EU lawmakers talked most about Ukraine, Russia and China Countries colored by the share of European Parliament plenary speeches in which they were nominally mentioned: Choropleth where countries with more mentions are more darkly coloured - Russia is darkest" class="wp-block-jetpack-slideshow_image wp-image-31318" data-id="31318" data-aspect-ratio="1598 / 1102" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png 1598w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png?w=150&amp;h=103 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png?w=300&amp;h=207 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png?w=768&amp;h=530 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png?w=1024&amp;h=706 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png?w=1440&amp;h=993 1440w" sizes="(max-width: 1598px) 100vw, 1598px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">DW <a href="https://www.dw.com/en/eu-foreign-policy-which-countries-dominate-the-agenda/a-69236555">chose a choropleth map</a> to compare countries by their mentions &#8211; this is probably a less effective method than a bar chart.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1270" height="1004" alt="Chart showing most-used words in the 2017 election manifestos: 'Labour' and 'people' are the most common" class="wp-block-jetpack-slideshow_image wp-image-31317" data-id="31317" data-aspect-ratio="1270 / 1004" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png 1270w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png?w=150&amp;h=119 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png?w=300&amp;h=237 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png?w=768&amp;h=607 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png?w=1024&amp;h=810 1024w" sizes="(max-width: 1270px) 100vw, 1270px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Guardian used circles to compare the most-used terms in a <a href="https://www.theguardian.com/politics/datablog/2017/may/20/general-election-2017-manifesto-word-count-in-data">story on election manifestos</a></figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1308" height="1082" alt="Les mots les plus prononcés par Emmanuel Macron en 2018" class="wp-block-jetpack-slideshow_image wp-image-31393" data-id="31393" data-aspect-ratio="1308 / 1082" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png 1308w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png?w=150&amp;h=124 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png?w=300&amp;h=248 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png?w=768&amp;h=635 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png?w=1024&amp;h=847 1024w" sizes="(max-width: 1308px) 100vw, 1308px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Ranking words was a key part of Paris Match&#8217;s <a href="https://www.parismatch.com/Actu/Politique/Les-mots-du-president-a-l-epreuve-1598267">Weight of Words</a> project, analysing speeches by President Macron </figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">Even exploratory features can focus on exploring different rankings: the <strong>Pudding</strong> feature <a href="https://pudding.cool/2019/05/names-in-songs/"><em>Sing My Name</em></a> moves through sections ranking the most popular names in songs, which songs have the most repeat mentions of the same name, and what songs contain the most names, among others.</p>



<h2 class="wp-block-heading">The scale of an issue — revealed in language<a href="https://github.com/paulbradshaw/dealingwithdocuments/edit/master/textstories.md#stories-revealing-scale-in-text"></a></h2>



<p class="wp-block-paragraph">Text analysis can allow journalists to reveal the <strong>scale</strong> of a problem which is not visible in more traditional statistics. </p>



<p class="wp-block-paragraph">Most of the text-based stories I&#8217;ve worked on at the BBC fall into this category: there are no official statistics on the outcomes of police misconduct cases, but I was able to <a href="https://www.bbc.co.uk/news/uk-59594712">analyse police watchdog reports</a> for a story that revealed &#8220;Half of police employees who committed gross misconduct were not dismissed&#8221;. </p>



<p class="wp-block-paragraph">Along similar lines, classifying text in festival line-ups led to the story <em><a href="https://www.bbc.co.uk/news/newsbeat-61512053">Music festivals: Only 13% of UK headliners in 2022 are female</a></em> and declarations of interest were analysed to <a href="https://www.bbc.co.uk/news/uk-england-40709220">establish</a> that &#8220;One in five MPs continue to employ a member of their family&#8221;.</p>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(625 / 475)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="474" alt="How much do female characters speak in Game of Thrones? Pictogram chart showing 22% dialogue is female" class="wp-block-jetpack-slideshow_image wp-image-31322" data-id="31322" data-aspect-ratio="625 / 475" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption"><a href="https://www.bbc.co.uk/news/entertainment-arts-48335099">This BBC story</a> uses a pictogram chart to communicate the scale of female dialogue. The margin of error is detailed in a footnote.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="596" alt="Female and non-binary artists are underrepresented as UK festival headliners: out of 200 headline acts at the biggest UK festivals, only 26 were an all-female band or solo artist" class="wp-block-jetpack-slideshow_image wp-image-31321" data-id="31321" data-aspect-ratio="625 / 596" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The scale of representation of female and non-binary artists is the focus of <a href="https://www.bbc.co.uk/news/newsbeat-61512053">this piece of BBC data journalism</a></figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="438" alt="A minute-by-minute look at one hour of hockey on ESPN Our AI detected gambling promoted in 40 of the 60 minutes we analyzed during a Jan. 23 NHL game between Tampa Bay and Chicago." class="wp-block-jetpack-slideshow_image wp-image-31457" data-id="31457" data-aspect-ratio="625 / 438" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Washington Post <a href="https://www.washingtonpost.com/investigations/interactive/2026/05/19/post-ai-analysis-sports-tv-detected-an-excess-gambling-ads/?pwapi_token=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJyZWFzb24iOiJnaWZ0IiwibmJmIjoxNzc5MTYzMjAwLCJpc3MiOiJzdWJzY3JpcHRpb25zIiwiZXhwIjoxNzgwNTQ1NTk5LCJpYXQiOjE3NzkxNjMyMDAsImp0aSI6IjkwOTY5ZDQ5LTIzYzQtNGE4My1hMzJhLTgwZmJkNmY2NWU1OSIsInVybCI6Imh0dHBzOi8vd3d3Lndhc2hpbmd0b25wb3N0LmNvbS9pbnZlc3RpZ2F0aW9ucy9pbnRlcmFjdGl2ZS8yMDI2LzA1LzE5L3Bvc3QtYWktYW5hbHlzaXMtc3BvcnRzLXR2LWRldGVjdGVkLWFuLWV4Y2Vzcy1nYW1ibGluZy1hZHMvIn0.9Fpaq6B-aXqJnOZJcGdRpxTmd9zWwUtKzY9sQQxi8Iw">used AI to detect text in sports broadcasts</a></figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph"><strong>Social media</strong> analysis might focus on the scale of a problem on particular platforms: 3 million tweets were analysed for the story&nbsp;<em><a href="https://www.bbc.co.uk/news/uk-63330885">Scale of abuse of politicians on Twitter revealed</a></em> and similar techniques were <a href="https://www.newstatesman.com/science-tech/2017/09/we-tracked-25688-abusive-tweets-sent-women-mps-half-were-directed-diane-abbott">used by Amnesty</a> and <a href="https://www.turing.ac.uk/news/seven-ten-premier-league-footballers-face-twitter-abuse">by the Turing Institute</a> to establish the scale of abuse of particular groups. </p>



<p class="wp-block-paragraph">Scale is a useful fallback angle if you are quantifying text, because no one else will have analysed the text before. The <a href="https://www.theguardian.com/world/2025/sep/28/far-right-facebook-groups-are-engine-of-radicalisation-in-uk-data-investigation-suggests">Guardian&#8217;s investigation into extremist Facebook groups</a>, for example, leads on the finding that the network &#8220;exposes <em>hundreds of thousands</em> of Britons to racist language, conspiracy and disinformation&#8221; (my emphasis), while USA Today analysed campaign rally speech transcripts for <a href="https://eu.usatoday.com/story/news/politics/elections/2019/08/08/trump-immigrants-rhetoric-criticized-el-paso-dayton-shootings/1936742001/"><em>Trump used words like invasion, killer to discuss immigrants 500 times</em></a>. </p>



<p class="wp-block-paragraph">It can also be a useful approach for factchecking or putting a news event into context, as in Der Spiegel&#8217;s story <a href="https://www.spiegel.de/kultur/musik/till-lindemann-wie-viel-gewalt-steckt-in-rammsteins-texten-a-b6e0820c-2b2f-4366-b6aa-84cc162f5665?giftToken=3a86d996-6673-405c-9a6e-899f1aa05f68"><em>How much violence is in Rammstein&#8217;s lyrics?</em></a> following allegations of sexual assault against the band&#8217;s singer (the investigation was later dropped)</p>



<h2 class="wp-block-heading">Women versus men and other variation stories</h2>



<p class="wp-block-paragraph"><a href="https://github.com/paulbradshaw/dealingwithdocuments/edit/master/textstories.md#variation-stories"></a>Variation stories <a href="https://onlinejournalismblog.com/2025/12/08/telling-stories-with-data-more-on-the-difference-between-variation-stories-and-ranking-angles/">rely on an expectation of fairness, equality or parity</a>. This limits the opportunities for this angle, but it can work especially well where language reveals implicit biases in society that are not quantified anywhere else.</p>



<p class="wp-block-paragraph">The <strong>LA Times</strong>&#8216;s scrollytell <em><a href="https://www.latimes.com/projects/star-wars-movies-female-character-analysis/">There are more women than ever in Star Wars. Men still do most of the talking</a></em> is just one of a number of pieces of data journalism looking at gender variation using text. Others include <em><a href="https://julienassouline.github.io/data-studios-projects/Star_Wars/">The Gender Divide in Star Wars Scripts</a></em>, <em><a href="https://www.theguardian.com/global/2020/jan/18/meghan-gets-more-than-twice-as-many-negative-headlines-as-positive">Meghan gets twice as many negative headlines as positive, analysis finds</a></em>, The New York Times&#8217;s <em><a href="https://www.nytimes.com/interactive/2017/11/07/upshot/modern-love-what-we-write-when-we-write-about-love.html">The Words Men and Women Use When They Write About Love</a></em>, and The Pudding&#8217;s&nbsp;<a href="https://pudding.cool/2020/07/gendered-descriptions/"><em>The physical traits that define men &amp; women in literature</em></a>:</p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">&#8220;Do authors really mention particular body parts more for men than for women? Are women’s bodies described using different adjectives than those attributed to men?&#8221;</p>
</blockquote>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(625 / 553)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="552" alt="Twitter's algorithm does not seem to silence conservatives: lollipop chart showing that language such as 'Donald' tend to be treated more positively and 'Democratic' more negatively when served to &quot;a clone of Donald Trump's account&quot;" class="wp-block-jetpack-slideshow_image wp-image-31323" data-id="31323" data-aspect-ratio="625 / 553" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png 884w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Economist <a href="https://www.economist.com/graphic-detail/2020/08/01/twitters-algorithm-does-not-seem-to-silence-conservatives 
">used a lollipop chart</a> to show the variation between chronological and algorithmic newsfeeds</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="625" height="502" alt="Two bar charts showing differences in word frequency between two candidates" class="wp-block-jetpack-slideshow_image wp-image-31403" data-id="31403" data-aspect-ratio="625 / 503" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=625" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=625 625w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=1250 1250w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=768 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=1024 1024w" sizes="(max-width: 625px) 100vw, 625px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Paris Match <a href="https://www.parismatch.com/Actu/Politique/Macron-Le-Pen-Leurs-silences-en-disent-long-1250597">visualised</a> variation in word frequency between two politicians</figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">Variation along lines of <strong>race</strong> is revealed in <em><a href="https://www.independent.co.uk/sport/football/news/football-commentary-racism-racial-bias-study-pfa-runrepeat-a9592276.html">Football commentary racially biased, study finds</a></em> and HuffPost UK&#8217;s <em><a href="https://www.huffingtonpost.co.uk/entry/metropolitan-police_uk_603fa18ec5b617a7e411ffc5">The Met Police Are More Likely To Publish Your Mugshot If You&#8217;re Black</a></em>, which classified press releases on criminal sentencing based on mentions of ethnicity, and compared the proportions to full data on sentencing by ethnicity.</p>



<p class="wp-block-paragraph">Examples of <strong>political and ideological variation</strong> can be found in The Economist&#8217;s&nbsp;<a href="https://www.economist.com/graphic-detail/2020/08/01/twitters-algorithm-does-not-seem-to-silence-conservatives"><em>Twitter’s algorithm does not seem to silence conservatives</em></a> and Paris Match&#8217;s political speech comparison <a href="https://www.parismatch.com/Actu/Politique/Macron-Le-Pen-Leurs-silences-en-disent-long-1250597"><em>Macron-Le Pen: Their silences speak volumes</em></a>.</p>



<p class="wp-block-paragraph">The Washington Post&#8217;s <em><a href="https://www.washingtonpost.com/news/monkey-cage/wp/2017/08/31/almost-all-news-coverage-of-the-barcelona-attack-mentioned-terrorism-very-little-coverage-of-charlottesville-did/">Almost all news coverage of the Barcelona attack mentioned terrorism. Very little coverage of Charlottesville did</a></em>: &#8220;Even before we did our study,&#8221; the story notes, &#8220;research showed disproportionately high media coverage of terrorism committed by Muslims — even though right-wing extremist groups have committed more attacks&#8221;. </p>



<p class="wp-block-paragraph">Given that databases of media coverage (such as <a href="https://www.lexisnexis.co.uk/products/nexis.html">Nexis</a>) provide an accessible source of text data, this feels like an area ripe for more analysis.</p>



<h2 class="wp-block-heading">&#8216;Meta&#8217; data stories are all methodologies</h2>



<p class="wp-block-paragraph">It is increasingly rare to find &#8216;meta&#8217; data journalism angles: stories about a lack of data, poor data, or &#8216;<em>Get the data</em>&#8216; articles that share data for others to analyse. And the exclusive nature of text data makes it even more unlikely. </p>



<p class="wp-block-paragraph">But the complexity of quantifying text for analysis means that there is sometimes a need to tell the <strong>story-behind-the-story</strong> about the methodologies that were employed: <em><a href="https://news.sky.com/story/how-sky-news-investigated-xs-algorithm-for-political-bias-13463916">How Sky News investigated X&#8217;s algorithm for political bias</a></em>, along with The Guardian&#8217;s <em><a href="https://www.theguardian.com/world/2019/mar/06/how-we-combed-leaders-speeches-to-gauge-populist-rise">How we measured the rise of populist rhetoric</a></em>, and <em><a href="https://www.theguardian.com/world/2025/sep/28/reading-the-post-riot-posts-how-we-traced-far-right-radicalisation-across-51000-facebook-messages">Reading the post-riot posts: how we traced far-right radicalisation across 51,000 Facebook messages</a></em> are typical examples, while <em><a href="https://www.theguardian.com/politics/2026/feb/25/behind-the-guardians-analysis-of-100-years-of-mps-language-on-immigration">Behind the Guardian&#8217;s analysis of 100 years of MPs&#8217; language on immigration</a></em> shows another approach.</p>



<h2 class="wp-block-heading">Connections, similarities, and correlations revealed by text analysis</h2>



<div class="wp-block-jetpack-slideshow aligncenter" data-effect="slide" style="--aspect-ratio:calc(1030 / 1104)"><div class="wp-block-jetpack-slideshow_container swiper"><ul class="wp-block-jetpack-slideshow_swiper-wrapper swiper-wrapper"><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1030" height="1104" alt="" class="wp-block-jetpack-slideshow_image wp-image-31212" data-id="31212" data-aspect-ratio="1030 / 1104" src="https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png 1030w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png?w=140&amp;h=150 140w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png?w=280&amp;h=300 280w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png?w=768&amp;h=823 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png?w=955&amp;h=1024 955w" sizes="(max-width: 1030px) 100vw, 1030px" /></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1278" height="1034" alt="Four histograms showing frequency of emails between Epstein or his assistants and four people in positions of power" class="wp-block-jetpack-slideshow_image wp-image-31327" data-id="31327" data-aspect-ratio="1278 / 1034" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png 1278w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png?w=150&amp;h=121 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png?w=300&amp;h=243 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png?w=768&amp;h=621 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png?w=1024&amp;h=828 1024w" sizes="(max-width: 1278px) 100vw, 1278px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Guardian&#8217;s <a href="https://www.theguardian.com/us-news/ng-interactive/2026/mar/18/jeffrey-epsteins-elite-relationships-visualised-the-prince-the-sultan-and-the-politicians">story on Epstein&#8217;s relationships</a> uses a &#8216;small multiple&#8217; of histograms to compare four connections over time.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1746" height="1036" alt="Chart showing quantity of emails between Epstein or Ghislaine Maxwell and Andrew Mountbatten, Sarah Ferguson and their assistants plotted from 2001 to 2011" class="wp-block-jetpack-slideshow_image wp-image-31326" data-id="31326" data-aspect-ratio="1746 / 1036" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png 1746w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png?w=150&amp;h=89 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png?w=300&amp;h=178 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png?w=768&amp;h=456 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png?w=1024&amp;h=608 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png?w=1440&amp;h=854 1440w" sizes="(max-width: 1746px) 100vw, 1746px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Guardian&#8217;s <a href="https://www.theguardian.com/us-news/ng-interactive/2026/mar/18/jeffrey-epsteins-elite-relationships-visualised-the-prince-the-sultan-and-the-politicians">story on Epstein&#8217;s relationships</a> uses a bubble chart with one axis to visualise connections with one person over time.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="2086" height="1186" alt="Network diagram with Donald Trump as the highlighted node. On the right a timeline lists each connection." class="wp-block-jetpack-slideshow_image wp-image-31325" data-id="31325" data-aspect-ratio="2086 / 1186" src="https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png 2086w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png?w=150&amp;h=85 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png?w=300&amp;h=171 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png?w=768&amp;h=437 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png?w=1024&amp;h=582 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png?w=1440&amp;h=819 1440w" sizes="(max-width: 2086px) 100vw, 2086px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption"><a href="https://epsteinvisualizer.com/">The Epstein Network</a> shows clusters of connections and allows you to click on a node and browse through connections in the data.</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="2232" height="1318" alt="When we separate his speeches into two types - teleprompter speeches and those which are off the cuff - it's clear that Trump is reading from a script in almost all of his most populist addresses" class="wp-block-jetpack-slideshow_image wp-image-31391" data-id="31391" data-aspect-ratio="2232 / 1318" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png 2232w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png?w=150&amp;h=89 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png?w=300&amp;h=177 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png?w=768&amp;h=454 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png?w=1024&amp;h=605 1024w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png?w=1440&amp;h=850 1440w" sizes="(max-width: 2232px) 100vw, 2232px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">The Guardian quantified political speech to identify a relationship between scripts and populism</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1106" height="1156" alt="Les mots associés à &quot;système&quot; par Jean-Luc Mélenchon - bar chart showing the words most associated with 'system' for one candidate" class="wp-block-jetpack-slideshow_image wp-image-31407" data-id="31407" data-aspect-ratio="1106 / 1156" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png 1106w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png?w=144&amp;h=150 144w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png?w=287&amp;h=300 287w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png?w=768&amp;h=803 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png?w=980&amp;h=1024 980w" sizes="(max-width: 1106px) 100vw, 1106px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">Paris Match <a href="https://www.parismatch.com/Actu/Politique/Qui-visent-les-candidats-quand-ils-parlent-de-systeme-1238246">used co-occurence analysis</a> to quantify the words politicians most associated with &#8216;system&#8217;</figcaption></figure></li><li class="wp-block-jetpack-slideshow_slide swiper-slide"><figure><img loading="lazy" width="1202" height="1014" alt="Linke und AfD seltener erwähnt
" class="wp-block-jetpack-slideshow_image wp-image-31489" data-id="31489" data-aspect-ratio="1202 / 1014" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/linke-und-afd-seltener-erwahnt.png" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/linke-und-afd-seltener-erwahnt.png 1202w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/linke-und-afd-seltener-erwahnt.png?w=150&amp;h=127 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/linke-und-afd-seltener-erwahnt.png?w=300&amp;h=253 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/linke-und-afd-seltener-erwahnt.png?w=768&amp;h=648 768w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/linke-und-afd-seltener-erwahnt.png?w=1024&amp;h=864 1024w" sizes="(max-width: 1202px) 100vw, 1202px" /><figcaption class="wp-block-jetpack-slideshow_caption gallery-caption">From <a href="https://www.swr.de/swraktuell/landtagswahl-2026-wie-ki-chatbots-deine-wahlentscheidung-beeinflussen-100.html">Infos zur Wahl: Warum Chatbots nicht neutral informieren</a></figcaption></figure></li></ul><a class="wp-block-jetpack-slideshow_button-prev swiper-button-prev swiper-button-white" role="button"></a><a class="wp-block-jetpack-slideshow_button-next swiper-button-next swiper-button-white" role="button"></a><a aria-label="Pause Slideshow" class="wp-block-jetpack-slideshow_button-pause" role="button"></a><div class="wp-block-jetpack-slideshow_pagination swiper-pagination swiper-pagination-white"></div></div></div>



<p class="wp-block-paragraph">Relationship angles were the least common in the examples of text-based data journalism I found. The few stories that did adopt this focus, however, pointed to at least three possible types of relationships that could be used for stories: <strong>correlations, similarities and connections</strong>.</p>



<p class="wp-block-paragraph">A <strong>correlation</strong> angle can be seen in The Economist’s <a href="https://www.economist.com/graphic-detail/2019/05/18/why-beer-snobs-guzzle-lagers-they-claim-to-dislike"><em>Why beer snobs guzzle lagers they claim to dislike</em></a>, which visualises the relationship between word frequencies and user ratings. </p>



<p class="wp-block-paragraph">The example points to the types of sources that might provide material for similar relationship stories: <strong>where text appears alongside a numerical measure</strong>.&nbsp;</p>



<p class="wp-block-paragraph">Reviews are just one such category of text data. <strong>Social media updates</strong> (which appear alongside numbers of likes, shares or views, and timestamps) and cultural texts such as <strong>books, songs and scripts</strong> (figures on sales, streams and views) have the same qualities.&nbsp;</p>



<p class="wp-block-paragraph">An alternative approach is to quantify text yourself, and combine it with other data: The Guardian <a href="https://www.theguardian.com/world/ng-interactive/2019/mar/07/the-teleprompter-test-why-trumps-populism-is-often-scripted">used this to establish a relationship between Donald Trump&#8217;s use of a teleprompter and populism</a>, while Paris Match <a href="https://www.parismatch.com/Actu/Politique/Une-campagne-folle-vue-depuis-Google-Trends-1250405">used it</a> to establish a <strong>lack of relationship</strong> between search interest and candidates&#8217; mentions of an issue.</p>



<p class="wp-block-paragraph"><strong>Similarities</strong> formed the focus of the network graph-led<a href="https://web.archive.org/web/20140107030355/http://elms.wordpress.com/2008/03/04/lexical-distance-among-languages-of-europe"> <em>A Map of Lexical Distances Between Europe’s Languages</em></a> and<a href="https://alternativetransport.wordpress.com/2015/05/05/34/"> <em>Lexical Distance Among Languages of Europe 2015</em></a>. The<a href="https://marcinciura.wordpress.com/2019/08/07/visualizing-lexical-distance-in-three-dimensions/"> methodology</a> points to how the &#8216;distance&#8217; between words can be quantified in order to identify patterns.</p>



<p class="wp-block-paragraph"><strong>Connections</strong> are the focus for <a href="https://qz.com/650796/mathematicians-mapped-out-every-game-of-thrones-relationship-to-find-the-main-character"><em>Mathematicians mapped out every “Game of Thrones” relationship to find the main character</em></a>, which quantifies text by classifying two characters being mentioned in the same sentence as a connection. Although this is a ranking story (the focus is on identifying the &#8216;main&#8217; character), the analysis could equally have been used to tell a story about relationships, and the same method could be adapted for any text where entities (people, companies, locations) are mentioned.&nbsp;Paris Match, for example, <a href="https://www.parismatch.com/Actu/Politique/Qui-visent-les-candidats-quand-ils-parlent-de-systeme-1238246">used co-occurrence analysis</a> to identify what words candidates associated with &#8216;system&#8217;. </p>



<p class="wp-block-paragraph">Connections encoded in large language models provide another potential source: Economist data journalist, Sondre Solstad, uses this to look at the <a href="https://www.economist.com/interactive/culture/2025/03/20/what-is-in-a-name">connotations of baby names</a>.</p>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<iframe loading="lazy" class="youtube-player" width="625" height="352" src="https://www.youtube.com/embed/4eGEoXaWOmY?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe>
</div></figure>



<p class="wp-block-paragraph">The <strong>Epstein files</strong> provide further ideas for working with text: The Economist&#8217;s<a href="https://www.economist.com/interactive/international/2026/02/12/inside-epsteins-network"> <em>Inside Epstein&#8217;s Network</em></a>, for example, also chooses a ranking angle revealing the &#8220;500 people who appear most frequently&#8221;, and while The Guardian headlined their story as “<a href="https://www.theguardian.com/us-news/ng-interactive/2026/mar/18/jeffrey-epsteins-elite-relationships-visualised-the-prince-the-sultan-and-the-politicians">Jeffrey Epstein’s elite relationships visualised</a>”, what is specifically visualised is the relationships between correspondence and time. What is missing from coverage generally are the clusters, cliques and bridges that can be generated by the<a href="https://epsteinvisualizer.com/"> Epstein Document Network Explorer</a>.</p>



<h2 class="wp-block-heading">Leads from analysis of text</h2>



<p class="wp-block-paragraph">The Epstein files also provide an example of how data-driven approaches to text can provide useful story <strong>leads</strong>. The New York Times, for example, <a href="https://reutersinstitute.politics.ox.ac.uk/news/epstein-files-investigative-journalism-prince-andrew-arrest">used</a> AI to help &#8220;identify clusters, patterns and themes,&#8221; although it may be that this is being underused. &#8220;The brunt of the work,&#8221; they admit, &#8220;is being done by a team of editors and reporters across beats and bureaus who have been preparing to dig into the files long before they were published.&#8221;</p>



<figure class="wp-block-image size-large"><a href="https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png"><img loading="lazy" width="960" height="540" data-attachment-id="31380" data-permalink="https://onlinejournalismblog.com/2026/05/14/how-to-tell-stories-about-documents-and-text-using-data-journalism/working-with-text-and-documents/" data-orig-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png" data-orig-size="960,540" data-comments-opened="1" data-image-meta="{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}" data-image-title="Working with text and documents" data-image-description="" data-image-caption="" data-large-file="https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png?w=625" src="https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png?w=960" alt="topic modelling of police misconduct reports shows 10 clusters of terms" class="wp-image-31380" srcset="https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png 960w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png?w=150 150w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png?w=300 300w, https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png?w=768 768w" sizes="auto, (max-width: 960px) 100vw, 960px" /></a><figcaption class="wp-element-caption">Using topic modelling with police misconduct reports allowed me to identify common themes</figcaption></figure>



<p class="wp-block-paragraph">Once you have collected a corpus of text, exploratory analysis can surface potential stories you may not have considered. <a href="https://github.com/paulbradshaw/dealingwithdocuments/edit/master/textstories.md#leads-from-data-analysis"></a></p>



<p class="wp-block-paragraph">When I worked on the BBC story&nbsp;<a href="https://github.com/BBC-Data-Unit/pay-to-work-dbs">DBS background checks mean NHS staff &#8216;paying to work&#8217;</a>, for example,&nbsp;the angle came as a result of identifying phrases which appeared multiple times in NHS job ads (we then went on to use the data to establish the scale of this). And once I&#8217;d collected documents on police misconduct, topic modelling allowed me to see common themes across a large number of documents, any one of which could have been an avenue for further reporting.</p>



<p class="wp-block-paragraph"><strong><em>I&#8217;d love to know of any examples where text analysis has been used to identify story leads. If you&#8217;ve been involved in a text analysis project, please let me know in the comments or <a href="https://www.linkedin.com/in/paulbradshawuk/">on LinkedIn</a></em></strong>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://onlinejournalismblog.com/2026/05/14/how-to-tell-stories-about-documents-and-text-using-data-journalism/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">31215</post-id>
		<media:content url="https://0.gravatar.com/avatar/3e60435c09b44f66a8f2b3f74c8725c4412847d4385077948734b7d7fad54c8b?s=96&#38;d=identicon&#38;r=G" medium="image">
			<media:title type="html">paulbradshawuk</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/speeches_gender.png" medium="image">
			<media:title type="html">What 1.2 million parliamentary speeches can teach us about gender representation</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/applekeynote.png" medium="image">
			<media:title type="html">The hidden structure of the Apple keynote</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/these-graphs-show-how-much-words-of-the-year-actually-get-used-after-they-join-the-dictionary.png" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/mumsnet_charts.png" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/russiansocialmediamood.png" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/washpo_change_investigation.png" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/changewordsfr.png" medium="image">
			<media:title type="html">&#034;Laïcité&#034;, mot rare chez le président Nous avons comparé les occurrences du mot «laïcité» dans les discours du président Macron en 2018, les discours du chef de l&#039;Etat en 2017 et les discours du candidat Macron.</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/pudding_ranking.png" medium="image">
			<media:title type="html">Alright, so we already know that “Baby” is the most popular name by far, but check out the other most popular names. Apart from Jesus, these names have a long history of popularity in the US. While many of these names hit their peak popularity in the 1950s, many are still popular today, with Michael and John still ranking as the 8th and 26th and most common boy’s names in the US in 2016, respectively.</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/dw_rankingmap.png" medium="image">
			<media:title type="html">EU lawmakers talked most about Ukraine, Russia and China Countries colored by the share of European Parliament plenary speeches in which they were nominally mentioned: Choropleth where countries with more mentions are more darkly coloured - Russia is darkest</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_wordcounts_election.png" medium="image">
			<media:title type="html">Chart showing most-used words in the 2017 election manifestos: &#039;Labour&#039; and &#039;people&#039; are the most common</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/rankingwordsfr.png" medium="image">
			<media:title type="html">Les mots les plus prononcés par Emmanuel Macron en 2018</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-got-scale.png?w=625" medium="image">
			<media:title type="html">How much do female characters speak in Game of Thrones? Pictogram chart showing 22% dialogue is female</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/bbc-festivals-gender-scale.png?w=625" medium="image">
			<media:title type="html">Female and non-binary artists are underrepresented as UK festival headliners: out of 200 headline acts at the biggest UK festivals, only 26 were an all-female band or solo artist</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-21-at-19.50.23.png?w=625" medium="image">
			<media:title type="html">A minute-by-minute look at one hour of hockey on ESPN Our AI detected gambling promoted in 40 of the 60 minutes we analyzed during a Jan. 23 NHL game between Tampa Bay and Chicago.</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/economist-variation.png?w=625" medium="image">
			<media:title type="html">Twitter&#039;s algorithm does not seem to silence conservatives: lollipop chart showing that language such as &#039;Donald&#039; tend to be treated more positively and &#039;Democratic&#039; more negatively when served to &#034;a clone of Donald Trump&#039;s account&#034;</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/variationfr2.png?w=625" medium="image">
			<media:title type="html">Two bar charts showing differences in word frequency between two candidates</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/03/scatterplot_beer.png" medium="image" />

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_histo.png" medium="image">
			<media:title type="html">Four histograms showing frequency of emails between Epstein or his assistants and four people in positions of power</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/guardian_network_bubble.png" medium="image">
			<media:title type="html">Chart showing quantity of emails between Epstein or Ghislaine Maxwell and Andrew Mountbatten, Sarah Ferguson and their assistants plotted from 2001 to 2011</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/04/epsteinnetworktrump.png" medium="image">
			<media:title type="html">Network diagram with Donald Trump as the highlighted node. On the right a timeline lists each connection.</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/screenshot-2026-05-13-at-10.30.28.png" medium="image">
			<media:title type="html">When we separate his speeches into two types - teleprompter speeches and those which are off the cuff - it&#039;s clear that Trump is reading from a script in almost all of his most populist addresses</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/relfr.png" medium="image">
			<media:title type="html">Les mots associés à &#034;système&#034; par Jean-Luc Mélenchon - bar chart showing the words most associated with &#039;system&#039; for one candidate</media:title>
		</media:content>

		<media:content url="https://onlinejournalismblog.com/wp-content/uploads/2026/05/working-with-text-and-documents.png?w=960" medium="image">
			<media:title type="html">topic modelling of police misconduct reports shows 10 clusters of terms</media:title>
		</media:content>
	</item>
	</channel>
</rss>
