Rick On the Road

Cluster analysis for evaluation purposes, and how a hybrid Human-LLM approach can help

2026-05-31T15:39:48.144+00:00

Friend to Groucho Marx: “Life is difficult”

Groucho Marx to Friend: “Compared to what?”

Comparison is intrinsic to evaluation. Not only between one thing and another, but more often, between one category of things and another. Typologies are involved in many of the comparisons evaluators have to make (between people, activities, outcomes, locations, et cetera). Typologies can spring forth from our minds, but there are also systematic methods for developing them, broadly described as clustering methods. It is this second grouping that I want to discuss here.

What types of clustering methods are there?

Using one prompt to start off with, Claude identified for me 33 different clustering methods, which were organised into 9 different categories. These categories were defined largely on the basis of differences in computational mechanisms and mathematical principles which are involved in the operation of the clustering methods. But there were a couple of groups which were organised using different criteria relating to who does the task (human or algorithm) and what the method is for.

I then did some experimentation of my own, getting Claude to cluster the methods, using two different clustering methods. One is called a “maximum spanning tree”, which connects methods according to which method is most similar to which other method. You can see this in figure 1 below, which is best read by double clicking on the image to get greater magnification. The second experiment used an agglomerative hierarchical clustering to produce what is called a dendrogram i.e. a tree structure displaying nested categories of methods. You can see this in figure 2 below, again probably best inspected by magnifying the image first. I like both of these, for reasons that will become clearer below.

Figure 1: Maximum Spanning Tree

Figure 2: Dendrogram

Introducing the Text Cluster Analysis Lab
I developed this mini-app recently, in May this year, with very substantial help from Claude AI. You can view (and copy) the app by following this link. You will see that one of the eight workflow tabs builds a dendrogram.

How does it compare?

When I asked Claude to compare this app with the list of 33 clustering methods, it identified similarities with a number of methods, including hierarchical clustering, latent class analysis, and pile sorting. But its best fitting answer was that the method is a hybrid. “Its uniqueness comes precisely from joining a human-sorting logic (criteria generated from the material, framed by the researcher's domain statement) to an algorithmic back end (binary scoring + agglomerative clustering). No single method in the set occupies that position”. The Lab's nearest relatives — hierarchical clustering, LCA, pile/card sorting — sit in three different families i.e. separate branches of the spanning tree and dendrogram. In this respect the Lab can claim novelty. Hopefully the basis of this claim will become clearer as I explain the Lab in some detail.

The inputs

By Claude's estimate around 75% of cluster analysis methods work with quantitative data only, the rest work with text only (15%) or text or numbers (10%). The lab works with multiple bodies of texts. More specifically, the focus of the Lab is on similarities and differences between texts, not on internal structure within a text, as is the case with much thematic coding. The texts I have used, while testing the mini-app, are a set of storylines about alternative futures, collaboratively developed by participants in a ParEvo.org exercise. I also have plans to analyse a set of Most Significant Change stories.

Another important input is the choices users make about various settings, in each of the eight tabs making up the staged workflow. These include what version of Claude to use at various stages of the analysis, the temperature setting for each stage, and most importantly, the precise wording of the prompt that Claude will respond to, which tells it what to look for. In the Reliability and Cluster tabs choices are made about which identified sorting criteria to include, and what clustering settings to use. All these choices are important, as will be discussed below. They are what makes this approach hybrid: part human, part automated processes.

The outputs

The first is a list of attributes, which differentiate one group of texts from another. And a matrix, where rows list the texts, columns describe the attributes of those texts and cells describe their presence or absence in each text. The text attributes are identified by a Claude AI search and comparison of the texts contents, operating within user-defined context and scope settings. Attributes that fail to discriminate between the texts (those present in all of them or none of them) are filtered out automatically, because they contribute nothing to distinguishing one text from another.

A second matrix is constructed as the results of a "back-translation" type of rater reliability test. This enables users to remove from use those attributes which are unreliably identifiable and to identify texts whose analyses are less reliable than others.

The third output is a dendrogram, representing a nested classification of the texts, using the user's selection of relevant text attributes. This is built using an agglomerative process, firstly finding pairs of text which are most similar, then pairs of pairs of texts which are most similar, et cetera. When building the tree, the user also chooses how similarity between texts is measured (using either a Hamming or a Jaccard distance) and how clusters are linked together as they merge. Each branch of the tree structure includes tool-tip information on how the attributes of that cluster of texts differ from its sibling branch. This is systematically identified by the app, and humanly verifiable. It is not a subjective judgement — as is the case with a number of other types of clustering processes.

The fourth output is an open-ended Claude chat-type query facility, where the user can ask questions about: (a) a cluster of texts on their own, or (b) in comparison to their sibling cluster in the dendrogram, or (c) in comparison to the whole set of texts. With or without uploading of additional context information.

The fifth output is a set of exportable products, including a detailed provenance statement describing how all the products have been produced, a listing of all the identified text attributes, a copy of the text-by-attributes matrix, the rater reliability assessment, a copy of the dendrogram, a copy of any query dialogue, and a breakdown of token use and token costs for each stage of each analysis. The details of the analysis process, including settings and contents generated, can also be saved as a JSON file and reimported for reuse.

A sixth output supports the choice of criteria itself. Because the criteria matter more than any other input choice (they define the matrix from which everything else follows) it is worth being able to compare different sets of them and being able to choose deliberately from within these, rather than accepting whatever a single run happens to produce. The Selection tab lets you place the criteria from two or more runs side by side, each shown with quantitative measures of its performance: how reliably it was identified, how many texts it applies to, and how much its work overlaps that of the other criteria. Alongside these per-criterion figures are measures for the set as a whole, including its overall reliability and a measure of how well the set as a whole tells the texts apart. From this you can hand-pick a set of criteria drawn from across the runs, add criteria of your own wording if a distinction you care about was missed, and then re-score the texts against just that curated set. The result is a new analysis like any other, which can be clustered, reliability-tested, and queried in turn. A short built-in guide explains how these measures relate to one another — for instance, why a criterion that applies to very few texts can look misleadingly distinctive, or why minimising overlap is not always the right aim — and points to the wider feature-selection literature for those who want to follow the ideas further.

For much more detail, go to this Introduction tab, on the Lab site

If this is the solution, what was the problem?

Eight of the nine clustering categories (29 of the 33 methods) identified by Claude can produce identifiable clusters of texts through replicable, transparent, deterministic processes — but then require a subjective human judgement to name and describe those clusters. That seems weirdly self-contradictory: a rigorous sorting process handed off, at the last step, to an undocumented act of interpretation. The remaining four methods (pile sorting, card sorting, Q Methodology, and Repertory Grid) are subjective throughout, because of their ethnographic orientation, so there is not such a visible internal contradiction.

A second problem concerns the interpretability of the dimensions within which clusters are located. At least eight of the methods rely on abstract derived axes — dimensions that exist and could in principle be examined, but which are statistical composites a non-statistician can't readily interpret (PCA, Factor analysis, ICA, t-SNE, UMAP, MDS, Spectral clustering, and SOM). The difficulty is not just that these axes are hard to read; it is that many of them can't be explained simply to a non-specialist without misrepresenting what they are.

A third problem is unaccountable input choices, such as the number of clusters, topics, classes, or dimensions to be identified. Yes, it is true that every method involves choices, including the Lab. But what matters is is not the number of choices but how defensible each one is. Choices can differ on at least four ways: whether the choice is visible in the output or buried behind it; whether it is checkable against some standard after the fact; whether a poor choice fails visibly rather than silently; and whether it is documented. A choice that is buried, uncheckable, silently failing, and undisclosed is the least defensible; one that is visible, checkable, self-signalling, and recorded is the most.

My assessment

In summary, the Lab addresses the naming problem and the dimensions problem directly, and improves on, without completely escaping, the input-choices problem.

On naming, the dendrogram's tooltips identify how each cluster differs from its sibling by reporting which attributes it has and that the sibling lacks. This is read straight off the matrix; it is systematic and humanly verifiable, not a subjective label imposed after the fact. In addition, the query facility can be used to generate names for the clusters based on their contents, which can then be tested for their reliability through another form of back translation. But it should be remembered that this is a reliability test, not validity test. That is, a back-translation can confirm a name is applied consistently across the texts, but not that it is the right name for what they share. That needs human checking.

On interpretability of dimensions, the dendrogram works at two levels. Its tree structure (a nested classification of groups within groups) is intuitively understandable to most readers without any statistical background; they can see directly which texts joined first and which groups merged later. Its horizontal axis, representing degree of similarity, is less immediately self-explanatory, but it can be explained in one plain sentence: similarity is the degree to which two texts share the same set of attributes. That similarity is computed using one of two standard measures: Hamming or Jaccard distance.

On input choices, as noted above the Lab provides many of these, like others do. The dendrogram's similarity and linkage choices are inherited from its hierarchical-clustering parentage and are similar to others in the degree of technical knowldge involved. But its single most consequential input choice ( the criteria used to rate the texts) is much more defensible, in four ways: the criteria are visible (they are the matrix columns, in plain language), checkable (the reliability test back-translates them and flags unreliable ones), self-signalling (a non-discriminating criterion is filtered out automatically rather than quietly distorting the result), and documented (they are recorded in the provenance statement).

And more recently, with the development of the Selection tab, now also improvable: alternative sets of criteria can be compared on these measures and a better one converged upon, rather than living with whatever a single run produced. The Lab does not remove the input-choices problem; but its strengths lie where it most matters, where the choice can actually be inspected and challenged.The criteria matter most because they define the matrix from which everything else follows: the dendrogram, the sibling comparisons, and the reliability test are all computed from the presence and absence of those attributes, so a choice made well or badly there propagates through every later step.

Returning to the bigger picture

One other criterion for assessing cluster analysis methods is utility or usefulness. A dendrogram is much more useful structure than a simple list of clusters, each of which is different from the other. A dendrogram is a nested set of groups and subgroups, and as such can provide macro, meso, and micro level perspective on the entities that have been grouped.

In addition, the binary branching structure means that at each branching point a pair comparison can be made. Pair comparisons, as distinct from comparisons of large numbers of entities, are most useful when the entities being compared are multidimensional. This is because the mind can only hold so many differences in view at once: comparing two entities that differ on many dimensions is manageable, but comparing ten such entities simultaneously is not, because the number of differences to track at once becomes overwhelming. The fewer dimensions of difference between entities, the more of them that can be compared at any one time. The constraint is cognitive, a limit on human attention, rather than mathematical.

Comparisons of sibling groups in the dendrogram can be past or future oriented in the type of question being asked. Comparison questions can be quantitative, as in which of these is more or less x, or qualitative, as in how is this group different from this group, in terms of x. Comparisons can be retrospective and evaluative, or prospective and planning orientated, and it is this second, forward-looking use that turns the structure from a record of how things are into a guide for what to do next. Most importantly, a dendrogram can be used as a kind of decision tree, not simply seen as a passive classification. They can be used by humans, with and without the assistance of AI models like Claude.

What's next?

When I first wrote this, the open question was optimisation: if the criteria are the most consequential choice, can the Lab help you make that choice well rather than just make it visible? Three things seemed worth optimising — the reliability of the set of criteria, their resolution (their ability to tell the texts apart), and the number of criteria needed to do so. I noted that the last two might be in tension: you can always tell more texts apart by adding more criteria, so a high-resolution set achieved lazily, by piling criteria on, is no great achievement.

The Selection tab, described above, now addresses the first two directly — it puts reliability and resolution in front of you as measures you can read and act on. Looking into the near future, I think the tension between resolution and parsimony will be best handled not by having the tool pick an answer, but by making the trade-off visible: for a set of criteria you have curated, how few of them do you actually need, and what does each additional one buy you in resolution? A set might reach ninety per cent of its discriminating power with four criteria and crawl to a hundred only by adding six more — in which case the four may be the better choice, depending on what those criteria mean. That is a frontier of diminishing returns, and seeing its shape is more useful than being handed a single "optimal" set, because the choice of where to stop on it is a judgement about the criteria's meaning that only the analyst can make. A specification for this exists; building it is the next step.

There is a wider point here worth recording. This problem — the smallest set of things that adequately covers a set of requirements — is a recognised one, the set covering problem, and I have met it before in a different guise: an earlier tool of mine, the Coverage Optimiser, searches for the smallest combination of foresight methods that covers a set of evaluation questions. Choosing the fewest criteria that still tell a set of texts apart is the same shape of problem, transposed. It also has close relatives in Qualitative Comparative Analysis and in the feature-selection methods used in machine learning. The Lab's expected contribution is not to solve the set covering problem automatically, but to make the avalable solution(s) something an evaluator can see, inspect, and decide on.

The patron saint of evaluation?

Footnote: The term "decision tree" can mean two related but distinct things. In machine learning it is an algorithm used to induce prediction rules from data, to classify cases or predict values (the classification and regression-tree family). In decision analysis and operations research, the older sense, it is a decision-support model in which people map out choices, chance events, and their consequences to compare the expected value or utility of competing options. It is this second, human-facing sense I am referring to when suggesting that a dendrogram can be read as a decision tree: not a model that makes the decision, but a structure that lays out the comparisons a decision-maker can make, at macro, meso and micro levels of detail. See Wikipedia, "Decision tree" and "Decision tree learning" and The Decision Lab, "Decision Tree Analysis."

Rethinking how we share evaluation methods

2026-04-07T13:07:00.008+00:00

This year I have been experimenting with a different approach to making evaluation methods more accessible and reusable. Working with Claude AI I've developed a series of mini-apps, each implementing a specific method of analysis as a single self-contained HTM webpage

What makes this approach worth sharing?

1. No dependency on AI services to use them. Once built, the app runs entirely in the browser, anyone can copy the html file and use it independently with no subscription, no login, no connectivity requirements

2. Data stays with the user. There's no server, no database, no cloud storage. Data is uploaded from, and downloaded to, the user's own device as a JSON file. For work involving sensitive information this matters.

3. Surprisingly fast to build with Claude AI. Turning a method of analysis into a working customised tool takes a fraction of the time you might expect, even for non developers.

4. Collective potential. If practioners share analysis methods AND the tools to implement them then others can use them directly or adapt them with AI assistance for the for their own context. The barriers to entry is low.

5. Easy to check for viruses. Being only a single webpage most widely used virus protection software should be able to scan any such mini-apps very quickly and thoroughly.

I've documented a number of examples so far, at mandenews.blogspot.com/2026.

if you're working on evaluation methods and curious as to whether this model fits your context I'm happy to discuss.

Postscript: Some tips on desirable features of a mini-app for the above purposes:

1. App name, that captures its purpose

2. Authors name along with copyright symbol and year. A Comment text in the underlying html should spell out under what conditions the mini-app code can be edited and redistributed

3. Tabs across the top that give access to a discrete set of processing steps, in sequence. All preceded by an Introduction tab that spells out the purpose and workings of the min-app, tab by tab.

4. Import and Export icons for managing json or csv files

5. json file with an example data set that can be imported into .the app, which is explained in the Introduction tab.

...to be continued...

When rankings tell different stories: an introduction to the Rank Explorer

2026-03-28T12:48:00.005+00:00

Caveat Emptor: I delegated the writing of this blog posting to Claude AI (Sonnet 4.6), based on an extended prior dialogue on the subject below, then a summary prompt of what was wanted. My post-production edits were quite limited.

--o0o--

There is a situation that turns up repeatedly in evaluation and research practice that is easy to overlook precisely because it looks like a data analysis problem rather than a methodological one. The situation is this: you have a set of cases that have been ranked on multiple factors, along with a ranking of their outcomes, and you want to understand what the relationship is between the factors and the outcome.

This sounds straightforward. And to a degree it is — you can correlate each factor against the outcome, identify the strongest relationships, and build a composite ranking that aggregates them. These are useful things to do. But they share a common assumption that is not always warranted: that the factors combine additively, that each one contributes independently to the outcome, and that a case that scores well on several factors will therefore tend to score well overall.

That assumption is often wrong. And when it is wrong, the gap between your composite ranking and the actual outcomes is telling you something important.

The problem with additive aggregation

Consider a concrete example. You have 63 local authority areas. You have ranked them on ten factors thought to be associated with population-level physical activity — access to green space, deprivation levels, sports facility density, and others. You have also ranked them on an outcome measure. You build a composite ranking from the factors, correlate it with the outcome, and find it predicts reasonably well — perhaps a Spearman r of 0.75.

That is a decent result. But it is hiding something. Some areas with strong factor rankings are performing poorly on the outcome; others with weak factor rankings are performing well. If you look closely, there are two or three quite different combinations of factors that each seem sufficient, on their own, to predict a good outcome. These are not variations on the same story — they are distinct causal pathways. However an additive composite averages across them, and in doing so obscures the structure.

This is what researchers in the QCA tradition call equifinality — multiple routes to the same outcome. Additive methods cannot find it. A decision tree can.

What the Rank Explorer does

The Rank Explorer is a browser-based companion tool to The Ethnographic Explorer. It is designed to import the rankings data that TEE generates from its Contrast tab, though it will also accept ranking data from any other source in the same CSV format.

[▶ Try the Rank Explorer]

The tool has four analysis tabs. The first three — Individual Factors, Composite Builder, and Scatter Plot — provide standard additive analysis: Spearman correlations for each factor, composite rankings using several aggregation methods (equal-weight, correlation-weighted, stepwise greedy, and exhaustive search), and a scatter plot that visualises how well your composite ranking predicts the outcome, with adjustable classification thresholds.

The fourth tab, Pathway Explorer, is where the configurational logic comes in. It builds an optimal classification tree over your data using exhaustive search: at each node, every available factor and every possible rank cut-off is tested, and the split that best separates high-outcome from low-outcome cases is chosen. The result is a tree that shows which specific combinations of factor ranks distinguish the two groups, displayed in a row-by-level icicle layout that makes the branching structure easy to follow.

The tree is not just a visual. Each leaf node shows which cases ended up there, whether they were correctly classified, and the conditions that led to that grouping. A pathway summary below the tree lists the conditions for each leaf in plain language — for instance, "Active travel infrastructure: rank 5 or better AND Deprivation index: rank 8 or better."

When the gap between additive and configurational results is itself a finding

One of the more useful diagnostics the tool enables is comparing the classification accuracy of the best composite ranking against that of the decision tree. If both are similar, the additive story is probably adequate. If the tree substantially outperforms the composite — reaching, say, 90% or 100% accuracy where the composite only reached 75% — that gap is a finding in itself: the causal structure in the data is better described by conjunctions of conditions than by sums of contributions.

This matters for intervention design. If a high-outcome classification requires both good green space access and good active travel infrastructure (rather than either being substitutable for the other), then improving one without the other may produce no discernible effect. Additive analysis will not surface that conclusion; configurational analysis will.

A note on scale, depth, and selective deepening

The tree-building algorithm uses exhaustive search, which is thorough but computationally intensive. With datasets of 60–70 cases and 10 factors, a depth-3 tree typically builds in a few seconds. Depth 4 or beyond is a different matter: computation time increases steeply, and more importantly, deeper trees on small datasets will often find spurious distinctions — patterns that reflect the quirks of the sample rather than anything real.

The Rank Explorer addresses this through a subgroup analysis feature that may be a modest innovation in decision tree practice for small-N datasets. Once a tree has been built, each leaf node displays an Analyse subgroup → button. Clicking it filters the dataset to only the cases in that leaf and opens a fresh analysis session for that group alone — with all four tabs, including the Pathway Explorer, reconfigured for the subgroup. The outcome cut-off resets automatically to the median of the subgroup's outcome ranks, so the high/low distinction remains balanced within the smaller group.

This allows selective deepening of a specific branch without rebuilding the entire tree at greater depth. If one leaf contains 24 cases that the main tree could not separate further, the subgroup analysis asks a different and analytically legitimate question: within this group, what distinguishes the relatively better-performing cases from the worse ones? The answer applies conditionally — only to cases that reached that leaf — but that conditionality is precisely what makes it interpretable. A banner remains visible throughout the subgroup session as a reminder that the high/low labels are relative to the subgroup, not the full dataset.

For larger or more complex datasets, the stepwise greedy method in the Composite Builder tab is a useful preliminary step: it adds factors to the composite one at a time, selecting whichever remaining factor most improves the correlation with the outcome at each step. The resulting path table shows the marginal contribution of each factor, making it straightforward to identify a smaller subset that carries most of the predictive weight — before running the Pathway Explorer on that reduced set.

Beyond TEE data

The tool is designed as a TEE companion but is not restricted to it. Any CSV with a column of case names and a set of ranking columns will load correctly. Evaluation practitioners who have generated case rankings through other means — expert scoring panels, secondary data, peer comparison exercises — can use the same analytical workflow.

Some framings that could map onto the same tool:

Programme portfolios: rank a set of projects on design-quality dimensions and an outcome measure, then identify which combinations of design features distinguish the most successful from the rest
Organisational assessments: rank a set of partner organisations on capability dimensions, use the tree to find which combinations are most predictive of delivery performance
Cross-country comparison: rank a set of countries or regions on contextual factors alongside a development indicator, and look for the configurational patterns that additive index approaches miss

In each case the structure is the same: cases, factor rankings, an outcome ranking, and the question of what the relationship looks like once you stop assuming it is additive.

An invitation to experiment

The tool is best explored with data you already have. If you have ever built a composite index and felt that it was not quite capturing something you could see in the data, or have had the experience of an outlier case that your model consistently misclassifies, the Pathway Explorer is a reasonable next step. Loading your own data, building a tree at depth 2 or 3, and comparing the pathway classification against your composite should take no more than a few minutes.

I am continuing to develop both tools and would welcome feedback on the approach, the interface, or uses I have not considered.

Accessing the code: The Rank Explorer runs entirely in your browser — no login required, no data is transmitted anywhere. To save your own copy of the code, open the tool, right-click, select View Page Source, copy the entire code, paste it into a text file, rename it to end in .html rather than .txt, and open it in any web browser.

Introducing Rank Order Counterfactuals (ROC)

2026-03-22T09:46:00.000+00:00

A counterfactual is a description of what would have happened, if an intervention had not taken place. The use of randomised control groups is one way to construct a counterfactual. A population of people are randomly assigned to either a control group or an intervention group. Differences in the outcomes of those populations are then compared. If the difference is sufficiently statistically significant then a plausible causal claim can be made that difference in outcomes is because of the intervention.

As might be expected, there are plenty of circumstances when social programs are designed and implemented, where it is simply not practical to organise a randomised control group. In addition, the comparisons that are made will be between averages of the two groups. However, in many social programmes such averages are of limited practical use, because the implementation contexts are so varied and no single “solution” is likely to be applicable. Average effects can still be informative at a high level, but they need to be complemented by methods that take contextual diversity seriously.

I'm currently working with an evaluation team that is examining a large-scale public health programme in the United Kingdom, covering many different locations and involving many different types of local partnerships. But with one common outcome of concern, which is to increase people's physical activity levels in their daily life. In their work the evaluation team is already making use of a causal configurational approach to the understanding of what works for whom in what circumstances. It is finding different configurations of causal conditions across these locations that are associated with changes in activity levels. This approach is consistent with the high level of diversity in locations partnerships and interventions.

But what it does not yet have is a counterfactual, a defensible description of what might have happened in these locations in the absence of this intervention. This is where the idea of a rank order counterfactual becomes relevant. By a rank‑order counterfactual I mean a very specific kind of “what would have happened otherwise.” Instead of trying to predict the exact outcome that would have been achieved in each location without an intervention, we can start by asking a simpler, comparative question: which location would probably have changed more, and which less, if the intervention had never existed? The answer will be in the form of a rank ordering of locations, from those with more to less expected change. That ranking would be constructed based on all available baseline information, trends, and contextual knowledge. This proposed approach falls into a category of counterfactuals known as "logically constructed counterfactuals", and it aligns well with configurational evaluation because it focuses on patterns of relative change across diverse contexts.

A subsequent evaluation of those same locations should also be able to generate a new rank ordering, which is based on observed outcomes. These counterfactual and actual rankings can then be compared, using a scatterplot and correlation measures. The scatterplot is also visually powerful for communication: it lets people see at a glance which locations behave as expected and which ones stand out as surprises. If the intervention had no effect we should see a linear relationship, the observed and counterfactual rankings should be the same. If the intervention had positive, or perhaps even negative effects, this should not happen. We might see various locations which are outliers from that expected trend. When locations we expected to be “natural leaders” did not improve much, and those we expected to be “natural laggards” moved to the top of the league table, that pattern is a signal that the intervention may have been influential, and it gives the evaluation team clear cases where alternative explanations should be probed. The task of the evaluation is then to probe those alternative explanations, not to assume the intervention is the only possible cause. The rankings are not a substitute for theory‑based evaluation; they are a way to make its claims sharper and more testable. The focus on ranking differences can convert a vague theory (“we think these factors matter”) into a concrete, specific prediction about which locations should do better.

The sensitivity of the rank comparison process will depend on the number of ranked items. The more rank positions there are, the more sensitivity there will be to differences in performance, which is good. But, as shown in research on sorting algorithms, the time required to generate a complete sorting, using any of the well-known methods, can be significant. Growing faster than proportionally to the number of items, though far slower than exponential growth. In addition to the extra time required, a rank order counterfactual will require a stronger evidential base where the number of rank positions is greater.

When a large number of locations are involved in an intervention one practical way of addressing this tension is to use a stratified random sample, and to generate the rankings for that sample only. Another approach to managing large numbers of locations is to think of ranked bands of locations rather than individual rankings for each location. What should be of interest, then, are systematic shifts in band membership between the counterfactual and actual observations – for example, locations expected to be in the “low‑change” band turning up in the “high‑change” band in practice.

In this way, rank‑order counterfactuals do not replace theory‑based evaluation, but sharpen it: they turn general expectations about context into explicit, testable predictions about who should have changed most in the absence of the programme. In work which I hope to document in the European Evaluation Society conference later this year I will explain how the use of the hierarchical card sorting process was used to generate argument and evidence based counterfactual rank orderings, and how an LLM was used to support this process.

Making implicit knowledge explicit, contestable and usable: an introduction to The Ethnographic Explorer

2026-03-07T14:40:00.023+00:00

--o0o--

There is a problem that turns up repeatedly in evaluation practice — and in many other fields — that many of us work around rather than solve directly. The problem is this: the people who know most about a programme, a portfolio, or a set of cases often cannot easily say what they know, or why they make the judgements they do.

This is not a failure of intelligence. It is the normal condition of what Michael Polanyi called tacit knowledge — the "we know more than we can tell" problem that is endemic to any field built on experience and judgement. The challenge for evaluation is to find structured ways of drawing it out.

The theoretical anchor: information as difference

The approach behind The Ethnographic Explorer draws on a deceptively simple idea from Gregory Bateson: information is a difference that makes a difference. What we notice, and consider significant, is always defined by contrast — not by properties of things in isolation. A project is "successful" relative to others; a group is "vulnerable" in comparison to other groups in a given context. And some of those differences have noticeable consequences, they "make a difference". Together they can be seen as simple "IF...THEN..." rules.

This suggests that a structured method for eliciting knowledge should be built around comparison — specifically, around asking people to identify and articulate differences, and then to explain what those differences imply. The Hierarchical Card Sort methodology is one way of doing this and is the basis of The Ethnographic Explorer's process of inquiry.

How the Ethnographic Explorer works

There are three linked stages.

In the Sort stage, a respondent is presented with a set of cases — projects, organisations, events, beneficiaries, or any entities they know well. They are asked to divide all the cases into two groups representing the most significant difference between them, from their point of view. Each group is named, the difference is recorded, and then the respondent is asked: "What difference does this difference make?" — a question that surfaces the consequence or implication of the distinction. The process repeats on each subgroup until every group contains a single case. The result is a binary tree: a structured, hierarchical map of the respondent's view of the case set.

The tree is informative in three ways: it reveals the contents of the distinctions the respondent considers important; it identifies the limits of their knowledge (where further differences cannot be found); and it indicates the direction of their attention (where further distinctions could usefully be explored).

In the Compare stage, the facilitator asks comparison questions at each split in the tree. Questions can be in degree ("Which group of cases is more likely to face sustainability problems?") or in kind ("How do these two groups differ in terms of their relationship with local government?"). Degree questions produce a ranking of all cases; kind questions produce descriptive contrasts. Both types build on the structure already revealed by the sort.

In the Contrast stage, any two degree-based rankings are plotted against each other in a scatter plot. The resulting quadrant analysis shows where the two rankings agree (cases high on both, or low on both) and where they diverge (cases high on one and low on the other). Adjustable cutoff sliders allow the facilitator to explore different thresholds, and a Spearman correlation coefficient summarises the overall relationship. Outlier cases — those that diverge most between the two rankings — are often the most analytically interesting. Strong relationships can be cast as potentially useful IF...THEN rules

The Ethnographic Explorer

With substantial coding help from Claude AI, I have been developing a browser-based implementation of this methodology — The Ethnographic Explorer — as a standalone single-file application. This new version supersedes an earlier WordPress-based tool at ethnographic.mande.co.uk, which required a server to run. The new version requires nothing beyond a web browser.

▶ Try The Ethnographic Explorer [When you get there, click on Introduction for guidance on how to explore the tool's functions]

The tool is designed for use in a shared-screen video call with a single respondent, or screened to multiple participants in a workshop setting. The facilitator drives the interface; the respondent provides the knowledge. A typical exercise with 8–12 cases takes between 45 minutes and two hours, depending on range of comparisons made.

A worked example: 12 Largest cities

Seven jscon files are are available alongside the app to allow you to explore the use of the app. The first is simply a list, which you can then sort and compare and contrast. The second is the same list, already sorted by Claude AI at my request, on the basis of " their appeal to international tourists, as seen from 6 different perspectives.

To load either dataset, click Import JSON in the top-right toolbar and select the file (after being downloaded into your computer).

Download 12 Largest Cities (unsorted)

Download 12 Largest Cities - Budget Backpacker view (sorted)

Download 12 Largest Cities - Food and Culture view (sorted)

Download 12 Largest Cities - Safety Concious view (sorted)

Download 12 Largest Cities - Sustainable Tourism view (sorted)

Download 12 Largest Cities - Travel Journalist view (sorted)

Download 12 Largest Cities - Heritage Specialist view (sorted)

What the tool produces

Each exercise generates a structured record of the respondent's knowledge about a set of cases:

a binary sort tree, with each split labelled by the most significant difference and its consequence
one or more named rankings, each derived from a sequence of binary degree judgements working from the root set down through the subsets
one or more descriptive contrasts, capturing how subgroups differ in kind on specified attributes
a scatter plot for any pair of degree rankings, with quadrant analysis and Spearman correlation
a qualitative responses panel, collecting the in-kind descriptions at each split

The full exercise exports as a JSON file, preserving the complete tree structure, all difference descriptions, all responses, and the exercise metadata. Files can be imported to resume exactly where you left off, or shared with a colleague for further analysis.

Beyond evaluation

The underlying structure of the method applies wherever you have a set of cases and a respondent who has differentiated knowledge about them. Some other framings that could be explored with the same tool:

Organisational learning: a team reviews a set of completed projects, identifying the distinctions that, in retrospect, most predicted success or failure
Capacity assessment: a trainer sorts a cohort of staff by their readiness for different kinds of work, making the basis for those judgements explicit and discussable
Stakeholder analysis: a key informant sorts a set of stakeholder organisations, revealing the distinctions they consider most consequential for programme implementation
Policy analysis: a policy analyst sorts a set of interventions by their perceived effectiveness, then compares that ranking against a ranking of their political feasibility

In each case, the sort structure is the same, the comparison logic is identical, and only the cases, the domain framing, and the named differences change.

An invitation to experiment

The tool is best understood by using it. I would encourage anyone engaged in evaluation, learning, or knowledge management work to try loading in a set of cases they know well — even with a rough sort to start — and see what structure emerges. The Compare and Contrast stages are particularly useful for surfacing assumptions that are rarely made explicit in standard reporting.

I am continuing to develop the tool and would welcome feedback on the methodology, the interface, or applications I have not yet considered.

Accessing the code: In addition to trying the app online, you can download a copy of the code and run it independently. Go to the app in yur directory, right-click, select View Page Source, copy the entire code, paste it into a text file, rename the file to end in .html rather than .txt, and open it in a web browser. The tool runs entirely in your browser — no login required, no data sent anywhere.

Optimising method selection: an introduction to the set covering problem — and a tool to help

2026-03-03T08:55:00.013+00:00

Caveat Emptor: I delegated the writing of this blog posting to Claude AI(Sonnet 4.6), based on an extended prior dialogue on the subject below, then a summary prompt of what was wanted. My post production edits were quite limited

--0O0--

There is a class of problem that turns up repeatedly in evaluation planning, and in many other fields, that most of us solve informally and imprecisely. The problem is this: given a set of needs to be addressed, what is the smallest combination of responses that covers all of them?

In evaluation this might be: given a set of evaluation questions I need to answer, what is the minimum combination of methods that gives me adequate coverage across all of them? In public health it might be: which combination of clinics ensures every neighbourhood has access to at least one? In software testing: which set of test cases exercises every code path? In logistics: which depots can serve every delivery zone?

These are all instances of what computer scientists call the Set Covering Problem — a well-studied problem in combinatorial optimisation that dates back to the 1970s. The formal version asks: given a collection of sets, find the smallest sub-collection whose union covers all elements of a target universe.

Why it matters for evaluation practice

When designing an evaluation, practitioners face exactly this structure. We have a list of evaluation questions (or dimensions of quality, or stakeholder concerns), and a repertoire of methods — each of which can address some but not all of those questions to a satisfactory level. Choosing methods one at a time, based on familiarity or habit, rarely produces the most efficient combination. We either end up with redundant overlaps in some areas and blind spots in others, or with far more methods than the budget or timeline can support.

A more systematic approach asks: which combination of methods is both complete (covers all questions at an adequate level) and minimal (uses as few methods, and as little resource, as possible)?

How the optimisation works

For a small number of methods and questions, you could check all possible combinations by hand. But the number of combinations grows exponentially — with ten methods, there are over a thousand possible subsets to evaluate. This is where an algorithm helps.

The simplest approach is a greedy algorithm: at each step, pick the method that covers the most currently-uncovered questions, then repeat until everything is covered. This is fast and usually finds a good solution, but not necessarily the best one.

A more thorough approach — exhaustive search — systematically checks all combinations up to a specified size and returns every minimal solution. This is slower but reveals the full landscape of equally-good options, which is often more useful than a single answer, particularly when cost or other practical constraints come into play.

The Coverage Optimiser

With substantial coding help from Claude AI, I have been developing a small browser-based tool — the Coverage Optimiser — that applies both approaches to a user-defined matrix of methods and question types.

▶ Try the Coverage Optimiser [When you get there click on Introduction, to get suggestions on how to explore the functions of the app]

The default example matrix uses ten foresight methods (Scenario Planning, Delphi, Horizon Scanning, and others) rated against five question types — Descriptive, Valuative, Explanatory, Predictive, and Prescriptive — at HIGH, MEDIUM, or LOW in terms of their usefulness in building a futures perspective into an evaluation. The tool finds the minimal combinations of methods that achieve the desired coverage level across all question types.

Each solution is displayed with:

the methods involved, each with its cost rating
a question-by-question coverage check
the total cost of the combination
an overlap score — the number of questions covered by more than one method in the solution

Overlap is worth attending to: it indicates redundancy, which in evaluation terms means resilience. If one method proves impractical in the field, a solution with higher overlap is more likely to remain viable.

The matrix itself is fully editable in a dedicated Matrix Editor tab. You can rename methods and question types, adjust HIGH/MEDIUM/LOW ratings on a choosen criteria, set importance weights for each question type (0–10), set cost ratings for each method (0–10), sort rows by any column, import and export as JSON, and print or save the matrix or results as PDF.

Note: The Coverage Optimiser runs entirely in your browser — no login required, no data sent anywhere.

Beyond methods and questions

The tool is not limited to foresight methods or evaluation questions. The underlying logic applies wherever you have a set of options and a set of requirements, and where each option partially addresses some requirements. Some other framings that could be loaded into the same tool:

Solutions × Problems: which combination of policy interventions covers the broadest range of identified problems?
Stakeholders × Information needs: which combination of engagement activities ensures all key stakeholder groups have their core information needs met?
Data sources × Indicators: which combination of data collection instruments covers all required indicators, at minimum cost?
Partners × Geographic areas: which combination of implementing partners ensures all target districts are reached?

In each case the matrix structure is the same, the optimisation logic is identical, and only the labels and rating criteria change.

An invitation to experiment

The tool is best understood by using it. I would encourage anyone planning a multi-method evaluation, or foresight exercise, to load in their own methods and questions — even with rough ratings to start — and see what combinations emerge. The exhaustive search mode is particularly useful for revealing that several equally-minimal combinations exist, which opens up a more deliberate conversation about which is preferable given cost, feasibility, or complementarity.

I am continuing to develop the tool and would welcome feedback on the matrix structure, the rating scales, or applications I have not yet considered.

Those who have used EvalC3 will find a family resemblance here: both tools use systematic search to find efficient combinations — EvalC3 searching for attribute combinations that predict specific individual outcomes, the Coverage Optimiser searching for method combinations that cover multiple requirements. The underlying computational logic is related, even if the problems look different on the surface.

Accessing the code: In addition to trying out the app, as it already exists online, you can also download a copy of the code and have it working independently on your own website. Go to the app online, right click your mouse, select View Page Source, copy the entire code, paste it into a txt file, rename the text file to end in .html not .txt, click on that file in a directory to open it up in a web browser. Simples, yes?

Exracting additional knowledge and performance from a configurational model that already has wide coverage

2025-12-25T01:38:00.016+00:00

A decision tree algorithm, as available within EvalC3, can generate a classification tree (a set of predictive models) of the kind shown here.

Some of the models (each branch is a model) are very detailed (i.e. has lots of attributes) and have narrow coverage. Such as HasQuotas+NotPost Conflict Situation+High Level of Human Development+ Low Womens Status = Low levels of womens representation in Parliament, which covers two cases (Senegal, Tanzania).

Others are quite simple, with only two or three attributes and can have much wider coverage. Such as HasQuotas+ IsPostConflict = High levels of womens representation in Parliament, which covers two six cases (Burundi, Ethiopia, Mozambique, Namibia, South Africa, Uganda)

These wide coverage models may have unexplored potential, in the form of unexploited information content within the cases they cover. The raw (i.e. numerical) outcome data for the cases they cover only can be examined and recalibrated i.e re-dichotomised into two new sub-groups representing relatively higher versus lower outcome values within that set only.

A new configurational analysis can then focus on that sub-set of cases to see if (a) any of the pre-existing attrubutes could predict membership of the two sub-groups, or (b) if any additional attributes, based on other knowledge of these cases, could do so.The ability to predict such finer grained performance differences would be a significant improvement.

This analytic step is a complementary move to that known as "pruning", where the removal of a mode attribute improves coverage, at the cost of precision. Here an extra attribute is sought that will improve precision but at the cost of coverage. Perhaps it could be called "grafting"...

Postscript: But how significant will this addition to the model be? If, as above, there are six cases involved, there are 2^6 possible binary groupings of these case i.e 32. So any one grouping of two sets of cases has a 1/32 or 3.125% chance of occuring randomly (if the cases are causally independent).

Objectives as data: The potential uses of updatable outcome targets

2025-12-02T22:44:22.363+00:00

The context

A specialist agency is funding more than 40 different partner organisations, each working in a different part of the country but with the same overall objective of increasing people's levels of physical activity (because of the positive health consequences). These partners are often working with quite different communities, and all have substantial degree of independence about how they work towards the overall objective.

Some agency representatives have asked about the nature of the target that program as a whole is working towards, and have emphasised how essential it is that there be clarity in this area. By target they mean an actual number. Specifically the percentage of people self-reporting that they achieve a certain level of physical activity each week, as identified by an annual survey that is already underway and will be repeated in future.

Possible responses

In principle it would be possible to set a target for the proportion of the population reporting being physically active. Such as 75%. But it would be very hard to identify an optimal target percentage, given the diversity of partner localities, and the communities within these.

Relative targets may be more appropriate. Such as a 25% increase in reported activity levels. Especially if partners were each asked to identify what they think are achievable percentage increases in their own localities within the next survey period. This estimation would take place in a context where these partners already have experience working in those locations, identifying some of the things that work and dont work. My hypothesis, yet to be tested, would be that these partners will make quite conservative estimates. And if so, this might come as some surprise to the donor and perhaps lead to some revision of their own expectations

Taking this idea further, partners could be periodically asked if they wanted to adjust their expectations upwards or downwards , of the change that could be achieved - in the time remaining in the interventions lifespan. Subject to being able to explain the rationale for doing so. My second hypothesis is that this number, and commentary, could be a valid and useful form of progress reporting in its own right.

Making sense of the responses

An assessment of overall progress over longer time scale would need to consider both the scale of ambitions and the extent of their achievement. These can't be combined into one number based on a simple formula because any such number could be achieved by adjustment of expectations and or performance. However it could be usefully represented by a scatterplot, with data points reflecting each of the partners, of the kind shown below.

The location of partners in different quadrants suggests different implications about how the different partners should be managed

High ambition/low achievement: May need additional support, capacity building, or problem-solving
Low ambition/low achievement: May need fundamental partnership restructuring or exit considerations
High ambition/high achievement: Candidates for scaling, sharing learning, reduced support intensity
Modest ambition/high achievement: Opportunities to stretch ambitions

This framework also provides plenty of potentially useful analytic questions

Are ambitions increasing or decreasing?
Is the gap between expected and actual narrowing or widening?
For a given level of actual achievement did differences in expectations have any role or consequences
For a given level of expected change what might explain the differences in the partners actual achievements
How do individual partners positions within this matrix change over time? Are there distinct types of trajectories and how can these differences be explained?

In summary

A single numerical value based on the data in this matrix will provide a meaningless simplification.

In contrast, a scatterplot visualisation can generate multiple potentially useful perspectives.

It is more useful to see targets as necessarily malleable responses to changing conditions, than as unarguable reference points.

Postscript

There is a type of reinforcement learning algorithm known as Temporal Difference Learning (TDD), that embodies a very similar process. It is described as "a model-free reinforcement learning method that learns by updating its predictions based on the difference between future predictions and the current estimate". Model-free means it has no built in model of the world it is working in.

When implemented as a human process it is vulnerable to gaming, because the agents (humans) are aware of the system's mechanics, unlike the neural networks or simplified agents typically used in computational TD learning. But one adaptation, suggested by Gemini AI, is to "reward partners not just for the +/-gap, but for the accuracy of their final predictions over multiple cycles". Relatively higher accuracy, over multiple time periods, might be indicative of potentially generalisable / replicable delivery capacity, usable beyound the current context.

On two types of Theories of Change: Temporal and atemporal, and how they might be bridged

2024-06-18T14:40:00.008+00:00

There are two quite different ways of representing theories of change – of the kind that might be useful when planning and monitoring development programmes of one kind or another.

The first kind is seen in representational devices such as the Logical Framework, Logic Models and boxes-and-arrows type diagrams. These differentiate events according to their location at different points over time, taking place between the initial provision of funding, its allocation and use and then it's subsequent effects and final impacts. These are temporal models.

The second kind, seen much less often, are seen in the analyses generated by Qualitative Comparative Analysis (QCA) and simple machine learning methods known as Decision Trees or Classifiers. Here the theory is in the form of multiple configurations of different attributes that are associated with desired outcome, and its absence. Those attributes may be of the intervention and/or its context. The defining feature of this approach is the focus on cases and differences between cases, rather than different points or periods in time. These cases are often geographical entities, or groups or persons, which have some persistence over time. They are effectively atemporal models.

Each of these approaches have their own merits. Theory of change which describes the sequence of expected events over time and how they relate to each other is useful for planning, monitoring and evaluation purposes. But it runs the risk of assuming a homogeneity of effects across all locations where it is is implemented. On the other hand, a QCA-type configurational approach helps us identify diversity in contexts and implementations, and its consequences. But it may not have any immediate management consequences, about what needs to be done when.

One of my current interests is exploring the possibility of combining these two approaches, such that we have theories of change that differentiate those events over time, while also differentiating cases across space where those events may or may not be happening.

One paper which I've just been told about is exploring these possibilities, as seen from a QCA starting point:Pagliarin, S., & Gerrits, L. (2020). Trajectory-based Qualitative Comparative Analysis: Accounting for case-based time dynamics. Methodological Innovations, 13. In this paper the authors introduce the innovative idea of cases as different periods of time in the same location, where each of those subsequent periods of time may have various attributes of interest present or absent, along with an outcome of interest being present or absent. This approach seems to have potential for enabling a systematic approach to within-case investigations complementing what might have been prior cross-case investigations. There is the potential to identify specific attributes, or combinations of these, which are necessary or sufficient for changes to take place within a given case.

Somewhat tangentially...

The same paper reminded me reminded me of some evaluation fieldwork I did in Burkina Faso in 1992, where I was interviewing farmers about the history of their development of a small market garden using irrigation water obtained from a nearby lake. Looking back at the history of the market, which I think was about six years old at the time, I asked them to identify the most significant change that had taken place during the period of time. They identified installation of the water pump in year 198?, and pointed out how it expanded the scale of their cultivation thereafter. I can remember also asking, but with less recall of what they then said, follow-up questions about the most significant change that it taken place in each smaller time period either side of that event, and then its consequences. I was in effect asking them to carve up the history of the garden into segments, and sub-segments, of time not defined by calendar, but by key events – each of which had consequences. These were in effect temporal "cases". Each of these had a configuration of multiple attributes, i.e. being attributes of the nested set of time periods that it belonged to. Associated with each of these were differenting judgements about the the productivity of the market garden. But with our team's time being short supply, I never got the opportunity to gather a full data set, so to speak.

Another of my current interests, prompted by the above conjectures, is the possible use of specific form of Hierarchical Card Sorting (HCS) as a means introducing a temporal element into case-based configurational analysis. The HCS process generates a tree structure of nested binary distinctions between cases. It is concievable that different broad criteria could be introduced for the type of differences being identified at each level of the branching structure. For example, at the top level the "most significant difference" being sought could be specified as being "in terms of funding received", then at the next level, "in terms of outputs generated" , and so on (Criteria 1,2,3 etc in Figure 1 below) .

Figure 1 below

Developing and using a Configurational Theory of Change within an evaluation

2024-04-24T12:04:00.030+00:00

Figure 1

Ho-hum, yet another evaluation brand being promoted in an already crowded marketplace.FFS...

Yes, I think this reaction is understandable, but I think there is something here captured under this title (Developing and using a Configurational Theory of Change... ) which has potential value. I will try to explain...

Many evaluators make use of theories of change, as part of a theory-based approach to evaluation. Many theories of change are described in some type of diagrammatic form. And a typical feature of those diagrams is their convergent nature. That is, they start of with a range of different types of inputs and activities which follow various causal pathways towards a limited number of final outcomes.

This image is almost the complete opposite of what happens in actual practice on the ground. Financial inputs come from a limited number of sources, these become available to a small range of partners who carry out their own range of activities, in a variety of different locations each with their own populations, including those intended and not intended to be affected. This description is of course a simplification, but it applies to many development aid programme designs. The point I'm making here is that this in-reality process is not convergent it is divergent! It seems like the diagrammatic theories of change I have described are a type of Procrustean bed.

This blog posting has been most immediately prompted by a report I have just reviewed on potential evaluation strategies for a large national level climate finance strategy (CFS). The theory of change describes multiple causal pathways connecting the initial provision of government finance through to four expected types of expected impacts. With two of these causal pathways alone the number of projects being funded is in the hundreds. The report struggled with the issue of how to measure the expected impacts given the scale and likely diversity of events on the ground. And the corresponding challenge of how to sample those projects. Part of my diagnosis of the problem here was the evaluation team's measurement-led approach. And the weakness of the conceptual framework i.e. the incapacity of the theory of change to capture the diversity of what was taking place.

Describing the alternative to my client is now my challenge. I think the alternative has two parts. Firstly, one should start at the beginning, where the money becomes available, and then follow the money (and the people responsible) as it gets distributed according to its intended purposes. If things are not happening as expected early on in this process then this affects expectations of what might and might not be observable later on in the form of 'outcomes' or 'impacts'. Put crudely, there is no point trying to observe the impact of something that has not yet been delivered. And in the case of strategies like the CFS, a large part of success can simply be gettting the money where it should be spent.

Secondly, as money is distributed from a central fund, decisions are going to be made about how it should be parcelled out in different amounts for different purposes through different institutions. Each time that happens the decisions that have been made about how to do this are hopefully not random. Evaluating how those decisions were made may not necessarily be all that useful, because often there will be opaque mini, meso and macro political processes involved. But the announced decisions may include some intentionally explicit expectations about the official purposes of different allocations. Interviews those responsible for those allocations might also elicit more informal and more current expectations about what might be the short and longer terms effects of some of these allocations, when compared to others.The point I am emphasising here is that sometimes we can come to evaluative judgements not through the use of any overriding predetermined criteria, but by using a more inductive process, where we compare one option to another. This is an excuse for me to quote Marx (G):

Friend says to Marx – 'Life is difficult'.

Marx replies to friend – 'Compared to what?'

This type of inductive comparative evaluation doesn't have to be completely free form. It is conceivable for example that we could look at two tranches of government climate finance funding and ask (those with proximate responsibilities for that funding) what difference there might between those blocks of funding in terms of how each might meet one or more of the OECD criteria (These range in their concerns from the more immediate issues of coherence and efficiency to later concerns with effectiveness and impact). Respondents answers in the form of expectations can be seen as mini theories a.k.a. hypotheses that then might be testable through the gathering of relevant data.

Before these questions can be posed the cases that are going to be compared would need to be identified. The 'cases' in this example would be particular blocks of funding. Further along the implementation process the cases could be partners who are receiving funding, or activities that those partners implementing, or communities those activities are directed towards. Nevertheless, at any point along this chain there is still a challenge, which is how to select cases for comparison. For example, if we are looking at a particular budget document which distributes funding into multiple purpose categories we will be faced with the question of which of these categories to compare.

One way forward is to let the interviewed person decide, especially if they have responsibilities in this area. Using hierarchical card sorting (HCS) the interviewer starts with a request, which is phrased like this: 'What is the most significant difference between all these budget categories in terms of how they will achieve the objectives of the climate Finance strategy? Please sort the budget categories into two piles according to this difference and then explain it to me". Having identified ppiles of types of cases that can be compared the respondent can then be asked for details about their expectations of the cases in one pile versus the other (See FN1).The same question can then be reiterated by focusing on each of those two piles in turn and getting the respondent to break them into two smaller sub- piles. When their answers are followed by explanations this will help differentiate expectations in further detail.

Figure 2 (click on to enlarge)

Figure 2 shows the results of such an exercise, where the respondents were NGO staff responsible for the development and management of a portfolio of projects. They were asked to sort the projects into two piles according to "What they saw as the most significant difference between the projects, of a kind that would make a difference to what they could achieve". Their choices generated the tree structure. They were then asked to make a series of binary choices at each branching point, indentifying which of the two types of projects described there that "they expected to be most successful, in terms of the extent to which they will contribute to the achievement of the overall objectives of the portfolio" . Their choices are shown by the red links. In this diagram their responses have been sorted such that the preferred red option is always shown above the non-preferred option. The aggregate result is a ranked set of 8 types of projects, with the highest rank (1) at the top. Each of these types is not an isolated category of its own, but part of a configuration that can be read along each branch, from left to right.

Here are some of the type descriptions and the reasons why one versus the other was selected most likely to contirbute to the portfolio objectives. Further discusison would be needed to esytablish how the presence/absence of these characteristics could be identified on the ground.

Wider focus	Aim to influence wider policy and environment, and have more sustainable and wider impact beyond children and their families.
	Likely to be more successful: Because it will have a wider reach and be more sustainable

Local focus	More hands on work with children on a day to day basis. Impact may be sustained but it will be limited to children and their families.
	Likely to be less successful:

Locally driven	Partner and the projects are locally rooted, driven by local needs and priorities. They are more likely to “get it right”. They can’t walk away when Comic Relief funding ends. More likely to be sustainable.
	Likely to be more successful: More embedded in the context, will outlast the project, be more responsive.

UK driven	UK driven projects, almost sub-contracting. They have a set end-point.
	Likely to be less successful:

There is a larger question here of course that also relates to sampling. Who are you going to interview in this way? The suggestion above was 'to follow the money '. In other words, to follow lines of responsibility and interview people about the domains of activity they are responsible for, using HCS as a means of structuring the discussion. There is a strategy choice here between what is known as a breadth-first search versus a depth-first search strategies. From a given point in a flow of funds (and of responsibilities) there can be distributions going in different directions, each of which could all be explored. Following all of these is a form of breadth-first search. Alternatively the focus could be just on one of those developments, and following the subsequent distribution of funding and responsibility further down one (or few) line. This is a form of depth-first search. Which of those search strategies to pursue is probably a matter to be decided by the evaluation client. But may also need to be adaptive, informed by what was found by the evaluation team in prior interviews.

Courtesy Jacky Lieu: Comparison of Breadth-First Search and Depth-First Search: Understanding Their Methods and Uses

But what about aggregation?

If you followed my suggested approach, the closer you got to the people whose lives were of final/main concern, the small the segments of all the funding you would be looking at. These would be more comparable than when looking at as part of a larger group, with more customised context specific assessments of expected and actual impact. But how would you / the evaluation team then be able to make any overall statement about the strategy as a whole?

The way forward is to think of performance measurement in slightly different terms, than just using a simple indicator based measure. Imagine a scatter plot, with one dimension X describing relative i.e. ranked expectations of achievement and the other dimension Y describing ranked actual/observed/assessed achievements. The entities in the scatter plot are the groups of cases in the smallest available sub-categories that were developed. Their rank position, relative to each other, is evident when all the binary assessments of expected performance are generated through the process described above. See here for more on how this is done. The scatter plot can in turn be summarised in at least two different ways: using a measure of rank correlation (or how achievement relates to expectations) and using Classification Accuracy, if and when a minimum rank position of achievement is identfiied. Equally importantly, qualitative descriptions can be given of cases that exemplify performance that most meets expectations, and the reverse, along with positive and negative deviants (outliers).

What we could end up with is a tree structure documenting multiple routes to both high and low performance, implemented in varyingly different contexts (describable at different levels of scale).

Other scatter plot designs are more relevant to assessments of strategies. The ranking generated by Figure 2 was plotted against the age of the projects and their grant size, which might be expected to be influenced by the contents of a funding strategy. Neither of these two measures showed any relationship to perceived strategic priorities!

To be continued....

PS1: When asking about expected effects of one type of allocation versus another, it may make sense to encourage a focus on more immediately expected effects first, and then later ones. They may be more likely, more easily articulated and more evaluable.

PS2: Hughes-McLure, S. (2022). Follow the money. Environment and Planning A: Economy and Space, 54(7), 1299–1322. https://doi.org/10.1177/0308518X221103267

Using the Confusion Matrix as a general-purpose analytic framework

2023-12-22T06:18:00.024+00:00

Background

This posting has been prompted by work I have done this year for the World Food Programme (WFP) as member of their Evaluation Methods Advisory Panel (EMAP). One task was to carry out a review, along with colleague Mike Reynolds, of the methods used in the 2023 Country Strategic Plans evaluations. You will be able to read about these, and related work, in a forthcoming report on the panel's work, which I will link to here when it becomes available.

One of the many findings of potential interest was: "there were relatively very few references to how data would be analysed, especially compared to the detailed description of data collection methods". In my own experience, this problem is widespread, found well beyond WFP. In the same report I proposed the use of what is known as the Confusion Matrix, as a general purpose analytic framework. Not as the only framework, but as one that could be used alongside more specific frameworks associated with particular intervention theories such as those derived from the social sciences.

What is a Confusion Matrix?

A Confusion Matrix is a type of truth table, i.e., a table representing all the logically possible combinations of two variables or characteristics. In an evaluation context these two characteristics could be the presence and absence of an intervention, and the presence and absence of an outcome. An intervention represents a specific theory (aka model), which includes a prediction that a specific type of outcome will occur if the intervention is implemented. In the 2 x 2 version you can see above, there are four types of possibilities:

The intervention is present and the outcome is present. Cases like this are known as True Positives
The intervention is present but the outcome is absent. Cases like this are known as False Positives.
The intervention is absent and the outcome is absent. Cases like this are known as True Negatives
The intervention is absent but the outcome is present. Cases like this are known as False Negatives.

Common uses of the Confusion Matrix

The use of Confusion Matrices is most commony associated with the field of machine learning and predictive analytics, but it has much wider application. These include the fields of medical diagnostic testing, predictive maintenance, fraud detection, customer churn prediction, remote sensing and geospatial analysis, cyber security, computer vision, and natural language processing. In these applications the Confusion Matrix is populated by the number of cases falling into each of the four categories. These numbers are in turn the basis of a wide range of performance measures, which are described in detail in the Wikipedia article on the Confusion Matrix. A selection of these is described here, in this blog on the use of the EvalC3 Excel app

The claim

Although the use of a Confusion Matrix is commonly associated with quantitative analyses of performance, such as the accuracy of predictive models, it can also be a useful framework for thinking in more qualitative terms. This is a less well known and publicised use, which I elaborate on below. It is the inclusion of this wider potential use that is the basis of my claim that the Confusion Matrix can be seen as a general-purpose analytic framework.

The supporting arguments

The claim has at least four main arguments:

The structure of the Confusion Matrix serves as a useful reminder and checklist, that at least four different kinds of cases should be sought after, when constructing and/or evaluating a claim that X (e.g. an intervention) lead to Y (e.g an outcome).

True Positive cases, which we will usually start looking for first of all. At worst, this is all we look for.
False Positive cases, which we are often advised to do, but often dont invest much time in actually doing so. Here we can learn what does not work and why so.
False Negative cases, which we probably do even less often. Here we can learn what else works, and perhaps why so,
True Negative cases, because sometimes there are asymmetric causes at play i.e not just the absence of the expected causes

The contents of the Confusion Matrix helps us to identify interventions that are necessary, sufficient or both. This can be practically useful knowledge

If there are no FP cases, this suggests an intervention is sufficient for the outcome to occur. The more cases we investigate , without still finding a TP, the stronger this suggestion is. But if only one FP is found, that tells us the intervention is not sufficient. Single cases can be informative. Large numbers of cases are not aways needed.
If there are no FN cases, this suggests an intervention is necessary for the outcome to occur. The more cases we investigate , without still finding a FN, the stronger this suggestion is. But if only one FN is found, that tells us the intervention is not necessary.
If there are no FP or FN cases, this suggests an intervention is sufficient and necessary for the outcome to occur. The more cases we investigate, without still finding a TP or FN, the stronger this suggestion is. But if only one FP, or FN is found, that tells us that the intervention is not sufficient or not necessary, respectively.

The contents of the Confusion Matrix help us identify the type and scale of errors and their acceptability. FP and FN cases are two different types of error that have different consequences in different contexts. A brain surgeon will be looking for an intervention that has a very low FP rate, because errors in brain surgery can be fatal, so cannot be recovered. On the other hand, a stockmarket investor is likely to be looking for a more general purpose model, with few FNs. However, it only has to be right 55% of the time to still make them money. So a high rate of FPs may not be a big concern. They can recover their losses through further trading. In the field of humanitarian assistance the corresponding concerns are with coverage (reaching all those in need, i.e minimising False Negatives) and leakage (minimising inclusion of those not in need i.e False Positives). There are Confusion Matrix based performance measures for both kinds error and for the degree that both kinds of error are balanced (See the Wikipedia entry)
The contents of the Confusion Matrix can help us identify usefull case studies for comparison purposes. These can include

Cases which exemplify the True Positive results, where the model (e.g an intervention) correctly predicted the presence of the outcome. Look within these cases to find any likely causal mechanisms connecting the intervention and outcome. Two sub-types can be useful to compare:

Modal cases, which represent the most common characteristics seen in this group, taking all comparable attributes into account, not just those within the prediction model.
Outlier cases, which represent those which were most dissimilar to all other cases in this group, apart from having the same prediction model characteristics

Cases which exemplify the False Positives, where the model incorrectly predicted the presence of the outcome.There are at least two possible explanations that can be explored:

In the False Positive cases, there are one or more other factors that all the cases have in common, which are blocking the model configuration from working i.e. delivering the outcome
In the True Positive cases, there are one or more other factors that all the cases have in common, which are enabling the model configuration from working i.e. delivering the outcome, but which are absent in the False Positive cases

Note: For comparisons with TPs cases, TP and FP cases should be maximally similar in their case attributes. I think this is called MSDO (most similar, different outcome) based case selection

Cases which exemplify the False Negatives, where the outcome occurred despite the absence the attributes of the model. There are three possibilities of interest here:

There may be some False Negative cases that have all but one of the attributes found in the prediction model. These cases would be worth examining, in order to understand why the absence of a particular attribute that is part of the predictive model does not prevent the outcome from occurring. There may be some counter-balancing enabling factor at work, enabling the outcome.
It is possible that some cases have been classed as FNs because they missed specific data on crucial attributes that would have otherwise classed them as TPs.
Other cases may represent genuine alternatives, which need within-case investigation to identify the attributes that appear to make them successful

Cases which exemplify the True Negatives, where the absence the attributes of the model is associated with the absence of the outcome.

Normally this are seen as not being of much interest. But there may cases here with all but one of the intervention attributes. If found then the missing attribute may be viewed as:

A necessary attribute, without which the outcome can occur
An INUS attribute i.e. an attribute that is Insufficient but Necessary in a configuration that is Unnecessary but Sufficient for the outcome (See Befani, 2016). It would then be worth investigating how these critical attributes have their effects by doing a detailed within-case analysis of the cases with the critical missing attribute.

Cases may become TNs for two reasons. The first, and most expected, is that the causes of positive outcomes are absent. The second, which is worth investigating, is that there are additional and different causes at work which are causing the outcome to be absent. The first of these is described as causal symmetry, the second of these is described as causal asymmetry. Because of the second possibility is worthwhile paying close attention to TN cases to identify the extent to which symmetrical causes or asymmetrical causes are at work. The findings could have significant implications for any intervention that is being designed. Here a useful comparision would be between maximally similar TP and TN cases.

Resources

Some of you may know that I have built the Confusion Matrix into the design of EvalC3, an Excel app for cross-case analysis, that combines measurement concepts from the disparate fields of machine learning and QCA (Qualitative Comparative Analysis). With fair winds. this should become available as a free to use web app in early 2024, courtesy of a team at Sheffield Hallam University. There you will be able to explore and exploit the uses of the Confusion Matrix for both quantative and qualitative analyses.

Beyond summarisation by AI and/or editors- Readers can now interrogate full transcripts of meeting discussions

2023-10-28T13:34:00.002+00:00

Over the last two months, a small group of us have been managing a MSC Monthly Online Gathering. In each meeting we have recorded the discussions, then generated a transcript, both using Otter.AI. Then I have used Claude AI, to generate a one-page summary of each discussion. That itself seems likely to be useful to both attendeess and non-attendees. (Though I have yet obtain feedback on this meeting output). You can view two AI summaries of discussions in the October meeting, here:

https://mande.co.uk/wp-content/uploads/2023/10/18th-October-MSC-AM-Rick.pdf

https://mande.co.uk/wp-content/uploads/2023/10/18th-October-PM-Konny.pdf

But why not jump ahead and give people more than a simple feedback opportunity. Let's enable them to question the full text of the transcript, in their own individual way, albeit after being informed about the overall topics covered during the discussion via the AI summaries above. This is now possible using a third party app known as Pickaxe. Here you can design an AI prompt that can then be made publically usable, preloaded with a given discussion transcript.

Here is a link to the two very simple Pickaxe public prompts I have developed that you can now use to interrogate the two discussions.

AM session
https://beta.pickaxeproject.com/axe?id=Interrogate_the_transcript_of_a_meeting_WA76O
PM session
https://beta.pickaxeproject.com/axe?id=Explore_issues_discussed_in_the_18th_October_MSC_Monthly_Online_Gathering_PM_session_1OPYP

You can ask follow up questions, click on "Go to Chat"

If you try these out, I will get feedback, in the form of a visible record of how you used it. You could also provide feedback on this experience, using the Comment function below

Give it a go, now...!

Postscript 31 October

I think the performance of Pickaxe on this task is poor, compared to that of Claue AI on the same task. I will be disabling this implementation in the next day or so

Evaluating thematic coding and text summarisation work done by artificial intelligence (LLM)

2023-08-31T13:30:00.247+00:00

Evaluation is a core part of the workings of artificial intelligence algorithms. It is something that can be built in, in the shape of specific segments of code. But it is also an additional human element which needs to complement and inform the subsequent use of any outputs of artificial intelligence systems.

If we take supervised machine learning algorithms as one of the simpler forms of artificial intelligence, all of these have a very simple basic structure. Their operations involve the reiteration of search followed by evaluation. For example, we have a dataset which describes a number of cases, which could be different locations where a particular development intervention is taking place. Each of these cases have a number of attributes which we think may be useful predictors of an outcome we are interested in. And in addition, some of those predictors (or combinations thereof) might reflect some underlying causal mechanisms which would be useful for us to know about. The simplest form of machine learning will involve what is called an exhaustive or brute force search of each possible combination of those attributes (defined in terms of their presence or absence, in this simple example). Taking one combination at a time, the algorithm will evaluate whether it predicted the outcome or not, and then store that judgement. Reiterating that process, it will then compare the next judgement to this earlier judgement and replace that earlier judgement if the new one is better. And so on until all possible combinations have been evaluated and compared to previous judgement. In more complex machine learning algorithms involving artificial neural networks the evaluation and feedback processes can be much more complex, but the abstract description still fits.

What I'm interested in talking about here is what happens outside the block of code that does this type of processing. Specifically, with the products that are produced and how we humans can evaluate its value. This is territory where a lot of effort has already been expended, most noticeably on the subject of algorithmic fairness and what is known as the alignment problem. These could be crudely described as representing both short and long-term concerns respectively. I won't be exploring that literature here, interesting and important as it is.

What I will be talking about here is two examples of my own current experiments with the use of one AI application known as Claude AI, used to do some forms of qualitative data analysis. In the field that I work in, which is largely to do with international development aid programs, a huge amount of qualitative data i.e text is generated and I think it is fair to say that its analysis is a lot more problematic than when we are dealing with many forms of quantitative data. So the arrival of large language model (LLM) versions of artificial intelligence appears to offer some interesting opportunities for making some usable progress in this difficult area.

The text data that I have been working with as been generated by participants in a participatory scenario planning process, carried out using ParEvo.org, and implemented by the International Civil Society Centre in Germany this year. The full details of that exercise will be available soon in a ICSC publication. The exercise generated a branching tree structure of storylines about the future, built with 109 paragraphs of text, contributed by 15 participants, over eight iterations.What I will be describing here concerns two types of analysis of that text data.

Text summarisation

[this section has been redrafted] The first was a text summarisation task, where I asked Claude AI to produce one sentence headline summaries of each of these 109 texts. Text summarisation is a very common application of LLMs. This it did quickly, as usual, and the results looked plausible. But but by now I had also learned to be appropriately sceptical and was asking myself how 'accurate' these headlines were. I could examine each headline and its associated text, but this would take time. So I tried another approach.

I opened up a new prompt window in Claude AI and uploaded 2 files. One containing the headlines, and the other containing each of the 109 texts preceded by an identification number. I then asked Claude AI to match each headline with the text that it best described, and to display the results using the ID number of the text (rather than its full contents) and the predicted associated headline. This process has some similarities with back translation. What I was interested in here was how well it could reassign the headlines to their original texts. If it did well this would give me some confidence in the accuracy of its analytic processes, and might obviate the need for a manual check up of the headlines' fit with content.

My first attempt was a clear failure, with classification accuracy of 21%, being far worse than chance. On examination this was caused by the way I had formated the uploaded data. The second attempt, using two separated data files, was more successful This time the classification accuracy was 63%. Given that the 27% error could occur at two stages (headline creation and headline matching) it could be argued that the classification error was more like half this value i.e 13.5% and so the classication accuracy was more like 76.5%. At this point it seemed worthwhile to also examine the misclassifications ( a back translation stage called reconciliation) - what headline was mismatched with what headline. An examination of the false classifications suggested that around 40% of the mismatches may have been because of words they had in common, despote the full headline being different.

Where does that leave me? With some confidence in the headline generation process, but could we do better? Could we find a better way to generate reproducable headlines...See further below where I talk about ensemble methods.

Content analysis

The second task was a type of content analysis. Because of a specific interest, I had separated a subset of the hundred nine paragraphs into two groups, the first of which had been the subject to further narrative development by the participants (aka surviving storylines), and the second being others which were not developed any further (aka extinct storylines). I asked Claude AI to analyse the subset of the texts in terms of three attributes: the vocabulary, the style of writing, and the genre. Then for each attribute, to sort the texts into two groups, and describe what each group had in common and how they differed from the other group. It then did so. Here is an image of its output.

But how can I evaluate this output? If I looked at one of the texts in a particular group would I find the attributes that Claude AI was telling me that the group it belonged to possessed? In order to make this form of verification easier, and smaller in scale, I gave Claude AI a follow-up task: for each of the two groups under each of the three attributes of the text Claude AI should provide the ID number of an exemplar body of text which best represented the presence of the characteristics that were described. This it was able to do, and in my first use of the specific case examples I found that 9/10 did fit the summary description provided for the group. This strategy is similar to another one which I've used with GPT4, when trying to extract specific information about evaluation methods used in a set of evaluation reports. There I have asked it to provide page or paragraph references for any claim about what methods are being used in the evaluation. Broadly speaking, in a large majority of cases, these page references pointed to relevant sections of text.

My second strategy was another version of back translation, connecting concrete instances with pre-existing abstract descriptions. This time I opened a new prompt session, still within Claude AI, and uploaded a file containing the same subset of paragraphs, and then in the prompt window I copy and pasted the description of the attributes of the three sets of two groups identified earlier (without information on which text belnged to which group). I then asked Claude AI to identify which paragraphs of text fitted which of the 3 x 2 groups, which it did. I then collated the results of the two tasks in an Excel file, which you can see here below (click on image to magnify it). The green cells are where the predicted group matches the original group, and the yellow cells are where there were mismatches. The overall classification accuracy was 67%, whch is better than chance but not great either. I should also add that this was done with prompt information that included the IDs of the exemplars mentioned above (a format called "one-shot learning")

What was I evaluating when I was doing these "reverse translations"? It could probably be described as a test of, or search for, some form of construct validity. Was there any stable concept involved?

Ensemble methods

Given the two results reported above, which were better than chance, but not much better, what else could be done? There is one possible way forward, which might give us more confidence in the products generated by LLM analyses. Both Claude AI and ChatGPT4, and probably others, allow users to hit a Retry button, to generate another response to the same prompt. These will usually vary, and the degree of variation can be controlled by a parameter known as "temperature".

An ensemble approach in this context would be to generate multiple responses using the same prompt and then use some type of aggregation process to find the best result. Similar to 'wisdom of crowds" processes. In its simplest form this would, for example, involve counting the number of times each different headlines were proposed for the same item of text, and selecting one with the highest count. This approach will work where you have predefined categories as "targets". Those categories could have been developed inductively (as above) or deductively, from prior theory. It may even be possible to design a prompt script that include multiple genetration steps, and even the aggregation and evaluation stages.

But to begin with i will probably focus on testing a manual version of the process. I will report on some experiments with this approach in the next few days....

Update 02/09/23: A yet to be tested draft prompt that could automate the process

A lesson learned on the way: I initially wrote out a rough draft of a Claude AI prompt that might help automate the process I've described above. I then ask Claude AI to convert this into a prompt which would be understood and generate reliable and interpretable results. When it did this it was clear that part of my intentions had not been understood correctly (however you interpret the word understood). This could be just an epiphenomenon, in the sense of it only being generated by this particular enquiry. Or, it could point to a deeper or more structurally embedded analytic risk that would have consequences if I actually ask Claude AI to implement the rough draft in its original form (as distinct from simply refine that text as a prompt). The latter possibility concerned me, so I edited the prompt text that had been revised by Claude AI to remove the misunderstood part of the process. The version you see above is Claude AIs interpretation of my revised version, which I think will now work. Lets see,,,!

Update 03/09/23: It looks like the ensemble method may work as expected. Using 10 iterations only, which is a small number compared to how they are normally used, the classication accuracy increased to 84%. In the data displayed about numbers of time each predicted headline was matched to a given text there were 4 instances where there were ties. There were also 8 instances where the best match was still only found in less than 5 of the 10 iterations. More iterations might generate more definitive best matches and increase the accuracy rate. The correct match was already visible in the second and third ranking best matches of 4 of the 18 incorrectly matches headlines.

Another lesson learned, perhaps: Careful wording of prompts is important, the more explicit the instructions are the better. I learned to preface the word "match" with a more specific "analyze the content of all the numbered texts in File 1 and identify which one the headline best describes" . And careful formating of the text data files was also potentially important. making it clear where each text began and ended and removing any formating artifacts that could cause confusion.

And because of experiences with such sensitivities, I think i should re-do the whole analysis, to see if I generate the same or similar results!!!

Ensembles of brittle prompts?

I just came across this glimpse of a paper "Prompt Ensembles Make LLMs More Reliable" which is a different version of the idea I explored above. Here the prompt that is in use is also varied, from iteration to iteration.

Ensembles of brittel prompts?

Finding useful distinctions between different futures

2023-05-26T10:56:00.005+00:00

This blog posting is a response to Joseph Voros's informative blog posting about the Futures Cone. It is a useful contribution in as much as it helps us think about the future in terms of different sets of possibilities. Here is a copy of his edited version.

Figure 1: Voros, 2017

My alternative, shown below, was developed in the context of supporting ParEvo.org explorations of alternative futures. It has some similarities and differences. For a start, here is the diagram.

Figure 2: Sets and sub-sets of alternative futures Davies, 2023

I will now list Joseph's explanation of each of the terms he used, and how they might relate to mine (in red)

Possible – these are those futures that we think ‘might’ happen, based on some future knowledge we do not yet possess, but which we might possess someday (e.g., warp drive). I think these fall in the grey area above (which also contain the dark and light green).
Plausible – those we think ‘could’ happen based on our current understanding of how the world works (physical laws, social processes, etc).I think these fall somewhere within the green matrix
Probable – those we think are ‘likely to’ happen, usually based on (in many cases, quantitative) current trends. These probably fall within the Likely row of the green matrix
Preferable – those we think ‘should’ or ‘ought to’ happen: normative value judgements as opposed to the mostly cognitive, above. There is also of course the associated converse class—the un-preferred futures—a ‘shadow’ form of anti-normative futures that we think should not happen nor ever be allowed to happen (e.g., global climate change scenarios comes to mind).These probably fall within the Desirable column of the green matrix
Projected – the (singular) default, business as usual, ‘baseline’, extrapolated ‘continuation of the past through the present’ future. This single future could also be considered as being ‘the most probable’ of the Probable futures. As suggested above, probably at the most likely end of the Likely row in the above green matrix
(Predicted) – the future that someone claims ‘will’ happen. I briefly toyed with using this category for a few years quite some time ago now, but I ended up not using it anymore because it tends to cloud the openness to possibilities (or, more usefully, the ‘preposter-abilities’!) that using the full Futures Cone is intended to engender. Probably also at the most likely end of the Likely row in the above green matrix

Preposterious events are not really covered. Perhaps they are at the extreme end of the Unlikely events with known probabilities i.e zero likelihood.

Though lacking in alliteration my schema does have some more practically useful features

The primary additional feature is that for each different kind of future there are some conjectured consequences in terms of likely appropriate responses. Some of these are shown red:

Organisational "slack" i.e. uncommitted resources or reserves that could enable responses to the unforeseen (though, of course, not every kind of unforseen event)
Fringe investments, such as blue sky research, can be appropriate where a possibility is in sight but its likelihood of happening is far from clear
Robust responses are those that might work, though not necessarily be the most effective or most efficient, across a span of possibilities having varying probabilities and desirabilities
Customised responses are those more tailored to specific combinations of un/likely and un/desirable events. The following more detailed version of the green martix describes some major possible variations of this kind

Figure 3

Where to next?

I would like to hear from readers their views on the possible utility of these distinctions. And whether any other distinctions could be added to or replace those I have used.

Postscript 2025 01 22

I came across this useful matrix view in
Luís, A., Garnett, K., Pollard, S. J. T., Lickorish, F., Jude, S., & Leinster, P. (2021). Fusing strategic risk and futures methods to inform long-term strategic planning: Case of water utilities. Environment Systems & Decisions, 41(4), 523–540. https://doi.org/10.1007/s10669-021-09815-1

How can evaluators practically think about multiple Theories of Change in a particular context?

2023-03-06T12:02:00.004+00:00

This blog posting is been prompted by participation in two recent events. One was some work I was doing with the ICRC, reviewing Terms of Reference for an evaluation. The other was listening in as a participant to this week's European Investment Bank conference titled "Picking up the pace: Evaluation in a rapidly changing world".

When I was reviewing some Terms of Reference for an evaluation I noticed a gap which I have seen many times before. While there was a reasonable discussion of the types of information that would need to be gathered there was a conspicuous absence of any discussion of how that data would be analysed. My feedback included the suggestion that the Terms of Reference needed to ask the evaluation team for a description of the analytical framework they would use to analyse the data they were collecting.

The first two sessions of this week's EIB conference were on the subject of foresight and evaluation. In other words how evaluators can think more creatively and usefully about possible futures – a subject of considerable interest to me. You might notice that I've referred to futures rather than the future, intentionally emphasising the fact that there may be many different kinds of futures, and with some exceptions (e.g. climate change) is not easy to identify which of these will actually eventuate.

To be honest, I wasn't too impressed with the ideas that came up in this morning's discussion about how evaluators could pay more attention to the plurality of possible futures. On the other hand, I did feel some sympathy for the panel members who were put on the spot to answer some quite difficult questions on this topic.

Benefiting from the luxury of more time to think about this topic, I would like to make a suggestion that might be practically usable by evaluators, and worth considering by commissioners of evaluations. The suggestion is how an evaluation team could realistically give attention not just to a single "official" Theory Of Change about an intervention, but to multiple relevant Theories Of Change about an intervention and its expected outcomes. In doing so I hope to address both issues I have raised above: (a) the need for an evaluation team to have a conceptual framework structuring how it will analyse the data it collects, and (b) the need to think about more than one possible future and how that might be realised i.e. more than one Theory of Change.

The core idea is to make use of something which I have discussed many times previously in this blog, known as the Confusion Matrix – to those involved in machine learning, and more generally described simply as a truth table - one that describes four types of possibilities. It takes the following form:

In the field of machine learning the main interest in the Confusion Matrix is the associated performance measures that can be generated, and used to analyse and assess the performance of different predictive models. While these are of interest, what I want to talk about here is how we can use the same framework to think about different types of theories, as distinct from different types of observed results.

There are four different types of Theories of Change that can be seen in the Confusion Matrix. The first (1) describes what is happening when intervention is present and the expected outcome of that intervention is present. This is the familiar territory of the kind of Theories of Change that an evaluator will be asked to examine.

The second (2) describes what is happening when intervention is present and the expected outcome of that intervention is absent. This theory would describe what additional conditions are present, or what expected conditions are absent, which will make a difference – leading to the expected outcome being absent. When it comes to analysing data on what actually happened identifying these conditions can lead to modification of the first (1) Theory of Change such that it becomes a better predictor of the outcome and there are fewer False Positives (found in cell 2). Ideally the less False Positives the better. But from a theory development point of view there should always be some situations described in cell 2 because there will never be an all-encompassing theory that works everywhere. There will always be boundary conditions beyond which the theory is not expected to work. So an important part of an evaluation is not just to refine the theory about what works (1) but also to refine the theory of the circumstances in which it will not be expected to work (2), sometimes known as conditions or boundary conditions.

The third theory (3) describes what is happening when the intervention is absent but nevertheless the outcome is present. Consideration of this possibility involves recognition of what is known as "multi-finality" i.e. that some events can arise from multiple alternative causal conditions (or combinations of causal conditions). It's not uncommon to find advice to evaluators that they should consider alternative theories to those they are currently focused on. For example in the literature on contribution analysis. But it strikes me that this is often close to a ritualistic requirement, or at least treated that way in practice. In this perspective alternative theories are a potential threat to the theory being focused on (1). But a much more useful perspective would be to treat these alternative theories as potentially useful other courses of an action that an agent could take, which warrant serious attention in their own right. And if they are shown to have some validity this does not by definition mean that the main theory of change (1) is wrong. It' simply means that there are alternative ways of achieving the outcome, which can only be a bonus finding.

The fourth theory describes what is happening when intervention is absent and the outcome is also absent (4). In its simplest interpretation, it may be that the actual absence of the attributes of the intervention is the reason why the outcome is not present. But this can't be assumed. There may be other factors which have been more important causes. For example the presence of an earthquake, or the holding of a very contested election. This possibility is captured by the term "asymmetric causality" i.e. that the causes of something not happening may not simply be the absence of the causes of something happening. Knowing about these other possible causes of desired outcome not happening is surely important, in addition to and alongside knowing about how an intervention does cause the outcome. Knowing more about these causes might help other parties with other interventions in mind move cases with this experience from being True Negatives (4) to being False Negatives (3)

In summary, I think there is an argument for evaluators not being too myopic when they are thinking about Theories of Change they need to pay attention to. It should not be all about testing the first (1) type of Theory of Change, and considering all the other possibility is simply as challengers, which may or may not then be dismissed Each of those other types of theories (2-3-4) are important and useful in their own right and deserve attention.

Four types of futures that should be covered by a Theories of Change

2022-10-18T10:58:00.001+00:00

ParEvo.org is a web app that enables the collaborative exploration of alternative futures, online. In the evaluation stage, participants are asked to identify which of the surviving storylines fall into each of these categories:

Most desirable
Least desirable
Most likely
Least likely

In one part of the analysis of storylines generated during a ParEvo exercise the storylines are plotted on scatter plot, where the two dimensions are likelihood and desirability, as seen in this example

Most Theories of Change that I have come across, when working as an evaluator, focus on a future that is seen as desirable and likely (as in expected). At best, the undesirable futures will be mentioned in an accompanying section on risks and their management.

A less myopic approach might be useful, one which would orient the users of the Theory of Change to a more adaptive stance towards the future.

One way forward would be to think of a four-part Theory of Change, each of which has different implications. as follows

The top right cell may already be covered by a Theory of Change. In the desirable but unlikely, and undesirable but likely two cells it would be useful to have ordered lists that describe events, what needs to be done before they happen, and what needs to be done after they happen. In the unlikely and undesirable cell plans for monitoring the status of these events need to be spelled out, and updated on an ongoing basis

We need more doubt and uncertainty!

2022-10-13T15:30:00.016+00:00

This week the Swedish Evaluation Society (SVUK) is holding its annual conference. I took part in a session today on Theories of Change. The first part of my presentation summarised the points I made in a 2018 CEDIL Inception Report titled 'Theories of Change: Technical Challenges with Evaluation Consequences'. Following the presentation I was asked by Gustav Petersson, the discussant, whether we should pay more attention to the process of generating diagrammatic Theories of Change. I could only agree, reflecting that for example it was not uncommon that a representative of a conference working group might summarise a very comprehensive and in-depth discussion in all too brief and succinct terms when reporting back to a plenary. Leaving out, or understating, the uncertainties , ambiguities and disagreements. Similarly the completed version of a diagrammatic Theory of Change is likely to suffer from the same limitations ... being an overly simplified version of a much more complex and nuanced discussions between those involved in its construction that went on beforehand.

Later in the day I was reminded of this section in the Hitchhiker's Guide to the Galaxy where Vroomfondel, representing a group of striking philosophers said '"That's right!" and shouted , "we demand rigidly defined areas of doubt and uncertainty!"

I'm inclined to make a similar kind of request of those developing Theories of Change. And of those subsequently charged with assessing the evaluability of the associated intervention, including its Theory of Change. What I mean is that the description of the Theory of Change should make it clear which various parts of the theory the owner(s) of that theory are more confident in verses less confident. Along with descriptions of the nature of the doubt or uncertainty and its causes e.g. first-hand experience, or supporting evidence (or lack of) from other sources.

Those undertaking an evaluability assessment could go a step further and convert various specific forms of doubt and uncertainty into evaluation questions that could form an important part of the Terms of Reference for an evaluation. This might go some way to remedying another problem discussed during the session, which is the all too common (in my experience) phenomena of Terms of Reference only making generic references to an intervention's Theory of Change. For example, by asking in broad terms about "what works and in what circumstances". Rather than the testing of various specific parts of that theory, which would arguably be more useful, and better use of limited time and resources.

The bottom line: The articulation of a Theory of Change should conclude with a list of important evaluation questions. Unless there are good reasons to the contrary, those questions should then appear in the Terms of Reference for a subsequent evaluation

PS: Vroomfondel is a philosopher. He appears in chapter 25 of The Hitchhiker's Guide to the Galaxy, along with his collegue Majikthise, as a representative of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons (AUPSLOTP; the BBC TV version inserts 'Professional' before 'Thinking'). The Union is protesting about Deep Thought, the computer which is being asked to determine the Answer to the Ultimate Question of Life, the Universe and Everything. See https://hitchhikers.fandom.com/wiki/Vroomfondel

Using ParEvo to conduct thought experiments

2022-06-30T14:34:00.010+00:00

I have just had an interesting conversation with an NGO network who have been developing some criteria to: (a) help speed up the approval and release of funding in humanitarian emergencies, but (b) at same time minimising risk of poor use of those funds.

They think these criteria are useful but are not entirely sure whether those seeking funding will agree. So they are exploring ways of testing out their applicability through a wider consultation process.

One way doing this, which we have been discussing, involves the use of ParEvo.org. The plan is that a group of participants representing potential grantees will develop a set of storylines which starts off with a particular organisation seeking funding for a particular humanitarian emergency. Then a branching structure of possible subsequent storyline developments will be articulated through the usual ParEvo process

After those storylines been developed there will be an evaluation phase, as is common practice now with most ParEvo exercises. At this point the participants will be asked two generic types of questions ( and variations on these), as described below:

1. Which of the criteria in the current framework would be most likely to help avoid or mitigate the problems seen in storyline X? (Answer=Description & Explanation)

and if the answer is none, are there any other criteria that could be included in the framework that might have helped?

2. Which of the storylines in the current exercise would have most benefited by criteria X in the current framework, in the sense of problems described there would have been avoided or mitigated. (Answer=Description & Explanation)

and if the answer is none, does this suggest that the criteria is irrelevant and could be removed?

Postscript: One interesting thing about this type of thought experiment is that the theory (the proposed funding criteria) and the possible realities that they may be applied to (where the theories may or may not work there as expected) are constructed by different parties who are independent from each other. This is not usually the case with thought experiments, and could be seen as a positive variation.

Stay tuned for if and when this idea flies, then soars or crashes

Courtesy https://xkcd.com/

For more on thought experiments, see Armchair science

Thought experiments played a crucial role in the history of science. But do they tell us anything about the real world?

Alternative futures as "search strategies"

2022-06-17T17:12:00.006+00:00

When you read the phrase "search strategy' this may bring to mind what you need when you are doing a literature search on the Internet. Or you may be thinking about different forms of supervised machine learning, which involve different types of search strategies. For example in my Excel-based EvalC3 prediction modelling app there are four different search strategies that users can choose from, to help find the most accurate predictive model describing what combinations of attributes are the best predictor of a particular outcome. Or you may have heard of James March, an organisational theorist who in 1981 wrote a paper called 'A model of adaptive organizational search ' where he talks about how organisations find the right new technologies to develop and explore.This is probably the closest thing to the type of search process that I'm describing below.

Right now I am in the process of helping some other consultants design a ParEvo exercise, in which recipients of research grants from the same foundation will collaboratively develop a number of alternative storylines describing how their efforts to ensure the uptake and use of the research findings takes place (and sometimes fails to take place) over the coming three years. Because these are descriptions of possible futures they are inherently a form of fiction. But please note they are not an attempt at "predicting" fiction. Rather, they are more like a form of 'preparedness enabling ' fiction.

As part of the planning process for this exercise we have had to articulate our expectations of what will come out of this exercise, in terms of possible desirable benefits for both the participants and the foundation. In other words the beginnings of a Theory of Change, which needs to be supplemented by details of how the exercise will be best be run in this particular instance, and thus hopefully deliver these results.

When thinking about reasonable expectations for this exercise I came up with the following possibilities, which are now under discussion:

1 Participants will hear different interpretations and views of

What other participants mean when they use the term "research uptake '
What successful, and unsuccessful, research uptake looks like in its various forms, to various participants
How the process of research uptake can be facilitated, and inhibited, by a range of factors – some within researchers control and some beyond their control.

2. This experience may then inform how each of the participants proceed with their own work on facilitating research uptake

3. The storylines that are generated by the end of the exercise will provide the participants and the XXXX trust with a flexible set of expectations against which actual progress with research uptake can be compared at a later date.

So, my current thinking is that what we have here is a description of a particular kind of search strategy where both the objectives worth pursuing, and the means of achieving them, are both being explored at the same time, at least within the ParEvo exercise. Though other things will also be happening after the exercise, hopefully involving some use of the ideas generated during exercise (see possibility 2)

There is also another facet of the idea of search strategies which needs to be mentioned here. When search is used in a machine learning context it is always accompanied by an evaluation function which determines whether the search continues or comes to a stop because the best possibility has now been identified (a stopping rule, I think is the term involved). So, in the three possibilities listed above the last one describes the possibility of an evaluation function. Exactly how it will work needs more thinking, but I think it will be along the lines of asking participants in the prior exercise to identify the extent to which their experience in the interim period has fitted any of the storylines that were developed earlier, and in what ways it has and has not, and why so in both cases. Stay tuned...

Budgets as theories

2022-04-28T11:02:00.009+00:00

A government has a new climate policy. It outlines how climate investments will be spread through a number of different ministries, and implemented by those ministries using a range of modalities. Some funding will be channelled to various multilateral organisations. Some will be spent directly by the ministries. Some will be channelled on to the private-sector. At some stage in the future this government wants to evaluate the impact of this climate policy. But before then it is been suggested that an evaluability assessment might be useful, to ask if how and when such an evaluation might be feasible.

This could be a challenge to those with the task of undertaking the evaluability assessment. And even for those planning the Terms of Reference for that evaluability assessment. The climate policy is not yet finalised. And if the history of most government policy statements (that I have seen) has any lessons it is that you can't expect to see a very clearly articulated Theory of Change of the kind that you might expect to find in the design of a particular aid programme.

My provisional suggestion at this stage is that the evaluability assessment should treat the government's budget, particularly those parts involving funding of climate investments, as a theory of what is intended. And to treat the actual flows of funding that subsequently occur as the implementation of that theory. My naïve understanding of the budget is that it consists of categories of funding, along with subcategories and sub- subcategories, et cetera. In other words a type of tree structure involving a nested series of choices about where more versus less funds should go. So, the first task of an evaluability assessment would be to map out the theory i.e. the intentions as captured by budget statements at different levels of detail, moving from national to ministerial and then to small units thereafter. And to comment on the adequacy of these descriptions and and gaps that need to be addressed.

This exercise on its own will not be sufficient as an explication of the climate policy theory because it will not tell us how these different flows of funding are expected to do their work. One option would be to follow each flow down to its 'final recipient', if such a thing can actually be identified. But that would be a lot of work and probably leave us with a huge diversity of detailed mechanisms. Alternatively, one might do this on sampling basis, but how would appropriate samples be selected?

There is an alternative which could be seen as a necessity that could then be complemented by a sampling process. This would involve examining each binary choice, starting from the very top of the budget structure and asking 'key informants" questions about why climate funding was present in one category but not the other, or more in one category than the other. This question on its own might have limited value because budgeting decisions are likely to have a complex and often muddy history, and the responses received might have a substantial element of 'constructed rationality' . Nevertheless the answers could provide some useful context.

A more useful follow-up question would be to then ask the same informants about their expectations of differences in performance of the amount of climate financing via category X versus category Y. Followed by a question about how they expect to hear about the achievement of that performance, if at all. Followed by a question about what they would most like to know about performance in this area. Here performance could be seen in terms of the continuum of behaviours, ranging from simple delivery of the amount of funds as originally planned, to their complete expenditure, followed by some form of reporting on outputs and outcomes, and maybe even some form of evaluation, reporting some form of changes.

These three follow-up questions would address three facets of an evaluability assessments (EA): a) The ToC - about expected changes, b) Data availability , c) Stakeholder interests. Questions would involve two types of comparisons: funding versus no funding, and more versus less funding. The fourth EA question, about the surrounding institutional context, typically asks about the factors that may enable and/or limit an evaluation of what actually happened (more on evaluability assessments here).

There will of course be complications in this sort of approach.. Budget documents will not simply be a nested series of binary choices, at each level their work may be multiple categories available rather than just two. However informants could be asked to identify 'the most significant difference 'between all these categories, in effect introducing an intermediary binary category. There could also be a great number of different levels to the budget documents, with each new level in effect doubling the number of choices and associated questions that need to be asked. Prioritisation of enquiries would be needed, possibly based on a 'follow the (biggest amount of) money 'principle. It is also possible that quite a few informants will have limited ideas or information about the binary comparisons they are asked about. A wider selection of informants might help fill that gap. Finally there is the question of how to 'validate" the views expressed about expected differences in performance, availability of performance information and relevant questions about performance. Validation might take the form of a survey of a wider constituency of stakeholders within the organisation of interest, of the views expressed by the informants.

PS: Re this comment in the third para above: "And to treat the actual flows of funding that subsequently occur as the implementation of that theory" One challenge the EA team might find is that while it may have accessed to detailed budget documents, in many places it may not yet be clear where funds have been tagged as climate finance spending. That itself would be an important EA finding.

To be continued...

Making small samples of large populations useful

2022-04-24T16:05:00.022+00:00

I was recently contacted by someone who is working for a consulting firm that has a contract to evaluate the implementation of a large-scale health program covering a huge number of countries. Their client had questioned their choice of 6 countries as case studies. They were encouraging the consultancy firm to expand the number of country case studies, apparently because they thought this would make this sample of country cases more representative of the population of countries as a whole. However, the consulting firm wasn't planning to aggregate results of the six country case studies and then make a claim about generalisability of findings across the whole population of countries. Quite the opposite, the intention was that each country case study would provide a detailed perspective on one or more particular issues that was well exemplified by that case.

In our discussions, I ended up suggesting a strategy that might satisfy both parties in that it addressed to some extent the question of generalisable knowledge at the same time was designed to exploit the particularities of individual country cases. My suggestion was relatively simple, although implementing might take a bit of work making use of whatever data is available on the full population of countries. The suggestion was that for each individual case study the first step in the process would be to identify and explain the interesting particularities of that case, within the context of the evaluation's objectives. Then the evaluation team would look through whatever data is available on the whole population of countries, with the aim of identifying a sub-set of other countries that had similar characteristics (perhaps both generic {political, socio-economic indicators} and issue specific) with the case study country. These would then be assumed to be the countries where the case study findings and recommendations could be most relevant.

A shown in the diagram below, it is possible that the sub-set of countries relevant to each case study county might overlap to some extent. Even when one case study country is examined it is possible that it might have more than one particularity of interest, each of whose analysis might be usefully generalised to a limited number of other countries. And those different sub-sets of countries may themselves overlap to some extent (not shown below).

Green nodes = case study countries
Red nodes = remainder of the whole population
Red nodes connected to green nodes = countries that might find green node country case study findings relevant
Unconnected red nodes = Parts of whole population where case study findings not expected to have any relevance

Another possibility, perhaps seen as unadvisable in normal circumstances, would be to identify the relevant countries to any case study analysis after the fact, not necessarily or only before. After the case study had actually been carried out there would be much more information available on the relevant particularities of the case study country that might make it easier to identify which other countries these finding were most relevant to. However the client of the evaluation might need to be given some reassurance in advance. For example, by ensuring that at least some of these (red node) countries were identified at the beginning, before the case studies were underway.

PS: It is possible to quantify the nature of this kind of sampling. For example, in the above diagram

Total number of cases = 37 (red and green).

Case study cases = 5 (14%) of all cases

Relevant-to-case-study cases = 17 (46%) of all cases

Relevant-to->1-case-study cases = 3 (8%) of all cases

Not-relevant-to-case-study* cases = 15 (40%) of all cases

*Bear in mind also that in many evaluations case studies will not be the only means of inquiry. For example, there are probably internal and external data sets that can be gathered and analysed re the whole set of 37 countries.

Conclusion: We should not be thinking in terms of binary options. It is not true that either a case is part of a representative sample of a whole population, or it is representative and of interest only to itself. It can be relevant to a sub-set of the population.

Choosing between simpler and more complex versions of a Theory of Change

2021-11-25T11:00:00.001+00:00

Background: Over the last few months I have been involved as a member of the Evaluation Task Force, convened by the Association of Professional Futurists. Futurists being people who explore alternative futures using various foresight and scenario planning methods. The intention is to help strengthen the evaluation capacity of those doing this kind of work

One part of this work will involve the development of various forms of introductory materials and guidelines documents. These will inevitably include discussion of the use of Theories of Change, and questions about appropriate levels of detail and complexity that they should involve.

In my dialogues with other Task Force members I have recently made the following comments, which may be of wider interest:

As already noted, a ToC can take various forms, from very simple linear versions to very complex network versions.

I have a hypothesis that may be useful when we are developing guidance on use of ToC by futurists. In fact I have two hypotheses:

H1: A simple linear ToC is more likely to be appropriate when dealing with outcomes that are closest in time to a given foresight activity of interest. Outcomes that are more distant in time, happening long after the foresight activity has finished, would be better represented in a ToC that took a more complex network (i.e. systems map type) form

Why so?: As time passes after a foresight activity, more and more other forces, or various kinds, are likely to come into play and influence the longer term outcome of interest. As a proportion of all influences, the foresight activity will grow progressively smaller and smaller. A type of ToC that takes into account this widening set of influences would seem essential

H2: This need for progressively more complex ToC, as the outcome of interest is located further away in time, can be moderated by a second variable, which is the social distance between those involved in the foresight activity and those involved in the outcome of interest . [Social distance is measured in social network analysis (SNA) terms by units known as "degree", i.e, the number of person-to-person linkages needed for information to flow between one person and another]. So, if the outcome is a change in the functioning of the same organization that the foresight exercise participants they themselves belong to, this distance will be short, relative to an outcome relating to another organisation altogether - where there may be few if any direct links between the exercise participants and staff of that organisation

The implications of these two perspectives could be graphically represented in a scatter plot or two-by-two matrix e.g.

On reflection, this view probably needs some more articulation. Social distance will probably not be present in the form of a single pathway through a network of actors. Especially given that any foresight activity will typically involve multiple participants, each with their own access to relevant networks. So there may be a third relevant dimension here to think about, which is the diversity of the participants. Greater diversity being plausibly associated with a greater range of social (and causal) pathways to the outcome of interest. And thus the need for more complex representations of the Theory of Change.

Exploring counterfactual histories of an intervention

2021-11-01T17:57:00.041+00:00

Background

I and others are providing technical advisory support to the evaluation of a large complex multilateral health intervention, one which is still underway. The intervention has multiple parts implemented by different partners and the surrounding context is changing. The intervention design is being adapted as time moves on. Probably not a unique situation.

The proposal

As one of a number of parts of a multi-year evaluation process I might suggest the following:

1. A timeline is developed describing what the partners in the intervention see as key moments in its history, defined as where decisions have been taken to change , stop or continue, a particular course(s) of action.

2. Take each of those moments as a kind of case study, which the evaluation team then elaborates in detail: (a) the options that were discussed, and others not, (d) the rationales for choosing for and against those discussed options at the time, (d) the current merit of those assessments, as seen in the light of subsequent events. [See more details below]

The objective?

To identify (a) how well the intervention has responded to changing circumstances and (b) any lessons that might be relevant to the future of the intervention, or generalisable to other similar intervention.

This seems like a form of (contemporary and micro-level) historical research, investigating actual versus possible causal pathways. It seems different from a Theory of Change (ToC) based evaluation, where the focus is on what was expected to happen and then what did happen. Whereas with this proposed historical research in to decisions taken the primary reference point is what did happen, then what could have happened.

It also seems different from what I understand is a process tracing form of inquiry where, I think, the focus is on particular hypothesised causal pathway. Not the consideration of multiple alternative possible pathways, as would be the case within each of a series of decision making case studies proposed here. There may be a breadth rather than depth of inquiry difference here.[Though I may be over-emphasising the difference here, ...I am just reading Mahoney, 2013 on use of process tracing in historical research]

The multiple possible alternatives that could have been chosen are the counterfactuals I am referring to here.

The challenges?

As Henry Ford allegedly said "History is just one damn thing after another" There are way too many events in most interventions where alternative histories could have taken off in a different direction. For example, at a normally trivial level, someone might have missed their train. So to be practical but also systematic and transparent the process of inquiry would need to focus on specific types of events, involving particular kinds of actors. Preferably where decisions were made about courses of action. Such as Board Meetings.

And in such individual settings how wide should the should the evaluation focus be? For example, only on where decisions were made to change something, or also where decisions were made to continue doing something? And what about the absence of decisions being even considered, when they might have been expected to be considered. That is, decisions about decisions.

Reading some of the literature about counter-factual history, written by historians, there is clearly a risk of developing historical counterfactuals that stray too far from what is known to have happened, in terms of imagined consequences of consequences, etc. In response, some historians talk about the need to limit inquiries to"constrained counterfactuals" and the use of a "Minimal Rewrite Rule". [I will find out more about these]

Perhaps another way forward is to talk about counter-factual reasoning, rather than counterfactual history (Schatzberg, 2014) . This seems to be more like what the proposed line of inquiry might be all about i.e. how the alternatives to what actually was decided and happened were considered (or not even considered) by the intervening agency. But even then, the evaluators' assessments of these reasonings would seem to necessarily involve some exploration of consequences of these decisions, and only some of which will have been observable, and others only conjectured.

The merits?

When compared to a ToC testing approach this historical approach does seem to have some merit. One of the problems of a ToC approach, particularly when applied to a complex intervention is the multiplicity of possible causal pathways, relative to the limited time and resources available available to an evaluation team. Choices usually need to be made, because not all avenues can be explored (unless some can excluded by machine learning explorations or other quantitative processes of analysis).

However, on reflection, the contrast with a historical analysis of the reality of what actually happened is not so black and white. In large complex programmes there are typically many people working away in parallel, generating their own local and sometimes intersecting histories. There is not just one history from within which to sample decision making events In this context a well articulated ToC may be a useful map, a means of identifying where to look for those histories in the making.

Where next

I have since found that that the evaluation team has been thinking along similar lines to myself i.e. about the need to document and analyse the history of key decisions made. If so, the focus now should be on elaborating questions that would be practically useful, some of which are touched on above. Including:

1. How to identify and sample important decision making points

At least two options here:

1. Identify a specific type of event where it is know that relevant decisions are made. E.g, Board Meetings. This is a top-down deductive approach. Risk here is that many decisions will (and have to be) made outside and prior to these events, and just receive official authorisation at these meetings. Snowball sampling backwards to original decisions may be possible...

2. Try using the HCS method to partition the total span of time of interest into smaller (and nested) periods of time. Then identify decisions that have generated the differences observed between these periods (which will sought about the intervention strategy). This is a more bottom-up inductive approach.

2. How to analyses individual decisions.

The latter includes interesting issues such as how much use should be made of prior/borrowed theories about what constitute good decision making, versus using a more inductive approach that emphasises understanding how the decisions were made within their own particular context. I am more in favor of the latter at present

Here is a VERY provisional framework/checklist for what could be examined, when looking at each decision making event:

In this context it may also be useful to think about a wider set of relevant ideas like the role of "path dependency" and "sunk costs"

3. How to aggregate/synthesise/summarise the analysis of multiple individual decision making cases

This is still being thought about, so caveat emptor:

Objectives were to identify:

(a) how well the intervention has responded to changing circumstances.

Possible summarising device? Rate each decision making event on degree to which it was optimal under the circumstances. Backed by a rubric explaining rating values.

Cross tabulate these ratings against a ratings of the subsequent impact of the decision that was made? An "Increased/decreased potential impact" scale.? Likewise supported by a rubric (i.e. annotated scale).

(b) any lessons that might be relevant to the future of the intervention, or generalisable to other similar intervention.

Text summary of implications identified from the analysis of each decision making event, with priority to more impactful/consequential decisions?

Lot more thinking yet to be done here...

Miscellaneous points of note hereafter...

Postscript 1: There must be a wider literature on this type of analysis, where there may be some useful experiences. "“Ninety per cent of problems have already been solved in some other field. You just have to find them.” McCaffrey, T. (2015) New Scientist.

Postscript 2: I just came across the idea of an "even if..." type counterfactual. As in "Even if I did catch the train, I would still not have got the job". This is where when an imagined action, different from what really what happened, still leads to the same outcome as when the real action took place.

Wishmi, CC BY-SA 4.0, via Wikimedia Commons

Reconciling the need for both horizontal and vertical dimensions in a Theory of Change diagram

2021-08-22T12:08:00.007+00:00

In their usual table form, Logical Frameworks are strong on their horizontal dimension but weak on their vertical dimension. On the horizontal dimension is the explanation of what kind of data will be collected and used to measure the changes that are described. This is good for accountability. On the vertical dimension is the explanation of how events at one level will connect and cause events at another level. This is good for learning. But unfortunately LogFrames often simply provide lists of events at each level, with relatively little description of which event will connect to which, especially where multiple and mixed sets of connections might be expected between events. On the other hand diagrammatic versions of a Theory of Change tend to be much better at explicating the various causal pathways at work, but weak on the information they provide on the horizontal dimension - on how various events will be observed and measures. Both of these problems reflect both a lack of space to do both things and different relative priorities pursued within those constraints.

The Donor Committee for Enterprise Development (DCED) has produced a web page based Theory of Change to explain its way of working, which I think points the way to reconciling these conflicting needs. At first glance here is what you see, when you visit this page of their website.

The different causal pathways are quite visible, more so than within a standard LogFrame table format. But another common weakness of diagrammatic versions of Theories of Change is the lack of explanation of what is going on within each of these pathways. The DCED addressed this problem by allowing visitors to click on a link and be taken to another web page, where visitors get a detailed text description of the available evidence, plus any assumptions, about the causal process(es) that connect the events connected by the arrow.

The one weakness in this DCED ToC diagram is the lack of detail about the horizontal dimension- how the various events described in the diagram will be observed./ measured and by who and when and where. But this is clearly resolvable by using the same approach with the links: enable users to click on any event and be taken to a web page where this information is provided for that specific event. As shown below:

Diversity and complexity? Where should we focus our attention?

2021-07-19T14:46:00.012+00:00

This posting has been promoted by Michael Bamberger's recent two blog postings on "Building complexity into development evaluation" on the 3ie website: Evidence Matters: Towards equitable, inclusive and sustainable development

I started to make some comments underneath each of the two postings but have now decided to try to organise and extend my thoughts here.

My starting point is an ongoing concern about how unproductive the discussion has been about complexity (especially in relation to evaluation). Like an elephant giving birth to a mouse, has been my chosen metaphor in the past. There probably is some rhetorical overkill here, but it does express some of my felt frustration with the limited value of the now quite extended discussion.

Michael's blog postings have not allayed these concerns. My concerns start with the idea of measuring complexity, both how you do it and how measuring would in fact be useful. Measuring complexity is Michael's proposed first step in a "practical five-step approach to address complexity-responsive evaluation in a systematic way" A lot of ink has already been spilled on the topic of measurement, which is the first of the five steps. A useful summary can be found in Melanie Mitchel's widely read Complexity: A guided Tour (2009:94-114) and Loyd, 1998, who counted at least 40 different ways. But I cant see any references back to any of these methods, suggesting that not much is being learned from past efforts, which is a pity.

Along with the challenge of how to do it is the question of why you would want to do it, ...how might it be useful? The second blog posting explains that " In addition to providing stakeholders with an understanding of what complexity means for their program, the checklist also helps decide whether the program is sufficiently complex to merit the additional investment of time and resources required to conduct a complexity-focused evaluation"

The second of these outcomes might be a more observable consequence, so my first question here is where is the cut-off point in a checklist derived score that would at least inform such a decision, and how is that cut-off point justified. The checklist has 25 attribute questions spread over 4 dimensions. This has not yet been made clear.

My next question is how do the results of this measurement exercise inform the next of the five steps: "Breaking the project into evaluable components and identifying the units of analysis for the component evaluation ". So far, I have not found any answers to this question either. PS 2021 07 29: Michael did reply to a Comment of mine raising the same issue, and suggested that high versus low scores on rated complexity might be one way.

Another concern which I have already written about in my comment on the blog postings is that complexity seems to be very much “in the eye of the beholder”, i.e. depending on who you are and what you are looking for. My partner sees much more complexity in the design and behaviour of moths and butterflies than I do. A friend of mine sees much more complexity in the performance of classical music than I do. Such observations prompt me to thinking that perhaps we should not put too much effort into trying to objectively measure complexity. Rather, perhaps we should take a more ethnographic perspective on complexity – i.e. we should pay attention to where people are seeing complexity and where they are not, and what are the implications thereof.

If we accept this suggestion it is still the case that the challenge of identifying complexity is still with us, but in a different form. So, I have another suggestion, which is to pay much more attention to diversity, as an important and related concept to complexity. As Scott Page has well described, there is a close and complicated relationship between diversity and complexity. Nevertheless, there are some practically useful points to note about the concept of diversity.

Firstly the presence of diversity is indicative of the absence of a common constraint, and the presence of many different causal influences. So can be treated as a proxy - indicating the presence of complex processes.

Secondly, there is extensive and more internally consistent and practically useful set of ways in which diversity can be measured. These mainly have their origins in the studies of biodiversity but have a much wider applicability. Diversity can also be measured in other spheres, in human relationships (using social network analysis tools) and how people see the world (using forms of ethnographic enquiry known as pile or card sorting).

Thirdly, diversity has some potentially global relevance as a value and as objective. Diversity of behaviour can be indicative of choice and empowerment.

Fourthly, diversity can also be seen as an important independent variable as well, enabling adaptation and creativity.

All this is not to say that diversity cannot also be problematic. Some forms of diversity in the present could severely limit the extent of diversity in the future. For example, if within a population there was a wide range of different types of guns held by households and many different views on how and when they could legitimately be used. At the more mundane level, within organisations different kinds of tasks may benefit from different levels of diversity within the groups addressing those tasks. So diversity presents a useful and important problematic in a way that the concept of complexity does not. What forms of diversity we want to see, and see sustained over time, and how can they be enabled? Where do we want choice and where should be accept restriction?

Arguing for more attention to diversity, rather than complexity, does not mean there also needs to be a whole new school of evaluation developed around this idea (Get Brand X Evaluation, just hot off the press! Uhh... No). It is consistent with a number of ideas already found useful, including the idea of equifinality (An outcome can arise from multiple different causes) and multifinality ( A cause can have multiple different outcomes), and the idea of multiple conjunctural causation. It is also compatible with a social constructionist and interpretive perspective on social reality.

Rick On the Road

Cluster analysis for evaluation purposes, and how a hybrid Human-LLM approach can help

Rethinking how we share evaluation methods

When rankings tell different stories: an introduction to the Rank Explorer

The problem with additive aggregation

What the Rank Explorer does

When the gap between additive and configurational results is itself a finding

A note on scale, depth, and selective deepening

Beyond TEE data

An invitation to experiment

Further reading

Introducing Rank Order Counterfactuals (ROC)

Making implicit knowledge explicit, contestable and usable: an introduction to The Ethnographic Explorer

The theoretical anchor: information as difference

How the Ethnographic Explorer works

The Ethnographic Explorer

A worked example: 12 Largest cities

What the tool produces

Beyond evaluation

An invitation to experiment

Further reading

Optimising method selection: an introduction to the set covering problem — and a tool to help

Why it matters for evaluation practice

How the optimisation works

The Coverage Optimiser

Beyond methods and questions

An invitation to experiment

Further reading

Exracting additional knowledge and performance from a configurational model that already has wide coverage

Objectives as data: The potential uses of updatable outcome targets

On two types of Theories of Change: Temporal and atemporal, and how they might be bridged

Developing and using a Configurational Theory of Change within an evaluation

But what about aggregation?

To be continued....

Using the Confusion Matrix as a general-purpose analytic framework

Beyond summarisation by AI and/or editors- Readers can now interrogate full transcripts of meeting discussions

Evaluating thematic coding and text summarisation work done by artificial intelligence (LLM)

Finding useful distinctions between different futures

Postscript 2025 01 22

How can evaluators practically think about multiple Theories of Change in a particular context?

Four types of futures that should be covered by a Theories of Change

We need more doubt and uncertainty!

Using ParEvo to conduct thought experiments

Alternative futures as "search strategies"

Budgets as theories

Making small samples of large populations useful

Choosing between simpler and more complex versions of a Theory of Change

Exploring counterfactual histories of an intervention

Reconciling the need for both horizontal and vertical dimensions in a Theory of Change diagram

Diversity and complexity? Where should we focus our attention?