There's a realistic possibility you'll read this
How likely is 'likely'?
How likely is ‘likely’? Does ‘likely’ have a higher probability than ‘probable’? How realistic is a ‘realistic possibility’?
As I’ve written about previously, there are some nice illustrative datasets out there on human perceptions of probability-based phrases, often collected via online polls. But all are relatively small; the biggest openly available one I could find is just over a hundred people. Datasets typically also focus on absolute estimates (e.g. what probability corresponds to ‘likely’?), rather than comparative judgements (e.g. does ‘likely’ or ‘probable’ have a higher probability?)
So a few weeks ago, I thought it would be interesting to run a larger quiz that included both absolute and comparative judgements, and allowed participants to see where they sit on the distribution.
Since the quiz launched, it seems to have really resonated with people. There have been over 5000 participants and counting, so thanks for everyone who’s taken part.
Now I’m making the underlying data available for others to explore too. I’m calling it CAPphrase (Comparative and Absolute Probability phrase dataset); this is an open access dataset containing over 150,000 probability-based language judgements.
The quiz itself included two parts: pairwise comparisons and absolute probability estimates, with presentation ordering randomised. As well as these two sets of estimates, there is also some top-level metadata for quiz participants: age band, English language background, highest education level, country of residence. More details on the methods can be found at the end here.
So what did the quiz find?
Probability estimates by phrase
Different people can interpret the same words very differently. Although lots of previous studies have found this, we can now look in more detail at more phrases. The below plot shows the distribution of absolute numerical probability estimates for each phrase (i.e. from Part 2 of the quiz), ordered by mean value.
We can see this variation more clearly if we order the phrases by the variance in responses. The below also shows how phrases compare with common ‘yardstick’ guidelines for official documents by different organisations (you can read more in papers like this one on intelligence usage or this on IPCC usage):
Of all the terms included in the quiz, ‘realistic possibility’ had the highest variance; it really seems to mean a lot of different things to different people. But it’s also a phrase that’s commonly used in worlds of intelligence and prediction (in the UK, the official definition is a 40-50% chance of occurring). And therefore it appears a lot in headlines, such as these from the last year:
Round round
As you’ve probably spotted above, quiz participants liked giving round numbers as their estimates: more than half the absolute percentage estimates given were a multiple of 10, and almost a third ended in a 5. Some people were also particular rounding fans; one in five people who did the quiz rounded all the estimates they gave:
Comparative judgement
As well as asking for absolute judgements, the quiz asked people to compare pairs of words. When shown two phrases and asked which conveys a higher probability, how often do respondents agree? The heatmap below shows the proportion choosing the row term over the column term:
In the bottom left corner, you can see that nobody thought ‘almost no chance’ is more likely than ‘will happen’. But there’s a lot more disagreement in the middle: ‘could happen’ and ‘might happen’ get almost a 50/50 split in terms of which is deemed more likely.
We can also look at how often respondents’ pairwise choices in Part 1 conflicted with their own numerical estimates in Part 2. The below shows the level of consistency in these comparative vs absolute judgements:
It might not have been obvious when taking the quiz, but all respondents were also shown one pair twice, at the start and end of the Part 1 comparisons, but with the order reverse. The below shows the percentage of times that participants disagreed with themselves on the pair ordering for the phrases they were shown:
Who are these people anyway?
People who did the quiz were invited to share some broad demographic data. The below shows the distribution of respondents across age bands, education levels, English language backgrounds, and countries of residence:
How do these demographic factors relate to estimates?
To look at how these factors related to probability estimates, I used a quick beta regression model to estimate variation in responses in terms of fold differences from a reference group (i.e. age 35-44, postgraduate, English first language, UK). Because the phrase and position in which it was presented might influence what people said, fixed effects were used for phrase and quiz position. Some people might also have systematically higher or lower preferences, so random effects were used for individuals.
The below shows the results. There seems to be some trend with education level, and very slightly higher estimates for US vs UK, as well as for phrases that happened to appear later in the randomised list presented to quiz participants.
Of course, a double-edged sword with a large dataset is lots of things will be ‘statistically significant’, but the real question is whether it’s meaningful in reality. Does it make a difference if, all things being equal, an American will give a probability estimate 1.03x higher than a Brit for a given phrase? Or someone age 55-64 will give an estimate 0.97x the value a 35-44 year old would give on average?
These are just some initial visualisations, but hopefully there’s lots more in the full dataset for people to explore and build on in future.











And then there are the scientists who write, in peer reviewed journal papers, that something is the "most likely" explanation for something without any evidence whatsoever. And the even more common practice, also in journal articles, of authors saying that something "might have", "may be" or "could be" the result of something for which they have no evidence at all. But they really want you to believe it. And obviously many smart people continue to fall for it.
The 2019 Paper you linked on Intelligence use of tis didn’t link back to the much earlier work by Heuer (https://www.cia.gov/resources/csi/static/Pyschology-of-Intelligence-Analysis.pdf). It may be interesting to compare Heuer's findings from decades ago to your more current sample.