Do we need more 'aluminium-standard' evidence?
Digging into a new book on scientific evidence
In 1971, a doctor named Archie Cochrane showed a group of cardiologists the results of a clinical trial he had been running. The response was furious: ‘Archie, we always thought you were unethical. You must stop this trial at once’.
The reason for their anger? Cochrane had been testing whether it was better for patients to recover from a heart attack at home or in hospital; his results showed hospital was safer. Why on earth had he thought it ethical to leave people at risk at home, when the alternative was obviously better?
But the trial wasn’t quite what it seemed. When the anger died down, Cochrane revealed he had played a trick. The study had actually shown that people who recovered at home had better outcomes than those who remained in hospital. Given they apparently placed so much weight on such a trial, Cochrane challenged them to argue with the same fervour that coronary care units should be stopped immediately1.
‘There was dead silence,’ Cochrane recalled, ‘and I felt rather sick because they were, after all, my medical colleagues.’
I was reminded of this anecdote after reading Helen Pearson’s fascinating new book Beyond Belief: How Evidence Shows What Really Works.
Cochrane appears early in the book, running his first clinical trial as a prisoner of war in 1941. Looking at the swollen legs of his sickly fellow prisoners, he wondered if the cause was beriberi, which resulted from a lack of the B vitamin thiamin. He got hold of some yeast (which contained the vitamin) and vitamin C tablets (which did not) on the camp black market, then recruited twenty prisoners to test whether the yeast would solve the problem. Sure enough, those taking yeast saw their swollen legs resolve.
In the decades that followed, Cochrane and others would challenge a medical culture that often relied on hierarchy and anecdote, a case of ‘eminence-based medicine’. Instead, they argued for more rigorous, evidence-based decision-making. This movement would lead to the creation of the Cochrane Collaboration, which has produced thousands of systematic reviews of available evidence. It would also inspire similar efforts across other fields, from education to conservation.
Pearson’s book shows how stuttering progress can be. Good studies often fail to change behaviour if the right incentives aren’t in place. There are several tales of frustrated researchers, who belatedly realise that their carefully curated evidence isn’t leading to effective change.
I’ve written a wider review of Beyond Belief for this month’s Literary Review, but in this post, I wanted to briefly focus more on one element of the book that stood out for me.
When it comes to evidence, I’ve noticed some groups will demand ‘gold standard’ evidence for interventions and policies they disagree with, while accepting anecdote and speculation for topics they support. If they want to oppose something, they set the bar higher and higher, insisting that unless this bar can be reached, we do not know anything and cannot do anything. In effect, it’s a form of weaponised certainty.
Decision-making is always at risk of the politician’s fallacy (i.e. ‘We must do something. This is something. Therefore, we must do this.’) But I’ve noticed that demanding more and more data can become a way for decision-makers to avoid making a decision. And after all, in a fast-moving situation, deciding not to act when some emerging evidence is available is a decision in itself.
For many problems that matter, we cannot run a neat randomised controlled trial like we can for a medical treatment against a common disease in a well-defined population. And even if we could, for a dynamic, complex challenge, the results wouldn’t necessarily translate from one location to another, or from one time period to another.
In Beyond Belief, Pearson talks about ‘aluminium-standard’ evidence: faster, cheaper and often more feasible than the ‘gold-standard’, but still good enough to promptly inform the decision being made. One example recounted in the book involves woodland caribou. In 2021, Parks Canada was deciding whether to invest $24 million in a major caribou breeding programme in the Rocky Mountains. The problem was that there weren’t any evaluations or experiments looking at whether captive-breeding was effective for caribou.
So instead the researchers split the one big question into smaller hypotheses:
The caribou population will be lost if they do nothing.
Other threats to the population, such as wolves, have been mitigated so that if the caribou are restored they should thrive.
Other rescue strategies, like moving the caribou elsewhere or fencing them off from predators, are unfeasible or unlikely to be effective.
The breeding strategy is technically feasible.
They couldn’t say conclusively whether the programme would work, but if these underlying hypotheses were supported by evidence, it would provide confidence in the overall decision. After hundreds of person-hours of evidence review and synthesis, the team concluded that there was sufficient evidence in favour of the four hypotheses to green-light the programme.
An aluminium-standard approach can come with more caveats and uncertainties, but it shows that progress can still be made by thinking through the underlying processes behind the outcome we’re interested in. As Chris Whitty once put it: ‘An 80% right paper before a policy decision is made is worth ten 95% right papers afterwards, provided the methodological limitations imposed by doing it fast are made clear’.
Late last year, after Beyond Belief would have gone to print, the first calves were born in captivity. There’s a long way to go, but it’s a promising early sign, and – as Pearson notes – the team deserve credit for taking a disciplined approach to such a difficult problem.
When I was working on Proof, I spoke to economist Rocío Titiunik, who made the point that there are broadly two types of researchers when it comes to understanding cause-and-effect. The first type starts with a question and remains committed to answering it, even if they can’t run a ‘gold-standard’ RCT or gain insight from a ‘natural experiment’. The second type starts by identifying research methods that are ‘gold-standard’, then looks at the subset of questions these methods can be applied to.
Both have an important role in research, but not always the same role in decision-making. There is a difference between answering the questions that need to be answered and answering the questions that can be conclusively answered. As well as striving for better evidence in theory, we should also strive for better outcomes in practice.
Cover image: Christoph Nolte
Note: this was back in the 1970s, and the hospital vs home finding would not be the same today.


Absolutely true. A tactic of lobbyists to oppose policy change .. say on climate .. deny the problem exists. Say they will change their minds when the evidence comes in. Then set a high bar, for what standard of evidence they'll accept... so high they know it's impossible for scientists to ever get it. Policy makers and media often fell for it. Naomi Oreskes, "Merchants of Doubt".
In praise for Ignaz Semmelweis
Ignaz Philipp Semmelweis was a Hungarian physician and scientist of German descent who was an early pioneer of antiseptic procedures and was described as the "saviour of mothers".
Postpartum infection, also known as puerperal fever or childbed fever, consists of any bacterial infection of the reproductive tract following birth and in the 19th century was common and often fatal.
Semmelweis demonstrated that the incidence of infection could be drastically reduced by requiring healthcare workers in obstetrical clinics to disinfect their hands.
In 1847, he proposed hand washing with chlorinated lime solutions at Vienna General Hospital's First Obstetrical Clinic, where doctors' wards had thrice the mortality of midwives' wards. The maternal mortality rate dropped from 18% to less than 2%, and he published a book of his findings, Etiology, Concept and Prophylaxis of Childbed Fever, in 1861.
Semmelweis's observations conflicted with the established scientific and medical opinions of the time and his ideas were rejected by the medical community. He could offer no theoretical explanation for his findings of reduced mortality due to hand-washing, and some doctors were offended at the suggestion that they should wash their hands and mocked him for it.
While his hygiene requirements (hand washing with chlorinated lime) were sometimes mockingly referred to as a "Jewish superstition" or "Jewish" by opponents, this was a derogatory, anti-Semitic slur used to discredit him and his ideas.
In 1865, the increasingly outspoken Semmelweis allegedly suffered a nervous breakdown and was committed to an asylum by his colleagues. In the asylum, he was beaten by the guards. He died 14 days later from a gangrenous wound on his right hand that may have been caused by the beating.
His findings earned widespread acceptance only years after his death, when Louis Pasteur confirmed the germ theory of disease, giving Semmelweis's observations a theoretical and scientific explanation, and Joseph Lister, acting on Pasteur's research, practised and operated using hygienic methods with great success.