Back to Blogs

Cherry Picking

How choosing only the evidence that agrees with us can make a weak case look like a certain one.

Cherry Picking

Three heart attacks that weren't in the paper

In November 2000, the New England Journal of Medicine published the results of a large trial called VIGOR. It compared rofecoxib, a new painkiller sold by the drug company Merck under the name Vioxx, with an older painkiller, naproxen, in about eight thousand people with rheumatoid arthritis. The headline finding was good news for Vioxx: it caused significantly fewer serious stomach problems, such as bleeding ulcers. The paper did report that more people on Vioxx had heart attacks, and it discussed the possibility that naproxen might protect the heart, which would make Vioxx look worse by comparison. Vioxx went on to be prescribed to millions of people.

In September 2004, Merck withdrew Vioxx worldwide after another trial found that people taking it for a long time had a higher risk of heart attacks and strokes. A year later, in December 2005, the journal's editors published an unusual note, an "expression of concern" about their own paper from 2000. They wrote that three additional heart attacks in the Vioxx group had been known to some of the authors before publication but were not included in the article, and that data about them had been deleted from a draft shortly before it was submitted. Merck and the authors said these events happened after a cut-off date that had been set in advance for counting heart problems. The editors' view was that leaving them out made the difference in heart risk between the two drugs look smaller than it was.

I'm not in a position to judge every detail of that dispute, and courts and journals argued about it for years. But it shows, in a very high-stakes form, how much depends on which evidence makes it onto the page. Three events out of thousands of patients sound like nothing. In a comparison of heart risk, they mattered.

This is cherry picking: presenting only the evidence that supports your conclusion and quietly leaving out the evidence that doesn't. Logic textbooks also call it the fallacy of incomplete evidence or of suppressed evidence. The name comes from fruit picking. Someone who picks only the ripest cherries can show you a perfect basket, but the basket tells you nothing about the tree. What makes it slippery is that every piece of evidence shown can be completely true. The distortion is in the selection.

Where are the ones who drowned?

In 1620, the English philosopher Francis Bacon wrote Novum Organum, a book about how people could reason better about the natural world. In it, he retells an old story, which goes back through Cicero to the Greek thinker Diagoras. A man is taken to a temple and shown a wall of paintings offered by sailors who had prayed during a storm at sea and survived. Surely, he is told, this proves the power of the gods. He asks where the paintings are of those who prayed and drowned anyway. Bacon's point was that the human mind is drawn to the cases that confirm what it already believes, and passes over the cases that don't, even when those are more numerous.

Cherry picking has many costumes. Quote mining takes a sentence out of a long argument so that it seems to say the opposite of what its author meant. Film posters can do a mild version of this: a review calling a film "a brilliant waste of a brilliant cast" could end up on the poster as simply "brilliant". Choosing a convenient time window is another: start a chart in an unusually bad month and almost any line looks like a recovery. In 1937, the American Institute for Propaganda Analysis listed a version of it, which they called "card stacking", among seven common tricks of persuasion: using only the facts and illustrations that favour your side, like a dealer arranging a deck.

Sometimes the cherry picking isn't done by any one person but by a whole system. In 1979, the psychologist Robert Rosenthal described the "file drawer problem". Studies that find an exciting result get published. Studies that find nothing tend to stay in a drawer. Each published paper might be honest, but the published literature as a whole becomes a basket of the ripest cherries. There's a close cousin, survivorship bias, where the evidence that didn't survive has vanished before anyone can choose, and it will get a post of its own later in this series.

Why a full basket feels convincing

In 1979, the psychologists Charles Lord, Lee Ross and Mark Lepper at Stanford gave students descriptions of two studies about whether the death penalty deters crime. One study appeared to show that it did; the other appeared to show that it didn't. Half the students already supported the death penalty and half opposed it. You might expect that reading balanced, mixed evidence would pull both groups towards the middle. Instead, each group rated the study that agreed with them as more convincing and better designed, picked holes in the other one, and came away more confident in what it already believed. Nobody was hiding anything. Their minds were sorting the evidence for them.

That's what makes cherry picking so common. Most of the time it isn't a deliberate deception. It's what an ordinary mind does when it wants something to be true. The examples that agree with us feel vivid and relevant. The ones that don't feel like exceptions, or flawed, or not quite comparable. I'll come back to this pull, the tendency to look for what confirms us, in a later post on confirmation bias.

There's also a structural reason. Every summary must select. A twelve-slide deck can't contain every interview, and a three-minute pitch can't include every data point. So the act of choosing is unavoidable, and the line between choosing well and choosing what flatters you is easy to cross without noticing. Incentives do the rest. Drug companies want approvals, researchers want publications, and designers want their ideas on the roadmap.

The trials that never made it into print

In 2008, the psychiatrist Erick Turner and his colleagues published a study in the New England Journal of Medicine that showed cherry picking at the scale of a whole field. When drug companies in the United States seek approval for a medicine, they must submit the results of their trials to the Food and Drug Administration, so the FDA had a complete record of 74 trials of twelve antidepressants. Turner's team compared that record with what had appeared in medical journals. Of the trials the FDA regarded as positive, almost all had been published. Of the trials with negative or doubtful results, most had either never been published or had been written up in a way that made the outcome sound positive. A doctor reading the journals would have seen that 94% of the trials were positive. According to the FDA's own reviews, only about half were.

The drugs weren't useless. But the published evidence made them look more effective than the full evidence did. Since then, many journals have required trials to be registered before they begin, so that results which disappear can at least be noticed.

India has an everyday version that most students will recognise. For years, coaching institutes for exams like the JEE engineering entrance and the civil services have advertised with rows of smiling toppers, sometimes with the same successful candidate appearing in ads for more than one institute. A topper may have attended only a short test series or an interview programme there. The thousands who enrolled and didn't get through never appear on the billboard. In 2024, India's Central Consumer Protection Authority issued guidelines on coaching advertisements that, among other things, ask institutes to state clearly which course a featured successful candidate actually took.

Picture a research synthesis with a favourite idea

Picture a team that has just finished twelve interviews about how small shop owners keep track of credit given to regular customers. The product manager already believes the answer is a reminder feature that sends payment nudges on WhatsApp. The synthesis session begins with a wall of sticky notes, and over two hours a story takes shape. When the findings deck goes out, it has four striking quotes about forgetting who owes what, all from two talkative participants. One chart of app usage starts in March, conveniently just after a dip in February.

Nobody in that room lied. Every quote is real and the chart is accurate. But the deck leaves out the seven participants who said they kept a perfectly good paper ledger, the three who said they would feel awkward sending a reminder to a neighbour, and the two months before March. Anyone reading it would think the case for reminders was overwhelming. The full wall of sticky notes says something more interesting: forgetting is a real problem for some shop owners, and asking for money is a delicate social act for many of them.

The honest version doesn't have to be dull. It says how many people raised each theme, "four of twelve", so readers can judge the weight of a quote. It includes at least one strong quote that cuts against the team's favourite idea. It shows the chart's full range, or explains why it starts where it does. And it keeps the raw notes somewhere the whole team can see, so anyone can check what was left on the wall.

When choosing is the whole point

Not all selection is cherry picking. A portfolio shows a designer's best work, and everyone knows that's what a portfolio is. A highlights reel of a cricket match shows the sixes and the wickets, not every dot ball, and nobody is fooled. Choosing a vivid example to illustrate a point is good writing, as long as the example is typical, or you say plainly that it isn't. Even removing data can be fine: researchers often exclude results from broken equipment or participants who misunderstood the task, as long as the rules for doing so were set before anyone saw which way the results went.

The difference comes down to two questions. Does the audience know that a selection has been made, and on what basis? And would the conclusion change if the evidence left out were put back in? A portfolio passes both tests. A deck with four hand-picked quotes presented as "what users told us" fails both.

How I try to catch it

Charles Darwin described a habit in his autobiography that I've tried to copy. Whenever he came across a fact or an observation that went against his general conclusions, he wrote it down straight away, because he'd noticed that such facts slipped from memory much more easily than the welcome ones. So the first thing I do is keep a "doesn't fit" list during research, and I make myself put at least one item from it in the final deck.

The second is to decide what counts before I look. If I'm judging whether a feature worked, I choose the metric and the time window first, and then I look at the numbers. It's the designer's version of registering a trial before it starts.

The third is a question I ask whenever I'm shown a neat basket of evidence, including my own: what does the evidence that was left out say? Sometimes the answer is "nothing different", and the case stands. Sometimes it's three heart attacks.

In the next post I'll look at a quieter trick, one that hides not in which evidence we choose but in which meaning of a word we use: equivocation.

Further reading: Francis Bacon, Novum Organum (1620) · Erick Turner et al., "Selective Publication of Antidepressant Trials and Its Influence on Apparent Efficacy" (2008) · Ben Goldacre, Bad Pharma (2012) · Darrell Huff, How to Lie with Statistics (1954) · Charles Darwin, The Autobiography of Charles Darwin (1887)

The question to askWhat does the evidence that was left out say?