Back to Blogs

False Equivalence

When two things that share one feature are treated as if they were equal in every way that matters.

False Equivalence

A one-minute statement in a biology class

In January 2005, ninth-grade biology students in the small town of Dover, Pennsylvania, heard a short statement read aloud in class. Their teachers had refused to read it, so school administrators came in to do it instead. The statement, required by a vote of the school board the previous October, told the students that Darwin's theory of evolution was a theory, not a fact, that it had gaps, and that intelligent design was an explanation of the origin of life that differed from Darwin's view. Students who wanted to know more were pointed to a book in the school library, Of Pandas and People.

On its face, the statement sounded fair. Here are two explanations. Keep an open mind. Read both. But eleven parents, led by Tammy Kitzmiller, sued the school district. In a six-week trial in Harrisburg in 2005, the court heard from scientists, philosophers and historians. One of the most memorable pieces of evidence came from the drafts of Of Pandas and People itself. Earlier versions had used the word "creationists". Later versions replaced it with "design proponents", and one draft caught the swap halfway, leaving the strange hybrid "cdesign proponentsists" in the text.

On 20 December 2005, Judge John E. Jones III ruled that intelligent design was not science and that the school board's policy was unconstitutional. His opinion ran to 139 pages. The statement read to the students had taken about a minute.

That gap is the subject of this post. The Dover statement put two things side by side as if they were the same kind of thing: two "theories", two "explanations", two sides of a question. One was a scientific theory supported by more than a century of evidence from fossils, genetics and direct observation. The other was a proposal that the court found did not work as science at all. Treating them as equals is called false equivalence: claiming that two things are equal, or comparable in the way that matters, because they share some feature, when in fact they differ greatly in size, evidence, seriousness or kind. In journalism, the same mistake goes by the name false balance.

Sharing a feature isn't sharing a weight

At its core, false equivalence takes a real similarity and stretches it too far. Both have sugar. Both are opinions. Both made mistakes. Both are bugs. Each of these can be perfectly true and still lead to a wrong conclusion, because the shared feature isn't the one that matters for the question being asked.

Logic students meet the skeleton of this mistake early. "All cats have four legs. My dog has four legs. So my dog is a cat." Medieval logicians called this the fallacy of the undistributed middle. The shared term, having four legs, connects the dog and the cat without telling you whether the dog belongs with the cats. Sharing a property is not the same as belonging to the same category, let alone being equal.

In everyday arguments the stretch is usually about size. A neighbour who parks badly once and a neighbour who blocks the gate every morning both "cause problems". An app that crashed for five minutes last year and an app that loses your data both "have had reliability issues". A snack that has a little sugar and a drink that is mostly sugar both "contain sugar", a line I've heard used to defend both. The words line the two things up neatly. The magnitudes do not line up at all.

A close relative is the move people often call whataboutism. Someone is criticised for a serious failing and replies by pointing to a minor failing of the critic, as if the two cancelled out. The reply may be true. It's the equivalence that is false.

Why equal time feels like fairness

Most of us grow up being told to hear both sides, and that's a good rule for disagreements between people. But it slides easily into assuming that every question has two sides of roughly equal weight. Presenting two positions with the same airtime, the same font size or the same number of slides looks neutral. It feels like the fair thing to do. In fact it's a decision, and often a misleading one.

Our minds also struggle with scale. In a study from the early 1990s, William Desvousges and colleagues asked people how much they would pay to protect migrating birds from drowning in uncovered oil ponds. Different groups were told the ponds threatened 2,000, 20,000 or 200,000 birds. The average amounts people offered were almost identical, around $80 in each case. People responded to the picture of a single oily bird, not to the number. Psychologists call this scope insensitivity, and it's the soil false equivalence grows in. If we don't feel the difference between two thousand and two hundred thousand, it's easy to accept that two very different problems deserve the same attention.

Labels finish the job. Once two things are filed under the same word, "a theory", "a risk", "a bug", "an opinion", they borrow each other's status. The label is shared, so the weight seems shared too.

Two sides, one weight of evidence

The clearest documented case of false balance comes from coverage of climate change. In 2004 the researchers Maxwell Boykoff and Jules Boykoff published a study called "Balance as bias". They looked at articles about global warming in four of the most influential American newspapers, The New York Times, The Washington Post, the Los Angeles Times and The Wall Street Journal, between 1988 and 2002. In more than half of them, the view that humans were contributing to warming and the view that natural variation explained it were given roughly equal attention. At the time, the scientific assessments were already clear about which view the evidence supported. Journalists following a good rule, present both sides, had ended up giving readers a misleading picture of where the science stood.

Broadcasters have struggled with the same problem. In 2011 the BBC Trust published a review of the BBC's science coverage by the geneticist Steve Jones. He warned that an over-rigid reading of the rules on impartiality could give undue attention to marginal opinions, setting a lone dissenter against a broad scientific consensus as if they were equals. In 2018 the BBC sent guidance to its journalists on reporting climate change, telling them that they did not need to include a sceptic to balance every discussion of the established science.

None of this means fringe views should be silenced, or that consensus is always right. Science changes, and dissenters are sometimes vindicated. It means that the shape of the coverage should match the shape of the evidence. Two chairs on a television set say "these positions are equal". Sometimes they are. Often they aren't.

The sticky notes that were all the same size

Picture a team that has just finished a round of usability tests on a banking app. In the synthesis session, every finding goes on a sticky note, all the same size and the same yellow. "The icons look dated." "Wants a dark mode." "Couldn't find where to change her address." "Sent money to the wrong person, because the confirmation screen showed the payee's nickname instead of their name and account number."

The team dot-votes on what to fix first. The dated icons win easily. Nearly every participant mentioned them, and they're easy to fix. The wrong-payee mistake gets one vote. Only one participant ran into it.

Nothing about the session was dishonest, yet the outcome is wrong, and the false equivalence happened before anyone voted. Putting every finding on an identical note said, visually, that every finding was the same size. An outdated icon and money sent to a stranger became one unit each. Dot voting then measured how many people noticed something, which is a different question from how much it matters.

The honest version separates the questions. How often does each problem happen, how bad is it when it does, and can people recover on their own? Jakob Nielsen proposed rating usability problems on exactly those factors, frequency, impact and persistence, and many teams use a simple severity scale from cosmetic to catastrophic. Rated that way, the wrong-payee problem goes straight to the top, because a single occurrence can cost someone real money and can't be undone. The icons are still worth fixing. They just aren't the same size of problem.

The same trap shows up when teams compare AI models. Two models can both score 92% on a test set and still be very different, if one gets harmless questions wrong and the other gets dangerous ones wrong. The single number makes them look equivalent. Reading the failures tells you they aren't.

When like really is like

Comparison is not the problem. Much of good thinking consists of noticing that two things really are alike in the way that matters. Aristotle, in the Nicomachean Ethics, describes justice partly in these terms: treating equals equally and unequals unequally, in proportion to the differences that are relevant. The hard part is that final word, relevant.

Indian constitutional law has a well-known test for exactly this question. Article 14 of the Constitution guarantees equality before the law, but since the early 1950s the Supreme Court has held that it does not forbid the state from treating groups differently. What it forbids is unreasonable classification. A classification is allowed when it rests on a real difference that can be understood, which the courts call an intelligible differentia, and when that difference has a rational connection to the purpose of the law. Treating two groups the same when they differ in a way that matters can be as unjust as treating them differently when they don't.

That's a good test outside a courtroom too. Two designs, two bugs, two candidates or two arguments can legitimately be compared and even treated identically, as long as the feature they share is the one the decision depends on. When the comparison rests on a feature that doesn't matter for the question, it's no longer a comparison. It's a sleight of hand.

How I try to catch it

The first question I ask when two things are presented as equal is: similar in what way, and does that way matter here? Usually the shared feature is real. The question is whether it's the one the conclusion depends on.

The second habit is to put both things on the same scale. How many people, how much money, how likely, how reversible? A sentence like "both teams missed targets" often hides a difference of ten times, and writing the numbers next to each other is usually enough to reveal it.

The third is to separate "both" from "equally". Both options have risks. Both sides have made mistakes. These sentences are almost always true and almost never the end of the analysis. The useful follow-up is: and how do they compare?

The statement read to those students in Dover took about a minute and sounded like fairness. Untangling it took a court six weeks and a judge 139 pages. False equivalence is quick to say and slow to undo, which is why it's worth catching early. In the next post, on false analogy, I'll look at its close cousin: an argument that two things are alike in one respect, and so must be alike in another.

Further reading: Maxwell T. Boykoff and Jules M. Boykoff, "Balance as bias: global warming and the US prestige press" (2004) · Edward Humes, Monkey Girl: Evolution, Education, Religion, and the Battle for America's Soul (2007) · Daniel Kahneman, Thinking, Fast and Slow (2011) · Jakob Nielsen, "Severity Ratings for Usability Problems" (1994) · Steve Jones, BBC Trust Review of Impartiality and Accuracy of the BBC's Coverage of Science (2011)

The question to askSimilar in what way, and does that way matter here?