
The Division Fallacy
Why what is true of a group, an average or a whole system isn't automatically true of each part.
The Division Fallacy
The pilot who didn't exist
In the late 1940s, the United States Air Force had a problem it couldn't explain. Its planes were crashing far too often, and not because the engines failed. Pilots were losing control of aircraft that, mechanically, were fine. The cockpits had been designed in 1926 around the measurements of hundreds of pilots, averaged into a single standard body. Perhaps, the engineers wondered, pilots had simply grown bigger since then.
In 1950, researchers at Wright-Patterson Air Force Base in Ohio set out to update the numbers. They measured more than four thousand pilots on 140 dimensions of the body, from thumb length to the distance between eye and ear. A young lieutenant on the team, Gilbert S. Daniels, who had studied physical anthropology, asked a different question. How many pilots were actually average? He picked the ten dimensions that mattered most for fitting a cockpit, and counted a pilot as average on a dimension if he fell within the middle 30 per cent of the range. Then he checked how many of the 4,063 pilots were average on all ten.

The answer was none. Not one. Even on just three of the dimensions, only a few per cent of pilots were average on all three. A cockpit built for the average pilot fitted no actual pilot. As Todd Rose tells the story in The End of Average (2016), the Air Force eventually responded by demanding something that now seems obvious: adjustable seats, pedals and helmet straps, so that equipment fitted the range of real bodies rather than an imaginary one.
The mistake Daniels exposed is the division fallacy. It's the belief that because a whole, a group or an average has a property, each part or member must have it too. The average pilot had arms of a certain length, so each pilot was assumed to have arms close enough. It's the mirror image of the composition fallacy, the subject of the last post, where every star player was excellent and so the team was assumed to be excellent. There we reasoned up, from parts to whole. Here we reason down.
Five is odd and even
Aristotle's version, in On Sophistical Refutations, was a word puzzle. Five is two and three. Two is even, and three is odd. So five is both even and odd. The trick lies in sliding from a statement about the number as a whole to a statement about its parts, or back again, and Aristotle grouped such tricks under "division" and "combination". The medieval logicians who translated him gave us the Latin names, and over time the idea grew into the version we use today.
Logicians now explain it with a distinction between two ways of using a word about a group. A property is distributive when it applies to each member: "the guests are vegetarian" usually means every guest is. A property is collective when it applies only to the group taken together: "the guests filled the hall" doesn't mean each guest filled it. Division happens when a collective property is quietly read as a distributive one. An orchestra can play a symphony. No single violinist can.

One of Aesop's fables, in the form usually told to children, makes the same point in reverse. An old man gives his quarrelling sons a bundle of sticks and asks them to break it. None of them can. Then he unties the bundle and hands them the sticks one at a time, and they snap each one easily. The strength belongs to the bundle, not to the sticks. Anyone who concluded from the unbreakable bundle that each stick was unbreakable would be committing the division fallacy, and would be surprised very quickly.
Averages are the most common way the fallacy shows up today. The Belgian statistician Adolphe Quetelet, in 1835, popularised the idea of l'homme moyen, the average man, as a kind of ideal type of a population. Averages are extremely useful for describing groups. The trouble starts when we treat the average as a description of a person, as the Air Force did.
States that read, people who didn't
The most famous demonstration of division in the social sciences came, by coincidence, in the same year as Daniels' study. In 1950, the sociologist William S. Robinson published a short paper using figures from the 1930 United States census. He looked at the 48 states and compared two numbers for each: the share of the population born abroad, and the share who could read. Across states, the relationship was clearly positive. States with more immigrants had higher literacy.

It would be natural to conclude that immigrants were more likely to read than people born in the country. Robinson showed that the opposite was true. When he looked at individuals rather than states, people born abroad were slightly less likely to be literate. The state-level pattern appeared because immigrants had tended to settle in states, mostly in the industrial north, where the people already living there were highly literate. The states' literacy came largely from their native-born residents, not from the immigrants who lived among them.
Robinson's point was that a correlation between groups, which he called an ecological correlation, cannot be assumed to hold for the individuals inside them. The sociologist Hanan Selvin later named the mistake the ecological fallacy, and it has haunted research ever since, from studies that compare countries' diets with their disease rates to maps of how neighbourhoods voted. A district that votes heavily for one party still contains plenty of people who voted for another.
Why group labels stick to people
We fall for division because sorting things into groups is one of the most useful things our minds do. The psychologist Gordon Allport, in The Nature of Prejudice (1954), argued that thinking in categories is natural and unavoidable. We couldn't function if we treated every chair, dog and stranger as entirely new. The danger he described is that once a category is formed, we let it stand in for the individual, and stop looking at the individual at all.
I see this most clearly in Indian hiring. The name of a college often does an enormous amount of work on a CV. A degree from one of the IITs, the IIMs or NID tells you something real about a group, because getting in is fiercely competitive and the average graduate is strong. But it's a fact about the group. It says much less about the particular person in front of you, in both directions. There are average graduates of excellent institutions, and excellent graduates of institutions nobody has heard of. An interview that stops at the college name is reasoning from the whole to the part.
The same shortcut shows up with brands and companies. If a company is known for good design, we assume each of its screens is well designed and skim past the one that isn't. If a team is known for being slow, the fast person on it gets read as slow too. The group's reputation is easy to see. The individual takes effort.
A 92 per cent that hid a 61
Picture a team that runs a quarterly usability benchmark on its banking app. Two hundred participants complete five core tasks, from checking a balance to paying a bill. The headline number is excellent: 92 per cent of tasks completed successfully, up from last quarter. The app's store rating sits at 4.6. In the review, the conclusion is clear. Users are happy, the app works, and the team can move on to new features.
Then a researcher breaks the results apart. Participants using the app in English completed 95 per cent of tasks. Participants using the Hindi interface, a fifth of the sample, completed 61 per cent of the bill-payment task, because a key button label had been translated into a phrase most of them didn't recognise. People using screen readers struggled with the same flow for a different reason. None of this was visible in the average, because the large, happy group swamped the small, struggling ones.

The division step is the jump from "the app performs well" to "the app performs well for each kind of user and each task". The whole really did perform well. The parts did not all share that property.
The same jump happens with AI. A model that scores highly on a public benchmark is good on average across thousands of benchmark questions. That doesn't mean it's good at your task, in your language, with your users' messy inputs. When I evaluate an AI feature now, I want to see results for the specific cases we care about, not just the headline score.
The honest version of the review reports the average and the spread. "Overall task success was 92 per cent. It was lowest for bill payment in Hindi, at 61 per cent, and for screen-reader users. Here's what we think is causing it." It's a less comfortable slide, and a far more useful one.
When the whole does speak for its parts
Not every move from whole to part is a fallacy. Distributive properties carry down perfectly well. If a crate holds only Alphonso mangoes, every mango in it is an Alphonso. If an entire building is made of brick, any wall you pick is brick. If every member of a committee was elected, then each member was elected. The question is always whether the property belongs to each member or only to the group.
Group-level facts are also perfectly good for group-level decisions. A city can plan its bus routes from average demand. A public health campaign can be aimed at the districts where a disease is most common. An insurer can price risk across thousands of people. None of these claims to know about any one person, and that's what keeps them honest.
Even for individuals, a group fact can be a sensible starting guess when you have nothing better. If you know nothing about a stranger except that they live in a city where most people speak Kannada, guessing they speak Kannada isn't foolish. The fallacy is treating that guess as certain, or refusing to update it once you can see the person. When you can look at the part directly, look at the part.
How I try to catch it
My first question is whether a claim describes each member or only the group as a whole. "Our users love the app" is usually a statement about an average. I try to translate it into what it actually says: most users, on average, rate it highly. Put that way, it's obvious that some may not.
The second habit is to ask for the spread, not just the centre. Who is at the edges? Which segment, language, device or task is doing worst? A single number hides exactly the people a designer most needs to see.
The third is about design itself. Wherever I can, I try to design for the range rather than the middle: text that can grow, layouts that adapt, flows that work for the first-time user as well as the expert. It's the same answer the Air Force reached.
Daniels never found the average pilot because there wasn't one. The average was a true fact about four thousand men, and a false description of every single one of them. The adjustable seat was the admission that the group and the person are different things. In the next post, on the appeal to authority, I'll look at another way we let a reputation speak for a specific claim.
Further reading: Todd Rose, The End of Average (2016) · William S. Robinson, "Ecological Correlations and the Behavior of Individuals" (1950) · Gilbert S. Daniels, "The 'Average Man'?" (1952) · Gordon Allport, The Nature of Prejudice (1954) · Aristotle, On Sophistical Refutations
The question to askIs this true of every part, or only of the whole?