
Moving the Goalposts
What happens when the bar for success keeps shifting until the answer we wanted is the only one left.
Moving the Goalposts
The day chess stopped counting
On 11 May 1997, in a room on the thirty-fifth floor of a building in midtown Manhattan, the world chess champion Garry Kasparov resigned the sixth and final game of his rematch against Deep Blue, a computer built by IBM, after just nineteen moves. A year earlier in Philadelphia, he had beaten an earlier version of the machine. This time Deep Blue won the match three and a half points to two and a half. It was the first time a computer had beaten a reigning world champion in a match under standard tournament conditions.
For decades before that, chess had been treated as one of the great tests of machine intelligence. In 1958, the researchers Herbert Simon and Allen Newell predicted that within ten years a computer would be the world chess champion. Playing chess well seemed to need foresight, planning and judgement. If a machine could do that, the argument went, we would have to take it seriously.

Then a machine did it, and the verdict changed almost overnight. Deep Blue, critics pointed out, didn't understand chess. It searched around two hundred million positions a second using specialised hardware and a hand-tuned scoring formula. That was brute force, not intelligence. The real test, people began to say, was Go, a far more intuitive game with too many possibilities to search. In March 2016, DeepMind's AlphaGo beat Lee Sedol, one of the strongest players in the world, four games to one in Seoul. The test moved again, to conversation, then to reasoning. Pamela McCorduck, in her history of the field, Machines Who Think, described the pattern decades ago: each time a machine learned to do something, that thing was quietly reclassified as not real thinking. A line usually credited to the computer scientist Larry Tesler, which he said he'd originally phrased slightly differently, puts it neatly: intelligence is whatever machines haven't done yet.
I'll come back to whether those critics were wrong. But the shape of the argument has a name. Moving the goalposts means changing the standard of evidence or success after it has been met, so that a claim you dislike can never be accepted, or a claim you like can never be rejected. The name comes from football, where shifting the posts while the ball is in the air would make scoring impossible. It's sometimes also called raising the bar.
Saving a theory from the facts
The clearest account of why this matters comes from the philosopher Karl Popper. In the essays collected in Conjectures and Refutations (1963), he described his growing unease, as a young man in Vienna, with theories that seemed able to explain absolutely anything. He contrasted them with Einstein's theory of gravity, which made a risky prediction about starlight bending around the sun that could have failed when it was tested in 1919. Popper argued that a scientific claim has to say in advance what would count against it. He also noticed that when a theory's predictions did fail, its defenders often reinterpreted either the theory or the evidence so that the failure no longer counted. Each rescue, he argued, made the theory safer but emptier.

The philosopher Imre Lakatos later softened this picture. He pointed out that real science always patches theories when they meet awkward facts, and that this isn't automatically dishonest. What matters is whether the patches lead somewhere, predicting new things that turn out to be true, or simply absorb each embarrassment as it arrives. He called the second kind a degenerating research programme.
In everyday arguments, goalpost moving comes in a few familiar flavours. One is the escalating demand, where each piece of evidence is met with a request for a different, harder kind. Another is the redefinition, where the meaning of a key word drifts until the evidence no longer fits it, a close cousin of what I described in the post on equivocation. A third works in the opposite direction. Instead of raising the bar for a claim you dislike, you lower it or swap it for a claim you like, so that whatever happened can be counted as success.
Why our own bar is so flexible
In a 1992 study, Peter Ditto and David Lopez told students they were being tested for a made-up enzyme deficiency, supposedly linked to later pancreatic problems. Each student dabbed saliva on a strip of test paper and waited for it to change colour. Some were told that a colour change meant they were healthy, and others that it meant they had the condition. In fact it never could. Students who believed the unchanged strip was good news accepted it quickly. Students who believed it was bad news waited longer for it to change, were more likely to test themselves again, and later rated the test as less accurate. Same evidence, different bar.

The psychologist Ziva Kunda, in a 1990 review titled "The Case for Motivated Reasoning", argued that we rarely believe whatever we like, but when we want a conclusion, we search for and weigh evidence in ways that make reaching it easier. Thomas Gilovich summed up the asymmetry in How We Know What Isn't So (1991): for a conclusion we like, we ask whether we can believe it, and for one we don't, whether we must. Moving the goalposts is what that asymmetry looks like when it is spoken out loud.
Workplaces make it easier still. Success criteria are often never written down, so after the results arrive, everyone remembers the goal slightly differently, and usually in their own favour. Careers and budgets depend on projects being judged as successes. And from the inside, moving the goalposts rarely feels like cheating. It feels like sincerely noticing that the original goal was too narrow.
When the reasons changed after the notes did
A large public example from India is the demonetisation of November 2016. On the evening of 8 November, the government announced that ₹500 and ₹1,000 notes, together around eighty-six per cent of the currency in circulation by value, would stop being legal tender within hours. The aims stated in that announcement centred on black money, counterfeit notes and the financing of terrorism. A widely discussed expectation was that a large share of the cancelled notes, the unaccounted cash, would never come back to the banks.
Within weeks, much of the government's public messaging placed its emphasis elsewhere, on moving India towards a cashless or "less-cash" economy and widening the tax base. In August 2018, the Reserve Bank of India's annual report confirmed that about 99.3 per cent of the cancelled notes by value had been returned. Supporters of the policy pointed to a rise in digital payments and in tax filings as evidence that it had worked. Critics pointed out that those had not been the headline aims at the start.
I'm not trying to settle whether demonetisation was a good policy. Economists still argue about its costs and benefits, and digital payments in India grew for many reasons, including UPI. The point is narrower. When the measure that a decision will be judged on shifts after the results come in, the original question, did it do what it set out to do, quietly disappears. Companies do this every day on a smaller scale.
A redesign that succeeded at something
Picture a team at a budgeting app redesigning its onboarding. The kickoff deck says the goal is activation: the share of new users who connect a bank account and set up a first budget within seven days. Activation sits at 38 per cent, and the team hopes the new flow, with fewer steps and friendlier illustrations, will lift it to 43.
Four weeks after launch, activation is at 38.4 per cent, within the normal weekly wobble. In the review, the lead designer says activation is really a lagging indicator, and points to time spent on the welcome screens, which is up by forty per cent. Nobody asks whether more time on welcome screens is good or bad. The conversation moves on to a small survey, where new users describe the app as "modern" and "friendly" more often than before. By the end of the meeting, the summary slide reads: "New onboarding improves brand perception and engagement."

Each step sounded reasonable, and brand perception does matter. But every new measure was chosen after the team had seen the numbers, and each one was chosen because it had moved. If activation had jumped to 45 per cent, nobody would have argued that it was a lagging indicator.
It works in the other direction too. Imagine the same designer presenting research from eight usability sessions to a sceptical stakeholder. "Eight people isn't enough." So the team runs a survey of six hundred. "Surveys only tell you what people say." So they run an A/B test. "The test ran during a holiday week." Each objection may be fair. But when no possible result would change someone's mind, the objections are about the decision, not the evidence.
The honest version starts before launch. The team writes a short document saying which number they expect to move, by roughly how much, over what period, and what they'll do if it doesn't. They add one or two guardrail measures that must not get worse. When the results come in, they report against that document first, even when the news is dull. If the review teaches them that activation was the wrong goal, they say so plainly, keep the original result on the record, and treat the new goal as a hypothesis to test on the next release, not a verdict on this one.
Sometimes the goal really was wrong
Back to the chess critics. In one sense, they were moving the goalposts, and many did it with a straight face. In another sense, they had learned something real. Chess turned out to be a game that could be beaten largely by very fast search, which meant it was a worse test of general intelligence than people had assumed. Changing the test in response to that discovery isn't a fallacy. It's how measurement improves.
Goodhart's law, named after the economist Charles Goodhart and often summarised in the anthropologist Marilyn Strathern's phrasing, says that when a measure becomes a target, it ceases to be a good measure. Reward a team for time spent in the app and someone will eventually find a way to keep people there that has nothing to do with value. Abandoning a broken metric like that is responsible, not slippery.
So the difference isn't whether the goal changes, but how. Was the change made before or after seeing the results? Is it announced openly, with the old result still visible, or slipped in quietly? Does the new standard apply in both directions, so that it could also count against the conclusion its proposer prefers? And, borrowing from Lakatos, does the new goal predict something we can go and check, or does it only explain away the last disappointment? Since 2013, a growing number of journals have offered registered reports, in which researchers submit their questions and methods for review before collecting data, fixing the goalposts in public before anyone kicks a ball.
How I try to catch it
The first habit is writing the goal down before I look. For anything I ship, I try to note what success would look like and what would tell me I was wrong, even if it's just three lines in the project doc. It feels bureaucratic until the review.
The second is a question I ask when a goal changes: if the first number had come out well, would we be changing it now? If the honest answer is no, we're probably moving the posts. If the answer is yes, because we've learned the measure was flawed whatever it showed, then the change is probably sound.
The third is for when I'm the one being asked for more evidence. I try to ask, kindly, what result would change the other person's mind. If they can name one, we have a plan. If they can't, more research won't settle it, and we need to talk about what's really behind the doubt, which often has nothing to do with the data.
The chess critics of 1997 were partly right and partly protecting a belief they weren't ready to give up. The two often come together, which is exactly why it helps to decide what counts before the game starts. In the next post, on the Texas sharpshooter, I'll look at the stranger trick of drawing the target after the shots.
Further reading: Karl Popper, Conjectures and Refutations (1963) · Peter Ditto and David Lopez, "Motivated Skepticism" (1992) · Ziva Kunda, "The Case for Motivated Reasoning" (1990) · Pamela McCorduck, Machines Who Think (1979) · Thomas Gilovich, How We Know What Isn't So (1991)
The question to askWhat did we say success looked like before we saw the results?