Three doors in front of you: switch, or stay?
The rules, on the table: three doors — behind one, a car; behind two, goats. You pick door 1. The host, who knows where the car is, opens door 3 — a goat bleats at you. He grins: 'Care to switch to door 2?'
Don't read on yet. Answer in your head: switch? Stay? Doesn't matter?
If your answer was 'two doors left, fifty-fifty, doesn't matter' — congratulations, you just fell into the same pit as nearly a thousand PhDs. The right answer: switch, and you win with probability 2/3; stay, and it's only 1/3.
The proof takes three lines: your first pick lands on the car with probability 1/3 — in that case switching loses for sure; your first pick lands on a goat with probability 2/3 — and then the host is forced to open the other goat door, so switching wins for sure. So 'switching wins' = 'your first pick was wrong' = 2/3. Statistician Steve Selvin posed the problem and gave this exact solution in a 1975 letter to The American Statistician — a full 15 years before the whole country started shouting about it.

The witness says it was blue: what probability would you put on it?
Next stop, the courtroom. In a certain city, 85% of the cabs belong to the Green company and 15% to the Blue company. One night a cab is involved in a hit-and-run. A witness swears: 'It was blue.' The court tests him under comparable conditions: he identifies colors correctly 80% of the time. Question: what's the probability the cab really was blue? Name a number in your head first.
Hands up if you said 80% — when Tversky & Kahneman ran this experiment, the median and modal answer were both 80%. Bayes' answer: about 41%. P(blue | says blue) = (0.15×0.80) / (0.15×0.80 + 0.85×0.20) = 0.12/0.29 ≈ 41% — in other words, the identification is more likely wrong than right. People took 'witness reliability' straight for the answer and threw away the base rate — blue cabs are only 15% of the fleet. Vivid, concrete case evidence steamrolls abstract statistical background: that's base-rate neglect.
Think only juries fall for this? In 1978, Casscells and colleagues reported in the New England Journal of Medicine a little quiz at Harvard teaching hospitals: a disease has a prevalence of 1/1000, the test has a 5% false-positive rate (and catches nearly every true case); a random individual tests positive — what's the probability he actually has the disease? Answer first, then read on. Among 60 medical students, residents, and attending physicians, the most common answer was 95% (45% said so), and only 18% got it right: about 2%. Here's the math: out of 1,000 people, roughly 1 true positive; among the 999 healthy, roughly 50 false positives; so among the positives, the truly ill = 1/51 ≈ 2%. It stings more that a 2014 replication in JAMA Internal Medicine found the correct-answer rate still only about 1 in 4, with 95% still the most common answer.
The root cause has a name — confusion of the inverse: taking P(evidence | hypothesis) directly for P(hypothesis | evidence). Doctors read '95% of the diseased test positive' as '95% of positives are diseased'; the courtroom version is the prosecutor's fallacy — 'if innocent, a DNA match would be extremely unlikely' gets read as 'a match means all but certainly guilty.'

Black 26 times in a row: is red 'due'?
August 18, 1913, the Monte Carlo casino. On one roulette table, black has already come up 15 times in a row. You're clutching your chips — do you bet red?
The gamblers of the day voted with real money: as Huff & Geis famously recount, from about the 15th black onward, people piled heavy bets on red — 'black has come up so many times, red is due' — and kept doubling down after every loss. Black went on to hit 26 in a row, and the casino raked in millions of francs.
What went wrong? Every spin of the wheel is an independent event — it doesn't remember, and it doesn't owe any color anything. Yes, the prior probability of a given color hitting 26 straight is a minuscule 2^-26 ≈ 1 in 67 million — but that's the probability computed before the run starts; once black has come up 25 times, the chance of red on spin 26 is not one bit higher. This intuition that 'random sequences self-correct' is the gambler's fallacy (a.k.a. the Monte Carlo fallacy); Tversky & Kahneman traced it to belief in the 'law of small numbers' — wrongly expecting short sequences to mirror long-run proportions.
The gambler's fallacy has a mirror twin: the hot hand — here people bet on continuation instead of reversal. In 1985, Gilovich, Vallone & Tversky surveyed 100 basketball fans: 91% believed a player is 'more likely to score after hitting two or three in a row than after missing two or three,' and 84% said the ball should go to whoever just hit several straight; fans estimated a 50%-shooter would hit 61% right after a make and 42% right after a miss. But the actual 1980-81 records of 9 Philadelphia 76ers regulars: a weighted average of just 46% after 3 straight makes, versus 56% after 3 straight misses (51% after 1 make, 54% after 1 miss); Celtics free-throw data and a 100-shot controlled experiment with Cornell varsity players likewise showed no positive correlation. GVT's conclusion: people badly underestimate how often pure randomness naturally produces streaks, and mistake normal clustering for a 'state.' The gambler's fallacy bets on reversal, the hot-hand fallacy bets on continuation — two faces of the same coin: neither accepts that randomness streaks all by itself.

自测 · 学完检查一下
想真正动手做题、记进度、攒连胜?到互动课里练。
Office raffle: three boxes, one holds the grand prize, two are empty. You pick box A. The host, who knows what's inside each, opens box C on the spot — empty — then asks if you want to switch to box B. What's the probability that switching wins the prize?
答案:2/3 — you picked wrong with probability 2/3, and whenever you did, switching wins for sure
Switching = winning whenever your first pick was wrong. Your first pick lands on the prize with probability 1/3 (switching loses for sure); it lands on an empty box with probability 2/3 (the host, knowing the contents, is forced to open the other empty box, so switching wins for sure) — so switching wins 2/3 of the time, double the odds of staying. 'Two left, so half-half' fails because it treats the host's opening as random: he only ever opens an empty box, and that act carries information. This is the Monty Hall problem; statistician Steve Selvin first posed it and gave the correct solution in a 1975 letter to The American Statistician. (Sources: Selvin 1975; Krauss & Wang 2003)
Why does switching win 2/3, rather than 'two doors left, 1/2 each'? Which explanation hits the mark?
答案:The host knows where the car is and only ever opens a goat door — which door he opens is constrained by your first pick and carries information, so the entire 2/3 probability from outside your first pick is pressed onto the one unopened door
The host isn't opening doors blindly: he knows where the car is, only ever opens a goat door, and which one he opens is constrained by your first pick — a conditional act that carries information, so the 2/3 probability outside your first pick lands entirely on the remaining unopened door. Conversely, if the host didn't know and happened to reveal a goat at random, then switching and staying would each be 1/2 — change the premise and the answer changes. As for vos Savant: she never conceded, and on July 21, 1991 The New York Times confirmed on its front page that she was right. (Sources: Selvin 1975; Tierney, NYT 1991)
True or False: Swap the host — this one has no idea where the car is, opens a door at random, and it just happens to show a goat. Switching still wins with probability 2/3.
答案:False
False. The 2/3 advantage rests on a key premise: the host knows what's behind the doors, must open a goat door, and must offer the switch — his opening is a constrained, conditional act. If the host merely opens a door at random and it happens to show a goat, switching and staying are each 1/2. Same physical act, completely different probabilities depending on whether the opener knows — this premise is how you spot every 'fake Monty Hall' scenario at a glance. (Source: Selvin 1975, The American Statistician)
In 1990 vos Savant wrote in her column that 'switching doubles your odds to 2/3.' She then received about ten thousand letters — nearly a thousand from PhDs — with 92% of general readers' letters against her. How did the battle end?
答案:In 1991 The New York Times confirmed on its front page that she was right; she rallied classrooms across America to run the experiment, and the simulations uniformly supported switching at 2/3
On July 21, 1991, The New York Times ran John Tierney's front-page story confirming vos Savant was right; she also rallied classrooms across America to run the experiment, and the simulations uniformly supported switching at 2/3. Look back at those ten thousand letters: 92% of general readers opposed, 65% of university letters opposed, nearly a thousand PhDs adamant — professional training does not immunize against conditional-probability illusions, which is exactly why the affair became the landmark case of 'collective failure of probabilistic intuition.' (Sources: Tierney, NYT 1991; vos Savant, Game Show Problem)
A rare disease affects 1 in 1,000 people. Your company's annual checkup uses a test that catches nearly every true case but has a 5% false-positive rate. Your colleague tests positive and falls apart. The probability he actually has the disease is closest to —
答案:About 2% — out of 1,000 people, roughly 1 true positive, but about 50 false positives among the 999 healthy, so only about 1 in 51 positives is truly ill
Lay out the natural frequencies and it's obvious: of 1,000 people, roughly 1 is truly ill (and the test almost surely flags him); of the 999 healthy, about 5% — roughly 50 — test falsely positive. So among the ~51 positives, only 1 is truly ill: about 2%. In 1978, Casscells and colleagues put these very numbers to 60 medical students, residents, and attendings at Harvard teaching hospitals: 45% answered 95%, and only 18% got it right; a 2014 JAMA Internal Medicine replication found the correct-answer rate still only about 1 in 4, with 95% still the most common answer. The 1/1000 base rate is the protagonist — a '5% false-positive rate' never means '95% of positives are ill.' (Sources: Casscells, Schoenberger & Graboys 1978, NEJM; Manrai et al. 2014)
In court, the prosecutor argues: 'If the defendant were innocent, the odds of his DNA matching by sheer chance would be minuscule — so given the match, he is all but certainly guilty.' Which trap did this step into?
答案:Confusion of the inverse — taking P(evidence | innocent) directly for P(innocent | evidence), skipping the base rate of how many people in the population might match by chance
This is the courtroom edition of confusion of the inverse — the prosecutor's fallacy: 'if innocent, a match would be extremely unlikely' states P(evidence | innocent), which is not P(innocent | evidence); between them sits the base rate — how many people in the population might match by chance. It shares its root with doctors reading '95% of the diseased test positive' as '95% of positives are diseased.' The gambler's fallacy, the hot hand, and the law of small numbers are all about misreading random sequences, not about flipping the direction of a conditional probability. (Sources: Gigerenzer & Hoffrage 1995; Eddy 1982)
Same breast-cancer Bayes problem, but Gigerenzer & Hoffrage (1995) changed how it was asked, and the correct-answer rate jumped from 16% to 46%. What did they do?
答案:Rewrote the probability format (1% prevalence, 80% sensitivity, 9.6% false-positive rate) as natural frequencies — 'of 1,000 women, 10 have cancer, 8 of whom test positive; of the 990 without cancer, about 95 also test positive'
They rewrote the probability format as natural frequencies, and the correct-answer rate went from 16% to 46% (up to 50% with the shorter frequency format). The magic of frequencies: the base rate (10 out of 1,000) becomes visible, and the Bayesian calculation collapses into one step — 8/(8+95). Gigerenzer & Hoffrage's conclusion: human cognition evolved to process case-by-case natural frequencies, not single-event probabilities — when intuition fails, it's usually not that people are dumb, but that the representation format doesn't match the mental algorithm; put it in frequencies and ordinary people do Bayes just fine. (Source: Gigerenzer & Hoffrage 1995, Psychological Review)
True or False: The roulette wheel has come up red 10 times in a row. Since red and black each take half in the long run, black is now more likely — so piling chips on black is the smart bet.
答案:False
False. Every spin is independent — the wheel has no memory and owes no color anything: after 10 straight reds, the chance of black on the next spin is not one bit higher. On August 18, 1913, black came up 26 times in a row at the Monte Carlo casino: as Huff & Geis recount, from about the 15th black onward gamblers piled heavy bets on red and kept doubling down through their losses, and the casino ended up raking in millions of francs. The prior probability of a given color hitting 26 straight is about 2^-26 ≈ 1 in 67 million, but the part already spun changes nothing about the next spin's conditional probability. That's the gambler's fallacy, which Tversky & Kahneman traced to belief in the 'law of small numbers.' (Sources: Huff & Geis 1959; Tversky & Kahneman 1971)
Timeout. The coach makes the call: 'Number 9 just hit 3 in a row — he's hot, next possession goes to him!' According to Gilovich, Vallone & Tversky's 1985 analysis of the 76ers' actual shooting records, this call —
答案:Runs into the data pointing the other way: the 76ers' regulars hit a weighted average of about 46% after 3 straight makes versus about 56% after 3 straight misses — from which GVT concluded the 'hot hand' was a misreading of random sequences
GVT 1985 found the data pointing the opposite way: 9 Philadelphia 76ers regulars hit a weighted average of just 46% after 3 straight makes versus 56% after 3 straight misses (51% after 1 make, 54% after 1 miss); Celtics free throws and a 100-shot controlled experiment with Cornell varsity players likewise showed no positive correlation — while 91% of fans believed makes breed makes and 84% wanted the ball fed to the hot player. GVT's conclusion: people underestimate how often random sequences naturally streak and mistake normal clustering for a state. One caveat: a 2018 follow-up proposed a correction to GVT's measure (see the next question), so don't canonize 'the hot hand is pure illusion' either. (Source: Gilovich, Vallone & Tversky 1985, Cognitive Psychology)
True or False: In 2018, Miller & Sanjurjo showed that the measure GVT used in 1985 suffers from 'streak selection bias' — flip a fair coin 100 times, and the expected sample proportion of 'heads right after 3 straight heads' is about 46%, not 50%; once corrected, GVT's own controlled Cornell shooting data actually shows a significant hot-hand effect.
答案:True
True. That is streak selection bias: in finite sequences, even with fully independent trials, the expected sample proportion of 'same outcome right after a streak of k' falls below the true probability. GVT had treated 'hit rate after makes ≈ after misses' as evidence of no hot hand — landing exactly in this bias; corrected, their Cornell controlled data shows a significant hot hand (point estimate around +13 percentage points, 95% CI roughly 4%-22%). Handle with care: this later-proposed correction mainly concerns the controlled experimental data, and the true size of the hot hand in live NBA play remains debated. The tastiest lesson: the fans who believed in the hot hand weren't necessarily wrong — the researchers who ran a biased measure for thirty years were. (Sources: Miller & Sanjurjo 2018, Econometrica; Data Colada [88])