Should you armor where the bullet holes are? — Survivorship bias and Wald's inversion
World War II. Bombers return to base riddled with bullet holes: wings and fuselage densely peppered, the area around the engines relatively clean. Armor is limited — add too much and the plane can't fly. So where should it go?
The intuitive answer: where the holes are — that's where the planes get hit most. According to Wallis (1980), the military indeed leaned toward reinforcing the most-hit parts of the returning planes.
Abraham Wald of Columbia University's Statistical Research Group (SRG, founded in the summer of 1942) pointed out the fatal flaw in that reasoning: the only damage data available came from planes that returned (the survivors); the planes that were shot down were entirely unobserved — using the survivors' bullet-hole distribution directly to assess aircraft vulnerability produces a systematic bias. Under the assumption that hits were roughly uniformly distributed, he inverted the inference: planes hit in vital parts were less likely to make it back, so the survivor sample shows fewer holes exactly there — the area around the engines looks 'clean' not because it wasn't hit, but because the planes hit there never came home.
Wald's actual method is far more hardcore than the legend: he wrote 8 memoranda, 'A Method of Estimating Plane Vulnerability Based on Damage of Survivors,' introducing the conditional probability p_i = 'the probability of being downed by the i-th hit, given survival of the first i−1 hits,' and building equations under explicitly stated assumptions such as 'planes are downed only by enemy fire (L0=0 — an unhit plane doesn't fall),' to infer the hit distribution of the planes that never returned from the data of those that did — one memo specifically estimates the vulnerability of each section of the aircraft. The method remained in use through the Korean and Vietnam wars; the memoranda were only reprinted publicly in 1980 by the Center for Naval Analyses (CNA, report CRC 432).

Every department treats women fairly — so why does the total flip? — Simpson's paradox
Fall 1973, UC Berkeley graduate admissions, aggregate data: 8,442 male applicants, about 44% admitted; 4,321 female applicants, about 35% admitted — a 9-point gap that looks like open-and-shut sex discrimination.
Then Bickel, Hammel, and O'Connell (Science, 1975) analyzed the data department by department, and the plot reversed: in most departments, women's admission rates were not lower. In the commonly taught 'six largest departments' subset — 2,691 men applied with 1,198 admitted, 1,835 women applied with 557 admitted — women had the higher admission rate in 4 of the 6 departments; properly pooled by department, there was even a small but statistically significant bias in favor of women.
The aggregate and the groups point in opposite directions, and the culprit is the confounding variable 'which department you applied to': women applied disproportionately to competitive departments with low overall admission rates, while men applied more to departments with high admission rates. And one more myth to bust while we're here: 'Berkeley was sued over this' — it never happened. There was no lawsuit; the associate dean of the graduate division, worried the university might be sued, asked statisticians to examine the data.
This is Simpson's paradox: an association or trend that holds within every subgroup (or exists in none of them) can reverse direction or vanish once the subgroups are pooled. The phenomenon was first noticed by Pearson (1899) and Yule (1903); Simpson's 1951 paper used a 2×2×2 contingency table to show that even without second-order interaction, you cannot mechanically collapse it into a 2×2 table for testing; the name 'Simpson's paradox' was coined by Blyth in 1972.

Does scolding work better than praise? — Regression to the mean, small samples, and charts that lie
Kahneman was teaching Israeli air force flight instructors that reward works better than punishment for skill learning, when a senior instructor objected: after I praise a cadet for a beautiful maneuver, the next attempt is usually worse; after I scream at a cadet for a terrible one, the next attempt is usually better — so punishment works and praise backfires.
Sounds airtight? Here's the crack: the instructor praises only after extremely good performances and screams only after extremely bad ones — and after an extreme performance, no matter what the instructor does, the next attempt will most likely regress toward the average. The instructor misread the inevitable fall-back/bounce-back of random fluctuation as the causal effect of his own rewards and punishments — this is the regression fallacy. Kahneman called this insight one of the most satisfying eureka moments of his career.
The statistical origin of 'regression': Galton (1886) measured 205 sets of parents (using 'mid-parent' height) and their 928 adult children, and found that the more the parents deviated from the population mean, the children on average deviated only about 2/3 as much (in the original, 'as 2 to 3') — he called it 'regression towards mediocrity,' and the statistical term 'regression' comes from this very paper. Two key points: it is a purely statistical phenomenon, not biological 'degeneration' — it appears between any two imperfectly correlated variables (|r|<1) — and it is symmetric in direction (the parents of exceptionally tall children are, on average, not that tall either). Whenever an intervention happens to occur right after an extreme value ('took the medicine when sickest, then improved'; 'fixed the accident blackspot, then accidents fell'), rule out regression to the mean before crediting the intervention.
Two more quick traps.
Sample-size neglect: a town has a large hospital delivering about 45 babies a day and a small one delivering about 15, with boys at about 50%. Over a year, which hospital records more days on which more than 60% of the babies are boys? Of Tversky & Kahneman's (1974) 95 subjects, 53 (56%) answered 'about the same,' and only 21 (about 22%) got it right: the small hospital. The smaller the sample, the larger the sampling fluctuation of a proportion — and the easier it is to cross the 60% threshold. Trusting 'representativeness' and assuming small and large samples reflect the population equally well is what Tversky & Kahneman called belief in the 'law of small numbers.' The everyday version: drawing conclusions from three to five users or a few days of data is exactly the same mistake.
The truncated Y axis (the Gee-Whiz Graph): Huff (1954) demonstrated that if a line chart's Y axis doesn't start at zero and the vertical scale is stretched, a 1% change can be drawn to look like a surge — the numbers aren't faked, but the impression entirely is. Pandey et al.'s crowdsourced experiment at CHI 2015 confirmed that this genuinely deceives people: subjects viewing truncated-axis bar charts significantly exaggerated the perceived size of differences, and 'message-reversal' distortions such as an inverted Y axis fooled people at similarly high rates. Tufte supplied the yardstick: Lie Factor = (size of effect shown in the graphic) / (size of effect in the data) — ideally about 1.

自测 · 学完检查一下
想真正动手做题、记进度、攒连胜?到互动课里练。
WWII bombers return with wings and fuselage riddled with bullet holes, but the area around the engines relatively clean. Under Wald's analysis, where should limited armor be considered first?
答案:The engines and other sections with few holes in the returning sample — planes hit in vital parts mostly never made it back, so the survivor data shows fewer holes exactly there
The only damage data available came from planes that returned (the survivors); the downed planes were entirely unobserved — using the survivors' bullet-hole distribution directly to assess vulnerability produces a systematic bias, which is survivorship bias. Per Wallis (1980): the military leaned toward reinforcing the most-hit parts of returning planes, while Wald, assuming roughly uniform hits, inferred that planes hit in vital parts were less likely to return — so the survivor sample shows fewer holes there. Note the last option is wrong too: the data was not unusable — Wald built precisely the method for inferring the non-returners' hit distribution from survivor data. (Sources: Mangel & Samaniego, JASA 1984; Wallis, JASA 1980 rejoinder)
In the popular story, Wald fires back at the military on the spot: 'The armor goes on the engines, where there are no bullet holes!' After the historical check, what did Wald's 8 memoranda of 1943 actually contain?
答案:Pure technical derivation: introducing the conditional probability p_i (downed by the i-th hit, given survival of the first i−1 hits) and, under explicitly stated assumptions such as 'planes are downed only by enemy fire (L0=0),' inferring the hit distribution of the planes that never returned from survivor data — with no recommendation whatsoever about where to place armor
The historical check (Casselman, AMS 2016): Wald's 8 memoranda are pure technical derivation, containing no recommendation about armor placement anywhere — SRG convention was to answer statistical questions only, never to make decisions for the military. The famous quote has no primary source; the vivid dialogue is a later 'plausible reconstruction'; the only primary sources on this work are the memoranda themselves and two brief, unnamed mentions in Wallis's memoir. The true core of the story appears in Wallis's 1980 JASA rejoinder: the military leaned toward protecting the most-hit sections, and Wald, under the uniform-hits assumption, inferred that planes hit in vital parts rarely returned. When citing this story, keep the 'documented method' and the 'legendary details' clearly separated. (Sources: Casselman, AMS Feature Column 2016; Wallis, JASA 1980)
True or False: The widely shared diagram of an airplane covered in red bullet-hole dots is not a historical artifact from Wald's time — it is a modern illustration drawn by designer Cameron Moll around 2005 for a talk (the 2016 Wikipedia redraw is the most widely circulated version).
答案:True
True. By the artist's own account, Cameron Moll created the red-dot airplane diagram around 2005 as a modern illustration for a talk — it is not Wald's; the version circulating most widely today is Wikipedia's 2016 redraw. And here lies this lesson's best self-referential twist: the go-to image for explaining survivorship bias is itself a live specimen of survivorship bias — what 'survives' in circulation is the most dramatic version (the quote plus the red-dot diagram), while the dry primary memoranda go unseen. Before you share, ask: where is the primary source for this 'historical artifact'? (Sources: Moll, 'Abraham Wald and the airplane diagram — origin story'; Casselman, AMS 2016)
Fall 1973, Berkeley graduate admissions: 8,442 male applicants with about 44% admitted versus 4,321 female applicants with about 35% — apparently discrimination against women; yet Bickel et al.'s department-level analysis found most departments' admission rates for women were not lower, and in 4 of the six largest departments women's rates were actually higher. Why do the total and the departments point in opposite directions?
答案:The confounding variable 'which department you applied to' is at work: women applied disproportionately to competitive departments with low overall admission rates, while men applied more to high-admission-rate departments — properly pooled by department, there was even a small but statistically significant bias in favor of women
This is the most famous real-data case of Simpson's paradox: the aggregate and the groups disagree because of the confounding variable 'which department you applied to' — women applied disproportionately to competitive, low-admission-rate departments. In the six-department subset, 2,691 men applied with 1,198 admitted and 1,835 women applied with 557 admitted, with women's rate higher in 4 departments; properly pooled by department, there was even a small but statistically significant bias in favor of women. The small-sample explanation fails (with over ten thousand applicants, a 9-point gap is far from random noise). And another common myth: 'Berkeley was sued over this' — there was no lawsuit; the associate dean of the graduate division, worried about a possible suit, asked statisticians to examine the data. (Source: Bickel, Hammel & O'Connell, Science 1975)
True or False: When facing Simpson's paradox, whether to trust the pooled data or the grouped data is a purely mathematical question — calculate carefully enough, and the math alone will give you the answer.
答案:False
False. Whether to trust the pooled or the grouped data is not something mathematics can answer on its own; it depends on the causal role of the grouping variable: a confounder should be stratified on or adjusted for (like the departments in the Berkeley case), while a mediator of the treatment's effect should not — Pearl formalized this judgment with causal graphs. Background for completeness: the phenomenon was first noticed by Pearson (1899) and Yule (1903); Simpson (1951) used a 2×2×2 contingency table to show that even without second-order interaction, you cannot mechanically collapse it into a 2×2 table for testing; the name 'Simpson's paradox' was coined by Blyth in 1972. (Sources: Simpson, JRSS B 1951; Blyth, JASA 1972; Pearl, UCLA R-414)
Israeli air force flight instructors observed: after a cadet is praised for a beautiful maneuver, the next attempt is usually worse; after being screamed at for a terrible one, the next attempt is usually better — so they concluded 'punishment works, praise backfires.' What did Kahneman say is wrong with this reasoning?
答案:Regression to the mean: praise comes only after extremely good performances and scolding only after extremely bad ones — and after an extreme performance, whatever the instructor does, the next attempt will most likely regress toward the average; the instructors misread the inevitable fall-back/bounce-back of random fluctuation as the causal effect of their rewards and punishments
This is the textbook case of the regression fallacy: instructors praise only at extreme highs and scold only at extreme lows, and after an extreme the next attempt was going to regress toward the mean anyway — the 'getting worse' and 'getting better' have nothing to do with the rewards or punishments. Kahneman called this insight one of the most satisfying eureka moments of his career in Thinking, Fast and Slow; the flight-training case is recorded in the 'misconceptions of regression' section of Tversky & Kahneman's 1974 Science paper. Whenever an intervention happens right after an extreme value (taking medicine when sickest and then improving; fixing an accident blackspot and then seeing fewer accidents), rule out regression to the mean before crediting the intervention. (Sources: Tversky & Kahneman, Science 1974; Kahneman, Thinking, Fast and Slow ch.17)
The origin of the statistical term 'regression': in 1886 Galton measured 205 sets of parents (using 'mid-parent' height) and their 928 adult children. What did he find?
答案:The more the parents deviate from the population mean, the children on average deviate only about 2/3 as much: children of tall parents are on average shorter than their parents, and children of short parents taller — he called it 'regression towards mediocrity,' and the term 'regression' was born
Galton found that the more parents deviated from the population mean, the children on average deviated only about 2/3 as much (in the original, 'as 2 to 3'), and called it 'regression towards mediocrity' — the statistical term 'regression' comes from this very paper. Two key points: (1) it is a purely statistical phenomenon, not biological 'degeneration' — it appears between any two imperfectly correlated variables (|r|<1); (2) it is symmetric in direction — the parents of exceptionally tall children are, on average, not that tall either. Whenever an extreme performance is followed by a fall-back (a slump after a career high, a champion turning ordinary the next year), check for regression to the mean first instead of rushing to a causal explanation. (Source: Galton, Journal of the Anthropological Institute 1886)
A town's large hospital delivers about 45 babies a day and its small hospital about 15, with boys at about 50%. Over a year, which hospital records more days on which more than 60% of the babies born are boys?
答案:The small hospital — the smaller the sample, the larger the sampling fluctuation of a proportion, and the easier it is to cross the 60% threshold
The correct answer is the small hospital: the smaller the sample, the larger the sampling fluctuation of a proportion, and the easier it is to cross the 60% threshold. In Tversky & Kahneman's 1974 Science paper, 53 of 95 subjects (56%) answered 'about the same,' and only 21 (about 22%) got it right — people trust 'representativeness' and assume small and large samples reflect the population equally well, an insensitivity to sample size that Tversky & Kahneman called belief in the 'law of small numbers' (mistakenly expecting small samples to obey the law of large numbers). The everyday version of the trap: drawing conclusions from three to five users or a few days of data is exactly the same mistake. (Sources: Tversky & Kahneman, Science 1974; Kahneman & Tversky, Cognitive Psychology 1972)
True or False: As long as the numbers in a chart are not themselves fabricated, the chart cannot be misleading — truncating the Y axis is merely a layout choice that doesn't affect readers' judgment.
答案:False
False. In 'The Gee-Whiz Graph' chapter of How to Lie with Statistics, Huff demonstrated that starting a line chart's Y axis above zero and stretching the vertical scale can make a 1% change look like a surge — the numbers aren't faked, but the impression entirely is. And this is no hypothetical worry: Pandey et al.'s crowdsourced experiment at CHI 2015 confirmed that subjects viewing truncated-axis bar charts significantly exaggerated the perceived size of differences, and 'message-reversal' distortions such as an inverted Y axis fooled people at similarly high rates. So the first step in reading any chart is always to check the axis start and scale. (Sources: Huff 1954; Pandey et al., Proc. ACM CHI 2015)
Which of the following is a correct practical rule against misleading charts?
答案:Bar charts encode value by length, so their baseline must be zero; line charts may use a non-zero baseline, but it should be made explicit; read the axis start and scale first — Tufte's Lie Factor (size of effect shown in the graphic / size of effect in the data) should ideally be about 1
The practical rules: bar charts encode value by length, so their baseline must be zero; line charts may use a non-zero baseline, but it should be made explicit; when reading a chart, check the axis start and scale first. Tufte defined the Lie Factor = (size of effect shown in the graphic) / (size of effect in the data), which should ideally be about 1 — a value far above 1 means the graphic is exaggerating the data. Note the first option overshoots: 'line charts must start at zero' is not the rule — line charts are allowed a non-zero baseline (made explicit); the zero-baseline requirement applies to bar charts, which encode value by length. (Sources: Tufte, The Visual Display of Quantitative Information 1983; Huff 1954; Pandey et al., CHI 2015)