Which is more probable? The Linda problem and the discarded base rate
Take this question first — answer honestly, don't skip.
Linda is 31, single, outspoken, and very bright. She majored in philosophy. As a student, she was deeply concerned with issues of discrimination and social justice, and she participated in anti-nuclear demonstrations.
Which is more probable?
- (A) Linda is a bank teller
- (B) Linda is a bank teller and is active in the feminist movement
Made your choice? If you picked B — congratulations, you've been caught, and you're in good company: Tversky & Kahneman (1983) gave the most direct, binary version to 142 UBC undergraduates, and 85% chose B. But B can never be more probable than A: B is a subset of A — every 'teller and feminist' Linda must first be a 'teller.' The probability of a conjunction can never exceed the probability of its components: P(A and B) ≤ P(A). This iron law is the conjunction rule, and violating it is the conjunction fallacy.
What stings more: statistical training barely helps. In the eight-item probability-ranking version, grouped by statistical background: statistically naive subjects violated the rule at 89%, graduate students who had taken statistics courses at 90%, and Stanford decision-science doctoral students at 85%; the direct test showed 88% violation overall.
Why? T&K's explanation is the representativeness heuristic: 'teller and feminist' looks more like Linda — the correlation between probability rankings and representativeness rankings was as high as .98. Your brain quietly swaps 'does it resemble?' for 'is it?' — substituting similarity for probability.
A popular defense: 'Subjects merely read "bank teller" as "teller but not a feminist" — a misreading, not a fallacy.' T&K ruled this out with a control at the time: another group of 119 subjects rated the two options item by item on 9-point probability scales — in a rating task there is no reason for an exclusive reading — and 82% still gave 'teller and feminist' the higher score (means 5.6 vs 3.5). Yet the wording does matter: switch to a frequency format ('Of 100 women like Linda, how many are bank tellers? How many are tellers and active in the feminist movement?') and the fallacy rate plunges from around eighty percent to around twenty. The fallacy is real, but its strength depends heavily on how the question is asked.
The same mechanism has another face: base-rate neglect. Kahneman & Tversky (1973) showed subjects personality sketches 'drawn at random from 100 professionals': one group was told the sample held 70 engineers and 30 lawyers; the other group, the reverse — 30:70. By Bayesian reasoning, the same sketch should yield clearly different 'this person is an engineer' probabilities in the two groups; in fact the two groups' judgments were nearly identical — subjects looked only at whether the description 'resembled an engineer,' and the base rate was ignored. The most glaring contrast is 'Dick,' a deliberately zero-information description (30 years old, married with no children, high ability and motivation, well liked by colleagues): both groups' median probability was 0.50. Handed a worthless description, people throw the base rate away; with no description at all, subjects used the 70%/30% correctly.

You think you're testing — you're confirming: 2-4-6 and the four cards
Round two. The experimenter has a rule in mind, and the sequence 2, 4, 6 conforms to it. You may propose any triples you like; for each one you're told whether it conforms; announce the rule once you feel certain. Think: which triple would you try first?
Most people already have a hypothesis — 'add 2 each time' — and start proposing 8, 10, 12, 20, 22, 24... All 'conform,' confidence swells, and they announce: add 2 each time. Wrong. The true rule is almost insultingly simple: 'three increasing numbers.' Of Wason's (1960) 29 psychology undergraduates, only 6 (21%) were correct on their first announcement; 13 were wrong once first, 9 were wrong twice or more, and 1 never managed to announce any correct rule.
The classic failure mode: after forming a hypothesis narrower than the true rule, like 'add 2 each time,' subjects generate only positive instances of their own hypothesis to 'confirm' it, and almost never propose triples that violate it (such as 1, 2, 3) to falsify it. Wason called this the 'failure to eliminate hypotheses' — the first experimental demonstration of confirmation bias.
But don't teach it wrong: subjects had no stake whatsoever in a number rule — nothing they 'wanted to believe' — yet they still systematically tested only positive cases. The heart of confirmation bias is the positive test strategy, a cognitive default, not motivated self-deception. Klayman & Ha (1987) went further: positive testing is actually efficient, even near-optimal, in most real environments; it turns fatal only when the true rule is broader than your hypothesis — and 2-4-6 is precisely such a deliberately constructed trap: the true rule 'increasing' fully contains the hypothesis 'add 2,' so positive tests always come back 'yes.' The antidote: deliberately design the negative test — 'if I'm wrong, what would I see?' Propose 1, 2, 3; if it also 'conforms,' your hypothesis goes bankrupt on the spot — that one step is worth more than a hundred confirmations.
Round three: four cards on the table — E, K, 4, 7 — each with a letter on one side and a number on the other. The rule: 'If a card has a vowel on one side, then it has an even number on the other side.' Which cards must you turn over to test whether the rule is true or false? Choose first, then read on.
The logically correct answer is E and 7: only a 'vowel + odd number' combination can falsify the rule — if E hides an odd number, or 7 hides a vowel, the rule breaks; and whatever is behind 4, the rule is not violated (it never says an even number must have a vowel on the back). Actual performance (Wason 1968 and replications): nearly everyone picks E, 60–75% wrongly pick 4 (affirming the consequent), and only a few pick 7; the standard abstract version is usually solved by fewer than 10% — Dawes (1975) reported that only 1 of 5 mathematical psychology PhDs got it right. Yet recast the same structure in familiar content and the result flips at once: in Griggs & Cox's (1982) drinking-age version — rule 'if a person is drinking beer, they must be over 19,' cards = drinking beer / drinking cola / 16 years old / 22 years old — 74% answered correctly. Identical logical structure; content and experience decide the outcome. That's the 'content effect': what you lack is often not logic, but a translation of the abstract rule into 'who should be checked for cheating.'

Easy to recall ≠ frequent: the letter K and causes of death overestimated 350-fold
Last round — an easy one (or is it?): if you draw a random word from an English text, is the letter K more likely to appear as the first letter, or in the third-letter position?
Tversky & Kahneman (1973) asked 152 subjects (about K, L, N, R, and V in turn, with ratio estimates). All five consonants are actually more common in the third position — about twice as common; yet 105 subjects (69%) judged the first position more common for a majority of the letters, with a median estimated ratio of about 2:1 toward the first position — the direction entirely reversed. The mechanism is naked: retrieving words by their first letter (words starting with K) is far easier than retrieving them by their third letter, and 'ease of retrieval' was taken for 'objective frequency' — the defining demonstration of the availability heuristic.
Availability doesn't only trap you in the lab. Lichtenstein, Slovic, Fischhoff, Layman & Combs (1978) ran 5 experiments with about 660 adults, comparing 106 pairs of causes of death — 'which kills more people per year in the US?' — with ratio estimates. Two systematic biases surfaced: first, rare but dramatic causes are overestimated while common causes are underestimated — vivid rarities like botulism and tornadoes were greatly overestimated, while common killers like stroke and diabetes were greatly underestimated (tornadoes were judged a more frequent killer than asthma, though asthma actually kills far more); second, the judged ratios were wildly off — motor-vehicle accidents actually kill only about 1.5 times as many people as diabetes, yet subjects judged the ratio at 350 times on average. Causes that get heavy media coverage and are easy to recall and imagine have their 'mental availability' impersonate real frequency. The 'danger ranking of the world' in your head is largely edited by news editors.

自测 · 学完检查一下
想真正动手做题、记进度、攒连胜?到互动课里练。
Linda is 31, single, outspoken, and very bright. She majored in philosophy. As a student she was deeply concerned with discrimination and social justice, and participated in anti-nuclear demonstrations. Which is more probable?
答案:Linda is a bank teller
The conjunction rule: P(teller and feminist) ≤ P(teller) — 'teller and feminist' is a subset of 'teller'; every Linda who is a 'teller and feminist' must first be a 'teller,' so no matter how fitting the description feels, the conjunction can never be the more probable option. In the most direct binary version, 85% of 142 UBC undergraduates chose 'teller and feminist' — the representativeness heuristic at work: the conjunction resembles Linda more (probability rankings correlated .98 with representativeness rankings), and similarity got substituted for probability. (Source: Tversky & Kahneman 1983 — Extensional versus intuitive reasoning, Psychological Review 90(4))
Old Zhou is 55, has smoked for thirty years, and coughs all year round. Which is more probable?
答案:Old Zhou will be hospitalized next year
This is a structural twin of the Linda problem: 'hospitalized for a lung disease' is a subset of 'hospitalized' — whatever puts Old Zhou in the hospital counts as 'hospitalized,' so P(hospitalized for a lung disease) ≤ P(hospitalized), and the iron law doesn't care about the description. 'Thirty years of smoking, coughing all year' makes 'lung disease' feel highly representative, luring you into swapping 'does it resemble?' for 'is it?' How to catch it: whenever wording adds detail to an event — 'and,' 'because,' 'specifically' — recite the conjunction rule first, then ask: 'for the broad category alone, what's the probability?' (Source: Tversky & Kahneman 1983 — the representativeness mechanism of the conjunction fallacy, Psychological Review 90(4))
True or False: In the eight-item ranking version of the Linda problem, graduate students who had taken statistics courses violated the conjunction rule at 90% and Stanford decision-science doctoral students at 85% — barely different from statistically naive subjects at 89%. Statistical training does almost nothing to remove the conjunction fallacy.
答案:True
True. Grouped by statistical background, the results were: statistically naive subjects 89%, graduate students with statistics coursework 90%, Stanford decision-science doctoral students 85% — with 88% violation overall in the direct test. The differences from training are negligible, showing that the conjunction fallacy is not a knowledge gap of 'never learned probability' but the representativeness heuristic running on autopilot. What does slash the fallacy rate is changing the wording: in a frequency format ('of 100 women like Linda, how many...'), the rate plunges from around eighty percent to around twenty. (Sources: Tversky & Kahneman 1983; Hertwig & Gigerenzer 1999; Fiedler 1988)
A defense of the Linda problem goes: 'Subjects merely understood "bank teller" as "teller but not a feminist" — a misreading, not a fallacy.' Which piece of evidence most directly rebuts this defense?
答案:T&K had another group of 119 subjects rate the two options item by item on 9-point probability scales — with no reason for an exclusive reading there, 82% still scored 'teller and feminist' higher (means 5.6 vs 3.5)
The misreading defense claims subjects read T exclusively as 'T and not F,' so choosing T&F wouldn't violate the conjunction rule. The direct rebuttal is the item-by-item rating control: 119 subjects scored T and T&F separately on 9-point scales — a task with no reason for an exclusive reading — and 82% still gave T&F the higher score (means 5.6 vs 3.5), so misreading cannot save the transparent version. Yet wording genuinely matters: a frequency format cuts the fallacy from around eighty percent to around twenty — so state the full conclusion: the fallacy is real, and its strength depends heavily on single-event vs frequency phrasing. The other options: the doctoral-student data show training doesn't help, and the .98 correlation shows the mechanism — neither directly addresses 'was it a misreading.' (Sources: Tversky & Kahneman 1983; Fiedler 1988; Hertwig & Gigerenzer 1999)
The 2-4-6 task: the experimenter has a rule in mind, and '2, 4, 6' conforms to it. Your hypothesis is 'add 2 each time.' Which triple is the most informative next test?
答案:1, 2, 3
The first three are all positive instances of 'add 2 each time' — if the true rule is broader than your hypothesis (say, 'three increasing numbers'), positive instances always come back 'conforms,' and a hundred of them will never reveal you're wrong. '1, 2, 3' violates 'add 2' but is still increasing: if the experimenter answers 'conforms,' your hypothesis is instantly proven too narrow — one negative test carries more information than a hundred confirmations. Wason's (1960) subjects failed exactly by testing only positive instances: of 29, just 6 (21%) announced the true rule 'three increasing numbers' correctly on the first try. (Sources: Wason 1960 — On the failure to eliminate hypotheses, QJEP 12(3); Klayman & Ha 1987)
True or False: In the 2-4-6 experiment, subjects tested only positive instances because they were emotionally invested in their hypothesis and wanted to prove what they wished to believe — confirmation bias is essentially motivated self-deception.
答案:False
False. Subjects had no stake whatsoever in a number rule — nothing they 'wanted to believe' — yet they still systematically tested only positive instances. The heart of confirmation bias is the positive test strategy, a cognitive default, not motivated favoritism or self-deception. Klayman & Ha (1987) went further: positive testing is efficient, even near-optimal, in most real environments; it turns fatal only when the true rule is broader than your hypothesis — and 2-4-6 is precisely such a deliberately constructed trap (the true rule 'increasing' fully contains the hypothesis 'add 2,' so positive tests always come back 'yes'). The antidote is not 'have no stance,' but deliberately designing the negative test: 'if I'm wrong, what would I see?' (Sources: Klayman & Ha 1987 — Psychological Review 94(2); Wason 1960)
Four cards on the table: E, K, 4, 7 — each with a letter on one side and a number on the other. The rule: 'If a card has a vowel on one side, then it has an even number on the other side.' Which cards must you turn over to test whether the rule is true or false?
答案:E and 7
Only a 'vowel + odd number' combination can falsify the rule, so you must turn E (which may hide an odd number) and 7 (which may hide a vowel); whatever is behind 4, the rule is not violated — it never says an even number must have a vowel on the back, so turning 4 or K is an uninformative wasted move, and 'all four' is not an answer to 'which must you turn over.' Actual performance: nearly everyone picks E, 60–75% wrongly pick 4 (affirming the consequent), and only a few pick 7; the standard abstract version is usually solved by fewer than 10% — Dawes (1975) reported only 1 of 5 mathematical psychology PhDs got it right. To test 'if P then Q,' always go after P and not-Q. (Sources: Wason 1968 — Reasoning about a rule, QJEP 20(3); Griggs & Cox 1983)
True or False: Recast the four-card task in a logically identical drinking-age setting — rule 'if a person is drinking beer, they must be over 19,' cards = drinking beer / drinking cola / 16 years old / 22 years old — and accuracy leaps from under 10% on the abstract version to 74%: failing the abstract version does not mean failing to understand 'if... then...' logic.
答案:True
True. Griggs & Cox's (1982) drinking-age version is structurally identical to the abstract one: what must be checked is still P and not-Q — 'drinking beer' (may be under 19 on the back) and '16 years old' (may be drinking beer on the back); 'drinking cola' and '22 years old' cannot violate the rule whatever is on their backs. The same kind of subjects usually solve fewer than 10% of the abstract letter-number version, but 74% solved this one — content and experience decide the outcome. That's the 'content effect': what people lack is often not logical ability itself, but a translation of the abstract rule into an intuition of 'who should be checked for cheating.' (Sources: Wason 1968; Griggs & Cox 1982, British Journal of Psychology 73; Griggs & Cox 1983)
Draw a random word from an English text: is the letter K more likely to appear as the first letter, or in the third-letter position?
答案:Third position — actually about twice as likely as the first
K is actually more common in the third position — about twice as common. Tversky & Kahneman (1973) asked 152 subjects about the five consonants K, L, N, R, V — all five are actually more common in the third position, yet 105 subjects (69%) judged the first position more common for a majority of the letters, with a median estimated ratio of about 2:1 toward the first position — the direction entirely reversed. Mechanism: retrieving words by first letter is far easier than by third letter, and 'ease of retrieval' was taken for 'objective frequency' — the availability heuristic. The real-world version of the same mechanism: heavily reported, vivid causes of death get hugely overestimated — car crashes actually kill only about 1.5 times as many people as diabetes, yet subjects judged the ratio at 350 times on average. Don't overstate it, though: T&K deliberately picked exceptional letters that are more common in third position (most consonants are more common word-initially), and Sedlmeier et al. (1998) found subjects judged the 'first-position-common' letters rather accurately — the sound conclusion is 'ease of retrieval systematically biases frequency judgments,' not 'people know nothing about letter frequencies.' (Sources: Tversky & Kahneman 1973 — Availability, Cognitive Psychology 5(2); Lichtenstein et al. 1978; Sedlmeier et al. 1998)
Kahneman & Tversky (1973) showed two groups the same personality sketch, 'drawn at random from 100 professionals': one group was told the sample held 70 engineers and 30 lawyers, the other that it was 30 engineers and 70 lawyers. The sketch's subject, Dick: 30 years old, married with no children, high ability and motivation, well liked by colleagues — written to carry zero information. How did the two groups judge the probability that 'Dick is an engineer'?
答案:Both groups' medians were 0.50 — one worthless description made both groups throw the 70%/30% base rate away entirely
Both groups' medians were 0.50. By Bayesian reasoning, the same description should yield clearly different probabilities in the 70:30 and 30:70 groups; in fact subjects looked only at whether the description 'resembled an engineer' (representativeness), and the base rate was ignored — most glaringly, a zero-information filler description made people discard the base rate, while with no description at all, subjects used the 70%/30% correctly. Representativeness = substituting similarity for probability: the Linda problem's conjunction fallacy and this base-rate neglect are two faces of the same mechanism. How to catch it: for any 'profile + judge the probability' question, ask for the base rate first, then look at the description. (Source: Kahneman & Tversky 1973 — On the psychology of prediction, Psychological Review 80(4); the median-0.50 data also reported in Gigerenzer 1991)