❤️ 15/15

Which is more probable? The Linda problem and the discarded base rate

Take this question first — answer honestly, don't skip.

Linda is 31, single, outspoken, and very bright. She majored in philosophy. As a student, she was deeply concerned with issues of discrimination and social justice, and she participated in anti-nuclear demonstrations.

Which is more probable?

- (A) Linda is a bank teller
- (B) Linda is a bank teller and is active in the feminist movement

Made your choice? If you picked B — congratulations, you've been caught, and you're in good company: Tversky & Kahneman (1983) gave the most direct, binary version to 142 UBC undergraduates, and 85% chose B. But B can never be more probable than A: B is a subset of A — every 'teller and feminist' Linda must first be a 'teller.' The probability of a conjunction can never exceed the probability of its components: P(A and B) ≤ P(A). This iron law is the conjunction rule, and violating it is the conjunction fallacy.

What stings more: statistical training barely helps. In the eight-item probability-ranking version, grouped by statistical background: statistically naive subjects violated the rule at 89%, graduate students who had taken statistics courses at 90%, and Stanford decision-science doctoral students at 85%; the direct test showed 88% violation overall.

Why? T&K's explanation is the representativeness heuristic: 'teller and feminist' looks more like Linda — the correlation between probability rankings and representativeness rankings was as high as .98. Your brain quietly swaps 'does it resemble?' for 'is it?' — substituting similarity for probability.

A popular defense: 'Subjects merely read "bank teller" as "teller but not a feminist" — a misreading, not a fallacy.' T&K ruled this out with a control at the time: another group of 119 subjects rated the two options item by item on 9-point probability scales — in a rating task there is no reason for an exclusive reading — and 82% still gave 'teller and feminist' the higher score (means 5.6 vs 3.5). Yet the wording does matter: switch to a frequency format ('Of 100 women like Linda, how many are bank tellers? How many are tellers and active in the feminist movement?') and the fallacy rate plunges from around eighty percent to around twenty. The fallacy is real, but its strength depends heavily on how the question is asked.

The same mechanism has another face: base-rate neglect. Kahneman & Tversky (1973) showed subjects personality sketches 'drawn at random from 100 professionals': one group was told the sample held 70 engineers and 30 lawyers; the other group, the reverse — 30:70. By Bayesian reasoning, the same sketch should yield clearly different 'this person is an engineer' probabilities in the two groups; in fact the two groups' judgments were nearly identical — subjects looked only at whether the description 'resembled an engineer,' and the base rate was ignored. The most glaring contrast is 'Dick,' a deliberately zero-information description (30 years old, married with no children, high ability and motivation, well liked by colleagues): both groups' median probability was 0.50. Handed a worthless description, people throw the base rate away; with no description at all, subjects used the 70%/30% correctly.

⚠️One line to remember: 'resembles' is not 'is.' The more detailed a story, the more vivid and 'fitting' it feels — but every added 'and' can only lower the probability. Next time you hear 'he's X, and surely Y too,' ask first: with X alone, what's the probability? And what's the base rate?
The conjunction fallacy: 'teller and feminist' is a subset of 'teller' — a subset can never be more probable than the whole

Test yourself