🪤 Thinking Traps · Classic Traps

Cognitive Reflection Test: 3 Easy Questions Smart People Miss

From bat-and-ball to doubling lily pads — three CRT 'grade-school' questions fooled half of MIT; meet System 1's blurting, System 2's laziness, and the cognitive miser, and learn to hit the slow-thinking switch when it counts

一句话先懂 · TL;DR

Take the bat-and-ball test, meet System 1 vs System 2 and the cognitive miser, and learn when to trust intuition — real CRT data from Frederick (2005).

Don't peek at the answer: one bat, one ball, and half of MIT goes down

Let's open the course with a quiz. Relax — it's grade-school arithmetic:

A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?

Don't scroll down yet — say your answer in your head within 3 seconds.

...

Got it? If what popped up was '10 cents' — congratulations, you're standing with more than half the students at Harvard, MIT, and Princeton: all wrong. One quick check exposes it: if the ball were 10 cents, the bat ($1.00 more) would be $1.10, totaling $1.20 — too much. The correct answer is 5 cents (ball $0.05 + bat $1.05 = $1.10). This is the first item of Frederick's (2005) Cognitive Reflection Test (CRT), and the single most famous question in all of 'intuition trap' research.

Why is almost everyone's first reaction 10 cents? Because two processes run in your head:

- System 1: automatic, fast, nearly effortless, impossible to switch off — like recognizing a friend's face at a glance;
- System 2: slow, serial, effortful, demanding motivation and attention — like computing √19163 without a calculator.

The labels were coined by Stanovich and West in 2000, then borrowed and popularized by Kahneman's Thinking, Fast and Slow (2011) as the framework of the whole book. The intuition trap lives exactly at their handoff: System 1 blurts out a fluent-sounding candidate ('$1.10 splits into $1 and 10 cents — so smooth!'), and System 2 often signs off without so much as a glance. So the CRT was never testing arithmetic — it tests whether you can hold back the answer that leaps to your tongue.

⚠️Getting fooled is genuinely nothing to be ashamed of. Frederick (2005) administered the three CRT items to 3,428 respondents over 26 months across 35 separate studies: the overall average was just 1.24 correct out of 3; 33% got all three wrong, and only 17% got all three right. By campus it stings even more: MIT topped the table at 2.18, Princeton 1.63, Harvard 1.43, the University of Toledo 0.57 — even at MIT, only 48% got all three right. Three pieces of 'grade-school arithmetic' floored half of the world's top engineering students.
System 1 blurts the answer while System 2 dozes off and rubber-stamps it — most errors start here

The machines question: why you feel absolutely nothing at the moment you get fooled

One more — same 3-second rule, answer in your head:

If it takes 5 machines 5 minutes to make 5 widgets, how long would it take 100 machines to make 100 widgets?

...

Hands up if you said '100 minutes.' The correct answer is 5 minutes: each machine takes 5 minutes to make 1 widget, so 100 machines running at once turn out 100 widgets in 5 minutes. That's CRT item two. Frederick also spotted a lovely detail on the answer sheets: even people who answered correctly often thought of the wrong answer first — 10 cents crossed out and replaced with 5 cents shows up all the time, while the reverse was never seen. The intuitive answer pops up for everyone; the only difference is whether you let it out the door.

Behind this is your brain's factory default: the cognitive miser — proposed by Fiske and Taylor in Social Cognition (1984): when processing information, people reach for the least-effort shortcuts and ready-made schemas, rather than gathering the full evidence and reasoning it through as the earlier 'naive scientist' model assumed. Kahneman calls the same phenomenon the law of least effort: System 2 is lazy by nature, and the fluent answers System 1 hands up are often accepted without review. Stanovich goes further, treating this 'miserly processing' as the core mechanism behind smart people making dumb decisions (dysrationalia).

Note: miserliness is not a defect — it's an energy-saving design. If your brain deliberated over everything, picking breakfast would require a board meeting. Most of the time the shortcuts work fine; trap questions just land precisely where they fail. Evidence? Frederick rewrote the problem as a structurally identical version that triggers no intuition (a banana and a bagel cost 37 cents; the banana costs 13 cents more) — and the error rate dropped sharply. The math didn't change; only the seductive fake answer disappeared.

⚠️The scariest data point of all. Frederick asked respondents to estimate 'what share of other people would answer the bat-and-ball problem correctly' — those who gave the wrong answer, '10 cents,' estimated on average that 92% of people would get it right; those who correctly answered '5 cents' estimated only 62%. In other words: the people who got it wrong found the question easier than the people who got it right. At the moment you get fooled there is no subjective alarm whatsoever — no 'hmm, something feels off,' only the serene glow of 'well, that was easy.' That is exactly what makes intuition traps dangerous: being wrong comes bundled with a feeling of fluency.
The brain is a miser: the smooth shortcut often leads straight into the trap

The lily pads question: so can intuition ever be trusted?

Last one — you know the drill, 3 seconds:

In a lake, a patch of lily pads doubles in size every day. It takes 48 days to cover the entire lake. How many days does it take to cover half the lake?

...

'24 days,' right? So symmetric, so elegant. Alas — 47 days. Doubling has to be thought backwards: if the patch doubles every day, then the day after it covers half, it covers the whole lake — so 48 − 1 = 47. You've now collected the full CRT set: intuitive answers 10 cents / 100 minutes / 24 days, correct answers 5 cents / 5 minutes / 47 days. Frederick tallied the errors: among all wrong answers, the numbers 10, 100, and 24 utterly dominate — people aren't erring at random, they're marching into the same pit in perfect formation.

By now, two thoughts may be forming. Thought one: 'I'm pretty smart, I should be fine.' — Sadly, no. Stanovich and West (2008) found across 7 studies that classic biases — anchoring, framing effects, conjunction effects, sunk cost, base-rate neglect — are essentially uncorrelated with cognitive ability (SAT scores). West, Meserve, and Stanovich (2012) sting even harder: the 'bias blind spot' — believing others are more biased than you are — is not attenuated by cognitive ability; people with higher cognitive ability show an even larger bias blind spot. Intelligence gives you horsepower when you use it; it doesn't guarantee you'll think to use it.

Thought two: 'So System 1 is the villain — from now on I'll think everything through slowly!' — Also wrong. Kahneman himself is explicit: System 1 is the source of many of our errors, and also the source of most of what we do right — and 'doing right' makes up the vast majority of what we do. Having System 2 double-check everything is impossible (attention and computing power are finite) and would grind life to a halt.

So when can intuition be trusted? Kahneman and Klein (2009) — one spent a career on how intuition fails, the other on how it works wonders — co-authored a paper giving two necessary conditions: (1) the environment offers enough predictability and regularity (a high-validity environment); (2) the judge has had the chance to learn those regularities through prolonged practice with timely, clear feedback. When both hold (the immediate calls of chess players, fireground commanders, clinical nurses), intuition is a real skill; when either is missing (long-term political forecasting, stock picking), years of seniority grow only confidence, never accuracy — and 'subjective experience is not a reliable indicator of judgment accuracy.'

So the right takeaway is not 'always think slow' but to install a situational switch: the answer feels suspiciously smooth, the problem involves probability or statistics, or the money or consequences are large — spot any of these three cues, then press the slow-thinking button.

Lily pads double daily: half the pond only on day 47 — exponential intuition fails hardest

自测 · 学完检查一下

想真正动手做题、记进度、攒连胜?到互动课里练。

A coffee and a mug cost $22 in total. The coffee costs $20 more than the mug. How much does the mug cost? (Hold back the answer that pops up in the first 3 seconds.)

答案:$1

Intuition blurts out '$2' — splitting 22 into 20 and 2, suspiciously smooth. Check it and it falls apart: if the mug were $2, the coffee ($20 more) would be $22, totaling $24 — too much. Let the mug be x: x + (x + 20) = 22, so x = $1 (coffee $21). This is a structural twin of the bat-and-ball problem from Frederick's (2005) CRT: bat and ball cost $1.10 with the bat $1.00 more — intuitive answer 10 cents, correct answer 5 cents, and 10 cents utterly dominates the wrong answers, showing everyone falls into the same pit. When you see a smooth 'split the total in two' answer, verify before you speak. (Source: Frederick, S. (2005). Cognitive Reflection and Decision Making, Journal of Economic Perspectives 19(4))

8 printers take 8 minutes to print 8 booklets. How long would 24 of the same printers take to print 24 booklets?

答案:8 minutes

Each printer takes 8 minutes to print 1 booklet; 24 printers running in parallel print 24 booklets in 8 minutes. The blurted '24 minutes' mistakes numerical symmetry for an answer — the same pit as the CRT machines item (5 machines, 5 minutes, 5 widgets; 100 machines for 100 widgets: intuitive 100 minutes, correct 5 minutes). More machines don't make each unit slower — capacity runs in parallel. Frederick's evidence shows 100 utterly dominates the wrong answers, a textbook impulsive intuition. (Source: Frederick, S. (2005). Cognitive Reflection and Decision Making, Journal of Economic Perspectives 19(4))

A patch of mold on a loaf of bread doubles in size every day, and covers the whole loaf on exactly day 30. On which day does it cover half the loaf?

答案:Day 29

Doubling must be thought backwards: the day after the mold covers half, it covers the whole loaf, so 30 − 1 = 29. The intuitive 'Day 15' simply halves 30 — the same pit as the CRT lily-pads item (doubling daily, 48 days to cover the lake; half takes 47 days, not 24): exponential growth doesn't spread evenly, and the final day's doubling equals everything that came before. Frederick's tally shows 24 utterly dominates the wrong answers on the original — 'cut it in half' is exactly the impulse that pops up first for everyone. (Source: Frederick, S. (2005). Cognitive Reflection and Decision Making, Journal of Economic Perspectives 19(4))

Which statement about System 1 and System 2 matches dual-process theory?

答案:System 1 is automatic, fast, nearly effortless, and cannot simply be switched off; System 2 is slow, serial, and effortful, demanding attention — recognizing a friend's face at a glance uses the former, computing √19163 without a calculator uses the latter

The 'System 1 / System 2' labels were coined by Stanovich and West in 2000 and borrowed and popularized by Kahneman in Thinking, Fast and Slow (2011): System 1 is automatic, fast, nearly attention-free, and cannot be switched off; System 2 is slow, serial, and demands effort, motivation, and attention. It is a functional distinction between two types of processing, not a left-brain/right-brain anatomy map; nor is System 1 a synonym for 'emotion.' As for 'System 2 supervises everything' — quite the opposite: the intuition-trap mechanism is precisely System 1 tossing up a fluent candidate answer and System 2 often endorsing it without review. (Sources: Stanovich & West (2000), Behavioral and Brain Sciences; Kahneman (2011), Thinking, Fast and Slow)

True or False: The mechanism of intuition traps is that System 1 tosses up a fluent candidate answer first, and System 2 often signs off on it without any review.

答案:True

True. This is exactly how dual-process theory explains intuition traps: System 1 automatically and rapidly produces a fluent candidate, and System 2, lazy by nature, often waves it through. Frederick's (2005) CRT paper opens with this very distinction to define what the test measures: whether you can suppress the wrong answer that leaps to mind. The answer sheets back it up: even correct answerers often thought of the wrong answer first — 10 cents crossed out and changed to 5 cents is common, while the reverse was never seen — showing the intuitive candidate pops up for everyone; the difference is whether System 2 gets up to check. (Sources: Stanovich & West (2000); Frederick (2005), JEP 19(4); Kahneman (2011), Thinking, Fast and Slow)

What does the 'cognitive miser' refer to?

答案:That when processing information, people default to the least-effort shortcuts and think no harder than necessary — an energy-saving default design that works fine most of the time, though trap questions land exactly where it fails

The 'cognitive miser' was proposed by Fiske and Taylor in Social Cognition (1984): people processing information reach for least-effort shortcuts and ready-made schemas, rather than gathering full evidence and reasoning it through as the earlier 'naive scientist' model assumed. It is neither a moral judgment nor a defect, but a default energy-saving design that is usually good enough. Kahneman calls the same phenomenon the 'law of least effort'; Stanovich treats 'miserly processing' as the core mechanism behind smart people making dumb decisions (dysrationalia) — so 'the highly educated aren't misers' fails too, and the CRT tests precisely whether you can overcome the miser at the crucial moment. (Sources: Fiske & Taylor (1984), Social Cognition; Stanovich (2009), What Intelligence Tests Miss; Kahneman (2011))

Frederick (2005) asked people who had answered the bat-and-ball problem to estimate 'what share of other people would get this question right.' What did he find?

答案:Those who gave the wrong answer '10 cents' estimated on average that 92% of people would get it right — finding the question easier than the correct answerers did (who estimated 62%), with no subjective alarm at the moment of being fooled

Frederick's (2005, pp. 27-28) finding runs in exactly the opposite direction from the distractor: those who answered '10 cents' estimated on average that 92% of people would get it right — they never noticed the trap at all; those who correctly answered '5 cents' estimated only 62% (both groups substantially overestimated the actual accuracy). This explains why intuition traps are dangerous: the wrong answer comes bundled with a 'this is easy' feeling of fluency, there is no subjective alarm at the moment of error, and the wrong are even more confident than the right — echoing Kahneman & Klein's (2009) point that 'subjective experience is not a reliable indicator of judgment accuracy.' (Sources: Frederick (2005), JEP 19(4): 25-42, pp. 27-28; Kahneman & Klein (2009), American Psychologist 64(6))

By Kahneman and Klein's (2009) two necessary conditions for trustworthy intuition — a high-validity environment plus prolonged practice with timely, clear feedback — which of these 'intuitions built on years of experience' is most likely to be reliable?

答案:A clinical nurse's immediate sense that a patient has suddenly taken a turn for the worse — a regular environment, with years of practice and timely, clear feedback

Kahneman & Klein's (2009) two necessary conditions: (1) the judgment environment offers enough predictability and regularity (a high-validity environment); (2) the judge has had ample opportunity to learn those regularities through prolonged practice with timely, clear feedback. A clinical nurse's immediate calls (like those of chess players and fireground commanders) satisfy both, so the intuition is a real skill; long-term political forecasting and stock picking sit in low-validity environments, where years of seniority grow only confidence, not reliable intuition; lottery numbers have no regularities to learn at all. And remember the paper's warning: 'subjective experience is not a reliable indicator of judgment accuracy' — feeling sure is not being right. (Source: Kahneman & Klein (2009). Conditions for Intuitive Expertise, American Psychologist 64(6): 515-526)

True or False: The stronger your cognitive ability (say, a higher SAT score), the smaller your 'bias blind spot' — the tendency to think others are more biased than you — so smart people see their own biases more clearly.

答案:False

False — the direction is exactly the opposite. West, Meserve, and Stanovich (2012) found the bias blind spot is not attenuated by cognitive ability or thinking dispositions; higher cognitive ability actually comes with a larger bias blind spot. Stanovich and West (2008) also found across 7 studies that many classic biases — anchoring, framing effects, conjunction effects, sunk cost, base-rate neglect — are essentially uncorrelated with cognitive ability (SAT). The CRT data point the same way: the top-scoring MIT sample averaged only 2.18 out of 3, and 52% of MIT respondents missed at least one item. Intelligence provides horsepower when you use it; it doesn't guarantee you'll think to use it. (Sources: West, Meserve & Stanovich (2012), JPSP 103(3): 506-519; Stanovich & West (2008), JPSP 94(4): 672-695; Frederick (2005), Table 1)

True or False: Since intuition gets fooled so easily, the most rational way to live is to treat System 1 as the enemy — switch off intuition and think everything through slowly with System 2.

答案:False

False. Kahneman is explicit in Thinking, Fast and Slow: System 1 is the source of many of our errors and also the source of most of what we do right — and 'doing right' makes up the vast majority of what we do; having System 2 double-check everything is impossible (attention and computing power are finite) and would grind life to a halt. Besides, System 1 can't be switched off even if you try. The workable strategy is a situational switch: shift to slow thinking when you spot the cues — 'the answer feels suspiciously smooth,' 'probability or statistics involved,' 'large sums or high stakes.' And on the flip side, Kahneman & Klein (2009) note that under a high-validity environment plus ample practice with timely feedback, intuition itself is a genuine skill. (Sources: Kahneman (2011), Thinking, Fast and Slow; Kahneman & Klein (2009), American Psychologist 64(6))

想边练边学,而不只是读?

到互动课里答题、记进度、攒连胜——游客即可试学,无需注册。

进入互动课程 →

Learn something new — don't miss updates

New courses, features and learning tips. Occasional emails, unsubscribe anytime.