🧠 Psychology, Made Simple · Understand Yourself & Others

Base Rate Fallacy Example: Medical Testing Walkthrough and MCAT-Style Quiz

Start with a hypothetical 10,000-person screening and separate sensitivity, false-positive rate, and positive predictive value

一句话先懂 · TL;DR

What is the base rate fallacy? Work through a medical testing example with natural frequencies, then apply it in six MCAT-style quiz questions.

Before Asking How Good the Test Is, Ask How Common the Disease Is

Suppose you take a hypothetical screening test for a rare disease. The disease prevalence is 1%. The test catches 90% of people who have it, giving it 90% sensitivity. It also returns a positive result for 5% of healthy people, giving it 95% specificity and a 5% false-positive rate.

Your result is positive. It is tempting to grab the 90% figure and conclude that you are 90% likely to be sick. But 90% answers, “If someone has the disease, how likely are they to test positive?” You need the reverse: “If someone tests positive, how likely are they to have the disease?” Reversing the direction can radically change the answer.

This is a common form of the base rate fallacy: vivid case-specific evidence pushes aside the event's original frequency, or base rate.

These measures are not interchangeable. Sensitivity is P(positive | disease). Specificity is P(negative | no disease). The false-positive rate is P(positive | no disease), which equals 1 minus specificity. Positive predictive value is the quantity you actually want: P(disease | positive). Overall accuracy is the proportion of all classifications that are correct, and it changes with the mix of diseased and healthy people in the sample. Hearing that a test is “99% accurate” without knowing the sample composition or confusion matrix is generally not enough to calculate the probability of disease after a positive result.

🔆Think of a test as a nightclub bouncer. Sensitivity asks how many actual gate-crashers the bouncer stops. Specificity asks how many legitimate guests the bouncer lets through. Positive predictive value asks how many of the stopped people are truly gate-crashers. If almost no gate-crashers show up that night, even a sharp-eyed bouncer may stop a group containing plenty of innocent guests.

The 10,000-Person Table: Make the Denominator Visible

Return to the hypothetical screening example. Instead of juggling three percentages, send 10,000 people through a table. At 1% prevalence, 100 have the disease and 9,900 do not. With 90% sensitivity, 90 of the 100 patients test positive. With a 5% false-positive rate, 495 of the 9,900 healthy people also test positive.

Now look only at the positive group: 90 true positives plus 495 false positives gives 585 positive results. The positive predictive value is 90 ÷ 585 ≈ 15.4%.

The result genuinely raises the probability of disease from 1% to about 15.4%, so the test provides important evidence. It simply does not raise it to 90%. We can also check overall accuracy: 90 true positives plus 9,405 true negatives gives 9,495 correct classifications, or 94.95% accuracy. That is neither the 90% sensitivity nor the 15.4% positive predictive value.

For an MCAT-style scenario, draw four cells: disease and positive, disease and negative, healthy and positive, healthy and negative. If the question asks about disease after a positive result, the denominator must include everyone who tested positive.

💡Finding the denominator is like arranging a group photo. If the question says “among people who tested positive,” invite every positive result into the frame first, then count how many actually have the disease. Do not accidentally use the separate photo containing all patients.

Choose the Right Base Rate—and Do Not Count the Same Evidence Twice

Suppose a battery brand has a company-wide failure rate of just 0.2%, but your battery came from a batch with a documented assembly problem and a 4% failure rate. Your device has now issued an overheating alert. Reassuring yourself with 0.2% would use a reference class that is too broad. If the batch information is reliable and applies to your unit, 4% is the more relevant starting point.

Choosing a base rate does not mean always reaching for the largest dataset or picking the number that suits your mood. The reference class should match the current case in relevant features such as time, location, source, and causal mechanism while still having enough observations to be stable.

Then let the evidence update that starting point. How sensitive is the alert to real failures? How often does it fire for normal batteries? Does a follow-up test actually supply new information?

If two apps merely read the same physical sensor, two pop-ups are not two independent pieces of evidence. A check based on a different mechanism with different error sources generally provides a stronger update. Even when two results are independent, however, “two positives” does not justify claiming 100% certainty. You still need the base rate and each result's conditional probabilities.

⚠️Do not count the same witness twice just because they changed jackets. Two signals on the screen do not necessarily mean twice as much information. First ask whether they share the same data source, error, or trigger.

自测 · 学完检查一下

想真正动手做题、记进度、攒连胜?到互动课里练。

On a secondhand marketplace, about 0.5% of listings are fraudulent. A risk system flags 80% of fraudulent listings and incorrectly flags 2% of legitimate listings. After one listing is flagged, an analyst says, “There is an 80% chance it is fraudulent.” What is the analyst's main error?

答案:He took the alert's hit rate as the final answer: the probability he wants also depends on the 0.5% fraud share and the 2% false-flag rate, so his conditional runs in the wrong direction

Correct: 80% is P(flagged | fraud), while the analyst needs P(fraud | flagged). The latter also depends on the fraud base rate and the false-flag rate among legitimate listings. The most tempting wrong answer is to keep 0.5% as the final probability. That recognizes the base rate but throws away the new evidence supplied by the flag; the right move is to update the base rate, not freeze it. The 98% figure is the probability that a legitimate listing avoids a false flag, while subtracting 2% from 80% combines rates defined over different groups.

A mountain rescue system monitors 10,000 hiker check-ins. Historically, about 2% correspond to real emergencies. The system alerts on 90% of real emergencies and falsely alerts on 4% of non-emergencies. After an alert, what is the closest probability that there is a real emergency?

答案:About 31%

Correct: among 10,000 check-ins, about 200 involve real emergencies, producing 180 true alerts. Of the remaining 9,800 non-emergencies, 4% produce false alerts, or 392. That gives about 572 alerts in total, with real emergencies accounting for 180 ÷ 572 ≈ 31%. The most tempting wrong answer is “About 90%,” but 90% is the probability of an alert given an emergency, not the probability of an emergency given an alert. “About 86%” incorrectly subtracts 4% from 90%, while “About 2%” fails to update the base rate at all.

A content platform announces only that its harmful-content classifier achieved 97% overall accuracy on an evaluation set. The model then flags a post. Which conclusion is the most rigorous?

答案:Overall accuracy alone cannot give the probability that this post is harmful; we also need the deployment base rate and class-specific performance

Correct: overall accuracy blends true positives and true negatives in an average weighted by the evaluation set's class proportions, so the same 97% can hide very different sensitivities and false-flag rates. The most tempting wrong answer claims the flag is 97% reliable “because accuracy already averages both error types”—but averaging is exactly the problem: it collapses the class-specific information you would need to unpack. The 3% figure is the error rate over all evaluated examples, not the error rate among flagged posts, and a classifier can reach 97% accuracy by excelling on a dominant class, so no 97% floor on sensitivity or specificity follows.

In a hypothetical MCAT-style question, the same test has 80% sensitivity and 90% specificity in two populations. Disease prevalence is 20% in Population A and 2% in Population B. Because a positive result is rarer and more surprising in Population B, a positive person in B is actually more likely to have the disease.

答案:False

Why “False” is correct: with identical test performance, the population with the higher disease base rate has the higher positive predictive value. Per 1,000 people, Population A produces about 160 true positives and 80 false positives, for a positive predictive value near 67%. Population B produces about 16 true positives and 98 false positives, or roughly 14%. The tempting intuition treats “more surprising” as “more trustworthy,” but a rarer disease supplies fewer true cases, allowing false positives to dominate the positive group.

A courier company is assessing whether a high-value parcel was stolen after it passed through Hub X overnight before a gap appeared in its scan history during the holiday rush. Which base rate is the best starting point for interpreting later evidence?

答案:The proportion of scan-interrupted parcels later confirmed stolen at Hub X, during comparable holiday periods and under the same type of shipping service

Correct: this reference class matches the current parcel on location, season, service type, and the known scan interruption. The most tempting wrong answer is the nationwide annual theft rate. It may use more observations, but it mixes parcels with different values, routes, processes, and time periods, making it too broad for this case. Looking only at parcels already reported missing conditions on an outcome and answers a different question, while news coverage is selectively sampled and supplies no stable denominator.

A hiring committee receives two glowing reference letters from two of a candidate's former colleagues. One committee member says, “Two letters are two independent pieces of evidence, so we should update the probability that this candidate is excellent twice.” Which new fact most directly weakens the independence assumption?

答案:Both referees formed their entire impression of the candidate from the same single flagship project the candidate led

Correct: both letters share one upstream source—that project's success. If the project's shine owed something to luck or to other people's work, that distortion feeds into both letters at once, so the second letter adds no fresh, independent observation. The most tempting wrong answer is that such candidates are rare: a low base rate lowers the final probability, but it answers the starting-point question and says nothing about whether the two pieces of evidence are independent. Differing strictness only means the referees use different judgment thresholds, while their sources can still overlap; and a one-week gap in arrival is just a delivery delay, not a new observation.

想边练边学,而不只是读?

到互动课里答题、记进度、攒连胜——游客即可试学,无需注册。

进入互动课程 →

Learn something new — don't miss updates

New courses, features and learning tips. Occasional emails, unsubscribe anytime.