Watson Glaser Practice Test: 20 Questions in 15 Minutes

All five Watson Glaser sections in one timed paper: rate inferences on the five-point scale, spot unstated assumptions, verify deductions, judge interpretations, and weigh arguments as strong or weak. Original questions, 15 minutes, per-section report on submission.

The Watson Glaser Critical Thinking Appraisal is the assessment law firms and professional-services employers use to screen applicants, and its five sections are always the same: drawing inferences on a five-point scale, recognizing unstated assumptions, testing deductions, judging interpretations, and rating arguments strong or weak. This practice paper follows that structure section for section with 20 original questions in 15 minutes.

To be clear about what this is: an independent practice paper in the Watson Glaser style, written for this site. It is not affiliated with, endorsed by, or sourced from Pearson, the publisher of the official assessment — which also means none of these questions can turn up in your real sitting, and practicing here spoils nothing.

Scoring is one point per question, no negative marking, nothing revealed until you submit. The report breaks accuracy down by section — the five-point inference scale and the strong/weak argument call are where most first attempts bleed marks — and every question carries an explanation of the judgment standard, not just the answer.

Timed mock test

Sit the timed paper

The five-section law-firm screening format — original questions, timed.

Questions
20 questions
Time limit
15:00
Scoring
1 point per question, no penalty for wrong answers
Feedback
No right/wrong shown mid-test

The clock starts the moment you hit start. You can jump between questions and change answers.

The five sections, and what each one is actually testing

Every Watson Glaser paper is built from the same five sections, always in the same order. They are not five flavours of the same question — each one asks you to switch to a different standard of proof, and most of the marks are lost by carrying the previous section's standard into the next.

1. Inference. You get a short passage of facts, then a proposed inference, and rate it on five points: True, Probably True, Insufficient Data, Probably False, False. The standard here is probability given only the passage. The trap is that "Insufficient Data" is correct far more often than candidates expect — if the passage does not bear on the statement at all, that is the answer, even when the statement is obviously true in the real world.

2. Recognition of Assumptions. A statement is given, then a proposed assumption, and you answer Assumption Made or Assumption Not Made. The standard is necessity: would the speaker's statement collapse without it? Not "is this plausible", not "would a reasonable person also believe this". If the statement still stands with the assumption removed, it was not made.

3. Deduction. Premises, then a conclusion: Conclusion Follows or Does Not Follow. This is the only section with a strict formal standard. Real-world truth is irrelevant — a conclusion can be factually absurd and still follow. Treat the premises as complete and closed.

4. Interpretation. A passage, then a proposed conclusion: Conclusion Follows or Does Not Follow. Weaker than Deduction — you are asked whether the conclusion follows beyond reasonable doubt, not with logical necessity. But it is far stricter than Inference. Generalising from a sample to a whole population usually fails here.

5. Evaluation of Arguments. A question, then an argument: Strong or Weak. An argument is strong only if it is both important and directly relevant to the question. Arguments that are true but trivial are weak. Arguments that appeal to what most people think are weak.

The standard-of-proof ladder, which is where most marks are lost

The five sections form a ladder from loosest to strictest, and knowing where you are on it is worth more than any individual technique:

  • Inference — probability, judged only against the passage
  • Interpretation — beyond reasonable doubt
  • Deduction — strict logical necessity
  • Assumptions — necessity, but of a premise rather than a conclusion
  • Evaluation — relevance and importance, not validity at all

The classic failure is arriving at Deduction still in Inference mode and marking a conclusion as following because it is probably true. Deduction does not care what is probable. The reverse failure is just as common: candidates who tighten up during Deduction stay tight through Interpretation and reject conclusions that do meet the beyond-reasonable-doubt standard.

When you practise, do not just check whether you got the item right. Check whether you applied the right standard. Two candidates with the same raw score can have completely different problems.

Who uses it, and what score you actually need

Watson Glaser is most strongly associated with law — in the UK it is a standard part of training-contract screening at large commercial firms — but it is also used for management consulting, professional services, graduate schemes, and internal promotion into analytical roles.

There is no universal pass mark, and any site quoting one is guessing. Employers set their own threshold, usually as a percentile against a norm group rather than a raw score. Two things follow from that:

  • The same raw score can pass at one employer and fail at another, because the norm group differs. Being compared against "graduate applicants" is easier than against "qualified solicitors".
  • A widely repeated figure for competitive law screening is around the 70th–80th percentile, but treat that as folklore unless your employer publishes it. Ask the recruiter what norm group they use — most will tell you.

What is reliably true: the test is time-pressured, sections are not equally weighted in difficulty, and unanswered items score zero. Finishing matters.

Watson Glaser vs the LSAT and other reasoning tests

These get confused because they all describe themselves as testing critical thinking, but they measure different things and reward different preparation.

Versus the LSAT. LSAT Logical Reasoning gives you one argument and asks a specific question about it — find the flaw, strengthen it, identify the assumption. Watson Glaser gives you five different tasks and asks you to switch standards between them. LSAT rewards depth on a single argument; Watson Glaser rewards accuracy about which standard applies. If you have done LSAT prep, your Deduction and Assumptions sections will be strong and your Inference section will probably be your weakest — LSAT trains you to reason past a gap, and Inference punishes exactly that.

Versus SHL and Kenexa reasoning tests. Those are usually verbal, numerical, or inductive, with a single question type repeated. Watson Glaser is the only common employer test that deliberately switches the standard of proof mid-paper.

Versus the GMAT. GMAT Critical Reasoning is closer to the LSAT in structure. If you are preparing for both, the overlap is real but partial — see the GMAT critical reasoning questions page for that format.

The five mistakes that cost the most marks

Importing outside knowledge. Every section is judged against the passage alone. If you know from experience that a proposed inference is true, that is irrelevant unless the passage supports it. This is the single largest source of lost marks, and it gets worse the more expert you are in the subject matter.

Avoiding "Insufficient Data". Candidates treat it as a cop-out and force a judgement. It is a real answer and often the correct one.

Confusing necessary with plausible in Assumptions. Test it by negation: remove the assumption and see whether the statement still works. If it does, the assumption was not made.

Judging arguments by their conclusion in Evaluation. An argument for a position you agree with can be weak, and one for a position you reject can be strong. Rate the reasoning, not the side.

Running out of time on the last section. Evaluation of Arguments comes last and is often rushed. Unanswered items score zero, so a quick judged answer beats a blank.

A worked example: one inference item, all five verdicts

The inference section is where the standard-of-proof ladder bites, so here is one original item worked all the way through. Statement: A city library extended its weekday opening hours by two hours in March. In April, weekday borrowing was 12% higher than in February. Weekend borrowing, where hours did not change, was unchanged.

Proposed inference: The extra hours caused the rise in weekday borrowing. Walk the ladder. True requires the passage to state or strictly entail it — it does not; nothing rules out a new-books campaign that also started in March. False requires the passage to contradict it — it does not. So the answer is one of the three middle rungs, and the question is whether the passage makes the inference more likely than not. The weekend figure is the deciding detail: hours changed only on weekdays, and only weekday borrowing rose. That is a control, not a coincidence, and it tips the inference to Probably true. Marking it Insufficient data is the most common miss, and it comes from demanding proof where the ladder only asks for likelihood.

Two more inferences from the same passage, for contrast. "Total borrowing rose in April"Probably true: weekdays rose 12% and weekends were flat, so the total went up unless weekday volume is trivial, which the passage gives no reason to assume. "The library also extended weekend hours"False: the passage says weekend hours did not change. Note that False here rests on an explicit statement, not on absence of information — absence of information is what Insufficient data is for.

  • The same discipline applies to the assumption section. Argument: "Move the team meeting to 8 a.m. so more people can attend." Proposed assumption: "Some people cannot attend at the current time"made, because the argument collapses without it. Proposed assumption: "8 a.m. meetings are shorter"not made; it may be true, but the argument does not need it.

The paper above has an item of each type in every section, and the report tells you, for each miss, which rung you chose and which the passage supported.

How to prepare in the week before

Watson Glaser rewards calibration far more than knowledge, so a short focused run is more useful than weeks of reading.

  • Do one full timed paper first, before any study, to find which of the five sections is actually your weak one. Most people guess wrong about this.
  • Review by section, not by score. For every item you missed, write down which standard of proof you applied and which one you should have applied.
  • Practise Inference separately. It is the section most people are worst at and the one least like anything else they have done.
  • Do the last full practice at the same time of day as the real test, and finish every item even if you are unsure.

The test above follows the same five-section structure and is timed. Nothing is marked until you submit, and the report breaks your score down by section so you can see which standard you are misapplying.

Where to go after the report

Frequently asked questions

Is this the official Watson Glaser test?

No. It is an original practice paper in the same five-section format, independently written for this site and not affiliated with Pearson. That independence is a feature for practice: no question here can appear in your real assessment.

Which employers use the Watson Glaser?

It is best known as the law-firm screener — many UK and international firms use it for training-contract and vacation-scheme applications — and it also appears in consulting, finance, and government selection. The five-section skills transfer directly to any critical-thinking assessment.

How is it scored and what should I aim for?

One point per question here. Real Watson Glaser sittings are typically norm-referenced against a comparison group, and competitive law-firm cohorts are strong — treat 75 percent-plus as the zone to work toward, and use the per-section report to find the section holding you back.

What is the hardest section?

By a distance, inference — because it is scored on a five-point scale where 'probably true' and 'insufficient data' sit close together. The discipline is to ask what the facts force, then how far short of certainty they stop. The explanations here drill exactly that boundary.

How should I use this paper with the critical thinking test?

They are the same format with disjoint questions — a natural Paper 1 and Paper 2. Sit one cold for a baseline, study its explanations, then sit the other to measure whether the judgment standards stuck. Two per-section reports side by side show your trend.