The Cognitive Reflection Test is a 3-question measure of whether you check your gut answer. The questions, average scores, and what it says about IQ.
Dr. Russell T. WarneChief Scientist
Share
The Cognitive Reflection Test, or CRT, is a three-question test introduced by the behavioral economist Shane Frederick in 2005. It does not measure how much you know or how fast you compute. It measures whether you catch yourself: each question invites a quick, confident answer that happens to be wrong, and the test scores your "ability or disposition to resist reporting the response that first comes to mind." The most famous item, the bat and ball problem, has become one of the most widely shared puzzles in psychology. This article gives the three questions, explains what the test really measures, reports how people actually score, and covers the test's biggest modern weakness: too many people have seen the answers.
The three questions, and why the obvious answers are wrong
• The bat and ball: A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost? The answer that springs to mind is 10 cents. But then the bat would cost $1.10 and the total $1.20. The ball costs 5 cents; the bat, $1.05.
• The widget machines: If it takes 5 machines 5 minutes to make 5 widgets, how long would it take 100 machines to make 100 widgets? The intuitive answer is 100 minutes. But each machine makes one widget in 5 minutes, so 100 machines make 100 widgets in 5 minutes.
• The lily pads: A patch of lily pads doubles in size every day and covers a lake in 48 days. How long to cover half the lake? The pull is to say 24. Since the patch doubles daily, the lake was half covered one day before it was full: 47 days.
Frederick noted that the items are "easy" in a particular sense: the solution "is easily understood when explained," yet getting there requires suppressing an answer that arrives "impulsively." In his data, catching the error was nearly the whole game: almost everyone who rejected "10 cents" went on to give the right answer.
What the CRT actually measures
Psychologists describe two modes of thinking, often called "System 1" and "System 2": fast, automatic processing that generates instant impressions, and slow, effortful processing that can check them. The CRT is a test of whether System 2 shows up for work. Daniel Kahneman built the opening chapters of Thinking, Fast and Slow around exactly this machinery, with the bat and ball as a centerpiece.
Later research sharpened the picture. Toplak, West, and Stanovich found that while the CRT correlates substantially with cognitive ability, it uniquely predicts performance on "heuristics-and-biases" tasks even after controlling for intelligence and executive function, because standard tests do not measure the tendency toward "miserly" processing, accepting the first plausible answer instead of checking it. Frederick also documented behavioral correlates: higher scorers were more patient in intertemporal choices and more willing to take favorable gambles.
How people actually score
Frederick administered the test to 3,428 respondents across 35 studies. The overall mean was 1.24 of 3. A third of respondents scored zero, and only 17 percent answered all three correctly. Even at MIT, the highest-scoring group he tested, the mean was 2.18, with fewer than half of students getting a perfect score; Princeton students averaged 1.63 and Carnegie Mellon students 1.51. Frederick cautioned that his samples leaned on selective universities, so population differences are likely larger than his tables show. If you scored 2 of 3 cold, you did fine.
The familiarity problem
The CRT's fame became its weakness. In one study of online research participants, 62.6 to 94 percent had seen at least one CRT item before. Haigh found that participants who had previously encountered the problems scored 2.36 on average versus 1.48 for newcomers, and people who had seen a specific item nearly always solved it. A separate study of 2,272 people found 44 percent were already familiar with the tasks.
Researchers responded on two fronts. New versions appeared, including a four-item "CRT-2" with less numeric content (sample: "If you're running a race and you pass the person in second place, what place are you in?"). And Bialek and Pennycook showed across six studies that repeated exposure inflates scores without destroying the test's predictive correlations. The practical upshot for readers: if you have seen the bat and ball before, a perfect score says little about your reflectiveness.
Is the Cognitive Reflection Test an IQ test?
No, and its author never claimed it was. CRT scores correlate moderately with ability measures, about .44 with SAT scores and .43 with the Wonderlic in Frederick's data, which is remarkable efficiency for three items; he noted its predictive validity "equals or exceeds" tests with up to 215 items. But a moderate correlation is not identity, and three items is a noisy basis for judging any individual, a limitation we cover in our guide to what a quick IQ test can and cannot tell you. The CRT and the Wonderlic occupy the short end of the spectrum; a full battery sits at the other. The Reasoning and Intelligence Online Test from RIOT IQ measures six abilities separately, from verbal and fluid reasoning to working memory and reaction time, with norms from real test-takers. If the bat and ball made you curious where your reasoning actually stands, you can take a free IQ test and get a real profile instead of a three-item verdict.
Frequently asked questions
What does the Cognitive Reflection Test measure?
The disposition to override an intuitive first answer and check it, sometimes called resistance to miserly processing. It is related to intelligence but distinct from it.
What are the answers to the three CRT questions?
5 cents, 5 minutes, and 47 days. The intuitive wrong answers are 10 cents, 100 minutes, and 24 days.
What is a good CRT score?
In Frederick's 3,428-person dataset the average was 1.24 of 3, and only 17 percent scored a perfect 3. MIT students averaged 2.18.
Why do smart people get the bat and ball problem wrong?
Because System 1 supplies a plausible answer instantly and System 2 often declines to check it. The error is about reflection, not arithmetic skill.
Does it still count if I have seen the questions before?
Familiarity inflates scores substantially, so a perfect score on known items means little. Alternate versions like the CRT-2 exist partly for this reason.
References
1. Frederick, S. (2005). Cognitive reflection and decision making. Journal of Economic Perspectives, 19(4), 25-42. doi.org/10.1257/089533005775196732
2. Toplak, M. E., West, R. F., & Stanovich, K. E. (2011). The Cognitive Reflection Test as a predictor of performance on heuristics-and-biases tasks. Memory & Cognition, 39(7), 1275-1289. doi.org/10.3758/s13421-011-0104-1
3. Thomson, K. S., & Oppenheimer, D. M. (2016). Investigating an alternate form of the cognitive reflection test. Judgment and Decision Making, 11(1), 99-113. sjdm.org
4. Haigh, M. (2016). Has the standard Cognitive Reflection Test become a victim of its own success? Advances in Cognitive Psychology, 12(3), 145-149. doi.org/10.5709/acp-0193-5
5. Stieger, S., & Reips, U.-D. (2016). A limitation of the Cognitive Reflection Test: Familiarity. PeerJ, 4, e2395. pmc.ncbi.nlm.nih.gov
6. Bialek, M., & Pennycook, G. (2018). The cognitive reflection test is robust to multiple exposures. Behavior Research Methods, 50(5), 1953-1959. doi.org/10.3758/s13428-017-0963-x