The hardest IQ tests include unsupervised high-range puzzles and high-ceiling tests like the WAIS-5 and SB-5. See why harder rarely means more accurate.
Dr. Russell T. WarneChief Scientist
Share
There is no officially recognized hardest IQ test, because "hardest" can mean two different things. If it means the most difficult items, the famous candidates are unsupervised "high-range" tests such as the Mega Test, which Omni magazine billed in 1985 as the world's most difficult IQ test. If it means the highest score a test can actually measure, the answer lies with professionally normed instruments such as the Stanford-Binet Fifth Edition with its extended scoring and the Wechsler scales with extended norms. This article explains what makes an IQ test hard, why the most difficult tests are rarely the most trustworthy, and which mainstream tests can measure the highest scores.
What makes an IQ test hard
Psychologists separate two ideas that everyday language blurs together. The first is "item difficulty," which is simply the proportion of test takers who answer a question correctly. A vocabulary word that 90 percent of adults know is an easy item; a logic puzzle that fewer than 1 percent solve is a very hard one. The second idea is the test "ceiling," the highest score an instrument can report before it runs out of room to distinguish one strong performer from another.
The two are related but far from identical. A test stuffed with brutally hard items feels harder, but that alone does not make it a better measuring tool. Good tests match item difficulty to the ability range they need to describe. Most professional tests concentrate their items around average difficulty, because that is where most people score and where precision matters for the majority of decisions. A test built only from extreme items would tell you almost nothing about a typical adult, since nearly everyone would miss nearly everything.
When a capable person answers every item a test offers, the test can no longer tell how much higher that person might have scored. This is the well documented ceiling effect, and it is the reason "hardest" questions and "highest ceiling" are worth discussing separately.
High-range tests: the famous "world's hardest" puzzles
The tests most often called the hardest in the world come from the high-range testing community rather than from clinical psychology. The best known is the Mega Test, created by Dr. Ronald K. Hoeflin. According to the Mega Society, interest in the society grew sharply after the Mega Test appeared in Omni magazine in April 1985, where it was presented as the "World's Most Difficult IQ Test." The society describes itself as open to people who have scored at the one-in-a-million level on a test credibly claimed to discriminate at that level, and its notable Mega Test takers include the columnist Marilyn vos Savant.
Tests of this type consist of a few dozen extremely difficult verbal analogies, number sequences, and spatial puzzles, worked on alone with no proctor and no fixed time limit. Later examples include Hoeflin's Titan, Ultra, and Power tests. Similar instruments serve as admission routes for various high-IQ societies that set cutoffs well beyond the 98th percentile.
The genre has a persistent practical problem. The Mega Society itself reports that the Mega Test was compromised, so scores earned after 1994 are no longer accepted, and that the Titan Test was retired in 2020 for the same reason. Once answers circulate, an unsupervised test loses whatever measurement value it had.
Why unsupervised high-range tests fall short
High-range tests are impressive puzzle collections, and the communities around them are sincere about intellectual challenge. As measurement instruments, though, they fall short of mainstream standards in several specific ways.
• No supervision: nobody verifies who took the test, how long they worked, or what help they used. The repeated compromises of the Mega Test, the Titan Test, and the earlier Langdon Adult Intelligence Test show how fragile unproctored scoring is.
• Self-selected samples: a "norm sample," the reference group that converts raw performance into an IQ score, should represent the general population. High-range tests are normed on the volunteers who happen to submit answer sheets, a group that is small and unusually able. Claims about one-in-a-million rarity cannot be established from a few thousand self-selected participants.
• Thin validity evidence: "validity" refers to the documented evidence that scores mean what they are claimed to mean. The Standards for Educational and Psychological Testing, published jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, expect test makers to document validity, reliability, and norming. High-range tests rarely publish evidence of this kind in peer-reviewed outlets.
• No institutional acceptance: American Mensa states that qualifying tests must be administered by a neutral and qualified third party in a traditional testing environment, and that it does not accept unsupervised testing, specifically including tests administered over the internet. Schools, clinics, and courts take the same position.
None of this makes high-range tests worthless as recreation. It does mean their scores should be read as puzzle performance, and the same caution applies to any unsupervised test that reports numbers far above the range mainstream instruments can support.
The mainstream tests with the highest ceilings
If "hardest" means the highest score a validated test can measure, the answer comes from the major clinical batteries.
The current adult standard is the WAIS, now in its fifth edition. Pearson published the WAIS-5 in 2024 for ages 16 through 90, with a full scale IQ derived from seven subtests in about 45 minutes. Its official score reports display the full scale IQ and index scores on a scale from 40 to 160, four standard deviations either side of the mean of 100. A score of 160 already corresponds to roughly the top 3 people in 100,000, so this ceiling covers almost everyone the test will ever meet.
The Stanford-Binet Intelligence Scales, Fifth Edition, authored by Gale H. Roid and published in 2003, has long been a preferred choice for gifted assessment. The publisher notes that the test was normed on 4,800 people ages 2 to 85 and up, and that its manual introduces an Extended IQ scale supporting full scale IQ scores substantially higher than 160, along with scores well below 40 at the other extreme.
For children, Pearson released extended norms for the WISC-V in a 2019 technical report by Susan Engi Raiford and colleagues. According to the report, the extended norms exist to more clearly identify highly gifted children with composite scores far above 130. They apply when a child reaches the maximum scaled score of 19 on one or more subtests, and they raise the subtest maximum to 28 points and composite scores to 210 points. The same report cautions that the extended norms are not useful for most children.
Even these figures deserve humility. Very few members of any norm sample score at the extremes, so estimates far above 145 rest on thinner data than estimates near the middle, and the confidence interval around an extreme score is wide. Extended scoring extends the ruler; it does not make the far end of the ruler as precise as the middle.
Why the hardest test is not the most accurate
Accuracy in testing means "reliability," the consistency of scores across items and occasions, combined with validity and sound norms. None of those properties follows from difficulty. A maximally hard test is actually a poor measure for almost everyone, because most people would score near the bottom, where the test cannot distinguish among them at all.
Mainstream psychometricians, including Dr. Russell T. Warne, consistently emphasize that a good test is one whose difficulty matches the people it measures, whose scores are stable, and whose norms come from representative samples. By that standard, the most accurate tests for a typical adult are well constructed instruments centered on the middle of the ability range, with enough headroom to describe strong performers. Distinctions between scores above 160 are rarely dependable, and no major decision in education or employment turns on them. An extremely hard test can flatter or frustrate; it cannot substitute for evidence of sound measurement.
How to get a score you can trust
For most people the practical question is a trustworthy score, and that means a professionally built test with published norms rather than the most punishing puzzle set available. The RIOT IQ test, the Reasoning and Intelligence Online Test, was developed by RIOT IQ with psychometrician Dr. Russell T. Warne for adults 18 and older. It uses 15 subtests across six cognitive indices (verbal reasoning, fluid reasoning, spatial ability, working memory, processing speed, and reaction time), takes about 52 minutes, and reports scores on the familiar mean-100, standard-deviation-15 scale. Its length is a deliberate reliability choice, since more items mean less measurement error. It does not replace an individually administered diagnostic evaluation, and no online test qualifies you for high-IQ society membership. If you want a rigorous, research-based estimate of your IQ, you can take the RIOT IQ test online today.
Frequently asked questions
What is the hardest IQ test in the world?
There is no official answer. The Mega Test was billed by Omni magazine in 1985 as the world's most difficult IQ test, but its items being hard did not make it a validated measure, and its scores after 1994 are no longer accepted even by the Mega Society.
Which IQ test has the highest ceiling?
Among professionally normed tests, the WISC-V extended norms reach composite scores of 210 for children, the Stanford-Binet Fifth Edition's Extended IQ scale supports scores above 160, and the WAIS-5 reports adult scores up to 160.
Are harder IQ tests more accurate?
No. Accuracy depends on reliability, validity, and representative norms. A test made only of extreme items measures most people poorly because nearly everyone scores near its floor.
Can a hard online IQ test qualify me for Mensa?
No. American Mensa requires supervised administration by a neutral, qualified third party and does not accept unsupervised or internet-based testing. The RIOT IQ test does not qualify you for Mensa membership either.
Is the Mega Test still used?
The Mega Society reports that the test was compromised and that scores earned after 1994 are not accepted. Its successor, the Titan Test, was retired in 2020 for the same reason.
Do IQ scores above 160 mean anything?
They sit at or beyond the ceiling of most professional tests, and the uncertainty around them is large. A very high score describes performance on one occasion; it is a description of a score, never a full description of a person.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
2. American Mensa. (n.d.). Qualifying test scores.. us.mensa.org
3. Mega Society. (n.d.). About the Mega Society.. megasociety.org
6. Raiford, S. E., Courville, T., Peters, D., Gilman, B. J., & Silverman, L. (2019). WISC-V technical report #6: Extended norms. NCS Pearson. pearsonassessments.com