🇺🇸The official website of Riot IQ
Log in
  • Home
  • About

Measure your
intelligence online.

Google

Assessments

  • All IQ Tests
  • Basic IQ Test
  • Full IQ Test
  • Custom IQ Test
  • Free IQ Test

Our Socials

  • X
  • YouTube
  • Facebook
  • LinkedIn

Other IQ Tests

  • WAIS-V
  • SB-5
  • Raven's 2
  • RIAS-2
  • CogAT 9
  • WISC-V

Community

  • Join Subreddit
  • Join Discord

Other Pages

  • Test Manual
  • Administer IQ Tests
  • About Us
  • Articles
  • Data
  • FAQ

Research

  • What do polygenic scores really predict?
  • Working speed and ability on the RIOT

Intelligence Journals & Organizations

  • Human Intelligence Research & Education (HIRE) Foundation
  • International Society for Intelligence Research (ISIR)
  • Intelligence & Cognitive Abilities Journal (ICA)
  • Intelligence Journal
  • Mensa Foundation

Contact

  • Email
  • Support

News & Press

  • International Society for Intelligence Research
  • American Thinker
  • Mensa Foundation (1/2)
  • Mensa Northern New Jersey
  • The University of Western Australia
  • Prolific
  • Quillette
  • Brainz

Our Articles

  • Gmatclub (1/2)
  • ApolloTechnical
  • LessWrong
  • Psychreg
  • Study in Switzerland
  • SuccessConsciousness
  • Creative Organizational Design (1/2)
  • ABNewsWire
  • Vanderbilt University

Our Articles

  • The Globe and Mail
  • Barchart
  • Journal
  • Mensa Foundation (2/2)
  • Psychologs
  • Creative Organizational Design (2/2)
  • AZBigMedia
  • Thoughts on Life and Love
  • Anxiety and Depression Association of America

Our Articles

  • Before It's News
  • Siglo XXI
  • TechBullion
  • Medium
  • Gmatclub (2/2)
  • MSN
  • National Review
  • Minding the Campus
  • Launching Next

Our Articles

  • Comparing Cronbach’s Alpha and McDonald’s Omega Reliability
  • Breaking the Intelligence & IQ Taboo
  • What is the Flynn Effect?
  • A Comprehensive History of IQ Tests
  • The 15 Subtests of the RIOT

Our Articles

  • How to Take an IQ Test
  • How to Calculate IQ
  • What is the RIOT IQ Test?
  • The Pro-Human Aspects of Intelligence Research
  • What is an IQ Test? A Beginner's Guide.

Our Articles

  • 5 Best IQ Tests in 2025
  • Cognitive Profiles on the RIOT IQ Test Results
  • 6 Cognitive Abilities of the RIOT
  • Are There Any Professional and Real Online IQ Tests?

Our Articles

  • Resources to Learn About IQ and Intelligence
  • Studying IQ Matters
  • The Search for Albert Einstein's IQ
  • Do Non-g Gains from the Flynn Effect Matter?

Riot IQ © 2026

  • Terms of Service
  • Privacy Policy
  • BAA Agreement
  • Test Administrator Terms
  • Terms of Service
  • •Privacy Policy
  • •BAA Agreement
  • •Test Administrator Terms

Table of Contents

  • Two different concepts wearing one name
  • The standing critique of laboratory cognitive tests
  • Executive function measures, where the critique bites hardest
  • Performance-based functional assessments
  • Why the critique is interesting rather than decisive
  • Frequently asked questions
  • What is the difference between ecological validity and external validity?
  • Is a test with low ecological validity useless?
  • Why do researchers criticise the term itself?
  • How do clinicians improve ecological validity in cognitive assessment?
  • References
Sep 26, 2026·Accuracy, Reliability & Criticism

What is ecological validity? The two meanings, and what they mean for IQ tests

Ecological validity is how far findings or test scores generalise to real life. The term has two meanings, and both matter for reading cognitive tests.

Dr. Russell T. WarneChief Scientist
Share
What is ecological validity? The two meanings, and what they mean for IQ tests
Ecological validity is the degree to which findings from a study, or scores from a test, generalise to real-world settings and everyday functioning. The complication is that the phrase carries two distinct meanings. Egon Brunswik coined it in the 1950s as a precise statistic, the correlation between a perceptual cue and the thing in the world that cue signals. The sense almost everyone now uses arrived later and means something looser, roughly "resembles real life." That looseness is itself a documented criticism in the measurement literature. This article separates the two senses, then works through what the ecological validity critique actually establishes about cognitive and IQ testing, and what it does not.


Two different concepts wearing one name

Brunswik's original usage was narrow and quantitative. In his lens model of perception, a "cue" is a proximal signal available to the observer, and its ecological validity is the correlation between that cue and the distal state of the world it stands for. It is a number attached to a cue, not a compliment paid to a study.

The modern usage descends from a separate adaptation. Kihlstrom traces it to Martin Orne, who used the phrase for the generalisation of experimental findings to the world outside the laboratory, and notes that the two meanings have coexisted ever since without most authors distinguishing them.

• The drift is documented, not alleged: Holleman, Hooge, Kemner and Hessels reviewed the term's use and found it applied to at least twelve different referents, including stimuli, tasks, settings, conditions, results, theories, designs and paradigms. They describe it as having been largely detached from its original parentage, and frequently conflated with Brunswik's separate idea of representative design.

• It is rarely defined where it is used: the same review found that researchers who invoke ecological validity seldom say what they mean by it, which makes the term difficult to argue with in either direction.

• It is dimensional rather than binary: Schmuckler's analysis separates the naturalness of the setting, the stimuli and the response demanded, and points out that a study can be realistic on one dimension while being artificial on another.

The practical upshot for a reader of test research is a single question: when a paper claims a measure lacks ecological validity, is it making a claim about a correlation, about the physical realism of the task, or about whether results generalise? Those need different evidence.


The standing critique of laboratory cognitive tests

The substantive version of the critique says that competence displayed in ordinary life is not well captured by tasks performed at a desk under standardised conditions.

The most-cited demonstration is Ceci and Liker's study of racetrack handicappers. Thirty avid patrons were split into fourteen experts and sixteen non-experts by their ability to predict post-time odds, and the expert group turned out to be using a mental model with multiple interactions and nonlinear terms. Mean measured IQ was 99.3 for the experts and 100.8 for the non-experts, so a complex reasoning skill was on display that the IQ score gave no hint of.

The finding is contested rather than settled. Detterman and Spry reanalysed the data and reported that IQ correlated .35 with a measure of handicapping success within the expert group and negative .25 within the novice group, arguing that an unreliable measure of expertise had produced the original null result. Ceci and Liker replied in the same issue. With fourteen and sixteen people per group, neither analysis carries much weight on its own.

A cleaner line of evidence comes from context. Carraher, Carraher and Schliemann studied child street vendors in Recife who handled commercial arithmetic daily, and found that the same children solved problems embedded in their working context more successfully than they solved the equivalent school-style word problems and context-free computations using identical numbers and operations. Performance moved with the setting while the arithmetic stayed the same.


Executive function measures, where the critique bites hardest

Neuropsychology has taken this question more seriously than any other field, largely because clinicians are asked to predict whether a patient can manage a household rather than whether they can sort cards.

Burgess and colleagues made the historical argument directly: the standard executive tests, including the Wisconsin Card Sorting Test and the Stroop task, are adaptations of procedures that emerged almost coincidentally from conceptual frameworks far removed from those now in favour, and may therefore not be optimal for the clinical purpose they are put to. Their proposed remedy is a "function-led" development programme, designing tasks backwards from the everyday behaviour of interest, exemplified by the Multiple Errands Test and the Six Element Test.

The empirical picture supports their concern. Toplak, West and Stanovich reviewed twenty studies comparing performance-based executive tests against rating-scale measures of everyday executive behaviour. Of 286 reported correlations, only 68, or 24 percent, reached statistical significance, and the median correlation across all studies was .19. Broken out by instrument, the median was .18 for the Behavior Rating Inventory of Executive Function, .14 for the dysexecutive questionnaire, and .25 for impulsivity ratings. The authors note that these values are probably generous, since non-significant correlations go unreported.

Chaytor and Schmitter-Edgecombe's review of the wider literature reached a more moderate verdict, concluding that many neuropsychological tests show a moderate level of ecological validity for predicting everyday cognitive functioning, with the strongest relationships appearing when the outcome measure matched the cognitive domain the test assessed. Who rated the everyday outcome, a clinician or a family member, moved the results as well.


Performance-based functional assessments

The practical response has been to build measures that sit between a laboratory task and real life. Patterson and colleagues developed the UCSD Performance-Based Skills Assessment, in which a person carries out standardised role-plays across household chores, communication, finance, transportation and recreational planning, scored on what they actually do.

These instruments show how much the choice of criterion matters. Twamley and colleagues, studying 111 older patients with psychosis, found the assessment correlated .61 with a global neuropsychological screening measure. More tellingly, a fuller neuropsychological battery correlated .54 with patients' real-world level of independent living, while the brief screening measure managed only .30 against the same outcome.

The ceiling is still low. In the first phase of the VALERO study, Harvey and colleagues examined 198 adults with schizophrenia and found that the best-performing real-world rating scale accounted for 24 percent of the variance in the underlying ability trait measured by cognitive and functional-capacity tests. Three-quarters of what determines how someone actually lives sat outside what the ability measures captured, which is a fair summary of the ecological validity problem in one statistic.


Why the critique is interesting rather than decisive

If laboratory cognitive tasks were as detached from ordinary life as the strong version of the critique implies, they should not predict much. They do. Deary, Strand, Smith and Fernandes followed more than 70,000 English children and found a correlation of .81 between latent cognitive ability at age 11 and latent educational achievement at 16. Roth and colleagues, pooling 240 samples and 105,185 participants, reported a corrected correlation of .54 between intelligence tests and school grades. Strenze's longitudinal meta-analysis found corrected correlations of .56 with educational attainment and .45 with occupational attainment.

That is the genuinely interesting part. Abstract matrix puzzles, with no surface resemblance to a classroom or an office, forecast classroom and office outcomes reasonably well, which suggests surface realism is a poor guide to what a measure can tell you. The criterion validity evidence and the ecological validity critique are answering different questions, and a test can score well on the first while leaving the second open.

Where the critique keeps its force is in the inference from a score to a person's everyday competence in a specific domain. A prediction that holds across thousands of people constrains what can be said about one of them, and the executive function literature shows how weak the test-to-daily-life link can get when the outcome is a single individual's behaviour at home. Our overview of whether IQ tests are valid covers the broader question, and construct validity covers how the whole family of validity evidence fits together.

A well-built cognitive test is honest about this boundary. The Reasoning and Intelligence Online Test, developed by RIOT IQ with Dr. Russell T. Warne, is a professionally developed IQ test that reports scores with confidence intervals across several cognitive domains for adults 18 and over, and it is not a functional assessment of how anyone manages daily life.


Frequently asked questions

What is the difference between ecological validity and external validity?

External validity is the general question of whether results hold beyond the specific sample, setting and materials studied. Ecological validity in the modern sense is the narrower part of that concerning everyday, real-world settings. In Brunswik's original sense it is a correlation and not a form of external validity at all.

Is a test with low ecological validity useless?

No. Ecological validity is about resemblance to and generalisation toward everyday settings, while usefulness depends on whether scores predict the outcome a decision turns on. Abstract reasoning tasks predict educational and occupational outcomes despite looking nothing like them.

Why do researchers criticise the term itself?

Because it is used inconsistently. One review found the phrase attached to at least twelve different things and rarely defined by the authors using it, which allows the same words to carry a precise statistical claim and a vague impression of realism.

How do clinicians improve ecological validity in cognitive assessment?

Mainly by adding performance-based functional measures and structured informant reports alongside the standard battery, and by matching the outcome measure to the cognitive domain being tested, which is where the published relationships are strongest.


References

1. Burgess, P. W., Alderman, N., Forbes, C., Costello, A., Coates, L. M.-A., Dawson, D. R., Anderson, N. D., Gilbert, S. J., Dumontheil, I., & Channon, S. (2006). The case for the development and use of "ecologically valid" measures of executive function in experimental and clinical neuropsychology. Journal of the International Neuropsychological Society, 12(2), 194-209. doi.org

2. Carraher, T. N., Carraher, D. W., & Schliemann, A. D. (1985). Mathematics in the streets and in schools. British Journal of Developmental Psychology, 3(1), 21-29. doi.org

3. Ceci, S. J., & Liker, J. K. (1986). A day at the races: A study of IQ, expertise, and cognitive complexity. Journal of Experimental Psychology: General, 115(3), 255-266. doi.org

4. Ceci, S. J., & Liker, J. K. (1988). Stalking the IQ-expertise relation: When the critics go fishing. Journal of Experimental Psychology: General, 117(1), 96-100. doi.org

5. Chaytor, N., & Schmitter-Edgecombe, M. (2003). The ecological validity of neuropsychological tests: A review of the literature on everyday cognitive skills. Neuropsychology Review, 13(4), 181-197. doi.org

6. Deary, I. J., Strand, S., Smith, P., & Fernandes, C. (2007). Intelligence and educational achievement. Intelligence, 35(1), 13-21. doi.org

7. Detterman, D. K., & Spry, K. M. (1988). Is it smart to play the horses? Comment on "A day at the races: A study of IQ, expertise, and cognitive complexity". Journal of Experimental Psychology: General, 117(1), 91-95. doi.org

8. Harvey, P. D., Raykov, T., Twamley, E. W., Vella, L., Heaton, R. K., & Patterson, T. L. (2011). Validating the measurement of real-world functional outcomes: Phase I results of the VALERO study. American Journal of Psychiatry, 168(11), 1195-1201. pmc.ncbi.nlm.nih.gov

9. Holleman, G. A., Hooge, I. T. C., Kemner, C., & Hessels, R. S. (2020). The "real-world approach" and its problems: A critique of the term ecological validity. Frontiers in Psychology, 11, 721. doi.org

10. Kihlstrom, J. F. (2021). Ecological validity and "ecological validity". Perspectives on Psychological Science, 16(2), 466-471. doi.org

11. Patterson, T. L., Goldman, S., McKibbin, C. L., Hughs, T., & Jeste, D. V. (2001). UCSD Performance-Based Skills Assessment: Development of a new measure of everyday functioning for severely mentally ill adults. Schizophrenia Bulletin, 27(2), 235-245. doi.org

12. Roth, B., Becker, N., Romeyke, S., Schäfer, S., Domnick, F., & Spinath, F. M. (2015). Intelligence and school grades: A meta-analysis. Intelligence, 53, 118-137. doi.org

13. Schmuckler, M. A. (2001). What is ecological validity? A dimensional analysis. Infancy, 2(4), 419-436. doi.org

14. Strenze, T. (2007). Intelligence and socioeconomic success: A meta-analytic review of longitudinal research. Intelligence, 35(5), 401-426. doi.org

15. Toplak, M. E., West, R. F., & Stanovich, K. E. (2013). Practitioner review: Do performance-based measures and ratings of executive function assess the same construct? Journal of Child Psychology and Psychiatry, 54(2), 131-143. doi.org

16. Twamley, E. W., Doshi, R. R., Nayak, G. V., Palmer, B. W., Golshan, S., Heaton, R. K., Patterson, T. L., & Jeste, D. V. (2002). Generalized cognitive impairments, ability to perform everyday tasks, and level of independence in community living situations of older patients with psychosis. American Journal of Psychiatry, 159(12), 2013-2020. doi.org

Hero image: fish vendor at the Los Cocos fish market, by Wilfredor, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.

Take our professional IQ test

Want to know your IQ? Try the first ever professional online IQ test.

Try our IQ test
Author
Dr. Russell T. WarneChief Scientist

Contact

Table of Contents

  • Two different concepts wearing one name
  • The standing critique of laboratory cognitive tests
  • Executive function measures, where the critique bites hardest
  • Performance-based functional assessments
  • Why the critique is interesting rather than decisive
  • Frequently asked questions
  • What is the difference between ecological validity and external validity?
  • Is a test with low ecological validity useless?
  • Why do researchers criticise the term itself?
  • How do clinicians improve ecological validity in cognitive assessment?
  • References
Article Categories
All ArticlesUnderstanding IQ ScoresTaking an IQ TestRIOT-Specific InformationGeneral IQ & IntelligenceAdvanced Topics & ResearchIQ Scores & InterpretationMensa & High-IQ SocietiesOnline IQ Tests IQ Test Basics & FundamentalsAverage IQ & DemographicsFamous People & IQHistory & Origins Of IQ TestingAccuracy, Reliability & CriticismSpecial Population & Related ConditionsImproving IQ / PreparationSpecific IQ Tests & FormatsIQ Testing for HR & RecruitmentSkills Assessment
Related Articles
Standard error of measurement: what it is and how it is calculatedWhat is internal consistency? Reliability from a single test sittingWhat is test-retest reliability? Score stability across two testingsWhat is face validity? Why a test that looks right can still be worthlessWhat is ecological validity? The two meanings, and what they mean for IQ testsWhat is criterion validity? Concurrent and predictive evidence explainedWhat is content validity? Sampling the domain a test claims to coverWhat is inter-rater reliability? Agreement between two scorersWhat is construct validity? How we know an IQ test measures intelligenceThe Mozart Effect: Does Listening to Music Raise IQ?How to Spot a Fake Online IQ TestWhat Is the Average IQ in the UK?Is Gen Z IQ Dropping?How to Tell If an Online IQ Test Is LegitimateWhy a Norm Sample Matters for IQ Test AccuracyCan Amateur IQ Tests Give Accurate Scores?How Accurate Are IQ Tests?What Is an IQ Confidence Interval? Why Scores Are Ranges7 Common Myths About IQ Tests DebunkedWhat Makes an IQ Test Scientifically Valid?Are IQ Tests Racist?Are IQ Tests Biased?How Reliable are IQ Tests?The IQ of Artificial IntelligenceChatGPT’s IQWhy Are IQ Tests Flawed?Are Online IQ Tests Legit?Is There an Official IQ Test?What is the Most Accurate IQ Test?Are IQ Tests Valid?Are IQ Tests Reliable?Are IQ Tests Good Measures of Intelligence?Are IQ Tests Accurate?
Take our IQ tests

Basic IQ Test

5 subtests + 5 cognitive abilities

Take the IQ test

Features

  • ~13 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±5.6 IQ margin of error

5/15 Subtests

Learn more
Vocabulary
Matrix Reasoning
SToVeS
Visual Reversal
Symbol Search
Most comprehensive

Full IQ Test

15 subtests + all cognitive abilities

Take the IQ test

Features

  • ~52 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±3.7 IQ margin of error

15/15 Subtests

Learn more
Vocabulary, Information, Analogies
Matrix Reasoning, Visual Puzzles, Figure Weights
Object Rotation, SToVeS, Spatial Orientation
Computation Span, Exposure Memory, Visual Reversal
Symbol Search, Abstract Matching
Simple Reaction Time, Choice Reaction Time
Compare all tests