🇺🇸The official website of Riot IQ
Log in
  • Home
  • About

Measure your
intelligence online.

Google

Assessments

  • All IQ Tests
  • Basic IQ Test
  • Full IQ Test
  • Custom IQ Test
  • Free IQ Test

Our Socials

  • X
  • YouTube
  • Facebook
  • LinkedIn

Other IQ Tests

  • WAIS-V
  • SB-5
  • Raven's 2
  • RIAS-2
  • CogAT 9
  • WISC-V

Community

  • Join Subreddit
  • Join Discord

Other Pages

  • Test Manual
  • Administer IQ Tests
  • About Us
  • Articles
  • Data
  • FAQ

Research

  • What do polygenic scores really predict?
  • Working speed and ability on the RIOT

Intelligence Journals & Organizations

  • Human Intelligence Research & Education (HIRE) Foundation
  • International Society for Intelligence Research (ISIR)
  • Intelligence & Cognitive Abilities Journal (ICA)
  • Intelligence Journal
  • Mensa Foundation

Contact

  • Email
  • Support

News & Press

  • International Society for Intelligence Research
  • American Thinker
  • Mensa Foundation (1/2)
  • Mensa Northern New Jersey
  • The University of Western Australia
  • Prolific
  • Quillette
  • Brainz

Our Articles

  • Gmatclub (1/2)
  • ApolloTechnical
  • LessWrong
  • Psychreg
  • Study in Switzerland
  • SuccessConsciousness
  • Creative Organizational Design (1/2)
  • ABNewsWire
  • Vanderbilt University

Our Articles

  • The Globe and Mail
  • Barchart
  • Journal
  • Mensa Foundation (2/2)
  • Psychologs
  • Creative Organizational Design (2/2)
  • AZBigMedia
  • Thoughts on Life and Love
  • Anxiety and Depression Association of America

Our Articles

  • Before It's News
  • Siglo XXI
  • TechBullion
  • Medium
  • Gmatclub (2/2)
  • MSN
  • National Review
  • Minding the Campus
  • Launching Next

Our Articles

  • Comparing Cronbach’s Alpha and McDonald’s Omega Reliability
  • Breaking the Intelligence & IQ Taboo
  • What is the Flynn Effect?
  • A Comprehensive History of IQ Tests
  • The 15 Subtests of the RIOT

Our Articles

  • How to Take an IQ Test
  • How to Calculate IQ
  • What is the RIOT IQ Test?
  • The Pro-Human Aspects of Intelligence Research
  • What is an IQ Test? A Beginner's Guide.

Our Articles

  • 5 Best IQ Tests in 2025
  • Cognitive Profiles on the RIOT IQ Test Results
  • 6 Cognitive Abilities of the RIOT
  • Are There Any Professional and Real Online IQ Tests?

Our Articles

  • Resources to Learn About IQ and Intelligence
  • Studying IQ Matters
  • The Search for Albert Einstein's IQ
  • Do Non-g Gains from the Flynn Effect Matter?

Riot IQ © 2026

  • Terms of Service
  • Privacy Policy
  • BAA Agreement
  • Test Administrator Terms
  • Terms of Service
  • •Privacy Policy
  • •BAA Agreement
  • •Test Administrator Terms

Table of Contents

  • Validity is one thing, and the evidence arrives from five directions
  • Convergent and discriminant evidence
  • Internal structure: g, CHC, and an argument still running
  • Evidence from outside the test
  • What weak construct validity actually looks like
  • Frequently asked questions
  • What is construct validity in simple terms?
  • How is construct validity different from content validity?
  • What are convergent and discriminant validity?
  • Do IQ tests have good construct validity?
  • Can a test be reliable but lack construct validity?
  • References
Sep 26, 2026·Accuracy, Reliability & Criticism

What is construct validity? How we know an IQ test measures intelligence

Construct validity is the evidence that a test measures the theoretical attribute it claims to measure. Here is how that case is built for real IQ tests.

Dr. Russell T. WarneChief Scientist
Share
What is construct validity? How we know an IQ test measures intelligence
Construct validity is the degree to which accumulated evidence supports the claim that a test measures the theoretical attribute it says it measures. For a cognitive test, the claim under examination is that a score reflects reasoning ability rather than schooling, motivation, or practice with a particular item format, and the evidence for it has to come from several independent directions at once. This page is the hub for the validity family: it covers the modern unified view of validity, the convergent and discriminant evidence at its centre, and how that evidence looks for the Wechsler scales, the Stanford-Binet and Raven's matrices.


Validity is one thing, and the evidence arrives from five directions

The term comes from Cronbach and Meehl (1955), who described a "construct" as a postulated attribute of people assumed to be reflected in test performance, and argued that such an attribute is validated by placing it inside a network of predicted relationships and then checking whether the predictions hold. Messick (1995) extended that into a unified account in which content, criterion and construct evidence are not separate species of validity but strands of a single argument about score meaning.

That unified view is now the governing position of the field. The AERA, APA and NCME Standards for Educational and Psychological Testing (2014) define validity as "the degree to which evidence and theory support the interpretations of test scores for proposed uses of tests" (p. 11), and add a warning that most casual writing about IQ tests ignores: "It is the interpretations of test scores for proposed uses that are evaluated, not the test itself... It is incorrect to use the unqualified phrase 'the validity of the test.'"

The Standards then set out five sources of validity evidence, and state plainly that these "do not represent distinct types of validity. Validity is a unitary concept."

• Evidence based on test content: whether the items sample the domain the test claims to cover, judged by test specifications and expert review. Covered in full on our page on content validity.

• Evidence based on response processes: whether test takers are actually doing the mental work the test assumes. If a matrix-reasoning item can be solved by elimination rather than by inference, the response process does not match the construct.

• Evidence based on internal structure: whether the correlations among items and subtests match the theory the test is built on.

• Evidence based on relations to other variables: correlations with other tests of the same construct, with tests of different constructs, and with external outcomes. The outcome half of this is treated separately as criterion validity.

• Evidence for validity and consequences of testing: whether observed consequences of using the test trace back to a flaw in score meaning rather than to social policy.

One consequence of the unified view is that older labels have lost their technical standing. The Standards say directly that their treatment "does not follow historical nomenclature (i.e., the use of the terms content validity or predictive validity)." Two familiar labels have fared worse still: face validity appears nowhere in the 2014 Standards, and ecological validity is a borrowing from experimental design rather than a measurement term.


Convergent and discriminant evidence

Campbell and Fiske (1959) gave the field its sharpest tool here. Their "multitrait-multimethod matrix" requires measuring several traits by several methods and then reading the pattern of correlations. A test earns convergent evidence when it agrees with other measures of the same trait taken by different methods, and discriminant evidence when it agrees less with measures of different traits, including measures that share its own method. Agreement alone is not enough, because two tests can agree simply by sharing a format.

The Standards use the same logic: relationships with measures of the same construct "provide convergent evidence, whereas relationships between test scores and measures purportedly of different constructs provide discriminant evidence."

Pearson's research summary for the WISC-V reports what this looks like in cognitive assessment. In a counterbalanced study of 89 children aged 6 to 16, corrected correlations between the WISC-V Full Scale IQ and the Kaufman Assessment Battery for Children, Second Edition were .77 with its Fluid-Crystallized Index and .81 with its Mental Processing Index. Two batteries built by different authors from different theoretical starting points landed on close to the same rank ordering of children.

The discriminant half shows up in the subscore pattern, where corresponding composites correlated from .50 to .74, with the highest value between the WISC-V Verbal Comprehension Index and the KABC-II crystallised-knowledge composite. A cleaner illustration comes from the WISC-V Integrated study, in which the Multiple Choice Verbal Comprehension Index correlated most highly with the Verbal Comprehension Index at .69, then with quantitative reasoning at .61, auditory working memory at .53 and fluid reasoning at .52. The ordering is the prediction the test's structure implies, and it held.


Internal structure: g, CHC, and an argument still running

The oldest construct-validity evidence in intelligence research is factor-analytic. Carroll's 1993 survey of several hundred factor-analytic datasets produced the three-stratum model, later merged with the Cattell-Horn framework into Cattell-Horn-Carroll theory, which now supplies the architecture for the Woodcock-Johnson, the Stanford-Binet 5 and, less explicitly, the Wechsler scales (McGrew, 2009, 2023). Our explainer on the g factor covers the general factor itself.

The strongest single piece of evidence that these batteries converge on one attribute comes from Johnson, Bouchard, Krueger, McGue and Gottesman (2004). They gave three different mental ability batteries to 436 adults, extracted a general factor from each, and found the three general factors were correlated .99, .99 and 1.00. Whatever the batteries disagreed about, they did not disagree about what sat at the top. A later replication across five batteries reached the same conclusion (Johnson, te Nijenhuis, & Bouchard, 2008).

Below the general factor the picture is genuinely contested. Canivez, Watkins and Dombrowski (2016) reanalysed the full WISC-V standardisation sample and found support for four first-order factors rather than the publisher's five, with the hierarchical general factor accounting for most of the common variance; they concluded that interpretation "should be primarily, if not exclusively," at the general level. Gignac (2015) found across several large samples that Raven's Progressive Matrices shares roughly 50% of its variance with the general factor, about 10% with a fluid-reasoning group factor, and carries around 25% test-specific reliable variance, which undercuts the common description of Raven's as a pure measure of general ability.


Evidence from outside the test

Relations to variables that are not tests carry a lot of weight, because they are hard to manufacture by test design.

• Developmental: cognitive abilities should follow the age trajectories theory predicts. Salthouse (2009), analysing large cross-sectional and longitudinal samples, found that reasoning, memory and speed measures begin declining in healthy adults from the twenties and thirties while knowledge-based measures continue rising, a dissociation that fluid and crystallised accounts predict and a single undifferentiated "test-taking skill" account does not.

• Biological: Pietschnig, Penke, Wicherts, Zeiler and Voracek (2015) meta-analysed 88 studies covering 148 samples and more than 8,000 people, and found brain volume correlated .24 with IQ, generalising across age, sex and IQ domain. That is a modest association, and the authors show the literature had overstated it, but it is a real relationship with an external biological variable.

• Criterion: correlations with schooling, training and job performance are covered on the criterion validity page rather than repeated here.


What weak construct validity actually looks like

The Standards name the two ways a test can fail. "Construct underrepresentation" is a test that measures less than it claims, so a battery of nothing but matrix puzzles cannot support a general intelligence interpretation. "Construct-irrelevant variance" is a test that measures more than it claims, so reading load, time pressure or item familiarity push scores around for reasons that have nothing to do with the attribute. Both are matters of degree, and both are why a responsible publisher reports factor analyses, correlations with rival batteries, and the limits of its own norms.

That is also the sharpest practical test a reader can apply. An unnormed quiz on a website has no published structure, no convergent study and no norm sample, so there is no construct-validity argument to inspect. Our guide to what makes an IQ test scientifically valid works through the consumer version of that check.

If you want a score from an instrument whose structure and norms are documented, the Reasoning and Intelligence Online Test from RIOT IQ is an online IQ test built by psychometricians for adults aged 18 and over. It reports six cognitive indices on the familiar mean-100, standard-deviation-15 scale. It is not a substitute for an individually administered clinical evaluation.


Frequently asked questions

What is construct validity in simple terms?

It is the evidence that a test measures the invisible attribute it claims to measure. Because the attribute cannot be observed directly, the case is built from many partial pieces: internal structure, agreement with other tests, disagreement with tests of other things, and links to outcomes.

How is construct validity different from content validity?

Content evidence asks whether the items sample the right domain, judged mainly by expert review of the items themselves. Construct validity is the whole argument that the scores mean what they are said to mean, and content evidence is one input to it.

What are convergent and discriminant validity?

Convergent evidence is high agreement with other measures of the same construct; discriminant evidence is lower agreement with measures of different constructs. Campbell and Fiske (1959) argued that neither is informative alone, since two tests can agree just by sharing a method.

Do IQ tests have good construct validity?

For the major individually administered batteries the evidence is unusually strong by the standards of psychological measurement, with cross-battery general factor correlations near 1.00. The evidence for interpreting the narrower index scores is weaker and actively debated.

Can a test be reliable but lack construct validity?

Yes. Reliability describes consistency, not meaning. A test can produce nearly identical scores every time and still be measuring reading speed or test familiarity rather than reasoning.


References

1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net

2. Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302. doi.org

3. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. doi.org

4. Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105. doi.org

5. Pearson. (2018). WISC-V efficacy research report. Pearson. pearson.com

6. Johnson, W., Bouchard, T. J., Krueger, R. F., McGue, M., & Gottesman, I. I. (2004). Just one g: Consistent results from three test batteries. Intelligence, 32(1), 95-107. doi.org

7. Johnson, W., te Nijenhuis, J., & Bouchard, T. J. (2008). Still just 1 g: Consistent results from five test batteries. Intelligence, 36(1), 81-95. doi.org

8. McGrew, K. S. (2009). CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1), 1-10. doi.org

9. McGrew, K. S. (2023). Carroll's three-stratum cognitive ability theory at 30 years. Journal of Intelligence, 11(2), 32. doi.org

10. Canivez, G. L., Watkins, M. W., & Dombrowski, S. C. (2016). Factor structure of the Wechsler Intelligence Scale for Children-Fifth Edition: Exploratory factor analyses with the 16 primary and secondary subtests. Psychological Assessment, 28(8), 975-986. doi.org

11. Gignac, G. E. (2015). Raven's is not a pure measure of general intelligence: Implications for g factor theory and the brief measurement of g. Intelligence, 52, 71-79. doi.org

12. Salthouse, T. A. (2009). When does age-related cognitive decline begin? Neurobiology of Aging, 30(4), 507-514. doi.org

13. Pietschnig, J., Penke, L., Wicherts, J. M., Zeiler, M., & Voracek, M. (2015). Meta-analysis of associations between human brain volume and intelligence differences: How strong are they and what do they mean? Neuroscience and Biobehavioral Reviews, 57, 411-432. doi.org

14. Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The concept of validity. Psychological Review, 111(4), 1061-1071. doi.org

Hero image: archery target with a grouped round, by Hemant23071999, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.

Take our professional IQ test

Want to know your IQ? Try the first ever professional online IQ test.

Try our IQ test
Author
Dr. Russell T. WarneChief Scientist

Contact

Table of Contents

  • Validity is one thing, and the evidence arrives from five directions
  • Convergent and discriminant evidence
  • Internal structure: g, CHC, and an argument still running
  • Evidence from outside the test
  • What weak construct validity actually looks like
  • Frequently asked questions
  • What is construct validity in simple terms?
  • How is construct validity different from content validity?
  • What are convergent and discriminant validity?
  • Do IQ tests have good construct validity?
  • Can a test be reliable but lack construct validity?
  • References
Article Categories
All ArticlesUnderstanding IQ ScoresTaking an IQ TestRIOT-Specific InformationGeneral IQ & IntelligenceAdvanced Topics & ResearchIQ Scores & InterpretationMensa & High-IQ SocietiesOnline IQ Tests IQ Test Basics & FundamentalsAverage IQ & DemographicsFamous People & IQHistory & Origins Of IQ TestingAccuracy, Reliability & CriticismSpecial Population & Related ConditionsImproving IQ / PreparationSpecific IQ Tests & FormatsIQ Testing for HR & RecruitmentSkills Assessment
Related Articles
Standard error of measurement: what it is and how it is calculatedWhat is internal consistency? Reliability from a single test sittingWhat is test-retest reliability? Score stability across two testingsWhat is face validity? Why a test that looks right can still be worthlessWhat is ecological validity? The two meanings, and what they mean for IQ testsWhat is criterion validity? Concurrent and predictive evidence explainedWhat is content validity? Sampling the domain a test claims to coverWhat is inter-rater reliability? Agreement between two scorersWhat is construct validity? How we know an IQ test measures intelligenceThe Mozart Effect: Does Listening to Music Raise IQ?How to Spot a Fake Online IQ TestWhat Is the Average IQ in the UK?Is Gen Z IQ Dropping?How to Tell If an Online IQ Test Is LegitimateWhy a Norm Sample Matters for IQ Test AccuracyCan Amateur IQ Tests Give Accurate Scores?How Accurate Are IQ Tests?What Is an IQ Confidence Interval? Why Scores Are Ranges7 Common Myths About IQ Tests DebunkedWhat Makes an IQ Test Scientifically Valid?Are IQ Tests Racist?Are IQ Tests Biased?How Reliable are IQ Tests?The IQ of Artificial IntelligenceChatGPT’s IQWhy Are IQ Tests Flawed?Are Online IQ Tests Legit?Is There an Official IQ Test?What is the Most Accurate IQ Test?Are IQ Tests Valid?Are IQ Tests Reliable?Are IQ Tests Good Measures of Intelligence?Are IQ Tests Accurate?
Take our IQ tests

Basic IQ Test

5 subtests + 5 cognitive abilities

Take the IQ test

Features

  • ~13 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±5.6 IQ margin of error

5/15 Subtests

Learn more
Vocabulary
Matrix Reasoning
SToVeS
Visual Reversal
Symbol Search
Most comprehensive

Full IQ Test

15 subtests + all cognitive abilities

Take the IQ test

Features

  • ~52 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±3.7 IQ margin of error

15/15 Subtests

Learn more
Vocabulary, Information, Analogies
Matrix Reasoning, Visual Puzzles, Figure Weights
Object Rotation, SToVeS, Spatial Orientation
Computation Span, Exposure Memory, Visual Reversal
Symbol Search, Abstract Matching
Simple Reaction Time, Choice Reaction Time
Compare all tests