🇺🇸The official website of Riot IQ
Log in
  • Home
  • About

Measure your
intelligence online.

Google

Assessments

  • All IQ Tests
  • Basic IQ Test
  • Full IQ Test
  • Custom IQ Test
  • Free IQ Test

Our Socials

  • X
  • YouTube
  • Facebook
  • LinkedIn

Other IQ Tests

  • WAIS-V
  • SB-5
  • Raven's 2
  • RIAS-2
  • CogAT 9
  • WISC-V

Community

  • Join Subreddit
  • Join Discord

Other Pages

  • Test Manual
  • Administer IQ Tests
  • About Us
  • Articles
  • Data
  • FAQ

Research

  • What do polygenic scores really predict?
  • Working speed and ability on the RIOT

Intelligence Journals & Organizations

  • Human Intelligence Research & Education (HIRE) Foundation
  • International Society for Intelligence Research (ISIR)
  • Intelligence & Cognitive Abilities Journal (ICA)
  • Intelligence Journal
  • Mensa Foundation

Contact

  • Email
  • Support

News & Press

  • International Society for Intelligence Research
  • American Thinker
  • Mensa Foundation (1/2)
  • Mensa Northern New Jersey
  • The University of Western Australia
  • Prolific
  • Quillette
  • Brainz

Our Articles

  • Gmatclub (1/2)
  • ApolloTechnical
  • LessWrong
  • Psychreg
  • Study in Switzerland
  • SuccessConsciousness
  • Creative Organizational Design (1/2)
  • ABNewsWire
  • Vanderbilt University

Our Articles

  • The Globe and Mail
  • Barchart
  • Journal
  • Mensa Foundation (2/2)
  • Psychologs
  • Creative Organizational Design (2/2)
  • AZBigMedia
  • Thoughts on Life and Love
  • Anxiety and Depression Association of America

Our Articles

  • Before It's News
  • Siglo XXI
  • TechBullion
  • Medium
  • Gmatclub (2/2)
  • MSN
  • National Review
  • Minding the Campus
  • Launching Next

Our Articles

  • Comparing Cronbach’s Alpha and McDonald’s Omega Reliability
  • Breaking the Intelligence & IQ Taboo
  • What is the Flynn Effect?
  • A Comprehensive History of IQ Tests
  • The 15 Subtests of the RIOT

Our Articles

  • How to Take an IQ Test
  • How to Calculate IQ
  • What is the RIOT IQ Test?
  • The Pro-Human Aspects of Intelligence Research
  • What is an IQ Test? A Beginner's Guide.

Our Articles

  • 5 Best IQ Tests in 2025
  • Cognitive Profiles on the RIOT IQ Test Results
  • 6 Cognitive Abilities of the RIOT
  • Are There Any Professional and Real Online IQ Tests?

Our Articles

  • Resources to Learn About IQ and Intelligence
  • Studying IQ Matters
  • The Search for Albert Einstein's IQ
  • Do Non-g Gains from the Flynn Effect Matter?

Riot IQ © 2026

  • Terms of Service
  • Privacy Policy
  • BAA Agreement
  • Test Administrator Terms
  • Terms of Service
  • •Privacy Policy
  • •BAA Agreement
  • •Test Administrator Terms

Table of Contents

  • The two questions a single score can answer
  • What changes when a test is built for one purpose or the other
  • The Woodcock-Johnson runs both scales at once
  • Why intelligence tests are norm-referenced nearly all the way down
  • Where the labels get attached to the wrong thing
  • Frequently asked questions
  • Is an IQ test norm-referenced or criterion-referenced?
  • Can the same test give both kinds of score?
  • Does criterion-referenced mean the test has no norm sample?
  • Why can a criterion-referenced test show poor reliability?
  • Which type should a school prefer?
  • References
Sep 27, 2026·IQ Scores & Interpretation

Norm referenced test vs criterion referenced test: how they differ

A norm referenced test reports where a score falls among other people, while a criterion referenced test measures performance against a fixed standard.

Dr. Russell T. WarneChief Scientist
Share
Norm referenced test vs criterion referenced test: how they differ
A "norm referenced test" reports where a score falls relative to other people, and a "criterion referenced test" reports what a person can do measured against a fixed standard of performance. The Standards for Educational and Psychological Testing, the governing document for test quality in the United States, locates the difference in the interpretation rather than in the paper. When scores are norm-referenced, "relative score interpretations are of primary interest," and a score "is ranked within a distribution of scores or compared with the average performance of test takers in a reference population." When interpretations are criterion-referenced, "absolute score interpretations are of primary interest," and "the test score conveys directly a level of competence in some defined criterion domain."

Intelligence tests sit almost entirely on the norm-referenced side. This page compares the two philosophies, explains what changes in how a test is built and evaluated under each, and identifies the places where the labels get attached to the wrong thing.


The two questions a single score can answer

Robert Glaser introduced the term criterion-referenced measurement in American Psychologist in 1963, arguing that ranking students answered the wrong question for instruction. Popham's later account of that moment is that Glaser saw a practical problem: when programmed instruction compressed student scores at the high end of the scale, "the possibility of useful student-to-student comparisons instantly evaporated." Six decades later the vocabulary is standard and the boundary is still misread, because the distinction lives in what the score is anchored to.

• A norm-referenced score is anchored to people: An IQ of 115 means the examinee outperformed roughly 84 percent of the reference sample. It carries no information about which tasks were solved. Our explainer on IQ percentiles covers how that ranking is expressed.

• A criterion-referenced score is anchored to tasks: A score reporting that a reader decodes two-syllable words with 92 percent accuracy states a level of skill. Nobody else's performance enters into the claim.

The Standards are explicit that a test form does not belong to one camp. "Both criterion-referenced and norm-referenced scales may be developed and used with the same test scores if appropriate methods are used to validate each type of interpretation." The document also notes that a scale built for one purpose can drift into supporting the other as experience accumulates, because research eventually reveals what particular score levels imply about real capability.


What changes when a test is built for one purpose or the other

Popham and Husek showed in 1969 that the psychometric properties which make a test good under one framework can make it look bad under the other, and their argument still governs test construction. The pivot is score variance.

• Norm-referenced construction chases spread: Ranking people precisely requires items that separate them, so developers favour items of middling difficulty with strong discrimination and they retain items on which the sample splits. An item everyone passes contributes nothing to a percentile rank.

• Criterion-referenced construction chases domain coverage: The items have to represent a defined body of skill, so an item that almost everyone passes is retained when the skill it samples belongs in the domain. Popham and Husek's blunt summary was that the philosophy underlying criterion-referenced testing rejects the relevance of variance altogether.

• The reliability statistics behave differently: Classical coefficients are ratios of true-score variance to observed-score variance, so they shrink when a group is homogeneous. The Standards acknowledge this directly, noting that traditional norm-referenced reliability coefficients were developed to evaluate the precision of relative standing, and that the range of indices has since grown because score uses expanded into classification. A well-built mastery test given to a class that has mostly mastered the material can post a weak coefficient while classifying almost everyone correctly.

• The scaling differs: Norm-referenced reporting converts raw counts into a common metric whose mean and spread are properties of the reference sample. Criterion-referenced reporting usually needs an equal-interval ability scale instead, so that a stated level of proficiency has a fixed meaning independent of who was sampled.


The Woodcock-Johnson runs both scales at once

The clearest worked example in cognitive assessment is the Woodcock-Johnson V, published in 2025. Its examiner's manual organises score types into four interpretive levels, and two of those levels are precisely the distinction this page is about. Level 3 is labelled Proficiency and marked criterion-referenced. Level 4 is labelled Relative Standing in a Group and marked norm-referenced, holding the standard scores and percentile ranks that most reports quote.

The Level 3 machinery is worth following, because it shows how a criterion-referenced score gets built on a modern ability scale. All Woodcock-Johnson scores rest on the "W scale," an equal-interval metric derived from the Rasch measurement model, centred at 500. For each age and grade group in the norming sample the median W ability is identified, which is the difficulty at which that group would answer half the items correctly. The "relative proficiency index" then moves the reference point 20 W units below that median, to the difficulty level at which 90 percent of average peers succeed, because 90 percent rather than 50 percent is what education normally treats as proficient. The resulting score is a fraction. An index of 60/90 means the examinee is expected to be about 60 percent successful on tasks that typical peers handle with 90 percent success.

Notice what is still in that sentence. The criterion was located using the norming sample. Criterion-referenced does not mean norm-free, and on this battery the proficiency scores and the percentile ranks come out of the same standardisation data. Our overview of the Woodcock-Johnson test covers the batteries themselves.


Why intelligence tests are norm-referenced nearly all the way down

Criterion-referenced interpretation needs a defined criterion domain, which means an agreed body of tasks that the score is about. Reading has one. Fourth-grade mathematics has one. General cognitive ability does not. There is no syllabus of reasoning, no list of matrices a competent adult should complete, and no defensible external standard for how much working memory counts as enough. The only stable anchor available is other people's performance.

That is why the familiar IQ metric is defined by its reference sample rather than by any task. A mean of 100 and a standard deviation of 15 are properties of the norming distribution, assigned to it by construction. The Wechsler Intelligence Scale for Children, Fifth Edition, derives its Full Scale IQ from a standardisation sample of 2,200 children in 11 age groups, matched to 2012 United States census figures on race and ethnicity, parent education level and geographic region. Change that sample and every score changes, which is the whole reason the Standards warn that "the validity of norm-referenced interpretations depends in part on the appropriateness of the reference group to which test scores are compared." Our explainer on the norm group in IQ testing covers how that group is defined.

The achievement halves of the same co-normed batteries can offer criterion-referenced scores because reading and mathematics have definable domains. The cognitive halves mostly cannot.


Where the labels get attached to the wrong thing

Two confusions account for most misuse of these terms, and both are addressed in the Standards.

The first is the assumption that a cut score makes a test criterion-referenced. It does not. The Standards note that cut-score interpretations "may likewise be either criterion referenced or norm referenced," and that when a test is used for selection and the cut is placed so as to admit a prespecified proportion of candidates, "the cut score interpretation is norm referenced." A high-IQ society admitting at the 98th percentile is running a norm-referenced cut. A licensure board whose expert panel judged what a minimally qualified candidate should answer correctly is running a criterion-referenced one. The arithmetic looks identical from outside.

The second is treating criterion-referenced as a synonym for fair, absolute or objective. Neither framework escapes judgement. A norm-referenced score inherits every decision made about who entered the reference sample. A criterion-referenced score inherits every decision made about what belongs in the domain and where proficiency begins. The Standards add a caution that applies to both, that "the likelihood of misclassification will generally be relatively high for persons with scores close to the cut scores."

For a cognitive score the practical question is therefore narrow: which reference sample produced the table, how recent is it, and does it resemble the person being tested. A test that publishes its normative sample, its stratification targets and its collection dates can be checked on those points. That documentation is the reason to look for it before trusting any number, and it is what separates a professionally developed instrument such as the Reasoning and Intelligence Online Test from an unnormed quiz. If you want a score you can actually place, take a full-length online IQ test built against published norms.


Frequently asked questions

Is an IQ test norm-referenced or criterion-referenced?

Norm-referenced. An IQ score states standing relative to a reference sample of the same age, and it carries no direct claim about which tasks the examinee can perform.

Can the same test give both kinds of score?

Yes, and several cognitive batteries do. The Woodcock-Johnson V reports criterion-referenced proficiency scores and norm-referenced standard scores from the same administration, and the Standards explicitly permit both interpretations from one set of scores when each is separately validated.

Does criterion-referenced mean the test has no norm sample?

No. The relative proficiency index on the Woodcock-Johnson locates its 90 percent criterion using the median ability of age or grade peers in the norming sample. The criterion is stated in task terms, and it was found using normative data.

Why can a criterion-referenced test show poor reliability?

Because the usual coefficients depend on how much scores vary. When a group has largely mastered the material, variance is small and the coefficient falls, even though the test may be sorting masters from non-masters accurately. Classification-consistency indices are the appropriate evidence in that case.

Which type should a school prefer?

It depends on the decision. Placement into a service designed for the top few percent of learners is inherently comparative and needs norm-referenced data. Deciding whether a pupil has learned this term's material is a domain question and needs criterion-referenced data.


References

1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net

2. Glaser, R. (1963). Instructional technology and the measurement of learning outcomes: Some questions. American Psychologist, 18(8), 519-521. Reprinted in Educational Measurement: Issues and Practice, 13(4), 6-8. eric.ed.gov

3. Popham, W. J., & Husek, T. R. (1969). Implications of criterion-referenced measurement. Journal of Educational Measurement, 6(1), 1-9. eric.ed.gov

4. Popham, W. J. (2014). Criterion-referenced measurement: Half a century wasted? Educational Leadership, 71(6). ASCD. ascd.org

5. Jaffe, L. E. (2025). Development, interpretation, and application of the W score and the relative proficiency index. Riverside Assessments. info.riversideinsights.com

6. LaForte, E. M., Dailey, D., & McGrew, K. S. (2025). WJ V technical abstract. Riverside Assessments. info.riversideinsights.com

7. Pearson. (2018). WISC-V efficacy research report. Pearson Education. pearson.com

8. Mensa International. Getting your IQ tested: frequently asked questions. Mensa International. mensa.org

Hero image: high jump at the Pan African Games, Lagos, 1973, by Aart Rietveld (ASC Leiden, Rietveld Collection), licensed CC BY-SA 4.0 (creativecommons.org/licenses/by-sa/4.0). Via Wikimedia Commons.

Take our professional IQ test

Want to know your IQ? Try the first ever professional online IQ test.

Try our IQ test
Author
Dr. Russell T. WarneChief Scientist

Contact

Table of Contents

  • The two questions a single score can answer
  • What changes when a test is built for one purpose or the other
  • The Woodcock-Johnson runs both scales at once
  • Why intelligence tests are norm-referenced nearly all the way down
  • Where the labels get attached to the wrong thing
  • Frequently asked questions
  • Is an IQ test norm-referenced or criterion-referenced?
  • Can the same test give both kinds of score?
  • Does criterion-referenced mean the test has no norm sample?
  • Why can a criterion-referenced test show poor reliability?
  • Which type should a school prefer?
  • References
Article Categories
All ArticlesUnderstanding IQ ScoresTaking an IQ TestRIOT-Specific InformationGeneral IQ & IntelligenceAdvanced Topics & ResearchIQ Scores & InterpretationMensa & High-IQ SocietiesOnline IQ Tests IQ Test Basics & FundamentalsAverage IQ & DemographicsFamous People & IQHistory & Origins Of IQ TestingAccuracy, Reliability & CriticismSpecial Population & Related ConditionsImproving IQ / PreparationSpecific IQ Tests & FormatsIQ Testing for HR & RecruitmentSkills Assessment
Related Articles
Test norms: what they are and how a norming study builds themStandard setting: how a test's cut score is actually decidedNorm referenced test vs criterion referenced test: how they differStanine scores explained: the nine-point standard scaleWhat is the General Ability Index (GAI)? The Wechsler composite explainedWhat is the Processing Speed Index (PSI)? Reading the Wechsler scoreWhat is the Perceptual Reasoning Index (PRI)? A legacy Wechsler score explainedWhat is the Visual Spatial Index (VSI)? The Wechsler score explainedWhat is the Verbal Comprehension Index (VCI)? Reading the score on a Wechsler reportWhat is the Fluid Reasoning Index (FRI)? The Wechsler composite explainedWhat is the Cognitive Proficiency Index? The Wechsler CPI explainedT score to percentile: how to read a T scoreScaled score to percentile: converting Wechsler subtest scoresStandard score to percentile: how the conversion worksWhat Does an IQ of 165 Mean?What Does an IQ of 200 Mean?What Does an IQ of 180 Mean?What Does an IQ of 155 Mean?What Does an IQ of 95 Mean?What Does an IQ of 90 Mean?What Does an IQ of 80 Mean?What Does an IQ of 75 Mean?What Does an IQ of 70 Mean?What Does an IQ of 105 Mean?What Does an IQ of 85 Mean?What Does an IQ of 128 Mean?What Does an IQ of 136 Mean?SAT to IQ Conversion: What the Correlation Really ShowsWhat Does an IQ of 110 Mean?What Does an IQ of 115 Mean?What Does an IQ of 150 Mean?What Does an IQ of 160 Mean?IQ Percentiles ExplainedWhat Does an IQ of 125 Mean?What Does an IQ of 135 Mean?What Does an IQ of 140 Mean?What Does an IQ of 130 Mean?The IQ Bell Curve ExplainedThe IQ Scale Explained: Ranges, Chart, and What Scores MeanACT to IQ Conversion Chart ExplainedWhat IQ Do Gifted People Have?What Does IQ Predict?What Is the Difference Between IQ and Emotional Intelligence?IQ vs. EQ: Which One Matters More for Career Success?What Is a Ceiling Effect on an IQ Test?What Is a Norm Group in IQ Testing?What Is Full Scale IQ? How to Read the Overall ScoreIs 145 a High IQ Score?What Is a 'Good' IQ Score?Is a 120 IQ Considered Gifted?What Is the Dumbest IQ Score?Is an IQ of 170 considered a genius?Is My IQ Score Good?What is a Good IQ Score?What is Considered a High IQ?Understanding Your Free IQ Test Results in 2026Why is 100 the Average IQ?Does IQ Change With Age?What Does an IQ of X Mean?Does IQ Matter?What is the IQ Scale/Range?What is the Highest Possible IQ?What is a Genius IQ Score?Is a 97 IQ considered dumb?
Take our IQ tests

Basic IQ Test

5 subtests + 5 cognitive abilities

Take the IQ test

Features

  • ~13 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±5.6 IQ margin of error

5/15 Subtests

Learn more
Vocabulary
Matrix Reasoning
SToVeS
Visual Reversal
Symbol Search
Most comprehensive

Full IQ Test

15 subtests + all cognitive abilities

Take the IQ test

Features

  • ~52 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±3.7 IQ margin of error

15/15 Subtests

Learn more
Vocabulary, Information, Analogies
Matrix Reasoning, Visual Puzzles, Figure Weights
Object Rotation, SToVeS, Spatial Orientation
Computation Span, Exposure Memory, Visual Reversal
Symbol Search, Abstract Matching
Simple Reaction Time, Choice Reaction Time
Compare all tests