Test norms: what they are and how a norming study builds them
Test norms are the score distributions collected from a reference sample, plus the tables that convert a raw score count into an age-relative IQ score.
Dr. Russell T. WarneChief Scientist
Share
"Test norms" are the score distributions obtained from a carefully selected reference sample, together with the tables that convert a raw count of correct answers into a score expressed relative to that sample. Every IQ point is a lookup in such a table. The number 100 does not describe a quantity of reasoning; it describes the median performance of the people who sat the test during its norming study, and every other score is positioned against them.
This page is about the sample and the table. It covers what a norming study actually does, how stratification to census targets works, how age blocking is decided, and why norms decay and have to be rebuilt. For what a reference group is and how to judge whether a given one applies to you, see our explainer on the norm group in IQ testing.
What a norming study produces
A norming study is a data collection exercise whose only purpose is to describe a population's performance well enough to rank future examinees against it. The Standards for Educational and Psychological Testing set out what the resulting documentation must contain: "precise specification of the population that was sampled, sampling procedures and participation rates, any weighting of the sample, the dates of testing, and descriptive statistics," plus an indication of the precision of the norms themselves. The sample itself must be "technically sound, representative" and "of sufficient size."
Raw scores are then mapped onto a common metric, usually a mean of 10 with a standard deviation of 3 for subtests and a mean of 100 with a standard deviation of 15 for composites, and percentile ranks are attached to each point on it. Reliability and error statistics come out of the same data, which is why the standard error printed in a manual is a property of the norming sample rather than of the person tested.
The samples behind the tests clinicians actually use
Published normative sample sizes vary by an order of magnitude, and the spread is mostly explained by how many tests a battery contains and how wide an age range it covers.
• Wechsler Intelligence Scale for Children, Fifth Edition: 2,200 children in 11 age groups, collected from April 2013 to March 2014, with each age group matched to 2012 United States census figures on race and ethnicity, parent education level and geographic region, and balanced on sex.
• Wechsler Adult Intelligence Scale, Fifth Edition: Published in 2024 for ages 16 to 90, with the sample recruited to represent the English-speaking United States population on education level, race and ethnicity and sex according to 2022 census data. A published re-analysis of the standardisation data reports 1,660 participants across the 11 adult age groups running from 20 to 24 through 85 to 90, collected between February 2023 and January 2024.
• Woodcock-Johnson V: 5,837 individuals, collected from February 2022 through August 2023, against a sampling plan that called for 6,000 examinees across 24 sampling age groups with a target of 250 per group. The obtained average was 245 cases per age group from ages 3 through 79, falling to 190 for the 80-and-over group, which the publisher attributes to the difficulty of recruiting older adults around the pandemic.
• Woodcock-Johnson IV: 7,416 individuals from 46 states and the District of Columbia, ages 2 to 90 and over, collected between December 2009 and January 2012.
• Stanford-Binet Intelligence Scales, Fifth Edition: 4,800 individuals aged 2 to 85 and over, published in 2003 and still the current edition.
Stratification means matching the sample to census targets
Publishers do not recruit at random from the population, because a random sample of a few thousand people will drift away from the national profile on the variables that predict test performance. Instead they set quotas. The Woodcock-Johnson V sampling plan was "stratified to control for census region, sex, ethnicity, race, and parent education level (for children) or examinee education level (for adults)," with the obtained distributions reported against the 2020 census. The fourth edition added community type, distinguishing metropolitan, micropolitan and rural recruitment, and country of birth.
Education level is the variable doing the most work in that list. It stands in for socioeconomic status, it correlates substantially with cognitive test performance, and a sample that over-recruits graduates produces norms that make the general population look below average. The Wechsler manuals use parent education for children and examinee education for adults for that reason.
Quotas are never filled exactly, so the remaining gap is closed arithmetically. The Woodcock-Johnson IV norming study assigned each participant a weight built from partial weights for each sampling variable, so that the final tables were "based on a sample with characteristics proportional to the U.S. population distribution." A reader comparing two manuals should check whether weights were applied and how large they were. The partial weight of 2.2 that the fourth edition needed for examinees born outside the United States marks a badly under-recruited cell.
Age blocking and why the intervals are uneven
Norms are not a single table. They are a stack of tables, one per age band, because the same raw score means something different at 6 and at 16. How finely those bands are cut is a design decision with real consequences.
Developmental rate drives it. The Woodcock-Johnson V used one-year groups from ages 3 through 19, ten-year groups from 20 through 79, and a single twenty-year group for ages 80 and over, and the publisher states plainly that the higher density of examinees between 3 and 19 "reflects the need to collect more concentrated data for ages when the abilities measured by the WJ V undergo the greatest rate of growth." The Wechsler children's scale takes the same view from the other direction, splitting ages 6 to 16 into 11 groups of 200 children each.
Coarse blocking at the top of the range reflects recruitment cost and the slower rate of change rather than a claim that ability is static after 20, and it means adult norms carry a wider effective age window than children's norms do. Interpolation within a band handles the rest, which is why manuals report norms in increments of years and months.
How norms go stale, and what renorming does
Norms have a shelf life because population performance moves. Scores on cognitive tests have risen across the twentieth century, and a meta-analysis of 285 studies since 1951 estimated the gain at 2.31 standard-score points per decade overall, and at 2.93 IQ points per decade for comparisons involving modern Stanford-Binet and Wechsler tests. Our page on the Flynn effect and why IQ scores are rising covers the explanations that have been offered for it.
The consequence for norms is arithmetic. A table built in 2013 and used in 2026 flatters the examinee, because the comparison group is drawn from a cohort that scored lower than today's. The mechanism cuts hard in the other direction when a test is replaced. Kanaya, Scullin and Ceci followed longitudinal records from nine sites and found that students in the borderline and mild range lost an average of 5.6 points when retested on a renormed test, and were more likely to be classified as having intellectual disability than peers retested on the same edition. Nothing about those children changed. The reference sample did.
The Standards place the duty squarely on publishers. Standard 5.11 states that "as long as the test remains in print, it is the test publisher's responsibility to renorm the test with sufficient frequency to permit continued accurate and appropriate score interpretations," and the accompanying commentary assigns test users the complementary duty of avoiding norms that are out of date. Actual cycles run roughly ten to twenty years. The Wechsler children's scale was renormed in 2003 and again in 2014, the adult scale in 2008 and again in 2024, and the Woodcock-Johnson in 2014 and again in 2025. The Stanford-Binet has run on 2003 norms for more than two decades, which is a fact worth knowing before quoting a score from it.
Four questions therefore settle most of what a reader needs. When was the sample collected, rather than when was the test published. Which census year were the quotas set against. Which variables were stratified, and does education appear among them. How many cases sit in the relevant age band. A test that answers all four in public documentation can be evaluated, and one that reports only a headline total cannot. This is why the Reasoning and Intelligence Online Test publishes its normative methodology alongside the instrument, and why an unnormed web quiz cannot produce an IQ score at all. To get a number anchored to a documented reference sample, take an online IQ test built by psychometricians.
Frequently asked questions
How large does a norming sample need to be?
The Standards set no numeric minimum, requiring only a technically sound and representative sample of sufficient size. In practice individually administered cognitive batteries report totals between roughly 2,000 and 7,500, and the figure that matters for an individual score is the number of cases in the relevant age band rather than the headline total.
How often are IQ tests renormed?
Roughly every ten to twenty years for the major batteries. The Wechsler children's scale moved from 2003 norms to 2014 norms, the adult scale from 2008 to 2024, and the Woodcock-Johnson from 2014 to 2025.
Do old norms make a score too high or too low?
Too high, in general. Because population performance has risen, an ageing table compares the examinee with a lower-scoring cohort. The reported effect is roughly 2 to 3 points per decade since the norms were collected.
Are norms the same in every country?
No. A norm table describes the population that was sampled, and most cognitive batteries are normed nationally. Applying United States norms to an examinee elsewhere introduces an unknown amount of error, which is why publishers commission separate standardisation studies for local editions.
What does stratification actually control for?
Typically census region, sex, race and ethnicity, and education level, with community type and country of birth added by some publishers. Parent education is used for children and examinee education for adults, because it is the strongest available proxy for socioeconomic background.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. LaForte, E. M., Dailey, D., & McGrew, K. S. (2025). WJ V technical abstract. Riverside Assessments. info.riversideinsights.com
3. LaForte, E. M., McGrew, K. S., & Schrank, F. A. (2014). WJ IV technical abstract (Woodcock-Johnson IV Assessment Service Bulletin No. 2). Riverside Publishing. info.riversideinsights.com
4. Pearson. (2018). WISC-V efficacy research report. Pearson Education. pearson.com
5. Winter, E. L., Dale, B. A., Maharjan, S., Lando, C. R., Larsen, C. M., Courville, T., & Kaufman, A. S. (2025). Cognitive aging revisited: A cross-sectional analysis of the WAIS-5. Journal of Intelligence, 13(7), 85. pmc.ncbi.nlm.nih.gov
8. Trahan, L. H., Stuebing, K. K., Fletcher, J. M., & Hiscock, M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin, 140(5), 1332-1360. pmc.ncbi.nlm.nih.gov
9. Kanaya, T., Scullin, M. H., & Ceci, S. J. (2003). The Flynn effect and U.S. policies: The impact of rising IQ scores on American society via mental retardation diagnoses. American Psychologist, 58(10), 778-790. pubmed.ncbi.nlm.nih.gov
Hero image: people waiting in line at the Stena Line ferry terminal, Frederikshavn, by W.carter, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.