What is an achievement test? How school skills are measured and scored
An achievement test measures what a person has already learned in reading, writing or math, scored against a norm group of same-age or same-grade peers.
Dr. Russell T. WarneChief Scientist
Share
An achievement test measures what a person has already learned in a defined academic domain, usually reading, written expression or mathematics, and reports how that learning compares with a representative sample of same-age or same-grade peers. It answers a backward-looking question about acquired skill, which is why a reading or arithmetic battery sits alongside a cognitive test in almost every school evaluation instead of replacing it.
This page covers how an achievement test differs from an aptitude or cognitive test, how individually administered batteries differ from the group tests given in classrooms, why grade and age equivalents mislead almost everyone, and what to take from a score report.
What an achievement test measures
The working distinction is one of content. An achievement battery samples skills schooling is supposed to install: decoding printed words, spelling, computing, reasoning about a word problem. A cognitive or aptitude battery samples reasoning tasks chosen to depend as little as possible on any particular curriculum. Beaujean and Parkin put the achievement side carefully in their review of the Wechsler achievement battery: psychologists treat academic achievement as the basic competencies members of a society typically acquire, and build the instruments so the content is not tied to one curriculum. That is what separates a norm-referenced achievement battery from a curriculum-based assessment or a licensing exam.
That is a difference of degree rather than of kind, and the correlations show it. Roth and colleagues published a psychometric meta-analysis of the relationship between standardized intelligence tests and school grades covering 240 independent samples and 105,185 participants. The mean observed correlation weighted by sample size was .44. After correcting for sampling error, unreliability in the predictor and indirect range restriction, the population correlation was .54, with a 95 percent confidence interval from .51 to .57. The figure at the top of this page plots their subgroup results.
Three features of it are worth pausing on. Mathematics and science grades correlated most strongly with tested ability at .49 and sports grades least at .09, so the relationship tracks the cognitive content of a subject rather than school success generally. The association strengthened with age, from .45 in elementary school to .58 in high school. And the authors were explicit that only 31.7 percent of the variance across studies was explained by the artifacts they corrected for, so the overall figure should not be treated as a constant.
Two cautions about carrying those numbers across to achievement testing. The outcome was school grades, which carry teacher judgement and classroom behaviour along with skill, and the coefficients are corrected upward from the observed .44. Our explainer on how an aptitude test differs from an IQ test covers the predictor side of the same relationship.
Individual batteries, group batteries and the state test
The word "achievement test" covers three quite different things, and the differences matter more than the label.
• Individually administered norm-referenced batteries: One examiner, one examinee, standardised administration, scores referenced to a national norm sample. The Wechsler Individual Achievement Test, fourth edition, published by Pearson in 2020 for ages 4 through 50, is a worked example: twenty subtests feeding core composites for reading, written expression and mathematics, a Total Achievement score, processing composites for phonological and orthographic skills, and a Dyslexia Index, normed on 1,832 participants. The Woodcock-Johnson V is the main alternative, and its publisher describes it as the only platform co-norming cognitive ability, academic achievement and oral language on one sample, a claim resting on the publisher rather than independent review. Our page on the Wechsler Individual Achievement Test goes through its structure subtest by subtest.
• Group-administered survey batteries: Machine-scored booklets or online forms given to a whole classroom, such as the Iowa Assessments. These are efficient and useful for screening, and cannot support the observational side of an individual evaluation. Publishers often sell an ability battery alongside the achievement battery for the same grades, which is how many schools end up with both kinds of score on one child.
• State accountability assessments: Required at 34 CFR §200.5(a)(1), which mandates annual reading or language arts and mathematics assessments in each of grades 3 through 8 and at least once in grades 9 through 12, plus science once in each of grades 3 to 5, 6 to 9 and 10 to 12. These report against a state's own proficiency standards rather than a national norm group, so a "proficient" label and a standard score of 100 are not the same kind of statement.
Grade equivalents and age equivalents, and why they mislead
Almost every individually administered battery can print a grade equivalent and an age equivalent, and both are the most reliably misread numbers on a report. A "grade equivalent" is nothing more than the median raw score earned by pupils at that grade level in the norm sample, which is not a statement about what curriculum the examinee has mastered.
Pearson's own guidance gives a concrete illustration from the Peabody Picture Vocabulary Test, third edition. A raw score of 50 yields an age equivalent of 4 years 0 months and 55 yields 4 years 4 months, so five raw points move it by four months. Higher up the same scale, 165 yields 16 years 4 months and 170 yields 18 years 2 months, so the same five points move it by nearly two years. The scale is stretched and squashed in ways the reader cannot see, because skills are acquired quickly at young ages and slowly later.
Several consequences follow directly.
• A stated delay means different things at different ages: A six-month lag for a young child can represent a far larger difference in skill than the same lag for an adolescent, where it may reflect one or two raw points.
• They are not an interval or ratio scale: Grade and age equivalents cannot legitimately be added, subtracted or averaged, and their reliability grows worse for advanced test takers, which is why many instruments stop reporting them past a ceiling. The OWLS Written Expression Scale reports age equivalents only to age 12 and grade equivalents only to grade 6.
• A high equivalent does not mean readiness for that grade: A fifth grader with a grade equivalent of 8.0 has earned the raw score a median eighth grader earns on this test. Nothing in the number says the child has covered eighth-grade work or could do it.
• Fixed cutoffs built from them do not work: Reynolds argued in 1981 that "two years below grade level for age" is a fallacy as a diagnostic criterion for reading disorders. Pearson cites that paper as one reason these scores should not be used for diagnostic or placement decisions.
Standard scores and percentile ranks are the units to read instead, because they reference the distribution of scores rather than a median raw score. Even there some humility is warranted. Beaujean and Parkin examined the technical manual's claim that the Wechsler achievement standard scores are on an equal-interval scale, found the evidence insufficient for every composite, and recommended that psychologists avoid any use depending on equal intervals, including arithmetic comparisons between scores.
What to take from an achievement score
Achievement scores usually arrive on the familiar metric with a mean of 100 and a "standard deviation", a measure of how spread out scores are, of 15. Read the percentile rank and confidence interval before the point score, because a single number is an estimate with error around it.
Federal special education regulation is unusually direct about the limits of one score. 34 CFR §300.304(b)(2) forbids using any single measure as the sole criterion for deciding whether a child has a disability, and §300.304(c)(2) requires assessments tailored to specific areas of educational need rather than instruments designed only to produce a general intelligence quotient. Where achievement testing carries weight is §300.309(a)(1), which lists eight areas in which inadequate achievement can support eligibility: oral expression, listening comprehension, written expression, basic reading skill, reading fluency skills, reading comprehension, mathematics calculation and mathematics problem solving.
On whether the achievement score is supposed to be compared against an IQ score to produce a gap, the answer is no, and it has been no since 2004. We cover that history at length in our article on the rise and fall of the IQ-achievement discrepancy model.
For a parent or an adult reading a report, the useful questions are narrow: which specific skill is low and by how much relative to the confidence interval, whether that skill has already been taught properly, and what the report recommends teaching next. For a sense of what the other half of a school evaluation looks like, you can take an online IQ test built by psychometricians, the Reasoning and Intelligence Online Test, which reports index scores with intervals rather than a bare number.
Frequently asked questions
Is an achievement test an IQ test?
No. An achievement test samples taught academic skills, and an IQ test samples reasoning tasks designed to be less dependent on any curriculum. Their scores are often on the same metric and correlate strongly, which is why the two are easy to confuse on a report.
What counts as a good achievement test score?
On a battery with a mean of 100 and a standard deviation of 15, scores from about 90 to 109 fall in the average range. More informative is the percentile rank with its confidence interval, and the pattern across subtests rather than the composite alone.
How well do achievement and cognitive scores go together?
Strongly, and not interchangeably. The best current meta-analysis of tested ability against school grades reports an observed correlation of .44, corrected to .54, varying by subject from .49 in mathematics and science to .09 in sports.
Why did my child take both a cognitive test and an achievement test?
Because they answer different questions, and federal regulation forbids resting an eligibility decision on a single measure. The achievement battery establishes which academic skills are low; the cognitive measure helps rule out explanations such as intellectual disability.
Are state tests achievement tests?
They measure achievement, for a different purpose. State assessments report standing against state proficiency standards under 34 CFR §200.5 rather than against a national norm sample, so they do not substitute for an individually administered battery in an evaluation.
The takeaway
An achievement test tells you what a person has learned, in units that mean nothing until you know which norm group and which score scale produced them. Read standard scores and percentile ranks with their confidence intervals, treat grade and age equivalents as decorative, and expect the achievement result to correlate with a cognitive result at roughly .5 without duplicating it. The two together describe a learner, and neither alone does.
References
1. Roth, B., Becker, N., Romeyke, S., Schäfer, S., Domnick, F., & Spinath, F. M. (2015). Intelligence and school grades: A meta-analysis. Intelligence, 53, 118-137. doi.org
2. Beaujean, A. A., & Parkin, J. R. (2022). Evaluation of the Wechsler Individual Achievement Test-Fourth Edition as a measurement instrument. Journal of Intelligence, 10(2), 30. pmc.ncbi.nlm.nih.gov
3. Pearson. (n.d.). Interpretation problems of age and grade equivalents. pearsonassessments.com
4. Reynolds, C. R. (1981). The fallacy of "two years below grade level for age" as a diagnostic criterion for reading disorders. Journal of School Psychology, 19(4), 350-358. doi.org
7. U.S. Department of Education. (2016). Assessment administration, 34 CFR §200.5. ecfr.gov
8. U.S. Department of Education. (2006). Evaluation procedures, 34 CFR §300.304. ecfr.gov
9. U.S. Department of Education. (2017). Determining the existence of a specific learning disability, 34 CFR §300.309. ecfr.gov
10. Grigorenko, E. L., Compton, D. L., Fuchs, L. S., Wagner, R. K., Willcutt, E. G., & Fletcher, J. M. (2020). Understanding, educating, and supporting children with specific learning disabilities: 50 years of science and practice. American Psychologist, 75(1), 37-51. doi.org
Figure by Riot IQ. Data from Roth, Becker, Romeyke, Schafer, Domnick and Spinath (2015), Intelligence and school grades: A meta-analysis, Intelligence, 53, 118-137, Table 1.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.