Standard error of measurement: what it is and how it is calculated
The standard error of measurement is the typical size of the random error in a score: SEM = SD x the square root of 1 minus reliability, or 3 IQ points.
Dr. Russell T. WarneChief Scientist
Share
The "standard error of measurement" is the typical size of the random error carried by a single test score, in the units of the score itself. It is calculated from two numbers the publisher already reports: SEM = SD x the square root of (1 minus reliability), where SD is the standard deviation of the score scale and the reliability coefficient comes from the technical manual. On the IQ metric, where the standard deviation is 15 points, a Full Scale IQ with a published reliability of .96 has an SEM of 3.0 points. That single line of arithmetic is where the familiar claim that a good IQ score is precise to within about three points comes from.
This page stays on the input side of score precision: what the SEM is, how the formula works on published coefficients, why it is larger for some index scores than others, and why it varies across the scale. The output side, turning an SEM into a reported range around an obtained score, belongs to our explainer on the IQ confidence interval.
Measurement error and the idea of a true score
Every observed score is treated as a stable component plus a random one. The Standards for Educational and Psychological Testing, the governing document for test quality in the United States, defines the error-free component as the person's "true score," the hypothetical average score obtained over an infinite set of replications of the testing procedure. Because a procedure cannot actually be replicated many times on one person, the error is estimated over a population instead, and that average is the SEM. A relatively large SEM indicates relatively low reliability and precision. Two properties of the definition matter for how the number should be read.
• It describes random error only: Fatigue on the day, a lucky guess, a rater's lapse and the particular sample of items a test contains are all random with respect to the person's standing. Systematic error, the kind that shifts every score in the same direction, does not show up in the SEM at all. A test with outdated norms can be perfectly precise and still be biased.
• It is an average, not a personal quantity: The SEM published in a manual is the average error across a reference population, not a statement that one examinee's score is off by that much.
Reliability, the other input, is not a single thing. Internal consistency, temporal stability and scorer agreement each answer a different question about what counts as error, so each produces a different coefficient and a different SEM. Our pages on internal consistency and test-retest reliability cover how each is estimated.
Working the formula on the real IQ metric
The formula multiplies the scale's standard deviation by the square root of the unreliable portion of score variance. On the IQ metric the standard deviation is fixed at 15, so the only thing that moves is the coefficient. These are published values for two heavily researched batteries.
• WISC-V Full Scale IQ, reliability .96 to .97: Miller and McGill's review reports internal consistency estimates across the eleven age groups ranging from .96 to .97 for the Full Scale IQ. At .96 the formula gives 15 x sqrt(.04), or 3.00 points; at .97 it gives 15 x sqrt(.03), or 2.60 points. The average SEM the publisher reports for the Full Scale IQ, 2.90, sits inside that band.
• WAIS-IV Full Scale IQ, reliability .97 to .98: Canivez's review of the adult battery reports internal consistency estimates across all thirteen age groups ranging from .97 to .98. At .98 the formula gives 15 x sqrt(.02), or 2.12 points.
• A short scale, reliability .81: Reliability drops as a scale gets shorter, and Miller and McGill report WISC-V subtest coefficients from .81 to .94. A .81 coefficient on a scale with a standard deviation of 15 gives an SEM of 6.54 points, which is why no responsible report treats a single subtest as an ability estimate.
Notice how unforgiving the square root is. Moving from a reliability of .98 to .96 grows the error by more than 40 percent, and the deterioration accelerates as coefficients fall below .90. The widely repeated "about three points" figure therefore belongs to a specific score from a specific battery rather than to IQ testing in general.
One caveat travels with every manual-reported SEM. Both reviews note that figures based on internal consistency are best-case estimates, because they ignore other real sources of error such as long-term instability and scoring mistakes. Whitaker's analysis of the low IQ range is concrete: combining stability, scorer error and internal consistency into one effective reliability produces a margin of error far wider than the value printed in a manual.
Why the SEM differs from one index to the next
A Full Scale IQ is an aggregate of several subtests, and aggregation is what buys precision. Every additional subtest adds construct-relevant variance faster than it adds error, so the composite is more reliable than any of its parts. A narrow index built from two subtests is therefore measurably less precise, and the published numbers show it.
On the WISC-V, index reliabilities run from .88 to .95 against .96 to .97 for the Full Scale IQ, and average composite standard errors of measurement run from 2.90 for the Full Scale IQ up to 5.24 for the Processing Speed Index. On the WAIS-IV, factor index coefficients span .87 to .98 while the Full Scale IQ holds at .97 to .98. An index at .87 has an SEM of 5.41 points, close to twice the Full Scale figure.
Miller and McGill supply a clean demonstration of the mechanism. They note that the WISC-V Verbal Comprehension Index coefficients are lower than those for the same index on the WISC-IV, and attribute it to the fifth edition's version containing only two subtests where the fourth edition's contained three. Nothing about the construct changed. The index got shorter, and its precision fell with it. The practical consequence is that index scores deserve wider ranges than Full Scale scores, and that a gap between two indexes has to clear the error in both before it means anything. Our overview of how reliable IQ tests are puts those coefficients in context.
Precision is not constant across the score range
A single SEM for a whole test is a convenient average that conceals real variation. The "conditional standard error of measurement" is the standard deviation of the measurement errors affecting test takers at one specified score level, and it is usually larger toward the extremes than near the mean. A test has most of its items targeted at the middle of the ability distribution, so it carries the most information there; in the tails there are fewer items at the right difficulty, and each score point rests on less evidence.
The Standards treat this as a reporting obligation. Standard 2.13 requires the standard error of measurement, both overall and conditional, in the units of each reported score, and Standard 2.14 requires conditional standard errors at several score levels when possible and appropriate. The Standards note that these can be far more informative than a single average, and that item response theory supplies a workable way to estimate them, an approach Price and colleagues applied to composite scores on a Wechsler preschool battery.
Publishers outside the clinical world say the same thing in plainer language. The Australian Council for Educational Research, describing the margins of error printed in its achievement-test norm tables, advises that small score differences should not be given more importance than they deserve, and states directly that error margins are larger for very high and very low scores.
This is where the issue bites. Identification decisions for giftedness and for intellectual disability both happen in the tails, where precision is weakest, and both are commonly made against a fixed cut score. Whitaker concluded that the true margin of error in the low range is substantially greater than the few points suggested by test manuals. A conditional SEM makes that visible; a single average SEM hides it.
From the SEM to a usable score range
On its own the SEM is a diagnostic about the instrument rather than a statement about a person. It becomes an interpretive tool once it is used to build a band around an obtained score, which involves a further choice between an obtained-score interval and an estimated-true-score interval, and a choice of confidence level. Those mechanics live on our page about the IQ confidence interval.
Before any interval is drawn, ask which reliability coefficient produced the SEM, whether the score came from a full-length composite or a two-subtest index, and whether it sits near the mean or out where precision thins. A well-built instrument publishes all three answers, which is part of what distinguishes a normed, documented assessment from an unnormed quiz. The Reasoning and Intelligence Online Test was developed on that basis, and readers who want to take a full-length online IQ test with documented psychometric properties can do so directly.
Frequently asked questions
What is a good standard error of measurement?
There is no universal threshold, because the SEM is in the units of the score. On the IQ metric, full-length composites from major individually administered batteries have SEMs of roughly 2 to 3 points, and narrower indexes run to about 5. Judge any SEM against the scale's standard deviation.
Is the standard error of measurement the same as the standard error of the mean?
No. The standard error of the mean describes how much a sample average varies across repeated samples from a population. The standard error of measurement describes how much one person's observed score would vary across repeated administrations of the same procedure.
How do I calculate the SEM if I only know the reliability?
You also need the standard deviation of the score scale. Subtract the reliability from 1, take the square root, and multiply by the standard deviation. On the IQ metric a reliability of .90 gives 15 x sqrt(.10), or about 4.74 points.
Does a higher SEM mean the test is measuring the wrong thing?
No. The SEM speaks to precision, not to what is being measured. A test can be precise and still measure a construct other than the one intended, which is a question of validity.
Why do two manuals report different SEMs for the same test?
Because they are usually based on different reliability coefficients. An SEM derived from an internal consistency coefficient is smaller than one derived from a test-retest coefficient, since the two count different things as error.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Miller, D. C., & McGill, R. J. (2016). Review of the WISC-V. In A. S. Kaufman, S. E. Raiford, & D. L. Coalson (Eds.), Intelligent testing with the WISC-V (pp. 645-662). Wiley. rjmcgill.com
3. Canivez, G. L. (2010). Review of the Wechsler Adult Intelligence Scale-Fourth Edition. In R. A. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook. Buros Center for Testing. ux1.eiu.edu
4. Price, L. R., Raju, N., Lurie, A., Wilkins, C., & Zhu, J. (2006). Conditional standard errors of measurement for composite scores on the Wechsler Preschool and Primary Scale of Intelligence-Third Edition. Psychological Reports, 98(1), 237-252. pubmed.ncbi.nlm.nih.gov
5. Whitaker, S. (2010). Error in the estimation of intellectual ability in the low range using the WISC-IV and WAIS-III. Personality and Individual Differences, 48(5), 517-521. eprints.hud.ac.uk
6. Australian Council for Educational Research. (2011). Interpreting ACER test results. ACER. acer.org
7. Pearson. Wechsler Intelligence Scale for Children | Fifth Edition (WISC-V). Pearson Assessments. pearsonassessments.com
Figure: original illustration created for RIOT IQ, computing the standard error of measurement as 15 times the square root of one minus reliability, at reliability values reported by Canivez (2010) and Miller and McGill (2016). It is not a reproduction of any published figure.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.