Standard setting: how a test's cut score is actually decided
Standard setting is the structured judgement process that places a cut score on a test. The number is chosen by an expert panel, not found in the data.
Dr. Russell T. WarneChief Scientist
Share
"Standard setting" is the structured process by which a panel of qualified judges decides where to place a cut score on a test. The number is chosen rather than discovered. No statistical procedure can locate the point at which a candidate becomes competent or a score becomes gifted, because those categories are defined by human purposes rather than by any discontinuity in the data. Psychometrics supplies a defensible procedure for making the judgement and a record of how it was made.
The Standards for Educational and Psychological Testing treat this as a documentation obligation. Standard 5.21 requires that "the rationale and procedures used for establishing cut scores should be documented clearly," and notes that where examinees are sorted into categories with no pre-established quota, "the standard-setting method must be documented in more detail."
A cut score is a judgement, and the Standards say so
Standard 5.22 requires that "the judgmental process should be designed so that the participants providing the judgments can bring their knowledge and experience to bear in a reasonable way." The commentary lists the supports that make that possible: familiarity with the descriptions of each proficiency level, practice at judging item difficulty with feedback on accuracy, and feedback on the pass rates a provisional standard would produce.
Two further requirements are where weak standard-setting studies fail.
• Disagreement among judges has to be quantified: The Standards ask for "an estimate of the amount of variation in cut scores that might be expected if the standard-setting procedure were replicated with a comparable standard-setting panel." A cut reported without that estimate is a point with an unknown margin.
• Empirical evidence should inform the judgement where it exists: Standard 5.23 asks that cut scores defining categories with distinct interpretations "be informed by sound empirical data concerning the relation of test performance to the relevant criteria." Judgement is unavoidable, and it does not have to be uninformed.
The final call is usually not the panel's. In the Standards' worked example of a state achievement test, expert committees "recommend cut scores," and "the final decision about the cut scores is a policy decision typically made by a policy body such as the board of education for the state."
The named methods, and which ones are current practice
Methods divide into two families. Test-centred methods ask a panel how a hypothetical borderline candidate would perform on the items. Examinee-centred methods start from real people whose competence has been classified by an outside criterion and work back to the score that best separates them.
• Modified Angoff, the dominant method: Judges estimate, item by item, the probability that a minimally qualified candidate would answer correctly, and the summed averages give the cut. The Institute for Credentialing Excellence reports that modified Angoff was the most commonly used approach on credentialing examinations and sees little reason to think that has changed, adding that it is more accurate to speak of the Angoff family than of one method. The modification that matters is iteration: judges rate independently, see the empirical item difficulties, then rate again.
• Bookmark, the runner-up and the schools favourite: Items are ordered from easiest to hardest and each judge marks the point separating items the borderline candidate should get right from those they should not. The same review calls it "perhaps the second most prevalent standard-setting method used in practice behind the Angoff method," popular because several cut scores can be set at once, as a basic, proficient and advanced reporting scheme requires. It needs an item response theory difficulty scale and a chosen response probability, conventionally 50 or 67 percent, and that choice moves the cut materially.
• Contrasting groups and borderline group, the empirical family: Candidates are sorted into qualified and unqualified on an external criterion, the two score distributions are compared, and the cut is placed where misclassification is minimised. The borderline variant averages the scores of candidates judged barely qualified. Both need a trustworthy external criterion, and frequently none exists.
• Ebel and Nedelsky, largely historical: Ebel has judges classify items on relevance as well as difficulty; Nedelsky works from which distractors a borderline candidate could eliminate. Both still appear in comparative research, including recent health-professions work, and neither features in the credentialing survey of current practice.
The methods disagree. One recalculation of a university cut score produced 27.87 by Angoff against values between 18.9 and 25.2 by bookmark, depending on the response probability and model used. A cut score is a property of a procedure as much as of a test.
The gifted threshold and Mensa's 98th percentile
Gifted identification is where most readers meet a cognitive cut score, and the number is looser than it looks. A threshold of 130 on the Wechsler and Stanford-Binet scales, roughly the 98th percentile, is the most widely cited figure, while state policy varies and thresholds of 120 and 125 are also in use, some set on achievement or creativity measures rather than on ability.
The National Association for Gifted Children has pushed back on rigid use of one composite. Its position statement on the fifth edition of the Wechsler children's scale argues that requiring the Full Scale IQ "undermines the identification of many gifted students," because gifted, bilingual and twice-exceptional children often show discrepancies large enough that the composite is not an interpretable single construct, and because processing-speed weaknesses on timed tasks can pull it below a cutoff for reasons irrelevant to advanced academic work. Its recommendation is that any of several reasoning-weighted index scores should be acceptable "if it falls within the confidence interval of the required score for admission," which turns a point into a band. Our page on how gifted testing works covers identification in practice.
High-IQ society admission shows the same threshold set purely by rank. Mensa admits candidates at or above the 98th percentile on an approved test of intelligence, and explains why it works in percentiles: "there are a large number of tests with different scales," and "a result on one test of 132 can be the same as a score of 148 on another test." Its qualifying-score list makes that concrete. The same percentile is a Full Scale IQ of 130 on the Wechsler scales and the Stanford-Binet Fifth Edition, 132 on the earlier Stanford-Binet, 131 on the fourth-edition Woodcock-Johnson cognitive battery, and 148 on the Cattell scale, whose standard deviation is 24. Our explainer on the Mensa IQ test covers the admission routes.
What the DSM-5 actually did with the intellectual disability criterion
The intellectual disability criterion is the most consequential cut score in cognitive assessment, and the usual account of it is half right. The DSM-5 did not abolish the number, it moved it.
Criterion A requires deficits in intellectual functions confirmed by both clinical assessment and individualised, standardised intelligence testing. The American Psychiatric Association states that the manual removed IQ test scores from the diagnostic criteria while "still including them in the text description," so that they "are not overemphasized as the defining factor of a person's overall ability, without adequately considering functioning levels." The same document states that intellectual disability "is considered to be approximately two standard deviations or more below the population, which equals an IQ score of about 70 or below." Severity is graded on adaptive functioning across conceptual, social and practical domains rather than on the score.
Two corrections follow. The text range is 65 to 75 rather than a bare 70, and the DSM-5-TR of 2022 revised the diagnostic features section "to communicate the idea that although one should not be bound narrowly to the 65-75 IQ score range, the diagnosis would not be appropriate for those with substantially higher IQ scores." The second is nomenclature: the current term is intellectual developmental disorder, with intellectual disability in parentheses, aligning the manual with the eleventh revision of the International Classification of Diseases.
Every cut score needs an error band around it
A cut applied to an obtained score without regard for measurement error will misclassify people, and the Standards say so: "the likelihood of misclassification will generally be relatively high for persons with scores close to the cut scores." They add that adequate precision where a cut sits is a prerequisite for reliable classification, which is a design argument. Our page on the standard error of measurement sets out how that precision is quantified and why it is weakest in the tails, which is exactly where gifted and intellectual disability decisions are made.
Renorming compounds the problem. Kanaya, Scullin and Ceci found that students in the borderline and mild range lost an average of 5.6 points when retested on a renormed edition, and were more likely to be classified as having intellectual disability than peers retested on the same edition. A fixed cut applied across a norm change is not a fixed standard.
The practical conclusion is the one the gifted-education position statement reached. Apply a cut score to a range rather than a point, keep the test's published error statistics on the table while doing it, and do not let one administration carry a placement decision. A test that publishes its reliability coefficients and confidence intervals lets you do that arithmetic, which is how the Reasoning and Intelligence Online Test reports results. For a score you can bracket properly, take a professionally developed IQ test.
Frequently asked questions
Is a cut score scientific?
The procedure can be, and the number itself is a value judgement. Psychometrics can say how precisely a test measures near a proposed cut and how much two panels would disagree, and it cannot say where competence begins.
Does DSM-5 use an IQ cutoff of 70?
Not as a criterion. The IQ figure was moved out of the diagnostic criteria into the descriptive text, where the range given is 65 to 75, and severity is graded on adaptive functioning. The DSM-5-TR added that a diagnosis would still not be appropriate at substantially higher scores.
Why is the gifted threshold 130 on some tests and 132 on others?
Because the tests use different standard deviations. The 98th percentile sits two standard deviations above the mean, which is 130 on a scale with a standard deviation of 15 and 132 on a scale with a standard deviation of 16.
Can a cut score be norm-referenced?
Yes. Mensa's 98th percentile and any selection rule admitting a fixed proportion of candidates are norm-referenced cuts. A pass mark set by an expert panel judging item content is criterion-referenced.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Fabrey, L. J., Dwyer, A. C., Friedman, C., Reid, J. B., & Young, P. (2018). Standard setting overview for credentialing programs. Institute for Credentialing Excellence. credentialingexcellence.org
3. Cetin, S., & Gelbal, S. (2013). A comparison of bookmark and Angoff standard setting methods. Educational Sciences: Theory & Practice, 13(4), 2169-2175. files.eric.ed.gov
4. Park, J., Ahn, D. S., Yim, M. K., & Lee, J. (2018). Comparison of standard-setting methods for the Korea Radiological Technologist Licensing Examination: Angoff, Ebel, bookmark, and Hofstee. Journal of Educational Evaluation for Health Professions, 15, 32. jeehp.org
5. National Association for Gifted Children. (2018). Use of the WISC-V for gifted and twice exceptional identification (Position statement). NAGC. portal.nagc.org
8. Mensa International. Getting your IQ tested: frequently asked questions. Mensa International. mensa.org
9. American Mensa. Qualifying test scores for Mensa membership. American Mensa. us.mensa.org
10. Kanaya, T., Scullin, M. H., & Ceci, S. J. (2003). The Flynn effect and U.S. policies: The impact of rising IQ scores on American society via mental retardation diagnoses. American Psychologist, 58(10), 778-790. pubmed.ncbi.nlm.nih.gov
Hero image: Rhine water gauge at Mannheim in flood, by Hubert Berberich (HubiB), licensed CC BY 3.0 (creativecommons.org/licenses/by/3.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.