What is inter-rater reliability? Agreement between two scorers
Inter-rater reliability is how closely two independent scorers agree on one performance. Interrater reliability matters on IQ tests more than expected.
Dr. Russell T. WarneChief Scientist
Share
Inter-rater reliability is the degree to which two or more independent scorers arrive at the same result when they judge the same performance. It is the reliability question that arises whenever a human being has to decide how good an answer was, and it is usually assumed to be irrelevant to intelligence testing. That assumption is wrong. Several of the most heavily weighted subtests on the Wechsler scales require the examiner to judge open-ended spoken answers, and the research on what examiners actually do with that discretion is not reassuring.
Agreement, consistency, and what the terms mean
The 2014 Standards for Educational and Psychological Testing separates two ideas that everyday usage runs together. Its glossary defines "interrater agreement," also called interrater consistency, as "the level of consistency with which two or more judges rate the work or performance of test takers," and "interrater reliability" as "the level of consistency in rank ordering of ratings across raters." The first asks whether two scorers gave the same score. The second asks whether they put the same people in the same order.
The two can come apart. If one scorer is systematically two points more generous than another, agreement is poor while the rank ordering is identical. Both spellings of the term are in circulation, since the Standards writes "interrater" as one word while the hyphenated "inter-rater" is commoner in general use.
The Standards sets a requirement in Standard 2.7: "When subjective judgment enters into test scoring, evidence should be provided on both interrater consistency in scoring and within-examinee consistency over repeated measurements."
Why an IQ test needs this at all
Matrix Reasoning has a key. Coding has a stopwatch. Vocabulary does not. On the Wechsler scales, the examiner asks what a word means, writes down what the person says, and then assigns 0, 1 or 2 points by comparing that verbatim response against sample responses in the manual. Similarities and Comprehension work the same way. These are individually administered clinical instruments, and a meaningful share of the Full Scale IQ rests on an examiner's in-the-moment judgment about whether an answer was good enough.
The publishers know this and study it. For the WISC-V, every record form in the standardization sample was scored twice by independent scorers, producing agreement coefficients of .98 to .99. Because the Verbal Comprehension subtests are the judgment-heavy ones, Pearson ran a separate study: 60 record forms drawn at random from the normative sample, scored independently by nine raters who were in doctoral clinical psychology programs and had no prior training on WISC-V scoring criteria. The intraclass correlations came out at .98 for Similarities, .97 for Vocabulary, .99 for Information, and .97 for Comprehension. The WAIS-IV figures are similar, with three graduate-student raters on 60 cases producing intraclass correlations from .91 to .97 on the four judgment subtests.
Those are good numbers. Pearson's own report attaches the caveat that matters: "Given the extensive training, feedback and support provided to the scorers participating in the study, it is not clear whether the estimated interrater agreement rates would apply to the typical clinician who does not receive this type of feedback and support."
Does that hold away from the publisher's training program?
It does not. The independent literature on examiner scoring errors in individually administered IQ tests is large, consistent, and almost entirely absent from public discussion of whether IQ tests are reliable.
• The errors concentrate exactly where the judgment is. Mrazik and colleagues examined six consecutive WISC-IV administrations by each of 19 graduate students and found 511 errors across 94% of protocols, a mean of 4.48 per protocol. Vocabulary, Similarities and Comprehension accounted for 80% of them. Performance did not improve across the six administrations.
• Experienced scorers err too. Brazelton and colleagues had 126 scorers of mixed training and seniority score three WISC-III protocols. No participant scored all three without error. The share making at least one scoring error was 98% on Comprehension, 96% on Vocabulary and 75% on Similarities, against 16% on Block Design.
• The same protocol can yield very different IQs. Ryan and Schnakenberg-Ott gave the identical two WAIS-III protocols to 19 psychologists and 19 graduate students. On the second protocol the psychologists' Verbal IQ scores spanned 12 points and the students' spanned 25. Across both protocols, only 42% of psychologists and 32% of students reproduced the correct Full Scale IQ exactly, and for the Verbal Comprehension Index the figures were 37% and 21%.
• Part of a real-world IQ score belongs to the examiner. McDermott, Watkins and Rhoad analyzed 2,783 children assessed by 448 school psychologists in routine special-education evaluations. The proportion of score variance attributable to which psychologist happened to do the testing was 12.5% for Full Scale IQ, 10.0% for the Verbal Comprehension Index, 14.3% for Vocabulary and 10.7% for Comprehension. Matrix Reasoning, which is objectively scored, showed 2.8% and no statistically significant effect.
That last contrast is the cleanest evidence available that the problem lies in the judgment rather than in carelessness generally. Styck and Walsh's meta-analysis of the Wechsler examiner-error literature reached the same place from a different direction. Averaged across studies, 99.7% of protocols contained at least one examiner error when failure to record a response was counted, and 41.2% did when errors of omission were ignored. On average, "73.1% of Full-Scale IQ (FSIQ) scores changed as a result of examiner errors." Their conclusion is worth quoting in full: "current estimates for the standard error of measurement of popular IQ tests may not adequately capture the variance due to the examiner."
Cohen's kappa or the intraclass correlation
Which statistic to use depends on what kind of judgment the raters made.
• Use kappa for categories. Cohen's kappa, introduced in 1960, measures agreement on nominal decisions, such as whether a response is codable as a particular error type or whether a diagnosis is present. Its value is the agreement observed beyond what chance assignment would produce. Landis and Koch's widely used bands treat .00 to .20 as slight, .21 to .40 as fair, .41 to .60 as moderate, .61 to .80 as substantial and .81 to 1.00 as almost perfect.
• Use the intraclass correlation for scores. As Hallgren puts it in a standard tutorial, kappa and its variants suit nominal variables while "the intra-class correlation (ICC) is one of the most commonly-used statistics for assessing IRR for ordinal, interval, and ratio variables." A Vocabulary item scored 0, 1 or 2 is ordinal, which is why Pearson's verbal-subtest study reports intraclass correlations rather than kappas.
Two cautions travel with these numbers. Kappa is sensitive to how common the category is, and a lopsided distribution can produce very high raw agreement with a near-zero kappa. Viera and Garrett work through a table with 85% observed agreement and a kappa of .04, and cite clinical data in which two examiners agreed 85% of the time on a physical sign with a kappa of .01. For the intraclass correlation, Koo and Li recommend judging the 95% confidence interval rather than the point estimate, and offer the bands that values "less than 0.5 are indicative of poor reliability, values between 0.5 and 0.75 indicate moderate reliability, values between 0.75 and 0.9 indicate good reliability, and values greater than 0.90 indicate excellent reliability."
What high agreement does not buy you
The Standards adds a warning that is easy to miss. Its commentary on Standard 2.7 notes that "high interrater consistency does not imply high examinee consistency from task to task. Therefore, interrater agreement does not guarantee high reliability of examinee scores."
In other words, two scorers agreeing perfectly tells you the scoring rules are being applied consistently. It does not tell you the person would perform the same way on a different set of tasks, which is the domain of internal consistency, or on a different day, which is the domain of test-retest reliability. Rater error is one contributor to the standard error of measurement, and on the evidence above it is a contributor that published standard errors from the Wechsler manuals probably understate.
None of this argues against individually administered testing, which buys clinical observation that no automated format can. It argues for double-scoring verbal subtests and for treating a single Vocabulary or Comprehension scaled score as provisional. It also matters that a task scored by rule rather than by judgment has no inter-rater variance to report, which is one reason the Reasoning and Intelligence Online Test from RIOT IQ uses objectively scored item formats throughout. It is a way to take a full-length online IQ test for adults 18 and over, and it does not replace an individually administered diagnostic evaluation.
Frequently asked questions
What is a good inter-rater reliability value?
For scores used in decisions about individuals, intraclass correlations above .90 are the usual expectation, and Koo and Li treat .75 to .90 as good. On the kappa scale, Landis and Koch call .61 to .80 substantial.
Is interrater reliability the same as inter-rater agreement?
The Standards distinguishes them. Agreement asks whether scorers gave the same score; reliability asks whether they ranked people in the same order. Two scorers can rank identically while disagreeing on every absolute value.
Does inter-rater reliability apply to IQ tests?
Yes, on any subtest where the examiner judges an open response. On the Wechsler scales that includes Vocabulary, Similarities and Comprehension, and independent research finds those subtests carry the most examiner error.
When should I use kappa instead of the intraclass correlation?
Use kappa when raters assign cases to unordered categories. Use the intraclass correlation when they assign numbers with meaningful order or spacing, such as item scores of 0, 1 or 2.
Why can agreement be 85% and kappa still be near zero?
Because kappa subtracts the agreement expected by chance. When almost every case falls into one category, chance agreement is already very high, so a large raw percentage leaves little for kappa to credit.
Do computer-scored tests need inter-rater reliability evidence?
Not for items scored by rule, since a scoring algorithm applied twice returns the same answer. Automated scoring of free text is a different case and does need agreement evidence against human raters.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Pearson. (2018). Efficacy research report: WISC-V. Pearson. pearson.com
3. Canivez, G. L. (2010). Review of the Wechsler Adult Intelligence Scale-Fourth Edition. In R. A. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook (pp. 684-688). Buros Center for Testing. ux1.eiu.edu
4. Mrazik, M., Janzen, T. M., Dombrowski, S. C., Barford, S. W., & Krawchuk, L. L. (2012). Administration and scoring errors of graduate students learning the WISC-IV: Issues and controversies. Canadian Journal of School Psychology, 27(4), 279-290. doi.org
5. Brazelton, E. W., Jackson, R., Buckhalt, J., Shapiro, S., & Byrd, D. (2003). Scoring errors on the WISC-III: A study across levels of education, degree fields, and current professional positions. The Professional Educator, 25(2), 1-8. files.eric.ed.gov
6. Ryan, J. J., & Schnakenberg-Ott, S. D. (2003). Scoring reliability on the Wechsler Adult Intelligence Scale-Third Edition (WAIS-III). Assessment, 10(2), 151-159. doi.org
7. McDermott, P. A., Watkins, M. W., & Rhoad, A. M. (2014). Whose IQ is it? Assessor bias variance in high-stakes psychological assessment. Psychological Assessment, 26(1), 207-214. edpsychassociates.com
8. Styck, K. M., & Walsh, S. M. (2016). Evaluating the prevalence and impact of examiner errors on the Wechsler scales of intelligence: A meta-analysis. Psychological Assessment, 28(1), 3-17. doi.org
9. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46. doi.org
10. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159-174. doi.org
11. Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23-34. doi.org
12. Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155-163. doi.org
13. Viera, A. J., & Garrett, J. M. (2005). Understanding interobserver agreement: The kappa statistic. Family Medicine, 37(5), 360-363. pubmed.ncbi.nlm.nih.gov
Hero image: certified kata judges scoring a competition, by Gotcha2, licensed CC BY-SA 3.0 (creativecommons.org/licenses/by-sa/3.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.