What is internal consistency? Reliability from a single test sitting
Internal consistency is reliability estimated from a single administration, from how far a test's items agree. On the WISC-V, Full Scale IQ reaches .96.
Dr. Russell T. WarneChief Scientist
Share
Internal consistency is reliability estimated from a single test administration, by looking at how far the items within the test agree with one another. If a set of items is measuring one thing, people who do well on some of them should tend to do well on the rest, and the strength of that tendency is what an internal-consistency coefficient reports. On the fifth edition of the Wechsler Intelligence Scale for Children, the internal-consistency estimate for the Full Scale IQ is .96. This page covers how those numbers are produced, which ones the major cognitive batteries publish, and the two warnings that most often go unsaid.
An estimate built from one sitting
The 2014 Standards for Educational and Psychological Testing lists internal consistency as one of three families of reliability coefficient, defining it as coefficients "based on the relationships/interactions among scores derived from individual items or subsets of the items within a test, all data accruing from a single administration."
That last clause is the whole appeal. Administering a test twice to a representative sample is expensive and slow. An internal-consistency coefficient needs one sitting and can be computed from data the publisher already has.
The Standards is careful to call these "less direct estimates" of reliability. They work by treating the agreement between parts of one test as a stand-in for the agreement you would see between two parallel forms. That substitution is reasonable, and it is also the source of every limitation described further down this page.
The glossary of the Standards gives the plain definition: an internal-consistency coefficient is "an index of the reliability of test scores derived from the statistical interrelationships among item responses or scores on separate parts of a test."
Split-half reliability, the Spearman-Brown correction, and alpha
The oldest version of the idea is "split-half reliability." Divide the test into two halves, most often odd-numbered against even-numbered items, score each half separately, and correlate the two scores.
That correlation understates the reliability of the whole test, because each half is only half as long. The fix arrived in 1910, in two papers printed back to back in the same issue of the British Journal of Psychology. Charles Spearman set out the problem of estimating "how much this reliability coefficient will probably be increased by any given additional number of measurements." William Brown, on the facing pages, gave the algebra: for n copies of a test whose single-copy reliability is r, the corrected value is n times r, divided by 1 plus (n minus 1) times r. The formula, Brown wrote, "furnishes a ready means of determining from the reliability coefficient of a single test, the number of applications of the test which would be necessary to give an amalgamated result of any desired degree of reliability."
Applied to a split half, n is 2, so a half-test correlation of .80 becomes a full-test estimate of .89. The result is the Spearman-Brown corrected split-half coefficient, and it is still the working method behind published IQ test manuals. Canivez and Watkins record that WISC-V subtest reliabilities were produced this way for every subtest except the speeded ones, for which short-term test-retest coefficients were substituted instead.
That exception is not a technicality. The Standards warns that "when a test is designed to reflect rate of work, internal-consistency estimates of reliability (particularly by the odd-even method) are likely to yield inflated estimates of reliability for highly speeded tests." On a timed task, an examinee who ran out of time misses the last odd items and the last even items together, which manufactures agreement between halves that has nothing to do with the ability being measured.
"Coefficient alpha," usually called Cronbach's alpha, generalizes the split-half idea. Cronbach showed in 1951 that alpha is "the mean of all split-half coefficients resulting from different splittings of a test," which removes the arbitrariness of choosing one particular split.
Alpha is not the only option, and for many cognitive tests it is not the best one. We have compared alpha against McDonald's omega on real subtest data in a separate article on Cronbach's alpha and McDonald's omega, and that is where the argument between them belongs.
What the cognitive batteries actually report
These are published figures, not illustrations.
• WISC-V: average internal-consistency coefficients for the composites run from .88 for the Processing Speed Index to .96 for the Full Scale IQ and the General Ability Index. Across the 11 age groups, Full Scale IQ estimates sit at .96 to .97. Subtests run from .81 for Symbol Search to .94 for Figure Weights.
• WAIS-IV: across the 13 age groups, Full Scale IQ estimates run .97 to .98, factor index scores .87 to .98, and subtests .71 to .96.
• Woodcock-Johnson IV Tests of Cognitive Abilities: the General Intellectual Ability cluster has a median reliability of .97, and of the 39 median test-level coefficients reported, 38 are .80 or higher.
• Raven's Standard Progressive Matrices: a Portuguese community sample of 522 people aged 12 to 95 produced an alpha of .94, and the earlier literature summarized in the same paper reports a modal value near .91.
• The RIOT IQ test: the 12 non-speeded subtests of the Reasoning and Intelligence Online Test have alphas from .771 for Figure Weights to .944 for Analogies.
One pattern repeats across all of them. Composites outperform their own subtests, and composites built from more subtests outperform those built from fewer, which is why the seven-subtest Full Scale IQ tops every list above. That is not a coincidence, and the reason is the subject of the next section.
The two warnings that usually go missing
A high alpha is regularly reported as though it proved the test measures one thing. It does not.
• Alpha is not a measure of unidimensionality. Schmitt made the distinction cleanly: "Internal consistency is certainly necessary for homogeneity, but it is not sufficient." He constructed two six-item matrices with the same standardized alpha of .86, one of which is clearly two-factor and one of which is not. Sijtsma put it more bluntly, concluding that "alpha is unrelated to the internal structure of the test" and that high and low values alike "can go either with unidimensionality or multidimensionality of the data." Whether a test measures one construct is a question for factor analysis and for construct validity, not for a reliability coefficient.
• Alpha rises with test length. This is the Spearman-Brown relationship read in the other direction. Schmitt states it directly: "It is also the case that alpha increases as a function of test length." A weak twenty-item scale can post a higher alpha than a strong six-item scale. The WISC-V shows the effect in a published manual, since its two-subtest Verbal Comprehension Index reports a lower coefficient than the WISC-IV version did for no reason other than dropping from three subtests to two.
A third caution is worth adding. McNeish's 2018 review argues that alpha rests on "unrealistic assumptions" whose violation usually makes a test "look less reliable than they actually are." The number is a floor as often as a ceiling.
What internal consistency cannot tell you
Reliability coefficients are not interchangeable. The Standards is explicit on this point in Standard 2.6: "internal-consistency, alternate-form, and test-retest coefficients should not be considered equivalent, as each incorporates a unique definition of measurement error." An internal-consistency figure "may reflect only the internal consistency of item responses within an instrument and fail to reflect measurement error associated with day-to-day changes in examinee performance."
In practice that means an alpha says nothing about whether the score would repeat next month, which is the job of test-retest reliability, and nothing about whether two examiners would score the same responses the same way, which is the job of inter-rater reliability. What it does feed directly is the standard error of measurement, the score-scale quantity that turns a coefficient into a confidence band around an individual's result.
For a score from an instrument whose subtest reliabilities are published rather than asserted, the Reasoning and Intelligence Online Test from RIOT IQ is an online IQ test built by psychometricians for adults 18 and over. It does not replace an individually administered diagnostic evaluation.
Frequently asked questions
What is a good internal consistency value?
It depends on the decision the score supports. Nunnally's often-misquoted advice was that .70 suffices in early research while applied settings that hinge on an individual's exact score need .90 as a minimum and .95 as the desirable standard. Published IQ composites generally clear .90.
Is internal consistency the same as Cronbach's alpha?
No. Internal consistency is the property; alpha is one estimator of it, alongside split-half coefficients, KR-20 and omega. A test has internal consistency whether or not anyone computes alpha.
What is the Spearman-Brown formula used for?
It estimates what a test's reliability would be if the test were lengthened or shortened. Its most common use is correcting a split-half correlation upward to represent the full-length test.
Does a high alpha mean the test measures one thing?
No. Alpha can be high for a clearly multidimensional set of items and is not evidence of unidimensionality. That question requires factor-analytic evidence.
Why is internal consistency unsuitable for timed tests?
Because running out of time depresses performance on the late items in both halves at once, which inflates the agreement between halves. Publishers therefore report test-retest coefficients for speeded subtests instead.
Can a test have high internal consistency and still be a bad test?
Yes. Items that are near-duplicates of each other will produce an excellent coefficient while sampling a construct very narrowly. Reliability sets a ceiling on validity without guaranteeing it.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Spearman, C. (1910). Correlation calculated from faulty data. British Journal of Psychology, 3(3), 271-295. doi.org
3. Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3(3), 296-322. doi.org
4. Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. doi.org
5. Canivez, G. L., & Watkins, M. W. (2016). Review of the Wechsler Intelligence Scale for Children-Fifth Edition: Critique, commentary, and independent analyses. In A. S. Kaufman, S. E. Raiford, & D. L. Coalson, Intelligent testing with the WISC-V (pp. 683-702). Wiley. ux1.eiu.edu
6. Canivez, G. L. (2010). Review of the Wechsler Adult Intelligence Scale-Fourth Edition. In R. A. Spies, J. F. Carlson, & K. F. Geisinger (Eds.), The eighteenth mental measurements yearbook (pp. 684-688). Buros Center for Testing. ux1.eiu.edu
7. LaForte, E. M., McGrew, K. S., & Schrank, F. A. (2014). WJ IV technical abstract (Woodcock-Johnson IV Assessment Service Bulletin No. 2). Riverside. info.riversideinsights.com
8. Queiroz-Garcia, I., Espirito-Santo, H., & Pires, C. (2021). Psychometric properties of the Raven's Standard Progressive Matrices in a Portuguese sample. Revista Portuguesa de Investigacao Comportamental e Social, 7(1), 84-101. doi.org
9. Schmitt, N. (1996). Uses and abuses of coefficient alpha. Psychological Assessment, 8(4), 350-353. doi.org
10. Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach's alpha. Psychometrika, 74(1), 107-120. doi.org
11. McNeish, D. (2018). Thanks coefficient alpha, we'll take it from here. Psychological Methods, 23(3), 412-433. doi.org
12. Nunnally, J. C. (1978). Psychometric theory (2nd ed., pp. 245-246). McGraw-Hill. Quoted in Newsom, J. T. (2017), Empirical estimates of reliability. web.pdx.edu
Hero image: chambray weave detail, by Ajay Suresh, licensed CC BY 2.0 (creativecommons.org/licenses/by/2.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.