What is criterion validity? Concurrent and predictive evidence explained
Criterion validity is how closely test scores relate to an outcome a test should predict, such as grades or job performance. The real IQ coefficients.
Dr. Russell T. WarneChief Scientist
Share
Criterion validity is the degree to which scores on a test relate to an external outcome the test is meant to predict or stand in for. That outcome is the "criterion," meaning the real-world result being forecast, and the correlation between test and criterion is called the "validity coefficient." For a cognitive test the criterion is usually school grades, job performance ratings, training results, or attained education and income. This article defines the two named subtypes, concurrent validity and predictive validity, sets out the published coefficients for cognitive ability tests, and explains why those numbers are argued about.
Criterion validity and where it sits in modern validity theory
Administer a test, measure an outcome, correlate the two, and the size of that correlation is the criterion-related evidence for the test.
• The criterion is the outcome, not another test: a criterion is something the test is supposed to say something about. Grade point average, a supervisor's rating, completion of a training course, and diagnostic status are all criteria. Correlating one IQ battery against another is a different argument, covered under construct validity.
• The coefficient is a correlation, so read it as one: a validity coefficient of .30 means the test and the criterion share roughly 9 percent of their variance, and .50 means roughly 25 percent. Useful prediction at the group level is entirely compatible with large errors for individuals.
• It belongs to a use, not to a test: the 2014 Standards for Educational and Psychological Testing, published jointly by AERA, APA and NCME, treat validity as one unified concept supported by several sources of evidence, with test-criterion relationships sitting inside the broader category of evidence based on relations to other variables. Under that framework "criterion validity" is shorthand for one strand of an argument about a specific intended use.
A reasoning test can therefore carry strong criterion evidence for predicting training success and none at all for identifying a learning disorder, and the label on the box tells you nothing about which.
Concurrent validity and predictive validity
The two named subtypes of criterion validity differ in one respect only: when the criterion is collected.
Concurrent validity
Concurrent validity is evidence gathered when the test and the criterion are measured at roughly the same time. The classic employment example is the concurrent study, in which current employees sit the test and their supervisors rate them within the same window. The National Academy of Sciences review of the General Aptitude Test Battery rests largely on this design, and Sackett and colleagues report mean uncorrected validities of .25 in the older set of those studies and .21 in the newer set. Checking a brief measure against a full battery administered in the same session is the same design in a clinical setting.
The weakness is built into the design. Anyone available to be tested has already survived hiring, promotion, or admission, so the sample's scores are squeezed into a narrower band than the applicant pool's would be.
Predictive validity
Predictive validity is evidence gathered when the test comes first and the criterion arrives later, which is the form most people have in mind when they ask whether a test predicts anything.
The strongest cognitive example is longitudinal. Deary, Strand, Smith and Fernandes followed more than 70,000 English children from a cognitive ability test at age 11 to national examination results at age 16, and found a correlation of .81 between the latent intelligence factor and the latent educational achievement factor, with g accounting for 58.6 percent of the variance in mathematics results and 18.1 percent in art and design.
Predictive designs cost years, and they carry their own restriction problem whenever the test itself decided who got followed up. Neither subtype is inherently superior, and the Standards treat both labels as descriptions of when a criterion study collects its data rather than as two kinds of validity.
The coefficients cognitive tests actually produce
Criterion evidence for cognitive ability is unusually rich, and the headline numbers are worth stating precisely.
• School grades: Roth and colleagues meta-analysed 240 independent samples totalling 105,185 participants. The sample-size-weighted observed correlation between intelligence tests and school grades was .44, rising to a corrected population value of .54. By subject the corrected values ran from .49 for mathematics and science down to .31 for fine art and music and .09 for sports.
• Attained status: Strenze's meta-analysis of longitudinal studies reported corrected correlations of .56 with educational attainment, .45 with occupational attainment, and .23 with income among the studies best suited to a predictive reading.
• Work: Schmidt and Hunter's 1998 summary put the validity of general mental ability for job performance at .51 for jobs of medium complexity, which they described as covering 62 percent of United States jobs, with the figure running from .23 for completely unskilled work to .58 for professional and managerial work. For performance in job training programmes they reported .56.
Those employment figures were the field's reference points for two decades, and they are the ones that have since moved. Our overview of what IQ predicts covers the outcome side.
Range restriction and criterion unreliability, the two contested corrections
Published validity coefficients are almost never raw. Two statistical corrections are routinely applied before a number reaches print, and both are defensible in principle and difficult in practice.
• Range restriction: when the people studied span a narrower ability range than the population the test would be used on, the observed correlation is pushed down, and analysts correct upward to estimate what it would have been in the full group. The size of that correction is easy to see in Frey and Detterman's study of 103 undergraduates at a private university, where the observed correlation of .483 between SAT scores held in admissions records and Raven's Advanced Progressive Matrices became .72 once corrected for restricted range. Mean SAT in that sample was 1372 against 854 in the broader national sample they used for comparison.
• Criterion unreliability: the criterion itself is measured with error, which also pushes the observed correlation down. Supervisor ratings are the standard case. Viswesvaran, Ones and Schmidt estimated the mean interrater reliability of supervisory ratings of overall job performance at .52, and correcting a validity coefficient for a criterion that unreliable raises it substantially.
In 2022 Sackett, Zhang, Berry and Lievens argued that the artifact distributions used to apply the range restriction correction in past meta-analyses systematically overcorrected. That recalibration cut validity estimates for job performance by .10 to .20 across most predictors. The estimate for cognitive ability tests dropped from .51 to .31. Structured interviews fell from .51 to .42 and became the top-ranked predictor, work sample tests fell from .54 to .33, job knowledge tests from .48 to .40, and integrity tests from .41 to .31.
Their treatment of the GATB data shows the mechanics. They recorrected the mean observed validities of .25 and .21 for criterion unreliability using a reliability of .60 rather than the .80 the National Academy of Sciences had used, giving .32 and .27, and applied no range restriction correction on the grounds that GATB scores had not been used to select those workers.
The revision is not settled. Oh, Le and Roth challenged the reasoning behind declining to correct concurrent studies, and Sackett and colleagues replied in the same issue. What has changed is the burden of proof: a validity coefficient now has to arrive with an account of which corrections were applied and why.
How to read a criterion validity claim
A coefficient on its own is close to uninterpretable, and a few questions do most of the work.
• What was the criterion, and how well was it measured? A supervisor rating with an interrater reliability near .50 is a noisy target, and a criterion that is itself a test raises the question of whether anything outside the testing room was measured.
• Is the number observed or corrected, and corrected how? The gap between .31 and .51 for the same underlying literature is entirely a matter of correction choices.
• Who was in the sample? Restriction of range is the single largest source of disagreement in this area, and a sample of incumbents or undergraduates is restricted by construction.
• Does the criterion represent the behaviour anyone cares about? A test can predict a laboratory or classroom criterion cleanly and still leave open how far that generalises, which is the separate question handled by ecological validity.
Read with those caveats, the record for cognitive ability remains one of the more robust in applied psychology. Sackett and colleagues, having cut the numbers themselves, concluded that selection procedures remain useful. A revised .31 is still a real relationship, and cognitive tests predict schooling and work outcomes meaningfully without predicting them precisely for any one person.
The Reasoning and Intelligence Online Test is an online IQ test built by psychometricians, developed by RIOT IQ with Dr. Russell T. Warne for adults 18 and over, and reported on the familiar mean-100 scale with a confidence interval attached rather than as a bare number.
Frequently asked questions
What is the difference between criterion validity and construct validity?
Criterion validity asks whether scores relate to a specific external outcome. Construct validity asks whether the test measures the trait it claims to measure. Modern testing standards treat the first as one source of evidence contributing to the second.
What counts as a good criterion validity coefficient?
There is no threshold. In applied psychology, coefficients in the .30s are ordinary and those above .50 are strong, and the honest comparison is against the best available alternative predictor for the same decision.
Is concurrent validity weaker evidence than predictive validity?
Not inherently. Concurrent studies are faster and usually more restricted in range, while predictive studies take years and can suffer attrition. The two generally produce similar estimates once samples are comparable.
Why did the estimated validity of IQ tests for job performance fall?
Sackett and colleagues showed in 2022 that the statistical corrections for restricted samples used in earlier meta-analyses overcorrected, and their recalibration moved the estimate for cognitive ability tests from .51 to .31. Other predictors fell by similar amounts.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. testingstandards.net
2. Deary, I. J., Strand, S., Smith, P., & Fernandes, C. (2007). Intelligence and educational achievement. Intelligence, 35(1), 13-21. doi.org
3. Frey, M. C., & Detterman, D. K. (2004). Scholastic assessment or g? The relationship between the Scholastic Assessment Test and general cognitive ability. Psychological Science, 15(6), 373-378. doi.org
4. Oh, I.-S., Le, H., & Roth, P. L. (2023). Revisiting Sackett et al.'s (2022) rationale behind their recommendation against correcting for range restriction in concurrent validation studies. Journal of Applied Psychology, 108(8), 1300-1310. doi.org
5. Roth, B., Becker, N., Romeyke, S., Schäfer, S., Domnick, F., & Spinath, F. M. (2015). Intelligence and school grades: A meta-analysis. Intelligence, 53, 118-137. doi.org
6. Sackett, P. R., Berry, C. M., Lievens, F., & Zhang, C. (2023). Correcting for range restriction in meta-analysis: A reply to Oh et al. (2023). Journal of Applied Psychology, 108(8), 1311-1315. doi.org
7. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. doi.org
8. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262-274. doi.org
9. Strenze, T. (2007). Intelligence and socioeconomic success: A meta-analytic review of longitudinal research. Intelligence, 35(5), 401-426. doi.org
10. Viswesvaran, C., Ones, D. S., & Schmidt, F. L. (1996). Comparative analysis of the reliability of job performance ratings. Journal of Applied Psychology, 81(5), 557-574. doi.org
Figure: original illustration created for RIOT IQ, plotting the validity coefficients reported in Schmidt and Hunter (1998) alongside the revised estimates in Sackett, Zhang, Berry and Lievens (2022). It is not a reproduction of any published figure.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.