What is content validity? Sampling the domain a test claims to cover
Content validity is the evidence that a test's items adequately sample the domain it claims to cover, judged by test blueprints and expert review panels.
Content validity is the evidence that the items on a test adequately and relevantly sample the domain the test claims to cover. Haynes, Richard and Kubany (1995) give the working definition still in use: it is "the degree to which elements of an assessment instrument are relevant to and representative of the targeted construct for a particular assessment purpose." It is established before anyone takes the test, through a written blueprint and structured expert review, and it is the one form of validity evidence that is straightforward for an achievement test and genuinely awkward for an intelligence test.
What the term means, and why the Standards stopped using it
The AERA, APA and NCME Standards for Educational and Psychological Testing (2014) treat test content as one of five sources of validity evidence rather than as a validity in its own right. Their section on it describes the evidence available from "an analysis of the relationship between the content of a test and the construct it is intended to measure," where content means "the themes, wording, and format of the items, tasks, or questions on a test."
The Standards then drop the old label outright, saying their treatment "does not follow historical nomenclature (i.e., the use of the terms content validity or predictive validity)." Sireci (1998) reached a similar conclusion from the research side, arguing that the unitary view of validity argues against the term while leaving the underlying requirement untouched. The practical position is that "content validity" remains the standard shorthand, and what it names is a requirement no serious test can skip. It feeds the wider argument described on our page on construct validity.
Two failure modes define the space. The Standards call them "construct underrepresentation," a test that leaves out important aspects of what it claims to measure, and "construct-irrelevance," a test whose scores are pushed around by processes extraneous to its purpose. Content evidence is the first line of defence against both.
The blueprint comes first
Before items exist there is a set of "test specifications," the written plan that fixes what will be measured and how. Standard 4.1 requires specifications to "describe the purpose(s) of the test, the definition of the construct or domain measured, the intended examinee population, and interpretations for intended uses," and the accompanying comment is unusually direct about the standard to be met: "The domain definition should be sufficiently detailed and delimited to show clearly what dimensions of knowledge, skills, cognitive processes, attitudes, values, emotions, or behaviors are included and what dimensions are excluded."
Standard 4.2 extends the specifications to test length, item formats, target psychometric properties, item ordering, timing, administration procedure and scoring rules. A blueprint in this sense is not a topic list. It is the document against which every later content claim is checked.
Sireci and Faulkner-Bond (2014) break the resulting evaluation into four parts, which is the most usable framework available:
• Domain definition: how the attribute is operationally defined, and whether independent experts agree that definition matches the field's understanding of the domain.
• Domain representation: whether the assembled items cover the defined domain rather than clustering in the easy corners of it.
• Domain relevance: whether each individual item is actually relevant to the domain, as opposed to merely adjacent to it.
• Appropriateness of the test development process: whether quality control was in place, including technical review by content experts, item-writing review by measurement specialists, and sensitivity review.
Lawshe's content validity ratio
The best known quantitative method for expert review is Lawshe's (1975) "content validity ratio." A panel of subject-matter experts rates each item into one of three categories, essential, useful but not essential, or not necessary. The ratio is then CVR = (ne − N/2) / (N/2), where ne is the number of panellists calling the item essential and N is the panel size. Values run from −1 to +1, and anything above zero means more than half the panel judged the item essential.
A positive value is not enough on its own, because with a small panel a majority can arise by chance. Lawshe published a table of critical values, and its details have been argued over ever since. Wilson, Pan and Schumsky (2012) tried to reconstruct the original calculation and produced a revised table using a normal approximation. Ayre and Scally (2014) then recomputed the values from exact binomial probabilities and concluded that Wilson and colleagues had misread the spreadsheet function they used, returning one fewer than the true critical number of experts at every panel size.
The Ayre and Scally figures are the ones to work from. At a one-tailed alpha of .05, a panel of 10 needs 9 members to call an item essential, giving a critical ratio of .800; a panel of 20 needs 15, giving .500; a panel of 40 needs 26, giving .300. The table is also not monotonic, because panel size and vote count are both whole numbers. Going from 12 panellists to 13 lowers the required proportion from .833 to .769, while going from 13 to 14 raises it again to .786. Panel size is a design decision with real consequences, not a detail.
Item bias and sensitivity review
Content review is also where fairness problems are caught before norming. The Standards describe the mechanism plainly: "Expert and sensitivity reviews can serve to guard against construct-irrelevant language and images, including those that may offend some individuals or subgroups, and against construct-irrelevant context that may be more familiar to some than others," and note that publishers routinely run such reviews across all test material before a test becomes operational.
They also note that review by a diverse expert panel can point to sources of "irrelevant difficulty (or easiness)" that need further investigation. This is judgemental review, and it is distinct from the statistical detection of differential item functioning, which happens after data collection and belongs to the internal-structure evidence. Our page on whether IQ tests are biased covers the empirical side of that question.
The honest problem: an intelligence test has no syllabus
Here is the tension that makes this topic interesting rather than procedural. Content validation assumes a domain that can be written down. A third-grade mathematics test has one, because a curriculum exists, and the Standards describe this as alignment, "evaluating the correspondence between student learning standards and test content." A licensure exam has one, derived from a practice analysis of the job. In both cases the domain is external to the test and the panel can check the match.
An intelligence test has no equivalent. Its domain is a theoretical construct, and the definition of that construct is what the test is meant to help establish. Haynes and colleagues named the difficulty directly: "Content validation is particularly challenging for constructs with fuzzy definitional boundaries or inconsistent definitions."
The history bears this out. Boake (2002) traced the origins of the subtests in Wechsler's 1939 Wechsler-Bellevue scale and found they came from tests developed between 1880 and the First World War, drawn from anthropometrics, association psychology, the Binet-Simon scales, nonverbal testing of immigrants and schoolchildren, and group testing of military recruits. Wechsler's selection reflected his clinical experience as much as any theory, and the resulting structure "has remained almost unchanged through later revisions." Nobody derived those subtests from a specification of the domain of human intelligence, because no such specification existed.
The nearest thing the field now has to a blueprint is Cattell-Horn-Carroll theory, the hierarchical taxonomy of broad and narrow abilities distilled from decades of factor-analytic work (McGrew, 2009). Modern batteries built explicitly on it, including the Woodcock-Johnson, can at least point to a stated domain and show which broad abilities each subtest is meant to sample. That is a real improvement, and it is still a theory being tested rather than a syllabus being audited.
The consequence is that content evidence alone can never validate an intelligence test. Haynes and colleagues made the general point well: an instrument with poor content validity can still be reliable and can still predict outcomes, if the shared variance comes from elements outside the intended domain. For a cognitive test the weight therefore falls on internal structure, on convergence with other batteries, and on criterion validity, with content evidence doing the narrower job of showing that the items are defensible, relevant and fairly written.
If you want a test whose item development and scoring are documented rather than assumed, the Reasoning and Intelligence Online Test from RIOT IQ is a professionally developed IQ test for adults aged 18 and over, reporting six cognitive indices on the mean-100, standard-deviation-15 scale.
Frequently asked questions
What is content validity in simple terms?
It is whether a test's questions are a fair sample of the subject the test says it covers. If half a driving theory exam asked about engine chemistry, it would have a content problem.
How is content validity measured?
Through structured expert judgement rather than test-taker data. Panels of subject-matter experts rate items for relevance and importance against the written test specifications, often summarised with an index such as Lawshe's content validity ratio.
What is a good content validity ratio?
It depends on panel size. Using exact binomial values at a one-tailed alpha of .05, the minimum is about .800 for a panel of 10, .500 for a panel of 20 and .300 for a panel of 40.
Is content validity the same as face validity?
No, and confusing them is the most common error on this topic. Content evidence is a judgement made by qualified experts against a written domain specification. Face validity is only how the test looks to the person taking it.
Do the Standards still use the term content validity?
Not as a technical term. The 2014 Standards refer to "evidence based on test content" and explicitly set aside the older nomenclature, while keeping the requirement itself.
Why is content validity harder for IQ tests than for school exams?
A school exam has a curriculum to check against. An intelligence test's domain is a theoretical construct with contested boundaries, so there is no external syllabus for a panel to audit the items against.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Haynes, S. N., Richard, D. C. S., & Kubany, E. S. (1995). Content validity in psychological assessment: A functional approach to concepts and methods. Psychological Assessment, 7(3), 238-247. doi.org
3. Lawshe, C. H. (1975). A quantitative approach to content validity. Personnel Psychology, 28(4), 563-575. doi.org
4. Ayre, C., & Scally, A. J. (2014). Critical values for Lawshe's content validity ratio: Revisiting the original methods of calculation. Measurement and Evaluation in Counseling and Development, 47(1), 79-86. doi.org
5. Wilson, F. R., Pan, W., & Schumsky, D. A. (2012). Recalculation of the critical values for Lawshe's content validity ratio. Measurement and Evaluation in Counseling and Development, 45(3), 197-210. doi.org
6. Sireci, S. G. (1998). The construct of content validity. Social Indicators Research, 45(1-3), 83-117. doi.org
7. Sireci, S., & Faulkner-Bond, M. (2014). Validity evidence based on test content. Psicothema, 26(1), 100-107. doi.org
8. Boake, C. (2002). From the Binet-Simon to the Wechsler-Bellevue: Tracing the history of intelligence testing. Journal of Clinical and Experimental Neuropsychology, 24(3), 383-405. doi.org
9. McGrew, K. S. (2009). CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1), 1-10. doi.org
10. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. doi.org
Hero image: collection drawer at the Bohart Museum of Entomology, University of California, Davis, by Daderot, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.