What is face validity? Why a test that looks right can still be worthless
Face validity is whether a test only looks valid to the person taking it. It is not technical validity evidence, and online IQ quizzes have nothing else.
Dr. Russell T. WarneChief Scientist
Share
Face validity is the extent to which a test appears, to the person taking it, to measure what it claims to measure. It is a judgement about surface plausibility made by people who are usually not qualified to make it, and under the current professional standards it is not a form of validity evidence at all. It still matters, because it shapes whether test takers cooperate and whether institutions accept results, and it is the one property that unnormed online IQ quizzes reliably have.
A test's appearance of being valid
The term has an awkward history. Mosier (1947) examined how it was being used and found several incompatible meanings running under one label, including the assumption that a test is valid because it looks sensible, and argued that the ambiguity made the term dangerous rather than merely vague.
Nevo (1985) rehabilitated it by narrowing it. His paper defines face validity as a test's appearance of being valid, treats it as a property of the test that can itself be measured, and reports empirical evidence from examinees' perceptions of a college entrance examination that it can be measured reliably. That is the modern reading: face validity is not evidence about the construct, it is data about perception, and it should be collected from test takers deliberately rather than asserted by the test author.
The key word is "appears." A vocabulary item looks like a test of verbal ability to almost everyone, and that impression is correct. A digit-span task looks to many people like a memory trick with no bearing on intelligence, and that impression is wrong, since working-memory measures load substantially on general ability. Appearance and evidence move independently, in both directions.
The Standards do not recognise it
This is the point most treatments of the topic bury. The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014), the field's governing document, is a 240-page volume covering validity, reliability, fairness, test design, scoring, documentation and the rights of test takers. The phrase "face validity" does not appear in it once.
The five sources of validity evidence the Standards do recognise are test content, response processes, internal structure, relations to other variables, and consequences of testing, and they state that these sources "do not represent distinct types of validity. Validity is a unitary concept." That unified argument is covered on our page on construct validity. How a test looks to a layperson is not among the five. The closest neighbour is content validity, which is also a judgement about items, but it is made by qualified subject-matter experts against a written specification of the domain, and it is documented. Face validity has neither the experts nor the specification.
The Standards do address test-taker perceptions, in their discussion of consequences. Their worked example concerns an employment test that predicts job performance well but leads some applicants to form a negative opinion of the organisation. That is a real and consequential problem, the Standards say, "but one that is not due to a flaw in the intended interpretation of test scores." Perception and score meaning are handled separately, and deliberately so.
Why it still matters
None of this makes face validity unimportant. It has measurable effects on the process of testing, and one of them shows up directly in cognitive ability research.
Chan, Schmitt, DeShon, Clause and Delbridge (1997) had undergraduates complete two parallel cognitive ability tests plus a reactions measure. Test-taking motivation predicted performance on the second test even after controlling for race and for performance on the first. Face-validity perceptions affected subsequent performance too, but, in the authors' words, "only indirectly through test-taking motivation." A test that strikes takers as pointless gets less effort, and less effort shows up as a lower score, which is construct-irrelevant variance by any definition.
• Motivation and effort: this is the mechanism Chan and colleagues isolated, and it is the strongest practical argument for caring about appearance in low-stakes testing, where nothing compels a taker to try.
• Acceptance and reactions: Hausknecht, Day and Thomas (2004) meta-analysed applicant reactions to selection procedures and found perceived job-relatedness sitting among the consistent drivers of favourable reactions, which in turn relate to organisational attractiveness and intentions.
• Defensibility in practice: a test that decision-makers and candidates can see the point of is easier to defend when challenged, which is a matter of institutional and legal practicality rather than psychometrics. A test that looks arbitrary invites challenge even when its technical evidence is sound.
• The cost of chasing it: optimising for appearance can damage measurement. Items that look obviously relevant are often also obviously fakeable, and a test redesigned to please takers can end up sampling the domain worse than the version they disliked.
Allen, Robson and Iliescu (2023), editing the European Journal of Psychological Assessment, argued that face validity has been under-attended in scale construction precisely because of its low technical status, and that treating it as beneath notice has costs of its own. The defensible position is that face validity is worth measuring and reporting, and worth nothing as an argument that a test works.
Where online IQ quizzes live
This is the whole reason the concept deserves a page on a site about intelligence testing. An unnormed online IQ quiz has face validity and, in most cases, nothing else.
The formula is familiar. Present twenty pattern-matching items in a grid that resembles a matrix-reasoning task, add a countdown timer, return a three-digit number near 100, and the experience looks exactly like an intelligence test to someone who has never seen a real one. Every visible cue is right. Nothing underneath it is.
What is missing is everything the Standards require a publisher to document. There is no representative norm sample, so the number returned is not anchored to any population and cannot be a percentile of anything; our page on why a norm sample matters works through why that alone is disqualifying. There is no reliability coefficient and no standard error, so the score comes with no indication of its precision. There is no factor analysis, no convergent study against an established battery, and no technical manual. Standard 7.1 requires that "the rationale for a test, recommended uses of the test, support for such uses, and information that assists in score interpretation should be documented," and that cautions against reasonably anticipated misuses be specified. Quiz sites publish none of this, because there is none to publish.
Flattering feedback does the rest of the work. Forer (1949) demonstrated in a classroom that people rate a generic personality description as an accurate account of themselves when it is presented as an individual result, an effect that has been replicated for decades. A quiz that returns scores clustered comfortably above average is exploiting the same mechanism, and satisfied users read their own satisfaction as evidence the instrument worked.
The practical check is to ignore the interface and look for the documentation. Our guides on how to spot a fake online IQ test and whether online IQ tests are legitimate set out what a real one publishes. Face validity is the one thing you can assess in five seconds, and it is the one thing that tells you nothing.
If you would rather have a score from an instrument with published norms and reported measurement error, the Reasoning and Intelligence Online Test from RIOT IQ is a full-length online IQ test for adults aged 18 and over, reporting six cognitive indices on the mean-100, standard-deviation-15 scale. It is not a replacement for an individually administered clinical assessment.
Frequently asked questions
What is face validity in simple terms?
Whether a test looks like it measures what it says it measures, judged by the person taking it rather than by an expert or by data.
Is face validity a real type of validity?
Not in technical use. The 2014 AERA, APA and NCME Standards never mention it, and recognise five sources of evidence, none of which is lay appearance.
What is the difference between face validity and content validity?
Content evidence is expert judgement of test items against a written definition of the domain, and it is documented. Face validity is an untrained impression of how the test looks, and it is usually not documented at all.
Why do researchers care about face validity if it proves nothing?
Because it affects behaviour. It influences test-taking motivation, which influences scores, and it influences whether candidates and institutions accept a procedure at all.
Can a test have high face validity and low actual validity?
Yes, and that combination is the business model of most free online IQ quizzes. It also runs the other way: several genuinely valid cognitive subtests look trivial or irrelevant to the people taking them.
How do you measure face validity?
By asking test takers, using structured rating scales. Nevo (1985) showed that examinee perceptions of a test's apparent relevance can be collected reliably, which is the appropriate way to treat it.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Mosier, C. I. (1947). A critical examination of the concepts of face validity. Educational and Psychological Measurement, 7(2), 191-205. doi.org
3. Nevo, B. (1985). Face validity revisited. Journal of Educational Measurement, 22(4), 287-293. doi.org
4. Chan, D., Schmitt, N., DeShon, R. P., Clause, C. S., & Delbridge, K. (1997). Reactions to cognitive ability tests: The relationships between race, test performance, face validity perceptions, and test-taking motivation. Journal of Applied Psychology, 82(2), 300-310. doi.org
5. Hausknecht, J. P., Day, D. V., & Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57(3), 639-683. doi.org
6. Allen, M. S., Robson, D. A., & Iliescu, D. (2023). Face validity: A critical but ignored component of scale construction in psychological assessment. European Journal of Psychological Assessment, 39(3), 153-156. doi.org
7. Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. The Journal of Abnormal and Social Psychology, 44(1), 118-123. doi.org
8. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. doi.org
9. Sireci, S., & Faulkner-Bond, M. (2014). Validity evidence based on test content. Psicothema, 26(1), 100-107. doi.org
Hero image: hand holding a smartphone with a blank screen, by Santeri Viinamaki, licensed CC BY-SA 4.0 (creativecommons.org/licenses/by-sa/4.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.