Psychometrics: the science of building and evaluating mental tests
Psychometrics is the science of measuring mental attributes: how tests are built, scored, scaled and judged. Here is what the field covers and who does it.
Dr. Russell T. WarneChief Scientist
Share
"Psychometrics" is the branch of quantitative science concerned with measuring mental attributes: how a test is built, how its items behave, how raw responses become scores, how those scores are placed on a scale, and how anyone can tell whether the resulting numbers mean what they are said to mean. It is the discipline that sits underneath every IQ score, reading assessment, licensure exam and clinical rating scale in use.
The word is often used loosely to mean any workplace questionnaire, and that is a different subject covered by our page on what a psychometric test is. This page is about the discipline: where it came from, the two theories of test scores it runs on, its two central concerns, and the professional standards that govern its output.
What psychometricians actually do
The job is less about administering tests than about the arithmetic of whether a test works. A psychometrician writes or reviews candidate items, pilots them, estimates how difficult each one is and how well it separates stronger from weaker respondents, discards the ones that misbehave, decides how many scores the instrument can defensibly report, builds the tables converting raw performance into a scaled score, and quantifies how much of that number is signal.
Most of that work happens on cognitive and educational instruments, because those are the tests with the longest measurement traditions and the highest stakes. When Pearson reports that the WAIS-5, published in 2024, uses ten primary subtests to produce five index scores for adults from 16 to 90 years and 11 months, every element of that sentence is a psychometric decision: which subtests, how many composites, what age range the norms support.
The field also has an internal critique worth knowing about. Borsboom (2006) argued that psychometrics has developed sophisticated measurement models that substantive psychology has largely failed to adopt, with most psychological tests still resting on the older theory rather than the newer one, and with "construct validity" pressed into service as a catch-all label for problems that deserve separate treatment.
Where the field came from
Psychometrics began as a response to a measurement problem in intelligence research. Spearman (1904) demonstrated that scores on dissimilar cognitive tasks correlate positively and invented factor analysis to describe that structure, which established that an unobservable attribute could be studied through the pattern of correlations among observable indicators.
The discipline became a profession three decades later. The Psychometric Society records its first organizational meeting on 4 September 1935 at Ann Arbor, Michigan, during the American Psychological Association's session, founded by six men including Louis Thurstone, who became its first president. Its journal Psychometrika followed in 1936, and the society's founders came together to start the journal before they thought of starting the society. That institutional base is why psychometrics developed as a distinct quantitative discipline rather than remaining a set of techniques inside psychology.
The two theories of test scores
Almost all practical test construction runs on one of two measurement models, and the difference between them determines what a score can be used for.
• Classical test theory: The older framework treats an observed score as the sum of a true score and an error component, with reliability defined as the proportion of score variance that is true rather than error. Its statistics describe the whole test rather than individual items, and they depend on the sample they were computed in, so an item's difficulty index changes when the group changes. The practical output most readers have met is the standard error of measurement, which our page on the standard error of measurement covers in full.
• Item response theory: The newer framework models the probability of a particular response to a particular item as a function of the respondent's position on a latent trait and of parameters belonging to the item, typically its difficulty and how sharply it discriminates. Because person and item parameters are estimated on a common scale, the same ability estimate can be obtained from different sets of items, which is what makes adaptive testing and test equating possible. The AERA, APA and NCME Standards for Educational and Psychological Testing treat information functions from item response theory as one legitimate way of reporting precision.
Computerized adaptive testing is the most visible consequence. If items are calibrated on a common scale, a test can select each next item on the basis of the answers so far, reaching a given precision in far fewer items than a fixed-length form needs.
Reliability and validity, the field's two central concerns
Everything psychometrics measures about a test reduces to two questions: how consistent are the scores, and do they support the interpretation being placed on them.
• Reliability, or precision: The 2014 Standards deliberately widened the older vocabulary, using "reliability/precision" for the general notion of consistency across instances of a testing procedure and reserving "reliability coefficient" for the coefficients of classical test theory. They also decline to name a numeric threshold. The familiar floors of .80 for group decisions and .90 for individual ones are textbook conventions rather than requirements of the Standards, and the document instead argues that the precision needed rises with the consequences of the decision, so a score driving an irreversible placement demands more than one that will be corroborated elsewhere.
• Validity: The Standards define validity as "the degree to which evidence and theory support the interpretations of test scores for proposed uses of tests," and they are emphatic that "it is the interpretations of test scores for proposed uses that are evaluated, not the test itself." Validity is treated as a single unitary judgement supported by several distinct sources of evidence, covering test content, response processes, internal structure, relations to other variables, and the consequences of testing. Our page on construct validity covers why that older label is now considered dated nomenclature.
Internal structure is where the two halves of the field meet. When a publisher claims that a battery yields five separate index scores, the supporting evidence is a factor analysis, and Standard 1.14 requires that where composite scores are developed, "the basis and rationale for arriving at the composites should be given." Independent reanalyses of the WISC-V standardization sample of 2,200 children have disputed exactly that claim, finding four group factors rather than five and a dominant general factor, which is a live psychometric argument rather than a settled one.
Scaling, norming, fairness and the standards
Three further areas of the discipline get less attention than reliability and validity and matter just as much in practice.
• Scaling and norming: A raw score means nothing on its own, so psychometricians convert it to a metric with a known distribution. For IQ that metric has a mean of 100 and a standard deviation of 15, and the conversion depends entirely on a reference sample. Pearson reports that WAIS-5 norms were collected in 2023 and 2024 from a sample matched to the most current census on five demographic variables, and that kind of stratification is what makes a percentile claim meaningful. A percentile is a statement about a specific reference sample and nothing else, so a score from an instrument with no published norm group cannot be read as a percentile at all. Norms also go stale, which is why major batteries are restandardized every decade or so.
• Fairness and differential item functioning: The Standards state that "fairness is a fundamental validity issue and requires attention throughout all stages of test development and use." The technical machinery behind that principle is differential item functioning analysis, which asks whether respondents of equal overall ability from different groups have systematically different chances of answering a particular item correctly. The Standards are careful to note that differential functioning "is not always a flaw or weakness," since it can also reveal a kind of multidimensionality the test framework anticipated.
• Cut scores and decisions: Where a test is used to sort people, somebody has to choose the boundary, and the discipline treats that choice as a judgement to be documented rather than a quantity to be discovered. Gifted-identification thresholds and diagnostic criteria are examples, and the standard error of measurement around a cut score is the reason a single point should never decide an outcome on its own.
• The professional standards themselves: The Standards for Educational and Psychological Testing, issued jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, is the field's reference document. It is freely available, and reading the validity and fairness chapters is the quickest way to see what the discipline expects of a test.
The practical payoff of all this is modest and important: it tells you how much to trust a number. A score reported without a reference sample, a reliability estimate or a confidence interval is not a measurement in the sense this field means. That is the whole difference between a standardized instrument and an online quiz. Readers who want to see what the discipline's output looks like can take a professionally developed IQ test, the Reasoning and Intelligence Online Test, which reports index scores with published norms and documented precision.
Frequently asked questions
Is psychometrics a science or a set of techniques?
Both, in the sense that it is a quantitative discipline with its own journals, professional society and theory, whose subject matter is the behaviour of measurement instruments. It has been an organised field since 1935 and its results are published in the same way as any other branch of applied statistics.
What qualifications does a psychometrician have?
Usually a graduate degree in quantitative psychology, educational measurement or statistics. The work is mathematical, and in test publishing it sits alongside subject-matter experts who write the items and licensed psychologists who administer and interpret the instrument.
Is psychometrics the same as psychological testing?
No. Testing is the activity of administering an instrument and interpreting the result for a person. Psychometrics is the study of whether the instrument works, which happens before and after any individual is tested.
Does psychometrics apply outside intelligence testing?
Widely. The same models underpin educational achievement testing, licensure and certification exams, clinical symptom scales and personality inventories. Cognitive assessment remains the area where the methods were developed and where the evidence base is deepest.
Why do psychometricians talk about latent variables?
Because the attribute of interest is never observed directly. Reading comprehension and working memory are inferred from performance on tasks, so the models treat the attribute as a latent variable and the test responses as fallible indicators of it.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Spearman, C. (1904). "General intelligence," objectively determined and measured. The American Journal of Psychology, 15(2), 201-292. jstor.org
3. Psychometric Society. History of the Psychometric Society. Psychometric Society. psychometricsociety.org
4. Jones, L. V., & Thissen, D. (2016). A creation narrative for the Psychometric Society and Psychometrika: In the beginning there was Paul Horst. Psychometrika, 81(4), 1120-1128. link.springer.com
5. Borsboom, D. (2006). The attack of the psychometricians. Psychometrika, 71(3), 425-440. pmc.ncbi.nlm.nih.gov
6. Canivez, G. L., Watkins, M. W., & Dombrowski, S. C. (2017). Structural validity of the Wechsler Intelligence Scale for Children-Fifth Edition: Confirmatory factor analyses with the 16 primary and secondary subtests. Psychological Assessment, 29(4), 458-472. doi.org
8. McGrew, K. S. (2023). Carroll's three-stratum (3S) cognitive ability theory at 30 years. Journal of Intelligence, 11(2), 32. pmc.ncbi.nlm.nih.gov
Hero image: Hipp chronoscope, c. 1890, Museum of Science and Industry, Chicago, by Daderot, released under CC0 1.0 (creativecommons.org/publicdomain/zero/1.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.