What is a diagnostic assessment? How it differs from a screener
A diagnostic assessment is built to locate a specific skill gap or support a diagnosis, which makes it longer and finer grained than a screening test.
Dr. Russell T. WarneChief Scientist
Share
A diagnostic assessment is a test given to establish precisely which skills or symptoms a person does and does not have, so that the result can direct instruction or feed into a diagnosis. That purpose is what makes it longer, finer grained and often individually administered, and it is what separates it from a screener, whose only job is to flag who needs a closer look.
The phrase carries two meanings depending on who is using it. In education a "diagnostic assessment" is usually a skills inventory given before or during instruction to find out what a learner has not yet mastered. In clinical and school-psychology practice it usually means an instrument contributing evidence toward a formal diagnosis. This page covers both readings, places them inside the wider taxonomy of assessment purposes, and explains the psychometric reason a screener and a diagnostic instrument cannot be the same test.
Assessments are classified by purpose, not by format
A common error is treating "diagnostic", "formative" and "summative" as descriptions of what a test looks like. They describe what a result is used for. The same forty multiple-choice items can be formative on Monday and summative in June, and a test's technical requirements follow from the decision it supports rather than from its content.
• Screening: A brief measure given to everyone in a defined population to identify who needs further attention. The canonical definition, carried into the World Health Organization's screening monograph from a 1951 Commission on Chronic Illness conference, is "the presumptive identification of unrecognized disease or defect by the application of tests, examinations, or other procedures which can be applied rapidly." The monograph is explicit that "a screening test is not intended to be diagnostic."
• Diagnostic: A longer or finer-grained measure given to the smaller group a screener flagged, designed to say what specifically is wrong rather than whether something might be. Glover and Albers argue that screeners and diagnostic instruments must be judged against different criteria, because technical adequacy and usability trade off differently when a test goes to everyone than when it goes to a referred few.
• Formative: Assessment embedded in teaching, used to adjust what happens next. Black and Wiliam's review of the classroom-assessment literature established the modern usage. Perie, Marion and Gong describe its cycle as running from about five seconds to an hour, with no real interest in aggregating the information beyond one classroom.
• Summative: Assessment of what has been learned by the end of a period, reported so it can be aggregated. State accountability testing is the clearest example. 34 CFR §200.5(a)(1) requires annual reading or language arts and mathematics assessment in each of grades 3 through 8 and at least once in high school.
A fifth category sits between formative and summative. Perie, Marion and Gong proposed "interim assessment" as an umbrella term for the products districts buy under the labels benchmark, diagnostic and predictive, defined as assessments given during instruction to evaluate skills against a specified set of academic goals, with results reportable in aggregate. A vendor calling a product "diagnostic" is therefore making a claim about grain size rather than about diagnosis. The i-Ready Diagnostic is a widely used instrument of that kind.
The two meanings of "diagnostic assessment"
In the educational sense, a diagnostic assessment answers the question "what does this learner not yet know?" It is typically given before a unit or after a screener has flagged a concern, and it samples subskills densely enough that a teacher can see where a chain of reasoning breaks. Ketterlin-Geller and Yovanoff argue that designing such an instrument requires an explicit model of how the skill develops, because the inference concerns a location on a learning progression rather than standing in a norm group. Leighton and Gierl make the parallel psychometric argument: diagnostic inferences about a student's misconceptions require a cognitive model of the task, and most educational tests carry one only implicitly.
In the clinical sense, a diagnostic assessment is one component of an evaluation whose output is a diagnosis or an eligibility decision, and the instrument is almost never sufficient by itself. Federal special education regulation is blunt about this. 34 CFR §300.304(b)(1) requires a variety of assessment tools and strategies, and §300.304(b)(2) forbids using "any single measure or assessment as the sole criterion" for deciding whether a child has a disability. The clinical usage names a role in an argument rather than a property of a booklet.
Both readings share a structure. The diagnostic tier is entered only once something has raised a question, and it answers a narrower question than the one that got the person there rather than re-answering the screening question with more decimal places.
Why a screener and a diagnostic instrument are built differently
A screener is deliberately tuned to miss as few cases as possible, which means accepting a fair number of false alarms. Two terms carry this. "Sensitivity" is the proportion of people who genuinely have the condition that the test correctly flags. "Specificity" is the proportion of people without it that the test correctly clears. Moving a cut score to raise one lowers the other.
Compton and colleagues illustrated the trade-off in a study of 355 first graders screened in the autumn and assessed for reading difficulty at the end of second grade. With sensitivity fixed at .90, their base screening model achieved specificity of .85 and an area under the curve of .948, good performance by the standards of the field. They were candid about the price of that setting: "false positives undermine prevention efforts by burdening schools with the obligation to provide early intervention to an unnecessarily large percentage of the population."
Their solution was a two-stage gated procedure, and its logic is the logic of the whole taxonomy. A brief first-stage measure clears the children who are plainly fine, and only the remainder receive the full battery. A standard-score cut of 108 on phonemic decoding efficiency removed 154 true negatives and cut the number needing the long battery by 43.4 percent without degrading the model's accuracy. The expensive instrument is reserved for the ambiguous cases.
That is also why reliability requirements differ by tier. "Reliability" is the consistency of a score across occasions, items or raters, and how much you need depends on whose decision rests on the number. Bracken's widely cited technical standards set a floor near .80 for a subtest used to generate hypotheses and near .90 for a total score used in a decision about an individual child. A screener can sit below the higher threshold because its output is a referral, which is revisable. A diagnostic instrument usually cannot, because its output feeds a label and a placement.
The governing principle is Kane's. Validity is not a property a test possesses, it is the degree to which a particular interpretation and use of the scores is justified. A screener read as a diagnosis is an invalid use of a perfectly good test.
What a positive screen actually tells you
Less than most people assume, and the reason is arithmetic rather than instrument quality. Meehl and Rosen set this out in 1955 in what remains the clearest statement of the problem. When a condition is uncommon, a test with respectable sensitivity and specificity will still generate more false positives than true ones, because the false positives come from a much larger pool. In low base-rate settings, they showed, a cutting score can perform worse than predicting that nobody has the condition.
The quantity that matters is "positive predictive value", the probability that someone who screens positive genuinely has the condition. It depends on the "base rate", the prevalence of the condition in the population screened, and not only on the test. Two consequences follow.
• A positive screen is a question, not a finding: It licenses a closer look and nothing more. Reporting it as a probable diagnosis misuses the instrument and predictably alarms families.
• The same test performs differently in different settings: A screener validated in a high-prevalence clinic will show lower positive predictive value in a general population, even though its sensitivity and specificity are unchanged.
What a diagnostic result licenses, and what it does not
A diagnostic assessment narrows the field. It rarely closes it. In an educational evaluation the diagnostic tier is a battery rather than a single test, which is a subject in its own right. Our page on a psychoeducational evaluation walks through what such an evaluation contains and who assembles it, and our page on an achievement test covers the academic-skills half of that battery and the score metrics it reports.
Two questions keep the taxonomy honest. What decision is this result supposed to support, and was the instrument normed on people resembling the person in front of me? Read the confidence interval before the point score as well, because a diagnostic-tier decision made on a number without its error band rests on an illusion of precision.
One caveat about educational "diagnostic" products deserves stating. Kingston and Nash's meta-analysis of formative assessment found a weighted mean effect size near 0.20, well below the 0.40 to 0.70 range often quoted from earlier reviews, and noted how thin the evidence was for commercial systems specifically. Finer-grained information does not improve learning by itself.
For a sense of what a well-normed cognitive measure looks like as evidence, you can take a professionally developed IQ test, the Reasoning and Intelligence Online Test, which reports index scores with confidence intervals rather than a bare number. It measures reasoning and is not a diagnostic instrument, so reading its output is useful practice for reading any test report.
Frequently asked questions
Is a diagnostic assessment the same as a diagnosis?
No. A diagnostic assessment produces evidence that a qualified professional weighs alongside history, observation and other measures. In education the term often does not involve a diagnosis at all and simply means a detailed skills inventory.
What is the difference between a screening and a diagnostic assessment?
Purpose and grain. A screener is brief, given to everyone, and tuned to catch nearly all cases at the cost of false alarms. A diagnostic assessment is longer, given only to those flagged, and built to identify which specific skill or symptom is present.
Is a diagnostic assessment given before or after teaching?
In the educational sense, usually before a unit of instruction or immediately after a screener raises a concern, so the results can shape what is taught. A test given after teaching to certify what was learned is summative.
Is a benchmark test a diagnostic assessment?
Usually not, despite the marketing. Benchmark and interim tests are given periodically to track progress toward year-end goals and are designed to be aggregated, which is a different purpose from locating one learner's specific skill gap.
Why did my child's positive screening result turn out to be nothing?
Because screeners are designed that way. When a condition is uncommon, most positive screens are false positives even for an accurate test, which is exactly why the result triggers a second look rather than a conclusion.
The takeaway
A diagnostic assessment is defined by the decision it serves. It sits one tier below screening, it is built to say what is wrong rather than whether something might be, and it carries higher reliability demands because an individual decision rests on it. The word means a pre-instruction skills inventory in education and a contributor to a formal diagnosis in clinical practice, and the two readings are easy to confuse because vendors use the first label while parents hear the second. Before accepting any assessment result, ask what purpose the instrument was built for and whether the use in front of you matches it.
References
1. Wilson, J. M. G., & Jungner, G. (1968). Principles and practice of screening for disease (Public Health Papers No. 34). World Health Organization. iris.who.int
2. Glover, T. A., & Albers, C. A. (2007). Considerations for evaluating universal screening assessments. Journal of School Psychology, 45(2), 117-135. doi.org
3. Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. doi.org
4. Perie, M., Marion, S., & Gong, B. (2009). Moving toward a comprehensive assessment system: A framework for considering interim assessments. Educational Measurement: Issues and Practice, 28(3), 5-13. doi.org
5. Perie, M., Marion, S., & Gong, B. (2007). A framework for considering interim assessments. National Center for the Improvement of Educational Assessment. nciea.org
6. Compton, D. L., Fuchs, D., Fuchs, L. S., Bouton, B., Gilbert, J. K., Barquero, L. A., Cho, E., & Crouch, R. C. (2010). Selecting at-risk first-grade readers for early intervention: Eliminating false positives and exploring the promise of a two-stage gated screening process. Journal of Educational Psychology, 102(2), 327-340. pmc.ncbi.nlm.nih.gov
7. Meehl, P. E., & Rosen, A. (1955). Antecedent probability and the efficiency of psychometric signs, patterns, or cutting scores. Psychological Bulletin, 52(3), 194-216. doi.org
8. Bracken, B. A. (1987). Limitations of preschool instruments and standards for minimal levels of technical adequacy. Journal of Psychoeducational Assessment, 5(4), 313-326. doi.org
9. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1-73. doi.org
10. Ketterlin-Geller, L. R., & Yovanoff, P. (2019). Considerations for using mathematical learning progressions to design diagnostic assessments. Measurement: Interdisciplinary Research and Perspectives, 17(1), 1-22. doi.org
11. Leighton, J. P., & Gierl, M. J. (2007). Defining and evaluating models of cognition used in educational measurement to make inferences about examinees' thinking processes. Educational Measurement: Issues and Practice, 26(2), 3-16. doi.org
12. Kingston, N., & Nash, B. (2011). Formative assessment: A meta-analysis and a call for research. Educational Measurement: Issues and Practice, 30(4), 28-37. doi.org
13. U.S. Department of Education. (2016). Assessment administration, 34 CFR §200.5. ecfr.gov
14. U.S. Department of Education. (2006). Evaluation procedures, 34 CFR §300.304. ecfr.gov
Hero image: school children sitting scholarship examinations in a classroom, Queensland, 16 April 1940. State Library of Queensland, no known copyright restrictions (flickr.com/commons/usage). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.