What is developmental screening? The schedule, the tools, the arithmetic
Developmental screening is a brief standardized check given to all young children to flag who needs a closer look. It is not a diagnosis or an IQ test.
Dr. Russell T. WarneChief Scientist
Share
Developmental screening is the brief, standardized check given to every young child at set ages to work out which children need a closer look at their development. What it produces is a flag rather than a finding, so a positive screen tells you that a child should be evaluated and tells you almost nothing about what that child has.
Nearly every misunderstanding here comes from reading the flag as the finding. This page covers the arithmetic that makes screening harder than it looks, the instruments and ages used in the United States, and what follows a positive result.
Where developmental screening sits, and what it is not
Screening is the broadest and cheapest tier of identification. It goes to everyone in an age band regardless of concern, it costs minutes, and its only output is a decision about who moves up a tier. The diagnostic tier above it is longer, individually administered, and aimed at saying what is actually going on. One instrument cannot serve both purposes. Three boundaries are worth marking.
• A screener is not an IQ test: Most are parent-completed questionnaires asking whether a child performs specific everyday behaviors, and they report whether the child falls below a cut point in a domain such as communication or gross motor skill. Nothing they produce is an index score or an IQ.
• A positive screen is not a diagnosis: Federal early-intervention regulation says so in its own language. Under 34 CFR §303.320 the purpose of a screening is to determine whether a child is "suspected of having a disability," and a suspicion triggers an evaluation rather than a label.
• "Screening" means two different things: In schools it names the brief academic measure given to a whole class several times a year inside a multi-tiered support system, which our page on response to intervention describes. The medical sense covered here runs on a developmental timetable in a pediatric office and asks about milestones rather than reading fluency.
Screening is also distinct from surveillance, the continuous and unstructured eliciting of parent concerns that the American Academy of Pediatrics asks clinicians to do at every visit. Screening is scheduled and instrumented.
The arithmetic that makes screening hard
The most useful thing to understand about any screening program has nothing to do with instrument quality. When a condition is uncommon, a test with respectable sensitivity and specificity still produces far more false positives than true ones, because the false positives come from a much larger pool. Paul Meehl and Albert Rosen showed in 1955 that at sufficiently low base rates a cutting score can perform worse than predicting that nobody has the condition.
Work it through. Guthrie and colleagues followed 25,999 children screened for autism with the M-CHAT/F at 16-to-26-month well-child visits in one large primary care network, tracking them for four to eight years afterward. Autism prevalence in that cohort was 2.2 percent, sensitivity was 38.8 percent, and specificity was 94.9 percent. Apply those numbers to a hypothetical 10,000 toddlers.
• Children who have the condition: 220 of the 10,000. The screener correctly flags 38.8 percent of them, which is 85 children, and misses the other 135.
• Children who do not have it: 9,780. The screener wrongly flags 5.1 percent of them, which is 499 children.
• Positive predictive value: 85 plus 499 gives 584 positive screens, of which 85 are children who genuinely have autism. That is 85/584, or 14.6 percent, exactly the value the study reports. Roughly six of every seven positive screens were false alarms.
• Negative predictive value: 9,416 children screen negative and 135 of them have autism, giving 9,281/9,416, or 98.6 percent, again the published value.
The arithmetic reproduces both reported figures, which is the point. No instrument failure is needed to get a positive predictive value of 14.6 percent, because prevalence does most of the work. Raise the base rate and the same test looks transformed. At the roughly 20 percent prevalence seen among younger siblings of autistic children, 10,000 screens yield 776 true positives against 408 false ones, and positive predictive value climbs to about 66 percent with no change to the test. A meta-analysis covering 49,841 children found that pattern, with pooled positive predictive value of 75.6 percent in high-likelihood samples against 51.2 percent in low-risk ones.
The schedule and the instruments that fill it
The Bright Futures and AAP periodicity schedule, in the edition approved in December 2024 and published in February 2025, sets the timetable. General developmental screening tests are to be conducted at the 9-, 18-, and 30-month health supervision visits, and autism spectrum disorder screening tests at the 18- and 24-month visits. Screening should also happen whenever surveillance raises a concern. The schedule's footnotes point to the AAP statements carrying the reasoning, Lipkin and Macias on surveillance and screening and Hyman, Levy and Myers on autism.
The instruments in routine use are few.
• ASQ-3 (Ages and Stages Questionnaires, Third Edition): Age-specific parent questionnaires spanning 1 month to 5.5 years, covering communication, motor, problem-solving, and personal-social domains, with cutoffs set in standard deviations below the mean.
• PEDS (Parents' Evaluation of Developmental Status): A short instrument that elicits and categorizes caregiver concerns rather than scoring milestones directly.
• SWYC (Survey of Well-being of Young Children): A freely available package for 2 to 60 months combining milestones with an autism-specific component and family risk questions.
• M-CHAT-R/F: Twenty parent items for 16 to 30 months plus a structured follow-up interview that re-interrogates failed items before a positive is recorded. That follow-up step keeps the positive rate manageable and is frequently skipped in practice.
• Denver II: Outdated, and worth naming for that reason. Its restandardization was published in 1992 on a Colorado sample, so its norms are over three decades old. Glascoe and colleagues tested it against a full diagnostic battery and found 83 percent sensitivity against only 43 percent specificity, with more than half of typically developing children flagged. The alternative scoring rule lifted specificity to 80 percent while sensitivity collapsed to 56 percent.
Head-to-head comparison is rare. Sheldrick and colleagues gave the ASQ-3, PEDS, and SWYC Milestones to 1,495 families in ten pediatric practices and then administered full developmental testing. Among children under 42 months the ASQ-3 and SWYC reached specificity near 89 percent against 79.6 percent for the PEDS, agreement between instruments was only moderate, and sensitivity exceeded 70 percent only for severe delays. Different screeners flag substantially different children, as educational screeners such as the Brigance Early Childhood Screens also do.
From recommendation to referral, and where children are lost
Recommended practice and observed practice are far apart. In the 2023-2024 National Survey of Children's Health, 36.5 percent of children aged 9 through 35 months were reported to have received a developmental screening using a parent-completed tool in the past year (95 percent CI 34.9 to 38.2). The 2016 figure was 30.4 percent, with state variation spanning 40 percentage points, from 17.2 percent in Mississippi to 58.8 percent in Oregon. A 2019 AAP Periodic Survey found 98.1 percent of pediatricians saying they screen or surveil for developmental delay while only 59.0 percent used a standardized instrument.
A positive screen is meant to produce a referral. For a child under three the destination is the state's Part C early intervention program, and 34 CFR §303.303 requires primary referral sources, physicians among them, to refer no more than seven days after identification. 34 CFR §303.310 then gives the program 45 days from referral to complete the evaluation and the initial Individualized Family Service Plan meeting, and eligibility under 34 CFR §303.21 turns on measured delay in cognitive, physical, communication, social-emotional, or adaptive development. For a child three or older the route runs through the school system toward the kind of battery described on our page on a psychoeducational evaluation.
The attrition along that path is documented and large. Among 2,882 children who screened positive on the M-CHAT at a 16-to-30-month visit in one network, 40.2 percent received at least one recommended referral and 3.7 percent received all of them, with 11 percent referred for a comprehensive autism evaluation. A linked-records study of 14,710 children in a Denver safety-net system found that 18.7 percent of early-intervention-eligible children were referred at all and 26 percent of those referred went on to receive services, a net enrollment rate of 5 percent among eligible children. Referral rates and screener accuracy were both lower for children of color and for lower-income households. A screening program is only as good as the pathway it feeds.
Why the USPSTF has not endorsed universal autism screening
Here the field holds a real disagreement worth stating plainly. The AAP recommends autism-specific screening of all children at 18 and 24 months. The US Preventive Services Task Force does not. Its recommendation covering children aged 18 to 30 months for whom no concerns have been raised by parents or a clinician carries an "I" grade, meaning the Task Force judged the evidence insufficient to assess the balance of benefits and harms. That statement issued in February 2016 and remains the Task Force's operative position, with an update in draft development as of 2026.
The reasoning is narrower than it first appears. The Task Force found adequate evidence that available instruments can detect autism in this age range. What it found inadequate was direct evidence that screening children with no raised concerns improves their outcomes, because no trial has followed screen-detected children through treatment to outcome and the available treatment studies were small, mostly non-randomized, and conducted in clinically referred children. An I statement is a verdict on the evidence base rather than a recommendation against the practice.
Both positions can be held honestly at once. The AAP reasons from the value of early identification and the documented lag before diagnosis; the Task Force applies an evidentiary standard this literature has not met. Neither implies that a parent with a concern should wait.
If you are curious about cognitive measurement in its own right rather than about screening, you can take a full-length online IQ test, the Reasoning and Intelligence Online Test, which reports index scores with confidence intervals. It measures reasoning in adults and has nothing to do with screening toddlers, which shows how narrowly a test's purpose constrains its use.
Frequently asked questions
Is developmental screening an IQ test?
No. A screener checks a young child against age expectations in domains like communication and motor skill and returns a pass-or-refer decision. It is not validated or intended as a measure of general cognitive ability.
What happens if my child fails a developmental screening?
The result should prompt a referral, to the state's Part C early intervention program for a child under three or to a diagnostic evaluation. It is a reason to look more closely and is not a diagnosis.
At what ages is developmental screening recommended?
The Bright Futures and AAP periodicity schedule calls for general developmental screening at the 9-, 18-, and 30-month visits and autism-specific screening at the 18- and 24-month visits, plus screening whenever a concern arises.
Why are so many positive screens false alarms?
Because the conditions screened for are uncommon. With autism prevalence of 2.2 percent, sensitivity of 38.8 percent, and specificity of 94.9 percent, only about one positive screen in seven identifies a child who genuinely has the condition.
Is the Denver II still a good screening test?
No. Its norms date from a 1992 Colorado standardization, and an independent accuracy study found specificity of only 43 percent.
Can I ask for an evaluation if my child passed the screen?
Yes. Under 34 CFR §303.320(a)(3) a parent may request and consent to an evaluation at any point during the screening process, and it must be conducted even where the agency has decided the child is not suspected of a disability.
The takeaway
Developmental screening is a population-level sorting operation, and judging it by the standards of a diagnostic test guarantees disappointment. Its job is to move a manageable subset of children into evaluation, and it does that while generating a majority of false positives, because low base rates make that unavoidable for any instrument. The schedule is clear and the instrument list is short, the Denver II no longer belongs on it, and the real weakness sits between the positive screen and the completed evaluation rather than inside the questionnaire. Read a positive screen as an instruction to find out more, and raise any concern you have whatever the screen said.
References
1. American Academy of Pediatrics. (2025). Recommendations for preventive pediatric health care (Bright Futures/AAP periodicity schedule). downloads.aap.org
2. Lipkin, P. H., & Macias, M. M. (2020). Promoting optimal development: Identifying infants and young children with developmental disorders through developmental surveillance and screening. Pediatrics, 145(1), e20193449. doi.org
3. Hyman, S. L., Levy, S. E., & Myers, S. M. (2020). Identification, evaluation, and management of children with autism spectrum disorder. Pediatrics, 145(1), e20193447. doi.org
4. Meehl, P. E., & Rosen, A. (1955). Antecedent probability and the efficiency of psychometric signs, patterns, or cutting scores. Psychological Bulletin, 52(3), 194-216. doi.org
5. Guthrie, W., Wallis, K., Bennett, A., Brooks, E., Dudley, J., Gerdes, M., Pandey, J., Levy, S. E., Schultz, R. T., & Miller, J. S. (2019). Accuracy of autism screening in a large pediatric network. Pediatrics, 144(4), e20183963. doi.org
6. Aishworiya, R., Ma, V. K., Stewart, S., Hagerman, R., & Feldman, H. M. (2023). Meta-analysis of the Modified Checklist for Autism in Toddlers, Revised/Follow-up for screening. Pediatrics, 151(6), e2022059393. doi.org
7. Sheldrick, R. C., Marakovitz, S., Garfinkel, D., Carter, A. S., & Perrin, E. C. (2020). Comparative accuracy of developmental screening questionnaires. JAMA Pediatrics, 174(4), 366-374. doi.org
8. Glascoe, F. P., Byrne, K. E., Ashford, L. G., Johnson, K. L., Chang, B., & Strickland, B. (1992). Accuracy of the Denver-II in developmental screening. Pediatrics, 89(6), 1221-1225. doi.org
9. Frankenburg, W. K., Dodds, J., Archer, P., Shapiro, H., & Bresnick, B. (1992). The Denver II: A major revision and restandardization of the Denver Developmental Screening Test. Pediatrics, 89(1), 91-97. doi.org
10. Hirai, A. H., Kogan, M. D., Kandasamy, V., Reuland, C., & Bethell, C. (2018). Prevalence and variation of developmental screening and surveillance in early childhood. JAMA Pediatrics, 172(9), 857-866. doi.org
11. Child and Adolescent Health Measurement Initiative. (2025). Children who received developmental screening using a parent-completed screening tool, 2023-2024 National Survey of Children's Health. Data Resource Center for Child and Adolescent Health. nschdata.org
12. Coker, T. R., Gottschlich, E. A., Burr, W. H., & Lipkin, P. H. (2024). Early childhood screening practices and barriers: A national survey of primary care pediatricians. Pediatrics, 154(2), e2023065552. doi.org
13. Wallis, K. E., Guthrie, W., Bennett, A. E., Gerdes, M., Levy, S. E., Mandell, D. S., & Miller, J. S. (2020). Adherence to screening and referral guidelines for autism spectrum disorder in toddlers in pediatric primary care. PLOS ONE, 15(5), e0232335. doi.org
14. McManus, B. M., Richardson, Z., Schenkman, M., Murphy, N. J., Everhart, R. M., Hambidge, S., & Morrato, E. (2020). Child characteristics and early intervention referral and receipt of services: A retrospective cohort study. BMC Pediatrics, 20, 84. doi.org
15. Siu, A. L., & US Preventive Services Task Force. (2016). Screening for autism spectrum disorder in young children: US Preventive Services Task Force recommendation statement. JAMA, 315(7), 691-696. doi.org
16. US Preventive Services Task Force. (2026). Autism spectrum disorder in young children: Screening (recommendation summary, grade I). uspreventiveservicestaskforce.org
17. US Department of Education. (2024). Screening procedures (optional), 34 CFR §303.320. govinfo.gov
18. US Department of Education. (2024). Referral procedures, 34 CFR §303.303. govinfo.gov
19. US Department of Education. (2024). Post-referral timeline (45 days), 34 CFR §303.310. govinfo.gov
20. US Department of Education. (2024). Infant or toddler with a disability, 34 CFR §303.21. govinfo.gov
Hero image: a clinician examining a young child's leg, by Shixart1985, licensed CC BY 2.0 (creativecommons.org/licenses/by/2.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.