What is the ITPA-3? The Illinois Test of Psycholinguistic Abilities explained
The ITPA-3 is a 2001 PRO-ED test of spoken and written language for ages 5 to 12. See its 12 subtests, storied history, and why its aging norms matter.
The ITPA-3, the Illinois Test of Psycholinguistic Abilities, Third Edition, is an individually administered, "norm-referenced" test (one that compares a child's performance to a national sample of same-age peers) of spoken and written language for children ages 5 years 0 months through 12 years 11 months. Authored by Donald Hammill, Nancy Mather, and Rhia Roberts and published by PRO-ED in 2001, it takes about 45 to 60 minutes and yields 12 subtest scores that combine into 10 composites. Despite the word "psycholinguistic" in its name, the ITPA-3 is a measure of linguistic processing, and it is not an IQ test. This article covers what the ITPA-3 measures, its unusually influential history, and why the age of its norms is now a serious interpretive caution.
What the ITPA-3 measures
According to PRO-ED, the ITPA-3 "is an effective measure of children's spoken and written language," and all of its subtests measure some aspect of language, including oral language, writing, reading, and spelling. The battery is split evenly into two halves.
The six spoken language subtests sample oral skills. In Spoken Analogies the child completes a four-part analogy ("Birds fly, fish ___"). Spoken Vocabulary asks the child to name a noun from one of its attributes ("I am thinking of something with a roof"). Morphological Closure has the child finish word patterns such as "big, bigger, ___," while Syntactic Sentences asks the child to repeat sentences that are grammatical but nonsensical, such as "Red flowers are smart." Sound Deletion asks the child to remove words, syllables, or "phonemes" (individual speech sounds) from spoken words, and Rhyming Sequences has the child repeat strings of rhyming words that grow longer.
The six written language subtests move to print. Sentence Sequencing asks the child to order sentences into a sensible paragraph. Written Vocabulary asks for a written noun that fits an adjective prompt such as "A broken ___." Sight Decoding and Sight Spelling use words with irregular parts, like "would" and "laugh," while Sound Decoding and Sound Spelling use phonically regular "pseudowords" (made-up words such as "Flant" that follow normal spelling rules). This pairing lets an examiner see whether a child struggles more with memorized word forms or with sounding words out.
Composites and how evaluators use them
The 12 subtests combine into 10 composite scores. The publisher describes the General Language Composite, which draws on all 12 subtests, as the best single estimate of linguistic ability for most children. Several of the others are worth knowing by name.
• Spoken Language and Written Language: each combines six subtests, and the contrast between the two is central to the test's design. PRO-ED states that the oral language versus written language discrepancy can contribute to an accurate diagnosis of dyslexia, since children with dyslexia typically show adequate spoken language alongside poor word identification and spelling.
• Specific two-subtest composites: Semantics, Grammar, Phonology, Comprehension, and Spelling each pair two subtests to isolate one aspect of language, such as competency with speech sounds, including "phonemic awareness" (sensitivity to the individual sounds inside words).
• Sight-Symbol and Sound-Symbol Processing: these composites separate children with poor "orthographic coding" (reading and spelling words with irregular elements) from those with poor "phonological coding" (reading and spelling regular pseudowords), a distinction that can shape intervention plans.
The publisher reports that internal consistency, stability, and interscorer reliability coefficients exceed .90, high enough to support clinical judgments, and that studies found little evidence of bias across demographic groups.
A storied and complicated history
Few tests carry as much history as the ITPA. Samuel Kirk and his colleagues at the University of Illinois published the original edition in 1961, a nine-subtest experimental battery built on psychologist Charles Osgood's model of communication. According to the University of Illinois Alumni Association, Kirk had founded the world's first multidisciplinary research institute on exceptional children in 1952, and he is widely credited with launching the term "learning disabilities" at a 1963 Chicago conference, although Kirk himself noted that he did not invent the phrase. The ITPA, revised in 1968 with James McCarthy and Winifred Kirk, became the signature diagnostic tool of the young learning disabilities field. The idea was appealing: profile a child's separate psycholinguistic processes, find the weak ones, and train them directly.
That second step, known as "psycholinguistic training," did not survive scrutiny. In 1974, Donald Hammill and Stephen Larsen reviewed 38 intervention studies that used the ITPA itself as the measure of improvement and concluded that the effectiveness of psycholinguistic training had not been demonstrated. The debate ran for years in the journal Exceptional Children. Kenneth Kavale's 1981 meta-analysis found some trainable elements, and Larsen and colleagues answered that the evidence still failed to validate the practice. Mainstream special education ultimately moved away from training isolated processing abilities and toward direct instruction in academic skills.
The third edition responded to this history in a candid way. Hammill, the field's most prominent critic of psycholinguistic training, became the ITPA-3's lead author, and the redesigned battery dropped the old perceptual and memory channel subtests. Every ITPA-3 subtest now measures spoken or written language, and the publisher's stated assumptions are modest ones: language is measurable, it can be improved through instruction, and it matters for school subjects such as reading and writing. Those claims sit comfortably with modern reading science, even if the test's name still echoes an earlier era.
Aging norms and the Flynn effect
The ITPA-3's normative data were collected in 1999 and 2000, stratified to match the United States population of that time on characteristics such as region, ethnicity, parental education, and family income. That sample was well constructed, but it is now a quarter century old, and scores from any test are only as current as the peer group behind them. Our article on why norm samples age matters for the Flynn effect explains the general problem.
The "Flynn effect" is the long-run rise in average test performance across generations. A 2014 meta-analysis in Psychological Bulletin by Lisa Trahan and colleagues, covering 285 studies, estimated the gain at about 0.23 IQ points per year, roughly 2.3 points per decade, and showed that outdated norms inflate scores relative to current peers. The effect is documented most thoroughly for intelligence tests, and the exact drift on a language battery like the ITPA-3 has not been established. The interpretive caution still applies. A child tested today is being compared to children tested when dial-up internet was common, before more than two decades of change in reading instruction, screen exposure, and curriculum. An examiner should treat ITPA-3 standard scores as estimates anchored to a dated reference group, and should say so in reports.
Strengths and limitations
• Strengths: the ITPA-3 covers spoken and written language in one coherent battery, its paired regular and irregular word tasks map neatly onto how reading researchers think about decoding, the publisher reports reliability above .90, and administration is a manageable 45 to 60 minutes.
• Limitations: the norms date to 1999 and 2000, the age range stops at 12 years 11 months, and the test's storied name can mislead people into treating it as a measure of cognitive ability. Newer language and literacy instruments with fresher norms are usually preferable when a current peer comparison is the goal.
For the specific skills the ITPA-3 samples, evaluators today often reach for more recently normed tools, such as the CTOPP-2 for phonological processing or the TOWRE-2 for word reading efficiency, both of which use the same kind of real-word versus pseudoword contrast the ITPA-3 helped popularize.
Where the ITPA-3 fits today
The ITPA-3 remains in print through PRO-ED and still appears in school files, research literature, and training programs, so knowing how to read it matters even if you never administer it. Read its composites as statements about language processing, keep the age of the norms in view, and remember that no single score, from this or any test, settles a diagnostic question by itself.
The ITPA-3 measures language, and a well-built intelligence test answers a different question. For adults who want a rigorous picture of their own cognitive abilities, the RIOT IQ test (the Reasoning and Intelligence Online Test), developed by RIOT IQ with psychometrician Dr. Russell T. Warne, offers 15 subtests across six cognitive indices (verbal reasoning, fluid reasoning, spatial ability, working memory, processing speed, and reaction time) in about 52 minutes, with scores reported on the familiar mean-100, standard-deviation-15 scale and current norms behind them. It is designed for adults 18 and older and does not replace an individually administered diagnostic evaluation. You can take the RIOT IQ test at riotiq.com.
Frequently asked questions
Is the ITPA-3 an IQ test?
No. PRO-ED describes it as a measure of children's spoken and written language, and every subtest targets a linguistic skill. Earlier editions sampled broader processes, but the third edition is strictly a language and literacy battery.
What ages does the ITPA-3 cover?
Ages 5 years 0 months through 12 years 11 months, according to the publisher. It is administered individually and takes about 45 to 60 minutes.
Who wrote the ITPA-3, and when was it published?
Donald Hammill, Nancy Mather, and Rhia Roberts authored the third edition, published by PRO-ED in 2001. The original 1961 edition came from Samuel Kirk and colleagues at the University of Illinois.
Why was the original ITPA controversial?
It anchored "psycholinguistic training," a remediation approach that tried to train weak processing abilities directly. Reviews beginning with Hammill and Larsen in 1974 concluded the approach was not validated, and the field shifted toward direct academic instruction.
Are ITPA-3 scores still trustworthy?
The publisher reports strong reliability, but the norms were collected in 1999 and 2000. Scores compare a child to peers from a generation ago, so results should be interpreted cautiously and, where possible, corroborated with more recently normed measures.
Can the ITPA-3 diagnose dyslexia?
By itself, no single test can. PRO-ED states that the ITPA-3's oral language versus written language discrepancy can contribute to an accurate diagnosis of dyslexia within a broader evaluation conducted by a qualified professional.
References
1. Hammill, D. D., Mather, N., & Roberts, R. (2001). ITPA-3: Illinois Test of Psycholinguistic Abilities–Third Edition [Product description]. PRO-ED. proedinc.com
2. Hammill, D. D., & Larsen, S. C. (1974). The effectiveness of psycholinguistic training. Exceptional Children, 41(1), 5–14. eric.ed.gov
3. Kavale, K. (1981). Functions of the Illinois Test of Psycholinguistic Abilities (ITPA): Are they trainable? Exceptional Children, 47(7), 496–510. eric.ed.gov
4. Trahan, L. H., Stuebing, K. K., Hiscock, M. K., & Fletcher, J. M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin, 140(5), 1332–1360. pmc.ncbi.nlm.nih.gov
5. Bateman, B. D. (1968). Interpretation of the 1961 Illinois Test of Psycholinguistic Abilities. ERIC Clearinghouse (ED026771). eric.ed.gov
6. University of Illinois Alumni Association. (2021). Ingenious: The father of special education.. uiaa.org