What Is the Test of Word Knowledge? The TOWK Explained
The Test of Word Knowledge (TOWK) was a vocabulary and semantics assessment for ages 5 to 17. Pearson has now retired it. Here is what the TOWK measured.
Dr. Russell T. WarneChief Scientist
Share
The Test of Word Knowledge, usually shortened to TOWK, was a norm-referenced measure of vocabulary and "semantics" (the meaning system of language) for children and adolescents aged 5 to 17, given one-to-one by a speech-language pathologist, psychologist, or other qualified examiner and scored against published norms. It was never an IQ test. It sampled how well a student understood and used word meanings, which is one narrow strand of the verbal ability that a full intelligence battery covers, so clinicians used it alongside a cognitive measure rather than in place of one. Pearson has since retired the product, and this article covers what the TOWK measured, how it was administered, what its scores meant, and which current instruments took its place.
What the Test of Word Knowledge measured
The TOWK was written by Elisabeth H. Wiig and Wayne A. Secord and published in 1992 by The Psychological Corporation, the company whose clinical assessment catalog later became part of Pearson. Its purpose was to identify students who were behind, or unusually strong, in the word-level meaning skills that underpin classroom language, reading comprehension, and everyday communication.
Published descriptions of the test list four broad content areas:
• Receptive and expressive vocabulary: whether a student recognized a spoken word's meaning, and whether the student could retrieve and produce the right word.
• Knowledge of figurative language: understanding of idioms, metaphors, and other non-literal usage that a strictly literal reader tends to miss.
• Multiple meanings: whether a student could work out which sense of a familiar word a sentence called for, since common English words carry several meanings.
• Conjunctions and transition words: the small connective words that signal cause, contrast, and sequence, and that carry much of the logical weight in written text.
Research groups that used the instrument treated it as a fine-grained vocabulary measure. A randomized controlled trial of a school vocabulary and storytelling program in England, reported by Victoria Joffe and colleagues in 2019, used four TOWK subtests as outcome measures: receptive vocabulary, expressive vocabulary, comprehension of words in multiple contexts, and figurative language use.
How the TOWK was structured and administered
The test was published in two levels of developmentally appropriate tasks spanning the 5 to 17 age range, with subtests labelled by level. It was administered face to face by a trained examiner, one student at a time, with each task following a scripted procedure.
The subtest formats were straightforward and are documented in the peer-reviewed literature that used them:
• Receptive vocabulary: the examiner said a word and the student pointed to the matching picture from a set of four.
• Expressive vocabulary: the examiner presented a picture and the student supplied the name for it.
• Word opposites: the student chose, from three options, the word whose meaning was opposite to a target word.
• Synonyms: the student selected the word closest in meaning to a target word from three options.
Catalog descriptions of the test noted that stimuli were presented in both visual and auditory form, which allowed examiners to accommodate students who read poorly or who had difficulty holding spoken material in memory. The same descriptions record that the TOWK was used as a criterion measure of residual or recovered semantic knowledge after traumatic head injury or acquired aphasia, and in identifying gifted students with unusually strong semantic skills. Because subtests were self-contained, clinicians and researchers often administered only the two or four that answered the question in front of them rather than the whole battery.
What TOWK scores meant
TOWK subtest results were reported as "scaled scores", a metric with a mean of 10 and a standard deviation of 3. On that scale a score of 10 sits at the average for the student's age group, and roughly two-thirds of same-age students fall between 7 and 13. The 2019 Joffe trial reported its TOWK subtest data on exactly that metric, alongside a separate vocabulary test and a Wechsler index reported on the more familiar mean-100 scale.
Two cautions travelled with those numbers, and they apply to any language or ability test.
• Every score carries measurement error: a single scaled score is a best estimate, not a fixed quantity, and a confidence interval around it is the honest way to report it. A one-point or two-point difference between subtests rarely means anything on its own.
• A score describes performance, not a person: a low expressive vocabulary score says that a student named fewer pictures correctly than most age peers on one occasion. Working out why, and whether it reflects a language disorder, limited exposure to the tested words, hearing history, or something else, is clinical work that belongs to a qualified professional using multiple sources of evidence.
The TOWK never produced an IQ, and its scores were not interchangeable with a Full Scale IQ from an individually administered cognitive battery. If you want the longer version of that distinction, our explainer on what vocabulary tests actually measure walks through the difference between a word-knowledge score and a general ability score.
Why the TOWK is no longer available
Pearson has retired the product. The publisher's own product page states that while the TOWK was published and sold by Pearson in the past, it is no longer available, and the old catalog listing for the test now redirects to that retirement notice. Pearson's UK clinical site carries the same message. Anyone still describing the TOWK as a current, purchasable assessment is working from stale information.
In its place, Pearson points buyers to three current instruments:
• PPVT-5: the fifth edition of the Peabody Picture Vocabulary Test, an individually administered receptive vocabulary measure. Pearson lists it for ages 2 years 6 months through 90 and older, with a 10 to 15 minute administration and standard scores on the mean-100, standard-deviation-15 scale. Our guide to the Peabody Picture Vocabulary Test covers it in detail.
• EVT-3: the third edition of the Expressive Vocabulary Test, the expressive counterpart designed to be co-normed with the PPVT-5 so that a student's receptive and expressive vocabulary can be compared directly.
• CELF-5: the fifth edition of the Clinical Evaluation of Language Fundamentals, a broader language battery from the same author line as the TOWK. Our article on the CELF-5 explains what it covers.
Age is the obvious reason a 1992 test fell out of use. Norms collected more than three decades ago no longer describe today's student population, and the vocabulary sampled by a test written in the early 1990s drifts away from current classroom language. A systematic review of child language assessments led by Deborah Denman and published in Frontiers in Psychology in 2017 screened the TOWK out of consideration on exactly that ground, noting it had not been published within the previous 20 years. That review also found that the 15 current assessments it did evaluate all had limitations in their published evidence of psychometric quality, which is a useful reminder that recency alone does not settle the question of test quality.
Word knowledge, language testing, and measured intelligence
Vocabulary sits at the intersection of language and cognition, which is why word-based tasks appear both in language batteries like the TOWK and in intelligence batteries. A vocabulary score reflects accumulated learning: the words a person has encountered, the contexts in which they encountered them, and how efficiently they stored and can retrieve them. That makes it a reasonable indicator of the knowledge-based side of ability, and a poor stand-alone proxy for reasoning with novel material. The store of knowledge a person builds through schooling and reading is related to reasoning ability without being the same thing, and a careful assessment report keeps the two apart.
A test that samples only word meanings, however well built, tells you about one region of the ability map. A cognitive battery samples several. The RIOT IQ test, developed by RIOT IQ with psychometrician Dr. Russell T. Warne, takes the broader approach for adults 18 and over: 15 subtests across six cognitive indices, covering verbal reasoning, fluid reasoning, spatial ability, working memory, processing speed, and reaction time, in about 52 minutes, with scores reported on the same mean-100, standard-deviation-15 scale used by the major clinical batteries. It gives an adult a properly normed picture of general cognitive ability rather than a single verbal slice, and it does not replace an individually administered diagnostic evaluation when a specific language or learning concern is on the table. You can take the Reasoning and Intelligence Online Test at riotiq.com.
Frequently asked questions
Is the Test of Word Knowledge still available?
No. Pearson's product page states plainly that the TOWK is no longer available, and the legacy catalog link for the test now redirects to that notice. Copies may still sit in school and clinic cupboards, but the instrument is out of print and its norms are more than 30 years old.
Was the TOWK an IQ test?
No. It was a norm-referenced language assessment focused on vocabulary and semantic knowledge. It produced subtest scaled scores for word-level meaning skills, not an IQ, and it was typically administered alongside cognitive testing rather than instead of it.
What ages did the TOWK cover?
It was designed for children and adolescents from 5 to 17 years old, with two levels of tasks so that the material suited the student's developmental stage.
Who was qualified to administer it?
Speech-language pathologists, school and clinical psychologists, and similarly trained professionals. It was delivered one-to-one, with the examiner presenting scripted items and recording responses, so it was never a self-administered or group instrument.
What should be used instead of the TOWK today?
Pearson directs users toward the PPVT-5 for receptive vocabulary, the EVT-3 for expressive vocabulary, and the CELF-5 for a broader language profile. Which one fits depends on the referral question, and that choice belongs to the professional carrying out the assessment.
Can an online test substitute for the TOWK?
No. A well-normed online cognitive test can give an adult a reliable estimate of general ability, but it does not stand in for an individually administered language assessment, and no online instrument can determine whether a student has a language disorder. That determination belongs to qualified clinicians working from multiple sources of evidence.
3. Joffe, V. L., Rixon, L., & Hulme, C. (2019). Improving storytelling and vocabulary in secondary school students with language disorder: a randomized controlled trial. International Journal of Language and Communication Disorders, 54(4), 656-672. pmc.ncbi.nlm.nih.gov
4. Groen, M. A., Laws, G., Nation, K., & Bishop, D. V. M. (2006). A case of exceptional reading accuracy in a child with Down syndrome: Underlying skills and the relation to reading comprehension. Cognitive Neuropsychology, 23(8), 1190-1214. pmc.ncbi.nlm.nih.gov
5. Arkansas Communication Board, State Consultant for School-Based Speech-Language Pathology Services. Evaluations and assessments (N-Z). arcommunicationboard.com
6. Denman, D., Speyer, R., Munro, N., Pearce, W. M., Chen, Y.-W., & Cordier, R. (2017). Psychometric properties of language assessments for children aged 4-12 years: A systematic review. Frontiers in Psychology, 8, 1515. doi.org