What is the Fountas & Pinnell Benchmark Assessment System? Levels, editions, and the evidence debate
The Fountas & Pinnell Benchmark Assessment System is a one-on-one leveled reading assessment for grades K-8, not an IQ test. Levels, editions, evidence.
Dr. Russell T. WarneChief Scientist
Share
The Fountas & Pinnell Benchmark Assessment System (usually shortened to "BAS") is a one-on-one reading assessment published by Heinemann in which a teacher listens to a student read a short leveled book aloud, records the errors, and then talks with the student about the text to gauge comprehension. The result is a reading level on the Fountas & Pinnell A to Z text gradient, used to place students in small reading groups and to track growth across the year. It is not an IQ test, and it is not a diagnostic instrument for dyslexia or any other condition. This article covers what the BAS measures, how the current editions are organized, what the research shows, and why the system faces criticism and restrictions.
What the Benchmark Assessment System measures
The publisher describes the BAS as a "formative reading assessment," meaning its purpose is to inform teaching rather than to produce a score for accountability. According to the publisher's field study summary, it measures decoding, fluency, vocabulary, and comprehension for kindergarten through eighth grade. Every level of the text gradient has two short original books, one fiction and one nonfiction.
The core of the assessment is the reading record. The teacher reads a standardized introduction, the student reads aloud while the teacher codes each error and self-correction, and the two then hold a "comprehension conversation" using standardized prompts. From the accuracy, fluency, and comprehension scores the teacher identifies the student's "independent" level (text the student can read alone), "instructional" level (text the student can read with support), and "hard" level.
Fountas and Pinnell's research page is candid about what the system is not. It states that the BAS "does not provide national norms or percentiles" and "is not intended for national achievement testing." That sentence separates the BAS from norm-referenced screeners such as DIBELS or the i-Ready Diagnostic, and even more sharply from a cognitive ability test.
Editions currently in use: Third Edition and BAS 2.0
Heinemann released the BAS in 2007, according to APM Reports, followed by a second edition field-tested around 2011 to 2012 and the Third Edition in 2016. Its main change, in the publisher's words, was a more rigorous comprehension conversation with new scoring rubrics and the removal of an "extra point" available in the second edition. Heinemann warns schools not to mix editions and notes that second-edition materials "may result in a slightly higher score" than the third.
The Third Edition is split into two boxed kits:
• System 1: Grades K to 2, levels A to N, 28 books (14 fiction and 14 nonfiction), plus an assessment guide, recording forms, and professional development videos.
• System 2: Grades 3 to 8, levels L to Z, 30 books (15 fiction and 15 nonfiction), with the same supporting materials. Heinemann lists it as an institutional purchase only, with a list price of about $700.
In April 2024 the publisher announced a successor, Benchmark Assessment System 2.0, and in September 2024 reported it was in stock. BAS 2.0 keeps the one-on-one conference but widens the overlap between kits: System 1 now covers levels A to P and System 2 covers levels K to Z. Heinemann's product pages describe the grade bands slightly differently (System 1 through second grade on one page, through third on another), so the level ranges are the safer thing to quote. BAS 2.0 adds new texts with Lexile measures assigned by MetaMetrics, redesigned recording forms, and a digital subscription for the guide, rubrics, and forms. A parallel Spanish system, Sistema de evaluación de la lectura 2.0, covers levels A to Z.
How a BAS conference works in practice
The publisher's FAQ gives realistic timing: a full conference "may take 20 to 30 minutes" at the earliest levels and 30 to 40 minutes at the upper levels, per student. Heinemann suggests administering it at the beginning of the year, optionally at mid-year, and again near the end, and a short "Where-to-Start Word Test" gives a rough starting level.
Texts are leveled using ten characteristics, from genre and text structure to sentence complexity and print features. Heinemann calls the BAS a standardized assessment because administration, coding, and scoring follow fixed procedures. In testing language, though, "standardized" usually also implies a norm sample and derived scores such as percentiles. The BAS provides neither; its output is a letter level that a school compares against expectations it sets itself.
What the publisher's research shows
The only published validity study of the BAS is the publisher's field study of the second edition, conducted with 498 students in 22 schools across five U.S. regions. The executive summary reports that 84 percent of students read the fiction books in sequential order of difficulty within one level of their target, 85 percent did so for nonfiction, and 76 percent read the fiction and nonfiction books at similar levels.
The summary labels the correlation between a student's fiction-series level and nonfiction-series level as "test-retest reliability," reporting .93 for levels A to N, .94 for levels L to Z, and .97 across all levels. Readers with a psychometrics background will notice that this compares two different books rather than the same test given twice, which is closer to "alternate-form reliability," and that pooling students across the whole K-8 range inflates the overall coefficient. For validity, the summary reports correlations of .94 and .93 between System 1 and the Reading Recovery text-level assessment, and moderate correlations of .69 and .62 between System 2 and the Slosson word-reading test. Reading Recovery levels come from the same leveled-text tradition as the BAS, so that strong correlation says less than an independent measure would. The Buros Center for Testing lists the BAS as never reviewed in the Mental Measurements Yearbook as of its listing.
The science of reading criticism and restrictions, 2023 to 2025
Independent research has been less favorable. A 2015 peer-reviewed study in Reading & Writing Quarterly compared one-minute oral reading fluency with an informal reading inventory as screeners for 968 second and third graders in a rural Minnesota district. Oral reading fluency correctly classified 80 percent of students as at risk or not; the reading inventory correctly classified 54 percent. APM Reports, which identified that inventory as the BAS, reported that it caught only 31 percent of the struggling readers. A companion study in the Journal of School Psychology found that when 64 second and third graders read three books rated at their instructional level, the readings agreed only about 67 to 70 percent of the time, and more than half of the weakest readers read at a frustration level on books supposed to fit them.
In December 2023, APM Reports published an investigation concluding that the BAS is "widely used and often wrong," estimating that it is used in about one in six American elementary schools, and reporting that San Francisco Unified, Fort Worth, Baltimore County, and Nashua had dropped it as a district-wide assessment. Heinemann declined to answer questions, and its attorney called the 2015 study "limited and flawed."
The broader controversy concerns "three-cueing," the idea that beginning readers should use meaning, sentence structure, and visual information together to identify words; the BAS reading record codes errors using those three sources. In 2021 Fountas and Pinnell published a blog series defending the approach, and cognitive scientist Mark Seidenberg told APM Reports the position "doesn't square with what decades of scientific research has shown." The publisher's October 2022 "Get the Facts" post responded that its resources teach children to decode, include explicit and systematic phonics, and are "not Whole Language."
Policy has moved quickly. APM Reports counted at least 26 states that passed reading-instruction laws between late 2022 and October 2025, and at least 15 that banned cueing from parts of their education systems, following earlier bans in Arkansas, Louisiana, and Virginia. Ohio's approved list of K-3 reading diagnostic assessments for the 2026-2027 school year names six products, and the BAS is not among them. In December 2024, two Massachusetts parents filed a lawsuit against Fountas, Pinnell, Lucy Calkins, and their publishers alleging deceptive marketing of products described as research-backed. That filing is an allegation, not a finding.
How the BAS differs from an IQ test
The BAS is a curriculum-linked reading assessment: its content is a set of books, its output is a text level, and its purpose is placement and progress monitoring. An adaptive achievement test such as the MAP test adds a scale score and national percentile. Neither measures general cognitive ability.
An IQ test samples reasoning, memory, spatial ability, and processing efficiency across many item types, is normed on a representative sample, and reports a standard score with a mean of 100 and a standard deviation of 15, with a confidence interval that reflects measurement error. A child can hold a strong BAS level and an average IQ, or the reverse, because reading skill and general reasoning ability overlap without being the same thing.
For adults, the Reasoning and Intelligence Online Test, developed by RIOT IQ with psychometrician Dr. Russell T. Warne, is an example of the second category. The RIOT IQ test is for adults 18 and over, contains 15 subtests across six cognitive indices (verbal reasoning, fluid reasoning, spatial ability, working memory, processing speed, and reaction time), takes about 52 minutes, and reports scores on the mean-100, standard-deviation-15 scale. It does not replace an individually administered diagnostic evaluation. Anyone who wants to see how a normed reasoning test is built can take the RIOT IQ test at riotiq.com.
Frequently asked questions
Is the Fountas & Pinnell Benchmark Assessment an IQ test?
No. It is a leveled reading assessment that places a student on the A to Z text gradient for instructional purposes. It has no national norms or percentiles and does not measure reasoning or general cognitive ability.
Can the BAS diagnose dyslexia?
No. The publisher describes it as a formative assessment for informing instruction. Diagnosis of any reading disorder requires a comprehensive evaluation by a qualified professional, and independent research has found the BAS misses many struggling readers when used as a screener.
What do the BAS levels A to Z mean?
Each letter is a step on the Fountas & Pinnell text gradient, with A the easiest and Z the hardest. A student's level is the hardest text read with acceptable accuracy and comprehension.
What is the difference between BAS Third Edition and BAS 2.0?
The Third Edition (2016) has System 1 for levels A to N and System 2 for levels L to Z. BAS 2.0, in stock since September 2024, widens the overlap to A to P and K to Z, adds new texts with Lexile measures, redesigns the recording forms, and moves the guide and forms to a digital subscription.
Why do some states and districts restrict the BAS?
The system rests on the three-cueing model of word reading, which cognitive science has not supported, and independent studies have found low classification accuracy. Since 2023 many states have banned cueing-based instruction and narrowed approved screener lists, and several large districts have dropped the BAS as a district-wide assessment.
References
1. Fountas & Pinnell Literacy. (2024). What is Benchmark Assessment System (BAS) and how is BAS used? Heinemann. fountasandpinnell.com
3. Fountas, I. C., & Pinnell, G. S. (2012). Benchmark Assessment System 2nd edition: Executive summary of the field study of reliability and validity. Heinemann. fountasandpinnell.com
4. Fountas & Pinnell Literacy. (n.d.). BAS research. Heinemann. fountasandpinnell.com
5. Parker, D. C., Zaslofsky, A. F., Burns, M. K., Kanive, R., Hodgson, J., Scholin, S. E., & Klingbeil, D. A. (2015). A brief report of the diagnostic accuracy of oral reading fluency and reading inventory levels for reading failure risk among second- and third-grade students. Reading & Writing Quarterly, 31(1), 56-67. doi.org
6. Burns, M. K., Pulles, S. M., Maki, K. E., Kanive, R., Hodgson, J., Helman, L. A., McComas, J. J., & Preast, J. L. (2015). Accuracy of student performance while reading leveled books rated at their instructional level by a reading inventory. Journal of School Psychology, 53(6), 437-445. doi.org
7. Peak, C. (2023, December 11). Benchmark Assessment System reading test is widely used and often wrong. APM Reports. apmreports.org
8. Peak, C. (2025, October 16). New reading laws sweep the nation following Sold a Story. APM Reports. apmreports.org
9. Schwartz, S. (2024, December 6). Here's what happens next on the Calkins, Fountas & Pinnell curriculum lawsuit. Education Week. edweek.org
Hero photo: A child reading a picture book, seen from behind. Photo by Library of Congress Life (Flickr), via Wikimedia Commons, CC0 (cropped).
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.