Computer adaptive testing: how a test that rebuilds itself as you go works
Computer adaptive testing selects each question to match a test taker's current estimated ability, reaching the same precision in far fewer items.
Dr. Russell T. WarneChief Scientist
Share
Computer adaptive testing, usually shortened to CAT, is a method of test delivery in which a computer chooses each question from a calibrated pool based on how the test taker answered the previous ones, revising an estimate of ability after every response. Questions far too easy or far too hard for a person carry almost no information about that person, so a test that keeps its questions near the test taker's current ability level can reach the same precision as a fixed paper test while administering far fewer items.
This page covers the mechanics: how the next item gets chosen, how the test decides to stop, how publishers stop their best items wearing out, how content coverage is protected, and where adaptive delivery is used in cognitive and achievement testing.
Why item response theory makes adaptivity possible
Adaptive delivery rests on "item response theory" (IRT), a family of models that place item difficulty and person ability on one common scale. The property doing the work is parameter invariance: an item's estimated difficulty does not depend on which people happened to answer it, and a person's estimated ability does not depend on which items they happened to receive. NWEA's technical documentation for MAP Growth calls parameter invariance "perhaps the reason why computer adaptive testing is possible," since two students can take entirely different item sets and still be scored on the same scale (NWEA, 2026). A separate article on this site covers item response theory itself; one consequence is enough here. Under IRT a test is a pool of calibrated items plus rules for selecting a subset, and the subset need not be the same for everyone. The model varies by test: CAT-ASVAB, the computerized Armed Services Vocational Aptitude Battery, uses the three-parameter logistic model, describing an item by its difficulty, its discrimination and a lower asymptote for guessing, while MAP Growth uses the Rasch model (Defense Manpower Data Center, 2006; NWEA, 2026).
How the test chooses each question
The loop is short and repeats after every answer. CAT-ASVAB documents it in eight steps, and most operational adaptive tests follow a version of the same cycle. The four details below come from that battery's technical bulletin (Defense Manpower Data Center, 2006).
• Set a starting estimate: The provisional ability estimate, written in the literature as theta-hat and meaning no more than "our current best guess at this person's standing," begins at the mean of the assumed ability distribution. CAT-ASVAB sets it to zero on a scale with mean 0 and standard deviation 1, so the opening question suits a person of average ability.
• Select the most informative available item: Every item has an "information function" that peaks near its own difficulty and falls away on either side, and the rule takes the item with the largest information value at the current estimate. CAT-ASVAB precomputes this as an information table, ranking every item by information at each of 37 ability levels spaced evenly from -2.25 to +2.25.
• Update after every response: The answer is scored and the estimate revised with a computationally cheap sequential Bayesian procedure, with a final estimate computed once the test ends.
• Check the stopping rule: CAT-ASVAB is fixed length, ending a subtest after a set number of items or at its time limit. Its stated reasoning is that highly informative items cluster in a limited difficulty range, so under a precision-based rule people at the extremes receive long tests in which each extra item adds little.
The Standards for Educational and Psychological Testing describe the general case: adaptive test length "is determined by stopping rules, which may be based on a fixed number of test questions or may be based on a desired level of score precision" (AERA, APA, & NCME, 2014, p. 79). The nursing licensure examination NCLEX takes the precision route, running until it can conclude with 95 percent confidence that the candidate sits above or below the passing standard, with a floor of 85 items and a ceiling of 150 (NCSBN, 2026). Any precision-based rule is expressed in terms of the standard error of measurement around the score.
Exposure control and content balancing
Left alone, maximum-information selection creates two problems, and operational systems solve both explicitly.
• Exposure control: If everyone starts at the same estimate, the single most informative item is the opening question for every examinee, the second is one of only two possibilities, and the early sequence becomes predictable and heavily used. CAT-ASVAB's solution is the probabilistic algorithm developed by Sympson and Hetter: each item carries a control parameter, and once an item is selected the system draws a random number and administers the item only if the parameter clears that number. Its target rates were set to match the paper battery's, one third for the tests feeding the qualification composite (Defense Manpower Data Center, 2006).
• Content balancing: Precision is not the only requirement, because the items a person receives still have to represent the intended content. CAT-ASVAB's General Science test carries a fixed allocation vector that alternates Life Science and Physical Science items and inserts exactly one Chemistry item, drawing each from its own information table (Defense Manpower Data Center, 2006). MAP Growth uses Constrained CAT (Kingsbury & Zara, 1989), which partitions the pool by content category, finds a category short of its target count, and takes the most informative item from there, within a cap of 43 items per test event (NWEA, 2026). This is the blueprint logic that governs fixed forms when an IQ test is created, applied one question at a time.
Standard 4.3 of the Standards makes both of these documentation obligations, covering item selection, the starting point, the termination conditions and exposure control (AERA, APA, & NCME, 2014, p. 86).
Why a CAT reaches the same precision in fewer items
Stated in words: total test information is the sum of what each administered item contributes, and the standard error of the ability estimate is one divided by the square root of that total. An item whose difficulty sits far from a person's ability contributes information close to zero, lengthening the test without shrinking the error bar.
Published CAT-ASVAB simulations show the size of the effect on real aptitude content. Measured against paper form ASVAB-9A, a 15-item adaptive Word Knowledge test reached a reliability of about .91 where the 35-item paper test reached about .90, a 15-item adaptive Mathematics Knowledge test reached about .93 against .85 for 25 paper items, and a 10-item adaptive Paragraph Comprehension test reached about .85 against .76 for 15 paper items (Defense Manpower Data Center, 2006). The pattern holds at battery level today, with 135 questions on the computer version against 225 on the paper form, and the ASVAB program states that adaptive selection "results in higher levels of test-score precision and shorter test lengths than the P&P-ASVAB" (Official ASVAB Program, 2026a, 2026b). The Standards note that the precision gain is largest for high-scoring and low-scoring examinees (AERA, APA, & NCME, 2014, p. 86), which is where a fixed form runs out of suitable items.
Where adaptive testing is actually used
• Military aptitude selection: CAT-ASVAB is the longest-running large-scale example. Joint-Service development began in 1979, operational evaluation started in June 1992, and the field system was replaced by a next-generation version in 1996 (Defense Manpower Data Center, 2006).
• School achievement and growth measurement: MAP Growth is an interim adaptive test given several times a year in mathematics, reading, language usage and science, reported on the Rasch-based RIT scale, with pools deep enough that a student rarely meets the same item twice (NWEA, 2026).
• Graduate admissions, at the section level: The GRE General Test needs a precise description. Its Verbal Reasoning and Quantitative Reasoning measures are section-level adaptive: the first section of each is of average difficulty, and the difficulty of the second depends on performance on the first (ETS, 2026). That is multistage adaptive testing rather than item-by-item adaptation.
• Individually administered intelligence tests, without a computer: Tailoring predates the hardware. On the Stanford-Binet Intelligence Scales, Fifth Edition, the examiner opens with two routing subtests, Object Series/Matrices and Vocabulary, then uses those raw scores to decide which of six difficulty levels of the remaining subtests to give (Roid & Barram, 2004). The logic is a CAT's, run by a person with a conversion table.
What adaptive testing does not fix
Adaptive delivery buys efficiency with infrastructure. It needs a calibrated pool large enough that everyone receives appropriate items while the content blueprint is still met, which the Standards note usually means large numbers of items and often several separate pools (AERA, APA, & NCME, 2014, p. 81). It generally prevents a test taker from revisiting earlier questions, since the later ones were chosen on the assumption that those answers stand. It also does nothing for item quality or for the adequacy of the norm sample. The useful question about any test remains whether its item parameters, its norms and its reported precision are documented. Readers weighing the formats an IQ test can take can take the Reasoning and Intelligence Online Test, an online IQ test built by psychometricians, and read its published methodology alongside the score.
Frequently asked questions
Is a computer adaptive test harder than a paper test?
It tends to feel harder, because the algorithm converges on items the test taker has a moderate chance of answering correctly. Item difficulty is part of the scoring model, so nothing is lost by that.
Does getting a question wrong mean the next one will be easier?
Broadly yes, though not always immediately. Selection responds to the updated ability estimate rather than the last answer alone, and content balancing or exposure control can override the statistical choice on any given question.
How many questions will an adaptive test ask?
That depends on the stopping rule. A fixed-length adaptive test asks a set number, as CAT-ASVAB does with 15 items on several subtests. A precision-based test runs until the score is precise enough, which is why NCLEX candidates answer between 85 and 150 items.
Is the GRE a computer adaptive test?
It is adaptive at the section level rather than the item level. ETS documents that the second Verbal and Quantitative sections are selected on the basis of performance on the first, which makes the current GRE a multistage adaptive test.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Defense Manpower Data Center, Personnel Testing Division. (2006). ASVAB technical bulletin no. 1: CAT-ASVAB forms 1 & 2. Department of Defense. officialasvab.com
3. Official ASVAB Program. (2026a). The CAT-ASVAB. United States Military Entrance Processing Command. officialasvab.com
4. Official ASVAB Program. (2026b). What to expect when you take the ASVAB. United States Military Entrance Processing Command. officialasvab.com
6. Educational Testing Service. (2026). GRE General Test structure. ETS. ets.org
7. National Council of State Boards of Nursing. (2026). How the NCLEX works: Frequently asked questions. NCSBN. ncsbn.org
8. Weiss, D. J. (1982). Improving measurement quality and efficiency with adaptive testing. Applied Psychological Measurement, 6, 473-492. doi.org
9. Kingsbury, G. G., & Zara, A. R. (1989). Procedures for selecting items for computerized adaptive tests. Applied Measurement in Education, 2, 359-375. doi.org
10. Roid, G. H., & Barram, R. A. (2004). Essentials of Stanford-Binet Intelligence Scales (SB5) assessment [Chapter 1 excerpt]. Wiley. catalogimages.wiley.com
Hero image: clinician with a phoropter, by Ddcmis, licensed CC BY-SA 4.0 (creativecommons.org/licenses/by-sa/4.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.