What is an executive function test? What it measures and what it misses
An executive function test measures the mental control skills behind self-restraint and flexible thinking, using either timed tasks or behavior rating scales.
Dr. Russell T. WarneChief Scientist
Share
An executive function test is a standardized measure of the mental control processes that let a person hold information in mind, override a habitual response and switch flexibly between rules or tasks. It comes in two very different forms, timed performance tasks done in front of an examiner and rating scales completed by a parent, teacher or the person themselves, and the two forms agree with each other far less than most people expect.
This page is a hub. It explains what psychologists mean by "executive function", why the two kinds of test diverge, which batteries are in common use, how executive function relates to IQ, and the measurement problem that makes any single executive score hard to interpret. Individual instruments have their own pages, linked below, so this one stays with the concepts that apply to all of them.
What "executive function" means: one ability or several?
"Executive function" (EF) is an umbrella term for the processes that regulate other cognitive processes in the service of a goal. Diamond's review in the Annual Review of Psychology identifies three core components, and most current tests are built around some version of them.
• Inhibition: Holding back a strong or automatic response, and screening out distraction, so that a less habitual but more appropriate response can win.
• Working memory: Keeping information active and updating it as new information arrives.
• Cognitive flexibility: Shifting between mental sets, rules or perspectives when circumstances change.
The most influential evidence about how these components fit together comes from Miyake and colleagues' 2000 study of 137 college students. Using "latent variable analysis", a statistical method that extracts what several tasks have in common and discards what is specific to each, they found that shifting, updating and inhibition were moderately correlated with one another but clearly separable. They called this pattern the "unity and diversity" of executive functions. The same study showed that classic clinical tasks draw on the components unevenly: performance on the Wisconsin Card Sorting Test related most strongly to shifting, and operation span related most strongly to updating.
Later work refined the picture. Friedman and Miyake's 2017 review describes a model in which a "Common EF" factor runs through every executive task, with additional factors specific to updating and to shifting. In that model no separate inhibition-specific factor remains once the common factor is accounted for, which suggests that much of what tests label "inhibition" is the shared core of executive control.
Performance tests and rating scales measure different things
The first form of executive function test is a "performance-based" task. The person sorts cards by a rule they must infer, names ink colors while ignoring the printed word, connects numbered and lettered circles in alternation, or plans moves on a tower puzzle, all under standardized conditions with an examiner timing and scoring. The Wisconsin Card Sorting Test and the Stroop test are familiar examples, and the Trail Making Test is another.
The second form is a "rating scale", a questionnaire on which someone who knows the person well reports how often everyday executive problems occur: losing track of tasks, acting without thinking, struggling to start homework. The BRIEF-2 is a widely used rating scale of this kind for children and adolescents.
Both forms carry the same label, so it is natural to assume they measure the same thing. Toplak, West and Stanovich tested that assumption directly. They gathered 20 studies (13 child and 7 adult samples) that had given both kinds of measure to the same people. Only 68 of the 286 relevant correlations, about 24 percent, were statistically significant, and the median correlation was .19. A correlation that size means the two methods share only a few percent of their variance.
Their interpretation is the most useful single idea for anyone reading an EF report. Performance tests capture the efficiency of cognitive processing under optimal, structured conditions, with an examiner present, instructions explicit and distractions removed. Rating scales capture how successfully a person pursues goals in ordinary life, where nobody supplies the structure. A child can do well on a card-sorting task in a quiet room and still be rated as highly disorganized at home, and both results can be accurate.
The practical consequence is that a normal score on one form does not rule out a problem on the other. An earlier study led by Toplak gave both kinds of measure to adolescents with and without ADHD. Parent and teacher ratings predicted ADHD status, while the performance tasks accounted for little unique variance and overlapped little with the ratings.
The main executive function batteries
Most clinicians assemble EF evidence from a battery plus a rating scale. Three batteries account for much of the published research.
• D-KEFS: The Delis-Kaplan Executive Function System consists of nine stand-alone tests normed for ages 8 to 89, and Homack, Lee and Riccio describe it as the first set of executive tests co-normed on a large, representative national sample. Its design compares performance across conditions of the same task, so an examiner can see whether a slow switching score reflects switching itself or slow basic processing.
• NEPSY-II: A developmental neuropsychological battery for children aged 3 to 16 in which attention and executive functioning is one of several domains. Brooks, Sherman and Strauss's review covers its structure and psychometric properties.
• CANTAB: The Cambridge Neuropsychological Test Automated Battery is a computerized, touchscreen battery whose tests were adapted from methods used to study learning and memory in non-human primates. Robbins and colleagues' factor analysis of 787 healthy adults aged 55 to 80 is an early large-sample description of its structure.
Many single-construct tasks sit outside these batteries, and several have pages of their own on this site.
How executive function relates to IQ
Executive function and IQ overlap substantially without being the same thing, and the overlap is not spread evenly across the components.
Friedman and colleagues examined this in 2006 by measuring inhibiting, shifting and updating in young adults alongside fluid intelligence, crystallized intelligence and Wechsler IQ. Updating was highly correlated with every intelligence measure. Inhibiting and shifting were not, and once the correlations among the three EFs were controlled, their relations to intelligence were small and not significant. The authors concluded that standard intelligence tests do not equally assess the full range of executive control.
In the same research program, summarized in Friedman and Miyake's 2017 review, full-scale intelligence correlated about .51 with the Common EF factor and .49 with the updating-specific factor. The authors estimated that Common EF shares only about 25 percent of its variance with the general factor of intelligence, so the two are related but distinct. In children the overlap appears larger. Freis and colleagues analyzed over 11,000 nine- and ten-year-olds in the Adolescent Brain Cognitive Development Study and found that Common EF and updating-specific factors correlated with IQ at .64 to .81, with a genetic correlation between Common EF and IQ of .86.
The working-memory link is the part most relevant to IQ testing. Updating is close to what working-memory subtests on intelligence batteries measure, which is one reason those subtests sit comfortably inside an IQ score while a task like Stroop interference does not.
Task impurity: why one score rarely isolates one function
Every executive function test has a built-in measurement problem. By definition, EFs control lower-level processes, so any task designed to measure an EF must also involve the non-executive processes being controlled. Friedman and Miyake call this "task impurity" and describe it as an unavoidable quality of EF tasks. A slow score on a color-word interference task may reflect weak inhibition, slow naming speed, a reading problem or poor color vision, and the score alone cannot say which.
Reliability compounds the problem. Hedge, Powell and Sumner measured the test-retest reliability of seven classic cognitive tasks, including Stroop, flanker, stop-signal and go/no-go paradigms, and found values ranging from 0 to .82, surprisingly low for most tasks given how often they are used. The reason is instructive. These tasks became popular because their effects are robust, which in practice means people differ little from one another on them, and low between-person variability is exactly what makes a measure unreliable for comparing individuals.
Researchers handle task impurity by giving several tasks per construct and extracting what they share. Snyder, Miyake and Hankin call task impurity arguably the most vexing problem in measuring EF and recommend, for clinical research, several tasks per component that look different on the surface, combined so that the variance unrelated to the construct partly cancels out. Even a simple average of standardized scores across tasks helps. The same logic applies to reading an individual report: a pattern across tasks says more than one low number.
Frequently asked questions
What does an executive function test measure?
It measures the control processes that manage other thinking: inhibiting automatic responses, holding and updating information in working memory, and shifting between rules or tasks. Performance tests measure these under structured conditions, while rating scales measure how they show up in daily life.
Is an executive function test the same as an IQ test?
No. The two overlap, mostly through working memory and updating, but measures of inhibition and shifting relate only weakly to IQ once their shared variance is removed. A person can have an average IQ and marked executive difficulties, or the reverse.
Why did my child score normally on the tests but poorly on the questionnaire?
Because the two methods measure different things. Across 20 studies, performance tests and EF rating scales correlated at a median of only .19, so disagreement between them is common and usually informative.
Can an executive function test diagnose ADHD?
No single EF test diagnoses ADHD. Clinicians weigh EF measures alongside developmental history, symptom ratings and observation, and in research comparing the two methods, rating scales have predicted ADHD status better than performance tasks.
Who gives executive function tests?
Standardized EF batteries are given and interpreted by psychologists, most often neuropsychologists, clinical psychologists and school psychologists. If you have a concern, start with a psychologist who has training in neuropsychological assessment.
The takeaway
An executive function test measures goal-directed control, a family of related but separable abilities with a common core. The label covers two methods that correlate weakly: timed tasks show how efficiently someone can exercise control when the situation is structured for them, and rating scales show how well they manage when it is not. Executive function overlaps with IQ mainly through working memory and updating, and every EF task mixes the target function with other processes. Read any EF result as one piece of a pattern, and ask which method produced it.
If you want to see how a well-normed measure of reasoning reports its results, you can take an online IQ test built by psychometricians, the Reasoning and Intelligence Online Test, which reports index scores with confidence intervals. It measures intelligence and is not an executive function assessment, so treat it as a separate kind of evidence from the evaluations described here.
References
1. Miyake, A., Friedman, N. P., Emerson, M. J., Witzki, A. H., Howerter, A., & Wager, T. D. (2000). The unity and diversity of executive functions and their contributions to complex "frontal lobe" tasks: A latent variable analysis. Cognitive Psychology, 41(1), 49-100. doi.org
2. Diamond, A. (2013). Executive functions. Annual Review of Psychology, 64, 135-168. doi.org
3. Friedman, N. P., & Miyake, A. (2017). Unity and diversity of executive functions: Individual differences as a window on cognitive structure. Cortex, 86, 186-204. doi.org
4. Toplak, M. E., West, R. F., & Stanovich, K. E. (2013). Practitioner review: Do performance-based measures and ratings of executive function assess the same construct? Journal of Child Psychology and Psychiatry, 54(2), 131-143. doi.org
5. Toplak, M. E., Bucciarelli, S. M., Jain, U., & Tannock, R. (2009). Executive functions: Performance-based measures and the Behavior Rating Inventory of Executive Function (BRIEF) in adolescents with attention deficit/hyperactivity disorder (ADHD). Child Neuropsychology, 15(1), 53-72. doi.org
6. Homack, S., Lee, D., & Riccio, C. A. (2005). Test review: Delis-Kaplan Executive Function System. Journal of Clinical and Experimental Neuropsychology, 27(5), 599-609. doi.org
7. Brooks, B. L., Sherman, E. M. S., & Strauss, E. (2010). NEPSY-II: A developmental neuropsychological assessment, second edition. Child Neuropsychology, 16(1), 80-101. doi.org
8. Robbins, T. W., James, M., Owen, A. M., Sahakian, B. J., McInnes, L., & Rabbitt, P. (1994). Cambridge Neuropsychological Test Automated Battery (CANTAB): A factor analytic study of a large sample of normal elderly volunteers. Dementia, 5(5), 266-281. doi.org
9. Friedman, N. P., Miyake, A., Corley, R. P., Young, S. E., DeFries, J. C., & Hewitt, J. K. (2006). Not all executive functions are related to intelligence. Psychological Science, 17(2), 172-179. doi.org
10. Freis, S. M., Morrison, C. L., Lessem, J. M., Hewitt, J. K., & Friedman, N. P. (2022). Genetic and environmental influences on executive functions and intelligence in middle childhood. Developmental Science, 25(1), e13150. doi.org
11. Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166-1186. doi.org
12. Snyder, H. R., Miyake, A., & Hankin, B. L. (2015). Advancing understanding of executive function impairments and psychopathology: Bridging the gap between clinical and cognitive approaches. Frontiers in Psychology, 6, 328. doi.org
Hero image: playing chess in the Jardin du Luxembourg, Paris, by Jorge Royan, licensed CC BY-SA 3.0 (creativecommons.org/licenses/by-sa/3.0). Cropped. Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.