Differential item functioning: how item bias is actually detected
Differential item functioning means equally able test takers from different groups answer an item correctly at different rates. Here is how it is found.
Dr. Russell T. WarneChief Scientist
Share
"Differential item functioning" (DIF) is a statistical property of a single test question: it occurs when test takers who stand at the same level on the ability the test measures, but belong to different groups, answer that question correctly at different rates. The Standards for Educational and Psychological Testing call it "a statistical indicator of the extent to which different groups of test takers who are at the same ability level have different frequencies of correct responses" (AERA, APA, & NCME, 2014, p. 217). DIF analysis is how item bias in a cognitive or achievement test gets detected in the first place.
This page is about that machinery. Whether intelligence tests are biased is argued elsewhere on this site, including in the article on whether IQ tests are biased. The question here comes first: how would anyone actually know?
DIF is not the same thing as impact
"Impact" is the raw, unconditional difference. Suppose one group answers an item correctly 62 percent of the time and another group 48 percent of the time; that gap is impact. DIF is what remains after the two groups have been matched on the ability the test is measuring. Michael Zieky's primer for Educational Testing Service states the reasoning behind that step: groups matched on the relevant knowledge and skill can generally be expected to perform similarly on a test question (Zieky, 2003, p. 1).
The distinction matters because impact is not evidence of anything by itself. The Standards put it as a rule for developers: "Subgroup mean differences do not in and of themselves indicate lack of fairness, but such differences should trigger follow-up studies, where feasible, to identify the potential causes of such differences" (AERA, APA, & NCME, 2014, p. 65). An item on which two groups differ can be perfectly well behaved if the groups also differ on the underlying ability, and an item on which they do not differ can still carry DIF if one group's average ability is higher.
Matching on ability is the whole trick, and it is imperfect
In practice the matching variable is almost always the test taker's own score on the test or subtest under study. Zieky is candid about why that is uncomfortable: the score used to define "equal ability" came from the very instrument under suspicion. Two partial answers are standard.
• Criterion refinement: ETS runs a preliminary DIF analysis, removes the items with elevated DIF values from the score used for matching, then recomputes the DIF statistics on that cleaned matching variable (Zieky, 2003, p. 2).
• Multidimensionality, which breaks the matching: If a mathematics test is mostly algebra with a few geometry questions, matching on total score matches people largely on algebra, and the geometry items may show DIF for that reason alone (Zieky, 2003, p. 3). The Standards treat this as an ordinary finding rather than a defect, stating that "differential item functioning is not always a flaw or weakness" because items sharing a characteristic can legitimately function differently for equally scoring test takers (AERA, APA, & NCME, 2014, p. 16).
Sample size is the other constraint. ETS requires at least 200 members of the smaller group and 500 in total at the test-assembly stage, rising to 300 and 700 after an administration, because smaller samples give unstable results (Zwick, 2012).
Uniform and non-uniform DIF
Two patterns are distinguished, and they call for different statistics.
• Uniform DIF: One group is at a disadvantage on the item by roughly the same amount at every ability level. In item response theory terms the two groups' item characteristic curves are separated by a difference in difficulty and do not cross.
• Non-uniform DIF: The size or even the direction of the disadvantage changes with ability, so the two curves cross. This is an interaction between group membership and ability, and a method that only tests for an average difference can miss it entirely.
Both occur in real cognitive tests. Maller (2001) analysed six subtests from the national standardization sample of the Wechsler Intelligence Scale for Children, Third Edition (n = 2,200) for DIF between boys and girls, and found both patterns, affecting about a third of the items studied.
The four families of detection method
• Mantel-Haenszel: The most widely used approach, brought into DIF work by Holland and Thayer (1986). Test takers are sorted into score levels on the matching variable, within each level a two-by-two table is formed of group against right or wrong, and the odds ratios from those tables are pooled into one estimate. ETS expresses that estimate on its delta scale of item difficulty as the statistic MH D-DIF, which in words is minus 2.35 times the natural logarithm of the pooled odds ratio (Zwick, 2012). Items are then classified A, B or C. Zieky states the rule precisely: category A is items whose MH D-DIF is not significantly different from zero or is below 1.0 in absolute value, category C is items whose MH D-DIF is significantly greater than 1.0 and at least 1.5 in absolute value, and category B is everything else (Zieky, 2003, p. 4).
• Logistic regression: Swaminathan and Rogers (1990) proposed predicting the item response from the matching score, then adding group membership, then adding the product of group and score. A significant gain from the group term indicates uniform DIF, and a significant gain from the product term indicates non-uniform DIF. One framework handles both patterns, and the matching variable can be continuous.
• Item response theory approaches: The item's parameters are estimated separately in each group and compared. Raju (1988) derived the exact area between two item characteristic curves as a measure of how far apart they lie, with signed and unsigned versions that distinguish uniform from non-uniform DIF. Related procedures test the equality of item parameters directly, or compare the fit of models that do and do not let parameters differ by group.
• Nonparametric and bundle approaches: SIBTEST (Shealy & Stout, 1993) uses a latent-ability regression to separate DIF from genuine group differences in ability, and it can be applied to a bundle of related items rather than one at a time, which helps when the suspected cause is a shared feature such as passage topic or response format.
What test publishers do with a flag
The statistic starts a process rather than ending one. ETS test assemblers prefer category A items, may use category B when the blueprint requires it, and may use a category C item only where it is essential to the test specifications, with documented justification and review by people who did not work on the test, including external reviewers (Zieky, 2003, pp. 4-5).
Cognitive and achievement publishers report the same kind of screening. The Woodcock-Johnson IV technical manual documents DIF analyses across sex, race and ethnicity; most items showed no DIF, and most of those that did were kept out of the published forms (Canivez, 2017). Canivez notes a real limitation in that work: collapsing race into "White" and "non-White" can hide non-invariance that would show up in a specific group. NWEA reports that the vast majority of MAP Growth items show negligible DIF across gender and racial or ethnic groups (NWEA, 2026). Standard 4.10 requires the screening process itself to be documented (AERA, APA, & NCME, 2014, p. 88).
Why a DIF flag is not proof of bias
Zieky is blunt: "No statistic can determine whether or not a test question is biased" (Zieky, 2003, p. 3). The Standards agree: "The detection of DIF does not always indicate bias in an item; there needs to be a suitable, substantial explanation for the DIF to justify the conclusion that the item is biased" (AERA, APA, & NCME, 2014, p. 51). Whether an item is unfair depends on what the test is for. Zieky's example is a nursing licensure question about breast cancer that women find easier than matched men: fair on a test of what nurses must know, unfair on a test of general knowledge.
The statistics are also less decisive than their reputation suggests. Zwick (2012) found that the ETS category C rule "often displays low DIF detection rates even when samples are large," so an absence of flags is weak evidence of an absence of DIF. That is why the arguments in the articles on whether intelligence tests are biased against diverse populations and on whether IQ tests are racist rest on accumulated item-level and structural evidence rather than on any single coefficient. Anyone evaluating a test is entitled to ask which DIF method was used, against which matching variable, in which groups, at what sample size, and what became of the flagged items. The Reasoning and Intelligence Online Test publishes its methodology for that kind of reading, and can be taken as a professionally developed IQ test with documented item analysis.
Frequently asked questions
What is the difference between DIF and item bias?
DIF is a statistical finding: equally able members of two groups answer an item differently. Bias is a judgement that the difference comes from something irrelevant to what the item is meant to measure. A biased item should show DIF, while an item showing DIF may still be perfectly fair.
How large does DIF have to be before anyone acts?
There is no universal threshold. The best known convention is the ETS classification, which places an item in category C when its MH D-DIF statistic is significantly greater than 1.0 and at least 1.5 in absolute value on the delta scale of item difficulty.
Can DIF be tested on an individually administered IQ test?
Yes, where the standardization sample is large enough. Maller's analysis of the Wechsler Intelligence Scale for Children, Third Edition used its 2,200-case national standardization sample, and the Woodcock-Johnson IV technical manual reports DIF screening across sex, race and ethnicity.
Does a clean DIF analysis prove a test is fair?
No. Fairness in the Standards covers far more than item-level statistics, including access, administration conditions and the validity of score interpretations for each group. A clean DIF analysis is one piece of evidence, and conservative flagging rules mean some real DIF goes undetected.
Which DIF method is best?
They answer slightly different questions. Mantel-Haenszel is cheap and needs no fitted model, logistic regression handles non-uniform DIF, item response theory methods describe the whole difference between two item characteristic curves, and SIBTEST can test a bundle of items at once.
References
1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net
2. Zieky, M. (2003). A DIF primer. Educational Testing Service. praxis.ets.org
3. Zwick, R. (2012). A review of ETS differential item functioning assessment procedures: Flagging rules, minimum sample size requirements, and criterion refinement (Research Report RR-12-08). Educational Testing Service. files.eric.ed.gov
4. Holland, P. W., & Thayer, D. T. (1986). Differential item functioning and the Mantel-Haenszel procedure. ETS Research Report Series, 1986(2). doi.org
5. Swaminathan, H., & Rogers, H. J. (1990). Detecting differential item functioning using logistic regression procedures. Journal of Educational Measurement, 27, 361-370. doi.org
6. Raju, N. S. (1988). The area between two item characteristic curves. Psychometrika, 53, 495-502. doi.org
7. Shealy, R., & Stout, W. (1993). A model-based standardization approach that separates true bias/DIF from group ability differences and detects test bias/DTF as well as item bias/DIF. Psychometrika, 58, 159-194. doi.org
8. Maller, S. J. (2001). Differential item functioning in the WISC-III: Item parameters for boys and girls in the national standardization sample. Educational and Psychological Measurement, 61, 793-817. doi.org
9. Canivez, G. L. (2017). Test review of the Woodcock-Johnson IV. In Mental Measurements Yearbook. Buros Center for Testing. ux1.eiu.edu
Hero image: wheelchair access ramp over steps at the Hotel Montescot, Chartres, by Coyau, licensed CC BY-SA 3.0 (creativecommons.org/licenses/by-sa/3.0). Via Wikimedia Commons.
Take our professional IQ test
Want to know your IQ? Try the first ever professional online IQ test.