🇺🇸The official website of Riot IQ
Log in
  • Home
  • About

Measure your
intelligence online.

Google

Assessments

  • All IQ Tests
  • Basic IQ Test
  • Full IQ Test
  • Custom IQ Test
  • Free IQ Test

Our Socials

  • X
  • YouTube
  • Facebook
  • LinkedIn

Other IQ Tests

  • WAIS-V
  • SB-5
  • Raven's 2
  • RIAS-2
  • CogAT 9
  • WISC-V

Community

  • Join Subreddit
  • Join Discord

Other Pages

  • Test Manual
  • Administer IQ Tests
  • About Us
  • Articles
  • Data
  • FAQ

Research

  • What do polygenic scores really predict?
  • Working speed and ability on the RIOT

Intelligence Journals & Organizations

  • Human Intelligence Research & Education (HIRE) Foundation
  • International Society for Intelligence Research (ISIR)
  • Intelligence & Cognitive Abilities Journal (ICA)
  • Intelligence Journal
  • Mensa Foundation

Contact

  • Email
  • Support

News & Press

  • International Society for Intelligence Research
  • American Thinker
  • Mensa Foundation (1/2)
  • Mensa Northern New Jersey
  • The University of Western Australia
  • Prolific
  • Quillette
  • Brainz

Our Articles

  • Gmatclub (1/2)
  • ApolloTechnical
  • LessWrong
  • Psychreg
  • Study in Switzerland
  • SuccessConsciousness
  • Creative Organizational Design (1/2)
  • ABNewsWire
  • Vanderbilt University

Our Articles

  • The Globe and Mail
  • Barchart
  • Journal
  • Mensa Foundation (2/2)
  • Psychologs
  • Creative Organizational Design (2/2)
  • AZBigMedia
  • Thoughts on Life and Love
  • Anxiety and Depression Association of America

Our Articles

  • Before It's News
  • Siglo XXI
  • TechBullion
  • Medium
  • Gmatclub (2/2)
  • MSN
  • National Review
  • Minding the Campus
  • Launching Next

Our Articles

  • Comparing Cronbach’s Alpha and McDonald’s Omega Reliability
  • Breaking the Intelligence & IQ Taboo
  • What is the Flynn Effect?
  • A Comprehensive History of IQ Tests
  • The 15 Subtests of the RIOT

Our Articles

  • How to Take an IQ Test
  • How to Calculate IQ
  • What is the RIOT IQ Test?
  • The Pro-Human Aspects of Intelligence Research
  • What is an IQ Test? A Beginner's Guide.

Our Articles

  • 5 Best IQ Tests in 2025
  • Cognitive Profiles on the RIOT IQ Test Results
  • 6 Cognitive Abilities of the RIOT
  • Are There Any Professional and Real Online IQ Tests?

Our Articles

  • Resources to Learn About IQ and Intelligence
  • Studying IQ Matters
  • The Search for Albert Einstein's IQ
  • Do Non-g Gains from the Flynn Effect Matter?

Riot IQ © 2026

  • Terms of Service
  • Privacy Policy
  • BAA Agreement
  • Test Administrator Terms
  • Terms of Service
  • •Privacy Policy
  • •BAA Agreement
  • •Test Administrator Terms

Table of Contents

  • DIF is not the same thing as impact
  • Matching on ability is the whole trick, and it is imperfect
  • Uniform and non-uniform DIF
  • The four families of detection method
  • What test publishers do with a flag
  • Why a DIF flag is not proof of bias
  • Frequently asked questions
  • What is the difference between DIF and item bias?
  • How large does DIF have to be before anyone acts?
  • Can DIF be tested on an individually administered IQ test?
  • Does a clean DIF analysis prove a test is fair?
  • Which DIF method is best?
  • References
Sep 27, 2026·Accuracy, Reliability & Criticism

Differential item functioning: how item bias is actually detected

Differential item functioning means equally able test takers from different groups answer an item correctly at different rates. Here is how it is found.

Dr. Russell T. WarneChief Scientist
Share
Differential item functioning: how item bias is actually detected
"Differential item functioning" (DIF) is a statistical property of a single test question: it occurs when test takers who stand at the same level on the ability the test measures, but belong to different groups, answer that question correctly at different rates. The Standards for Educational and Psychological Testing call it "a statistical indicator of the extent to which different groups of test takers who are at the same ability level have different frequencies of correct responses" (AERA, APA, & NCME, 2014, p. 217). DIF analysis is how item bias in a cognitive or achievement test gets detected in the first place.

This page is about that machinery. Whether intelligence tests are biased is argued elsewhere on this site, including in the article on whether IQ tests are biased. The question here comes first: how would anyone actually know?


DIF is not the same thing as impact

"Impact" is the raw, unconditional difference. Suppose one group answers an item correctly 62 percent of the time and another group 48 percent of the time; that gap is impact. DIF is what remains after the two groups have been matched on the ability the test is measuring. Michael Zieky's primer for Educational Testing Service states the reasoning behind that step: groups matched on the relevant knowledge and skill can generally be expected to perform similarly on a test question (Zieky, 2003, p. 1).

The distinction matters because impact is not evidence of anything by itself. The Standards put it as a rule for developers: "Subgroup mean differences do not in and of themselves indicate lack of fairness, but such differences should trigger follow-up studies, where feasible, to identify the potential causes of such differences" (AERA, APA, & NCME, 2014, p. 65). An item on which two groups differ can be perfectly well behaved if the groups also differ on the underlying ability, and an item on which they do not differ can still carry DIF if one group's average ability is higher.


Matching on ability is the whole trick, and it is imperfect

In practice the matching variable is almost always the test taker's own score on the test or subtest under study. Zieky is candid about why that is uncomfortable: the score used to define "equal ability" came from the very instrument under suspicion. Two partial answers are standard.

• Criterion refinement: ETS runs a preliminary DIF analysis, removes the items with elevated DIF values from the score used for matching, then recomputes the DIF statistics on that cleaned matching variable (Zieky, 2003, p. 2).

• Multidimensionality, which breaks the matching: If a mathematics test is mostly algebra with a few geometry questions, matching on total score matches people largely on algebra, and the geometry items may show DIF for that reason alone (Zieky, 2003, p. 3). The Standards treat this as an ordinary finding rather than a defect, stating that "differential item functioning is not always a flaw or weakness" because items sharing a characteristic can legitimately function differently for equally scoring test takers (AERA, APA, & NCME, 2014, p. 16).

Sample size is the other constraint. ETS requires at least 200 members of the smaller group and 500 in total at the test-assembly stage, rising to 300 and 700 after an administration, because smaller samples give unstable results (Zwick, 2012).


Uniform and non-uniform DIF

Two patterns are distinguished, and they call for different statistics.

• Uniform DIF: One group is at a disadvantage on the item by roughly the same amount at every ability level. In item response theory terms the two groups' item characteristic curves are separated by a difference in difficulty and do not cross.

• Non-uniform DIF: The size or even the direction of the disadvantage changes with ability, so the two curves cross. This is an interaction between group membership and ability, and a method that only tests for an average difference can miss it entirely.

Both occur in real cognitive tests. Maller (2001) analysed six subtests from the national standardization sample of the Wechsler Intelligence Scale for Children, Third Edition (n = 2,200) for DIF between boys and girls, and found both patterns, affecting about a third of the items studied.


The four families of detection method

• Mantel-Haenszel: The most widely used approach, brought into DIF work by Holland and Thayer (1986). Test takers are sorted into score levels on the matching variable, within each level a two-by-two table is formed of group against right or wrong, and the odds ratios from those tables are pooled into one estimate. ETS expresses that estimate on its delta scale of item difficulty as the statistic MH D-DIF, which in words is minus 2.35 times the natural logarithm of the pooled odds ratio (Zwick, 2012). Items are then classified A, B or C. Zieky states the rule precisely: category A is items whose MH D-DIF is not significantly different from zero or is below 1.0 in absolute value, category C is items whose MH D-DIF is significantly greater than 1.0 and at least 1.5 in absolute value, and category B is everything else (Zieky, 2003, p. 4).

• Logistic regression: Swaminathan and Rogers (1990) proposed predicting the item response from the matching score, then adding group membership, then adding the product of group and score. A significant gain from the group term indicates uniform DIF, and a significant gain from the product term indicates non-uniform DIF. One framework handles both patterns, and the matching variable can be continuous.

• Item response theory approaches: The item's parameters are estimated separately in each group and compared. Raju (1988) derived the exact area between two item characteristic curves as a measure of how far apart they lie, with signed and unsigned versions that distinguish uniform from non-uniform DIF. Related procedures test the equality of item parameters directly, or compare the fit of models that do and do not let parameters differ by group.

• Nonparametric and bundle approaches: SIBTEST (Shealy & Stout, 1993) uses a latent-ability regression to separate DIF from genuine group differences in ability, and it can be applied to a bundle of related items rather than one at a time, which helps when the suspected cause is a shared feature such as passage topic or response format.


What test publishers do with a flag

The statistic starts a process rather than ending one. ETS test assemblers prefer category A items, may use category B when the blueprint requires it, and may use a category C item only where it is essential to the test specifications, with documented justification and review by people who did not work on the test, including external reviewers (Zieky, 2003, pp. 4-5).

Cognitive and achievement publishers report the same kind of screening. The Woodcock-Johnson IV technical manual documents DIF analyses across sex, race and ethnicity; most items showed no DIF, and most of those that did were kept out of the published forms (Canivez, 2017). Canivez notes a real limitation in that work: collapsing race into "White" and "non-White" can hide non-invariance that would show up in a specific group. NWEA reports that the vast majority of MAP Growth items show negligible DIF across gender and racial or ethnic groups (NWEA, 2026). Standard 4.10 requires the screening process itself to be documented (AERA, APA, & NCME, 2014, p. 88).


Why a DIF flag is not proof of bias

Zieky is blunt: "No statistic can determine whether or not a test question is biased" (Zieky, 2003, p. 3). The Standards agree: "The detection of DIF does not always indicate bias in an item; there needs to be a suitable, substantial explanation for the DIF to justify the conclusion that the item is biased" (AERA, APA, & NCME, 2014, p. 51). Whether an item is unfair depends on what the test is for. Zieky's example is a nursing licensure question about breast cancer that women find easier than matched men: fair on a test of what nurses must know, unfair on a test of general knowledge.

The statistics are also less decisive than their reputation suggests. Zwick (2012) found that the ETS category C rule "often displays low DIF detection rates even when samples are large," so an absence of flags is weak evidence of an absence of DIF. That is why the arguments in the articles on whether intelligence tests are biased against diverse populations and on whether IQ tests are racist rest on accumulated item-level and structural evidence rather than on any single coefficient. Anyone evaluating a test is entitled to ask which DIF method was used, against which matching variable, in which groups, at what sample size, and what became of the flagged items. The Reasoning and Intelligence Online Test publishes its methodology for that kind of reading, and can be taken as a professionally developed IQ test with documented item analysis.


Frequently asked questions

What is the difference between DIF and item bias?

DIF is a statistical finding: equally able members of two groups answer an item differently. Bias is a judgement that the difference comes from something irrelevant to what the item is meant to measure. A biased item should show DIF, while an item showing DIF may still be perfectly fair.

How large does DIF have to be before anyone acts?

There is no universal threshold. The best known convention is the ETS classification, which places an item in category C when its MH D-DIF statistic is significantly greater than 1.0 and at least 1.5 in absolute value on the delta scale of item difficulty.

Can DIF be tested on an individually administered IQ test?

Yes, where the standardization sample is large enough. Maller's analysis of the Wechsler Intelligence Scale for Children, Third Edition used its 2,200-case national standardization sample, and the Woodcock-Johnson IV technical manual reports DIF screening across sex, race and ethnicity.

Does a clean DIF analysis prove a test is fair?

No. Fairness in the Standards covers far more than item-level statistics, including access, administration conditions and the validity of score interpretations for each group. A clean DIF analysis is one piece of evidence, and conservative flagging rules mean some real DIF goes undetected.

Which DIF method is best?

They answer slightly different questions. Mantel-Haenszel is cheap and needs no fitted model, logistic regression handles non-uniform DIF, item response theory methods describe the whole difference between two item characteristic curves, and SIBTEST can test a bundle of items at once.


References

1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA. testingstandards.net

2. Zieky, M. (2003). A DIF primer. Educational Testing Service. praxis.ets.org

3. Zwick, R. (2012). A review of ETS differential item functioning assessment procedures: Flagging rules, minimum sample size requirements, and criterion refinement (Research Report RR-12-08). Educational Testing Service. files.eric.ed.gov

4. Holland, P. W., & Thayer, D. T. (1986). Differential item functioning and the Mantel-Haenszel procedure. ETS Research Report Series, 1986(2). doi.org

5. Swaminathan, H., & Rogers, H. J. (1990). Detecting differential item functioning using logistic regression procedures. Journal of Educational Measurement, 27, 361-370. doi.org

6. Raju, N. S. (1988). The area between two item characteristic curves. Psychometrika, 53, 495-502. doi.org

7. Shealy, R., & Stout, W. (1993). A model-based standardization approach that separates true bias/DIF from group ability differences and detects test bias/DTF as well as item bias/DIF. Psychometrika, 58, 159-194. doi.org

8. Maller, S. J. (2001). Differential item functioning in the WISC-III: Item parameters for boys and girls in the national standardization sample. Educational and Psychological Measurement, 61, 793-817. doi.org

9. Canivez, G. L. (2017). Test review of the Woodcock-Johnson IV. In Mental Measurements Yearbook. Buros Center for Testing. ux1.eiu.edu

10. NWEA. (2026). MAP Growth technical report for 2024-2025. NWEA. nwea.org

Hero image: wheelchair access ramp over steps at the Hotel Montescot, Chartres, by Coyau, licensed CC BY-SA 3.0 (creativecommons.org/licenses/by-sa/3.0). Via Wikimedia Commons.

Take our professional IQ test

Want to know your IQ? Try the first ever professional online IQ test.

Try our IQ test
Author
Dr. Russell T. WarneChief Scientist

Contact

Table of Contents

  • DIF is not the same thing as impact
  • Matching on ability is the whole trick, and it is imperfect
  • Uniform and non-uniform DIF
  • The four families of detection method
  • What test publishers do with a flag
  • Why a DIF flag is not proof of bias
  • Frequently asked questions
  • What is the difference between DIF and item bias?
  • How large does DIF have to be before anyone acts?
  • Can DIF be tested on an individually administered IQ test?
  • Does a clean DIF analysis prove a test is fair?
  • Which DIF method is best?
  • References
Article Categories
All ArticlesUnderstanding IQ ScoresTaking an IQ TestRIOT-Specific InformationGeneral IQ & IntelligenceAdvanced Topics & ResearchIQ Scores & InterpretationMensa & High-IQ SocietiesOnline IQ Tests IQ Test Basics & FundamentalsAverage IQ & DemographicsFamous People & IQHistory & Origins Of IQ TestingAccuracy, Reliability & CriticismSpecial Population & Related ConditionsImproving IQ / PreparationSpecific IQ Tests & FormatsIQ Testing for HR & RecruitmentSkills Assessment
Related Articles
Differential item functioning: how item bias is actually detectedStandard error of measurement: what it is and how it is calculatedWhat is internal consistency? Reliability from a single test sittingWhat is test-retest reliability? Score stability across two testingsWhat is face validity? Why a test that looks right can still be worthlessWhat is ecological validity? The two meanings, and what they mean for IQ testsWhat is criterion validity? Concurrent and predictive evidence explainedWhat is content validity? Sampling the domain a test claims to coverWhat is inter-rater reliability? Agreement between two scorersWhat is construct validity? How we know an IQ test measures intelligenceThe Mozart Effect: Does Listening to Music Raise IQ?How to Spot a Fake Online IQ TestWhat Is the Average IQ in the UK?Is Gen Z IQ Dropping?How to Tell If an Online IQ Test Is LegitimateWhy a Norm Sample Matters for IQ Test AccuracyCan Amateur IQ Tests Give Accurate Scores?How Accurate Are IQ Tests?What Is an IQ Confidence Interval? Why Scores Are Ranges7 Common Myths About IQ Tests DebunkedWhat Makes an IQ Test Scientifically Valid?Are IQ Tests Racist?Are IQ Tests Biased?How Reliable are IQ Tests?The IQ of Artificial IntelligenceChatGPT’s IQWhy Are IQ Tests Flawed?Are Online IQ Tests Legit?Is There an Official IQ Test?What is the Most Accurate IQ Test?Are IQ Tests Valid?Are IQ Tests Reliable?Are IQ Tests Good Measures of Intelligence?Are IQ Tests Accurate?
Take our IQ tests

Basic IQ Test

5 subtests + 5 cognitive abilities

Take the IQ test

Features

  • ~13 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±5.6 IQ margin of error

5/15 Subtests

Learn more
Vocabulary
Matrix Reasoning
SToVeS
Visual Reversal
Symbol Search
Most comprehensive

Full IQ Test

15 subtests + all cognitive abilities

Take the IQ test

Features

  • ~52 Minutes
  • IQ score
  • Cognitive abilities breakdown
  • ±3.7 IQ margin of error

15/15 Subtests

Learn more
Vocabulary, Information, Analogies
Matrix Reasoning, Visual Puzzles, Figure Weights
Object Rotation, SToVeS, Spatial Orientation
Computation Span, Exposure Memory, Visual Reversal
Symbol Search, Abstract Matching
Simple Reaction Time, Choice Reaction Time
Compare all tests