A situational judgement test presents realistic work scenarios and asks which responses are most effective. Learn how SJTs are scored and what they measure.
Dr. Russell T. WarneChief Scientist
Share
A situational judgement test (SJT) presents realistic work or study scenarios and asks you to judge which of several possible responses are most (and sometimes least) effective. Employers, medical schools, and government recruiters use SJTs as early screening tools because the scenarios can be tailored to a job and delivered cheaply online. What surprises many candidates is how much these tests overlap with cognitive ability. This article explains how SJTs work, what they actually measure, how they are scored, and how they differ from an IQ test.
How a situational judgement test works
Each SJT item presents a scenario and a set of potential responses or actions. Scenarios can be delivered in written, verbal, or video form, and the response format varies: pick the best option, pick the best and worst, rate every option's effectiveness, or rank the options in order.
A subtle design choice changes what the test measures. Researchers distinguish two instruction types. "Knowledge instructions" ask which option is most effective, a judgment about what one should do. "Behavioral tendency instructions" ask which option you would most likely perform. Meta-analytic work by McDaniel and colleagues shows that should-do versions correlate more strongly with cognitive ability, while would-do versions correlate more with personality traits. Two SJTs can look identical on screen and still measure different things.
What an SJT actually measures
An SJT is best understood as a measurement method rather than a single construct. Most SJTs target leadership and interpersonal skill, with others aimed at teamwork, personality, or job knowledge. Under the hood, scores blend general cognitive ability, personality, and job knowledge in proportions that depend on the content and the instructions.
The cognitive share is substantial. Across 79 correlations involving 16,984 people, SJT scores correlated with cognitive ability at a meta-analytic rho of .46. Correlations with personality dimensions such as agreeableness, conscientiousness, and emotional stability sit roughly in the .25 to .31 range. So while an SJT has little in common visually with a numerical reasoning test or an abstract reasoning test, a meaningful part of what it captures is the same general reasoning ability those tests measure directly.
SJTs also predict job performance at useful levels. The classic meta-analysis reported a validity of .34 across 102 coefficients and 10,640 people; the more conservative Sackett re-analysis puts SJTs at about .26, below structured interviews (.42) and close to the revised .31 estimate for general cognitive ability tests. One reason organizations like them is that SJTs tend to show smaller subgroup differences than pure cognitive tests, although research finds those differences grow as an SJT's correlation with cognitive ability grows.
Where you will meet one, and how it is scored
• University admissions: The UCAT, used for medical and dental admissions in the UK, Australia, and New Zealand, includes a Situational Judgement subtest of 69 questions in 26 minutes. It is scored in Bands 1 to 4, with Band 1 meaning your judgements matched a panel of experts in most cases, separate from the 300 to 900 scale used for the cognitive subtests.
• Government hiring: The UK Civil Service Judgement Test is an untimed online sift; official guidance says most people take two to four minutes per scenario. Candidates rate responses from counterproductive to effective, some scenarios play as short videos, a self-assessment section contributes 15% of the score, and results are reported as a percentile against other applicants.
• Medical careers: SJTs are used in postgraduate selection for doctors in the UK, Australia, and Canada. In UK general-practice selection, an SJT was found to be the best single predictor of supervisor ratings of junior doctors' job performance in their first year.
• Corporate hiring: Graduate schemes and volume hiring use written and video SJTs as unsupervised online screens, often blended with ability tests such as the SHL test battery. They are typically used early in the funnel to screen out low-scoring applicants.
Scoring keys are usually built from expert consensus or empirical keying, and results are reported normatively as bands or percentiles. Because different SJTs measure different blends of constructs, a score on one does not translate to another. Unlike IQ scores, there is no common SJT scale.
Can you prepare for an SJT?
The evidence is genuinely mixed. One study of Belgian medical-school admissions found commercial coaching was associated with a substantial score gain. Other research concluded that coaching has little effect on SJT scores, because the strategies needed to identify correct responses are more complex than on other tests. On faking, should-do SJTs are hard to distort, since applicants tend to respond to the best of their ability regardless of pressure, and even would-do versions appear less fakeable than personality questionnaires.
Familiarization clearly helps at the margin, and official bodies encourage it: the Civil Service directs candidates to a practice test, and the UCAT publishes full practice materials. For the cognitively loaded portion, retest research on cognitive tests shows average gains of about a quarter of a standard deviation, so treat practice as insurance against unfamiliarity rather than a way to transform your standing.
How an SJT differs from an IQ test
An IQ test measures reasoning abilities directly, on a common scale, with a known margin of error. An SJT embeds some of that reasoning inside job-specific scenarios and reports how your judgements compare with an expert key or an applicant pool. That makes SJT scores useful for one selection decision and hard to interpret anywhere else, which is also true of critical-thinking screens such as the Watson Glaser test.
If the reasoning component itself is what you want measured, use an instrument built for it. The RIOT IQ test (the Reasoning and Intelligence Online Test) reports IQ from six ability areas with a stated margin of error at every length: the free version takes about 8 minutes at ±14.9 IQ points, the Basic version takes about 13 minutes at ±5.6, and the Full version runs 15 subtests over about 52 minutes at ±3.7. A good SJT tells an employer how you weigh workplace trade-offs. A good IQ test tells you, with quantified precision, how strong the reasoning behind those judgements is.
Frequently asked questions
What is a situational judgement test?
A test that presents realistic work or study scenarios with possible actions and asks you to judge which responses are most, and sometimes least, effective.
Do SJTs correlate with IQ?
Yes. Meta-analytically, SJT scores correlate with general cognitive ability at about .46, and the correlation is higher when the test asks what you should do rather than what you would do.
Are there right answers on an SJT?
Usually, yes. Most high-stakes SJTs score you against a key built from expert consensus. The idea that there are no right answers applies to personality questionnaires, not SJTs.
Is the Civil Service Judgement Test timed?
No. The official guidance describes it as untimed, with most people spending two to four minutes per scenario, and it includes a self-assessment worth 15% of the score.
Are SJTs good predictors of job performance?
Useful but mid-tier. Classic estimates put validity at about .34, revised to roughly .26 in newer analyses, below structured interviews and comparable to many standard selection tools.
References
1. Christian, M. S., Edwards, B. D., & Bradley, J. C. (2010). Situational judgment tests: Constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology, 63(1), 83-117. unc.edu
2. Hausknecht, J. P., Halpert, J. A., Di Paolo, N. T., & Moriarty Gerrard, M. O. (2007). Retesting in selection: A meta-analysis of coaching and practice effects for tests of cognitive ability. Journal of Applied Psychology, 92(2), 373-385. doi.org/10.1037/0021-9010.92.2.373
3. McDaniel, M. A., Hartman, N. S., Whetzel, D. L., & Grubb, W. L., III. (2007). Situational judgment tests, response instructions, and validity: A meta-analysis. Personnel Psychology, 60(1), 63-91. doi.org/10.1111/j.1744-6570.2007.00065.x
4. McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology, 86(4), 730-740. doi.org/10.1037/0021-9010.86.4.730
5. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. doi.org/10.1037/apl0000994
6. UCAT Consortium. (n.d.). Test format. ucat.ac.uk
7. GOV.UK. (n.d.). Preparing for the Civil Service Judgement Test. gov.uk