You are leading a project. A teammate has just missed a second deadline, and a client presentation is now at risk. Below the situation sit four responses: talk to the teammate privately about what is getting in the way, escalate the problem to their manager, quietly redistribute the remaining work across the team, or raise the missed deadlines at the next team meeting. Pick the best response. Now pick the worst.
That is a situational judgment test item, and a full test is a dozen or more of them in a row. Situational judgment tests, SJTs from here on, are assessments used for hiring and promotion. Each item presents a small job situation and a handful of plausible actions, and the candidate judges among them, most often by choosing the best and worst options. Answers are scored against the judgment of experienced performers, and employers have used the format in selection since the 1940s. Delivery is usually text, sometimes video. That is the whole machine.
Why does this format look familiar?
Because you build it for a living. Strip the labels off an SJT item and what remains is the skeleton of scenario-based learning: a situation flows into a decision point, which has options. A project. A teammate in trouble. A moment that demands judgment rather than recall. Four responses a real person might actually choose. Any instructional designer who has written a scenario decision would recognize the anatomy at a glance.
Now notice what is missing. The candidate never learns which response the experts favored. The story never continues past the choice; no client reacts, no teammate responds. Nobody explains why the private conversation beats the public callout. There is no retry, no mentor, no narrative. An SJT is a scenario decision point with everything but the judgment stripped away. It measures and stops, on purpose: an assessment that coached its candidates would stop being an assessment.
Here is why that should interest anyone who designs scenarios. SJTs are assessments with no feedback, no consequences, and no story. Scenarios add the learning loop. Which means the two formats share their load-bearing structure, and HR has spent decades doing something scenario designers rarely get funding for: rigorously measuring whether that bare structure, a situation plus a judgment among plausible options, tells you anything real about how a person performs. It does. When you build a decision point, you are building on validated assessment bones, and everything you add on top is the learning.
What has the hiring research established?
Four findings matter most, and each one carries a design lesson. A quick translation first, because the research literature leans on one word constantly: when hiring researchers call a test "valid," they mean its scores line up with how people later perform on the job. The strength of that lineup is reported as a correlation, where zero means the test tells you nothing and 1.0 would mean it predicts perfectly.
Judging small situations predicts real performance. The classic meta-analysis, which is a study that pools the results of many prior studies, comes from McDaniel and colleagues. Across 102 separate correlations covering 10,640 people, SJT scores correlated about .34 with job performance (McDaniel et al., 2001). In everyday terms: people who judge these stripped-down situations well go on to do the job better, reliably enough that companies stake actual hiring and promotion decisions on the scores. For learning designers, this is the foundational reassurance. The decision point format is not an engagement gimmick. Situation plus plausible options is a measurement instrument with a hiring-grade track record.
Teamwork and leadership judgment predict best. Christian, Edwards, and Bradley (2010) sorted the SJT literature by what each test was built to measure and found that SJTs most often assess leadership (38%) and interpersonal skills (13%) among the tests that reported what they measured. More useful still, tests targeting teamwork skills and leadership showed relatively high correlations with overall job performance, .38 and .28 in their pooled estimates. Look at that territory: leading, collaborating, handling people. It is precisely the judgment-heavy ground most workplace scenarios are built on. The situations hardest to cover with a policy document are the ones where judging well is most tightly linked to performing well.
Matching the measure to the outcome sharpens it. The same meta-analysis found evidence, from an admittedly small set of studies, that prediction improved when the skill a test measured was matched to the facet of performance being predicted, such as a leadership-focused test predicting managerial performance specifically (Christian et al., 2010). The design translation: decide what each decision point is about, and align it with the outcome you actually care about. A decision point that is "about" everything measures nothing in particular. That discipline also determines what your decision data can later tell you, the subject of our guide to what learning analytics should measure.
Richer presentation tended to measure better. Christian and colleagues also compared delivery formats while holding constant what the tests measured, and video-based SJTs tended to show stronger correlations with performance than pencil-and-paper versions; for tests of interpersonal skills, the pooled estimates were .47 for video against .27 for paper, though the video figure rests on only a couple of studies (Christian et al., 2010). Treat it as early evidence rather than settled fact, but the direction is striking: even for pure assessment, where nothing needs to be engaging, presenting the situation more like reality made the measurement better. Fidelity is not decoration. Seeing and hearing a tense conversation engages the same social judgment the job will demand in a way a paragraph cannot.
One more detail rewards attention: the ask itself. SJT instructions come in two families. Knowledge instructions ask the candidate to evaluate the options, picking the best, or the best and worst. Behavioral-tendency instructions ask what the candidate would most likely do. McDaniel and colleagues' follow-up work found the two families measure different things, with knowledge instructions leaning toward general reasoning ability and tendency instructions leaning toward personality (McDaniel et al., 2007, via the same summary). So the phrasing of a decision prompt is a genuine design lever. And the field's habit of asking for best and worst treats the option set as a spectrum of quality rather than one right answer surrounded by filler, the same principle behind grading every option that our guide to designing decision points builds its craft on. That guide owns the how; this article's point is that the assessment world independently landed on the same shape.
What do scenarios add that SJTs never had?
An SJT item
A realistic situation
A decision point
Plausible options
→ a score. Full stop.
A scenario decision point
A realistic situation
A decision point
Plausible options
+ the learning loop
Consequences play out · coaching at the choice · try again · the story continues
Everything the SJT strips out on purpose, the scenario builds in on purpose. That is the cleanest way to see what scenario-based learning actually is: the same validated bones, plus the learning loop.
The loop has four parts. Feedback arrives at the moment of choice, while the learner's reasoning is still in working memory, rather than never arriving at all. Consequences play out: the teammate reacts to being called out in the meeting, the client hears about the slipped deadline, and the learner watches the situation move. Coaching explains the why, the expert reasoning that separates the private conversation from the public one, which is the craft covered in our guide to decision-point mentoring. And the story continues, carrying the learner into the next situation with the last decision's weight still on them, often with the chance to return and choose better.
There is a second benefit, and it deserves plain language. A scenario decision point has the same structure as an SJT item, so the choices learners make in a scenario are the same kind of evidence companies already trust when they use SJT scores to decide who to hire. When a learner judges a realistic situation inside a scenario, that choice is not a soft engagement metric. It is the same class of evidence an SJT collects, gathered in an environment that also develops the capability it measures.
For teams that want this without assembling it from parts, AliveSim's Guided Scenarios are built as exactly this combination: decision points with the anatomy the assessment research validated, situations flowing into judgments among plausible, graded options, wrapped in the learning loop the SJT never had. A mentor responds to the specific choice made, consequences play out safely, and the learner returns to the decision to find the stronger approaches, while every judgment is captured as decision data.
The next time a stakeholder asks whether choices in a scenario really measure anything, the answer comes from an unexpected direction. HR answered it decades ago, one stripped-down decision point at a time. The bones are proven. The loop is yours.
References
- Christian, M. S., Edwards, B. D., & Bradley, J. C. (2010). Situational judgment tests: Constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology, 63(1), 83–117.
- McDaniel, M. A., Hartman, N. S., Whetzel, D. L., & Grubb, W. L. (2007). Situational judgment tests, response instructions, and validity: A meta-analysis. Personnel Psychology, 60(1), 63–91.
- McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology, 86(4), 730–740.
- Paul, M. (2020). Umbrella summary: Situational judgment tests. Quality Improvement Center for Workforce Development. https://www.qic-wd.org/umbrella/situational-judgment-tests
Related questions
What is a situational judgment test?
A situational judgment test (SJT) is an assessment used in hiring and promotion, built from short descriptions of job situations. Each item presents a situation, such as a teammate missing a second deadline while a client presentation is at risk, followed by four or five plausible responses. The candidate judges the responses, typically picking the best and the worst. Scores are compared against the judgments of experienced performers, and employers use the results to inform selection decisions. SJTs are usually delivered in text or video form and have been used in employee selection since the 1940s.
Do situational judgment tests predict job performance?
Yes, and the evidence base is large. A meta-analysis by McDaniel and colleagues pooled 102 separate correlations covering 10,640 people and found that SJT scores correlate about .34 with job performance, a strong result by the standards of hiring research, where employers routinely base selection decisions on instruments at this level. A later meta-analysis by Christian, Edwards, and Bradley found the prediction is strongest when the test targets teamwork or leadership judgment, and that prediction improves further when the skill the test measures is matched to the kind of performance being predicted.
How are SJTs different from scenario-based learning?
They share the same skeleton and differ in everything wrapped around it. Both present a realistic situation that flows into a decision point with plausible options. An SJT stops there: the candidate judges, the test scores, and nothing else happens, because helping the candidate would defeat the purpose of an assessment. A learning scenario adds the learning loop, meaning feedback at the moment of choice, consequences that play out, coaching on the reasoning, a story that continues, and the chance to choose again. An SJT measures judgment; a scenario measures it and then develops it.
Why ask for the best and worst response instead of one right answer?
Because it treats the option set as a spectrum of quality, which is what real situations offer. Asking only for the single right answer implies every other option is equally wrong, which experienced people know is false. Asking for best and worst requires two separate acts of judgment, discriminating among the strong options and recognizing which plausible-looking move causes the most damage. Research on SJT response instructions shows the wording of the ask changes what an item measures, so the question format is a genuine design decision, not a cosmetic one.
Published July 17, 2026 · 8 min read