In 1950, two University of Chicago researchers set out to help college students who kept doing poorly on multiple-choice exams (and one researcher you probably know well). First, they collected think-aloud transcripts from students who scored well. Then they had the struggling students think aloud through the same items, handed each one their own transcript beside a transcript from a successful student, and asked a single question: what was the difference? They offered no advice and no instruction. The students found their own gaps. Some noticed that when they hit an item they did not know, they gave up, while the successful student shifted into problem solving and started eliminating options. Others saw that the strong performers kept returning to the problem statement to correct their own misreadings. Performance improved significantly (Bloom & Broder, 1950, as described in Klein, Hintze & Saab, 2013).
The lead researcher was Benjamin S. Bloom. The same Bloom whose taxonomy sits in every instructional designer's toolkit. He ran this experiment six years before the taxonomy was published, and few in the field have ever heard of it. Notice what the design contains: no lecture, no feedback in the usual sense, just a learner's own thinking laid next to a stronger performer's, with the comparison doing all the work.
The idea did not stay in 1950. Decades later, a research field with no connection to instructional design rediscovered it, gave it a method and a name, and put firefighter data behind it.
What is naturalistic decision-making?
That field is naturalistic decision-making, and its founding observations came from firegrounds. Gary Klein and his colleagues studied how commanders decide under time pressure and found that they do not weigh options the way decision textbooks prescribe. They size up the situation, recognize a pattern from experience, and commit to a course of action, evaluating it on its own terms rather than lining it up against alternatives. Klein's Sources of Power traces this recognition-driven judgment through the expert domains the field studied, chess masters among them (Klein, 1998).
The finding creates a training problem. If expert decision quality lives in recognition built from years of situations, how does a newcomer acquire it without the years? One answer: let trainees see situations through the eyes of experts.
What is the ShadowBox method?
Neil Hintze, a New York City fire battalion chief, built that answer for his 2008 master's thesis. Newly promoted fire officers face situations they may never have encountered during routine operations, so Hintze developed four challenging scenarios and gave them to fourteen experienced New York City Fire Department officers, who described how they would handle each decision and the rationale behind it. Those responses were synthesized into materials that put the expert mindset on paper (Klein, Hintze & Saab, 2013).
The method that grew out of the thesis works like this. The trainee receives a scenario booklet. At each predetermined decision point, a prompt asks a question, and the trainee writes an answer and rationale into a small printed box, usually one inch square, with about two and a half minutes to do it. Once finished, the trainee can never turn back; all they have to go on is what they wrote down. On the next page, they see what the expert panel agreed should go in that box, along with the panel's rationale, and they describe the differences between the two. The boxes take different forms: an Attention box for information worth retaining, an Action Priority box for ranking a short list of possible actions and recording the top three, an Information box for the one question the trainee would most want answered at that moment. Because the expert responses are captured ahead of time, no facilitator needs to be present.
The name is literal: trainees shadow the thinking of experts who are not in the room, writing their decisions in the one-inch-square boxes on the scenario materials, which is why Klein and his coauthors titled the paper "Thinking Inside the Box."
One detail of the expert panel deserves attention. Hintze found strong consensus for many of the boxes but never one hundred percent convergence, and where the experts split, he told trainees about any strong minority position and made clear that no answer carried ground truth.
The validation numbers are the reason the method travels. An experimental group of fourteen recently promoted New York State fire officers worked through the materials; a matched control group of fifteen officers did not. On a scoring key with a maximum of 100 points, the trained group averaged 86.9 against 73.6 for the controls, an 18% improvement in decision-making performance (p < .001).
What should scenario designers steal?
Three techniques transfer directly from ShadowBox to anyone building scenario-based learning, whatever the tool.
Commit before you compare. The learner writes their decision down before seeing any expert view, so their real reasoning is on the table. Bloom's students produced transcripts before seeing a stronger one; Hintze's trainees filled in the box before turning the page, with no way back. Without that commitment, comparison collapses into hindsight, and every learner quietly concludes they would have chosen well.
Use an expert panel, not an answer key. Real experts disagree, and the method treats that as information rather than noise. Several defensible choices are shown as the panel actually held them, including the strongest minority opinion. A learning designer who smooths the disagreement into a single sanctioned answer is teaching a picture of the work that experienced audiences will recognize as false. It's a multiple-choice mindset. In real situations, more than one good choice is common.
Compare reasoning against expert reasoning. The trainee sees why the experts chose, not just what they chose, and measures their own thinking against it. The rationale is where the expert mindset actually lives; an answer without it is a verdict the learner can memorize but not learn from.
This article deliberately stops short of procedure. How to find decision points, build plausible choices, and write the expert rationale that accompanies them is craft covered in two companion guides: designing decision points and decision-point mentoring.
Is decision-making skill really trainable?
A fair question sits underneath all of this: is decision-making a skill you can train, or a trait you hire for? Ralph Keeney, writing from the decision analysis tradition, offers a wide-angle answer (Keeney, 2004). Sketching how a representative 10,000 decisions get made, he estimates that about 7,000 carry consequences too small to merit thought and another 2,000 are no-brainers, but roughly 1,000 are worth thinking about, and of those only about 40 receive systematic thought. Almost nobody has ever been trained in decision-making. We learn it by doing, starting around age two, and pick up bad habits along the way.
Keeney's argument is that this is fixable, because decision-making decomposes the way tennis does. A tennis game breaks into serve, forehand, backhand; a serve breaks into positioning, toss, and contact; each element can be learned and improved on its own, then integrated. Decisions decompose the same way, into elements such as defining the problem, specifying objectives, and creating alternatives, and he argues decision-making should be treated as a primary skill. For designers of workplace learning, the implication is direct: decision quality is a design target, not a fixed property of the learner.
Why does the convergence matter?
Learning science
Bloom & Broder · Ericsson
Compare your thinking against successful thinking, then refine
Cognitive apprenticeship
Collins, Brown & Holum
Make expert reasoning visible at the moment of work
Naturalistic decision-making
Klein · ShadowBox
Commit to a decision, then compare against an expert panel
A realistic situation · a committed choice · expert reasoning at the decision
Step back and count the traditions in this story. Learning science, running from Bloom and Broder's comparison study through the research summarized in our guide to deliberate practice, keeps finding that expertise grows from focused attempts with expert models and feedback. Cognitive apprenticeship built its methods around making expert thinking visible and coaching learners on the specific events that arise as they work. And naturalistic decision-making, starting from firegrounds rather than classrooms, produced ShadowBox: realistic scenarios, committed decisions, expert reasoning revealed at each decision point.
Three fields, different founding questions, different methods, different journals, and the same architecture at the end: a realistic situation, a committed choice, and expert reasoning revealed at the decision. When fields that do not read each other's literature arrive independently at one design, that convergence is the strongest form of evidence applied science produces. The architecture belongs to no single method and no single maker; it is settled ground that any scenario designer can build on with confidence.
AliveSim's Guided Scenarios are one embodiment of that shared architecture, with a deliberate refinement: the expert reasoning moves into the moment of choice. Where ShadowBox reveals the panel's thinking after the learner turns the page, a Guided Scenario delivers it as mentoring the instant the learner commits, responding to the specific choice they made while their own reasoning is still fresh, and then lets them choose again. The architecture came first; the platform is one way of building it.
The three steals need no platform at all. Commit before comparing, expert panels over answer keys, reasoning against expert reasoning: Bloom proved the idea with transcripts and a question, and Hintze proved it with a booklet and a one-inch box.
References
- Bloom, B. S., & Broder, L. J. (1950). Problem-solving processes of college students: An exploratory investigation. University of Chicago Press. (As described in Klein, Hintze & Saab, 2013.)
- Keeney, R. L. (2004). Making better decision makers. Decision Analysis, 1(4), 193–204.
- Klein, G. (1998). Sources of Power: How People Make Decisions. MIT Press.
- Klein, G., Hintze, N., & Saab, D. (2013). Thinking inside the box: The ShadowBox method for cognitive skill development. Proceedings of the 11th International Conference on Naturalistic Decision Making. Marseille, France.
Related questions
What is the ShadowBox method?
ShadowBox is a scenario-based method for developing cognitive skill, described by Klein, Hintze, and Saab (2013). Trainees work through a scenario booklet and, at predetermined decision points, write their answers and rationale into small printed boxes, usually one inch square, before turning the page to see what a panel of experts wrote in the same boxes and why. The name is literal: trainees shadow the thinking of experts who are not in the room. Because the expert responses are captured in advance, the method needs no facilitator. Its validation study with newly promoted fire officers showed an 18% improvement in decision-making performance over a matched control group (p < .001).
What is naturalistic decision-making?
Naturalistic decision-making (NDM) is the research field that studies how people actually decide in real settings, under time pressure, uncertainty, and shifting conditions, rather than in laboratory puzzles. Its signature finding came from Gary Klein's studies of fireground commanders: experienced decision makers rarely line up options and weigh them the way decision textbooks prescribe. They size up the situation, recognize a pattern from experience, and commit to a course of action (Klein, 1998). That finding reframes decision training as a problem of building recognition, which is the problem the ShadowBox method was designed to solve.
Why show learners minority expert opinions?
Because hiding them misrepresents the work. When Hintze built the ShadowBox expert panel from interviews with fourteen experienced fire officers, he found strong consensus on many decision points but never total convergence, so he told trainees about any strong minority position and made clear that no answer carried ground truth. A single sanctioned answer teaches learners that real situations have one, which experienced audiences know is false. Showing the minority view alongside the consensus gives learners the true shape of the option space and lets them weigh competing expert rationales, which is much closer to the judgment the job actually demands.
Did the Bloom of Bloom's Taxonomy really study decision training?
Yes. Benjamin S. Bloom, with Lois Broder at the University of Chicago, published the study in 1950, six years before the taxonomy that made his name. Bloom and Broder collected think-aloud transcripts from students who performed well on multiple-choice exams, had under-performing students think aloud through the same items, then showed each struggling student their own transcript beside a successful student's and asked what the difference was. They gave no advice. The students identified their own gaps, and performance improved significantly. Klein, Hintze, and Saab (2013) cite the study as a direct ancestor of the ShadowBox method.
Published July 17, 2026 · 8 min read