Every learning program already runs a pre-test and a post-test. It is the most universal measurement ritual in education, and almost every instance of it measures the same thing: knowledge, what learners can recall before the activity and after it. The interesting question is whether the same familiar instrument can measure something far more valuable, whether people changed how they would perform in the situations the education targets. It can. The design that makes it true has three parts, and most programs are missing two of them.
Part one: the tests are vignettes
A vignette is a mini-situation, not a quiz item. It has context: a presentation with enough specifics that the decision is real rather than abstract. It has a decision: something the learner must choose to do, not something to define. And it has plausible options, which is where the craft lives. Good options are drawn from what people in this audience actually do, which means the tempting suboptimal choice is on the list, worded with the same care as the optimal one. An option nobody would pick is not a distractor; it is filler.
This changes the measurement target. A recall item asks whether the learner can retrieve content. A vignette asks what the learner would do, in context, with the real alternatives in view. That is a measure of judgment in the situation, and judgment is what the education was built to change.
The principle has strong validation behind it. When Peabody and colleagues compared measurement methods against unannounced standardized patients, the gold standard in that literature, clinical vignettes, case-based decision measures with controlled context, consistently scored closer to the standard than chart abstraction did, in the original 2000 JAMA study and again in the 2004 multisite replication. The 2004 study concluded that vignettes are a valid tool for measuring the quality of clinical practice and may be useful for longitudinal evaluations of interventions intended to change it. Two boundaries travel with that citation. The validation holds at the group level, for cohort claims, not for individual scorecards. And it attaches to the principle, case-based measurement with controlled context, not to any specific response format.
Part two: the activity contains similar situations, never identical ones
This is the heart of the design, and the part almost everyone misses.
Between the two tests sits the learning activity, and for the post-test to mean anything about performance, the activity must contain situations built on the same concept the vignettes measure, but never the vignettes themselves. The learner should meet the concept in a messier form than the test presents it: more ambiguity, less structure, more going on. Where the vignette isolates the decision, the scenario embeds it, in an unfolding case with incomplete information, competing considerations, and expert mentoring at each decision point. The learner experiences applying the concept there, choice by choice.
The reason for the never-identical rule is what each design would train. If the activity contains the exact situations the test presents, the learner meets the vignette twice, and the second meeting is recognition. The post-test gain would measure rote memory of a specific item, which is a knowledge result mislabeled as a performance result. Similar-but-different situations force something else: the learner has to work with the concept, because the surface details do not match anything they can simply recall. When the post-test then presents its vignette, the learner faces a situation they were never shown, and a changed response means the concept transferred.
Concretely, "similar but never identical" means the same decision logic under different particulars: a different presentation, a different mix of complicating factors, a different point in the case where the decision arises. The alignment runs concept to concept, not surface to surface. Getting that alignment explicit, this vignette measures this judgment, these scenario situations exercise it, is the design work of the whole instrument.
1 · Pre-test
A vignette: a mini-situation
Context, a decision, plausible options. Measures judgment, not recall, before the program.
2 · The guided scenario
The same concept, messier
More ambiguity, less structure, mentoring at each decision. The learner experiences applying the concept the vignette measures.
Similar, never identical
3 · Post-test
The same instrument, a changed response
The learner now handles the vignette differently: measured change in how they would perform.
Part three: a changed response is a changed performance
Now the post-test carries a different claim. When a learner who chose the tempting suboptimal option on the pretest chooses the appropriate response on the post-test, the program has observed a change in how that learner handles the situation: same judgment required, different particulars, better decision. Aggregated across a cohort, that is a measured change in performance capability, captured on an instrument every stakeholder already accepts and already budgets for. No new data source, no follow-up survey, no chart pull. The claim's strength as evidence of practice change, and how to report it, is the subject of our companion guide on Moore's Level 5 surrogates, which explains why the claim stays "surrogate" rather than proof of practice change. The point here is what the instrument itself can support when all three parts are in place.
Two lenses on the same answer
Syandus has measured the educational impact of this design directly, through two independent instruments.
The first lens is the outside instrument, pre/post vignette performance. Across 15,197 posttest learner responses spanning 14 therapeutic areas, with all data in the analysis period included (no selection), clinicians performed 2.58 times better on clinical vignette questions after learning through AliveSim's Guided Scenarios. The effect size is very large by any conventional standard (Cohen's d = 2.71, where 0.8 is conventionally considered large; p < 0.001). One methods note belongs alongside a figure that size: each clinician is compared against their own pretest, on the same test items before and after. The test items are the same at both ends; it is the scenario situations in between that are similar but never identical, which is why the gain measures applied judgment rather than rote memory.
The second lens is the inside instrument, and it replicates the finding within the simulation itself. In a matched decision-point analysis of 31,673 decisions across seven programs, clinicians showed a 2.85x improvement in decision-making when they encountered similar situations later in the simulation (d = 3.95, p < 0.001): evidence that expert mentoring was applied, not merely received.
Notice that the inside lens runs on the same similar-but-never-identical logic as the whole design. The improvement is measured across similar situations inside the activity, so it is a transfer measure too, not a repetition measure. Two different instruments, measuring at different points, one outside the activity and one inside it, converging on the same answer. That convergence is what "it works" looks like in data.
The design diagnoses itself
The same data that evaluates the program also grades the instruments, and this may be the most useful property of the whole design. There are three diagnostic reads, and each one is a quality measure for the vignettes and the scenario themselves.
- Pretest scores come back too high. The vignette question is too easy. It is not reaching the judgment where the gap lives, so it cannot register the change the program produces. The item needs a harder situation or a more tempting suboptimal option.
- Most learners' first in-scenario choices are already optimal. The scenario is not addressing a significant gap for this audience. The struggle it was designed around is not there, and education aimed where the audience already performs well is education misdirected.
- The pre/post gain is not substantial. One of two things is true: the pretest is not aligned with the concept the scenario develops, so the instrument is measuring beside the intervention, or the pretest is aligned and the scenario was not effective at closing the gap, in which case the scenario needs revision.
Read this way, every program becomes an instrument-calibration exercise as well as an educational one. The data shows you which vignettes reached the real gap and which did not, which scenario situations produced the struggle you designed for and which the audience walked through untouched. Each program's results make the next program's pretest questions sharper and its scenarios better aimed. The measurement system improves the design system, which is exactly what a measurement system is for.
The design is demanding: vignettes built from real decision behavior, scenario situations aligned to the same concepts but never duplicating them, and item writing disciplined enough that gains mean what they claim. But every piece of it is craft, not infrastructure. The pre/post test is already in every program. Built this way, it measures what the program was for. (For what the in-activity data stream itself should capture, see what learning analytics should measure.)
References
- Peabody, J. W., Luck, J., Glassman, P., Dresselhaus, T. R., & Lee, M. (2000). Comparison of vignettes, standardized patients, and chart abstraction: A prospective validation study of 3 methods for measuring quality. JAMA, 283(13), 1715-1722.
- Peabody, J. W., et al. (2004). Measuring the quality of physician practice by using clinical vignettes: A prospective validation study. Annals of Internal Medicine, 141(10), 771-780.
Related questions
Why do most pre/post tests only measure knowledge?
Because both halves of the design default to knowledge. The test items are usually recall questions, which measure whether the learner can retrieve content, and the activity between the tests usually delivers content rather than situations, so retrieval is all there is to improve. A gain on that design is real, but it is a knowledge gain. Measuring changed performance requires changing both halves: the tests become vignettes, mini-situations that measure judgment, and the activity becomes a set of situations in which the learner applies the same concept the vignettes measure, in messier form, with expert mentoring at the decisions.
What is a vignette in outcomes measurement?
A vignette is a mini-situation used as a test item: enough context to make the decision real, a decision point, and plausible response options drawn from what people in the audience actually do, including the tempting suboptimal choice. What it measures is the learner's judgment in the situation rather than recall of content. The principle has strong validation behind it: Peabody and colleagues showed that case-based measures with controlled context track the quality of actual clinical practice more closely than chart abstraction does, in a 2000 JAMA study and a 2004 multisite replication. That validation holds at the group level, for cohort claims, and it attaches to the principle of case-based, context-controlled measurement rather than to any specific response format.
Why must the scenario situations differ from the test vignettes?
Because identical situations train recognition. If the learner meets the exact vignette inside the activity, the post-test measures whether they remember the answer to that item, and the gain is rote memory. When the activity instead presents the same concept in situations that are similar but never identical, messier, more ambiguous, with more going on, the learner has to work with the concept itself rather than a memorized response. A changed answer on the post-test vignette then means the concept transferred to a situation the learner was never shown, which is the difference between measuring memory and measuring the ability to apply.
How do you know if your pretest questions are good?
The same data that evaluates the program evaluates the instruments, and there are three diagnostic reads. If pretest scores come back too high, the vignette is too easy: it is not reaching the judgment where the gap lives. If most learners' first in-scenario choices are already optimal, the scenario is not addressing a significant gap for this audience, because the struggle it was designed around is not there. If the pre/post gain is not substantial, either the pretest is not aligned with the concept the scenario develops, or the scenario was not effective at closing the gap and needs revision. Read this way, every program's results show you how to craft better pretest questions and better scenarios for the next one.
Published July 17, 2026 · 7 min read