Every proposed learning program needs to answer the same question: how can you show it worked? That defines whether it is worth the investment. The question arrives in a budget review, an accreditation report, a board update, or a course evaluation, and it is the same question in each. A traditional e-learning module answers it with testing, and testing proves something real: scores went up, so learners understood the material. What it leaves unanswered is the part everyone actually cares about. When a real situation arrives, can the learner apply what they were trained on? That is the return on education, and this guide is about how to prove it.
Understanding is not application
A pre-test and a post-test are worth something. They show that learners took in the material and that their knowledge moved. But knowing the material and being able to act on it in a real situation are two different claims, and a test only supports the first one.
The major evaluation frameworks draw exactly this line. Kirkpatrick's framework, common in corporate learning, reserves a level of its own for whether people use the training on the job, the level it calls Transfer. Moore's framework, used in medical education, places performing in an authentic situation at its higher levels, Competence and Performance (Byerly). Different vocabularies, same point: recalling knowledge is one level, and doing something with it is a higher one. A post-test lives on the lower level. Whoever is asking you to show the program worked is asking about the higher one.
The proof is built into the design
A Guided Scenario answers the harder question by design, not by hope. The way one is built, every learner works through realistic situations and arrives at a successful solution to each. Along the way they commit to their own first approach, then are coached where that approach, or what they thought was appropriate, was not. They experience what optimal looks like, and they come to understand why it is best rather than being told. And every step of that path is recorded as it happens.
That structure is the proof, and it reads the same wherever the program runs: learners demonstrably applied the knowledge in the situations they were being prepared for, and the record of it can be put in front of whoever signs off.
Step 1
The learner's own first approach
Each learner commits to what they would do now. This is the baseline.
Step 2
Coached where the approach fell short
Where a current approach, or what seemed appropriate, was not, the learner is coached to adapt their own mental model.
Step 3
Arrives at the successful solution
Every learner reaches the successful solution, experiences what optimal looks like, and sees why it is best.
Recorded at every step
Business and operations
Evidence that people applied it in the real situations they face.
Education
Evidence that learners actually applied the learning.
What the numbers show
Syandus has measured this directly. In a matched decision-point analysis of 31,673 decisions across seven programs, learners showed a 2.85x improvement in decision-making when they met the same decision points in changed situations later in the experience. The strength of that effect: Cohen's d, the standard measure of effect size, came to 3.95, with p < 0.001. For scale, an effect size of 0.8 is conventionally considered large; 3.95 is several times beyond that threshold. "Matched decision points" means the same underlying gap, with the same set of options, met again in a different situation. Because the later situations differ from the earlier ones, the gain is not memory of a particular scenario; what improved is the handling of the gap itself. That is evidence that the coaching was applied, not merely received.
What kind of evidence is this? Observed evidence. Each decision was recorded as the learner made it, in a carefully controlled simulation of an authentic situation, and scored against that learner's own first choices. Currier drew the line that matters here: an after-the-fact survey asking whether practice changed is not a measurement, it is an informed guess, because it is not based on observing what the person actually did (Currier, 2007). Recorded decisions are the observation. And observed performance in authentic simulated situations is a potent surrogate for real-world performance: it demonstrates that learners applied the knowledge successfully in the situations they will face. Our guide on Moore's Level 5 and the surrogate argument makes that case in full, and it is the backbone of this one.
The retention dividend
The proof is the headline, and the return keeps compounding behind it. An organization that built a program has already paid for the knowledge it delivers, and knowledge that is delivered but never applied recedes, taking the investment with it. Adding the application step changes that arithmetic twice. It produces the proof described above, and the very act of applying converts the knowledge into the form that stays. That claim is supported by learning science, and the evidence is in our guide on why trained knowledge recedes and what holds it in place. For a return-on-education conversation, the point is direct: the same component that demonstrates the program worked also increases how much of the program survives in the learner, which raises the return on everything already spent.
Adoption is low-risk
The last reason this is an easy case to make is that it asks very little of what you already have. You are not rewriting your content. You are adding a component to it, one that demonstrates and improves the value of the program you already built. If you are building something new, you can go further and build the knowledge and the application together from the start, so the two arrive as one experience. Either way the move is additive, and the easiest program to get approved is the one that adds value without discarding prior investment. For the retrofit path in detail, see adding an application step to a program you already own.
Where the proof comes from
The strongest answer to "show it worked" is observation: put learners in realistic situations, record the decisions they make, coach them to the successful solution, and keep the whole path as evidence. That is the approach, and it is what AliveSim Guided Scenarios are built to do. Every learner arrives at the successful solution, having seen what optimal looks like and why, with every decision captured along the way. For how the measurement itself is designed and reported, see our guide to measuring behavior change.
References
- Byerly, B. Moore's Levels for CME: When Do They Work, and When Do They Fail? LinkedIn. (Framing: Transfer in the Kirkpatrick levels versus Competence and Performance in Moore's levels, and the parallel between corporate and medical-education evaluation.)
- Currier, R. L. (2007). Interactive simulations of clinical case studies: A cost-effective method for measuring changes in clinical practice. CE Measure, 1(2), 54–58.
Related questions
How do you prove learners can apply what they learned, not just that they understood it?
Understanding and application are different claims, and a pre-test and post-test only support the first one. To show application, you need learners making real decisions in realistic situations, with each decision recorded as it happens. A Guided Scenario is built so every learner commits to a first approach, is coached where that approach falls short, and arrives at the successful solution, with the whole path captured. That record is direct, observed evidence that the learner applied the knowledge in the situation, not a survey asking them to recall whether they think they did.
Is measured performance inside a simulation real proof?
It is the strongest proof most programs can obtain. The decisions are real, observed as they happen, in carefully controlled simulations of the authentic situations learners are being prepared for. That makes the record a surrogate for real-world performance: not a workplace study, but a demonstration that learners can apply the knowledge in the situations that matter. Our guide on Moore's Level 5 makes the full argument that observed simulated performance is the most potent surrogate available for real-world performance, precisely because the alternatives, recall tests and self-report surveys, observe nothing at all.
How is this different from a pre-test and post-test?
A pre/post test measures whether knowledge went up. It is answered by recall, and it tells you nothing about what a learner would actually do in a live situation. Currier makes the point plainly for the other common alternative: an after-the-fact survey asking whether practice changed is not a measurement, it is an informed guess, because it is not based on observing what the person actually did (Currier, 2007). Recorded in-scenario decisions are observation. That is what qualifies them to carry the application claim that testing and surveys cannot.
Is adding this to an existing program disruptive?
It does not require rewriting your content. You add a component to the program you already own, so the existing investment stays intact and gains both a way to demonstrate its value and a boost to how much of it learners retain. If you are building something new, you can go further and build the knowledge and the application together from the start. Either way the change is additive, which makes it one of the easier improvements to get approved.
Published July 20, 2026 · 5 min read