Every learning program eventually faces the same question from a sponsor, an executive, or an accreditor: did it change what people do? Completions, satisfaction scores, and quiz results are the answers most programs have on hand, and none of them answers the question. This guide covers what standard metrics actually measure, what the major outcomes frameworks define as real evidence, and how realistic scenarios turn behavior change from something inferred into something observed.
What do completions, satisfaction scores, and quizzes actually measure?
The three metrics most learning programs report are records of activity, opinion, and recall.
- Completions confirm that learners showed up and clicked through. They measure participation, nothing more.
- Satisfaction scores capture how learners felt about the experience. Reaction data is useful for improving a program's design and marketing, but it is a poor predictor of whether anyone behaves differently afterward.
- Quiz scores measure recall of declarative knowledge at the moment of maximum freshness, minutes after the content was presented. They say nothing about whether the learner can use that knowledge when a real situation demands it, or whether they still can a month later.
The gap between these metrics and actual behavior is well quantified. Saks and Belcourt (2006), surveying training professionals across 150 organizations, found that 62% of employees apply training material immediately after a program, 44% at six months, and 34% at one year. An end-of-course quiz samples learners at the top of that curve and is blind to everything that follows. Wilson Learning's review of the transfer literature relays a starker consensus estimate: only about 15 to 20% of learning investments result in measurable work performance change.
So a program can post excellent completions, glowing satisfaction, and high quiz averages while changing very little. The metrics are not wrong; they are aimed at the wrong level. Stakeholders are asking a behavior question, and the standard dashboard answers an activity question.
What do outcomes frameworks say counts as evidence?
Two frameworks dominate how organizations think about learning evaluation, and both draw the same line between measuring learning and measuring behavior.
Kirkpatrick's four levels, the standard model in corporate learning, run from Reaction (Level 1) through Learning (Level 2), Behavior (Level 3), and Results (Level 4). Wilson Learning's critique of common evaluation practice is that the model devotes three levels to measuring learning outcomes and only one to measuring performance outcomes, which is where the original investment case for the program lives.
Moore's outcomes framework (Moore, Green, and Gallis, Journal of Continuing Education in the Health Professions, 2009), developed for continuing medical education and used by accreditors such as the ACCME, expands the ladder into seven levels: participation, satisfaction, learning (split into declarative and procedural knowledge), competence, performance, and then patient health and community health. The framework works like a staircase: each level presupposes the one below it. Its most useful contribution for instructional designers is the distinction between competence, demonstrated when a learner shows how to apply knowledge in a realistic educational setting, and performance, demonstrated when the learner actually changes what they do in practice.
| Metric | Kirkpatrick | Moore | What it tells you | What it cannot tell you |
|---|---|---|---|---|
| Enrollment, completion | (pre-Level 1 usage) | Level 1: Participation | Reach of the program | Whether anyone learned or changed |
| Satisfaction survey | Level 1: Reaction | Level 2: Satisfaction | How learners felt about it | Whether it worked |
| Quiz, knowledge test | Level 2: Learning | Level 3: Learning (declarative/procedural) | What learners can recall or describe | Whether they can apply it in context |
| Decisions in realistic scenarios | (bridge to Level 3) | Level 4: Competence | Whether learners apply knowledge when situations demand it | Whether workplace conditions will allow it |
| On-the-job behavior data | Level 3: Behavior | Level 5: Performance | Whether practice actually changed | Why it did or did not change |
| Business or health outcomes | Level 4: Results | Levels 6-7: Patient and community health | Whether the change mattered | Which decisions drove it |
The frameworks agree on the direction of value: the evidence stakeholders care about lives above the learning level. Outcomes-measurement guidance built on Moore's framework adds two methodological principles that travel well beyond medicine: every measure should align with a predefined learning objective, and demonstrating change requires data from both before and after the educational intervention.
Why is behavior so hard to measure directly?
If behavior and results are what matter, why do so few programs measure them? Because direct observation of performance is expensive, slow, and often ambiguous.
On-the-job observation does not scale beyond small cohorts. System-of-record data, such as patient charts, EHR extracts, CRM activity, or quality audits, arrives months after the program and strips away context: it shows what was done but not what the person was weighing, what alternatives they considered, or why they chose as they did. And self-reported behavior change, the most common fallback, is so widely questioned that outcomes-measurement guidance in medical education recommends labeling it a "surrogate performance" measure rather than treating it as performance data.
The result is a systematic mismatch: programs report the levels that are easy to measure instead of the levels that matter. But the deeper issue is instructional, not analytic. A program built entirely from content delivery and quizzes generates no behavioral signal to capture in the first place. There is nothing to observe because learners are never asked to do anything except recall. Measurement at the competence level has to be designed into the learning experience itself.
How do realistic scenarios make behavior observable?
The core methodological insight is simple: a decision made in a realistic situation is behavior. When a learner works through a scenario that mirrors the situations the program targets, a patient presenting with ambiguous symptoms, a customer raising a pricing objection, a direct report pushing back on feedback, every choice they make is an observable act of judgment rather than a recall response. The scenario converts an invisible cognitive skill into a stream of recordable decisions.
That stream is what decision-level performance data looks like in practice:
- Which option each learner chose at each decision point, across every situation in the program.
- First-choice patterns, sometimes called initial competence: the share of learners who selected an optimal option on their first attempt, before receiving any feedback. This is the program's true baseline.
- Response to coaching: whether a learner recognized the optimal approach after expert mentoring, and how quickly.
- Change across situations: when learners face several similar situations in sequence, the shift in first-choice accuracy from the first situation to the last is a change-over-time measure taken on behavior, analogous to the before-and-after evidence Moore's framework calls for.
- Cohort and preference patterns: decisions can be segmented by role, specialty, or experience level, and where several options are defensible, the data shows which one each group actually favors.
One condition makes all of this work: comparability. Every learner has to encounter the same expert-designed situations and decision points, or one learner's choices cannot be lined up against another's and the analytics dissolve into anecdotes.
Published outcomes show what this data yields. In a continuing medical education activity on severe atopic dermatitis, described in the ACEhp Almanac (Seifert, 2021), 36% of learner decisions were initially competent in the first clinical situation. By the second situation initial competence reached 73%, and by the third it reached 89%, a 147% relative gain, with 97% of learners reporting greater confidence in developing treatment strategies. A companion program on chronic lymphocytic leukemia showed initial competency increasing fourfold across three situations on one practice gap, and an 85% relative gain on another. These outcomes are measured directly and objectively during the program, months before any practice-level data could exist. By strict framework mapping they report as Moore's Level 4 competence. Because they are observed decisions made in full clinical context rather than self-reports, however, there is a serious argument that they stand as a strong surrogate for Level 5 performance: at least as strong as the self-reported practice-change measures commonly accepted at that level, and stronger on the appropriateness dimension than practice data stripped of case context. That argument deserves its own treatment, and a forthcoming guide will make it in full.
How does just-in-time mentoring double as measurement?
There is a second, less obvious payoff to designing expert guidance into scenarios: the mentoring itself is a measurement instrument.
When corrective mentoring is delivered at the moment a learner makes a suboptimal choice, the same event serves both purposes at once. The learner gets the expert rationale in context, reflects, and revises their approach, which is the learning function. And the program records exactly where the correction was needed, which is the measurement function. Every mentoring event is a flag marking a specific situation, a specific decision, and a specific gap between what learners currently do and what experts do.
Aggregated, those flags produce a diagnostic map that no quiz can generate. The atopic dermatitis and CLL programs cited above were built around identified practice gaps, and the mentoring data showed precisely which clinical contexts clinicians struggled with before recognizing an optimal path, which gaps closed over successive situations, and where additional education was still needed. Distinguishing decisions that were competent on the first attempt from decisions that became competent only after mentoring turns a single program into a longitudinal record of skill developing, learner by learner and cohort by cohort.
This resolves an old tension in assessment design. Pure assessment produces clean data but leaves the learner uncoached; pure coaching helps the learner but produces no comparable record. Mentoring at the decision point does both.
How can instructional designers build measurement into a program?
Behavior-level measurement is a design decision made at the start of a program, not an analysis run at the end. Five practices follow directly from the frameworks and evidence above.
- Define objectives as observable decisions. Replace "understand" and "appreciate" with statements of the form "given this situation, the learner selects this course of action." An objective phrased as a decision can be measured; an objective phrased as a mental state cannot. This also satisfies Moore's requirement that every measure align with a predefined objective.
- Baseline before intervening. Learners' first choices, made before any feedback, are the baseline. In the atopic dermatitis program, the 36% initial-competence figure was not a disappointment; it was the denominator that made the eventual 89% meaningful.
- Measure change, not snapshots. A single competence score proves little. Repeated, comparable situations produce the pre/post trajectory that demonstrates the program itself moved the number.
- Instrument the feedback. Track where corrective mentoring was needed, not just final outcomes. That data tells the program team where gaps persist and what to build next.
- Report in framework language. Executives and accreditors already think in Kirkpatrick and Moore terms. Presenting decision-level data explicitly as competence evidence (Moore Level 4, the bridge to Kirkpatrick Level 3) connects learning analytics to the outcomes conversation stakeholders are actually having.
This is the measurement model that platforms built for the application step operationalize. AliveSim's Guided Scenarios place every learner in the same expert-designed situations and capture each decision, each first-choice pattern, and each response to just-in-time mentoring as comparable analytics. Across organizations using the platform, learners show a 2.5x improvement in decision-making performance against their own baseline first choices, and because that gain is captured inside the learning experience itself, it is a claim a program team can substantiate from its own data rather than from follow-up surveys.
None of this replaces on-the-job performance data where it exists; Moore's Level 5 and Kirkpatrick's Levels 3 and 4 remain the destination. But decision-level data from realistic scenarios is the earliest direct evidence that a program changed how people decide and act in the situations it targets, available while the program is still running, attributable to the program itself, and expressed in the currency stakeholders asked about in the first place: behavior.
References
- Kirkpatrick, D. L., & Kirkpatrick, J. D. (2006). Evaluating training programs: The four levels (3rd ed.). Berrett-Koehler.
- Leimbach, M. Learning transfer model: A research-driven approach to enhancing learning effectiveness. Wilson Learning Worldwide.
- Moore, D. E., Green, J. S., & Gallis, H. A. (2009). Achieving desired results and improved outcomes: Integrating planning and assessment throughout learning activities. Journal of Continuing Education in the Health Professions, 29(1), 1–15.
- Saks, A. M., & Belcourt, M. (2006). An investigation of training activities and transfer of training in organizations. Human Resource Management, 45(4), 629–648.
- Seifert, D. (2021). Incorporating skill development in CME via corrective mentoring. Alliance for Continuing Education in the Health Professions Almanac.
Related questions
What are Kirkpatrick levels 3 and 4?
In Kirkpatrick's four-level model, Level 3 (Behavior) asks whether learners actually changed what they do on the job after training, and Level 4 (Results) asks whether that change moved a business or organizational metric. They sit above Level 1 (Reaction, typically satisfaction surveys) and Level 2 (Learning, typically knowledge tests). Levels 3 and 4 are what stakeholders usually mean when they ask whether training worked, yet they are the levels least often measured, because observing on-the-job behavior and isolating training's effect on results both require deliberate measurement design rather than end-of-course data collection.
What is Moore's outcomes framework?
Moore's outcomes framework (Moore, Green, and Gallis, 2009) is an expanded taxonomy for evaluating education, developed for continuing medical education and adopted by accreditors. It defines seven levels: participation, satisfaction, learning (split into declarative and procedural knowledge), competence, performance, patient health, and community health. Its key contribution is separating competence, meaning the learner can show how to apply knowledge in a realistic educational setting, from performance, meaning the learner actually does it in practice. That distinction gives educators a measurable, observable outcome that can be captured during a program, before practice-level data exists.
Why are quiz scores a poor measure of training effectiveness?
Quiz scores measure recall of declarative knowledge at the moment of peak freshness, immediately after content is delivered. They sit at the learning level of both Kirkpatrick's and Moore's frameworks, below competence, behavior, and results. A learner can score perfectly on a quiz and still fail to apply the material in a real situation, because recognizing a correct answer is a different skill from choosing well amid realistic pressures and tradeoffs. Transfer research shows the gap: application of training decays from 62% immediately after a program to 34% a year later, a decline that end-of-course quizzes never detect.
What analytics should scenario-based learning capture?
Scenario-based learning should capture decision-level data: which option each learner chose at each decision point, the share of first choices that were optimal before any feedback (a baseline of initial competence), whether learners recognized the optimal approach after corrective mentoring, and how first-choice accuracy changed across successive similar situations (a pre/post gain). Cohort breakdowns by role, specialty, or experience level show where gaps concentrate, and preference data shows which of several acceptable options learners favor. For this data to be comparable across learners, every learner must face the same expert-designed situations and decision points.
Published September 16, 2025 · Updated July 10, 2026 · 11 min read