Ask any accredited provider what they can show at Moore's Level 5, meaning performance, evidence that clinicians changed what they do in practice (Moore, Green, and Gallis, 2009), and you will hear about one of two measurement routes. The field is reporting higher-level outcomes more than it used to: the ACCME's 2024 Annual Data Report shows 46% of activities now assessing performance, against 95% assessing competence. But both routes into that 46% deserve more scrutiny than they usually get, because both are weaker than the field's reporting conventions suggest. This article makes the case for a third evidence type, observed decision-making in simulated clinical scenarios, as the most potent Level 5 surrogate available, and gives you the reporting language to use it. (For the Moore framework itself, and how it maps to Kirkpatrick, see our guide to measuring behavior change.)
Route 1: Practice data is activity without context
The aspirational route is real-world practice data: EHR extracts, administrative claims, chart abstraction, registries. The field's own guidance is candid about the access problem. The Alliance's Beginner's Guide to Measuring Educational Outcomes in CEhp notes that clinical care documents "tend to be outside of the control of the CME planner"; that EHR limitations "have been the limited interoperability available for extracting data, and the lack of complete, or codified data"; that claims data omits physiological elements entirely; and that chart abstraction is "laborious and costly." A NAM Perspectives discussion paper reached the same conclusion at the system level: assessing outcomes beyond satisfaction, pre/post knowledge, and commitment to change "is time-consuming and may be logistically or methodologically challenging and is thus underutilized" (Price, Davis, and Filerman, 2021). In practice, providers frequently cannot get the data at all. Educational platforms often cannot even share participant contact information for follow-up, let alone practice records.
The deeper problem is what the data means when you do get it. Utilization data shows what was done: more of drug X prescribed, more of procedure Y ordered. It carries no presentation, no comorbidities, no contraindications, and no view of the alternatives the clinician weighed. It cannot say whether the decision was appropriate for that patient. This is a documented measurement failure, not an intuition. Claims and administrative data "are not suitable to capture the nuances of appropriateness of care," and can misclassify clinically appropriate decisions precisely because patient context is integral to judging them (Isaac et al., 2018). A count of prescriptions is a measure of activity, not of judgment. Yet judgment, the right decision for the given patient, is what education targets and what quality actually is.
Even as a record of what happened, the chart drops information. When Peabody and colleagues measured physicians' quality of care against unannounced standardized patients, the gold standard in that literature, chart abstraction scored lowest of the three methods tested: 65.6% against the standardized-patient benchmark of 76.2% in the 2000 study, and 63% against 73% in the 2004 multisite replication. Clinical vignettes, case-based decision measures with controlled context, consistently produced scores closer to the standardized-patient benchmark than chart abstraction did, a pattern robust across conditions, case complexity, site, and physician training level (Peabody et al., 2000; 2004). The 2004 study concluded that vignettes are "a valid tool for measuring the quality of clinical practice" and "may be useful for longitudinal evaluations of interventions intended to change clinical practice," which is precisely CME's job description.
Route 2: Self-report, the surrogate the field already accepts
The practical route runs through two instruments that reporting often blurs together, and the distinction matters. Commitment-to-change statements collected at activity close capture intent, and intent is competence territory; a plan to change leans Level 4. The actual Level 5 claim rests on the follow-up survey sent 30 to 90 days later, asking clinicians what they in fact changed. The Alliance's own assessment guidance lists, as the Level 5 question type, rating-scale items "designed to assess incorporation of planned practice changes." That is self-report. Notably, in the same guidance table, the observed-measure column at Level 5 is empty.
Two problems compound here. The first is accuracy: a systematic review in JAMA found physician self-assessment correlates poorly with observed measures of competence and performance (Davis et al., 2006). The outcomes community has internalized this, which is why the recommendation is to label self-reported performance change "surrogate performance," because its accuracy is "broadly questioned within the profession." The second is coverage. Anyone who has fielded one of these follow-up surveys knows the response rates: often 10% of learners or less, skewed toward the most engaged. The Alliance's Beginner's Guide acknowledges that "fewer learners will continue to participate in and respond to the testing" as time passes, and its own worked case study makes the point vividly: in an AAFP program that added a delayed retention test, roughly 6% of learners completed the full test sequence, a figure the authors offered to show "how difficult it can be to obtain robust sample sizes in this type of measurement."
So the current state of Level 5 measurement is a choice between context-free activity data that is hard to obtain and demonstrably incomplete, and recollections that the field itself flags as unreliable, gathered from a small self-selected slice of the audience.
What observed simulated performance offers
Observed simulated performance is exactly what the name says: clinical decision-making captured while it happens, inside an authentic simulated scenario. The clinician evaluates a fully specified patient presentation, weighs therapeutic options, commits to decisions, and responds to expert mentoring, repeatedly, across varied presentations. Every decision is:
- Observed, not recalled. Captured as it happens, not reconstructed on a survey weeks later. This is the exact weakness Davis et al. documented in self-report.
- Context-complete. The case is authored, so appropriateness is measurable by design. This is the exact dimension Isaac and colleagues document claims data missing.
- Baselined. First choices before any feedback establish where each clinician started; change is measured against it.
- Standardized across learners. Every clinician faces the same decisions, so the measure is identical across cohorts, specialties, and experience levels, not each learner reporting on a different practice.
| Self-reported change | Utilization data | Observed simulated decisions | |
|---|---|---|---|
| Observed as it happens? | |||
| Full case context included? | |||
| Baselined against first choices? | |||
| Comparable across a cohort? |
This is performance in every sense but the venue: real decision behavior, on case-based measures of the kind Peabody validated for measuring the quality of clinical practice, minus only the practice setting. The field has already begun moving in this direction. Published CME studies using virtual patient simulation argue explicitly that treatment decisions replicating deliberate practice in a clinical environment provide "objective and in-depth insight into clinician practice behaviour" (Lucero et al., 2020).
The position: observed decision-making within authentic simulated scenarios is the most potent surrogate for Moore's Level 5. It is at least as strong as the self-reported measures routinely accepted at that level, and stronger than context-free utilization data on the dimension that matters most, appropriateness. Level 5's definition predates learning environments where clinicians perform rather than report; the framework has not yet caught up with what can now be measured.
The evidence, through two lenses
Syandus has measured the educational impact of this approach directly, through two independent instruments.
The first lens is pre/post vignette performance. Across 15,197 posttest learner responses spanning 14 therapeutic areas, with all data in the analysis period included (no selection), clinicians performed 2.58 times better on clinical vignette questions after learning through AliveSim's Guided Scenarios. The effect size is very large by any conventional standard (Cohen's d = 2.71, where 0.8 is conventionally considered large; p < 0.001). For calibration, the benchmarks CME outcomes practitioners cite put the pooled effect of live CME on knowledge outcomes at roughly d = 0.79, and internet-based learning in the health professions at about 1.0 (Olivieri, 2018).
The second lens replicates the finding inside the simulation itself. In a matched decision-point analysis of 31,673 decisions across seven programs, clinicians showed a 2.85x improvement in decision-making when they met the same decisions in changed situations later in the simulation (d = 3.95, p < 0.001). Because those later situations differ from the earlier ones, the improvement is not rote memory of a scenario; what improved is the handling of the gap itself, evidence that expert mentoring was applied, not merely received.
Program-level data shows the same shape, with initial competence rising across sequential scenarios while required corrective mentoring falls. In the published severe atopic dermatitis program, that meant a 1.5x relative increase in initial competency alongside a 57 to 61% relative decrease in mentoring required.
Two different instruments, measuring at different points, converging on the same answer. And because every learner faces the same expert-designed decision points, the dataset supports analyses unavailable to either surveys or utilization data: baseline decision patterns, including where clinicians initially select inappropriate clinical options and for which presentations; how decision quality improves after expert mentoring, decision by decision; which presentations require the most intensive mentoring before optimal choices emerge; how clinicians select among multiple valid therapeutic options and how those preferences evolve; and where cohorts diverge from predicted patterns, flagging where additional education is most valuable.
Reporting guidance
Get the framing right, because the framing is the argument. Nobody directly observes clinicians in their practices at scale. Every form of Level 5 evidence in common use is therefore a surrogate: self-reported practice change, which the field itself labels "surrogate performance," and utilization data with no case context. The question an outcomes report should answer is not "is this Level 4 or Level 5?" It is "which surrogate for practice performance carries the most evidentiary weight?" On that question, observed, context-complete, baselined clinical decisions are an equal or superior surrogate to both accepted alternatives.
Proposed language for outcomes reports and partner conversations:
"Participants demonstrated measured improvement in clinical decision-making performance within authentic simulated scenarios. Because clinicians cannot be directly observed in practice at scale, all common Level 5 evidence is surrogate evidence: self-reported practice change, whose accuracy the field itself questions (Davis et al., JAMA 2006), or utilization data that shows what was done without the case context to judge whether it was appropriate (Isaac et al., 2018). We report observed simulated performance as a Level 5 surrogate of equal or greater evidentiary strength: decisions that are directly observed rather than recalled, context-complete rather than context-free, and taken on case-based measures validated against the quality of actual clinical practice (Peabody et al., 2000, 2004)."
Two boundaries worth keeping
First, Peabody is cited here for the validated principle, namely that case-based measurement with controlled context outperforms the medical record for measuring quality of care at the group level, not as a validation of any specific instrument format. His validated instruments were open-ended, full-visit simulations. Sequential in-scenario decisions with consequences are a richer measure than static multiple-choice items, but that equivalence has not been independently validated, and a careful outcomes report should not imply it has.
Second, the word is surrogate, and it stays that way. The know-do gap literature is a standing reminder that measured capability can overestimate what happens in a clinician's practice, and vignette-to-practice fidelity is context-dependent. The claim is "most potent surrogate," never "equivalent," and the claim is enough: measured against the alternatives providers actually report at Level 5, thousands of observed, context-complete, baselined clinical decisions win on the merits. The field's frameworks should evolve to say where such evidence belongs.
References
- Accreditation Council for Continuing Medical Education. (2025). 2024 Annual Data Report: Evolving Impact in a Shifting Landscape.
- Bird, G. C., et al. (2015). Beginner's guide to measuring educational outcomes in CEhp. Alliance for Continuing Education in the Health Professions Almanac series.
- Davis, D. A., et al. (2006). Accuracy of physician self-assessment compared with observed measures of competence: A systematic review. JAMA, 296(9), 1094–1102.
- Isaac, T., et al. (2018). Measuring overuse with electronic health records data. American Journal of Managed Care, 24(1), 19–25.
- Lucero, K. S., et al. (2020). Virtual patient simulation in continuing education. Journal of European CME, 9(1), 1836865.
- Moore, D. E., Green, J. S., & Gallis, H. A. (2009). Achieving desired results and improved outcomes: Integrating planning and assessment throughout learning activities. Journal of Continuing Education in the Health Professions, 29(1), 1–15.
- Olivieri, J. (2018). Outcomes à la carte: A menu of frequently asked questions regarding CME outcomes. Alliance for Continuing Education in the Health Professions 2018 Annual Conference.
- Outcomes Standardization Project. Standardized outcomes glossary: Moore's Level 5, Performance.
- Peabody, J. W., Luck, J., Glassman, P., Dresselhaus, T. R., & Lee, M. (2000). Comparison of vignettes, standardized patients, and chart abstraction: A prospective validation study of 3 methods for measuring quality. JAMA, 283(13), 1715–1722.
- Peabody, J. W., et al. (2004). Measuring the quality of physician practice by using clinical vignettes: A prospective validation study. Annals of Internal Medicine, 141(10), 771–780.
- Price, D. W., Davis, D. A., & Filerman, G. L. (2021). Systems-integrated CME. NAM Perspectives.
Related questions
What counts as Moore's Level 5 evidence?
Moore's Level 5 (performance) asks whether clinicians changed what they do in practice after an educational activity. In common use, two evidence types support the claim: real-world practice data (EHR extracts, claims, chart abstraction, registries) and self-reported practice change collected by follow-up survey weeks after the activity. Both are surrogates for direct observation, which does not happen at scale. Practice data is hard to obtain and carries no case context, so it cannot establish whether a decision was appropriate for the patient. Self-report is accessible but unreliable enough that outcomes guidance recommends labeling it surrogate performance. Observed decision-making in authentic simulated scenarios is a third surrogate: directly observed, context-complete, and baselined, which is why there is a strong argument for reporting it at Level 5 with the framing made explicit.
Why is self-reported practice change considered weak evidence?
Two reasons, one about accuracy and one about coverage. On accuracy, a systematic review in JAMA (Davis et al., 2006) found that physician self-assessment correlates poorly with observed measures of competence and performance, which is why outcomes guidance recommends calling self-reported practice change a surrogate performance measure rather than performance data. On coverage, the Level 5 version of self-report is the follow-up survey sent 30 to 90 days after the activity, and in practice those surveys draw a small, self-selected fraction of the cohort: practitioners commonly see follow-up response rates near 10% of learners or below, and in one Alliance case study only about 6% of learners completed a full delayed test sequence. The result is an unreliable instrument answered by the most engaged few.
Are clinical vignettes a valid measure of real practice?
At the group level, yes, and the validation is strong. Peabody and colleagues compared three measurement methods against unannounced standardized patients, the gold standard for measuring quality of care. Clinical vignettes scored consistently closer to that standard than chart abstraction did, in both the original 2000 JAMA study (vignettes 71.0% vs. charts 65.6%, against a standardized-patient benchmark of 76.2%) and the 2004 multisite replication (68% vs. 63%, against 73%). The 2004 study concluded vignettes are a valid tool for measuring the quality of clinical practice and may be useful for longitudinal evaluations of interventions intended to change it. The validated principle is case-based measurement with controlled context; Peabody's instruments were open-ended full-visit simulations, so the validation attaches to the principle rather than to any specific response format.
How should a provider report simulation outcomes at Level 5?
Report them as a Level 5 surrogate, with the surrogate framing stated rather than hidden, because the framing is the argument. All common Level 5 evidence is surrogate evidence: self-reported practice change or context-free utilization data. Language that works: participants demonstrated measured improvement in clinical decision-making performance within authentic simulated scenarios, reported as a Level 5 surrogate of equal or greater evidentiary strength, because the decisions are directly observed rather than recalled, context-complete rather than context-free, and taken on case-based measures validated against the quality of actual clinical practice. Keep the claims at the cohort level, and say potent surrogate, not equivalent.
Published July 17, 2026 · 11 min read