A conversation is playing out in front of you between 3D avatar characters, and it feels real, like one that would actually happen where you work. Then it stops and turns to you: how should this situation be handled? The call is not obvious. Several approaches all look like they could work, and you are not sure which way to go.
You never say a line. Yet you drive everything: the pace, the flow, each decision, what happens next, the consequences of your choices, and the feedback you get. You work through it until you can see for yourself what great looks like and what not-so-great looks like, arrive at the best approach, watch it play out, and understand why it works.
That experience is the reason to put more than one character in a scene, and it exists to do one job: set up situations where learners apply what they know and get feedback on their decision-making. It is easy to assume there are only two ways to use a character in learning. Either the character talks at you, which is a lecture with a face on it, or you talk to the character, which rehearses how to say something. The third option is the one people tend to miss: a conversation among several characters sets up the situation, the learner has to make the call, and the learner drives the whole experience, the conversation, the involvement, the feedback. It is a designed experience with measurable outcomes. Everything below explains it.
What the learner is actually deciding
Start with the part that is easiest to get wrong. The learner's job is not to decide where the conversation goes. This is not choose-your-own-adventure dialogue steering, picking the next line and nudging the chat down a branch.
The learner is deciding how best to handle the situation. The conversation sets the situation up; the decision is about the situation itself. Sometimes the conversation is the essential element, a difficult exchange the learner is guiding, and they watch it evolve with each choice. Often it is not: in a clinical scenario, the conversation may set the stage, and the decision is about what to do next for the patient. Either way the learner weighs a finite set of approaches, gets feedback on which are strong and which fall short, and decides which to move forward with. The point is the strategic call and the feedback on it, carried by a natural conversation, the kind they could have in an office with a colleague, a mentor, or a coach.
1
The situation
A conversation among the characters sets it up, and the learner is involved from the start.
2
The decision
The learner makes the call, choosing among approaches. More than one can be optimal.
3
The consequences
The choice plays out. The learner sees what great and not-so-great look like.
4
The feedback
Nuanced, domain-specific, and about the why, at the moment of the choice.
Then the next situation begins.
How it works: the conversation engine
What makes this possible is a conversation engine. The conversation among the characters is dynamic, not a fixed clip of two characters talking. It sets up the situation, responds to the learner's decisions, and returns feedback based on what the learner chose. The situation is set up, the learner faces the decision point, makes the call, sees the consequences, learns from the feedback, and moves on to the next situation: a learning loop.
The characters also involve the learner directly. They can look at the learner, address them, and ask for their read, so the experience never feels like a video playing at a distance. A character who speaks to the learner in a natural, conversational register invites the social processing that a formal narration never gets, which is Mayer's personalization principle and his work on social agency and embodiment (Mayer, 2008, 2024). A character can even say something privately to the learner alone, which has uses of its own, covered in breaking the fourth wall. And the learning that happens while characters converse is genuine: observers of a good tutorial dialogue have been found to learn as much as the people in it, with probing why-questions as the active ingredient (Chi, Roy & Hausmann, 2008; Craig, Gholson, Ventura & Graesser, 2000; Craig, Sullins, Witherspoon & Gholson, 2006).
This is what more than one character buys: a situation with real people in it, a decision that belongs to the learner, and feedback that arrives inside the scene.
Two things this gets confused with
Two other formats look similar from a distance, and it is worth drawing both lines clearly.
An AI-generated video is still a lecture. A generated conversation between characters, or a generated character talking at the learner, is a lecture no matter how it is produced. That holds even when the character looks straight at the learner and delivers the line like a talking head. There is no situation to handle, nothing to decide, and no feedback that responds to a decision. What we described above is an experience the learner drives.
An AI role-play is a valuable tool for a different job. You may be asking: is the learner actually speaking to the characters? In an AI role-play, yes, and that is a genuinely different application. The character plays a role and speaks in it, and these applications carry design of their own: a character with an opinion, an approach, a demeanor, and often a rubric behind the scoring. What they are built for is rehearsing delivery, how to say something. The content of the exchange is freeform conversation rather than instructional design that sets up a specific situation, so the domain information the character carries is limited, every learner's exchange is different and unknowable in advance, and the analytics describe the speech. Deciding what to do in a nuanced situation, with domain-specific expert feedback on the decision, is the other job, and the full argument between the two lives in scenario-based learning versus AI role-play.
Why the choices are finite
Each decision point offers a finite set of options, and that is a strength, not a limitation. The goal is not freeform, do-whatever-you-want expression. It is to help the learner see the approaches that are typically taken in that kind of situation, recognize which options are optimal for this particular one, and understand why. More than one option can be optimal, and an option that is optimal here may be the weaker call in the next situation, which is exactly the judgment the design is meant to build.
What multiple characters add
Two things come along with the multi-avatar approach, beyond the experience itself.
Retention rises when knowledge is applied rather than only received, which is a point in its own right and is carried by the guide on training and knowledge retention.
And this can be added alongside an existing program rather than replacing it. It complements what is already there, giving learners a place to apply the knowledge and giving the organization evidence that they applied it in the real situations they will actually face. The guide on proving the return on education is where that evidence story lives.
How AliveSim builds this
Set up realistic situations, let learners apply what they know, and land feedback on every decision: that is the approach, and the most natural way to carry it is a conversation among 3D avatar characters that the learner is inside of. AliveSim Guided Scenarios run exactly that: multiple characters in a dynamic conversation, situations drawn from the learner's own work, a finite set of options at each decision, several of which can be optimal, and coaching that arrives at the moment of choice. For how that fits together, see the AliveSim platform.
The takeaway
Start where this article started: the job is to set up situations where learners apply their knowledge and get feedback on their decision-making. A single character talking at a learner cannot do that job. A one-on-one role-play does a different one, rehearsing delivery. A conversation among several characters can do it: it sets up the situation, involves the learner in it, hands them the decisions, and brings the feedback to the moment of choice. The learner sees what great looks like and what not-so-great looks like, and why the best approach is best, and they walk away having applied the knowledge, not just heard it.
References
- Chi, M. T. H., Roy, M., & Hausmann, R. G. M. (2008). Observing tutorial dialogues collaboratively: Insights about human tutoring effectiveness from vicarious learning. Cognitive Science, 32(2), 301–341.
- Craig, S. D., Gholson, B., Ventura, M., & Graesser, A. C. (2000). Overhearing dialogues and monologues in virtual tutoring sessions: Effects on questioning and vicarious learning. International Journal of Artificial Intelligence in Education, 11, 242–253.
- Craig, S. D., Sullins, J., Witherspoon, A., & Gholson, B. (2006). The deep-level-reasoning-question effect: The role of dialogue and deep-level-reasoning questions during vicarious learning. Cognition and Instruction, 24(4), 565–591.
- Mayer, R. E. (2008). Applying the science of learning: Evidence-based principles for the design of multimedia instruction. American Psychologist, 63(8), 760–769.
- Mayer, R. E. (2024). The past, present, and future of the cognitive theory of multimedia learning. Educational Psychology Review, 36, 8.
Related questions
What does the learner actually do in a multi-avatar scenario?
They apply their knowledge. A conversation among the characters sets up a realistic situation, and the learner decides how to handle it: they weigh the approaches on offer at each decision point, commit to one, see the consequences play out, and receive feedback on the call they made. They also control the pace, and the characters involve them directly, looking at them, addressing them, asking for their read. The learner has no speaking part and loses nothing by it: every meaningful act in the experience, every decision, belongs to them.
How is this different from an AI-generated video of characters talking?
A generated video is watched from the outside, however good it looks. Fast text-to-video and talking-head tools can make one character or several speak from a script, and the character can even look straight at the learner while delivering the line. There is still no situation to handle, no decision to make, and no feedback that responds to a choice. A multi-avatar scenario is driven by the learner: the conversation sets up the situation, reacts to the decisions, and returns feedback on each one.
Is this the same as an AI role-play app where you talk to a character?
No, and both are valid tools with different jobs. In an AI role-play the learner speaks with a character that plays a role, and these applications carry real design of their own: a character with an opinion, an approach, a demeanor, and often a rubric behind the scoring. What they are built for is rehearsing delivery, how to say something. The content of the exchange is freeform conversation rather than instructional design that sets up a specific situation, so the domain information the character carries is limited, every learner's exchange is different and unknowable in advance, and the analytics describe the speech. Deciding what to do in a nuanced situation, with domain-specific expert feedback on the decision, is a different job, and the full comparison lives in our guide on scenario-based learning versus AI role-play.
What is the learner deciding, if not where the conversation goes?
How best to handle the situation. This is not choose-your-own-adventure dialogue steering. The conversation sets up the situation; the learner then chooses among a finite set of approaches, gets feedback on the strength of each, and decides which to move forward with. Sometimes the conversation itself is the essential element, and the learner watches it evolve with their choices. Often, as in a clinical decision, the conversation sets the stage and the decision is about what happens next in the situation. More than one option can be optimal, and the goal is recognizing which options are optimal here, and why an option that is optimal in this situation might not be in another.
Published July 20, 2026 · 7 min read