When Roxana Moreno was a graduate student in Richard Mayer's lab, she brought him an idea he later called preposterous: that writing an online science lesson in casual, conversational language, rather than formal prose, would make people learn more. Mayer was skeptical. They ran the experiments anyway, and every one came back positive. After the effect replicated across topic after topic, he accepted it and named it the personalization principle (Mayer, 2024). It also changed his theory. Learning from a voice on a screen, he concluded, is not only a cognitive event but a social one.
That story matters for anyone building scenarios with characters who talk, because it turns something most people treat as a matter of taste into one of the most-replicated results in the science of learning. Conversational wording beat formal wording in all eleven experiments Mayer reviewed in 2008, with a large effect (d = 1.11), and it still holds at d = 1.00 in his 2024 review (Mayer, 2008; Mayer, 2024). A character speaking in plain, human language is not a style choice. It is one of the best-supported moves in the literature.
Conversation, though, is more than word choice. This guide takes four things a good conversational scenario does and shows what Mayer's evidence says about each. It draws on three sources: Mayer and Moreno's 2003 account of how cognitive load works, Mayer's 2008 summary of evidence-based design principles, and his 2024 review of the theory after four decades. The four characteristics are simple to state. The learner controls the pace through their own actions. The spoken words and the visuals work together rather than compete. The conversation stays casual rather than formal. And everything the learner needs for a decision is on the table before the decision arrives.
Four characteristics of a conversational scenario
Each is a design choice, and each is backed by one of Mayer's evidence-based principles.
1
The learner controls the pace
The learner advances the scene by clicking and deciding. Nothing moves on a timer, so each choice is a place to consolidate.
Segmenting
2
Spoken words and visuals work together
The voice carries the words; the screen shows the scene. A visual goes up only when it reinforces what is said, never a caption that competes.
Modality, multimedia, contiguity, signaling
3
The conversation stays casual
The character speaks in natural, human language, in a human voice, and behaves like a person rather than a diagram.
Personalization, voice, embodiment
4
Everything is on the table first
The terms and the situation arrive before the decision that turns on them. When the learner decides, the talking stops.
Pretraining
The learner controls the pace through their own actions
In a good scenario the learner is the one who moves it forward: a click to continue, a click on a character to hear more, a decision among options. Nothing advances until the learner advances it. Mayer's segmenting principle is exactly this, and his tested mechanic is a button. In his lightning study, a continuous narrated animation was broken into sixteen segments, and a CONTINUE button appeared at the end of each; the next segment played only when the learner clicked it (Mayer, 2008; Mayer & Moreno, 2003). Learners who advanced the lesson themselves outperformed those who watched the same content run straight through, in three of three experiments (d = 0.98) and by as much as d = 1.36 in a single study, and the effect holds in the 2024 review (Mayer, 2008; Mayer & Moreno, 2003; Mayer, 2024).
The reason is timing. When material is rich and arrives at a fixed pace, the learner can take in the incoming words and images but has no time left to organize and connect them before the next stretch begins (Mayer & Moreno, 2003). A learner-set pace gives that time back. A conversational scenario has the mechanism built in: its decisions are its segment boundaries. Each choice is a natural place to stop and consolidate what just happened, so the scene should wait for the learner rather than run on a timer.
The spoken words and the visuals work together, never compete
To see why this matters, you have to be precise about how a lesson gets in, and the split is not words against pictures. It is which sense does the intake. A multimedia message enters through the eyes and the ears: printed words and graphics are held first in visual memory, and spoken words in auditory memory (Mayer, 2024). So spoken words ride the ears into one channel, while everything on the screen, the picture and any printed text alike, enters through the eyes into the other. Printed words are still words, but because they come in through the eyes they land in the visual channel and compete with the picture there. That is the real reason on-screen text can hurt: not because it is words, but because it is words crowding a picture for one channel.
A character speaking to the learner uses both channels the way the evidence rewards. The voice carries the words through the ears; the scene occupies the eyes. Mayer's modality principle, the single most-supported result in his program, found that people learn better from graphics with spoken words than from graphics with printed words, in seventeen of seventeen experiments (d = 1.02) (Mayer, 2008; Mayer, 2024). Reprinting the character's spoken line as a caption at the bottom of the screen undoes that gain: the caption duplicates a stream the learner is already hearing and competes with the scene for the eyes.
This is the redundancy principle, and its boundary is exact. What Mayer tested was a verbatim caption of the narration running while a picture competed for the eyes; that version did not help and hurt when the lesson was fast-paced (d = 0.10 in the 2024 review, d = 0.72 in 2008) (Mayer, 2008; Mayer, 2024). The principle does not ban all on-screen text. Mayer and Moreno report the opposite case plainly: when no picture is competing, concurrent narration and on-screen text can beat narration alone, because the text has nothing to fight for in the visual channel (Mayer & Moreno, 2003). Captions provided for accessibility are a separate and necessary matter. What the evidence warns against is a duplicate caption of the dialogue running over a scene that needs the eyes.
So a visual is not the enemy of speech. The question is whether it reinforces what is being said or competes with it, and four of Mayer's principles say when a visual earns its place while a character talks.
- A diagram the character talks through is the picture that goes with the words. The multimedia principle, that people learn better from words and pictures than from words alone, is Mayer's single strongest reason to add a visual, not withhold one (d = 1.39 in 2008, eleven of eleven; d = 1.35 in 2024) (Mayer, 2008; Mayer, 2024).
- Timing decides whether it helps. Temporal contiguity finds that corresponding words and pictures should appear at the same time, not one after the other (d = 1.31, eight of eight) (Mayer, 2008). The diagram element, the defined term, or the study name should come up while the character is saying it, not before or after.
- A study name or a key term appearing on screen at the moment it is spoken is a signal. Signaling means adding a cue that points the learner to the essential parts and shows how they fit together, such as stressing a key word, a heading, or a colored arrow on a diagram (d = 0.52 in 2008, six of six; d = 0.69 in 2024) (Mayer, 2008; Mayer, 2024).
- Put the label next to the thing it names. Spatial contiguity finds that words and the part of the picture they describe should sit near each other, not in a caption bar off to the side (d = 1.12 in 2008, five of five) (Mayer, 2008).
The same logic decides what to leave off the screen. Mayer's coherence principle, which he also calls weeding, found that cutting interesting but irrelevant detail improves understanding (d = 0.97 in 2008, thirteen of fourteen) (Mayer, 2008). While a character speaks, only what reinforces the point should go up; decoration added to dress up the scene competes for the same channel and should go.
The line between a reinforcing label and a redundant caption is the difference between two principles. A short cue that highlights the essential item is signaling, a positive effect. A verbatim stream of the dialogue printed over a competing picture is redundancy, the one to avoid. Mayer's own experiments tested the verbatim-caption case, not the short-label case, so treat the permission for a reinforcing label as the signaling principle applied together with the no-competing-picture exception, rather than a directly tested result.
The conversation stays casual, not formal
Here is where the opening story pays off. Personalization, the finding that conversational wording beats formal wording, is not only large (d = 1.11 in 2008, eleven of eleven; d = 1.00 in 2024) but among the most-replicated effects Mayer has studied, and it is one he resisted until his own lab forced his hand (Mayer, 2008; Mayer, 2024). His example is almost mundane: take a script about the respiratory system and change each "the" to "your," and learning improves (Mayer, 2008). The voice principle adds the same lesson from the audio side, that a friendly human voice beats a machine voice (d = 0.74 in 2024) (Mayer, 2024). Both are principles Mayer added after 2003, once the replications convinced him that learning from someone on a screen is a social act as much as a cognitive one. A character who talks to the learner in natural language, in a human voice, is doing what these two principles reward.
Two other findings are often misread as evidence that the character itself does not matter. They do not say that. The image principle found that parking a static picture of the speaker on the screen does not reliably help (d = 0.20). The immersion principle found that the same lesson in 3D virtual reality does not beat a matched 2D version (d = -0.10, essentially zero and slightly negative) (Mayer, 2024). Both are about scenery: a frozen face, a fancier display. Neither is about a character who behaves like a person. The principle that is about behavior points the other way. Embodiment found that an on-screen character who gestures, moves, and makes eye contact helps meaningfully more than a flat, inert one (d = 0.58) (Mayer, 2024). A cardboard cutout is exactly the low-embodiment case that loses. So the evidence does not flatten the character. It tells you what makes one earn its place: not that it is visible, but that it acts like someone in the room. An animated, conversational character that gestures and reacts is the high-embodiment case the evidence rewards, not the static face it is indifferent to.
Everything the learner needs is on the table before the decision
The last characteristic is about sequence. When a decision comes up, the learner should already know the terms, the players, and the situation the decision turns on. Mayer's pretraining principle is the evidence: people learn better from a complex lesson when they already know the names and characteristics of its main concepts, in five of five experiments (d = 0.85) (Mayer, 2008). The mechanism is capacity. A learner still working out what a term means has less left for the reasoning the decision requires; a learner who already knows the term can spend everything on the judgment (Mayer & Moreno, 2003).
The word "pretraining" suggests a separate lesson beforehand, but the principle does not require one. Mayer and Moreno define it as "a specific sequencing strategy in which components are presented before a causal system is presented" (Mayer & Moreno, 2003), and Mayer's 2024 statement is simply that people do better when they know the names and characteristics of the main concepts (Mayer, 2024): a fact about what the learner knows when the material arrives, not about delivery format. A scenario can own that sequence itself. It can roll out the terms and the situation earlier in the same scenario, before the decision that depends on them. One caveat is worth stating plainly: every pretraining experiment used a discrete lesson ahead of the main animation, so applying the definition to a scenario's internal sequence follows the mechanism Mayer measured, not a study that tested that exact arrangement.
A design point follows from the same mechanism, offered as guidance rather than as a Mayer result. When the decision is up, the characters should stop talking. New speech arriving while the learner decides is fresh material competing for the capacity the decision itself needs, the overload Mayer and Moreno describe when new content keeps coming before the last is processed (Mayer & Moreno, 2003). Mayer never ran this experiment, so this is his mechanism extended to a decision moment, not a finding. The rule is plain: put everything on the table, then let nothing compete with the act of deciding.
One boundary is worth naming and setting aside. Mayer and Moreno list a ninth cognitive-load method, individualizing, which is not a design rule at all but an observation that high-spatial learners benefit more from some layouts than low-spatial learners do (d = 1.13). They say so directly, that it "is not technically a design method for reducing cognitive load but rather a way to select individual learners" (Mayer & Moreno, 2003). It is a caveat about who learns, not a characteristic to design for.
The evidence behind each characteristic
| Characteristic | Mayer's evidence |
|---|---|
| The learner controls the pace | Segmenting: learner-clicked CONTINUE, d = 0.98 across three of three (Mayer, 2008), as high as 1.36 in a single study (Mayer & Moreno, 2003). |
| Spoken words and visuals work together | Modality (d = 1.02, seventeen of seventeen), multimedia (1.39), temporal contiguity (1.31), signaling (0.52), spatial contiguity (1.12), coherence/weeding (0.97) (Mayer, 2008); redundancy (0.10, fast-paced) marks the one thing to avoid (Mayer, 2024). |
| The conversation stays casual | Personalization (1.11, eleven of eleven), voice (0.74), embodiment (0.58) (Mayer, 2008; Mayer, 2024). |
| Everything is on the table first | Pretraining (0.85, five of five) framed as sequencing, not a separate lesson (Mayer, 2008; Mayer & Moreno, 2003). |
Where this leaves the build
Stated together, these four characteristics describe a way of presenting a lesson, not a particular product: any format can honor them or ignore them. A scenario that honors them lets a character speak in natural language without a duplicate caption fighting the scene for the eyes, advances at the pace the learner sets by deciding, brings a reinforcing visual up in sync when it helps, and puts the terms and the situation in front of the learner before the decision that turns on them.
AliveSim's Guided Scenarios are built this way. The learner enters a realistic situation with 3D avatar characters who speak in natural, conversational language; reads the situation and makes real decisions among plausible options; and gets mentoring at the moment of each choice, so the learner can learn to apply the concept while the reasoning is still live. You can see how the pieces fit together on the AliveSim platform.
One last distinction is worth drawing. Mayer's principles are about cognitive load, how much of the learner's limited attention the presentation consumes. They are not about how hard the learner is thinking, which is a separate question covered in what more engaging actually means. A scenario can be clean in its delivery and still fail to ask enough of the learner. Getting the conversation right keeps the delivery out of the way; the design still has to make the thinking worth doing.
References
- Mayer, R. E. (2008). Applying the science of learning: Evidence-based principles for the design of multimedia instruction. American Psychologist, 63(8), 760–769.
- Mayer, R. E. (2024). The past, present, and future of the cognitive theory of multimedia learning. Educational Psychology Review, 36, 8.
- Mayer, R. E., & Moreno, R. (2003). Nine ways to reduce cognitive load in multimedia learning. Educational Psychologist, 38(1), 43–52.
Related questions
Is conversational dialogue in a scenario a style choice, or does it actually help learning?
It helps, and the evidence is unusually strong. Mayer's personalization principle finds that people learn more from a lesson written in conversational style than from the same lesson in formal style. It held in all eleven experiments Mayer reviewed in 2008, with a large effect (d = 1.11), and it still holds at d = 1.00 in his 2024 review. Mayer himself was skeptical of the idea until his lab replicated it across topic after topic, and it changed his theory: learning from a voice on a screen is a social event, not only a cognitive one. His voice principle adds the same lesson from the audio side, that a friendly human voice beats a machine voice. So a character speaking to the learner in plain, human language is doing what the evidence most strongly rewards.
Should a character's spoken words also appear on screen as text?
Not as a duplicate caption. Mayer's modality principle, the most-supported result in his program (17 of 17 experiments, d = 1.02), finds that people learn better from graphics with spoken words than from graphics with printed words, because printed text competes with the scene for the visual channel. His redundancy principle finds that printing the narration verbatim on screen while a picture competes does not help and hurts when the lesson is fast-paced. The exception is real and Mayer states it: when no picture is competing for the eyes, concurrent narration and on-screen text can beat narration alone. And captions provided for accessibility are a separate, necessary matter. The rule of thumb for a scenario: let the voice carry the dialogue and keep the screen for the situation, rather than reprinting the line the learner is already hearing.
When does a visual help while a character is talking, and when does it get in the way?
A visual helps when it reinforces what is being said and appears in step with it. Mayer's multimedia principle is his strongest reason to add a picture: people learn better from words and pictures than from words alone (d = 1.39, 11 of 11). A diagram the character talks through, a defined term or a study name that appears on screen at the moment it is spoken, or a label placed next to the part it names are all supported moves, governed by the multimedia, temporal contiguity, signaling, and spatial contiguity principles. A visual gets in the way when it competes: a verbatim caption of the dialogue running over a scene that needs the eyes, or decoration added to spice up the screen. The line to hold is that a short reinforcing cue is signaling, a positive effect, while a duplicate stream of the dialogue over a competing picture is redundancy, the one to avoid.
Does a learner have to be taught everything in a separate lesson before the scenario?
No. Mayer's pretraining principle finds that people learn better from a complex lesson when they already know the names and characteristics of its main concepts (d = 0.85, 5 of 5). What the principle requires is a sequence, that the concepts a decision depends on are already known when that decision arrives, not a physically separate lesson beforehand. Mayer and Moreno define pretraining as a sequencing strategy in which components are presented before the system that uses them. A scenario can own that sequence itself, rolling out the terms and the situation earlier in the same scenario, before the decision that turns on them. One caveat: the experiments used a discrete lesson ahead of the main material, so applying the definition to a scenario's internal sequence follows the mechanism Mayer measured rather than a study that tested that exact arrangement.
Published July 19, 2026 · 13 min read