"Make it more engaging" is the most common note a course draft receives, and it arrives with a sibling: "make it more interactive." Both words carry approval without carrying a definition. A drag-and-drop claims to be interactive. A quiz claims it. So do a branching click-through and a full simulation. When one word covers all of these, it cannot guide a design decision.
In 2014, Michelene Chi and Ruth Wylie gave the vague words a precise replacement. Their ICAP framework sorts everything a learner visibly does with learning material into four modes, Interactive, Constructive, Active, and Passive, and predicts that learning deepens with each step up the ladder (Chi & Wylie, 2014). ICAP measures cognitive engagement, meaning what the learner's mind is doing with the material as read from overt behavior. The other everyday sense of engaging, holding attention or being enjoyable, is a motivational question the authors deliberately set aside.
The framework also carries an uncomfortable reversal for anyone who has written "interactive" on a feature list. By ICAP's own text, most of what the industry calls interactive is merely active. Addressing so-called interactive computer systems directly, the authors classify a student's response of picking an answer from a menu of choices as active, not interactive, because "the student does not generate a product" (Chi & Wylie, 2014). Clicking and selecting is not the same as generating. The rest of this article walks the ladder rung by rung, so that the next time someone asks for more engagement, you can ask which rung they mean.
What are the four modes of engagement?
Passive: receiving. The learner is oriented toward the material and takes it in, doing nothing else. A financial analyst watches a thirty-minute anti-money-laundering video straight through. ICAP predicts knowledge received this way is stored in isolation, retrievable mainly when the same cue reappears.
Active: manipulating. The learner performs some physical action focused on the material: highlighting sentences, pausing and replaying a video, copying solution steps, dragging items, or selecting from given options. The same analyst now sorts sample transactions into "flag" and "clear" bins. Manipulating focuses attention and helps connect new information to prior knowledge, real progress over receiving. But nothing new has been produced.
Constructive: generating. The learner produces an output that goes beyond what the materials provided: an explanation in their own words, a justification, a prediction, a diagram that was not given. The test is generation. If a learner's "summary" is assembled by copying sentences, the behavior drops back to active. Our analyst writes three sentences explaining why one borderline transaction must be flagged, drawing on reasoning the module never stated outright.
Interactive: dialoguing. The most misused word gets the strictest definition. Chi and Wylie count an exchange as interactive only when it meets two criteria: both partners' contributions must be primarily constructive, and the exchange must involve a sufficient degree of turn-taking. The partner can be a peer, an instructor, or a computer agent, provided it responds in a content-relevant way. The analyst works the borderline case through with a compliance officer who asks for her reasons, challenges one of them, and adds context she lacked, over many turns. Neither party could have produced the resulting understanding alone.
The modes are nested: interactive behavior includes constructive work, constructive includes active manipulation, and active includes passive attention. From that hierarchy comes the ICAP hypothesis: interactive engagement produces more learning than constructive, which beats active, which beats passive.
Interactive
Co-generating through dialogue
Working a decision through with a partner who responds
Constructive
Generating beyond what was given
Explaining why you chose, in your own words
Active
Manipulating what was given
← most “interactive” e-learning sits here
Selecting, sorting, clicking through
Passive
Receiving
Watching, listening, reading
How strong is the evidence for the ladder?
The hypothesis was derived conceptually first, and much of its support comes from reinterpreting existing studies through the ICAP lens rather than from one decisive trial. The authors present it that way themselves.
The strongest single test came from Chi's own lab. A materials-science study assigned learners to all four modes: reading a passage (passive), reading and highlighting (active), interpreting a graph of the same content (constructive), and interpreting it jointly with a peer (interactive). Each mode outperformed the one below it by roughly 8 to 10 percent, in the predicted order (Menekse, Stump, Krause & Chi, 2013, reported in Chi & Wylie, 2014). Two earlier studies whose conditions spanned three modes showed the same ordering. And when a single activity is implemented at two different rungs, the higher implementation wins: notes in one's own words beat copy-and-paste notes, generating the links in a concept map beat selecting them from a list, and explaining solution steps with a partner beat explaining them alone (all assembled in Chi & Wylie, 2014).
For the top rung, converging evidence comes from research on tutoring. Chi, Roy and Hausmann (2008) compared five ways of learning to solve physics problems, and only the three built on dialogue produced statistically significant gains: being tutored one-on-one, observing a tutoring dialogue in pairs, and collaborating with a peer. Pairs who talked through the problems while observing learned to solve them as effectively as the tutees themselves; students who observed alone or studied alone did not gain significantly.
The framework's authors also mark its boundaries. Shallow tests hide the differences: in the four-mode study, easy multiple-choice items showed no separation between conditions, and the ladder emerged only on harder items. Domains governed by arbitrary rules, where there is no deeper rationale to generate, may not reward constructive work. And the intended mode is not always the enacted one: a learner asked to summarize can quietly reduce the task to copying, turning a constructive design into an active performance.
Where do familiar formats actually sit?
Run the diagnostic on the formats a content developer meets every week, grading each by what the learner overtly does:
- A narrated slide course or explainer video: the learner receives. Passive.
- A knowledge check: the learner selects an answer from given options. Active.
- A drag-and-drop exercise: the learner manipulates given items. Active.
- A branching click-through: at each juncture the learner selects from authored options. The branching structure changes what happens next, not what the learner does, which is selecting. Active.
- A discussion prompt answered in the learner's own words: the learner generates. Constructive. A live group discussion, though, is constructive only for the few people talking; Chi and Wylie classify whole-class discussion as passive for the silent majority.
There is the reversal in full. Nearly everything the industry labels interactive lands on the second rung. None of these formats is worthless: active reliably beats passive, and a course that moves learners from watching to manipulating has gained something real. The problem is stopping at the second rung while using the top rung's word, because the label persuades everyone the engagement question is already solved.
What does one rung up look like?
The step from active to constructive is the cheapest upgrade in learning design: after any choice, ask the learner to explain it in their own words. The explanation must be generated, not selected. Chi and Wylie classify choosing a justification from a menu as active, so a picklist of reasons does not clear the bar. A free-text response, a predicted consequence, a one-sentence rationale: each forces the learner to produce ideas the screen never showed.
The step from constructive to interactive requires a responding counterpart, and ICAP's two criteria are the checklist. A counterpart that only confirms or denies is not contributing constructively, and an exchange without real turn-taking is two monologues. The counterpart need not be human; a computer character qualifies when it responds in a content-relevant way. What the counterpart says matters as much as its presence. In studies of observed tutoring dialogues, the strongest results came when each piece of content was preceded by a deep-level-reasoning question: a question about causes, consequences, or mechanisms. Learners who watched those dialogues significantly outperformed every other condition tested, including learners who used the tutoring system directly (Craig, Sullins, Witherspoon & Gholson, 2006). The active ingredient of dialogue is that it draws generation out of the learner, and a well-designed exchange does that on every turn.
Where does a mentored decision dialogue sit on the ladder?
The approach comes first, and it stands on its own: place the learner's decisions inside a conversation. Characters in the scenario respond to what the learner decides, and a mentor character coaches at the decision itself, asking for and supplying reasons in dialogue rather than delivering a verdict afterward. AliveSim was built around this pattern of mentored decision dialogue, with 3D avatar characters carrying the situation, the responses, and the coaching. On ICAP's ladder, the design moves learners well up from watching: the learner commits to a decision at every pivotal moment, and the scenario answers it. The framework's own text keeps the top-rung claim precise, though. A selection by itself counts as active, however rich the response it triggers, so the climb depends on what the dialogue draws out of the learner. Coaching that has the learner weigh reasons before committing is working in constructive territory, and an exchange earns ICAP's interactive label only where both of the framework's criteria are genuinely met, with the learner contributing constructively and real turn-taking on both sides.
For where the learner should sit in that conversation, see our guide to the hot seat versus the advisor seat; for how the coaching works at each choice, see decision-point mentoring; for the wider divide between consuming content and acting in situations, see immersive learning vs. e-learning.
The takeaway
"Make it more engaging" is only an empty instruction if engagement stays undefined. ICAP defines it, and the definition fits in one question learning leaders can ask in any design review: what does the learner generate, and who responds to it? If the answer is "nothing, and no one," the course is passive or active, whatever its feature list says. One rung up is usually within reach of the design you already have: a choice becomes constructive when the learner must explain it, and it starts up toward interactive when something in the experience answers back constructively, turn after turn. The word "interactive" has a real meaning. It is worth reserving for designs that earn it.
References
- Chi, M. T. H., Roy, M., & Hausmann, R. G. M. (2008). Observing tutorial dialogues collaboratively: Insights about human tutoring effectiveness from vicarious learning. Cognitive Science, 32(2), 301–341.
- Chi, M. T. H., & Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4), 219–243.
- Craig, S. D., Sullins, J., Witherspoon, A., & Gholson, B. (2006). The deep-level-reasoning-question effect: The role of dialogue and deep-level-reasoning questions during vicarious learning. Cognition and Instruction, 24(4), 565–591.
Related questions
What is the ICAP framework?
The ICAP framework, published by Michelene Chi and Ruth Wylie in 2014, classifies a learner's engagement with material into four modes based on visible behavior: Interactive (dialoguing), Constructive (generating), Active (manipulating), and Passive (receiving). The ICAP hypothesis predicts that learning deepens as engagement moves up from passive to active to constructive to interactive, and the authors support the prediction with laboratory and classroom studies, including one that compared all four modes directly (Chi & Wylie, 2014). Its usefulness for instructional designers is diagnostic: any activity can be graded by what the learner overtly does with it.
What is the difference between active and interactive learning?
In ICAP's terms, active engagement means physically manipulating the material: highlighting, dragging, pausing and replaying, or selecting an answer from given options. The learner focuses attention but adds nothing new. Interactive engagement is defined much more strictly, as dialogue that meets two criteria: both partners' contributions must be primarily constructive, and there must be a sufficient degree of turn-taking (Chi & Wylie, 2014). Between the two sits the constructive mode, in which a learner working alone generates something beyond what was given, such as an explanation in their own words. Much e-learning that is described as interactive is, by these definitions, active.
Is a quiz interactive?
Not by ICAP's definition. The framework addresses this case directly: when a student's response to a computer-based system consists of selecting an answer from a menu of choices, the selection is classified as active, because the student does not generate a product (Chi & Wylie, 2014). That does not make quizzes worthless. Active engagement reliably beats passive receiving, and a well-placed question focuses attention. But a quiz reaches the constructive mode only when the learner must produce reasoning in their own words, and it reaches the interactive mode only when a counterpart responds constructively to what the learner said and the exchange continues over multiple turns.
How do you make e-learning more engaging without gamification?
Change what the learner does with the content rather than adding points, badges, or a leaderboard around the same activity. The ICAP framework points to two moves. First, make the learner generate: have them explain a choice in their own words, predict an outcome before it is revealed, or summarize a case without copying, all of which reach the constructive mode. Second, give the learner's decisions a responding counterpart: a colleague, a coach, or a computer character that replies in a content-relevant way and keeps the exchange going. Evidence gathered by Chi and Wylie (2014) shows the same activity produces more learning when it is implemented one mode higher.
Published July 17, 2026 · 9 min read