Building to Mayer's Multimedia Principles with AI
An instructional designer asks an AI tool to "make the cybersecurity module more engaging." It returns a slide where a synthetic narrator reads a paragraph aloud while the exact same paragraph scrolls on screen, a stock animation of a glowing lock pulses in the corner, upbeat music plays under the voice, and a cartoon hacker waves from the margin. It feels lively. It tests terribly. Three weeks later the quiz scores on that module are lower than the boring text version it replaced, and nobody can say why. The why is a fifty-year-old body of cognitive science the AI cheerfully ignored, and learning it is the difference between media that decorates and media that teaches.
Why "Engaging" and "Effective" Are Not the Same Word
The single most expensive misunderstanding in AI-assisted media is the belief that more is better: more visuals, more motion, more narration, more music equals more learning. The science says the opposite, and it has said so for decades. Mayer's Cognitive Theory of Multimedia Learning is a research-backed framework, refined across hundreds of controlled experiments, describing how people actually learn from words and pictures together. Why you care: it is the closest thing the field has to physics for media design, and an AI generator, optimizing for output that looks rich and finished, will violate its principles by default unless you hold the line. The theory rests on three premises a learning professional should be able to recite. People process visual and verbal information through two separate channels. Each channel has a strictly limited capacity at any moment. And real learning is the active work of selecting, organizing, and integrating that information into memory.
That third premise is the one the "make it engaging" instinct forgets. Learning is effortful selection and integration, not passive reception of stimulation. The relevant constraint here is cognitive load, the total amount of mental work a learner's limited working memory is doing at once. When media adds elements that do not serve the objective, a decorative animation, a redundant on-screen duplicate of the narration, a swelling soundtrack, it does not add engagement to a learner with spare capacity. It spends working memory the learner needed for the actual content. The cartoon hacker did not make anyone learn cybersecurity. It ate the working memory they needed to encode the lesson. AI is a load-adding machine by temperament, because rich, busy, "finished-looking" output is exactly what it is rewarded for producing.
An AI generator optimizes for output that looks rich and finished. Mayer's principles optimize for what a limited working memory can actually encode. Those two targets point in opposite directions, and you are the one who decides which one wins.
The Twelve Principles, and the Three AI Loves to Break
Mayer's framework is usually taught as twelve principles, grouped by the job they do. Some reduce extraneous processing (the wasted load decoration causes), some manage essential processing (the load the material itself imposes), and some foster generative processing (the productive effort that builds understanding). The full list is worth knowing, but three of them are where AI-generated media fails most reliably and most expensively. Learn these three cold, because they are the line you will defend on almost every build.
| Principle | What it says | How AI media violates it |
|---|---|---|
| Coherence | Exclude words, pictures, and sounds that do not serve the objective | Adds decorative stock animation, background music, and "engaging" visuals that look finished but spend working memory |
| Signaling | Cue the learner to what matters with highlights and structure | Generates flat, uniform slides where nothing is emphasized, so the learner cannot tell the load-bearing point from the filler |
| Redundancy | Do not narrate on-screen text word for word; use narration plus visuals, not narration plus identical text | Defaults to a narrator reading the exact text already on screen, forcing both channels to process the same words |
| Modality | Present words as narration alongside a visual, not as on-screen text competing with a visual | Dumps dense paragraphs onto a slide and narrates separately, overloading the visual channel |
| Segmenting | Break content into learner-paced chunks | Produces one long, unbroken auto-generated video with no natural stopping points |
Coherence: The Cartoon Hacker Problem
The coherence principle says that people learn better when extraneous words, pictures, and sounds are excluded, not included. The seductive lie of AI media is that adding things improves it. A generator asked to make a module "more engaging" will add background music, motion graphics, and decorative imagery, every one of which is extraneous load if it does not advance the objective. The discipline is subtractive. The right prompt is not "make this more engaging" but "remove everything that does not serve the learning objective, and justify every visual element by the point it teaches." Coherence is the principle that turns "more is better" on its head, and it is the one a learning professional defends most often against a stakeholder who thinks a busier slide is a better slide.
Redundancy: The Narrator Reading the Slide
The redundancy principle says that when a visual is accompanied by narration, adding identical on-screen text harms learning, because the learner's verbal channel tries to process the spoken and the written words at the same time and competes with itself. This is the single most common failure of AI-generated video, because the default behavior of an avatar or narration tool is to read aloud the exact text already on the slide. It feels thorough. It is actively worse than narration alone with a supporting visual. The fix is to separate the channels: narrate the explanation, show a relevant visual, and reserve on-screen text for short labels and signals, not a transcript of the voice. A crucial caveat, which is where redundancy meets accessibility: captions are not a redundancy violation. Captions are an accessibility requirement for learners who cannot hear, presented as a separate, learner-controlled track, not as primary instructional text competing with the narration. The redundancy principle governs your instructional design choices, never your accessibility obligations.
Signaling: The Flat, Uniform Slide
The signaling principle says that cueing the learner to the essential material, through headings, highlights, arrows, and structure, improves learning by guiding attention. AI-generated slides tend to be flat: every bullet weighted the same, no emphasis, no visual hierarchy, because the model has no idea which of the eight points is the one a technician's life depends on. You do. Signaling is where your judgment about what matters becomes a design instruction, and it is invisible work the AI cannot do for you, because it requires knowing the objective and the consequence, not just the content.
The Other Nine Still Matter
The three above are where AI fails most reliably, but the rest of the twelve are still your responsibility, and several have AI-specific traps worth naming. The modality principle says words are better as narration alongside a visual than as on-screen text competing with that visual, and a model asked to "put the content on the slide" will dump a paragraph onto the screen and overload the visual channel; the fix is to move the explanation into narration and leave the screen to the diagram. The segmenting principle says break content into learner-paced chunks, and an auto-generated single-take video has no natural stopping points unless you ask for them. The spatial and temporal contiguity principles say put related words and pictures near each other in space and time, and an AI tool that captions an image two screens away from the visual it describes quietly breaks this. The personalization principle says a conversational tone outperforms a formal one, which is one place AI's default voice can actually help, though it can also slide into chattiness that becomes its own extraneous load. Knowing the full set lets you name precisely which principle a draft violated, instead of saying vaguely that the media "feels off."
One more idea ties the whole framework together: the three kinds of processing the principles manage. Extraneous processing is the wasted load decoration causes, and coherence, redundancy, signaling, spatial and temporal contiguity all exist to reduce it. Essential processing is the unavoidable load the material itself imposes, and segmenting and modality exist to manage it so it does not exceed capacity. Generative processing is the productive effort that actually builds understanding, the work you want the learner spending their capacity on. The whole point of holding the line is to spend a finite working memory on generative processing instead of letting AI's default richness burn it on extraneous load. When you cut a decorative animation, you are not making the module duller; you are returning the working memory the learner needs to think.
Holding the Line: Prompts That Respect the Science
Holding the line is not a matter of fixing AI media after the fact. It is a matter of instructing the generator in the language of the principles from the start, and then verifying the output against them. The shift is from a prompt that asks for richness to a prompt that asks for restraint and alignment. Compare the two.
The load-adding prompt: "Create an engaging, visually rich training video on data handling with animations and music to keep learners interested." This prompt instructs the model to violate coherence, invites redundancy, and asks for engagement as if it were the goal. The output will look impressive in a demo and underperform in a quiz.
The principle-aligned prompt: "Draft a narrated explainer on data handling for the objective: the learner classifies a record at the correct sensitivity level. Narrate the explanation; on screen, show only a labeled decision diagram, no decorative imagery, no background music, no on-screen duplication of the narration. Signal the one classification rule that learners most often get wrong. Segment into three learner-paced parts." This prompt encodes coherence (no decoration, no music), redundancy avoidance (no on-screen duplication), signaling (cue the common error), modality (narration plus diagram), and segmenting. It produces a less flashy and far more teachable draft, and it gives you a checklist to verify the output against.
Then you verify. Watch the draft once as a designer and ask: is there anything on screen or in the audio that does not serve the objective? Is the narration reading the slide word for word? Can I tell, in three seconds, what the most important point on each screen is? Are there natural stopping points? Every yes-to-decoration and yes-to-redundancy is a cut, not a polish. The verification is fast, but it is not optional, because the generator's default is to fail these checks.
Notice that this is the same discipline the whole program teaches, applied to media. AI assists by producing a draft in seconds; the human verifies it against a standard; the human owns whether it ships. With content, the standard is the source of truth and the failure mode is the hallucinated claim. With media, the standard is the cognitive-load science and the failure mode is the extraneous element that looks like a feature. In both cases, the speed is real and the temptation is to accept the polished-looking output unread. In both cases, the value you add is the verification the tool cannot perform on itself, because the tool does not know your objective, your learner, or the consequence of the point being missed. A stakeholder who pushes back with "but it looks great" is describing exactly the surface the principles warn you to distrust.
A Worked Example: Before and After
Return to the cybersecurity module from the opening, and watch the same content built two ways.
Before (engagement as the goal). The designer prompts for an engaging, visually rich module. The AI delivers a synthetic narrator reading each slide's paragraph aloud while the identical paragraph sits on screen (a textbook redundancy violation), a pulsing lock animation and a waving cartoon hacker in the margins (coherence violations), and an upbeat music bed under the narration (more extraneous load). Every slide is a flat wall of equally weighted bullets (a signaling failure), and the whole thing is one unbroken nine-minute video (no segmenting). It looks like real e-learning. The learner's working memory spends itself reconciling the spoken and written words, filtering out the music and the hacker, and hunting for the point with no cue to find it. Comprehension drops. The quiz scores fall below the plain-text version, and the team blames the learners.
After (the principles as the spec). The same content is rebuilt to the principle-aligned prompt. The narrator explains while the screen shows a clean sensitivity-classification decision tree, no music, no mascot, no scrolling transcript. The one rule learners most often miss is highlighted and called out in the narration (signaling). On-screen text is reduced to short labels (redundancy avoided, modality respected). The video is broken into three short, learner-paced segments with a check between them (segmenting). Accurate captions ride as a separate accessibility track, not as primary text. It is visually quieter and it teaches more, and the quiz scores recover and then exceed the original. The content did not change. The cognitive load did, because the principles were the spec instead of an afterthought.
The lesson is not that AI cannot make good media. It is that AI makes media that looks finished and busy by default, and "looks finished and busy" is precisely what Mayer's research warns against. The learning professional's value is being the person in the room who knows that a quieter, more disciplined screen is usually the more effective one, and who can defend that against a stakeholder who wants more sparkle. Speed is the AI's gift. Restraint grounded in the science is yours.
Key Takeaways
- "Engaging" and "effective" are not the same word: more visuals, motion, narration, and music usually spend the working memory a learner needs for the actual content, not add to engagement.
- Mayer's Cognitive Theory of Multimedia Learning rests on three premises: two separate channels (visual and verbal), each with limited capacity, and learning as the active work of selecting, organizing, and integrating, not passive reception.
- AI is a load-adding machine by temperament, because rich, busy, finished-looking output is exactly what it is rewarded for producing, so it violates the principles by default.
- The three principles AI breaks most reliably are coherence (it adds decoration), redundancy (it narrates the on-screen text word for word), and signaling (it produces flat slides with nothing emphasized).
- Captions are not a redundancy violation: they are an accessibility requirement presented as a separate learner-controlled track, and the redundancy principle governs instructional choices, never accessibility obligations.
- Hold the line at the prompt: instruct the generator in the language of the principles, no decoration, no music, no on-screen duplication, signal the key point, segment the content, instead of asking for richness.
- Then verify: watch as a designer and cut everything that does not serve the objective, because the generator's default is to fail the coherence, redundancy, and signaling checks.
- The learning professional's value is restraint grounded in the science: knowing the quieter, more disciplined screen is usually the more effective one, and defending it against the demand for more sparkle.
Skill.re