โ†
AI for Instructors & Learning Professionals
Aware ยท M4 ยท lesson 4 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI in Content and Media Production
๐Ÿ“–
now learning

AI in Content and Media Production

15 min

It is 4 p.m. on a Thursday, and an e-learning developer has a finished-looking course in front of her. An AI authoring assistant generated the screens from a product brief in nine minutes. An AI video tool turned the script into a polished presenter video with a synthetic avatar in another twelve. An AI voice tool narrated the whole thing in a warm, confident tone for the cost of a coffee. The module looks like 30,000 dollars of agency work and took an afternoon. Then the accessibility reviewer opens it, turns on a screen reader, and within ninety seconds has a list of seven failures that will keep this course out of the LMS. The production was fast. The build was not done. This lesson is about the gap between those two sentences, and the three AI categories that live inside it.

The Three Production Categories, Named Cleanly

When a learning professional says "we used AI to make the content," they are almost always describing three very different machines doing three very different jobs, each with its own strengths and its own way of breaking. Conflate them and you will trust one to do work it never did. Separate them and you can place each on the build, attach the right review, and answer for the output. The three categories are authoring assistants (AI that drafts the course itself), AI video and avatars (AI that produces a synthetic presenter or animated explainer), and AI voice (AI that narrates in a synthetic voice). They are taught here only as an orientation map, never as an endorsement; the named tools change every quarter, the judgment does not.

Before going further, anchor a term that runs through the whole lesson. An accessibility reviewer is the person who checks that a learning experience can be used by someone who is blind, deaf, has low vision, a motor impairment, or a cognitive difference, measured against a published standard. Why you care: a polished, AI-produced course that a reviewer rejects does not ship, no matter how fast it was built, and "the tool generated it that way" is not an answer a reviewer accepts. The standard they hold you to is WCAG 2.2 AA (Web Content Accessibility Guidelines, version 2.2, conformance level AA, a W3C Recommendation since 5 October 2023), and in US federal and federally funded contexts, Section 508, which incorporates WCAG by reference. Speed does not exempt a course from that gate. Nothing does.

A course produced in an afternoon and rejected in ninety seconds did not save you time. It moved the cost from production to rework, and added a failed review to the record.

Authoring Assistants: What They Do and Where They Break

An authoring assistant is AI built into a course-authoring tool that drafts the thing you would otherwise build by hand: screens of text, a knowledge-check quiz, a summary, a set of captions, sometimes a whole module from a single brief or an uploaded document. Examples in the category, for orientation only, include the AI features inside tools like Articulate Rise 360, Adobe Captivate, and iSpring. What they genuinely do well is collapse the blank-page problem. A first draft that used to take a week of writing and screen-building can appear in minutes, structured into a recognizable course shape with headings, body copy, and a quiz already attached.

That speed is real, and it is where the trouble hides. The authoring assistant is doing generation, the AI job of writing something new, and generation's signature failure mode is hallucination: fluent, confident content that is simply false. When the assistant drafts a compliance module and writes that "records must be retained for five years," it may have generated that number from its training data, not retrieved it from your actual policy, which might say seven. The sentence reads perfectly. It is wrong. The model does not flag the difference between a fact it looked up and a fact it invented, because to the model there is no difference; both are just plausible text.

The Four Things an Authoring Assistant Quietly Gets Wrong

Beyond outright hallucination, authoring assistants tend to break in four specific, predictable ways, and naming them turns a vague unease into a checklist. First, the unsourced regulated claim: a threshold, a deadline, a procedure step that traces to nothing you can point at. Second, reading-level drift: the draft creeps to a more complex reading level than your frontline audience can use, because the model writes the way its training text was written, not the way your warehouse staff read. Third, the look-right-measure-nothing quiz item: a knowledge check that is well-formed and tests recall of a sentence on the previous screen rather than the actual job skill the objective demands. Fourth, structural accessibility gaps baked into the generated output: missing heading structure, images dropped in without alt text, color used as the only way to signal meaning. The assistant produced a course; it did not produce a verified, aligned, accessible course, and the difference is the entire job.

AI Video and Avatars: The Synthetic Presenter

The second category turns a script into video without a camera, a studio, or a presenter. AI video and avatar tools generate a synthetic on-screen presenter (an avatar that appears to speak your script) or an animated explainer assembled from text. For orientation, the category includes tools like Synthesia, HeyGen, Colossyan, and Vyond. The appeal is obvious and genuine: a localized, updatable presenter video that once cost thousands of dollars and a shoot day now costs a subscription and an hour, and when the policy changes you edit the script and regenerate instead of rebooking a studio.

Here the breakage is less about a wrong fact and more about three things a reviewer and a legal team both care about. The first is synthetic-media disclosure: the question of whether learners are told the presenter is AI-generated and not a real person. Many organizations, and a growing set of regulations, expect that disclosure, and a synthetic human delivering compliance guidance without it is a trust problem waiting to surface. The second is accessibility of the video itself, which AI generation does not handle for you by default. The third is the uncanny, subtle errors avatars produce: a mispronounced product name, a gesture that does not match the words, a tone that lands wrong for a sensitive topic like a harassment policy or a layoff communication.

Why an AI Video Still Has to Pass WCAG 2.2 AA

A video, AI-generated or not, has to meet specific accessibility requirements before it ships, and AI tools routinely produce video that fails them. It needs accurate captions for learners who are deaf or hard of hearing, and AI auto-captions are a draft, not a finished artifact; they mis-transcribe product names, acronyms, and exactly the regulated terms that matter most. It needs a transcript for screen-reader users and for anyone who learns better by reading. If the video conveys information visually that the narration does not speak aloud, it needs audio description or a redesign so the spoken track carries the meaning. And any on-screen text and graphics have to meet contrast requirements. None of this is automatic. The avatar looks human and the production looks finished, and that polish is precisely what fools a team into skipping the conformance check the reviewer will not skip.

AI Voice: The Narration That Sounds Finished

AI voice tools generate spoken narration in a synthetic voice, including cloned voices that imitate a specific person, from a text script. For orientation, the category includes tools like ElevenLabs and Descript. The draw is the same shape as the others: studio-quality narration in any of dozens of languages, regenerated in seconds when the script changes, without booking a voice actor. For a global compliance course that has to ship in eleven languages, this is a genuine transformation of the production economics.

The risks are specific and easy to underestimate. AI voice mispronounces the words that matter most: your company name, a drug name, a chemical, a legal term, a procedure that has to be heard correctly because someone will act on it. It can place emphasis in a way that subtly changes meaning, stressing a word in a safety instruction so the warning lands as optional. Cloned voices raise a consent question that is not optional: using a real person's voice, even an internal executive's, requires their permission, and a cloned voice deployed without consent is a legal exposure, not a convenience. And like avatars, synthetic narration invites the same disclosure question, because a learner has a reasonable interest in knowing whether the calm, authoritative voice walking them through a safety procedure is a person or a generated track.

The mispronunciation risk deserves a closer look, because it is the one teams most often wave away as cosmetic. A human narrator who hits an unfamiliar chemical name will usually slow down, sound it out, or flag that they are unsure, and a reviewer listening to the recording will catch the stumble. AI voice does the opposite: it pronounces every word, right or wrong, with the same fluent confidence, so a mangled compound name or a wrong stress on a dosage sails by sounding exactly as authoritative as the correct narration around it. There is no audible signal that something is off, which means the error is not just present, it is camouflaged. For audio that a learner will act on, in a pharmacy, on a loading dock, at a chemical cabinet, that camouflaged confidence is precisely the danger, and it is why a subject-matter expert has to listen to the regulated terms in the actual generated track, not just read the script and assume the voice will say it right.

A synthetic voice that mispronounces the one chemical name in a hazmat module is not a typo. It is a hearing learner being told the wrong thing in a confident tone, at scale.

The Category Map: What Each Does, Where It Breaks, Who Catches It

Here is the one-page artifact worth keeping. Read the right-hand column down the page: a human always catches the failure, and the human is rarely the person who clicked generate.

CategoryWhat it actually doesWhere it breaksWho catches it, and against what
Authoring assistantDrafts screens, quizzes, summaries, captions from a brief or documentHallucinated facts, unsourced regulated claims, reading-level drift, invalid quiz items, missing heading and alt-text structureThe SME against the source of truth; the designer against the objective; the reviewer against WCAG 2.2 AA
AI video and avatarGenerates a synthetic presenter or animated explainer from a scriptNo synthetic-media disclosure, inaccurate auto-captions, no transcript, missing audio description, mismatched tone or gestureThe accessibility reviewer against captions, transcript, contrast; legal against disclosure rules
AI voiceNarrates a script in a synthetic or cloned voice, in many languagesMispronounced regulated terms, misplaced emphasis, voice-cloning consent gaps, disclosure gapsThe SME against the script and pronunciation; legal against consent; the reviewer against the transcript

Notice that a single AI authoring platform can do all three jobs in one run: draft the screens, generate the avatar presenter, and narrate it in a cloned voice, then hand you a "finished course." That seamlessness is the trap. The platform blended three jobs with three different failure modes and three different reviewers into one polished export, and your job as the learning professional is to mentally un-blend it back into its parts, because the reviewer, the SME, and the legal team will. Naming the category does not transfer the obligation to the vendor. The course ships under your organization's name, not the tool's.

A Worked Example: Before and After

Return to the Thursday-afternoon course and watch two versions of the same workflow.

Before (the blob of polish). The team uploads a product brief, clicks generate, and a 25-screen module appears with a synthetic presenter and warm AI narration. It looks like agency work, so it goes straight into the LMS for a global rollout to 4,000 employees. The accessibility reviewer was never in the loop because the course "looked done." Two problems surface within a month. The auto-captions transcribed the flagship product name wrong in every one of forty videos, so deaf learners across eleven languages received the brand name incorrectly. And screen 14, a data-handling step, stated a retention period the authoring assistant generated from training data rather than from the actual policy; it was off by two years. The narration was a cloned version of a VP's voice, used without asking her. When a learner using a screen reader files a complaint and the VP hears her own voice in a course she never recorded, the questions start. Who verified the retention period? Who approved the voice clone? Who checked the captions? The room goes quiet, and the quiet is the liability.

After (three categories, three reviews). The same team runs the same tools but treats the export as three jobs. Authoring: the draft is generated, then the SME checks every regulated claim against the source policy and catches the retention period, the designer confirms each quiz item measures its objective, and the reading level is checked against the frontline audience. Video: the synthetic-media disclosure is added, the auto-captions are corrected against the script with special attention to product and regulated terms, a transcript is attached, and the reviewer confirms contrast and that the spoken track carries all the meaning. Voice: the VP signs a consent form before her voice is cloned, the pronunciation of the chemical and product names is verified, and the narration is checked against the corrected transcript. The accessibility reviewer runs a conformance pass before launch, not after. When the same complaint scenario is imagined now, every question has an answer on file: here is the SME sign-off on the retention period, here is the signed voice consent, here is the corrected caption file and the conformance report. Same tools, same speed, completely different fate, because the production was un-blended into its real jobs and each was verified for what it actually was.

The lesson is not that AI production is dangerous. It is that an undifferentiated "we made the content with AI" is dangerous, and a precisely named "here is the authoring draft the SME verified, here is the video the reviewer cleared, here is the voice the person consented to" is defensible. The course did not change between the two versions. The accountability did.

Key Takeaways

  • AI content production is three different machines, not one: authoring assistants that draft the course, AI video and avatar tools that generate a synthetic presenter, and AI voice tools that narrate. Each has a distinct failure mode and a distinct reviewer.
  • Authoring assistants do generation, so their signature risk is hallucination: an unsourced regulated claim, reading-level drift, an invalid quiz item, or accessibility gaps baked into the generated output.
  • AI video looks finished but rarely is: it needs synthetic-media disclosure, accurate corrected captions, a transcript, audio description where visuals carry meaning, and sufficient contrast, none of which the tool guarantees.
  • AI voice mispronounces exactly the regulated terms that matter, can misplace emphasis in a safety instruction, and raises a hard consent question whenever it clones a real person's voice.
  • WCAG 2.2 AA is the conformance gate (Section 508 incorporates it by reference), and speed never exempts an AI-produced course from passing it; a fast course rejected in review cost time, it did not save it.
  • A single platform can blend all three jobs into one polished export; the skill is mentally un-blending it back into authoring, video, and voice so each reaches the right reviewer.
  • Naming a vendor never makes the output correct, and the obligation never transfers to the tool: the course ships under your organization's name, and a human still answers for every claim, caption, and consent.
  • The before/after lesson holds: same tools and same speed produce either a liability shipped at scale or a defensible build, and the only difference is whether the production was un-blended and verified.