โ†
AI for Instructors & Learning Professionals
Capable ยท M17 ยท lesson 17 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Recognizing Bad AI Output in Learning Work
๐Ÿ“–
now learning

Recognizing Bad AI Output in Learning Work

15 min

An L&D manager has forty minutes before a build review and an AI-drafted module sitting in front of her: a refreshed data-privacy course, 32 screens, an 8-item quiz, narration script attached. It looks polished. Her instinct is to skim it for tone, nod, and send it on. That instinct is the most expensive habit in AI-assisted learning. Because somewhere in those 32 screens is a fabricated retention period, a quiz question that tests reading instead of judgment, an alt-text field that says "image," and a sentence pitched three grades above her audience. None of those will jump out at a skim. All of them will surface later: in an audit, in a failed accessibility review, in a learner who passed the quiz and still mishandled the data. Reading an AI draft like an auditor, not a reader, is the skill this lesson builds.

Why a Skim Is the Wrong Read

A human writer's errors come with tells. When a person is unsure, their prose gets hedged, vague, or visibly thin; you can feel the author losing confidence, and that feeling is a signal to look closer. An AI model has no such tell. Its confidence is constant. The sentence where it states a true, sourced fact and the sentence where it invents a number out of nothing are written with identical fluency, identical authority, identical polish. This is the single most important thing to understand about reviewing AI output: the usual cues you have spent a career developing to spot a weak passage do not fire, because the model is never visibly weak. It is always confident, including when it is confidently wrong.

That is why a skim, the read that works fine on a colleague's draft, is exactly the wrong read for an AI draft. A skim relies on those missing tells. It catches the passage that feels off. But an AI draft has no passage that feels off; the fabricated retention period feels exactly as solid as the real one. To review AI output you have to switch modes entirely, from reader to auditor. A reader asks "does this flow and make sense?" An auditor asks "what is the claim, where is the evidence, and does this actually do the job it says it does?" The auditor does not trust the prose. The auditor checks. That mode switch is the whole discipline, and it is unnatural, because the draft is engineered to look like it deserves trust.

You cannot proofread your way to catching an AI error. A human's mistakes hedge and wobble; a model's mistakes are written in the same confident hand as its truths. Stop reading like a reader. Read like an auditor.

The Four Failure Families to Hunt

An auditor does not read randomly; an auditor works a checklist, because a checklist catches what attention drifts past. AI learning output fails in four recognizable families, and once you know the families, you know what to hunt for on every draft. They are invented facts, misaligned assessment items, accessibility traps, and reading-level drift. Each has its own signature, its own danger, and its own check. Learn the four and you have converted a vague dread ("is this draft trustworthy?") into four specific, answerable questions you can run in order.

Invented Facts

The first and most dangerous family is the invented fact: a hallucination, the model's tendency to produce fluent, confident content that is simply false. In learning work this shows up as a fabricated policy threshold ("data must be retained for seven years" when the real period is three), an invented procedure step, a made-up statistic dropped into a slide to sound authoritative, or a citation to a standard or regulation that does not exist. Why you care: in a compliance or safety module, a single invented fact ships to thousands of learners with your organization's name on it, and "the AI wrote it" is not a defense to a regulator. The auditor's check for this family is the trace test: for every factual claim, threshold, procedure, or citation, ask "where does this trace to, and have I confirmed it against the approved source?" A claim that traces to nothing is a fabrication until proven otherwise. Treat every number as a number to verify, not a number to repeat.

Misaligned Assessment Items

The second family is the misaligned assessment item: a question that is well-written and grammatical and tests the wrong thing, or nothing. The most common version is a question that tests reading comprehension or recall when the objective demanded judgment or application. The objective says "the learner can decide whether a given data request is permissible," but the AI-drafted item asks "which of the following is the definition of personal data?" The item is clean. It is also invalid, because assessment validity, the degree to which an item actually measures the objective it claims to, has been broken: a learner can pass this item by recognizing a definition and still be unable to make the real decision the job requires. Why you care: an invalid item quietly certifies people as competent when they are not, which is the assessment equivalent of a hallucinated fact. The auditor's check is the alignment test: for every item, name the objective it claims to measure and the cognitive level that objective demands, then ask "could a learner answer this correctly without being able to do the actual task?" If yes, the item is misaligned, however polished it reads.

Accessibility Traps

The third family is the accessibility trap: a draft that silently fails the conformance standard your content has to meet. AI media and content generation is full of these, because the model optimizes for what looks finished, not for what an assistive technology can use. The signatures are specific and learnable: alt-text fields that are empty or say "image" or "decorative" when the image carries meaning, a narration script with no caption track, a video with no transcript, color used as the only way to convey meaning ("click the green button"), reading order that breaks when navigated by keyboard, and contrast that fails for low-vision learners. The standard these violate is WCAG 2.2 AA (Web Content Accessibility Guidelines 2.2, level AA), the conformance target for learning content and the version Section 508 incorporates by reference. Why you care: an AI-generated experience that fails WCAG 2.2 AA does not ship, full stop; accessibility is a legal gate and a moral one, not a polish step. The auditor's check is the assistive-technology test: for every image, ask "does the alt text convey the meaning?"; for every video, ask "are there captions and a transcript?"; for every instruction, ask "does it work without relying on color or a mouse?"

Reading-Level Drift

The fourth family is the quietest: reading-level drift. The model's natural register is articulate and abstract, a notch or three above where most workplace learners read, and unless you constrained it, the draft drifts upward into long sentences, abstract nouns, and unexplained jargon. This is not a cosmetic issue. Content pitched above the learner's reading level teaches nothing, no matter how accurate it is, because the learner cannot process it; the cognitive load of decoding the prose crowds out the learning. The signatures are long compound sentences, nominalizations ("the utilization of the aforementioned methodology"), undefined acronyms, and a tone that reads like a policy memo rather than a teaching tool. Why you care: you can ship a module that is perfectly accurate, perfectly aligned, and perfectly accessible, and still fail to change behavior because nobody on the floor could read it. The auditor's check is the audience test: pick the lowest-reading-level learner in your real audience and ask "would this person finish this screen and understand it?" If you are unsure, run a readability estimate and compare it to your audience's level, treating the score as a number to verify against the real readers, not a target to game.

The Skeptic's Checklist

Put the four families into a single pass you run on every AI draft before it moves forward. This is the artifact to keep beside your review screen.

Failure familySignature to spotThe auditor's check
Invented factsThresholds, procedures, statistics, and citations stated with confidenceTrace test: every claim traces to the approved source, or it is a fabrication until proven
Misaligned itemsA clean question that tests recall when the objective demands judgmentAlignment test: could a learner pass this without being able to do the real task?
Accessibility trapsEmpty or meaningless alt text, missing captions or transcript, color-only meaning, broken order or contrastAssistive-technology test: does it work for a screen reader, a keyboard, and a low-vision learner?
Reading-level driftLong sentences, abstract nouns, undefined jargon, a policy-memo toneAudience test: would the lowest-reading-level learner finish and understand this?

Run the checklist in this order on every draft, and notice what it does to your forty minutes. Instead of one anxious skim that catches nothing specific, you have four targeted passes, each hunting one family of error, each with a clear pass-or-fail question. The checklist is not slower than a skim; it is faster than the rework, the audit finding, and the incident that a skim invites. And it converts the impossible task ("spot whatever is wrong in this confident draft") into four possible ones. That conversion, from dread to checklist, is what lets a human verify AI output at the speed AI produces it.

A Worked Audit: The Data-Privacy Module

Return to the 32-screen data-privacy module and run the skeptic's checklist instead of the skim.

Invented facts. Screen 9 states that "personal data must be retained for no longer than seven years." The manager runs the trace test: she opens the approved data-retention policy and finds the real period is three years for this data class. The "seven years" traces to nothing in her source; the model imported it from the average of other organizations' policies. Caught. Screen 14 cites "GDPR Article 47" as the basis for a rule; she checks, and Article 47 is about binding corporate rules, not the rule on the screen. A misapplied citation, caught by the trace test. Two invented facts that a skim would have waved through because both read as confidently as the true content around them.

Misaligned items. The objective for the assessment section is "the learner can determine whether a specific data-sharing request is permissible." Quiz item 3 asks the learner to select the correct definition of "data controller." The manager runs the alignment test: a learner can recognize that definition and still have no idea whether a real sharing request is permissible. The item tests recall; the objective demands application. Misaligned, caught, flagged for a rewrite to a scenario item that puts the learner in front of an actual request. Item 6 has two defensible correct answers, which is a different validity defect, also caught only because she read for alignment rather than tone.

Accessibility traps. The module has eight images. Five have alt text that reads "image" or is blank; two of those images carry instructional meaning, a workflow diagram and a decision tree, so their blank alt text makes the content invisible to a screen-reader user. The narration script has no caption track attached. One screen instructs the learner to "click the red box to report a breach," with color as the only cue. The assistive-technology test catches all three: the diagrams need real alt text, the narration needs captions and a transcript, and the instruction needs a non-color cue. None of these would have surfaced in a skim, and all of them would have failed a 508 review weeks later, at far higher cost.

Reading-level drift. The audience is a general employee population, target reading level around grade 8. Screen 2 opens: "The categorization of data assets in accordance with the established taxonomic framework necessitates the application of the appropriate retention schedule." The audience test fails on contact: a typical employee will not parse that sentence, and the whole screen reads like a policy memo. The manager flags the module for a plain-language pass: "Sort each type of data using the chart, then apply the matching retention rule." Same meaning, readable by the actual audience. The drift was invisible to a tone-skim because it sounded sophisticated, which is exactly the trap.

Four passes, roughly thirty minutes, and the module that "looked polished" turned out to carry two fabricated facts, two invalid items, three accessibility failures, and a pervasive reading-level problem. Every one of them was real, shippable, and invisible to a skim. The checklist did not make the manager slower. It made her the auditor the draft needed, instead of the reader the draft was designed to fool.

The Auditor's Mindset and the Iron Rule

The deepest shift this lesson asks for is not a technique but a stance. The draft in front of you is not a finished product to approve; it is a set of claims to verify, a set of items to validate, an experience to test, and a register to check. The polish is not evidence of quality. The polish is the thing you have to see past, because the model's fluency is constant whether it is right or wrong, and that constant fluency is precisely what makes the bad output dangerous: it does not look bad. An auditor assumes nothing from appearance and confirms everything from evidence, and that is the only posture that survives contact with a tool that is always confident.

This is the iron rule of the program, applied to the moment of review: AI assists, the human verifies, the human owns the decision, and "the AI wrote it" is never a defense to a compliance officer, an accessibility auditor, or a CFO. Recognizing bad AI output is the verification half of that rule made into a daily practice. The model assisted by producing a draft fast. You verify it by running the four-family checklist, because you, not the model, will answer for the retention period, the quiz item, the alt text, and the reading level when someone asks. The skeptic's checklist is not cynicism about AI. It is the discipline that lets you use AI's speed without inheriting its failures, by reading every draft the way the auditor, the accessibility reviewer, and the learner who has to actually use it eventually will.

Key Takeaways

  • An AI model's confidence is constant: the sentence stating a true fact and the sentence inventing one are written with identical fluency, so the tells you use to spot a weak human passage never fire.
  • A skim is the wrong read for an AI draft because it relies on those missing tells; you must switch from reader ("does this flow?") to auditor ("what is the claim, where is the evidence, does it do the job?").
  • AI learning output fails in four hunt-able families: invented facts, misaligned assessment items, accessibility traps, and reading-level drift, each with a signature and a specific check.
  • Invented facts (hallucinations) are caught by the trace test: every threshold, procedure, statistic, and citation must trace to the approved source, or it is a fabrication until proven; treat every number as a number to verify.
  • Misaligned items are caught by the alignment test: a clean question that tests recall when the objective demands judgment is invalid, because a learner can pass it without being able to do the real task.
  • Accessibility traps are caught by the assistive-technology test against WCAG 2.2 AA: real alt text, captions and transcripts, and meaning that does not depend on color or a mouse; a failing experience does not ship.
  • Reading-level drift is caught by the audience test: content pitched above the learner's reading level teaches nothing however accurate, because decoding the prose crowds out the learning.
  • The skeptic's checklist converts the impossible task of spotting whatever is wrong into four answerable passes, and it is the verification half of the iron rule made into a daily practice: AI assists, the human verifies, the human owns the decision.