โ†
AI for Instructors & Learning Professionals
Proficient ยท M15 ยท lesson 15 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Mapping the Design-to-Evaluation Process
๐Ÿ“–
now learning

Mapping the Design-to-Evaluation Process

15 min

It is a Monday standup, and a head of learning drops a number on the table: the new-hire compliance curriculum, eleven modules, has to be rebuilt and live in five weeks because a policy changed. The room does the old math in their heads, six weeks per module, and someone exhales. Then the AI-integrated designer in the corner says something different. She says, "I can have grounded drafts of all eleven by Thursday, but I need the SME calendar locked for the three regulated modules and an accessibility reviewer booked for week four, because those are the gates, and the gates are the schedule." She is not promising speed. She is promising a map: which steps AI can run at full pace, and which steps a human has to stop and own. That map is the difference between rebuilding eleven modules in five weeks and shipping eleven liabilities in three.

Why the Workflow Is the Real Deliverable

At Level 1 you learned to read an AI learning claim skeptically. At Level 2 you learned to draft a single grounded objective, a single verified module, a single validated item. Level 3 is where those isolated skills become a system. The unit of work is no longer the prompt or the page. It is the workflow: the full path a piece of learning travels from a performance gap to an evaluation plan, with every step labeled by who runs it and who owns it. A workflow is just the design lifecycle made explicit, drawn so that anyone can see where speed is safe and where judgment must stay.

Here is the trap that makes this lesson necessary. AI made every individual step faster, so it is tempting to believe the whole pipeline got faster by the same multiple. It did not. A pipeline moves at the speed of its slowest mandatory gate, and AI does not touch the gates. A gate is a step where a named human has to verify, validate, or approve before the work can continue, and where "the AI did it" is not an acceptable answer to the person who later asks "who checked this." The drafting got ten times faster. The SME verification of a regulated claim did not, because it cannot: verification is the thing AI cannot do for itself. So if you map the workflow honestly, you discover that the build time is now dominated not by the typing but by the checking, and the checking is exactly where you must spend your scarce human attention. The map is what tells you where to spend it.

An instructional designer who cannot draw this map will do one of two bad things. She will either treat the whole pipeline as AI-safe and ship unverified content at speed, which is the liability we have warned about since Level 1, or she will treat the whole pipeline as human-only and throw away the speed AI legitimately offers, which gets her replaced by someone who did not. The map is the third path: full speed where it is safe, full stop where it is not, and a clear line between the two that a compliance officer can read.

A pipeline moves at the speed of its slowest mandatory gate. AI accelerates the steps between the gates and never the gates themselves, so the workflow map is really a map of where your human attention has to go.

The Three Kinds of Step

Map any learning lifecycle and every step falls into one of three buckets. Naming the buckets is the whole discipline, because once a step is named, the right way to run it is obvious.

AI-Ready Steps: Where Speed Is Safe

An AI-ready step is one where AI can run at full pace and a mistake is cheap, visible, and reversible before it reaches a learner. Drafting a storyboard from a source you control, generating fifteen candidate distractors so you can cull the weak ones, summarizing a forty-page SOP into a study guide outline, producing first-pass alt text for a diagram, drafting a launch email, reformatting prose into the content blocks your authoring tool accepts. These are not low-value steps. They are the steps that used to eat days, and AI returns those days to you. The reason they are safe is not that AI is accurate here; it is that the output is a draft that a human will see and shape before anyone is taught from it. Speed is safe when the next step is a human looking at the result.

Human-in-the-Loop Steps: Where AI Drafts and a Human Shapes

A human-in-the-loop step is one where AI produces a candidate and a human's judgment is required to turn it into something usable, but the stakes do not yet demand a formal sign-off. Aligning a drafted objective to the real performance gap. Deciding whether a generated scenario is realistic enough to teach from. Choosing which three of fifteen distractors are genuinely plausible. Judging whether the reading level fits the audience. Here the human is not rubber-stamping; she is doing the actual instructional-design work, with AI handling the typing and the volume. The failure mode is the seductive draft: an output so fluent and well-formatted that the human stops shaping and starts accepting. The discipline is to treat every AI output at these steps as a proposal from a fast, confident, and occasionally wrong junior who has never met your learners.

Human-Only Gates: Where Judgment Must Stay

A human-only gate is a step where a named person must verify and approve, the approval must be recorded, and the work cannot proceed without it. Verifying a regulated, safety, or policy claim against an approved source of truth. Validating that an assessment item actually measures its objective before it certifies anyone. Confirming an experience meets WCAG 2.2 AA, the Web Content Accessibility Guidelines at the AA conformance level, which is the accessibility standard learning content is held to. Bias-checking a scenario about people before it ships. Owning the pass or fail decision on a credential. These are the steps where "the AI did it" is not a defense to a compliance officer, an accessibility auditor, or a CFO, and so they are the steps a human must own with her name attached. AI may assist at a gate, drafting the item or proposing the alt text, but it can never be the gate. The gate is a human, full stop.

The art of workflow design is sorting every step into the right bucket and then resisting the two temptations: do not demote a gate to an AI-ready step because you are in a hurry, and do not promote an AI-ready step to a gate because you are nervous. A misplaced gate is either a liability or a bottleneck, and a good map has exactly the gates it needs and not one more.

The Lifecycle Map, Step by Step

Here is the artifact worth printing and pinning to the wall: the full design-to-evaluation lifecycle, every step labeled with its bucket, the human who owns it, and the failure you are preventing. This is the map the designer in the opening scene was reading from. Read the "bucket" column first, then notice how the gates cluster around regulated claims, assessment validity, and accessibility, exactly the three places where being wrong is expensive.

Lifecycle stepBucketWho owns itFailure prevented
Needs and task analysisHuman-in-the-loopDesignerBuilding a course for a gap that does not exist
Cluster and summarize source inputsAI-readyDesigner (reviews)Drowning in raw interview and document data
Write measurable objectivesHuman-in-the-loopDesignerObjectives at the wrong Bloom's level for the job
Draft storyboard and module from a source of truthAI-readyDesigner (reviews)The slow blank-page first draft
Verify every regulated or safety claim against the sourceHuman-only gateSMEA hallucinated policy threshold or wrong procedure shipping at scale
Draft assessment items and distractorsAI-readyDesigner (reviews)Weeks spent writing and rewriting questions
Validate each item measures its objectiveHuman-only gateDesigner or assessment leadAn invalid item certifying people who cannot do the task
Draft media: script, narration, visuals, captions, alt textAI-readyDesigner (reviews)The slow, expensive media build
Confirm WCAG 2.2 AA conformanceHuman-only gateAccessibility reviewerAn AI-narrated experience that silently fails a 508 audit
Bias-check any scenario about peopleHuman-only gateDesigner plus a second reviewerA stereotype baked into a DEI or hiring module
Record SME and reviewer sign-offHuman-only gateDesigner (logs)The unanswerable "who verified this?"
Tag content and package for the LMSAI-readyLearning technologist (confirms)Broken reporting and mis-routed content
Design the Kirkpatrick measurement planHuman-in-the-loopDesigner plus stakeholderShipping with no way to prove it worked
Own the pass or fail credential decisionHuman-only gateDesigner or program ownerAI certifying competence it cannot judge

Look at the shape of it. Roughly half the steps are AI-ready, and those are where your five-week timeline gets its speed. But every regulated claim, every assessment-validity question, every accessibility check, and the final credential decision sit behind a human-only gate. The map does not slow you down; it tells you precisely where you are allowed to go fast, which is most places, and where you must stop, which is exactly the places an auditor will inspect. Kirkpatrick, by the way, is the four-level model of training evaluation, Reaction, Learning, Behavior, and Results, and it appears on the map because in a Level 3 workflow the measurement plan is designed in from the start, not bolted on after launch.

Where the Gates Cluster and Why

Notice that the human-only gates are not scattered randomly. They cluster at three predictable places, and understanding why lets you find the gates in any new workflow without memorizing a list.

The first cluster is the regulated claim. Anywhere a piece of content asserts a fact that a learner will act on and a regulator could audit, a policy threshold, a safety step, a legal requirement, a clinical dosage, a code-of-conduct rule, you have a gate, and a SME owns it. The reason is that the cost of being wrong is not a typo; it is an incident, an audit finding, or a lawsuit, shipped to thousands at once. AI drafting these claims from a grounded source is fine and fast. AI being the final word on whether the claim is correct is the single most dangerous move in learning, because a generation model produces a confident wrong threshold with exactly the same fluency as a correct one. The gate exists because fluency is not truth.

The second cluster is assessment validity. Anywhere an output decides, or contributes to deciding, whether a learner is competent, you have a gate. A well-written question can measure nothing at all, testing reading comprehension or trivia instead of the skill, and an AI item generator produces beautifully formatted invalid items at scale. So every item passes a human-only validity check before it can certify anyone, and a human owns the final pass or fail. The bright line of the whole program lives here: AI does not certify a learner as competent. It may draft the item and the feedback; a human validates the assessment and owns the decision, full stop.

The third cluster is the experience itself: accessibility and bias. An AI-generated video can omit captions, produce a transcript that drifts from the narration, choose a contrast ratio that fails, or auto-generate alt text that describes the wrong thing, and any of these fails WCAG 2.2 AA. A generated scenario about people can encode a stereotype that turns a harassment-prevention module into the very liability it was meant to prevent. These are gates because an accessibility audit can reject the course and a biased scenario can become a legal exposure, and neither risk is one AI can be trusted to clear on its own. An AI-generated experience that fails WCAG 2.2 AA does not ship, full stop. An AI-generated scenario about people is bias-checked before it ships, full stop.

Find the gates by finding the expense. Wherever being wrong costs an incident, an invalid credential, a failed audit, or a discrimination claim, a named human owns the step and AI only assists.

A Worked Example: Before and After

Return to the eleven-module compliance rebuild and watch two versions of the same five weeks.

Before (the unmapped pipeline). The team treats "AI rebuilds the course" as one fast step. They feed the changed policy and the old modules into an authoring assistant, generate eleven new modules with quizzes by Wednesday, lightly proofread for tone, and push everything live in week two to celebrate beating the deadline. Three of the eleven modules cover regulated procedures. In module seven, the AI updated a threshold from the new policy correctly but, in the surrounding narrative it generated to "explain" the change, introduced a second threshold that was never in any policy, a plausible-sounding number the model invented. Nobody verified the narrative against the source because the workflow had no verification step; it had only a proofread. Two modules also shipped AI-narrated video with auto-captions that the team never checked against WCAG 2.2 AA, and one of them fails contrast on the on-screen text. Six weeks later a compliance officer pulls module seven for a routine review, finds the invented threshold sitting in a record that 3,000 employees completed, and asks the question that has no good answer: "Who verified this before it went live?" The deadline was beaten. The audit was not.

After (the mapped workflow). The same team draws the map first. Eight of the eleven modules are general-skills content with no regulated claims; those run almost entirely through AI-ready steps and human-in-the-loop shaping, and grounded drafts are ready by Thursday, exactly as promised. The three regulated modules follow the same fast drafting path but then hit the gates. Each regulated claim in those modules is traced back to a line in the new policy, and the SME for that domain verifies it; the invented second threshold in module seven is caught here, at the gate, because the designer cannot find a source for it and the SME confirms it does not exist. Every assessment item across all eleven modules passes a validity check against its objective before the item bank is locked. The two AI-narrated videos go to an accessibility reviewer who catches the contrast failure and the caption drift in week four, while there is still time to fix them. Every gate clearance is logged with a name and a date. The curriculum ships in five weeks, the same speed the unmapped team thought they had, but now when the compliance officer pulls module seven, the lead opens the sign-off log and answers in one sentence: "Every claim traces to the new policy, here is the SME who verified module seven on the fourteenth, and here is the accessibility conformance report." Same eleven modules, same five weeks, opposite outcome at the audit. The map did not cost time. It spent the time where it mattered.

The lesson is not that the mapped team was slower. It was not; both teams hit the deadline. The difference is that the unmapped team spent its speed everywhere, including the three places where speed was a liability, and the mapped team spent its speed on the eight safe modules and its human attention on the three dangerous ones. Same budget of hours, radically different distribution, and the distribution is the whole skill.

Drawing Your Own Map

You will not always be rebuilding compliance training. The point of the map is that it generalizes: any learning build, from a microlearning drip to a leadership program, can be drawn the same way. Here is the procedure.

First, list the steps of your actual lifecycle, end to end, from the performance gap to the evaluation plan. Do not start from a tool's feature list; start from the work. Second, label each step with one of the three buckets by asking a single question of each step: if AI got this wrong and nobody caught it, what is the cost, and is it cheap, visible, and reversible before a learner sees it, or expensive, invisible, and live? Cheap and reversible means AI-ready. Requires judgment to be usable means human-in-the-loop. Expensive, auditable, or credential-deciding means human-only gate. Third, for every gate, name the human who owns it and decide how the approval gets recorded, because a gate with no named owner and no record is not a gate; it is a hope. Fourth, sanity-check the gate count: too many gates and you have thrown away AI's speed and created a bottleneck the business will route around; too few and you have left a liability unguarded. The right number is the smallest set that covers every regulated claim, every validity decision, every accessibility and bias check, and the final credential call.

The map is a living artifact, not a one-time diagram. When a tool improves, a step may move from human-in-the-loop toward AI-ready, and you let it, because the bucket is defined by the cost of error, not by loyalty to old habits. But a gate almost never moves, because the cost of a wrong regulated claim, an invalid credential, or a failed audit does not fall just because the model got better at sounding confident. Confidence was never the problem. Verified truth, valid measurement, and accessible delivery were always the job, and they stay human-owned no matter how good the drafting gets.

Key Takeaways

  • The Level 3 unit of work is the workflow, not the prompt: the full path from performance gap to evaluation plan, with every step labeled by who runs it and who owns it.
  • A pipeline moves at the speed of its slowest mandatory gate, and AI accelerates the steps between gates but never the gates themselves, so the build time is now dominated by verification, not typing.
  • Every step falls into one of three buckets: AI-ready (speed is safe because a mistake is cheap and reversible), human-in-the-loop (AI drafts, a human shapes), or human-only gate (a named person verifies, approves, and records before work proceeds).
  • Human-only gates cluster at three predictable places: the regulated claim (owned by a SME), assessment validity (owned by the designer or assessment lead), and the experience itself, meaning accessibility and bias (owned by an accessibility and bias reviewer).
  • Find the gates by finding the expense: wherever being wrong costs an incident, an invalid credential, a failed audit, or a discrimination claim, a human owns the step and AI only assists.
  • The map does not slow you down; it tells you precisely where you are allowed to go fast (most places) and where you must stop (exactly the places an auditor will inspect), so you spend the same hours but distribute them correctly.
  • A gate with no named owner and no recorded approval is not a gate; it is a hope, and "the AI did it" is not an answer to "who verified this?"
  • The map is living: a step can move toward AI-ready as tools improve, but a gate almost never moves, because the cost of a wrong claim, an invalid credential, or a failed audit does not fall just because the model sounds more confident.