Running Learning-AI Pilots to Real Evidence
It is the readout meeting for a ninety-day learning-AI pilot, and the head of learning has a slide that says the new AI tutor was used by 4,000 people, scored 4.6 out of 5 on the feedback survey, and cut content-build time by 60 percent. The room nods. Then the CFO asks a quiet question: did anyone actually do the job better afterward, and can you show me the evidence in a way my auditor and your accessibility reviewer would both accept. The slide has no answer. The pilot ran for three months, produced a warm feeling and a vendor-friendly number, and proved nothing a leader can act on. This lesson is about the discipline that would have made that meeting go differently, the pilot that produces evidence instead of enthusiasm, and that satisfies accessibility, validity, and impact at the same time rather than one at a time.
The Pilot That Proves Nothing
Most learning-AI pilots are theater. They are designed to generate a favorable anecdote, not a defensible finding, and you can spot one by what it measures: usage, satisfaction, and speed. All three feel like evidence and none of them are evidence of the thing leadership actually needs to know, which is whether the AI-assisted approach makes people do their job better, safely, accessibly, and at a cost worth paying. A pilot that reports 4,000 users and a 4.6 smile-sheet score has measured that people showed up and did not hate it. That is a smile sheet, a Kirkpatrick Level 1 reaction measure, and the oldest trap in learning evaluation: it tells you the room was warm, not that anything changed. Why you care: a warm room is exactly what a vendor demo produces on purpose, and a transformation leader who scales on a warm room is scaling on nothing.
The deeper problem is that a badly designed pilot does not just fail to prove the good case. It actively hides the bad one. If the AI tutor quietly gave a wrong answer to 3 percent of benefits questions, the smile-sheet score would not move, because a confident wrong answer feels exactly as satisfying as a right one to the person receiving it. If the synthetic-video refresh failed WCAG 2.2 AA for screen-reader users, the usage number would look fine, because the learners who could not use it simply do not appear in the completion data as a problem, they appear as an absence nobody counted. A pilot that measures reaction and usage is structurally blind to the two failure modes that matter most at scale: the confident wrong output and the inaccessible experience. It is not neutral. It is a machine for manufacturing false confidence.
The fix is not more measurement for its own sake. It is measurement designed backward from the decision the pilot exists to inform. Before a single learner touches the AI, the leader writes down the decision the pilot will produce, scale it, kill it, or redesign it, and the specific evidence that would justify each choice. A pilot without a pre-committed decision and a pre-committed evidence bar is not a pilot. It is a demo with a longer runway, and it will end the way demos end: with a warm feeling and a vendor logo, and a CFO asking the one question the slide cannot answer.
A pilot that measures usage, satisfaction, and speed has proven that people showed up and did not hate it. It has not proven that anyone did the job better, that the content was correct, or that everyone could use it. Those are the only three questions worth ninety days.
The Three Gates That Must Pass Together
A defensible learning-AI pilot has to clear three gates, and the entire discipline of this lesson is that it must clear all three at once, not sequentially and not selectively. The three are accessibility, validity, and impact. Teams fail because they treat these as separate workstreams owned by separate people at separate times: the accessibility check happens at the end if at all, the validity question is assumed away because the content "looks right," and impact is measured with a smile sheet because behavior measurement is hard. A pilot that passes one gate and quietly skips the other two is not a smaller version of a good pilot. It is a different thing that produces a misleading result, because the three failure modes interact.
Gate One: Accessibility, a Precondition, Not a Polish Step
The accessibility gate asks whether every learner can actually use the AI-generated experience, measured against WCAG 2.2 AA, the accessibility conformance standard that learning content is expected to meet. Why you care: an AI-narrated video with no captions, a chat tutor that fails keyboard navigation, or a synthetic avatar with no transcript is not a course with a minor defect. It is a course that excludes a population of learners and, in many jurisdictions, a legal exposure with the organization's name on it. The gate is a precondition because an inaccessible pilot has already failed for the people it excluded, and no impact number computed over the learners who could use it can undo that. In a defensible pilot, accessibility is verified before launch and re-verified on every AI-generated artifact, because generation is exactly where captions get dropped, contrast fails, and color-only meaning creeps in.
Gate Two: Validity, Is the Content Correct and the Assessment Sound
The validity gate asks two linked questions. First, is the AI-generated content correct, every claim, threshold, and procedure traced to an approved source rather than the model's training data. Second, if the pilot includes any assessment, does the item actually measure the objective, or is it a well-formed question that tests nothing? Assessment validity means an item measures the competency it claims to measure. Why you care: an AI can write a fluent, plausible test item that certifies people as competent when they are not, and a smile sheet will never catch it because the item looks professional. The validity gate is where a pilot proves that the content is grounded and the assessment is sound, with a human sign-off log behind every regulated claim, before anyone counts a completion. Skip it and your impact number is measuring the effect of possibly-wrong content, which is worse than measuring nothing.
Gate Three: Impact, Did Behavior Actually Change
The impact gate asks whether the AI-assisted approach changed what people do, not what they felt. This is Kirkpatrick Level 3, behavior, the level most pilots skip because it is the hardest and the only one leadership truly cares about. A defensible pilot designs the behavior measure before launch: a specific, observable on-the-job action, a way to measure it at baseline and after, and ideally a comparison group that did not get the AI-assisted version so you can attribute the change rather than assume it. Where the case warrants it, the pilot reaches toward Kirkpatrick Level 4, results, and only for the rare program does it compute a Phillips Level 5 ROI. The point of the impact gate is to replace "people liked it" with "people do the job differently, here is the measure, here is the baseline, here is the change." That sentence is the one the CFO asked for, and it is the only one worth building a pilot to produce.
Why the Gates Must Pass at Once, Not in Sequence
The reason all three gates must be designed in from the start, rather than bolted on in order, is that they are not independent. They interact, and a pilot that stages them sequentially discovers the interactions too late to fix them. Consider what happens when a team runs impact first and accessibility last, the most common sequencing. They spend ninety days measuring behavior change on a cohort, produce a beautiful Level 3 result, and then the accessibility review fails the synthetic video. Now the impact number is meaningless, because it was measured only over the learners who could use an experience that excluded others, and the whole pilot has to be redesigned and rerun. The sequence did not save time. It cost a full cycle.
The interactions run in every direction. An accessibility failure invalidates the impact measure, because the population that could not access the content is missing from the behavior data, so the change you measured is not the change the full workforce would see. A validity failure poisons the impact measure, because if the content was subtly wrong, you measured the behavior change caused by wrong content, and scaling that is worse than doing nothing. And an impact measure with no baseline invalidates itself, because "people scored 80 percent after" means nothing without "people scored 55 percent before" and a comparison group who scored 78 percent without the AI. The gates are a single interlocking system, and the discipline is to build the measurement plan for all three before launch, so the pilot produces one coherent finding instead of three findings that contradict each other.
| Gate | The question it answers | The measure | What a skipped gate hides |
|---|---|---|---|
| Accessibility | Can every learner use the AI-generated experience? | WCAG 2.2 AA conformance, verified before launch and per artifact | An excluded population absent from the data, and a legal exposure |
| Validity (content) | Is every claim correct and traced to an approved source? | SME sign-off log, source trace on every regulated claim | A confident wrong output a smile sheet cannot detect |
| Validity (assessment) | Does the item measure the objective? | Item-objective alignment review, item discrimination | People certified competent who cannot do the task |
| Impact | Did behavior actually change, and can you attribute it? | Kirkpatrick L3 behavior measure, baseline plus comparison group | A warm smile-sheet score standing in for a result that never happened |
A Worked Example: The Same Pilot, Run Two Ways
Return to the AI tutor and watch the same ninety-day pilot run first as theater and then as evidence, so the difference is concrete rather than abstract.
Before (the theater pilot). The team stands up the AI tutor, opens it to 4,000 employees, and lets it run for a quarter. At the end they pull three numbers: 4,000 users, a 4.6 satisfaction score, and a 60 percent reduction in the time it took to build the tutor's content versus a traditional module. The readout slide looks triumphant. Then the questions come. The CFO asks whether anyone did the job better, and the team has no behavior measure, only satisfaction. The compliance officer asks who verified the tutor's answers about the leave policy, and the team realizes the tutor generated some answers from training data rather than retrieving from the approved policy, and nobody logged a sign-off. The accessibility reviewer asks for the conformance report on the tutor's interface, and there is none, because accessibility was going to be checked "before the real rollout." The pilot proved that people used a tool and did not complain. It cannot answer a single question that would justify scaling it, and worse, it surfaced two failures, ungrounded answers and unverified accessibility, that a disciplined pilot would have caught in week one instead of at the board readout.
After (the evidence pilot). The same team runs the same tutor, but designs the three gates in before launch. Accessibility: the tutor interface is tested against WCAG 2.2 AA before a single learner touches it, keyboard navigation and screen-reader support are confirmed, and every AI-generated help article carries a transcript, so the pilot is usable by the whole cohort, not a subset. Validity: the tutor is grounded on the approved policy library and cites its source on every answer, a sample of answers is checked against the source each week, and any regulated claim carries a SME sign-off in a log; when a benefits answer is found to be ungrounded in week one, it is caught and fixed, not shipped. Impact: before launch the team defines the behavior they expect to change, the rate at which agents resolve a benefits question correctly without escalating, measures it at baseline for both a pilot group and a matched comparison group, and measures it again at sixty days. The readout is different in kind. It says: the tutor is WCAG 2.2 AA conformant, here is the report; every regulated answer traces to an approved source, here is the sign-off log and the weekly citation-accuracy rate; and the pilot group's correct-resolution rate rose from 71 percent to 86 percent while the comparison group held at 72 percent, a 14-point attributable gain. The CFO does not have a quiet question, because the slide already answered it.
Notice what changed and what did not. The tool was identical. The runtime was identical. The difference was entirely in the design of the pilot: whether the three gates were pre-committed and measured together, or discovered one at a time after the fact. The theater pilot cost a quarter and produced a warm feeling and two hidden failures. The evidence pilot cost the same quarter and produced a finding a leader can scale, a compliance officer can accept, and an accessibility auditor can sign. Same ninety days, opposite value, because one was designed backward from the decision and the other forward from the demo.
The Pilot Design Checklist a Leader Commits Before Launch
A defensible pilot is won or lost in the design phase, before any learner is involved, because the gates cannot be added retroactively without invalidating the result. The following commitments are written down and agreed before launch, and a pilot that cannot answer them is not ready to run.
- The decision and the evidence bar. State the decision the pilot will produce (scale, kill, or redesign) and the specific evidence that justifies each outcome. No pre-committed decision, no pilot.
- The behavior measure and its baseline. Name one observable on-the-job action the pilot should change, measure it at baseline, and where possible run a matched comparison group so the change can be attributed rather than assumed.
- The accessibility gate, before launch. Confirm WCAG 2.2 AA conformance of the AI-generated experience before any learner uses it, and re-verify on every AI-generated artifact, because generation is where accessibility silently breaks.
- The grounding and sign-off trail. Ground the AI on the approved source, require a citation on every load-bearing answer, sample the answers for accuracy on a schedule, and log a human sign-off on every regulated claim.
- The assessment-validity check, if the pilot certifies anything. Confirm every item measures its objective and that a human, not the AI, owns any pass, fail, or completion decision, because AI does not certify a learner as competent.
- The scale, time, and cost boundary. Fix the cohort size, the duration, and the cost so the pilot is bounded and reversible, and so its blast radius stays contained while it is still proving itself.
- The kill criteria. Define in advance what result ends the pilot, an ungrounded-answer rate above a threshold, an accessibility failure, or no attributable behavior change, so the pilot can fail honestly instead of being talked into a scale-up.
The checklist is not bureaucracy. It is the difference between a pilot that answers the CFO and a pilot that gets ambushed by the CFO. A leader who commits these seven items before launch runs a study. A leader who skips them runs a demo and calls it a study, and the gap between the two surfaces at exactly the worst moment, in front of the people whose trust the whole transformation depends on.
Reading the Evidence Honestly, Including When It Says Stop
The final discipline is the hardest, because it runs against every incentive in the room: reading the pilot's evidence honestly, including when it says the case does not work. A pilot exists to produce a decision, and one of the legitimate decisions is "do not scale this." A transformation leader who can only ever conclude "scale it" is not running pilots, they are running a rubber stamp, and the organization learns to distrust every green light they give. The credibility of a scale-up decision comes entirely from the visible willingness to have killed it. When you can point to the pilot you stopped because the behavior measure did not move, or the one you redesigned because the accessibility gate failed, the pilot you did scale carries weight, because the process demonstrably has teeth.
Honest reading also means resisting the two seductions that turn evidence back into theater at the last moment. The first is the vanity metric that reappears when the impact number disappoints: the behavior change was flat, so the slide quietly leads with the 4.6 satisfaction score again, and the pilot is scaled on reaction after all. The second is the attribution shortcut: the pilot group improved, there was no comparison group, and the improvement is credited to the AI when it might have come from the extra attention, the new manager, or the quarter. A leader who reports "the pilot group improved and the comparison group did not, here is the gap" is offering evidence. A leader who reports "the pilot group improved" alone is offering a hope dressed as a finding. The whole point of ninety days of disciplined work is to be able to tell the difference, out loud, in front of the CFO, and to let the evidence decide even when it decides against the idea you were excited to ship.
Key Takeaways
- Most learning-AI pilots are theater: they measure usage, satisfaction, and speed, which prove people showed up and did not complain, not that anyone did the job better, that the content was correct, or that everyone could use it.
- A pilot that measures only reaction is structurally blind to the two failure modes that matter most at scale, the confident wrong output and the inaccessible experience, because a wrong answer and an unusable course both leave the smile-sheet score untouched.
- A defensible pilot clears three gates and must clear them together: accessibility (WCAG 2.2 AA, a precondition), validity (grounded content plus sound assessment), and impact (Kirkpatrick Level 3 behavior change, attributable).
- The gates interact, so sequencing them wastes a cycle: an accessibility failure invalidates the impact number, a validity failure poisons it, and an impact measure with no baseline or comparison group invalidates itself.
- Design the pilot backward from the decision it will produce, not forward from the demo: pre-commit the decision, the evidence bar, the behavior measure and baseline, the accessibility gate, the grounding and sign-off trail, and the kill criteria before any learner touches the AI.
- The evidence-pilot readout replaces "4,000 users and 4.6 satisfaction" with "the pilot group's correct-resolution rate rose from 71 to 86 percent while the comparison group held at 72, and here are the conformance report and the sign-off log."
- Read the evidence honestly, including when it says stop; the credibility of every scale-up decision comes from the visible willingness to have killed the pilot, and a leader who can only ever say "scale it" is running a rubber stamp, not a study.
- AI still does not certify a learner as competent, and the human owns the sign-off and the pass-fail decision; the pilot's job is to produce evidence a compliance officer, an accessibility auditor, and a CFO can all accept, not a warm feeling and a vendor logo.
Skill.re