Assessing Your Function's AI Readiness
The CEO read a headline on a flight and forwarded it to the head of learning with one line: "If AI makes training ten times faster, rebuild the whole catalog by Q3." It is now Tuesday. The head of learning is looking at a function of nine people, a content library nobody has audited in three years, an LMS that reports completions and almost nothing else, and a verification habit that exists in two designers' heads and nowhere on paper. She has roughly forty minutes before a leadership call where she will be asked for a date. The honest answer is not a date. It is a readiness assessment, and the gap between those two answers is the difference between a transformation that ships and a promise that detonates in an audit eighteen months from now.
Why Readiness Comes Before the Roadmap
Every failed learning-AI program this decade shares an origin story: a leader promised an outcome before anyone measured the function's capacity to deliver it safely. The promise was speed. The unmeasured thing was whether the function could verify what AI produced fast enough to keep up with the speed it unlocked. That is the trap. AI compresses production time dramatically, but it does not compress verification time at all, and a function that cannot verify at the new speed does not get faster. It gets more dangerous, because now wrong content ships at the same velocity as right content, to the same thousands of employees, into the same compliance record an auditor can pull.
A readiness assessment is an honest, evidence-based baseline of what your learning function can actually support before you commit to a transformation. Why you care: it converts a leadership conversation from "when will it be done" into "here is where we are strong, here is where we are exposed, and here is the sequence that captures the upside without inheriting the disaster." It is the artifact that lets a head of learning say no to a reckless date and yes to a defensible plan in the same breath. Without it, you are negotiating a deadline blind, and the person who negotiates a deadline blind owns every consequence of guessing wrong.
The strategist's job at Level 4 is not to be the most enthusiastic person in the room about AI. It is to be the most accurate person in the room about the function. Enthusiasm is cheap and abundant in 2026; the Josh Bersin Company frames AI as disrupting a corporate-learning market it sizes at roughly 400 billion dollars, and every vendor on the conference floor will sell you the upside. Accuracy about your own readiness is scarce, and it is the thing that turns the upside into a result instead of a liability. The readiness assessment is how you manufacture that accuracy on purpose.
AI collapses production time and leaves verification time untouched. A function that cannot verify at the new speed does not get faster. It ships its mistakes faster.
The Four Dimensions of Readiness
A learning function is not ready or unready as a single fact. It is ready along four separate axes, each of which can be strong or weak independently, and each of which gates a different category of AI use. Assess them separately and you get a map of exactly where to start. Smear them into a single "are we ready" feeling and you get a guess. The four dimensions are content maturity, data, verification culture, and accessibility posture. The first principle of the assessment is that the weakest dimension, not the strongest, determines what you can safely promise.
Content Maturity: Is There a Source of Truth to Ground On?
The single most important precondition for safe learning AI is a verified source of truth the model can be grounded on. Grounded generation, often called RAG (retrieval-augmented generation), means forcing the model to draft from your approved policies, SOPs, and SME-verified material rather than its own training data. Why you care: an ungrounded model invents plausible facts, and a plausible invented fact in a compliance module is the failure mode that ends careers. Grounding is only possible if there is something approved and current to ground on.
So content maturity asks blunt questions. Do you have a canonical, current source of truth for your regulated and safety content, or does the "real" policy live in a SME's inbox and a three-year-old PDF that disagree with each other? Is your content tagged, structured, and findable, or is it a sprawl of orphaned SCORM packages nobody can map to a competency? When a regulation changed last quarter, did the affected modules get updated, and can you prove which ones? A function with a clean, current, governed content base can ground AI on day one. A function whose content is stale and scattered has to fix the source of truth first, because grounding AI on a wrong source produces wrong content faster, which is worse than slow.
Data: Can You Measure Anything Beyond Completions?
The data dimension asks what your learning data can actually prove. Most functions can report that a course was completed and that learners rated it well on a smile sheet, the satisfaction survey at the end of a course, which is Kirkpatrick Level 1 (Reaction) and proves almost nothing about whether anyone learned or changed. Why you care: if you cannot measure behavior, you cannot prove AI-assisted learning worked, and a transformation you cannot measure is a transformation you cannot defend to a CFO who already suspects training is a cost center.
Maturity here is a ladder. Can you measure learning (Level 2, did knowledge change), behavior (Level 3, did the job change), and ideally results (Level 4, did the business outcome move)? Do you emit xAPI statements, the standard records that capture learner behavior beyond the LMS, or are you blind the moment a learner leaves a course page? Is your learner data clean enough to trust, and have you decided who is allowed to use it, including whether a vendor's tool can train its model on it? A function with real behavioral data can prove AI's impact and detect when an AI-built course quietly underperforms. A function with only completions and smile sheets is flying on vanity metrics, and vanity metrics cannot survive a real business case.
Verification Culture: Does the Habit Exist on Paper?
This is the dimension leaders most often overrate, because it feels like something the team already does. Verification culture asks whether checking AI output against a source is a logged, repeatable, named-owner practice, or an informal habit living in a couple of careful people's heads. Why you care: the iron rule of this entire program is that AI assists, the human verifies, the human owns the decision, and "the AI wrote it" is never a defense to a compliance officer, an accessibility auditor, or a CFO. That rule only protects you if verification is an enforced control, not a personality trait.
Concretely: when a designer accepts an AI-drafted compliance claim, is there a SME sign-off log, a tamper-evident record of who verified which claim against which source and when, or does the verification disappear the moment the screen looks finished? Do you have a written checklist for catching invented facts, misaligned items, and accessibility failures before sign-off? If your two best designers left tomorrow, would the verification discipline leave with them? A mature verification culture means the control survives the people. An immature one means you are one resignation away from shipping unverified AI content at scale, and you would not know until the audit.
Accessibility Posture: Is the Gate Already in Place?
The fourth dimension asks whether accessibility is a gate or an afterthought. WCAG 2.2 AA (the Web Content Accessibility Guidelines, version 2.2, conformance level AA, a W3C Recommendation since 5 October 2023) is the conformance target for learning content, and Section 508 incorporates WCAG by reference. Why you care: an AI-generated experience that fails accessibility does not ship, full stop, and AI-generated media fails in specific, repeatable ways, missing captions, bad alt text, poor contrast, that an audit hunts for.
Maturity here is whether you have a real conformance process before launch, ideally producing a VPAT or ACR (a Voluntary Product Accessibility Template, or the Accessibility Conformance Report it generates, the document that states how a product meets the standard), or whether accessibility is a rushed retrofit after legal complains. A function that already gates on accessibility can let AI accelerate production without multiplying accessibility debt. A function that treats accessibility as polish will find that AI multiplies the debt at the same rate it multiplies the output, and the auditor arrives at the same time as the lawsuit.
Scoring the Baseline: From Feeling to Evidence
An assessment that produces a feeling is worthless to a leadership team. An assessment that produces a defensible score per dimension, with evidence behind each score, is the artifact that earns the function the right to set its own timeline. Score each dimension on a simple, honest scale and force yourself to attach evidence, not optimism, to every number.
| Dimension | Level 1: Exposed | Level 2: Developing | Level 3: Ready | The evidence that proves the score |
|---|---|---|---|---|
| Content maturity | Source of truth scattered or stale; content untagged | Key regulated content governed; rest is mixed | Canonical, current, tagged source of truth AI can ground on | Point to the governed source for a named compliance topic and its last review date |
| Data | Completions and smile sheets only | Some Level 2 learning data; xAPI in places | Behavior (Level 3) measurable; clean data; usage rights decided | Show one program with a real Level 3 behavior measure |
| Verification culture | Informal, in people's heads, unlogged | Checklist exists; sign-off inconsistent | Logged SME sign-off; control survives staff turnover | Produce the sign-off log for a recently shipped regulated module |
| Accessibility posture | Retrofit after complaints | Checked late, sometimes skipped under deadline | Gated before launch; VPAT/ACR produced | Produce the conformance report for the last AI-touched course |
The discipline that makes this honest is the evidence column. Anyone can claim Level 3 verification culture; only a function that actually has it can produce the sign-off log for a real, recently shipped module on demand. If you cannot produce the evidence, the score is not Level 3, no matter how the team feels about its own carefulness. This is the same standard an auditor will apply, which is precisely why applying it to yourself first is the strategist's most valuable move. You want to find the gap in a planning meeting, not in a regulator's conference room.
If you cannot produce the artifact on demand, you do not have the capability. The auditor scores on evidence, so score yourself on evidence first.
A Worked Baseline: Before and After
Return to the head of learning with forty minutes before the leadership call, and watch two versions of how she handles the CEO's "rebuild the catalog by Q3."
Before (the promise). She wants to look capable, so she says yes to Q3 and commits the team to "AI-accelerated catalog rebuild." Production does get faster; the team drafts modules in days that used to take weeks. But the content base was never governed, so AI grounds on stale and conflicting sources and several modules inherit a wrong policy threshold nobody catches, because verification was informal and the team is now moving too fast for two careful people to check everything. Accessibility was a retrofit, so the AI-narrated videos ship without proper captions. Eighteen months later a compliance audit pulls the safety refresh, finds a fabricated threshold that reached 4,000 employees, and asks the question that has no good answer: "who verified this before it shipped?" The catalog was rebuilt on schedule. It was also indefensible, and the speed is now the prosecution's exhibit.
After (the baseline). She spends the forty minutes building the readiness assessment instead of guessing a date. She scores content maturity at Level 1 for the safety library (stale, scattered) but Level 3 for the recently governed onboarding content. Data is Level 1 across the board, completions only. Verification culture is Level 2, a checklist exists but sign-off is inconsistent. Accessibility is Level 2, gated but skipped under deadline. On the call she does not refuse the CEO; she reframes. "We can move at AI speed where we are ready, and onboarding is ready now. The safety catalog is our highest risk and our weakest content base, so AI there would ship faster mistakes into a compliance record. Here is the sequence: AI-accelerate onboarding this quarter to prove the model and the savings, fix the safety source of truth and the sign-off log in parallel, and bring AI to the regulated catalog once verification can keep pace. You get a visible win in Q3 and a defensible catalog by year-end, instead of a fast catalog you cannot defend." The CEO gets a date for the win and a reason for the sequence. Same enthusiasm, completely different fate, because the baseline replaced the guess.
The difference between the two heads of learning is not talent or courage. It is that one had a readiness assessment and one had a feeling. The assessment is what let her say no to the dangerous part of the request and yes to the valuable part, with evidence behind both. That is the whole function of the artifact: it gives a leader the standing to sequence the transformation by what the function can actually support, rather than by what a headline on a flight made a CEO want.
Turning the Baseline Into a Leadership Conversation
A readiness assessment that lives in a spreadsheet changes nothing. Its value is realized only when it reframes the leadership conversation from a date to a sequence, and that reframing is a skill in itself. The move is never "we are not ready," which sounds like an excuse and invites someone to overrule you. The move is "here is exactly where we are ready, here is where we are exposed, and here is the order that captures the speed leadership wants without inheriting the risk the function would own."
This is why the assessment is scored per dimension rather than as a single grade. A single grade forces a binary, go or no-go, that a leadership team will always resolve toward go. Four dimensions with evidence let you say yes and no at the same time, surgically: yes to AI in the well-governed, low-stakes content where you are ready, no to AI in the stale, high-stakes safety catalog until the source of truth and the verification log are fixed. That precision is what makes a head of learning sound like a strategist instead of a brake. The CEO does not want to hear that the function is slow. The CEO wants to hear that the function knows exactly what it is doing, and the readiness assessment is the document that proves it does.
One discipline keeps the assessment honest over time: re-baseline before every major commitment, because readiness moves. A source of truth you governed last year drifts as policies change. A verification culture you scored Level 3 erodes when the two careful designers leave. An accessibility process degrades under a tight deadline. The strategist who treats the readiness assessment as a living instrument, re-run before each roadmap phase, keeps the function's promises calibrated to its real capacity. The one who scores it once and files it is back to negotiating deadlines blind within a year.
Key Takeaways
- A readiness assessment is the honest baseline you build before promising leadership a transformation; it converts "when will it be done" into "here is where we are strong, where we are exposed, and the sequence that captures the upside safely."
- AI collapses production time but leaves verification time untouched, so a function that cannot verify at the new speed does not get faster, it ships its mistakes faster.
- Readiness has four independent dimensions: content maturity (is there a source of truth to ground on), data (can you measure beyond completions), verification culture (is the habit logged and named), and accessibility posture (is the gate already in place).
- The weakest dimension, not the strongest, determines what you can safely promise; grounding AI on a stale source produces wrong content faster, which is worse than slow.
- Score each dimension on evidence, not optimism: if you cannot produce the sign-off log or the conformance report on demand, the capability is not there, because that is exactly the standard an auditor will apply.
- The assessment's value is realized in the leadership conversation, where four scored dimensions let you say yes and no surgically instead of resolving a binary toward a reckless go.
- Re-baseline before every major commitment, because readiness drifts as content goes stale, careful people leave, and accessibility erodes under deadline pressure.
- The strategist's job is not to be the most enthusiastic person about AI but the most accurate person about the function, because accuracy is what turns the upside into a result instead of a liability.
Skill.re