โ†
AI for Instructors & Learning Professionals
Strategic ยท M9 ยท lesson 9 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Evaluating Learning-AI Vendors for Verifiability and Accessibility
๐Ÿ“–
now learning

Evaluating Learning-AI Vendors for Verifiability and Accessibility

15 min

A head of learning sits across a demo screen as a vendor clicks "generate course." Forty seconds later a finished module appears, narrated, branched, quizzed. It is beautiful. So she asks the only question that matters: "Show me where that procedure step came from." The rep smiles and says the model "draws on best practices." There is no source button, no citation, no document the step traces to. The tool just generated a regulated claim from its training data and presented it as truth. She closes the laptop. That single missing button, the inability to show a source, is the difference between an audit-grade tool and a convenient one, and it is the whole subject of this lesson.

The Demo Is Not the Product

Every learning-AI demo is engineered to make you feel the speed and ignore the seams. The module appears in seconds, the avatar speaks fluently, the quiz looks finished. What the demo never shows you is the moment six months later when a compliance officer points at screen 18 and asks where the lockout/tagout step came from, or when an accessibility auditor runs the course through a screen reader and the AI-generated interaction collapses into silence. Procurement is the one chance you have to make the vendor prove the boring things before the logo ends up on a failed course with your name on the sign-off.

Here is the term that anchors this lesson. Verifiability is the property that lets a human trace any claim a tool produces back to an approved source, fast enough to be practical at the scale you operate. Why you care: a tool that cannot show a source is not saving you time, it is moving the verification burden downstream to the moment of greatest cost, the audit, the incident, the lawsuit. The L4 strategist's job is not to be dazzled by output. It is to interrogate the tool the way a SME, an accessibility auditor, and a CFO eventually will, and to do it in the procurement window when you still have leverage.

The reason this falls to L&D and not just to IT procurement is that the failure modes are learning-specific. IT will check the security posture and the SSO integration. IT will not ask whether the AI generates a policy threshold from training data instead of retrieving it from your SOP, whether the auto-generated captions meet WCAG 2.2 AA, or whether an AI-drafted assessment item actually measures the objective it claims to. Those are learning questions, and the person who owns what the workforce is taught is the only one in the room equipped to ask them.

A tool that cannot show you a source is not a productivity gain. It is a verification debt with interest, and the audit is when the bill comes due.

Grounding and Source Transparency: The First Gate

The single most important question in any learning-AI evaluation is this: when the tool states a fact, can it show you the source, and is that source one you control? This is the difference between grounded generation and ungrounded generation, and it is not a nice-to-have. It is the line between a tool that can be defended in a regulated environment and one that cannot.

Let me define the terms in L&D language. Grounding, also called RAG (retrieval-augmented generation), means the tool answers from a source you loaded and approved, your policy, your SOP, your SME transcript, rather than from the open web or the model's training data. Why you care: a grounded tool produces a claim with a receipt attached, and a receipt is what an auditor wants. Source transparency means the tool actually exposes that receipt to you: a citation, a highlighted passage, a link back to the document and ideally the line. A tool can be grounded internally but opaque to you, which is almost as bad, because you cannot verify what you cannot see.

The procurement scene to run is concrete. Ask the vendor to load one of your real documents, a safety procedure or a policy, and then ask the tool a question whose answer is a specific threshold or step. Watch what comes back. Does the tool quote the source and show you where it pulled the answer, or does it produce a fluent paragraph with no provenance? Then ask the harder version: ask it something your document does not cover, and see whether it refuses and says the source does not address it, or whether it confidently invents an answer. A tool that cites when it can and refuses when it cannot is doing the job. A tool that always answers, source or no source, is a liability engine wearing a friendly interface.

The Questions That Expose Ungrounded Generation

The rubric below is built around questions a vendor cannot answer with a slide. Each one is designed to surface a specific failure mode before it ships.

  • Can it cite? When the tool states a regulated fact, does it show the source document and location, or does it produce prose with no traceable origin?
  • Does it refuse? When you ask something outside the loaded sources, does the tool say it cannot answer from the approved material, or does it hallucinate a confident answer?
  • Whose data grounds it? Is the tool grounded on your controlled corpus, or on a general knowledge base the vendor curated, or on the open web? Only the first is yours to defend.
  • Can you see the prompt and the retrieval? Can an administrator inspect what the tool retrieved and how it built the answer, or is it a black box that emits finished content?
  • Does it version the source? When your policy updates, does the tool re-ground on the new version, and can it show which version a given claim came from?

Accessibility Conformance: The Gate Vendors Skip

The second gate is the one vendors are most likely to wave past with a single word: "accessible." Press on that word and it usually dissolves. The standard is not "accessible," it is WCAG 2.2 AA, the W3C Recommendation finalized 5 October 2023 and the conformance target that Section 508 incorporates by reference. The vocabulary you need in the room is the VPAT (Voluntary Product Accessibility Template) and its filled-in form, the ACR (Accessibility Conformance Report). Why you care: the VPAT/ACR is the document where a vendor states, claim by claim, how their product meets each WCAG success criterion. A vendor who cannot produce a current one is telling you they have not done the work.

The trap specific to AI is that the tool generates the experience on the fly, so accessibility is not a fixed property you can audit once. An AI that auto-generates captions can produce captions that are grammatically fluent and factually wrong, which is worse than no captions because they look trustworthy. An AI that generates alt text can describe an image confidently and incorrectly. An AI that builds an interaction can build one that a keyboard user cannot reach and a screen reader cannot announce. The conformance question is therefore not only "is the player accessible" but "is the AI-generated content accessible by construction, and who verifies that on every generation."

Vendor claimWhat it usually meansWhat you must verify
"Our product is accessible"The marketing site passed a checker onceA current VPAT/ACR mapped to WCAG 2.2 AA, dated and version-specific
"We auto-generate captions"Speech-to-text with no accuracy guaranteeCaption accuracy on your accented, jargon-heavy content, and a human-edit workflow
"AI writes the alt text"Generated descriptions, sometimes wrongWhether a human reviews alt text before publish, and how errors are caught
"Fully keyboard navigable"The shell is; generated interactions may not beKeyboard and screen-reader testing on AI-generated interactions, not just the chrome
"508 compliant"A claim with no evidence behind itThe ACR, the testing methodology, and the date of last audit

The bright line is non-negotiable and worth stating the behavioral-health way: an AI-generated experience that fails WCAG 2.2 AA does not ship. Accessibility is a gate, not a polish step. In procurement that means you do not accept "we are working toward conformance" as an answer. You ask for the ACR, you ask how generated content is checked, and you treat a missing or stale report as a failed criterion, because that is exactly how a 508 audit will treat it.

Data Terms: The Clause That Decides Ownership

The third gate lives in the contract, not the demo, and it is the one most likely to be skipped because it is boring and legal and nobody on the L&D side feels qualified to read it. Read it anyway, because a single clause decides whether your learners' data trains the vendor's model, whether your proprietary content becomes part of someone else's product, and whether you can get your data out when you leave. The full treatment of the model-training question is the next lesson; here the point is narrower: data terms are a procurement criterion, and a tool that is excellent on grounding and accessibility but ruinous on data terms is still a failed evaluation.

The questions to put in the rubric are specific. Does the contract say your content and your learner data are used only to provide the service to you, or does it grant the vendor a license to use your data to improve their models? Is there an explicit opt-out from model training, and is it the default or a setting you have to find? What happens to your data when the contract ends: is it returned and deleted, or retained? Where is it processed, and does that satisfy your GDPR-style obligations? A vendor who answers these crisply has thought about it. A vendor who says "we take privacy seriously" and changes the subject has told you the answer is in the fine print and you will not like it.

The feature list is the romance. The data-processing clause is the prenup. Read the prenup, because that is the document that governs what happens when it goes wrong.

A Worked Example: Two Vendors, Same Demo

Two vendors give the identical dazzling demo to the same head of learning evaluating a tool for a regulated safety curriculum. Watch how the rubric separates them.

Vendor A, the convenient tool. The demo is flawless. Asked to show a source for a generated procedure step, the rep says the model uses "industry best practices" and there is no citation feature. Asked for a VPAT, they send a marketing one-pager that says "accessible" with no WCAG mapping. Asked about model training, the contract grants them a broad license to "use customer data to improve our services," with no opt-out surfaced. Every answer is smooth, and every answer is a deferral of risk onto the buyer. On the rubric, Vendor A fails the grounding gate (cannot cite, does not refuse), fails the accessibility gate (no ACR), and fails the data gate (training license, no opt-out). The speed was real. The defensibility was zero.

Vendor B, the audit-grade tool. The demo is no faster. But when asked for a source, the tool quotes the loaded SOP and highlights the exact line, and when asked something outside the document it says the source does not cover it. The vendor produces a current ACR mapped to WCAG 2.2 AA, names the AI-specific gaps honestly (auto-captions need human review for accented speech), and describes the human-in-the-loop check. The contract states customer data is used only to provide the service, model training is off by default with a written opt-out, and data is returned and deleted at termination. Vendor B passes all three gates. The strategist can stand in front of a compliance officer, an accessibility auditor, and a CFO and answer every question with a document.

The lesson is not that one vendor was faster. They were equally fast. The lesson is that speed is the easy thing to demo and the worthless thing to buy on. What you are actually purchasing is the ability to answer "show me the source," "show me the conformance report," and "show me the data clause" without flinching. The rubric is how you find out which vendor can do that before you sign, not after the audit.

There is a subtler test buried in that worked example that is worth naming on its own: the way each vendor handled the limits of its own product. Vendor A claimed perfection, the AI was accurate, the captions were fine, the tool just worked. Vendor B volunteered that auto-captions need human review for accented speech and described the workflow that catches it. Counterintuitively, the vendor who admitted a weakness is the more trustworthy one, because no generative tool is free of hallucination or caption error, so a claim of perfection is either ignorance or marketing, and both predict that the verification burden will land silently on you. A vendor who can tell you precisely what their AI gets wrong has almost certainly built the human checks to catch it, and that honesty is itself a procurement signal worth scoring.

Turning the Rubric Into a Contract

A rubric that lives only in the evaluation is half a rubric. The vendor's reassuring answers in the room have no force once the ink is dry unless they are written into the agreement as warranties, so the final discipline of an audit-grade procurement is to bind the passing answers into the contract itself. The source-citation capability becomes a warranted feature, not a demo nicety. The WCAG 2.2 AA conformance becomes a maintained obligation with a current ACR and the human-review workflow named, not a one-time claim. The data terms, the no-training default, the processing region, the return-and-delete, become enforceable clauses rather than verbal comfort. The reason is simple: at the audit, you will be defending the contract, not your memory of the demo, and a promise that is not in the contract is a promise the vendor is free to forget.

This also reframes what procurement is for. It is not the moment you choose the most impressive tool; it is the moment you convert a vendor's marketing into legally durable commitments while you still have the leverage of an unsigned buyer. The three gates, grounding, accessibility, and data, are the structure of that conversion. Run them as a scored, weighted, pass-or-fail rubric, bind the passing answers into warranties, and treat any single-gate failure as disqualifying, and you have done the one thing a glossy demo can never do for you: you have made the tool defensible before it ever touches a learner, which is the only state in which an L&D strategist should ever let a learning-AI tool into the building.

The One-Page Rubric to Keep

Distill the evaluation to a scored sheet you run on every learning-AI tool, and weight the three gates so that a failure on any one of them is disqualifying, not averaged away by a high feature score.

  • Grounding and source transparency. Cites sources, refuses outside the corpus, grounds on your controlled data, exposes retrieval, versions the source. Any "no" on regulated content is a fail.
  • Accessibility conformance. Current VPAT/ACR mapped to WCAG 2.2 AA, AI-generated content checked by a human, keyboard and screen-reader tested on generated interactions. Missing ACR is a fail.
  • Data terms. Data used only to serve you, model-training opt-out by default, defined processing location, return-and-delete at exit. A training license with no opt-out is a fail.
  • Vendor honesty. Names its own limitations versus papering over them. A vendor who tells you what their AI gets wrong is more trustworthy than one who claims it never errs.

Key Takeaways

  • The demo is engineered to sell speed and hide the seams; procurement is your one chance to make the vendor prove verifiability, accessibility, and data terms before the logo lands on a failed course.
  • Verifiability means any claim traces back to an approved source fast enough to be practical; a tool that cannot show a source is not saving time, it is deferring verification cost to the audit.
  • The grounding gate: the tool must cite when it can, refuse when the source does not cover the question, and ground on data you control, not the open web or the vendor's curated base.
  • The accessibility gate is WCAG 2.2 AA, evidenced by a current VPAT/ACR; AI-generated captions, alt text, and interactions must be checked by a human because fluent wrong output is worse than none.
  • The data gate lives in the contract: confirm your data serves only you, model training is opt-out by default, processing location is defined, and data is returned and deleted at exit.
  • Score the three gates so a failure on any one is disqualifying, not averaged away by a glossy feature list; a high feature score does not buy back a missing ACR or a training license.
  • A vendor who names their own AI's limitations is more trustworthy than one who claims it never errs; honesty about failure modes is itself a procurement signal.
  • The obligation never transfers to the platform: when the auditor asks where a claim came from, "the vendor's tool generated it" is not an answer, and the iron rule still holds: AI assists, the human verifies, the human owns the decision.