โ†
AI for Instructors & Learning Professionals
Capable ยท M20 ยท lesson 20 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
The Invalid-Item Trap
๐Ÿ“–
now learning

The Invalid-Item Trap

15 min

An anti-bribery certification went live across a 4,000-person sales organization. The pass rate was 96%. Leadership was delighted. Eight months later, a regional team closed a deal with a facilitation payment that the policy plainly forbids, and in the review that followed someone pulled the certification questions. Every single one was well written. Every one had a correct answer. And every one could be answered correctly by a person who did not understand the policy at all. The test had certified 4,000 people as competent and measured none of them. This lesson is about how a question that looks completely fine can test absolutely nothing, and the specific checks that expose it before it ships.

How a Good-Looking Question Tests Nothing

The unsettling truth at the center of this chapter is that the qualities that make an item look valid, clean grammar, the right vocabulary, a clearly correct answer, plausible-seeming wrong answers, have almost nothing to do with whether it is valid. Recall the core term from the previous lesson. Validity is the degree to which an item measures the specific competence it claims to measure, and nothing else. An item is invalid when something other than the target competence determines whether a learner gets it right. The cruel part is that invalidity is usually invisible on the surface, because the thing that makes the item answerable for the wrong reason is hidden inside the structure of the question, not in its wording.

The general name for the problem is construct-irrelevant variance: differences in who passes and who fails that are driven by something other than the skill being tested. Why you care: every bit of construct-irrelevant variance is a leak in your measurement, a way for a learner to get the right answer without the right knowledge, or the wrong answer despite having it. In a certification, those leaks are not academic. They are the exact mechanism by which an unqualified person walks out certified. The job of this lesson is to teach you to see the leaks, because AI is a prolific source of them and they are catastrophically easy to miss.

AI makes the trap worse for a precise reason. A generative model is optimized to produce text that looks like a good test question, because looking like a good test question is what its training data rewards. It is not optimized to produce a question that discriminates between people who have the skill and people who do not, because the model has no access to your learners and no concept of the underlying competence. So the model is, in effect, a machine for producing items that pass the eye test and fail the validity test. The polish is the danger. A rough-looking item invites scrutiny. A polished invalid item sails through review precisely because it looks finished.

It is worth sitting with why this is so disorienting, because the instinct it violates is a deep one. In almost every other part of a learning build, surface quality and underlying quality move together. A well-written paragraph usually reflects clear thinking. A clean storyboard usually reflects a sound design. So a reviewer's eye, trained on years of "if it reads well it probably is well," carries that habit into the item bank, where it stops being true. The correlation between how an item reads and whether it measures anything is close to zero, and may even run backwards, because the AI items that read most smoothly are often the ones where the correct answer was lifted cleanly from the source, which is exactly the verbatim echo that makes the item answerable without knowledge. The reviewer's competence becomes a liability. The better your eye for prose, the more confidently you will wave through an item that prose cannot judge.

A question that looks like a test and a question that is a test are different objects. AI is very good at the first and has no idea about the second.

The Three Classic Leaks

Most invalid AI items fail in one of three recognizable ways. Learn the three and you can scan a bank quickly and catch the majority of the damage.

Leak One: Cueing

Cueing means the item contains a signal that points to the correct answer independent of the knowledge being tested. Why you care: a cued item can be answered by a test-taking trick, so it measures test-taking, not competence. AI produces cued items constantly, because the patterns that cue an answer are statistically common in its training data. The most frequent cue is the verbatim echo: the correct option repeats the language of the stem or the source slide, while the distractors use unrelated phrasing, so the right answer is simply the one that "sounds like" the question. Another classic cue is the grammatically consistent option: three distractors that do not fit the stem's grammar and one that does, so the learner picks the only one that reads smoothly. A third is the longest, most qualified option: models love to write the correct answer with the most hedging and detail, so the longest option is usually right, and learners know it. Each of these lets a person score points with zero subject knowledge.

Leak Two: Testing Recall When the Objective Is Application

The second leak is the cognitive-level mismatch covered in the previous lesson, and it is worth naming again here as a validity failure, not just an alignment one. When the objective requires application, judging a situation, making a decision, performing a procedure, and the item only requires recall, recognizing a fact from a list, the item is invalid for that objective no matter how correct and clean it is. The reason this is a validity failure and not a stylistic preference: a learner who can recall the fact but cannot apply it will pass the recall item and fail the job. The item certifies a competence the learner does not have. AI drifts toward recall by default, so this leak is present in a large fraction of unguided machine-drafted banks, and it is invisible unless you are explicitly comparing the item's cognitive demand to the objective's.

Leak Three: The Implausible Distractor Set

A distractor is a wrong answer option in a multiple-choice item. Why you care: distractors are not filler, they are the working machinery of the question, because a learner who lacks the skill has to be genuinely tempted by a wrong option for the item to discriminate. When AI generates distractors, it tends to produce obviously wrong, unrelated, or absurd options, the patient is "over sixty," the room is "too cold," the answer is "purple", because plausible distractors require understanding the real misconceptions a learner holds, and the model is not reasoning about your learners' misconceptions. The result is an item where the correct answer is obvious by elimination. A learner who knows nothing can still rule out the three silly options and land on the right one. The item has the form of a four-option question and the discriminating power of a true-or-false with an obvious answer.

Item Discrimination: The Number That Tells the Truth

There is a measurement concept that cuts through all of this, and every learning professional working with AI-generated assessments should hold it. Item discrimination is the degree to which an item separates learners who know the material from those who do not. Why you care: it is the single most useful empirical signal of whether an item is doing its job, because it is computed from real learner responses, not from how the item looks. A well-functioning item is one that strong learners tend to get right and weak learners tend to get wrong. An item with poor discrimination is one where strong and weak learners pass at roughly the same rate, which means the item is not measuring the competence that separates them. It is measuring something else, or nothing.

The practical version, without any statistics, is a question you can ask of any item: would a person who has the skill and a person who lacks the skill answer this differently? If a slide-skimmer who never understood the policy can get the item right by cueing or elimination, the item does not discriminate, and a non-discriminating item is an invalid item by definition, because it is not separating competence from its absence. This is the deepest reason the eye test fails. An item can look rigorous and have zero discriminating power, and you cannot see discrimination by reading the item. You see it by asking whether the right and wrong learners would be sorted by it, and, once the item is live, by checking whether they actually are.

Once an item has been answered by real learners, discrimination stops being a thought experiment and becomes a number you can read. The simplest version: split your learners into the group that scored high on the assessment overall and the group that scored low, and look at how each group did on the single item in question. A healthy item is one the high group gets right far more often than the low group, because the same competence that earned high scores overall is the competence the item rewards. An item where the two groups perform about the same is telling you something blunt and important: this item is not measuring whatever separates your strong learners from your weak ones. Worse, an item the low group passes more often than the high group is actively miscued, rewarding a trick that test-wise weak learners exploit and strong learners overthink. You do not need to compute a coefficient to act on this. The pattern of who passes is the item confessing what it measures.

This is also why discrimination is the concept that finally disciplines AI-generated banks at scale. You cannot read 500 machine-drafted items for validity by eye without your prose instinct betraying you, but you can watch the response data and let the weak items raise their hands. Items that everyone passes regardless of overall performance are the cued and eliminable ones the model loves to produce, and they surface in the data even when they survived the read. The lesson is not to skip the editorial checklist; it is that the checklist and the response data are two views of the same question, and a high-stakes certification deserves both, because the cost of a single invalid item is measured in people certified to do something they cannot do.

The question that matters is not "is this a good question?" It is "would the person who can do the job and the person who cannot answer this differently?" If not, it measures nothing.

The Detection Checklist

Put the leaks and the discrimination question together into a scan you can run on any AI-generated item before it goes near a learner. This is the artifact to keep.

CheckWhat you are hunting forThe item fails if
Cognitive levelMatch between the item's demand and the objective's verbThe objective says apply or analyze and the item only requires recall
Verbatim echoThe correct option mirroring the stem or source wordingA learner could pick the answer by matching words to the slide
Grammatical or length cueThe correct option being the only smooth fit, or the longest and most hedgedA test-wise learner could pick it by structure alone
Distractor plausibilityWrong options built from real learner misconceptionsThe distractors are absurd or unrelated and removable by elimination
Discrimination questionWhether a skilled and an unskilled learner would answer differentlyA person without the competence could get it right by trick or elimination
Job fidelityWhether a pass predicts on-the-job performanceThe item tests something never done in isolation on the actual job

Any item that fails any row is suspect and gets repaired or discarded before it can certify anyone. Notice that none of these checks can be performed by reading the item for whether it "seems right." Every one requires comparing the item to something outside it: the objective, the source, the real misconceptions, the actual learners. That comparison is the work, and it is human work, because the model that produced the item has access to none of those external referents.

A Worked Walkthrough: The Bribery Item

Return to the anti-bribery certification and run one of its items through the checklist, before and after.

Before (the invalid item that shipped). The objective was at the apply level: "Given a business scenario, determine whether a proposed payment violates the anti-bribery policy." The AI drafted: "What does the anti-bribery policy prohibit? A. Offering anything of value to a government official to obtain or retain business. B. Sending internal emails after 6pm. C. Holding meetings in rooms without windows. D. Using a personal laptop for work." Run the checklist. Cognitive level: the objective says apply, the item tests recall, fail. Verbatim echo: option A is the policy's headline sentence almost word for word, fail. Distractor plausibility: B, C, and D are absurd and unrelated, removable instantly by elimination, fail. Discrimination: a person who has never read the policy can pick A by recognizing the only serious-sounding option, so a skilled and an unskilled learner answer identically, fail. Job fidelity: on the job, nobody is handed four options with three jokes; they face a real, ambiguous payment decision, fail. The item failed every row, and yet it read as a perfectly professional certification question. That is the trap in one object.

After (the valid item). "A foreign customs official says your shipment will clear faster if you pay a 200 dollar 'expediting fee' directly to him in cash. Your contract has a hard deadline and a penalty for late delivery. Under the anti-bribery policy, can you make the payment, and what is the basis for your answer?" Run the checklist again. Cognitive level: the learner must apply the policy to a novel situation, pass. Verbatim echo: nothing in the answer mirrors a slide sentence, pass. Distractors, if it is multiple choice, are built from real misconceptions, that a small payment is allowed, that a facilitation payment is exempt, that a deadline creates an exception, each one a wrong belief a real employee actually holds, pass. Discrimination: a person who understands the policy distinguishes a prohibited bribe from a permitted courtesy and the unskilled person does not, so they answer differently, pass. Job fidelity: this is exactly the decision the employee faces in the field, pass. Same policy, same objective, an item that now does what the certification claimed all 4,000 items did.

The lesson is not that the AI was useless. The AI drafted both items in seconds, and the second, valid item is just as fast to produce once a human specifies the level, forbids the verbatim echo, and demands misconception-based distractors. The lesson is that the first item shipped because nobody ran the checklist, and the absence of that ten-second human check is what turned a fast build into a 4,000-person false certification.

The Iron Rule and What It Costs to Skip It

The bright line holds without exception: AI does not certify a learner as competent. AI may draft the item; a human validates that the item discriminates competence from its absence and owns the pass or fail decision, full stop. The invalid-item trap is the reason this rule exists. An invalid item is not a small quality issue. It is a silent failure of the entire purpose of assessment, and it is silent precisely because the item looks fine. The cost is not visible at launch, when the pass rate is high and everyone is pleased. It surfaces later, in the incident, the audit, the lawsuit, when someone certified as competent does the thing they were certified not to do, and a reviewer pulls the questions and finds that the test never measured the competence at all.

So when a compliance officer asks the question that should keep you up at night, "can someone pass this test without actually being able to do the job," your answer has to be grounded in the checklist, not in the pass rate. You say: every item was checked against its objective's cognitive level, screened for verbatim and structural cues, built with distractors from real misconceptions, and confirmed to discriminate between learners who have the skill and learners who do not. That answer is the difference between a certification that means something and a 96% pass rate that means nothing.

Key Takeaways

  • The qualities that make an item look valid, clean grammar, right vocabulary, a clear correct answer, have almost nothing to do with whether it is valid; invalidity is usually invisible on the surface.
  • Construct-irrelevant variance is any reason a learner gets an item right or wrong other than the target competence; in a certification, each leak is a path for an unqualified person to pass.
  • AI is optimized to produce items that look like good questions, not items that discriminate competence, so the polish itself is the danger and a polished invalid item sails through review.
  • The three classic leaks are cueing (verbatim echo, grammar fit, longest option), testing recall when the objective is application, and implausible distractors that are removable by elimination.
  • A distractor is the working machinery of an item; plausible distractors require understanding real learner misconceptions, which the model does not have, so AI distractors are often absurd and non-discriminating.
  • Item discrimination, whether skilled and unskilled learners answer differently, is the truest signal of validity because it comes from real responses, not from how the item looks; a non-discriminating item is invalid by definition.
  • The detection checklist (cognitive level, verbatim echo, structural cues, distractor plausibility, the discrimination question, job fidelity) requires comparing the item to the objective, the source, the misconceptions, and the learners, which is human work.
  • The iron rule holds: AI may draft the item, but AI does not certify a learner as competent; a human confirms the item discriminates competence and owns the pass or fail decision, full stop.