Generating Assessment Items That Measure the Objective
A designer asks an AI tool for ten quiz questions on a forklift safety objective. Ninety seconds later she has a clean, grammatical, professional-looking item bank. Every question is well written. Every one has a defensible-looking right answer. And not a single one of them measures whether a person can safely operate a forklift. They measure whether the learner read the slide. The bank looks finished, it looks rigorous, and it certifies nobody. This lesson is about the difference between a question that looks like a test and a question that is one.
The Objective Is the Contract
Before a single item is drafted, one thing has to be settled, and AI cannot settle it for you: what the assessment is actually a promise about. Start with two words this whole chapter turns on. Validity is the degree to which a question measures the specific thing it claims to measure, and nothing else. Why you care: an invalid item gives you a score that means something other than what you think it means, and you will make a pass or fail decision on that wrong meaning. Reliability is the degree to which a question gives a consistent result across learners and occasions. Why you care: a reliable but invalid item consistently measures the wrong thing, which is worse than noise, because it looks trustworthy. The whole job of this chapter is validity. A test can be perfectly reliable and completely useless.
The instrument that makes validity possible is the learning objective: a specific, measurable statement of what the learner will be able to do after the instruction, written with an observable verb. "Understand forklift safety" is not an objective, because you cannot observe "understand." "Identify the three conditions that require a forklift to be removed from service before the next shift" is an objective, because you can watch a person do it or fail to. The objective is the contract the assessment has to honor. An item is valid when it asks the learner to perform the action in the objective, at the cognitive level the objective demands, under conditions close enough to the job that a pass actually predicts on-the-job competence.
This idea has a name worth keeping. Constructive alignment is the principle that the objective, the assessment, and the instruction all point at the same performance, so that what you teach, what you ask, and what the learner has to do on the job are one thing, not three. Why you care: when alignment breaks, you can have excellent content and an excellent-looking test and still certify people who cannot do the work, because the test quietly drifted to measuring something easier than the objective. AI is extraordinarily good at producing that drift while making it invisible, because the drifted item still reads beautifully.
An item is not valid because it is well written. It is valid because it makes the learner do the thing the objective promised, at the level the job demands.
Bloom's Level Is Not a Decoration
The most common way an AI item silently betrays its objective is by sliding down the cognitive scale. Bloom's taxonomy is a ladder of cognitive demand, from remembering and understanding at the bottom, through applying and analyzing, up to evaluating and creating at the top, each rung named by the kind of thinking it requires. Why you care: an objective lives on a specific rung, and an item has to test the learner on that same rung. If the objective says "apply" and the item tests "remember," the item is invalid no matter how polished it is, because it is measuring an easier thing than the one the job requires.
Here is the trap in its purest form. The objective is at the apply level: "Given a spill scenario, select the correct containment procedure." A learner who can do this can stand in a real spill and act. Now watch the AI draft a question. Asked for a quiz item on that objective, a model will very often produce something like: "What is the first step of the spill containment procedure?" with four options, one of which is the textbook first step. That item reads as on-topic. It uses the right vocabulary. It feels like it tests the procedure. It does not. It tests whether the learner can recall the first step from the page they just read. A person can memorize "step one is contain the source" and still freeze in front of an actual chemical spill, because choosing the first step from a list is a recall task and running a containment procedure in a novel situation is an application task. The item dropped two full rungs below the objective, and the drop is invisible unless you are looking for it.
The reason AI does this so reliably is mechanical, and worth understanding so you stop being surprised by it. A generative model produces the most statistically likely text given your prompt. The most likely "quiz question about a procedure" in its training data is a recall question, because recall questions are the most common kind of quiz question ever written. So when you ask for "a quiz item on this objective," the model regresses to the most common shape of quiz question, not to the cognitive level your specific objective demands. It is not reasoning about your objective. It is matching the pattern of "quiz question." Left unguided, it will always tend to drift toward recall, because recall is where the gravity of the training data pulls.
Naming the Level in the Prompt
The first defense is to stop asking for "a question" and start asking for a question at a named cognitive level, with the action spelled out. Instead of "write a quiz question about spill containment," the disciplined prompt is: "Write a scenario-based item that requires the learner to apply the containment procedure to a situation they have not seen before. Give them a novel spill scenario with specific conditions, and ask them to determine the correct containment action. Do not ask them to recall a step from a list." This does not guarantee a valid item. It does dramatically raise the odds that the draft lands on the right rung, because you have replaced the model's default pattern ("quiz question") with the shape you actually need ("apply-level scenario item"). The model still has no idea whether the result is valid. You named the target so the draft would at least aim at it.
Even then, the verb in the objective and the verb in the item have to be checked against each other by a human, every time. The cheapest, most powerful habit in this whole chapter is to put the objective and the item side by side and ask one question: does the action the item requires match the action the objective promised? If the objective says "diagnose" and the item says "list," the item is invalid. If the objective says "select the correct procedure in a novel case" and the item says "name the first step," the item is invalid. The match is not about topic. It is about the verb.
Testing the Skill, Not Reading Comprehension
There is a second, sneakier way an AI item measures the wrong thing, and it has nothing to do with Bloom's level. The item can quietly become a reading test. This is the single most common defect in machine-drafted assessments, and it is worth seeing in detail, because once you see it you cannot unsee it.
When a model drafts a question and its options from a piece of content, it tends to lift the correct answer almost verbatim from the source text and to write the wrong options in noticeably different language. The result is an item a learner can answer correctly without knowing anything about the subject, purely by matching the wording. The right answer is the option that sounds most like the paragraph. A skilled test-taker, or anyone who skimmed the slide thirty seconds ago, will pick it by surface pattern, not by understanding. The item is measuring construct-irrelevant variance: differences in scores that come from something other than the skill being tested. Why you care: every point of construct-irrelevant variance is a point of false signal, and false signal in a certification is exactly how you certify people who cannot do the job.
Watch a concrete before and after. The source slide says: "A defibrillator must not be used if the patient is lying in standing water, because the current can travel through the water and injure the rescuer." A model drafts:
- Question: "When must a defibrillator not be used?"
- A. If the patient is lying in standing water, because the current can travel through the water and injure the rescuer. (correct)
- B. If the patient is over the age of sixty.
- C. If the room temperature is below normal.
- D. If more than one rescuer is present.
Option A is the slide sentence, almost word for word. The three distractors are about unrelated things in unrelated language. A learner who never understood why water is dangerous can still pick A, because A is the option that matches the text they just saw. This item certifies recognition of a sentence, not understanding of an electrical hazard. It is invalid, and it is invalid in a way that a quick read will miss entirely, because the question is grammatical, on-topic, and has a genuinely correct answer.
The after version forces the skill. "A patient collapses on a poolside deck. There is a thin film of water under and around them from a recent splash, and the patient's torso is dry. You have a defibrillator. What should you do before delivering a shock, and why?" Now the learner cannot pattern-match a sentence. They have to reason about the hazard, apply the rule to a situation the slide did not literally describe, and justify the action. The wording of the correct answer no longer mirrors the source. The construct-irrelevant variance is squeezed out, and what is left is the skill the objective actually cares about. Same fact, same objective, completely different measurement.
If a learner can answer the question by matching words instead of using the skill, the question is measuring reading, and reading was not the objective.
A Worked Walkthrough: From Objective to Valid Item
Put the whole discipline together on one objective and watch a designer take an AI draft from invalid to valid. The objective, drawn from a data-handling compliance module, reads: "Given a customer request involving personal data, determine whether the request can be fulfilled under the company's data-retention policy, and identify the approval required." This is an apply and analyze objective. The learner has to reason about a real request against a real policy, not recite the policy.
The designer prompts the AI for ten items and gets back, among them, this draft: "How long does the company retain customer personal data? A. 30 days B. 90 days C. 12 months D. 7 years." It is clean. It has a correct answer. It is also a pure recall item that drops the objective from analyze all the way to remember, and it tests a single isolated number rather than the judgment the objective promised. A learner can memorize "12 months" and still mishandle every real request that crosses their desk, because real requests are messy and the policy has conditions. This is the classic AI item: right-looking, on-topic, and measuring almost nothing the objective asked for.
The table tracks the repair, defect by defect.
| What the designer checked | The AI draft (invalid) | The repaired item (valid) |
|---|---|---|
| Cognitive level vs. objective | Remember (recall one number) | Analyze (judge a request against the policy) |
| Action required of the learner | Recognize a retention period | Determine fulfillability and name the approval |
| Construct-irrelevant variance | Answer matches the slide verbatim | Answer requires reasoning, not word-matching |
| Fidelity to the job | Tests a fact never used in isolation on the job | Tests a decision the learner makes weekly on the job |
| What a pass actually predicts | The learner read the slide | The learner can handle a real request safely |
The repaired item reads: "A customer emails asking you to delete all of their account data immediately, but they have an open dispute on their last invoice. Under the data-retention policy, can you fulfill the deletion request right now, and if not, what approval or condition is required first?" This item lives at the objective's level. It cannot be answered by recall. It mirrors a decision the learner will actually make. And critically, the AI could draft a candidate of this shape in seconds once the designer told it the level, the action, and the prohibition on word-matching. The speed is real. The validity came from the human who knew what the objective was a contract for and refused to ship the draft that broke it.
Notice what the designer did not do. She did not trust that ten clean-looking items were ten valid items. She did not assume that because the model used the right vocabulary the question tested the right skill. She read each item against its objective with the validity question in hand, and she found that the most professional-looking draft in the set was the most quietly broken. That is the work. The drafting got faster. The judgment did not move.
The Iron Rule for Item Generation
This chapter sits on a bright line the whole program enforces: AI does not certify a learner as competent. AI may draft the item, the options, and the feedback. A human validates that the item measures the objective and owns the pass or fail decision, full stop. When you let an AI generate an item bank and push it live without checking each item against its objective, you have not saved time. You have outsourced the validity of your certification to a pattern-matching system that has no concept of validity, and you have done it silently, which means nobody will catch it until a certified person fails on the job.
The practical version of the rule is a short, non-negotiable sequence. Start from a real, observable objective at a known Bloom's level. Prompt for items at that named level, with the action spelled out and word-matching forbidden. Then, for every single item, put it beside its objective and ask: does this require the action the objective promised, at the right level, without letting the learner pass by recognizing words instead of using the skill. Keep the items that pass. Repair or discard the rest. The model accelerates the first two steps. The third step is yours, and it is the step that decides whether your assessment means anything at all.
So when an auditor or a head of learning asks the question that ends careers, "how do you know this test actually measures competence," you do not point at the volume of items or the speed of the build. You point at the objective each item traces to, the cognitive level it was written and checked against, and the human who validated the match. That answer, not the clean formatting of the bank, is what makes a certification defensible.
Key Takeaways
- Validity is whether an item measures the specific thing it claims to measure; reliability is whether it does so consistently. A reliable but invalid item is dangerous because it looks trustworthy while measuring the wrong thing.
- The learning objective is the contract the item must honor: an item is valid only when it makes the learner perform the objective's action, at the objective's cognitive level, under job-like conditions.
- Constructive alignment means the objective, the assessment, and the instruction all point at the same performance; AI is very good at breaking that alignment invisibly while the item still reads beautifully.
- AI items drift toward recall because the model produces the most common shape of "quiz question," which is a recall question, not the cognitive level your specific objective demands.
- Naming the Bloom's level and the required action in the prompt, and forbidding word-matching, raises the odds of a valid draft, but never guarantees one.
- Construct-irrelevant variance, especially when the correct answer mirrors the source text, turns an item into a reading test; a learner who matches words instead of using the skill passes a question that certifies nothing.
- The validity check is a human habit: put each item beside its objective and confirm the verb, the level, and the absence of word-matching, every time, for every item.
- The iron rule holds: AI may draft the item, but AI does not certify a learner as competent. The human validates the item against the objective and owns the pass or fail decision, full stop.
Skill.re