Aligned Objectives and a Validated Item Bank in One Flow
A medical-device manufacturer needs a certification exam for technicians who service infusion pumps, and the old way took a psychometrician three weeks to build forty items. The new designer runs it as one flow: the model drafts the objectives and the items together, then she culls eleven items that look right and measure nothing, and validates every survivor against the objective it claims to test. Two days later the item bank is done, and every question in it can answer the one challenge that ends a certification program: prove this item measures the skill it certifies. That single flow, objectives and items generated together and validated as a pair, is Stage 2 of the pipeline, and it is where a course stops being content and becomes an assessment a regulator will accept.
Why Objectives and Items Belong in One Flow
This is Stage 2 of the source-to-certified-course pipeline, and it builds directly on the grounded draft from Stage 1. The grounded content told the learner what is true. Stage 2 builds the part of the course that decides whether the learner can do the job: the objectives and the assessment that measures them. The reason these two things must be built in one flow, not in separate passes weeks apart, is the principle that governs all sound assessment design.
That principle is constructive alignment, and it deserves a plain definition. Constructive alignment means the objective, the assessment, and the content all point at the same target: what the objective says the learner will be able to do is exactly what the assessment measures and exactly what the content teaches. Why you care: when alignment breaks, you get the most dangerous artifact in learning, a course that looks complete and certifies the wrong people, because the test measures something other than what the objective promised. A learning objective is a precise statement of what the learner will be able to do after the course, written with an observable verb at a specific cognitive level. The cognitive level comes from Bloom's taxonomy, the verb-based ladder from remembering at the bottom through understanding, applying, analyzing, evaluating, and creating at the top. Why you care: an objective that says "the technician will understand the calibration procedure" and an item that asks the technician to perform the calibration are at different Bloom's levels, and the mismatch means the test does not measure the objective even when both look fine in isolation.
Building objectives and items in separate passes is how alignment quietly breaks. The objective gets written, approved, and filed. Weeks later, someone writes items "to cover the objective," and without the discipline of generating them as a pair, the items drift to whatever is easy to ask rather than what the objective demanded. Generating them in one flow, item drafted against objective, in the same breath, keeps the pair locked together from the start, which is exactly the alignment an auditor checks for.
An objective and the item that tests it are a single design decision split across two artifacts. Build them apart and they drift; build them together and they stay aligned.
The Validity Question That Fails Most AI Items
The core risk Stage 2 manages is not a wrong fact; Stage 1 handled grounding. The risk here is the invalid item: a question that is grammatically clean, plausible, and well-formatted, and that measures nothing the objective cares about. Assessment validity is the property that an item actually measures the skill the objective names. Why you care: an invalid item is more dangerous than an obviously broken one, because it passes review on looks and then quietly certifies people who cannot do the task, which is the exact failure that surfaces in an incident or an audit. AI is a prolific source of invalid items precisely because it is so good at producing items that look right.
Here are the specific ways an AI-drafted item looks valid and is not, the patterns you learn to hunt.
The Recall Shortcut
The objective demands application, the ability to perform or decide in a real situation, but the item asks the learner to recall a definition. "Which of the following is the definition of lockout/tagout" tests reading, not the ability to actually de-energize a panel. The item is valid for a remembering objective and invalid for an applying one, and AI defaults to recall items because they are the easiest to generate cleanly. The check is to read the objective's verb and confirm the item demands the same cognitive action.
The Cue Leak
The correct answer is signaled by a clue the question accidentally contains: it is the longest option, the only grammatically matching one, the one using the same unusual word as the stem, or the only one that is plausibly true. A test-wise learner picks it without knowing the content. The item discriminates on test-taking skill, not competence. AI produces cue leaks constantly because it writes the correct answer carefully and the distractors carelessly. The check is to ask whether someone ignorant of the topic could still pick the answer from the surface features.
The Implausible Distractor
A distractor is a wrong answer option, and a good one is wrong but tempting, representing a real misconception a learner might hold. An AI-drafted item often surrounds the correct answer with obviously absurd options nobody would choose, which makes the item easy regardless of competence. The item has the form of a four-option question and the difficulty of a one-option one. The check is whether each distractor represents a plausible error a real learner could make.
The Construct-Irrelevant Difficulty
The item is hard, but for the wrong reason: convoluted wording, a double negative, an unnecessarily complex scenario that taxes reading comprehension rather than the target skill. A competent technician fails because the sentence was a maze, not because they lacked the skill. The difficulty is real and irrelevant to the construct being measured. The check is whether the difficulty comes from the skill or from the language. This pattern is also where accessibility and validity meet: an item that is hard because of dense, convoluted language fails the learner with lower reading fluency or a cognitive disability for reasons that have nothing to do with the competence being certified, which is both an invalidity problem and an equity problem.
Notice what these four patterns have in common. In every case the item is well-formed on its surface and broken underneath, and the break is always a gap between what the item appears to measure and what it actually measures. This is why an invalid item cannot be caught by reading it for quality the way you would proofread prose. A grammatically perfect, on-topic, professionally worded item can be completely invalid, and a slightly awkward item can be perfectly valid. Quality of writing and validity of measurement are independent properties, and AI maximizes the first while having no inherent grip on the second. The entire skill of Stage 2 is learning to read an item for what it measures rather than how it reads, because the model has already optimized how it reads and left the measurement to you.
The One Flow: Generate, Cull, Validate
Stage 2 is a three-move flow, and the moves are sequential for a reason: each one reduces the work of the next, so the human spends judgment where it counts and not on items that should never have survived.
Generate as pairs. Instruct the model, grounded on the Stage 1 source, to produce each objective with the items that measure it, at the stated Bloom's level, and to label each item with the objective it tests and the cognitive action it requires. The output is not a pile of objectives and a separate pile of items; it is a set of locked pairs. This is constructive alignment built in, not inspected for later. Demand more items per objective than you need, because the next move throws many away.
Cull the invalid. Pass every item against the four checks above and discard, do not fix, the ones that fail on construct. Culling rather than repairing is deliberate: a cleanly written invalid item tempts you to tweak it into validity, and tweaking an item whose whole premise is a recall shortcut usually just produces a prettier recall shortcut. It is faster and safer to generate ten and keep four than to rehabilitate the six. The cull is where the bank shrinks from looks-right to is-right.
Validate every survivor against its objective. For each item that survives the cull, confirm the explicit chain: this item, at this Bloom's level, measures this objective, which traces to this Stage 1 source. This is the artifact an auditor wants. Validation is not "the item looks good"; it is "here is the objective, here is the item, here is why the item measures the objective and not something adjacent to it." When this chain is recorded as you go, the item bank ships with its own validity argument attached.
| Move | What it produces | What it prevents |
|---|---|---|
| Generate as pairs | Locked objective-item pairs at a stated Bloom's level | Items drifting to what is easy to ask weeks after the objective was filed |
| Cull the invalid | A bank where every item survived the four validity checks | A clean-looking item that measures recall, a cue, or reading instead of the skill |
| Validate every survivor | An explicit chain from item to objective to source | A bank that cannot answer 'prove this item measures the skill it certifies' |
A Worked Example: The Pump-Technician Exam
Watch the same exam built two ways.
Before, the cover-the-content sweep. The objective on file says "the technician will calibrate the infusion pump to specification." Under deadline, the designer asks the AI to "write forty quiz questions covering pump calibration." It returns forty clean items. Twenty-eight of them ask the technician to identify, define, or recognize calibration facts: "Which value indicates correct calibration," "What is the definition of occlusion pressure." They are remembering items. The objective demanded applying, performing the calibration. Several distractors are absurd, several stems leak the answer through length. The exam passes a quick read because every item is grammatical and on-topic. It ships. Six months later a calibration error in the field triggers a review, and the question arrives that the exam cannot answer: how did certified technicians who passed this exam fail to calibrate correctly. The answer is that the exam never measured calibration; it measured whether you could recognize calibration vocabulary, and the two are not the same skill. The exam certified the wrong competence, fluently.
After, the one flow. The designer runs generate-cull-validate. She instructs the model to write, for the objective "calibrate the infusion pump to specification" at Bloom's applying level, twelve scenario items that require the technician to decide or perform a calibration step in a described situation, each labeled with the objective and the cognitive action. The model drafts twelve. The cull removes five: two were recall items mislabeled as application, two had implausible distractors, one was hard only because of a double negative. Seven survive. For each survivor she validates the chain: this scenario item requires the technician to choose the correct calibration adjustment given a described fault, which is the applying action the objective names, grounded on the calibration SOP from Stage 1. She repeats across objectives until the bank is full of validated survivors. The exam now measures the skill it certifies. When the field review later asks how the exam relates to real calibration competence, she shows the validity chain item by item: objective, Bloom's level, the action the item demands, and the source the content traces to. The exam defends itself.
Both exams were drafted by the same model in minutes. The difference was the flow: the first asked for coverage and got recall dressed as a test, the second asked for aligned pairs, culled the invalid, and validated the survivors. The first was faster to ship and certified the wrong people. The second was the genuine two-days-instead-of-three-weeks win, because it was fast and valid, which is the only kind of fast an assessment is allowed to be.
A test that measures the wrong thing is not a faster test. It is a credential that lies, issued at scale, and AI does not certify a learner as competent, full stop.
The Human Owns the Pass-Fail
Stage 2 ends where the iron rule is most absolute. AI may draft the item, the distractors, and the feedback; a human validates the assessment and owns the pass-fail decision, full stop. This is not a stylistic preference; it is a bright line. The reason is what an assessment does: it certifies a person as competent or not, and that certification has consequences, a technician cleared to service a device, an employee cleared to handle hazardous material, a manager cleared to make hiring decisions. AI can accelerate the drafting of the instrument that produces that judgment. It cannot make the judgment, and it cannot validate the instrument, because validation is the claim that the test measures the right thing, and that claim is a human's to make and to answer for.
This means three things in practice. First, the validity chain is signed by a human, not asserted by the tool: a named person confirms each item measures its objective. Second, the pass-fail cut, the score that separates competent from not, is a human decision informed by the stakes, never a default the tool picked. Third, when a certified person fails in the field, the question "was the assessment valid" is answered by the human who validated it, with the chain they recorded, not by pointing at the model. The whole flow exists to make that human's answer fast, sourced, and defensible. Generate at AI speed, cull without mercy, validate every survivor, and own the decision, because the model produced the questions but the human certifies the people.
One more discipline keeps the human's ownership honest rather than ceremonial: the cut score should be set from what competence actually requires, not from a target pass rate. It is tempting, when AI makes item generation cheap and stakeholders want high completion numbers, to tune an exam until the pass rate looks good, which is a quiet way of letting the desired outcome set the standard rather than the standard setting the outcome. A hazardous-material certification does not get safer because more people passed it; it gets safer because the people who passed it can actually do the task. The human who owns the pass-fail decision owns it precisely so that the cut reflects the consequence of being wrong, the technician at the panel, the operator on the line, not the convenience of a clean dashboard. This is the same accountability that runs through the whole pipeline: the number is allowed to be inconvenient, because the alternative is a credential that certifies the wrong people to make the dashboard look better.
It is worth saying plainly how Stage 2 connects to the stage before and the stage after it, because the pipeline is a chain and the item bank is a link in it. The validated items rest on the grounded content from Stage 1: an item can only be valid if the fact it tests is true and sourced, so a validated item bank built on ungrounded content is validated against a possible fabrication, which is no validation at all. And the validity chain you record here, item by item, becomes part of the accountability record assembled in Stage 4: the same way a regulated content claim carries its source and approver, a certification item carries its objective, its validity argument, and the human who signed it. When an auditor later asks not "is this fact right" but "does this exam measure what it certifies," the answer comes from the chain Stage 2 produced. Building the chain as you go is what makes that answer a lookup instead of a defense improvised under pressure.
Key Takeaways
- Stage 2 of the pipeline builds objectives and assessment items together in one flow, because constructive alignment, the objective, item, and content pointing at one target, breaks when objectives and items are written in separate passes weeks apart.
- A learning objective names what the learner will do with an observable verb at a Bloom's level; an item at a different cognitive level than its objective does not measure it, even when both look fine alone.
- The core Stage 2 risk is the invalid item: clean, plausible, and measuring nothing the objective cares about, which is more dangerous than an obviously broken one because it passes review on looks and certifies the wrong people.
- Hunt four invalidity patterns: the recall shortcut (asks recall for an application objective), the cue leak (answer signaled by surface features), the implausible distractor (absurd wrong options), and construct-irrelevant difficulty (hard because of language, not skill).
- The flow is generate-cull-validate: generate locked objective-item pairs, cull rather than repair the invalid ones, and validate every survivor with an explicit chain from item to objective to Stage 1 source.
- Cull do not fix: tweaking a cleanly written invalid item usually just produces a prettier invalid item, so generate more than you need and keep only the survivors.
- The pump-technician exam shows the stakes: a cover-the-content sweep produced recall items that certified vocabulary recognition, not calibration, and certified the wrong competence fluently.
- The iron rule is most absolute here: AI may draft the item and feedback, but a human validates the assessment and owns the pass-fail decision, because AI does not certify a learner as competent, full stop.
Skill.re