Adaptive Practice and Item-Bank Health
A learning manager pulls the certification report for the annual food-safety exam and sees a number that should not be possible: a 94% pass rate, up from 71% the year before, with no change in the training. Nothing got better. The item bank rotted. Over eighteen months, an AI tool had grown the bank from 60 questions to 400, and nobody checked the new items for health. A third of them were so easy that everyone got them right, a handful had two defensible answers, and the adaptive engine, dutifully serving questions by difficulty, had been feeding learners the easy ones. The exam still looked rigorous. It certified almost everyone, including the people who could not actually hold a safe holding temperature. The pass rate went up because the measurement broke, and a broken measurement that certifies the wrong people is worse than no exam at all.
Why an AI-Grown Item Bank Rots
AI changed the economics of the item bank the same way it changed everything else in learning: it made production nearly free. A model can draft fifty plausible questions in the time it used to take to write three. That is genuinely useful, because a healthy bank needs depth, you cannot run adaptive practice or secure certification off a dozen items everyone memorizes. But cheap production has a shadow. The cost of writing an item collapsed; the cost of an item being bad did not. And a bad item in a bank is not a neutral dud. It is an active corruption of the measurement, because every learner it touches gets a slightly wrong signal about whether they know the material.
It is worth dwelling on why this asymmetry is so counterintuitive, because the intuition is what lets the rot happen. Everyone understands that writing a bad sentence in a module is a problem you can fix by editing the sentence. So people extend that intuition to items: a bad item is just a bad question, you will catch it eventually, no harm done. But an item is not a sentence. It is a measuring instrument, and a broken instrument does not announce itself by looking broken. A thermometer that reads three degrees high looks exactly like a working thermometer; you only discover the error by comparing it against a known reference. A bad item is the same. It looks like a question. It reads fine. The only way to know it is broken is to compare how learners perform on it against what you would expect from a working item, which is precisely the comparison nobody runs when the bank is "AI-generated and reviewed for accuracy."
An item bank is the pool of assessment questions a course or certification draws from. Why you care: it is the instrument you measure competence with, and like any instrument it drifts out of calibration unless someone maintains it. When a human wrote every item slowly, the bank stayed small and was implicitly curated by the pain of authoring. When AI writes items in bulk, the bank grows fast and nobody feels the pain that used to force quality. The result is a bank that looks impressive (400 items!) and measures poorly, and the only way to know is to look at how the items actually perform when learners answer them. That looking has a name: psychometric hygiene, the routine maintenance that keeps an assessment instrument trustworthy. It is the unglamorous discipline that separates a real exam from a quiz that happens to have a passing score.
A bad item is not a neutral dud sitting in a bank. It is an active corruption of the measurement, and an AI that writes fifty items an hour can corrupt the instrument faster than anyone is checking it.
The Three Numbers That Tell You an Item Is Healthy
You do not judge an item's health by reading it. A question can be beautifully written and measure nothing. You judge it by how learners actually perform on it, using three statistics that any honest item analysis reports. None of them require a statistician to understand; each answers a plain question.
P-Value: Is This Item the Right Difficulty?
The p-value of an item is simply the proportion of learners who got it right. Despite the confusing name (it is unrelated to the p-value of a hypothesis test), it is just difficulty expressed as a fraction: a p-value of 0.95 means 95% answered correctly, so the item is very easy; a p-value of 0.30 means only 30% did, so it is hard. Why you care: an item everyone passes (p near 1.0) and an item almost everyone fails (p near 0) both tell you almost nothing, because they do not separate people who know the material from people who do not. Most useful items sit in a middle band, often cited around 0.3 to 0.9 depending on the stakes, with a security-critical certification deliberately holding some harder items. The food-safety bank rotted in part because a third of the AI-written items had p-values above 0.95: technically correct, instructionally empty.
Item Discrimination: Does This Item Separate the Strong from the Weak?
Item discrimination is the most important number and the one teams skip. It measures whether the learners who do well on the assessment overall also tend to get this particular item right. A common form is the point-biserial correlation, reported between roughly minus 1 and plus 1. A high positive discrimination (say, above 0.3) means the item behaves the way a good item should: strong learners get it, weak learners miss it, so the item helps sort competence. A discrimination near zero means the item is noise: knowing the material does not predict getting it right. And a negative discrimination is an alarm: the strong learners are getting it wrong and the weak ones right, which almost always means the item is flawed, ambiguous, or the answer key is incorrect. Why you care: an item with negative discrimination is actively punishing the people who know the most, and an AI-drafted item with a subtly miskeyed answer is a classic source of exactly this.
Distractor Analysis: Are the Wrong Answers Doing Their Job?
A multiple-choice item is only as good as its wrong answers, its distractors. Distractor analysis looks at how often each wrong option was chosen and by whom. A healthy distractor is plausible enough that some learners pick it, and it is picked more by weak learners than strong ones. A distractor nobody ever selects is dead weight: it makes the item easier than it looks, because the learner effectively chooses among fewer real options. A distractor that strong learners pick more than weak ones is a warning that the option may actually be defensible, meaning the item has two arguably correct answers. Why you care: AI loves to generate distractors that are obviously wrong (the "all of the above" filler, the joke option, the one that is a different category entirely), which inflates the apparent difficulty and quietly makes the item easier and less valid than its p-value suggests.
Retiring Weak Items Is the Job, Not an Afterthought
Here is the discipline that keeps a bank from rotting: items are not written once and trusted forever. They are monitored, and the weak ones are retired, pulled from active use, after enough learners have answered them to judge. This is the maintenance cycle AI makes more necessary, not less, because a bigger, faster-growing bank has more places for rot to hide. The cycle is simple to state and easy to neglect: generate candidate items, pilot them, read the three numbers, keep the healthy ones, fix or retire the rest, and re-check periodically because an item's difficulty can drift as the workforce and the content change.
| Signal | What it usually means | The action |
|---|---|---|
| P-value above ~0.95 | Item is too easy; everyone passes it, so it separates no one | Retire or rewrite harder, unless it is a deliberate warm-up item |
| P-value below ~0.20 | Item is too hard or confusingly worded | Review wording and key; retire if it is genuinely a trick rather than a measure |
| Discrimination near zero | Item is noise; competence does not predict getting it right | Investigate and usually retire; it adds length without measuring |
| Negative discrimination | Strong learners miss it, weak learners pass it: likely flawed or miskeyed | Pull immediately and check the answer key; a classic AI miskey symptom |
| A dead distractor (never chosen) | Option is implausible; item is easier than it looks | Replace the distractor with a plausible one |
| A distractor strong learners pick | Possible second defensible answer; validity threat | Review the item; it may have two correct answers |
Notice that AI is genuinely helpful on the left of this cycle, drafting candidate items and even proposing plausible distractors, and genuinely cannot be trusted on the right, the decision to keep or retire. That decision is a judgment about whether the instrument measures competence, and AI does not make that call. This is the same iron rule the whole program runs on, applied to the item bank: AI assists with production, the human owns the measurement.
One subtlety in the retirement decision deserves attention, because it is where judgment beats a mechanical rule. The three numbers flag a candidate for review; they do not, by themselves, decide its fate. An item with a p-value of 0.96 looks like a retire candidate, but if it is a deliberate gateway item, one you want every competent learner to pass to confirm a non-negotiable basic, its easiness is the point and you keep it as a warm-up rather than a discriminator. An item with a negative discrimination almost always indicates a real flaw, but the right move is to investigate why before pulling it, because the cause, a miskeyed answer, an ambiguous stem, a second defensible option, tells you whether to fix it or kill it. The statistics are an alarm system, not an autopilot. They tell you which items to look at; a human who understands the objective decides what to do, which is exactly why this cannot be handed to the AI. The model can compute the numbers and flag the outliers all day. It cannot weigh whether an easy item is a flaw or a deliberate floor, because that weighing depends on what the assessment is for, and the purpose of the assessment is a human decision.
This also reframes what an item bank actually is. It is tempting to think of the bank as an asset that grows in value as it grows in size, the way a content library does. It is closer to a living system that needs tending. Items age. The workforce that answers them changes. Content gets updated and an item that perfectly matched the old SOP now tests a superseded step. A bank is healthy not when it is large but when every item in it is currently earning its place in the measurement, and that health is a state you maintain, not a milestone you reach. AI makes it trivially easy to add items and does nothing to help you tend the ones already there, which is precisely the imbalance that produces a large, impressive, rotting bank.
Adaptive Practice Makes Bank Health Non-Negotiable
Adaptive practice serves each learner items matched to their estimated ability: get one right, the next is a little harder; miss one, the next is a little easier. Done well, it is efficient and motivating, the learner spends time at the edge of what they can do instead of grinding through items that are too easy or too hard. But adaptive practice has a brutal dependency: it is only as good as the difficulty labels on the items, and those labels come from the p-values. If the bank's difficulty estimates are wrong, the engine serves the wrong items with total confidence. An item the bank thinks is hard but is actually trivial will be served to advanced learners, who breeze through it and get told they have mastered something they were never tested on.
This is exactly what happened in the food-safety exam. The adaptive engine trusted the bank's difficulty metadata. The metadata was wrong because nobody had run item analysis on the 340 AI-generated additions. So the engine, behaving correctly, served plausible-looking easy items as if they were the right challenge level, and the pass rate climbed while competence did not. The engine was not broken. The instrument under it was, and an adaptive engine on a rotten bank does not fail loudly. It fails by certifying the wrong people smoothly, which is the failure that surfaces in an incident rather than a report.
There is a second-order trap worth naming, because it bites teams that do start measuring. When you first run item analysis on an AI-grown bank, the difficulty numbers you get are based on whatever mix of learners has answered so far, and early data on a new item is noisy. An item that looks like it has a p-value of 0.6 after twenty learners might settle at 0.85 after two hundred. The discipline, then, is not a one-time analysis but a cycle: pilot new items on enough learners to get a stable read before they count toward anyone's score, treat early statistics as provisional, and re-check as the sample grows. Adaptive practice intensifies this need, because the engine starts routing on the difficulty estimate immediately, so a noisy early estimate becomes a routing decision before it is trustworthy. The defensible pattern is to hold new AI-generated items in a piloting state, gathering performance data without letting them drive certification or adaptive routing, until the three numbers are stable enough to trust. Speed of generation does not buy you speed of validation; the learners still have to answer the item before you know if it works.
A Worked Example: Before and After
Watch the same food-safety bank handled two ways.
Before (grow and trust). The team uses an AI tool to expand the bank from 60 to 400 items over eighteen months to feed adaptive practice and reduce item exposure. The items read well, so they go live as they are drafted. No item analysis runs, because the bank is "AI-generated and reviewed for accuracy," which checks that each item is factually true but never checks whether it measures. Eighteen months later the pass rate is 94%, three points of managers are quietly certified who cannot hold a safe temperature, and the first sign of trouble is a failed health inspection traced back to a "certified" employee. The exam was never validated as an instrument; it was proofread as content. The two are not the same job.
After (grow and maintain). Same AI tool, same 400 items, but every new item is piloted before it counts, and item analysis runs on a schedule. The p-value report flags 130 items above 0.95 and 12 below 0.20; the easy ones are rewritten or reserved as warm-ups, the impossibly hard ones are reviewed. Discrimination flags 8 items with negative point-biserial; three turn out to be miskeyed by the AI and are corrected, five are retired. Distractor analysis kills 40 dead options and surfaces 6 items with a second defensible answer, which are rewritten. The difficulty metadata is now real, so the adaptive engine serves genuine challenge. The pass rate settles at 78%, lower than the rotten 94% and far more trustworthy, because now it means something. When an auditor asks "how do you know this certification measures competence," the answer is an item-analysis report with p-values, discrimination indices, and a retirement log, not a pass-rate chart and a hope.
The lesson is that growing a bank and maintaining a bank are different jobs, and AI only does the first. A bank that grows without psychometric hygiene does not get stronger. It rots quietly, and a rotten bank under an adaptive engine certifies the wrong people with a smile.
Key Takeaways
- AI made item production nearly free, but the cost of a bad item did not fall: a bad item is an active corruption of the measurement, not a neutral dud.
- You judge an item by how learners perform on it, not by reading it; three numbers tell the story: p-value (difficulty), item discrimination, and distractor analysis.
- P-value is the proportion who got the item right; items near 1.0 or near 0 separate no one, and most useful items sit in a middle difficulty band.
- Item discrimination is the most important and most skipped number: it measures whether strong learners get the item right, and a negative value signals a flawed or miskeyed item punishing the people who know most.
- Distractor analysis checks that wrong answers are plausible and chosen more by weak learners; dead distractors inflate apparent difficulty and a distractor strong learners pick signals a second defensible answer.
- Retiring weak items is the job, not an afterthought: generate, pilot, read the numbers, keep the healthy, retire the rest, and re-check as difficulty drifts.
- AI is trustworthy on production (drafting items and distractors) and not trustworthy on the keep-or-retire decision, which is a human judgment about whether the instrument measures competence.
- Adaptive practice depends entirely on correct difficulty labels, so an unmaintained AI-grown bank fails silently by certifying the wrong people smoothly, and the proof of a valid certification is an item-analysis report, not a pass-rate chart.
Skill.re