Recognizing Bad AI Output in a Clinical Setting
A hospital pharmacist verifying overnight orders ran a complex regimen through an AI assistant to double-check for interactions. The tool returned a clean, well-organized summary and, near the bottom, flagged a serious interaction between two of the patient's medications, complete with a mechanism and a recommendation to separate the doses. It read like something out of a reference. The pharmacist almost acted on it, then paused, because one thing nagged: the mechanism it described was the mechanism for a different, similarly named drug, not the one the patient was actually on. She opened the interaction reference herself. The real drug had no such interaction. The AI had blended two drugs that sounded alike into a single confident warning, a fabricated interaction wearing the costume of a real one. Acting on it would have meant separating doses unnecessarily, or worse, second-guessing a safe regimen on the strength of an invention. What caught it was not a tool and not a rule in a manual; it was a trained skeptic's reflex, a pharmacist who knew the tells of bad AI output and ran the check before the output reached the patient. This lesson builds that reflex. Level 1 named the five forms a clinical hallucination takes and made you a skeptical reader; the last two lessons taught you to prompt and ground for accurate output. This one assumes bad output will still appear, because it will, and teaches the skeptic's checklist for catching it: the tells, the red flags, and the specific moves that stop a fabrication before it becomes a patient-safety event.
Why Bad Output Still Appears, Even with Good Prompting
Grounding and careful prompting reduce bad output; they do not eliminate it, and a pharmacist who believes good technique makes verification optional has misunderstood the tool. The reason is structural, and it is the same reason from Level 1: a generative model produces fluent, probable-sounding text whether or not that text is true, and it produces the true and the false in the identical confident register. Even a well-grounded model can misread the source it was given, blend two similar entries, or assert a claim that goes slightly beyond what the source supports. Even a tool with retrieval can retrieve the wrong passage or stitch a real passage to an invented detail. The fluency is constant; the truth is variable; and the model gives you no reliable internal signal of which is which. That is why verification is not a fallback for when prompting fails, it is a permanent control, and why recognizing bad output is a core clinical skill rather than a beginner's crutch you outgrow.
It helps to hold the failure forms from Level 1 in mind, because recognizing bad output means knowing what bad output looks like. The five clinical forms a hallucination takes are: the fabricated clinical fact (a failed therapy or diagnosis the record does not support), the wrong dose (a number that is plausible but incorrect, often a renal or pediatric adjustment), the fabricated or mismatched interaction (a warning for a drug the patient is not on, or a blend of two drugs), the invented coverage criterion (a payer rule the payer never published), and the fabricated citation (a reference, guideline, or study that does not exist or does not say what the model claims). Each of these arrives inside clean, confident, professional-looking output, which is exactly why none of them announces itself. The skeptic's job is to know these shapes well enough to feel the prickle of suspicion when one appears, and then to run the check that confirms or kills it.
Good prompting reduces bad output; it never removes the need to verify. The model produces truth and fabrication in the same confident register, so verification is a permanent clinical control, not a fallback. Recognizing bad output is the skill that catches what grounding misses.
The Tells: What Bad Output Feels Like
Experienced verifiers develop a feel for output that is "off," and that feel can be named and taught so you do not have to acquire it the slow way. The tells are not proof of error; they are triggers for a harder look, and learning them turns vague unease into a specific check. The first tell is suspicious convenience: an answer that is exactly what you hoped for, the justification that perfectly satisfies the criterion, the interaction check that comes back conveniently clean, the dose that happens to match the order. Real clinical pictures are messy; output that is too tidy has often been smoothed by the model filling a gap with a plausible invention. The second tell is specificity without a source: a precise claim, a named guideline, an exact criterion, a particular dose, asserted with no traceable origin. Precision is not accuracy; a fabricated detail is often more specific than a real one, because the model generates the kind of specificity that sounds authoritative.
The third tell is a uniform confident tone across a claim that should carry uncertainty. When a model states a borderline renal dose or a contested interaction with the same flat confidence it uses for a textbook fact, the absence of hedging where hedging belongs is itself a warning, because a careful human source would qualify what is genuinely uncertain. The fourth tell is the near-miss: a claim that is right for a neighbor, the right interaction for a similarly named drug, the right criterion for a different plan, the right dose for the adult when the patient is a child. This is the most dangerous tell because it is the hardest to catch, the answer is real, just attached to the wrong patient, drug, or plan. The fifth tell is drift across a long answer: an output that starts grounded and accurate and then, as it continues, wanders into claims the source never supported, because the further the model generates from its grounding, the more it falls back on the plausible average. When you feel any of these tells, you do not yet know the output is wrong, but you know exactly where to point the check.
The Skeptic's Checklist for Clinical Output
The tells trigger suspicion; the checklist resolves it. The point of a checklist is that it does not depend on suspicion, you run it on every clinically load-bearing output, suspicious or not, because the most dangerous fabrications are the ones that did not trip a tell. The checklist is organized around the load-bearing facts, the ones where an error reaches a patient, and it is short enough to run fast.
Doses and Adjustments
For any dose the model states, confirm three things against an authoritative source: the drug, the dose, and the adjustment for this patient's specifics. The most common dose failure is the adjustment, a model that gives the standard adult dose when the patient's renal function, age, or weight demands a change, or that gives a plausible adjusted dose that is simply wrong. Trace every dose to the package insert or an authoritative dosing reference, and confirm the adjustment against this patient's actual labs and parameters. Never accept a dose because it appears in a clean, confident output; the confidence is constant whether the dose is right or invented.
Interactions
For any interaction the model flags or clears, confirm it is real, confirm it applies to the drugs this patient is actually on, and confirm the model did not miss one. The fabricated-interaction failure has two faces: the invented warning (an interaction that does not exist, or exists for a similarly named drug) and the false clear (a clean result that missed a real interaction). Verify flagged interactions against an interaction reference, and never treat a clean interaction check as authoritative, because a missed interaction is as dangerous as an invented one and harder to notice. The drug-name near-miss is especially common here, so confirm the interaction is for the exact drug, not a soundalike.
Coverage Criteria
For any payer criterion the model cites in a PA or appeal, open the actual current payer policy and confirm the criterion is stated as the model claims and that this patient's documented history genuinely satisfies it. The invented-criterion failure produces a submission built on a rule the payer never published, which causes an avoidable denial and a delayed patient. Quote-match the criterion to the published policy, and match the patient's documented facts to it line by line, refusing any justification that relies on a fact the chart does not support.
Citations and Guidelines
For any reference, guideline, or study the model cites, confirm it exists and says what the model claims it says. Fabricated citations are common and dangerous because they borrow the authority of a real literature. Do not accept a citation you have not opened; a reference that cannot be found or that does not contain the claimed statement is a fabrication, and any conclusion resting on it is unsupported. A model that cites confidently is not a model that cites correctly.
The Moves That Catch Fabrication and Drift
Beyond the checklist of what to verify, there are specific moves that actively flush out bad output, techniques that make a fabrication reveal itself. The first move is ask for the source and watch what happens. A grounded claim has a source the model can produce on request; an invented one does not. When you ask "show me the chart line that documents this failed therapy" or "quote the exact policy sentence for this criterion," a real claim yields the source and a fabricated one yields a dodge, a restatement, a new and different answer, or a quiet retraction. The model's inability to produce the source when pressed is often the cleanest signal that the claim was never grounded, and it costs you one prompt to find out.
The second move is cross-check the load-bearing fact against the independent truth, not against the model. The model is not a witness to its own accuracy; confirming a claim by asking the model again just gets you a second confident assertion, possibly the same fabrication restated. The verification must go to the actual source: the chart, the package insert, the payer policy, the interaction reference. This is the deepest discipline of the lesson and the one most often skipped under time pressure, the temptation is to let the fluent output stand because checking feels redundant against something that reads so authoritatively. It is not redundant; it is the entire safety control. The third move is re-read the long answer for drift: the end of a lengthy output is where fabrication concentrates, so give the final claims of any long response extra scrutiny, since that is where the model most likely wandered off its grounding. The fourth move is verify the negative: a clean interaction check, a "no contraindication," a "criterion met" deserve the same skepticism as a positive claim, because a comforting absence can hide a missed real signal, and the false clear is uniquely dangerous precisely because it produces no alarm.
One move deserves singling out because it is the cardinal rule in action: the AI flag is a prompt to think, never a verdict to act on. When the model surfaces a signal, an interaction, a renal concern, an off-formulary status, the correct response is to investigate the signal against the source, not to act on the model's word. This cuts both ways. A flag is a reason to look, not a reason to act; and a clean result is a reason to confirm, not a reason to relax. The pharmacist who treats AI output as decision support, input to a human judgment that verifies it, catches what the model gets wrong. The pharmacist who treats AI output as a decision, a verdict to rubber-stamp, ships the model's errors to patients under a professional credential. The entire difference between safe and unsafe AI-assisted practice lives in that distinction.
Building the Skeptic's Reflex Into Daily Practice
Recognizing bad output cannot be an occasional effort summoned when you happen to feel uneasy; it has to be a reflex that runs on every clinically load-bearing AI output, because the most dangerous fabrications are the ones that never tripped your unease. The way you build the reflex is the way you build any clinical habit: you make the check non-negotiable, you make it fast, and you make it specific. Non-negotiable means the rule is "every dose, every interaction, every criterion, every citation the AI touches gets traced to its source before it informs a clinical decision," with no exception for output that looks clean, because looking clean is what dangerous output does. Fast means the check is targeted, you are not re-deriving the answer, you are confirming the load-bearing claim against the source the model should have cited, which takes seconds when the output is grounded and is itself a red flag when it is not. Specific means you know the failure forms and the tells well enough to aim the check, dose to the package insert, criterion to the policy, interaction to the reference, citation to the literature.
There is a cultural dimension to this that matters in a real pharmacy. The pressure that erodes verification is not laziness; it is speed, the queue is long, the output is fluent, the afternoon is slipping, and the fabrication that reads perfectly is precisely the one your tired judgment most wants to wave through. A pharmacy that wants safe AI-assisted practice builds verification into the workflow so it does not depend on heroic individual diligence at 4 p.m.: standing instructions that require sources, output formats that surface citations, and a shared norm that a clean check is confirmed, not assumed. The asymmetry that anchors this whole program lands here with full force. The seconds you spend confirming a dose, a criterion, or an interaction against the source are small; the harm a single confident fabrication can do, a wrong renal dose dispensed, a real interaction missed, a fabricated criterion submitted under your credential, is not an efficiency miss but a patient-safety event. Recognizing bad output is the skill that stands between the model's confident error and the patient, and a pharmacist who has made it a reflex has become the kind of AI-assisted practitioner the whole program is built to produce: fast because the tool is fast, and safe because the skeptic never goes off duty.
Key Takeaways
- Good prompting and grounding reduce bad output but never eliminate the need to verify, because a generative model produces truth and fabrication in the identical confident register; verification is a permanent clinical control, not a fallback for failed prompting.
- Bad clinical output takes five forms from Level 1: the fabricated clinical fact, the wrong dose (often a renal or pediatric adjustment), the fabricated or mismatched interaction, the invented coverage criterion, and the fabricated citation, each arriving inside clean, confident output that never announces itself.
- The tells that should trigger a harder look are suspicious convenience (an answer too tidy to be real), specificity without a source (precision is not accuracy), a uniform confident tone where uncertainty belongs, the near-miss (right for a soundalike drug, a different plan, or the wrong patient), and drift across a long answer.
- Run the skeptic's checklist on every load-bearing output regardless of suspicion: trace every dose and its adjustment to the package insert and the patient's labs, verify every flagged or cleared interaction against an interaction reference, quote-match every payer criterion to the current policy, and open every citation to confirm it exists and says what is claimed.
- The moves that catch fabrication: ask for the source and watch whether the model produces it or dodges; cross-check the fact against the independent source, never against the model itself; re-read long answers for end-of-output drift; and verify the negative, because a false clear is as dangerous as an invented warning and triggers no alarm.
- The cardinal rule in action: an AI flag is a prompt to think, never a verdict to act on, and a clean result is a reason to confirm, not to relax; treating AI output as decision support catches the model's errors, treating it as a decision ships them to patients.
- Build the reflex by making verification non-negotiable (no exception for clean-looking output), fast (a targeted confirmation against the cited source), and specific (aim each check at the right source for the failure form).
- The eroding pressure is speed, not laziness, so pharmacies should build verification into the workflow rather than rely on individual diligence; the seconds spent confirming a dose, criterion, or interaction are small against the patient-safety event a single confident fabrication can cause.
Skill.re