AI for Pharmacy
Aware · M2 · lesson 2 of 19 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Hallucinations in a Clinical Context
📖
now learning

AI Hallucinations in a Clinical Context

15 min

A specialty pharmacist named Tomás was working an appeal for a patient whose biologic had just been denied. The AI assistant drafted a tight, persuasive appeal letter in under a minute, and one sentence in it was a gift: "Per the patient's chart, the trial of adalimumab from January produced an inadequate response with documented disease progression." It was exactly the clinical argument the appeal needed. Tomás almost sent it. Then he did the thing this lesson is about: he opened the chart and looked for the January adalimumab trial. There was no adalimumab trial. The patient had been on a different agent entirely, and the documented response had been partial, not a clear failure. The AI had not retrieved a fact; it had generated the fact the letter needed, the clinically perfect sentence, sourced from nowhere. Had Tomás sent it, he would have submitted a fabricated treatment history to a payer in a legal document, on behalf of a patient, under his professional credential. The previous lessons established that hallucinations happen and that in pharmacy they are patient-safety events. This lesson gets specific and tactical: it walks through the exact clinical forms a hallucination takes, shows you what each one looks like in the wild, and gives you the precise check that catches it. Vague awareness that "AI can be wrong" does not protect a patient. Knowing the four or five specific ways it goes wrong, and exactly where to look, does.

The Invented Clinical Fact

The first and most dangerous form is the one that caught Tomás: the model generates a specific clinical fact about the patient that is not in the record. A failed therapy that never happened. A diagnosis the patient does not carry. A lab value, a date, a symptom, a response to treatment, all fabricated to fit the shape of the document the model is writing. This form is uniquely seductive because the fabricated fact is usually exactly the fact you wanted to be true. The appeal needed a documented first-line failure, so the model produced one. The justification needed a qualifying diagnosis, so the model supplied it. The model is not malicious and it is not guessing randomly; it is generating the most plausible content for the document type, and in clinical documents the most plausible content is precisely the clinically convenient fact. That is what makes this form so hazardous: it does not look like an error, it looks like good news, and good news lowers your guard at the exact moment you need it raised.

The check is a hard rule with no exceptions: every patient-specific clinical fact in an AI-generated document must be traced to the chart before the document is used. Not "does it sound right," not "is it plausible," but "can I point to where in this patient's record this specific fact appears." If the letter says the patient failed adalimumab in January, you open the record and find the adalimumab trial and the documented inadequate response, or the sentence does not go out. A patient-specific fact that cannot be located in the record is a fabrication until proven otherwise, regardless of how clinically perfect it is. This single check, applied without exception to every patient claim, is the highest-value verification habit in clinical AI use, because the invented clinical fact is both the most common serious hallucination and the most consequential, especially in prior authorization, where it becomes a misrepresentation to a payer in the patient's name.

Every patient-specific clinical fact in an AI-generated document is a fabrication until you have traced it to a specific place in that patient's record. The clinically perfect fact is the one to distrust most.

The Wrong Number: Doses, Thresholds, and Adjustments

The second form is the wrong quantitative fact: a dose, a maximum, a frequency, a renal or hepatic adjustment threshold, a conversion between two drugs or two routes. This form is dangerous in a different way from the invented fact, because the number is often not patient-specific and so cannot be caught by tracing to the chart; it has to be checked against an authoritative drug-information source. And it is dangerous because of the precision problem covered earlier: the model tends to produce a number that is in the right neighborhood, which means the errors that survive are the plausible near-misses, a maximum that is real for a related drug, an adjustment that applies to a different degree of renal impairment, a conversion ratio that is close but not correct. A wildly wrong number triggers suspicion; a number that is wrong by a clinically meaningful but superficially plausible amount sails through.

The check is to treat every quantitative clinical value the model produces as unverified until confirmed against an authoritative reference, and to be most careful precisely when the number looks reasonable. This is counterintuitive, because a reasonable-looking number invites trust, but the reasonable-looking wrong number is exactly the one your reference check exists to catch. The habit that protects patients here is to refuse to let a dose or a threshold reach a clinical decision on the model's authority alone, no matter how confident or how plausible. The model can be a useful prompt, "you may want to consider a renal adjustment here", but the specific number always comes from a source you trust, not from the generator. A particularly important sub-case is the unit and the route: models can produce the right number with the wrong unit, or a dose correct for one route stated for another, and these errors are easy to skim past because the number itself looks fine. The check covers the whole quantitative statement, value, unit, route, and frequency, not just the digits.

The Interaction That Isn't, and the One That Is

The third form concerns drug interactions, contraindications, and warnings, and it is uniquely treacherous because it fails in two opposite directions, each with its own harm. The false positive is the model asserting an interaction, contraindication, or warning that is not real. The harm is a clinically unnecessary action: a therapy changed, a drug withheld, a prescriber called and a patient alarmed over a danger that does not exist, and every unnecessary change introduces its own risk. The false negative is the model failing to surface a real interaction or contraindication, and this is the more insidious of the two, because it is an error of absence. A fabricated warning is at least visible; you can check it and find it false. A missing warning is silence, and silence does not announce itself. The model that should have flagged a dangerous interaction and did not has produced an output that looks complete and reassuring precisely because the danger is not mentioned.

This two-directional failure has a specific implication that pharmacists must internalize: an AI interaction screen can never be treated as the interaction screen. The presence of a flag must be verified (is this interaction real and clinically significant for this patient?), and, far more importantly, the absence of a flag must never be read as the absence of a problem. The model's silence is not a clean bill of health. The clinical interaction check, against a trusted interaction resource and the pharmacist's own knowledge, remains the actual safety control; the AI is at most a prompt that may surface something to look into, never the authority that confirms there is nothing to find. A pharmacy that lets an AI interaction screen replace, rather than supplement, the real check has converted a tool that should add a layer of safety into one that quietly removes a layer, because the staff now trust a silence the model was never qualified to give.

The Fabricated Source and the Confident Citation

The fourth form is the fabricated reference, and it deserves its own treatment because it is so convincing and so common. Asked to support a clinical claim, the model produces a citation, a guideline name, a study, a journal article, complete with plausible authors, a year, and a title that sounds exactly like a real paper. The shape is perfect because the model has seen thousands of real citations and generating the form of one is precisely what it is good at. The existence is fictional. In a clinical or prior-authorization context, this matters enormously: a justification that cites a guideline recommendation is only as good as the guideline actually saying what is claimed, and a fabricated or misremembered citation can assert a level of evidence that does not exist. The danger is amplified because a citation feels like the end of verification, you see a source and relax, when in fact a model-produced citation is the beginning of verification.

The check is to treat every citation the model produces as a claim to confirm, not a source to trust. The guideline must be located and confirmed to say what the document claims; the study must be confirmed to exist and to support the assertion. This is true even when the citation is to something real, because a second failure mode is the real-source-misrepresented: the model cites an actual guideline but states a recommendation the guideline does not actually make, or applies it to a population it does not cover. The discipline is the same as for every other form: the load-bearing element, here the cited evidence, is confirmed against the actual source before it is allowed to carry weight in a clinical document. A pharmacist who has internalized that a confident citation is the least, not the most, trustworthy-feeling part of an AI output is a pharmacist who will not be fooled by the form of authority into skipping the substance of verification.

The Summary That Quietly Drifts

A fifth form is subtler than the others and easy to miss precisely because nothing in it looks fabricated: the summary or simplification that drifts from the source. Ask a model to condense a long discharge summary, summarize a patient's medication history, or simplify a complex warning for a patient, and it will produce something fluent and reasonable that has quietly altered the substance. A nuance is dropped. A "consider" becomes a "should." A qualified finding becomes a definite one. A list of five medications becomes a list of four because one was on a page the model weighted less. Nothing here is a dramatic invented fact; the summary is mostly right, which is exactly what makes the drift dangerous. You read it, it matches your general sense of the patient, and you accept it, never noticing that the one altered detail is the one that mattered.

This form is common in two high-value pharmacy tasks: orienting yourself to a long record, and producing patient-facing education. In the first, the risk is that you make a clinical decision based on a summary that subtly misrepresents the record, so the check is that any fact from the summary that will drive a decision gets confirmed against the source, with the summary treated as a finding aid, not the finding. In the second, patient education, the risk is that a simplified explanation drops or softens a real warning, turning safe-but-complex language into friendly-but-incomplete language, so the check is to fact-check the simplified version against the full, true content before the patient hears it. The unifying principle is that summarization is not a neutral, lossless operation when a generative model does it; it is another act of generation, with its own opportunity to drift, and the load-bearing facts inside a summary need the same confirmation as facts anywhere else.

Building the Catching Reflex

The point of naming these forms precisely is to convert verification from a vague intention into a specific, fast, repeatable scan that you run on any AI-touched clinical output. The scan has a shape: for each clinical statement, identify which form of hallucination could be hiding in it, and apply that form's check. Is it a patient-specific fact? Trace it to the chart. Is it a number, a dose, a threshold, a conversion? Confirm it against an authoritative reference, including unit and route. Is it an interaction, contraindication, or warning, present or absent? Run the real clinical check and never trust the model's silence. Is it a citation? Confirm the source exists and says what is claimed. Is it a summary or simplification? Treat it as a finding aid and confirm any decision-driving fact against the source. Five questions, five checks, run against the specific facts that matter, and the great majority of dangerous clinical hallucinations are caught before they reach a patient.

There is a reason to run the scan in this explicit, almost mechanical way rather than relying on a general feeling of carefulness. Under time pressure, a vague intention to "be careful with AI" collapses into doing nothing different, because carefulness without a procedure has no handles to grab. A named scan has handles. You can teach it, you can checklist it, you can audit whether it happened, and you can do it the same way at 2 a.m. on your fortieth order as you did fresh at 9 a.m. on your first. This is also exactly what the later governance and URAC lessons will ask a pharmacy to demonstrate: not that staff are generally thoughtful, but that a specific, repeatable verification procedure is applied to AI-touched clinical work. The five-form scan is the front-line version of that procedure, the thing an individual pharmacist actually does, and building it into a reflex now is what makes the organizational version credible later.

What makes this reflex powerful is that it is targeted rather than exhausting. You are not re-deriving the entire document from scratch; you are aiming verification precisely at the load-bearing clinical facts where a hallucination would do harm, and letting the model have the parts that are genuinely just drafting, the structure, the phrasing, the formatting, where a hallucination is harmless. Tomás did not rewrite the appeal letter; he checked one sentence, the patient-specific clinical claim, found it fabricated, and saved himself and his patient from a serious error in about ninety seconds. That is the whole skill in miniature: know the forms, find the load-bearing facts, apply the right check, and let the tool do the rest. A pharmacist who runs this scan as a reflex gets the full speed of AI on the drafting while giving up none of the safety on the substance, which is exactly the balance this entire program is built to produce.

Key Takeaways

  • Clinical hallucinations take specific, recognizable forms, and knowing them precisely, not just "AI can be wrong," is what actually protects a patient.
  • The invented clinical fact (a fabricated failed therapy, diagnosis, lab, or response) is the most dangerous form because the fabrication is usually the clinically convenient fact you wanted, which lowers your guard; the check is to trace every patient-specific fact to the chart, treating any untraceable fact as a fabrication.
  • The wrong number (dose, threshold, adjustment, conversion) survives as a plausible near-miss; confirm every quantitative value, including unit and route, against an authoritative reference, and be most careful when the number looks reasonable.
  • Interaction and contraindication errors fail in two directions: the false positive (an asserted danger that is not real) and the more insidious false negative (a real danger left unmentioned); the model's silence is never a clean bill of health, so the real clinical check always stands.
  • The fabricated citation is convincing because the model excels at producing the form of a real reference; treat every citation as a claim to confirm against the actual source, and watch for real sources that are misrepresented.
  • The drifting summary is the subtlest form: summarization by a generative model is another act of generation, not a neutral lossless operation, so a condensed record or simplified warning can quietly alter the one detail that mattered; confirm any decision-driving fact against the source and fact-check simplified patient education against the full content.
  • The catching reflex is a fast scan: for each clinical statement, identify which of the five forms could be hiding in it and apply that form's check (trace to chart, confirm against reference, run the real interaction check, confirm the source, verify the summary against the record).
  • The reflex is targeted, not exhausting: aim verification at the load-bearing clinical facts and let the model have the drafting, which delivers the full speed of AI with none of the safety surrendered.