Getting Accurate Clinical Output
A specialty pharmacist was assembling a prior authorization (PA) for a patient on a biologic and asked the AI tool a reasonable-sounding question: "What is the step-therapy requirement for this drug under this plan?" The model answered immediately, with the calm authority these tools always carry: the plan required a documented inadequate response to two conventional agents over a minimum of twelve weeks. It was specific, it was plausible, and it was wrong. The plan's published policy required one conventional agent over eight weeks, not two over twelve. The AI had not consulted the plan; it had produced the most statistically average step-therapy requirement it had ever seen, dressed in the confident tone it uses whether it is right or inventing. Had she submitted on that basis, the patient's documented eight-week trial would have looked insufficient against a criterion the plan never had, and the request would have bounced. What saved her was not luck. It was a habit: she had asked the wrong way, gotten a fluent answer, and then refused to trust it until she had forced the model onto the actual policy and onto the actual chart. This lesson is about that forcing. Level 1 taught you to read AI output skeptically and named the five forms a clinical hallucination takes. The previous lesson taught you to prompt with structure. This lesson takes the next step: how to make the model produce output you can actually check, by grounding it in the formulary and the record, demanding sources and uncertainty, and refining iteratively until the answer is specific, traceable, and safe to verify.
Generic Output Is the Default Failure
The pharmacist's wrong step-therapy answer was not a freak event; it was the model behaving exactly as a generative model behaves. Recall the machinery from Level 1: the model produces the most probable continuation of the text it is given, drawn from everything it has absorbed. When you ask it a clinical question without anchoring it to a specific source, the most probable continuation is the statistical average of all the similar things it has seen, a plausible blend, not a fact retrieved from a place you can point to. For a question about step therapy, the average is some common-looking requirement that fits the general pattern of biologic policies, which is precisely why it sounded right and was nonetheless wrong for this plan. The default behavior of an ungrounded model is to give you the average answer with the confidence of a specific one, and in a clinical context the average answer is not good enough, because the patient in front of you is not the average patient and the plan in front of you is not the average plan.
This is the single most important reframe of the lesson: an ungrounded clinical answer is a hazard not because it is always wrong but because it is plausibly average, and average is undetectable until you check it against the specific truth. A wildly wrong answer is easy to catch; the dangerous answer is the one that is off by exactly enough to matter, a dose that is close, a criterion that is almost right, an interaction that is real for a related drug but not this one. The way you defend against the plausible-average failure is to stop asking the model to recall and start forcing it to consult, to ground every clinically load-bearing answer in the actual formulary, the actual payer policy, and the actual patient record, so that the output is not a blend of everything but a specific claim about a specific source you can open and confirm.
An ungrounded clinical answer is the statistical average dressed as a specific fact. The danger is not wild error; it is plausible error, off by just enough to matter. The defense is to force the model onto the actual source and the actual record, every time.
Grounding the Model onto the Source of Truth
Grounding means making the model's answer depend on a specific document you control rather than on its diffuse memory of everything. In practice, this takes two forms, and a skilled pharmacist uses both. The first is to bring the source into the conversation: paste the relevant formulary entry, the payer's published criteria, or the chart excerpt directly into the prompt, and instruct the model to answer only from that text. Instead of "What is the step-therapy requirement?", the grounded version is "Here is the plan's published policy section. State the step-therapy requirement exactly as written here, and quote the sentence it comes from." Now the model is not recalling the average requirement; it is reading the one in front of it, and it can show you the sentence so you can confirm. The same move works for a renal dose ("Here is the package insert's renal dosing table; state the adjustment for a creatinine clearance of 35"), an interaction ("Here are this patient's active medications; flag interactions only among these drugs"), and a coverage criterion ("Quote the exact criterion from this policy that applies").
The second form of grounding is structural, built into the tools your pharmacy adopts: retrieval-augmented generation (RAG), where the system automatically retrieves the relevant formulary or policy passage from an authoritative database and feeds it to the model before it answers. You will meet RAG in depth in Level 3; for now the point is that a well-built clinical AI tool grounds its answers on a current, authoritative source rather than on the model's training memory, and that the difference between a tool that does this and one that does not is the difference between an answer you can trust to verify and an answer that is a confident guess. When you evaluate or use any clinical AI tool, the first question is always the grounding question: where does this answer come from, and can I see the source it came from? If the answer is "the model's general knowledge," you are back to the plausible average, and you must supply the grounding yourself by bringing the source into the prompt.
Ground on the Record, Not Just the Rule
Grounding the model onto the payer's policy fixes half the problem; the other half is grounding it onto the patient's actual record. A justification can cite the correct criterion and still fail if it asserts a clinical fact the chart does not support, a failed therapy that did not happen, a diagnosis the patient does not carry, a lab value off by a digit. So the discipline is two-sided: force the model onto the rule (the formulary, the policy, the package insert) and force it onto the record (this patient's documented history, labs, and therapies). The prompt that does both might read: "Using only the attached chart excerpt and the attached policy section, draft a justification that matches this patient's documented history to the policy's criterion. Cite the chart line for every clinical fact and quote the policy criterion. If the chart does not contain a fact the criterion requires, say so explicitly rather than supplying it." That last instruction is the safety valve: it tells the model to surface a gap rather than paper over it with an invented fact, which is exactly the fabrication you most need to catch.
Demand Sources and Honest Uncertainty
A model that grounds its answer is more trustworthy; a model that also shows its source and flags its uncertainty is far easier to verify, because it tells you where to look and where to look hardest. So the second pillar of accurate clinical output, after grounding, is demanding that the model expose two things: the source of each claim and its own confidence in each claim. Demanding sources is straightforward and transformative: instruct the model to attach, to every clinical assertion, the specific place it came from, the chart line, the policy sentence, the table row in the package insert. This does two things. It forces the model toward groundedness, because a model required to cite is less free to invent. And it converts verification from a hunt into a check: instead of re-deriving the whole answer, you open the cited source and confirm the claim matches it, which is fast precisely because the model told you where to go.
Demanding uncertainty is the subtler and equally important half. A generative model will, by default, answer everything in the same confident register, the right answers and the invented ones in identical tones, which is the property that makes hallucinations dangerous. You can partly counter this by instructing the model to flag what it is unsure of: "Mark any part of this answer you are not confident about, and say what additional information would resolve it." A model so prompted will often, though not always, surface its own weak points, the criterion it is guessing at, the dose it is unsure applies, the interaction it flagged on thin evidence. This is not a guarantee, a model can be confidently wrong and can also fail to flag a real uncertainty, but it meaningfully improves the raw material, because a flagged uncertainty is a signpost telling you exactly which claim to verify first. The skilled operator treats every unflagged claim as still requiring verification and every flagged one as requiring it urgently.
There is a clinical subtlety here worth stating plainly: a model that refuses or hedges when it lacks grounding is doing the right thing, and you should reward that behavior rather than push past it. If you ask for a renal dose and the model says, "I do not have this drug's renal dosing in the provided sources; please supply the package insert," that is the model declining to invent, which is exactly what you want. The failure mode is the opposite: a model that, lacking the source, supplies a plausible dose anyway. So when you build prompts, build them to make refusal acceptable: tell the model that "I cannot determine this from the provided sources" is a valid and preferred answer when the grounding is absent. A pharmacy that prompts for honest refusal gets fewer confident fabrications, which is the whole game.
Iterative Refinement to a Checkable Answer
Accurate clinical output is rarely produced by a single perfect prompt; it is produced by a short conversation in which you refine the model toward an answer you can check. The first response is a draft, not a verdict, and the skilled operator reads it not as the answer but as raw material to sharpen. The refinement loop has a rhythm. You ground and ask. You read the response and find the parts that are vague, unsourced, or suspicious. You push back specifically: "You cited the criterion but did not quote it, quote the exact sentence." "This dose is not in the table I gave you, where did it come from?" "You asserted a prior adalimumab failure, show me the chart line that documents it." Each push narrows the output toward something traceable, and a model that cannot produce the source when pressed has just told you the claim was ungrounded, which is itself a valuable finding.
The goal of the loop is not a longer answer; it is a checkable answer, one specific enough and sourced enough that verification becomes a fast confirmation rather than a slow re-derivation. There is a useful test you can apply to any clinical output before you accept it: can I trace every load-bearing claim to a specific source in seconds? If yes, the output is checkable and you proceed to verify it. If no, the output is not yet done, and you refine until it is, or you discard it. This reframes iteration as a quality gate rather than a chore: you are not refining to make the model happy, you are refining until the output passes the checkable test, because output that fails that test cannot be safely verified and therefore cannot be safely used. A pharmacist who internalizes this stops accepting fluent first drafts and starts driving the conversation to a specific, sourced, refusable, checkable answer, which is the entire skill.
A Worked Refinement
Return to the step-therapy question and watch the loop run correctly. The pharmacist pastes the plan's published policy section and asks: "State the step-therapy requirement for this drug from this policy text, and quote the exact sentence." The model answers: one conventional agent, eight weeks, and quotes the sentence. She opens the policy and confirms the quoted sentence is real, the first verification done in seconds because the model pointed her to it. Next she pastes the chart excerpt: "Using only this chart, confirm whether this patient's documented history satisfies that requirement, and cite the chart line for the trial and its duration. If the duration is not documented, say so." The model finds a documented eight-week trial of a conventional agent and cites the line; it also notes that the response outcome is recorded but the exact start date is ambiguous, flagging its own uncertainty. She checks the cited line, confirms the trial, and resolves the date ambiguity by opening the dispensing record. Three short turns produced a justification grounded on the real policy and the real chart, with every clinical claim traced to a source and the one soft spot surfaced for her to close. That is accurate clinical output: not a magic first answer, but a refined, grounded, sourced, checkable one.
Building the Grounding into a Reusable Habit
Doing this once is a skill; doing it the same way every time is a discipline, and clinical safety depends on the discipline. The way to make grounding reliable is to convert these moves into a standing set of instructions you apply to every clinical prompt, so you are not reinventing the technique under time pressure on a busy afternoon. The standing instructions are short and they encode this whole lesson: answer only from the sources I provide; quote or cite the exact source for every clinical claim; if a required fact is not in the sources, say so rather than supplying it; flag anything you are uncertain about and say what would resolve it; and treat "I cannot determine this from the provided sources" as a valid answer. A pharmacist who opens every clinical conversation with those instructions, or who works in a tool that has them built in, has front-loaded the safety into the prompt, so the default output is grounded, sourced, and refusable rather than fluent and unverifiable.
This habit pays off most exactly where the stakes are highest. On a routine question, a generic answer that you then look up costs you a little time. On a specialty PA for a ten-thousand-dollar-a-month therapy, on a renal dose for a patient with failing kidneys, on an interaction check for a patient on eight medications, the grounded-and-sourced answer is the difference between a fast, defensible decision and a confident error that reaches a patient. The asymmetry that runs through this entire program lives here too: the time you spend forcing the model onto the source is small, and the harm you prevent by doing so, an avoidable denial that delays therapy, a wrong dose, a missed interaction, is large. Getting accurate clinical output is not a separate skill from patient safety; it is patient safety, practiced at the keyboard, one grounded prompt at a time. The next lesson turns to the other side of the same coin: even with good grounding, bad output will sometimes appear, and you need a skeptic's checklist to catch it before it reaches a patient.
Key Takeaways
- An ungrounded clinical answer is the statistical average of everything the model has seen, dressed in the confident tone of a specific fact; the danger is not wild error but plausible error, off by just enough to matter, which is undetectable until you check it against the specific source.
- Grounding means making the answer depend on a specific document you control: bring the formulary entry, payer policy, package-insert table, or chart excerpt into the prompt and instruct the model to answer only from that text and quote the source.
- Ground on both the rule and the record: force the model onto the formulary or policy (so the cited criterion is real) and onto the patient's chart (so every clinical fact is supported), and tell it to flag a missing required fact rather than invent one.
- Retrieval-augmented generation (RAG) is the structural version of grounding, where a well-built tool retrieves the authoritative passage before answering; when using any clinical AI, the first question is always where the answer came from and whether you can see the source.
- Demand sources (cite the chart line, the policy sentence, the table row) to force groundedness and turn verification from a re-derivation into a fast confirmation; demand uncertainty (flag what you are unsure of) to get signposts to the claims that need verifying first.
- Reward refusal: a model that says "I cannot determine this from the provided sources" is doing the right thing; build prompts that make honest refusal a valid, preferred answer so you get fewer confident fabrications.
- Accurate output comes from iterative refinement, not a perfect first prompt; push back specifically on vague, unsourced, or suspicious claims until the output passes the checkable test (every load-bearing claim traceable to a source in seconds), and discard or refine anything that fails it.
- Convert these moves into standing instructions applied to every clinical prompt, so grounding, sourcing, and refusal are the default under time pressure; the small time spent forcing the model onto the source prevents the large harm of an avoidable denial, a wrong dose, or a missed interaction.
Skill.re