Grounding AI on Local Protocols and Formularies
A resident in the emergency department, working up a patient with suspected community-acquired pneumonia, asks a general-purpose model: "What is the first-line antibiotic for community-acquired pneumonia?" The answer comes back clean and confident: a specific drug, a specific dose, a tidy sentence about coverage. It reads like a guideline. The resident, reassured, is about to enter the order. What the model cannot know, and did not say, is that this hospital's antibiogram shows worrying local resistance to exactly that agent, that the drug it named is non-formulary and will trigger a pharmacy callback, and that the institution's own pneumonia pathway routes this patient to a different combination entirely. The model gave a textbook answer. The patient is not in a textbook. The patient is in this hospital, tonight, and this hospital has a protocol. This lesson is about closing that gap: making the AI answer from your institution's real guidelines, order sets, and formulary rather than from a generic average that may be outdated, may belong to a different country, and almost certainly is not yours.
The Generic Answer Is an Average of Strangers
To use AI safely for anything protocol-driven, you have to understand where a default answer actually comes from. A language model does not consult your hospital. It does not open your pharmacy's formulary, read your sepsis bundle, or check what your stewardship committee decided last quarter. When you ask a bare question, the model produces an answer assembled from the statistical average of everything it absorbed during training: textbooks, published guidelines from several countries and several years, forum posts, review articles, and a great deal of material whose provenance you cannot see. The answer that emerges is a kind of composite, a plausible consensus of strangers, none of whom practice where you practice.
That composite has three quiet problems baked into it, and each one can reach a patient. First, it may be outdated. Training data has a cutoff, and guidelines move; the "first-line" agent the model names may have been superseded by a newer recommendation, a new warning, or a formulary change the model never saw. Second, it may belong to a different system entirely. Dosing conventions, available formulations, brand names, and even which drug counts as first-line differ between countries and between health systems, and the model has no idea which one you are standing in. Third, and most importantly, it does not match your specific institution. Your hospital has a formulary, an antibiogram, an order set, and a set of local pathways that were built by your own committees for your own patient population and your own resistance patterns. The generic answer cannot reflect any of that, because it was never given any of that. It is answering a question about medicine in general when you asked a question about medicine here.
What Grounding Actually Means
Grounding is the practice of making the model answer from a real source you provide rather than from its own parametric memory. That distinction is the whole lesson, so it is worth slowing down on. A model has two very different ways to produce an answer. It can reach into its parametric memory, the diffuse statistical knowledge baked into its weights during training, and generate the most plausible-sounding continuation. Or it can be handed an actual document and told, in effect, "answer using this, and only this." The first mode is the average of strangers. The second mode is your protocol, read back to you. Grounding is deliberately steering the model out of the first mode and into the second.
In practice, grounding takes one of two concrete forms, and every clinician can do the first one today. The simplest is to paste the real source into the prompt: copy your institution's actual pneumonia pathway, or the relevant page of the formulary, or the current sepsis order set, directly into the conversation, and then ask your question with an explicit instruction to answer from that text. "Here is our hospital's community-acquired pneumonia pathway. Using only this document, tell me the recommended empiric regimen for a patient admitted to the general ward with no risk factors for resistant organisms." Now the model is not guessing from a global average. It is reading your pathway and reporting what it says. The second form is a connected tool: some institutional AI deployments are wired directly into the EHR, the formulary database, or a curated document library, so that the model retrieves the relevant institutional source automatically. That is more powerful and more convenient, but the principle is identical, and so is the clinician's responsibility. Whether you pasted the source or the system retrieved it, you are still accountable for confirming the answer actually reflects it.
The Difference You Can Feel in the Answer
A grounded answer feels different from an ungrounded one, and learning to feel that difference is a practical skill. An ungrounded answer is smooth, general, and rootless: it names a drug but cannot tell you why that drug, in your setting, over an alternative. A well-grounded answer points at something: "According to the pathway you provided, the empiric regimen for this ward-admitted patient is X, and the document notes Y as the alternative for penicillin allergy." It quotes, it references, it stays inside the four corners of the source. When an answer stops pointing at your source and starts sounding like a general lecture again, that is your signal that the model has drifted back to its parametric default, and the grounding has quietly failed.
There is a subtle craft to the instruction that produces this behavior, and it is worth being explicit about because the wording does real work. Compare three prompts. "What is the first-line antibiotic for community-acquired pneumonia?" invites a pure parametric answer. "Here is our pathway, what is the first-line antibiotic?" is better but still leaves the model free to blend the pathway with its memory. "Here is our pathway. Using only this document, state the recommended regimen, quote the line it comes from, and if the document does not address my question, say so rather than answering from general knowledge" is the version that actually pins the model to the source. The last clause matters as much as the first: without an explicit permission to say "the document does not cover this," a helpful model will reach into its parametric memory to fill any silence, and it will do so invisibly. You are not just handing the model a source; you are closing the exits through which its default knowledge would otherwise leak back in.
An ungrounded model answers about medicine in general. A grounded model answers about medicine here. The gap between those two is where a patient gets a non-formulary drug, a superseded dose, or a bundle that is not the one your hospital actually runs.
The Generic-Versus-Local Gap Is a Patient-Safety Issue
It is tempting to treat this as a matter of convenience, as if the only cost of a generic answer is an annoyed pharmacist and an extra phone call. That undersells the hazard badly. The gap between the generic answer and your local protocol is a genuine patient-safety issue, and it shows up in ways that are easy to miss precisely because the generic answer looks so correct.
Consider antibiotics, where the gap is sharpest. A model's default first-line agent for a given infection may not be on your formulary at all, which at best delays care while pharmacy substitutes and at worst leads to an off-pathway choice made under time pressure. Even when the drug is stocked, it may conflict with your local antibiogram, the institution-specific map of which organisms are resistant to which agents in your patient population. A drug that is a fine empiric choice nationally can be a poor one in a hospital where local resistance has climbed, and only your antibiogram knows that. The generic answer cannot. Layer on antimicrobial stewardship, the institutional program that governs which antibiotics are used when in order to preserve their effectiveness and reduce resistance, and the gap widens further: your stewardship committee may deliberately restrict the very agent the model recommends, reserving it for cases the model knows nothing about.
The same pattern holds beyond antibiotics. Your institution's sepsis bundle, the specific time-anchored set of interventions your hospital commits to for sepsis, may differ in its details from the textbook version the model recites, and in sepsis the details are the care. Your anticoagulation protocol, your DKA order set, your pain pathway, your discharge criteria: each of these is a local decision, and the model's default is a global average that was never asked to match any of them. When the generic answer and the local protocol diverge, following the generic answer is not a shortcut. It is a deviation from your own standard of care, made to look authoritative by a fluent machine.
It is worth naming why the local protocol exists at all, because that is what makes the deviation a safety issue rather than a stylistic preference. A local pathway is not an arbitrary house style laid over universal medicine. It is the encoded output of your own committees weighing your own data: your antibiogram's resistance patterns, your formulary's available agents and their contracted forms, your stewardship program's restrictions, your population's comorbidities, your pharmacy's dosing conventions, and the incidents your own quality process has learned from. When a hospital routes community-acquired pneumonia to a particular combination, that choice usually carries the memory of a local resistance trend or an adverse event the committee is deliberately steering around. The generic answer has none of that memory. So the divergence between the model's default and your pathway is rarely random noise; it is frequently the exact place where your institution has made a considered, local decision that the global average could not know about. Overriding it on the strength of a fluent machine is overriding your own institution's clinical judgment without realizing you have done so.
The gap widens further when you remember that a formulary is not only a list of which drugs exist but a set of local decisions about which form of a drug is stocked, at which concentration, under which restriction, and with which required monitoring. A model may recommend an agent your hospital carries only in an oral formulation when your patient needs the intravenous route, or a dose that assumes a concentration your pharmacy does not compound, or an agent that is on formulary but restricted to infectious-disease approval. Each of these is invisible to the generic answer because none of it is a fact about medicine in general; it is a fact about your building's supply chain and governance. A clinician who enters the model's default order discovers the mismatch only when the order rejects, the pharmacy calls back, or, in the worst case, a workaround substitution is made hurriedly at the bedside without the deliberate reasoning the formulary was designed to force. The point is not that the model is careless. The point is that the model was answering a question it was never equipped to answer, because the answer lives in documents it was never shown.
There is a further, quieter reason the local source matters that becomes visible only when you consider a patient who does not fit the average. Protocols and formularies frequently encode the exceptions your institution has decided to handle in a specific way: the penicillin-allergic pathway, the renal-dosing table, the pregnancy-adjusted regimen, the pediatric weight-based calculation, the alternative for a patient already colonized with a resistant organism. A generic answer tends to give you the modal case, the answer for the average patient, because that is what dominates its training data. Your local pathway, by contrast, was written by people who knew they would have to care for the non-average patient in front of you and built the exception into the document. Grounding on that document is therefore not only about getting the common answer right; it is about surfacing the institution's considered handling of the uncommon case, which is precisely the situation where a confident generic answer is most likely to be both wrong and dangerous. The average of strangers has no opinion about your specific outlier. Your protocol does, and grounding is how you reach it.
There is also a documentation dimension that clinicians feel later, at chart review. If you follow your local pathway, the record shows care consistent with your institution's own standard, which is the standard you will be measured against by a coding audit, a Joint Commission surveyor, or a plaintiff's expert. If you follow a generic AI answer that happened to diverge from the pathway, you have created a gap between what you did and what your own institution says it does, and that gap is precisely the kind of finding a review process is built to surface. The pathway is not only a clinical instrument; it is the yardstick your documentation will be held against, which is one more reason the local source, not the global average, is the thing to answer from.
Verify the Model Actually Used Your Source
Here is the trap that catches careful people, and it is subtle enough to deserve its own section. Grounding a model is not a switch that, once flipped, guarantees a grounded answer. Even when you paste the correct, current protocol into the prompt, the model can still blend that source with its own parametric defaults and hand you a hybrid: mostly your protocol, with a stray generic detail woven in so smoothly you would never spot the seam. It might report your pathway's regimen correctly but attach a dose from its own memory that your protocol does not specify. It might follow your source for four steps and then, where your document is silent, quietly fill the gap with a global default and present the whole thing as if it all came from you.
So the discipline does not end when you provide the source. You have to verify that the model actually used it. The practical checks are quick and they are non-negotiable. Ask the model to quote or cite the specific part of your source that supports each key claim, so a statement with no anchor in your document becomes visible. Scan for any specific value the model produced that is not actually in your source: a dose, a threshold, a timeframe that reads precisely but appears nowhere in the text you pasted is a fabrication, and it is more dangerous than a vague answer because it looks grounded. And do the thing that no prompt can do for you: open your actual protocol and confirm, with your own eyes, that the answer matches the current document. Grounding makes that final check fast and focused. It does not remove it. "The model was reading our protocol" is not verification any more than "the model said so" is, because the model may have been reading your protocol and its own memory at the same time and could not tell you which was which.
Ground on the Right Source, and the Current Version
Grounding solves one problem and quietly creates another, and mishandling the new one can be worse than never grounding at all. If you make the model answer from a provided source, then the source becomes the answer. Which means an outdated source grounds you, firmly and convincingly, in the wrong answer. Paste last year's antibiotic pathway and the model will faithfully report last year's recommendation, with all the fluent confidence of a grounded reply, and now the error wears two uniforms: the authority of the machine and the authority of "our own protocol." That is a harder error to catch than a generic guess, because it feels institutional and current when it is neither.
So grounding shifts the critical question from "does the model know" to "did I feed it the right document, in its current version." Before you trust a grounded answer, confirm two things about the source itself. First, is it the right source for this question. The general medicine pneumonia pathway is not the same as the immunocompromised-host pathway; the adult sepsis bundle is not the pediatric one. Grounding on a real document that is the wrong real document is its own failure mode. Second, is it the current version. Protocols are versioned and dated for a reason. The copy someone pasted into a shared prompt template six months ago, or the PDF saved to a desktop last year, may have been superseded by a revision you have not seen. Check the version and the effective date the way you would check any order set before you rely on it. Grounding is only as trustworthy as the source you ground on, and keeping that source correct and current is a human responsibility that no amount of clever prompting transfers to the machine.
A Worked Example: The Same Question, Grounded and Not
Watch the gap open and then close. A hospitalist is admitting a patient with suspected sepsis and wants to move fast. The ungrounded attempt is the natural one, typed into a general model: "What should I order for this septic patient?" Back comes a competent, textbook sepsis response: broad-spectrum antibiotics named generically, a fluid target, lactate monitoring, the familiar bundle. It is not wrong in the abstract. But it is an average of strangers. The antibiotics it names may not match this hospital's antibiogram or its stewardship restrictions; the fluid target and the timing may not match this institution's specific bundle; and the whole answer arrives with no way to tell which parts are your hospital's policy and which are the model's global default. If the hospitalist enters those orders on the strength of that answer, the model's memory has just written a piece of this patient's care.
Now the grounded version. The hospitalist pastes the hospital's current sepsis bundle, the actual document, version-dated this year, into the prompt and asks: "Here is our institution's sepsis bundle, effective this quarter. Using only this document, list the empiric antibiotic regimen it specifies for a patient with suspected sepsis and no beta-lactam allergy, the fluid resuscitation target, and the required timing for each step. For each item, quote the line of the bundle it comes from. If the bundle does not specify something I asked about, say so rather than filling it in." The answer that returns is anchored: each order points at a line of the real bundle, the antibiotics named are the ones this hospital actually uses given its own antibiogram and stewardship rules, and where the bundle is silent, the model says so instead of importing a default. The hospitalist then does the final check: opens the bundle, confirms the version is current, and confirms each quoted line is really there and really says that. Same clinical question, same thirty seconds of typing, entirely different safety profile. The first answer came from the average of strangers. The second came from this hospital, verified by this clinician, and the record can show exactly which document grounded the decision.
Notice what the grounded prompt accomplished beyond the correct drug. By forcing the model to quote the source and to declare its silences, the hospitalist turned an opaque answer into an auditable one. A colleague, a pharmacist, or a quality reviewer can now see not just what was recommended but what it rested on, and can catch a wrong-version bundle or an unsupported detail before it reaches the patient. That is the difference between an answer you hope is right and an answer you can defend.
Key Takeaways
- A model's default answer is a statistical average of its training data: it may be outdated, may reflect a different country or health system, and almost certainly does not match your hospital's specific formulary, antibiogram, order set, or protocol.
- Grounding means making the model answer from a real source you provide rather than from its parametric memory; the simplest form is pasting your actual protocol or formulary into the prompt and instructing the model to answer only from it.
- The generic-versus-local gap is a patient-safety issue, not a convenience issue: a generic first-line antibiotic may be non-formulary, may conflict with your local antibiogram or stewardship restrictions, and your sepsis bundle may differ from the textbook in exactly the details that are the care.
- Providing a source does not guarantee the model used it; a grounded model can still blend your protocol with its own defaults and produce a hybrid whose seams are invisible, so you must verify the source was actually used.
- Verify grounding by asking the model to quote the specific part of your source behind each claim, scanning for any specific value that is not actually in the source, and opening the real protocol to confirm the match with your own eyes.
- Grounding makes the source the answer, so an outdated or wrong source grounds you convincingly in the wrong answer; this error is harder to catch because it wears the authority of both the machine and "our own protocol."
- Before trusting a grounded answer, confirm you fed the model the right source for this specific question and the current, version-dated version of it; keeping the source correct and current is a human responsibility no prompt can transfer to the machine.
- A grounded, quoted answer is also an auditable one: the record can show which document and which version grounded the decision, turning an answer you hope is right into one you can defend.
Skill.re