โ†
AI for Healthcare & Clinical Practice
Visionary ยท M11 ยท lesson 11 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Clinical AI Pilots to Evidence
๐Ÿ“–
now learning

Running Clinical AI Pilots to Evidence

15 min

Six months into a well-loved pilot, the innovation director stands in front of the clinical AI governance committee with a slide that glows: 340 clinicians enrolled, 92% still using the tool weekly, average note-writing time down eleven minutes, satisfaction scores through the roof. The committee is delighted. Then the chief quality officer, who has read a lot of malpractice depositions, asks the one question that empties the room's good mood: "This tells me they like it and it is fast. What does it tell me about whether it is accurate, and whether it is accurate for every kind of patient we serve?" The slide has no answer, because the pilot was never designed to produce one. Eleven minutes and a smile are not evidence. They are the trap this lesson exists to help you avoid.

The Difference Between a Pilot and a Demo That Lasted Six Months

Most clinical AI pilots are not pilots. They are extended demos, run to build enthusiasm and momentum, and they succeed at exactly that while proving nothing that a safety committee or a regulator can rely on. The tell is what they measure. A demo-that-lasted-six-months measures adoption, speed, and satisfaction, all of which are real and none of which answer the question that actually governs whether the tool should scale: is it accurate, is it safe, and does it perform comparably across the patients you serve? A pilot, in the disciplined sense this lesson means, is an experiment designed before it starts to produce evidence on those questions, evidence you could hand to a skeptical quality officer, a Joint Commission surveyor operating under the new RUAIH expectations, or a plaintiff's expert, and have it hold.

The distinction is not academic, because the two kinds of pilot fail in opposite ways. The extended demo fails silently: it generates confidence untethered to evidence, and that confidence carries the tool into full deployment where its unmeasured error rate finally reaches patients at scale. The disciplined pilot can fail loudly and early, which is the point; it is built to surface the tool's real accuracy and its real disparate performance while the tool is still contained, so that a no-go decision costs you a pilot rather than a patient. If you take one idea from this lesson, take this: the purpose of a pilot is not to prove the tool works. It is to find out whether it works, on your data, in your workflow, for your patients, in time to stop it if it does not.

Pre-Specify, or You Are Just Collecting Anecdotes

The single discipline that separates evidence from anecdote is pre-specification: writing down, before the pilot begins, what success looks like, what failure looks like, and what you will do about each. This is borrowed directly from clinical trial methodology, and for the same reason it exists there. If you decide what counts as success only after you see the results, you will unconsciously define success as whatever the results happened to show, and your pilot becomes an elaborate mechanism for confirming what you already wanted to believe. Pre-specification is the guardrail against your own motivated reasoning, and it is the first thing a rigorous reviewer will look for.

A pre-specified pilot names its numbers in advance. It sets a success criterion: the accuracy, safety, or performance threshold the tool must hit to justify scaling, stated as a number, not a vibe. It sets a stopping criterion: the level of error, harm, or disparate performance at which you halt the pilot immediately, regardless of how much everyone likes the tool. And it commits, in advance, to a go/no-go decision tied to those criteria, so that when the data arrives the decision is already made in principle and only has to be executed. The stopping criterion is the one people most want to skip and the one that matters most, because a pilot with no pre-specified stop is a pilot that will rationalize continuing no matter what it sees. You are far more honest about a threshold you set before you were emotionally invested in the outcome.

It is worth being blunt about the psychology here, because it is the whole reason the discipline exists. By month three of a pilot, the tool has a fan base. Clinicians have reorganized their day around it, the vendor has become a partner, the innovation team has staked its credibility on the launch, and everyone in the room has a quiet stake in the answer being yes. That is a perfectly human situation and it is exactly the situation in which a threshold invented on the spot will be generous, a worrying signal will be explained away, and a subgroup problem will be labeled an artifact of small numbers rather than a warning. Pre-specification is how a sober version of you, the one who existed before any of that investment accrued, ties the hands of the invested version of you who will be in the room when the data lands. The criteria are a letter you write to your future self, and the more uncomfortable it feels to write a hard stopping rule at the start, the more you needed to write it.

There is also a governance benefit that leaders underrate. When criteria are written and agreed at the outset, the eventual decision belongs to the evidence rather than to the most senior or most enthusiastic voice in the room. A committee that pre-specified its thresholds is far more resistant to being talked into a bad go by a charismatic champion, because the argument is no longer about impressions; it is about whether a measured number cleared a line everyone agreed to before they knew the result. That is how you keep an innovation function honest without turning it into a source of endless friction: the friction is spent up front, once, in defining what would count, and then the data is allowed to speak.

It is worth being concrete about what a pre-specified plan actually contains, because "pre-specify" is empty until you can name the fields. A usable pilot plan states, in writing and before enrollment: the specific clinical question and the population it applies to; the primary outcome measure and exactly how it will be computed; the sample size and how long the pilot will run to reach it; who reviews the outputs and whether they are blinded to the tool; the subgroups the results will be stratified by; the numeric success threshold; the numeric stopping threshold and who has the authority to pull the trigger; and the go/no-go rule that maps the measured numbers onto a decision. If any of those is missing, the pilot has a hole a motivated reading will later crawl through. The plan should be short enough to fit on two pages and specific enough that two different reviewers, handed the same results, would reach the same verdict. That last property, that the decision is reproducible from the plan and the data alone, is the real test of whether you pre-specified or merely gestured at rigor.

A pilot that only proves people adopted the tool and it saved time has proven the two things that were never in doubt and measured nothing that keeps a patient safe.

Measure Accuracy AND Equity, Prospectively and Locally

What you measure is the whole game, and two words carry most of the weight: prospectively and locally. Prospectively means you measure the tool's performance as it runs, on cases as they occur, rather than reconstructing a flattering story from the cases you happen to remember. Locally means you measure it on your patients, in your workflow, not on the vendor's validation set from some other population. The vendor's numbers describe the vendor's population; they are a reason to start a pilot, never a substitute for one. The whole point of a local prospective measurement is that it is the only thing that tells you how the tool behaves in the exact conditions where it will actually be used.

A note on why local matters so much, because leaders often accept a vendor number they would never accept from an internal analyst. Performance in clinical AI is not a fixed property of the model; it is a property of the model meeting a specific population and a specific workflow. The same ambient scribe that is excellent in a quiet primary-care room can degrade in a noisy emergency department, with overlapping speakers and interrupted trains of thought. The same risk model validated on an academic center's insured population can behave differently on a safety-net panel with more missing data and more comorbidity. When a vendor says the tool is 94% accurate, the honest translation is that it was 94% accurate on someone else's patients doing something slightly different from what your patients will do. That number is a hypothesis about your setting, not a finding in it, and the pilot exists precisely to test the hypothesis.

And you must measure two things, not one. The first is accuracy: how often the tool is right, and, more revealingly, the shape of its errors when it is wrong. A note tool's accuracy is not a single number; it is the rate of confabulated findings, dropped pertinent negatives, wrong laterality, and fabricated values, each of which is a different clinical hazard. The second, which the extended demo almost always omits, is equity: does the accuracy hold across the subgroups you serve, or does an impressive aggregate number hide a tool that works well for the majority and poorly for the patients already underserved? A model trained on a non-representative population underperforms for exactly those patients, and an aggregate accuracy figure will conceal it perfectly. Measuring equity means stratifying your accuracy results by the populations that matter, by race and ethnicity, by language, by age, by the axes along which your patients differ, and looking hard at whether the tool is quietly worse for some of them. A pilot that measures accuracy but not equity has done half the job and produced a number that can be dangerously reassuring.

The mechanics of measuring accuracy honestly deserve to be spelled out, because a sloppy method can flatter a tool as easily as a good method exposes it. The gold-standard approach for a note or summary tool is a blinded chart review: a prospective sample of real outputs is pulled, a clinician reviewer who does not know which notes came from the AI compares each against the source encounter, and the errors are categorized by type and severity, not lumped into one pass or fail. Blinding matters because a reviewer who knows a note is AI-generated will read it differently, either more forgivingly or more harshly, and either way the number stops being trustworthy. Sampling matters because you cannot review everything, so you draw a prospective, representative sample rather than the cases someone remembers. And a defined error taxonomy matters because "accuracy 94%" hides whether the 6% of errors were harmless phrasing slips or fabricated exam findings that change management. Choose a comparator too: the honest bar is usually the current clinician baseline, because the question is not whether the tool is perfect but whether it is at least as safe as what it replaces. A pilot that skips blinding, skips sampling discipline, or reports one undifferentiated accuracy number has measured something, but not the thing a safety committee can rely on.

Shadow Mode: The Safest Place to Find Out

There is a way to run a pilot that measures real accuracy on real cases without letting a single unproven output touch a patient, and mature programs use it as a default. It is called shadow mode: the AI runs on live clinical cases and generates its outputs, but those outputs are held back from care and compared against what the accountable clinicians actually did. The tool predicts, drafts, or scores in parallel with normal care, and no patient is exposed to its judgment; you simply collect, day after day, the record of where the AI agreed with the clinicians and where it diverged, and you inspect the divergences to learn what the tool gets wrong before you ever let it act.

Shadow mode is powerful because it dissolves the central tension of piloting, that you cannot know if a tool is safe until you use it, but you should not use it until you know it is safe. Running silently, the tool generates exactly the prospective, local accuracy and equity data you need, at full clinical fidelity, with zero patient risk, because the safety net of ordinary care never comes down. It is especially suited to predictive and scoring tools, where you can compare the AI's prediction against the outcome that actually unfolded. For a deterioration or sepsis model, shadow mode answers the question the enthusiastic demo never could: on the patients this model would have flagged, and the ones it would have missed, what actually happened? That is evidence a safety committee can weigh, generated without gambling a single patient to get it.

The mechanics of a shadow-mode evaluation are worth making concrete, because the value comes from measuring the right comparison. For a predictive score, you log the model's output on every eligible patient, then wait for the actual outcome to unfold and compare. This lets you build the confusion matrix that a demo never produces: the patients the model flagged who did deteriorate (true positives), the ones it flagged who did not (false positives, the alarm-fatigue tax), the ones it missed who deteriorated anyway (false negatives, the dangerous ones), and the ones it correctly left alone. From that you can compute the sensitivity, specificity, and, crucially, the positive predictive value in your population, which is the number that tells clinicians how often a flag actually means something, and which shifts dramatically with how common the condition is in your patients. A model with excellent sensitivity but a 5% positive predictive value will flood a unit with false alarms, and shadow mode surfaces that before a single nurse is trained to trust the alert. You also stratify the shadow results by subgroup, exactly as you would in a live pilot, so the equity question is answered before exposure, not after. Shadow mode, run this way, is not a vague "let it watch for a while"; it is a structured prospective accuracy and equity study with the safety net of ordinary care still fully up.

Shadow mode is not free, and part of running pilots well is knowing its limits. It cannot measure how the tool changes clinician behavior, because in shadow mode no clinician sees the output, so effects like automation bias and workflow disruption remain invisible until the tool goes live. It can also be gamed by the passage of time: a model that looks accurate in a silent run may drift once the population or the documentation patterns shift, which is why the monitoring you build for shadow mode has to survive into deployment rather than being switched off the day the tool goes live. Used well, shadow mode is the first of two phases, a silent accuracy-and-equity phase that establishes the tool is worth exposing to patients at all, followed by a carefully instrumented limited live phase that measures the human-factors effects shadow mode cannot see. Skipping the first phase gambles patients on unproven accuracy; skipping the second ships a tool whose real-world behavior you have never actually observed. The disciplined program runs both, in that order, and treats the transition between them as its own go/no-go with its own pre-specified criteria.

A Worked Example: The Pilot That Was Almost a Triumph

Return to the glowing slide from the opening and rebuild that pilot the disciplined way, so the contrast is concrete. Before enrolling anyone, the team writes the plan. Success criterion: the ambient note tool must produce a signed-note error rate no higher than the current clinician baseline, with no increase in clinically significant errors (fabricated findings, wrong laterality, dropped critical negatives), measured on a prospectively sampled, blindly reviewed set of notes, and it must hold this within each major patient subgroup. Stopping criterion: any clinically significant fabricated finding rate above a pre-set floor, or any subgroup where the error rate is materially worse, halts the pilot for review. Go/no-go: scaling requires meeting the success criterion overall and in every measured subgroup; failing either is a no-go regardless of adoption or speed.

Now the pilot runs, and it collects the measures that matter alongside the ones that feel good. Adoption and time saved are still tracked, because they are real benefits, but they are not the verdict. The verdict comes from a blinded chart review of a prospective sample, stratified by subgroup. Suppose the review finds the tool excellent overall, error rate at baseline, clinicians thrilled, but in the subset of visits conducted through a medical interpreter, the fabricated-finding rate is three times higher, because the tool handles interpreted, code-switched speech poorly. The glowing demo would have shipped this tool and quietly degraded documentation quality for exactly the patients a language barrier already disadvantages. The disciplined pilot catches it, triggers the stopping criterion for that subgroup, and turns a near-miss into a no-go-until-fixed. Same tool, same clinicians, same six months. The only difference is that one pilot was designed to find the truth and the other was designed to feel good, and only one of them protected the patients it was supposed to serve.

Notice how each design choice earned its place in catching that failure. The prospective sample meant the interpreted visits were actually in the reviewed set rather than filtered out by whoever picked the cases. The blinding meant the reviewer did not unconsciously grade the AI-drafted interpreted notes more gently. The subgroup stratification meant the interpreter cohort was reported as its own line rather than diluted into a reassuring average. And the pre-specified subgroup stopping rule meant the finding forced a decision instead of becoming a bullet point someone promised to "keep an eye on." Remove any one of those, and the tool ships. That is the practical lesson: rigor is not one grand gesture but a chain of small, boring methodological choices, each of which is the difference between seeing the failure and shipping it. The discipline is cheap to specify at the start and nearly impossible to reconstruct after the fact, which is exactly why it has to be written down before the pilot runs, not assembled from memory when the results are already in hand and everyone already wants the answer to be yes.

Evidence, Not Enthusiasm, and the Clean Decision

The reason all of this rigor pays off is that it produces a decision you can actually defend and actually execute. When a pilot is pre-specified, measured prospectively and locally on accuracy and equity, and run in shadow mode where the risk warrants, the go/no-go decision at the end is clean: you compare the evidence to the criteria you set before you were invested, and you act. There is no agonized committee meeting where enthusiasm wrestles with a vague unease, because the unease has been made concrete, into thresholds, into subgroup tables, into a documented decision rule. This is what it means to run a pilot to evidence rather than to a feeling. And notice that this rigor is not the enemy of speed; it is what makes speed safe. A program that can stop a bad case cheaply and scale a good one with conviction moves faster overall than one that reverses a full deployment after harm, which is the most expensive and slowest thing a health system can be forced to do.

A word on the metric that gets gamed, because it is the most common way a disciplined-looking pilot quietly becomes a demo. When a tool's success is tied to a single easy number, people optimize the number rather than the outcome it was supposed to stand for. A sepsis alert judged only on "cases flagged" can be tuned to flag almost everyone, hitting its sensitivity target while burying clinicians in false alarms that the pilot never counted. A note tool judged only on "minutes saved" can save minutes precisely by encouraging clinicians to stop reading, which is the exact failure the primary-care coding audit exposed elsewhere in this program. A documentation tool scored on "clinician satisfaction" can be loved because it is fast and wrong. The defense against a gamed metric is to pair every efficiency or adoption number with a safety-and-accuracy number measured independently, and to watch for the tell: a metric that improves while the thing it was supposed to represent gets worse. If minutes-saved climbs while chart-review error rate also climbs, the minutes-saved metric has been gamed, and a pilot that only tracked the first number would have called that a triumph. Pre-specifying the accuracy and equity measures alongside the feel-good ones is what makes gaming visible instead of invisible.

It also produces the artifact that matters most when the tool eventually errs at scale, as every deployed tool eventually does. A defensible pilot leaves behind a record: here is what we required, here is what we measured, here is how it performed across our populations, here is why we decided to scale or not. That record is the difference between a program that can show a surveyor or a court that it deployed responsibly and one that can only show a satisfaction survey. Enthusiasm evaporates the moment something goes wrong; evidence is still there in the file. Beware, above all, the pilot that proves only adoption and speed, because it is the most dangerous artifact in clinical AI: it looks exactly like success, it feels exactly like validation, and it has told you nothing about whether the tool is safe. The whole discipline of this lesson can be compressed into one sentence you should carry into every governance meeting: a pilot that only proves people liked it and it was fast has not been a pilot at all.

Key Takeaways

  • Most clinical AI pilots are extended demos that measure adoption, speed, and satisfaction, which are real but never answer the governing question: is the tool accurate and safe, and does it perform comparably across the patients you serve.
  • The purpose of a pilot is not to prove the tool works but to find out whether it works, on your data, in your workflow, for your patients, in time to stop it if it does not.
  • Pre-specify before you start: a numeric success criterion, a stopping criterion, and a go/no-go decision tied to them, so success is not defined after the fact as whatever the results happened to show.
  • The stopping criterion is the one people most want to skip and the one that matters most; a pilot with no pre-specified stop will rationalize continuing no matter what it sees.
  • Measure accuracy AND equity, prospectively and locally: the vendor's validation set describes the vendor's population, and an impressive aggregate accuracy can hide a tool that works poorly for the patients already underserved.
  • Stratify accuracy by the subgroups that matter (race and ethnicity, language, age); a pilot that measures accuracy but not equity has done half the job and produced a dangerously reassuring number.
  • Use shadow mode where risk warrants: run the AI on live cases with its outputs held back from care and compared to what clinicians actually did, generating real prospective accuracy and equity data at zero patient risk.
  • Run the pilot to evidence, not enthusiasm; the pilot that proves only adoption and speed is the most dangerous artifact in clinical AI, because it looks exactly like success and has told you nothing about safety.