โ†
AI for Healthcare & Clinical Practice
Aware ยท M8 ยท lesson 8 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Separating Real Capability from Demo Magic
๐Ÿ“–
now learning

Separating Real Capability from Demo Magic

15 min

The demo is flawless. The AI listens to a simulated patient with a clear, unhurried voice, and out flows a perfect SOAP note. It reads a textbook chest film and flags the obvious finding in green. It answers a clean clinical question with a confident, well-formatted paragraph. Everyone in the room nods. And you, if you have learned anything from this level, feel a small, useful discomfort, because you know something the slide will never show you: a demo is a curated best case, and your patients live in the long tail. This lesson gives you the handful of questions that deflate a glossy pilot and reveal what a tool can actually do when the case is hard, the chart is messy, and the stakes are real.

A Demo Is a Best Case, By Design

Start by being fair to the demo. It is not usually a lie. It is a selection. A vendor building a demo naturally chooses the input that makes the tool shine: the clearly enunciated visit, the classic radiographic finding, the well-documented chart, the question with a clean answer. No one sets out to demo the atypical presentation, the patient with a thick accent and a productive cough, the chart with fifteen years of contradictory notes, or the question whose honest answer is "it depends." The demo is the top of the bell curve. Your job is care across the entire distribution, including the thin, dangerous tail where the tool was never shown to you performing.

This is the single most important reframe in evaluating any clinical AI: the gap between a demo and a deployment is the gap between the best case and the long tail. A tool can be genuinely excellent on the eighty percent of cases that look like the demo and quietly, dangerously unreliable on the twenty percent that do not, and it is precisely that twenty percent, the atypical, the complex, the sick, where a clinician's judgment matters most and where a confident wrong answer does the most harm. So the questions that matter are not the ones the demo answers. They are the ones the demo is built to avoid.

An analogy: the listing photo and the walkthrough

It helps to have a picture for the gap. A demo is a real-estate listing photo: shot with a wide lens on a sunny afternoon, the clutter pushed out of frame, the one flattering angle chosen from forty. It is not a lie. Every pixel is real. It is a selection, and the selection is the point. A deployment is the walkthrough on a rainy Tuesday, when you open the closet the photo never showed, notice the slope in the floor, and smell the basement. The listing photo tells you the house can look beautiful. It does not tell you how the house lives. When a vendor demos a clinical tool, you are looking at the listing photo. Your patients are the rainy Tuesday, and your job is to insist on the walkthrough before you sign anything.

The table below sketches the same gap across the five questions this lesson will build. On the left is the signal a polished demo sends. On the right is the reality that only shows up in deployment, once you ask the question the demo was staged to skip.

What the demo signalsWhat deployment actually decides
The tool handles this clean case beautifullyHow the tool handles the messy, atypical case you were never shown
A single headline accuracy figureSensitivity, specificity, and positive predictive value at your prevalence
It worked on the data in the studyIt works prospectively, in live workflow, on unselected patients
Look how confident and fast the output isWhat the tool does on the hard case that makes a clinician nervous
A human is in the loop, so it is safeWhich named human checks what, when, with the source and the time to do it

A demo shows you the best case the vendor could stage. Your patients are the long tail. The only useful questions are the ones the demo is designed not to raise.

Question One: On What Population Was It Validated?

The first question deflates more glossy pilots than any other: on what population was this tool validated, and does that population include patients like mine? A tool validated on a large academic center's patients may behave differently in a rural clinic, a safety-net hospital, or a pediatric population. A model trained mostly on one demographic can underperform, sometimes badly, on the patients it saw least, which is not a hypothetical but a documented pattern in clinical AI, and it lands hardest on the patients already underserved. If the vendor cannot name the validation population, that is not a small gap in the sales deck. It is the answer. A tool whose validation population is undisclosed is a tool you cannot yet trust for anyone, because you have no idea whose best case you are looking at.

Push further than "was it validated." Ask where, on whom, how many, and how recently. A model validated five years ago on one system's patients is making a promise about a world that has since changed. And ask the question that vendors least like: does the validation set include the patients I actually see, the elderly, the multimorbid, the non-English-speaking, the ones whose presentations are messy? If the honest answer is that the tool was validated on a population that looks nothing like your panel, then the impressive demo is evidence about someone else's patients, not yours.

There is a subtle trap here that catches even careful buyers. A vendor can answer "yes, we validated on a diverse national population" and still leave you exposed, because a population that is diverse in aggregate can be thin in exactly the subgroup you serve. A tool tested across many hospitals may have seen only a handful of patients like the ones who fill your clinic, and a handful is not validation. So the question is not merely whether the validation population was large or diverse, but whether it was adequately powered for the specific patients whose care you are deciding. The answer you want is not a reassuring adjective. It is a number, broken down by the subgroups that matter to you, and a vendor who has done the work can give it to you.

The clearance trap: what FDA authorization does and does not promise

Here is where the population question and the marketing slide collide most often. A vendor will project a badge and say the tool is FDA cleared, and the room relaxes, because clearance sounds like a promise of safety. It is not. Clearance is a regulatory authorization for a device with a defined intended use, granted because the manufacturer showed the device is safe and effective for that specific use, usually by demonstrating it is substantially equivalent to something already on the market. The regulator authorized a device to do a stated thing for a stated population. It did not certify that the device is safe in your workflow, on your patients, at your prevalence, wired into your electronic record the way you will wire it. Those are exactly the things the five questions are for, and clearance answers none of them.

The scale of this matters because clearance is now common enough to feel like a default. By early 2026, well over 1,350 AI and machine-learning enabled devices had been authorized, and the field is heavily concentrated: radiology alone accounted for roughly seventy-four percent of the AI and machine-learning device authorizations in 2024. Treat those figures the way this lesson treats every vendor number: as something to verify against the current regulatory database, not to repeat because it sounds impressive. The point of citing them is not the exact count. It is that "FDA cleared" no longer distinguishes a tool at all, and it never told you what you actually needed to know.

Consider two clinicians shown the same cleared tool. The first treats clearance as the finish line. She hears "FDA cleared," stops asking, and deploys the tool on her unit, quietly assuming a regulator has vouched for its behavior on her patients. Months later it misreads an atypical presentation from a population the clearance study barely contained, and she is surprised, though she should not be. The second clinician treats clearance as the starting line. She reads the intended-use statement out loud and asks a single question: does that stated use match what I am about to do with it, on the patients I actually see? She notices the cleared use is narrower than the vendor's pitch, that the validation population skews younger and less comorbid than her panel, and that the workflow the vendor imagines is not hers. She asks for the population breakdown, the prospective evidence, and the verification step. Same badge, same tool, opposite safety, and the only difference is that one clinician read clearance as a defined authorization and the other read it as a blanket guarantee it was never meant to be.

FDA clearance does guaranteeFDA clearance does not guarantee
A regulator reviewed the device for a defined intended useThe tool is safe in your specific workflow or on your population
The manufacturer showed safety and effectiveness for that stated useThe validation population resembles the patients you actually see
The device met a regulatory bar, often substantial equivalenceThe positive predictive value holds at your prevalence
There is a documented, checkable intended-use statementThe vendor's pitch stays inside that intended use

Question Two: What Is the Actual Error Rate, Not the Accuracy?

The word "accuracy" is where many clinicians get quietly misled, because it sounds precise and is often nearly meaningless. A tool that says "no cancer" to every image is 99 percent accurate in a population where cancer prevalence is 1 percent, and it is also completely useless and dangerous. Accuracy conflates the two kinds of error that matter to you and hides the base rate. What you actually need is the language you already know from screening tests: sensitivity, specificity, and, most of all, positive predictive value in your prevalence.

Sensitivity tells you how often the tool catches the thing when it is truly there, which governs the misses. Specificity tells you how often it correctly stays quiet when the thing is absent, which governs the false alarms. But the number that changes how you should actually treat an output is positive predictive value, and PPV depends on prevalence, on how common the condition is in your population. A tool with excellent sensitivity and specificity can still have a low PPV in a low-prevalence setting, meaning most of its positive flags are false, which is exactly the recipe for alarm fatigue. So the question is not "how accurate is it," a question with no useful answer, but "what are its sensitivity and specificity, and what does its positive predictive value become at my prevalence?" A vendor who can answer that is selling a tool. A vendor who only offers a single glossy accuracy figure is selling a demo.

Question Three: Was the Evidence Prospective and Real-World?

Now ask how the evidence was generated, because there is a world of difference between a tool that performed well on a retrospective, cleaned-up dataset the developers chose and a tool that performed well prospectively, in real workflow, on the patients who walked through the door. Retrospective validation on a curated dataset is the research equivalent of a demo: the developers looked backward at data they controlled and reported how the tool would have done. Prospective, real-world evidence is far harder to produce and far more honest, because it captures the messy inputs, the workflow interruptions, the atypical presentations, and the cases no one curated away.

The tell to listen for is the difference between "in our study the model achieved" and "in prospective deployment across these sites the tool achieved." The former can be, and often is, a best case assembled after the fact. The latter is a claim about the real world. Ask specifically: was this tested prospectively, in a live clinical setting, on unselected patients, and were the results published and independent, or is the evidence a vendor white paper describing a retrospective analysis of data the vendor curated? The strength of the evidence is not a technicality. It is the difference between a promise about your patients and a story about the vendor's spreadsheet.

Watch too for the gap between how a model performed in a validation study and how it performs once it is embedded in a real workflow, because the two can diverge sharply even for the same tool. A model can be accurate on paper and still fail in practice because clinicians game it, ignore it, or feed it inputs the study never contained. The famous cautionary tales in clinical AI are rarely models that were mathematically wrong in the lab; they are models that were mathematically fine and then behaved badly in the messy reality of a live unit, drifting as the population shifted or firing so often that staff learned to tune them out. Retrospective accuracy is a necessary condition, never a sufficient one. The only evidence that fully answers the question is evidence generated in the same kind of live, uncurated setting where you intend to use the tool.

Model drift deserves its own line, because it is the failure that a demo can never show and a one-time validation can never rule out. A tool validated on last year's patients is making a promise about this year's, and the world underneath it keeps moving: coding practices change, a new documentation template shifts what the inputs look like, a referral pattern sends you a sicker mix than the model ever saw, a lab swaps assays. None of that appears in the frozen dataset the vendor validated against, and none of it appears in the demo, which was staged once and never has to face a Tuesday six months from now. So the honest version of question three is not only "was it prospective," but "how will you know when it stops working, and who is watching." A vendor who has thought about deployment has an answer about ongoing monitoring. A vendor who has only thought about the sale has a validation number and a hope.

Question Four: What Happens on the Hard Case?

This is the question that separates clinicians from spectators, because only a clinician thinks to ask it. Watch what the tool does not on the clean demo input but on the case that makes you nervous. What does the ambient scribe produce when the patient has a heavy accent, speaks over the clinician, and describes three problems at once? What does the imaging tool do with the atypical presentation, the subtle finding, the poor-quality film, the anatomy it was not trained on? What does the summarizer do with the chart that contradicts itself across fifteen years of notes? What does the clinical-question tool do with the question whose real answer is uncertain, where a confident paragraph is itself a form of error?

Insist on seeing the hard case, and be suspicious when a vendor cannot or will not show you one. A tool that is only ever demonstrated on clean inputs is a tool whose behavior on dirty inputs is unknown, and unknown behavior on the hard case is exactly the risk you are being paid to manage. The most revealing question you can ask in any demo is disarmingly simple: "show me one it gets wrong." How a vendor responds tells you almost everything. A serious vendor has thought hard about the failure cases and will walk you through them, because they know their tool's edges better than you do. A vendor who cannot produce a single failure case has either not looked or does not want you to, and both answers should worry you more than any accuracy figure could reassure you.

Question Five: Who Stays Accountable When It Is Wrong?

The last question is the one this whole program keeps returning to, and it belongs in every evaluation. When this tool is wrong, and it will be, who is accountable, and specifically, who verifies what, and when? Watch closely for the phrase "human in the loop," because it is often used to reassure without meaning anything. A human in the loop is not a safety feature unless you can say exactly which human, checking exactly what, at exactly what point, with the time and information to actually do it. A tool that flashes an output on a crowded screen and calls the overwhelmed clinician who glances at it "the human in the loop" has a human in name only. The loop is decorative.

Consider how differently two tools with identical demos can behave once this question is answered. Both flag suspicious findings; both show a flawless example. But one is built so that every flag opens the underlying image, the measurement, and the comparison prior, inviting the radiologist to confirm or dismiss it in seconds, while the other simply appends a confident label to the worklist with nothing behind it. The first tool makes the human check faster and easier, which is what a well-designed clinical AI does. The second makes the human check harder, because there is nothing to check against, so the path of least resistance is to accept the label. Same demo, opposite safety profile, and the difference is invisible until you ask who verifies what, with what, and when. That single question can matter more to patient safety than any headline accuracy figure on the slide.

So make the claim concrete. If the vendor says a clinician reviews the output, ask what the clinician is expected to check, whether they have the source material to check it against, and whether the workflow gives them the seconds or minutes that checking actually requires. Ask whether the tool is designed to make verification easy, by showing its source, its confidence, the evidence behind a flag, or whether it is designed to be trusted, by presenting a confident answer with nothing to check it against. The first kind of tool respects the human check. The second kind quietly dismantles it while claiming to preserve it. And remember the cardinal rule that no demo can change: whatever the tool, accountability stays human, so the only "human in the loop" that protects a patient is a real one, named, resourced, and expected to look.

A Worked Example: Deflating the Pilot

Picture a vendor presenting an AI tool that flags patients at high risk of deterioration, promising to save lives by catching decline early. The slide says "94 percent accuracy" over a photograph of a grateful family. The room is impressed. Then a clinician who has done this level asks five quiet questions. On what population was it validated? The vendor names a single large academic hospital. Does that include patients like ours, in a community setting with an older, sicker, more diverse panel? The vendor is not sure. What is the positive predictive value at our deterioration prevalence, not the headline accuracy? The vendor has only the accuracy figure. Was it tested prospectively in live workflow, or retrospectively on curated data? Retrospectively. When it fires a false alarm, who verifies, checking what, with what time, before anyone acts?

Notice what those five questions did. They did not require a data-science degree. They required only the discipline to refuse to be impressed by a curated best case and to ask, instead, about the long tail where your patients live. The "94 percent accuracy" did not survive contact with a single one of them, not because the tool is necessarily bad, but because the demo was never evidence about your patients in the first place. This is the whole skill. A demo is a story a vendor tells. Your five questions turn it back into a question about the only thing that matters, which is whether this tool is safe for the patient you will actually see, on the day the case is hard.

It is worth being clear about the goal, because these questions are not a rejection ritual. The point is not to sink every pilot or to treat every vendor as an adversary; plenty of clinical AI is genuinely good and genuinely helpful, and a tool that answers all five questions well deserves your serious attention. The point is to move the decision from the vendor's ground, where a polished demo is the strongest evidence, onto yours, where validated performance on your patients is. A good vendor welcomes these questions, because a good product survives them, and the questions actually help the vendor prove the value they are claiming. The clinician who asks them is not the obstacle in the room. They are the person doing the one job the demo cannot do for anyone: separating what a tool can really do from what it was staged to appear to do, so that when it reaches a patient, it reaches them as a verified capability rather than a hopeful guess.

Key Takeaways

  • A demo is a curated best case by design; the vendor selects the input that makes the tool shine. Your patients are the long tail of atypical, complex, and messy cases the demo is built to avoid.
  • Question one, population: on what population was it validated, and does that include patients like mine? An undisclosed validation population is itself the answer, because you cannot know whose best case you are seeing.
  • Question two, error rate: "accuracy" is often meaningless because it hides the base rate. Ask for sensitivity, specificity, and above all positive predictive value at your prevalence, because a high accuracy figure can mask useless or dangerous performance.
  • Question three, evidence quality: was it validated prospectively in live workflow on unselected patients, or retrospectively on a dataset the vendor curated? Retrospective, curated evidence is the research equivalent of a demo.
  • Question four, the hard case: insist on seeing the tool handle the accented speaker, the messy chart, the atypical film, the uncertain question. "Show me one it gets wrong" is the most revealing question in any demo.
  • Question five, accountability: "human in the loop" means nothing unless you can name which human checks exactly what, when, with the source material and the time to actually do it. A decorative loop is no loop at all.
  • The five questions require clinical judgment, not data science, and they turn a vendor's story back into the only question that matters: is this tool safe for the patient I will actually see when the case is hard?
  • A vague accuracy figure, an unnamed validation population, testimonials in place of data, and an unspecified human in the loop are the four warning signs that you are being shown demo magic, not real capability.