โ†
AI for Healthcare & Clinical Practice
Visionary ยท M5 ยท lesson 5 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Identifying Novel, Defensible Clinical AI Use Cases
๐Ÿ“–
now learning

Identifying Novel, Defensible Clinical AI Use Cases

15 min

A vendor sits across the table from your innovation committee and demos an AI that predicts, from a patient's unstructured chart, which of your admitted patients will deteriorate in the next twelve hours. The room leans in. It is novel, it is impressive, and it would be the first of its kind in your market. Then your quietest committee member, a hospitalist who has been burned before, asks three questions: On what population was this validated? What is the false negative rate on the patients who look stable but are not? And when it is wrong, who owns the miss? The demo does not have clean answers. The excitement in the room does not change the fact that you have just been shown a case that is novel, expensive, and, on this evidence, not defensible. Learning to feel the difference between those two things, in the moment, in the room, is the core skill of this lesson.

The Question Behind the Question

At Level 5 you are no longer asking whether AI belongs in care. That argument is settled: roughly 75% of health systems now run at least one AI application, 71% of hospitals report predictive AI in the EHR, and US physician adoption crossed 63% in the 2026 Doximity report. The question that separates a mature innovation function from a reckless one is narrower and harder. Of all the things AI could plausibly do in your system, which few are worth actually doing, and which of those can you defend later to a safety committee, a regulator, a plaintiff's attorney, and a family? A defensible clinical AI use case is not the most novel one, the most fundable one, or the one that photographs best in a board deck. It is the one you can explain, validate, monitor, and stand behind when it is wrong, because it will sometimes be wrong.

Notice the frame. We are not screening for cases that will succeed; no one can promise that. We are screening for cases where, if the tool fails, you can still show that you deployed it responsibly. That is what defensible means. It is the difference between a program that survives its first bad outcome and one that does not, and it is a fundamentally different question from the one vendors and enthusiasts are asking. They ask what is possible. You are paid to ask what is responsible, and to say no to the impressive thing that is not.

Where the Next High-Value Case Actually Lives

The highest-value cases share a shape, and once you learn to see it you will find it repeatedly. The best clinical AI opportunities cluster where three conditions overlap: there is a large, expensive, repetitive burden on your workforce; the AI output is verifiable by the human who owns it before it reaches a patient; and the failure mode, when it happens, is bounded and catchable rather than silent and catastrophic. Ambient documentation is the canonical example precisely because it fits this shape, which is why it scaled first. The burden is enormous (documentation is the top driver of a physician burnout rate near 42%, with 20.9% of physicians doing eight-plus hours of after-hours EHR work). The output is a draft note a clinician reads and attests before it becomes the legal record. And the failure mode, a confabulated finding or a dropped negative, is catchable at the point of attestation by the person accountable for the chart.

Contrast that with the deterioration-prediction demo from the opening. The burden is real and the value is high, but the output is not cleanly verifiable before it acts, the failure mode is a silent false negative on a patient who looks fine, and the accountability for the miss is diffuse. Same enthusiasm, opposite defensibility. High value is necessary but nowhere near sufficient. The value tells you the case is worth examining; it tells you nothing about whether the risk is worth the reward. That second judgment is where most innovation programs go wrong, because value is easy to feel and risk is easy to discount, especially when the tool is novel and the room is excited.

It helps to hold the shape as a single mental picture. Imagine two funnels sitting side by side. In the first, work flows from the AI to a human who reads it, and only then to the patient; the human is a gate the output must pass, and every error the model makes is presented for inspection before it can act. In the second, work flows from the AI straight to the patient, and the human, if there is one, arrives after the fact to reconstruct what happened. The first funnel is where defensible cases live, because the verification that makes the tool safe is built into the flow rather than bolted on afterward. The second funnel can still be made defensible, but only by deliberately inserting a gate, and if you cannot find a place to put that gate without destroying the value, you are looking at a case whose risk is structurally hard to bound. Learning to sketch this funnel in your head during a demo is worth more than any feature list, because it tells you in thirty seconds whether the human check has anywhere to stand.

A defensible case is not the one you can build. It is the one you can explain, validate, monitor, and still stand behind on the day it is wrong.

The Four-Screen Filter

Rather than argue case by case, mature programs run every candidate through the same four screens, in order, and kill anything that fails one. The discipline of a fixed screen matters more than any single criterion, because it protects you from the specific failure of falling in love with a novel case and reverse-engineering a justification for it. Run the screens honestly and the weak cases eliminate themselves before they consume a year and a safety committee's patience. The table below is the filter in one view; each screen has a question, a piece of evidence you should be able to point to, and a disqualifying answer that should end the conversation regardless of how exciting the case is.

ScreenThe question it asksEvidence you should seeDisqualifying answer
1. ValueIs the burden large, expensive, and quantifiable, with a measured baseline?Hours, harm events, or dollars, sized before you build"We are not sure, but the technology is cool"
2. Bounded, verifiable riskWhen it is wrong, is the harm recoverable, and can the accountable human check the output before it acts?A human gate the output passes before reaching a patientAn autonomous action no one sees until after harm
3. Data readinessDo you have the data, at the quality and representativeness the model needs?HTI-1 source attributes and a stated training populationVendor cannot or will not show the training population
4. EquityWill it perform comparably across the patients you actually serve?Subgroup performance, or a firm commitment to test it prospectivelyBest for the majority group, untested elsewhere

Notice the screens are ordered so the cheapest disqualifier comes first. A case with no real burden dies at screen one before anyone spends a committee's time on risk analysis. A high-value case with unbounded risk dies at screen two before you invest in a data assessment. The ordering is not arbitrary; it is how you spend the least effort to reach the most no's, and it is why a fixed sequence beats a case-by-case debate that always seems to start from the feature that excites the room.

Screen One: Value That Is Real, Sized, and Not a Vanity Metric

Is there a large, expensive, quantifiable burden this would relieve, and can you state the number before you build? Not "clinicians would love this" but hours reclaimed, harm events prevented, or dollars of avoidable cost, with a baseline you measured rather than assumed. The trap here is the vanity case: novel, demo-friendly, beloved by the committee, but relieving a burden that was never actually large. Novelty for its own sake fails this screen every time. If the honest answer to "how big is the problem" is "we are not sure, but the technology is cool," you have your answer.

Screen Two: Bounded, Verifiable Risk

This is the screen the enthusiasts skip and the one that most often should stop a case. Two sub-questions. First, is the risk bounded: when the tool is wrong, is the harm limited and recoverable, or is it silent, downstream, and potentially catastrophic? Second, is the risk verifiable: can the accountable human actually check the output before it acts, or does the design ask them to trust it? A case where a clinician reads and attests a draft before it becomes real has verifiable risk. A case where an autonomous score changes management before any human sees it does not. The FDA authorizing a device does not resolve this for you; over 1,350 AI/ML-enabled devices are now authorized, but clearance certifies an intended use, not that the risk is bounded and verifiable in your workflow and your population. You still have to answer this screen yourself.

Screen Three: Data Readiness

Does the tool need data you actually have, at the quality it needs, representative of the patients it will serve? A model that requires clean structured problem lists in a system where those lists are notoriously incomplete is not ready, however elegant the model. Data readiness also folds in the ONC HTI-1 transparency you can now demand: for a predictive DSI, you are entitled to the source attributes, the nutrition-label facts about what the intervention was trained on and how it was validated. If the vendor cannot or will not show you the population the model was built on, you cannot assess readiness, and that alone is often disqualifying. Concretely, the source attributes are a nutrition-label-style set of facts: what data the intervention was developed on, the demographic makeup of that data, how the output should be interpreted, the intended use and the uses to avoid, and how the tool was validated and maintained. Since certified health IT had to meet the DSI criteria by the end of 2024 with ongoing maintenance from 2025, a predictive DSI in your EHR is supposed to expose these. Reading them is not a compliance chore; it is the fastest way to fail a case early. A model developed on a population that looks nothing like yours, or whose intended use is narrower than the use you were about to put it to, tells you at the identification stage what you would otherwise learn only after deployment, and every fact on that label is a claim to verify against your own data rather than to accept on faith.

Screen Four: Equity

Will this tool perform comparably across the patients you actually serve, or will it work best for the majority group in its training data and worst for the patients already underserved? This is not a compliance afterthought; it is a screen with veto power. A model trained on a non-representative population underperforms for exactly the people a safety-net or diverse system most needs to protect, and disparate performance is both a clinical harm and a legal exposure. A case that cannot demonstrate, or at minimum commit to prospectively testing, comparable performance across your populations is not defensible no matter how high its aggregate value looks. Equity is where a high-value case most often reveals that it was never as good as the aggregate number suggested, because an impressive average can hide a group for whom the tool performs badly, and that group is frequently the very one your safety-net or diverse system most needs to protect.

A Worked Example: Two Candidates, One Survivor

Bring the filter to life with two real-feeling candidates competing for the same innovation dollars. Candidate A is an AI that drafts the plain-language after-visit summary and the patient-portal reply, surfacing the clinical content the clinician then reviews, edits, and sends. Candidate B is an autonomous imaging triage tool that reprioritizes the radiology worklist by predicted acuity, moving studies up or down the queue before a radiologist has looked.

Run the screens. Candidate A: value is real and sized (inbox and after-visit documentation are a measured, heavy burden). Risk is bounded and verifiable (a licensed clinician reads and sends every message, satisfying both the safety logic and California AB 3030's disclosure regime). Data readiness is good (it works from the visit content you already capture). Equity is testable (you can measure reading-level and accuracy of drafts across language and literacy groups before scale). Candidate A passes all four and is defensible: you can explain it, validate it, monitor it, and stand behind it when a draft is imperfect, because a human caught it.

Sit for a moment with why Candidate A is not merely the safe choice but genuinely the innovative one. It changes a heavy, low-joy part of clinical work, it is measurable, and it can be improved iteration by iteration precisely because every output passes a human who can flag what went wrong. A case you can watch, correct, and refine is a case that gets better in your hands; a case that acts before you can see it is a case you can only audit after harm. Defensible cases are also, quietly, the more improvable ones, which is a reason the disciplined path tends to compound in value over time rather than merely avoid disaster.

Candidate B: value is real (worklist prioritization saves time and can speed critical reads). But screen two bites hard. The risk is not cleanly verifiable, because the reprioritization acts before a human reviews it, and it is not obviously bounded, because a study wrongly deprioritized as low-acuity is a silent delay on a patient who needed to be seen first, precisely the catastrophic-and-quiet failure the screen exists to catch. Equity compounds it: if the acuity model underperforms for a subgroup, that subgroup is systematically pushed down the queue. Candidate B is more novel and arguably more exciting, and it fails the filter. The mature decision is not "never," it is "not like this": Candidate B could become defensible if redesigned so that no study is delayed below a safety floor without human review, and if subgroup performance is validated first. The filter did not just reject a case; it told you exactly what would have to change to make the case defensible.

The Novelty Trap and the Me-Too Trap

Two opposite errors kill innovation programs, and a mature leader has to hold the line against both. The first is the novelty trap: chasing the newest, most impressive capability because it is new, mistaking being first for being right. Novelty is not value. A case is not better because no one has done it; often no one has done it because the risk is not worth the reward, and you are being offered the chance to discover that at your patients' expense. When you notice the room's excitement is running ahead of the evidence, treat that gap as the warning it is.

The second, quieter error is the me-too trap: adopting a case only because a competitor announced it, importing their risk without their validation. A tool that another system deployed is not thereby validated for yours; their population, their workflow, and their monitoring do not travel with the press release. The me-too trap is more seductive than the novelty trap precisely because it feels prudent. Copying a peer looks like the conservative move, the opposite of reckless novelty, and that disguise is what makes it dangerous. You tell yourself you are de-risking by following a proven adopter, when in reality you are inheriting a case whose one piece of genuinely load-bearing evidence, the local validation, is the piece that did not come in the box. The competitor's success is a reason to look harder at the case, never a substitute for looking. The next lesson on pilots and the one after on scaling are both, at heart, about refusing to inherit someone else's evidence. At the identification stage, the antidote to both traps is the same: the four screens do not care whether a case is novel or familiar, only whether it is valuable, bounded, ready, and equitable. Let the filter, not the fashion, decide.

There is a practical way to make the filter resistant to hype: require the requester to write the answers down before the meeting, not improvise them in it. A structured intake that asks, in advance, for the sized burden, the exact point where a human verifies the output, the training population and its representativeness, and the plan to test subgroup performance, does two things at once. It forces the enthusiasm to survive contact with specifics, and it produces the paper trail that later proves you screened the case honestly. Cases fueled by hype tend to collapse at exactly the questions the intake makes unavoidable: the burden turns out to be unsized, the human gate turns out to have no place to stand, the training population turns out to be undisclosed. A verbal pitch can dodge those questions; a written intake cannot. The discipline is not to be clever in the room. It is to make the room unnecessary by asking the four questions on paper, where a weak answer has nowhere to hide.

What Defensible Buys You

It is worth being concrete about why this discipline pays for itself, because saying no to exciting cases is politically costly and you will need to defend the practice of defending. A defensible case is one you can survive. When the tool is wrong (and every tool is eventually wrong), the defensible case lets you show a governance body, a Joint Commission surveyor operating under the new RUAIH expectations, or a court that you selected the use case responsibly: you sized the value, you bounded and verified the risk, you confirmed the data, you tested for disparate performance, and you built the monitoring before you scaled. That record is the difference between a bad outcome that is a tragedy and a bad outcome that is also negligence.

It helps to be precise about what the word defensible actually contains, because it is not a mood or a level of caution; it is a set of things you can produce on demand. Four pillars carry the weight, and a case is defensible only when you can show evidence under all four. The table names them, states the question each answers, and names the artifact you would hand a surveyor or a court.

PillarThe questionThe artifact you can produce
ValueWas the burden real and sized before you built?A measured baseline and the burden estimate
SafetyWas the risk bounded and verifiable, with a human gate?The workflow design and the attestation step
EquityDid you test for disparate performance across your populations?Subgroup validation results and the monitoring plan
EvidenceDid you validate on your own data, not the vendor's national numbers?The local validation report and HTI-1 source attributes

The distinction between the vendor's evidence and yours is the one that most often decides a case after the fact. A vendor's national accuracy figure is a claim about a population that is not yours, gathered under a workflow that is not yours; treat it as a number to verify, never a result to repeat, and insist on the source attributes that HTI-1 now entitles you to demand. When you can point to a local validation on your own encounters, subgroup results across the patients you serve, and a monitoring plan with an owner and a threshold, you have the four artifacts that turn a bad outcome from a tragedy into a defensible one. When you cannot, the four pillars collapse into the same admission.

The indefensible case offers none of that shelter. When it fails, and novel high-risk cases fail more often, you are left explaining why you deployed a tool whose risk you never bounded, on data you never checked, across populations you never tested, because it was exciting and a competitor had one. "The model recommended it" is not a defense, and neither is "it was innovative." The cardinal rule of this whole program scales up cleanly to the enterprise: AI assists, the clinician decides, the record proves it, and at Level 5 the record that matters most is the one showing you chose the case with your eyes open. Identifying the right cases is not the cautious alternative to innovation. It is what serious innovation actually is.

There is a second, quieter payoff that experienced leaders come to value even more than the legal shelter, and it is worth naming because it changes how a committee behaves. A team that trusts its own filter can move faster on the cases that pass, not slower. When everyone in the room knows that any case reaching a build decision has already survived the four screens, the endless relitigating stops; the argument is no longer whether to trust the tool but how to deploy the trustworthy one well. The screens do not just protect you from the bad case; they liberate you to commit to the good one without the low hum of unspoken doubt that makes organizations hesitate. Discipline at the front end buys conviction at the back end. The programs that scale clinical AI successfully are not the ones that say yes to everything; they are the ones that say no cleanly and early, and therefore get to say a wholehearted, well-documented yes to the handful of cases that deserve it. That is the posture this chapter is building toward, and the next two lessons, on running pilots to real evidence and on scaling across settings without breaking safety, are simply what happens after a defensible case has earned its yes.

Key Takeaways

  • At the enterprise level the question is no longer whether AI belongs in care but which few cases are worth doing and which of those you can defend later to a safety committee, a regulator, and a family.
  • A defensible case is not the most novel or fundable one; it is the one you can explain, validate, monitor, and still stand behind on the day it is wrong.
  • The highest-value cases share a shape: a large expensive repetitive burden, an output the accountable human can verify before it reaches a patient, and a failure mode that is bounded and catchable rather than silent and catastrophic.
  • High value is necessary but never sufficient; value tells you a case is worth examining, not whether the risk is worth the reward.
  • Run every candidate through four screens in order: real sized value, bounded and verifiable risk, data readiness, and equity. Kill anything that fails one, and let the filter rather than the fashion decide.
  • Bounded, verifiable risk is the screen enthusiasts skip: ask whether the harm is limited and recoverable and whether the accountable human can actually check the output before it acts.
  • FDA clearance and a competitor's announcement are not validation for your system; clearance certifies an intended use, and a press release imports risk without evidence.
  • Avoid both the novelty trap (chasing new because it is new) and the me-too trap (copying a competitor without their validation); the same four screens defeat both, and a defensible case is the only kind you can survive when it is wrong.