Prioritizing Use Cases by Value and Risk
The governance committee had a whiteboard and a hard afternoon ahead. On the table were four AI proposals competing for the same finite pool of validation effort, integration hours, and clinician patience: an ambient scribe for the primary care clinics, an inbox drafting assistant for the patient portal, a sepsis prediction model the ICU physicians were excited about, and an autonomous imaging triage tool a vendor swore would read chest films faster than the night radiologist. Every proposal had a champion, a demo, and a slide claiming a stunning number. The CMIO asked the only question that mattered: "If we can only do two of these well this year, which two, and how do you defend that choice to the board, the safety committee, and a family whose loved one was hurt by the third?" The room reached for adjectives. This lesson gives leaders the discipline that replaces adjectives with a defensible ranking: value against risk, and the courage to start where value is high and risk is bounded and verifiable.
Why Ranking Is the Real Leadership Job
Choosing which AI to build or buy is not the hard part. Vendors will bring you more good-looking opportunities than you could staff in a decade. The hard part, the part that is actually a leadership act rather than a purchasing act, is ranking: deciding which opportunities go first, which go later, and which do not go at all, and being able to defend that order to people who will second-guess it. A ranking that lives only in your head is not a strategy. A ranking you can put on a page, with the reasoning visible, is the difference between a governance committee that steers and one that simply reacts to whoever demos last.
The reason ranking is hard is that the two things you must weigh do not share a unit. On one side is value: how much a capability advances an outcome the organization already cares about, whether that is clinician burnout, throughput, quality, avoidable harm, equity, or margin. On the other side is risk: how much damage the capability can do to a patient, to a population, or to the organization's legal and regulatory standing if it fails, and, crucially, how easily that failure can be caught before it reaches a patient. Value pulls you toward the biggest, most exciting bets. Risk should pull you toward the ones you can deploy safely and prove. A mature ranking holds both in view at once, and it refuses to let a large value number silence a large risk number.
This is why the naive ranking, sort every opportunity by projected return on investment and fund from the top, is dangerous in healthcare in a way it is not in most industries. In a retail company, the worst case of a wrong AI bet is a wasted quarter. In a health system, the worst case is a patient harmed by an output no one verified, a population underserved by a model that never worked for them, a fraud finding when a coding tool upcoded in the background, or a malpractice case in which the defense is the indefensible sentence "the model recommended it." Return on investment is a real axis, but it belongs on the value side of a two-dimensional judgment, not as the sole sort key. The moment a leader sorts purely by value, they have quietly decided that risk does not count, and risk in this domain is measured in patients.
What Value Actually Means, and Why It Is Not Just ROI
Value in a clinical AI portfolio is broader than a finance spreadsheet, and a leader who reduces it to dollars will systematically undervalue the capabilities that matter most to a care-delivery organization. A disciplined value assessment weighs several strands and names them explicitly so the ranking can be audited later.
The first strand is clinical benefit: does the capability improve outcomes, reduce avoidable harm, or catch something clinicians reliably miss? This is the strand most likely to be overstated in a vendor demo and most in need of your own verification against your own population. The second strand is operational and financial benefit: throughput, length of stay, denials, after-hours documentation time, staffing pressure relieved. The third strand, easy to forget and often the largest in practice, is clinician experience: a tool that gives a burned-out primary care physician back an hour of after-hours charting has real value even if it never touches a diagnosis, because a 2025 multi-system study found burnout falling from roughly 52 percent to roughly 39 percent within thirty days on an ambient documentation tool. Treat that figure as a number to verify against your own baseline, not a promise to paste into a business case, but do not pretend clinician relief is not value. The fourth strand is strategic fit: does this advance a goal the board already holds, or is it a shiny orphan that advances nothing anyone in the C-suite can name?
Two disciplines keep the value axis honest. First, insist on a baseline and a target before ranking, not after. "Improve documentation" is not a value claim; "reduce after-hours EHR time for primary care from the current measured baseline to a defined target within two quarters, measured by the EHR's own time logs" is. Second, discount vendor-supplied value by your own ability to realize it. A model that produced a beautiful result in an academic center with a data-science team and a homogeneous population may deliver a fraction of that in your community hospital. The value that belongs in your ranking is the value you can actually capture and prove, not the value on the vendor's slide.
It helps to make the four strands of value explicit as a scoring grid rather than a single impression, because a tool can be strong on one strand and empty on the others, and the impression tends to be dominated by whichever strand the demo emphasized. Scoring each strand separately forces the committee to name where the value actually is.
| Value strand | What it captures | How to verify it locally |
|---|---|---|
| Clinical benefit | Better outcomes, fewer misses, avoidable harm reduced | Demand subgroup performance on your population; treat the vendor number as a hypothesis to test. |
| Operational and financial | Throughput, length of stay, denials, after-hours documentation time, staffing pressure | Tie to an existing operational metric with a measured baseline, not a projected saving. |
| Clinician experience | Time and cognitive burden returned to burned-out staff | Measure after-hours EHR time or a validated burnout instrument before and after, on your clinicians. |
| Strategic fit | Whether it advances a goal the board already holds | Name the specific stated priority it serves; if no one in the C-suite can name it, the fit is weak. |
A tool that scores high on clinician experience and strategic fit but is untested on clinical benefit is not worthless; it is a documentation or administrative tool whose value is real and whose clinical risk is low, which is precisely why such tools sequence early. A tool that claims enormous clinical benefit but scores that benefit only from a vendor slide has not earned a high value number yet. It has earned a demand for local evidence.
What Risk Actually Means: Severity, Population, and Verifiability
Risk in a clinical AI portfolio is not a single number either, and the leaders who rank well decompose it into strands they can reason about separately. The most important insight in this entire lesson is that not all risk is equal, and the strand that should drive your sequencing hardest is not how bad a failure would be but how easily you can catch it before it reaches a patient.
The first strand is severity: if the AI is wrong, how much harm can reach a patient? A scribe that confabulates a physical exam finding the physician never performed is a serious record-integrity and liability problem. A diagnostic model that misses a bleed is a direct patient-safety catastrophe. Severity tracks how close the output sits to an irreversible clinical decision. The second strand is autonomy: how much does the workflow let the AI act without a human check? An output that a clinician must read and sign is bounded by that human step; an output that fires an order or a message with no human in the loop is not. The third strand, and the one this lesson elevates, is verifiability: can a clinician readily confirm or refute the output at the point of use? An ambient note can be read against the encounter the clinician just conducted; the clinician was in the room. A black-box risk score that says "this patient is high risk" with no reasoning a clinician can check is far harder to verify, which means the human check that is supposed to bound the risk may be hollow.
Human-in-the-loop, the design principle that a qualified human reviews and is accountable for every AI output that touches a patient, is only as strong as the verifiability of the thing being reviewed. This is the trap that catches sophisticated systems: they deploy a high-severity model, they put a human in the loop to feel safe, and then the human cannot actually verify the output under real conditions, so the loop becomes a rubber stamp powered by automation bias, the well-documented human tendency to accept an authoritative machine output under time pressure without the usual scrutiny. A loop you cannot verify is not a safety control. It is a liability with a signature line.
Because these three strands move independently, it is worth scoring them separately for each candidate rather than collapsing them into a single sense of dread. The combination that should worry a leader most is high severity paired with low verifiability, because that is exactly where the human check that everyone points to as the safeguard is weakest. The table below shows how the same tool categories map onto the strands, and why documentation lands in a very different place than a black-box score.
| Risk strand | The question it answers | Low-risk example | High-risk example |
|---|---|---|---|
| Severity | If it is wrong, how much harm can reach a patient? | A scribe drops a word the clinician re-reads | A diagnostic model misses an intracranial bleed |
| Autonomy | How much can it act without a human check? | Output must be read and signed before it counts | Output fires an order or message with no human in the loop |
| Verifiability | Can a clinician readily confirm or refute it at the point of use? | An ambient note checked against the encounter the clinician just conducted | A black-box risk score with no reasoning to check |
Read the table across a row and the sequencing logic falls out on its own. A tool that is low on all three strands is safe to deploy and easy to govern. A tool that is high on severity but keeps a human in the loop and stays verifiable is manageable, because the human check is real. The dangerous cell is high severity, high autonomy, low verifiability all at once, and any capability that sits there should never lead a roadmap, no matter how good its aggregate numbers look.
Rank risk by how easily a failure is caught, not just by how bad it would be. A bounded, verifiable error a clinician can see and correct is a manageable risk. A catastrophic error hidden inside a black box, behind a human check that is really a rubber stamp, is the risk that ends up in a deposition.
Equity Risk: The Strand That Hides Until It Is a Headline
There is one strand of risk so important, and so easy to miss in a value-versus-risk sort, that it deserves its own section: equity risk, the danger that a model performs worse for some patient populations than others, quietly widening the very disparities the organization says it wants to close. The technical name for the phenomenon is disparate performance: a model that achieves excellent accuracy on average can be meaningfully worse for a subgroup, typically the patients already underserved, because the data it learned from underrepresented them or encoded historical inequities in access and treatment.
Equity risk is treacherous precisely because it is invisible in the metrics most people look at. A vendor reports one aggregate accuracy number, and it looks great. But aggregate accuracy is a weighted average, and a model can be superb for the majority population and poor for a minority one while the single headline number stays high. To see the risk, a leader has to demand performance broken out by subgroup: by race and ethnicity, by primary language, by sex, by age, by payer, by the axes along which your patients actually differ. Two clinical measures make this concrete. Positive predictive value (PPV) is the probability that a patient the model flags as positive truly is; negative predictive value (NPV) is the probability that a patient the model clears truly is fine. A care-gap model with high PPV for insured suburban patients and low PPV for uninsured urban ones will send your outreach staff chasing false alarms in one group while missing real gaps in another, and it will do so while its aggregate numbers look perfectly respectable.
The equity strand changes the ranking in a specific way: a high-value capability whose disparate performance you have not tested is not a high-value capability. It is an untested risk wearing a high-value costume. Before a model earns its place in the ranking, someone must ask, for every subgroup that matters, does it work, and how do we know? If the vendor cannot show subgroup performance and you cannot generate it on your own population, that opportunity drops in the ranking, no matter how good its aggregate story sounds, because deploying it means gambling with the patients you are most obligated to protect. The ONC HTI-1 transparency rule gives leaders a lever here: certified health IT must now expose source attributes for a predictive decision support intervention (predictive DSI), a nutrition-label-style set of facts including how the intervention was developed and validated. A leader can and should demand to see those attributes and ask specifically what populations the validation covered.
The Value-vs-Risk Matrix
The tool that makes the ranking visible and defensible is a simple two-by-two: value on one axis, risk on the other, four quadrants that tell you not just where an opportunity sits but what to do about it. The power of the matrix is that it converts a vague argument ("this one feels important") into a positioned claim anyone on the committee can challenge or endorse. Here is the matrix, with the leadership action each quadrant implies.
| Quadrant | Value | Risk / verifiability | Examples | Leadership action |
|---|---|---|---|---|
| Start here | High | Low severity, high verifiability, human bounds it | Ambient scribe, inbox draft assistant, coding suggestion, prior-auth letter drafting, ambient handoff summary | Deploy first. Build the governance and measurement muscle here. Value is real and every error is visible and correctable at the point of use. |
| Earn the right | High | High severity, lower verifiability, autonomy tempting | Sepsis or deterioration prediction, diagnostic decision support, imaging triage, treatment recommendation | Defer until the safety machinery is proven. Require local validation, subgroup performance, real-world monitoring, and a genuine (not rubber-stamp) human-in-the-loop workflow. |
| Fill only if free | Low | Low | A niche scheduling helper, a minor UI convenience | Deploy only if it is nearly effortless. It will not move a strategic goal, so it must not spend scarce validation or integration attention. |
| Avoid | Low | High | An autonomous high-stakes tool solving a problem you do not have, a black-box score no one will act on safely | Do not build or buy. High risk with no offsetting value is the easiest no on the board and the clearest one to defend. |
Two features of this matrix separate leaders who use it well from those who use it as decoration. First, the risk axis is really about verifiability and boundedness, not just severity: the reason documentation and administrative use cases sit in the "start here" quadrant is not that errors are impossible but that the errors are visible to the clinician who was in the room and can be corrected before they reach a patient. The clinician can read the note against the encounter, catch the confabulated finding, and fix it. That is a bounded, verifiable risk. Second, the "earn the right" quadrant is not a rejection. It is a sequence. High-value, high-stakes clinical decision support is exactly what a mature program should reach, but only after it has proven, on the safe quadrant, that it can validate, monitor, train, and govern. Starting in the top-right quadrant is not ambition; it is skipping the reps that make the ambition survivable.
The Iron Rule Holds in Every Quadrant
It would be a serious misreading of this matrix to think the "start here" quadrant means "unsupervised." It does not. The iron rule of this entire program applies with equal force in all four quadrants: every AI output that touches a patient or the record must be verified, and "the AI said so" is not verification. Accountability stays human. The matrix does not tell you where verification is optional. It tells you where verification is easy and reliable versus where it is hard and fragile, and it tells you to build your program on the ground where verification is easy before you venture onto the ground where it is hard.
This distinction matters at the bedside and in the record. A physician who signs an ambient note is attesting to it; the note becomes the legal record and the standard of care, the level of care a reasonably prudent clinician would provide, is measured against what that record shows. The evolving standard now cuts both ways: a clinician can be liable for following a wrong AI recommendation and for ignoring an accurate one. A short attestation habit, a clinician noting why they agreed or disagreed with an AI suggestion, materially strengthens the record in either direction. The matrix helps a leader choose battles the clinicians can actually win at the point of care. It never excuses the organization from the verification the iron rule demands. A tool in the safe quadrant is a tool whose verification is cheap and dependable, which is exactly why it belongs first.
A Worked Example: Ranking the Four Proposals
Return to the whiteboard and the four competing proposals. Watch the value-versus-risk discipline turn a shouting match into a defensible ranking the CMIO can present without flinching.
Proposal one, the ambient scribe. Value: high, and directly aimed at the board's top-named priority, clinician burnout, with a plausible mechanism and a measurable baseline in after-hours EHR time. Risk: bounded and verifiable. The failure modes (a confabulated exam finding, wrong laterality, a dropped pertinent negative) are real, but the clinician was in the room and reads and attests to the note before it becomes the record. Every error is catchable at the point of use. Quadrant: start here. This is phase one.
Proposal two, the inbox drafting assistant. Value: solid, aimed at portal message burden and clinician experience. Risk: bounded, with one live regulatory wrinkle. California AB 3030 requires a disclaimer and instructions to contact a human when generative AI produces patient clinical communications, unless a licensed provider reviews the message. As long as a clinician reads and edits before sending, the message is reviewed, the exposure is bounded, and the error is visible. Quadrant: start here, with a disclosure control attached. This joins phase one or early phase two.
Proposal three, the sepsis prediction model. Value: potentially high, aimed at avoidable harm. Risk: high severity (a missed or false deterioration alert changes escalation), lower verifiability (the score is not something a nurse can independently confirm at a glance), and real equity exposure (deterioration models have a documented history of disparate performance across populations if trained on non-representative data). The committee asks for subgroup PPV and NPV on the local population and discovers the vendor has only aggregate numbers. Quadrant: earn the right. It goes to phase three, gated behind local validation, subgroup testing, real-world monitoring, and a designed human-in-the-loop workflow that a bedside nurse can actually operate without automation bias. Not a no. A not yet.
Proposal four, the autonomous imaging triage tool. Value: modest for this system, which already has adequate overnight radiology coverage and no throughput crisis in imaging. Risk: high, because the vendor is pitching autonomy (reading films without a radiologist in the loop overnight) and the failure mode is a missed finding on a study no human reviewed until morning. High risk, low local value, high autonomy. Quadrant: avoid, or at minimum defer indefinitely. This is the easiest no of the four to defend, precisely because the value is low and the risk is high; there is no painful tradeoff to litigate.
The CMIO now presents not four adjectives but a ranked plan: fund the scribe and the inbox assistant now (high value, bounded and verifiable risk), place the sepsis model in a gated phase-three slot with explicit validation and equity requirements, and decline the imaging tool with a one-sentence rationale any board member can repeat. When the safety committee asks why the sepsis model, the most clinically exciting proposal, is not first, the answer is ready: because we have not yet proven we can validate its subgroup performance or verify its outputs at the bedside, and deploying a high-severity, low-verifiability model before we can do those things would be gambling with our most vulnerable patients. That sentence is the whole lesson. It is what a defensible ranking sounds like out loud.
Key Takeaways
- Ranking, not choosing, is the leadership act. Vendors supply endless good opportunities; the job is to put them in a defensible order and be able to explain the order to a board, a safety committee, and a harmed family.
- Value and risk do not share a unit, so never sort purely by ROI. In healthcare the worst case of a wrong bet is a harmed patient, so risk belongs on its own axis, not folded into a return number.
- Value is broader than dollars: clinical benefit, operational and financial benefit, clinician experience, and strategic fit. Insist on a baseline and target before ranking, and discount vendor value by your own ability to realize it.
- Rank risk by verifiability, not just severity. The strand that should drive sequencing hardest is how easily a clinician can catch a failure at the point of use, because a human-in-the-loop you cannot verify is a rubber stamp powered by automation bias.
- Equity risk hides behind aggregate accuracy. Demand disparate-performance testing (subgroup PPV and NPV) and the ONC predictive DSI source attributes; an untested-for-equity capability is an untested risk, not a high-value one, no matter its headline number.
- Use the value-versus-risk matrix: start where value is high and risk is bounded and verifiable (documentation and administrative use), earn the right to high-stakes clinical decision support later, fill low-value low-risk slots only if effortless, and avoid low-value high-risk tools outright.
- Starting in the safe quadrant is not timidity; it builds the validation, monitoring, training, and governance muscle that high-stakes clinical use requires. Starting in the top-right quadrant skips the reps that make ambition survivable.
- The iron rule holds in every quadrant: every AI output touching a patient or the record must be verified, "the AI said so" is never verification, and accountability stays human. The matrix tells you where verification is easy versus fragile, never where it is optional.
Skill.re