AI for Healthcare & Clinical Practice
Strategic · M8 · lesson 8 of 23 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
From Point Tools to a Clinical AI Roadmap
📖
now learning

From Point Tools to a Clinical AI Roadmap

15 min

The slide was titled "Our AI Journey," and it made the CMIO wince. Nineteen logos scattered across it, a scribe here, a sepsis model there, an imaging triage tool, three inbox pilots, a prior-auth bot the revenue cycle team bought without telling anyone, and a chatbot from an innovation grant that had run out of money in the spring. Each had a champion. None had a number. When the board chair asked the simple question, "So what did all of this do for our patients and our margin?", the room went quiet. That silence is the sound of a health system that bought a lot of AI and never built a strategy. This lesson is about the difference between the two, and about how a leader turns a pile of point tools into a roadmap that a board, a survey team, and a bedside nurse can all recognize as coherent.

Pilot Purgatory and How Systems Get There

Most health systems do not decide to accumulate scattered AI. They drift into it. A well-meaning department buys a tool that solves a real local pain, a service line wins a grant, a vendor lands a friendly pilot with a sympathetic physician champion, and the innovation office greenlights three experiments to "learn." Each decision is defensible in isolation. The sum is not. The industry has a name for the resulting state: pilot purgatory, the condition where an organization is perpetually piloting and never scaling, where dozens of small proofs of concept run in parallel, none reaching the size or the integration where it changes outcomes, and none quite dying either. The pilots consume attention, budget, security review, and clinician goodwill, and they return anecdotes instead of results.

Pilot purgatory is expensive in ways that do not show up on any single line item. Every pilot demands a privacy and security review, a business associate agreement, an EHR integration conversation, and clinician time to train and give feedback. When ten pilots run at once, the governance committee drowns, the security team triages by whoever shouts loudest, and the clinicians develop pilot fatigue, the weary sense that another shiny tool will appear next quarter and vanish the quarter after. Worse, because none of the pilots is tied to a defined outcome with a baseline and a target, none can be honestly declared a success or a failure. They simply persist, a standing tax on the organization's finite capacity to absorb change. A leader who inherits this landscape has not inherited a portfolio. They have inherited a mess with momentum.

The root cause is almost never bad tools. It is the absence of a shared answer to a prior question: what are we actually trying to accomplish, in what order, and how will we know it worked? Without that answer, every individual purchase looks reasonable and the aggregate looks like the wince-inducing slide. The move from point tools to a roadmap is the move from answering "is this tool good?" one gadget at a time to answering "what is our clinical AI strategy?" once, and then judging every tool against it.

A Roadmap, Not a Shopping List

The central discipline of this lesson can be compressed into one contrast: a strategy is a roadmap, not a shopping list. A shopping list is a set of things you intend to acquire. It answers the question "what do we want to buy?" A roadmap answers a harder and more useful set of questions: what outcomes are we pursuing, which capabilities get us there, in what sequence, dependent on what readiness, measured by what metrics, and stopped by what guardrails? A shopping list is organized by vendor and by whoever is excited. A roadmap is organized by clinical and financial outcome and by the order in which the organization can actually absorb and prove each step.

The distinction matters because AI in a health system is not a consumer purchase; it is a change to how care is delivered and documented inside a regulated, audited, human-accountable environment. A new ambient scribe changes the clinical note, which is the legal record. A predictive model changes what clinicians attend to and when. A summary tool changes what information a hospitalist trusts at three in the morning. None of these is a widget you plug in and forget. Each reshapes a workflow, creates new failure modes, and carries a patient-safety and liability tail. Treating that as a shopping decision is a category error, and it is exactly the error that produces nineteen logos and no numbers.

A shopping list asks what you want to buy. A roadmap asks what outcomes you are chasing, in what order your organization can prove them, and what would make you stop. Nineteen logos is a shopping list. A sequenced set of outcomes with baselines, targets, and guardrails is a strategy.

The good news is that building a roadmap does not require throwing away the point tools you already have. It requires reorganizing them, and the future ones, around outcomes and sequence. That reorganization has three moves: take an honest inventory of what you are already running, align each item to a strategic goal and prune what aligns to none, and then sequence what survives by value and readiness. The rest of this lesson walks through each.

Move One: Inventory What You Actually Run

You cannot sequence what you cannot see. The first move is a candid inventory of every AI capability already touching your patients, your clinicians, or your record, including the ones nobody told you about. In most systems this exercise is genuinely surprising, because AI enters through many doors: the EHR vendor ships predictive models that light up by default, a department buys a niche tool on a departmental budget, a grant funds a pilot that never got a governance sign-off, and a consumer-grade tool creeps into a clinic because it was free and helpful. A 2026 estimate that roughly three-quarters of health systems run at least one AI application understates the reality for most large systems, which run many, often uncounted. Verify that figure for your own house rather than repeating it, because the number that matters is not the industry average; it is yours.

A useful inventory captures, for each tool, a short set of facts a leader can act on: what it does, which patients and clinicians it touches, who owns it, whether it has a business associate agreement and a completed security review, whether it is an FDA-regulated device or a predictive decision support intervention subject to the ONC transparency criteria, what outcome it was supposed to improve, and whether anyone is measuring that outcome. The last two columns are where most tools fail the test. A tool with no defined outcome and no measurement is not a strategic asset; it is an unmanaged risk with a login page. The inventory is not busywork. It is the map you cannot build the roadmap without, and it frequently pays for itself immediately by surfacing a tool with PHI flowing to a vendor that never signed a BAA, or a default-on predictive model no one validated on the local population.

The columns are not arbitrary. Each one maps to a question a compliance officer, a surveyor, or a plaintiff's attorney can ask you next year, and a blank cell is an answer you do not want to give under oath. A practical inventory row looks like the table below, filled in for every capability, no matter how small or how someone insists it does not count.

Inventory fieldWhat it recordsWhy a blank cell is dangerous
Function and workflowWhat the tool does and where it sits in the clinical or administrative flowIf you cannot describe the workflow, you cannot describe the failure mode that reaches a patient.
Population touchedWhich patients and which clinicians it affects, and at what volumeA tool touching thousands of encounters is a different risk than a single-clinic pilot; scale sets priority.
OwnerThe named accountable person, not a departmentOwnerless tools are the ones no one turns off, updates, or monitors.
BAA and security reviewWhether PHI is covered by a signed BAA and whether security cleared itPHI to a vendor with no BAA is a HIPAA exposure the system owns regardless of who bought the tool.
Regulatory statusFDA-authorized device and intended use, or predictive DSI with ONC source attributes, or neitherOff-label device use and unexamined predictive DSI are both survey and liability findings waiting to happen.
Target outcomeThe strategic goal it was meant to moveNo named outcome means the tool cannot be defended, prioritized, or judged.
MeasurementThe baseline, target, and who is watching the metricWithout measurement you can never declare success or failure, so the tool drifts forever.

Run this against a real system and the pattern is consistent: a handful of tools have every cell filled, most have the last two blank, and a few have blanks in the BAA or regulatory column that trigger action the same week. That distribution is itself diagnostic. It tells you the organization has been buying capability faster than it has been building the discipline to account for it, which is exactly the gap the roadmap closes.

The Shadow AI Problem

The inventory almost always uncovers shadow AI: capabilities running without governance knowing. This is not usually malice; it is the natural result of AI being cheap, useful, and easy to adopt at the edge. A resident pastes a de-identified-looking case into a consumer chatbot to draft a plan. A clinic manager subscribes to a scheduling assistant. A service line runs a vendor pilot on a handshake. Each is a real exposure, because obligations do not disappear just because leadership did not authorize the tool. The health system still owns the HIPAA breach if PHI leaked, still owns the standard-of-care question if a clinician acted on an unvalidated output, and still owns the survey finding if an accreditor asks how AI is governed and the honest answer is "we are not sure what we have." Surfacing shadow AI is not a witch hunt. It is the precondition for governing the real, current footprint rather than an imagined tidy one.

Move Two: Align Each Tool to a Strategic Goal

An inventory tells you what you have. Alignment tells you what deserves to survive. The second move is to lay the inventory against the organization's actual strategic priorities, the ones the board and the executive team already care about, and ask of each tool a blunt question: which strategic goal does this advance, and by how much? The strategic goals are not AI goals. They are the goals the system has anyway: reduce clinician burnout and documentation burden, improve throughput and length of stay, close care gaps and improve quality measures, reduce avoidable harm, protect and grow margin, advance equity. AI is not a goal. It is a means, and a means earns its place only by moving one of those needles in a way you can demonstrate.

This is the step where the shopping list gets shorter, and it should. A tool that advances a named strategic goal, with a plausible mechanism and a way to measure it, stays as a candidate. A tool that advances no strategic goal, that exists because it was interesting or because a vendor was persuasive, is a candidate for pruning even if it works, because organizational attention is the scarcest resource in the building and every tool spends it. Alignment also reveals gaps: strategic goals with no tool pointed at them, which is where the roadmap's future investments belong. The output of this move is a much smaller, much clearer set: the AI capabilities, current and proposed, that each connect to something the organization is already committed to achieving.

A subtle but critical part of alignment is naming the metric before you sequence. For each surviving candidate, write down the outcome it targets, the current baseline, the target, and how you will measure it. "Reduce documentation burden" becomes "reduce after-hours EHR time for primary care from the current baseline to a defined target within two quarters, measured by the EHR's own time logs." A 2025 multi-system study reporting that burnout fell from roughly 52 percent to roughly 39 percent within thirty days on an ambient documentation tool is a useful anchor for what is plausible, but it is a number to verify against your own baseline and population, not a promise to paste into a business case. Without a baseline and a metric, you cannot sequence honestly and you cannot ever declare victory, which is how tools drift back into purgatory.

Governance Is the Connective Tissue, Not a Speed Bump

There is a temptation to treat governance as the thing that slows the roadmap down, a committee that clinicians dread and vendors route around. That framing is exactly backward. In a coherent roadmap, governance is the connective tissue that lets the point tools become a portfolio instead of a pile. It is the single process every capability passes through, so that the inventory stays current, the alignment stays honest, the validation gets done, and the monitoring keeps running after the launch excitement fades. Without that shared process, every tool is governed differently or not at all, which is precisely how shadow AI accumulates and how a system ends up unable to answer a surveyor. The Joint Commission and CHAI guidance released in September 2025 names a designated governance structure as one of its foundational elements for exactly this reason: a named body, a repeatable intake, and a validation and monitoring expectation that applies to every tool, not just the ones someone remembered to bring to the committee.

The practical form of connective tissue is a single intake gate. Any new AI capability, whether a vendor pitch, a grant pilot, or an EHR feature about to be switched on, enters through one door that asks the same questions in the same order: what outcome does it target, what is the baseline, is there a BAA and a security review, what is its regulatory status, what populations was it validated on, and who will monitor it after go-live. The gate does not have to be slow. Bounded, low-risk documentation tools can pass in days; a high-stakes predictive model earns a longer look. What the gate guarantees is that nothing important reaches a patient without the roadmap's discipline touching it first. That is how governance stops being a speed bump and becomes the mechanism that makes the whole sequence trustworthy.

Move Three: Sequence by Value and Readiness

Now the roadmap takes shape. Sequencing orders the aligned candidates along two axes at once: value, meaning how much the capability advances a strategic goal, and readiness, meaning how prepared the organization actually is to deploy it safely and prove it. Value without readiness is the trap that produces failed high-stakes projects; readiness without value is the trap that produces safe, well-run tools that no one needed. The art is to sequence so that early wins are both valuable and achievable, building the governance muscle, the clinician trust, and the measurement discipline that the harder, higher-stakes projects will require.

Readiness is concrete, not vibes. It includes data readiness (is the data available, representative, and clean enough for the tool to work on your population), technical readiness (can it integrate with the EHR without a heroic effort), governance readiness (do you have the committee, the validation process, and the monitoring to deploy it responsibly), and workforce readiness (are the clinicians who must use it trained, consulted, and willing). A capability can be enormously valuable and completely unready, and pretending otherwise is how a system lands a flagship predictive project that fails not because the model was bad but because the organization could not absorb it. Sequencing respects readiness so that the roadmap is a series of steps the organization can actually take, in an order that compounds.

Because readiness has four dimensions, a single tool can be green on three and red on one, and the red one governs. A deterioration model may have clean data and an easy integration, but if you have no committee that can validate it and no monitoring to watch it drift, its governance readiness is red and the whole capability is not ready, whatever the other columns say. The discipline is to score each dimension explicitly rather than average them into a comforting middle, because averaging hides the one deficit that will sink the deployment. The table below turns readiness from an argument into a checklist.

Readiness dimensionThe question it answersA red flag that blocks sequencing
DataIs the data available, representative of your population, and clean enough?The model was trained on a population unlike yours and cannot be revalidated locally.
TechnicalCan it integrate with the EHR without a heroic, multi-quarter effort?Integration requires custom interfaces no one has budgeted or staffed.
GovernanceIs there a committee, a validation process, and monitoring to deploy responsibly?No body owns validation, and nobody is assigned to watch performance after go-live.
WorkforceAre the clinicians who must use it trained, consulted, and willing?Frontline clinicians were never consulted and see the tool as one more imposed pilot.

Value and readiness together produce a simple sequencing logic that a board can follow. High value and high readiness goes first, because it delivers and it is achievable. High value and low readiness is not abandoned; it is scheduled later and paired with an explicit plan to build the missing readiness, so the roadmap says not just what but when and what has to be true first. Low value is deprioritized regardless of readiness, because even an easy deployment spends attention that a higher-value goal needs. The point of naming readiness is not to say no to hard things; it is to sequence hard things behind the specific capabilities that make them survivable.

The general shape most systems converge on, and the reason the next lesson exists, is to sequence the bounded-risk, high-readiness, clearly measurable capabilities first, documentation and administrative support where the value is real and the failure modes are visible and verifiable, and to reach the high-stakes clinical-decision uses later, once the governance and monitoring machinery is proven. That is not timidity; it is how you earn the organizational credibility and the safety infrastructure that the hard cases demand. A roadmap that starts with an ambient scribe and builds toward diagnostic decision support is sequenced by value and readiness. A roadmap that starts with autonomous diagnosis because it sounded impressive is sequenced by ego, and it will stall.

A Worked Example: Turning the Wince Slide Into a Roadmap

Return to the CMIO and the nineteen logos. Watch the three moves convert the mess into a strategy. The inventory takes six weeks and turns up twenty-three AI capabilities, not nineteen, because it surfaces four shadow tools: a resident-favored chatbot with no BAA, a clinic scheduling assistant, a default-on EHR readmission model no one validated locally, and a coding-suggestion feature quietly upcoding in the background. Two of the four are shut off within a week, one for the PHI exposure and one for the coding risk, and that alone justifies the exercise to the compliance officer.

Alignment then maps the survivors against the system's five stated priorities: burnout, throughput, quality, harm reduction, and margin. Eight capabilities align cleanly, six align weakly, and nine align to nothing anyone can name. The nine are not all switched off, but they lose their claim on scarce integration and governance attention, and three are formally retired. Crucially, alignment reveals that the system's top-stated priority, clinician burnout, has exactly one tool pointed at it, an under-deployed ambient scribe, while three redundant tools crowd a low-priority goal. The strategy was upside down, and no one could see it until the inventory and alignment made it visible.

Sequencing then produces a roadmap with a spine. Phase one: fully deploy and measure the ambient scribe against a documented burnout and after-hours-EHR baseline, because it is high value, high readiness, and bounded in risk, and it will build the measurement and governance habits the later phases need. Phase two: add AI chart summarization and inbox drafting, again bounded and measurable, reusing the governance process phase one proved. Phase three, only after the machinery is trusted: pilot a predictive deterioration model with a formal validation on the local population, real-world monitoring, and a defined human-in-the-loop workflow, because it is high value but high stakes and low current readiness. The board now sees not nineteen logos but three phases, each tied to a named outcome, a baseline, a target, and a guardrail. When the chair asks what the AI did for patients and margin, the CMIO has a slide that answers.

Notice what the roadmap did that the shopping list could not. It made the invisible visible (shadow AI and the upside-down priorities), it created a basis for saying no (alignment retired tools that worked but did not matter), and it sequenced the surviving work so that early, safe wins funded the credibility for the hard, high-stakes bets later. That is the whole job. A roadmap is not a longer shopping list. It is a different kind of object, one that a board can fund, a survey team can respect, and a clinician can trust, because it is organized around outcomes and safety rather than logos and enthusiasm.

Key Takeaways

  • Scattered pilots are not a strategy. Pilot purgatory, perpetually piloting and never scaling, is a standing tax on governance, security, and clinician goodwill, and it returns anecdotes instead of outcomes because no pilot is tied to a baseline and a target.
  • A strategy is a roadmap, not a shopping list. A shopping list asks what you want to buy; a roadmap asks what outcomes you are chasing, in what order your organization can prove them, and what guardrails would make you stop.
  • Move one is an honest inventory of every AI capability already touching patients, clinicians, or the record, including shadow AI that governance never authorized, because obligations do not disappear just because leadership did not approve the tool.
  • Move two aligns each tool to a strategic goal the organization already holds, burnout, throughput, quality, harm, equity, margin, and prunes what aligns to nothing, because organizational attention is the scarcest resource and every tool spends it.
  • Name the metric before you sequence: for each survivor, write the outcome, the current baseline, the target, and the measurement, so you can honestly declare success or failure instead of drifting back into purgatory.
  • Move three sequences the survivors by value and readiness together, so early wins are both valuable and achievable and build the governance, trust, and measurement muscle the hard cases need.
  • Sequence bounded-risk, high-readiness, measurable capabilities first (documentation and administrative support) and reach high-stakes clinical-decision uses later, once the safety machinery is proven; that is prudence, not timidity.
  • Treat every adoption and ROI figure, including a reported burnout drop from roughly 52 percent to 39 percent on ambient documentation or the claim that about three-quarters of systems run AI, as a number to verify against your own baseline and population, not a promise to repeat.