Funding and the ROI Story for Clinical AI
The CFO had a red pen and a single question. The clinical AI lead had walked into the budget meeting with a beautiful slide: an ambient documentation tool, a projected two hundred thousand dollars a year in "physician time saved," and a hockey-stick chart. The CFO read it twice, then set the pen down and said, "Saved for whom? If a doctor finishes charting twenty minutes earlier and goes home, that is a wonderful thing for the doctor and I am glad for her, but it does not put a dollar back in this system's account. Show me the money I can actually count, and show me the money I do not lose. Then we will talk." The room went quiet. That silence is the sound of an ROI story that counted saved minutes and forgot everything else. This lesson is about building the case that survives a CFO with a red pen: the dual-axis case that pairs time and cost saved with risk and harm avoided, and never pretends efficiency alone is the whole story.
Why the Efficiency-Only Pitch Dies in the Budget Meeting
Most clinical AI business cases are built on a single axis: efficiency. The tool saves clinician time, so multiply the minutes saved by an hourly cost, annualize it, and there is your return. It is a clean story and it is usually wrong, or at least dangerously incomplete. The trouble is that saved minutes are what a CFO calls soft dollars: value that is real to the person experiencing it but does not convert into cash the organization can spend unless something structural changes. If a physician saves twenty minutes of documentation and uses it to leave on time, morale improves and burnout eases, both of which matter enormously, but the general ledger does not move. Soft-dollar savings become hard dollars (cash that actually shows up in a budget line) only when the saved capacity is deliberately converted: the clinic adds visits into the freed time, a scribe contract is cancelled, a vacancy goes unfilled because the remaining staff can now absorb the work, or overtime falls. An efficiency pitch that never explains the conversion is asking the CFO to fund a feeling.
It helps to make the conversion explicit as a mechanism, because that is the exact link a CFO tests. The table below shows how a soft signal turns into a hard-dollar line, and, just as importantly, the caveat that limits each one. Read it as the bridge between what clinicians feel and what finance can book.
| Soft signal (real but not cash) | Conversion mechanism | Hard-dollar line it becomes | Caveat that limits the claim |
|---|---|---|---|
| Twenty minutes less charting per clinic session | Add a defined number of incremental visits into the freed slot | Contribution margin on the new visits | Only real if patient demand and a fillable schedule exist; operations sizes it, not the vendor |
| Physicians no longer need a per-encounter human scribe | Cancel or do not renew the scribe or transcription contract | Eliminated contract spend | Only counts once, and only for scribes the tool genuinely replaces |
| Remaining staff can absorb the documentation workload | Leave an open FTE or backfill vacancy unfilled | Avoided salary and benefits on the unfilled line | Requires that service levels hold without the position; do not assume it |
| Lower after-hours EHR time and less weekend catch-up | Reduce paid overtime and expensive locum or agency coverage | Reduced overtime and temporary-staffing spend | Must be measured against a real baseline, not asserted |
| Improved burnout or intent-to-stay scores in a strained group | Prevent one to two physician departures per year | Avoided recruitment, onboarding, and locum-gap replacement cost | Claim only the measured or credibly modeled retention delta, never the whole burnout drop |
There is a second, deeper reason the efficiency-only pitch fails, and it is the one this lesson exists to fix. Efficiency is only half of what clinical AI does for a health system. The other half is risk avoided: the harm that does not happen, the rework that is not required, the audit that is not lost, the malpractice claim that is not filed, the readmission penalty that is not levied. These are enormous numbers, often larger than the time savings, and an ROI story that ignores them is not conservative; it is incomplete in a way that both undersells the tool and, worse, hides the safety and change-management costs that make the whole case honest. A leader who walks in with only the efficiency axis has brought half a case and will be treated accordingly.
An ROI story that counts only saved minutes and ignores avoided harm is only half the case. Efficiency is one axis. Risk avoided is the other. A case built on one axis is not conservative, it is incomplete.
The discipline this lesson teaches is to build the case on two axes at once, to quantify both honestly, and to subtract the real costs that vendors leave off their decks. Do that and the CFO's red pen goes back in the drawer, because you have answered both halves of the question: the money you can count, and the money you do not lose.
The Dual-Axis Frame: Time and Cost Saved, Risk and Harm Avoided
Picture the case as a table with two columns. The left column is time and cost saved: the efficiency axis, the money that flows in or the cost that flows out because the work gets done faster or with fewer resources. The right column is risk and harm avoided: the money that never leaves the building because a bad thing did not happen. A complete business case fills both columns and then subtracts a third thing, the cost of doing the deployment safely and getting people to use it, which we will come to. Most vendor decks fill the left column lavishly, leave the right column blank, and pretend the third thing does not exist. Your job as a leader is to fill all three.
The two axes are not interchangeable, and it matters which kind of dollar you are claiming. Time and cost saved is usually a probabilistic-but-steady stream: if the tool genuinely reduces after-hours EHR time, that reduction recurs every day across every user, and you can measure it directly from the EHR's own time logs. Risk and harm avoided is a different beast: it is a distribution of low-frequency, high-severity events, a single malpractice claim, a single lost coding audit, a single serious safety event. You cannot point to the specific claim that did not happen, but you can estimate the expected value: the probability of the event multiplied by its cost, reduced by the fraction the tool plausibly prevents. Finance does exactly this every day; it is how insurance, reserves, and risk-adjusted capital work. Speaking in expected-value terms is how you make the right column legible to a CFO who lives in that language.
The Two Columns, Side by Side
| Axis 1: Time and Cost Saved (efficiency) | Axis 2: Risk and Harm Avoided (protection) |
|---|---|
| Reduced after-hours EHR time ("pajama time"), the documentation clinicians finish at home at night | Avoided patient harm from a missed finding, and the associated care and liability cost |
| Higher throughput: more visits or procedures in the same clinic hours | Reduced malpractice exposure from a more complete, defensible, contemporaneous record |
| Shorter length of stay when a tool speeds a decision or a discharge step | Avoided rework: notes, orders, and claims that do not have to be redone because they were right the first time |
| Cancelled or reduced human scribe and transcription contracts | Fewer denied claims and lost RAC or coding audits (a RAC audit is a Recovery Audit Contractor review that can claw back payment for undocumented or miscoded care) |
| Avoided overtime and reduced reliance on expensive temporary staffing | Avoided readmission and quality-measure penalties that reduce reimbursement |
| Lower recruitment and turnover cost as burnout eases and clinicians stay | Fewer survey findings and the cost of remediation when an accreditor or regulator inspects |
Read the table as a leader, not an accountant. The left column is where the tool pays for itself in visible, countable ways. The right column is where the tool protects the enterprise from the catastrophic and the corrosive: the single event that costs millions, and the steady drip of rework and denials that costs a fortune in aggregate. A case that presents both columns tells the CFO the truth: this investment both earns and protects. A case that presents only the left column invites the exact question that killed the ambient-scribe slide.
Quantifying the Efficiency Axis Honestly
The efficiency axis is the easier of the two to quantify, which is precisely why it gets abused. The abuse takes a predictable form: multiply an optimistic minutes-saved figure by a fully loaded hourly rate, apply it to every clinician in the building, and annualize. The result is a spectacular number that no CFO believes, because it assumes universal adoption, sustained savings, and, most importantly, that every saved minute converts to cash. Honest quantification is more modest and far more persuasive.
Start with the metric the EHR already measures. After-hours EHR time, the charting clinicians do outside scheduled hours, is logged natively by every major EHR and is the cleanest proxy for documentation burden. Establish a baseline before deployment, measure the same cohort after, and claim only the delta you can see. If a documentation tool moves a primary-care group's after-hours time down meaningfully, that is real and defensible. Then, and this is the step vendors skip, state the conversion mechanism. Does the freed time become additional visits (hard revenue, if the demand and the schedule exist to fill them)? Does it let a clinic run without a per-visit scribe (a hard cost avoided)? Does it reduce the turnover that was costing you a fortune in locum coverage and recruitment? Each conversion is a real hard-dollar line; the raw minutes are not. A business case that says "we saved forty thousand clinician-hours" is soft. A case that says "we cancelled a scribe contract worth X, added Y incremental visits, and reduced turnover by Z" is hard, and it is a fraction of the vaunted forty thousand hours, which is exactly why it is credible.
Throughput and length of stay belong here too, but with discipline. If a tool genuinely lets a department see more patients or discharge a day sooner, the value is real and often large, because a freed inpatient bed-day or an incremental procedure carries substantial contribution margin. But throughput gains are notoriously easy to overclaim, because they depend on downstream capacity you may not have: an extra discharge only helps if there is a patient waiting for the bed, and an extra clinic slot only helps if it gets filled. Claim throughput only where the downstream demand and capacity actually exist to absorb it, and let operations, not the vendor, size the number.
The Burnout and Retention Line: Soft Signal, Hard Consequence
Burnout reduction is where the two axes start to blur, and it deserves careful handling because it is both the most emotionally compelling and the most easily hand-waved part of the case. The emotional story is easy: documentation burden is the top driver of physician burnout, and tools that lift it make clinicians' lives better. A widely cited 2025 multi-system study reported that clinician burnout fell from roughly 52 percent to roughly 39 percent within thirty days of adopting an ambient documentation tool. Treat that as a number to verify against your own baseline and population, not a figure to paste into a business case as if it were your own result. It is a plausibility anchor, not a promise.
The hard consequence, the part a CFO can bank, is retention. Physician turnover is astonishingly expensive once you count recruitment, onboarding, lost productivity during ramp, and locum coverage during the gap; the fully loaded cost of replacing a single physician runs into hundreds of thousands of dollars and, for some specialties, well beyond. If a documentation tool measurably improves retention in a strained group, even a small reduction in turnover pays for a great deal of software. The move is to connect the soft signal (a burnout score, an engagement survey, a stated intent to leave) to the hard consequence (an avoided replacement cost) with a defensible, conservative estimate. Do not claim the whole burnout drop as dollars. Claim the retention delta you can measure or credibly model, and present the burnout improvement as the mechanism and as a workforce-stability benefit in its own right.
Quantifying the Risk-Avoided Axis Without Fantasy
The risk-avoided axis is where the biggest numbers and the biggest temptations to fabricate both live. The right way to quantify it is the expected-value method finance already trusts: for each category of avoidable bad event, estimate the annual frequency, the cost per event, and the fraction the tool plausibly prevents, then multiply. The output is not a claim that a specific disaster was averted; it is a risk-adjusted expected saving, stated as such, with the assumptions visible so the CFO can push on them. Visible assumptions are a feature, not a weakness, because they signal that you are estimating honestly rather than conjuring.
Consider the categories. Avoided rework is the most tractable: notes returned for clarification, claims denied for insufficient documentation, orders that had to be corrected. These have frequencies and costs your revenue-cycle and quality teams already track, and a tool that improves first-pass accuracy reduces them measurably. Coding and audit exposure is next: a more complete, better-documented record supports appropriate coding and survives a RAC or payer audit that would otherwise claw back payment. Note the double edge here, because it is a governance point as much as a financial one: a tool that improves documentation legitimately reduces audit loss, but a tool that quietly inflates coding to boost that number is manufacturing fraud exposure, not avoiding it. The right column must be built on defensible documentation, never on upcoding dressed as efficiency.
Avoided harm and malpractice exposure is the largest and the hardest. A single serious safety event carries direct care costs, potential penalties, reputational damage, and a malpractice tail that can run into the millions. You cannot honestly claim a tool eliminates these, and you should never try. What you can do is estimate conservatively: if a tool improves the completeness of the record, or surfaces a deteriorating patient earlier, or reduces the omissions that drive a class of claims, apply a modest prevented-fraction to the expected annual cost of that class of events. Even a small fraction of a large, low-frequency, high-severity cost is a meaningful number, and stating it as an expected value rather than a certainty is what keeps the case honest. The iron rule of the program applies with full force here: every AI output that touches a patient must still be verified by a human, and accountability stays human. The risk-avoided axis is not a claim that AI made care safe by itself; it is a claim that a well-governed, human-verified AI workflow reduces a measurable class of expensive failures.
The worksheet below shows the shape of the calculation for a mid-sized group. Every number in it is illustrative only, a structure to verify against your own data, never a figure to repeat. Fill each cell from your revenue-cycle, quality, and risk-management records, then let the CFO push on the prevented-fraction column, which is where the honesty lives. Notice that the expected annual saving is frequency x cost per event x prevented fraction, so a conservative prevented fraction keeps a large gross number from becoming a fantasy.
| Event category | Annual frequency (illustrative) | Cost per event (illustrative) | Prevented fraction (illustrative) | Expected annual saving (verify, do not repeat) |
|---|---|---|---|---|
| Documentation-driven claim denials and rework | 1,200 | $180 | 0.15 | $32,400 |
| Coding-audit clawback on defensibly documented care | 40 | $3,500 | 0.20 | $28,000 |
| Readmission or quality-measure penalty exposure | 1 (aggregate) | $250,000 | 0.05 | $12,500 |
| Documentation-omission malpractice claim class | 0.5 (once per two years) | $900,000 | 0.04 | $18,000 |
The point of the worksheet is not the total. It is the discipline: name the category, source the frequency and cost from your own data, apply a prevented fraction you can defend out loud, and label the result as risk-adjusted expected value. A CFO who sees a 0.04 prevented fraction on a malpractice class understands immediately that you are not claiming the tool eliminates lawsuits; you are claiming a small, defensible reduction in a large, rare cost. That restraint is what makes the whole right column believable.
Subtract the Costs Vendors Leave Off the Deck
A dual-axis case is still a fantasy if it counts only benefits. The third move, and the one that separates a credible leader from an enthusiastic one, is to subtract the full cost of doing this safely. Vendors quote a license fee. The real number is the total cost of ownership: the sum of every cost required to deploy, govern, integrate, and sustain the tool over its life, most of which never appears on the vendor's price sheet.
The line items that get left off are predictable and large. Change management, the work of training clinicians, redesigning workflows, managing the dip in productivity while people learn, and sustaining adoption after the novelty fades, is frequently the single biggest cost of a clinical AI deployment and is almost never in the vendor's ROI deck. A tool that no one adopts returns zero regardless of how good it is, and adoption is bought with change-management effort, not license fees. Safety and governance cost is the next omission: the validation of the tool on your own population before go-live, the ongoing monitoring for drift and disparate performance, the committee time, the incident-response process, and the human verification step that the iron rule requires. That verification is not free; it is clinician attention, and a case that assumes the AI output is trusted blindly has both understated the cost and violated the safety principle. Integration and IT, security review, the business associate agreement, EHR interface work, and internal staff time round out the picture.
Laying the line items out as a total-cost-of-ownership ledger makes the omissions obvious. A vendor deck typically shows the first row and stops. The rows below it are where the real money and the real safety obligations live, and a case that leaves them blank is not cheaper, it is dishonest.
| Line item | What it actually covers | Where it usually hides |
|---|---|---|
| License and subscription | The per-user or per-encounter fee the vendor quotes | On the vendor slide, and often the only number shown |
| Change management | Training, workflow redesign, the productivity dip during ramp, and sustaining adoption after the novelty fades | Almost never in the deck; frequently the single biggest line |
| Validation before go-live | Testing the tool on your own note types, specialties, and patient mix, including checking for disparate performance | Assumed to be the vendor's job; it is yours |
| Ongoing monitoring | Watching for model drift and disparate performance after deployment, with a defined cadence and owner | Treated as optional; it is a recurring safety cost |
| Integration, IT, and BAA | EHR interface work, security review, the business associate agreement, and internal staff time | Underscoped as a one-time IT task |
| Human verification time | The clinician attention required to review and attest to every AI output before it enters the legal record | Silently assumed to be zero, which both understates cost and violates the iron rule |
An ROI story that ignores safety cost and change-management cost is not a business case, it is a sales brochure. The verification step the iron rule demands is a real cost, and a case that pretends the AI is trusted blindly has hidden both a cost and a liability.
There is a governance reason, not just a financial one, to insist on these lines. A case that omits the safety and verification cost is implicitly promising that the AI will be trusted without checking, which is the precise setup for automation bias and a patient-safety event. Putting the verification cost on the ledger does two things at once: it makes the ROI honest, and it forces the organization to fund the human check that keeps accountability where it belongs. The most dangerous business case is not the one that is too expensive; it is the one that looks cheap because it quietly assumed the humans would stop verifying.
A Worked Example: Building the Dual-Axis Case for Ambient Documentation
Return to the CFO and the red pen. Watch the clinical AI lead rebuild the case the right way, for a proposed ambient documentation deployment across a two-hundred-physician primary-care and specialty group. The original slide claimed two hundred thousand dollars in "time saved" and nothing else. The rebuilt case has three parts, and it is smaller in its efficiency claim, larger overall, and vastly more credible.
Axis one, time and cost saved, stated as hard dollars only. Rather than monetizing raw minutes, the lead commits to three convertible lines. First, the group cancels a per-encounter human scribe pilot the tool replaces, a documented contract saving. Second, operations agrees that a defined subset of physicians will add a modest number of incremental visits per week into freed capacity where real patient demand exists, and finance books only the contribution margin on visits it is confident will fill. Third, the lead models a conservative reduction in physician turnover in the two most strained specialties, converting a measured improvement in intent-to-stay into one to two avoided physician replacements a year at a fully loaded replacement cost the CFO's own team supplies. The efficiency axis now totals less than the original two hundred thousand dollars in raw "time saved," but every dollar is a line the CFO can trace.
Axis two, risk and harm avoided, stated as expected value. Working with revenue cycle and quality, the lead builds three risk lines. First, avoided documentation-driven claim denials and rework, sized from the group's actual denial rate and the fraction plausibly improved by more complete, contemporaneous notes. Second, reduced coding-audit exposure, built strictly on better documentation supporting codes that were always appropriate, never on inflating them, with a conservative prevented-loss estimate. Third, a deliberately modest malpractice-exposure line: applying a small prevented-fraction to the expected annual cost of documentation-omission claims for a group this size, stated explicitly as a risk-adjusted expected value with the assumptions on the slide. This axis, even estimated conservatively, rivals or exceeds the efficiency axis, which is exactly the point the original slide missed.
The subtraction, total cost of ownership. Against those benefits the lead subtracts the full cost: the license, yes, but also a substantial change-management budget for training and workflow redesign and the productivity dip during ramp; the validation of the tool on the group's own note types and patient mix before go-live; ongoing monitoring for drift and for disparate performance across patient populations; the committee and incident-response overhead; and, named explicitly as a recurring cost, the clinician time to verify and attest to every AI-drafted note before it enters the legal record. The net case is positive, but it is positive after honesty, not before it.
Collapsed to a single net-case summary, the rebuilt slide looks like the table below. The figures are illustrative placeholders to show the shape, not numbers to repeat; the discipline is that efficiency is stated as hard dollars only, risk as expected value with visible assumptions, and cost as the full total of ownership. The net is positive after honesty, which is the only kind of positive worth presenting.
| Line | Basis | Illustrative annual value (verify, do not repeat) |
|---|---|---|
| Efficiency axis: cancelled scribe contract, incremental visits, retention delta | Hard dollars only | +$140,000 |
| Risk axis: denials and rework, coding-audit, penalty, malpractice class | Expected value with visible assumptions | +$90,000 |
| Total cost of ownership: license, change management, validation, monitoring, integration, verification time | Full cost of doing it safely | -$150,000 |
| Net case | Positive after honesty | +$80,000 |
When the lead presents it, the CFO picks the pen back up, but this time to sign. The difference is not that the number got bigger. The difference is that the case answered both of the CFO's questions, the money you can count and the money you do not lose, and it did so after subtracting the cost of doing it safely. The market context, an AI-in-healthcare market estimated somewhere between roughly thirty-seven and fifty-six billion dollars in 2026 and roughly three-quarters of health systems already running at least one AI application, told the CFO this was not a fringe bet. But it was the dual-axis structure, not the market hype, that won the funding. Every one of those market and study figures, the burnout drop, the market size, the adoption rate, was presented as a number to verify against the group's own data, never as a borrowed promise.
Key Takeaways
- Efficiency alone is half a case. Saved minutes are soft dollars that move the ledger only when deliberately converted into added visits, cancelled contracts, avoided overtime, or reduced turnover; an efficiency pitch that never names the conversion is asking finance to fund a feeling.
- Build the case on two axes at once: time and cost saved (the money you can count) and risk and harm avoided (the money you do not lose). A case built on one axis is not conservative, it is incomplete, and it both undersells the tool and hides the costs that make it honest.
- Quantify the efficiency axis from the EHR's own after-hours-time logs against a real baseline, claim only the delta you can see, and state the hard-dollar conversion mechanism explicitly rather than monetizing raw hours.
- Treat burnout reduction as the mechanism and retention as the bankable consequence: connect a measured burnout or intent-to-stay signal to a conservative avoided-physician-replacement cost rather than claiming the whole burnout drop as dollars.
- Quantify the risk axis with expected value (frequency times cost times prevented-fraction), covering avoided rework, denials, RAC and coding-audit exposure, readmission and quality penalties, and a deliberately modest malpractice-exposure line, with assumptions visible.
- Never build the audit or coding line on upcoding: a tool that inflates codes manufactures fraud exposure rather than avoiding it. The right column must rest on defensible documentation only.
- Subtract the real total cost of ownership, especially change management and the safety and verification cost the vendor deck omits. An ROI story that ignores safety and change-management cost is a fantasy, and the human verification the iron rule requires is a recurring, fundable line, not a free assumption.
- Treat every cited figure, a burnout drop from roughly 52 to 39 percent, a market between roughly 37 and 56 billion dollars in 2026, roughly three-quarters of systems running AI, as a number to verify against your own baseline and population, never a borrowed promise to paste into a business case.
Skill.re