AI for Healthcare & Clinical Practice
Strategic · M7 · lesson 7 of 23 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Evaluating Clinical AI Vendors for Safety and Transparency
📖
now learning

Evaluating Clinical AI Vendors for Safety and Transparency

15 min

The demo was flawless. On a wall-sized screen, the vendor's sepsis-prediction model lit up a deteriorating patient forty minutes before the bedside team would have caught it, and the room of executives leaned in. Then your CMIO asked a quiet question that changed the temperature: "What was the sensitivity and the positive predictive value in a population that looks like ours, and how many false alarms per real case did the nurses see?" The vendor's slide deck did not have that number. It had a testimonial, a logo wall, and a single retrospective AUROC from an academic center two thousand miles away. That gap, between what was on the screen and what your CMIO asked for, is the entire subject of this lesson. Vendor evaluation is not about being impressed. It is about deciding, before any contract is signed, whether a tool is assurable: something you can validate, monitor, defend to a surveyor, and stand behind when it touches a patient.

The Job Is Diligence, Not Shopping

When a clinical AI tool enters your institution, it does not arrive as a neutral piece of software. It arrives as a component of care that your clinicians will act on, that your record will reflect, and that your organization will answer for when it is wrong. That reframes the entire evaluation. You are not a shopper comparing features; you are a fiduciary deciding whether to import a source of risk into a system where the accountability, by law and by ethics, stays with you. The vendor's job in the sales cycle is to make the tool look inevitable. Your job is to make it prove that it is safe, fair, and monitorable in your hands, on your patients, in your workflow. Those two jobs are not the same, and a good evaluation keeps them clearly separated.

The reason this matters so much in 2026 is that the market is enormous and uneven. Roughly seventy-five percent of health systems now run at least one AI application, seventy-one percent of hospitals report predictive AI in the EHR, and the healthcare-AI market sits somewhere around thirty-seven to fifty-six billion dollars. In a market that large and that fast, a great deal of the marketing has outrun the evidence. Some vendors have done the hard, expensive work of prospective validation, bias testing, and honest disclosure of limits. Others have a beautiful interface, a compelling retrospective number, and a legal team that has thought carefully about how little they are obligated to tell you. The whole point of structured diligence is to tell those two apart before money and patients are on the line, because after go-live the cost of being wrong is measured in harm, not in a refund.

One artifact makes the ONC transparency rule concrete at the point of purchase: the source attributes. Under HTI-1, certified health IT that includes a predictive decision support intervention must expose a defined set of facts about the intervention, a nutrition-label-style disclosure covering how it was developed, on what data it was trained and validated, how it performs, how fairness was addressed, and how it should be used and maintained. When a vendor sells you a predictive DSI, you are entitled to see these attributes, and a vendor who cannot or will not surface them for a tool that is supposed to carry them has already failed a baseline expectation. The table below maps the categories of source attributes you should insist on and the diligence question each one lets you ask.

Source attribute categoryWhat it disclosesThe question it lets you press
Purpose and intended useWhat the intervention is for and the population it targetsIs my population inside the intended use, or am I off-label?
Development dataThe data and cohorts used to build the modelDo those cohorts resemble my patients at all?
Validation and performanceHow and where it was validated, and its measured performanceWas it validated prospectively, and on people like mine?
Fairness and biasHow disparate performance was evaluatedShow me the subgroup numbers, not an aggregate.
Use and maintenanceHow to use it safely and how it is updated and monitoredWhat do I have to watch, and how will I know when it changes?

These attributes do not replace your own diligence; a nutrition label tells you what is in the box but not whether the meal is right for your patient. They give you a right to ask, backed by regulation, and they give a good vendor a structured way to answer. Treat a complete, candid set of source attributes as a strong positive signal, and treat their absence, evasiveness, or a shrug of "that does not apply to us" as a finding you record and weigh, not a formality you skip.

The Joint Commission and CHAI made this explicit in the Responsible Use of AI in Healthcare guidance released in September 2025. Among its seven foundational elements is a plain expectation that vendors disclose the known risks, limitations, and biases of their tools, and that the deploying organization evaluate risk and bias before and after deployment on representative data. Read that carefully: the accreditor does not treat the vendor's confidence as evidence. It treats the vendor's disclosure of what the tool cannot do as the evidence, and it puts the burden of validation on you. A vendor who will not disclose limits is not passing your diligence; they are failing the first question the accreditor will ask you at survey.

The Five Questions That Separate Assurable from Glossy

Underneath every polished pitch, five questions do the real work. They are not adversarial for the sake of it; a good vendor will have crisp answers to all five and will respect you for asking. A vendor who deflects, redirects to a testimonial, or treats the questions as an insult has told you something important about what happens after you sign.

Question one: on whom was this validated?

Every performance claim is a claim about a specific population, and the number is only meaningful if that population resembles yours. Ask precisely: what was the age distribution, the racial and ethnic composition, the payer mix, the comorbidity burden, the care setting, and the geography of the validation cohort? A sepsis model validated on a tertiary academic center with a young, well-insured, largely white population may perform very differently on your safety-net community hospital with older patients, more uncontrolled diabetes, and a different mix of presentations. The question is not "is the model good," it is "is the model good on patients like mine," and only a described validation population lets you answer it. If the vendor cannot describe the validation population in detail, they cannot claim the tool will work for yours, and you should treat the headline number as a hypothesis, not a result.

Question two: what are the real error rates?

A single accuracy figure or an AUROC is a marketing number, not a safety number. What a clinical leader needs is the confusion matrix in clinical terms: sensitivity (of the patients who truly have the condition, what fraction does the tool catch), specificity (of those who do not, what fraction does it correctly clear), positive predictive value (when the tool fires, how often is it actually right), and negative predictive value. PPV is the one vendors most often hide, because it depends on how common the condition actually is in your population, and in a low-prevalence setting even a very sensitive model can be wrong most of the time it alarms. If a deterioration model has a PPV of fifteen percent, that means roughly six of every seven alerts your exhausted nurses chase are false, which is not a detail. It is the difference between a tool that saves lives and a tool that trains your staff to ignore it. Demand the numbers, demand the population they came from, and demand to know how false positives and false negatives are distributed, because a model can hit an impressive average while failing badly on the subgroup you most need it to catch.

Question three: is the evidence prospective or retrospective?

This distinction is where a great deal of clinical AI marketing quietly lives. A retrospective study runs the model backward over data that already exists and asks how well it would have done. A prospective study deploys the model in real time and measures what actually happened, including how clinicians behaved and whether outcomes changed. Retrospective performance is necessary but nowhere near sufficient, because it cannot capture the two things that matter most at the bedside: how the model behaves on data as it arrives messily in real time, and how real clinicians act on its output under real conditions. A model that looks brilliant retrospectively can fail prospectively because clinicians ignore it, or over-trust it, or because live data differs from the clean training set. Ask which studies are prospective, whether any measured a change in patient outcomes rather than just model accuracy, and whether they were done at sites like yours. Retrospective-only evidence is a reason to run your own prospective pilot, which the next lessons will build.

Question four: was it tested for disparate performance?

A model trained on a non-representative population underperforms for the patients already underserved, and that underperformance is both a clinical harm and a legal exposure. The question is direct: did you measure performance separately by race, ethnicity, sex, age, language, and payer, and can you show me those subgroup numbers? A tool with excellent aggregate performance can have a sensitivity that collapses for Black patients, or for women, or for the very elderly, and the aggregate number will hide it completely. This is not a theoretical concern; it is one of the best-documented failure modes in clinical AI, and it is exactly what the CHAI guidance asks you to evaluate before and after deployment. If a vendor has never run subgroup analysis, they have not tested for the harm most likely to land on your most vulnerable patients, and you now know something the demo did not tell you.

Question five: will it be monitorable after go-live?

A model is not a finished object; it drifts. As your population, your practice patterns, your coding, and your upstream data change, a model that was accurate at launch degrades silently, and the only defense is ongoing monitoring. So ask what the vendor gives you to monitor it: does the tool expose the performance metrics and alert volumes you need to watch drift, will they notify you when they change the model, do they support local re-validation, and can you get the data out to audit it yourself? A vendor whose answer is "trust us, we monitor it centrally" is asking you to accept a black box that you are nonetheless accountable for. The assurable tool comes with the instruments to watch it. The glossy one comes with a dashboard designed to reassure, not to reveal.

The five questions are easier to hold in a room if you carry the answer you actually need for each one and the deflection that should raise an alarm. A vendor who deflects on any single question has not disqualified themselves; a vendor who deflects on most of them has told you what the partnership will feel like the first time something goes wrong.

Diligence questionThe answer you needThe deflection that should worry you
On whom was it validated?A described cohort: age, race and ethnicity, payer mix, comorbidity, setting, geography"A large, diverse dataset" with no breakdown you can inspect.
What are the real error rates?Sensitivity, specificity, PPV and NPV, with the prevalence they assumeA single accuracy figure or an AUROC offered as if it were a safety number.
Prospective or retrospective?At least one prospective study, ideally showing a change in patient outcomesRetrospective-only evidence presented as proof it works live.
Tested for disparate performance?Subgroup results by race, ethnicity, sex, age, language, and payer"We do not use race, so bias is not an issue," or no subgroup analysis at all.
Monitorable after go-live?Drift metrics, model-change notification, local re-validation support, exportable data"We handle monitoring on our back end," offering reassurance instead of instruments.

Notice that none of these answers requires you to be a data scientist to evaluate. Each is a document a serious vendor already has and a glossy one does not. The instrument you are really testing is not the model; it is whether the vendor built the evidence a regulated buyer needs, or built a demo. That distinction survives every change in the underlying technology, which is why these questions age well even as the models improve.

A demo shows you the best case the vendor could stage. Diligence asks about the worst case you will have to answer for. Buy on the second, never the first.

How a Vendor Discloses Risk Tells You Who They Are

There is a single signal that predicts more than any other whether a partnership will go well, and it is how the vendor talks about what their tool cannot do. A mature clinical AI vendor volunteers the limits. They will tell you the populations where performance is weaker, the failure modes they have seen, the conditions the model was never trained on, the settings where it should not be used, and the monitoring you will need to run. That candor is not a weakness in their pitch; it is the strongest evidence you will get that they understand their own product and that they have a safety culture rather than a sales culture. It also happens to be exactly what the Joint Commission and CHAI expect a vendor to provide.

The opposite posture is the tell. A vendor who has no documented limitations, who answers every risk question with a reassurance, who treats "where does this fail" as a hostile question, or who cannot produce a written statement of known risks and biases, is either unaware of their own failure modes or unwilling to put them in writing where you could hold them to it. Both are disqualifying for a tool that touches patients. You are not looking for a vendor with no limitations, because no such tool exists. You are looking for a vendor who knows their limitations well enough to name them, because those are the limitations you will otherwise discover the hard way, in a chart review or a survey finding or a harmed patient. Insist that the disclosure of risks and limits be a written artifact you can keep, reference, and attach to your governance file, not a verbal reassurance in a sales meeting that evaporates the moment something goes wrong.

A Worked Example: Two Vendors, the Same Beautiful Demo

Consider two vendors selling an AI tool that flags patients at risk of readmission so care managers can intervene. Both give an identical, impressive demo: the same slick interface, the same success story, the same claim of "ninety percent accuracy." A shopper would call it a tie and pick on price. A clinical leader running diligence pulls them apart in twenty minutes.

Vendor A, asked about population, produces a one-page validation summary: the cohort was one hundred forty thousand admissions across six hospitals including two safety-net sites, with the age, race, payer, and comorbidity breakdown attached. Asked about error rates, they hand over sensitivity, specificity, and PPV by subgroup and note candidly that PPV drops in the lowest-acuity tier, so the tool should be aimed at medium-and-high-risk patients. Asked about evidence, they cite one retrospective and one prospective study, the prospective one showing a real reduction in thirty-day readmissions. Asked about bias, they show subgroup performance and flag that sensitivity is somewhat lower for patients whose primary language is not English, with a monitoring plan to watch it. Asked about monitoring, they provide a drift dashboard, a commitment to notify you before any model change, and an export for your own audit. Their written limitations document is three pages long.

Vendor B, asked the same five questions, says the validation was done "on a large diverse dataset" but cannot break it down, offers "ninety percent accuracy" but no PPV and no subgroup numbers, has only retrospective evidence, has "not specifically" tested for disparate performance because "the model does not use race," and answers the monitoring question with "our team handles that on the back end." Their limitations document does not exist; they offer to "put something together." On price, Vendor B is cheaper. On assurability, Vendor B is a liability you would be importing with your own signature on it. The demo told you nothing. The five questions told you everything, and notice that at no point did you need to be a data scientist. You needed to know which questions a tool that touches patients must be able to answer, and to refuse to be satisfied until they were answered.

Turning Diligence into a Repeatable Instrument

A one-off interrogation is good; a standing instrument is far better, because it makes the diligence survive the enthusiasm of the executive who fell in love with the demo. Build a written vendor safety-and-transparency questionnaire that every clinical AI vendor must complete before procurement will advance, and score it. It should demand, in writing and with evidence attached, the validation population, the real error rates including PPV by subgroup, the prospective evidence, the bias testing, the monitoring and model-change notification commitments, the regulatory status (whether it is an FDA-authorized device and under what intended use, or a non-device tool), the ONC source attributes if it is a predictive decision-support intervention, and a signed statement of known risks and limits. Make the questionnaire a gate: incomplete answers stop the process, and the answers become part of the permanent governance record for that tool.

This does two things at once. It protects patients, obviously, by keeping unassurable tools out. But it also protects your institution the day something goes wrong, because when a surveyor, a plaintiff's attorney, or your own board asks "how did you decide this tool was safe to deploy," you can produce a completed, evidence-backed diligence file rather than a memory of an impressive presentation. The vendors who can complete it well are, almost by definition, the ones worth partnering with, and the ones who cannot have selected themselves out before they could do harm. The questionnaire is not bureaucracy. It is the difference between a decision you can defend and a decision you merely made.

There is one boundary the questionnaire must make explicit, because it is the boundary vendors most want to blur: a strong vendor reduces your risk, but it never assumes your obligations. The accountability for a patient touched by AI stays with the licensed clinician and the covered entity, no matter how much the contract talks about the vendor's quality program. A useful diligence file names, for each obligation, what the vendor can supply and what remains yours, so that no one on your team mistakes a good vendor for a transfer of duty.

ObligationWhat a strong vendor can supplyWhat stays with you and never transfers
Verification of outputsExplainable outputs and reasoning a clinician can checkThe clinician's duty to verify before acting; "the AI said so" is never a defense.
Fitness for your populationValidation data and subgroup performanceConfirming the tool works on your patients, in your workflow, before you rely on it.
HIPAA and PHI protectionA signed BAA and security controlsThe covered entity still owns a breach if PHI is mishandled downstream.
Disclosure to patientsDocumentation of AI use and its limitsMeeting state disclosure law, such as Texas TRAIGA effective in 2026, remains the provider's duty.
Standard of careEvidence and guidance about appropriate useThe clinical decision, and liability for following a wrong output or ignoring an accurate one, stays human.
Ongoing safetyDrift metrics and model-change noticesActing on the monitoring, pausing a drifting tool, and governing it over time.

The practical value of naming this boundary is that it changes how you read a contract. A clause promising the vendor will "ensure clinical safety" or "handle compliance" is not a shield; the regulator and the plaintiff will still come to you. What protects the institution is not a sentence in a vendor agreement but the evidence file and the human verification behind every output. Read the vendor's promises as a supplement to your accountability, never a substitute for it, and the whole diligence exercise falls into its correct place: it makes your unavoidable responsibility easier to discharge, not something you can hand away.

Key Takeaways

  • Vendor evaluation is diligence, not shopping: you are deciding whether a tool is assurable (validatable, monitorable, and defensible on your patients), because the accountability for its errors stays with your institution no matter who built it.
  • Five questions separate an assurable tool from a glossy one: on whom was it validated, what are the real error rates (sensitivity, specificity, and especially PPV), is the evidence prospective or only retrospective, was it tested for disparate performance by subgroup, and will it be monitorable after go-live.
  • PPV is the number vendors most often hide, because in a low-prevalence setting even a sensitive model can be wrong most times it alarms, which is what trains staff to ignore it.
  • Retrospective performance is necessary but not sufficient; only prospective evaluation captures how live data and real clinician behavior change the result, which is why you still run your own pilot.
  • Aggregate accuracy hides subgroup failure; demand performance broken out by race, ethnicity, sex, age, language, and payer, because disparate performance is the harm most likely to land on your most vulnerable patients.
  • How a vendor discloses limits is the single strongest signal of who they are: the Joint Commission and CHAI expect vendors to disclose known risks, limitations, and biases in writing, and a vendor with no documented limitations is disqualifying, not reassuring.
  • Insist that risk and limitation disclosures be a written artifact you can keep and attach to governance, never a verbal reassurance in a sales meeting.
  • Turn diligence into a standing, scored vendor safety-and-transparency questionnaire that gates procurement, so the decision is repeatable, survives executive enthusiasm, and produces a defensible file the day someone asks how you decided the tool was safe.