AI in Healthcare Delivery
Learning Objectives
By the end of this lecture you will be able to (1) map the federal healthcare AI landscape across VA, CMS, FDA, CDC, NIH, HRSA, IHS, ONC, and SAMHSA, with the specific authorities each exercises; (2) describe the Food and Drug Administration's regulatory framework for Software as a Medical Device (SaMD), including the 2019 AI/ML-based SaMD Action Plan, the January 2021 AI/ML SaMD Update, the 2023 guiding principles for Good Machine Learning Practice (GMLP) co-developed with Health Canada and the UK MHRA, and the predetermined change control plan under Section 3060 of the 21st Century Cures Act; (3) apply HIPAA Privacy and Security Rules, the Common Rule, and the 42 CFR Part 2 substance use protections to AI projects involving protected health information; (4) analyze three widely cited AI failures in healthcare (Epic Sepsis Model; IBM Watson for Oncology; Optum chronic-care algorithm) and translate lessons to federal program design; (5) evaluate real federal AI deployments at VA (REACH VET suicide risk), CMS (risk adjustment audits, fraud detection), FDA (adverse-event surveillance), CDC (syndromic surveillance, CFA Center for Forecasting and Outbreak Analytics), NIH (Bridge2AI, All of Us), and the Indian Health Service; (6) align deployments with OMB M-24-10 rights-impacting and safety-impacting minimum practices, NIST AI RMF, EO 14110, and forthcoming HHS AI strategy guidance; and (7) design clinical-governance structures, including model oversight committees, physician informaticist roles, patient advisory participation, and ongoing monitoring that meets Joint Commission and CMS Conditions of Participation expectations.
The Federal Healthcare AI Landscape
Federal involvement in health AI spans payers, providers, regulators, researchers, and public-health authorities. CMS operates Medicare, Medicaid, CHIP, and the Marketplace, touching nearly half the U.S. population through payment and coverage decisions. The Veterans Health Administration is the largest integrated health system in the country, with more than 170 medical centers and 1,000+ outpatient clinics. FDA regulates medical devices including SaMD under 21 U.S.C. 360 and the 2016 21st Century Cures Act. CDC conducts surveillance through the Center for Forecasting and Outbreak Analytics (CFA) established 2021 and the National Syndromic Surveillance Program. NIH funds biomedical research including the Bridge2AI program launched 2022, the All of Us Research Program, and AIM-AHEAD for health equity AI. HRSA administers community health programs. IHS serves tribal populations. ONC (Office of the National Coordinator for Health IT) sets electronic health record standards including the HTI-1 Final Rule (January 2024) which includes algorithm transparency requirements for certified health IT. HHS Office for Civil Rights enforces HIPAA and Section 1557 (nondiscrimination in federally-funded health programs), including 2024 Final Rule provisions on AI and clinical decision support. SAMHSA addresses mental health and substance use. Each arm has distinct authority, risk profile, and governance; coordinated through the HHS AI Council chaired by the HHS CAIO. Under OMB M-24-10, HHS designated a CAIO and published its first use case inventory in 2024.
FDA Regulation of AI/ML Medical Devices
FDA regulates medical devices under the Federal Food, Drug, and Cosmetic Act. Software qualifies as a device if intended for medical purposes; SaMD stands for Software as a Medical Device. AI/ML-based SaMD introduces novel regulatory challenges because models are not static: retraining, fine-tuning, and distributional shift change behavior over time. FDA's April 2019 discussion paper 'Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device' introduced the Predetermined Change Control Plan concept. The January 2021 Action Plan updated this thinking. October 2021 Good Machine Learning Practice Guiding Principles, co-authored with Health Canada and the UK MHRA, set ten foundational principles including multi-disciplinary expertise applied throughout, sound software engineering and security practices, representative clinical study participants and datasets, independence of training and test data, human-in-the-loop consideration, testing demonstrating device performance during clinically relevant conditions, clear information for users, performance monitoring of deployed models, and clear communications of updates. The April 2023 guidance on Predetermined Change Control Plans operationalized a pathway under Section 3060 of the 21st Century Cures Act. As of 2024, more than 900 AI-enabled medical devices have been authorized via 510(k), De Novo, or PMA pathways. Federal program managers procuring AI medical devices should ensure FDA authorization where applicable, understand the predicate device basis, and evaluate labeling for intended use and population.
HIPAA, Privacy, and Security for Healthcare AI
HIPAA has three rules most relevant to AI. The Privacy Rule governs uses and disclosures of protected health information (PHI). The Security Rule specifies administrative, physical, and technical safeguards for electronic PHI. The Breach Notification Rule addresses incident response. AI projects touch PHI in three ways: training data, inference inputs, and outputs that become part of the medical record. Each creates compliance obligations. For training on PHI, the Privacy Rule permits uses for 'health care operations' broadly, and research uses typically require IRB approval with waiver or authorization, or de-identification. De-identification can follow Safe Harbor or Expert Determination; both have limitations for AI, as high-dimensional health data frequently remains re-identifiable. For inference, PHI inputs require Business Associate Agreements with any vendor processing PHI, covered under 45 CFR 164.504(e). For outputs that become part of the record, model cards and intended-use documentation should accompany. The 2024 HHS Office for Civil Rights Final Rule under Section 1557 extended nondiscrimination requirements to 'patient care decision support tools' including AI, requiring covered entities to make reasonable efforts to identify and mitigate discrimination risk. State laws add complexity: California CMIA, Texas Medical Records Privacy Act, and New York SHIELD Act impose requirements beyond HIPAA. Federal programs must also consider 42 CFR Part 2 for substance use disorder records, which has stricter re-disclosure rules.
Case Study: Epic Sepsis Model
The Epic Sepsis Model was deployed at hundreds of hospitals in the United States using Epic Systems' EHR. Epic marketed the model with a reported area-under-receiver-operating-characteristic (AUROC) of 0.76-0.83 for predicting sepsis. In June 2021, Wong et al. published a study in JAMA Internal Medicine evaluating the model at University of Michigan, finding much lower performance: AUROC of 0.63, sensitivity of 33 percent at typical operating thresholds, and a numbers-needed-to-evaluate ratio that created significant alert fatigue. The lower performance relative to vendor claims was attributed to differences between development and deployment populations, concept drift, and changes in clinical workflow. Epic subsequently worked with institutions to adjust the model and released improvements. Lessons for federal program managers: (1) vendor-reported performance metrics must be independently validated before deployment; (2) performance depends heavily on local population, EHR configuration, and workflow, not just the model itself; (3) alert fatigue is a safety concern, not just a usability issue; (4) ongoing monitoring with publication of real-world performance is the standard of care; (5) the VA, which runs its own clinical decision support pipeline, and the Military Health System, which runs MHS GENESIS on Cerner/Oracle, should not assume vendor models are production-ready without local validation.
Case Study: IBM Watson for Oncology
IBM Watson for Oncology was marketed as an AI system to assist oncologists in treatment planning, developed in collaboration with Memorial Sloan Kettering Cancer Center and deployed at MD Anderson Cancer Center beginning 2013. Internal documents reported by STAT in 2017 and 2018 indicated the system had been trained on hypothetical cases rather than real patient data, and physicians at multiple sites reported 'unsafe and incorrect' treatment recommendations. MD Anderson terminated its relationship in 2017 after spending more than $60 million. IBM subsequently wound down the product. Lessons for federal program managers: (1) data provenance is not optional; training on synthetic or hypothetical data without rigorous external validation is a design defect; (2) marketing claims should be scrutinized against clinical evidence; (3) expensive failures can damage institutional trust long after the product is retired; (4) procurement safeguards including pilot evaluation, independent validation, and staged deployment would likely have surfaced the problem earlier; (5) the Department of Veterans Affairs explicitly cited Watson's failure in designing the National AI Institute's validation requirements.
Case Study: Optum Chronic Care Algorithm
In October 2019, Obermeyer, Powers, Vogeli, and Mullainathan published in Science an analysis of a widely used chronic care management algorithm affecting more than 200 million Americans. The algorithm, attributed to Optum (a UnitedHealth Group subsidiary), ranked patients for high-risk care management by predicted future healthcare costs. Because Black patients historically received less spending for similar medical need (systemic inequity in access and utilization), using costs as a proxy for need systematically understated the needs of Black patients. The authors estimated that changing the label from cost to actual illness burden would have more than doubled the share of Black patients automatically identified for additional care. Optum and the authors subsequently collaborated on remediation. Lessons for federal program managers: (1) proxy labels encode the inequities of the processes that generate them; (2) algorithmic audits should specifically evaluate performance across race, ethnicity, disability, age, geography, and intersectional groups; (3) CMS, Medicaid, and VA risk stratification algorithms must be audited against this pattern; (4) Section 1557 of the Affordable Care Act applies to algorithmic discrimination in federally-funded health programs; (5) HHS OCR's 2024 Final Rule codifies this expectation.
VA REACH VET and Veterans Health AI
REACH VET (Recovery Engagement and Coordination for Health - Veterans Enhanced Treatment), launched nationally by VA in 2017, is a statistical model that identifies veterans at elevated risk of suicide, overdose, or other adverse events using EHR data. The top 0.1 percent of veterans by risk score are flagged for proactive outreach by local clinicians. Academic evaluation (Kessler et al., Bossarte et al.) found reductions in healthcare utilization associated with the intervention. REACH VET represents a mature federal AI deployment with several model features: it was developed by VA researchers with in-house data and validation; it operates at a very high precision threshold rather than a low-threshold alert; it is paired with a specific clinical workflow (outreach by a provider who knows the veteran) rather than standalone automation; it is monitored through performance reporting; and it has been the subject of published peer-reviewed evaluations. VA's National AI Institute, founded 2019, continues work in suicide prevention, cardiology, radiology, and oncology. Federal program managers should treat REACH VET as a case study in responsible deployment rather than a generic template: its strength comes from in-house development, clinical workflow integration, and conservative operating thresholds.
CMS Fraud Detection and Risk Adjustment
CMS uses AI and advanced analytics extensively. The Fraud Prevention System, operated through the Center for Program Integrity, screens Medicare fee-for-service claims for fraud indicators. The Predictive Learning Analytics Tracking Outcomes (PLATO) system and related tools assist in risk adjustment for Medicare Advantage. GAO has published multiple audits noting both the benefits and the concerns, including the 2019 and 2023 reports on Medicare Advantage risk adjustment errors. AI in payment integrity raises several issues: (1) disparate impact, where algorithmic selection may concentrate audits on particular geographic or provider populations; (2) due process, where providers must have opportunity to contest; (3) adversarial adaptation, as providers adjust billing practices in response to detection algorithms; (4) data quality and concept drift, as claim coding changes over time. CMS has published some aspects of its approach in federal register notices but much detail remains internal. The CMS Innovation Center has piloted AI for care quality measurement and beneficiary outreach. Under OMB M-24-10, CMS is responsible for classifying its AI uses and applying minimum practices where rights- or safety-impacting.
CDC Surveillance and Outbreak Analytics
CDC applies AI to public health surveillance. The Center for Forecasting and Outbreak Analytics, established in 2021 with $200 million in American Rescue Plan funding and $50 million annual appropriation, integrates machine learning with traditional epidemiology for respiratory disease forecasting, wastewater surveillance analysis, and outbreak detection. The National Syndromic Surveillance Program aggregates de-identified emergency department data from hospitals across the country. AI supports anomaly detection, genomic epidemiology (SARS-CoV-2 variant tracking), and vaccine effectiveness monitoring. Challenges include data sharing limitations with state and local public health, variable EHR data quality, privacy frameworks under HIPAA and state public health exceptions, and the need for rapid model adaptation during novel outbreaks. Lessons from COVID-19 informed CFA's design: durable data pipelines, academic partnerships through the Insight Net, publicly available code and evaluation, and integration with CDC's MMWR (Morbidity and Mortality Weekly Report) communication pipeline. Federal program managers building surveillance AI should consult CFA's published materials for infrastructure patterns.
Clinical Governance for Healthcare AI
Responsible deployment of AI in healthcare requires governance structures beyond generic AI governance. A Clinical AI Oversight Committee should include medical leadership (CMO or designee), informatics leadership (CMIO), quality and safety, nursing leadership, pharmacy, patient advocacy, bioethics, legal, privacy (CPO/SAOP), security (CISO), and data science/engineering. This committee should review pre-deployment evidence including intended use, clinical validation, fairness audit across relevant subgroups, implementation plan, monitoring plan, incident response plan, and deactivation criteria. Post-deployment, the committee should receive periodic performance reports and incident reports. Physician informaticists should serve as embedded owners for each clinical AI system. Patient advisory participation should be structural, not token. Alignment with CMS Conditions of Participation, Joint Commission standards, and institutional bylaws is essential. Specific attention to alert fatigue, documentation burden, liability allocation between algorithm developer and clinician, and to the 21st Century Cures Act's anti-information-blocking rules is required. State medical board guidance increasingly addresses AI liability and scope of practice.
Health Equity and Access
Federal healthcare AI must serve all populations equitably. Section 1557 of the ACA prohibits discrimination by race, color, national origin, sex, age, or disability in federally-funded health programs; the 2024 HHS OCR Final Rule explicitly extends this to patient care decision support tools. The NIH AIM-AHEAD program funds research on AI and health equity in minority-serving institutions. The Indian Health Service and tribal epidemiology centers serve Native American populations with unique data sovereignty considerations under tribal sovereignty principles and CARE Principles for Indigenous Data Governance. Rural health (through HRSA programs, Medicare rural adjustments, and the National Rural Health Association) requires deliberate inclusion in training and validation populations. Limited English Proficiency (LEP) populations are covered under Executive Order 13166 and require translation and culturally appropriate design. Disability accessibility (ADA, Section 508) applies to AI interfaces. Federal AI healthcare programs should specify equity-specific evaluation criteria upfront, not as an afterthought, and incorporate community input through federally qualified health centers, tribal governments, and patient advocacy organizations.
Procurement and Contracting for Healthcare AI
Healthcare AI procurement in federal practice must integrate FAR, HIPAA Business Associate Agreements, FedRAMP for cloud services, FDA authorization where applicable, and clinical acceptance criteria. Common vehicles include the VA's FedRAMP Moderate cloud, the IHS and CMS enterprise agreements, GSA Multiple Award Schedule AI SINs, and for DoD medical, DHA-specific vehicles. Contract language should specify (a) data rights: who owns fine-tuned weights trained on PHI; (b) training data restrictions: vendor may not use agency PHI to train shared models without specific authorization; (c) performance evidence: delivery of real-world performance metrics stratified by subgroup; (d) change control: notification and re-validation on material model updates; (e) exit: data return and destruction on contract conclusion; (f) incident response: notification obligations meeting HIPAA Breach Notification timelines; (g) transparency: model cards, data statements, intended-use documentation. The Department of Defense Business Associate Agreement template and HHS OCR's sample BAA provide starting points. For research uses, Certificates of Confidentiality under 42 U.S.C. 241(d) may apply. FDA-regulated devices have additional requirements under Quality System Regulation (21 CFR 820).
Healthcare AI Anti-Patterns
(1) Deploy-and-Forget: procurement ends at go-live, no monitoring, no incident response planning. Epic Sepsis is the canonical example. (2) Vendor Claim Trust: deploying based on vendor-published metrics without local validation. Again Epic Sepsis and Watson Oncology. (3) Demographic Blindness: failing to stratify performance across race, ethnicity, disability, language, and intersectional groups. Optum chronic care is the canonical example. (4) Cost as Proxy: using utilization or spending as a label when access inequities exist; Optum chronic care. (5) Alert Fatigue: low-threshold alerts that overwhelm clinicians, reducing safety. Epic Sepsis contributed. (6) Missing BAA: vendor touches PHI without Business Associate Agreement; OCR enforcement follows. (7) Shadow AI: clinicians or administrators using unapproved consumer AI tools (e.g., pasting PHI into public LLM interfaces). (8) Single-Site Validation: deploying an AI system without validating on the population it will serve. (9) Opaque Labels: predicting a label whose definition is not clinically grounded. (10) Missing Deactivation Plan: no criteria for turning the AI off if it underperforms. Each anti-pattern maps to specific real-world failures; federal healthcare AI programs should audit against this list.
Summary and Next Steps
Federal healthcare AI spans VA clinical care, CMS payment integrity, FDA device regulation, CDC surveillance, NIH research, IHS tribal health, HRSA community health, ONC interoperability, and HHS OCR civil rights. HIPAA, Section 1557, 21st Century Cures, FDA SaMD rules, OMB M-24-10 minimum practices, and NIST AI RMF all apply. Three widely-cited failures (Epic Sepsis, Watson Oncology, Optum chronic care) illustrate the data-provenance, validation, and fairness disciplines required. VA REACH VET, CMS Fraud Prevention System, CDC CFA, and NIH Bridge2AI illustrate mature or maturing deployments. Clinical governance including multi-disciplinary oversight committees, physician informaticist ownership, patient advisory participation, and alignment with CMS Conditions of Participation and Joint Commission standards is essential. Equity, procurement rigor, and anti-pattern avoidance round out the discipline. The next lecture, AI in Social Services, applies analogous patterns to human services programs including TANF, SNAP, child welfare, and workforce services.
Skill.re