AI for Care Management and Care Coordination
It is 7:40 on a Monday morning, and a care manager named Dana opens her laptop to a fresh list. The AI has done its work overnight: forty patients ranked by predicted risk, the ones her panel management model says are most likely to land in the hospital in the next ninety days, sorted neatly from most urgent to least, each with a one-line reason and a suggested outreach action. It is a beautiful artifact. It is also, this particular Monday, quietly wrong at the top and quietly wrong at the bottom, and the entire question of whether Dana's morning helps her patients or harms them turns on whether she treats that list as an input she will verify or a verdict she will simply execute.
What Care Management Actually Is, and Where AI Enters
Care management and care coordination are the connective tissue of modern healthcare, and they are almost invisible to people who have never needed them. A care manager owns a panel, a defined group of patients, often the sickest and most complex, and their job is to keep those patients from falling through the cracks between visits, between specialists, between the hospital and home. They chase the open referral that never got scheduled. They call the newly discharged heart-failure patient to make sure the new diuretic dose is being taken and the follow-up is booked. They close the care gaps, the overdue mammogram, the missing A1c, the lapsed medication, that quietly accumulate into the next crisis. Coordination is the same work aimed outward: making sure the cardiologist, the primary care physician, the home health nurse, and the family are all working from the same understanding of the same patient, so that nothing gets dropped in the handoffs between them.
This work has always been rationed by a brutal arithmetic. A care manager cannot call everyone. The panel is large, the day is finite, and the whole discipline is really about one scarce resource: attention. Who gets the outreach call today, and who waits until next week or next month? Historically that triage was done by hand and by memory, informed by whichever patient happened to be top of mind, whichever chart happened to cross the desk, whichever crisis was already loud enough to demand a response. It was reactive, uneven, and heavily dependent on the individual care manager's knowledge of their own panel. Sick patients who were quiet, who did not call, who did not show up in the emergency department yet, could go unnoticed until they were no longer quiet.
This is exactly the seam into which AI has been inserted, and on paper the fit is superb. A risk stratification model can look across the entire panel at once, ingest claims, labs, vitals, diagnoses, utilization history, and medication data, and produce a ranked prediction of who is most likely to deteriorate. It can generate the outreach list Dana opened this morning. It can summarize a complex chart in seconds so she does not spend twenty minutes reconstructing a patient's story before a five-minute call. It can flag every open care gap on a patient in one view. It can even draft the outreach script, the words she might use to explain a medication change or coax a reluctant patient into scheduling a colonoscopy. Used well, this is a genuine multiplier: it points a finite amount of human attention at the patients most likely to benefit from it, and it removes hours of preparatory grind so the human can spend her scarce minutes on the human part of the work.
It is worth being precise about how much of this is genuinely new and how much is the same old work wearing a faster engine. Care managers have always stratified their panels; what the model changes is the speed, the breadth, and above all the apparent authority of the stratification. A human triaging from memory holds maybe a dozen patients clearly in mind at once and knows she is guessing about the rest. A model holds the whole panel and every data point at once and presents its ranking with no visible hesitation. That combination of breadth and confidence is the gift and the trap in a single object. The gift is that a patient who would have stayed invisible to a tired human, quiet, non-complaining, buried in a panel of four hundred, can now be surfaced by a model that never forgets anyone. The trap is that the same confident surface makes it psychologically easy to forget that the model is guessing too, only faster and with better production values.
The Named Human Who Owns the Panel
Here is the load-bearing principle, and it is the same principle that runs through this entire program, wearing new clothes for a new setting: the AI generates the list, but a named human care manager owns the panel. Not the model. Not the vendor. Not the abstract system. A specific person, whose name is on the panel, who is accountable for what happens to those patients, and who decides, every day, who gets the call and who does not. The AI is an extraordinarily useful assistant to that person. It is never a substitute for that person's judgment, and it is never the entity that bears responsibility when a decision goes wrong.
This matters because of a subtle and dangerous shift in how a ranked list changes behavior. When a human triages a panel from memory, she knows the judgment is hers, provisional, and fallible, so she stays alert. When the same human opens an authoritative, precisely ordered, algorithmically generated list, something in the psychology changes. The list looks like an answer. It has the crisp confidence of a machine that has considered more data than any person could hold in their head. The temptation is to stop thinking of it as a suggestion to weigh and start treating it as a work queue to clear from top to bottom. That is the moment care management goes wrong, not because the model is evil or even usually inaccurate, but because a fallible input has been silently promoted to an unquestioned verdict, and the human check that was supposed to catch the model's errors has quietly switched off.
The discipline, then, is to keep the list in its proper place. It is an input, one voice at the table, a well-informed but imperfect colleague offering an opinion. The care manager's job is to receive that opinion, test it against what she knows and what the chart actually says, and then decide. She verifies the list against reality. She notices the patient the model ranked third whom she spoke with on Friday and who is doing fine, and she notices the patient the model did not surface at all whose spouse called the front desk in tears yesterday. She owns the panel. The list serves her. It does not command her.
It helps to make the contrast concrete, because the difference between the two postures is not a matter of attitude but of what actually happens to a patient. The table below sets the two side by side.
| The list as verdict | The list as input |
|---|---|
| The ranking is a work queue to clear from top to bottom. | The ranking is a hypothesis to test before acting. |
| The human check switches off; the model's errors pass straight through to the patient. | The human check stays on; the model's errors are caught before they reach the patient. |
| Staleness and bias are invisible, because nobody is looking past the order the model produced. | Staleness and bias are surfaced, because the human compares the ranking to what she and her team actually know. |
| When a patient is missed, the record shows a queue was executed. | When a patient is missed, the record shows a person weighed the evidence and made a defensible call. |
| Accountability drifts toward the model and the vendor, where it cannot legally rest. | Accountability stays with the named human who owns the panel, where it belongs. |
Nothing in the right-hand column requires distrusting the model or throwing it away. It requires only that the care manager keep her own judgment switched on while she uses it, which is a small habit with enormous consequences.
A risk list is not a diagnosis of your panel. It is a hypothesis about your panel, generated by a model that has never met your patients, and every hypothesis has to be tested against the reality only you can see.
The Central Danger: A Biased List Steers Care Away From the Underserved
Now we reach the specific danger that makes AI in care management different in kind, not just in degree, from a busy care manager triaging by memory. A biased or badly designed risk model does not fail randomly. It fails systematically, and it tends to fail in the same direction, steering scarce care-management resources away from exactly the patients who are already underserved. This is not a hypothetical worry. It is one of the best-documented findings in the entire field of clinical AI, and every care manager working from a risk list should understand it in their bones.
The now-classic example comes from a widely used commercial algorithm designed to identify patients with complex health needs for extra care-management support. The model did not predict illness directly. Instead it used a proxy: it predicted future healthcare costs, on the reasonable-sounding theory that sicker patients cost more, so the highest-cost patients must be the sickest and most in need of help. The logic is seductive and it is wrong, because cost is not a clean measure of sickness. It is a measure of sickness filtered through access, through trust, through insurance, through every barrier that stands between a patient and the healthcare system. Black patients in the study, at any given level of actual illness, generated lower costs than white patients, not because they were healthier but because they had less access to care, faced more barriers, and received fewer services for the same underlying disease. The model, trained to chase cost, therefore concluded that these sicker patients were healthier than they were, and systematically under-identified them for the extra care they most needed. Researchers estimated that correcting the bias would have more than doubled the proportion of Black patients flagged for additional help.
Sit with the shape of that failure, because it is the template for a whole class of dangers. The model was not obviously broken. It ran smoothly, produced confident rankings, and looked exactly as authoritative as a well-behaved model. Its bias was invisible in the output and lived entirely in the choice of what it was trained to predict. And its effect was not neutral: it took the patients who already got too little and gave them less, because it read their under-treatment as wellness. A care manager clearing that list top to bottom would have spent her scarce attention on the patients the system already served well, while the patients the system was already failing stayed invisible, one more time, now with an algorithm's authority behind their invisibility. This is what disparate performance means in the flesh: a model that works less well for some populations than others, deployed into a workflow that turns its blind spot into a pattern of who gets cared for and who does not.
The Stale List Problem: Yesterday's Emergency, Today's Miss
Bias is the deep structural danger. There is a second, more mundane danger that will bite far more often, and it is the reason Dana's list this morning was wrong at both ends. A risk model is only as current as the data it was last fed and the moment it last ran. Health, especially the health of complex high-risk patients, changes faster than most model refresh cycles. The result is a stale list: a ranking that reflects a snapshot of reality that has already moved.
Staleness produces two failures, and they are mirror images. The first is the false positive that will not die. A patient is ranked at the top because six weeks ago he was admitted, unstable, and clearly high-risk. Since then he has stabilized, his medications are optimized, his follow-up happened, and he is, right now, doing well. But the model still sees the six-week-old signal and still ranks him first. A care manager who works the list mechanically spends her morning on a patient who no longer needs the intervention, attention drained on a problem that has already resolved. The second failure is worse: the newly decompensating patient the list cannot see. Somewhere on that panel is a woman who was stable at the last data pull and so sits quietly in the middle of the list, but whose condition began sliding three days ago, after the model last ran. She is the single most important call of the day, and the list points away from her. The stale list surfaces the patient who no longer needs help and buries the patient who just started to.
The defense against staleness is not a better model, though better refresh cadence helps. The defense is the care manager's own current knowledge of her panel, deliberately kept in the loop. She knows the top-ranked patient stabilized because she talked to him last week. She hears about the newly decompensating patient because a front-desk colleague mentioned the tearful spouse, or because she keeps a habit of scanning recent messages and results that the overnight model run could not have included. The list is a starting point built from yesterday's data. The living, current truth of the panel lives in the human, in the team, and in the parts of the record the model has not yet ingested. Verifying the list against that living truth is not optional polish. It is the core of the job.
A Worked Example: The List Versus the Truth
Return to Dana's Monday morning and watch two versions of it unfold, because the contrast is the whole lesson.
In the first version, Dana treats the list as a verdict. Ranked first is Mr. Alvarez, flagged for a heart-failure admission six weeks ago. She spends forty minutes preparing for and making a careful outreach call, only to find he is doing well, already has his cardiology follow-up booked, and is faintly puzzled about why she is checking in again. Ranked second and third are two more patients recovering nicely from resolved events. She works down the list, diligent and thorough, clearing it top to bottom like a queue. She never reaches the middle of the list, because her morning is spent. She never calls Mrs. Booker, ranked twenty-second on stable data from ten days ago, whose home blood-pressure readings have been climbing all week and who stopped her own diuretic three days ago because it made her dizzy. The list was internally coherent and externally wrong, and Dana, by trusting it as an answer, executed its errors faithfully. Mrs. Booker is in the emergency department by Thursday.
In the second version, Dana treats the list as an input. She opens it and, before touching the phone, spends five minutes cross-checking the top of the list against what she actually knows and what the recent record shows. Mr. Alvarez, ranked first: she remembers his follow-up was booked, sees the cardiology note from last week, and demotes him with a quick glance, no forty-minute call required. She scans recent messages and results the overnight run could not have seen and catches the string of rising home blood pressures on Mrs. Booker, and the flag that her diuretic fills stopped. She pulls Mrs. Booker to the top of her own working list, above where the model placed her. She uses the AI to summarize Mrs. Booker's complex chart in thirty seconds, uses it to draft an outreach script about the stopped medication, verifies both against the source, and makes the call that actually matters. The same tool, the same list, the same morning. The difference is entirely in whether the human owned the panel or the list did.
Notice what Dana did not do in the good version. She did not throw the AI away; the summary and the script and even the ranking saved her real time and pointed her usefully. She did not distrust it reflexively into uselessness. She held it exactly where it belongs: a fast, informed, fallible assistant whose output she verifies against a reality only she can see, before she decides. That posture, neither blind trust nor blanket rejection, is the entire craft of working an AI-assisted panel.
The Outreach That Is Really a Clinical Decision
There is one more thing hiding inside Dana's good morning that deserves to be pulled into the light, because it is where care management quietly crosses into clinical care, and it is the place careless AI use does the most damage. When the AI drafted an outreach script about Mrs. Booker's stopped diuretic, it did not draft a friendly reminder to book an appointment. It drafted a communication about a medication and a symptom. A message that says, in effect, you stopped your water pill because it made you dizzy, please restart it or come in, is not administrative outreach. It is a clinical act. It touches a drug, a dose, and a symptom, and if the words are wrong, the consequence lands on a patient's body, not on a schedule.
Consider how many ways a plausible-looking drafted script can be clinically wrong. It might tell the patient to simply restart a diuretic that was in fact intentionally held by her cardiologist two days ago for a rising creatinine, in which case the model, working from stale or partial data, has just drafted advice that could harm her. It might name the wrong medication, or the wrong dose, or attribute the dizziness to the diuretic when the real driver was a new blood-pressure agent. It might reassure a patient whose symptom actually warranted urgent evaluation, or alarm one whose symptom was benign. Each of these is a clinical error dressed as a routine message. The AI does not know which; it produces fluent, confident text either way, and fluent confident text about a medication change is exactly the kind of output automation bias makes easy to send without a second look.
This is why the iron rule applies to the outreach draft with the same force it applies to a note or a summary. Any outreach that touches a medication, a symptom, a result, or a change in the plan is a clinical communication, and a clinical communication has to be verified against the source before it goes out. Verified means the care manager checks the actual current medication list, the actual recent notes, the actual reason a drug was started or stopped, and confirms that what the draft says is true for this patient right now. If the outreach carries clinical content she is not authorized or positioned to decide, restarting a held medication, interpreting a symptom, changing a dose, the safe move is not to send a polished AI draft but to route the decision to the clinician who owns it, then communicate the verified plan. The draft saves the typing. It never inherits the authority. The rule to carry is simple: if the message could change what a patient does with a drug or a symptom, treat drafting it and sending it as a clinical decision, not a clerical one, and verify before you press send.
Testing the Model, Monitoring for Drift, and Proving the Human Decided
Everything so far is about the care manager at her desk. But the biased-list danger is not something an individual can fully defend against alone, because she cannot see the model's disparate performance from her single panel. That defense lives at the level of the organization that selects, validates, and monitors the model, and every care manager should know it exists and should feel entitled to ask about it.
A risk model must be tested for disparate performance before deployment, meaning validated on data representative of the actual patient population, with its accuracy checked separately across racial, ethnic, socioeconomic, language, and other groups, not just in aggregate. Aggregate accuracy can look excellent while hiding a serious failure for a subgroup, exactly as the cost-proxy algorithm did. And testing once is not enough, because models drift: as populations, coding practices, and care patterns change, a model that was accurate at deployment can silently degrade. So the model must also be monitored after deployment, with its performance and its equity re-checked on an ongoing basis. This is precisely the kind of before-and-after bias and risk evaluation that emerging accreditation guidance now names as a foundational element of responsible AI use, and it is a fair question for any frontline clinician to put to their governance committee: how do we know this risk model performs equitably for our patients, and how are we watching it over time?
A frontline care manager does not need to be a data scientist to ask sharp governance questions. A short, concrete list of questions is often enough to tell whether a model has been treated as a serious clinical instrument or waved through on a demo. Reasonable questions to raise include:
- What does the model actually predict, and if it predicts cost or utilization rather than illness, how has the risk of the cost-proxy bias been checked and corrected?
- Was the model validated on data that looks like our patients, including our underserved patients, and not just on the vendor's development population?
- Was accuracy reported separately by racial, ethnic, socioeconomic, and language subgroups, and did any subgroup show worse performance than the aggregate?
- Who monitors this model for drift after go-live, how often, and what happens when its equity or accuracy slips?
- What are the known limits and failure modes the vendor has disclosed, and are those limits written down where a care manager can see them?
- If I believe the list is steering care away from patients who need it, who do I tell, and does that feedback actually reach the people who can change the model?
None of these questions require special standing. They are the ordinary due diligence of a professional who is being asked to act on a tool's output, and asking them is itself part of keeping accountability human.
What Verification Actually Looks Like on a List
Verification is not a vague virtue; it is a short, repeatable set of moves a care manager runs before she commits her scarce calls. In practice it looks like this:
- Scan the top for resolved cases. For each high-ranked patient, ask whether anything you already know, a recent call, a booked follow-up, a note from last week, means this person has stabilized since the model last ran. Demote the false positives quickly.
- Pull in what the model could not see. Scan recent messages, results, and team chatter for the newly decompensating patient the overnight run missed, and promote that patient above the position the model gave them.
- Sanity-check the reasons, not just the ranks. Read the one-line reason the model attached to each flag and confirm it matches the current chart, so you catch a flag built on stale or wrong data before you act on it.
- Verify any clinical content before it goes out. If the suggested action or drafted script touches a medication, symptom, or result, check it against the source record and, if it carries a clinical decision, route that decision to the clinician who owns it.
- Watch for the pattern, not just the patient. If the same kind of patient keeps being under-ranked, or a whole clinic rarely appears, treat that as a possible bias signal worth reporting upward, because you may be seeing disparate performance from your seat.
- Document the adjustments and the reasons. Record what you promoted, what you demoted, and why, so the record shows a human steered the panel.
And then there is the record, the quiet third leg of the iron rule. When Dana demotes Mr. Alvarez and promotes Mrs. Booker above where the model placed her, that decision, and the human reasoning behind it, should be visible in the record. Not because of paperwork for its own sake, but because the record is what proves a human owned the panel. It shows that the ranking was reviewed and adjusted by a named professional exercising judgment, not executed blindly by a queue-clearing machine. If a patient is missed, the record distinguishes a system where a human weighed the evidence and made a defensible call from one where nobody was really steering and the algorithm's output was simply obeyed. The record is how "the human decided" stops being a claim and becomes a fact.
The Iron Rule, in the Language of the Panel
Strip away the specifics and you are left with the same three-part rule that governs every safe use of AI in healthcare, rendered here in the dialect of care management. AI assists: it stratifies the panel, summarizes the chart, flags the care gaps, suggests who to call, and drafts the outreach, and it does all of this fast enough to give the human back her most precious resource, time. The human decides: a named care manager owns the panel, verifies the ranked list against the living reality of her patients, corrects for the staleness the model cannot see and stays alert to the bias the model may carry, and chooses who actually gets today's scarce attention. The record proves it: the adjustments, the reasoning, and the human ownership are documented, so that the panel is demonstrably steered by a person and not merely processed by a model.
The reason this ordering is not negotiable is the stakes of getting it backwards. When the list becomes the verdict, its two failure modes, systematic bias against the already-underserved and staleness that hides the newly sick, are not caught by anyone, because the one human positioned to catch them has abdicated to the machine. When the list stays an input, those same failures become survivable, because a person who owns the panel is standing exactly where the model's errors would otherwise reach the patient. The AI does not make the care manager obsolete. It makes her more necessary, because someone has to be the reality check on a confident machine, and that someone has a name.
Key Takeaways
- Care management rations one scarce resource, human attention, across a panel of complex patients, and AI enters exactly at the point of deciding who gets the call: risk stratification, chart summaries, care-gap flags, and drafted outreach scripts.
- The load-bearing principle is that a named human care manager owns the panel and decides who gets attention. The AI-generated list is an input to verify, not a verdict to execute.
- The central danger is that a biased model fails systematically, steering scarce care-management resources away from patients already underserved, as in the classic cost-proxy algorithm that read Black patients' under-treatment as wellness and under-identified them for extra care.
- Disparate performance means a model works less well for some populations than others; deployed into a workflow, that blind spot becomes a pattern of who gets cared for and who does not.
- A stale list fails in mirror images: it surfaces a patient who has already stabilized and no longer needs help, while burying the newly decompensating patient whose slide began after the model last ran.
- The defense against staleness is the care manager's own current knowledge of her panel, kept deliberately in the loop, plus the parts of the record the overnight model run could not yet see.
- Risk models must be tested for disparate performance on representative data before deployment and monitored for drift and equity after, not just checked in aggregate; frontline clinicians are entitled to ask how their model is validated and watched.
- The iron rule in the language of the panel: AI assists (stratify, summarize, flag, draft), the human decides (own the panel, verify against reality, correct for bias and staleness), and the record proves a named person, not a machine, steered the care.
Skill.re