Predictive, Generative, and Assistive AI - Three Different Tools
Three AI tools land on the same nurse's screen in a single shift. A sepsis model flashes a red risk score on a patient in bed 12. An ambient scribe drafts her admission note. A chatbot in the corner of the EHR offers to answer a question about a heparin drip. She is told all three are "AI," and if she extends the same trust to each, she will misuse at least two of them, because these are not three flavors of one thing. They are three fundamentally different kinds of tool, with different jobs, different failure modes, different regulators watching them, and three completely different questions you must ask before you rely on them. This lesson is about telling them apart, because the trust you owe a risk score is not the trust you owe a note draft, and confusing the two is how good clinicians get burned.
Why the Category Decides the Trust
In the last lessons we took AI apart by its underlying machinery. Now we look at it the way it actually arrives at your workflow: as a product in one of three broad categories. Predictive AI tells you something is likely. Generative AI writes something for you. Assistive AI acts alongside you inside a task. Each category earns a different kind of trust because each fails differently and carries a different consequence when it does. A clinician who learns to instantly place a new tool into one of these three buckets has a ready-made set of the right questions, and asking the right question is most of safety.
The reason this matters so much is that the failure of each category looks nothing like the others. A predictive tool fails by being miscalibrated or biased, quietly, on a whole population. A generative tool fails by inventing a specific fact, loudly plausible, on a single output. An assistive tool fails by doing part of a task subtly wrong and letting the human's guard down through sheer convenience. If you carry one mental model of "AI error" and apply it everywhere, you will look for the generative failure in the predictive tool and miss the one that was actually there. The categories are not academic. They tell you where to look.
Predictive AI: The Tool That Tells You Something Is Likely
Predictive AI takes data and outputs a probability or a classification about the future or the unseen: a sepsis risk score, a deterioration alert, a no-show predictor, a readmission risk, a flag on a chest film. This is the oldest and most heavily studied category in clinical AI, and in 2026 it is also the one the regulators moved on first. Under the federal HTI-1 rule, this family, called predictive decision support intervention, now carries transparency obligations: the tool is supposed to disclose a defined set of facts about how it was built, what data trained it, how it performed, and where it may not apply. That transparency exists precisely because of how predictive AI fails.
A predictive model does not invent a dramatic falsehood. It fails in three quieter ways, and each is dangerous exactly because it is undramatic. It can be miscalibrated, so its 80 percent does not really mean 80 percent in your hospital. It can be biased, performing worse for a subgroup that was underrepresented when it was trained, so it silently underserves the patients already underserved. And it can drift, degrading over months as your population, your documentation habits, or your lab vendors change out from under it. None of these announces itself on the screen. The number looks just as crisp when it is wrong as when it is right.
So the trust question for predictive AI is not "does this sound right." It is "what are its numbers, and do they apply to this patient." What was its sensitivity and specificity, and on whom. Was it validated on a population like yours. When it fires, what is the actual chance the thing it predicts is true, its positive predictive value, which in a low-prevalence setting can be far lower than the impressive-sounding accuracy implies. And above all, the cardinal move for this category: the score is a probability about a population and an input to your judgment, never a verdict about the person in the bed. When the sepsis model fires on bed 12, the right response is to go look at the patient, not to treat the number.
The low-prevalence trap is worth making concrete, because it is where confident clinicians are most often surprised. Suppose a deterioration model advertises 90 percent sensitivity and 90 percent specificity, numbers a vendor would put on a slide with pride. Now run it on a ward where only two of every hundred patients will actually deteriorate. Of the ninety-eight who will not, roughly ten still trip the alert, while fewer than two of the true events light up. The alert you are staring at is wrong far more often than it is right, not because the model is bad, but because the event is rare, and rare events punish even a strong classifier's positive predictive value. This is not an argument to ignore the alert. It is an argument to treat it as a prompt to look, weigh it against everything else you know about the patient, and never let a single flag override a clinical picture that disagrees with it. A clinician who understands why the crisp 90 percent does not mean what it sounds like is far harder to stampede.
Generative AI: The Tool That Writes for You
Generative AI produces new content: a note, a patient message, an answer, a letter, a summary. We spent a whole lesson on its machinery, so here the point is how its trust profile differs from the predictive tool sitting one window over. Where the predictive model fails quietly across a population, the generative model fails loudly on a single output, by inventing a specific, plausible, wrong fact: a lab value, a citation, a medication, a symptom the patient denied or affirmed. Its danger is not miscalibration; it is fabrication delivered in flawless prose.
The regulators treat it differently too, and a clinician should feel the difference. Because generative tools increasingly speak to patients in the clinician's voice, the newest state laws target this category specifically. California's AB 3030 requires that when generative AI produces a patient's clinical communication, the patient be told and be given a way to reach a human, unless a licensed clinician reviewed the message first. Texas now requires disclosure when AI is used in diagnosis or treatment. The through-line is that generative output aimed at a patient is a regulated act, and the safe habit is to assume disclosure is required and to keep a human in the loop by default.
So the trust question for generative AI is the one from the previous lesson: is this reshaping content I gave it, in which case verify against the source quickly, or is it asserting a specific external fact, in which case verify against a trusted source before it informs care. The same tool can do both in one session, so the question is asked per output, not per tool.
Consider how differently the two failures land on a chart. A miscalibrated sepsis score that reads 60 percent when the real risk is 40 percent produces no artifact anyone will ever point to in a deposition; it simply nudges attention slightly wrong across hundreds of patients, and no single note carries the fingerprint. A generative fabrication is the opposite: it leaves a specific, quotable line in the legal record, a penicillin allergy rendered as "no known drug allergies," a lateral tear documented on the wrong knee, a dose that was never ordered. That line is exactly what a coding auditor, a plaintiff's attorney, or a Joint Commission surveyor reads back to you two years later, and "the AI drafted it" is not a defense, because your signature attested to it. The predictive failure is a population problem you manage with statistics and monitoring; the generative failure is a single-document problem you manage by reading the high-stakes claims before you sign. Same word, "error," two entirely different jobs.
A predictive tool asks for your skepticism about its numbers. A generative tool asks for your skepticism about its facts. An assistive tool asks for your skepticism about your own attention. Three tools, three different guards.
Assistive AI: The Tool That Acts Alongside You
The third category is the fastest-growing and the trickiest, because it hides. Assistive AI does not just tell you something or write you something; it takes part in a task with you, often stitching prediction and generation together and threading itself into the workflow. The ambient scribe that listens, transcribes, structures, and drafts is assistive. The inbox tool that reads messages, sorts them, and proposes replies is assistive. The emerging agentic tools that can take multi-step actions, drafting an order, queuing a referral, populating a form, are assistive at the far end of the spectrum. These tools are wonderful precisely because they remove friction, and that is exactly what makes them dangerous.
The failure mode of assistive AI is not primarily technical. It is human, and it has a name we will keep returning to: automation bias. When a tool does most of a task smoothly and reliably for weeks, the human doing the last-mile check relaxes. The convenience that makes assistive AI valuable is the same force that erodes the vigilance that keeps it safe. An assistive tool that is right 95 percent of the time is, in a strange way, more dangerous than one that is right 70 percent, because the 95 percent trains you to stop looking, and then the improbable error in the other 5 percent sails straight through your lowered guard and into the record or the order.
So the trust question for assistive AI is different again, and it is uncomfortable because it points at you. Not "is the tool accurate," but "where exactly is my verification step, and is it strong enough to survive my own complacency." The safe use of assistive AI is a designed verification habit that does not depend on your mood or how good the last hundred outputs were. It is the checkpoint you keep even when, especially when, the tool has earned your trust. We will build these checkpoints deliberately in later levels; here the job is to recognize when you are holding an assistive tool and therefore that your own attention is the thing most at risk.
There is a further wrinkle that makes assistive tools the category to watch most carefully, and it is the direction the market is moving. The early assistive tools mostly drafted something and stopped, leaving a clear artifact for you to inspect before you signed. The newer ones increasingly act: they queue an order, populate a referral, file a form, close a loop, and sometimes do several of these in a chain before you see any of it. When a tool takes an action rather than producing a draft, the verification step is no longer a comfortable pause between output and signature; it may be the only thing standing between an automated action and the record. The more of a task a tool completes on its own, the earlier and firmer your checkpoint has to be, because there is less friction left to catch the error on its way through. The design question for these tools is not "did it write something good" but "at which exact step does a human have to look, and what is that human actually able to see when they look." A verification step that arrives after the action is not a verification step; it is an incident review.
A Regulator for Each: Why the Rulebooks Split
One of the clearest signs that these categories are genuinely different is that the people who write the rules treat them differently. If AI were one thing, there would be one rulebook. Instead, the oversight has split along exactly the category lines we have been drawing, because each category creates a different kind of harm that a different mechanism has to catch. You do not need to memorize the statutes, but seeing how they map onto the three categories makes the distinctions concrete and shows you that this framework is not a teaching device someone invented; it is the shape the whole field is organizing itself around.
| Category | What it does | Signature failure | Who is watching in 2026 |
|---|---|---|---|
| Predictive | Outputs a probability or classification (risk score, alert, flag) | Quiet miscalibration, bias, and drift across a population | ONC HTI-1 predictive decision support intervention transparency; FDA when the tool is a regulated device |
| Generative | Writes new content (note, message, answer, summary) | Loud fabrication of a specific plausible fact in one output | State disclosure laws such as California AB 3030 and Texas TRAIGA when output reaches patients |
| Assistive | Acts alongside you in a task, often combining the above | Automation bias eroding the human verification step | Joint Commission and CHAI responsible-use expectations; inherits both regimes above |
Read the table as a map of accountability. The predictive column is watched because its failures are invisible to the user, so the rule forces the tool to disclose how it was built and where it may not apply. The generative column is watched because its output can reach a patient in your voice, so the rule forces disclosure and a path to a human. The assistive column inherits both concerns and adds the human-factors one, which is why the first accrediting-body guidance leans so hard on governance and workforce training. The pattern to carry away is simple: the regulators are not confused about what AI is. They have already sorted it into these buckets, and so should you.
Telling Them Apart in Five Seconds, Then One Shift That Uses All Three
You will rarely be told which category a tool belongs to; the login screen just says the vendor's name. So here is a fast field test you can run on anything that lands in your workflow. Ask what the tool hands back to you. If it hands you a number, a probability, a flag, or a category, it is predictive, and your guard is its numbers and whether they apply to this patient. If it hands you written content you did not write, a note, a message, an answer, a summary, it is generative, and your guard is verifying what it asserts. If it does part of a task for you and moves it forward, listening and drafting, sorting and replying, populating and queuing, it is assistive, and your guard is your own verification step, held firm against the comfort of its convenience.
Many real tools are blends, and that is fine; the test still works because you apply every guard that fits. An ambient scribe hands you written content (generative guard: check the facts) inside a task it moved forward for you (assistive guard: keep your read-before-sign step). A risk-stratification dashboard that also drafts an outreach message is predictive (scrutinize the score) and generative (verify the message) at once. You are not trying to force each tool into a single box. You are collecting the full set of guards the tool earns, and the five-second test is how you assemble that set without a manual, a committee, or a vendor's reassurance.
Notice what the test deliberately ignores. It does not ask the vendor's name, the funding round, the market-share bar, or how many health systems have signed. Those are the very facts a demo leads with, and none of them changes which guard the tool has earned. A tool used by twenty thousand physicians and a tool piloted on one unit last month, if both hand you a drafted note, both earn the generative guard in full. Reputation buys the tool a place on your screen; it buys the output nothing. This is the small discipline that keeps automation bias from creeping in at the level of the brand rather than the single output: you owe the category's questions to every tool in it, including the ones you like.
Now watch a calibrated professional run the whole test live, because the five-second habit only earns its keep under real time pressure. Return to our nurse and follow her through a single shift with all three tools. The sepsis model fires on bed 12. She does not dismiss it and she does not treat the number; she recognizes a predictive tool, asks whether this patient resembles the ones it works on, and goes to the bedside to gather the real signals. The patient is warm, tachycardic, mottled at the knees, and trending the wrong way on a lactate she pulls herself, and the alert has earned action, so she escalates, using the score as the nudge it is meant to be rather than the verdict it never was.
The ambient scribe drafts her admission note. She recognizes a generative tool wrapped in an assistive workflow, and she reads the high-stakes claims before she signs: the medication list, the allergies, the pertinent negatives. It has rendered a penicillin allergy as "no known drug allergies," a classic confident fabrication, and she catches and fixes it because she knew a generative tool can invent and an assistive one can lull her into not checking. Then the chatbot offers guidance on the heparin drip. She recognizes a generative tool asserting a specific external fact, a dose and a protocol, and she treats it as an unverified claim, confirming the actual number against the pump protocol before she touches the drip. Three tools, three different questions, three different guards, one shift. That fluency, knowing instantly which kind of tool is in front of her and therefore which risk to watch, is the entire skill of this lesson.
The Cost of Getting the Category Wrong
It is worth making the failure concrete, because the danger of miscategorizing a tool is not abstract. Picture a clinician who mentally files the ambient scribe as "just a generative writing tool" and applies the generative guard alone: they diligently check whether the facts in the note are true. Good, but incomplete, because they have missed that it is also assistive, and so they never noticed that their own read-before-sign habit had quietly decayed over three months of reliable output. The facts they spot-check are fine; the allergy the tool dropped on a rushed afternoon is not, and it slips through because the guard they skipped was the one this category most needed. The error was not ignorance of AI. It was applying the right guard for the wrong category and feeling safe while doing it.
The reverse is just as costly. A clinician who treats a predictive sepsis score as if it were a generative claim goes looking for a fabricated fact, finds none, concludes the tool is trustworthy, and never asks the questions that actually matter for a predictor: was it validated on patients like mine, what is its positive predictive value here, is it drifting. Feeling reassured by the wrong check is more dangerous than no check at all, because it manufactures false confidence. This is precisely why the categories are worth the small effort to learn. They are not trivia about how AI is built. They are a map of where the danger lives in each tool, and a clinician who reads that map correctly spends their limited attention exactly where a patient is most likely to be harmed.
There is a subtler version of the same error that catches experienced clinicians specifically. The more fluent you are with one category, the more tempting it is to apply its habits everywhere, because those habits have served you well. A radiologist steeped in imaging AI learns to scrutinize intended-use boundaries and may carry that same reflex to an ambient scribe, where it is beside the point. A hospitalist who has been burned by a hallucinated citation may over-index on fact-checking a predictive dashboard whose real risk is drift, not fabrication. Expertise in one district does not transfer to another; if anything, it can disguise the gap, because competence in the wrong frame feels like competence. The remedy is not more suspicion in general. It is the deliberate first step of naming the category before you choose the guard, so that your considerable skill gets pointed at the risk that is actually present rather than the one you happen to be practiced at spotting.
Key Takeaways
- AI arrives in three product categories, and each earns a different kind of trust: predictive (tells you something is likely), generative (writes something for you), and assistive (acts alongside you in a task).
- Predictive AI fails quietly across a population through miscalibration, bias, and drift. Ask for its numbers, whether it was validated on a population like yours, and its positive predictive value, and treat the score as an input, never a verdict.
- Predictive tools are now governed as predictive decision support interventions under the federal HTI-1 rule, which requires transparency about how they were built and where they may not apply.
- Generative AI fails loudly on single outputs by inventing plausible, specific facts. Verify transformation against the source and any asserted external fact against a trusted source.
- Generative output aimed at patients is a regulated act. California's AB 3030 and Texas TRAIGA require disclosure, so assume disclosure is needed and keep a human in the loop by default.
- Assistive AI fails through your attention, not just its accuracy. Its convenience erodes vigilance, and a tool that is usually right is what trains you to stop checking, so the guard is a designed verification step you keep regardless of mood.
- The reusable skill is instant categorization: place any new tool into predictive, generative, or assistive, and the right questions come with it.
- Different tools, different guards: skepticism about the numbers, skepticism about the facts, and skepticism about your own attention.
Skill.re