HIPAA, PHI, and AI Tools
A hospitalist, three patients behind and staring at a discharge that will not write itself, opens a consumer AI chatbot in a browser tab and pastes in the whole messy story: the patient's name, the admission date, the room number, the biopsy result, the family situation, all of it, and types "turn this into a clean discharge summary." Ten seconds later a beautiful paragraph appears. It is also, in that instant, a HIPAA breach. Not a gray area, not a technicality that a good lawyer could argue away. The protected health information of a real, identifiable patient has left the covered entity and entered a system with no contract governing what happens to it, and nothing the hospitalist does next can un-send it.
What PHI Actually Is, in Plain Terms
Protected Health Information, PHI, is individually identifiable health information: any health data that is tied to, or could reasonably be tied to, a specific person. That coupling of two things, a health fact and an identity, is the whole idea. A lab value floating free of any name is not PHI. A lab value attached to Mr. Alvarez in room 412 is. HIPAA lists eighteen categories of identifiers that make information identifiable, and the list is broader than most clinicians assume: not just the obvious ones like name, medical record number, and Social Security number, but also dates more specific than a year (admission date, date of birth, date of death), all geographic subdivisions smaller than a state, phone and fax numbers, email addresses, device identifiers, full-face photographs, and a catch-all for any other unique identifying number or characteristic. The reason the list matters is that clinicians routinely think they have de-identified a story when they have only removed the name, leaving a dozen other identifiers behind that, in combination, point straight back to one human being.
This is why the free-text clinical narrative is so treacherous. A structured field is easy to strip. A paragraph of prose is not, because identity hides inside it in ways you do not notice while you are typing: the patient who is a well-known local figure, the rare diagnosis in a small town that narrows the field to one person, the phrase "her husband, the fire chief," the exact date of the accident that made the news. You can delete the name and still have written something that identifies the patient to anyone who knows the community. The eighteen-identifier list is not a checklist you can run in your head at speed under pressure. It is a warning that identifiability is slippery, and that the safest assumption about any real patient's story is that it is PHI until proven otherwise.
The Eighteen, as They Show Up in Your Day
It helps to see the list not as a legal appendix but as the ordinary furniture of a clinical note. Names, yes. Geographic subdivisions smaller than a state, which quietly means the town, the county, the street, and the ZIP code you wrote without thinking. All dates tied to an individual that are more specific than a year: birth date, admission date, discharge date, the date of the procedure, the date of death, and even ages over eighty-nine, which HIPAA singles out because a very old age in a small population points to few people. Telephone and fax numbers. Email addresses. Social Security numbers. Medical record numbers. Health plan beneficiary numbers. Account numbers. Certificate and license numbers. Vehicle identifiers including license plates. Device identifiers and serial numbers, which is why the pump or the pacemaker in your note can carry identity with it. Web URLs. IP addresses. Biometric identifiers such as fingerprints and voiceprints. Full-face photographs and comparable images. And the eighteenth, the one that swallows the rest: any other unique identifying number, characteristic, or code. That last category is the reason you can never fully relax. The "retired opera singer with a transplanted kidney and a service dog named after a composer" has no name in the sentence and is nonetheless one person.
Consider a single line a hospitalist might dictate: "Mr. Okafor, seen 3/14 in the Blue Ridge clinic, is a 91-year-old former state senator readmitted after his LVAD alarmed." Strip the name and you still have a date, a named clinic, an age over eighty-nine, a public role, and a distinctive device. Five of the eighteen categories are still in the sentence, and any local reader knows exactly who this is. The lesson is not to memorize the list and run it at speed. It is to internalize that identity is sticky, and that the honest default for any real patient's story is to treat it as PHI.
The Breach, and What a BAA Actually Changes
Here is the rule that anchors everything in this lesson: pasting a patient's identifiable story or identifiers into a consumer AI tool that has no Business Associate Agreement with your organization is a HIPAA breach, full stop. Not because the tool is malicious, and not because anyone necessarily reads what you typed. It is a breach because you disclosed PHI to a third party that is not permitted to receive it, and HIPAA governs the disclosure itself, not just what the recipient does afterward. The harm is legal and structural, and it exists the moment the data leaves. Whether that data is later used to train a model, cached on a server, seen by a contractor, or exposed in that vendor's own breach is entirely outside your control, which is exactly the point: you have handed the patient's information to a party with no obligation to you or to them.
A Business Associate Agreement, the BAA, is the contract that changes this. A business associate is a vendor that handles PHI on behalf of a covered entity, and the BAA is the legal instrument that binds that vendor to HIPAA's rules: to safeguard the data, to use it only for the permitted purpose, to not sell it or train on it without authorization, to report breaches, and to accept liability for its own failures. When your organization has a signed BAA with an AI vendor, that vendor is inside the compliance perimeter, and sending PHI to it for a permitted purpose is lawful. When there is no BAA, the vendor is a stranger, and sending PHI to it is a disclosure to an unauthorized party. The BAA is the single line that separates a sanctioned enterprise tool from a consumer chatbot, and it is invisible from the user interface. The same-looking chat box can be either one. You cannot tell which by looking at it, which is precisely why the question has to be answered before, not after.
What the BAA Does, and What It Does Not
Picture the same underlying model reached through two doors. Through the first door, an enterprise agreement your health system signed, the vendor has promised in writing that your inputs will not be used to train its models, that data is segregated and encrypted, that access is logged, that sub-contractors are bound by the same terms, and that if anything leaks the vendor tells you and shares liability. Through the second door, the free consumer app, the only thing you have agreed to is a terms-of-service page you did not read, which in many cases reserves the right to review conversations to improve the product. The words on the screen are identical. The legal reality behind them is not even close. A BAA is not a magic shield that makes any use safe; it is a defined allocation of duties. It does not authorize you to disclose more than the permitted purpose requires, it does not turn off the minimum necessary standard, and it does not make the vendor infallible. What it does is bring the vendor into the same accountable system you already work inside, so that when PHI flows to it, the flow is governed rather than orphaned.
There is a second, quieter thing a BAA changes: who answers when something goes wrong. Without a BAA, an impermissible disclosure is entirely your organization's problem, and the vendor owes you nothing. With one, the vendor has contractual obligations to safeguard, to report, and to cooperate, and it carries defined liability for its own failures. The BAA does not erase your responsibility, but it stops the buck from disappearing into a company that never agreed to be responsible in the first place. That is why the phrase compliance perimeter is apt. The perimeter is the boundary of people and vendors who have all agreed to the same rules. A BAA moves a vendor from outside that boundary to inside it. Pasting PHI into a tool outside the boundary is not a smaller version of the same act; it is a categorically different act, because the receiver has taken on none of the duties that make the disclosure lawful.
The data you never put into a tool cannot leak from it. Minimum necessary is not a compliance nicety; it is the one control that works even when every other safeguard fails.
The First Question: Enterprise or Consumer
Before any real patient data touches any tool, the first question is not "is this AI any good" but "is this a BAA-covered tool my organization has sanctioned for PHI, or a consumer tool that is not." Everything downstream depends on the answer, and the two categories can be nearly indistinguishable on screen. An enterprise deployment of a large language model, purchased by your health system, configured under a BAA, with data handling terms that forbid training on your inputs, may present the exact same chat interface as the free public version of the same underlying model that has no such protections. The interface is not the contract. The clinician who assumes that because the tool is "the same AI everyone uses" it must be fine has made a category error with legal consequences.
This is where individual judgment meets organizational governance, and where the clinician's job is mostly to know which lane they are in. Your institution's compliance and IT functions decide which tools are approved for PHI, negotiate the BAAs, and publish the list. Your job is to use only what is on that list for anything involving a real patient, and to treat everything else, every personal app, every free chatbot, every browser extension, every tool you found yourself and liked, as off-limits for PHI until it has been sanctioned. The correct instinct when a genuinely useful but unapproved tool appears is not to quietly start feeding it patient data; it is to route it to the people whose job is to evaluate and, if appropriate, contract for it. Shadow IT, the unapproved tool a well-meaning clinician adopts on their own, is one of the most common ways PHI walks out of a covered entity in 2026.
The Shadow IT Moment, Up Close
Shadow IT rarely looks like recklessness. It looks like a tired clinician trying to do right by patients. A nurse manager on a short-staffed medical-surgical floor finds a free browser extension that summarizes long handoff notes into a tidy SBAR, and it is genuinely good. She starts using it during report because it saves her ten minutes she does not have. She has told no one, signed nothing, and has no idea where the text goes after she pastes it. That extension is now a business associate her organization never contracted with, receiving PHI it has no authority to hold. The failure here is not bad intent; it is a governance gap that a well-meaning person filled with a tool she trusted because it worked. Multiply her by a few hundred staff across a system and you have described how PHI quietly leaves covered entities every day.
The right reflex is small and specific. When you find something that genuinely helps, treat that as a signal worth escalating, not a secret worth keeping. The message to your informatics or compliance team is short: "I found a tool that saves real time on handoffs; can we evaluate it and get a BAA in place?" That sentence does two useful things at once. It gets a valuable tool onto the path toward safe, sanctioned use, and it keeps you out of the incident report. The clinician who says nothing and keeps pasting is not being efficient; she is running an unmonitored disclosure pipeline with her name on every entry. Governance is not the enemy of the useful tool. It is the only route by which a useful tool becomes a safe one.
Minimum Necessary: The Habit That Protects You
HIPAA's minimum necessary standard says you should use, disclose, or request only the least PHI needed to accomplish the task at hand. It is one of the oldest ideas in the rule and one of the least glamorous, which is exactly why it gets overlooked in favor of flashier controls like encryption and access logs. But it has a property those controls lack: it reduces risk by reducing exposure at the source rather than by managing exposure after it exists. It was written for a paper-and-fax world, but it turns out to be the single most powerful mental habit for the age of AI tools, because of a simple asymmetry: the data you never put into a tool cannot leak from it, cannot be trained on, cannot surface in someone else's breach, and cannot appear in a subpoena. Every identifier you leave out is a risk that ceases to exist rather than a risk you are hoping the vendor manages well. Minimum necessary turns privacy from something you delegate to a contract into something you control at the keyboard. It is the one privacy control that keeps working even if the BAA is somehow flawed, the vendor is later breached, or the tool's data-handling turns out to be worse than promised, because a piece of information that was never entered cannot be exposed by any downstream failure at all.
In practice this reshapes how a careful clinician uses even a fully sanctioned, BAA-covered tool. If you want an AI to help you phrase patient-education language about managing new-onset atrial fibrillation, you do not need to include the patient's name, date of birth, or medical record number to get that help; the clinical question stands on its own without a single identifier. If you want help structuring a differential, you can describe the presentation without the identifying specifics that make it this patient rather than a teaching case. The goal is to get into the reflex of asking, before you type, "what is the least I can include and still get what I need?" On an approved tool, minimum necessary is defense in depth: even inside the perimeter, less exposure is less risk. On an unapproved tool, minimum necessary is the difference between a breach and a clinical question that never involved PHI at all.
De-Identification Is Real, and Easy to Botch
There is a legitimate path by which health information leaves HIPAA's scope entirely: de-identification. HIPAA recognizes two methods. The first, Safe Harbor, requires removing all eighteen categories of identifiers, at which point the information is no longer considered identifiable and is no longer PHI. The second, expert determination, uses a qualified statistician to certify that the risk of re-identification is very small. Properly de-identified data can be used more freely, including in some AI tools, because it is no longer protected health information. This sounds like a clean escape hatch, and in structured, carefully handled datasets it can be. The danger is believing you have done it when you have not.
The classic failure is free-text re-identification. A clinician deletes the patient's name from a narrative and believes the story is now de-identified, but the eighteen identifiers include dates, locations, and any unique characteristic, and prose is full of these in ways that resist quick redaction. "A 34-year-old woman admitted last Tuesday after the pileup on the interstate, whose case was in the local paper," names no one and identifies exactly one person. Removing the name is not de-identification; it is a cosmetic gesture that leaves the re-identification risk almost fully intact. Safe Harbor is a demanding, complete removal of a long list, not a quick scrub of the obvious. The honest posture for a frontline clinician is to treat de-identification of free-text patient stories as something you almost certainly cannot do reliably at the keyboard in the flow of work, and therefore to lean on minimum necessary and approved tools instead of talking yourself into believing a lightly edited paragraph is safe.
The distinction between the two lawful methods matters because clinicians conflate them. Safe Harbor is a rule you can in principle follow yourself: remove every one of the eighteen categories, in full, and add the requirement that you have no actual knowledge the remaining information could re-identify anyone. Expert determination is a judgment call made by someone qualified in statistical disclosure methods, who analyzes the specific dataset and its context and certifies that the risk of re-identification is very small. Neither is a name-deletion pass done in ninety seconds between patients. A dataset that has been through proper Safe Harbor or expert determination is genuinely outside HIPAA and can, for example, be used in tools that would be forbidden for PHI. But that outcome is earned by rigorous process, not asserted by an optimistic clinician who removed the obvious labels and hoped for the best.
Removing the name is not de-identification. It is redecorating the front of a house whose address is still printed on the mailbox.
There is a subtler trap worth naming: the illusion that de-identifying makes verification unnecessary. It does not. Even when data is genuinely outside HIPAA, the iron rule of clinical AI still holds. Any output that will touch a patient or the record must be verified by a human, because a de-identified input does not guarantee a correct output, and "the model produced it" is never a substitute for a clinician confirming it. De-identification is a privacy control, not a quality control. The two problems are separate, and you owe the patient both.
What a Breach Actually Sets in Motion
It helps to understand what "a breach, full stop" actually triggers, because the phrase can sound abstract until you see the machinery it starts. Under HIPAA's Breach Notification Rule, an impermissible disclosure of unsecured PHI is presumed to be a reportable breach unless the covered entity can demonstrate, through a formal risk assessment, a low probability that the information was compromised. That assessment weighs the nature and extent of the PHI involved, who received it, whether it was actually acquired or viewed, and the extent to which the risk has been mitigated. Pasting a patient's identifiable story into a public AI tool fails most of those factors: sensitive clinical detail, an unauthorized recipient with no BAA, no ability to confirm the data was not retained or seen, and no meaningful way to claw it back. In practice, that is very hard to argue down to low probability, which is precisely why the safe assumption is that it is reportable.
Reportable means real, escalating obligations. The organization must notify affected individuals, and depending on the number involved, notify the Department of Health and Human Services and potentially the media. There is the direct cost of investigation and notification, the reputational cost, the erosion of patient trust, and the possibility of civil monetary penalties that scale with the degree of culpability, with the harshest tier reserved for willful neglect. For the clinician who caused it, there is the internal consequence: the incident report, the conversation with compliance, the mark on a professional record, the mandatory retraining. None of this requires that any patient was ever actually harmed by the disclosure. The breach is complete and consequential the moment the PHI left, regardless of outcome, which is exactly the mental model to carry: you are not gambling on whether the data gets misused. You have already lost control of it, and the loss of control is the injury HIPAA names.
A Worked Example: Two Ways to Ask the Same Question
Watch the same clinical need handled two ways. A nurse practitioner wants help drafting patient-friendly instructions for a newly diagnosed diabetic. The unsafe version: she opens a personal, free chatbot on her phone, and types, "Write discharge instructions for Maria Delgado, DOB 3/12/1961, MRN 00814537, seen today at Riverside Clinic, new type 2 diabetes, A1c 9.4, starting metformin, lives alone, limited English." That single message disclosed a name, a date of birth, a medical record number, a facility, a visit date, and a clinical picture to a tool with no BAA. It is a breach, and it happened in the time it took to type it. The output may be lovely and the patient may never be harmed, and none of that changes the fact that identifiable PHI was disclosed to an unauthorized third party.
The safe version: she uses her health system's BAA-covered, sanctioned AI tool, and even there she applies minimum necessary, typing, "Draft plain-language instructions for a patient with newly diagnosed type 2 diabetes starting metformin, at roughly a 6th-grade reading level, covering what the medication is for, how to take it, common side effects, and when to call. Keep it under 200 words." No name, no date of birth, no MRN, no facility, no visit date. The clinical request is intact; the identity is absent. She gets the same useful draft, verifies it for accuracy before it reaches the patient, and adds the patient-specific details herself inside the record. Two paths to the same paragraph. One created a reportable breach and a permanent loss of control over a patient's data. The other never let PHI leave the perimeter at all. The difference was not the AI. It was two decisions made before the first keystroke: which tool, and how much to include.
Verify, Do Not Repeat Blindly
The safe version does not end when the draft appears. Watch what a careful nurse practitioner does with that paragraph before it becomes patient-facing. The model wrote, "Take metformin twice daily with meals; a common early side effect is low blood sugar." She stops. Metformin as monotherapy does not typically cause hypoglycemia; that sentence is a confident error the model produced because diabetes and low blood sugar travel together in its training text. She corrects it to name the actual common effects, gastrointestinal upset that usually eases over a week or two, and she keeps the instruction to call if certain symptoms appear. This is the whole discipline in miniature. The AI is a fast first drafter, not a source of truth, and "the AI said so" is not verification. Her signature on the final instructions is an attestation that a licensed human read every line and stands behind it. The privacy decision protected the patient's identity; the verification decision protects the patient's body. A gold-standard clinician makes both, every time, and the record shows she did.
The same posture applies to the far more common ambient-scribe workflow, where a sanctioned, BAA-covered tool listens to the visit and drafts the note. The privacy question is already answered by the enterprise contract, but the accountability question is not. If the draft note records a normal neurologic exam you did not perform, or flips the laterality of a lesion, or drops the one abnormal value that changes the plan, signing it unread makes that error yours. Minimum necessary keeps identity from leaking to the wrong place; verification keeps confabulation from leaking into the legal record. Both failures are chart-review findings waiting to happen, and both are prevented by a human who refuses to let the AI have the last word.
Key Takeaways
- PHI is individually identifiable health information: a health fact coupled to an identity. HIPAA's eighteen identifier categories go well beyond names to include dates, locations, and any unique characteristic, which is why free-text patient stories are so easily identifiable even after you delete the name.
- Pasting an identifiable patient story or identifiers into a consumer AI tool with no Business Associate Agreement is a HIPAA breach, full stop, because the disclosure to an unauthorized party is itself the violation, regardless of what the vendor later does.
- A BAA is the contract that brings a vendor inside HIPAA's rules: safeguard the data, use it only for the permitted purpose, do not train on or sell it without authorization, report breaches. It is the line between a sanctioned enterprise tool and a consumer chatbot, and it is invisible from the interface.
- The first question before any real patient data touches any tool is enterprise-or-consumer: is this a BAA-covered tool my organization approved for PHI, or not. The same-looking chat box can be either. Use only the approved list for real patients.
- Minimum necessary, using only the least PHI needed for the task, is the most powerful habit in the AI era, because the data you never enter cannot leak, be trained on, or surface in someone else's breach. It is a control you hold at the keyboard.
- De-identification can take data outside HIPAA, but Safe Harbor requires removing all eighteen identifier categories, and deleting only the name from a narrative is not de-identification. Free-text re-identification risk is high; do not talk yourself into believing a lightly edited story is safe.
- Shadow IT, the unapproved tool a well-meaning clinician adopts alone, is a leading way PHI leaves a covered entity. When a useful but unsanctioned tool appears, route it to compliance and IT rather than feeding it patient data.
- Even on an approved tool, apply minimum necessary as defense in depth. The safest clinical AI habit is two decisions made before the first keystroke: which tool, and how little to include.
Skill.re