Data Governance and PHI at Scale
A specialty pharmacy enterprise discovered, during a routine accreditation prep, that it did not actually know where its patients' data went. Patient charts flowed into an AI prior-authorization tool that processed them in one cloud region, passed extracted clinical facts to a second vendor's summarization service, stored the drafts in a third system, and surfaced them back to pharmacists across forty sites. Somewhere in that chain, a sub-processor the enterprise had never heard of was retaining inputs for an unspecified period in a jurisdiction no one had reviewed. None of this was malicious. It was the ordinary sprawl of an enterprise that had adopted AI tool by tool, each with its own contract and its own data flow, until the protected health information (PHI) of hundreds of thousands of patients moved through a system no single person could draw on a whiteboard. The Health Insurance Portability and Accountability Act (HIPAA), the federal law governing PHI protection, made the enterprise responsible for every hop in that chain, and the enterprise could not even name the hops. This lesson is about the foundation underneath every other piece of enterprise pharmacy AI: data governance and PHI protection at scale, the business associate agreements, the data flows, and the protections that have to hold across an entire enterprise, not just one tool at one site.
Why Data Governance Is the Foundation
Every other capability in an enterprise pharmacy AI program rests on data, and specifically on PHI, the patient's identity connected to their health data. The prior-authorization tool reads the chart. The verification-support layer reads the labs and the renal function. The counseling engine reads the medication profile. The clinical-decision support reads the patient's history. There is no pharmacy AI that does not touch patient data, which means there is no pharmacy AI program whose foundation is not data governance. If the foundation is unsound, everything built on it inherits the unsoundness: a brilliant PA workflow running on an ungoverned data flow is a breach with a fast turnaround, and the speed does not redeem the exposure. This is why data governance is treated not as one workstream among many but as the substrate the whole program stands on.
It helps to be precise about why an unsound data foundation contaminates everything above it, because the point is easy to nod along to and hard to feel. Consider the prior-authorization workflow that the whole program celebrates, the one that collapses turnaround from roughly twenty-five minutes to about five and gets patients on therapy faster. Every bit of that value is real, and every bit of it depends on the patient's chart flowing into the tool that assembles the justification. If that flow is ungoverned, if the tool has no business associate agreement, processes the chart in an unreviewed jurisdiction, or hands it to a sub-processor that retains it, then the celebrated workflow is not a faster PA, it is a faster way to expose PHI, and the enterprise has built its proudest capability on a breach. The value and the exposure travel together on the same data flow, which is why you cannot govern the value without governing the flow. Data governance is not a constraint bolted onto the program after the fact; it is the thing that determines whether the program's wins are wins at all or liabilities wearing the costume of wins.
The reason it gets dramatically harder at enterprise scale is multiplication. One pharmacist using one sanctioned tool is a manageable data-governance problem: one flow, one agreement, one place the PHI goes. An enterprise running many tools across retail, hospital, and specialty, each with its own vendor, its own data flow, and often its own chain of sub-processors, faces that problem multiplied by every tool and every site, and the multiplication is not additive but combinatorial, because the flows interact, hand data to each other, and create paths no one designed on purpose. The enterprise that adopted AI tool by tool without governing the data centrally ends up exactly where the specialty pharmacy did: legally responsible for a data estate it cannot map. As the covered entity under HIPAA, the entity legally obligated to protect PHI, the enterprise carries that responsibility whether or not it can see what it is responsible for, which is why seeing it, mapping it, and governing it is the foundational work.
There is no pharmacy AI that does not touch patient data, so there is no pharmacy AI program whose foundation is not data governance. A brilliant workflow on an ungoverned data flow is a breach with a fast turnaround.
The Business Associate Agreement at Scale
The business associate agreement (BAA) is the legal mechanism that lets PHI flow to a vendor safely: when a covered entity shares PHI with an outside vendor performing services on its behalf, that vendor must be a business associate bound by a BAA that legally obligates it to protect the PHI and restricts how it can be used. At the scale of a single tool, the BAA is a checkbox: is there one, yes or no. At enterprise scale, the BAA becomes a discipline, because the enterprise is not managing one agreement but a portfolio of them, and a portfolio has failure modes a single agreement does not. The first is coverage: every AI tool that touches PHI anywhere in the enterprise needs a BAA, and the enterprise has to know, as a continuously maintained fact rather than an occasional audit finding, that every such tool is covered. A single uncovered tool at a single site is an unbound flow of PHI, and the enterprise's responsibility does not shrink because the gap was small or local.
The second failure mode is the sub-processor chain, which is where enterprise BAAs most often fall short of the real data flow. A vendor the enterprise has a BAA with may pass PHI to its own sub-processors, other services that store, process, or transmit the data downstream, and the protection has to extend the whole length of that chain. A BAA that binds the primary vendor but is silent on its sub-processors leaves PHI protected for one hop and exposed for the rest, which is precisely the gap that hid the unknown retaining sub-processor in the opening story. Enterprise BAA discipline means the agreement accounts for the sub-processors, the enterprise knows who they are, and the protective obligations flow all the way down, because PHI that is protected only partway through its journey is not protected.
There is a subtler coverage gap that the sub-processor problem creates: the model-training question. A BAA may bind a vendor to protect PHI and still leave ambiguous whether the vendor or its sub-processors may use the patient data to train or improve their models. At the scale of one tool, this is a question to ask once; at enterprise scale, across a portfolio of vendors with evolving terms, it is a standing exposure, because a vendor that trains on the enterprise's PHI has effectively absorbed that patient data into a system the enterprise cannot retrieve it from or control. The enterprise BAA discipline has to make the training restriction explicit and verified for every vendor and every sub-processor in the chain, not assumed, because the default in much of the AI market is to retain and learn from inputs unless contractually forbidden. PHI that has been trained into a model is the most permanent form of exposure there is, and the only place to stop it is in the agreement, before the data ever flows.
The third failure mode is drift: BAAs and the data postures behind them are not static. A vendor changes its sub-processors, moves data to a new region, updates its model-training terms, or alters its retention policy, and a BAA that was adequate at signing can quietly become inadequate. At enterprise scale, with a portfolio of vendors each evolving, the discipline is ongoing rather than one-time: the enterprise reviews the agreements and the actual data postures on a cycle, catches the drift, and renegotiates or replaces what has fallen out of compliance. A BAA managed as a one-time signature is a snapshot of a protection that has since moved; a BAA managed as a living obligation is the protection the enterprise can actually stand behind. Treating vendor and contract terms as benchmarks to verify and re-verify, never as guarantees frozen at signing, is the posture that keeps the portfolio sound.
Mapping the Data Flows
You cannot govern what you cannot see, so the practical heart of enterprise data governance is the data-flow map: a clear, maintained picture of where PHI goes, through which tools and vendors and sub-processors, where it is processed, where it is stored, and where it comes back. The opening enterprise could not answer the most basic governance question, where does our patient data go, because it had never drawn that map, and an enterprise that cannot draw the map cannot protect the data, cannot answer an accreditor, and cannot respond to an incident, because all three require knowing the territory. Mapping the flows is the work that converts an invisible, sprawling data estate into a governable one.
A useful map answers a specific set of questions for every AI tool that touches PHI: what patient data enters the tool, where that data is processed and in what jurisdiction, where it is stored and for how long, whether it is used to train the vendor's models and under what restriction, which sub-processors handle it downstream, and where the output, which is often still PHI because output derived from PHI carries the same protection, ultimately goes. Notice that these are the same data-posture questions that govern a single tool, asked across the whole estate and connected into flows rather than left as isolated facts. The enterprise value of the map is that it makes the connections visible: it shows where two tools hand PHI to each other, where a flow crosses into a jurisdiction the enterprise did not intend, where data lands in a store that more people can reach than should, and where the output of one AI step becomes the unguarded input of the next.
The map is also what makes the rest of governance operational. It tells the BAA discipline exactly which vendors and sub-processors must be covered. It tells the security review where the sensitive stores and transfers actually are. It tells incident response, when something goes wrong, exactly what data was where and who could have touched it, which is the difference between a contained incident and a guess. And it tells the enterprise, when a new tool is proposed, whether the new flow it introduces is acceptable before the tool is adopted rather than after the PHI is already moving. An enterprise that maintains its data-flow map governs PHI proactively; an enterprise without one governs by reaction, discovering its flows only when an audit or a breach forces the discovery, which is the most expensive moment to learn where your patient data has been going.
It is worth dwelling on one consequence of the map that executives consistently underestimate: jurisdiction and retention are governance facts, not technical footnotes. Where PHI is processed and stored determines which legal regimes touch it, and an enterprise that lets a vendor quietly shift processing to a new region has changed its compliance exposure without deciding to. How long PHI is retained determines how large the standing pool of exposed data is at any moment, and a vendor that retains inputs indefinitely is accumulating the enterprise's liability on the enterprise's behalf. The map surfaces both, and surfacing them is what lets the enterprise govern them: decide where processing is acceptable, set retention expectations in the BAA, and verify that the actual flow matches the decision. Without the map, these facts are set by whichever vendor's default happened to apply, which is to say they are not governed by the enterprise at all, even though the enterprise is the one HIPAA holds responsible for them.
Protections That Hold Across the Enterprise
Knowing where the data goes is necessary but not sufficient; the protections themselves have to hold across every flow the map reveals, uniformly, because a protection that is strong in one division and weak in another protects only the patients in the strong division. The foundational protection is the sanctioned-tools rule applied enterprise-wide: only tools that have cleared the enterprise's evaluation, with a BAA, a reviewed data posture, and adequate security, may touch PHI anywhere in the organization, and the front-line rule that follows from it is the same in every building, that patient information goes only into sanctioned tools and never into unsanctioned public ones. At enterprise scale this rule needs enforcement, not just publication, because the most likely breach remains the convenient one: a well-meaning pharmacist or technician at any of forty sites pasting a chart into a public chatbot that has no agreement and may retain or train on the input. The enterprise protects against the multiplied version of that single-site risk by making the sanctioned path available everywhere and the unsanctioned path closed everywhere.
Beyond the tool boundary, enterprise protections include the disciplines that scale with data volume: access controls that ensure PHI in any store is reachable only by those who need it, so a flow that lands data in a shared system does not quietly widen who can read it; minimization, so tools receive only the patient data they actually need rather than the whole chart by default, shrinking the exposure at every hop; and attention to output, because AI output derived from PHI is often still PHI, so where the output flows matters as much as where the input came from, and an enterprise that guards its inputs while letting PHI-laden outputs land in unprotected places has secured the front door and left the back one open. The de-identification temptation deserves the same enterprise-level caution it gets at the front line: removing a name does not make patient data safe for an ungoverned tool, because HIPAA identifiers extend well beyond the name to dates, geography, record numbers, and rare-condition combinations, and true de-identification is a deliberate, verified process, not a quick edit that licenses a shortcut.
What ties these protections into a foundation rather than a list is that they are governed centrally and applied uniformly, with the data-flow map showing where each must hold and the BAA portfolio ensuring the legal protection travels with the data. This is also the evidence the enterprise brings to accreditation: the Utilization Review Accreditation Commission (URAC), which launched the first national Health Care AI Accreditation with separate developer and user tracks, expects a user organization to demonstrate exactly this, that it knows where PHI goes, binds its vendors, protects the flows, and can show the documentation. A coherent data-governance foundation is what lets the enterprise capture everything AI offers, the collapsed prior-authorization turnaround, the supported clinical review, the faster patient access, while keeping faith with every patient's trust that their information is protected at every hop, in every division, across the whole enterprise. Get the foundation right, and the rest of the program is buildable; get it wrong, and everything above it is a breach waiting to be discovered.
Key Takeaways
- There is no pharmacy AI that does not touch PHI, so data governance is the foundation the entire enterprise AI program rests on; a brilliant workflow running on an ungoverned data flow is a breach with a fast turnaround, and the speed does not redeem the exposure.
- Data governance gets combinatorially harder at scale: many tools across retail, hospital, and specialty, each with its own vendor and sub-processor chain, create interacting flows no one designed, leaving the enterprise legally responsible (as the covered entity) for a data estate it often cannot map.
- The business associate agreement (BAA) becomes a portfolio discipline at scale, with three failure modes: coverage gaps (every PHI-touching tool needs one), the sub-processor chain (protection must extend the full length of the data's journey), and drift (postures change, so BAAs must be reviewed on a cycle, not signed once).
- You cannot govern what you cannot see, so the practical heart of enterprise data governance is a maintained data-flow map showing, for every tool, what PHI enters, where it is processed and stored, whether it trains models, which sub-processors touch it, and where the output (often still PHI) goes.
- The data-flow map operationalizes everything else: it tells the BAA discipline which vendors to cover, the security review where the sensitive stores are, incident response exactly what was where, and the enterprise whether a proposed new tool's flow is acceptable before adoption rather than after.
- Protections must hold uniformly across every flow: the enterprise-wide sanctioned-tools rule (only evaluated, BAA-bound tools touch PHI), the front-line rule (never PHI into unsanctioned public tools), access controls, data minimization, and attention to PHI-derived output.
- De-identification by name-deletion is a trap at enterprise scale too: HIPAA identifiers extend far beyond the name to dates, geography, record numbers, and rare-condition combinations, so true de-identification is a deliberate, verified process, never a quick edit that licenses an ungoverned shortcut.
- A centrally governed, uniformly applied data foundation is the evidence URAC's user track expects and the precondition for capturing AI's value; get the foundation right and the program is buildable, get it wrong and everything above it is a breach waiting to be discovered.
Skill.re