The BAA, the DPA, and the Subprocessor Map
Jordan finally read the contract. Not the summary, not the sales deck's "fully HIPAA compliant" slide: the actual Business Associate Agreement, late on a Tuesday, the compliance officer's annotations in the margin. The BAA was real. It bound the vendor. And on page nine, in the subprocessor exhibit, sat the sentence hiding in plain sight: the vendor passes data to a model provider that does not sign a BAA at the tier Jordan's vendor purchases. The agreement covered the vendor and stopped there, while the transcripts kept going. This lesson teaches you to read the three documents that govern where your clients' words actually travel (the BAA, the Data Processing Agreement, the subprocessor list) and to draw from them the artifact most diligence never produces: the subprocessor map. You will read three BAA categories side by side (per-clinician scribes like Mentalyc and Upheal, an enterprise platform like Eleos Health), learn what each document does and does not promise, and finish with a Subprocessor Map Worksheet whose acceptance criteria your practice signs before any vendor does.
Chain of Custody: The Organizing Analogy
Hold one analogy through this lesson: PHI moving through an AI vendor's stack is evidence moving through a chain of custody. What matters in a chain of custody is not whether the first officer was trustworthy; it is that every hand the evidence passes through is identified, authorized, and accountable, and that the chain is documented end to end. One undocumented handoff compromises the whole chain, however impeccable the other links. Your clients' session transcripts work the same way: the transcript leaves the session, enters the vendor's application, transits to a model provider, lands in cloud storage, may touch an analytics pipeline and a support engineer's screen. Each handoff is a link. The BAA binds one link. The question this lesson answers is who binds the rest.
Clinicians are trained in a version of this already. When you release records to a primary care doctor, you think about what that doctor may do with them; when 42 CFR Part 2 applies, the redisclosure prohibition makes the chain explicit, because Part 2 records carry their restrictions to the recipient. The instinct that makes you careful about the second disclosure is the instinct vendor diligence requires, scaled to a stack you cannot see. The vendor is the first recipient, the model provider the second, the cloud platform the third. Most practices vet the first link and assume the rest, and the verdict deserves repeating: the subprocessor question is the one most diligence misses.
Why is it missed? Because the documents are split. The promises a practice cares about live in three instruments with three jobs, and a practice that reads only the BAA, the one document with HIPAA's name on it, reads a third of the chain. We take the documents in order, then assemble the map.
Document One: What the BAA Does, and Where It Stops
The Business Associate Agreement is the HIPAA instrument. When your practice (a covered entity) lets a vendor create, receive, maintain, or transmit PHI on its behalf, the vendor is a business associate and HIPAA requires the BAA. Read one and you find a standard skeleton: permitted uses and disclosures (PHI used only to provide the service); safeguards; breach notification within a stated window, so your own clocks can start; access and amendment support for client rights requests; return or destruction at termination; and, the clause this lesson turns on, the subcontractor flow-down: the business associate must obtain equivalent written assurances from its subcontractors that handle PHI. In chain-of-custody terms, the flow-down clause is the BAA promising that the next link will also be bound.
Here is where careful reading pays. A flow-down clause is a promise about paper, and three failure patterns hide inside it. Pattern one: the unbound subprocessor, Jordan's case. The BAA contains a perfectly standard flow-down clause, and the vendor's actual model provider does not sign a BAA at the tier the vendor purchases, so the flow-down promise is either broken or satisfied by something weaker than the clause implies. Pattern two: the de-identification escape hatch. Some BAAs permit the vendor to de-identify PHI and use the result freely, moving the question from HIPAA to de-identification quality: a transcript "de-identified" by stripping eighteen identifiers may still tell a recognizable story. Pattern three: the scope carve-out, where the BAA covers the core service but the definitions section quietly excludes the analytics module, usage data, or "service improvement" processing. The BAA you want states which services and flows it covers, names the breach window in days, and contains a flow-down clause you can verify rather than merely admire.
And note what the BAA does not do, even at its best. It does not name the subprocessors, say what country your data sits in, how long the model provider retains it, or whether anyone trains on it. It binds one party and gestures at the rest. For the rest, you need the second document.
Document Two: The DPA, the Document That Names Names
The Data Processing Agreement comes from a different legal tradition: it is the instrument privacy regimes like the GDPR built to govern processors, and US vendors adopted it because their customers are global. It matters because it holds the operational specifics the BAA omits: categories of data processed, purposes of processing, concrete security measures, international transfer mechanisms, and, critically, the subprocessor regime: an exhibit or linked page listing subprocessors by name, a promise to notify customers before adding or replacing one, and, in stronger DPAs, a right to object.
Read the subprocessor clause with the care you gave the flow-down. The strong version: a current, dated list of named subprocessors with functions and locations; written notice (30 days is common) before any change; and an objection right with a real remedy, typically termination without penalty if the objection cannot be resolved. The weak version: a link to a webpage "as updated from time to time," no notice, no objection right, which translates to "the chain of custody can change without your knowledge." The difference is not academic: a vendor that swaps model providers has changed who holds your clients' trauma narratives, and without notice rights you learn about it never. Read together, the two documents divide the labor cleanly: the BAA establishes the HIPAA obligations of the first link; the DPA tells you who the other links are and how the list can change. Most practices read the first and skim the second; the second is where Jordan's problem was visible.
One practical note on reading three vendor stacks side by side. Take a per-clinician scribe in the Mentalyc category, another in the Upheal category, and an enterprise platform in the Eleos Health category, and lay their BAAs and DPAs on one table. You will see the same skeletons with different muscles: per-clinician products offer standardized, non-negotiable paper with a published subprocessor page, while enterprise paper is negotiated, so the flow-down language, notice windows, and objection rights are places your counsel can push. Three side by side teaches you the market's range, the only way to know whether the clause in front of you is standard, generous, or quietly deficient.
The BAA binds the first link; the subprocessor map is the whole chain. Most diligence reads the first document and assumes the rest, and the assumption is where the breach lives.
Document Three: The Subprocessor List, Where Tier Is Everything
The subprocessor list is the shortest of the three documents and decides the most. It names the third parties the vendor passes data to, and in an AI scribe's stack the names cluster into layers: the model providers (OpenAI, Anthropic), the cloud platforms (AWS, Azure, GCP, which may host the vendor's infrastructure, the model itself, or both, as when a vendor reaches Anthropic through AWS Bedrock or OpenAI through Azure OpenAI Service), and the supporting cast: storage, logging, analytics, support tooling, email. Every name is a hand in the chain of custody, and for each you need three facts: what data reaches it, what agreement binds it, and at what tier.
Tier is the word this chapter keeps returning to, because it is the word that undid Jordan. Major model and cloud providers sell at multiple commercial tiers, and the legal commitments differ by tier: whether the provider signs a BAA, whether zero-data-retention applies, whether inputs can train models, what human review is permitted. The same company name on a subprocessor list can mean a fully bound, ZDR, no-training enterprise arrangement or a default API arrangement with none of those properties, and the list does not say which unless you make it say. So the question is never "is OpenAI or Anthropic on the list," as if the name were the verdict; it is "at what tier does this vendor purchase, and does that tier carry a BAA, ZDR, and a training prohibition." A vendor on a healthcare-grade tier of a major provider can be a stronger chain than one running its own models with weak controls; a vendor on a default tier of the most reputable provider on earth is an unbound link wearing a famous name.
The cloud layer has its own tier question. AWS, Azure, and GCP all sign BAAs, but a cloud BAA covers designated services configured in designated ways; a vendor that says "we are covered by our AWS BAA" must be using covered services within that scope, and routing model traffic through a cloud marketplace (Bedrock, Azure OpenAI) changes which entity's tier terms govern the model hop. None of this requires you to become a cloud architect. It requires you to ask, for each named subprocessor, the three facts, in writing, and to refuse to call the chain complete until every link has them.
Building the Map: From Lists to a Picture
Now assemble the artifact. A subprocessor map is a one-page table that follows the data, hop by hop, from the therapy room to every system that touches it, with the binding agreement and tier recorded at each hop. Build it in rows, one per link. Link 1, the vendor application (binding: your BAA, scope and breach window noted). Link 2, the model provider(s) (the vendor's agreement with the provider, at a named tier, with BAA yes/no, ZDR yes/no, training prohibition yes/no). Link 3, the cloud platform(s) (whose BAA, covering which services, and whether the model hop routes through the cloud marketplace). Links 4 and beyond, the supporting cast, recording what data reaches each: a logging vendor receiving de-identified telemetry is a different risk than a support tool showing full transcripts to agents.
Filling the map is an exercise in refusing to accept blanks. The published subprocessor page gives the names; the DPA gives the change-notice regime; the RFI responses (questions two and seven) give tiers and retention; and where a cell stays empty, you ask, in writing, until it is filled or until the refusal itself becomes your answer. Expect two frictions. First, the vendor who answers tier questions with the provider's general marketing ("OpenAI offers zero data retention") rather than its own purchasing ("we are on the tier that includes it, see attached"); the cells are about this vendor's actual agreements, and only their documentation fills them. Second, the vendor who treats the supporting cast as beneath diligence; it is not, because breach history across industries is full of incidents that came through the logging tool and the support desk, not the headline model provider.
When the map is full, read it the way you read a genogram: the picture shows what the individual facts did not. Five links bound and one blank is a chain with a hole, and the hole, not the average, is the risk. Every link bound but a "no" in the training-prohibition column at the model layer is a consent problem in a compliance costume: the BAA chain can be perfect while clients' narratives flow into a training corpus the consent addendum never described. And a map the vendor cannot help you complete is itself the finding: a company that does not know, or will not say, where your clients' words go has answered the only question that mattered.
Acceptance Criteria: What Your Practice Can Live With
A map without a decision rule is trivia, so the worksheet ends with acceptance criteria: the practice's written standards for what a completed map must show before any contract is signed. Write them as testable statements, before you map any vendor, for the same pre-commitment reason the RFI had gates and the pilot had signed thresholds. A defensible baseline set: every link receiving identifiable PHI is bound by a BAA or equivalent written obligation at the tier actually purchased, verified by documentation, not marketing. The model-provider link carries zero-data-retention or an explicitly disclosed retention the practice has affirmatively accepted and reflected in its consent language. No link may use identifiable client data for model training; "de-identified" training is acceptable only if the methodology addressed narrative re-identification and the consent addendum discloses it. The DPA provides named subprocessors, change notice of at least 30 days, and an objection right with a termination remedy. And supporting-cast links that receive transcripts or PHI (support tooling above all) meet the headline standard.
Then add the two criteria practices forget. Location: where each link processes and stores data, because your consent language, breach obligations, and some state laws care. And change governance: the map is dated, the acceptance applies to this map, and a subprocessor change reopens the acceptance review; the DPA's notice clause keeps the map current, and that one sentence costs nothing now and preserves your leverage forever.
Run Jordan's vendor through these criteria and the outcome is no longer a Tuesday-night discovery; it is a row failing a written test. Model provider link: BAA at purchased tier, no. Criterion one fails, the map is rejected, and the conversation changes from "are you HIPAA compliant" (which invites the technically true "yes") to "your model-provider link is unbound at your tier; bind it, upgrade it, or we are done," a sentence a vendor can act on. That is the map's quiet power: it turns diffuse compliance anxiety into specific, negotiable, named deficiencies.
The Side-by-Side Reading: Three Vendor Categories
Close the loop with the exercise this chapter prescribes: three BAAs side by side, three subprocessor regimes, one afternoon. Take the per-clinician scribes (the Mentalyc and Upheal categories) and the enterprise platform (the Eleos Health category) and read with four questions in hand. Where does each BAA's scope start and stop, and which services or flows are excluded? What is each breach-notification window, in days? What does each flow-down clause require of subcontractors, and can you verify it against the published subprocessor list? What does each DPA give you on change notice and objection rights?
The comparative reading teaches three durable lessons. First, standardization cuts both ways: per-clinician take-it-or-leave-it paper is consistent and published, making mapping fast, but nothing in it is negotiable, so your only lever is selection; enterprise negotiated paper is slower to map but movable, so the deficiencies your map surfaces become redlines rather than dealbreakers. Second, the documents disagree more often than you expect: a flow-down promising equivalent obligations, beside a subprocessor page listing a provider that does not offer such obligations at any plausible tier the vendor could be buying, is a discrepancy you can now see, name, and ask about. Third, and most important: after three side-by-side readings, the documents stop being intimidating. The skeleton repeats. The clauses have jobs. The blanks have names. An owner who has built one subprocessor map can read any vendor's paper in an hour, and that capability, not any single vendor decision, is what this lesson is for.
The Applied Problem: Complete the Subprocessor Map Worksheet
Your artifact is the Subprocessor Map Worksheet: a one-page map template plus a signed acceptance-criteria page, ready for any vendor. Four steps.
Step one, build the map table with one row per link and six columns: Link (name and function), Data Received (full transcripts, PHI fields, de-identified telemetry), Binding Agreement (BAA, DPA flow-down, other, none), Tier (the named commercial tier, with BAA yes/no, ZDR yes/no, training prohibition yes/no at that tier), Location (processing and storage), and Evidence (the document proving the cell, by name and date). Pre-populate the rows for the standard AI scribe stack: vendor application, model provider(s), cloud platform(s), storage, logging/analytics, support tooling, and anything else the published list names.
Step two, write the acceptance criteria as numbered, testable statements: every PHI-receiving link bound at the purchased tier with documentary evidence; model layer ZDR or affirmatively accepted disclosed retention; no identifiable-data training anywhere, de-identified training only with narrative-aware methodology and consent disclosure; DPA with named subprocessors, 30-day change notice, and objection-with-termination remedy; transcript-touching supporting links held to the headline standard; locations recorded; and the change-governance clause: this acceptance applies to the map as dated, and any subprocessor change reopens review. The owner and compliance officer sign and date the criteria before any vendor's map is attempted.
Step three, run the worksheet against one real vendor (the one your RFI scored highest is the natural candidate). Pull the published subprocessor list, the BAA, and the DPA; fill every cell you can from documents; and draft the written follow-up for every blank: "For [provider] at the tier you purchase: does that tier include a BAA, zero-data-retention, and a training prohibition? Attach the agreement excerpt or attestation." Log which cells the vendor fills, which it evades, and which it refuses; refusals go to the front of the next conversation.
Step four, the verification pass. Read the completed map against the signed criteria, criterion by criterion, marking pass or fail; any fail produces a named deficiency sentence the vendor can act on ("your model-provider link is unbound at your tier"). Done looks like this: a stranger reading the worksheet can trace a transcript from therapy room through every hand that touches it, say which links are bound and at what tier, state in one sentence what would make your practice walk away, and would have caught Jordan's page-nine problem in ten minutes instead of late on a Tuesday.
Key Takeaways
- PHI in an AI vendor's stack is evidence in a chain of custody: every hand it passes through (vendor application, model provider, cloud platform, storage, logging, support tooling) must be identified, bound, and documented, and one undocumented handoff compromises the chain no matter how strong the other links are.
- The BAA binds the first link: permitted uses, safeguards, breach notification with a stated window, termination return-or-destruction, and the subcontractor flow-down clause. Its three hiding places are the unbound subprocessor (Jordan's case), the de-identification escape hatch, and the scope carve-out excluding analytics or service-improvement processing.
- The DPA is where names live: subprocessors listed with functions and locations, change notice (30 days is the standard worth demanding), and an objection right with a termination remedy. A weak DPA ("as updated from time to time," no notice) means the chain of custody can change without your knowledge.
- Tier is everything at the model and cloud layers: the same provider name (OpenAI, Anthropic, AWS, Azure, GCP) can mean a bound, zero-data-retention, no-training arrangement or a default arrangement with none of those properties. The question is never whether a famous name is on the list; it is what tier the vendor purchases and what that tier carries, proven by the vendor's own paper.
- The map is a one-page, row-per-link table recording data received, binding agreement, tier (BAA/ZDR/training), location, and evidence, read like a genogram: the hole, not the average, is the risk, and a vendor that cannot help you complete it has answered the question.
- Acceptance criteria are written and signed before mapping any vendor: every PHI link bound at the purchased tier, model-layer ZDR or accepted disclosed retention, no identifiable-data training, DPA notice and objection rights, supporting tools held to the headline standard, locations recorded, and a change-governance clause reopening review when the list changes.
- Reading three BAAs side by side (the Mentalyc, Upheal, and Eleos Health categories) teaches the market's range: standardized per-clinician paper is fast to map but non-negotiable, enterprise paper is slower but movable, and discrepancies between a flow-down promise and a published subprocessor list become findable. The subprocessor question is the one most diligence misses; the map is how you stop missing it.
Skill.re