The Vocabulary You Need: LLM, BAA, ZDR, PHI, Psychotherapy Notes, Part 2 Records
Jordan's compliance officer asks one question in the Friday meeting: "Does our scribe vendor's BAA cover zero data retention, and do their subprocessors touch Part 2 records?" Twelve words, four terms, and the room goes quiet, because the practice owner who cannot parse that sentence cannot evaluate a vendor contract, answer a malpractice questionnaire, or explain to a board investigator what protections were in place when the breach happened. Behavioral health AI runs on two vocabularies at once, the machine words (LLM, foundation model, RAG, embedding, prompt) and the law words (BAA, ZDR, PHI, psychotherapy notes, Part 2 records, minor's records), and the expensive mistakes happen exactly where clinicians know one vocabulary and bluff the other. This lesson builds your single-page glossary, the reference sheet every later lesson in this certification will point back to. By the end, the compliance officer's question will read like plain English, because it is.
Two Keyrings: The Machine Words and the Law Words
Here is the controlling analogy. Think of your AI vocabulary as two keyrings. The first keyring holds the machine keys: terms that unlock how the technology works, so that vendor claims and product behaviors stop being magic. The second holds the law keys: terms that unlock what the data is, who may touch it, and under what paper. A clinician carrying only the machine keys can use the tools fluently and walk straight into a privacy violation; a clinician carrying only the law keys can recite the rules and still be conned by a demo. You need both rings on one carabiner, and the carabiner is this lesson's glossary.
Why one carabiner? Because every real decision uses keys from both rings at once. "Can Carmen run a session recording through her $59/month scribe?" requires a machine key (what does the model do with the audio and transcript?) and three law keys (is this PHI? is there a BAA? does the vendor retain data?). The compliance officer's twelve-word question used BAA, ZDR, subprocessor, and Part 2 in a single breath because real compliance questions are always compound. Clinicians who learn the terms as two separate lists freeze on compound questions. Clinicians who learn them as one working vocabulary answer them at the speed of conversation. That is the bar for this lesson: not recognition, fluency.
The Machine Keys: LLM, Foundation Model, Prompt, Embedding, RAG
An LLM, large language model, is the text-prediction engine from the mental-model lesson: software trained on enormous text corpora that generates output one token at a time by predicting what plausibly comes next. Mentalyc, Upheal, Heidi, Twofold, and TherapyNotes' built-in AI all run on LLMs; so do ChatGPT and the companion chatbots. The term tells you the mechanism and therefore the failure modes: plausibility without verification, gaps filled with the field's averages.
A foundation model is the big, general-purpose LLM that vendors build on rather than build themselves: a handful of labs train these enormous base models, and behavioral health products typically layer their own tuning, prompts, and safeguards on top. This key matters for one reason above all: when a vendor says "our AI," they usually mean "our wrapper around someone else's foundation model," which means there is a subprocessor, another company whose servers may touch your client data. Jordan's blind spot in the insider story is exactly this: the EHR vendor's subprocessor list includes a model provider that does not sign a BAA at the tier Jordan is paying for. The foundation-model key teaches you the question: "Whose model is under the hood, and does your BAA chain cover them?"
A prompt is everything you place in front of the model for a given request: instructions, shorthand, transcript, format requirements. Prompts are how you constrain generation ("use only the facts provided"), and prompts containing client material are data disclosures, which is where the law keys take over. An embedding is a numerical representation of text that captures meaning-similarity: the machinery that lets software find that "can't get out of bed" and "no energy lately" are related content. Embeddings power search and theme-finding across sessions, and the key clinical fact is that embeddings derived from client text are still derived from client data: converting words to numbers is not de-identification. RAG, retrieval-augmented generation, is the architecture that fetches relevant stored material (prior notes, your templates, clinical references) and places it into the prompt automatically, so the model generates with that context on its desk. RAG is how a product "remembers" your client across sessions, and the key question it should trigger is now familiar: retrieval from where? A RAG feature means a stored repository of client material exists somewhere, with retention, security, and BAA implications. Every machine key, you will notice, ends by handing you to the law ring. That is not an accident. In behavioral health, every architecture question matures into a data-handling question.
The Law Keys, Part One: PHI, De-Identified, BAA, ZDR
PHI, protected health information, is individually identifiable health information held by a covered entity or its business associates: not just names, but the constellation of details (dates, locations, circumstances, diagnosis, employer, the custody dispute with the named ex) that can identify a person. The working clinical rule: assume session-derived text is PHI until proven otherwise. De-identified data is information from which identifiers have been removed to the point that re-identification is not reasonably possible, and the bar is higher than most clinicians think: stripping the name from a narrative about "a 56-year-old C-PTSD client in Oakland who works at the port and is divorcing a firefighter" has not de-identified anything. This distinction decides which tools are even candidates: PHI may enter only tools under appropriate agreements; truly de-identified content has more latitude. The expensive failure is calling something de-identified because it feels anonymous.
A BAA, Business Associate Agreement, is the HIPAA-required contract between a covered entity (you or your practice) and any vendor that creates, receives, maintains, or transmits PHI on your behalf. The BAA obligates the vendor to safeguard PHI, report breaches, and bind its own subcontractors. No BAA, no PHI: a free consumer chatbot tier without a BAA is not a place client material can go, full stop, which is the precise reason Maria closed the ChatGPT tab at 9:54 PM. And the BAA must chain: if your scribe vendor sends data to a foundation-model subprocessor, the protections must follow the data. "We have a BAA" is the beginning of diligence, not the end; "your BAA covers which entities, at which tier, for which data flows?" is the actual question.
ZDR, zero data retention, is the commitment that the vendor (and critically, its model providers) does not store your inputs and outputs beyond processing, and does not use them to train models. ZDR is a contract term, not a settings toggle you hope is on. The two law keys travel together but answer different questions: the BAA answers "is the vendor legally bound to protect this data?"; ZDR answers "does the data persist at all, and can it teach the model?" A vendor can hold a signed BAA and still retain transcripts for years; a vendor can promise deletion and have no BAA binding them to anything. Carmen's $59/month subscription decision, made alone at her kitchen table, was actually a BAA-plus-ZDR evaluation she did not know she was performing. The glossary exists so the next Carmen performs it knowingly.
Every machine term ends in a data question, and every data question ends in a paper question: what is this data, who touches it, and what contract binds them? Clinicians who can ask that chain out loud do not get surprised by their vendors.
The Law Keys, Part Two: Psychotherapy Notes, the Narrowest Term in the Room
Here is the term clinicians misuse most, and the misuse is dangerous in both directions. Under HIPAA, psychotherapy notes is a narrow term of art defined at 45 CFR §164.501: notes recorded by a mental health professional documenting or analyzing the contents of conversation during a counseling session, that are kept separate from the rest of the medical record. These are your process notes, your hypotheses, your countertransference observations, the material you would never want in the chart. HIPAA gives them extra protection: under 45 CFR §164.508(a)(2), most uses and disclosures of psychotherapy notes require the client's specific, separate authorization, beyond the general consent that covers ordinary treatment records.
What psychotherapy notes are not, and this is the part that surprises working clinicians: your progress notes are not psychotherapy notes. The regulation explicitly excludes medication prescription and monitoring, session start and stop times, modalities and frequencies of treatment, results of clinical tests, and any summary of diagnosis, functional status, the treatment plan, symptoms, prognosis, and progress to date. That excluded list is, almost exactly, the contents of a SOAP or BIRP note. So the progress note your AI scribe drafts is an ordinary treatment record, protected as PHI but not under the psychotherapy-notes carve-out, while the separate process journal you keep is psychotherapy notes and should, as a strong default, never enter an AI tool at all: the carve-out's protection logic (separate, rarely disclosed, specially authorized) is defeated the moment that material flows into a vendor pipeline. The two-direction danger: clinicians who think all their notes are "psychotherapy notes" overestimate their legal protection and resist legitimate disclosures; clinicians who do not know the category exists feed their most sensitive reflections into tools as if they were ordinary documentation. The key opens both doors: know which document you are holding before you decide where it can go.
The Law Keys, Part Three: Part 2 Records and Minor's Records
Part 2 records are substance use disorder treatment records governed by 42 CFR Part 2, the federal confidentiality regulation that applies to programs that hold themselves out as providing SUD diagnosis, treatment, or referral. Part 2 has always been stricter than HIPAA, born from the recognition that SUD treatment collapses if records leak into prosecutions, custody fights, and employment decisions. The 2024 final rule realigned Part 2 closer to HIPAA: it permits a single consent for all future uses and disclosures for treatment, payment, and operations, and it changed redisclosure rules, but the heightened baseline remains, and the practical AI consequence is blunt: a vendor adequate for general mental health PHI is not automatically adequate for Part 2 records. If your program or caseload includes SUD treatment, the vendor conversation has a second, harder round: how are Part 2 records segmented, what consent language covers AI processing, and how does the redisclosure prohibition follow the data, especially in the common scenario where the same client also has a release to a primary care doctor who is not a Part 2 program. This certification devotes a full lesson to the 2024 final rule later; the glossary entry's job is to make sure you never treat "we're HIPAA compliant" as an answer to a Part 2 question.
Minor's records carry the third special status, and here the rule is that there is no single rule: state minor-consent statutes determine when a minor can consent to their own mental health treatment, and with that consent comes a measure of control over the records and over who must authorize disclosures, including disclosures to an AI vendor's pipeline. California Family Code §6924 allows minors of a defined age and maturity to consent to outpatient mental health treatment; New York Mental Hygiene Law §33.21 governs minors' mental health records and consent in New York; Texas Family Code §32.004 sets Texas's narrower rules on a minor's consent to counseling. The statutes differ significantly, which is exactly the point: recording a minor's session into an AI scribe raises the question "whose consent satisfies my state's statute?", and any vendor or colleague who answers it without naming a statute is guessing. The glossary discipline, here as everywhere, is to refuse the generalized answer: jurisdiction first, statute second, workflow third.
Reading a Compound Sentence With Both Keyrings
Now return to the compliance officer's question and read it like the bilingual clinician you are becoming: "Does our scribe vendor's BAA cover zero data retention, and do their subprocessors touch Part 2 records?" Decoded: (1) Is there a HIPAA contract binding the vendor to protect PHI? (2) Does that contract include the commitment that inputs and outputs are not stored or used for training? (3) The vendor builds on a foundation model, so other companies sit in the data path; are they bound by the BAA chain? (4) Our caseload includes SUD treatment, whose records carry stricter-than-HIPAA federal protection; does the data flow keep those records segmented and consented? Four questions, each answerable, each pointing to a specific document or vendor representation. That is what vocabulary buys you: the difference between a vague unease about AI and a checklist you can run against a contract.
Practice the same decoding on vendor marketing, because marketing is written to be skimmed by clinicians holding neither keyring. "HIPAA-compliant AI": which entity is compliant, under what BAA, covering which subprocessors? Compliance is a property of a data flow, not a logo. "Your data is never used for training": is that ZDR in the contract, or a privacy-policy sentence that can change next quarter? "Bank-level encryption": encryption is necessary and says nothing about retention, training use, or the BAA chain; it is a machine answer to a law question. "Works great for SUD programs": show me the Part 2 analysis, the consent language, and the redisclosure handling. None of this requires hostility toward vendors; the good ones answer these questions crisply and in writing, and their crispness is itself a screening signal. The vocabulary turns you from an audience into a counterparty.
The Applied Problem: The Single-Page Glossary Sheet
Your artifact is the Single-Page Glossary Sheet, and unlike a definitions list you copy from this lesson, you will build it in a format designed for use under pressure: term, plain-English definition in your own words, and, the column that makes it yours, the question this term teaches me to ask. This sheet is referenced through every level of this certification, L1 to L5; it will sit beside your Sorting Card in vendor demos and beside your supervision agenda when an associate asks what ZDR means.
Build it as a two-section table. Section one, Machine Keys, five rows: LLM, foundation model, prompt, embedding, RAG. Section two, Law Keys, seven rows: PHI, de-identified, BAA, ZDR, psychotherapy notes (with the §164.501 separation requirement and the §164.508(a)(2) authorization rule noted), Part 2 records (with "2024 final rule: single consent, redisclosure changes, stricter than HIPAA" as the memory hook), and minor's records (with your own state's statute written in, CA Fam Code §6924, NY MHL §33.21, or TX Family Code §32.004 if you practice in those states, otherwise the citation you look up for your jurisdiction). For each row, the third column is the test: LLM teaches "what are this output's failure modes?"; foundation model teaches "whose model is under the hood and does the BAA chain cover them?"; RAG teaches "retrieval from where, retained how long?"; PHI teaches "is this identifiable, really?"; BAA teaches "which entities, which tier, which data flows?"; ZDR teaches "is non-retention in the contract?"; psychotherapy notes teaches "which document am I holding, chart note or process note?"; Part 2 teaches "is HIPAA-compliant even the right standard here?"; minor's records teaches "whose consent does my state's statute require?"
Then run the verification pass, which doubles as the first real use of the sheet. Take one actual document, your current scribe's terms of service, your EHR vendor's AI feature announcement, or, if you use no AI yet, any scribe vendor's public security page, and annotate it against your glossary: circle every glossary term that appears, underline every claim that uses a machine word to answer a law question (the "bank-level encryption" move), and write in the margin the unanswered question from your third column. If your annotation produces no margin questions, you have either found an unusually transparent vendor or you have not pressed hard enough; in this field, the second is more common.
"Done" looks like this: one page, twelve terms, every definition in your own words, your state's minor-consent statute filled in, and one real vendor document annotated as proof the sheet works as an instrument rather than a souvenir. Date it and file it as governance artifact number four, beside the Consultation-Group Paragraph, the Annotated Shorthand-to-Note Test, and the Three-Behavior Sorting Card. Chapter one is complete, and you now hold what most of your consultation group does not: the line, the mechanism, the behaviors, and the words.
Key Takeaways
- Behavioral health AI fluency requires two keyrings carried together: machine keys (LLM, foundation model, prompt, embedding, RAG) that decode how tools work, and law keys (PHI, de-identified, BAA, ZDR, psychotherapy notes, Part 2 records, minor's records) that decode what the data is and what paper must bind whoever touches it. Real compliance questions are compound and use both rings in one sentence.
- Every machine term matures into a data question: foundation models mean subprocessors in your data path, embeddings derived from client text are still client data, and RAG means a stored repository of client material exists somewhere with retention and BAA implications. "Whose model is under the hood, and does the BAA chain cover them?" is the question Jordan's practice never asked.
- PHI is identifiable health information, and the working rule is to assume session-derived text is PHI until proven otherwise; de-identification is a high bar, and stripping a name from a re-identifiable narrative does not clear it. No BAA, no PHI: a Business Associate Agreement is the HIPAA-required contract for any vendor touching PHI, and it must chain to subcontractors.
- ZDR (zero data retention) and the BAA answer different questions: the BAA asks whether the vendor is legally bound to protect the data; ZDR asks whether the data persists at all and whether it can train models. A vendor can have either without the other; diligence requires both answers in writing.
- Psychotherapy notes is the narrowest term in the room: under 45 CFR §164.501, only process notes kept separate from the medical record qualify, and they get heightened protection (separate authorization under §164.508(a)(2)). Your SOAP/BIRP progress notes are ordinary treatment records; your process notes should, as a strong default, never enter an AI pipeline at all.
- Part 2 records (SUD treatment, 42 CFR Part 2) remain stricter than HIPAA even after the 2024 final rule's alignment (single consent for TPO, redisclosure changes), so "HIPAA-compliant" is not an answer to a Part 2 question. Minor's records have no general rule: consent and disclosure control follow your state's statute, CA Fam Code §6924, NY MHL §33.21, TX Family Code §32.004, and any answer that names no statute is a guess.
- The artifact is the Single-Page Glossary Sheet: twelve terms, definitions in your own words, the question each term teaches you to ask, your state's minor-consent statute filled in, and one real vendor document annotated as proof of use. It is governance artifact number four and the reference sheet for every later level of this certification.
Skill.re