AI for Mental & Behavioral Health Clinicians
Capable · M10 · lesson 10 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Few-Shot Prompting with De-Identified Examples
📖
now learning

Few-Shot Prompting with De-Identified Examples

15 min

Maria's AI drafts are accurate now that her system prompt is installed, but they do not sound like her. Her own notes say "Client identified a connection between conflict with her adult daughter and the resurgence of intrusive memories"; the AI writes "Patient verbalized insight regarding interpersonal stressors." Both are clinically fine. Only one will pass the eyeball test of a payer reviewer, a subpoena, or her own continuity of care when she rereads the chart in eight months, because a record that changes voice mid-chart reads like a record written by two different people, and in a sense it was. The fix is few-shot prompting: giving the model three of your own prior notes, fully de-identified, as worked examples of the target voice and structure. This lesson teaches you the few shot prompting clinical workflow end to end: why examples outperform descriptions, the Safe Harbor de-identification pass that must happen before any prior note leaves your EHR, the three-example block itself with a complete annotated sample, and the contract verification (BAA, DPA, data-handling addendum) that confirms the model does not retain or train on what you show it. By the end you will have a written de-identification and few-shot protocol you can run in twenty minutes and defend in an audit.

Why Three Examples Beat Three Paragraphs of Description

In the last lesson you wrote standing orders: the system prompt that fixes role, format, and guardrails. Standing orders tell a new scribe the rules. They do not teach the scribe your voice. For that, every supervisor in history has done the same thing: "Here, read three of my notes before you draft your first one." That is few-shot prompting. Instead of describing your style in adjectives ("concise, behavioral, strengths-based"), you show the model two to four complete examples of input and desired output, and the model infers the pattern: your sentence length, your verb choices, how you phrase progress toward goals, how much Subjective you write versus Assessment, whether you write "Client" or "Ct," how you document homework.

Language models are pattern-completion engines, which is why demonstration outperforms description so reliably. "Write concisely" is an instruction the model interprets against its training average; a 140-word Subjective section is a specification it can measure itself against. Style adjectives are ambiguous: your "concise" and the model's "concise" can differ by a hundred words a section. Examples collapse that ambiguity. In practical terms, a system prompt alone gets you a correct note in a generic voice; a system prompt plus three of your notes gets you a correct note in your voice, and the difference is visible in the first draft.

The number three is not arbitrary. One example teaches a template but not a range; the model over-fits to that single note's quirks, including its accidents. Two examples begin to show variation. Three examples spanning your real range, a routine maintenance session, a more clinically active session, and a session that includes a documented (clinician-assessed) risk discussion, teach the pattern and its boundaries without bloating the prompt. Past four or five, you pay in context length and the marginal style gain flattens. Three well-chosen, fully de-identified notes is the working standard this chapter uses, and it is what the L2 capstone's four progress notes will be drafted against.

The Non-Negotiable Gate: No Prior Note Leaves the EHR Identified

Here is where the convenience of few-shot prompting collides with the law, and the law wins. Your prior notes are protected health information. Pasting a real progress note into any AI tool, even a BAA-covered one, when the purpose is "teaching the model my style" rather than treatment, payment, or operations for that client, is a use you should not be making with identified data when a de-identified version accomplishes the identical goal. The model does not need to know the client's name to learn your voice. It does not need the real city, the real employer, the real daughter's age. Every identifier you leave in is risk you took for zero benefit, and "I left her name in because I was tired" is not a sentence anyone wants to say to an OCR investigator or their board.

So the rule is absolute: every example note is de-identified before it leaves the EHR, every time, no exceptions for "safe" tools. The standard to work from is the HIPAA Safe Harbor method (45 CFR 164.514(b)(2)): remove all eighteen identifier categories, names, geographic subdivisions smaller than a state, all elements of dates except year (and all ages over 89), phone numbers, email addresses, record numbers, and the rest, plus the catch-all that matters most in psychotherapy notes: any other unique identifying characteristic. In a behavioral health note, the dangerous identifiers are rarely the name. They are the story details: "client's husband is the fire chief in a small Sierra foothills town," "client is one of two pediatric oncologists at the regional hospital," "client's son was the subject of a local news story." A note can be name-free and still identify a human in two sentences. Your de-identification pass reads for re-identification risk, not just for the eighteen categories.

And there is a Part 2 tripwire: if the note comes from substance use disorder treatment covered by 42 CFR Part 2, the disclosure rules are stricter than HIPAA's, and the cleanest practice is to not use Part 2 records as few-shot examples at all. Choose style examples from non-Part 2 charts; your voice is the same.

The Two-Pass De-Identification Protocol

Do this by hand, outside any AI tool, in a plain text editor. Pass one is mechanical: work down the Safe Harbor list. Replace the name with "Client." Strip dates to a relative form ("session 6 of the current episode" instead of the date). Generalize geography to nothing smaller than the state, and usually drop it entirely. Replace ages with a band ("client in her 50s"). Remove provider names, clinic names, payer names, record numbers, medication pharmacy details. Replace family members' names with roles ("adult daughter," "spouse"). This pass takes about five minutes per note once you have done it twice.

Pass two is narrative, and it is the clinical-judgment pass no checklist replaces. Read the note as a stranger from the client's town and ask: could anyone who knows this person recognize them? Unusual occupation: generalize it ("works in a public-safety leadership role" becomes "works in a demanding supervisory role," or drop it if the style example does not need it). Distinctive events: abstract them ("the lawsuit against her former employer" becomes "an ongoing legal stressor"). Rare diagnoses or famous-in-a-small-pond details: swap or remove. The test is not "did I remove the eighteen categories"; the test is "would this client, reading this note on a projector at a conference, feel exposed?" If yes, keep editing. When in doubt, you may also lightly fictionalize non-clinical surface details (a different hobby, a different family configuration) since the example's job is to carry your voice and structure, not this client's facts; just never alter the clinical reasoning style, because that is exactly what you are trying to teach.

Then the verification step that makes this a protocol instead of a habit: a second read, ideally the next day or by a colleague in a group practice, against a written checklist (the eighteen categories plus "re-identification narrative risk: yes/no"). Date it, initial it, and keep the de-identified example set in a dedicated file. That file, three notes, checklist initialed, dated, is itself an audit artifact: proof that what entered the AI tool was de-identified by a documented process, not by vibes at 10 PM.

The model does not need to know who the client is to learn how you write. Every identifier you leave in a few-shot example is risk you took for zero benefit.

The Three-Example Few-Shot Block, Annotated

Here is the structure of the block you append after your system prompt. It has a framing line, three example pairs, and a closing instruction. Framing line: "Below are three de-identified examples of my prior notes. Match their voice, structure, section length, and level of clinical specificity. Do not reuse their clinical content; they are style examples only." That last sentence matters more than it looks: without it, models sometimes import facts from the examples into the new note, and a borrowed PHQ-9 score or a borrowed family detail appearing in a different client's chart is a serious documentation error.

Then the examples, each as an input/output pair so the model learns the transformation, not just the destination. EXAMPLE 1 (routine session). Input shorthand: "53 min, 90837. Anxiety re: job change, used cognitive restructuring on catastrophizing thought re: failing in new role, client generated alternative thought, mood improved in session, HW: thought record. PHQ-9 11, down from 14." Output note: "S: Client reported anticipatory anxiety related to an upcoming job transition, describing recurrent catastrophizing thoughts about failing in the new role. PHQ-9 today: 11, decreased from 14 at last administration. O: Client was alert, engaged, affect mildly anxious at session start with visible easing by mid-session. A: Client responded well to cognitive restructuring targeting catastrophic predictions; she independently generated a balanced alternative thought, consistent with treatment plan goal 2 (reduce anxiety-driven avoidance). Symptom trajectory supports continued weekly treatment at current intensity. P: Continue weekly 90837. Homework: daily thought record. Session 53 minutes, 4:02-4:55 PM." The annotation you should notice: every output fact traces to the input; the voice is specific ("she independently generated") rather than generic ("client demonstrated insight").

EXAMPLE 2 (clinically active session) shows a denser Assessment: grief work, an in-session abreaction, your way of writing affect and containment, still fully sourced from its input line. EXAMPLE 3 (risk-documented session) is the most important teaching example: its input includes "passive SI reported, no plan/intent, clinician assessed risk as low, safety plan reviewed and updated, means counseling re: medication storage," and its output documents exactly that, with the risk assessment attributed to the clinician ("Clinician assessed current risk as low based on absence of plan, intent, or preparatory behavior; safety plan reviewed and updated in session"). Including a risk example teaches the model what risk documentation looks like when the clinician has already made the call, which reinforces, rather than undermines, the system prompt's rule that the model never makes it.

Closing instruction: "New session input follows. Produce one note in the same format and voice. All facts must come from the new input only; flag gaps with [MISSING]." Total block: usually 700-1,000 words. Save it with your system prompt as a single reusable preamble; in a tool with project files or custom instructions, install it once.

Before You Paste Anything: Verify the BAA, the DPA, and the Data-Handling Addendum

De-identification protects the client if the tool misbehaves. The contract stack determines whether the tool is allowed to misbehave. Even with de-identified examples, you are establishing a workflow that will, in daily use, carry real session shorthand about real clients, so the playbook's instruction is to verify three documents before this workflow goes live. First, the BAA (Business Associate Agreement): does one exist for your tier of the tool, and does it cover this use? No BAA, no client-related use, full stop; that decision was made in earlier chapters and nothing here reopens it.

Second, the DPA (Data Processing Agreement or addendum): the document that says what the vendor may do with the data you submit. You are looking for three specific commitments in writing: inputs and outputs are not used to train or improve the models; retention is limited and defined (zero data retention, or a stated deletion window); and subprocessors are listed, because "we don't train on your data, but our model provider does" is a real failure pattern hiding one contract layer down. Jordan's practice in Sacramento learned this the expensive way: the EHR vendor said yes to AI, and the vendor's subprocessor list included a model provider that does not sign a BAA at the tier Jordan was paying for.

Third, the data-handling or AI addendum many vendors now publish: where training-use and retention commitments actually live, often distinct from the marketing page. Read the document, not the FAQ. The specific question for few-shot work: "Are prompt contents, including examples I provide, retained or used for model training at my subscription tier?" Get the answer in the contract or in writing from the vendor, screenshot or PDF it, date it, and file it next to your de-identified example set. Settings matter too: in general-purpose tools, confirm training/data-sharing toggles are off at the workspace level, and note that consumer free tiers typically retain and may train, which is one more reason they were excluded back when the BAA question was settled. Re-verify annually and at every plan change, because tiers and terms move.

Failure Modes: Content Bleed, Style Lock, and the Stale Example Set

Few-shot prompting has its own failure modes, and naming them is how you catch them. Content bleed is the model importing facts from your examples into a new note: Example 1's PHQ-9 of 11 surfacing in a different client's draft, or Example 3's safety-plan language appearing in a session where no safety planning occurred. The defenses are the "style examples only" framing line, the closing "facts from new input only" instruction, and your line-by-line read before signing, looking specifically for anything that feels familiar from the examples. During your first week on this workflow, keep the three examples open beside the draft and check.

Style lock is subtler: the model matches your examples so well that every note starts sounding identical, the same Assessment skeleton, the same transition phrases, eight notes a day. Identical-sounding notes are their own audit flag; payer reviewers and board investigators both know what cloned documentation looks like. The defense is choosing three examples that span your real range (routine, active, risk-documented) rather than three near-twins, and editing each draft enough that your actual session is what varies the note, which it will if your input shorthand is specific.

Finally, the stale example set: your documentation evolves, your payer mix changes, you add a modality, and the examples still teach 2024-you. Refresh the set when your format or practice changes materially, and re-run de-identification from scratch on any new example; never "lightly edit" an identified note inside the AI tool, because the identified version just left your EHR, which is the exact event this whole protocol exists to prevent. The example file gets a version number and a date, same discipline as the system prompt it rides with.

Supervision, Group Practices, and Whose Voice the Model Learns

Two governance wrinkles before you build. For pre-licensed clinicians like Carmen, the few-shot example set is a supervision topic, not a private hack. Whose notes go in the block matters: if Carmen seeds the model with her supervisor's notes, the AI drafts in a voice that overstates her independent judgment; if she seeds it with her own early notes, she is teaching the model her current weaknesses. The defensible version is examples chosen with the supervisor, from Carmen's own best supervisor-approved notes, with the de-identification checklist co-signed, and the whole protocol named in the supervision agreement's AI clause. That single paragraph protects both licenses and turns a risk into a teaching exercise.

For group practices like Jordan's, the question is standardization versus voice. A practice can ship one shared example set that encodes the practice's documentation standard (useful for audit consistency and onboarding), or let each clinician maintain a personal set under a shared de-identification protocol. The hybrid usually wins: a practice-standard system prompt, a shared de-identification checklist that compliance signs off on, and personal example sets built under it. What a practice cannot defensibly allow is what Jordan currently has: twelve clinicians pasting whatever into whatever, with no protocol, no checklist, and no contract verification on file. The artifact you build below is, at group scale, the policy fix for exactly that.

The Applied Problem: Your De-Identification and Few-Shot Protocol

Your artifact is the De-Identification + Few-Shot Protocol: a one-page written procedure plus the assembled example block, ready to run and ready to show. It is the second component of your prompt stack (system prompt, then this), and the L2 capstone, the complete documentation kit for one de-identified client, depends on it twice: the capstone's four progress notes are drafted in your voice because of this block, and the capstone client's de-identification uses this same two-pass protocol.

Step one: select three prior notes spanning your range, one routine, one clinically active, one with clinician-assessed risk documentation, from non-Part 2 charts. Step two: run the two-pass de-identification in a plain text editor outside any AI tool: mechanical pass against the Safe Harbor eighteen plus the catch-all, then the narrative pass against the conference-projector test. Step three: second read against the written checklist, dated and initialed (by you the next day, or a colleague). Step four: verify the contract stack and file the evidence: BAA covering your tier, DPA or AI addendum stating no training on inputs and defined retention, subprocessor list reviewed, training toggles off, screenshots dated. Step five: assemble the block, framing line, three input/output pairs, closing instruction, and install it with your system prompt. Step six: test once with a fictional session shorthand and read the draft against the examples, hunting specifically for content bleed.

The one-page protocol document records all six steps as your standing procedure, with the checklist embedded and a refresh rule (re-verify contracts annually and at plan changes; rebuild examples when your documentation materially changes). "Done" looks like: a dated protocol page, a de-identified example file with an initialed checklist, contract evidence in the same folder, and a first AI draft that a colleague who knows your charts would attribute to you. When your supervisor, compliance officer, or an auditor asks "you fed your notes to an AI?", your answer is a folder, not an apology.

Key Takeaways

  • Few-shot prompting teaches the model your documentation voice by demonstration: two to four complete example notes outperform any paragraph of style adjectives, because models are pattern-completion engines and examples collapse the ambiguity that descriptions leave open. Three examples spanning your range (routine, clinically active, risk-documented) is the working standard.
  • The gate is absolute: every example note is de-identified before it leaves the EHR, every time, regardless of how trustworthy the tool is. The model does not need the client's identity to learn your voice; every identifier left in is risk taken for zero benefit. Notes from 42 CFR Part 2 charts are not used as examples at all.
  • De-identification is two passes: a mechanical pass against the HIPAA Safe Harbor categories under 45 CFR 164.514(b)(2) (names, dates beyond year, geography below state, ages over 89, and the rest), then a narrative pass for re-identification risk, because behavioral health notes identify people through story details (the fire-chief husband, the small-town pediatric oncologist), not just through names. The test is the conference-projector test, verified by a second, dated, initialed read.
  • Before the workflow goes live, verify three documents and file the evidence: the BAA for your tier, the DPA or AI addendum committing in writing that inputs are not used for training and retention is defined, and the subprocessor list, because "our model provider trains" hides one contract layer down. Confirm training toggles are off; re-verify annually and at plan changes.
  • The block itself is a framing line ("style examples only, do not reuse their clinical content"), three input/output pairs so the model learns the transformation, and a closing instruction ("facts from the new input only; flag gaps with [MISSING]"). The risk-documented example teaches the model what clinician-assessed risk documentation looks like, reinforcing that the assessment is never the model's to make.
  • Watch the three failure modes: content bleed (example facts surfacing in new drafts; defend with framing lines and a line-by-line read), style lock (cloned-sounding notes, an audit flag in themselves; defend with range-spanning examples), and the stale example set (refresh on material change, always re-de-identifying from the EHR, never editing an identified note inside the tool).
  • Your artifact is the De-Identification + Few-Shot Protocol: a one-page procedure, the de-identified example file with an initialed checklist, and dated contract evidence. For associates it is chosen with the supervisor and named in the supervision agreement; for group practices it is the policy fix for twelve clinicians pasting whatever into whatever. The L2 capstone's four progress notes are drafted under it.