AI for Mental & Behavioral Health Clinicians
Capable · M19 · lesson 19 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
System Prompts for a Clinical Practice
📖
now learning

System Prompts for a Clinical Practice

15 min

Maria, the Oakland LCSW with eight clients a day, types the same forty words into her AI tool every night: "Write a SOAP note, I'm an LCSW in California, don't make anything up..." On the nights she is too tired to retype it, the output drifts: the note comes back DAP when she needed SOAP, the modality says "CBT" when she ran an EMDR resourcing session, and one Thursday the draft confidently included a coping skill she never taught. Every one of those drifts is a prompt-architecture failure, not a model failure. This lesson teaches the highest-leverage piece of prompt engineering for behavioral health clinicians: the system prompt, a standing instruction block encoding your modality, your note format, your state's documentation rules, your CPT preferences, and a hard "never invent" guardrail, written once and saved as a reusable template. By the end you will have a complete clinical system prompt for therapy notes you can copy, adapt, and save, and you will see why everything in Level 3 of this certification stands on it.

What a System Prompt Is, and Why It Outranks Everything You Type Later

Every modern AI tool, whether it is ChatGPT with a Business Associate Agreement on an enterprise plan, Claude, or the model humming inside Mentalyc, Upheal, or SimplePractice's Sidekick, processes two layers of instruction. The first is the system prompt: standing instructions that define who the model is, what rules it follows, and what it must never do, set before the conversation starts. The second is the user prompt: the thing you type in the moment, tonight's session shorthand. The model treats the system layer as constitution and the user layer as legislation; when the two conflict, a well-built tool resolves toward the system prompt. That hierarchy is the entire reason this lesson exists.

Here is the controlling analogy for this lesson, and it carries through all four lessons in this chapter: a system prompt is the standing orders you would write for a brand-new scribe on their first day. You would not re-explain your note format, license type, and documentation rules at every handoff. You would write it down once: "You work for a California LCSW. We use SOAP. We never write a diagnosis I did not state. We document time in and time out because I bill 90837 and the payer audits it." After that, every handoff can be short, because the standing orders carry the weight. A clinician who retypes her requirements every night is re-hiring and re-training the scribe eight times a day, and the predictable result is the inconsistency Maria sees: format drift, modality errors, and invented content on the nights the instructions got abbreviated.

There is a second, quieter benefit. A written system prompt is an auditable artifact. When Jordan's group practice in Sacramento writes its AI policy, the question "what instructions does our AI operate under?" has a one-page answer the compliance officer, the malpractice carrier, and a board investigator can read. "Whatever each clinician typed that night" is not a defensible answer. The system prompt is to your AI workflow what the informed consent form is to your treatment relationship: the document proving the rules existed before the incident did.

The Five Components Every Clinical System Prompt Must Encode

A clinical system prompt has five load-bearing components, and the order matters because models weight early instructions heavily. First, identity and scope: who you are (license type, state, setting) and what the AI's role is (a documentation assistant that drafts for clinician review, never a clinical decision-maker). This paragraph keeps the tool on the right side of the Illinois WOPR Act and Nevada AB 406 line: AI drafts, the licensed clinician decides and signs.

Second, modality and clinical frame: the approaches you actually practice (CBT, EMDR, DBT-informed, ACT, psychodynamic) so the model describes interventions in your modality's language instead of defaulting to generic CBT phrasing, the most common drift in AI-drafted therapy notes. If you run EMDR, your note should say "resourcing," "bilateral stimulation," and "SUD rating," not "challenged cognitive distortions."

Third, note format and documentation rules: your format (SOAP, DAP, BIRP, GIRP), your section order, your state's requirements, and your payer realities. A California clinician encodes that risk documentation reflects the clinician's own assessment and that duty-to-protect analysis under Civil Code §43.92 is never the model's to make. A clinician billing 90837 encodes that every note carries start time, stop time, and total minutes, because high-frequency 90837 billing is what payer algorithms flag, and time-in-session is the verifiable detail the AI cannot supply: it must come from you, and the prompt must demand it rather than fabricate it.

Fourth, CPT and billing preferences: the codes you actually use (90791 for diagnostic evaluation, 90834 and 90837 for individual psychotherapy, 90846/90847 for family work, 96127 for brief behavioral assessments) and the rule that the AI may format billing information you provide but never selects or upgrades a code. Code selection is a clinical-billing judgment with fraud exposure attached; it stays human.

Fifth, the "never invent" guardrail: an explicit prohibition on fabricating quotes, symptoms, interventions, scores, diagnoses, or risk determinations, with the instruction to flag gaps instead of filling them. This component is so important it gets its own lesson at the end of this chapter, where you build the five-line anti-hallucination suffix: standing rule in the system prompt, per-request enforcement in the suffix. Belt and suspenders, because the failure it prevents, fabricated clinical content under your signature, is the one a license does not survive gracefully.

The Complete Clinical System Prompt, Ready to Copy and Adapt

Here is the full template. Read it slowly, because every line is doing work. This version is written for a California LCSW running a CBT and EMDR practice billing through SimplePractice; yours will differ in the brackets, not the architecture.

ROLE AND SCOPE. "You are a clinical documentation assistant for a licensed clinical social worker (LCSW) practicing psychotherapy in California. Your only function is to draft documentation from information the clinician provides. You are not a clinician. You never diagnose, never assess risk, never select billing codes, never recommend treatment, never communicate with clients. Every draft will be reviewed, edited, and signed by the clinician, whose signature is a legal attestation."

CLINICAL FRAME. "The clinician practices cognitive behavioral therapy (CBT) and EMDR. Describe interventions using the terminology of the modality the clinician names for that session. Do not default to CBT language for non-CBT sessions. Do not describe interventions the clinician did not state were used."

NOTE FORMAT. "Default format is SOAP: Subjective (client's report, with direct quotes only if the clinician supplies them verbatim), Objective (clinician's observations of presentation, affect, and mental status, exactly as stated), Assessment (the clinician's stated clinical assessment, progress toward treatment plan goals, and medical-necessity language tying the session to the diagnosis and plan), Plan (next steps as stated by the clinician). Include session start time, stop time, total minutes, CPT code, and diagnosis code only when the clinician provides them; if any is missing, write [MISSING: field] rather than a value."

DOCUMENTATION RULES. "Notes must support medical necessity: connect symptoms to diagnosis, intervention to treatment plan goal, and client response to continued care. Write in past tense, third person ('Client reported,' 'Clinician utilized'). No process-note material: the clinician's private impressions, hypotheses, and countertransference observations belong in separate psychotherapy notes under 45 CFR 164.501, not in the progress note."

RISK CONTENT. "If the input mentions suicidal ideation, homicidal ideation, abuse, neglect, or danger to others, document only the clinician's stated assessment, stated risk level, and stated actions. Never generate, score, or imply a risk level, a CSSRS score, a duty-to-protect determination, or a mandated-reporting decision. If risk content appears without a clinician-stated assessment, stop and write [RISK CONTENT PRESENT: clinician assessment required] at the top of the draft."

NEVER INVENT. "Use only facts present in the clinician's input. Do not infer, embellish, or add plausible details. Do not invent quotes, symptoms, interventions, homework, scores, or dates. Flag every gap with [MISSING] rather than filling it. When uncertain whether something was stated, omit and flag."

That is roughly 380 words. It fits in the custom-instructions field of ChatGPT, the project instructions of Claude, or the settings template of most clinical scribes, and once saved, it governs every request without being retyped. The model cannot quietly drift into clinical judgment, because the constitution forbids it before any session content arrives.

A system prompt is the standing orders you write for your scribe on day one. The clinician who retypes her rules every night is re-hiring and re-training the scribe eight times a day, and the output drift proves it.

Adapting the Brackets: State, License, Modality, Payer

The template's architecture is universal; the brackets are not, and adapting them is where your clinical judgment enters. Start with jurisdiction. The California version above assumes Civ Code §43.92's duty to protect; a New York clinician's risk paragraph references their own framework and never imports California's. Copying a colleague's system prompt across state lines is copying their compliance posture without their statutes. Cite your own board's documentation standards, your own retention rules, your own mandated-reporting framework; if you do not know them cold, that gap is a continuing-education problem the prompt cannot fix.

License type changes the identity line and sometimes the scope. A PMHNP's system prompt adds medication-management rules (medications discussed as stated, side effects as reported, never an AI-suggested dosage) and the psychotherapy add-on codes 90833/90836/90838 alongside E/M codes. A BCBA's prompt swaps SOAP for session-data documentation aligned with skills acquisition programs and CPT 97153/97155, under the BACB Ethics Code. An AMFT like Carmen in Fresno adds a line her supervisor will appreciate: "All drafts are prepared for review by both the associate and the supervising LMFT; nothing is finalized until both have reviewed," because her supervisor's signature exposure is real and the prompt should reflect the supervision structure she actually works under.

Payer reality shapes the documentation-rules paragraph hardest. If you bill 90837 at high frequency, your prompt demands time documentation and a stated rationale for the extended session, because "why 53 minutes and not 45" is the first question in a United Healthcare high-utilization review. If you take Medicaid, encode your state's golden-thread requirement linking every note to the treatment plan. If you run measurement-based care, instruct the model to include the PHQ-9 or GAD-7 score only when you supply the number, never to estimate one from the narrative: a fabricated score in a payer file is not a typo, it is a false record.

One adaptation discipline: change one bracket at a time and test before trusting the result. Prompt engineering for clinicians is closer to titrating a medication than redecorating a website: small changes, observed effects, documented versions.

Where the System Prompt Lives in Each Tool

The concept is portable; the mechanics differ by tool, and knowing where your standing orders actually live is part of owning them. In ChatGPT (Team or Enterprise, the tiers where a BAA is even possible), the system prompt goes into Custom Instructions or a Project's instructions; the free consumer tier is off the table for anything touching client material regardless of how good your prompt is, because no prompt fixes a missing BAA. In Claude, a Project's custom instructions serve the same role. In purpose-built scribes like Mentalyc, Upheal, Twofold, or Heidi, much of the system prompt is baked in, but most expose settings for note format, modality, and style; your job is to verify their baked-in rules match your standing orders rather than assuming they do, by running one de-identified test session against your components.

In EHR-embedded AI (SimplePractice's Sidekick, TherapyNotes AI, Valant's documentation tools) the system prompt is largely invisible to you, which makes your verification pass more important, not less. Ask the vendor directly: what instructions govern the model, can I customize the note format, what happens when my input contains risk content, and does the system refuse to fabricate missing fields or quietly fill them? A vendor who cannot answer those four questions has not thought about the problem at the depth your license requires.

Wherever it lives, version your prompt like the clinical document it is. Keep a dated master copy outside the tool, increment it when you change a line, and note why. When the tool updates (vendors change models under the hood without telling you), rerun your de-identified test session and confirm the standing orders still hold. The prompt you cannot produce on request is a prompt you do not really have.

The Four Failure Modes a Good System Prompt Prevents

Watch what happens without standing orders, because each failure mode maps to a missing component. Format drift: Maria's Thursday note comes back DAP instead of SOAP because the hurried prompt never specified and the model defaulted to whatever its training favored; the note-format component prevents it. Modality mislabeling: the model writes "challenged cognitive distortions" for an EMDR resourcing session because generic therapy language is CBT-shaped; in a payer review, interventions that do not match the treatment plan's stated modality read as sloppy documentation or a different service than billed. The clinical-frame component prevents it.

Fabricated specificity: the most dangerous one. Asked for a "complete" note from thin shorthand, an unconstrained model fills gaps with plausible clinical furniture: a coping skill never taught, a quote never said, a "denied SI" the clinician never asked about. That last example deserves a full stop. A note that says "client denied suicidal ideation" when the screening never happened is a fabricated clinical assessment in a legal record; if that client attempts suicide the following week, the note becomes evidence that you assessed and were wrong, when the truth is you never assessed. The never-invent component, plus the risk-content rule demanding [RISK CONTENT PRESENT] flags, prevents it.

Scope creep: over weeks, the clinician starts asking the same chat window things the tool should refuse: "what diagnosis fits this?", "is this client high risk?", "should I bill 90837 or 90834?" An unconstrained model answers all three fluently. A model under your role-and-scope paragraph responds that diagnosis, risk assessment, and code selection are the clinician's determinations, and offers to format whatever you decide. That refusal is the tool keeping you on the legal side of the WOPR line at 11 PM when your judgment about what to ask is at its weakest. The best system prompts protect the clinician from the clinician's tired self.

Testing the Prompt Before You Trust It

A system prompt is a clinical instrument, and you would not use an unvalidated instrument on a caseload. Run a three-case acceptance test before the prompt touches real work, using fictional or fully de-identified material. Case one, the clean session: a routine summary with all fields present; check format compliance, modality language, tense, and that every fact in the output traces to a fact in your input. Case two, the gappy session: shorthand with the minutes, the PHQ-9 score, and the homework deliberately omitted. A passing output shows [MISSING: total minutes], [MISSING: PHQ-9 score], [MISSING: homework]; a failing output shows the model's best guesses, and a model that guesses minutes will eventually guess a risk assessment.

Case three, the risk session: shorthand including "client mentioned passive SI" with no clinician assessment attached. The only passing output begins with [RISK CONTENT PRESENT: clinician assessment required] and refuses to characterize the risk. If the model writes "risk assessed as low," it just assigned a risk level and your system prompt failed its most important job; tighten the risk-content paragraph and rerun until the refusal is reliable. Log all three results with dates in the same file as your versioned prompt; that test log is the difference between "I use AI carefully" and being able to prove it.

Re-test quarterly and on triggers: after any vendor model update you learn about, and after any edit to the prompt itself. This is the same logic as fidelity checks in manualized treatment: the protocol existed; the question is whether it is still being followed.

The Applied Problem: Your Saved, Versioned Practice System Prompt

Your artifact for this lesson is the Practice System Prompt, v1.0: a complete, saved, dated system prompt for your actual practice, built on the template above and validated against the three-case acceptance test. This artifact is load-bearing: the few-shot examples in the next lesson, the EHR paste-back schema after that, and the anti-hallucination suffix that closes this chapter all bolt onto it, and the Level 2 capstone, a complete documentation kit for one de-identified client (intake, diagnostic formulation, treatment plan, four progress notes, a prior auth letter, and an ROI, all AI-drafted and clinician-verified line by line), is produced entirely under it.

Step one: copy the template and rewrite every bracket for your reality. Your license and state in the identity line. Your actual modalities, not aspirational ones. Your real note format, your real CPT codes, your state's actual risk framework. Where the template says California and §43.92, you say your jurisdiction or you delete the claim; never carry someone else's statute. Budget thirty focused minutes; the brackets force decisions about your own documentation standards most clinicians have never written down, which is precisely the value.

Step two: install it in your tool's system-prompt location and save the master copy as a dated plain-text file, practice-system-prompt-v1.0-2026-06.txt, in the same folder as your consent addendum and AI policy documents. Step three: run the three-case acceptance test (clean, gappy, risk) with de-identified material and log the results under the prompt text. If the risk case does not produce a refusal, revise and rerun before any real use; that test is non-negotiable.

"Done" looks like this: a one-page file containing your versioned prompt, your three test results with dates, and a change log with one entry, producible in sixty seconds if your supervisor, your group practice owner, or a board investigator asks what instructions your AI operates under. And tonight, when you draft your first real note under it, the request you type is two lines instead of forty words of retyped boilerplate, because the standing orders are finally doing their job.

Key Takeaways

  • A system prompt is the standing instruction layer that governs every request before session content arrives, and the model weights it above whatever you type in the moment. It is the standing orders you would write for a new scribe on day one; retyping requirements nightly produces the format drift, modality errors, and fabricated content that follow.
  • Every clinical system prompt needs six components in order: role and scope (documentation assistant, never clinician), clinical frame (your actual modalities), note format (SOAP/DAP/BIRP with section rules), documentation rules (medical necessity, tense, psychotherapy-notes exclusion under 45 CFR 164.501), risk content (never score, never assign, flag and stop), and never invent (flag gaps with [MISSING] instead of filling them).
  • The risk-content paragraph is the most important line you will write: if input contains SI, HI, abuse, or danger content without a clinician-stated assessment, the only acceptable output begins with [RISK CONTENT PRESENT: clinician assessment required]. AI never scores a CSSRS, never assigns a risk level, never makes the duty-to-protect or mandated-report call; in California the duty under Civ Code §43.92 is a duty to protect, and it belongs to a licensed human.
  • Adapt brackets, never architecture: jurisdiction (your statutes, not a colleague's), license type (PMHNP medication rules, BCBA's BACB code and 97153/97155, an associate's dual-review line), and payer reality (90837 time documentation, Medicaid golden thread, scores only when you supply the number).
  • Know where the prompt lives in your tool: custom or project instructions in general-purpose models on BAA-eligible tiers, settings in scribes like Mentalyc and Upheal, largely invisible in EHR-embedded AI like SimplePractice Sidekick or TherapyNotes AI, which makes vendor questions and your own test pass more important.
  • Validate before trusting: the three-case acceptance test (clean, gappy, risk) with de-identified material, logged with dates, rerun quarterly and after any prompt edit or vendor model change. A prompt without a test log is a vibe, not a control.
  • Your artifact is the Practice System Prompt v1.0: versioned, dated, tested, stored with your policy documents, producible in sixty seconds on request. The next three lessons, and the Level 2 capstone documentation kit, all build directly on it.