AI for Mental & Behavioral Health Clinicians
Capable · M12 · lesson 12 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Getting Useful Clinical Output: The Specificity Lever
📖
now learning

Getting Useful Clinical Output: The Specificity Lever

15 min

Two clinicians use the same AI tool on the same night. One gets a draft so generic she rewrites it from scratch, concluding AI is overhyped. The other gets a draft that names her intervention, carries her PHQ-9 delta, and survives her supervisor's review with two edits. The difference is not the tool, the subscription tier, or luck. It is specificity: the precision of what each clinician told the model about the session. In the last lesson you built the five-part prompt skeleton; this lesson is about loading it. You will run a side-by-side experiment, a generic "write a note" prompt versus a prompt that names the modality, the session focus, the assessment scores, and the note format, and you will learn the four words that move output quality more than any others: modality, focus, intervention, plan. By the end you will own a Specificity Upgrade Worksheet that turns any vague prompt into a precise one in under three minutes. This is the working core of prompt engineering for therapists.

The Specificity Lever: Why Detail In Equals Defensibility Out

Picture the model's output as water finding a channel. A generic prompt is flat ground: the water spreads into the widest, shallowest pattern available, which in clinical writing means the most statistically common therapy note in the model's training data. That note mentions "anxiety" generically, "coping skills" generically, "progress" generically. It is not wrong, exactly. It is unfalsifiable, and unfalsifiable is the problem: a note that could describe any client in any session documents no medical necessity for this client in this session. Payer reviewers are trained to spot exactly this pattern, and so is your clinical supervisor.

Specificity cuts a channel. Every precise term you put in the context block, the modality by name, the worksheet you actually used, the score you actually collected, eliminates thousands of generic alternatives the model would otherwise reach for. Tell it "therapy session" and it can write anything. Tell it "session 6 of DBT skills training, distress tolerance module, TIPP skill taught and practiced in session" and the space of plausible sentences collapses to a narrow band that sounds like your session, because it is built from your session. The lever is mechanical, not magical: the model predicts the next word from the words you gave it, so clinical words in produce clinical words out.

Carry one analogy through this lesson: the prompt is a referral with attachments. Lesson one taught you the referral letter's structure; this lesson teaches you what to attach. A referral that says "please see this client" wastes the consultant's time. A referral that attaches the score, the modality history, and the specific question gets a useful consult on the first pass. Specificity is not extra work added to AI use. It is the two minutes of clinical recall that makes the other fifteen minutes unnecessary.

The Side-by-Side: One Session, Two Prompts

Here is the experiment, with a fabricated, de-identified case you can replicate today. The client, "T.," is an adult in week six of treatment for generalized anxiety with panic features. Today's 45-minute session was CBT, focused on interoceptive exposure: T. practiced induced dizziness (30 seconds of spinning) in session, rated peak anxiety 80 SUDS, watched it fall to 35 within four minutes without escaping or using a safety behavior, and connected the experience to the catastrophic thought "these sensations mean I am dying." GAD-7 today: 12, down from 17 at intake. Homework: one interoceptive exercise daily with a SUDS log. Plan: continue weekly, begin in-vivo exposure hierarchy next session. You inquired about and the client denied any suicidal ideation; that assessment was yours.

Prompt A, the generic version: "Write a progress note for an anxiety therapy session."

A representative excerpt of what comes back: "Client attended session and discussed ongoing anxiety symptoms. Therapist utilized cognitive behavioral techniques to help client manage worry. Client was receptive and engaged. Client is making progress toward treatment goals. Plan: continue current treatment approach and monitor symptoms."

Read it as a reviewer. Which cognitive behavioral techniques? There is no intervention named, so there is nothing to bill against. "Receptive and engaged" is a posture, not a response to treatment. "Making progress" toward which goal, measured how? "Continue current treatment approach" is a plan that plans nothing. This note would be equally true, which is to say equally empty, for every anxiety client in your caseload. If you submitted ten of these under 90834, a utilization reviewer would see ten identical sessions and start asking why the treatment never changes and never ends.

Prompt B, the specific version (built on your lesson-one skeleton, context and format loaded): "You are an experienced licensed clinician writing audit-ready outpatient psychotherapy notes. Draft a progress note for the session below. Context: adult client, week 6, generalized anxiety with panic features. 45-minute individual CBT session. Focus: interoceptive exposure. Intervention: induced dizziness exercise (30 seconds spinning) conducted in session; peak SUDS 80, declining to 35 within 4 minutes; no escape or safety behaviors; linked to catastrophic thought 'these sensations mean I am dying.' GAD-7 today: 12 (intake: 17). Clinician inquired directly about suicidal ideation; client denied SI, plan, and intent. Homework: one interoceptive exercise daily with SUDS log. Plan: continue weekly sessions; begin in-vivo exposure hierarchy next session. Format: SOAP, labeled sections, past tense, refer to client as 'Client.' Constraints: use only facts provided; no invented quotes, symptoms, or risk findings; do not state a risk level; write [CLINICIAN TO COMPLETE] for anything required but not provided. Maximum 250 words."

The Prompt B excerpt: "O: Conducted interoceptive exposure (induced dizziness, 30 seconds). Client's peak anxiety reached 80 SUDS and decreased to 35 within four minutes without escape or safety behaviors. GAD-7 administered: 12, decreased from 17 at intake. [CLINICIAN TO COMPLETE: MSE observations]. A: Client demonstrated tolerance of feared interoceptive sensations and identified the associated catastrophic cognition, consistent with continued response to CBT for generalized anxiety with panic features. P: Continue weekly individual CBT; initiate in-vivo exposure hierarchy next session; client to complete one interoceptive exercise daily with SUDS log."

Same model. Same night. One draft is a liability; the other is a working document with a flagged gap waiting for your MSE. The entire difference traveled through the context block.

The Four Words That Move Quality Most: Modality, Focus, Intervention, Plan

If you remember nothing else from this lesson, remember four words. When clinicians upgrade prompts, four additions account for most of the quality gain, and they map directly onto what a payer audit looks for.

Modality. Name the treatment model by its real name: CBT, DBT, EMDR, ACT, IFS, MI, CPT for PTSD. Each modality carries its own technical vocabulary, and naming it switches the model into that vocabulary. A prompt that says "CPT for PTSD, session 5, worked on the impact statement and the assimilated stuck point 'the assault was my fault'" produces stuck points, Socratic dialogue, and ABC worksheets. A prompt that says "trauma therapy" produces candles-and-journaling prose. The modality word is also your first audit anchor: medical necessity reviews ask whether a recognized treatment is being delivered, and the note can only say so if the prompt did.

Focus. What was this session about, in one phrase? "Distress tolerance module," "interoceptive exposure," "preparing the impact statement," "values clarification around parenting." Focus is what distinguishes session 6 from session 5 and session 7. Notes without a focus are the identical-ten-notes problem; notes with one show a treatment that is going somewhere.

Intervention. The specific thing you did, with the client's specific response. Not "used CBT techniques" but "conducted induced-dizziness interoceptive exposure; SUDS 80 to 35 in four minutes without safety behaviors." Intervention-plus-response is the clinical heart of the note, the line that proves a service was rendered and that it did something. It is also the line AI absolutely cannot supply, because only you know what you did and what happened next. This is where specificity is not optional: a model given no intervention will invent one, and an invented intervention in a signed note is a misrepresentation of the service billed.

Plan. What happens next, concretely: homework with parameters, the next clinical target, frequency, any coordination. "Continue treatment" is a non-plan; "begin in-vivo exposure hierarchy next session; daily interoceptive practice with SUDS log" is a plan a concurrent-review nurse can see continuing care justified by.

Modality, focus, intervention, plan: the four words that turn "a therapy note" into "the note from this session." Everything generic in an AI draft traces back to one of these four being missing from the prompt.

Two Force Multipliers: The Named Format and the Carried Score

Beyond the four words, two further specifics multiply everything else. The first is the named format. SOAP, DAP, BIRP, and GIRP are not interchangeable labels; they distribute the same clinical content differently, and your setting usually dictates which one your reviewer expects. Naming the format in the prompt does more than arrange headings: it creates slots the model must fill, and empty slots are where the [CLINICIAN TO COMPLETE] placeholders surface. A SOAP request exposes a missing Objective; a BIRP request exposes a missing Response. The format is a checklist disguised as a structure, and asking for it explicitly is how you make the model run the checklist for you. You will go deep on choosing among these formats in the next chapter; for now, the rule is simply that the prompt always names one.

The second multiplier is the carried score. An assessment number with its prior value, GAD-7 of 12 against an intake 17, a PHQ-9 of 11 against 14, is the densest evidence in any behavioral health note. It is objective, it is trended, and it is the kind of verifiable detail no model can fabricate safely because the true value exists only in your records. Put the score and its comparison point in every prompt where you have one. The draft will weave it into the Objective and Assessment sections, and your note acquires the one sentence utilization reviewers most want to see: measurable change attributable to treatment. The discipline this builds runs in both directions, too. Clinicians who prompt with scores start collecting scores more consistently, because the prompt has a slot that wants filling. The worksheet you build at the end of this lesson includes a scores line for exactly that reason.

A caution that belongs here, before habit forms: specificity means specific clinical detail, never identifying detail. The model needs "adult client, week 6, generalized anxiety with panic features." It never needs a name, an employer, a city, a date of birth, or the unusual occupation that makes a client identifiable in a town of 9,000. The specificity lever moves on clinical structure: diagnosis, modality, focus, intervention, response, scores, plan. Identity adds zero output quality and unbounded risk. If a detail would help a stranger recognize the client, it does not belong in the prompt, no matter what platform you are using.

Running the Experiment in Your Own Practice

Reading a side-by-side is persuasive; running one is converting. Here is the protocol, and it uses only fabricated material, so you can do it tonight on any tool without a single compliance question. Take the T. case above, or invent your own analog from a modality you practice: a DBT skills session, an EMDR reprocessing session at SUD 6 dropping to 2, an ACT session on cognitive defusion, an MI session working ambivalence about drinking with a change-talk summary. Write Prompt A in one line, the way a tired clinician types at 9:54 PM. Then write Prompt B from your lesson-one template, loading the four words plus format plus score. Run both. Print or paste the outputs side by side.

Now grade them with a highlighter and three colors. First color: every phrase in each output that could appear in any client's note, "engaged in session," "making progress," "continue current approach." Second color: every phrase traceable to a specific fact in your prompt, the SUDS values, the named exercise, the score delta. Third color: anything in the output traceable to nothing, which in a well-constrained Prompt B should be zero and in Prompt A is often most of the note. The visual result is the lesson: Prompt A comes back mostly first-color generic with a dangerous scatter of third-color invention; Prompt B comes back dominated by second-color traceable content with a placeholder where your MSE belongs.

Then do the step most clinicians skip: ask what Prompt B still got subtly wrong. Maybe it called the exercise "exposure therapy" generically instead of "interoceptive exposure" in one sentence. Maybe it softened "no safety behaviors" into "managed anxiety well." These small drifts matter, because they blur exactly the precision you paid for, and catching them rehearses the line-by-line review you owe every draft anyway. Specificity gets you to a strong draft; it never gets you out of reading it. The clinician signs the note, and the signature attests every word, including the ones the model softened.

Specificity Is Not Volume: The Curated Context

A misunderstanding to kill early: the specificity lever is not "paste everything." Clinicians discover the lever and start dumping entire intake summaries, prior notes, and stream-of-consciousness session recollections into the context block, on the theory that more input must mean better output. It does not. Irrelevant context dilutes the signal: the model gives weight to whatever you include, so three paragraphs of history can crowd this session's intervention out of the draft, and you get a note that re-narrates the intake instead of documenting today. Worse, wholesale pasting is how identifying details slip in, because nobody proofreads a dump.

The skill is curation: the smallest set of facts that fully determines today's note. In practice that is eight to twelve lines, the same slots your lesson-one template already has: diagnosis, session number and length, modality, focus, intervention with response, scores with priors, symptoms reported this week, the risk inquiry you conducted and its result, homework, plan. If a fact does not change what this note should say, it does not go in the prompt. This is the same editorial judgment you already exercise when you present a case in consultation group in ninety seconds instead of nine minutes, and it is why experienced clinicians get better AI output than technically savvy non-clinicians: the lever rewards clinical thinking, not typing.

Curation also keeps you fast. The whole economic case for AI-assisted documentation, the case that matters to Maria with seven notes left, is minutes. A curated context block takes two to three minutes to write and saves ten to fifteen of drafting; an everything-dump takes ten minutes to assemble, produces a worse draft, and still requires the full read. Specific and small beats comprehensive and bloated every night of the week.

The Applied Problem: Your Specificity Upgrade Worksheet

Your artifact from this lesson is the Specificity Upgrade Worksheet: a one-page form that converts any vague documentation prompt into a loaded one. Build it in four steps, and keep it next to the five-part template from lesson one; together they are your complete prompting kit for this level.

Step 1: Draw the form. Two columns. Left column header: "What I would have typed" with three blank lines, room for the one-line generic prompt. Right column header: "The upgrade" with ten labeled rows: Modality (by name: CBT, DBT, EMDR, ACT, IFS, MI, CPT); Session focus (one phrase); Intervention used (the specific exercise or technique); Client response (observable, with numbers where you have them: SUDS, belief ratings, minutes); Scores (today's value and the prior value); Symptoms reported this week; Risk inquiry conducted and result; Homework with parameters; Plan (next target, frequency); Format (SOAP, DAP, BIRP, or GIRP). Mark the four highest-leverage rows, modality, focus, intervention, plan, with a star; if you have only ninety seconds, those four are the upgrade.

Step 2: Test it on the worked case. Fill the worksheet from the T. interoceptive-exposure session in this lesson. Left column: "Write a progress note for an anxiety therapy session." Right column: the ten facts. Then transfer the right column into your five-part template's context block and run it. The output passes if every clinical sentence is traceable to a worksheet row and the MSE slot shows a placeholder.

Step 3: Run the highlighter grade. Three colors on both outputs, generic-anywhere phrases, traceable phrases, untraceable phrases, exactly as described above. Staple the graded pages to the worksheet the first time; it is the evidence that converts the habit.

Step 4: Define done. The worksheet is done when you can fill all ten rows for a real session from memory in under three minutes, when the starred four rows are reflexive, and when a colleague using your worksheet on their own fabricated session produces a draft with zero untraceable sentences. Save it as v1; lesson three will add the matching review checklist for the other side of the keyboard.

Key Takeaways

  • Specificity is the lever: the model predicts output from your input, so a generic prompt produces the statistically average note, which is unfalsifiable, identical across clients, and exactly what payer reviewers and supervisors flag. Precise clinical terms collapse the output space to your actual session.
  • Four words move quality most: modality, focus, intervention, plan. Name the treatment model (CBT, DBT, EMDR, ACT, IFS, MI, CPT for PTSD), state what this session was about, describe the specific intervention with the client's measurable response, and give a concrete next-step plan.
  • Two force multipliers complete the upgrade: the named note format (SOAP, DAP, BIRP, GIRP), which creates slots that expose gaps as [CLINICIAN TO COMPLETE] placeholders, and the carried score (GAD-7 of 12 against intake 17), the densest, most audit-valuable sentence in any behavioral health note.
  • The side-by-side experiment is the proof: the generic prompt returned "utilized cognitive behavioral techniques" and "continue current treatment approach," while the specific prompt returned the named interoceptive exposure, the SUDS 80-to-35 trajectory, and the in-vivo hierarchy plan. Same model, same night; the difference traveled through the context block.
  • Specific means clinically specific, never identifying. Diagnosis, modality, focus, intervention, response, scores, and plan raise output quality; names, employers, cities, and recognizable details add zero quality and unbounded risk. If a stranger could recognize the client from it, it stays out of the prompt.
  • Specificity is curation, not volume. Eight to twelve curated lines outperform a pasted intake summary, because irrelevant context dilutes the signal, slows you down, and smuggles in identifiers. The lever rewards clinical thinking, which is why clinicians out-prompt technologists.
  • Specificity gets you a strong draft; it never gets you out of reading it. Small drifts survive even good prompts (softened response language, genericized intervention names), and the clinician signs the note: the signature is a legal attestation of every word, so read every word.