AI for Mental & Behavioral Health Clinicians
Capable · M15 · lesson 15 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Recognizing Bad AI Output Before You Sign It
📖
now learning

Recognizing Bad AI Output Before You Sign It

15 min

The most dangerous AI note is not the bad one. It is the fluent one: clean grammar, confident clinical register, perfect SOAP headings, and three sentences that never happened. At 9:54 PM, with six notes behind you and your judgment running on fumes, fluency reads as accuracy, and that is exactly when a fabricated quote or an inflated medical-necessity claim slides under your signature and becomes your legal attestation. The first two lessons taught you to build prompts that prevent most failures; this lesson teaches you to catch the ones that get through. You will learn to diagnose the ten failure modes of AI-drafted clinical text, see each one in actual before/after note excerpts, and build the artifact that will protect your license for the rest of your career: a personal seven-item, line-by-line, pre-sign review checklist. Every clinician using ChatGPT for therapy notes, an AI scribe, or an EHR drafting feature needs this skill before the next note gets signed.

Why Fluency Fools Tired Clinicians

Your training equipped you to spot bad clinical writing: the vague note, the missing section, the intern's wandering narrative. Nothing in your training equipped you for plausible wrong writing, because until now, producing a sentence like "Client reported a 40 percent reduction in panic frequency since initiating interoceptive exposure" required someone who knew what those words meant. A language model produces that sentence with no knowledge of your client, your session, or whether any panic was ever measured. The words are well-formed; the claim is unmoored. This is the core inversion of the AI era: polish no longer correlates with accuracy, and every review instinct you calibrated on human writing needs recalibration.

Think of yourself, from this lesson forward, as the radiologist of your own documentation. A radiologist does not glance at a film and ask "does this look like a chest?" They run a systematic search pattern, the same anatomy in the same order every time, precisely because the dangerous finding hides in the corner the eye skips when it is tired. Your AI draft review works the same way: not "does this read like my session?" but a fixed sequence of checks, run in the same order on every draft, so that the fabricated quote in the Subjective and the wrong code in the billing line get caught by procedure rather than by luck. The ten failure modes below are your pathology atlas; the seven-item checklist at the end is your search pattern.

One framing rule before the atlas. These failures are not exotic malfunctions; they are the model doing what it always does, predicting plausible text, in places where plausible is not the standard. The standard is true. That gap, between plausible and true, is the entire job of your review, and it is why the cardinal rule of this program sits at the center of this lesson: the clinician signs the note, and the signature is a legal attestation that every word is accurate, not a formatting step at the end of a workflow.

Failure Modes 1-3: The Fabrications

Failure mode 1: fabricated quotes. The model puts words in your client's mouth, inside quotation marks, because training data taught it that good notes contain quotes. Before: your context said "client reported feeling overwhelmed by coparenting conflict." After, in the draft: Client stated, "I can't keep living like this, something has to give." That sentence was never spoken. In a custody case, a subpoenaed record, or a board complaint, a quotation mark is a claim of verbatim speech, and a fabricated one is indefensible. Detection rule: every quotation mark in a draft is guilty until you personally remember or have recorded the client saying those words. When in doubt, convert to paraphrase or delete.

Failure mode 2: invented diagnoses. The draft upgrades, adds, or swaps diagnoses to match its statistical priors. You documented adjustment disorder with depressed mood, F43.21; the draft writes major depressive disorder, or appends "with anxious distress," or casually mentions "the client's PTSD." Each invented diagnosis is a clinical determination you never made, now sitting in a legal record that follows the client into insurance histories, custody evaluations, and security clearances. Detection rule: every diagnosis and code in the draft must match your actual diagnostic formulation character for character.

Failure mode 3: inflated medical necessity. The subtlest fabrication, because it flatters you. The model has learned that strong notes justify treatment, so it overstates: "client's severe functional impairment across occupational and social domains necessitates continued intensive weekly psychotherapy" for a client you assessed as moderately impaired and improving. Inflation feels like the model helping you get paid. It is the model drafting a misrepresentation to a payer over your signature, and a pattern of inflated notes is precisely what turns a routine audit into a fraud referral. Detection rule: every severity word (severe, significant, marked, intensive) must trace to your own assessment; if you would not say it under oath, it does not stay in the note.

Failure Modes 4-5: The Wrong-Content Errors

Failure mode 4: generic CBT language not used in session. The model's training is saturated with CBT, so CBT is its default filler: "challenged cognitive distortions," "practiced thought stopping," "reviewed coping skills." If you ran an EMDR reprocessing session, an IFS parts dialogue, or an MI conversation about ambivalence, and the draft says you "utilized cognitive restructuring techniques," the note now documents an intervention that did not occur, in a modality you were not practicing. Beyond inaccuracy, this creates a consistency problem: your treatment plan says EMDR, your notes say CBT boilerplate, and a reviewer reasonably concludes that nobody is reading these notes, which is the worst possible posture entering an audit. Detection rule: every intervention named in the draft must be one you actually delivered, in the vocabulary of the modality you actually practice.

Failure mode 5: the wrong CPT code. Drafts that mention billing love to assert codes: a 53-minute session labeled 90834 instead of 90837, a family session tagged 90791, a code invented outright. The model does not know your session length, your service type, or your payer's rules; it knows which codes appear near which words in training text. A wrong code in a signed note is a billing error you attested to, and systematic wrong codes are the fastest route to recoupment. Detection rule: the model never chooses the code. You verify session type and the actual timed minutes against the code yourself, every time, and your prompt should not ask the model to code at all. The minutes in the room are a fact only you possess, and a 90837 without documented time supporting it is an audit flag whether AI wrote it or you did.

Polish no longer correlates with accuracy. The model's job is plausible; your signature's job is true; the review is where the gap closes, line by line, before you sign.

Failure Modes 6-8: The Dangerous Omissions

Failure mode 6: missing risk assessment. The draft simply says nothing about risk, because your context said nothing about risk, and silence propagates. But a psychotherapy note with no risk documentation, for a client where risk inquiry was clinically indicated, is a gap that gets read backwards after a bad outcome: the absence of documentation becomes the alleged absence of assessment. The deeper rule, repeated throughout this program: AI never performs the risk assessment, never scores an instrument like the CSSRS, never assigns a risk level. You conduct the inquiry, you make the determination, and the note documents what you did. The draft's job is at most to carry your sentence: "Clinician inquired directly about suicidal ideation; client denied SI, plan, and intent." If that sentence is missing and the inquiry happened, you add it. If the inquiry did not happen, no sentence gets invented; that is a clinical matter, not a documentation one.

Failure mode 7: missing MSE. You met this in lesson one and it remains the most common structural gap: the draft has no mental status content, or worse, has fabricated MSE content. The MSE is observational data that exists only in your perception of the client in the room: appearance, behavior, speech, observed affect, thought process and content, insight, judgment. A well-constrained prompt yields a [CLINICIAN TO COMPLETE] placeholder here; an unconstrained one yields fiction like "affect congruent, thought process linear and goal-directed" for a session the model never observed. Detection rule: the MSE in a signed note is written or dictated by you, from your observations, full stop.

Failure mode 8: missing plan. The draft trails off after the Assessment, or closes with the non-plan you learned to spot in the last lesson: "continue current treatment approach." A note without a concrete plan fails on three fronts at once: clinically (no documented direction), legally (no evidence of clinical reasoning about next steps), and financially (nothing justifying the next authorized session). Detection rule: the Plan section must contain at least the next clinical target, the frequency, and any homework with parameters, all sourced from your prompt, never improvised by the model.

Failure Modes 9-10: The Tense Tells

Failure mode 9: present-tense client speech. The draft renders reports as standing facts: "Client feels hopeless and sees no way forward" instead of "Client reported feeling hopeless during the session." The difference is not stylistic. Past-tense reported speech documents what was said in a bounded clinical encounter; present tense asserts the client's current ongoing state, which you cannot know and did not claim. In a risk-adjacent sentence the distinction can be the whole ballgame: "Client feels hopeless" in a record read after an adverse event sounds like a clinician who documented active hopelessness and did nothing. Detection rule: client-reported content is always attributed and past tense: reported, stated, described, denied.

Failure mode 10: tense drift. Even when the draft starts correctly, models drift: Subjective in past tense, Objective sliding into present, Plan wobbling between "will practice" and "practices." Drift is partly cosmetic, but it is also the single most reliable tell, to a supervisor, an auditor, or opposing counsel, that a note was machine-drafted and never read. A note with three tenses announces that no clinician's eye passed over it. Detection rule: one read-through purely for tense, which takes fifteen seconds and doubles as the pass where failure mode 9 surfaces.

Notice the architecture of the ten. Three fabrications (quotes, diagnoses, inflated necessity), two wrong-content errors (generic CBT filler, wrong CPT), three omissions (risk, MSE, plan), and two tense tells (present-tense speech, drift). Fabrications make the note say false things; omissions make it silent where it must speak; the tense tells expose that nobody checked. Your checklist must catch all three classes, which is exactly how it is built.

The Before/After Clinic: One Paragraph, Six Failures

Here is a single fabricated draft excerpt, the kind that arrives looking signable, followed by the corrected version. The session: 53 minutes, EMDR reprocessing for a PTSD client; you supplied target memory work, SUD dropping 7 to 3, your direct SI inquiry with denial, and next-session plan to continue the same target.

Before (the AI draft): "Client states, 'The nightmares are destroying me.' Client presents with severe PTSD and major depressive disorder causing marked impairment in all life domains. Therapist utilized cognitive restructuring and thought-stopping techniques to challenge the client's distorted beliefs. Client feels safer and is improving. Session billed as 90834. Will continue current approach."

Run the atlas: a fabricated quote (mode 1); an invented MDD diagnosis and inflated "severe... all life domains" necessity language (modes 2 and 3); CBT boilerplate replacing the actual EMDR reprocessing (mode 4); 90834 asserted for a 53-minute session that is 90837 territory, decided by you, not the model (mode 5); no risk documentation despite your inquiry (mode 6); no MSE (mode 7); a non-plan (mode 8); "Client feels safer" as present-tense state (mode 9); and tense wobble throughout (mode 10). Six-plus failures in five sentences, every one fluent.

After (corrected under your hand): "Client reported continued trauma-related nightmares, approximately four per week. Conducted EMDR reprocessing on the established target memory; SUD rating decreased from 7 to 3 across sets. [MSE: written by clinician from in-session observations.] Clinician inquired directly about suicidal ideation; client denied SI, plan, and intent. Diagnosis remains PTSD per existing formulation. Plan: continue reprocessing the same target next session; weekly frequency; client to log nightmare frequency. Time in session: 53 minutes, documented by clinician." Every sentence traceable, every omission filled by you, every judgment yours. That is what a signable AI-assisted note looks like, and the distance between the two paragraphs is your review.

The Seven-Item Line-by-Line Review Checklist

Ten failure modes do not require ten checks, because the modes cluster. Seven items, run line by line in fixed order, cover all ten. This is the search pattern; in the Applied Problem you will personalize it.

1. Quotes and facts trace. Every quotation mark, every reported symptom, every event in the draft appears in your prompt or your memory of the session. Catches mode 1 and all general fabrication.

2. Diagnosis exact. Every diagnostic term and ICD-10 code matches your formulation character for character; nothing added, upgraded, or "with" appended. Catches mode 2.

3. Severity honest. Every severity and impairment word survives the under-oath test against your actual assessment. Catches mode 3.

4. Interventions real, in your modality's language. Every named technique is one you delivered this session, in the vocabulary of the modality you practice; no CBT filler in an EMDR note. Catches mode 4.

5. Code and minutes yours. The CPT code, if present, was chosen by you against the actual session type and the timed minutes you recorded; the model never codes. Catches mode 5.

6. The three clinician sections present and authored by you. Risk documentation reflecting the inquiry you actually conducted; an MSE written from your observations; a concrete plan with target, frequency, and homework. One check, three slots. Catches modes 6, 7, and 8.

7. Tense pass. One straight read for tense only: all client content attributed and past tense, no present-tense state claims, no drift between sections. Catches modes 9 and 10.

Run without shortcuts, the seven items take ninety seconds to three minutes on a one-page note. That is the price of signing, and it is non-negotiable at any hour. The checklist is also your answer to the question every supervisor and every malpractice carrier is starting to ask: "How do you verify AI-drafted documentation?" You will have a literal, numbered answer.

The Applied Problem: Your Personal Pre-Sign Review Checklist

Your artifact from this lesson is the Personal Pre-Sign Review Checklist: the seven items above, adapted to your modality, your setting, and your known weak spots, in a form you will actually run at 9:54 PM. Build it in four steps.

Step 1: Personalize the seven items. Copy the seven-item checklist into a document. Under item 4, write the actual vocabulary of your primary modalities (your EMDR terms, your DBT skills names, your MI processes) so "in my modality's language" has a concrete reference. Under item 5, write your own coding rules for your common services (your 90834 versus 90837 time thresholds, your add-on codes if you prescribe). Under item 6, write your standard risk-documentation sentence verbatim so you paste, never improvise, the documentation of an inquiry you conducted. Add one personal item slot at the bottom for your own recurring weak spot once you discover it; everyone has one.

Step 2: Calibrate it on the before/after clinic. Take the "Before" paragraph from this lesson and run your checklist against it cold, marking which item catches which sentence. You should net all ten failure modes across the seven items; if any failure slips through your personalized wording, tighten the item that should have caught it. This calibration is the proof your checklist works before it ever touches a real note.

Step 3: Run it live on fabricated material. Generate a fresh draft from your lesson-two specific prompt on a fabricated session, then review it with the checklist at full speed, timing yourself. Target: under three minutes with every item actually checked, not skimmed. If you cannot finish in three minutes, your note is too long or your checklist wording is too vague; fix the wording, not the rigor.

Step 4: Define done and install it. The checklist is done when it is printed or pinned where you sign notes, when you can recite the seven items from memory, and when you have run it on five consecutive drafts without skipping an item. Then make it policy for yourself: no AI-assisted note gets your signature without the seven passes. If you are pre-licensed, bring the checklist to supervision and ask your supervisor to countersign the practice; it protects their license alongside yours, and it converts "I use AI carefully" from a claim into a procedure you can show anyone, from a supervisor to a board investigator.

Key Takeaways

  • Fluency is not accuracy. AI drafts fail in confident, well-formed prose, which defeats review instincts calibrated on human writing; the answer is a fixed search pattern run on every draft, like a radiologist's systematic read, not a vibe check of whether the note "sounds right."
  • The ten failure modes come in three classes. Fabrications: invented quotes, invented diagnoses, inflated medical necessity. Wrong content: generic CBT language for sessions that were not CBT, and wrong CPT codes. Omissions: missing risk documentation, missing MSE, missing plan. Plus two tense tells: present-tense client speech and tense drift, the signatures of an unread machine draft.
  • Quotation marks are claims of verbatim speech. A fabricated quote in a signed note is indefensible in a custody record request, a subpoena, or a board complaint; every quote is guilty until you remember or recorded the client saying it, and doubt converts it to paraphrase or deletes it.
  • The clinician-only triad never delegates: AI never performs the risk assessment, never scores a CSSRS, never assigns a risk level; the MSE is written from your in-room observations; and the CPT code is chosen by you against the session type and the actual timed minutes only you possess.
  • Inflated medical necessity is the failure that flatters you, and the one that converts an audit into a fraud question. Every severity word (severe, marked, intensive) must survive the under-oath test against your real assessment of this client.
  • The seven-item line-by-line checklist covers all ten modes: quotes and facts trace; diagnosis exact; severity honest; interventions real in your modality's language; code and minutes yours; the three clinician sections (risk, MSE, plan) present and authored by you; and a dedicated tense pass. Ninety seconds to three minutes per note, every note.
  • The cardinal rule closes every loop: the clinician signs the note, and the signature is a legal attestation of every word, not a formatting step. The checklist is what makes that attestation true at 9:54 PM, and a personalized, practiced version of it is the artifact this lesson leaves in your hands.