AI for Mental & Behavioral Health Clinicians
Proficient · M3 · lesson 3 of 30 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Building a Pre-Sign Note Verification Checklist
📖
now learning

Building a Pre-Sign Note Verification Checklist

15 min

It is 9:54 PM and Maria has seven AI-drafted notes open in SimplePractice, each one fluent, well organized, and unread. The drafts look finished, which is exactly the danger: a polished paragraph invites a polished signature, and a signature is not a formatting step, it is a legal attestation that every word above it is true. One fabricated quote, one inferred intervention, one risk sentence the AI softened, and that attestation becomes the exhibit in a recoupment letter or a board complaint. By the end of this lesson you will have a 10-item pre-sign note verification checklist that takes about ninety seconds per note, plus an edit-distance audit log that proves, months later, that a human clinician actually reviewed and revised every AI draft before signing. This is the ai note audit checklist that turns "I always read my notes" from a claim into a record.

Why the Signature Is the Whole Game

Start with what the signature legally is, because everything in this lesson hangs on it. When you sign a progress note, you are not closing a document. You are attesting, as a licensed clinician, that the clinical content is accurate, that the services described were rendered as described, that the time billed was the time spent, and that the risk assessment reflects your judgment. Payers treat the signed note as the claim's evidence. Boards treat it as your professional account. Attorneys treat it as your testimony, written in advance. No payer, board, or court has a category called "the AI wrote that part." There is only your name and your license number.

AI drafting changes the failure mode of documentation, and most clinicians have not updated their habits to match. Before AI, the typical defect in a note was omission: the thin 9:54 PM note that says "processed trauma material, will follow up" and fails to establish medical necessity for a 90837. With AI, the typical defect is fabrication that reads like competence: an intervention the model inferred but you never used, a client quote assembled from plausible phrasing, a number of minutes the model defaulted to because the prompt did not supply one. A thin note loses a therapy documentation audit. A fabricated note can lose a license, because the auditor's question shifts from "was this enough?" to "was this true?"

That is why the verification step cannot live in your intentions. It has to live in a fixed, written sequence you run the same way every time, the way an anesthesiologist runs the same machine check before every induction regardless of how routine the case looks. The routine cases are precisely where checklists earn their keep, because vigilance fades on the eighth note of the night and the checklist does not.

The Controlling Analogy: A Preflight Checklist, Not a Proofread

Hold one picture through this whole lesson: the pre-sign checklist is a preflight inspection, not a proofread. A proofreader asks whether the sentences are good. A pilot walking around the aircraft does not care whether the fuselage is attractive; she checks specific failure points in a specific order because each item on the card corresponds to a known way the aircraft has killed people. Flaps, fuel, control surfaces, pitot cover. The card exists because experienced pilots, left to memory, skip steps on the flights that feel routine, and the National Transportation Safety Board reports are full of routine flights.

Your AI-drafted note has known failure points too, and they are remarkably consistent across tools, whether you use Mentalyc, Upheal, Heidi, Twofold, or a general model behind a BAA. The model fabricates specifics it was not given. It infers interventions from context. It normalizes risk language toward the statistical middle, softening "client described a plan" into "client endorsed some distressing thoughts." It defaults CPT-relevant details like session minutes. It produces quotes that sound like your client because it has read ten thousand clients who sound similar. Each of those failure modes gets one line on the card. You walk the note the way the pilot walks the plane: same order, every time, no skipping because the draft "looks clean." Looking clean is what AI drafts do. It is their one reliable talent.

The analogy also tells you what the checklist is not. It is not a quality rubric for great writing, not a place to wordsmith, and not a review of your clinical decisions, which were made in the room. It is a defect detector. Ninety seconds, ten items, sign or fix. If an item fails, you do not annotate it for later; you fix the note now or you do not sign it tonight.

The Ten Items, One at a Time

Here is the card, in the order you run it. Items one through five verify clinical truth; items six through ten verify billing and regulatory integrity. Take them slowly the first twenty times. After that, the sequence becomes muscle memory and the ninety-second estimate becomes real.

1. Facts. Every concrete fact in the note happened in the session: events the client reported, dates, names, medication changes, homework reviewed. AI fills factual gaps with plausible inventions, and plausible is the problem. Read each factual claim and ask "did I hear this on Tuesday, or does it merely sound like this client?" 2. Mental status exam. The MSE describes the client you actually saw: affect, mood, speech, thought process, orientation, insight, judgment. Scribes love a template MSE ("alert and oriented x4, affect congruent") that may flatly contradict the tearful, pressured client in front of you. A boilerplate MSE that contradicts the narrative section is an internal-inconsistency flag any auditor can spot. 3. Risk. The risk language is yours, in your words, reflecting your determination. AI never scores risk, never assigns a level, never decides that ideation was passive. You made the risk call in the room; the note must state your call, not the model's paraphrase of it. If the draft contains any risk sentence you did not author or explicitly verify against your own assessment, rewrite it by hand. 4. Intervention. The interventions named are the ones you used, named the way you used them. "Utilized CBT techniques" when you actually did EMDR resourcing is a fabrication, not a simplification. 5. Plan. The plan reflects what you and the client actually agreed to: the homework assigned, the next-session focus, the referral discussed, the measure due. A model will happily generate a generic plan, and a generic plan tells a reviewer the treatment is generic.

6. CPT alignment. The code matches the service: a 90837 note must reflect a 53-plus-minute psychotherapy session, with the minutes stated, and you are the only one who held the clock. 7. Medical necessity. The note answers why this session, now, for this diagnosis: symptom severity, functional impairment, the measurement delta, the connection to the treatment plan goal. This is the sentence the UnitedHealthcare reviewer reads first in a 90837 audit defense. 8. No fabrication. A final sweep for anything asserted that did not occur: a phone call never made, a coordination contact never attempted, a symptom never reported. 9. No hallucinated quotes. Every quotation mark in the note brackets words the client actually said. If you cannot vouch for the exact words, remove the quotation marks and paraphrase, because a fabricated quote attributed to a client is the single worst artifact to explain under oath. 10. Timeliness. The note will be signed inside the regulatory and payer window that applies to you, with same-day as the house rule. A pattern of late signatures is an audit finding all by itself.

The Three Items That Carry the Most Risk

All ten items matter, but three of them, risk, quotes, and CPT alignment, account for the catastrophic outcomes, so give them disproportionate attention while the habit forms.

Risk first, because it is the bright line of this entire program. The model is a language engine, and language engines regress toward typical language. Most therapy sessions do not involve active suicidal ideation, so when ideation appears in a transcript, the statistically "expected" continuation is reassuring language, and drafts drift reassuring. The clinician decides the risk level; the clinician authors the risk documentation; AI may format and structure only after that determination exists in your words. If you remember one rule from Level 3, make it this one: a risk paragraph you did not write is a risk paragraph you cannot defend.

Quotes second, because they are uniquely damning. A reviewer who finds one demonstrably fabricated client quote stops trusting the entire chart, and reasonably so. Some clinicians adopt a personal rule worth copying: quotation marks appear in a note only when the words came from a verified transcript or were written down in session; everything else is paraphrase. It costs you nothing clinically and removes a whole category of exposure.

CPT alignment third, because it is where the money lives. The verifiable detail AI cannot supply is the worked example to internalize: a 90837 requires the actual minutes ("53 minutes, 4:02 to 4:55 PM"), the medical-necessity line requires the actual measurement delta ("PHQ-9 today 13, down from 17 four weeks ago"), and the intervention line requires the modality you actually named in session. Those three specifics come from you or they do not exist. A note that is fluent everywhere and vague exactly there reads, to an experienced auditor, like an AI draft nobody verified, which is precisely the inference you are building this checklist to prevent.

The checklist does not exist to make your notes better. It exists to make your signature true.

The Edit-Distance Audit Log: Proving the Human Was There

The checklist protects each note. The audit log protects you across all of them, and this is the piece most clinicians skip because nobody told them why it matters. Imagine the question a board investigator or a Recovery Audit Contractor will eventually ask: "You say you review every AI draft before signing. Show me." If your answer is "you have my word," you have nothing. If your answer is a log showing that, across 312 notes this quarter, you made substantive edits to 87 percent of drafts, rewrote the risk section by hand in every note where risk content appeared, and your median review time supports actual reading, you have evidence. The log converts diligence from a character claim into a documented practice.

Edit distance, for this purpose, does not require software that computes Levenshtein distance, although some EHR-integrated scribes are beginning to surface revision metrics. A practical clinician-grade version is a simple three-level code you record per note: E0, signed with no changes or trivial changes (typos, formatting); E1, moderate edits (sentences corrected, details added, the minutes and measure scores inserted); E2, substantive rewrite (risk language re-authored, interventions corrected, factual errors removed). Each row of the log captures date of service, client identifier (your internal ID, never a name in a side document), the tool used, the edit level, which checklist items failed if any, and the sign date. Five fields, fifteen seconds, done in the same breath as the signature.

Two warnings about what the log will reveal. First, a long run of E0 entries is not a sign the tool is excellent; it is a sign you have stopped reading, because no scribe on the market drafts a clinically perfect note often enough to justify months of zero-edit signatures, and an auditor will read that pattern exactly that way. Second, the log is discoverable, and that is a feature, not a bug. A discoverable record of consistent, substantive human review is the best possible thing for opposing counsel to find, because it is the documentary opposite of negligence. You are not hiding your AI use; you are documenting your supervision of it. That posture, supervised tool use with receipts, is the one that survives a therapy documentation audit, and it is the same posture the next lesson extends from per-note checking into quarterly quality measurement.

A Story: The Audit That Turned on Version History

Here is how this plays out when it goes wrong, and then when it goes right. A peer of Maria's, an LPC two towns over, drew a high-utilization 90837 review: 30 notes requested, standard fare. The notes were articulate, which initially looked good. Then the reviewer noticed that 11 of the 30 contained an identical sentence pattern in the intervention section and that two notes referenced "continued grief processing regarding the client's father" for a client whose intake clearly documented both parents living. The clinician had been signing AI drafts unread for months. The EHR's version history, which she had never thought about, showed each note created and signed within the same two-minute window, at 11 PM, in batches of six. The version history did not show AI use directly; it showed the absence of review, which was worse. The recoupment demand covered every sampled note extrapolated across the audit period, and the payer referred the fabricated-content question onward.

Now run the same audit against a clinician with the system from this lesson. The notes requested arrive with specifics in exactly the places AI cannot supply: minutes, measure deltas, named modalities. The version history shows a draft created at 5:10 PM and a signed final at 5:14 PM with visible revisions in between, including hand-edited risk language on the two sessions where ideation was assessed. The edit-distance log, produced on request, shows the quarter's pattern: mostly E1, E2 on every risk-containing note, three drafts rejected entirely with the checklist item that failed. The reviewer is not evaluating whether this clinician used AI. The reviewer is looking at evidence that a licensed professional verified every clinical assertion before attesting to it. The first audit turned on version history and ended a practice. The second turned on version history and ended in a closed file. Same technology, same payer, opposite outcomes, and the difference was about two minutes per note.

Running It at Volume: Eight Clients a Day Without Cutting Corners

A checklist that only works on light days is decoration. Maria sees eight clients a day; the system has to survive Thursday. Three operating rules make it survivable. First, verify in the gap, not in the batch. The checklist takes ninety seconds when the session is hours old and fifteen minutes when it is days old, because every uncertain fact sends you back to memory that no longer exists. Run the card in the ten-minute gap after each session while the AI draft is fresh and so is your recall. The 9:54 PM batch is where item 1 and item 9 failures get signed, because tired clinicians grade fluency instead of truth.

Second, fix or refuse, never defer. A note that fails an item gets corrected on the spot or it does not get signed in that sitting. The most dangerous workflow is the mental note to "come back and fix the risk paragraph," because the EHR does not distinguish a note you meant to fix from a note you attested to. If you truly cannot fix it now, leave it unsigned and log why; an unsigned note tonight beats a false attestation forever.

Third, let the failures teach you. Every checklist failure is data about your tool and your prompts. If item 4 fails weekly because the scribe keeps inferring "CBT techniques," your note prompt needs a hard instruction: name only interventions stated by the clinician, and write "intervention: [clinician to specify]" when none was stated. If item 2 keeps producing boilerplate MSEs, require the prompt to draw MSE content only from explicitly dictated observations. Tally which items fail and how often, per tool. That tally is the raw material for the quarterly hallucination audit in the next lesson, and it is also how you discover, before a payer does, that a tool update quietly changed your scribe's behavior. Checklists catch defects; logged checklists catch trends.

The Applied Problem: Build Your Pre-Sign Checklist and Edit-Distance Log

Your deliverable is a two-part artifact: the Pre-Sign Note Verification Checklist, a single card with the ten items, and the Edit-Distance Audit Log, a running table you fill at signature time. Build both tonight; calibrate them over the next ten notes.

Step one, draft the card. Open a blank document titled "Pre-Sign Note Verification Checklist v1, [your name], [date]." List the ten items in run order, each as a yes/no question in your own clinical vocabulary: 1. Facts: every factual claim occurred in session. 2. MSE: describes the client I actually saw. 3. Risk: risk language is mine, reflects my determination, AI scored nothing. 4. Intervention: only modalities I actually used, named correctly. 5. Plan: matches what we actually agreed. 6. CPT: code matches service; minutes stated and true. 7. Medical necessity: severity, impairment, and measure delta present. 8. No fabrication: nothing asserted that did not occur. 9. No hallucinated quotes: every quote verified or converted to paraphrase. 10. Timeliness: signing inside my regulatory and payer window. Add one footer line: "Any NO = fix now or do not sign." Print it. Tape it to the monitor bezel. A checklist that lives in a drawer is a proofread.

Step two, build the log. Create a spreadsheet (inside your HIPAA-compliant environment, internal client IDs only) with six columns: Date of service, Client ID, Tool used, Edit level (E0/E1/E2), Checklist items failed, Date signed. Define the three edit levels at the top of the sheet exactly as this lesson defines them so a future reader, possibly an auditor, can interpret the codes without you in the room. If your scribe or EHR surfaces version history or revision metrics, note in the header that version history is retained in the EHR and the log is the index to it.

Step three, run the calibration. For your next ten AI-drafted notes, run the full card out loud or with a pen, slowly, and time yourself. Expect three to four minutes per note at first, dropping toward ninety seconds by note ten. Record every failure faithfully, including the embarrassing ones; the failures are the point. "Done" looks like this: a printed card at your workstation, a log with ten honest rows, at least one E2 entry where you re-authored risk language by hand, and a one-line note to yourself about which item fails most with your current tool. That last line is your entry ticket to the quarterly quality protocol in the next lesson.

Key Takeaways

  • A signature on a progress note is a legal attestation that every word is true, not a formatting step. Payers, boards, and courts have no category for "the AI wrote that part"; there is only your license, so verification must happen before the signature, every time.
  • AI shifts the dominant documentation defect from omission to fluent fabrication: inferred interventions, plausible invented facts, template MSEs, softened risk language, and assembled quotes. Polished drafts invite unread signatures, which is exactly why a fixed checklist beats vigilance.
  • The pre-sign checklist is a preflight inspection, not a proofread: ten items in a fixed order (facts, MSE, risk, intervention, plan, CPT alignment, medical necessity, no fabrication, no hallucinated quotes, timeliness), each mapped to a known AI failure mode. Any failed item means fix now or do not sign.
  • Risk is the bright line: AI never scores risk, never assigns a level, and never softens your determination. The clinician decides risk in the room and authors the risk language by hand; a risk paragraph you did not write is one you cannot defend.
  • The verifiable details only you can supply, session minutes for the 90837, the measurement delta like a PHQ-9 drop from 17 to 13, and the modality you actually named, are the heart of medical necessity and the first things a payer reviewer checks in a 90837 audit defense.
  • The edit-distance audit log (date, client ID, tool, E0/E1/E2 edit level, failed items, sign date) converts "I always review my drafts" into documentary evidence. A long run of E0 entries means you have stopped reading, and an auditor will read the pattern the same way.
  • Version history cuts both ways: batch-signed two-minute notes at 11 PM prove the absence of review, while visible draft-to-final revisions plus a consistent log prove supervised, attestation-worthy AI use. The difference in an audit is roughly two minutes per note, invested before signing instead of surrendered after the recoupment letter.