Verifying an AI-Drafted Note Line by Line
The promise of an AI scribe is the hour it gives back; the risk is what slips into the chart while you are busy enjoying that hour. Every lesson in this chapter has ended at the same gate: read every word before you sign. This lesson builds the gate itself, a seven-pass verification protocol that checks an AI-drafted note for factual accuracy, MSE integrity, risk-assessment integrity, intervention specificity, plan quality, CPT match, and fabricated content, in a fixed order, with a stopwatch running. The stopwatch matters as much as the passes: a verification ritual that takes fifteen minutes will be abandoned by Thursday, while one that takes four will outlast your career. Most clinicians, once the protocol is internalized, land at 4 to 7 minutes per note, total, verification included, and this lesson ends with you timing yourself against that target and building a personal time-per-note baseline. This is the lesson where AI-assisted documentation stops being a leap of faith and becomes a procedure, the difference between trusting a tool and verifying an AI progress note like the licensed professional whose signature sits under it.
Why a Fixed-Order Protocol Beats Careful Reading
The intuitive approach to checking an AI draft, read it carefully and fix what looks wrong, fails for a documented human-factors reason: fluent text suppresses scrutiny. AI drafts are grammatical, confident, and clinically toned, which means your reading brain glides; the same fabricated PHQ-9 score that would jump off a page of your own rough typing nestles invisibly inside polished prose. The countermeasure is the one every safety-critical field converged on independently: the checklist run in fixed order. A pilot does not "look the plane over carefully"; she runs the same passes in the same sequence every flight, because sequence is what defeats glide. That is the controlling analogy for this lesson: the seven passes are your preflight, the note is the aircraft, and the signature is takeoff, after which problems stop being correctable and start being survivable or not.
Fixed order has a second virtue: each pass primes the next. You check facts before fabrication because a fact-verified note narrows what fabrication can hide in; you check risk before intervention because risk findings can change what the intervention documentation must say; you check CPT last-but-one because the code can only be matched after the content is settled. And a fixed order makes the protocol timeable, which makes it improvable: when your stopwatch says pass four is eating two minutes every note, you have found a shorthand-capture problem, not a reading-speed problem, and you can fix it upstream.
One scope note before the passes: this protocol assumes the chapter's architecture is already in place. The draft came from your shorthand (or your dictation, or a BAA-covered scribe's transcript) through a constrained prompt that placeholders gaps rather than inventing content. The protocol is the downstream filter, not the only filter; run on the output of an unconstrained "write me a note" prompt, it will still work, but every pass will run longer and catch more, which is the tool telling you to fix your prompt.
Pass One: Factual Accuracy. Pass Two: MSE Integrity
Pass one, factual accuracy, roughly 60 to 90 seconds: every checkable fact in the draft, traced to your source. Numbers first, because numbers are where hallucination does the most damage: scores and their deltas (PHQ-9 16, not 19, not 15), counts (two missed workdays, two exposure trials, one panic episode), times and dates (the session's clock times, the date of service), dosages if you are a prescriber. Then quotes: verbatim against your shorthand, quotation marks intact, no model "improvements." Then names and identifiers: the right client, the right diagnosis code character by character (F33.1 is not F32.1), the right session number. The discipline is digit-level: read each number aloud against the source if you are tired, because transposition is precisely the error fatigue produces and fatigue misses. Anything unverifiable gets struck or flagged, not pardoned; in this protocol, a fact without a source is treated as false until you can source it from your own memory of the session, deliberately.
Pass two, MSE integrity, roughly 30 to 45 seconds: the mental status content stands or falls on one question, did I observe this? AI scribes reliably hallucinate the four MSE elements clinicians dictate least: appearance, affect range, insight, and judgment, because training-data notes always contain them and your shorthand frequently does not. Scan the draft's observational claims (appeared adequately groomed, affect congruent, insight fair, judgment intact) and sort each into three bins: traceable to my shorthand, keep; not captured but I did observe it and can add it now as my own deliberate documentation, keep with ownership; observed by nobody, strike. The third bin is non-negotiable, and it is where the transcriptionist analogy from the start of this chapter earns its keep: a transcriptionist who added physical-exam findings would have been fired the same day, and an MSE line nobody observed is the same offense in psychiatric clothing. If the resulting MSE is sparse, let it be sparse; a two-element MSE you observed beats a ten-element MSE the model decorated.
Pass Three: Risk Assessment Integrity. Pass Four: Intervention Specificity
Pass three, risk-assessment integrity, roughly 30 to 60 seconds, and this is the pass with the hard rule over it: AI never scores risk, never assigns a risk level, never makes the duty-to-protect determination; it may only format and transcribe what you determined. The pass asks three questions in order. One: does the draft claim a risk assessment occurred? If yes, did it, in the room, by you? A claimed screen that never happened is the gravest fabrication a note can carry, and it gets struck without ceremony. Two: if a screen happened, is it documented with its result and, ideally, the rationale that prompted it ("asked directly given elevated job-search stress; client denied SI"), because the rationale is what survives board discovery? Three: did anything risk-relevant happen in session that the draft omits? The model cannot know that the client's joke about "not being around for the holidays" registered in your clinical gut; if it registered and you assessed, the note must carry it, and if it registered and you did not assess, that is a clinical follow-up for tomorrow, not a sentence to backfill tonight. Notes about suicidality, abuse disclosures, IPV, and mandated-report territory get this pass run twice, slowly; everything else in the note can survive an error better than this section can.
Pass four, intervention specificity, roughly 30 to 45 seconds: find the sentence that says what you did, and test it against three requirements. Named, at technique level: cognitive restructuring with a thought record, interoceptive exposure via straw breathing, not "CBT techniques were utilized." Targeted: at what cognition, behavior, or skill ("targeting the automatic thought 'nothing I do matters'"). Actually performed: the generic-intervention hallucination, where the model narrates Socratic questioning and downward-arrow work you never did, is the most common fabrication in psychotherapy drafts because intervention language is the most templated text in the training data. The cross-check is your shorthand's intervention line; if the draft's intervention paragraph is longer and more decorated than what you captured, the decoration is the suspect. This pass is also where the payer earns or loses: the previous lessons established that skilled, named, targeted intervention is what reviewers reimburse, so pass four is simultaneously a fabrication check and a claim-strength check, which is why it gets its own pass instead of riding inside pass one.
Fluent text suppresses scrutiny, and a checklist run in fixed order is the only reliable countermeasure: the seven passes are your preflight, the note is the aircraft, and the signature is takeoff, after which errors stop being correctable.
Pass Five: Plan Quality. Pass Six: CPT Match. Pass Seven: Fabrication Sweep
Pass five, plan quality, roughly 20 to 30 seconds: the Plan must contain a next step you actually intend (next session date or interval, the planned focus), the homework actually assigned in the words you assigned it, any coordination or referral actually initiated, and a reason attached to continuation rather than "continue current treatment" standing alone. The model's characteristic plan failure is plausible inertia: it writes the plan that usually follows this kind of session, which may not be the plan you made. Thirty seconds against your shorthand's plan line settles it.
Pass six, CPT match, roughly 20 to 30 seconds, and it is pure clinician territory: does the code on the claim match the service the note now describes? A 90834 needs the note's time consistent with the 45-minute band; a 90837 needs 53-plus minutes documented from your calendar, with clock times, because the time is the first thing a high-frequency review checks; a 90853 group note needs the individualized content the format lessons established; an add-on 90833/90836/90838 for prescribers needs psychotherapy time and content documented distinctly from the E/M work. The model cannot verify any of this because the truth lives in your calendar and your room. Wrong-CPT drafting is one of the classic failure modes from earlier in this level, and pass six is its scheduled execution.
Pass seven, the fabrication sweep, roughly 30 to 60 seconds: a final read of the whole note asking one question of every sentence: where did this come from? Six sources are legitimate: your shorthand, your dictation, the instrument you administered, your calendar, the chart you are writing into, and your own present, deliberate memory of the session. Everything else is the model, and everything from the model that asserts a clinical fact gets struck. This pass also catches the residue the targeted passes miss: the invented therapeutic-alliance sentence, the unearned "client appeared more hopeful," the tense drift into present-tense client speech, the subtle inflation that turned "missed two days" into "ongoing inability to work." Pass seven is deliberately redundant with passes one through six, the same way a surgical count is deliberately redundant: the cost of the redundancy is thirty seconds, and the cost of the thing it catches is your license.
The Stopwatch, the Baseline, and the 4-to-7-Minute Target
Now the part most training skips: timing it, because an unsustainable protocol is an abandoned protocol, and an abandoned protocol is worse than none, since the clinician believes verification is happening. Add the pass budgets: 60-90 + 30-45 + 30-60 + 30-45 + 20-30 + 20-30 + 30-60 seconds comes to roughly three and a half to six minutes of pure verification on a standard outpatient note. Stack that on a sub-minute prompt run and the two minutes of shorthand you already spent at session time, and the whole documented encounter costs five to eight minutes in your first weeks. With repetition the passes compress, not because you check less but because your shorthand improves (fewer placeholders to resolve), your prompt stabilizes (fewer regenerations), and your eye learns each pass's failure signatures. Most clinicians land at 4 to 7 minutes per note, total, once the protocol is internalized. That is the published experience of this chapter's whole workflow, and it is the number that makes the economics undeniable: against a 15-to-20-minute cold draft, a clinician with Maria's eight-client day recovers roughly an hour and a half, every working day, with a chart that is stronger than the one the cold drafts produced.
Build the baseline empirically, not aspirationally. For ten consecutive notes, log four numbers: total minutes, minutes in verification, the pass that ran longest, and the count of strikes (fabrications or unverifiable claims removed). The log diagnoses your system. Verification over seven minutes with pass one longest means your shorthand is not capturing numbers cleanly; pass four longest means your intervention capture is thin and the model is decorating; a strike count that is not falling across ten notes means your prompt constraints are weak or your tool ignores them, and the tool, not your stamina, is the problem to fix. A strike count of zero from the start is its own flag: either your system is unusually clean or your passes are gliding, and you should seed a test, run one fabricated practice note through the protocol and confirm you catch the planted errors, the same way a TSA scanner is tested with planted contraband.
Set your personal target as a range, not a single number: standard outpatient note, 4 to 7 minutes; risk-relevant session, plus two minutes with pass three doubled; group session, the per-member protocol from the format lesson, budgeted per note, eight members meaning eight timed runs. Put the target where you will see it, on the comparison sheet from the previous lesson, and re-log a week of notes every quarter, because drift is silent: tools update, prompts rot, and the clinician who verified rigorously in March is skimming by August unless the stopwatch says otherwise.
When a Pass Fails: The Escalation Rules
A protocol needs failure rules, and this one has three. Rule one, the strike rule: any single fabricated clinical fact (a score, a screen, an observation, an intervention) is struck and the note proceeds; that is the protocol working. Rule two, the regenerate rule: three or more strikes in one note means the draft is not worth repairing; fix the input (the shorthand or the prompt) and regenerate, because line-editing a deeply fabricated draft risks missing the fourth fabrication, and the time math favors regeneration anyway. Rule three, the tool rule: a fabrication pattern that survives prompt repair across multiple notes, a tool that keeps inventing MSE elements after being instructed to placeholder, that fabricates session minutes, that hallucinates risk screens, is disqualifying, and the finding goes to whoever governs your tool stack: yourself if solo, your supervisor if you are an associate (Carmen's $59-a-month Upheal subscription is exactly the kind of tool a supervisor must be able to see this data about), the practice owner if you are one of Jordan's twenty-five clinicians, because a tool that fabricates against instructions is not a personal workflow problem, it is a practice risk.
And the rule above the rules, stated one final time because this chapter ends here: the clinician signs the note. The seven passes are not a ceremony performed to honor the signature; they are the factual basis that makes the signature true. When the board investigator, the RAC auditor, or the concurrent-review nurse reads your chart, none of them will ask what tool drafted it, and all of them will hold you to every word, because the attestation at the bottom says you did. The protocol is how that attestation stays honest at scale, eight clients a day, five days a week, at 9:54 PM and at 9:54 AM alike.
The Applied Problem: Your Seven-Pass Verification Protocol with Time Target
Your artifact, the capstone of this chapter, is a one-page Seven-Pass Verification Protocol card with your personal time target. Step one: write the seven passes in fixed order with their budgets and their one-line test: 1. Factual accuracy (60-90s): every number, quote, and code traced to source, digit level. 2. MSE integrity (30-45s): three bins, observed-and-captured, observed-and-added-deliberately, observed-by-nobody-struck; watch appearance, affect range, insight, judgment. 3. Risk integrity (30-60s): no claimed screen that did not happen, every performed screen documented with result and rationale, nothing risk-relevant omitted; AI never scores risk; run twice on risk-heavy notes. 4. Intervention specificity (30-45s): named, targeted, actually performed; decoration beyond the shorthand is the suspect. 5. Plan quality (20-30s): your actual next step, your actual homework, a reason attached. 6. CPT match (20-30s): code matches content and time, minutes and clock times clinician-supplied. 7. Fabrication sweep (30-60s): every sentence sourced to one of the six legitimate origins or struck.
Step two: add the escalation rules verbatim (one strike: strike and proceed; three strikes: regenerate from fixed input; persistent pattern: disqualify the tool and report up), and your dead-phrase blacklist from the comparison sheet as pass seven's quick-scan list. Step three: run the protocol live. Take one fabricated practice note into which you have deliberately planted five errors (a transposed score, an invented MSE element, an unperformed screen, a decorated intervention, a wrong CPT), run all seven passes with a stopwatch, and confirm you catch all five; if you catch four, find the fifth before you trust the card. Then run ten real notes, log the four baseline numbers, and write your personal target range on the card: most clinicians should expect to write "4-7 minutes" within a few weeks. Done looks like: a card you can execute from memory, a passed five-error seed test, a ten-note log, and a target range in your own handwriting. File it with the conversion workflow, the decision tree, the template bank, and the comparison sheet: five artifacts, one system, and a signature you can stand behind in any room that ever asks.
Key Takeaways
- Careful reading fails against AI drafts because fluent text suppresses scrutiny; the countermeasure is the safety-critical checklist: seven passes in fixed order, timed, like a preflight. The signature is takeoff, after which errors stop being correctable.
- The seven passes in order: factual accuracy (every number, quote, and code traced digit-level to source), MSE integrity (three bins; watch the four elements AI hallucinates most: appearance, affect range, insight, judgment), risk-assessment integrity (AI never scores risk; no claimed screen that did not happen; rationale documented; run twice on risk-heavy notes), intervention specificity (named, targeted, actually performed), plan quality (your actual next step with a reason), CPT match (minutes and clock times from your calendar, 53-plus for a 90837), and the fabrication sweep (every sentence sourced to one of six legitimate origins or struck).
- Pass seven's deliberate redundancy is the surgical count of documentation: thirty seconds that catches the residue the targeted passes miss, the unearned hopefulness sentence, the tense drift, the quiet inflation of impairment language.
- The stopwatch is part of the protocol, not an accessory: pure verification budgets to roughly three and a half to six minutes, the whole documented encounter to five to eight minutes at first, and most clinicians land at 4 to 7 minutes per note once internalized, recovering about ninety minutes a day against cold drafting on an eight-client caseload.
- The ten-note baseline log (total minutes, verification minutes, longest pass, strike count) diagnoses the system: a long pass one means weak number capture, a long pass four means the model is decorating thin intervention shorthand, a non-falling strike count means weak constraints or a non-compliant tool, and a zero strike count demands a seeded five-error test to prove the passes are not gliding.
- The escalation rules keep the protocol honest: one fabrication is struck and the note proceeds; three strikes means regenerate from corrected input rather than line-editing a deeply fabricated draft; a fabrication pattern that survives prompt repair disqualifies the tool and gets reported to whoever governs the stack, supervisor, practice owner, or yourself.
- The passes are not ceremony; they are the factual basis that makes the signature true. No auditor, reviewer, or board investigator will ask which tool drafted the note, and every one of them will hold the signing clinician to every word, which is the cardinal rule this entire chapter exists to operationalize.
Skill.re