Catching Hallucinations and Bias: The Quarterly Quality Protocol
Maria's pre-sign checklist worked beautifully for five months, and that was the problem. In month six her edit log went quiet: E0 after E0, almost nothing flagged, and she read the silence as proof the tool had matured. Then a colleague reviewing a shared client's chart asked why a note referenced a sister the client does not have. Maria had not gotten better at catching hallucinations; she had gotten slower, and nothing in her daily workflow could tell her which. Per-note checking cannot measure its own decay, and it cannot see bias at all, because bias does not live in any single note; it lives in the pattern across notes. By the end of this lesson you will have a quarterly quality protocol: a 10-note hallucination audit that quantifies your personal catch rate, a 5-formulation bias audit across race, language, and gender presentations, and, most importantly, a written personal stop-use threshold, the number below which you stop using the tool, decided now, while you are calm, instead of later, while you are rationalizing.
Why Daily Checking Is Not Quality Assurance
The pre-sign checklist from the last lesson is a control, and controls degrade. Every quality system in every safety-critical field assumes this: the inspector's eye dulls, the tool's behavior drifts after a vendor update, and the two failures hide each other perfectly. If your scribe starts hallucinating more and you start catching less, your edit log looks identical to a world where the scribe improved and your reviews stayed sharp. The daily workflow produces no signal that can distinguish those two worlds, and one of them ends in a fabricated sister in a signed chart.
This is why mature quality systems separate the control from the measurement of the control. A hospital does not just require hand hygiene; it sends an observer to count compliance, because the requirement and the rate are different facts. Your pre-sign checklist is the requirement. The quarterly protocol is the rate. Without it, "I verify every note" is a belief about yourself, and beliefs about one's own vigilance are exactly the kind of data clinicians are trained to distrust when a client offers them.
There is a second blind spot no per-note review can reach: bias drift. A hallucination is visible inside a single note if you look. Bias is a between-notes phenomenon. If your AI-assisted formulations describe your white clients as "perfectionistic and self-critical" and your Black clients with identical presentations as "irritable and noncompliant," no individual note looks wrong. The pattern only appears when you lay five formulations side by side and ask the question the NASW cultural competence standards and the APA Multicultural Guidelines require you to ask of your own work: is the framing tracking the clinical picture, or the demographics? AI did not invent this problem; biased language in human-written charts is well documented. But AI can industrialize it, reproducing a framing pattern across your entire caseload at a consistency no tired human could match.
The Controlling Analogy: Testing the Smoke Detector
Carry one picture through this lesson: the quarterly protocol is pressing the test button on the smoke detector. You do not press it because you smell smoke. You press it because a detector with a dead battery looks exactly like a working one, silently, for months, and the only time you discover the difference without testing is during the fire. Your verification habit is the detector. The quarterly audit is the button. The stop-use threshold is the rule, written on the wall in advance, about what you do when the test fails: replace the battery, not negotiate with it.
The analogy carries three lessons. First, the test is scheduled, not triggered. If you only audit after something scares you, you are measuring your luck, not your system; the calendar, not an incident, fires the protocol. Second, the test uses a known signal: you design it so the result is countable, which is why the protocol produces a percentage, not an impression. Third, the response is pre-committed. Nobody stands under a silent detector during a fire drill debating whether batteries are necessary. The value of a stop-use threshold is that you set it before the quarter when crossing it would be inconvenient, because that quarter will always be inconvenient. Maria's catch rate did not slip in a slow month; it slipped in the month she added two evening clients and needed the scribe most. That is not coincidence. It is the mechanism: dependence grows exactly when scrutiny shrinks.
Part One: The 10-Note Hallucination Audit
Once a quarter, you formally audit ten AI-drafted notes against ground truth. Here is the procedure, slowly. Select ten notes from the past quarter using a method that prevents cherry-picking: every Nth note from your log, or a random-number draw against your session list. Include at least two notes containing risk content, two from your longest sessions, and one from each AI tool you use, because hallucination rates differ by tool, session length, and clinical complexity. For each note, pull the ground truth you have: session memory if recent, your dictated recap, the verified transcript where one exists, the prior chart, and the measures actually administered.
Then read each signed note line by line against ground truth and classify every discrepancy into four tallied categories: fabricated facts (events, people, dates, calls that did not occur, the client's nonexistent sister); inferred interventions (modalities named that you did not use); invented or altered quotes (quotation marks around words the client did not say); and distorted risk or symptom language (softened, amplified, or templated clinical content, the most dangerous category because AI never scores risk and the clinician's determination must survive into the record intact). For each defect, note whether it appears in the AI draft only or in the signed final. A defect edited out is a catch. A defect that survived into the signed note is a miss.
Now compute the number this whole exercise exists to produce: your catch rate, catches divided by total defects found across the ten drafts, expressed as a percentage. If the ten drafts contained 14 defects and 12 were edited out before signature, your catch rate is 86 percent and you have two misses to remediate by addendum or correction in the chart, today, because a known false statement in a signed note does not get to wait for next quarter. Record the catch rate, the defect count per category, and the per-tool breakdown on the protocol sheet. One refinement worth adopting once the basic audit is habit: seed the test. Take one AI draft, deliberately plant a plausible error before your normal review, and see whether your everyday checklist catches it. A seeded error found is reassurance; a seeded error signed would have been a miss with a known denominator, which is the cleanest possible measurement of the detector battery.
Part Two: The 5-Formulation Bias Audit
The second audit examines not truth but framing. Once a quarter, pull five AI-assisted case formulations or conceptualization-bearing documents (treatment plan narratives, BPS summaries, formulation paragraphs in intakes) chosen to vary across race, language, and gender presentations in your caseload. If your caseload allows it, include at least one client who used interpretation services or for whom English is a second language, because language-mediated sessions are where AI summarization most often flattens nuance into pathology.
Read the five side by side and interrogate the language with specific questions. Attribution: when two clients miss homework, does one "struggle with follow-through due to overwhelming circumstances" while the other "demonstrates poor compliance"? Agency: who "is working hard in treatment" and who "was resistant"? Pathologizing of culture: is a client's deference to family framed as enmeshment, is guardedness in a client with documented reasons to distrust systems framed as paranoia, is expressive grief framed as lability? Strength asymmetry: which formulations mention resilience, resources, and protective factors, and which are wall-to-wall deficits? Diagnostic drift: are similar symptom pictures pulling toward different diagnostic framings along demographic lines, the pattern the disparities literature has long documented in human practice? You are not asking whether any sentence is false. You are asking whether the lens changes with the client's demographics while the clinical picture stays constant.
When you find biased framing, and across enough quarters you will, the remediation is concrete: rewrite the biased passage in the chart by addendum where it affects clinical meaning, log the pattern on the protocol sheet, and add a counter-instruction to your prompt block ("describe behavior concretely; do not characterize compliance, resistance, or motivation; note strengths and protective factors for every client"). Then re-test the same scenario next quarter to see whether the instruction held. Ground this practice where it professionally lives: the NASW cultural competence standards and the APA Multicultural Guidelines already obligate you to examine your own framing for cultural bias; the bias audit simply extends an existing ethical duty to a tool that now writes first drafts of your clinical voice. The duty did not change. The volume did.
Set the number that makes you stop before the quarter that makes you rationalize. A threshold chosen during the fire is not a threshold; it is a negotiation.
The Personal Stop-Use Threshold
Now the part most clinicians skip, and the part that makes this a protocol instead of a ritual: the pre-committed threshold. You will write, on the protocol sheet, the specific numerical conditions under which you stop using the tool, and you will write them this week, in a calm month, with no audit pending. A workable starter set, which you should adjust to your own risk tolerance and caseload acuity, looks like this. Stop-use, immediate: any fabricated risk or safety content in a signed note; any hallucinated quote that survived to signature in a risk-bearing note; catch rate below 80 percent in a quarterly audit. Probation, heightened review: catch rate between 80 and 90 percent; three or more fabricated facts per ten drafts from a single tool; any recurring bias pattern that survives one quarter of prompt correction. Continue: catch rate at or above 90 percent, no surviving risk-content defects, bias findings remediated and not recurring.
Understand precisely what "stop-use" means, because vagueness here is how thresholds die. It means you suspend the tool for drafting and return to manual or dictation-first documentation while you investigate: was the degradation in the tool (a model update, a new note template) or in you (caseload creep, batch reviewing, the 9:54 PM signing window returning)? You resume only after a defined re-entry test: run the tool in parallel for ten notes, full verification, re-measure. The threshold is yours, but it must be written, dated, and specific, because the quarter you cross it will be the quarter you are busiest, most dependent on the tool, and best at explaining to yourself why this quarter does not count. Maria's slipped quarter is the case study: catch rate 73 percent, discovered only because the audit was on the calendar, and her first instinct, she will tell you herself, was to re-score the borderline defects. The written threshold is what overruled her. That is its entire job: to be the version of you that decided before the pressure arrived.
One more boundary, stated plainly because this is the chapter's spine: nothing in this protocol delegates clinical judgment. AI never scores risk, never assigns a risk level, never makes the duty-to-protect or mandated-report call, and never decides what a formulation means. The audits measure how faithfully the tool records and frames what you decided, and how reliably you catch it when the tool fails. The clinician decides; the signature attests; the protocol measures whether the attestation is staying true.
A Story: The Quarter Maria's Catch Rate Slipped
Walk through Maria's audit, because the numbers teach better than the rules. Quarter one: ten notes, nine defects in the drafts (five inferred interventions, two fabricated minor facts, one altered quote, one softened risk phrase), eight edited out before signing. Catch rate: 89 percent. The one miss, a quote, she corrected by addendum and added the no-unverified-quotes instruction to her prompt block. Quarter two: eleven defects, ten caught, 91 percent, and the quote category dropped to zero. The system was working, and the protocol sheet showed it working, which mattered later.
Quarter three was the slip. She had added two evening clients, her review window had quietly migrated from the post-session gap to a 9:40 PM block, and her scribe vendor had shipped a mid-quarter model update that made drafts noticeably more fluent. The audit found fifteen defects across ten drafts, including the nonexistent sister, and only eleven caught: 73 percent, below her written 80 percent line, with one miss in a risk-adjacent note. Her first reaction was negotiation: two defects were arguably trivial; re-scoring them would lift her to 79.8, practically 80. She caught herself because the sheet's footer, in her own handwriting from a calmer month, said: "No re-scoring after the count. The threshold exists for this exact moment." She suspended drafting for three weeks, returned to dictation-first, corrected the misses by addendum, and investigated. The diagnosis was both failure modes at once: the vendor update had increased fabricated facts per draft, and her batch reviewing had dulled the catch. She rebuilt the post-session review gap, added a fabrication-specific prompt instruction, ran the ten-note re-entry test, and resumed at a measured 92 percent. The quarter cost her some evenings. A payer audit, a board response, or a client's trust would have cost far more, and the protocol sheet now documents not just the failure but the detection and the repair, which is exactly what a reviewer means by a functioning quality program.
Logging the Protocol and What the Numbers Buy You
Every quarterly run produces one page, and the page is the point. The protocol sheet records: the date and quarter; the ten note IDs and five formulation IDs sampled and the selection method; defects found by category and by tool; catch rate with the quarter-over-quarter trend; bias findings with the specific language patterns identified; remediations performed (addenda filed, prompt instructions added, tools placed on probation); the threshold status (continue, probation, stop-use) against your written numbers; and your signature with date. Fifteen minutes of writing after roughly ninety minutes of auditing, four times a year. Store it with your compliance documents, not your clinical charts.
What do six honest sheets buy you? Three things. Clinically, a trend line: you will know whether a tool update degraded your drafts within one quarter instead of one complaint. Professionally, alignment with the ethics infrastructure you already answer to: the NASW standards and APA Multicultural Guidelines on examining bias in your own practice, now applied to the tool that drafts in your voice. And defensively: when a payer, board, or plaintiff's attorney asks "what quality controls governed your AI use?", the answer is a dated stack of protocol sheets showing scheduled audits, quantified catch rates, a written stop-use threshold, and at least one quarter where you enforced it against your own convenience. A clinician who can produce that stack is not defending AI use; she is demonstrating supervision of it. The next lesson slots this sheet into the full four-artifact audit binder. This lesson's job is to make sure the sheet exists and the numbers on it are real.
The Applied Problem: Build Your Quarterly Quality Protocol Sheet
Your deliverable is the Quarterly Quality Protocol Sheet: one template page you will complete four times a year, with your personal stop-use threshold written into it before the first audit runs. Build it now, in three steps, and schedule the first audit before you close this lesson.
Step one, draft the template. Title it "Quarterly AI Quality Protocol, [your name], v1, [date]." Create five blocks. Block A, Hallucination Audit: fields for quarter, selection method, the ten note IDs, a four-row tally table (fabricated facts, inferred interventions, invented/altered quotes, distorted risk or symptom language) with columns for defects-in-draft, caught-before-signing, survived-to-signature, and tool; then a catch-rate line: catches divided by total defects, as a percentage, with last quarter's number beside it. Block B, Bias Audit: the five formulation IDs with the demographic spread sampled (race, language, gender presentation), the five interrogation lenses (attribution, agency, pathologizing of culture, strength asymmetry, diagnostic drift), findings, and rewrites filed. Block C, Remediation: addenda filed today for any surviving defect, prompt-block instructions added, tools placed on probation. Block D, Threshold: copy your written stop-use, probation, and continue conditions verbatim, then circle the status this quarter earned. Block E, signature and date, plus the footer line in your own handwriting: "No re-scoring after the count."
Step two, set the threshold before you have a score. Write your three tiers now: immediate stop-use conditions (any fabricated risk content signed; any surviving hallucinated quote in a risk-bearing note; catch rate below your floor, 80 percent is a defensible starting line), probation conditions (80 to 90 percent; three-plus fabrications per ten drafts from one tool; a bias pattern recurring after one quarter of correction), and continue conditions (90 percent or better, no surviving risk defects, bias findings remediated). Date the threshold. It cannot be edited mid-quarter; revisions happen only at quarter boundaries, in writing, with the old version retained.
Step three, run audit zero this week as a half-scale rehearsal: five notes instead of ten, three formulations instead of five, full procedure otherwise, ground-truth pull, four-category tally, catch-rate math, side-by-side bias read. Expect it to take about an hour and to be humbling; the rehearsal calibrates the template and your honesty before the numbers count. "Done" looks like: a finished template, a dated threshold you did not soften after seeing your rehearsal score, one calendar recurrence set for the first full quarterly run, and, if the rehearsal found any surviving defect, the addendum already filed. That sheet, completed quarterly, is artifact number three of the audit binder you will assemble in the next lesson.
Key Takeaways
- Per-note verification cannot measure its own decay: a degrading tool and a dulling reviewer produce the same quiet edit log. The quarterly protocol separates the control (your checklist) from the measurement of the control (your catch rate), the way every mature quality system does.
- The 10-note hallucination audit samples notes without cherry-picking, classifies defects into four categories (fabricated facts, inferred interventions, invented or altered quotes, distorted risk or symptom language), distinguishes catches from misses, and produces one number: your catch rate, with surviving defects corrected by addendum the same day.
- Bias is a between-notes phenomenon no single-note review can see. The 5-formulation audit reads conceptualizations side by side across race, language, and gender presentations, interrogating attribution, agency, pathologizing of culture, strength asymmetry, and diagnostic drift, and rewrites biased framing on the spot.
- This is an extension of an existing duty, not a new one: the NASW cultural competence standards and the APA Multicultural Guidelines already require you to examine your own framing for bias; AI simply industrialized the volume of framing you are responsible for.
- The personal stop-use threshold is written, dated, and numerical, set in a calm month: for example, immediate stop-use on any fabricated risk content signed or a catch rate below 80 percent, probation between 80 and 90, continue at 90-plus. It cannot be re-scored or renegotiated mid-quarter, because the quarter you cross it will be the quarter you most want to.
- Nothing in the protocol delegates judgment: AI never scores risk, never assigns a level, never makes the duty-to-protect or report call. The audits measure how faithfully the tool records what the clinician decided, and how reliably the clinician catches it when the tool fails; the signature remains a legal attestation.
- Each quarterly run yields one signed protocol sheet: sample, defect tallies, catch rate trend, bias findings, remediations, and threshold status. A dated stack of these sheets is what converts "I use AI carefully" into a demonstrable quality program when a payer, board, or attorney asks what controls governed your AI use.
Skill.re