Recognizing Bad AI Output in a Case Note
The note looked finished. A child-welfare caseworker carrying twenty-six families had run her field notes from an afternoon home visit through the agency's AI documentation tool, and ninety seconds later a clean, complete case note sat on her screen, formatted the way the case-management system wanted it, written in the careful professional register she would have used herself on a good day with time to spare. She was tired. It was 6:40 in the evening, she had two more notes to write before she could go home, and the draft in front of her was, on first read, good. That first read is exactly where families get harmed. Because somewhere in a draft that is ninety-five percent accurate, the tool had written that the mother "acknowledged ongoing difficulty managing the children's bedtime routine," a small, plausible, characterizing sentence that the mother had never said and the worker had never written down. The skill this lesson teaches is the one that separates the worker who catches that sentence from the worker who files it: the trained, skeptical second read that treats every AI draft as guilty until verified against the record.
Why a Skeptic's Eye Is a Trained Skill, Not a Mood
Most caseworkers, handed an AI documentation tool for the first time, are told to "review the output carefully." That instruction is nearly useless, and it is worth understanding why. Careful reading, the kind you do when you proofread, is built to catch the errors that human writers make: typos, dropped words, a wrong date you happen to remember, a sentence that does not quite parse. AI-generated text contains almost none of those errors. Large language models (LLMs, the AI systems that generate text by predicting the most likely next word) produce prose that is grammatically clean, tonally consistent, and contextually plausible. The errors they introduce do not look like errors. They look like competent professional writing. So a worker applying ordinary careful-reading habits to an AI draft is using the wrong detector for the actual fault.
Recognizing bad AI output is therefore a distinct, learnable skill with its own technique, not a general attitude of carefulness. It has a specific object (the factual claim), a specific method (tracing each claim to an independent source), and a specific mindset (the assumption that a fluent, confident sentence carries no evidence of its own truth). A caseworker who develops this skill reads an AI draft the way a fact-checker reads a reporter's copy or an auditor reads a ledger: not for whether it reads well, but for whether each assertion can be backed. The draft reading well is not reassuring. It is the precise condition under which a fabricated claim hides.
Put the stakes in caseload terms. A unit of fifteen workers, each producing perhaps fifteen to twenty AI-assisted notes and reports a week, generates several hundred AI-drafted documents weekly that will enter case records, inform safety assessments, and travel to court. If even a small fraction carry an undetected fabricated observation, a misapplied policy rule, or an invented prior history, that is a steady stream of false statements entering legal records under workers' names. The skeptic's eye is the control that stands between that stream and the families on the other side of it. It is also the skill that protects the worker, because the name on the note is theirs, and "the AI wrote it" is not a defense in a court, a fair hearing, or a licensing review.
An AI draft reading well is not evidence that it is accurate. Fluency and truth are produced by two different processes, and the model only controls the first.
The Three Things You Are Always Hunting
A skeptic's read is not a vague suspicion of the whole document. It is a targeted search for three specific kinds of fault, the three failure modes that appear inside a case record. Knowing exactly what you are hunting makes the search faster and more reliable, because you stop scanning for "something wrong" and start checking three defined categories. Every claim in an AI-drafted case document falls into one of them.
Invented Observations: Did I See This?
The first and most dangerous category is the invented observation: a specific, concrete detail about what the worker saw, heard, or was told during a contact that was never part of the worker's actual notes or experience. The bruising in the gold-standard example, the bedtime sentence in this lesson's opener, a description of a "cluttered and unsanitary" kitchen, a child's "flat affect," a parent who "appeared agitated." These are the sentences that read most naturally, because the model has been trained on enormous volumes of real case notes and knows exactly what observations sound like. When your field notes are thin or ambiguous, the model fills the gap with the statistically likely continuation, and the result is indistinguishable in tone from the observations you actually made.
The test for an invented observation is a single question asked of every observational sentence: did I observe this, and is it in my raw field notes? Not "is this plausible," not "does this sound like the visit," not "could this have happened." Plausibility is exactly what the model manufactures. The only acceptable answer is a specific source: your own contemporaneous notes from the visit. If a sentence describes something you did not write down and do not independently and specifically recall, it is an addition, and it comes out. The case note documents what was observed, not what a model inferred was probably observed. A worker who carries a caseload of thirty families cannot reliably remember which of forty observations across a week's visits were real and which were the model's plausible filler. That is why the comparison is always draft-against-notes, never draft-against-memory.
Wrong Facts and Misapplied Policy: Says Who?
The second category covers factual claims and policy statements that are checkable against an authority outside the worker's own observation: dates, names, ages, dollar figures, eligibility rules, statutory thresholds, program criteria. Here the model's failure is not inventing an experience but generating a plausible-but-wrong fact or applying the wrong rule. A caseworker checking eligibility for the Supplemental Nutrition Assistance Program (SNAP, the federal food-assistance benefit) might get a confident determination that a family fails the gross-income test, when the family has a member receiving Supplemental Security Income (SSI) and is therefore categorically eligible under a separate provision that makes the gross-income test irrelevant. The output cites a regulation. It uses the correct program name. It is wrong, and a family loses food because of it.
The test for this category is the auditor's question: says who? Every checkable fact and every policy claim must trace to an independent, current source. For a date or a name, that source is the case-management record or the document itself. For a policy or eligibility rule, it is the actual state policy manual, the federal regulation, or the agency procedure as it currently stands, never the AI's own confirmation of the rule it just cited. The model that misapplied the policy will happily reaffirm the same wrong rule when asked. This matters most for rules that move: income thresholds reindexed to the federal poverty level each year, state benefit rules changed by a legislative session, eligibility criteria modified after the model's training data was collected. A model trained eighteen months ago may apply a threshold that no longer exists, and it will do so with a confident citation.
Fabricated History: Is It in the File?
The third category is fabricated history: a reference to a prior incident, prior service, prior contact, or prior finding that is not in the case record and did not happen. This failure mode is most likely when the model is given a partial record and asked to draft a section, such as the background or prior-history portion of a court report, that it "knows" from its training should exist. The model knows court reports usually contain a history section, so it produces one, and if the record it was given is thin, it backfills with plausible-sounding history: a "2024 TANF (Temporary Assistance for Needy Families, the cash-assistance program) sanction for noncompliance," a prior CPS (child protective services) report with its later unsubstantiation quietly omitted, a substance-use evaluation from a different, closed case cited as current.
Fabricated history is the most insidious of the three because it is the one a tired reviewer is most likely to skim. The recent observations get checked; the current service plan gets checked; the history section "looks familiar" and gets a pass. That skim is where the invented prior incident lives, and these are not neutral facts. "A family with prior CPS involvement," "a history of benefits noncompliance," "documented substance-use concerns" are characterizations that change how a judge reads the report, how a guardian ad litem advocates, and how the next worker approaches the family. The test is the file question: can I point to the specific record entry that supports this historical claim? A specific intake record for "a report was received in March 2024." A documented service record for "the family completed a parenting class." A document in the file for "an assessment was completed and returned negative." A claim with no traceable entry is removed.
The Second-Read Protocol
The skill becomes reliable when it becomes a protocol, the same sequence of moves every time, so that fatigue and time pressure do not quietly erode it. Here is the second read as a repeatable practice, the one the worker in the opener needed at 6:40 in the evening.
First, change your posture. The first read is for sense and flow, and you may have already done it. The second read is adversarial. You are no longer reading to understand the document; you are reading to find the sentence the model invented, on the working assumption that there is one until you have proven there is not. This posture shift is the whole game. A reviewer who expects the draft to be fine reads to confirm it; a reviewer who expects a fault reads to find it, and only the second one catches anything.
Second, read claim by claim, not paragraph by paragraph. Slow down to the level of the individual assertion. Each observational sentence: is it in my field notes? Each checkable fact: does it match the record or the document? Each policy statement: does it match the current manual? Each historical reference: can I point to the file entry? You are not reading prose; you are auditing a list of claims that happen to be arranged as prose. It helps to physically work with two windows open: the AI draft and the source (your field notes, the case-management system, the policy manual), moving your eyes between them for each claim rather than reading the draft straight through.
Third, treat unverifiable as unfilable. A claim you cannot trace to a source is not a claim to leave in "because it is probably right." It is a claim to remove or to mark for follow-up before the document is filed. The standard is not "I have no specific reason to doubt this." The standard is "I have a specific source that supports this." The burden of proof sits on the claim, not on your doubt, because the document is going into a legal record where the same standard applies.
Fourth, watch the seams. Hallucinations cluster where the model had the least to work from: thin field notes, gaps in the record, the transition between sections, the prior-history paragraph, the part of an eligibility determination where the rule gets complicated. When you know your field notes were sparse for a given visit, read that note's draft with extra suspicion, because that is exactly where the model did the most filling-in.
A Worked Second Read
Return to the opener. The worker has her ninety-second draft and her field notes. The adversarial second read takes her perhaps eight to ten minutes, against the twenty-five to thirty-five minutes the note would have taken to write from scratch, so the time dividend is real even with full verification. She goes claim by claim. "Worker arrived at 3:15 PM": in her notes, verified. "Two children present, ages 7 and 9": in her notes, verified. "Home was clean and adequately furnished": in her notes, verified. "Mother acknowledged ongoing difficulty managing the children's bedtime routine": she stops. She checks her field notes. Nothing about bedtime. She checks her memory: the mother said the mornings were hard because of the bus schedule, nothing about bedtime, nothing acknowledged as a difficulty she was struggling to manage. The model took "mornings are hard" and produced the more clinical, more characterizing "acknowledged ongoing difficulty managing the bedtime routine," a sentence that subtly recasts a logistics complaint as a parenting deficit. She deletes it and writes what was actually said. The eight minutes she spent reading adversarially is the difference between a record that documents a bus-schedule frustration and a record that documents a parenting concern that never existed.
Signals That Should Raise Your Suspicion
Because hallucinations are tonally indistinguishable from accurate output, there is no reliable visual tell that marks a sentence as false. That bears repeating, because workers naturally hope for a giveaway: a confident sentence is not more likely to be true or false than a hedged one, and a well-written passage is not safer than an awkward one. There is no shortcut around claim-by-claim verification. That said, certain patterns should raise your suspicion and make you check those passages with particular care, even though their absence proves nothing.
- Specificity you do not remember providing. A precise detail, an exact time, a specific clinical descriptor, a named diagnosis, where your input was vague. The model may have sharpened your ambiguity into a false precision.
- Characterizing language about a person. Sentences that assess rather than describe: "the mother was resistant," "the father appeared disengaged," "the child seemed fearful." Description traces to observation; characterization is where the model's training-data patterns are most likely to add a judgment you did not make.
- A complete prior-history section built from a thin record. If you gave the tool limited background and it returned a full, fluent history, treat every sentence in it as unverified until traced to a file entry.
- A clean policy citation on a complex rule. A confident regulatory cite, especially on eligibility, categorical exceptions, or anything with nested rules, is a flag to open the actual manual, not a sign the model got it right.
- Smoothness across a known gap. When you know information was missing and the draft reads as though nothing was missing, the model filled the gap. Find what it filled it with.
- Quotes and attributed statements. Any sentence that puts words in a client's mouth ("the mother stated that...") must match what was actually said. Invented or paraphrased quotes that harden a client's words into something more incriminating are a recurring failure.
None of these signals is a verdict. Their job is to tell you where to spend the extra thirty seconds, not to let you skip verification anywhere else. The discipline is to verify everything and to verify the flagged passages twice.
When to Stop, Rewrite, or Escalate
Recognizing bad output is not only about catching individual sentences. It is also about reading the draft as a whole and making a judgment about whether the tool helped or whether it is generating more risk than it is saving. A skilled reviewer knows three different responses, and choosing the right one is part of the skill.
Fix and file. The normal case: the draft is mostly grounded, you find and correct a small number of issues, and the verified document is accurate and faster to produce than writing from scratch. This is the workflow working as intended, and it is genuinely valuable. The point of this lesson is not to make you distrust AI documentation to the point of abandoning it; it is to make the fixing reliable so the value is real and the record stays clean.
Reject and rewrite. Sometimes the draft is so far from the record, with several fabricated observations or an entire invented history, that fixing it claim by claim would take longer than writing the note yourself, and worse, the act of patching a heavily fabricated draft risks leaving an invented sentence in by accident. When a draft is fundamentally unmoored from your notes, the right move is to discard it and write from your field notes directly. A draft that needs more correction than composition is not a time-saver; it is a trap.
Escalate. When you find a hallucination that already reached a filed record, a court report on its way to a hearing, or a pattern across multiple drafts that suggests the tool is systematically generating a particular kind of false content, that is not a quiet fix. It is a supervisor conversation and, where the agency has one, an incident or feedback channel. A fabricated observation that was filed may already be informing a decision, and a systematic failure mode in the tool affects every worker using it. The individual catch protects one document; the escalation protects the unit and the families it serves, and it creates the record that lets the agency hold its tool and its vendor to account.
The job is no longer to produce the draft. The job is to verify it, and to know when a draft is too compromised to verify and must be discarded or escalated.
Key Takeaways
- Recognizing bad AI output is a distinct, trained skill, not a general attitude of carefulness. Ordinary careful reading catches human errors (typos, wrong dates); it does not catch AI errors, which are grammatically clean, tonally consistent, and contextually plausible. The skill has its own object (the factual claim), method (tracing each claim to a source), and mindset (a fluent sentence carries no evidence of its own truth).
- You are hunting three specific faults: invented observations (did I see this, and is it in my field notes?), wrong facts and misapplied policy (says who, against an independent current source?), and fabricated history (is it in the file, traceable to a specific record entry?). Every claim falls into one of these categories.
- The comparison for observations is always draft-against-field-notes, never draft-against-memory. A worker carrying thirty families cannot reliably recall which observations across a week were real and which were the model's plausible filler.
- Verify policy and eligibility claims against the current policy manual, regulation, or procedure, never against the AI's confirmation of its own citation. Rules that change annually (income thresholds, state benefit rules) are the highest-risk category because the model may apply an outdated rule with full confidence.
- The second-read protocol is adversarial: assume there is an invented sentence until proven otherwise, read claim by claim with the source open beside the draft, treat any unverifiable claim as unfilable, and watch the seams where the model had the least to work from.
- Suspicion signals (unremembered specificity, characterizing language, a full history from a thin record, a clean cite on a complex rule, smoothness across a known gap, attributed quotes) tell you where to look harder. They never license skipping verification, because hallucinations have no reliable visual tell.
- Know the three responses: fix and file (the normal, valuable case), reject and rewrite (when patching a heavily fabricated draft is slower and riskier than writing from notes), and escalate (when a hallucination reached a filed record or a pattern suggests a systematic tool failure).
- The caseworker whose name is on the note owns every claim in it. "The AI wrote it" is not a defense in a court, a fair hearing, or a licensing review, which is why the skeptic's second read is a basic requirement of professional practice, not an optional enhancement.
Skill.re