Verifying Every Factual Claim
The court report was due at 4:00 PM, and it was 2:40 when the caseworker finished her last home visit and opened her laptop in the parking lot. She had eight families on her caseload that were active in dependency court, and this was the third report she had drafted that week using the agency's AI documentation tool. The draft was good. It was clean, it was organized, it read like something a careful professional had written over two unhurried hours. She skimmed it once, it sounded right, and her thumb hovered over the button that would attach it to the court filing. Then she stopped, because a supervisor had told her something during training that had stuck: a draft that reads well is not a draft that is true, and the only way to tell the difference is to check it line by line against the record. She pulled up her field notes and the case-management system, set a timer, and began the work that this lesson is entirely about. Twenty-two minutes later she had found two errors that nobody skimming would ever have caught, and one of them, left in place, could have changed how a judge saw a mother.
Why Verification Is the Job Now
For most of the history of this profession, the hard part of documentation was producing it. A caseworker carrying twenty to thirty families spends a large share of every day, often half or more, writing case notes, court reports, intake assessments, and eligibility records instead of being with the people they serve. That production burden is the single biggest driver of burnout and turnover in the field. AI documentation tools, the transcription and summarization tools that draft a clean case note from a home visit in minutes, change that equation. They take the production burden and shrink it. A note that took forty-five minutes to write can come back from the model in two.
But the burden does not disappear. It moves. When a model can draft a court report faster than you can read one, the scarce and load-bearing human task is no longer writing the draft. It is verifying it. The job shifted from "produce the document" to "verify the document," and that shift is the most important thing an AI-assisted caseworker has to internalize. The hours AI returns are not a gift to be spent on more cases or a faster clock. A meaningful share of them has to be spent on verification, because the document the model produced is a draft from a fluent, confident, and fundamentally unreliable source.
Here is the discipline stated plainly. Every factual claim in an AI-assisted document must be traced to an independent source before that document is filed. Not skimmed for sense. Not re-read for flow. Traced, claim by claim, to the place the claim came from: the worker's own field notes, the policy manual or regulation, the entry in the case-management system. A claim that cannot be traced to a source is removed. That is the entire practice, and the rest of this lesson is about how to do it well, why each piece matters, and what it costs in time and in consequence when it is skipped.
The model made producing the draft cheap. It did not make the draft true. Verification is the part of the work that AI cannot do for you, because the whole problem is that the model cannot tell its accurate sentences from its invented ones.
Why must a human do this, and not a second AI? Because the failure being guarded against is structural. A large language model (LLM, the AI system that generates text from prompts and documents) works by predicting the most statistically probable next word given everything it has seen. It is not a retrieval system that pulls verified facts from a database and checks them against the source. It generates fluent, plausible prose, and it cannot distinguish between text it accurately extracted from your notes and text it produced because that kind of text typically appears in documents like the one it is drafting. A second model asked to check the first one shares the same blind spot. The independent source is the record. The independent checker is a person who reads against the record.
The Three Claim Types and Where Each Is Verified
Verification is more reliable when it is targeted, because not every claim is verified the same way or against the same source. The previous lesson in this chapter taught the three ways an AI hallucinates inside a case record: invented observations, misapplied or wrong policy, and fabricated history. Those three failure modes map directly onto three claim types, and each claim type has its own independent source. A caseworker who can sort the sentences in a draft into these three buckets, and who knows where each bucket gets checked, verifies faster and misses less.
Observations, checked against field notes
An observation is a claim about what the worker saw, heard, or did during a contact: the condition of a home, the appearance or behavior of a child, what a parent said, whether a safety plan was reviewed. The independent source for an observation is the worker's own field notes from that contact, not the worker's general memory of it. This distinction is not pedantic. Memory is reconstructive and fallible, especially across a caseload of similar visits. A fabricated observation that is broadly consistent with the kind of visit you remember will slide right past a memory-based check. It will not slide past a comparison against what you actually wrote down in the moment.
Work it concretely. The draft says, "Worker observed that the kitchen was clean and adequately stocked, and the two children appeared well-nourished and engaged." You go to your field notes. The notes say "kitchen clean, fridge had food, kids fine." The "clean kitchen" checks. The "adequately stocked" is a reasonable rendering of "fridge had food," and you let it stand because it traces. But "well-nourished and engaged" is a clinical characterization your notes do not support; you wrote "kids fine," which is a general impression, not an assessment of nutrition. So you soften the draft to match what you actually observed, because "well-nourished" is a word a pediatrician earns and a case note should not borrow from a model. That is verification of an observation: each detail traced to the field note, anything that cannot be traced either removed or pulled back to what the notes support.
This is why good AI-assisted documentation starts before the visit, with field notes specific enough to verify against. If your raw notes are three words and the rest lives in memory, the verification step is compromised at the source. A worker who knows the model will draft from these notes and that the draft will be checked against them writes notes that can carry that weight: concrete, specific, and honest about what was and was not observed.
Policy, checked against the current source
A policy claim is any statement that applies a rule: an eligibility threshold, a substantiation standard, a regulatory citation, a program requirement. The independent source is the actual, current policy: the state policy manual, the federal regulation, the agency operating procedure, the statute. It is never the AI confirming the rule it just stated, because the model that misapplied the rule will generate a fluent confirmation of the same wrong rule.
Consider a SNAP (Supplemental Nutrition Assistance Program, the federal food-assistance benefit known as food stamps) determination. The draft says the household fails the gross income test under a cited regulation. You go to the current eligibility rules, not back to the model. You find that this household has a member receiving Supplemental Security Income (SSI), which makes the household categorically eligible, so the gross income test does not apply to them at all. The model cited a real regulation and applied it to a situation it does not govern. Verifying that the cited rule exists would not have caught this; the rule exists, it is just the wrong rule for this family. Policy verification means checking that the rule is current, that it applies to this situation, and that it was applied correctly, which is three checks, not one.
Currency matters because rules change. State benefit rules update through legislative sessions and administrative actions. Income thresholds reindex to the federal poverty level each year. A model trained on data from eighteen months ago will confidently generate the threshold that was correct then. The worker's obligation is to the rule in force today, regardless of when the model's training data was collected. In child welfare the same discipline applies to substantiation standards and mandated-reporter thresholds: verify against the statute or manual, not the model.
History, checked against the case record
A historical claim refers to a prior incident, prior service, prior contact, or prior finding. The independent source is the case-management system: the actual record of what happened in this case. Verifying history means opening the system alongside the draft and tracing each historical reference to a specific entry. "A CPS (child protective services) report was received in March 2024" must correspond to a specific intake record. "The family completed a parenting class last fall" must correspond to a documented service record. "A substance use assessment returned negative" must correspond to a specific document in the file. A claim that cannot be traced to an entry is removed.
History verification is the slowest of the three and the most often shortchanged, because a court report's history section spans months or years and it looks familiar on a skim. That familiarity is exactly the trap. The fabricated prior sanction, the invented earlier report, the service that was never actually delivered: these live in the section a tired worker skims because it reads like every history section they have ever written. The history section gets the most careful trace, not the least, precisely because it is the one your eyes want to glide over.
The Line-by-Line Method
Verification is a skill with a procedure, not a vague instruction to be careful. "Be careful" does not catch a hallucination, because the hallucination is well-constructed, grammatically clean, and contextually plausible. It does not look like an error. What catches it is a method applied the same way every time, even when the draft looks perfect, especially when the draft looks perfect.
The method is a claim-by-claim trace. Read the draft one sentence at a time. For each sentence, ask: is this a factual claim, and if so, which type? An observation, a policy statement, or a historical reference. Then go to the independent source for that type and confirm the claim is supported. If it is, move on. If it is not, the claim is removed or corrected to match the source. You do not keep a claim because it is plausible. You do not keep it because it is probably true. You do not keep it because removing it makes the document weaker. You keep it only if it traces.
Set up the workspace so the trace is fast: the AI draft on one side, your field notes and the case-management system on the other, the relevant policy source open in a third place. The friction you remove from getting to the source is friction removed from the verification, and friction is what makes a worker under a deadline skip the step. Open the sources first, then read the draft.
What does this cost? In the opening scene the worker set a timer and the trace took twenty-two minutes on a court report that the model had drafted in two. That ratio is the honest economics of AI-assisted documentation. The model saved her perhaps forty minutes of writing; she spent twenty-two of them verifying, and she still came out ahead on time and far ahead on accuracy. A worker who pockets all forty minutes and skips the trace has not saved time. They have shifted risk into the case record and called it efficiency. The trace is not overhead on the real work. The trace is the real work now.
What the Trace Found
Return to the two errors she caught. The first was a historical claim: the draft's background section stated the mother had "a prior CPS referral in 2023 that was screened in for investigation." The worker traced it to the case-management system and found a 2023 referral that had been screened out, not in, and never investigated. A one-word difference, "in" versus "out," that reframes a mother from someone with a prior investigated allegation to someone with a single referral that an agency declined to pursue. A judge reads those two histories differently. The model had generated the more typical phrasing, because referrals that get written about in court reports are more often the ones that were screened in. She corrected it to match the record.
The second was an observation. The draft said the mother "appeared overwhelmed and tearful when discussing the children's school attendance." Her field notes said the mother "got emotional talking about the kids missing school, said she's been trying." The model had compressed a complicated human moment into a clinical-sounding characterization, "overwhelmed," that carried a weight the worker had not intended and her notes did not support. The mother had been emotional and self-advocating, not overwhelmed. She rewrote the sentence to match her notes. Neither error was a typo. Neither would have been caught by a worker reading for sense. Both were caught by a worker reading against the source.
The Court-Record Standard
The bar for this verification is not "good enough for an internal note." It is the court-record standard: the document must be accurate enough to withstand a judge reading it as evidence, an attorney cross-examining the worker on it, and an advocate challenging it on behalf of the family. This standard is higher than ordinary professional carefulness because of what these documents do. A case note is a legal record. A court report informs decisions about removal, placement, reunification, and the termination of parental rights, decisions among the most consequential any government makes, bound by due process and the right to challenge. An eligibility determination decides whether a family eats or has shelter. The accuracy of the record is a due-process safeguard, and verification is how that safeguard is actually maintained.
The court-record standard has a hard consequence for testimony. Picture the cross-examination. An attorney asks the worker about a specific observation in her report, the kind of pointed question a thorough advocate asks. If the worker verified, she can say where the observation came from and point to the field note that supports it. If she did not, she may discover on the stand that the sentence with her signature under it came from the model and not from her notes, and that she cannot source it. That is a due-process problem for the family and a credibility problem for the worker, and it cannot be cured after the fact. The verification that would have prevented it had to happen before the report was filed.
This is also why "the AI wrote it" is not a defense. The worker who files an AI-assisted document is professionally and legally accountable for its contents regardless of what tool produced the first draft. A licensing review, an agency investigation, or a court does not transfer that accountability to a vendor. The signature on the document is a representation that its contents are accurate, and verification is how the worker earns the right to make that representation. The obligation to verify is not an enhancement to professional practice. It is professional practice.
Grounding Helps But Does Not Replace Verification
Some AI documentation tools include grounding, also called retrieval-augmented generation (RAG, a technique that connects the model to a specific document set before it generates output), and present each output statement with a citation to a source passage. These features genuinely reduce hallucination risk and they make verification faster, because the citation tells the worker where to look. They do not eliminate the need to verify. A grounded model can still produce a citation that links to a passage that does not actually support the claim, and it can still generate content in the gaps where the source is thin. The worker clicks through the citation and confirms the passage says what the model claims it says. Grounding turns a blind trace into a guided one. It does not turn verification into a feature you can buy and stop doing.
Building Verification Into the Workflow
Individual discipline is necessary and not sufficient. A unit of fifteen caseworkers, each carrying twenty to thirty families, each drafting notes and reports under constant deadline pressure, cannot be the agency's only defense against hallucination when the workers are tired, the caseload spikes, and ninety-five percent of the draft is correct. The fifth case at 4:30 on a Friday is where the trace gets skipped, and that is the case that ends up in front of a judge. Verification has to be built into the workflow as a required step, not left to the willpower of an exhausted worker.
That means a reusable verification checklist tied to the three claim types, so the trace is a procedure a worker follows rather than a habit they have to summon. It means supervisory review that treats an AI-assisted document the way it treats all documentation, as a draft to be checked, with the supervisor spot-checking that claims trace to sources rather than assuming the worker did it. It means clear written policy about which documents require verification before filing and what that verification must include. And it means honest math on capacity: if the documentation load is so heavy that the only way to keep up is to trust the draft and file it, the verification step will be skipped, and the agency has not reduced its documentation risk. It has moved the risk into a new and more dangerous form, because now the errors are fluent and confident and wear a professional's signature.
The time AI returns is the budget that makes verification possible. An agency that deploys AI documentation tools and then absorbs the saved time into higher caseloads has spent its own safety margin. The hours have to be protected for two things: verification, and the direct time with families that was the reason anyone took this job. A program that can show a judge or an advocate that every AI-assisted document was traced claim by claim to its source, that humans made every consequential decision, and that the saved hours went back to families, has built something defensible. A program that bought a tool and skipped the discipline has built a faster way to put errors into the record.
Key Takeaways
- AI documentation tools make producing a draft cheap but do not make it true, so the load-bearing human task has shifted from writing the document to verifying it. A meaningful share of the hours AI returns must be spent on verification.
- The discipline is a claim-by-claim trace: every factual claim in an AI-assisted document is traced to an independent source before filing, and any claim that cannot be traced is removed or corrected. Skimming for sense and re-reading for flow do not catch hallucinations, because the false sentences are fluent, confident, and contextually plausible.
- Verification is targeted by claim type. Observations are checked against the worker's own field notes (not memory); policy claims are checked against the current, applicable policy source (not the AI confirming itself); historical claims are checked against specific entries in the case-management system.
- The history section deserves the most careful trace, not the least. It looks familiar on a skim, and that familiarity is where fabricated prior incidents and never-delivered services hide.
- The standard is the court-record standard: accurate enough to withstand a judge, a cross-examining attorney, and an advocate challenging the record. A worker who did not verify may be unable to source a claim under oath, which harms the family's due process and the worker's credibility, and cannot be cured after filing.
- The worker who files an AI-assisted document is accountable for its contents. "The AI wrote it" is not a defense in a court, a licensing review, or an agency investigation. Verification is how the worker earns the right to put a signature on the document.
- Grounding and citation features (RAG) reduce hallucination risk and make verification faster by showing where to look, but they do not replace verification. A grounded model can still cite a passage that does not support the claim, so the worker still confirms the source.
- Verification must be built into the workflow with a reusable checklist, supervisory review that treats AI drafts as drafts to be checked, written policy, and protected capacity. If the load is too heavy to verify every draft, the agency has moved documentation risk into a more dangerous form, not reduced it.
Skill.re