โ†
AI for Social Work & Human Services
Proficient ยท M9 ยท lesson 9 of 18 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Grounding AI on the Case Record
๐Ÿ“–
now learning

Grounding AI on the Case Record

15 min

The two caseworkers used the same AI tool to draft the same kind of document, a court report background section, on the same Tuesday afternoon. The first worker opened a general chatbot, pasted in three of her recent case notes, and typed: "Write the background section of a dependency court report for the Alvarez family." The draft came back polished in ninety seconds, and it read beautifully. It also referenced a prior substance use evaluation that was never ordered, a 2024 missed visitation that never happened, and a county policy citation that did not exist. The second worker used a tool her agency had configured to retrieve only from the Alvarez case file in the state system and the current policy manual, with every sentence linked to a specific source record. Her draft was slower, plainer, and in two places it simply said the record contained no information on a point. It was also true. The difference between those two drafts was not the model. It was the same underlying model. The difference was grounding: whether the AI was allowed to answer from the open expanse of everything it had ever read, or was confined to the case record in front of it. In this field, that difference is the difference between a document you can file and a document that can separate a family.

What Grounding Actually Means

Grounding is the practice of confining an AI system to a specific, trusted set of source documents so that it answers from those documents rather than from its general training. The technical name for the most common grounding method is retrieval-augmented generation (RAG, a technique that retrieves relevant passages from a defined document set and places them in front of the model before it generates a single word). The plain-language version: instead of asking the model "what do you know about this," you build a workflow that says "here is the Alvarez case file and the current policy manual, and you may only use these."

To understand why this matters, recall what a large language model (LLM, the AI systems that generate text from prompts and documents) actually is. An LLM is not a database of facts that it looks up and reports. It is a next-token prediction system trained on billions of pages of text, and it generates the most statistically plausible continuation of whatever it is given. When you ask an ungrounded model about the Alvarez family, it has no Alvarez family in its memory. It has the general shape of what court reports about families look like, learned from its training. So it produces a court report shaped like the ones it learned from, populated with details that are plausible for such a report but were never true of this family. The fabricated substance use evaluation is not a malfunction. It is the model doing exactly what it was built to do: generating text that fits the pattern.

Grounding changes the inputs to that prediction process. When the relevant passages from the actual case file are placed directly in front of the model, the most statistically plausible continuation shifts toward what is actually in those passages. The model is still predicting tokens, but now it is predicting them against a backdrop that contains the real facts. This does not make hallucination impossible, a point this lesson will return to repeatedly, but it changes the odds dramatically. A grounded model drafting the Alvarez background has the real prior history in front of it and is far more likely to summarize that history than to invent a different one.

The contrast that defines this lesson is grounding against the case record versus grounding against the open web, or against nothing at all. An ungrounded model, or a model grounded on the open internet, answers a question about a family from everything it has ever read. A model grounded on the case record answers from the file, the policy manual, and the relevant statute, and from nothing else. For a caseworker carrying 22 families, each with a court date, a service plan, and a due-process clock running, that boundary is not a technical preference. It is the line between an AI that can mislead a judge and an AI that can save an hour of charting without manufacturing evidence.

Ungrounded, the model answers from everything it has ever read. Grounded on the case record, it answers from the file in front of it. In this field, that boundary is a due-process safeguard.

The Three Sources Worth Grounding On

Grounding is only as good as the document set you ground on. Putting garbage in front of the model produces grounded garbage. So before any caseworker or agency builds a grounded workflow, the question is precise: which sources, and which alone, should the AI be allowed to draw from? In human-services documentation there are three, and each one earns its place for a different reason.

The Case Record Itself

The first source is the case file in the case-management system: the CCWIS (Comprehensive Child Welfare Information System, the federally aligned record system many child-welfare agencies use), the FAMCare or Casebook instance, the eligibility file. This is the spine of grounding because almost every factual claim in a case document is a claim about this specific case: what was observed, what services were delivered, what the prior history is, what the current plan requires. When an AI drafts a court report background and grounds on the case record, the prior-history section is generated from the actual intake records, the actual service logs, the actual prior findings. The fabricated 2024 visitation cannot appear, because the model is drawing from a source where it does not exist.

Consider the hours this protects. A worker preparing a six-month review report for a family open since January has to assemble a history that spans dozens of contacts, three service referrals, two prior court hearings, and a change of placement. Ungrounded, that history is a hallucination risk on every line, and verifying it means tracing every claim back to the system by hand, which can take 45 minutes for a single complex report. Grounded on the case record, the draft is built from the real entries, and verification becomes a check that each sentence matches its linked source rather than a hunt for inventions that may or may not be there. The work shifts from "find the fabrications" to "confirm the citations," which is faster and far more reliable.

The Current Policy and Eligibility Manual

The second source is the agency's current policy manual and the governing regulations: the state child-welfare policy manual, the SNAP (Supplemental Nutrition Assistance Program, the federal food-assistance benefit) eligibility rules, the Medicaid and TANF (Temporary Assistance for Needy Families) policy as the state administers it. This source matters because a large share of human-services errors are not factual inventions about a family but misapplications of policy, and an ungrounded model is a notorious source of confident, wrong policy.

The reason grounding on the current manual is non-negotiable comes down to a timing problem. An LLM's training data has a cutoff. The policy it absorbed is the policy as it existed when the training data was collected, which may be 12 or 18 months before the model was deployed, and longer by the time a worker uses it. State benefit rules change every legislative session. Income thresholds are reindexed to the federal poverty level annually. A category of eligibility is expanded by an administrative action in March, and the ungrounded model, trained the prior year, will cheerfully apply the superseded rule. Ground the model on the current manual, the version the agency updated last month, and the retrieved passage is the current rule. The model drafts the eligibility rationale from the rule that actually governs, not the rule it half-remembers from training.

The due-process consequence of getting this wrong is direct. A SNAP determination drafted from a superseded income threshold can deny a family that qualifies under the current expanded rule. That family loses food assistance, may lack the time or language access to request a fair hearing, and may never learn that the denial rested on a rule that no longer applied. Grounding the determination support on the current manual closes that specific gap, because the model is reasoning from the rule that is actually in force.

The Relevant Statute and Court Standard

The third source is the controlling legal material: the dependency statute, the substantiation standard, the mandated-reporter threshold, the procedural rules that govern what a court report must contain and how a determination must be justified. This source matters most for the documents that go directly to a judge, because the legal framing of a case report carries weight that a factual summary does not, and a misstated standard can distort how a court reads the entire document.

An ungrounded model asked to frame a safety assessment against the legal threshold for removal will produce a confident-sounding statement of that threshold drawn from the general patterns of legal text it learned, which may blend standards from different states, cite a superseded version, or invent a threshold that sounds right. Grounded on the actual statute and the agency's procedural guidance, the model frames the assessment against the standard that actually applies in this jurisdiction. The worker still owns the legal judgment, always, but the draft is built on the real legal scaffolding rather than a plausible imitation of it.

Why Grounding Is Not a Cure

It would be a serious error to leave this lesson believing that grounding solves hallucination. It does not. It reduces it, sometimes dramatically, and it changes the character of the remaining errors, but a grounded system still requires the full verification discipline the program teaches. Understanding precisely how a grounded model still fails is what keeps a caseworker from the most dangerous mistake of all: trusting the output because the tool "uses the case record."

A grounded model fails in several specific ways. First, the retrieval step can pull the wrong passages. RAG works by finding the passages it judges most relevant to the prompt, and that judgment is itself imperfect. If the retrieval misses the intake record that documents a prior unsubstantiated report, the model drafts without it, and the omission can be as distorting as an invention. The history section that leaves out a prior finding misrepresents the case just as surely as one that adds a false finding.

Second, the model can hallucinate even with correct passages in front of it. Grounding makes the real facts available; it does not force the model to use only them. A grounded model drafting a court report can still generate a sentence that goes beyond what the retrieved passages support, blending a real prior service with an invented detail about its outcome, or stating a conclusion the source records do not actually establish. The grounding lowers the probability of this, but the next-token prediction architecture that produces hallucination is still running underneath.

Third, and most treacherously, a grounded tool can produce a citation that does not support the claim it is attached to. Citation features create a powerful illusion of verification: every sentence has a little source link, so the document looks checked. But the link points to a passage that the model selected as plausibly related, and the model can attach a citation to a passage that does not actually contain the claim. A worker who sees citations and assumes the document is therefore grounded and accurate has been given a more convincing version of the same failure, not protection from it. The citation has to be opened and read against the claim, every time, or it is decoration.

Consider a concrete case. A supervisor reviews a grounded, citation-enabled court report from a worker carrying 25 families. The report states the parent completed a substance use assessment, with a citation linking to a record in the system. The supervisor, trusting the grounding, signs it. At the hearing, the parent's attorney notes the assessment was scheduled but never completed; the linked record was the referral, not a completion. The grounded tool retrieved a real, relevant record and the model overstated what it showed. The citation was accurate as a pointer and wrong as evidence for the claim. Only a worker who opened the citation and read it against the sentence would have caught it.

A citation is a pointer, not a proof. Grounding tells you where the model looked; it does not tell you that what it wrote is what the source says.

Building a Grounded Workflow on the Record

For a practitioner at this level, the goal is not to understand grounding in the abstract but to run a grounded documentation workflow that returns hours and holds up to a court. That workflow has a small number of load-bearing parts, and each one should be deliberate.

Define the Document Set and Lock the Boundary

The first move is to define precisely what the AI may draw from and to make that boundary enforced rather than suggested. The document set for a given task is the specific case file, the current policy manual, and the relevant statute, and nothing else. The open web is excluded, because a model that can reach the open web can pull in a news article about a different family, an outdated policy summary, or a forum post, and present any of it in the confident voice of a case record. Other families' files are excluded, both for privacy and because cross-contamination between cases is a catastrophic failure in this field. The boundary is not "prefer the case record." It is "use only the case record, the current manual, and the statute."

For an agency, this is a procurement and configuration requirement, not a user preference. The tool must support a closed document set scoped to a single case, must not silently fall back to general knowledge when retrieval comes up empty, and must make clear when it is answering from the record versus when the record is silent. A worker should be able to tell the difference between "the record shows no prior CPS (child protective services) reports" and "I could not find prior CPS reports," because the first is a grounded finding and the second is a retrieval gap that requires a manual check.

Prompt the Model to Stay on the Record

Grounding the document set is necessary but not sufficient; the instruction to the model has to reinforce the boundary. A strong grounded prompt for case documentation tells the model, in effect: draft this section using only the retrieved case records and policy passages; for any point where the record contains no information, state that the record is silent rather than supplying a plausible answer; attach to each factual claim the specific source it came from; do not add observations, history, or policy that does not appear in the provided sources. This is the practice taught elsewhere in the program as constraining the model to the documented facts, applied specifically to a grounded retrieval setup.

The instruction to say "the record is silent" rather than fill the gap is the single most valuable line in a grounded prompt for this field. Hallucination thrives in gaps. A model that has learned court reports contain a prior-history section will generate one even when the retrieval returns nothing, because the pattern demands it. An explicit instruction to surface the gap rather than paper over it turns a hidden hallucination risk into a visible flag the worker can act on. A draft that says "the case record contains no documentation of services between March and June" is honest and verifiable; a draft that invents services to fill that window is a fabrication waiting to mislead a court.

Verify Against the Source, Not the Citation

The verification step in a grounded workflow is more efficient than in an ungrounded one, but it is not eliminated, and it is performed against the source, never against the citation alone. For each factual claim, the worker opens the linked source record and confirms the claim is what the source actually says. For observations, the source is the worker's own field notes inside the record. For policy, the source is the current manual passage, checked to confirm it both exists and applies to this situation. For history, the source is the specific intake, service, or finding record in the case-management system.

The time math here is what makes grounding worth the investment. In an ungrounded workflow, verification is a hunt: the worker must read every line suspecting invention, with no map of where the model might have gone off the record. A complex court report can take 45 minutes to verify this way, and the hunt is unreliable because tonal indistinguishability means the fabrications do not announce themselves. In a grounded workflow, each claim arrives with a source pointer, so verification becomes a structured pass: open the link, read the source, confirm the match, move on. The same report can be verified in perhaps half the time, and the verification is more complete because the worker is checking specific claims against specific sources rather than scanning for inventions. The hours saved are real, and they should go to the families and to the verification itself, not to absorbing more cases.

Log What the Model Was Grounded On

The final part of a grounded workflow is the audit trail, because in this field a workflow that cannot be explained to a court and an advocate is not finished. The log records what document set the AI was grounded on, what it drafted, what the worker verified, what the worker changed, and that the worker, not the model, made every consequential judgment. When an advocate later asks how a particular sentence entered the court report, the answer is a record: the sentence was drafted by a tool grounded on the case file, linked to this specific source record, verified by this worker on this date, and adopted into the report by a human who is accountable for it. That is a defensible answer. "The AI generated it from the case record" without a verification log is not.

Grounding as an Equity and Due-Process Safeguard

It is tempting to file grounding under "accuracy," a technical quality concern. In human services it is more than that. Grounding is an equity and due-process safeguard, and treating it as such changes how an agency deploys it.

Start with due process. The people in this system have the right to notice, to a fair hearing, and to challenge the record used against them. A challenge is only meaningful if the record can be traced to a real source. A court report built from an ungrounded model is a document whose claims may have no source at all, which makes a family's right to challenge it hollow, because there is nothing to examine behind a fabricated sentence. A grounded, logged document gives the family and their advocate something to test: this claim, this source, this verification. Grounding makes the record contestable, and a contestable record is what due process requires.

The equity dimension is subtler and just as important. Ungrounded models carry the biases of their training data, and when they fill gaps in a case with plausible-sounding inventions, those inventions are drawn from statistical patterns that can encode stereotype. A model invited to imagine the history of a family from a particular neighborhood or background can supply a history that reflects what its training data associated with such families rather than what is true of this family. The Allegheny Family Screening Tool debate and the failures of automated benefits systems like Michigan's MiDAS and the Dutch childcare-benefits scandal all share a root: a system reasoning from patterns rather than from the actual facts of the actual person, with the burden falling hardest on the people the patterns most malign. Grounding on the case record is a direct counter to this, because it forces the AI to reason from this family's documented reality rather than from the population-level patterns that can carry bias. It does not by itself make a system equitable, equity auditing remains a continuous practice, but it removes one of the most insidious channels through which an AI can import bias into an individual case.

Read together, these two points reframe the whole practice. Grounding on the case record is not primarily about making the model more convenient or even more accurate in a narrow sense. It is about keeping the AI tethered to the documented truth of the specific human being in front of the worker, so that the record stays contestable and the family's particular reality, not a statistical shadow of it, is what the document describes. That is why grounding belongs in the same sentence as due process and equity, and why an agency that deploys AI documentation without grounding has skipped a safeguard, not just an optimization.

Key Takeaways

  • Grounding confines an AI system to a specific, trusted set of source documents so it answers from those documents rather than from its general training. Retrieval-augmented generation (RAG) is the most common method, placing relevant passages from a defined document set in front of the model before it generates output.
  • An ungrounded large language model (LLM) answers a question about a family from everything it has ever read, generating plausible details that were never true of this case. A model grounded on the case record answers from the file, the current policy manual, and the relevant statute, and from nothing else.
  • The three sources worth grounding on are the case record itself (for facts about this case), the current policy and eligibility manual (because an LLM's training cutoff means it half-remembers superseded rules), and the controlling statute and court standard (because a misstated legal threshold distorts how a court reads the document).
  • Grounding reduces hallucination but does not cure it. Retrieval can miss the right passage, the model can still generate beyond what the passages support, and a citation can point to a source that does not actually contain the claim. A citation is a pointer, not a proof.
  • A grounded workflow has four load-bearing parts: define and lock the document set to the case record, current manual, and statute (excluding the open web and other families' files); prompt the model to use only the retrieved sources and to say the record is silent rather than fill a gap; verify each claim against the source record, not the citation alone; and log what the model was grounded on for a court-ready audit trail.
  • The instruction to say "the record is silent" rather than supply a plausible answer is the highest-value line in a grounded prompt, because hallucination thrives in gaps and this turns a hidden risk into a visible flag.
  • Grounding cuts verification time roughly in half by turning an unreliable hunt for inventions into a structured pass that confirms each claim against its linked source. The hours saved should go to families and to verification, not to absorbing more cases.
  • Grounding is a due-process and equity safeguard, not just an accuracy feature. It keeps the record contestable so a family can challenge it, and it tethers the AI to this family's documented reality rather than the population-level patterns that can carry bias.