System Prompts for Human-Services Contexts
The first time the unit tried AI for case notes, it worked beautifully for a week, and then it failed in a way nobody saw coming. Every worker had been trained the same way: paste your field notes, ask the tool to draft the note, verify before filing. The trouble was that each worker typed the instructions slightly differently every time. One worker wrote "make it professional," and the tool, helpfully, smoothed the rough edges of the visit into something cleaner and more conclusive than the facts supported. Another wrote "summarize what happened," and the tool, filling gaps the way models do, added an inference about the parent's mood that the worker had never observed. A third, in a hurry, just pasted the notes with no instruction at all, and got back a polished narrative that quietly invented a sentence about prior services. Three workers, three different sets of instructions typed from memory, three different failure modes. The verification step caught most of it. It did not catch all of it. And the supervisor, reviewing the misses, realized the problem was not the workers and not even really the tool. The problem was that the most important instructions, the rules that should never change, were being retyped from memory by a tired human a hundred times a week. This lesson is about fixing that with a system prompt: writing the constraints down once, correctly, so they travel with every single use instead of depending on whether someone remembered them at 6 p.m.
What a System Prompt Is, in Plain Terms
A large language model (LLM, the AI system that generates text from instructions and documents) receives two broad kinds of input. There is the thing you type each time, the specific request: here are my field notes, draft the home-visit note. And there is, in many tools, a separate standing instruction that sits above the conversation and shapes how the model behaves across every request. That standing instruction is the system prompt. Think of the difference between the question you ask a colleague today and the job description, training, and standing policies that govern how that colleague answers every question. The per-request message is the question. The system prompt is the standing instructions the model carries into every answer.
In practical terms, the system prompt is where you write down the rules that should never change: the model's role, the constraints it must always honor, what it is forbidden to do, and the format you need back. Some tools expose this directly as a "system prompt" or "custom instructions" field. Some agency-deployed tools have it configured by an administrator so it is the same for every worker. Some let an individual save it as a reusable template. The mechanism varies. The principle is constant: the load-bearing constraints belong somewhere durable, written once and applied every time, not retyped from memory in the per-request message where they will be inconsistent, incomplete, and forgotten under pressure.
Why does this matter so much in human services specifically? Because the constraints here are not stylistic preferences. They are the rules that keep an AI-drafted note from becoming a false legal record. "Only what was documented." "Do not infer or add observations." "AI drafts, the worker decides." These are the non-negotiables of the field, and the entire benefit of writing them into a system prompt is that the non-negotiables stop depending on a tired human's memory. The constraint that protects a family should not be one a worker has to remember to type. It should be there before the worker types anything.
The constraints that protect a family are too important to retype from memory a hundred times a week. Write them once, correctly, and let them travel with every use.
The Two Constraints That Must Always Travel
Two constraints are the heart of every human-services system prompt, the two the program returns to again and again because they are the two that protect people. The first is grounding: only what is documented. The second is the decision-aid boundary: AI drafts and organizes, the human decides. A system prompt that locks these two in has done most of its job. Everything else is refinement.
Constraint One: Only What Is Documented
The single most dangerous thing an AI documentation tool does is fill a gap with a plausible invention: an observation the worker never made, a mood the worker never noted, a service that never happened. The grounding constraint exists to forbid exactly that. Written into the system prompt, it instructs the model to draw only from the source material the worker provides, to never add observations, inferences, or characterizations that are not in that source, and to mark anything missing as missing rather than fill it in.
A weak instruction here invites the failure. "Summarize this home visit" tells the model to produce a summary, and a model producing a summary will, where the source is thin, generate the connective tissue that summaries usually have, including inferences the worker never made. A strong grounding instruction is specific and restrictive. It says, in effect: use only the facts present in the notes I provide; do not add any observation, characterization, diagnosis, or inference that is not explicitly in those notes; if a section of the note has no supporting information, write that the information was not documented rather than supplying it. The difference between those two instructions is the difference between a draft you have to scrub for inventions and a draft built to contain none. Neither removes the verification step, but the second gives verification far less to catch.
Consider the concrete payoff in a unit of fifteen caseworkers, each carrying twenty to thirty families, each drafting several notes a week. Without a grounding constraint in the system prompt, every worker's verification has to hunt for inventions the loose instruction invited, and the burden multiplies across hundreds of notes a week. With a strong grounding constraint applied uniformly, the drafts arrive cleaner, the inventions are rarer, and the verification time per note drops because there is less to catch. The constraint does not replace verification. It reduces the volume of errors verification has to find, across every worker, every time, which is exactly the kind of leverage a single well-written instruction can provide.
Constraint Two: The Decision-Aid Boundary
The second constraint encodes the field's cardinal rule, that AI informs and humans decide, directly into the model's standing instructions. The decision to substantiate a report, to recommend a removal, to deny benefits, to find that a safety threshold is met: these are human decisions bound by due process, and the model must never make them, recommend them as conclusions, or phrase its output as though it had. The system prompt instructs the model to draft, organize, and summarize, and to stop short of the determination.
This matters because a model left to its own helpful tendencies will reach for conclusions. Ask an unconstrained model to draft a safety assessment and it may, drawing on the patterns of the assessments it was trained on, produce a recommended disposition: "the home presents a safety risk requiring removal." That conclusory sentence is exactly what the worker, supervisor, and court must own, and it must not be pre-written by a text-prediction system. A decision-aid constraint in the system prompt instructs the model to present the documented facts and organize them, and to leave the determination blank for the human, never to characterize the overall situation as safe or unsafe, never to recommend a disposition, never to state a conclusion the human is responsible for. The model lays out what is documented. The human reads it and decides.
There is a subtle version of this failure worth naming, because it slips past easily. A model can respect the letter of the boundary, declining to write "I recommend removal," while violating its spirit through loaded framing: selecting and ordering the facts so that one conclusion feels inevitable, using characterizing adjectives that tilt the read, emphasizing some observations and burying others. The system prompt should address this too, instructing the model toward neutral, non-conclusory language and balanced presentation, so the human encounters the documented facts rather than a conclusion dressed as a summary.
The Anatomy of a Human-Services System Prompt
A complete system prompt for this work has a small number of parts, each doing a specific job. Writing it as named parts rather than a paragraph of good intentions makes it auditable: a supervisor can check that each part is present, and a worker can see what protection each part provides.
Role. State what the model is and, more importantly, what it is not. "You are a documentation assistant that drafts case notes and reports from the caseworker's provided notes and case record. You are not a decision-maker and you do not make eligibility, safety, or placement determinations." The role sentence sets the frame for everything below it.
Grounding rules. The "only what is documented" constraint, stated explicitly: draw only from provided source material, add no observation or inference not present in it, and flag missing information as missing rather than supplying it.
Decision-aid rules. The boundary, stated explicitly: draft and organize, never conclude, never recommend a disposition, never characterize the overall situation, use neutral non-conclusory language, leave determinations to the human.
Privacy rules. An instruction reflecting the field's perimeter: handle the personally identifiable information (PII, the data that identifies a specific person, such as names, dates of birth, and case numbers) and sensitive details only as the agency's policy permits, and do not generate or speculate about details of vulnerable individuals beyond what the source documents. (The agency's actual data-handling rules govern what may be entered into any tool in the first place; the prompt reinforces care, it does not replace policy.)
Output format. The structure the case-management system and the reader expect: the sections of a note in order, the headings, the level of detail, so the draft arrives in a shape the worker can verify and file rather than reformat.
Uncertainty and missing-information behavior. What the model does when the source is thin: state that information was not documented, do not guess, do not smooth over the gap. This part is what turns "the model invented a sentence to fill the silence" into "the model flagged the silence."
Here is how those parts read assembled, in the plain imperative language a model follows best:
You are a case-documentation assistant. You draft notes and reports from the material the caseworker provides. You are not a decision-maker. Use only the facts present in the provided notes and record. Do not add any observation, characterization, diagnosis, inference, or history that is not explicitly in the source. If a section has no supporting information, write "not documented" rather than supplying content. Do not state conclusions, recommend dispositions, or characterize the overall situation as safe, unsafe, eligible, or ineligible; those determinations belong to the caseworker, supervisor, and court. Use neutral, non-conclusory language and present the documented facts in a balanced way. Handle names and identifying details only as provided; do not invent or speculate about personal details. Return the draft in these sections, in order, leaving any determination field blank for the worker to complete.
Why the Prompt Does Not Replace Verification
A system prompt is a powerful control, and the temptation it creates is to treat it as a guarantee. It is not. The same structural truth from earlier in this program holds: an LLM generates statistically plausible text and cannot reliably distinguish what it accurately extracted from what it invented. A system prompt shapes the model's behavior and reduces the rate and severity of failures. It does not change the underlying architecture. A model told firmly to add nothing not in the source can still, on a given draft, add something not in the source, because following an instruction is itself a probabilistic behavior, not a hard constraint enforced by the machine.
So the relationship between the system prompt and verification is layered, not either-or. The prompt is the upstream control that reduces how many errors are produced. Verification is the downstream control that catches the errors that get through. Removing either one is a mistake. A unit that writes an excellent system prompt and then relaxes verification because "the prompt handles it" has fooled itself: the prompt made errors rarer, which makes the surviving errors harder to spot precisely because the reviewer expects fewer of them. The prompt and the verification checklist are a pair, and the next lessons in this chapter build the structured-output and verification-checklist halves of that pair deliberately.
This is also the honest answer to a worker who asks whether a good enough system prompt means they can finally trust the draft. The answer is that a good system prompt means the draft is cleaner and the verification is faster, not that the verification is optional. The professional and legal accountability for what gets filed under the worker's name does not transfer to the prompt any more than it transfers to the tool. "The system prompt told it not to" is no more a defense than "the AI wrote it." The worker still owns every word that enters the record.
Building, Testing, and Governing the Prompt
A system prompt is itself a small piece of safety-critical configuration, and it deserves to be treated like one rather than scribbled and forgotten. The discipline around the prompt is what makes it dependable across a unit and over time.
Write It Once, Centrally
The whole point is consistency, so the prompt should be authored once, ideally with supervisory and, where relevant, legal input, and made the default for everyone rather than left to each worker to compose. The unit in the opening story failed because the load-bearing instructions lived in fifteen separate memories. The fix is a single authored prompt, deployed uniformly: configured by an administrator in an agency tool, or distributed as a required template where workers paste it. When the constraints live in one authored place, they can be reviewed, improved, and trusted. When they live in fifteen memories, they are fifteen different prompts, and the protection is only as good as the most tired worker's recollection.
Test It Adversarially
Before a prompt is trusted, it should be tested against the failures it is meant to prevent, not just the happy path. Feed it deliberately thin notes and check whether it flags the gaps or fills them. Feed it a visit with an ambiguous detail and check whether it stays neutral or reaches for a conclusion. Feed it a case where the right output is "not documented" in several sections and confirm it writes that rather than inventing content. Testing a prompt is checking whether the constraints actually hold under the conditions that produce hallucinations, because a prompt that works on a clean, complete case but breaks on a thin one is a prompt that fails exactly when it matters most.
Version and Review It
The prompt will need to change: a new failure mode shows up, a policy updates, a model behaves differently after a vendor update. Treat the prompt as a versioned document with an owner, a change history, and periodic review, so a change is deliberate and traceable rather than an undocumented edit that quietly alters how every note in the unit is drafted. When a court or an advocate asks how the agency's AI-assisted documentation is governed, "here is the system prompt that constrains every draft, here is its version history, here is who reviews it" is a defensible answer. "Each worker types their own instructions" is not.
The Prompt as Part of the Audit Trail
Because the system prompt defines how AI participated in producing a record, it is part of the transparency the work owes a court and an advocate. An agency that can show the standing constraints under which every AI-assisted draft was generated has a far stronger account of its practice than one that cannot. The prompt is not only an operational control; it is evidence of the discipline, the written proof that the agency built the non-negotiables into the tool rather than hoping each worker remembered them. That is the through-line of this lesson: the system prompt turns the field's rules from things people must remember into things the tool always does, and that durability is exactly what makes the practice defensible.
Key Takeaways
- A system prompt is the standing instruction that sits above every request and shapes how a large language model (LLM) behaves across all of them, distinct from the per-request message you type each time. It is where the rules that should never change belong, written once and applied every time rather than retyped from memory under pressure.
- The opening failure is the case for the practice: when load-bearing constraints live in fifteen workers' memories, they become fifteen inconsistent prompts, and protection depends on the most tired worker's recollection. The fix is to write the constraints down once, correctly, so they travel with every use.
- Two constraints are the heart of any human-services system prompt: grounding ("only what is documented," forbidding invented observations, inferences, and history and flagging gaps as not documented) and the decision-aid boundary ("AI drafts and organizes, the human decides," forbidding conclusions, recommended dispositions, and conclusory characterizations).
- The decision-aid boundary has a subtle failure: a model can avoid an explicit recommendation while still tilting the read through loaded framing, selective emphasis, and characterizing adjectives. The prompt should require neutral, non-conclusory, balanced presentation.
- A complete prompt has named parts: role (what the model is and is not), grounding rules, decision-aid rules, privacy rules covering personally identifiable information (PII), output format matching the case-management system, and explicit missing-information behavior that flags gaps instead of filling them.
- A system prompt is an upstream control that reduces how many errors are produced; verification is the downstream control that catches what gets through. They are a pair. A unit that relaxes verification because "the prompt handles it" makes the surviving errors harder to catch, not safer.
- Accountability does not transfer to the prompt. "The system prompt told it not to" is no more a defense than "the AI wrote it." The worker owns every word that enters the record under their name.
- Treat the prompt as safety-critical configuration: author it once and centrally with supervisory and legal input, test it adversarially against thin and ambiguous inputs that produce hallucinations, and version, review, and govern it. A documented, versioned system prompt is part of the audit trail that makes AI-assisted documentation defensible to a court and an advocate.
Skill.re