โ†
AI for Social Work & Human Services
Aware ยท M10 ยท lesson 10 of 18 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Privacy and Sensitive Data
๐Ÿ“–
now learning

Privacy and Sensitive Data

15 min

It is 4:47 in the afternoon, and Marisela has just closed the door to her car in the parking lot of the housing authority. She has her phone in one hand and, on the seat beside her, a spiral notebook with three pages of handwritten observations from a home visit that went long. The family she visited is in a domestic violence situation. The mother has given Marisela the address of the safe house where she and the children will go if he comes back tonight. There is a child-abuse referral in the file, currently under investigation. There are behavioral-health (BH) notes from a contracted counselor. The mother has undisclosed immigration status, which she shared in confidence because she needed Marisela to understand why she cannot call law enforcement directly. Marisela is exhausted and running behind on three other case notes. Her supervisor has been asking the unit to try the new AI drafting tool the agency recently licensed. She opens the app, taps "New note," and pauses, finger hovering over the text field. She is about to paste everything in. This lesson is about what happens in the moment she stops and asks: what exactly am I about to send, and where is it going?

The Anatomy of Sensitive Casework Data

Before we can protect sensitive information, we have to understand what we are actually holding. Human-services casework concentrates some of the most sensitive personal information that any government or organization collects. It is worth being specific, because "sensitive data" in the abstract is easy to dismiss. In casework, it is not abstract at all.

Categories That Require the Highest Protection

Child-abuse and neglect reports and substantiation records. When someone makes a report to CPS (child protective services), the identity of the reporter is typically protected by state confidentiality statute. The contents of the report, the investigation record, and the substantiation or unsubstantiation determination are similarly protected. These records can destroy a family relationship, affect custody in a divorce proceeding, affect employment in positions working with children, and carry a stigma that can follow a person for decades. Every state has a statute governing who may access these records, under what conditions, and with what penalties for unauthorized disclosure. Pasting a child-abuse investigation note into a general-purpose AI tool is a disclosure of that information to a third party, and most state confidentiality statutes did not contemplate that disclosure when they were written.

Domestic violence addresses and safety plans. The address of a domestic violence shelter or safe house is perhaps the single most acutely dangerous piece of information in a casework file. Disclosure of a shelter address to an abusive partner can directly result in violence or death. Many states have address-confidentiality programs specifically designed to protect this information. A case note that includes the address of the place where a mother and her children are hiding must never enter any system whose data handling is not verifiably controlled and understood.

Behavioral-health notes. Mental health treatment records and substance use disorder (SUD) records carry both HIPAA (Health Insurance Portability and Accountability Act) protection and, for federally assisted treatment programs, the additional protection of 42 CFR Part 2, a federal confidentiality regulation that predates and is in some respects stricter than HIPAA. A caseworker who receives behavioral-health notes from a contracted counselor or a treatment program as part of a case file is handling information that the person sharing it had a legal and clinical expectation would stay within a defined circle of care. Those expectations do not evaporate because the information is now in a case management system.

Immigration status. Immigration status is not covered by HIPAA, but it is among the most sensitive information a person can disclose in a service context. For a person who is undocumented or has a complex immigration situation, disclosure of their status to an unauthorized party can have immediate and irreversible consequences: immigration enforcement contact, detention, deportation, family separation. A person who shares their immigration status with a caseworker does so in a position of extraordinary vulnerability and trust. The caseworker's obligation to protect that information is not only ethical; in many jurisdictions it is codified in policy and law.

Benefits enrollment and financial records. SNAP (Supplemental Nutrition Assistance Program), TANF (Temporary Assistance for Needy Families), Medicaid, and housing assistance enrollment records reveal a person's financial circumstances in detail. They also, in combination, create a profile that can be used for financial exploitation. An older adult whose benefits file includes their payment dates, amounts, bank account routing information (captured during direct-deposit setup), and home address is at risk if that information leaves a controlled system.

PII (personally identifiable information) and its combinations. Name, date of birth, Social Security number, and address individually can be sensitive. In combination, they become an identity-theft kit. A case file routinely contains all of them, plus medical diagnoses, behavioral-health history, employment status, and income. The combination is far more powerful and far more dangerous than any single element alone.

A case file is a portrait of a person at their most vulnerable. Every piece of that portrait carries the person's trust that the agency receiving it will protect it. That trust is a legal obligation and an ethical one, and it does not end when the information is handed to an AI tool.

What Happens When Data Enters an AI Tool

This is the section Marisela needed before she opened that app. The risks are concrete, and they flow from how modern AI tools actually work. A caseworker who understands them can make a reasoned decision; one who does not is making an uninformed one, which in this context is not acceptable.

Data Retention by Vendors

When a user sends text to a cloud-based AI tool, that text is transmitted to a server operated by the vendor. How long does it stay there? What is done with it? The answers vary enormously by vendor and by contract, and they are almost never obvious from the consumer interface. Many commercial AI products, in their default configurations, retain conversation history indefinitely or for extended periods. Some allow the vendor's safety and quality teams to review prompts. Some make prompt data available to model improvement processes. The default is often not "your data disappears after your session." The default is often "your data is stored, for purposes you may not have read in a terms-of-service agreement."

For a caseworker using a general-purpose AI tool (one not specifically licensed for their agency's use under a data-processing agreement), the terms of service they agreed to when they created a personal account are the governing document. Those terms typically include sweeping data-use permissions. They were not written for child-welfare data or domestic-violence addresses. They were written for consumer use cases, and they make room for the vendor to do many things with the text that flows through the system.

Training on Inputs

LLMs (large language models), the AI systems that power most AI writing and drafting tools, are trained on large corpora of text. After initial training, many vendors continue to fine-tune their models using new data, including, in many cases, the conversations users have with the model. If a caseworker sends a case note draft to an AI tool whose terms of service allow training on user inputs, the content of that note, including names, addresses, diagnoses, and immigration status, may become part of the model's training data.

The practical implication is not that the AI will "remember" and report that specific information to another user in recognizable form, though that is a risk in some architectures. The more common concern is that sensitive information enters a vendor's data infrastructure in a way that is hard to audit, hard to retract, and potentially persistent. If the person whose information was sent later invokes a right to have their data deleted, the agency may have no mechanism to ensure deletion from a vendor's training pipeline. That is a data-rights problem as well as a confidentiality one.

Vendor Access and the Subprocessor Chain

A cloud-based AI tool is rarely a single company. Typically it involves a vendor whose product the agency or the caseworker is using, that vendor's cloud infrastructure provider (often one of the major public cloud platforms), and potentially additional subprocessors for safety review, model inference, or data storage. Each link in that chain is a potential access point for sensitive data. An agency's HIPAA-compliant business associate agreement (BAA), if it has one, may cover the primary vendor but not every subprocessor. A data processing agreement negotiated with a vendor may have representations about the vendor's practices without equivalent representations from every party in the subprocessor chain.

This is not a hypothetical concern. Multiple documented incidents in healthcare and human services have involved sensitive data appearing in unexpected places because a subprocessor had access the primary vendor relationship had not made explicit. The complexity of modern SaaS (software as a service) infrastructure makes it genuinely difficult to know where data goes once it leaves an agency system. That difficulty is a reason for caution, not a reason for complacency.

Re-Identification Risk

A caseworker who tries to protect privacy by removing names before sending a note to an AI tool is doing something sensible, but it is not sufficient. Research in data privacy has consistently shown that even substantially anonymized records can be re-identified when enough auxiliary information is present. A note that says "34-year-old Latina mother, two children ages 4 and 6, currently receiving SNAP and TANF in ZIP code [redacted], referred to CPS for neglect on [month, year], behavioral-health history of anxiety and depression, immigration status undisclosed" contains enough information that a determined party with access to other databases could likely identify the person. The combination of demographics, geography, program enrollment, and history makes anonymization much harder than it appears.

This is why de-identification, the formal process of removing or abstracting information to a standard that has been tested against re-identification risk, is the standard being aimed for, not just "remove the name." And for the most sensitive categories of information, even de-identification is not sufficient for AI processing; the better practice is not to include the information at all unless the task genuinely requires it.

Caseworkers operate inside a complex web of confidentiality obligations, many of which they learned in professional training but which they may not have mapped to AI use specifically. This section does that mapping.

HIPAA and HIPAA-Adjacent Obligations

HIPAA applies to covered entities (health plans, healthcare clearinghouses, and healthcare providers) and their business associates. Many human-services agencies are not themselves HIPAA-covered entities, but they receive information from covered entities (hospital discharge summaries, behavioral-health treatment records, SUD records) as part of case management. The caseworker's obligation to treat that information appropriately derives from the confidentiality terms under which it was shared, the agency's data-use agreements with the covered entities that shared it, and in some cases state statutes that mirror or extend HIPAA protections.

When a caseworker sends a behavioral-health note received from a contracted therapist to an AI tool, the question is not only "does HIPAA directly apply to me?" It is: "Did I receive this information under terms that restrict how I can share it further? Does sending it to a third-party AI vendor constitute a disclosure that requires the person's authorization or the covered entity's consent? Does the vendor qualify as a business associate, and do we have a BAA?" For most general-purpose consumer AI tools, the answer to the last question is no, and the vendor has not agreed to HIPAA's requirements for handling protected health information (PHI).

42 CFR Part 2 is stricter than HIPAA in important ways. It governs records of patients in federally assisted substance use disorder programs and prohibits disclosure of those records without patient consent except in very narrow circumstances. A case file that includes SUD treatment records from a program that receives any federal funding is subject to 42 CFR Part 2, and the restrictions travel with the records. Sending those records to an AI tool is a disclosure, and it requires the patient's written consent or a specific exception. Most AI drafting use cases do not qualify for an exception.

State Confidentiality Statutes

Every state has confidentiality statutes governing the types of records generated in human services work. Child-welfare records are protected in all 50 states; the specific provisions vary. Domestic violence address-confidentiality programs exist in most states. Mental health records and juvenile records typically have their own statutory protections. Many states have enacted specific regulations about what data can leave state agency systems and under what conditions.

These statutes were written before AI tools existed. Their application to AI use is, in many cases, an open question that agencies and their legal counsel are working through right now. But the general principle is clear: a state confidentiality statute that prohibits the disclosure of child-welfare records to unauthorized parties does not have a carve-out for AI tools. Sending a child-welfare record to a vendor whose systems are not authorized to receive it is a disclosure that implicates the statute, regardless of the purpose.

Agency Policy and Data Governance

Most agencies have IT security policies and data governance frameworks that govern what data can be sent to external systems. In 2026, many of these policies are being updated to specifically address AI tools, but the update cycle is slow relative to the pace of AI adoption. A caseworker who uses an AI tool that is not approved under the agency's IT governance process may be violating agency policy whether or not the policy specifically names AI tools. The general principle that personally identifiable information (PII) and protected information should not be sent to unauthorized external systems almost always covers AI tools.

The practical implication for the caseworker: before using any AI tool for case-related work, confirm that it is on the agency's approved vendor list and that it has been evaluated under the agency's data-governance process. If it is not, the right next step is to raise the question with a supervisor, not to proceed on a best-guess assumption that it is probably fine.

The Obligations That Stay with the Agency

There is a temptation, when a vendor provides a tool, to assume the vendor has taken responsibility for compliance. This is not how data law works. The agency's obligations under HIPAA, state confidentiality statutes, and its own data-governance policies do not transfer to the vendor when the vendor processes the agency's data. The agency is still responsible for ensuring that any disclosure of information to a vendor is authorized, that the vendor is qualified to receive and protect it, and that the people whose information was shared would have access to redress if something went wrong. This is the principle sometimes called "obligations that stay with the agency," and it governs AI vendor relationships just as it governs any other data processing relationship.

The specific obligations that do not transfer include:

  • The duty to authorize disclosures. If a disclosure requires client consent or a statutory exception, the agency must obtain that consent or confirm the exception before sharing the data with any vendor, including an AI vendor. The vendor's willingness to process the data does not make the disclosure authorized.
  • The duty to ensure minimum-necessary sharing. The minimum-necessary standard under HIPAA and its analogs in state law requires that only the information necessary for the specific purpose be disclosed. Sending an entire case file to an AI tool when only one section is needed for the drafting task violates minimum-necessary, even if the vendor processes only the relevant section. The disclosure happened.
  • The duty to enter a qualified data-processing agreement. A BAA (business associate agreement) or data-processing agreement (DPA) is a legal contract in which the vendor agrees to protect the information according to defined standards. Without one, the agency has no contractual basis for ensuring the vendor handles the data appropriately, and no remedy when it does not. The absence of a BAA or DPA is not a technical oversight; it is a compliance failure.
  • The duty to log and audit AI-assisted work. If AI tools are used in case documentation, the agency's audit and record-management obligations apply to the AI-generated content just as they do to human-generated content. If a court or an advocate asks how a case note was prepared, the fact that AI was used is material, and the agency's documentation practices need to capture it.

Concrete Safeguards for Casework Settings

The previous sections have been about understanding the risk. This section is about what to actually do. These are the safeguards that distinguish responsible AI use in a casework setting from a privacy incident waiting to happen.

Vendor Due Diligence Before Any Use

Before using any AI tool for case-related work, the agency (not the individual caseworker) should conduct vendor due diligence that addresses the following questions:

  • Does the vendor offer a BAA or DPA that covers the categories of sensitive information the agency holds? If the agency handles PHI, a BAA is not optional; it is legally required before PHI can be shared with the vendor.
  • Does the vendor offer a "no-train" guarantee, meaning a contractual commitment that prompt inputs will not be used to train or fine-tune the vendor's models? Consumer-grade AI tools almost never do. Enterprise agreements often can be negotiated to include this.
  • What is the vendor's data retention policy for prompts and outputs? Is there a contractual retention limit (for example, 30 days) that matches the agency's data-governance requirements?
  • Who are the vendor's subprocessors, and are they subject to the same contractual protections as the primary vendor?
  • Is the vendor's infrastructure compliant with FedRAMP (Federal Risk and Authorization Management Program), SOC 2 (System and Organization Controls 2), or other security certifications relevant to government data?
  • Has the vendor specifically represented its compliance posture under HIPAA, 42 CFR Part 2, and the state confidentiality statutes that govern the agency's records?

This is the due diligence a procurement team should complete before the tool reaches any caseworker's desk. A caseworker who is being handed an AI tool by their agency should be able to ask: "Was this evaluated under our data-governance process?" and receive a yes or a no. If the answer is no, the responsible path is to flag that and wait.

De-Identification and Minimum-Necessary

When an approved AI tool is in use, the discipline of de-identification and minimum-necessary sharing reduces the risk to the floor it can actually reach.

De-identification means removing or abstracting information that could identify the person before it enters the AI tool. At a minimum, this means removing the person's name, Social Security number (SSN), date of birth, precise address, and phone number. For high-sensitivity categories (domestic violence safety plans, immigration status, SUD records), the standard should be higher: if the information is not required for the specific drafting task, it does not enter the prompt at all.

It is important to be realistic about what de-identification achieves. As noted earlier, even substantially de-identified records can sometimes be re-identified. De-identification reduces risk; it does not eliminate it. For the most sensitive categories, such as a domestic violence address or an immigration disclosure, the minimum-necessary principle should be applied strictly: that information should not be in the prompt regardless of whether the name has been removed. If the AI drafting task is "write a contact note summarizing the topics we discussed," the address of the safe house is not a topic that the AI needs in order to complete that task. Leave it out.

Minimum-necessary means sharing only what is required for the specific task. If the task is drafting a contact note summarizing a home visit, the prompt should include the visit observations. It should not include the full case history, prior investigation records, benefits enrollment details, and immigration notes that happen to be in the file. Paste only what is needed. The temptation to paste everything, because it is faster and the AI might produce a better draft with more context, is real, but it trades convenience for privacy, and the data belonging to a vulnerable person is not the caseworker's to trade.

No-Train Guarantees and Contractual Controls

A no-train guarantee is a contractual term in which the vendor commits that data submitted through a designated API or enterprise channel will not be used to train or improve the underlying model. For agencies using AI tools on an enterprise basis, negotiating this term is non-negotiable. Without it, the agency cannot represent to the people in its system that their information is not being used to improve a commercial product.

No-train guarantees are now standard in enterprise AI agreements from major vendors, and they are typically accessible by entering into an enterprise contract rather than using the consumer product. The gap between the consumer version of an AI tool (which often permits training on inputs) and the enterprise version (which contractually prohibits it) is significant. Agencies should ensure that any AI tool used for case-related work is provisioned under an enterprise agreement that includes this protection, not through individual caseworker accounts created under consumer terms of service.

Related contractual controls include:

  • Data residency clauses specifying where data is processed and stored, to ensure it does not leave jurisdictions covered by the agency's legal authority.
  • Deletion rights giving the agency the ability to request deletion of specific data from vendor systems, which is essential when a person exercises a data-deletion right or when a record must be expunged.
  • Audit rights giving the agency the ability to review vendor data-handling practices, relevant to demonstrating compliance to oversight bodies.
  • Breach notification requirements specifying the timeframe and form of notification in the event of a security incident involving agency data.

The Approved-Channel Discipline

One of the most practical safeguards for caseworkers is the approved-channel discipline: use only AI tools that have been approved and provisioned by the agency, and use them through the approved interface. This sounds obvious, but it runs directly against a natural human instinct. A caseworker who has seen a better AI tool in a personal context, or who has heard about a tool from a colleague, or who simply finds the agency's approved tool frustrating, has an understandable incentive to use what works. The problem is that the better personal tool likely has consumer terms of service, no BAA, no no-train guarantee, and no contractual controls. Using it with case data is a privacy incident, even if nothing visible goes wrong.

Supervisors and unit leaders play an important role here. When a caseworker is under caseload pressure and documentation is piling up, the temptation to use whatever tool is fastest is high. Building a culture where the approved-channel discipline is understood as protecting the families in the system, not as bureaucratic obstruction, is part of responsible AI leadership in a human-services unit. The alternative, where caseworkers routinely use unapproved tools because the approved path is too slow or the approved tools are not good enough, is a privacy-incident risk distributed across the whole caseload.

Heightened Protection for the Most Sensitive Categories

Even within an approved AI tool operating under a strong enterprise agreement, some categories of information should never enter the AI drafting workflow at all. These are:

  • Domestic violence addresses and safety plan details. A safe house address, a shelter address, or the location details of a safety plan should not be in any AI prompt, ever. The harm from disclosure is irreversible and potentially fatal. The task of drafting a contact note about a domestic violence case can be completed without including the address. If it cannot, the task should be completed without AI assistance.
  • 42 CFR Part 2 SUD records unless the agency has obtained patient consent and the vendor has specifically represented compliance with Part 2's requirements. The strictness of Part 2, which prohibits disclosure even to other treating providers without consent except in narrow circumstances, makes it ill-suited for general AI processing.
  • Identifying immigration-status disclosures. A client's immigration status, shared in confidence, should not be in an AI prompt without the client's knowledge and consent. Even in an approved enterprise tool, the sensitivity of this information is such that including it in a prompt goes beyond what the minimum-necessary standard permits in most drafting tasks.
  • Child-abuse report source identity. The identity of a person who made a child-abuse report is protected by statute in virtually every state. A case note that includes identifying information about a reporter should not enter an AI drafting workflow. The reporter's identity can be omitted without compromising the note's usefulness.

What Marisela Does Next

Let us return to the parking lot. Marisela is looking at her phone. She is tired and running behind. The AI tool on her phone is a general-purpose consumer product she heard about from another caseworker. It is not the tool her agency approved.

She does not paste everything in. She knows, because her agency has trained on this, that the general-purpose tool on her phone does not have a BAA, does not have a no-train guarantee, and that using it with the family's data is a disclosure she is not authorized to make. She also knows that her agency has an approved AI drafting tool, provisioned through the agency's IT system, which has the contractual protections in place. She will use that one, tomorrow, when she is back at her desk.

Tonight, she drafts the contact note without AI: a few short paragraphs capturing the visit's observable facts, the current safety plan status, and the next steps. She leaves out the safe house address, because the address is not a fact that belongs in a case note at all: it is held in a separate, access-controlled location in the case management system, where only authorized staff can see it. She leaves out the immigration status, because it is not relevant to tonight's contact note. What she writes is an accurate, professional record of what happened during the visit.

It takes her twenty-five minutes. Tomorrow, with the approved tool and with de-identified notes that capture the observations without names or identifying details, she could draft the same note in eight minutes and spend the other seventeen minutes returning a call to a family that needs her. That is the genuine promise of AI for case documentation: time returned to human connection. But the promise only holds when the data disciplines that protect the people in the system are in place. Shortcutting them does not save time; it converts a documentation benefit into a privacy liability.

The agency's obligation, and the supervisor's obligation, is to make the approved path not just available but genuinely usable. If the approved tool is slow, cumbersome, or inaccessible from the field, caseworkers will route around it, and the privacy disciplines will be bypassed in the gap. Getting the approved path right, so that caseworkers can use it quickly and confidently, is part of what it means for an agency to deploy AI responsibly.

Key Takeaways

  • Human-services casework concentrates some of the most sensitive personal information that exists: child-abuse and neglect records, domestic violence addresses, behavioral-health (BH) notes, immigration status, and benefits enrollment data. Each category carries specific legal protections, and each becomes dramatically more sensitive when combined with others in a single case file.
  • When case data enters a cloud-based AI tool, it crosses the boundary of the agency's controlled environment and enters a vendor's infrastructure. The risks include data retention beyond the session, use of prompts in vendor training pipelines, vendor and subprocessor access to sensitive content, and re-identification of records that appear de-identified.
  • HIPAA (Health Insurance Portability and Accountability Act) and, for substance use disorder records, 42 CFR Part 2 impose legal requirements on how health-related information in a case file can be shared. Sending PHI (protected health information) to an AI vendor without a BAA (business associate agreement) is a HIPAA violation. 42 CFR Part 2 records require patient consent before disclosure in almost all circumstances.
  • State confidentiality statutes protect child-welfare records, domestic violence information, and other human-services data. These statutes do not have carve-outs for AI tools. Sending protected records to a vendor not authorized to receive them is a disclosure that implicates the statute.
  • The agency's confidentiality and data-protection obligations do not transfer to the AI vendor when the vendor processes agency data. The agency remains responsible for authorizing disclosures, applying the minimum-necessary standard, executing BAAs or DPAs, and logging AI-assisted work.
  • Effective safeguards include vendor due diligence (BAA, no-train guarantee, data retention limits, subprocessor controls), strict de-identification and minimum-necessary sharing in every prompt, and the approved-channel discipline: using only AI tools the agency has evaluated and provisioned under appropriate agreements.
  • Certain categories of information should never enter an AI prompt regardless of the vendor controls in place: domestic violence addresses and safety plan locations, 42 CFR Part 2 SUD records without consent, identifying immigration-status disclosures, and the identity of child-abuse report sources.
  • The genuine benefit of AI in casework documentation, time returned to human connection, depends on getting the privacy disciplines right. An agency that skips the disciplines in the name of speed does not save time; it converts a documentation benefit into a privacy liability and a potential harm to the very people the agency exists to protect.