Data Sensitivity and Classification
Learning Objectives
After completing this lecture, you will be able to:
- Understand the key concepts of data sensitivity and classification in a government context
- Connect data sensitivity and classification to your agency's AI initiatives
- Identify next steps for applying these concepts in your role
Key Topics Covered
- CUI, PII, PHI, classified vs
- What data categories exist and how AI changes handling requirements
- Government context for data sensitivity and classification
- Practical applications and next steps
Why This Matters for Government
Government agencies face unique challenges when it comes to AI adoption. This lecture addresses these challenges head-on by providing all government employees with the knowledge and frameworks needed to navigate AI in the public sector responsibly and effectively.
As part of the L1 (AI Aware) curriculum, this lecture builds on the foundational principle that every AI system in government ultimately serves citizens. Whether you are working with AI tools daily or setting strategy for your agency, understanding data sensitivity and classification is essential for responsible, effective government AI adoption.
Lecture URL: https://skill.re/learn/govt/data-sensitivity-and-classification.php
======================================================================
TRANSCRIPT: Data Sensitivity and Classification
======================================================================
What you will learn: Data classification frameworks—CUI, PII, PHI, classified vs. unclassified. What categories exist and how AI changes data handling requirements.
Welcome. If you work in government, you've almost certainly heard people talk about "sensitive data" or "classified information." But what do those terms really mean? And more importantly for our purposes: how does using AI change the way we need to handle different types of data?
In this lecture, we're going to walk through the landscape of data classification in government. We'll talk about what different categories of data mean, why they matter, and most critically, how the introduction of AI into your workflows changes the handling requirements for each category.
This might sound like a dry, bureaucratic topic. But it's not. Data sensitivity matters because it's the first line of defense against harm. If you don't know what data you have, how sensitive it is, and how to handle it properly, no amount of AI policy will save you from a breach, a scandal, or worse—harm to the people whose data you're responsible for.
WHY THIS MATTERS FOR GOVERNMENT
Data classification is foundational to government AI governance for several reasons:
- Legal requirements: Government agencies are subject to laws and regulations about data handling. Different types of data require different levels of protection. Using AI with data you're not supposed to use that way can violate law.
- Operational risk: A data breach involving sensitive personal information can cripple an agency's operations, damage public trust, and result in legal liability.
- Citizen privacy: Government collects information from citizens under the implicit promise that the data will be handled responsibly and securely. Using that data for purposes citizens don't expect, or failing to protect it, violates that trust.
- AI-specific risks: AI systems change how data is used and combined. A dataset that was safe to share in one context might become dangerous in an AI system where it's combined with other datasets and processed in new ways.
Understanding data sensitivity and classification is your first responsibility as someone working with AI in government.
CLASSIFICATION FRAMEWORK
Let's walk through the main data categories you'll encounter in government work.
Category 1: Unclassified/Public Data
Unclassified data is information that poses no known risk if released to the public. It's already public, or it could be public without harm.
Examples: Weather data, published economic statistics, agency press releases, general information about government programs (eligibility requirements, how to apply), publicly available organizational charts.
How AI changes the handling: Even public data can become sensitive when combined with other data in an AI system. For example, neighborhood median income is public data. Home addresses might be publicly available. But an AI system that combines them to predict property values might create privacy concerns. You need to think about not just whether individual datasets are public, but what emerges when you combine them in an AI system.
Handling requirements: Low risk. No special access controls needed. Can typically be shared with external partners, used in public-facing AI systems.
Category 2: CUI (Controlled Unclassified Information)
CUI is information that is not classified but that needs to be protected for reasons other than national security. It's sensitive, but not at the level of classified information.
The U.S. has official CUI markings and handling procedures. Other countries have equivalent categories (Official Sensitive, Internal Only, etc.).
Examples: Certain law enforcement records, pre-decisional deliberations, proprietary business information, certain personal information that doesn't rise to PII status, agency internal communications that could cause embarrassment or operational disruption if released.
How AI changes the handling: CUI often can't be used in AI systems without explicit safeguards. If you want to use CUI in an AI system, you need to:
- Anonymize or de-identify the data
- Limit access to the AI system itself
- Document who has access and why
- Ensure the AI system's outputs are handled as CUI
Handling requirements: Moderate to high risk. Access limited to authorized personnel. Storage in secure systems. Encryption during transmission. No external sharing without authorization. Clear documentation of who accessed what and when.
Category 3: PII (Personally Identifiable Information)
PII is information about an individual that can be used to identify that person. This is one of the most commonly regulated categories in government AI work.
Examples: Name, social security number, date of birth, driver's license number, passport number, email address, home address, phone number, biometric data (fingerprints, iris scans, facial recognition data), financial account numbers, medical information, criminal history, employment records.
How AI changes the handling: PII is often the most valuable data for AI systems—it's specific, personal, and predictive. But it's also the most tightly regulated. If you want to use PII in an AI system, you need to:
- Have explicit legal authority to use it for that purpose
- Notify individuals that their data is being used this way (usually)
- Implement strong security measures
- Have a plan for data retention and deletion
- Test the system for privacy leakage (can the system outputs reveal the original PII?)
- Consider consent requirements (does the person have to agree to this use?)
Handling requirements: High risk. PII should never be used in unapproved AI systems. It should be encrypted at rest and in transit. Access should be strictly limited to essential personnel. Never, ever upload PII to cloud AI systems (like public ChatGPT) unless you have explicit authorization and security review.
Category 4: PHI (Protected Health Information)
PHI is a subset of PII: medical and health information about individuals. In many countries, PHI is subject to specific privacy laws (HIPAA in the U.S., for example).
Examples: Medical diagnoses, treatment history, prescription information, mental health records, genetic information, health insurance information, immunization records.
How AI changes the handling: PHI is among the most sensitive data in government. If you work in health, veterans benefits, or similar agencies, you need to understand PHI handling thoroughly.
Using PHI in AI systems requires:
- Explicit legal authority under health privacy laws
- Special security requirements (often higher than for other PII)
- Individual consent in many contexts
- Audit trails documenting every access
- Strong de-identification if you want to reduce risk (de-identification must be done carefully so the data can't be re-identified)
Handling requirements: Highest risk. PHI should be used in AI systems only when absolutely necessary and only with explicit legal authority and individual consent. Encryption is mandatory. Audit logging is mandatory. Access is restricted to minimum necessary personnel. Never share PHI with external partners without ironclad legal agreements.
Category 5: Classified Information
Classified information is material related to national security. It's marked at different levels (Top Secret, Secret, Confidential in the U.S., or equivalent levels in other countries).
Examples: Intelligence information, military capabilities and tactics, diplomatic cables, nuclear weapons information, some cybersecurity information.
How AI changes the handling: Classified information should generally NOT be used in AI systems at all, and certainly not in commercial AI systems or cloud systems. If you work in an agency that handles classified material and you're thinking about using AI, you need to consult with your security and legal teams first.
The risk is extreme: classified information in an AI system could compromise national security.
Handling requirements: The strictest possible. Classified information can only be processed on classified systems with multiple layers of security. It should not be submitted to any AI system that isn't specifically approved for classified work and physically secured at the appropriate level.
DE-IDENTIFICATION AND ANONYMIZATION
One strategy for handling sensitive data in AI systems is to remove the information that makes the data "personally identifiable." This is called de-identification or anonymization.
The idea is: if you remove names, addresses, and other direct identifiers, the data is no longer personal data—it's just data about characteristics, behaviors, or outcomes.
But here's the critical nuance: de-identification is not a simple switch. It's a process, and it can fail.
True anonymization requires that:
- Direct identifiers are removed (name, SSN, etc.)
- The data cannot be re-identified by combining it with other data
- There's no residual risk of re-identification through pattern matching or linkage
That third requirement is hard. In the age of big data and AI, re-identification is increasingly possible. If your de-identified dataset includes age, gender, zip code, and medical condition, researchers can often re-identify individuals by linking to other public datasets that have the same combination of characteristics.
In government, you should treat de-identification with caution:
- Don't assume de-identified data is safe to use in any AI system
- Have a data scientist or security expert review your de-identification approach
- Test whether re-identification is possible by trying to link your de-identified data to other known datasets
- Document your de-identification process and the assumptions underlying it
- Plan for the possibility that your de-identification might fail—what's your fallback?
ANTI-PATTERNS / MISUSE RISKS
Here are common ways organizations mishandle data classification in AI work.
Anti-Pattern 1: Not Classifying Data at All
An organization starts building an AI system and collects data without explicitly thinking about sensitivity levels. "We need this data for the AI." They don't document what data they're using or how sensitive it is.
The risk: They end up using data they're not legally authorized to use, or mishandling sensitive data. When discovered, they face legal liability and have to shut down the system.
Anti-Pattern 2: Incorrectly De-identifying Data
An organization removes names from a dataset and assumes the data is now de-identified and safe to use anywhere. They don't realize that age, gender, zip code, and medical condition together can re-identify individuals in many cases.
The risk: The data is less anonymized than they think. Individuals can be re-identified. Privacy expectations are violated.
Anti-Pattern 3: Uploading PII to Unauthorized Systems
An employee is working on an AI project. They need to test the system. They copy a dataset containing PII and upload it to a cloud AI service (like public ChatGPT) to experiment. This violates every security and privacy policy.
The risk: The data is now outside government control. It can be accessed by foreign entities, used to train other models, leaked. If PII is exposed, the agency faces legal liability, reputational damage, and the need to notify affected individuals.
Anti-Pattern 4: Not Updating Classification When AI Involves New Data Combinations
An organization has two datasets, both classified as low-sensitivity. They combine them in an AI system. But the combination creates new insights that are more sensitive than either dataset alone.
The risk: They treat the output as low-sensitivity when it should be high-sensitivity. They share it with people who shouldn't see it.
PRACTICE / REFLECTION PROMPTS
- Think about data your agency collects or processes. Pick one dataset. What sensitivity level would you assign to it? Why? Are there any parts of it that are more sensitive than others?
- If your agency wanted to use that data in an AI system, what would need to be true for that to be safe? What legal authority would be needed? What security measures?
- Have you ever been surprised by how sensitive data could become when combined with other information? Describe the situation. What did you learn about data combination?
- In your current role, what data do you handle? Are the classification and handling procedures clear? If someone asked you "Is this data safe to use in an AI system?", what would you tell them?
KEY TAKEAWAYS
- Data classification is not optional bureaucracy—it's the foundation of responsible AI. You can't govern AI appropriately if you don't know what data you're using and how sensitive it is.
- Different data categories require different handling approaches. Unclassified public data can be used broadly. PII requires strict controls. Classified information requires special systems.
- AI changes data handling requirements. Even safe individual datasets can become risky when combined in an AI system. De-identified data can be re-identified when linked to other datasets.
- De-identification is not a simple fix. It's a process that requires expertise and testing. Don't assume that removing names makes data safe.
- PII and PHI should never be uploaded to unauthorized cloud AI systems. This is not a gray area—it's a clear violation of security and privacy standards.
- Document your data classifications and your reasoning. If something goes wrong later, documentation is your protection.
- When in doubt about whether data is safe to use in an AI system, ask. Talk to your security team, privacy office, or legal counsel. It's better to ask before you cause a problem than to find out after.
TERMS / GLOSSARY ITEMS
CUI (Controlled Unclassified Information): Unclassified information that requires protection for reasons other than national security (law enforcement sensitivity, pre-decisional information, etc.).
PII (Personally Identifiable Information): Information that can be used to identify an individual (name, SSN, address, etc.).
PHI (Protected Health Information): Medical or health-related information about an individual, subject to special privacy protections.
De-identification: The process of removing direct identifiers from data to make it non-personally identifiable.
Anonymization: The process of removing all information that could identify an individual, either directly or through combination with other data.
Re-identification: The process of determining the identity of individuals in supposedly de-identified data through linkage to other datasets or pattern matching.
Data Classification: The process of assigning sensitivity levels to data based on the harm that would result from unauthorized disclosure.
Imagine your agency wants to use AI to improve services delivery. You want to predict which eligible citizens are at risk of not applying for benefits they qualify for, so you can reach out to help them.
This is a good use of AI—it could help people get services they're entitled to. But what data do you need?
You'd need: demographic information (age, location, family size), employment information, previous benefit application history, maybe economic indicators for their area.
Most of this is sensitive—either PII or CUI. Here's how you'd think through data sensitivity:
- What data do you need? Employment history, family size, address. All PII.
- Do you have legal authority to use this data for this purpose? Check your agency's organic statute (your founding law) and relevant regulations. Can you use benefit application data to predict unmet need? Usually yes—it's within your mission. Can you use it to train an AI? You need to check. Maybe you need to notify people first.
- How will you handle it? Store it encrypted. Limit access to the specific team building the model. Don't put it in a public cloud. Audit who accesses it. Have a data retention policy (delete it when the project ends). Test the model to make sure it doesn't reveal PII.
- How will you handle outputs? The model produces a list of people to contact. This is sensitive information—it reveals something about them (suspected need for benefits). Handle it as you would any sensitive benefit information.
- What if you want to test the model? Use test data, not real data with real PII. Or use real data but in a controlled, secure environment, not in public AI systems.
That's how you operationalize data sensitivity in AI work.
Take 10 minutes. Write down:
- Three datasets your agency uses
- The sensitivity level of each (unclassified/public, CUI, PII, PHI, classified)
- Your reasoning for that classification
- What you'd need to do differently if you wanted to use each one in an AI system
For any where you're uncertain about the classification, flag it for follow-up with your security or privacy office.
Data sensitivity and classification might seem like a compliance checkbox—something to worry about and then forget. But it's really foundational to trustworthy AI. If you can't keep data secure and handle it appropriately, no amount of algorithmic sophistication matters. The AI system is built on a foundation of sand.
As you work with AI, keep data sensitivity front and center. Classify your data. Handle it appropriately. Don't cut corners to move faster. A data breach or privacy violation will cost you far more time and credibility than the extra weeks spent implementing proper data governance.
Thank you for taking data responsibility seriously.
Government AI CLUB Certification Program
Level 1: AI Aware | Government AI Policy Landscape | Lecture 2.3
A GOVT.CLUB initiative.
<- 1.2.1 Government AI Policy Landscape 1.2.3 PII and AI: The Bright Red Lines ->
Start Your CLUB Certification
This lecture is part of L1: AI Aware—8 hours of comprehensive government AI training.
Explore CLUB Certification
Related Lectures
L1 1.2.1—Government AI Policy Landscape 20 min - Video + Reading
L1 1.2.3—PII and AI: The Bright Red Lines 15 min - Video + Scenarios
L1 1.2.4—The Blueprint for an AI Bill of Rights 20 min - Video + Reading
Frequently Asked Questions
What will I learn in Data Sensitivity and Classification?
In this 15 min video + checklist lecture, you will CUI, PII, PHI, classified vs. unclassified. What data categories exist and how AI changes handling requirements
What level is Data Sensitivity and Classification?
This is a Level 1 (AI Aware) lecture, part of Chapter 1.2 \u2014 Responsible AI Use. It is designed for all government employees.
How long is lecture 1.2.2?
Lecture 1.2.2 (Data Sensitivity and Classification) takes 15 min. It is delivered as a video + checklist format.
Do I need prerequisites for Data Sensitivity and Classification?
This lecture is part of L1 (AI Aware). Prerequisites: None.
What is the CLUB Certification?
CLUB (Community Leading Unified Benchmarks) is a maturity-based AI certification for government professionals with 5 levels (L1-L5), 215 lectures, and 25 chapters aligned with NIST AI RMF, OMB, and GAO frameworks.
Skill.re