What AI Does Well and Where It Fails
Learning Objectives
After completing this lecture, you will be able to:
- Understand the key concepts of what ai does well and where it fails in a government context
- Apply knowledge of summarization, classification, generation, translation
- Apply knowledge of limitations
- Analyze real-world case studies from government agencies
Key Topics Covered
- Capabilities: summarization, classification, generation, translation
- Limitations: reasoning, factuality, context, common sense
- Government context for what ai does well and where it fails
- Practical applications and next steps
Why This Matters for Government
Government agencies face unique challenges when it comes to AI adoption. This lecture addresses these challenges head-on by providing all government employees with the knowledge and frameworks needed to navigate AI in the public sector responsibly and effectively.
As part of the L1 (AI Aware) curriculum, this lecture builds on the foundational principle that every AI system in government ultimately serves citizens. Whether you are working with AI tools daily or setting strategy for your agency, understanding what ai does well and where it fails is essential for responsible, effective government AI adoption.
======================================================================
TRANSCRIPT: What AI Does Well and Where It Fails
======================================================================
What you will learn: What AI excels at (pattern matching, at scale, fast). Where it struggles (reasoning, factuality without sources, contextual judgment). How to match tasks to capabilities.
Now that you understand how AI works—pattern matching from training data—and what types exist, let's get practical: what can AI actually DO well, and what should you never ask it to do?
This is where many government initiatives fail. A well-intentioned agency adopts AI for the wrong problem. The system performs technically but doesn't solve the actual challenge because the task isn't suited to AI. Or worse: it automates something that should never be automated.
This lecture teaches you to evaluate tasks. By the end, you'll be able to look at any government process and say: "AI could help here, but not here. And here it would be wrong to use it at all."
Purpose
AI is powerful within its domain but useless or harmful outside it. Understanding which tasks are "AI problems" and which aren't is perhaps the most practical skill you'll develop in this course. It prevents wasted money and failed projects.
Why This Matters for Government
Government agencies often adopt technology because it's trendy, not because it solves a problem. We've seen it with "blockchains for government" and "IoT for everything." AI is the same. Some tasks are genuinely improved by AI. Many are not. Government has a duty to citizens to spend resources wisely. That duty starts with asking: "Is AI actually the right tool?"
Core Concepts
- What AI Does Brilliantly
Pattern Recognition at Scale
AI's superpower is finding patterns in huge datasets that humans could never analyze manually. Thousands of documents. Millions of transactions. Terabytes of sensor data.
A human auditor could review 100 financial documents per day. An AI system reviews 100,000. A human might catch irregularities in 5 of them. An AI system might catch 50. That's the win: speed and scale.
Example: An agency receives millions of small claims annually. They can't manually review each one for fraud. An AI system flags suspicious patterns (transactions at odd hours, unusual amounts, new vendors from high-risk regions) for human investigation. The system doesn't make decisions; it prioritizes human attention.
Classification and Categorization
Given a document, image, or transaction, AI can assign it to a category remarkably well. Not perfectly. But consistently and at scale.
Example: An agency receives thousands of emails daily from citizens. An AI system sorts them: "Benefits question," "Complaint about service," "Fraud report," "General inquiry." Humans then handle each category appropriately.
Summarization and Extraction
AI can read documents and pull out key information: names, dates, amounts, decisions, main points.
Example: An agency has 20 years of case files, mostly unstructured written notes. Lawyers need to know what happened in each case for legal defense. An AI system reads the notes and extracts: "Accused stole $50,000 from agency. Evidence includes bank records and witness statements. Case was settled for $30,000." Human lawyer reviews.
Translation
AI can translate between languages reasonably well, especially for high-resource language pairs (English-Spanish) better than for low-resource pairs.
Example: A federal agency serves immigrants who speak multiple languages. An AI system translates documents and interactions. Not perfect—context gets lost—but it enables service delivery that otherwise wouldn't happen.
Prediction (with Caveats)
AI can predict outcomes from historical patterns. Not because it understands cause-and-effect, but because it found correlations.
Example: A school district wants to identify students at risk of dropping out. An AI system trained on historical student data (attendance, grades, demographic factors) flags students likely to drop out. School counselors intervene. Caveat: the system's predictions are only as good as the historical patterns. If past data shows that low-income students dropped out more, the system will flag low-income students more. That correlation might be real (poverty affects education) or it might reflect past neglect of low-income students by the school. The system doesn't tell you which.
- Where AI Struggles
Reasoning and Causal Logic
AI doesn't reason. It doesn't understand cause-and-effect. It finds correlations.
This matters when the task requires thinking like "If we do X, then Y will follow" or "Because of Z, we should do W."
Example: An agency wants AI to help develop policy. "We want to reduce homelessness. What should we do?" An AI system trained on homelessness data might say: "People in cities with mild climates are less likely to be homeless." Should policy be "move homeless people to California"? Obviously not. The AI found a correlation (mild climate, less homelessness) but didn't reason about causation (does climate cause homelessness, or do people migrate to mild climates when homeless?). It also ignored causal pathways the data didn't show (housing prices, job availability, social networks).
Policy requires reasoning about causation. AI is bad at this.
Factuality Without External Sources
Language models especially can't reliably distinguish true from false. They're trained to generate plausible-sounding text. Plausibility and truthfulness aren't the same.
Example: An LLM is asked about a Supreme Court ruling. It confidently explains a ruling that doesn't exist. Or it misrepresents a real ruling. It sounds authoritative. A human reading it might believe it. But it's wrong.
This is especially problematic for government because official guidance needs to be factually correct. An agency can't publish an LLM-generated explanation of legal obligations without verifying it against the actual law.
Contextual Judgment
AI doesn't understand context. It can't adapt to "this is a special situation."
Example: A welfare agency has a rules-based system: "If income is below X, approve benefits." A person applies. Their income is above X. But they're suddenly unemployed and savings are depleting. "I know the rules say no, but in this case, it makes sense." A human can make that judgment. An AI system can't, unless explicitly programmed to.
Government decisions often require context. "Does this person deserve a waiver?" "Is this situation an exception?" "What's the right approach given these unusual circumstances?" These require human judgment.
Common Sense Reasoning
AI lacks human intuition and common sense. A human knows that "you can't put a living dog in the freezer to preserve it" even if no one explicitly told them. An AI system trained on data about food preservation might suggest it.
Example: An AI system optimizes a government schedule. It assigns an elderly caseworker to 8 office visits per day across the city. The system met its efficiency targets. But it didn't account for human limitations: the caseworker is exhausted by day 3 and service quality suffers.
Understanding Natural Language Nuance
AI struggles with sarcasm, irony, metaphor, and cultural context. Sarcasm especially trips up text AI systems.
Example: A citizen says "Great job on the service here—only three weeks to get an approval." Sarcasm. An AI sentiment analyzer might flag it as "positive" and mark the issue as resolved. A human would understand it as a complaint.
- The "Grey Zone": Tasks That Seem Like They Should Be AI But Aren't
High-Stakes Decisions Without Clear Rules
Tasks that sound algorithmic but require judgment.
Example: "An AI system will decide whether to approve disability benefits." Benefits determination requires assessing someone's capacity to work given their condition, age, job market, geographic location, etc. This isn't a classification problem. It's a judgment call. Frameworks exist (the Social Security Administration has detailed rules), but they're complex and contextual. An AI system could assist (extracting information, flagging cases that clearly meet or don't meet criteria) but shouldn't replace human judgment.
Tasks With Changing Definitions
If what "success" means changes over time, AI struggles because it's trained on historical patterns.
Example: "An AI system will detect fraud." Fraud evolves. Criminals adapt. Last year's fraud patterns are not this year's patterns. An AI system trained on 2023 fraud data will miss 2025 fraud techniques. It needs continuous retraining.
Tasks With Insufficient or Biased Historical Data
If you don't have good training data, don't use ML.
Example: "We want to predict which state employees will become leaders." You don't have decades of data on this (what makes someone a leader is changing). You don't have unbiased data (past promotions reflected the biases of past decision-makers). AI is likely to perpetuate those biases and not predict actual leadership potential.
Tasks Where Failure is Catastrophic
If the consequences of the AI being wrong are severe, human judgment should make the decision.
Example: "AI will recommend whether to approve new medications." This is life-or-death. AI can assist (analyzing clinical trial data) but shouldn't make the call.
Practical Use Cases
Good Use of AI: Mail Sorting (Post Office)
A postal service needs to sort millions of pieces of mail. OCR AI reads the address. A classifier assigns it to a region. This is: clear task, high volume, low individual stakes (if one piece is mislabeled, the system catches it eventually), high value (saves enormous labor).
Bad Use of AI: Eligibility Determination (Government Benefits)
Someone applies for benefits. An AI system determines eligibility. Why is this bad? High stakes (affects someone's income), requires judgment (special circumstances), political sensitivity (if a demographic group sees lower approval rates, there's a liability), no room for appeal ("the algorithm said no" isn't an explanation).
The right approach: AI flags clearly eligible cases and clearly ineligible cases for instant processing. Middle cases go to a human caseworker who can consider context.
Good Use of AI: Research Assistance (Policy Team)
A policy team is researching housing costs across different states. An AI system reads housing studies, extracts key statistics, identifies trends. A human then writes the analysis, interpreting what the numbers mean for policy.
Why is this good? The AI handles tedious work (reading hundreds of studies). Humans do the judgment (what does it mean? What should we recommend?).
Bad Use of AI: Policy Recommendations
"Use AI to determine the best housing policy." AI can't do this because it requires judgment, understanding causality, weighing tradeoffs, considering political feasibility, and ethical decisions. All of these require human reasoning.
Anti-Patterns / Misuse Risks
Anti-Pattern 1: Expecting AI to Reason About Cause-and-Effect
Risk: Using AI to answer "why" questions when it can only answer "what" questions.
What Goes Wrong: Policy makers ask AI: "Why is poverty increasing in this region?" AI analyzes data and says: "Areas with more industrial closures have more poverty." This is a correlation. But people interpret it as cause-and-effect and recommend "the government should support manufacturing." Maybe that's right. Maybe the actual problem is something else that correlates with both manufacturing and poverty. The AI didn't think; it just found patterns.
How to Avoid: Use AI for what/where/when questions (which regions, how many cases, what's trending). Use human analysis for why and what-should-we-do questions.
Anti-Pattern 2: Automating Decisions That Require Judgment
Risk: Removing humans from decisions that need context, fairness assessment, or exception-handling.
What Goes Wrong: A system automatically denies benefit applications where income exceeds threshold. A person is denied even though they have severe medical expenses and high debts. The system can't consider "fairness" because it doesn't understand the concept.
How to Avoid: Automate only decisions with clear rules and no exceptions. Use AI to assist with judgment calls, not replace them.
Anti-Pattern 3: Trusting AI Predictions Without Understanding Limitations
Risk: Assuming an AI prediction is more reliable than human judgment because it's "data-driven."
What Goes Wrong: An AI system predicts student dropout risk. A counselor sees that a student is flagged as high-risk. The counselor doesn't intervene because "the AI probably knows better." The student drops out. The system was based on historical data that might not apply to this student.
How to Avoid: Understand what the AI prediction is based on. Is it recent data? Is it unbiased? Compare AI predictions to human expert judgment. Use both.
Practice / Reflection Prompts
- Matching Tasks to Capabilities: Identify five tasks in your agency. For each, determine: Does AI's strength in pattern recognition help? Is human reasoning required? Are there exceptions that require judgment?
- Fact vs. Pattern: Find a news article about an AI system claiming to have "discovered" something. Does the article confuse correlation with causation? Write one sentence clarifying the distinction.
- Failure Analysis: Identify a government AI system. What would happen if it made a wrong decision? Is that acceptable? Would a human need to review?
- Grey Zone Analysis: Identify a task in your agency that might seem like a good AI candidate but actually requires human judgment. Explain why.
- Hybrid Approach Design: For a process you know, design a hybrid: what would AI do, and what would humans do? How would you divide responsibility?
Key Takeaways
- AI excels at pattern recognition at scale—processing volumes that humans can't, identifying subtle patterns in data, but only patterns that existed in training data.
- Classification and categorization are AI strengths—assigning documents to categories, routing requests, but the categories must be clear and non-contentious.
- Summarization and extraction work well when structure is consistent—pulling information from documents, but context and nuance get lost.
- Translation is possible but imperfect—useful as assistance but shouldn't replace human translation for sensitive documents.
- AI is terrible at reasoning about cause-and-effect—it finds correlations, which people naturally interpret as causation, leading to wrong policy conclusions.
- Large language models generate plausible-sounding text that might be completely false—never use them as authoritative sources; always verify.
- Contextual judgment, exceptions, and fairness assessment require human involvement—AI can't replace that, only assist.
Terms / Glossary Items
Pattern Recognition: AI's ability to find correlations and relationships in data that humans might miss, especially at scale.
Correlation: When two variables tend to occur together. Doesn't imply causation.
Causation: When one variable directly causes changes in another. Much harder to establish than correlation.
Contextual Judgment: Decision-making that requires understanding unique circumstances and adjusting rules accordingly.
Hallucination: When an AI system (especially language models) generates false information while sounding confident.
Human-in-the-Loop: A system where AI provides input but humans make final decisions.
Automation Bias: The tendency of humans to favor decisions made by automated systems, even when they should question them.
Task-Appropriate: Whether a tool is well-suited to the problem it's being used on.
You now know AI's real strengths and real limitations. You can look at a government task and assess: Is this a good AI problem?
The pattern is becoming clear: AI works best for high-volume, low-stakes assistance tasks. It's dangerous when used for high-stakes decisions, especially those requiring judgment or fairness assessment.
Next lectures zoom into specific government contexts: policy, tools, security threats, and fairness. Each assumes you now know what AI can and can't do.
Think of a major pain point in your agency. Something that takes a lot of staff time or causes bottlenecks. Is it something AI could help with?
- High volume? (Good for AI)
- Clear rules? (Good for AI)
- Requires judgment? (Bad for AI alone; good for AI assistance)
- People-focused? (Maybe; depends on the details)
Draft a proposal for how AI could help, and what humans would still need to do.
That's the reality check. AI is powerful. But it's not magic. It doesn't understand. It doesn't reason. It doesn't care about fairness.
Use it where it's strong. Don't use it where it's weak. Next: real government deployments. You'll see both good examples and cautionary tales.
Government AI CLUB Certification Program
Level 1: AI Aware | What AI Is and Is Not | Lecture 1.1.4
A GOVT.CLUB initiative.
<- 1.1.3 Types of AI Systems 1.1.5 AI in Government Today ->
Start Your CLUB Certification
This lecture is part of L1: AI Aware—8 hours of comprehensive government AI training.
Explore CLUB Certification
Related Lectures
L1 1.1.1—What AI Is and Is Not 20 min - Video + Reading
L1 1.1.2—How AI Actually Works 20 min - Video + Interactive
L1 1.1.3—Types of AI Systems 20 min - Video + Reading
Frequently Asked Questions
What will I learn in What AI Does Well and Where It Fails?
In this 20 min video + case studies lecture, you will Capabilities: summarization, classification, generation, translation. Limitations: reasoning, factuality, context, common sense
What level is What AI Does Well and Where It Fails?
This is a Level 1 (AI Aware) lecture, part of Chapter 1.1 \u2014 AI Foundations. It is designed for all government employees.
How long is lecture 1.1.4?
Lecture 1.1.4 (What AI Does Well and Where It Fails) takes 20 min. It is delivered as a video + case studies format.
Do I need prerequisites for What AI Does Well and Where It Fails?
This lecture is part of L1 (AI Aware). Prerequisites: None.
What is the CLUB Certification?
CLUB (Community Leading Unified Benchmarks) is a maturity-based AI certification for government professionals with 5 levels (L1-L5), 215 lectures, and 25 chapters aligned with NIST AI RMF, OMB, and GAO frameworks.
Skill.re