Evaluating AI Outputs
Learning Objectives
After completing this lecture, you will be able to:
- Understand the key concepts of evaluating ai outputs in a government context
- Apply knowledge of validate facts, examine sources, review for bias, inspect logic, find errors, yield judgment
- Apply knowledge of hands-on fact-checking exercises
- Complete hands-on exercises that reinforce practical skills
Key Topics Covered
- The VERIFY method: Validate facts, Examine sources, Review for bias, Inspect logic, Find errors, Yield judgment
- Hands-on fact-checking exercises
- Government context for evaluating ai outputs
- Practical applications and next steps
Why This Matters for Government
Government agencies face unique challenges when it comes to AI adoption. This lecture addresses these challenges head-on by providing all government employees with the knowledge and frameworks needed to navigate AI in the public sector responsibly and effectively.
As part of the L1 (AI Aware) curriculum, this lecture builds on the foundational principle that every AI system in government ultimately serves citizens. Whether you are working with AI tools daily or setting strategy for your agency, understanding evaluating ai outputs is essential for responsible, effective government AI adoption.
Lecture URL: https://skill.re/learn/govt/evaluating-ai-outputs.php
======================================================================
TRANSCRIPT: Evaluating AI Outputs
======================================================================
What you will learn: The VERIFY method for evaluating AI outputs. Validate facts, Examine sources, Review for bias, Inspect logic, Find errors, Yield judgment.
You ask an AI system a question. It provides an answer. The answer sounds plausible. Should you trust it?
Not automatically. You need to evaluate the output critically. You need a systematic way to check whether what the AI produced is accurate, reasonable, and appropriate to rely on.
In this lecture, I'm going to teach you the VERIFY method. This is a systematic framework for evaluating what an AI system tells you.
WHY THIS MATTERS FOR GOVERNMENT
Government decisions affect people. If you base a decision on information from an AI system, and that information is wrong, people are harmed and you're liable.
Evaluating AI outputs critically is not optional. It's fundamental to responsible use.
THE VERIFY METHOD
Here's a framework for evaluating AI outputs. VERIFY stands for:
V - Validate Facts
Does the AI's output contain factual claims? Verify them.
If the AI says "The Affordable Care Act was passed in 2010," can you verify that? (Yes—it was 2010.)
If the AI says "Studies show that X intervention reduces Y outcome by 50%," can you verify that specific claim? (You'd need to read the studies.)
For important facts, verify. For routine facts, you might spot-check samples.
E - Examine Sources
Did the AI cite sources? Are those sources credible?
Good: "Research by [name of researcher], published in [journal], found..."
Bad: "Studies show..." with no specific citation
If the AI provides citations, follow them up. Read the actual source. Check whether the AI characterized it correctly.
R - Review for Bias
Is the AI's output biased in any direction?
Loaded language? Presenting one perspective as obviously true without acknowledging alternatives?
Ask yourself: What perspective is this output reflecting? Are there other legitimate perspectives not represented?
I - Inspect Logic
Does the reasoning make sense?
Follow the logical steps. Are the premises sound? Do the conclusions follow from the premises?
Example: "People who graduate from University X are more likely to succeed in politics. Therefore, University X produces better political leaders."
Wait—is success in politics proof of being a "better" political leader? Logical gap.
F - Find Errors
Read carefully for errors: factual errors, grammatical errors, internal contradictions, unsupported claims.
Y - Yield Judgment
After VERIFY, make a judgment: Is this output reliable enough to use? For what purpose? What caveats apply?
High-stakes decisions require higher confidence. Low-stakes uses allow lower confidence.
APPLYING VERIFY
Example: Evaluating a Policy Recommendation
The AI says: "To reduce unemployment in your region, implement a job training program focused on technology skills. Research shows that technology training increases employment by 40% and leads to higher wages. Cities that have implemented similar programs have seen success."
Apply VERIFY:
V - Validate Facts:
- "research shows 40% increase"—specific? Citation needed. Find the research. Does it actually show 40% for your population?
- Vague: "leads to higher wages"—higher than what? By how much? For which jobs?
E - Examine Sources:
- AI provided no specific citations. Get them. Look them up. Check if they support the claims.
R - Review for Bias:
- The recommendation is for technology training. Is that biased toward certain demographics? Accessible to people with disabilities? Appropriate for your region's economy?
I - Inspect Logic:
- Logic seems sound: training -> employment. But what about people who can't access training? What about people for whom training isn't appropriate?
F - Find Errors:
- The general claim is reasonable, but lacking specificity.
Y - Yield Judgment:
- The AI's recommendation has merit, but I need more specific research before deciding. It's a good starting point for further investigation, not a basis for policy decision.
SPECIAL CASE
Sometimes AI systems generate false information with confidence. This is called "hallucination."
Examples:
- Making up a research study that sounds real but doesn't exist
- Citing a law that doesn't exist
- Describing an event that didn't happen
- Creating statistics from thin air
Hallucinations are particularly dangerous because the AI states them with confidence.
How to Detect Hallucinations:
- Verify citations. The AI cites a study. Look it up. If it doesn't exist, hallucination.
- Check specific facts. Dates, numbers, names—if they're wrong, you know the AI hallucinated.
- Ask follow-up questions. If you're suspicious, ask the AI to provide more detail or the source. Sometimes AI will admit it can't find the source.
- Be especially suspicious of:
- Very specific statistics without sources
- Quotes attributed to famous people
- References to specific studies with specific results
ANTI-PATTERNS / MISUSE RISKS
Anti-Pattern 1: Assuming AI Outputs Are Accurate
You ask the AI something. It provides an answer with confidence. You assume it's correct.
Risk: The AI might have hallucinated or made an error.
Anti-Pattern 2: Skipping Verification for Seemingly Obvious Facts
"The AI says the capital of France is Paris. Everyone knows that. No need to verify."
But what if the AI was asked "What's a city that was incorrectly described as the capital of France in 1800?" and it misunderstood? Spot-check even obvious-seeming facts.
Anti-Pattern 3: Not Checking for Bias in Policy Recommendations
The AI recommends a policy. It sounds good. You don't stop to ask: Is this biased? Does it disadvantage any groups? Are there legitimate alternative perspectives?
Risk: You implement a biased policy.
Anti-Pattern 4: Relying on AI for High-Stakes Decisions
You ask the AI "Should we deny this person's benefit?" and use the AI's reasoning as the basis for the decision.
Risk: The decision is too important for AI to handle alone.
PRACTICE / REFLECTION PROMPTS
- Ask an approved AI tool a substantive question relevant to your work. Apply VERIFY. How confident are you in the output?
- Find an instance where you've used information without fully verifying it. What would VERIFY have revealed?
- What types of AI outputs in your work are high-stakes enough to require rigorous VERIFY?
KEY TAKEAWAYS
- Never trust AI output without evaluation. Use VERIFY systematically.
- Validate facts, especially specific claims with numbers or citations.
- Hallucinations are real. AI can generate false information confidently. Verify especially suspicious-sounding claims.
- Check sources. If the AI cites something, follow up. Read the actual source.
- Review for bias. Is the output presenting one perspective as obviously true? Are there legitimate alternatives?
- Make a judgment call. After VERIFY, decide: Is this reliable enough for my purpose?
TERMS / GLOSSARY ITEMS
Hallucination: When an AI generates false information.
Citation: A reference to a source of information.
Bias: Systematic tendency toward a particular perspective or conclusion.
Logic: Reasoning that follows clear rules and sound premises.
Confidence: The degree to which something seems reliable or trustworthy.
You ask an AI: "What are the most effective interventions to reduce recidivism among people released from incarceration?"
The AI responds with a list of approaches and claims: "Research shows that employment support reduces recidivism by 35%. Educational programs reduce it by 28%. Cognitive behavioral therapy reduces it by 15%."
Apply VERIFY:
V: Do those specific percentages actually exist in research? Find the studies. Check if they show those numbers.
E: Are the studies credible? Peer-reviewed? Large sample sizes?
R: Does the list have bias? Are all approaches equally evidence-based? Are some more speculative?
I: The logic (helping people succeed in life reduces crime) is sound. But the specificity of percentages needs validation.
F: One error: the AI says "cognitive behavioral therapy reduces it by 15%" but doesn't note this is for specific populations.
Y: The AI's general recommendation has merit. But I wouldn't rely on the specific percentages without reading the original research. I'd use this as a starting point, then research further.
15 minutes.
Ask an AI system a question about something you know well. Evaluate the output using VERIFY. Did VERIFY reveal any problems with the output?
Evaluating AI outputs is not about being paranoid or distrustful. It's about being professional. Use VERIFY. Trust, but verify. Make your decisions based on information you've personally evaluated.
Government AI CLUB Certification Program
Level 1: AI Aware | Your Agency's Approved AI Tools | Lecture 3.5
A GOVT.CLUB initiative.
<- 1.3.2 Prompt Engineering Basics 1.3.4 AI for Government Tasks: Summarization and Drafting ->
Start Your CLUB Certification
This lecture is part of L1: AI Aware—8 hours of comprehensive government AI training.
Explore CLUB Certification
Related Lectures
L1 1.3.1—Your Agency's Approved AI Tools 20 min - Hands-On Lab
L1 1.3.2—Prompt Engineering Basics 25 min - Hands-On Lab
L1 1.3.4—AI for Government Tasks: Summarization and Drafting 20 min - Hands-On Lab
Frequently Asked Questions
What will I learn in Evaluating AI Outputs?
In this 20 min hands-on lab lecture, you will The VERIFY method: Validate facts, Examine sources, Review for bias, Inspect logic, Find errors, Yield judgment. Hands-on fact-checking exercises
What level is Evaluating AI Outputs?
This is a Level 1 (AI Aware) lecture, part of Chapter 1.3 \u2014 Practical AI Skills. It is designed for all government employees.
How long is lecture 1.3.3?
Lecture 1.3.3 (Evaluating AI Outputs) takes 20 min. It is delivered as a hands-on lab format.
Do I need prerequisites for Evaluating AI Outputs?
This lecture is part of L1 (AI Aware). Prerequisites: None.
What is the CLUB Certification?
CLUB (Community Leading Unified Benchmarks) is a maturity-based AI certification for government professionals with 5 levels (L1-L5), 215 lectures, and 25 chapters aligned with NIST AI RMF, OMB, and GAO frameworks.
Skill.re