AI for Government
Capable · M25 · lesson 25 of 43 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Measuring AI Impact
📖
now learning

Measuring AI Impact

15 min

Learning Objectives

After completing this lecture, you will be able to:

  • Understand the key concepts of measuring ai impact in a government context
  • Participate in structured workshop activities with real-world scenarios
  • Connect measuring ai impact to your agency's AI initiatives
  • Identify next steps for applying these concepts in your role

Key Topics Covered

-
Defining KPIs for AI projects

-
Before/after measurement

-
Time savings, quality improvements, cost reduction, citizen satisfaction

Why This Matters for Government

Government agencies face unique challenges when it comes to AI adoption. This lecture addresses these challenges head-on by providing analysts, project leads, team supervisors with the knowledge and frameworks needed to navigate AI in the public sector responsibly and effectively.

As part of the L2 (AI Practitioner) curriculum, this lecture builds on the foundational principle that every AI system in government ultimately serves citizens. Whether you are working with AI tools daily or setting strategy for your agency, understanding measuring ai impact is essential for responsible, effective government AI adoption.

======================================================================

TRANSCRIPT: Measuring AI Impact

======================================================================

Chapter: 4 -- AI Project Management

What you will learn:

  • How to define Key Performance Indicators (KPIs) for AI projects
  • Before/after measurement frameworks
  • How to isolate AI impact from other factors
  • How to measure different types of value (efficiency, quality, equity)
  • How to report results to stakeholders
  • Common measurement pitfalls and how to avoid them

You've deployed an AI system. Leadership asks: "Is it working? Are we getting the value we expected?" You need clear answers based on data, not opinions. This lecture teaches how to measure AI impact systematically.

PURPOSE STATEMENT

Impact measurement proves that AI systems deliver the value they promised. Without measurement, you can't demonstrate ROI, can't improve systems, can't make data-driven decisions about scaling. This lecture teaches measurement frameworks that provide evidence of impact.

WHY THIS MATTERS FOR GOVERNMENT

Government is accountable for how it spends money and how it deploys systems. Measurement demonstrates that systems work, that taxpayer money was well-spent, and that citizens benefit. Measurement also drives continuous improvement--you identify what's working and what needs adjustment.

KPIs FOR AI SYSTEMS

Key Performance Indicators are measurable metrics that show whether the system is delivering its intended value.

CATEGORIES OF KPIs:

Efficiency KPIs:

  • Time savings: Hours of manual work eliminated per case
  • Cost savings: Cost per decision before and after AI
  • Throughput: Cases processed per day
  • Staff productivity: Output per staff member

Quality KPIs:

  • Accuracy: Percentage of correct decisions
  • Consistency: How much do similar cases get similar outcomes?
  • Rework rate: Percentage of decisions that get overturned or redone
  • Quality of outcomes: Downstream impact (e.g., eligibility errors that cause problems)

Equity KPIs:

  • Demographic parity: Are outcomes equal across groups?
  • Disparate impact analysis: Any illegal discrimination?
  • Equity in access: Are all populations served?
  • Equity in speed: Do all groups get equal wait times?

Adoption KPIs:

  • Usage rate: What percentage of cases go through the system?
  • User confidence: Do staff trust the system?
  • Training completion: What percentage of staff completed training?
  • Sustainability: System working without constant central support?

Business KPIs:

  • ROI: Cost savings vs. investment
  • Timeliness: Processing time for decisions
  • Stakeholder satisfaction: Are customers satisfied with decisions?
  • Risk reduction: Cases that would have caused problems caught early

BEFORE/AFTER MEASUREMENT

Overview

The classic approach to measuring impact: Measure the baseline (before AI), deploy the system, measure again (after AI), compare.

BASELINE MEASUREMENT

Before deploying the system:

  • Define what you're measuring (e.g., average time to process a case)
  • Measure it for a period (e.g., 4 weeks of manual processing)
  • Get multiple measurement points, not just one
  • Document any unusual circumstances affecting baseline

Example baseline:

  • Average processing time without AI: 45 minutes per case
  • Cost per case: $18
  • Accuracy rate: 87% (of manual decisions, how many are eventually correct?)
  • Processing capacity: 50 cases per day

POST-DEPLOYMENT MEASUREMENT

After deploying the system (allow time for adoption first--2-4 months):

  • Measure the same metrics
  • Compare to baseline
  • Account for differences (staffing changes, volume changes, case mix changes)
  • Measure consistently (same time period, same methodology)

Example post-deployment:

  • Average processing time with AI: 18 minutes per case (60% reduction)
  • Cost per case: $7 (61% reduction)
  • Accuracy rate: 91% (4 point improvement)
  • Processing capacity: 120 cases per day (140% increase)

ATTRIBUTING CHANGE TO AI

This is the hard part. You see improvements after deploying AI, but were they caused by AI or by other factors?

Control for other factors:

  • Volume changes (if cases increased, some time savings might be due to economies of scale)
  • Staffing changes (if you hired more staff, processing improvements might be due to that)
  • Process improvements (if you also changed workflows, which change drove improvements?)
  • Seasonal factors (some work has natural seasonal variation)

Solution: Use a control group if possible.

  • Deploy AI to office A but not office B (which is similar)
  • Measure both offices
  • Attribute improvements in office A to AI, baseline drift in office B to external factors
  • If office B also improves, that drift is common to both and not due to AI

If control group isn't possible:

  • Document all other changes you made at the same time
  • Account for them in your analysis
  • Make conservative claims about AI's contribution
  • Be transparent about confounding factors

MEASUREMENT APPROACHES FOR DIFFERENT VALUE TYPES

EFFICIENCY VALUE MEASUREMENT

Goal: Show that the system saves time and money

Metrics:

  • Processing time: Measure hours per case before and after
  • Cost per decision: Calculate labor cost, system cost, overhead
  • Throughput: Cases processed per staff member per day
  • Rework rate: Percentage of decisions that need to be redone

Calculation example:

  • 50 cases/day without AI at $18/case = $900/day total cost
  • 120 cases/day with AI at $7/case = $840/day total cost
  • Net savings: $60/day even with more output
  • Annual savings: ~$15,000 (assuming 250 work days)

QUALITY VALUE MEASUREMENT

Goal: Show that the system produces better outcomes

Metrics:

  • Accuracy: Percentage of correct decisions
  • Consistency: How often do similar cases get similar outcomes?
  • Appeals/overturns: Percentage of decisions that are challenged
  • Downstream quality: Are eligible people getting benefits? Are ineligible people denied?

Measurement: This is complex because you need to know the true "right answer"

  • Do spot checks of decisions
  • Have experts review sample of decisions
  • Track downstream outcomes (if someone gets benefits through AI system, are they using it correctly?)
  • Track complaints/appeals

EQUITY VALUE MEASUREMENT

Goal: Show that the system treats all groups fairly

Metrics:

  • Demographic parity: Are approval rates equal across groups?
  • Equalized odds: Do false positive rates equal across groups?
  • Wait time equity: Do all groups wait the same time?
  • Outcome quality by group: Are all groups getting good decisions?

Measurement: This requires tracking demographic data

  • Link decisions to demographic characteristics
  • Compare outcomes across groups
  • Test for disparate impact
  • Compare to legal thresholds

COMMON MEASUREMENT PITFALLS

PITFALL 1

You measure processing time, which improves, but the goal was to improve accuracy. You succeed at the wrong metric.

How to avoid: Define success metrics upfront as part of requirements. Measure what matters, not what's easy to measure.

PITFALL 2

You deploy the system, measure results after 2 weeks. Adoption is low, results are disappointing. You conclude the system doesn't work.

How to avoid: Allow time for adoption (2-4 months) before declaring success or failure. Growth takes time.

PITFALL 3

You see improvements after deploying AI. You claim success. You don't account for the fact that you also hired 5 new staff members.

How to avoid: Document other changes happening at the same time. Use control groups if possible. Make conservative claims.

PITFALL 4

Staff learn what you're measuring and game the metrics. Processing time metric improves because staff are rushing through cases, not because AI is helping.

How to avoid: Measure quality alongside efficiency. Don't incentivize staff to optimize single metrics at the expense of overall outcomes.

PITFALL 5

You easily measure processing time. You can't easily measure fairness or equity, so you skip it.

How to avoid: Invest in measurement infrastructure for the metrics that matter, even if they're hard. Fairness matters.

PRACTICE PROMPTS

EXERCISE 1

Design a KPI framework for measuring the impact of an AI hiring assistance system:

  • What efficiency gains would you measure?
  • What quality metrics matter?
  • What equity metrics are critical?
  • How would you establish baseline?
  • What's your post-deployment measurement plan?

EXERCISE 2

You deploy an AI system and measure improvements:

  • Processing time down 30%
  • Accuracy up 5%
  • Diversity of hires up 15%

What questions would you ask to attribute these changes to the AI system vs. other factors?

EXERCISE 3

Design a dashboard for reporting impact quarterly to leadership. What metrics would you include? How would you present results to make the case for continued investment?

KEY TAKEAWAYS

  • Define KPIs before deployment. Know what success looks like in measurable terms.
  • Measure baselines before deploying. You need before/after comparison to show impact.
  • Account for confounding factors. Isolate AI's contribution from other changes.
  • Measure multiple dimensions. Efficiency alone doesn't prove success if quality dropped.
  • Measure equity explicitly. Don't assume fairness; verify it.
  • Allow time for adoption. Don't measure impact until 2-4 months post-deployment.
  • Make measurement actionable. Use results to improve the system.
  • Report results transparently. Show both successes and areas for improvement.

GLOSSARY

KPI (Key Performance Indicator): Measurable metric that indicates whether an initiative is delivering its intended value or objectives.

Baseline: Measurement of current state before intervention, used for comparison to show impact of changes.

Disparate Impact: Statistical difference in outcomes for protected groups that may indicate discrimination, even without intentional bias.

Control Group: Population similar to the group receiving intervention but not receiving it, used to distinguish effects of intervention from other factors.

Metric Gaming: Behavior where performance is optimized for the metric measured at the expense of true objectives, distorting results.

Impact measurement transforms AI from "seems like it's working" to "here's the evidence." Good measurement drives three important outcomes: It demonstrates value (ROI, outcomes), it drives improvement (here's what needs to get better), and it builds trust (we're measuring and transparent).

Think of a change you've implemented--in work or personally. How would you measure whether it actually worked? What would you measure? What baseline would you establish? Use this reflection to develop intuition for impact measurement.

You've learned how to measure AI impact. The next lecture focuses on communicating that impact to leadership--how to translate technical results into business language that executives understand and care about. See you there.

Government AI CLUB Certification Program

Level 2: AI Ready | Measuring AI Impact | Lecture 2.4.7

A GOVT.CLUB initiative

Visit: https://govt.club/learn/lectures/l2/247-measuring-impact.html

======================================================================

<- 2.4.5 Change Management for AI Adoption
2.4.7 Communicating AI Projects to Leadership ->

Start Your CLUB Certification

This lecture is part of L2: AI Practitioner -- 40 hours of comprehensive government AI training.

Explore CLUB Certification

L2
2.4.1 -- How AI Projects Differ from Traditional IT
60 min - Video + Comparison

L2
2.4.2 -- Requirements Gathering for AI
60 min - Workshop

L2
2.4.3 -- Working with AI Vendors and Contractors
60 min - Video + Checklist