AI for Government
Visionary · M13 · lesson 13 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI in Public Safety
📖
now learning

AI in Public Safety

15 min

Learning Objectives

By the end of this lecture, government leaders and policy strategists will be able to: (1) articulate why AI in public safety is classified as both safety-impacting and rights-impacting under OMB Memorandum M-24-10, and what that classification requires of a CFO Act agency; (2) map the policy stack that governs public-safety AI in the United States, including Executive Order 14110, the NIST AI Risk Management Framework (AI RMF 1.0) and its Generative AI Profile, the DHS AI Use Case Inventory requirements, and relevant DOJ guidance on face recognition technology; (3) evaluate real deployments, FEMA's use of machine learning for damage assessment after Hurricanes Ian and Idalia, NYPD's Domain Awareness System, Chicago's historical Strategic Subject List, New Orleans' earlier Palantir predictive-policing engagement, and ShotSpotter gunshot-detection deployments, against a structured harm-and-benefit framework; (4) design an agency-level oversight architecture that includes a Chief AI Officer under OMB M-24-10, an AI Governance Board, a risk-management officer, independent red-teaming, and clear escalation to Inspectors General and civil-rights counsel; and (5) draft the talking points, public-notice language, and legislative engagement strategy needed to build legitimacy for high-stakes AI in policing, 911/CAD systems, wildfire prediction, and emergency response.

Key Topics Covered

This lecture covers six tightly linked domains. First, emergency management AI: FEMA's post-disaster imagery triage, damage proxy models built with the Oak Ridge National Laboratory, and the integration of AI into National Response Framework workflows. Second, wildfire and natural-hazard prediction: USGS, NOAA, and CAL FIRE's use of satellite-plus-ML pipelines, and the role of AI in the Wildfire Crisis Strategy run by the U.S. Forest Service. Third, law-enforcement oversight: face recognition technology (FRT) as used by FBI's NGI-IPS, CBP biometric entry-exit, and state-level deployments, evaluated against the 2023 DOJ guidance, the Vermont and San Francisco FRT bans, and the NIST FRVT vendor tests. Fourth, predictive policing and risk scoring: PredPol/Geolitica, HunchLab, Chicago's Strategic Subject List that the RAND Corporation evaluated in 2016, COMPAS pretrial risk assessment and the ProPublica critique, and the ethical tradeoffs the 2020 Santa Cruz ban surfaced. Fifth, 911/CAD automation and non-emergency response: Carbyne-style caller triage, natural-language transcription, and drone-as-first-responder (DFR) programs in Chula Vista and Brookhaven. Sixth, oversight mechanics: how Inspectors General, state Attorneys General, the GAO, and civilian oversight boards audit public-safety AI under FISMA, the Privacy Act, and 28 C.F.R. Part 23 for criminal intelligence systems.

Why This Matters for Government

Public-safety AI is the highest-stakes domain for government AI because it combines coercive state power (arrest, detention, surveillance, border enforcement, emergency response prioritization) with life-safety consequence (false arrest, denied disaster assistance, missed rescue) in systems that affect people who rarely have the ability to choose another provider or walk away. Federal, state, local, tribal, and territorial agencies deploying AI in policing, corrections, homeland security, emergency management, fire services, maritime safety, border security, transportation security, and public health emergency response operate under a unique constellation of authorities: the Fourth Amendment and Fourteenth Amendment of the US Constitution, the Civil Rights Act of 1964, the Posse Comitatus Act, the Stafford Act for disaster response, the Homeland Security Act, the Aviation and Transportation Security Act, the SAFETY Act (Support Anti-Terrorism by Fostering Effective Technologies Act of 2002), the First Step Act at the federal corrections level, state criminal procedure codes, state and local consent decrees, and (for federal law enforcement) OMB Memorandum M-24-10 rights-impacting AI provisions. Each of those authorities places boundaries on what an AI system can do, who is accountable when it fails, and how affected individuals can obtain redress.

The stakes are concrete and asymmetric. A predictive-policing model that over-weights historically over-policed neighborhoods reinforces the feedback loop documented by the RAND Corporation's 2016 evaluation of the Chicago Strategic Subject List and the Chicago Office of Inspector General's 2020 follow-up, producing real arrests of real people. A facial-recognition match treated as probable cause produced the wrongful arrest of Robert Williams by the Detroit Police Department in January 2020, the first publicly documented US arrest based on a false face-recognition match, and has produced at least six additional publicly known false-match arrests since then, disproportionately affecting Black men. The NIST Face Recognition Vendor Test Part 3 (Demographic Effects, NIST IR 8280, December 2019) found false-positive differentials of up to two orders of magnitude across demographic groups on some algorithms. TSA facial-verification pilots at more than twenty-five airports and US Customs and Border Protection's Biometric Entry/Exit program have drawn sustained congressional and GAO scrutiny (GAO-22-106154) for oversight gaps. The ShotSpotter gunshot-detection system deployed in Chicago, New York, and other jurisdictions has been the subject of ongoing litigation and Office of Inspector General reports questioning accuracy and community impact. Predictive-victimization models, risk assessment instruments like Compas (State v. Loomis, Wisconsin 2016), and pretrial risk tools used by federal and state courts have all been challenged on due-process, equal-protection, and transparency grounds. At the emergency-management end, FEMA's Office of Response and Recovery has reported that imagery-triage tools shorten initial damage estimates after hurricanes from weeks to under seventy-two hours, but errors in that pipeline translate directly into denied Individual Assistance applications and delayed rebuild grants that can cascade for years.

The governance framework is both rich and incomplete. OMB Memorandum M-24-10 (March 2024) classifies law-enforcement biometric identification and risk assessment tools as presumptively rights-impacting, requiring AI Impact Assessments, disparate-impact testing, ongoing monitoring, and documented waiver processes when minimum practices cannot be met. The NIST AI Risk Management Framework 1.0 (January 2023) GOVERN, MAP, MEASURE, and MANAGE functions apply with particular force to public-safety deployments because consequences of failure are both severe and irreversible. Executive Order 14110 (October 2023) places additional obligations on federal law enforcement uses of AI and directs the Department of Justice to develop best practices. The Department of Homeland Security has issued its own AI policies including the DHS AI Strategy (2023) and subsequent use-case inventories across ICE, CBP, TSA, FEMA, USCG, USCIS, and the Secret Service. State-level frameworks vary widely: Washington State's face-recognition law (RCW 43.386) requires accountability reports; Illinois's Biometric Information Privacy Act (BIPA) governs private-sector biometrics with downstream effects on government vendors including Clearview AI; California's AB 1215 imposed a moratorium on face recognition in police body cameras (expired 2023); New York City's Local Law 144 and the NYC AI playbook set procurement-side disclosure requirements. The European Union AI Act (entered into force August 2024) designates law enforcement, migration, and border control AI as high-risk or prohibited under specific conditions, providing a comparative yardstick that US agencies with international partners must understand.

For the Chief AI Officer, the police chief, the fire marshal, the emergency manager, the corrections director, and the agency general counsel, the operational implications are concrete. First, never deploy rights-impacting public-safety AI without a completed AI Impact Assessment that documents statutory authority, use-case scope, affected populations, disparate-impact testing tied to NIST AI RMF MEASURE 2.11 and DOJ guidance, and a human-in-the-loop protocol where facial recognition, predictive assessment, or automated triage informs coercive action. Second, treat vendor claims skeptically: request independent evaluation anchored to NIST benchmarks, audit disparate-impact results, require published model cards and data sheets, and obtain contractual audit rights. Third, engage communities: public-safety AI deployed without community input has a documented failure pattern from the Los Angeles PredPol contract termination to the Santa Cruz police department's abandonment of predictive policing in 2020 to the Pittsburgh Allegheny County Family Screening Tool controversies. Fourth, build redress: an affected person must have a realistic path to contest the outcome, consistent with due process and agency complaint procedures. Fifth, coordinate across federal partners including DOJ Civil Rights Division, DHS Office for Civil Rights and Civil Liberties, CISA, the FBI Criminal Justice Information Services division, NIST, GAO, and the Privacy and Civil Liberties Oversight Board. Sixth, prepare for incident response: document, disclose appropriately, brief Congress, coordinate with OMB, and feed lessons back into policy. The remainder of this seminar operationalizes those directives with case studies from Detroit, Chicago, New Orleans, Houston HISD, CBP, TSA, FEMA, and peer jurisdictions.

The Stakes Are Concrete and Asymmetric

Public-safety AI is where the two most politically sensitive properties of AI systems, coercive power and life-safety consequence, meet in a single deployment. When a damage-assessment model undercounts destruction in a census tract after a hurricane, federal Individual Assistance funds can be misrouted for years; FEMA's Office of Response and Recovery has reported that imagery-triage tools shorten initial damage estimates from weeks to under 72 hours, but errors in that pipeline translate directly into denied applications and delayed rebuild grants. When a predictive-policing model over-weights historically over-policed census blocks, the feedback loop documented by the RAND Corporation's 2016 evaluation of Chicago's Strategic Subject List is no longer theoretical. It is the operational reality that the Chicago Office of Inspector General flagged in its 2020 follow-up report. When NYPD, Detroit PD, or the New Orleans Police Department use face recognition against a witness photo, misidentification risk is not evenly distributed: the NIST Face Recognition Vendor Test Part 3 (Demographic Effects, 2019) found false-positive differentials across demographic groups of up to two orders of magnitude on some algorithms. Robert Williams' wrongful arrest by Detroit PD in January 2020, the first publicly documented U.S. arrest based on a false face-recognition match, is the canonical agency-scale failure. At the federal level, the Transportation Security Administration's face-verification pilots at 25+ airports and U.S. Customs and Border Protection's biometric entry-exit program have drawn congressional and GAO scrutiny (GAO-22-106154) for exactly the oversight gaps this lecture addresses.

The Policy Stack That Governs Public-Safety AI

Effective L5 leaders operate fluently across seven overlapping instruments. (1) Executive Order 14110 (October 30, 2023) on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence directs DHS, DOJ, and the National Security Council to issue guidance for AI in law enforcement and to coordinate with the AI Safety Institute at NIST. (2) OMB Memorandum M-24-10 (March 28, 2024) designates law-enforcement AI, emergency-response dispatch prioritization, and critical-infrastructure monitoring as both 'safety-impacting' and 'rights-impacting,' triggering minimum practices including pre-deployment testing, impact assessment, ongoing monitoring, public notice, and the right to human alternative and appeal. Waivers require Chief AI Officer approval with notice to OMB. (3) OMB M-24-18 on procurement layers vendor disclosure, testing evidence, and post-award monitoring obligations. (4) The NIST AI Risk Management Framework (AI RMF 1.0) and its Generative AI Profile (NIST AI 600-1, July 2024) structure the Govern-Map-Measure-Manage functions and are now the default reference for agency AI governance boards. (5) The DHS AI Use Case Inventory (published annually per Section 7225 of the FY21 NDAA and OMB guidance) forces disclosure of every AI use case and has become the go-to primary source for oversight journalism. (6) 28 C.F.R. Part 23 continues to govern criminal intelligence systems and is now interpreted by DOJ's Office of Justice Programs to apply to AI-augmented intelligence products. (7) For biometric use specifically, the 2023 DOJ Interim Guidance on Facial Recognition and the NIST FRVT benchmarks establish accuracy and operational-use expectations that plaintiffs, defense counsel, and Inspectors General will cite in court and in audits.

Agency-Scale Scenarios L5 Leaders Will Actually Face

Scenario A - The 'we already bought it' problem. A mid-sized city police department procured a ShotSpotter contract in 2019 using asset-forfeiture funds. The mayor now asks you, as a deputy CAO working with OMB's Office of the Federal CIO on grants, whether to condition the next Byrne JAG award on demonstrating M-24-10 minimum practices. You must weigh the MacArthur Justice Center's 2021 Chicago analysis (false-positive rate findings), the City of Chicago's 2024 decision to let the contract lapse, and the counter-evidence from other deployments, and you must do it inside the constraints of Byrne JAG's existing authorities. Scenario B - The FEMA damage-assessment pipeline. After Hurricane Helene, FEMA's National Response Coordination Center uses a Remote Sensing and Damage Assessment pipeline trained on pre- and post-event imagery. You chair the interagency AI governance board. A county in western North Carolina reports that entire hollows were missed because cloud cover defeated the classifier and manual teams could not reach the area. How do you adjust the rule that triggers human re-inspection, and how do you communicate that adjustment without undermining field trust in the tool? Scenario C - The face-recognition request. A state fusion center asks your agency for access to a face-recognition search against driver's-license photos to identify a suspect in a violent-crime investigation. Your state lacks explicit FRT enabling legislation, and the DOJ 2023 interim guidance sets a minimum bar but leaves judgment on specific queries to agencies. What is your decision procedure, who signs off, what is logged, and what is disclosed to the defense under Brady obligations? Scenario D, 911/CAD prioritization. A county CAD vendor proposes an ML model to triage incoming 911 calls by likely severity. The union raises concerns that the model will be used to justify reduced dispatcher headcount. How do you scope the pilot under OMB M-24-10's 'rights-impacting' requirements, what is the minimum acceptable monitoring regime, and what is the rollback trigger?

Case Studies: What Has Actually Gone Wrong, and Why

Michigan's MiDAS unemployment fraud system (deployed 2013-2015) is the canonical U.S. cautionary tale for safety- and rights-impacting automated decisions in government: the state later acknowledged that the system wrongly accused tens of thousands of claimants of fraud at rates later reported above 90 percent error on the contested determinations, with minimal human review. It is not, strictly, a public-safety system, but every principle it violated, inadequate testing, no meaningful human-in-the-loop, no appeal pathway calibrated to the stakes, is exactly what OMB M-24-10 now requires for public-safety AI. The Netherlands' SyRI welfare-risk system (struck down by the District Court of The Hague in February 2020 under Article 8 ECHR) and the Dutch childcare benefits (toeslagenaffaire) scandal that brought down the Rutte III cabinet in January 2021 offer the most developed European jurisprudence on why algorithmic risk scoring in enforcement contexts demands a dramatically higher evidentiary standard. The COMPAS pretrial risk assessment tool, examined by ProPublica in 2016 and defended by Northpointe (now Equivant), illustrates that competing fairness definitions, predictive parity versus equal false-positive rates, cannot all be satisfied at once, and that the choice among them is a policy decision that belongs with elected officials, not with a vendor. The Houston Independent School District's 'value-added' teacher-evaluation system (invalidated by a federal court in 2017 after the Houston Federation of Teachers suit) reinforces the due-process baseline. The Ohio BMV's pandemic-era automated license suspensions similarly underscore that 'automated' is not a defense to constitutional process. Robert Williams (Detroit, 2020), Nijeer Parks (New Jersey, 2019), and Randal Reid (Louisiana/Georgia, 2022) together establish that face-recognition false positives are not edge cases. L5 leaders must be able to name these cases in public, cite what specifically went wrong, and explain how their proposed governance avoids the same failure mode.

Institutional Architecture: Who Does What

A credible public-safety AI program requires a clear division of labor. The Chief AI Officer (required under OMB M-24-10 for CFO Act agencies and strongly recommended for state/local governments at scale) owns the governance program and signs waivers. The AI Governance Board, typically chaired by the CAO with General Counsel, CIO, CISO, CPO (privacy), civil-rights/equity lead, mission owner, and a line attorney from the Inspector General's office as observer, reviews new use cases, sets minimum testing and monitoring standards, and approves waivers. Independent red-teaming, consistent with NIST AI 600-1 guidance and the AI Safety Institute's evaluation practices, tests both model behavior and operator workflow. Inspectors General (at federal agencies, per the Inspector General Act of 1978 as amended, and at state/local levels where equivalents exist) audit against agency-defined standards and produce the public record that Congress, courts, and journalists rely on. GAO reports, such as GAO-21-518SP (AI accountability framework), GAO-22-106154 (CBP biometrics), and GAO-23-105923 (federal law-enforcement FRT), are the external counterweight. Civil-society actors, the ACLU, Electronic Frontier Foundation, Electronic Privacy Information Center, NAACP Legal Defense Fund, Leadership Conference on Civil and Human Rights, Upturn, AI Now Institute, and the Center for Democracy & Technology, provide the outside pressure that keeps the system honest. The vendor ecosystem, Axon, Motorola Solutions, Palantir, Clearview AI, IDEMIA, Thomson Reuters CLEAR, NEC, Dataminr, ShotSpotter/SoundThinking, is neither enemy nor ally; it is a set of counterparties to be procured against under FAR Part 39 and OMB M-24-18 with exit rights, testing evidence, and FedRAMP Moderate authorization where SaaS is used.

Strategic Choices and Tradeoffs

Five decisions define a public-safety AI program. (1) Deployment posture: ban certain uses outright (as San Francisco did for FRT in 2019 and Portland, Oregon did in 2020), moratorium with sunset (as the state of Washington has debated), or permissive with oversight (the federal posture under EO 14110 and M-24-10). Each is defensible; each has costs. (2) Disclosure posture: minimum legal compliance, full use-case inventory plus impact assessments (the OMB M-24-10 baseline), or proactive community engagement (Seattle's Surveillance Ordinance model). (3) Testing regime: vendor-provided evidence only, independent pre-deployment red-team, or continuous monitoring with statistical early-warning triggers. (4) Human-in-the-loop design: reviewer-on-the-loop (human reviews flagged outputs), reviewer-in-the-loop (human must concur before action), or reviewer-after-the-loop (human only reviews a sample post-hoc). For face-recognition and rights-impacting uses, the M-24-10 default is reviewer-in-the-loop with documented independent corroboration before any enforcement action. (5) Sunset and reauthorization: no sunset, periodic reauthorization by the legislative body, or performance-contingent continuation. L5 leaders should resist the temptation to treat any of these as a technical question; each is fundamentally political and requires legitimacy built through stakeholder engagement, not imposed by fiat.

Common Traps at the L5 Level

Trap 1 - The vendor demo trap. A compelling demo of a face-recognition or predictive model shown in an ideal environment (good lighting, curated imagery, balanced dataset) becomes the basis for procurement; operational performance on body-worn camera footage in urban night conditions is 5-20x worse, as NIST FRVT Part 3 and independent academic work consistently find. Defense: require vendor to submit NIST FRVT results on the specific algorithm version and require in-situ testing on the agency's own data before deployment. Trap 2 - The 'it's just a tip' defense. Officials claim the AI output is 'only a lead' and therefore does not require oversight comparable to evidence. This defense collapses at trial and under Brady/Giglio obligations; Detroit's Williams case and New Orleans' Palantir engagement both showed that 'just a lead' in practice became the primary basis for action. Trap 3 - The procurement shortcut. Using asset-forfeiture funds, exigent sole-source authorities, or pilot carve-outs to avoid competitive procurement and governance review. Trap 4 - The 'federated' fig leaf. Calling a system 'federated' or 'de-identified' without cryptographic or statistical backing to justify reduced oversight. Trap 5 - The measurement trap. Tracking 'alerts generated' or 'matches returned' rather than end-to-end outcomes (lawful arrests, convictions, errors avoided, demographic disparity). Trap 6 - The 'we already committed publicly' trap. Political leaders announce a system before the governance review; leaders below them are then told to make it work. The defense is to insert governance review into the announcement pathway, not after it.

Reflection and Application

Before advancing, work through the following against a specific agency you lead or advise. (1) List every public-safety AI system your agency operates or contracts for. Cross-check against the DHS AI Use Case Inventory or equivalent. Identify gaps. (2) For each system, classify it under OMB M-24-10 as safety-impacting, rights-impacting, both, or neither. Write down the reasoning; expect to defend it in front of an IG. (3) For each safety- or rights-impacting system, document the minimum practices: impact assessment, pre-deployment testing evidence, ongoing monitoring, public notice, human alternative/appeal. Identify which are missing and build a remediation plan with dates. (4) For the highest-stakes system, draft the public-facing notice a citizen would see if they were subject to it. Test it with three people who do not work in government. (5) Identify the three stakeholders most likely to sue or investigate you, usually a civil-rights organization, the IG, and a legislative oversight committee, and brief them before they learn about the system from the press. (6) Set a reauthorization date and an early-warning metric. If the metric breaches, the system stops until governance reapproves it. This is the operational meaning of 'trustworthy AI' in public safety.

Key Terms

SAFETY-IMPACTING AI (OMB M-24-10): AI whose output could meaningfully affect human safety, including public-safety dispatch, emergency alerts, and critical-infrastructure monitoring. RIGHTS-IMPACTING AI (OMB M-24-10): AI whose output could meaningfully affect civil rights, civil liberties, or access to critical services. MINIMUM PRACTICES: The floor of testing, monitoring, notice, and appeal required for safety- or rights-impacting AI, waivable only by the Chief AI Officer with notice to OMB. NIST AI RMF: The NIST Artificial Intelligence Risk Management Framework structured around Govern-Map-Measure-Manage. FRVT: NIST's Face Recognition Vendor Test, the authoritative accuracy and demographic-differential benchmark. 28 C.F.R. PART 23: The federal regulation governing operation of federally funded criminal intelligence systems. FEDRAMP: The Federal Risk and Authorization Management Program; FedRAMP Moderate is the default baseline for SaaS handling CUI. AI USE CASE INVENTORY: The annual public disclosure required of federal agencies by OMB under EO 14110 and prior authorities. HUMAN-IN-THE-LOOP / ON-THE-LOOP / AFTER-THE-LOOP: Three distinct oversight postures with different evidentiary weight in court and in audit.

L5 5.3.1 - AI for Mission-Critical Government Functions (240 min, Seminar + Cases)
L5 5.3.2 - AI in Defense and National Security (180 min, Seminar)
L5 5.3.3 - AI in Healthcare Delivery (180 min, Seminar + Cases)
L5 5.3.4 - AI in Infrastructure (180 min, Seminar)
L5 5.3.7 - AI in Education (180 min, Seminar + Cases)