Success Metrics for Public-Safety AI
The deputy chief walked into the quarterly review with a single slide: "Average report time down 73%." The room applauded. The city manager nodded. The IT director smiled. Nobody asked the follow-up question, which was the one that mattered: had accuracy held, had disclosure compliance kept pace, and were any of those faster reports already creating problems downstream in the courts? Six months later, the answer came from the prosecutor's office in a terse letter, not a press release. Three AI-assisted reports in a homicide case had been flagged for undisclosed AI involvement, and the defense was filing a motion.
Why Speed Alone Is a Trap
Time saved is a real metric. Officers at agencies piloting Axon Draft One, which drafts police report narratives from body-worn camera (BWC, the recording device on the officer's uniform) audio, reported an 82% decrease in report-writing time. That is not a vendor talking point to dismiss. That is a genuine productivity gain that returns hours to patrol, to community engagement, to investigations. When officers currently spend 30 to 40 percent of every shift on paperwork, an 82% reduction in the report-writing portion of that load is transformative.
But a metric is only as good as what it does not measure. A speed metric that runs alone, without a paired integrity metric, does something dangerous inside an agency: it creates institutional pressure to be fast. That pressure is not cynical. It is ordinary organizational behavior. When leadership celebrates throughput, the people doing the work optimize for throughput. When that optimization reaches the verification step, the verification step shortens. When the verification step shortens on an AI-assisted report, the gap-fill that should have been caught does not get caught. The report is fast. The report is wrong. The report is sworn. And no metric on anyone's dashboard shows it.
The key performance indicator (KPI, a tracked measure used to evaluate progress toward an organizational goal) for report time is not wrong. The error is treating it as sufficient. Success in public-safety AI is not a single number. It is a minimum of two axes measured in parallel, and command staff who understand this protect their agencies from a class of failures that a speed-only dashboard will never surface.
Measure what you want, and you will get what you measure. Measure speed alone in a context where accuracy is constitutional, and you will eventually get fast, accurate-looking, wrong reports.
The Dual-Axis Scorecard
The framework that works is a dual-axis scorecard: one axis for efficiency, one axis for integrity. Neither axis subordinates the other. Both are reported together, to the same audience, in the same briefing. When they diverge, the divergence is the most important data point on the board.
Axis One: Efficiency Metrics
Efficiency metrics measure the operational benefit that AI is delivering. They are the numbers that justify the investment and demonstrate the program's contribution to the agency's mission. For a report-writing AI program, the core efficiency metrics are:
Report-writing time per incident. This is the elapsed time from incident closure to report submission, measured for AI-assisted reports and compared against a baseline from before the program. For agencies using Axon Draft One or comparable tools, the 82% time reduction figure provides a benchmark, but every agency's reduction will vary based on incident complexity, officer experience, and how thoroughly the verification step is being conducted. The number you are looking for is the agency-specific figure, measured consistently across a representative sample, not the vendor's headline number applied to your situation without verification.
Shift time recovered. If report writing occupied 30 to 40 percent of a shift, and if AI-assisted tools are reducing that by a meaningful fraction, how many officer-hours per week are being returned to other activities? Translate the report-time reduction into shift-time-recovered figures. This is the number that resonates with city councils: "We are returning approximately 12 officer-hours per patrol shift to community contact and response" is a statement about public safety impact, not just software performance.
Report queue age. In records units, the AI benefit often shows up as a reduction in the age of the report queue: the time between when an incident happens and when the complete report is in the records management system (RMS, the agency's database for case files and reports). A reduction in queue age has downstream benefits for prosecutors, for public-records requests, and for investigation timelines. Measure it before and after.
Redaction throughput. For agencies using AI-assisted redaction in body-camera footage release workflows, redaction throughput, the number of footage packages processed per week or month, is an efficiency metric worth tracking. The backlog is real in most agencies, and AI-assisted redaction is one of the higher-value, lower-risk applications of the technology.
Axis Two: Integrity Metrics
Integrity metrics measure whether the AI-assisted output is accurate, complete, and legally compliant. They are what keep the efficiency gains from becoming a liability. Without integrity metrics, you cannot know whether the time savings are coming from genuine efficiency or from a shortening of the verification step that should not be shortening. The core integrity metrics are:
Report accuracy rate. This is the percentage of AI-assisted reports that, when reviewed in the quality-assurance (QA) process, contain no unresolved factual errors. An "unresolved factual error" is a claim in the sworn report that cannot be verified against the BWC footage, the computer-aided dispatch (CAD) entry, or other documented sources, and that was not corrected before submission. This metric requires a sampling protocol: you cannot review every report, but you can review a statistically meaningful random sample and track the accuracy rate over time. A rate that is declining while report-writing speed is increasing is a specific and actionable warning signal.
Verification completion rate. Does the agency have a record that the verification pass was completed for AI-assisted reports? If the workflow requires officers to log the verification step, that log is a compliance metric. The verification completion rate is the percentage of AI-assisted reports that have a logged, completed verification pass on file. A gap between "reports submitted with AI assistance" and "reports with a logged verification pass" is a governance finding, not just a training issue.
Disclosure compliance rate. When agency policy and prosecutor agreements require that AI assistance be disclosed in the report or in the case file, is that disclosure happening? Brady v. Maryland (the 1963 Supreme Court decision requiring prosecutors to disclose exculpatory evidence to the defense) and Giglio v. United States (the 1972 decision requiring disclosure of evidence affecting witness credibility) create a constitutional disclosure framework that AI use can implicate directly. The disclosure compliance rate is the percentage of AI-assisted reports that include the required disclosure language. This is not optional bookkeeping. It is constitutional compliance.
AI-originated error catch rate. When errors are found in AI-assisted reports, what percentage were caught during the officer's own verification pass versus later, in supervisor review, in QA, by a prosecutor, or by the defense? The earlier in the process errors are caught, the lower the legal and institutional cost. An error caught during the officer's verification pass costs a correction and a logged note. An error caught by the defense in a suppression motion costs the case.
Post-submission challenge rate. Track how many AI-assisted reports receive a legal challenge (suppression motion, Brady challenge, Giglio disclosure request, discovery dispute) compared to a baseline period before AI assistance. If the challenge rate is rising while the submission rate is also rising, that is an early warning sign that the integrity axis of the program is not holding pace with the efficiency axis.
Defining What Counts as a Success
A well-designed metrics program requires clear definitions before deployment, not after the first problem surfaces. The definitions that matter most:
Defining an Unresolved Error
An unresolved error is a factual claim in a submitted, sworn AI-assisted report that cannot be supported by the evidentiary record (BWC footage, CAD entry, field notes, or other documented sources) and that was submitted without correction. An error that was caught during verification and corrected is not an unresolved error. It is a catch, and it should be counted as a positive data point in the AI-originated error catch rate. The distinction matters because a metrics program that treats all AI errors as failures, including the ones that were caught and corrected, will undercount success and may discourage officers from using the verification process honestly.
Defining a Disclosure Event
A disclosure event is any situation where the use of AI in generating a report or investigative document was required to be disclosed under agency policy, a prosecution agreement, or applicable law, and where that disclosure was made. Each agency's disclosure requirements will differ depending on the agreements in place with the local prosecutor's office, the jurisdiction's open-records rules, and the agency's own policy. The metric is whether the disclosure happened when it was required, not whether AI was used at all.
The King County Governance Line
The King County, Washington, prosecutor's office barred AI-written police reports in a direct governance response to the disclosure and accuracy concerns they were seeing. That decision is not a cautionary tale to be dismissed as overcaution. It is a concrete signal of what happens when an agency deploys a report-writing AI without a credible, documented answer to the questions that a prosecutor must ask: Was this AI-assisted? Was it verified? Is the verification documented? Who is the accountable author?
The EFF (Electronic Frontier Foundation, a digital-rights advocacy organization) has raised parallel transparency concerns: that AI-assisted police reports, when not disclosed, obscure the provenance of evidence in a way that undermines the defendant's right to challenge that evidence. A metrics program that includes a disclosure compliance rate and an accuracy rate is the evidence-first answer to both the King County objection and the EFF transparency concern. It does not just assert that the program is operating responsibly. It demonstrates it, with numbers, on a recurring schedule.
Building the Measurement Infrastructure
Metrics require infrastructure. Without a data collection mechanism, the numbers you need are either unavailable or only available through expensive manual review. The infrastructure investment is not optional if the agency wants a metrics program that is credible under scrutiny.
The Audit Log as the Foundation
The foundation of the integrity metrics infrastructure is the AI audit log: a persistent, tamper-evident record of every AI-assisted report, including which tool was used, which version of the tool, when the AI draft was generated, when the officer opened the verification screen, when the verification was logged as complete, what corrections were made, and when the report was submitted. Most modern BWC evidence platforms that include AI drafting capabilities produce some version of this log. The question is whether the agency has configured it, is retaining it, and is using it as a measurement input.
The audit log answers questions that other data sources cannot. If a report is challenged in court and the defense asks whether the verification pass was actually completed, the audit log is the answer. If the agency wants to know whether verification completion rates are higher for certain officers or certain incident types, the audit log provides the analysis data. The CJIS (Criminal Justice Information Services) Security Policy, which governs handling of criminal justice information, places data-handling obligations on the agency, not the vendor. The audit log is agency data, and the agency is responsible for its security, retention, and availability.
Sampling Protocols for QA
No agency will have the capacity to QA-review every AI-assisted report. A sampling protocol creates a defensible, consistent process that surfaces systemic problems without requiring universal review. The sampling design should address:
Sample size. For agencies processing hundreds of reports per month with AI assistance, a 5 to 10 percent random sample is generally sufficient to produce statistically meaningful accuracy rates. Smaller agencies may need a higher percentage to get meaningful numbers. Consult with your city attorney or legal counsel about what sample size provides a defensible basis for accuracy claims.
Stratification. Not all incident types carry the same risk. Use-of-force reports, arrest reports, and reports in cases that have been charged (where a prosecutor is looking at the file) should be sampled at a higher rate than property-only reports or minor infractions. The sampling design should reflect the risk profile of the incident types, not treat all reports as equivalent.
Independence. The QA review should not be conducted by the officer who submitted the report or by that officer's immediate supervisor. An independent review function, whether a dedicated QA unit or a rotating review panel, provides a check that self-review cannot.
Baseline Before Deployment
The single most common measurement failure in public-safety AI programs is deploying the tool before collecting baseline data. Without a pre-deployment baseline, there is no "before" to compare against. The agency cannot show how much report time has decreased because nobody measured it before. The agency cannot show how accuracy compares to the pre-AI period because nobody measured it before. The agency cannot show disclosure compliance has improved because nobody measured whether there was disclosure before.
Before deploying AI-assisted reporting tools, spend at least 60 days collecting baseline measurements on the metrics that matter: average report completion time, error rates in supervisor QA reviews, disclosure compliance where it is already required, and report queue age. These numbers will be the "before" of every subsequent "before and after" comparison the agency makes to command, council, and the public.
Reporting Rhythm and Thresholds
Metrics without a reporting rhythm are just data. The program needs a defined cadence for surfacing the numbers to the people who need to act on them.
Weekly operational reporting. Verification completion rates and report queue age should be surfaced weekly to supervisors. These are operational metrics that can reveal a problem in near-real-time. A supervisor who sees that verification completion rates dropped from 95% to 78% in a single week can investigate immediately, before the pipeline fills with unverified reports.
Monthly program reporting. Report accuracy rates, AI-originated error catch rates, and disclosure compliance rates should be reported monthly to the program manager or the command-level AI lead. Monthly reporting provides enough data volume for the numbers to be meaningful while keeping the reporting interval short enough to catch trends before they become entrenched.
Quarterly review to command and council. The dual-axis scorecard, efficiency and integrity together, should be presented quarterly to command staff and, in an appropriate form, to the city council or oversight board. The quarterly review is the accountability moment: it is when leadership sees whether the efficiency gains are holding and whether the integrity metrics are holding alongside them.
Thresholds for escalation. The metrics program needs defined thresholds that trigger escalation. A disclosure compliance rate below 90% is not just a number: it is a policy failure that requires immediate investigation. A post-submission challenge rate that has doubled over two quarters is not a fluctuation: it is a signal that something in the workflow is broken. Define the thresholds before deployment, in writing, with clear escalation paths. Thresholds set after a problem surfaces look like they were set to cover the problem.
Communicating the Metrics Story
The dual-axis scorecard has to be communicated in a way that different audiences can act on. Command staff, city councils, prosecutors, and oversight boards all care about the numbers, but they ask different questions about them.
Command staff want to know whether the program is operating safely and whether it is delivering the promised operational benefit. The command briefing should lead with the efficiency gain (here is the time we have recovered, here is what that time is going back to), follow with the integrity picture (here is the accuracy rate, here is the disclosure compliance rate, here is where we are catching errors), and close with any open risks or concerning trends that require command attention or resource investment.
City councils and city managers want to know whether the investment is justified and whether the agency is protecting the city from legal exposure. The council presentation should translate the time-saved numbers into dollar figures and service-delivery improvements, then pair that with a clear statement about the governance controls that protect the city from the risks the council will hear about from advocacy groups and the media. "We are saving approximately 1,200 officer-hours per month in report writing and returning those hours to patrol, and we have a disclosure compliance rate of 97% with a documented QA process that any audit can examine" is a statement a council member can act on and defend to constituents.
Prosecutors want to know whether the AI-assisted reports are accurate and properly disclosed. Before a prosecutor will engage seriously with an agency's AI program, they need credible answers to two questions: "Is the AI involvement in this report disclosed to me and to the defense?" and "Is this report accurate?" The metrics program provides the evidence-based answer to both questions. A prosecutor who sees consistent documentation of disclosure compliance and a QA-verified accuracy rate has a basis for confidence that an agency asserting "we verify everything" without data does not provide.
Key Takeaways
- Speed is a real and valuable metric for public-safety AI programs, but measuring it alone creates institutional pressure to skip verification, which turns faster reporting into faster errors in sworn evidence.
- The dual-axis scorecard pairs an efficiency axis (report-writing time, shift time recovered, queue age) with an integrity axis (accuracy rate, verification completion rate, disclosure compliance rate, error catch rate, post-submission challenge rate) and reports both together on a defined schedule.
- Axon Draft One and similar tools report an 82% reduction in report-writing time for tested officers; agencies should verify their own reduction against a pre-deployment baseline, not apply the vendor benchmark directly.
- Disclosure compliance is a constitutional metric, not optional bookkeeping: Brady v. Maryland and Giglio v. United States create obligations that AI-assisted reporting directly implicates, and a disclosure compliance rate below threshold is a legal risk, not just a process gap.
- The King County prosecutor's ban on AI-written police reports and the EFF's transparency concerns are the predictable consequence of deploying AI-assisted reporting without credible, documented integrity metrics; the dual-axis scorecard is the evidence-first answer to both.
- The audit log, produced by the AI drafting platform and retained by the agency as CJIS-governed data, is the foundation of the integrity metrics infrastructure: it is the record that proves the verification pass happened, that errors were caught, and that disclosure was made.
- Baseline data must be collected before deployment; without it, there is no credible before-and-after comparison for any metric the agency presents to command, council, or oversight bodies.
- Define escalation thresholds for every metric before deployment, in writing, with clear response protocols: a threshold set after a problem surfaces looks like it was set to manage the problem rather than prevent it.
Skill.re