Avoiding Metrics That Erode Trust
In the ninth month of the agency's AI report-writing program, a supervisor noticed something in the weekly data. Report submission times were still fast, faster than ever. But the number of AI-originated errors being caught at the verification pass had dropped not because there were fewer errors, but because the verification completion times had compressed so dramatically that genuine verification was no longer happening. The dashboard was green. The program was quietly failing.
How a Good Metric Becomes a Bad Incentive
A metric does not have to be wrong to be dangerous. The most damaging metrics in public-safety AI programs are the ones that are technically accurate and institutionally destructive at the same time. Speed, measured as report submission time from incident close, is an accurate metric. It is also a metric that, in the absence of paired integrity measures, teaches the people doing the work exactly what the organization values. And organizations that value speed get speed. What they also get, in the context of sworn evidence, is AI errors that the verification pass no longer has time to catch.
This is not a failure of individual officers. It is a failure of incentive design. When the metric on the supervisor's dashboard is report submission time, and when that metric is what the supervisor is evaluated on, and when the supervisor is the person who tells officers whether they are performing well or poorly, the direction of that system is not ambiguous. Officers who submit quickly are performing well on the measure that counts. Officers who spend the time to run a genuine verification pass are producing the same green dashboard number as officers who log a 90-second verification on a 45-minute incident. There is no signal in the current metric that distinguishes between the two.
The key performance indicator (KPI, a tracked measure used to evaluate progress toward an organizational goal) problem in public-safety AI is not that agencies measure the wrong things. It is that they measure too few things, and the things they measure create pressure against the behaviors they need most.
A metric that pressures officers to skip verification is not a productivity tool. It is a machine for producing sworn errors at scale.
The Four Metrics That Erode Trust
There are four specific metric patterns that public-safety AI programs use that reliably produce trust erosion over time. They do not appear dangerous at first. They look like efficiency gains, volume achievements, and process improvements. The erosion happens quietly, inside the verification step and the disclosure step, until a legal challenge or an oversight audit makes it visible.
Throughput Without Accuracy
The throughput-only metric counts reports submitted, calls processed, records released, or cases documented, without any paired measure of whether those outputs are accurate. In a non-evidence context, throughput is a reasonable measure. In a context where the output is a sworn statement that will be read by a defense attorney, tested at trial, and potentially the basis of someone's conviction or acquittal, throughput without accuracy is a measure of how fast the agency is generating potentially compromised evidence.
Axon Draft One (the tool that drafts police report narratives from body-worn camera (BWC, the recording device on the officer's uniform) audio) testing showed an 82% decrease in report-writing time. That is a throughput story. The corresponding accuracy story, which tells whether the reports produced in 18% of the previous time are as accurate as the reports produced in the previous 100% of the time, is the story that determines whether the throughput gain is safe to claim. Without both numbers, the throughput number is not just incomplete. It is actively misleading about the program's effect on report quality.
The specific failure mode throughput-only metrics create is the following sequence: the agency celebrates reports-per-shift increasing; supervisors feel pressure to maintain that number; supervisors implicitly or explicitly signal to officers that fast submission is valued; officers internalize the signal; the verification pass shortens to match the pressure; AI errors pass into sworn reports; over time, the post-submission challenge rate rises; the organization attributes the challenge rate to external factors (active defense bar, complicated cases) rather than to the metric pressure that created the conditions for uncaught errors. By the time the causal connection is visible, multiple cases have been affected and the erosion of prosecutorial and community trust is already underway.
Speed Without Disclosure
The speed-without-disclosure failure is the disclosure compliance version of the throughput problem. An agency might track how quickly officers submit AI-assisted reports and celebrate the reduction in submission time, while having no metric at all for whether those reports disclose that AI was used. The case for disclosure is constitutional: Brady v. Maryland (the 1963 Supreme Court decision requiring prosecutors to disclose exculpatory evidence to the defense) and Giglio v. United States (the 1972 decision requiring disclosure of evidence affecting witness credibility, including about officers) frame AI use as a disclosure matter when the AI output is part of the evidentiary record. The EFF (Electronic Frontier Foundation, a digital-rights advocacy organization) has raised transparency concerns specifically about undisclosed AI involvement in police reports, arguing it undermines the defendant's right to challenge the evidence.
The King County, Washington, prosecutor's office barred AI-written police reports in direct response to disclosure concerns. The bar came after the prosecutor's office was receiving AI-assisted reports without a clear way to determine that AI was involved or how the AI output had been reviewed. That is the end state of a speed-without-disclosure metrics program: the disclosure that should have been built into the protocol from the start becomes the reason an oversight body removes the tool entirely.
An agency's trust with its prosecutor's office is not rebuilt quickly. If the prosecutor's office has been receiving undisclosed AI-assisted reports and discovers it, the agency is not in a conversation about how to fix the disclosure protocol. It is in a conversation about whether the agency can be trusted to manage the tool at all. The speed gain that came from skipping the disclosure step costs more than the time it saved.
Activity Metrics Without Outcome Metrics
Activity metrics measure what the AI tool is doing: how many drafts it has generated, how many redactions it has proposed, how many call summaries it has produced. Outcome metrics measure what those activities have actually produced: how accurate the drafts were, how correct the redactions were, whether the call summaries were tied to what the caller actually said. A program that measures activity without measuring outcome has data on how busy the tool is, not on whether the tool is helping.
This distinction matters most in redaction workflows. An agency using AI-assisted redaction of body-camera footage for public-records release can measure its redaction throughput (number of footage packages processed per month) without measuring its redaction accuracy (percentage of required redactions correctly made, percentage of packages reviewed before release, number of missed faces or plates identified in QA). A throughput number that is rising while accuracy is unmeasured is a program that may be releasing footage with unredacted faces, unredacted license plates, or private information that was supposed to be withheld. The missed redaction is a privacy violation. It is also a public-records lawsuit. It is also the kind of visible failure that triggers an oversight audit and a community trust crisis that years of efficient processing cannot repair.
CJIS (Criminal Justice Information Services) Security Policy, which governs how criminal justice information is handled, places the data-handling obligations on the agency, not the vendor. A missed redaction in a release is an agency failure, not a vendor failure. The metric that would have caught it, a redaction accuracy rate from QA sampling of released packages, was not being tracked. The activity metric (packages processed) was green. The outcome was a violation.
Averages That Hide the Distribution
The fourth pattern is the one that looks the most like a well-designed metric and is the most sophisticated failure mode. An accuracy rate of 89% sounds like a strong program. And if that 89% rate is uniformly distributed across all report types, it probably represents a strong program. But if the 89% overall rate is composed of a 97% accuracy rate on property crime reports and a 71% accuracy rate on use-of-force and arrest reports, the average is hiding the problem where it matters most.
Use-of-force reports are the reports that face the highest legal scrutiny, the most intensive defense review, and the most severe consequences when inaccurate. A suppression motion, a civil rights claim, a Brady challenge, and an oversight investigation all land most heavily on use-of-force documentation. An accuracy rate that averages together low-stakes property reports with high-stakes use-of-force reports is not giving leadership the information they need to protect the agency and the people involved in those incidents.
The same averaging problem appears in verification completion rates. An agency that reports an overall verification completion rate of 93% may have a 98% rate for routine reports and a 81% rate for use-of-force and officer-involved incidents, precisely where verification is most critical and most difficult. If the dashboard shows only the aggregate, the rate looks strong. If the rate is broken out by incident type, a trend line on the use-of-force verification rate tells a very different story.
The fix is stratification: reporting accuracy rates, verification completion rates, and disclosure compliance rates broken out by incident type, with use-of-force, arrest, and charged-case reports in their own reporting tier. That stratification tells leadership what the program is actually doing where it matters most.
The Metric That Pressures Verification
The specific failure mode this lesson focuses on is the metric that creates pressure to skip or compress the verification pass. It is worth examining in detail because it is the most direct path from a well-intentioned efficiency program to a constitutional accountability problem.
The verification pass is the step in the AI-assisted report workflow where the officer checks every factual claim in the AI draft against the BWC footage, the computer-aided dispatch (CAD) entry, and field notes, and corrects any claim that is unsupported. The verification pass is time-consuming by design. A genuine verification pass on a 45-minute incident with a use-of-force component takes meaningful time: the officer must locate the relevant footage timestamps, review each factual claim at the corresponding timestamp, and make corrections to any claim the footage does not support. That time cannot be compressed without compressing the verification itself.
When the agency's primary metric is report submission time, the verification pass is the part of the workflow that creates variance. Officers who want to perform well on the submission time metric have a strong incentive to do what makes submission faster, and the verification pass is what makes submission slower. The metric is not punishing bad verification. It is punishing slow submission. In a context where bad verification and fast submission can coexist (because the metric measures only submission time), the rational response is fast submission with minimal verification.
What does this look like in practice? It looks like a verification log that shows completion times of 60 to 120 seconds for incidents that ran 30 to 60 minutes. It looks like use-of-force reports with boilerplate language that matches across multiple incidents regardless of the specifics of each encounter. It looks like a rising post-submission challenge rate that coincides with a falling verification completion time and rising submission speed. These patterns are detectable in the audit log data if someone is looking for them. The problem is that in a speed-only metrics culture, the data that would show the problem is not being analyzed because the only number that matters is the one that is green.
The Brady Exposure from Compressed Verification
The legal consequence of a compressed verification pass is a Brady v. Maryland exposure that the agency did not intend to create and cannot easily undo. Brady requires disclosure of exculpatory evidence. A factual error in an AI-assisted report, where the report says the subject's hands were raised when the footage shows one hand raised and one at the subject's side, is not Brady-exculpatory evidence in the strict sense if it was an innocent error. But if the error was the result of a verification pass that was skipped because the submission time metric created pressure to submit quickly, and if the defense can show that pattern through the audit log, the agency is now in a conversation about whether its verification protocol was a genuine control or a compliance checkbox.
A defense attorney who can show that the agency's officers routinely completed verification passes in 90 seconds for 45-minute incidents has evidence that the verification protocol was not being followed in substance. That evidence is relevant to every AI-assisted report the officer ever submitted, not just the one in the case before the court. The Giglio disclosure obligation (which requires disclosure of evidence affecting witness credibility) extends to patterns of behavior that affect the reliability of an officer's reports. A pattern of de-facto verification skipping, visible in the audit log, is evidence that affects witness credibility within the Giglio framework.
Building Metrics That Protect Instead of Erode
The solution is not to abandon speed metrics. Speed is a legitimate, important measure of the operational benefit that AI is delivering. The solution is to pair every speed metric with a quality metric that measures the same output from a different angle, and to weight the quality metric heavily enough that it cannot be ignored by officers and supervisors optimizing for speed alone.
The Paired Metric Principle
Every efficiency metric in a public-safety AI program should have a paired integrity metric that measures the quality of what the efficiency metric is counting. The pairs are:
Report submission time (efficiency) paired with QA accuracy rate (integrity). When both move favorably, submission is fast and reports are accurate. When submission time falls while accuracy rate falls, the program is producing fast errors. When submission time holds while accuracy rate rises, the program is producing better reports at the same speed. The paired view tells the story that neither metric tells alone.
Verification completion rate (efficiency proxy) paired with verification completion time analysis (integrity check). A 95% verification completion rate and an average completion time of 90 seconds per incident is a different program than a 95% completion rate with completion times that scale with incident complexity. The time analysis is not a separate metric system. It is an audit log query that the program manager can run as part of the monthly review and report to command as a signal of whether the verification pass is genuine.
Records released per month (efficiency) paired with redaction accuracy rate (integrity). The volume of records released under public-records statutes is a legitimate efficiency metric. Whether those records were correctly redacted before release is the integrity counterpart. Both numbers, reported together, tell the oversight body and the community whether the agency is clearing its backlog safely or clearing it carelessly.
AI draft utilization rate (activity) paired with officer correction rate (outcome). The percentage of eligible reports using AI assistance is an activity metric. The percentage of AI drafts that required at least one officer correction before submission is an outcome metric that provides indirect evidence of AI output quality and verification thoroughness. If the correction rate drops to near zero, one of two things is true: either the AI is producing very accurate drafts, or officers are no longer correcting them. The combination of other metrics (accuracy rates from QA, verification completion times) tells which interpretation is accurate.
Weighting and Reporting Structure
Paired metrics only work if the integrity metric carries real weight in the evaluation and reporting structure. If the efficiency metric is on the dashboard that supervisors see every morning and the integrity metric is in a monthly report that the program manager reviews, the effective weight is not equal. The integrity metrics must be surfaced to supervisors at the same frequency and in the same context as the efficiency metrics, or the incentive structure will default to valuing what is most visible.
The verification completion time analysis, the accuracy rate stratified by incident type, and the disclosure compliance rate should be on the same weekly supervisor dashboard as the submission time metric. If a supervisor sees that submission time is down and verification completion times are compressed, and if those two data points are in the same view, the supervisor has the information to investigate. If the submission time is in the daily dashboard and the verification completion time is in a quarterly report, the supervisor cannot act on information they are not seeing.
What Community Trust Actually Measures
Community trust in law enforcement AI is not measured by a single metric. It is the aggregate of many small signals that the community is receiving about how the agency is operating. A speed-only metrics program sends a specific signal: the agency is optimizing for throughput. A dual-axis metrics program with genuine integrity weight sends a different signal: the agency is optimizing for accuracy and accountability alongside throughput.
The community signal that erodes trust most quickly is not the AI error that reaches the news. It is the pattern that the AI error reveals: that the agency had a system for producing AI-assisted reports quickly, and that the system for checking those reports was not operating at the same quality level as the system for producing them. The King County prosecutor's ban sent a signal to the community that the prosecution's office was not confident in the accuracy or disclosure posture of the AI-assisted reports it was receiving. That signal reached the community directly, through news coverage, and it framed AI-assisted policing as a governance failure rather than an efficiency gain.
The agency that avoids this outcome is not the one with the best AI tool. It is the one whose metrics program would have caught the problem before the prosecutor did, whose audit log would have documented the verification passes that prevented errors from reaching the file, and whose disclosure compliance rate would have ensured the prosecutor and the defense both knew AI was involved from the moment the report was filed. That agency is not faster than the agency that skips verification. It is slower, by the amount of time the verification pass takes. And it is the one that keeps the trust it takes years to build.
Key Takeaways
- A metric that creates pressure to skip or compress the verification pass is not a productivity tool: it is a mechanism for producing sworn errors at scale, and the harm is institutional and constitutional rather than just operational.
- The four specific metric patterns that erode trust are: throughput without accuracy, speed without disclosure, activity without outcome, and averages that hide the distribution by blending high-stakes and low-stakes incident types.
- The verification pass cannot be meaningfully compressed: a genuine pass on a 45-minute incident takes real time, and a log entry showing 90-second completion is evidence that the pass was not genuine, with Brady and Giglio implications for the officer's entire report history under audit.
- Every efficiency metric needs a paired integrity metric reported at the same frequency and in the same supervisory view: report submission time paired with QA accuracy rate, verification completion rate paired with completion time analysis, records released paired with redaction accuracy, and AI draft utilization paired with officer correction rate.
- Averages across incident types are the most sophisticated failure mode: an 89% overall accuracy rate composed of 97% on property reports and 71% on use-of-force reports is hiding the problem exactly where it matters most, and stratification by incident type is the required fix.
- The King County prosecutor's ban and the EFF transparency concerns are the community-trust consequence of a speed-without-disclosure metrics program: the end state is not a faster-and-safer program, it is a moratorium and a conversation about whether the agency can be trusted to operate the tool at all.
- CJIS obligations and Brady and Giglio disclosure requirements stay with the agency regardless of speed gains; a speed metric that drives down disclosure compliance is trading a constitutional obligation for a dashboard number, and the trade does not hold up when the prosecutor or the defense examines the file.
- Community trust is the long-run outcome of many small metric signals: an agency that demonstrably measures accuracy and disclosure alongside speed is sending the signal that accountability is built into the system, which is the only signal that sustains trust through an AI-related incident.
Skill.re