โ†
AI for Public Safety & First Responders
Visionary ยท M7 ยท lesson 7 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Measuring Transformation Agency-Wide
๐Ÿ“–
now learning

Measuring Transformation Agency-Wide

15 min

The compstat meeting had been running for forty minutes when the assistant chief pulled up the AI performance dashboard. Report completion time was down 71 percent. Officer adoption was at 91 percent. Cost per report had dropped to $14 from $52. The numbers were good and she knew it. What she did not know, and what the dashboard did not show, was that two weeks earlier the agency's prosecutor liaison had received the first of what would become seven consecutive suppression motions from the same defense firm, all arguing that AI-generated narratives in the agency's reports contained unsupported factual claims. The efficiency story on the screen was real. The integrity story it was hiding was also real. And the community trust story, the fact that a local civil rights organization had filed a public records request for every AI-assisted report filed in the previous twelve months, had not been mentioned at compstat at all.

Why a Single-Axis Dashboard Fails

Every public safety agency deploying AI faces a measurement temptation: speed is easy to measure, speed improvement is easy to celebrate, and speed improvement looks good in a city council presentation. The problem is not that speed matters. It does. Officers spending 30 to 40 percent of every shift on paperwork is a documented productivity problem with a real dollar cost. Axon Draft One (the AI product that drafts police narratives from body-worn camera, or BWC, audio and has been reported to reduce report-writing time by 82 percent in officer testing) is producing real time savings in the agencies that have deployed it. Those savings deserve to be measured and reported honestly.

The problem is what a speed-only dashboard does not show. It does not show whether officers are running the footage-grounded verification pass before adopting the AI draft as their sworn account. It does not show whether the gap-fill rate, the fraction of AI-generated narratives containing unsupported factual claims that the AI produced to complete incomplete audio, is growing, shrinking, or holding steady. It does not show whether the prosecutor's office is beginning to push back on reports. It does not show whether the community's trust in the agency's documentation process is holding or eroding. And it does not show whether the agency's disclosure obligations under Brady v. Maryland (the 1963 Supreme Court case requiring disclosure of exculpatory evidence to the defense) and Giglio v. United States (the 1972 case requiring disclosure of impeachment evidence including information affecting officer credibility) are being met consistently.

An agency that manages its AI program on the efficiency axis alone is managing the axis it controls most easily and the axis that most visibly justifies the investment. It is not managing the program. A program managed only on the efficiency axis will eventually produce the integrity failure or the trust collapse that forces a reactive response, usually at the worst possible time, usually in a council chamber or a courtroom, usually after the damage has already been done.

An AI program that improves speed while degrading integrity has not improved anything. It has made the agency faster at producing unreliable evidence.

The Three-Axis Scorecard

A complete measurement framework for an agency-wide AI transformation tracks three axes: efficiency, integrity, and trust. Each axis has specific metrics. Each axis has specific thresholds that trigger intervention. And the three axes must be reported together, at the same cadence, to the same audience, so that leadership sees the complete performance picture rather than the most favorable slice of it.

The Efficiency Axis

The efficiency axis measures what AI deployment returned to the agency in terms of time, cost, and capacity. The core metrics are: average report completion time before and after AI adoption; officer time returned to patrol per shift per officer; report backlog reduction (the number of reports pending beyond the agency's standard submission window); and overtime attributable to documentation burden. These metrics answer the question "Did we get the time back?"

The efficiency metrics must be reported with the adoption rate: what fraction of officers and dispatchers are using the AI-assisted workflow on what fraction of applicable incidents. A 71 percent reduction in report completion time from a 40 percent adoption rate is a very different program than a 71 percent reduction from a 91 percent adoption rate. The numerator and the denominator both matter.

The efficiency axis also includes cost metrics: cost per report before and after AI adoption, total AI program cost as a fraction of the annual documentation burden cost (using the loaded labor cost the agency established in its investment case), and the annualized return on investment, including the risk adjustment component. An efficiency dashboard that does not show the total program cost alongside the time savings is a dashboard that tells the council the savings without the investment, which is not a complete financial picture.

The Integrity Axis

The integrity axis is the most important axis and the hardest to measure. It answers the question "Is the AI-assisted documentation meeting the evidentiary standard the program was designed to meet?" The core integrity metrics are:

The verification compliance rate: what fraction of AI-assisted reports can be confirmed, through the audit trail, to have received the footage-grounded verification pass before the officer adopted the draft as their sworn account. A verification compliance rate below 100 percent is not a target to celebrate. It is a number to understand and close, because every AI-assisted report that reached submission without verification is a report that may contain an uncorrected gap-fill in the sworn evidentiary record. The CJIS (Criminal Justice Information Services) Security Policy obligations, which govern how criminal justice information is handled and stored, sit with the agency. So does the evidentiary standard. Neither the vendor's accuracy claim nor the aggregate efficiency improvement is a substitute for a 100 percent verification compliance rate.

The gap-fill detection rate: what fraction of AI-generated narratives contained at least one AI-produced claim that was not supported by the BWC footage, the CAD (computer-aided dispatch, the system that logs dispatch events) entry, or field notes, and that the officer identified and corrected during the verification pass. A higher gap-fill detection rate, counterintuitively, is the better outcome: it means officers are finding and correcting errors before they reach the sworn record. A declining gap-fill detection rate in a mature program might indicate that verification is becoming more cursory over time, officers are becoming complacent, or the AI model is improving. The interpretation requires investigation.

The disclosure compliance rate: what fraction of applicable reports, meaning reports where AI was used in a manner that requires disclosure under the agency's disclosure policy and the prosecutor's accepted format, received the required disclosure within the required window. This metric directly addresses the Brady and Giglio obligations. A disclosure compliance rate below 100 percent is a legal compliance problem, not a performance benchmark to optimize. The prosecutor's office should receive this metric at the same cadence the chief does.

The suppression motion rate: how many suppression motions in the reporting period cited AI-related grounds, and what was the outcome. A single sustained suppression motion on AI grounds is not a data point to bury. It is an early warning signal that deserves immediate investigation of the verification and disclosure practices in the affected unit.

The error correction log: a systematic record of AI-generated claims that officers identified as incorrect during the verification pass, categorized by error type (gap-fill detail, softened fact, invented quote, scene description without footage support, use-of-force boilerplate), unit, and incident type. This log, reviewed monthly, is the earliest indicator the agency has of whether the AI model is producing more or fewer errors over time, and of whether certain incident types or units have higher gap-fill rates than others.

The Trust Axis

The trust axis measures whether the agency's AI program is producing or consuming community trust. It is the axis most easily dismissed as soft or unmeasurable. It is not soft and it is measurable. The Electronic Frontier Foundation (EFF, the civil liberties organization that has raised transparency concerns about AI police reports) and community oversight bodies have made clear that opacity about AI in police reports is a trust problem. The agency that treats community trust as a lagging indicator that only becomes visible after a crisis has not been managing it. It has been ignoring it.

The trust axis metrics include: the community complaint rate specifically citing AI-related concerns in police reports or records; public records request volume for AI-assisted reports, which is a leading indicator of community concern even before formal complaints; civilian oversight body findings related to the AI program; and a periodic community engagement survey that asks specifically about awareness of and confidence in the agency's AI program safeguards.

The King County, Washington, prosecutor's decision to bar AI-written police reports is the institutional trust failure that happens when an agency has not managed the trust axis. The King County bar did not arise in a vacuum. It arose because the prosecutor's office concluded that the AI-assisted reports it was receiving did not meet the disclosure and verification standard required for constitutional discovery. That is a trust collapse between the agency and its primary partner in the justice system. A trust axis metric that tracks prosecutor satisfaction with the agency's disclosure format, measured at least quarterly, is an early warning system for the King County problem.

Cadence, Audience, and Accountability

A measurement framework without a reporting cadence, a designated audience, and an accountability mechanism is not a measurement framework. It is a data collection exercise. The three-axis scorecard must be presented to specific audiences at specific intervals, with specific accountability for metrics that fall below threshold.

The recommended cadence is: a weekly operational metric review for the agency AI lead and the disclosure coordinator, covering verification compliance rate and disclosure compliance rate; a monthly performance review for command staff, covering all three axes with trend analysis; a quarterly governance review for the chief and the prosecutor's liaison, covering the integrity and trust axes with specific attention to suppression motions and disclosure findings; and a semi-annual community report, in plain language, covering the trust axis metrics and an honest account of the program's performance on all three axes.

The semi-annual community report is the EFF transparency answer. It is also the civilian oversight board's working document for their review of the program. It should be published proactively, on the agency's public website, without being prompted by a public records request. An agency that publishes its AI performance data proactively has converted the EFF's transparency concern into a program feature. An agency that makes community members or civil liberties organizations file records requests to get the same information has not addressed the concern; it has confirmed it.

The Intervention Threshold

Each metric on the three-axis scorecard should have a documented intervention threshold: the metric value at which the agency AI lead is required to escalate to command, and the metric value at which the chief is required to brief the council and the civilian oversight body. Intervention thresholds make the measurement framework actionable rather than decorative.

For the verification compliance rate, the intervention threshold might be: below 95 percent triggers a unit-level review by the agency AI lead; below 90 percent triggers a command-level review and a temporary enhanced verification requirement; below 85 percent triggers a chief-level briefing and a program pause for the affected unit pending remediation. These are specific thresholds with specific consequences. They are not policies that can be quietly ignored when the dashboard shows a number nobody wants to escalate.

The suppression motion rate is a zero-tolerance metric: any AI-related suppression motion that is sustained triggers an immediate review of the affected report, the verification log, the disclosure documentation, and the officer's and supervisor's records on verification compliance. A sustained suppression motion is not an isolated incident. It is evidence that the verification or disclosure standard failed in at least one specific case, and the investigation of that case is the agency's opportunity to understand whether the failure was individual, unit-level, or systemic.

Telling the Transformation Story With Data

The three-axis scorecard is the data foundation for the three-audience alignment the agency's transformation requires. Command staff sees the full scorecard and is accountable for the integrity and trust metrics in their units. The city council sees the efficiency story alongside the integrity and trust story, in accessible language, in the same document. The community sees the plain-language version that gives them factual information about what the program is doing, what it is finding, and what the agency is doing about the findings.

The transformation story told with three-axis data is fundamentally more credible than the transformation story told with efficiency data alone. A chief who stands before the council and says "report time is down 71 percent and our disclosure compliance rate is 98.7 percent and our gap-fill detection rate shows officers are catching errors at a 94 percent rate" is telling a complete story that can be scrutinized and is more defensible under scrutiny. A chief who says "report time is down 71 percent" is telling half a story that will be challenged as soon as a defense attorney files a suppression motion.

The transformation is not complete because the efficiency numbers look good. The transformation is complete when efficiency, integrity, and trust are all moving in the right direction, when the program is producing time savings and maintaining evidentiary quality and earning community confidence, and when the data that demonstrates all three is available to every audience that has a right to see it. That is the measure of an agency that has not just deployed AI. It has transformed.

When a Metric Collapses

A metric collapse, meaning a rapid deterioration in one of the three axes, is the most important situation the measurement framework must be designed to surface quickly. The assistant chief's compstat scenario in this lesson's opening is a metric collapse in slow motion: the efficiency metrics are high, but the integrity metrics (seven consecutive suppression motions) and the trust metrics (the civil rights organization's records request) are collapsing at the same time and are invisible to the leadership team because they are not on the dashboard.

A metric collapse in the integrity axis typically has a leading indicator: the verification compliance rate starts declining before the suppression motions arrive. Officers become more familiar with the tool and begin to treat the verification pass as optional on routine calls. The gap-fill detection rate may decline before error rates in submitted reports increase, because declining detection means errors are reaching the sworn record uncorrected. The error correction log's monthly review is the earliest warning signal, and it is only useful if someone is assigned to review it and escalate what they find.

A metric collapse in the trust axis typically has a leading indicator: the public records request volume for AI-assisted reports increases before the formal complaint arrives. An uptick in records requests specifically for AI-assisted reports, or a request from a known civil rights organization for all AI-assisted reports in a specified period, is a signal that requires a proactive response, not a passive records production. The proactive response is: review the disclosure compliance rate for the period covered by the request, review the verification compliance rate, and if both are strong, proactively share that data with the requesting organization before they have to file a follow-up request. If either is not strong, the agency needs to address the gap before the organization publishes its findings.

Key Takeaways

  • A single-axis dashboard that measures only efficiency will eventually produce the integrity failure or trust collapse it was not designed to surface. An AI program that improves speed while degrading integrity has not improved anything; it has made the agency faster at producing unreliable evidence.
  • The three-axis scorecard tracks efficiency (time returned, cost per report, adoption rate, ROI), integrity (verification compliance rate, gap-fill detection rate, disclosure compliance rate, suppression motion rate, error correction log), and trust (community complaint rate, public records request volume, oversight findings, community survey, prosecutor satisfaction).
  • The verification compliance rate and the disclosure compliance rate are not targets to optimize. They are compliance floors. The verification compliance rate below 100 percent is a number to understand and close, because every unverified report may contain an uncorrected gap-fill in the sworn evidentiary record.
  • The King County, Washington, prosecutor's bar on AI-written reports is the institutional trust failure the trust axis is designed to prevent proactively. A quarterly metric tracking prosecutor satisfaction with disclosure is an early warning system for this specific failure mode.
  • Intervention thresholds make the measurement framework actionable. Specific metric values must trigger specific actions: a unit-level review, a command briefing, a program pause, or a chief-level report to the council and oversight body.
  • The semi-annual community report published proactively on the agency's public website, covering all three axes in plain language, is the EFF transparency answer and the civilian oversight board's working document. It converts the transparency concern from a vulnerability into a program strength.
  • Metric collapses have leading indicators: the verification compliance rate declining before suppression motions arrive, and the public records request volume increasing before formal complaints. The measurement framework must be designed to surface leading indicators, not just outcomes, so the agency can intervene before the damage is done.
  • The transformation is complete not when the efficiency numbers are good, but when efficiency, integrity, and trust are all moving in the right direction and the data demonstrating all three is available to every audience that has a right to see it.