โ†
AI for Public Safety & First Responders
Proficient ยท M15 ยท lesson 15 of 18 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Quality Metrics for AI-Assisted Reports
๐Ÿ“–
now learning

Quality Metrics for AI-Assisted Reports

15 min

Commander Patricia Osei had been fielding the same question from two different directions for six months. From the chief's office: "Is the AI saving us time?" From the county prosecutor's office: "Are the AI-assisted reports holding up?" She had a confident answer to the first question: the body-worn camera (BWC, the recording device worn on the officer's uniform) platform vendor had provided a dashboard showing an average 78 percent reduction in time-to-submit for reports generated with Axon Draft One. She had almost no answer to the second question. Her records management system (RMS, the digital system that stores, searches, and manages reports and case records) did not distinguish between AI-assisted and traditionally typed reports. Her supervisors were reviewing reports without a standard checklist tied to AI-specific risks. Her disclosure tracking was inconsistent. She knew, anecdotally, that the tool was saving real time. She had no systematic way to tell the prosecutor whether the reports it was producing were as accurate, complete, and legally defensible as the ones her officers had been typing by hand. The prosecutor was making a reasonable demand: prove it. Commander Osei had no data to answer with.

Why Quality Metrics Matter in an AI-Assisted Workflow

The time-saving argument for AI-assisted report writing is easy to make and easy to measure. An AI tool that produces an 80-word-per-minute first draft from audio, compared to an officer typing at 40 words per minute after a demanding call, generates a measurable reduction in time-to-submit that any supervisor can observe and any dashboard can display. Vendors like Axon, whose bundled contracts can reach approximately $45 million over ten years across cameras, drones, cloud storage, and AI tools, have strong commercial incentives to lead with the time-saving metric. Officers who spend 30 to 40 percent of every shift on paperwork have strong personal incentives to embrace it.

The quality argument is harder to make and harder to measure, but it is the argument that matters in the venues where police reports are actually tested: the deposition, the suppression motion, the civil rights claim, the administrative review, and the public-records request. An agency that can prove its AI-assisted reports save time but cannot prove they are accurate, complete, and legally defensible has made a compelling procurement case and an incomplete governance case. The prosecutor's office is asking the governance question, not the procurement question. So is the oversight board. So is the defense attorney who has just been handed a report that contains a phrase the body-camera footage does not support.

Quality metrics for AI-assisted reports are the systematic answer to the governance question. They are the data layer that sits between the vendor's time-saving dashboard and the agency's accountability obligations. They tell the agency not just how fast the reports are coming in, but whether the reports coming in are good enough to stand on under the evidentiary standard that police reports must meet.

Time-to-submit is a vendor metric. Accuracy, completeness, and legal defensibility are agency metrics. An agency that measures only the first has outsourced quality assurance to whoever challenges the report in court.

The Five Quality Dimensions

Quality in an AI-assisted report is multidimensional. An agency that measures it with a single metric, whether time-to-submit, supervisor approval rate, or correction frequency, is measuring a shadow of the real picture. A complete quality framework covers five dimensions: accuracy, completeness, verification integrity, disclosure compliance, and legal defensibility.

Dimension One: Accuracy

Accuracy is the foundational dimension: does the report describe what actually happened? In an AI-assisted workflow, accuracy has a specific sub-question that traditional report quality review does not surface: does the AI draft describe what actually happened, or does it describe what the model's training data suggests typically happens on calls of this type? The gap between those two descriptions is the gap-fill problem, and accuracy metrics must be designed to detect it.

Measuring accuracy in an AI-assisted workflow requires comparing the adopted report against the evidentiary record: the BWC footage, the computer-aided dispatch (CAD, the timestamped record of the call, assigned units, and dispatcher notes) entry, and the case file. A meaningful accuracy metric tracks the rate at which adopted reports contain claims that cannot be verified against the evidentiary record, the rate at which the original AI draft and the adopted report differ in factual claims (a proxy for how many errors were caught in verification), and the rate at which reports are flagged, returned, or challenged on accuracy grounds post-submission.

The base rate for corrections should be neither zero nor excessively high. A zero-correction rate on a large volume of complex reports suggests that verification passes are not being run, not that the AI is perfect. An excessively high correction rate, say, more than 30 to 40 percent of sentences requiring changes in a typical arrest report, may indicate that the prompt configuration or the audio quality feeding the model is poor, and that the agency is spending more time correcting than it is saving in drafting. The target range for a well-functioning pipeline is a moderate, consistent correction rate that reflects genuine verification, concentrated in the predictable gap-fill zones: use-of-force language, attributed statements, and detailed scene descriptions.

Dimension Two: Completeness

Completeness asks whether the report contains all the elements required by agency policy, state statute, and prosecutorial standards. A police report that is accurate in what it says but missing essential elements, the disposition of recovered property, the basis for probable cause, the specific force type used and its legal justification, is an incomplete report that may fail at prosecution even if every sentence it contains is factually grounded.

AI-assisted drafts can have a characteristic completeness problem that differs from the completeness failures of hand-typed reports. A hand-typed report is more likely to be missing elements because the officer ran out of time or energy at the end of a long shift. An AI-assisted draft may produce fluent, well-structured prose that covers the narrative arc of the incident while omitting elements the officer would have included because they were on the required-elements checklist: specific statutory citations, witness contact information in the correct format, documentation of property disposition, or the required language for specific charge types. The model is drafting from audio; it does not have access to the required-elements checklist unless that checklist is built into the prompt.

Completeness metrics should track required-element completion rates across AI-assisted reports, compared to a baseline from traditionally typed reports. If AI-assisted reports are systematically missing specific elements, the gap is a prompt engineering problem: the element is not being requested or required in the system prompt that drives the AI's output. Fixing it at the prompt level fixes it for every subsequent report, not just the one that was flagged.

Dimension Three: Verification Integrity

Verification integrity is the dimension that tracks whether the verification pass is actually being run, not just whether the submitted report is accurate. An agency could have high accuracy on submitted reports because its officers are running thorough verification passes, or because most calls are simple enough that the AI's draft happens to be accurate without a genuine pass. Those two situations look the same in the submitted report, but they are very different in terms of governance risk: the second situation means the agency is one complex, contested call away from an adopted report that contains a significant error the officer did not catch.

Verification integrity metrics track the process, not just the outcome. They include: the percentage of AI-assisted reports with a completed verification log (a critical metric for identifying officers who are submitting without completing the pass), the average duration of verification passes across report types and call complexities, the distribution of corrections across the three verification tiers (routine narrative, disputed content, use-of-force), and the supervisor's attestation that the verification documentation was reviewed before the report was finalized.

An agency that tracks verification integrity will find a distribution it does not expect: a subset of officers who run consistent, documented passes; a larger group whose compliance is inconsistent; and a tail of officers who are completing verification log fields without running a genuine pass (detectable because their correction rates are zero across all call types including complex arrests and use-of-force incidents). The tail is the governance risk. Identifying it requires measuring verification integrity, not just accuracy of submitted reports.

Dimension Four: Disclosure Compliance

Disclosure compliance tracks whether AI assistance is being disclosed in the report as agency policy and prosecutorial guidance require. The King County, Washington, prosecutor's office barred AI-written police reports specifically because disclosure was inconsistent: some reports indicated AI involvement, many did not, and the prosecutor's office could not tell which reports to scrutinize. A disclosure compliance metric that tracks the percentage of AI-assisted reports with a properly formatted disclosure notation, and that tracks the notation's content against the agency's standard, tells the agency at a glance whether it is meeting its transparency obligations or creating a discovery liability.

Disclosure compliance also connects to the Brady and Giglio framework. Brady v. Maryland (1963) requires the prosecution to disclose exculpatory evidence to the defense. Giglio v. United States (1972) extends disclosure to impeachment evidence about officers. If an agency's AI-assisted reports are being disclosed to the defense without indicating AI involvement, the defense has not received the full picture of how the report was generated, and the agency may be creating a retroactive disclosure problem in every case touched by those reports. Tracking disclosure compliance before case filings, not after, is the only way to close that problem systematically rather than case by case.

Legal defensibility is the lagging indicator: it measures how AI-assisted reports perform in the legal proceedings where they are tested. Tracking data in this dimension includes the rate at which AI-assisted reports are challenged in suppression motions, the outcomes of those motions, the rate at which AI-assisted reports are flagged during discovery as potentially problematic, the rate at which officers with AI-assisted reports require extended deposition preparation time related to the report's provenance, and any instances where a report's AI origin was raised by defense counsel and had a material effect on the proceeding.

Legal defensibility data is the hardest to collect because it requires tracking reports through the criminal justice system (Criminal Justice Information Services, or CJIS, governs the handling of this information, and the obligations stay with the agency), which can take months to years. But it is the most important long-term quality indicator, because it is the only metric that directly measures whether the agency's AI-assisted reports are meeting the evidentiary standard they are required to meet.

An agency that has been using AI-assisted reporting for eighteen months without tracking legal defensibility metrics has made a decision, possibly an unconscious one, to learn about quality failures from the defense bar rather than from its own quality assurance program. The first suppression motion that succeeds, or the first deposition at which an officer cannot answer basic questions about their verification process, will cost more in actual attorney hours, administrative time, and reputational damage than a year of systematic quality tracking would have required.

Building the Measurement Program

A quality metrics program for AI-assisted reports does not require a new system, a dedicated analyst, or a significant budget. It requires three things: a consistent data collection method, a reporting cadence, and a designated owner who reviews the data and acts on what it shows.

Data Collection and Sources

Most of the data needed for a quality metrics program already exists or can be produced with minor workflow additions. The RMS contains the submitted reports, the submission timestamps, and the supervisor review records. The BWC platform contains the original AI drafts (in systems that preserve them), the footage clips, and the association between clips and reports. The verification log, if the agency has standardized its format, is either in the RMS, in the BWC platform, or in a designated notes field. Corrections, if they are being logged, are in the officer's field notes or in a structured correction-tracking field.

The first measurement step is an audit of what data currently exists and what gaps need to be closed. An agency that is logging verification completion, preserving original drafts, and tracking corrections already has most of what it needs. An agency that is not yet doing those things must start building the foundation: standardize the verification log format, require correction logging, require disclosure notation, and configure the BWC platform to preserve original drafts. These are operational workflow decisions that cost time in the short term and produce the measurement substrate that quality assurance requires.

The second measurement step is defining the baseline. Before AI-assisted reports can be evaluated against a quality standard, the agency needs to know its quality baseline for traditionally typed reports: what is the current rate of required-element completion? What is the rate of supervisor returns for accuracy problems? What is the supervisor approval rate on first submission? These baselines are typically available in the RMS for the period before AI drafting was deployed. Comparing AI-assisted report quality against the pre-AI baseline is the most credible way to answer the prosecutor's question: "Are the AI-assisted reports as good as the ones your officers were typing by hand?"

Reporting Cadence and Benchmarks

Quality data that is collected but not reviewed on a regular cadence is not a quality program. It is a data archive. The reporting cadence should match the agency's decision-making cycle: a monthly summary for the commander level covering the five quality dimensions and trends over time, a weekly summary for the supervisor level covering verification compliance and correction rates by team, and immediate alerts for outlier events, such as an officer with a zero-correction rate across ten or more complex reports, a report that was returned multiple times for accuracy problems, or a suppression motion that cited the AI origin of the report.

Benchmarks give the data meaning. Some benchmarks can be drawn from the research literature: the 82 percent time reduction figure from Axon's testing is a public benchmark that agencies can compare their own time savings against. Others must be agency-specific: a target verification-log completion rate of 95 percent or higher for all AI-assisted reports, a target correction rate in the 10 to 25 percent range for complex arrest reports to signal active verification, a target disclosure compliance rate of 100 percent. The specific numbers matter less than the commitment to measure against them and act when the data diverges from the target.

Acting on What the Data Shows

A quality metrics program that produces a monthly report which the commander files without action is a reporting program, not a quality program. Quality improvement requires that the data drives decisions: decisions about officer training, decisions about prompt configuration, decisions about supervisor review checklists, and decisions about vendor contract terms.

When accuracy data shows a recurring gap-fill pattern in a specific call type, the response is a prompt engineering fix: modify the system prompt for that call type to require the model to flag when it is inferring rather than transcribing, and to suppress the boilerplate language that the model applies to fill the gap. This is a technical fix, not a training problem, and it should be documented and implemented by whoever manages the agency's AI tool configuration.

When verification integrity data shows a cohort of officers with consistently low correction rates across complex calls, the response is targeted retraining on the verification pass: why the gap-fill happens, what the three gap-fill signatures look like, and what the personal accountability consequences are of adopting an AI draft that contains an unsupported claim. The training is not punitive. It is the information the officer needs to understand that a zero-correction rate on a complex arrest report is evidence that the pass was not run, not evidence that the AI is perfect.

When legal defensibility data shows a specific report type or call type generating disproportionate legal challenge, the response is a workflow review: are the verification tiers being applied correctly for that call type? Is the disclosure notation in the right location and format? Is the audit trail complete? The review produces specific workflow adjustments that address the specific legal challenge, rather than a general exhortation to be more careful.

The Accountability Conversation

A complete quality metrics program, covering the five dimensions with consistent data, a regular reporting cadence, and documented responses to quality gaps, enables the accountability conversation that agencies with AI-assisted reporting programs will inevitably need to have: with prosecutors, with oversight boards, with city councils, and with the public.

The King County prosecutor's bar on AI-written reports was not a permanent regulatory posture. It was a response to a specific, demonstrable failure of transparency and quality assurance. An agency that approaches that prosecutor's office with a complete quality metrics program, demonstrating that its verification compliance rate is above 95 percent, that its disclosure notation appears in 100 percent of AI-assisted reports, that its use-of-force section correction rates demonstrate active verification, and that its reports have not generated a disproportionate rate of suppression motions, is having a different conversation than the agency that can only offer: "We reviewed them."

The Electronic Frontier Foundation's (EFF, a digital-rights organization) transparency concerns about AI policing records are, at their core, an accountability argument: agencies using AI to generate evidence documents should be able to demonstrate, specifically and publicly, what that use looks like and whether it is meeting the standards that evidence documents require. A quality metrics program does not require the agency to share its internal data publicly. But it does require the agency to have the data, which is the precondition for having the accountability conversation rather than avoiding it.

Community trust in AI-assisted policing is not separate from quality metrics. It is a downstream function of them. A community that is told "our AI-assisted reports are as accurate as our hand-typed ones" is receiving a claim. A community that can be shown "here is how we measure accuracy, here is what our verification compliance rate is, and here is how our AI-assisted reports have performed in the courts" is receiving evidence. The difference between those two conversations is the difference between an agency that adopted a technology and an agency that governs one.

Key Takeaways

  • Time-to-submit is a vendor metric that measures drafting speed. Accuracy, completeness, verification integrity, disclosure compliance, and legal defensibility are agency metrics that measure whether the AI-assisted reports are meeting the evidentiary standard they are required to meet. Both categories must be tracked.
  • Accuracy metrics in an AI-assisted workflow must surface gap-fill failures specifically: the rate at which adopted reports contain claims that cannot be verified against the BWC footage, the CAD entry, or the case file. A zero-correction rate on complex calls is a warning sign, not a success metric.
  • Completeness failures in AI-assisted drafts tend to differ from completeness failures in hand-typed reports. The AI produces fluent prose that covers the narrative arc while potentially omitting required statutory elements that the officer would have added from a checklist. Completeness metrics identify these gaps so they can be fixed at the prompt level.
  • Verification integrity metrics track the process, not just the outcome: whether the verification log was completed, the duration of the pass, the distribution of corrections across verification tiers, and the supervisor's attestation. These metrics identify the tail of officers who are submitting without running a genuine pass.
  • Disclosure compliance must reach 100 percent. The King County prosecutor's bar on AI-generated reports was a response to inconsistent disclosure. An agency that cannot demonstrate consistent disclosure compliance is creating a discovery liability in every case those reports touch.
  • Legal defensibility data, gathered through tracking suppression motions, discovery challenges, and deposition outcomes related to AI-generated reports, is the lagging indicator that tells the agency whether its pipeline is working in the venues that matter. Waiting to learn this from the defense bar is an expensive alternative to tracking it internally.
  • The quality metrics program requires consistent data collection, a regular reporting cadence, and documented responses to quality gaps. Data collected and filed without action is an archive, not a program. Each quality gap drives a specific response: prompt engineering for accuracy patterns, targeted retraining for verification integrity gaps, workflow adjustments for legal defensibility failures.
  • A complete quality metrics program is the foundation of the accountability conversation with prosecutors, oversight boards, and the public. Showing a prosecutor a 95 percent verification compliance rate, 100 percent disclosure notation coverage, and a clean suppression motion record is a different conversation than saying "our officers reviewed their reports." The data is what makes the accountability claim credible rather than asserted.