โ†
AI for Public Safety & First Responders
Strategic ยท M8 ยท lesson 8 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Evaluating Public-Safety AI Vendors
๐Ÿ“–
now learning

Evaluating Public-Safety AI Vendors

15 min

The demo lasted forty-five minutes. The vendor's sales team had flown in, the projector was working, and the system did everything they asked. It transcribed a body-worn camera (BWC, the recording device worn on an officer's uniform) clip in under a minute. It drafted a patrol narrative that looked professional. It flagged three items the sergeant had identified as potential Brady disclosures. The chief was impressed. The procurement officer was ready to recommend award. The city attorney was on her phone. Nobody in the room had asked the question that mattered most: what happens when this tool produces a wrong answer, and how will we know?

Why the Demo Is the Wrong Test

Vendor demonstrations for public-safety artificial intelligence are carefully constructed performances. That is not a criticism and it is not unique to this industry. Every software vendor shows their product at its best, on clean data, under favorable conditions, with a knowledgeable operator at the keyboard. The demo is designed to answer the question: can this system do impressive things? The answer is almost always yes. That is not the question a command staff officer, a procurement lead, or an agency attorney should be asking.

The questions that matter for public-safety AI are different in kind. They are questions about failure modes, accountability structures, and legal obligations. They are questions about what the system does when the footage is poor quality, when the incident is high-stakes, when the output is wrong, and when a defense attorney subpoenas the original AI draft. They are questions about what happens after the contract is signed and the vendor's sales team has moved on to the next agency.

Axon's Draft One, the most widely deployed AI report-drafting tool in American law enforcement as of 2026, drafts police report narratives directly from body-worn camera audio. Officers in test agencies reported an 82 percent decrease in report-writing time. That figure is real and significant. Officers currently spend roughly 30 to 40 percent of every shift on paperwork. Getting that time back is a genuine operational win. But the 82 percent figure is a benchmark from a controlled study, not a guarantee, and it says nothing about output accuracy, about what the system does with degraded audio, or about how errors surface and get corrected when the report is later read in court.

A procurement process that stops at "does it work in the demo" is a process that may select a tool based entirely on the vendor's best-case scenario. The evaluation framework in this lesson is designed to go further: to ask the questions that separate a vendor who has built accountability into the product from one who has not.

The demo shows you what the tool can do. The evaluation shows you what the tool does when it fails, and who is accountable when it does.

The Five Domains of Vendor Evaluation

A structured vendor evaluation for public-safety AI covers five distinct domains. Each domain has specific questions, and each question has a minimum acceptable answer. Vendors who cannot answer these questions clearly are telling you something important about their product and their organization.

Domain One: Accuracy and Verification Controls

The first and most important domain is accuracy and the controls the vendor has built to surface and correct inaccurate output. A police report is not a memo. It is evidence, disclosed to the defense, read in depositions, and tested at trial. Under Brady v. Maryland (the Supreme Court ruling that requires prosecutors to disclose exculpatory evidence to the defense), any material inaccuracy in an AI-drafted report that is not caught and corrected before submission could create a disclosure obligation or a suppression problem. Under Giglio v. United States (the ruling that requires disclosure of evidence that could impeach a witness's credibility, including the officer's), evidence that an officer routinely adopted AI-drafted reports without verifying them could be used to attack that officer's credibility in every subsequent case.

The questions to ask in this domain: What is the vendor's published accuracy rate, and on what data set was it measured? Does the accuracy rate vary by incident type, by audio quality, or by demographic of the incident participants? How does the system flag its own uncertainty? What does it do when the audio is degraded or the incident is complex? Does the system produce confidence scores or uncertainty indicators, or does it produce confident prose with no signal about reliability? What is the system's specific failure mode: does it produce plausible-sounding but wrong details (hallucinations in the AI literature, gap-fills in the lesson language of this program), or does it decline to produce output when uncertain?

The minimum acceptable answer: the vendor can describe their accuracy measurement methodology, can show you examples of system output on poor-quality audio, and has built a human review step into the workflow that is not optional. Any vendor who tells you the system is accurate enough that detailed human review is not required is describing a product that cannot be safely used in a law enforcement context.

Domain Two: Human Review Architecture

The second domain is the design of the human review step. This is where many public-safety AI deployments fail, not because the AI is unusually bad, but because the workflow around it has been designed to move fast rather than to move accurately. The EFF (Electronic Frontier Foundation, the civil-liberties organization that has raised transparency concerns about AI police reports) has specifically flagged the risk that speed incentives built into AI report tools can erode the review discipline that keeps those tools safe. When an officer is looking at a screen that shows the time remaining before the report is overdue, the review becomes a race rather than a verification.

The questions to ask: Does the workflow require the officer to explicitly verify each factual claim, or does it ask only for a general review and sign-off? Is there a documented verification protocol built into the platform, or is verification entirely up to the officer and not tracked? Does the platform log which elements were reviewed, what was changed, and by whom? Can the platform produce an audit trail that shows the original AI output, the officer's corrections, and the final submitted report? Is the review step timed in any way that could create speed pressure?

The minimum acceptable answer: the vendor has a documented human review protocol, that protocol is enforced in the platform (not just recommended in a training manual), and the platform produces an audit trail that can be produced in discovery. An agency that cannot show a defense attorney the specific review steps the officer took is in a weak position at the suppression hearing.

Domain Three: Disclosure and Transparency Features

The third domain is disclosure. The King County, Washington, prosecutor's office made national news when it issued a policy barring AI-written police reports from use in prosecution. The reasoning was clear: the prosecutor could not certify to the court what role AI had played in drafting the report, whether the output had been reviewed and to what standard, or whether the original AI draft was available for discovery. The King County policy is a teachable line. It shows exactly what a vendor evaluation must address.

The questions to ask: Does the platform automatically document that AI was used to draft the report, in a form that can be produced to the prosecutor and the defense? Does the disclosure documentation include the version of the AI model used, the timestamp, and the review steps? Can the original AI draft be retrieved after the officer has made corrections? Is the disclosure formatted to meet the requirements of the agency's local prosecutor's office? Has the vendor engaged with any prosecutor's offices to develop disclosure standards, and which offices?

The minimum acceptable answer: the platform produces automatic, retrievable documentation of AI assistance in a format that can be disclosed without additional work by the agency. A vendor who says "the officer can note in the report that AI was used" has not built a disclosure feature. They have passed the disclosure burden back to the officer and the agency.

Domain Four: Data Security and CJIS Compliance

The fourth domain is data security and compliance with the CJIS (Criminal Justice Information Services) Security Policy, the FBI's framework that governs how agencies handle criminal justice information. CJIS compliance is a non-negotiable requirement for any vendor handling data that touches a law enforcement investigation, and it has a specific and important feature: CJIS obligations stay with the agency, not the vendor. A vendor can certify that their systems meet CJIS technical requirements. That certification does not relieve the agency of its own CJIS obligations. The agency remains responsible for ensuring that data is handled correctly at every point in the pipeline, including within the vendor's systems.

The questions to ask: Is the vendor CJIS-compliant, and what specific CJIS controls have been audited? Where is agency data stored: in a dedicated instance, in a multi-tenant cloud environment, or in a vendor-managed shared environment? Is the vendor willing to submit to an agency-conducted CJIS audit, or only to provide self-certification? What is the data retention policy: how long is footage, audio, and draft text stored, and how is it deleted? Does the vendor have a data breach response plan, and does it meet the notification requirements applicable to criminal justice information?

The minimum acceptable answer: the vendor holds current CJIS compliance certification, submits to third-party audits, stores data in an environment the agency can inspect, has a documented data retention and deletion policy, and has a breach response plan. An agency that signs a contract with a vendor on the basis of self-certification of CJIS compliance, without auditing the vendor's controls, has transferred data handling responsibility in practice while retaining legal responsibility on paper. That is a gap that a single breach or audit finding will expose.

Domain Five: Contract Terms, Exit Rights, and Portability

The fifth domain is the contract itself: its duration, its termination provisions, its data portability guarantees, and the rights it reserves for the vendor versus the agency. This domain is addressed in depth in the next lesson (The 10-Year-Contract Trap), but vendor evaluation cannot be complete without raising the contract questions. The most capable tool in the world becomes a liability if the contract governing it traps the agency for a decade or gives the vendor rights over the agency's data that conflict with the agency's legal obligations.

The questions to ask at the evaluation stage: What is the contract term being offered, and is a shorter term available? What happens to the agency's data (footage, reports, AI drafts, audit trails) if the agency leaves the contract? Is the data returned in a format the agency can use, or is it locked in a proprietary format? Is the pricing fixed for the contract term, or can the vendor increase prices? Does the contract include a technology-change provision: what happens if the vendor's AI model changes in ways that affect output quality or accuracy during the contract term?

The minimum acceptable answer: the contract is available in terms shorter than five years, data portability is guaranteed in a usable format, pricing is fixed or capped, and the agency retains all data rights regardless of contract status. A vendor who will not discuss contract terms until after a procurement decision has been made is telling you something about the terms you will eventually be asked to accept.

The Questions Vendors Cannot Answer, and What They Mean

In practice, evaluating multiple vendors across these five domains will produce a clear pattern: some questions will be answered confidently and specifically, and others will be deflected, generalized, or answered with a redirect to the sales engineer. The deflections are informative. They are not always bad faith. They sometimes reflect genuine gaps in what the vendor has built. But they are gaps that an agency needs to understand before, not after, the contract is signed.

Common deflections and what they signal: "Our accuracy is industry-leading" without a methodology means the vendor has not measured accuracy in a way they are willing to be held to. "Human review is built into our recommended workflow" without platform enforcement means the vendor knows human review is necessary but has not made it mandatory in the product. "We are fully CJIS compliant" without an audit trail means the compliance may be self-certified. "Our contracts are standard" when asked about exit rights and data portability means the standard terms likely favor the vendor on those points.

These deflections are not disqualifying by themselves. They are the start of a negotiation. The agency's goal is not to find a vendor with perfect answers to every question. It is to understand exactly what each vendor has built, what they have not built, and what the agency will need to build itself to fill the gaps. An honest vendor with gaps the agency can fill is preferable to a vendor who claims to have no gaps and has not been tested.

Building a Structured Evaluation Process

A structured evaluation process creates a record that protects the agency regardless of the procurement outcome. It shows that the decision was made on documented criteria, not on the strength of a vendor's demo day presentation. It provides a baseline against which the vendor's actual performance can be measured after deployment. And it creates a paper trail that demonstrates due diligence if an AI-assisted report later becomes the subject of a legal challenge.

The Written Requirements Document

Before soliciting vendor responses, the agency should produce a written requirements document that specifies the evaluation criteria across the five domains. This document should include the minimum acceptable answers described above, the specific features the agency requires (audit trails, disclosure documentation, CJIS compliance certification, review protocol enforcement), and the contract terms the agency considers non-negotiable. The requirements document is not a wish list. It is the standard against which vendors will be scored, and vendors who do not meet the minimum should not advance in the evaluation regardless of how impressive their demo is.

The requirements document should be written before any vendor presentations. Writing it after seeing a demo invites the common procurement error of designing the requirements to fit the vendor rather than designing the vendor selection to fit the requirements. An agency that writes its AI requirements document after a single vendor has briefed command staff is not running an evaluation. It is conducting a post-hoc justification for a decision already made in the conference room.

Pilot Before Production

Every serious evaluation should include a pilot deployment before a production commitment. A pilot runs the vendor's tool on real incidents, with real officers, against real body-worn camera footage from the agency's own evidence archive. The pilot should be large enough to surface failure modes: at a minimum, it should include incidents with poor audio quality, use-of-force incidents, incidents involving ambiguous facts, and incidents that were later contested in court or subject to public-records requests.

The pilot should be evaluated by a panel that includes a patrol sergeant, a detective, a records supervisor, and the agency's legal counsel or city attorney. It should produce a written assessment of output accuracy on the pilot set, an assessment of the human review workflow in practice, and an assessment of the disclosure documentation the platform produces. The pilot assessment should be compared directly against the vendor's claims in the evaluation process. Gaps between vendor claims and pilot performance are the most important finding in any evaluation.

Vendors sometimes resist pilots, particularly pilots that involve use-of-force incidents or contested cases. That resistance is worth noting. A vendor who is confident in their product's performance on difficult cases should welcome the opportunity to demonstrate it. A vendor who seeks to limit the pilot to clean, routine calls is telling you something about how the tool performs on the calls that matter most.

Engaging the Prosecuting Authority Early

The King County experience demonstrates that the prosecutor's office is a critical stakeholder in any AI report-writing deployment. An agency that deploys AI-assisted reporting without engaging the local prosecutor's office in advance may find, as some agencies have, that the prosecutor will not use AI-drafted reports in court, or will use them only with specific disclosure conditions the agency did not anticipate. The right time to engage the prosecutor is before the contract is signed, not after the first case is charged.

The questions to bring to the prosecutor early: What disclosure language does the prosecutor require when AI assisted in drafting a report? What documentation of the review process does the prosecutor need to use an AI-assisted report in court? Does the prosecutor want access to the original AI draft, the correction log, or both? Are there case types (juvenile cases, cases involving confidential informants, homicide cases) for which the prosecutor will not accept AI-drafted reports regardless of the review standard? Getting these answers before procurement means the agency can evaluate whether a vendor's disclosure documentation matches what the local prosecutor will actually accept.

After the Award: Ongoing Evaluation

Evaluation does not end at contract award. A vendor who performs well in the pilot may perform differently at scale, with the full agency's incident volume, across a wider range of incident types, and under the time pressures of actual operations. An ongoing evaluation cadence ensures the agency does not discover performance gaps only when they surface in a courtroom.

A minimum ongoing evaluation cadence for a deployed public-safety AI tool includes: a quarterly review of a random sample of AI-drafted reports against the corresponding footage, a monthly review of correction logs to identify patterns in what officers are correcting (systematic corrections to the same types of claims are a signal of a systematic accuracy problem), an annual full-scale accuracy assessment similar to the pilot evaluation, and a standing mechanism for officers to flag output quality concerns without administrative penalty. The last item is often overlooked. Officers who are concerned that a tool is producing poor output need a channel to raise that concern that does not feel like a complaint against a tool the chief has publicly committed to.

The RMS (records management system, the platform that stores and manages police reports and case files) and CAD (computer-aided dispatch) integration also require ongoing attention. An AI tool that integrates cleanly with the RMS at deployment may create compatibility problems when either system is updated. Tracking integration-related errors separately from output-quality errors allows the agency to identify and address the right problem when things go wrong.

Key Takeaways

  • A vendor demonstration shows the tool at its best. The evaluation must go beyond the demo to ask what happens when the tool fails, who is accountable, and how the failure is detected and corrected before it reaches a sworn report.
  • The five evaluation domains are: accuracy and verification controls, human review architecture, disclosure and transparency features, data security and CJIS (Criminal Justice Information Services) compliance, and contract terms including exit rights and data portability.
  • Brady v. Maryland and Giglio v. United States frame AI-assisted reporting as a constitutional disclosure matter. A vendor's product must support the agency's ability to meet its Brady and Giglio obligations, not just produce fast drafts.
  • CJIS obligations stay with the agency, not the vendor. A vendor's CJIS certification does not relieve the agency of its own obligations. The agency must audit the vendor's controls rather than accepting self-certification.
  • The King County, Washington, prosecutor's policy barring AI-written police reports is a real and teachable governance line. Engaging the prosecuting authority before procurement allows the agency to select a vendor whose disclosure documentation the prosecutor will actually accept.
  • Every evaluation should include a pilot on real incidents, including use-of-force incidents and contested cases. A vendor who resists a pilot on difficult cases is signaling something important about how the tool performs on them.
  • The written requirements document should be produced before any vendor presentations. Requirements written after demos are more likely to describe the vendor already in the room than to specify what the agency actually needs.
  • Evaluation does not end at contract award. A quarterly and annual cadence of accuracy review, correction-log analysis, and officer feedback ensures the agency detects performance gaps before they surface in court rather than after.