Running Pilots to Evidence
The deputy chief handed the assistant chief a one-page results summary from the agency's six-month AI report-writing pilot and said: "Council wants to know whether we scale this." The assistant chief read the page and then asked the question that distinguished her from her predecessor in situations like this: "What was the evidence standard, and did we meet it?" The deputy chief did not have an answer. The pilot had tracked time saved. It had not tracked error rates, disclosure compliance rates, corrections made during verification, or the number of cases where the AI draft had to be substantially rewritten. The summary said officers were satisfied. It did not say anything a prosecutor, a civil-liberties advocate, or a council member skeptical of AI would find persuasive. The assistant chief sent the summary back and said: "Run it again, and this time tell me what we would need to see to be confident this is safe to scale."
What Running a Pilot to Evidence Means
In the context of agency AI programs, the term "pilot" is used loosely. Sometimes it means a vendor demonstration. Sometimes it means a trial deployment with no defined success criteria, no control condition, and no plan for how the results will inform the scale decision. Sometimes it means a limited rollout that was always going to become a full rollout regardless of the results, because the contract was already signed. None of these is a pilot in the sense that matters for executive governance.
A pilot to evidence, as the phrase is used in this lesson, is a bounded deployment with a defined scope, a pre-specified evidence standard, a data-collection plan, an independent review structure, and a scale decision that genuinely depends on whether the evidence standard is met. Every element of that definition matters. A pilot without a defined scope is a full deployment called a pilot. A pilot without a pre-specified evidence standard can only produce results the agency was already prepared to call a success. A pilot without independent review is a self-assessment. And a scale decision that was made before the evidence was collected is not a decision at all.
Running a pilot to evidence is demanding. It requires the agency to commit, in advance, to the conditions under which it would decline to scale. It requires data collection infrastructure that most agencies do not have in place. It requires an honest assessment process that may produce results the command staff does not want to share with the council or the community. These are real costs. They are also the cost of the credibility that comes from being able to say, truthfully, "we piloted this, we set a standard, the standard was met, and here is the evidence." That statement is worth more, before a council, a prosecutor, and an oversight board, than any vendor reference.
A pilot that was always going to succeed is not a pilot. It is a procurement with extra steps. The evidence standard must be set before the data comes in, and the scale decision must genuinely depend on it.
Designing the Pilot: Scope and Evidence Standard
The design of a pilot to evidence begins with two questions that must be answered before the pilot begins: what is the scope, and what is the evidence standard?
Scope
Scope defines the boundary of the pilot in terms of people, places, call types, time period, and the specific AI capability being tested. A well-scoped pilot is specific. It covers, for example, a defined set of patrol officers in a defined geographic area, handling a defined set of call types (excluding, perhaps, use-of-force incidents in a first phase), over a defined period of three to six months. It tests a specific AI capability, such as report drafting from body-worn camera (BWC) footage, not a bundle of capabilities simultaneously. It has a defined comparison condition: either a before-after comparison using the same officers, or a comparison between officers using the AI tool and a comparable group not using it, or both.
Scope also defines what the pilot does not test. If the pilot does not include use-of-force incidents, the evidence it produces does not support a conclusion about the safety of using the tool for use-of-force reports. If the pilot is conducted during a period of low call volume, the evidence it produces does not speak to performance during a high-volume period. An honest pilot report acknowledges the scope of the evidence produced. An agency that presents a narrowly scoped pilot as evidence of a broadly applicable conclusion has made a claim its data cannot support, and a sophisticated oversight board or civil-liberties group will notice.
Evidence Standard
The evidence standard defines what success looks like, before the data comes in. For an AI report-writing pilot, the evidence standard might include: a time-saving benchmark (for instance, a reduction in report-writing time of at least 25% compared to the pre-pilot baseline, not the vendor's claimed 82%, but a threshold that the agency has independently defined as meaningful given its staffing and workload context); an accuracy threshold (a rate of AI-draft errors that are caught and corrected during verification of no more than a defined number per 100 reports); a verification compliance rate (a percentage of AI-drafted reports for which the officer completed and documented a footage-grounded verification pass, as taught in this program); and a disclosure compliance rate (a percentage of AI-assisted reports for which the disclosure was properly documented per the agency's policy).
The evidence standard should also include conditions under which the pilot would be paused or terminated early: a specific error type that would trigger immediate review (for example, an AI draft that contained a factual claim contradicting the BWC footage in a use-of-force report that was submitted without correction), a threshold rate of disclosure non-compliance, or a complaint from the prosecutor's office about an AI-assisted case file. These are the pilot's kill criteria. They exist because a responsible pilot acknowledges that the AI may fail in ways that require a response, and that response should be defined in advance rather than improvised after the fact.
Data Collection for an Evidence-Grade Pilot
An evidence-grade pilot requires data that the agency may not currently be collecting. Identifying what data is needed, how it will be collected, and who will analyze it is part of the pilot design, not an afterthought.
Time Data
Time savings are the most commonly cited benefit of AI report-writing tools, and the 82% reduction figure from testing of Axon Draft One has been cited repeatedly. Measuring time savings requires a baseline: how long did it take these officers to write these types of reports before the AI tool was introduced? That baseline should be collected before the pilot begins, not reconstructed from memory afterward. Time data should capture both the AI-draft generation time (fast) and the verification-and-review time (the discipline this program teaches), because the total officer time invested is the relevant figure. An AI tool that generates a draft in two minutes but requires a thirty-minute verification pass to be used responsibly does not save the same time as one that generates a draft in two minutes and can be verified in five. The honest time measurement captures the complete workflow.
Accuracy Data
Accuracy data is harder to collect than time data, and it is more important. Accuracy means: how often does the AI draft contain a factual error, and what type of error is it? The categories that matter most for a public-safety context are the gap-fill error (the model invents a detail not supported by the footage), the softened-fact error (the model presents a contested or ambiguous fact more definitively than the footage supports), and the invented-quote error (the model attributes a statement to a person that the person did not say, or did not say in those words). Each of these is a potential Brady disclosure problem, as Brady v. Maryland requires the prosecution to disclose exculpatory evidence to the defense; each is a potential Giglio impeachment problem, as Giglio v. United States requires disclosure of information that could be used to impeach a witness including an officer.
Collecting accuracy data requires someone to compare the AI draft to the footage and flag errors, before the officer submits the report. This is a sampling exercise, not a comprehensive audit of every report in the pilot. A sample of 10 to 20% of pilot reports, reviewed by a trained reviewer (not the drafting officer) against the BWC footage and the CAD (computer-aided dispatch) entry, will produce a statistically usable error rate. That error rate, compared to the error rate for a comparable sample of hand-typed reports, is a genuine measure of whether the AI tool makes the agency's reports more or less accurate. If the error rate is lower for AI-assisted reports that went through the verification pass, that is a meaningful finding. If the error rate is comparable, or higher, that is a finding that must be disclosed and addressed before scaling.
Process Compliance Data
Process compliance data captures whether officers are doing the things the agency's AI policy requires: completing the verification pass, documenting corrections, and recording the disclosure statement. This data can be collected through periodic supervisory review of a sample of AI-assisted reports, through the audit trail that AI platforms like Axon's generate, or through self-report with periodic audits to check the self-report. The compliance rate is both an outcome measure (are officers using the tool correctly?) and a governance measure (does the agency's policy and training produce the behavior it intends?). A low compliance rate is not an officer problem until it has been confirmed that the policy is clear, the training is adequate, and the supervisory structure provides the right incentives. A persistent low compliance rate despite clear policy and adequate training is a system design problem, and scaling a system with a compliance problem is how the problem becomes a crisis.
The Independent Review Requirement
A pilot to evidence requires independent review, meaning review by someone who is not a member of the command staff and does not have a direct interest in the pilot's success. This is not a statement of distrust. It is a structural requirement for credibility. The agency's command staff, however honest, cannot be the sole evaluators of a program they commissioned, procured, and are invested in seeing succeed. The council, the oversight board, and the community know this. An independent review does not eliminate that tension. It addresses it.
Independent review can take several forms. It can be an external audit by a contractor with public safety AI expertise. It can be a review by the agency's legal counsel and the prosecutor's office together, focused on the Brady/Giglio implications and the disclosure record from the pilot. It can be a review by a civilian oversight body, with the pilot's data provided in a format the oversight body can analyze. It can be some combination of these. What it cannot be is a review conducted entirely by the people who designed and operated the pilot, even if those people are competent and honest.
The independent review should produce a written report that is shared with command, with the oversight body, and, in at least summary form, with the public. The report should address the evidence standard (was it met, and how was that determination made?), the failure modes observed during the pilot (including any early-termination triggers that were approached or reached), the compliance data, and a recommendation about scaling. An independent review that recommends against scaling, or recommends scaling with additional conditions, is a valuable output, not a failure. It means the pilot worked: it produced evidence, and the evidence informed the decision. A pilot where the independent review always recommends scaling because the incentives and the relationship structure make any other recommendation unlikely is not independent. It is a ratification process.
What a Pilot Failure Teaches and What It Requires
A pilot that fails to meet its evidence standard, or that produces evidence strong enough to trigger a pause or termination, is a difficult outcome for an executive. It means telling the council that the agency spent resources on a pilot and did not advance to scale. It means telling the vendor that the trial did not support a full deployment. It may mean telling the workforce that a tool many officers were enthusiastic about will not be deployed, or will require additional conditions before deployment. None of these conversations is easy. All of them are the right conversation.
An agency that pauses or terminates a pilot because the evidence standard was not met has demonstrated the governance discipline that builds long-term credibility. It has shown the council, the oversight board, and the community that the pre-specified standard was real, not performative, and that the agency is not in the business of deploying technology that has not been shown to be safe and effective. That demonstration is an asset. It makes it easier to bring the next pilot to the council with genuine credibility rather than with the implied question "will you also approve this one before the evidence is in?"
A pilot failure also teaches specific lessons about why the tool or the deployment conditions did not produce the expected results. Was the time savings lower than expected because the verification pass took longer than estimated? That is a workflow design lesson. Was the accuracy lower than expected because officers were not completing verification passes? That is a training and compliance lesson. Was there a specific call type or incident type where the error rate was significantly higher? That is a scope lesson, and it may mean that the tool can be deployed with a narrower scope, excluding the high-error call types, rather than not deployed at all. A pilot that fails at the aggregate level may succeed at a more refined scope. The evidence from the failed pilot makes that refinement possible.
Disclosing Pilot Results
Pilot results, including results that are unfavorable or mixed, should be disclosed. The disclosure obligation here is not a legal one in most jurisdictions, though some public records laws may reach pilot documentation. It is a governance obligation that follows from the agency's commitment to transparency with the oversight body and the community. An agency that only reports favorable pilot results, and quietly shelves the unfavorable ones, is not operating a responsible-innovation program. It is cherry-picking evidence to support predetermined decisions. That pattern, if it becomes visible, destroys the credibility that responsible piloting is designed to build.
The disclosure of pilot results should include the evidence standard that was used, the data that was collected, the comparison condition, the error rate, the compliance rate, and the recommendation from the independent review. It should acknowledge the limitations of the pilot: the scope, the time period, the conditions under which it was conducted, and what additional evidence would be needed to extend the findings to a broader deployment or a different use case. A transparent pilot disclosure is a governance document. It tells the community and the oversight body what the agency looked for, what it found, and what it concluded. That is how trust is built: not through the technology, but through the honesty of the process around it.
Key Takeaways
- A pilot to evidence is a bounded deployment with a defined scope, a pre-specified evidence standard, an independent review structure, and a scale decision that genuinely depends on whether the evidence standard is met. A pilot without all four elements is not a pilot in the governance sense.
- The evidence standard must be set before the data comes in. Pre-specifying the threshold for success, and the conditions that would trigger early termination, is what separates a genuine pilot from a ratification process for a decision already made.
- Time savings are the most commonly cited pilot metric but are not the most important one for governance purposes. Accuracy (gap-fill error rates, softened-fact rates, invented-quote rates) and process compliance (verification pass completion, disclosure documentation) are the metrics that matter for Brady/Giglio credibility and oversight-body confidence.
- Collecting accuracy data requires sampling: comparing the AI draft to the BWC footage and CAD entry for a sample of pilot reports, reviewed by someone other than the drafting officer. A 10 to 20% sample is statistically sufficient and practically manageable.
- Independent review means review by someone without a direct interest in the pilot's success. External audit, prosecutor review, or civilian oversight body review each qualify. Internal-only review by command staff does not, regardless of the reviewers' competence or honesty.
- A pilot that fails its evidence standard is a governance success: it demonstrates that the standard was real, not performative, and builds the credibility that makes the next pilot more trustworthy. Failure data also teaches specific lessons about why and where the tool fell short, enabling refined future deployments.
- Pilot results, including unfavorable or mixed results, must be disclosed to the oversight body and, in summary, to the public. Selective disclosure of favorable results destroys the credibility that responsible piloting is designed to build.
- CJIS (Criminal Justice Information Services) compliance, Brady v. Maryland disclosure obligations, and Giglio v. United States impeachment-evidence obligations apply to AI pilots just as they apply to full deployments. The legal review should happen before the pilot begins, with prosecutor involvement, not after the pilot produces a scale recommendation.
Skill.re