AI Hallucinations in a Report Context
Officer Diallo had nine years on the job and a reputation for detailed, accurate reports. She had been using her department's AI drafting tool for four months, and she had found it genuinely useful: what used to take forty-five minutes now took about eight. On the night of the Mercer Street incident, a domestic disturbance that escalated to an assault, she processed the scene, secured the suspect, and completed her report using the AI draft as her base. She read through it, corrected a spelling, adjusted the timeline by two minutes, and submitted it. Three months later, in the discovery phase before trial, the defense attorney pulled the body-worn camera (BWC, the recording device officers wear on their uniform to document encounters) footage and found three things in Officer Diallo's report that were not in the footage: a specific statement attributed to the suspect ("I didn't touch her"), a description of a physical object in the suspect's hand that the footage showed was not present, and a sequence placing the victim at the kitchen doorway before Officer Diallo had even entered the apartment. None of these details were fabrications in the sense of deliberate falsehood. Officer Diallo had not invented them. The AI had. And because Officer Diallo had been using the tool for four months with generally excellent results, she had trusted the draft enough to accept it without checking each claim against the footage. Those three details became the foundation of the defense's Brady motion, a suppression hearing, and a case that eventually ended with the charges dismissed. The model did not know it was wrong. That is the entire problem.
What a Hallucination Is in a Report Context
A hallucination, in the technical sense used for AI language models, is output that is internally coherent, grammatically correct, and confident in tone but factually unsupported, wrong, or fabricated. The term can sound exotic, but the mechanism is straightforward. Language models work by predicting the next most-probable sequence of words given everything they have seen: their training data, the system instructions, and the input document or transcript. They do not retrieve facts and verify them. They generate language that is statistically consistent with patterns learned during training.
When Axon Draft One or any similar tool transcribes BWC audio and generates a narrative, it is doing something like pattern completion: given this transcript, this call type, and these incident characteristics, what is the most probable narrative? The model has seen thousands of domestic disturbance reports. It knows how they are structured, what observations typically appear, how officers describe subjects and sequences of events. When the audio transcript is incomplete, unclear, or missing an element that normally appears in a domestic disturbance report, the model tends to supply the missing element from its training data, because that is what pattern completion does. It fills gaps with what fits.
In a word-processing document, a fabricated sentence is a typo or a factual error. In a police report, a fabricated sentence is evidence. It is disclosed to the defense. It can be cross-examined. It can be used to impeach the officer's credibility as a reporter of facts under Giglio v. United States (the 1972 Supreme Court case requiring disclosure of impeachment evidence against government witnesses). And if the fabricated detail would have been favorable to the defense in its accurate form, it is a Brady v. Maryland (the 1963 Supreme Court case requiring disclosure of evidence favorable to the defense) violation. The same behavior of the model that makes it useful, confident pattern completion, is the behavior that makes its failure modes legally dangerous.
The Three Failure Modes
In the context of AI-assisted police report drafting, hallucinations tend to fall into three categories. Understanding each one concretely is the foundation of the verification discipline that the next lesson in this chapter makes explicit.
The Gap-Fill Detail
The gap-fill detail is the most common and the most dangerous. It occurs when the BWC audio transcript is incomplete or silent on a point that a report would normally address, and the model fills the gap with a detail that is statistically plausible given the incident type but not supported by the footage.
Consider a domestic disturbance call. A complete report of such a call typically includes the suspect's physical position when officers arrived, any objects in the subject's immediate environment, any statements made by the subject or the victim, and the sequence of events leading to the moment of arrest. If the BWC audio captured only fragments of what occurred (background noise, overlapping voices, the officer speaking into the radio), the transcript will have gaps. The AI, trained on thousands of domestic disturbance reports, knows what details are typically present. It fills those gaps from its training pattern. The result is a report that reads smoothly and completely but contains details the footage does not support.
The gap-fill detail is invisible without verification. The report does not announce which sentences came from the audio transcript and which came from the model's pattern-completion function. Every sentence reads the same way: confidently, in the officer's professional register, as if it were observed fact. The officer who reads through the draft without checking each factual claim against the BWC footage will not see the gap-fill detail. It looks right. It fits the incident. And in Officer Diallo's case, the physical object that appeared in the suspect's hand in the AI draft, but not in the footage, was the detail the defense used to argue the report was unreliable.
The Softened Fact
The softened fact is the failure mode where the AI accurately captures an event but then subtly modifies the characterization in a direction that makes the event less legally significant. This can happen in either direction: the AI may soften an aggressive action (describing a punch as a push, an explicit threat as an angry statement, a controlled substance as "an unknown substance"), or it may escalate a neutral action (describing a stopped subject as "appearing to flee," a raised voice as "shouting," a casual gesture as "reaching aggressively").
The softened fact is particularly dangerous because it is harder to detect than the gap-fill detail. The gap-fill detail involves an event or object that is simply not in the footage; checking the footage will reveal its absence. The softened fact involves an event that is in the footage, but with a characterization that does not accurately reflect the legal or factual significance of what was observed. An officer reviewing the draft may see the event on the footage and confirm that yes, there was a physical altercation, without catching that the model described a felony-level contact as something that sounds like a misdemeanor.
The softened fact is a Brady problem in either direction. If the model softened a violent action in a way that understates the crime, the report misrepresents the severity of the incident. If the model escalated a neutral action in a way that makes the subject appear more threatening than they were, the report contains a characterization the defense will challenge, the footage will not support, and the officer will not be able to defend under cross-examination. In both cases, the accurate version of the event, shown in the footage, is material to the case.
The Invented Quote
The invented quote is the most legally serious hallucination failure mode in a police report context. It occurs when the AI attributes a specific statement to a subject, a victim, or a witness that does not appear in the audio transcript. The statement may be plausible given the incident, may fit the subject's apparent demeanor, and may match the kind of statement that typically appears in similar reports. But if the subject did not say it, or if what they said was materially different, an attributed quote in a sworn report is a fabrication that the BWC audio will disprove.
In Officer Diallo's case, the attributed statement ("I didn't touch her") appeared nowhere in the BWC audio transcript. The AI generated it because attributed denials by suspects are statistically common in domestic disturbance reports that proceed to arrest. The model has seen that pattern enough times to treat it as something that belongs in this type of report. The actual statement the suspect made, captured clearly on the audio, was different in a way that was material to the defense: the suspect had said nothing at all at the moment the report placed the denial, and had then made an admission later in the encounter that the report did not include. The AI had invented the denial and omitted the admission.
An invented quote in a sworn report is a Brady violation if the actual quote would have been favorable to the defense. It is a Giglio problem because it calls the officer's credibility as a reporter of facts into question. And in a case where the charge depends on the subject's state of mind or intent, an invented or incorrect attributed statement is not a footnote; it is potentially the difference between conviction and acquittal.
Why the Model Cannot Warn You
One of the most important things to understand about AI hallucinations in a report context is that the model produces them with the same confident tone as accurate output. There is no asterisk, no confidence score displayed to the officer, no sentence that says "this detail was not in the audio transcript." The model's output is presented as a unified narrative, and the gap-fill detail looks exactly like the detail that was transcribed from the footage.
This is not a fixable user-interface problem. It reflects something fundamental about how language models work. The model does not have a separate "I am sure about this" and "I am guessing about this" output state. It generates the most probable next token given its context. When it fills a gap from training patterns, that is, from its prior knowledge of what domestic disturbance reports look like, it is doing exactly what it was designed to do. The confident tone is not overconfidence in the usual sense; it is just the way language model output is structured.
The officer is the error-detection mechanism. This is not a weakness of any particular tool; it is the defining characteristic of the officer-AI verification relationship in report drafting. The AI generates a draft. The officer verifies the draft against the record (the footage, the CAD (computer-aided dispatch, the system dispatchers use to log calls and track incident details), the actual observations made at scene). That verification step is where the gap-fill details, the softened facts, and the invented quotes are caught or missed. There is no substitute for it. There is no AI tool that checks another AI tool's output reliably in this context. The footage and the officer's direct knowledge of the incident are the only ground truth available.
Hallucination Frequency: What the Evidence Says
A question officers often ask at this point is: how common are these failures? If the error rate is very low, is the verification burden proportionate?
The honest answer is that the error rate in AI-assisted police report drafting is not well-characterized in peer-reviewed literature as of mid-2026. The field is moving faster than the research. The Axon Draft One deployment figures (82 percent reduction in report-writing time) come from officer self-reporting, not from systematic error-rate studies. Vendors have not published failure-mode frequency data that would allow a precise answer. What we know from the broader AI literature on language model hallucinations is that the rate is not zero, it varies by task type, it varies by input quality (a clean audio transcript produces fewer gap-fills than a noisy one), and it varies by how familiar the incident type is to the model's training distribution.
What we also know from the legal and professional context is that the frequency of errors is not the primary consideration for verification discipline. A police report is a legal document that may determine whether someone goes to prison. The standard for verification is not "errors are rare enough to accept without checking." The standard is "verify because a defense attorney, a prosecutor, and a judge will read this." At that standard, a one-percent error rate distributed across thousands of reports per year, in an agency with three hundred officers, means dozens of reports per year containing unsupported details that could trigger Brady motions, suppression hearings, or dismissed charges. The math does not require a high error rate to make verification a professional obligation.
The Brady Problem and the Giglio Problem Made Concrete
It is worth grounding the abstract legal standards in concrete form, because the practical consequences of an AI hallucination in a police report are not abstract.
Brady violation from an invented quote: The suspect in a robbery case allegedly said "I don't have anything on me" when the officer asked about weapons. This statement, if true, is relevant to whether the officer's subsequent search was lawful. The AI draft included this statement. The BWC audio shows the suspect said nothing about weapons at that point in the encounter. The actual statement ("I don't have anything on me") never occurred. The fabricated statement allowed the officer to document a consent predicate for a search that the footage suggests was not consented to in those terms. The report goes to the prosecutor. The prosecutor, relying on the report, does not disclose the BWC footage to the defense before trial. The defense, having obtained the footage independently through a public records request, presents the discrepancy at trial. The result is a Brady violation and a case that ends without a conviction. The officer, the prosecutor, and the agency are all implicated, and the AI produced the originating error.
Giglio problem from a softened escalation: The subject of a trespassing call was described in the AI draft as "making threatening gestures" when approached. The BWC footage shows the subject raising their hands in a non-threatening manner, consistent with compliance, not with threat. The officer adopted the draft without verifying the characterization against the footage. When the defense cross-examines the officer and plays the BWC footage, the gap between the report and the footage impeaches the officer's credibility as a reporter of facts. The jury hears that the officer signed a report that said "threatening gestures" about footage that clearly shows hands raised in a compliant posture. The officer's testimony about the rest of the incident becomes unreliable in the jury's assessment. The case is lost not because the officer's underlying observation was wrong, but because the AI draft's characterization was not verified, and the officer adopted language that the footage contradicts.
These are not hypothetical worst-case scenarios constructed to frighten officers away from useful tools. They are the predictable downstream consequences of the gap-fill detail and the softened fact failure modes, played out in the adversarial setting of a criminal trial where the footage is available and the defense has the time and the legal right to examine every word of every report.
Key Takeaways
- An AI hallucination in a police report is output that is internally coherent and confident in tone but factually unsupported, wrong, or fabricated. Language models generate it because they work through pattern completion, not fact retrieval and verification.
- The three specific failure modes in AI-assisted report drafting are: the gap-fill detail (the model fills a transcript gap with a plausible but unsupported detail), the softened fact (the model modifies a characterization to be less or more legally significant than the footage shows), and the invented quote (the model attributes a statement to a subject or witness that does not appear in the audio transcript).
- The gap-fill detail is invisible without verification: the report does not signal which sentences came from the audio transcript and which came from the model's training patterns. Every sentence reads with the same confident professional tone.
- The softened fact is harder to detect than the gap-fill detail because the underlying event is real; only the characterization is wrong. The BWC footage shows the event; the officer must specifically verify that the report's characterization accurately represents what was observed.
- The invented quote is the most legally serious hallucination failure mode because an attributed statement to a subject or witness that is fabricated is a direct Brady-Giglio problem: if the actual statement would have been favorable to the defense, it is a Brady violation; if it is not disclosed, it impeaches the officer's credibility as a fact reporter under Giglio.
- The model cannot warn the officer when it is hallucinating. The confident tone of hallucinated output is identical to the confident tone of accurate output. There is no asterisk, no confidence score, no signal from the model. The officer is the error-detection mechanism.
- The verification standard for AI-assisted police reports is not proportionate to the estimated hallucination frequency. It is calibrated to the legal consequences: verify because a defense attorney, a prosecutor, and a judge will read this, and the failure modes have Brady and Giglio implications regardless of how rarely they occur.
- A one-percent error rate across hundreds of officers generating reports per year produces dozens of reports annually with legally problematic hallucinations. The math makes verification a professional obligation even in a high-accuracy tool deployment.
Skill.re