Catching Hallucinations Against the Record
The deposition started at 9 a.m. on a Tuesday. Senior Detective Rosa Alvarez had testified dozens of times. She was not nervous. She was confident in the arrest, confident in the investigation, and confident in the report she had reviewed and submitted. Then the defense attorney put a sentence on the screen: "The suspect stated that he had not been in the area for several days." The attorney asked Alvarez to identify the timestamp in the original interview recording that supported that statement. Alvarez looked at the sentence. She had reviewed the AI-drafted case summary. She had reviewed it. She had not gone to the recording. The interview was four hours and eleven minutes long. She could not point to the timestamp because she had not checked it. The sentence was not in the recording. The suspect had said he had not been in that neighborhood specifically, for that specific reason. The broader claim in the AI summary was a paraphrase that crossed the line into fabrication. The defense attorney used the discrepancy to undermine every other quote attributed to her suspect in the case summary. The case survived, but narrowly, and Alvarez spent two days of testimony re-establishing facts that should have been uncontested. The AI had not lied. It had generalized. She had not caught it.
What a Hallucination Costs in This Context
The word "hallucination" comes from the machine learning field, where it describes a model generating content that is fluent, grammatically correct, and confidently presented but not grounded in the source material the model was asked to use. In a public safety context, that technical definition collides with a legal one. When an AI-generated police report narrative, case summary, or interview transcript summary contains a claim that is not supported by the evidentiary record, that claim is not merely a software error. It is a potential Brady violation, a potential Giglio problem, and a potential basis for suppression of evidence.
Brady v. Maryland (1963) requires the prosecution to disclose material exculpatory evidence to the defense. If an AI-generated summary omits or distorts an exculpatory statement the suspect actually made, and that distortion is not caught before the report is submitted, the agency has a Brady problem. Brady (named for the case that established the rule) is not a technicality. A Brady violation can result in dismissal of charges, reversal of conviction, and civil exposure for the agency and the individuals involved.
Giglio v. United States (1972) extends Brady to include impeachment evidence, meaning evidence that would undermine the credibility of a witness, including the officer who authored the report. If a defense attorney can show that an officer submitted a case summary containing a statement attributed to a suspect that the suspect never actually made, that fact is now Giglio material in every case that officer touches going forward. The officer's credibility is damaged not just in this case but in every future case where they testify. That is the actual cost of a hallucinated quote that gets through the review process.
The CJIS (Criminal Justice Information Services) Security Policy, administered by the FBI, governs how criminal justice information is handled, stored, and transmitted. CJIS obligations stay with the agency, not with the AI vendor. The vendor does not inherit responsibility for the accuracy of the records your agency submits to prosecutorial teams, defense counsel, or the court. That responsibility is yours, and has always been yours. AI-assisted drafting does not change the ownership of that obligation.
A hallucination in a sworn report is not a software failure. It is an unsigned check drawn on your credibility as an officer, submitted to a court that will cash it.
The Taxonomy of Hallucination Types in Public Safety Records
To catch hallucinations systematically, you need to know what they look like. There are six distinct failure modes that appear in AI-generated public safety documents, and each one has a different signature. Knowing the taxonomy is the foundation of a systematic detection practice.
The Gap-Fill
The gap-fill is the most common hallucination type. It occurs when the audio or source material has a gap, a moment where no relevant information was captured, and the model fills that gap with content that is statistically plausible given the context. A BWC (body-worn camera) recording with degraded audio, a period of silence during a stop, a moment where voices overlap and individual statements cannot be distinguished: all of these create conditions where the model must choose between acknowledging the gap or filling it. Models prefer to fill. A complete, coherent narrative is what the model is optimized to produce. Incompleteness is penalized; fluency is rewarded. The gap-fill produces text that reads well and is, in whole or in part, invented.
The gap-fill signature: the draft contains a specific claim about an action, statement, or observation for which there is no corresponding moment in the source recording. When you look for the timestamp that supports the claim, it does not exist. The claim was assembled from surrounding context, not from the specific moment in question.
The Softened Fact
The softened fact occurs when the model encounters something difficult in the source material, a clear statement of resistance, a specific admission, a sharp moment of confrontation, and produces a version that is semantically similar but legally less precise. "I didn't do it and you can't prove it" becomes "the subject denied involvement." "Get your hands off me" becomes "the subject indicated unwillingness to comply." The softened version is not exactly wrong. But the original version is more probative, more specific, and more useful to both prosecution and defense in understanding exactly what happened. The softened version loses information that matters.
The softened-fact signature: a statement in the draft that is vague where the original was specific, or general where the original was precise. If the source material contains a clear, emphatic, exact statement and the draft contains a smooth paraphrase, verify the original. The smooth paraphrase is often a cover for softening.
The Invented Quote
The invented quote is the most dangerous variant because it presents fabricated content with the maximum legal authority: quotation marks. In legal contexts, a direct quote is a representation of verbatim accuracy. It is not a paraphrase. When an attorney or a judge sees a direct quote in a sworn document, they are entitled to assume those are the exact words the person said. A model that produces a direct quote that is a plausible reconstruction of something the speaker might have said, rather than a verbatim transcription of something the speaker actually said, has produced a document that is legally false on its face.
The invented-quote signature: any direct quotation in the draft that you cannot verify verbatim against the audio. If you cannot play the recording to the relevant timestamp and hear those exact words, the quote is unverified. An unverified direct quote does not belong in a sworn document. It should be converted to paraphrase language with an explicit notation that the exact words are paraphrased from the recording.
The Sequence Error
The sequence error occurs when the model correctly identifies events that occurred but places them in the wrong order. The model has encountered audio that includes references to both event A and event B, but the audio does not make the sequence entirely clear, and the model produces the sequence that is most common in incidents of this type, not necessarily the sequence of this specific incident. Sequence matters enormously in law enforcement reporting. Whether an officer drew their weapon before or after a suspect reached into their waistband is not a narrative detail. It is a legal determination. A sequence error in an AI-drafted use-of-force narrative can reframe the entire force decision as proactive rather than reactive.
The sequence-error signature: a narrative where the order of events in the draft does not match your recollection or your notes, or where the draft places an event at a timestamp that is inconsistent with the footage or the CAD (computer-aided dispatch) entry. Sequence errors are often subtle because the individual events are accurately described. Only the order is wrong. That is why temporal anchoring, matching each event in the draft to a specific CAD or footage timestamp, is essential to catching them.
The Attribution Error
The attribution error occurs when the model correctly captures something that was said or done but attributes it to the wrong person. In a multi-party incident with multiple voices on the recording, the model tracks who said what based on audio cues: voice pitch, direction, proximity to the microphone, speaking patterns. In chaotic situations, those cues are unreliable. The model may assign a statement to the officer that was actually made by a bystander, or assign a statement to suspect A that was actually made by suspect B. Attribution errors are particularly dangerous when they affect who gave consent, who made an admission, or who issued a threat.
The attribution-error signature: a statement attributed to a specific person in the draft that does not match your recollection of who actually made that statement. In multi-party incidents, verify attribution for every significant statement: confirm not just that the statement was made but that the person the draft identifies as the speaker was in fact the person you heard make it.
The Contextual Collapse
The contextual collapse occurs when the model summarizes a long interview or incident report and loses the nuance that distinguishes what was said from what was implied, what was established from what was alleged, or what the suspect said from what the investigator characterized. A four-hour interview produces an AI summary that, if not verified, may collapse the distinctions between proven facts and investigative theories, between the suspect's words and the detective's interpretations of those words, or between what was established on the record and what was said in an off-the-record conversation. The summary reads efficiently and loses the nuance that a defense attorney will immediately identify as missing from the original recording.
The contextual-collapse signature: a summary that presents investigative inferences as established facts, or that collapses the difference between what a witness said and what an officer concluded from what the witness said. If the summary reads cleaner and more definitive than the underlying material, that is a flag. Certainty in a summary that was not in the original is usually a sign of contextual collapse.
Building a Systematic Detection Practice
The word "systematic" is load-bearing here. Catching hallucinations reliably is not about vigilance as a personal virtue. It is about process design. The officer who catches hallucinations consistently does so because they have a defined process they run every time, not because they are more skeptical or more careful than their peers. Heroic attention to detail is not a sustainable operational standard. Process is.
The Source Map
Before running any verification pass on an AI-generated document, build a source map. A source map is a simple reference document that lists every source material that went into the AI's context: the BWC footage and its runtime, any dashcam footage, the CAD entry and its timestamps, field notes, any supplemental recordings, any physical evidence documentation. This takes approximately three minutes to produce and is the foundation of every verification decision that follows.
The source map answers the question every verification decision requires: does the source material exist that could support this claim? If a claim in the draft is about something that happened at 22:14, and the BWC was not activated until 22:17, the source map tells you immediately that the 22:14 claim has no footage source. It may have a CAD source, a field-notes source, or no source at all. The source map surfaces that problem before you spend twenty minutes looking for footage that does not exist.
The Three-Pass Structure
Run the verification in three passes, not one. Each pass has a different focus, and they build on each other.
The first pass is the structural pass. Read the entire document without comparing it to any source material. Your goal in the first pass is to identify the claims that are most likely to be contested or consequential: use-of-force descriptions, direct quotes, specific times and locations, attributions of statements to specific individuals, and anything that establishes the legal basis for an arrest, a search, or a use of force. Mark those sections. They are the priority verification targets in the second pass.
The second pass is the source verification pass. For each marked claim, go to the source material and find the support. Open the footage, navigate to the timestamp, and confirm. Check the CAD entry for time and location claims. Verify quotes against the audio, verbatim for anything in quotation marks, substance for close paraphrases. If a claim has no source, mark it as unverified. Do not correct it yet. Finish the second pass completely before you start editing. You want to know the full scope of what needs correction before you start making changes.
The third pass is the correction and documentation pass. For each unverified claim, make one of three decisions: correct it to what the source material shows, convert it to explicitly sourced language (noting the source, including personal observation where appropriate), or remove it from the document. Document every correction: what the draft said, what the source material showed, and what the corrected text says. This documentation is your audit trail. It is also your protection if the original AI draft is ever reviewed.
The Quote Protocol
Quotes require a separate protocol because the legal stakes of a misquote are higher than the stakes of a paraphrased error. The quote protocol has three rules.
Rule one: if quotation marks appear in the draft, verify verbatim against the recording. No exceptions. Not "close enough." Verbatim. If the recording says "I haven't been there" and the draft says "I was not there," those are not the same statement, and they should not appear in the same quotation marks.
Rule two: if the audio is too degraded to verify verbatim, convert the quote to paraphrase language and note the audio quality. "The subject stated words to the effect that he had not been in the location" is a stronger document than a direct quote you cannot verify. The paraphrase, honestly attributed as a paraphrase, cannot be disproved by playing the recording. The fabricated direct quote can.
Rule three: if the model has attributed a statement to the wrong person, correct the attribution before anything else. An accurate statement attributed to the wrong person is an incorrect statement. In an evidentiary document, attribution is not a secondary consideration.
The Temporal Anchor
For any document that involves a sequence of events, build the temporal anchor before the verification pass. A temporal anchor is a timeline of the incident built from the CAD entry, the RMS (records management system) timestamp data, and the footage timestamps. It shows, in order, when the call was received, when units were dispatched, when footage began rolling, and what the footage shows at each key timestamp.
With the temporal anchor in hand, verify each event in the AI draft against the timeline. Does the draft place events in the correct order? Does it attribute events to the correct time period? A sequence error that assigns the wrong temporal relationship to two events may not be visible if you read the draft as a narrative. It becomes visible immediately when you compare the draft's sequence to the anchored timeline.
The Deposition Scenario and What It Demands
The standard of practice for hallucination detection should be set by imagining the deposition. Not the routine case, not the case that settles, not the case that never goes to trial. The case that goes to a full adversarial proceeding with a skilled defense attorney who has read every document in the file, listened to every recording, and identified every discrepancy.
That attorney will find the gap-fill. They will find the softened fact. They will find the invented quote. These are not hypothetical risks. They are the things a competent defense attorney is specifically trained to find, because those discrepancies are exactly the kind of material that can establish reasonable doubt, impeach a witness, or support a motion to suppress. The question is not whether the discrepancies will be found. It is whether they will be in your document when the attorney finds them, or whether your verification pass found them first and corrected them before they became part of the sworn record.
The deposition answer you are building toward is: "The AI generated an initial draft from the audio and footage. I reviewed that draft against the recording and my field notes. I made specific corrections where the draft did not match the recording. The corrections are documented in my workflow log. The report I submitted reflects what the recording shows." That answer is not defensive. It is accurate. It demonstrates professional practice. It protects you because it is true.
The answer you are trying to avoid is: "I reviewed the draft and thought it was accurate." That answer invites the follow-up: "When you reviewed the draft, did you compare it to the recording?" If the answer to that question is no, the attorney has established that the report you submitted may contain claims you cannot personally verify against the source material you were given. That is the foundation of impeachment.
Agency-Level Systemization
Individual verification discipline is necessary but not sufficient. An agency that relies on each officer's personal discipline to catch every hallucination has an inconsistent quality-control system that will fail when officers are fatigued, when caseloads are high, or when the shift ends and the report needs to be filed. The verification practice needs to be built into the agency's workflows, not left to individual initiative.
Axon's Draft One, which drafts police report narratives from BWC audio, reported that testing officers saw an 82% decrease in report-writing time. That figure is significant. It is also a baseline, not a guarantee, and it describes the drafting time, not the verification time. The verification pass adds time back. The question is not whether the overall time savings survive the verification pass (they do, meaningfully) but whether the agency has accounted for verification time in its workflow design, its staffing models, and its training programs.
Officers in agencies where AI-assisted report writing has been deployed without a formal verification protocol are spending the saved time in ways that have not been designed with verification in mind. That is a governance gap. The 30 to 40% of shift time that officers historically spent on paperwork does not disappear when the drafting is automated. Some portion of it needs to be reallocated to verification. An agency that captures the full 30 to 40% as pure efficiency without accounting for verification is running a quality-control deficit that will show up in court.
A formal verification protocol at the agency level includes three components: a written standard operating procedure that specifies the verification steps required before an AI-assisted document can be submitted; a workflow log that creates a record of the verification for every AI-assisted document; and a supervisory review cadence that samples completed documents to confirm the verification standard is being met. The King County, Washington, prosecutor who barred AI-written police reports from their cases did so in part because the agency could not demonstrate a verification standard. A formal protocol is the evidence that a standard exists.
The EFF Concern and What It Means for Your Practice
The Electronic Frontier Foundation (EFF) has raised transparency concerns about AI-assisted police report writing that are worth understanding directly, not as obstacles to dismiss but as the legitimate criticisms of a careful civil-liberties organization that has been watching these systems deploy in real agencies.
The EFF's core concern is that AI-assisted reports make it difficult or impossible for defendants to understand how the report was generated, what the source material was, and whether the content of the report accurately reflects that source material. When a defendant receives a police report in discovery, they are entitled to understand what that document is: who wrote it, when, based on what observations. If the report was AI-generated from BWC audio, and that fact is not disclosed, the defendant is receiving a document whose authorship and factual basis are obscured. That is a transparency problem and potentially a Brady problem.
The practical implication for your verification practice is that transparency is not just a courtesy. It is a requirement. Every AI-assisted document that goes into the criminal justice process should be clearly identified as AI-assisted, should document the review process that was applied, and should be accompanied by the source material that was used to generate it (or by clear documentation that the source material is available for review). This is not bureaucratic overhead. It is the disclosure architecture that allows the defense to meaningfully evaluate the reliability of the document. That evaluation is a constitutional right, and it is one that the EFF is watching closely.
Agencies that have deployed AI-assisted report writing with bundled, multi-year, sole-vendor contracts (contracts in the range of approximately $45 million over ten years have been reported in several procurements) have taken on disclosure obligations that extend for the life of those contracts. Every report generated under those contracts is potentially subject to disclosure questions for the duration of the contract period and beyond, because convictions obtained with AI-assisted reports will be subject to challenge during appeals and post-conviction proceedings that may extend years into the future. The verification and documentation practices you build today are the practices that will be examined in proceedings that may not occur until long after the specific incident is forgotten.
Key Takeaways
- AI hallucinations in public safety records take six specific forms: gap-fills, softened facts, invented quotes, sequence errors, attribution errors, and contextual collapses. Each has a recognizable signature, and knowing the taxonomy is the foundation of systematic detection.
- A hallucinated quote in a sworn document is not a software error with a simple fix. It is a potential Brady violation, a Giglio problem, and a basis for suppression that can follow an officer through every case they testify in afterward.
- Catching hallucinations consistently requires process, not just vigilance. Build a source map before the verification pass, run the three-pass structure (structural, source-verification, correction-and-documentation), and apply the quote protocol to every direct quotation.
- Sequence errors and attribution errors are the most likely to survive a casual review because the underlying events are accurate. Temporal anchoring and explicit attribution verification are the tools that catch them.
- The CJIS Security Policy obligations and report accuracy obligations stay with the agency, not the AI vendor. No contract transfers the legal responsibility for the accuracy of sworn documents from the officer and agency to the technology provider.
- Agency-level systemization, written SOPs, workflow logs, and supervisory review cadences, is necessary because individual verification discipline alone will not hold under operational pressure. The King County example is a real governance line: agencies that cannot demonstrate a verification standard are vulnerable to having AI-assisted records barred from their cases.
- EFF transparency concerns are legitimate and should be addressed directly: every AI-assisted document that enters the criminal justice process should be identified as AI-assisted, should document the review applied, and should preserve access to the source material for defense review.
- The deposition answer that protects you is specific and documented: you reviewed the draft against the recording, you found and corrected specific discrepancies, and those corrections are in your workflow log. The answer that exposes you is that you thought the draft was accurate without verifying it.
Skill.re