Common Failure Modes in AI Reports
Detective Marcus Okafor had a confession that almost wasn't. Three months into his agency's rollout of an AI report-drafting platform, he sat in a pretrial conference while a defense attorney slid a single page across the table: the AI-assisted arrest narrative his patrol officer had filed on a domestic battery case. The attorney's finger rested on one sentence. "Your officer wrote that my client 'appeared calm and cooperative throughout the contact.' I have the body-worn camera footage. My client is screaming for two of the four minutes. Who decided he was calm, Detective, the officer, or the software?" Okafor did not have an answer that night. The narrative had been generated by the tool, lightly reviewed, and adopted. The footage said one thing. The report said another. And the gap between them was now the most important fact in the case.
That sentence was not a typo, and it was not a lie the officer told on purpose. It was a failure mode. AI report-drafting systems, including widely deployed tools like Axon Draft One, do not fabricate randomly. They fail in specific, recognizable, repeatable ways. Once you learn the shapes those failures take, they stop being surprises that ambush you in a pretrial conference and become things you spot during your verification pass, the same way you spot a missing date or a transposed plate number. This lesson is about three of those shapes: the gap-fill detail, the softened fact, and the invented quote. For each one we will look at how it gets into the report, how a careful officer catches it against the record, and what happens at trial if nobody does.
Why AI Reports Fail in Patterns, Not at Random
To catch the failure modes, you have to understand why they happen, because the cause is what tells you where to look. A large language model drafting a police report from body-worn camera (BWC, the camera clipped to an officer's chest that records audio and video through the contact) audio is doing something narrower than it appears. It is predicting the most likely next words given everything it has seen before: your transcript, your computer-aided dispatch (CAD, the system that logs the call and its timestamps) notes, and an enormous body of training text about how incidents are usually described. When your source material is rich and specific, the model leans on your material. When your source material goes quiet, the model leans on the training text. It does not announce the switch. The prose reads identically either way.
That single mechanism produces all three failure modes in this lesson. A gap-fill happens when the footage is silent and the model fills the silence with what usually goes there. A softened fact happens when the model's training distribution, drawn from countless ordinary descriptions, smooths a sharp edge toward the bland statistical middle. An invented quote happens when the model reconstructs what a person "would have said" in this situation rather than transcribing what they did say. None of these is the model malfunctioning. This is the model working exactly as designed, and that is precisely why the output is so dangerous: it is fluent, confident, professionally formatted, and wrong in ways that look right.
This matters because of the one fact that reorders everything in this field. The output is evidence. A police report is not a memo. It is disclosed to the defense in discovery, read aloud in depositions, projected on a courtroom screen next to the footage, and tested under cross-examination. Officers spend roughly 30 to 40% of every shift on paperwork, and Axon's Draft One testing showed officers reporting an 82% decrease in report-writing time, so the pull toward fast adoption is real and strong. But the 82% saving lands on the drafting side. The verification side is still yours, and the failure modes below are exactly what the verification side exists to catch.
The model does not fail randomly. It fails by leaning on its training text whenever your record goes quiet, and it never tells you when it switched.
The Verification Mindset for Failure Modes
There is a mental shift that separates officers who catch these failures from officers who get caught by them. When you wrote reports by hand, the act of writing was itself a form of verification: every sentence came from your memory of the scene. When you review an AI draft, you are reading a document engineered to read as if it has no errors. You cannot rely on the prose to alert you, because the prose is uniformly smooth across the true parts and the invented parts. The only reliable method is to stop trusting the narrative and start interrogating it against the record: the BWC footage, the CAD entry, the records management system (RMS, the agency's central case-file database) file, and your own field notes. The narrative is the suspect. The record is the witness. You believe the witness.
Failure Mode One: The Gap-Fill Detail
The gap-fill is the most dangerous of the three because it is the hardest to see. It is a fact the model invents to fill a hole that the footage never shows. The model does not know it is inventing. From the inside, generating a plausible detail for a quiet stretch of audio feels identical to transcribing a detail that was actually spoken. The result is a sentence that sounds like every other sentence in the report but has no source behind it.
Worked Example: The Direction of Travel
Consider a foot-pursuit narrative. The officer chased a suspect who fled on foot from a traffic stop. The BWC audio captures the officer shouting commands, heavy breathing, and the radio traffic, but the visual frame is chaotic and the suspect leaves the camera's view for several seconds near an alley. The AI draft produces this sentence:
"The subject fled westbound through the alley and discarded a dark object near the dumpster before being detained."
Read it cold and it is unremarkable. It is the kind of sentence that appears in a thousand pursuit reports. But where did "westbound" come from? Where did "discarded a dark object near the dumpster" come from? If you go to the footage, you may find that the camera never captured the direction of travel during those seconds, and never captured any object being discarded. The model produced both details because in its training data, fleeing suspects in alleys frequently discard objects, and reports frequently specify a compass direction. The model filled the visual gap with the statistically typical content. That dark object, if a weapon was later recovered nearby, just became the centerpiece of a probable-cause argument built on a detail no camera ever recorded.
The Catch
The verification move for a gap-fill is the discipline of source-tracing every factual claim. Go through the draft and, for each fact (a direction, an object, an action, a location, a time), ask one question: where does this come from? Point to the timestamp in the BWC footage, the line in the CAD entry, the field in the RMS, or your own direct observation. If you can point to a source, the fact is verified. If you cannot, you have found a gap-fill. In the example, you would scrub to the alley segment, watch it, and discover the camera shows nothing about direction or a discarded object during the gap. That sentence cannot stand as written. You either correct it to what the footage actually shows ("The subject left the camera's field of view for approximately four seconds near the alley"), add what you personally observed with explicit attribution ("Observed by the responding officer, not captured on BWC: the subject turned west"), or remove the unsupported claim entirely.
The Consequence
The gap-fill detail is an evidentiary fault line. If the "discarded dark object" sentence helped establish probable cause and the footage does not support it, the defense has a suppression argument: the search or seizure rested on a fact with no evidentiary foundation. Worse, once a defense attorney demonstrates that one sentence in your report describes something that never happened on camera, every other sentence is now suspect. That is the impeachment dynamic that Giglio v. United States (405 U.S. 150, requiring disclosure of evidence that can impeach a witness, including the testifying officer) puts in motion. The officer's credibility, which is the currency of the entire case, takes the hit. An unverified gap-fill is also a Brady (Brady v. Maryland, 373 U.S. 83, requiring disclosure of material exculpatory evidence) problem in reverse: the report asserts an inculpatory fact the record cannot back, which is exactly the kind of unreliability the defense is entitled to expose.
Failure Mode Two: The Softened Fact
The softened fact is the subtlest of the three, and the one that the opening story turned on. Where the gap-fill invents a detail in a silence, the softened fact takes a real, sharp-edged event from the record and describes it more mildly than it actually was. Force becomes "contact." A struggle becomes "noncompliance." A visible injury becomes "minor." Active, sustained resistance becomes "some hesitation." The event is real and the description is in the neighborhood of true, which is exactly why it slips past a casual read. It is not a lie. It is a drift toward the bland statistical center of how things are usually described, and that drift runs in the direction that does the most damage in court.
Worked Example: "Appeared Calm and Cooperative"
Return to Detective Okafor's case. The AI draft of the domestic battery arrest contained this line:
"The subject appeared calm and cooperative throughout the contact."
The BWC footage tells a different story. For roughly half of the four-minute contact, the subject is shouting, refusing commands, and physically tensing against the officers. Why did the model write "calm and cooperative"? Because the bulk of routine arrest narratives in its training data describe subjects who comply, and because the phrase "calm and cooperative" is one of the most common boilerplate constructions in the genre. The model reached for the high-frequency phrase and smoothed over the parts of the audio where the subject was anything but calm. It also softened in the direction that, in this case, helped no one: it understated the resistance, which undercut the officer's own justification for the level of control used and handed the defense a contradiction between the sworn report and the recording.
Softening cuts both ways and both are dangerous. A draft that understates a suspect's resistance can gut the officer's use-of-force justification. A draft that softens the severity of a victim's injury ("minor abrasions" when the footage and medical record show something far worse) can weaken the charge and, critically, can suppress exculpatory or aggravating facts the parties are entitled to. The model is not choosing a side. It is regressing toward the mean, and the mean is rarely the truth of a specific violent moment.
The Catch
The verification move for a softened fact is to treat every characterizing word as a claim that must match the footage at full intensity, not approximately. Adjectives and intensity words are the tells: "calm," "minor," "brief," "cooperative," "some," "appeared to." When you hit one, do not read past it. Go to the footage and ask whether the word matches what the camera recorded at the actual level of severity. Watch the contact. If the subject is shouting and tensing for two of four minutes, "calm and cooperative throughout" fails. You replace it with what the record shows: "The subject was initially noncompliant, shouting and physically tensing against officers for approximately the first two minutes of the contact, before complying with commands." Cross-check injury characterizations against the footage and against the medical or RMS record, not against the model's adjective. The CAD entry, the footage timestamps, and any photos are your intensity gauge. The model's word choice is not.
The Consequence
A softened fact that survives to the sworn report is a direct impeachment gift. When the defense plays the footage of a screaming, resisting subject next to your report's "calm and cooperative throughout," the jury does not just disbelieve that one sentence. They watch the officer's report fail to match reality on a question the camera answers definitively, and the credibility damage spreads to everything the officer says the camera did not capture. This is the Giglio impeachment scenario in its purest form. If the softening understated a victim's injury or a suspect's exculpatory statement, it is squarely a Brady problem: the sworn account suppressed material the defense was constitutionally entitled to. Cases have been weakened, charges reduced, and convictions put at risk over exactly this kind of mismatch between the report's adjectives and the footage's reality.
The softened fact is real but understated. Treat every adjective as a claim, and verify it against the footage at full intensity, never approximately.
Failure Mode Three: The Invented or Smoothed Quote
The third failure mode lives inside quotation marks, which is precisely what makes it so corrosive. Quotation marks are a representation to every future reader that the words between them are what the person actually said, verbatim. An AI draft can violate that representation in two ways. It can smooth a real statement into cleaner language than the person used, or it can invent a statement the person never made at all, reconstructing what someone "would have said" in the situation. Both put fabricated words in a real person's mouth under the authority of your sworn signature.
Worked Example: The Cleaned-Up Admission
An officer responds to a shoplifting call and detains a subject who, on the BWC audio, says: "I ain't got nothin on me, I didn't take nothin, you can check." The AI draft renders the statement this way:
The subject stated, "I do not have anything on me. I did not take anything. You may check."
Every fact is preserved. The grammar is fixed. And the quote is now false, because those are not the words the subject said. The model smoothed nonstandard grammar into clean prose because clean prose is what its training data overwhelmingly contains inside quotation marks in formal documents. In a harder version of this failure, the model invents a quote with no audio behind it at all. Suppose the footage shows the subject mumbling inaudibly during a pat-down, and the draft confidently reports: The subject stated, "Okay, you got me." If no such statement exists in the audio, the model has manufactured a confession. That fabricated admission, sitting in a sworn report, could drive a charging decision and a plea.
The Catch
The verification move for quotes is the strictest in the entire review, and it admits no shortcuts: every set of quotation marks gets checked verbatim against the audio. This is not a paraphrase check or a substance check. You go to the exact moment in the BWC footage, you listen, and you confirm that the words between the marks are the words on the recording, exactly. "I do not have anything" does not match "I ain't got nothin," so the quotation marks have to go. You either transcribe what the subject actually said, word for word, or you convert it to an attributed paraphrase that drops the quotation marks ("The subject denied possessing or taking any items and invited officers to search"). For the invented confession, you scrub to the pat-down, find no such audio, and delete the fabricated quote entirely. If a statement was made off-camera and you heard it yourself, you attribute it explicitly as your own observation, not as captured audio, and you never wrap your reconstruction in quotation marks as if it were verbatim.
The Consequence
The invented or smoothed quote is the failure mode most likely to end a case and a career. Tense, grammar, and exact word choice are the raw material of cross-examination. When the defense plays audio of "I ain't got nothin" against a report that quotes "I do not have anything," the officer is now explaining on the stand why a sworn report contains words a recording proves were never spoken. A fully invented quote is far worse: a fabricated admission in a sworn document is the kind of thing that triggers suppression of the statement, dismissal of charges, and a Giglio entry that follows the officer to every future case. The Electronic Frontier Foundation (EFF, a digital-rights organization that has raised transparency concerns about AI-drafted police reports) has pointed to exactly this risk, that automated systems generate language a defendant is then accused of saying. A quote you did not verify against the audio is not the suspect's statement. It is the model's guess, dressed in quotation marks and signed by you.
Building the Failure-Mode Pass Into Your Review
Knowing the three shapes is not enough on its own. The shapes have to become a habit you run on every AI-touched report, the same way you check the incident number and the date without being told. The good news is that the three failure modes map onto three concrete verification moves you can sequence into your existing review.
The Three Moves
- Source-trace every fact. For each factual claim (direction, object, action, location, time), point to the specific source in the BWC footage, CAD entry, RMS, or your own observation. Anything you cannot trace is a candidate gap-fill. Correct it to the record, attribute it as personal observation, or remove it.
- Pressure-test every characterizing word. Stop at every adjective and intensity word ("calm," "minor," "brief," "cooperative," "some"). Verify it against the footage at full severity, not approximately. If the word understates what the camera shows, replace it with the accurate description.
- Verify every quote verbatim. For each set of quotation marks, go to the exact audio moment and confirm the words match exactly. If they do not, transcribe accurately or drop the quotation marks for an attributed paraphrase. Delete any quote with no audio behind it.
These three moves are not extra work bolted onto the job. They are the job now. The 82% time saving from AI drafting is real, but it is the drafting that got faster, not the verification. The discipline that catches the gap-fill, the softened fact, and the invented quote is the discipline that lets you stand behind the report when a defense attorney slides it across the table.
Governance and the Bigger Picture
These failure modes are also why prosecutors and oversight bodies are watching AI reports so closely. The King County, Washington prosecutor's office drew a hard governance line, barring AI-written police reports from its charging process absent specific documentation of the review. That position is not anti-technology. It is a direct response to exactly the failures in this lesson: a prosecutor who has seen a gap-fill or an invented quote blow up a case in discovery does not want unverified AI prose in their charging file. The officer who can document a real failure-mode verification pass, claim by claim, against the footage, is the officer whose reports survive that scrutiny. The officer who adopts the draft on a quick read is the officer who learns about the gap-fill from the defense.
When you are asked, in a deposition or a hearing or a review-board meeting, how you know the report is accurate, the answer is not "the tool generated it" and it is not "I read it over." The answer is the failure-mode pass: "I traced every fact to a source in the record, I verified every characterization against the footage at its actual severity, and I checked every quotation against the audio verbatim. Where the draft did not match the record, I corrected it." That answer is true, documented, and defensible. The three failure modes stop being things that surprise you and become things you caught.
Key Takeaways
- AI report tools fail in three recognizable patterns, not at random: the model leans on its training text whenever your record goes quiet, and it never signals the switch. Learn the shapes and the failures become things you spot rather than things that ambush you.
- The gap-fill detail is an invented fact filling a hole the footage never shows (a direction, a discarded object, an action). Catch it by source-tracing every claim; an untraceable inculpatory fact is a suppression and Brady-reliability problem.
- The softened fact takes a real, sharp event and describes it more mildly than the record ("calm and cooperative" over a screaming, resisting subject; "minor" over a serious injury). Catch it by verifying every adjective against the footage at full intensity.
- The invented or smoothed quote puts fabricated words inside quotation marks, either by cleaning up nonstandard grammar or by manufacturing a statement that was never made. Verify every quote verbatim against the audio; a fabricated admission can drive suppression and dismissal.
- All three are Giglio impeachment gifts: once the footage contradicts one sentence, the officer's credibility on every uncaptured claim collapses. A softened or invented detail that suppresses material the defense is owed is also a Brady violation.
- The verification mindset is to distrust the narrative and interrogate it against the record. The prose reads identically across the true parts and the invented parts, so smoothness is no signal; only the BWC footage, CAD entry, RMS file, and your notes are.
- The three verification moves (source-trace every fact, pressure-test every characterizing word, verify every quote verbatim) sequence into your existing review and are the work, not extra work. The 82% saving is on the drafting side; verification is still yours.
- The King County prosecutor's bar and the EFF's transparency concerns target exactly these failures. The officer who can document a failure-mode pass, claim by claim against the footage, is the officer whose reports survive discovery, deposition, and cross-examination.
Skill.re