Recognizing Bad AI Output in a Report
The report was nineteen paragraphs long and read like a textbook example of good police writing. The narrative was organized chronologically, the officer's actions were described with appropriate specificity, and the subject's statements were presented in clean, attributed indirect speech. The supervising sergeant approved it in four minutes. It went to the prosecutor. The prosecutor disclosed it to the defense. The defense attorney read it and asked to see the body-worn camera footage. She had the footage transcript in her hand and the report on her desk, and she found the problem in forty minutes. Three facts in the report, stated with full confidence, did not appear anywhere in the footage. One of them was directly contradicted by the footage. The prosecution's case collapsed before it reached trial.
That sergeant approved a bad AI output. Not because she was careless. Because she did not have a checklist. She read the report the way you read a report written by a competent officer: looking for logical coherence and completeness. The document was logically coherent and appeared complete. What it was not was accurate. The AI had gap-filled three details with plausible inventions, and nothing about the surface of the document telegraphed that. The only way to find those errors was to compare the document against the footage, the CAD entry, and the case file, claim by claim. That is the skeptic's read, and this lesson is how to do it systematically.
The Three Failure Modes to Recognize
Bad AI output in a public-safety report almost always takes one of three forms. Understanding the forms before you read the document is like understanding how a counterfeit bill is constructed before you try to spot one: you know what you are looking for instead of trying to sense something is wrong.
Failure Mode 1: The Gap-Fill Detail. This is the most common failure. A gap-fill is a fact the model invented because the source material was silent on that point. The footage may have been partially obscured. The CAD notes may not have included a detail about the subject's demeanor. The transcript may have been ambiguous about the direction of travel. Where a careful human writer would write "direction of travel not observed" or "subject's demeanor not fully captured on BWC (body-worn camera)," the AI model writes what typically happens in incidents of this type: the subject walked north, the subject appeared calm, the subject complied without incident. The invented detail is consistent with the incident. It just was not observed.
Failure Mode 2: The Softened Fact. This is the subtlest failure and the most legally dangerous. A softened fact is a real fact from the footage that the model restates in language that is less specific, less forceful, or less legally significant than the original. A subject who "attempted to strike the officer" becomes a subject who "made a sudden movement." A subject who "stated that he would not comply" becomes a subject who "expressed hesitation." The softening is not random. Models trained on large volumes of writing develop a tendency to produce polished, professional language, and polished professional language in a police context often means measured, bureaucratic language. The problem is that the legally operative detail (the attempted strike, the explicit refusal to comply) is the detail the charge depends on. Softening it can mean the charge is unsupportable.
Failure Mode 3: The Invented Quote. This is the rarest failure and the most catastrophic. An invented quote is an attributed statement that the model generated because it is the kind of thing a person in this situation typically says, not because the person actually said it. A witness "stated that she had seen the subject in the area before." A subject "stated that he had not been at the location." A bystander "stated she heard arguing but did not see what happened." Each of these is a plausible statement. In the model's training distribution, these are exactly the kinds of statements people make in incidents of these types. But if none of them appears in the transcript of the recorded interview, all of them are fabrications. In a criminal case, a fabricated statement attributed to a witness or subject is a Brady violation and potentially perjury if it is incorporated into sworn testimony.
The surface of a bad AI output looks identical to the surface of a good one. The only way to tell them apart is to compare the document against the record, claim by claim.
Why AI Output Is Designed to Look Right
The gap-fill detail, the softened fact, and the invented quote are all produced by the same underlying process: the model generates text that is statistically consistent with its training data and with the inputs you provided. Statistically consistent text looks professional, reads smoothly, and conforms to the genre conventions of the document type it is producing. That is what makes these errors hard to catch on a surface read. They are not sloppy. They are not obviously wrong. They are exactly what the document should look like, built from statistical patterns rather than from the actual record.
This is fundamentally different from the errors a tired officer makes when writing from memory. A tired officer might misspell a name, forget a detail, or get a time wrong by fifteen minutes. Those errors are usually apparent on review and usually trivial in consequence. A tired officer is unlikely to invent a witness statement from whole cloth because the officer knows what the witness actually said and knows the witness statement is foundational to the case. The AI model does not know what the witness said. It has a transcript. Sometimes it synthesizes the transcript accurately. Sometimes it generates a plausible statement the transcript does not support, because a plausible statement is what its generation process produces.
The professional posture for reviewing AI-assisted reports is therefore not "assume this is right unless something looks wrong." It is "assume this document may contain errors that look indistinguishable from correct information, and apply a systematic check." That systematic check is the skeptic's read.
The Skeptic's Checklist for Facts, Quotes, and Sequence
The skeptic's checklist is a structured read of the AI output against the source record. It has three passes, each targeting a different failure mode. Done in sequence, the three passes take between eight and twenty minutes depending on the length of the document and the length of the source record. They are the minimum verification discipline for any AI-assisted document that will be disclosed to a prosecutor or defense attorney.
Pass 1: The Facts Pass. Read the AI output and highlight every specific factual claim: times, locations, directions of travel, physical descriptions, vehicle descriptions, property descriptions, and any factual assertion about what was or was not present at the scene. For each highlighted claim, locate the corresponding support in the source material (the BWC transcript, the CAD (computer-aided dispatch) entry, or the case file). If you cannot locate support within thirty seconds, flag the claim for deeper review. If you cannot locate support after a thorough search, the claim is unverified. Mark it for deletion or correction.
The Facts Pass catches gap-fill details. These are the claims that the model inserted to fill a gap in the source material. They are usually specific and concrete: "the subject was wearing a blue hoodie" when the footage was too dark to determine clothing color, or "the contact occurred at 2247 hours" when the CAD notes show 2251 hours. These discrepancies are not minor. Clothing descriptions are evidentiary. Time discrepancies can affect the credibility of the timeline at trial.
Pass 2: The Quotes Pass. Read the AI output again and identify every attributed statement: any sentence that says a party "stated," "said," "reported," "indicated," or "expressed" something. For each attributed statement, locate the corresponding audio or transcript point in the source recording. The match does not need to be verbatim if the statement is rendered in indirect speech ("the subject stated that he had not been at the location" should correspond to the subject having said something to that effect). But the gist must be present in the record. If the attributed statement is not in the record, mark it for deletion.
The Quotes Pass catches invented quotes. These are attributed statements the model generated from its training distribution. They are often plausible and sometimes nearly accurate, but if they are not in the record, they cannot appear in the sworn document. For statements rendered in direct quotes (with quotation marks), the match must be exact or the quotation marks must be removed and the statement must be rewritten in indirect speech that accurately reflects what the record shows.
Pass 3: The Sequence Pass. Read the AI output one more time, this time tracking the sequence of events as stated. Compare the sequence to the timestamps in the CAD entry and the footage. The CAD entry is the authoritative record of the call's timeline. If the narrative states that Officer A arrived before Officer B but the CAD shows the reverse, that is a sequence error. If the narrative states that a statement was made before the officer announced, but the footage shows it happened after, that is a sequence error. Sequence errors are particularly dangerous in use-of-force reports and in any case where the order of events is legally significant (for example, whether a consent to search preceded or followed a period of detention).
The Sequence Pass catches timeline errors and the specific failure mode of the model rearranging events in a more narratively logical order. Models are trained on narrative text. Narrative logic often puts events in an order that makes the story clearer, not necessarily in the order they actually occurred. The CAD timestamps are your anchor.
The Specific Tells of Each Failure Mode
With practice, reviewers learn to recognize the textual patterns that often precede a failure. These are not infallible signals. They are indicators that a closer look is warranted.
For the gap-fill detail, watch for: claims about sensory observations (what was seen, heard, smelled, or felt) that are more specific than the footage quality would support; physical descriptions of persons or property when the footage was at night or at a distance; exact counts or measurements that the footage could not have captured (a gap-fill might say "approximately eight to ten bystanders" when the footage shows a crowd that was never specifically characterized); and transition phrases like "the officer observed" or "it was apparent that" preceding a claim that is asserted confidently despite being based on inference.
For the softened fact, watch for: any description of a physical encounter that uses passive or ambiguous language ("contact occurred," "a struggle ensued," "a physical intervention was necessary") where the footage shows specific actions that should be described with specificity; descriptions of a subject's verbal conduct that use general emotional language ("the subject was agitated," "the subject expressed concern") where the footage shows specific statements or actions; and any description of a use-of-force incident that describes the officer's actions in more general terms than the footage supports. Specificity is protective in a use-of-force context. Vagueness is a vulnerability.
For the invented quote, watch for: direct quotations (text in quotation marks) in any AI output, which should always trigger a verification against the exact audio or transcript; indirect speech attributions ("the subject stated that," "the witness reported that") for statements that seem too clean or too directly responsive to the investigation's central questions; and any attributed statement that is presented without a corresponding source reference or timestamp. A well-prompted AI output should either have timestamps on attributed statements or should use generic indirect speech. A clean direct quote without a timestamp is almost always a red flag.
The Fast Read Before Submission
Not every AI-assisted document has a thirty-minute verification window before it needs to be submitted. A patrol officer clearing a minor property damage call at the end of a shift does not have the time budget that a detective has before presenting to a grand jury. The skeptic's full three-pass checklist is the standard for high-stakes documents: felony arrests, use-of-force incidents, complex investigations. For lower-stakes documents, a condensed version is appropriate.
The fast read is a five to eight minute single pass that hits the highest-risk elements:
First, read every attributed statement (anything the AI says a party "stated," "said," or "indicated") and verify each one against the recording or transcript. Quotes are the single highest-risk element because they are the most likely to be fabricated and the most legally significant if they are. This step alone takes two to three minutes for a short report.
Second, check every time and address reference against the CAD export. Times and addresses are objective facts that the CAD records accurately. If the AI narrative has a time that does not match the CAD, that is an error regardless of how minor it seems. A defense attorney will find it and use it.
Third, read the use-of-force and resistance description (if any) against the footage, specifically looking for softened language. If the footage shows an officer applying a control hold and the report says "contact was made to gain compliance," that is a softening that needs to be corrected before submission. The narrative must reflect what actually occurred with appropriate specificity.
That is the minimum fast read: quotes, times and addresses, use of force. Three checks. Five to eight minutes. For a low-stakes document, this minimum standard is defensible. For anything involving a felony arrest, a use-of-force incident, a juvenile contact, a victim with special status, or a case that is likely to result in criminal proceedings, the full three-pass checklist is the standard.
What to Do When You Find a Problem
The skeptic's read is valuable only if you act on what it finds. Here is the decision tree for each type of problem.
For an unverified factual claim: If the claim is minor and its absence does not affect the completeness of the report, delete it. If the claim is potentially important but you cannot verify it from the source materials, check your own field notes and memory. If you personally observed the fact in question, rewrite the claim to attribute it explicitly to your personal observation, not to the AI's source synthesis: "Officer observed [fact], not captured on BWC." If you cannot verify it from any source, delete it.
For a softened fact: Rewrite the sentence to reflect the specific action or statement captured on footage. If the footage shows the subject grabbed the officer's arm, the report says the subject grabbed the officer's arm, not "the subject made physical contact." The specificity is not optional. The specificity is the evidence.
For an invented quote: If the quote is in direct speech, delete it. If you have a record of what the party actually said, rewrite in accurate indirect speech with a source reference. If you do not have a record, delete the attributed statement entirely. A missing attributed statement is a smaller problem than an invented one.
For a sequence error: Reorder the narrative to match the CAD timestamps and footage record. If the reordering is significant (it changes the apparent story of the incident), note in the verification documentation that the AI draft had an incorrect sequence and that you corrected it. The CJIS (Criminal Justice Information Services) Security Policy requires agencies to maintain the integrity of criminal justice information. A known sequence error in an AI draft that was submitted without correction is a known accuracy failure.
Every correction you make to an AI draft is an act of authorship. The corrected report is yours. You reviewed it, corrected it, and adopted it. "The AI wrote the uncorrected version" is not information that belongs in the report. The final signed narrative is your sworn account. What belongs in the report is the accurate version of events, verified against the record.
Building the Habit Across the Agency
Individual officers applying the skeptic's checklist are the foundation, but an agency's verification standard is only as strong as its weakest reviewer. The King County (Washington) prosecutor's decision to bar AI-written police reports came from a recognition that the verification standard was not uniform: some officers were applying rigorous checks, others were submitting AI drafts with minimal review. The response was not to ban the technology. The response was to establish a governance standard. Knowing that standard exists, and knowing that prosecutors and defense attorneys are looking for the failures the standard is designed to prevent, is the professional context for everything in this lesson.
Supervisors reviewing AI-assisted reports should apply the same three-pass checklist as the submitting officer. Not as a punitive measure, but as a quality gate. The supervisor who approves a bad AI output shares the accountability for that output when it reaches discovery. A supervisor who cannot perform the skeptic's checklist is a supervisor who cannot review AI-assisted work. That is a skills gap that needs to be filled.
The Electronic Frontier Foundation (EFF) has specifically raised transparency concerns about AI-assisted police reports: the public has a right to know when a report was AI-assisted, how it was reviewed, and what verification standard the agency applies. Those concerns are legitimate. The best answer to them is not to hide the AI assistance. The best answer is to demonstrate the discipline: here is the prompting standard, here is the verification checklist, here is the disclosure language in the report, and here is the audit trail that shows the officer reviewed the draft against the footage before submission. The skeptic's read, documented and disclosed, is the professional answer to the EFF's concern and the prosecutor's concern and the defense attorney's concern. All three concerns are answered by the same discipline.
Giglio v. United States extended Brady to include evidence that could impeach a witness, which in an AI-assisted report context includes the officer's own verification failures. An officer who submits an AI draft without a verification pass, and whose report later turns out to contain invented details, has a Giglio (named for the 1972 Supreme Court case establishing that witness credibility evidence must be disclosed) problem: the fact of insufficient verification is itself impeachment evidence that must be disclosed. The discipline of the skeptic's read is not just a quality practice. It is a protection for the officer's own credibility.
Key Takeaways
- Bad AI output in a public-safety report takes three forms: the gap-fill detail (invented fact), the softened fact (real fact made legally less specific), and the invented quote (fabricated attributed statement). All three look professional on the surface.
- The skeptic's checklist has three passes: the Facts Pass (verify every specific factual claim against the source), the Quotes Pass (verify every attributed statement against the recording), and the Sequence Pass (verify the timeline against CAD timestamps). Done in sequence, they catch all three failure modes.
- For the gap-fill, watch for sensory observations more specific than the footage quality supports, exact counts the footage could not capture, and transition phrases preceding confident inferences.
- For the softened fact, watch for passive or vague language in physical encounter descriptions, general emotional characterizations where specific statements or actions were on record, and use-of-force narratives that lack the specificity the footage shows.
- For the invented quote, all direct quotations in AI output require timestamp verification against the exact audio. Attributed indirect speech without a source reference is a red flag.
- The fast read (quotes, times and addresses, use of force) is the minimum for lower-stakes documents. The full three-pass checklist is mandatory for felony arrests, use-of-force incidents, and any document likely to be contested at trial.
- Every error the skeptic's read finds should be corrected before submission. Softened facts are rewritten with specificity. Invented quotes are deleted or replaced with accurate indirect speech. Sequence errors are corrected against the CAD record.
- Applying the skeptic's checklist and documenting that you did is also the best available answer to King County-style objections, EFF transparency concerns, and Brady and Giglio disclosure obligations. The discipline and the documentation protect both the case and the officer.
Skill.re