AI-Assisted Interview and Evidence Summarization
Detective Alicia Torres had twelve hours of recorded interview audio sitting on her desk, four witnesses, two suspects, and a trial date six weeks out. The recordings had been transcribed overnight by her agency's AI platform. The summaries were clean, coherent, and readable. Then she opened the first one and noticed a sentence she could not place: a witness was described as saying the vehicle "turned left onto Maplewood before the shots." She pulled up the original recording and scrubbed to the timestamp. The witness never said that. The model had inferred it from context, from something a second witness had said about the route, and had placed the statement in the wrong interview entirely. The summary was confident. It was professional. And it had just put words in a witness's mouth that the witness never uttered.
Why Interview Summarization Is the Hardest AI Task in Investigations
A police report from a single officer's body-worn camera (BWC, the small camera mounted on the chest or lapel of a uniformed officer, recording audio and video of the officer's interactions) is difficult enough to verify against the record. An interview summary from twelve hours of audio across four witnesses is an order of magnitude harder. The volume is real: a major-crimes detective handling a homicide or a sexual assault may accumulate dozens of recorded interviews over the life of a case, each interview running from thirty minutes to several hours. Transcribing and summarizing that material by hand is not just time-consuming. It is, in practice, often incomplete. Detectives make notes, clip key passages, and work from memory for the rest.
AI tools that transcribe and summarize interview audio represent a genuine capability leap. The same platforms that transcribe body-camera audio for patrol officers can process recorded interviews, produce timestamped transcripts, extract key statements, and generate narrative summaries. The time return is real. A detective who previously spent two days processing interview transcripts can now receive a searchable, organized summary within hours. The 82% reduction in report-writing time that Axon's Draft One has reported in patrol contexts translates, in broad terms, to similar savings in any documentation-heavy investigative task.
But the investigative context introduces a failure mode that does not exist in the same form in patrol report writing: the invented or misattributed quote. When an AI summarizes a police report from body-camera footage, the fabricated detail is bad. It may be a gap-fill about a sequence of events, a softened description of force used, or an incorrectly timed event. Each of those is serious. But when an AI summarizes an interview and invents, blends, or misattributes a quote from a witness or a suspect, the failure mode is case-ending.
What Makes a Fabricated Witness Quote Different
To understand why the misattributed or fabricated quote is a different category of error, it helps to think through what happens when it reaches the case file. Brady v. Maryland (1963) requires the prosecution to disclose to the defense any exculpatory evidence it has. Giglio v. United States (1972) extends that obligation to evidence that could be used to impeach the credibility of a witness, including evidence that a witness's prior statements are inconsistent. These are constitutional obligations, not procedural preferences. A fabricated or blended quote in a case summary can be a Brady or Giglio problem in either direction: if the model softens an inculpatory statement to sound less damning, the defense should have seen the full statement but received a weaker version; if the model strengthens a statement that was actually tentative or hedged, the defense is now impeaching a witness on a statement the witness did not fully make.
Brady v. Maryland requires prosecutors to turn over exculpatory material. Giglio v. United States requires disclosure of material that could impeach a government witness. Both obligations attach to the investigative file that AI touched. A fabricated quote, once it enters the case summary, becomes part of the record the prosecution is working from. If the trial testimony differs from the summary, the defense attorney will ask why. If the difference traces back to AI summarization that invented or misattributed language, the detective may be facing a Giglio problem and possibly a motion to suppress based on prosecutorial misconduct in the construction of the file.
Beyond Brady and Giglio, there is the direct evidentiary consequence. A defendant's statement is one of the most powerful pieces of evidence in a criminal case. If an AI tool slightly paraphrases or strengthens what a suspect said, and the detective's notes and the case summary both reflect the AI version rather than the verbatim recording, the defense will subpoena the recording, compare it to the summary, and present the discrepancy to the jury. The detective, who wrote or adopted the summary, will be cross-examined on the difference. "Detective, did the defendant actually say the word 'definitely,' or is that how your AI tool phrased what it heard?" is a question that no detective wants to answer in a courtroom without a timestamped citation to the recording.
How AI Interview Summarization Actually Works and Where It Fails
Understanding the failure mode requires a working knowledge of what the model is actually doing. Modern interview summarization tools use a pipeline that typically involves: (1) speech-to-text transcription, which converts the audio to a text file; (2) segmentation, which organizes the transcript by speaker and by topic or time block; (3) extraction, which identifies key statements, facts, and admissions; and (4) synthesis, which generates a readable summary narrative or key-points list from the extracted material.
Each stage has its own failure mode. Transcription errors are the most visible: a word is misheard, a name is misspelled, a numeral is wrong. These are relatively easy to catch against the recording because the error is usually localized and obvious on review. The more dangerous failures occur in the extraction and synthesis stages.
In extraction, the model identifies which statements are "key." This is a judgment call the model is making based on patterns from its training. It may decide that a hedged, uncertain statement from a witness is less important than a confident one, and exclude or minimize the hedge. A witness who said "I think it was him, I'm not totally sure, but it looked like him" may come out of the extraction stage as "identified the suspect." The uncertainty, which is legally important for both direct examination and impeachment, has been smoothed away.
In synthesis, the model generates a narrative from the extracted material. Narrative generation is where blending and confabulation are most likely. The model is producing coherent prose from a set of data points. If two witnesses described the same event in slightly different ways, the model may combine their accounts into a single composite statement without indicating that it has done so. If a witness described something ambiguously, the model may resolve the ambiguity in the direction that makes the narrative flow better. These are not malfunctions. They are exactly what a language model is designed to do: produce coherent, fluent, organized text. The problem is that in an investigative context, coherence and accuracy are not the same thing, and the model has no way to know that the coherent version it produced is different from what actually occurred.
The Multi-Interview Blending Problem
The scenario in the opening of this lesson illustrates the multi-interview blending problem, which is the failure mode most specific to the investigative context. When an AI tool processes multiple interviews from the same case simultaneously, or when it has access to all of the case material in a single session, it may blend statements from different witnesses into a single witness's summary. The model knows that a vehicle turned left onto Maplewood because a different witness said so. It places that statement in Witness A's summary because it fits the narrative of Witness A's account of the sequence of events. The result is a professionally formatted summary that attributes a statement to Witness A that Witness A never made.
This is not a random error. It is the model doing its job: synthesizing information from the available material to produce the most coherent account. The problem is that in an investigative context, each witness's account must stand alone. The question is not "what happened, synthesized from all available accounts?" It is "what did this specific witness say, in this specific interview, at this specific point in their statement?" Those are fundamentally different questions, and the AI tool that is optimizing for narrative coherence may systematically produce the wrong answer to the second question even while producing a reasonable answer to the first.
The Quote Verification Standard
The verification standard for AI-assisted interview summaries is more demanding than the verification standard for patrol report review, because the stakes attached to a misattributed quote are higher than the stakes attached to an incorrectly sequenced event in a use-of-force narrative. The standard can be stated simply and then worked through in practice.
Every direct quote and every paraphrase attributed to a specific person in an AI-assisted interview summary must be verified against the timestamped recording before that summary is placed in the investigative file.
This is not a quality standard. It is an evidence standard. The investigative file is the foundation of the prosecution's case. Defense discovery includes the entire file. A quote that cannot be verified against the recording does not belong in the file as a quote, and a paraphrase that the model has strengthened, weakened, or blended needs to be corrected before it becomes the version of record.
In practice, the quote verification standard works like this. When the AI produces a summary of a recorded interview, the detective treats every piece of direct or paraphrased speech as provisional. The verification process involves: (1) identifying every statement in the summary that is attributed to a specific person, whether in direct quotation marks or in paraphrase form; (2) using the transcript's timestamps to locate the corresponding passage in the recording; (3) listening to or reading the verbatim transcript of that passage; and (4) confirming that the summary version accurately represents what the person said, including any hedging, uncertainty, or qualification that the person expressed.
The last step, confirming that qualifications and hedges are preserved, is where many reviews fall short. A detective who confirms that a witness said something approximating the words in the summary but does not catch that the model stripped out "I think" or "I'm not sure" has caught a transcription error but missed a legal one. Hedging language is substantive. It affects how a statement can be used at trial. "I saw him do it" and "I think I saw him do it" are different statements for impeachment purposes, and the model will routinely prefer the confident version because it produces better narrative flow.
Practical Workflow for Quote Verification
A practical workflow for quote verification in an AI-assisted interview summary looks like this. Start with the AI-generated transcript, which should be timestamped. Print or display the AI-generated summary alongside the transcript. Go through the summary statement by statement. For each attribution, locate the corresponding transcript timestamp. Read the verbatim text at that timestamp. Compare the verbatim text to the summary version. Mark discrepancies: additions, omissions, smoothed language, strengthened or weakened qualifications. Correct the summary to match the verbatim record. Document the review: the timestamp, the verbatim version, the correction made.
This process is more time-consuming than a standard review of an AI-drafted patrol report, but it is substantially faster than producing the interview summary from scratch. The return on investment is real: the detective who previously spent three days producing interview summaries from raw transcripts may now spend three to four hours using AI to generate the summaries and then another three to four hours verifying them. That is still a significant reduction in labor, and it is a reduction that produces a verified, court-ready document rather than a summary that depends on the detective's memory of twelve hours of audio.
The documentation of the review is important for a different reason. When a defense attorney files a discovery motion that includes all case documents and asks about AI use, the investigative file needs to show not just the AI-generated summary but the verification pass: which statements were checked, when the check was performed, and what corrections were made. That record is the detective's answer to the cross-examination question. "We used an AI tool to produce a first-pass summary. I reviewed every attributed statement against the timestamped recording. Here are the corrections I made. The summary in the file reflects the verified record, not the AI's first draft."
Evidence Summarization Beyond Interviews
The same principles apply when AI tools are used to summarize other forms of evidence: surveillance footage, digital communications, financial records, and large document sets. Each presents its own version of the quote verification problem.
Surveillance footage summarization is emerging as a capability in major investigations. An AI tool that processes hours of surveillance footage from multiple cameras can produce a timeline of events, identify persons of interest, and summarize movements and interactions. The failure mode is the same as in interview summarization: the model may blend footage from different cameras, misattribute an action to a specific person when the footage is ambiguous, or describe a movement as definitive when the footage is unclear. A detective who adopts an AI-generated surveillance summary without verifying the key events against the actual footage has the same problem as the detective who adopts an AI interview summary with an invented quote: the verified record is the footage, not the summary, and the summary must be corrected to match it.
Digital communications, including text messages, emails, and social media content, are frequently summarized in major investigations because the volume of material can be extraordinary. A phone extraction in a drug trafficking case may produce thousands of text messages. An AI tool that summarizes the relevant communications and produces a chronological narrative is genuinely useful. The risk is the same: the model may quote a message inaccurately, combine two separate messages into a single paraphrase, or describe an intent that the message implies but does not state. Every communication attributed in a case summary needs to be traceable to the specific message in the original extraction, with the verbatim text preserved and accessible.
Physical Evidence and Forensic Report Summarization
Physical evidence summaries and forensic report summaries present a different but related risk. When an AI tool summarizes a ballistics report, a toxicology report, or a forensic examination, the failure mode is less likely to be a fabricated quote and more likely to be a misrepresented conclusion. A forensic report that says "consistent with" is a different statement than one that says "matches." A toxicology report that identifies a substance at a concentration "below the threshold for impairment" is a different statement than one that says the subject "was not impaired." Language models that synthesize forensic material are at risk of hardening qualified conclusions: they read "consistent with" and write "matches" because "matches" is a more typical way to describe a positive forensic result in the kind of documents they were trained on.
The verification standard for forensic summaries is: every conclusion in the AI-generated summary must be traceable to the exact language of the underlying forensic report, and the qualified language in the report must be preserved verbatim in the summary. A detective who softens or strengthens a forensic conclusion in a case summary, even if the AI did the softening or strengthening, has produced a misleading document. That document will be disclosed to the defense. The defense will compare it to the original forensic report. The discrepancy will come out at trial.
Disclosure When AI Summarized the Interviews
The investigative file in an AI-assisted case needs to contain more than the AI-generated summaries, even after verification. It needs to contain the documentation of AI use, the nature of the AI tool's involvement, and the verification standard applied. This is the disclosure obligation that Brady v. Maryland and the growing body of prosecutorial-discovery case law requires when AI has touched the evidentiary chain.
What does disclosure look like in practice? The investigative report notes that AI-assisted transcription and summarization tools were used in the processing of witness interviews. It identifies the specific recordings summarized, the tool used, and the review process applied. It states that every attributed statement was verified against the timestamped recording and that corrections were made where the AI summary differed from the verbatim record. It preserves both the original AI output (the first draft) and the verified corrected version (the document of record), so the defense can, if they choose, review the changes and ask questions about the verification process.
Some agencies, following the lead of prosecutors like King County, Washington, which barred AI-written police reports over concerns about accuracy and disclosure, are developing specific disclosure language for AI-assisted investigative documents. The Electronic Frontier Foundation (EFF, a civil liberties organization focused on digital rights) has raised concerns about the lack of transparency when AI tools are used in criminal investigations, and those concerns are not going away. A detective who has built verification and disclosure into the workflow is not just protecting the case. They are answering the EFF concern and the King County concern head-on: "We used the tool. We verified the output. We disclosed the process. The record is clean."
CJIS (Criminal Justice Information Services) Security Policy, which governs how criminal justice data must be handled and secured, also has implications for AI-assisted interview processing. The audio recordings of witness and suspect interviews are criminal justice information subject to CJIS controls. Any AI tool that processes that material must meet the CJIS requirements for data handling, storage, and access. The obligations stay with the agency, not the vendor. A detective who sends interview recordings to a cloud-based AI summarization service without confirming that the service meets CJIS standards has created a data-handling problem, independent of the accuracy question. Check the agency's approved tools list before using any AI platform to process recorded interviews.
Building the Habit: Evidence-First in Every Summary
The most important shift that AI-assisted interview summarization requires from an investigator is not a technical skill. It is a mental model. The AI summary is a starting point, not a destination. It is the tool's best attempt to organize the material, and it will often be close. But "close" in an investigative summary is not "verified," and "verified" is the only standard that holds in a deposition, in a suppression hearing, or in front of a jury.
The evidence-first habit means treating the recorded interview, the verbatim transcript, and the physical or digital evidence as the source of truth, and treating the AI summary as a draft that must be corrected to match the source of truth. It means preserving the hedges, the qualifications, the uncertainties, and the contradictions in witness statements rather than letting the AI smooth them into a clean narrative. It means flagging the places where the AI's synthesis has resolved an ambiguity that the detective cannot independently resolve from the record, and noting those unresolved ambiguities in the summary rather than inheriting the AI's resolution of them.
This habit protects the case. A defense attorney cannot effectively impeach a verified summary, because the verified summary matches the record. A defense attorney who gets the AI's first draft, by contrast, and finds a statement that the recording does not support, has an impeachment opportunity that the detective handed them. The extra hours spent on quote verification are not bureaucratic overhead. They are the work that makes the case file bulletproof at trial.
For investigations involving multiple witnesses, consider building an attribution map as you verify: a simple reference document that lists each statement in the case summary, the witness who made it, the recording it came from, and the timestamp where it can be found. This map becomes the detective's working record of the verification process and the ready answer to any discovery request that asks how AI-generated summaries were reviewed. It also becomes a tool for the prosecutor: a prosecutor who receives a case file with an attribution map can quickly confirm that the statements in the summary match the recordings, and can use the map to prepare witnesses for trial.
Key Takeaways
- AI tools that transcribe and summarize recorded interviews offer real time savings in investigations, but the failure modes in investigative summarization are more serious than in patrol report writing because a fabricated or misattributed witness quote is a Brady and Giglio problem that can sink a prosecution.
- The multi-interview blending problem is the failure mode most specific to investigations: when an AI processes multiple interviews simultaneously, it may place a statement from one witness in a different witness's summary, producing a professionally formatted document that puts words in a witness's mouth.
- Every direct quote and every paraphrase attributed to a specific person in an AI-assisted summary must be verified against the timestamped recording before that summary enters the investigative file. This is an evidence standard, not a quality preference.
- Hedging language matters. A model that smooths "I think I saw him do it" into "I saw him do it" has altered the statement in a way that affects how it can be used at trial. The verification pass must preserve qualifications and uncertainties, not just confirm that the general content is approximately right.
- The same verification standard applies to surveillance footage summaries, digital communications summaries, and forensic report summaries: every conclusion must trace back to the specific source material, and the qualified language of forensic reports must be preserved verbatim.
- Disclosure of AI use in investigative files is a Brady and Giglio obligation. The investigative report must document which tools were used, which material they processed, and what verification standard was applied. Both the AI first draft and the verified corrected version should be preserved in the file.
- CJIS Security Policy governs the handling of recorded interview material. Detectives must confirm that any AI summarization tool used to process witness or suspect recordings meets CJIS requirements before using it. The obligation stays with the agency, not the vendor.
- The evidence-first habit treats the AI summary as a draft to be corrected against the record, not a product to be adopted. The time savings are real, but they must be taken at the draft stage. The verification pass is the work that makes the summary court-ready.
Skill.re