AI for Researchers
Proficient · M14 · lesson 14 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

3.5: Reporting AI Use in Systematic Reviews

15 min

Overview

The transparency imperative in systematic review methodology extends fully to AI-assisted workflows. When AI tools assist with searching, screening, data extraction, quality assessment, synthesis, or writing, readers and peer reviewers need to know which steps were AI-assisted, what the AI produced, how outputs were verified, and what role human judgment played at each stage. Emerging reporting standards, including PRISMA-S extensions for AI-assisted searching, guidelines from major systematic review organizations, and journal-specific policies, are rapidly formalizing what this disclosure must include. Researchers who learn to report AI use rigorously will produce more credible, reproducible work and be better prepared as these standards tighten.

Title

Lesson 3.5: Reporting AI Use in Systematic Reviews

Purpose

This lesson teaches researchers how to transparently report AI use in systematic reviews following emerging standards like PRISMA-S (Preferred Reporting Items for Systematic reviews and Meta-Analyses for Searching), creating comprehensive documentation enabling readers to understand what AI did and didn't do, and supporting reproducibility through detailed methodological reporting.


Why Reporting AI Use Matters for Systematic Review Credibility

Systematic reviews achieve their special epistemic status, sitting at the top of evidence hierarchies, because of methodological transparency: every decision that shaped the evidence base is documented, every included study can be traced, every exclusion can be questioned. When AI tools enter this workflow, the transparency imperative doesn't diminish; it extends.

Readers need to know whether an AI tool was used for title and abstract screening because AI screening has known error rates that affect the completeness of the included evidence base. If AI handled data extraction, readers need to know how accuracy was verified, because systematic extraction errors can distort synthesis findings. If AI drafted the synthesis narrative, readers need to know whether the text was carefully reviewed against source studies, because AI can generate plausible but inaccurate summaries. If AI assisted with writing, questions of authorship, attribution, and responsibility arise.

Beyond reader needs, reporting serves the research community's ability to replicate, evaluate, and extend the review. A review that uses AI tools but provides no information about which tools, which model versions, how they were configured, and what prompts were used is not reproducible, even in principle. Given that AI models change rapidly (updated weights, deprecated versions, changed behaviors), documenting the specific tool version and date of use is as important as documenting the database version searched.

Finally, reporting AI use is increasingly an ethical obligation. Systematic reviews inform clinical guidelines, policy decisions, and resource allocation. Failing to report AI use that could have introduced systematic errors or biases is a form of methodological misrepresentation that undermines the review's validity and the decisions made based on it.

What Needs to Be Reported at Each Workflow Stage

Reporting requirements differ by workflow stage, reflecting the different ways AI can influence review quality at each point.

For search assistance: Report the AI tool used (name, version, date of use), the nature of assistance (synonym generation, Boolean logic structuring, database-specific query translation), and how the AI-generated search strategy was validated (e.g., tested against a set of known relevant studies, reviewed by an information specialist). Do not present an AI-assisted search strategy as if it were developed entirely by a human librarian.

For title and abstract screening: Report the tool name and version, the calibration procedure (how many training examples were provided, what the sensitivity/specificity was on a calibration set), the human review procedure for AI-uncertain cases, and the procedure for auditing AI exclusions (sample size of audit, estimated false exclusion rate). Present the screening section of the PRISMA flow diagram as AI-assisted, with explicit notation.

For data extraction: Report whether AI performed primary extraction, what fields were extracted by AI versus human, the accuracy verification procedure (spot-check percentage, error rate on checked extractions), and any systematic errors identified and corrected.

For quality/risk of bias assessment: Report whether AI extracted quality-relevant information from papers to support human assessment or made quality judgments autonomously (the former is acceptable; the latter requires very careful documentation of verification procedures).

For synthesis: Report whether AI drafted any component of the narrative synthesis, what review procedures were applied to AI-drafted text, and how the final synthesis text was verified against source studies.

For writing: Report whether AI assisted with drafting the Methods, Results, or Discussion sections, whether AI-generated text was reviewed and modified, and whether any verbatim AI text appears in the manuscript.

PRISMA-S and Emerging Reporting Standards

PRISMA-S (Preferred Reporting Items for Systematic reviews and Meta-Analyses for Searching) is an extension of the PRISMA 2020 guidelines specifically addressing search reporting, and several of its items are directly relevant to AI-assisted searches. Item 22 of PRISMA-S covers the reporting of search assistance, including AI tools. At minimum, PRISMA-S-compliant AI search reporting should include: the specific tool and version, the nature of assistance, any validation steps, and the final database-specific search strategies (which must be full, reproducible strings, not just a description of the approach).

The Cochrane Collaboration has issued guidelines on AI use in Cochrane reviews, requiring disclosure of AI tool use at each review stage and specifying that AI cannot replace the human judgment required for eligibility assessment or risk of bias assessment. These guidelines are being updated as AI capabilities evolve, and researchers should check for current versions before submitting to Cochrane.

Several major journals in health sciences have implemented specific author declaration requirements for AI use: Nature, BMJ, JAMA, and Lancet all require that AI assistance be disclosed in the Methods section. Some journals require disclosure in the author contributions statement; others require a dedicated AI use statement. Checking the target journal's current policy before submission is essential, as these policies are actively evolving.

JBI (Joanna Briggs Institute) systematic review guidance and the Campbell Collaboration have both addressed AI use in their methodological handbooks, generally following a framework of: disclose fully, verify independently, and ensure human accountability for all methodological decisions. The consensus across these organizations is that AI tools cannot be listed as authors and that human researchers bear full responsibility for the review's conclusions regardless of what AI assistance was used.

Practical Documentation Practices During the Review

Retrospective documentation of AI use, trying to remember after the review is complete which tools were used and how, produces incomplete and potentially inaccurate reporting. Prospective documentation, maintained throughout the review, is both easier and more accurate.

The most practical approach is to maintain a running AI use log, a document separate from the review data that records: date of use, tool name and version (or model identifier), the exact prompt or instruction used, the AI's output (or a representative sample), what verification step was applied, and what the researcher's final decision or modification was. This log is the source material for the Methods section and also serves as the audit trail for any queries about the review's methods.

For screening tools specifically, record the calibration procedure, the performance metrics on the calibration set, and the date of any updates to the calibration training data. If the AI screening tool was used over multiple weeks, record any changes in tool behavior you observed and how you responded.

For AI-drafted text components, maintain a version history that shows the original AI output alongside the researcher's edited version. This demonstrates that the final text is the researcher's intellectual work and that AI output was reviewed and modified, not accepted verbatim.

Finally, record the model or tool version at the time of use. AI models are updated frequently, and a model queried in March may behave differently than the same-named model queried in June. Version documentation enables future researchers to understand the specific tool capabilities that generated your outputs and to assess whether the tool's behavior may have changed in ways relevant to the review.

How to Write About AI Use in Your Methods Section

The Methods section description of AI use should be specific, honest, and structured to address three questions: what AI did, what humans did, and how outputs were verified. Generic language like 'AI tools were used to assist with the review process' provides no useful information and will increasingly be insufficient for peer review.

For abstract screening, a well-formed reporting statement might read: 'Title and abstract screening was conducted using [Tool Name, version X.X, accessed [date]], calibrated with [n] confirmed include and [n] confirmed exclude examples. The tool achieved [sensitivity]% sensitivity and [specificity]% specificity on a [n]-record calibration set. AI-uncertain cases (classified with probability between 0.3 and 0.7) were reviewed by [reviewer]; a random sample of [n] AI-excluded records was independently reviewed and showed an estimated exclusion error rate of [x]%.'

For data extraction: 'Initial data extraction for [specified fields] was AI-assisted using [tool]. A second reviewer independently verified [x]% of extracted records, with disagreements resolved through [procedure]. The error rate on verified extractions was [x]%.'

For synthesis drafting: 'Preliminary drafts of the narrative synthesis sections were generated by [AI tool] based on structured summaries of included studies. Each draft paragraph was reviewed against the source studies by [reviewer] and substantially revised to correct inaccuracies and integrate verbatim quotations. No AI-generated text appears verbatim in the final manuscript.'

These specific statements do several things: they enable readers to assess the risk introduced by AI assistance, they demonstrate that verification occurred, they document the specific tool and configuration, and they make clear that human judgment remained in control of the final output. They also create a clear record in case of future queries about the review's methods.

Authorship, Attribution, and Responsibility

The question of AI authorship in systematic reviews has been settled by all major journals and research organizations: AI tools cannot be listed as authors. Authorship carries responsibilities, responding to queries, standing accountable for the work, consenting to publication, that AI systems cannot fulfill. This consensus is universal and appears stable regardless of how capable AI tools become.

However, the consensus on authorship coexists with a growing recognition that the nature of human intellectual contribution is changing. When AI assists substantially with multiple review stages, the human authors' contributions must be clearly delineated to clarify what each person actually did. Author contribution statements should specify which authors were responsible for which AI-assisted stages, what verification each author conducted, and who bears accountability for the quality of those stages.

Responsibility is the key principle: human authors are fully responsible for the review's methodology and conclusions, regardless of what AI assistance was used. This means that if AI made systematic errors in screening that affected the included evidence base, the human authors are responsible for having had inadequate verification procedures. If AI-drafted synthesis text contains inaccuracies that appear in the final publication, the human authors are responsible for insufficient review. This full accountability framework is why the verification procedures documented in the Methods section are not merely procedural requirements. They are the evidence that human researchers discharged their responsibility to ensure the quality of AI-assisted outputs.

Finally, consider the downstream use of your review. Clinical guideline developers, policy analysts, and future systematic reviewers will use your review as a source. The credibility of their work depends partly on the credibility of yours. Transparent, thorough reporting of AI use is not just about satisfying journal requirements; it is about contributing to an evidence ecosystem in which the quality of each component is clearly labeled so that consumers can calibrate their reliance appropriately.