โ†
AI for Researchers
Visionary ยท M10 ยท lesson 10 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
๐Ÿ“–
in this lesson

3.2: AI and the Future of Peer Review

15 min

The Peer Review Crisis: Capacity, Burnout, and Systemic Failure

Peer review is the mechanism by which scientific communities validate knowledge claims, and it is operating under conditions of structural stress that predate AI but are being dramatically amplified by it. Understanding the precise nature of this crisis is essential for institutional leaders who participate in editorial governance and whose faculty carry significant reviewer burdens that affect productivity and wellbeing.

The numbers are stark. An estimated 3.5 million manuscripts are submitted to academic journals annually, and each submitted manuscript typically requires two to three reviewers. Even accounting for desk-rejection rates of 50 to 70 percent at selective journals, the system still requires approximately 1.8 million individual reviews per year to function. But reviewer response rates to invitation emails average around 25 percent, meaning journals must invite four reviewers for every one who actually agrees to review. The implied pool of reviewer invitations is therefore closer to 7 million per year. This burden falls disproportionately on senior researchers, whose names appear most frequently in reviewer databases, and who are simultaneously facing increasing administrative and grant-writing demands.

Average review time, from submission to first editorial decision, runs three to six months at most journals. In fast-moving fields like AI, machine learning, and infectious disease, a six-month review cycle means that by the time a paper is accepted, the field may have moved significantly beyond the contribution. This mismatch between review speed and research velocity is one of the primary drivers of the preprint movement: researchers post preprints on arXiv, bioRxiv, or SSRN immediately upon completion to establish priority and gather community feedback while the formal review process grinds forward.

Reviewer burnout is not a soft or anecdotal problem. It has measurable effects on review quality and editorial system functioning. Studies of editorial database records show that individual reviewers are declining an increasing proportion of review invitations over time, and that review quality is declining as reviewers who do accept are stretched across more simultaneous commitments. The pandemic accelerated these trends, as researchers absorbed additional institutional duties, and post-pandemic recovery has been partial at best.

AI is being proposed simultaneously as the primary solution to and the primary threat to peer review. As solution: AI can automate routine elements of review, pre-screen manuscripts for obvious methodological problems, match reviewers to manuscripts more effectively, and generate preliminary review frameworks that reduce the cognitive load on human reviewers. As threat: AI enables researchers to produce more manuscripts faster, amplifying the submission volume that already overwhelms the system, and AI-generated reviews could displace the expert human judgment that makes peer review valuable. Institutional leaders navigating this landscape need clarity on which applications of AI are beneficial, which are harmful, and what policies are needed to govern the transition.

Current AI-Assist Applications: From Spell-Check to Statistical Audit

Before examining the more transformative and contested applications of AI in peer review, it is important to understand the applications that are already in widespread use and that represent a relatively uncontroversial first layer of AI-assisted review. These applications address specific, well-defined elements of the review process where automation provides clear value and the risks of error are manageable.

Language and clarity review tools are the most widely adopted category. Grammarly Academic, Writefull, and TXYZ.AI provide manuscript-level language quality assessment that is particularly valuable for non-native English-speaking authors, who have historically faced documented discrimination in peer review processes based on language quality rather than scientific merit. When journals provide these tools to authors before submission, they reduce the language-related workload on peer reviewers, who can focus on scientific substance rather than grammatical correction. Several publishers, including Springer Nature and Wiley, now offer Writefull integration as part of their submission portal.

Reference checking automation addresses a problem that has plagued the literature for decades: incorrect citations. Studies have estimated that 10 to 20 percent of citations in published papers contain errors ranging from wrong page numbers to completely misrepresented claims about what the cited paper demonstrated. Automated DOI verification catches papers cited that do not exist or have been retracted. Citation context checking, matching the claim made about a cited paper against the actual content of the cited paper, is a more sophisticated AI application that several journals are beginning to pilot. This catches the particularly problematic class of citation errors where a paper is cited to support a claim it does not actually support.

Statistical review automation is arguably the most impactful current application. The StatReviewer platform, developed specifically for biomedical journals, automated checks of statistical reporting against established guidelines: checking that effect sizes are reported with confidence intervals, that multiple comparisons are appropriately corrected, that sample sizes match descriptions across the paper. The GRIM (Granularity-Related Inconsistency of Means) test, implemented as an automated tool, identifies impossible mean values given reported sample sizes and integer-constrained measurements, a check that has caught numerous fraudulent data sets. SPRITE (Sample Parameter Reconstruction via Iterative TEchniques) similarly checks whether reported statistical parameters are consistent with plausible underlying datasets.

Methodology auditing, checking that the methods described in a paper actually match the claims made in the results and discussion, is an emerging AI application with significant potential. Large language models can be instructed to compare the methods section of a manuscript against its results and flag apparent inconsistencies. While false positive rates are currently too high for this to replace human review, using AI-flagged inconsistencies as prompts for human reviewers focuses expert attention on the most likely problem areas.

Plagiarism and self-plagiarism detection through iThenticate and Crosscheck are now standard at most major journals and many conference proceedings. These tools have significantly reduced the burden of identifying copied text, though sophisticated text manipulation can still evade detection. The AI arms race between detection and evasion is an ongoing concern that institutions need to monitor.

AI as Primary Reviewer: Experiments, Results, and Honest Assessment

The most provocative and contested application of AI in peer review is its use as a primary or co-equal reviewer, a system in which an AI generates a structured evaluation of a manuscript that either replaces or substantially informs the human review. Several major publishers and conference organizers conducted experiments in this space in 2024 and 2025, and the results provide important data for institutional leaders formulating policy positions.

Nature Portfolio experimented with LLM-generated preliminary review reports for a subset of submissions in 2024. The experimental design was carefully controlled: manuscripts were reviewed independently by both LLM and human reviewers, and editorial decisions were made based on human reviews alone, with the LLM reports analyzed separately. The findings were mixed. LLM reviewers performed well on formal aspects of manuscripts, identifying missing methodological details, flagging absent statistical information, noting inconsistent terminology, but performed significantly worse than expert humans on evaluating the significance and novelty of the contribution. The false positive rate for methodology flaws was notably high: AI reviewers frequently flagged methodological choices as problematic when they were actually appropriate for the research context, generating noise that would have required significant human effort to filter.

PLOS ONE piloted automated technical review for statistical reporting in 2024, using an AI system to check statistical methodology against PLOS ONE's explicit reporting requirements before manuscripts were sent to human reviewers. This narrower application, AI reviewing against a defined checklist rather than providing open-ended scientific judgment, proved significantly more reliable. Manuscripts returned to authors with AI-identified statistical reporting gaps came back revised in ways that reduced the statistical review burden on human reviewers by approximately 30 percent in the pilot.

The ACL 2024 conference (Association for Computational Linguistics) conducted a carefully designed experiment with GPT-4-generated reviews for a sample of submitted papers. Reviewers were asked to evaluate both human and AI-generated reviews for their usefulness, and submitting authors were asked to evaluate the feedback they received. The findings aligned with Nature's: AI reviews were rated as more thorough in coverage of formal aspects but less useful in identifying the central scientific contribution and evaluating its significance. Authors found AI feedback useful for identifying presentation problems but insufficient for determining whether their core scientific claims were sound.

The key finding across these experiments is that AI peer review, in its current form, functions better as a pre-screening and formatting tool than as a scientific judgment system. The aspects of peer review that require deep domain expertise, evaluating whether a claimed finding is actually novel, whether a methodology is appropriate for a specific research question, whether the interpretation of results is warranted, remain beyond current AI capability in the sense that AI produces confident-sounding but unreliable judgments in these areas. Institutional leaders should be cautious about any publisher or system claiming that AI review has solved the reviewer capacity problem.

Bias Perpetuation in AI Peer Review: Systematizing Historical Inequity

Among the most serious risks of AI adoption in peer review is the potential to systematize and scale existing biases in scholarly communication. Peer review has always contained bias, against women, against researchers from non-elite institutions, against non-English-speaking authors, against qualitative and interpretive research methods, and AI systems trained on historical review data inherit and potentially amplify these biases. Institutional leaders need to understand the specific mechanisms through which this happens.

Training on historical reviews encodes historical biases in a direct technical sense: if an AI peer review system is trained on a corpus of historical review decisions and the text of historical reviews, it learns the patterns in those decisions, including patterns that reflect bias rather than scientific merit. If qualitative research in a field has historically received more skeptical reviews than quantitative research, a well-documented pattern in psychology, education research, and some health sciences, an AI trained on that history will be more skeptical of qualitative submissions. If papers from non-elite institutions have historically received more negative reviews even controlling for quality, also documented, an AI trained on that history will reproduce that pattern.

Non-English-speaking authors face compounding disadvantages in an AI peer review environment. Language models assess writing quality as a proxy for other qualities, and writing in English as a second language creates systematic markers that AI systems may associate with lower-quality reasoning even when the underlying science is equivalent. While tools like Writefull partially address this by helping authors improve language quality before submission, the underlying pattern of AI systems rating non-native English writing as less sophisticated remains a bias to monitor.

Institutional prestige signaling is a subtler bias pathway. In nominally blind review, clues to institutional affiliation often appear in the text: methodology descriptions that reference specific platforms, collaboration acknowledgments, citations to the authors' own prior work, even stylistic patterns that cluster by training environment. AI systems capable of identifying authors from writing style with approximately 80 percent accuracy in published studies can make double-blind review functionally meaningless, and if the AI uses institutional affiliation signals in its evaluation, it will perpetuate prestige bias.

The risk of systematizing existing bias at scale is the critical concern: individual human reviewers who are biased introduce that bias into one review; an AI system that is biased introduces it into every review it generates. The scale of AI peer review application could entrench biases that decades of diversity, equity, and inclusion efforts in academic publishing have been slowly eroding. Institutional leaders whose faculty include researchers from historically underrepresented groups, who study populations underrepresented in mainstream research, or who work in methodological traditions that have been historically undervalued should be particularly attentive to this risk and should actively advocate within professional societies and journal editorial boards for bias auditing requirements before AI peer review tools are deployed at scale.

Reviewer-Manuscript Matching with AI: Where the Technology Works

While AI-as-reviewer remains contested, AI-assisted reviewer-manuscript matching is one of the clearest success stories in the application of AI to peer review infrastructure. The bottleneck of identifying suitable reviewers for specialized manuscripts, particularly in interdisciplinary fields and rapidly evolving areas, has been a persistent editorial burden that AI can genuinely reduce.

Semantic Scholar's reviewer recommendation system uses dense vector representations of manuscript content and researcher publication histories to identify reviewers whose published work is semantically similar to a submitted manuscript. Unlike keyword-based matching, which fails when relevant researchers use different terminology for the same concepts, semantic matching captures conceptual similarity across vocabulary differences. A paper on 'transformer attention mechanisms applied to protein folding' might match semantically with researchers publishing on 'neural network architectures for structural biology' even if those exact terms don't appear in each other's papers.

The Toronto Paper Matching System (TPMS), developed at the University of Toronto and adopted by several major conferences including NIPS and ICML, uses topic modeling and author profile analysis to generate reviewer-paper match scores. TPMS has been shown to identify well-matched reviewers significantly faster than human editors working from memory and personal networks, and to surface relevant reviewers outside the editor's immediate awareness, particularly valuable for identifying reviewers from different geographic regions or disciplinary sub-communities than the editor typically interacts with.

Deep learning approaches to reviewer recommendation outperform keyword matching in several specific ways. They handle papers at the intersection of multiple fields better, because they can identify researchers with expertise in both areas even when those researchers don't appear in any single keyword search. They handle new sub-field terminology better, because they learn from usage patterns rather than requiring manual curation of keyword lists. And they generalize across language variation in a way that keyword matching cannot.

Conflict of interest detection is an area where AI adds particular value. Collaboration network analysis can systematically identify co-authorship relationships, institutional co-affiliation periods, grant co-investigatorship, and other relationships that constitute conflicts of interest, and do so across a researcher's full career history, not just the last five years that editors typically track manually. AI tools that maintain and query collaboration graphs can flag potential conflicts before reviewer invitations are issued, reducing the embarrassing and ethics-problematic situations that arise when conflicts are disclosed after review.

The practical recommendation for institutional leaders who sit on journal editorial boards or lead society publications is to advocate for the adoption of AI reviewer matching while establishing clear requirements for how matched reviewer lists are presented, as recommendations for human editorial judgment, not as automated assignments, and for how conflicts of interest are verified before invitations are sent.

Open Peer Review, AI Transparency, and the Future of Scholarly Accountability

The rise of open peer review, in which reviews are published alongside papers, authors and reviewers are known to each other, and the review process is visible rather than hidden, creates a new context for AI integration in peer review with distinctive accountability implications. The tension between transparency norms in open peer review and the opacity of many AI systems creates challenges that institutions need to anticipate.

Signed peer review, in which reviewers are identified alongside their reviews, has been shown to correlate with more collegial and constructive reviewer behavior, reviewers who know their names will appear are less likely to be gratuitously negative. When AI assists with or generates elements of signed reviews, however, accountability becomes complex: if a human reviewer's signed review was substantially generated by an AI, to whom is the review accountable? Nature's policy on AI disclosure in peer review requires that reviewers declare AI tool use, but the specific implications for attribution of review content remain unclear.

Public review histories, as practiced by journals including eLife, EMBO Journal, and F1000Research, make the full review correspondence between editors, reviewers, and authors visible to readers after publication. This transparency creates a powerful accountability mechanism: readers can evaluate whether review identified the paper's actual weaknesses, whether author responses adequately addressed reviewer concerns, and whether editorial decisions were justified. When AI is incorporated into these documented processes, its contributions need to be captured in the public record. eLife's 'publish, then review' model, in which preprints are published immediately and then receive community review, creates an environment where AI pre-screening of methodological issues could be deployed before community review begins, reducing the burden on community reviewers while maintaining the transparency of the review record.

Nature's AI policy for peer review, which requires disclosure of AI tool use by both authors and reviewers, represents the most explicit current framework from a major publisher. The policy requires that reviewers disclose if they used AI assistance in preparing their reviews and prohibits sharing confidential manuscript content with AI systems, a policy that has practical implications for the use of general-purpose AI assistants like Claude or ChatGPT in reviewing, since submitting a manuscript to these systems potentially violates confidentiality agreements.

For institutional policy, leaders need to develop clear guidance for their faculty who serve as peer reviewers: which AI tools are permitted in the review process, what must be disclosed to editors, and how to avoid confidentiality violations. The range of acceptable use needs to be explicit: using AI for grammar checking of a review you have already written is categorically different from submitting the manuscript content to an AI and using its output as your review, and institutional policy should distinguish between these clearly.

Preprint Ecosystems and AI: Screening, Discovery, and the New Publishing Landscape

The growth of preprint repositories, bioRxiv, medRxiv, arXiv, SSRN, ChemRxiv, and others, represents the most significant structural change in scholarly communication in decades. Preprints provide immediate dissemination of research findings before peer review, establishing priority, enabling rapid community feedback, and supporting the pace of research in fast-moving fields. AI is reshaping preprint ecosystems in multiple ways, some of which institutional leaders need to actively shape through policy and resource allocation.

BioRxiv's AI screening for methodological red flags before posting represents an early deployment of AI as a quality floor rather than a quality ceiling. The screening does not assess scientific significance, bioRxiv explicitly does not peer review, but it automatically checks for obvious integrity issues: suspicious image duplications using forensic image analysis tools, statistical patterns inconsistent with claimed experimental design, and reference to retracted papers as supporting evidence. This basic screening prevents the most egregious integrity violations from appearing on a platform that is widely read before peer review, without attempting to replace the substantive assessment that peer review provides.

arXiv's use of AI for moderation is more extensive and has a different focus: identifying submissions that are off-topic for the claimed subject category, detecting duplicate submissions across arXiv categories, and flagging submissions that appear to be generated primarily by AI without original scientific contribution. The arXiv moderation challenge is significant. It receives over 200,000 new submissions per year across physics, mathematics, computer science, quantitative biology, and economics, and its volunteer moderators are a finite resource. AI-assisted moderation allows the human moderator pool to focus on judgment calls rather than routine categorization.

SSRN, the primary preprint server for economics, law, and social science, has integrated AI-enhanced discovery features that allow readers to find papers relevant to their interests beyond simple keyword search. Using semantic similarity to connect papers on conceptually related topics across disciplinary silos, SSRN's AI discovery features are increasing the cross-disciplinary reach of social science preprints, a meaningful impact metric in itself.

The rise of AI-powered preprint recommendation systems represents a potential reshaping of how researchers keep up with the literature. Systems that track a researcher's reading patterns and publication history and recommend newly posted preprints on relevant topics are already operating at scale through services like Semantic Scholar Feeds and connected tools. For institutional leaders, the implication is that the traditional journal-subscriptions-plus-search model of literature engagement is being supplemented by algorithmic curation, and that the institutions and tools doing the curating will have increasing influence over which research gets read and cited.

Double-Blind Review in the AI Era: The Unraveling of Anonymity

Double-blind peer review, in which the identities of both authors and reviewers are concealed from each other, has been advocated as a mechanism for reducing bias in publication decisions. The evidence for its effectiveness is mixed, but it remains the standard practice at many journals in the social sciences, education research, and some natural science fields. AI capabilities for de-anonymization are fundamentally challenging whether double-blind review can remain meaningful.

AI de-anonymization of research manuscripts is more accurate than most researchers appreciate. Studies of stylometric analysis, the identification of individual writing style fingerprints, using modern language models have demonstrated author identification accuracy of approximately 80 percent from manuscript text alone, even when authors attempt to obscure their writing style. The signals that give away authorship include vocabulary distribution, sentence complexity patterns, citation network characteristics (authors cite their own prior work, and even when specific titles are redacted, the pattern of the research they cite reveals their location in the literature), and methodological signatures (researchers have recognizable methodological preferences that persist across papers).

The practical implication is that double-blind review, at journals where AI-assisted reviewer tools are available, may be functionally single-blind or open review in practice, regardless of formal policy. A reviewer using an AI assistant to help evaluate a manuscript may inadvertently or deliberately allow the AI to identify authors from the manuscript text, rendering the blinding meaningless. This is not hypothetical. It is a predictable consequence of deploying AI tools in review processes where de-anonymization capability has not been audited.

Institutional response strategies need to reckon honestly with this reality. Options include: accepting that double-blind review is functionally over and moving to fully open review with appropriate accountability mechanisms; implementing technical controls that prevent AI tool use by reviewers (practically difficult to enforce and potentially counterproductive given the legitimate benefits of AI review assistance); or redesigning double-blind submission processes to remove more identifying information: redacting acknowledgments, funding sources, institutional affiliations of collaborators, and prior work citations entirely, which creates significant preparation burden on authors.

New protocols for truly blind review in the AI era would need to go beyond current practices in ways that most journals have not yet seriously planned. The scholarly communication community is at an early stage of grappling with these implications, and institutional leaders whose faculty are active in journal editorial governance should be raising these questions in editorial board conversations now, rather than after the issue has caused a significant bias incident.

Institutional Policy on AI in Peer Review: Governance That Protects Integrity

Institutional leaders face the challenge of developing AI peer review policy before the scholarly communication community has reached consensus, but waiting for consensus to emerge before developing institutional policy is also a strategy with costs, since faculty are already using AI tools in ways that may create problems if they are not guided. A proactive institutional policy framework, developed with input from faculty governance and research integrity offices, is better than either prohibition or laissez-faire permissiveness.

The first policy dimension is permitted uses: what AI tools may faculty use when serving as peer reviewers? A principled framework distinguishes between AI assistance for comprehension, using AI to clarify unfamiliar terminology, to check one's understanding of an analytical technique, to review relevant literature that will inform the review, and AI assistance for evaluation, submitting the manuscript content to an AI and using its output as the basis for the review. The former is analogous to consulting a reference book or a colleague, which has always been permitted. The latter raises questions of confidentiality, attribution, and whether the review represents the reviewer's own expert judgment.

The second policy dimension is disclosure requirements: what must faculty disclose to editors when they have used AI assistance in preparing a review? Institutional policy should align with the policies of major journals in the faculty member's field and should encourage transparency. A simple statement of principle, 'disclose AI tool use of any kind in peer review to the editor handling the manuscript', is better than attempting to enumerate specific permitted and prohibited tools, which becomes outdated quickly.

Confidentiality is the most legally significant policy dimension. Peer review involves receiving confidential research findings that have not yet been published. Submitting those findings to a commercial AI service whose terms of service allow training on submitted content potentially violates the confidentiality obligations reviewers accept when they agree to review. Institutions that provide AI services through enterprise agreements with privacy protections, as many universities now provide through Microsoft 365 Copilot, Google Workspace AI features, or similar enterprise tools, can create a cleaner confidentiality boundary than the use of general-purpose commercial AI tools.

Implications for reviewer credit systems represent an emerging governance challenge. Several initiatives, including ORCID's reviewer recognition feature, Publons, and F1000's open reviewer recognition, are building systems for formally crediting peer review contributions in ways that can appear in academic CVs and promotion dossiers. If AI substantially generates review content, the question of whether the reviewer has performed a creditable scholarly service analogous to human review becomes significant. Institutional promotion and tenure policies should begin addressing this question before it becomes a contested case.