End-to-End AI Workflow for Literature Surveillance
A pharmacovigilance scientist opens the weekly literature queue and finds four hundred and twelve articles returned by the standing search across PubMed, EMBASE, and the regional databases the company is obligated to monitor. Three of those articles describe a possible new adverse-reaction signal for a marketed biologic, and they are buried among hundreds of duplicates, off-label discussions, pharmacokinetic studies with no case-level data, and review articles that mention the product in passing. The task is to find every article that contains a reportable individual case or a safety signal and to reject the rest with a documented rationale, within the week, every week, without ever missing the one case that matters. This is the structural inversion that makes literature surveillance unlike almost every other AI task in the program: here the governing metric is recall, the fraction of truly relevant articles the workflow catches, and a workflow that is fast, precise, and confident but misses a single reportable case has failed at the one thing it exists to do. This lesson designs the end-to-end AI workflow for literature surveillance with recall as the structural control, the per-article rejection audit log as the defensibility spine, and the EudraVigilance Article 57 literature-monitoring overlay as the regulatory frame.
Why Recall Is the Governing Metric, Not Precision
Most AI tasks in regulated writing optimize implicitly for precision, the fraction of the workflow's outputs that are correct, because a fabricated citation or a wrong number is the failure that hurts. Literature surveillance inverts that priority, because the cost structure is asymmetric in the opposite direction: a false positive, an irrelevant article that the workflow flags for human review, costs a few minutes of a reviewer's time, while a false negative, a reportable case that the workflow rejects as irrelevant, is a missed adverse-event report that the company was legally obligated to detect and submit. The regulatory expectation is that the marketing-authorization holder monitors the scientific literature and identifies every individual case and every signal it contains, and the failure that draws a Good Vigilance Practice finding is the missed case, not the over-inclusion. This is why recall, the fraction of truly relevant articles the workflow catches, is the metric the entire workflow is engineered around, and why a high-precision, low-recall AI triage that produces a clean, confident shortlist is precisely the dangerous configuration: it looks efficient and it loses cases.
The consequence for the workflow's design is that the AI is positioned as a recall-protective filter rather than a decision-maker, and the human review effort is concentrated on the boundary cases rather than spread thinly across all four hundred and twelve. The model's job is not to decide which three articles are reportable; its job is to triage the queue so that no relevant article is silently dropped, which means the workflow must be tuned to over-include rather than under-include, accepting a higher false-positive rate as the deliberate price of protecting recall. An AI triage configured to maximize precision will exclude the ambiguous article, the case report buried in a methods section, the off-label use that nonetheless describes a serious reaction, and each of those exclusions is the exact failure the workflow exists to prevent. The workflow therefore treats every AI exclusion as a claim that must be defensible, and it treats recall, measured against a validated reference set, as the acceptance criterion that the workflow must meet before it can be trusted.
Query Construction and the Database-Coverage Problem
The workflow begins before any article is triaged, at the construction of the search query itself, because recall is determined first by whether the relevant article was retrieved at all, and an article that the query never returned cannot be triaged, accepted, or rejected; it is simply invisible. The query must be built against the right databases, PubMed and EMBASE for the global literature, and the regional and local databases the company is obligated to monitor for products marketed in specific jurisdictions, because a product authorized in a market whose local literature is not searched has a coverage gap that no downstream triage can repair. AI assists query construction by expanding the product and substance terms to include synonyms, brand and generic names, salt forms, and known misspellings, by incorporating the relevant reaction and class terms, and by mapping the query to each database's indexing vocabulary, EMBASE's Emtree and PubMed's MeSH, so that the search captures the indexed and the free-text mentions alike. The danger is that an over-narrowed query silently shrinks the recall ceiling: a query that omits a synonym or a database is a recall failure that occurs before triage and is invisible in the triage metrics, because the missed article never entered the queue.
The workflow therefore treats query construction as the first recall control, validating the query against a known set of relevant articles to confirm that the query retrieves them, and documenting the database list, the search terms, the indexing-vocabulary mappings, and the search dates so that the coverage of the search is itself auditable. This matters because an inspector evaluating the literature-monitoring system reads the search strategy as the foundation of the entire process, and a search strategy that cannot be shown to cover the obligated databases and the full term space is a deficiency in the system, not merely in a week's queue. The AI accelerates the construction of a comprehensive query, but the adequacy of the database coverage and the completeness of the term space are judgments a qualified scientist confirms, because the recall ceiling is set here, at the query, before the model ever reads an abstract.
Relevance Triage and the Per-Article Rejection Audit Log
With the queue retrieved, the workflow performs AI-assisted relevance triage, classifying each article by whether it plausibly contains a reportable individual case or a safety signal, and routing the queue into accept, reject, and uncertain bands so that human review concentrates where the model is least confident. The triage is genuinely useful at this volume, because the model is good at recognizing that a pharmacokinetic modeling study with no patients contains no case, or that a duplicate of an already-processed article is a duplicate, and clearing those high-confidence rejections frees the scientist to focus on the ambiguous remainder. The discipline that makes the triage defensible is that every rejection is logged with its rationale, because the regulatory requirement is not merely to find the relevant articles but to be able to demonstrate, article by article, why each rejected article was rejected, and a literature-monitoring system that cannot produce a rejection rationale for a challenged article cannot defend its exclusions.
The per-article rejection audit log is therefore the defensibility spine of the workflow, and it is the artifact an inspector asks for: for any article in the week's retrieval that was not escalated, the log records the article identifier, the triage decision, the rationale for rejection, the model and version that proposed it, and the human reviewer who confirmed or overrode the rejection. The log converts the workflow's exclusions from silent disappearances into documented, traceable decisions, which is exactly what distinguishes a defensible literature-monitoring system from an efficient one that cannot prove what it discarded. The boundary the workflow draws is that high-confidence rejections may be confirmed in review at the band level, but every uncertain article and every article a downstream check flags is individually adjudicated by a qualified scientist, and the rejection of any article that later proves to contain a case must be reconstructable from the log so the system can be corrected and the gap explained. A literature-monitoring workflow without a per-article rejection log is not a regulated workflow; it is an unaccountable filter.
ICSR-Eligible Case Extraction and the Eligibility Determination
For every article the triage escalates as plausibly relevant, the workflow performs case extraction, reading the full text to identify whether it contains an individual case that meets the minimum criteria for an ICSR, an identifiable patient, a reporter, a suspect product, and an event, and extracting the case elements when it does. AI assists this extraction by reading the article and pulling the candidate case details, the patient characteristics, the suspect and concomitant products, the reaction, the dose, and the outcome, into a structured form that feeds the ICSR workflow downstream, and this is a real acceleration over manual full-text reading of every escalated article. The model is good at recognizing the presence of a case and assembling its elements, and a structured extraction that hands a clean case skeleton to the ICSR intake process compresses the time from literature retrieval to case creation substantially.
The eligibility determination, however, is a judgment the workflow keeps human-owned, because deciding that an article does or does not contain a reportable case is the literature equivalent of the seriousness determination in the ICSR workflow: a model that reads a case report and concludes it describes no identifiable patient, when the patient is identifiable enough by the article's standards to constitute a valid case, has rejected a reportable case under the guise of an extraction result. The workflow therefore positions the model as the extractor and surfacer of candidate cases, with a qualified scientist confirming the ICSR eligibility of each escalated article and the completeness of the extracted case, and with the eligibility decision and its rationale recorded. The extraction also feeds the duplicate check, because a case described in a literature article may already exist in the safety database from a spontaneous report, and the workflow surfaces candidate duplicates for human confirmation rather than auto-merging, because a wrongly merged case loses information and a missed duplicate inflates the count, the same trade-off the ICSR workflow confronts.
Measuring Recall and the Reference-Set Problem
Recall is only a usable control if it is measured, and measuring recall requires a reference set, a curated collection of articles whose true relevance is known, against which the workflow's accept and reject decisions can be scored to compute the fraction of truly relevant articles the workflow caught. Building that reference set is the unglamorous foundation of a defensible literature-monitoring system, because the recall number that the validation spec depends on is only as trustworthy as the reference set that produced it, and a reference set that under-represents the hard cases, the buried case reports, the off-label serious reactions, the ambiguous abstracts, will report a flattering recall that the live queue does not deliver. The workflow therefore constructs the reference set to over-weight the boundary cases that are the model's known failure modes, so that the measured recall reflects performance on exactly the articles most likely to be lost, rather than on the easy rejections the model handles trivially.
The reference-set discipline also governs how the workflow is monitored over time, because a model's triage behavior can drift as the literature shifts, as new products enter the portfolio, and as the term space evolves, so recall is not measured once at validation and assumed thereafter but re-measured on a defined cadence and after any material change to the model, the query, or the scope. This is the ongoing-monitoring expectation of the FDA-EMA principles applied to the literature workflow: the recall that qualified the system is a claim about a moment, and a system that does not re-measure recall against a maintained reference set has stopped knowing whether it still catches the cases it was validated to catch. The workflow records each recall measurement, the reference set version it used, and the date, so that the recall history is itself an auditable record of the system's continued fitness for purpose.
The Article 57 Overlay, the Validation Gate, and What the Inspector Asks
The EudraVigilance Article 57 literature-monitoring overlay shapes the workflow for products within its scope, because the EMA conducts centralized monitoring of the medical literature for a defined list of active substances and enters the resulting cases into EudraVigilance directly, which means a marketing-authorization holder whose substance is on that list must not duplicate those cases and must reconcile its own monitoring against the EMA's coverage to avoid double-reporting while still covering the literature the EMA does not. The workflow therefore incorporates the scope of the EMA's centralized monitoring as a routing rule, distinguishing the literature the holder must monitor and report from the literature the EMA covers, so that the holder's obligations are met without duplicate submission. This overlay is a configuration the workflow must encode explicitly, because a holder that mis-scopes its monitoring either double-reports cases the EMA already entered or, far worse, assumes the EMA covers literature it does not and leaves a true gap, and the gap is the failure that recall exists to prevent.
The Level 3 deliverable is the validated workflow and the audit trail that make the literature-monitoring system defensible to a Good Vigilance Practice inspection, not a single week's triage. The validation spec states the intended use, the monitoring of the obligated literature for reportable cases and signals, the fitness-for-purpose statement aligned to the FDA-EMA principle, and the acceptance criteria, which are dominated by a measured recall against a validated reference set of known-relevant articles, alongside query-coverage completeness, per-article rejection-log completeness, and Article 57 scope correctness. The audit trail captures the search strategy with its databases, terms, and dates, the full retrieval, the triage decisions with rationales, the model and version, the human confirmations and overrides, the extracted cases, and the duplicate reconciliations. The human handoff is unambiguous: the model retrieves, triages, and extracts, but a qualified scientist owns the eligibility determinations, confirms the boundary rejections, and signs the week's monitoring, with recall as the metric the system is validated against and the rejection log as the evidence of what was discarded and why. When an inspector asks the holder to demonstrate that the literature for a product was monitored completely and that a specific rejected article was correctly excluded, the search strategy, the recall validation, and the per-article rejection log already hold the answer. The model triages the queue; the named scientist owns the recall and certifies that nothing reportable was lost.
Key Takeaways
- Recall is the governing metric, and the dangerous configuration is high precision with low recall. A false positive costs a reviewer minutes; a false negative is a missed adverse-event report the company was obligated to detect, so the workflow is deliberately tuned to over-include and is validated against a measured recall on a reference set of known-relevant articles.
- The recall ceiling is set at query construction, before any article is triaged. An article the query never retrieves is invisible and cannot be recovered downstream, so AI expands the term space across synonyms, brand and generic names, and indexing vocabularies like Emtree and MeSH, while a qualified scientist confirms database coverage and validates the query against known-relevant articles.
- The per-article rejection audit log is the defensibility spine. The regulatory requirement is not only to find the relevant articles but to demonstrate, article by article, why each rejected article was rejected, so the log records the identifier, decision, rationale, model and version, and human confirmation, converting silent exclusions into traceable decisions.
- Case extraction accelerates the work, but the ICSR-eligibility determination stays human-owned. The model reads escalated articles and assembles candidate case elements, but deciding that an article does or does not contain a reportable case is the literature equivalent of the seriousness determination, so a qualified scientist confirms eligibility, completeness, and candidate duplicates.
- The Article 57 overlay and the validation gate make the system defensible to a GVP inspection. The EudraVigilance centralized-monitoring scope is encoded as a routing rule to avoid both double-reporting and a true coverage gap, and measured recall, query coverage, rejection-log completeness, and Article 57 scope are the record that answers an inspector on whether a product's literature was monitored completely.
Skill.re