โ†
AI for Pharma & Life Sciences
Aware ยท M12 ยท lesson 12 of 17 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
The Five AI Modalities You Will Use Before NDA Filing
๐Ÿ“–
now learning

The Five AI Modalities You Will Use Before NDA Filing

15 min

People talk about "using AI" in a submission as if it were one thing, the way people once talked about "using a computer." It is not one thing. Between the day a Phase 3 database locks and the day an NDA is filed, a regulatory and clinical team will lean on five mechanically different kinds of AI, and each one fails in its own way, saves time in its own way, and places the line of human accountability in a different place. A medical writer drafting a Module 2.5 Clinical Overview is using a different modality than a safety associate triaging an incoming case, who is using a different modality than the literature analyst pulling adverse-event mentions out of four hundred PubMed abstracts, who is using a different modality than the publishing lead asking the Veeva Vault RIM AI Agent to assemble a sequence. If you cannot tell these apart, you cannot reason about where the risk lives, and you certainly cannot write the AI section of the cover letter. This lesson names the five modalities, anchors each to a real artifact on the road to filing, and tells you exactly where the human has to stand in each one.

Modality One: Generation, the Modality That Drafts the Dossier

Generation is the modality most people mean when they say "AI." It is a model producing new text: a draft Module 2.5.4 efficacy section from a CSR, a Clinical Study Protocol section from a target product profile, a case narrative from intake fields, a cover-letter paragraph from a submission plan. The output is original prose that did not exist before, assembled token by token from the statistical patterns of the training corpus and whatever sources sit in the context window.

Generation is the highest-value modality and the highest-risk one, for the same reason: it creates content rather than sorting or finding it. When a writer asks Certara CoAuthor or the enterprise model to draft the efficacy summary, the tool can produce a structurally perfect section in seconds, and it can also invent a hazard ratio, a table cross-reference, or a subgroup claim with exactly the same fluency. The value and the danger are inseparable, because both flow from the model's willingness to produce a plausible continuation whether or not it is grounded.

The human accountability line in generation sits at every factual and structural claim. The model proposes; the named author disposes, by reconciling each claim to a source before the text is allowed to stand. This is why generation is the modality that demands the most verification discipline per output: the writer is not editing a found fact, they are validating a manufactured one. On the road to NDA filing, generation touches the CSR-to-2.5 mapping, the protocol, the Investigator's Brochure update, the ICF, the briefing book, the 483 response, and the cover letter. It is everywhere a human used to face a blank page, and the blank page is exactly where invention is most tempting and least checkable.

Modality Two: Classification, the Modality That Sorts the Inflow

Classification is the model assigning an input to one of a fixed set of categories. It does not write; it labels. The canonical pharma example is incoming case triage in pharmacovigilance: is this report serious or non-serious, expected or unexpected, a valid ICSR or noise? Under ICH E2D, seriousness and expectedness are defined categories with regulatory consequences, and a classifier can read an incoming report and propose the label in milliseconds, where a human takes minutes.

Classification feels safer than generation because the output is bounded; the model can only choose among the allowed labels, not invent a new one. But the bounded output hides a subtler risk: a confident wrong label. A classifier that tags a serious case as non-serious does not produce an obvious fabrication; it produces a clean, plausible category that happens to be wrong, and a wrong seriousness call collapses the entire downstream clock, because a serious unexpected case may be a 15-day expedited report and a non-serious one is not. The error is invisible precisely because the output looks like every correct output.

The human accountability line in classification sits on the consequential boundary cases and on the audit of the label distribution over time. A safety physician does not re-review every case the classifier calls non-serious, but the workflow must surface the borderline calls for human confirmation and must monitor whether the classifier's serious-rate is drifting against the historical baseline. Classification also appears in eCTD section assignment, in protocol-deviation major-versus-minor sorting, and in MSL insight categorization. In each, the discipline is the same: trust the classifier for volume, verify it on the boundary, and watch the distribution for drift.

Modality Three: Extraction, the Modality That Reads at Scale

Extraction, often called named-entity recognition or information extraction, is the model pulling specific structured facts out of unstructured text. Out of a four-hundred-abstract literature surveillance queue, it identifies which abstracts mention the product, an adverse event, a patient, and a causal link. Out of a case report, it pulls the suspect drug, the reaction term, the dose, the onset date. Out of a protocol, it pulls the eligibility criteria into a structured table. The output is not new prose and not a single label; it is a set of fields lifted from the source and mapped onto a schema.

Extraction is the workhorse modality of pharmacovigilance and literature monitoring under the volume pressures of 2026, where a single PV team may face hundreds of articles a week and an ICSR queue running over a hundred cases a day. It is genuinely transformative because it turns reading from a linear human bottleneck into a parallel machine pass. But extraction has two characteristic failures: it can miss a fact that is present, a false negative, and it can assert a fact the source does not support, a false positive. In a safety context the false negative is the one that ends careers, because a missed adverse-event signal in the literature is a missed signal, full stop.

The human accountability line in extraction sits on recall for safety-relevant facts and on the rejection rationale for what was filtered out. When a tool triages four hundred abstracts down to ten relevant cases, the defensible workflow does not just keep the ten; it records why each of the three hundred ninety was rejected, so that an inspector can see the surveillance was systematic and not lucky. The named MedDRA coder still owns the Lowest Level Term assignment; the extractor proposes, the coder decides. Extraction proposes structure; the human owns the consequences of what the structure left out.

Modality Four: Retrieval, the Modality That Grounds the Answer

Retrieval is the modality that fetches relevant source material and puts it in front of the model, or in front of you, so that an answer is grounded in your documents rather than in the model's general training. Its most important form is retrieval-augmented generation, where a system searches a store of your Investigator's Brochures, prior submissions, CSRs, and SOPs, pulls the relevant passages, and feeds them into the context window before the model generates. Retrieval is what lets a model answer "what did we commit to at the End-of-Phase-2 meeting" from your actual meeting minutes rather than from a plausible guess.

Retrieval is the modality that, done well, suppresses the failures of the others, because grounding a generation in real retrieved sources is the single most effective way to reduce hallucination. But retrieval has its own quiet failure mode, and it is one of the most dangerous in the whole stack: retrieving the wrong source and grounding confidently in it. If the retriever pulls a superseded version of an SOP, or a draft IB instead of the approved one, or last cycle's protocol amendment, the model will produce an answer that is fluent, grounded, cited, and wrong, and it will be harder to catch than a free hallucination because it carries the authority of a real citation to a real document that simply should not have been used.

The human accountability line in retrieval sits on the provenance and currency of what was retrieved. The question is never only "is the answer supported by a source," it is "is it supported by the right source, the current and approved one." This is why a Level 3 chapter is devoted to RAG over the submission library and to detecting when grounding fails. For the Level 1 learner, the rule is to treat a citation as a starting point for checking, not as proof, and to confirm that the retrieved document is the version of record before trusting an answer built on it.

Modality Five: Agentic AI, the Modality That Takes Steps

Agentic AI is the modality where the model does not just produce one output but plans and executes a sequence of steps, often calling tools, querying systems, and chaining its own outputs into the next action. The Veeva Vault RIM AI Agents, which moved into the Vault platform from late 2025 with the RIM rollout planned for 2026, can take a submission-planning intent and assemble correspondence, tag documents, and propose a sequence structure across multiple steps. The Medable CRA Agent can take a monitoring context and work through site data to surface what a CRA should look at. An agent is generation, classification, extraction, and retrieval wired together with the authority to act, not just to answer.

Agentic AI is the modality with the highest leverage and the most diffuse accountability, and the second property is the dangerous one. When a single output is wrong, you inspect one output. When an agent takes nine steps and the third step quietly retrieved a superseded document, the error is laundered through six downstream steps before it reaches you, and the final result looks like the clean product of a competent process. The accountability question shifts from "is this output correct" to "is each step in this process inspectable, logged, and reversible." An agent that cannot show its work is an audit-trail problem before it is anything else.

The human accountability line in agentic AI sits on the design of the checkpoints, not on the final output alone. A defensible agentic workflow has named human gates at the consequential transitions, logs every step in a form that satisfies 21 CFR Part 11, and never lets the agent close a step whose reversal would be impossible. On the road to NDA filing, agentic AI is the newest and fastest-moving modality, and it is precisely where the FDA-EMA principles of accountability and human oversight bite hardest, because the technology's whole appeal is that it reduces the number of times a human touches the work. The discipline is to put the human touches back exactly where the consequences concentrate.

Why Naming the Modality Is the First Skill

The reason to learn these five names is not taxonomy for its own sake. It is that the right verification step is different for each modality, and a writer who cannot name the modality will apply the wrong control. You verify a generation by reconciling every claim to a source. You verify a classification by auditing the boundary cases and watching the label distribution for drift. You verify an extraction by checking recall on the safety-critical facts and recording the rejection rationale. You verify a retrieval by confirming the provenance and currency of the retrieved source. You verify an agentic workflow by inspecting the logged steps and the human checkpoints. Apply a generation control to an extraction problem and you will check the wrong thing while the real failure slips past.

Most real tools blend modalities, and that is exactly why naming them matters. When a PV platform reads a case, classifies its seriousness, extracts the MedDRA terms, retrieves the product's listedness from the reference safety information, and generates the narrative, it has used four modalities in one pass, and a defensible review checks each at its own seam. The skill you are building is the habit of decomposing a tool's behavior into its modalities and placing the right human check at each one. That habit is what separates a writer who "uses AI" from a writer who can stand in front of an inspector and account for it.

Why the Modalities Blur, and Why That Is the Trap

The reason this taxonomy is worth memorizing is that the tools you actually buy do not announce their modalities. They present a single smooth surface, a box you type into and a result that comes back, and the modalities are wired together underneath where you cannot see the seams. This is convenient and it is the trap, because the seams are exactly where the failures live, and a tool that hides them hides its own risk surface. When a literature-surveillance platform shows you ten relevant cases out of four hundred abstracts, you are looking at the output of an extraction pass, a classification pass, and possibly a generation pass that wrote the summary, and you cannot apply the right control unless you mentally pull those passes apart.

Consider a concrete blur that bites in practice. A safety platform is asked to produce a case narrative, and a key field, the time to onset, is missing from the structured intake. A pure extraction tool would leave the field blank or flag it as missing, which is the safe behavior. But a tool that silently slides from extraction into generation will fill the gap with a plausible interval, because generation abhors a blank, and now an invented onset time sits inside a narrative that reads as if every field were extracted from the source. The failure is invisible precisely because the modality switched without telling you. The defense is a discipline that follows directly from the taxonomy: require a source locator for every extracted field, and treat any field that cannot point to a location in the source as missing rather than filled. That rule only occurs to a reviewer who knows that extraction and generation are different modalities with different trust profiles.

This is why the first skill is decomposition and the second is suspicion of the smooth surface. A tool that proudly cannot tell you which modality produced which part of its output is a tool you cannot fully account for to an inspector, and that is a procurement signal as much as a technical one. The most defensible tools are the ones that expose their seams, that show you the retrieved source, the extracted field with its locator, the classification with its confidence, and the generated text marked as generated. Transparency about modality is not a nicety; it is the precondition for placing the human correctly, which is the whole job.

The Five Modalities on the Road to Filing

Picture the last six weeks before an NDA lock as a single integrated scene, and you can see all five modalities working at once. In medical writing, generation is drafting the Module 2.5 and 2.7 summaries while a writer reconciles every TLF citation. In safety, classification is triaging the daily ICSR inflow while a physician confirms the borderline seriousness calls, and extraction is mining the weekly literature while an analyst records why each rejected article was set aside. In regulatory operations, retrieval is grounding the answers to internal questions in the prior-submission library while a publisher confirms each retrieved document is the version of record. And across the whole submission, an agentic Vault RIM workflow is assembling the sequence while a human owns the checkpoints and the logs.

No one of these is "the AI." They are five tools with five risk profiles and five places where a named human has to stand. The cover letter that discloses AI involvement, the subject of a Level 2 lesson, is in essence a statement of which modalities were used where and how each was verified. You cannot write that letter honestly until you can name the modalities, and you cannot file with confidence until you have placed the human correctly in each one. That is the whole point of starting here.

Key Takeaways

  • Five mechanically distinct AI modalities touch a submission before NDA filing: generation (drafting the 2.5, protocol, narratives), classification (ICSR seriousness triage under ICH E2D), extraction (literature and case mining), retrieval (grounding answers in your submission library), and agentic AI (multi-step assembly via tools like Veeva Vault RIM AI Agents and the Medable CRA Agent).
  • Each modality fails differently, so each needs a different control. Generation invents plausible claims; classification produces confident wrong labels; extraction misses present facts or asserts unsupported ones; retrieval grounds confidently in the wrong or superseded source; agentic AI launders an early error through later steps.
  • The human accountability line moves with the modality. Reconcile claims (generation), audit boundary cases and watch drift (classification), guard recall and record rejection rationale (extraction), confirm provenance and currency (retrieval), and inspect logged steps and checkpoints (agentic).
  • Retrieval's quiet failure is the most deceptive: a fluent, cited, grounded answer built on a superseded SOP or a draft IB carries the authority of a real citation and is harder to catch than a free hallucination. Confirm the version of record before trusting an answer.
  • Naming the modality is the first operational skill, because real tools blend several in one pass. A PV platform may classify, extract, retrieve, and generate in a single case; a defensible review decomposes the behavior and places the right human check at each seam. The AI section of the cover letter is, in effect, that decomposition written down.