AI for Pharma & Life Sciences
Proficient · M28 · lesson 28 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Mapping a Regulated Workflow for AI Integration
📖
now learning

Mapping a Regulated Workflow for AI Integration

15 min

A program director hands you a 14-week Clinical Study Report development cycle for a Phase 3 oncology pivotal and asks a deceptively simple question: where, in this process, are we allowed to let the AI work? Most teams answer that question by tool, naming Certara CoAuthor or the Veeva Vault RIM AI Agent and declaring that the tool is "approved" or "not approved." That is the wrong axis entirely. A regulated AI integration is not approved at the level of the tool; it is approved at the level of the step. The same model that is fully defensible turning a TLF shell into a first-draft narrative is flatly prohibited from authoring the benefit-risk integration in Module 2.5.6, and the difference is not the software, it is the nature of the cognitive work and who owns the conclusion. This lesson teaches you to process-map a regulated workflow at step granularity, to classify each step as AI-ready, AI-assistive, or AI-prohibited, and to locate the human handoff at every transition, because a workflow you cannot map at this resolution is a workflow you cannot validate, cannot audit, and cannot defend on Day 74.

Why the Unit of Analysis Is the Step, Not the Tool

The instinct to govern AI at the level of the tool comes from the way enterprise software has always been procured and validated, where you qualify a system once and then use it broadly. Generative AI breaks that model because the same model invoked on two different steps of the same workflow carries radically different risk, and the risk lives in the task, not the binary. When you ask the model to convert a finalized Table 14.2.1.4 into a paragraph of efficacy prose, the ground truth exists, it is loaded, and verification is a reconciliation against a fixed source. When you ask the same model to weigh the magnitude of a progression-free survival benefit against a numerically higher rate of grade 3 hepatotoxicity and render an integrated benefit-risk conclusion, there is no source table to reconcile against, because the conclusion is the thing being created, and a named human must own it under the FDA-EMA accountability principle.

This is why the foundational move of Level 3 is to refuse the tool-level question and insist on the step-level question. You are not asking "is CoAuthor validated for Module 2.5"; you are asking "for step 7 of the 2.5 production cycle, what is the intended use of AI, what is the ground truth it reconciles against, who verifies, and who signs." A single tool may be AI-ready on step 3, AI-assistive on step 9, and AI-prohibited on step 14, all within one document's lifecycle. The validation artifact you will build in the IQ/OQ/PQ lesson later in this chapter is scoped to the step and its intended use, not to the product name, and an inspector who asks "where did AI touch this submission" expects an answer at step resolution, not a vendor list.

The practical consequence is that your first deliverable as an AI-integrated submission professional is a process map, and the map must be drawn at the granularity where the cognitive nature of the work changes. Too coarse, and you hide a prohibited judgment step inside an "AI-assisted drafting" block. Too fine, and you drown in transitions that carry no real handoff. The right resolution is the resolution at which the answer to "does ground truth exist for this step, and who owns the output" changes, and learning to draw the map at exactly that resolution is the skill this lesson builds.

Process-Mapping the 14-Week CSR Development Cycle

Take the pivotal oncology CSR cycle and lay it out as a sequence of discrete steps with named inputs and named outputs. Weeks 1 to 2 are the database-lock-to-TLF-shell phase, where the Statistical Analysis Plan, the finalized analysis datasets, and the TLF shells exist but the populated tables do not. Weeks 3 to 5 are TLF production and dry-run review, where the statistical programming team generates and QCs the actual Tables, Listings, and Figures against the SAP. Weeks 6 to 9 are the CSR body draft, where the medical writer builds Sections 9 through 14 of the ICH E3 report, the efficacy and safety results sections, against the now-finalized TLF package. Weeks 10 to 11 are SME and co-author review, where clinical, statistical, and clinical-pharmacology subject-matter experts mark up the draft. Weeks 12 to 13 are quality control and reference QC, where every cross-reference is reconciled and consistency across sections is checked. Week 14 is finalization, sign-off, and the handoff to the Module 2 summary writers and the publishing team.

Each of these phases decomposes further into steps, and it is at the step level that classification happens. Inside weeks 6 to 9, for example, there is a step that takes a finalized efficacy table and produces a first-draft results paragraph, a different step that drafts the interpretive text linking that result to the clinical context, and a third step that drafts the integrated discussion in Section 13 that weighs efficacy against safety. These three steps live in the same phase, are touched by the same writer, and may even use the same tool, yet they sit in three different risk classes because the ground truth available to each is different. The mapping discipline forces you to see that the phase label "CSR body draft" conceals a spectrum from pure reconciliation to pure judgment, and the AI integration must respect that spectrum rather than the phase boundary.

The map is not a flowchart for decoration; it is the spine of every downstream control. The audit trail you write to Veeva Vault QualityDocs is organized by step. The verification gate you design in the next lesson sits at the transition between steps. The IQ/OQ/PQ acceptance criteria you write attach to the intended use of AI at a specific step. When you later build the end-to-end Module 2.5 workflow in Chapter 2, you will discover that it is literally this map, extended downstream into the summary lifecycle, with the same three-class logic applied. A team that skips the mapping discipline and jumps straight to "let CoAuthor draft the 2.5" has no spine to hang controls on, which is precisely why their AI use collapses under an Information Request.

The Three Classes: AI-Ready, AI-Assistive, AI-Prohibited

An AI-ready step is one where ground truth exists in a loadable, fixed form, the AI's task is to transform that ground truth into a different representation, and verification is a deterministic reconciliation against the source. Turning a finalized TLF table into a first-draft narrative paragraph is the canonical AI-ready step, because Table 14.2.1.4 is the source of truth, the model's job is to render its contents as prose, and the writer verifies by reconciling every number, population, and direction in the paragraph against the cell it came from. AI-ready does not mean unverified; it means the verification is bounded, mechanical, and complete, because there is a fixed answer to check against. These are the steps where AI delivers its largest, safest acceleration, and they are the steps you push hardest to integrate.

An AI-assistive step is one where the AI augments human cognition on a task that has no single fixed answer but does have checkable properties, and the human remains the author of the judgment while the AI surfaces, organizes, or stress-tests. A cross-section consistency check is the canonical AI-assistive step: the AI reads Sections 11 and 13 of the CSR and flags that a hazard ratio stated in the efficacy results does not match the value carried into the discussion, but the AI does not decide which value is correct, it raises the discrepancy for a human to adjudicate against the TLF. The AI is a second set of eyes that never tires across 600 pages, which is a genuine strength, but the resolution of every flag is a human act. AI-assistive steps are where most teams over-trust, because the output looks like an answer when it is actually a prompt for human judgment.

An AI-prohibited step is one where the output is a regulated conclusion that a named individual must own, where no loadable ground truth exists because the conclusion is the artifact being created, and where delegating the cognition to a pattern-completer would violate the accountability principle and, in substance, the meaning of authorship. The integrated benefit-risk conclusion in Module 2.5.6 is the defining example, and the limits-of-reasoning lesson in Level 1 established why: the model can structure the argument, lay out the efficacy magnitude and the safety profile in a balanced frame, and even draft the connective prose, but the act of concluding that the benefit outweighs the risk in the indicated population is a judgment the sponsor's named signatory renders and defends, not a token the model sampled. Causality determinations in pharmacovigilance, comparability conclusions under ICH Q5E, and the final benefit-risk integration share this property: the AI may build the scaffold, but a human owns the verdict, and the step is prohibited from AI authorship of the conclusion itself.

The Test That Classifies a Step

Classification is not a matter of taste, and you need a repeatable test so that two trained professionals classify the same step the same way. The test has three questions asked in order. First, does loadable ground truth exist for this step, a fixed source against which the output can be reconciled? If yes, and the task is transformation of that source into another form, the step is at least a candidate for AI-ready. Second, if no single fixed answer exists, does the step nonetheless have checkable properties, such as internal consistency, completeness against a checklist, or alignment to a named structure like ICH E3, where the AI can surface candidates for a human to adjudicate? If yes, the step is AI-assistive. Third, is the output of this step a regulated conclusion that a named human must own and defend, with no source to reconcile against because the conclusion is itself the artifact? If yes, the step is AI-prohibited for the authoring of that conclusion, regardless of how capable the model is.

Apply the test to a borderline case to see how it disciplines judgment. Drafting the Section 11 efficacy results text seems like pure transformation, hence AI-ready, and the numeric rendering portion is. But the same step often includes a sentence of interpretation, "these results demonstrate a clinically meaningful improvement," and that clause has no source table to reconcile against, because clinical meaningfulness is a judgment. The disciplined response is to split the step: the numeric rendering is AI-ready and the model drafts it, while the interpretive clause is AI-assistive at most, where the model may propose language but the writer owns the claim, and any conclusion that the benefit is clinically meaningful in a way that drives the benefit-risk verdict migrates toward prohibited. The test, applied honestly, forces you to split steps that hide a class boundary inside them, which is the single most common error in naive AI integration.

The test also protects you from the opposite error, which is reflexively prohibiting AI from steps that are genuinely AI-ready because they feel important. Generating the first-draft narrative of a primary efficacy table feels high-stakes, and it is, but the stakes are managed by reconciliation, not by exclusion, because the ground truth exists and the verification is complete. Treating an AI-ready step as if it were prohibited wastes the largest safe acceleration available to a regulated writer and pushes teams toward shadow AI use that escapes the audit trail entirely. The goal is not maximal caution; it is correct classification, because a correctly classified AI-ready step is both faster and more defensible than the manual alternative, while a misclassified prohibited step is a credibility catastrophe waiting for an inspector.

The Human Handoff at Each Transition

Between every two steps in the map there is a transition, and a transition in a regulated AI workflow is never a silent data pass; it is a handoff with a named owner, a verification act, and a logged record. The handoff is where the human takes possession of the AI's output, performs the verification appropriate to the step's class, and either accepts the output into the controlled record or rejects and reworks it. At the transition out of an AI-ready table-to-narrative step, the handoff is a reconciliation: the writer confirms every numeric and structural claim against the source table and signs the reconciliation before the paragraph advances. At the transition out of an AI-assistive consistency check, the handoff is an adjudication: the human resolves every flag the AI raised, deciding which value is correct against the TLF, and records the disposition of each flag. At the transition into an AI-prohibited step, the handoff is an exclusion boundary: the AI's contribution stops at the scaffold, and the human authors the conclusion, with the record showing that the conclusion originated with the named author.

The reason the handoff must be explicit and logged is that the entire defensibility of the workflow rests on demonstrating, after the fact, that a competent human stood between the model and the submission at every point where it mattered. An inspector reading the Veeva Vault QualityDocs audit trail on Day 74 is not reassured by "AI-assisted, human-reviewed" stamped on the document; they want to see, step by step, what the AI produced, what the human did with it, who that human was, and when. The handoff is the unit at which that evidence is created. A workflow with vague handoffs produces an audit trail that says human review happened without showing what it consisted of, and that is indistinguishable, to an inspector, from no review at all. The next lesson in this chapter is devoted entirely to designing these handoffs as formal verification gates, because the handoff is where Level 3 lives.

There is a subtler design point in the handoff: the verification effort must match the step's class, and a well-mapped workflow allocates human attention accordingly. Spending equal verification effort on every step is both wasteful and dangerous, because it under-protects the high-risk transitions and over-burdens the low-risk ones, driving writers to cut corners everywhere. The map lets you concentrate scrutiny where the class demands it, light reconciliation on a bounded AI-ready transformation, heavy adjudication on an AI-assistive consistency sweep, and full human authorship at a prohibited boundary. This is the operational meaning of the FDA-EMA risk-based principle applied inside a single document's lifecycle: risk is assessed and controls are sized at the step, and the map is what makes that sizing possible.

Where the Map Lives in the Audit Trail

A process map that lives only in a slide deck is governance theater. The map becomes a control when it is bound to the records system, and in a Veeva Vault QualityDocs environment that binding is concrete. Each AI-touched step corresponds to an action in the document lifecycle, and the audit trail captures, for that action, the model and version, the system prompt identity, the sources loaded, the human verifier, the verification disposition, and the timestamp, organized so that the sequence of actions reconstructs the map. When the AI use rises to the level that disclosure is warranted, a summary of the workflow and its controls is written to the Module 1.14 administrative section as the sponsor's account of how AI participated in producing the submission, and that summary is coherent precisely because it is generated from a map that already classifies every step.

This is why the mapping discipline is the prerequisite for everything downstream and not merely a planning nicety. The verification gate design in the next lesson attaches gates to the transitions the map defines. The IQ/OQ/PQ validation in the third lesson writes acceptance criteria against the intended use the map assigns to each AI-ready and AI-assistive step. The end-to-end Module 2.5 workflow in Chapter 2 is this map extended and instantiated in a real tool stack with a real audit trail. A team that has internalized the step-level, three-class, handoff-at-every-transition discipline can answer the inspector's Day 74 question in one sentence and then substantiate it line by line; a team that governs by tool name can only gesture at "approved software," which is the answer that loses an Information Request and, with it, the credibility of the submission.

Key Takeaways

  • Govern AI at the level of the step, not the tool. The same model is fully defensible on a TLF-to-narrative transformation and flatly prohibited on the Module 2.5.6 benefit-risk conclusion, because risk lives in the cognitive nature of the task and in who owns the output, not in the software's name.
  • Process-map the 14-week CSR cycle at the resolution where ground truth and ownership change. The phase label "CSR body draft" hides a spectrum from pure reconciliation to pure judgment, and the map must be drawn fine enough to expose every class boundary inside a phase.
  • Classify every step as AI-ready, AI-assistive, or AI-prohibited using the three-question test. Does loadable ground truth exist; does the step have checkable properties for a human to adjudicate; or is the output a regulated conclusion a named human must own? Honest application forces you to split steps that hide a class boundary.
  • Design an explicit, logged human handoff at every transition. Reconciliation out of AI-ready steps, adjudication out of AI-assistive steps, and an exclusion boundary into AI-prohibited steps, each with a named owner and a recorded disposition, because "human-reviewed" without evidence is indistinguishable from no review to an inspector.
  • Bind the map to the records system so it becomes a control, not a slide. The Veeva Vault QualityDocs audit trail is organized by step, the Module 1.14 AI disclosure is generated from the classified map, and the map is the spine on which the verification gates, IQ/OQ/PQ criteria, and the end-to-end Module 2.5 workflow all hang.