โ†
AI for Pharma & Life Sciences
Visionary ยท M14 ยท lesson 14 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Cross-Functional AI Pilots: Discovery to Regulatory
๐Ÿ“–
now learning

Running Cross-Functional AI Pilots: Discovery to Regulatory

15 min

In the second quarter of 2026, a Chief Medical Officer, a Chief Data Officer, and a Head of Regulatory Affairs sat in a steering review and discovered that their organization was running thirty-one AI pilots, and that not one of them had a written kill criterion. Discovery had a generative chemistry pilot that everyone loved and no one could retire. Translational had a biomarker-clustering pilot that had quietly become a production dependency without ever passing a validation gate. Clinical had three central-monitoring pilots from three vendors solving the same QTL-excursion problem. Regulatory had a Module 2.5 drafting pilot whose outputs were already in two live submissions, with an AI use-log that no one was confident would survive an Office of New Drugs Information Request on Day 74. Thirty-one pilots, zero discipline, and a board that had just asked for the AI ROI number. This lesson is about the governance pattern that turns that chaos into a portfolio: how a Level 5 leader runs cross-functional AI pilots that span discovery, translational, clinical, and regulatory, with scope discipline, named kill criteria, and named scale criteria, so that the pilots that work can scale into a GxP-defensible enterprise and the pilots that do not can be retired before they become liabilities you cannot see.

Why the Pilot Is the Unit of Enterprise AI Risk

The instinct of most organizations is to treat an AI pilot as a low-stakes experiment, a sandbox where the normal rules are relaxed so that innovation can happen. In a regulated drug-development enterprise this instinct is precisely inverted, because the pilot is the moment at which AI first touches a GxP artifact, and a pilot without governance is not a safe experiment, it is an undocumented production system waiting to be discovered by an inspector. The biomarker-clustering pilot that became a production dependency is the canonical failure: it was never validated because it was "just a pilot," and it was never retired because it had become load-bearing, so it occupied the worst possible quadrant, a system that the organization relied on and could not defend. The FDA-EMA Guiding Principles of Good AI Practice, released jointly on 14 January 2026, are explicit that AI used to generate, review, or support regulatory content is in scope for accountability and lifecycle management from the first use, which means the pilot is not outside the regulatory floor. The Level 5 leader's first move is therefore conceptual: the pilot is the unit of enterprise AI risk, and it must be governed as a controlled experiment with a defined intended use, a defined scope, a defined success measure, and a defined exit, not as an unmonitored playground.

This reframing has an immediate organizational consequence. If the pilot is the unit of risk, then the count of ungoverned pilots is a risk metric the board should see, in the same way that open Form 483 observations or overdue CAPAs are risk metrics the board sees. The thirty-one-pilot discovery was not a sign of innovation health; it was a sign of governance debt, and the leader who can name that debt and put a structure around it is doing the actual work of enterprise transformation. The remainder of this lesson builds that structure: a pilot charter that scopes the experiment, a cross-functional design that spans the lifecycle without losing accountability at the handoffs, kill criteria that let you stop, and scale criteria that let you grow without inheriting undocumented risk.

The Pilot Charter: Scope Discipline as the First Control

Every governed pilot begins with a one-page charter, and the charter is not bureaucracy, it is the single artifact that prevents the most common pilot failure, which is scope creep into a regulated decision the pilot was never designed or validated to make. The charter names the intended use in fit-for-purpose language drawn directly from the FDA-EMA "fitness for purpose" principle, and it draws a hard boundary around what the AI is allowed to touch. A discovery generative-chemistry pilot might be scoped to "propose candidate structures for human medicinal-chemist review, with no structure advancing to synthesis without named-chemist sign-off," and that boundary is what keeps the pilot from quietly becoming the decision-maker. A regulatory Module 2.5 drafting pilot might be scoped to "produce draft efficacy-summary text with every claim source-linked to a TLF table, for named-writer reconciliation, with no AI-drafted text entering a submission sequence until the claim-reconciliation log is complete and signed." The scope is the control, and the writer of the charter is exercising the most underrated leadership skill in enterprise AI, which is the discipline to say what the pilot will not do.

The charter also names the artifacts the pilot produces and the regulatory regime each artifact lives under, because a pilot that produces an internal slide deck and a pilot that produces a paragraph destined for a Module 2.5 Clinical Overview sit at opposite ends of the risk spectrum even if they use the same underlying model. A useful charter taxonomy mirrors the enterprise AI policy categories that a later lesson develops in full: internal-use (no GxP artifact touched), regulated (the output becomes part of a submission, a CSR, an ICSR, or any record an inspector may read), and high-risk (the output influences a benefit-risk, causality, or comparability conclusion, or falls under EU AI Act high-risk-system obligations). The charter assigns the pilot to one of these categories on day one, because the category determines the evidence the pilot must generate, the validation rigor it must meet to scale, and the kill criteria that govern it. A pilot whose category is unclear is a pilot whose risk is unmanaged, and the leader's job is to force that clarity before the first prompt is ever run.

Designing a Cross-Functional Pilot Without Losing Accountability at the Handoffs

The pilots that create the most enterprise value are the ones that span the lifecycle, because the value of AI in drug development is rarely contained inside one function. Consider a pilot that threads a single mechanism-of-action narrative from discovery through to regulatory: a generative model proposes a target hypothesis in discovery, a translational team uses AI to cluster the biomarker evidence that supports the hypothesis, a clinical team uses AI to draft the protocol's scientific rationale from that evidence, and a regulatory team uses AI to assemble the Module 2.4 nonclinical overview and the Module 2.5 introduction that carry the mechanism story into the dossier. The cross-functional pilot is powerful precisely because the same evidence flows end to end, but that power is also its danger, because accountability tends to evaporate at the handoffs. The translational team assumes discovery validated the hypothesis; the clinical team assumes translational verified the biomarker clustering; the regulatory writer inherits a narrative four hands deep and cannot trace any single claim to its source. This is the cross-functional version of the fabricated-cross-reference problem, scaled across an organization.

The design discipline that prevents this is the named handoff gate, borrowed directly from the human-AI handoff pattern that Level 3 teaches for a single workflow and elevated to the cross-functional level. At each functional boundary, the pilot defines what is handed off, who verifies it, what is logged, and who signs, so that the regulatory writer at the end of the chain receives not a four-hands-deep narrative but a traceable evidence package in which every upstream AI contribution carries a named human attestation and a source link. The discovery handoff attests that proposed structures were reviewed by a named chemist; the translational handoff attests that biomarker clusters were verified against the underlying data by a named scientist; the clinical handoff attests that the protocol rationale reconciles to the verified clusters. The Level 5 leader is not designing the prompts; the leader is designing the accountability architecture that makes a cross-functional pilot defensible, ensuring that ALCOA+ attributability survives every handoff and that no claim arrives at the regulatory boundary without a traceable human owner upstream.

Kill Criteria: The Discipline to Stop

A pilot without a kill criterion does not end; it metastasizes, and the thirty-one-pilot portfolio is what an organization that cannot stop looks like. Kill criteria are the conditions, defined in the charter before the pilot starts, under which the organization will retire the pilot regardless of how attached anyone has become to it, and writing them in advance is the only way to make them real, because after a team has invested six months and built emotional and political capital, no one will volunteer to kill their own project. A well-constructed kill criterion is specific and measurable: "retire if the AI's source-linked citation accuracy, measured against named-writer reconciliation across fifty drafted sections, falls below ninety-eight percent," or "retire if the central-monitoring signal pilot produces a false-positive QTL-excursion rate that consumes more CRA hours than it saves, measured over one full monitoring cycle." The criterion converts an emotional decision into an evidentiary one, and the leader's role is to hold the organization to the evidence when the team wants to argue from sunk cost.

There is a second, subtler class of kill criteria that the regulated context demands: criteria that retire a pilot not because it failed to perform but because it cannot be made defensible. A generative pilot that performs beautifully but whose outputs cannot be reconciled to source, whose run metadata cannot be captured, or whose vendor cannot meet the zero-data-retention and audit-trail requirements of a 21 CFR Part 11 workflow, must be killed even if its accuracy is dazzling, because performance without defensibility is a Form 483 observation waiting to happen. This is the hardest kill for an organization to accept, because the pilot works, and stopping something that works feels like leaving value on the table. The Level 5 leader's job is to insist that in a regulated enterprise, a capability that cannot be governed is not value, it is latent liability, and that retiring it is the responsible act. The kill criteria, taken together, are what allow an organization to run many pilots cheaply, because a portfolio you can prune is a portfolio you can afford to grow, and an organization that cannot kill a pilot will eventually stop starting them.

Scale Criteria: The Evidence Required to Grow

If kill criteria define when to stop, scale criteria define what a pilot must prove before it earns the right to expand beyond its sandbox, and they are the mirror image of the kill criteria, written with equal specificity in the same charter. A pilot does not scale because executives are excited about it; it scales because it has generated the evidence its risk category requires. For a regulated-category pilot, the scale evidence includes a documented validation against its fit-for-purpose statement, a demonstrated audit trail that captures the prompt, system prompt, model and version, temperature, sources, and human-verified artifact for every output, a claim-reconciliation record that survives reviewer-grade scrutiny, and a defined human-accountability gate at every point where AI output enters a GxP artifact. For a high-risk-category pilot, the evidence bar rises to include alignment with EU AI Act high-risk-system obligations where applicable, a model card and data card documenting the system's intended use and data lineage, and an impact assessment that names the worst-case failure and the controls that bound it.

The scale criteria do something organizationally important beyond gating any single pilot: they create a common currency of evidence across functions, so that a discovery pilot and a pharmacovigilance pilot are held to the same standard of defensibility even though they touch completely different artifacts. This common currency is what makes a portfolio governable, because the steering committee can compare pilots not by how impressive their demos were but by how much of their scale evidence they have generated, which is the only comparison that predicts whether a pilot will survive contact with an inspector. The leader who establishes this common evidentiary currency has done something more durable than approving any individual pilot: they have built the rails on which the entire enterprise AI transformation will run, and the next lesson, on scaling without triggering an inspection finding, is entirely about traveling those rails from a single successful pilot to enterprise production without losing the defensibility the scale criteria established.

The Portfolio View and the Board Conversation

A Level 5 leader does not manage pilots one at a time; they manage a portfolio, and the portfolio view is the artifact that turns thirty-one ungoverned experiments into a governable strategy. The portfolio is a simple but powerful matrix: each pilot plotted by risk category against scale-evidence maturity, with its kill criteria, its named owner, its decision date, and its current evidentiary status visible at a glance. The pilots in the high-evidence, defensible quadrant are candidates to scale; the pilots that have stalled against their kill criteria are candidates to retire; the pilots that are load-bearing but undocumented, the biomarker-clustering quadrant, are the urgent remediation list, because they represent production risk the organization has not acknowledged. This single view is what the leader brings to the audit committee, and it transforms the board conversation from "how much are we spending on AI" to "here is our governed AI portfolio, here is the evidence behind what we are scaling, and here is what we are deliberately retiring."

The portfolio view also reframes the ROI question, which is the question the board will always ask, in a way that survives regulatory scrutiny. The naive ROI number, "AI saved us thirty million dollars in medical-writing hours," is the number that consulting firms publish and sponsors rarely reproduce, because it counts the hours AI removed without counting the verification hours AI added or the latent liability of ungoverned pilots. The governed-portfolio ROI is a more honest and more defensible number: the value of the pilots that passed their scale criteria, net of the verification cost their risk category requires, with the retired pilots counted not as failures but as cheaply avoided liabilities. A leader who can present ROI this way is doing something a panel at the DIA Annual Meeting cannot teach, which is connecting the daily discipline of pilot governance to the board-level story of value creation, and it is precisely this connection that defines the Level 5 voice: the ability to hold the inspector's standard and the investor's question in the same sentence and answer both without contradiction.

What This Means for the Leader on Monday

The transformation does not begin with a new model or a new vendor; it begins with an inventory and a charter template. The leader's first Monday action is to count the ungoverned pilots, because you cannot govern what you have not named, and the count is almost always larger and more load-bearing than the organization believes. The second action is to issue a one-page charter requirement, so that no new pilot starts without an intended use, a risk category, named kill criteria, and named scale criteria, and so that the charter requirement applies retroactively to the pilots already running, including the ones that have quietly become production. The third action is to stand up the portfolio matrix and bring it to the next steering review, where the hardest and most valuable conversation will be about the load-bearing undocumented pilots that must either be validated to a scale standard or retired, because they cannot be allowed to remain in the quadrant where the organization depends on them and cannot defend them.

None of this slows innovation; it is what makes innovation affordable, because an organization that can charter a pilot in a day, prune it cheaply when it stalls, and scale it confidently when it proves itself will run more experiments and ship more durable capability than an organization paralyzed by the fear that any pilot might become an undocumented liability. The cross-functional pilots that span discovery to regulatory are the highest-value experiments in the enterprise precisely because they thread real evidence across the whole lifecycle, and they are governable only if the handoff gates preserve accountability at every functional boundary. The leader who builds this discipline is not the person who approves the most exciting demo; they are the person who can stand in front of a regulator and a board on the same day and show that every AI capability in production was chartered, validated, and reconciled by a named human, and that every capability that could not meet that standard was deliberately retired. That posture is the foundation on which the rest of Level 5 is built.

Key Takeaways

  • The pilot is the unit of enterprise AI risk, not a rules-free sandbox. Under the FDA-EMA Guiding Principles, AI that touches a GxP artifact is in scope from first use, so a pilot without governance is an undocumented production system, and the count of ungoverned pilots is a board-level risk metric like open Form 483 observations.
  • A one-page charter scoping intended use, risk category, kill criteria, and scale criteria is the first control. Scope discipline, the discipline to say what the pilot will not do, is what keeps a discovery or Module 2.5 pilot from quietly becoming the decision-maker, and the internal-use, regulated, and high-risk category determines the evidence the pilot must generate.
  • Cross-functional pilots that span discovery to regulatory create the most value and lose accountability at the handoffs. The named handoff gate, elevated from the Level 3 human-AI handoff pattern, ensures every upstream AI contribution carries a named human attestation and a source link, so ALCOA+ attributability survives every functional boundary.
  • Kill criteria must be written before the pilot starts and must include defensibility, not just performance. A pilot that performs beautifully but cannot be reconciled to source, cannot capture run metadata, or cannot meet Part 11 requirements must be retired even when its accuracy is dazzling, because performance without defensibility is a Form 483 waiting to happen.
  • Scale criteria create a common currency of evidence that makes the portfolio governable. The portfolio matrix, plotting each pilot by risk category against scale-evidence maturity, turns ungoverned experiments into a governable strategy and reframes the board's ROI question as the net value of governed, validated capability minus deliberately retired liability.