Assessing AI Readiness in a Regulatory or Medical Writing Function
The Chief Regulatory Officer has thirty minutes on the quarterly governance call, and the question on the slide is blunt: "Are we ready to put AI into the submission machine, yes or no?" The Head of Regulatory Writing who answers "yes" because Certara CoAuthor produced an impressive Module 2.5 demo last Tuesday has misunderstood the question. The CRO is not asking whether a model can draft a paragraph. The CRO is asking whether the function, the people and the process and the data and the validation posture and the vendor footprint, can deploy that model in a way that survives a 21 CFR Part 11 audit-trail review, an EMA GCP inspection, and a Day 74 Office of New Drugs Information Request without the sponsor's signature becoming indefensible. Readiness is not a model capability. It is a function-level property, and it is measurable. This lesson gives you the five-dimension capability audit that turns "are we ready" from a gut feeling into a scored, defensible, board-presentable assessment, anchored to the FDA-EMA Guiding Principles' ninth principle on governance and documentation, and built so that the score itself is the artifact you hand to the CRO.
Why Readiness Is a Function Property, Not a Tool Feature
The single most common strategic error at Level 4 is to confuse tool readiness with function readiness. A vendor demo shows you that the tool can do the task under ideal conditions, with curated inputs, on a stage, by an engineer who has run the demo two hundred times. It tells you nothing about whether your function can operate that tool, in your environment, on your worst Tuesday, under audit. The gap between "the tool works" and "we are ready to deploy the tool" is precisely the gap that this audit measures, and it is wide. A function with a flawless tool and no validation posture is not ready. A function with a brilliant power-user and no change-control process is not ready. Readiness is the joint property of five dimensions, and the function is only as ready as its weakest one, because an inspector, an auditor, and a reviewer all read to the weakest link.
This is why you assess the function, not the model. The FDA-EMA Guiding Principles of Good AI Practice, released jointly by CDER and the EMA on 14 January 2026, make this explicit in their ninth principle on governance and documentation: the obligation to maintain governance structures and documentation sits with the sponsor organization, not with the AI system. A model cannot hold a governance structure. A function can. So the question the CRO is really asking, and the question this audit answers, is whether the function has built the organizational scaffolding that makes a probabilistic tool produce a deterministic, attributable, inspectable record. That scaffolding is the deliverable. The model is a component inside it.
Dimension One: Process Maturity
The first dimension asks whether your function's processes are documented, repeatable, and measurable enough that you could insert an AI step and know exactly where it sits, what it consumes, what it produces, and who verifies its output. A function whose CSR-to-Module-2.5 workflow lives in the heads of three senior writers cannot insert AI cleanly, because there is no documented process map to insert it into; the AI step would float in an undefined space with no defined verification gate, which is exactly the configuration an inspector flags first. A function with a versioned, swim-laned process map of its 14-week submission cycle, with named handoffs and named quality gates, can drop an AI-ready step into a specific lane and define its verification gate against an existing control. The maturity question is therefore not "do you write good documents" but "is your process legible enough to be safely augmented."
Score this dimension on a five-level scale that any CRO will recognize from CMMI-style maturity models: Level 1 is ad hoc and undocumented; Level 2 is documented but inconsistently followed; Level 3 is documented, followed, and measured; Level 4 is managed with quantitative process control; Level 5 is optimizing with continuous improvement. Most regulatory writing functions in 2026 sit at Level 2 or low Level 3, and that is a finding, not a failure, but it bounds what you can responsibly deploy. A Level 2 function should deploy AI only in drafting-assist mode with heavy human verification, because it lacks the process control to defend anything more autonomous. A Level 4 function can consider integrated lifecycle workflows. The honest scoring of process maturity is the single best predictor of how fast the roadmap in the next lesson can move.
Dimension Two: Data Infrastructure
The second dimension asks whether the function's source documents, the CSRs, the TLF packages, the Investigator's Brochures, the prior submissions, the global safety database, exist in a form an AI workflow can actually retrieve, ground against, and cite. This is where most readiness assessments are too optimistic. A retrieval-augmented workflow that drafts a Module 2.5.4 efficacy section is only as trustworthy as its access to the final TLF package, and if the TLF package lives as a flat PDF on a shared drive with no machine-readable table structure, the workflow cannot reliably ground a cross-reference against it. The data infrastructure question is concrete: are your source-of-truth artifacts stored in a system of record (Veeva Vault RIM, Vault QualityDocs) with version control, are they structured or at least consistently formatted, and is there a clean boundary between validated source content and working drafts?
Assess three sub-properties. First, accessibility: can an approved AI workflow reach the source content through a governed, logged channel, or would it require pasting confidential content into an ungoverned tool, which is itself a finding under the confidentiality principle. Second, quality and structure: is the content clean enough that grounding works, or is it riddled with scanned tables, inconsistent numbering, and orphaned cross-references that will degrade retrieval. Third, lineage: can you trace any given source artifact to its version, its approval, and its place in the document hierarchy, because an AI output grounded against an outdated CSR is worse than no AI at all. A function with mature data infrastructure has already done most of the hard work of AI readiness, because grounding is where defensibility lives, and grounding is a data problem before it is a model problem.
Dimension Three: GxP Validation Posture
The third dimension is the one that separates pharma from every other industry, and it is the one tool vendors most want you to skip. It asks whether your function and your organization have the validation machinery to qualify an AI tool as fit for its intended use, under 21 CFR Part 11, EU Annex 11, and the GAMP 5 framework, including its Second Edition AI-specific appendix. The question is not whether you can validate software in general; most pharma functions can. The question is whether you have a validation approach for a tool whose output is non-deterministic, whose behavior depends on a system prompt you may not control, and whose vendor ships updates on a release cadence you do not set. The Veeva Vault RIM AI Agents reaching general availability on the August 2026 Vault release, and the Certara CoAuthor and Veeva Vault RIM integration moving into production rollouts across sponsors through 2026, mean this is not hypothetical; these are tools your function will be asked to qualify on a real timeline.
Score the validation posture on whether the function has, or can readily access, four capabilities. First, an intended-use and fitness-for-purpose discipline that can write a defensible statement of what the tool is and is not qualified to do, aligned to the FDA-EMA "fitness for purpose" principle. Second, a Computer Software Assurance approach, following FDA's CSA guidance finalized 24 September 2025, that can scale validation rigor to risk rather than validating everything to the same heavy standard. Third, an ongoing performance monitoring capability that treats a learning or updated AI tool the way the FDA's January 2025 PCCP guidance treats adaptive SaMD, with defined acceptance criteria and re-validation triggers. Fourth, an audit-trail architecture that captures the specific run, prompt, system prompt, model and version, sources loaded, and human verification, so that "the AI helped" becomes a Part 11 record. A function strong on tooling but weak on validation posture is the most dangerous configuration of all, because it will deploy confidently and fail at inspection.
Dimension Four: Talent and Literacy
The fourth dimension asks whether the people in the function have the literacy to use AI well and the judgment to catch it when it is wrong. This is not a question of whether your writers are smart; they are. It is a question of whether they understand the mechanics well enough to know that a fluent draft is the start of the work, not the end, and whether they have the domain depth to catch a plausible-but-wrong hazard ratio or a fabricated TLF cross-reference that survives spell-check. A function staffed with brilliant senior writers who treat AI output as either magic or garbage is not ready; both postures fail. The ready function has writers who treat AI output as a competent but unreliable junior colleague whose every factual claim must be reconciled to source, and who can do that reconciliation efficiently because they know the source.
Assess talent across three layers. The power-user layer: do you have the first three to six people who can build, test, and refine AI workflows, the champions that a later chapter addresses, without whom no deployment scales. The practitioner layer: can the broad body of the function execute AI-assisted drafting with documented verification, the Level 2 competency this program certifies. The leadership layer: do the RA managers, PV leads, and medical-writing directors understand AI well enough to govern it, set verification expectations, and defend the function's choices to the CRO, because a leadership team that cannot articulate why a workflow is defensible cannot defend it under pressure. A talent gap is the most fixable of the five dimensions through training, which is why honest scoring here directly feeds the change-management and internal-certification investments in your roadmap.
Dimension Five: Vendor Footprint
The fifth dimension asks what AI you already have, whether you know it, and whether it is governed. Almost every function in 2026 has a larger AI footprint than its leadership believes, because AI capabilities arrive embedded inside tools the function already owns. Your Veeva Vault may have AI agents available. Your Certara writing tools may have GenAI capabilities. Your PV platform, whether ArisGlobal LifeSphere NavaX or Oracle Argus AI, may already be doing case-narrative drafting. Your clinical operations stack, Medidata Acorn or Saama, may be generating monitoring signals. The vendor-footprint audit is a discovery exercise: enumerate every tool in the function, identify which have AI capabilities active or available, and determine for each whether it has been validated, whether its data handling meets the confidentiality floor of a Business Associate Agreement with zero data retention, and whether its outputs are entering regulated artifacts through a governed path or a shadow one.
The most dangerous finding in a vendor-footprint audit is shadow AI: capabilities switched on or accessible inside owned tools, used by individuals, with no validation and no audit trail, entering regulated content. A senior writer quietly using an embedded summarization feature to compress a CSR section, with no documentation, has created an undocumented AI dependency inside a submission, and that is a finding waiting for an inspector. The vendor-footprint dimension also assesses concentration and exit risk: a function entirely dependent on a single vendor's AI roadmap has a strategic exposure that the CRO will want quantified, because vendor lock-in in a regulated function is not just a commercial risk, it is a continuity-of-compliance risk. Scoring this dimension produces both a readiness input and an immediate remediation list, because shadow AI must be governed before any new deployment is responsible.
The Cross-Dimension Interactions That Most Audits Miss
The five dimensions are not independent, and the readiness audit that treats them as five separate scores misses the interactions that actually determine deployment risk. The most important interaction is between data infrastructure and validation posture: a function with unstructured source content cannot build a credible fitness-for-purpose argument for a grounded workflow, because the workflow's grounding cannot be qualified against sources that are not machine-readable, so a weak data dimension silently caps the validation dimension regardless of how mature the validation machinery is on paper. A second interaction runs between talent and vendor footprint: a function with strong power-users will discover its shadow AI quickly, because the people who understand AI notice the embedded features their colleagues are using, while a function with low literacy will not even know what to audit. A third runs between process maturity and every other dimension, because an illegible process gives the AI step nowhere defensible to live no matter how strong the tool, the data, or the validation approach.
These interactions matter because they change the remediation sequence. A naive audit that scores data at 2 and validation at 2 might recommend investing equally in both, but the interaction tells you that the data investment must come first, because validation cannot rise until the data it grounds against is structured. Sequencing remediation against the dependency graph, rather than against the raw scores, is what separates a readiness audit that accelerates the function from one that funds parallel investments that block each other. The audit should therefore present not just five scores but a short dependency analysis: which weak dimensions are gating which others, and therefore which investment unlocks the most downstream movement. This is the analysis that turns the readiness profile from a snapshot into a sequencing plan, and it is the bridge to the roadmap, because the roadmap's pace is set precisely by the order in which these gating dependencies can be cleared.
Assembling the Readiness Score and the CRO Conversation
The five dimensions combine into a readiness profile, not a single number, and the profile is the artifact. Score each dimension one to five, but resist the temptation to average them, because averaging hides the weakest-link property that actually governs deployment risk. A function scoring 4-4-2-4-3 is not a "3.4 function"; it is a function gated by its validation posture at 2, and the honest readout to the CRO is that the function can deploy AI in drafting-assist mode now but cannot responsibly move to integrated workflows until the validation gap closes. This weakest-link framing is what makes the assessment defensible rather than promotional, and it is what distinguishes a Level 4 strategist from a vendor's account executive. The CRO has heard the promotional version; what earns trust is the version that names the constraint.
Present the readiness profile to the CRO as a decision instrument, not a report card. Pair each dimension's score with the specific deployment it permits and the specific investment that would raise it, so the conversation moves immediately from "where are we" to "what do we fund to move." Tie the whole assessment explicitly to the FDA-EMA governance-and-documentation principle, so the CRO understands that the readiness profile is itself the kind of governance artifact the principle expects the sponsor to maintain. And be honest about the asymmetry of error: a function that over-states its readiness and deploys an integrated workflow it cannot validate is exposed at the next inspection, while a function that under-states its readiness merely moves slightly slower. In a regulated environment, the cost of false readiness vastly exceeds the cost of conservative readiness, and saying that out loud is what makes the CRO trust the rest of your strategy. The readiness audit is the foundation the roadmap is built on, and a roadmap built on an honest audit is the only kind that survives contact with an inspector.
Key Takeaways
- Readiness is a function-level property, not a tool feature, and the function is only as ready as its weakest dimension. A vendor demo proves the tool works under ideal conditions; it says nothing about whether your people, process, data, validation posture, and vendor governance can deploy it under audit. Assess the function, not the model, because the FDA-EMA governance-and-documentation principle places the obligation on the sponsor organization.
- The five dimensions are process maturity, data infrastructure, GxP validation posture, talent and literacy, and vendor footprint. Score each one to five on a maturity scale and do not average them; a 4-4-2-4-3 function is gated at the 2, and the honest readout names that constraint rather than hiding it behind a mean.
- Data infrastructure is where defensibility lives, because grounding is a data problem before it is a model problem. If the final TLF package is an unstructured PDF with no machine-readable tables, no retrieval workflow can reliably ground a cross-reference against it, no matter how good the model is.
- The vendor-footprint audit almost always uncovers shadow AI: embedded capabilities used without validation or audit trail, entering regulated content. Shadow AI must be governed before any new deployment is responsible, and vendor concentration is a continuity-of-compliance risk the CRO will want quantified, not just a commercial one.
- The readiness profile is the artifact you hand the CRO, framed as a decision instrument, not a report card. Pair each score with the deployment it permits and the investment that would raise it, and state the asymmetry of error openly: in a regulated function, the cost of false readiness vastly exceeds the cost of conservative readiness.
Skill.re