21 CFR Part 11, EU Annex 11, ALCOA+, and GxP, Applied to AI Tools
Picture an FDA investigator sitting across the table from a regulatory operations lead during a Pre-Approval Inspection, pointing at a sentence in a Module 2.5 Clinical Overview and asking a deceptively simple question: who created this electronic record, when, against which sources, and how do I know it has not been altered since. In a fully manual world the answer lives in a validated document-management system with a tamper-evident audit trail and an electronic signature bound to a named human. Now add the detail that an enterprise large language model drafted the first version of that sentence, and the investigator's question does not get easier, it gets sharper, because 21 CFR Part 11 and EU Annex 11 do not care that a model produced the words. They care that the record is attributable, traceable, secured, and signed, exactly as they always have. This lesson takes the two rules that govern electronic records in regulated life sciences, lays them next to the reality of LLM-based tools, and shows you precisely why a consumer chatbot as shipped cannot meet the standard, what makes an AI-assisted workflow defensible instead, and how GAMP 5 software categories tell you how much validation a given tool actually demands. It is the rule-by-rule companion to the ALCOA+ lesson before it: ALCOA+ defines what a trustworthy record is, and Part 11 and Annex 11 are the regulations that make those properties enforceable.
What 21 CFR Part 11 and EU Annex 11 Actually Require
21 CFR Part 11 is the United States Food and Drug Administration regulation that governs electronic records and electronic signatures in FDA-regulated activities. In plain terms, it sets the conditions under which an electronic record can be trusted to the same degree as a paper record with a wet-ink signature. Its core demands are a secure, computer-generated, time-stamped audit trail that records who did what and when and does not let users overwrite their own history; access controls that limit system use to authorized individuals; electronic signatures that are uniquely bound to one person and cannot be reused or transferred; and validation of the system to ensure it does what it is supposed to do, consistently and reliably. Part 11 is the reason a regulated electronic document carries a who, a when, and a tamper-evident record of every change, and it is the standard against which any tool that touches a regulated record is measured.
EU Annex 11 is the European Union's parallel instrument, the annex to the EU Good Manufacturing Practice guidelines that governs computerised systems used in GMP-regulated activities. It covers much of the same ground from a slightly different angle: risk management across the system lifecycle, validation proportionate to risk, supplier and service-provider assessment, data integrity, audit trails, and the controls around electronic signatures. Annex 11 is more explicit than Part 11 about risk-based thinking and about the obligations that flow to the regulated company when a system is supplied or operated by a third party, which is precisely the situation when you use a vendor's AI tool. Together, Part 11 on the US side and Annex 11 on the EU side form the regulatory floor for any computerised system, AI or otherwise, that creates, modifies, maintains, or transmits records supporting a regulated decision.
The critical move for this program is to recognize that an LLM-based drafting tool is a computerised system that creates and modifies regulated records, which places it squarely inside the scope of both rules. The model's fluency does not exempt it; if its output ends up in a submission, a batch record, a case file, or a trial document, the tool that produced it is subject to the same audit-trail, access-control, signature, and validation expectations as any other system in that chain. This is the single most important reframing of the lesson: the question is never whether Part 11 and Annex 11 apply to your AI, but whether your AI deployment is built to satisfy them.
Why a Consumer Chatbot As Shipped Is Not Part 11 Compliant
Take the most common shortcut in the field, a writer opening a public consumer chatbot in a personal browser tab and pasting in a draft, and run it against the four pillars of Part 11. It fails all four, and understanding why is more instructive than memorizing that it does. On the audit trail, a consumer tool keeps a conversation history for the convenience of the user, not a secure, tamper-evident, computer-generated audit trail in the regulatory sense; the user can delete the thread, the history is not bound to the regulated record it helped produce, and there is no immutable log an inspector can retrieve and trust. On access control, the account is personal, often outside the company's identity and access management, so there is no enforceable link between the system and the authorized, trained, named individuals the regulation expects.
On electronic signatures, the consumer tool offers nothing that qualifies as a Part 11 electronic signature, no unique binding of a signing act to a named person under controls that prevent reuse or repudiation, because it was never built to be a system of record. On validation, the regulated company has performed none, has no documented intended use, no acceptance criteria, no qualification evidence, and crucially no control over when the vendor changes the underlying model, which means the system's behavior can shift without notice and without any change-control record. A tool that can change underneath you with no notification cannot be in a validated state, by definition. This is the precise, technical meaning of the claim that a consumer chatbot as shipped is not Part 11 compliant: it is not that the output is necessarily wrong, it is that the system around the output satisfies none of the controls that make an electronic record trustworthy, and no amount of careful prompting changes that.
The corollary matters as much as the claim. The same underlying model, deployed inside a validated enterprise environment with single sign-on, role-based access, immutable logging, a documented intended use, qualification evidence, and a managed change-control process for model updates, can sit inside a Part 11 compliant workflow. The compliance lives in the deployment and the surrounding controls, not in the raw model. This is why the answer to AI in regulated work is almost never "ban it" and almost never "paste freely," but "use the governed deployment," and it is why vendor selection and validation, treated in depth in the applied levels, are not bureaucratic overhead but the mechanism by which AI use becomes lawful at all.
The Audit Trail for an AI-Assisted Record
Part 11 and Annex 11 both turn on the audit trail, and AI-assisted work places a specific and unusual demand on it. A conventional audit trail records the human actions on a record: created, edited, reviewed, approved, each with a user and a timestamp. An AI-assisted record needs all of that and one more layer, the provenance of the AI contribution itself, because a future inspector or reviewer may ask not only who approved the section but how it was generated. The defensible audit trail therefore captures the run that produced the draft: the prompt, the system prompt or its versioned identity, the model and version, the temperature or generation settings, the timestamp, the exact sources loaded into the context window, and the human verification that followed, all linked to the specific version of the document under change control.
This is where the Contemporaneous attribute from the ALCOA+ lesson becomes a Part 11 obligation rather than a best practice. Because model output varies run to run, the provenance must be captured at the moment of generation; reconstructed three weeks later it is a guess, not a record, and a guess is exactly what Part 11's audit-trail requirement exists to prevent. A well-built enterprise AI deployment automates this capture so the writer does not have to assemble it by hand, which is both more reliable and more defensible than a manual log, and it stores that provenance in the same controlled, retained, retrievable system as the document itself, so that the AI involvement endures and is available for the full retention period the way Annex 11 and Part 11 expect of any regulated record.
The failure mode here is subtle and common. A team adopts an enterprise AI tool, believes that because it is enterprise it is automatically compliant, and never confirms that the tool actually captures and retains the generation provenance in an inspectable form. Enterprise branding is not the same as a Part 11 audit trail. The diligent question, asked at adoption and confirmed at validation, is whether the tool logs the generation event, binds that log to the output, secures it against alteration, and retains it for the required period, because if it does not, the workflow has an audit-trail gap precisely at the point where the AI touched the record, and that gap is what an inspector at a Pre-Approval Inspection is increasingly trained to look for.
Electronic Signatures and the Named Author Who Cannot Delegate to a Model
An electronic signature under Part 11 is not a typed name or a scanned image; it is a controlled act that binds a specific, authenticated person to a specific record, with controls ensuring it cannot be reused, transferred, or repudiated, and with the meaning of the signature, authorship, review, or approval, made explicit. Annex 11 carries the equivalent expectation for signatures on computerised records in the EU GMP context. The signature is the mechanism by which accountability, the load-bearing principle from the FDA-EMA lesson, becomes a concrete, auditable event: a named human attesting that this record is theirs and they stand behind it.
AI changes nothing about who can hold a signature and everything about what the signature now attests. A model cannot sign a Part 11 record, because it is not a person and accountability cannot transfer to it; only the named human author or approver signs. But when that human signs an AI-assisted section, they are attesting to a record whose first draft they did not write, which means the signature is only as trustworthy as the verification behind it. Signing an unverified AI draft is not a clerical act, it is a false attestation, because the signer is asserting ownership of claims they have not reconciled to source. The entire weight of the previous lessons, load the sources, verify each claim by type, capture the run, lands here: the verification is what makes the signature honest, and the signature is what makes the record compliant.
This reframes a temptation that AI introduces. Because an AI draft looks finished, there is pressure to treat the human step as a quick formality and sign quickly. Part 11 does not recognize a fast formality; it recognizes an attestation, and an attestation carries the same legal and regulatory weight whether the words underneath it came from a senior writer or a model. The signer owns the gap between how finished the draft looks and how verified it actually is, and the signature is the point at which that ownership becomes a recorded, inspectable, and personally attributable fact.
Validation and the Moving Model Problem
Validation is the documented evidence that a system does what it is intended to do, reliably and reproducibly, and it is the Part 11 and Annex 11 requirement that AI strains most, because the classic validation model assumes the software is stable between validated states. You qualify the system, you place it under change control, and you re-qualify when it changes. A traditional rules-based application changes only when someone deliberately ships a new version, so this model works cleanly. An LLM-based tool breaks the assumption in two ways: the model can be updated by the vendor, sometimes silently, changing its behavior without a deliberate act by the regulated company, and even at a fixed version the output is probabilistic rather than deterministic, so the same input can yield different output across runs.
The response is not to abandon validation but to adapt it, and the adaptation has a name and a shape that later levels build out in full. First, validation of an AI tool is intended-use and fitness-for-purpose driven: you do not validate "the model" in the abstract, you qualify it for a specific task with specific acceptance criteria, the way fitness for purpose in the FDA-EMA principles demands. Second, because the output is probabilistic, performance is evaluated statistically against those acceptance criteria rather than by expecting a single correct output, which is the role of evals. Third, because the model can move, change control must include awareness and management of vendor model updates, with re-qualification triggered when a material change occurs, and ongoing performance monitoring to detect drift between formal re-qualifications. This is the operational reason the FDA-EMA principles pair model performance monitoring with ongoing lifecycle monitoring: a one-time qualification of a moving, probabilistic system is not validation, it is a snapshot, and the snapshot expires.
There is a 2026 development that sharpens this. The FDA finalized its Computer Software Assurance guidance, Computer Software Assurance for Production and Quality System Software, on 24 September 2025, superseding Section 6 of the 2002 general principles of software validation guidance and promoting a risk-based, critical-thinking approach to assurance over exhaustive scripted testing of low-risk functions. CSA reshapes how production and quality-system software is assured, and it is a forcing function for re-baselining vendor validation across the industry. One precise caveat the diligent professional must hold: CSA explicitly does not apply to software in or as a medical device, so a tool that is itself a device function is governed by the device pathway and the PCCP framework of the next lesson, not by CSA. For the LLM-based drafting and operations tools at the center of this program, which are not medical devices, CSA's risk-based assurance posture is directly relevant and broadly welcomed.
GAMP 5 Categories and Where LLM-Based Tools Sit
The practical question every quality and validation lead asks is how much validation a given tool actually demands, and the standard framework for answering it is GAMP 5, the widely adopted Good Automated Manufacturing Practice guide whose second edition includes an AI-specific appendix. GAMP 5 sorts software into categories that scale validation rigor to how configured or custom the software is, and the relevant ones run from Category 3 to Category 5. Category 3 is non-configured, off-the-shelf software used as supplied; Category 4 is configured software, where the company sets up the product within its intended design to fit its process; and Category 5 is custom or bespoke software built specifically for the company, which carries the highest validation burden because there is no vendor track record and every behavior is unique.
Mapping LLM-based tools onto these categories is where the judgment lives, and it is genuinely a judgment, not a lookup. A general-purpose model used essentially as supplied through a vendor interface leans toward Category 3, but the moment the company adds a configured system prompt, a retrieval layer over its own documents, role-based access, and intended-use constraints, it has moved toward Category 4, because it has configured the product within its design to fit a regulated process. A bespoke model fine-tuned on the company's data or an in-house system built on a foundation model leans toward Category 5, with the heaviest validation expectations and the obligation to qualify behavior the company itself created. The GAMP 5 Second Edition AI appendix exists precisely because the probabilistic, evolving nature of these tools complicates the classic categories, and it guides how to apply risk-based, lifecycle validation to them. For the individual user, the takeaway is not to perform the categorization, that is the quality function's job, but to understand that the tool you use sits in a category that determined how heavily it was validated, and that your verification discipline operates on top of, not instead of, that validation.
The Monday Demands of Part 11 and Annex 11 on AI Work
All of this resolves into a short set of demands a professional can carry into any AI-assisted regulated task. Use the governed deployment, never the personal consumer account, because the deployment is where the audit trail, access control, signature capability, and validation live, and the personal account satisfies none of them. Confirm, do not assume, that the tool captures and retains generation provenance in an inspectable, retained, retrievable form, because enterprise branding is not the same as a Part 11 audit trail. Treat your signature as the attestation it legally is, and never apply it to an AI draft you have not reconciled to source, because the signature carries the same weight regardless of who, or what, produced the words.
Understand that the tool you use was validated for a specific intended use under a GAMP 5 category, and that using it outside that intended use, for a task it was never qualified for, breaks the validation that makes it defensible. And hold the moving-model awareness that the regulation now demands: a tool that performed well at adoption can change after a vendor update, so a stable workflow includes someone watching for that change rather than assuming permanence. None of these demands slows good work; they are simply the shape that good AI-assisted work takes when the record it produces has to survive a Pre-Approval Inspection, an EMA GMP inspection, or a Form 483 response. The previous lesson taught you what a trustworthy record is; this one tells you which regulations enforce it and how a defensible AI deployment is built to satisfy them, and the next turns to the framework that governs AI which keeps learning after you deploy it.
Key Takeaways
- An LLM-based drafting or operations tool is a computerised system that creates and modifies regulated records, so 21 CFR Part 11 and EU Annex 11 apply to it in full. The question is never whether the rules apply to your AI, but whether your AI deployment is built to satisfy the audit-trail, access-control, electronic-signature, and validation requirements they impose.
- A consumer chatbot as shipped fails all four Part 11 pillars: no secure tamper-evident audit trail, no enforceable access control, no compliant electronic signature, and no validation, made worse by a model that can change under you with no change-control record. The same model inside a governed enterprise deployment can be compliant, because compliance lives in the deployment and surrounding controls, not in the raw model.
- The audit trail for an AI-assisted record must capture the generation provenance, not just the human actions: prompt, system-prompt version, model and version, settings, timestamp, sources loaded, and verification, captured contemporaneously and retained in the controlled system, because reconstructed-after-the-fact provenance is a guess Part 11 exists to prevent.
- A model cannot hold a Part 11 electronic signature; only the named human can, and the signature is an attestation, not a formality. Signing an unverified AI draft is a false attestation, so the verification of every claim is what makes the signature honest and the record compliant.
- Validation must adapt to a moving, probabilistic system using intended-use qualification, statistical performance evaluation (evals), and change control over vendor model updates with drift monitoring. GAMP 5 categories 3 to 5 scale that validation rigor, the FDA's CSA final guidance (24 September 2025) promotes risk-based assurance but explicitly excludes software in or as a medical device, and your verification operates on top of the validation, never instead of it.
Skill.re