Designing a Part 11 + Annex 11 Validation Protocol for an AI Tool
A Director of Regulatory Operations has been told by the Chief Quality Officer that the AI tool the toxicology writing group has been quietly using for nine months to draft IND-enabling single-dose and repeat-dose toxicology study reports is going to be inspected, whether by an FDA investigator at a Pre-Approval Inspection or by the company's own internal audit, and that there is currently no validation package behind it. The tool drafts the tabulated summaries, the narrative descriptions of clinical signs, the histopathology incidence text, and the no-observed-adverse-effect-level discussion that feed Module 4.2.3 and ultimately the Module 2.6.6 toxicology written summary. It works well. The writers trust it. And there is nothing on paper that says why anyone should. This lesson builds the validation protocol that closes that gap, anchored to 21 CFR Part 11 and EU Annex 11, written in the language of intended use, fitness for purpose, acceptance criteria, performance qualification, and ongoing performance monitoring. It is the artifact a strategist signs, hands to QA, and defends to a regulator. The whole point of Level 4 is that you can produce it, and produce it so it holds.
Why the Validation Protocol Is the Strategic Artifact, Not the Paperwork
A validation protocol is often treated as the document QA forces on a function after the fact, a compliance tax paid in tables and signatures. At Level 4 you have to invert that posture completely, because the protocol is the single document where strategy becomes defensible. It is the place where the loose claim "our AI tool is reliable" is decomposed into specific, testable assertions about what the tool is allowed to do, what counts as acceptable performance, and what evidence proves it. Everything upstream of the protocol, the readiness assessment, the vendor scorecard, the roadmap, is preparation; everything downstream, the inspection defense, the ROI narrative, the governance reporting, leans on it. If the protocol is vague, every downstream defense is vague, and a vague defense in front of an FDA investigator is indistinguishable from no defense at all.
The reason the protocol carries this weight is that it is the bridge between two languages that do not naturally translate. On one side is the probabilistic, non-deterministic reality of a large language model that completes patterns and varies run to run. On the other is the deterministic, evidence-based world of 21 CFR Part 11 and EU Annex 11, which were written for computerized systems that behave the same way every time and whose correctness can be demonstrated by showing that a given input always produces a given output. The validation protocol is where you reconcile these, and the reconciliation is not cosmetic. It changes what you test, how you define passing, and what you commit to monitor forever. A strategist who cannot write this document cannot actually deploy AI in a regulated function; they can only hope it is never inspected.
There is a second strategic reason the protocol matters, which is scope control. Writing the intended-use statement forces the function to decide, explicitly and on the record, what the tool is and is not for. That decision is the most consequential one in the entire deployment, because every acceptance criterion, every test, and every monitoring metric flows from it. A tool scoped as "drafts the histopathology incidence narrative for human verification" is validated against a manageable set of claims. The same tool scoped as "produces the toxicology summary" implies a claim about scientific conclusions that no LLM can meet and no protocol can validate. The protocol is where ambition is disciplined into something testable, and a strategist who lets the scope inflate has written a document that cannot pass.
Anchoring the Protocol to the IND-Enabling Toxicology Report
This protocol is not abstract; it validates an AI tool used to draft IND-enabling toxicology study reports, and the specificity of that artifact is what makes the protocol writable. The tool sits inside a nonclinical writing workflow where a GLP toxicology study has been completed, the study director has signed the GLP-compliant study report, and the data, the individual animal data, the summary tables of body weights and organ weights, the clinical-pathology results, the histopathology findings, exist as the source of truth. The AI tool's job is to turn that validated source data into the prose and tabulated text that populate the nonclinical study report sections and feed Module 4.2.3 of the eCTD and, through it, the Module 2.6.6 toxicology written summary and the Module 2.6.7 tabulated summary. It is not generating data. It is generating language that describes data.
That distinction governs the entire protocol. Because the source data are validated GLP outputs under an established quality system, the protocol does not have to validate the science; it has to validate the fidelity of the translation from data to language. The risk is not that the study found the wrong NOAEL; the study director owns that. The risk is that the AI tool, drafting the narrative, transcribes a mean organ-weight change incorrectly, attaches a histopathology finding to the wrong dose group, describes a finding as "non-adverse" when the study director classified it as adverse, or invents a cross-reference to a summary table that does not carry that result. Each of these is a fidelity failure, and fidelity failures are exactly what a validation protocol can be built to detect, because each one is checkable against the validated source.
Anchoring to this artifact also tells you which regulations bite and how hard. The toxicology report is a GLP document, so 21 CFR Part 58 sits in the background, but the AI tool itself is a computerized system that creates electronic records destined for a submission, which puts it squarely under 21 CFR Part 11 and, for any European filing or EU-based operation, EU Annex 11. The records the tool produces, the drafts, the edits, the final verified text, must be attributable, contemporaneous, and traceable, and the audit trail must show who did what and when. The protocol therefore validates two intertwined things at once: that the tool produces faithful content, and that the system around the tool produces a defensible record of how that content came to be. Both are in scope, and a protocol that validates only the first fails Part 11 on the second.
The Intended-Use Statement: The Sentence Everything Hangs On
The intended-use statement is the foundational clause of the protocol, and it should be written with the care of a claim in a label, because functionally that is what it is. A strong intended-use statement for this tool reads something like: "The tool is used to generate draft narrative and tabulated summary text for IND-enabling GLP toxicology study reports, from validated study source data provided as structured input, for review and verification by a qualified nonclinical writer and approval by the study director, with no scientific conclusion accepted without human verification against source." Every phrase in that sentence is load-bearing. "Draft" means the output is never final. "From validated study source data provided as structured input" means the tool is grounded, not free-generating. "For review and verification by a qualified nonclinical writer" names the human accountability. "No scientific conclusion accepted without human verification against source" forecloses the over-reach that would make the tool unvalidatable.
The intended-use statement also has to state, explicitly, what the tool is not for, because the negative space is where inspectors probe and where scope creep does its damage. It is not for generating the NOAEL determination, classifying findings as adverse or non-adverse, or making any toxicological judgment; those are study-director decisions captured in the signed GLP report. It is not for use on source data that have not been verified. It is not for any document type outside the validated scope; a tool validated for toxicology reports is not thereby validated for clinical study reports, and the protocol says so. This matters because the FDA-EMA Guiding Principles' fitness-for-purpose principle is precisely a demand that the tool be characterized against a defined intended use, and a tool used outside its intended use is, by definition, operating unvalidated, which is the finding you most want to avoid.
One discipline separates a professional intended-use statement from an amateur one: it is written so that compliance with it is observable. "The tool assists writers" is not observable; you cannot tell from a record whether it was followed. "The tool generates draft text that a named writer verifies against source before the study director approves, with the verification captured in the audit trail" is observable; the record either shows the verification or it does not. The protocol's later sections, the acceptance criteria and the performance qualification, are essentially a test of whether the intended-use statement is being honored, so the statement must be phrased in terms that those tests can actually check. A strategist who writes an aspirational intended-use statement has built a protocol that cannot prove anything.
Fitness for Purpose and the Risk Assessment That Sets the Rigor
Fitness for purpose is the bridge between the intended use and the amount of validation rigor you commit. It is not a yes-or-no judgment; it is a risk-based determination of how much evidence the intended use demands, and it is where the protocol either calibrates effort sensibly or wastes it. The core question is the consequence of an undetected error in the tool's output, assessed in the context of the controls that surround it. For the toxicology drafting tool, the consequence of an undetected fidelity error is potentially severe, because a misattributed histopathology finding or a mis-transcribed NOAEL-relevant value that survives into Module 2.6.6 can mislead an IND reviewer about the safety margin supporting first-in-human dosing. That is a high-consequence output, and the risk assessment must say so plainly.
But consequence is only half the calculation. The other half is the probability that an error reaches the patient-relevant decision despite the controls, and here the human verification step is the dominant control. Because the intended use requires a qualified writer to verify every claim against source and the study director to approve the final report, an error in the raw AI draft has to pass through two independent human checks to reach a reviewer. The risk assessment must evaluate how reliable those checks actually are, not assume they are perfect, because the most common validation failure is a protocol that claims human review as a control without testing whether the reviewers can actually catch the errors the tool makes. If the tool produces a fluent, confident draft and the verification is a light read rather than a claim-by-claim reconciliation, the control is weaker than the protocol assumes, and the rigor must rise to compensate.
The output of the risk assessment is a tiering of the tool's output types by consequence, and that tiering drives differentiated acceptance criteria. The structural scaffolding of the report, the headings, the standard descriptive phrasing, is low consequence and can be validated lightly. The transcribed quantitative values and the finding-to-dose-group attributions are high consequence and demand tight acceptance criteria and dense testing. The cross-references to summary tables are highest consequence, because, as the mechanics of language models make clear, a cross-reference is a claim about a document structure the model often cannot see, and a fabricated table reference survives every check except deliberate reconciliation. The risk assessment that does not separate these tiers will either over-test the headings or under-test the cross-references, and the second error is the one that fails an inspection.
Acceptance Criteria That Survive Non-Determinism
Acceptance criteria are where the probabilistic nature of the tool collides hardest with the deterministic instincts of classical validation, and getting them right is the technical heart of the protocol. A classical computerized-system protocol writes acceptance criteria as input-output matches: given this input, the system shall produce this exact output. That formulation is impossible for an LLM, because the same input over the same sources can produce different acceptable phrasings at any temperature above zero, and an acceptance criterion that demands an exact output will fail a correct draft for using a synonym. The protocol must therefore shift from output-matching to property-checking: it asserts invariants that every acceptable output must satisfy, regardless of how the prose happens to be worded on a given run.
For the toxicology tool, the invariants are concrete and checkable. Every quantitative value in the draft must match the corresponding value in the validated source within zero tolerance; a mean organ-weight change of 12 percent in the source must read as 12 percent in the draft, and any deviation is a failure. Every histopathology finding must be attributed to the dose group the source assigns it to; an arm or group swap is a failure regardless of how fluent the surrounding sentence is. Every cross-reference to a summary table must resolve to a table that exists and contains the cited result; an unresolvable reference is a failure. No finding the study director classified as adverse may be described as non-adverse, or the reverse. These are properties, not phrasings, and they hold across the output distribution, which is exactly what non-determinism requires.
The acceptance criteria must also set the bar quantitatively, because "the tool is accurate" is not a criterion anyone can pass or fail against. The protocol commits to a target, for example that across the performance-qualification test set the tool produces zero unresolved cross-references and zero quantitative-value errors, and that any finding-attribution error rate above a defined threshold fails the qualification. Setting these numbers is a strategic act: set them too loose and the validation proves nothing, set them too tight against an inherently probabilistic system and nothing passes. The defensible move is zero tolerance on the highest-consequence properties, the values, the attributions, the cross-references, because those are the ones that mislead a reviewer, paired with a documented residual-risk acceptance that the human verification step is the control that catches the rare miss the tool will inevitably produce. The protocol does not pretend the tool is perfect; it proves the system around the tool is safe.
Performance Qualification: Proving the Claim With a Real Test Set
Performance qualification is where the protocol stops asserting and starts proving, and its quality is determined almost entirely by the test set. The classical IQ/OQ/PQ frame still applies: installation qualification confirms the tool is deployed in the validated configuration, operational qualification confirms it functions and that its controls, the audit trail, the access controls, the grounding configuration, behave as specified, and performance qualification confirms that in the hands of real users on real work it meets the acceptance criteria. But for an AI tool the performance qualification carries unusual weight, because the operational qualification cannot demonstrate correctness the way it can for deterministic software; correctness is a statistical property of the output distribution, and only a representative test set can characterize it.
The test set must be built from real, validated toxicology study data spanning the range of cases the tool will encounter in production: single-dose and repeat-dose studies, studies with clean findings and studies with complex, dose-dependent histopathology, studies with subtle adverse-versus-non-adverse classifications that stress the tool's fidelity. For each case, the source data and the human-verified correct output are known, so each AI draft can be scored against the acceptance criteria property by property. The test set must be large enough that the measured error rates are meaningful rather than anecdotal, and it must include the hard cases deliberately, because a performance qualification run only on clean studies proves the tool works when nothing is difficult and tells you nothing about the cases where it fails. A strategist who lets the test set drift toward the easy cases has built a qualification that passes and a tool that is not actually qualified for the work.
The performance qualification must also be run on the configuration that will actually be deployed, including the specific model version, the system prompt, the grounding and retrieval setup, and the temperature, because all of these shape the output distribution and a qualification of a different configuration validates a tool you are not using. This is why version pinning is a validation requirement and not an IT preference: if the qualification was run on model version X at temperature 0.2 with a specific grounding configuration, then production must run that same configuration, and any change to it, a vendor model update, a system-prompt edit, a retrieval change, is a change to the validated state that triggers re-assessment. The performance qualification produces the evidence; the change-control discipline preserves the conditions under which that evidence remains valid. Without the second, the first expires silently the next time the vendor ships an update.
Ongoing Performance Monitoring and the Part 11 Record
The protocol does not end at the performance qualification, because a validated state for an AI tool is a maintained condition, not a one-time achievement, and the section that commits to ongoing performance monitoring is what makes the validation durable. A point-in-time qualification proves the tool met the acceptance criteria on a given day with a given configuration; it says nothing about next quarter, when the vendor may have updated the model, the writing group may have started using the tool on a new study type, or subtle drift may have crept into the output quality. The monitoring section defines the metrics that are tracked in production, the cross-reference resolution rate, the human-edit rate on quantitative values, the frequency of finding-attribution corrections caught at verification, and the thresholds on those metrics that trigger investigation or re-qualification. This is the bridge to the ongoing-monitoring lesson that follows, and it is what the fitness-for-purpose and lifecycle-monitoring principles of the FDA-EMA framework demand.
The monitoring is only possible because the system is built to produce the record in the first place, which returns the protocol to 21 CFR Part 11 and EU Annex 11. The audit trail must capture, for every report the tool touches, the source data provided, the model and version, the system prompt and configuration, the temperature, the raw draft, the human edits and the identity of the verifier, and the final approved text, all attributable and contemporaneous and enduring. This is not optional metadata; it is the evidence that the intended use was honored, the means by which monitoring metrics are computed, and the record an investigator reads at a Pre-Approval Inspection to confirm that what the protocol promised actually happened on the real reports in the submission. A protocol that defines beautiful acceptance criteria but does not require the system to log the evidence has validated a tool whose compliance cannot be demonstrated, which is to say it has validated nothing an inspector will accept.
The strategist's final responsibility in the protocol is to make the monitoring actionable rather than ceremonial. A metric that is collected and never reviewed is worse than no metric, because it creates a record that the function knew the data existed and did not act on it. The protocol therefore assigns ownership of the monitoring review, sets the cadence, and defines the escalation path when a threshold is crossed, tying the AI tool's monitoring into the function's existing quality system rather than bolting on a parallel process no one runs. Done this way, the validation protocol is not a document that sits in a binder; it is a living control that produces evidence continuously, catches degradation before it reaches a submission, and gives the strategist a defensible, current answer when the Chief Quality Officer or the FDA investigator asks the question that started this lesson: why should anyone trust this tool?
Key Takeaways
- The validation protocol is the strategic artifact where ambition becomes testable and defense becomes possible. It is the bridge between the probabilistic reality of an LLM and the deterministic evidence world of 21 CFR Part 11 and EU Annex 11, and every downstream defense, inspection, ROI claim, governance report, leans on its specificity. A vague protocol is indistinguishable from no protocol in front of an investigator.
- The intended-use statement is the load-bearing sentence, and it must be written so compliance with it is observable. It scopes the tool to drafting from validated source data for human verification, names the human accountability, and explicitly forecloses scientific conclusions, because the fitness-for-purpose principle validates the tool only against its defined intended use and operating outside it is operating unvalidated.
- Acceptance criteria must shift from output-matching to property-checking to survive non-determinism. Instead of demanding an exact output, the protocol asserts invariants every acceptable draft satisfies: zero-tolerance quantitative fidelity, correct finding-to-dose-group attribution, resolvable cross-references, and preserved adverse classifications, with the human verification step accepted as the control for the residual rare miss.
- Performance qualification proves the claim only if the test set is representative and the configuration is the deployed one. The test set must include the hard cases deliberately and be scored property by property against validated correct outputs, and the qualification must run the exact model version, system prompt, grounding, and temperature that production uses, which is why version pinning is a validation requirement and any change triggers re-assessment.
- A validated state is a maintained condition, sustained by ongoing monitoring built on a complete Part 11 record. The audit trail must capture source, model and version, configuration, raw draft, human edits and verifier identity, and final text, both as the evidence an inspector reads and as the data that feeds the monitoring metrics, with ownership, cadence, and escalation assigned so the monitoring is actionable rather than ceremonial.
Skill.re