Computer Software Assurance (CSA) vs Validation (CSV) for AI Tools
For two decades, the answer to "how do we validate this software" in a regulated pharma function was a reflex: write the protocol, script the test cases, execute every one, screenshot the evidence, and build a binder thick enough to reassure anyone who never reads it. That reflex produced an industry-wide pattern of over-documenting low-risk software and, paradoxically, under-thinking high-risk software, because the energy went into paperwork volume rather than into the question of what could actually go wrong. On 24 September 2025, the FDA finalized "Computer Software Assurance for Production and Quality System Software," a CDRH and CBER guidance that supersedes Section 6 of the 2002 General Principles of Software Validation, and it codifies a different reflex: think first about risk, then test the way the risk demands, and document only what the assurance actually requires. For an AI Function Strategist standing up a portfolio of LLM-based tools, CSA is the most useful regulatory development of the year, because it gives you a defensible way to spend validation effort where the risk is and stop spending it where the risk is not. This lesson explains what CSA actually changes, where it earns its keep for AI tools, where it explicitly does not reach, and what the function has to do in 2026 to re-baseline its AI-tool inventory against the final guidance.
What CSA Actually Changes, and What It Does Not
The most common misreading of CSA is that it is "less validation," a permission slip to do less work. That framing will get a function into trouble, because CSA does not lower the assurance bar; it reallocates the effort. Computer Software Validation as practiced under the old reflex front-loaded scripted testing and documentation uniformly, regardless of risk, so a spreadsheet template and a system controlling a sterilization cycle could end up with comparably heavy test binders. CSA's central move is to make risk the first question and the determinant of everything that follows: you scope the intended use of the software, you assess what could go wrong and how bad it would be, and only then do you choose the assurance activities and the documentation proportionate to that risk. The high-risk software gets more and better testing than it did before; the low-risk software gets less ceremony and more critical thinking.
The second change is the legitimization of testing methods that the old reflex treated as second-class. CSA explicitly recognizes unscripted testing, ad hoc and exploratory and error-guessing approaches, as valid assurance activities, not just the fully scripted, pre-written, step-by-step test cases that dominated CSV practice. For lower-risk features, a skilled tester exercising the software thoughtfully and recording what they did and found can be entirely sufficient, and that is a profound efficiency unlock because scripted test authoring and execution is where the old approach burned most of its hours. The guidance also accepts that vendor activities, supplier testing, and existing evidence can be leveraged rather than re-performed, so the function is not re-validating what a credible vendor already validated. None of this is a license to skip assurance; it is a license to stop manufacturing evidence that the risk never called for.
What CSA does not change is the destination. The records still have to satisfy 21 CFR Part 11, the system still has to do what it is intended to do, and the function still has to be able to demonstrate that to an investigator. CSA changes how you get there and how you justify the route, not whether you arrive. A strategist who sells CSA internally as "we can do less" has set up the function to under-assure something high-risk and then discover, at an inspection, that "we adopted CSA" is not a defense for having skipped the critical thinking. CSA is more thinking and less typing, and the thinking is the part that is hard to fake.
The Risk-Based Critical-Thinking Pattern, Step by Step
CSA operationalizes as a repeatable pattern, and learning the pattern is what lets a strategist apply it consistently across a portfolio rather than re-arguing it tool by tool. The first step is intended-use scoping: state precisely what the software feature does in your process and what decision or record depends on it. This is the same discipline as the intended-use statement in a validation protocol, and it is the foundation, because risk is meaningless in the abstract; the risk of an AI drafting tool depends entirely on what its output feeds and who checks it before it matters. A feature whose output is always human-verified against source before it influences anything is a different risk than a feature whose output flows directly into a released record.
The second step is the risk determination: for the scoped intended use, what could go wrong, and what would be the consequence if it did, specifically the consequence for product quality or patient safety, which is the axis CSA cares about. This is where CSA asks the question the old reflex skipped, namely whether a failure of this feature could affect the safety or quality of the product, and it sorts features into higher and lower process-risk accordingly. A feature controlling a process parameter that determines product sterility is high process-risk; a feature that drafts text a qualified human fully verifies is lower process-risk, not because the text does not matter but because the human verification stands between the feature and any product-quality consequence. The honesty of this step is everything, because inflating or deflating the risk corrupts every downstream decision.
The third step chooses the assurance activity proportionate to the risk: high process-risk features get rigorous, often scripted testing with robust evidence; lower process-risk features can be assured through unscripted or exploratory testing or by leveraging vendor and existing evidence. The fourth step records the assurance commensurately, capturing what was tested, by whom, what was found, and the basis for the risk determination, without the reflexive volume that the old approach demanded everywhere. The pattern is deliberately the same every time, because consistency is what makes it defensible: an investigator who sees the function apply intended-use scoping, honest risk determination, proportionate testing, and commensurate documentation uniformly across its tools is looking at a system, not a set of ad hoc choices, and a system is what passes.
Where CSA Reduces Validation Effort for AI Tools
For a large class of AI tools in a regulated function, CSA is a genuine and defensible reduction in effort, and recognizing which tools qualify is a core strategic skill. The clearest case is the AI tool whose output is always human-verified against source before it influences any released record, the drafting assistant for a Module 2.5 section, the toxicology narrative drafter, the ICSR narrative first-drafter, where a qualified person reconciles every claim before the content goes anywhere. For these tools, the human verification step is a powerful risk control that sits between the AI output and any product-quality or patient-safety consequence, and CSA's logic recognizes that a strong downstream control lowers the process-risk of the feature itself, which in turn justifies lighter, often unscripted assurance of the tool.
This is the most important conceptual move for the strategist to internalize: under CSA you assess the risk of the feature in the context of its controls, not in isolation. An AI drafting tool considered alone looks frightening, because it is non-deterministic and can hallucinate. The same tool considered inside a workflow with mandatory claim-by-claim human reconciliation and a Part 11 audit trail is a lower process-risk feature, because the failure mode that matters, an undetected error reaching a released record, is gated by a control. CSA lets you take credit for that control in the risk determination, which means the validation effort can concentrate on confirming the control actually works, the verification is happening and being recorded, rather than on exhaustively scripting the inherently unscriptable behavior of the model. That is both more efficient and more honest about where the real assurance comes from.
The efficiency also comes from leveraging vendor evidence. A credible AI vendor that has validated its own platform, characterized its model behavior, and supports a documented configuration lets the function assure the tool partly by reference rather than by re-performing the vendor's work. The strategist's job is to evaluate whether the vendor's evidence is actually credible and applicable to the function's intended use, which is exactly the vendor-evaluation discipline of an earlier chapter, and then to assure the part that is genuinely the function's own: the configuration, the grounding sources, the system prompt, the workflow controls, and the verification step. CSA does not eliminate assurance of the AI tool; it lets the function stop re-validating what was validated elsewhere and focus its scripted, rigorous testing on the few things that are genuinely high process-risk and genuinely its own responsibility.
Where CSV-Level Rigor Still Applies for Product-Impacting AI
The flip side of CSA's efficiency is the discipline to recognize where it does not apply, and a strategist who only learns the effort-reduction half of CSA has learned the dangerous half. Where an AI feature's output can affect product quality or patient safety without a reliable human control standing in the way, the process-risk is high and the assurance must be rigorous, scripted, and thorough, which is to say CSV-level. This is the case for any AI whose output flows into a released record, a control decision, or a regulatory conclusion without effective independent verification, and the strategist's responsibility is to identify these honestly rather than reaching for the lighter path because it is cheaper. The whole value of CSA's risk-based logic is destroyed if the risk determination is biased toward the answer that saves work.
The specific trap with AI tools is the temptation to claim human verification as a control when the verification is not actually effective. CSA lets a strong control lower the process-risk, but the control has to be real, and as the validation-protocol lesson established, fluent AI output invites light reading rather than genuine reconciliation, so a verification step that exists on paper but is not reliably performed does not lower the risk in practice. If the function cannot demonstrate that the human verification genuinely catches the errors the AI makes, then the AI feature must be assured as if that control were weak, which pushes it back toward rigorous testing. The honest risk determination forces the function to either make the control real, with evidence, or accept the higher assurance burden, and pretending the control is real when it is not is the exact failure that converts a CSA efficiency into an inspection finding.
There is one boundary that the strategist must never blur, because it is a hard line in the guidance itself: CSA applies to software used in production and the quality system, and it explicitly does not apply to Software-as-a-Medical-Device or Software-in-a-Medical-Device. If the function's AI is, or becomes, a medical device software function, it is governed by the device frameworks, including the predetermined-change-control-plan pathway for AI-enabled device software functions, not by CSA's production-software logic. Most AI used to draft submission content sits comfortably on the production and quality-system side of that line, but the strategist owns the determination, and a tool that drifts toward making a clinical or diagnostic claim can cross into SaMD territory where a different regulatory regime applies. Knowing which side of the line a tool sits on is part of scoping its intended use, and getting it wrong applies the wrong assurance framework entirely.
Re-Baselining the AI-Tool Inventory in 2026
The finalization of CSA in September 2025 created a concrete 2026 work item that the strategist owns: re-baselining the function's AI-tool inventory against the final guidance. Many functions validated their early AI tools under the old CSV reflex, producing heavy scripted-test binders for tools whose actual process-risk, given the human verification around them, did not warrant that volume, while sometimes under-thinking the genuinely higher-risk tools. The arrival of the final guidance is the natural trigger to revisit the whole inventory: re-scope each tool's intended use, re-assess its process-risk in the context of its current controls, and re-allocate assurance effort accordingly, lightening where the old approach over-documented and strengthening where it under-assured. This is not busywork; it is the moment the function converts its validation spend from uniform to risk-proportionate.
The re-baselining is also a cross-functional act, which is why it lands on the strategist rather than on QA alone. The QA function owns the quality-system interpretation of CSA, IT owns the configuration and version control that the assurance depends on, and the CSV or CSA leads own the test design, but the AI Function Strategist owns the intended-use scoping and the risk-determination logic for the AI tools specifically, because those determinations require understanding how the AI behaves, what its failure modes are, and how the workflow controls actually function. A re-baselining run by QA without the strategist's input tends to mis-scope the AI risk, either treating every LLM as terrifyingly high-risk or waving them all through as low-risk office software, and both errors produce an inventory that fails on contact with an inspector who understands AI.
The output of a good re-baselining is an inventory where every AI tool has a documented intended-use statement, an honest process-risk determination tied to its specific controls, an assurance approach proportionate to that risk, and a clear record of why the chosen approach is sufficient. That inventory is itself a strategic asset, because it is the document the strategist hands to a Quality Council, an internal audit, or an FDA investigator to demonstrate that the function's AI estate is under control. It also feeds directly into the ongoing-monitoring discipline of the next lesson, because the process-risk determination for each tool sets how intensively that tool's production performance needs to be watched. CSA, applied honestly, is not a one-time validation event; it is the risk-based logic that organizes the function's entire approach to assuring its AI, from the first scoping decision through the monitoring that keeps the assurance current.
Defending the CSA Approach to the Quality Council
The strategist will have to defend the CSA approach to a Quality Council or a Chief Quality Officer who may be cautious about anything that looks like reduced validation, and the defense has to be framed correctly to land. The wrong frame is "CSA lets us do less validation," which triggers exactly the caution that kills the proposal, because it sounds like cutting corners on compliance. The right frame is "CSA lets us put our validation effort where the risk actually is, which means our high-risk AI tools get more rigorous assurance than they would have under uniform CSV, and our low-risk tools stop consuming effort that produced documentation no one used." That frame is both true and aligned with what a quality leader actually wants, which is risk-proportionate control, not maximum paperwork.
The defense is strongest when it is anchored to the guidance and to the FDA's own intent, because CSA is the agency's framework, not the function's invention. The strategist can point out that CSA supersedes Section 6 of the 2002 software-validation guidance, that the FDA itself observed that the old approach drove over-documentation of low-risk software at the expense of critical thinking, and that adopting CSA aligns the function with the current regulatory expectation rather than departing from it. This reframes the proposal from "we want to do less" to "we want to do what the FDA finalized as good practice," which is a fundamentally easier conversation. It also connects to the FDA-EMA fitness-for-purpose and risk-based principles, so the AI-specific application of CSA sits inside a coherent governance story rather than looking like a special pleading for the AI tools.
Finally, the defense must concede honestly what CSA does not do, because a quality leader trusts a strategist who names the limits. CSA does not lower the Part 11 bar, does not apply to SaMD, does not excuse a dishonest risk determination, and does not work without the cross-functional discipline to apply it consistently. A strategist who presents CSA as a balanced reallocation, with explicit boundaries and an honest account of where rigor increases, earns the Council's confidence in a way that an over-promised efficiency story never could. The re-baselined AI inventory, with its honest risk determinations and proportionate assurance, is the proof that the function can wield CSA responsibly, and that proof is what converts a cautious Quality Council from an obstacle into a sponsor of the function's AI program.
Key Takeaways
- CSA reallocates validation effort by risk; it does not lower the assurance bar. The FDA's final guidance of 24 September 2025 supersedes Section 6 of the 2002 software-validation guidance and makes risk the first question, so high-risk software gets more and better testing while low-risk software gets less ceremony and more critical thinking. Selling CSA internally as "do less" sets the function up to under-assure something high-risk and discover at inspection that adopting CSA is not a defense.
- The critical-thinking pattern is repeatable: scope the intended use, determine the process-risk honestly, choose proportionate assurance, and document commensurately. Applied uniformly across the portfolio, the pattern is what makes the approach defensible, because an investigator sees a system rather than ad hoc choices, and the honesty of the risk determination is the part that cannot be faked.
- CSA reduces effort for AI tools whose output is always human-verified, because it lets you assess risk in the context of controls. A drafting tool considered alone looks frightening; the same tool inside a workflow with mandatory claim-by-claim reconciliation and a Part 11 audit trail is lower process-risk, so assurance can concentrate on confirming the control works and on leveraging credible vendor evidence rather than scripting the unscriptable.
- CSV-level rigor still applies where AI output can affect product quality or patient safety without an effective human control, and CSA does not apply to SaMD or SiMD at all. The specific trap is claiming human verification as a control when it is not actually effective; if the function cannot demonstrate the verification catches the errors, the AI must be assured as if the control were weak, and a tool that crosses into a clinical or diagnostic claim leaves CSA's production-software scope entirely.
- Re-baselining the AI-tool inventory against the final guidance is the strategist's 2026 work item and a cross-functional act. Re-scope every tool, re-assess process-risk against current controls, and re-allocate assurance, producing an inventory with documented intended use, honest risk determinations, and proportionate assurance, which becomes the asset handed to a Quality Council or investigator and the input that sets each tool's ongoing-monitoring intensity.
Skill.re