The Limits of AI Reasoning in Causality, Sameness, and Benefit-Risk
There is a tempting and dangerous idea that as models get better, the list of things only a human can do will shrink toward nothing, and that the judgments we reserve for people today are reserved only because the technology is not yet good enough. For most of the structure-bound work in the previous lesson, that idea is roughly right; the models are already excellent and getting better. But there is a specific class of judgments in drug development where the reservation is not about model capability at all, and where a frontier model that drafts a flawless Module 2.5 around the judgment still cannot make the judgment itself. Causality assessment, comparability and sameness arguments, and benefit-risk integration are the three clearest examples, and understanding why they resist delegation, even to a very good model, is what keeps a professional from the most seductive error in this whole field: handing the machine the one decision that the regulation, the science, and the accountability structure all insist a named human must own. This lesson explains the why, because the why is what makes the boundary stable rather than a temporary limitation waiting to be overtaken.
What These Three Judgments Have in Common
Before taking the three in turn, it helps to see what unites them, because the common thread is the reason the boundary holds. Each of these judgments requires integrating evidence that is incomplete, weighing considerations that do not reduce to a formula, and accepting personal accountability for a conclusion that reasonable experts could dispute. They are not lookups, where a correct answer exists in a source and the task is to find it. They are not even consistency checks, where the task is to detect divergence in existing material. They are acts of judgment under uncertainty, where the evidence underdetermines the answer and a qualified human has to decide and stand behind the decision.
A model can do something that looks remarkably like this. It can produce the prose of a causality assessment, structure the argument of a comparability conclusion, and assemble the components of a benefit-risk statement, because it has seen thousands of examples of each and can complete the pattern fluently. But producing the prose of a judgment is not the same as making the judgment, and the difference is invisible on the page, which is exactly what makes it dangerous. The model's benefit-risk paragraph and the medical officer's benefit-risk paragraph can read identically; only one of them is backed by a qualified person who has weighed the evidence and accepts accountability for the conclusion. The regulation cares about that backing, not about the prose, and so should the writer. The skill is to let the model assemble the structure and the supporting material while ensuring the actual judgment, and the accountability for it, rests with the named human who is qualified to hold it.
Causality Assessment: Why the WHO-UMC Call Stays Human
When a serious adverse event occurs in a trial or in the post-market setting, someone has to assess whether the drug caused it, and the assessment drives reporting obligations, label changes, and signal decisions. Frameworks like WHO-UMC and the Naranjo algorithm exist to structure this judgment, and their existence can mislead a newcomer into thinking causality is a calculation: feed in the temporal relationship, the dechallenge, the rechallenge, the plausibility, and out comes the category. But the frameworks structure the judgment; they do not replace it. The temporal relationship is often ambiguous, the dechallenge is confounded by other changes in the patient's care, the biological plausibility is a matter of expert interpretation, and the same case can reasonably be assessed as possible by one experienced assessor and probable by another.
A model can apply the framework's structure and even propose a category, and that proposal can be a useful starting point, the way a checklist is useful. But the assessment requires medical judgment about a specific patient with incomplete information, integrating the clinical narrative, the concomitant medications, the underlying disease, and the literature, in a way that the assessor must be qualified to perform and willing to defend. When that assessment determines whether a case is a serious unexpected reaction requiring an expedited report, the consequence of the call is regulatory and the accountability is personal, and a category proposed by a pattern-completer cannot carry that weight. The model can organize the evidence for the assessor; it cannot be the assessor. This is why pharmacovigilance workflows that use AI for narrative drafting and triage still route the causality and the listedness determination to the qualified human, every time, and why "the tool assessed it as unlikely related" is never an acceptable basis for the call.
Sameness and Comparability: Why ICH Q5E Conclusions Resist Automation
The second judgment lives in the world of biologics manufacturing and biosimilars, and it is even less amenable to automation because the science is genuinely hard. When a manufacturer changes a biologic's process, ICH Q5E governs the comparability exercise that establishes whether the pre-change and post-change product are comparable, meaning highly similar with no adverse impact on safety or efficacy. When a biosimilar developer builds a 351(k) application, an analytical-similarity exercise establishes whether the proposed product is sufficiently similar to the reference. In both cases, the conclusion rests on a totality-of-evidence judgment that integrates analytical data, functional assays, and where needed clinical data, against acceptance criteria that themselves require justification.
The reason this resists automation is that the conclusion is not a threshold test that data either passes or fails; it is an argument that a body of evidence, taken together, supports a scientific claim about similarity, and the FDA reviewer at the Office of Pharmaceutical Quality is evaluating the quality of that argument. A model can draft the comparability protocol, structure the data presentation, and even assemble a first-pass narrative, and these are real contributions. But the judgment about whether the totality of evidence actually supports comparability, whether an observed analytical difference is or is not meaningful for the product's mechanism, whether the acceptance criteria were appropriately set, is a scientific judgment that a qualified person must make and defend. The characteristic failure of a model here is the one the program flags elsewhere: it can produce a comparability argument that tests everything the manufacturer already knew how to test and calls the result comparable, which is exactly the argument that earns a major deficiency, because the reviewer is judging the adequacy of the evidence, not the fluency of the narrative. The model cannot make that adequacy judgment, because it requires knowing what the evidence does not show, which is precisely what a pattern-completer cannot reliably surface.
Benefit-Risk Integration: The Judgment at the Center of the Dossier
The third judgment is the one the entire submission exists to support: whether, for the proposed indication and population, the benefits of the drug outweigh its risks. It lives most visibly in the Module 2.5.6 benefit-risk conclusion, and it is the judgment a reviewer, an advisory committee, and ultimately a regulator's signature are all oriented around. It integrates the efficacy results, the safety profile, the severity of the disease, the available alternatives, the uncertainties in the data, and the manageability of the risks, into a conclusion that no formula produces and that reasonable experts can and do dispute, which is why advisory committees vote rather than calculate.
This is the judgment that most clearly cannot be delegated, and the reason is the convergence of everything in this lesson. The evidence is incomplete and must be weighed, the considerations do not reduce to a formula, and the accountability is the most consequential in the entire enterprise, because a wrong benefit-risk conclusion is a public-health error. A model can assemble every input to this judgment, the efficacy summary, the integrated safety analysis, the comparison to alternatives, and that assembly is genuinely valuable, because organizing the evidence for the human who must decide is real work that AI does well. But the integration itself, the act of weighing and concluding, must be performed and owned by qualified humans, and the FDA-EMA principles' accountability principle is explicit that this responsibility does not transfer to a tool or a vendor. When a writer lets a model draft the 2.5.6 benefit-risk paragraph, the writer must understand that the model has drafted the prose of a conclusion, and that the conclusion itself still has to be made, verified against the evidence, and owned by the named accountable person, or the most important sentence in the submission is unbacked.
The Seductive Error: Mistaking Fluent Assembly for Judgment
The common failure across all three is a single seductive error: mistaking the model's fluent assembly of a judgment's components for the judgment itself. The error is seductive precisely because the model is so good at the assembly. It produces a causality narrative that reads like an expert wrote it, a comparability conclusion that reads like a scientist defended it, a benefit-risk statement that reads like a medical officer integrated it, and the fluency creates a powerful illusion that the thinking has been done. It has not. The thinking, the weighing of incomplete evidence and the acceptance of accountability, is exactly the part the model cannot do, and the better the assembly, the stronger the illusion that it can.
Protecting against this error is not about distrusting the model's output; it is about understanding what the output is. When the model produces a benefit-risk paragraph, the correct mental model is that it has produced a draft of how the conclusion might be expressed, populated from the available evidence, which a qualified human must now treat as a starting point for the actual integration, not as the integration. The verification is therefore different in kind from verifying a number or a citation. You do not check a benefit-risk conclusion against a source cell, because there is no cell; you check whether the conclusion is one a qualified person, having weighed the evidence, is prepared to make and defend. That check can only be performed by such a person, which is the entire point. The model is the most capable assistant a regulatory or medical writer has ever had, and it is categorically not the decision-maker for the judgments at the center of the dossier, and a professional who holds that distinction clearly will use AI aggressively for assembly and never once let it hold a judgment it cannot be accountable for.
Where the Line Gets Tested: Significance, Listedness, and the Tempting Middle
The three judgments above are the clear cases, but a professional has to handle the middle ground, where a task looks more delegable than it is, and the middle is exactly where careless delegation happens. Consider statistical significance interpretation. A model can correctly report that a subgroup p-value was below 0.05, which is a fact a reviewer can check, and from there it is a short and dangerous step to letting the model characterize that result as a meaningful treatment effect. But whether a nominal subgroup p-value supports a claim depends on the multiplicity structure of the analysis, the pre-specification, and the consistency with the overall result, and that characterization is a judgment, not the lookup it resembles. The model reporting the number is on safe ground; the model interpreting the number is over the line, and the two can sit in the same sentence.
Listedness in pharmacovigilance is a second tempting middle. Determining whether an adverse reaction is listed in the reference safety information looks like a matching task, and a model can genuinely help by retrieving the relevant section of the company core data sheet. But the determination of whether a specific reported event is or is not covered by a listed term, with all the clinical nuance of how the event was described and how the term is defined, is the judgment that drives expectedness and therefore expedited reporting, and it stays with the qualified safety reviewer. The model retrieves the candidate listing; the human decides whether it applies. The pattern across both middles is the same diagnostic question: is the task retrieving or reporting a fact, which AI may do with verification, or is it interpreting or concluding from facts under uncertainty, which the qualified human must own? Asking that question of every task that feels delegable is what keeps the line from quietly eroding one convenient shortcut at a time.
Why This Boundary Is Stable, Not Temporary
It is worth being precise about why this boundary will not simply erode as models improve, because that precision is what lets a function build durable policy rather than a stopgap. The boundary is not primarily about capability; it is about accountability and the nature of judgment under uncertainty. Even a model that could weigh evidence as well as a human expert, which is not today's model, would still not be a qualified person who can be held accountable for a regulatory conclusion, because accountability attaches to a person with a name, a qualification, and a professional and legal responsibility, and a tool has none of these. The FDA-EMA principles encode this directly: sponsor accountability does not transfer to a vendor or a system. The named author owns the 2.5; the qualified person owns the causality call; the medical officer owns the benefit-risk; and no improvement in model capability changes who can be held responsible when the conclusion is wrong.
This is liberating rather than limiting, because it tells a function exactly where to invest its human expertise and exactly where to deploy AI without hesitation. Pour the model into the assembly, the summarization, the structuring, the consistency-checking, the drafting around the judgment, and pour human expertise into the judgments themselves and the verification that the assembled evidence actually supports them. A function that draws the line here is not being cautious for caution's sake; it is aligning its workflow with the immovable fact that judgment under uncertainty, and accountability for it, belong to qualified people, which is the same fact the regulators built their principles around. The next lesson takes this from principle to the practical floor, the ALCOA+ standard that governs how AI output, including the assembled drafts around these human judgments, has to be handled to be defensible at all.
Key Takeaways
- Three judgments resist delegation even to a very good model: causality assessment (WHO-UMC, Naranjo), comparability and sameness arguments (ICH Q5E, biosimilar 351(k) analytical similarity), and benefit-risk integration (Module 2.5.6). The reservation is not about model capability but about judgment under uncertainty and personal accountability.
- All three share a structure: integrating incomplete evidence, weighing considerations that do not reduce to a formula, and accepting accountability for a conclusion reasonable experts could dispute. They are acts of judgment, not lookups or consistency checks, and a model can produce the prose of the judgment without making it.
- The frameworks structure the judgment; they do not replace it. WHO-UMC categories and ICH Q5E comparability are not calculations; the temporal relationship is ambiguous, the totality-of-evidence conclusion is an argument a reviewer judges, and the same case can reasonably be assessed differently by qualified experts.
- The characteristic CMC failure is the model that tests what the manufacturer already knew how to test and calls it comparable, which earns a major deficiency because the reviewer judges the adequacy of the evidence, not the fluency of the narrative, and the model cannot surface what the evidence does not show.
- The boundary is stable, not temporary, because it rests on accountability, not capability. A tool has no name, qualification, or legal responsibility, and the FDA-EMA principles state that sponsor accountability does not transfer to a vendor. Deploy AI aggressively for assembly around the judgment; keep the judgment and its ownership with the qualified human.
Skill.re