Bias in Clinical AI: Trial Enrollment, RWE Cohorts, and Adverse Event Coding
A clinical operations lead at a mid-size oncology sponsor is staring at a ranked list of 140 candidate investigator sites for a Phase 3 trial in metastatic non-small-cell lung cancer. An AI site-selection tool, trained on the enrollment performance of the sponsor's last six trials, has scored every site on predicted enrollment velocity, and the top thirty are almost entirely large academic centers in a handful of affluent metropolitan areas. The model is confident, the dashboard is clean, and the recommendation is internally consistent. It is also quietly reproducing a pattern that has skewed oncology enrollment for two decades, because the historical data the model learned from was generated by a recruitment machine that already over-selected those exact centers and under-enrolled community sites, rural populations, and several racial and ethnic groups. The model did not invent the bias. It inherited it, polished it, and handed it back as a prediction. This lesson is about how biased training data becomes biased clinical decisions across three specific artifacts, the site-selection model, the MedDRA coding suggestion, and the subgroup analysis, and why the FDA-EMA principle of data quality and lifecycle management makes this your problem, not the vendor's.
Where Bias Actually Enters a Clinical AI System
Bias in clinical AI is not a moral failing of the model and it is not a bug in the code. It is a faithful statistical reflection of the data the model was trained on, surfaced as a recommendation that looks neutral. A machine learning model learns the patterns present in its training corpus, and if that corpus encodes a historical pattern of who got enrolled, who got coded a certain way, or whose subgroup was large enough to analyze, the model reproduces that pattern as if it were a law of nature. The danger is precisely that the output does not look biased. It looks like data. A ranked list of sites, a suggested Lowest Level Term, a proposed subgroup split, all arrive with the same even confidence whether they are equitable or skewed.
The FDA-EMA Guiding Principles of Good AI Practice in Drug Development, released on 14 January 2026, name this directly under the principle of data quality and lifecycle management. The principle holds that the data used to develop and operate an AI system must be representative, relevant, and fit for the context in which the system is used, and that its quality must be managed across the whole lifecycle, not certified once at launch and then forgotten. The reason the principle is framed around lifecycle rather than a one-time check is that data drifts: the population enrolling in trials changes, coding conventions evolve across MedDRA versions, and a model frozen on yesterday's distribution becomes progressively less fit for today's. Representativeness is not a property you verify once. It is a property you monitor.
For the clinical professional, this reframes the question. The question is never simply "is the model accurate?" because a model can be highly accurate at predicting the biased outcome it was trained on. The question is "accurate at what, measured against which population, and does that population match the one I am actually deciding about?" A site-selection model that is 90 percent accurate at predicting which sites enrolled fast historically is 90 percent accurate at reproducing the historical enrollment geography, including everything wrong with it. Accuracy against a biased target is not a defense. It is the mechanism of harm.
Biased Site Selection and the Enrollment Skew It Produces
Site selection is the first place bias becomes a clinical decision with downstream consequences for the entire trial. AI site-selection and predictive-enrollment tools, the kind embedded in platforms across the clinical operations landscape, rank candidate sites by predicted enrollment velocity, predicted data quality, and predicted retention, using features drawn from historical trial performance, claims data, and electronic health record footprints. The intended value is real: a sponsor that picks high-performing sites finishes enrollment faster and spends less. But the features that predict fast historical enrollment are heavily correlated with site size, urban location, academic affiliation, and a patient population that has been recruited into trials before, and all of those correlate in turn with the demographic skew that has made oncology trial populations persistently unrepresentative of the patients who will eventually take the drug.
Across 2024 to 2026, this pattern has been documented repeatedly as site-selection skew in oncology AI tools, where models trained on past enrollment systematically down-rank community oncology practices, safety-net hospitals, and sites serving predominantly Black, Hispanic, rural, and lower-income populations, not because those sites perform poorly but because they were historically under-used and therefore under-represented in the training data. The model sees a community site with little trial history and predicts low enrollment, which is partly a self-fulfilling artifact: the site has little history because it was rarely selected, and the model now recommends not selecting it again, deepening the very gap it is measuring. This is a feedback loop, and a feedback loop is the most dangerous shape bias can take, because each cycle makes the skew look more justified by the data.
The regulatory stakes are concrete and rising. The FDA's Diversity Action Plan expectations, strengthened under the Food and Drug Omnibus Reform Act, require sponsors to set and justify enrollment goals for clinically relevant demographic subgroups, and a site-selection process that an AI tool steered toward a homogeneous set of centers is a process the sponsor must be able to defend against those expectations. A clinical operations lead who accepts the model's top thirty sites without examining the demographic reach of that set has not made a faster decision. They have made an indefensible one, because the resulting enrollment may fail to support the subgroup conclusions the label will need, and that failure surfaces late, after the sites are activated and the money is spent.
Biased MedDRA Coding Suggestions at the LLT Level
The second artifact is the adverse event code, and here bias operates at a finer grain that is easy to miss. MedDRA, the Medical Dictionary for Regulatory Activities, organizes adverse event terms in a hierarchy that runs from the Lowest Level Term, the LLT, up to the Preferred Term, the PT, and ultimately to the System Organ Class, the SOC. When a pharmacovigilance case arrives describing a patient's reaction in the reporter's own words, that verbatim text must be coded to an LLT, and the choice of LLT determines how the event aggregates, how it signals, and how it appears in the safety database and the PSUR. AI coding tools now suggest the LLT from the verbatim, and the suggestion is fast and usually plausible.
The bias risk is that the model's mapping from verbatim language to LLT was learned from a historical coding corpus, and that corpus carries the linguistic and demographic patterns of who reported and how. Verbatim descriptions phrased in non-clinical language, in translated text, or in the idiom of populations under-represented in the training data may be mapped less accurately, defaulting to a more generic or a systematically different LLT than the equivalent description phrased in the clinical English the model saw most often. A symptom described by a patient in one community may be coded to a precise, signal-relevant LLT, while the same symptom described differently is coded to a vaguer term that dilutes the signal. Because coding decisions aggregate, a small, consistent skew at the LLT level can suppress or distort a safety signal for exactly the subpopulation least able to afford it.
This is why the LLT suggestion is a draft and never a decision. The coder owns the mapping, and the discipline is to read the verbatim against the suggested LLT and ask whether the term truly captures the reported event or whether the model reached for the statistically common code rather than the correct one. MedDRA coding is governed by the MedDRA term selection points to consider, a published convention, and the human coder applies that convention; the model only proposes. A coder who accepts AI LLT suggestions without reading the verbatim is allowing a historical coding bias to propagate into the signal-detection layer of the safety system, where it is far harder to see and far more consequential.
Biased RWE Cohorts in Komodo, TriNetX, and Aetion
The third place bias compounds is the real-world evidence cohort, and the stakes are higher here because the underlying data was never collected for research. Real-world evidence platforms such as Komodo Health, TriNetX, and Aetion build cohorts from claims data, electronic health records, and linked datasets, and AI is woven through the cohort construction, the phenotype definition, and the analysis. These platforms surface patterns across populations no manual analysis could reach, which is a genuine and growing capability. But a cohort is only as representative as the data that feeds it, and claims and EHR data systematically under-capture the uninsured, the under-insured, populations with fragmented care, and groups that interact less with the documented healthcare system.
When an AI-constructed cohort inherits these gaps, a real-world finding can be confidently wrong about the population it claims to describe. A treatment effect estimated in a TriNetX or Komodo cohort that skews toward commercially insured, urban, well-documented patients may not hold for the patients excluded by the data's own structure, and a comparative analysis run in Aetion is only valid for the population the data actually represents. The bias is not introduced by the platform; it is present in the source data and faithfully carried into the cohort, which is exactly the lifecycle data-quality concern the FDA-EMA principle names. The analyst's first obligation is to characterize what the cohort represents and, just as importantly, what it omits, before any conclusion is drawn about effectiveness, safety, or the product's real-world value.
The interpretive judgment is irreducibly human. A pattern in a real-world cohort is a hypothesis shaped by the data's reach, not a fact about the disease, and turning it into a defensible claim for a payer dossier, a label discussion, or a publication requires a validity assessment the model cannot perform. The RWE platforms are powerful precisely because they process scale, and they are dangerous for the same reason: scale makes a biased finding look authoritative. The discipline is to treat every AI-generated cohort as a population question first and an analysis second, asking who is in, who is out, and whether the conclusion survives that accounting.
Biased Subgroup Analysis Suggestions and the Statistics of Who Gets Studied
The fourth artifact is the subgroup analysis, and bias here is subtle because it hides inside a legitimate statistical concern. When an AI tool suggests which subgroups to analyze in an efficacy or safety dataset, it tends to propose the subgroups that are large enough to yield statistically stable estimates, which means it proposes the subgroups that were well enrolled and de-prioritizes the ones that were not. This is statistically reasonable and demographically corrosive at the same time: the subgroups most at risk of being under-studied are exactly the ones the historical enrollment skew left small, and an AI that optimizes for statistical power recommends studying the already-well-studied and skipping the rest.
The result is a quiet narrowing of the evidence base for under-represented populations, dressed as methodological rigor. A subgroup-analysis suggestion that omits a racial or ethnic group because the enrolled count is low is not neutral; it is the enrollment bias from the first section reappearing one analysis downstream, where it determines what the label can say about whom. And when an AI tool does propose a small-subgroup analysis, it may not adequately flag the instability of the estimate, presenting a hazard ratio from a thin cell with the same confidence as one from a robust cell, which invites over-interpretation in the opposite direction. Either failure, omission or false precision, distorts the benefit-risk picture for the very populations the Diversity Action Plan framework is meant to protect.
The named author of the statistical analysis plan and the clinical summary owns these choices. The decision about which subgroups to analyze, how to caveat thin cells, and how to characterize a finding for an under-represented group is a scientific and regulatory judgment that does not transfer to the tool. An AI can enumerate the statistically convenient subgroups in seconds; the human decides which subgroups the regulation, the population, and the science require, and labels the limits of each honestly.
The Lifecycle Discipline That Makes Bias Defensible
The thread connecting all four artifacts is that bias is not detected by looking at the output, because the output looks like data. It is detected by interrogating the data behind the output and monitoring it over time, which is precisely what the FDA-EMA data quality and lifecycle management principle requires. For the site model, that means examining the demographic and geographic reach of the recommended set against the population the drug will serve, not just the predicted enrollment velocity. For MedDRA coding, it means reading the verbatim against the suggested LLT rather than accepting the mapping. For RWE cohorts, it means characterizing inclusion and exclusion before interpreting effect. For subgroup analysis, it means deciding which subgroups the science requires rather than which the data made convenient.
Lifecycle is the operative word because each of these systems drifts. The enrolling population changes, MedDRA versions update terms, claims and EHR coverage shifts, and a model validated once on a past distribution silently degrades. The defensible posture is to treat representativeness as a monitored property: to record what population a tool was trained and validated on, to check periodically whether the current decision population still matches, and to document the human review that caught and corrected the skew. This is not anti-AI. The tools genuinely accelerate site selection, coding, cohort construction, and analysis, and the professional who uses them well moves faster than one who does not. The discipline is to use them with the data-quality question always in front of the recommendation, because a biased decision made quickly is still a biased decision, and on these four artifacts it is one the sponsor, not the vendor, will be asked to defend.
Key Takeaways
- Bias in clinical AI is a faithful reflection of biased training data, surfaced as a recommendation that looks neutral. A model can be highly accurate at reproducing a biased historical outcome; accuracy against a biased target is the mechanism of harm, not a defense. The FDA-EMA principle of data quality and lifecycle management names representativeness as a property to monitor, not certify once.
- AI site-selection tools reproduce documented oncology enrollment skew by down-ranking community, rural, and demographically diverse sites that were historically under-used. This is a self-fulfilling feedback loop, and a homogeneous selected set undermines the subgroup conclusions a label needs and the FDA Diversity Action Plan expectations the sponsor must defend.
- Biased MedDRA LLT coding suggestions can mis-map verbatim text from under-represented populations to vaguer or systematically different Lowest Level Terms, diluting safety signals where they aggregate. The coder reads the verbatim against the suggested LLT and owns the mapping; the model only proposes a draft.
- RWE cohorts in Komodo Health, TriNetX, and Aetion inherit the gaps of claims and EHR data, which under-capture uninsured and fragmented-care populations. The analyst must characterize who the cohort represents and who it omits before interpreting any effect, because scale makes a biased finding look authoritative.
- AI subgroup-analysis suggestions optimize for statistical power, which recommends studying the well-enrolled and skipping the under-enrolled, re-expressing enrollment bias one analysis downstream. The named author of the SAP and clinical summary owns which subgroups the science and regulation require and labels thin-cell estimates honestly.
Skill.re