AI Hallucinations: When AI Invents a Trial, a Dose, or a TLF Reference
There is a specific kind of mistake that AI makes which is unlike any mistake a human colleague makes, and learning to fear the right one is the difference between a writer who is safe with these tools and a writer who is dangerous with them. A junior medical writer who is unsure of a number leaves it blank, or writes "TBC," or asks. A large language model that is unsure of a number writes a confident, specific, plausible number, formatted exactly like a real one, with no signal whatsoever that it was invented. This behavior is called hallucination, and in a regulated submission it is not a quirk to be charmed by; it is the central risk the entire verification discipline exists to contain. This lesson is a careful catalog of how AI hallucinates in a life-sciences context, drawn from the failure modes documented across 2024 to 2026, organized so that you can recognize each one before it reaches a dossier. The goal is not to make you distrust the tool. It is to make you distrust the tool in exactly the right places, because a fabricated grammatical flourish is harmless and a fabricated dose is a patient-safety statement that should never have existed.
Why Hallucination Is Not a Bug You Can Wait to Be Fixed
The first thing to understand is that hallucination is not an engineering defect that a future version will patch away. It is a direct consequence of how the model works. A large language model generates the most plausible next token given everything before it, and "plausible" is a statement about the shape of the training data, not about the truth of your trial. When the model has seen ten thousand efficacy sections that cite a table, the most plausible continuation of your efficacy section is a table citation, whether or not the table exists. When it has seen countless dose statements of the form "administered at 10 mg/kg every two weeks," that pattern is available to complete even when your protocol used a different dose. The model is not lying, because lying requires knowing the truth and choosing against it. The model has no access to the truth; it has access to the shape of plausible text.
This matters because it sets the right expectation. Better models hallucinate less, because their sense of plausibility is more refined and grounding techniques constrain them more tightly, but no model hallucinates never, and a function that waits for the hallucination-free model is waiting for something the architecture cannot deliver. The correct response is not to hope the problem goes away but to build the detection and verification that assumes it will not. Every later lesson on verification, audit trail, and human-in-the-loop is, at bottom, a response to this one fact: the tool will sometimes produce a confident falsehood, and your job is to catch it every time it matters.
The Fabricated Citation: The Most Dangerous Pattern in the Catalog
The single most dangerous hallucination in regulated writing is the fabricated cross-reference, and it deserves the top of the catalog because of a property the others lack: it is invisible to the ordinary defenses. A fabricated citation to Table 14.2.1.4 in a Module 2.5 efficacy section passes spell-check, because it is spelled correctly. It passes a grammar review, because it is grammatically perfect. It passes a hasty read by a busy co-author, because it looks exactly like every real table citation in the document. It is caught only when someone goes to the actual table package and finds that Table 14.2.1.4 does not exist, or exists and says something different, and that check happens late in the publishing cycle, by which point the false citation may have been copied into Module 2.7.3 and the integrated summary.
The reason this pattern is so persistent is that cross-references are claims about the structure of documents the model frequently cannot see. If the final TLF package was not in the context window, the model has no way to verify the table number against anything, but the pattern of an efficacy section demands a citation, so it produces a plausible one. The same failure produces fabricated journal citations in literature summaries, where a model "remembers" a paper that sounds exactly like real papers in the field and does not exist, and fabricated bulletin or guidance numbers, where it cites a section of an ICH guideline that is real-sounding and wrong. In every case the danger is identical: a structural claim that wears the costume of a verified fact and is detected only by going to the source. The discipline that contains it is equally identical: treat every citation as a starting point for checking, not as proof, and treat any citation you cannot check as wrong until proven right.
The Invented Dose and the Invented Number: When Fabrication Becomes a Safety Statement
The second pattern raises the stakes from credibility to safety. A model summarizing a protocol or a CSR can produce a dose, a frequency, a route of administration, or a result that was never in the source, and the consequence depends entirely on where that statement lands. A fabricated dose in an internal orientation note is an error to be corrected; a fabricated dose in an Informed Consent Form, an Investigator's Brochure, or a Module 2.5 safety section is a statement that could, if it reached a clinician or a patient, cause harm. This is the pattern that most clearly distinguishes life-sciences hallucination from hallucination in a low-stakes domain, because the artifacts here touch patients.
The invented number is the broader version of the same pattern. A hazard ratio of 0.68 where the true value is 0.71, a median progression-free survival of 11.2 months where the true value is 10.4, a confidence interval that is plausible and wrong, a subject count that is off by two. None of these announces itself. The model produces 0.68 with exactly the confidence it would produce 0.71, and on the page they are indistinguishable. The particular trap is that a plausible wrong number is harder to catch than an absurd one; a hazard ratio of 4.2 in a positive trial would jump out, but 0.68 instead of 0.71 sits comfortably in the range a reviewer expects, so only reconciliation against the actual TLF cell catches it. This is why the verification standard for any quantitative claim is not "does it look reasonable" but "does it match the source," and why the source has to be loaded before the claim can be trusted.
The Invented Relationship: Plausible Reasoning That Connects the Wrong Dots
The third pattern is subtler and harder to teach because it does not involve a single wrong token but a wrong connection between true ones. A model can take two real facts and assert a relationship between them that the source does not support: that an adverse event was treatment-related when the CSR characterized it as unlikely related, that a subgroup benefit was statistically significant when it was nominal and not multiplicity-controlled, that a finding in one study confirms a finding in another when the studies are not comparable. Each component fact may be real; the asserted relationship is the fabrication.
This pattern is dangerous precisely because it mimics reasoning, and reasoning is what reviewers and writers are trained to trust. A sentence that says "the consistent benefit across subgroups, together with the favorable safety profile, supports a positive benefit-risk" reads like an argument, and if the subgroup consistency was overstated or the safety profile was characterized more favorably than the data support, the argument is built on a fabricated relationship while every individual word is defensible. The defense here is different from the defense against an invented number, because there is no single cell to check. The defense is to verify the relationships, not just the facts: to ask of every causal, comparative, or evaluative claim whether the source actually supports that connection, and to keep the originating judgments, the causality call, the significance interpretation, the benefit-risk conclusion, with the named human who is accountable for them. The next lesson takes up exactly why those judgments cannot be delegated.
The Confident Misalignment With Guidance: Wrong About the Rules
The fourth pattern is hallucination about the regulatory framework itself, and it is easy to miss because it sounds authoritative. Asked how a section should be structured or what a guideline requires, a model can state a requirement that is plausible and wrong: that ICH E3 requires a particular subsection it does not, that a Type C meeting has a different time clock than it does, that an eCTD module accepts content it does not. The model has absorbed a great deal of regulatory text, so its statements about the rules carry the surface authority of someone who has read the guidances, but it can confidently assert a misremembered rule with the same fluency as a correct one.
This pattern is particularly insidious because the people most likely to rely on it are those least equipped to catch it, the newer professionals using AI to learn the framework. A senior RA director will notice that the model has the refuse-to-file clock wrong; a first-year associate will not, and will carry the error into a gap-closure plan. The defense is to treat the model as a fast but unreliable index to the guidance, never as the guidance itself, and to confirm any specific regulatory requirement against the actual document of record before acting on it. The same playbook treats named windows and clocks, the kind of precise figures this program is careful to source, as facts to verify rather than to recall, because a confidently misremembered clock is exactly the kind of plausible error the architecture produces.
Where Hallucinations Cluster: The Gap, the Edge, and the Handoff
Hallucinations are not distributed evenly across a document; they concentrate in predictable places, and knowing where they cluster lets a writer aim verification rather than spread it thin. The first cluster is the gap, any point where the source is silent on something the output format expects. When a template has a field for time-to-onset and the intake did not capture it, when an efficacy section conventionally cites a table and the table package was not loaded, when a summary calls for a cross-study comparison the reports do not actually support, the model meets a blank that the pattern says should be filled, and filling blanks is exactly where invention happens. The practical signal is that any field the format demands but the source did not supply is a hallucination hotspot, and the defense is to require a source locator for that field and treat its absence as missing rather than letting the model paper over it.
The second cluster is the edge, the boundary of the model's reliable knowledge, where specific named facts live: exact regulatory clocks, precise statistical thresholds, particular form numbers, individual study identifiers. These are facts the model has seen many similar versions of, which is precisely what makes it confident and unreliable about the exact one, because the many similar versions blur into a plausible average that may not match your specific case. The defense at the edge is the verify-not-recall rule: any precise named figure is confirmed against the document of record, never trusted from the model's memory.
The third cluster is the handoff, the seam where one person's or one tool's output becomes another's input. A fabricated citation born in a Module 2.5 draft becomes dangerous at the handoff to the 2.7.3 writer who copies it, and an extracted field that silently became a generated one is most invisible at the moment it passes into the next step looking like clean structured data. The defense at the handoff is to verify at the seam, to treat every point where content changes hands as a checkpoint rather than a pass-through, because an unverified output that crosses a handoff inherits the credibility of the receiving step and becomes much harder to trace back. Together, the gap, the edge, and the handoff are a map of where to look, and a writer who patrols those three places catches the large majority of hallucinations with a fraction of the effort of reading everything with equal suspicion.
Why the Fabricated Table Outranks the Typo: A Ranking of Consequence
Not all hallucinations are equally dangerous, and a writer who treats them as a uniform threat will spend verification effort in the wrong places. The ranking that matters runs by two axes: how invisible the error is to ordinary review, and how consequential it is if it survives. A fabricated TLF cross-reference scores high on both, invisible to spell-check and grammar review, and consequential because it propagates across modules and surfaces as a reviewer Information Request. An invented dose in a patient-facing document scores at the top of consequence even if it is more catchable, because the harm is direct. An invented relationship scores high on invisibility because it mimics reasoning. A confident misalignment with guidance scores high on invisibility for junior users specifically. A genuine typo or an awkward phrasing scores low on both and barely deserves the name hallucination.
The practical lesson is to spend verification where the product of invisibility and consequence is highest. That means every quantitative claim and every cross-reference in a submission document gets reconciled to source, every causal and evaluative relationship gets checked against what the source actually supports, every regulatory requirement gets confirmed against the document of record, and the patient-facing artifacts get the heaviest scrutiny of all. It also means that the verification is not a vague injunction to "review carefully" but a specific, claim-typed discipline: this is a number, check the cell; this is a citation, check the target; this is a relationship, check the support; this is a rule, check the guidance. That specificity is what turns an awareness of hallucination into a habit that actually catches it, and it is the bridge from this lesson to every workflow in the rest of the program.
One closing reframing helps the discipline stick. The goal of all this is not to make you slower or more anxious; it is to make your trust precise. An experienced writer using AI well does not read every sentence with paralyzing suspicion, because that is unsustainable and would surrender the speed that makes the tool worth using. Instead, they trust the model exactly where it is trustworthy, the structure, the phrasing, the compression, and they withhold trust exactly where it is not, the specific number, the specific citation, the asserted relationship, the stated rule. Calibrated trust, not blanket suspicion, is the mark of the professional, and it is learnable precisely because the dangerous patterns are finite and predictable. Once you can name the four hallucination types and the three places they cluster, verification stops being a vague burden and becomes a targeted, fast, and even satisfying act of professional control over a powerful but fallible tool.
Key Takeaways
- Hallucination is not a bug awaiting a patch; it is intrinsic to next-token generation. The model produces the most plausible continuation, and plausible is about the shape of training data, not the truth of your trial. Better models hallucinate less, none hallucinate never, so the response is detection and verification, not waiting.
- The fabricated cross-reference is the most dangerous pattern because it is invisible to ordinary defenses. A citation to a nonexistent Table 14.2.1.4 passes spell-check, grammar review, and a hasty read, and is caught only at the source, often after it has propagated into Module 2.7.3. Treat every citation as a starting point for checking and any uncheckable one as wrong.
- An invented dose or number is a safety and integrity failure, and a plausible wrong value is harder to catch than an absurd one. A hazard ratio of 0.68 instead of 0.71 sits in the expected range; only reconciliation against the actual TLF cell catches it. The standard is "does it match the source," not "does it look reasonable."
- The invented relationship fabricates a connection between true facts, mimicking reasoning. Overstated subgroup significance or a too-favorable safety characterization can build a defensible-sounding benefit-risk argument on a false relationship. Verify the relationships, not just the facts, and keep causality, significance, and benefit-risk judgments with the named human.
- Spend verification where invisibility times consequence is highest, with a claim-typed discipline: this is a number, check the cell; a citation, check the target; a relationship, check the support; a rule, check the guidance. Patient-facing artifacts get the heaviest scrutiny, and any specific regulatory clock or requirement is verified against the document of record, never recalled from the model.
Skill.re