โ†
AI for Construction & AEC
Aware ยท M1 ยท lesson 1 of 17 ยท in progress
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI Hallucinations in Construction Documents
๐Ÿ“–
now learning

AI Hallucinations in Construction Documents

15 min

A hallucination is not the AI being broken. It is the AI working exactly as designed, producing the most plausible-sounding text, on a question where plausible and true happen to diverge. In casual use this is harmless; if it invents a fake restaurant recommendation you shrug. In construction documents it is a landmine, because a fabricated spec section, a phantom code citation, or a made-up UL listing does not look wrong. It looks exactly like the real thing, sails through a busy reviewer, and surfaces only when a plan checker, an arbitrator, or an inspector catches it at the worst possible moment. This lesson teaches you to see the specific flavors of fabrication that show up in AEC documents, why each one happens, and how to catch them before they cost you a rejection, a claim, or a stamp.

Why Fabrication Is the Default, Not the Exception

You met the mechanism in the LLM lesson, and it is worth stating bluntly here because it changes how you read every AI-drafted document: the model does not know the difference between a fact it retrieved and a fact it invented, because it does not retrieve facts at all. It predicts plausible text. When you ask for a spec section governing traffic-bearing waterproofing, the model produces a string that has the shape of a real section number, formatted correctly, sitting in a grammatically perfect sentence, and whether that exact number exists in your project manual is simply not something the model checked, because checking is not what it does.

This means hallucination is not a rare malfunction you will occasionally encounter; it is the baseline behavior on every factual claim the model makes without a retrieval system behind it. The fluent, correct-looking citation is not evidence the model "knew" the answer. It is evidence the model is good at producing fluent, correct-looking citations, which it is, whether or not the underlying fact is real. Internalize this and you stop being surprised, and more importantly you stop being fooled, because you read every specific citation in an AI draft as a claim to verify rather than a fact to trust. The danger is entirely in the gap between how confident it looks and how unverified it actually is, and in construction that gap is where rejections and claims live.

A hallucinated citation does not look like an error. It looks like a real citation, formatted perfectly, in a confident sentence. That is exactly why it is dangerous: it is built to pass the glance that a busy reviewer gives it.

The Flavors of Fabrication You Will Actually See

Construction documents have a particular vocabulary of references, and the model fabricates each kind in a recognizable way. Knowing the flavors lets you aim your verification at the highest-risk spots rather than re-reading everything.

The first and most common is the invented spec section. The model produces a MasterFormat-style number, something like "Section 09 81 17," that is formatted exactly like a real section and does not exist in MasterFormat 2024 or in your project manual. It is convincing precisely because the format is right; the six-digit structure, the division logic, all of it looks correct, and only checking the actual specification reveals that the section is not there or does not say what the model claims.

The second is the fabricated contract-clause number. Ask about notice requirements and the model may cite "AIA A201 Section 4.5.2" with total confidence when the clause it is describing actually lives elsewhere, or when that number governs something entirely different. This flavor is especially dangerous because contract clauses drive notice windows and claim rights, so a wrong clause number in a notice can undermine the very entitlement the notice was meant to preserve. We have a whole later lesson on the specific A201 clauses people conflate, and the reason that lesson exists is that the model conflates them constantly.

The third is the made-up standard or test citation: a phantom ASTM designation, a non-existent ASHRAE section, a fabricated NEC article, or an invented UL listing number. These are catnip for the model because the formats are so regular, ASTM E followed by digits, NEC article numbers, that producing a plausible fake is trivial, and these citations often go into submittals and product approvals where a fake listing could wave through a non-compliant product. The fourth flavor is subtler and arguably worse: the conflation, where the model does not invent a fake reference but misapplies a real one, summarizing OSHA by blending 1910 (general industry) with 1926 (construction) as if they were one standard, or describing a real code provision but attaching the wrong section number or the wrong edition. The reference exists; the application is wrong; and that is harder to catch because a quick check confirms the reference is real without revealing that it does not say what the draft claims.

The Real Cases That Make This Concrete

This is not theoretical. The pattern that recurs across firms is the AI-drafted code narrative that cites IBC sections which do not exist. An intern racing a permit set asks for a means-of-egress narrative, the model produces clean, professional prose, and salted through it are section numbers that sound authoritative and are fabricated, the infamous egress section that appears confidently and exists in no edition of the code. The narrative reads beautifully. A plan checker who knows Chapter 10 cold bounces it within the hour, because to anyone who actually knows the code the fake sections are glaring even though to the busy intern they looked fine.

The OSHA conflation is equally common and more dangerous because it touches safety. Ask a general model to summarize the fall-protection or scaffolding requirements and it will happily blend 29 CFR 1910, which governs general industry, with 29 CFR 1926, which governs construction, presenting a tidy summary that is wrong about which standard applies to your jobsite. Since the two standards differ in real, consequential ways, a pre-task plan or toolbox talk built on the conflated summary can instruct a crew to the wrong requirement, which is a safety failure dressed as a documentation convenience. The lesson in both cases is the same: the output was fluent, professional, and confidently wrong, and only a human who knew the source caught it.

How to Catch Them Without Re-Doing the Work

The instinct after hearing all this is either to distrust AI entirely or to laboriously re-verify everything, and both throw away the value. The disciplined middle is targeted verification: you do not re-read the whole draft for tone or grammar, you hunt specifically for the reference claims and check each one at its source. Every spec section, every clause number, every standard citation, every code section, every product listing, gets confirmed against the actual document, and everything else, the prose, the structure, the argument, you treat as the fast first draft it is.

This is far less work than it sounds, because the reference claims are a small fraction of any document and they cluster in predictable places: the citations in a code narrative, the section references in an RFI, the standard numbers in a submittal. You learn to scan for them the way a proofreader scans for a particular kind of error. A practical habit that helps enormously is to require the model, in your prompt, to either cite only references it can ground in a document you provided or to explicitly say "I cannot find a governing section" instead of inventing one, which a retrieval-based tool can actually honor and a plain chatbot cannot, telling you immediately which kind of tool you are dealing with. The verification still happens, but the honest tool fabricates far less to begin with.

One more defense matters: never let the model both write a citation and confirm it. Asking the same model "are you sure that section is correct?" produces more prediction, and it will often double down confidently on its own fabrication, because agreeing that the section is right is just as plausible a continuation as inventing it was. Verification has to come from outside the model, from the published code, the actual contract, the real standard, every time. The model is not a witness to its own accuracy.

What It Costs When a Fabrication Survives the Glance

It helps to follow a single hallucination all the way to its consequence, because the abstract risk only becomes real when you trace where a fake citation ends up. Picture a fabricated spec section that nobody caught. It went into an RFI response, which was logged, which a sub relied on to order material, which was installed, and now the installed condition references a spec section that does not exist, so when the owner's inspector asks for the basis of the installation there is no governing section to point to. The fabrication did not stay a documentation error; it became a physical condition and a contractual exposure, and unwinding it costs far more than the thirty seconds of verification that would have caught it at the source.

The contract-clause version is worse because it can quietly destroy an entitlement. A delay notice that cites the wrong A201 clause, or asserts a notice window that does not apply to the clause it named, can be challenged on exactly that basis, and an owner's counsel who finds a fabricated or misapplied citation in your notice now has a thread to pull on the whole claim. The notice was supposed to preserve your rights; the hallucination inside it handed the other side an argument. And the safety version is the gravest of all: a pre-task plan built on a conflated OSHA standard that instructs a crew to a general-industry requirement on a construction site is not just wrong on paper, it can put a worker in a condition the correct standard would have prevented, and "the AI summarized it that way" is not a defense that helps anyone in a hospital or a deposition.

None of these consequences announce themselves at the moment of the error. That is the entire problem with hallucination in construction documents: the cost is deferred and amplified, hidden at the moment of drafting and revealed at the moment of inspection, dispute, or injury, which is precisely when it is most expensive to fix. The verification at the source is cheap and immediate; the consequence of skipping it is expensive and delayed. That asymmetry is the whole economic argument for the discipline.

Why Better Models Reduce This But Never Erase It

A fair objection at this point is that the models keep getting better, so surely hallucination is a problem that will solve itself. It is getting less frequent, and retrieval-based tools that ground answers in real documents truly reduce it, sometimes dramatically. But it does not go to zero, and the reason is structural rather than a temporary limitation. As long as the underlying engine generates plausible text rather than retrieving verified fact, there is always some probability it produces a confident, well-formed claim that is not true, and a lower rate of fabrication is in one way more dangerous, not less, because rarer errors lull you into dropping your guard.

Think about it from the verification side. If a tool fabricates one citation in three, you stay sharp because you are catching errors constantly. If a better tool fabricates one in fifty, you relax, and the one fabrication in fifty sails through precisely because you stopped looking, and it is no less wrong than the one in three. The improvement in the model does not change the consequence of the single error that gets through; it only changes how complacent you have become when it does. This is why the verification discipline cannot be tied to how good the model is. It has to be a fixed step in the process regardless of the tool's accuracy, because the step exists to catch the rare residual error, and the rare residual error is exactly the one your vigilance has stopped expecting. The better the tool, the more disciplined, not less, you have to be about the step, because the tool has quietly removed the steady drumbeat of obvious errors that used to keep you honest.

The Applied Problem: Mark Every Fabrication in a Code Narrative

Here is the exercise that builds the reflex. Take an AI-generated two-page code narrative on means of egress, the kind of output a general chatbot produces in seconds for an IBC question, and audit it against the published IBC 2024 Chapter 10 with the actual code open beside you. Do not skim for plausibility; check every single citation.

For each section the narrative cites, go to the published code and verify three things in order: does the section exist, does it say what the narrative claims, and is it from the correct adopted edition. Mark each fabrication with the published-code reference that disproves it, "narrative cites 1029.6.4; no such section exists in IBC 2024; the governing egress-width provision is at the actual section," because that markup is both your correction and your evidence. You will typically find a mix: some citations real and correctly applied, some real but misapplied (the conflation flavor), and at least one pure invention. Cataloging which flavor each error is trains you to predict where the next draft will fail.

The deliverable is a marked-up narrative plus a short tally: how many citations total, how many fabricated, how many misapplied, how many correct. That tally is quietly powerful, because the first time you run it you will be startled by how many confident citations were wrong, and that startle is the lesson landing. From then on you will never read an AI-drafted citation as a fact again, only as a claim awaiting the source, which is exactly the posture that keeps a fabricated section out of a permit set, a fake listing out of a submittal, and a conflated OSHA standard out of a toolbox talk. The reflex you build here is the one that protects every document the rest of this program teaches you to draft.

Key Takeaways

  • A hallucination is the model working as designed: producing plausible text on a question where plausible and true diverge. It is the default behavior on every factual claim made without a retrieval system, not a rare malfunction.
  • A fabricated citation does not look wrong. It looks like a real citation, formatted perfectly, in a confident sentence, built to pass the glance a busy reviewer gives it. The danger lives in the gap between how confident it looks and how unverified it is.
  • The flavors recur: invented spec sections (a perfectly formatted "Section 09 81 17" that does not exist), fabricated AIA clause numbers (which can undermine the notice they appear in), made-up ASTM/NEC/UL/ASHRAE citations, and conflation (misapplying a real reference, like blending OSHA 1910 and 1926).
  • The real cases are common: AI code narratives citing IBC sections that exist in no edition, and OSHA summaries that conflate general-industry and construction standards, which can route a crew to the wrong safety requirement.
  • Catch them with targeted verification: hunt the reference claims (a small, predictable fraction of any document) and check each at its source, treating the prose as a fast draft. Require the model to cite only grounded references or say it cannot find one.
  • Never let the model confirm its own citation; asking "are you sure?" produces more prediction and often a confident doubling-down. Verification comes from the published source, every time.
  • The artifact: audit a two-page code narrative against IBC 2024 Chapter 10, mark every fabrication with the disproving reference, and tally total, fabricated, misapplied, and correct. The startle from that first tally is the reflex that protects every document you draft.