โ†
AI for Manufacturing
Aware ยท M1 ยท lesson 1 of 19 ยท in progress
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI Hallucinations on the Floor
๐Ÿ“–
now learning

AI Hallucinations on the Floor

15 min

It was a Tuesday on the third shift, and a new process tech named Marco had a torque value he could not find. The work instruction for a bracket assembly referenced a fastener spec that lived in a drawing nobody could locate at 2 a.m., and the line was down waiting on it. Marco did what a lot of people now do when the answer is not on the shelf: he asked a generative AI assistant. He typed in the part number, the bolt size, the material, and asked for the assembly torque. In about four seconds the model gave him a clean, confident answer: 24 newton-meters, with a tidy little note that this was "standard for an M8 fastener in this grade." It looked right. It read like something an engineer would say. Marco torqued forty units to 24 Nm and released them to the next station. The problem is that the actual print called for 14 Nm, because the joint clamped a soft aluminum boss that crushes above 16. The model had not read the drawing. It had never seen the drawing. It produced a number that was statistically plausible for an M8 bolt in general and had nothing to do with this joint. Three days later the customer found cracked bosses in receiving inspection, and a containment that cost the plant about $38,000 in sort, rework, and expedited freight started with four confident seconds of a machine that does not know what it does not know. That is a hallucination, and on a manufacturing floor it does not stay an abstract AI quirk. It ships.

What a Hallucination Actually Is on the Floor

The word "hallucination" sounds like a malfunction, like the model broke. It did not break. A generative AI model produced exactly what it is built to produce: the most statistically likely next words given everything it was asked. A large language model (the kind of AI behind a chat assistant) does not store facts the way a drawing stores a torque spec or the way your historian stores a tag. The historian (the database that logs every sensor reading on the floor, also called the data historian) holds a real number that a real sensor measured at a real timestamp. The model holds patterns about how words and numbers tend to go together. When you ask it for a torque value, it is not looking up your joint. It is predicting what a plausible answer looks like, the way autocomplete predicts the rest of your sentence, except it will finish the sentence with a number and present it with the same calm confidence whether the number is right or invented.

That is the part that catches working people off guard. A wrong answer from a person usually comes with a tell. The new tech hesitates, says "I think it's around 20, let me check the print." A hallucination has no tell. The model does not know it is wrong, so it cannot sound unsure. It produces the fabrication in the same tone, the same format, the same crisp engineering register as a correct answer. There is no flag, no error code, no red light. On a floor where people are trained to trust a clean, well-formatted output, that smooth confidence is the danger, not a feature.

A hallucination is not the model breaking. It is the model doing exactly what it does, which is predict plausible text, and there is no built-in tell that separates a true answer from an invented one.

Here is the analogy that lands on the floor. Picture the most confident talker you ever worked with, the guy who had an answer for everything and was right maybe eighty percent of the time but never, ever said "I don't know." The other twenty percent he just made something up that sounded right, and he said it with the exact same certainty as the things he actually knew. A generative model is that guy, scaled up and given a clean font. The skill that kept you safe around that coworker is the exact same skill that keeps you safe around AI: you learned which claims to take at face value and which ones to walk over to the machine and verify. The model never earns the "take at face value" status on a technical number. Not because it is usually wrong, but because you cannot tell from the answer alone which kind of answer you got.

This matters more in 2026 than it did even a year ago, because adoption is no longer a pilot. Roughly 47 percent of manufacturers now use AI in quality work, up from about 33 percent the prior year. The model is already in your plant, drafting, summarizing, answering. The question is not whether your crew will lean on it. They already are, the way Marco did at 2 a.m. The question is whether they know the three specific ways it lies on technical work, and how to catch each one before it reaches the line.

Failure One: The Invented Spec

The first failure mode is the one that bit Marco: the invented spec. You ask the model for a number that is supposed to be exact, a torque value, a tolerance, a temperature, a cure time, a clearance, a pressure, and it hands you a number that is plausible for the general category and wrong for your specific part. This is the most dangerous failure because numbers carry an aura of precision. A paragraph of prose invites skepticism. A number like "14.2 Nm" feels like it came from somewhere measured. It did not. It came from the average of every M8 torque value the model ever saw in its training, blended into something that looks like a measurement.

Consider the worked example fully, because the dollars are the lesson. Marco's joint needed 14 Nm. The model said 24 Nm. That is not a rounding error. That is a 71 percent overshoot on a soft aluminum boss with a crush limit around 16 Nm. Forty units went out at 24. The defect did not show up at Marco's station, because an over-torqued boss does not always crack immediately; it cracks later, under thermal cycling or vibration, which is exactly why the customer caught it three days downstream. By then those forty parts were inside subassemblies. The containment had to pull every unit built in that window, sort the good from the suspect, rework or scrap the cracked ones, and air-freight clean parts to keep the customer's line running. Sort labor, rework, scrap, freight, and the engineering hours to write the 8D (the eight-discipline corrective-action report a customer requires after an escape) added up to about $38,000. The model's answer cost four seconds to get and $38,000 to clean up.

The reason the invented spec is so common is structural, not occasional. The model has seen thousands of torque specs across the internet and technical documents. M8 fasteners in steel often run in the 20-to-25 Nm range, so when you ask for an M8 torque without telling the model that this joint clamps soft aluminum with a crush limit, the most statistically likely answer it can produce is in that 20-to-25 band. It is not guessing wildly. It is giving you the population average for the category, which is precisely wrong for the specific case where your joint differs from the average. The more your real part deviates from the textbook norm, and manufacturing is nothing but parts that deviate from the norm, the more confidently the model will hand you the norm.

How a working engineer catches it. The catch is not clever; it is disciplined. Any number that will be torqued, cut, heated, pressed, or measured against gets traced to the controlled source before it is used. The controlled source is the drawing, the released spec, the supplier's documented value, the standard the customer flowed down. The rule is brutally simple: a generative model is never the source of a spec, only a pointer toward where the real spec might live. If the model says 24 Nm, the response is not "great" and it is not "wrong," it is "show me the print." When you cannot find the print, you do not torque to the model's number. You hold the line and escalate, because a held line costs a known amount of money per hour and a wrong spec costs an unknown and usually larger amount later.

Failure Two: The Fabricated Procedure

The second failure mode is the fabricated procedure. You ask the model for the steps to do something, a lockout sequence, a changeover, a calibration routine, a startup, a purge, and it produces a clean, numbered, professional-looking procedure. The steps are in a sensible order. They use the right verbs. They reference the right kinds of components. And they are wrong for your machine, because the model has never seen your machine. It assembled a generic procedure for the category of machine and presented it as if it were the procedure for your specific asset.

This one is more insidious than the invented spec, because a procedure is long and mostly plausible. The number 24 stood out as something to verify. A ten-step lockout sequence reads like authority. Nine of the steps may be perfectly fine. Step four, the one that says to bleed the accumulator before opening the guard, might be missing entirely, or placed after a step that exposes the operator to stored hydraulic energy. The danger is not that the whole procedure is garbage. The danger is that it is mostly right, which is exactly what makes the one wrong or missing step easy to miss, and on a lockout that one step is the difference between a routine changeover and someone in the hospital.

Here is a floor-grade example without a single piece of equipment named. A maintenance tech asks the model how to safely relieve pressure on a hydraulic clamp before servicing it. The model produces a confident sequence: isolate power, tag out, relieve pressure at the manifold, verify zero on the gauge, proceed. It sounds complete. What it does not know, because it has never seen this clamp, is that this particular unit has an accumulator that holds pressure after the pump is isolated, and the gauge the model told the tech to check reads line pressure, not accumulator pressure. The tech follows the confident steps, sees zero on the gauge, opens the clamp, and the stored energy in the accumulator drives the clamp closed. The model did not lie on purpose. It pattern-matched to a generic hydraulic system that does not have this accumulator, and it had no way to know yours does.

A fabricated procedure is dangerous precisely because it is mostly right. The nine correct steps lend false authority to the one wrong or missing step, and on a safety procedure that one step is the whole point.

How a working engineer catches it. A procedure that touches energy, motion, or anything that can hurt a person is verified against the machine's own documentation and validated by someone who knows the asset, before it is followed, every time, with no exception for time pressure. The standard work, the OEM (original equipment manufacturer, the company that built the machine) manual, the lockout-tagout (LOTO) procedure on file, and the tribal knowledge of the person who has serviced this asset are the sources. The model can draft a starting point that saves typing. It cannot be the final authority on a step that, if wrong, ends a shift in the ER. The rule of thumb: the model can write the draft, but a procedure that protects a body gets verified against the actual asset and signed by a human who knows it.

Failure Three: The Confidently Wrong Root Cause

The third failure mode is the most seductive, because it flatters your own thinking. You feed the model a defect description, a symptom, a downtime event, and you ask it for the root cause. It gives you a clean, logical, well-organized root-cause analysis. It might even format it as a fishbone (the cause-and-effect diagram that sorts potential causes into categories like machine, method, material, man, measurement, environment) or a 5-Whys (the technique of asking "why" repeatedly until you reach the underlying cause). It reads like a quality engineer wrote it. And the cause it lands on is frequently the most common cause for that symptom in general, not the actual cause in your process, which may be something specific to your tooling, your material lot, your shift pattern, or your machine's particular wear.

This is the confidently wrong root cause, and it is dangerous in a way the other two are not, because it does not just produce a bad output, it short-circuits the investigation. When a model hands a green crew a polished, confident root cause, the natural human move is to stop digging. Why run the fishbone yourself when the machine already drew one? Why pull the historian trace and the material certs when the answer is right there, formatted, footnoted-looking, and plausible? The model does not just risk being wrong. It risks convincing a thin, time-pressured crew to close an investigation on a fabricated conclusion, which means the real cause stays in the process and the defect comes back.

Work the example. A plant sees a spike in a dimensional defect on a machined part: a bore coming in oversized. The quality engineer asks the model for likely root causes. The model produces a beautiful answer: tool wear is the most common driver of oversized bores, recommend checking tool life and replacing the cutting tool, verify with a capability study. Every word of that is reasonable. Tool wear is the most common cause of oversized bores across all the machining the model ever read about. But in this plant, the bore went oversized the same week a new coolant concentration was dialed in, the coolant change shifted the thermal behavior of the part during machining, and the part was growing on the gauge after it cooled. The tool was fine. The crew, trusting the confident answer, changed three perfectly good tools, kept making oversized bores, scrapped another two days of production at roughly $4,200 a day, and only found the coolant link when a veteran finally ignored the AI and pulled the process change log. The model gave the textbook cause. The textbook cause was not the cause. And the confidence of the answer cost the plant two extra days of scrap and three tools because it stopped the real investigation before it started.

How a working engineer catches it. A root cause from a model is treated as a hypothesis to test, never as a conclusion to act on. The discipline that catches it is the discipline you already know from a real 8D: the cause has to be confirmed against evidence specific to your event before any corrective action is taken. Pull the historian trace for the window. Check what changed: material lot, process settings, tooling, shift, environment. Confirm the suspected cause can be turned off and the defect goes away, and turned back on and it returns. The model can help you assemble the evidence and suggest categories to check, which is genuine value. It cannot tell you which category is yours, because it has never measured your process. The rule: the model proposes the hypothesis; your evidence confirms the cause; corrective action follows the evidence, not the model.

Why all three share the same root

The invented spec, the fabricated procedure, and the confidently wrong root cause look like three different problems, but they are the same problem in three costumes. In every case the model substituted the general for the specific. It gave the population average torque for your specific joint, the generic procedure for your specific machine, the textbook cause for your specific event. That is the whole failure pattern, and once you see it you can predict where the model will fail before it does. It fails hardest exactly where your reality deviates from the textbook, and a manufacturing plant is a giant collection of deviations from the textbook. Your soft aluminum boss, your accumulator, your coolant change. The model knows the textbook. It does not know your plant.

Why the Floor Is Uniquely Exposed

Every industry that uses generative AI faces hallucinations. The manufacturing floor is uniquely exposed for three reasons that compound, and naming them is how you take the risk seriously without taking it personally.

The crew is thinner and greener than it used to be. This is the through-line of everything in this program. Roughly two million manufacturing workers need reskilling by 2026 against about half a million unfilled roles, and 85 percent of manufacturers say staffing shortages are already hurting product quality. The practical effect on hallucinations is direct: the people most likely to lean on an AI answer at 2 a.m. are the newest people, the ones who do not yet have the gut sense that 24 Nm is wrong for a soft boss. A thirty-year veteran would have squinted at that number and said "no chance, that'll crush it." Marco did not have thirty years. He had four seconds and a confident machine. The thinner the crew, the fewer people in the building who can catch the lie by instinct, and the more the verification has to be built into the process instead of living in someone's head.

The output is physical and often irreversible. In an office, a hallucinated fact in a draft email gets caught in review and deleted. On the floor, a hallucinated torque value gets torqued into forty parts and shipped. You cannot un-crack a boss. You cannot un-injure the tech who opened the clamp. The cost of a hallucination in most industries is embarrassment and a correction. The cost on the floor is scrap, containment, customer escapes, and sometimes a body. The stakes raise the bar on verification from "good practice" to "the job."

The plant runs on confident-sounding documents already. A floor is a place where authority comes in the form of a clean document: the work instruction, the spec, the SOP (standard operating procedure), the traveler (the paper or digital record that follows a part through the process). People are trained, correctly, to follow the controlled document. A hallucination arrives wearing the exact same uniform. It is formatted like a spec, written like an SOP, organized like a root-cause report. The cultural muscle that says "follow the document" is the muscle the hallucination exploits. That is why the defense cannot be "be more skeptical in general." It has to be specific: know which documents are controlled and traceable, and know that an AI output is never one of them until a human has tied it back to the controlled source.

The floor is uniquely exposed because the output is physical, the crew that catches the lie by instinct is thinning, and a hallucination arrives wearing the same clean uniform as a controlled document.

The Verification Habit That Actually Holds

Knowing the three failure modes is not enough. Under time pressure, knowledge evaporates and habits remain. The goal is a verification habit so simple it survives a 2 a.m. line-down with one tech short. The habit has three moves, and they map directly onto the three failure modes.

Move one: separate the draft from the fact. Decide, before you ask the model anything, which category your question falls into. Are you asking it to draft something (write a first pass of a work instruction, summarize a long log, reorganize notes), or are you asking it for a fact (a spec, a procedure step, a root cause)? Drafting is where the model is genuinely strong and the risk is low, because a human reads and edits the draft before it matters. Facts are where the model is dangerous, because a wrong fact can flow straight to the line. The single most useful mental move is to refuse to let a model be the source of a fact. Let it draft all day. Never let it be the last word on a number, a step, or a cause.

Move two: trace every load-bearing claim to a controlled source. A load-bearing claim is anything that, if wrong, ships a defect or hurts someone: the torque, the tolerance, the lockout step, the disposition, the cause. Each load-bearing claim gets traced back to the drawing, the released spec, the OEM manual, the historian trace, or the validated standard before it is used. If the source cannot be found, the claim does not get used, and the line holds. This is the same instinct a good auditor has and the same instinct the next lessons in this program build into a repeatable checklist. The model can point you toward where the source might live. It is never the source.

Move three: keep the human accountable, on the record. When an AI-touched output goes into a controlled record, a quality disposition, a work instruction, a corrective action, a named human reviews it and signs it, and that review is documented. This is the bridge to the next lesson, the cardinal rule that the customer audits you, not the vendor. "The model flagged it" is never an acceptable answer to an auditor or a customer, and "the model gave me the spec" is never an acceptable answer when the boss cracks. The human who signs owns the answer. That is not a burden the AI added. It is the job, performed with a faster drafting tool.

Notice what this habit does and does not do. It does not slow the plant down to a crawl, because most of what the model does well, drafting and summarizing, carries low risk and needs only a normal read. It concentrates the verification effort precisely on the load-bearing claims, the handful of numbers and steps and causes that can actually ship a defect. A disciplined verification on Marco's one torque value would have taken him the two minutes it takes to find a print or escalate, against a $38,000 containment. That is the entire economic case for the habit, and it is not close.

What the habit is not

The verification habit is not a reason to keep AI off the floor, and that distinction matters for your career. The plants that win in a thin-crew environment are the ones whose people use AI heavily and verify ruthlessly, not the ones who ban it out of fear or trust it out of laziness. A generative model that drafts a clean first pass of a work instruction, summarizes a forty-page maintenance log into the three events that matter, or reorganizes a tech's shorthand into a searchable record is a real multiplier for a crew that is too thin to do all of that by hand. The skill is not avoidance. The skill is knowing exactly where the line sits between "let it draft" and "verify before it ships," and never letting time pressure move that line.

Key Takeaways

  • A hallucination is not the model malfunctioning; it is a generative model doing exactly what it does, predicting plausible text, with no built-in tell that separates a true answer from an invented one. The smooth confidence is the danger, not a sign of correctness.
  • There are three failure modes on technical work: the invented spec (a plausible number that is wrong for your specific part), the fabricated procedure (a mostly-right sequence missing the one step that matters), and the confidently wrong root cause (the textbook cause that is not your event's actual cause).
  • All three are the same failure in different costumes: the model substitutes the general for the specific. It fails hardest exactly where your real part, machine, or process deviates from the textbook, and a plant is nothing but deviations from the textbook.
  • The invented spec cost a real plant about $38,000 when 24 Nm crushed a soft boss that needed 14 Nm; the confidently wrong root cause cost two extra days of scrap at roughly $4,200 a day plus three good tools because it stopped the real investigation before it started.
  • The floor is uniquely exposed because the output is physical and often irreversible, the thinning crew has fewer veterans who catch the lie by instinct, and a hallucination arrives wearing the same clean uniform as a controlled document.
  • The verification habit has three moves: separate the draft from the fact and never let the model be the source of a fact; trace every load-bearing claim to a controlled source before it is used; keep a named human accountable and on the record for any AI-touched controlled output.
  • A model is never the source of a spec, a procedure, or a root cause. It is a pointer toward where the real source lives and a drafting tool that saves typing. The drawing, the OEM manual, the released standard, and the historian are the sources.
  • Verification concentrates effort on the handful of load-bearing claims that can ship a defect, not on everything, so it does not slow the plant; two minutes of tracing one torque value would have prevented a $38,000 containment, which is the entire economic case for the habit.