Catching Hallucinations in Technical Output
The torque spec was 47 Newton-meters, and it was completely invented. A maintenance planner named Renee had asked an AI tool to draft a reassembly procedure for a gearbox after a bearing change, and the model delivered a tidy, professional document that specified torquing the housing bolts to 47 Nm in a star pattern. It looked exactly like every other procedure in the plant's library. It had the right structure, the right voice, the right confident specificity. The only problem was that nobody, anywhere, had ever specified 47 Nm for those bolts. The actual figure on the manufacturer's drawing was 62 Nm. A bolt torqued to 47 instead of 62 is under-clamped by nearly a quarter, and on a gearbox that runs hot and vibrates, an under-clamped housing bolt walks loose, the joint leaks, and eventually the bearing you just replaced fails again from contamination. Renee caught it only because she had a habit: before any AI-drafted spec reached a work order, she pulled the drawing and checked the number against print. That habit, a thirty-second cross-reference, stood between a clean-looking document and a repeat failure that would have cost a second four-hour outage at roughly 1,800 dollars an hour. This lesson is about that habit, made systematic, so it catches the fabrication every time and not just on the nights you happen to be careful.
What a Hallucination Actually Is, in Floor Terms
The word "hallucination" makes it sound like a glitch, something that happens when the AI breaks. It is the opposite. A hallucination is the model working exactly as designed. A generative model predicts the most plausible next piece of text given everything before it. It is a probability engine for what a competent answer usually looks like, not a lookup table for what is true. When it does not have your specific number, it does not stop and say "I do not have that." It produces the number that would most plausibly appear in a document like the one you asked for. For housing bolts of that size, something in the 40-to-65 Nm range is plausible, so it confidently writes 47, and it writes it in the same voice it uses for facts it actually has.
This is why hallucinations are so dangerous on technical work specifically. The model is fluent in the format of a torque spec, a tolerance, a procedure, and an 8D (the eight-discipline structured problem-solving report customers require after a defect escape). Fluency in the format is exactly what makes a fabricated value blend in. A wrong torque value does not arrive with a warning label. It arrives looking identical to a right one, embedded in a procedure that is otherwise correct, which is the worst possible camouflage. The error rate on the prose around it might be low; the error sits in the one number that determines whether the joint holds.
There is a useful distinction to hold onto. The model can be wrong in two different ways, and they fail differently. The first is a plausible fabrication: a number or procedure that does not exist in your records but sounds right, like the 47 Nm. The second is a confidently wrong synthesis: the model takes real information and combines it incorrectly, like correctly reading two tolerances from a drawing and then computing a stack-up that violates the assembly fit. The first type is caught by checking the value against the source. The second is caught by checking the reasoning against the physics. You need both checks, because a verification habit that only confirms quoted numbers will miss the synthesis error entirely.
A hallucination is not the model breaking. It is the model doing exactly what it was built to do: produce the most plausible text, whether or not it is true.
The Three Fabrications That Cost the Most
On a shop floor, generative AI fails in three specific shapes, and each one has a characteristic price tag. Knowing the shapes is the first half of catching them, because you verify differently depending on what kind of error you are hunting.
The invented specification
This is the 47 Nm case. A torque, a tolerance, a clearance, a temperature, a pressure, a cure time, a material grade: any hard number that should trace to a drawing or a standard, and does not. The invented spec is the most common and often the most expensive, because a single wrong dimension can scrap a lot or, worse, pass inspection and fail in the field. The verification is direct: every hard number must trace to a named source. If the model cannot tell you which drawing and revision the 47 Nm came from, the number does not exist until you find it on print yourself.
The fabricated procedure
This is the bore-reaming class of error, where the model produces a sequence of steps that is internally coherent but physically or logically wrong for your process. It might tell an operator to apply a coating before a surface prep step that the coating requires, or to torque in a sequence that warps the part, or to rework a feature in a way that cannot achieve the result. The fabricated procedure is dangerous because it reads like competence; the steps are specific and ordered, which signals expertise. The verification is to walk the procedure against the actual process with someone who runs it, and against the controlling work instruction, checking that each step is both real and in the right order.
The confidently wrong root cause
When you feed an AI the symptoms of a failure and ask for the cause, it will give you one, and it will sound authoritative. The trouble is that it reasons from the average of how failures like this are usually described, not from your machine's actual history. It will tell you the porosity is from gas entrapment when your historian shows the real driver was a die-temperature excursion the model never saw. The confidently wrong root cause is the most insidious of the three, because acting on it sends your corrective action in the wrong direction. You spend the kaizen fixing the wrong thing, the defect recurs, and now you have burned both the containment cost and the improvement effort. The verification is to demand that every claimed cause point to evidence in your data: the historian trace, the CMMS history, the actual measurements, not the general pattern.
Put a number on the stakes. A plant that runs AI-drafted root causes into its 8D process without verification will, sooner or later, ship a corrective action built on a fabricated cause to a customer. When the defect recurs because the real cause was never addressed, the customer escalation, the repeat containment, and the controlled-shipping status that follows a repeat escape commonly run past 50,000 dollars and put the supplier's quality rating at risk. The confidently wrong root cause is cheap to produce and ruinously expensive to act on. That asymmetry is the whole argument for verification.
The mixed fabrication, where two types hide together
The real floor is messier than three clean categories. The most expensive escapes often carry two fabrications at once, and the second hides behind the first. Picture an AI-drafted 8D for a recurring leak at a press station. The model states the root cause as a worn seal (a confidently wrong cause, because your historian actually shows a clamp-pressure drop), and in the corrective action it specifies replacing the seal and torquing the retainer to 35 Nm (an invented spec, because the drawing calls for 28 Nm). A reviewer who only checks causes might catch the seal error and feel satisfied, never noticing the fabricated torque riding along inside the corrective action. A reviewer who only checks numbers might confirm the torque against the wrong drawing and miss that the whole cause is wrong. The mixed fabrication is why verification cannot be a single glance. It has to check the cause against the data and every number against the drawing, as two separate passes, because the two errors fail in different places and a single sweep misses one of them.
The defense against the mixed fabrication is to verify by class, not by document. Walk the entire 8D once asking only "does every claimed cause match real evidence in the historian and CMMS," and walk it a second time asking only "does every hard number trace to the current controlled drawing." Two narrow passes catch more than one wide one, because each pass has a single question and cannot be distracted by the part of the document that looks fine. On a recurring leak that has already cost two containments, the few extra minutes of a second pass are trivial against the third containment you are trying to prevent.
The Cross-Reference Discipline: Drawing, Standard, Historian
Renee's thirty-second habit, generalized, is the single most valuable verification skill in this entire program. The discipline is to cross-reference every AI-touched technical claim against the authoritative source for that kind of claim, before it reaches anything that moves or ships. There are three authoritative sources, and each owns a different class of claim.
The drawing owns dimensions and specs. Every tolerance, torque, clearance, finish, and material callout traces to a controlled drawing at a specific revision. When the AI states a spec, your job is not to assess whether it sounds right. Sounding right is exactly the trap; the 47 Nm sounded right. Your job is to open the drawing and read the number off print. If the model gave you a drawing and revision, confirm it is the current controlled revision, because a model can cite a real but superseded drawing and quote a value that was correct two revisions ago and is now wrong. The drawing is the authority. The model's number is a claim about the drawing, nothing more.
The standard owns procedures and acceptance criteria. When the AI drafts a procedure or states an acceptance limit, the authority is the controlling work instruction, the control plan, and the applicable standard, whether that is IATF 16949 in automotive (the standard governing how suppliers control processes and defects) or AS9100 in aerospace. You check the AI's steps against the documented method, not against your memory of it, because under time pressure your memory drifts toward what the confident document just told you. This is a real cognitive trap: a fluent wrong procedure can overwrite your own recollection of the right one. Reading the controlled document, not recalling it, is the defense.
The historian and the CMMS own what actually happened. When the AI claims a cause, a trend, or a failure pattern, the authority is your plant's own data: the historian tags (the time-stamped sensor and process values your control system logs), the CMMS maintenance history, and the actual inspection measurements. A claimed root cause that cannot be matched to a real excursion in the historian or a real pattern in the maintenance history is a story, not a finding. The historian does not flatter the model. If the temperature trace is flat at the moment the model blames a temperature excursion, the model is wrong, and the trace will tell you so in seconds.
The economics of this habit are lopsided in your favor. The cross-reference for a single spec takes about thirty seconds to two minutes. The cost of one escaped fabrication ranges from a scrapped lot in the low thousands to a customer containment in the tens of thousands. Even if only one AI-drafted claim in two hundred is a costly fabrication, verifying all two hundred at two minutes each costs under seven hours of attention to prevent a five-figure loss. No quality control you can buy has a better return than that.
Forcing the Model to Make Itself Checkable
You can make verification dramatically faster by changing how you ask, so that the model exposes its own fabrications instead of hiding them in fluent prose. The principle is to force the model to separate what it knows from what it is guessing, and to attach a checkable source to every hard claim.
The most powerful single instruction is to require a citation for every specification: "For every torque, tolerance, or dimension, state the drawing number and revision it comes from. If you do not have the source, write SOURCE NOT PROVIDED instead of a number." This works because a model that has been told to cite either gives you a real, checkable pointer or visibly flags its own gap. The 47 Nm would have arrived as either "47 Nm per drawing 88102 rev B," which Renee could check in thirty seconds, or as "SOURCE NOT PROVIDED," which tells her instantly that the number is a guess. Either way the fabrication is no longer camouflaged. The fluent prose that hid it has been replaced by a claim with a visible source field, and an empty source field is a loud alarm.
A second technique is to ask the model to flag its own confidence and separate fact from inference. Instruct it: "Mark each statement as VERIFIED if it comes from a source I provided, or INFERRED if it is your general knowledge." A model forced to label its inferences will mark the invented torque as INFERRED, because it has to, and an INFERRED hard spec is a stop sign. This does not make the model honest in some deep sense; it makes the structure of the answer reveal where the model is operating without grounding.
A third, and underused, technique is to give the model the source and ask it only to extract, never to supply. Instead of "what is the housing bolt torque," paste the relevant page of the drawing or the work instruction and ask "what does this document say the housing bolt torque is." Now the model's job is reading, not generating, and reading is far less prone to fabrication than generating from scratch. When you must rely on a number, narrow the model's task from inventing to quoting, and verify the quote against the page you pasted. This is the practical face of grounding: the tighter you bind the model to a document you control, the less room it has to hallucinate.
None of these techniques replace the human cross-reference; they make it cheaper. A citation you still confirm against the controlled drawing. An INFERRED label still tells you to go check. A quote still gets matched to the source page. The instructions shrink the haystack so the needle is easy to find. They do not let you skip looking.
Building the Habit Into the Workflow So It Survives a Bad Night
A verification habit that depends on a careful person being careful will fail the night that person is tired, short-staffed, and three work orders behind. The whole point of this program is that the plant is thinner and greener than it used to be, which means you cannot rely on a Renee being on shift. The habit has to live in the workflow, not in a person's memory, so it runs even when nobody is feeling diligent.
The mechanism is a verification gate: a required step between an AI draft and the system of record where the draft becomes real. For a spec, the gate is a field that cannot be left blank, forcing the planner to enter the drawing and revision the number was confirmed against. For a procedure, the gate is a sign-off from someone who runs the process. For a root cause in an 8D, the gate is a required evidence reference for each claimed cause, a historian tag or a measurement, with no cause accepted that has no evidence. The gate turns the skeptic's instinct into a control that the workflow enforces, so the verification happens because the system will not advance without it, not because someone remembered.
Make the gate fast or it will be bypassed. A verification step that takes fifteen minutes per spec will quietly be skipped at 2 a.m., and a control that gets skipped is worse than no control because it creates a false record of diligence. The combination that holds is the citation-forcing prompt feeding a thirty-second cross-reference feeding a one-field gate. The prompt makes the number checkable, the cross-reference confirms it, and the gate records that it happened. Total cost: under two minutes for the kind of claim that, fabricated and escaped, costs five figures.
One more discipline closes the loop. Every time a fabrication gets caught, log it, briefly: what the model invented, where, and how it was caught. Over a few months this log becomes a map of where your AI tools fabricate most, which is almost never random. You will find the model hallucinates torque specs more than tolerances, or invents causes for one failure mode more than others, and that map tells you exactly where to tighten the prompts and harden the gates. The caught fabrication is not just a near miss to feel good about. It is data about your AI's specific weaknesses, and a thinning plant that mines that data gets steadily safer with AI instead of steadily luckier.
Key Takeaways
- A hallucination is not a glitch; it is the model working as designed, producing the most plausible text rather than the true value. That is why a fabricated 47 Nm torque blends in perfectly with correct prose and is so dangerous on technical work.
- The two failure types fail differently: a plausible fabrication (a number that does not exist but sounds right) is caught by checking the value against the source; a confidently wrong synthesis (real numbers combined incorrectly) is caught by checking the reasoning against the physics. You need both checks.
- Three fabrications cost the most: the invented specification (a wrong torque or tolerance), the fabricated procedure (coherent but physically wrong steps), and the confidently wrong root cause, which is the most insidious because acting on it sends the corrective action in the wrong direction and routinely costs past 50,000 dollars when the defect recurs.
- Cross-reference every AI-touched claim against its authoritative source: the drawing owns dimensions and specs (read the number off print, confirm the revision is current), the standard and work instruction own procedures, and the historian and CMMS own what actually happened.
- Sounding right is the trap, not the test. The 47 Nm sounded right. Verification means opening the controlled drawing and reading the number, never assessing plausibility, because a fluent wrong procedure can overwrite your own memory of the right one.
- Force the model to make itself checkable: require a drawing-and-revision citation for every spec or a SOURCE NOT PROVIDED flag, ask it to label statements VERIFIED or INFERRED, and prefer giving it the document and asking it to quote rather than supply. These shrink the haystack; they do not replace the human check.
- Build verification into the workflow as a fast gate, not a careful person's habit, because the plant is too thin to rely on diligence at 2 a.m. A citation-forcing prompt plus a thirty-second cross-reference plus a one-field gate costs under two minutes against a five-figure escaped fabrication.
- Log every caught fabrication. Over months the log maps where your AI tools fabricate most, telling you exactly where to tighten prompts and harden gates so the plant gets steadily safer with AI instead of steadily luckier.
Skill.re