Building Verification Checklists for Floor AI
It is 6:40 in the morning and Maria, a process engineer on second shift turned first-shift fill-in, is staring at a work instruction her AI assistant drafted overnight for a torque-down operation on a bracket assembly. The draft looks beautiful. It is formatted cleanly, it cites the operation number, it names the tool, and it confidently specifies 38 newton-meters of torque on the M8 fasteners. The problem is that the print on her desk, the controlled engineering drawing the customer audits against, calls for 25 newton-meters. Nobody told the model the real number. It produced 38 because 38 is a plausible-sounding torque for an M8 bolt, and plausible is exactly what a language model is built to produce. If Maria signs that instruction and walks away, the line will over-torque a few thousand brackets, strip threads, create a field-failure risk, and hand the customer a containment that costs more than a month of her department's budget. The only thing standing between that draft and that disaster is whether Maria has a repeatable way to catch the invented number before it reaches the floor. That repeatable way is a verification checklist, and building one is the single highest-leverage habit a plant can install around AI.
Why a Checklist and Not Just a Careful Engineer
The obvious objection comes up in every kaizen on this topic: "We hire good engineers. They are careful. Why do we need a checklist to tell a smart person to read the drawing?" The answer is the same reason a surgeon who has done ten thousand procedures still runs a pre-incision checklist, and the same reason a pilot with twenty thousand hours still reads the before-takeoff list out loud. A checklist is not an insult to expertise. It is a defense against the specific way expertise fails, which is that experienced people skip steps when they are busy, tired, confident, or interrupted, and they skip them precisely on the items that almost never go wrong, which are the items that bite hardest when they finally do.
AI raises the stakes on this human pattern in a way that ordinary engineering work does not. When a junior engineer drafts a work instruction by hand, the errors tend to look like errors. A blank field, a missing step, a number that is obviously a guess, a sentence that trails off. Those wrong drafts trigger the reviewer's suspicion because they look unfinished. An AI draft does the opposite. It is complete, fluent, formatted, and confident. The fabricated torque spec sits in a clean table next to four correct specs, in identical formatting, with no visual tell that one of them was invented and four were retrieved. The fluency is the danger. Research and field reports through 2026 keep landing on the same lesson: the job did not get easier when AI started producing the draft, it shifted. The work moved from "produce the document" to "verify the document against the drawing, the standard, and the historian." A checklist is how you make that verification step survive a bad morning.
Consider the floor math. A skipped verification on one over-torque spec, applied to a 3,000-part run before anyone notices, can generate scrap, rework, a customer containment, and a corrective-action report (an 8D, the eight-discipline structured problem-solving report customers require) that consumes two engineers for a week. Call it a conservative 60,000 dollars all-in when you add the parts, the labor, the expedited reinspection, and the customer's lost confidence. The checklist that would have caught it takes ninety seconds to run. That is the trade the whole lesson turns on: ninety seconds of structured doubt against a five-figure escape. Run that trade a few hundred times a year across every shift and the checklist is not overhead, it is one of the cheapest insurance policies the plant owns.
A verification checklist is not a tax on good engineers. It is the thing that makes a good engineer reliable on the morning everything is on fire.
The Three Failure Modes a Checklist Must Catch
You cannot build a checklist that catches everything, and trying to will produce a list so long nobody runs it. The art is targeting the specific ways AI output goes wrong on the floor. There are three, and every line on your checklist should map to one of them.
The invented spec
The invented spec is the failure Maria almost shipped. The model produces a number, a setting, a material callout, or a tolerance that sounds right and is wrong because it was generated from the pattern of plausible values rather than retrieved from the controlled source. Torque values, oven temperatures, cure times, gap and flush tolerances, gas flow rates, fastener grades, and inspection sample sizes are all classic targets. The tell is that the value is specific and confident and there is no traceable source for it inside the prompt. The checklist defense is a hard rule: every number that drives a physical action gets traced back to the drawing, the spec, the work order, or the historian tag before the document is released. No source, no signature.
The fabricated procedure
The fabricated procedure is a sequence of steps that reads like a real procedure but does not match how the process actually runs. The model fills gaps with industry-generic steps, invents a fixture that does not exist on this line, references a fault-reset sequence from a different control platform, or orders the steps in a way that is logical in the abstract and wrong for this machine. A green operator following a fabricated procedure does not know which step is invented, so the procedure has to be verified by someone who has actually run the operation. The checklist defense is a walk-through requirement: a person who has performed the task confirms each step matches reality, with special attention to any step the model added that the source material did not contain.
The confidently wrong root cause
The confidently wrong root cause is the most seductive failure because it shows up in the analytical work where AI feels most helpful. Asked to assemble a fishbone or a 5-Whys for a defect, the model produces a clean, well-reasoned chain that points at a cause, and the chain is internally coherent and unsupported by the plant's actual data. It blames a worn die when the historian shows the die was changed two days before the defects started. It names a humidity excursion that the environmental log does not contain. The checklist defense is grounding: every causal claim must be checked against the actual evidence, the traveler, the historian trace, the maintenance history, and the measurement data, and any claim that cannot be tied to a record gets demoted from "cause" to "hypothesis to test." The model is allowed to suggest; it is not allowed to conclude.
Notice that all three failures share a root: the model fills a gap with something plausible instead of leaving the gap visible. Your checklist, at its core, is a machine for re-exposing the gaps the model papered over. Every line is a question that forces a comparison between the confident output and an authoritative source.
The Anatomy of a Floor-Grade Checklist
A checklist that works on the floor has a shape, and most failed checklists fail because they ignore it. Borrowing from aviation and surgical practice, where checklists have been refined under real pressure for decades, a floor-grade verification checklist has six properties.
It is short. Five to nine items is the working range. The moment a checklist crosses into the teens, people start pencil-whipping it, signing the whole thing without reading, which is worse than no checklist because it manufactures false assurance and a paper trail that lies to the auditor. If you have twenty things to verify, you have several checklists for several document types, not one monster list.
Every item is a yes/no the verifier can actually answer. "Is the output high quality?" is not a checklist item; it is a wish. "Does every torque, temperature, and tolerance in this document trace to the controlled drawing or spec?" is a checklist item, because a person can put a yes or a no next to it and be held to that answer.
Every item names its authoritative source. A verification step is only as good as what it checks against. The line should not say "verify the torque is correct." It should say "verify each torque against drawing rev currently released in the document control system." Correct against what is the whole game. On the floor the authoritative source is almost never the engineer's memory; it is the controlled drawing, the customer spec, the work order, the historian tag, the CMMS (computerized maintenance management system, the database of work orders and equipment history) record, or the validated process parameter sheet.
It is killer-item ordered. Put the items that prevent a customer escape or a safety event first, while attention is freshest. The torque trace and the safety-step check go above the formatting check. If the verifier is interrupted halfway through, you want the half that got done to be the half that matters.
It is read-and-respond, not recall. The verifier reads the item and acts, rather than trying to remember a list. This is why the checklist lives next to the work, in the template, in the MES (manufacturing execution system, the software that manages and records production on the floor) record, on a laminated card at the station, not in a binder nobody opens.
It ends in accountable sign-off. The last line is a named human taking responsibility: verified by, date, time, and an affirmation that the items above are true. This is the line the customer audits. "The model drafted it" is never the answer to an auditor; "I, the named verifier, confirmed each spec against the released drawing on this date" is. Accountability for an AI-touched quality or maintenance decision stays with the plant and the person who signs the record, and the sign-off line is where that accountability becomes concrete.
A Worked Checklist for an AI-Drafted Work Instruction
Abstractions do not survive contact with a 6:40 morning, so here is a concrete, ready-to-run checklist for the most common AI deliverable on the floor, the drafted or revised work instruction. Maria would run this on her bracket torque-down draft, and it would catch the 38 versus 25 newton-meter error on line one.
1. Spec trace. Does every number that drives a physical action (torque, temperature, time, pressure, flow, speed, dimension, tolerance) trace to a named, currently released source: drawing rev, customer spec section, or validated parameter sheet? List the source next to each. If any number has no source, stop. This is the line that catches the invented spec.
2. Revision currency. Is the source you traced to the current released revision, not a superseded one the model may have learned from? A correct number against last year's rev is still a defect. Confirm the rev in document control.
3. Step reality. Has a person who actually runs this operation confirmed each step matches the real process, including fixtures, tools, and machine-specific reset or setup sequences? Flag any step the model added that the source did not contain.
4. Safety and EHS steps. Are all required lockout, PPE, guarding, and ergonomic steps present and correct, and did the model not silently drop one to make the procedure read more smoothly? Safety steps are killer items; they go high on the list.
5. Units and conversions. Are all units explicit and consistent with the controlled source, with no silent conversion errors (newton-meters versus foot-pounds, Celsius versus Fahrenheit, millimeters versus inches)? Unit drift is a quiet, classic AI error.
6. Completeness. Does the instruction cover the full operation with no gap the model skipped or summarized away, including inspection and acceptance criteria?
7. Accountable sign-off. Verified by (name), date, time. By signing, the verifier confirms items 1 through 6 are true and takes responsibility for the released document.
Seven items. Under two minutes for a routine instruction once the verifier is practiced. The first line alone, run honestly, would have turned Maria's beautiful, confident, wrong draft back before it ever reached the torque gun. And the sign-off line is what makes the whole thing audit-grade: when the customer's quality auditor pulls this work instruction in an IATF 16949 audit (the automotive quality management standard customers enforce), the verification record shows a named human checked the controlled source, which is exactly what the standard expects and exactly what "the model flagged it" fails to provide.
Checklists for the Other Floor Deliverables
The work instruction is one of several AI deliverables that hit the floor, and each needs its own short, targeted list. The principle is identical, the killer items differ. A few worked examples show the pattern.
The AI-drafted 8D or root-cause analysis
The killer item here is grounding, because the failure mode is the confidently wrong root cause. The checklist asks: is every causal claim tied to a specific record (traveler, historian trace, CMMS history, measurement data), and has any claim without a record been demoted to a hypothesis to test? It asks whether the timeline in the analysis matches the historian and maintenance logs, because a model will happily blame a worn die that the CMMS shows was replaced before the defects began. It asks whether the corrective actions actually address the verified cause rather than a plausible-sounding cause. A worked example: a defect-escape investigation on a stamping line where the AI draft blamed material lot variation. The grounding line forced a check against the historian, which showed the press tonnage had drifted 8 percent over the same window. The real cause was a tooling wear issue the model never saw because nobody fed it the tonnage trace. Grounding turned a wrong, expensive corrective action into the right one and saved a recurrence of the escape.
The AI-summarized vision-system or inspection result
The killer item is the honest reading of false rejects versus escapes. The checklist asks whether the summary distinguishes false rejects (good parts the system called bad) from real catches, and whether it reports the false-reject rate alongside the catch rate rather than burying it. It asks whether any disposition decision was left to a named human rather than auto-accepted from the model, because an operator who has been burned by a false alarm will eventually disable the green light, and a summary that hides the false-reject cost makes that disabling more likely, not less. The false-reject rate is dollars, not an abstraction: a vision system that scraps good parts at even a few percent on a high-volume line can quietly cost more than the escapes it prevents.
The AI-drafted maintenance work order from a predictive alert
The killer item is the action and the source signal. The checklist asks whether the work order names the specific asset and the specific signal that triggered it (the historian tag, the vibration band, the temperature trend) rather than a generic "model predicts failure." It asks whether a reliability person confirmed the recommended action matches the failure mode the signal indicates, because a prediction is not a diagnosis. It asks whether the work order is prioritized against real risk rather than dumped into the queue at the same priority as everything else, since alert fatigue is how a green crew learns to ignore the one alert that was real.
Across all of these, the pattern holds: identify the deliverable's specific failure mode, write five to nine yes/no items that each name an authoritative source, order them killer-first, and end in an accountable signature. The checklist is the verification step made repeatable, portable across shifts, and visible to the auditor.
Making the Checklist Actually Get Run
A perfect checklist that lives in a forgotten folder catches nothing. The hardest part of this work is not writing the list, it is making it survive contact with three shifts, a staffing shortage, and the end-of-month push. Several practices, drawn from plants that made it stick, separate a checklist that works from a checklist that becomes a compliance fiction.
Build it into the workflow, not alongside it. The checklist should be a required step in the template or the MES record, not a separate form a busy verifier can route around. If the AI-drafting tool or the document template will not let you release the document until the verification items are answered, the checklist runs every time by default. The goal is to make the right path the path of least resistance.
Keep it honest by making yes mean yes. Pencil-whipping is the death of checklists. The countermeasure is partly cultural and partly structural: keep the list short so reading it is faster than faking it, require the verifier to record the source they checked against rather than just ticking a box, and audit the audits. A periodic spot-check where a lead re-verifies a sample of signed-off documents tells you fast whether the checklist is being run or rubber-stamped. When you find a pencil-whipped sign-off, treat it as the serious event it is, because a false verification record is worse than no record.
Tune it from the misses. A checklist is a living document. Every time an AI error gets past the floor and becomes a defect or a near miss, ask whether a checklist item would have caught it, and if not, add the item, then prune something low-value to keep the list short. This is the same continuous-improvement loop the plant already runs on its processes, pointed at its own verification. Over a year, the checklist evolves from a generic template into a sharp instrument shaped by your plant's actual failure modes.
Train the why, not just the what. A verifier who understands that the model produces plausible numbers rather than retrieved numbers will catch errors the checklist did not anticipate. A verifier who just ticks boxes will catch only what is on the list. Spend the ten minutes in onboarding explaining the three failure modes and showing a real fabricated spec, because structured training is what makes the habit transfer. Plants with structured AI training programs see substantially higher real adoption than plants that leave it to self-directed learning, and a checklist is exactly the kind of structured practice that transfers across a thinning, greening crew.
Make the sign-off mean accountability, not blame. The verifier is signing that they checked, not promising the document is perfect forever. If a verified document later has a problem the checklist did not cover, the fix is to improve the checklist, not to punish the verifier who ran it honestly. Punishing honest verification teaches people to stop verifying. The accountability you want is the accountability of "I ran the check and recorded what I found," which is exactly the accountability an auditor respects and a customer trusts.
The thinning crew is the reason all of this matters now. With roughly 2 million manufacturing workers needing reskilling against about 500,000 unfilled roles, and 85 percent of manufacturers reporting that staffing shortages are hurting product quality, the plant cannot rely on a deep bench of experienced engineers to catch every AI error by feel. The checklist is how a thinner, greener crew inherits the catching instinct of the experts who are retiring. It is institutional skepticism, written down, so it does not walk out the door in November when Dave the inspector does.
Key Takeaways
- AI output fails in three specific ways on the floor: the invented spec, the fabricated procedure, and the confidently wrong root cause. Every checklist item should target one of them.
- The danger of AI drafts is their fluency. A fabricated torque spec sits in clean formatting next to four correct ones with no visual tell, so suspicion must be structured, not left to instinct.
- A floor-grade checklist is short (five to nine items), made of yes/no questions, each naming an authoritative source, ordered killer-item first, read-and-respond, and ending in accountable sign-off.
- Correct against what is the whole game. Every number that drives a physical action must trace to the currently released drawing, spec, work order, historian tag, or validated parameter sheet before release.
- The trade is ninety seconds of structured doubt against a five-figure escape: a single skipped verification on an over-torque spec across a 3,000-part run can cost on the order of 60,000 dollars all-in.
- Each deliverable needs its own list with its own killer item: grounding for an 8D, honest false-reject reporting for a vision summary, named signal and confirmed action for a predictive work order.
- A checklist only counts if it gets run. Build it into the workflow, keep yes meaning yes by recording the source checked, tune it from every miss, and make sign-off mean accountability without blame.
- Accountability stays with the plant and the human who signs. "The model drafted it" never answers an auditor; "I verified each spec against the released drawing on this date" does, and the sign-off line is where that becomes audit-grade.
Skill.re