โ†
AI for Manufacturing
Visionary ยท M11 ยท lesson 11 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Plant AI Pilots to Evidence
๐Ÿ“–
now learning

Running Plant AI Pilots to Evidence

15 min

The Tuesday morning review in the corporate operations room had a familiar smell to it, the smell of a project quietly dying. On the screen was a slide from the Ohio plant: a machine-vision pilot on the wheel-hub line that, eight months earlier, had been the star of the quarterly town hall. The vendor demo had been beautiful. The camera caught a hairline crack the day shift would have missed, the plant manager clapped, and the VP of Operations told the board the plant was "going AI." Now the slide said the pilot was "on hold pending resource availability," which everyone in the room understood as the corporate phrase for dead. Nobody could say what the false-reject rate had been. Nobody could point to a single defect escape it had prevented in a number a customer would accept. There had been no control line to compare against, no agreed definition of success written down before the camera went live, and no owner once the vendor's two free engineers went home. It had produced a great demo and zero evidence. Down the hall, a second plant, the one in Tennessee, had run a quieter pilot on a bearing-press cell that nobody clapped for, and that one was about to be scaled to four more lines, because it had done the unglamorous thing the Ohio pilot never did: it had turned a promising demo into evidence strong enough to defend in front of a plant manager, a quality auditor, and a finance director who trusts nothing he cannot tie to a dollar. This lesson is about the difference between those two pilots. The difference is not the technology. It is the discipline.

Why Most Plant AI Pilots Die in the Demo

A demo and a pilot are not the same thing, and conflating them is the original sin of plant AI. A demo is a performance: a vendor brings their best camera, their cleanest lighting, a tray of pre-selected parts, and an engineer who knows exactly which knob to turn when the model wobbles. A demo is designed to make you feel a thing. A pilot is an experiment: it runs on your line, under your lighting that changes between shifts, on your parts that come in wet, with your operators who have been burned before, and it is designed to produce a number you can defend. The Ohio pilot died because it was a demo that someone scheduled for eight months and called a pilot.

The pattern repeats across the industry because the incentives push toward demos. The VP wants something to show the board after touring a competitor's smart factory. The vendor wants a logo and a reference, so the vendor optimizes for the wow moment, not the durable result. The plant is short three maintenance techs and cannot spare anyone to babysit a science project, so the pilot has no real owner. And nobody wrote down, before the camera went live, what number would count as success and what number would count as failure. Without that line drawn in advance, every result becomes negotiable. A false-reject rate of 9% gets described as "early days," a missed escape gets described as "an edge case," and the whole thing drifts until the budget cycle quietly absorbs it.

There is a hard cost to this churn that rarely makes it onto a slide. A serious vision or predictive pilot consumes real money before it ever proves anything: figure roughly 30,000 to 80,000 dollars in vendor time, integration labor, sensors or cameras, and the engineering hours pulled off other work, plus the opportunity cost of the line time. A plant that runs four pilots a year and scales none of them has spent a quarter of a million dollars to produce four slides that say "on hold." That is not an AI problem. It is a pilot-discipline problem, and it is the single most expensive habit in plant AI.

The talent cliff makes this worse, not better. With roughly 2 million manufacturing workers needing reskilling by 2026 against about 500,000 unfilled roles, the plant cannot afford to burn its scarce engineering attention on experiments that were never designed to conclude. Every hour an over-stretched process engineer spends nursing a directionless pilot is an hour not spent capturing the knowledge of an inspector who retires in November. Pilot discipline is, in the end, a way of respecting a thin crew's time.

A demo is built to be applauded. A pilot is built to be defended. If you cannot say in advance what number would make you walk away, you are running a demo and calling it a pilot.

Setting the Evidence Bar Before You Switch It On

The first act of pilot discipline happens before any hardware arrives. You write the success criteria down, you get the people who will later judge the pilot to sign the same page, and you do it while everyone is still calm and nobody is defending a number they already spent money on. This document is short, one page, and it is the most valuable artifact the pilot produces even if the pilot fails, because it is what turns a result into evidence instead of an opinion.

A defensible evidence bar names four things. The baseline: what the line does today, measured before the AI touches it. If you are piloting machine vision (a camera plus a model that grades parts at line speed) on a cosmetic-defect line, your baseline is the current escape rate, the current scrap and rework, and the current manual inspection time, each as a real number with a date on it. You cannot prove improvement against a baseline you never measured, and "it feels better" is not a baseline a finance director accepts.

The target: the specific, numeric result that would justify scaling. Not "improve quality," but "reduce defect escapes on this line by at least 40% while holding false-reject rate at or below 2%." The false-reject rate (the share of good parts the system wrongly rejects) belongs in the target as a hard ceiling, because a vision system that catches every defect by rejecting a fifth of good parts is a system the operators will quietly disable. False rejects are real money: at 30 seconds of operator handling per false reject and a few hundred false rejects a shift, you can burn an inspector's entire day chasing parts that were fine, which is exactly the labor you were trying to save.

The kill criteria: the number that means stop. This is the part everyone skips and the part that saves the most money. Write it as plainly as the target: "If after six weeks the false-reject rate cannot be held below 4% without dropping escape detection below baseline, we end the pilot." Kill criteria are not pessimism. They are how you free the next quarter's budget and engineering hours to spend on a use case that can actually work. The Ohio plant had no kill criteria, which is precisely why its pilot could not die cleanly and instead haunted the budget for eight months.

The decision date and the decision owner: who decides, by when, with what authority. A pilot without a decision date drifts forever. A pilot without a named owner who can say "scale it" or "stop it" becomes a committee artifact that nobody can end. The decision owner is usually the plant manager or the operations director, not the vendor and not the data scientist, because the person accountable for the line's quality and uptime is the person who must own the call.

The Control Line and Honest Measurement

The reason the Tennessee bearing-press pilot produced evidence and the Ohio vision pilot produced a feeling comes down to one word: control. The Tennessee team ran the predictive-maintenance model on one press cell and kept an identical second cell on the old reactive schedule, then compared the two over the same twelve weeks under the same production load. When the model flagged a spindle bearing trending toward failure and the work order landed before the breakdown, they could point at the control cell, which threw the same fault three weeks later and stopped the line for six hours, and say in plain numbers what the AI had bought.

A control line is the difference between correlation and evidence. Plants change a dozen things a quarter: a new lot of material, a different operator, warmer weather, a tweaked recipe. If you run an AI pilot on a line and yield goes up, you do not actually know the AI did it unless you have something to compare against that did not get the AI. The control does not have to be a whole separate line. It can be a matched cell, alternating shifts, or a holdout set of parts the model scores but does not act on, so you can check the model's call against what the inspector actually found. The principle is the same: hold everything constant except the AI, so the result points at one cause.

Honest measurement also means measuring the things that embarrass the vendor, not just the things that flatter it. For a vision pilot, that means tracking the false-reject rate as carefully as the catch rate, and tracking both by shift and by lighting condition, because a model that looks great on the day shift can fall apart on the night shift when the overhead lights are different and the parts come off a colder press. A confusion matrix, which is just a simple two-by-two count of caught defects, missed defects, false rejects, and correctly-passed parts, is the floor-level tool here. It fits on an index card and it tells you in four numbers whether the system is helping or quietly costing you.

A worked example. Suppose the cosmetic-defect line runs 5,000 parts a shift and historically escapes 25 defects per shift, each escape carrying an expected cost of about 200 dollars when you blend the rare containment against the common minor return. That is 5,000 dollars a shift of escape risk. The vision pilot, measured honestly over six weeks against a control, catches 80% of those escapes, removing 4,000 dollars a shift of risk. But it also false-rejects 1.5% of good parts, which is 75 parts a shift, each costing 30 seconds of operator handling and re-inspection. At a loaded labor rate, 75 false rejects might cost 40 dollars a shift in handling. Net, the pilot removes about 3,960 dollars of cost a shift. That is evidence. Now suppose the false-reject rate had been 12% instead: 600 parts a shift, enough to consume an inspector and slow the line, and the operators would have switched it off in a week. Same camera, same model, completely different verdict, and only honest measurement of the unflattering number tells you which world you are in.

Running a Pilot in a Brownfield Plant

Most pilots do not fail on the model. They fail on the plumbing, because the plant is brownfield: a 1990s PLC (programmable logic controller, the industrial computer that actually runs the machine), a historian (the database that logs sensor tags over time) that nobody has queried in years, and an MES (manufacturing execution system, the software that tracks what is being made on which line) that does not talk cleanly to anything. Greenfield plants deploy AI 40 to 60% faster than brownfield plants for exactly this reason, and pretending your brownfield line is greenfield is how a pilot's timeline doubles and its credibility evaporates.

The pilot plan has to budget for the plumbing as a first-class line item, not an afterthought. Before the model, you need to answer: can we actually get the data off this machine, at the rate the model needs, without touching the control loop? For a predictive-maintenance pilot, that often means adding a vibration or temperature sensor on the OT (operational technology, the systems that run physical equipment) side and getting its readings to the model without bridging the OT network straight onto the IT network, because that boundary is a hard security constraint, not a suggestion. Recall that 78% of OT networks lack centralized monitoring, which means on most plant floors you cannot even fully see the network you would be attaching a model to. A pilot that quietly punches a hole in the OT boundary to get its data is a pilot that should be killed on a security basis alone, no matter how good its numbers look.

Keep the pilot advisory. In a brownfield pilot, the AI flags and recommends; a human acts. The vision system lights a screen and an operator makes the disposition. The predictive model writes a draft work order and a maintenance planner approves it. The model does not reach into the PLC and stop the press, and it does not auto-reject parts into a bin with no human in the loop. Advisory mode is not timidity. It is what lets you run a real experiment on a real line without the model being able to cause the very downtime or scrap you are trying to prevent, and it is what keeps the human accountable for every decision the pilot touches, which is exactly what a customer auditor will want to see.

Plan for drift from day one. A vision model degrades as lighting shifts between shifts, as the camera angle creeps from a forklift bump, and as the material changes lot to lot. A pilot that does not include a drift check, a periodic re-measure of the confusion matrix against fresh labeled parts, will look great in week one and quietly rot by week six, and you will not know which until an escape gets out. Building the drift check into the pilot is also building the evidence that the system can be maintained after scale, which is a question the decision owner will absolutely ask.

The Operator, the Green Light, and the Save Log

A pilot that the operators do not trust is dead even if its numbers are perfect, because the operators are the ones who decide, shift after shift, whether the green light gets believed or ignored. The fastest way to kill operator trust is a false alarm that wastes their time and makes them look bad in front of a supervisor. An operator who gets burned once by a vision system that rejected a perfectly good part, then had to defend the scrap number to the line lead, will learn to override the system, and once that habit forms the pilot is measuring a system nobody actually uses.

So operator trust is a design requirement of the pilot, not a soft afterthought. That means involving the operators before the camera goes live, letting them see the confusion matrix so the system is not a black box, giving them a fast and blameless way to flag a false reject, and holding the false-reject rate to a level that respects their time. It also means being honest with them that the AI is there to make a thinner, greener crew safe and effective, not to replace them or to grade them. The 85% of manufacturers who say staffing shortages are hurting product quality are not going to fix that by handing operators a tool they resent.

On the maintenance side, the equivalent of operator trust is the logged save. A predictive-maintenance pilot (PdM, using sensor and historian data to predict a failure before it happens) lives or dies on whether it can show a real save written into the CMMS (computerized maintenance management system, the software that holds work orders and equipment history). When the model flags the spindle bearing and the planner schedules the swap during a planned changeover instead of eating an unplanned six-hour line stop, that avoided downtime has to be logged as a save with a number on it, tied to the work order, ideally with the control cell's later failure as the counterfactual. A dashboard that lights up amber is not evidence. A CMMS record that says "predicted bearing failure on Press 2, swapped during changeover, avoided an estimated six hours of unplanned downtime worth roughly 18,000 dollars" is evidence a finance director can take to a budget meeting.

The save-log discipline matters because of how leadership counts. A plant manager staring at a downtime Pareto where "unplanned" is the tallest bar does not care that a model has 94% recall on a benchmark. He cares that last quarter the line stopped four times and this quarter, on the piloted cell, it stopped once and the model called the one it caught. The pilot's job is to convert a model metric into an operations metric the plant already lives by: hours of unplanned downtime, first-pass yield, escape rate, OEE (overall equipment effectiveness, the combined measure of availability, performance, and quality). Evidence is a translation job as much as a measurement job.

Packaging the Evidence and the Scale Decision

When the decision date arrives, the pilot owner has to put something in front of the plant manager, the quality lead, and the finance director that survives their skepticism. This is the evidence package, and it is the deliverable the whole pilot exists to produce. A weak pilot hands over a vendor slide and an anecdote. A disciplined pilot hands over a short, defensible package built from the bar that was set on day one.

The evidence package has six parts. The baseline and the target, exactly as written before the pilot, so nobody can move the goalposts after the fact. The result against the control, stated in the operations metrics that matter: escape rate, false-reject rate, unplanned downtime hours, yield delta, each with the control comparison so the result points at the AI and not the weather. The dollar math, the worked number that nets the value created against the cost of false rejects, false alarms, integration, and ongoing maintenance, so finance sees the real figure and not the gross one. The drift and durability evidence, showing the system held its numbers across shifts and lighting and lots over the pilot window, which is the proof that scale will not collapse in month two. The OT and audit trail, showing the pilot stayed advisory, respected the OT boundary, and logged every AI-touched decision, because the customer audits the plant, not the vendor, and "the model flagged it" is never a sufficient answer to an auditor. And the scale plan and its cost, what it takes to roll from one line to four, including the unglamorous integration and the ongoing drift-monitoring labor.

The decision itself is one of three, and a disciplined pilot is comfortable with all three. Scale when the result clears the target against the control and the dollar math holds at scale. Stop when the kill criteria were hit, and stop cleanly, freeing the budget and the engineering hours without shame, because a pilot that fails fast and proves a use case does not work is a success of the discipline, not a failure of the plant. And iterate when the result is promising but short, when a defined change, better lighting, a re-trained model, a tighter sensor placement, has a clear shot at clearing the bar in a bounded second window with its own kill criteria. What you do not do is the Ohio thing: leave it "on hold," undecided, consuming attention, producing nothing.

Across a multi-site network, this discipline compounds. A plant that runs five disciplined pilots a year and scales the two that clear the bar, kills the two that hit kill criteria, and iterates the one that is close has spent its pilot budget on evidence and freed itself from the graveyard of half-dead demos. Structured programs of this kind see 3 to 4 times higher adoption than self-directed, ad-hoc experimentation, for the simple reason that they produce results people can trust and repeat. The credential the graduate carries out of this is not "we tried AI." It is "here is the bar we set, here is the control we ran, here is the false-reject number, here is the logged save, and here is the line we scaled because the evidence held."

Key Takeaways

  • A demo is built to be applauded and a pilot is built to be defended. Most plant AI pilots die because someone ran a vendor demo for eight months and called it a pilot, with no baseline, no control, no kill criteria, and no owner.
  • Set the evidence bar before the hardware arrives: a measured baseline, a numeric target with a false-reject ceiling, explicit kill criteria, and a named decision owner with a decision date. The one-page bar is the most valuable artifact the pilot produces even if it fails.
  • Run a control: a matched cell, alternating shifts, or a scored holdout, so the result points at the AI and not at a new material lot or warmer weather. Without a control you have correlation, not evidence.
  • Measure the unflattering numbers as carefully as the flattering ones. Track false-reject rate by shift and lighting, use a confusion matrix, and net the value against the real cost of false rejects, which can quietly exceed the value of the escapes caught.
  • Budget for the brownfield plumbing as a first-class line item, keep the AI advisory and out of the control loop, respect the OT boundary that 78% of plants cannot fully monitor, and build a drift check in from day one so the system survives past week six.
  • Operator trust is a design requirement, not an afterthought. A false alarm that wastes an operator's time gets the green light overridden, and a system nobody uses produces no evidence. On maintenance, the equivalent is a real save logged in the CMMS with a dollar figure and a control counterfactual.
  • Translate model metrics into the operations metrics leadership already lives by: unplanned downtime hours, first-pass yield, escape rate, and OEE. Evidence is a translation job as much as a measurement job.
  • End every pilot with a scale, stop, or iterate decision backed by a six-part evidence package, baseline and target, result against control, dollar math, drift and durability, OT and audit trail, and the scale plan with its cost. A clean stop on kill criteria is a win, because it frees budget and attention for a use case that can actually hold.