Success Metrics for Plant AI
Eighteen months into the plant's AI push, the corporate slide deck looks triumphant. Four models deployed. A vision system on the assembly line, a predictive-maintenance model on the gearboxes, a root-cause assistant in the quality office, and a knowledge base built from interviews with Dave the inspector before he retired in November. The VP who toured a competitor's smart factory is happy. The slide says "4 AI deployments live." And yet the plant manager, standing in the morning production meeting, cannot answer the only question that matters when the CFO calls: did any of it actually move the plant? First-pass yield is flat. The downtime Pareto looks the same as it did before, with "unplanned" still the tallest bar. The operators on second shift quietly stopped trusting the vision system's green light two months ago because it kept rejecting good parts, so they wave parts through and the model's numbers look great while the real escape rate creeps up. The deck measured the wrong thing. It measured activity, "models deployed," when it should have measured results: the four numbers that prove a plant-AI program is working or quietly is not. This lesson is about those four numbers, why they are the only honest scoreboard, and how to keep a vanity metric from telling you a comforting lie while the plant slides.
The Four Numbers That Matter
A plant-AI program touches a hundred things, but it justifies itself on four. Get these four right and you can defend the program to a CFO, a customer, and a skeptical operator. Get them wrong, or worse, ignore them in favor of activity counts, and you will run a program that looks busy and changes nothing.
Number one: Overall Equipment Effectiveness (OEE). OEE is the master gauge of how well a line actually runs, expressed as a single percentage that multiplies three factors together: availability (was the machine running when it was scheduled to?), performance (was it running at its rated speed?), and quality (were the parts good?). A line that is available 90 percent of the time, runs at 95 percent of rated speed, and makes 98 percent good parts has an OEE of 0.90 times 0.95 times 0.98, which is about 84 percent. World-class is often cited near 85 percent, and most real lines live well below that. OEE matters as an AI metric because it is the one number that catches a program cheating on one factor while wrecking another. A vision system that boosts the quality factor by catching defects but slows the line and tanks the performance factor can leave OEE flat or lower. If your AI moved a sub-metric but OEE did not move, the program did not actually help the line.
Number two: First-Pass Yield (FPY). FPY is the percentage of units that make it through the process correctly the first time, with no rework and no scrap. It is the cleanest single measure of quality health, and it is where AI in quality is supposed to pay off, since 47 percent of manufacturers now use AI in quality, up from 33 percent the year before. The trap is that FPY can be gamed by a vision system that rejects too aggressively: parts get pulled, reworked, and re-run, so the line stays busy but the true yield is hidden behind a rework loop. The honest FPY counts a unit as a first-pass success only if it passed clean the first time, which is exactly why it exposes a vision program that is trading false rejects for a comforting defect count.
Number three: Unplanned Downtime. This is the tallest bar on most plants' loss charts and the headline target for predictive maintenance. Measured in hours, or better in dollars of lost contribution margin, unplanned downtime is the number a predictive-maintenance program must actually reduce to justify itself. The subtlety is that a model can generate hundreds of alerts and a beautiful dashboard while unplanned downtime does not budge, because nobody turned the alerts into work orders that prevented failures. Unplanned downtime is the metric that refuses to be impressed by a dashboard. Either the breakdowns went down or they did not.
Number four: False-Reject Rate. This is the metric the other three programs do not have and the one most plants forget to track. The false-reject rate is the percentage of good parts the vision system wrongly rejects. It is pure cost: every false reject is a good part scrapped or reworked for nothing, plus the operator time to handle it, plus the slow poison of eroding trust. False-reject economics are real money, and a vision system's false-reject rate can quietly cost more than the escapes it catches. It is the canary for the whole quality-AI program, because the moment it climbs, operators start disabling the green light, and a disabled green light makes every other quality number a fiction.
Work the false-reject math once and you will never ignore it again. Suppose a line runs 5,000 parts a shift and the vision system carries a 2 percent false-reject rate. That is 100 good parts a shift pulled for no reason. If each false reject burns five minutes of operator handling plus a re-inspection, that is over eight hours of wasted labor a shift, roughly a full extra person doing nothing but chasing the model's mistakes. If even a fraction of those good parts get scrapped rather than recovered, the material loss stacks on top. Now compare that to the escape side: the same model might catch a handful of real defects a week. It is entirely possible for the false-reject cost to dwarf the escape-prevention benefit, and the only way you would ever know is by tracking both numbers side by side. A program that reports its catch rate but not its false-reject rate is showing you one side of a ledger and hiding the side that may be bleeding.
OEE, first-pass yield, unplanned downtime, and false-reject rate. If the program cannot move these four, it did not work, no matter how many models are live.
Why "Models Deployed" Is Not a Result
The most seductive lie in a plant-AI program is the activity metric. "Four models deployed." "Twelve use cases in production." "Two thousand alerts generated this quarter." Every one of these counts effort, not outcome, and a program managed to activity metrics will optimize for activity: more models, more alerts, more dashboards, none of which the line feels.
The mechanism of the lie is worth understanding because it is so easy to fall for. Deploying a model is visible, finite, and satisfying. It produces a launch date, a screenshot, a slide. Moving OEE two points is slow, contested, and hard to attribute. So a program under pressure to show progress reaches for the visible thing. The vendor reinforces this because the vendor sells deployments, not yield. Treat vendor and research performance figures as benchmarks to verify, never as results you have achieved. A vendor's claim that the model is "99 percent accurate on the test set" is not your false-reject rate on your wet, variable, drifting line. The only number that counts is the one measured on your floor, after deployment, against your baseline.
Here is the worked contrast. Plant A reports "predictive-maintenance model live on 12 assets, 1,400 alerts generated." Plant B reports "unplanned downtime on the two pilot lines down 11 percent over the prior six months, three failures averted and logged in the CMMS." Plant A has more impressive activity. Plant B has the only thing a CFO will fund again. And note the trap inside Plant A's number: 1,400 alerts on 12 assets is roughly two alerts per asset per week, which is almost certainly alert fatigue. A crew buried in alerts stops reading them, which means the program is generating cost (attention) and no benefit (averted failures). The activity metric not only failed to prove value; it actively hid the fact that the program had become noise.
The discipline is to refuse to report an activity number as if it were a result. "Models deployed" belongs on a project-status update, not a value scorecard. The value scorecard carries the four numbers and nothing that can be inflated by simply doing more. When leadership asks "how is the AI program going," the honest answer is a before-and-after on OEE, FPY, unplanned downtime, and false-reject rate, with the baseline stated, not a count of what was launched.
There is a quieter reason activity metrics are dangerous: they shape behavior. What you measure is what the team optimizes, and a team rewarded for deployments will ship deployments whether or not the line needs them. You end up with twelve thin pilots that each move nothing instead of two deep ones that each move a number, because twelve looks better on the activity slide. The same dynamic corrupts vendor relationships. A vendor paid and renewed on deployments has every incentive to expand the footprint and no incentive to prove the footprint earned its keep. The moment you switch the scorecard from activity to the four results, the incentives flip: now the team and the vendor both have to make a number move, and the thin pilots that were padding the count get exposed as the cost centers they always were. Changing the metric is not a reporting cleanup. It is the single most powerful lever a plant strategist has to redirect a program from looking busy to being useful.
The Baseline and the Counterfactual
A number means nothing without the number it replaced. The most common way a plant-AI program fools itself is by reporting a current value with no honest baseline, or by claiming credit for an improvement the AI did not cause.
Set the baseline before you deploy, not after. The baseline is the four numbers measured over a representative period, long enough to average out the good weeks and the bad, before the AI touches the line. If you wait until after deployment to define the baseline, you will, consciously or not, pick a comparison period that flatters the program. A disciplined baseline is dated, documented, and agreed with the people who will later judge the program. For a seasonal plant, the baseline must span the season; comparing a slow summer post-deployment to a busy winter pre-deployment will manufacture an improvement that is really just demand.
Beware the counterfactual. The hard question a good CFO will ask is: would the number have moved anyway? If unplanned downtime fell 11 percent, but you also rebuilt two machines and added a maintenance tech in the same window, how much of the gain is the AI and how much is the steel and the headcount? You will rarely isolate this perfectly on a running plant, and pretending you can is its own dishonesty. The practical answer is to log the specific saves the AI caused, the individual predicted failures that were averted and written up in the CMMS (Computerized Maintenance Management System, the software that holds work orders and maintenance history). A logged save with a date, an asset, a prediction, and an averted-downtime estimate is defensible attribution. A plant-wide trend line is suggestive but contestable. The strongest case pairs the trend (the four numbers moved) with the receipts (here are the eleven specific saves that drove it).
Worked example of honest attribution. The predictive-maintenance program reports unplanned downtime down 11 percent, and backs it with a CMMS log of three averted failures: a gearbox bearing caught six weeks early, a pump seal caught before a leak, and a motor flagged before a winding failure. Each entry estimates the downtime avoided based on what that failure mode historically costs, say eight hours at five thousand dollars an hour of contribution margin, which is forty thousand dollars per averted event. Three events is on the order of a hundred and twenty thousand dollars of avoided downtime, against a program cost the plant can name. That is a defensible ROI conversation. The same 11 percent with no logged saves is a number the CFO can wave away as luck, and should.
Leading Versus Lagging Indicators, and the Trust Metric
The four headline numbers are mostly lagging indicators: they tell you what already happened. OEE, FPY, and unplanned downtime move slowly and confirm a result after the fact. To run the program week to week, you also need leading indicators, the early signals that predict whether the lagging numbers will move, and one of them is so important it deserves to be treated as a metric in its own right: operator trust.
The false-reject rate is the leading indicator for the whole quality-AI program. It moves before FPY does, and it predicts whether operators will keep using the system. Watch it weekly. A climbing false-reject rate is the early warning that the model is drifting, that lighting or camera angle or material changed between shifts, and that you are days away from operators disabling the green light. By the time FPY reflects the damage, the trust is already gone and the harder problem is winning it back.
Operator trust is a metric you can actually measure, and you must. The cleanest proxy is the override rate: how often operators ignore, disable, or work around the AI. If second shift is waving parts past a vision system they have stopped believing, the model's reported accuracy is fiction, because it is only judging the parts the operators still let it see. An override rate that climbs is a louder alarm than any dashboard, because it means the humans who actually run the line have voted no confidence. A plant-AI program with great model metrics and a 40 percent operator override rate has failed, and the model metrics are hiding it. Measuring trust, through override rate, through a quick monthly operator pulse, through how many techs actually action the maintenance alerts, is how you catch the failure the lagging numbers will not show you for months.
The relationship between the metrics is a chain. A drifting model raises the false-reject rate (leading), which raises the operator override rate (trust), which corrupts the reported quality numbers and eventually shows up as falling true FPY and a quiet rise in escapes (lagging). The plant that watches only the lagging end of this chain learns about the failure last, after parts have shipped. The plant that watches the leading and trust end catches it first, while it is still a tuning problem and not a containment.
Building the Metrics Dashboard That Survives a CFO
A scorecard that survives leadership scrutiny has a specific shape. It is short, it is honest about baselines, it separates the four results from the supporting detail, and it never lets an activity count masquerade as a result.
The top of the dashboard carries the four numbers, each shown as baseline, current, and delta, with the baseline period named. OEE: was 71 percent over the prior twelve months, now 74 percent. FPY: was 94.2 percent, now 95.6 percent. Unplanned downtime: was 6.1 percent of scheduled time, now 5.4 percent. False-reject rate: was not measured before, now 1.8 percent and falling. That top row is the whole program's report card, and it is legible to a CFO in ten seconds.
Below the headline sit the leading and trust indicators that explain whether the headline will hold: the weekly false-reject trend, the operator override rate, the count of maintenance alerts actioned versus generated. These are the early-warning band. A CFO does not need to read them every month, but a plant manager does, because they are where the program is saved or lost.
At the bottom sit the receipts: the logged saves with dates and dollar estimates, the specific averted failures, the documented yield improvements tied to specific defect modes the vision system now catches. This is the attribution evidence that turns a trend into a defensible ROI claim. Crucially, the activity counts, models deployed, alerts generated, do not appear on this scorecard at all, because they belong on a separate project-status page where they cannot be mistaken for value.
Worked example of the dashboard catching a problem. Month nine, the headline looks fine: FPY is up, OEE is up. But the early-warning band shows the false-reject rate ticked from 1.8 to 3.1 percent over three weeks and the second-shift override rate jumped to 22 percent. The plant manager reads the early-warning band, not just the headline, and acts: the team finds that a maintenance crew adjusted the line lighting during a recent shutdown, the model drifted, and operators were starting to wave parts. They retune, the false-reject rate falls back, and the override rate recovers, all before the lagging FPY number ever turned negative. The dashboard earned its keep not by reporting the result but by surfacing the leading signal in time to protect it. That is the difference between a scorecard that decorates a meeting and one that runs a program.
Key Takeaways
- Four numbers prove a plant-AI program is working or quietly is not: Overall Equipment Effectiveness (OEE), First-Pass Yield (FPY), unplanned downtime, and false-reject rate. If the program cannot move these, it did not work, regardless of how many models are live.
- OEE is the master gauge because it multiplies availability, performance, and quality, so it catches a program that boosts one factor while wrecking another, such as a vision system that improves quality but slows the line.
- "Models deployed," "use cases live," and "alerts generated" are activity metrics, not results. A program managed to activity will optimize for activity the line never feels. Vendor accuracy claims are benchmarks to verify, never results you have achieved on your floor.
- A number means nothing without an honest baseline measured before deployment, dated and agreed, and spanning the full season for a seasonal plant. Reverse-engineering a flattering comparison period is the most common self-deception in a plant-AI program.
- Beware the counterfactual: would the number have moved anyway? Defend attribution with logged CMMS saves, specific averted failures with dates and dollar estimates, paired with the trend. A logged save is defensible; a trend line alone is contestable.
- The false-reject rate is the leading indicator for the whole quality-AI program; it climbs before FPY falls and predicts when operators will disable the green light. Watch it weekly.
- Operator trust is a measurable metric, best proxied by the override rate. A program with great model numbers and a high override rate has failed, and the model numbers are hiding it because they only judge the parts operators still let the system see.
- A scorecard that survives a CFO shows the four numbers as baseline, current, and delta with the baseline named, an early-warning band of leading and trust indicators, and the logged-save receipts for attribution, with activity counts kept off it entirely.
Skill.re