Avoiding Vanity Metrics
The corporate AI steering committee is reviewing the plant on a Thursday afternoon, and the slide that goes up is glowing. Twelve models deployed. Eighteen thousand alerts generated last quarter. Ninety-four percent model accuracy in validation. Three lines instrumented. A heat map with a lot of green. The VP nods, the plant gets its budget renewed, and everyone files out feeling good. Down on the floor, none of it is true in the way that matters. First-pass yield is exactly where it was a year ago. The downtime Pareto has the same tallest bar it had last January. The maintenance crew quietly stopped opening the predictive alerts in week three because most of them were noise. Twelve models are deployed and not one of them has moved a number that shows up on the plant manager's loss chart. That gap, between the metrics that win the meeting and the metrics that move the business, is the entire subject of this lesson. The slide measured activity. Nobody measured outcome.
The Difference Between Activity and Outcome
A vanity metric is a number that goes up reliably, looks impressive in a deck, and tells you almost nothing about whether the thing you built is working. It is seductive precisely because it is easy to produce and always trends in the flattering direction. "Models deployed" only ever increases. "Alerts generated" only ever increases. "Accuracy in validation" was measured on the vendor's clean test set and will never embarrass you. These numbers are real, in the sense that you can count them, but they are not results. They are evidence that work happened, not evidence that the work mattered.
The opposite of a vanity metric is an outcome metric, a number tied directly to the loss chart the plant actually runs on. In manufacturing those numbers are not mysterious, because the floor has measured them for decades: first-pass yield (FPY, the percentage of parts that pass all quality steps the first time without rework), OEE (Overall Equipment Effectiveness, the product of availability, performance, and quality that summarizes how well a line actually runs), unplanned downtime hours, scrap and rework cost, the false-reject rate on a vision cell, and the customer escape rate. These are the bars on the loss chart. The whole point of putting AI on the floor was to move one of them. A metric that does not connect to one of them is not measuring your AI program; it is decorating it.
Here is the test that separates the two, and it is worth memorizing because it cuts through almost every glowing dashboard. Ask of any metric: if this number doubled tomorrow, would the plant be measurably better off? If "models deployed" doubled from twelve to twenty-four, would yield rise or downtime fall? Not necessarily, and possibly the opposite, because more half-baked models means more false alarms and more crew distrust. If "alerts generated" doubled, would the plant be better? Almost certainly worse, because the crew is already drowning. Now run the test on an outcome metric. If unplanned downtime hours dropped by half, is the plant better off? Unambiguously, measurably, in dollars. That asymmetry is how you tell a vanity metric from a real one. A vanity metric can double while the business gets worse. An outcome metric cannot.
If the number can double while the plant gets worse, it is a vanity metric. The only metrics worth reporting are the ones that cannot.
The Loss Chart Is the Scoreboard
Every plant already has the right scoreboard, and most AI reporting ignores it. The loss chart, the ranked list of where the plant bleeds money, is usually dominated by two bars: unplanned downtime and scrap or rework. The reason those bars are so tall in 2026 is not primarily technical, it is demographic. With 85 percent of manufacturers saying staffing shortages are hurting product quality, and the most experienced inspectors and techs retiring, the losses on that chart are getting worse for reasons no amount of "models deployed" addresses. The job of plant AI is to push those specific bars down. So the job of plant AI measurement is to prove, with a before and an after, that a specific bar moved because of a specific deployment.
This reframes reporting entirely. Instead of leading with how many models you deployed, you lead with the bar you targeted and what happened to it. Consider a real shape of report. Before deployment, Line 2 averaged 38 hours of unplanned downtime a month, the tallest bar on the plant's chart. The team deployed a predictive maintenance model on the three motors that drove most of that downtime. Six months later, Line 2 averages 22 hours of unplanned downtime a month. That is a 16-hour-per-month reduction. If a downtime hour on Line 2 costs 4,000 dollars in lost throughput, scrap at restart, and labor, 16 hours a month is 64,000 dollars a month, roughly 768,000 dollars a year. That sentence is worth more than every vanity slide combined, because it names the bar, shows the before and after, attaches a dollar figure, and survives the only question that matters: "How do you know the AI did that?"
Notice what the outcome report forces you to have that the vanity slide never required: a baseline. You cannot claim a 16-hour reduction unless you measured the 38 hours before you started. The single most common reason plants cannot prove AI ROI is that nobody wrote down the baseline, so when the number improves there is no honest way to attribute the change, and when it does not improve there is no way to notice. Measuring the baseline before you deploy is not bureaucracy. It is the only thing that makes any later claim defensible. If you take one operational habit from this lesson, make it this: never deploy a floor AI system without first recording, in writing, the current value of the loss-chart bar you intend to move.
The Vanity Metrics That Fool Good People
These metrics are not stupid, and the people who report them are not lazy. They are tempting because each one feels like it should correlate with value, and sometimes loosely does. Naming them precisely is how you stop quoting them. Here is the catalog of the worst offenders on a plant floor.
Models deployed. The flagship vanity metric. It measures effort, not effect. A plant with twelve mediocre models that nobody trusts is worse off than a plant with one excellent model that logged a real save, but the first plant has a better-looking slide. Deployment is a cost, not a benefit. Reporting it as an achievement is reporting your spending as if it were your earnings.
Alerts generated. The most actively harmful vanity metric, because the number you want is usually the opposite of high. A predictive maintenance system that generated 18,000 alerts last quarter, of which the crew acted on 200 and ignored the rest, is not eighteen thousand units of value. It is a system actively training the crew to ignore it. The right metric is not alerts generated; it is the alert-to-action ratio and the count of confirmed saves. A system that fires five precise alerts a week, all five investigated, two of which prevented a real failure, beats the 18,000-alert system on every dimension that matters and loses every beauty contest.
Validation accuracy. The most respectable-looking vanity metric, because it has a decimal point and sounds scientific. The trap is that 94 percent accuracy on the vendor's clean validation set tells you nothing about performance on your wet parts under changing shift lighting on a drifting camera. Worse, accuracy is the wrong measure for an imbalanced problem. If 2 percent of your parts are defective, a model that blindly passes everything is 98 percent accurate and catches zero defects. Accuracy can be high while the model is useless. The numbers that matter are precision, recall, the false-reject rate, and the escape rate measured on your line, in production, over time, not a single accuracy figure measured once in a lab.
Dashboards built and lines instrumented. Infrastructure metrics dressed as outcomes. A dashboard is a window, not a result. Instrumenting a line is a precondition for value, not value itself. Reporting "three lines instrumented" answers the question "what did you buy" rather than "what did it earn," and leadership, reasonably, eventually starts asking the second question.
User logins and prompts run. For generative and knowledge tools, the equivalent trap. That operators ran 4,000 prompts last month tells you the tool was opened, not that any work instruction it drafted was verified, used, and correct. Usage is not value. A single verified SOP that a green crew used to run a machine safely is worth more than ten thousand unread chatbot sessions.
The common thread is that every one of these counts an input or an activity, and dresses it as an output. The discipline is to keep asking, of every number on the slide, "is this something we spent or something we earned?" Spending is not an achievement. The plant did not put AI on the floor to deploy models. It put AI on the floor to move the loss chart.
Leading, Lagging, and the Trap of the Easy Proxy
There is a subtler version of the vanity-metric problem that catches sophisticated teams, the ones who already know not to report "models deployed." It is the trap of the easy proxy: choosing a metric not because it best measures the outcome but because it is the easiest to pull from the system. The classic example is a quality team that reports "defects detected by the vision system" as a success metric. That sounds like an outcome, it is closer than "models deployed," and it is wrong, because detected defects is partly a function of how many defects there were to detect. A great upstream process that produces fewer defects makes the AI's "defects detected" number fall, which would make a good plant look like a failing one. You have chosen a proxy that moves for reasons unrelated to the AI's value.
The fix is to understand the difference between a leading metric and a lagging metric, and to insist on the lagging one as the verdict while using the leading one only as an early-warning signal. A lagging metric is the outcome itself, measured after the fact: the customer escape rate, the unplanned downtime hours, the first-pass yield for the quarter. It is the truth, but it arrives late, sometimes a quarter late, which is why teams are tempted to substitute something faster. A leading metric is an early indicator that should predict the lagging one: the false-reject rate this week, the alert-to-action ratio this week, the drift-monitor reading on the vision model. Leading metrics are useful for steering between reporting cycles, catching a problem before it shows up in the quarterly escape rate. They are dangerous when they get promoted into the headline result, because a leading metric can look fine while the lagging metric it was supposed to predict goes the wrong way.
Worked example of the trap and the fix. A maintenance team reports "predictive model precision is 88 percent" as its headline quarterly result, a leading metric, and the deck is green. But the lagging metric, unplanned downtime hours, did not move at all, because the model was precise about failures the crew was already catching on the existing PM walks and missed the two big surprise failures that actually drove the downtime bar. The precision number was true and useless as a verdict. The honest report leads with the lagging metric, downtime hours unchanged, names the precision figure as context, and concludes the model is not yet earning its keep on the bar that matters. That is the discipline: the lagging loss-chart number is the verdict, leading metrics are the dashboard you watch between verdicts, and you never let a leading metric stand in for an outcome just because it arrived sooner.
This also reframes how often to report. Vanity programs report on a fixed calendar because the numbers always look good, so the cadence is just a ritual. An outcome program reports the lagging metric on the cycle the loss chart actually moves on, monthly for downtime, quarterly for escape trends, and uses the leading metrics continuously so that a deployment in trouble is caught in week three, not at the quarterly review when three months of a drifting model have already leaked into the scrap bucket. The cadence follows the physics of the loss, not the calendar of the steering committee.
Building Metrics That Survive the Loss-Chart Question
Replacing vanity metrics with real ones is not about banning the easy numbers, it is about building a small set of outcome metrics that survive scrutiny. A defensible plant AI metric has four properties, and you can check any proposed metric against them in about a minute.
First, it ties to a loss-chart bar. The metric moves a number leadership already cares about: yield, downtime, scrap, false-reject cost, escape rate, OEE. If you cannot draw a straight line from the metric to a bar on the loss chart, it is decoration. This is the property that kills "models deployed" instantly, because deployment connects to no bar.
Second, it has a baseline and a before-and-after. The metric is reported as a change from a recorded starting point, not as an absolute that floats free of history. "Downtime is 22 hours" means nothing. "Downtime fell from 38 to 22 hours after the model went in" is a claim you can defend. The baseline is what converts a number into evidence.
Third, it carries a dollar figure. Leadership funds the program in dollars and should see the result in dollars. Sixteen fewer downtime hours a month is good; 64,000 dollars a month saved is fundable. The translation from operational units to money is what lets a plant manager defend the budget and a graduate of this program defend their own role. Always do that translation, and always show the per-unit assumption (what a downtime hour costs, what a scrapped part costs) so the number is auditable rather than magic.
Fourth, it can be attributed honestly. The hardest and most important property. When yield improves, was it the AI, or was it the new supplier lot, the process tweak, the seasonal change, or the fact that the best operator happened to be on days that month? Honest attribution means you can rule out the obvious confounders, ideally because you ran a controlled comparison: the instrumented line versus a sister line that did not get the model, or the same line before and after with nothing else changed. A metric you cannot attribute is a coincidence you are taking credit for, and the first time leadership catches that, every number you report afterward is suspect.
Worked example of the four properties together. A vision cell on Line 5 had a customer escape problem: an average of two defect escapes a quarter, each one a containment that cost roughly 80,000 dollars in sorting, expedited replacement, and customer-team time, or about 640,000 dollars a year. After the vision system went in, with a daily independent re-check as a safeguard, escapes dropped to zero over the following two quarters, and the false-reject rate was measured and held at 1.5 percent rather than allowed to run wild. The report writes itself: it ties to the escape-rate bar, it shows before (two a quarter) and after (zero), it carries a dollar figure (about 640,000 dollars a year of containments avoided), and it can be attributed because the change was isolated to the vision deployment with the false-reject cost explicitly measured so the number is net of the system's own waste. That is a metric a board can fund and an auditor can trust.
The Honesty of Reporting What Did Not Work
There is one more property of a real metrics program that vanity reporting structurally cannot have: it can report failure. A dashboard built to make the program look good will never show a model that did not pay off, because its whole purpose is to glow. But a metrics program built on the loss chart will sometimes have to say "we deployed a scheduling optimizer on Line 4 and OEE did not move, so we are turning it off." Far from being a weakness, that sentence is the single strongest signal that your reporting is honest, and honest reporting is the thing that makes the good numbers believable.
Consider why this matters in front of a skeptical VP or a board. If every quarterly report is uniformly green, a thoughtful leader stops believing it, because no real program of a dozen experiments succeeds twelve times out of twelve. The all-green deck signals either that the team is measuring vanity metrics that cannot fail, or that it is hiding the failures. Either way, trust erodes. A report that says "three deployments moved the loss chart and saved a combined 1.4 million dollars, two are still proving out, and one we killed because it did not work" is far more credible, because it shows the team is measuring outcomes ruthlessly enough that failures are visible. Killing a model that did not earn its keep is not a black mark. It is proof the measurement system works, and it frees the budget and the crew's patience for the deployments that do.
This connects directly to alert fatigue and operator trust, the social contract underneath the whole floor-AI effort. Operators and techs can tell the difference between a program that measures whether the AI actually helps them and one that measures whether the slide looks good. When you report alert-to-action ratios and confirmed saves, you are telling the crew you care whether the system earns its place on their line. When you report alerts generated, you are telling them the opposite. The metrics you choose are not just a corporate reporting decision; they are a message to the floor about whether you are serious. The plant that reports outcomes, baselines, dollars, honest attribution, and the occasional honest failure is the plant whose crew trusts the AI, whose leadership keeps funding it, and whose customer audit goes smoothly because every claim on the slide is one the plant can actually back up. The slide that won the Thursday meeting with twelve models and eighteen thousand alerts buys none of that. It buys a renewal that lasts exactly until someone asks what moved.
Key Takeaways
- A vanity metric goes up reliably, looks impressive, and tells you nothing about whether the AI works. The test: if this number doubled tomorrow, would the plant be measurably better off? A vanity metric can double while the business gets worse; an outcome metric cannot.
- "Models deployed," "alerts generated," "validation accuracy," "dashboards built," "lines instrumented," and "prompts run" are the worst offenders. Each counts an input or an activity and dresses it as an output. Deployment is spending, not earning.
- The loss chart is the scoreboard. Real plant AI metrics tie to bars the plant already runs on: first-pass yield, OEE, unplanned downtime hours, scrap and rework cost, false-reject rate, and customer escape rate. If a metric connects to no bar, it is decoration.
- Never deploy a floor AI system without first recording the baseline value of the loss-chart bar you intend to move. The most common reason plants cannot prove AI ROI is that nobody wrote down the before, so no later claim can be defended.
- A defensible metric has four properties: it ties to a loss-chart bar, it shows a before-and-after against a baseline, it carries a dollar figure with an auditable per-unit assumption, and it can be attributed honestly by ruling out confounders, ideally with a sister-line or before-and-after comparison.
- Validation accuracy is especially deceptive: 94 percent on a clean vendor set says nothing about your drifting camera, and on an imbalanced defect problem a model can be 98 percent accurate while catching zero defects. Report precision, recall, false-reject rate, and escape rate measured on your line over time instead.
- The worked numbers: a PdM model cutting Line 2 downtime from 38 to 22 hours a month at 4,000 dollars an hour is about 768,000 dollars a year; a vision cell taking Line 5 escapes from two a quarter to zero avoids roughly 640,000 dollars a year in containments. Those sentences fund a program; vanity slides do not.
- A real metrics program can report failure, and that is its strength. An all-green deck signals hidden failures or unfalsifiable vanity metrics and erodes trust. Killing a model that did not move the loss chart proves the measurement works and frees budget and crew patience for the deployments that do.
Skill.re