โ†
AI for Manufacturing
Strategic ยท M14 ยท lesson 14 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
PoC Design on a Real Line
๐Ÿ“–
now learning

PoC Design on a Real Line

15 min

The vendor's demo was flawless. On a thirty-inch monitor in a clean conference room, the AI vision system found every scratch, every short shot, every flash line on a tray of sample parts the salesperson had brought in a padded case. The plant manager nodded. The quality manager nodded. Somebody said the word "obviously" twice. Three months and one signed purchase order later, the same system was bolted over the real line, and it was rejecting eleven percent of good parts because the overhead lights in that bay were a different color temperature than the demo booth, the conveyor vibrated the camera mount a few thousandths every cycle, and the night shift ran a slightly glossier resin from a second supplier. The escapes it was bought to catch? It caught some. But the operators had already taped a piece of cardboard over the reject chute light because the false alarms were stopping the line every four minutes, and once the green light is taped over, you do not have a quality system, you have an expensive paperweight. The plant did not buy a bad product. The plant bought a demo and called it a proof. A proof of concept (PoC, a small, time-boxed test that proves whether a thing actually works in your conditions before you commit real money) is the single cheapest insurance you will ever buy against that outcome, and most plants run it wrong: they let the vendor prove the vendor can demo, when the only question that matters is whether the system holds up on your line, with your parts, your lighting, your crew, and your definition of a good part.

The Difference Between a Demo and a Proof

A demo is designed to make you say yes. A proof is designed to find out whether yes is the right answer. Those are opposite jobs, and the vendor cannot do both for you. The vendor's demo runs on hand-picked golden parts, in controlled light, with a model the vendor's own engineers tuned overnight on data the vendor selected. None of those conditions exist on your floor at 2 a.m. on a hot Thursday in August when the line is behind schedule and the new operator has never seen this fault before. The demo answers "can this technology work somewhere?" The proof answers "will this technology work here, run by us, on the loss we are actually trying to kill?"

The reframe that saves plants money is this: a PoC is not a sales step you tolerate to humor a vendor. It is an experiment you design, you control, and you grade against a number you wrote down before the vendor ever touched your line. If the vendor designs the PoC, the vendor designs it to pass. If you design it, you design it to tell the truth. Greenfield plants, the brand-new facilities built from a blank slab, deploy AI roughly 40 to 60 percent faster than brownfield plants precisely because they get to define conditions once and lock them. You are almost certainly brownfield: a 1990s programmable logic controller (PLC, the industrial computer that actually runs the machine's logic), a historian nobody has queried in years, mixed lighting, and three suppliers feeding one line. The brownfield plant cannot skip the proof. The brownfield plant is the proof.

A demo proves the vendor can sell. A proof of concept proves the system survives your line. Never confuse the first for the second.

Consider the math on getting this wrong. A mid-line vision QA system, installed, integrated, and trained, runs somewhere in the range of 80,000 to 250,000 dollars of capital plus the integration labor. If you sign on the demo and discover the false-reject problem in month two, you have not just lost the capital. You have lost the line time the false rejects ate, the operator trust you will spend a year trying to rebuild, and the credibility of the next AI project you bring to the floor. A well-designed PoC that costs you three weeks of engineering attention and a small data-collection effort is cheap against a six-figure mistake you cannot fully unwind. The PoC is not the expensive part. The PoC is what keeps the expensive part from being a loss.

Define Success Before You Sign

The most important sentence in any PoC lives in a document you write before the vendor is in the building, and it reads like this: "This proof of concept succeeds if, and only if, it achieves X on metric Y measured by method Z over period P." Every word of that sentence is load bearing. If you cannot fill in all four blanks, you are not ready to run a PoC, because you have no way to fail it, and a test you cannot fail is a test that proves nothing.

Tie the success metric to a real loss with a real dollar figure. Walk the loss chart first. The two tallest bars on almost every plant's loss chart are unplanned downtime and scrap plus rework, and 85 percent of manufacturers say staffing shortages are actively hurting product quality, which means the loss you pick is probably a quality escape or a breakdown that your thinning, greener crew can no longer reliably catch by eye. Say your line ships first-pass yield (FPY, the percentage of parts that pass inspection the first time with no rework) of 94 percent, and the two-point drop last quarter to that number is costing you 240,000 dollars a year in scrap, rework, and one near-containment. That is your loss. Your PoC success metric is now anchored to it.

The four blanks, filled in

Metric Y is the thing you measure, and for a vision QA PoC it is almost never "accuracy." Accuracy is a vanity number that hides the two failure modes that actually cost money. You want two metrics: the escape rate (the fraction of bad parts the system passes as good, the defect that becomes a customer containment) and the false-reject rate (the fraction of good parts the system rejects, the false alarm that stops the line and burns operator trust). These two pull against each other. A system tuned to catch every escape will reject more good parts; a system tuned to never false-reject will let more escapes through. The PoC has to measure both, in dollars, because they are both dollars.

Target X is the threshold that makes the project worth doing, and you set it from the loss, not from the vendor's brochure. If your current manual inspection catches 90 percent of escapes and false-rejects 3 percent of good parts, the system has to beat that on at least one axis without losing ground on the other, or it is not an improvement, it is a lateral move with a capital cost. Write down: "Catches at least 97 percent of seeded defects with a false-reject rate at or below 2 percent, measured on the holdout set." Now you have a number the vendor cannot argue with after the fact.

Method Z is how you measure, and this is where most PoCs quietly cheat. The only honest method is a holdout test: a set of parts, including known good parts and known defective parts, that the vendor's model never saw during tuning, graded by a human inspector or a coordinate measuring machine independently of the AI, with the AI's calls compared against that ground truth after the fact. If the parts you grade the system on are the same parts the vendor tuned on, you are measuring memorization, not detection, and you will be shocked on the live line. The holdout set is non-negotiable, and you control it, not the vendor.

Period P is long enough to catch the conditions that break systems. A PoC that runs for one shift in good light proves nothing about the night shift, the second resin supplier, the humid week, or the camera mount loosening over time. Run the PoC across at least two weeks and across every shift, deliberately including the conditions you know are hard: the glossier material, the dimmer bay, the changeover, the operator who is new this month. You are not trying to make the system look good. You are trying to find the condition that breaks it, on your dime, before you have spent the capital.

Run It on the Real Line, Not the Booth

The phrase "on a real line" is the whole point of this lesson. A PoC run in a lab, on a bench, or in the vendor's facility is a demo wearing a lab coat. The conditions that decide whether floor AI survives are exactly the conditions a lab strips out: the lighting that changes when the bay doors open on a sunny afternoon, the camera angle that drifts as the mount vibrates, the material variation between two qualified suppliers, the wet or oily part, the dust, the operator who loads the fixture a quarter inch off because they are running behind.

Run the PoC on the actual production line, during actual production, with the actual crew, on the actual parts. If that is operationally impossible because you cannot risk the running line, run it in parallel: the system watches and logs its calls, but the existing inspection process remains in control and ships the parts. Parallel running is the safest PoC architecture on a quality-critical line because it lets you measure the AI's escapes and false-rejects against reality with zero risk to the customer, and it respects the cardinal rule of floor AI: the AI stays advisory until it has earned the right to be trusted, and the human and the existing process stay accountable for what ships.

This is also where the operational technology boundary becomes a hard constraint. The operational technology (OT, the networks and controllers that run the physical plant, as opposed to information technology or IT, the office and data systems) is where the PLC and the line controls live, and 78 percent of OT networks lack centralized monitoring, which means most plants cannot even see what is talking to what on the line. A PoC must never give the vendor's system write access to the line controls. It reads, it logs, it advises. It does not push a reject signal directly to the PLC during the PoC, and arguably not after either until governance is mature. If a vendor's PoC design requires plugging their box directly into your control network with the ability to stop the line, that is a security decision dressed as a productivity decision, and the answer during a PoC is no.

Instrument the conditions, not just the parts

A PoC that only logs pass or fail calls throws away the most valuable data it generates. Log the conditions alongside the calls: the shift, the supplier lot, the ambient light reading if you can get one, the line speed, the operator, the time since last changeover. When the false-reject rate spikes from 2 percent to 9 percent on Tuesday night, the condition log is what tells you it was the glossier resin and not random noise. That is the difference between a PoC that says "it failed" and a PoC that says "it works except on supplier B's gloss finish, which we can fix with a polarizing filter or a separate model, here is the cost." The second answer lets you make a real buying decision. The first answer just makes everyone argue.

The Data Question That Kills Most PoCs

Before you design the test, ask the question that quietly disqualifies more PoCs than any other: do you have labeled data? A vision model that grades parts learns from examples of good parts and defective parts that a human has already labeled correctly. If your defect of interest happens twice a month, you do not have enough defective examples to train or honestly test a model, and no PoC will fix that this quarter. This is one of the written kill criteria for floor AI: no labeled data, no PoC, not yet.

There are three honest paths when data is thin. First, seed defects deliberately: with engineering and quality sign-off, introduce known defects of known types and severities into the PoC stream so you have ground truth to grade against. This is how you measure escape rate honestly, because you know exactly how many bad parts you put in. Second, mine history: pull every defect image, every rejected part record, every scrap photo the plant already has, even if they are scattered across a quality spreadsheet and three operators' phones. Most plants have more labeled history than they think, written down but never read. Third, extend the period: if the defect is rare, the PoC has to run longer to see enough of it, which is a cost you decide on consciously rather than discovering after you sign.

Worked example. A plant wants to catch a hairline crack that escapes maybe four times a month and has historically cost about 60,000 dollars per escape when it reaches the customer as a containment. Over a two-week PoC the line will naturally produce maybe two such cracks, which is not enough to grade an escape rate on. So quality seeds 40 cracked parts of graded severity into the stream across the two weeks, mixed with the natural production. Now the PoC can say: of 42 cracked parts, the system caught 40, missed two, both at the faintest severity grade, while false-rejecting 1.8 percent of the roughly 9,000 good parts it saw. That is a result you can take to a capital committee. "It caught every crack a human could see and missed two that were near the limit of detection, at a false-reject cost we modeled at 14,000 dollars a year against escapes that cost us 240,000 a year." That sentence is what a PoC is for.

Grade It Honestly and Write the Decision Down

When the period closes, you grade against the success sentence you wrote at the start, and you resist the two temptations that ruin PoC discipline. The first temptation is to move the goalposts: the system hit 94 percent escape detection against a 97 percent target, and someone says "that is basically there, let us just buy it." It is not basically there if 97 percent was the number that made the economics work; the gap between 94 and 97 might be exactly the escapes that become containments. The second temptation is sunk cost: you have spent three weeks and built a relationship with the vendor and it feels wasteful to walk away. The three weeks were the price of finding out. Walking away from a system that failed the test is the PoC working, not the PoC wasting your time.

Grade with a confusion matrix in dollars, because precision and recall are not abstractions on the floor, they are money. Lay out four cells: parts the system correctly passed (good, no cost), parts the system correctly rejected (caught defects, value saved), parts the system wrongly rejected (false rejects, each one a stopped-line cost plus the good part scrapped or re-inspected), and parts the system wrongly passed (escapes, each one a potential containment). Multiply each cell by its real dollar value and sum. If the saved-defect value minus the false-reject cost minus the integration and capital cost is positive over the payback period you require, the system passes the economics. If it is negative, it fails, no matter how impressive the demo was.

Then write the decision down with its evidence. A one-page PoC decision memo names the loss, the success sentence, the measured results with the condition breakdown, the dollar confusion matrix, the conditions that broke the system and the cost to fix them, and the go or no-go with the reasoning. This memo is not bureaucracy. It is the artifact that protects you three ways: it justifies the capital to leadership with numbers instead of a vendor's slides, it gives you a baseline to hold the vendor to after purchase ("the PoC showed 1.8 percent false-reject, we are now at 6 percent, what changed"), and it builds the audit trail the customer will eventually ask for, because the customer audits you, not the vendor, and "the model flagged it" is never a sufficient answer to an auditor or a customer.

A worked grading example, all the way to the decision

Bring the four PoC stages together on one concrete case so the grading discipline is unmistakable. A stamping line ships first-pass yield of 96 percent, and the two-point gap below the 98 percent the customer expects is driven by an intermittent burr defect that escapes manual inspection roughly three times a week, each escape costing about 18,000 dollars in sorting, rework, and the occasional customer return. The annual cost of the burr escape is therefore on the order of 2.8 million stamped parts producing enough escapes to total close to 280,000 dollars a year. That is the loss the PoC has to move.

The success sentence written before the vendor arrived: "This PoC succeeds if the system catches at least 96 percent of burr defects (metric: escape rate against seeded and naturally occurring burrs) with a false-reject rate at or below 1.5 percent (metric: false rejects against the good-part stream), measured on a holdout set the vendor never tuned on and graded by the lead inspector and a coordinate measuring machine, over three weeks spanning all three shifts including the high-speed run and the secondary-supplier coil." Every blank is filled, and the sentence can fail.

The PoC runs. Quality seeds 90 graded burrs across the three weeks because natural production would only throw about 9 in that window, far too few to grade an escape rate honestly. At the close, the holdout grading shows the system caught 93 of 99 total burrs, an escape rate miss of 6 percent against the 4 percent the success sentence allowed, and a false-reject rate of 1.2 percent, comfortably under the 1.5 percent target. The condition log shows that 5 of the 6 missed burrs occurred during the high-speed run on the secondary-supplier coil, where the burr presents differently. The system missed the escape target, but the failure has a name and a likely fix.

Now the dollar confusion matrix decides it, not the disappointment of a missed target. The system catches enough burrs to cut the escape cost from roughly 280,000 dollars a year to roughly 95,000 dollars a year, a saved-defect value near 185,000 dollars. The false rejects at 1.2 percent across the line's volume cost about 22,000 dollars a year in stopped-line moments and re-inspected good parts. Amortized capital and integration run about 70,000 dollars a year over the required payback. Net: 185,000 minus 22,000 minus 70,000 equals roughly 93,000 dollars positive in year one, and that is before fixing the high-speed-coil miss. The decision memo records all of it, recommends a conditional go contingent on the vendor resolving the secondary-coil miss in a follow-up window, and sets 1.2 percent as the false-reject baseline the vendor will be held to after purchase. That is a PoC that told the truth and produced a defensible decision, missed target and all.

The PoC also tests the relationship

A PoC is not only a test of the technology. It is a test of how the vendor behaves when the numbers are not flattering. The right vendor, shown a 9 percent false-reject spike on the glossy resin, digs in with you and proposes a fix. The wrong vendor explains why your test was unfair, why the resin is your problem not theirs, and why the demo numbers are the real numbers. How a vendor responds to a failed PoC condition tells you exactly how they will respond to a production problem after the check clears. That behavioral signal is worth as much as the metrics, and a well-run PoC surfaces it for free.

Key Takeaways

  • A demo is designed to make you say yes; a proof of concept is designed to find out whether yes is the right answer. The vendor cannot do both jobs for you, so you design and grade the PoC, not the vendor.
  • Write the success sentence before you sign: "This PoC succeeds if it achieves X on metric Y measured by method Z over period P." If you cannot fill all four blanks, you cannot fail the test, and a test you cannot fail proves nothing.
  • Measure escape rate and false-reject rate, both in dollars, not "accuracy." Accuracy hides the two failure modes that actually cost money: the escape that becomes a containment and the false alarm that stops the line and burns operator trust.
  • Use a holdout set the vendor never tuned on, graded independently by a human or a measuring machine. Grading the system on the parts it was tuned on measures memorization, not detection, and guarantees a shock on the live line.
  • Run it on the real line across at least two weeks and every shift, deliberately including the hard conditions: glossier material, dimmer light, changeover, the new operator. Run parallel and advisory so the AI stays out of direct line control and the existing process keeps shipping. Never give the vendor write access to the PLC during a PoC.
  • Check the data question first. No labeled data is a kill criterion. Seed defects with sign-off, mine the plant's scattered history, or extend the period, but do not pretend a model can be honestly tested on a defect it has barely seen.
  • Grade against the original success sentence with a dollar confusion matrix, and resist moving the goalposts or surrendering to sunk cost. A system that fails the test you designed is the PoC working, not the PoC wasting your time.
  • Write a one-page decision memo with the loss, the metrics, the condition breakdown, the dollar math, and the go or no-go. It justifies the capital, holds the vendor to the PoC baseline after purchase, and starts the audit trail, because the customer audits you, not the vendor.