โ†
AI for Manufacturing
Strategic ยท M9 ยท lesson 9 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Evaluating Floor-AI Vendors Without Lock-In
๐Ÿ“–
now learning

Evaluating Floor-AI Vendors Without Lock-In

15 min

The demo ran flawlessly. On the projector in the plant conference room, the vendor's vision system caught every cosmetic defect on the sample parts they had brought, drew a tidy red box around each scratch, and flashed a green light on the good ones. The first-pass yield number on their slide read 99.4 percent. The sales engineer was sharp, the dashboard was beautiful, and the VP who had toured a competitor's "smart factory" last month was nodding. Then the plant's own quality engineer asked a single question: "Can I bring three of our actual reject parts from last night's shift, the wet ones off Line 4, and run them through right now?" The room went quiet. The vendor said the parts would need to be "added to the training set" first, that the model was "tuned for the demo dataset," and that a proper deployment would take "a few weeks of data collection." That pause, that one honest hesitation, told the quality engineer more than the entire forty-slide deck. The demo was a polished pilot. It was not a system that would survive her real line. This lesson is about how to tell the difference before you sign, because the cost of finding out afterward is measured in dead pilots, stranded data, and a crew that stops trusting the green light. The single skill that protects you is the ability to ask the question that breaks the demo, and to read the answer for what it reveals about lock-in.

Why Lock-In Is the Real Risk, Not the Price Tag

When a plant evaluates a floor-AI vendor, the instinct is to compare prices, feature lists, and accuracy claims side by side, the same way you would buy a torque wrench or a pallet jack. That instinct is wrong, and it is the most expensive mistake in the room. The price on the quote is the smallest number you will ever pay. The real cost is what it takes to leave, and the vendors who understand this design their product so that leaving is unthinkable.

Lock-in is the situation where switching away from a vendor, or even just integrating a second tool alongside theirs, costs so much in money, time, and risk that you stay even when the product stops serving you. On the plant floor it shows up in five specific places, and a vendor can lock you in through any one of them. Data lock-in is when the images, sensor traces, labels, and quality records your line generates get stored in the vendor's proprietary format inside their cloud, and you cannot export them in a usable form. You spent two years labeling 40,000 defect images, and the day you want to switch you discover the labels live in their schema and walk out the door with them. Model lock-in is when the trained model itself, the thing your data paid to create, belongs to the vendor and cannot be moved, inspected, or retrained anywhere else. Integration lock-in is when the connectors into your MES (the Manufacturing Execution System, the software that tracks what gets built on the floor and when), your historian (the database that records every sensor tag over time), and your CMMS (the Computerized Maintenance Management System, where work orders and maintenance history live) are custom, undocumented, and only the vendor can maintain them. Workflow lock-in is when your operators, quality engineers, and maintenance techs have built their entire daily routine around one vendor's screens, so the switching cost is retraining a whole crew. And contract lock-in is the multi-year auto-renewing agreement with a termination clause that costs more than a year of the subscription.

Here is the worked example that makes this concrete. A mid-market plant signs a three-year deal for a vision-QA system at 8,000 dollars per line per month across four lines, so 384,000 dollars per year. The system works adequately for eighteen months. Then a better, cheaper option appears, or the vendor doubles the price at renewal, which happens. The plant does the math on leaving and finds: 40,000 labeled images locked in a proprietary format that a competing vendor will not accept, so 14 months of labeling labor to recreate, valued at roughly 90,000 dollars. Three custom MES connectors only the original vendor can touch, so a 60,000 dollar integration rebuild. A trained model the plant cannot extract, so the new vendor starts from zero and the plant eats three months of degraded false-reject performance during reboarding. And a crew of 22 operators across shifts who need retraining. The "switching cost" the plant never put on the original quote is north of 200,000 dollars and a quarter of disrupted quality. So the plant renews at the doubled price. That is lock-in, and it was decided not at renewal but at signing, when nobody asked where the data and the model would live.

The price on the quote is the cheapest number in the deal. The real cost is what it takes to leave, and that cost is decided the day you sign, not the day you want out.

The reason lock-in is the dominant risk in 2026 specifically is the talent cliff underneath everything. With roughly 2 million manufacturing workers needing AI reskilling against about 500,000 unfilled roles, and 85 percent of manufacturers saying staffing shortages are already hurting product quality, you do not have a spare team to rip and replace a failed AI deployment. The labor to recover from a bad vendor choice is exactly the labor you do not have. A greenfield plant with a full digital team can absorb a switching project; greenfield sites deploy AI 40 to 60 percent faster than brownfield for this reason. Your brownfield plant, running a 1990s PLC (the Programmable Logic Controller, the industrial computer that actually runs the machine) and a historian nobody has queried in years, cannot. So the cost of a wrong choice is higher for you than for the glossy reference customer in the vendor's case study, which means your evaluation has to be sharper, not more trusting.

The Category Map: Knowing What You Are Actually Buying

Before you can evaluate a vendor you have to know which category they live in, because vendors blur the lines on purpose. A salesperson will happily let you believe their narrow vision tool is a "plant AI platform" if it closes the deal. The category map below is the first filter. It tells you what a given product can and cannot do, and therefore what questions even apply.

Category one: point solutions. These do one job well. A vision-QA system that grades parts at line speed. A predictive-maintenance (PdM, using sensor data to flag a machine trending toward failure before it breaks) tool that watches vibration on your spindle bearings. A scheduling optimizer. The strength of a point solution is depth and a fast time to value; it does the one thing it was built for, and you can often pilot it in weeks. The weakness is that you will end up with five of them, each with its own login, its own data island, and its own integration into your MES, and none of them talk to each other. The integration tax is the hidden cost of a point-solution strategy.

Category two: platforms. These are broad systems that promise quality, maintenance, and operations under one roof, usually from a large automation vendor. The strength is a single integration surface and a unified data model; in theory you connect once and everything inside the platform shares data. The weakness is the deepest lock-in available anywhere in the market, because the platform is engineered to be the whole world. Platform vendors run academies that teach their platform as the destination, not the toolbox. The platform may also be a mediocre version of each individual capability, strong on integration breadth and weak on the depth a point solution gives you in any single domain.

Category three: foundation-model and build-your-own tooling. This is where you use general-purpose AI building blocks, a large language model for knowledge capture and work-instruction drafting, an open vision framework for inspection, and you assemble the workflow yourself or with a systems integrator. The strength is maximum flexibility and minimum lock-in, because the components are interchangeable and you own the assembly. The weakness is that it demands internal capability you may not have on a thin crew, and "build your own" can quietly become "maintain your own forever." For a knowledge-capture or operator-support use case grounded in your own documents, this category is increasingly viable; for safety-adjacent vision at line speed it usually is not, yet.

Category four: systems integrators and consultancies. These are not products, they are the labor that connects the products to your brownfield reality. The strength is that they meet a plant where it is, with the 1990s PLC and the unqueried historian, and they do the unglamorous integration work. The weakness is that strategy-deck-heavy SI engagements can produce a roadmap and a slide library rather than a deployed, working line, and a bad integrator becomes its own form of lock-in when only they understand the spaghetti they built.

The worked judgment here: a plant with one screaming loss, say a defect escape that became a customer containment last quarter, should usually start with a point solution aimed precisely at that loss, not a platform. The defect escape that shipped 4,000 parts is one problem; solve that one with a deep tool, measure the false-reject rate and the yield delta, and earn the credibility to expand. The platform sale tries to reverse this, asking you to buy the whole plant's worth of capability to solve one line's problem, which front-loads cost and lock-in before you have proven a single dollar of value. With 47 percent of manufacturers now using AI in quality, up from 33 percent the year before, the proof points for narrow quality tools are abundant; you do not need to be a platform pioneer to solve a quality escape.

The Demo-Busting Questions

A demo is a controlled performance. The vendor chose the parts, the lighting, the dataset, and the script. Your job in the room is not to admire the performance; it is to break the script in the specific ways your real line will break it, and to read what the answers reveal. These are the questions that separate a vendor who will survive your line from one selling a polished pilot. Ask them in this order, and watch the hesitations.

"Can we run it on my parts, right now, that I brought?" This is the single most powerful question and it is the one from the opening story. The demo dataset is curated. Your reject parts from last night's shift are wet, scratched in unfamiliar ways, lit by your shift lighting, and presented at your camera angle. A vendor confident in their system will say yes and accept a degraded number on unseen parts as honest. A vendor selling a pilot will explain why your parts need to be "added to the training set" first. Both answers are informative. The honest one tells you how the system behaves in week one on your floor, which is the only number that matters.

"What is your false-reject rate, and how was it measured?" A vision system that catches every defect by rejecting half the good parts is worthless, because a false-reject rate (the rate at which the system flags a good part as defective) is real money in scrapped good parts, in operator time spent re-inspecting, and in the operator who, after the tenth false alarm before lunch, disables the green light entirely and the system becomes shelfware. Vendors love to quote recall, the share of true defects caught, because it sounds like 99 percent. Push for both precision and recall, and for the dollar cost of the false rejects at your volume. Worked example: a system catching 99 percent of defects but with a 4 percent false-reject rate on a line running 10,000 good parts a shift is scrapping or re-handling 400 good parts every shift. If a good part is worth 12 dollars in material and labor, that is 4,800 dollars a shift of false-reject cost, potentially more than the escapes it prevents. The false-reject rate is the number that decides whether your operators keep the system on.

"How do you detect and handle drift?" Drift is the slow degradation of model accuracy as the real world changes, the lighting shifts between day and night shift, the camera mount vibrates loose a few degrees over months, the supplier changes the surface finish on the raw material. A vendor who has deployed on real floors has a drift-monitoring answer: they watch the score distribution, they alert when it shifts, they have a retraining cadence. A vendor who has only demoed will treat drift as a surprise. The plant that does not ask this question discovers drift the hard way, three months in, when the false-reject rate has crept up and the operators have already lost trust.

"Where does my data live, in what format, and can I export all of it tomorrow?" This is the data lock-in question, and it must be asked before signing because it cannot be renegotiated after. The good answer is specific: your images and labels live in your cloud tenant or on-premises, in an open format (standard image files plus labels in a documented schema), and you can export everything including the trained model with a documented API. The bad answer is vague, mentions "proprietary optimization," or reveals that the labels and model are the vendor's intellectual property. Remember the 40,000-image, 90,000-dollar relabeling trap; this one question is what prevents it.

"Who can integrate with my MES, historian, and CMMS, and is that integration documented and open?" The integration is where AI becomes real, and it is also where lock-in hides. Ask whether the connectors use open standards (OPC UA for OT data, REST APIs for IT systems, the SQL interface your historian already exposes) or proprietary middleware only the vendor can maintain. Ask whether you or any competent integrator can read and rebuild the connection, or whether you are buying a black box wired into your plant's nervous system. The OT/IT boundary (OT is operational technology, the systems that run the machines; IT is information technology, the business systems) is a hard constraint here, and a vendor who waves it away has not deployed in a real plant.

"Show me a reference customer in my industry, my plant size, and my brownfield reality, and let me call them." A reference customer who is a greenfield enterprise flagship tells you nothing about your brownfield job shop. Ask for a customer with your constraints: the old PLC, the thin crew, the messy data. Then ask that customer the real questions: what broke in the first month, what did the integration actually cost versus the quote, and what is the false-reject rate today, not on day one. The vendor's chosen references are a performance too; the value is in the unscripted phone call.

Reading the Vendor on Accountability and the OT Boundary

There is one answer that should end an evaluation, and it is the answer that tries to take accountability off your plate. A vendor who says, in any form, "our model handles the quality decision so you don't have to worry about it," has just disqualified themselves, because they either do not understand your world or they are willing to mislead you about it. The cardinal rule of the floor is that the customer audits you, not the vendor. When your automotive customer runs an IATF 16949 audit (the global automotive quality management standard), or your aerospace customer runs AS9100, the auditor stands in your plant and asks your quality engineer why a part was passed or failed. "The model flagged it" is never a sufficient answer. Accountability for an AI-touched quality decision stays with the plant and the human who signs the record, and no contract clause transfers it to the vendor. So when you evaluate, you are not buying away your responsibility; you are buying a tool you remain responsible for, and the right vendor talks that way.

The audit-trail question follows directly: "When your system passes or rejects a part, what does it log, and can I produce that record for a customer auditor?" The good vendor shows you a per-part record with the image, the score, the threshold, the disposition, and the human sign-off where one occurred. The bad vendor shows you an aggregate dashboard and cannot reconstruct an individual decision. An aggregate dashboard does not survive an audit; a per-decision log does. This is also a governance question about keeping AI advisory rather than in direct control. With 78 percent of OT networks lacking centralized monitoring, you often cannot fully see the environment the AI runs in, which is exactly why AI on the floor should stay advisory and out of direct control of anything that moves until it is properly governed. A safety-critical control loop is not where AI goes first, and a vendor pushing closed-loop autonomous control onto your line before you can even monitor your OT network is selling you risk you cannot see.

The security due-diligence question is the third leg: "How does your system connect to my OT network, what ports does it open, where does data leave the plant, and how is that connection secured and monitored?" A vendor who installs an agent on a machine adjacent to the PLC, opens an outbound connection to their cloud, and cannot tell you exactly what crosses that boundary has just expanded your attack surface in a plant you already cannot fully monitor. The security obligation does not transfer to the vendor either; if their connector becomes the path into your controls network, the incident is yours to answer for. Treat every vendor connection as a new door into the OT environment, and require the vendor to document and justify each one.

Here is the worked scenario that ties accountability to lock-in. A plant deploys a platform vendor's vision system in closed-loop mode, where the system itself rejects parts and diverts them without a human in the loop, because the demo made it look effortless and the thin crew welcomed the labor relief. Six months in, a drift event causes the system to start passing a subtle defect it used to catch. Because the loop is closed and the audit log is an aggregate dashboard, the plant does not notice until a customer containment lands: 4,000 parts shipped with the defect. The customer's auditor asks for the per-part inspection record. The plant cannot produce one, because the vendor's system did not log per-part decisions in an exportable form, and the plant cannot independently reconstruct the model's behavior because the model is locked in the vendor's cloud. The containment cost, the lost customer confidence, and the scramble to add human inspection back into a loop the plant had removed all trace to two evaluation failures: accepting closed-loop control before earning it, and accepting an audit log the customer would never accept. Both were preventable with two questions in the demo room.

Structuring the Evaluation So the Demo Cannot Win Alone

A demo is designed to win the room emotionally before anyone runs the numbers. The defense is a structured evaluation that forces the decision out of the conference room and onto your actual line, with success criteria written down before the vendor knows them. The structure has three stages, and skipping any stage is how plants end up locked into a tool that never proved itself.

Stage one: the written requirement and the scorecard. Before you see a single demo, write down what success means for your specific loss. If the loss is the defect escape, success is a measured recall above your threshold at a false-reject rate your operators will tolerate, on your parts, with an exportable audit log and your data in your tenant. Build a scorecard with weighted categories: fit to your loss, false-reject economics, data and model portability, integration openness, accountability and audit support, OT security, total cost including the switching cost, and reference quality. Weight portability and false-reject economics heavily, because those are the two areas where a great demo hides the worst long-term pain. Score every vendor on the same sheet so the polished demo competes on the same axis as the honest one.

Stage two: the proof of concept on a real line, with exit defined first. The next lesson covers PoC design in depth, so here the evaluation point is narrow but vital: define the success criteria and the exit before you start, and make the PoC use your data on your line, not the vendor's curated dataset. A pilot that runs on the vendor's data proves nothing about your floor. A pilot with no pre-defined success threshold becomes a sunk-cost trap, where the plant keeps extending because nobody wrote down what "good enough to buy" looks like, and the vendor is happy to keep the pilot running because a live pilot is a soft form of lock-in. Write the number down first: "we proceed if the false-reject rate stays under 1.5 percent across two weeks on Line 4 with our parts and our shift lighting." Then the demo cannot override the data.

Stage three: the contract terms that protect the exit. The contract is where lock-in is either prevented or cemented. Insist on a data-portability clause that guarantees export of all your data including labels in an open format on demand. Insist on model portability or, at minimum, a clear statement of what you can and cannot take with you. Avoid multi-year auto-renew with punitive termination; prefer a shorter initial term with the right to renew, so the vendor has to keep earning the relationship. Pin the price for the integration work and require that integration documentation be delivered to you, so a future integrator can maintain it. Each of these terms maps directly to one of the five lock-in types, and each is far cheaper to negotiate before signing than to litigate after.

The worked payback of this discipline: the mid-market plant from the first example, had it run this structure, would have caught the proprietary-format data trap in stage one, required the export clause in stage three, and at the eighteen-month decision point would have been able to leave for the better, cheaper option, taking its 40,000 labeled images with it. The structure does not slow you down meaningfully; the scorecard takes a day, the real-line PoC runs in weeks, and the contract terms are a negotiation you were going to have anyway. What it buys you is the freedom that the talent cliff makes precious: the ability to change your mind without a 200,000-dollar penalty and a quarter of disrupted quality you do not have the crew to absorb.

Key Takeaways

  • The price on the quote is the cheapest number in the deal; the real cost is what it takes to leave, and that switching cost is decided at signing, not at renewal. A plant that never asked where its data and model live can face a 200,000-dollar-plus exit penalty it never put on the original quote.
  • Lock-in hides in five places: data (proprietary formats for your labeled images), model (the trained model belongs to the vendor), integration (custom connectors only the vendor maintains), workflow (a whole crew built around one set of screens), and contract (multi-year auto-renew with punitive termination). A vendor can trap you through any one of them.
  • Know the category before you evaluate: point solutions (deep, fast, but you end up with five data islands), platforms (one integration surface but the deepest lock-in), foundation-model build-your-own (maximum flexibility, demands internal capability), and systems integrators (the labor that meets brownfield where it is). A single screaming loss usually calls for a point solution, not a platform.
  • The demo is a controlled performance; break the script with your parts. "Can we run it on my reject parts from last night's shift, right now?" reveals in one answer whether you are looking at a working system or a polished pilot tuned to a curated dataset.
  • Push past recall to the false-reject rate and its dollar cost. A 4 percent false-reject rate on 10,000 good parts a shift at 12 dollars each is 4,800 dollars a shift, often more than the escapes the system prevents, and it is the number that decides whether operators keep the green light on.
  • Accountability never transfers to the vendor. The customer audits you, not the vendor, so "the model flagged it" is never a sufficient answer under IATF 16949 or AS9100. Require a per-decision audit log you can export, not an aggregate dashboard, and keep AI advisory rather than in closed-loop control until it is properly governed.
  • The OT boundary is a hard constraint and the security obligation stays with you. With 78 percent of OT networks lacking centralized monitoring, treat every vendor connection as a new door into your controls network and require it to be documented, justified, and monitored.
  • Structure the evaluation so the demo cannot win alone: a written scorecard weighting portability and false-reject economics, a proof of concept on your real line with the success threshold and exit written down first, and contract terms (data portability, model portability, no punitive auto-renew, delivered integration documentation) that protect your ability to leave. On a thin crew in a brownfield plant, the freedom to change your mind without a crippling penalty is the whole point.