AI for Energy & Utilities
Strategic · M10 · lesson 10 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Evaluating Grid AI Vendors Without Buying Lock-In
📖
now learning

Evaluating Grid AI Vendors Without Buying Lock-In

15 min

A vendor walks into your boardroom with a live demo: a slick map of your service territory, congestion clearing in real time, a forecast accuracy number that beats your ARIMA model by three percentage points. The demo is polished. The price is large. And somewhere in the contract is a data-export clause that will cost you six figures to exercise. This lesson is your defense.

The Vendor Landscape as Orientation, Not a Shopping List

Before you can evaluate a grid AI vendor, you need a mental map of the space. The vendor category map taught in this program is an orientation tool, not an endorsement ranking. Each category solves a different problem, and a utility that conflates them ends up buying a forecasting platform when it needed a topology optimizer, or paying for a demand-response orchestrator when its actual constraint is queue-study throughput.

Here is how to read the map. Demand-side and DER orchestration platforms (AutoGrid, now part of Uplight; Virtual Peaker; Itron's DER management line) focus on aggregating flexible load, managing dispatch signals, and settling demand-response events. Their differentiation is usually in device-layer communication breadth and the speed of response cycles. Grid operations and topology optimization platforms (New Grid, Schneider EcoStruxure Grid, GE Vernova GridOS) work at the transmission or distribution system level, recommending switching sequences or redispatch actions to relieve congestion. Their core value proposition is N-1 simulation depth and operator-interface fidelity. Integrated utility analytics suites (IBM Maximo-adjacent offerings, ABB Ability) tend to wrap asset health, work management, and some forecasting into enterprise platforms that already live in your IT environment. Specialized grid AI firms (Emerald AI, which built the flexible-load optimization work with National Grid as a reference deployment) occupy narrower niches but sometimes represent the most aggressive capability advances precisely because they are not carrying legacy product lines.

None of those descriptions should be read as a recommendation. What they should do is help you define the category you are actually shopping in before a vendor tells you that their platform "does everything." A platform that does everything usually does several things adequately and nothing exceptionally. For a regulated, reliability-first utility, adequate is not good enough when the question is real-time switching or peak-day procurement.

The Questions That Separate Capability from Demo

The demo environment is always perfect. It runs on clean, pre-loaded historical data. It does not have the telemetry gaps your EMS actually produces at two in the morning. It does not have the three substations whose SCADA integration has been "in progress" for eighteen months. It does not have the 400 MW data-center interconnection that went live last quarter and broke your net-load profile in ways the vendor's reference dataset has never seen.

The capability questions that matter are the ones the vendor's sales team has no incentive to raise. You have to raise them yourself.

Data Provenance and Integration Questions

Start with data. Ask: what are the specific data feeds your platform requires, and what happens when one is missing or delayed? A vendor who says "we handle any data quality" is not answering the question. Push for a concrete answer about degraded-mode behavior. An AI load forecasting platform that silently falls back to a prior-day average when a weather feed is late is a very different risk profile from one that flags the gap, alerts the forecaster, and holds the current forecast pending review.

Ask about the training data geography and vintage. If the platform was trained on ISO New England data and you operate in the Desert Southwest with a very different summer peak profile and a rapidly growing behind-the-meter solar fleet, you need to understand how much fine-tuning was done on your system and whether there is ongoing retraining as your load shape changes. The data-center step-load problem, where a model trained on gradual historical growth encounters a 300 MW overnight interconnection, is the extreme version of this. But smaller mismatches in training data destroy forecast accuracy in ways that do not show up in a vendor's reference-customer MAPE until you are deep into a pilot.

Ask for the data dictionary. Every integration point your utility would need to wire up should be documented in a vendor-provided schema. If the vendor cannot hand you that document in a pre-sales conversation, you have learned something important about their operational maturity.

Model Interpretability and Override Questions

The operator accountability boundary runs through this question. When your topology optimization platform recommends a switching sequence, can an operator see why? Not a feature-importance ranking buried in a PDF, but a human-readable rationale accessible in the same interface where the recommendation appears. The operator who has 90 seconds to accept or reject a congestion-relief action cannot open a separate interpretability dashboard.

Ask: what is the override mechanism, and how is it logged? An AI recommendation that gets overridden by an experienced operator is not a failure, it is the system working correctly. But the override has to be logged, timestamped, and attributable. That log is your compliance evidence. If the vendor's platform does not support structured override logging, that is a procurement risk.

Ask whether the model explains its confidence. A forecast with a 95% confidence interval that is genuinely tight is useful. A forecast that always shows 95% confidence regardless of input quality is dangerous. Vendors who cannot show you real confidence-interval behavior on degraded-input scenarios have not tested the thing they are selling you.

Lock-In and Portability Questions

Lock-in in grid AI is not primarily about software licensing, it is about data. If the vendor ingests your EMS/SCADA telemetry, your asset registry, your weather-station feeds, and your historical interval data for model training, and then stores that in a proprietary format or a cloud environment you do not control, you have created a dependency that will cost you to exit.

Ask for the data-export clause to be shown to you in the contract before the pilot ends, not after. Ask specifically: in what format can our data be exported, on what timeline, and at what cost? Ask whether the trained model weights are yours or the vendor's intellectual property. In many contracts, the vendor retains the model because the weights were trained on their base model that happens to be fine-tuned on your data. That fine-tuned model has significant value and you may not be able to take it with you.

Ask about API openness. A platform that surfaces all outputs through documented, versioned APIs gives you the optionality to substitute components later. A platform that requires you to use its proprietary front-end, its proprietary alerting layer, and its proprietary reporting module is creating dependency at every layer.

The test of a vendor is not how well their demo works on their data. It is how clearly they can tell you what happens when your data is missing, late, or wrong.

Regulatory and Compliance Due Diligence Before You Buy

Grid AI procurement in 2026 carries regulatory obligations that did not exist five years ago. CIP-003-9, enforceable April 1, 2026, specifically addresses vendor electronic remote access and supply-chain security for low-impact BES Cyber Assets, a significant expansion that directly implicates how AI vendors connect to your OT systems. CIP-012-2 (effective July 1, 2026, adding availability requirements to the confidentiality and integrity protections established by the prior CIP-012-1) protects real-time operational data communicated between control centers. Neither of those obligations disappears because you purchased a vendor platform that handles data flows you used to handle internally.

The critical framing is this: compliance accountability does not transfer with the contract. You are the registered entity. NERC cites you, not the vendor, if a finding emerges. A vendor who tells you in the sales process that their platform is "CIP-compliant" is making a marketing statement, not a legal commitment. What you need is a documented understanding of exactly where the vendor's system sits in relation to your Electronic Security Perimeter (ESP), whether it receives or transmits data that falls within your CIP scope, and what controls they have in place to meet your obligations, not their own.

During due diligence, ask the vendor to map their system against your ESP and your Protected Cyber Assets (PCAs). If they cannot do this before contract signing, you are taking on undocumented compliance risk. Ask whether they have undergone a third-party CIP audit and whether they will share the findings under NDA. Ask for their incident-response process specifically for a breach involving data within your CIP scope.

For platforms that touch real-time operational data, CIP-012-2 requires that data communicated between control centers be protected for confidentiality, integrity, and availability. If the vendor's platform sits in a data pathway between your control center and an ISO or between two of your own operating centers, you need to document that pathway and confirm the controls are in place. The fact that the vendor "takes care of encryption" is not sufficient documentation for a NERC audit.

The Reference-Customer Discipline

Every vendor has reference customers. Not every vendor has reference customers who are honest with you about what did not work.

The reference conversation you should request is not a vendor-arranged call with the customer's director of digital transformation who championed the purchase. You want to speak with the person who had to integrate the platform into their EMS, the person who ran the first failing pilot, and the compliance lead who had to document the system for a NERC audit. Those conversations often require going around the vendor. That is appropriate diligence.

When you speak with reference customers, ask specific questions rather than open-ended ones. How long did the integration actually take versus what the vendor said? What data quality problems surfaced in the first six months that the vendor had not prepared you for? Has the platform ever produced an output that was wrong in a way that could have caused an operational incident if an operator had acted on it without verifying? If the reference customer has never seen a wrong output, either the platform is extraordinary or the customer is not using it for anything consequential enough to matter.

Ask about support quality after go-live. The vendor's pre-sales engineering team is always excellent. The post-sale support organization is where commitments get tested. Downtime during a peak-day event, a missed forecast that drove an expensive procurement, a SCADA integration that dropped telemetry for four hours without alerting anyone. Those are the stories that tell you whether a vendor is a partner or a product vendor.

A Scoring Framework That Survives a Reliability Review

Vendor evaluation at a utility is not a pure procurement exercise. It is also a reliability and compliance exercise. That means the scoring framework has to include dimensions that a procurement team working from a standard software-acquisition rubric would not include.

A workable framework has five dimensions. Technical capability is the first: does the platform actually do what you need on your data, verified through a structured pilot or at minimum a vendor demonstration on your actual historical data, not their reference data? The second is integration readiness: can the platform connect to your EMS, ADMS, OMS, and GIS without requiring you to build a custom integration layer that will break every time you upgrade your operational systems? The third is compliance posture: can the vendor demonstrate their security controls against your CIP scope, and are they willing to sign a contractual commitment that includes their obligations in a compliance finding? The fourth is vendor viability: is this a company that will exist in five years? Grid AI firms are raising capital at high valuations on thin revenue bases. Evaluate whether the vendor's business model is sustainable without continuing to raise money. The fifth is portability: can you exit without losing your data, your trained model, and your institutional knowledge of how the platform was configured?

Weight these dimensions by the risk profile of the use case. For a forecasting tool that provides advisory output to a human planner, you can weight technical capability highest because the human is still the last line of defense. For a topology optimization platform that makes switching recommendations inside an operational timeframe, compliance posture and integration readiness move to the top because the consequences of an integration failure in a high-load event are not recoverable with a spreadsheet backup.

Worked Example: The Topology Optimizer That Almost Locked Us In

Consider this scenario, which is constructed from patterns that appear repeatedly in utility procurement. A mid-size IOU in a rapidly growing service territory is experiencing peak-hour congestion on three transmission corridors simultaneously. The traditional approach, operator judgment plus pre-studied switching sequences, is hitting its limits as the number of concurrent DER injections and large-load events grows. The utility decides to evaluate grid-topology optimization platforms.

Two vendors advance to a structured evaluation. Vendor A has a polished platform with impressive reference customers at larger utilities. Their demo runs on their reference data and shows 4% reduction in congestion costs. Vendor B has a more modest interface and a shorter reference list but is willing to run their demo on six months of the utility's actual EMS telemetry, including the three months immediately following the first major data-center interconnection.

The evaluation team runs both demos. Vendor A produces clean results on their reference data. Vendor B's demo on the utility's actual data surfaces a problem immediately: the platform's congestion model has trouble with the step-load pattern from the data center because its training data predates that type of load. Vendor B's team identifies this within the first two hours and proposes a retraining approach. Vendor A, asked to run the same scenario, declines to demonstrate on the utility's data until after contract signing.

That single data point, the willingness to test on real data before the contract is signed, is more diagnostic than any feature comparison. The evaluation team also discovers during contract review that Vendor A's agreement includes a clause that the trained model weights are Vendor A's intellectual property and cannot be exported. Switching vendors would mean retraining from scratch on utility data. Vendor B's agreement includes a data-portability clause and an explicit commitment that the model weights trained on utility data are the utility's property.

The utility chooses Vendor B. The integration takes four months longer than projected, and the first six months of operation require significant configuration work as the platform learns the utility's specific data patterns. But by month twelve, the platform is producing consistent results on real operational scenarios, and the utility has full portability rights. Two years later, when Vendor B is acquired by a larger firm and the support quality declines, the utility is able to migrate their trained model to a new platform without losing their investment in data preparation and model tuning.

The lesson is not that smaller vendors are better. The lesson is that the questions you ask before signing the contract determine whether you retain optionality after signing it.

Key Takeaways

  • The vendor category map (AutoGrid, now part of Uplight; Itron; GE Vernova GridOS; Schneider EcoStruxure; New Grid; Emerald AI; IBM; ABB; Virtual Peaker) is an orientation tool. Know which category solves your actual problem before the vendor tells you their platform solves everything.
  • The diagnostic question is not "how does your demo perform?" but "what happens when our data is missing, late, or wrong?" A vendor who cannot answer that concretely has not tested operational failure modes.
  • Lock-in in grid AI is primarily about data and model weights, not software licensing. Get the data-export clause and intellectual property terms reviewed before the pilot ends, not after.
  • CIP compliance accountability does not transfer with the contract. A vendor claiming their platform is "CIP-compliant" is not a defense in a NERC audit. Map the vendor's system against your ESP and PCAs before signing.
  • Reference customers should be interrogated on what failed, not on what worked. The most useful reference is the person who ran the first failing pilot or had to document the system for an audit.
  • Weight your evaluation dimensions by use-case risk: compliance posture and integration readiness outweigh raw capability for operational systems; capability can lead for advisory planning tools.
  • The willingness of a vendor to demonstrate on your actual historical data, including the hard scenarios, before contract signing is one of the most diagnostic signals available to a procurement team.