Build vs. Buy vs. Partner for Grid AI
Every utility eventually faces the question that looks simple until you try to answer it: should we build this AI capability ourselves, buy it from a vendor, or join something in between? The answer is not "it depends" without a framework. It is a structured analysis tied to your specific capabilities, your regulatory obligations, your data posture, and the rate of change in both the problem and the technology. Get it wrong and you spend several years and significant capital either building software your core team was never equipped to maintain, or locked into a vendor platform that no longer fits your system.
The Spectrum from Build to Consortium
The build-buy-partner question is not binary. It is a spectrum with at least four meaningful positions, and utilities often occupy different positions for different use cases simultaneously.
At one end is the internal build: your utility writes the code, trains the models, operates the infrastructure, and maintains the capability. You own everything and depend on nothing external for the capability to function. The constraint is that you need the staff, the data infrastructure, and the engineering culture to sustain what you build. Most utilities do not have a large enough machine-learning engineering team to maintain sophisticated AI models across multiple use cases. Those that do, typically large investor-owned utilities with dedicated data science organizations, can build selectively for use cases where their operational data is a genuine moat and where vendor solutions do not yet reach the required performance level.
Just inside the build end is the "build with tools" position: your utility uses open-source frameworks (TensorFlow, PyTorch, scikit-learn) or cloud AI infrastructure (AWS SageMaker, Azure Machine Learning, Google Vertex AI) to build custom models without writing everything from scratch. This reduces the development burden but maintains control over the model and data. It is appropriate when your use case is well-defined, your team has at least a small ML engineering capability, and no vendor product adequately addresses your specific requirements.
In the middle is the configure-and-extend model: you purchase a vendor platform that provides the underlying AI infrastructure and reference models, and your team configures it for your specific system, feeders, and load profiles. The vendor maintains the platform; you maintain the configuration. This is the most common operational posture for utilities that have made AI investments, because it provides a faster path to production than building while preserving significant customization capability.
At the buy end is a fully managed vendor solution: you purchase a platform, provide the data, and the vendor manages the model training, deployment, and maintenance. You interact with outputs, not with the model internals. This is appropriate for use cases where vendor expertise substantially exceeds what your team can build, where the use case is not strategically differentiated, and where your primary need is outputs rather than capability ownership.
The partner model, including utility consortia, joint ventures, and ISO-led shared platforms, occupies a different dimension. Instead of buying from a vendor or building internally, you co-invest with peer utilities to develop or operate shared AI capabilities. The EPRI utility consortia model and several ISO-driven queue-automation initiatives represent this approach. It combines the cost-sharing benefits of buying with more control over requirements and intellectual property than a standard vendor purchase.
When Utilities Build: The Four Conditions
The decision to build is appropriate when four conditions are simultaneously true. First, you have a genuinely unique data advantage: your system has operational characteristics, customer mix, or DER profile that is sufficiently different from the vendor's reference customer base that a vendor-trained model will never perform as well as one trained exclusively on your data. Second, you have or can hire the engineering capability to build and sustain the model. Building a load forecasting model once is much easier than maintaining and retraining it as your load profile changes over several years. The maintenance capability is often underestimated. Third, the use case is strategically differentiating: it provides a competitive or operational advantage that you do not want to share with peer utilities who might buy the same vendor platform. Fourth, the use case is stable enough to justify the development investment: rapidly evolving AI capabilities mean a use case where vendor products are advancing quickly may be better addressed through purchase than through internal development that will be obsolete before it is deployed.
The classic build case in utility AI is the site-specific load forecasting model. A large IOU with many years of AMI data, a complex DER mix, and a rapidly evolving industrial customer portfolio may have a data asset that is genuinely richer than any vendor's reference data. A model trained exclusively on that utility's data, by engineers who understand the operational context of every unusual load pattern, may perform better than a vendor's model fine-tuned on the same data. This is a legitimate build case.
The classic bad build case is when a utility's project team decides to build a topology optimization engine because they are uncomfortable relying on a vendor for something as important as switching recommendations. The discomfort is understandable. The build decision is not: topology optimization requires deep integration with power-flow solvers, real-time state estimation, and N-1 contingency modeling. These are engineering challenges that specialized vendors have been working on for years. A utility without a dedicated power-systems AI research team will not build a better topology optimizer than a focused vendor in a shorter timeframe at a lower cost.
When Utilities Buy: The Four Signals
The decision to buy is appropriate when vendor capability has reached a level of maturity that exceeds what your team can build in a reasonable timeframe, when the use case is not strategically differentiating, when your team's comparative advantage is in operating and verifying AI outputs rather than building models, and when the procurement process includes adequate portability protections so that buying does not mean permanent lock-in.
The strongest buy signals are use cases where: the vendor category is mature and competitive (multiple vendors competing on performance produces better products than any single utility could build); the vendor's training data advantage is substantial (a DER orchestration platform that has learned from thousands of device types across dozens of utility deployments has a training data asset no single utility can replicate); the use case requires capabilities outside the utility's core competency (sophisticated computer vision for vegetation management, for example, is not a core competency of most distribution planning teams); and the operational support requirement is 24/7 availability that the utility does not staff for software maintenance.
The vendor category map taught throughout this program, with AutoGrid (now part of Uplight), Itron, GE Vernova GridOS, Schneider EcoStruxure, New Grid, Emerald AI, ABB, IBM, and Virtual Peaker as orientation markers, represents the buy side of the spectrum. Each of these vendors has invested years of product development in specific categories. No utility builds what they have built at the scale they have built it. The buy decision is the right decision for the majority of utility AI use cases, provided the procurement process is rigorous.
Build where you have a data moat and engineering depth. Buy where the vendor's training data advantage and platform maturity exceed what you can build in your procurement window. Partner where the cost and data requirements exceed any individual utility's capacity and the problem is shared across the industry.
When Utilities Partner: The Consortium Logic
The partner model makes sense when two conditions hold simultaneously: the problem requires more data than any single utility has, and the cost of building the solution exceeds what any single utility should absorb for what is effectively a shared industry problem. Queue-study automation at scale is a good example. The interconnection queue has 2,060 GW and growing. The AI models that will most effectively accelerate queue studies need to have seen thousands of interconnection applications with diverse project types, grid configurations, and study outcomes. No single utility's queue data is large enough to train a high-performing model in isolation. A consortium of utilities contributing application data and study outcomes, co-developing the automation platform, and sharing both the cost and the intellectual property is a superior structure for this use case.
The EPRI model for utility AI consortia reflects this logic. EPRI operates as a shared research organization for utilities. Its white papers, research programs, and early-stage AI capability development represent a kind of consortium investment where each utility's membership fee funds research that all members share. The limitation of the EPRI model is that it produces research artifacts, not deployable production systems. Consortia that want production systems need a different governance structure, typically a joint procurement of a shared vendor platform with co-investment in customization, or a jointly owned utility subsidiary that operates the shared capability.
ISO and RTO-led shared platforms represent another form of the partner model. Several ISOs have deployed AI-assisted queue management platforms that all utilities in their footprint use for first-pass interconnection study support. In these cases, the ISO is effectively the "builder" on behalf of the utility community, and the utilities are contributing data and funding through their interconnection study fees. This model works well for use cases where the ISO's system-level view is necessary for the AI to be useful: a queue-study automation platform that works across all utilities in a region needs access to the full regional network model, which only the ISO controls.
The governance challenge of the partner model is proportional representation of stakeholder interests. A consortium where large IOUs dominate the steering committee will prioritize use cases and requirements that fit large IOU operations, potentially at the expense of smaller co-ops and municipal utilities whose data patterns and operational needs are different. Successful utility AI consortia explicitly address this governance challenge in their founding documents, typically by setting steering committee voting weights proportional to data contribution rather than equity contribution.
Applying the Framework: Four Use Cases
The framework applies differently to each of the four major grid AI use case categories. For load forecasting, the decision typically favors buy-and-configure for most utilities, with build reserved for utilities with large AMI datasets, significant behind-the-meter solar penetration, and unique industrial customer mixes that vendors cannot adequately represent. For topology optimization, the decision almost always favors buy, because the engineering depth required to build a production-grade topology optimizer with correct N-1 modeling is beyond the capability of most utility planning teams, and several vendors have reached maturity in this space. For DER orchestration, the buy decision is typically correct for the device-layer dispatch infrastructure, but partner or build may make sense for the utility-specific portfolio optimization layer that sits above it. For queue-study automation, the partner model is increasingly the right answer as ISOs develop shared platforms, with individual utility builds reserved for utilities that operate outside ISO footprints or have unusual queue characteristics that shared platforms do not address.
There is a fifth category that complicates the framework: document drafting, regulatory support, and compliance documentation assistance. These use cases sit entirely in the IT domain, require no OT data access, and can be addressed with general-purpose AI tools rather than purpose-built grid AI platforms. For this category, the buy decision is typically correct using general-purpose AI tools, with configuration focused on prompting and context rather than model training. No utility should be building their own large language model for regulatory document drafting.
Worked Example: A Co-op Consortium Outsmarts the Market
Twelve rural electric cooperatives in a single state are all facing the same problem: their AMI infrastructure is generating interval data that their existing forecasting tools cannot fully leverage, and their rooftop solar penetration is growing faster than their models can track. Each co-op has roughly 40,000 to 80,000 meters. Individually, none of them has enough data to train a high-performing AI forecasting model. Together, they have 600,000 to 900,000 meters of interval data covering a diverse geography and customer mix.
They approach this problem as a consortium. They retain a neutral technology advisor (not a vendor) to evaluate the build-buy-partner options. The analysis finds: buying the same vendor platform twelve times would cost each co-op significantly more in licensing fees than their usage volume justifies; building individually is not feasible given that none of the co-ops has a data science team; and a consortium can negotiate better licensing terms, contribute pooled data to a shared model, and split the implementation cost.
The consortium negotiates a shared platform agreement with a forecasting vendor, contributing pooled AMI data for model training while negotiating that each co-op's data remains its own intellectual property and is never shared with competing utilities. The vendor trains a regional model on the pooled data that performs significantly better than what any individual co-op's data alone would support. The consortium's governing agreement specifies that model weights trained on the pooled data are owned jointly by the consortium members, not by the vendor, and that any member can export their share of the model configuration upon exit.
Two years later, the consortium model is producing day-ahead forecasts at 1.8% MAPE on the pooled validation set, compared to the 3.5% the co-ops were averaging with their prior statistical models. The cost per co-op is 40% lower than the individual licensing price the vendor offered to the co-ops before the consortium negotiated. The build-buy-partner analysis took three months and cost less than the first year of licensing fees. It was the highest-leverage planning decision the consortium made.
Key Takeaways
- Build when you have a genuine data moat, sustainable engineering depth, and a strategically differentiating use case. Do not build what vendors have spent years developing at scale, such as topology optimization or device-layer DER dispatch.
- Buy when vendor capability exceeds what you can build in your procurement window, when the use case is not strategically differentiating, and when portability provisions protect you from lock-in. Most utility AI use cases fall in this category.
- Partner when data requirements exceed any single utility's capacity and the problem is genuinely shared across the industry, such as queue-study automation that requires regional network models or AI models that need pooled data to perform adequately.
- The partner model requires explicit governance on data ownership, IP, and voting rights. A consortium dominated by large IOUs will not serve co-ops and municipal utilities well unless voting is weighted by data contribution rather than equity.
- The build-buy-partner decision should be made separately for each use case, not applied as a uniform philosophy. A utility may build its load forecasting model, buy its topology optimizer, and partner for queue automation, and all three decisions can be correct simultaneously.
- Document drafting and regulatory support AI tools belong in a different evaluation framework from operational grid AI. They are general-purpose IT tools, not specialized grid AI platforms, and should be procured accordingly.
- The consortium model, demonstrated in the co-op example, can unlock AI capability for utilities that individually lack the data volume or budget to access it, while preserving data ownership and IP rights through careful governance design.
Skill.re