โ†
AI for Energy & Utilities
Proficient ยท M20 ยท lesson 20 of 20 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Topology Optimization in the Loop
๐Ÿ“–
now learning

Topology Optimization in the Loop

15 min

The alarm sounds at 2:47 a.m.: a 138 kV line in the northeast corner of the service territory has reached 98 percent of its thermal limit, and the day-ahead schedule did not anticipate the temperature inversion that killed the wind forecast. The energy management system (EMS) has already flagged the congestion. The topology optimizer has already ranked four switching sequences that would relieve it. The operator has ninety seconds to decide whether to act, ask for more context, or override entirely. This lesson is about that ninety seconds and everything that must be engineered before it arrives.

What Topology Optimization Actually Does

Topology optimization is the practice of reconfiguring which breakers and switches are open or closed in a transmission or distribution network to redistribute power flows, relieve thermal overloads, reduce losses, or restore service after a fault. Before AI tools entered the picture, an experienced operator would look at a SCADA display, recall years of pattern recognition, consult the switching desk, and manually identify a sequence of switching actions. The knowledge was real, deep, and extraordinarily difficult to document.

AI-powered topology optimization engines (category examples for orientation: New Grid's PowerOps, Schneider Electric's EcoStruxure Grid operations layer, GE Vernova GridOS) work by solving a constrained optimization problem over the current network state in near-real-time. They ingest real-time SCADA telemetry, the network model (the detailed electrical representation of every bus, line, transformer, and breaker), current load and generation, weather data, and N-1 contingency requirements. The output is a ranked list of switching actions with predicted post-switching flows, estimated congestion relief in MW, and a validity timestamp.

The validity timestamp matters enormously. A switching recommendation that was computed thirty seconds ago against a network state that has since shifted is not just stale: it may be actively dangerous. Modern systems refresh recommendations on cycle times of ten to sixty seconds depending on the grid's pace of change. Operators must understand this clock and know to request a fresh recommendation if the system state has changed materially since the last computation.

The Network Model Dependency

Every topology optimization result is only as good as the network model it runs on. The network model is the mathematical representation of the grid that lives inside the EMS, updated continuously from SCADA status points. If a breaker status is wrong (stuck-closed reported as open, or vice versa), the optimizer will propose switching sequences based on a false picture of the network and the operator will face recommendations that do not match what is physically in front of them. This is not a hypothetical. Model-state mismatches are among the most common root causes of operator confusion during topology optimization trials.

Before trusting any switching recommendation, the operator's first question should be: is my real-time network model current and consistent with what I see on the board? This means checking for recent breaker status updates, confirming that maintenance clearances (which take equipment out of service and must be reflected in the model) are properly tagged, and verifying that any recent switching already executed has been captured. A recommendation derived from a stale or incorrect model is a recommendation to be overridden, not executed.

How Congestion-Relief Recommendations Reach the Operator

In a fully integrated deployment, the topology optimizer sits adjacent to the EMS and SCADA environment (not inside the protected cyber asset boundary unless it has met all CIP governance requirements). It watches real-time telemetry through a data feed, computes recommendations continuously, and surfaces them through an advisory panel on the operator display. The panel typically shows the overloaded element, the magnitude of the overload, the recommended switching actions in sequence, predicted post-action loading on all affected lines, and an indication of whether any proposed action creates an N-1 violation elsewhere.

That last point deserves emphasis. A good topology optimizer does not just solve the presenting problem: it checks whether the solution creates a new vulnerability. Relieving a 98 percent loading on Line A by closing a normally open switch that now puts Line B at 91 percent may pass the immediate thermal test while creating a precarious N-1 situation. The operator must be able to see this second-order effect in the display, not infer it from memory. Any recommendation that is silent on downstream contingency exposure is incomplete and should prompt the operator to run a manual contingency check before acting.

The Advisory, Not Autonomous, Boundary

The cardinal rule of topology optimization in 2026 is that the system advises; the operator decides. This is not a philosophical preference: it is a regulatory and safety boundary enforced by NERC reliability standards, utility operating procedures, and the accountability structure of the control room. The operator who executes a switching sequence owns that decision. If the sequence trips a breaker that should not have been operated, or causes a fault, or drops load, the investigation will ask what the operator knew, what they verified, and why they acted. "The optimizer said so" is not a defense. It never has been. It never will be.

The model proposes. The operator decides. The log proves it. If those three sentences do not describe your topology optimization workflow, fix the workflow before you deploy the model.

This boundary has a practical design implication: the operator must have a genuine, friction-reduced override path. If overriding the system requires filling out a lengthy justification form or escalating to a supervisor in real time, operators under time pressure will take the path of least resistance and execute recommendations they have not fully verified. The override mechanism should be a single clearly labeled action, with a brief free-text note field, and no operational penalty for using it. Override rates are a health metric: too few overrides often means operators are not critically evaluating recommendations; too many means the model's quality or operator training needs attention.

Designing the Operator Override Discipline

Override discipline is the structured practice of evaluating a recommendation before executing it, and documenting both the evaluation and the decision. It is a skill that must be trained, not assumed. Consider what a well-trained operator does in those ninety seconds when a congestion-relief recommendation arrives:

  1. Confirm network model currency. Are the recent breaker statuses accurate? Are any clearances tagged? Has anything changed in the last two minutes that would invalidate the recommendation?
  2. Read the full recommendation, not just the top action. What sequence of steps is proposed? What are the predicted post-action flows on all affected elements? What contingency violations, if any, does the system flag?
  3. Apply local knowledge. Is there anything the model cannot know? A crew working near the proposed switching location? A ground rod in place that was not captured in a tag-out? A feeder serving a critical load that is not flagged as such in the model?
  4. Decide and document. Execute, modify, or override. Log the decision and the reasoning, even briefly. This log is your audit trail, your defense in a post-event review, and the training data for improving the model.

Step three deserves a longer pause. Topology optimizers operate on the network model. Local, transient, human-only knowledge sits outside the model. The crew working a distribution lateral that physically cannot be switched without endangering a worker is not in the SCADA telemetry. The critical hospital that an operator knows from experience drops power for three seconds during a certain switching sequence is not coded as a critical load unless someone entered it. The experienced operator carries a map of these invisible constraints. The new operator does not. This is where the Great Crew Change creates real risk: when 25 percent or more of the workforce is retirement-eligible, the institutional knowledge that lives in experienced operators' heads is the most important data asset the utility owns, and it is not in the optimizer.

What the Optimizer Cannot Know

Topology optimizers are powerful precisely because they can compute solutions across hundreds or thousands of switching combinations in seconds, something no human can match. But their knowledge boundary is sharply defined by what is in the network model, the telemetry, and the training data. Several categories of real-world knowledge fall outside this boundary and must be supplied by the operator:

Crew and equipment positions in the field. A switching recommendation may call for operating a switch that has a lineworker standing ten feet away on a job that was dispatched thirty minutes ago. The EMS sees a breaker status. It does not see a human being. Every switching action recommended by an optimizer must be verified against active crew locations, and the dispatch system must be the authoritative source for this check.

Non-standard equipment states. Field equipment may be in a condition that telemetry does not fully capture: an aging recloser that the district supervisor knows is sticky and needs to be operated slowly, a transformer that has been flagged for replacement and is running near the top of its allowed temperature range, a cable section with a known incipient fault that the asset team has been monitoring. These constraints belong in formal tagging and clearance systems, but in practice they often live in oral tradition among experienced crews.

Downstream customer impacts. The optimizer models load flows and thermal limits. It does not model which bus serves the water treatment plant that cannot tolerate a momentary interruption, or which feeder has a 100-unit apartment complex whose residents have already called the service center twice this week. Customer sensitivity data, when it exists, is typically in the customer information system (CIS) or in separate critical-load registries. Integrating those registries into the optimizer's decision framework is a deliberate engineering and data governance project, not an automatic capability.

Weather and environmental conditions at the switching location. A recommended switching action may involve a substation in a flood-prone area that is currently inaccessible by the standard route due to a road closure the EMS does not know about. Or it may involve outdoor switching in lightning conditions that the grid-wide weather feed has not yet flagged at that specific location. The operator's situational awareness extends beyond the display.

Worked Example: A Tale of Two Operators

Consider a scenario that plays out in real control rooms. It is a hot August afternoon. A 115 kV line serving a suburban load pocket has reached 104 percent of its emergency rating following the unplanned outage of a neighboring line. The topology optimizer surfaces two options:

Option A (ranked first by the optimizer): Close normally-open switch SW-247 between Bus 14 and Bus 22, redistributing approximately 45 MW onto a parallel path. Predicted post-action loading on the overloaded line: 87 percent. No N-1 violations flagged.

Option B (ranked second): Redispatch 30 MW of generation on Bus 9 (a gas peaker), reducing flows on the overloaded line to 91 percent. Higher cost; N-1 clean.

Operator 1 has twelve years of experience. She reads Option A, notes that SW-247 is a manually operated switch at a substation that had a maintenance crew working there earlier in the shift. She checks the dispatch system. The crew cleared the job forty minutes ago and has returned to the district office. She calls the substation anyway to confirm no one is near the switch. Confirmed clear. She executes Option A, noting in the log: "Verified SW-247 area clear via dispatch and substation call before closing per optimizer recommendation."

Operator 2 is eight months into the job after coming from a technology background. He sees Option A ranked first and executes it without the crew check. In this scenario, the crew had, in fact, cleared the job and no one was hurt. But the log reads: "Closed SW-247 per system recommendation." In a post-event review, that log entry is a finding: operator executed without documented verification. The recommendation is not wrong; the protocol is broken.

The difference between these two operators is not the optimizer. It is the override discipline: the trained habit of verification before action. Deploying topology optimization without investing in that training is deploying the car without teaching the driver.

Integration with EMS, ADMS, and the OT Boundary

EMS (Energy Management System) and ADMS (Advanced Distribution Management System) are the operational platforms that topology optimization engines integrate with. The EMS handles transmission-level operations; ADMS handles distribution. Both are considered Operational Technology (OT) environments subject to NERC CIP cybersecurity standards at the appropriate BES (Bulk Electric System) applicability levels.

NERC CIP-003-9, enforceable since April 1, 2026, establishes cybersecurity obligations for low-impact BES cyber systems, and CIP-012-2 protects real-time data communicated between control centers. An AI topology optimizer that is integrated into the EMS data feed and whose outputs are displayed in the control room is touching the OT boundary. Whether the optimizer itself is classified as a BES Cyber System depends on whether it performs or directly affects the reliable operation of the BES. That determination is not a software vendor's call: it belongs to the utility's NERC compliance team, documented in the system's applicable system list (ASL) or asset identification process.

The practical approach that most utilities took in early deployments is to run the optimizer as a read-only advisory system: it ingests real-time data but cannot issue any command to the network. Outputs are advisory displays on the operator workstation, not automatic control actions. This architecture keeps the optimizer outside the highest CIP tiers while still delivering operational value. As confidence builds and governance matures, some utilities are exploring tighter integration, but those projects carry compliance obligations that must be addressed explicitly before go-live.

Building a Topology Optimization Governance Framework

Deploying a topology optimizer without governance is like deploying a load forecast without a MAPE target: you have a number, but you do not know if it is trustworthy. Governance for topology optimization covers four areas:

Model accuracy tracking. How often do post-action measurements match the optimizer's predicted flows? A mismatch rate above a defined threshold (typically set by the operations engineering team after initial deployment) should trigger a model review. The network model may have drifted from physical reality due to equipment changes that were not updated in the EMS, or the optimizer's algorithms may need recalibration to new grid conditions.

Override logging and review. Every override should be logged with a reason code and reviewed in a weekly or monthly operations meeting. Patterns in override reasons are diagnostic. If operators are overriding because "I didn't trust the network model," that is a network model currency problem. If they are overriding because "the recommendation would have affected an unlisted critical load," that is a data integration problem. If they are overriding because "the model keeps recommending actions I already know won't work," that may indicate the optimizer's training data is stale or its constraint set is incomplete.

Training and certification. Control room operators, shift supervisors, and the operations engineering team that maintains the network model all need role-specific training. Operators need to understand what the system shows and what it does not show. Engineers need to understand how the optimizer uses the network model and how to keep it current. Supervisors need to understand the governance metrics and what action thresholds mean.

Change management for network model updates. Every time a new line is commissioned, a substation is reconfigured, or a significant piece of equipment is replaced, the network model must be updated. This seems obvious, but in practice it requires a defined process: who is responsible for the update, what is the verification step, and what is the timeline from physical change to model currency? A topology optimizer running on a model that is six weeks behind physical reality is providing recommendations for a grid that no longer exists in that form.

Key Takeaways

  • Topology optimization engines compute ranked switching recommendations by solving constrained optimization over the real-time network model; the quality of the recommendation is bounded by the currency and accuracy of that model.
  • Every recommendation has a validity timestamp; operators must understand the system's refresh cycle and treat stale recommendations as requiring re-computation before action.
  • The advisory-not-autonomous boundary is a hard regulatory and safety line: the operator who executes a switching sequence owns that decision, and the audit log must prove the verification steps taken before execution.
  • Topology optimizers cannot know crew positions, non-standard equipment states, unlisted critical loads, or local environmental conditions; these gaps are the operator's responsibility to fill before acting on any recommendation.
  • Override discipline is a trained skill: a structured habit of model-currency check, full-recommendation review, local-knowledge overlay, and documented decision that must be built into operator training before deployment.
  • Integration with the OT boundary carries NERC CIP obligations that the utility's compliance team must evaluate; most early deployments use read-only advisory architecture to keep the optimizer outside the highest CIP tiers.
  • Governance requires tracking model accuracy against post-action measurements, reviewing override patterns diagnostically, and maintaining a defined process for keeping the network model current after every physical change.