AI for Energy & Utilities
Capable · M18 · lesson 18 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Recognizing Bad AI Output in a Grid Context
📖
now learning

Recognizing Bad AI Output in a Grid Context

15 min

The distribution engineer had been using AI for six months when the output that almost caused a problem landed on his desk. It was a restoration estimate for a storm-damaged feeder, structured like every previous brief he had seen, with a customer count, a circuit ID, and an 8 PM estimated restoration time. The crew on site had no idea what circuit the brief was describing. The substation name existed, but the feeder designation was a plausible-sounding invention. The 8 PM ERT was not from any field note. The brief looked right and was wrong in the ways that would have mattered most during a major storm. This lesson is about building the eye that catches outputs like that before they drive a decision, a dispatch, or a regulatory filing.

Why Bad AI Output Looks So Good

The hardest problem with AI output quality is not identifying obviously wrong answers. It is identifying subtly wrong answers that look authoritative. A generative model is optimized to produce fluent, coherent, professionally formatted text. It has seen more utility documents than most utility professionals read in a career. When it gets something wrong, it gets it wrong in the vocabulary and format of someone who knows what they are talking about. That is what makes the skeptic's checklist necessary: you cannot rely on prose quality or structural polish as a proxy for accuracy.

There are three specific qualities that make bad AI output hard to catch on a quick read:

Internal coherence. A bad AI output is usually internally consistent. The fabricated substation name appears in the same service territory as the real ones. The invented regulatory citation follows the numbering convention of real standards. The wrong unit (MWh instead of MW) produces a plausible-looking number at the right order of magnitude. Internal coherence suppresses the reviewer's instinct to dig deeper, because the output "fits."

Confidence calibration. An AI model does not flag its uncertain outputs with a different tone. The sentence about a fabricated circuit ID reads with the same confidence as the sentence about a real one. Without the cite-or-refuse instruction, there is no linguistic signal distinguishing a verified fact from a plausible invention. The absence of hesitation is not evidence of accuracy.

Plausible magnitude. When an AI produces a wrong number, it rarely produces one that is absurd. A restoration estimate that should be 10 hours comes back as 8 hours or 12 hours, not 48. A capacity number that should be 300 MW comes back as 280 or 330. The order-of-magnitude is usually right. The specific value may be wrong in ways that matter: an 8 PM ERT that is actually 6 AM is a massive customer service and regulatory consequence, but both are plausible-looking numbers.

Understanding these three qualities changes how you approach an AI output review. You are not looking for obvious nonsense. You are looking for the specific categories of error that are most likely in the type of output you are reviewing, using a systematic checklist rather than a general readthrough.

The Skeptic's Checklist for a Load Forecast Output

AI-assisted load forecasts are among the most financially consequential outputs in utility planning. A wrong growth rate propagates through resource planning, interconnection studies, and rate cases. A missed step-load event from a data-center interconnection can produce a capacity shortfall that costs ratepayers hundreds of millions of dollars in emergency procurement. The skeptic's checklist for a forecast output has six items.

Six Items on the Forecast Skeptic's Checklist

  1. Data vintage. What time period is the historical data drawn from? If the model was trained on or asked to reason from data that predates a major load-shaping event (a data-center campus interconnection, a large industrial closure, behind-the-meter solar adoption), the forecast baseline is wrong regardless of how sophisticated the extrapolation is. Ask the AI: "What is the most recent data point in your analysis, and is it from my provided data or your training data?" If the answer is the latter, the forecast is ungrounded.
  2. Step-load exposure. Does the forecast acknowledge the data-center step-load problem? The interconnection queue has over 2,060 GW of backlog at the end of 2025, with approximately 90 GW attributable to data centers. A forecast that extrapolates smooth historical growth without accounting for queued large-load interconnections is systematically wrong for any service territory absorbing these loads. Ask: "Does this forecast include any provision for step-load events from data-center interconnections above 50 MW? If not, what would the five-year peak look like if 300 MW of data-center load comes online in year two?"
  3. Unit confirmation. Is the forecast in summer peak MW, winter peak MW, or annual MWh? These are different metrics with different planning consequences. A summer peak MW forecast drives capacity adequacy analysis; a winter peak forecast may be more important for utilities with high electric heating load; an annual MWh forecast drives fuel and supply planning. If the unit is not stated, ask for explicit confirmation. If the AI says "load growth of 3 percent" without specifying the metric, the number is not actionable.
  4. Net vs. gross load. Does the forecast treat solar generation as reducing load (net load) or ignore it (gross load)? In territories with high rooftop solar penetration, net load at peak can be materially different from gross load. A forecast that mixes these metrics or is unclear which it represents can produce a planning number that overstates or understates the capacity requirement. Ask specifically: "Is this forecast for gross load or net load? Does it account for behind-the-meter generation at the distribution level?"
  5. Uncertainty quantification. Does the forecast include a range or confidence interval? A point estimate without an uncertainty band is not a planning tool; it is a guess. Real forecasting systems produce high, medium, and low scenarios. An AI output that produces a single number without a range is either working from an excessively simplified model or has been asked a question that does not elicit uncertainty. Ask: "What is the high and low scenario for this forecast? What are the two or three factors that most affect which scenario materializes?"
  6. Source traceability. Can every number in the forecast be traced to either (a) the data you provided or (b) an explicitly labeled inference? If the AI produced a 3.5 percent CAGR without showing you where that number came from, it invented it from training data. The right output shows: "Based on the IRP excerpt you provided, the base case growth rate is 1.8 percent. I have added 0.5 percent for data-center load based on the queue data you provided, producing a total growth rate of 2.3 percent. [INFERENCE: I have assumed 50 percent of queued data-center projects complete on schedule, based on historical queue completion rates from DOE research; this should be verified.]"

The Skeptic's Checklist for a Restoration Estimate

Restoration estimates drive crew dispatch, customer communications, and regulatory reporting. An ERT that is four hours too optimistic leads to under-dispatched crews and irate customers; an ERT that is four hours too pessimistic leads to wasted overtime and regulatory scrutiny of why the utility understated its restoration capability. The stakes are high during major weather events, and the AI output that looks most authoritative is often the one produced under the most time pressure.

Five items on the restoration estimate checklist:

  1. Field-note basis for the ERT. Does the ERT in the output trace to an actual field crew note in the OMS data provided? If the prompt included an OMS export and the ERT is for a specific circuit, find the circuit in the OMS data and confirm there is a field note from a crew assigning that time. If the ERT is not in the OMS data, the AI invented it. Write [ERT: NO FIELD BASIS] and make a phone call.
  2. Asset identifier verification. Does every circuit, substation, and equipment designation in the brief match the identifiers in the OMS export? Open the OMS data and find each designation. A designation that does not appear in the OMS data is a fabricated identifier. This is not a rare failure mode; it is one of the most common errors in AI-assisted outage briefs, because the model fills in plausible-sounding identifiers when the provided data is incomplete.
  3. Customer count currency. Is the customer count in the brief the current count or the count at the start of the event? Outage events are dynamic; customers restore as crews work. An AI output that was generated from an OMS snapshot taken four hours ago will have a customer count that is four hours out of date. Ask: "What is the timestamp of the OMS data I provided, and does the customer count in this brief reflect that timestamp or a current estimate?"
  4. Cause code confirmation. Does the stated cause of the outage match the cause code in the OMS data? AI models sometimes substitute a plausible cause (equipment failure during a storm) for the actual cause code in the data (transmission relay operation). The cause code matters for regulatory reporting and for the post-event review that determines whether a reliability event must be reported to NERC.
  5. Communication-safe language. Is the brief suitable for customer communication? Customer-facing language about restoration times must not state a specific ERT unless a crew has confirmed it. An AI output that produces customer communication language including a specific restoration time, without a field crew's confirmation, creates a utility liability if the ERT is missed by more than the state commission's service quality standards allow. Check that any ERT stated in customer-facing language is either field-confirmed or replaced with a range or an "as soon as safely possible" formulation.

The Skeptic's Checklist for a Regulatory Summary

Regulatory summaries are the output category with the highest potential for harm that is hardest to detect. A wrong load forecast number is usually caught when it is compared to other numbers in the planning process. A wrong regulatory citation may not be caught until a NERC auditor or commission staff identifies it, at which point the correction is expensive and public.

Six items on the regulatory summary checklist:

  1. Standard version verification. Is the standard cited in the summary the current effective version? Check the NERC standards website (or FERC's eFiling for tariff references) for the version number and the effective date. If the summary cites CIP-003-8 and the effective standard is CIP-003-9, every compliance conclusion in the summary is wrong. This is a 30-second check that prevents a compliance program failure.
  2. Applicability verification. Does the summary correctly identify which registered entity types are subject to the cited requirements? Pull the applicability table from the standard and confirm. An AI that reads the requirement text without the applicability table frequently assigns obligations to the wrong entity type, producing an analysis that either over-imposes obligations or misses them entirely.
  3. Effective date verification. Does the summary correctly state when requirements became or become effective? CIP-003-9 became enforceable April 1, 2026. The NERC CLE registry category was committed for December 31, 2026. A summary that states the wrong enforcement date for either of these can produce a compliance calendar that misses an obligation or triggers an unnecessary early compliance effort. Verify the effective date against the NERC standards website and the relevant FERC order.
  4. Definition reliance check. Does any conclusion in the summary depend on a defined term? If so, does the summary use the correct definition? Regulatory standards are full of defined terms (Bulk Electric System, Electronic Security Perimeter, Cyber Asset, Reliable Operation) that have very specific meanings in the NERC glossary. A model that uses the common-language meaning of a defined term instead of the glossary definition can produce a completely wrong applicability or compliance conclusion.
  5. Jurisdiction cross-check. Does the summary appropriately separate federal (FERC/NERC) and state (PUC) obligations, or does it blend them? A regulatory summary that attributes a state commission obligation to FERC, or a FERC obligation to the state commission, creates a compliance gap or an incorrect regulatory strategy. Ask: "For each obligation in this summary, which regulatory authority requires it, and under what filing or standard is it established?"
  6. Recent action check. Does the summary account for any regulatory action in the last 12 months? This is the hardest check because it requires the reviewer to know what has changed. For 2026, the relevant recent actions include: FERC large-load interconnection rulemaking (April 2026 deadline), NERC CLE registry commitment (March 2026), NERC CIP-003-9 enforcement (April 1, 2026), and NERC Level 3 Alert on data-center load (May 2026). An AI output that does not mention these developments in a summary about large-load service, CIP compliance, or interconnection policy is likely based on pre-2026 training data and is missing the most important recent developments.

The worst regulatory summary is not the obviously incomplete one. It is the one that covers 14 of 15 recent developments and confidently misses the one that matters most for your filing.

Building a Team Skeptic Culture

The skeptic's checklist is most effective when it becomes a team habit rather than an individual practice. A single analyst with a sharp eye will catch bad AI output for their own work. A team where every member uses the checklist creates an organization where bad output rarely makes it into a filed document, an operational decision, or a customer communication.

Building the team culture requires three elements:

Shared failure documentation. When a team member catches a bad AI output using the checklist, the failure should be documented in the team's failure log with enough detail for every other team member to recognize the same pattern if it appears in their work. The failure log entry should note: the type of error (stale citation, fabricated asset ID, wrong unit), the specific checklist item that caught it, and the consequence if it had not been caught. Over time, the failure log reveals which error types are most common in the team's specific workflow, allowing targeted prompt improvements.

Peer review for high-stakes outputs. Any AI-assisted output going into a regulatory filing, a NERC compliance document, or a customer communication should have a second reviewer run the applicable checklist before it leaves the team. This is not a full document review; it is a 15-minute checklist application on the highest-risk elements. The peer review step catches errors that the author missed because familiarity with the content makes errors harder to see.

Checklist integration into templates. The prompt library templates described in the previous lessons should include the applicable skeptic's checklist as part of the output template. When the AI produces an outage brief, the output format should include a section labeled "Reviewer Checklist" that lists the five items to verify before the brief is used. This integration ensures the checklist is applied at the right moment rather than being a separate step that is easy to skip under time pressure.

The Floor Below Which AI Should Not Operate Without Human Confirmation

Not every grid task is suitable for AI assistance with a checklist-and-continue workflow. Some outputs require human confirmation before the checklist can even be applied, because the consequences of acting on a wrong output are immediate and irreversible. Knowing where this floor is for your specific role is as important as knowing the checklist.

The floor for grid AI work is defined by two criteria: reversibility and immediacy. If an action based on an AI output is difficult or impossible to reverse (a switching action that opens a transmission circuit, a NERC self-certification that is filed with the organization), and if the action is time-sensitive in a way that compresses the verification window, the floor is lower and human confirmation is required before any checklist is applied.

For most planning, compliance, and documentation work, the checklist-and-continue workflow is appropriate because the action is a draft or an analysis, not a real-time operational decision. For real-time operational recommendations (topology optimization suggestions, restoration sequence recommendations), the floor requires that the operator confirm the recommendation against the actual EMS/SCADA display before any switching action, regardless of how confident the AI output appears. The checklist is not a substitute for this confirmation; it is a supplement to it.

Key Takeaways

  • Bad AI output looks good because it is internally coherent, uses the right vocabulary, and produces plausible-magnitude numbers. You cannot rely on prose quality or structural polish as a proxy for accuracy.
  • The six-item forecast skeptic's checklist covers: data vintage, step-load exposure (especially data-center interconnections), unit confirmation, net versus gross load, uncertainty quantification, and source traceability for every number.
  • The five-item restoration estimate checklist covers: field-note basis for every ERT, asset identifier verification against the OMS data, customer count currency, cause code confirmation, and communication-safe language for customer-facing ERTs.
  • The six-item regulatory summary checklist covers: standard version, applicability, effective dates, definition reliance, jurisdiction separation, and recent regulatory action. For 2026, the recent action check must include FERC large-load rulemaking, NERC CLE registry, CIP-003-9 enforcement, and the NERC Level 3 Alert.
  • Building a team skeptic culture requires shared failure documentation, peer review for high-stakes outputs, and checklist integration into the output templates that prompt AI tools produce.
  • The floor below which AI should not operate without human confirmation is defined by reversibility and immediacy. Real-time operational recommendations require operator confirmation against the actual EMS/SCADA display before any switching action, regardless of checklist results.
  • The skeptic's checklist is not a sign of distrust in AI tools. It is a sign of professional judgment about where AI adds value and where human verification is the last line of defense in a safety-critical, regulated environment.