AI for Energy & Utilities
Capable · M14 · lesson 14 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Building Verification Checklists for Energy AI
📖
now learning

Building Verification Checklists for Energy AI

15 min

A compliance lead at a regional utility spent two hours on a Sunday afternoon explaining to her NERC auditor why a section number cited in the utility's most recent CIP self-certification did not exist in the actual standard. The AI assistant had generated it. There was no verification checklist. There was no step between "AI produced this" and "we filed this." Two hours of uncomfortable explanation could have been avoided by five minutes of structured review, applied systematically to every piece of AI-assisted regulatory work.

Why a Checklist, and Why Now

The energy industry has operated on checklists for decades. Pre-switching checklists. Pre-energization checklists. Storm preparation checklists. The checklist is the industry's primary defense against the gap between "I know how to do this" and "I did it correctly this time, under these conditions, with this crew." Checklists work because they externalize memory and sequence, and because they create a documented record that a step was performed.

AI introduces a new category of step that needs to be on the checklist: output verification. The AI produces output that looks authoritative. It uses industry vocabulary, it formats numbers correctly, and it cites documents that sound real. The risk is not that the output is obviously wrong; obvious errors are caught by anyone who reads the output. The risk is that the output is subtly wrong in a way that requires active checking to detect: a citation that is slightly off, an asset ID that does not exist in your records, a forecast value that is within plausible range but from the wrong data vintage.

The verification checklist is the mechanism that makes active checking a systematic habit rather than an optional extra step that happens when someone has time. It should be built once, validated against the failure modes you actually see in your workflows, and applied consistently across every AI-assisted deliverable before it moves to the next step in the process.

The checklist does not make you slow. It makes you the person who never has to explain a fabricated citation to an auditor.

The Three Failure Modes to Design Against

A well-built verification checklist targets specific failure modes, not general quality concerns. Three failure modes are most common and most consequential in grid AI workflows, and each requires a different type of check.

Failure Mode 1: Hallucinated Assets

The model generates asset identifiers, substation names, line segment IDs, generator IDs, or bus numbers that do not exist in your actual asset registry. These hallucinated assets are particularly dangerous because they are not obviously wrong: they follow the naming conventions of your system, they sound like real assets, and they may even match real assets in other parts of your territory or in other utilities' systems.

The check for hallucinated assets is the cross-reference check: compare every asset identifier in the AI output against your authoritative asset registry (GIS for distribution and transmission topology, the model-of-record for power flow studies, the NERC registry for generator and interconnection IDs). Any identifier that does not appear in the registry is a potential hallucination and must be confirmed or removed before the output is used.

The hallucinated-asset check is particularly important for: interconnection study reports (where asset IDs from the study go into formal filings), switching step sequences (where a wrong device tag creates an operational risk), and compliance evidence narratives (where a referenced facility must be a real, registered asset).

Failure Mode 2: Fabricated Citations

The model generates regulatory citations, standard section numbers, FERC order references, or state PUC docket numbers that do not exist or are incorrect. Fabricated citations are the most common and most embarrassing failure mode in regulatory and compliance work. They occur because the model's training data includes thousands of regulatory documents with similar formatting, and the model learns to generate plausible-sounding citations even when it does not have access to the specific document needed.

The check for fabricated citations is the primary-source verification: look up every cited document, section, and version against the actual regulatory source. For NERC standards, go to the NERC website and confirm the standard version and section number. For FERC orders, confirm the order number and the specific provision cited. For state PUC rules, confirm the docket and rule number.

Two common fabrication patterns to know: (1) the slightly-wrong section number, where the standard is real but the section reference is off by one or two numbers; (2) the plausible-but-nonexistent order number, where the FERC order number sounds right for the topic area but does not correspond to a real order. Both patterns pass a casual read because they look authoritative.

Failure Mode 3: Unbounded Forecasts

The model generates forecast values, cost estimates, or planning projections without uncertainty bounds, or with uncertainty bounds that are unrealistically narrow. This failure mode is less about fabrication and more about misrepresentation: the values may be drawn from real data, but presenting a point estimate as if it were a reliable single number, without communicating the uncertainty band, misleads the decision-maker.

In energy planning, unbounded forecasts are particularly dangerous for: peak demand forecasts used in resource adequacy determinations (where missing the uncertainty band can lead to under-procurement or over-building), cost estimates used in rate-case filings (where a point estimate without a confidence interval may mislead the commission about the reliability of the projection), and step-load additions to load growth forecasts (where the timing uncertainty of a data-center interconnection can easily span two to three years).

The check for unbounded forecasts: confirm that every forecast or projection in the AI output includes an uncertainty band, a confidence range, or an explicit statement of the assumptions that would cause the value to be higher or lower. If the AI output presents a single number without qualification, add the qualification before filing or presenting the value.

Building Your Reusable Checklist

A reusable verification checklist has three layers: a universal layer that applies to every AI-assisted output regardless of workflow, a workflow-specific layer that applies to a particular output type (forecast narrative, compliance evidence, switching plan), and a use-case layer that captures the specific gotchas you have actually seen in your environment. Build the universal layer first, then add the workflow-specific layers as you encounter each output type.

Universal Layer: Every AI Output

These checks apply to every piece of AI output before it is used for any purpose. They are fast (two to five minutes total) and catch the most common failures.

Check What to Look For Pass Criterion
Citation verification Any regulatory document, standard section, FERC order, or rule cited in the output Each citation verified against the primary source document; section numbers confirmed
Asset ID cross-reference Any substation name, line ID, generator ID, bus number, or device tag Each identifier found in GIS, the model-of-record, or the NERC registry
Unit consistency Any quantity with a unit (MW, MVA, kV, MWh, $/MWh) All units match the system prompt convention; no silent unit changes between sections
Uncertainty present Any forecast, estimate, or projection An uncertainty band, confidence range, or assumption statement accompanies each projection
Scope boundary Any recommendation, decision, or action item in the output No switching orders, protection-setting decisions, or work-permit approvals; any such content is removed before use

Workflow-Specific Layers

Each workflow adds a short layer of checks specific to the failure modes for that output type.

Forecast narratives:

  • Peak demand value sourced from an identified dataset (not from AI memory)
  • Step-load additions explicitly named and sourced, or explicitly noted as not included
  • MAPE or accuracy metric sourced from a holdout evaluation, not a training metric
  • Weather assumptions stated (normal year, extreme year, scenario)
  • DER adjustments: method described and DER data source identified

Compliance evidence narratives:

  • Standard version and effective date confirmed (e.g., CIP-003-9, effective April 1, 2026, not CIP-003-8)
  • Every requirement description verified against the actual standard text, not a paraphrase
  • All referenced evidence documents confirmed to exist in the evidence package
  • No self-certification language (the narrative describes evidence; it does not substitute for the attesting officer's signature)
  • Registered entity name and ID match the NERC registry record

Interconnection study reports:

  • POI identifier confirmed against the transmission owner's asset records
  • Study process (FERC Order 2023 cluster or tariff provision) correctly identified
  • Nameplate capacity matches the interconnection agreement or application
  • Thermal loading percentages use normal rating as the denominator (not emergency rating)
  • All mitigation requirements cited with the applicable standard (e.g., TPL-001-5)

Switching sequences:

  • Every device tag in the sequence verified against GIS or the EMS model
  • Starting state and ending state match the work-order description
  • Hold points present for telephone confirmations and protection checks
  • Verification reminder present in the document (not to be executed without sign-off)
  • Rollback steps identified for each critical switching action

Making the Checklist Stick

A checklist only works if it is actually used. The three most common reasons checklists fail to stick in energy organizations are: they are too long to use under time pressure, they live in a document no one opens, and they produce no evidence of completion. Address each problem directly.

Keep It Short Enough to Use Every Time

The universal layer should be completable in five minutes or fewer. If it takes longer, it will be skipped under deadline. The workflow-specific layers should add no more than five additional checks. A checklist with 30 items will be completed once and then abandoned. A checklist with 10 items will be completed every time. If your workflow genuinely requires more than 10 checks, split it into two checklists: a "before filing" checklist and a "before use in analysis" checklist, with different personnel responsible for each.

Embed It in the Workflow, Not in a Separate Document

The checklist should appear at the point of use: at the bottom of the document template, as a tab in the planning workbook, as a field in the project tracking system. If completing the checklist requires opening a separate document, the friction will cause it to be skipped. The completion evidence should be in the same record as the output: a checked field, a date, and the reviewer's initials, stored alongside the AI-assisted document.

Make Completion Evidence Explicit

For any AI-assisted document that will be filed with a regulator or used in a formal proceeding, the checklist completion is the audit trail that shows human oversight was applied. Each check should have a date and a reviewer identifier. For NERC compliance evidence, the checklist completion record is part of the evidence package. For rate-case filings, the checklist record documents that the AI-assisted testimony was verified by a qualified professional before submission.

Worked Example: Applying the Checklist

Consider a compliance analyst who has used AI to draft an evidence narrative for NERC CIP-003-9 (effective April 1, 2026). The draft looks polished: it describes the utility's physical access controls for low-impact BES Cyber Systems, cites the relevant standard requirements, and references the evidence documents. The analyst applies the checklist.

Universal layer:

Citation verification: The draft cites "CIP-003-9, Requirement R1, Attachment 1, Section 1.2." The analyst opens CIP-003-9 on the NERC website. Attachment 1 in the actual standard has five sections. Section 1.2 is "Physical Security Controls." The citation is correct. Check passes.

Asset ID cross-reference: The draft references "Ridgeline Control Center" as the registered entity location. The analyst verifies this against the NERC registry. The registered entity name is "Ridgeline Substation Control Building." The names do not match exactly. This is a potential issue: the audit record should use the NERC-registered name precisely. The analyst corrects the name before filing. Check reveals a discrepancy; corrected.

Unit consistency: No quantitative values in this narrative. Check not applicable.

Uncertainty present: No forecasts or projections in a compliance narrative. Check not applicable.

Scope boundary: The draft does not include any switching orders or protection decisions. Check passes.

Compliance evidence layer:

Standard version and effective date: The draft correctly identifies CIP-003-9 as effective April 1, 2026. Check passes.

Requirement description verified: The draft describes the physical access control requirement. The analyst reads the actual CIP-003-9 Attachment 1, Section 1.2, sentence by sentence. The AI description includes a phrase that does not appear in the standard: "regular inspection of physical access controls." The actual standard requires a process for controlling physical access but does not specifically require "regular inspection" at this section. The AI has paraphrased with an addition. The analyst removes the extra phrase. Check reveals a discrepancy; corrected.

Evidence documents confirmed: The draft references three evidence documents. The analyst checks the evidence folder. Two of the three documents are present. The third, a "Physical Access Log Q1 2026," does not yet exist in the folder. The analyst notes that this evidence item must be obtained before filing. Check reveals a missing document; flagged for follow-up.

Three checks were completed; two revealed real issues: an incorrect registered entity name and a paraphrased requirement with an unsupported addition. Both were corrected before filing. The third revealed a missing evidence document that would have caused a compliance gap if the narrative had been filed without follow-up. This is what a working verification checklist looks like: not bureaucracy, but a systematic catch of real problems.

Adapting the Checklist as Your Workflows Evolve

A verification checklist is a living document. As you use AI-assisted workflows, you will encounter failure modes specific to your tools, your data, and your regulatory environment that are not covered by the universal layer. Add them. Keep a log of every AI error your team has caught, describe the check that caught it or the check that should have caught it, and add that check to the relevant workflow layer.

Two events in 2026 should trigger checklist updates for utilities using AI for compliance and interconnection work. First, CIP-003-9 became enforceable April 1, 2026: any compliance evidence checklist that still references CIP-003-8 is using the wrong version. Update the standard version reference in the checklist's compliance evidence layer. Second, the NERC Computational Load Entity registry is committed for delivery by December 31, 2026: any checklist for large-load interconnection work should add a check confirming whether the prospective customer may be subject to CLE registration requirements and whether that determination has been made.

Similarly, when a new AI tool is deployed or an existing one is updated, run a brief red-team exercise: give the new tool a task you know well, check the output against primary sources, and see what the tool gets wrong. The errors it makes become your first use-case-layer additions to the checklist. A team that does this consistently builds a checklist that is calibrated to the actual failure modes of its actual tools, rather than a generic checklist that catches some errors and misses the ones specific to your environment.

Every error your checklist catches is a problem that did not make it into a filing, a switching order, or a compliance binder. That is the return on the investment.

Key Takeaways

  • Three failure modes drive the checklist design: hallucinated assets (cross-reference check), fabricated citations (primary-source verification), and unbounded forecasts (uncertainty check). Target the checklist at these specifically.
  • The universal layer (five checks, five minutes) applies to every AI output: citation verification, asset ID cross-reference, unit consistency, uncertainty present, and scope boundary.
  • Workflow-specific layers add four to five checks for each output type: forecast narratives, compliance evidence, interconnection study reports, and switching sequences each have distinct failure modes.
  • Keep the checklist short enough to complete every time: 10 checks maximum per workflow layer. A checklist that is skipped under deadline is worthless; a short checklist that is always completed catches more errors than a comprehensive one that is used occasionally.
  • Embed the checklist at the point of use and capture completion evidence (date, reviewer) in the same record as the AI-assisted output. For regulatory filings, the checklist completion record is the audit trail for human oversight.
  • Update the checklist when standards change (CIP-003-9 in April 2026, CLE registry in December 2026), when new tools are deployed, and whenever your team catches a new AI error that was not covered by an existing check.
  • The compliance lead who catches a fabricated citation at the checklist stage has a two-minute correction. The one who catches it during an auditor review has a two-hour conversation. The checklist is not overhead: it is the difference between those two outcomes.