AI for Energy & Utilities
Capable · M15 · lesson 15 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Getting Accurate Energy Output from AI
📖
now learning

Getting Accurate Energy Output from AI

15 min

An interconnection engineer at a regional transmission organization pasted a paragraph from a developer's study request into an AI tool and asked for the applicable NERC standard. The model returned "FAC-001-2, Requirement R3" with the confidence of a senior compliance attorney. She forwarded it to the developer. The problem: the current effective version is FAC-001-3, and the specific requirement the developer needed to understand was R4, not R3. The version difference was not minor; the two requirements have different applicability dates and different evidence obligations. The developer spent two weeks building a response to the wrong requirement. That story is about how to stop the model from inventing or misremembering citations, and how to force it to anchor every claim in a specific, verifiable document you provide.

The Citation Problem in Energy AI

Generative AI is trained to produce fluent, coherent, plausible text. In the energy domain, fluent and plausible often means: the right vocabulary (NERC, FERC, MW, IRP, N-1), the right structure (a numbered list of requirements, a compliance calendar), and the right tone of authority. What it does not mean is: grounded in the specific document version, the specific tariff schedule, or the specific asset record that your question is actually about.

The citation problem has three distinct forms in utility AI work. Understanding which form you are dealing with changes how you address it in the prompt.

The first form is the stale citation: the model cites a real standard but the wrong version. FAC-001-2 versus FAC-001-3. CIP-003-8 versus CIP-003-9 (effective April 1, 2026). The standard exists, the model has seen it, but it is not the current effective version. This is the most common citation problem and the hardest to spot, because the wrong version looks right to anyone who is not already familiar with the specific change between versions.

The second form is the invented citation: the model produces a plausible-sounding standard number, tariff schedule, or requirement reference that does not exist. "NERC PRC-025-3" when the actual standard is PRC-025-2. "Schedule 26-A" when your tariff has no such schedule. "Requirement R7" when the standard has only six requirements. Invented citations are caught more easily than stale ones, but only if the reviewer knows the standard well enough to notice the gap.

The third form is the misattributed citation: the model cites a real, current document but gets the requirement wrong. It attributes an obligation to R3 when the actual obligation is in R5. It cites the wrong section of the tariff. It assigns an applicability to a Transmission Owner when the requirement applies to a Generator Owner. This form is the most dangerous because it is hardest to catch without reading both the AI output and the cited document in parallel.

All three forms share a single root cause: the model does not know what it does not know. It produces citations by pattern-matching from training data, not by looking up a document. The only structural fix is to provide the document and require the model to reason from it rather than from memory.

There is a fourth failure mode worth naming separately: the correct citation, wrong conclusion. The model cites the right standard, the right version, the right requirement number, and then draws a conclusion that the cited text does not support. "NERC FAC-001-3, R4 requires the Transmission Planner to..." followed by a paraphrase that subtly changes the obligation. This failure is the most insidious because a reviewer who checks that R4 exists and that the standard version is correct may not re-read R4 carefully enough to catch the wrong paraphrase. The citation spot-check described later in this lesson addresses this specifically.

Forcing the Model to Cite the Standard

The core technique for accurate regulatory output is simple to describe and requires discipline to maintain: paste the relevant section of the document you want the model to reason from, and instruct the model to cite section numbers and never go beyond what the pasted text supports.

This is not just a best practice. It is the only reliable method for getting legally accurate output from a generative AI model in a regulated industry context. The model's training data may include NERC standards, FERC orders, and state commission rules, but it learned from whatever versions, excerpts, and summaries were in its training corpus. None of that is traceable or auditable. A model reasoning from text you provided is different: you know what it read, you know the version date, and you can show the audit trail to a NERC auditor or commission staff.

The Paste-and-Cite Technique

The paste-and-cite technique works in four steps:

  1. Retrieve the current effective document. Go to the primary source: NERC's posted standards page for reliability standards, FERC's eFiling system (eLibrary) for orders and tariffs, your state commission's docket management system for commission orders, your utility's filed OATT for tariff provisions. Do not use a secondary summary, a vendor white paper, or a prior-year compliance calendar. The current effective version from the primary source is the only version that matters for a compliance or regulatory determination. Print the effective date at the top of what you paste.
  2. Identify and paste the relevant section. You do not need to paste an entire 50-page standard. Identify the specific requirements, definitions, and applicability tables that bear on your question. Paste that section into the prompt with a clear label: "The following is NERC FAC-001-3, Requirements R4 and R5, and the Applicability section, retrieved from the NERC standards website on [date]. Reason only from this text for all requirement-related claims." For context length reasons, paste the minimum necessary section and note that the model should ask you for additional sections rather than reasoning from its training data.
  3. State the cite-and-refuse instruction. Add: "For every claim in your response that relates to a regulatory requirement, cite the specific requirement number and subsection from the text I have pasted. If your answer requires information not present in the pasted text, say so explicitly and describe what is missing rather than supplementing from your training data. Do not paraphrase requirements in ways that change their meaning; quote the relevant language directly when it matters."
  4. Verify the output against the source. When you receive the response, open the original document alongside it and check that every cited requirement number says what the model claims it says. This step takes five to ten minutes for most compliance analysis tasks and catches misattributed and paraphrased citations before they drive a compliance decision. It is the last human check in the workflow and cannot be skipped.

A practical note on the paste-and-cite workflow for NERC standards: the NERC standards website organizes standards by category, and each standard includes a table of applicability that lists which registered entity types are subject to which requirements. Always paste both the requirement text and the applicability table. The most common misattribution error is the model assigning a requirement to the wrong entity type, because it read the requirement but not the applicability table.

Forcing the Model to Cite the Tariff

Tariff citations present a different challenge than standard citations. NERC standards are numbered and versioned with public effective dates accessible on a single website. Tariffs are utility-specific, filed with FERC or a state commission, and revised through docketed proceedings with alphanumeric docket numbers that can be hard to track. The current effective tariff for your utility is not in the model's training data in a useful form: it may have been trained on an older version, a filed version that was later amended, or a version from a similar utility in a different jurisdiction.

The paste-and-cite technique applies directly to tariff work, with a complication: tariffs are complex, interconnected documents. A rate question may require the applicable rate schedule, the relevant definition, and the service agreement form, all of which interact. A model given only the rate schedule may produce output that is correct as far as it goes but misses an obligation in the definition or the service agreement.

For tariff work, the prompt should include: the rate schedule section, the definitions for any term that appears in the rate schedule and could change the answer, and an explicit scope instruction: "I am pasting Schedule LCS and the Definitions section. If your analysis requires a section I have not provided (such as the OATT or the service agreement appendix), tell me which section you need rather than reasoning from a section I have not pasted."

Worked Example: Large-Load Service Rate Question

Situation: A rate analyst needs to determine whether a new data-center customer contracting for 200 MW qualifies for the utility's Large Customer Service schedule or must take service under a general commercial schedule.

Weak prompt:

Does a 200 MW data center qualify for large load service rates at our utility?

What you get: a generic answer about large-load service rate structures that may apply to a notional average utility, not to your specific tariff. The model may cite a threshold of 1 MW, 5 MW, or 50 MW depending on which utilities' tariffs were most common in its training data. It may also omit the definition of "Large Load Customer" in the tariff's Definitions section that determines whether demand is measured at the customer's meter or at the point of delivery on the transmission system. None of these details may match your tariff.

Strong prompt:

I am a rate analyst at an investor-owned utility. The following is our current effective Large Customer Service Schedule (Schedule LCS) and the definition of "Large Load Customer" from our tariff's Definitions section, both as filed with the [state] Public Utilities Commission and effective as of [date]. Using only this text, determine whether a customer with a contracted demand of 200 MW qualifies for service under Schedule LCS. Cite the threshold requirement and the definition from the pasted text. If the tariff text does not address a 200 MW contract specifically or if the answer depends on a provision not in the pasted sections, identify what I need to review and with whom. [Schedule LCS and Definitions text follows]

What you get: an analysis grounded in your actual tariff, citing the exact threshold and definition, with a clear flag if a part of the question requires additional tariff sections or legal review. This output is the starting point for a defensible rate determination.

Forcing the Model to Cite the Asset Record

Asset references are a category of citation failure that is particularly dangerous for operational work. When a model is asked about a specific substation, circuit, or piece of equipment, and it does not have the asset record in front of it, it may produce plausible-sounding asset details drawn from the naming conventions of utilities in its training data. The names sound right. The identifiers follow the format you expect. They may not match anything in your actual GIS or asset register.

This failure mode appears most often in three types of work:

  • Work order drafts: the model invents circuit IDs or equipment designation numbers that are not in the provided telemetry or work management system export. A switching order with a fabricated circuit ID is an operational safety risk: it can direct a field crew to the wrong equipment or fail to identify the correct clearance boundaries.
  • Outage briefs: the model invents substation names that sound plausible for the service territory but do not match the OMS data. An outage brief with a wrong substation name creates confusion in customer communications, crew dispatch, and regulatory reporting.
  • Interconnection study narratives: the model invents point-of-interconnection identifiers, bus names, or transformer designations that do not appear in the study data. A study narrative with wrong asset references fails engineering review and may need to be completely rewritten.

The fix is the same structural approach as for standards and tariffs: provide the asset record and instruct the model not to go beyond it. The specific instruction for each use case:

  • Work order: "Use only the equipment IDs, circuit designations, and clearance tag numbers from the work management system export I have pasted. If the switching sequence requires an asset not listed in my data, write [ASSET NOT IN PROVIDED DATA] and describe which asset you need the record for."
  • Outage brief: "All circuit and substation references must match exactly the identifiers in the OMS export I have provided. Do not abbreviate, substitute shorthand, or use informal names that are not present in the data."
  • Study narrative: "Use only the bus names, transformer IDs, and point-of-interconnection designations from the study data I have provided. If a section of the narrative requires an asset identifier not in the data, write [ASSET ID NEEDED] and describe which asset you need."

An invented asset ID in a switching order is not just an accuracy problem. It is a safety event waiting to happen. The model that creates it never has to answer to a field crew standing in front of the wrong panel.

The Verification Workflow After the Prompt

The paste-and-cite technique and the [UNVERIFIED] flag instruction reduce the probability of wrong output. They do not eliminate it. Even a model reasoning from pasted text can misattribute a citation (claim R4 supports a conclusion that R4 does not actually support) or produce a paraphrase that changes the meaning of the cited requirement. The verification workflow after the prompt is the last professional check before AI output influences a real decision.

The verification workflow has three components, each addressing a different failure mode:

  1. Citation spot-check. Pick three to five cited requirements, tariff sections, or asset references from the AI output. Open the original document and confirm that the cited section says what the model claims it says, in the language the model claims. This is not a full document review; it is a targeted check on the specific citations that will drive the most consequential conclusions. Five minutes of spot-checking catches misattributed and paraphrased citations before they enter a filing or an operational procedure. Document which sections you spot-checked and what you confirmed.
  2. Flag resolution. Review every [UNVERIFIED], [INFERENCE], and [ASSET NOT PROVIDED] flag in the output. For each flag, decide: (a) can this be resolved by providing additional source material in a follow-up prompt, (b) does it require independent research from a primary source that you then add to the working document, or (c) does it need to go to legal, engineering, or compliance review? Document how each flag was resolved, who resolved it, and what source was used. This documentation becomes the audit trail if the output is later incorporated into a filing or an operational procedure.
  3. Completeness check. Ask: is there a part of the applicable document I did not paste that could change this analysis? For a NERC standard, this might be the definitions section, the applicability table, or the implementation guidance. For a tariff, it might be the service agreement form, the OATT, or a recent revision. For an IRP document, it might be the methodology appendix. The completeness check is a professional judgment call that requires knowing the document structure, which is why it cannot be automated and must remain a human-in-the-loop step at the end of every AI-assisted regulatory or compliance session.

When to Ask the Model to Refuse

There are categories of question in utility work where the right answer is "I cannot determine this from the information provided, and you should not ask me to guess." Teaching the model to refuse is as important as teaching it to cite. Forcing the model to answer a question it cannot answer well is worse than leaving the question open, because a fabricated answer carries an authority that an honest "I don't know" does not.

The categories where the prompt should explicitly invite refusal:

  • Legal interpretation questions. "Does this contract language create a firm service obligation?" is a question for a licensed attorney, not a language model. The model can organize the relevant contract language and identify the provisions that create ambiguity; it should not reach a legal conclusion. The prompt should say: "Do not state a legal conclusion. Identify the provisions that bear on this question, note where the language is ambiguous, and flag any precedent question that requires legal review."
  • Engineering judgment questions. "Is this transformer overloaded under N-1 contingency?" requires an actual power-flow calculation with the relevant network model, not a language model's pattern-matched response to a description of the transformer. The model can summarize the applicable NERC reliability standard and describe the calculation methodology; it should not produce a numeric result that mimics a real contingency analysis. The prompt should say: "Describe the applicable reliability standard and the steps of the required analysis, but do not produce a numeric result. I will run the power flow analysis separately."
  • Unanswered regulatory questions. If a standard does not address a specific scenario and no FERC or NERC staff guidance has been issued, the correct output is to flag the gap rather than infer an answer from analogous provisions. The prompt should say: "If the pasted text does not resolve this question, say so and identify the specific gap, citing the provision that is closest to the question and explaining why it does not answer it fully. Do not infer an answer from analogous provisions unless you label the inference clearly and explain why the analogy applies."

The refuse-and-flag output is the most professionally valuable behavior in a high-stakes regulated context. A model that says "Section R5(ii) does not address this scenario; the ambiguity is between the applicability language for Transmission Owners and the separate applicability for Transmission Planners, and this should be taken to compliance counsel" has done its job. It has organized the relevant text, identified the specific gap, and referred the question to the right human. That is exactly the role AI should play in a regulatory workflow.

Key Takeaways

  • The three citation failure modes for utility AI are stale citations (wrong version), invented citations (nonexistent standard or section), and misattributed citations (real document, wrong requirement). A fourth mode, correct citation with wrong conclusion, requires the citation spot-check to catch.
  • The paste-and-cite technique is the only reliable method for accurate regulatory AI output. Retrieve the current effective document from the primary source, paste the relevant sections, and instruct the model to cite section numbers and nothing beyond what you pasted.
  • Tariff citations require a scope instruction: tell the model which sections you have provided and instruct it to ask for missing sections rather than reasoning from sections you did not paste.
  • Asset references must be anchored to the records you provide. The instruction not to invent equipment IDs or circuit designations is a safety requirement in operational work, not a style preference.
  • The three-step verification workflow (citation spot-check, flag resolution, completeness check) is the professional's last line of defense before AI output enters a decision. It cannot be skipped in compliance, rate case, or operational contexts.
  • Teaching the model to refuse is as important as teaching it to cite. Legal interpretation, engineering judgment, and unanswered regulatory gaps should produce a flag for expert review, not an AI-generated answer that fills the silence with confident inference.
  • The accountability principle is unchanged: providing source documents and requiring citations creates the audit trail that enables a human to verify the output and take accountability for the decision. It does not move accountability to the model.