AI for Energy & Utilities
Capable · M21 · lesson 21 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Verifying a Forecast Before It Drives a Dollar
📖
now learning

Verifying a Forecast Before It Drives a Dollar

15 min

A peak-day procurement analyst at a Southwest utility reviewed the AI load forecast at 6:30 a.m. on a day the operations center had flagged as a potential record high. The forecast showed a peak of 11,200 MW. She ran the five-step verification workflow she had been drilling for two months. She found three things: the model had not been retrained in seven months; a 450 MW data-center complex had completed interconnection six weeks earlier; and the weather inputs were from a forecast file generated at 4 a.m. that had since been updated with a heat advisory raising the expected high by 4 degrees. She adjusted, documented, and forwarded a revised estimate of 11,720 MW. The actual peak came in at 11,740 MW. Without the workflow, the desk would have committed to capacity coverage for 11,200 MW and been 540 MW short in the hottest part of the afternoon.

Why Verification Is Not Optional: The Stakes Define the Workflow

Load forecast verification is the discipline of checking an AI output against the conditions it was designed to handle, the conditions it was not, and the real-world changes that have occurred since it was calibrated, before that output drives a financial or operational commitment. For a daily forecast used in low-stakes scheduling, a quick sanity check may be sufficient. For a peak-day procurement decision that commits millions of dollars of capacity, a structured verification workflow is not a best practice. It is the professional standard.

The argument for verification rests on a simple asymmetry: the cost of running a 30-minute verification workflow before a peak-day procurement is near zero relative to the value of the decision being made. The cost of a 540 MW reserve shortfall on the hottest afternoon of the year is not near zero. In a regional ISO market, real-time prices during a capacity emergency can run at 10 to 50 times the day-ahead price. Emergency bilateral capacity purchases at the last minute can cost tens of millions of dollars more than a properly-committed day-ahead position. Every dollar of that premium is a direct consequence of the unverified forecast.

This lesson presents the complete verification workflow: the steps, the checks, the documentation requirements, and the decision logic that determines when an AI forecast can be trusted as-is and when it needs an adjustment. The workflow is not complicated. It is disciplined. And discipline under time pressure is exactly what the professional standard requires.

The Five-Step Verification Workflow

The complete pre-procurement forecast verification workflow consists of five steps. They are designed to be executed in sequence, to take less than 30 minutes on a routine day and perhaps 45 minutes on a day requiring a significant adjustment, and to produce a documented record that supports the procurement decision.

Step 1: Confirm Model Currency

The first step is to confirm that the AI model's training data is reasonably current relative to today's load conditions. Ask: what is the model's training cutoff date? When was the model last retrained or fine-tuned? Has the load shape changed materially since that date?

The materiality threshold is practical: if the system has not changed significantly, a model trained within the last year is usually acceptable for day-ahead forecasting. If a large new load (over 50 MW) has come online since the training cutoff, you have a potential step-load miss that requires investigation. Check the interconnection queue completion records or your interconnection team's notification system for any energizations in the past 90 to 180 days that exceed your materiality threshold. If you find one, record it and proceed to the holdout test in step 3, which will confirm whether the step-load effect is visible in the error pattern.

Step 2: Validate Weather Inputs

The second step is to confirm that the weather inputs the model was run on are current and appropriate for the forecast day. Ask: what weather source was used? When was the weather file last updated? Is the forecast high temperature for today consistent with the current National Weather Service or commercial weather service update?

On peak-risk days, weather files can change significantly between 4 a.m. (when many automated forecasting systems generate their day-ahead output) and 7 a.m. (when the analyst reviews it before the procurement window). A heat dome advisory, a dew point revision, or a significant temperature upgrade can shift expected peak load by 100 to 200 MW. If you find that the weather inputs are more than 3 hours old on a day with active weather updates, re-run the forecast with the current weather file or apply a manual adjustment based on the temperature sensitivity calculation.

Temperature sensitivity (MW per degree Fahrenheit) is a number you should know for your system and for the season. In the summer, typical large-system temperature sensitivities run from 30 to 80 MW per degree Fahrenheit. If the weather file has been updated by 3 degrees since the forecast was run, the corresponding load adjustment is 3 times your temperature sensitivity. Document the adjustment and its basis.

Step 3: Run the Holdout Check

The third step is to run a rapid holdout check against recent comparable days. For a summer peak, this means checking the model's performance on the last five to ten days when temperatures exceeded the threshold comparable to today's forecast. You are looking for one thing: is the model consistently off in the same direction?

A consistent directional miss (forecast too low by 200 MW every hot day for the past two weeks) is the signature of a step-load problem or model calibration drift. Random scatter around zero means the model is well-calibrated and the variation is weather noise. The holdout check answers the question "is the model still working?" in about 10 minutes if the data is accessible. If the mean directional error over the holdout period is greater than 1% of expected peak load, an adjustment is warranted.

Step 4: Apply the DER Sanity Check

The fourth step applies specifically to territories with significant behind-the-meter solar, storage, or EV load. Check the model's expected midday load profile against the five tells described in the prior lesson: is the midday dip appropriately sized for today's solar forecast? Is the evening ramp shape consistent with recent actuals in a solar-and-EV territory? Have any new DER interconnections occurred since the model's training cutoff that might need an incremental adjustment?

This step can be abbreviated on days without significant solar generation (overcast days, winter days) and expanded on bright summer days in high-penetration territories. The goal is not to rebuild the DER analysis from scratch every morning. It is to catch the large obvious DER errors before they propagate into the procurement commitment.

Step 5: Document and Decide

The fifth step is the most important and the most frequently skipped: write it down. For every peak-day verification, the record should include: the model version and training cutoff date; the weather source and last update time; the holdout check results (days tested, mean error, direction); any DER issues found; any manual adjustments applied with their basis and magnitude; the final verified forecast with uncertainty band; and the name of the analyst who completed the verification and the time it was completed.

This record serves multiple purposes. It protects the analyst: if a question arises later about why a procurement decision was made, the verification record shows that the professional standard was followed. It protects the utility: in a rate case or NERC audit, the verification record demonstrates prudent process. And it creates institutional learning: a log of verification records over time reveals patterns (recurring step-load problems, recurring weather input staleness) that can be addressed systemically rather than discovered each morning.

What "Confidently Wrong" Looks Like

The phrase "confidently wrong" describes the specific failure mode that verification is designed to catch. An AI forecast is confidently wrong when it produces a specific, precise-looking number with high apparent accuracy that is systematically incorrect because the model's training does not reflect the current conditions.

The hallmarks of a confidently wrong forecast are: the model has not been retrained since a significant change (a large load addition, a major DER deployment, a structural change in the load curve); the model's recent accuracy on comparable days has deteriorated compared to its historical track record; and the forecast has not been stress-tested with weather sensitivity runs or holdout checks. The forecast does not look wrong. It produces a number and presents it with the same computational confidence it would use to forecast any other day. The confidence is in the mechanism, not in the accuracy of the result.

The consequences of acting on a confidently wrong forecast are disproportionate to the magnitude of the error because the error is directional. A random forecast error can be over or under; on average it nets out. A systematic directional error from a step-load miss is always in the same direction. You are always short. And you are always short on the days that matter most (peak days, high-load days) because the step load is constant but the market stress is concentrated on high-demand days.

The model's confidence is a measure of its internal consistency, not a measure of its accuracy against reality. Verification measures accuracy against reality. These are different things, and confusing them is how a confidently wrong forecast drives a real dollar into the wrong place.

Building the Verification Habit: From Procedure to Reflex

A verification workflow written in a manual is not the same as a verification workflow that runs reliably at 6:30 a.m. before a peak-day procurement commitment. The difference is habit, and habit is built through repetition under actual conditions, not through reading about the workflow. This section is about how teams can institutionalize verification so it becomes a reflex rather than a procedure.

The key design principle is to make the default behavior the correct behavior. If the verification workflow requires the analyst to proactively go find data that is not in front of them, it will be skipped under time pressure. If the workflow is embedded in a checklist that is part of the daily opening procedure, and if the data needed for steps one through four is accessible in a single dashboard or log, then execution becomes the path of least resistance.

Practically, this means: the model's training cutoff date and last retraining date should be visible on the forecast output, not buried in a separate system. The weather source and last update time should be visible alongside the forecast. The holdout comparison for the past ten comparable days should be available in the same tool that shows the current forecast. And the DER registry summary (total installed capacity, last update) should be accessible in under one minute. None of these require new systems. They require that existing data be organized for verification access, not just for forecast generation.

In the Great Crew Change context, this institutionalization becomes urgent. When the experienced analyst who has been running mental holdout checks for fifteen years retires, the institutional knowledge of how to catch a step-load problem goes with them. The verification workflow must be documented, tool-supported, and drilled with new analysts before the experienced cohort leaves. A team that builds the habit into the work environment is less dependent on any individual's intuition and more resilient to workforce transition.

The Decision Logic: When to Proceed, When to Adjust, When to Escalate

After running the five steps, the analyst must make one of three decisions: proceed with the AI forecast as-is, apply a documented adjustment and proceed, or escalate for additional review before a procurement commitment is made.

Proceed as-is when: the model is reasonably current (retrained within six months and no major load additions since then), the weather inputs are current (updated within three hours), the holdout check shows a mean error below 1% of expected peak with no consistent directional bias, and the DER sanity check finds no obvious calibration issues. Under these conditions, the AI forecast is a verified output that a professional can stand behind.

Apply and document an adjustment when: the model has a known step-load miss of quantifiable size, or the weather inputs are stale by more than three hours and a temperature update has occurred, or the holdout check reveals a consistent directional miss. Each of these cases is correctable with a documented manual adjustment. Apply the adjustment, record the basis, update the uncertainty band to reflect any increased uncertainty in the adjustment, and proceed. A well-documented manual adjustment is not a sign of AI model failure. It is a sign of professional judgment working correctly.

Escalate when: the holdout check reveals a large systematic error whose source you cannot identify (more than 3% of expected peak), or the model appears to have not been retrained for over a year and multiple large loads are unaccounted for, or there is a conflict between the AI forecast and the outputs of an independent statistical model that is too large to resolve within the time available. Escalation means getting a senior forecaster or supervisor involved before the commitment is made, not after. The commitment window can usually accommodate a brief escalation if the rest of the workflow has been run efficiently.

Verification in the ISO Market Context: Timing Is Everything

For utilities operating in deregulated ISO markets, the verification workflow must fit within the market's day-ahead and real-time timelines. The day-ahead energy market typically closes at noon or 1 p.m. for the following operating day. The capacity commitment window may be even earlier, requiring a verified forecast by 9 or 10 a.m. to give the procurement desk time to execute. The verification workflow must therefore be designed to complete well before the market closes, not just before actual peak demand arrives.

In this context, the five-step workflow becomes a morning routine, not an emergency response. The analyst who waits until 10 a.m. to run a verification that should have started at 6:30 a.m. may find that the weather file has already been updated twice, that a new large load came online overnight, and that the procurement window is now 45 minutes away. The institutional answer is simple: the verification workflow starts when the day-ahead forecast is generated, not when the analyst gets around to it.

ISO market rules also create specific documentation requirements. In many RTOs, the load-serving entity's day-ahead load forecast submission is a formal market action that creates obligations. If the submitted forecast is materially different from actual load, the entity may face imbalance settlement charges. A verified forecast that includes a documented adjustment for a step-load miss can reduce the magnitude of imbalance settlement relative to an unverified forecast that was simply wrong. The verification documentation also supports any challenge to imbalance charges in the settlement dispute process.

For real-time operations, the verification concern shifts to the latest available inputs. If the day-ahead commitment was made on a verified forecast, and actual temperature has since exceeded the P90 scenario assumption, the operations team needs to know whether additional capacity is available intra-day. This requires a real-time forecast update using the most current weather inputs, and a rapid holdout check comparing the updated forecast to actual intervals from the current day. The same five-step logic applies, compressed to perhaps 10 minutes and supported by a real-time dashboard rather than a spreadsheet.

The principle is consistent across all market contexts: verification is not a document-everything bureaucratic exercise. It is a disciplined, time-bounded professional process that extracts the highest-quality decision from the available information before the commitment window closes. The investment is 30 minutes. The protection is proportional to the stakes of the decision, which on a peak-day procurement can easily exceed $10 million.

Key Takeaways

  • A 30-minute structured verification workflow before a peak-day procurement commitment costs nearly nothing relative to the value of the decision and can prevent a shortfall that costs millions in emergency capacity premiums.
  • The five-step verification workflow covers: confirming model currency, validating weather inputs and their age, running a holdout check on recent comparable days, applying the DER sanity check in high-penetration territories, and documenting everything before the commitment is made.
  • A "confidently wrong" forecast is one that produces a specific, precise-looking number that is systematically incorrect because the model's training does not reflect current conditions; the model's computational confidence is not evidence of accuracy against reality.
  • A consistent directional miss in the holdout check (always too low, not random noise) is the strongest signal that a step-load problem or calibration drift exists and that a manual adjustment is warranted before procurement commits.
  • The decision logic after verification is: proceed as-is (verified, no issues), apply and document an adjustment (known correctable issue), or escalate (large unidentified error or conflicting models too close to the commitment window to resolve alone).
  • Documentation is not bureaucracy. It is the record that protects the analyst, defends the procurement decision in a rate case or NERC audit, and creates the institutional learning that improves future forecasts.
  • Building the verification workflow into daily habit, supported by a dashboard that makes the needed data accessible in under five minutes, is the organizational intervention that prevents a single experienced analyst's retirement from ending the verification practice.