AI for Energy & Utilities
Strategic · M7 · lesson 7 of 22 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Capturing Institutional Knowledge Before It Retires
📖
now learning

Capturing Institutional Knowledge Before It Retires

15 min

The superintendent retires in six weeks. In his head: thirty years of feeder quirks, the exact load mix that stresses the 115-kV line after a cold snap, and the verbal protocol every crew chief follows when the SCADA telemetry goes dark on that circuit. None of it is in the EMS. None of it is in the ADMS. It lives in one person who is about to leave, and nobody has built a system to catch it.

The Great Crew Change Is Already Here

Energy professionals have heard the phrase "Great Crew Change" for over a decade, often treated as a distant problem on someone else's watch. In 2026, it is not distant. Industry surveys consistently show that more than 25 percent of utility workers are retirement-eligible within the next few years, concentrated in the roles that carry the most irreplaceable operational knowledge: system operators, distribution superintendents, protection engineers, and senior planners who have lived through every major grid event of the past three decades.

EPRI's workforce analysis projects more than 30 percent growth in digital and analytical utility roles through 2030. That is not a coincidence. The same period that empties the institutional-knowledge vault is the period when AI-assisted grid work requires deeper, better-structured knowledge to function well. A forecasting model is only as smart as the operating context you feed it. A topology optimizer can only recommend sound switching sequences if the rules of the local grid are encoded somewhere it can reach. When the person who holds those rules walks out the door, the AI does not automatically learn them.

This lesson is about closing that gap before it opens. Not in an abstract HR sense, but operationally: building the interviews, the structured capture formats, the knowledge repositories, and the embedding workflows that turn a retiring superintendent's expertise into durable input for both AI systems and the next generation of operators.

What Institutional Knowledge Actually Is (and Why It Is Hard to Capture)

Institutional knowledge sounds vague until you try to write it down. Then you realize it has at least three distinct layers, each with a different capture method and a different failure mode.

Layer One: Explicit Procedural Knowledge

This is the layer people think of first: the switching order for a specific outage scenario, the load factor calculation for an industrial tariff class, the checklist for restoring a feeder after a major storm event. Some of this exists in procedures manuals, operations logs, and training materials. Much of it is outdated, incomplete, or filed somewhere nobody has opened in years. The first capture task is locating and auditing what is already written down, then verifying it against current practice. The superintendent you are interviewing will often say "that procedure is in the binder, but we stopped doing it that way five years ago." That gap is the first artifact you want to capture.

Layer Two: Contextual Judgment

This is harder. It is the knowledge of when a rule applies and when it does not. An operator who has worked the same control territory for twenty years does not just know the switching sequence; they know that on hot July afternoons after a prior outage on the parallel circuit, the load transfer limit on feeder F-14 is about 80 percent of nameplate, not 100 percent, because the aging transformer at the junction substation runs hot. That limit is not in the EMS. It is in the operator's head, learned from a failure that happened twelve years ago that was never formally written up because the situation was resolved before it became a reportable event.

Contextual judgment lives at the intersection of system topology, equipment history, and operational experience. Capturing it requires structured interviews that surface specific scenarios, not general principles. "Walk me through the last three times you deviated from standard procedure, and tell me exactly why" is a far better capture prompt than "what do you know that the manual doesn't say."

Layer Three: Relational Knowledge

Who calls whom when there is an unusual event on the transmission side that might affect the distribution feeder. Which neighboring utility's system operator can be trusted to give accurate information on intertie status at 2 AM. Which contractor crew is genuinely capable of working energized lines safely and which crew's safety record is more complicated than their reputation suggests. This layer is invisible in any formal system and often does not survive a single retirement. It is also the layer that most directly affects response time in a major event.

Structured Knowledge Capture Methods for Utility Environments

Capturing institutional knowledge is a discipline with a body of practice. The challenge for utilities is that most formal knowledge-management frameworks were built for software companies or consulting firms, not for organizations where the knowledge is inseparable from physical infrastructure and regulatory accountability. Adapting these methods to a grid context requires some specific choices.

Scenario-Driven Interviews

The most productive interviews are built around specific past events, not abstract competency questions. Collect the operator's log entries, the outage management system (OMS) event records, and the post-event reports for the five or ten most significant events the person managed in the last decade. Use those as the spine of the interview. For each event, ask: what did you know at the start that the official records did not capture? What was the first indicator that told you this was going to be unusual? What did you do that was not in any procedure? What would you tell a new operator who faces this same situation?

Document the responses in a structured template that tags each insight to a specific asset, circuit, or operating condition. That tagging is what makes the knowledge usable for AI later. A paragraph of undifferentiated narrative is difficult to retrieve at the moment of need. A structured record that says "asset: feeder F-14, condition: hot summer afternoon after parallel outage, rule: derate transfer limit to 80% of nameplate, reason: T-junction transformer thermal history" is something an AI-assisted operations support tool can surface as a relevant advisory when that specific combination of conditions recurs.

Shadowing and Think-Aloud Protocols

For knowledge that is procedural but not written, the most effective capture method is direct observation combined with verbal narration. Have a documentation specialist shadow the experienced operator or engineer during a representative shift, asking them to narrate their decision process in real time. This produces a qualitatively different kind of record than retrospective interviews, because it captures the monitoring attention, the quick environmental checks, and the small adjustments that the operator makes automatically but cannot easily articulate in a structured interview.

The output of a think-aloud session is typically a rough transcript that requires significant editing to turn into structured knowledge. Budget for that work. A single well-conducted think-aloud session followed by two hours of structured documentation is worth more than ten general-knowledge interviews that produce narrative summaries nobody ever consults.

Knowledge Graphs and Structured Tagging

The goal is not a document library. It is a queryable knowledge structure. Modern AI systems work best with knowledge that is organized as connected entities: asset, condition, rule, exception, source, and date. A simple spreadsheet with those columns is already substantially better than unstructured text. A knowledge graph that explicitly links a transformer's thermal history to the operating rules that derive from it is better still.

Vendors in the utility knowledge management space offer tools ranging from simple templates to enterprise knowledge graphs. The specific tool matters less than the commitment to structured tagging from the first interview. Retrofitting structure onto a library of untagged narrative documents is painful and expensive. Building structure in from the start adds perhaps 30 percent to the initial documentation effort and pays back on the first day a new operator needs to find something quickly.

Turning Captured Knowledge into Usable AI Input

This is where the workforce conversation intersects with the AI deployment conversation. An AI-assisted grid operations tool can be an extraordinary amplifier of institutional knowledge, but only if that knowledge exists in a form the tool can access and apply.

Grounding the AI with Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) is the technical pattern behind most production-grade AI tools that need to reference authoritative local information rather than general training data. In a RAG-based operations support tool, when an operator asks "what are the constraints on restoring feeder F-14 after a parallel outage?", the system does not guess from its training. It retrieves the relevant entries from the structured knowledge repository and uses them to generate a contextually accurate advisory.

This only works if the knowledge repository exists, is structured, and is kept current. The superintendent's tacit knowledge, once captured in the structured format described above, becomes the grounding layer for exactly these queries. Without it, the AI tool operates on generic grid knowledge and misses the local operational rules that define reliable operation of that specific system.

Think of it this way: the model is a very well-read junior colleague who has studied every grid textbook ever written. The knowledge repository is the veteran beside them saying "that's the textbook answer, but here's what actually happens on this feeder." Both are necessary. The AI without the local knowledge gives generic advice. The local knowledge without the AI is locked in a document nobody reads. Together, they produce contextually accurate, retrievable, current operational guidance.

Validation, Currency, and Human Sign-Off

Captured knowledge is not static. Grid topology changes. Equipment ages and gets replaced. Operating rules evolve with load growth, new generation, and regulatory changes. A knowledge repository that was accurate when built in 2024 may contain significant errors by 2028 if nobody is maintaining it. Build into the capture process a maintenance workflow: who is responsible for reviewing each category of knowledge asset, on what schedule, and what triggers an out-of-cycle review?

Critically, any knowledge that enters an AI-assisted operational workflow must carry a verification timestamp and a named human reviewer. This is not bureaucratic overhead; it is the evidence trail that a reliability organization or a NERC auditor can follow if a decision is ever questioned. "The AI recommended it" is never an acceptable defense. "The AI retrieved a knowledge asset that was reviewed and approved by Operations Superintendent Jane Smith on March 15, 2026" is a substantially different and defensible answer.

Building the Capture Program: A Practical Playbook

Knowledge capture at the scale the Great Crew Change demands is a program, not a project. It requires an ongoing organizational capability, not a one-time interview sprint.

Identifying At-Risk Knowledge Holders

Start with a simple analysis: which retirement-eligible employees hold knowledge that is not documented anywhere and that, if lost, would directly affect operational reliability or regulatory compliance? Rank them by proximity to retirement and by the criticality and uniqueness of what they know. This produces a prioritized list that tells you who to interview first.

In a typical medium-to-large utility, this analysis will identify somewhere between fifteen and fifty individuals who should be considered high-priority for structured knowledge capture. The number is manageable. The failure mode is not having this list and doing nothing until the person has already submitted their retirement notice.

The Transition Window

The most effective knowledge transfer happens while the retiring employee is still working, over a period of three to twelve months before their departure date. This is the window for structured interviews, shadowing sessions, and mentored knowledge transfer to specific successors. A one-week "knowledge transfer" meeting in the final days before retirement is nearly useless for anything beyond the most explicit procedural knowledge.

Utilities that have done this well typically build a formal "transition phase" into the retirement agreement itself, often as part of a phased-retirement or emeritus-consultant arrangement. The employee is compensated specifically for structured knowledge transfer activities, which both motivates participation and creates a formal organizational record that the transfer occurred.

Successor Pairing and Active Testing

Knowledge transfer is not complete when the document is written. It is complete when a successor can perform the relevant function independently and accurately. Build into the capture program an active testing phase: the successor is asked to handle a simulated version of the specific scenarios that the knowledge repository addresses, with the experienced employee available to correct and annotate. The documentation is revised based on what the successor gets wrong. This is the quality-control loop that separates a knowledge library people actually use from one that sits on a shared drive.

Worked Example: A Feeder Supervisor Transition

Consider a concrete scenario. A distribution supervisor with 28 years of experience on a specific territory in the Pacific Northwest is retiring in four months. His territory includes two industrial feeders serving large manufacturing loads, a section of underground cable installed in the late 1980s with a repair history that is partially in paper records, and a normally-open tie point to a neighboring distribution circuit that is used for emergency restoration but has a non-obvious load management constraint that the operations team learned about through two near-misses in the 1990s.

The utility has a new ADMS that is being configured for AI-assisted fault location and switching recommendations on this territory. The ADMS vendor has configured the topological model correctly based on GIS data. But the system does not know about the cable repair history that makes one specific segment significantly more likely to fail under high-load conditions. It does not know about the tie-point load constraint. And it does not know the load profiles of the industrial customers well enough to anticipate the unusual demand pattern that occurs when both manufacturing plants run extended shifts during a particular production season.

Without the supervisor's knowledge, the ADMS will give technically correct but operationally naive recommendations for this territory. With it, the ADMS's recommendations will be grounded in three decades of operational learning. The capture process for this supervisor includes: a review of all OMS event records on his territory for the past ten years, a three-session structured interview covering the ten most significant events and the five most common non-standard operating conditions, two shadowing sessions during the highest-load period of the spring industrial season, and documentation of the cable segment risk profile using maintenance records augmented by his annotated recollections.

The output is not a report. It is a set of structured knowledge assets tagged to specific assets and conditions, ingested into the ADMS knowledge layer, reviewed by the incoming supervisor, and scheduled for annual review by the operations engineering team. On the day the retiring supervisor's last shift ends, his knowledge does not retire with him. It continues to inform every switching recommendation the ADMS makes on that territory.

Key Takeaways

  • More than 25 percent of utility workers are retirement-eligible within the coming years, concentrated in the roles with the deepest operational knowledge; this is not a future problem but a current one requiring active programs now.
  • Institutional knowledge has three layers: explicit procedural rules, contextual judgment about when rules apply, and relational knowledge about who to call; each layer requires a different capture method, and the deepest layer is the most commonly lost.
  • Scenario-driven interviews built around specific past events consistently produce higher-quality, more actionable knowledge artifacts than competency-based or general-knowledge interviews.
  • The goal of capture is not a document library but a structured, tagged, queryable knowledge repository that AI-assisted operations tools can use as a retrieval-augmented grounding layer for locally accurate operational guidance.
  • Every knowledge asset that enters an AI-assisted workflow must carry a verification timestamp and a named human reviewer; this is both a quality control requirement and the audit trail that supports regulatory accountability.
  • Effective knowledge transfer is confirmed only when a successor can independently and accurately handle the scenarios the knowledge addresses; the write-down is an intermediate step, not the end state.
  • A formal transition-phase arrangement, beginning three to twelve months before retirement, is substantially more effective than last-minute knowledge transfer sessions and should be built into the utility's workforce planning process as a standard practice.