Incident Response for a Shipped Critical Error
It is 9:14 on a Tuesday morning when the email lands, and your stomach drops before you finish reading the subject line. The client is a pharmaceutical manufacturer, the account is worth more than a quarter of your language-service provider's revenue, and the message is three sentences long. A patient-facing dosing leaflet you localized into German and delivered eleven days ago tells patients to take the medication twice daily. The English source says once daily. A pharmacist in Hamburg caught it, escalated it to the client's regulatory affairs team, and now the client wants to know, quote, "how this happened, what else is affected, and what you are doing about it, by end of day." The file passed your delivery gate. The evaluation logged zero Criticals. The quality record has your post-editor's sign-off on it. And a fluent, grammatical, confident German sentence that means the opposite of the source has been in patients' hands for eleven days. This lesson is about the next eight hours, and the eight weeks after them, because what you do now decides whether this becomes a hard incident you contained and learned from, or the incident that ends the relationship. Incident response is the disciplined, pre-planned sequence a localization operation runs the moment a Critical error is discovered after delivery: detect, contain, assess, root-cause, communicate, remediate, and prevent. This is the strategist's version of the discipline the whole program has been building toward, and the honest truth it opens with is this: your gate will fail eventually, and the measure of a mature operation is not that it never ships a Critical, but how it behaves in the hours after it does.
The Incident You Hoped Would Never Come
Every earlier lesson in this program has been about prevention: risk-tiered intake, grounded post-editing, the severity-scored evaluation, the one-Critical-fails delivery gate, the sign-off. If prevention worked perfectly, this lesson would not need to exist. It exists because prevention is never perfect, and a strategist who believes their gate is infallible has built a fragile operation that will handle its first escaped Critical by improvising in a panic. The mature posture is the opposite: assume that despite every control, one silent Critical will eventually escape, and build the machinery to respond to it before you need it. A fire drill run for the first time during an actual fire is not a drill; it is a second disaster layered on the first.
Let us fix the vocabulary precisely, because in an incident the words you use with a client are load-bearing. A Critical error is the highest severity band in the MQM (Multidimensional Quality Metrics) and ISO 5060 error typology: an error that makes the content dangerous, legally exposed, unusable, or actively misleading on a point that matters, such as a flipped dosage, a dropped negation, an inverted safety instruction, or a reversed indemnity clause. MQM is the analytic error-typology framework that classifies translation errors by dimension (accuracy, terminology, locale, fluency) and severity. ISO 5060:2024 is the international standard, published in 2024, that formalizes the MQM-aligned model for human evaluation of translation output. ISO 18587 is the standard for machine-translation post-editing, currently under revision to cover AI and LLM output and to insist the post-editor hold full professional-translator competence. A silent error is a mistranslation delivered in fluent, grammatical prose that does not trip the eye, which is exactly why it survives review: it reads perfectly and means the wrong thing. MT is machine translation; an LLM is a large language model, a general-purpose text predictor that translates as a side effect; MTPE is machine-translation post-editing, the workflow where a human edits machine output. A segment is the unit a translation tool works in, usually a sentence, the row in a CAT tool (computer-assisted translation tool). A TMS is the translation-management system the file moves through, and a TM (translation memory) is the store of previously approved translations that can propagate an error into future projects. A quality record is the segment-level documentation of a delivery: the MT source, the post-edit, the term decisions, the evaluation, the disposition of every flagged error, and the sign-off. That record, built for audit, turns out to be the single most valuable asset you have in an incident, and much of this lesson is about why.
A mature operation is not one that never ships a Critical. It is one whose behavior in the eight hours after it ships a Critical is already written down, rehearsed, and calm.
Why Incident Response Is a Different Discipline From the Gate
The delivery gate and incident response are cousins, but they are not the same skill, and conflating them is a strategic error. The gate is a control that runs before release: it is preventive, it operates on a file you still hold, and its output is a binary go/no-go. Incident response runs after release: it is corrective, it operates on content that is already in the world causing whatever harm it causes, and its output is not a single decision but a coordinated sequence across containment, investigation, communication, and prevention. The gate answers "should this ship?" Incident response answers a harder set of questions all at once: "how do we stop the bleeding, how do we tell the client without making it worse, how did our controls let this through, and how do we make sure it never happens again?" A strategist who has mastered the gate has mastered prevention. Mastering incident response is mastering what happens when prevention fails, and that is the capability that actually protects the business, because prevention failing is not a hypothetical. It is a certainty over a long enough time horizon.
The Seven-Stage Incident-Response Playbook
An incident handled by instinct is an incident handled badly, because instinct under pressure pulls in exactly the wrong directions: it wants to minimize, to reassure prematurely, to fix the one visible symptom and move on, and above all to avoid the uncomfortable client call. A written playbook exists to override those instincts with a disciplined sequence. The seven stages are: detect, contain, assess impact, root-cause, communicate, remediate, and prevent recurrence. Hold the shape in your head as a single line: detect, contain, assess, root-cause, communicate, remediate, prevent. They are ordered deliberately, but not strictly serial. Containment starts the instant you detect, before you know the root cause, because stopping ongoing harm cannot wait for a full investigation. Communication runs in parallel with assessment and root cause, because the client needs to hear from you long before you have every answer. The ordering is a priority ranking as much as a timeline: it tells you what to do first when everything feels urgent at once.
Stage One: Detect and Declare
Detection is the moment the incident becomes known, and it matters strategically because the source of detection tells you something uncomfortable about your operation. In the best case, you detect the escaped Critical yourself, through a downstream QE sweep, a TM cleanup, a spot re-evaluation, or a linguist noticing a pattern. In the worst and most common case, the client detects it, or the client's customer detects it, or a regulator detects it, which is exactly the Hamburg pharmacist in the opening. Self-detection buys you the chance to contain and communicate on your own terms. Client-detection means you are already on the back foot, responding to their alarm rather than leading with your control. Either way, the first disciplined act is to declare an incident: to formally name this as an incident rather than treat it as a routine correction, which triggers the playbook, assigns an incident owner, and starts the clock and the log. The declaration is not bureaucracy. It is the switch that moves the organization from "someone should look at this" to "we are running the response, and one named person owns it." An escaped Critical that gets quietly patched by whoever happened to see it, with no declaration, no owner, and no log, is an incident you will be unable to explain when the client asks how it happened, because nobody actually investigated. Declare, assign an owner, open the incident log, and record the timestamp of detection. Everything after this stage gets logged against that clock.
Stage Two: Contain and Recall
Containment is stopping the harm from spreading while you still do not know its full extent, and it is the stage where speed matters most and where instinct most wants to skip ahead to explanation. Do not explain yet. Stop the bleeding first. Containment in localization has several concrete moves, and the strategist runs the ones that apply:
- Halt any in-flight work that shares the defect. If the same content, the same engine, the same term, or the same post-editor is on other active files right now, freeze those files before they ship the same error. The escaped Critical is a signal about a source of errors, not a one-off.
- Quarantine the poisoned linguistic assets. This is the containment move most operations forget, and it is the one that turns a single incident into a recurring one. If the erroneous segment was written back into the translation memory, it will now propagate: the next project that gets a fuzzy match on that segment inherits the flipped dosage. Pull the bad segment out of the TM immediately, and check whether it fed any termbase entry.
- Recall or flag the delivered content where you can. You rarely control the client's published artifact, but you can and must tell the client exactly which deliverable, which file version, which segments, and which locales are affected, so they can execute their own recall, republication, or field-safety action. Your job is to hand them a precise, bounded target, not a vague apology.
- Preserve the evidence. Before anyone corrects anything, snapshot the delivered file, the quality record, the evaluation log, and the engine and workflow configuration as they were at delivery. If you fix the segment before you preserve the record, you have destroyed the root-cause evidence, and you will be reconstructing what happened from memory.
Containment before explanation. You do not need to know why the dosage flipped to know that you must pull the poisoned segment out of the TM before it flips the next file too.
The recall-versus-patch tension is real and worth naming. There is always pressure to just fix the one segment quietly and move on, especially if the client has not noticed. Resist it in a high-liability incident. A flipped dosage in a patient leaflet is not a typo to correct on the next update; it is a field-safety matter the client's regulatory team may be legally required to act on, and concealing it, even by omission, converts a quality incident into a breach of trust and possibly a breach of contract. Containment is honest and bounded: here is exactly what is affected, here is what we have frozen, here is what we need you to be able to act on.
Stage Three: Assess Impact and Blast Radius
Impact assessment answers the question the client asked in the opening email: "what else is affected?" It is the disciplined determination of the blast radius, the full scope of content that shares the defect, as opposed to the single segment that got noticed. A strategist who reports only the one caught error has answered the wrong question, because the client is not worried about the one error a pharmacist already found; they are worried about the errors nobody has found yet. Blast-radius assessment asks, systematically:
- Same file, other segments. Does the same defect appear elsewhere in the same deliverable? If the engine flipped one dosage, re-check every dosage, every number, every negation in that file against source, because the failure mode is now known and searchable.
- Same defect, other files. Was the same engine, term, or post-editor used on other deliveries to this client or others? If the root is "this engine mishandles this construction," every file it touched with that construction is suspect.
- Propagation through the TM. Did the bad segment leverage into other projects as a fuzzy or exact match before you quarantined it? Trace the TM's usage.
- Time and reach. How long was the content live, how many locales, how many downstream artifacts, and for high-liability content, how many end users could have acted on it? Eleven days of a patient leaflet is a different blast radius than eleven minutes of a draft.
The output of this stage is a bounded, defensible scope statement: precisely which content is affected and precisely which is confirmed clean, with the method you used to confirm it. "We found one error and fixed it" is not a scope statement; it is a hope. "We re-evaluated all 340 dosage-bearing segments across the four files this engine touched for this client in the affected period, found one additional Major and zero further Criticals, and confirmed the German, French, and Spanish leaflets clean by full re-check" is a scope statement. The difference is the difference between a client who trusts your investigation and a client who assumes you missed more.
Stage Four: Root Cause via the Quality Record
Root cause is the disciplined determination of why the Critical escaped, not just what the error was. The error was a flipped dosage. The root cause is the chain of failures that let a flipped dosage pass every control you built. Those are completely different findings, and the client, the auditor, and your own prevention work all need the second one. A root-cause investigation that concludes "the post-editor made a mistake" has stopped one question too early, because it does not explain why the gate did not catch the mistake, and a prevention plan built on "the post-editor should be more careful" will prevent nothing.
This is where the quality record becomes the most valuable artifact you own, and where every hour you spent building segment-level documentation for audit purposes pays back at once. Without a quality record, a root-cause investigation is archaeology: you are guessing at what happened from a finished file, reconstructing decisions from memory, and unable to prove anything. With a quality record, root cause becomes a reading exercise. For the flipped-dosage segment, the record tells you, in minutes:
- What the MT engine produced. Did the engine itself flip the dosage, or did it render it correctly? If the raw MT was correct and the delivered target is wrong, the error entered during post-editing, and you are looking at a human or process failure, not an engine failure. If the raw MT was already wrong, the engine is the source.
- What the post-editor did. Did they touch that segment? Did they accept a suspiciously fluent MT output without checking it against the source number? A high fuzzy match or a clean-looking segment is exactly where a busy post-editor's eye slides past.
- What the evaluation saw. Was the segment sampled by the evaluation at all, or did it fall outside the sample? Was it flagged and then mis-dispositioned, or never flagged? A Critical that was flagged and waved through is a disposition failure; a Critical that was never flagged is a sampling or evaluator-competence failure. These lead to different fixes.
- What the gate did. Did the file pass the gate with the Critical undetected, meaning the gate's coverage has a hole, or did the gate not run properly at all?
Notice how the record turns a vague "how did this happen" into a precise, sequential map of exactly which control failed and why. In the opening incident, suppose the record shows the raw MT rendered "once daily" correctly, the post-editor changed it to "twice daily" while smoothing an awkward sentence and did not re-check the number against source, the segment fell outside the evaluation's sample because the sample was random rather than risk-weighted, and so the gate never saw it. That is a four-link failure chain: an unforced post-editing error, a re-check that did not happen, a sample that did not cover a high-consequence segment, and a gate blind to what the sample missed. Every link is a prevention target. None of them is "the post-editor should try harder." That is what a quality record buys you: a root cause specific enough to fix.
The quality record is the difference between a root cause you can read in an hour and one you have to guess at for a week. The audit artifact you resented building is the investigation you are grateful to have.
Stage Five: Communicate to the Client
Client communication is the stage that most often turns a survivable incident into a lost account, not because the error was unforgivable, but because the communication was. A client can forgive a Critical if they trust your response; they cannot forgive being managed, minimized, or misled while patients are exposed. This stage deserves its own full treatment, and it gets one in the next section. For now, place it in the sequence: you communicate early, before you have all the answers, with an honest acknowledgment and a clear statement of what you are doing, and you keep communicating as assessment and root cause produce findings. Communication is not a single message at the end; it is a channel you open at the start and keep open until the incident closes.
Stage Six: Remediate
Remediation is fixing the actual content, and by this stage it is almost the easy part, because the blast-radius assessment already told you exactly what to fix. Remediation means correcting every affected segment across every affected file and locale, re-verifying each fix against the source (a rushed fix under incident pressure is itself a fresh Critical risk), re-running the delivery gate on the corrected deliverables, and re-issuing them to the client with a clear version marking so nobody confuses the corrected file with the defective one. Two disciplines matter here. First, remediate the whole blast radius, not just the caught segment: fixing only the Hamburg dosage while leaving an unfound error in the Spanish leaflet means a second incident is already scheduled. Second, remediate the poisoned assets, not just the deliverables: purge the bad segment from the TM permanently, correct or add the termbase entry, and confirm the correction propagates to any in-flight work. A remediation that fixes the files but leaves the TM poisoned has fixed the past and left the future broken.
Stage Seven: Prevent Recurrence
Prevention is where an incident stops being a loss and becomes an investment, and it is the stage that separates an operation that has one incident from an operation that has the same incident repeatedly. Prevention takes the specific failure chain the root cause exposed and closes each link so that this class of error cannot recur. In the worked example, the four-link chain yields four concrete prevention actions: a mandatory source-number re-check step in post-editing for dosage-bearing segments so the unforced error is caught at the source; a risk-weighted evaluation sample that guarantees every high-consequence segment (dosages, negations, obligations) is evaluated rather than left to random chance; a gate coverage rule that no high-liability file passes without a targeted numeric-and-negation check independent of the sample; and a feedback entry that this engine mishandles this construction, so future intake routes it accordingly. Each action maps to a link. Prevention that does not map to the actual failure chain is theater, and the classic theater is "we told everyone to be more careful," which changes nothing because carelessness was never the root cause. The strategist closes the loop by writing prevention actions into the pipeline as controls, assigning owners and dates, and verifying weeks later that they actually took, because a prevention plan filed and never implemented is the reason the second incident looks exactly like the first.
Client Communication: The Do's and Don'ts That Decide the Relationship
The technical stages contain the harm, but the communication decides whether you keep the client, so it deserves the most careful treatment in this lesson. A localization strategist is, in an incident, a crisis communicator, and the instincts that feel protective in the moment are almost all wrong. The governing principle is simple to state and hard to hold under pressure: the client will forgive an error far sooner than they will forgive a cover-up, a minimization, or a surprise. Every do and don't below flows from that principle.
The Do's
- Communicate early, before you have every answer. The client needs to hear from you the moment you have confirmed an incident, not after you have finished the investigation. An early message that says "we have confirmed a Critical error in your German leaflet, we have contained it, we are assessing full scope and will update you by 2 p.m." is infinitely stronger than silence followed by a complete report a day later. Silence reads as either ignorance or concealment, and both are worse than the error.
- Acknowledge plainly and take ownership. Name the error accurately: "the German target instructed twice-daily dosing where the source specifies once daily." Do not soften it into "a discrepancy" or "a minor inconsistency." Owning it plainly is what earns you the credibility to be believed about everything else you say.
- Lead with containment and impact, not with explanation. The client's first two questions are "is it still causing harm?" and "what else is affected?" Answer those first: here is what we have frozen, here is the confirmed blast radius, here is what is confirmed clean. The root cause comes after, because to a client with patients exposed, your containment matters more than your excuse.
- Give a bounded scope and a timeline you will hit. Precise, defensible scope ("340 segments re-evaluated, one additional Major found and fixed, all leaflets confirmed clean") plus a timeline you actually meet rebuilds the trust the error spent.
- Report root cause as a system failure and prevention as a system fix. When you do explain, explain it as a failure of controls you are strengthening, not as a person you are blaming. "Our evaluation sample did not guarantee coverage of dosage segments; we have changed it to risk-weighted sampling" tells the client you found the real cause and fixed the machine. It is both more honest and more reassuring than a scapegoat.
- Close the loop in writing. End the incident with a written summary: what happened, root cause, remediation, and the specific prevention actions with dates. This is the artifact the client's own regulatory or quality team needs, and offering it unprompted signals a mature operation.
The Don'ts
- Do not blame the engine. "The AI produced it" is the single most damaging thing you can say, because it tells the client you do not understand your own accountability. The entire program's cardinal rule is that accountability never transfers to the machine; a client who hears you blame the engine hears an operation that will do it again. You post-edited it, evaluated it, gated it, and signed off on it. Own that.
- Do not minimize or hedge. "It was only one segment" and "technically it still reads correctly" and "these things happen with MT" all read as an operation trying to shrink its responsibility. Minimization in a high-liability incident is not reassuring; it is alarming, because it suggests you do not grasp the stakes.
- Do not speculate about cause before you know it. Guessing at root cause in the first message ("we think the engine had an update") and being wrong later destroys credibility precisely when you need it. Say what you know, say what you are still determining, and never present a guess as a finding.
- Do not overpromise the impossible. "This will never happen again" is a promise no operation can keep and every client knows it. Promise the specific, verifiable controls you are adding, which is a stronger claim because it is a true one.
- Do not go silent between updates. A gap in communication during an open incident is read as the worst possible interpretation. If you said 2 p.m., send something at 2 p.m., even if it is "still assessing, next update at 4."
- Do not let the error propagate while you are talking. Communication is not a substitute for containment. If you are on the phone reassuring the client while the poisoned segment is still leveraging into the next file from an un-quarantined TM, your words are undermined by your own pipeline.
The error is a wound the relationship can survive. The cover-up, the minimization, and the surprise are the infection that kills it. Communicate early, own it plainly, and never blame the machine.
A Worked Incident, Detection to Prevention
Return to the Hamburg leaflet and run the whole playbook, so the sequence stops being abstract. It is 9:14 Tuesday; the client's email is on the screen; the flipped dosage has been live for eleven days. Here is the response, stage by stage, with the reasoning showing.
Detect and Contain
Detection was external, the Hamburg pharmacist through the client, which means you are on the back foot and speed and honesty matter double. You declare the incident at 9:20: you name an incident owner (yourself, the localization strategist), open the incident log, and record detection at 9:14. Containment starts immediately, before you understand anything. You freeze every in-flight file that used the same engine and post-editor for this client, three active projects, at 9:35. You check the TM and find the flipped segment was written back on delivery day; you quarantine it at 9:45 so it cannot leverage into the three frozen files or anything else, and you confirm no termbase entry was corrupted. You snapshot the delivered German file, its quality record, the evaluation log, and the engine configuration as of the delivery date, preserving the evidence before a single correction is made. By 10:00 the bleeding is stopped and the evidence is safe, and you have not yet said a word to the client, which is correct: you needed something true to tell them.
First Client Contact
At 10:05 you send the first message, early and honest: "We have confirmed a Critical accuracy error in the German leaflet delivered on the 4th: the target instructs twice-daily dosing where your source specifies once daily. We have contained it, frozen related work, and quarantined the affected translation memory so it cannot spread. We are now assessing the full scope across all files and locales and will send you a confirmed impact statement by 2 p.m. today." Note what this message does and does not do. It acknowledges plainly, names the error accurately, leads with containment, gives a timeline, and does not speculate about cause, does not blame the engine, and does not minimize. It buys you the afternoon to do the assessment properly, on your terms, with the client's trust intact for now.
Assess and Root-Cause via the Record
Impact assessment runs through the morning. You re-evaluate every dosage-bearing, negation-bearing, and numeric segment across the four files this engine and post-editor touched for this client in the affected window, 340 segments in all. You find one additional Major (a unit-formatting slip, not a safety error) and zero further Criticals. You confirm the French and Spanish leaflets clean by full re-check. You trace the TM and confirm the poisoned segment did not leverage anywhere before the 9:45 quarantine. The blast radius is now bounded and defensible.
Root cause runs in parallel, and here the quality record earns its entire existence. You open the record for the flipped segment and read the chain in under an hour. The raw MT rendered "once daily" correctly. The post-editor changed it to "twice daily" while restructuring an awkward German sentence and did not re-check the number against the source. The segment was not in the evaluation's sample because the sample was random, not risk-weighted, so it never reached an evaluator's eye. The gate, relying on the sample, never saw it. Four links: an unforced post-editing error, a missing source-number re-check, a sample that skipped a high-consequence segment, and a gate blind to what the sample missed. Without the record, you would be guessing at every one of these. With it, you can state each with evidence.
The 2 p.m. Update and the Close
At 2 p.m., on time, you send the full impact statement: the confirmed blast radius (one Critical, one additional Major, all leaflets now confirmed clean, no TM propagation), the root cause stated as a system failure ("our evaluation sample did not guarantee coverage of dosage-bearing segments, and our post-editing step lacked a mandatory source-number re-check"), and the remediation status. You do not blame the post-editor by name to the client; you present it as a control gap, which is both the honest framing and the useful one. You remediate: correct the segment, re-verify against source, re-run the gate on all four files, re-issue with clear version marking, and purge the poisoned segment from the TM. Then you close in writing with the prevention plan and dates.
Prevention: Closing Every Link
The prevention plan maps one action to each failure link, and this precision is what convinces the client's quality team you actually found the cause:
- Link one, the unforced edit: a mandatory source-number and negation re-check step in post-editing for all high-liability segments, so a changed number must be confirmed against source before the segment closes.
- Link two, the blind sample: the evaluation sample moves from random to risk-weighted, guaranteeing that every dosage, unit, negation, and obligation in high-liability content is evaluated, not left to chance.
- Link three, the gate hole: a gate rule that no high-liability file passes without a targeted numeric-and-negation verification independent of the sample.
- Link four, the engine signal: a feedback entry that this engine and content type produced a post-edit-induced numeric flip, routed to intake so this content class gets stricter handling going forward.
You assign each action an owner and a date, and you schedule a check four weeks out to confirm each control actually took. The incident that arrived at 9:14 as a potential account-ender closes as a documented, contained, root-caused, remediated, and prevented event, with a client who watched you handle it like a professional. That is the whole difference. The Critical was the same in both worlds. Your response is what decided which world you ended up in.
Two operations shipped the identical flipped dosage. One improvised, minimized, blamed the engine, and lost the account. The other declared, contained, read its quality record, communicated honestly, and closed every link. Same error. Opposite outcome. The playbook is the difference.
Building the Response Before You Need It
Everything above describes what to do during an incident, but the strategist's real work happens before the incident, because a playbook you read for the first time at 9:14 on a Tuesday is a playbook you will execute badly. The capability that separates a mature operation is not knowing the seven stages in the abstract; it is having them pre-built, pre-assigned, and pre-rehearsed so that when the email lands, the response is muscle memory rather than invention. Three things must exist on paper before the first incident, and standing them up is a governance responsibility, not a reaction to a crisis.
First, a named incident owner and an escalation path must be defined in advance. When the Hamburg email arrives, there can be no ambiguity about who runs the response, who has the authority to freeze in-flight work, who contacts the client, and who can renegotiate a deadline. An incident is not the moment to discover that nobody has that authority. Second, the quality record must already be the default output of every high-liability delivery, not something you wish you had once the investigation starts. The record cannot be retrofitted after the fact; it is either captured at delivery time as a matter of routine or it is gone. This is why the audit-readiness discipline covered elsewhere in this program is also incident-readiness: the same segment-level record that satisfies a certifier is the record that makes root cause an hour of reading instead of a week of guessing. Third, the failure path with the client should be pre-agreed: the quality agreement should already state that a discovered Critical triggers a disclosure, a containment action, and a remediation timeline, so the incident conversation is executing a procedure the client signed up for rather than springing a crisis on them.
The mature operation writes its incident response before the incident. When the email lands, it executes a rehearsed procedure with a named owner and a quality record already in hand, not an improvisation.
There is one last strategic reframe worth holding. An incident, handled well, is not purely a loss. It is the most credible demonstration of competence a client will ever see, because anyone can look good on a delivery that goes smoothly; only a genuinely mature operation looks good on a delivery that went wrong. A client who watches you declare, contain, root-cause honestly through your quality record, communicate without minimizing or blaming the machine, remediate the full blast radius, and close every failure link has just watched you prove that your quality claims are real under the one condition that actually tests them. That is a relationship that often emerges from an incident stronger than it entered, which is the opposite of what the panic instinct expects, and it is available only to the operation that built the response before it needed it.
Key Takeaways
- Incident response is the pre-planned sequence a localization operation runs when a Critical error is discovered after delivery: detect, contain, assess impact, root-cause, communicate, remediate, prevent. A mature operation is measured not by never shipping a Critical, which is impossible over time, but by how calmly and honestly it responds when one escapes.
- The seven stages are ordered as a priority ranking, not a strict serial timeline: containment starts the instant you detect, before root cause is known, and communication runs in parallel with assessment, because the client must hear from you long before you have every answer.
- Containment means stop the bleeding before you explain: freeze in-flight work sharing the defect, quarantine the poisoned TM segment so it cannot propagate into future files, hand the client a precise recall target, and preserve the evidence before any correction destroys it.
- Impact assessment determines the blast radius, not the one caught error: same file, same defect across other files, TM propagation, and time-and-reach, producing a bounded, defensible scope statement of exactly what is affected and what is confirmed clean.
- The quality record is the most valuable artifact in an incident: it turns root cause from week-long archaeology into an hour-long reading exercise, showing exactly which control failed, whether the engine or the post-edit introduced the error, whether the evaluation sampled it, and whether the gate saw it.
- A real root cause is the failure chain, not the error and not a person: "the post-editor should be more careful" prevents nothing, while a specific multi-link chain (unforced edit, missing re-check, blind sample, gate hole) gives a prevention target for every link.
- Client communication decides whether you keep the account: communicate early before you have all answers, acknowledge plainly, lead with containment and impact over explanation, give bounded scope and a timeline you hit, and close in writing. Never blame the engine, never minimize, never speculate, never overpromise "this will never happen again," and never go silent between updates.
- Prevention is where the incident becomes an investment: map one concrete control to each failure link, write it into the pipeline with an owner and a date, and verify weeks later that it took, because a prevention plan filed and never implemented is why the second incident looks exactly like the first.
Skill.re