Training the Humans Who Will Work With the Agent
An agent goes live on Monday. By Friday, half the human reviewers are batch-approving without reading. By the following Friday, two of them have publicly told their managers they "don't trust this thing." By the end of week three, the escalation owner has not actually owned an escalation because the runbook lives in a Notion page nobody has reopened since the kickoff. The agent is working perfectly. The humans are not โ because nobody trained them, and the orientation slides emailed before launch were treated the way orientation slides are always treated. This lesson is the 90-minute live onboarding for the humans who will work with the agent: reviewers, escalation owners, stakeholders. Designed to pre-empt the "I don't trust it" backlash by replacing abstract reassurance with concrete competence. Built in 2026 from the trial-and-error of teams who launched agents in 2024 and 2025 and watched their human side fail before their agent did.
Why 90 Minutes, Why Live, Why Mandatory
The default approach to onboarding humans around an agent is wrong in three predictable ways. It is asynchronous (a Loom video, a Notion page, a wiki link). It is one-time (you attend during onboarding, then never again). It is voluntary (the comms email says "please review when you have a chance").
Each of these defaults is wrong, and the corrections are non-negotiable.
Live, not async
An asynchronous training is fine for transferring knowledge about a stable, well-documented system. An agent is neither stable nor well-documented in week one. The questions that matter most are the ones that emerge from the room โ the questions the trainer did not anticipate, the questions other attendees did not know to ask, the half-formed concerns that need to be drawn out by a facilitator paying attention to facial expressions.
The 90 minutes are live. Camera-on for remote sessions. Voice-on. Recorded for those who could not attend, but the recording exists as a fallback, not as the primary delivery vehicle. Anyone whose only exposure to the training is the recording starts at a disadvantage and has to make it up some other way.
90 minutes, not 60, not 180
Sixty minutes is the standard meeting block in corporate culture. Sixty minutes is wrong for this. The first 20 minutes are warm-up โ context, scope, "what this is and is not." Real learning happens in the middle 60 minutes, when attendees are working with the agent themselves, asking questions, watching live demos go wrong. The last 10 minutes are cool-down โ what happens next, contact information, escalation paths.
Squeeze the 90 minutes into 60 and you cut the middle, which is where the value is. Stretch to 180 and attendees disengage in the back half. Ninety is the empirical sweet spot. Block 120 on the calendar to allow for natural overrun, but design for 90.
Mandatory, not optional
Optional training has a known attendance pattern: the people who attend are the people who would have figured it out anyway. The people who would have benefited most do not attend. This is the population that, three weeks later, becomes the source of "I don't trust it" backlash.
Mandatory training has a different attendance pattern: the population matches the population of people whose work will intersect the agent. The training surfaces concerns from the skeptics, the cautious, the uncertain โ exactly the people whose buy-in determines whether the rollout succeeds.
Mandatory means: the manager is on the hook for attendance, completion is tracked, and inability to operate without the training is a real workflow consequence (not a policy threat but a practical one โ reviewers without the training cannot get access to the review queue).
The training is not about the agent. It is about the humans. The reviewer who leaves the room with concrete competence in three review patterns, knowing exactly who to escalate to and when, is a different reviewer than one who left with general enthusiasm. General enthusiasm decays in two weeks. Concrete competence compounds.
Three Audiences, Three Tracks
The 90-minute training is not the same content for everyone. There are three distinct audiences with three distinct jobs, and the training delivers a customized track for each.
Track one: reviewers
The largest audience. The humans who will spend hours per week reading the agent's output and approving, editing, or rejecting it. They need to leave the training competent in three patterns: how to approve confidently, how to edit precisely, how to reject and provide feedback that improves the agent.
The reviewer track focuses on muscle memory. They need to see the review interface, do five live reviews with coaching, learn the keyboard shortcuts, internalize the rubric, and practice making the call on borderline cases.
Track two: escalation owners
A smaller audience. The humans who own the agent's edge cases, who get paged when the agent does something it should not, who handle customer-facing fallout, who make the calls about pausing the agent in degraded operation.
The escalation track focuses on judgment. They need to see the alerting console, walk through the runbook for the top five incident patterns, practice the pause-the-agent procedure on a non-production environment, and rehearse the customer-facing communications.
Track three: stakeholders
The smallest audience by attendance but the largest by influence. Product managers whose features depend on the agent. Sales engineers who reference the agent in customer conversations. Executives whose reports include agent metrics. Customer success managers who handle escalations from the customer side.
The stakeholder track focuses on mental model. They need to know what the agent can and cannot do, how to read the dashboard, what the metrics mean, when to bring concerns to whom, and how to talk about the agent externally.
Why separate tracks
The instinct to combine all three audiences into one training is wrong for the same reason combining all three audiences into one announcement email is wrong. Each track has different questions, different anxieties, different success criteria. A combined training pleases no one โ reviewers find the executive-level mental model too high-level, executives find the keyboard shortcuts irrelevant, escalation owners get lost between the two.
Three 90-minute sessions, ideally on the same week to keep teams aligned but separately attended, work much better than one 120-minute combined session.
The Reviewer Track Curriculum (90 minutes)
The detailed agenda for the largest and highest-leverage track.
Minutes 0-15: context and scope
What the agent does. What it does not do. Where it sits in the workflow. The pilot numbers (approval rate, escalation rate, time savings โ concrete numbers, not "performance was strong"). The reviewer's role in continued performance โ every review is a training signal.
End with the team's commitments: response time SLA on rejections (the agent improvement loop depends on reviewer feedback being incorporated), escalation response time, weekly metrics review the reviewer can attend.
Minutes 15-25: the rubric
How to decide approve, edit, or reject. The decision framework presented as a clear flowchart, not a list of guidelines. The five common borderline cases and how to think about each. The rule for when to call the escalation owner instead of deciding alone.
Include the explicit rubric for adversarial input: when reviewing an output that responded to suspicious input, what red flags to look for, what to do when you see them. Prompt injection awareness goes here.
Minutes 25-55: live practice
This is the load-bearing 30 minutes. Five practice cases, each three to five minutes including discussion. The cases are drawn from real pilot data, sanitized, representing different scenarios:
- The clear approve: the agent did exactly the right thing; reviewer practices the muscle of confident approval without over-scrutinizing.
- The clear edit: the agent did mostly the right thing with one fixable issue; reviewer practices identifying the issue and editing efficiently.
- The clear reject: the agent did the wrong thing; reviewer practices the rejection-with-feedback workflow that drives agent improvement.
- The borderline: the agent did something defensible but maybe not optimal; reviewer practices judgment, attendees compare answers, facilitator reveals what the team's calibration is.
- The escalation: the agent did something the reviewer cannot evaluate alone โ sensitive customer, unusual case, edge of the agent's capability; reviewer practices the escalation handoff.
For each case, attendees make the call individually before discussion. The discussion reveals the spread of answers, which is itself the most valuable artifact โ attendees see that other reviewers find these cases hard too, that there is genuine ambiguity, that the team's calibration is shared, not assumed.
Minutes 55-70: the review interface tour
Live walkthrough of the actual interface the reviewer will use. Keyboard shortcuts. The escalation button. The pause button. The history view. The feedback channel for systemic issues. The personal preferences (notification settings, queue filters).
This section deliberately covers practical mechanics that an async video could also cover โ but the value of doing it live is that questions surface. The reviewer asks "what about the case where..." and the facilitator answers, and the answer is visible to the whole room.
Minutes 70-80: when to trust, when to doubt
The section that most matters for the "I don't trust it" question.
The honest message: do not trust the agent. Trust the system. The agent is one component; you are another. The system works because of the combination. Your judgment matters. Your edits matter. Your rejections matter. The agent gets better because of you, not in spite of you.
Specific cases where reviewer skepticism has caught real errors in the pilot โ concrete examples, not vague reassurance. Acknowledgment of cases where the agent will look confident and be wrong, and how the reviewer learns to spot those cases. Acknowledgment that the reviewer will sometimes be wrong, and the system catches reviewer errors too (escalation paths, peer review on sampled cases, the eval set re-running periodically).
Minutes 80-90: contact, cadence, close
Who to contact for what. Slack channel (real channel name, not a future placeholder). Office hours (real time on the calendar). Weekly metrics review (real meeting on the calendar). Quarterly listening session (real date).
Close with a commitment: the agent owner commits to acting on reviewer feedback within a defined SLA, sharing aggregated metrics weekly, and bringing every systemic issue to a transparent decision. Reviewers leave the room with the contact information memorized and the cadence in their calendar.
The Escalation Owner Track (90 minutes)
A different curriculum for a smaller audience with a different job.
Minutes 0-15: incident lifecycle
The lifecycle of an agent incident from detection through postmortem. Where the escalation owner sits in the lifecycle. The on-call rotation. The expected response time at each stage.
Minutes 15-30: the top five incident patterns
Drawn from the pilot and from industry. For each: how it presents in alerting, the diagnostic steps, the containment options, the communications template, the postmortem implications.
The five patterns vary by domain but typical 2026 examples include: model regression after silent provider update, tool call failing with unexpected error, prompt injection succeeding past guardrails, latency spike from upstream dependency, runaway loop hitting cost cap.
Minutes 30-50: the pause-the-agent procedure
Live walkthrough on a non-production environment. Every escalation owner gets hands-on practice pausing the agent. Includes the rollback procedure if pausing reveals a problem with the pause mechanism itself.
Includes the communications procedure: who gets paged, who gets a notification, who is expected to acknowledge. The Slack templates for incident channel creation. The customer-facing communications template (and the rule that customer-facing comms always go through the on-call communications lead, never the escalation owner unilaterally).
Minutes 50-65: customer-facing fallout
What happens when the agent did something a customer is angry about. The escalation script. The empathy framework. The bounds of what the escalation owner can offer (a refund? a credit? a personal call from a senior leader?). The handoff to customer success or legal when the situation exceeds the escalation owner's authority.
This section is heavily roleplayed. Two attendees take turns being the angry customer and the escalation owner. The room provides feedback. The facilitator models the response when attendees freeze or default to corporate speak.
Minutes 65-80: the postmortem
What goes in an agent postmortem. The two new fields beyond SRE (version drift and eval gap). The escalation owner's responsibility in producing the postmortem. The cross-team learning expectation โ every postmortem is distributed to the agent owner, the AI Platform team, and other agent teams.
Minutes 80-90: contact, cadence, close
Same as reviewer track but with escalation-specific contacts: PagerDuty rotation, incident channel template, legal on-call, customer success on-call.
The Stakeholder Track (90 minutes)
The third curriculum, for the audience whose job is mental-model accuracy, not operational execution.
Minutes 0-20: the agent in context
What the agent does and where it fits in the broader product or operation. The pilot numbers. The roadmap of what the agent will do next quarter, what it will not do next quarter, what is being considered. The boundary line between "agent's job" and "human's job" with concrete examples.
Minutes 20-40: how to read the dashboard
Live walkthrough of the dashboard. What each metric means. What "good" looks like, what "concerning" looks like, what "incident" looks like. The difference between leading and lagging indicators. The lag time before a regression shows in the metric.
Critical: the stakeholders who will be asked "how is the agent doing?" need to be equipped to answer accurately, not just optimistically.
Minutes 40-55: what can the agent do, what can it not
Concrete examples of within-scope and out-of-scope work. The scope boundary. The reasons for the boundary (cost, safety, latency, accuracy thresholds). When the boundary will expand and how that decision is made.
For stakeholders whose features may depend on agent capabilities, this section is the basis of their planning. Better to learn the scope clearly in the training than to ship a feature that assumes capability the agent does not have.
Minutes 55-70: how to talk about the agent
External-facing language. What to say to customers. What to say to prospects. What to say to investors. The bright lines (do not over-claim, do not under-claim, do not anthropomorphize the agent in customer-facing language). The approval process for public statements about the agent.
Minutes 70-85: feedback loops
How the stakeholder gives input back to the agent owner. The product backlog process for feature requests against the agent. The escalation path for concerns. The cadence of the cross-functional agent review.
Minutes 85-90: contact, cadence, close
Standard close with stakeholder-relevant channels.
Pre-empting the "I Don't Trust It" Backlash
The most common post-launch failure is the reviewer or stakeholder who, three weeks in, becomes vocally skeptical of the agent. The skepticism is often correct โ the agent does have failure modes, and the skeptic is often the one paying attention. But uncoordinated skepticism becomes corrosive. The training is one of the strongest tools to channel skepticism productively.
Where the backlash comes from
Backlash has three roots:
- Unaddressed concerns. The reviewer had a concern, did not voice it (or voiced it and was not heard), and the concern festered. The agent eventually does the thing the reviewer was worried about, and the reviewer says "I told you so" โ except they did not, because nobody heard.
- Surprise failures. The reviewer expected the agent to be more capable than it is, encountered a failure mode they did not know to expect, and concluded the agent is unreliable. The concern is real; the framing is wrong (the agent has bounded reliability; the reviewer needed to know the bounds).
- Identity threat. The reviewer's role identity is changing โ from primary author to reviewer โ and the change is being processed as loss. The agent becomes the symbol of the loss. Skepticism toward the agent becomes a way to defend identity.
The training addresses each root.
Channels for unaddressed concerns
The training establishes explicit, low-friction channels for concerns. Slack channel. Office hours. Anonymous feedback. The training names specific people who answer (and the SLA they commit to). The training models how to raise a concern productively: specific case, expected behavior, observed behavior, suggested response.
The architect's job is to make raising concerns feel like contribution, not complaint. Reviewers who raise concerns get acknowledged in the weekly metrics review, get attributed when the concern leads to an agent change, get credit in the broader rollup.
Calibration of expectations
The "what can the agent do, what can it not" segment is calibration work. Attendees leave with a realistic mental model. The agent is not magic. The agent has known failure modes. The reviewer is part of the safety net for the failure modes. The known failure modes are listed explicitly in a printed handout.
This is harder than it sounds because the architect has to suppress their own enthusiasm. The architect built the thing; they want to talk about everything it does well. The training is not about that. The training is about giving the reviewer an accurate mental model, which often means dwelling on the failure modes more than the success modes.
Honoring the role transition
The identity-threat root requires the most care. The training acknowledges, explicitly: the reviewer's role is changing. The reviewer is becoming someone who applies judgment to drafts rather than someone who produces drafts from scratch. This is a different skill. It is not a lesser skill. In many cases it is a more senior skill โ the reviewer is now operating at the level of editor rather than writer, manager rather than individual contributor.
The training names the transition. The training does not pretend the transition is trivial. The training offers a career path for the reviewer-role identity (Senior Reviewer, Reviewer Lead, Agent Quality Manager โ whatever the org calls it).
For organizations where the role transition genuinely is a downgrade โ where the reviewer used to do work that paid more and was more interesting and now does less interesting work for the same pay โ the training cannot paper over the reality. Either the role is changed materially (new responsibilities added, compensation reviewed) or the rollout will struggle. Pretending the transition is a promotion when it is not is worse than acknowledging the reality.
The Trainer and the Facilitator
Two roles in the training that look similar but are different.
The trainer
The technical expert. Knows the agent inside and out. Has built it, evaluated it, debugged it. Answers the hard questions about how it works.
The trainer is usually the agent architect or a senior engineer on the agent team. They deliver the substantive content. Their failure mode: too deep in technical detail, too dismissive of "obvious" questions, too defensive when the agent's failures are surfaced.
The facilitator
The communication expert. Manages the room. Reads facial expressions. Draws out the quiet attendee. Surfaces the half-formed concern. Keeps the trainer from steamrolling.
The facilitator is usually HR, an internal communications lead, or a dedicated training professional. They are not expected to know the agent in technical detail โ they are expected to know the room.
The pairing matters. A trainer running solo will optimize for substance and lose the room. A facilitator running solo will optimize for connection and miss the substantive questions. Both in the room covers both.
For organizations without a dedicated facilitator role, the agent architect should partner with their HR business partner for the training. The HRBP attends, manages the soft side, and gives the architect feedback after the session.
Post-Training Cadence
The 90-minute training is not the whole training program. It is the kickoff. The post-training cadence determines whether the initial competence sticks.
Week one: office hours
An hour of office hours the week after launch. The agent architect or technical lead is available; reviewers, escalation owners, and stakeholders can drop in with questions. Attendance is voluntary. The session is recorded with consent.
This session catches the questions that did not surface during training because the attendee had not yet experienced the agent in their own work.
Week four: first reviewer retrospective
Thirty minutes with the reviewer cohort. What is working. What is not. What needs to change. The architect listens; the architect commits to specific action items with dates.
The reviewer retrospective is the load-bearing meeting of the first month. It is where the rollout either earns trust (specific actions accepted, dates committed, follow-up evident) or starts losing it (vague acknowledgment, no actions, drift).
Month three: refresher training
Sixty minutes, reviewer cohort. Re-grounding in the rubric, updates on agent capabilities, discussion of cases that proved harder than expected in production. Attendance is required.
Quarter-by-quarter: listening sessions
Forty-five minutes per quarter with each cohort separately. What changed this quarter, what is changing next quarter, what we are hearing, what we are doing about it.
The new-hire onboarding
The 90-minute training is also the new-hire onboarding for anyone joining the affected team after launch. The training content is maintained as a living document. The recording of the most recent live session is the asynchronous fallback for new hires; the architect commits to running a fresh live session at least quarterly so the recording is never more than a quarter stale.
Signals That the Training Worked
How does the architect know the training landed? Specific signals to track:
- Approval-rate calibration. In the first two weeks, the reviewer cohort's approval rates should cluster within a defined range. Reviewers approving 99% of cases were either insufficiently trained on rejection or are not engaging. Reviewers rejecting 60% are either insufficiently trained on approval or have unrealistic expectations.
- Rejection feedback quality. The rejection-with-feedback workflow should produce structured feedback that the agent team can act on. Vague rejections ("this is wrong") indicate the rejection segment of the training did not land.
- Escalation precision. The escalation owner should be escalated the right things โ not everything, not nothing. If everything gets escalated, the reviewer track did not communicate the rubric. If nothing gets escalated when there are real issues, the escalation path is not trusted.
- Office hours attendance. First-month office hours attendance is a positive signal โ questions are surfacing. Drop-off after the first month is normal. If office hours go empty in week two, either the training was so thorough no questions remain (unlikely) or attendees do not trust the channel.
- Anonymous feedback content. Specific, action-oriented anonymous feedback is good โ the channel works. Vague complaints or silence is concerning.
- Slack channel engagement. Steady traffic in the dedicated channel โ questions, observations, kudos โ indicates the channel is alive. Silence indicates the channel is dead, and discussions are happening elsewhere (almost certainly DMs).
Anti-Patterns to Avoid
The slide deck that becomes the training
The architect builds a slide deck. The slide deck becomes the training. The training is the slide deck. Attendees see ninety slides over ninety minutes; nobody learns anything because nobody practices anything.
Fix: build the training as a series of activities, not a sequence of slides. Slides exist to anchor the activity, not to replace it.
The "trust the agent" message
The trainer, wanting to drive adoption, leans hard on "trust the agent." Reviewers hear: "your judgment is not needed." Approval rates skyrocket because reviewers stop reviewing. The first incident exposes the over-trust. The whole rollout reverses.
Fix: explicit "do not trust the agent, trust the system" message. Reviewer judgment is centered. The agent is presented as a component, not the answer.
The training that ignores the role transition
The training treats the reviewer's job as the same job it always was, just with an agent draft as input. Reviewers process the change as loss anyway, but without a frame for talking about it. The unprocessed loss becomes resentment.
Fix: explicit acknowledgment of the role transition. Naming the new skill. Offering the career path.
The single combined session
The architect, to save time, runs one 120-minute session for all three audiences. Each audience leaves with about a third of what they needed. The training is checked off; the competence is not built.
Fix: three 90-minute sessions, one per audience. Yes, this costs the architect 4.5 hours instead of 2. Yes, it is worth it.
The training that the architect alone runs
The architect runs the training solo. The architect is excellent at substance and weak at facilitation. The questions that needed to be drawn out are not drawn out. The training is technically complete and operationally fragile.
Fix: trainer + facilitator pairing. The architect partners with HR or comms for every session.
The training-and-done approach
The 90-minute training happens. The architect declares training complete. The reviewer cohort, three months in, has questions nobody is answering. Competence decays. Trust frays.
Fix: post-training cadence. Office hours, retrospective, refresher, listening sessions, new-hire onboarding. The training is the kickoff, not the conclusion.
Key Takeaways
- Default approach (async, one-time, voluntary) is wrong in three ways. The training must be live, 90 minutes, mandatory. Each correction is non-negotiable.
- Three audiences, three tracks: reviewers (operational muscle memory), escalation owners (incident judgment), stakeholders (mental model accuracy). Combining audiences into one session pleases nobody.
- The reviewer track's load-bearing 30 minutes is live practice on five real cases: clear approve, clear edit, clear reject, borderline, escalation. Attendees make the call individually, then compare โ the spread of answers is itself the artifact.
- The escalation owner track centers incident judgment: top five patterns, hands-on pause procedure, customer-facing roleplay, postmortem responsibility.
- The stakeholder track centers mental model: dashboard literacy, scope boundaries, external-facing language, feedback loops.
- The "I don't trust it" backlash has three roots: unaddressed concerns, surprise failures, identity threat. The training addresses each by establishing concern channels, calibrating expectations, and honoring the role transition.
- Trainer-plus-facilitator pairing. Trainer owns substance; facilitator owns the room. Solo trainer optimizes for content and loses the room; solo facilitator optimizes for connection and misses substance.
- The training is the kickoff, not the conclusion. Post-training cadence: week-one office hours, week-four retrospective, month-three refresher, quarterly listening sessions, ongoing new-hire onboarding.
- Signals that the training worked: approval-rate calibration in a defined range, structured rejection feedback, escalation precision, office-hours attendance, anonymous feedback quality, Slack channel engagement.
- Anti-patterns: slide deck as training, "trust the agent" message, ignoring role transition, single combined session, solo architect delivery, training-and-done. Each has a fix; the architect plans for each before launch.
Skill.re