Identifying AI-Ready Steps vs. Human-Only Steps
Mapping the 38 steps was the substrate. This lesson is the decision logic on top of it. Which steps does AI run cleanly, which does it run partially, which does it stay out of entirely โ and how does the same step move between categories depending on ticket size, regulatory exposure, customer risk, and the role of the human in the loop? The answer is not a single rule. It is six decision tests applied per step. Repair authorization above $1,500 is human; a call summary on the same job is AI. The financing soft-pull explanation is human; the kitchen-table follow-up text the next afternoon is AI. The recall classification suggestion is AI; the final SLA-tier assignment is human. This lesson gives the manager the six tests, applies them across the 38 steps, and produces the operating logic the team uses on the dispatch board, the CSR row, the tech tablet, the advisor's prep, and the owner's Friday review. By the end the manager has decision rules they can defend to an owner asking why Step 19 has AI on the trigger but not the explanation, to a dispatcher asking why Step 8 has AI on the recommendation but not the final assignment, and to a Wrench Group HQ asking why corporate-mandated Step 32 review-response AI is not being auto-posted at this location.
Why Decision Rules, Not Categorical Bans
The first instinct of a manager building an AI workflow is to draw a hard line. AI handles inbound; humans handle the kitchen table. AI writes the summary; humans write the proposal. The hard-line approach is wrong in 2026 because every step contains sub-decisions on different sides of the line. Inbound is AI for routine, human for a 14-degree no-heat with asthmatic kids in the house. Kitchen-table close is human on the close itself, AI on the proposal narrative, AI on the financing math, AI on the rebate calc, human on the 25C credit explanation, AI on the follow-up text. Six sub-decisions in one step.
Categorical bans also break with scale. A 1-truck owner does every step personally; AI augments their own decisions. A 60-truck shop has a CSR floor lead, dispatch lead, tech captain, three advisors, service manager, marketing manager, owner โ the same step has different humans at different sub-decisions, and the AI posture reflects the org chart. The rules are scale-elastic; categorical bans are not.
The L3 discipline is to move from "AI or human" to "AI for this sub-decision, human for that one, AI-with-verification for the third." The six tests give the manager the criteria per step, per sub-decision, per shop scale. They produce defensible decisions because the criteria are explicit, and trainable decisions because the team can apply the same six tests to whatever new AI feature ships from ServiceTitan, Avoca, or Rilla next quarter. The tests outlast any specific decision.
The Six Decision Tests
Apply these six tests in order to every step on the 38-step map. The first test that yields a "human-only" answer wins; the step is human or AI-with-human-verification, not pure AI. If all six tests pass for AI, the step is a clean AI plug-in. If some pass and some fail, the step is clumsy โ AI on the parts that pass, human on the parts that don't.
Test One: Regulatory Exposure
Does this step touch language, decisions, or disclosures regulated by FTC, EPA, Reg Z, FCRA, TCPA, state contractor board, or another regulator? If yes, AI is excluded from the regulated language unless locked verbatim to verified inputs. The financing soft-pull explanation (Step 19) is regulated under Reg Z and FCRA; AI cannot draft APR, term, payment, or disclosure language โ training data is stale and hallucinated rates are regulatory events. AI triggers and surfaces the portal's approved-up-to amount; the human delivers the explanation reading portal output verbatim. Two-party-consent recordings (Step 31 review request, Step 33 inbound recall) are regulated under state wiretap law; AI's disclosure script must be the shop's vetted language locked verbatim. EPA 608 refrigerant handling (Step 15) is licensing-bound; AI cannot recommend a refrigerant action, only structure the tech's post-action notes.
The test is binary at the language level. If the language is regulated, AI cannot draft it from open generation; AI assembles only from verified locked inputs. The constraint enforces this โ "do not invent APR, do not invent rebate amounts, do not cite NEC code sections, do not draft warranty language." The constraint is not optional; it is the manager's regulatory firewall.
Test Two: Ticket Size and Customer Financial Stake
How much money is at risk for the customer? Below $1,500 the stake is modest and the AI's draft is recoverable. Above $1,500 the stake compounds โ financing terms, warranty obligations, replacement decisions โ and AI-only on customer-facing output creates dispute exposure the shop cannot recover from. As ticket size rises, the human verification step moves closer to the customer. A $400 repair quote: AI drafts, tech delivers. A $4,000 repair: AI drafts the talk-track, tech delivers, advisor confirms financing math before the customer signs. A $24,000 replacement: AI drafts the proposal, advisor reads at the truck, advisor verifies the 25C credit and Wisetack APR against portals, advisor delivers โ human at the financial decision, AI at the prep.
Shop-specific thresholds apply. A $35K-average-replacement shop applies the $1,500 rule to add-ons, not the base ticket. A roofing shop running $14K average tear-off applies it to the first-line repair quote, not the re-roof. Principle constant: rising customer stake, falling AI autonomy on customer-facing output.
Test Three: Licensing and Trade Judgment
Does this step require a licensed tradesperson to perform, sign, or supervise? If yes, AI is in support โ never in execution. The licensed tech performs the EPA 608 work, the journeyman electrician sizes the panel, the master plumber re-pipes the home. AI structures the voice notes, drafts the customer-facing narrative on what was done, surfaces the warranty registration form. The licensed pro signs.
Regulatory in some states, operational in all. Licensed work performed on AI suggestion that the unlicensed tech executed creates state contractor board exposure. The shop's license is the asset. The test asks: if this step goes wrong and the board reviews, does the licensed pro have a defensible role at the decision point? Yes โ AI supports. No โ AI is out.
Test Four: Customer Relationship and Emotional Judgment
Does this step require reading mood, hesitation, body language, or emotional state? If yes, AI is out at the moment of contact and in for prep and post-contact. Intro at the door (Step 13) โ human at the moment. Symptom interview (Step 14) โ human at the conversation; AI structures the notes after. Kitchen-table close (Step 22) โ human at the close; AI preps the talk-track and rebuttal library, Rilla scores after, but the close is human. Dispute apology (Step 35 above threshold) โ human at the language the customer reads; AI drafts and the owner edits.
Timing relative to customer interaction. Before: AI preps extensively. During: AI is out. After: AI structures and acts on what happened. The dispatcher hears tone the AI transcript loses. The advisor sees the hand drift toward the phone the AI cannot see. Customer-emotional moments are human territory; prep and post-record are AI.
Test Five: Data Completeness and Verification Cost
If AI produces the output, what does it take a human to verify? 5-second skim, AI is in. 5-minute rewrite, AI is out. The test forces honesty about whether AI saves time or just shifts it. Call summary (Step 7): 5-second skim โ clean. Proposal narrative (Step 22 prep): 90-second read with financing-math line-item check โ AI works but verification is non-trivial. Warranty terms section (Step 27): impossible to verify without pulling manufacturer language โ fails the test; shop pulls manufacturer language directly.
Data completeness corollary. If the prompt context contained verified inputs (portal APR, manufacturer warranty terms, EPA-approved refrigerant, AHJ-cited NEC section), AI assembles from verified inputs and verification is 5 seconds. If the context was incomplete, AI fills gaps with plausible-sounding output and verification becomes research. The manager designs prompts so verified inputs are in context โ collapsing verification cost to the level that makes AI worthwhile.
Test Six: Irreversibility and Blast Radius
If this step goes wrong, can the shop recover? Yes โ AI is in with verification. No โ AI is on suggestion only or out. Wrong on-my-way text: recoverable, dispatcher resends. Wrong recall classification: recoverable in 24 hours if Step 33 SLA review catches it; AI suggests, human confirms. Wrong $4,000 dispute refund: irreversible once posted; owner signs. Wrong refrigerant recommendation: irreversible until next service call and creates safety risk; AI does not recommend, the licensed tech does.
Blast radius scales with customers touched. One wrong on-my-way text โ one customer. A wrong review-response template across 200 reviews โ 200 customers. A wrong nurture sequence across 1,200 stale leads โ 1,200 customers. High-blast-radius applications get template-level review before deployment; low-blast-radius applications get individual review post-deployment.
Applying the Tests Step-by-Step
Here is the application of the six tests across the 38 steps. The result is the AI plug-in classification for the workflow map. Steps grouped by the test outcomes.
Clean AI Wins (All Six Tests Pass)
Step 2 (Answer rate): No regulatory language on the answer itself; no ticket size at this step; no licensing; no customer-emotional moment at the ring; verification is 4 p.m. CSR review; reversible. Clean. Step 6 (Confirmation send): Brand-voice locked; no regulatory; small blast radius. Clean. Step 7 (Customer record update / call summary): Internal-only; 5-second verification; reversible. Clean. Step 9-10 (Pre-arrival texts): Brand-voice locked; no regulatory; small blast radius. Clean. Step 12 (On-my-way text): Personalization at scale; reversible; small blast radius per text. Clean. Step 16 (Photo annotation): Internal-only at this step; tech reviews before the photo enters the proposal. Clean. Step 17 (Voice notes): 5-section structured output; tech verifies before submit. Clean with constraint (no invented readings). Step 28 (Departure notes): Same as Step 17 โ clean with constraint. Step 31 (Review request): Brand-voice locked text; reversible; opt-out compliant. Clean. Step 34 (Root-cause clustering): Internal-only analytical output; service manager reviews. Clean. Step 37 (Renewal touch): Personalization at scale; marketing manager edits the batch; opt-out compliant. Clean.
AI with Human Confirmation (Most Tests Pass, Judgment Needed)
Step 3 (Intent classification): AI tags; CSR or AI router uses tag for triage decisions. Step 5 (Slot offer and booking): AI offers from board; CSR or AI receptionist captures; 4 p.m. verification on slot, name, address, no fabricated promises. Step 8 (Dispatch board): AI re-evaluates every 10 minutes; dispatcher overrides 5-15% with documented reason. Step 11 (Tech pre-route prep): AI surfaces history and parts; tech confirms before route. Step 18 (Repair-or-replace frame): AI assembles 3-option talk-track from verified inputs; tech delivers, advisor confirms financing math. Step 21 (Membership pitch): AI drafts 30-second tailored pitch; tech delivers, customer signs membership form. Step 32 (Review response): AI drafts within 24 hours; service or marketing manager approves and posts within 48 hours. Step 33 (Recall classification): AI classifier tags with confidence score; service manager confirms within 8 business hours; SLA clock starts at confirmed tag. Step 35 (Customer recovery, routine): AI drafts apology and remedy for issues below threshold; manager edits and sends.
Human-Only (Test Failure on Licensing, Customer Relationship, or Regulatory)
Step 4 (Triage decision): Test 4 โ emergency triage requires dispatcher's read on volume, weather, crew. Step 13 (Arrival and intro): Test 4 โ human chemistry at the door. Step 14 (Symptom interview): Test 4 โ active listening, pattern recognition. Step 15 (System inspection): Test 3 โ licensed work, EPA 608, state contractor board. Step 19 (Soft-pull explanation): Test 1 โ Reg Z and FCRA on disclosure language; AI triggers and surfaces, human explains. Step 20 (Authorization above $1,500): Test 2 โ customer financial stake above threshold; customer signs. Step 22 (Kitchen-table close, the close itself): Test 4 โ eye contact, pacing, hesitation. Step 23 (Install execution): Test 3 โ licensed crew judgment. Step 24 (Parts decisions at the truck level): Test 3 โ field judgment on what fits. Step 26 (Customer signoff and payment): Test 1 and Test 4 โ Reg E and PCI; human witnessing. Step 27 (Warranty signing): Test 3 โ licensed installer signs. Step 35 (Customer recovery above threshold): Test 4 and Test 6 โ irreversibility; owner judgment.
AI Makes Things Worse Without Aggressive Constraint
Step 19 (Soft-pull explanation, if AI drafts the explanation): Reg Z and FCRA exposure. Constraint must lock AI to portal numbers; never draft APR or term language. Step 22 (AI-scripted close attempts): Breaks the human read at the table. Constraint must restrict AI to prep and post-call; never script the close moment. Step 25 (AI-fabricated readings in documentation): Becomes the warranty trail. Constraint must forbid invented readings; AI structures verified inputs only. Step 27 (AI-invented warranty language): Contradicts manufacturer terms. Constraint must lock to manufacturer's verbatim warranty language; AI does not draft warranty terms. Step 32 (Auto-posted review responses): Google flags inauthentic. Constraint requires human approval before post; never auto-post. Step 33 (AI-only recall classification): Mis-tag becomes liability. Constraint requires manager confirmation within 8 business hours; never auto-tag SLA. Step 35 (AI-drafted disputes without owner edit): Admits fault the shop doesn't own. Constraint requires owner sign-off above refund threshold. Step 38 (AI auto-closed renewals): Misses retention conversation. Constraint requires senior CSR call on flagged churners; never auto-cancel without contact attempt.
Worked Example: Financing Soft-Pull Explanation (Step 19)
This step illustrates the six tests most clearly because it sits across the AI / human line within a single 90-second interaction. The advisor stands at the door. The tech just finished diagnostic and the customer is open to replacement. The advisor's tablet shows the Wisetack approved-up-to that AI just surfaced โ $22,000 at 84 months, 8.99% APR, soft-pull only, no credit impact. The customer asks "what is a soft-pull and what does this do to my credit?"
Test 1 (regulatory): Reg Z and FCRA on the explanation โ AI cannot draft what a soft-pull is, FCRA-defined credit impact, or disclosure obligations. Fail. Test 2 (ticket size): $22,000 โ high stake. Fail. Test 3 (licensing): Lender licensing met by Wisetack. Pass. Test 4 (emotional): Customer is asking a direct question about their own credit. Fail. Test 5 (verification cost): Verifying an AI-drafted explanation against Reg Z, FCRA, and Wisetack portal takes 3+ minutes per draft. Fail. Test 6 (irreversibility): Misstatement is a regulatory exposure the shop cannot recover from without settlement or correction. Fail.
Five of six fail. AI is out on the explanation. What AI does is bounded: trigger the soft-pull, surface the approved-up-to amount, display Wisetack's disclosure verbatim on the tablet, and assemble the surrounding talk-track. The advisor reads the portal disclosure verbatim, answers the credit question from training, and pivots back to the AI-prepped proposal. AI did 70% of the prep work; human did 100% of the regulated language and customer-emotional moment.
Worked Example: Kitchen-Table Follow-Up Text the Next Afternoon
Different step, similar customer, completely different test outcomes. The Bel Air close was Tuesday evening. The customer said "let me sleep on it." The advisor's next move is the Wednesday-afternoon follow-up text. Should AI draft it?
Test 1 (regulatory): No regulated language. Pass. Test 2 (ticket size): Customer stake exists but the text commits to no financial decision โ it surfaces the option, references the proposal sent Tuesday, and offers a phone call. Pass. Test 3 (licensing): None. Pass. Test 4 (emotional): The emotional moment was Tuesday at the table; Wednesday's text is prep for the next moment, not the moment itself. Pass with voice-locked, constraint-bounded draft. Test 5 (verification): 30-second skim. Pass. Test 6 (irreversibility): Text can be deleted, re-sent, or corrected; blast radius one customer. Pass.
Six of six pass. AI drafts. The constraint locks voice ("direct, respectful, no scarcity, no first-name address, no exclamation points, under 90 words"), references the Tuesday proposal, names the financing payment math the customer saw, ends with one open question. Advisor reviews in 12 seconds and sends. Same customer, same $22K ticket โ but one step later with the emotional moment over and the regulated language locked in the proposal that already went out, AI is appropriate.
The Decision Rule Quick Card
For the manager who wants the rule on a Post-it: AI runs the step if all six tests pass. Test 1 (regulatory) fail โ AI out on regulated language. Test 2 (ticket size) fail above threshold โ AI out on customer-facing output without verification. Test 3 (licensing) fail โ AI in support only, never execution. Test 4 (emotional) fail โ AI out at the moment of contact, in for prep and post-record. Test 5 (verification) fail โ step is human. Test 6 (irreversibility) fail โ AI on suggestion only with human confirmation required.
The card lives next to the workflow map. New steps that land next quarter โ a ServiceTitan AI proposal feature, an Avoca outbound retention voice agent, a Birdeye auto-poster โ get the six tests applied before they go live. Tests outlast vendor decisions, vendor turnover, regulatory updates, and AI capabilities the shop has not seen yet.
The Manager's Defense in Three Conversations
By the end of this lesson the manager can defend AI plug-in decisions to three audiences. To the owner asking why we pay for Avoca if the CSR still verifies bookings: Steps 1-5 are AI (Tests 1-6 all pass for answer and book) and Step 7 is AI with 4 p.m. CSR verification (Test 5 โ 5 seconds; AI saves time net of verification). To a Wrench Group / Authority Brands / Apex Service Partners HQ asking why the corporate-mandated review-response AI auto-poster is off here: Step 32 fails Test 6 (irreversibility) and Test 1 (FTC endorsement guidelines on AI-generated reviews); local override memo cites the framework. To a Nexstar peer call asking how the shop moved close rate 8 points without changing comp: Step 22 is human at the close, AI at the prep and Rilla scorecard, with daily Step 30 service-manager review added โ the framework let AI in on the right sub-decisions without disrupting the kitchen-table moment.
Three audiences. Same framework. Tests produce defensible decisions because the criteria are explicit. The map produces defensible discipline because the steps are explicit. Together they make L3 โ manager owns the design โ qualitatively different from L2 where the role-holder owns the prompts. The design has to defend itself in front of audiences that did not see the work.
Key Takeaways
- Six decision tests applied per step, per sub-decision. Regulatory exposure, ticket size, licensing, customer emotional moment, verification cost, irreversibility. The first failing test determines whether the step is AI, AI-with-human-confirmation, or human-only.
- Categorical bans break with scale and sub-decisions. The same step contains 4-6 sub-decisions that fall on different sides of the AI line. The manager's job is to name the sub-decisions and apply the tests to each.
- Repair authorization above $1,500 is human; the call summary on the same job is AI. Test 2 fails on the authorization; all tests pass on the summary. Same call, two postures.
- Financing soft-pull explanation is human; the kitchen-table follow-up text is AI. Tests 1, 2, 4, 5, 6 fail on the explanation; all six tests pass on the next-afternoon text. Same customer, different step, different posture.
- AI runs in support, never in execution, on licensed work. EPA 608, state contractor board, financing-lender licensing. The licensed pro performs and signs; AI structures the notes and surfaces the prep.
- Customer-emotional moments are human at the moment, AI before and after. Intro, symptom interview, kitchen-table close, dispute apology. AI does the prep extensively and the post-record extensively; AI is out at the moment of contact.
- The Quick Card outlasts vendor decisions. New AI features ship every quarter; the six tests stay constant. The shop applies the tests to new capabilities before they go live, and the framework survives every vendor turnover and regulatory update.
- The manager defends the design to three audiences with the same framework. Owner, franchise HQ, peer-group call. Tests explicit, criteria explicit, decisions explicit. This is what distinguishes L3 (workflow design) from L2 (prompt usage) and L1 (Cardinal Rule).
Skill.re