AI for Skilled Trades & Home Services
Proficient · M6 · lesson 6 of 29 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI-Powered RC&D Triage — Recall, Callback, or Warranty
📖
now learning

AI-Powered RC&D Triage — Recall, Callback, or Warranty

15 min

Every shop in 2026 has a "go back" pile. Three tickets last Tuesday, one this morning, two on Saturday. Some of them are recalls — the original problem the customer paid us to fix is still broken, the customer is hot, and the clock started ticking the moment they hung up. Some are callbacks — the tech missed something on the first visit, the shop eats it, the customer is annoyed but not lost. Some are warranties — a manufacturer part failed inside the term, the shop bills the manufacturer, the customer waits in line. Industry median true-recall rate is 4-7% of completed service calls. Top-quartile shops run under 2%. The 200-basis-point gap is roughly $42,000-$68,000 of margin a year at a 6-truck residential shop, plus the brand-damage tail that does not show up on the P&L. This lesson builds the AI classifier that reads every "go back" job, tags it as recall, callback, or warranty in 8 seconds, and routes it to the right SLA tier — 24-hour response on recalls, 48-hour response on callbacks, 5-business-day response on warranties. By the end of the lesson the service manager has a prompt that runs at 4 p.m. daily, a dashboard the owner reads in 90 seconds, and a recall % that drops to under 2% inside 60 days because the leaks stop being invisible.

Why the Go-Back Pile Is the Most Ignored Leak in the Shop

The go-back pile lives in the dispatcher's head, the service manager's spreadsheet, and the customer's last text. It does not live in ServiceTitan as a labeled category. ServiceTitan, Sera, Housecall Pro, FieldEdge — none of them ship a default "this is a recall vs. callback vs. warranty" classifier. The dispatcher tags the return visit as a recall when the dispatch board prompts them; the service manager re-tags some of them when reviewing the Friday standup; the warranty admin moves the manufacturer-covered ones into a different bucket two weeks later. By the time anyone runs a report on recall %, the pile has been re-shuffled three times, and the number that lands in front of the owner is roughly directional and definitely undercounted. The 4-7% industry-median recall number is itself a charitable read of an underlying mess.

The undercount is structural. Front-line incentives push tickets out of the recall bucket. Dispatchers prefer "callback" because it carries less heat. Techs prefer "warranty" because the manufacturer covers labor. Service managers prefer "courtesy revisit" because it does not flag the tech's scorecard. None of these biases are malicious; they are how the shop's tagging has always worked. The cost is that the actual recall pattern — same tech, same part, same symptom — never surfaces because events get re-bucketed before they cluster. A shop running a stated 3.1% recall % may actually be running 5.4% — the median industry leak, undetected, costing $40K+ a year.

AI is built for this. Classifier prompts read the call notes, dispatch notes, original ticket, and customer message verbatim; they do not have a personal scorecard at stake, do not care which tech is named, and apply the definitional test consistently across all 50 go-back jobs this month. Output is a tagged, SLA-bound queue the service manager can act on by 4:15 p.m. The leak goes from invisible to visible. From there, the workflow that closes the leak is a separate question — the next two lessons in this chapter cover root-cause analysis and customer-recovery scripting. This lesson is the classifier and the SLA gate. Get the tagging right and the rest of the chapter compounds; skip it and every downstream report is dirty data.

The Three Buckets and the Definitional Tests That Separate Them

The classifier does one job: take a go-back ticket and put it in exactly one of three buckets. The definitions have to be tight enough that two service managers reading the same ticket land on the same bucket nine times out of ten. Loose definitions are the source of the undercount. Here are the operating definitions the classifier prompt enforces.

Bucket One: Recall — The Original Complaint Is Not Fixed

A recall is a return visit for the same complaint the customer paid the shop to fix. AC blowing warm, AC still blowing warm after the visit — recall. Drain backed up, drain still backed up after the snake — recall. Furnace short-cycling, furnace still short-cycling after the board swap — recall. Two-part test: (1) same customer, same address, same symptom, and (2) inside 30 days of the original ticket. The shop ate the diagnostic and repair labor originally, and now eats the return. Both customer perception and shop P&L treat it as internal failure. Customer is owed apology and same-day or next-day resolution; shop is owed a root-cause analysis to prevent the pattern.

What is not a recall: a new complaint at the same address ("you fixed my AC last week, now my furnace is acting up"), a different symptom on the same equipment, or a return more than 30 days out where the original complaint was resolved. Those are new tickets with their own diagnostic and pricing. The classifier makes this explicit so a confused message ("hey it's making a noise again") does not silently get bucketed as recall when it is actually new.

Bucket Two: Callback — The Tech Missed Something, Shop Eats It

A callback is a return triggered by something the tech missed, mis-installed, or did not communicate, where the original complaint was resolved but the visit produced a downstream issue. Tech replaced the capacitor, AC is cooling, but the wire was not properly seated and the unit shorts three days later. Tech cleared the drain, drain flowing, but the cleanout cap was not tightened and the basement leaks two days later. Tech swapped the thermostat, heat is on, but no schedule walkthrough and the homeowner calls back annoyed at 3 a.m. heat. Original complaint resolved; workmanship or communication produced the return; shop eats the labor.

The difference between recall and callback is not blame — both are internal failure. The difference is whether the original problem was fixed. Recall: original problem still broken — "you didn't fix what I paid you to fix." Callback: original fixed but something else broke downstream — "you fixed it but now something else is wrong because of how you fixed it." Callback gets a 48-hour SLA rather than 24-hour. Both hit the tech's scorecard, but recalls carry roughly twice the weight in coaching reviews.

Bucket Three: Warranty — Manufacturer Part Failure Inside the Term

A warranty visit is a return triggered by a manufacturer-covered part failing inside the manufacturer's term. Goodman compressor at month 18 on 10-year compressor coverage — warranty. Trane coil leaks at month 36 on 5-year coverage — warranty. Carrier blower motor inside 1-year parts — warranty. Two-part test: (1) failure is a defect attributable to the manufacturer (not the installer), and (2) inside the documented term for that component on that equipment. Manufacturer covers parts and sometimes a labor allowance; shop bills the manufacturer for what is covered and the customer for what is not. Customer is owed transparency on coverage, exclusions, and timeline — manufacturer claims take 5-10 business days, which is why the customer-facing SLA is 5 business days, not 24 hours.

What is not a warranty: installer-error failure (lineset brazed badly, leaks at 9 months — that is a callback, workmanship not manufacturer defect), out-of-term failure (compressor at year 12 on 10-year — customer pays), or excluded consumables (filters, belts, refrigerant top-offs). The classifier reads the ticket to identify whether the failure is on the manufacturer's covered list and inside the term; if either fails, ticket reclassifies to callback (workmanship) or full-pay (out of term). Misclassifying installer-error as warranty is one of the biggest hidden margin leaks — manufacturer denies, shop eats the labor anyway, and the coaching loop on the tech is broken because the ticket got tagged warranty rather than callback.

The SLA Tiers and Why They Are Tiered the Way They Are

Each bucket carries a service-level agreement on the customer-facing response time. The tiers are not arbitrary — they reflect what is reasonable to the customer, what is operationally executable, and what the shop's reputation can afford.

Recall: 24-hour response, same-day where physically possible. Customer paid for a fix; fix did not hold. In peak season the response has to be same-day or the customer goes online — Google review, Nextdoor post, BBB complaint. The 24-hour SLA is the floor; same-day is the target. Dispatcher pulls a slot for the originating tech (recalls go back to who did the original work — coaching plus customer trust); senior tech if originator unavailable. No rookie on a recall, ever.

Callback: 48-hour response, same-week target. Original problem resolved; customer annoyed about the downstream issue. 48 hours preserves trust; same-week is executable without breaking the board. Dispatcher routes to the originating tech where appropriate (wire-not-seated) or to a different tech where the original was the source (no-thermostat-orientation might route to service manager for the orientation, not the tech). 48-hour SLA gives the dispatcher planning room without leaving the customer feeling ignored.

Warranty: 5-business-day resolution with day-one communication. Manufacturer claim takes 5-10 business days; customer needs to hear from the shop on day one that the claim is filed, what is covered, the realistic timeline, and the workaround if the equipment is critical (loaner, repair-now-credit-later, temporary fix at no charge). The 5-day SLA is the resolution; the communication is same-day. No day-one touch is the second-most-common reason warranty customers go to a competitor mid-claim. Brand survives the wait if the customer hears early; brand does not survive silence.

The SLA enforcement is operational, not aspirational. The classifier prompt outputs a target response time on every ticket; the dispatcher's morning huddle reads the queue; the service manager's 4 p.m. review confirms every SLA is met for the day. Misses get logged and escalated to the owner if any single ticket breaches by more than 12 hours. The escalation rule is the discipline; without it, SLAs are wall-posters.

The Classifier Prompt — Built, Tuned, and Run at 4 p.m. Daily

Here is the working classifier prompt, written in the 5-part trades-AI structure from L2. The service manager runs it at 4 p.m. against the day's go-back pile, the AI returns a tagged, SLA-bound queue, the dispatcher pulls tomorrow's slots, and the customer-facing messages go out before the close of business. Total time from open-prompt to messages-sent: 18 minutes.

Role: Service Manager at a 6-truck residential HVAC shop running a 4 p.m. RC&D triage on today's go-back pile. AI is acting as the classifier; final tagging authority rests with the service manager, but the AI's first-pass label is the working draft for the dispatcher's morning routing.
Context: Industry-median recall rate is 4-7% of completed service calls; shop is targeting under 2%. Three buckets — recall (same customer, same address, same symptom, inside 30 days, original problem still broken — shop eats), callback (original problem resolved but tech missed something downstream — shop eats), warranty (manufacturer-covered part failed inside the term — manufacturer pays, day-one communication to customer). SLA tiers: recall = 24-hour response (same-day target), callback = 48-hour response (same-week target), warranty = 5-business-day resolution (same-day communication). I will paste the day's go-back tickets below — each ticket includes original visit date, original complaint, current customer message, equipment make/model/age, and the originating tech's name.
Task: For each ticket, output (1) bucket tag, (2) the specific definitional test that drove the tag, (3) target response time per the SLA tier, (4) recommended dispatch route (originating tech vs. senior tech vs. service manager vs. warranty admin), (5) any disqualifier or escalation flag (e.g., out-of-term warranty claim that needs re-bucketing to full-pay).
Format: Output as a table with one row per ticket; columns are Ticket ID, Bucket, Definitional Test, SLA, Route, Flag. No preamble. Below the table, a 3-line summary: today's recall count, today's callback count, today's warranty count.
Constraint: Do not invent ticket fields I did not paste. Do not assume a manufacturer warranty term — flag if the equipment age vs. term is unclear and route to the warranty admin for verification. Do not soften the recall count — call it as the definitional tests resolve. If a ticket message is ambiguous (could be recall or new ticket), flag it for service manager review rather than bucketing it.

The discipline is in the constraints. "Do not invent ticket fields" prevents the AI filling in equipment age or original visit date when missing — a frequent fabrication on incomplete records. "Do not assume a manufacturer warranty term" forces the warranty admin into verification rather than the AI hallucinating a 10-year compressor on a 5-year-coverage SKU. "Do not soften the recall count" is the structural defense against the undercount habit. "Flag if ambiguous" is the safety valve — ambiguity goes to the human, not the bucket.

Tuning happens at the Friday standup. The service manager reviews tickets where the AI's first-pass tag was overridden, identifies the pattern (AI consistently tagging brazing failures as warranty when they are actually callbacks), and updates the context or constraint line. Versions dated and archived; quarterly review confirms tuning. By month three, AI first-pass agreement with the service manager's final tag runs 92-96%, and manual override time per day drops from 18 minutes to 6.

Surfacing the Real Recall % and the Dashboard the Owner Reads

Once the classifier runs daily, recall % moves from a quarterly directional guess to a weekly real number. The Friday recap divides recalls by completed service calls for the period. The discipline is consistency. A shop running a stated 3.1% under dispatcher tagging routinely lands at 5.0-6.5% under the classifier — not because the work got worse, but because recalls stopped escaping into other buckets. The owner's first reaction is often defensive ("our recall % went up since we started AI"); the service manager reframes: the count did not go up, the bucketing got accurate. The shop was running 5.4% all along; now we can see it.

The dashboard has four lines. Line one: today's go-back count, by bucket. Line two: week-to-date recall % vs. target (under 2%) and vs. median (4-7%). Line three: SLA compliance — recalls answered within 24 hours, callbacks within 48, warranties communicated same-day. Line four: top three tech-by-part-by-symptom clusters from the AI cluster analysis (the next lesson). Ninety seconds to read, every Friday by 4:30 p.m. The owner walks into Monday's huddle knowing exactly which tech, which part, and which symptom drove the week.

Recall % is a leading indicator with a six-week lag on margin. A shop that lets recall % drift from 2% to 4% over a quarter sees gross margin per truck drop 3-5 points the following quarter — recall labor, rework parts, customer-recovery credits, lost referrals, and elevated review-response work compound silently. The owner reading the dashboard weekly catches the drift inside the lag and reverses it before margin shows bruised on the P&L. Classifier is the input; dashboard is the output; weekly review is the cadence.

The Handoffs and the Named Roles in the RC&D Workflow

The RC&D workflow has four named roles, each with one job. The service manager runs the 4 p.m. classifier and owns the queue. The dispatcher reads the tagged queue at the 7 a.m. huddle and routes the day's responses (originating tech where applicable, senior tech for recalls when originator is unavailable, warranty admin for manufacturer claims). The warranty admin owns the manufacturer claim filing and the day-one customer communication on warranty bucket tickets. The owner reads the Friday dashboard and signs off on any single-ticket SLA breach over 12 hours.

The handoffs are explicit and timed. Service manager classifies by 4:15 p.m.; queue is in the dispatcher's hands by 4:30; dispatcher's morning huddle covers the queue by 7:15 a.m.; warranty admin's day-one communications go out by 10 a.m.; SLA compliance reviewed at the 4 p.m. service-manager standup. Any ticket past its SLA tier escalates to the owner with a short note: ticket, bucket, hours over, reason. No exceptions — the discipline is the rule.

The owner's job is not to triage individual tickets. The owner reads the weekly dashboard, holds the service manager accountable for the under-2% target, holds the dispatcher accountable for SLA compliance, and breaks ties when a customer-recovery decision requires authority beyond the service manager (refund threshold, manufacturer dispute escalation, customer who has gone public on Google or Nextdoor). The handoff hierarchy keeps the owner out of daily triage and keeps daily triage from devolving into roll-ups — the operating failure mode in most shops without a real RC&D workflow.

What the Classifier Does Not Do, and What Comes Next

The classifier is necessary, not sufficient. It tags tickets and routes them to SLA tiers. It does not analyze why the recalls are happening — that is the next lesson, root-cause analysis with AI on recall tickets, where the named cluster cuts (by tech, by part, by symptom, by install date) surface the patterns the classifier creates the dataset for. It does not handle the customer-recovery scripting on a warranty exception or a contested recall — that is the third lesson in this chapter, the customer recovery playbook with AI-drafted apology and remedy scripts, escalation routing, and the refund-vs-credit decision tree. The three lessons compose the RC&D workflow as a unit; the classifier is the foundation under both of the other two.

The classifier also does not replace the service manager's judgment on edge cases. About 4-8% of go-back tickets land in the AI's flag-for-review bucket — ambiguous customer messages, missing equipment data, unclear original-visit attribution, warranty terms the AI cannot verify against the SKU. The service manager handles those personally, often with a 60-second call to the customer or originating tech. AI handles the 92-96% where definitional tests are clear; the human handles the residual. Without AI, the service manager spends 90 minutes on the queue every afternoon; with AI, 18 minutes on the 92% and a focused 25 minutes on the 8% that genuinely needs human attention.

The classifier's quiet effect is on shop culture. Once recall %, callback %, and warranty % become weekly numbers the owner reads, the dispatcher stops tagging recalls as callbacks, the tech stops protecting their scorecard at the expense of the data, and the warranty admin stops re-bucketing installer errors into manufacturer claims. Data quality improves because the data is suddenly visible. Visibility is the intervention; the classifier produces it. The shop that runs the classifier for 90 days runs a different culture by month four — and a different recall % by month six.

Key Takeaways

  • The go-back pile is the most ignored leak in the shop — industry-median recall % is 4-7%, top-quartile is under 2%, and most shops are running 5-7% with a stated 2-3% because the buckets get re-shuffled before anyone sees the real number. The 200-basis-point gap is $42K-$68K of margin per year at a 6-truck residential shop.
  • Three buckets, three definitional tests. Recall = same customer, same address, same symptom, inside 30 days, original problem still broken. Callback = original problem resolved but tech missed something downstream. Warranty = manufacturer-covered part failed inside the term. The classifier prompt enforces the tests consistently across every ticket.
  • SLA tiers: 24-hour recall, 48-hour callback, 5-business-day warranty with same-day communication. The tiers reflect what the customer expects, what the shop can operationally deliver, and what the brand can afford. Misses over 12 hours escalate to the owner.
  • The classifier prompt runs at 4 p.m. daily in the 5-part Role/Context/Task/Format/Constraint structure. Constraints prevent fabricated warranty terms, soften-the-recall bias, and over-confident classification on ambiguous tickets. Time from open to dispatched: 18 minutes. By month three, AI first-pass agreement with the service manager's final tag is 92-96%.
  • The Friday dashboard the owner reads is four lines — today's bucketed count, week-to-date recall % vs. target and median, SLA compliance percentages, top three tech-by-part-by-symptom clusters. Ninety seconds to read; six-week lead time on margin movement.
  • Four named roles, explicit handoffs, no overlap. Service manager classifies; dispatcher routes; warranty admin files and communicates; owner reads the dashboard and signs off on breaches. The hierarchy keeps the owner out of daily triage and keeps daily triage from escalating to the owner.
  • Recalls go to originating tech where possible, senior tech where not, never a rookie. Recalls carry roughly twice the scorecard weight of callbacks in the tech-coaching reviews. The dispatcher's morning routing enforces both rules.
  • The classifier creates the dataset for root-cause analysis — recall-by-tech-by-part-by-symptom clustering only works once recalls are accurately bucketed. The next lesson uses the classifier's output as the input to cluster recalls and find the pattern in 10 minutes, not 10 weeks.
  • Misclassifying installer-error as manufacturer warranty is a hidden margin leak — manufacturer denies the claim, shop eats the labor anyway, but the coaching loop on the tech is broken because the ticket was tagged warranty rather than callback. The classifier's warranty test catches this on day one.
  • Visibility is the operating intervention. The classifier's quiet effect is on shop culture — once recall %, callback %, and warranty % become weekly numbers the owner reads, the bucketing biases collapse. Month four runs a different culture; month six runs a different recall %.
  • The classifier is necessary, not sufficient. Tagging is the foundation; root-cause clustering (next lesson) and customer-recovery scripting (third lesson) ride on top. Together they are the RC&D workflow as a unit. Apart, they are pieces.