The Shop AI Readiness Audit
Every owner who has read three Avoca case studies, watched a Rilla demo, and sat through a ServiceTitan Pantheon keynote arrives at the same Tuesday morning question: where do we actually start, and what do we fix first? The answer is not the tool. The answer is the audit. Before a single contract gets signed, before the marketing manager opens another vendor deck, the owner spends 90 minutes scoring the shop across 20 questions on a 1-to-5 scale. Data quality in ServiceTitan / Sera / Housecall Pro. CSR call volume and booking percent. Tech-scorecard maturity and the daily ride-along cadence. Marketing attribution depth. Owner time commitment. Financing close baseline. The composite score (out of 100) tells the owner three things: which AI bets pay back in 90 days, which require 6 months of plumbing before AI helps, and which the shop has no business buying yet. This lesson is the audit, the scoring, the diagnostic interpretation, and the roadmap-construction logic โ the artifact the owner walks into the L4 capstone with and the conversation a coach, peer group, or PE partner expects to have in the first 30 minutes.
Why the Readiness Audit Comes Before the Roadmap
The pattern of failed AI rollouts in trades shops in 2026 is consistent. The owner buys Avoca because the missed-call story is compelling. Six weeks in, Avoca is answering 100% of the calls but only booking 58% โ because the slot-offer logic depends on dispatch capacity data ServiceTitan was never set up to publish cleanly. The owner buys Rilla because the 18% close-rate lift is documented. Three months in, the close rate has not moved โ because the service manager has no daily routine for reviewing transcripts and the advisors have no comp tie to the scorecard. The owner buys Hatch because the stale-lead reactivation is real. Sixty days in, Hatch surfaces 1,800 stale leads but the CSR floor cannot absorb the call-back volume โ because the CSR pod was already stretched answering Avoca's transferred complex calls. None of these tools failed. The shop's readiness failed.
The readiness audit is the diagnostic that prevents this. It scores six dimensions on a 1-to-5 scale, names the gaps, and sequences the fixes against the AI bets that depend on each gap closing. A shop that scores 4-5 on data quality, CSR maturity, and tech scorecards can light Phase 1 of the roadmap in week one with a 90-day payback. A shop that scores 2-3 on the same dimensions needs three months of plumbing (CRM hygiene, CSR comp redesign, scorecard adoption) before the AI bet pays back. The audit is the owner's protection against buying a tool the shop cannot operate.
The audit is also the artifact a coach, peer group, or PE partner expects on the table. Nexstar, CertainPath, BDR, Service Champions, the Wrench Group portfolio reviews, and the PE diligence calls all open the same way: "show me your readiness score and the roadmap built off it." The owner who walks in with the scored audit and the dated roadmap has the credibility to talk about Avoca, Rilla, Dispatch Pro, Hatch, NiceJob, Podium AI Employee, Birdeye AI Employee, and Ryze AI as specific decisions with named owners and named metrics. The owner who walks in without it gets the same question every quarter and never finishes the conversation.
The Six Dimensions of Shop AI Readiness
The audit covers six dimensions. Each dimension has three to four scored questions for a total of 20. Each question scores 1 (broken) to 5 (top quartile). Total possible score: 100. The composite groups into four bands that map to roadmap pacing.
Dimension 1: FSM Data Quality (4 questions, 20 points). The data living in ServiceTitan, Sera Systems, or Housecall Pro is the substrate every AI tool reads from. If equipment records are missing, address fields are inconsistent, ticket statuses are stale, or the dispatch board is overridden without logging โ AI tools optimize on bad inputs and produce confidently wrong outputs. Question 1: Are equipment records (make, model, serial, install date, warranty) populated on 90%+ of active customer profiles? Question 2: Are dispatch overrides logged with a documented reason on 95%+ of overrides? Question 3: Is the RC&D (recall / callback / warranty) tagging discipline applied to 100% of go-back tickets within 24 hours? Question 4: Is the membership tier (member / non-member / lapsed) accurate on 98%+ of customer records?
Dimension 2: CSR Call Volume and Booking Maturity (3 questions, 15 points). Avoca, Jobber AI Receptionist, Housecall Pro AI Agents, and ServiceTitan Voice all depend on a CSR floor that can absorb the transferred complex calls AI bumps up. Question 5: What is the current booking percent on inbound service calls? (1 = below 55%, 5 = above 80%.) Question 6: What is the current missed-call percent? (1 = above 25%, 5 = below 8%.) Question 7: Is the CSR comp plan tied to booking percent, show rate, or just hourly wage? (1 = hourly only, 5 = comp tied to booking and show rate with monthly review.)
Dimension 3: Tech Scorecard and Ride-Along Maturity (4 questions, 20 points). Rilla, ResponsiBid, and ServiceTitan close-rate analytics depend on the service manager having a working scorecard routine and the advisors being open to coaching. Question 8: Is there a documented tech scorecard with weekly review? Question 9: Does the service manager run 30+ ride-along reviews per week (live or virtual)? Question 10: Is the close-rate baseline measured per advisor with a 90-day rolling average? Question 11: Is the comp plan tied to close rate, average ticket, or MPR (membership penetration rate)?
Dimension 4: Marketing Attribution Depth (3 questions, 15 points). CallRail Conversation Intelligence, GLSA AI bidding via Ryze AI, and Hatch nurture all require a marketing attribution chain that connects lead source to closed revenue. Question 12: Does every inbound call route through a tracked number with source attribution? Question 13: Does GLSA cost-per-booked-call get computed weekly by source? Question 14: Does the marketing manager produce a Friday recap that tracks channel ROAS and cost-per-booked-call against last week?
Dimension 5: Owner Time and Routine Commitment (3 questions, 15 points). The 12-metric dashboard, the daily 8-minute routine, the Friday 25-minute recap, the 90-minute quarterly review โ none of it operates without owner discipline. Question 15: Does the owner have a daily 8-minute dashboard routine running today? Question 16: Does the owner block 25 minutes Friday for the weekly recap? Question 17: Does the owner block 90 minutes per quarter for the strategic review?
Dimension 6: Financing Close Baseline (3 questions, 15 points). Wisetack, GreenSky, and Synchrony soft-pull integration plus the financing close rate on $5K+ jobs is the high-leverage AI bet. Question 18: Is financing offered at the kitchen table on 95%+ of $5K+ proposals? Question 19: Is the financing close rate on $5K+ jobs above 28%? Question 20: Is the soft-pull approved-up-to amount surfaced to the advisor before the close conversation begins?
The Scoring Rubric and the Four Bands
Each question scores 1 to 5. Total score: 100. The bands map to roadmap pacing.
Band A: 80-100 โ AI-Ready (top quartile). The shop has clean FSM data, a comp-aligned CSR floor at 75%+ booking, a working tech scorecard with weekly ride-along review, a Friday recap routine, an owner dashboard, and a financing baseline above 28%. Recommendation: light Phase 1 of the roadmap in week one. Avoca / Jobber AI Receptionist / HCP AI Agents goes live month one, ServiceTitan call summaries go live month two, Dispatch Pro / Sera scheduling activates month three. 90-day payback realistic on three of four Phase 1 bets.
Band B: 65-79 โ Mostly Ready (median-plus). The shop has decent FSM data, a CSR floor that books 65-74%, a tech scorecard that is reviewed monthly rather than weekly, partial marketing attribution, an inconsistent owner routine, and financing at 18-27%. Recommendation: 30 days of focused remediation on the lowest-scoring dimension before Phase 1 launches. Most common 30-day fix: CSR comp redesign (booking % tied) + service-manager scorecard cadence (daily 15 minutes). Then light Phase 1 month two. 6-month payback on the full Phase 1 stack.
Band C: 45-64 โ Foundational Gaps (median-minus). The shop has FSM data gaps (equipment records, RC&D tagging), a CSR floor at 55-64% booking with hourly comp, no working tech scorecard, weak marketing attribution, no owner routine, and financing below 18%. Recommendation: 90 days of foundational fixes before any AI tool gets bought. Specifically: FSM data hygiene project (equipment field completion + RC&D tagging discipline), CSR comp redesign and 30-day re-training, service-manager scorecard build-out, marketing attribution wiring (CallRail install plus GLSA tracking). Then Phase 1 launches month four. 9-month payback on Phase 1.
Band D: Below 45 โ Not Ready. The shop has broken FSM data, a CSR floor with no booking discipline, no service-manager review cadence, no marketing attribution, no owner dashboard, and a financing baseline that suggests no working soft-pull workflow. Recommendation: do not buy AI tools yet. Spend 4-6 months on operational fundamentals before the audit gets re-run. The most common driver of a Band D score is an owner pulled into the field who has not had time to build operating infrastructure. AI tools at this stage accelerate the chaos rather than the revenue.
Walking Through a Sample Audit โ a 6-Truck HVAC Shop
Take a real-shaped example. A 6-truck residential HVAC shop in Phoenix, $4.2M revenue, owner-operated, runs ServiceTitan with three CSRs on the phones, two service managers, six service techs, two Comfort Advisors, and a part-time marketing coordinator. The owner runs the readiness audit on a Tuesday morning. The scoring lands as follows.
FSM Data Quality: equipment records on 64% of customers (score 3), dispatch overrides logged on 40% (score 2), RC&D tagging on 70% of go-back tickets (score 3), membership accuracy on 92% (score 4). Subtotal: 12 of 20. CSR Maturity: booking percent at 68% (score 3), missed-call percent at 19% (score 2), CSR comp hourly only (score 2). Subtotal: 7 of 15. Tech Scorecard: scorecard exists but reviewed monthly (score 2), ride-along reviews at 8 per week (score 2), close-rate baseline measured per advisor (score 4), comp tied to close rate (score 4). Subtotal: 12 of 20. Marketing Attribution: tracked numbers on 80% of inbound (score 3), GLSA cost-per-booked-call computed monthly (score 3), Friday recap done inconsistently (score 2). Subtotal: 8 of 15. Owner Routine: no daily dashboard routine (score 1), Friday recap blocked sometimes (score 2), quarterly review skipped twice (score 2). Subtotal: 5 of 15. Financing: offered at 80% of $5K+ jobs (score 3), close rate at 22% (score 3), soft-pull surfaced inconsistently (score 2). Subtotal: 8 of 15. Total: 52 of 100. Band C.
The diagnostic is clear. The shop is not ready for Avoca or Rilla on day one. The 90-day remediation: complete equipment-record cleanup (target 90% by day 30), build the dispatch-override log discipline (target 95% by day 45), CSR comp redesign with booking-tied incentive (rollout day 30, comp month 2 paycheck), service-manager daily 15-minute scorecard routine (start day 1), CallRail full install with source attribution on 100% of numbers (day 60), owner's 8-minute daily routine starts day 1. Re-audit day 90: target band B (65+). Phase 1 of the roadmap launches day 100 with Avoca pilot on the residential queue.
The owner who walks into the Nexstar peer call with this audit, this diagnostic, and this 90-day plan has the credibility to talk about the AI roadmap. The owner who walks in saying "we should buy Avoca" does not.
How the Audit Feeds the 12-Month Roadmap
The audit is the input. The 12-month roadmap (Lesson 2) is the output. The translation is mechanical once the audit is scored.
Phase 1 of the roadmap (months 1-3) targets a booking percent lift from 65% to 80% and missed-call recovery via Avoca / Jobber AI Receptionist / Housecall Pro AI Agents plus ServiceTitan call summaries. The audit's CSR Maturity dimension (Dimension 2) and FSM Data Quality dimension (Dimension 1) are the gating scores. Below 12 of 20 on either, Phase 1 launches month 2 or month 3 instead of month 1. Below 8 on either, Phase 1 launches month 4 after remediation.
Phase 2 (months 4-6) targets a Revenue Per Truck (RPT) lift of 12-18% via Dispatch Pro / Sera / FieldEdge plus the service-manager scorecard cadence. The audit's Tech Scorecard dimension (Dimension 3) is the gating score. Below 12 of 20, Phase 2 requires service-manager training on the scorecard routine before the dispatch AI tool launches.
Phase 3 (months 7-9) targets a close-rate lift of 8-14 points and a recall percent reduction below 3% via Rilla, ResponsiBid, and AI-powered RC&D triage. The audit's Tech Scorecard dimension (Dimension 3) plus Financing dimension (Dimension 6) are the gating scores. Below 12 on Dimension 3 or below 8 on Dimension 6, Phase 3 launches month 10 or 11 instead of month 7.
Phase 4 (months 10-12) targets a GLSA ROAS lift of 30-50% on top of a 3-4x baseline plus the full owner dashboard via Ryze AI, CallRail, NiceJob, Podium AI Employee, Birdeye AI Employee, and Hatch. The audit's Marketing Attribution dimension (Dimension 4) plus Owner Routine dimension (Dimension 5) are the gating scores. Below 8 on either, Phase 4 launches month 13 or later โ or runs with a partial scope (CallRail and NiceJob only, deferring Ryze AI bid-management until attribution is complete).
The audit-to-roadmap translation is the difference between an aspirational plan and a defendable one. The owner who scores the audit honestly produces a roadmap that paces against the shop's real readiness. The owner who scores generously produces a roadmap that misses every milestone by 60-90 days and burns peer-group credibility in the process.
Re-Auditing Quarterly and the Trend Line
The audit is not a one-time exercise. It re-runs every 90 days. Quarter-over-quarter score movement is the trend line the owner watches as carefully as booking percent or RPT.
The Q1 score sets the starting band. The Q2 score reveals whether the remediation worked. A Band C shop that re-audits at Band B by Q2 has executed the remediation plan and is ready for the next Phase. A Band C shop still at Band C by Q2 has a remediation problem, not an AI problem โ the owner is buying tools that the shop cannot operate, and the conversation pivots to operational fundamentals before another contract gets signed.
The Q3 and Q4 audits track the AI investment against the roadmap. By Q3, a Band B shop that has executed Phase 1 and Phase 2 should be at Band A on Dimensions 1, 2, and 3 โ FSM data is clean, CSR is at 75%+ booking, tech scorecard is in daily rhythm. By Q4, Phase 3 and Phase 4 are in flight; Dimensions 4 and 5 (Marketing Attribution, Owner Routine) move from 8-10 to 12-14 as the dashboard and the Friday recap become habit. The composite score moves from 52 (start) to 78 (end of year one) on a successful execution.
The audit's compounding value is the quarterly accountability artifact. The owner reports the score to the coach, the peer group, the franchise HQ, or the PE partner. The trajectory tells the story. A 20-point year-one lift on the readiness composite predicts a 6-9 month payback on the AI stack and a 1.5-2x RPT lift by month 24. The audit is not just a starting diagnostic โ it is the operating-cadence artifact that compounds across the L4 capstone defense and the L5 platform reporting if the shop ever crosses into multi-location.
Key Takeaways
- The readiness audit comes before the tool decision. Most failed AI rollouts in 2026 trades shops are not tool failures โ they are readiness failures. Avoca, Rilla, and Hatch all assume operational substrate the audit names and scores before any contract gets signed.
- 20 questions across six dimensions, scored 1-5, total 100. FSM Data Quality (20 pts), CSR Maturity (15 pts), Tech Scorecard (20 pts), Marketing Attribution (15 pts), Owner Routine (15 pts), Financing Baseline (15 pts).
- Four bands map to roadmap pacing. Band A (80-100): light Phase 1 week one, 90-day payback. Band B (65-79): 30 days of remediation, Phase 1 month two, 6-month payback. Band C (45-64): 90 days of foundational fixes, Phase 1 month four, 9-month payback. Band D (under 45): do not buy AI tools yet, 4-6 months of operational fundamentals first.
- The audit-to-roadmap translation is mechanical. Phase 1 gates on Dimensions 1 and 2. Phase 2 gates on Dimension 3. Phase 3 gates on Dimensions 3 and 6. Phase 4 gates on Dimensions 4 and 5. Below the gating threshold, the Phase moves later or runs with partial scope.
- The audit re-runs every 90 days. Q1 sets the starting band; Q2 reveals whether remediation worked; Q3 and Q4 track AI investment against roadmap. A 20-point year-one lift predicts 6-9 month payback on the AI stack and 1.5-2x RPT lift by month 24.
- A 6-truck shop sample audit lands at Band C (52). Diagnostic: 90-day remediation on FSM data, CSR comp redesign, service-manager scorecard cadence, CallRail full install, owner's 8-minute daily routine. Re-audit day 90 targets Band B. Phase 1 launches day 100.
- The audit is the artifact a coach, peer group, or PE partner expects. Nexstar, CertainPath, BDR, Service Champions, Wrench Group portfolio reviews, and PE diligence calls all open with "show me your readiness score and the roadmap built off it." The owner who walks in without it has the same conversation every quarter and never finishes it.
- Honest scoring beats generous scoring. A roadmap built on a generous audit misses every milestone by 60-90 days and burns peer-group credibility. A roadmap built on an honest audit paces against real readiness and defends in front of the partner.
Skill.re