โ†
AI for Skilled Trades & Home Services
Visionary ยท M14 ยท lesson 14 of 20 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Scaling Wins Across Locations Without Losing the Win
๐Ÿ“–
now learning

Scaling Wins Across Locations Without Losing the Win

15 min

A Granite Comfort booking-percentage lift, a Yost & Campbell missed-call recovery, an HL Bowman cost-per-conversion drop from $350 to $215 โ€” these are the 2025-2026 case studies trades AI vendors put on the cover of every deck. The platform CEO's job is not to celebrate them; it is to figure out which wins replicate across the rest of the portfolio and which are one-shop stories. Wave 1 of the rollout framework is where the platform finds out โ€” and where most platforms learn, expensively, that the pilot site's lift was 30-60% structural and 40-70% the GM, the local market, the CSR floor's maturity, the technician bench, the comp plan, or some other site-specific factor. This lesson is the discipline for scaling wins without losing them: the 4 failure modes โ€” local market difference, tech-stack version drift, CSR-floor maturity, owner sponsorship loss โ€” that explain almost every wave-1 underperformance in 2025-2026. The diagnostic that surfaces which mode is in play at which site. The remediation playbook by mode. The wave-2 segmentation that turns wave-1 spread analysis into a 30-, 100-, or 450-location rollout that holds the win at scale. The board-defendable version of "we are scaling AI without losing the original lift."

Why Scaling Wins Is the Actual Hard Part

The pilot stage produces a clean lift number at 3 sites under controlled conditions. The wave-1 stage exposes the lift to portfolio heterogeneity โ€” 5-10 sites across different markets, brands, GM tenures, CSR floors, and tech stacks. The honest 2025-2026 data: pilot-stage lift estimates carry forward at roughly 60-80% of the pilot value at wave 1 once heterogeneity is in play. The 20-40% lift-erosion gap is the cost of scale and the source of most wave-1 disappointment.

Platforms that walked into wave 1 expecting pilot-equivalent lift consistently underdelivered and lost board confidence. Platforms that walked into wave 1 with a 20-30% risk-adjusted discount on pilot estimates (the discipline from Lesson 1) and an active diagnostic for the 4 failure modes consistently met or exceeded the risk-adjusted target. The difference is not analytical sophistication; it is the discipline of expecting heterogeneity and having a remediation playbook before wave-1 deployment begins.

The Granite Comfort story illustrates. The 2025 Avoca deployment produced a missed-call rate drop from 24% baseline to 6% steady-state at the pilot location โ€” 18 percentage points, supporting wave-1 rollout to 5 additional locations. At wave 1, the spread surfaced. Two of 5 sites matched the pilot lift within 2 points. Two produced 9-12 point improvement (half the pilot lift). One produced only 4 points and required mid-wave-1 intervention. Aggregate wave-1 lift was 11 percentage points โ€” well below the pilot's 18, well above pre-pilot baseline, and exactly what the risk-adjusted commitment had targeted. The platform met the commitment because the discount and diagnostic were applied. The HL Bowman and Yost & Campbell stories follow similar logic. The lesson is not "Avoca works" โ€” it is "the platform's discipline at wave 1 determines whether the pilot win scales." For the platform CEO walking into the wave-1 stage-gate review, the board's question is precise: did we hold the pilot lift at wave 1, and if not, why not, and what is the remediation. The 4 failure-mode diagnostic produces the "why not" with named root cause; the remediation playbook produces the "what next" with named owner and timeline. Without the framework, the board hears excuses; with it, the board hears operating discipline.

Failure Mode One โ€” Local Market Difference

Local market difference surfaces when a wave-1 site's customer base, competitive density, household income mix, or seasonal demand pattern materially differs from the pilot site's. A pilot site in Houston with above-median household income and an evenly-distributed HVAC service-call profile produces an Avoca lift that a wave-1 site in El Paso with below-median income and peak-summer concentrated demand cannot replicate. The workflow's value depends on customer mix the wave-1 site does not have.

The diagnostic looks at four indicators. First, baseline primary-metric distribution โ€” if the wave-1 site's baseline missed-call rate is materially higher (32% vs. pilot's 22%) or lower (14% vs. 22%), lift potential differs. Second, household income mix โ€” Avoca's cost-per-conversion lift is larger when recovered calls convert to higher-ticket jobs; lower-ticket markets produce smaller dollar-lift even at equivalent percentage improvement. Third, competitive density and review velocity โ€” more competition means more leakage on missed calls and proportionally larger Avoca lift; weaker competition means smaller lift because the customer would have called back anyway. Fourth, seasonal demand pattern โ€” peak-concentrated demand produces different missed-call patterns than year-round demand.

Local market difference is the hardest failure mode to remediate because it is structural. The remediation is not "fix the deployment"; it is "adjust expected lift at this site to reflect market reality and segment wave 2 accordingly." For a Wrench Group operator running wave 1 across Texas (Houston, Austin, El Paso, San Antonio, Dallas), the diagnostic surfaces El Paso's structurally lower lift potential; the wave-2 plan deploys El Paso in a later tranche with adjusted expected EBITDA contribution. For Authority Brands, local market difference surfaces across brand-and-geography combinations โ€” a Mister Sparky franchisee in a Phoenix outer-suburb with below-median income and high competitive density produces different Avoca economics than a Benjamin Franklin franchisee in a Pacific Northwest market with above-median income. The franchise opt-in offer (L5 Ch1 Lesson 3) reflects the franchisee's local market. For Apex Service Partners' post-acquisition integration, local market difference is built into the acquisition diligence โ€” the M&A AI workflow (L5 Ch5 Lesson 1) ingests ServiceTitan/HCP data plus call recordings and estimates local market characteristics before close; expected post-close Avoca lift is sized against those characteristics, not platform average. The remediation is fundamentally segmentation: top-tier markets get full-workflow with full-lift expectations; median-tier markets get full-workflow with risk-adjusted expectations; below-median markets get scope-reduced deployment or deferred deployment with explicit re-evaluation criteria.

Failure Mode Two โ€” Tech-Stack Version Drift

Tech-stack version drift surfaces when wave-1 sites run different versions, configurations, or integration depths of the operator's standard tech stack (ServiceTitan, CallRail, NiceJob, financing portals) than the pilot site ran. The pilot site ran ServiceTitan with full Dispatch Pro configuration, CallRail Conversation Intelligence enabled, and full integration plumbing. A wave-1 site running ServiceTitan with only partial Dispatch Pro, no Conversation Intelligence subscription, and integration plumbing built 18 months ago by a prior CTO who left produces materially different workflow performance.

The diagnostic inventories the stack at each wave-1 site and identifies gaps. First, version compatibility โ€” is the ServiceTitan instance on the current release with the modules the AI workflow integrates with. Second, configuration depth โ€” does the site have required ServiceTitan configurations (job tagging, status codes, custom fields, automated workflows) at production discipline. Third, integration plumbing โ€” does the call routing, CallRail, financing portal, and ServiceTitan integration work at the data quality the workflow requires. Fourth, vendor subscription completeness across CallRail CI, financing portals, review platforms.

Tech-stack version drift is the most recoverable failure mode because gaps are addressable through targeted investment. Remediation is stack remediation before workflow deployment. For version gaps, the platform's IT team prioritizes the upgrade. For configuration gaps, a 2-4 week configuration project closes the gap. For integration gaps, the platform's integration plumbing team rebuilds or refreshes connections. For subscription gaps, procurement adds the missing subscriptions at master-agreement pricing.

The wave-1 stage-gate discipline is to require stack readiness as a deployment precondition. A wave-1 site that is not stack-ready does not deploy until gaps are closed. The wave-1 deployment plan includes named stack-readiness checkpoints at 30 and 60 days before workflow deployment. Sites that miss the checkpoint get deferred to a later wave-2 tranche with explicit stack-remediation funding allocated. For Apex Service Partners' acquired-brand integration, tech-stack drift is acute because acquired shops typically run different versions; the post-acquisition 90-day integration plan includes stack-remediation as a Day 1-45 priority, with AI workflow deployment scheduled for Day 45-75 once the stack is at platform standard. Skipping stack-remediation is the most common failure mode in less-disciplined acquisition integrations. For Wrench Group's owned-portfolio rollout, the platform's stack-modernization initiative runs in parallel with the AI rollout โ€” sites due for modernization get the modernization-plus-workflow deployment; sites already at standard get the workflow directly.

Failure Mode Three โ€” CSR-Floor Maturity

CSR-floor maturity surfaces when wave-1 sites have CSR floors at materially different tenure, training, or operational discipline than the pilot. The pilot site had a CSR floor with 18-month average tenure, weekly scorecard review, monthly Rilla-equivalent coaching, and a documented escalation playbook. A wave-1 site with a 6-month average tenure floor, sporadic scorecard review, no coaching cadence, and informal escalation produces materially different Avoca performance โ€” AI receptionist handoffs get mishandled, exception rate spikes, CX degrades, wave-1 lift erodes.

The diagnostic looks at five indicators. First, average CSR tenure โ€” sites under 12 months average tenure are typically not floor-ready. Second, scorecard cadence โ€” weekly vs. monthly. Third, coaching cadence โ€” structured coaching cycle (Rilla-equivalent or shop-specific) or coaching-by-exception. Fourth, escalation playbook discipline โ€” documented escalation tree for customer issues, billing disputes, recall calls. Fifth, exception handling โ€” can the floor absorb a 10-15% workflow exception rate without spiking customer-friction metrics.

CSR-floor maturity remediation is the slowest of the four failure modes because it requires investment in the floor itself โ€” hiring, training, retention, coaching, scorecard discipline. The remediation is not a 30-60 day fix; it is a 6-9 month floor-maturation effort. For sites with floor immaturity, the wave-1 plan defers AI deployment until the floor reaches maturity targets โ€” the platform invests in floor maturation as a precondition for AI deployment.

The Avoca illustration: deployed at a mature floor, Avoca produces the documented HL Bowman-style lift because the floor absorbs handoffs cleanly, manages exceptions efficiently, and reinforces the workflow's customer experience. Deployed at an immature floor, Avoca produces an apparent missed-call rate improvement that erodes within 6-8 weeks because the floor cannot maintain the discipline the handoff design requires. The wave-1 lift decays over the measurement window; the Day 90 stage-gate read is misleadingly positive. The Director of AI Operations runs the floor-maturity diagnostic at each wave-1 candidate site 60 days before deployment; immature-floor sites move to the floor-maturation pipeline. For Authority Brands' franchise rollout, CSR-floor maturity is the franchisee's responsibility under the franchise agreement; the platform's role is governance. For Apex Service Partners' acquired brands, the M&A AI workflow ingests CSR data and call recordings to estimate the acquired floor's maturity score; mature-floor shops deploy in the Day 45-75 integration window, immature floors get the Day 1-180 maturation program.

Failure Mode Four โ€” Owner or GM Sponsorship Loss

Sponsorship loss surfaces when wave-1 sites have GMs, Regional Directors, or owners who are skeptical, ambivalent, or actively resistant to the AI deployment. The pilot site had a volunteer GM who advocated for the workflow and made the on-floor decisions that supported the deployment. A wave-1 site with an assigned-not-volunteer GM produces deployment that is technically completed but operationally undermined โ€” the GM's signal to the floor is "do the minimum to get HQ off our back," the floor reflects the signal, and the lift never materializes.

The diagnostic looks at four indicators. First, GM volunteer-vs-assigned status โ€” assigned-not-volunteer GMs are at high sponsorship-loss risk. Second, GM's prior history with platform-mandated initiatives โ€” has this GM successfully executed prior initiatives, or consistently underdelivered. Third, pre-deployment engagement โ€” does the GM attend deployment planning meetings, ask questions, propose adjustments, or just show up to kickoff. Fourth, post-deployment week-2 signal โ€” does the GM's daily standup reference the workflow, advocate with the floor, address exceptions, or treat the workflow as someone else's problem. Sponsorship loss is the most expensive failure mode because it is recoverable only with intentional sponsorship investment before deployment. Once a GM has signaled "this is HQ's project, not mine" to the floor, recovery requires GM replacement, comp tied to workflow outcome, or HQ-level direct intervention โ€” all organizationally expensive and slow.

The remediation playbook starts at site selection. Wave-1 sites are selected for high-sponsorship probability (volunteer GMs, prior-initiative success record, pre-deployment engagement). For sites with sponsorship risk, the remediation is intentional sponsorship investment 30-60 days before deployment โ€” GM 1:1 with the platform CEO or COO to align on the workflow's value, GM comp adjustment to tie a portion of variable comp to workflow outcome, GM engagement in deployment planning at leadership-decision level, GM participation in the cross-portfolio wave-1 GM cohort that builds peer accountability.

For Wrench Group's owned-portfolio rollout, GM sponsorship is managed through Regional Directors who hold accountability for wave-1 deployment across 5-10 sites; Regional Director quarterly comp reflects deployment outcomes. For Authority Brands, owner sponsorship is at the franchisee level โ€” franchisee opt-in is the structural mechanism that filters for sponsorship; franchisees who opt in have it, those who do not are not deployed against. For Apex Service Partners, owner sponsorship is at the local sales leadership level โ€” the seller's leadership team that stays through the integration; acquisition diligence assesses leadership-stay probability, integration plan invests in leadership engagement at Day 1-30, AI workflow deployment at Day 45-75 reflects the leadership team's actual sponsorship signal.

The Wave-1 Diagnostic and Spread Analysis

The wave-1 stage-gate review at Day 90 produces the spread analysis โ€” the distribution of primary-metric lift across the 5-10 wave-1 sites. Aggregate lift is the headline; spread is the diagnostic data. A wave-1 with 8 sites all producing 14-18 point lift is a clean rollout; a wave-1 with 8 sites producing 4-22 point lift with high variance requires the 4-failure-mode diagnostic at each below-target site.

The diagnostic runs in 4 passes. Pass 1 reviews local market characteristics โ€” household income, competitive density, seasonal demand pattern, baseline metric distribution. Sites with material difference get categorized as Mode-1. Pass 2 reviews tech stack โ€” ServiceTitan version, CallRail config, integration plumbing, subscription completeness. Mode-2 sites flagged. Pass 3 reviews CSR floor maturity โ€” tenure, scorecard cadence, coaching cycle, escalation playbook, exception handling. Mode-3 sites flagged. Pass 4 reviews GM and owner sponsorship โ€” volunteer-vs-assigned, prior initiative success, pre-deployment engagement, post-deployment signal. Mode-4 sites flagged.

Most below-target sites surface 2-3 failure modes simultaneously. A site with local-market difference (Mode 1) and floor immaturity (Mode 3) presents a different remediation challenge than one with stack drift (Mode 2) alone. The diagnostic's value is identifying the combination, not just the individual mode. The remediation playbook addresses each mode in sequence โ€” stack remediation first (most recoverable), then sponsorship investment, then floor maturation, then market-adjusted lift expectations as the structural baseline. The wave-1 stage-gate board memo presents the spread analysis with the diagnostic. For each below-target site, the memo identifies modes in play, the remediation plan with named owner and timeline, expected lift recovery at remediation milestone, and the wave-2 segmentation implication. The board reads not a wave-1 disappointment but a diagnostic with remediation discipline.

Wave-2 Segmentation by Readiness Tier

Wave-2 deployment to the remaining 25-440 sites runs on tier-segmented cadence informed by the wave-1 spread analysis. Tier A sites have local-market characteristics matching high-lift pilot/wave-1 sites, tech stack at platform standard, mature CSR floors at 18+ months average tenure with scorecard and coaching discipline, and sponsored GMs with volunteer or prior-success status. Tier A deploys first at full-workflow scope with full-lift expectations.

Tier B sites have mixed readiness โ€” strong on 2-3 of the 4 dimensions, with 1-2 requiring remediation. Tier B deploys in tranche 2 with remediation built into the plan: stack drift remediated in the deployment window; sponsorship gaps invested in 30-60 days before deployment; floor immaturity gets 6-9 month floor-maturation runway with deployment scheduled at the floor-readiness milestone. Tier C sites have structural disadvantages on 2-3 dimensions โ€” typically local-market difference plus floor immaturity, often plus sponsorship risk. Tier C deploys in tranche 3 with scope-reduced workflow (the components that work at this site type) and adjusted lift expectations reflecting structural constraints; Tier C may also defer until constraints relax.

The tier segmentation discipline produces wave-2 deployment that holds aggregate lift at scale. Tier A deploys first and builds the wave-2 financial proof point quickly. Tier B follows with remediation-built deployment producing risk-adjusted lift. Tier C deploys last or defers, sized appropriately. The platform's wave-2 EBITDA contribution commitment is the sum of tier-appropriate expected lift โ€” defensible to the board because segmentation reflects operational reality rather than optimistic platform-average assumption. For Wrench Group's 100-site wave 2, the tier breakdown is typically 25-35 Tier A, 40-55 Tier B, 20-30 Tier C. Deployment cadence: 8-12 weeks for tier A across 25 sites, 12-16 weeks for tier B across 50 sites (with remediation in parallel), 16-24 weeks for tier C across 25 sites. Total timeline: 9-12 months. Aggregate lift: roughly 65-80% of pilot lift (Tier A 85-95%, Tier B 60-75%, Tier C 35-50%). Authority Brands' franchise wave 2 reflects franchisee opt-in plus readiness; Tier A franchisees opt in early, Tier B after seeing Tier A results, Tier C may not opt in until references mature. For Apex Service Partners, wave 2 is the integration cadence across newly-acquired brands plus existing portfolio โ€” each acquired brand enters at the tier acquisition diligence identified; the acquisition cadence (8-12 deals per quarter) drives continuous wave-2 deployment rather than a batched rollout.

How the Eight Platforms Actually Run This in 2026

Wrench Group runs the 4-failure-mode diagnostic at every wave-1 stage gate across HVAC and plumbing portfolios. The Director of AI Operations and Regional Directors share the diagnostic responsibility โ€” Director owns methodology and wave-2 segmentation; Regional Directors own site-level remediation. Wrench's 2025-2026 Avoca and Dispatch Pro rollouts both ran this discipline; wave-2 aggregate lift held at 70-80% of pilot lift, meeting the risk-adjusted commitment.

Authority Brands runs the diagnostic at franchise-system scale with franchise opt-in as the primary tier-segmentation mechanism. The platform's franchise consultants carry the diagnostic responsibility for each franchisee's wave-1 read. The override request memo provides the franchisee's path to deploy ahead of franchise-system pace; the franchise-standard mandate provides the platform's path to deploy when system economics demand.

Apex Service Partners builds the diagnostic into acquisition diligence and post-close 90-day integration. The M&A AI workflow estimates failure-mode risk pre-close; the integration plan addresses identified risks in the 90-day window; AI deployment at Day 45-75 reflects integration's success at risk mitigation. Apex's 2025-2026 acquisition pace (40+ acquisitions per year at peak) is enabled by this discipline; less-disciplined platforms cannot absorb acquisition pace because wave-1 failure rate compounds.

Sila Services runs the diagnostic in its cleanest form across owned Northeast HVAC โ€” no franchisee constraint, no acquisition pressure. The teaching reference for the methodology. Path Light Pro adapts the diagnostic for commercial electrical work (local market becomes commercial pipeline density, CSR floor becomes project-management maturity, GM sponsorship becomes Regional Director sponsorship of mission-critical bids). Redwood Services and ARS-Rescue Rooter each run portfolio-specific variations. Framework is consistent; calibrations differ. For the 5-15 location multi-shop independent, the same diagnostic applies at smaller scale: pilot at 2 sites, wave 1 at next 2-3, wave 2 at remaining 4-10. The 4 failure modes surface at small scale exactly as at platform scale; operator and COO/GM share diagnostic and remediation responsibility. The framework's scaling discipline produces the independent's compounding advantage against undisciplined competitors.

Key Takeaways

  • Pilot lift carries forward at roughly 60-80% at wave 1 once portfolio heterogeneity is in play. The 20-40% lift-erosion gap is the cost of scale. Platforms that expect pilot-equivalent lift consistently underdeliver; platforms that apply a 20-30% risk-adjusted discount and an active 4-failure-mode diagnostic consistently meet risk-adjusted targets.
  • Failure Mode 1 โ€” Local Market Difference: customer mix, competitive density, household income, seasonal demand pattern differ from pilot. Hardest to remediate because structural. Remediation is wave-2 segmentation with sites tiered by market characteristics and lift sized to market reality.
  • Failure Mode 2 โ€” Tech-Stack Version Drift: ServiceTitan version, CallRail config, integration plumbing, vendor subscriptions differ from pilot standard. Most recoverable. Remediation is stack readiness as deployment precondition โ€” sites not stack-ready get deferred until 30-60-day stack remediation closes the gap.
  • Failure Mode 3 โ€” CSR-Floor Maturity: CSR tenure, scorecard cadence, coaching cycle, escalation playbook, exception handling differ from pilot's mature floor. Slowest to remediate (6-9 month maturation). Critical because immature floors produce apparent lift that decays over the wave-1 measurement window.
  • Failure Mode 4 โ€” Owner or GM Sponsorship Loss: assigned-not-volunteer GMs, prior-initiative-underperformance history, low pre-deployment engagement. Most expensive because recoverable only with intentional sponsorship investment before deployment. Remediation: site-selection discipline plus 30-60 day pre-deployment investment.
  • The wave-1 diagnostic runs 4 passes โ€” local market, tech stack, CSR floor, sponsorship. Most below-target sites surface 2-3 modes simultaneously; the diagnostic identifies the combination and the playbook addresses each in sequence (stack first, sponsorship next, floor third, market-adjusted expectations as structural baseline).
  • Wave-2 segmentation by readiness tier: Tier A (full lift, deploy first) 25-35%; Tier B (risk-adjusted with remediation) 40-55%; Tier C (scope-reduced or deferred) 20-30%. Aggregate wave-2 lift averages 65-80% of pilot lift โ€” defensible because segmentation reflects operational reality.
  • Granite Comfort, Yost & Campbell, HL Bowman wins replicate only with the diagnostic discipline. The 2025-2026 case studies hold at wave 1 when the framework is applied; they erode when expectation is set at pilot-equivalent lift without diagnostic infrastructure.
  • Wrench, Sila, Apex run the diagnostic straight; Authority Brands adapts for franchisee opt-in; Apex builds it into acquisition diligence; Path Light Pro adapts categories for commercial electrical. Framework is consistent; calibrations differ.
  • The framework scales down to 5-15 location independents. Pilot at 2 sites, wave 1 at next 2-3, wave 2 at remaining 4-10. The 4 failure modes surface at small scale exactly as at platform scale. The discipline is the compounding advantage against undisciplined competitors.