The Platform AI Transformation Framework (Franchise Multi-Unit Rollout Playbook)
A 25-, 100-, or 450-location platform CEO does not "deploy AI." They run a sequenced rollout across four stages โ pilot location, reference location, wave 1 (5-10 sites), wave 2 (full portfolio) โ with stage-gate metrics that decide whether the next wave proceeds, holds, or rolls back. This is the operating playbook that Wrench Group runs across its HVAC and plumbing portfolio. Authority Brands runs it across the One Hour Heating & Air, Benjamin Franklin, and Mister Sparky franchise networks. Apex Service Partners runs it across acquired residential HVAC and plumbing brands. Sila Services runs it across the Northeast HVAC portfolio. Path Light Pro runs it across the electrical platform. Redwood Services runs it across the residential portfolio. ARS-Rescue Rooter runs it across the national service footprint. The playbook is not theoretical โ it is the structure that turned the 100-location Avoca deployment, the 80-location Rilla deployment, and the 60-location Dispatch Pro deployment into board-defended EBITDA contribution rather than 18-month tech debt write-offs. This lesson is the four-stage framework with stage-gate metrics, named owners, and the rollback criteria a PE board will accept. The board-defendable version of the AI rollout. The version the platform CEO walks into the quarterly review with on slide 7.
Why Platforms Fail Without the Stage-Gate Framework
The 2024-2025 platform AI graveyard is full of $1M-$3M write-offs. A 40-location HVAC platform signs a national agreement with an AI voice vendor, mandates simultaneous rollout across all 40 sites, and 90 days later discovers booking percentages dropped at 14 locations because the local CSR-floor maturity could not support the deployment cadence. The platform pulls back, the vendor relationship sours, and the operator-bench is now skeptical of the next AI initiative. The pattern repeats at three named platforms in 2024-2025 โ names blurred in coach circles but discussed openly at the IRA Service Industry Roundtable.
The failure mode is structural, not vendor-specific. The platform skipped the pilot stage, skipped the reference location, and went directly from contract signature to full-portfolio rollout. Without staged proof, the platform had no defense when the first locations showed early friction. Without stage-gate metrics, the platform could not distinguish "expected ramp friction" from "real deployment failure." Without rollback criteria, the platform either over-committed to a failing rollout or pulled back too aggressively from a recoverable one. The four-stage framework exists because every PE-backed trades platform that has run a successful AI rollout in 2025-2026 โ Wrench, Authority, Apex, Sila, Path Light, Redwood, ARS โ has converged on a similar structure. The naming differs across platforms; the structure does not.
The framework's purpose is operational, not theoretical. At each stage gate, the platform CEO answers three questions in front of the executive team and (every quarter) the PE board: did the stage produce the metric movement the thesis required, what changes does the next stage need based on what we learned, and what is the rollback trigger if the next stage falters. The framework forces the discipline that converts AI vendor adoption into EBITDA contribution. Without it, the platform is gambling capital against vendor pitch decks.
Stage One โ The Pilot Location
The pilot location is a single site chosen for one reason: it produces the cleanest signal. The criteria for pilot selection are non-negotiable. Site must have an experienced GM who can isolate AI's contribution from operational noise. Site must run the platform's standard tech stack (ServiceTitan, CallRail, NiceJob baseline) so AI deployment is incremental, not foundational. Site must have a CSR floor at minimum tenure of 6 months on the standard workflow. Site must have an above-platform-median baseline on the metric the AI is targeting (so any movement is attributable to AI, not to ramp-up). And the site must be one the GM volunteered, not one HQ assigned โ volunteers run pilots; conscripts run resistance.
The pilot duration is 60-90 days. Shorter than 60 days does not produce enough cycles to distinguish signal from noise; longer than 90 stalls the platform's rollout velocity. The pilot scope is bounded to one AI workflow at a time โ Avoca for missed-call recovery, or Rilla for ride-along coaching, or Dispatch Pro for board optimization, or Hatch for stale-lead reactivation. Bundled pilots produce ambiguous attribution. Single-workflow pilots produce clean metric movement the platform can defend in the board deck. The 60-90 day window also matches what most vendors will support with a dedicated success manager โ the lift is real only with the vendor's deployment discipline applied; the vendor's incentive to support that discipline ends roughly at day 90.
The pilot stage-gate metrics depend on the workflow. For Avoca: missed-call percentage (target movement from 22% baseline to under 8%), after-hours capture (from 0-15% baseline to 60%+), cost per booked call (track for ROI defensibility). For Rilla: close-rate lift on Comfort Advisor sales (target 8-18 points across the advisor team), virtual ride-alongs per manager per day (30-40 target), advisor feedback on coaching quality. For Dispatch Pro: revenue per truck per day (target $400-$800 lift), dispatch override rate (target falling toward 15% from 40% baseline), tech satisfaction score. For Hatch: stale-lead reactivation rate (target 30-45%), close rate on reactivated leads (target 18-28%), incremental closed revenue against the dormant pile.
The pilot stage gate produces a binary decision at day 90. Did the metric move enough to justify the reference-location investment? If yes, the platform advances. If no, the platform either kills the workflow (vendor switched, scope changed, thesis invalidated) or extends the pilot for another 30 days with a documented hypothesis on what needs to change. The "extend without hypothesis" path is what kills platform AI rollouts; the discipline at the first stage gate is the discipline that protects every subsequent stage.
Stage Two โ The Reference Location
The reference location is the pilot's stress test. The site chosen looks deliberately different from the pilot โ different geography, different brand (if multi-brand platform), different CSR floor maturity, different competitive density. The reference's purpose is to prove the lift is replicable across the heterogeneity that defines the rest of the portfolio. Wrench Group's reference for a Texas HVAC pilot becomes a Florida site under a different brand. Authority Brands' reference for a Mister Sparky pilot becomes a One Hour Heating & Air location in a different state. Apex Service Partners' reference for a plumbing pilot becomes an HVAC location at a recently-acquired shop.
The reference duration is 60 days, faster than the pilot because the deployment playbook is now informed by pilot learnings. Vendor success manager is the same person (continuity matters). Platform-side deployment lead is a Director of AI Operations or equivalent โ the role L5 Ch4 covers in detail. The deployment artifacts from the pilot become the templates: the CSR training script, the GM onboarding deck, the weekly stage-gate review format, the rollback playbook. The reference produces those artifacts in production form so wave 1 can use them without rewriting.
The reference stage-gate metrics are the same as the pilot's plus three additions. First, deployment time-to-value (how many days from contract to first lift measurement) โ target under 30 days for the second deployment vs. 60-90 for the pilot. Second, deployment cost variance (how much over or under the pilot's deployment cost ran the second site) โ target within 15%. Third, GM-to-GM transferability score (does the reference GM, surveyed at day 60, say they could run this rollout themselves on a third site without HQ deep involvement). The transferability score is the single most important number in the reference stage โ if a GM cannot operate the AI workflow without HQ holding their hand, wave 1 will fail at scale.
The reference stage gate decision is binary: do we have a replicable playbook that the next 5-10 sites can deploy with HQ playing a supporting (not driving) role? If yes, wave 1 proceeds. If no, the reference is extended or a second reference is added in a different geography to surface what made the first reference unrepresentative. The discipline at this gate is the discipline that protects the rollout from becoming dependent on HQ-level talent that does not scale to 50, 100, or 450 locations.
Stage Three โ Wave One and the Five-to-Ten Site Rollout
Wave 1 is the first scaled deployment. 5-10 sites simultaneously, chosen for portfolio diversity (mix of geographies, brands, CSR floor maturities, competitive densities). The wave runs on a 90-day cadence with weekly stage-gate reviews. Each site has a named deployment owner from HQ (the Director of AI Operations or a designated Regional Director), each site reports the same metric set on the same cadence, and each site escalates exceptions through the same governance structure.
The wave 1 stage-gate metrics expand from the pilot and reference set. Beyond the workflow's primary metric (missed-call %, close rate, RPT, recall %), wave 1 measures distribution โ what is the spread across the 10 sites on the lift measurement, what is the bottom-quartile site doing relative to the top-quartile site, and which deployment variables explain the spread. The spread analysis is what informs wave 2. If the bottom-quartile site is 40% of the top-quartile lift, the platform needs to understand whether the difference is CSR-floor maturity (correctable via training), local market difference (uncorrectable, requires segmentation in wave 2 expectations), GM sponsorship (correctable via comp tie-in), or tech-stack version drift (correctable via standardization investment).
Wave 1 also produces the platform's first true financial proof point. With 5-10 sites deployed, the platform CEO can present an aggregate margin contribution number at the next quarterly board review. Avoca lift ร 7 sites ร annual contribution = a defensible EBITDA line item. Rilla lift ร 8 sites ร close-rate ร ticket = a defensible incremental revenue line item. Dispatch Pro lift ร 6 sites ร yield ร days = a defensible incremental margin line item. The wave 1 financial proof is what unlocks the capital for wave 2; without it, the PE board treats AI as an expense rather than a margin lever.
The wave 1 stage gate produces three decisions at day 90. First, does the platform commit to wave 2 full rollout, hold at the wave 1 footprint, or roll back specific sites? Second, what segmentation does wave 2 require โ do we deploy to all remaining sites, or do we group sites by readiness tier and deploy in tranches? Third, what changes to the deployment playbook based on wave 1 spread analysis? The decisions get documented in a wave 2 plan that the platform CEO presents at the PE board meeting and that becomes the operating commitment for the next 6-9 months.
Stage Four โ Wave Two and the Full Portfolio Rollout
Wave 2 is the rest of the portfolio. For a 40-location platform, that is 30-32 sites. For a 100-location platform, 85-90 sites. For Wrench Group, Authority Brands, or Apex Service Partners at 200-450 locations, wave 2 is 180-420 sites deployed over 9-12 months in segmented tranches. The wave 2 deployment is not a single push โ it is 3-5 sub-waves of 20-50 sites each, sequenced by readiness tier (top-quartile sites first, then median, then below-median with remediation built into the deployment plan).
The wave 2 governance shifts. The Director of AI Operations does not deploy each site directly; they govern the regional deployment leads who deploy each tranche. The stage-gate review cadence becomes monthly per tranche, quarterly across the portfolio. The financial reporting becomes monthly board-deck material โ Avoca contribution by region, Rilla contribution by brand, Dispatch Pro contribution by truck-fleet type. The platform's quarterly synergy synthesis (L5 Ch5 covers this) ingests wave 2 deployment progress as a quarterly board KPI.
The wave 2 stage-gate metrics add three more layers. First, deployment velocity โ sites deployed per quarter against the deployment-velocity target the wave 2 plan committed. Second, lift consistency โ the spread between top-quartile and bottom-quartile sites narrowing as the playbook refines. Third, tool sprawl management โ is the platform also adding new workflows (a second AI tool, a new financing-AI integration, a new AEO publishing engine) at a velocity the operating bench can absorb. The tool-sprawl metric is the discipline that protects the deployment from becoming a perpetually-half-finished initiative.
Wave 2 ends in a transition rather than a stage gate. At full portfolio coverage, the AI workflow moves from "rollout" to "ongoing operations." Governance shifts from the Director of AI Operations as deployment leader to the Director of AI Operations as steady-state operator overseeing platform-wide quarterly tuning, vendor relationship management, and the next workflow's pilot stage. The full-portfolio AI deployment becomes a platform competency, not a project. The board's quarterly review shifts from "are we deploying" to "are we operating at the lift the deployment promised."
The Stage-Gate Metrics and the Rollback Playbook
Stage-gate metrics are the platform's defense against vendor capture and against the sunk-cost fallacy that kills bad deployments slowly. Each stage gate has a binary decision (proceed, hold, roll back) and three or four metrics that drive the decision. The metrics must be defined before the stage starts; defining them after the stage runs is how platforms rationalize failing deployments. The metrics must be tied to the platform's existing KPI dashboard (booking %, MPR, financing close %, recall %, GLSA ROAS, RPT) so AI contribution is measurable against the same yardstick the board uses for everything else.
The rollback playbook is the discipline most platforms don't write down until they need it. The rollback playbook covers: which sites get rolled back, who decides, what the customer-facing language is (if customer behavior changed during the AI deployment), what the vendor contract termination clauses are, what the recovery timeline is (typically 30-60 days), and what the post-mortem requirement is (what did the platform learn that the next workflow's pilot stage will reflect). Written rollback playbooks are common at Wrench, Authority, and Apex; the absence of one is a flag that the platform is not yet running the framework at the discipline level required.
The framework also produces a second-order benefit: it shifts the platform's vendor negotiation posture. A vendor selling into a platform that runs stage-gated rollouts commits to deployment success metrics, not just contract revenue. The vendor's customer success manager becomes accountable for the wave 1 lift measurement, not just for the contract renewal. The vendor's roadmap commitments tie to the platform's stage-gate cadence rather than to generic quarterly releases. The platform's purchasing leverage at the contract stage compounds quarterly as the framework demonstrates the platform's discipline to vendors and the platform's value to PE capital.
How the Eight Platforms Actually Run This in 2026
Wrench Group runs the framework across its HVAC and plumbing portfolio with ServiceTitan as the platform standard and Avoca as the primary voice AI deployment. The Wrench Group AI rollout in 2025-2026 hit wave 2 across most of the 100-plus location portfolio; the platform's quarterly board reviews now include an AI-ROI line on the EBITDA waterfall by region.
Authority Brands runs the framework across One Hour Heating & Air, Benjamin Franklin, and Mister Sparky franchise networks. The complexity is greater because franchise locations are independently owned โ the rollout requires both HQ-level standardization and franchisee-level adoption commitment. The wave structure adapts: pilot at a corporate-owned location, reference at a high-performing franchisee, wave 1 at the platform's top-quartile franchisees who opted in voluntarily, wave 2 across the remainder with HQ mandating selected tools as franchise-standard while allowing franchisee override requests on others (the override framework is the L5 Ch1 Lesson 3 topic).
Apex Service Partners runs the framework across acquired residential HVAC and plumbing brands with ServiceTitan plus Rilla plus CallRail as the standard stack. The framework runs differently at Apex because acquisitions add new locations to the portfolio quarterly โ the wave structure has to absorb the integration cadence. The post-acquisition 90-day integration plan includes the AI stack deployment, effectively treating each acquisition as a single-site wave that joins the platform's broader rollout cadence.
Sila Services has standardized on a ServiceTitan-and-Avoca core across its Northeast HVAC portfolio. Path Light Pro runs the electrical platform with a similar discipline. Redwood Services and ARS-Rescue Rooter each run portfolio-specific variations of the same four-stage structure. The names differ; the framework converges. The 2026 platform AI rollout playbook is no longer theoretical โ it is the operating standard across the eight named platforms running ~60% of trades PE deal flow.
For the multi-shop operator who is not yet at PE-platform scale (the 5-15 location independent), the framework still applies. Pilot at one location, reference at the second, wave 1 at the next 4-6, wave 2 at the remainder. The stage gates, the metrics, the rollback playbook โ all scale down. The discipline that protects a 450-location platform's AI rollout protects a 5-location independent's rollout equally well. The framework is the operating playbook, not the platform-only artifact.
Key Takeaways
- The four-stage framework: pilot location (60-90 days, one workflow, volunteer GM, above-median baseline site) โ reference location (60 days, deliberately different from pilot, produces replicable playbook) โ wave 1 (5-10 sites, 90 days, weekly stage-gate reviews, first board-defendable financial proof) โ wave 2 (rest of portfolio, 9-12 months, segmented tranches, monthly per tranche / quarterly portfolio governance).
- Each stage has binary stage-gate decisions: proceed, hold, or roll back. The metrics that drive the decision are defined before the stage starts, tied to the platform's existing KPI dashboard, and reported in the same format the PE board uses for every other operating review.
- The pilot's job is clean signal; bounded to one AI workflow, run at a volunteer-GM site with above-median baseline, with a binary day-90 decision. Extending a pilot without a documented hypothesis is the failure mode that kills platform AI rollouts.
- The reference's job is replicability; chosen deliberately different from the pilot, produces production-form deployment artifacts (CSR training, GM onboarding deck, weekly review format, rollback playbook), and is gated by GM-to-GM transferability score โ can the reference GM run a third deployment without HQ deep involvement?
- Wave 1's job is the first scaled financial proof; 5-10 sites, 90 days, spread analysis (top vs. bottom quartile), aggregate margin contribution defended at the PE board, and a documented wave 2 plan with sites segmented by readiness tier.
- Wave 2's job is full portfolio coverage in 3-5 sub-waves of 20-50 sites over 9-12 months. Governance shifts to monthly per tranche, quarterly across portfolio. Three additional metrics emerge: deployment velocity, lift consistency (spread narrowing), and tool-sprawl management.
- Rollback playbooks are non-negotiable: which sites get rolled back, who decides, customer-facing language, vendor contract termination clauses, recovery timeline, post-mortem requirement. Written rollback playbooks are common at Wrench, Authority, and Apex; absence is a discipline flag.
- The eight platforms converge on this structure: Wrench Group, Authority Brands, Apex Service Partners, Sila Services, Path Light Pro, Redwood Services, Leap Partners, ARS-Rescue Rooter. Names differ; framework converges. Authority Brands adapts for franchisee override; Apex adapts for acquisition cadence; Sila and Path Light run the structure straight.
- The framework scales down to 5-15 location independents. The discipline that protects a 450-location platform's rollout protects a 5-location independent's rollout equally well. The framework is operating playbook, not platform-only artifact.
- The framework's second-order benefit is vendor-negotiation leverage. A vendor selling into a stage-gated platform commits to deployment success metrics, not just contract revenue. Purchasing leverage compounds quarterly as the framework demonstrates platform discipline to vendors and value to PE capital.
Skill.re