Running a 3-Location AI Pilot That a 30-Location Operator Can Trust
A 30-location operator does not buy AI on a slide deck. They buy it on a pilot that survives their CFO, their COO, and their PE-board operating partner reading the same numbers without a vendor in the room. A 3-location pilot is the right scope โ small enough to run with discipline, large enough to surface heterogeneity, statistically defensible enough to ground a 27-site capital commitment. This lesson is the pilot-design playbook an Apex Service Partners regional director, a Wrench Group operating partner, an Authority Brands franchisee council member, or an independent 30-location HVAC operator can run end-to-end in 90 days. Control vs. test architecture. The 6 metrics that matter. The 90-day cycle with weekly stage gates. The 30 / 60 / 90 kill rules that prevent sunk-cost rationalization. The statistical-significance math that holds up under a 3-location sample. The defense memo that turns the pilot result into a 30-site rollout commitment a board will sign.
Why a 3-Location Pilot Is the Right Design for a 30-Location Operator
A single-site pilot has an inherent attribution problem. One location's lift can be the GM's energy, a seasonal tailwind, a competitor's webpage outage, a temporary CSR-floor staffing event, or operational noise the 60-90 day window cannot distinguish from AI contribution. The board's reasonable response to a one-site result is "show me it isn't the GM." A 10-location pilot has the opposite problem โ deployment cost balloons to $400K-$1.2M, operator-bench runs out of bandwidth managing 10 simultaneous deployments, vendor success-management capacity gets thin, and the 90-day window does not mature clean comparison data at 10 sites with their own ramp curves. By the time the 10-site pilot reads cleanly, the operator has spent 18-24 months and the 27 remaining sites are still waiting.
The 3-site pilot produces statistically defensible signal in 90 days at $80K-$300K deployment cost (workflow-dependent), with bandwidth one Director of AI Operations can manage cleanly. Three sites surface heterogeneity (market difference, GM difference, CSR-floor difference) without overwhelming the deployment apparatus. They produce six data streams against the 6 metrics. They let the operator stage rollouts โ site 1 at week 0, site 2 at week 2, site 3 at week 4 โ so site 1 learnings reflect in site 2's and site 3's deployments.
The 3-site pilot also satisfies the social architecture of the rollout decision. The board sees 3 reads, not 1. The operator argues from a range (worst, average, best) rather than a point estimate. The other 27 GMs see sites that look enough like their own that wave-1 commitment is operationally credible โ not "the corporate flagship's result" but "the result we got across our portfolio." The vendor sees deployment discipline that justifies a deeper investment at the next-stage rollout.
Control vs. Test Architecture and Pilot Site Selection
The pilot's statistical defense depends on the control vs. test architecture. Without a control set, every lift number is contested. The CFO asks "what would the metric have done without the AI" and the operator's only answer is the prior 12-month baseline โ wrong because it ignores seasonal, competitive, and operational changes during the pilot window. The right answer is "the control sites' metric movement during the same 90-day window." Control vs. test converts a soft baseline comparison into a defensible attribution claim.
The control set is typically 3-6 sites drawn from the operator's portfolio that match the test sites on observable characteristics โ market type, brand, baseline metric range, CSR-floor maturity, GM tenure, fleet size. Control sites continue operating without the AI workflow during the 90-day window. The pilot's read is not "test sites moved X percent" but "test sites moved X percent vs. control sites moved Y percent during the same 90-day window, with delta (X-Y) attributable to the AI workflow after controlling for cohort drift." This is the language a CFO can underwrite and a board can defend.
Pilot site selection follows specific rules. Test sites should be middle-of-the-distribution on the baseline metric โ not top-quartile (lift suppressed by being already strong) and not bottom-quartile (lift confounded by operational issues unrelated to AI). Volunteer GMs, not assigned GMs. Standard tech stack โ ServiceTitan, CallRail, NiceJob baseline โ so AI deployment is incremental. CSR floors at minimum 6 months tenure. No other major operational changes scheduled during the window (comp plan reset, brand relaunch, major commercial contract starting). Control site selection mirrors test selection. Match documented before pilot starts; modifying control set during the pilot to chase a better result destroys statistical credibility. A Wrench Group 3-site Avoca pilot: 3 test sites in Texas across 2 brands at baseline missed-call rates 19-24%; 4 control sites in Texas at 18-25% baseline. An Authority Brands franchisee 3-site Rilla pilot: 3 test franchisees across One Hour, Benjamin Franklin, Mister Sparky at baseline close rates 41-46%; 4 control franchisees at 40-47%. Matching documentation goes into the pilot binder and the pre-registered analytical plan.
The Six Metrics and Why Each Matters
The 3-site pilot tracks 6 metrics across two layers โ the primary workflow metric and 5 secondary metrics that catch failure modes the primary alone would miss. Tracking only the primary produces ambiguous reads ("missed-call rate dropped, but bookings didn't move"); the 6-metric set forces full operational exposure.
Metric 1 is the primary workflow metric โ what the AI workflow is supposed to move. Avoca: missed-call percentage (target 22% baseline to under 8%). Rilla: close-rate lift on Comfort Advisor sales (target 8-18 points). Dispatch Pro: revenue per truck per day (target $400-$800 lift). Hatch: stale-lead reactivation rate (target 30-45%). The primary metric is the workflow's thesis; the pilot is a test of the thesis.
Metric 2 is the downstream conversion metric โ what the primary is supposed to enable. Avoca improves missed-call recovery; downstream is bookings per recovered call. Rilla improves close-rate; downstream is revenue per advisor-week. Dispatch Pro improves yield; downstream is RPT. Hatch improves reactivation; downstream is closed revenue per reactivated lead. Metric 3 is the customer experience metric โ does the workflow degrade CX while improving the operational metric. Track NPS at test sites, customer-complaint rate, review velocity during the pilot. A workflow that improves the primary while degrading NPS is one the operator should not roll out โ medium-term revenue loss exceeds short-term efficiency gain.
Metric 4 is the operator experience metric โ does the workflow create friction for the CSRs, dispatchers, techs, or advisors. Track operator satisfaction (5-point survey at weeks 4, 8, 12), workflow exception rate (how often AI output gets overridden), and operator turnover. A workflow that requires 20 minutes per shift of CSR override management, or causes a dispatcher to quit because AI board decisions are unreviewable, is a workflow that will not scale. Metric 5 is the deployment economics metric โ fully-loaded cost at steady state. Track per-site monthly tooling, support, time-to-steady-state, platform overhead allocation. A workflow producing 30% primary-metric lift at fully-loaded cost consuming 60% of the lift is a workflow whose ROI does not justify rollout. The deployment-economics read is the CFO's number; it must come from the pilot, not from a vendor quote.
Metric 6 is the deployment-discipline metric โ is the operator's apparatus ready to scale this workflow to 27 more sites. Track deployment time per site (target falling from 60-90 days at site 1 to 15-30 days at sites 2 and 3), deployment-cost variance (within 15% of plan), and GM-to-GM transferability score (can test-site GMs train next-wave GMs without HQ holding the bag). Without deployment-discipline readiness, the 30-site rollout fails at scale even if the primary metric is strong at pilot sites.
The 90-Day Cycle and the Weekly Stage Gates
The 90-day pilot runs on a structured cycle with weekly stage gates and three explicit decision points at days 30, 60, 90. The cycle is not "deploy and measure" โ it is "deploy, measure weekly, adjust within bounds, decide at stage gates." Weekly cadence prevents drift into 90 days of unsupervised activity that produces a Day-90 surprise.
Weeks 1-2 deploy site 1. Vendor success manager onboards the site, configures the workflow, integrates the tech stack, trains the CSR floor or advisor team, and runs the first week of live operation. The Director of AI Operations shadows the deployment and documents cost variance, operator-friction points, deployment-discipline observations. Weeks 3-4 deploy site 2 with site 1 in early steady-state. Site 2's deployment incorporates site 1's learnings โ CSR training refined, integration steps sequenced more efficiently, operator friction addressed in the deployment kit. Site 1's steady-state metrics track weekly; first weeks of test/control comparison data accumulate. Weeks 5-6 deploy site 3 with sites 1 and 2 in steady-state. The deployment kit is now production form. Three test sites running; three control sites running; six data streams accumulating against the 6 metrics.
Day 30 is the first stage-gate decision: were deployments on time and within budget? Did test sites show early signal on the primary metric? Did CX or operator experience degrade? Did any site need rollback? The Day-30 decision is binary per site: continue, hold, or roll back. A Day-30 hold or rollback at site 1 may delay or cancel site 3; documented and communicated to the board operating partner.
Weeks 7-9 are mid-pilot steady-state across all three sites. Data accumulates; primary-metric reads mature; secondary metrics surface failure modes or strengths. Weekly stage gates review data, operator-friction signals, vendor performance, emerging concerns. Adjustments within bounds (workflow tweaks, training refresh, additional vendor engagement) are made; structural changes (workflow scope, metric definition, control set) are not โ those invalidate statistical defense. Day 60 is the second stage-gate decision. Sites 1 and 2 have 4-6 weeks of steady-state data; site 3 has 2-3. Primary-metric movement is visible; secondary signals are surfacing. Is the pilot on track for a clean Day-90 read, or are interim signals suggesting kill, extend, or scope-adjust?
Weeks 10-12 are late-pilot data maturation. Primary reads stabilize; secondary metrics produce full signal; deployment economics finalize; deployment-discipline score completes. The board defense memo drafts in week 11, CFO and COO review in week 12, presentation at the Day-90 board meeting. Day 90 is the rollout-decision stage gate: commit to wave 1 (next 5-10 sites) at full investment, hold at the 3-site footprint for another quarter, or kill the workflow. Wave-1 commitment is the upside; hold is the middle path; kill is the discipline that prevents sunk-cost rollouts.
The 30 / 60 / 90 Kill Rules
The kill rules are the operator's defense against the sunk-cost fallacy. Without explicit kill rules, the 90-day pilot becomes a 180-day pilot becomes a 12-month pilot becomes a deployed workflow that never produced clean evidence. Kill rules are written before the pilot starts; modifying them during the pilot converts disciplined evaluation into vendor capture.
Day 30 kill rule: if any test site shows customer-experience degradation (NPS down 8 points, complaint rate up 50%, review velocity down 20%) or operator-experience collapse (CSR turnover exceeding 20% in 30 days, or daily exception rate exceeding 40% of workflow events), that site is rolled back immediately. The pilot may continue at the other 2 sites if site-specific, or the entire pilot may be killed if workflow-level. Day-30 kills are recoverable โ the workflow can be redesigned and re-piloted in 6-9 months; customer and operator-team relationships are protected.
Day 60 kill rule: if the primary metric has not moved at 2 of 3 sites by Day 60 (defined as moving in the right direction by at least 30% of target lift), the pilot is killed. The thesis has failed; further extension produces sunk-cost rationalization. The Day-60 kill is documented with the data; the post-mortem runs; the vendor's contract terminates per the rollback playbook; the operator captures learnings and applies them to the next workflow's pilot design. The Day-60 dual-trigger structure also fires if Site 3 lift falls below 25% of Site 2 lift โ that pattern proves portfolio-replicability failure, not workflow-ceiling failure, and kills wave 1 economics either way.
Day 90 kill rule: if the primary metric moved but downstream conversion (Metric 2) did not, or if deployment economics produce a fully-loaded ROI under 2.0x on a 24-month payback model, the workflow does not roll out. A 30% missed-call recovery improvement that does not produce additional bookings is not a workflow worth scaling to 27 sites. A close-rate improvement that doesn't translate to revenue is not a workflow worth scaling. The Day-90 kill protects the operator from committing $1M-$3M in rollout capital against a workflow whose thesis only half-works. The kill rules have explicit non-kill conditions. A Day-30 noise read (one site underperforming on a noisy weekly metric) is hold-and-watch. A Day-60 deployment-discipline issue (one site's deployment ran long) is a process-refinement flag, not thesis kill. A Day-90 primary read positive but below target range is hold-or-roll-forward-with-scope-reduction, not automatic kill. Kill rules are sharp; non-kill conditions prevent over-reactive killing of recoverable pilots.
Statistical Significance for 3-Location Samples
The 3-location sample is the most common source of CFO skepticism about pilot reads. With 3 test sites and 3-6 control sites, can the data actually produce a defensible attribution claim? Yes โ with the right analytical model, the right metric definitions, and the right sample-size math at the data-point level rather than the site level.
The first key insight is that the unit of analysis is not the site โ it is the data point. A 3-site pilot tracking missed-call rate over 12 weeks at each site produces 36 site-weeks per arm (test and control), or 72 site-weeks total. At each site-week, missed-call rate is calculated from typically 200-800 inbound calls โ so the call-level sample is on the order of 50,000-150,000 calls per arm over the 12-week pilot. This is more than enough to detect a 5-15 percentage-point shift with 95% confidence.
The second key insight is that the analytical model is difference-in-differences (DiD). The model computes the change in the primary metric at test sites from pre-pilot baseline to pilot steady-state, and compares it to the change at control sites over the same window. The DiD estimate is (test change) minus (control change) โ test sites' incremental movement net of background drift. Standard errors are computed using clustered standard errors at the site level to account for within-site correlation. The DiD estimate with clustered SEs gives the CFO the statistical defense.
The third key insight is that secondary-metric movement reinforces or undermines the primary read. If the primary moved by the DiD estimate AND downstream conversion moved in proportion AND CX held steady AND operator experience held steady, the pilot produced a coherent, defensible read. If the primary moved but downstream did not, or CX degraded, the pilot produced an ambiguous read that the kill rules should catch. The 6-metric framework is not just for tracking โ it is for triangulation. For the operator without an in-house statistician, ServiceTitan reporting produces site-week-level data; Excel or Google Sheets compute DiD with clustered SEs using built-in regression functions; vendor success-management teams often have a customer-analytics function. A 20-hour freelance analytics engagement to validate the model runs $5K-$15K โ rounding error against the rollout commitment. The statistical-defense bar is not "the pilot proves the workflow works in every market" โ that bar is unreachable in 90 days. The bar is "the pilot's data is consistent with the workflow producing the target lift, and the operator's downside risk in committing wave 1 capital is bounded by the kill rules and the rollback playbook." That bar is reachable.
Defending Pilot Evidence to a 30-Location Operator
The pilot's final deliverable is the board defense memo โ 8-12 pages, structured the way the operator's board reads operating reviews. It opens with the recommendation (commit to wave 1, hold, or kill); shows the data; documents the kill rules tested; presents the deployment-economics ROI; identifies risks and the rollback plan; commits the wave 1 plan with sites segmented by readiness tier; ends with the next-quarter operating commitment.
The opening recommendation is the most important sentence. A wave 1 commitment recommendation comes with the case โ primary-metric read with DiD estimate and confidence interval, secondary-metric coherence, deployment-economics ROI, deployment-discipline score. A hold recommendation comes with specific data justifying the additional quarter and the criteria that would convert hold to wave 1 (or to kill). A kill recommendation comes with the specific data that triggered the kill rule and the lessons learned the next pilot will reflect. The data section presents the DiD estimate for the primary metric with 95% confidence interval and clustered standard error. Charts show metric movement at test vs. control sites over the 12-week window. Secondary-metric charts follow. Operator-experience and customer-experience metrics are presented with the same rigor as the primary โ not as afterthought signals but as gating evidence.
The deployment-economics section presents per-site fully-loaded cost (tooling, support, integration, internal labor), per-site projected lift in EBITDA dollars at steady state, payback period, 24-month and 36-month NPV at the platform's hurdle rate. The CFO reads this section first; numbers must reconcile to the existing financial model without explanation. The wave 1 plan section commits the next 5-10 sites by name, deployment sequence, named owner, and expected lift. Sites segmented by readiness tier โ top-quartile first, median in tranche 2, below-median in tranche 3 with explicit remediation. Expected lift sized at the DiD estimate, discounted 20-30% for heterogeneity. The risk section identifies what could go wrong in wave 1 and the rollback plan โ CSR-floor maturity gaps, deployment-discipline failures, vendor-side risk (success-manager turnover, roadmap changes). Rollback covers which sites roll back, who decides, customer-facing language, vendor contract termination clauses, recovery timeline, post-mortem requirement. The memo closes with the next-quarter operating commitment: who owns wave 1, weekly stage-gate cadence, board check-in schedule, wave 1 stage-gate criteria, wave 2 trigger. The operator walks out of the Day-90 meeting with a signed wave 1 commitment, a documented governance cadence, and a vendor relationship that has demonstrated discipline.
How the Eight Platforms Actually Run This in 2026
Wrench Group runs 3-site pilots at the brand level (HVAC separate from plumbing) with Texas or Arizona as typical test geography, the 6-metric framework as standard, and a Director of AI Operations managing the deployment cadence. Wrench's 2025-2026 Avoca and Dispatch Pro rollouts both started as 3-site pilots that produced wave 1 commitment within 100 days of pilot start.
Authority Brands adapts for franchise networks. The pilot includes one corporate-owned location and 2 high-performing franchisees who volunteered, drawn from One Hour Heating & Air, Benjamin Franklin, and Mister Sparky. Control draws from comparable franchisee locations in the same regions. The Day-90 decision triggers the franchise-standard mandate (for HQ-mandated tools) or the franchise opt-in offer (for elective tools). The override request memo (L5 Ch1 Lesson 3) is the franchisee's path to add the piloted tool ahead of franchise-standard adoption.
Apex Service Partners runs 3-site pilots within the post-acquisition 90-day integration window. Acquired-brand sites become the test set; comparable platform-portfolio sites become the control set; the Day-90 decision converts into the Day-100 rollout commitment. The pace is faster because acquisition cadence demands it; discipline is the same. Sila Services runs the framework's cleanest form โ owned Northeast HVAC, no franchisee constraint, no acquisition pressure โ the teaching reference. Path Light Pro adapts metric definitions for commercial electrical (commercial vs. residential pipeline metrics, prevailing-wage analytics, commercial financing close rates). Redwood Services and ARS-Rescue Rooter run portfolio-specific variations. For the 5-15 location independent, the framework scales down to a 2-site pilot with 2-3 control sites. Discipline is the same; the "board" is the operator and the CFO or operating advisor. Scale changes; rigor does not.
Key Takeaways
- The 3-site pilot is the design point for a 30-location operator. Single-site pilots have attribution problems; 10-site pilots have cost and capacity problems. 3 sites surface heterogeneity, produce statistically defensible signal in 90 days, fit one Director of AI Operations' bandwidth, and give the board a range of reads to underwrite the 27-site rollout commitment against.
- Control vs. test architecture is non-negotiable. 3 test sites + 3-6 control sites matched on market type, brand, baseline metric, CSR maturity, GM tenure, and absence of confounding events. Matching documented before pilot starts; modifying control set mid-pilot destroys statistical credibility.
- The 6 metrics: primary workflow metric, downstream conversion metric, customer experience, operator experience, deployment economics, deployment discipline. Tracking only the primary produces ambiguous reads; the 6-metric set forces full operational exposure and triangulates the workflow's contribution.
- The 90-day cycle runs weekly stage gates with explicit Day 30, Day 60, Day 90 decisions. Weeks 1-2 deploy site 1; weeks 3-4 deploy site 2; weeks 5-6 deploy site 3; weeks 7-9 mid-pilot steady-state; weeks 10-12 late-pilot data maturation and board memo drafting.
- The 30 / 60 / 90 kill rules: Day 30 kills any site with CX degradation or operator collapse; Day 60 kills the pilot if primary metric hasn't moved at 2 of 3 sites OR Site 3 lift is under 25% of Site 2 lift; Day 90 doesn't roll out if downstream conversion didn't follow or fully-loaded ROI is under 2.0x. Kill rules are written before pilot starts; modification during pilot is sunk-cost rationalization.
- Statistical defense at 3 sites works because the unit of analysis is the data point, not the site. 36 site-weeks per arm ร 200-800 calls per site-week = 50,000-150,000 call-level observations per arm โ more than enough for difference-in-differences (DiD) estimation with clustered standard errors. CFO defense is the DiD estimate with 95% CI, not the 3-site point estimate.
- The board defense memo is 8-12 pages: opening recommendation, data section with DiD estimate, deployment-economics ROI, wave 1 plan with sites segmented by readiness tier, risk section with named rollback plans, next-quarter operating commitment. The memo converts a 3-site pilot into a 27-site rollout commitment.
- Wrench, Apex, Sila run the framework straight; Authority Brands adapts for franchise opt-in and override; Path Light Pro adapts metric definitions for commercial electrical; Redwood and ARS run portfolio-specific variations. The framework scales down to 5-15 location independents as a 2-site pilot with 2-3 control sites โ discipline is the same; the "board" is the operator and the CFO or advisor.
- The pilot's job is conversion, not validation. The 3-site pilot does not prove the workflow works at every site in the portfolio โ that bar is unreachable in 90 days. The pilot produces evidence that bounds downside risk, kill rules that catch failure modes, and a wave 1 commitment the board can sign. That is the bar a 30-location operator's pilot must clear.
Skill.re