The 30-Second Verify Habit
The L1 Cardinal Rule says verify everything that touches a customer. This lesson is the habit that operationalizes it. Thirty seconds. Five checkpoints โ numbers, names, parts, warranty terms, financing language โ applied every time AI generates a customer-facing, estimate, part-order, permit, or regulatory artifact. Every role applies it at their verify surface: CSR, dispatcher, tech, Comfort Advisor, service manager, marketing manager, owner. Shops that survive AI in 2026 build this habit in week one of every deployment. Shops that try to install it in month six compose the 88% non-embedded half of the ServiceTitan 2026 State of AI in the Trades report โ tool worked, workflow was thin, verify discipline absent, fabrication landed, owner pulled the plug, AI initiative dead for 12-18 months. This lesson is the muscle memory. Build it now.
Why Thirty Seconds and Not Three
The verify pass is named "30-second" on purpose. Five seconds is too short to catch fluent fabrications inside well-formatted prose โ the SEER rating two points high, the warranty year off by one, the APR a hundred basis points wrong, the NEC citation that looks right and reads wrong. Three minutes is too long to sustain across daily artifact volume in a $5M-$25M shop. The CSR row processes 60-120 inbound calls per CSR per day; the advisor ships 4-8 AI-drafted proposals; the dispatcher reviews 30-50 Dispatch Pro suggestions; the service manager reviews 20-40 Rilla scorecards across the week; the marketing manager queues 8-20 AI-drafted review responses on NiceJob, Podium AI Employee, Birdeye AI Employee, or Yelp AI. Three minutes per artifact compounds to hours of unbillable verify time per role. The discipline collapses by week two.
Thirty seconds is the sweet spot. Long enough to walk the five checkpoints; short enough to apply on every artifact without breaking the daily cadence. Practiced for 30 days, it becomes the muscle memory the artifact passes through on the way out โ the CSR does it as the Avoca confirmation loads, the advisor does it as the proposal renders, the marketing manager does it as the review response queues. The verify happens inside the flow, not as an interruption to it. The number is also forensic: when the owner audits why a fabrication shipped, the question is "did the 30-second pass happen?" Five seconds means it didn't; three minutes means the team is gaming the timer.
Checkpoint One โ Numbers
Numbers is the most common AI fabrication category and the highest-dollar exposure. Every number AI produces gets cross-referenced against source-of-truth before the artifact ships. Six number families a 2026 shop verifies daily, each with a named source-of-truth.
Equipment ratings. SEER, SEER2, AFUE, EER2, HSPF2 โ verified against the manufacturer's published spec sheet for the specific model and tonnage. A Trane XR16 at 16 SEER2 is not 18 because the AI wrote 18; it is 16 because the data plate says 16. Comfort Advisors verify the rating on every line of every replacement quote.
Refrigerant charge weights. R-454B is the 2026 verification minefield โ training data on the A2L transition is thinnest from 2023-2025 corpora; AI produces weights wrong by 1-3 lbs, and the certified tech who follows under-charges or over-charges. The fix is the EPA 608 compliance binder โ current manufacturer charge tables for every R-454B SKU in truck inventory, verified before the tech opens the gauges. R-410A is more forgiving because corpora are dense, but the rule is the same.
Federal tax credits. Section 25C (Energy Efficient Home Improvement Credit) caps at $1,200-$2,000/year depending on equipment category in 2026; Section 25D (Residential Clean Energy Credit) covers 30% uncapped through 2032. Both have changed annually since 2022. AI's memory is stale; the IRS guidance page is the source-of-truth. Customers hold the shop to the proposal's quoted credit; a $400 over-quote on a 25C heat-pump line produces a 1-star Google review six months later when the customer's CPA strips it.
State and utility rebates. PG&E, Con Ed, Duke Energy, Florida Power & Light, TVA, Xcel โ each utility runs its own current heat-pump rebate. State energy offices run additional programs (NYSERDA, Mass Save, Efficiency Vermont, Energy Trust of Oregon, AustinEnergy, Ameren). Rebate amounts change quarterly; the current 2026 program page is the source-of-truth, not AI memory. The fix is the 60-second tab open in the proposal workflow before the rebate goes into the proposal.
Financing payment math. Wisetack, GreenSky, Synchrony portal output is the only acceptable source. AI never calculates monthly payment, quotes APR, or produces term length. The advisor pulls portal output and copy-pastes verbatim. AI may write surrounding context ("here's why this financing fits your replacement") but never the regulated number.
Pricing and discounts. The pricebook is source-of-truth. AI does not invent line-item prices. If AI shows a $1,420 inducer-motor line and the pricebook shows $1,540, pricebook wins and the proposal is corrected before send. AI-generated discount language ("10% off through Friday") must match what the shop is actually running, not what AI thinks promotions look like.
Checkpoint Two โ Names
Names are the highest-frequency offender on AI voice-agent platforms โ Avoca, Jobber AI Receptionist, Housecall Pro AI Agents, ServiceTitan Voice โ and the most expensive when missed. The AI hears "Smith on Oak Lane" and produces "Smyth on Oak Avenue." The dispatcher routes to Oak Avenue. The tech arrives at the wrong house. The actual Smith on Oak Lane is on the phone wondering where the truck is. Two service slots gone; one 1-star review by 6:30 p.m.
The names checkpoint covers seven fields: customer first and last name; service address (street, suite, city, state, ZIP); phone on file; equipment make, model, serial; prior-call history (last visit date, last tech, last work). The CSR's 5-second skim runs the Avoca confirmation against caller-ID and the FSM platform's record; mismatches surface in 5 seconds. The dispatcher does a second 5-second pass when accepting โ confirming address ZIP is in territory and equipment matches tech-skill assignment.
AI-drafted customer summaries on the record carry the same checkpoint. If the AI summary says "inducer motor replacement on Goodman GMVC96" but the actual ticket was a capacitor on a Trane condenser, the next CSR reads bad context, the next dispatcher assigns the wrong skill, the next advisor walks into the kitchen with a fabricated history. Names accuracy compounds across the record; the 30-second pass at AI-summary generation is the brake on cascade. The 4 p.m. daily review of all AI-booked calls catches drift before tomorrow's complaints; the Friday standup surfaces patterns ("Avoca keeps mis-hearing apartment numbers") that become Monday system-prompt updates.
Checkpoint Three โ Parts
AI never orders parts. The rule is binary because the failure mode is binary โ wrong part arrives, truck rolls back, no-cool call is in day three, diagnostic fee gets refunded, tech eats $387 of unbilled work. Parts come from the supplier catalog (Carrier Enterprise, Johnstone Supply, RE Michel, Munch's Supply, ABCO HVACR), the manufacturer's part-lookup database, or the FSM platform's parts integration. AI suggests; the catalog confirms; the order ships only after cross-reference.
The checkpoint covers four sub-categories. Direct replacement parts (inducer motors, blower motors, capacitors, contactors, ignition modules, gas valves, expansion valves) โ supplier catalog every order. Refrigerant matches โ R-410A vs. R-454B is the catastrophic 2026 mismatch; data plate and AI suggestion both read against manufacturer spec before the cylinder connects. Filter dimensions โ off-by-one inch wastes the trip. Breaker amperage and wire gauge โ a 30A breaker on a 40A circuit fails inspection and creates a fire-risk citation; the licensed electrician verifies against load calc and current NEC jurisdiction adoption.
The 5-second supplier-catalog cross-reference is the discipline. Open the Johnstone tab, search AI's suggested part, confirm match to the data plate, place the order. Five seconds beats three hours of recovery on a wrong-part rebook. Shops building this in week one report 60-80% drop in wrong-part truck rolls within 30 days.
Checkpoint Four โ Warranty Terms
Warranty is the checkpoint that bites a shop in year 11 of a 10-year term. The Comfort Advisor's AI-drafted proposal says "12-year compressor coverage." Manufacturer's actual current term is 10. Customer signs. Three years later compressor fails. Customer is on the phone with the proposal saying year 12; the shop eats $4,200 because the proposal is the operative document, not the term sheet.
The manufacturer term sheet is source-of-truth. Goodman, Trane, Carrier, Lennox, Bryant, Rheem, York, Daikin, Mitsubishi, LG, Fujitsu, American Standard, Amana โ each publishes a current 2026 sheet. AI's memory is structurally stale; manufacturers revise registration windows, extended-coverage tiers, parts-only vs. parts-and-labor distinctions, and the training corpus sits 12-18 months behind.
The verify pattern is 5 seconds per warranty paragraph. The shop's pre-built warranty matrix โ a single PDF or shared spreadsheet that maps "manufacturer ร model ร installed-year ร registration-status โ warranty term" โ is open in the advisor's tablet. Every AI-drafted warranty line in the proposal gets read against the matrix; mismatches are corrected before send. The matrix is refreshed quarterly. The same checkpoint catches AI fabricating "lifetime heat exchanger" coverage on equipment that carries 20-year or limited-lifetime terms with specific exclusions. "Lifetime" without the qualifying clauses ("limited," "to original homeowner," "subject to registration within 60 days," "excluding commercial application") is a commitment the shop will be held to. The matrix carries qualifying language verbatim; the proposal copies the matrix verbatim; AI is not the source.
Checkpoint Five โ Financing Language
The financing checkpoint is binary and strictest of the five. AI does not draft financing language with specific numbers. Period. Wisetack, GreenSky, or Synchrony portal output is copy-pasted verbatim. FCRA adverse-action notices on soft-pull declines use the lender's signed template โ no AI generation, no exception. Two-party-consent disclosure language is the counsel-reviewed shop standard, configured per-state in the vendor (Avoca, CallRail, Rilla, ServiceTitan). AI may draft surrounding context โ "here's why this financing fits your replacement" โ but never regulated APR, term, fee, or disclosure language.
The exposure is the reason for the binary rule. Reg Z (TILA) statutory damages run $500-$5,000 per violation plus attorney fees plus CFPB exposure plus state TILA-analog exposure. A 100-bp APR drift on $18K over 84 months is roughly $1,260 in extra interest the consumer did not consent to โ the bait-and-switch the statute prevents. FCRA adverse-action notices have four required components (action, bureau ID with contact, score range with reason codes, dispute-rights notice); missing one is $100-$1,000 per violation plus actual damages plus attorney fees. Two-party-consent statutory damages run $1,000-$10,000 per call in CA, PA, FL, IL, MA, MD, CT, DE, MT, WA, NH, DC.
The fix is structural. The shop's financing-language template has a literal "[paste portal output here]" placeholder; the advisor pulls portal output, copy-pastes, ships, never edits placeholder content. AI's role is bounded to surrounding sales context. The same placeholder appears in the FCRA adverse-action template โ lender's signed template only. Two-party-consent disclosure is per-state vendor configuration confirmed at pilot (Avoca, Rilla, CallRail, ServiceTitan); operator confirms against the state map; quarterly audit re-confirms. AI never generates the disclosure on the fly.
Role by Role โ Where the Checklist Lands
The five checkpoints apply identically across roles, but each role's daily artifact mix concentrates on different checkpoints. Knowing which checkpoints dominate the role builds the muscle memory faster.
CSR โ Names Dominate, Numbers Second
The CSR's verify surface is dominated by names (Avoca/Jobber/HCP AI Agents/ServiceTitan Voice booking confirmations) and secondarily by numbers (dispatch-fee quote, service-window quote, after-hours pricing). The 5-second skim per booking confirms caller-ID matches AI-captured name; address ZIP is in territory; equipment make matches caller statement; no forbidden promises in the AI summary ("free estimate" on a $79-diagnosis shop, "same-day guaranteed" when the territory is booked). The 4 p.m. daily review of all AI-booked calls catches drift before tomorrow's complaints. Patterns drive system-prompt updates.
Dispatcher โ Names, Routing, Override Log
The dispatcher's primary verify is on ServiceTitan Dispatch Pro / Sera Systems / FieldEdge AI routing suggestions. Does the suggested tech actually fit the call (skill, comp plan, callback originator)? Are stacked install crews preserved? Is the recall on the originating tech? Each override gets a documented reason โ comp-plan, install crew, recall risk, customer request โ logged at every override. Target override rate 5-15%. Above 25% suggests Dispatch Pro is mis-configured; below 3% suggests the dispatcher is rubber-stamping. The override log is reviewed at Friday standup; patterns drive system configuration and tech-skill-tag updates.
Tech โ Parts and Numbers
The tech's verify is concentrated on parts (supplier catalog cross-reference) and numbers (charge weights, voltage readings, code-cited amperage). AI-drafted voice notes get a 30-second pass for equipment correctness, condition stated accurately, recommendation defensible. Repair-vs-replace prompts get a 30-second pass for repair-cost number, financing payment, warranty term, rebate amount โ all against source-of-truth. The tech reads the prompt to the customer; the verify happens in the 30 seconds before reading it.
Comfort Advisor โ All Five Checkpoints Every Time
The Comfort Advisor faces the highest-stakes verify surface โ $8K-$45K residential, $25K-$250K commercial proposals. Every line of every AI-drafted proposal gets all five checkpoints. SEER/AFUE against spec sheet. Equipment names against the order. Parts against supplier. Warranty against the manufacturer matrix. Financing against the portal. The owner has final authority on customer-facing proposals above $5K โ meaning the proposal does not leave the advisor's tablet without a verify-pass timestamp; on $14K+ replacement work, an owner or service manager signoff. The advisor's job at the table is to close; the verify lives in the 60 seconds at the truck before they walk in.
Service Manager โ Rilla Scorecards, RC&D Triage
The service manager verifies Rilla scorecards and AI-drafted ride-along commentary plus AI-tagged RC&D triage. Does the Rilla coaching commentary match what happened in the ride-along? Is the flagged moment genuinely a coaching opportunity? Is the RC&D tag correct (true recall vs. callback vs. warranty)? Is the SLA tier right (24-hour for recall, 48-hour for callback, 5-business-day for warranty)? Without the verify, AI-tagged recalls flood the wrong SLA bucket and the shop's true recall rate stays invisible.
Marketing Manager โ Review Responses, Outbound Consent
The marketing manager's verify surface is NiceJob, Podium AI Employee, Birdeye AI Employee, Yelp AI review responses plus Hatch nurture sequences. The 30-second pass on review responses confirms voice matches the shop, specific complaint is acknowledged, no commitment language ("we will refund," "we guarantee," "this won't happen again"), no fabricated remedies. Either manager pre-post review or 24-hour holding window. On Hatch outbound, opt-out and DNC status confirmed from CRM source-of-truth at send time, not local cache.
Owner โ Verify the Verify
The owner's verify surface is the discipline itself. The owner does not verify every artifact โ the owner verifies that the verify discipline runs. Monday-morning AI dashboard surfaces missed-call recovery, recall flags, dispatch overrides, complaint flags. The weekly AI failure-log review (30 minutes, service manager prepares) shows artifacts caught at verify, patterns emerging, system-prompt updates queued. The quarterly governance review (90 minutes) decides tool retention, vendor changes, policy updates.
Building the Habit in Thirty Days
The 30-second habit is built in 30 days, not 30 minutes. Same protocol every new role hired, every new tool deployed. Day one: team reads the failure-mode lesson (L1 Ch2 L2), the Cardinal Rule (L1 Ch2 L3), and this lesson aloud. Five checkpoints recited as a team. Wall poster up. The "AI Caught a Hallucination" bulletin board (Lesson 2 next) cleared and dated.
Day two through five: every AI artifact every team member touches gets the 30-second pass, observed by a buddy. CSR shadowed by service manager; dispatcher by ops manager or owner; Comfort Advisor by owner or sales manager; marketing manager by owner. Five days of buddy-observed verify builds muscle memory under supervision. Catches go on the board with date, role, artifact, remediation.
Day six: team meets 20 minutes. What did the pass catch? What patterns emerged? What system-prompt updates are queued for Monday? The 20-minute review converts individual catches into institutional learning. Day seven: habit is institutional; buddy shadow drops to weekly random sampling.
Day eight through thirty: muscle memory. The CSR's 4 p.m. daily review is on the calendar. The dispatcher's override log is logged at every override. The advisor's pre-kitchen-table verify is in the workflow. The service manager's Rilla and RC&D verify is end-of-day. The marketing manager's review-response verify is at posting or 24-hour hold. The owner reads the weekly AI failure log on Friday. By end of week four, the habit is institutional; new hires are onboarded into it as non-negotiable.
The 30-day window is the discipline cliff. Behavior literature converges on 21-30 days for repeated behavior to become automatic; in trades shops the variable is daily-artifact volume โ at 60-120 artifacts per role per day, the pass becomes muscle memory in 30 days or it doesn't. After 30 days without the habit, the default has shifted to "AI handles it," fabrications have shipped, commitments have been honored, and re-installing requires undoing internalized behavior. Install in week one or pay double in month six.
The Pass and the Published Lifts
The verify pass is not friction โ it is the discipline that allows the published lifts to land. Avoca's HL Bowman case (100% answer rate, 70% YoY revenue growth, cost per conversion $350 to $215) compounds in shops with the pass running and rolls back in shops without it. Rilla's 18% close-rate lift compounds where the service manager verifies scorecard commentary before coaching huddles and rolls back where the manager rubber-stamps. ServiceTitan Dispatch Pro's 12-18% yield lift compounds with override-log discipline and stalls without it. Hatch's 30-45% stale-lead reactivation compounds where outbound consent is enforced at send and produces a TCPA class action where it is not.
The 60-point gap in the ServiceTitan 2026 report (72% relevance, 12% embedded) is primarily a verify-discipline gap. Tools exist; workflows are documented. Shops that build the habit in week one capture the lifts; shops that skip produce the rollback statistics. The 30-second pass converts vendor capability into shop-level metric movement. Build it now; compound for twelve months.
Key Takeaways
- Thirty seconds, five checkpoints โ Long enough to catch fluent fabrications, short enough to scale across daily artifact volume. Five seconds misses; three minutes collapses by week two.
- Checkpoint one โ Numbers โ SEER/AFUE against spec sheet, R-454B charge against manufacturer table (thinnest training data in 2026), Section 25C/25D caps against current IRS, state/utility rebates against current 2026 program pages, financing payment math against Wisetack/GreenSky/Synchrony portal, line-item pricing against pricebook.
- Checkpoint two โ Names โ Customer name, address, phone, equipment make/model/serial, prior-call history. Avoca confirmation skim by CSR; second pass by dispatcher accepting the booking; 4 p.m. daily review of all AI-booked calls.
- Checkpoint three โ Parts โ AI never orders parts. Carrier Enterprise, Johnstone Supply, RE Michel, Munch's Supply, ABCO HVACR โ supplier catalog cross-reference every order. R-410A vs. R-454B mismatch is catastrophic; filter sizes and breaker amperage are inspection failures.
- Checkpoint four โ Warranty terms โ Manufacturer term sheet source-of-truth (Goodman, Trane, Carrier, Lennox, Bryant, Rheem, York, Daikin, Mitsubishi, LG, Fujitsu, American Standard, Amana). Shop's pre-built warranty matrix open in advisor's tablet; 5 seconds per warranty paragraph.
- Checkpoint five โ Financing language โ Binary. Wisetack/GreenSky/Synchrony portal output copy-pasted verbatim. FCRA adverse-action lender's signed template only. Two-party-consent per-state vendor configuration. AI never touches regulated APR, term, fee, or disclosure language.
- Role concentration โ CSR (names); Dispatcher (names + override log); Tech (parts + numbers); Comfort Advisor (all five every time); Service Manager (Rilla + RC&D); Marketing Manager (review responses + outbound consent); Owner (verify the verify).
- Owner threshold โ Final authority on $5K+ customer commitments, financing disputes, regulatory filings, recall escalations. Below $5K is operations; above $5K is owner judgment.
- 30-day discipline cliff โ Habit installs in 30 days or it doesn't. Day 1 read aloud, day 2-5 buddy shadow, day 6 20-min team review, day 7 institutional, day 8-30 muscle memory. After 30 days without the habit, fabrications have shipped, customer commitments have been honored, default has shifted to "AI handles it."
- The pass enables the lifts โ Avoca's HL Bowman 70% YoY growth, Rilla's 18% close-rate lift, Dispatch Pro's 12-18% yield lift, Hatch's 30-45% stale-lead reactivation โ all compound in shops with the verify pass running, all roll back in shops without it.
- The 60-point gap is a discipline gap โ 72% relevant, 12% embedded; the 60-point delta is verify discipline, not tooling.
Skill.re