Spotting a Bad AI Answer in 10 Seconds
The five-part prompt makes the AI's draft better. Context discipline keeps the data flow clean. Now the third skill in this chapter: catching the bad answer before it reaches the customer. The AI hallucinates with the same friendly voice it uses when it is right โ confident SEER ratings that do not exist, R-454B refrigerant charges off by 30%, Wisetack APRs from 2024 quoted as current, NEC code sections that were never written, Trane warranty terms that ended two product cycles ago. The journeyman who cannot spot the bad answer in 10 seconds will sign off on hallucinations until the kitchen-table moment when the homeowner waves the proposal and the SEER number is wrong. This lesson is the smell-test discipline that turns a 30-minute verify into a 10-second skim. The CSR's 5-second pass on Avoca bookings. The advisor's 30-second verify on the kitchen-table proposal. The tech's gut check on the spec sheet. None of it is intuitive. All of it is trainable. Build it in week one of any AI deployment and it becomes muscle memory by month two.
The Five Smell Tests
Every AI hallucination in a trades shop comes from one of five categories. Memorize the five; the smell test for each is ten seconds or less. The categories are: hallucinated part numbers, fake efficiency ratings (SEER, SEER2, AFUE, HSPF), invented financing terms (APR, term length, promo window), wrong warranty language (year count, coverage scope), and fabricated code citations (NEC, IRC, IMC, IECC, EPA 608). Every one of these has a source-of-truth a journeyman can check in under 10 seconds with the right reflex. The reflexes are what we are building.
The reason five and not seven: these are the five categories where the AI's failure mode is structural, not occasional. The training corpus is thin or stale on each. Part numbers change quarterly at the manufacturer level. SEER ratings on specific equipment models drift as manufacturers release SKU variants. Financing APRs and promo windows turn over every quarter at Wisetack, GreenSky, and Synchrony. Warranty terms revise periodically at Goodman, Trane, Carrier, Lennox, Bryant, Rheem, and York. Code sections vary by AHJ amendment and version year. The AI does not know which version it is operating against, and it does not tell you when it does not know. The five-category checklist is the operator's compensating discipline.
Hallucinated Part Numbers
The AI will produce part numbers that look real, have the right character format for the manufacturer's catalog, and do not exist in any supplier's system. A confident "Goodman CAPF3636B6" that is not in Carrier Enterprise's catalog. A "Trane 4TWR4036G1000A" that is two product cycles old and discontinued. A "Rheem RA1424AJ1NA" that should be RA1430AJ1NA. The number looks right; the number is not right. The tech orders the part; the supplier flags it as unknown; the tech rebooks the call; the customer's heat does not come back on that day.
The 10-second smell test for parts: did the AI provide a part number for any line item in the recommendation, the proposal, or the order? If yes, the tech opens the supplier's catalog or the FSM platform's parts integration and pastes the number. Match in 5 seconds, mismatch in 5 seconds. Mismatch means the AI guessed; the actual part number comes from the supplier site or the manufacturer's serial-number lookup. Never accept an AI-provided part number on its face. Cross-reference is non-negotiable for every order.
Fake Efficiency Ratings (SEER, SEER2, AFUE, HSPF)
The AI loves SEER ratings. Confident "16 SEER2 with 9.5 HSPF2" on a Carrier system whose actual current SKU is 15.2 SEER2 with 8.1 HSPF2. The hallucination is most common on the 2023-2024 SEER2 transition โ training data is mixed between pre-transition SEER ratings and post-transition SEER2 ratings, and the AI averages across both producing numbers that look right and are not. The Comfort Advisor reads the proposal at the kitchen table; the homeowner Googles "Carrier 25HCC524A SEER" while sitting on the couch; the actual rating is one number, the proposal is another, the close collapses.
The 10-second smell test for ratings: the proposal has a number; the advisor opens the manufacturer's actual model-specific spec sheet (Carrier, Trane, Goodman, Lennox, Bryant, Rheem, York, Mitsubishi, Daikin all publish them) and verifies the number. SEER2 vs. SEER matters; HSPF2 vs. HSPF matters; AFUE on a furnace vs. EER on a condenser matters. The advisor's discipline is to never paste an AI-generated rating into a proposal without the spec-sheet cross-check. The spec sheet is on the manufacturer's site or in the shop's pricebook integration; 10 seconds, every time.
Invented Financing Terms (APR, Term Length, Promo Window)
The AI knows the names Wisetack, GreenSky, and Synchrony. It does not know the current rates. It will produce "Wisetack 0% for 18 months" when the actual 2026 Wisetack promo is 0% for 12 months on $1,000+ tickets, or "GreenSky 84 months at 7.99%" when the actual current GreenSky band for that credit tier is 84 months at 9.49%. The Comfort Advisor reads the AI-drafted talk-track; the homeowner asks the specific monthly payment; the math comes back different from what the portal calculates; the close trust collapses; worse, the financing language now sits in the proposal as a Reg Z misstatement.
The 10-second smell test for financing: any number that names a lender, an APR, a term length, or a promo window must come from the lender's portal verbatim, never from the AI's draft. The advisor's discipline is to draft the surrounding talk-track with the AI but copy-paste the actual rates, terms, and disclosure language from the Wisetack, GreenSky, or Synchrony portal output. The portal is the source of truth. The AI surrounds the portal output with the explanation. The portal output never gets paraphrased or restated by the AI. The 10-second check is a glance at the portal screenshot or printed approval-tier output that sits next to the proposal.
Wrong Warranty Language
The AI's memory of warranty terms is structurally stale. Manufacturers revise warranty terms periodically โ extended coverage promotions, reduced coverage on certain SKUs, registration-required vs. registration-not-required variations โ and the training corpus may sit a year or more behind. The AI confidently writes "12-year compressor coverage on the Trane XR16" when Trane's current term sheet shows 10 years on that SKU. The Comfort Advisor's proposal includes the AI-drafted warranty paragraph; the proposal is signed; year 11 the compressor fails; the homeowner expects coverage; the shop now owes the cost out of pocket because the proposal said 12 years.
The 10-second smell test for warranty: any warranty year count, any coverage scope description, any registration requirement language goes against the manufacturer's current term sheet. Goodman, Trane, Carrier, Lennox, Bryant, Rheem, and York all publish current term sheets โ the shop's pricebook integration or the manufacturer's dealer portal carries them. The advisor's discipline is to paste warranty paragraphs verbatim from the term sheet, or from the shop's pre-built warranty matrix that the service manager maintains and updates monthly. The AI never originates warranty language; the AI surrounds verified warranty language with explanation.
Fabricated Code Citations (NEC, IRC, IMC, IECC, EPA 608)
The most dangerous category. The AI will cite "NEC 230.71" or "IRC R310.2" or "EPA 608 Section 82.156" with confidence. Sometimes the section exists. Sometimes the section number is right but the content is wrong because the AHJ amended it. Sometimes the section was renumbered between code-cycle editions. Sometimes the citation is entirely fabricated. The tech reads the AI's recommendation, accepts the code citation, executes the work; the permit inspection fails; the rework costs the shop $1,800 and a Yelp review; the licensed installer's name is on the permit, and the contractor's board now has a complaint open.
The 10-second smell test for code citations: any code citation in an AI output is suspect by default. The licensed tech, advisor, or master electrician pulls the actual code book or the AHJ-published amended edition and verifies the section, the year, and the local amendment. This is the one category where the smell test cannot compress to under 30 seconds โ the verification requires a real source check. The compensating discipline is to instruct the AI in the Constraint field of every prompt to not cite code sections; the licensed pro adds citations manually when needed. "Do not cite NEC, IRC, IMC, IECC, or EPA sections โ the licensed tech will add citations from the AHJ" is the most important constraint in any code-adjacent prompt.
The CSR's 5-Second Skim
The CSR is the highest-volume AI verifier in the shop. Every Avoca booking, every AI-drafted call summary, every CSR-AI rebuttal preview crosses her screen. She cannot run a 30-second verify on every one โ she would never take a call. She runs a 5-second skim. The 5-second skim has four checkpoints, in order: name, address, slot, forbidden-promise scan. If the booking passes all four in 5 seconds, it accepts. If any one fails, she pulls it for the full verify or routes to a colleague.
Name: the AI heard the caller's name correctly. Avoca's most common error is "Smith on Oak Lane" becoming "Smyth on Oak Avenue" โ phonetically close, operationally wrong. The CSR catches it at the booking confirmation; if it slips, dispatch routes the tech to the wrong street and two slots are lost.
Address: the AI captured the full street address, the city, and the ZIP correctly, and the ZIP is actually in the shop's service area. Avoca-confidence on ZIPs is the failure pattern โ the AI says "yes we service 99352" when 99352 is in a different state. The CSR's catch is whether the ZIP matches the service area map; if not, the booking is unrunnable and the CSR re-engages the customer with a real explanation.
Slot: the AI booked a slot the dispatch board can actually staff. Avoca pulls from ServiceTitan API slot availability, but the API can show ghost availability if the platform's dispatch logic and the AI's slot view fall out of sync. The CSR's catch is a 2-second glance at the dispatch board to confirm a tech is actually available for that window.
Forbidden-promise scan: the AI did not promise something the shop cannot deliver. "Free estimate" when the shop is $79 diagnostic. "Same day" when the territory is booked. "Lifetime warranty" when the actual policy is one year. "We service that area" when the ZIP is outside. The forbidden-promise list lives in the shop's prompt library and the CSR's verify cheat sheet; the 5-second scan is whether the booking confirmation contains any of those exact phrases.
The 5-second skim runs on every AI booking, in real time. The 4 p.m. daily review batches them for the service manager's audit โ false negatives caught at 4 p.m. become bulletin-board entries and constraint updates for the next day's templates. The discipline scales because the per-call cost is 5 seconds.
The Advisor's 30-Second Verify on the Kitchen-Table Proposal
The Comfort Advisor's verify surface is the highest-stakes in the shop. Every AI-drafted proposal at $8K to $45K residential, $25K to $250K commercial replacement goes through the full 30-second verify before the advisor walks into the kitchen. The five checkpoints from L1 (numbers, names, parts, warranty, financing/regulatory) apply with high precision; the L2 smell-test discipline overlays specific reflexes on the high-frequency categories.
SEER and AFUE on every equipment line: open the spec sheet for the specific SKU on the manufacturer's site or in the pricebook, verify the rating. Any rating that "feels round" (16.0, 18.0, 20.0) deserves an extra look โ the AI rounds, the actual SKU rarely does (15.2, 16.6, 17.8 are more common in 2026 SEER2 ratings). Refrigerant charge weight: open the spec sheet, verify the R-454B charge against the manufacturer's published weight; 2026 hallucination rates are highest here because the corpus is thin. Financing payment math: open the Wisetack, GreenSky, or Synchrony portal that produced the approval; verify the APR, term, and monthly payment match the portal output verbatim. Warranty paragraph: open the manufacturer's term sheet; verify year count, coverage scope, and registration requirement. Rebate amount: open the current state-utility program page (Mass Save, Energy Trust of Oregon, NV Energy, ConEd, etc.); verify the rebate dollar amount and program window.
Thirty seconds for five checks is six seconds each. The discipline is real-time. The advisor verifies at the truck before walking into the kitchen, never at the kitchen table โ verify at the table breaks the close. If any check fails, the advisor pulls the AI-drafted line and replaces it with the verified source-of-truth output. The proposal that enters the kitchen has zero hallucinations and a defensible answer to every number on every page.
The Tech's Gut Check โ "Does This Match the Spec Sheet?"
The tech's verify discipline is shorter and more reflexive than the advisor's because the tech is moving faster and the stakes per artifact are lower. The tech's primary AI outputs are voice-notes structures (from L1) and repair-vs-replace prompts (the talk-track read off the tablet). The gut check the tech runs in 5-10 seconds is "does this match the spec sheet?" โ applied to every number the AI surfaces.
Equipment age: the AI inferred "14-year-old install" from voice notes; the FSM platform's customer record says 11 years; the tech catches the mismatch and the recommendation logic shifts (11-year systems carry different rebate eligibility and warranty math than 14-year systems). Repair cost: the AI's repair estimate includes a line item the tech did not authorize; the tech opens the shop's flat-rate book and verifies the line item exists; if not, it gets cut. Replacement options: the AI surfaced a 17 SEER2 dual-fuel option; the tech opens the pricebook and verifies the option is in stock and currently priced โ Q3 2026 supply-chain shifts on R-454B equipment have made several SKUs unavailable in certain regions. Financing presentation: the AI's "good/better/best" payment options reference financing terms the tech reads to the customer; the tech runs the 10-second smell test on the APR and term against the morning's Wisetack soft-pull output.
The tech's catches also feed the prompt library โ hallucinated SEER, out-of-date pricebook reference, fabricated NEC section all go on the bulletin board and update the constraint or context for next time.
The Service Manager's Pattern Detection
The CSR catches in real time. The advisor catches at the truck. The tech catches at the tablet. The service manager catches the patterns across all of them. Every Friday the service manager runs a 30-minute review of the week's bulletin-board catches: how many were hallucinated SEER numbers, how many were wrong warranty terms, how many were forbidden-promise hits, how many were code citations the licensed tech caught at the AHJ check. The pattern tells the service manager which prompt templates need a constraint update.
If SEER hallucinations cluster on Trane SKUs, the Trane equipment-tag context block needs more specificity. If financing hallucinations cluster on the Q3 promo window, the Wisetack constraint needs "use only the rates from the portal approval pasted in context." If code citation catches all come from the rookie tech, his template needs the "do not cite code sections" constraint and a 1-on-1 on the discipline. Pattern detection makes verify an improving system, not a static checklist.
The Bulletin Board as Feedback Loop
The L1 "AI Caught a Hallucination" bulletin board is not just a visible artifact โ it is the shop's most important feedback loop on AI quality. Every catch goes on the board with date, role, artifact type, source-of-truth used, specific fabrication, remediation, and cost avoided. By month three, the board carries 30-80 entries depending on shop size, and the patterns are obvious. The service manager mines the board for the prompt-library updates that compound the discipline. The owner reviews the board quarterly for the AI vendor renewal decisions.
The board also reframes mistakes as wins. The CSR who catches a wrong-ZIP Avoca booking is "the CSR who caught one," not "the CSR who almost let a bad call through." The advisor who catches a hallucinated SEER at the truck is "the advisor who saved the close." The tech who catches a fabricated NEC citation is "the tech who saved the permit." Verify discipline lives in voluntary catches, not enforced rules. The board makes the catch the recognized win.
When to Trust the AI Without Full Verifying
Not every AI artifact gets the 30-second verify; not every artifact needs it. The discipline is calibrated by stakes. An internal Friday recap that only the owner reads: 5-second skim for the wrong number, but no full verify required โ the owner is the verify. A CSR rebuttal that goes on the cheat sheet and gets read 50 times before being updated: full 30-second verify when first written, then 5-second skim each use. A proposal at the kitchen table: full 30-second verify every time, no exceptions, because the customer is the auditor. An internal call summary on the customer record: 10-second skim because future operators will act on it, but full verify is not required if the next role-holder also runs the smell tests.
The shop's verify policy lives in the prompt library next to each template โ the policy is "5-second skim for internal-only," "10-second smell test for in-shop handoff," "30-second full verify for customer-facing." The discipline calibrates effort to stakes. The verify is non-negotiable on customer-facing artifacts; on internal artifacts the calibration prevents the verify from becoming a bottleneck. Shops that overcalibrate (30-second verify on every artifact) burn out the team in 6 weeks; shops that undercalibrate (5-second skim on the kitchen-table proposal) get caught by a hallucination at month three. The middle path is calibrated by artifact stakes.
The Three AI Tools That Most Commonly Fail Each Smell Test
By 2026 the trades industry has accumulated enough data to know which tools fail which smell tests most often. ChatGPT, Claude, and Gemini โ the public LLMs โ fail every smell test occasionally because they are trained on broad corpora with no trades-specific verification layer. The shop's discipline on these tools is full verify on every customer-facing output. Avoca, Rilla, and Titan Intelligence โ the trades-tuned tools โ fail less often but still fail; their integrations with FSM data reduce some categories (parts and customer history) but do not eliminate the others (financing, warranty, code). The shop's discipline on these is calibrated verify based on which categories the tool's tuning covers.
By 2026: ChatGPT and Claude fail most on financing (training stale on promo windows), code (national-default sometimes wrong for AHJ), and R-454B charge weights (corpus thin). Avoca fails most on names and addresses (speech-to-text errors). Rilla fails most on subtle tone scoring (coaching commentary occasionally invents moments). Titan Intelligence fails least because FSM data grounds outputs, but still hallucinates financing when the prompt does not reference the portal. Verify discipline calibrates per tool.
Building the 10-Second Reflex
The 10-second smell test is a reflex, not a checklist. Reflexes are built by repetition under guidance. Week one of any AI deployment, every role's verify pass is observed by a buddy โ the CSR's 5-second skim is shadowed by the service manager, the advisor's 30-second verify is shadowed by the owner, the tech's gut check is shadowed by the lead tech. The buddy is not auditing โ they are training the reflex. By day five, the reflex is the role-holder's; by day seven, the buddy steps back to weekly random sampling.
Week two through four: the reflex runs in the daily rhythm. The CSR catches a wrong ZIP at 9:14 a.m. and the booking gets reissued; the catch goes on the board. The advisor catches a hallucinated SEER at the truck at 6:25 p.m. and the proposal gets edited; the catch goes on the board. The tech catches a fabricated NEC citation at 2:40 p.m. and the prompt's constraint gets updated; the catch goes on the board. By end of month one, the board has 15-30 entries and the patterns inform next month's template updates. By end of month three, the reflex is institutional and new hires onboard against the catches already documented.
Shops that install the reflex at month six fail at the same rate as shops installing the 30-second verify at month six. The window is the first 30 days. After that, the team has internalized "the AI is probably right" and re-installing requires undoing the default. The 30-day discipline cliff: this lesson belongs to week one of every AI tool deployment.
Key Takeaways
- Five smell tests cover every common AI hallucination in a trades shop: hallucinated part numbers, fake efficiency ratings (SEER/SEER2/AFUE/HSPF), invented financing terms (APR/term/promo), wrong warranty language (year count/coverage/registration), fabricated code citations (NEC/IRC/IMC/IECC/EPA 608).
- Each smell test has a 10-second source-of-truth cross-check: supplier catalog for parts, manufacturer spec sheet for ratings, lender portal for financing, manufacturer term sheet for warranty, code book or AHJ for code citations.
- The CSR's 5-second skim has four checkpoints: name, address, slot, forbidden-promise scan. Real-time, every AI booking. 4 p.m. daily review batches the false negatives.
- The advisor's 30-second verify happens at the truck, not at the kitchen table. Verify at the table breaks the close. Five checkpoints, six seconds each, against five sources of truth (spec sheet, portal, term sheet, utility page, AHJ).
- The tech's gut check is "does this match the spec sheet?" Equipment age vs. FSM record, repair line items vs. flat-rate book, replacement options vs. pricebook, financing presentation vs. morning's portal output. 5-10 seconds, reflexive.
- The service manager runs pattern detection on Friday. 30-minute review of the bulletin-board catches; cluster patterns drive prompt template updates; this is what makes the verify discipline an improving system rather than a static checklist.
- Calibrate verify effort to artifact stakes. 5-second skim for internal-only, 10-second smell test for in-shop handoff, 30-second full verify for customer-facing. Overcalibration burns out the team; undercalibration catches a hallucination at month three. The reflex is built in week one and becomes institutional by month three.
Skill.re