Generative AI for Trades, in 8 Minutes
You will hear seven words at every AHR Expo booth, every ServiceTitan webinar, every Avoca demo, and every Rilla pitch from a trade-show sponsorship table for the rest of the decade โ token, prompt, model, temperature, context window, system prompt, fine-tune. Most of the people using those words at vendor demos do not entirely know what they mean. The technicians in your shop definitely do not. That is okay. Your job is not to write the AI; your job is to operate it. This lesson is the eight-minute version of the technical vocabulary every CSR, dispatcher, tech, advisor, and owner needs to use Titan Intelligence, Avoca, Jobber Copilot, Sera AI, or Housecall Pro AI Agents intelligently. Eight minutes to read, eight minutes to learn, eight minutes you will never have to spend again. Bookmark this one. Tape it to the wall above the shop coffee pot.
The Token: The Only Unit the AI Counts
A token is a piece of a word. Not a word, not a letter โ a piece. "Compressor" might be one token. "Refrigerant" might be two. "R-454B" might be three or four depending on the model. The AI does not see your sentences. It sees a stream of tokens. When the AI generates an answer, it produces one token at a time, each one chosen because it is the most likely token to follow the ones before. That is the entire machinery, dressed up.
Why does this matter on the truck? Because every AI tool you buy charges by tokens. Avoca's per-minute pricing includes tokens behind the scenes. ChatGPT and Claude API calls are billed per million tokens in and per million tokens out. ServiceTitan's Titan Intelligence has token-based limits inside its plan tier. The owner who copies and pastes a 30-page manufacturer technical bulletin into a chat to "ask the AI about it" is burning tokens the way the apprentice burns flux on a copper joint โ wastefully and possibly setting something on fire. The owner who pastes only the two relevant pages and asks one specific question burns a tenth of the tokens for a better answer.
A rough trades-shop rule: 750 words is about 1,000 tokens. A typical customer email is 100-200 tokens. A typical call summary is 200-400 tokens. A 10-page manufacturer manual is around 5,000-10,000 tokens. The largest 2026 models accept context windows of 100K-2M tokens, which sounds infinite โ but cost and quality both degrade as you stuff more into the window. Less is usually more. The CSR with a 50-token, focused question gets a better answer than the owner with a 50,000-token wall of text. This is the discipline that pays for itself: write tight, focus the question, paste only what matters.
The Prompt: What You Actually Write
A prompt is the instructions you give the AI. That is it. The marketing world has turned "prompt engineering" into a $200 course; the trades version is much simpler. A prompt is a set of words that tells the AI what role to play, what context to know, what task to do, what format to produce, and what limits to respect. Five things. We will spend two lessons on this in Level 2. The L1 version is: a prompt is your instruction.
The CSR's prompt at 8:32 a.m. might be: "Draft a rebuttal to a homeowner who says 'I'm just calling around for prices' on a no-cool call. The homeowner has a 14-year-old AC unit. Keep it under 60 words. Sound like a friendly receptionist at a Texas-based HVAC company." The dispatcher's prompt at 11:45 a.m. might be: "Marin Park called and cancelled their 2 p.m. AC install. Jose is the only senior tech available. We have two warranty recalls and one no-heat call queued for the afternoon. What is the highest-revenue redeployment, and what is the customer experience risk?" The Comfort Advisor's prompt at 6:50 p.m. might be: "The homeowner said 'I need to think about it' on a $18,400 furnace replacement. Their wife was supportive of the financing. They expressed concern about the warranty length. Give me three follow-up text drafts I can send within the next 90 minutes that don't sound desperate." Five things in every prompt: role, context, task, format, constraint. Repeat them and the AI gives you usable output. Skip them and the AI gives you generic output.
The Model: Which Brain You Rented
A model is the specific AI you are talking to. GPT-5, Claude Opus 4.6, Claude Sonnet 4.6, Gemini 2.5, Llama 4, Mistral Large โ each is a different model with different strengths, different speeds, different costs per million tokens, and different blind spots. You will not pick the model directly very often as a trades operator; ServiceTitan picks it for Titan Intelligence, Avoca picks it for their voice agent, Rilla picks it for their coaching tool. But you should know that the picks differ and the picks affect what you experience.
The two things that matter to you: speed and reliability on your specific task. Avoca picked the model that produces voice responses fastest without obvious hallucination on missed-call answering. ServiceTitan's Dispatch Pro picked a smaller, faster model that runs the every-10-minute re-evaluation without bogging down the dispatch board. Rilla picked a model strong on multi-speaker transcription and long-context summarization. None of these are "the best AI." They are the right tool for the apprentice job.
When you evaluate a new vendor in a 30-day pilot, you do not need to know the model name. You do need to know two things: how often it produces output your shop has to fix (the hallucination rate, in shop terms), and whether the speed matches the work pace. A 12-second AI response feels fast on an email; it feels glacial on a live phone call. The right model for the right apprentice job.
Temperature: How Creative Is the Apprentice?
Temperature is a number between 0 and 1 (sometimes higher) that controls how random the AI's output is. Low temperature (near 0) means the AI picks the most likely next token every time โ predictable, repetitive, accurate on factual recall. High temperature (near 1) means the AI samples from the top several likely tokens โ varied, more "creative," more prone to wandering.
For trades work, you almost always want low temperature on customer-facing artifacts (invoices, financing language, code citations, warranty terms) because predictability is the value. You can stomach higher temperature on marketing copy, blog posts, social media drafts, and creative angles in proposals because variety is the value. Most consumer AI tools (ChatGPT, Claude) default around 0.7-1.0, which is fine for brainstorming and bad for warranty disclosures. Vendor tools (Avoca, Rilla, Titan Intelligence) usually lock the temperature low because they know what kind of output the shop needs. When you do roll your own custom GPT or a Make/Zapier workflow that calls the API, you control the dial โ set it low for anything customer-facing.
Context Window: How Much the AI Can See at Once
Context window is the maximum amount of input the AI can hold in mind at once โ measured in tokens. GPT-5 is around 200K-400K tokens. Claude Opus 4.6 is around 200K-1M. Gemini 2.5 is up to 2M. That means in principle you can paste a whole manufacturer service manual, a full 12-month ServiceTitan export, and three weeks of CallRail transcripts into one prompt and ask a question.
In practice, you should not. Quality degrades. Cost rises. The AI "forgets" things in the middle of long contexts โ a documented phenomenon called the "lost in the middle" effect. The pragmatic trades discipline: keep prompts tight, paste only what is relevant, and use the AI's built-in tools (file uploads, retrieval, integrations) for large documents rather than copy-paste. Avoca, Rilla, Titan Intelligence, and Jobber Copilot all use retrieval under the hood โ they fetch the relevant chunk from your data, not the entire dataset, and feed only that to the model. You should mimic that discipline manually when you operate ChatGPT or Claude directly.
The System Prompt: The Permanent Instructions
A system prompt is the AI's standing orders โ the role, rules, and constraints that apply to every conversation until you change them. Avoca's system prompt for an HVAC shop's voice agent might say: "You are a friendly receptionist for [Shop Name]. You book service calls. You never quote prices for repair or replacement. You always offer the next available slot first. You never promise a tech by name. If asked about warranty, transfer to a human." The CSR's daily questions never see that system prompt; it sits underneath every conversation invisibly.
The reason this matters: when you buy Avoca, Rilla, ResponsiBid, or any other trades AI tool, you are essentially renting their system prompt plus their model. The system prompt is what they spent six months tuning to make the AI sound right for trades. The difference between Avoca and ChatGPT answering your missed call is not the underlying model โ it might be the same OpenAI or Anthropic backbone โ but Avoca's system prompt is shaped by thousands of hours of HVAC, plumbing, and electrical call data and a product team that knows what a CSR should and should not say. ChatGPT's system prompt is generic. That is what you are paying for. Knowing this lets you ask vendors the right question: "What does your system prompt do that I could not write myself?" The good ones have a real answer.
Fine-Tune vs. RAG vs. Prompt: The Three Ways AI Gets Customized
You will hear three customization methods in vendor pitches. They are very different in cost, quality, and risk.
Prompt engineering is the cheapest: change the system prompt, the AI behaves differently. Avoca tuning their voice agent for HVAC vs. plumbing is mostly prompt engineering. This is what 80% of "AI products" in 2026 actually are.
RAG (retrieval augmented generation) is the next tier: the AI is given access to a database โ your ServiceTitan history, your manufacturer manuals, your shop's call recordings โ and retrieves the relevant chunk before answering. CallRail Conversation Intelligence does RAG on call transcripts. Titan Intelligence does RAG on the shop's CRM data. This is where most serious 2026 trades AI lives. The model itself is not changed; the AI is given better context. Critically, RAG is also how the AI gets to know your specific shop without expensive retraining.
Fine-tuning is the heaviest: the underlying model is retrained on a domain-specific dataset (every HVAC service call from 50 shops, say) so the model itself thinks differently about trades work. It is expensive, slow, and only some vendors do it. When they do, they say so loudly. Most do not, because RAG plus prompt engineering covers 90% of what trades shops need at 5% of the cost.
When a vendor says "our AI is trained on the trades," the L1 graduate asks: trained how? Prompt engineering, RAG, or full fine-tune? The answer tells you what you are paying for.
Hallucination: The Failure Mode, Named
Hallucination is the technical term for when the AI confidently produces fabricated content. The 12-lb R-454B charge that should be 9.3 lbs. The Section 25C credit cap that is wrong by $3,000. The warranty year that does not match the manufacturer's actual term. The plumbing-code reference for IPC Section 802.6 that does not exist.
Hallucination is not a bug. It is the predictable consequence of statistical text generation. The AI is producing the most likely next words. When the right answer is statistically dense in the training data, the most-likely tokens land on truth. When the answer is rare or specific or the training data was thin, the most-likely tokens land on plausible-sounding fiction. The AI cannot tell which is which. You can. The 30-second verify is your hallucination filter.
Hallucination rate is roughly proportional to how rare and specific the question is. CSR rebuttal scripts have a low hallucination rate (the training data is dense). Section 25C IRS guidance for a specific 2026 AGI bracket has a high hallucination rate (the training data is sparse and changes year to year). Plumbing-code citations for a specific city's amended IPC have a very high hallucination rate (training data essentially does not include local amendments). Knowing where the hallucination risk lives helps you allocate verification effort.
Agents vs. Chatbots: The 2026 Distinction
The word "agent" gets thrown around a lot in 2026. Here is the distinction that matters. A chatbot answers questions. An agent takes actions. Avoca is an agent: it answers calls and books appointments in ServiceTitan. Hatch is an agent: it sends nurture texts and books warm leads. ServiceTitan Dispatch Pro is an agent: it reshuffles the board. ChatGPT is a chatbot: it produces text and stops. The agent has hands; the chatbot has only a mouth.
Agents are higher-leverage and higher-risk. Higher-leverage because they save you the human step of taking the AI's output and turning it into action. Higher-risk because an agent that hallucinates can book unrunnable calls, send wrong texts to customers, or reshuffle a dispatch board in ways that strand a crew. Every agent your shop buys in 2026 needs an action-level verify discipline โ not just "is the text good" but "should this action have happened." The 4 p.m. CSR review of Avoca-booked calls is exactly that action audit, applied daily.
The Vocabulary Test
The L1 graduate of this lesson can sit in a vendor demo, listen to "our proprietary AI is a fine-tuned model with retrieval augmentation across a 200K context window and a temperature-locked system prompt deployed as an agent for the HVAC vertical," and translate it in their head to: "they wrote a system prompt for our trade, plugged the AI into a database of our type of shop, kept the temperature low for predictability, and the tool takes actions in ServiceTitan rather than just suggesting them." That is the bar. You do not need to know how transformer attention works. You need to know what each piece means in operational terms so the demo cannot baffle you and so you can ask the questions that actually predict pilot success.
Every word in this lesson maps to a buying decision. Token count maps to monthly cost. Prompt quality maps to output quality. Model choice maps to speed and reliability. Temperature maps to risk. Context window maps to data-input strategy. System prompt maps to vendor moat. RAG vs. fine-tune vs. prompt maps to customization depth. Hallucination maps to verify discipline. Agent vs. chatbot maps to action-level risk. Nine concepts. Eight minutes. Every Monday morning standup, every Service World Expo demo, every Nexstar peer call where a fellow owner says "we're piloting AI" โ you now have the vocabulary to follow the conversation and to drive it.
What the CSR, the Dispatcher, the Tech Need to Know
Not everyone in the shop needs the full vocabulary. The owner needs all of it. The service manager needs all of it. The marketing manager needs prompt, model, temperature, hallucination, and agent. The Comfort Advisor needs prompt, hallucination, and agent. The CSR needs prompt and hallucination. The tech needs hallucination โ that's it, that's the one. If your tech understands hallucination ("the AI will sometimes confidently invent a part number โ verify on the manufacturer site before you order"), you have done the L1 job for that role.
The discipline of teaching only the words that matter to the role keeps L1 from feeling like a CS degree. Most of your team's vocabulary load in 2026 is one word: hallucination. Teach that one and the 30-second verify habit lands. Add prompt for the CSR, advisor, and marketer. Add model and temperature for the owner. That is the program in eight minutes, and it is sufficient for a shop to run AI competently in 2026 without needing a single engineer on staff.
Key Takeaways
- Token = the unit AI counts. 750 words โ 1,000 tokens. Every vendor charges by it. Less context, focused question, better answer.
- Prompt = your instruction. Five parts every time: role, context, task, format, constraint. Skip them and you get generic output.
- Model = which brain you rented. Vendors pick for you. Evaluate on speed and hallucination rate on your specific apprentice task, not on model name.
- Temperature = how random/creative. Low for customer-facing artifacts (invoices, financing, warranty, code). Higher for brainstorming and marketing copy.
- Context window = how much AI can see at once. Large in 2026 (up to 2M tokens) but quality degrades in long contexts. Paste only what's relevant.
- System prompt = the vendor's moat. Avoca's, Rilla's, Titan Intelligence's months of tuning live here. Ask any vendor: what does your system prompt do that we couldn't write ourselves?
- RAG > fine-tune > prompt only, for shop-specific customization. Most serious 2026 trades AI is RAG plus a tuned system prompt; full fine-tune is rare and announced loudly.
- Hallucination = predictable failure mode. Rate is roughly proportional to how rare/specific/recent the question is. Code citations, fresh tax credits, local amendments hallucinate worst.
- Agent vs. chatbot: agents take actions in your shop's systems (Avoca, Hatch, Dispatch Pro). Action-level verify discipline is non-negotiable on every agent.
- Vocabulary load by role: tech needs hallucination. CSR needs prompt + hallucination. Advisor adds agent. Marketer adds model + temperature. Owner needs all nine.
Skill.re