Few-Shot Prompting with Operations Examples
Overview
You ask an AI to write a new SOP section for vendor communication. It comes back generic. Bureaucratic. Not how your team actually works. You ask again with more specific instructions. Still doesn't hit the mark. Then you try something different: you show the AI three examples of vendor communication emails your team has actually written, real, strong examples from your operations. You ask it to write a new one in that style. The result is immediately better. It captured the tone, the level of detail, the professionalism, and the practical approach of your actual team work.
That's few-shot prompting at work.
Few-shot prompting means giving the AI examples of what good looks like and asking it to follow that pattern. Instead of describing quality in words, you show it. Most people know this works. They just don't know how much it transforms operations output, or how to use it systematically.
Why Examples Beat Instructions in Operations Work
Here's the fundamental problem with most AI prompts: they assume the AI knows what you mean by "good" and "professional" and "thorough." It doesn't. Those words mean different things to different people, organizations, and contexts. Professional for a Fortune 500 compliance department looks different from professional for a startup. Thorough for financial auditing looks different from thorough for vendor risk assessment.
When you show the AI examples, you eliminate ambiguity. You're not saying "write a professional vendor review." You're showing what a professional vendor review looks like in your environment, with your standards, at your detail level, using your vocabulary and your process expectations.
Consider what happens with a typical prompt:
> "Write a vendor compliance review. Check their certifications and identify risks."
The AI will write something. It might be thorough or surface-level. It might match your quality bar or miss the mark. You won't know until you see it. You'll probably get something between 600-1200 words that covers the basics but doesn't match how your team thinks. You'll spend 30-45 minutes editing it to match your standard.
Now consider a few-shot prompt:
> "Write a vendor compliance review following this format and depth:
> EXAMPLE: [Your previous vendor review, the one your team agreed was solid, complete, and representative of your standard. 800-1200 words, covers the specific areas you care about, reaches the conclusions your team would reach.]
> Write a similar review for [New vendor], using the same structure, depth, and decision-making criteria."
The AI has a concrete example to match. It will write something much closer to what you actually need. Instead of 30-45 minutes of editing, you'll need 10-15 minutes. Instead of rewriting sections, you'll be adjusting tone and adding one or two specific details.
In operations, where consistency matters and standards need to be met repeatedly, this difference is massive. Few-shot prompting turns AI from a brainstorming tool into a work-production tool. The difference between "AI draft that needs significant rewriting" and "AI draft that needs minor polish" is the difference between a technique that saves time and a technique that saves time plus improves quality.
The Power of 2-3 Examples: Why More Isn't Always Better
Here's the counterintuitive part: you don't need many examples. One or two good examples are usually enough. Three is often optimal. Five or six starts hitting diminishing returns.
Why? Because the AI's job is to extract the pattern, not memorize specific examples. Too many examples can actually confuse it. Give it the clearest, most representative examples you have. One strong SOP section is better than five average ones.
For operations work, here's the rule of thumb:
- 1 example: Works for simple, straightforward tasks (e.g., "Write a vendor evaluation summary like this one")
- 2-3 examples: Best for most operational work (SOPs, process docs, compliance reviews). Shows variation while staying consistent.
- 4+ examples: Use only if the task is complex or has high variability (e.g., you want the AI to handle different types of vendor relationships, and each type needs a different approach)
Quality matters more than quantity. A poorly written example will pull the AI toward poor output. Choose examples you're actually proud of, work that represents your team at its best.
Few-Shot Prompting in Action: Three Operational Use Cases
Use Case 1: SOP and Process Documentation
You want the AI to draft a new SOP for a process. Instead of describing what you want, show it what you want.
Without examples (weak):
> "Write an SOP for the vendor onboarding process. Include all necessary steps and compliance checks."
The result might be generic, missing your specific compliance requirements, or structured differently than your existing SOPs.
With examples (strong):
> "Write an SOP for the vendor offboarding process. Use this format and depth:
> [PASTE IN YOUR BEST EXISTING SOP - the vendor onboarding SOP you wrote last year. 800-1200 words. Shows what "complete" and "well-structured" means in your environment.]
> The offboarding SOP should include: final payment processes, data deletion requirements, relationship wind-down, document archival, and compliance closeout. Use the same structure and level of detail as the onboarding SOP above."
Now the AI understands your tone, your level of detail, your compliance requirements, and your structural preferences. The output will be dramatically better, and you'll recognize it as something written by your team.
Tip: When using SOPs as examples, include the actual SOP you use, not a cleaned-up version you wish you had written. Real, imperfect examples that your team actually uses will produce output that works in your real environment. Aspirational examples often produce unrealistic results.
Use Case 2: Vendor and Supplier Analysis
You're evaluating a new vendor. Instead of asking generic questions, show the AI how you analyze vendors.
Without examples:
> "Evaluate this vendor on cost, quality, and reliability. Create a summary."
With examples (strong):
> "Evaluate this vendor using our standard evaluation framework:
> EXAMPLE 1: [Paste your evaluation of Vendor A from six months ago. The one your team agreed was thorough and useful. Include the structure you used, the questions you asked, the risk assessment, and your recommendation.]
> EXAMPLE 2: [Paste your evaluation of Vendor B from three months ago. Different vendor type. Different conclusion. Shows how you structure analysis for both "good vendor" and "risky vendor" scenarios.]
> Using the same structure, depth, and evaluation framework, analyze [New Vendor]. Include cost analysis, quality assessment, compliance verification, risk flags, and your recommendation."
With two vendor analyses to learn from, the AI will structure analysis the same way, ask similar questions, and apply your risk thresholds. You're not getting generic vendor analysis. You're getting your vendor analysis, extended to a new vendor.
Use Case 3: Risk Assessment and Compliance Documentation
You need to assess risk for a process change or new system. Show the AI how you do it.
With examples (strong):
> "Assess the risks of this proposed change [DESCRIBE CHANGE] using our risk assessment framework:
> EXAMPLE: [Paste your risk assessment from a previous process change. Show how you categorize risks (operational, compliance, financial, etc.), how you assign probability and impact ratings, how you identify mitigations, and how you structure the final recommendation.]
> Structure your assessment the same way. Identify all operational, compliance, financial, and strategic risks. Rate probability and impact. Propose mitigations. Give a recommendation on whether we should proceed."
Now the AI won't miss categories of risk you typically consider. It won't minimize concerns you'd flag. It will think like your risk team.
Building Your Few-Shot Example Library
Instead of hunting for examples every time, build a library of your best work.
What makes a good few-shot example?
- It's recent. Examples from last month beat examples from two years ago. Standards evolve, compliance changes, your team improves. Use current examples.
- It's actually good work. Not perfect, not aspirational. Work your team agreed was solid, useful, and met the standard. Something a team member would point to and say "yeah, that's how we do it."
- It's complete. Don't use partial examples. The AI learns from the whole thing, structure, detail level, tone, conclusions. Partial examples are confusing.
- It's representative of your standard. Not your best day. Not your worst. Representative of what you actually produce.
- It's varied (if using multiple examples). If you're giving two vendor analyses, give one of a vendor you approved and one you rejected. If you're giving two SOPs, give one for a simple process and one for a complex one. Variety shows the AI how to handle different situations.
Building your library:
Start with the work you do most often. For most operations teams, that's:
- Standard SOPs or runbooks (pick 2-3 good ones)
- Vendor evaluations (pick 2 of different types: one approved, one rejected)
- Process change assessments or risk analyses (pick 2: one low-risk, one high-risk)
- Compliance or audit responses (pick 1-2: typical format)
- Status reports or executive summaries (pick 1: your standard format)
Store these in a document your team can access. Label them clearly: "EXAMPLE: Vendor Evaluation (Approved Vendor)" so when you need them, you know what you're grabbing.
Structuring Few-Shot Prompts for Maximum Impact
The way you present examples matters. Here's the structure that works best:
Structure 1: Simple Pattern Replication
Use this for straightforward tasks where you have one clear example.
> "I'm going to show you an example of [TASK] that represents our standard. Study the structure, tone, and level of detail. Then do the same thing for [NEW TASK].
> EXAMPLE: [Paste example]
> Now, [SPECIFIC INSTRUCTION]."
Example: "Write a vendor risk assessment for [Vendor Name] following the same structure and depth as the example above."
Structure 2: Comparative Examples (Show Variation)
Use this when you want to show the AI different scenarios or outcomes.
> "Here are two examples of [TASK] showing different scenarios:
> EXAMPLE 1 (SCENARIO A): [Example of situation/vendor/process that led to this outcome]
> EXAMPLE 2 (SCENARIO B): [Different example showing different outcome or approach]
> Notice how [WHAT TO OBSERVE]. Now apply the same approach to [NEW SCENARIO]."
Example: "Here are two vendor evaluations, one where we approved the vendor and one where we rejected them. Notice how the risk assessment changes the recommendation. Now evaluate [New Vendor]."
Structure 3: Template + Example
Use this when you have both a formal template and a completed example.
> "Use this template structure:
> [TEMPLATE WITH HEADINGS]
> Here's how we filled it out for a similar [TASK]:
> [COMPLETED EXAMPLE FOLLOWING THE TEMPLATE]
> Now complete the same template for [NEW TASK]."
Try This Now: Create Your First Few-Shot Prompt
Pick a task you do regularly. Give yourself 20 minutes.
Step 1: Identify the task. What's something your team does repeatedly? (Vendor evaluation, SOP section, process assessment, compliance review, change request, incident summary?)
Step 2: Find your best example. What's the most recent, solid work you did for this task? Find it. This is your example. If you don't have a recent example, do one manually now, and use that as your example for future AI prompts.
Step 3: Anonymize if needed. Does it contain sensitive details? Replace vendor names with "Vendor A," financial figures with "[AMOUNT]," customer names with "[CUSTOMER]." Keep the structure and substance intact. You're protecting confidentiality while keeping the model for the AI.
Step 4: Write the few-shot prompt. Use this template:
> "I need you to [TASK] for [NEW SITUATION]. Here's an example of how we approach this:
> EXAMPLE: [Paste your example]
> Apply the same structure, depth, and approach to [NEW SITUATION]. Pay special attention to [SPECIFIC ELEMENTS that matter to your organization]."
Step 5: Test it. Give it to your AI. Compare the output to your example. Ask yourself: Does it capture the tone and structure? Does it hit your quality bar? Does it make the same judgment calls you would make? If yes, you've got a reusable few-shot prompt. If no, refine the example or the instructions and test again. Sometimes the example needs to be more recent or represent a clearer standard. Sometimes the instruction needs to be more specific about what matters.
Step 6: Save it. Once the prompt works, save it in a document your team can access. Label it clearly: "EXAMPLE: Vendor Evaluation using Few-Shot Prompting" or "TEMPLATE: Process Improvement Plan Analysis." Make it easy for others on your team to use.
Important: If your examples contain sensitive information (specific vendor names, financial details, customer names), you have two options: (1) Use real examples but strip out sensitive details while keeping the structure and approach, or (2) Use anonymized examples where you've replaced actual vendor names with "Vendor A" and actual contract amounts with "[AMOUNT]". Few-shot prompting works with anonymized examples. The pattern is what matters, not the specific data.
Advanced: Chaining Examples (Multi-Step Tasks)
Sometimes you need the AI to do something complex with multiple steps. You can show examples of each step. This is where few-shot prompting becomes really powerful. You're not just teaching the AI a format, you're teaching it a decision-making framework.
Example: Creating a process improvement plan
> "I need you to analyze a process and create an improvement plan in three steps:
> STEP 1 - Current state assessment:
> Here's how we documented the current state for [Process A]:
[Example of current state assessment, include what you measured, what metrics you used, how you described bottlenecks, and how you identified root causes]
> STEP 2 - Gap and opportunity analysis:
> Here's how we analyzed gaps in [Process A]:
[Example of gap analysis, show how you identified the gaps between current state and target state, how you prioritized gaps by business impact, and how you framed opportunities]
> STEP 3 - Improvement recommendations:
> Here's how we structured recommendations for [Process A]:
[Example of recommendations, show how you tied recommendations back to gaps, how you estimated impact, and how you framed implementation barriers and mitigation]
> Now apply this three-step approach to [New Process]. Use the same rigor in current state description, the same prioritization logic in gap analysis, and the same recommendation structure with estimated impact."
By showing an example of each step, you're teaching the AI not just the overall structure, but the thinking process at each stage. You're showing how your organization analyzes problems: what data you gather, what you measure, what you prioritize, how you think about trade-offs. The AI learns not just your format but your reasoning framework.
This is especially valuable for complex operational decisions where the process is as important as the output. You want the AI to think like your operations team, not just produce outputs in your template.
Failure Modes: When Few-Shot Prompting Backfires
Few-shot prompting is powerful, but it can fail in specific ways. Understanding these helps you avoid them.
Failure Mode 1: Using an outdated example
You used a vendor evaluation from 2023. Your evaluation criteria have changed. Your compliance requirements have tightened. Your budget has shifted. The AI learns from the old standard, not your new standard. The result meets the example's bar, which is no longer your bar. The fix: Update your examples quarterly. If your team makes significant process changes, update immediately. If you adopted new compliance requirements, update your examples to reflect the new standard. Your examples should always reflect current practice.
Failure Mode 2: Using an aspirational example instead of a real one
You have the "ideal" vendor evaluation, the one you wish you always produced. It's perfect. But your team doesn't actually work that way. When you use it as an example, the AI produces idealized output that your team can't follow up on. The output is better than your actual standard, which creates frustration when people try to implement it. The fix: Use real work that your team produced and agreed was solid. Not the best day. Representative of your actual standard. That's what teaches the AI to produce work your team can actually use.
Failure Mode 3: Using an incomplete example
You show the AI half an SOP. It doesn't understand what you're showing. The structure is unclear. The thinking is incomplete. The output is confused because the example was confusing. The fix: Use complete, finished work. The AI learns from the entire artifact, structure, detail level, tone, conclusions. Partial examples muddy the water.
Failure Mode 4: Mixing inconsistent examples when using multiple
You give the AI two vendor evaluations with completely different formats. One is structured as a matrix. One is narrative. One concludes with a recommendation. One just states facts. The AI doesn't know which pattern to follow. It produces inconsistent output because your examples were inconsistent. The fix: Keep your examples consistent in structure even if they reach different conclusions. Show variation in outcome (approved vs. rejected vendor), not variation in method.
Failure Mode 5: Forgetting to verify the AI's output
You ran a few-shot prompt and got output that matched your example perfectly. You assume it's correct and send it to leadership. But the AI misread a key detail or made a different judgment call than you would have. Few-shot prompting gets you most of the way, but it's not "set it and forget it." The fix: Read the output. Spot-check it against your judgment. If the AI made a different interpretation than you would have, decide if it's actually better or if you need to adjust. Few-shot prompting is acceleration, not automation.
Example completed few-shot prompt (vendor evaluation):
> "I need you to evaluate a new logistics vendor for our operations. Here's an example of how we structure vendor evaluations:
> EXAMPLE: [Paste your previous evaluation of Current Logistics Vendor, anonymized: 1200 words, includes cost analysis, service level assessment, compliance verification, risk identification, and recommendation]
> Using the same structure and level of detail, evaluate [New Vendor Name]. Include cost comparison, service capabilities, compliance status, operational risks, and your recommendation."
What to Do Monday Morning
- Find 2-3 pieces of your recent work. SOP you wrote. Vendor evaluation. Process assessment. Incident summary. Work you're proud of and that your team agrees represents your standard. That's your example library starter.
- Create one few-shot prompt using an example of your best work and a current task that needs doing. Follow the five-step process from the "Try This Now" section.
- Compare the AI output to your original example. Does it match the structure? The tone? The depth? Does it meet your quality bar? If not, which aspect diverged, was it the example or the instruction?
- Refine and test again. If the output needed adjustment, refine either the example (maybe it's not recent enough) or the instructions (maybe you weren't specific enough about what matters). Test again.
- Build your example library. By next month, have 3-4 solid examples ready. Vendor eval. SOP. Risk assessment. Process doc. Incident analysis. Label them clearly with dates and what makes them good so you grab the right one when you need it.
- Share with your team. Show one prompt and output to your team. Explain why the example-based approach worked better. Get their feedback on the quality. Encourage them to try it.
Key Takeaways
- Show, don't tell. Examples teach the AI what "good" looks like better than descriptions ever will. One strong example beats five pages of instructions. The AI learns by pattern-matching to your example, not by interpreting abstract principles.
- Use 1-3 examples for best results. For most operational tasks, two good examples are optimal. One solid example works for simpler tasks. More than three adds noise and dilutes the pattern the AI should follow.
- Make examples recent and representative, not aspirational. Use real work from the last month or two. Use work your team actually produced and agreed was solid. Use current standards, not the standard you wish you had. The closer the example matches your actual output quality, the better the AI will match your actual needs.
- Build a reusable example library and update it regularly. Store your best vendor evaluations, SOPs, and process assessments where your team can access them. Label them clearly with dates. Use them again and again. Update quarterly or when processes change significantly.
- Strip sensitive details but keep the substance. Vendor names become "Vendor A," amounts become "[AMOUNT]," but the structure and approach stay intact. The pattern is what the AI learns from, not the specific data.
- Few-shot prompting is acceleration, not automation. The AI's output will be much closer to your standard, but always verify it. Read what was produced. Spot-check key judgments. Make sure the AI understood what mattered about your example.
Frequently Asked Questions
Q: If I show the AI two different vendor evaluations, will it get confused about which format to use?
Only if they're very different. If both follow the same structure (Cost Analysis, Service Capabilities, Compliance, Risks, Recommendation) but reach different conclusions, the AI will understand: this is the structure we use, and different vendors lead to different recommendations. What confuses it is inconsistent structure across examples. Keep your examples consistent in form, varied in outcome.
Q: Can I use examples from other companies or from the internet?
Technically yes, but they're less effective than your own examples. Generic examples teach the AI generic patterns. Your own examples teach it your standards. If you don't have recent examples of your own, create a template first, do one draft yourself, then use that draft as your few-shot example going forward.
Q: What if my best example is actually pretty mediocre? Is there a better example I should use?
Use the best one you actually have. If your mediocre example represents your actual standard, that's what the AI should replicate. Don't use an idealized or aspirational example. Your team won't be able to follow the output if it's better than what you actually produce. If your examples are all mediocre, that's useful information, your team should invest in raising the quality of the work you do, and then use those improved versions as examples going forward.
Q: How often should I update my few-shot examples?
When your processes change significantly, update them. When your team's standards improve, update them. Generally: quarterly review, update as needed. If you made a major process change, update immediately. If you realized a previous example had a compliance issue, replace it. But don't update constantly. Stability matters. Use an example for a few months, then refresh it.
Q: Does few-shot prompting work with short tasks, or is it only useful for long documents?
It works for both. Few-shot prompting improves a brief vendor risk summary just as much as a long SOP. The principle is the same: show the AI what you want. Whether it's 200 words or 2000, an example teaches better than a description.
Skill.re