The Cardinal Rule: Verify Everything Before It Touches a Process
Overview
A marketing team uses AI to generate social media content. An error in the AI output goes live: A product feature that was supposed to launch next month goes live in today's posts. The error gets 500 shares before it's corrected. Customers start pre-ordering a product that isn't available. The company's support team gets hammered with inquiries, then has to send follow-up messages explaining the error. It's messy but recoverable.
The same company's operations team uses AI to help design a new vendor payment process. The AI makes an error in the verification step logic. It suggests vendors get paid before their deliverables are confirmed instead of after. The error gets baked into the SOP and deployed. For three weeks, vendors get paid regardless of delivery quality. By the time the error is caught, two vendors have already exploited the gap, delivering partial work and getting full payment. The company loses money, vendor relationships are strained, and there's now a compliance issue because the process violated company policy.
Both teams trusted AI output. Only one team faced serious consequences.
This is the fundamental difference between operations and other functions: Processes touch everything. A broken process affects every transaction that runs through it, every person who relies on it, and every stakeholder connected to those people. The stakes are higher. The blast radius is wider. A single error in a process that runs 1,000 times a month doesn't cause one problem. It causes problems at scale.
This is why the cardinal rule of using AI in operations is: Verify everything before it touches a process. Not most things. Everything. And I mean verify rigorously, not a glance, but actual verification against source material and operational reality.
Why Operations Verification Is Different From Other Functions
Operations professionals understand that process design is consequential. A process change affects everyone who uses it, hundreds or thousands of times. Small errors get magnified across volume. And once a process is live, people stop thinking about whether it's right. They just follow it.
Let's compare stakes across functions:
Marketing: An AI generates a social post with an error. It goes live, gets criticized, gets corrected. Reputational damage is real but time-limited. The cost is embarrassment and brief customer confusion.
Sales: An AI generates an inaccurate competitive analysis. A salesperson uses it in a pitch. The prospect catches the error and loses trust. The cost is one lost deal.
Finance: An AI generates a journal entry with the wrong account coding. It gets reviewed by an accountant before posting. The error is caught and corrected. No damage.
Operations: An AI helps design a process with an error. The error gets built into the SOP. The process runs 1,000 times a month. Each time, the error creates a small problem: a payment that should be conditional is unconditional, an approval that should be required gets skipped, a documentation step that should happen at the end doesn't happen. After three weeks, you have 3,000 small problems. Cumulatively, they're not small anymore.
The difference is scale and systemic impact. Operations is the function where errors compound through volume. An operational error isn't one bad outcome. It's the same bad outcome happening repeatedly until someone finally notices the pattern.
This is why verification for operations is non-negotiable. You're not verifying AI output because you don't trust AI (you don't). You're verifying because the consequences of getting it wrong are distributed across hundreds or thousands of operational transactions.
Important: Operational Errors Scale Through Volume
A single operational error doesn't cost you one incident. It costs you one incident per instance the process runs. If you implement an AI-designed process with an error and it runs 500 times before you catch the error, you've got 500 problems. This is why verification before deployment is non-negotiable, not optional.
What Needs Verification Before It Touches Operations
Let me be specific about what "before it touches a process" means. Here's what needs verification:
Any AI-generated process flow or SOP: If an AI creates or substantially contributes to a process design, verify every major step before the process goes live. This isn't a read-through. It's a line-by-line evaluation against how you actually work and what your requirements actually are.
Any numerical requirement or metric the AI suggests: If an AI recommends "vendor response time should be 4 hours," you need to verify that 4 hours is actually appropriate for your context, that it's achievable, and that it's based on something real (not hallucinated).
Any compliance or policy claim: If an AI says "industry standard requires X," you need to verify this against actual industry standards, regulations, and your organization's actual requirements. Don't assume something is a standard just because an AI said it confidently.
Any comparative analysis used for decision-making: If an AI creates a vendor scorecard, SLA summary, or compliance comparison that you're using to make decisions, verify the key claims against source material before the decision is final.
Any responsibility assignment or approval requirement: If an AI designs who approves what or who's responsible for what, verify it against your actual organizational structure and decision patterns. RACI matrices from AI are famous hallucinations.
Any requirement claimed to be from a contract or existing document: If an AI cites something from your vendor contracts, policies, or SLAs, spot-check it. AI frequently misquotes or extracts from the wrong section. "Section 3.2 requires X", go look at Section 3.2 and verify it actually says X.
What doesn't necessarily need rigorous verification: Initial drafts clearly labeled as drafts. Summaries of unstructured information where you understand the source and purpose. Brainstorm ideas. Explanations or educational content where accuracy doesn't drive decisions. The verification requirement is proportional to how much the output will influence operational decisions or be deployed as truth.
Verification Strategies That Actually Catch Errors
Verification is not a single technique. It's multiple overlapping approaches that collectively catch errors at different levels. Here's how to design a verification process:
Source Verification: Trace Claims Back to Originals
When an AI makes a specific claim about something in source material (contract terms, existing process steps, policy requirements, vendor SLA commitments), verify by going back to the source. Don't verify by asking a different AI. Go to the actual contract, actual SOP, actual policy, actual vendor commitment.
This is tedious. That's fine. Tedium is the price of verification. For high-stakes AI output, you should expect verification to take 20-40% of the time it took the AI to generate the content.
Practical approach: Pick the top 5-10 claims in the AI output (the ones with highest operational impact). For each one, trace it back to source and verify it's accurate. If you find hallucinations in the top claims, you know there might be issues elsewhere, and you need more comprehensive verification.
Structural Verification: Compare Against Known Patterns
When an AI generates something with structure (a process flow, a responsibility matrix, a compliance checklist), compare it against known patterns from similar situations in your organization. Does it match the pattern of how you actually structure things? Or does it diverge in ways that seem off?
Example: An AI generates a vendor onboarding SOP with 15 steps. You look at your actual vendor onboarding and you currently do about 10 major things. The AI added 5 steps that sound reasonable but don't match your current practice. Are these improvements the AI is suggesting, or are these hallucinations? You need to figure out which. The ones that match your current practice are probably accurate. The ones that don't might be invented.
Cross-Reference Against Multiple Sources
If the AI claims something is a requirement or standard (like "SOC 2 certification is standard for technology vendors"), don't just take its word. Check multiple sources: your actual contracts, your industry guidelines, actual vendor commitments, RFP standards in your field. If the "requirement" doesn't appear in multiple independent sources, it's either not actually standard or it's not standard for your context.
Reality Check: Does This Match How We Actually Work?
An AI-generated process might be theoretically sound but operationally unrealistic. Does it require decisions to be made by people or at times when those people are typically unavailable? Does it require resources your organization doesn't have? Does it assume cultural behaviors (like strict adherence to processes) that don't match your actual organization?
These reality checks require someone who knows your organization. A process that works for a formal, structured enterprise might be total fiction for an informal, fast-moving startup. An AI can't know your organization's actual operating culture. You can.
Assumption Verification: Make Implicit Assumptions Explicit
Every AI-generated operational content contains implicit assumptions. About how decisions are made. About what information is available when decisions need to be made. About how compliant people are with processes. About what counts as success or failure.
Make these assumptions explicit and check them: "This process assumes Manager X is available to approve within 4 hours. Is that realistic?" "This assumes vendors will fill out this form completely and accurately. What's our actual experience?" "This assumes the approval gate prevents bad vendors from proceeding. What happens if someone bypasses it?"
Verification isn't about proving the AI wrong. It's about surfacing whether the underlying assumptions are valid in your context.
Tip: Build a Verification Template for Your Most Common AI Tasks
Create templates for verification of the AI outputs your team uses most often. For vendor analysis: "Spot-check top 3 claims against contracts. Cross-reference vendor SLA metrics. Does assessment match your actual experience?" For process design: "Trace every requirement to its source. Reality-check against current practice. Test assumptions with process owners." Use these templates consistently so verification becomes routine.
Real Operational Failures From Unverified AI Output
Let me walk through specific failures that happened because verification didn't happen:
Case 1: The Unverified Compliance Checklist
A COO asked an AI to generate a comprehensive vendor compliance checklist. The AI created a 25-item checklist: SOC 2 Type II certification, ISO 27001, annual penetration testing, liability insurance (amount not specified), compliance attestation, data processing agreement, privacy policy, security questionnaire, audit rights, incident response plan (vendor's), business continuity plan, disaster recovery plan, legal right to audit, annual security review, compliance certification from vendor, third-party audit results, code of conduct, supply chain transparency, conflict minerals attestation, and others.
Without verifying the checklist against what vendors could actually provide or what the company actually needed, it was implemented. The first vendor to be screened against it failed 12 items. The company found that most vendors don't undergo annual penetration testing (it's expensive). The "compliance certification" item was too vague, vendors kept providing various certifications the company didn't need. The liability insurance amount was never specified, so vendors provided everything from $1M to $50M, none of it guidance from the company.
Outcome: The checklist wasn't useful because it was unverified. It was a theoretical best-practice list, not a list of requirements that actually mattered to this company for this category of vendor. Months of procurement delays and vendor frustration while the company figured out what it actually needed.
How it should have been handled: Verify the checklist against actual vendor contracts the company had, cross-reference against what their customers actually required, test against vendors in the category to see what's feasible. Build the checklist from reality, not from an AI's pattern-matching of compliance best practices.
Case 2: The Process With Hidden Errors
A company used AI to redesign their invoice approval process. The AI took the old process (PO โ Goods Received โ Invoice โ Manager Approval โ Finance Review โ Payment) and "improved" it to (PO โ Invoice โ Manager Approval โ Goods Received โ Finance Review โ Payment). The change moved invoice verification before goods receipt verification to "speed up payment."
This looked reasonable in the redesigned flowchart. But nobody verified the logic against actual business risk. What's the consequence of approving payment for an invoice before you've confirmed the goods actually arrived? You now have financial liability before you have the goods. You're paying vendors before you've verified they delivered.
The process ran for six weeks before a finance manager caught the error. In those six weeks, the company overpaid for partial deliveries, paid for goods that never arrived, and had to go back and renegotiate with vendors about what was actually delivered.
How it should have been handled: Before deploying a redesigned process, the designer (human or AI-assisted) should have worked through the risk implications of every change. What's different about the new order? What's the business consequence if each step fails? Why is the new order better? These questions would have surfaced the error immediately.
Case 3: The Vendor Performance Dashboard That Wasn't
An operations manager asked an AI to create a vendor performance dashboard spec based on her vendor contracts. The AI extracted SLA metrics and suggested a dashboard that tracked: Uptime percentage, response time in hours, resolution time in hours, customer satisfaction (CSAT) score, cost per transaction, and on-time delivery rate.
The manager built the dashboard without verifying whether the metrics were actually in the contracts or if the company had the capability to measure them. When she tried to populate the dashboard: Vendor A's contract mentioned uptime but didn't specify 99%, so "uptime percentage" was hard to define. Vendor B's response time commitment was "reasonable response time within business hours," not a number. The company had no way to measure CSAT for vendor services. Cost per transaction wasn't tracked in the company's systems. On-time delivery rate had different meanings for different vendors.
The dashboard couldn't be populated because the AI had recommended metrics that looked standard but weren't actually in the contracts or measurable in the company's systems.
How it should have been handled: Before accepting the AI's metric recommendations, verify each one. For each metric: Is it explicitly in the contract? Does the company have the data to measure it? Is it defined consistently across vendors? Only build the dashboard from verified metrics.
Verification at Different Operational Levels
How rigorously should you verify? It depends on the stakes:
Level 1: Low-Stakes Drafts (10-15% Verification)
Documentation drafts you're clearly going to rewrite. Status summaries. First-pass incident analyses. For these, light verification: Skim it, catch obvious errors, look for anything that's completely wrong. You're not fact-checking everything because you know the draft will be heavily revised.
Level 2: Medium-Stakes Analysis (30-50% Verification)
Vendor performance summaries. Process documentation that will be deployed. SLA analysis. Compliance summaries. These are being used to make real decisions, so verify key claims. Spot-check top metrics. Cross-reference sample claims against sources. You're not verifying 100%, but you're systematically checking the most important things.
Level 3: High-Stakes Process Deployments (50-100% Verification)
New processes, policy changes, anything that will run repeatedly and affect multiple stakeholders. These need rigorous verification. Trace every requirement. Reality-check every assumption. Have subject matter experts review. Make sure the process actually works the way it's documented. This is not quick work, but it's necessary work.
Level 4: Critical/Compliance Work (Avoid AI or Minimal AI Role)
For work that has genuine compliance implications, safety implications, or major financial consequences, AI should have a minimal role. You can use AI for research, analysis, and drafting, but the actual design and decision needs to be human-led and expert-reviewed. Don't use AI to design the controls that prevent fraud. Don't use AI to design safety-critical processes. These require human expertise and accountability.
Building Verification Into Your Workflow
Verification isn't something you add at the end. It's part of the workflow from the beginning. Here's how to structure it:
Before You Use AI: Define Verification Requirements
Before asking an AI to help, decide: How will I verify this output? What checks will I run? How much time will I budget? If verification seems impossible or would take longer than creating from scratch, don't use AI. Save AI for situations where verification is feasible.
During AI Execution: Set Up for Verification
Ask the AI to cite sources. Ask it to flag assumptions. Ask it to note areas of uncertainty. These requests make verification easier, instead of guessing where claims came from, the AI tells you. You still verify independently, but at least you know what to look for.
After AI Generates Output: Execute Verification
Don't rush to use the output. Set aside time for verification. Follow your verification template or checklist. Spot-check key claims. Have others review. Note any uncertainties or areas that need rework.
Only After Verification: Deploy or Use
Only after you've verified the output and addressed any issues should it be deployed, used in decisions, or shared beyond your immediate team. A process that looks reasonable but hasn't been verified is a risk waiting to happen.
After Deployment: Monitor for Errors
Once something from AI is in production (a new process, updated SOP, compliance checklist), monitor it. Are people able to follow it? Are there issues? When you catch issues (and you will), use them to improve both the process and your verification approaches.
What to Do Monday Morning
- For every AI-generated operational output you're currently using or considering, define verification requirements. What would it take to verify this is correct? What are the highest-stakes claims? What sources would you check? Be specific. Write it down.
- Create a verification template for your most common AI operational tasks. If your team regularly uses AI to generate process flows, create a "verify process design" template. If you use AI for vendor analysis, create a "verify vendor claims" template. Use these templates consistently.
- Identify one critical operational process that's been recently updated (either with AI help or otherwise). Audit it. Does it match reality? Would you be comfortable deploying it again? If not, what's wrong? Fix it and use the insights to improve your verification approach going forward.
- For your next AI-generated operational content, execute rigorous verification. Trace claims. Cross-reference. Reality-check. Document what you find. See how long verification actually takes. You'll calibrate better expectations going forward.
- Train your team on verification expectations. If you're using AI to help with operational work, your team needs to understand that verification is non-negotiable, not optional. Make it part of the culture. Reward people who catch verification issues, don't penalize them for slowing things down.
Key Takeaways
- Understand why operations verification is different from other functions. Operational errors scale through volume. One error in a process affects every instance the process runs. Verification is how you prevent one error from becoming 500 errors.
- Make verification part of the workflow, not an afterthought. Define verification requirements before using AI. Set up the output for verification while the AI runs. Execute verification before deployment. Monitor after deployment. Verification is integrated into the process, not added at the end.
- Match verification rigor to stakes. Light verification for low-stakes drafts. Medium verification for analysis that influences decisions. Rigorous verification for processes that will be deployed. No AI role for compliance-critical work without extensive human oversight.
- Trace claims back to sources, don't trust confidence. If an AI claims something is in a contract, go look at the contract. If it claims something is industry standard, check independent sources. Don't assume confidence in an AI output means the claim is verified. Verification requires independent checking.
- Use verification to improve your AI prompting over time. When you find verification issues, note them. Use them to improve how you ask AI for help in the future. "This time, we should ask the AI to cite sources." "Next time, we should reality-check assumptions with department heads first." Verification feedback improves the whole process.
Frequently Asked Questions
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "How much time should verification take compared to AI generation?",
"acceptedAnswer": {
"@type": "Answer",
"text": "It depends on stakes. For low-stakes drafts: 10-15% of generation time. For medium-stakes analysis: 30-50% of generation time. For high-stakes process deployments: 50-100% or more. If verification is taking longer than original creation would have taken, the value prop doesn't work, don't use AI for that task. Verification time is a legitimate operational cost."
}
},
{
"@type": "Question",
"name": "Can I use spot-checking instead of verifying everything?",
"acceptedAnswer": {
"@type": "Answer",
"text": "For lower-stakes work, yes. Spot-check the top 5-7 highest-stakes claims rather than verifying 100%. If your spot-checks reveal hallucinations, you know to do more comprehensive verification. For critical process work, spot-checking isn't enough, verify the major logic and requirements comprehensively before deployment."
}
},
{
"@type": "Question",
"name": "What's the difference between verification and validation?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Verification is checking that something is accurate and matches source material. Validation is checking that something works as intended in your operational context. Both matter. Verify the vendor SLA numbers against contracts (accuracy). Validate the SLA metrics against actual vendor performance (do they work to measure what we care about?). You need both."
}
},
{
"@type": "Question",
"name": "What do I do if AI output is mostly correct but has some errors?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Fix the errors before deployment. Don't rationalize that 'most of it is right.' In operational processes, systematic correctness matters. Even small errors can compound through volume. If you find hallucinations or inaccuracies, fix them all before the output goes into production. Use the discovery of errors to recalibrate your verification approach for next time."
}
},
{
"@type": "Question",
"name": "How do I convince leadership that verification time is worth the cost?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Use historical examples. Show what happens when processes have errors: a process error that ran 500 times costs more to fix retroactively than verification would have cost upfront. Calculate the operational cost of rework, vendor disputes, or compliance issues from unverified AI output. Frame verification as insurance against expensive operational failures. The cost of verification is always less than the cost of fixing an error that scaled through volume."
}
}
]
}
Skill.re