Running HR AI Pilots: From Idea to Evidence to Scale
Overview
You've identified a novel AI application: predictive retention. It could save your company millions. You could gain competitive advantage. Your CEO is interested. So you start a pilot.
Six months later, you've spent $300K, trained 150 people, and... you still don't know if the thing works.
This is the fate of most HR AI pilots. They're ambiguous. They generate activity, not clarity. They consume budget without generating the evidence you need to make a scale/kill decision.
The best pilots follow a different approach: They're tight. They have a hypothesis. They generate clear data. They produce a decision, not a report.
>
Executive Summary: Effective HR AI pilots follow a disciplined methodology: hypothesis-driven design, tight scope, clear success metrics, built-in decision criteria, and ruthless measurement. Bad pilots are exploratory, open-ended, and designed to justify what was decided before they started. Good pilots produce evidence that informs whether to scale, adjust, or kill. The difference is framework.
Purpose Statement
By the end of this lesson, you'll know how to design and run pilots that generate evidence rather than activity, how to set kill criteria that you'll actually use, and how to scale what works.
Why This Matters for HR Executives
From L1-L4, you've learned to assess and implement AI solutions that are proven elsewhere. From Lesson 1, you've learned to identify novel applications that could create competitive advantage.
But there's a gap: How do you move from "This could work" to "This definitely works" to "Let's scale it"?
Pilots are the bridge. But most pilots are poorly designed. Here's why:
Mistake 1: Open-ended scope. "Let's pilot predictive retention to see what happens." See what happens? You'll generate a lot of data and still not know whether to scale.
Mistake 2: Wrong success metrics. You measure what's easy to measure (training completion, tool adoption) instead of what matters (did we actually predict retention? Did retention improve?).
Mistake 3: No kill criteria. If a pilot underperforms, you're supposed to kill it. But most teams don't, because they've built an emotional investment. Instead, they say, "Well, the model wasn't quite right, but we learned valuable things." That's code for "we wasted money and are hoping you don't notice."
Mistake 4: Misaligned participants. You run a pilot with people who want it to work. You get confirmation bias. You get artificially high adoption. You can't tell what real-world adoption looks like.
Mistake 5: No clear decision process. At the end of the pilot, what's your decision? Scale? Adjust? Kill? Who decides? Based on what criteria? If this isn't clear upfront, you'll argue about it when results come in.
This lesson teaches the discipline that prevents these mistakes.
The Pilot Methodology: Five Phases
Phase 1: Hypothesis and Design (Weeks 1-3)
Before you build anything, get clear on what you're testing.
Write your hypothesis:
Not "Let's see if predictive retention works." But:
"We hypothesize that a machine learning model trained on 5 years of historical exit data can predict with 75%+ accuracy whether an at-risk employee will leave in the next 6 months, enabling managers to intervene and retain 40% of identified at-risk employees."
This is specific. It's testable. It sets success criteria. It's different from "let's explore."
Define your scope:
A good pilot is narrow.
- Population: "Sales team in North America" (not "entire company")
- Duration: "12 weeks" (not "open-ended")
- Success metrics: "Model accuracy, adoption rate, retention improvement" (not "engagement" or "vague learning")
Example bad pilot scope: "Let's try predictive retention across the organization and see if we can improve overall retention."
Example good pilot scope: "We'll build a model to predict attrition in our engineering team, deploy it to 20 engineering managers in the US, measure model accuracy and whether they intervene and whether retained employees actually stay, over 12 weeks."
Build your success rubric:
Before you start, define what success looks like. Example:
SUCCESS RUBRIC FOR RETENTION PREDICTION PILOT
Model Accuracy: SUCCESS if 75%+, ADJUST if 65-75%, KILL if <65%
Manager Adoption: SUCCESS if 70%+ use the tool, ADJUST if 50-70%, KILL if <50%
Manager Intervention: SUCCESS if 60%+ engage in development conversations with identified at-risk employees, ADJUST if 40-60%, KILL if <40%
Retention Improvement: SUCCESS if 50%+ of at-risk employees identified stay, ADJUST if 30-50%, KILL if <30%
OVERALL DECISION:
- All SUCCESS or mostly SUCCESS = Scale
- Mix of SUCCESS and ADJUST = Adjust and run extended pilot
- More than one KILL = Kill pilot entirely
This is your decision framework. Agree on it upfront. Don't negotiate it when results come in.
Resource the pilot properly:
Under-resourced pilots fail and tell you nothing. Example: You assign a junior data analyst 20% time to run a predictive retention pilot. That person doesn't have the skills. You're not surprised when it underperforms.
Good pilot budgeting:
- Data work: $30-50K
- Tool/platform: $10-20K (or use existing vendor)
- Change management and training: $20-30K
- Measurement and analysis: $10-20K
- Leadership time and effort: ~5 hours/week for 12 weeks
- Total: ~$70-120K
If a pilot is costing you less than this, you're probably under-resourced. If it's costing more, your scope is too big.
Phase 2: Build and Prepare (Weeks 3-6)
You've got your hypothesis and scope. Now build and prepare to test it.
Build the minimum viable pilot:
Don't build perfection. Build enough to test the hypothesis.
Example: For predictive retention, you need:
- Historical exit data (who left, when)
- A model trained on that data
- A way for managers to see predictions
- A way to record whether they intervened
- A way to track whether the at-risk person stayed
You don't need: Beautiful UI, mobile app, enterprise security, 100% data quality, integration with every system.
Build fast. Test the hypothesis. Iterate based on learning.
Train your pilot participants:
Your participants need to understand:
- Why this matters (business context)
- What you're testing
- What you're asking them to do
- What support is available
Example training agenda (2-3 hours):
1. Business context: Why retention in this team matters (30 min)
2. How the model works: What it's predicting, what it's not, what to trust and not trust (45 min)
3. How to use it: Here's the interface, here's how to interpret a prediction (30 min)
4. What happens next: We'll track whether you use it, whether you intervene, what happens to retention (15 min)
Don't over-train. People learn by doing.
Set up measurement infrastructure:
You need to measure:
- Usage data: Are people actually using the tool?
- Intervention data: When a prediction is shown, what does the manager do?
- Outcome data: Did the predicted person stay?
- Qualitative data: What do managers think of the model? Is it helpful? Frustrating?
Set this up before the pilot starts. Don't try to figure it out halfway through.
Phase 3: Execute and Observe (Weeks 6-10)
The pilot is running. Your job now is to observe, support, and measure, not to convince people it's working.
Weekly measurement rhythms:
- Monday: Tool usage dashboard (who used it, how much)
- Wednesday: Manager feedback (what problems are emerging)
- Friday: Leadership sync (are we on track, do we need to adjust)
Intervention points: If something's clearly broken, fix it. If 70% of managers say the predictions don't make sense, that's a signal the model needs adjustment. Fix it.
But here's the key: Don't change your success criteria based on what you're seeing. If you said "70% adoption is success," don't change that to 40% when adoption is running at 45%. It's okay to adjust the mechanism (how you're measuring, what data you're tracking) but not the success criteria.
Weekly manager check-ins:
- "Are you using the tool?"
- "Is it helpful?"
- "What's confusing?"
- "What would make it more useful?"
This isn't selling. This is learning. You're gathering qualitative data on adoption friction and value perception.
Phase 4: Measure and Analyze (Weeks 10-12)
The pilot is wrapping up. Time to measure.
Quantitative analysis:
- Model accuracy: Did the model predict accurately?
- Adoption: What percentage used it? How much?
- Intervention: When predictions were shown, what did managers do?
- Outcome: Did the at-risk people stay or leave?
Example output:
MODEL ACCURACY: 78% (SUCCESS - meets 75% target)
โโ True Positive Rate: 82% (of people who left, 82% had been flagged)
โโ False Positive Rate: 15% (of people flagged who didn't leave)
โโ Implication: Model is reliable. When it says someone's at risk, pay attention.
ADOPTION: 68% (ADJUST - below 70% target but close)
โโ Usage frequency: 2.3 times/month average
โโ Variance: 40% of managers use weekly; 30% use monthly; 30% never used
โโ Implication: Adoption is real but not universal. Usage patterns matter more than average.
MANAGER INTERVENTION: 52% (ADJUST - below 60% target)
โโ Development conversations: 35% had explicit career conversations
โโ Compensation conversations: 12% adjusted compensation
โโ Role/opportunity conversations: 5% changed role or opportunity
โโ Implication: Managers see the predictions but hesitate to act. Need better guidance on what to do.
RETENTION OUTCOME: 48% of at-risk identified stayed (ADJUST - below 50% but close)
โโ Of those who stayed, 73% reported a development or comp conversation
โโ Of those who left, only 18% reported a conversation
โโ Implication: Conversations work. Interventions move the needle. But we need higher intervention rate.
Qualitative analysis:
Interview 8-10 managers from the pilot group. Ask:
- Did you find the predictions accurate?
- Did it help you have better conversations?
- What made it hard to act on the predictions?
- Would you use this if we continued?
- What would make it more useful?
Synthesize what you hear. You're looking for patterns. "80% of managers said the model was accurate but didn't know what to do when a prediction showed up." That's data. "Most thought it was threatening to their autonomy as managers." That's different data.
Phase 5: Decide and Communicate (Week 12-13)
You've got your data. Time to decide.
Compare against your success rubric:
Model Accuracy: 78% = SUCCESS
Adoption: 68% = ADJUST (but operational)
Manager Intervention: 52% = ADJUST
Retention Impact: 48% = ADJUST but directional
What do these results mean?
The model works. That's the good news. Adoption is real, and there's a clear connection between manager intervention and retention.
The issue is that managers aren't intervening at a high enough rate. Why? Training analysis shows managers don't know what to do when they see a prediction.
Decision: Adjust and scale to extended pilot
Not: "Kill it, it didn't work."
Not: "Scale it now."
But: "We've proven the core hypothesis (model accuracy + retention impact). We need to improve intervention capability. Let's run a 6-month extended pilot with better training and manager support structure."
Extended pilot plan:
- Manager coaching: Manager of each predicted employee gets a 30-min coaching call: "Here's what the data says. Here's what conversations to have. Here's outcomes we've seen."
- Career conversation template: We provide a structure for the conversation (not a script, a structure)
- Weekly check-in: Manager reports back on whether conversation happened, what the person said
- Outcome tracking: Did they stay?
- Go/no-go decision at 6 months
Communicate the result:
To your leadership:
"Predictive retention pilot shows that our model can accurately predict attrition with 78% accuracy. When managers intervene, retention improves. Current adoption and intervention is suboptimal (68% adoption, 52% intervention) due to training gaps. We're running a 6-month extended pilot with enhanced manager support. Clear scale/kill decision will come in 6 months."
This is honest. It's clear. It sets expectations. You're not overselling. You're also not burying the finding.
Kill Criteria That Actually Work
The hardest part of piloting is being willing to kill when something doesn't work.
Here's how to make this real:
1. Establish kill criteria upfront.
Don't wait until results come in to figure out when something is bad enough to kill.
Examples:
- Model accuracy below 60%
- Adoption below 40% after training
- Manager intervention rate below 30%
- Retention of identified at-risk employees below 20%
Pick 3-4 kill criteria. If you hit any of them, you kill the pilot.
2. Assign a decision-maker.
Who calls kill? Is it the CHRO? The VP of Analytics? A committee? If it's a committee, you'll argue and fudge. Make it one person.
"The VP of HR owns the scale/adjust/kill decision. She'll make the call based on the rubric we've established."
3. Make it easy to kill.
If you've invested $100K and it looks like you need to kill, you'll find reasons not to. Set the emotional tone early: "Some of our pilots will fail. That's okay. We're learning. Killing a pilot is a success, not a failure."
Build a celebration moment: "Thank you to everyone who ran this pilot. We learned that X doesn't work for our organization. That's valuable. Here's what we're doing differently next time."
4. Have a next step ready.
When you kill a pilot, you don't want to leave your participants demoralized. Have a transition plan: "The predictive retention pilot isn't scaling, but we're investing in manager training instead, and we'd love your input on that. Here's the pivot..."
The Pilot Design Template
Use this template for every pilot:
PILOT DESIGN
HYPOTHESIS:
[What are we testing? Be specific and measurable.]
SCOPE:
โโ Population: [Who?]
โโ Duration: [12 weeks]
โโ Success metrics: [What matters?]
โโ Kill criteria: [When do we stop?]
RESOURCES:
โโ Budget: $X
โโ Team: Who?
โโ Timeline: Build (3 weeks), execute (7 weeks), measure (2 weeks)
โโ Leadership sponsor: [Who?]
SUCCESS RUBRIC:
โโ Metric 1: SUCCESS if X, ADJUST if Y, KILL if Z
โโ Metric 2: SUCCESS if X, ADJUST if Y, KILL if Z
โโ Overall decision logic: [X=Scale, Y=Adjust, Z=Kill]
MEASUREMENT PLAN:
โโ Quantitative: [What data will you track?]
โโ Qualitative: [Who will you interview? When?]
โโ Weekly reporting: [What goes to leadership?]
SCALE PLAN (if successful):
โโ Timeline: [When would you roll out?]
โโ Population: [How many people?]
โโ Resources required: [Budget? Team?]
โโ Risk mitigation: [What could go wrong?]
What to Do Monday Morning
Pick one pilot to run. Preferably one that's novel (not just better/faster/cheaper).
Write your hypothesis. Make it specific, measurable, testable.
Define your scope. Narrow is better. 12 weeks is better than open-ended.
Write your success rubric. Before you start, what does winning look like? What does losing look like?
Identify your kill criteria. When would you stop?
Get leadership alignment. Show them the rubric. Make sure they're comfortable with the kill criteria.
Key Takeaways
- Pilots are about generating evidence, not justifying decisions. If you know what you want to prove before you start, you're not really piloting.
- Narrow scope is better. Small, tight pilots teach you more than sprawling ones.
- Kill criteria matter. If you don't define them upfront, you'll rationalize failure when results come in.
- Measurement is not optional. Without tight measurement, you don't know whether to scale or kill.
- Be willing to kill. Some pilots will fail. That's not a waste of time. That's learning.
FAQ
Q: How long should a pilot actually be?
A: 8-12 weeks for most HR applications. Longer than 12 weeks and you're delaying the decision. Shorter than 8 weeks and you don't get enough data.
Q: What if we're pressured to scale before the pilot is done?
A: Show the leadership team the success rubric and the measurement data. "Here's what we said success looked like. Here's where we are. When we hit these targets, we scale." Don't cave.
Q: What if the pilot works but it's expensive?
A: Then you've got a different decision to make: Is the ROI strong enough to justify the cost? That's a business decision, not a pilot decision. At least you know it works.
Q: Should we run multiple pilots in parallel?
A: Not in year 1. Focus. Run one tight pilot, learn, decide, and then move to the next. Multiple pilots create cognitive load and diffuse resources.
What's Next
You've run pilots and you've got evidence. Some work. Some don't. Now you need to know how to scale what works without losing the rigor and care you had in the pilot. That's the focus of the next lesson: Scaling HR AI Initiatives.
Skill.re