AI for IT Certification
Aware · M50 · lesson 50 of 120 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Defining Success Metrics
📖
now learning

Defining Success Metrics

15 min

Overview

Your team implemented an AI-assisted ticket routing system three months ago. Everyone says it's working great. But when the CFO asks "What's the ROI?", you don't have a clear answer.

"It seems faster," you say.

"Compared to what?" she asks.

"... Compared to before."

"By how much? How are we measuring?"

Silence.

This is the challenge of measuring AI impact in IT: everyone feels it's working, but if you can't quantify it, you can't defend the investment, iterate on it, or convince leadership to fund the next AI initiative.

By the end of this lesson, you'll have a framework for choosing the right metrics, setting baselines, and measuring what actually matters.

Purpose

Metrics serve two functions:

  • For you: To understand whether the AI initiative is working and where to improve
    - For leadership: To justify the investment and drive future funding decisions

Bad metrics feel wrong (you know something's working, but the number doesn't show it) or misleading (the number looks great but doesn't match reality).

Good metrics tell you the truth and align with business value.

Why This Matters

Without clear metrics:

  • You can't answer the question "Is this working?"
  • You can't compare AI initiatives to decide which ones to fund
  • You can't improve (you don't know what to improve)
  • Leadership loses confidence (you're asking for money but can't show results)
  • You're vulnerable to politics (whoever tells the best story wins, not whoever delivers value)

With clear metrics:

  • You know whether initiatives are working
  • You can prioritize future investments
  • You can iterate and optimize
  • Leadership trusts your judgment
  • You can scale what works

Core Concepts

Key Insight: Leading vs. Lagging Indicators

Metrics come in two types. You need both.

Lagging indicators measure the final outcome. They answer "Did we succeed?"

  • Customer satisfaction improved by 15%
  • Tickets resolved 25% faster
  • Cost per incident dropped 20%
  • Employee satisfaction increased by 10%

Lagging indicators are trustworthy but slow. You might not know the results until weeks or months after implementing something.

Leading indicators measure activities that predict outcomes. They answer "Are we on track?"

  • 75% of eligible tickets are going through AI routing
  • Team members are using the AI tool daily
  • Adoption rate is 80%+
  • New team members are trained within 2 days

Leading indicators are actionable quickly. If leading indicators are bad, you can fix things before lagging indicators show problems.

Both matter:

  • Use leading indicators to know if you're on track (weekly/daily)
  • Use lagging indicators to know if you succeeded (monthly/quarterly)

If your leading indicators are great but lagging indicators are bad, something's wrong with your leading indicators. Adjust.

Key Insight: Quantitative vs. Qualitative Metrics

Most people think of quantitative metrics (numbers). But qualitative metrics are equally important.

Quantitative metrics:

  • Time-to-resolution
  • Ticket deflection rate
  • Accuracy rate
  • False positive rate
  • Cost per incident

These are specific and measurable. Easy to track and compare.

Qualitative metrics:

  • How confident are teams in AI recommendations?
  • What problems is the tool causing?
  • Where do teams struggle?
  • What's the biggest blocker to adoption?
  • Would you recommend this tool to another team?

These are harder to quantify but often more important. Qualitative feedback tells you why quantitative metrics are what they are.

Both matter:

  • Quantitative metrics show whether you succeeded
  • Qualitative metrics show why you succeeded or failed, and what to do next

Measure both.

Key Insight: The Core Metrics for IT AI Initiatives

Different AI initiatives measure different things, but most IT AI projects care about these categories:

Category 1: Productivity

  • Mean time to resolution (MTTR)
  • Mean time to detect (MTTD)
  • Time per ticket
  • Cost per ticket resolved
  • Tickets handled per person per day

These measure whether the AI made people faster.

Category 2: Quality

  • First-time resolution rate
  • Escalation rate
  • Rework rate
  • False positive rate (for detection/alerting tools)
  • Accuracy of AI recommendations

These measure whether the AI improved quality or just speed.

Category 3: Adoption

  • % of eligible work using the AI tool
  • % of team members trained
  • Daily active users
  • Feature utilization rate
  • User satisfaction with the tool

These measure whether people are actually using what you built.

Category 4: Compliance/Risk

  • Compliance violations involving the AI tool
  • Security incidents caused by the tool
  • False negative rate (problems missed by the AI)
  • Human override rate (how often people correctly reject the AI)

These measure whether the AI is safe and reliable.

Category 5: Financial

  • Total cost of ownership
  • Cost savings
  • Revenue impact
  • ROI (return on investment)
  • Cost per unit improvement (e.g., cost per 1% improvement in MTTR)

These measure whether the investment was worthwhile.

For each initiative, you probably care about 2-3 of these categories most. Pick those and measure them well, rather than trying to measure everything.

Key Insight: Setting Baselines

You can't measure improvement without a baseline. Baselines are your before-state.

What to measure before you implement:

  • Same metrics you'll track after
  • Historical data (at least 4-8 weeks, ideally 3 months)
  • Capture context: Were there disruptions? Was the team operating normally?

Example:

"In the 8 weeks before AI implementation:

  • Average MTTR: 4.2 hours
  • First-time resolution: 72%
  • Cost per ticket: $47
  • Team utilization: 78%"

Common baseline mistakes:

  • Measuring only one week (too much variance)
  • Measuring during a disruption (infrastructure changes, holidays, new people)
  • Choosing metrics different from what you'll track after (makes comparison impossible)
  • Assuming baseline will stay the same (it usually improves naturally over time)

A good baseline is:

  • At least 4-8 weeks of historical data
  • Captured during normal operations
  • Using the same measurement methodology you'll use after
  • Documented with context (any unusual events?)

Key Insight: Avoiding Vanity Metrics

Vanity metrics feel good but don't tell you whether the initiative actually works.

Examples of vanity metrics:

  • "We deployed the tool to 500 people" (did they use it? did it help?)
  • "Tool usage increased 40%" (from what baseline? is increased usage good?)
  • "User satisfaction: 4.2/5 stars" (satisfied with what? the tool? the experience?)
  • "Time saved: 10,000 hours per year" (estimated or measured? is this realistic?)

Vanity metrics often sound good in a PowerPoint but don't survive scrutiny.

How to avoid vanity metrics:


  • Ask "So what?"

"Tool adoption is 75%." So what? Is that good? What impact does it have?

"Adoption is 75%, and teams using it resolve tickets 25% faster." That's meaningful.


  • Focus on outcomes, not activities

Don't measure "people trained." Measure "people using it correctly."

Don't measure "time spent in system." Measure "tickets resolved faster."


  • Compare to baseline

Don't just say "metric X is Y." Say "metric X improved from Y1 to Y2, an improvement of Z%."


  • Measure impact, not correlation

"Tickets resolved faster and we deployed the tool" is not proof of impact. Measure actual impact:

  • Before: 4.2 hours MTTR
    - After: 3.1 hours MTTR
    - Attribution: Analysis shows the AI tool accounts for 40% of the improvement, other factors 60%

Key Insight: Small Team Challenges

IT operations teams are usually small. That makes measurement harder.

Challenge 1: Not enough variance to measure

You have 8 people. Two are sick. That's 25% of capacity gone. It's hard to see the signal of the AI tool over the noise of normal variation.

Solution: Measure longer (3 months instead of 4 weeks). Use statistical methods (compare trends, not individual data points). Measure the change in trend, not the change in absolute numbers.

Challenge 2: Confounding variables

You implement the AI tool. Meanwhile, you hire a new person and upgrade infrastructure. What drove the improvement? Hard to say.

Solution: Document what else changed. Try to isolate the AI's contribution through controlled comparison (if possible) or honest assessment ("50% of improvement is the AI tool, 30% is new hire, 20% is infrastructure").

Challenge 3: Measurement overhead

Measuring takes time. On a small team, that overhead is real.

Solution: Automate measurement where possible (pull metrics from tools, don't manually collect). Measure monthly, not daily. Keep metrics simple.

Practical Use Cases

Use Case 1: Choosing Metrics for an AI Incident Response Tool

Scenario: You're deploying AI-assisted incident response. What should you measure?

Your context:

  • 15-person ops team
  • Currently handle ~500 incidents per month
  • Want to improve MTTR and reduce on-call burden
  • Have existing monitoring and incident tracking systems

Metric selection:

Primary metrics (track these closely):

  1. MTTR (Mean Time To Resolution)
  • Baseline: Currently 4.2 hours
  • Target: 3.0 hours (29% improvement)
  • Measurement: Automated from incident tracking system
  • Frequency: Weekly review
  • First-time resolution rate
    - Baseline: Currently 68%
    - Target: 78% (reduces rework)
    - Measurement: Manual classification in incident tracking
    -
    Frequency: Monthly analysis

  • Adoption rate
  • Baseline: N/A (pre-launch)
    - Target: 80% of eligible incidents using the AI tool
    - Measurement: Incident tracking system
    - Frequency: Weekly

Secondary metrics (track less frequently):

  1. Cost per incident (MTTR × hourly cost)
  2. On-call satisfaction (do people feel less burdened?)
  3. Escalation rate (do fewer incidents get escalated?)

Measurement plan:

  • Baseline: Measure current state for 8 weeks before launch
  • Launch: Week 0
  • Tracking: Weekly metrics review, monthly detailed analysis
  • Evaluation: 3 months post-launch (formal assessment)

You're not measuring everything. You're measuring what matters for your goals.

Use Case 2: Establishing a Baseline When You Don't Have Historical Data

Scenario: You're implementing AI for a process that's never been formally measured. You don't have historical data.

Approach:

Option 1: Create baseline by measurement

  • Measure the current process intensively for 4 weeks
  • Track metrics manually if needed
  • This becomes your baseline
  • Cost: 1-2 weeks of measurement overhead

Option 2: Use expert estimation

  • Ask experienced team members: "How long does this typically take?"
  • Get multiple estimates
  • Average them
  • Note this as "estimated baseline, not measured"
  • Less accurate but faster

Option 3: Parallel measurement

  • Run the AI tool alongside the old process for 2 weeks
  • Measure both
  • Compare them
  • This is expensive but gives you the most accurate comparison

Option 4: Hybrid approach

  • Do option 1 (4 weeks baseline measurement)
  • Run both processes for 1 week in parallel
  • Compare baselines to reality
  • This validates your baseline

Choose based on your resources and accuracy requirements.

Use Case 3: Addressing a Metric That Doesn't Match Reality

Scenario: Your AI tool shows 25% improvement in MTTR in the metrics. But the team says it doesn't feel faster.

Diagnosis:

  • Check the measurement
    - How is MTTR defined? (detection to full resolution?)
    - Was the baseline accurate?
    -
    Is anything gaming the metric? (counting time differently?)

  • Look at the data
  • Did the top performers improve more or less than average?
    - Did specific types of incidents improve more? Which ones?
    -
    Did improvement happen all at once or gradually?

  • Ask the team
  • "The metrics show improvement. But you say it doesn't feel faster. What's going on?"
    - Often: "The AI makes easy tickets faster, but we're spending more time on hard tickets, so it evens out."
    -
    Or: "The metrics are measuring time from detection. Detection was already fast. We're slower at resolution."

  • Investigate deeper
  • Maybe the metric is the problem (it doesn't capture what matters)
    - Maybe implementation is the problem (the tool isn't being used well)
    -
    Maybe the tool is the problem (it's actually not helping)

  • Adjust
  • Change the metric if it doesn't measure what you care about
    - Change how the tool is used if it's not being applied well
    - Accept that the tool has limitations and measure what it actually does

Lesson: Trust both the metrics and the team's gut. If they don't match, investigate. Usually there's a mismatch in how something is defined, not a problem with the tool.

Examples

Example 1: A Metrics Dashboard for IT AI Initiatives

Title: AI Incident Response Initiative, 3-Month Review

Executive Summary

  • MTTR: 4.2h → 3.0h (-29%, target 3.1h) ✓
  • Adoption: 78% of eligible incidents (target 80%) ✓
  • First-time resolution: 68% → 74% (+9%, target 75%) ✓
  • Team satisfaction: 72% report reduced burden (qualitative) ✓
  • Cost per incident: $47 → $33 (-30%)
  • Status: On track. Initiative delivering expected value.

Primary Metrics

Metric
Baseline
Target
Month 1
Month 2
Month 3
Trend
Status

MTTR (hours)
4.2
3.1
4.1
3.5
3.0

Adoption (%)
N/A
80%
45%
62%
78%

First-time resolution (%)
68%
75%
68%
71%
74%

~

Cost per incident ($)
$47
$35
$47
$40
$33

Secondary Metrics

Metric
Baseline
Target
Actual
Status

Escalation rate
22%
<18%
19%

On-call satisfaction
Not measured
70%
72%

Tool accuracy (AI recommendations followed correctly)
N/A
75%+
76%

Adoption Metrics

Metric
Month 1
Month 2
Month 3
Trend

Daily active users
8/15 (53%)
11/15 (73%)
13/15 (87%)

Avg incidents per day using AI
12/35 (34%)
20/35 (57%)
28/36 (78%)

Training completion
40%
80%
100%

Qualitative Feedback

  • "The AI catches things I miss when I'm busy", Ops engineer
    - "I'm more confident that tickets are going to the right people", Senior ops
    - "Sometimes the categorization is wrong and I have to override", Ops engineer
    - "On-call is less stressful. I trust the AI to not miss critical things": On-call rotation manager

Financial Impact

  • Cost of AI tool: $50K/year
    - Annual cost savings (MTTR improvement): $45K
    - Annual cost savings (fewer escalations): $12K
    - Total annual benefit: $57K
    - Net ROI: 14% in year 1 (not great, but expected to improve as adoption increases)

Findings & Recommendations

What's working:

  • MTTR improvement exceeded expectations
  • Adoption is progressing well (78% at month 3, target was 80%)
  • Cost savings starting to accrue

What needs improvement:

  • First-time resolution increase is slow (currently 74%, target 75%)
  • Action: Additional training on complex ticket analysis
  • Two team members still not using the tool
  • Action: 1-on-1 coaching on their specific use case

Next steps:

  • Month 4: Focus on quality (first-time resolution)
  • Month 5-6: Document best practices and develop training for new hires
  • Month 6: Evaluate whether to expand AI to other functions

Example 2: A Statistical Approach to Small Sample Sizes

Problem: You have 8 people on the ops team. One person is sick, so you're down to 7. One person is out on vacation. Do your metrics still mean anything?

Solution: Measure trends, not absolute numbers

Instead of:

"This week, tickets resolved: 45. Last week, 42. We're up 7%!"

Do this:

"Over the last 8 weeks, we've resolved an average of 43 tickets per week. Trend is slightly upward (roughly 0.5 tickets per week improvement). The variation in the data is about ±8 tickets, so this trend is within normal variation."

Then:

"The AI tool launched in week 4. We'd expect to see a noticeable shift in the trend if it were working. After 4 weeks (week 8), we're seeing a shift of roughly 2 tickets per week improvement, which is larger than normal variation. Preliminary indication that the tool is helping, but we need more data before concluding."

This approach:

  • Acknowledges that small teams have natural variation
  • Doesn't over-interpret small changes
  • Looks for signal in the noise (the trend)
  • Waits for sufficient data before concluding

Example 3: Attribution When Multiple Changes Happen At Once

Situation:

  • Month 1: Deploy AI tool
  • Month 1: Hire new senior ops engineer
  • Month 2: Upgrade infrastructure
  • Improvement in MTTR: 25%

Question: How much of the improvement is the AI tool?

Analysis:

Method 1: Expert judgment

  • "The AI tool accounts for 40% of the improvement"
  • "The new hire is responsible for 35%"
  • "The infrastructure upgrade is responsible for 25%"
  • Note: This is a rough estimate

Method 2: Controlled comparison

  • If possible, measure a team without the AI tool but with the new hire and infrastructure upgrade
  • See how much they improved
  • The difference is the AI tool's contribution
  • Often not feasible in small organizations

Method 3: Ask the team

  • "What's making the biggest difference?"
  • Senior engineer: "The infrastructure upgrade was important, but the AI tool really helps with initial triage. I'd say 40% of my speedup is the tool."
  • Junior engineer: "I'm using the tool for almost everything. Maybe 50% of my improvement is the tool."

Method 4: Phased rollout

  • Deploy the tool to one team first
  • Deploy new hire to a different team
  • Compare the teams
  • Shows which change drove improvement

Be transparent: "The improvement is due to multiple changes. Our estimate is that the AI tool contributed about 40% of the improvement, but we can't be 100% certain."

Anti-Patterns

Anti-Pattern 1: Measuring Too Many Metrics

The trap: You try to measure everything. You end up with a dashboard with 40 metrics. Nobody reads it. You can't tell what's working.

Why it fails: Measuring is expensive. Too many metrics means you're measuring things that don't matter. Decision-making suffers.

Fix: Pick 3-5 metrics that actually matter. Measure those well. Add more only if you need them.

Anti-Pattern 2: Metrics Without Context

The trap: "MTTR is 3.2 hours." Is that good? Compared to what? Better than before?

Why it fails: A number without context is meaningless. You don't know if you've succeeded.

Fix: Always provide context. "MTTR was 4.2 hours, now 3.2 hours (24% improvement, target was 25%)."

Anti-Pattern 3: Vanity Metrics

The trap: "We've saved 10,000 hours per year!" Sounds great, but is it real? How did you calculate it?

Why it fails: Inflated numbers look good until scrutiny reveals the flaws. Then credibility is lost.

Fix: Be conservative. Measure what you can verify. Estimate what you can't.

Anti-Pattern 4: Not Measuring Qualitative Feedback

The trap: You have great quantitative metrics but no idea why they're great or bad.

Why it fails: You optimize for the metric and miss what's actually going on. Numbers and reality diverge.

Fix: Always get qualitative feedback. Ask teams what's working and what's not. Use that to interpret numbers.

Human Judgment Checkpoints

  • Does each metric pass the "so what?" test? If you can't explain why the metric matters, remove it.
    - Can you explain the baseline? If you can't defend where 4.2 hours came from, your measurement is weak.
    - Do the team and metrics agree? If team says "This isn't working" and metrics say "It's great," investigate.
    - Are you measuring what you care about? Or are you measuring what's easy?

Key Takeaways


  • Use both leading and lagging indicators. Leading indicators (adoption, usage) tell you if you're on track. Lagging indicators (MTTR, cost, quality) tell you if you succeeded.

  • Measure what matters for your goals. If your goal is speed, measure MTTR. If it's quality, measure first-time resolution. Don't measure everything.

  • Establish baselines before you start. You can't measure improvement without knowing the starting point. Baselines are 4-8 weeks of historical data.

  • Avoid vanity metrics. Focus on outcomes (tickets resolved faster) not activities (people trained). Measure impact, not correlation.

  • Combine quantitative and qualitative. Numbers show you succeeded. Qualitative feedback shows you why.

  • Account for small team variance. With small teams, measure longer, look for trends, and acknowledge that natural variation is large.

  • Be transparent about attribution. When multiple changes happen at once, honestly assess which one drove the improvement.

  • Keep metrics simple. 3-5 core metrics you understand deeply are better than 30 metrics you don't.