โ†
AI for Instructors & Learning Professionals
Strategic ยท M19 ยท lesson 19 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
The Measurement Stack: Kirkpatrick L1 to L4 and Phillips ROI
๐Ÿ“–
now learning

The Measurement Stack: Kirkpatrick L1 to L4 and Phillips ROI

15 min

The board meeting is in twenty minutes, and a head of learning is holding two numbers. The first is glorious: 4,200 employees completed the new compliance refresh, a 94% completion rate, average satisfaction 4.6 out of 5. The second is the one the CFO will actually ask about: did fewer people violate the policy after the training than before. She has the first number cold. She does not have the second, has never measured it, and is about to learn that the most expensive mistake in learning measurement is mistaking the cheap number for the important one. Every level of the measurement stack costs something different to climb, proves something different at the top, and the whole discipline is knowing which level your question actually needs.

Why The Measurement Stack Exists At All

Learning measurement has a framework because, left to instinct, organizations measure what is easy and call it proof. The easy thing to measure is whether people showed up and whether they liked it. The hard thing to measure is whether anything they do at work changed, and whether that change moved a number the business cares about. The gap between those two is the entire reason a stack exists: it forces you to name the level of claim you are making and to stop letting a completion rate masquerade as evidence of impact.

The Kirkpatrick Four Levels are the spine. Donald Kirkpatrick first laid them out in a 1959 article series, drawn from his 1954 dissertation, and they have outlasted nearly every framework that came after because they answer the four questions a sponsor actually asks, in order of rising cost and rising value. The New World Kirkpatrick Model (2016), built by his son James and Wendy Kirkpatrick, did not replace the levels; it sharpened them, adding the idea that you design the evaluation backward from the result you want and that Level 3 behavior needs ongoing reinforcement and accountability to survive, not a one-time post-test. Phillips ROI Methodology sits one floor above, adding a fifth level that converts results into a financial benefit-cost ratio. Why you care: when you say "the training worked," a precise listener hears a question, which level of worked, and the stack is how you answer without bluffing.

Define one term up front, because the rest of the lesson leans on it. A smile sheet is the end-of-course satisfaction survey, the form that asks whether the learner found the session useful, the facilitator engaging, the pace right. It is a Level 1 instrument. It is not worthless, but it measures reaction, not learning and certainly not behavior, and the smile-sheet trap is the habit of treating a warm smile sheet as if it proved the program changed how people work. It did not. It proved people enjoyed an hour. The entire measurement discipline is, in one sense, a structured refusal to fall into that trap.

The Four Levels: What Each Proves And What It Costs

Walk the stack one floor at a time. Each level answers a different question, demands a different instrument, and costs a different amount of effort. The discipline is matching the level to the stakes, not always climbing to the top.

Level 1: Reaction

Level 1 asks: did the learners react well. Did they find it relevant, engaging, worth their time. The instrument is the smile sheet, usually a five-point scale and a comment box. It is cheap, fast, and almost universally collected. Its honest value is real but narrow: a program people find irrelevant and tedious will struggle to change anything, so a terrible Level 1 is a genuine early warning. The New World model upgraded Level 1 by adding two reaction dimensions worth capturing: relevance (did this apply to my actual job) and confidence (do I believe I can now do it), which predict downstream behavior far better than a raw satisfaction star. But Level 1 proves only that the experience landed. It says nothing about whether anyone learned, changed, or produced a result. A 4.6 satisfaction score is a Level 1 number, and presenting it as impact is the original sin of L&D reporting.

Level 2: Learning

Level 2 asks: did knowledge, skill, attitude, confidence, or commitment change as a result of the program. The instrument is assessment, ideally a pre-test and post-test so you can attribute the change to the learning rather than to what people already knew. This is where a valid item bank earns its keep: a Level 2 result is only as trustworthy as the validity of the items measuring it, and an AI-drafted question that looks right but tests nothing will produce a confident, false Level 2 gain. Level 2 costs more than Level 1 because it requires designed assessment and, for real rigor, a baseline. It proves people can demonstrate the capability in a test environment. It still does not prove they will use it on the job. The gap between "passed the test on Friday" and "did the task differently on Monday" is the chasm Level 3 exists to cross.

Level 3: Behavior

Level 3 asks the question the business actually cares about: did on-the-job behavior change. Are people doing the task differently, applying the skill, following the procedure, having the conversation. This is the level where learning stops being an event and starts being a performance change, and it is where most L&D measurement quietly stops because Level 3 is expensive and slow. You cannot measure behavior in the classroom; you measure it in the field, weeks or months later, through manager observation, performance data, work-product review, system logs, xAPI statements capturing real actions in real tools, and structured follow-up. The New World model's central insight lives here: behavior change does not survive on a post-test, it survives on required drivers, the reinforcement, coaching, and accountability that keep the new behavior alive after the course ends. A Level 3 measure is the first one a CFO finds genuinely persuasive, because it is the first one about what people do, not what they felt or recalled.

Level 4: Results

Level 4 asks: did the behavior change move a business result. Fewer safety incidents, lower error rates, faster ramp time, higher sales, reduced complaints, lower regrettable attrition. This is the level leadership cares most about and the level hardest to attribute cleanly, because results have many causes and your training is only one of them. Level 4 proves the program connected to an outcome the organization measures in its own reporting, independent of L&D. It costs the most to measure honestly because it requires you to isolate the program's contribution from everything else moving the same number. Done with rigor, Level 4 is the most valuable evidence L&D can produce. Done carelessly, it is the most dangerous, because claiming a results number you cannot defend is how a learning function loses its credibility in a single CFO meeting.

LevelQuestion it answersTypical instrumentCost to measureWhat it does NOT prove
1 ReactionDid they react well, find it relevant?Smile sheet, relevance and confidence itemsLowThat anyone learned or changed
2 LearningDid knowledge or skill change?Pre-test and post-test, validated itemsMediumThat they use it on the job
3 BehaviorDid on-the-job behavior change?Observation, performance data, xAPI, follow-upHighThat it moved a business result
4 ResultsDid a business outcome move?Operational metrics, isolated contributionHighestThe financial return ratio (that is Level 5)
5 ROI (Phillips)Did the financial benefit exceed the cost?Monetized results, benefit-cost ratioHighest plus conversion effortAnything Level 4 did not already establish

Phillips Level 5: ROI, And When It Is Theater

Jack Phillips added a fifth level on top of Kirkpatrick: convert the Level 4 results into money, subtract the fully loaded cost of the program, and express the result as a ratio or a percentage return. The ROI formula is plain: net program benefits divided by program costs, times one hundred, gives a percentage; a 1.5 benefit-cost ratio means a dollar fifty back for every dollar spent. The appeal is obvious. It speaks the CFO's native language and turns "the training helped" into "the training returned 240%." When it is earned, it is the most powerful sentence a learning leader can say.

Here is the honest part the vendors skip. Level 5 is warranted for only a small share of programs, and Phillips himself recommends evaluating ROI on a minority of your portfolio, not everything. The reason is that Level 5 inherits every weakness of Level 4 and adds two more: you must monetize the result (assign a defensible dollar value to a fewer-incidents or faster-ramp number) and you must isolate the training's contribution (separate it from the new manager, the process change, the market shift, and the tool rollout that happened the same quarter). Each of those steps is an assumption, and a chain of optimistic assumptions produces a number that is precise, impressive, and indefensible. That number is theater: an ROI figure calculated to look rigorous that collapses the moment a skeptical CFO pulls one thread of the isolation logic.

A Level 5 ROI you cannot defend under questioning is worse than no ROI at all. It does not prove the program worked. It proves you will report a number you cannot stand behind, and that is the one thing that ends a learning function's credibility with finance.

So when is Level 5 worth it? When three conditions hold together. The program is expensive or high-visibility enough that leadership will demand a financial answer regardless. The Level 4 result is measurable in the organization's own data, so you are monetizing a real number, not a survey. And you can credibly isolate the contribution, ideally with a control or comparison group, a trend line, or at minimum a defensible estimation method you disclose. A safety program that demonstrably cut recordable incidents, a sales-enablement program with a matched control region, a costly leadership program the CEO is personally sponsoring: those earn a Level 5. A monthly compliance refresh measured only by completion does not, and forcing an ROI onto it produces theater. The discipline is not "always calculate ROI." It is "calculate ROI only where you can defend it, and report behavior honestly everywhere else."

There is a counterintuitive professional move buried in this, and it is one of the marks of a mature learning leader: sometimes the strongest thing you can say to a CFO is "we did not calculate ROI on this program, and here is why." That sentence sounds like a confession of weakness and is actually a demonstration of rigor. It tells the CFO that your function knows the difference between a number it can defend and a number it cannot, which is exactly the discernment finance trusts. A learning team that reports an ROI on everything is a team that does not understand isolation, and a CFO who has seen a few budget cycles knows it. A team that reserves ROI for the few programs that genuinely earn it, and reports behavior honestly on the rest, is a team whose numbers can be believed when they do appear. Restraint, here, is credibility. The willingness to not produce a number is what makes your numbers worth something.

Leading Versus Lagging Measures, And Where AI Fits

One more pair of terms separates the professionals from the dashboard-makers. A lagging measure reports an outcome after the fact: last quarter's incident count, this period's error rate, the closed-won revenue. It is the truth, but it arrives too late to act on. A leading measure is an earlier signal that predicts the lagging one: the rate of correct procedure completion in the field, the percentage of reps using the new objection-handling move, the observed coaching frequency. Why you care: leading measures let you intervene before the lagging result is locked in, and they are usually Level 3 behaviors. The mature measurement plan pairs them, watching leading behavior to forecast and protect the lagging result.

This is the precise place AI enters the measurement stack, and it enters as an analyst, never as a judge. AI can read thousands of xAPI statements and surface which behaviors actually preceded the good outcome. It can cluster open-ended Level 3 survey comments into themes a human would take days to code by hand. It can draft the isolation narrative and flag where the data is too thin to support a Level 5 claim. What it cannot do is decide what the evidence proves, monetize a result on its own authority, or sign the number that goes to the board. The iron rule of the whole program holds with full force at the top of the stack: AI assists the analysis, the human verifies the logic, the human owns the claim, and "the model calculated it" is never a defense to a CFO who asks how you isolated the effect. An AI-generated ROI is a draft with a confidence problem until a human has interrogated every assumption inside it.

A Worked Example: The Compliance Refresh, Before And After

Return to the head of learning from the opening and run her program two ways.

Before (the smile-sheet trap). The annual anti-bribery refresh ships to 4,200 employees. Measurement consists of one thing: the LMS completion report and the end-of-course satisfaction survey. The deck to the board reads "94% completion, 4.6 satisfaction, on time and under budget." It is confetti. When a board member asks the obvious question, "are people actually less likely to violate the policy now," the honest answer is that nobody knows, because nothing was measured above Level 1. The program may have worked beautifully or done nothing at all; the measurement cannot tell the difference. Worse, the next year finance proposes cutting the training budget, and L&D has no evidence to defend it, because a completion rate is not a result and everyone in the room knows it.

After (the stack, matched to the stakes). The same refresh is measured deliberately, level by level. Level 1: the smile sheet keeps the two New World items, relevance and confidence, and flags one business unit where relevance scored low, an early warning acted on. Level 2: a short validated scenario assessment, pre and post, shows the percentage correctly identifying a prohibited gift scenario rose from 61% to 89%, a real learning gain, with items a human validated so the gain is trustworthy. Level 3: sixty days out, the team measures behavior, not memory: the rate at which employees correctly logged and escalated gift-and-hospitality entries in the compliance system, an xAPI-and-system-data leading measure, rose meaningfully, and AI helped surface that the rise concentrated where managers had reinforced the behavior, exactly the New World required-driver effect. Level 4: over the year, the count of policy exceptions requiring legal review dropped, a lagging result the legal team already tracks independently of L&D. Level 5: the team does not force an ROI, because they cannot cleanly isolate the training from a parallel policy change, and they say so plainly. The board deck now reads: people can identify violations they could not before, they are behaving differently in the system, the exception count is down, and here is exactly what we can and cannot attribute to the training. That deck survives questions. The confetti deck did not.

The lesson is not that you must always reach Level 4. It is that you measure to the level your claim requires and you never present a lower level as if it were a higher one. A completion rate dressed up as impact is the trap. A behavior measure stated honestly, with its limits named, is the credibility.

Key Takeaways

  • The measurement stack exists to stop organizations from measuring what is easy (showed up, liked it) and calling it proof of impact; each level names the kind of claim you are making.
  • Kirkpatrick's four levels rise in cost and value: Level 1 Reaction (smile sheet), Level 2 Learning (validated pre and post-test), Level 3 Behavior (on-the-job change), Level 4 Results (a business outcome moved).
  • The smile-sheet trap is treating a warm Level 1 satisfaction score as evidence the program changed behavior; it proves only that people enjoyed an hour.
  • Level 3 behavior is the first level a CFO finds genuinely persuasive because it is about what people do, not what they felt or recalled, and the New World model shows it survives on reinforcement and accountability, not on a post-test.
  • Phillips Level 5 ROI converts results to a benefit-cost ratio, but it inherits all of Level 4's attribution weakness and adds monetization and isolation assumptions, so it is warranted for only a small share of programs.
  • An ROI you cannot defend under a CFO's questioning is theater and is worse than no ROI; calculate Level 5 only when the program is high-stakes, the result is in the organization's own data, and you can credibly isolate the contribution.
  • Leading measures (early behavior signals) predict and protect lagging results (after-the-fact outcomes); the mature plan pairs them, and most leading measures are Level 3 behaviors.
  • AI enters the stack as an analyst, not a judge: it reads xAPI at scale, clusters qualitative data, and drafts the isolation narrative, but the human verifies the logic and owns the number, because "the model calculated it" is never a defense to finance.