Skip to main content

Connor Wright

Growth at Yoodli

The Kirkpatrick Model: Four Levels of Training Evaluation

September 25, 2026

•

6 min read

What Is the Kirkpatrick Model?

The Kirkpatrick Model is a four-level framework for evaluating training programs, developed by Donald Kirkpatrick in the 1950s and still the most widely used evaluation structure in corporate learning and development today. It breaks the question of whether training worked into four separate, increasingly specific questions: did people like it, did they learn anything, did their behavior change on the job, and did that change produce a measurable business result. Each level builds on the one before it, and each level gets harder to measure than the last. A full explanation of the model, including its history and later additions like the New World Kirkpatrick Model, is available from Kirkpatrick Partners, the organization Donald Kirkpatrick’s family still runs.

For an enablement or L&D team justifying a training budget, or trying to show that a program actually changed something, the four levels give a shared vocabulary for the conversation. Here is what each level looks like in a corporate training context, and where most programs stop measuring before they reach the levels that matter most to a CFO.

Level 1: Reaction

Level 1 measures how participants felt about the training. Did they find it relevant, engaging, and worth their time. Teams collect this almost entirely through post-session surveys: a handful of questions asking a rep to rate the trainer, the content, and their overall satisfaction on a simple scale.

A sales onboarding cohort that rates a new objection-handling workshop 4.6 out of 5 has given a team Level 1 data. That score confirms the room did not dislike the session. It says nothing about whether anyone can handle an objection differently as a result.

Level 2: Learning

Level 2 measures whether participants absorbed the knowledge or skill the training was meant to build. A knowledge check, a quiz, a role-specific certification exam, or a scored practice exercise all fall under this level. The defining feature of Level 2 is timing: it happens at or near the end of the training itself, before the person goes back to real work.

A support team that runs new hires through a product certification exam and requires an 85% pass rate before they take live tickets is measuring Level 2. So is a sales team that has reps complete a scored pitch practice session and requires a passing grade before they run live discovery calls.

Level 3: Behavior

Level 3 asks whether the training changed what people actually do on the job, weeks or months after the session ended. This is where most programs start to lose the thread. Behavior change has to be observed in the field, not in the training room, and observing it at scale takes real effort.

A manager sitting in on a rep’s live discovery calls and noting whether the rep now asks the qualifying questions taught in training is a Level 3 measurement. A QA team scoring a sample of live support calls for whether agents use a new de-escalation script is doing the same thing. Both require someone to watch real work happen, which is why Level 3 measurement in most organizations comes down to a manager’s occasional spot check rather than a consistent data set.

Level 4: Results

Level 4 connects training back to a business outcome: win rate, average handle time, ramp time to full productivity, quota attainment, customer retention. This is the level executives care about most, and the level with the most confounding variables. A quarter’s win rate moves for a dozen reasons that have nothing to do with a training program, so isolating the training’s contribution to a Level 4 number takes real tracking discipline.

A revenue enablement team that can show new reps trained under a redesigned onboarding program reach full quota three weeks faster than reps trained under the old program has a Level 4 result. Getting there requires tracking a cohort over months and controlling for enough other variables that the comparison holds up.

Where Most Training Programs Get Stuck

Level 1 data is cheap, and nearly every training program collects it. Level 2 data is achievable with a quiz or a certification test built into the session. The drop-off happens at Level 3. Most L&D and enablement teams have no practical way to observe hundreds of reps’ actual field behavior on an ongoing basis, so Level 3 gets reduced to anecdote: a manager mentioning in a 1:1 that a rep seems to be using the framework more, which is not data a team can put in front of a board.

Level 4 gets even less attention. Attributing a specific business result to a specific training intervention takes a level of tracking discipline that most programs never build, so the metric tends to stay unmeasured rather than measured poorly. The result is a familiar pattern across L&D teams: a stack of satisfaction scores, a handful of certification pass rates, and a gap where behavior and business impact data should sit.

Why AI Roleplay Practice Produces More Level 2 and Level 3 Data

A traditional workshop generates one data point per person per session: a single quiz score or a single satisfaction rating. AI roleplay practice generates a data point for every repetition, because every practice conversation gets scored against the same rubric.

Yoodli’s AI roleplays give reps a scored conversation to practice against, and because the rubric stays consistent across attempts, a team can track completion rates, scores on a specific skill like discovery questioning or objection handling, and how those scores move across five, ten, or twenty reps. That is Level 2 evaluation with a far larger sample size than a single end-of-session quiz produces.

It also starts to close the Level 3 gap. When a rep’s roleplay scores on a specific objection improve across successive attempts, and a manager can see that trend alongside notes from real calls, the picture gets closer to actual behavior change than an occasional spot check allows. Harness saw a 75% reduction in the time it takes to review sales training submissions once scoring moved from manual review to AI roleplay data, and RingCentral cut call-center certification time by 90% the same way. Neither number is a Level 4 business result by itself, but both show how much more Level 2 and Level 3 data a team can generate once practice itself produces a score instead of a workshop rating. For more on how that kind of data connects to a Level 4 metric enablement leaders already track, see this breakdown of AI sales training and time to productivity.

None of this replaces Level 4 measurement. A team still has to connect practice data to win rates, ramp time, or retention to answer the question a CFO actually asks. A training program that can show a rep’s scored performance improving over ten roleplay attempts, tied to one named skill, gives a manager something concrete to point to in a 1:1 that a satisfaction survey never could.

Using the Four Levels Together

The Kirkpatrick Model works best used as a full stack, with all four levels tracked together rather than one level standing in for the rest. Reaction data shows whether people will sit through the training again. Learning data shows whether the content landed. Behavior data shows whether it changed anything reps actually do. Results data shows whether that change was worth the investment.

Most programs can build the first two levels without much difficulty. The last two require deciding, before a training program launches, which behavior and which business metric it is supposed to move, then building a way to track both over time. Teams that make that decision early, rather than trying to reconstruct it after a program has already run, end up with a Level 4 story to tell when the budget conversation comes around. If you want to see what that looks like with a live cohort of reps, talk to the Yoodli team about setting up scored AI roleplay practice for your next training rollout.

Bring Yoodli to your team