The four decisions a Level 3 plan is made of — behaviour, rater, timing, threshold — with a worked observation rubric and what the evidence says.
Most Level 3 plans I have been handed are one sentence long. “Manager survey at 90 days.”
That is not a plan. It is a placeholder that has learned to look like one, and it will produce a number that nobody in the room knows how to interpret — including you, because six months earlier nobody decided what the number was supposed to mean.
A Level 3 plan is four decisions. Which behaviour, rated by whom, at what point, against what threshold. Each one changes the result you will eventually report, and three of the four are usually made by accident.
This piece is about writing those four fields. It is not an explanation of the four levels; that exists everywhere and you do not need another one.
Worth being precise, because the sloppiness starts here. Jim and Wendy Kirkpatrick define Level 3 as “the degree to which participants apply what they learned during training when they are back on the job” (An Introduction to The New World Kirkpatrick® Model, p. 6).
Two words in that sentence do the work. Apply — not understand, not intend, not value. And job — not classroom, not simulation, not a follow-up survey about the classroom.
The same document introduces the term that makes the plan writable: critical behaviors, defined as “the few, specific actions, which, if performed consistently on the job, will have the biggest impact on the desired results” (p. 6). Few and specific are the constraint. A Level 3 plan with nine behaviours in it is a Level 3 plan that will be abandoned in month two.
It is also worth knowing that this is the level the field mostly skips. Kirkpatrick Partners’ own introduction reports Level 1 being measured in roughly 78% of training events and Levels 3 and 4 at 25% and 15% — citing a 2009 ASTD study, The Value of Evaluation, which I have not read and am therefore quoting at second hand. Take the exact figures loosely. The direction is not in dispute by anyone who has worked in this field.
Write one behaviour. Two if you must. Each one needs three parts: the action, the frequency, and the context.
The test is whether two people watching the same manager for the same three months would give the same answer. If they would not, you have written a topic, not a behaviour.
Where the behaviour comes from matters as much as how it is worded. It comes out of the diagnosis, and the frame you chose in diagnosis has already decided the unit you can intervene at — an individual, an intact team, or a coalition. If the frame and the unit were chosen carelessly, no wording will save the measurement, because you will be observing the wrong people carefully.
This is the decision most plans make by default, and the default is the trainee. It is the cheapest option, it has the best response rate, and it is not a weaker version of the measurement you wanted. It is a different measurement.
Blume, Ford, Baldwin and Huang’s meta-analysis of transfer research — 89 studies, 93 independent samples, N = 24,493 (Journal of Management, 36(4), 1065–1105, p. 1075) — coded who supplied the transfer measure and then tested whether it changed the answer. It did:
“With motivation, when transfer was measured by the trainee (self), the r was .33, compared with .11 when transfer was measured objectively or by others.” (p. 1082)
The same pattern holds elsewhere in the paper: pre-training self-efficacy correlates with transfer at .29 when the trainee reports the transfer and .14 when someone else does; the work environment relationship moves from .28 self-rated to .20 other-rated (p. 1082).
Read that as a designer rather than as a researcher. The relationships that make your programme look like it worked are consistently larger when the person who attended the programme is the one saying whether it worked. That is not evidence of dishonesty. It is what self-report does.
Who supplies the rating changes the answer
In every pair, the self-rated figure is the larger one. That is not evidence of dishonesty — it is what self-report does. Choosing the trainee as your rater is not a cheaper Level 3; it is a different measurement, and the direction of the difference is known in advance. Source: Blume, Ford, Baldwin & Huang (2010), Journal of Management 36(4), p. 1082.
Nor is the fix “just ask the manager instead,” as though supervisor ratings were the true score. Harris and Schaubroeck’s meta-analysis found “a relatively high correlation between peer and supervisor ratings, but only a moderate correlation between self-supervisor and self-peer ratings” (Personnel Psychology, 41(1), 43–62; ERIC EJ371657). Peers and supervisors largely agree with each other. The person themselves is the outlier. So if you have to pick one rater who is not the participant, a peer is closer to a supervisor’s view than the participant is — which is useful, because peers are often easier to get than a supervisor’s time.
What to write in field 2. Name the rater, and name what the rater’s view is worth. A plan that says “self-report at 90 days, triangulated against a manager rating for a 20% sample” is honest and affordable. A plan that says “survey” has silently chosen the most flattering instrument available and has not told anyone.
The usual reasoning is that you should measure early before the training wears off. The evidence points the other way.
Taylor, Russ-Eft and Chan’s meta-analysis of 117 behaviour-modelling studies found that “although BMT effects on declarative knowledge decayed over time, training effects on skills and job behavior remained stable or even increased” (Journal of Applied Psychology, 90(4), 692–709; ERIC EJ936609). What fades is what people can recite. What holds — and sometimes grows — is what they do.
So measuring Level 3 late is not the risk. Measuring it too early is, and for a plainer reason than decay: the behaviour has not had enough occasions to happen yet. If the behaviour is quarterly, a 30-day measure is asking whether something that occurs four times a year happened in the last month.
Blume et al. add a second reason to be wary of early measurement — “pretraining self-efficacy is influenced by when transfer is measured, that is, r = .32 when taken immediately after training, versus r = .21 when there is a time lag” (p. 1082). Immediate measures run warm.
What to write in field 3. Derive the date from the behaviour’s own frequency, not from a convention. Three occurrences is the minimum worth reporting on. A quarterly behaviour is a nine-month measure or it is nothing, and if the client needs an answer at ninety days then the plan needs a different, more frequent behaviour — which is a design decision, made now, not a reporting problem discovered later.
Decide, in writing, before the programme runs, what result would count as it having worked. Not a target for a proposal. A number you would be willing to see fail.
“At nine months, at least 60% of participating managers will show the documented conversation in place for at least three-quarters of their direct reports, against a baseline of 15%.”
Three properties make that sentence do its job. It has a baseline, so it can move. It has a proportion of people rather than an average score, because an average hides the fact that eight enthusiasts can carry forty. And it is falsifiable, which is the only property that makes anyone believe the result when it lands in your favour.
The baseline is the part that has to be agreed in the room, early, while the client is still describing the problem rather than approving a plan — which is a conversation with its own failure modes. Any measure that cannot be baselined before you start is a story you will be telling from memory.
This is the part that is usually missing, so here is one, complete, for the behaviour above. It is deliberately three rows. A rubric a busy manager will not finish is a rubric that produces no data.
Rated by: the participant’s manager, for each direct report conversation observed or reviewed in the period.
Rater instruction: Rate what you saw or read, not what you believe the person is capable of. If you did not see it, mark “no evidence” — that is a valid and useful answer, and it will not be held against the person.
| 0 — no evidence | 1 — partial | 2 — met | |
|---|---|---|---|
| A gap was named | No specific capability gap appears in the conversation or the record. | A gap is implied or described in general terms (“communication”, “presence”). | A specific gap is named, with an example of when it showed up. |
| The impact was stated | No consequence is connected to the gap. | Consequence stated generically (“it affects the team”). | Consequence stated concretely, with who is affected and how. |
| The person had a route to respond | The conversation is one-directional. | The person responds but no next step is agreed. | The person responds, and an action with a date is agreed and recorded. |
Reported as: the proportion of managers scoring 2 on all three rows for at least three-quarters of their direct reports. Not a mean. Means across a rubric like this are arithmetic on labels and they smuggle a partial into looking like most of a met.
Two things about that rubric are worth stealing regardless of your subject. The “no evidence” option is what protects the data from politeness — without it, every rater rounds up. And the anchors describe what is on the page, not what the rater infers about the person, which is the whole difference between an instrument two raters agree on and a survey of opinions about a person.
Every Level 3 plan depends on things outside the programme. Kirkpatrick’s model has a name for them: required drivers, defined as “processes and systems that reinforce, monitor, encourage, and reward performance of critical behaviors on the job,” with examples including “job aids, coaching, work review, pay-for-performance systems, and recognition for a job well done” (p. 6).
The reason to take that seriously is that it is not only the model’s own assertion. Taylor, Russ-Eft and Chan report that transfer was greatest “when trainees’ superiors were also trained, and when rewards and sanctions were instituted in trainees’ work environments” (ERIC EJ936609). Training the participant’s manager is not a nice-to-have that gets cut in scoping. It is one of the few design decisions with meta-analytic support behind it, and it is usually the cheapest thing in the plan.
So write the conditions as a table with a name against each one, and put it in the design document, not the appendix:
| Condition | Owner | Status at sign-off |
|---|---|---|
| Conversations have a place to be recorded | Systems owner | |
| Managers have the time in the calendar | Operations lead | |
| The participant’s own manager is trained on the same behaviour | Sponsor | |
| The manager’s manager asks about these conversations in one-to-ones | Functional director | |
| Naming a gap carries no career cost | Nobody can own this outright | Stated as a risk |
A condition with no owner will not hold. Writing that down before the programme runs is what separates a measurement plan from a hopeful one, and the last row — the one nobody can own — is the one that most often explains the result.
Worth naming, because designers borrow measurement approaches from the wrong place. Blume et al. separate closed skills, where “the trainee was taught to respond in a particular way according to a set of rules or procedures,” from open skills, where “the trainee had considerable latitude in deciding a course of action.” Their examples are explicit: technical and software training are typical closed skills, and “typical open skills training included leadership and interpersonal skills training” (p. 1076).
Almost everything we design is open. There is no single correct action to observe, which is why the rubric above scores the presence of three elements rather than adherence to a script — and why a compliance-style checklist imported from technical training will mark good judgement as non-compliance.
The last section of a Level 3 plan should be the honest one, and it should be written at design time when it is a limitation, rather than at reporting time when it looks like an excuse.
For the plan above:
I have never seen a client react badly to that list. What clients react badly to is finding out in month nine that the number they were promised was never going to exist.
Four fields, and then two blocks:
The plan, on one page
Four decisions and two blocks. Each of the four changes the number you eventually report, and three of the four are usually made by accident. It fits on a page, and it can be argued with before anything is built.
That is a Level 3 plan. It fits on a page, it can be argued with before anything is built, and it is the difference between reporting a number and reporting a finding.
Try it on a live brief. The Diagnostic Conversation Coach is free, needs no signup, and works from a gap you describe in your own words — including the behaviour it points to and what you would have to measure to know it changed. If it does not sharpen the way you run a discovery call, nothing else here will interest you.