All perspectives Measurement

What gets measured six months later is what actually changed

Transformation is usually declared at the closing meeting — the one moment when every person in the room is incentivised to call it a success.

Written by

Kavita Deshpande

Practice

Assessment & Total Rewards

Reading time

5 min read

Published

January 2026

The closing meeting is a wonderful thing. The deck is good. The feedback scores are strong. The sponsor thanks the team, the consultants thank the sponsor, and everyone leaves believing something has changed.

It is also the least reliable moment in the entire engagement to form a judgement — because it is the one moment when nobody in the room has any incentive to look closely.

Everyone is pulling the same way

Consider who is present. The consulting firm wants a reference and a renewal. The sponsor spent budget and staked credibility on the programme working. The participants enjoyed two days away from their inbox and would like to be invited again. The HR team needs a win to point at in the board pack.

There is no one in the room whose job it is to say we don't know yet.

A closing meeting doesn't measure whether the work succeeded. It measures whether everyone would like it to have.

Satisfaction is not evidence

The standard instrument is the end-of-programme feedback form, and it reliably measures three things: the facilitator's charisma, the quality of the venue, and the lunch.

None of these correlate with behaviour change. Some correlate negatively — the most comfortable programmes are frequently the ones that asked least of anyone. A leadership workshop that leaves people slightly unsettled, having discovered something unflattering, will score worse and work better.

Reaction data is not worthless. It tells you whether the room was with you. It simply cannot tell you what you actually want to know, which is whether anyone does anything differently on the Tuesday three months later, when the consultant is gone and the old pressures have returned.

The discipline has known this for sixty years. Kirkpatrick's four levels — reaction, learning, behaviour, results — are printed in every L&D textbook in the country. Level one is the feedback form. Almost every organisation measures level one. Very few measure level three, and level three is the only one that means anything, because it is the first level at which the organisation, rather than the participant, has changed.

The reason is not ignorance. It is that levels one and two can be collected in the room before everyone disperses, and levels three and four require someone to come back.

The six-month rule

Six months is not arbitrary. It is roughly the point at which every temporary support has been withdrawn: the novelty, the visible executive attention, the consultant in the corridor. What remains at six months is whatever the organisation is genuinely capable of sustaining on its own.

That is the only thing worth calling a result.

And it requires something most engagements skip: a baseline taken before anyone is trained. Without a measurement from before the work started, the six-month number is uninterpretable. You cannot show a shift from a point you never recorded. The most expensive mistake in this discipline costs almost nothing to avoid — you simply have to do it first, when it feels premature.

What to actually measure

Two layers, and both are needed.

The behaviour. Specific, observable, and named before the design begins. Not "improved collaboration" — that is a hope, not a measure. Instead: do cross-functional escalations now get resolved at the level below the MD? Do one-to-ones happen at the stated cadence, and do people report they are about development rather than status? These can be surveyed, sampled and observed.

The business metric the behaviour is supposed to move. Cycle time. Regretted attrition in the first year. Forecast accuracy. Never claim the programme caused the movement — too many things move at once, and attribution in an open system is mostly storytelling. But if the behaviour changed and the metric didn't, that is worth knowing. It usually means the behaviour was the wrong lever, which is a finding rather than a failure.

Be honest about attribution

Here is where measurement practice most often loses its nerve, and overclaims instead.

A programme runs. Six months later regretted attrition in the target cadre is down four points. The temptation — commercially enormous — is to put that on the slide as a result of the work. But in the same six months the market cooled, a competitor froze hiring, and the organisation revised its salary bands. Any of those could account for four points. Possibly all of them.

The honest position is narrower and considerably more useful: the behaviour we set out to change did change, here is the evidence, and the metric it was intended to influence moved in the expected direction alongside several other factors we do not control. That sentence wins fewer awards. It survives scrutiny, which the alternative does not.

The overclaim is not only a credibility problem. It actively destroys the feedback loop. If every programme is recorded as a success, nothing is ever learned, and the same design gets sold into the next organisation with the same confident deck.

A finding that the intervention didn't work is worth more than a claim that it did. It is the only one of the two that improves the next engagement.

Why almost no one does this

Because it is commercially irrational.

A firm that returns at six months is volunteering to discover — in writing, in front of the client — that its work didn't hold. The rational commercial move is to close warmly at the applause and let memory do the rest. Memory is generous; it remembers the good deck and the good lunch.

We do it anyway, for a reason that isn't nobility. Programmes fail in patterns. The manager who was never given time to coach. The framework that contradicted the incentive scheme. The behaviour that was trained but never once inspected. You only see those patterns if you go back and look, and if you never look, you will sell the same failure again next year with complete sincerity.


If you take one thing from this: put the review date in the contract. Not as a courtesy call — as a scheduled obligation with an agreed measure and a baseline taken before the first session.

The firms that resist will tell you it isn't necessary. What they mean is that it isn't comfortable.

Kavita Deshpande · Assessment & Total Rewards · Belvane Partners

Ask us what we found at six months.

Share the challenge. We'll tell you honestly whether — and how — we can help.

Start a conversation