Blog

Ask Your Team What Really Gets You Promoted Here. Ask Them Separately.

Written by Olaf Bach | 31 Aug 2026

Companies keep announcing AI transformation and keep reporting thin results. A paper from 1975 explains much of that gap without mentioning technology once: organizations reward A while hoping for B, and people, sensibly, do A. Here is what Steven Kerr found, and how a leadership team can read its own reward system honestly.

Kerr's article is an odd classic. There is no experiment in it and very little data. What it has instead is a catalogue: politics, the war in Vietnam, medicine, orphanages, universities, business. In each case he set the behavior an institution said it wanted next to the behavior its rewards actually paid for, and the two pointed in different directions. His conclusion was not that people are cynical. It was that they are alert. Fifty years on, the executive wondering why an AI program produced activity and no results is looking at the same mechanism from the inside.

What did Kerr actually find?

The examples carry the argument.

  • In the Second World War, American soldiers went home when the war was won, which tied a soldier's exit to the organization's goal. In Vietnam, rotation ran on a fixed tour, so what got rewarded was surviving twelve months, whatever happened to the mission.

  • Orphanages were funded according to the number of children in their care while hoping for adoption, which made a full institution a solvent one.

  • Universities announced that teaching mattered and promoted on research and publications;

  • companies asked for long-term growth and paid on short-run sales and earnings.

Two organizational cases sit closer to home.

  • In a large insurance company, underpaid claims produced complaints and were counted, while overpaid claims were accepted in silence and were not. The operating rule staff learned, in Kerr's words, was "when in doubt, pay it out".

  • In a Midwestern manufacturer he surveyed employees on which behaviors the company said it wanted and which it actually rewarded. What earned approval, a substantial share answered, was agreeing with the boss and going along with the majority, while that same management complained about the caution and the yes-manning it was seeing. Nobody there was confused. The reward system had simply been read correctly.

Why do sensible organizations keep doing this?

Kerr gives four reasons and they are still the ones you meet.

  1. First, a fascination with objective criteria: a number that can be defended in a meeting beats a judgment that cannot, even when the number measures the wrong thing.

  2. Second, an overemphasis on visible behavior: output is easy to see, while cooperation, judgment and the decision not to do something are close to invisible, so they go unrewarded by default.

  3. Third, hypocrisy: occasionally the rewarder is getting exactly what he wants and the stated goal is for public consumption.

  4. Fourth, a preference for equity over efficiency: treating everyone the same is comfortable and defensible, and it flattens the very differences an incentive is meant to create.

Only the first two are honest mistakes, and in our experience they are also the common ones. That matters, because if the cause were hypocrisy the work ahead would be a values discussion. It usually isn't. The work is duller and far more tractable: find the visible, defensible metrics your system is actually paying for and accept that they are teaching your people something you never meant to teach. The reward system is the one strategy document nobody skims.

How does a leadership team find out what it actually rewards?

Not by asking itself what it values. That question returns the proclaimed answer, the same trap we described in our post on the culture you proclaim. The honest answer is already written down, just in other documents: the scorecard weights, the last ten promotions, the agenda of the monthly operating review. Here is a loop a leadership team can run on its own live work, in one prepared session and then on a rhythm.

  1. Name the B. One behavior, described concretely enough that you could recognize it happening in a meeting. "More collaboration" is not a B. "A unit hands a customer to another unit when that unit is better placed to serve them" is.
  2. Reconstruct the A from records, not opinions. Take the last ten promotions and the current bonus scorecard and ask what behavior each of them paid for. Take last quarter's operating review and ask what got the first twenty minutes.
  3. Put the newcomer question to each member separately, in writing. A capable, ambitious person joins your unit on Monday and wants to be promoted within two years. What do they actually do? Ask separately, because a group answers with the official version. The individual answers tend to be candid, since nobody is confessing about themselves, and the disagreement between them is the finding.
  4. Take the sharpest contradiction and change one mechanism, not the messaging. A weight in the scorecard. A criterion that now has to be argued in every promotion decision. What the operating review asks for first.
  5. Come back to the same question a quarter later, with the same records and the same newcomer.

In our work with leadership teams, step three is the one that changes the room: people describe their own system with startling accuracy the moment the question is about a hypothetical newcomer rather than about them. We also treat the exercise as recurring rather than diagnostic, because reward systems drift as new metrics get added in good faith.

A decision rule we find useful: if you announced a priority more than two quarters ago and cannot name one thing that has become harder to get promoted for, or one metric that lost weight in the scorecard, you have not changed your incentives. You have changed your communication. Reconstructing the A is a prepared half-day, not an afternoon of goodwill, and changing one mechanism and seeing behavior move takes roughly a quarter. If nothing has moved by then, you changed a mechanism nobody was reading.

Where does this meet AI transformation?

Most AI programs are, in Kerr's terms, elaborate machines for rewarding A. The visible, defensible metric is adoption: seats, logins, an AI usage score by department. The hoped-for B is different work, meaning faster decisions, redesigned handovers, fewer steps. Adoption is easy to count, so adoption is what gets paid for, and we end up where survey after survey of the last two years has put us: a large majority of companies report that they use AI, a much smaller minority report results.

There is a sharper version of the same problem. Consider a competent employee who works out how to do a third of their job in a fraction of the time. In most reward systems, the return on announcing that is more work at the same salary, and the return on staying quiet is an easier quarter. We are, in other words, rewarding the concealment of exactly the productivity we say we are chasing. The same logic governs failed experiments, which is what learning from AI adoption actually runs on and which almost no bonus system pays for. That takes the kind of speak-up climate we looked at in our post on Edmondson's original psychological safety study. If you want the B, somebody has to be visibly better off for having produced it. That is a design question about your incentives, not about your tools.

The counter-argument: rewarding A while hoping for B is a genuine trade-off

Kerr's article is a classic because most readers can immediately think of similar cases from the organization's they have seen themselves. However, Kerr's article is an argument illustrated by cases, not a study: the survey material covers two organizations, is self-report and is fifty years old. It survives because its mechanism is easy to test against your own system, not because it was ever demonstrated in a trial.

It also drew a serious counter-argument. Boettger and Greer, writing in Organization Science in 1994 under the title "On the wisdom of rewarding A while hoping for B", argued that organizations routinely carry goals that genuinely conflict, so a system that pays for one while leadership hopes for the other is sometimes an honest reflection of a real trade-off rather than a design fault. Where that is the case, the work is to help people handle the inconsistency, not to pretend it away. Nor will incentives carry a change on their own: tighten a metric hard enough and you get gaming rather than performance, which is the aim-rather-than-volume problem we wrote about in our post on feedback interventions.

Which brings us back to his catalogue. Kerr never suggested the institutions in it were badly run. Most were run by capable people who had, at some defensible moment, picked a criterion they could stand behind in a meeting. Revisiting the argument twenty years later, he remarked on how little had changed, which says something about the pull toward a measurable A. That is why this deserves a session with your own scorecard on the table rather than a nod of recognition. Your people have already read it.

Want to see how your leadership team could run this audit on its own scorecard? Schedule a short demo.