The Rubric Pays You to Leave Your Team
Once a year I open a form with four boxes on it. Each box is a competency the company has decided to grade me on: Influence, Teamwork, Results Focus, Customer Focus. I’m supposed to write a paragraph in each about how I embodied it, attach evidence, and hand it to a manager who will argue it in a room against other managers arguing for other people.
I have filled this form out enough times to notice something that took me a while to say out loud: two of those boxes are at war with each other. And the way we score them — a calibrated stack ranking — doesn’t just fail to referee that war. It picks a side.
The two boxes
Start with how “Influence” actually gets operationalized here, because the word is doing a lot of quiet work. On paper it means “impact beyond your immediate scope.” In practice, in the hallway and in the calibration room, it means build your own brand. Be known outside your team. Drive the cross-org initiative. Give the talk. Be the name that comes up in a room you’re not in.
Now read the box right next to it. “Teamwork” means the opposite kind of thing: subordinate yourself to the group, share credit, unblock the person next to you, do the unglamorous work that makes the team’s number go up rather than yours.
One box rewards you for making your contribution individually legible and attributable. The other rewards you for dissolving your contribution into a collective. You are being asked, on the same form, to both maximize and minimize how much of the work has your name on it. That’s not a balanced scorecard. That’s a contradiction with a staple through it.
Why this isn’t just me being cynical
There’s a genuinely canonical piece of economics sitting under this, and it’s worth knowing the name. In 1991 Bengt Holmström and Paul Milgrom published Multitask Principal-Agent Analyses — the paper that later helped win Holmström a Nobel. Its result, stripped of the math: when a job has several tasks competing for one person’s finite effort, paying harder for the measurable task doesn’t just add effort, it moves effort — off the tasks you can’t measure well and onto the one you can. Push the incentive far enough and the model says the optimal thing to do is stop paying for performance at all, because a flat wage distorts less than a sharp incentive pointed at the wrong dimension.
Translate that into our four boxes. Influence — cross-org, visible, individually attributable — is the measurable task. Teamwork — diffuse, shared, much of it invisible — is the one that’s hard to measure. Holmström and Milgrom’s result says a rational engineer, spending a finite year, will drift their effort toward the legible box and away from the illegible one. Not because they’re a bad teammate. Because that’s what the incentive gradient is for.
And the direction of that gradient isn’t a coincidence either. The stuff that scores Influence points is, by construction, the visible stuff. The stuff that holds a team together — the mentoring, the code review, the “let me pair with you on this,” the doc nobody asked for — is what Tanya Reilly named glue work: essential, and nearly invisible on a promotion packet. The academic version is Babcock, Recalde, Vesterlund and Weingart’s “non-promotable tasks” — work that benefits the organization but does nothing for the individual’s advancement, which everyone would prefer someone else do. The glue is real work. It just doesn’t fit in the box that’s easy to defend.
The calibration room only sees the legible box
Here’s where our specific mechanism turns the screw. We don’t rate people in isolation; we calibrate. A room of managers fits everyone to a distribution and argues relative position. And a manager in that room can only spend one kind of currency on your behalf: legible, attributable evidence they can say out loud.
“Shipped the cross-org migration, everyone knows their name” is ammunition. It survives contact with a skeptical room. “Held the team together, great to work with, unblocked three people” evaporates against a peer’s shiny individually-owned win — not because it’s worth less, but because it’s harder to defend to managers who never saw it.
And they never saw it. That’s the part I keep coming back to. When calibration pools people across teams, the other managers in that room have no visibility into your within-team contribution. They didn’t see your reviews. They didn’t see you stay late to unstick someone. The only thing legible across a team boundary is the work you did across that boundary — your Influence. So the room that sets your rank can, almost by information-theoretic necessity, see your Influence and cannot see your Teamwork.
Sit with the loop that closes: the deciding room can only see out-of-team work, so it rewards out-of-team work, so you do out-of-team work — and your Teamwork score gets set by people structurally incapable of observing your teamwork.
Stack ranking makes Teamwork a curve
Now add the ranking. A forced distribution means that even if every person on a team collaborates flawlessly, the math still requires someone to land at the bottom. Your teammate isn’t only the person you’re graded on collaborating with. They’re your direct competitor for a fixed number of good ratings.
That converts every genuinely helpful act into a personal cost. Unblock them, share the context you worked out, document the thing only you know — and you’ve handed relative position to someone you’re ranked against. The system doesn’t merely forget to reward collaboration. It prices it as a loss.
This is the one place I don’t have to reason from theory, because someone ran the experiment. Loberg, Nüesch and Foege tested forced distribution rating systems in controlled experiments in 2021. Individuals working alone got faster under forced ranking. But teams got slower, task-relevant communication dropped by about a third, and the authors are explicit that the damage “goes beyond sabotage” into people quietly declining to cooperate. That’s not a morale anecdote. That’s the mechanism, measured: rank people against each other and they stop talking to each other.
Hoarding is the rational move, and it doubles as Influence
Put the two together and you get the behavior that made me want to write this down.
Under a personal-brand definition of Influence, your knowledge is your brand. Sharing it — documenting it, teaching it, handing it off — spends down your own valuation. Under a stack ranking, sharing it also lifts a competitor. So both incentives point at the same conclusion: hoard. Don’t write the doc that makes you replaceable. Be the single point of contact. Stay the only person who understands the thing.
Hoarding is a documented response, not a hypothetical — it’s one of the named gaming behaviors Aboubichr and Conway catalogued in a real performance-management system, right alongside “playing safe” and “cooking the books.” When you measure individually-legible output, people optimize for individually-legible output, and hiding what you know is a very effective optimization.
And here’s the cruel part, the part that makes it a trap rather than just a bad trade-off: being the irreplaceable owner of a critical thing is itself high Influence. The hoarding doesn’t cost you on the scorecard. It pays twice — protects your rank, builds your brand — while quietly raising the whole team’s bus factor to one. One action, scored positive on Influence and negative on Teamwork, and a rational optimizer takes the Influence points every time, because those are attributable and the Teamwork ding is diffuse.
The slow-walk is the same move, quieter
Hoarding has a subtler cousin, and once you see it you can’t stop seeing it: the response that never quite comes. The Slack message read at 9:14am and answered — vaguely, partially — at 4:50pm. The code review that sits for three days. The “let me get back to you on that” that never gets back to you.
Organizational psychology has a name for this. Connelly, Zweig, Webster and Trougakos coined “knowledge hiding”, and split it into three flavors. One of them, evasive hiding, is exactly this behavior: giving incomplete information, or promising a fuller answer later that never arrives. It’s distinct from hoarding — hoarding is refusing to share what you have; evasive hiding is the slow, partial, plausibly-deniable non-answer. The kind you can’t be called out for, because technically you replied.
Under a curve, answering a teammate’s question well is a pure gift: it costs you an hour of focus you could have spent on visible work, and it lifts a person you’re ranked against. The scorecard-optimal move isn’t to refuse — refusal is legible and contestable. It’s to let it sit. Not “no.” Just… later. Forever. Loberg’s experiment already showed forced ranking cuts task-relevant communication by about a third; evasive hiding is the individual-level texture of that number.
Here’s the caveat, and it’s the important one: most slow replies are not sabotage. They’re overload, or someone guarding a block of deep-work focus, which is legitimate and often exactly the right call. I’m not claiming your unresponsive colleague is a schemer. I’m claiming something quieter and worse — the rubric makes the ambiguous case, the message you could answer now or answer later, resolve toward “later” every single time, because responsiveness is invisible glue work and the thing you’d do instead is visible and yours. The incentive doesn’t manufacture villains. It tilts a thousand small judgment calls in the same direction, and calls the result a performance distribution.
You grade the people you’re competing with
Now the mechanism that closes the circle, because it’s the most direct one of all. Hoarding and slow-walking are indirect — you reallocate your own effort. Peer review is not indirect. Once a year the company requires you to collect a handful of peer reviews: colleagues writing evaluations that feed the same calibration that ranks you against them. Sit with the shape of that. The company hands you a form that lets you shape the standing of your direct competitors, and makes filling it out mandatory.
Under a curve, that form is a lever, and the incentive is not subtle about which way to pull it. But here’s where the honest version parts company with the cynical one: the response usually isn’t the overt hit piece. An overt hit piece is detectable, it invites retaliation, and you have to keep working with the person on Monday. So the corruption is quieter, and it runs two ways at once.
For your allies, the review inflates — specific, warm, vivid, the kind of endorsement that survives a calibration room, and often reciprocal: I write you a great one, you write me a great one. Aboubichr and Conway found exactly these “collusive alliances” in a real system.
For the person who threatens your rank, the review doesn’t attack. It goes thin. “Solid engineer. Reliable. Works hard.” Faint, generic praise that reads as perfectly fine and quietly withholds the specific, memorable evidence calibration actually runs on. You didn’t lie. You didn’t knife anyone. You just declined to be their advocate — and in a zero-sum ranking, declining to help a competitor is arithmetically the same as pushing them down. This is the sabotage literature arriving through the politest possible channel: Harbring and Irlenbusch showed people harm rivals more as the reward spread widens, and Lazear showed decades ago that competition breeds “industrial politics” — here delivered on a required feedback form, with impeccable manners.
I want to be careful here, because the honest counter-evidence cuts the other way and it matters. The better-documented failure of peer review isn’t the hit job — it’s leniency, everyone rating everyone highly to dodge conflict. A curve doesn’t cleanly cause people to trash each other. What it does is take an already-soft, already-political signal and bend it, so the warmth flows toward allies and the faint praise toward threats. The peer-review data was never clean. The ranking just gives its distortion a direction — and points it straight at the teammates you most need to outrank.
The honest part
I want to be careful here rather than oversell, because the cynical version of this argument is easy and mostly wrong.
The strong claim I can defend is about incentives, not minds. I can’t prove this rubric makes anyone distracted in some measurable, cognitive sense. Nobody has wired up an engineer and watched their attention wander toward their personal brand. What the research supports is the weaker, sturdier claim: the rubric creates a misaligned incentive gradient, and rational people follow gradients. Read the “distraction” as a prediction from the incentive structure, not a measured fact.
The experiments are mostly lab experiments, and the cleanest field studies of ranking come from fruit-pickers and garment plants, not software teams. The transfer to us is an analogy — a well-motivated one, but say the word “analogy” out loud.
And forced ranking genuinely works somewhere. Gerhart and Fang make the real counter-case: individual pay-for-performance raises productivity, partly through incentives and partly by sorting — attracting and keeping stronger people. Forced ranking has real short-term motivational upside. The honest synthesis isn’t “ranking is evil.” It’s that ranking’s benefits show up in individual, low-interdependence work — and get inverted precisely where work is interdependent and hard to attribute, which is the exact description of a software team. Moon, Scullen and Latham make this case directly: the harms grow over time and under interdependence. We are the worst-case cell in that table.
The part where the industry already told us
None of this is new, which is the most damning thing about it. Microsoft ran a stack ranking through the 2000s, and when the company went looking for why a decade of talent produced so little, the ranking system was named as a prime suspect — engineers describing colleagues as competitors to be managed rather than teammates to be helped. Microsoft dropped it in 2013. GE, which invented the modern version, walked it back around 2015. A generation of companies ran this experiment on themselves and quietly filed out of the room.
We kept the room. We put four boxes on the wall, told everyone to fill all four, and then built a scoring system that can only see one of them and mathematically requires half the team to lose. Then we act surprised when the strongest engineers spend the year building a brand somewhere other than the team that needs them.
The rubric isn’t neutral. Three of its four boxes point their legibility gradient away from the team, the fourth is graded by people who can’t see it, and the whole thing rides on a curve that makes your teammate your competitor. It doesn’t have a teamwork problem despite the incentives. It has a teamwork problem because of them — and it’s paying, in real ratings, for exactly the behavior it claims to be measuring against.