Why the tools get worse exactly when they matter.
Three separate research literatures, on three different instruments, built by different people who mostly do not cite each other, agree on one thing.
Every one of these tools becomes less reliable when a consequence is attached to it.
Evidence Supervisor ratings. A meta-analysis of how much two supervisors agree when rating the same person reports .61 when the ratings were collected for research, and .45 when they were collected to administer something real (Salgado & Moscoso, 2019).
Evidence Interviews. In the large interview meta-analysis, validity against criteria gathered for research came out at .47. Against criteria gathered for administrative use, .36 (McDaniel et al., 1994).
Evidence Feedback. A meta-analysis of multisource feedback found ratings improved afterwards by d = .28 when the programme was for development, and d = .09 when it fed an administrative decision (Smither, London & Reilly, 2005).
Three instruments. Three literatures. Same direction, every time. The measurement is at its best when nothing is riding on it, and degrades at the moment you decide to use it for something.
Why this happens is not mysterious
Nothing exotic is required. When a rating decides a bonus, a promotion or a redundancy, everyone in the process starts optimising against it. The rater knows what the number will do and adjusts. The person being rated knows what is being measured and performs for it. Someone above both of them wants the distribution to look a certain way. The instrument does not change. What it is measuring does.
A number that people are judged on stops describing what it used to describe, because everyone starts playing to the number. That is an old observation. What is new here is seeing it as a measured, repeated pattern in the tools most firms run every year.
What this means practically, and it is not "stop measuring"
Hypothesis If you want to know how someone is doing so you can help them, that is one system. If you need to decide pay, promotion or exits, that is a different system. Running them through the same form on the same day is how you get the degraded version of both. The development conversation becomes a negotiation, and the administrative decision rests on a number nobody was honest into.
Firms combine them because one form is efficient, and because a single record feels fair. The numbers above are what that efficiency costs. And the honest caveats: the supervisor figures rest on 18 samples for the administrative side against 201 for research, the feedback effects are small with intervals that cross zero for most rater sources, and the comparison across the three literatures is mine, not the authors’. Three arrows pointing the same way is a pattern worth naming, not a law.
Where this stops. These are three separate findings pointing the same way, not one study of one thing. Each was collected differently and the comparison across them is mine, not the authors’. The mechanism I have offered is reasoning, not something these papers tested.
What survived checking, and what did not. Salgado and Moscoso’s .45 comes from 18 samples while the .61 comes from 201, and their own recommendation is to use three or more supervisors rather than to abandon ratings. Smither’s headline effects are small and their confidence intervals include zero for three of four rater sources, and the meta-analysis is not independent of its authors, whose own study is the largest one in it.
One number I am not using. The widely quoted interview validity figures of .44 and .33 appear in older sources and rest on a statistical correction that has since been discarded as a substantial overestimate. See the piece on that.