Leah Gerber

2026.08.21 · Workplace series, part three

The interview numbers everyone quotes have been revised. Almost nobody noticed.

If you have read anything about hiring in the last twenty years you have met these numbers. Structured interviews predict job performance at around .44. Unstructured ones at .33. They appear in books, in training decks, in vendor pitches, and they are usually the strongest quantitative thing in the room.

Where this ends up: both numbers rest on a statistical correction that has since been examined and discarded, and the revised figures make the case for structuring your interviews stronger rather than weaker. A second, related piece of advice turns out to be backwards.

Evidence They come from a 1994 meta-analysis, a study that pools the results of many earlier studies, and they are correctly transcribed from it (McDaniel, Whetzel, Schmidt & Maurer, 1994). The numbers were copied correctly. The problem is that those figures are not raw results, they are estimates adjusted upward by a statistical correction, and that correction has since been examined and rejected.

Evidence A 2022 paper worked through the corrections used across this whole literature and concluded the adjustment applied here was built from 14 of 245 studies and then applied to study designs it did not fit, which they expect to produce a substantial overestimate (Sackett, Zhang, Berry & Lievens, 2022). Their current estimates are .42 for structured interviews and .19 for unstructured. The 1994 paper’s own uncorrected figures were .31 and .23.

The practical advice gets stronger, not weaker

This is the part worth sitting with, because the instinct is to hear a downgrade.

Under the old numbers, structuring your interviews bought you .11. Under the revised ones it buys you .23. The gap roughly doubled. Everything anyone told you about writing the questions in advance, asking every candidate the same ones, and scoring against a rubric is better supported now than it was under the numbers people still quote.

What changed is the ceiling. Under the old figures an unstructured conversation looked like a reasonable instrument on its own. At .19 it does not.

The one that reverses

Suggestion A related claim goes further and says situational questions, the what-would-you-do-in-this-scenario kind, predict better than asking about past behaviour. That ordering does appear in the 1994 paper. It does not survive contact with the paper itself.

Every situational interview in that dataset was coded as structured and most behavioural ones were not, so question type and structure are the same cut of the data and cannot be separated. The authors put behaviour-description questions into the comparison bin, said the choice was arguable, and asked explicitly for them to be analysed separately. When that was done, behaviour description came out ahead, .51 against .43 (Huffcutt, Conway, Roth & Klehe, 2004).

So the popular version of this advice is backwards, and it is backwards because people quoted a table without reading the paragraph underneath it asking them not to.

One number has to travel with all of this. The same 2022 paper reports that structured interviews produce smaller differences between Black and white candidates than unstructured ones, .23 against .32 in standardised terms, and both are far below cognitive tests at .79. Structured interviews look better on that axis too. But .23 is not zero, and a validity number quoted without the group-difference number from the same table is half a finding.

What none of these numbers say

Every figure above is a correlation across a population of hiring decisions. It describes how much better than chance a method ranks candidates in aggregate. It says nothing about whether it will rank this candidate correctly, and the spread around these averages is large enough that at the low end of each distribution the ranking between methods reverses entirely.

An interview also measures something specific: how a person performs in a room somebody else chose, at a time somebody else set, answering questions somebody else wrote. That it predicts later ratings at all is interesting. Treating it as a measurement of the person, rather than of the person in that room on that day, is the mistake.

Where this stops. Both papers were read in full during verification. The 2022 revision is itself contested by other researchers in the field, and I have not read that exchange, so treat the revised figures as the current best estimate rather than a settled one.

The number that must travel with these. The same 2022 paper reports Black-White standardised mean differences of .23 for structured interviews and .32 for unstructured. Structured interviews look better on that axis too, but .23 is not zero, and a validity number quoted without the group-difference number from the same table is half a finding.

More claims like this on the tripwires page.