Leah Gerber

2026.08.21

Sentences that mean nobody read the study.

Checking research for this site left me with a side effect: I now know, with the receipts, exactly how a couple of dozen famous findings get miscited. That turns out to be more useful than the findings themselves, because each one works as a test you can run on any book, article or post in about ninety seconds.

If a source states the wrong version of one of these, you have learned something specific: whoever wrote it repeated a number without opening the paper it came from. That does not make the rest worthless. It tells you how much weight to put on the parts you cannot check yourself.

At work

“Putting people in work that fits them improves performance.”

Across 172 studies, fit correlates with job satisfaction at .56 and with commitment at .51. Against actual job performance it is .07, and where different people rated each side rather than one person rating both, it is -.02. Fit is well evidenced for how work feels, which is not nothing. Anyone selling the performance claim is quoting the satisfaction evidence and hoping you do not look.

Kristof-Brown, Zimmerman & Johnson, 2005. Full essay: does a job that suits you make you better at it, or just happier?

“Structured interviews predict performance at .44, unstructured at .33.”

Those figures come from a 1994 meta-analysis and were correctly copied from it. They rest on a statistical correction that a 2022 reanalysis discarded as a substantial overestimate. The current estimates are .42 and .19. The practical advice gets stronger: the gap between structured and unstructured roughly doubled.

McDaniel et al., 1994; Sackett et al., 2022. Full essay: the interview numbers everyone quotes have been revised.

“Ask what-would-you-do questions. Situational interviews predict best.”

In the dataset that ordering comes from, every situational interview was coded as structured and most behavioural ones were not, so question type and structure cannot be separated. The authors put tell-me-about-a-time questions in the comparison bin and asked for them to be analysed separately. When someone did, behaviour description came out ahead, .51 against .43.

Huffcutt, Conway, Roth & Klehe, 2004.

“A third of 360 feedback programmes make performance worse.”

The paper this comes from is not a 360 study. Its search terms were feedback and knowledge of results, and its overall average effect was positive. Over a third of its 607 pooled effect sizes fell below zero, which is a fact about a heterogeneous set of effect sizes, many without control groups, not a measured rate at which feedback harmed anyone. The miscitation chain has been traced in print.

Kluger & DeNisi, 1996, via Bracken, Rose & Church, 2016. See Smither, London & Reilly, 2005 for what multisource feedback actually does.

“Performance ratings explain only 21 percent of actual performance.”

The 21 percent is a share of variance for the person being rated, in a model fitted to the ratings themselves. There is no independent measure of performance anywhere in that study. What the same study does show is that 53 to 62 percent of rating variance belonged to the individual rater, which is a different and better-supported sentence.

Scullen, Mount & Goff, 2000.

About people in general

“People vary about eight times more within themselves than between each other.”

The 7.85 to 1 ratio. The between-person figure in it is a spread of group averages, one per moment, with the same people in every average, so the differences between people cancel out exactly. The ratio tracks roughly the square root of how many people went into each average. I wanted this one to be true and it was the best number I had. I got part of the mechanism wrong the first time and the correction is on the essay.

Full essay: do you vary more than you differ from other people?

“Sleeping on a problem boosts your insight.”

The famous 2004 study is about one specific hidden-rule task. The direction survived replication; the size fell from 59 percent to 18. Studies using other kinds of puzzle found nothing. Separate work suggests the help comes from time away from the problem rather than sleep itself.

Full essay: how good is the evidence that sleeping on a problem helps?

“The bigger the brainstorming group, the worse it gets.”

Group size was never manipulated in the research this comes from. It is a correlation across study-level summaries, tangled with era, task and time limit, and partly true by definition, since pooling more separate lists raises the total arithmetically. The underlying finding does hold: people working alone and pooling afterwards list more non-duplicate ideas than the same people brainstorming together.

Mullen, Johnson & Salas, 1991.

“The jam study proved fewer choices sell more and satisfy more.”

Among shoppers who stopped, 30 percent bought at the six-jam table against 3 percent at the 24-jam table. But 60 percent of passers-by stopped at the big display against 40 percent at the small one, the study measured no satisfaction at all, and a meta-analysis of 50 experiments found the average effect of assortment size was zero.

Iyengar & Lepper, 2000; Scheibehenne et al., 2010.

“You cannot extinguish a bad habit, only replace it, and the replacement sticks.”

Extinction works and does not erase the original learning, which returns when the context changes. Neither does replacement. In one human experiment the old response came back in 17 of 18 people after a pure replacement procedure. The first half of the saying is roughly right and the second half is not.

Bouton, 2000; Jackson & Kestner, 2026.

“Brain scans of a thousand people disproved left-brained and right-brained personalities.”

The study found no global left or right pattern in resting brain connectivity across 1,011 people. It collected no personality or cognitive data, so it does not test the personality claim, and the paper says so. Individual brain regions are robustly lateralised. What is unsupported is sorting people into two types on that basis.

Nielsen et al., 2013.

“You remember it vividly, so it must have encoded accurately.”

Memories of hearing about September 11th were no more consistent over time than everyday memories from the same period. Confidence and vividness stayed high while consistency fell like anything else, and across ten years confidence and consistency were essentially unrelated.

Talarico & Rubin, 2003; Hirst et al., 2015. Full essay: a vivid memory is not a more accurate one.

Health and diagnosis

“There are hundreds of thousands of ways to meet the criteria for one diagnosis.”

That is a calculation of how many symptom combinations could satisfy a definition. It is not a count of what anyone has. In measured data across 155,474 participants, the one percent most common presentations covered between a third and four fifths of each sample.

Spiller et al., 2024. Full essay: there are three questions about psychiatric diagnosis, not one.

“Clinicians disagree on three in five patients.”

Nobody counted patients. That sentence is a misreading of an agreement coefficient called kappa, which is not a percentage of people. A middling kappa is compatible with two clinicians agreeing about most of the patients they see. It is the most harmful sentence available on this topic, because a reader hears it as a statement about their own file.

See the diagnosis essay and Reed et al., 2018 for agreement measured on real patients.

Home and city

“Clutter makes your brain work harder.”

This traces to a brain-imaging study of 24 people searching busy photographs for objects. The words clutter, declutter, mood and tidy do not appear in it. Participants were no slower or less accurate. The senior author has said publicly that not all clutter is bad.

Seidl, Peelen & Kastner, 2012.

“Schelling showed that mild preferences produce complete segregation.”

He moved coins by hand on a grid under the rule that each piece wants at least half its neighbours to match, which is not mild. The result was neighbourhoods 80 to 90 percent own-colour, striking but not total. He wrote himself that the figures have no quantitative analogue in real cities and ranked discrimination and economic sorting as larger causes.

Schelling, 1971.

“Contact reduces prejudice, but only under the right conditions.”

The meta-analysis this is attributed to concluded the opposite: that the classic conditions are not essential and are best seen as a facilitating bundle. Separately, 95 percent of its 515 studies were non-randomised, and randomised studies trend toward zero as they get larger.

Paluck, Green & Green, 2019, reporting Pettigrew & Tropp, 2006.

“Moving house is a window for changing habits.”

The evidence is a snapshot survey of 18,053 people who had moved at different times, read as though a process had been watched over time. The paper itself says it cannot account for behaviour before the move.

Thomas, Poortinga & Sautkina, 2016.

The tell

“Psychologists say. Therapists note. Studies show.”

A plural profession with no name, no year and no paper attached. I have five specimens using this identical move on five unrelated topics, and in every case the sentence would have been checkable if a single name had been attached, which is why one never is. When the source is a profession, there is no source.

Specimens in the notebook.

The general form, which is the one worth keeping

All six are the same habit. Ask what the number is a number of.

Here is what that catches. A review of workplace research reports that one factor comes in at 14 percent. That reads like the factor explains 14 percent of whether something worked. It does not. The researchers went through 47 studies and counted every time any factor was mentioned, and this one accounted for 14 percent of the mentions. It is a measure of what people chose to study. It says nothing about what helped. The paper states this plainly and the number gets quoted anyway.

Once you have seen that move you see it everywhere. A list of the accommodations people most commonly asked for, quoted as the list of ones that work. A median cost calculated only among the people who could name a cost. A spread of averages, standing in for a spread of people.

The question is always the same three parts. What was counted, among whom, and what was left out.

And how to read a popular book without inheriting its errors

Do not read it for the findings. Read it for the bibliography, then check the sources yourself. A popular book is a secondary source, and secondary is the layer where overclaiming enters, every time, in the same direction. But the citation list is genuinely valuable, because somebody already did the searching for you.

Where this stops. Every item traces to a verification pass on this project in which somebody opened the source named. Where a full essay exists it is linked; the rest rest on the reading-list entry, which says what was read and what was not. This page grows as passes finish and shrinks if any item is shown to be wrong.

What this is not. A claim that any of these findings are worthless. Several are real in narrower form, and I say so in each case. The tripwire is the overstated version, not the research.

If I am wrong about one of these, I want to know. Each is a specific, checkable claim, which is the point.