Leah Gerber

Resources

A working library, with a note on what each one actually shows and where it stops. These come out of five verified research passes run in August 2026. Every pass also failed to cover things, and those gaps are recorded alongside the findings rather than smoothed over. Sources marked free can be read without a university login. For the paywalled ones, a public library card often gets you in, and many libraries hand them out online in a few minutes.

Unfamiliar with a term? Jump to the glossary at the foot of this page.

Within-person variation

  • Fleeson (2001), Journal of Personality and Social Psychology author copy free Shows: sampling people five times a day for weeks, one person’s variation in personality states rivals the differences between people’s averages. Stops: small student samples, self-report, and the result flips when measured against standard questionnaires. His own abstract also shows person averages are almost perfectly stable, which most citations leave out.
  • Fisher, Medaglia & Jeronimus (2018), Proceedings of the National Academy of Sciences free · PMC6142277 Shows: group and individual estimates diverge in spread and in correlations. Stops: the famous 7.85-to-1 ratio it is cited for is contested — the between-person figures appear to be spreads of averages, not of people, and I do not cite that number.
  • Adolf & Fried (2019) and the authors’ reply, Proceedings of the National Academy of Sciences free · PMC6452692, PMC6452707 Shows: the argument narrowing in real time. The reply concedes group-to-individual generalizability lies on a continuum. Stops: it is an exchange of letters, not a settled result.
  • Hamaker & Ryan (2019), “A squared standard error is not a measure of individual differences,” PNAS 116(14):6544, and the authors’ reply, 116(14):6546 free · PMC6452686, PMC6452651 Shows: a letter arguing the 2018 paper’s across-occasion spread is a proxy for the squared standard error of a cross-sectional estimate, and “a SE is not supposed to measure variability across individuals.” The authors replied that their between-group figure “is fundamentally not a representation of within-subject processes over time” and that “it may be inappropriate to use one estimate to represent or draw inferences about the other.” Stops: the exchange concerns the correlation results. No published letter addresses the spread table quoted on this site; applying the same reading there is this site’s inference, checked by three independent readings of the paper.
  • Anvari et al. (2025), Personality and Social Psychology Bulletin free · PMC12044207 Shows: what people believe about their own variability and what repeated measurement finds correlate at only .14 to .30. Stops: seven days of sampling, so nothing about longer horizons.
  • Molenaar (2004), Measurement paywalled Shows: the mathematical argument that group statistics do not apply to individuals by default. Stops: it is an entailment about inference, not a measurement of how much anyone varies. I have only reached the abstract so far.

Sleep and insight

  • Wagner et al. (2004), Nature author copy free · counts at PMC3902672 Shows: 59% of a sleep group found a hidden rule against about 23% awake. Stops: 22 people per cell, one omnibus test at p = .03, no effect size or confidence interval anywhere in the paper.
  • Schonauer et al. (2018), Frontiers in Human Neuroscience free · PMC5834438 Shows: no sleep effect on classic insight puzzles, effect sizes near zero. Stops: a three-hour daytime sleep window rather than a full night.
  • Brodt et al. (2018), Sleep paywalled Shows: time away from a problem helped, and spending it asleep added nothing. The single most relevant paper here. Stops: I have read it through a secondary source and say so until I get the full text.
  • Lacaux et al. (2021), Science Advances free · PMC8654287 Shows: people who drifted into early-stage sleep found the rule far more often. Stops: nobody was assigned to a sleep stage, so it cannot separate the stage causing insight from insight-prone people drifting there.
  • Lowe et al. (2025), PLoS Biology free · PMC12200826 Shows: a preregistered failure to replicate the sleep-onset claim, with an unplanned deeper-sleep effect instead. Stops: the new effect was exploratory and carries the same sorting problem.
  • Cordi & Rasch (2021), Current Opinion in Neurobiology partly free Shows: the sleep-and-memory literature revised downward by one of its own principal authors. Stops: a review, not new data.
  • A 2019 systematic review of the number reduction task free · PMC6779511 Shows: almost all studies using this exact task found sleep helping people state the hidden rule, so the direction of the 2004 result does hold. Stops: a review of a single task, and it does not settle the size of the effect. I have not recorded the authors, so it is cited here by description rather than by name.
  • A 2023 study running the same task across sleep and wake free · PMC10722168 Shows: 18% found the rule after a night of quiet sleep, four people out of 22, against nobody in the wake group. Stops: it was built to test pink noise during sleep and reports these groups alongside that, and 22 per cell is small. Authors not recorded, so cited by description.
  • Debarnot, Rossi, Faraguna, Schwartz & Sebastiani (2017), Neurobiology of Learning and Memory paywalled · abstract free Shows: insight emerged significantly less often after a night of sleep in older adults than in young, and within the older group sleeping was no better than staying awake. Stops: I could not reach the full text, so the sample size and the age-by-condition statistic are unconfirmed. The abstract does not name the task. One study, attributed, not settled.

Idea generation

  • Mullen, Johnson & Salas (1991), Basic and Applied Social Psychology paywalled Shows: pooling 34 tests, people brainstorming face to face listed fewer non-duplicate ideas than the same number working alone and pooling afterwards. Stops: the widely quoted .633 is not the pooled effect, it is the subset where an experimenter was present. The quality half rests on nine tests, not 34. Heterogeneity was large. And the measured thing is a count of ideas produced inside a short lab session, not the worth of any of them. Two critical commentaries run immediately after it in the same issue, which anyone citing it should read.
  • Diehl & Stroebe (1987), Journal of Personality and Social Psychology paywalled Shows: production blocking, simply waiting your turn to speak, is enough on its own to cut the idea count. Stops: Experiment 4 had three groups per condition, and the simulation overshot, since blocked individuals produced fewer ideas than real groups did. No quality measure was taken.

Matching and attraction

  • Joel, Eastwick & Finkel (2017), Psychological Science paywalled · DOI 10.1177/0956797617714580 Shows: machine learning on more than 100 questionnaire measures, collected before 350 people speed-dated, partly predicted how much each person would desire dates in general and be desired in general, and predicted essentially none of the desire for one specific partner over another. Cross-sample correlations for the pair-specific part were -.06 and .02. Measures taken after the dates explained 16 to 29 percent of the same component. Stops: two samples from one research team and campus, young students, four-minute dates, and the authors say it addresses long-term compatibility only obliquely. A 2025 independent re-execution audit found minor errors and let the core conclusions stand.
  • Finkel, Eastwick, Karney, Reis & Sprecher (2012), Psychological Science in the Public Interest paywalled Shows: the major review of online dating, concluding the scientific compatibility claims of matching sites were unsupported by published evidence. Stops: shares two authors with the 2017 study above, so the two are one research voice, not independent confirmation of each other.

Interviews, ratings and memory

  • Spiller et al. (2024), JAMA Psychiatry abstract level Shows: across 155,474 participants in four datasets, the one percent most common symptom presentations covered 33 to 79 percent of each sample, while most theoretically possible combinations were each reported by under one percent. Stops: read at abstract level in our pass; it answers the measured question and says nothing about whether categories are valid.
  • Altmann, Fleischer, Tse & Haslam (2024), PLOS Mental Health free · PMC12798618 Shows: in two online experiments (N = 261 and 684), a diagnostic label attached to a vignette raised judged suitability for treatment, empathy and support for accommodations, and also made difficulties seem more lasting and less controllable. Stops: members of the public rating a fictional stranger; the second study’s effects were small and not significant for every vignette; the senior author originated the concept-creep thesis, though the treatment result runs against that stake. One passage of the discussion is excluded from this site deliberately.
  • Mickelberg, Walker, Ecker & Fay (2024), Acta Psychologica paywalled · registered report Shows: 560 participants; a diagnostic label shaped judgements of a described person and a clear retraction did not fully undo it. Stops: no effect size could be verified; the lab studies misinformation, not psychiatry, and has published nulls in the same paradigm for non-psychiatric information, so the persistence is attributed to prior stigma beliefs, not to corrections failing in general.
  • Pitt, Kilbride, Welford, Nothard & Morrison (2009), Psychiatric Bulletin free Shows: a user-led interview study of eight people with experience of psychosis, for whom diagnosis was variously a means of access and a cause of disempowerment. Stops: eight people; the authors’ stated main limitation is the inability to generalise, and their conclusion is about how a diagnosis is communicated.
  • Reed, Sharan, Rebello et al. (2018), World Psychiatry free · PMC5980512 Shows: two clinicians assessing the same real patient, 1,806 patients, 339 clinicians, 28 centres, 13 countries, in local languages. Agreement ran from about .45 to .88 depending on the condition, with schizophrenia at .87 and bipolar I at .84. Stops: this measures whether two clinicians reach the same category, which is a different question from whether the category corresponds to anything, and a different question again from whether using it helps anyone. Several authors helped write the guidelines they were testing.
  • Flinchum, Kreamer, Rogelberg & Gooty (2023), Organizational Psychology Review paywalled · DOI 10.1177/20413866221097570 Shows: the field’s own agenda-setting review of one-on-one meetings, stating they have not been studied empirically as a focal topic. Stops: a conceptual review with no original data. The widely quoted claim that 47 percent of meetings are one-on-ones comes from its plain-language summary and rests on unpublished internal figures from two companies that sell collaboration tools; do not repeat it as a finding.
  • Sin, Nahrgang & Morgeson (2009), Journal of Applied Psychology paywalled Shows: meta-analytic agreement between a leader and a follower describing the quality of the same relationship, about .37. Stops: a rater-agreement statistic, not an effect of anything on anything. Read at abstract level in our pass.
  • Martin, Guillaume, Thomas, Lee & Epitropaki (2016), Personnel Psychology free author manuscript Shows: leader-member relationship quality correlates with task performance at about .30 overall. When the leader supplies both ratings the figure is .58; when performance is measured by someone else it is .14, on six samples and 722 people. Stops: the paper’s own sample counts do not fully reconcile across sections, no publication-bias analysis was run, and the .14 cell is small. Read in full.
  • McDaniel, Whetzel, Schmidt & Maurer (1994), Journal of Applied Psychology paywalled Shows: the meta-analysis behind the famous interview validity figures, structured .44 and unstructured .33, plus validity of .47 against research criteria and .36 against administrative ones. Stops: those headline figures rest on a range-restriction correction that Sackett et al. 2022 discard as a substantial overestimate. The paper's own uncorrected figures are .31 and .23. Its situational-versus-behavioural ordering cannot be separated from structure, and the authors themselves asked for behaviour-description interviews to be analysed separately.
  • Sackett, Zhang, Berry & Lievens (2022), Journal of Applied Psychology paywalled Shows: a rework of the corrections used across the selection literature, revising structured interviews to .42 and unstructured to .19. Also reports Black-White standardised mean differences of .23 for structured interviews and .32 for unstructured. Stops: the revision is contested by other researchers in the field, and that exchange was not read in our verification pass.
  • Huffcutt, Conway, Roth & Klehe (2004) paywalled Shows: when behaviour-description interviews are analysed separately as McDaniel et al. requested, they come out ahead of situational questions, .51 against .43. Stops: reached through a verification pass rather than read in full here.
  • Salgado & Moscoso (2019), Frontiers in Psychology free · PMC6813221 Shows: agreement between two supervisors rating the same person, .61 when ratings were collected for research and .45 when collected to administer something. Stops: the .45 rests on 18 samples and the .61 on 201, the abstract and results sections give different sample counts, and the paper is partly a defence of supervisory ratings whose own recommendation is to use three or more supervisors rather than abandon them.
  • Smither, London & Reilly (2005), Personnel Psychology free author copy Shows: multisource feedback improves later ratings by about d = .15 overall, and by d = .28 when the programme is developmental against d = .09 when it is administrative. Also tested directly whether adding rater sources helps, and found upward-only and full 360 statistically indistinguishable. Stops: confidence intervals include zero for three of four rater sources, and the meta-analysis is not independent of its authors, whose own study is the largest one included.
  • Scullen, Mount & Goff (2000), Journal of Applied Psychology paywalled · abstract only Shows: in developmental multisource ratings of managers, idiosyncratic rater effects took 62% and 53% of rating variance across two datasets, with the ratee component at 21% and 25%. Stops: the model was fitted to the ratings themselves, so there is no independent measure of performance anywhere in the design, and the widely quoted 21% is a variance component rather than an account of how much a rating reflects real performance. We could not open the full text.
  • Talarico & Rubin (2003) paywalled Shows: memories of hearing about September 11th were no more consistent over time than everyday memories from the same period. Confidence and vividness stayed high while consistency fell like anything else. Stops: 54 people, and consistency is measured against the participant's own earlier account rather than any external record.
  • Hirst et al. (2015) paywalled Shows: across ten years, confidence in a flashbulb memory and its consistency were essentially unrelated, r = .07. Stops: 202 people, same self-comparison limit as above, and it addresses calibration rather than whether any given account is true.

Leads checked for the tripwires page

  • Iyengar & Lepper (2000), Journal of Personality and Social Psychology read in full Shows: among shoppers who stopped, 30% bought at a six-jam table against 3% at a 24-jam table. Stops: 60% of passers-by stopped at the large display against 40% at the small one, the study measured no satisfaction, and a direct supermarket replication found d = 0.02.
  • Scheibehenne, Greifeneder & Todd (2010), Journal of Consumer Research read in full Shows: across 50 experiments and 5,036 participants, the mean effect of assortment size on choice and satisfaction was d = 0.02, indistinguishable from zero. Stops: a later meta-analysis argues for a conditional effect under specific circumstances; that paper was not opened.
  • Schelling (1971), Journal of Mathematical Sociology read in full Shows: coins moved by hand on a 13 by 16 grid, each wanting at least half its neighbours to match, produced neighbourhoods 80 to 90 percent own-colour from a random start. Stops: Schelling wrote that the figures have no quantitative analogue in real cities and ranked organised discrimination and economic sorting as larger causes. The word mild is later commentary.
  • Paluck, Green & Green (2019), Behavioural Public Policy free · read in full Shows: across 27 randomised contact studies with delayed outcomes, d = 0.39, falling to 0.25 for racial, ethnic and religious prejudice and near zero in the three pre-registered studies. Also documents that 95% of the 515 studies in the famous 2006 meta-analysis were non-randomised, and that its authors called Allport’s conditions not essential. Stops: no randomised study of interracial contact in adults over 25 exists.
  • Bouton (2000), Health Psychology, and Jackson & Kestner (2026), Journal of the Experimental Analysis of Behavior abstracts Shows: extinction reduces a response without erasing the original learning, which returns when context changes; replacing the response with another is the same kind of non-erasing learning, and in a 2026 human experiment the old response returned in 17 of 18 people after a pure replacement procedure. Stops: abstract level; no synthesis comparing replacement with extinction was found.
  • Nielsen et al. (2013), PLOS ONE free · read from XML Shows: in resting brain scans of 1,011 people aged 7 to 29, individual networks are lateralised but people do not sort into left-brained or right-brained types; lateralisation of one network did not predict another. Stops: no personality or cognitive data were collected, so it does not directly test the personality version of the claim, and the paper says so.
  • Seidl, Peelen & Kastner (2012), Journal of Neuroscience free · PMC3443854 Shows: in 24 adults searching busy photographs for objects, visual cortex carried less information about a previously relevant category. Stops: the words clutter, declutter, mood and tidy do not appear, participants were no slower or less accurate, and the study is about photographs, not rooms. The most miscited paper in the home-environment literature.
  • Lopez et al. (2024), Journal of Experimental Psychology: General abstract · PMID 37917442 Shows: a preregistered replication of the self-refilling soup bowl with 464 participants found people ate more and did not believe they had, d = 0.45 against 0.84 computed from the 2005 original. Stops: the 2005 paper is not retracted, whatever circulates; one lab sitting, no follow-up.
  • Thomas, Poortinga & Sautkina (2016), PLOS ONE free · read Shows: in 18,053 UK survey respondents, attitudes predicted commute mode most strongly among those who had moved within the past 24 months. Stops: a cross-sectional snapshot of different people who moved at different times, not a process watched over time, and the paper says it cannot account for behaviour before the move.

Person-environment fit

  • Kristof-Brown, Zimmerman & Johnson (2005), Personnel Psychology paywalled · widely available through a library Shows: fit predicts how people feel about work, .44 to .56, and barely predicts what they produce, .07. Stops: correlational throughout, and the effects shrink hard when different people rate each side.
  • Kristof-Brown, Schneider & Su (2023), Personnel Psychology free preprint Shows: the field’s own stocktaking, eighteen years on, revising none of the numbers and calling causal direction unresolved. Stops: only a handful of within-person studies existed to review.
  • Bloom, Moen and colleagues, the STAR trial free · PMC6719311 Shows: a randomized workplace trial that gave employees schedule control. Performance barely moved. Stops: one intervention in one firm, and the outcomes were self-reported.
  • Edwards (1994), Organizational Behavior and Human Decision Processes free copy on the author site Shows: the field’s own methods critique — difference scores confound what they claim to measure. Stops: methodological, so it bounds other findings rather than making its own.

Judging people, official versions

  • National Research Council (2003), The Polygraph and Lie Detection free at nationalacademies.org Shows: the polygraph beats chance on specific incidents and fails for security screening, because base rates defeat it. Also that the states it measures arise without deception. Stops: 2003, and lab-heavy.
  • Bond & DePaulo (2006), Personality and Social Psychology Review paywalled · abstract free Shows: people average 54% at spotting lies in real time, professionals included, and do better with ears than eyes. Stops: lab studies at even base rates, so not an operational estimate.
  • GAO reviews of TSA behavior detection (2010–2017) free at gao.gov Shows: 98% of the sources the agency cited for its behavioral indicators did not provide valid evidence. Stops: reviews of the evidence base, not of whether the program ever caught anyone.
  • Lenzenweger (2015), Journal of Personality Assessment free author copy Shows: the famous 1948 OSS assessment structure does not reproduce with modern methods. Stops: a reanalysis of the same published matrix, and it says nothing either way about whether the assessments predicted field performance. Nothing verified does.

Judgment and sequence

  • Danziger, Levav & Avnaim-Pesso (2011), Proceedings of the National Academy of Sciences free · PMC3084045 Shows: the famous hungry-judges pattern, favorable parole rulings falling across a session. Stops: archival, cases were not randomly ordered, and the reported effect is about fifty times larger than the mechanism proposed to explain it. I report it as contested and do not cite it as causal.
  • Glöckner (2016), Judgment and Decision Making free Shows: the effect-size arithmetic, d of roughly 1.96 against a proposed mechanism whose registered replication sits at 0.04, plus simulations where rational scheduling produces part of the curve. Stops: simulation, without access to the raw data.
  • Chen, Moskowitz & Shue (2016), Quarterly Journal of Economics free working paper Shows: position in a sequence moves real decisions even when order is randomized and the merits are held fixed — loan officers, umpires, asylum judges. Stops: the mechanism is the gambler’s fallacy, not fatigue, and popular coverage merges the two.

Behind each entry sits a claim-by-claim audit trail with verbatim quotes, so anything I say from these can be traced to a page. If a note above overstates what a source shows, tell me and I will correct it in public.

The numbers you will see

These five carry most of the weight in a results section.

Correlation r

How much two things move together, on a scale from minus one to one. Zero means no relationship at all. One means perfectly locked together.

To feel how big it is, square it. That gives you roughly the share of the variation it accounts for. An r of .56 is about 31%. An r of .07 is about 0.5%.

In your reading. Fit and job satisfaction is .56. Fit and job performance is .07. Squaring them is the clearest way to see why that gap matters. One explains a third of what is going on. The other explains almost nothing.

Standard deviation SD

The typical distance from the average. If average sleep is seven hours with a standard deviation of one hour, most people land between six and eight.

Small standard deviation means everyone is bunched together. Large means they are spread out.

Standard error, and the square root of n

This is the one behind the argument about the 7.85 ratio, so it is worth a minute.

Individual scores bounce around a lot. Averages of many scores bounce around much less, because averaging cancels out the extremes. How much less? Divide by the square root of how many things you averaged together.

Average 4 things, the spread halves, because the square root of 4 is 2. Average 100 things, the spread shrinks tenfold. Average 78 things and it shrinks by about 8.8.

In your reading. The famous 7.85 to 1 ratio compared one person’s raw scores against the spread of a set of averages. Those averages each pooled 78 observations, and the square root of 78 is 8.83. That may be most of what the ratio is measuring. Contested, and not settled.

p-value p

The probability of seeing a result at least this big if nothing real were going on. A p of .03 means 3%. By convention anything under .05 gets called significant, which is a bad word for it.

Two things it does not tell you, and the second catches nearly everybody. It says nothing about how big the effect is. And it is not the probability that there is nothing there. It only says how surprising your data would look if there were nothing there. A tiny meaningless effect measured in enough people produces a beautiful p-value.

Red flag. A paper that reports p-values and never reports an effect size is telling you it found something without telling you how much. Wagner 2004 did exactly that.

Effect size

How big, as opposed to whether. The p-value answers is there anything here. The effect size answers is it worth caring about.

They come in several currencies. Correlation is one. So is Cohen’s d, and partial eta-squared, which is the share of variation one factor accounts for.

In your reading. The study that found no sleep effect on insight reported a partial eta-squared of .005. That is 0.5%. It is the numerical way of saying nothing happened.

Confidence interval, and credibility interval CI

A confidence interval is a range for how uncertain one estimate is. If it crosses zero, then no effect at all is still on the table.

A credibility interval is a different animal that gets confused with it. When many studies are pooled, it describes how much the true effect varies from setting to setting rather than how uncertain the average is. A wide one means the effect really does differ by context.

In your reading. Fit and performance had a credibility interval running from -.04 to .44. Because it includes zero, that means in some workplaces the effect is nothing.

N and k

N is how many people. Small n sometimes means how many in one group.

k is how many studies, and you only see it when someone is pooling studies together.

In your reading. Fit and satisfaction rested on k of 65 studies and N of nearly 43,000 people. Fit and performance rested on k of 22 and N of under 6,000. The weak result is also the thinly evidenced one.

How a study is built

The design decides what the study is allowed to conclude. This matters more than the numbers.

Cross-sectional

One snapshot. Many people, measured once. Cheap, common, and it cannot tell you what caused what or how anything changed.

Longitudinal

The same people followed over time. Better for direction, because you can at least see which thing moved first.

Between-person, and within-person

Between-person compares different people to each other. Within-person compares one person to themselves at another moment.

They answer different questions and are constantly confused for one another. This confusion is the subject of your whole project.

In your reading. The fit literature’s own 2023 review said the research is mostly cross-sectional and between-person, and that only a handful of studies were within-person. That single sentence is why nobody can say fit causes anything.

Randomized

The researcher decides by chance who gets what. This is the thing that licenses saying one caused the other. Without it you have a pattern, not a cause.

Post-hoc sorting

Dividing people into groups afterward, based on what happened to them rather than what you assigned. It produces tables that look identical to a real experiment and proves far less.

In your reading. Both sleep-stage studies sorted people by whichever stage they happened to drift into. So they cannot separate that stage causing insight from insight-prone people being the ones who drift there. The giveaway is uneven group sizes, 49, 24 and 14. Nobody designs that.

Preregistration

Writing down what you predict, and how you will test it, before you collect the data. It stops a researcher quietly changing the question once they see the results.

A preregistered result carries more weight than one that was not. And a finding a paper stumbles into afterward is exploratory, which means suggestive rather than established, even inside a preregistered study.

In your reading. Lowe 2025 preregistered a prediction about N1 sleep and it failed. They then found something at N2 that they had not predicted. The first result is strong evidence. The second is a lead.

Replication

Someone running it again. Direct means the same design as closely as possible. Conceptual means the same idea tested a different way.

A finding that has never been directly replicated is a finding that has been agreed with, not confirmed.

Null result

The study found nothing. This is real information and it is much harder to publish, which means the published record leans toward things having worked.

Meta-analysis

Pooling many studies into one estimate. Stronger than any single study, and it inherits every flaw in the studies it pools. If they all measured the thing badly, the meta-analysis measures it badly with more decimal places.

Corrected correlation

A correlation adjusted upward to account for imperfect measurement. Always larger than the raw number.

Not cheating, but it means you cannot compare a corrected number from one paper to an uncorrected one from another.

The traps

These are the ways a result can look real and not be.

Confound

Something else that could explain the result. The two things you measured both move, and a third thing you did not measure is moving them.

In your reading. Your own example. Sleep drops and ideas arrive during mania. Sleep and ideas are correlated, and the state is moving both. Neither one is causing the other.

Common-method bias, or same-source bias

When one person supplies both measurements, their answers agree with each other more than reality warrants. People are consistent about their own opinions.

In your reading. The clearest example you have. Fit and performance is .22 when the same person rates both, and -.02 when different people do. Nearly all of the apparent effect tracked with who was doing the rating.

Moderator

Something that changes the size of an effect. Not a cause of the outcome. A dial on how strongly the cause works.

In your reading. Age moderates the sleep and insight effect. It showed up in young adults and was much weaker in older ones. The effect did not vanish, it depends.

Underpowered

Too few people to detect the effect reliably. If it finds nothing, that is not evidence the effect is absent, only that this study could not see it. If it finds something, the size is probably overstated.

In your reading. 22 people per group, in the study behind every sleep-on-it article.

Who paid for it

Near the end of most papers there is a funding statement and a conflict-of-interest declaration. Read them. They tell you who sponsored the work and whether the authors have a stake in the answer.

Funding does not make a finding false, and independent work can still be wrong. What a sponsor buys is usually subtler than faked results. It shapes which questions get asked, which comparisons get run, and which findings get submitted at all. So the check is not who paid, it is whether the design quietly favors the payer — the comparison chosen, the outcome measured, the follow-up length.

In your reading. TSA's evidence for behavior detection was reviewed by GAO, an auditor with no stake in the program surviving. The agency’s own report to Congress was written to release withheld funding — and still conceded the evidence was missing, which is exactly why it is the most quotable document in the file. An admission against interest is the strongest kind.

Regression to the mean

Extremes drift back toward average on their own. Measure people on a terrible day and they will look better next time whatever you do to them.

It is the reason a self-experiment started because things were going badly will tend to show improvement.

The words specific to your project

Nomothetic, and idiographic

Nomothetic means studying what is true of people in general. Idiographic means studying one person in depth.

Almost all of psychology is the first. Your project is arguing for more of the second.

Ergodicity

A borrowed term from physics. A process is ergodic if what is true across a group at one moment is also true of each individual over time.

The conditions it requires are strict and human behaviour often fails them. The argument is not that group findings are always wrong about you. It is that you cannot assume they hold for you without checking, and almost nobody checks.

Careful. This is a mathematical point about what you are permitted to infer. It is not a measurement of how much people vary. Those get mixed up constantly and the distinction is what keeps your version honest.

Intraclass correlation ICC

How much of the total variation sits between people rather than within them. An ICC of .7 means 70% of the variation is differences between people and 30% is people differing from themselves.

Experience sampling, and EMA

Pinging people several times a day during ordinary life and asking what is happening right now. It beats asking someone to remember last month.

EMA stands for ecological momentary assessment, which is the same thing in a lab coat.

N-of-1

An experiment with one participant, run properly. Not a diary. It means repeating the switch between conditions several times, ideally without knowing which one you are in, so the pattern cannot be wishful thinking.

Heritability

The share of the differences among people, in one population at one time, that tracks with genetic differences among them.

It is not the share of you that came from your genes. There is no such quantity. And the figure moves when the environment changes, without any genes changing.

Getting hold of the papers

Open access, paywalled, preprint, green copy

Open access means free and legal. Paywalled means the journal wants money, usually $30 to $50 for one article.

A preprint is the version before peer review, posted free. Useful, and it may differ from the final paper, so check.

A green copy is the author’s own accepted version, posted on a university site. Legal, free, same content, different formatting.

DOI, PMC, and PMID

A DOI is a permanent address for a paper, starting with 10-point-something. It is the thing to search when a link dies.

PMC and PMID are identifiers in the free US National Library of Medicine database. A PMC number means a free full text exists.

Where to look, in order

Unpaywall first, which you already have. Then PubMed Central, then Semantic Scholar, then the author’s own university page. Then a public library card, which gets you into databases from home.

Then email the author. They own the right to send you a copy and most are pleased to be asked.