Leah Gerber

2026.09.01

There are three questions about psychiatric diagnosis, not one.

I have a psychiatric diagnosis and I take medication for it. Once a month I answer the same twenty questions so my psychiatrist can see whether anything has moved. I am writing this from inside the system, not above it, and I am not a clinician. Nothing on this page is a reason for anyone to seek, avoid, start, stop or change anything, and I could not tell you what to do with any of it if I wanted to.

What I can do is report where the research actually is, because I went and checked, and what I found was not what the popular critique says.

Where this ends up: whether two clinicians agree, whether a category corresponds to something real, and whether using it helps anyone are three separate questions, and the research answers them at very different levels of confidence. Agreement with real patients is much better than the popular version says. And the categories are how people reach treatment, coverage and accommodation, which is a fact that has to sit in the same breath as everything else here.

The three questions

Almost every argument about diagnostic categories runs three separate questions together and answers the first as if it settled the other two.

Do two clinicians agree? That is a question about reliability, and it has a number attached, usually a coefficient called kappa.

Does the category correspond to something real? That is a question about validity, and it is a different question. Two people can agree perfectly about a category that carves nothing at a natural joint, and disagree about one that does.

Does using the category help anyone? That is a question about usefulness, and it is different again. A category can be fuzzy at the edges and still be the thing that gets someone treatment, coverage, an accommodation at work, or a name for what has been happening to them.

The literature answers these three at very different levels of confidence. Treating one answer as all three is the most common mistake on this topic, and the popular critique makes it constantly.

On agreement, which is better than the story says

Evidence The low agreement numbers that circulate mostly come from studies where clinicians read a written case description and sort it. In one large study of that kind, two doctors reading the same written case landed on the same diagnosis about 55 percent of the time. That is a real finding about how a paragraph gets sorted, and the disagreements were systematic, not random.

Evidence With real patients the picture is different. In a study where two clinicians each assessed the same person, 1,806 patients across 28 centres and 13 countries in local languages, agreement ran from about .45 to .88 depending on the condition, with schizophrenia at .87 and bipolar I at .84 on a scale where 1 is perfect (Reed et al., 2018).

Two cautions travel with that. Several of the authors helped write the guidelines they were testing, which is a stake in a good result. And a kappa is not a percentage of patients. A coefficient of .5 does not mean clinicians disagree about half the people they see; a middling kappa is compatible with two clinicians agreeing about most of them. The sentence "they disagree on three in five patients" is false and it is the most harmful sentence available on this topic, because a reader hears it as a statement about their own file.

On heterogeneity, where a calculation gets mistaken for a finding

Evidence A figure that circulates says there are hundreds of thousands of ways to meet the criteria for one diagnosis. That number is a calculation of how many symptom combinations could satisfy a definition. It is not a count of what anyone has. When researchers looked at what people actually present with, across 155,474 participants in four datasets, the one percent most common presentations covered between a third and four fifths of each sample, while most of the theoretically possible combinations were each reported by fewer than one percent of people, if at all (Spiller et al., 2024).

So the categories are broad, and people cluster inside them far more than the arithmetic implies. Both things are true. Only one of them gets quoted.

On what a label does, which runs both ways at once

Evidence Here the research is clearer and less comfortable for either side. In two online experiments, people read a short description of a person with mild difficulties. When the description was preceded by a sentence saying the person had a diagnosis, readers judged them more suitable for professional help, felt more empathy, and were more supportive of accommodations. The same readers also judged the difficulties as more lasting and less within the person’s control (Altmann et al., 2024). The authors call it a mixed blessing. The raters were members of the public judging a fictional stranger, not clinicians, and not the person themselves.

Evidence A separate registered report found that a label shaped how participants judged a character, and that a clear retraction of the label did not fully undo it (Mickelberg et al., 2024). The authors’ own explanation is that the label persisted because it fit beliefs the readers had already brought with them. The effect lives in the people doing the reading.

Observation And from the other side of the table: in a small user-led study, eight people with experience of psychosis described what diagnosis had done for them. For some it had been a means of access. For others, a cause of disempowerment. Both were in the same room (Pitt et al., 2009). Eight people cannot generalise to anyone, and the authors say so; the point is that the question "is diagnosis good or bad for the person" does not have one answer even at a table of eight.

What the categories are for, stated plainly

Diagnostic categories are how people reach treatment, insurance coverage, workplace accommodation, disability support and research. In many systems the category is the thing the paperwork runs on. Nothing unresolved about how symptoms cluster statistically takes anything away from anyone’s claim on those things, and a page like this one is quotable by an insurer or an employer against a person whose diagnosis is load-bearing. So I am saying it in the same breath: the questions above are questions about research methods. They are not questions about whether anybody’s diagnosis is real.

Where this connects to everything else here

Hypothesis Medicine compares you to a population to diagnose, and to yourself to monitor. Those are different questions and it uses different tools for them, which is more than most fields that judge people manage. A diagnostic category is a between-person instrument doing a between-person job, deciding whether this person qualifies for this. It does that job. The trouble starts when the category gets read as an explanation of a person rather than a decision about their care, and that reading is done by the people around them at least as often as by anyone with a licence.

My monthly twenty questions exist because of a category. They work because they compare me to myself.

Where this stops. Two verification passes, sixteen claims, four survived, and every one of the four needed narrowing before it could go here. The popular critique of psychiatric classification did not survive contact with its own sources; the material that did is narrower and more interesting. Several sources were read at abstract level only and are marked so on the reading list.

What is deliberately not here. Any diagnostic criterion, symptom count, threshold or cut-off. Any figure about treatment response. Any comparison between what the research shows and what a reader’s own clinician told them. Every sentence on this page takes a study or a literature as its subject, on purpose.

What I am not. A clinician. This is a page about measurement, written by a patient who wanted to know where the evidence actually stood.