What counts as noise depends on what you are asking.
Signal is what you are trying to detect. Noise is everything else that moves your measurement. On an old radio, the station is the signal and the static is the noise, and turning the volume up makes both louder, which is why it never helped.
Where this ends up: signal and noise are not properties of your data, they are properties of your question, and the same variation between people is scatter to be averaged away or the entire point depending on what you asked. Along the way, what averaging can fix and what it cannot.
That much is a definition. The useful part is that signal and noise are not properties of the data. They are properties of your question. The same numbers change roles depending on what you asked.
The three ways to improve a measurement
There are only three. Measure more times, because errors that push in random directions partly cancel while the thing you are measuring does not. Use a better instrument. Or remove sources of interference, which is what you are doing when you weigh yourself at the same time of day in the same clothes.
Evidence The first one has a rule attached, and it is mathematics rather than a finding about people. The precision of an average improves in proportion to the square root of how many measurements went into it. So halving your error takes four times as many observations, and cutting it to a tenth takes a hundred times as many.
This is a theorem. It cannot fail to replicate. It can only have its assumptions broken, and three of them matter.
The spread has to be finite. There are distributions where averaging a million observations leaves you exactly as uncertain as one. Rare in practice, worth knowing it is possible.
The observations have to be unrelated to each other. This is the one that breaks constantly when measuring people. If your errors share a source, precision does not keep improving. It approaches a floor, and past a point more observations buy you almost nothing.
You have to be sampling from a large pool. Once you have measured nearly everyone, you do better than the rule promises.
The second condition is the one that should worry you
Hypothesis Averaging fixes randomness. It does nothing at all about bias.
If a manager is consistently generous, ten of their ratings are not better than one. You have measured their generosity ten times, very precisely. If every entry in a diary is written by the same person, alone, at the end of the day, in the same format, then thirty entries are not thirty independent looks at that person. They are one way of looking, sampled thirty times, and they will converge confidently on a picture of that person writing in a notebook at night.
Which is a real problem for anyone whose method is repeated self-observation, including mine. The practical answer is to deliberately break the pattern sometimes. Answer out loud into a voice memo instead of writing. Have someone else ask the weekly questions. Write one entry in the morning about the day before. Not for thoroughness. To find out whether what you are seeing survives a change of context, or only exists in a notebook at 11pm.
The part that is actually about this site
Here is where the definition stops being a definition.
Ask what sleep does for people on average, and the fact that people differ from each other is noise. It is scatter, and you average it away to see the effect underneath.
Ask what sleep does for one particular person, and that same variation between people is the entire signal, and the group average is the thing hiding it.
Hypothesis Same table. Same numbers. Opposite roles. And a century of research picked the first question, called the second one error, and averaged it out of existence. Not because anyone was careless, but because the tools were built for the first question and the second one did not have a name.
The most important thing on this page
A shrinking margin of error says that a group average has been pinned down precisely. It says nothing about whether any individual’s account is accurate, and nothing about how much people genuinely differ from one another. Those are separate numbers that move independently.
So when someone answers a person’s account of their own experience with that is within the noise, they have swapped one for the other. A person whose experience sits far from the average is part of the spread, which is a real feature of the world, not an error to be averaged away.
That sentence is the reason this page exists. Everything else on it is arithmetic.
Where this stops. The square root rule and its three conditions are mathematics, verified against textbook statements rather than studies, and I have given them with the conditions attached because the conditions are where the interest is. The claim that shared error is common in measuring people is reasoning, not a measured fact.
What I have deliberately left out. Any number for how strongly one person’s mood correlates from day to day. I went looking and could not verify one, so there is none on this page.
What I would need help with. Anything requiring a projection of how many observations are enough for a particular measurement. The shape of that curve is a theorem; where you sit on it depends on quantities that have to be estimated for each instrument, and that needs a statistician rather than me.