40
Chain
Stimpunks × More Realms · Zine No. 40

Five Sigma

how much evidence is enough before you do something to a child


L★S
Love You Down To Your Star Stuff
open edition · print freely
CERN · 4 July 2012

The word they would not say


For years, physicists at the LHC had a bump in their data where the Higgs boson should be. They would not call it a discovery, and the reason was a number.

The convention is exact, and Louis Lyons — who wrote the paper on where it came from — states it plainly:

"It has become the convention in Particle Physics that in order to claim a discovery of some form of New Physics, the chance of just the background having a statistical fluctuation at least as large as the observed effect is equivalent to the area beyond 5σ in one tail of a normalised Gaussian distribution, or smaller; i.e. the p-value is no larger than 3 × 10⁻⁷."Louis Lyons, "Discovering the Significance of 5σ," arXiv:1310.1284, October 2013

Three in ten million. About one chance in three and a half million that the background alone could have thrown up something this big.

And the discipline is enforced socially, not just statistically. Lyons notes that journals are very reluctant to allow the word 'discovery' to appear in the title, abstract or conclusions of a paper, unless the p-value is that small.

Hold that beside a fact about what was at stake. If the physicists had been wrong, the universe would have been unaffected. No particle would have suffered. The cost of a false claim was a retraction and some embarrassment.

They set the bar at one in three and a half million anyway.

documented  the convention → why 5 and not 3

Where five sigma sits A normal distribution with the region beyond two sigma shaded and labelled p equals 0.05, and the region beyond five sigma marked far out in the tail and labelled three times ten to the minus seven. p = 0.05 3 × 10⁻⁷ one in 3,500,000 before you may say “discovery”
Figure 1 · the tail, and the far tail
physicshuman sci
1 · the Higgs · 5σ
The reason · empirical, not principled

A field that kept being wrong


The 5σ bar was not derived from theory. It was set because the field kept announcing things that turned out not to exist.

"In the past there have been many 'phenomena' that corresponded to 3 or 4σ effects that have gone away when more data were collected. The 5σ criterion is designed to reduce the number of such false claims."Lyons, 2013 — the first argument he lists

Read that as an institutional confession. Three sigma is already odds of about one in a thousand. Four sigma is about one in sixteen thousand. Physics found that even those were producing claims that evaporated, and so raised its own bar against its own enthusiasm.

That is the move worth stealing. Not the specific number — the willingness to notice your field's error rate and respond by making your own life harder.

And it went further. Lyons prints, in his own paper, the objection that the whole thing is unsound:

"Statisticians who hear about the 5σ criterion are very skeptical about it being sensible."Lyons, 2013 — quoted by the man defending the criterion

He does not bury that. He states it, gives the statisticians' reasoning — that probability distributions rarely describe reality accurately that far out in the tails — and then concedes: "Thus the Statisticians' criticism may well be valid."

A field that publishes the case against its own gold standard is doing science. Keep hold of that sentence; the last third of this zine is about what it looks like when a field does not.

documented  false claims → the look-elsewhere effect

Claims that evaporated A row of signals at three and four sigma, most of them fading to nothing as more data arrives, with the threshold moved rightward to five sigma in response. announced at 3–4σ after more data so the field raised its own bar against its own enthusiasm
Figure 2 · an institutional confession
physicshuman sci
2 · why 5, not 3
Dial one · multiplicity

The look-elsewhere effect


Search a wide enough space and something improbable will turn up somewhere. That is not a discovery. That is arithmetic.

Physics has a name and a correction for this. The local p-value asks how unlikely a fluctuation this big is at the mass where you found it. The global p-value asks how unlikely it is anywhere you were looking — and it is always larger, sometimes enormously so.

The correction depends on how honestly you define "anywhere," and Lyons enumerates the choices: any reasonable mass, anywhere in the current analysis, or anywhere in analyses by anyone in the whole collaboration. Then this:

"It is clearly important to decide in advance of the analysis what procedure will be used; and to specify in any publication exactly how the global p-values are defined."Lyons, 2013

That is pre-registration, argued as ordinary practice, in a physics note, in 2013 — while psychology was still discovering it needed the idea.

Now count the search space on the other side. Thousands of researchers, hundreds of candidate interventions, dozens of outcome measures each, decades of trials, and a publication system that prints the positives. The look-elsewhere effect there is not a correction term. It is the dominant feature of the landscape.

And almost nobody applies it.

documented  multiplicity → extraordinary claims

Local versus global A wide noisy spectrum with one tall peak. An arrow marked local asks how unlikely the peak is at that point; an arrow marked global asks how unlikely such a peak is anywhere across the whole spectrum. local: unlikely here global: expected somewhere
Figure 3 · the same peak, two questions
physicshuman sci
3 · look elsewhere
Dial two · degree of surprise

Extraordinary claims


How much evidence you need depends on what you are claiming. Physics builds this in deliberately, and calls it the subconscious Bayes' factor.

The reasoning is straightforward once stated. The evidence in front of you multiplies your prior belief; it does not replace it. So a result that would be convincing for an expected effect is not convincing for an impossible one.

"even if the likelihood ratio favours H₁, we would still prefer H₀ if our prior belief in H₁ was very low. An example would be that, in order to claim that we had discovered energy non-conservation in proton-proton collisions at the LHC, we would require extremely strong evidence from the data because our prior belief in energy non-conservation is very low."Lyons, 2013

He is explicit that this is a Bayesian argument sitting inside a frequentist practice, and defends it anyway:

"this type of reasoning does and should play a role in requiring a high standard of evidence before we reject well-established theories. There is sense to the oft-quoted maxim 'Extraordinary claims require extraordinary evidence'."Lyons, 2013

So hold the question open, because the whole zine turns on it. Is it an ordinary claim or an extraordinary one to say that a standardised intervention, delivered to wildly heterogeneous children, produces a durable improvement in their lives?

Physics would want to know your prior. Ours, after a century of this collection's first chain, is not high.

documented  priors → systematics

Evidence multiplies the prior Two identical bars of evidence applied to two different priors: a high prior reaches the threshold, a very low prior does not. expected effect prior evidence → believed extraordinary claim tiny prior same evidence → not yet the bar moves with the claim so: how surprising is “this will fix the child”?
Figure 4 · dial two
physicshuman sci
4 · degree of surprise
Dial three · the errors you cannot count

The errors you cannot count


Statistical error is the part you can compute from your sample. Systematic error is everything else — the miscalibrated instrument, the selection you did not notice, the assumption baked into the model.

It is much harder to estimate, and physics treats that difficulty as a first-class problem rather than a footnote. Lyons puts a number on the exposure:

"a 5σ claim should be reduced to merely 2.5σ if the systematic uncertainties had been underestimated by a factor of 2; this corresponds to the p-value increasing from 3 × 10⁻⁷ by a dramatic factor of 2 × 10⁴."Lyons, 2013

Read that twice. Being wrong by a factor of two about your unmeasured errors turns a one-in-3.5-million result into a one-in-eighty result. The entire distance between "discovery" and "shrug" is inside a factor of two on the thing you did not measure.

And Lyons is clear that the honest response is to model the uncertainty, not to paper over it with a bigger number:

"It is far better to have a procedure which allows in some way for uncertain systematics than to raise the significance level for all experiments as a way of coping with this."Lyons, 2013

Keep that factor of 2 × 10⁴ in your hand for the next few spreads. It is the answer to a question this zine is going to ask about a literature that does not measure its harms at all — because you cannot underestimate a systematic by a factor of two if you never put a number on it in the first place.

documented  the four dials → Lyons' table

A factor of two A five sigma result collapsing to two point five sigma when systematic uncertainty is doubled, with the p-value rising by twenty thousand times. 3 × 10⁻⁷ systematics × 2 2.5σ ≈ 0.006 p-value up 20,000× from one unmeasured thing
Figure 5 · dial three
physicshuman sci
5 · systematics
The centre of this zine · Lyons, Table 1

The threshold is not one number


Here is the part almost nobody outside physics knows, and it is the reason this zine exists. Lyons does not merely explain 5σ. He argues against using it uniformly — and publishes a table of what different claims should require instead.

His section heading is literally "Why not 5σ?", and it opens: "There are several reasons why it is not sensible to use a uniform criterion of 5σ for all searches for new physics."

The dials are the ones we have just walked: degree of surprise, impact, look-elsewhere effect, systematics. Turn them, and the required significance moves:

after Lyons 2013, Table 1 — selected rows
SearchSurpriseImpactσ
Single topNoLow3
HiggsMediumVery high5
Dark matter (direct)MediumHigh5
SupersymmetryYesVery high7
Gravitational wavesNoHigh7
Neutrinos faster than lightEnormousEnormous>8

Single top gets — predicted by theory, no surprise, low impact, no look-elsewhere problem. Gravitational waves get despite being unsurprising, because the search space is enormous. Superluminal neutrinos get more than 8σ, because the claim would break causality.

And Lyons is careful about the status of his own numbers:

"The spirit of these numbers is to provoke discussion of this issue (rather than being a rigid set of rules), by suggesting a graded set of significance levels for different types of experiments."Lyons, 2013

This is a worked-out theory of how much evidence is enough, calibrated to what you are claiming and what it would cost to be wrong. It exists. It is short, readable and free.

documented  the graded table, as published

Four dials Four labelled dials — surprise, impact, look elsewhere, systematics — feeding into a single output marked required sigma, which ranges from three to more than eight. surprise impact look-elsewhere systematics required σ 358+ a theory of how much is enough short · readable · free
Figure 6 · the instrument, before we borrow it
physicshuman sci
6 · the graded table
The other rail · 1925

Convenient


There is a second number in the world, and almost every claim ever made about a child rests on it.

In Statistical Methods for Research Workers (1925), Ronald Fisher wrote the sentence that set it:

"The value for which P=0.05, or 1 in 20 is 1.96 or nearly 2; it is convenient to take this point as a limit in judging whether a deviation is to be considered significant or not."R. A. Fisher, Statistical Methods for Research Workers, 1925

Two things in that sentence deserve a much longer look than they get.

First: "convenient." Not derived, not calibrated, not justified by an error-rate analysis. A working statistician picking a round number because the arithmetic was tidy and the tables were printable. He said so.

Second: "1.96, or nearly 2." The most consequential threshold in the human sciences is, in the units this zine has been using all along, two sigma.

Physics looked at 3σ and 4σ results, watched them evaporate, and moved up. The human sciences settled at 2σ and stayed for a hundred years.

One field treated its threshold as a hypothesis about its own reliability and kept testing it. The other treated a convenience as a law.the two rails, and they never meet

It is worth saying that Fisher himself did not intend a bright line — he treated significance as one input to judgement, and elsewhere used other levels. The hardening is not his fault. But it hardened, and it hardened into the gate that decides what gets published, funded, mandated, and done.

documented  Fisher's convenience → the hardened rule

Two thresholds, two directions A timeline where the physics threshold moves upward from three sigma to five sigma while the human sciences threshold stays flat at just under two sigma from 1925 to the present. 1925 now physics 1.96σ everything about people one bar moved. one did not. “it is convenient to take this point”
Figure 7 · a hundred years apart
physicshuman sci
the other rail · Fisher 1925
The distance between the rails

About a hundred and seventy thousand


Put the two numbers side by side and divide.

3×10⁻⁷
5σ · to say
a particle exists
0.05
≈2σ · to say
a child must change

The ratio is about 170,000. That is the factor by which one field is more willing to be wrong than the other — and it points the wrong way.

Because now ask what a false positive costs on each side.

Physics gets it wrong: a paper is retracted. A conference is awkward. More data is collected. The particle was never there and never minded either way. The universe is not altered by our having been mistaken about it.

We get it wrong: a child spends forty hours a week for years being trained out of behaviours that were regulating them. A reader is drilled on phonics while the thing that was actually hard goes unexamined. A teenager is told their problem is insufficient grit. And there is no more data to collect, because the childhood already happened.

Physics demands more evidence to say a thing exists than we demand to say a child must be changed. The asymmetry is not subtle, and it runs backwards.the argument of this zine, in one

You cannot un-apply an intervention to a childhood. There is no revision, no fuller dataset, no second run at ages four to nine. Irreversibility is a property of the claim, and it belongs on the dial marked impact.

documented  the two thresholds, and their ratio

The cost of being wrong Two columns. Physics: a wrong claim leads to a retraction and more data, and the particle is unaffected. Us: a wrong claim leads to years of a childhood, with no more data available. physics is wrong retraction more data corrected the particle never minded we are wrong 40 hrs / week for years ages 2–9 no more data to collect the strict standard is on the reversible one
Figure 8 · the asymmetry, running backwards
physicshuman sci
the gap · 170,000×
Our move, not his

Run the dials on a child


Lyons built an instrument for deciding how much evidence a claim needs. Let us point it somewhere he did not, and be honest that we are the ones pointing it.

The claim: this intervention should be delivered to this child, intensively, for years, during the years in which they are becoming a person. Turn the four dials.

the same four columns, one new row
DialSettingWhy
SurpriseHighThat one protocol produces durable gains across radically heterogeneous minds is a strong claim, not a modest one
ImpactEnormousYears of a childhood, and irreversible — there is no second run
Look-elsewhereEnormousThousands of studies, dozens of outcome measures each, decades, and a literature that prints positives
SystematicsEnormousUnblinded delivery, outcomes often rated by the people delivering or paying for the intervention
By Lyons' own table≫ 7σ
What is actually required1.96σ

Every one of those settings is at or above the level Lyons assigns to supersymmetry — a claim he thinks needs 7σ. Nothing here is stretched to get there. Enormous look-elsewhere and enormous impact are simply what the situation is.

Marked contested, and here is precisely what is contested. Lyons was writing about detectors, and would not necessarily endorse this transfer. What we claim is narrower than "physics standards should govern education": it is that surprise, impact, multiplicity and unmeasured error are properties of evidence, not properties of physics, and that a field which has thought hard about all four has something to lend a field that has not. The transfer is arguable. The next spread argues the other side of it, at full strength, because a chain that only argues its own case is the thing this collection exists to refuse.

contested  physics' evidentiary framework → human interventions

All four dials at maximum The same four dials from the earlier figure, now all turned to maximum, feeding an output far past five sigma — beside a much lower marker showing the significance actually required. surprise impact look-elsewhere systematics required used the gap is the whole subject and this transfer is ours, not Lyons'
Figure 9 · borrowed instrument, new target
physicshuman sci
the transfer that never happened
The objection · at full strength

Where this analogy breaks


The previous spread is the most persuasive thing in this zine, which is exactly why it gets a spread arguing against it. Here is the case we would have to answer.

One: physics is asking a different kind of question. A 5σ threshold is for an existence claim — is the particle there at all? Most claims about interventions are about effect size in a system where something is already happening. "Does this exist" and "how much does this help" are not the same question and do not obviously take the same machinery.

Two: the null hypothesis is not neutral. In physics, "no new particle" is a genuine, costless default. With a distressed child there is no null option — doing nothing is also an intervention, with its own risks, and a threshold so high that nothing ever clears it is a decision to do nothing, made by arithmetic and disguised as rigour.

Three: physics has repeatable identical trials. Every proton collision is like every other. No two children are, and demanding particle-physics certainty about heterogeneous humans could be a demand that can never in principle be met — which would make it a rhetorical device, not a standard.

Four: the cost of a false negative is real too. Withholding genuine support from a child who needed it is also a harm, and it is one that a very high evidential bar produces systematically.

All four objections are good. Here is what they do not touch.the answer

None of them argues for 1.96σ. Every one of them is a reason the threshold should be thought about — which is Lyons' actual position, and the opposite of an unexamined convenience inherited from 1925. Objection two is the strongest, and it cuts toward our argument as much as against it: if doing nothing is also an intervention, then it needs an evidence base too, and the ecosystemic answer — change the conditions, not the child — is the option that keeps getting scored as "nothing."

And none of the four explains the finding on the next spread, which is not about thresholds at all.

documented  the objections, and what they leave standing

What the objections reach Four objections drawn as arrows striking a band marked "how high should the bar be" but none of them reaching a lower band marked 1.96 sigma, unexamined since 1925. existence vs effect no neutral null no identical trials false negatives how high should it be? 1.96σ unexamined since 1925 no objection reaches down here they argue for thinking · not for 0.05
Figure 10 · the strongest case against us
physicshuman sci
the objection · argued in place
The finding that is not about thresholds

The column is empty


Everything so far has been an argument about where to set a bar. This is not that. This is about a column that has nothing in it at all.

Bottema-Beutel, Crowley, Sandbank and Woynaroski reviewed 150 reports of group-design, non-pharmacological intervention studies for young Autistic children, and asked a simple question: did the study report whether anyone was harmed?

150
intervention
studies reviewed
11
mentioned adverse
events at all

Eleven. Out of a hundred and fifty. And of the 54 studies that gave reasons for participants withdrawing, ten reported reasons that could be categorised as adverse events, and twelve more were described too vaguely to tell.

Go back to spread six and the factor of 2 × 10⁴. Lyons showed what happens to a 5σ claim when you underestimate your systematic uncertainty by a factor of two.

You cannot underestimate a systematic by a factor of two if you never assigned it a number. An unmeasured harm does not enter the calculation as a large uncertainty. It enters as a zero.why this spread is not an argument about thresholds

A drug trial that did not record adverse events would not be a weak trial; it would not be a trial. In physics, an analysis that omitted its dominant systematic would not be a marginal result; it would be unpublishable. Here it was the overwhelming majority of the literature, and that literature is what "evidence-based" points at.

The same group later put the point in a title: Evidence-b(i)ased practice: selective and inadequate reporting in early childhood autism intervention research.

documented  the reporting gap → what gets called evidence-based

One hundred and fifty studies A grid of 150 marks, of which 11 are highlighted as having mentioned adverse events and 139 are left blank. 150 intervention studies
Figure 11 · gold = reported harms at all
physicshuman sci
the empty column · 150 / 11
Three interventions · one standard

Three interventions, and all of them compulsory for somebody


It is not one literature. The same evidentiary standard is doing the same work in three places at once.

Behavioural intervention. The literature reviewed above — where the harms column is largely blank, and where the outcome measure inherited from 1987 was indistinguishability.

Growth mindset. Macnamara and Burgoyne meta-analysed 63 studies, 97,672 participants, and reported evidence of publication bias, evidence that authors with a financial incentive to find positive effects were more likely to report them, and major threats to internal validity. Their finding on quality is the one to hold: higher-quality studies were less likely to show a benefit. That is the signature of an effect that is not there.

The Science of Reading. Presented as settled science and legislated in state after state, while reading researchers point out that the case rests on a small number of concepts taken from a few simple, dated studies (Seidenberg), and that the movement's certainty outruns its evidence.

Three interventions. None of them near 5σ. All of them mandatory for children who did not consent and cannot opt out.the compulsory part is the part that changes the maths

That last clause is where the evidentiary argument becomes an ethical one. A weak finding published in a journal harms nobody. A weak finding written into a state statute is applied to every child in the state. The look-elsewhere effect made the finding; the legislature made it compulsory; and the child has no standing in either process.

This is what measuring the surface, badly produces. Not merely a wrong answer — a wrong answer with the force of policy behind it, and the word science on the front.

documented  three evidence bases → the same standard

Three interventions against the scale A significance scale from two to eight sigma. Behavioural intervention, growth mindset and the science of reading all sit at the low end near two sigma, far below the five sigma line and the seven sigma marker for what the dials require. physics floor the dials say behavioural intervention growth mindset science of reading all three are compulsory for somebody
Figure 12 · the scale, with everything on it
physicshuman sci
three interventions
The name for it

Scientism is not too much science


There is a word for using the authority of science without the discipline of science, and it is not science.

Scientism is the aesthetics: the p-value, the effect size, the phrase evidence-based, the citation, the graph. What it leaves behind is the practice — pre-registration, blinding, adverse-event reporting, replication, publishing your nulls, printing the case against your own standard, and calibrating your threshold to what you are claiming and what being wrong would cost.

Every device in the first half of this zine is science protecting itself from its own enthusiasm. Scientism keeps the enthusiasm and discards the protection.the distinction the whole collection runs on

And it has a characteristic victim, which is why this argument belongs to us and not only to methodologists. When a field claims scientific authority while ignoring inconvenient findings and lived experience, the people who have the inconvenient experience are not merely disbelieved — they are disbelieved on scientific grounds, which is a much harder thing to answer.

That is epistemic injustice (Fricker), and Robert Chapman has traced how it operates specifically against neurodivergent people: our testimony about our own lives arrives already discounted, because the apparatus has spoken.

We have watched this happen twice already in this collection. In No. 9 the border between disorder and difference is redrawn by committee while the people either side of it stay the same. In No. 39 people asking about the air are told they are catastrophising, by institutions resting on a number nobody had checked.

Same shape every time. The tools of science, pointed at human beings, used to flatten exactly the variation that the science elsewhere insists is real.

contested  the evidentiary gap → epistemic injustice

What gets kept and what gets dropped Two columns. Kept: p-values, effect sizes, the phrase evidence-based, citations. Dropped: pre-registration, blinding, adverse event reporting, replication, published nulls, calibrated thresholds. kept p-values effect sizes “evidence-based” citations graphs dropped pre-registration blinding harm reporting replication published nulls calibrated bar the discipline was the science the rest was the look of it
Figure 13 · the aesthetics, without the practice
physicshuman sci
what it costs · epistemic injustice
The turn

We are asking for more science, not less


Everything in this zine is borrowed from physicists. That is the point, and it is worth being unmistakable about it.

We did not arrive with a philosophical objection to measurement. We arrived with a physicist's own paper, and asked why the standard it describes — calibrate your threshold to how surprising the claim is, how much impact it carries, how wide you searched, and how badly you understand your own errors — stops at the laboratory door.

Eugenics and behaviourism did not fail because they were too scientific. They failed because they took real mathematics — a curve for observational error, a law of effect measured honestly in cats — and applied it to human beings without the discipline that made it work in the first place. That is the whole subject of this collection.

Reclaiming science from scientism is not a retreat from evidence. It is a demand for the parts that got left behind: the harms column, the pre-registration, the null result, the published objection to your own gold standard.what How We Got Here is for

And there is a version of this that is not a complaint but a design. Set the bar by what it costs to be wrong. Where an intervention is reversible, low-stakes and chosen, a modest evidence base is fine. Where it is compulsory, intensive, delivered during childhood and impossible to undo, ask for something closer to what you would want before announcing a particle — and until you have it, change the conditions rather than the child, which is the option that has been scored as "no intervention" this whole time.

That is not anti-scientific. A field that raised its own bar because its own results kept evaporating would recognise it immediately.

Not:that statistics are the enemy. The best material in this zine is a statistics paper. The enemy is a threshold nobody has examined since 1925 doing work nobody chose it for.
Not:that 5σ should govern education. We ran physics' dials on a human question and marked that transfer contested, then spent a spread arguing against ourselves. The claim is that the dials exist and nobody turns them.
Not:that doing nothing is safe. Withholding real support is a harm too. That is an argument for measuring both columns, not for leaving one blank.
Not:a proof. A chain — eleven documented joints and two contested, and no leaps at all. That is unusual here, and it is not tidiness: every gap in this one is published and counted rather than inferred.
The whole chain, complete The full two-rail chain: a physics rail building a graded standard, a human sciences rail running from Fisher to compulsory intervention, and a single dashed rung between them marked the transfer that never happened. A legend gives nine documented joints, one contested, and zero leaps. physics · builds a graded standard Fisher 1925 · → compulsory the transfer that never happened documented · 11 contested · 2 leap · 0 — and that is a finding
Figure 14 · two rails, one rung, no leaps
L★S

Set the bar by what it costs to be wrong. Then look at who is standing under it.

No. 9 The Lines We Drew — the constructed border, and epistemic injustice
No. 38 Eight Tenths of a Second — an instrument with no endpoint
No. 39 Five Microns — a number that answered a different question
No. 40 Five Sigma — how much evidence is enough ← you are here
Reflection

What was decided about you on evidence nobody showed you?

Where in your work does a threshold operate that no one has examined — and who chose it, and when?

If the harms column were compulsory, which practices around you would still be called evidence-based?

What would you need to know before doing something irreversible to a child? Write the number down.

Sources

The spine is one paper, read at the primary: Louis Lyons, "Discovering the Significance of 5σ," arXiv:1310.1284v1, 4 October 2013 (Blackett Lab., Imperial College, and Particle Physics, Oxford). Every quotation on spreads 2–7 is verbatim from it: the definition of the convention and its p-value of 3 × 10⁻⁷; journals' reluctance to permit the word "discovery"; the History argument about 3σ and 4σ effects that "have gone away when more data were collected"; the statisticians' scepticism and Lyons' concession that it "may well be valid"; the look-elsewhere effect and the instruction to "decide in advance of the analysis what procedure will be used"; the subconscious Bayes' factor, the energy non-conservation example, and "Extraordinary claims require extraordinary evidence"; the systematics arithmetic (5σ → 2.5σ on a factor-of-2 underestimate, p rising by 2 × 10⁴); "There are several reasons why it is not sensible to use a uniform criterion of 5σ"; and Table 1, from which the σ values quoted for single top (3), the Higgs (5), direct dark matter (5), SUSY (7), gravitational waves (7) and superluminal neutrinos (>8) are taken, along with his note that their spirit "is to provoke discussion … rather than being a rigid set of rules."

The other threshold. R. A. Fisher, Statistical Methods for Research Workers (1925): "The value for which P=0.05, or 1 in 20 is 1.96 or nearly 2; it is convenient to take this point as a limit in judging whether a deviation is to be considered significant or not." The ~170,000× figure is simply 0.05 ÷ (3 × 10⁻⁷), and is stated as approximate because both thresholds are conventions rather than exact quantities.

The empty column. Kristen Bottema-Beutel, Shannon Crowley, Micheal Sandbank & Tiffany G. Woynaroski, "Adverse event reporting in intervention research for young autistic children," Autism 25(2), 2021, 322–335 — 150 reports of group-design non-pharmacological intervention studies; 11 mentioned adverse events; of 54 reporting reasons for withdrawal, 10 gave reasons categorisable as adverse events and 12 were too vague to classify. The follow-on title quoted is Micheal Sandbank, Kristen Bottema-Beutel, Ya-Cing Syu, Nicolette Caldwell, Jacob I. Feldman & Tiffany Woynaroski, "Evidence-b(i)ased practice: Selective and inadequate reporting in early childhood autism intervention research," Autism, 2024.

Growth mindset. Brooke N. Macnamara & Alexander P. Burgoyne, "Do growth mindset interventions impact students' academic achievement? A systematic review and meta-analysis with recommendations for best practices," Psychological Bulletin (2023; online first 3 November 2022) — 63 studies, N = 97,672; publication bias, authors with financial incentives more likely to report positive effects, major threats to internal validity, and higher-quality studies less likely to show benefit.

Scientism and the surface. On the Problems with Science of Reading at Stimpunks, which supplies both the framing and the phrase measuring the surface, badly (from the Stimpunks glossary entry on behaviorism); Mark S. Seidenberg on the science of reading resting on "a small number of concepts taken from a few simple, dated studies." Epistemic injustice is Miranda Fricker, Epistemic Injustice: Power and the Ethics of Knowing (2007); its application to neurodivergent people here follows Robert Chapman, "Neurodiversity, epistemic injustice, and the good human life," Journal of Social Philosophy (2022). Pathology paradigm, where it is implied, remains Nick Walker's.

Held honestly. Spread 10 applies Lyons' framework to a question he did not write about; it is marked contested and spread 11 argues the case against it rather than hedging in a footnote. The dial settings on spread 10 are our judgements, not measurements, and are presented as such. The Higgs narrative on spread 2 is context around Lyons' own text rather than a separate cited source. This chain contains no leap joints — eleven documented, two contested — which is unusual for the form and is reported rather than smoothed: every gap in it is published and counted, not inferred.