Backporch
Research · the statistics of contested claims

The statistics of contested claims

The hardest test of an inferential method isn't a clean problem — it's a contested one: a weak signal, strong incentives, and a community that disagrees about what the evidence says. This corner uses the most contested literature I know — parapsychology — not to argue the phenomena are real, but because it is an unusually honest mirror for the questions every empirical field faces.

The stance, stated plainly. I do not claim that psi, PK, or remote viewing are established — the scientific consensus is that they are not, and nothing here overturns it. What I claim is narrower and, I think, more interesting: the statistical machinery people have built to argue about these claims — and to argue back — is real, instructive, and transfers directly to fields nobody doubts.

The unlikely catalyst

In 2011 a respected journal (JPSP) published Daryl Bem's “Feeling the Future,” reporting precognition across nine experiments. It was not the result that mattered most — it was that the paper used standard, accepted methods. If those methods could certify time-reversed causation, the methods were the problem. Bem's paper helped spark the replication crisis: the failures to replicate, and the re-analyses that followed, reshaped how all of psychology and much of science now thinks about evidence. Extraordinary claims did the field a favor.

Four questions extraordinary claims force

Each is a live methodological problem with a real literature. None is specific to psi — they govern clinical trials, genomics, and cosmology just as hard.

Optional stopping · researcher degrees of freedom
When may you stop collecting data?
Publication bias · the file drawer
What never got published?
Prior sensitivity · Bayes factors
How much should weak evidence move you?
Meta-analysis · heterogeneity
Does the pile of studies actually agree?

A case the government actually adjudicated

When the CIA/DIA's remote-viewing program (Stargate) was reviewed in 1995, two statisticians read the same evidence and disagreed in public: Jessica Utts argued the effect sizes were real and consistent; Ray Hyman argued the methodology and replication couldn't bear the weight. Both reports are worth reading precisely because the disagreement is statistical, not rhetorical — it's about controls, replication, and what an effect size means when the mechanism is unknown. That is the genre of argument I want to be excellent at.

Why it belongs in a statistics program

The same four questions decide whether a drug works, whether a gene associates with a disease, whether a faint astronomical signal is a planet or noise. Contested claims are simply the place where weak signal and strong incentive meet most sharply — so the inferential mistakes are largest and easiest to study. The skills transfer wholesale to astrostatistics, biostatistics, and the neuro-cognitive sciences, where I actually want to work.

This sits next to the Judgement Lab (where sound statistics gets told dishonestly) and the calibration thread in QBism as decision theory: three angles on one question — how should a careful person change their mind?

Sources, if you want to read along

If you work on meta-analysis, bias correction, or the foundations of evidence — or you just like arguing about this honestly — I'd like to hear from you: research@backporch.studio.