Judgement Lab
A reference tells you what a method is. This is the other half — instruments for the part a search box can't teach: where sound statistics gets told dishonestly, and where I build my own things to see. Move the sliders; the deception (and the topology) is live.
Where the telling goes wrong
The math below is exactly right. Every lie lives in what gets shown and what gets left out — the gap between a sound result and an honest telling.
Torture the data — p-hacking
Every dot is a test on pure noise — there is no real effect anywhere here. Under the null, p-values are uniform, so some land below 0.05 on their own. With 20 tries the chance of at least one “discovery” is 64%. The statistics is flawless. The lie is reporting the amber dot and never mentioning the 19 grey ones.
Same data, opposite story — Simpson's paradox
Both groups trend upward (positive slopes), yet pull them apart and the overall line tips the other way. Same points, opposite stories. Nobody falsified anything — the lie is choosing whether to show the amber line or the two dashed ones.
Same statistics, different worlds — the Datasaurus
Flip between the six. The scatter changes completely — yet the means and standard deviations don't move; they're identical to two decimals (≈ 54.26, 47.83, 16.76, 26.93) by construction. This is the Datasaurus (Alberto Cairo; Justin Matejka & George Fitzmaurice, 2017), the heir to Anscombe's quartet. The moral every analyst relearns: a summary can hide anything — including a dinosaur. A dashboard that reports only means is a place lies hide. Always plot the data.
Same GPA, different students — the transcript Datasaurus
Started at 1.8, finished at 3.6 — someone who figured it out. The trend is the story, and probably the best bet of the five. A GPA cutoff would rank them even with the student who fell apart.
Five students, one GPA: the average is identical — 2.80 for every one — but the trend (teal) and the consistency are not, and those are the features a real decision needs. The mean is where you start asking, never where you stop. More in The average hides the student.
A 1-credit A is not a 4-credit A — credit weighting
The registrar computes the weighted one — every grade counts in proportion to its credit hours. Drag the credits: the same five grades give a GPA anywhere from ~2.3 to ~3.5 depending only on where the A's sit. Put your A's in the four-credit courses and your low marks in the one-credit lab and you look like a different student — same letters, different weight. An unweighted average treats a Friday lab like a capstone.
How you'd actually decide — a holistic, weighted model
All five students have the same 2.80 GPA, so GPA itself can't rank them — it's constant. A holistic model scores what's left: the trend, the recent form, and the consistency. Move the weights and the order reshuffles — because there is no single right answer. That's the whole point: a real decision is an optimization over many weighted signals, tuned to what you actually care about — not a threshold on one mean. How do you decide? You decide what to weigh.
Something of my own — persistent homology
Not a borrowed demo. This is the topology I actually work in, made playable: a point cloud, a growing scale, and the holes that persist — computed live, Betti numbers and barcode and all.
- Every feature is born at one ε and dies at a larger one. Its bar length — equivalently its height above the diagonal — is its persistence.
- Long features are real structure; short ones are noise from how the points happened to land. Persistence is the signal-to-noise of shape.
- Blue · H₀ = connected pieces (clusters). Violet · H₁ = holes (loops). The faint blue bar that never ends is the whole dataset, finally one piece.
- Try Clusters: three long blue bars — three groups — collapsing into one. Circle: one long violet bar, the hole. Figure-eight: two.
Drag a point, click empty space to add one, double-click to delete — it all recomputes on your cloud. Real persistence, computed live: union-find for b₀, a GF(2) boundary-matrix rank for b₁.
Where this points: knots are made of cords with crossings, and the configuration of a tangle is exactly the kind of point cloud this instrument was built to read. The same diagram has a second, algebraic shadow — the Jones polynomial, computed live next door — so a knot can be measured two ways at once: by its topology here and its polynomial there. That pairing — persistence on knotted cord, against an invariant that can't be fooled by how it's drawn — is the thread I keep pulling.
See also — Contested claims: inference under weak signal and bias, worked through the parapsychology literature.