The distribution foundry
The densities in the catalog look conjured — here, there, as if a crystal ball were required. It wasn't. Every is one of three bricks: an ordering counted, a sphere measured, or a latent scale integrated out. Each brick below assembles live from raw draws.
Companions: the t & F forge (change of variables), generating functions (the same laws as transforms), and the transforms table (Γ is Mellin's creature: , the Mellin transform of ).
1 · The ordering brick — waits convolve into Gamma
Wait for the -th arrival of a rate- Poisson process. The arrivals before time must land in order — and the volume of that ordered region is the whole mystery:
Given T, the two earlier arrivals are uniform on the triangle t₁ < t₂ < T. Its area is T²/2! — the (k−1)! is a simplex volume.
Wait for the 3rd arrival, 6,000 runs — against the exact Gamma(3, 1) density.
Brick I. continued off the integers — a simplex volume, i.e. a count of orderings. The same brick puts the in Poisson's pmf.
2 · The sphere brick — Gaussians square-sum into χ²
Square-sum standard normals. The Gaussian cloud has no preferred direction, so its integral only sees the radius — and pays a sphere-area factor for each shell:
Two of the ν coordinates. The cloud has no preferred direction — so the integral only sees the radius, sphere by sphere.
‖Z‖² = Z₁² + ⋯ + Z3², 6,000 draws — against the exact χ²(3) density.
Brick II. is the sphere-area function. It haunts everything born from Gaussians — χ², Rayleigh, Maxwell — because polar coordinates always route through it. And in high dimensions that same factor concentrates the Gaussian onto a thin shell — a sphere persistent homology can read: the Gaussian is a sphere →
3 · The ratio brick — Gammas trade to Beta
Take independent , and keep only the share . In the quadrant, is the ray and is the radius — and they separate:
Color = the ratio w = x/(x+y). Every ray from the origin is one color: the ray forgets the radius.
W = X/(X+Y), 6,000 draws — against the exact Beta(2, 5) density.
Brick III. The famous identity is not clever — it is forced: the radius integral factors out as because ray ⟂ radius. Order statistics of uniforms land on the same density by counting who falls left and right.
4 · The mixture finale — a Normal with a random scale is t
Student's is bricks II and III composed: a Normal whose variance is itself random (-distributed). Integrate the latent scale out and the Γ-ratio is the residue:
T = Z / √(V/ν) with V ~ χ²(3) — solid: exact t(3); dashed: N(0,1). Small ν = fat tails; ν → ∞ recovers the normal. Built by hand in the forge; here the point is where the Γ-ratio comes from.
The mixture brick. Every “magic” Γ-ratio — , , Negative Binomial with real — is a latent variable someone integrated out. The constants are receipts, not incantations.
5 · The three bricks
| Brick | Move | Where Γ comes from | Lands on |
|---|---|---|---|
| Ordering | convolve waits | simplex volume | Gamma, Erlang, Poisson's |
| Sphere | square-sum normals | χ², Rayleigh, Maxwell | |
| Mixture / ratio | integrate a latent scale | the -integral = | Beta, t, F, NegBin |
One function, three doors — and all three are the same door from higher up: is the Mellin transform of , the normalizer of the multiplicative world. That is the door the Mellin bridge walks through toward and modular forms: — the factor in the functional equation — is the Mellin transform of the Gaussian.
6 · The missing brick — dependence is its own object
The three bricks forge one-variable laws — and margins never determine the joint. Strip any continuous margin to uniform (the probability integral transform, ) and what remains is the copula: dependence, isolated. Five different joints below — with the identical margin every time:
Elliptical dependence, no tail clustering: λ = 0 for every ρ < 1.
The margin, X = −log(1−U), against the exact Exp(1) density — identical under every preset. Margins are blind to the joint.
The missing brick. The right panel never moves — margins are blind to the joint. And tail dependence is structure a correlation coefficient cannot carry: Gaussian ρ has , while Clayton's shared-Gamma frailty (the mixture brick, again) keeps . Pricing joint defaults with a Gaussian copula is how that lesson got learned in 2008.
7 · The workbench — diagnose a dataset, build a law
Two directions through one machine. Diagnose runs the derivation checklist on a dataset: support → moment ladder → variance scan → moment fit → KS verdict. Build runs the three moves on atoms: tilt (which generates the exponential family), convolve (which walks to the bell), mix (which inflates variance — NegBin and t are mixtures caught in the act). Every curve is a closed form, not a simulation.
- Support. 400 values on ℤ≥0 → count laws (min 0.00, max 9.00)
- Ladder. x̄ = 1.84 · s² = 3.03 · skew g₁ = 1.28 · excess kurtosis g₂ = 1.81
- Variance scan. D = s²/x̄ = 1.64 > 1 — overdispersed → mixture (NegBin)
- Fit (moments). NegBin(r=2.86, p=0.61) vs Poisson(λ=1.84)
- Verify (KS). NegBin: 0.029 · Poisson: 0.090 → NegBin wins
Solid: NegBin(r=2.86, p=0.61) (KS 0.029) · dashed: Poisson(λ=1.84) (KS 0.090).
PP plot for the winner — model CDF vs empirical. On the diagonal = the story holds.
Samples are synthetic and seeded (reproducible); every step runs in your browser. Verdicts read “consistent with”, never “is” — the checklist narrows the shelf, the KS number keeps it honest.
The skill, named. Morris's theorem: a natural exponential family with variance quadratic in the mean must be one of six laws — so the variance scan in the checklist isn't a heuristic, it's reading off a classification. Pick the shelf by support, read , tilt/mix to fit, verify by KS. That is “deriving a distribution on demand.” (v1 uses built-in seeded samples; pasting your own data arrives next.)
8 · Where probability lives — the Fisher–Rao sphere
Write and “probability requirements satisfied” becomes the unit-sphere equation . That is Rao's theorem: the Fisher information metric on distributions is the round sphere metric in these coordinates — so Hellinger distance is the chord and the Bhattacharyya coefficient is the cosine of the angle at the origin. Drag and read the same contours in both charts:
arc = 0.286 chord = 0.143 cosine = 0.990
Arc, chord, and cosine are three readings of the same separation from p to the uniform center (grey). The flat chart distorts them toward the boundary; in √p coordinates the geometry is exactly the round unit sphere — “probability requirements satisfied” is .
The sphere brick, globalized. §2 met the sphere locally (Gaussian in polar coordinates); here the sphere is the space of distributions itself. The same square-root map is the Born rule read backwards — , quantum amplitudes as the unit sphere (the QBism door) — and high-dimensional uniform sphere measure projects to the Gaussian (Poincaré–Borel).
9 · Where the mean dies — compactness vs central tendency
On a compact space the mean must be intrinsic: the Fréchet mean, the minimizer of expected squared distance. Von Mises data on the circle shows its engine — the resultant vector — and its death: as concentration the resultant collapses and the bootstrap spray of mean directions widens from an arc to the whole rim.
resultant length |R̄| = 0.838
circular variance 1 − |R̄| = 0.162
The arrow points at the mean direction; the violet spray is its sampling uncertainty, still a tight arc.
Limits of central tendency, dispersion, and convergence — named honestly. This death is compactness and curvature, not the Hairy Ball: the circle combs perfectly (see §10) yet its mean still dies at maximal dispersion. Near uniformity even the convergence rate degrades — the smeary CLT of Eltzner–Huckemann (2019) replaces with — which is what the widening violet spray is showing you.
10 · Comb the sphere — every flow has a still point
The Hairy Ball theorem is about vector fields, not measures: a continuous tangent field on must vanish, and Poincaré–Hopf says precisely how much — the indices of its zeros always sum to . Comb it, spin it, or hoard the winding in one point: the budget of 2 is conserved.
Poincaré–Hopf: Σ indices = +1 +1 = 2 = χ(S²)
Comb everything toward +e₁: the hair must part somewhere. Two zeros, +1 each.
The gold rings are where the flow stops — no continuous field on S² avoids them, however you comb. The budget of 2 can move and merge, never vanish.
Decision flows. Read as a preference gradient or an excess-demand field: Walras's law makes demand tangent to the price sphere, so “an equilibrium exists” is “this field has a zero” — the same index-theory family as the ball you can't comb. That, and not the mean's death in §9, is where the Hairy Ball genuinely constrains decisions.
Next in this dive: paste your own data into the workbench (v2).