Generating functions & moments
One function can carry an entire distribution. Differentiate it and the moments fall out; multiply two of them and you've added independent variables; take its logarithm and you get the cumulants. Push it to the imaginary axis and it becomes the Fourier transform of the density — the object that drives the central limit theorem. This deep dive builds that machine and wires it to the rest of the site.
Cross-references, so nothing is re-derived twice: the combinatorial generating-functions proof and the law of large numbers live in Proofs; the distribution catalog and the expected-value glossary live in the Library; Fourier itself is in change of basis.
1 · Package a distribution into one function
A generating function turns a whole sequence into a single object. For a counting sequence it's the ordinary generating function proved over in Proofs (Fibonacci → Binet's closed form); for a random variable there are three flavors:
The PGF is natural for counts, the MGF for moments, the characteristic function for limit theorems. The mean and variance of every distribution in the Library catalog are the first two moments these functions encode.
2 · Every shape number, by differentiation
Differentiate the MGF at the origin and the raw moments fall straight out — the MGF is the exponential generating function of the moments:
But the raw moments are not yet the names you want. is the mean — but is not the variance; you have to centre it, . The clean route is the cumulant generating function : its derivatives are the named shape numbers — mean and variance are , then standardized skewness and kurtosis are , and the higher cumulants are the corrections Edgeworth and Cornish–Fisher expansions lean on. Hover a rung to see what it measures; then slide the shape toward the normal.
The coefficient on is the raw moment . Pull it out, centre and standardize it, and it becomes a named number — the ladder.
- 1mean2.00
- 2variance2.00
- 3skewness1.414
- 4kurtosis3.000
- 55th cumulant8.485
- 66th cumulant30.000
Read it across: the series hands you raw moments as coefficients; take and its coefficients are the cumulants — the mean, variance, standardized skewness and kurtosis on the ladder. Hover any rung or term to light up its meaning on the shape. Then slide k up: the Gamma melts onto its normal twin and every number past variance falls to zero — the fingerprint of the normal, and the reason it is the one distribution with no shape left to name.
3 · Sums become products — and the bell appears
Here's why these functions earn their keep. For independent :
Adding random variables convolves their densities — a painful integral — but it just multiplies their MGFs (the same convolution-to-product move Fourier provides in change of basis). Take the logarithm and the product becomes a sum, so cumulants add: . Stack copies and that single fact is the whole central limit theorem — the standardized shape numbers shrink and every distribution melts onto the same bell.
Read across: , so adding the variables multiplies the series. But a product mixes orders — the coefficient of picks up the cross terms (highlighted), not just . That mixing is exactly why moments don't add. Take the and the product collapses into a clean sum — the cumulant series of §4, where every order adds on its own and nothing crosses between columns.
4 · Cumulants — the numbers that add
Take the logarithm of the MGF and you get the cumulant generating function ; its coefficients are the cumulants . These are the named numbers from §2 — and they have a property raw moments lack: over independent sums they simply add. Build two distributions and watch the bars stack.
Read down each column: one order, one color. Taking of the MGF turns the messy product from §3 into a clean sum of these two series — so the cumulants simply add, order by order, with nothing leaking between columns. Make a distribution Normal and its skew and kurt chips read 0: the bell adds nothing past variance. That order-by-order cleanliness is the whole reason cumulants — not moments — are the natural coordinates for sums.
5 · The characteristic function is the Fourier transform
Push the MGF to the imaginary axis and it becomes the Fourier transform of the density:
Unlike the MGF, it always exists — for every , because rides the unit circle instead of racing off. It determines the law uniquely by inversion, and pointwise convergence of characteristic functions implies convergence in distribution (Lévy's continuity theorem) — which is exactly the march to the bell you just watched in section 3, and the real engine behind the law of large numbers and the CLT. This is the same inner-product / Fourier geometry as the L² deep dive, now applied to the density itself.
6 · What moments don't pin down
A caution that keeps everything honest: the first two moments are not the distribution. The Datasaurus dozen in the Judgement Lab all share the same mean, variance, and correlation while looking nothing alike — a dinosaur and a star with identical summary statistics.
Even the entire moment sequence can fail to determine a law (the lognormal is famously moment-indeterminate — different distributions, identical moments). That's exactly why the characteristic function matters: it always exists and always determines the distribution. Moments describe; the characteristic function decides.
7 · The calculator — every law, four ways
One engine for all of it. Pick a distribution and read off its four generating functions at once — the PGF, MGF, CGF, and characteristic function — as closed forms and as the color-coded power series of moments and cumulants. The menu is a set of legos; the counterexamples are in it on purpose — pick Cauchy and watch the MGF refuse to exist while the characteristic function stays perfectly smooth.
Four transforms, one law. The PGF marks counts, the MGF dispenses moments, its log — the CGF — gives the cumulants (the coordinates that add), and the characteristic function is the Fourier twin that always exists. Moving between them is change of basis; , . Every cumulant equals λ — the cleanest fingerprint there is. The moments are the messy Bell polynomials of it.
8 · The calculator — combine two laws, X + Y
Now the operations. Pick X and Y and add them: convolving densities is hard, but an operation on the variables is algebra on their transforms — the MGF, PGF, and characteristic function multiply while the CGF adds, so the cumulants add order by order. When the sum lands back in a named family the calculator names it; pick two Cauchys and watch a law with no moments stay closed and stable — the warning that averaging can fail to tame a tail.
Adding independent variables convolves their densities — hard. But it multiplies their MGFs (§3), and taking turns that product into a sum of CGFs (§4), so the cumulants add order by order while the moments mix. That is the calculator's one idea: an operation on the variables is algebra on the transforms, and each of the four transforms (§7) is one coordinate system for the same law. When the sum closes, it lands back in a named family; when X or Y is Cauchy, no MGF exists at all — but the characteristic function never flinches, and stability means averaging never tames it. The squared-distance L² metric on these variables is the next floor down.
9 · Partitions — one product, every way to bucket n
The combinatorial side, and the Ramanujan lego. A partition of is a way to drop balls into buckets; one infinite product is the generating function for every such grouping at once. Cycle through them — the balls regroup, the bucket sizes redraw as a treemap, and the coefficient of counts them. Label the balls and the same picture becomes what is a bucket? — set partitions, counted by Bell numbers.
Every way to drop identical balls into buckets is a partition — and the single product is the generating function for all of them at once: the coefficient of is p(6) = 11. Hardy and Ramanujan found how fast that count grows; the pentagonal-number theorem hides inside the same product. Now label the balls and the buckets become sets — that is what is a bucket?, counted by the Bell number B(6) = 203 instead. Same picture, distinguishable balls. The same grouping is a TDA cluster, an equivalence class, a quotient — buckets all the way down.
Now Ramanujan's miracle, hiding in those counts: lay in a grid and a whole column is divisible by 5 — and by 7, and 11. Watch the column light up:
These are the same partition numbers — and Ramanujan saw that the column is entirely divisible by 5: . Concretely p(4)=5, p(9)=30, p(14)=135, p(19)=490, p(24)=1,575 — every one a multiple of 5. It happens for 5, 7, and 11 only (nothing so clean for 13 or beyond). Why those three? The generating function is almost a modular form — with — and the proof rides the Hecke operators mod you met in the Eigenbook: the partition function is a shadow of a modular form, and these congruences are that shadow's arithmetic. The same machinery, one floor down — from buckets of balls to to Frobenius.
And here is the engine that computes every one of those — Euler's pentagonal number theorem. Multiply out and almost everything cancels; slide the factors in and watch the coefficients lock into a sparse pattern:
The exponents that survive — 1, 2, 5, 7, 12, 15, 22, 26, … — are the generalized pentagonal numbers , and every surviving coefficient is just . An infinite product collapsing to almost all zeros is the marvel. And it is not a curiosity: this product is exactly , so multiplying out and matching coefficients gives a recurrence for the partition numbers,the very engine behind every in the grid above — e.g. , the same 11 that opens Ramanujan's column. The same product, dressed with , is the modular form — the bridge to the Eigenbook.
10 · The Mellin bridge — from q-series to modular forms
The door out. The same coefficients that ride on powers of as a modular form ride on powers of as a Dirichlet series — and the Mellin transform is the bridge between them. This is Hecke's correspondence: the modular law of becomes the functional equation of . Toggle Ramanujan's Δ (a cusp form — entire, symmetric about ) against Eisenstein E₄ (whose has a pole at the abscissa — where Bohr–Landau and the Laurent expansion live).
One set of numbers, τ(n), read two ways: stacked on powers of q it is a modular form (a generating function); stacked on powers of 1/n it is a Dirichlet series. The Mellin transform is the bridge, and the modular law of becomes the analytic law of . Because Δ is a cusp form, its is entire and perfectly symmetric about s = 6 — Hecke's functional equation, drawn. Square the spectrum of the Eigenbook and this is the analytic spine it climbs.
Study these — derivation cards
Reconstruct before you flip. The matching spaced-repetition deck lives in the Obsidian vault.