Paradigms
Before a method, a stance — but not a rivalry to crown a winner from. That's the trap the word versus sets. Every framing below is an answer to one of three questions: what structure does the object carry, in what representation do I express it, and under what stance do I count something as true? Nodes in a lattice, not a list — connected by translations, each one fingerprinted by what it leaves unchanged.
A paradigm taking shape: Architectures — the math of generative AI (attention as kernel regression, diffusion, the transformer block), each a live lab.
Part I — the objectStructure: what does it carry?5
Start with the tower. A bare set knows nothing; add structure a layer at a time — measurable, topological, smooth, metric — and each layer is defined, in Klein's sense, by the transformations that leave it alone. “vs” here means a step up (enrich) or down (forget) the tower. Climb only as high as your question needs, and claim only the invariants your rung can support.
Worked as derivations: what is a bucket? (a forgetful step down the tower), what is a process? (a map whose fibers are buckets), what is a shape made of? (cells in, holes out), and does order matter? (the group of symmetries).
Topological vs smooth vs Riemannian manifoldseach rung adds one structure and one invariance — homeomorphism → diffeomorphism → isometry.
The archetypal tower. The underlying space never changes; each rung equips it with one more layer of structure, and that structure is exactly what makes a new class of operations well-defined — each pinned down by the group it must stay invariant under (homeomorphism → diffeomorphism → isometry). Read every row as a capability the structure unlocks, not as points of the space.
| Rung | Topological | Smooth | Riemannian |
|---|---|---|---|
| Structure added | an atlas of charts to , glued by homeomorphisms | + smooth () transition maps | + a metric tensor — an inner product on each |
| Now well-defined | nearness, continuity, dimension, compactness | tangent vectors, derivatives, differential forms | length, angle, distance, curvature |
| Not yet | differentiation — charts agree only continuously | length & angle — no inner product on yet | — top rung here |
| Integration | none intrinsic | integrate top-degree forms (given an orientation); no canonical measure | a canonical volume measure from |
| Invariant under | homeomorphism — Betti numbers | diffeomorphism — de Rham cohomology | isometry — curvature, the Laplacian spectrum |
| Tells apart | sphere vs torus | distinct smooth structures (exotic ) | round vs flat metric on one smooth shape |
Two cautions the tidy tower hides. The climb isn't automatic — not every topological manifold admits a smooth structure, and some admit many inequivalent ones (exotic ); yet once a space is smooth it always carries a metric, usually infinitely many. And “measure” is overloaded: the verb in the table is measuring a length or angle — a metric act needing — not a measure in the measure-theoretic sense, which is the separate tower one step down (measure vs probability space). The metric's is precisely what ties those two senses together.
The bottom rung is where persistent homology lives — it reads only the homeomorphism-level invariants, the features no continuous deformation can remove. The top rung is where the spectral work reaches: the spectrum of the Laplace–Beltrami operator is an isometry invariant — the same “shape from a spectrum” move, read on a curved space.
Measure space vs probability spacea measure space with its total mass pinned to 1 — probability is measure theory wearing one rule.
Another step on the tower — but this one adds a constraint, not new geometry. A probability space is a measure space whose total mass is pinned to . Everything in probability is measure theory wearing that one rule — and the rule unlocks a vocabulary (events, expectation, independence, conditioning) that plain measure theory never needs to name.
| Concept | Measure space | Probability space |
|---|---|---|
| Total mass | — may be infinite | — normalized |
| The ground set | a space / domain | the sample space of outcomes |
| The -algebra | measurable sets | events — yes/no questions |
| A measurable map | measurable function | random variable |
| Integration | expectation | |
| Density | Radon–Nikodym | pmf / pdf |
| Only here | — | independence, conditioning , the distribution as a pushforward |
The right column is the left plus a normalization — yet that one rule is what lets you speak of a distribution (the measure pushed through , exactly the pushforward from the pillars) and of independence, which has no analog in a generic measure space. The constraint is the content.
The probability tower — what each rung unlocksthe manifold tower run again inside probability — add integrability a rung at a time, a counterexample policing each gap.
The manifold tower, run again inside probability. Start from a bare law and add integrability a rung at a time; each rung admits new transforms and new inference, and a famous counterexample polices every gap. Same Klein test as above: a rung is what acts on it.
| Rung | Law (CDF) | + density | + two moments (L²) | + MGF (light tails) | + exponential family |
|---|---|---|---|---|---|
| Extra data | a distribution function | (Radon–Nikodym) | — joins | finite near 0 | log-density affine in : |
| You now have | order, quantiles, sampling by inversion — and , always | likelihood, MLE, entropy, KL, Bayes | mean & variance, , Chebyshev, the CLT | all cumulants, Chernoff bounds, large deviations | sufficiency, conjugate priors, Fisher geometry, GLMs |
| You can't yet | write a likelihood | trust moments — they may diverge | control tails past Chebyshev | pin a family's structure | keep heavy tails; close under maxima |
| Transform that arrives | quantile , PIT | the likelihood function | covariance as inner product | ; Legendre Cramér's rate | tilting ; mean map |
| The law that stops here | Cantor — continuous , no density | Cauchy — density, no mean | — variance, heavy tails; lognormal — all moments, no MGF, and the moments don't determine it | two-Normal mixtures — MGF fine, no finite-dim sufficiency | — (top rung here; extremes climb a different tower) |
The counterexamples are the load-bearing part — each one is a law that passes a rung and fails the next, so the tower can't be collapsed. Where the rungs come from is worked live in the distribution foundry; what moments fail to pin down is gen-fn §6.
Topology vs analysisone intuition, two primitives for nearness — a metric's ε–δ, or the open sets alone.
Not a rung but a span: two different primitives for one intuition — nearness. Analysis starts from a distance, a metric , and builds with and . Topology throws the ruler away and keeps only the open sets. Neither sits above the other; they meet at metric spaces, where a distance induces a topology and the two agree.
| Concept | Analysis — metric, ε–δ | Topology — open sets |
|---|---|---|
| Continuity | preimage of every open set is open | |
| Distance | the metric is primitive | forgotten; "metrizability" asks when a topology even comes from a |
| Boundary | limit points of sequences, via | , purely set-theoretic |
| Convergence | – with distances | nets & filters: eventually inside every neighborhood |
| Compactness | closed & bounded (Heine–Borel in ) | every open cover has a finite subcover |
The payoff is knowing which invariants you're allowed to use. Distance is a choice — the same lesson as change of basis, and the same lesson a QBist draws about a frame. Topology keeps only what survives any continuous deformation; analysis adds back the measuring stick when you need a rate.
Local vs globalcompute on simple local patches, then glue; when the glue fails, cohomology measures the gap.
The deepest structural lens of all, and the one the page badly needed. Local is where you can compute — near a point everything looks like , or like a single completion of a field. Global is where the truth lives — the whole object at once. The bridge is always the same: cover the object by simple local pieces, describe each, then try to glue. When the glue fails, the failure is measured by cohomology.
| Aspect | Local | Global |
|---|---|---|
| What you see | behavior on one patch / near a point | the whole object at once |
| Difficulty | easy — looks like a model space | the hard part |
| The bridge | a cover by patches + transition maps | gluing the patches consistently |
| Obstruction | — | cohomology: |
| In number theory | solve over each and | solve over |
| Examples | charts, germs, Taylor series, the p-adics | the manifold, the variety, the rational points |
This is one idea wearing three coats. It is the manifold idea (charts are local; the manifold is global); it is the local–global principle in number theory (solve everywhere locally, then ask if it glues — the failure is the Tate–Shafarevich group); and it is the heart of P9, where a nonzero first Čech cohomology of a sheaf of local policies is exactly a global coordination obstruction. Local is where you compute; global is where the answer hides; cohomology measures the gap.
Part II — the coordinatesRepresentation: in what coordinates?3
Fix the object; now choose how to write it. Every move here is one idea — re-express in a basis where a hard operation goes easy (differentiation becomes multiplication, a convolution becomes a product). The frame is yours to pick; the invariant is what survives the change. This is the change-of-basis pillar, industrialized.
Worked end to end as derivations: Taylor (into the powers), correlation (a projection),Student's t (a rotation), a Fourier series (into frequencies), and holomorphic functions (rotation-and-scale itself).
Transforms — Fourier vs Laplace vs Legendre vs Mellin vs Zfour are one move — a linear transform into an operator's eigenfunctions; Legendre is the convex odd one out.
Four of these five are the same move (a linear transform into the eigenfunctions of some operator); Legendre is the odd one out — not a change of basis but a convex duality that swaps a variable for its conjugate slope.
| Transform | Kernel / rule | Suited to | Turns the hard thing easy | Lives in | In probability |
|---|---|---|---|---|---|
| Fourier | oscillation, the line & circle | ; convolution → product | signals, PDEs, the characteristic function, QM | the characteristic function — exists for every law; Lévy continuity drives the CLT | |
| Laplace | causal, transient , growth | ODEs+initial data → algebra in | control theory, circuits, ODEs | the MGF — light tails only (Cauchy has none) | |
| Legendre | convex functions | trades variable for slope | Lagrangian ↔ Hamiltonian, thermodynamics, convex duality | Cramér's large-deviations rate — the Legendre dual of the CGF | |
| Mellin | scale-invariant, multiplicative | scaling → a phase | ζ & Dirichlet series, asymptotics, RH | for — products & ratios; where the foundry's Γ's live | |
| Z | discrete sequences | shift / delay → ; recurrences → algebra | DSP, digital filters, discrete control | the PGF — counts, branching processes |
They're relatives, not strangers: substitute and the Mellin transform becomes a two-sided Laplace; set and the Z-transform is Laplace sampled in discrete time. The one that opens onto the number-theory work is Mellin: it carries 's functional equation, which is why it, not Fourier, is the transform standing behind the Riemann Hypothesis. And probability runs all five at once — one law is a (Fourier), an MGF (Laplace), a PGF (Z), a Mellin moment function, and its tail rate is the Legendre dual of its log-MGF. The last column is not an analogy; it is the same five operators acting on a density.
Chebyshev — and the company it keepsto approximate you must choose a basis — Chebyshev is minimax on an interval, a Fourier cosine series in disguise.
To approximate a function you must choose a basis, and the choice is everything. The Chebyshev polynomials are the quiet champions on a closed interval — and the reason is a single substitution: under they are a Fourier cosine series, so they inherit the FFT's speed and Fourier's exponential accuracy for smooth functions.
| Approach | What it is | Wins when | Watch out |
|---|---|---|---|
| Equispaced / monomials | on evenly spaced nodes | quick, low degree | Runge phenomenon — wild edge oscillation; ill-conditioned |
| Taylor | derivatives at one point | local, near the center | only local; dies past the radius of convergence |
| Chebyshev | on nodes clustered at the ends | uniform accuracy on | needs the interval mapped to |
| Fourier (cosine) | on a periodic domain | periodic signals; FFT | Gibbs ringing at jumps; needs periodicity |
| Legendre / Hermite / Laguerre | orthogonal under a weight | the domain & weight match ( on ℝ, …) | not minimax; weight-specific |
The deep fact: among all monic degree- polynomials, has the smallest possible peak on — the precise sense in which Chebyshev is minimax, “best in the worst case.” And the endpoint-clustered nodes are exactly what cancels the Runge blow-up that dooms equispaced interpolation. Same lesson, one level up: pick the basis where the hard thing is easy.
Duality — every object has a mirrorpair an object with the maps out of it; the mirror holds the same information turned inside out (V ≅ V**).
The pattern hiding behind half of this part. Pair an object with “the maps out of it” and you get a dual that carries the same information turned inside out — and doing it twice brings you home. That round trip, , is the invariant; the art is asking your question in whichever description makes it linear.
| Object | Its dual | What the pairing flips |
|---|---|---|
| Vector space | functionals | vectors ↔ measurements; |
| Function | spectrum (Fourier) | time ↔ frequency; |
| Convex | conjugate (Legendre) | value ↔ slope |
| Abelian group | characters (Pontryagin) | |
| Homology | cohomology (Poincaré) | cycles ↔ cocycles |
| LP: | resources ↔ prices (strong duality) |
A problem that is opaque in one description is often transparent in its dual: a hard convolution is an easy product after Fourier; a hard primal program is its dual's shadow price. Duality is why the transforms above are not a grab-bag — Fourier and Legendre are both dualities, one linear and one convex. The mirror is not the thing; but you can read the thing off the mirror.
Part III — the knowerStance: what do you count as true?3
The first two parts were about the object. This one is about you. What counts as a proof, as existence, as a probability — these are commitments the mathematician makes, not facts read off the world. Here is where the Axiom of Choice sits next to QBism, exactly where it belongs.
Foundations & the Axiom of Choicethe Axiom of Choice — a genuine stance, not a deduction; the same bedrock as ‘every space has a basis.’
What it is. Given any collection of non-empty sets — even uncountably many — the Axiom of Choice asserts there is a choice function that picks one element from each at once, with no rule telling you how. Harmless-sounding; it is anything but. It is equivalent to Zorn's Lemma, to the Well-Ordering Theorem, to Tychonoff's theorem — and to “every vector space has a basis.”
Why it's on this page. Read that last equivalence again against the spine of this whole site. Your through-line is change of basis — and in infinite dimensions, the existence of a basis to change is the Axiom of Choice (a Hamel basis of over cannot be built, only chosen). The same axiom manufactures non-measurable sets (Vitali, Banach–Tarski) — which is precisely why the measure lens must restrict to a -algebra instead of measuring everything. Choice is not a footnote to the page; it is the bedrock under it.
| Question | Classical (ZFC) | Constructive / intuitionist |
|---|---|---|
| To exist is to… | be free of contradiction | be exhibited — show a witness |
| Excluded middle | assumed | not in general |
| Axiom of Choice | full | weak or none |
| A basis for every vector space | yes | not in general |
| Non-measurable sets | exist (Vitali) | never arise |
| Prove by contradiction | valid | rejected |
Choice is independent of the other axioms — Gödel showed you can't refute it, Cohen that you can't prove it — so adopting it is a genuine stance, not a deduction. And it is the same stance, one level down, that a QBist takes about probability: what is true depends on what you are willing to commit to, and the honesty is in owning the commitment.
The axioms you're free to choose — a MECE mapevery “Axiom of ___” is a fork; here is the whole map, grouped so the families don't overlap and together cover the base.
Choice is only the most famous fork. Here is the catalog of the “Axiom of ___” decisions a mathematician can make — each one independent (provable by no one, refutable by no one), each one a stance under which a different mathematics unfolds. Grouped so the families are mutually exclusive and, together, collectively exhaustive of the base.
| “Axiom of ___” | What it pins down | Deny it, and you get… |
|---|---|---|
| I · The Zermelo–Fraenkel core (ZF) — what a set is | ||
| Extensionality | a set is its members — same elements equal | sets with identity beyond their contents |
| Empty set · Pairing · Union · Power set | the constructions: | no way to build new sets from old |
| Infinity | one endless set exists — the naturals | only finite sets |
| Separation · Replacement | carve a subset by a predicate; take the image of a set (schemas) | Russell's paradox, or images that escape |
| Regularity (Foundation) | no -loops — the hierarchy is well-founded | a set that contains itself |
| II · Choice, graded — how much selection you allow | ||
| Choice (full) | pick one from each of infinitely many sets, with no rule; Zorn, well-ordering, “every space has a basis” | a vector space with no basis; every set measurable |
| Countable / Dependent Choice | the safe fragment analysis actually needs | even stops behaving sequentially |
| Ultrafilter lemma | a middle strength — prime ideals, compactness (Tychonoff) | no non-principal ultrafilters |
| Determinacy (AD) | the rival — every infinite game has a winner | with full Choice, undetermined games: |
| III · The undecidable frontier — independent of ZFC; pick your universe | ||
| Continuum Hypothesis | a size strictly between and ? | Cohen: either answer is consistent |
| Constructibility (V=L) | the leanest universe — settles CH (no) and Choice (yes) | room for large cardinals reopens |
| Martin's Axiom · large cardinals | assume bigger infinities: inaccessible, measurable, … | a shorter tower — less consistency strength |
| IV · Geometry — the first crisis | ||
| The Parallel Postulate | through a point, exactly one parallel (Euclid's 5th) | hyperbolic (many) or elliptic (none) — both consistent |
| V · Beyond sets — new foundations | ||
| Univalence | equivalent structures are equal: (HoTT) | equivalence and equality stay apart, as in set theory |
| Function extensionality | functions are equal iff equal at every input | functions that differ with no witnessing input |
MECE by construction: I fixes what a set is; II grades how freely you may choose (Choice lives here, not in I); III is everything ZFC leaves open; IV is the one geometric fork that started the whole “which axiom?” question; V moves to a foundation where equality itself is redefined. The section above lives in Family II; Univalence in Family V is its modern echo — equivalence to Choice's selection.
How much will you assume? — the modeling paradigmsbefore a distribution, a stance — how much shape are you willing to commit to?
Before a distribution, a stance: how much shape are you willing to commit to? The quiet choice underneath every model.
Parametric
Assume a distribution family; estimate a few numbers.
Commit to a shape — Normal(μ,σ), Poisson(λ), Weibull(k,λ) — and let the data fix its handful of parameters. Efficient and interpretable when the assumption holds; biased and overconfident when it doesn't.
Semi-parametric
Parametric where you trust it, free where you don't.
Pin down the part you care about with parameters and leave the nuisance part unspecified. The Cox model gives parametric covariate effects over a non-parametric baseline hazard; GAMs bolt smooth free-form terms onto a parametric backbone.
Non-parametric
Let the data choose the shape.
Make as few distributional assumptions as possible — estimate the density or the effect directly. Flexible and robust, but hungrier for data and easier to overfit. Most of modern machine learning lives here.
Bayesian vs Frequentist
A distribution over the parameter — or a fixed unknown.
Frequentists treat a parameter as a fixed number and reason about long-run sampling (confidence intervals, p-values). Bayesians put a probability distribution on the parameter — a prior, updated by data into a posterior (credible intervals). The Beta–Binomial pair is the cleanest bridge: a Beta prior on a rate becomes a Beta posterior.
Inference & interpretationwhat a probability is, what causation means, and how a claim earns trust.
What a probability is, what causation means, and how a claim earns trust.
QBism (Quantum Bayesianism)
A probability is an agent's own bet, not a fact in the world.
QBism reads the quantum state as one agent's personal degrees of belief, updated by experience, with the Born rule as a normative coherence constraint on those bets. The measurement 'paradox' dissolves: an outcome is an experience for the agent who acts. It's the philosophical engine under P2 — probability as something you act on, not something you find.
Causal inference
Correlation describes; causation says what happens if you intervene.
The move from P(Y | X) — observing — to P(Y | do(X)) — acting. Two languages: potential outcomes (Rubin) and structural causal models with do-calculus (Pearl), using DAGs to encode assumptions and decide what's even identifiable. Confounders, colliders, and backdoor adjustment are the grammar.
Proofs & counterexamples
A proof shows it must hold; a counterexample shows exactly where it breaks.
Two halves of one honesty: build the argument, then probe its edges. Learning to prove (Velleman's How to Prove It; Cummings' long-form Proofs) and learning where intuition fails (the Counterexamples shelf — Analysis, Topology, Probability, Measure) are the same discipline — testing a claim to destruction before you trust it.
Zoom out — when one lens isn't enoughWhen one lens isn't enough3
Structure, representation, stance are how you hold a single object. The hardest problems need more: you build a translation between whole worlds. Here the two great cultures fuse, new fields are born, and the immortal problems get their scaffolding.
Analysis vs algebra — the cultures that fusetwo temperaments that don't resolve into one — bounds vs exact structure; their fusion is where new fields are born.
Two temperaments more than two subjects, and — unlike the lenses above — not a structural relation at all. Analysis asks how big, how close, does it converge? and is content with bounds. Algebra asks what is the structure, up to isomorphism? and wants exact identities. The point of putting them last: they don't resolve into one being “under” the other — they fuse, and the fusion is where the new fields come from.
| Concept | Analysis | Algebra |
|---|---|---|
| Primitive | limit, approximation, the continuum | operation, structure, exactness |
| The question | how large? how near? does it converge? | what symmetry? what structure up to iso? |
| Tolerance | inequalities, error terms, | exact equalities, identities |
| Canonical object | integral, derivative, norm, measure | group, ring, field, module, morphism |
| Truth via | estimates & convergence | axioms & universal properties |
| Infinity | completion ( from ) | structural (extensions, closures) |
Analytic number theory uses analysis to interrogate the integers (, L-functions); representation theory uses algebra to organize analysis; the Langlands program is the grand reconciliation. The Riemann Hypothesis is exactly an analytic question about an arithmetic object — which is why it is so hard, and so central.
How a new paradigm is madepoint one lens at another's objects — the faithful translation is itself the new paradigm.
Where do whole new fields come from? Most often, one lens is pointed at the objects of another — a faithful translation that sends a hard problem in one world to a tractable one in the next. The translation itself is the paradigm. It's change of basis at the scale of a discipline.
- Algebraic topology = algebra ⟶ topology. Homology turns a space into a sequence of groups, so “can these shapes be deformed into each other?” becomes “are these groups isomorphic?” The functor is the whole idea — and it's the engine of the TDA work.
- Differential geometry = calculus ⟶ geometry. Lay a smooth, then metric, structure over a space (the manifold ladder) and curvature becomes something you can differentiate and compute.
- Analytic number theory = analysis ⟶ the integers. Encode arithmetic in a function (, L-functions) and bring the machinery of the continuum to bear — the home of RH.
- Algebraic geometry = algebra ⟶ geometry. Solutions of polynomials become a space; ideals become subspaces. The speculative “field with one element” (P6) is a bet on the next fusion.
The art is finding a translation that loses nothing essential. That is also why category theory keeps surfacing in this program: it is the mathematics of faithful translations — functors between worlds — the language in which “the same idea, seen twice” can be made precise. The honest verb, again, is =.
Paradigms that scaffold the immortalsthe reusable moves behind BSD, RH, and Langlands — each builds a functor to a world with more structure.
The immortal problems — BSD, the Riemann Hypothesis, Langlands — aren't attacked head-on. They're approached through a handful of reusable moves that recur across every serious assault: each one builds a functor to a world with more structure. These are the scaffolding, not the building.
Local–global principle
Understand a global object by assembling all of its local pieces.
Study a problem one prime at a time (and at ∞), then ask whether solvability everywhere locally forces a global solution. For quadratic forms it does (Hasse–Minkowski); in general it can fail, and the failure is itself the prize — for an elliptic curve that obstruction is the Tate–Shafarevich group, the term BSD's finer form is really about.
Analytic continuation & a functional equation
Define an object by a series in one region, extend it everywhere, exploit its symmetry.
A Dirichlet series converges only on a half-plane, but continues to a function on (almost) all of ℂ — and the continuation satisfies a reflection s ↔ 1−s. The entire meaning of ζ's nontrivial zeros, and of L(E,s) at the central point s = 1, lives in the continued region the original sum never reached.
Spectral interpretation
Realize arithmetic data as the spectrum of an operator.
The Hilbert–Pólya dream: if the nontrivial zeros of ζ were the eigenvalues of some self-adjoint operator, RH would follow from realness. No one has found the operator — but the zeros' spacing statistics do match those of a random Hermitian matrix (Montgomery–Dyson, the GUE link). Read a hard arithmetic object through a spectrum.
Parametrize the family, then study it
Don't study one object — build its moduli space and study the statistics of the whole family.
A single curve or L-function is hard; a family, ordered by conductor, has structure a single object can't show. Katz–Sarnak: a family's low-lying zeros follow one of three random-matrix symmetry types (unitary, orthogonal, symplectic), and the type predicts central vanishing. The object of study becomes the family, not the member.
Reciprocity — a dictionary between two worlds
Translate a number-theory question into harmonic analysis, where it becomes tractable.
The Langlands paradigm itself: a correspondence sending Galois representations to automorphic forms, so a question about symmetries of number fields becomes a question about spectra of automorphic objects. Class field theory is the abelian case already proven; the general dictionary is the program's horizon — and BSD and RH are entries in it.