Backporch
Library · the ideas the program runs on, hands-on

The library

More than a statistics shelf. This is a hands-on library of the ideas the whole program runs on: a cross-subject glossary — random matrices, L-functions, topology, decision theory, and the contemplative thread, each term tagged by subject and project — alongside interactive probability distributions you can drag, an inference glossary, an AI glossary, and a which-distribution guide. Built to learn out loud, and the same toolkit behind the studio's models.

New here, or want the fabric underneath? Start with The pillars — sample spaces and events, why a random variable is neither random nor a variable, measurement scales for any data type, and theorems as true statements.

Ideas & terms — a one-stop glossary

The vocabulary the research program runs on, in one place: random matrix theory, L-functions and number theory, topological data analysis, statistics, probability, decision theory, and the contemplative thread. Every term is tagged by subject and by the project it earns its keep in — filter by either, or search the lot.

Subject
Project

32 of 32 terms

Analytic rank & central valueL(½) and its order of vanishing there.

The central value L(½) (or L(E,1) for an elliptic curve) and the order to which the L-function vanishes at the centre — the analytic rank. Whether it vanishes is exactly what the family's symmetry type predicts.

Number theory & L-functionsP6P10
Automorphic L-functionAn L-function attached to an automorphic representation (Langlands).

The Langlands program conjectures every L-function arises from an automorphic representation of GL(n). Dirichlet, Hecke, and modular-form L-functions are the low-rank cases.

Number theory & L-functionsP6
Betti numbersCounts of holes by dimension: H₀ components, H₁ loops, H₂ voids.

βₖ = dim Hₖ(X) = dim(ker ∂ₖ / im ∂ₖ₊₁). β₀ counts connected pieces, β₁ independent loops, and so on — the coarse topological fingerprint persistence tracks across scale.

TDAP1P9
Birch–Swinnerton-Dyer (BSD)The rank of an elliptic curve equals the order of vanishing of its L-function at s = 1.

A Clay Millennium problem linking arithmetic (the rank of E(ℚ)) to analysis (vanishing of L(E,s) at the centre). Known for analytic rank ≤ 1 (Gross–Zagier, Kolyvagin). The reason 'vanishing' and 'spacing' are the same question seen twice.

Number theory & L-functionsP6
Bottleneck & Wasserstein distanceMetrics between persistence diagrams; the basis of stability.

Optimal-matching distances between diagrams. The stability theorem says small data perturbations move the diagram only a little in bottleneck distance — what makes persistence a trustworthy statistic.

TDAP5P9
CalibrationBeliefs that match frequencies: of the things you call 90% likely, 90% happen.

The honest target for a probabilistic agent and the heart of the QBist reading of P2 — a decision-maker's credences are good insofar as they are calibrated, not insofar as they feel confident.

Decision theoryProbabilityP2
Contextual fractionA sheaf-cohomology measure of how non-classical a system of observations is.

Abramsky–Brandenburger model contextuality as the failure of local data to glue into a global section; the contextual fraction quantifies it as a dimensionless effect size — the bridge between QBism, sheaf theory, and P9's conjectures.

Decision theoryRMTP2P9
Effect size vs p-valueHow big versus how surprising — report both.

A p-value measures incompatibility with the null; an effect size measures magnitude with units and a confidence interval. A significant tiny effect (the P9 ρ ≈ −0.22) can be real yet useless — which is why the program reports effect sizes, not just p.

StatisticsFoundations
Effective dimensionalityHow many directions in your data are real signal.

The count of eigenvalues above the Marchenko–Pastur edge — a principled, scale-free alternative to scree plots and the Kaiser rule. In real neural and galaxy-spectra data this came out near ~12, stably.

RMTStatisticsP10
Eigenvalue clipping & shrinkageDenoising a covariance by treating sub-edge eigenvalues as noise.

Replace the bulk (below the MP edge) eigenvalues with their average (clipping) or shrink toward a target (Ledoit–Wolf). Stays well-conditioned even when samples are scarcer than dimensions — where the raw covariance is singular and the Hartlap correction is undefined (P10/S3).

RMTStatisticsP10
GUE / GOE / GSE ensemblesThe three classical random-matrix ensembles (β = 2, 1, 4).

Gaussian Unitary (complex Hermitian, β=2), Orthogonal (real symmetric, β=1), and Symplectic (β=4) ensembles. They set the universal local spacing statistics: a real data covariance is GOE; the Riemann zeros behave like GUE.

RMTP10
Katz–Sarnak symmetry typesFamilies of L-functions show unitary, orthogonal, or symplectic statistics.

Ordered by conductor, a family's low-lying zeros match one of three ensembles (U / O / Sp). Proven over function fields (via monodromy), conjectural over ℚ. The symmetry type predicts central vanishing — the natural extension of P10's spectral zoo into arithmetic.

Number theory & L-functionsRMTP10
L-functionA Dirichlet series with an Euler product and a functional equation; ζ is the prototype.

An analytic object encoding arithmetic, of the form Σ aₙ/nˢ, with an Euler product over primes and a symmetry s ↔ 1−s. Its zeros control deep arithmetic; their statistics are conjecturally governed by random-matrix ensembles.

Number theory & L-functionsP6P10
Level repulsionGeneric spectra avoid crowding; nearby levels push apart.

The signature of random-matrix (and quantum-chaotic) spectra: the probability of a tiny gap vanishes. Its opposite is level clustering (Poisson, independent levels) and, at the far extreme, a rigid comb (equal gaps).

RMTP10
LMFDBThe L-functions and Modular Forms Database — the field's one-stop data source.

An open atlas of L-functions, modular forms, elliptic curves, and their zeros — the practical place to pull family data for spacing / low-lying-zero experiments (with Sage / pari for computation).

Number theory & L-functionsP6P10
Marchenko–Pastur lawThe noise band of sample-covariance eigenvalues; upper edge (1+√(p/n))².

For p variables and n samples of pure noise, the sample correlation eigenvalues fill [(1−√(p/n))², (1+√(p/n))²]. Eigenvalues above the upper edge cannot be explained by noise — they are candidate signal. This edge is the factor-selection decision instrument in P10.

RMTP10
Method of lociThe memory palace: bind ideas to places you can walk.

An ancient mnemonic — encode an ordered argument as a route through a remembered space. Used here as a study aid for derivations, and a small example of the contemplative-meets-computational practice the program runs on.

ContemplativeStatisticsFoundations
Montgomery–Dyson phenomenonThe spacings of the Riemann zeros match GUE eigenvalue statistics.

Montgomery (1973) computed the pair correlation of ζ-zeros; Dyson recognised the GUE form. P10 reproduces it with the spacing ratio: the zeros give ⟨r⟩ ≈ 0.60, the unitary corner of the spectral zoo.

Number theory & L-functionsRMTP6P10
Motivic L-functionAn L-function attached to a motive / Galois representation.

Built from arithmetic geometry — elliptic curves, modular forms, symmetric powers. These carry the deepest arithmetic invariants (ranks, special values) and are where vanishing meets BSD.

Number theory & L-functionsP6
Out-of-sample validationJudge a method on data it didn't see — the honest test.

Cross-validation / held-out testing: fit on one split, score on another. The whole P10 'earn its place' discipline rests on it — an in-sample correlation that vanishes out-of-sample (as in the P9 wedge) is not a finding.

StatisticsFoundationsP10
OverfittingFitting noise as if it were signal; great in-sample, poor out.

When a model has enough freedom to memorise quirks of the training data. Why keeping too many (noise) dimensions hurts — and why the Marchenko–Pastur edge's parsimony pays off when samples are scarce.

StatisticsFoundations
Persistence diagram / barcodeThe birth–death record of topological features.

Each feature is a point (birth, death) — or a bar of that length. Distance from the diagonal = persistence = how 'real' the feature is. The object every TDA summary statistic is computed from.

TDAP1P5P9
Persistence entropyA single number summarising how spread-out a diagram's lifetimes are.

Shannon entropy of the normalised persistence lifetimes — a scale-free scalar feature used to compare diagrams across objects (e.g. the khipu signatures in P9).

TDAP9
Persistent homologyMultiscale topology: which holes in a point cloud survive across scale.

Build a growing family of complexes from data and track when topological features (components, loops, voids) are born and die. Long-lived features are real structure; short-lived ones are noise — a topological echo of the signal/noise split.

TDAP1P5P9
Process-relational / interbeingReality as relations and processes (morphisms), not substances (objects).

The shared substrate of P9: Whitehead's process philosophy, category theory's primacy of morphisms, QBism's experience-first reading, and the contemplative 'interbeing' of Thich Nhat Hanh — read here as a stance you can try to make falsifiable, not a slogan.

ContemplativeP9
Random matrix theory (RMT)The statistics of eigenvalues of large random matrices — and their surprising universality.

RMT studies how the eigenvalues of large random matrices behave. Two facts power everything else: the bulk follows a fixed law (Marchenko–Pastur for covariance), and the local spacing statistics are universal — depending only on a symmetry class, not the matrix's details.

RMTP10
Riemann zeta & the Riemann Hypothesisζ(s) = Σ 1/nˢ; RH says its nontrivial zeros lie on Re(s) = ½.

The prototype L-function. The Riemann Hypothesis — all nontrivial zeros on the critical line — is the central open problem; its truth controls the error term in the prime-counting function.

Number theory & L-functionsP6P10
Spacing ratio ⟨r⟩An unfolding-free statistic for level repulsion (Atas et al.).

r = min(sₙ,sₙ₊₁)/max(sₙ,sₙ₊₁) on consecutive gaps; its mean is density-independent so it needs no 'unfolding'. Benchmarks: Poisson 0.386, GOE 0.536, GUE 0.600, rigid comb → 1. The backbone of P10's spectral-zoo comparisons.

RMTStatisticsP10
Tracy–Widom distributionThe fluctuation law of the largest eigenvalue at the spectral edge.

How far the top eigenvalue strays beyond the Marchenko–Pastur edge — the rigorous basis for deciding whether a borderline eigenvalue is signal or an edge fluctuation of noise.

RMTP10
Vanishing & non-vanishingWhether L(½) = 0 — tied to rank and to symmetry type.

Orthogonal families have a positive density of central zeros (extra vanishing); symplectic ones repel zeros from the centre; unitary ones are generic. Non-vanishing theorems (Iwaniec–Sarnak, Soundararajan) bound how often L(½) ≠ 0.

Number theory & L-functionsP10
Vietoris–Rips filtrationThe standard growing complex built from a point cloud.

Connect points within distance ε and fill in simplices; grow ε from 0 to ∞. The resulting nested complexes are what persistent homology is computed on.

TDAP1P9
Wigner surmiseThe level-spacing law showing repulsion: P(s) → 0 as s → 0.

An accurate closed form for the nearest-neighbour spacing distribution of RMT eigenvalues — e.g. GUE: P(s) ≈ (32/π²) s² e^(−4s²/π). The s² prefactor is level repulsion: levels almost never coincide, unlike independent (Poisson) levels.

RMTP10

Distributions & inference, hands-on

The interactive distribution catalog, the inference glossary, the bootstrap demo, and the which-distribution guide.

Distributions

Toggle PMF/PDF ↔ CDF and drag the sliders — every chart recomputes live.

Bernoulli

discrete

A single yes/no trial — the atom every other count is built from.

00.510.5
mean 0.5 = pvariance 0.25 = p(1−p)
supportk ∈ {0, 1}
pmf/pdf
when & whyOne trial, two outcomes, a fixed success probability p. The honest question it forces: is your event really binary, and is p really constant?
E-commerce Did this visitor convert — buy or not?
Healthcare Did the patient readmit within 30 days?
Manufacturing Did the unit pass or fail inspection?
binaryyes/nosuccessindicatorconversion

Binomial

discrete

How many successes in a fixed number of independent yes/no trials.

010200.184
mean 7 = npvariance 4.55 = np(1−p)
supportk ∈ {0, 1, …, n}
pmf/pdf
when & whyn fixed independent trials, each with the same p. Breaks when trials are correlated (clustering) or p drifts over time.
Manufacturing Defective units in a batch of n parts.
Marketing Clicks out of n impressions on a campaign.
Public health Vaccinated people in a sample of n.
Finance Loans that default out of n issued.
countssuccessestrialsproportiondefectsconversion

Geometric

discrete

How many failures before the first success — the waiting game, in trials.

016320.25
mean 3 = (1−p)/pvariance 12 = (1−p)/p²
supportk ∈ {0, 1, 2, …}
pmf/pdf
when & whyIndependent trials with constant p, counted until the first success. Memoryless — past failures don't make success any nearer.
Sales Cold calls before the first yes.
Reliability Cycles a part survives before first failure.
Support Retries before a request finally succeeds.
waitingfirst successtrials untilretrymemoryless

Negative Binomial

discrete

Counts with extra spread — the realistic cousin of Poisson when variance beats the mean.

015.5310.1
mean 7.5 = r(1−p)/pvariance 18.75 = r(1−p)/p²
supportk ∈ {0, 1, 2, …}
pmf/pdf
when & whyFailures before the r-th success — or, more usefully, a Poisson whose rate itself varies (a Poisson–Gamma mixture). The go-to when counts are overdispersed: variance > mean.
Insurance Claims per policy per year (classic overdispersed count).
Public sector 311 requests or crime incidents per district.
Transportation Crashes per intersection per year (traffic-safety standard).
Healthcare Infections per 1,000 patient-days (clustered, not Poisson).
countsoverdispersionclusteringclaimsincidentspoisson mixture

Poisson

discrete

How many events happen in a fixed window when they arrive at a steady average rate.

09180.195
mean 4 = λvariance 4 = λ
supportk ∈ {0, 1, 2, …}
pmf/pdf
when & whyEvents independent, a constant average rate, no two at exactly the same instant. Its signature — and limitation — is mean = variance. Real counts usually spread wider (see Negative Binomial).
Operations Calls arriving at a help desk per hour.
Manufacturing Defects per unit when the process is stable.
Web / telecom Requests hitting a server per second.
Sports Goals per soccer match (the textbook fit).
countsratearrivalsrare eventsper windowqueue

Discrete Uniform

discrete

Every outcome in a range equally likely — the fair die.

13.560.167
mean 3.5 = (a+b)/2variance 2.917 = ((b−a+1)²−1)/12
supportk ∈ {a, …, b}
pmf/pdf
when & whyA finite set of outcomes with nothing to distinguish them. The honest 'I have no reason to prefer any value' baseline.
Simulation Rolling a die or picking a random slot.
Sampling Selecting one record at random from a frame.
fairequally likelyrandom pickdiebaseline

Hypergeometric

discrete

Successes when you sample without replacement — Binomial's finite-population cousin.

05100.298
mean 3 = nK/Nvariance 1.714 = n·(K/N)(1−K/N)·(N−n)/(N−1)
supportk valid given N, K, n
pmf/pdf
when & whyDrawing n items without replacement from N, of which K are 'successes'. Use it instead of Binomial when the population is small enough that each draw changes the odds.
Quality / audit Defectives found when sampling a finite lot.
Elections / cards Drawing a hand from a fixed deck.
without replacementsamplingqualityauditfinite population

Continuous Uniform

continuous

Equal density across an interval — flat ignorance between two bounds.

0.164.58.840.143
mean 4.5 = (a+b)/2variance 4.083 = (b−a)²/12
supportx ∈ [a, b]
pmf/pdf
when & whyAny value in [a, b] equally plausible. The maximum-entropy choice when all you know is a range.
Simulation Random arrival time within a known window.
Pricing A value known only to lie between two bounds.
flatrangebaselinerandom numbermax entropy

Normal (Gaussian)

continuous

The bell curve — what sums of many small independent effects converge to.

-6060.266
mean 0 = μvariance 2.25 = σ²
supportx ∈ ℝ
pmf/pdf
when & whySymmetric, light-tailed, fully described by mean and variance. Justified by the Central Limit Theorem — but it under-models heavy tails and skew, so don't assume it for everything.
Manufacturing Machined dimensions around a target (SPC).
Finance Daily returns — a useful-but-imperfect first model.
Any Measurement error and sums of many small effects.
bell curvegaussianerrorscltmeasurementz-score

Exponential

continuous

The waiting time until the next event when events arrive at a steady rate.

04.2868.5710.7
mean 1.429 = 1/λvariance 2.041 = 1/λ²
supportx ≥ 0
pmf/pdf
when & whyMemoryless waiting between independent, constant-rate events — the continuous partner of Poisson. Time already waited tells you nothing about time remaining.
Operations Time between customer arrivals.
Reliability Time-to-failure when the rate is constant.
Telecom Gaps between packets on a line.
waiting timetime betweenlifetimememorylessqueuereliability

Gamma

continuous

Waiting time for several events to accumulate — positive, skewed, flexible.

05.46410.9280.271
mean 3 = variance 3 = kθ²
supportx > 0
pmf/pdf
when & whySum of k independent exponential waits; for integer k it's the Erlang. A versatile model for positive, right-skewed quantities — and the conjugate prior for a Poisson rate.
Insurance Claim size / severity (positive and right-skewed).
Utilities Rainfall totals and load durations.
Bayesian Prior on a Poisson or exponential rate.
positiveskewedwaitingrainfallinsurance severityprior

Beta

continuous

A distribution over a probability itself — your belief about a rate between 0 and 1.

00.512.458
mean 0.286 = α/(α+β)variance 0.026 = αβ / ((α+β)²(α+β+1))
supportx ∈ (0, 1)
pmf/pdf
when & whyLives on [0,1], so it models proportions and probabilities. As the conjugate prior for the Binomial, α−1 and β−1 read as prior successes and failures — belief you update with data.
Experimentation Posterior belief about a conversion rate.
Marketing Click-through rate with a prior, updated nightly.
Reliability Uncertain pass-rate of a component.
proportionprobabilityrateconjugate priorbayesianconversion rate

Log-normal

continuous

When the logarithm is Normal — positive, heavy-tailed, born from multiplying effects.

04.0838.1660.496
mean 2.065 = e^(μ+σ²/2)variance 1.211 = (e^(σ²)−1)e^(2μ+σ²)
supportx > 0
pmf/pdf
when & whyUse it when a quantity is a product of many positive factors (so its log is a sum → Normal). Right-skewed and strictly positive — incomes, sizes, durations.
Economics Incomes and wealth (long right tail).
Web / infra Request latencies and file sizes.
Insurance An alternative severity model to Gamma.
positiveskewedmultiplicativeincomefile sizeduration

Weibull

continuous

The reliability workhorse — failure times whose hazard can fall, hold, or climb.

03.6247.2490.373
mean 1.805 = λ·Γ(1+1/k)variance 1.503 = λ²[Γ(1+2/k) − Γ(1+1/k)²]
supportx ≥ 0
pmf/pdf
when & whyThe shape k sets the hazard's direction: k<1 is infant mortality (a falling failure rate), k=1 is constant (it collapses to the Exponential), k>1 is wear-out (a rising rate). The default model for time-to-failure and wind speed.
Reliability Time-to-failure of bearings, batteries, components.
Energy Wind-speed distribution for turbine siting.
Manufacturing Fatigue life and weakest-link breaking strength.
reliabilitysurvivalfailure timehazardwind speedweakest link

Chi-squared

continuous

Sum of squared standard Normals — the engine behind variance and goodness-of-fit tests.

09.65719.3140.184
mean 4 = νvariance 8 =
supportx > 0
pmf/pdf
when & whyThe distribution of a sum of ν squared independent standard Normals. Shows up wherever you test variances or compare observed-vs-expected counts.
Analytics Chi-squared test of independence in a contingency table.
Quality Inference about a process variance.
goodness of fitvariance testdfindependence testgamma special case

Student's t

continuous

The bell curve with heavier tails — what you use when the variance is estimated, not known.

-6060.375
mean 0 = 0 (ν>1)variance 2 = ν/(ν−2) (ν>2)
supportx ∈ ℝ
pmf/pdf
when & whyLike a Normal but with fatter tails that shrink toward Normal as df grow. The right reference when you estimate σ from a small sample — the basis of the t-test.
Experimentation Comparing two group means on small samples.
Finance A heavier-tailed model for returns than Normal.
t-testsmall sampleheavy tailsrobustmean inference

Compare two distributions

Overlay any two to see the shape difference — start with Poisson vs Negative Binomial (same mean, fatter tail).

015.5310.195
PoissonNegative Binomial

Foundations — the inference glossary

The ideas the distributions sit inside. The same search box filters these.

Expected valueThe probability-weighted long-run average.

E[X] = Σ k·P(k) for discrete, ∫ x·f(x) dx for continuous. It's the balance point of the distribution — and need not be a value X can actually take (a die's mean is 3.5).

meanaveragemoment
Variance & standard deviationHow far outcomes spread from the mean.

Var(X) = E[(X−μ)²]; the standard deviation is its square root, back in the original units. Variance adds for independent variables, which is why it shows up everywhere.

spreaddispersionsd
SupportThe set of outcomes a distribution actually puts mass on.

The support is where the PMF/PDF is positive — the values that can actually occur: supp(f) = closure of { x : f(x) > 0 }. It bounds where estimates and integrals live, and mismatched supports (a model assigning probability 0 to data you observed) quietly break likelihoods and importance sampling.

supportdensityfoundationsrigor
Probability vs likelihoodSame formula, opposite thing held fixed.

Probability fixes the parameters and asks about data: P(data | θ), summing to 1 over data. Likelihood fixes the observed data and varies the parameters: L(θ) = P(data | θ) read as a function of θ — it is NOT a probability distribution over θ.

likelihoodprobabilitymlefoundations
Probability distributionThe full map from outcomes to their probabilities.

A complete description of a random quantity: every outcome and how much probability (or density) it carries. Everything else — mean, variance, tail risk — is read off from it.

pmfpdfdistribution
PMF vs PDFMass for discrete, density for continuous.

A PMF gives P(X=k) directly and sums to 1. A PDF gives density, where probability is area under the curve; for continuous X, P(X = any exact point) is 0, so only intervals carry probability.

pmfpdfdensity
Likelihood function & MLEPick the parameters that make your data least surprising.

The likelihood L(θ) = ∏ f(xᵢ; θ). Maximum likelihood estimation chooses the θ that maximizes it — usually via the log-likelihood, which turns the product into a sum and the optimization into calculus.

mleestimationlikelihood
Estimator vs estimateThe rule versus the number it produces.

An estimator is a function of the data — itself a random variable with its own distribution (e.g. the sample mean). An estimate is the single value that rule returns for one dataset. Properties like bias and variance belong to the estimator.

estimatorestimatesampling distribution
BiasHow far an estimator is wrong on average.

Bias = E[estimator] − true value. Unbiased means zero on average across samples — but a slightly biased estimator with much lower variance is often the better bet (the bias–variance trade-off).

biasaccuracy
ConsistencyIt converges to the truth as data grow.

An estimator is consistent if it converges in probability to the true parameter as n → ∞. More data should pin it down; an inconsistent estimator never settles on the right answer no matter how much you collect.

consistencyasymptoticslarge n
SufficiencyA summary that loses nothing about the parameter.

A statistic is sufficient for θ if the data carry no extra information about θ once you know it (Fisher–Neyman factorization). The sample sum is sufficient for a Poisson rate — you can throw away the rest.

sufficiencyinformation
EfficiencyThe lowest-variance estimator you can get.

Among unbiased estimators, the efficient one has the smallest variance — bounded below by the Cramér–Rao limit. Efficiency is why we prefer one valid estimator over another.

efficiencycramer-raovariance
Hypothesis testingAssume nothing's going on, then see if the data argue otherwise.

A decision procedure: take H₀ as true, compute how surprising the data would be under it, and reject H₀ only if that surprise crosses a pre-set threshold. It controls error rates, not truth.

testinginferenceh0
Null & alternative hypothesesThe default vs what you'd conclude instead.

H₀ is the no-effect / status-quo claim you try to falsify; H₁ is the alternative you'd accept if the evidence is strong enough. You never 'prove' H₀ — you only fail to reject it.

h0h1testing
Test statisticA number whose behavior under H₀ you know.

A summary of the data (z, t, χ², F) whose sampling distribution is known when H₀ holds. You locate your observed value in that distribution to judge how extreme it is.

ztchi-squaredstatistic
p-valueHow surprising the data are if H₀ is true.

The probability, assuming H₀, of a test statistic at least as extreme as the one observed. It is NOT the probability that H₀ is true, nor the chance your result was luck — just a measure of incompatibility with H₀.

p-valuesignificancetesting
Significance level αThe false-positive rate you agree to tolerate.

Chosen before you look (often 0.05): the probability of rejecting a true H₀. Reject when p < α. It's a budget for Type-I errors, not a law of nature — pick it for the decision's stakes.

alphatype isignificance
Type I & Type II errorsFalse alarm vs missed signal.

Type I: rejecting a true H₀ (a false positive), rate α. Type II: failing to reject a false H₀ (a missed effect), rate β. Tightening one usually loosens the other for fixed sample size.

type itype iierrors
PowerThe chance of catching a real effect.

Power = 1 − β: the probability of correctly rejecting H₀ when H₁ is true. It rises with sample size, effect size, and lower noise — the thing a good experiment is designed around.

powersample sizedesign
Confidence intervalA procedure that brackets the truth most of the time.

A 95% CI is built by a method that, across repeated samples, contains the true parameter 95% of the time. The frequentist subtlety: it's a property of the procedure, not a 95% probability about this one interval.

ciintervaluncertainty
Central Limit TheoremAverages tend to a bell curve.

Sums and means of many independent, finite-variance variables approach a Normal distribution regardless of the original shape. It's why so much of inference leans on the Normal even when the data aren't.

cltnormalasymptotics
Law of Large NumbersSample averages settle on the expected value.

As n grows, the sample mean converges to E[X]. It's the guarantee that more data, gathered honestly, pulls your estimate toward the truth — the backbone of simulation and bootstrapping.

llnconvergenceaverages
BootstrappingResample your data to feel out its uncertainty.

Resample the dataset with replacement many times, recompute the statistic each time, and use the spread of results as an approximate sampling distribution — confidence intervals with almost no distributional assumptions.

bootstrapresamplingcinonparametric
BaggingAverage many models trained on resamples to cut variance.

Bootstrap AGGregating: fit a model on each bootstrap resample and average (or vote) their predictions. It tames variance in unstable learners — random forests are bagging plus feature subsampling.

baggingensemblerandom forestvariance
Bayesian vs frequentistBeliefs you update vs long-run frequencies.

Frequentists treat parameters as fixed unknowns and reason about repeated sampling. Bayesians put a probability distribution on the parameter — a prior updated by data into a posterior. Different questions, both useful; the Beta–Binomial pair is the classic bridge.

bayesianfrequentistpriorposterior
Conjugate prior (Beta–Binomial)A prior whose update keeps the same shape.

A Beta prior on a probability, combined with Binomial data, yields a Beta posterior — just add successes to α and failures to β. That tidy update is why Beta is the natural language for an uncertain rate.

conjugatebetabinomialbayesian

Bootstrapping, live

A confidence interval built by resampling — no formula, no distributional assumption.

1.2013.763
sample mean 2.36495% CI [1.713, 3.124]true mean 2

One fixed sample of 40 draws from a skewed (Exponential) population — we never see the rest of the world. Each resample draws 40 of those values with replacement and takes the mean; the teal histogram is the bootstrap distribution of the mean, and the dashed amber lines are the 2.5th/97.5th percentiles — a 95% confidence interval built with almost no assumptions. Slide B up and watch it steady (the Law of Large Numbers at work); the white line is the original sample mean.

Which distribution? — a click-through guide

Start from the question you're actually asking, then click a suggestion to jump to its card.

Counting events in a fixed window (time, space, batch)?
Poisson if the spread looks like the mean; Negative Binomial if the variance runs hotter (overdispersed) — which it usually does.
Counting successes in a fixed number of yes/no trials?
Binomial. A single trial is Bernoulli. Sampling without replacement from a small pool → Hypergeometric.
Counting trials until a success happens?
Geometric for the first success; Negative Binomial for the r-th.
Modeling a proportion or a probability between 0 and 1?
Beta — and it doubles as the prior you update with data.
A waiting time or time between events?
Exponential for the next event (constant rate); Gamma for several events to accumulate.
A measurement clustering symmetrically around a center?
Normal — or Student's t when you're estimating the spread from a small sample.
A positive, right-skewed quantity (size, income, duration)?
Log-normal if it's a product of effects; Gamma is a flexible alternative.
Truly no reason to prefer any outcome in a range?
Uniform — discrete for a finite set, continuous for an interval.

Deep dive — counting things (Poisson & Negative Binomial)

Counts are the most common data shape in real operations — events per window, per unit, per place. Poisson is the honest starting point (one rate, mean = variance), but real counts almost always spread wider than Poisson allows: a few busy days, a clustered outbreak, a bad batch. That overdispersion is exactly what the Negative Binomial absorbs (it's a Poisson whose rate itself varies). Everything below is a count problem — and where the variance beats the mean, reach for Negative Binomial.

Negative Binomial

discrete

Counts with extra spread — the realistic cousin of Poisson when variance beats the mean.

015.5310.1
mean 7.5 = r(1−p)/pvariance 18.75 = r(1−p)/p²
supportk ∈ {0, 1, 2, …}
pmf/pdf
when & whyFailures before the r-th success — or, more usefully, a Poisson whose rate itself varies (a Poisson–Gamma mixture). The go-to when counts are overdispersed: variance > mean.
Insurance Claims per policy per year (classic overdispersed count).
Public sector 311 requests or crime incidents per district.
Transportation Crashes per intersection per year (traffic-safety standard).
Healthcare Infections per 1,000 patient-days (clustered, not Poisson).
countsoverdispersionclusteringclaimsincidentspoisson mixture

Poisson

discrete

How many events happen in a fixed window when they arrive at a steady average rate.

09180.195
mean 4 = λvariance 4 = λ
supportk ∈ {0, 1, 2, …}
pmf/pdf
when & whyEvents independent, a constant average rate, no two at exactly the same instant. Its signature — and limitation — is mean = variance. Real counts usually spread wider (see Negative Binomial).
Operations Calls arriving at a help desk per hour.
Manufacturing Defects per unit when the process is stable.
Web / telecom Requests hitting a server per second.
Sports Goals per soccer match (the textbook fit).
countsratearrivalsrare eventsper windowqueue

Tweedie — the pure-premium bridge

Frequency (Poisson) × severity (Gamma) in a single model. Where the counts story meets the dollars.

Tweedie

simulated

One family spanning Poisson (p=1), Gamma (p=2), and everything between — with a real point mass at zero.

0 (spike)7.387
P(X=0) 0.042 = e⁻λmean 2.54 ≈ μVar/μᵖ 1.006 ≈ φ
variance
when & whyFor 1<p<2 it's a compound Poisson–Gamma: a Poisson count of Gamma-sized amounts. That gives an exact spike at zero (no events) plus a skewed positive tail (when events happen) — exactly insurance pure premium, where most policies cost nothing and a few cost a lot. There's no closed-form density in this range, so the chart is simulated.
Insurance Pure premium per policy — claim frequency × severity in one model.
Retail / demand Spend per customer: many zeros, a long right tail.
Machine learning XGBoost / LightGBM reg:tweedie for zero-heavy, right-skewed targets.
compound poisson-gammazero-inflatedpure premiumpower variancexgboost

The violet bar is the mass at exactly zero (no claim); the amber tail is the cost when something happens. Push p→1 and it leans Poisson (more, smaller events); p→2 and it leans Gamma (fewer, larger). λ ≈ 3.162 events expected per draw.

An example in every industry

Manufacturing
Defects per unit or per batch. Poisson when the line is stable; Negative Binomial when failures cluster around a drifting machine or shift.
Poisson → Negative Binomial
Insurance
Claims per policy per year — the textbook overdispersed count. Negative Binomial frequency × a severity model gives the loss.
Negative Binomial
Healthcare
ER arrivals per hour (Poisson); hospital-acquired infections per 1,000 patient-days (clustered → Negative Binomial).
Poisson → Negative Binomial
Public sector
311 service requests per day, or incidents per district — heavy clustering makes Negative Binomial the standard.
Negative Binomial
Transportation
Crashes per intersection per year — Negative Binomial is the accepted model in traffic safety.
Negative Binomial
Retail / e-commerce
Purchases per customer per month, or items per basket — overdispersed by loyal heavy buyers.
Negative Binomial
Web / telecom
Requests per second or failed logins per minute. Near-Poisson at steady state, bursty under load.
Poisson → Negative Binomial
Utilities
Outages or trouble tickets per feeder per month — weather clusters them beyond Poisson.
Negative Binomial
Agriculture
Pest or weed counts per plant/quadrat — aggregation in patches makes them overdispersed.
Negative Binomial
Sports
Goals per soccer match — the classic clean Poisson fit, the baseline every model is judged against.
Poisson
Finance
Defaults per portfolio per quarter, or trades per account — contagion pushes toward Negative Binomial.
Negative Binomial
Hospitality
No-shows per night or complaints per stay — small counts, occasional clusters.
Poisson → Negative Binomial

From distribution to model — bootstrapping, bagging, XGBoost

Poisson / Negative-Binomial regression (GLM)
The interpretable baseline: a log-link generalized linear model gives each driver a multiplicative effect on the rate. If a dispersion test fails (variance ≫ mean), switch the Poisson GLM for a Negative Binomial one — same coefficients to read, honest standard errors.
XGBoost with a count objective
For nonlinear interactions, gradient boosting with count:poisson (or reg:tweedie for zero-heavy counts) predicts a rate directly and captures structure a GLM misses. You trade some interpretability for fit — pair it with SHAP to win it back.
Bootstrapping the rate
Resample records with replacement, refit, and read the spread to get confidence intervals on a rate or a model metric without trusting a single distributional assumption — especially valuable when counts are messy and overdispersed.
Bagging / random forests
Averaging many trees over bootstrap resamples cuts the variance of a wobbly count model. It won't give you clean coefficients, but it's a sturdy, low-fuss predictor and a strong accuracy benchmark for the GLM to beat.

Activations & the distributions inside AI

Neural networks are stitched together from functions that turn raw scores into something useful — and most of them are a probability distribution in disguise. Drag the parameters; the curves are live.

Sigmoid (logistic)

(0, 1)
-661.077-0.08
f(x)
the stats insideIt is exactly the logistic distribution's CDF — which is why it turns a real-valued score into a probability, and why logistic regression and the Bernoulli likelihood are built on it.
used inBinary-classification heads; gates in LSTM/GRU cells.
probabilitylogistic cdfgatebernoulli

Tanh

(−1, 1)
-441.159-1.159
f(x)
the stats insideA rescaled, recentered sigmoid — tanh(x) = 2σ(2x) − 1. Zero-centered, which historically kept gradients healthier than the sigmoid.
used inClassic RNNs; bounded hidden units.
zero-centeredsigmoidbounded

Elliott (softsign)

(−1, 1)
-660.994-0.994
f(x)
the stats insideA rational stand-in for tanh — the same S-shape and (−1, 1) range with no exponential to compute, so it's cheap on constrained hardware. Also called softsign (David Elliott, 1993); its tails approach ±1 more slowly than tanh's.
used inLow-power / embedded inference where exp() is costly.
softsigntanh approxrationalcheap

ReLU

[0, ∞)
-444.32-0.32
f(x)
the stats insideNot a distribution itself, but it induces sparsity — roughly half its inputs map to exactly 0 — and a half-rectified output whose positive part is often modeled as half-normal or exponential.
used inThe default in most CNNs and MLPs.
sparsityrectifierdefault

Leaky ReLU

(−∞, ∞)
-444.352-0.752
f(x)
the stats insideKeeps a small gradient for negative inputs so neurons don't 'die' at zero. α is the leak.
used inGANs and deep CNNs where dead ReLUs hurt training.
leakdying relugradient

ELU

(−α, ∞)
-444.399-1.38
f(x)
the stats insideSmooth, saturating negative branch that nudges mean activations toward zero — a self-normalizing pressure (SELU is its scaled cousin).
used inDeep nets seeking stable activation statistics.
smoothmean-zeroselu

GELU

≈ (−0.17, ∞)
-444.333-0.504
f(x)
the stats insidex times the standard-Normal CDF Φ(x): each input is gated by the probability a Gaussian falls below it. The distribution is literally inside the activation.
used inTransformers — BERT, GPT, ViT.
gaussian cdftransformersmooth

SiLU / Swish

≈ (−0.28, ∞)
-555.386-0.698
f(x)
the stats insidex gated by a logistic sigmoid — a smooth, non-monotonic relative of GELU. β slides it from near-linear toward near-ReLU.
used inEfficientNet and many modern transformer variants.
swishgatesigmoidsmooth

Softplus

(0, ∞)
-444.34-0.321
f(x)
the stats insideA smooth ReLU whose derivative is exactly the sigmoid — the integral of the logistic curve. Used to force a parameter positive, e.g. a predicted variance σ or a Poisson rate λ.
used inPositive-output heads (predicted σ, rate λ).
smooth relupositivesigmoid derivative

Mish

≈ (−0.31, ∞)
-555.424-0.733
f(x)
the stats insideA smooth, non-monotonic self-gating activation in the Swish family.
used inSome vision backbones (e.g. YOLOv4).
smoothnon-monotonicself-gating

Distributions emerging in AI, CS & computational topology

Softmax = the Boltzmann/Gibbs distributionScores → a categorical distribution, with a temperature dial.

Softmax maps a vector of scores to probabilities pᵢ = e^(zᵢ/T) / Σ e^(zⱼ/T) — exactly the Gibbs/Boltzmann distribution from statistical physics. Temperature T flattens (high) or sharpens (low) the choice.

softmaxcategoricalboltzmanntemperature
Attention is a probability distributionEach head is a distribution over where to look.

An attention head softmaxes query·key scores into weights that sum to 1 — a distribution over positions. A transformer layer is a stack of these little distributions mixing values together.

attentiontransformersoftmax
Dropout = Bernoulli noiseA Bernoulli mask that behaves like bagging.

Dropout multiplies each activation by an independent Bernoulli(p) mask at train time — a stochastic ensemble that approximates bagging over exponentially many sub-networks.

dropoutbernoullibaggingregularization
Weight init = scaled NormalsXavier/He draw weights from variance-tuned Gaussians.

Glorot/Xavier and He initialization sample weights from a Normal (or uniform) whose variance is set by the layer's fan-in/out, so signal neither explodes nor vanishes with depth. The Gaussian is doing structural work.

initnormalxavierhevariance
Logits → Bernoulli via the logistic linkSigmoid + cross-entropy is the Bernoulli likelihood.

A binary classifier's logit passes through the sigmoid (the logistic CDF) to give P(y=1). Minimizing cross-entropy is exactly maximizing the Bernoulli log-likelihood.

sigmoidbernoullicross-entropylikelihood
Gaussian processesA Normal distribution over whole functions.

A GP places a joint Gaussian over function values: any finite set of points is multivariate Normal. It's the distribution-first view of regression — predictions arrive with calibrated uncertainty.

gaussian processnormalbayesianuncertainty
Temperature & sampling in LLMsGeneration samples from the softmax over tokens.

An LLM samples the next token from the softmax distribution over the vocabulary; temperature scales the logits, trading determinism (T→0, argmax) for diversity. Top-k / nucleus sampling truncate that distribution's tail.

temperaturesamplingllmsoftmax
Computational topology (TDA)Persistent homology turns shape into countable data.

Persistent homology summarizes data by the connected components and holes that persist across scales — Betti numbers and persistence diagrams. Those summaries are themselves data (counts and lifetimes) you can model statistically: the bridge from shape to inference.

tdapersistent homologytopologybetti

Modern AI — a working glossary

The architectures and moving parts of today's models — transformers, attention and its variants, the LLM stack. Plain definitions, kept honest.

TransformerAttention-only sequence model; the backbone of modern AI.

Introduced in “Attention Is All You Need” (2017), the transformer drops recurrence and convolution for stacked self-attention plus feed-forward layers. Because attention is fully parallel across positions, it trains at scale — the architecture under nearly every modern LLM, vision, and multimodal model.

transformerattentionarchitecture
Attention (Q, K, V)A weighted lookup: softmax of query·key, applied to values.

Each token forms a query and compares it to every token's key; the scaled dot products softmax into weights that mix the values — Attention(Q,K,V) = softmax(QKᵀ/√d)·V. Multi-head attention runs several such lookups in parallel subspaces and concatenates them.

attentionqkvmulti-headsoftmax
Architecture categoriesEncoder-only, decoder-only, encoder–decoder — plus the CNN/RNN/SSM families.

Transformers come in three shapes: encoder-only (BERT — bidirectional, for understanding/embeddings), decoder-only (GPT — autoregressive, for generation), and encoder–decoder (T5 — sequence-to-sequence). Alongside them live convolutional (CNN), recurrent (RNN/LSTM), and state-space (SSM/Mamba) families.

architectureencoderdecodertaxonomy
AgentsSystems that perceive, decide, and act toward a goal in a loop.

An agent senses its environment, chooses actions toward an objective, acts, observes, and repeats (plan → act → observe). LLM agents add tool use, memory, and multi-step autonomy. The rigor is decision-theoretic — acting well under uncertainty, not merely predicting the next token.

agenttool-usedecisionautonomy
RNN & LSTMSequence models with memory; LSTMs add gates against vanishing gradients.

A recurrent network carries a hidden state across time steps. Plain RNNs lose long-range context (vanishing gradients); the LSTM adds input/forget/output gates and a cell state to preserve information over long spans. Largely superseded by transformers for long range, still handy when compute or latency is tight.

rnnlstmsequencegates
CNNConvolutional nets — local filters with shared weights.

A CNN slides learned filters over grid-structured data (images, audio), exploiting locality and translation-equivariance with far fewer parameters than a dense net. Convolution is weight-sharing across position; pooling adds a measure of scale invariance.

cnnconvolutionvisionweight-sharing
GPTDecoder-only, autoregressive next-token prediction.

Generative Pre-trained Transformer: a decoder-only model trained to predict the next token given the previous ones. Scale plus this single objective yields broad capability; generation samples from the softmax over the vocabulary.

gptdecoderautoregressivellm
BERTEncoder-only, bidirectional, masked-language pretraining.

BERT reads the whole context at once (bidirectional) and pretrains by predicting masked-out tokens. Encoder-only, it excels at understanding and embedding tasks — classification, retrieval — rather than free-form generation.

bertencodermasked-lmembeddings
Encoder–decoderCompress to a representation, then generate from it — the information-theory pillars.

The encoder maps input to a latent representation; the decoder generates output from it. Read through information theory it is the whole pipeline — storing, coding, transmitting, de-noising, and extracting. BERT is encoder-only, GPT decoder-only, T5 the full encoder–decoder.

encoder-decoderseq2seqinformation-theoryt5
LLMA transformer trained at scale on next-token prediction.

A large language model is a (usually decoder-only) transformer trained on vast text to predict the next token. Scale brings emergent abilities — in-context learning, reasoning traces — and new failure modes: hallucination and calibration drift.

llmtransformerscalenext-token
Tokens & tokenizationModels read subword units, not characters or words.

Text is split into tokens — subword pieces from algorithms like BPE or WordPiece — each mapped to an integer id and an embedding. Token count drives context limits and cost, and odd tokenization explains many quirks (arithmetic, spelling).

tokensbpetokenizationcontext
Context & contextualizationThe window the model attends over, and embeddings that depend on it.

The context window is the span of tokens a model can attend to at once. Contextualization means a token's representation depends on its neighbors — the same word gets a different vector in a different sentence, unlike static (word2vec) embeddings.

contextwindowcontextual-embeddingattention
RAG (retrieval-augmented generation)Retrieve relevant documents, then condition generation on them.

RAG searches a corpus (usually by vector similarity) for passages relevant to a query, then feeds them to the model as grounding. It cuts hallucination and lets a frozen model use fresh or private knowledge without retraining.

ragretrievalvector-searchgrounding
Local AIModels that run on your own device or premises.

Running models locally (Ollama, llama.cpp, on-device) keeps data in your environment — no egress, lower latency, offline capability, privacy by construction. The same compute-to-data principle as keeping client data where it lives.

localon-deviceprivacyollama
Gated Linear Units (GLU)A learned gate multiplies one projection by an activation of another.

A GLU computes (xW) ⊗ σ(xV) — one linear path gated elementwise by a nonlinearity of another. The variants GeGLU and SwiGLU (GELU/Swish gates) reliably improve transformer feed-forward blocks.

gluswiglugatingffn
RoPE (Rotary Position Embedding)Encode position by rotating query/key vectors.

RoPE rotates pairs of query/key coordinates by an angle proportional to position, so the attention dot product depends only on relative position. It bakes in order without separate position vectors and extrapolates to longer contexts.

ropepositionrotationrelative
ALiBi (Attention with Linear Biases)Bias attention scores by distance — no position embeddings.

ALiBi adds a fixed linear penalty to attention scores proportional to the distance between tokens, favoring nearer ones. With no learned positional embeddings, it extrapolates gracefully to sequences far longer than seen in training.

alibipositionbiasextrapolation
Random Feature AttentionApproximate softmax attention in linear time.

Softmax attention costs O(n²) in sequence length. Random-feature methods (e.g., Performer's FAVOR+) approximate the softmax kernel with random features and factor it, so attention runs in linear time and memory — trading exactness for scale.

linear-attentionrandom-featuresperformerefficiency
Transfer learningPretrain broadly, then adapt to a specific task.

Learn general representations on a large corpus, then reuse them — full fine-tuning, adapters/LoRA, or prompting — for a downstream task with little data. The pretraining → adaptation paradigm is why one base model powers many applications.

transferpretrainingfine-tuninglora
Evaluation engineering for LLMsBuilding rigorous, honest measures of model quality.

Because benchmarks leak and single scores mislead, eval engineering designs held-out and contamination-controlled test sets, task-grounded metrics, rubrics, and LLM-as-judge with calibration — tracked over time. What you measure is what you ship; the discipline is making the measurement trustworthy.

evaluationbenchmarksllm-as-judgecalibration