The library
More than a statistics shelf. This is a hands-on library of the ideas the whole program runs on: a cross-subject glossary — random matrices, L-functions, topology, decision theory, and the contemplative thread, each term tagged by subject and project — alongside interactive probability distributions you can drag, an inference glossary, an AI glossary, and a which-distribution guide. Built to learn out loud, and the same toolkit behind the studio's models.
Ideas & terms — a one-stop glossary
The vocabulary the research program runs on, in one place: random matrix theory, L-functions and number theory, topological data analysis, statistics, probability, decision theory, and the contemplative thread. Every term is tagged by subject and by the project it earns its keep in — filter by either, or search the lot.
32 of 32 terms
Analytic rank & central valueL(½) and its order of vanishing there.
The central value L(½) (or L(E,1) for an elliptic curve) and the order to which the L-function vanishes at the centre — the analytic rank. Whether it vanishes is exactly what the family's symmetry type predicts.
Automorphic L-functionAn L-function attached to an automorphic representation (Langlands).
The Langlands program conjectures every L-function arises from an automorphic representation of GL(n). Dirichlet, Hecke, and modular-form L-functions are the low-rank cases.
Betti numbersCounts of holes by dimension: H₀ components, H₁ loops, H₂ voids.
βₖ = dim Hₖ(X) = dim(ker ∂ₖ / im ∂ₖ₊₁). β₀ counts connected pieces, β₁ independent loops, and so on — the coarse topological fingerprint persistence tracks across scale.
Birch–Swinnerton-Dyer (BSD)The rank of an elliptic curve equals the order of vanishing of its L-function at s = 1.
A Clay Millennium problem linking arithmetic (the rank of E(ℚ)) to analysis (vanishing of L(E,s) at the centre). Known for analytic rank ≤ 1 (Gross–Zagier, Kolyvagin). The reason 'vanishing' and 'spacing' are the same question seen twice.
Bottleneck & Wasserstein distanceMetrics between persistence diagrams; the basis of stability.
Optimal-matching distances between diagrams. The stability theorem says small data perturbations move the diagram only a little in bottleneck distance — what makes persistence a trustworthy statistic.
CalibrationBeliefs that match frequencies: of the things you call 90% likely, 90% happen.
The honest target for a probabilistic agent and the heart of the QBist reading of P2 — a decision-maker's credences are good insofar as they are calibrated, not insofar as they feel confident.
Contextual fractionA sheaf-cohomology measure of how non-classical a system of observations is.
Abramsky–Brandenburger model contextuality as the failure of local data to glue into a global section; the contextual fraction quantifies it as a dimensionless effect size — the bridge between QBism, sheaf theory, and P9's conjectures.
Effect size vs p-valueHow big versus how surprising — report both.
A p-value measures incompatibility with the null; an effect size measures magnitude with units and a confidence interval. A significant tiny effect (the P9 ρ ≈ −0.22) can be real yet useless — which is why the program reports effect sizes, not just p.
Effective dimensionalityHow many directions in your data are real signal.
The count of eigenvalues above the Marchenko–Pastur edge — a principled, scale-free alternative to scree plots and the Kaiser rule. In real neural and galaxy-spectra data this came out near ~12, stably.
Eigenvalue clipping & shrinkageDenoising a covariance by treating sub-edge eigenvalues as noise.
Replace the bulk (below the MP edge) eigenvalues with their average (clipping) or shrink toward a target (Ledoit–Wolf). Stays well-conditioned even when samples are scarcer than dimensions — where the raw covariance is singular and the Hartlap correction is undefined (P10/S3).
GUE / GOE / GSE ensemblesThe three classical random-matrix ensembles (β = 2, 1, 4).
Gaussian Unitary (complex Hermitian, β=2), Orthogonal (real symmetric, β=1), and Symplectic (β=4) ensembles. They set the universal local spacing statistics: a real data covariance is GOE; the Riemann zeros behave like GUE.
Katz–Sarnak symmetry typesFamilies of L-functions show unitary, orthogonal, or symplectic statistics.
Ordered by conductor, a family's low-lying zeros match one of three ensembles (U / O / Sp). Proven over function fields (via monodromy), conjectural over ℚ. The symmetry type predicts central vanishing — the natural extension of P10's spectral zoo into arithmetic.
L-functionA Dirichlet series with an Euler product and a functional equation; ζ is the prototype.
An analytic object encoding arithmetic, of the form Σ aₙ/nˢ, with an Euler product over primes and a symmetry s ↔ 1−s. Its zeros control deep arithmetic; their statistics are conjecturally governed by random-matrix ensembles.
Level repulsionGeneric spectra avoid crowding; nearby levels push apart.
The signature of random-matrix (and quantum-chaotic) spectra: the probability of a tiny gap vanishes. Its opposite is level clustering (Poisson, independent levels) and, at the far extreme, a rigid comb (equal gaps).
LMFDBThe L-functions and Modular Forms Database — the field's one-stop data source.
An open atlas of L-functions, modular forms, elliptic curves, and their zeros — the practical place to pull family data for spacing / low-lying-zero experiments (with Sage / pari for computation).
Marchenko–Pastur lawThe noise band of sample-covariance eigenvalues; upper edge (1+√(p/n))².
For p variables and n samples of pure noise, the sample correlation eigenvalues fill [(1−√(p/n))², (1+√(p/n))²]. Eigenvalues above the upper edge cannot be explained by noise — they are candidate signal. This edge is the factor-selection decision instrument in P10.
Method of lociThe memory palace: bind ideas to places you can walk.
An ancient mnemonic — encode an ordered argument as a route through a remembered space. Used here as a study aid for derivations, and a small example of the contemplative-meets-computational practice the program runs on.
Montgomery–Dyson phenomenonThe spacings of the Riemann zeros match GUE eigenvalue statistics.
Montgomery (1973) computed the pair correlation of ζ-zeros; Dyson recognised the GUE form. P10 reproduces it with the spacing ratio: the zeros give ⟨r⟩ ≈ 0.60, the unitary corner of the spectral zoo.
Motivic L-functionAn L-function attached to a motive / Galois representation.
Built from arithmetic geometry — elliptic curves, modular forms, symmetric powers. These carry the deepest arithmetic invariants (ranks, special values) and are where vanishing meets BSD.
Out-of-sample validationJudge a method on data it didn't see — the honest test.
Cross-validation / held-out testing: fit on one split, score on another. The whole P10 'earn its place' discipline rests on it — an in-sample correlation that vanishes out-of-sample (as in the P9 wedge) is not a finding.
OverfittingFitting noise as if it were signal; great in-sample, poor out.
When a model has enough freedom to memorise quirks of the training data. Why keeping too many (noise) dimensions hurts — and why the Marchenko–Pastur edge's parsimony pays off when samples are scarce.
Persistence diagram / barcodeThe birth–death record of topological features.
Each feature is a point (birth, death) — or a bar of that length. Distance from the diagonal = persistence = how 'real' the feature is. The object every TDA summary statistic is computed from.
Persistence entropyA single number summarising how spread-out a diagram's lifetimes are.
Shannon entropy of the normalised persistence lifetimes — a scale-free scalar feature used to compare diagrams across objects (e.g. the khipu signatures in P9).
Persistent homologyMultiscale topology: which holes in a point cloud survive across scale.
Build a growing family of complexes from data and track when topological features (components, loops, voids) are born and die. Long-lived features are real structure; short-lived ones are noise — a topological echo of the signal/noise split.
Process-relational / interbeingReality as relations and processes (morphisms), not substances (objects).
The shared substrate of P9: Whitehead's process philosophy, category theory's primacy of morphisms, QBism's experience-first reading, and the contemplative 'interbeing' of Thich Nhat Hanh — read here as a stance you can try to make falsifiable, not a slogan.
Random matrix theory (RMT)The statistics of eigenvalues of large random matrices — and their surprising universality.
RMT studies how the eigenvalues of large random matrices behave. Two facts power everything else: the bulk follows a fixed law (Marchenko–Pastur for covariance), and the local spacing statistics are universal — depending only on a symmetry class, not the matrix's details.
Riemann zeta & the Riemann Hypothesisζ(s) = Σ 1/nˢ; RH says its nontrivial zeros lie on Re(s) = ½.
The prototype L-function. The Riemann Hypothesis — all nontrivial zeros on the critical line — is the central open problem; its truth controls the error term in the prime-counting function.
Spacing ratio ⟨r⟩An unfolding-free statistic for level repulsion (Atas et al.).
r = min(sₙ,sₙ₊₁)/max(sₙ,sₙ₊₁) on consecutive gaps; its mean is density-independent so it needs no 'unfolding'. Benchmarks: Poisson 0.386, GOE 0.536, GUE 0.600, rigid comb → 1. The backbone of P10's spectral-zoo comparisons.
Tracy–Widom distributionThe fluctuation law of the largest eigenvalue at the spectral edge.
How far the top eigenvalue strays beyond the Marchenko–Pastur edge — the rigorous basis for deciding whether a borderline eigenvalue is signal or an edge fluctuation of noise.
Vanishing & non-vanishingWhether L(½) = 0 — tied to rank and to symmetry type.
Orthogonal families have a positive density of central zeros (extra vanishing); symplectic ones repel zeros from the centre; unitary ones are generic. Non-vanishing theorems (Iwaniec–Sarnak, Soundararajan) bound how often L(½) ≠ 0.
Vietoris–Rips filtrationThe standard growing complex built from a point cloud.
Connect points within distance ε and fill in simplices; grow ε from 0 to ∞. The resulting nested complexes are what persistent homology is computed on.
Wigner surmiseThe level-spacing law showing repulsion: P(s) → 0 as s → 0.
An accurate closed form for the nearest-neighbour spacing distribution of RMT eigenvalues — e.g. GUE: P(s) ≈ (32/π²) s² e^(−4s²/π). The s² prefactor is level repulsion: levels almost never coincide, unlike independent (Poisson) levels.
Distributions & inference, hands-on
The interactive distribution catalog, the inference glossary, the bootstrap demo, and the which-distribution guide.
Distributions
Toggle PMF/PDF ↔ CDF and drag the sliders — every chart recomputes live.
Bernoulli
discreteA single yes/no trial — the atom every other count is built from.
Binomial
discreteHow many successes in a fixed number of independent yes/no trials.
Geometric
discreteHow many failures before the first success — the waiting game, in trials.
Negative Binomial
discreteCounts with extra spread — the realistic cousin of Poisson when variance beats the mean.
Poisson
discreteHow many events happen in a fixed window when they arrive at a steady average rate.
Discrete Uniform
discreteEvery outcome in a range equally likely — the fair die.
Hypergeometric
discreteSuccesses when you sample without replacement — Binomial's finite-population cousin.
Continuous Uniform
continuousEqual density across an interval — flat ignorance between two bounds.
Normal (Gaussian)
continuousThe bell curve — what sums of many small independent effects converge to.
Exponential
continuousThe waiting time until the next event when events arrive at a steady rate.
Gamma
continuousWaiting time for several events to accumulate — positive, skewed, flexible.
Beta
continuousA distribution over a probability itself — your belief about a rate between 0 and 1.
Log-normal
continuousWhen the logarithm is Normal — positive, heavy-tailed, born from multiplying effects.
Weibull
continuousThe reliability workhorse — failure times whose hazard can fall, hold, or climb.
Chi-squared
continuousSum of squared standard Normals — the engine behind variance and goodness-of-fit tests.
Student's t
continuousThe bell curve with heavier tails — what you use when the variance is estimated, not known.
Compare two distributions
Overlay any two to see the shape difference — start with Poisson vs Negative Binomial (same mean, fatter tail).
Foundations — the inference glossary
The ideas the distributions sit inside. The same search box filters these.
Expected valueThe probability-weighted long-run average.
E[X] = Σ k·P(k) for discrete, ∫ x·f(x) dx for continuous. It's the balance point of the distribution — and need not be a value X can actually take (a die's mean is 3.5).
Variance & standard deviationHow far outcomes spread from the mean.
Var(X) = E[(X−μ)²]; the standard deviation is its square root, back in the original units. Variance adds for independent variables, which is why it shows up everywhere.
SupportThe set of outcomes a distribution actually puts mass on.
The support is where the PMF/PDF is positive — the values that can actually occur: supp(f) = closure of { x : f(x) > 0 }. It bounds where estimates and integrals live, and mismatched supports (a model assigning probability 0 to data you observed) quietly break likelihoods and importance sampling.
Probability vs likelihoodSame formula, opposite thing held fixed.
Probability fixes the parameters and asks about data: P(data | θ), summing to 1 over data. Likelihood fixes the observed data and varies the parameters: L(θ) = P(data | θ) read as a function of θ — it is NOT a probability distribution over θ.
Probability distributionThe full map from outcomes to their probabilities.
A complete description of a random quantity: every outcome and how much probability (or density) it carries. Everything else — mean, variance, tail risk — is read off from it.
PMF vs PDFMass for discrete, density for continuous.
A PMF gives P(X=k) directly and sums to 1. A PDF gives density, where probability is area under the curve; for continuous X, P(X = any exact point) is 0, so only intervals carry probability.
Likelihood function & MLEPick the parameters that make your data least surprising.
The likelihood L(θ) = ∏ f(xᵢ; θ). Maximum likelihood estimation chooses the θ that maximizes it — usually via the log-likelihood, which turns the product into a sum and the optimization into calculus.
Estimator vs estimateThe rule versus the number it produces.
An estimator is a function of the data — itself a random variable with its own distribution (e.g. the sample mean). An estimate is the single value that rule returns for one dataset. Properties like bias and variance belong to the estimator.
BiasHow far an estimator is wrong on average.
Bias = E[estimator] − true value. Unbiased means zero on average across samples — but a slightly biased estimator with much lower variance is often the better bet (the bias–variance trade-off).
ConsistencyIt converges to the truth as data grow.
An estimator is consistent if it converges in probability to the true parameter as n → ∞. More data should pin it down; an inconsistent estimator never settles on the right answer no matter how much you collect.
SufficiencyA summary that loses nothing about the parameter.
A statistic is sufficient for θ if the data carry no extra information about θ once you know it (Fisher–Neyman factorization). The sample sum is sufficient for a Poisson rate — you can throw away the rest.
EfficiencyThe lowest-variance estimator you can get.
Among unbiased estimators, the efficient one has the smallest variance — bounded below by the Cramér–Rao limit. Efficiency is why we prefer one valid estimator over another.
Hypothesis testingAssume nothing's going on, then see if the data argue otherwise.
A decision procedure: take H₀ as true, compute how surprising the data would be under it, and reject H₀ only if that surprise crosses a pre-set threshold. It controls error rates, not truth.
Null & alternative hypothesesThe default vs what you'd conclude instead.
H₀ is the no-effect / status-quo claim you try to falsify; H₁ is the alternative you'd accept if the evidence is strong enough. You never 'prove' H₀ — you only fail to reject it.
Test statisticA number whose behavior under H₀ you know.
A summary of the data (z, t, χ², F) whose sampling distribution is known when H₀ holds. You locate your observed value in that distribution to judge how extreme it is.
p-valueHow surprising the data are if H₀ is true.
The probability, assuming H₀, of a test statistic at least as extreme as the one observed. It is NOT the probability that H₀ is true, nor the chance your result was luck — just a measure of incompatibility with H₀.
Significance level αThe false-positive rate you agree to tolerate.
Chosen before you look (often 0.05): the probability of rejecting a true H₀. Reject when p < α. It's a budget for Type-I errors, not a law of nature — pick it for the decision's stakes.
Type I & Type II errorsFalse alarm vs missed signal.
Type I: rejecting a true H₀ (a false positive), rate α. Type II: failing to reject a false H₀ (a missed effect), rate β. Tightening one usually loosens the other for fixed sample size.
PowerThe chance of catching a real effect.
Power = 1 − β: the probability of correctly rejecting H₀ when H₁ is true. It rises with sample size, effect size, and lower noise — the thing a good experiment is designed around.
Confidence intervalA procedure that brackets the truth most of the time.
A 95% CI is built by a method that, across repeated samples, contains the true parameter 95% of the time. The frequentist subtlety: it's a property of the procedure, not a 95% probability about this one interval.
Central Limit TheoremAverages tend to a bell curve.
Sums and means of many independent, finite-variance variables approach a Normal distribution regardless of the original shape. It's why so much of inference leans on the Normal even when the data aren't.
Law of Large NumbersSample averages settle on the expected value.
As n grows, the sample mean converges to E[X]. It's the guarantee that more data, gathered honestly, pulls your estimate toward the truth — the backbone of simulation and bootstrapping.
BootstrappingResample your data to feel out its uncertainty.
Resample the dataset with replacement many times, recompute the statistic each time, and use the spread of results as an approximate sampling distribution — confidence intervals with almost no distributional assumptions.
BaggingAverage many models trained on resamples to cut variance.
Bootstrap AGGregating: fit a model on each bootstrap resample and average (or vote) their predictions. It tames variance in unstable learners — random forests are bagging plus feature subsampling.
Bayesian vs frequentistBeliefs you update vs long-run frequencies.
Frequentists treat parameters as fixed unknowns and reason about repeated sampling. Bayesians put a probability distribution on the parameter — a prior updated by data into a posterior. Different questions, both useful; the Beta–Binomial pair is the classic bridge.
Conjugate prior (Beta–Binomial)A prior whose update keeps the same shape.
A Beta prior on a probability, combined with Binomial data, yields a Beta posterior — just add successes to α and failures to β. That tidy update is why Beta is the natural language for an uncertain rate.
Bootstrapping, live
A confidence interval built by resampling — no formula, no distributional assumption.
One fixed sample of 40 draws from a skewed (Exponential) population — we never see the rest of the world. Each resample draws 40 of those values with replacement and takes the mean; the teal histogram is the bootstrap distribution of the mean, and the dashed amber lines are the 2.5th/97.5th percentiles — a 95% confidence interval built with almost no assumptions. Slide B up and watch it steady (the Law of Large Numbers at work); the white line is the original sample mean.
Which distribution? — a click-through guide
Start from the question you're actually asking, then click a suggestion to jump to its card.
Deep dive — counting things (Poisson & Negative Binomial)
Counts are the most common data shape in real operations — events per window, per unit, per place. Poisson is the honest starting point (one rate, mean = variance), but real counts almost always spread wider than Poisson allows: a few busy days, a clustered outbreak, a bad batch. That overdispersion is exactly what the Negative Binomial absorbs (it's a Poisson whose rate itself varies). Everything below is a count problem — and where the variance beats the mean, reach for Negative Binomial.
Negative Binomial
discreteCounts with extra spread — the realistic cousin of Poisson when variance beats the mean.
Poisson
discreteHow many events happen in a fixed window when they arrive at a steady average rate.
Tweedie — the pure-premium bridge
Frequency (Poisson) × severity (Gamma) in a single model. Where the counts story meets the dollars.
Tweedie
simulatedOne family spanning Poisson (p=1), Gamma (p=2), and everything between — with a real point mass at zero.
The violet bar is the mass at exactly zero (no claim); the amber tail is the cost when something happens. Push p→1 and it leans Poisson (more, smaller events); p→2 and it leans Gamma (fewer, larger). λ ≈ 3.162 events expected per draw.
An example in every industry
From distribution to model — bootstrapping, bagging, XGBoost
Activations & the distributions inside AI
Neural networks are stitched together from functions that turn raw scores into something useful — and most of them are a probability distribution in disguise. Drag the parameters; the curves are live.
Sigmoid (logistic)
(0, 1)Tanh
(−1, 1)Elliott (softsign)
(−1, 1)ReLU
[0, ∞)Leaky ReLU
(−∞, ∞)ELU
(−α, ∞)GELU
≈ (−0.17, ∞)SiLU / Swish
≈ (−0.28, ∞)Softplus
(0, ∞)Mish
≈ (−0.31, ∞)Distributions emerging in AI, CS & computational topology
Softmax = the Boltzmann/Gibbs distributionScores → a categorical distribution, with a temperature dial.
Softmax maps a vector of scores to probabilities pᵢ = e^(zᵢ/T) / Σ e^(zⱼ/T) — exactly the Gibbs/Boltzmann distribution from statistical physics. Temperature T flattens (high) or sharpens (low) the choice.
Attention is a probability distributionEach head is a distribution over where to look.
An attention head softmaxes query·key scores into weights that sum to 1 — a distribution over positions. A transformer layer is a stack of these little distributions mixing values together.
Dropout = Bernoulli noiseA Bernoulli mask that behaves like bagging.
Dropout multiplies each activation by an independent Bernoulli(p) mask at train time — a stochastic ensemble that approximates bagging over exponentially many sub-networks.
Weight init = scaled NormalsXavier/He draw weights from variance-tuned Gaussians.
Glorot/Xavier and He initialization sample weights from a Normal (or uniform) whose variance is set by the layer's fan-in/out, so signal neither explodes nor vanishes with depth. The Gaussian is doing structural work.
Logits → Bernoulli via the logistic linkSigmoid + cross-entropy is the Bernoulli likelihood.
A binary classifier's logit passes through the sigmoid (the logistic CDF) to give P(y=1). Minimizing cross-entropy is exactly maximizing the Bernoulli log-likelihood.
Gaussian processesA Normal distribution over whole functions.
A GP places a joint Gaussian over function values: any finite set of points is multivariate Normal. It's the distribution-first view of regression — predictions arrive with calibrated uncertainty.
Temperature & sampling in LLMsGeneration samples from the softmax over tokens.
An LLM samples the next token from the softmax distribution over the vocabulary; temperature scales the logits, trading determinism (T→0, argmax) for diversity. Top-k / nucleus sampling truncate that distribution's tail.
Computational topology (TDA)Persistent homology turns shape into countable data.
Persistent homology summarizes data by the connected components and holes that persist across scales — Betti numbers and persistence diagrams. Those summaries are themselves data (counts and lifetimes) you can model statistically: the bridge from shape to inference.
Modern AI — a working glossary
The architectures and moving parts of today's models — transformers, attention and its variants, the LLM stack. Plain definitions, kept honest.
TransformerAttention-only sequence model; the backbone of modern AI.
Introduced in “Attention Is All You Need” (2017), the transformer drops recurrence and convolution for stacked self-attention plus feed-forward layers. Because attention is fully parallel across positions, it trains at scale — the architecture under nearly every modern LLM, vision, and multimodal model.
Attention (Q, K, V)A weighted lookup: softmax of query·key, applied to values.
Each token forms a query and compares it to every token's key; the scaled dot products softmax into weights that mix the values — Attention(Q,K,V) = softmax(QKᵀ/√d)·V. Multi-head attention runs several such lookups in parallel subspaces and concatenates them.
Architecture categoriesEncoder-only, decoder-only, encoder–decoder — plus the CNN/RNN/SSM families.
Transformers come in three shapes: encoder-only (BERT — bidirectional, for understanding/embeddings), decoder-only (GPT — autoregressive, for generation), and encoder–decoder (T5 — sequence-to-sequence). Alongside them live convolutional (CNN), recurrent (RNN/LSTM), and state-space (SSM/Mamba) families.
AgentsSystems that perceive, decide, and act toward a goal in a loop.
An agent senses its environment, chooses actions toward an objective, acts, observes, and repeats (plan → act → observe). LLM agents add tool use, memory, and multi-step autonomy. The rigor is decision-theoretic — acting well under uncertainty, not merely predicting the next token.
RNN & LSTMSequence models with memory; LSTMs add gates against vanishing gradients.
A recurrent network carries a hidden state across time steps. Plain RNNs lose long-range context (vanishing gradients); the LSTM adds input/forget/output gates and a cell state to preserve information over long spans. Largely superseded by transformers for long range, still handy when compute or latency is tight.
CNNConvolutional nets — local filters with shared weights.
A CNN slides learned filters over grid-structured data (images, audio), exploiting locality and translation-equivariance with far fewer parameters than a dense net. Convolution is weight-sharing across position; pooling adds a measure of scale invariance.
GPTDecoder-only, autoregressive next-token prediction.
Generative Pre-trained Transformer: a decoder-only model trained to predict the next token given the previous ones. Scale plus this single objective yields broad capability; generation samples from the softmax over the vocabulary.
BERTEncoder-only, bidirectional, masked-language pretraining.
BERT reads the whole context at once (bidirectional) and pretrains by predicting masked-out tokens. Encoder-only, it excels at understanding and embedding tasks — classification, retrieval — rather than free-form generation.
Encoder–decoderCompress to a representation, then generate from it — the information-theory pillars.
The encoder maps input to a latent representation; the decoder generates output from it. Read through information theory it is the whole pipeline — storing, coding, transmitting, de-noising, and extracting. BERT is encoder-only, GPT decoder-only, T5 the full encoder–decoder.
LLMA transformer trained at scale on next-token prediction.
A large language model is a (usually decoder-only) transformer trained on vast text to predict the next token. Scale brings emergent abilities — in-context learning, reasoning traces — and new failure modes: hallucination and calibration drift.
Tokens & tokenizationModels read subword units, not characters or words.
Text is split into tokens — subword pieces from algorithms like BPE or WordPiece — each mapped to an integer id and an embedding. Token count drives context limits and cost, and odd tokenization explains many quirks (arithmetic, spelling).
Context & contextualizationThe window the model attends over, and embeddings that depend on it.
The context window is the span of tokens a model can attend to at once. Contextualization means a token's representation depends on its neighbors — the same word gets a different vector in a different sentence, unlike static (word2vec) embeddings.
RAG (retrieval-augmented generation)Retrieve relevant documents, then condition generation on them.
RAG searches a corpus (usually by vector similarity) for passages relevant to a query, then feeds them to the model as grounding. It cuts hallucination and lets a frozen model use fresh or private knowledge without retraining.
Local AIModels that run on your own device or premises.
Running models locally (Ollama, llama.cpp, on-device) keeps data in your environment — no egress, lower latency, offline capability, privacy by construction. The same compute-to-data principle as keeping client data where it lives.
Gated Linear Units (GLU)A learned gate multiplies one projection by an activation of another.
A GLU computes (xW) ⊗ σ(xV) — one linear path gated elementwise by a nonlinearity of another. The variants GeGLU and SwiGLU (GELU/Swish gates) reliably improve transformer feed-forward blocks.
RoPE (Rotary Position Embedding)Encode position by rotating query/key vectors.
RoPE rotates pairs of query/key coordinates by an angle proportional to position, so the attention dot product depends only on relative position. It bakes in order without separate position vectors and extrapolates to longer contexts.
ALiBi (Attention with Linear Biases)Bias attention scores by distance — no position embeddings.
ALiBi adds a fixed linear penalty to attention scores proportional to the distance between tokens, favoring nearer ones. With no learned positional embeddings, it extrapolates gracefully to sequences far longer than seen in training.
Random Feature AttentionApproximate softmax attention in linear time.
Softmax attention costs O(n²) in sequence length. Random-feature methods (e.g., Performer's FAVOR+) approximate the softmax kernel with random features and factor it, so attention runs in linear time and memory — trading exactness for scale.
Transfer learningPretrain broadly, then adapt to a specific task.
Learn general representations on a large corpus, then reuse them — full fine-tuning, adapters/LoRA, or prompting — for a downstream task with little data. The pretraining → adaptation paradigm is why one base model powers many applications.
Evaluation engineering for LLMsBuilding rigorous, honest measures of model quality.
Because benchmarks leak and single scores mislead, eval engineering designs held-out and contamination-controlled test sets, task-grounded metrics, rubrics, and LLM-as-judge with calibration — tracked over time. What you measure is what you ship; the discipline is making the measurement trustworthy.