\[\begin{align} M'(t) &= \frac{d}{dt} \int e^{tx}f(x)\,\textrm{d}x = \int \frac{d}{dt} e^{tx}f(x)\,\textrm{d}x = \int xe^{tx}f(x)\,\textrm{d}x \\ & \Rightarrow M´(0) = \int xf(x)\,\textrm{d}x = \textrm{E}(x). \qquad \textrm{By induction } \textrm{E}(X^p) = M^{(p)}(0). \end{align}\]
A standard empirical task is to “choose” a distribution that “fits” some data. This is not entirely satisfactory from a theoretical perspective.
A parametric model is a deliberate restriction \[ \{F_\theta : \theta \in \Theta \subset \mathbb{R}^k\}, \] a collection of distributions indexed by a (possibly vector-valued) parameter \(\theta\).
The trade-off:
We will build a small set of families from first principles: start with simple mechanisms (trials, counting, waiting time), apply sums, limits and transformations, and arrive at the standard parametric families.
Discrete distributions
Continuous time
Continuous measurements
Constrained supports
Consider a random experiment where
This is a Bernoulli trial. The pmf of \(X\) with outcomes \(\{0,1\}\) and success probability \(p\) is \[ f(x; p) = P(X = x) = p^x(1-p)^{1-x}. \]
One parameter, \(p \in [0,1] = \Theta \subset \mathbb{R}\). We write \(X \sim \text{Bernoulli}(p)\).
Discussion: calculate \(E(X) = \qquad\) and \(\text{Var}(X) = \qquad\)
Let \(X_1, \ldots, X_n \overset{\text{iid}}{\sim} \text{Bernoulli}(p)\) and let \(S_n = \sum_{i=1}^n X_i\) be the number of successes.
By the binomial theorem, the moment generating function is \[ M(t) = \bigl((1-p) + pe^t\bigr)^n = \bigl[M_{\text{Bernoulli}}(t)\bigr]^n. \]
Moment generating functions of sums are products of moment generating functions.
Derivation of the mgf, and \(E(X)\), Var\((X)\) from it, see notes, The binomial distribution.
Let \(X\) be the number of events in a section of time or space. Think of the section as many very small units (\(n\) big), each a Bernoulli trial with \(p\) small.
Formally, let \(g(x, h)\) be the probability of \(x\) events in an interval of length \(h\). If
then \[ P(X = x) = \frac{\lambda^x e^{-\lambda}}{x!}, \quad x = 0, 1, 2, \ldots, \qquad X \sim \text{Poisson}(\lambda). \]
Proof of both, via moment generating functions, see notes, Poisson.
Independent Bernoulli\((p)\) trials. Let \(T\) be the number of trials until the \(r\)-th success: we need \(r-1\) successes in the first \(k-1\) trials, then a success on the \(k\)-th, so \[ P(T = k) = \binom{k-1}{r-1} p^r (1-p)^{k-r}, \quad k = r, r+1, \ldots \] This is the negative binomial.
Both derivations, and the precise \(r \to \infty\) statement, see notes, Geometric and Negative binomial.
Let \(W\) be the time until the first event in a Poisson process with rate \(\lambda\). Taking \(x = 0\) in the Poisson pmf gives \(P(W > w) = e^{-\lambda w}\), so \[ g(w) = \lambda e^{-\lambda w}, \quad 0 < w < \infty, \qquad W \sim \text{Exp}(\lambda). \]
The continuous twin of the geometric, and like it memoryless. For \(s > t\), so that \(\{W > s\} \subset \{W > t\}\), \[ P(W > s \mid W > t) = \frac{P(W > s)}{P(W > t)} = \frac{e^{-\lambda s}}{e^{-\lambda t}} = e^{-\lambda(s-t)} = P(W > s - t). \]
How long you have already waited tells you nothing about how much longer you will wait.
\(E(W) = 1/\lambda\) and Var\((W) = 1/\lambda^2\), see notes, Exponential distribution.
The same argument for the \(k\)-th event in a Poisson process with rate \(\lambda\) gives the density \[ g(w) = \frac{\lambda^k w^{k-1} e^{-\lambda w}}{\Gamma(k)}, \quad 0 < w < \infty, \] where \(\Gamma(k) = \int_0^\infty y^{k-1} e^{-y}\,dy\) is the gamma function.
The common parameterization uses \(\alpha = k\) and \(\beta = 1/\lambda\): \(X \sim \text{Gamma}(\alpha, \beta)\), with \(E(X) = \alpha\beta\) and \(\text{Var}(X) = \alpha\beta^2\).
Density and moments: notes, Gamma distribution.
| one event | several events | |
|---|---|---|
| discrete trials | geometric | negative binomial |
| continuous time | exponential | gamma |
We will later prove that sums of random variables tend to a common distribution (almost) regardless of the distribution of the components — the Central Limit Theorem. That limit is the normal distribution, with density \[ f(x) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\!\left\{-\frac{(x-\mu)^2}{2\sigma^2}\right\}. \]
A feature of nature: “additive noise” is often the sum of many small influences, so the CLT makes it approximately normal.
A tool in inference: the statistics we compute are sums over observations, so the CLT makes those approximately normal — whatever the data look like.
\(\chi^2\): if \(Z_1, \ldots, Z_\nu \overset{\text{iid}}{\sim} N(0,1)\), then \(Q = \sum_{i=1}^\nu Z_i^2 \sim \chi^2_\nu\), with \(\nu\) degrees of freedom. The empirical variance of a sample is approximately \(\chi^2\).
(Is the \(\chi^2\) just a special case of the Gamma?)
Student \(t\): if \(Z \sim N(0,1)\), \(U \sim \chi^2_\nu\), independent, then \(T = Z/\sqrt{U/\nu}\) is \(t\)-distributed with \(\nu\) degrees of freedom. The \(t\)-statistic \((\bar{X} - \mu)/s\) is approximately \(t\).
\(F\): if \(U_1 \sim \chi^2_{\nu_1}\) and \(U_2 \sim \chi^2_{\nu_2}\) are independent, then \(F = \dfrac{U_1/\nu_1}{U_2/\nu_2}\) is \(F\)-distributed with \((\nu_1, \nu_2)\) degrees of freedom. The ratio of two empirical variances is approximately \(F\).
Beta: defined on \([0,1]\), so the natural model for proportions. If \(X_1 \sim \text{Gamma}(\alpha, \theta)\) and \(X_2 \sim \text{Gamma}(\beta, \theta)\) are independent — same scale, different shapes — then \(X_1/(X_1+X_2) \sim \text{Beta}(\alpha, \beta)\).
Lognormal: \(\exp(Z)\) for \(Z \sim N(\mu, \sigma^2)\). Additive noise becomes multiplicative: growth and compounding.
Uniform: \(U \sim \text{Uniform}(0,1)\) generates all the others — \(F^{-1}(U)\) has distribution \(F\), and \(F(X)\) is uniform.
The exponential family: a large class containing most of the above, for which general results can be proved once instead of case by case: \[ f(x; \theta) = h(x)\, c(\theta) \exp\!\left\{\sum_{j=1}^k w_j(\theta)\, t_j(x)\right\}. \]
The Beta construction and its density, see notes, Beta.
Create an overview with the following properties for each family:
Complement this with relations between the distributions: