2 · Distributions of Random Variables
Parametric families of distributions
We will build a small set of distribution families that cover basic empirical work.
In practice, a relatively standard empirical task is to “choose” a distribution that “fits” some data.
This is, however, not entirely satisfactory from a theoretical perspective.
We will here
- start with simple mechanisms (trials, counting, waiting time),
- use mathematical operations and arguments (sums, limits, transformations …), and
- end up with the standard parametric families.
Why parametric families exist at all
The true distribution of real data is rarely known except in very particular cases.
A parametric model is a restriction \(\{F_\theta : \theta \in \Theta \subset \mathbb{R}^k\}\), that is, we consider a collection of distributions that is indexed by some (possibly vector valued) parameter \(\theta\).
This is a deliberate trade-off that we often want to make:
- Fewer degrees of freedom \(\Rightarrow\) easier to estimate and to interpret.
- But mis-specification is possible (and non-trivial).
We will here see how some common distributions arise from first principles, and later introduce estimation, sampling distributions and statistical inference.
A roadmap
- Discrete distributions
- Bernoulli \(\to\) Binomial (sum of trials)
- Binomial \(\to\) Poisson (rare-event limit)
- Bernoulli trials \(\to\) Geometric / Negative binomial (waiting time)
- Continuous time
- Poisson process \(\to\) Exponential (first waiting time)
- Exponential sums \(\to\) Gamma (waiting time to multiple events)
- Continuous measurements
- Normal as the standard “additive noise” model (CLT proof comes later)
- \(\chi^2\), \(t\), \(F\) distributions derived from the normal distribution (used for inference later).
- Constrained supports
- Beta for proportions on \([0,1]\)
- Lognormal for positive multiplicative variables
Bernoulli: the atom of models
Consider a random experiment where
- the sample space contains exactly two elements, and
- the two outcomes happen with constant probabilities \(p\) and (by property 1 in the last lecture) \(1-p\), respectively.
This experiment is known as a Bernoulli trial.
The pmf of a random variable \(X\) resulting from a Bernoulli trial with outcomes \(\{0,1\}\) and success probability \(p\) is \[ f(x) = f(x; p) = P(X = x) = p^x (1-p)^{1-x}. \]
This is a distribution with one parameter \(p \in [0,1] = \Theta \subset \mathbb{R}\). We write \(X \sim \text{Bernoulli}(p)\).
Calculate: \[ E(X) = \qquad\qquad \operatorname{Var}(X) = \]
The expectation is a sum over the two possible values: \[ E(X) = 0 \cdot (1-p) + 1 \cdot p = p. \]
For the variance, note that \(X\) takes only the values \(0\) and \(1\), so that \(X^2 = X\) and therefore \(E(X^2) = E(X) = p\). Hence \[ \operatorname{Var}(X) = E(X^2) - \big(E(X)\big)^2 = p - p^2 = p(1-p). \] This is largest at \(p = 1/2\) and zero at \(p = 0\) and \(p = 1\).
The same two-term sum gives the mgf, which we need in the next section: \[ M(t) = E\!\left(e^{tX}\right) = (1-p)e^{0} + pe^{t} = (1-p) + pe^{t}. \]
The binomial distribution: a sum of Bernoulli trials
Let \(X_1, \ldots, X_n \overset{iid}{\sim} \text{Bernoulli}(p)\) and define the number of successes in \(n\) independent Bernoulli trials as \[ S_n = \sum_{i=1}^{n} X_i. \]
There are \(\binom{n}{k}\) equally probable ways to observe \(k\) successes (and thus \(n-k\) failures).
By independence, each of them has probability \(p^k (1-p)^{n-k}\).
Thus: \[ P(S_n = k) = \binom{n}{k} p^k (1-p)^{n-k}. \]
… and we write \(S_n \sim \text{Bin}(n, p)\) (or just \(X \sim \text{Bin}(n,p)\)).
The mgf of the binomial distribution is \[ \begin{aligned} M(t) = E\!\left(e^{tX}\right) &= \sum_{x=0}^{n} e^{tx} \binom{n}{x} p^x (1-p)^{n-x} = \sum_{x=0}^{n} \binom{n}{x} \left(pe^{t}\right)^{x} (1-p)^{n-x} \\ &= \left((1-p) + pe^{t}\right)^{n}, \end{aligned} \] because of the binomial theorem: \((a+b)^n = \sum_{x=0}^{n} \binom{n}{x} b^x a^{n-x}\).
Calculate \(E(X)\) and \(\operatorname{Var}(X)\):
Note first that the mgf is the Bernoulli mgf raised to the \(n\)-th power, which is what we should expect for a sum of \(n\) independent Bernoulli variables.
Write \(q = 1-p\) and differentiate: \[ M'(t) = n\left(q + pe^{t}\right)^{n-1} pe^{t}, \] so that \[ E(X) = M'(0) = n \cdot 1^{n-1} \cdot p = np. \]
Differentiating once more, by the product rule, \[ M''(t) = n(n-1)\left(q + pe^{t}\right)^{n-2}\left(pe^{t}\right)^{2} + n\left(q + pe^{t}\right)^{n-1} pe^{t}, \] so that \(E(X^2) = M''(0) = n(n-1)p^2 + np\), and therefore \[ \operatorname{Var}(X) = n(n-1)p^{2} + np - (np)^{2} = np - np^{2} = np(1-p). \]
Both answers are \(n\) times the Bernoulli values, as they must be for a sum of \(n\) independent copies, and setting \(n = 1\) returns the Bernoulli case.
Poisson: counts of events in continuous space or time
Let \(X\) be the number of events recorded in a section of time or space.
Think of the section as containing a lot of very small units (\(n\) very big), such that the occurrence of an event within that unit can be considered a Bernoulli trial with \(p\) very small.
In formal terms: let \(g(x, w)\) denote the probability of \(x\) events in each interval of length \(w\). If:
- \(g(1,h) = \lambda h + o(h)\), \(\lambda\) is a positive constant, \(h > 0\),
- \(\displaystyle\sum_{x=2}^{\infty} g(x,h) = o(h)\),
- the numbers of changes/events in non-overlapping intervals are independent,
then it can be shown that \[ g(x, w) = \frac{(\lambda w)^x e^{-\lambda w}}{x!}, \qquad x = 0, 1, 2, \ldots \]
- The proof is essentially about solving a simple differential equation.
- “\(o(h)\)” is the “little-o”-notation, and you can read this as “a quantity that tends to zero faster than \(h\)”.
For an integer random variable \(X\) that is Poisson distributed, we usually just write the probabilities using a single parameter \(\lambda\): \[ P(X = x) = \frac{\lambda^x e^{-\lambda}}{x!}, \qquad x = 0, 1, 2, \ldots \]
An important property of the Poisson distribution is that \[ E(X) = \operatorname{Var}(X) = \lambda. \]
It is also closed under summation: if \(X_1, \ldots, X_n\) are independent with \(X_i \sim \text{Poisson}(\lambda_i)\), then \[ Y = \sum_{i=1}^{n} X_i \sim \text{Poisson}\!\left(\sum_{i=1}^{n} \lambda_i\right). \]
Both claims follow from the moment generating function, which comes straight from the definition: \[ \begin{aligned} M(t) = E\!\left(e^{tX}\right) &= \sum_{x=0}^{\infty} e^{tx}\,\frac{\lambda^x e^{-\lambda}}{x!} = e^{-\lambda}\sum_{x=0}^{\infty} \frac{\left(\lambda e^{t}\right)^{x}}{x!} \\ &= e^{-\lambda}\,e^{\lambda e^{t}} = \exp\!\left\{\lambda\left(e^{t} - 1\right)\right\}. \end{aligned} \] The only step that needs a comment is the second one, which is the exponential series \(\sum_{k=0}^{\infty} a^k/k! = e^{a}\) applied with \(a = \lambda e^{t}\).
The mean and the variance. Differentiating once, and again by the product rule, \[ M'(t) = \lambda e^{t} M(t), \qquad M''(t) = \left(\lambda e^{t} + \lambda^2 e^{2t}\right) M(t). \] Since \(M(0) = 1\), this gives \[ E(X) = M'(0) = \lambda, \qquad E(X^2) = M''(0) = \lambda + \lambda^2, \] and therefore \[ \operatorname{Var}(X) = E(X^2) - \big(E(X)\big)^2 = \lambda + \lambda^2 - \lambda^2 = \lambda. \]
Closure under summation. Independence turns the moment generating function of a sum into the product of the moment generating functions, so \[ M_Y(t) = \prod_{i=1}^{n} \exp\!\left\{\lambda_i\left(e^{t}-1\right)\right\} = \exp\!\left\{\Big(\sum_{i=1}^{n}\lambda_i\Big)\left(e^{t}-1\right)\right\}. \] This is the moment generating function of a Poisson variable with parameter \(\sum_{i=1}^{n}\lambda_i\), so by the uniqueness theorem from the previous lecture, \(Y \sim \text{Poisson}\!\left(\sum_{i=1}^{n}\lambda_i\right)\).
Note that the last step never computes a probability for \(Y\). It recognises the moment generating function and appeals to uniqueness.
Geometric distribution: waiting time in Bernoulli trials
Let \(X_1, X_2, \ldots\) be a series of independent Bernoulli trials with success probability \(p\).
Let \(T\) be the number of trials until the first success.
That is, we observe \((k-1)\) failures, then one success.
Hence, the p.m.f. of \(T\) is \[ P(T = k) = p(1-p)^{k-1}, \qquad k = 1, 2, 3, \ldots \]
Negative binomial: general waiting times in Bernoulli trials
Same setting as above, but let \(T\) be the number of trials until the \(r\)’th success.
We need \(r-1\) successes in the first \(k-1\) trials and then a success on the \(k\)’th trial.
The probability of this is \[ P(T = k) = \underbrace{\binom{k-1}{r-1} p^{r-1}(1-p)^{k-r}}_{\substack{r-1 \text{ successes in } k-1 \text{ trials; this is the}\\ \text{binomial with } k-1 \text{ instead of } n,\ \text{evaluated at } r-1}} \times \underbrace{p}_{\substack{\text{the final}\\ \text{success}}} = \binom{k-1}{r-1} p^{r}(1-p)^{k-r}, \] for \(k = r, r+1, r+2, \ldots\)
If \(r = 1\), the negative binomial reduces to the geometric (obviously!): substituting \(r = 1\) gives \(\binom{k-1}{0} = 1\) and leaves \(P(T = k) = p(1-p)^{k-1}\).
Counted in failures rather than in trials, and with the mean held fixed, the negative binomial approaches the Poisson as \(r \to \infty\). Unlike the case \(r = 1\), this is not a substitution: it needs both the change of variable and the reparameterisation.
Negative binomial \(\to\) Poisson. Let \(Y = T - r\) be the number of failures before the \(r\)’th success, and let \(p = r/(r+\lambda)\), so that \(E(Y) = \lambda\) for every \(r\). Then for each fixed \(y\), \[ P(Y = y) \longrightarrow \frac{\lambda^{y} e^{-\lambda}}{y!} \qquad \text{as } r \to \infty. \]
Proof. Since \(1 - p = \lambda/(r+\lambda)\), \[ P(Y = y) = \binom{y+r-1}{y} p^{r}(1-p)^{y} = \frac{\lambda^{y}}{y!} \prod_{j=0}^{y-1}\frac{r+j}{r+\lambda} \left(1 + \frac{\lambda}{r}\right)^{-r}. \] The product has \(y\) factors, each tending to \(1\), and \((1+\lambda/r)^{-r} \to e^{-\lambda}\). \(\qquad \blacksquare\)
- Thus, the negative binomial is more flexible than the Poisson (2 vs. 1 parameters) and the mean and the variance are not necessarily equal.
\(\rightarrow\) Very useful for practical modeling of counting variables!
Exponential distribution: time to first event
Let us return to the argument leading to the Poisson distribution: the time/space is continuous. Let \(W\) be the time until the first event in a Poisson process with rate \(\lambda\).
\(W\) is a continuous variable.
We can then show (p. 158–159 in HMC) that the cdf of \(W\) is \[ G(w) = P(W \leq w) = \int_0^w \lambda e^{-\lambda u}\,du, \]
and hence that the pdf is \[ g(w) = \begin{cases} \lambda e^{-\lambda w} & 0 < w < \infty \\ 0 & \text{elsewhere.} \end{cases} \]
This is the exponential distribution.
Note the parallel to the geometric distribution; time to first event in a Poisson process in discrete and continuous time.
These two distributions also share the memoryless property: the probability of waiting a certain time for the event (failure, death, …) only depends on the interval length, not how long we have been waiting so far.
\(\rightarrow\) Mathematically, for the exponential distribution, taking \(s > t\) so that \(\{W > s\} \subset \{W > t\}\) (which is what makes the second equality below hold): \[ \begin{aligned} P(W > s \mid W > t) &= \frac{P(W > s,\, W > t)}{P(W > t)} = \frac{P(W > s)}{P(W > t)} = \frac{\int_s^{\infty} \lambda e^{-\lambda w}\,dw}{\int_t^{\infty} \lambda e^{-\lambda w}\,dw} \\ &= \frac{\exp(-s\lambda)}{\exp(-t\lambda)} = e^{-\lambda(s-t)} = P(W > s-t). \end{aligned} \]
Calculate: \[ E(W) = \qquad\qquad \operatorname{Var}(W) = \]
Start from the mgf: \[ M(t) = E\!\left(e^{tW}\right) = \int_0^{\infty} e^{tw}\lambda e^{-\lambda w}\,dw = \lambda \int_0^{\infty} e^{-(\lambda - t)w}\,dw = \frac{\lambda}{\lambda - t}. \] This requires \(t < \lambda\): for \(t \geq \lambda\) the exponent is no longer negative and the integral diverges. It is a concrete illustration of why the definition of the mgf asks only for an interval around zero, and never for all \(t\).
Differentiating, \[ M'(t) = \frac{\lambda}{(\lambda - t)^{2}}, \qquad M''(t) = \frac{2\lambda}{(\lambda - t)^{3}}, \] so that \[ E(W) = M'(0) = \frac{1}{\lambda}, \qquad E(W^2) = M''(0) = \frac{2}{\lambda^{2}}, \] and therefore \[ \operatorname{Var}(W) = \frac{2}{\lambda^{2}} - \frac{1}{\lambda^{2}} = \frac{1}{\lambda^{2}}. \]
The mean and the standard deviation are both \(1/\lambda\), whatever the rate.
Gamma distribution: time to \(k\)’th event
A more general version of the preceding argument gives rise to the gamma distribution:
The waiting time \(W\) until the \(k\)’th event in a Poisson process has density function \[ g(w) = \begin{cases} \dfrac{\lambda^k w^{k-1} e^{-\lambda w}}{\Gamma(k)} & 0 < w < \infty \\[6pt] 0 & \text{elsewhere,} \end{cases} \]
where \(\Gamma(k) = \int_0^{\infty} y^{k-1} e^{-y}\,dy\) is the gamma function.
The most common parameterization is \(\alpha = k\) and \(\beta = 1/\lambda\).
Example 3.3.1 in CB establishes this link from the perspective of already knowing the gamma distribution, which is a little bit easier than deriving it from the Poisson: if \(X \sim \text{Gamma}(\alpha, \beta)\) where \(\alpha\) is an integer, then \[ P(X \leq x) = P(Y \geq \alpha), \] where \(Y \sim \text{Poisson}(x/\beta)\); that is, the events \(X \leq x\) and \(Y \geq \alpha\) are the same. Since \(\alpha\) is an integer, we write \(\Gamma(\alpha) = (\alpha-1)!\) to get \[ \begin{aligned} P(X \leq x) &= \frac{1}{(\alpha-1)!\,\beta^{\alpha}} \int_0^x t^{\alpha-1} e^{-t/\beta}\,dt \\ &= \frac{1}{(\alpha-1)!\,\beta^{\alpha}} \left[ -t^{\alpha-1}\beta e^{-t/\beta} \Big|_0^x + \int_0^x (\alpha-1) t^{\alpha-2} \beta e^{-t/\beta}\,dt \right], \end{aligned} \] where we use integration by parts with \(u = t^{\alpha-1}\), \(dv = e^{-t/\beta}\,dt\). Continuing: \[ \begin{aligned} P(X \leq x) &= \frac{-1}{(\alpha-1)!\,\beta^{\alpha-1}} x^{\alpha-1} e^{-x/\beta} + \frac{1}{(\alpha-2)!\,\beta^{\alpha-1}} \int_0^x t^{\alpha-2} e^{-t/\beta}\,dt \\ &= \frac{1}{(\alpha-2)!\,\beta^{\alpha-1}} \int_0^x t^{\alpha-2} e^{-t/\beta}\,dt - P(Y = \alpha-1), \end{aligned} \] where \(Y \sim \text{Poisson}(x/\beta)\). Doing this \(\alpha-2\) more times we get the desired result.
- If \(X \sim \text{Gamma}(\alpha, \beta)\), then \[ E(X) = \alpha\beta, \qquad \operatorname{Var}(X) = \alpha\beta^2. \]
The normal / Gaussian distribution: the default model for additive noise
We will later prove a remarkable result: sums of random variables, properly normalized, tend to a common distribution (almost) regardless of the distribution of the components!
This is the Central Limit Theorem, and this limiting distribution is known as the normal / Gaussian distribution.
The density of the normal distribution with \(E(X) = \mu\) and \(\operatorname{Var}(X) = \sigma^2\) is \[ f(x) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left\{-\frac{(x-\mu)^2}{2\sigma^2}\right\}. \]
The fundamental nature of the normal distribution:
The normal distribution as a mathematical feature of nature:
- It is often natural to think of “additive noise” as the sum of many small influences.
- The CLT tells us that this noise will be approximately normally distributed.
The normal distribution as a tool in statistical inference:
- Taking sums over (functions of) observations is very useful for summarising data as an early step in statistical inference.
- The CLT tells us that these sums are approximately normally distributed.
From the normal to other useful inference distributions
\(\chi^2\) (chi-squared)
If \(Z_1, \ldots, Z_\nu \overset{iid}{\sim} N(0,1)\) then \(Q = \sum_{i=1}^{\nu} Z_i^2 \sim \chi^2_\nu\).
That is: a sum of \(n\) squared independent standard normal random variables is chi-squared distributed with \(n\) degrees of freedom.
Simply stated, this means that the empirical variance of samples will be approximately \(\chi^2\) distributed (we will come back to the exact details).
Note also that as \(n\) becomes big, the \(\chi^2\) distribution will approach the normal, due to the CLT.
Density function: \[ f(x) = \frac{x^{\frac{k}{2}-1} e^{-\frac{x}{2}}}{2^{\frac{k}{2}} \Gamma\!\left(\frac{k}{2}\right)}. \]
(What? Is it just a special case of the gamma?)
Student \(t\)
If \(Z \sim N(0,1)\), \(U \sim \chi^2_\nu\) and \(Z \perp U\) (independent), then \(T = Z / \sqrt{U/\nu}\) is \(t\)-distributed with \(\nu\) degrees of freedom.
\(\rightarrow\) In other words: the \(t\)-statistic \((\bar{X}-\mu)/s\) is approximately \(t\)-distributed (exact details come later).
\(F\)
If \(U_1 \sim \chi^2_{\nu_1}\), \(U_2 \sim \chi^2_{\nu_2}\) are independent, then \[ F = \frac{U_1/\nu_1}{U_2/\nu_2} \] is \(F\)-distributed with \((\nu_1, \nu_2)\) degrees of freedom.
\(\rightarrow\) That is, the ratio of two empirical variances is approximately \(F\)-distributed.
Beta: proportions of accumulated mass
(and one of few parametric families defined on \([0,1]\))
The “physical” argument
Let \(X_1\) and \(X_2\) be two “masses” that accumulate over time.
Each “addition” to the mass can be thought of as a smaller “accumulation” that is added to the total after being triggered by a Poisson process; hence they can be thought of as exponentially distributed waiting times.
Let \(X_1 \sim \text{Gamma}(\alpha, \theta)\) and \(X_2 \sim \text{Gamma}(\beta, \theta)\) be independent. Note that they share the scale \(\theta\) but have different shapes. The proportion of mass in \(X_1\) is then \[ \frac{X_1}{X_1 + X_2} \sim \text{Beta}(\alpha, \beta). \] The scale \(\theta\) cancels, which is why it does not appear on the right.
Deriving the density of the Beta distribution
See p. 164 in HMC. We need the bivariate version of the transformation formula (next lecture), where the derivative of the inverse transformation is replaced with the determinant of the Jacobian matrix.
Let \(X_1 \sim \text{Gamma}(\alpha, 1)\) and \(X_2 \sim \text{Gamma}(\beta, 1)\) be independent; the scale cancels in the ratio, so taking it to be \(1\) costs nothing. Define the bivariate transformation \[ Y_1 = X_1 + X_2, \qquad Y_2 = \frac{X_1}{X_1 + X_2}, \] for which the inverse transformation is \(x_1 = y_1 y_2\), \(x_2 = y_1(1 - y_2)\). The Jacobian determinant is \[ J = \begin{vmatrix} y_2 & y_1 \\ 1-y_2 & -y_1 \end{vmatrix} = -y_1. \]
This transformation is one-to-one and maps \((0,\infty) \times (0,\infty)\) onto \((0,\infty) \times (0,1)\). The joint density of \((Y_1, Y_2)\) is \[ \begin{aligned} g(y_1, y_2) &= y_1 \cdot \frac{1}{\Gamma(\alpha)\Gamma(\beta)} (y_1 y_2)^{\alpha-1} \big(y_1(1-y_2)\big)^{\beta-1} e^{-y_1} \\ &= \begin{cases} \dfrac{y_2^{\alpha-1}(1-y_2)^{\beta-1}}{\Gamma(\alpha)\Gamma(\beta)}\, y_1^{\alpha+\beta-1} e^{-y_1} & 0 < y_1 < \infty,\; 0 < y_2 < 1 \\[6pt] 0 & \text{elsewhere.} \end{cases} \end{aligned} \]
Since the joint density of \((y_1, y_2)\) can be written as \(h_1(y_1) h_2(y_2)\), the two variables are independent (Lemma 4.2.7 in CB, next lecture). The marginal pdf of \(Y_2\) is \[ \begin{aligned} g_2(y_2) &= \frac{y_2^{\alpha-1}(1-y_2)^{\beta-1}}{\Gamma(\alpha)\Gamma(\beta)} \int_0^{\infty} y_1^{\alpha+\beta-1} e^{-y_1}\,dy_1 \\ &= \begin{cases} \dfrac{\Gamma(\alpha+\beta)}{\Gamma(\alpha)\Gamma(\beta)}\, y_2^{\alpha-1}(1-y_2)^{\beta-1} & 0 < y_2 < 1 \\[6pt] 0 & \text{elsewhere,} \end{cases} \end{aligned} \] where \[ B(\alpha, \beta) = \int_0^1 x^{\alpha-1}(1-x)^{\beta-1}\,dx = \frac{\Gamma(\alpha)\Gamma(\beta)}{\Gamma(\alpha+\beta)} \] is the Beta function, which gives name to this distribution.
Note that we can view the Beta distribution in a similar manner as the Negative Binomial (“general waiting times in binomial trials, but generally useful as a flexible distribution for counting variables”):
Yes, we can derive the Beta from other distributions (“Bernoulli \(\to\) Binomial \(\to\) Poisson \(\to\) Exponential \(\to\) Gamma \(\to\) Beta”), but it is generally very useful to model random variables that live on \([0,1]\).
Check out Figures 3.3.3 and 3.3.4 in CB (p. 88) to appreciate the flexibility of this distribution.
Other common distributions
If \(Z\) is normal, then \(\exp(Z)\) is lognormal, useful for modeling growth and compounding.
If \(U\) has constant density \(f(u) = (b-a)^{-1}\) over the interval \([a,b]\), then \(U\) is uniform. This is extremely useful because all other random variables can be expressed in terms of a uniform:
- If \(U \sim \text{Uniform}(0,1)\), then \(Y = F^{-1}(U)\) has distribution \(F\)! And, conversely, \(F(X) \sim \text{Uniform}(0,1)\)!
- (See Theorem 2.1.10 in CB.)
The exponential family is a large class of distributions, including most (all?) of those we have discussed so far, for which many general results and methods can be expressed and proved. A pmf or pdf that can be expressed in the following form is a member of the exponential family: \[ f(x; \theta) = h(x)\,c(\theta) \exp\left\{\sum_{i=1}^{k} w_i(\theta)\,t_i(x)\right\}. \]
A task to complete your understanding
We have now looked at a few common parametric families of distributions. Create an overview with the following properties for each of them:
- pdf/pmf, cdf, mgf, \(E(X)\), \(\operatorname{Var}(X)\), and any additional properties that you deem important (e.g. a sum of independent Poisson variables is also Poisson).
- A complete overview will also include derivations of each property.
Complement this with relations between the distributions:
- Which are obvious (e.g. Bernoulli \(\to\) Binomial).
- Provide mathematical details where possible (Binomial \(\to\) Poisson \(\to\) Gamma, for example?).
- Indicate which relations are probably too difficult to prove right now (possibly \(\sum X_i \to\) normal \(\to \chi^2, t, F\)).