2. Distributions of Random Variables

Recap I

  • Random variable: \(X(c): \mathcal{B} \mapsto \mathbb{R}\).
  • The CDF: \(F_X(x) = P(X\leq x)\).
  • The CDF always exists and completely characterizes \(X\).
    • Calculating probabilities:
    \[P(X\in A) = \int_A \textrm{d}F_X(x) = \begin{cases} \int_A f_X(x) \,\textrm{d}x & \textrm{if cont.} \\ \sum_A p_X(x) & \textrm{if disc.}\end{cases}\]
    • If \(X \sim F_X\), then \(Y = g(X) \sim F_Y\), where we can set up explicit formulas for \(F_Y\) if \(g(\cdot)\) is one-to-one (i.e., if \(g^{-1}(\cdot)\) exists).

Recap II

  • Expectations:
    • The expectation of \(X\): \(\mu = \textrm{E}(X) = \int x\,\textrm{d}F_X(x) = \begin{cases} \int x f_X(x) \,\textrm{d}x & \textrm{if cont.} \\ \sum xp_X(x) & \textrm{if disc.}\end{cases}\)
    • The \(p\)’th moment: E\((X^p)\).
    • The variance: Var\((X) = \textrm{E}(X-\mu)^2 = \textrm{E}(X^2) - \mu^2\).
    • Functions in general: E\((g(X)) = \begin{cases} \int g(x)f_X(x) \,\textrm{d}x & \textrm{if cont.} \\ \sum g(x)p_X(x) & \textrm{if disc.}\end{cases}\)
  • A special expectation, the moment generating function: \[M_X(t) = \textrm{E}(e^{tX}), \qquad -h < t < h, \,\, h> 0.\]
  • The name comes from the following property:

\[\begin{align} M'(t) &= \frac{d}{dt} \int e^{tx}f(x)\,\textrm{d}x = \int \frac{d}{dt} e^{tx}f(x)\,\textrm{d}x = \int xe^{tx}f(x)\,\textrm{d}x \\ & \Rightarrow M´(0) = \int xf(x)\,\textrm{d}x = \textrm{E}(x). \qquad \textrm{By induction } \textrm{E}(X^p) = M^{(p)}(0). \end{align}\]

Why parametric families

  • A standard empirical task is to “choose” a distribution that “fits” some data. This is not entirely satisfactory from a theoretical perspective.

  • A parametric model is a deliberate restriction \[ \{F_\theta : \theta \in \Theta \subset \mathbb{R}^k\}, \] a collection of distributions indexed by a (possibly vector-valued) parameter \(\theta\).

  • The trade-off:

    • Fewer degrees of freedom \(\Rightarrow\) easier to estimate and to interpret.
    • But mis-specification is possible (and non-trivial).
  • We will build a small set of families from first principles: start with simple mechanisms (trials, counting, waiting time), apply sums, limits and transformations, and arrive at the standard parametric families.

A roadmap

Discrete distributions

  • Bernoulli \(\to\) Binomial (sum of trials)
  • Binomial \(\to\) Poisson (rare-event limit)
  • Bernoulli trials \(\to\) Geometric / Negative binomial (waiting time)

Continuous time

  • Poisson process \(\to\) Exponential (first waiting time)
  • Exponential sums \(\to\) Gamma (waiting time to multiple events)

Continuous measurements

  • Normal as the standard “additive noise” model (CLT proof comes later)
  • \(\chi^2\), \(t\), \(F\) derived from the normal (used for inference later)

Constrained supports

  • Beta for proportions on \([0,1]\)
  • Lognormal for positive multiplicative variables

Bernoulli: the atom of models

Consider a random experiment where

  1. the sample space contains exactly two elements, and
  2. the two outcomes happen with constant probabilities \(p\) and \(1-p\).

This is a Bernoulli trial. The pmf of \(X\) with outcomes \(\{0,1\}\) and success probability \(p\) is \[ f(x; p) = P(X = x) = p^x(1-p)^{1-x}. \]

One parameter, \(p \in [0,1] = \Theta \subset \mathbb{R}\). We write \(X \sim \text{Bernoulli}(p)\).

Discussion: calculate \(E(X) = \qquad\) and \(\text{Var}(X) = \qquad\)

The binomial distribution: a sum of Bernoulli trials

Let \(X_1, \ldots, X_n \overset{\text{iid}}{\sim} \text{Bernoulli}(p)\) and let \(S_n = \sum_{i=1}^n X_i\) be the number of successes.

  • There are \(\binom{n}{k}\) equally probable ways to observe \(k\) successes.
  • By independence, each has probability \(p^k(1-p)^{n-k}\). \[ P(S_n = k) = \binom{n}{k} p^k (1-p)^{n-k}, \quad k = 0, 1, \ldots, n, \] so \(S_n \sim \text{Bin}(n, p)\).

By the binomial theorem, the moment generating function is \[ M(t) = \bigl((1-p) + pe^t\bigr)^n = \bigl[M_{\text{Bernoulli}}(t)\bigr]^n. \]

Moment generating functions of sums are products of moment generating functions.

Derivation of the mgf, and \(E(X)\), Var\((X)\) from it, see notes, The binomial distribution.

Poisson: counts of events in continuous space or time

Let \(X\) be the number of events in a section of time or space. Think of the section as many very small units (\(n\) big), each a Bernoulli trial with \(p\) small.

Formally, let \(g(x, h)\) be the probability of \(x\) events in an interval of length \(h\). If

  1. \(g(1, h) = \lambda h + o(h)\), where \(\lambda > 0\),
  2. \(\sum_{x=2}^\infty g(x, h) = o(h)\),
  3. counts in non-overlapping intervals are independent,

then \[ P(X = x) = \frac{\lambda^x e^{-\lambda}}{x!}, \quad x = 0, 1, 2, \ldots, \qquad X \sim \text{Poisson}(\lambda). \]

  • \(E(X) = \text{Var}(X) = \lambda\) — mean equals variance, a testable restriction.
  • Closed under summation: independent \(X_i \sim \text{Poisson}(\lambda_i)\) give \(\sum_i X_i \sim \text{Poisson}\bigl(\sum_i \lambda_i\bigr)\).

Proof of both, via moment generating functions, see notes, Poisson.

Waiting times: geometric and negative binomial

Independent Bernoulli\((p)\) trials. Let \(T\) be the number of trials until the \(r\)-th success: we need \(r-1\) successes in the first \(k-1\) trials, then a success on the \(k\)-th, so \[ P(T = k) = \binom{k-1}{r-1} p^r (1-p)^{k-r}, \quad k = r, r+1, \ldots \] This is the negative binomial.

  • \(r = 1\) gives the geometric, \(P(T = k) = p(1-p)^{k-1}\), \(k = 1, 2, \ldots\) — the waiting time to the first success.
  • As \(r \to \infty\) the negative binomial approaches the Poisson, once the mean is held fixed.
  • Two parameters rather than one, so the mean and variance need not be equal \(\Rightarrow\) the standard repair when counts are more variable than a Poisson allows.

Both derivations, and the precise \(r \to \infty\) statement, see notes, Geometric and Negative binomial.

Exponential distribution, and memorylessness

Let \(W\) be the time until the first event in a Poisson process with rate \(\lambda\). Taking \(x = 0\) in the Poisson pmf gives \(P(W > w) = e^{-\lambda w}\), so \[ g(w) = \lambda e^{-\lambda w}, \quad 0 < w < \infty, \qquad W \sim \text{Exp}(\lambda). \]

The continuous twin of the geometric, and like it memoryless. For \(s > t\), so that \(\{W > s\} \subset \{W > t\}\), \[ P(W > s \mid W > t) = \frac{P(W > s)}{P(W > t)} = \frac{e^{-\lambda s}}{e^{-\lambda t}} = e^{-\lambda(s-t)} = P(W > s - t). \]

How long you have already waited tells you nothing about how much longer you will wait.

\(E(W) = 1/\lambda\) and Var\((W) = 1/\lambda^2\), see notes, Exponential distribution.

Gamma distribution: time to the \(k\)-th event

The same argument for the \(k\)-th event in a Poisson process with rate \(\lambda\) gives the density \[ g(w) = \frac{\lambda^k w^{k-1} e^{-\lambda w}}{\Gamma(k)}, \quad 0 < w < \infty, \] where \(\Gamma(k) = \int_0^\infty y^{k-1} e^{-y}\,dy\) is the gamma function.

The common parameterization uses \(\alpha = k\) and \(\beta = 1/\lambda\): \(X \sim \text{Gamma}(\alpha, \beta)\), with \(E(X) = \alpha\beta\) and \(\text{Var}(X) = \alpha\beta^2\).

Density and moments: notes, Gamma distribution.

  • Summary of the waiting-times distributions:
one event several events
discrete trials geometric negative binomial
continuous time exponential gamma

Normal distribution: the default model for additive noise

We will later prove that sums of random variables tend to a common distribution (almost) regardless of the distribution of the components — the Central Limit Theorem. That limit is the normal distribution, with density \[ f(x) = \frac{1}{\sqrt{2\pi\sigma^2}} \exp\!\left\{-\frac{(x-\mu)^2}{2\sigma^2}\right\}. \]

  1. A feature of nature: “additive noise” is often the sum of many small influences, so the CLT makes it approximately normal.

  2. A tool in inference: the statistics we compute are sums over observations, so the CLT makes those approximately normal — whatever the data look like.

From the normal to other useful inference distributions

\(\chi^2\): if \(Z_1, \ldots, Z_\nu \overset{\text{iid}}{\sim} N(0,1)\), then \(Q = \sum_{i=1}^\nu Z_i^2 \sim \chi^2_\nu\), with \(\nu\) degrees of freedom. The empirical variance of a sample is approximately \(\chi^2\).

(Is the \(\chi^2\) just a special case of the Gamma?)

Student \(t\): if \(Z \sim N(0,1)\), \(U \sim \chi^2_\nu\), independent, then \(T = Z/\sqrt{U/\nu}\) is \(t\)-distributed with \(\nu\) degrees of freedom. The \(t\)-statistic \((\bar{X} - \mu)/s\) is approximately \(t\).

\(F\): if \(U_1 \sim \chi^2_{\nu_1}\) and \(U_2 \sim \chi^2_{\nu_2}\) are independent, then \(F = \dfrac{U_1/\nu_1}{U_2/\nu_2}\) is \(F\)-distributed with \((\nu_1, \nu_2)\) degrees of freedom. The ratio of two empirical variances is approximately \(F\).

Other families

Beta: defined on \([0,1]\), so the natural model for proportions. If \(X_1 \sim \text{Gamma}(\alpha, \theta)\) and \(X_2 \sim \text{Gamma}(\beta, \theta)\) are independent — same scale, different shapes — then \(X_1/(X_1+X_2) \sim \text{Beta}(\alpha, \beta)\).

Lognormal: \(\exp(Z)\) for \(Z \sim N(\mu, \sigma^2)\). Additive noise becomes multiplicative: growth and compounding.

Uniform: \(U \sim \text{Uniform}(0,1)\) generates all the others — \(F^{-1}(U)\) has distribution \(F\), and \(F(X)\) is uniform.

The exponential family: a large class containing most of the above, for which general results can be proved once instead of case by case: \[ f(x; \theta) = h(x)\, c(\theta) \exp\!\left\{\sum_{j=1}^k w_j(\theta)\, t_j(x)\right\}. \]

The Beta construction and its density, see notes, Beta.

A task to complete your understanding

Create an overview with the following properties for each family:

  • pdf/pmf, cdf, mgf, \(E(X)\), \(\text{Var}(X)\), and any additional notable properties (e.g. a sum of independent Poisson variables is also Poisson).
  • A complete overview will also include derivations of each property.

Complement this with relations between the distributions:

  • Which are obvious (e.g. Bernoulli \(\to\) Binomial)?
  • Provide mathematical details where possible (Binomial \(\to\) Poisson \(\to\) Gamma, for example).
  • Indicate which relations are probably too difficult to prove right now (possibly \(\sum X_i \to\) Normal \(\to \chi^2, t, F\)).