Definition (Rudin 3.1). A sequence \(\{p_n\}\) of numbers converges if there exists a number \(p\) such that for every \(\varepsilon > 0\) there is an integer \(N\) with \[ n \geq N \;\Longrightarrow\; |p_n - p| < \varepsilon. \] We write \(p_n \to p\), or \(\displaystyle\lim_{n\to\infty} p_n = p\).
We can make the distance between \(p_n\) and \(p\) smaller than any positive number, for all \(n \geq N\), by choosing \(N\) large enough.
Why we need the same idea for random variables:
For sequences \(X_1, X_2, \ldots\) of random variables we will distinguish:
Important. The sequence \(X_1, X_2, \ldots\) is not an iid sample. The \(X_n\) have different distributions in general and are typically dependent. The whole point is to describe how the distribution of \(X_n\) changes as \(n\) grows.
From CB: “The distribution changes as the subscript changes and the convergence concepts describe different ways in which the distribution of \(X_n\) converges to some limiting distribution (of \(X\)) as the subscript becomes large.”
Definition 5.5.1. A sequence \(X_1, X_2, \ldots\) of random variables converges in probability to \(X\) if, for every \(\varepsilon > 0\), \[ \lim_{n\to\infty} P\big(|X_n - X| \geq \varepsilon\big) = 0, \] equivalently \(\lim_{n\to\infty} P\big(|X_n - X| < \varepsilon\big) = 1\).
Notation: \(X_n \xrightarrow{P} X\).
We are ready to state and prove our first major result, which formalizes the convergence of the sample mean to the population mean:
Theorem 5.5.2 (Weak Law of Large Numbers). Let \(X_1, X_2, \ldots\) be iid with \(E[X_i] = \mu\) and \(\operatorname{Var}(X_i) = \sigma^2 < \infty\), and set \[ \bar X_n = \frac{1}{n}\sum_{i=1}^n X_i. \] Then for every \(\varepsilon > 0\), \[ \lim_{n\to\infty} P\big(|\bar X_n - \mu| < \varepsilon\big) = 1, \] i.e. \(\bar X_n \xrightarrow{P} \mu\).
We know from Topic 4 (Theorem 5.2.6) that \(E(\bar X_n) = \mu\) and \(\operatorname{Var}(\bar X_n) = \sigma^2/n\). The centre stays at \(\mu\) while the spread shrinks. The curve is the \(N(\mu, \sigma^2/n)\) density, which is the distribution of \(\bar X_n\) for a normal population.
Even though \(\bar X_n\) is random, it becomes more and more like the constant \(\mu\): for every \(\varepsilon > 0\), the probability that \(\bar X_n\) lands more than \(\varepsilon\) away from \(\mu\) converges to zero as \(n \to \infty\).
Chebyshev’s inequality (Theorem 3.6.1). Let \(X\) be a random variable and \(g(x) \geq 0\). Then for any \(r > 0\), \[P\big(g(X) \geq r\big) \leq \frac{E\,g(X)}{r}.\]
Apply it with \(g(x) = (x - \mu)^2\) and \(r = \varepsilon^2\): \[ P\big(|\bar X_n - \mu| \geq \varepsilon\big) = P\big((\bar X_n - \mu)^2 \geq \varepsilon^2\big) \leq \frac{E(\bar X_n - \mu)^2}{\varepsilon^2}. \]
The numerator is \(\operatorname{Var}(\bar X_n) = \sigma^2/n\), by Theorem 5.2.6. Hence \[ P\big(|\bar X_n - \mu| < \varepsilon\big) \;\geq\; 1 - \frac{\sigma^2}{n\varepsilon^2} \;\xrightarrow{n\to\infty}\; 1. \qquad\blacksquare \]
The finite variance is not needed for the result, but it makes the proof this short.
Theorem 5.5.4 (Continuous mapping). If \(X_n \xrightarrow{P} X\) and \(h\) is a continuous function, then \[ h(X_n) \xrightarrow{P} h(X). \]
Convergence in probability is preserved under continuous transformations, e.g. \(\bar X_n \xrightarrow{P} \mu \Rightarrow \bar X_n^2 \xrightarrow{P} \mu^2\).
The same Chebyshev argument shows that \(S_n^2 \xrightarrow{P} \sigma^2\) (Example 5.5.3, problem P3): the sample variance is a consistent estimator of \(\sigma^2\). By continuous mapping, \(S_n \xrightarrow{P} \sigma\).
Definition 5.5.6. A sequence \(X_1, X_2, \ldots\) of random variables converges almost surely to \(X\) if, for every \(\varepsilon > 0\), \[ P\!\left(\lim_{n\to\infty} |X_n - X| < \varepsilon\right) = 1. \] Notation: \(X_n \xrightarrow{a.s.} X\).
Theorem 5.5.9 (Strong Law of Large Numbers). Under the conditions of the WLLN, \(\bar X_n \xrightarrow{a.s.} \mu\).
Why almost sure convergence implies convergence in probability: notes, Concept 2: Almost sure convergence.
Definition 5.5.10. A sequence \(X_1, X_2, \ldots\) converges in distribution to \(X\) if \[ \lim_{n\to\infty} F_{X_n}(x) = F_X(x) \] at every point \(x\) where \(F_X\) is continuous. Notation: \(X_n \xrightarrow{d} X\).
This is the weakest of the three, but the most important in practice, because we will soon prove that the sample mean converges in distribution to the normal distribution (almost) regardless of the distribution of the individual observations.
Let \(X_1, X_2, \ldots\) be iid Uniform\((0,1)\) and \(X_{(n)} = \max_{1 \leq i \leq n} X_i\).
Intuition: As we observe more uniforms, the maximum should approach \(1\).
Since \(X_{(n)} < 1\), the event \(|X_{(n)} - 1| \geq \varepsilon\) is the event \(X_{(n)} \leq 1 - \varepsilon\). For any \(\varepsilon \in (0,1)\), \[ \begin{align*} P\big(X_{(n)} \leq 1 - \varepsilon\big) &= P(X_1 \leq 1 - \varepsilon, \ldots, X_n \leq 1 - \varepsilon) && \text{(the max is below iff all are)}\\ &= \prod_{i=1}^n P(X_i \leq 1 - \varepsilon) && \text{(independence)}\\ &= (1 - \varepsilon)^n \to 0 && \text{(uniform cdf: } F(x) = x\text{)}. \end{align*} \]
So \(X_{(n)} \xrightarrow{P} 1\). (For \(\varepsilon \geq 1\) the probability is \(0\).)
Each of 2000 repetitions draws \(n\) uniforms and records the maximum.
Now fix \(t > 0\) and substitute \(\varepsilon = t/n\) (a specific sequence converging to \(0\) at a specific speed): \[ P\!\left(X_{(n)} \leq 1 - \tfrac{t}{n}\right) = \left(1 - \tfrac{t}{n}\right)^n \xrightarrow{n\to\infty} e^{-t}. \]
Equivalently, \[ P\big(n(1 - X_{(n)}) \leq t\big) \to 1 - e^{-t}, \] which is the CDF of an Exponential\((1)\) random variable: \(n(1 - X_{(n)}) \xrightarrow{d} \operatorname{Exp}(1)\).
Take some time to understand this result: \((1 - X_{(n)})\) goes to zero, and \(n\) goes to infinity. When we multiply them together we get something that does not go to zero, and also not to infinity. These two quantities go in opposite directions “at the same speed” in some sense, and we can say that the rate of convergence of \((1 - X_{(n)})\) is \(n\). See CB Exercise 5.42 for further exploration.
Each of 4000 repetitions draws \(n\) uniforms. Left: the average distance from the maximum to \(1\) against \(n\), on log scales, with a point for each \(n\) visited. Right: the distance multiplied by \(n\).
Theorem 5.5.14 (Central Limit Theorem). Let \(X_1, X_2, \ldots\) be iid with \(E[X_i] = \mu\), \(\operatorname{Var}(X_i) = \sigma^2 > 0\) (both finite), and suppose \(M_{X_i}(t)\) exists for \(|t| < h\), for some \(h > 0\). Let \[ G_n(x) = P\!\left(\frac{\sqrt{n}\,(\bar X_n - \mu)}{\sigma} \leq x\right). \] Then for every \(x \in \mathbb{R}\), \[ \lim_{n\to\infty} G_n(x) = \int_{-\infty}^x \frac{1}{\sqrt{2\pi}}\operatorname{exp}\left(-y^2/2\right)\,\textrm{d}y, \] that is, \(\dfrac{\sqrt{n}\,(\bar X_n - \mu)}{\sigma} \xrightarrow{d} N(0,1)\).
Each panel shows 3000 simulated values of \(\sqrt{n}\,(\bar X_n - \mu)/\sigma\) for samples of size \(n\) from the population named above it. The dashed curve is the \(N(0,1)\) density, and all panels share the same scales.
Let \(Y_i = (X_i - \mu)/\sigma\). Then \(E[Y_i] = 0\), \(\operatorname{Var}(Y_i) = 1\), and the common mgf \(M_Y(t)\) exists for \(|t| < \sigma h\). Rewrite the standardized mean in terms of the \(Y_i\): \[ \frac{\sqrt{n}\,(\bar X_n - \mu)}{\sigma} = \sqrt{n}\cdot\frac{1}{n}\sum_{i=1}^n \frac{X_i - \mu}{\sigma} = \frac{1}{\sqrt{n}}\sum_{i=1}^n Y_i. \]
Its mgf follows from two earlier results: \[ M_{\sum_{i} Y_i / \sqrt{n}}(t) \underset{\text{Thm 2.3.15}}{=} M_{\sum_{i} Y_i}\!\left(\frac{t}{\sqrt{n}}\right) \underset{\text{Thm 4.6.7}}{=} \Big(M_Y\big(t/\sqrt{n}\big)\Big)^{n}, \] first scaling (\(M_{aZ}(t) = M_Z(at)\)), then the mgf of a sum of independent variables.
Strategy. Show that \(\big(M_Y(t/\sqrt{n})\big)^n \to e^{t^2/2}\), the mgf of \(N(0,1)\). By Theorem 2.3.12, convergence of mgfs in a neighbourhood of \(0\) implies convergence of the cdfs at every continuity point, which is what the CLT claims.
Expand \(M_Y(t/\sqrt{n})\) as a Taylor series around \(0\), valid for \(|t| < \sqrt{n}\,\sigma h\): \[ M_Y(t/\sqrt{n}) = \sum_{k=0}^\infty M_Y^{(k)}(0)\,\frac{(t/\sqrt{n})^k}{k!}. \] If the mgf is finite in an interval around zero, it is differentiable at zero to all orders.
Using \(M_Y(0) = 1\), \(M_Y'(0) = E[Y] = 0\) and \(M_Y''(0) = E[Y^2] = 1\): \[ M_Y(t/\sqrt{n}) = 1 + \frac{(t/\sqrt{n})^2}{2} + R_Y(t/\sqrt{n}), \qquad R_Y(u) = \sum_{k=3}^\infty M_Y^{(k)}(0)\,\frac{u^k}{k!}. \]
The remainder vanishes faster than the last term kept (Theorem 5.5.21): \(R_Y(u)/u^2 \to 0\) as \(u \to 0\). With \(u = t/\sqrt{n}\) and \(t \neq 0\) fixed, \[ n\,R_Y(t/\sqrt{n}) = t^2\cdot\frac{R_Y(t/\sqrt{n})}{(t/\sqrt{n})^2} \;\xrightarrow{n\to\infty}\; t^2 \cdot 0 = 0, \] and at \(t = 0\) it is \(0\) for every \(n\).
Lemma 2.3.14. If \(a_n \to a\), then \(\displaystyle\lim_{n\to\infty}\left(1 + \frac{a_n}{n}\right)^n = e^a\).
Collect the terms into this form: \[ \big[M_Y(t/\sqrt{n})\big]^n = \left[1 + \frac{1}{n}\underbrace{\left(\frac{t^2}{2} + n\,R_Y(t/\sqrt{n})\right)}_{a_n \;\to\; t^2/2}\right]^n. \]
By the lemma, \(\big[M_Y(t/\sqrt{n})\big]^n \to e^{t^2/2}\), the mgf of \(N(0,1)\). By Theorem 2.3.12, \(G_n(x) \to \Phi(x)\) for every \(x\). \(\qquad\blacksquare\)
Rate of convergence. \(\bar X_n - \mu\) shrinks and \(\sqrt{n}\) grows, and their product behaves like a normal variable: the sample mean converges at rate \(\sqrt{n}\). This is the law of diminishing returns for data: going from 10 to 11 observations is worth more than going from 1000 to 1001.
Stephen Stigler counts this rate among The Seven Pillars of Statistical Wisdom. When Laplace discovered it (1810), “the implications were striking: if you wished to double the accuracy of an investigation, it was insufficient to double the effort; you must increase the effort fourfold.”
The CLT holds under much more general conditions: we can drop both “i”s in iid. Theorem 5.5.15 needs only finite variance, no mgf. The general versions are much harder to prove.
The following results are very useful in “practical” theoretical statistics.
Theorem 5.5.17 (Slutsky). Suppose \(X_n \xrightarrow{d} X\) and \(Y_n \xrightarrow{P} a\) for some constant \(a\). Then
Extremely useful: you can replace a nuisance quantity that converges to a constant (e.g. \(S_n \xrightarrow{P} \sigma\)) without changing the limit distribution.
From the previous topic, we have for a normal sample: \(\dfrac{\bar X_n - \mu}{S_n/\sqrt{n}} \sim t_{n-1}\) exactly, which gave the interval \(\bar X_n \pm t_{n-1,\alpha/2}\,S_n/\sqrt{n}\).
For any population with finite variance \(\sigma^2 > 0\): \[ \frac{\sqrt{n}\,(\bar X_n - \mu)}{S_n} \xrightarrow{d} N(0,1). \] The proof combines the CLT, the consistency of \(S_n\), and Slutsky: problem P10.
Theorem 5.5.24 (Delta method). Let \(Y_n\) satisfy \(\sqrt{n}\,(Y_n - \theta) \xrightarrow{d} N(0, \sigma^2)\). For a function \(g\), if \(g'(\theta)\) exists and \(g'(\theta) \neq 0\), then \[ \sqrt{n}\,\big(g(Y_n) - g(\theta)\big) \xrightarrow{d} N\!\big(0,\; \sigma^2 [g'(\theta)]^2\big). \]
Why \(g'(\theta) \neq 0\). If \(g'(\theta) = 0\), the limit is \(N(0, 0)\), a point mass at \(0\). That is true, but it says nothing about the distribution of \(g(Y_n)\). A second-order expansion is needed instead: problem A2.
Taylor-expand \(g\) around \(\theta\): \[ g(Y_n) = g(\theta) + g'(\theta)(Y_n - \theta) + R_n, \qquad \frac{R_n}{Y_n - \theta} \to 0 \text{ as } Y_n \to \theta. \]
1. \(Y_n \xrightarrow{P} \theta\). Write \(Y_n - \theta = \frac{1}{\sqrt{n}}\cdot\sqrt{n}(Y_n - \theta)\). The first factor tends to \(0\) and the second converges in distribution, so by Slutsky \(Y_n - \theta \xrightarrow{d} 0\). Convergence in distribution to a constant is convergence in probability (Theorem 5.5.13; problem P5).
2. The remainder is negligible. Write \(\sqrt{n}\,R_n = \dfrac{R_n}{Y_n - \theta}\cdot\sqrt{n}(Y_n - \theta)\). The first factor \(\xrightarrow{P} 0\) by step 1 and continuous mapping, the second converges in distribution, so by Slutsky \(\sqrt{n}\,R_n \xrightarrow{d} 0\), hence \(\xrightarrow{P} 0\).
3. Slutsky once more. With \(Z \sim N(0, \sigma^2)\), \[ \sqrt{n}\big(g(Y_n) - g(\theta)\big) = g'(\theta)\,\sqrt{n}(Y_n - \theta) + \sqrt{n}\,R_n \xrightarrow{d} g'(\theta)\,Z + 0 \sim N\!\big(0, \sigma^2[g'(\theta)]^2\big). \qquad\blacksquare \]
Let \(X_1, X_2, \ldots\) be iid exponential with mean \(\beta\), so the rate is \(\lambda = 1/\beta\). Then \(E X_i = \beta\) and \(\operatorname{Var} X_i = \beta^2\), so by the CLT \[ \sqrt{n}\,(\bar X_n - \beta) \xrightarrow{d} N(0, \beta^2). \]
The natural estimator of the rate is \(1/\bar X_n\). Take \(g(y) = 1/y\), so \(g'(\beta) = -1/\beta^2 \neq 0\), and \[ \sqrt{n}\left(\frac{1}{\bar X_n} - \lambda\right) \xrightarrow{d} N\!\left(0,\; \beta^2\cdot\frac{1}{\beta^4}\right) = N(0, \lambda^2). \]
So \(1/\bar X_n\) is approximately \(N(\lambda, \lambda^2/n)\), with standard error about \(\lambda/\sqrt{n}\). In Topic 6, \(1/\bar X_n\) turns out to be the maximum likelihood estimator of \(\lambda\), and \(\lambda^2\) reappears as the inverse of the Fisher information.
It is very common to use Taylor expansions in this way. When we want to understand the limiting behavior of “something complicated”, we use Taylor’s theorem to write
\[\textrm{Something complicated} = \textrm{Something much simpler} + \textrm{Something small}.\]
Many proofs then proceed by
The delta method also has a multivariate form (Theorem 5.5.28): if \(\sqrt{n}(Y_n - \theta) \xrightarrow{d} N(0, \Sigma)\) and \(g\) has Jacobian \(J = \nabla g(\theta)\), then \(\sqrt{n}\big(g(Y_n) - g(\theta)\big) \xrightarrow{d} N(0, J\Sigma J^{\top})\).
\[ X_n \xrightarrow{a.s.} X \quad\Longrightarrow\quad X_n \xrightarrow{P} X \quad\Longrightarrow\quad X_n \xrightarrow{d} X \]
| Mode | Statement about | Key result |
|---|---|---|
| Almost sure | \(P(\lim X_n = X) = 1\) | SLLN |
| In probability | \(\lim P(\lvert X_n - X\rvert \geq \varepsilon) = 0\) | WLLN |
| In distribution | \(\lim F_{X_n}(x) = F_X(x)\) at cts. points | CLT |
Toolbox: Continuous mapping, Slutsky, and the Delta method let us propagate these limits through the functions we actually care about.
Limit theory is how “approximately” is made precise.