5. Limit Theory

Convergence of numbers, and why we need it

Definition (Rudin 3.1). A sequence \(\{p_n\}\) of numbers converges if there exists a number \(p\) such that for every \(\varepsilon > 0\) there is an integer \(N\) with \[ n \geq N \;\Longrightarrow\; |p_n - p| < \varepsilon. \] We write \(p_n \to p\), or \(\displaystyle\lim_{n\to\infty} p_n = p\).

We can make the distance between \(p_n\) and \(p\) smaller than any positive number, for all \(n \geq N\), by choosing \(N\) large enough.

Why we need the same idea for random variables:

  • It describes precisely how sample statistics (e.g. \(\bar X_n\)) relate to population parameters (e.g. \(\mu\)), and what the sample size \(n\) does.
  • We have claimed that the sample mean is “approximately normal”. The Central Limit Theorem makes that claim precise.
  • If a property of \(X_n\) is hard to obtain, but \(X_n\) converges to some \(X\), we can approximate it by the corresponding property of \(X\).

Three concepts of convergence for random variables

For sequences \(X_1, X_2, \ldots\) of random variables we will distinguish:

  1. Convergence in probability \(\quad X_n \xrightarrow{P} X\)
  2. Almost sure convergence \(\quad X_n \xrightarrow{a.s.} X\)
  3. Convergence in distribution \(\quad X_n \xrightarrow{d} X\)

Important. The sequence \(X_1, X_2, \ldots\) is not an iid sample. The \(X_n\) have different distributions in general and are typically dependent. The whole point is to describe how the distribution of \(X_n\) changes as \(n\) grows.

From CB: “The distribution changes as the subscript changes and the convergence concepts describe different ways in which the distribution of \(X_n\) converges to some limiting distribution (of \(X\)) as the subscript becomes large.”

Concept 1 — Convergence in probability

Definition 5.5.1. A sequence \(X_1, X_2, \ldots\) of random variables converges in probability to \(X\) if, for every \(\varepsilon > 0\), \[ \lim_{n\to\infty} P\big(|X_n - X| \geq \varepsilon\big) = 0, \] equivalently \(\lim_{n\to\infty} P\big(|X_n - X| < \varepsilon\big) = 1\).

Notation: \(X_n \xrightarrow{P} X\).

Weak Law of Large Numbers

We are ready to state and prove our first major result, which formalizes the convergence of the sample mean to the population mean:

Theorem 5.5.2 (Weak Law of Large Numbers). Let \(X_1, X_2, \ldots\) be iid with \(E[X_i] = \mu\) and \(\operatorname{Var}(X_i) = \sigma^2 < \infty\), and set \[ \bar X_n = \frac{1}{n}\sum_{i=1}^n X_i. \] Then for every \(\varepsilon > 0\), \[ \lim_{n\to\infty} P\big(|\bar X_n - \mu| < \varepsilon\big) = 1, \] i.e. \(\bar X_n \xrightarrow{P} \mu\).

Why the WLLN should not be surprising

We know from Topic 4 (Theorem 5.2.6) that \(E(\bar X_n) = \mu\) and \(\operatorname{Var}(\bar X_n) = \sigma^2/n\). The centre stays at \(\mu\) while the spread shrinks. The curve is the \(N(\mu, \sigma^2/n)\) density, which is the distribution of \(\bar X_n\) for a normal population.

Even though \(\bar X_n\) is random, it becomes more and more like the constant \(\mu\): for every \(\varepsilon > 0\), the probability that \(\bar X_n\) lands more than \(\varepsilon\) away from \(\mu\) converges to zero as \(n \to \infty\).

Proof of the WLLN via Chebyshev

Chebyshev’s inequality (Theorem 3.6.1). Let \(X\) be a random variable and \(g(x) \geq 0\). Then for any \(r > 0\), \[P\big(g(X) \geq r\big) \leq \frac{E\,g(X)}{r}.\]

Apply it with \(g(x) = (x - \mu)^2\) and \(r = \varepsilon^2\): \[ P\big(|\bar X_n - \mu| \geq \varepsilon\big) = P\big((\bar X_n - \mu)^2 \geq \varepsilon^2\big) \leq \frac{E(\bar X_n - \mu)^2}{\varepsilon^2}. \]

The numerator is \(\operatorname{Var}(\bar X_n) = \sigma^2/n\), by Theorem 5.2.6. Hence \[ P\big(|\bar X_n - \mu| < \varepsilon\big) \;\geq\; 1 - \frac{\sigma^2}{n\varepsilon^2} \;\xrightarrow{n\to\infty}\; 1. \qquad\blacksquare \]

The finite variance is not needed for the result, but it makes the proof this short.

Continuous mapping

Theorem 5.5.4 (Continuous mapping). If \(X_n \xrightarrow{P} X\) and \(h\) is a continuous function, then \[ h(X_n) \xrightarrow{P} h(X). \]

Convergence in probability is preserved under continuous transformations, e.g. \(\bar X_n \xrightarrow{P} \mu \Rightarrow \bar X_n^2 \xrightarrow{P} \mu^2\).

The same Chebyshev argument shows that \(S_n^2 \xrightarrow{P} \sigma^2\) (Example 5.5.3, problem P3): the sample variance is a consistent estimator of \(\sigma^2\). By continuous mapping, \(S_n \xrightarrow{P} \sigma\).

Concept 2 — Almost sure convergence

Definition 5.5.6. A sequence \(X_1, X_2, \ldots\) of random variables converges almost surely to \(X\) if, for every \(\varepsilon > 0\), \[ P\!\left(\lim_{n\to\infty} |X_n - X| < \varepsilon\right) = 1. \] Notation: \(X_n \xrightarrow{a.s.} X\).

  • It looks as if the limit has moved inside the probability. But a limit of probabilities and the pointwise limit of the functions \(X_n(\omega)\) are different things.
  • “Almost” means that non-convergence is allowed on a set of probability zero.
  • Almost sure convergence implies convergence in probability. The converse is false: problem A1 gives a sequence that converges in probability, but for no \(\omega\).

Theorem 5.5.9 (Strong Law of Large Numbers). Under the conditions of the WLLN, \(\bar X_n \xrightarrow{a.s.} \mu\).

Why almost sure convergence implies convergence in probability: notes, Concept 2: Almost sure convergence.

Concept 3 — Convergence in distribution

Definition 5.5.10. A sequence \(X_1, X_2, \ldots\) converges in distribution to \(X\) if \[ \lim_{n\to\infty} F_{X_n}(x) = F_X(x) \] at every point \(x\) where \(F_X\) is continuous. Notation: \(X_n \xrightarrow{d} X\).

This is the weakest of the three, but the most important in practice, because we will soon prove that the sample mean converges in distribution to the normal distribution (almost) regardless of the distribution of the individual observations.

Example 5.5.11 — Maximum of uniforms

Let \(X_1, X_2, \ldots\) be iid Uniform\((0,1)\) and \(X_{(n)} = \max_{1 \leq i \leq n} X_i\).

Intuition: As we observe more uniforms, the maximum should approach \(1\).

Since \(X_{(n)} < 1\), the event \(|X_{(n)} - 1| \geq \varepsilon\) is the event \(X_{(n)} \leq 1 - \varepsilon\). For any \(\varepsilon \in (0,1)\), \[ \begin{align*} P\big(X_{(n)} \leq 1 - \varepsilon\big) &= P(X_1 \leq 1 - \varepsilon, \ldots, X_n \leq 1 - \varepsilon) && \text{(the max is below iff all are)}\\ &= \prod_{i=1}^n P(X_i \leq 1 - \varepsilon) && \text{(independence)}\\ &= (1 - \varepsilon)^n \to 0 && \text{(uniform cdf: } F(x) = x\text{)}. \end{align*} \]

So \(X_{(n)} \xrightarrow{P} 1\). (For \(\varepsilon \geq 1\) the probability is \(0\).)

Simulation: \(X_{(n)} \xrightarrow{P} 1\)

Each of 2000 repetitions draws \(n\) uniforms and records the maximum.

Example 5.5.11 — The rate of convergence

Now fix \(t > 0\) and substitute \(\varepsilon = t/n\) (a specific sequence converging to \(0\) at a specific speed): \[ P\!\left(X_{(n)} \leq 1 - \tfrac{t}{n}\right) = \left(1 - \tfrac{t}{n}\right)^n \xrightarrow{n\to\infty} e^{-t}. \]

Equivalently, \[ P\big(n(1 - X_{(n)}) \leq t\big) \to 1 - e^{-t}, \] which is the CDF of an Exponential\((1)\) random variable: \(n(1 - X_{(n)}) \xrightarrow{d} \operatorname{Exp}(1)\).

Take some time to understand this result: \((1 - X_{(n)})\) goes to zero, and \(n\) goes to infinity. When we multiply them together we get something that does not go to zero, and also not to infinity. These two quantities go in opposite directions “at the same speed” in some sense, and we can say that the rate of convergence of \((1 - X_{(n)})\) is \(n\). See CB Exercise 5.42 for further exploration.

Simulation: \(n(1 - X_{(n)}) \xrightarrow{d} \operatorname{Exp}(1)\)

Each of 4000 repetitions draws \(n\) uniforms. Left: the average distance from the maximum to \(1\) against \(n\), on log scales, with a point for each \(n\) visited. Right: the distance multiplied by \(n\).

The Central Limit Theorem

Theorem 5.5.14 (Central Limit Theorem). Let \(X_1, X_2, \ldots\) be iid with \(E[X_i] = \mu\), \(\operatorname{Var}(X_i) = \sigma^2 > 0\) (both finite), and suppose \(M_{X_i}(t)\) exists for \(|t| < h\), for some \(h > 0\). Let \[ G_n(x) = P\!\left(\frac{\sqrt{n}\,(\bar X_n - \mu)}{\sigma} \leq x\right). \] Then for every \(x \in \mathbb{R}\), \[ \lim_{n\to\infty} G_n(x) = \int_{-\infty}^x \frac{1}{\sqrt{2\pi}}\operatorname{exp}\left(-y^2/2\right)\,\textrm{d}y, \] that is, \(\dfrac{\sqrt{n}\,(\bar X_n - \mu)}{\sigma} \xrightarrow{d} N(0,1)\).

Simulation: \(\sqrt{n}\,(\bar X_n - \mu)/\sigma \xrightarrow{d} N(0,1)\)

Each panel shows 3000 simulated values of \(\sqrt{n}\,(\bar X_n - \mu)/\sigma\) for samples of size \(n\) from the population named above it. The dashed curve is the \(N(0,1)\) density, and all panels share the same scales.

Proof of the CLT — setup

Let \(Y_i = (X_i - \mu)/\sigma\). Then \(E[Y_i] = 0\), \(\operatorname{Var}(Y_i) = 1\), and the common mgf \(M_Y(t)\) exists for \(|t| < \sigma h\). Rewrite the standardized mean in terms of the \(Y_i\): \[ \frac{\sqrt{n}\,(\bar X_n - \mu)}{\sigma} = \sqrt{n}\cdot\frac{1}{n}\sum_{i=1}^n \frac{X_i - \mu}{\sigma} = \frac{1}{\sqrt{n}}\sum_{i=1}^n Y_i. \]

Its mgf follows from two earlier results: \[ M_{\sum_{i} Y_i / \sqrt{n}}(t) \underset{\text{Thm 2.3.15}}{=} M_{\sum_{i} Y_i}\!\left(\frac{t}{\sqrt{n}}\right) \underset{\text{Thm 4.6.7}}{=} \Big(M_Y\big(t/\sqrt{n}\big)\Big)^{n}, \] first scaling (\(M_{aZ}(t) = M_Z(at)\)), then the mgf of a sum of independent variables.

Strategy. Show that \(\big(M_Y(t/\sqrt{n})\big)^n \to e^{t^2/2}\), the mgf of \(N(0,1)\). By Theorem 2.3.12, convergence of mgfs in a neighbourhood of \(0\) implies convergence of the cdfs at every continuity point, which is what the CLT claims.

Proof of the CLT — Taylor expansion

Expand \(M_Y(t/\sqrt{n})\) as a Taylor series around \(0\), valid for \(|t| < \sqrt{n}\,\sigma h\): \[ M_Y(t/\sqrt{n}) = \sum_{k=0}^\infty M_Y^{(k)}(0)\,\frac{(t/\sqrt{n})^k}{k!}. \] If the mgf is finite in an interval around zero, it is differentiable at zero to all orders.

Using \(M_Y(0) = 1\), \(M_Y'(0) = E[Y] = 0\) and \(M_Y''(0) = E[Y^2] = 1\): \[ M_Y(t/\sqrt{n}) = 1 + \frac{(t/\sqrt{n})^2}{2} + R_Y(t/\sqrt{n}), \qquad R_Y(u) = \sum_{k=3}^\infty M_Y^{(k)}(0)\,\frac{u^k}{k!}. \]

The remainder vanishes faster than the last term kept (Theorem 5.5.21): \(R_Y(u)/u^2 \to 0\) as \(u \to 0\). With \(u = t/\sqrt{n}\) and \(t \neq 0\) fixed, \[ n\,R_Y(t/\sqrt{n}) = t^2\cdot\frac{R_Y(t/\sqrt{n})}{(t/\sqrt{n})^2} \;\xrightarrow{n\to\infty}\; t^2 \cdot 0 = 0, \] and at \(t = 0\) it is \(0\) for every \(n\).

Proof of the CLT — the limit

Lemma 2.3.14. If \(a_n \to a\), then \(\displaystyle\lim_{n\to\infty}\left(1 + \frac{a_n}{n}\right)^n = e^a\).

Collect the terms into this form: \[ \big[M_Y(t/\sqrt{n})\big]^n = \left[1 + \frac{1}{n}\underbrace{\left(\frac{t^2}{2} + n\,R_Y(t/\sqrt{n})\right)}_{a_n \;\to\; t^2/2}\right]^n. \]

By the lemma, \(\big[M_Y(t/\sqrt{n})\big]^n \to e^{t^2/2}\), the mgf of \(N(0,1)\). By Theorem 2.3.12, \(G_n(x) \to \Phi(x)\) for every \(x\). \(\qquad\blacksquare\)

What the CLT says

Rate of convergence. \(\bar X_n - \mu\) shrinks and \(\sqrt{n}\) grows, and their product behaves like a normal variable: the sample mean converges at rate \(\sqrt{n}\). This is the law of diminishing returns for data: going from 10 to 11 observations is worth more than going from 1000 to 1001.

Stephen Stigler counts this rate among The Seven Pillars of Statistical Wisdom. When Laplace discovered it (1810), “the implications were striking: if you wished to double the accuracy of an investigation, it was insufficient to double the effort; you must increase the effort fourfold.”

The CLT holds under much more general conditions: we can drop both “i”s in iid. Theorem 5.5.15 needs only finite variance, no mgf. The general versions are much harder to prove.

Slutsky’s theorem

The following results are very useful in “practical” theoretical statistics.

Theorem 5.5.17 (Slutsky). Suppose \(X_n \xrightarrow{d} X\) and \(Y_n \xrightarrow{P} a\) for some constant \(a\). Then

  1. \(\;Y_n X_n \xrightarrow{d} a\,X\),
  2. \(\;X_n + Y_n \xrightarrow{d} X + a\).

Extremely useful: you can replace a nuisance quantity that converges to a constant (e.g. \(S_n \xrightarrow{P} \sigma\)) without changing the limit distribution.

Slutsky in action: the \(t\)-statistic without normality

From the previous topic, we have for a normal sample: \(\dfrac{\bar X_n - \mu}{S_n/\sqrt{n}} \sim t_{n-1}\) exactly, which gave the interval \(\bar X_n \pm t_{n-1,\alpha/2}\,S_n/\sqrt{n}\).

For any population with finite variance \(\sigma^2 > 0\): \[ \frac{\sqrt{n}\,(\bar X_n - \mu)}{S_n} \xrightarrow{d} N(0,1). \] The proof combines the CLT, the consistency of \(S_n\), and Slutsky: problem P10.

  • So \(P\big(\bar X_n - z_{\alpha/2}\,S_n/\sqrt{n} \leq \mu \leq \bar X_n + z_{\alpha/2}\,S_n/\sqrt{n}\big) \to 1 - \alpha\), whatever the population.
  • Since \(t_{n-1,\alpha/2} \to z_{\alpha/2}\), the \(t\)-interval is approximately right too.
  • Normality made the interval exact. Limit theory makes it approximate, for every population.

The Delta method

Theorem 5.5.24 (Delta method). Let \(Y_n\) satisfy \(\sqrt{n}\,(Y_n - \theta) \xrightarrow{d} N(0, \sigma^2)\). For a function \(g\), if \(g'(\theta)\) exists and \(g'(\theta) \neq 0\), then \[ \sqrt{n}\,\big(g(Y_n) - g(\theta)\big) \xrightarrow{d} N\!\big(0,\; \sigma^2 [g'(\theta)]^2\big). \]

Why \(g'(\theta) \neq 0\). If \(g'(\theta) = 0\), the limit is \(N(0, 0)\), a point mass at \(0\). That is true, but it says nothing about the distribution of \(g(Y_n)\). A second-order expansion is needed instead: problem A2.

Proof of the Delta method

Taylor-expand \(g\) around \(\theta\): \[ g(Y_n) = g(\theta) + g'(\theta)(Y_n - \theta) + R_n, \qquad \frac{R_n}{Y_n - \theta} \to 0 \text{ as } Y_n \to \theta. \]

1. \(Y_n \xrightarrow{P} \theta\). Write \(Y_n - \theta = \frac{1}{\sqrt{n}}\cdot\sqrt{n}(Y_n - \theta)\). The first factor tends to \(0\) and the second converges in distribution, so by Slutsky \(Y_n - \theta \xrightarrow{d} 0\). Convergence in distribution to a constant is convergence in probability (Theorem 5.5.13; problem P5).

2. The remainder is negligible. Write \(\sqrt{n}\,R_n = \dfrac{R_n}{Y_n - \theta}\cdot\sqrt{n}(Y_n - \theta)\). The first factor \(\xrightarrow{P} 0\) by step 1 and continuous mapping, the second converges in distribution, so by Slutsky \(\sqrt{n}\,R_n \xrightarrow{d} 0\), hence \(\xrightarrow{P} 0\).

3. Slutsky once more. With \(Z \sim N(0, \sigma^2)\), \[ \sqrt{n}\big(g(Y_n) - g(\theta)\big) = g'(\theta)\,\sqrt{n}(Y_n - \theta) + \sqrt{n}\,R_n \xrightarrow{d} g'(\theta)\,Z + 0 \sim N\!\big(0, \sigma^2[g'(\theta)]^2\big). \qquad\blacksquare \]

Delta method: the exponential rate

Let \(X_1, X_2, \ldots\) be iid exponential with mean \(\beta\), so the rate is \(\lambda = 1/\beta\). Then \(E X_i = \beta\) and \(\operatorname{Var} X_i = \beta^2\), so by the CLT \[ \sqrt{n}\,(\bar X_n - \beta) \xrightarrow{d} N(0, \beta^2). \]

The natural estimator of the rate is \(1/\bar X_n\). Take \(g(y) = 1/y\), so \(g'(\beta) = -1/\beta^2 \neq 0\), and \[ \sqrt{n}\left(\frac{1}{\bar X_n} - \lambda\right) \xrightarrow{d} N\!\left(0,\; \beta^2\cdot\frac{1}{\beta^4}\right) = N(0, \lambda^2). \]

So \(1/\bar X_n\) is approximately \(N(\lambda, \lambda^2/n)\), with standard error about \(\lambda/\sqrt{n}\). In Topic 6, \(1/\bar X_n\) turns out to be the maximum likelihood estimator of \(\lambda\), and \(\lambda^2\) reappears as the inverse of the Fisher information.

A pattern to notice

It is very common to use Taylor expansions in this way. When we want to understand the limiting behavior of “something complicated”, we use Taylor’s theorem to write

\[\textrm{Something complicated} = \textrm{Something much simpler} + \textrm{Something small}.\]

Many proofs then proceed by

  1. deriving the desired property for the simple thing,
  2. showing that the small thing converges to zero, and
  3. applying Slutsky’s theorem to glue the pieces together.

The delta method also has a multivariate form (Theorem 5.5.28): if \(\sqrt{n}(Y_n - \theta) \xrightarrow{d} N(0, \Sigma)\) and \(g\) has Jacobian \(J = \nabla g(\theta)\), then \(\sqrt{n}\big(g(Y_n) - g(\theta)\big) \xrightarrow{d} N(0, J\Sigma J^{\top})\).

Summary

\[ X_n \xrightarrow{a.s.} X \quad\Longrightarrow\quad X_n \xrightarrow{P} X \quad\Longrightarrow\quad X_n \xrightarrow{d} X \]

Mode Statement about Key result
Almost sure \(P(\lim X_n = X) = 1\) SLLN
In probability \(\lim P(\lvert X_n - X\rvert \geq \varepsilon) = 0\) WLLN
In distribution \(\lim F_{X_n}(x) = F_X(x)\) at cts. points CLT

Toolbox: Continuous mapping, Slutsky, and the Delta method let us propagate these limits through the functions we actually care about.

Limit theory is how “approximately” is made precise.