BEA529 Probability & Statistical Inference
NHH · Autumn 2026

5 · Limit Theory

← Back to schedule

Let us warm up by recalling (?) what we mean when we say that a sequence of numbers converges:

Definition 3.1 in Baby Rudin. A sequence of numbers \(\{p_n\}\) is said to converge if there is a number \(p\) with the following property: for every \(\varepsilon > 0\) there is an integer \(N\) such that \(n \geq N\) implies that \(|p_n - p| < \varepsilon\).

We say that \(\{p_n\}\) converges to \(p\) or that \(p\) is the limit of \(\{p_n\}\), and we write \(p_n \to p\), or \[ \lim_{n \to \infty} p_n = p. \]

This concept is easy to understand: we can make the distance between \(p_n\) and \(p\) for all \(n \geq N\) smaller than any positive number by choosing \(N\) large enough.

We will next introduce several similar concepts (three, to be exact) for sequences of random variables instead of numbers.

Why is this useful? Because the notion of convergence in statistics gives us a more detailed understanding of the relationship / correspondence between sample statistics (such as the sample mean) and properties of the population (such as the population mean). For example, what is the role of the sample size \(n\)?

We have also mentioned the fact that the sample mean is “approximately normal”. The Central Limit Theorem is a statement about convergence that we need to formulate precisely.

More generally, we will see that statistical limit theory is of enormous practical use. If we are interested in a statistical property of the random variables \(X_n\) that is difficult to obtain, but that we also know that \(\{X_n\}\) “converges to” another random variable \(X\), we can approximate the intractable properties of \(X_n\) with the corresponding property of \(X\), as long as the word “convergence” means approximately the same thing (“close/similar for large \(n\)”).

Concept 1: Convergence in probability

Definition 5.5.1. A sequence of random variables \(X_1, X_2, \ldots\) converges in probability to a random variable \(X\) if, for every \(\varepsilon > 0\), \[ \lim_{n \to \infty} P\big(|X_n - X| \geq \varepsilon\big) = 0 \qquad\text{or, equivalently,}\qquad \lim_{n \to \infty} P\big(|X_n - X| < \varepsilon\big) = 1. \]

Note that the sequence \(X_1, X_2, \ldots\) does not represent a random (iid) sample; they do not share the same distribution and are typically not independent.

From CB: the distribution changes as the subscript changes, and the convergence concepts describe different ways in which the distribution of \(X_n\) converges to some limiting distribution (of \(X\)) as the subscript becomes large.

We are actually ready to state and prove our first major result, which formalizes the convergence of the sample mean to the population mean.

Theorem 5.5.2 (Weak Law of Large Numbers). Let \(X_1, X_2, \ldots\) be iid variables with \(E(X_i) = \mu\) and \(\operatorname{Var}(X_i) = \sigma^2 < \infty\). Define \(\bar{X}_n = n^{-1}\sum_{i=1}^{n} X_i\). Then, for every \(\varepsilon > 0\), \[ \lim_{n \to \infty} P\big(|\bar{X}_n - \mu| < \varepsilon\big) = 1, \] that is, \(\bar{X}_n\) converges in probability to \(\mu\).

Proof. Note that we assume \(\sigma^2 < \infty\). This is not necessary for the result, but it makes the proof straightforward using Chebychev’s inequality (Theorem 3.6.1): \[ P\big(|\bar{X}_n - \mu| \geq \varepsilon\big) = P\big((\bar{X}_n - \mu)^2 \geq \varepsilon^2\big) \leq \frac{E(\bar{X}_n - \mu)^2}{\varepsilon^2} = \frac{\operatorname{Var} \bar{X}_n}{\varepsilon^2} = \frac{\sigma^2}{n\varepsilon^2}. \] Hence, \[ P\big(|\bar{X}_n - \mu| < \varepsilon\big) = 1 - P\big(|\bar{X}_n - \mu| \geq \varepsilon\big) \geq 1 - \frac{\sigma^2}{n\varepsilon^2} \to 1 \quad\text{as } n \to \infty. \qquad \blacksquare \]

Example 5.5.3 uses the same technique to prove consistency of the sample variance, and Theorem 5.5.4 states that the WLLN is still valid under a continuous transformation:

Theorem 5.5.4. Suppose \(X_1, X_2, \ldots\) converges in probability to a random variable \(X\), and that \(h(\cdot)\) is a continuous function. Then \(h(X_1), h(X_2), \ldots\) converges in probability to \(h(X)\).

Concept 2: Almost sure convergence

Almost sure convergence is a stronger concept than convergence in probability, in the sense that if \(X_n\) converges to \(X\) almost surely, then \(X_n\) also converges to \(X\) in probability. The converse statement is not true in general.

Definition 5.5.6. A sequence of random variables \(X_1, X_2, \ldots\) converges almost surely to a random variable \(X\) if, for every \(\varepsilon > 0\), \[ P\left(\lim_{n \to \infty} |X_n - X| < \varepsilon\right) = 1. \]

Cosmetically it looks like we have “moved the limit” inside the probability operator, but note that the limit of probabilities and the (pointwise) limit of a function (recall that a random variable is a set function from the outcomes to the real numbers) are two quite different things.

The “almost” comes from the possibility of non-convergence on sets with probability zero.

We can indicate that almost sure convergence implies convergence in probability with the following argument:

Fix \(\varepsilon > 0\). Let \(B_n = \{|X_n - X| > \varepsilon\}\); we want to show that \(P(B_n) \to 0\).

Since \(X_n \overset{a.s.}{\to} X\), the set of “bad” outcomes \(\omega\) where \(X_n(\omega) \nrightarrow X(\omega)\) has probability zero. For every “good” \(\omega\) (and there are almost all of them), \(X_n(\omega)\) eventually stays within \(\varepsilon\) of \(X(\omega)\), meaning that \(\omega\) leaves \(B_n\) as soon as \(n\) becomes sufficiently large.

This means that the indicator function \(\mathbf{1}_{B_n}(\omega) \to 0\) for almost every \(\omega\). We can then use the dominated convergence theorem (see baby Rudin, Th. 11.32 for a general version) to write \[ P(B_n) = E\big(\mathbf{1}_{B_n}\big) \to 0. \]

Theorem 5.5.9 (Strong Law of Large Numbers). Let \(X_1, X_2, \ldots\) be iid random variables with \(E(X_i) = \mu\) and \(\operatorname{Var}(X_i) = \sigma^2 < \infty\), and define \(\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i\). Then, for every \(\varepsilon > 0\), \[ P\left(\lim_{n \to \infty} |\bar{X}_n - \mu| < \varepsilon\right) = 1. \]

This result can also be proved by some smart applications of Chebyshev’s inequality. It also holds without assuming \(\sigma^2 < \infty\), but is much harder to prove.

Concept 3: Convergence in distribution

This concept is weaker than convergence in probability.

Definition 5.5.10. A sequence of random variables \(X_1, X_2, \ldots\) converges in distribution to a random variable \(X\) if \[ \lim_{n \to \infty} F_{X_n}(x) = F_X(x) \] at all points where \(F_X(x)\) is continuous.

This is the convergence type that has the most practical importance to us, most importantly because we will soon prove that the sample mean converges in distribution to the normal distribution (almost) regardless of the distribution of the individual components.

Let us warm up with the following example, where we in particular must notice the concepts normalization and convergence speed.

Example 5.5.11 (maximum of uniforms). If \(X_1, X_2, \ldots\) are iid uniform, and \(X_{(n)} = \max_{1 \leq i \leq n} X_i\), let us examine if (and to where) \(X_{(n)}\) converges in distribution.

Intuition: as we observe more and more independent uniform \((0,1)\), we expect the maximum value to become closer and closer to 1.

\(X_{(n)}\) is always less than 1, so for any \(\varepsilon > 0\), we have that \[ P\big(|X_{(n)} - 1| \geq \varepsilon\big) = P\big(X_{(n)} \leq 1 - \varepsilon\big). \]

Since the observations are independent, we can continue: \[ P\big(X_{(n)} \leq 1 - \varepsilon\big) = P\big(X_i \leq 1 - \varepsilon \text{ for all } i\big) = (1 - \varepsilon)^n, \] because the probability for a \(U(0,1)\) variable to fall in a specific interval in \([0,1]\) is the length of that interval. Since \((1-\varepsilon)^n \to 0\) for any \(0 < \varepsilon < 1\) (and \(P(X_{(n)} \leq 1 - \varepsilon) \equiv 0\) if \(\varepsilon \geq 1\)), we have shown that \(X_{(n)}\) converges in probability to the constant 1.

If we let \(\varepsilon = t/n\) for any \(t > 0\) (that is, a specific sequence that converges at a specific speed), we have \[ P\left(X_{(n)} \leq 1 - \frac{t}{n}\right) = \left(1 - \frac{t}{n}\right)^n \to e^{-t}, \] which is the same as saying that \[ P\big(n(1 - X_{(n)}) \leq t\big) \to 1 - e^{-t}. \]

In other words: \(n(1 - X_{(n)})\) converges in distribution to an exponential\((1)\) random variable.

Take some time to understand this result: \((1 - X_{(n)})\) goes to zero, and \(n\) goes to infinity. When we multiply them together we get something that does not go to zero, and also not to infinity; these two quantities go in opposite directions at “the same speed” in some sense, and we can say that the rate of convergence for \((1 - X_{(n)})\) is \(1/n\).

\(\rightarrow\) See exercise 5.42 in CB for further exploration of this.

We are ready to state (a version of) the central limit theorem:

Theorem 5.5.14. Let \(X_1, X_2, \ldots\) be a sequence of iid random variables whose mgfs exist in a neighbourhood of \(0\) (that is, \(M_{X_i}(t)\) exists for \(|t| < h\) for some positive \(h\)). Let \(E X_i = \mu\) and \(\operatorname{Var} X_i = \sigma^2 > 0\) (both are finite because the mgf exists). Define \(\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i\). Let \(G_n(x)\) denote the cdf of \(\sqrt{n}(\bar{X}_n - \mu)/\sigma\). Then, for any \(x\), \[ \lim_{n \to \infty} G_n(x) = \int_{-\infty}^{x} \frac{1}{\sqrt{2\pi}} e^{-y^2/2}\,dy, \] that is, \(\sqrt{n}(\bar{X}_n - \mu)/\sigma\) has a limiting standard normal distribution.

Proof. The existence of moment generating functions together with the iid assumption make this proof rather straightforward. The result still holds under much more general conditions, but that would also make the proof much harder.

The idea is to show that the mgf of \(\sqrt{n}(\bar{X}_n - \mu)/\sigma\) converges to \(\exp\{t^2/2\}\), which is the mgf of a \(N(0,1)\) variable. By Theorem 2.3.12, convergence of mgfs in a neighbourhood of \(0\) implies convergence of the cdfs at every continuity point, which is what the theorem claims.

Define \(Y_i = (X_i - \mu)/\sigma\), and let \(M_Y(t)\) denote the common mgf for the \(Y_i\)s, which exists for \(|t| < \sigma h\). We have that \[ \frac{\sqrt{n}(\bar{X}_n - \mu)}{\sigma} = \frac{1}{\sqrt{n}}\sum_{i=1}^{n} Y_i, \] and from Theorems 2.3.15 and 4.6.7 we can write \[ M_{\sqrt{n}(\bar{X}_n - \mu)/\sigma}(t) = M_{\sum Y_i/\sqrt{n}}(t) = M_{\sum Y_i}\!\left(\frac{t}{\sqrt{n}}\right) = \left(M_Y\!\left(\frac{t}{\sqrt{n}}\right)\right)^n. \]

The next step is to expand \(M_Y(t/\sqrt{n})\) around \(0\) as a Taylor series: \[ M_Y\!\left(\frac{t}{\sqrt{n}}\right) = \sum_{k=0}^{\infty} M_Y^{(k)}(0) \frac{(t/\sqrt{n})^k}{k!}, \] where \(M_Y^{(k)}(0) = \frac{d^k}{dt^k} M_Y(t)\big|_{t=0}\) (note that if the mgf is finite in an interval around zero, it is also differentiable in zero at all orders). Since \(M_Y\) exists for \(|t| < \sigma h\), the power series expansion is valid for \(|t| < \sqrt{n}\,\sigma h\).

Using the facts that \(M_Y(0) = 1\), \(M_Y^{(1)}(0) = 0\) and \(M_Y^{(2)}(0) = 1\), we can simplify this to \[ M_Y\!\left(\frac{t}{\sqrt{n}}\right) = 1 + \frac{(t/\sqrt{n})^2}{2!} + R_Y\!\left(\frac{t}{\sqrt{n}}\right), \] where \(R_Y(t/\sqrt{n})\) is the remainder term in the Taylor expansion, \[ R_Y\!\left(\frac{t}{\sqrt{n}}\right) = \sum_{k=3}^{\infty} M_Y^{(k)}(0) \frac{(t/\sqrt{n})^k}{k!}. \]

One important property of the Taylor expansion is that the remainder always tends to zero faster than the preceding terms (see Th. 5.5.21 with \(r = 3\)), which means that \[ \begin{aligned} \lim_{n \to \infty} \frac{R_Y(t/\sqrt{n})}{(t/\sqrt{n})^2} &= 0 \\ \Longrightarrow\quad \lim_{n \to \infty} \frac{R_Y(t/\sqrt{n})}{(1/\sqrt{n})^2} &= \lim_{n \to \infty} n R_Y\!\left(\frac{t}{\sqrt{n}}\right) = 0, \qquad 0 < |t| < \sqrt{n}\,\sigma h. \end{aligned} \]

This is also true at \(t = 0\) (identically, without the limit), so for any \(|t| < \sqrt{n}\,\sigma h\), \[ \begin{aligned} \lim_{n \to \infty} \left(M_Y\!\left(\frac{t}{\sqrt{n}}\right)\right)^n &= \lim_{n \to \infty} \left[1 + \frac{(t/\sqrt{n})^2}{2!} + R_Y\!\left(\frac{t}{\sqrt{n}}\right)\right]^n \\ &= \lim_{n \to \infty} \left[1 + \frac{1}{n}\left(\frac{t^2}{2} + n R_Y\!\left(\frac{t}{\sqrt{n}}\right)\right)\right]^n \\ &= e^{t^2/2}. \end{aligned} \] The last step is Lemma 2.3.14: if \(a_n \to a\), then \((1 + a_n/n)^n \to e^a\). Here \(a_n = t^2/2 + nR_Y(t/\sqrt{n}) \to t^2/2\). By Theorem 2.3.12, the cdf of \(\sqrt{n}(\bar{X}_n - \mu)/\sigma\) therefore converges to the standard normal cdf at every \(x\). \(\qquad \blacksquare\)

Note, again, that we can read the convergence rate directly from the result: \((\bar{X} - \mu)/\sigma\) will become smaller, and \(\sqrt{n}\) becomes bigger as \(n \to \infty\), but in such a way that their product does neither, but rather behaves more and more like a \(N(0,1)\) random variable.

This rate is considered as one of “The Seven Pillars of Statistical Wisdom” by Stephen Stigler, who notes that, when this rate was discovered by Laplace (1810), “the implications were striking: if you wished to double the accuracy of an investigation, it was insufficient to double the effort; you must increase the effort fourfold.”

Note also that the CLT holds under much more general conditions: we can drop both “i”s in iid. This, however, makes the result much harder to prove.

The following results are very useful in “practical” theoretical statistics:

Theorem 5.5.17 (Slutsky’s theorem). If \(X_n \to X\) in distribution and \(Y_n \to a\), a constant, in probability, then

  1. \(Y_n X_n \to a X\) in distribution.
  2. \(X_n + Y_n \to X + a\) in distribution.

Theorem 5.5.24 (the delta method). Let \(Y_n\) be a sequence of random variables that satisfy \(\sqrt{n}(Y_n - \theta) \to N(0, \sigma^2)\) in distribution. For a given function \(g\) and a specific value of \(\theta\), suppose that \(g'(\theta)\) exists and is not zero. Then \[ \sqrt{n}\big(g(Y_n) - g(\theta)\big) \to N\!\left(0, \sigma^2 [g'(\theta)]^2\right) \quad\text{in distribution.} \]

Proof. The Taylor expansion of \(g(Y_n)\) around \(Y_n = \theta\) is \[ g(Y_n) = g(\theta) + g'(\theta)(Y_n - \theta) + R_n, \qquad \frac{R_n}{Y_n - \theta} \to 0 \text{ as } Y_n \to \theta. \] It is not enough that \(R_n \to 0\): after multiplying by \(\sqrt{n}\), we need \(\sqrt{n}\,R_n \to 0\). This takes three steps.

  1. \(Y_n \to \theta\) in probability. Write \(Y_n - \theta = n^{-1/2}\cdot\sqrt{n}(Y_n - \theta)\). The first factor tends to \(0\) and the second converges in distribution, so by Slutsky \(Y_n - \theta \to 0\) in distribution. Convergence in distribution to a constant is convergence in probability (Th. 5.5.13).
  2. The remainder is negligible. Write \(\sqrt{n}\,R_n = \dfrac{R_n}{Y_n - \theta}\cdot\sqrt{n}(Y_n - \theta)\). The first factor tends to \(0\) in probability by step 1 and continuous mapping, and the second converges in distribution, so by Slutsky \(\sqrt{n}\,R_n \to 0\) in distribution, and hence in probability.
  3. Slutsky once more. With \(Z \sim N(0, \sigma^2)\), \[\sqrt{n}\big(g(Y_n) - g(\theta)\big) = g'(\theta)\sqrt{n}(Y_n - \theta) + \sqrt{n}\,R_n \to g'(\theta) Z + 0\] in distribution, and \(g'(\theta)Z \sim N\big(0, \sigma^2[g'(\theta)]^2\big)\). \(\qquad \blacksquare\)

The condition \(g'(\theta) \neq 0\) is not needed for the argument to go through. If \(g'(\theta) = 0\), the limit is \(N(0, 0)\), a point mass at \(0\): true, but it says nothing about the distribution of \(g(Y_n)\). A second-order expansion is then needed instead; see problem A2.

Note. It is very common to use Taylor expansions in this way. When we want to understand the limiting behavior of “something complicated”, we use Taylor’s theorem to write

Something complicated = Something much simpler + Something small.

Many proofs will then proceed by (i) deriving/proving the desired property for the simple thing, and (ii) showing that the small thing converges to zero.

The delta method has a general multivariate form (Th. 5.5.28): suppose \(Y_n\) is a sequence of random vectors satisfying \(\sqrt{n}(Y_n - \theta) \overset{d}{\to} N(0, \Sigma)\), and let \(g : \mathbb{R}^k \to \mathbb{R}^m\) be differentiable at \(\theta\) with Jacobian \(J = \nabla g(\theta)\). Then \[ \sqrt{n}\big(g(Y_n) - g(\theta)\big) \overset{d}{\to} N\!\left(0, J \Sigma J^{\top}\right). \]