In this lecture we return to a theoretical view of statistical inference. We will look in some detail at two questions:
We will mostly use material from Chapter 6 in Casella & Berger. The central idea is that of data reduction: a statistic summarises a whole sample into something simpler, and we want to understand when such a summary throws away nothing that matters about the parameter.
Suppose we want to estimate the expectation \(\mu\) in a \(N(\mu, \sigma^2)\) distribution from an iid sample \(x = (x_1, \ldots, x_n)\). Several arguments tell us that \(\bar{x} = \tfrac{1}{n}\sum_{i=1}^n x_i\) is a sensible estimator of \(\mu\):
But note that the sample only enters the mean through the sum of the observations. If we had another sample \(y = (y_1, \ldots, y_n)\) with \(\sum_{i=1}^n y_i = \sum_{i=1}^n x_i\), then both samples would lead to the same conclusions about \(\mu\).
\[\longrightarrow \text{In some sense, the two samples carry the \textbf{same information} about } \mu.\]
Before formalising this idea, let us fix some notation.
Any statistic \(T(\cdot)\) defines a form of data reduction or data summary. Following the example above, \(T(x) = x_1 + \cdots + x_n\) does not report all of the values in the sample — only a single number, the sum.
We will discuss two ways to analyse the information content in \(x\) and \(T(x)\) with respect to an unknown parameter. The book presents a third as well.
A sufficient statistic for a parameter \(\theta\) is a statistic that, in a certain sense, captures all the information about \(\theta\) contained in the sample. This leads to the following principle.
The sufficiency principle. If \(T(X)\) is a sufficient statistic for \(\theta\), then any inference about \(\theta\) should depend on the sample only through the value \(T(X)\). That is, if \(x\) and \(y\) are two samples with \(T(x) = T(y)\), then the inference about \(\theta\) should be the same whether \(X = x\) or \(X = y\) is observed.
We now make the notion of “capturing all the information” precise.
Definition. A statistic \(T(X)\) is a sufficient statistic for \(\theta\) if the conditional distribution of the sample \(X\) given the value \(T(X)\) does not depend on \(\theta\).
The definition concerns the conditional probability \(P_\theta\bigl(X = x \mid T(X) = T(x)\bigr)\). If \(T\) is sufficient, this quantity does not depend on \(\theta\).
Intuitively: once we know \(T(x)\), the parameter \(\theta\) tells us nothing more about how the remaining randomness in the sample was distributed. Everything that \(\theta\) has to say about the data is already contained in \(T(x)\).
Casella & Berger (p. 222) explain the definition through a thought experiment.
Under sufficiency, the conditional distribution of the sample is fully known — it does not depend on the unknown \(\theta\). So Experimenter 2 can generate a new sample \(Y\) from this distribution.
Whether it was Experimenter 1’s \(X\) or Experimenter 2’s \(Y\) that led to \(T(x)\) does not matter, because one can show that \[P_\theta(X = x) = P_\theta(Y = x),\] meaning both experimenters have the same knowledge about \(\theta\) — their likelihood functions are identical.
\[\longrightarrow \text{It is } \underline{\text{enough}} \text{ (\textit{sufficient}) to observe the value of a sufficient statistic.}\]
Checking the definition directly is awkward. The following theorem gives a workable criterion.
Theorem 6.2.2. Let \(p(x \mid \theta)\) be the joint pdf or pmf of \(X\), and let \(q(t \mid \theta)\) be the pdf or pmf of \(T(X)\). Then \(T(X)\) is a sufficient statistic for \(\theta\) if, for every \(x\) in the sample space, the ratio \[\frac{p(x \mid \theta)}{q\bigl(T(x) \mid \theta\bigr)}\] is constant as a function of \(\theta\).
Proof idea. Start from the conditional probability and use that \(\{X = x\} \subseteq \{T(X) = T(x)\}\):
\[\begin{align*} P_\theta\bigl(X = x \mid T(X) = T(x)\bigr) &= \frac{P_\theta\bigl(X = x \text{ and } T(X) = T(x)\bigr)}{P_\theta\bigl(T(X) = T(x)\bigr)} = \frac{P_\theta(X = x)}{P_\theta\bigl(T(X) = T(x)\bigr)} \\ &= \frac{p(x \mid \theta)}{q\bigl(T(x) \mid \theta\bigr)}. \end{align*}\] Under sufficiency the left-hand side does not depend on \(\theta\), so neither does the ratio on the right.
Example 6.2.4. Let \(X_1, \ldots, X_n\) be iid \(N(\mu, \sigma^2)\) with \(\sigma^2\) known. We show that \(T(X) = \bar{X} = (X_1 + \cdots + X_n)/n\) is sufficient for \(\mu\). The joint pdf is \[f(x \mid \mu) = \prod_{i=1}^n (2\pi\sigma^2)^{-1/2}\exp\!\left\{-\frac{(x_i - \mu)^2}{2\sigma^2}\right\} = (2\pi\sigma^2)^{-n/2}\exp\!\left\{-\sum_{i=1}^n \frac{(x_i - \mu)^2}{2\sigma^2}\right\}.\]
Adding and subtracting \(\bar{x}\) inside the square, and using \(\sum_{i=1}^n (x_i - \bar{x}) = 0\): \[\begin{align*} f(x \mid \mu) &= (2\pi\sigma^2)^{-n/2}\exp\!\left\{-\sum_{i=1}^n \frac{(x_i - \bar{x} + \bar{x} - \mu)^2}{2\sigma^2}\right\} \\&= (2\pi\sigma^2)^{-n/2}\exp\!\left\{-\frac{\sum_{i=1}^n (x_i - \bar{x})^2 + n(\bar{x} - \mu)^2}{2\sigma^2}\right\}. \end{align*}\]
We have previously shown that \(\bar{X} \sim N(\mu, \sigma^2/n)\). The ratio of Theorem 6.2.2 is therefore \[\frac{f(x \mid \mu)}{q\bigl(T(x) \mid \mu\bigr)} = \frac{(2\pi\sigma^2)^{-n/2}\exp\!\left\{-\bigl(\sum_{i=1}^n (x_i - \bar{x})^2 + n(\bar{x} - \mu)^2\bigr)/(2\sigma^2)\right\}}{(2\pi\sigma^2/n)^{-1/2}\exp\!\left\{-n(\bar{x} - \mu)^2/(2\sigma^2)\right\}}.\]
The terms involving \(\mu\) cancel, leaving \[\frac{f(x \mid \mu)}{q\bigl(T(x) \mid \mu\bigr)} = n^{-1/2}(2\pi\sigma^2)^{-(n-1)/2}\exp\!\left\{-\sum_{i=1}^n \frac{(x_i - \bar{x})^2}{2\sigma^2}\right\},\]
which does not depend on \(\mu\). By Theorem 6.2.2, \(T(X) = \bar{X}\) is a sufficient statistic for \(\mu\).
In the motivating example we argued that the sum \(s = \sum_{i=1}^n x_i\) was somehow “sufficient” for estimating \(\mu\). We then defined sufficiency formally and showed that the mean \(\bar{x} = s/n\) is sufficient. So which is it — the sum or the mean?
One might think that just knowing the sum is not enough: to estimate \(\mu\) we also need the sample size \(n\), so \(s\) alone is not sufficient.
But this is wrong. The sample size \(n\) is normally not considered part of the data — it is part of the experimental setup, known in principle before any observations are made. Whether we multiply \(s\) by the known constant \(1/n\) does not change anything, and is just a matter of convenience.
The next theorem lets us find a sufficient statistic by simply inspecting the pdf or pmf.
Theorem 6.2.6 (Factorization theorem). Let \(f(x \mid \theta)\) denote the joint pdf or pmf of a sample \(X\). A statistic \(T(X)\) is a sufficient statistic for \(\theta\) if and only if there exist functions \(g(t \mid \theta)\) and \(h(x)\) such that, for all sample points \(x\) and all parameter points \(\theta\), \[f(x \mid \theta) = g\bigl(T(x) \mid \theta\bigr)\, h(x).\]
The parameter \(\theta\) interacts with the data only through \(T(x)\), in the factor \(g\). The factor \(h(x)\) carries the rest of the dependence on \(x\), free of \(\theta\).
Casella & Berger prove this formally in the discrete case, but it makes very good sense given what we know about maximum likelihood. The log-likelihood of the sample is \[\ell(\theta \mid x) = \log f(x \mid \theta) = \log g\bigl(T(x) \mid \theta\bigr) + \log h(x).\]
When differentiating with respect to \(\theta\), the term \(\log h(x)\) vanishes, so the MLE can be expressed solely in terms of \(T\): \[\frac{\partial \ell(\theta \mid x)}{\partial \theta} = \frac{\partial}{\partial \theta}\log g\bigl(T(x) \mid \theta\bigr).\]
The entire likelihood analysis — and hence the MLE — depends on the data only through \(T(x)\).
Example 6.2.7. From the normal example above, the pdf factors as \[f(x \mid \mu) = \underbrace{(2\pi\sigma^2)^{-n/2}\exp\!\left\{-\sum_{i=1}^n (x_i - \bar{x})^2/(2\sigma^2)\right\}}_{h(x)}\;\underbrace{\exp\!\left\{-n(\bar{x} - \mu)^2/(2\sigma^2)\right\}}_{g(T(x)\,\mid\,\mu)}.\]
Here \(\sigma^2\) is treated as a known constant, so it may go into \(h(x)\). The second factor depends on \(x\) only through \(\bar{x} = T(x)\). By the factorization theorem, \(T(X) = \bar{X}\) is immediately seen to be sufficient for \(\mu\) — no ratio computation needed.
Sometimes we cannot summarise the information in the data into just a single number, so the sufficient statistic takes the form of a vector \[T(X) = \bigl(T_1(X), \ldots, T_r(X)\bigr).\]
This usually happens when we have several parameters \(\theta = (\theta_1, \ldots, \theta_s)\). The typical case is \(r = s\), and the factorization theorem works in exactly the same way.
Example 6.2.9. Now both \(\mu\) and \(\sigma^2\) are unknown and must be estimated from data. When using the factorization theorem, any part of the joint pdf that depends on \(\mu\) or \(\sigma^2\) must be included in the \(g\) factor.
Working with the expression from before, the pdf depends on the sample \(x\) only through the two values \[T_1(x) = \bar{x} \qquad \text{and} \qquad T_2(x) = s^2 = \frac{1}{n-1}\sum_{i=1}^n (x_i - \bar{x})^2.\]
We can therefore set \(h(x) = 1\) and write \(f(x \mid \mu, \sigma^2) = g\bigl(T_1(x), T_2(x) \mid \mu, \sigma^2\bigr)\,h(x)\). Hence \[T(X) = \bigl(\bar{X}, S^2\bigr)\] is a sufficient statistic for \((\mu, \sigma^2)\) in the normal model.
A particularly clean result holds for the exponential family.
Theorem 6.2.10. Let \(X_1, \ldots, X_n\) be iid observations from a pdf or pmf \(f(x \mid \theta)\) belonging to an exponential family \[f(x \mid \theta) = h(x)\, c(\theta)\, \exp\!\left(\sum_{j=1}^k w_j(\theta)\, t_j(x)\right),\] where \(\theta = (\theta_1, \ldots, \theta_d)\) with \(d \leq k\). Then \[T(X) = \left(\sum_{i=1}^n t_1(X_i), \;\ldots,\; \sum_{i=1}^n t_k(X_i)\right)\] is a sufficient statistic for \(\theta\).
Most of our standard distributions are exponential families, so this single result delivers their sufficient statistics directly.
Recall from the previous lecture that the likelihood function \(L(\theta \mid x) = f(x \mid \theta)\), viewed as a function of \(\theta\) for the fixed, observed \(x\), is the central object of likelihood-based inference.
The likelihood principle says, roughly, that this function contains everything the data have to say about \(\theta\): nothing else about how the data arose should matter.
The likelihood principle. If \(x\) and \(y\) are two sample points such that \(L(\theta \mid x)\) is proportional to \(L(\theta \mid y)\) — that is, there exists a constant \(C(x, y)\) with \[L(\theta \mid x) = C(x, y)\, L(\theta \mid y) \quad \text{for all } \theta,\] then the conclusions drawn from \(x\) and \(y\) should be identical.
Note how close this is to the sufficiency principle. In fact, the factorization theorem shows that sufficiency is a consequence of the likelihood principle.
If \(T(X)\) is sufficient, then \(f(x \mid \theta) = g\bigl(T(x) \mid \theta\bigr)h(x)\), so \[L(\theta \mid x) = h(x)\, g\bigl(T(x) \mid \theta\bigr).\]
If \(T(x) = T(y)\), then \[\frac{L(\theta \mid x)}{L(\theta \mid y)} = \frac{h(x)\, g\bigl(T(x) \mid \theta\bigr)}{h(y)\, g\bigl(T(y) \mid \theta\bigr)} = \frac{h(x)}{h(y)},\]
which does not depend on \(\theta\). So two samples with the same value of a sufficient statistic automatically have proportional likelihoods.
(A version of) Example 6.3.5. Suppose we want to learn about \(\theta\), the probability that a coin shows heads.
Experiment 1: flip the coin \(n\) times and count \(x\) heads. Then \(X \sim \text{Binomial}(n, \theta)\) and \[L_1(\theta \mid x) = \binom{n}{x}\theta^x (1 - \theta)^{n - x}.\]
Experiment 2: keep flipping until the \(x\)th head appears, and it happens to take \(n\) flips. Then \(N \sim \text{Negative Binomial}(x, \theta)\) and \[L_2(\theta \mid n) = \binom{n - 1}{x - 1}\theta^x (1 - \theta)^{n - x}.\]
The two likelihoods are proportional (both \(\propto \theta^x (1-\theta)^{n-x}\)), so the likelihood principle says our conclusions about \(\theta\) should be identical in the two experiments.
Yet classical procedures such as \(p\)-values or confidence intervals can give different answers in the two cases. The reason: they average over “what could have happened” under hypothetical repetitions, and these differ for the two sampling schemes — one fixes \(n\), the other fixes \(x\).
The likelihood principle says these differences are irrelevant: once \(x\) and \(n\) are observed, the two experiments carry exactly the same information about \(\theta\).
This is a proper controversial point in statistics — likelihood-based and Bayesian methods respect the likelihood principle, while many classical frequentist procedures do not.
Casella & Berger present a third principle, the equivariance principle, which says that if two problems have identical formal structure — the same model up to a relabelling of \(\theta\) — then they should be analysed by the same procedure, applied to the relabelled quantities.
We have already seen the most important instance of this in Theorem 7.2.10, the invariance of the MLE under reparameterization: if \(\hat\theta\) is the MLE of \(\theta\), then \(\tau(\hat\theta)\) is the MLE of \(\tau(\theta)\). This is exactly the equivariance principle applied to maximum likelihood estimation.
We will not develop this idea any further here.