7 · Principles of Data Reduction
In this lecture we will return to a theoretical view of statistical inference, and we will do this by considering in some detail what is the information content in a random sample, and furthermore: what steps can we take to extract that information in an optimal way?
Example. Let us say that we want to estimate the expectation \(\mu\) in a \(N(\mu, \sigma^2)\) distribution based on an iid sample \(x = (x_1, \ldots, x_n)\). We have used several theoretical arguments to say that \(\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i\) is a sensible estimator for \(\mu\):
- it is unbiased: \(E(\bar{X}) = \mu\),
- it is consistent: \(\bar{X} \overset{a.s.}{\to} \mu\),
- it is both the method of moments and the maximum likelihood estimator for \(\mu\).
But note that the sample only enters the mean through the sum of the observations. If we had another sample \(y = (y_1, \ldots, y_n)\) for which \(\sum_{i=1}^{n} y_i = \sum_{i=1}^{n} x_i\), then both samples would lead to the same conclusions about \(\mu\). So in some sense, the two samples have the same information content regarding the unknown parameter \(\mu\).
We will formalize this idea in the following. We will mostly use material from Chapter 6 in Casella & Berger.
- \(X\) denotes the random variables \(X_1, \ldots, X_n\); \(x\) denotes the sample \(x_1, \ldots, x_n\).
- \(T(\cdot)\) denotes a statistic; \(T(X)\) is a random variable, \(T(x)\) a realized value.
Any statistic \(T(\cdot)\) defines a form of data reduction or data summary. Following the example above, \(T(x) = x_1 + \cdots + x_n\) does not report all of the values in the sample, but only a single number: the sum. We will discuss two ways to analyze the information content in \(x\) and \(T(x)\) with respect to an unknown parameter. The book presents a third as well.
The sufficiency principle
A sufficient statistic for a parameter \(\theta\) is a statistic that, in a certain sense, captures all the information about \(\theta\) contained in the sample. This leads to
The sufficiency principle. If \(T(X)\) is a sufficient statistic for \(\theta\), then any inference about \(\theta\) should depend on the sample only through the value \(T(X)\). That is, if \(x\) and \(y\) are two samples such that \(T(x) = T(y)\), then the inference about \(\theta\) should be the same whether \(X = x\) or \(X = y\) is observed.
We have the following formal definition of a sufficient statistic:
Definition. A statistic \(T(X)\) is a sufficient statistic for \(\theta\) if the conditional distribution of the sample \(X\) given \(T(X)\) does not depend on \(\theta\).
CB explains this definition in the following way (p. 222):
The definition above concerns the conditional probability \(P\big(X = x \mid T(X) = T(x)\big)\), and if \(T\) is sufficient, then this quantity does not depend on \(\theta\).
Assume:
- Experimenter 1 knows the full sample \(X = x\), and thus also the value \(T(x)\).
- Experimenter 2 only knows the realized value of the sufficient statistic \(T(x)\).
But under sufficiency, the distribution of the sample is now fully known: it does not depend on the unknown parameter \(\theta\).
\(\Rightarrow\) Experimenter 2 can generate a new sample \(Y\) from this distribution.
Whether it was Experimenter 1’s \(X\) or Experimenter 2’s \(Y\) that led us to observe \(T(x)\) does not matter, because they go on to show that \[ P_\theta(X = x) = P_\theta(Y = x), \] meaning that both experimenters have the same knowledge of the unknown parameter (their likelihood functions will be the same!).
\(\longrightarrow\) It is enough (sufficient) to observe the value of a sufficient statistic!
Theorem 6.2.2. If \(p(x \mid \theta)\) is the joint pdf or pmf of \(X\), and \(q(t \mid \theta)\) is the pdf or pmf of \(T(X)\), then \(T(X)\) is a sufficient statistic for \(\theta\) if, for every \(x\) in the sample space, the ratio \(p(x \mid \theta) / q\big(T(x) \mid \theta\big)\) is constant as a function of \(\theta\).
Proof. We start from the other direction: \[ \begin{aligned} P_\theta\big(X = x \mid T(X) = T(x)\big) &= \frac{P_\theta\big(X = x \text{ and } T(X) = T(x)\big)}{P_\theta\big(T(X) = T(x)\big)} \overset{(\dagger)}{=} \frac{P_\theta(X = x)}{P_\theta\big(T(X) = T(x)\big)} \\ &= \frac{p(x \mid \theta)}{q\big(T(x) \mid \theta\big)}, \end{aligned} \] where \((\dagger)\) holds because \(\{X = x\} \subset \{T(X) = T(x)\}\), and under sufficiency the left hand side does not depend on \(\theta\). \(\qquad \blacksquare\)
Example 6.2.4 (normal example, formalized). We continue the example from above, but formalize it according to what we have established so far. Let \(X_1, \ldots, X_n\) be iid \(N(\mu, \sigma^2)\) with \(\sigma^2\) known. We wish to show that the sample mean, \(T(X) = \bar{X} = (X_1 + \cdots + X_n)/n\), is a sufficient statistic for \(\mu\). The joint pdf of the sample \(X\) is \[ \begin{aligned} f(x \mid \mu) &= \prod_{i=1}^{n} (2\pi\sigma^2)^{-1/2} \exp\left\{-\frac{(x_i - \mu)^2}{2\sigma^2}\right\} = (2\pi\sigma^2)^{-n/2} \exp\left\{-\sum_{i=1}^{n}\frac{(x_i - \mu)^2}{2\sigma^2}\right\} \\ &\overset{(a)}{=} (2\pi\sigma^2)^{-n/2} \exp\left\{-\sum_{i=1}^{n}\frac{(x_i - \bar{x} + \bar{x} - \mu)^2}{2\sigma^2}\right\} \\ &\overset{(b)}{=} (2\pi\sigma^2)^{-n/2} \exp\left\{-\left(\sum_{i=1}^{n}(x_i - \bar{x})^2 + n(\bar{x} - \mu)^2\right)\Big/(2\sigma^2)\right\}, \end{aligned} \] where \((a)\) adds and subtracts \(\bar{x}\), and \((b)\) uses \(\sum_{i=1}^{n}(x_i - \bar{x}) = 0\).
We have previously shown that \(\bar{x} \sim N(\mu, \sigma^2/n)\). The ratio is therefore \[ \begin{aligned} \frac{f(x \mid \theta)}{q\big(T(x) \mid \theta\big)} &= \frac{(2\pi\sigma^2)^{-n/2} \exp\left\{-\left(\sum_{i=1}^{n}(x_i - \bar{x})^2 + n(\bar{x} - \mu)^2\right)/(2\sigma^2)\right\}} {(2\pi\sigma^2/n)^{-1/2}\exp\left\{-n(\bar{x} - \mu)^2/(2\sigma^2)\right\}} \\ &= n^{-1/2}(2\pi\sigma^2)^{-(n-1)/2} \exp\left\{-\sum_{i=1}^{n}\frac{(x_i - \bar{x})^2}{2\sigma^2}\right\}, \end{aligned} \] which does not depend on \(\mu\), so by Theorem 6.2.2, \(T(X) = \bar{X}\) is a sufficient statistic for \(\mu\).
A small detail. Note that in the motivating example, we argued that the sum \(s = \sum_{i=1}^{n} x_i\) is somehow “sufficient” for estimating \(\mu\) in a normal distribution. We then went on to define formally what we mean by “sufficiency” and then showed that the mean \(\bar{x} = s/n\) is sufficient, formally, to estimate \(\mu\).
So which one is it? The sum or the mean? One might think that just knowing the sum is not enough — in order to estimate \(\mu\) we also need to know the value of \(n\), the sample size, and hence that \(s\) is not sufficient.
But this is wrong. The sample size is normally not considered as part of the data, but a part of the experimental setup that is known, in principle, before we do any observations. Hence, whether we multiply \(s\) with the known constant \(1/n\) does not change anything, and is just a matter of convenience.
The following theorem allows us to find a sufficient statistic by simply inspecting the pdf or pmf of the sample:
Theorem 6.2.6. Let \(f(x \mid \theta)\) denote the joint pdf or pmf of a sample \(X\). A statistic \(T\) is a sufficient statistic for \(\theta\) if and only if there exist functions \(g(t \mid \theta)\) and \(h(x)\) such that for all sample points \(x\) and all parameter points \(\theta\), \[ f(x \mid \theta) = g\big(T(x) \mid \theta\big)\, h(x). \]
Casella & Berger prove this result formally in the discrete case, but it makes very good sense given what we know about maximum likelihood estimation. The log-likelihood of the sample is \[ \ell(\theta \mid x) = \log f(x \mid \theta) = \log g\big(T(x) \mid \theta\big) + \log h(x), \] so when differentiating with respect to \(\theta\), we see that we can express the MLE solely in terms of \(T\): \[ \frac{\partial \ell(\theta \mid x)}{\partial \theta} = \frac{\partial}{\partial \theta} \log g\big(T(x) \mid \theta\big). \]
Example 6.2.7 (further continuation of the normal example). We saw earlier that the pdf of the normal sample could be factored as \[ f(x \mid \mu) = \underbrace{(2\pi\sigma^2)^{-n/2} \exp\left\{-\sum_{i=1}^{n}(x_i - \bar{x})^2/(2\sigma^2)\right\}}_{h(x)}\; \underbrace{\exp\left\{-n(\bar{x} - \mu)^2/(2\sigma^2)\right\}}_{g(T(x)\,\mid\,\mu)}, \] because we treat \(\sigma^2\) as a known constant here. From the factorization theorem, we see immediately that \(T(X)\) is sufficient for \(\mu\).
Higher dimensions
Sometimes we are not able to summarise the information in the data into just a single number, so that the sufficient statistic takes the form of a vector: \(T(X) = (T_1(X), \ldots, T_r(X))\). This usually happens when we have several parameters \(\theta = (\theta_1, \ldots, \theta_s)\), and the typical case is that \(r = s\). The factorization theorem works in the same way:
Example 6.2.9 (normal example finally concluded). Now both \(\mu\) and \(\sigma^2\) are unknown and must be estimated from data. When using the factorization theorem, any part of the joint pdf that depends on \(\mu\) or \(\sigma^2\) must be included in the \(g\) function. Working with the expression above, we see that the pdf depends on the sample \(x\) only through the two values \(T_1(x) = \bar{x}\) and \(T_2(x) = s^2 = \sum_{i=1}^{n}(x_i - \bar{x})^2/(n-1)\). We can therefore define \(h(x) = 1\) and easily see that \[ f(x; \mu, \sigma^2) = g\big(T_1(x), T_2(x);\, \mu, \sigma^2\big)\, h(x), \] and thus that \(T(X) = (T_1(X), T_2(X)) = (\bar{X}, S^2)\) is a sufficient statistic for \((\mu, \sigma^2)\) in the normal model.
Note the following result, which gives sufficient statistics in the exponential family:
Theorem 6.2.10. Let \(X_1, \ldots, X_n\) be iid observations from a pdf or pmf \(f(x; \theta)\) which belongs to an exponential family given by \[ f(x; \theta) = h(x) c(\theta) \exp\left(\sum_{i=1}^{k} w_i(\theta) t_i(x)\right), \] where \(\theta = (\theta_1, \ldots, \theta_d)\), \(d \leq k\). Then \[ T(X) = \left(\sum_{i=1}^{n} t_1(X_i), \ldots, \sum_{i=1}^{n} t_k(X_i)\right) \] is a sufficient statistic for \(\theta\).
The likelihood principle
Recall from the previous lecture that the likelihood function \(L(\theta \mid x) = f(x \mid \theta)\), viewed as a function of \(\theta\) for the fixed, observed \(x\), is the central object of likelihood-based inference. The likelihood principle says, roughly, that this function contains everything the data has to say about \(\theta\): nothing else about how the data arose should matter.
The likelihood principle. If \(x\) and \(y\) are two sample points such that \(L(\theta \mid x)\) is proportional to \(L(\theta \mid y)\) — that is, there exists a constant \(C(x,y)\) such that \[ L(\theta \mid x) = C(x,y)\, L(\theta \mid y) \quad \text{for all } \theta, \] then the conclusions drawn from \(x\) and \(y\) should be identical.
Note how close this is to the sufficiency principle that we have just established. In fact, the factorization theorem shows that sufficiency is a consequence of the likelihood principle. If \(T(x)\) is sufficient, then \(f(x \mid \theta) = g(T(x) \mid \theta) h(x)\), so \[ L(\theta \mid x) = h(x)\, g\big(T(x) \mid \theta\big). \]
If \(T(x) = T(y)\), then \[ \frac{L(\theta \mid x)}{L(\theta \mid y)} = \frac{h(x)\, g\big(T(x) \mid \theta\big)}{h(y)\, g\big(T(y) \mid \theta\big)} = \frac{h(x)}{h(y)}, \] which does not depend on \(\theta\). So two samples with the same value of a sufficient statistic automatically have proportional likelihoods.
(A version of) Example 6.3.5. Suppose we want to learn about \(\theta\), the probability that a coin shows heads.
Experiment 1: flip the coin \(n\) times and count \(x\) heads. Then \(X \sim \text{Binomial}(n, \theta)\) and \[ L_1(\theta \mid x) = \binom{n}{x}\theta^x (1-\theta)^{n-x}. \]
Experiment 2: keep flipping until the \(x\)’th head appears, and it happens to take \(n\) flips. Then \(N \sim \text{Negative Binomial}(x, \theta)\) and \[ L_2(\theta \mid n) = \binom{n-1}{x-1}\theta^x (1-\theta)^{n-x}. \]
The two likelihoods are proportional, and the likelihood principle says that our conclusions about \(\theta\) should be identical in the two experiments. Yet classical procedures such as \(p\)-values or confidence intervals can give different answers in the two cases, because they average over “what could have happened” under hypothetical repetitions, which would be different for the two sampling schemes (one fixes \(n\), the other fixes \(x\)). The likelihood principle says that these differences are irrelevant: once \(x\) and \(n\) are observed, the two experiments carry exactly the same information about \(\theta\).
A small remark on equivariance
C&B present a third principle, the equivariance principle, which says that if two problems have identical formal structure (the same model up to a relabelling of \(\theta\)) then they should be analyzed by the same procedure, applied to the relabelled quantities.
We have already seen the most important instance of this in Theorem 7.2.10, which is the invariance of the MLE under reparameterization. This is exactly the invariance principle applied to maximum likelihood estimation.
We will not develop this idea any further here.