7. Principles of Data Reduction

A theoretical view of inference

In this lecture we return to a theoretical view of statistical inference. We will look in some detail at two questions:

  • What is the information content in a random sample?
  • What steps can we take to extract that information in an optimal way?

We will mostly use material from Chapter 6 in Casella & Berger. The central idea is that of data reduction: a statistic summarises a whole sample into something simpler, and we want to understand when such a summary throws away nothing that matters about the parameter.

A motivating example

Suppose we want to estimate the expectation \(\mu\) in a \(N(\mu, \sigma^2)\) distribution from an iid sample \(x = (x_1, \ldots, x_n)\). Several arguments tell us that \(\bar{x} = \tfrac{1}{n}\sum_{i=1}^n x_i\) is a sensible estimator of \(\mu\):

  • it is unbiased: \(E(\bar{X}) = \mu\),
  • it is consistent: \(\bar{X} \xrightarrow{\text{a.s.}} \mu\),
  • it is both the method of moments and the maximum likelihood estimator for \(\mu\).

But note that the sample only enters the mean through the sum of the observations. If we had another sample \(y = (y_1, \ldots, y_n)\) with \(\sum_{i=1}^n y_i = \sum_{i=1}^n x_i\), then both samples would lead to the same conclusions about \(\mu\).

\[\longrightarrow \text{In some sense, the two samples carry the \textbf{same information} about } \mu.\]

Notation and statistics

Before formalising this idea, let us fix some notation.

  • \(X\) denotes the random variables \(X_1, \ldots, X_n\); \(\;x\) denotes the observed sample \(x_1, \ldots, x_n\).
  • \(T(\cdot)\) denotes a statistic. Then \(T(X)\) is a random variable and \(T(x)\) is a realised value.

Any statistic \(T(\cdot)\) defines a form of data reduction or data summary. Following the example above, \(T(x) = x_1 + \cdots + x_n\) does not report all of the values in the sample — only a single number, the sum.

We will discuss two ways to analyse the information content in \(x\) and \(T(x)\) with respect to an unknown parameter. The book presents a third as well.

Sufficient statistics

A sufficient statistic for a parameter \(\theta\) is a statistic that, in a certain sense, captures all the information about \(\theta\) contained in the sample. This leads to the following principle.

The sufficiency principle. If \(T(X)\) is a sufficient statistic for \(\theta\), then any inference about \(\theta\) should depend on the sample only through the value \(T(X)\). That is, if \(x\) and \(y\) are two samples with \(T(x) = T(y)\), then the inference about \(\theta\) should be the same whether \(X = x\) or \(X = y\) is observed.

We now make the notion of “capturing all the information” precise.

Defining sufficiency

Definition. A statistic \(T(X)\) is a sufficient statistic for \(\theta\) if the conditional distribution of the sample \(X\) given the value \(T(X)\) does not depend on \(\theta\).

The definition concerns the conditional probability \(P_\theta\bigl(X = x \mid T(X) = T(x)\bigr)\). If \(T\) is sufficient, this quantity does not depend on \(\theta\).

Intuitively: once we know \(T(x)\), the parameter \(\theta\) tells us nothing more about how the remaining randomness in the sample was distributed. Everything that \(\theta\) has to say about the data is already contained in \(T(x)\).

Why “sufficient”? Two experimenters

Casella & Berger (p. 222) explain the definition through a thought experiment.

  • Experimenter 1 knows the full sample \(X = x\), and hence also the value \(T(x)\).
  • Experimenter 2 knows only the realised value of the sufficient statistic \(T(x)\).

Under sufficiency, the conditional distribution of the sample is fully known — it does not depend on the unknown \(\theta\). So Experimenter 2 can generate a new sample \(Y\) from this distribution.

Whether it was Experimenter 1’s \(X\) or Experimenter 2’s \(Y\) that led to \(T(x)\) does not matter, because one can show that \[P_\theta(X = x) = P_\theta(Y = x),\] meaning both experimenters have the same knowledge about \(\theta\) — their likelihood functions are identical.

\[\longrightarrow \text{It is } \underline{\text{enough}} \text{ (\textit{sufficient}) to observe the value of a sufficient statistic.}\]

A first characterisation

Checking the definition directly is awkward. The following theorem gives a workable criterion.

Theorem 6.2.2. Let \(p(x \mid \theta)\) be the joint pdf or pmf of \(X\), and let \(q(t \mid \theta)\) be the pdf or pmf of \(T(X)\). Then \(T(X)\) is a sufficient statistic for \(\theta\) if, for every \(x\) in the sample space, the ratio \[\frac{p(x \mid \theta)}{q\bigl(T(x) \mid \theta\bigr)}\] is constant as a function of \(\theta\).

Proof idea. Start from the conditional probability and use that \(\{X = x\} \subseteq \{T(X) = T(x)\}\):

\[\begin{align*} P_\theta\bigl(X = x \mid T(X) = T(x)\bigr) &= \frac{P_\theta\bigl(X = x \text{ and } T(X) = T(x)\bigr)}{P_\theta\bigl(T(X) = T(x)\bigr)} = \frac{P_\theta(X = x)}{P_\theta\bigl(T(X) = T(x)\bigr)} \\ &= \frac{p(x \mid \theta)}{q\bigl(T(x) \mid \theta\bigr)}. \end{align*}\] Under sufficiency the left-hand side does not depend on \(\theta\), so neither does the ratio on the right.

Example: the normal mean, formalised

Example 6.2.4. Let \(X_1, \ldots, X_n\) be iid \(N(\mu, \sigma^2)\) with \(\sigma^2\) known. We show that \(T(X) = \bar{X} = (X_1 + \cdots + X_n)/n\) is sufficient for \(\mu\). The joint pdf is \[f(x \mid \mu) = \prod_{i=1}^n (2\pi\sigma^2)^{-1/2}\exp\!\left\{-\frac{(x_i - \mu)^2}{2\sigma^2}\right\} = (2\pi\sigma^2)^{-n/2}\exp\!\left\{-\sum_{i=1}^n \frac{(x_i - \mu)^2}{2\sigma^2}\right\}.\]

Adding and subtracting \(\bar{x}\) inside the square, and using \(\sum_{i=1}^n (x_i - \bar{x}) = 0\): \[\begin{align*} f(x \mid \mu) &= (2\pi\sigma^2)^{-n/2}\exp\!\left\{-\sum_{i=1}^n \frac{(x_i - \bar{x} + \bar{x} - \mu)^2}{2\sigma^2}\right\} \\&= (2\pi\sigma^2)^{-n/2}\exp\!\left\{-\frac{\sum_{i=1}^n (x_i - \bar{x})^2 + n(\bar{x} - \mu)^2}{2\sigma^2}\right\}. \end{align*}\]

Example: the normal mean, formalised (cont.)

We have previously shown that \(\bar{X} \sim N(\mu, \sigma^2/n)\). The ratio of Theorem 6.2.2 is therefore \[\frac{f(x \mid \mu)}{q\bigl(T(x) \mid \mu\bigr)} = \frac{(2\pi\sigma^2)^{-n/2}\exp\!\left\{-\bigl(\sum_{i=1}^n (x_i - \bar{x})^2 + n(\bar{x} - \mu)^2\bigr)/(2\sigma^2)\right\}}{(2\pi\sigma^2/n)^{-1/2}\exp\!\left\{-n(\bar{x} - \mu)^2/(2\sigma^2)\right\}}.\]

The terms involving \(\mu\) cancel, leaving \[\frac{f(x \mid \mu)}{q\bigl(T(x) \mid \mu\bigr)} = n^{-1/2}(2\pi\sigma^2)^{-(n-1)/2}\exp\!\left\{-\sum_{i=1}^n \frac{(x_i - \bar{x})^2}{2\sigma^2}\right\},\]

which does not depend on \(\mu\). By Theorem 6.2.2, \(T(X) = \bar{X}\) is a sufficient statistic for \(\mu\).

A subtle point: sum or mean?

In the motivating example we argued that the sum \(s = \sum_{i=1}^n x_i\) was somehow “sufficient” for estimating \(\mu\). We then defined sufficiency formally and showed that the mean \(\bar{x} = s/n\) is sufficient. So which is it — the sum or the mean?

One might think that just knowing the sum is not enough: to estimate \(\mu\) we also need the sample size \(n\), so \(s\) alone is not sufficient.

But this is wrong. The sample size \(n\) is normally not considered part of the data — it is part of the experimental setup, known in principle before any observations are made. Whether we multiply \(s\) by the known constant \(1/n\) does not change anything, and is just a matter of convenience.

The factorization theorem

The next theorem lets us find a sufficient statistic by simply inspecting the pdf or pmf.

Theorem 6.2.6 (Factorization theorem). Let \(f(x \mid \theta)\) denote the joint pdf or pmf of a sample \(X\). A statistic \(T(X)\) is a sufficient statistic for \(\theta\) if and only if there exist functions \(g(t \mid \theta)\) and \(h(x)\) such that, for all sample points \(x\) and all parameter points \(\theta\), \[f(x \mid \theta) = g\bigl(T(x) \mid \theta\bigr)\, h(x).\]

The parameter \(\theta\) interacts with the data only through \(T(x)\), in the factor \(g\). The factor \(h(x)\) carries the rest of the dependence on \(x\), free of \(\theta\).

Why the factorization theorem makes sense

Casella & Berger prove this formally in the discrete case, but it makes very good sense given what we know about maximum likelihood. The log-likelihood of the sample is \[\ell(\theta \mid x) = \log f(x \mid \theta) = \log g\bigl(T(x) \mid \theta\bigr) + \log h(x).\]

When differentiating with respect to \(\theta\), the term \(\log h(x)\) vanishes, so the MLE can be expressed solely in terms of \(T\): \[\frac{\partial \ell(\theta \mid x)}{\partial \theta} = \frac{\partial}{\partial \theta}\log g\bigl(T(x) \mid \theta\bigr).\]

The entire likelihood analysis — and hence the MLE — depends on the data only through \(T(x)\).

Example: the normal mean via factorization

Example 6.2.7. From the normal example above, the pdf factors as \[f(x \mid \mu) = \underbrace{(2\pi\sigma^2)^{-n/2}\exp\!\left\{-\sum_{i=1}^n (x_i - \bar{x})^2/(2\sigma^2)\right\}}_{h(x)}\;\underbrace{\exp\!\left\{-n(\bar{x} - \mu)^2/(2\sigma^2)\right\}}_{g(T(x)\,\mid\,\mu)}.\]

Here \(\sigma^2\) is treated as a known constant, so it may go into \(h(x)\). The second factor depends on \(x\) only through \(\bar{x} = T(x)\). By the factorization theorem, \(T(X) = \bar{X}\) is immediately seen to be sufficient for \(\mu\) — no ratio computation needed.

Higher dimensions

Sometimes we cannot summarise the information in the data into just a single number, so the sufficient statistic takes the form of a vector \[T(X) = \bigl(T_1(X), \ldots, T_r(X)\bigr).\]

This usually happens when we have several parameters \(\theta = (\theta_1, \ldots, \theta_s)\). The typical case is \(r = s\), and the factorization theorem works in exactly the same way.

Example: the normal, both parameters unknown

Example 6.2.9. Now both \(\mu\) and \(\sigma^2\) are unknown and must be estimated from data. When using the factorization theorem, any part of the joint pdf that depends on \(\mu\) or \(\sigma^2\) must be included in the \(g\) factor.

Working with the expression from before, the pdf depends on the sample \(x\) only through the two values \[T_1(x) = \bar{x} \qquad \text{and} \qquad T_2(x) = s^2 = \frac{1}{n-1}\sum_{i=1}^n (x_i - \bar{x})^2.\]

We can therefore set \(h(x) = 1\) and write \(f(x \mid \mu, \sigma^2) = g\bigl(T_1(x), T_2(x) \mid \mu, \sigma^2\bigr)\,h(x)\). Hence \[T(X) = \bigl(\bar{X}, S^2\bigr)\] is a sufficient statistic for \((\mu, \sigma^2)\) in the normal model.

Sufficiency in the exponential family

A particularly clean result holds for the exponential family.

Theorem 6.2.10. Let \(X_1, \ldots, X_n\) be iid observations from a pdf or pmf \(f(x \mid \theta)\) belonging to an exponential family \[f(x \mid \theta) = h(x)\, c(\theta)\, \exp\!\left(\sum_{j=1}^k w_j(\theta)\, t_j(x)\right),\] where \(\theta = (\theta_1, \ldots, \theta_d)\) with \(d \leq k\). Then \[T(X) = \left(\sum_{i=1}^n t_1(X_i), \;\ldots,\; \sum_{i=1}^n t_k(X_i)\right)\] is a sufficient statistic for \(\theta\).

Most of our standard distributions are exponential families, so this single result delivers their sufficient statistics directly.

The likelihood principle

Recall from the previous lecture that the likelihood function \(L(\theta \mid x) = f(x \mid \theta)\), viewed as a function of \(\theta\) for the fixed, observed \(x\), is the central object of likelihood-based inference.

The likelihood principle says, roughly, that this function contains everything the data have to say about \(\theta\): nothing else about how the data arose should matter.

The likelihood principle. If \(x\) and \(y\) are two sample points such that \(L(\theta \mid x)\) is proportional to \(L(\theta \mid y)\) — that is, there exists a constant \(C(x, y)\) with \[L(\theta \mid x) = C(x, y)\, L(\theta \mid y) \quad \text{for all } \theta,\] then the conclusions drawn from \(x\) and \(y\) should be identical.

Sufficiency is a consequence

Note how close this is to the sufficiency principle. In fact, the factorization theorem shows that sufficiency is a consequence of the likelihood principle.

If \(T(X)\) is sufficient, then \(f(x \mid \theta) = g\bigl(T(x) \mid \theta\bigr)h(x)\), so \[L(\theta \mid x) = h(x)\, g\bigl(T(x) \mid \theta\bigr).\]

If \(T(x) = T(y)\), then \[\frac{L(\theta \mid x)}{L(\theta \mid y)} = \frac{h(x)\, g\bigl(T(x) \mid \theta\bigr)}{h(y)\, g\bigl(T(y) \mid \theta\bigr)} = \frac{h(x)}{h(y)},\]

which does not depend on \(\theta\). So two samples with the same value of a sufficient statistic automatically have proportional likelihoods.

Example: two ways to flip a coin

(A version of) Example 6.3.5. Suppose we want to learn about \(\theta\), the probability that a coin shows heads.

  • Experiment 1: flip the coin \(n\) times and count \(x\) heads. Then \(X \sim \text{Binomial}(n, \theta)\) and \[L_1(\theta \mid x) = \binom{n}{x}\theta^x (1 - \theta)^{n - x}.\]

  • Experiment 2: keep flipping until the \(x\)th head appears, and it happens to take \(n\) flips. Then \(N \sim \text{Negative Binomial}(x, \theta)\) and \[L_2(\theta \mid n) = \binom{n - 1}{x - 1}\theta^x (1 - \theta)^{n - x}.\]

The two likelihoods are proportional (both \(\propto \theta^x (1-\theta)^{n-x}\)), so the likelihood principle says our conclusions about \(\theta\) should be identical in the two experiments.

Example: two ways to flip a coin (cont.)

Yet classical procedures such as \(p\)-values or confidence intervals can give different answers in the two cases. The reason: they average over “what could have happened” under hypothetical repetitions, and these differ for the two sampling schemes — one fixes \(n\), the other fixes \(x\).

The likelihood principle says these differences are irrelevant: once \(x\) and \(n\) are observed, the two experiments carry exactly the same information about \(\theta\).

This is a proper controversial point in statistics — likelihood-based and Bayesian methods respect the likelihood principle, while many classical frequentist procedures do not.

The equivariance principle

Casella & Berger present a third principle, the equivariance principle, which says that if two problems have identical formal structure — the same model up to a relabelling of \(\theta\) — then they should be analysed by the same procedure, applied to the relabelled quantities.

We have already seen the most important instance of this in Theorem 7.2.10, the invariance of the MLE under reparameterization: if \(\hat\theta\) is the MLE of \(\theta\), then \(\tau(\hat\theta)\) is the MLE of \(\tau(\theta)\). This is exactly the equivariance principle applied to maximum likelihood estimation.

We will not develop this idea any further here.

Summary

  • A statistic \(T(X)\) reduces the data; the question is whether the reduction loses information about \(\theta\).
  • \(T(X)\) is sufficient if the conditional distribution of \(X\) given \(T(X)\) does not depend on \(\theta\) — then inference should depend on the data only through \(T(X)\).
  • The factorization theorem \(f(x \mid \theta) = g(T(x) \mid \theta)\,h(x)\) is the practical tool for finding sufficient statistics; exponential families deliver them automatically.
  • The likelihood principle says proportional likelihoods carry the same information; sufficiency follows from it.
  • The equivariance principle is reflected in the invariance of the MLE under reparameterization.