Observed data are treated as realizations of random variables.
Definition 5.1.1 (iid sample). The random variables \(X_1, \ldots, X_n\) are a random sample of size \(n\) from the population \(f(x)\) if they are mutually independent and each has the same marginal pdf or pmf \(f(x)\). The common term is iid: independent and identically distributed.
By independence, the joint pdf or pmf is the product of the marginals. For a parametric family \(f(x\mid\theta)\), \[f(x_1, \ldots, x_n \mid \theta) = \prod_{i=1}^n f(x_i \mid \theta).\]
Let \(X_1, \ldots, X_n\) be a random sample from an exponential population with mean \(\beta\) — for example, the time until failure for \(n\) identical units. The joint pdf is \[f(x_1, \ldots, x_n \mid \beta) = \prod_{i=1}^n \frac{1}{\beta}e^{-x_i/\beta} = \frac{1}{\beta^n} e^{-(x_1 + \cdots + x_n)/\beta}.\]
The probability that all of the units last more than 2 time units: \[\begin{align} P(X_1 > 2, \ldots, X_n > 2) &= \int_2^\infty \cdots \int_2^\infty \prod_{i=1}^n \frac{1}{\beta}e^{-x_i/\beta}\,dx_1\cdots dx_n \\ &= \left(\int_2^\infty \frac{1}{\beta}e^{-x/\beta}\,dx\right)^n = \left(e^{-2/\beta}\right)^n = e^{-2n/\beta}. \end{align}\]
If \(\beta\) is large relative to \(n\), this probability is close to one.
Assume the population belongs to a parametric family, but the parameter is unknown.
Can the observed sample \(X_1, \ldots, X_n\) be used to find a plausible value of the parameter?
One answer: choose the value of the parameter that maximizes the probability of the observed sample. This is maximum likelihood estimation (Topic 6).
Statistic. Any function \(T = T(X_1, \ldots, X_n)\) of the sample is a statistic.
Unbiasedness. A statistic \(T\) is an unbiased estimator of \(\theta\) if \(E(T) = \theta\).
Example. For a random sample with \(E(X) = \mu\), the sample mean is unbiased: \[E(\bar{X}) = E\!\left(\frac{1}{n}\sum_{i=1}^n X_i\right) = \frac{1}{n}\sum_{i=1}^n E(X_i) = \mu.\]
The sample average corresponds to the population expectation: an observable property of the sample, used to learn about an unobservable property of the population. Working out such correspondences precisely is statistical inference.
Definition 5.2.3. The sample variance is the statistic \[S^2 = \frac{1}{n-1}\sum_{i=1}^n (X_i - \bar{X})^2.\]
Definition 5.2.1 (second half). The probability distribution of a statistic \(T\) is its sampling distribution.
Theorem 5.2.4. Let \(x_1, \ldots, x_n\) be any numbers, and \(\bar{x} = (x_1 + \cdots + x_n)/n\).
\(\;\displaystyle\min_a \sum_{i=1}^n (x_i - a)^2 = \sum_{i=1}^n (x_i - \bar{x})^2\).
\(\;(n-1)s^2 = \displaystyle\sum_{i=1}^n (x_i - \bar{x})^2 = \sum_{i=1}^n x_i^2 - n\bar{x}^2\).
Proof of (a). Write \(x_i - a = (x_i - \bar{x}) + (\bar{x} - a)\) and expand: \[\begin{align} \sum_{i=1}^n(x_i - a)^2 &= \sum_{i=1}^n(x_i-\bar{x})^2 + 2(\bar{x}-a)\underbrace{\sum_{i=1}^n(x_i-\bar{x})}_{=\,0} + n(\bar{x}-a)^2 \\ &= \sum_{i=1}^n(x_i-\bar{x})^2 + n(\bar{x}-a)^2, \end{align}\] which is minimized when \(a = \bar{x}\). Part (b) follows in the same manner, since the cross-term is again zero. \(\blacksquare\)
Lemma 5.2.5. Let \(X_1, \ldots, X_n\) be a random sample and let \(g(x)\) be a function such that \(Eg(X_i)\) and \(\text{Var}(g(X_i))\) exist. Then \[E\!\left(\sum_{i=1}^n g(X_i)\right) = n\,Eg(X_1) \qquad \text{and} \qquad \text{Var}\!\left(\sum_{i=1}^n g(X_i)\right) = n\,\text{Var}\,g(X_1).\]
The expectation result does not require independence; the variance result does.
Theorem 5.2.6. Let \(X_1, \ldots, X_n\) be a random sample from a population with mean \(\mu\) and variance \(\sigma^2 < \infty\). Then
Parts (a) and (b). (a) was shown earlier. For (b), taking the sum out of the variance uses independence: \[\text{Var}(\bar{X}) = \text{Var}\!\left(\frac{1}{n}\sum_{i=1}^n X_i\right) = \frac{1}{n^2}\sum_{i=1}^n \text{Var}(X_i) = \frac{1}{n^2}\cdot n\sigma^2 = \frac{\sigma^2}{n}.\]
Part (c). By Theorem 5.2.4(b), with \(EX^2 = \sigma^2 + \mu^2\) and \(E\bar{X}^2 = \text{Var}(\bar{X}) + (E\bar{X})^2 = \sigma^2/n + \mu^2\): \[\begin{align} E(S^2) &= \frac{1}{n-1}\bigl(n\,EX^2 - n\,E\bar{X}^2\bigr) = \frac{1}{n-1}\left(n(\sigma^2 + \mu^2) - n\left(\frac{\sigma^2}{n} + \mu^2\right)\right) = \sigma^2. \qquad\blacksquare \end{align}\]
\(T_1 = \bar{X}\), \(T_2 = (X_1 + X_2)/2\) and \(T_3 = X_1\) are all unbiased for \(\mu\), with variances \(\sigma^2/n\), \(\sigma^2/2\) and \(\sigma^2\). Unbiasedness does not choose between them.
Mean squared error (CB §7.3.1). \(\;\text{MSE}(T) = E\bigl((T - \theta)^2\bigr) = \text{Var}(T) + \bigl(E(T) - \theta\bigr)^2\): variance plus squared bias.
For a normal sample, compare \(S^2\) with \(\hat\sigma^2 = \frac{n-1}{n}S^2\). Using \(\text{Var}(S^2) = 2\sigma^4/(n-1)\), which follows from Theorem 5.3.1(c), \[\text{MSE}(S^2) = \frac{2\sigma^4}{n-1}, \qquad \text{MSE}(\hat\sigma^2) = \frac{2n-1}{n^2}\,\sigma^4 < \frac{2\sigma^4}{n-1}.\]
\(\hat\sigma^2\) is biased, yet closer to \(\sigma^2\) on average for every \(n\). It is also the maximum likelihood estimator of \(\sigma^2\) (Topic 6).
Theorem 5.2.7. Let \(X_1, \ldots, X_n\) be a random sample from a population with mgf \(M_X(t)\). Then the mgf of the sample mean is \[M_{\bar{X}}(t) = \bigl(M_X(t/n)\bigr)^n.\]
Proof. \[\begin{align} M_{\bar{X}}(t) = E e^{t\bar{X}} &= E e^{t(X_1+\cdots+X_n)/n} = E\!\left[e^{(t/n)X_1}\,e^{(t/n)X_2}\cdots e^{(t/n)X_n}\right] \\ &= E e^{(t/n)X_1}\cdots E e^{(t/n)X_n} \qquad \text{(by independence)} \\ &= \bigl(M_X(t/n)\bigr)^n. \qquad\blacksquare \end{align}\]
Under normal sampling, the sampling distributions of the statistics used most in inference can be derived exactly. By the central limit theorem (Topic 5), they remain approximately correct far more generally.
Three tasks:
The arguments are constructive: not advanced, just long.
Lemma 5.3.2. We use the notation \(\chi^2_p\) for a chi-squared random variable with \(p\) degrees of freedom.
Proof of (a). From Examples 2.1.7 and 2.1.9 in CB, if \(Y = X^2\), then \[\begin{align*} F_Y(y) &= P(Y\leq y) = P(X^2 \leq y ) = P(-\sqrt{y} \leq X \leq \sqrt{y}) \\ &= F_X(\sqrt{y}) - F_X(-\sqrt{y})),\end{align*}\] so the density of \(Y\) is \(f_Y(y) = \frac{1}{2\sqrt{y}}(f_X(\sqrt{y}) + f_X(-\sqrt{y}))\). Applying this to \(Z \sim N(0,1)\): \[ f_Y(y) = \frac{1}{2\sqrt{y}}\cdot\frac{1}{\sqrt{2\pi}}\bigl(e^{-y/2} + e^{-y/2}\bigr) = \frac{1}{\sqrt{2\pi}}\frac{1}{\sqrt{y}}e^{-y/2}, \] which is the \(\chi^2_1\) pdf.
For (b) we can use that the MGF of the \(\chi^2_p\) is \(M_X(t) = (1-2t)^{-p/2}\), so the MGF of a sum of such independent variables is
\[\prod_{i=1}^n(1-2t^{-p_i/2}) = (1-2t)^{-\frac{1}{2}\sum_{i=1}^np_i} \sim \chi^2_{p1 + \cdots + p_n}.\]
Theorem 5.3.1. Let \(X_1, \ldots, X_n\) be a random sample from a \(N(\mu, \sigma^2)\) distribution. Then
The deviations from the sample mean always sum to zero: \[\sum_{i=1}^n (X_i - \bar{X}) = 0, \qquad\text{so}\qquad X_1 - \bar{X} = -\sum_{i=2}^n (X_i - \bar{X}).\]
Only \(n-1\) of the deviations are free. The same fact appears three times:
Take \(\mu = 0\) and \(\sigma^2 = 1\).
1. Reduce. We know that \(S^2 = \frac{1}{n-1}\sum_{i=1}^n (X_i - \bar{X})^2\). Is \(S^2\) a function of all \(n\) deviations \(X_i - \bar{X}\), \(i = 1, \ldots, n\)?
No: from the previous slide, the deviations sum to zero, so the first one is determined by the others, \[X_1 - \bar{X} = -\sum_{i=2}^n (X_i - \bar{X}). \qquad (*)\]
Substituting \((*)\) for the first term, \[S^2 = \frac{1}{n-1}\left(\Bigl[\sum_{i=2}^n (X_i - \bar{X})\Bigr]^2 + \sum_{i=2}^n (X_i - \bar{X})^2\right),\] so \(S^2\) is a function of the \(n-1\) variables \(X_2 - \bar{X}, \ldots, X_n - \bar{X}\) alone.
It is enough to show that \((X_2 - \bar{X}, \ldots, X_n - \bar{X})\) is independent of \(\bar{X}\), since a function of the first is then independent of \(\bar{X}\) (Theorem 4.6.12).
2. Transform. Put \[y_1 = \bar{x}, \qquad y_k = x_k - \bar{x}, \quad k = 2, \ldots, n.\]
The individual \(x\)’s are recovered from the \(y\)’s, using \((*)\) for the first one: \[\begin{align*} x_1 &= \bar{x} - \sum_{i=2}^n (x_i - \bar{x}) = y_1 - \sum_{i=2}^n y_i, \\ x_k &= \bar{x} + (x_k - \bar{x}) = y_1 + y_k, \qquad k = 2, \ldots, n. \end{align*}\]
The Jacobian of the inverse transformation is \[J = \begin{vmatrix} 1 & -1 & -1 & \cdots & -1 \\ 1 & 1 & 0 & \cdots & 0 \\ 1 & 0 & 1 & \cdots & 0 \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ 1 & 0 & 0 & \cdots & 1 \end{vmatrix} = n.\]
3. Substitute. The joint pdf is \(f(x_1, \ldots, x_n) = (2\pi)^{-n/2}\exp\bigl(-\frac{1}{2}\sum_i x_i^2\bigr)\), so we need the sum of squares in terms of the \(y\)’s: \[\begin{align*} \sum_{i=1}^n x_i^2 &= \Bigl(y_1 - \sum_{i=2}^n y_i\Bigr)^2 + \sum_{i=2}^n (y_1 + y_i)^2 \\ &= y_1^2 - 2y_1\!\sum_{i=2}^n y_i + \Bigl(\sum_{i=2}^n y_i\Bigr)^2 + (n-1)y_1^2 + 2y_1\!\sum_{i=2}^n y_i + \sum_{i=2}^n y_i^2 \\ &= n y_1^2 + \sum_{i=2}^n y_i^2 + \Bigl(\sum_{i=2}^n y_i\Bigr)^2, \qquad \text{(the cross-terms cancel).} \end{align*}\]
Multiplying by the Jacobian \(n = n^{1/2}\cdot n^{1/2}\) and splitting the exponential, \[f(y_1, \ldots, y_n) = \underbrace{\left[\left(\frac{n}{2\pi}\right)^{1/2}\!\exp\!\left(-\frac{n}{2}y_1^2\right)\right]}_{y_1 = \bar{x} \text{ only}}\underbrace{\left[\frac{n^{1/2}}{(2\pi)^{(n-1)/2}}\exp\!\left(-\frac{1}{2}\left(\sum_{i=2}^n y_i^2 + \Bigl(\sum_{i=2}^n y_i\Bigr)^{\!2}\right)\right)\right]}_{(y_2, \ldots, y_n) \text{ only}}.\]
4. Factor. The density factors into a function of \(y_1\) alone times a function of \((y_2, \ldots, y_n)\) alone, so by Theorem 4.6.11 these are independent, and therefore \(\bar{X}\) and \(S^2\) are independent. The first bracket is the \(N(0,1/n)\) density, which is part (b). \(\blacksquare\)
Take \(\mu = 0\) and \(\sigma^2 = 1\), and write \(\bar{X}_k\) and \(S_k^2\) for the mean and variance of the first \(k\) observations. The proof is by induction, using the recursion \[(n-1)S_n^2 = (n-2)S_{n-1}^2 + \frac{n-1}{n}\bigl(X_n - \bar{X}_{n-1}\bigr)^2. \qquad (*)\]
Base case, \(n = 2\). \(S_2^2 = \frac{1}{2}(X_2-X_1)^2\) and \((X_2-X_1)/\sqrt{2} \sim N(0,1)\), so \(S_2^2 \sim \chi^2_1\) by Lemma 5.3.2(a).
Inductive step. Assume \((k-1)S_k^2 \sim \chi^2_{k-1}\). By \((*)\), \[kS_{k+1}^2 = (k-1)S_k^2 + \frac{k}{k+1}\bigl(X_{k+1} - \bar{X}_k\bigr)^2.\]
By Lemma 5.3.2(b), \(kS_{k+1}^2 \sim \chi^2_k\). \(\blacksquare\)
Derivation of the recursion \((*)\): notes, Sampling from the normal distribution, and problem A2.
Theorem (Definition 5.3.4 in CB). Let \(X_1, \ldots, X_n\) be a random sample from \(N(\mu, \sigma^2)\). Then \(T = (\bar{X} - \mu)/(S/\sqrt{n})\) is \(t\)-distributed with \(p = n-1\) degrees of freedom, with density \[f_T(t) = \frac{\Gamma\bigl(\frac{p+1}{2}\bigr)}{\Gamma\bigl(\frac{p}{2}\bigr)\sqrt{p\pi}}\left(1 + \frac{t^2}{p}\right)^{-(p+1)/2}, \qquad -\infty < t < \infty.\]
Proof sketch. Write \[T = \frac{(\bar{X} - \mu)/(\sigma/\sqrt{n})}{\sqrt{S^2/\sigma^2}}.\] The numerator is \(N(0,1)\), the denominator is \(\sqrt{\chi^2_{n-1}/(n-1)}\), and the two are independent by Theorem 5.3.1(a). The density of such a ratio follows from a bivariate transformation: integrating out the \(\chi^2\) variable leaves a gamma kernel. \(\blacksquare\)
The transformation and the integral: notes, Sampling from the normal distribution.
Definition 5.3.6 (\(F\)-distribution). Let \(X_1, \ldots, X_n\) be a random sample from a \(N(\mu_X, \sigma_X^2)\) population, and let \(Y_1, \ldots, Y_m\) be a random sample from an independent \(N(\mu_Y, \sigma_Y^2)\) population. The random variable \[F = \frac{S_X^2/\sigma_X^2}{S_Y^2/\sigma_Y^2}\] is \(F\)-distributed with \((n-1, m-1)\) degrees of freedom. The pdf of an \(F_{p,q}\) variable is \[f_F(x) = \frac{\Gamma\!\left(\frac{p+q}{2}\right)}{\Gamma\!\left(\frac{p}{2}\right)\Gamma\!\left(\frac{q}{2}\right)}\left(\frac{p}{q}\right)^{p/2}\frac{x^{p/2-1}}{\left(1 + \frac{p}{q}x\right)^{(p+q)/2}}, \quad 0 < x < \infty.\]
The density \(f_F(x)\) can be derived in the same way as for the \(t\)-distribution.
These distributions are fundamental sampling distributions that provide quantiles for important statistics under common null hypotheses.
Interval for \(\mu\) with \(\sigma\) unknown. Since \(T = (\bar{X} - \mu)/(S/\sqrt{n}) \sim t_{n-1}\), \[P\bigl(-t_{n-1,\alpha/2} \leq T \leq t_{n-1,\alpha/2}\bigr) = 1 - \alpha \quad\Longrightarrow\quad \bar{X} \pm t_{n-1,\alpha/2}\,\frac{S}{\sqrt{n}}.\]
Interval for \(\sigma^2\). Since \((n-1)S^2/\sigma^2 \sim \chi^2_{n-1}\), \[\left[\frac{(n-1)S^2}{\chi^2_{n-1,\alpha/2}},\ \frac{(n-1)S^2}{\chi^2_{n-1,1-\alpha/2}}\right].\]
Test of \(\sigma_X^2 = \sigma_Y^2\). Under the hypothesis the \(F\) ratio reduces to \(S_X^2/S_Y^2 \sim F_{n-1,m-1}\), which can be computed from the data alone.
Every statistic so far has been an average, or built from averages. Some questions need something else.
The German tank problem. Captured tanks carry serial numbers. Treat the observed numbers 19, 40, 42 and 60 as a sample without replacement from \(\{1, \ldots, N\}\), and estimate \(N\).
Definition 5.4.1. The order statistics of a random sample \(X_1, \ldots, X_n\) are the sample values in ascending order, \(X_{(1)} \leq X_{(2)} \leq \cdots \leq X_{(n)}\). In particular \(X_{(1)} = \min_i X_i\) and \(X_{(n)} = \max_i X_i\).
The distribution of the maximum of \(X_1,\ldots,X_n \stackrel{\textrm{iid}}{\sim}F_X: X_{(n)} = \max_n X_i\): \[\begin{align} F_{X_{(n)}}(x) &= P(X_{(n)} \leq x) = P(\textrm{All } X_i \leq x) = (F(x))^n \\ & \Rightarrow f_{X_{(n)}}(x) = \frac{\textrm{d}}{\textrm{d}x}F_{X_{(n)}}(x) = nF^{n-1}(x)f(x). \end{align}\]
The general result for the \(j\)-th order statistic is Theorem 5.4.4 in CB, and problem A3.
For sampling without replacement from \(\{1, \ldots, N\}\), the maximum has \[P(X_{(n)} = x) = \frac{\binom{x-1}{n-1}}{\binom{N}{n}}, \qquad E X_{(n)} = \frac{n(N+1)}{n+1}.\]
Solving \(EX_{(n)} = m\), the observed maximum, for \(N\) gives the unbiased estimator \[\hat{N} = \frac{n+1}{n}\,m - 1, \qquad\text{here}\quad \hat{N} = \frac{5}{4}\cdot 60 - 1 = \mathbf{74}.\]
(For June 1940 to September 1942, conventional intelligence estimated German tank production at about 1,400 a month. The serial-number estimate was 246. German records captured after the war showed 245.)