3. Multiple Random Variables

Recap

  • We built up a set of parametric distributions from first principles, starting with the Bernoulli trial.
  • We also mentioned the Normal distribution, and its fundamental nature due to the Central Limit Theorem. We will build more things on top of that in the next topic (\(\chi^2\), \(t\), \(F\)).
  • We also mentioned the uniform which is another very fundamental distribution with the following remarkable property:

Let \(F\) be any cdf. Let \(U \sim U(0,1)\), and define \(X =F^{-1}(U)\). Then:

\[\begin{align} F_X(x) &= P(X\leq x) = P(F^{-1}(U) \leq x) = P(U\leq F(x)) \\ & = F(x). \end{align}\]

  • The distribution of \(F^{-1}(U)\) is \(F\). We can prove, in the same way, that the distribution of \(F(X)\) is \(U(0,1)\), if \(F\) is continuous.
  • \(\Rightarrow\) Any distribution \(F\) can be generated from the uniform, and the uniform can be generated from any (continuous) distribution!

Where we are

  • Lecture 1 was about the basics of probability using outcomes and sets.
  • In Lecture 2 we expressed that knowledge in terms of random variables and their distributions, which is much more useful.
  • We will now complete this initial tour of basic concepts by generalizing to the multivariate case.
  • We will for this topic closely follow CB, Chapter 4.

Random vectors and joint distributions

Definition 4.1.1. An \(n\)-dimensional random vector is a function from a sample space \(S\) into \(\mathbb{R}^n\), the \(n\)-dimensional Euclidean space.

Example. Let \(s \in S\) be a newborn lamb — the physical animal you can hold in your hands. Define \(X_1 =\) its weight and \(X_2 =\) its length (both continuous), and \(X_3 = 0\) if it dies, \(1\) if it lives (discrete). Then \(\mathbf{X} = (X_1, X_2, X_3)\) is a trivariate random vector.

Definition 4.1.3. Let \((X, Y)\) be a discrete bivariate random vector. The function \(f(x, y) = P(X = x, Y = y)\) is the joint probability mass function of \((X, Y)\).

It computes the probability of any event defined in terms of \((X,Y)\), and expectations behave as we expect: \[P\bigl((X, Y) \in A\bigr) = \sum_{(x,y) \in A} f(x, y), \qquad E\,g(X, Y) = \sum_{(x,y) \in \mathbb{R}^2} g(x, y)\,f(x, y).\]

As for a single variable, \(f(x,y) \geq 0\), the values sum to one, and \(E\) is linear.

Marginals, and what they leave out ★

Theorem 4.1.6 (Marginal pmf). Let \((X, Y)\) be a discrete bivariate random vector with joint pmf \(f_{X,Y}(x,y)\). Then the marginal pmfs of \(X\) and \(Y\) are \[f_X(x) = \sum_{y \in \mathbb{R}} f_{X,Y}(x, y) \qquad \text{and} \qquad f_Y(y) = \sum_{x \in \mathbb{R}} f_{X,Y}(x, y).\]

Random vectors can have the same marginals even when their joint distributions are as different as they can be:

\(X=0\) \(X=1\) \(f_Y\)
\(Y=0\) 0.45 0.05 0.50
\(Y=1\) 0.05 0.45 0.50
\(f_X\) 0.50 0.50
\(X=0\) \(X=1\) \(f_Y\)
\(Y=0\) 0.05 0.45 0.50
\(Y=1\) 0.45 0.05 0.50
\(f_X\) 0.50 0.50

The left puts almost all its mass on \(X = Y\), the right on \(X \neq Y\). See also Example 4.1.9 in CB.

Proof of Theorem 4.1.6: notes, The bivariate case: distributions and properties.

The continuous case, and the joint cdf/pdf

Definition 4.1.10. A function \(f(x, y)\) from \(\mathbb{R}^2\) into \(\mathbb{R}\) is a joint probability density function of the continuous bivariate random vector \((X, Y)\) if, for every \(A \subset \mathbb{R}^2\), \[P\bigl((X, Y) \in A\bigr) = \iint_A f(x, y)\,dx\,dy.\]

Expectations and marginals work as in the discrete case, with integrals in place of sums: \[E\,g(X, Y) = \iint g(x, y)\,f(x, y)\,dx\,dy, \qquad f_X(x) = \int_{-\infty}^\infty f_{X,Y}(x, y)\,dy.\]

The joint cumulative distribution function is \(F(x, y) = P(X \leq x,\, Y \leq y)\), related to the density by \[\frac{\partial^2 F(x, y)}{\partial x\,\partial y} = f(x, y).\]

Conditional distributions

From Lecture 1: \(P(A \mid B) = P(A \cap B)/P(B)\). We now express this in terms of random variables and their distributions.

Definitions 4.2.1 and 4.2.3. Let \((X, Y)\) be a bivariate random vector with joint pmf or pdf \(f(x,y)\). For any \(x\) with \(f_X(x) > 0\), the conditional pmf or pdf of \(Y\) given \(X = x\) is \[f(y \mid x) = \frac{f(x, y)}{f_X(x)}.\]

In the discrete case this is literally \(P(Y = y \mid X = x)\), and summing over \(y\) gives one, so it is a pmf. The continuous case needs more care, since \(P(X = x) = 0\), but a limiting argument gives the same formula.

Conditional expectations are defined as expected, and we may speak of \(Y \mid X = x\) as a random variable: \[E\bigl(g(Y) \mid X = x\bigr) = \sum_y g(y)\,f(y \mid x) \quad\text{or}\quad \int_{-\infty}^\infty g(y)\,f(y \mid x)\,dy.\]

Task: work carefully through Example 4.2.4 in CB.

That \(f(y \mid x)\) is a pmf, and the limiting argument for the continuous case: notes, Conditional distributions and independence.

Independence

Definition 4.2.5. Let \((X, Y)\) be a bivariate random vector with joint pdf or pmf \(f(x,y)\) and marginals \(f_X(x)\) and \(f_Y(y)\). Then \(X\) and \(Y\) are independent if, for all \(x\) and \(y\), \[f(x, y) = f_X(x)\,f_Y(y).\]

  • In the discrete case this is exactly \(P(A \cap B) = P(A)P(B)\).
  • If \(X\) and \(Y\) are independent, then \[f(y \mid x) = \frac{f(x,y)}{f_X(x)} = \frac{f_X(x)\,f_Y(y)}{f_X(x)} = f_Y(y).\]

“\(X\) and \(Y\) are independent if the conditional density of \(Y\) given \(X\) does not depend on \(X\), but is just the marginal density of \(Y\).”

Independence by factorization

Lemma 4.2.7. Let \((X, Y)\) be a bivariate random vector with joint pdf or pmf \(f(x,y)\). Then \(X\) and \(Y\) are independent if and only if there exist functions \(g(x)\) and \(h(y)\) such that for all \(x\) and \(y\), \(f(x, y) = g(x)\,h(y)\).

Proof. “\(\Rightarrow\)”: assume independence. Then \(g(x) = f_X(x)\) and \(h(y) = f_Y(y)\) by Definition 4.2.5.

“\(\Leftarrow\)”: assume \(f(x,y) = g(x)\,h(y)\). Define \(c = \int g(x)\,dx\) and \(d = \int h(y)\,dy\). Since \(\iint f(x,y)\,dx\,dy = 1\), we get \(cd = 1\).

Then \(f_X(x) = \int g(x)\,h(y)\,dy = g(x)\cdot d\) and \(f_Y(y) = \int g(x)\,h(y)\,dx = h(y)\cdot c\), so \[f_X(x)\,f_Y(y) = g(x)\,d\,h(y)\,c = g(x)\,h(y)\,cd = f(x,y). \qquad\blacksquare\]

Note the support: the factorization must hold for all \((x,y)\), so a region like \(0 < x < y < 1\) ties the variables together even when the formula factors.

Consequences of independence

Theorem 4.2.10. Let \(X\) and \(Y\) be independent.

  1. \(P(X \in A,\, Y \in B) = P(X \in A)\,P(Y \in B)\) for any \(A, B \subset \mathbb{R}\).

  2. For \(g\) a function only of \(x\) and \(h\) only of \(y\), \(\;E\bigl(g(X)\,h(Y)\bigr) = E\,g(X)\cdot E\,h(Y)\).

Theorem 4.2.12. If \(X\) and \(Y\) are independent with mgfs \(M_X(t)\) and \(M_Y(t)\), then \(Z = X + Y\) has mgf \(M_Z(t) = M_X(t)\,M_Y(t)\).

Proof. \(M_Z(t) = E e^{tZ} = E e^{t(X+Y)} = E\bigl(e^{tX} e^{tY}\bigr) = \bigl(E e^{tX}\bigr)\bigl(E e^{tY}\bigr) = M_X(t)\,M_Y(t)\), where the fourth equality is Theorem 4.2.10(b). \(\blacksquare\)

This is the result behind everything we did with mgfs in Lecture 2.

Proof of Theorem 4.2.10: notes, Conditional distributions and independence.

Bivariate transformations ★

Let \((X, Y)\) have joint pdf \(f_{X,Y}(x,y)\), and define \(U = g_1(X, Y)\) and \(V = g_2(X, Y)\). What is the joint pdf of \((U, V)\)?

Assuming the transformation is one-to-one, define the Jacobian: \[J = \begin{vmatrix} \dfrac{\partial x}{\partial u} & \dfrac{\partial x}{\partial v} \\[6pt] \dfrac{\partial y}{\partial u} & \dfrac{\partial y}{\partial v} \end{vmatrix} = \frac{\partial x}{\partial u}\frac{\partial y}{\partial v} - \frac{\partial y}{\partial u}\frac{\partial x}{\partial v}.\]

By a change-of-variables result from analysis (e.g. Theorem 10.9 in Rudin), \[f_{U,V}(u, v) = f_{X,Y}\bigl(g_1^{-1}(u,v),\, g_2^{-1}(u,v)\bigr)\,|J|.\]

If the transformation is many-to-one, partition the sample space into regions where it is one-to-one and add the results.

We already used this when deriving the Beta density in Lecture 2. Instructive examples in CB §4.3: 4.3.1 sum of Poissons, 4.3.3 product of Betas, 4.3.4 sums and differences of normals, 4.3.6 ratio of normals.

Covariance and correlation

We have said what it means for a relationship to be present or absent — independence. The next task is to quantify it.

Definitions 4.5.1 and 4.5.2. The covariance and correlation of \(X\) and \(Y\) are \[\text{Cov}(X, Y) = E\bigl((X - \mu_X)(Y - \mu_Y)\bigr), \qquad \rho_{XY} = \frac{\text{Cov}(X, Y)}{\sigma_X\,\sigma_Y}.\]

Theorem 4.5.3. \(\text{Cov}(X,Y) = E(XY) - \mu_X\mu_Y\).

Theorem 4.5.5. If \(X\) and \(Y\) are independent, then \(\text{Cov}(X,Y) = 0\) and \(\rho_{XY} = 0\).

Theorem 4.5.5 follows from 4.5.3 together with Theorem 4.2.10(b), which gives \(E(XY) = \mu_X\mu_Y\).

Proofs of Theorems 4.5.3 and 4.5.5: notes, Covariance and correlation.

Correlation measures linear dependence only ★

It is crucial that the converse of Theorem 4.5.5 is not true: \(X\) and \(Y\) can be dependent with zero correlation. Covariance and correlation measure one particular form of dependence — linear dependence.

Theorem 4.5.7. For any random variables \(X\) and \(Y\):

  1. \(-1 \leq \rho_{XY} \leq 1\).

  2. \(|\rho_{XY}| = 1\) if and only if there exist \(a \neq 0\) and \(b\) with \(P(Y = aX + b) = 1\). If \(\rho_{XY} = 1\) then \(a > 0\); if \(\rho_{XY} = -1\) then \(a < 0\).

So \(|\rho| = 1\) exactly when \(Y\) is a linear function of \(X\). Anything else — a perfect but non-linear relationship, say \(Y = X^2\) — can give \(\rho = 0\) while \(Y\) is completely determined by \(X\).

Examples 4.5.4, 4.5.8 and 4.5.9 in CB are very instructive.

Proof of the correlation bound in (a): problem P10.

The general multivariate case ★

Everything generalizes in the obvious way. For \(\mathbf{X} = (X_1, \ldots, X_p)\), the pdf/pmf/cdf are \(p\)-variate functions, probabilities come from summing or integrating, and marginals come from integrating out the other variables.

Definition 4.6.5. \(X_1, \ldots, X_p\) are mutually independent if, for every \((x_1,\ldots,x_p)\), \[f(x_1, \ldots, x_p) = \prod_{i=1}^p f_{X_i}(x_i).\]

The bivariate results carry over with \(p\) factors instead of two:

  • Theorem 4.6.6. \(E\bigl(g_1(X_1)\cdots g_p(X_p)\bigr) = E\,g_1(X_1)\cdots E\,g_p(X_p)\).
  • Theorem 4.6.7. \(Z = X_1 + \cdots + X_p\) has mgf \(M_Z(t) = M_{X_1}(t)\cdots M_{X_p}(t)\).
  • Theorem 4.6.11. Mutual independence holds if and only if \(f(x_1,\ldots,x_p) = \prod_i g_i(x_i)\) for some functions \(g_i\).
  • Transformations: the same formula, with \(|J|\) the determinant of the \(p \times p\) Jacobian.

Statements and arguments: notes, The general multivariate case.