3 · Multiple Random Variables
Fundamental terms
Lecture 1 was about the basics of probability using outcomes and sets.
In Lecture 2 we expressed that knowledge in terms of random variables and their distributions, which is much more useful.
We will now complete this initial tour of basic concepts by generalizing what we know to the multivariate case.
We will for this topic closely follow CB, Chapter 4.
General definition
Definition 4.1.1. An \(n\)-dimensional random vector is a function from a sample space \(S\) into \(\mathbb{R}^n\), the \(n\)-dimensional Euclidean space.
Example. Let \(S\) be the sample space of newborn sheep. Let \(s \in S\) be a specific outcome; the physical lamb that you can hold in your hands. Define the random variables
\[ \begin{aligned} X_1 &= \text{Weight of } s && \text{(cont.)} \\ X_2 &= \text{Length of } s \ \text{(measured in some standard way)} && \text{(cont.)} \\ X_3 &= 0 \text{ if dead}, \ 1 \text{ if lives} && \text{(discr.)} \end{aligned} \]
Then \(\mathbf{X} = (X_1, X_2, X_3)\) is a trivariate random vector.
The bivariate case: distributions and properties
We will concentrate on the bivariate case for now.
Definition 4.1.3. Let \((X,Y)\) be a discrete bivariate random vector (i.e. both \(X\) and \(Y\) are discrete random variables). The function \(f(x,y) : \mathbb{R}^2 \to \mathbb{R}\) defined by \(f(x,y) = P(X = x, Y = y)\) is called the joint probability mass function of \((X,Y)\).
The joint pmf can be used to compute the probability of any event defined in terms of \((X,Y)\): \[ P\big((X,Y) \in A\big) = \sum_{(x,y) \in A} f(x,y), \qquad A \subset \mathbb{R}^2. \]
Expectations are defined just as we expect (!): \[ E\,g(X,Y) = \sum_{(x,y) \in \mathbb{R}^2} g(x,y) f(x,y). \]
The expectation operator is obviously linear: \[ E\big(a g_1(X,Y) + b g_2(X,Y)\big) = a\,E g_1(X,Y) + b\,E g_2(X,Y). \]
Since \(f(x,y)\) is the probability \(P(X = x, Y = y)\): \[ f(x,y) \geq 0 \quad \forall\, (x,y) \in \mathbb{R}^2, \qquad \sum_{(x,y) \in \mathbb{R}^2} f(x,y) = 1. \]
Theorem 4.1.6 (marginal pmf). Let \((X,Y)\) be a discrete bivariate random vector with joint pmf \(f_{X,Y}(x,y)\). Then the marginal pmfs of \(X\) and \(Y\), \(f_X(x) = P(X = x)\) and \(f_Y(y) = P(Y = y)\), are given by \[ f_X(x) = \sum_{y \in \mathbb{R}} f_{X,Y}(x,y) \qquad\text{and}\qquad f_Y(y) = \sum_{x \in \mathbb{R}} f_{X,Y}(x,y). \]
Proof. \[ \begin{aligned} f_X(x) = P(X = x) &= P(X = x,\ -\infty < Y < \infty) \\ &= P\big((X,Y) \in A_x\big) \\ &= \sum_{(x,y) \in A_x} f_{X,Y}(x,y) = \sum_{y \in \mathbb{R}} f_{X,Y}(x,y), \end{aligned} \] where \(A_x = \{(x,y) : -\infty < y < \infty\}\), … and the same for \(f_Y(y)\). \(\qquad \blacksquare\)
Have a look at Example 4.1.9. It is important to understand that random vectors can have the same marginal distributions even if their joint distribution is (very!) different.
A simpler example:
\[ \begin{array}{c|cc|c} & X=0 & X=1 & f_Y(y) \\ \hline Y=0 & 0.45 & 0.05 & 0.5 \\ Y=1 & 0.05 & 0.45 & 0.5 \\ \hline f_X(x) & 0.5 & 0.5 & 1 \end{array} \qquad\qquad \begin{array}{c|cc|c} & X=0 & X=1 & f_Y(y) \\ \hline Y=0 & 0.05 & 0.45 & 0.5 \\ Y=1 & 0.45 & 0.05 & 0.5 \\ \hline f_X(x) & 0.5 & 0.5 & 1 \end{array} \]
Both joint pmfs have the same marginals, \(f_X = f_Y = (0.5, 0.5)\), but the joint distributions are as different as they can be: the left one puts almost all its mass on \(X = Y\), the right one on \(X \neq Y\).
Definition 4.1.10. A function \(f(x,y)\) from \(\mathbb{R}^2\) into \(\mathbb{R}\) is called a joint probability density function of the continuous bivariate random vector \((X,Y)\) if, for every \(A \subset \mathbb{R}^2\), \[ P\big((X,Y) \in A\big) = \iint_A f(x,y)\,dx\,dy. \]
The expectation works in the same way: \[ E\,g(X,Y) = \int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty} g(x,y) f(x,y)\,dx\,dy. \]
We calculate marginal density functions by “integrating out” the other variable (the proof is the same as in the discrete case): \[ f_X(x) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dy, \qquad f_Y(y) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dx. \]
The joint cumulative distribution function of \((X,Y)\) is defined by \[ F(x,y) = P(X \leq x,\ Y \leq y), \] which, by the fundamental theorem of calculus, is related to the joint density function via \[ \frac{\partial^2 F(x,y)}{\partial x\, \partial y} = f(x,y). \]
Conditional distributions and independence
Recall that we defined the conditional probability of event \(A\) given event \(B\) as \(P(A \mid B) = P(A \cap B)/P(B)\).
From this, we saw that the logical definition of independence between events \(A\) and \(B\) is \(P(A \mid B) = P(A)\), and hence, that \(P(A \cap B) = P(A)P(B)\).
We will express these concepts in terms of random variables and their distributions \(\rightarrow\) this is much more useful!
For two discrete random variables, the translation is straightforward and obvious:
Definition 4.2.1. Let \((X,Y)\) be a discrete bivariate random vector with joint pmf \(f(x,y)\) and marginal pmfs \(f_X(x)\) and \(f_Y(y)\). For any \(x\) such that \(P(X = x) = f_X(x) > 0\), the conditional pmf of \(Y\) given that \(X = x\) is the function \(f(y \mid x)\) defined by \[ f(y \mid x) = P(Y = y \mid X = x) = \frac{f(x,y)}{f_X(x)}. \]
We can easily check that \(f(y \mid x)\) is a pmf (p. 122 in CB): \[ \sum_y f(y \mid x) = \frac{\sum_y f(x,y)}{f_X(x)} = \frac{f_X(x)}{f_X(x)} = 1. \]
We can therefore treat \(Y \mid X = x\) as a random variable and say things like “the conditional distribution of \(Y\) given that \(X = x\) is \(f(y \mid x)\)”.
The continuous case is not that easy because \(P(X = x) = 0 \neq f_X(x)\). However, a limiting argument (see Misc. 4.9.3 in CB) reveals that a parallel definition is logically consistent with the properties of probabilities.
Definition 4.2.3. Let \((X,Y)\) be a bivariate random vector with joint pdf \(f(x,y)\) and marginal pdfs \(f_X(x)\) and \(f_Y(y)\). For any \(x\) such that \(f_X(x) > 0\), the conditional pdf of \(Y\) given that \(X = x\) is denoted by \(f(y \mid x)\) and given by \[ f(y \mid x) = \frac{f(x,y)}{f_X(x)}. \]
Conditional expectations are defined as expected: \[ \begin{aligned} E\big(g(Y) \mid X = x\big) &= \sum_y g(y) f(y \mid x) && \text{if } Y \text{ is discrete,} \\ E\big(g(Y) \mid X = x\big) &= \int_{-\infty}^{\infty} g(y) f(y \mid x)\,dy && \text{if } Y \text{ is continuous.} \end{aligned} \]
One instructive activity at this point is to carefully work through Example 4.2.4 in CB.
Definition 4.2.5 (Independence). Let \((X,Y)\) be a bivariate random vector with joint pdf or pmf \(f(x,y)\), and marginal pdfs or pmfs \(f_X(x)\) and \(f_Y(y)\). Then \(X\) and \(Y\) are called independent random variables if, for all \(x\) and \(y\), \[ f(x,y) = f_X(x) f_Y(y). \]
In the discrete case, this definition is obviously consistent with our definition of independence between events (\(P(A \cap B) = P(A)P(B)\)).
In the continuous case, we can build on our definition of conditional densities above and write, if \(X\) and \(Y\) are independent, \[ f(y \mid x) = \frac{f(x,y)}{f_X(x)} = \frac{f_X(x) f_Y(y)}{f_X(x)} = f_Y(y). \]
(“\(X\) and \(Y\) are independent if the conditional density of \(Y\) given \(X\) does not depend on \(X\), but is just the marginal density of \(Y\).”)
Lemma 4.2.7. Let \((X,Y)\) be a bivariate random vector with joint pdf or pmf \(f(x,y)\). Then \(X\) and \(Y\) are independent random variables if and only if there exist functions \(g(x)\) and \(h(y)\) such that for all \(x\) and \(y\), \(f(x,y) = g(x) h(y)\).
Proof.
“\(\Rightarrow\)” Assume independence. Then we can directly identify \(g(x) = f_X(x)\) and \(h(y) = f_Y(y)\) from Definition 4.2.5.
“\(\Leftarrow\)” Assume there exist \(g(x)\) and \(h(y)\) such that \(f(x,y) \overset{(\ast)}{=} g(x)h(y)\). Define \[ \int_{-\infty}^{\infty} g(x)\,dx = c \qquad\text{and}\qquad \int_{-\infty}^{\infty} h(y)\,dy = d, \] where it follows that \[ \begin{aligned} cd &= \left(\int_{-\infty}^{\infty} g(x)\,dx\right)\left(\int_{-\infty}^{\infty} h(y)\,dy\right) = \int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty} g(x)h(y)\,dx\,dy \\ &\overset{(\ast)}{=} \int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty} f(x,y)\,dx\,dy = 1. \end{aligned} \] But \(f_X(x) = \int_{-\infty}^{\infty} g(x)h(y)\,dy = g(x) \cdot d\) and \(f_Y(y) = \int_{-\infty}^{\infty} g(x)h(y)\,dx = h(y) \cdot c\), so \[ f(x,y) \overset{(\ast)}{=} g(x)h(y) = g(x)h(y)\,cd = f_X(x) f_Y(y). \qquad \blacksquare \]
Theorem 4.2.10. Let \(X\) and \(Y\) be independent random variables.
For any \(A \subset \mathbb{R}\) and \(B \subset \mathbb{R}\), \(P(X \in A, Y \in B) = P(X \in A)P(Y \in B)\), that is, the events \(X \in A\) and \(Y \in B\) are independent events.
Let \(g(x)\) be a function only of \(x\) and \(h(y)\) a function only of \(y\), then \(E\big(g(X)h(Y)\big) = E\big(g(X)\big) E\big(h(Y)\big)\).
The proof of part (b) is done separately for continuous variables and discrete variables in CB, using integrals and sums, respectively. Note, however, that the integral argument works for all types of random variables, including discrete and continuous, if we use the more general Lebesgue integrals (using which the probability measure is generally defined anyway).
Part (a) follows from part (b) by defining \(g(X) = \mathbf{1}\{X \in A\}\) and \(h(Y) = \mathbf{1}\{Y \in B\}\) and noting that \[ E\big(g(X)\big) = \int_{-\infty}^{\infty} g(x) f_X(x)\,dx = \int_A f_X(x)\,dx = P(X \in A), \] and similarly for \(h(\cdot)\).
Theorem 4.2.12. Let \(X\) and \(Y\) be independent random variables with moment generating functions \(M_X(t)\) and \(M_Y(t)\). Then the mgf of \(Z = X + Y\) is given by \(M_Z(t) = M_X(t) M_Y(t)\).
Proof by construction. \[ M_Z(t) = E e^{tZ} = E e^{t(X+Y)} = E e^{tX} e^{tY} = \left(E e^{tX}\right)\left(E e^{tY}\right) = M_X(t) M_Y(t). \qquad \blacksquare \]
Task: Work through Example 4.2.13 to calculate the pdf of the sum of two independent normal variables.
Bivariate transformations
We have already used this when deriving the pdf of the Beta distribution.
Let \((X,Y)\) be a random vector with joint pdf \(f(x,y)\). Define \(U = g_1(X,Y)\) and \(V = g_2(X,Y)\) — what is the joint pdf of \((U,V)\)?
Assuming that \(u = g_1(x,y)\) and \(v = g_2(x,y)\) defines a one-to-one transformation, define the Jacobian of the transformation \[ J = \begin{vmatrix} \dfrac{\partial x}{\partial u} & \dfrac{\partial x}{\partial v} \\[8pt] \dfrac{\partial y}{\partial u} & \dfrac{\partial y}{\partial v} \end{vmatrix} = \frac{\partial x}{\partial u}\frac{\partial y}{\partial v} - \frac{\partial y}{\partial u}\frac{\partial x}{\partial v}, \] which we assume is not identically zero on the sample space of \((U,V)\). By borrowing a multivariate change-of-variables result from mathematical analysis (see, e.g. Th. 10.9 in baby Rudin) we get that \[ f_{U,V}(u,v) = f_{X,Y}\big(g_1^{-1}(u,v),\, g_2^{-1}(u,v)\big)\, |J|, \] where \(|J|\) is the absolute value of \(J\).
If the transformation is many-to-one, we can consider one subset of the sample space at a time on which the transformation is one-to-one, and add the results together (assuming, of course, that such partitions exist … we can always calculate the cdf of \((U,V)\), though!).
Instructive examples in this section:
- Example 4.3.1: sum of Poisson variables
- Example 4.3.3: product of Beta variables
- Example 4.3.4: sums and differences of Normal variables
- Example 4.3.6: ratio of normal variables
Covariance and correlation
We have previously defined what we mean by the presence or absence of a relationship between (i) events and (ii) random variables.
That is: dependence and independence.
Our next task is to come up with ways to quantify this relationship.
One way is to use the covariance.
Define \(E(X) = \mu_X\), \(E(Y) = \mu_Y\), \(\operatorname{Var}(X) = \sigma_X^2\), \(\operatorname{Var}(Y) = \sigma_Y^2\), all assumed to exist.
Definition 4.5.1. The covariance of \(X\) and \(Y\) is \[ \operatorname{Cov}(X,Y) = E\big((X - \mu_X)(Y - \mu_Y)\big). \]
It is common to normalize the covariance so that we get the correlation:
Definition 4.5.2. The correlation of \(X\) and \(Y\) is \[ \operatorname{Cor}(X,Y) = \rho_{XY} = \frac{\operatorname{Cov}(X,Y)}{\sigma_X \sigma_Y}. \]
Theorem 4.5.3. For any random variables \(X\) and \(Y\), \[ \operatorname{Cov}(X,Y) = E(XY) - \mu_X \mu_Y. \]
Proof. \[ \begin{aligned} \operatorname{Cov}(X,Y) &= E\big((X - \mu_X)(Y - \mu_Y)\big) = E\big(XY - \mu_X Y - \mu_Y X + \mu_X \mu_Y\big) \\ &= E(XY) - \mu_X E(Y) - \mu_Y E(X) + \mu_X \mu_Y = E(XY) - \mu_X \mu_Y. \qquad \blacksquare \end{aligned} \]
Theorem 4.5.5. If \(X\) and \(Y\) are independent random variables then \(\operatorname{Cov}(X,Y) = 0\) and \(\rho_{XY} = 0\).
Proof. This follows directly from Theorems 4.5.3 and 4.2.10. \(\qquad \blacksquare\)
It is crucial to note that the converse statement is not necessarily true! \(X\) and \(Y\) can still be dependent even if the correlation (and thus, the covariance) between them is zero. Indeed, these two quantities are only measures of a particular form of dependence: linear dependence, as is indicated by the following result:
Theorem 4.5.7. For any random variables \(X\) and \(Y\),
- \(-1 \leq \rho_{XY} \leq 1\).
- \(|\rho_{XY}| = 1\) if and only if there exist \(a \neq 0\) and \(b\) such that \(P(Y = aX + b) = 1\). If \(\rho_{XY} = 1\) then \(a > 0\); if \(\rho_{XY} = -1\), then \(a < 0\).
Examples 4.5.4, 4.5.8 and 4.5.9 are very instructive!
The general multivariate case
The concepts and the results that we have encountered generalize to the multivariate case in the obvious ways:
\(\mathbf{X} = (X_1, X_2, \ldots, X_p)\) is a \(p\)-variate random vector.
The pdf/pmf/cdf of \(\mathbf{X}\) are \(p\)-variate functions.
If \(\mathbf{X}\) has a pmf or pdf, it must satisfy \[ \begin{aligned} &\sum_{(x_1, x_2, \ldots, x_p) \in \mathbb{R}^p} f(x_1, x_2, \ldots, x_p) = 1 \\ \text{or}\quad &\int_{-\infty}^{\infty}\!\!\int_{-\infty}^{\infty}\!\!\cdots\!\int_{-\infty}^{\infty} f(x_1, x_2, \ldots, x_p)\,dx_1 dx_2 \cdots dx_p = 1. \end{aligned} \]
Probabilities are calculated by summing or integrating the pmf or pdf over the relevant outcomes.
Expectations: \[ E\,g(\mathbf{X}) = \sum_{\mathbf{x} \in \mathbb{R}^p} g(\mathbf{x}) f(\mathbf{x}) \quad\text{or}\quad E\,g(\mathbf{X}) = \int_{-\infty}^{\infty}\!\!\cdots\!\int_{-\infty}^{\infty} g(\mathbf{x}) f(\mathbf{x})\,dx_1 dx_2 \cdots dx_p. \]
Marginal pmfs or pdfs of a subset of variables are obtained by integrating out (or summing over) all the other variables.
Definition 4.6.5. Let \(X_1, \ldots, X_p\) be random vectors with joint pdf or pmf \(f(x_1, \ldots, x_p)\). Let \(f_{X_i}(x_i)\) denote the marginal pdf or pmf of \(X_i\). Then \(X_1, \ldots, X_p\) are called mutually independent random vectors if, for every \((x_1, \ldots, x_p)\), \[ f(x_1, \ldots, x_p) = \prod_{i=1}^{p} f_{X_i}(x_i). \]
Theorem 4.6.6. If \(X_1, \ldots, X_p\) are mutually independent random variables, then \[ E\big(g_1(X_1) \cdots g_p(X_p)\big) = E g_1(X_1) \cdots E g_p(X_p). \]
Theorem 4.6.7. If \(X_1, \ldots, X_p\) are mutually independent random variables with moment generating functions \(M_{X_1}(t), \ldots, M_{X_p}(t)\), then the mgf of \(Z = X_1 + \cdots + X_p\) is \[ M_Z(t) = M_{X_1}(t) \cdots M_{X_p}(t). \]
Theorem 4.6.11. Let \(X_1, \ldots, X_p\) be random vectors. Then \(X_1, \ldots, X_p\) are mutually independent random vectors if and only if there exist functions \(g_i(x_i)\), \(i = 1, \ldots, p\), such that their joint pdf or pmf can be written as \[ f(x_1, \ldots, x_p) = \prod_{i=1}^{p} g_i(x_i). \]
None of these three results is proved here, and none needs a new idea. Theorems 4.6.6, 4.6.7 and 4.6.11 are the \(p\)-variable restatements of Theorems 4.2.10(b) and 4.2.12 and Lemma 4.2.7, and the arguments given for those go through unchanged with \(p\) factors in place of two.
Let \(\mathbf{U} = (U_1, \ldots, U_p) = \big(g_1(\mathbf{X}), \ldots, g_p(\mathbf{X})\big)\) be a one-to-one transformation of \(\mathbf{X} = (X_1, \ldots, X_p)\). The pdf of \(\mathbf{U}\) is \[ f_{\mathbf{U}}(u_1, \ldots, u_p) = f_{\mathbf{X}}\big(g_1^{-1}(\mathbf{u}), \ldots, g_p^{-1}(\mathbf{u})\big)\,|J|, \] where \(|J|\) is the determinant of the Jacobian of the transformation.