Let \(F\) be any cdf. Let \(U \sim U(0,1)\), and define \(X =F^{-1}(U)\). Then:
\[\begin{align} F_X(x) &= P(X\leq x) = P(F^{-1}(U) \leq x) = P(U\leq F(x)) \\ & = F(x). \end{align}\]
Definition 4.1.1. An \(n\)-dimensional random vector is a function from a sample space \(S\) into \(\mathbb{R}^n\), the \(n\)-dimensional Euclidean space.
Example. Let \(s \in S\) be a newborn lamb — the physical animal you can hold in your hands. Define \(X_1 =\) its weight and \(X_2 =\) its length (both continuous), and \(X_3 = 0\) if it dies, \(1\) if it lives (discrete). Then \(\mathbf{X} = (X_1, X_2, X_3)\) is a trivariate random vector.
Definition 4.1.3. Let \((X, Y)\) be a discrete bivariate random vector. The function \(f(x, y) = P(X = x, Y = y)\) is the joint probability mass function of \((X, Y)\).
It computes the probability of any event defined in terms of \((X,Y)\), and expectations behave as we expect: \[P\bigl((X, Y) \in A\bigr) = \sum_{(x,y) \in A} f(x, y), \qquad E\,g(X, Y) = \sum_{(x,y) \in \mathbb{R}^2} g(x, y)\,f(x, y).\]
As for a single variable, \(f(x,y) \geq 0\), the values sum to one, and \(E\) is linear.
Theorem 4.1.6 (Marginal pmf). Let \((X, Y)\) be a discrete bivariate random vector with joint pmf \(f_{X,Y}(x,y)\). Then the marginal pmfs of \(X\) and \(Y\) are \[f_X(x) = \sum_{y \in \mathbb{R}} f_{X,Y}(x, y) \qquad \text{and} \qquad f_Y(y) = \sum_{x \in \mathbb{R}} f_{X,Y}(x, y).\]
Random vectors can have the same marginals even when their joint distributions are as different as they can be:
| \(X=0\) | \(X=1\) | \(f_Y\) | |
|---|---|---|---|
| \(Y=0\) | 0.45 | 0.05 | 0.50 |
| \(Y=1\) | 0.05 | 0.45 | 0.50 |
| \(f_X\) | 0.50 | 0.50 |
| \(X=0\) | \(X=1\) | \(f_Y\) | |
|---|---|---|---|
| \(Y=0\) | 0.05 | 0.45 | 0.50 |
| \(Y=1\) | 0.45 | 0.05 | 0.50 |
| \(f_X\) | 0.50 | 0.50 |
The left puts almost all its mass on \(X = Y\), the right on \(X \neq Y\). See also Example 4.1.9 in CB.
Proof of Theorem 4.1.6: notes, The bivariate case: distributions and properties.
Definition 4.1.10. A function \(f(x, y)\) from \(\mathbb{R}^2\) into \(\mathbb{R}\) is a joint probability density function of the continuous bivariate random vector \((X, Y)\) if, for every \(A \subset \mathbb{R}^2\), \[P\bigl((X, Y) \in A\bigr) = \iint_A f(x, y)\,dx\,dy.\]
Expectations and marginals work as in the discrete case, with integrals in place of sums: \[E\,g(X, Y) = \iint g(x, y)\,f(x, y)\,dx\,dy, \qquad f_X(x) = \int_{-\infty}^\infty f_{X,Y}(x, y)\,dy.\]
The joint cumulative distribution function is \(F(x, y) = P(X \leq x,\, Y \leq y)\), related to the density by \[\frac{\partial^2 F(x, y)}{\partial x\,\partial y} = f(x, y).\]
From Lecture 1: \(P(A \mid B) = P(A \cap B)/P(B)\). We now express this in terms of random variables and their distributions.
Definitions 4.2.1 and 4.2.3. Let \((X, Y)\) be a bivariate random vector with joint pmf or pdf \(f(x,y)\). For any \(x\) with \(f_X(x) > 0\), the conditional pmf or pdf of \(Y\) given \(X = x\) is \[f(y \mid x) = \frac{f(x, y)}{f_X(x)}.\]
In the discrete case this is literally \(P(Y = y \mid X = x)\), and summing over \(y\) gives one, so it is a pmf. The continuous case needs more care, since \(P(X = x) = 0\), but a limiting argument gives the same formula.
Conditional expectations are defined as expected, and we may speak of \(Y \mid X = x\) as a random variable: \[E\bigl(g(Y) \mid X = x\bigr) = \sum_y g(y)\,f(y \mid x) \quad\text{or}\quad \int_{-\infty}^\infty g(y)\,f(y \mid x)\,dy.\]
Task: work carefully through Example 4.2.4 in CB.
That \(f(y \mid x)\) is a pmf, and the limiting argument for the continuous case: notes, Conditional distributions and independence.
Definition 4.2.5. Let \((X, Y)\) be a bivariate random vector with joint pdf or pmf \(f(x,y)\) and marginals \(f_X(x)\) and \(f_Y(y)\). Then \(X\) and \(Y\) are independent if, for all \(x\) and \(y\), \[f(x, y) = f_X(x)\,f_Y(y).\]
“\(X\) and \(Y\) are independent if the conditional density of \(Y\) given \(X\) does not depend on \(X\), but is just the marginal density of \(Y\).”
Lemma 4.2.7. Let \((X, Y)\) be a bivariate random vector with joint pdf or pmf \(f(x,y)\). Then \(X\) and \(Y\) are independent if and only if there exist functions \(g(x)\) and \(h(y)\) such that for all \(x\) and \(y\), \(f(x, y) = g(x)\,h(y)\).
Proof. “\(\Rightarrow\)”: assume independence. Then \(g(x) = f_X(x)\) and \(h(y) = f_Y(y)\) by Definition 4.2.5.
“\(\Leftarrow\)”: assume \(f(x,y) = g(x)\,h(y)\). Define \(c = \int g(x)\,dx\) and \(d = \int h(y)\,dy\). Since \(\iint f(x,y)\,dx\,dy = 1\), we get \(cd = 1\).
Then \(f_X(x) = \int g(x)\,h(y)\,dy = g(x)\cdot d\) and \(f_Y(y) = \int g(x)\,h(y)\,dx = h(y)\cdot c\), so \[f_X(x)\,f_Y(y) = g(x)\,d\,h(y)\,c = g(x)\,h(y)\,cd = f(x,y). \qquad\blacksquare\]
Note the support: the factorization must hold for all \((x,y)\), so a region like \(0 < x < y < 1\) ties the variables together even when the formula factors.
Theorem 4.2.10. Let \(X\) and \(Y\) be independent.
\(P(X \in A,\, Y \in B) = P(X \in A)\,P(Y \in B)\) for any \(A, B \subset \mathbb{R}\).
For \(g\) a function only of \(x\) and \(h\) only of \(y\), \(\;E\bigl(g(X)\,h(Y)\bigr) = E\,g(X)\cdot E\,h(Y)\).
Theorem 4.2.12. If \(X\) and \(Y\) are independent with mgfs \(M_X(t)\) and \(M_Y(t)\), then \(Z = X + Y\) has mgf \(M_Z(t) = M_X(t)\,M_Y(t)\).
Proof. \(M_Z(t) = E e^{tZ} = E e^{t(X+Y)} = E\bigl(e^{tX} e^{tY}\bigr) = \bigl(E e^{tX}\bigr)\bigl(E e^{tY}\bigr) = M_X(t)\,M_Y(t)\), where the fourth equality is Theorem 4.2.10(b). \(\blacksquare\)
This is the result behind everything we did with mgfs in Lecture 2.
Proof of Theorem 4.2.10: notes, Conditional distributions and independence.
Let \((X, Y)\) have joint pdf \(f_{X,Y}(x,y)\), and define \(U = g_1(X, Y)\) and \(V = g_2(X, Y)\). What is the joint pdf of \((U, V)\)?
Assuming the transformation is one-to-one, define the Jacobian: \[J = \begin{vmatrix} \dfrac{\partial x}{\partial u} & \dfrac{\partial x}{\partial v} \\[6pt] \dfrac{\partial y}{\partial u} & \dfrac{\partial y}{\partial v} \end{vmatrix} = \frac{\partial x}{\partial u}\frac{\partial y}{\partial v} - \frac{\partial y}{\partial u}\frac{\partial x}{\partial v}.\]
By a change-of-variables result from analysis (e.g. Theorem 10.9 in Rudin), \[f_{U,V}(u, v) = f_{X,Y}\bigl(g_1^{-1}(u,v),\, g_2^{-1}(u,v)\bigr)\,|J|.\]
If the transformation is many-to-one, partition the sample space into regions where it is one-to-one and add the results.
We already used this when deriving the Beta density in Lecture 2. Instructive examples in CB §4.3: 4.3.1 sum of Poissons, 4.3.3 product of Betas, 4.3.4 sums and differences of normals, 4.3.6 ratio of normals.
We have said what it means for a relationship to be present or absent — independence. The next task is to quantify it.
Definitions 4.5.1 and 4.5.2. The covariance and correlation of \(X\) and \(Y\) are \[\text{Cov}(X, Y) = E\bigl((X - \mu_X)(Y - \mu_Y)\bigr), \qquad \rho_{XY} = \frac{\text{Cov}(X, Y)}{\sigma_X\,\sigma_Y}.\]
Theorem 4.5.3. \(\text{Cov}(X,Y) = E(XY) - \mu_X\mu_Y\).
Theorem 4.5.5. If \(X\) and \(Y\) are independent, then \(\text{Cov}(X,Y) = 0\) and \(\rho_{XY} = 0\).
Theorem 4.5.5 follows from 4.5.3 together with Theorem 4.2.10(b), which gives \(E(XY) = \mu_X\mu_Y\).
Proofs of Theorems 4.5.3 and 4.5.5: notes, Covariance and correlation.
It is crucial that the converse of Theorem 4.5.5 is not true: \(X\) and \(Y\) can be dependent with zero correlation. Covariance and correlation measure one particular form of dependence — linear dependence.
Theorem 4.5.7. For any random variables \(X\) and \(Y\):
\(-1 \leq \rho_{XY} \leq 1\).
\(|\rho_{XY}| = 1\) if and only if there exist \(a \neq 0\) and \(b\) with \(P(Y = aX + b) = 1\). If \(\rho_{XY} = 1\) then \(a > 0\); if \(\rho_{XY} = -1\) then \(a < 0\).
So \(|\rho| = 1\) exactly when \(Y\) is a linear function of \(X\). Anything else — a perfect but non-linear relationship, say \(Y = X^2\) — can give \(\rho = 0\) while \(Y\) is completely determined by \(X\).
Examples 4.5.4, 4.5.8 and 4.5.9 in CB are very instructive.
Proof of the correlation bound in (a): problem P10.
Everything generalizes in the obvious way. For \(\mathbf{X} = (X_1, \ldots, X_p)\), the pdf/pmf/cdf are \(p\)-variate functions, probabilities come from summing or integrating, and marginals come from integrating out the other variables.
Definition 4.6.5. \(X_1, \ldots, X_p\) are mutually independent if, for every \((x_1,\ldots,x_p)\), \[f(x_1, \ldots, x_p) = \prod_{i=1}^p f_{X_i}(x_i).\]
The bivariate results carry over with \(p\) factors instead of two:
Statements and arguments: notes, The general multivariate case.