Probability · from the model to the data
Statistics · from the data to the model
Probability is the mathematical study of events that are understood in practice to be random:
The outcome of a process is not known beforehand, but various alternatives are more or less probable or likely, and the probability or likelihood of an outcome may depend on other things, such as the outcome of another event.
Statistics, or more precisely statistical inference, is the process of deriving general knowledge about the nature and properties of random processes, based on observed outcomes (data).
Define the following terms:
Random experiment
Sample space
(Relative) frequency
A bit harder: Probability.
A set \(C\) is a collection of elements. We can define a set using notation, e.g. \[ C = \{x : 0 \leq x \leq 1\}, \] and write, for example, \(\tfrac{1}{2} \in C\), \(2 \notin C\).
The set \(C\) is countable if its elements can be uniquely enumerated by the positive integers.
If each element of \(C_1\) is also an element of \(C_2\), then \(C_1\) is a subset of \(C_2\): \[C_1 \subset C_2.\]
If \(C_1 \subset C_2\) and \(C_2 \subset C_1\), then \(C_1 = C_2\).
Formally, a probability space is a special case of a measure space from mathematical analysis. It has three components, say \(\mathcal{X}\), \(\mathcal{B}\), and \(P\).
More informally:
\(\mathcal{X}\) is the set of all possible outcomes — the sample space.
\(\mathcal{B}\) is the collection of all combinations of outcomes (events) to which we can associate a probability — the measurable space, and this collection is known as a \(\sigma\)-algebra.
\(P\) is the probability function \(P : \mathcal{B} \to [0,1]\) that associates a probability to each event in \(\mathcal{B}\).
Definition 3.1 (Hogg, McKean, Craig). Let \(\mathcal{X}\) be a sample space and let \(\mathcal{B}\) be the set of events. Let \(P\) be a real-valued function on \(\mathcal{B}\). Then \(P\) is a probability set function if it satisfies:
These properties are derived from the three axioms.
For each event \(C \in \mathcal{B}\), \(\quad P(C) = 1 - P(C^c)\).
Proof. \(\mathcal{X} = C \cup C^c\), \(C \cap C^c = \emptyset\), so from axioms 2 and 3: \[1 = P(\mathcal{X}) = P(C \cup C^c) = P(C) + P(C^c). \qquad\blacksquare\]
\(P(\emptyset) = 0\).
Proof. Apply result 1 with \(C = \emptyset\), so \(C^c = \mathcal{X}\): \[P(\emptyset) = 1 - P(\mathcal{X}) = 1 - 1 = 0. \qquad\blacksquare\]
If \(C_1 \subset C_2\), then \(P(C_1) \leq P(C_2)\).
Proof. \(C_2 = C_1 \cup (C_1^c \cap C_2)\) with non-intersecting parts, so \(P(C_2) = P(C_1) + \underbrace{P(C_1^c \cap C_2)}_{\geq\,0\text{ from axiom 1}}\). \(\blacksquare\)
For each \(C \in \mathcal{B}\), \(\quad 0 \leq P(C) \leq 1\).
Proof. \(\emptyset \subset C \subset \mathcal{X}\); apply result 3. \(\blacksquare\)
\(P(C_1 \cup C_2) = P(C_1) + P(C_2) - P(C_1 \cap C_2)\).
Proof. Write \(C_1 \cup C_2 = C_1 \cup (C_1^c \cap C_2)\) and \(C_2 = (C_1 \cap C_2) \cup (C_1^c \cap C_2)\), both non-intersecting. This gives two expressions for \(P(C_1^c \cap C_2)\) that are equal, yielding the result. \(\blacksquare\)
The equally-likely framework reduces probability to counting. Key rules:
\(mn\) rule. If experiment 1 has \(m\) possible outcomes and experiment 2 has \(n\), then together they have \(mn\) possible outcomes.
Ordered selection. There are \[ n(n-1)\cdots(n-k+1) = \frac{n!}{(n-k)!} \] different ways to select \(k\) unique items in order from a set with \(n\) elements.
Unordered selection. If ordering does not matter, there are \[ \binom{n}{k} = \frac{n!}{k!\,(n-k)!} \] different ways to select \(k\) unique elements from a set of \(n\) elements. This is the binomial coefficient.
The conditional probability of \(C_2\) given \(C_1\), provided \(P(C_1) > 0\), is defined as \[ P(C_2 \mid C_1) = \frac{P(C_1 \cap C_2)}{P(C_1)}. \]
This definition immediately leads to the law of total probability: let \(C_1, C_2, \ldots, C_k\) be a partition of \(\mathcal{X}\). Then \[ P(C) = \sum_{i=1}^k P(C_i \cap C) = \sum_{i=1}^k P(C \mid C_i)\,P(C_i). \]
… which in turn leads to the important Bayes’ theorem:
Bayes’ Theorem. Let \(C_1, C_2, \ldots, C_k\) be a partition of \(\mathcal{X}\) with \(P(C_i) > 0\) for all \(i\). Then for any event \(C\) with \(P(C) > 0\), \[ P(C_j \mid C) = \frac{P(C_j)\,P(C \mid C_j)}{\displaystyle\sum_{i=1}^k P(C_i)\,P(C \mid C_i)}. \]
If knowing that \(C_2\) will happen does not change the probability of \(C_1\) happening, we have \[ P(C_1 \mid C_2) = P(C_1), \] and hence \(P(C_1 \cap C_2) = P(C_1 \mid C_2)\,P(C_2) = P(C_1)\,P(C_2)\).
We take this as the definition of independence between \(C_1\) and \(C_2\):
Events \(C_1\) and \(C_2\) are independent if \[ P(C_1 \cap C_2) = P(C_1)\,P(C_2). \]
So far we have discussed probabilities in terms of general “sets”. To use mathematics to deal with random outcomes, we add another layer of abstraction:
A random variable is a function that translates outcomes to numbers.
Formally: consider an experiment with sample space \(\mathcal{X}\). A function \(X\) which assigns to each element \(c \in \mathcal{X}\) one and only one number \(X(c) = x\) is called a random variable. The space or range of \(X\) is \[ \mathcal{D} = \{x : x = X(c),\; c \in \mathcal{X}\}. \]
Example (die). A die is cast; the outcome \(c\) is the physical die lying on the table.
Can you give an example of an outcome that gives rise to several different random variables?
Let \(X\) be a random variable. Its cumulative distribution function (or simply distribution function) is defined by \[ F_X(x) = P_X\!\big((-\infty, x]\big) = P\!\big(\{c \in \mathcal{X} : X(c) \leq x\}\big). \]
Let \(A \subseteq \mathbb{R}\). In the general measure-theoretic sense we can always write \[ P_X(X \in A) = \int_A dF_X(x), \] using the general Lebesgue integral. This reduces to familiar forms in practice:
Continuous random variable. If \(F(x)\) is differentiable over \(\mathcal{D}\), then \(f(x) = dF(x)/dx\) is the density function of \(X\), and \[ P(A) = \int_A f(x)\,dx \quad \text{(using the familiar Riemann integral)}. \]
Discrete random variable. If \(\mathcal{D}\) is countable, the Lebesgue integral becomes a sum: \[ P(A) = \sum_{x \in A} p_X(x), \] where \(p_X(x) = P(X = x)\) is the probability mass function of \(X\).
The random variable \(X\) has distribution function \(F_X\). What is the distribution of \(Y = g(X)\)?
Discrete, one-to-one \(g\): \(p_Y(y) = P(X = g^{-1}(y)) = p_X(g^{-1}(y))\).
Continuous, differentiable, one-to-one \(g\): Work via the CDF: \[ F_Y(y) = P(Y \leq y) = P(g(X) \leq y) = \begin{cases} F_X(g^{-1}(y)) & \text{if } g \text{ is increasing,} \\ 1 - F_X(g^{-1}(y)) & \text{if } g \text{ is decreasing.} \end{cases} \] Differentiating in either case (chain rule!): \[ f_Y(y) = f_X\!\big(g^{-1}(y)\big)\left|\frac{dx}{dy}\right|. \]
The general definition of the expected value of \(X\), if the integral exists, is \[ E(X) = \int_{\mathcal{D}} x\,dF_X(x). \]
This reduces to familiar forms: \[ E(X) = \begin{cases} \displaystyle\sum_{x \in \mathcal{D}} x\,P(X = x) & \text{(discrete)} \\[8pt] \displaystyle\int_{\mathcal{D}} x\,f(x)\,dx & \text{(continuous).} \end{cases} \]
Theorem (Law of the Unconscious Statistician). Let \(X\) be a random variable and let \(Y = g(X)\) for some function \(g(\cdot)\). The expected value of \(Y\) is \[ E(Y) = E(g(X)) = \int_{\mathcal{D}} g(x)\,dF_X(x), \] provided the integral exists.
This reduces to ordinary sums and integrals for discrete and continuous random variables. The proof is a change-of-variables operation (a non-trivial proof, that is, hence the term “unconscious”, because the result itself appears as totally obvious).
The expectation is a linear operator: \[ E\!\big(k_1 g_1(X) + k_2 g_2(X)\big) = k_1\,E(g_1(X)) + k_2\,E(g_2(X)). \]
The variance of a random variable is \[ \operatorname{Var}(X) = E\!\big((X - \mu)^2\big) = E(X^2) - \mu^2, \] otherwise known as the second central moment of \(X\).
More generally, \(E(X^p)\) is the \(p\)th moment of \(X\).
Let \(X\) be a random variable such that for some \(h > 0\), the expectation of \(\exp(tX)\) exists for \(-h < t < h\). The moment generating function of \(X\) is \[ M_X(t) = E\!\big(\exp(tX)\big), \quad -h < t < h. \]
Uniqueness. Let \(X\) and \(Y\) be random variables with MGFs \(M_X\) and \(M_Y\) existing in open intervals about \(0\). Then \(F_X(z) = F_Y(z)\) for all \(z \in \mathbb{R}\) if and only if \(M_X(t) = M_Y(t)\) for all \(t \in (-h, h)\) for some \(h > 0\).
Moments from the MGF. Let \(X\) be a random variable with MGF \(M_X(t)\). If \(m\) is a positive integer, then \(M_X^{(m)}(0) = E(X^m)\).