1. Review of Basic Probability

The road ahead

Probability  ·  from the model to the data

events1 probability1 random variables1 distributions2–3 samples4
the model sufficiency7 estimation6 approximation5

Statistics  ·  from the data to the model

Fundamental terms

  • Probability is the mathematical study of events that are understood in practice to be random:

    The outcome of a process is not known beforehand, but various alternatives are more or less probable or likely, and the probability or likelihood of an outcome may depend on other things, such as the outcome of another event.

  • Statistics, or more precisely statistical inference, is the process of deriving general knowledge about the nature and properties of random processes, based on observed outcomes (data).

Fundamental terms — discussion

Define the following terms:

  • Random experiment

  • Sample space

  • (Relative) frequency

A bit harder: Probability.

  • Frequentist: Probability is the long-run relative frequency of an event in a large number of repeated, identical experiments. (e.g., a fair coin lands heads 50% of the time in the limit.)
  • Classical (Laplace): Probability is the ratio of favorable outcomes to total equally likely outcomes. (e.g., rolling a die: P(1) = 1/6.)
  • Subjective (Bayesian): Probability is a degree of belief — a measure of how confident a rational agent is that an event will occur, updated as new evidence arrives.

Sets

  • A set \(C\) is a collection of elements. We can define a set using notation, e.g. \[ C = \{x : 0 \leq x \leq 1\}, \] and write, for example, \(\tfrac{1}{2} \in C\), \(2 \notin C\).

  • The set \(C\) is countable if its elements can be uniquely enumerated by the positive integers.

  • If each element of \(C_1\) is also an element of \(C_2\), then \(C_1\) is a subset of \(C_2\): \[C_1 \subset C_2.\]

  • If \(C_1 \subset C_2\) and \(C_2 \subset C_1\), then \(C_1 = C_2\).

Set operations

  • Union: \(C_1 \cup C_2 = \{x : x \in C_1 \text{ or } x \in C_2\}\)
  • Intersection: \(C_1 \cap C_2 = \{x : x \in C_1 \text{ and } x \in C_2\}\)
  • Complement: \(C^c = \{x : x \notin C\}\)
  • A set with no elements is the null set or empty set, written \(C = \emptyset\).
  • The union and intersection of several (possibly infinitely many) sets \(C_1, C_2, \ldots\): \[ \bigcup_{j=1}^{\infty} C_j = \{x : x \in C_j \text{ for at least one } j\}, \qquad \bigcap_{j=1}^{\infty} C_j = \{x : x \in C_j \text{ for all } j\}. \]
  • The set of all elements under consideration is typically called a space.

Probability spaces

Formally, a probability space is a special case of a measure space from mathematical analysis. It has three components, say \(\mathcal{X}\), \(\mathcal{B}\), and \(P\).

More informally:

  • \(\mathcal{X}\) is the set of all possible outcomes — the sample space.

  • \(\mathcal{B}\) is the collection of all combinations of outcomes (events) to which we can associate a probability — the measurable space, and this collection is known as a \(\sigma\)-algebra.

  • \(P\) is the probability function \(P : \mathcal{B} \to [0,1]\) that associates a probability to each event in \(\mathcal{B}\).

The probability set function

Definition 3.1 (Hogg, McKean, Craig). Let \(\mathcal{X}\) be a sample space and let \(\mathcal{B}\) be the set of events. Let \(P\) be a real-valued function on \(\mathcal{B}\). Then \(P\) is a probability set function if it satisfies:

  1. \(P(C) \geq 0\) for all \(C \in \mathcal{B}\),
  2. \(P(\mathcal{X}) = 1\),
  3. If \(\{C_n\}\) is a sequence of events in \(\mathcal{B}\) with \(C_m \cap C_n = \emptyset\) for \(m \neq n\), then \(\displaystyle P\!\left(\bigcup_{n=1}^{\infty} C_n\right) = \sum_{n=1}^{\infty} P(C_n)\).

Properties of \(P\) — complement and empty set

These properties are derived from the three axioms.

  1. For each event \(C \in \mathcal{B}\), \(\quad P(C) = 1 - P(C^c)\).

    Proof. \(\mathcal{X} = C \cup C^c\), \(C \cap C^c = \emptyset\), so from axioms 2 and 3: \[1 = P(\mathcal{X}) = P(C \cup C^c) = P(C) + P(C^c). \qquad\blacksquare\]

  2. \(P(\emptyset) = 0\).

    Proof. Apply result 1 with \(C = \emptyset\), so \(C^c = \mathcal{X}\): \[P(\emptyset) = 1 - P(\mathcal{X}) = 1 - 1 = 0. \qquad\blacksquare\]

Properties of \(P\) — monotonicity, bounds, and addition

  1. If \(C_1 \subset C_2\), then \(P(C_1) \leq P(C_2)\).

    Proof. \(C_2 = C_1 \cup (C_1^c \cap C_2)\) with non-intersecting parts, so \(P(C_2) = P(C_1) + \underbrace{P(C_1^c \cap C_2)}_{\geq\,0\text{ from axiom 1}}\). \(\blacksquare\)

  2. For each \(C \in \mathcal{B}\), \(\quad 0 \leq P(C) \leq 1\).

    Proof. \(\emptyset \subset C \subset \mathcal{X}\); apply result 3. \(\blacksquare\)

  3. \(P(C_1 \cup C_2) = P(C_1) + P(C_2) - P(C_1 \cap C_2)\).

    Proof. Write \(C_1 \cup C_2 = C_1 \cup (C_1^c \cap C_2)\) and \(C_2 = (C_1 \cap C_2) \cup (C_1^c \cap C_2)\), both non-intersecting. This gives two expressions for \(P(C_1^c \cap C_2)\) that are equal, yielding the result. \(\blacksquare\)

Equally likely outcomes

  1. Let \(C_1, C_2, \ldots, C_k\) be mutually exclusive and exhaustive events that are equally likely, i.e. \(P(C_j) = 1/k\) for all \(j \in \{1, \ldots, k\}\). Let \(E = \{C_\ell : \ell \in L \subset \{1,\ldots,k\}\}\) be a union of \(r \leq k\) of these events. Then: \[ P(E) = \sum_{\ell \in L} P(C_\ell) = \frac{r}{k}. \] The probability of \(E\) is the number of ways \(E\) can happen, divided by the total number of ways the experiment may terminate.

Counting rules

The equally-likely framework reduces probability to counting. Key rules:

  • \(mn\) rule. If experiment 1 has \(m\) possible outcomes and experiment 2 has \(n\), then together they have \(mn\) possible outcomes.

  • Ordered selection. There are \[ n(n-1)\cdots(n-k+1) = \frac{n!}{(n-k)!} \] different ways to select \(k\) unique items in order from a set with \(n\) elements.

  • Unordered selection. If ordering does not matter, there are \[ \binom{n}{k} = \frac{n!}{k!\,(n-k)!} \] different ways to select \(k\) unique elements from a set of \(n\) elements. This is the binomial coefficient.

Conditional probability

The conditional probability of \(C_2\) given \(C_1\), provided \(P(C_1) > 0\), is defined as \[ P(C_2 \mid C_1) = \frac{P(C_1 \cap C_2)}{P(C_1)}. \]

This definition immediately leads to the law of total probability: let \(C_1, C_2, \ldots, C_k\) be a partition of \(\mathcal{X}\). Then \[ P(C) = \sum_{i=1}^k P(C_i \cap C) = \sum_{i=1}^k P(C \mid C_i)\,P(C_i). \]

Bayes’ theorem

… which in turn leads to the important Bayes’ theorem:

Bayes’ Theorem. Let \(C_1, C_2, \ldots, C_k\) be a partition of \(\mathcal{X}\) with \(P(C_i) > 0\) for all \(i\). Then for any event \(C\) with \(P(C) > 0\), \[ P(C_j \mid C) = \frac{P(C_j)\,P(C \mid C_j)}{\displaystyle\sum_{i=1}^k P(C_i)\,P(C \mid C_i)}. \]

Independence

If knowing that \(C_2\) will happen does not change the probability of \(C_1\) happening, we have \[ P(C_1 \mid C_2) = P(C_1), \] and hence \(P(C_1 \cap C_2) = P(C_1 \mid C_2)\,P(C_2) = P(C_1)\,P(C_2)\).

We take this as the definition of independence between \(C_1\) and \(C_2\):

Events \(C_1\) and \(C_2\) are independent if \[ P(C_1 \cap C_2) = P(C_1)\,P(C_2). \]

Random variables

So far we have discussed probabilities in terms of general “sets”. To use mathematics to deal with random outcomes, we add another layer of abstraction:

A random variable is a function that translates outcomes to numbers.

Formally: consider an experiment with sample space \(\mathcal{X}\). A function \(X\) which assigns to each element \(c \in \mathcal{X}\) one and only one number \(X(c) = x\) is called a random variable. The space or range of \(X\) is \[ \mathcal{D} = \{x : x = X(c),\; c \in \mathcal{X}\}. \]

Random variables — example

Example (die). A die is cast; the outcome \(c\) is the physical die lying on the table.

  • The sample space \(\mathcal{X}\) is all possible ways a die can land on a horizontal surface.
  • Define the mapping \(X : \mathcal{X} \to \{1, 2, 3, 4, 5, 6\}\) such that \[\begin{align*} X(c) = &\text{"number of dots on the side of the die} \\ & \qquad \qquad \text{that faces away from the surface"}. \end{align*}\]

Can you give an example of an outcome that gives rise to several different random variables?

The cumulative distribution function

Let \(X\) be a random variable. Its cumulative distribution function (or simply distribution function) is defined by \[ F_X(x) = P_X\!\big((-\infty, x]\big) = P\!\big(\{c \in \mathcal{X} : X(c) \leq x\}\big). \]

Calculating probabilities

Let \(A \subseteq \mathbb{R}\). In the general measure-theoretic sense we can always write \[ P_X(X \in A) = \int_A dF_X(x), \] using the general Lebesgue integral. This reduces to familiar forms in practice:

  • Continuous random variable. If \(F(x)\) is differentiable over \(\mathcal{D}\), then \(f(x) = dF(x)/dx\) is the density function of \(X\), and \[ P(A) = \int_A f(x)\,dx \quad \text{(using the familiar Riemann integral)}. \]

  • Discrete random variable. If \(\mathcal{D}\) is countable, the Lebesgue integral becomes a sum: \[ P(A) = \sum_{x \in A} p_X(x), \] where \(p_X(x) = P(X = x)\) is the probability mass function of \(X\).

Transformations

The random variable \(X\) has distribution function \(F_X\). What is the distribution of \(Y = g(X)\)?

  • Discrete, one-to-one \(g\): \(p_Y(y) = P(X = g^{-1}(y)) = p_X(g^{-1}(y))\).

  • Continuous, differentiable, one-to-one \(g\): Work via the CDF: \[ F_Y(y) = P(Y \leq y) = P(g(X) \leq y) = \begin{cases} F_X(g^{-1}(y)) & \text{if } g \text{ is increasing,} \\ 1 - F_X(g^{-1}(y)) & \text{if } g \text{ is decreasing.} \end{cases} \] Differentiating in either case (chain rule!): \[ f_Y(y) = f_X\!\big(g^{-1}(y)\big)\left|\frac{dx}{dy}\right|. \]

Expectations

The general definition of the expected value of \(X\), if the integral exists, is \[ E(X) = \int_{\mathcal{D}} x\,dF_X(x). \]

This reduces to familiar forms: \[ E(X) = \begin{cases} \displaystyle\sum_{x \in \mathcal{D}} x\,P(X = x) & \text{(discrete)} \\[8pt] \displaystyle\int_{\mathcal{D}} x\,f(x)\,dx & \text{(continuous).} \end{cases} \]

Expectation of a function

Theorem (Law of the Unconscious Statistician). Let \(X\) be a random variable and let \(Y = g(X)\) for some function \(g(\cdot)\). The expected value of \(Y\) is \[ E(Y) = E(g(X)) = \int_{\mathcal{D}} g(x)\,dF_X(x), \] provided the integral exists.

This reduces to ordinary sums and integrals for discrete and continuous random variables. The proof is a change-of-variables operation (a non-trivial proof, that is, hence the term “unconscious”, because the result itself appears as totally obvious).

The expectation is a linear operator: \[ E\!\big(k_1 g_1(X) + k_2 g_2(X)\big) = k_1\,E(g_1(X)) + k_2\,E(g_2(X)). \]

Higher moments

The variance of a random variable is \[ \operatorname{Var}(X) = E\!\big((X - \mu)^2\big) = E(X^2) - \mu^2, \] otherwise known as the second central moment of \(X\).

More generally, \(E(X^p)\) is the \(p\)th moment of \(X\).

Moment generating functions

Let \(X\) be a random variable such that for some \(h > 0\), the expectation of \(\exp(tX)\) exists for \(-h < t < h\). The moment generating function of \(X\) is \[ M_X(t) = E\!\big(\exp(tX)\big), \quad -h < t < h. \]

Uniqueness. Let \(X\) and \(Y\) be random variables with MGFs \(M_X\) and \(M_Y\) existing in open intervals about \(0\). Then \(F_X(z) = F_Y(z)\) for all \(z \in \mathbb{R}\) if and only if \(M_X(t) = M_Y(t)\) for all \(t \in (-h, h)\) for some \(h > 0\).

Moments from the MGF. Let \(X\) be a random variable with MGF \(M_X(t)\). If \(m\) is a positive integer, then \(M_X^{(m)}(0) = E(X^m)\).