Aditya Makkar
A Note On Conditional Probability

Introduction

The concept of conditional probability is central to probability theory and excellent treatment of it can be found in many books. My aim with this blog post is to consolidate in one place some ideas around it which helped me form a better intuition. These ideas will be useful if you have already been exposed to this concept from a textbook and just want one more person's ramblings about it.

I will start by defining conditional expectation and stating some of its properties. It will be a grave injustice to claim my discussion of it is complete since I don't even prove its existence; this section exists solely for establishing notation. I will then spend some time discussing conditional probability, relating it to the traditional notion of

P(A∣B)=P(A∩B)P(B).\begin{aligned} \mathbb{P}(A \mid B) = \frac{\mathbb{P}(A \cap B)}{\mathbb{P}(B)}.\end{aligned}

These discussions will naturally lead to the notions of regular conditional probability and regular conditional distribution which I discuss next.

Conditional Expectation

Recall the concept of conditional expectation.

Theorem 1: Let (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) be a probability space, and XX a random variable with E(∣X∣)<∞.\mathbb{E}(|X|) < \infty. Let G\mathcal{G} be a sub-σ\sigma-algebra of F.\mathcal{F}. Then there exists a random variable YY such that

  1. YY is G\mathcal{G}-measurable,

  2. E(∣Y∣)<∞\mathbb{E}(|Y|) < \infty, and

  3. ∫GY dP=∫GX dP\int_G Y \,\mathrm{d}\mathbb{P} = \int_G X \,\mathrm{d}\mathbb{P} for every G∈G.G \in \mathcal{G}.

Remarks:

  1. It is easy to see from the π−λ\pi-\lambda theorem that the last condition can be relaxed such that ∫GY dP=∫GX dP\int_G Y \,\mathrm{d}\mathbb{P} = \int_G X \,\mathrm{d}\mathbb{P} for every GG in some π\pi-system which contains Ω\Omega and generates G.\mathcal{G}.

  2. If Y′Y' is another random variable with the three properties above then Y′=YY' = Y a.s.. Therefore, YY in the theorem above is called a version of the conditional expectation. The notation E ⁣(X ∣ G)\mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right) is used to denote this unique (up to a.e. equivalence) random variable.

The proof of this standard theorem can be found in any probability textbook; see (Williams, 1991) or (Kallenberg, 2021) for example.

Definition 1: In the setting of Theorem 1, if ZZ is a random variable, we write E ⁣(X ∣ Z)\mathbb{E}\!\left(\left. X\,\right\vert\, Z \right) for E ⁣(X ∣ σ(Z)).\mathbb{E}\!\left(\left. X\,\right\vert\, \sigma(Z) \right).

The fact that conditional expectation is defined as a random variable might come as a surprise, but the correspondence with the traditional usage of conditional expectation as a number becomes clear once you realize that here we are conditioning on a σ\sigma-algebra instead of a single event (we haven’t defined what conditioning on an event means, but think of the intuitive meaning for now). For example, consider the life expectancy of a new born baby conditioned on sex. This is a random variable that takes one value for males and another value for females.

Properties of conditional expectation

For completeness I state some useful properties of conditional expectation. You can find the proofs in [1,2] for example. Most of them are parallels to the well-known properties of (unconditional) expectation. Assume that all the XX's satisfy E(∣X∣)<∞\mathbb{E}(|X|) < \infty and let G,H\mathcal{G}, \mathcal{H} be sub-σ\sigma-algebras of F.\mathcal{F}.

  1. [Linearity] E ⁣(a1X1+a2X2 ∣ G)=a1E ⁣(X1 ∣ G)+a2E ⁣(X2 ∣ G)\mathbb{E}\!\left(\left. a_1 X_1 + a_2 X_2\,\right\vert\, \mathcal{G}\right) = a_1 \mathbb{E}\!\left(\left. X_1\,\right\vert\, \mathcal{G}\right) + a_2 \mathbb{E}\!\left(\left. X_2\,\right\vert\, \mathcal{G}\right) a.s. for real numbers a1a_1 and a2.a_2.

  2. [Positivity] If X≥0X \ge 0, then E ⁣(X ∣ G)≥0\mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right) \ge 0 a.s..

  3. [Monotone convergence theorem for conditional expectation] If E(∣Y∣)<∞\mathbb{E}(|Y|) < \infty and Y≤Xn↑XY \le X_n \uparrow X a.s., then E ⁣(Xn ∣ G)↑E ⁣(X ∣ G)\mathbb{E}\!\left(\left. X_n\,\right\vert\, \mathcal{G}\right) \uparrow \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right) a.s..

  4. [Fatou's lemma for conditional expectation] If E(∣Y∣)<∞\mathbb{E}(|Y|) < \infty and Y≤XnY \le X_n for all n≥1n \ge 1 a.s., then E ⁣(lim inf⁡n→∞Xn ∣ G)≤lim inf⁡n→∞E ⁣(Xn ∣ G)\mathbb{E}\!\left(\left. \liminf_{n \to \infty} X_n\,\right\vert\, \mathcal{G}\right) \le \liminf_{n \to \infty}\mathbb{E}\!\left(\left. X_n\,\right\vert\, \mathcal{G}\right) a.s..

  5. [Dominated convergence theorem for conditional expectation] If ∣Xn∣≤∣Y∣|X_n| \le |Y| for all n≥1n \ge 1, E(∣Y∣)<∞\mathbb{E}(|Y|) < \infty, and Xn→XX_n \to X a.s., then E ⁣(Xn ∣ G)→E ⁣(X ∣ G)\mathbb{E}\!\left(\left. X_n\,\right\vert\, \mathcal{G}\right) \to \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right) a.s..

  6. [Tower property] If H⊆G\mathcal{H} \subseteq \mathcal{G}, then E ⁣(E ⁣(X ∣ G) ∣ H)=E ⁣(E ⁣(X ∣ H) ∣ G)=E ⁣(X ∣ H)\mathbb{E}\!\left(\left. \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right)\,\right\vert\, \mathcal{H}\right) = \mathbb{E}\!\left(\left. \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{H}\right)\,\right\vert\, \mathcal{G}\right)=\mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{H}\right) a.s.

  7. [Taking out what's known] If YY is G\mathcal{G}-measurable and bounded, then E ⁣(YX ∣ G)=YE ⁣(X ∣ G) a.s..\begin{aligned} \mathbb{E}\!\left(\left. YX\,\right\vert\, \mathcal{G}\right) = Y \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right) \text{ a.s..}\end{aligned} If p>1p > 1, 1/p+1/q=11/p + 1/q = 1, X∈Lp(Ω,F,P)X \in L^p(\Omega, \mathcal{F}, \mathbb{P}) and Y∈Lq(Ω,G,P)Y \in L^q(\Omega, \mathcal{G}, \mathbb{P}), then (1)(1) again holds. If XX is a nonnegative F\mathcal{F}-measurable random variable, YY is a nonnegative G\mathcal{G}-measurable random variable, E(X)<∞\mathbb{E}(X) < \infty and E(XY)<∞\mathbb{E}(XY) < \infty, then also (1) holds.

  8. [Role of independence] If H\mathcal{H} is independent of σ(σ(X)∪G)\sigma(\sigma(X) \cup \mathcal{G}), then E ⁣(X ∣ σ(G∪H))=E ⁣(X ∣ G)\mathbb{E}\!\left(\left. X\,\right\vert\, \sigma(\mathcal{G} \cup \mathcal H)\right) = \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right) a.s.. In particular, if XX is independent of H\mathcal{H}, then E ⁣(X ∣ H)=E(X)\mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{H}\right) = \mathbb{E}(X) a.s..

Conditional Probability

Definition 2: In the setting of Theorem 1, if A∈FA \in \mathcal{F}, we let P ⁣(A ∣ G)\mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) to mean E ⁣(1A ∣ G)\mathbb{E}\!\left(\left. \mathbf{1}_A\,\right\vert\, \mathcal{G}\right) and call it the conditional probability of AA given G\mathcal{G}. Here 1A\mathbf{1}_A is the indicator random variable. If B∈FB \in \mathcal{F}, we let P ⁣(A ∣ B)\mathbb{P}\!\left(\left. A\,\right\vert\, B\right) to mean E ⁣(1A ∣ 1B).\mathbb{E}\!\left(\left. \mathbf{1}_A\,\right\vert\, \mathbf{1}_B\right).

Just like conditional expectation, conditional probability, as defined above, is a random variable! Unlike conditional expectation this isn't very palpable and deserves more rumination (Halmos, 1950). We have our probability space (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) and let A,B∈FA,B \in \mathcal{F} be such that P(B)≠0\mathbb{P}(B) \neq 0 and P(Bc)≠0.\mathbb{P}(B^\mathsf{c}) \neq 0. Then our traditional notion of conditional probability tells us that the conditional probability of AA given BB is defined by

PB(A)=P(A∩B)P(B).\begin{aligned} \mathbb{P}_B(A) = \frac{\mathbb{P}(A \cap B)}{\mathbb{P}(B)}.\end{aligned}

Let us investigate how PB(A)\mathbb{P}_B(A) depends on B.B. To this end, introduce the discrete measurable space (Λ,2Λ)(\Lambda, 2^\Lambda) with Λ={λ1,λ2}\Lambda = \{\lambda_1, \lambda_2\}, and a measurable mapping T ⁣:Ω→ΛT \colon \Omega \to \Lambda such that

T(ω)={λ1 if ω∈Bλ2 if ω∈Bc.\begin{aligned} T(\omega) = \begin{cases} \lambda_1 & \text{ if } \omega \in B \\ \lambda_2 & \text{ if } \omega \in B^\mathsf{c}. \end{cases}\end{aligned}

Define the two measures νA\nu_A and ν\nu on (Λ,2Λ)(\Lambda, 2^\Lambda) as follows for any E⊆ΛE \subseteq \Lambda,

νA(E)=P(A∩T−1(E))ν(E)=P(T−1(E)).\begin{aligned} \nu_A(E) &= \mathbb{P}(A \cap T^{-1}(E))\\ \nu(E) &= \mathbb{P}(T^{-1}(E)).\end{aligned}

Then it is easy to see that

PB(A)=νA({λ1})ν({λ1})PBc(A)=νA({λ2})ν({λ2}).\begin{aligned} \mathbb{P}_B(A) &= \frac{\nu_A(\{\lambda_1\})}{\nu(\{\lambda_1\})} \\ \mathbb{P}_{B^\mathsf{c}}(A) &= \frac{\nu_A(\{\lambda_2\})}{\nu(\{\lambda_2\})}.\end{aligned}

In other words conditional probability may be viewed as a measurable function on Λ.\Lambda.

This can easily be generalized to any finite setting as follows. Let {A1,…,An}⊆F\{A_1, \ldots, A_n\} \subseteq \mathcal{F} be a partition of Ω\Omega, i.e., Ai∩Aj=∅A_i \cap A_j = \varnothing for i≠ji \neq j and ⋃iAi=Ω.\bigcup_i A_i = \Omega. Introduce the discrete measurable space (Λ,2Λ)(\Lambda, 2^\Lambda) with Λ={λ1,…,λn}.\Lambda = \{\lambda_1, \ldots, \lambda_n\}. Define a measurable mapping T ⁣:Ω→ΛT \colon \Omega \to \Lambda such that T(ω)=λiT(\omega) = \lambda_i whenever ω∈Ai.\omega \in A_i. Define the measures νA1,…,νAn,ν\nu_{A_1}, \ldots, \nu_{A_n}, \nu on (Λ,2Λ)(\Lambda, 2^\Lambda) as follows for any E⊆ΛE \subseteq \Lambda,

νAi(E)=P(Ai∩T−1(E))for all i=1,…,nν(E)=P(T−1(E)).\begin{aligned} \nu_{A_i}(E) &= \mathbb{P}(A_i \cap T^{-1}(E))\quad \text{for all } i = 1, \ldots, n\\ \nu(E) &= \mathbb{P}(T^{-1}(E)).\end{aligned}
Then once again we have for any A∈FA \in \mathcal{F},
PAi(A)=P(A∩Ai)P(Ai)=νAi({λi})ν({λi})for all i=1,…,n.\begin{aligned} \mathbb{P}_{A_i}(A) = \frac{\mathbb{P}(A \cap A_i)}{\mathbb{P}(A_i)} = \frac{\nu_{A_i}(\{\lambda_i\})}{\nu(\{\lambda_i\})} \quad \text{for all } i = 1, \ldots, n.\end{aligned}

These considerations are what motivated the definition of conditional probability in general cases, as you see in Definition 2. If TT is any measurable mapping from (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) into an arbitrary measurable space (Λ,L)(\Lambda, \mathcal{L}), and if we write νA(E)=P(A∩T−1(E))\nu_A(E) = \mathbb{P}(A \cap T^{-1}(E)) where A∈FA \in \mathcal{F} and E∈LE \in \mathcal{L}, then it is clear that νE\nu_E and P∘T−1\mathbb{P} \circ T^{-1} are measures on L\mathcal{L} such that νA≪P∘T−1.\nu_A \ll \mathbb{P} \circ T^{-1}. Radon-Nikodym theorem now implies that there exists an P∘T−1\mathbb{P} \circ T^{-1}-integrable function pAp_A, unique upto P∘T−1\mathbb{P} \circ T^{-1}-a.e., such that

P(A∩T−1(E))=∫EpA(λ) (P∘T−1)(dλ)for all E∈L.\begin{aligned} \mathbb{P}(A \cap T^{-1}(E)) = \int_E p_A(\lambda) \; (\mathbb{P} \circ T^{-1})(\mathrm{d} \lambda) \quad \text{for all } E \in \mathcal{L}.\end{aligned}

We anoint pA(λ)p_A(\lambda) as the conditional probability of AA given λ∈Λ\lambda \in \Lambda or the conditional probability of AA given that T(ω)=λ.T(\omega) = \lambda. Note that here we are conditioning on a measurable mapping TT instead of a sub-σ\sigma-algebra, but this notion is related to conditioning on σ(T)\sigma(T) as will become clear ahead. Keep this "rumination" in mind when we discuss regular conditional distribution later.

Let's look at our definition of conditional probability from the other direction and show that P(A∣B)\mathbb{P}(A \mid B) as defined in Definition 2 conforms to our traditional usage. To start, note that σ(1B)={∅,B,Bc,Ω}\sigma(\mathbf{1}_B) =\{\varnothing, B, B^\mathsf{c}, \Omega\}, and since P(A∣B)\mathbb{P}(A \mid B) is σ(1B)\sigma(\mathbf{1}_B)-measurable, it must be constant on each of the sets B,BcB, B^\mathsf{c}, thereby necessitating

P(A∣B)(ω)={P(A∩B)P(B) if ω∈BP(A∩Bc)P(Bc) if ω∈Bc\begin{aligned} \mathbb{P}(A \mid B)(\omega) = \begin{cases} \frac{\mathbb{P}(A \cap B)}{\mathbb{P}(B)} & \text{ if } \omega \in B \\ \frac{\mathbb{P}(A \cap B^\mathsf{c})}{\mathbb{P}(B^\mathsf{c})} & \text{ if } \omega \in B^\mathsf{c} \end{cases}\end{aligned}
because of property 3. of Theorem 1 by taking GG to be BB and Bc.B^\mathsf{c}. Of course, if any of the sets BB or BcB^\mathsf{c} is of measure 00, then you can take the corresponding value for P(A∣B)(ω)\mathbb{P}(A \mid B)(\omega) to be anything in [0,1][0,1] and it won't matter since the concept of conditional probability is defined only up to sets of measure 0.0.

Properties of conditional probability

Positivity property (property 2. above) and monotone convergence property (property 3. above) of conditional expectation imply 0≤P ⁣(A ∣ G)≤10 \le\mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) \le 1 a.s. for any A∈FA \in \mathcal{F}, P ⁣(A ∣ G)=0\mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) = 0 a.s. if and only if P(A)=0\mathbb{P}(A) = 0, and P ⁣(A ∣ G)=1\mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) = 1 a.s. if and only if P(A)=1.\mathbb{P}(A) = 1.

Let A1,A2,…∈FA_1, A_2, \ldots \in \mathcal{F} be a sequence of disjoint sets. By linearity (property 1. above) and monotone convergence theorem for conditional expectation (property 3. above), we see that

P ⁣(⋃nAn ∣ G)=∑nP ⁣(An ∣ G)a.s..\begin{aligned} \mathbb{P}\!\left(\left. \bigcup_n A_n\,\right\vert\, \mathcal{G}\right) = \sum_n \mathbb{P}\!\left(\left. A_n\,\right\vert\, \mathcal{G}\right) \quad \text{a.s.}.\end{aligned}

If An∈FA_n \in \mathcal{F} for n≥1n \ge 1 and lim⁡n→∞An=A\lim_{n \to \infty} A_n = A, then we also have lim⁡n→∞P ⁣(An ∣ G)=P ⁣(A ∣ G)\lim_{n \to \infty} \mathbb{P}\!\left(\left. A_n\,\right\vert\, \mathcal{G}\right) = \mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) a.s..

It seems very tempting from the foregoing discussion to claim that P ⁣(⋅ ∣ G)(ω)\mathbb{P}\!\left(\left. \cdot\,\right\vert\, \mathcal{G}\right)(\omega) is a probability measure on F\mathcal{F} for almost all ω∈Ω\omega \in \Omega, but except for some nice spaces, which we will discuss below, this isn't true. Let us first try to see this intuitively (Chow and Teicher, 1997). Equation (2) holds for all ω∈Ω\omega \in \Omega EXCEPT for some null set which may well depend on the particular sequence {An}n∈N.\{A_n\}_{n \in \mathbb N}. It does NOT stipulate that there exists a fixed null set N∈FN \in \mathcal{F} such that

P ⁣(⋃nAn ∣ G)(ω)=∑nP ⁣(An ∣ G)(ω),ω∈Ω∖N\begin{aligned} \mathbb{P}\!\left(\left. \bigcup_n A_n\,\right\vert\, \mathcal{G}\right)(\omega) = \sum_n \mathbb{P}\!\left(\left. A_n\,\right\vert\, \mathcal{G}\right)(\omega), \quad \omega \in \Omega \setminus N\end{aligned}
for every disjoint sequence {An}n∈N⊆F.\{A_n\}_{n \in \mathbb N} \subseteq \mathcal{F}. Except in trivial cases, there are uncountably many disjoint sequences, and therefore we will need uncountable union of such null sets to be of measure 00, which of course may not even be defined let alone be of measure 0.0. To further drive this point home you can take a look at an explicit example of how this can fail in an exercise in (Halmos, 1950) on page 210.

Regular Conditional Probability

Motivated by our discussion above, we define regular conditional probability as follows.

Definition 3: Let (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) be a probability space and G,H\mathcal{G}, \mathcal{H} be sub-σ\sigma-algebras of F.\mathcal{F}. A regular conditional probability on H\mathcal{H} given G\mathcal{G} is a function P(⋅,⋅) ⁣:Ω×H→[0,1]\mathbb{P}(\cdot, \cdot) \colon \Omega \times \mathcal{H} \to [0,1] such that

  1. for a.e. ω∈Ω\omega \in \Omega, P(ω,⋅)\mathbb{P}(\omega, \cdot) is a probability measure on H\mathcal{H},

  2. for each A∈HA \in \mathcal{H}, P(⋅,A)\mathbb{P}(\cdot, A) is a G\mathcal{G}-measurable function on Ω\Omega coinciding with the conditional probability of AA given G\mathcal{G}, i.e., P(⋅,A)=P ⁣(A ∣ G)\mathbb{P}(\cdot, A) = \mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) a.s..

Let's show that this definition is not outrageous by showing that it agrees with our traditional notion of conditional pdf and conditional expectation. To that end, suppose that XX and YY are random variables which have a joint probability density function fX,Y(x,y).f_{X,Y}(x,y). This means that we are considering the probability space (R2,B(R2),P)(\mathbb R^2, \mathcal{B}(\mathbb R^2), \mathbb{P}) with XX and YY being the coordinate random variables, i.e. (x,y)↦x(x,y) \mapsto x and (x,y)↦y(x,y) \mapsto y respectively, and having an absolutely continuous distribution function FX,Y(x,y)F_{X,Y}(x,y) such that

FX,Y(x,y)=∫−∞y∫−∞xfX,Y(s,t) ds dt.\begin{aligned} F_{X,Y}(x,y) = \int_{-\infty}^y \int_{-\infty}^x f_{X,Y}(s,t) \; \mathrm{d} s \, \mathrm{d} t.\end{aligned}
We recall that fX(x)=∫RfX,Y(x,y) dyf_X(x) = \int_\mathbb R f_{X,Y}(x,y) \, \mathrm{d} y and fY(y)=∫RfX,Y(x,y) dxf_Y(y) = \int_\mathbb R f_{X,Y}(x,y) \, \mathrm{d} x act as probability density functions for XX and YY respectively, and
fX∣Y(x∣y)={fX,Y(x,y)fY(y) if fY(y)≠00 otherwise\begin{aligned} f_{X \mid Y}(x \mid y) = \begin{cases} \frac{f_{X,Y}(x,y)}{f_Y(y)} & \text{ if } f_Y(y) \neq 0 \\ 0 & \text{ otherwise} \end{cases}\end{aligned}
defines the elementary conditional pdf fX∣Yf_{X \mid Y} of XX given Y.Y. By Fubini's theorem fXf_X and fYf_Y are Borel functions on R\mathbb R and so fX∣Yf_{X \mid Y} is a Borel function on R2.\mathbb R^2. Let H=B(R2)=σ(X,Y)\mathcal{H} = \mathcal{B}(\mathbb R^2) = \sigma(X, Y) and G=σ(Y)=R×B(R).\mathcal{G} = \sigma(Y) = \mathbb R \times \mathcal{B}(\mathbb R). For A∈HA \in \mathcal{H} and ω=(x,y)∈R2\omega = (x,y) \in \mathbb R^2 we define
P(ω,A)=∫{s : (s,y)∈A}fX∣Y(s∣y) ds.\begin{aligned} \mathbb{P}(\omega, A) = \int_{\{s\,:\,(s,y) \in A\}} f_{X \mid Y}(s \mid y) \, \mathrm{d} s.\end{aligned}
Then for each ω∈R2\omega \in \mathbb R^2, P(ω,⋅)\mathbb{P}(\omega, \cdot) is a probability measure on H\mathcal{H}, and for each A∈HA \in \mathcal{H}, P(⋅,A)\mathbb{P}(\cdot, A) is a Borel function in yy and hence G\mathcal{G}-measurable. To verify that P(⋅,A)=P ⁣(A ∣ G)\mathbb{P}(\cdot, A) = \mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) a.s. for any A∈HA \in \mathcal{H} we just need to verify property 3. of Theorem 1. To this end, fix A∈HA \in \mathcal{H} and G∈GG \in \mathcal{G}, and note that GG must be of the form G=R×BG = \mathbb R \times B for B∈B(R).B \in \mathcal{B}(\mathbb R). Thus
∫GP(ω,A) dP(ω)=∫B∫RP((s,t),A)fX∣Y(s,t) ds dt(by absolute continuity and Fubini’s theorem)=∫B∫R[∫{u : (u,t)∈A}fX∣Y(u∣t) du]fX∣Y(s,t) ds dt=∫B[∫{u : (u,t)∈A}fX∣Y(u∣t) du]fY(t) dt=∫B∫{u : (u,t)∈A}fX,Y(u,t) du dt=∫B∫R1A(u,t)fX,Y(u,t) du dt=∫G1A(ω) dP(ω)\begin{aligned} \int_G \mathbb{P}(\omega, A) \,\mathrm{d}\mathbb{P}(\omega) &= \int_B \int_{\mathbb R} \mathbb{P}((s,t), A) f_{X \mid Y}(s,t) \, \mathrm{d} s \, \mathrm{d} t \quad \text{(by absolute continuity and Fubini's theorem)} \\ &= \int_B \int_{\mathbb R} \left[ \int_{\{u\,:\,(u,t) \in A\}} f_{X \mid Y}(u \mid t) \, \mathrm{d} u \right] f_{X \mid Y}(s,t) \, \mathrm{d} s \, \mathrm{d} t \\ &= \int_B \left[ \int_{\{u\,:\,(u,t) \in A\}} f_{X \mid Y}(u \mid t) \, \mathrm{d} u \right] f_{Y}(t) \, \mathrm{d} t \\ &= \int_B \int_{\{u\,:\,(u,t) \in A\}} f_{X, Y}(u, t) \, \mathrm{d} u \, \mathrm{d} t \\ &= \int_B \int_{\mathbb R} \mathbf{1}_{A}(u,t) f_{X, Y}(u, t) \, \mathrm{d} u \, \mathrm{d} t \\ &= \int_{G} \mathbf{1}_{A}(\omega) \, \mathrm{d}\mathbb{P}(\omega)\end{aligned}

and so P(⋅,A)=E ⁣(1A ∣ G)=P ⁣(A ∣ G)\mathbb{P}(\cdot, A) = \mathbb{E}\!\left(\left. \mathbf{1}_A\,\right\vert\, \mathcal{G}\right)= \mathbb{P}\!\left(\left. A\,\right\vert\, \mathcal{G}\right) a.s.. Hence, P(ω,A)\mathbb{P}(\omega, A) is a regular conditional probability on H\mathcal{H} given G.\mathcal{G}.

For the corresponding analysis for conditional expectation, let hh be a Borel function on R2\mathbb R^2 such that

E∣h(X,Y)∣=∫R∫R∣h(x,y)∣fX,Y(x,y) dx dy<∞.\begin{aligned} \mathbb{E}{|h(X,Y)|} = \int_\mathbb R \int_\mathbb R |h(x,y)| f_{X,Y}(x,y) \, \mathrm{d} x \, \mathrm{d} y < \infty.\end{aligned}
Set
g(y)=∫Rh(s,y)fX∣Y(s∣y) ds.\begin{aligned} g(y) = \int_\mathbb R h(s, y) f_{X \mid Y}(s \mid y) \, \mathrm{d} s.\end{aligned}
g(y)g(y) is the traditional conditional density of h(X,Y)h(X, Y) given Y=y.Y = y. Then the claim is that
g(Y)=E ⁣(h(X,Y) ∣ σ(Y))a.s..\begin{aligned} g(Y) = \mathbb{E}\!\left(\left. h(X,Y)\,\right\vert\, \sigma(Y)\right) \text{a.s.}.\end{aligned}
A typical element of σ(Y)\sigma(Y) has the form {ω∈R2 : Y(ω)∈B}\{\omega \in \mathbb R^2 \, : \, Y(\omega) \in B\}, where B∈B(R).B \in \mathcal{B}(\mathbb R). Hence, we must show that
L=E[h(X,Y)1B(Y)]=E[g(Y)1B(Y)]=R.\begin{aligned} L = \mathbb{E}\left[h(X,Y) \mathbf{1}_{B}(Y)\right] = \mathbb{E}\left[g(Y) \mathbf{1}_{B}(Y)\right] = R.\end{aligned}
But we can write LL and RR as
L=∫∫h(x,y)1B(y)fX,Y(x,y) dx dyR=∫g(y)1B(y)fY(y) dy\begin{aligned} L &= \int \int h(x,y) \mathbf{1}_{B}(y) f_{X,Y}(x,y) \, \mathrm{d} x \, \mathrm{d} y \\ R &= \int g(y) \mathbf{1}_{B}(y) f_Y(y) \, \mathrm{d} y\end{aligned}
and they are equal by Fubini's theorem.

In general, we have the following useful theorem (taken from (Chow and Teicher, 1997)) which allows us to view conditional expectations as ordinary expectations relative to the measure induced by regular conditional probability.

Theorem 2: Consider the setting of Definition 3 and denote Pω(⋅)=P(ω,⋅).\mathbb{P}_\omega(\cdot) = \mathbb{P}(\omega, \cdot). Let XX be an H\mathcal{H}-measurable function with E(X)<∞.\mathbb{E}(X) < \infty. Then E ⁣(X ∣ G)(ω)=∫ΩX dPωa.s..\begin{aligned} \mathbb{E}\!\left(\left. X\,\right\vert\, \mathcal{G}\right)(\omega) = \int_\Omega X \, \mathrm{d}\mathbb{P}_\omega \quad \text{a.s.}.\end{aligned}

Recall the monotone class theorem for functions:

Monotone Class Theorem for Functions: Let H\mathscr{H} be a family of nonnegative functions on Ω\Omega which contains all indicators of sets of some class H\mathcal{H} of subsets of Ω.\Omega. If either (i) H\mathcal{H} is a π\pi-class and H\mathscr{H} is a λ\lambda-system, or (ii) H\mathcal{H} is a σ\sigma-algebra and H\mathscr{H} is a monotone system, then H\mathscr{H} contains all nonnegative σ(H)\sigma(\mathcal{H})-measurable functions.

By separate considerations of X+X^+ and X−X^-, it may be supposed that X≥0.X \ge 0. Define

H={X : X≥0, X is H-measurable, and (3) holds for X}.\begin{aligned} \mathscr{H} = \{X \, : \, X \ge 0,\, X \text{ is } \mathcal{H} \text{-measurable, and (3) holds for } X\}.\end{aligned}
By the definition of regular conditional probability, 1A∈H\mathbf{1}_A \in \mathscr{H} for A∈H.A \in \mathcal{H}. H\mathcal{H} is already a σ\sigma-algebra. Let's show that H\mathscr{H} is a monotone system.

If X1,X2∈HX_1, X_2 \in \mathscr{H} and c1,c2≥0c_1, c_2 \ge 0, then c1X1+c2X2≥0c_1 X_1 + c_2 X_2 \ge 0, c1X1+c2X2c_1 X_1 + c_2 X_2 is H\mathcal{H}-measurable and Equation (3) holds because of linearity of expectation and conditional expectation, and thus c1X1+c2X2∈H.c_1 X_1 + c_2 X_2 \in \mathscr{H}. If {Xn}n∈N⊆H\{X_n\}_{n \in \mathbb N} \subseteq \mathscr{H} such that Xn↑XX_n \uparrow X, then X≥0X \ge 0, XX is H\mathcal{H}-measurable, and Equation (3) holds for XX because of monotone convergence theorem for expectation and conditional expectation, and thus X∈H.X \in \mathscr{H}. Therefore, by the monotone class theorem H\mathscr{H} contains all nonnegative H\mathcal{H}-measurable functions.

Regular Conditional Distribution

In some cases even the concept of regular conditional probability in inadequate, and that motivates the concept of regular conditional distributions.

Definition 4: Let (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) be a probability space, G⊆F\mathcal{G} \subseteq \mathcal{F} a σ\sigma-algebra, (Λ,L)(\Lambda, \mathcal{L}) a measurable space, and T ⁣:Ω→ΛT \colon \Omega \to \Lambda a measurable mapping. A regular conditional distribution for TT given G\mathcal{G} is a function PT ⁣:Ω×L→[0,1]\mathbb{P}_T \colon \Omega \times \mathcal{L} \to [0,1] such that

  1. for a.e. ω∈Ω\omega \in \Omega, PT(ω,⋅)\mathbb{P}_T(\omega, \cdot) is a probability measure on L\mathcal{L},

  2. for each A∈LA \in \mathcal{L}, PT(⋅,A)\mathbb{P}_T(\cdot, A) is a G\mathcal{G}-measurable function on Ω\Omega coinciding with the conditional probability of T−1(A)T^{-1}(A) given G\mathcal{G}, i.e., PT(⋅,A)=P ⁣(T−1(A) ∣ G)\mathbb{P}_T(\cdot, A) = \mathbb{P}\!\left(\left. T^{-1}(A)\,\right\vert\, \mathcal{G}\right) a.s..

It is clear that when Λ=Ω\Lambda = \Omega, L=H⊆F\mathcal{L} = \mathcal{H} \subseteq \mathcal{F} and TT is the identity map, PT\mathbb{P}_T is exactly the regular conditional probability as defined in Definition 3.

Now would be a good time to reread the first “rumination” in the section Conditional Probability and realize that the definition of regular conditional distribution is in fact well motivated.

A corresponding version of Theorem 2 exists, proof of which I'll leave as an easy exercise:

Theorem 3: In the setting of Definition 4, if PTω(A)=PT(ω,A)\mathbb{P}_T^\omega(A) = \mathbb{P}_T(\omega, A) and h ⁣:Λ→Rh \colon \Lambda \to \mathbb R is a Borel function with E(∣h(T)∣)<∞\mathbb{E}(|h(T)|) < \infty, then
E ⁣(h(T) ∣ G)(ω)=∫Λh(λ) PTω(dλ).\begin{aligned} \mathbb{E}\!\left(\left. h(T)\,\right\vert\, \mathcal{G}\right)(\omega) = \int_\Lambda h(\lambda) \, \mathbb{P}_T^\omega(\mathrm{d}\lambda).\end{aligned}

To see the power of thinking about conditional probabilities like this, let's give an unbelievably short proof of conditional Hölder's inequality that I took from (Chow and Teicher, 1997). Contrast it with other proofs.

Theorem 4: If X,YX,Y are random variables on (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}), G⊆F\mathcal{G} \subseteq \mathcal{F} is a σ\sigma-algebra, and 1<p<∞1 < p < \infty, 1/p+1/q=11/p + 1/q = 1, then
E ⁣(∣XY∣ ∣ G)≤[E ⁣(∣X∣p ∣ G)]1/p[E ⁣(∣Y∣q ∣ G)]1/q a.s..\begin{aligned} \mathbb{E}\!\left(\left. |XY|\,\right\vert\, \mathcal{G}\right) \le \left[ \mathbb{E}\!\left(\left. |X|^p\,\right\vert\, \mathcal{G}\right)\right]^{1/p}\left[ \mathbb{E}\!\left(\left. |Y|^q\,\right\vert\, \mathcal{G}\right)\right]^{1/q} \text{ a.s.}.\end{aligned}
For B∈B(R2)B \in \mathcal{B}(\mathbb R^2) and ω∈Ω\omega \in \Omega, let PX,Yω(B)=PX,Y(ω,B)\mathbb{P}_{X,Y}^\omega(B) = \mathbb{P}_{X,Y}(\omega, B) be the regular conditional distribution for (X,Y)(X,Y) given G.\mathcal{G}. Theorem 3 allows us to write
E ⁣(∣XY∣ ∣ G)(ω)=∫R2∣xy∣ PX,Yω(d(x,y))[E ⁣(∣X∣p ∣ G)]1/p(ω)=[∫R2∣x∣p PX,Yω(d(x,y))]1/p[E ⁣(∣Y∣q ∣ G)]1/q(ω)=[∫R2∣y∣q PX,Yω(d(x,y))]1/q.\begin{aligned} \mathbb{E}\!\left(\left. |XY|\,\right\vert\, \mathcal{G}\right)(\omega) &= \int_{\mathbb R^2} |x y| \,\mathbb{P}_{X,Y}^\omega(\mathrm{d} (x,y))\\ \left[ \mathbb{E}\!\left(\left. |X|^p\,\right\vert\, \mathcal{G}\right)\right]^{1/p}(\omega) &= \left[ \int_{\mathbb R^2} |x|^p \, \mathbb{P}_{X,Y}^\omega(\mathrm{d} (x,y)) \right]^{1/p} \\ \left[ \mathbb{E}\!\left(\left. |Y|^q\,\right\vert\, \mathcal{G}\right)\right]^{1/q}(\omega) &= \left[ \int_{\mathbb R^2} |y|^q \, \mathbb{P}_{X,Y}^\omega(\mathrm{d} (x,y)) \right]^{1/q}.\end{aligned}
And now our desired inequality follows immediately from the ordinary Hölder's inequality.

Existence of Regular Conditional Distribution

Before we discuss their existence, let us define the concept of standard Borel space (Encyclopedia of Mathematics).

Definition 5: Let (X,X)(X, \mathcal{X}) and (Y,Y)(Y, \mathcal{Y}) be two measurable spaces. They are called isomorphic if there exists a bijection f ⁣:X→Yf \colon X \to Y such that ff and its inverse f−1f^{-1} are both measurable. The function ff is called an isomorphism.

Definition 6: A measurable space (X,X)(X, \mathcal{X}) is called a standard Borel space if it satisfies any of the following equivalent conditions:

  1. (X,X)(X, \mathcal{X}) is isomorphic to some compact metric space with the Borel σ\sigma-algebra.

  2. (X,X)(X, \mathcal{X}) is isomorphic to some Polish space (i.e., a separable complete metric space) with the Borel σ\sigma-algebra.

  3. (X,X)(X, \mathcal{X}) is isomorphic to some Borel subset of some Polish space with the Borel σ\sigma-algebra.

As you can guess most spaces we deal with are standard Borel spaces. (Durrett, 2019) calls these space nice since we already have too many things named after Borel. I am not sure I agree with his reasoning but I like Durrett's terminology.

The next two theorems show the existence of regular conditional distribution and are taken from [(Durrett, 2019), Section 4.1.3]. See also [(Parthasarathy, 1967), Section V.8].

Theorem 5: Regular conditional distribution exists if (Λ,L)(\Lambda, \mathcal{L}) is nice.

A generalization of the last theorem:

Theorem 6: Suppose (Λ,L)(\Lambda, \mathcal{L}) is a nice space, TT and SS are measurable mappings from Ω\Omega to Λ\Lambda, and G=σ(S).\mathcal{G} = \sigma(S). Then there exists a function μ ⁣:Λ×L→[0,1]\mu \colon \Lambda \times \mathcal{L} \to [0,1] such that

  1. for a.e. ω∈Ω\omega \in \Omega, μ(S(ω),⋅)\mu(S(\omega), \cdot) is a probability measure on L\mathcal{L}, and

  2. for each A∈LA \in \mathcal{L}, μ(S(⋅),A)=P(T−1(A)∣G)\mu(S(\cdot), A) = \mathbb{P}(T^{-1}(A)\mid\mathcal{G}) a.s..

It is instructive to prove Theorem 5 in the special case when (Λ,L)=(Rn,B(Rn)).(\Lambda, \mathcal{L}) = (\mathbb R^n, \mathcal{B}(\mathbb R^n)). The theorem and the proof is taken from (Chow and Teicher, 1997).

Theorem 7: In the setting of Definition 4, let (Λ,L)=(Rn,B(Rn))(\Lambda, \mathcal{L}) = (\mathbb R^n, \mathcal{B}(\mathbb R^n)) and T=(T1,…,Tn) ⁣:Ω→Rn.T = (T_1, \ldots, T_n) \colon \Omega \to \mathbb R^n. Then there exists a regular conditional distribution for TT given G.\mathcal{G}.

Let's recall the definition of an nn-dimensional distribution function on Rn\mathbb R^n (the Russian convention of left-continuous distribution function).

An nn-dimensional distribution function on Rn\mathbb R^n is a function F ⁣:Rn→[0,1]F \colon \mathbb R^n \to [0,1] satisfying:
lim⁡xj→−∞F(x1,…,xn)=0,1≤j≤n;lim⁡xj→∞1≤j≤nF(x1,…,xn)=1;lim⁡yj↑xjF(x1,…,xj−1,yj,xj+1,…,xn)=F(x1,…,xj,…,xn),1≤j≤n; and\begin{aligned} \lim_{x_j \to - \infty} F(x_1, \ldots, x_n) &= 0, \quad 1 \le j \le n;\\ \lim_{\substack{x_j \to \infty \\ 1 \le j \le n}} F(x_1, \ldots, x_n) &= 1;\\ \lim_{y_j \uparrow x_j} F(x_1, \ldots, x_{j-1}, y_j, x_{j+1}, \ldots, x_n) &= F(x_1, \ldots, x_j, \ldots, x_n), \quad 1 \le j \le n; \text{ and}\end{aligned}
F(b1,…,bn)−∑j=1nF(b1,…,bj−1,aj,bj+1,…,bn)+∑1≤j<k≤nF(b1,…,bj−1,aj,bj+1,…,bk−1,ak,bk+1,…,bn)−⋯(−1)nF(a1,…,an)=:Δna,b≥0.\begin{aligned} &F(b_1, \ldots, b_n) - \sum_{j=1}^n F(b_1, \ldots, b_{j-1}, a_j, b_{j+1}, \ldots, b_n)+ \sum_{1 \le j < k \le n} F(b_1, \ldots, b_{j-1}, a_j, b_{j+1}, \ldots, b_{k-1}, a_k, b_{k+1}, \ldots, b_n) - \cdots (-1)^n F(a_1, \ldots, a_n)\\ & =: \Delta_n^{a,b} \ge 0.\end{aligned}

We will try to construct a distribution function on Rn.\mathbb R^n. To this end, for any rational number r1,…,rnr_1, \ldots, r_n and ω∈Ω\omega \in \Omega, define Fnω(r1,…,rn):=P ⁣(⋂i=1n{Ti<ri} ∣ G)(ω).\begin{aligned} F_n^\omega(r_1, \ldots, r_n) := \mathbb{P}\!\left(\left. \bigcap_{i=1}^n \{T_i < r_i\} \,\right\vert\, \mathcal{G}\right)(\omega).\end{aligned}

It's evident that the properties of conditional probability discussed above imply that there is a null set N∈GN \in \mathcal{G} such that for ω∈Ω∖N\omega \in \Omega \setminus N and all rational numbers ri,ri′,qi,mr_i, r_i', q_{i,m} the following holds

Fnω(r1,…,rn)≥Fnω(r1′,…,rn′) if ri>ri′, 1≤i≤n,Fnω(r1,…,rn)=lim⁡qi,m↑ri1≤i≤nFnω(q1,m,…,qn,m),lim⁡ri→−∞Fnω(r1,…,rn)=0,1≤i≤n,lim⁡ri→∞1≤i≤nFnω(r1,…,rn)=1, andΔnr,r′Fnω≥0 if r≤r′,\begin{aligned} F_n^\omega(r_1, \ldots, r_n) &\ge F_n^\omega(r_1', \ldots, r_n') \text{ if } r_i > r_i', \, 1 \le i \le n, \\ F_n^\omega(r_1, \ldots, r_n) &= \lim_{\substack{q_{i,m} \uparrow r_i \\ 1 \le i \le n}} F_n^\omega(q_{1, m}, \ldots, q_{n, m}), \\ \lim_{r_i \to -\infty} F_n^\omega(r_1, \ldots, r_n) &= 0, \quad 1 \le i \le n, \\ \lim_{\substack{r_i \to \infty \\ 1 \le i \le n}} F_n^\omega(r_1, \ldots, r_n) &= 1, \text{ and} \\ \Delta_n^{r, r'} F_n^\omega &\ge 0 \text{ if } r \le r',\end{aligned}
where r≤r′r \le r' means ri≤ri′r_i \le r_i' for all 1≤i≤n.1 \le i \le n. Having defined FnωF_n^\omega for rational values, extend to any real numbers x1,…,xnx_1, \ldots, x_n as follows Fnω(x1,…,xn)={lim⁡ri↑xiri∈Q1≤i≤nFnω(r1,…,rn) if ω∈Ω∖NP(⋂i=1n{Ti<ri}) if ω∈N.\begin{aligned} F_n^\omega(x_1, \ldots, x_n) = \begin{cases} \lim_{\substack{r_i \uparrow x_i \\ r_i \in \mathbb{Q} \\ 1 \le i \le n}} F_n^\omega(r_1, \ldots, r_n) & \text{ if } \omega \in \Omega \setminus N \\ \mathbb{P}\left(\bigcap_{i=1}^n \{T_i < r_i\}\right) & \text{ if } \omega \in N. \end{cases}\end{aligned}

Then for each ω∈Ω\omega \in \Omega, Fnω(x1,…,xn)F_n^\omega(x_1, \ldots, x_n) is an nn-dimensional distribution function and hence determines a Lebesgue-Stieltjes measure μω\mu_\omega on B(Rn)\mathcal{B}(\mathbb R^n) with μω(Rn)=1.\mu_\omega(\mathbb R^n) = 1. For B∈B(Rn)B \in \mathcal{B}(\mathbb R^n) and ω∈Ω\omega \in \Omega define

PT(ω,B)=μω(B).\begin{aligned} \mathbb{P}_T(\omega, B) = \mu_\omega(B).\end{aligned}
If
H={B∈B(Rn) : PT(⋅,B)=P(T−1(B)∣G) a.s.}D={B∈B(Rn) : B=[−∞,r1)×⋯×[−∞,rn), ri∈Q},\begin{aligned} \mathcal{H} &= \{B \in \mathcal{B}(\mathbb R^n)\,:\,\mathbb{P}_T(\cdot, B) = \mathbb{P}(T^{-1}(B)\mid\mathcal{G}) \text{ a.s.}\} \\ \mathcal{D} &= \{B \in \mathcal{B}(\mathbb R^n)\,:\,B = [-\infty, r_1) \times \cdots \times [-\infty, r_n), \, r_i \in \mathbb{Q}\},\end{aligned}
then a moment's reflection will convince you that that H\mathcal{H} is a λ\lambda-class, D\mathcal D is a π\pi-class, and H⊇D.\mathcal{H} \supseteq \mathcal D. Hence, by the π−λ\pi-\lambda theorem H⊇σ(D)=B(Rn)\mathcal{H} \supseteq \sigma(\mathcal D) = \mathcal{B}(\mathbb R^n), or in other words, PT(ω,B)\mathbb{P}_T(\omega, B) is a regular conditional distribution for TT given G.\mathcal{G}.

In fact, this theorem is easily extended to (R∞,B(R∞))(\mathbb R^\infty, \mathcal{B}(\mathbb R^\infty)) as follows: For all n≥1n \ge 1, define FnωF_n^\omega as in Equation (4). Select the null set N∈GN \in \mathcal{G} such that in addition to the conditions it satisfies above we also have the consistency condition

lim⁡rn+1→∞Fnω(r1,…,rn,rn+1)=Fnω(r1,…,rn),n≥1.\begin{aligned} \lim_{r_{n+1} \to \infty} F_n^\omega(r_1, \ldots, r_n, r_{n+1}) = F_n^\omega(r_1, \ldots, r_n), \quad n \ge 1.\end{aligned}

For reals x1,…,xnx_1, \ldots, x_n, define just like Equation (5). Then for each ω∈Ω\omega \in \Omega, {Fnω, n≥1}\{F_n^\omega, \, n \ge 1\} is a consistent family of distribution functions, and hence by the Kolmogorov extension theorem there exists a unique measure μω\mu_\omega on (R∞,B(R∞))(\mathbb R^\infty, \mathcal{B}(\mathbb R^\infty)) whose finite dimensional distributions are {Fnω, n≥1}.\{F_n^\omega, \, n \ge 1\}. Define PT(ω,B)=μω(B)\mathbb{P}_T(\omega, B) = \mu_\omega(B) for B∈B(R∞).B \in \mathcal{B}(\mathbb R^\infty). If

H={B∈B(R∞) : PT(⋅,B)=P(T−1(B)∣G) a.s.}D=⋃n=1∞{B∈B(R∞) : B=[−∞,r1)×⋯×[−∞,rn)×R×R×⋯ , ri∈Q},\begin{aligned} \mathcal{H} &= \{B \in \mathcal{B}(\mathbb R^\infty)\,:\,\mathbb{P}_T( \cdot, B) = \mathbb{P}(T^{-1}(B)\mid\mathcal{G}) \text{ a.s.}\} \\ \mathcal D &= \bigcup_{n=1}^\infty \{B \in \mathcal{B}(\mathbb R^\infty)\,:\,B = [-\infty, r_1) \times \cdots \times [-\infty, r_n) \times \mathbb R \times \mathbb R \times \cdots, \, r_i \in \mathbb{Q}\},\end{aligned}
then H\mathcal{H} is a λ\lambda-class, D\mathcal{D} is a π\pi-class, and H⊇D.\mathcal{H} \supseteq \mathcal{D}. Hence, by the π−λ\pi-\lambda theorem H⊇σ(D)=B(R∞).\mathcal{H} \supseteq \sigma(\mathcal{D}) = \mathcal{B}(\mathbb R^\infty).

References