The concept of conditional probability is central to probability theory and excellent treatment of it can be found in many books. My aim with this blog post is to consolidate in one place some ideas around it which helped me form a better intuition. These ideas will be useful if you have already been exposed to this concept from a textbook and just want one more person's ramblings about it.
I will start by defining conditional expectation and stating some of its properties. It will be a grave injustice to claim my discussion of it is complete since I don't even prove its existence; this section exists solely for establishing notation. I will then spend some time discussing conditional probability, relating it to the traditional notion of
P(A∣B)=P(B)P(A∩B).
These discussions will naturally lead to the notions of regular conditional probability and regular conditional distribution which I discuss next.
Theorem 1: Let (Ω,F,P) be a probability space, and X a random variable with E(∣X∣)<∞. Let G be a sub-σ-algebra of F. Then there exists a random variable Y such that
Y is G-measurable,
E(∣Y∣)<∞, and
∫GYdP=∫GXdP for every G∈G.
Remarks:
It is easy to see from the π−λ theorem that the last condition can be relaxed such that ∫GYdP=∫GXdP for every G in some π-system which contains Ω and generates G.
If Y′ is another random variable with the three properties above then Y′=Y a.s.. Therefore, Y in the theorem above is called a version of the conditional expectation. The notation E(X∣G) is used to denote this unique (up to a.e. equivalence) random variable.
Definition 1: In the setting of Theorem 1, if Z is a random variable, we write E(X∣Z) for E(X∣σ(Z)).
The fact that conditional expectation is defined as a random variable might come as a surprise, but the correspondence with the traditional usage of conditional expectation as a number becomes clear once you realize that here we are conditioning on a σ-algebra instead of a single event (we haven’t defined what conditioning on an event means, but think of the intuitive meaning for now). For example, consider the life expectancy of a new born baby conditioned on sex. This is a random variable that takes one value for males and another value for females.
For completeness I state some useful properties of conditional expectation. You can find the proofs in [1,2] for example. Most of them are parallels to the well-known properties of (unconditional) expectation. Assume that all the X's satisfy E(∣X∣)<∞ and let G,H be sub-σ-algebras of F.
[Linearity] E(a1X1+a2X2∣G)=a1E(X1∣G)+a2E(X2∣G) a.s. for real numbers a1 and a2.
[Positivity] If X≥0, then E(X∣G)≥0 a.s..
[Monotone convergence theorem for conditional expectation] If E(∣Y∣)<∞ and Y≤Xn↑X a.s., then E(Xn∣G)↑E(X∣G) a.s..
[Fatou's lemma for conditional expectation] If E(∣Y∣)<∞ and Y≤Xn for all n≥1 a.s., then E(liminfn→∞Xn∣G)≤liminfn→∞E(Xn∣G) a.s..
[Dominated convergence theorem for conditional expectation] If ∣Xn∣≤∣Y∣ for all n≥1, E(∣Y∣)<∞, and Xn→X a.s., then E(Xn∣G)→E(X∣G) a.s..
[Tower property] If H⊆G, then E(E(X∣G)∣H)=E(E(X∣H)∣G)=E(X∣H) a.s.
[Taking out what's known] If Y is G-measurable and bounded, then E(YX∣G)=YE(X∣G) a.s.. If p>1, 1/p+1/q=1, X∈Lp(Ω,F,P) and Y∈Lq(Ω,G,P), then (1) again holds. If X is a nonnegative F-measurable random variable, Y is a nonnegative G-measurable random variable, E(X)<∞ and E(XY)<∞, then also (1) holds.
[Role of independence] If H is independent of σ(σ(X)∪G), then E(X∣σ(G∪H))=E(X∣G) a.s.. In particular, if X is independent of H, then E(X∣H)=E(X) a.s..
Definition 2: In the setting of Theorem 1, if A∈F, we let P(A∣G) to mean E(1A∣G) and call it the conditional probability of A given G. Here 1A is the indicator random variable. If B∈F, we let P(A∣B) to mean E(1A∣1B).
Just like conditional expectation, conditional probability, as defined above, is a random variable! Unlike conditional expectation this isn't very palpable and deserves more rumination (Halmos, 1950). We have our probability space (Ω,F,P) and let A,B∈F be such that P(B)=0 and P(Bc)=0. Then our traditional notion of conditional probability tells us that the conditional probability of A given B is defined by
PB(A)=P(B)P(A∩B).
Let us investigate how PB(A) depends on B. To this end, introduce the discrete measurable space (Λ,2Λ) with Λ={λ1,λ2}, and a measurable mapping T:Ω→Λ such that
T(ω)={λ1λ2 if ω∈B if ω∈Bc.
Define the two measures νA and ν on (Λ,2Λ) as follows for any E⊆Λ,
In other words conditional probability may be viewed as a measurable function on Λ.
This can easily be generalized to any finite setting as follows. Let {A1,…,An}⊆F be a partition of Ω, i.e., Ai∩Aj=∅ for i=j and ⋃iAi=Ω. Introduce the discrete measurable space (Λ,2Λ) with Λ={λ1,…,λn}. Define a measurable mapping T:Ω→Λ such that T(ω)=λi whenever ω∈Ai. Define the measures νA1,…,νAn,ν on (Λ,2Λ) as follows for any E⊆Λ,
νAi(E)ν(E)=P(Ai∩T−1(E))for all i=1,…,n=P(T−1(E)).
Then once again we have for any A∈F,
PAi(A)=P(Ai)P(A∩Ai)=ν({λi})νAi({λi})for all i=1,…,n.
These considerations are what motivated the definition of conditional probability in general cases, as you see in Definition 2. If T is any measurable mapping from (Ω,F,P) into an arbitrary measurable space (Λ,L), and if we write νA(E)=P(A∩T−1(E)) where A∈F and E∈L, then it is clear that νE and P∘T−1 are measures on L such that νA≪P∘T−1. Radon-Nikodym theorem now implies that there exists an P∘T−1-integrable function pA, unique upto P∘T−1-a.e., such that
P(A∩T−1(E))=∫EpA(λ)(P∘T−1)(dλ)for all E∈L.
We anoint pA(λ) as the conditional probability of A given λ∈Λ or the conditional probability of A given that T(ω)=λ. Note that here we are conditioning on a measurable mapping T instead of a sub-σ-algebra, but this notion is related to conditioning on σ(T) as will become clear ahead. Keep this "rumination" in mind when we discuss regular conditional distribution later.
Let's look at our definition of conditional probability from the other direction and show that P(A∣B) as defined in Definition 2 conforms to our traditional usage. To start, note that σ(1B)={∅,B,Bc,Ω}, and since P(A∣B) is σ(1B)-measurable, it must be constant on each of the sets B,Bc, thereby necessitating
P(A∣B)(ω)={P(B)P(A∩B)P(Bc)P(A∩Bc) if ω∈B if ω∈Bc
because of property 3. of Theorem 1 by taking G to be B and Bc. Of course, if any of the sets B or Bc is of measure 0, then you can take the corresponding value for P(A∣B)(ω) to be anything in [0,1] and it won't matter since the concept of conditional probability is defined only up to sets of measure 0.
Positivity property (property 2. above) and monotone convergence property (property 3. above) of conditional expectation imply 0≤P(A∣G)≤1 a.s. for any A∈F, P(A∣G)=0 a.s. if and only if P(A)=0, and P(A∣G)=1 a.s. if and only if P(A)=1.
Let A1,A2,…∈F be a sequence of disjoint sets. By linearity (property 1. above) and monotone convergence theorem for conditional expectation (property 3. above), we see that
P(n⋃AnG)=n∑P(An∣G)a.s..
If An∈F for n≥1 and limn→∞An=A, then we also have limn→∞P(An∣G)=P(A∣G) a.s..
It seems very tempting from the foregoing discussion to claim that P(⋅∣G)(ω) is a probability measure on F for almost all ω∈Ω, but except for some nice spaces, which we will discuss below, this isn't true. Let us first try to see this intuitively (Chow and Teicher, 1997). Equation (2) holds for all ω∈Ω EXCEPT for some null set which may well depend on the particular sequence {An}n∈N. It does NOT stipulate that there exists a fixed null set N∈F such that
P(n⋃AnG)(ω)=n∑P(An∣G)(ω),ω∈Ω∖N
for every disjoint sequence {An}n∈N⊆F. Except in trivial cases, there are uncountably many disjoint sequences, and therefore we will need uncountable union of such null sets to be of measure 0, which of course may not even be defined let alone be of measure 0. To further drive this point home you can take a look at an explicit example of how this can fail in an exercise in (Halmos, 1950) on page 210.
Motivated by our discussion above, we define regular conditional probability as follows.
Definition 3: Let (Ω,F,P) be a probability space and G,H be sub-σ-algebras of F. A regular conditional probability on H given G is a function P(⋅,⋅):Ω×H→[0,1] such that
for a.e. ω∈Ω, P(ω,⋅) is a probability measure on H,
for each A∈H, P(⋅,A) is a G-measurable function on Ω coinciding with the conditional probability of A given G, i.e., P(⋅,A)=P(A∣G) a.s..
Let's show that this definition is not outrageous by showing that it agrees with our traditional notion of conditional pdf and conditional expectation. To that end, suppose that X and Y are random variables which have a joint probability density function fX,Y(x,y). This means that we are considering the probability space (R2,B(R2),P) with X and Y being the coordinate random variables, i.e. (x,y)↦x and (x,y)↦y respectively, and having an absolutely continuous distribution function FX,Y(x,y) such that
FX,Y(x,y)=∫−∞y∫−∞xfX,Y(s,t)dsdt.
We recall that fX(x)=∫RfX,Y(x,y)dy and fY(y)=∫RfX,Y(x,y)dx act as probability density functions for X and Y respectively, and
fX∣Y(x∣y)={fY(y)fX,Y(x,y)0 if fY(y)=0 otherwise
defines the elementary conditional pdf fX∣Y of X given Y. By Fubini's theorem fX and fY are Borel functions on R and so fX∣Y is a Borel function on R2. Let H=B(R2)=σ(X,Y) and G=σ(Y)=R×B(R). For A∈H and ω=(x,y)∈R2 we define
P(ω,A)=∫{s:(s,y)∈A}fX∣Y(s∣y)ds.
Then for each ω∈R2, P(ω,⋅) is a probability measure on H, and for each A∈H, P(⋅,A) is a Borel function in y and hence G-measurable. To verify that P(⋅,A)=P(A∣G) a.s. for any A∈H we just need to verify property 3. of Theorem 1. To this end, fix A∈H and G∈G, and note that G must be of the form G=R×B for B∈B(R). Thus
∫GP(ω,A)dP(ω)=∫B∫RP((s,t),A)fX∣Y(s,t)dsdt(by absolute continuity and Fubini’s theorem)=∫B∫R[∫{u:(u,t)∈A}fX∣Y(u∣t)du]fX∣Y(s,t)dsdt=∫B[∫{u:(u,t)∈A}fX∣Y(u∣t)du]fY(t)dt=∫B∫{u:(u,t)∈A}fX,Y(u,t)dudt=∫B∫R1A(u,t)fX,Y(u,t)dudt=∫G1A(ω)dP(ω)
and so P(⋅,A)=E(1A∣G)=P(A∣G) a.s.. Hence, P(ω,A) is a regular conditional probability on H given G.
For the corresponding analysis for conditional expectation, let h be a Borel function on R2 such that
E∣h(X,Y)∣=∫R∫R∣h(x,y)∣fX,Y(x,y)dxdy<∞.
Set
g(y)=∫Rh(s,y)fX∣Y(s∣y)ds.
g(y) is the traditional conditional density of h(X,Y) given Y=y. Then the claim is that
g(Y)=E(h(X,Y)∣σ(Y))a.s..
A typical element of σ(Y) has the form {ω∈R2:Y(ω)∈B}, where B∈B(R). Hence, we must show that
In general, we have the following useful theorem (taken from (Chow and Teicher, 1997)) which allows us to view conditional expectations as ordinary expectations relative to the measure induced by regular conditional probability.
Theorem 2: Consider the setting of Definition 3 and denote Pω(⋅)=P(ω,⋅). Let X be an H-measurable function with E(X)<∞. Then E(X∣G)(ω)=∫ΩXdPωa.s..
Recall the monotone class theorem for functions:
Monotone Class Theorem for Functions: Let H be a family of nonnegative functions on Ω which contains all indicators of sets of some class H of subsets of Ω. If either (i) H is a π-class and H is a λ-system, or (ii) H is a σ-algebra and H is a monotone system, then H contains all nonnegative σ(H)-measurable functions.
By separate considerations of X+ and X−, it may be supposed that X≥0. Define
H={X:X≥0,X is H-measurable, and (3) holds for X}.
By the definition of regular conditional probability, 1A∈H for A∈H.H is already a σ-algebra. Let's show that H is a monotone system.
If X1,X2∈H and c1,c2≥0, then c1X1+c2X2≥0, c1X1+c2X2 is H-measurable and Equation (3) holds because of linearity of expectation and conditional expectation, and thus c1X1+c2X2∈H. If {Xn}n∈N⊆H such that Xn↑X, then X≥0, X is H-measurable, and Equation (3) holds for X because of monotone convergence theorem for expectation and conditional expectation, and thus X∈H. Therefore, by the monotone class theorem H contains all nonnegative H-measurable functions.
In some cases even the concept of regular conditional probability in inadequate, and that motivates the concept of regular conditional distributions.
Definition 4: Let (Ω,F,P) be a probability space, G⊆F a σ-algebra, (Λ,L) a measurable space, and T:Ω→Λ a measurable mapping. A regular conditional distribution for T given G is a function PT:Ω×L→[0,1] such that
for a.e. ω∈Ω, PT(ω,⋅) is a probability measure on L,
for each A∈L, PT(⋅,A) is a G-measurable function on Ω coinciding with the conditional probability of T−1(A) given G, i.e., PT(⋅,A)=P(T−1(A)G) a.s..
It is clear that when Λ=Ω, L=H⊆F and T is the identity map, PT is exactly the regular conditional probability as defined in Definition 3.
Now would be a good time to reread the first “rumination” in the section Conditional Probability and realize that the definition of regular conditional distribution is in fact well motivated.
A corresponding version of Theorem 2 exists, proof of which I'll leave as an easy exercise:
Theorem 3: In the setting of Definition 4, if PTω(A)=PT(ω,A) and h:Λ→R is a Borel function with E(∣h(T)∣)<∞, then
E(h(T)∣G)(ω)=∫Λh(λ)PTω(dλ).
To see the power of thinking about conditional probabilities like this, let's give an unbelievably short proof of conditional Hölder's inequality that I took from (Chow and Teicher, 1997). Contrast it with other proofs.
Theorem 4: If X,Y are random variables on (Ω,F,P), G⊆F is a σ-algebra, and 1<p<∞, 1/p+1/q=1, then
E(∣XY∣∣G)≤[E(∣X∣p∣G)]1/p[E(∣Y∣q∣G)]1/q a.s..
For B∈B(R2) and ω∈Ω, let PX,Yω(B)=PX,Y(ω,B) be the regular conditional distribution for (X,Y) given G. Theorem 3 allows us to write
Definition 5: Let (X,X) and (Y,Y) be two measurable spaces. They are called isomorphic if there exists a bijection f:X→Y such that f and its inverse f−1 are both measurable. The function f is called an isomorphism.
Definition 6: A measurable space (X,X) is called a standard Borel space if it satisfies any of the following equivalent conditions:
(X,X) is isomorphic to some compact metric space with the Borel σ-algebra.
(X,X) is isomorphic to some Polish space (i.e., a separable complete metric space) with the Borel σ-algebra.
(X,X) is isomorphic to some Borel subset of some Polish space with the Borel σ-algebra.
As you can guess most spaces we deal with are standard Borel spaces. (Durrett, 2019) calls these space nice since we already have too many things named after Borel. I am not sure I agree with his reasoning but I like Durrett's terminology.
The next two theorems show the existence of regular conditional distribution and are taken from [(Durrett, 2019), Section 4.1.3]. See also [(Parthasarathy, 1967), Section V.8].
Theorem 5: Regular conditional distribution exists if (Λ,L) is nice.
A generalization of the last theorem:
Theorem 6: Suppose (Λ,L) is a nice space, T and S are measurable mappings from Ω to Λ, and G=σ(S). Then there exists a function μ:Λ×L→[0,1] such that
for a.e. ω∈Ω, μ(S(ω),⋅) is a probability measure on L, and
for each A∈L, μ(S(⋅),A)=P(T−1(A)∣G) a.s..
It is instructive to prove Theorem 5 in the special case when (Λ,L)=(Rn,B(Rn)). The theorem and the proof is taken from (Chow and Teicher, 1997).
Theorem 7: In the setting of Definition 4, let (Λ,L)=(Rn,B(Rn)) and T=(T1,…,Tn):Ω→Rn. Then there exists a regular conditional distribution for T given G.
Let's recall the definition of an n-dimensional distribution function on Rn (the Russian convention of left-continuous distribution function).
An n-dimensional distribution function on Rn is a function F:Rn→[0,1] satisfying:
We will try to construct a distribution function on Rn. To this end, for any rational number r1,…,rn and ω∈Ω, define Fnω(r1,…,rn):=P(i=1⋂n{Ti<ri}G)(ω).
It's evident that the properties of conditional probability discussed above imply that there is a null set N∈G such that for ω∈Ω∖N and all rational numbers ri,ri′,qi,m the following holds
Fnω(r1,…,rn)Fnω(r1,…,rn)ri→−∞limFnω(r1,…,rn)ri→∞1≤i≤nlimFnω(r1,…,rn)Δnr,r′Fnω≥Fnω(r1′,…,rn′) if ri>ri′,1≤i≤n,=qi,m↑ri1≤i≤nlimFnω(q1,m,…,qn,m),=0,1≤i≤n,=1, and≥0 if r≤r′,
where r≤r′ means ri≤ri′ for all 1≤i≤n. Having defined Fnω for rational values, extend to any real numbers x1,…,xn as follows Fnω(x1,…,xn)=⎩⎨⎧limri↑xiri∈Q1≤i≤nFnω(r1,…,rn)P(⋂i=1n{Ti<ri}) if ω∈Ω∖N if ω∈N.
Then for each ω∈Ω, Fnω(x1,…,xn) is an n-dimensional distribution function and hence determines a Lebesgue-Stieltjes measure μω on B(Rn) with μω(Rn)=1. For B∈B(Rn) and ω∈Ω define
then a moment's reflection will convince you that that H is a λ-class, D is a π-class, and H⊇D. Hence, by the π−λ theorem H⊇σ(D)=B(Rn), or in other words, PT(ω,B) is a regular conditional distribution for T given G.
In fact, this theorem is easily extended to (R∞,B(R∞)) as follows: For all n≥1, define Fnω as in Equation (4). Select the null set N∈G such that in addition to the conditions it satisfies above we also have the consistency condition
For reals x1,…,xn, define just like Equation (5). Then for each ω∈Ω, {Fnω,n≥1} is a consistent family of distribution functions, and hence by the Kolmogorov extension theorem there exists a unique measure μω on (R∞,B(R∞)) whose finite dimensional distributions are {Fnω,n≥1}. Define PT(ω,B)=μω(B) for B∈B(R∞). If