Aditya Makkar
General Theory of Processes - Part 2

This is part 2 of the sequence of blog posts on general theory of processes. For part 1 see here. I will be building upon the content presented there.

Introduction

“Mathematics exists solely for the honour of the human mind.” — Jacobi in a letter to Legendre after the death of Fourier. Fourier had the opinion that the principal aim of mathematics was public utility and explanation of natural phenomena.

We previously discussed Choquet's theory of capacities and its applications in measure theory. The goal of this blog post is to discuss the debut, section and projection theorems in stochastic processes. These theorems form the core of the “general theory of processes”, developed primarily by the Strasbourg school of probability under the tutelage of Paul-André Meyer. At the risk of being overly simplistic, general theory of processes is the study of filtrations and stopping times—martingales will be absent in our discussion.

Unlike part 1 where there were no prerequisites other than basic measure theory, this part assumes a good amount of familiarity with stochastic processes, at the level of (Karatzas and Shreve, 1998a), (Le Gall, 2016). Without that the material here will feel unmotivated and difficult. On the other hand, the theorems proved here are skipped even in advanced courses in stochastic processes as they aren't the most useful for applications, but rather are necessary to fill in gaps if a rigorous treatment is warranted. Nevertheless the theory presented is extremely profound and beautiful, and reading it will be an honour for the mind, if nothing else.

Setting The Stage

Completion of Measure Spaces and Filtrations

Definition 18: A measure space (E,E,μ)(E, \mathscr{E}, \mu) is called complete if BEB \in \mathscr{E} and μ(B)=0\mu(B) = 0 implies that every subset of BB is in E.\mathscr E. We call AEA \subseteq E a μ\mu-negligible or a μ\mu-null set if there exists BEB \in \mathscr E such that ABA \subseteq B and μ(B)=0.\mu(B) = 0. A given statement is said to hold μ\mu-almost everywhere, or simply μ\mu-a.e., if the set on which it fails is μ\mu-negligible.

Thus the measure space (E,E,μ)(E, \mathscr{E}, \mu) is complete if and only if every μ\mu-negligible subset of EE belongs to E.\mathscr E.

Recall that an outer measure on EE is a function μ ⁣:P(E)[0,]\mu^* \colon \mathfrak{P}(E) \to [0, \infty] such that

  1. μ()=0\mu^*(\varnothing) = 0,

  2. if ABEA \subseteq B \subseteq E, then μ(A)μ(B)\mu^*(A) \le \mu^*(B), and

  3. if {An}nN\{A_n\}_{n \in \mathbb N} is a sequence of subsets of EE, then

    μ(nNAn)nNμ(An).\begin{aligned} \mu^*\left( \bigcup_{n \in \mathbb N} A_n \right) \le \sum_{n \in \mathbb N} \mu^*(A_n).\end{aligned}

Also recall that a subset BEB \subseteq E is called μ\mu^*-measurable if

μ(A)=μ(AB)+μ(ABc), AE.\begin{aligned} \mu^*(A) = \mu^*(A \cap B) + \mu^*(A \cap B^{\mathsf{c}}), \quad \forall \; A \subseteq E.\end{aligned}

It is a simple exercise to show that if BEB \subseteq E is such that either μ(B)=0\mu^*(B) = 0 or μ(Bc)=0\mu^*(B^\mathsf{c}) = 0, then BB is μ\mu^*-measurable. Finally, recall that if Mμ\mathscr{M}_{\mu^*} denotes the collection of all μ\mu^*-measurable subsets of EE, then Mμ\mathscr{M}_{\mu^*} is a σ\sigma-algebra, and the restriction of μ\mu^* to Mμ\mathscr{M}_{\mu^*} is a measure on Mμ.\mathscr{M}_{\mu^*}. Therefore, it follows that the measure space (E,Mμ,μ)(E, \mathscr{M}_{\mu^*}, \mu^*) is complete. In particular, the Lebesgue measure on the σ\sigma-algebra of Lebesgue subsets of R\mathbb R is complete. It can be shown that the restriction of Lebesgue measure to the σ\sigma-algebra of Borel subsets of R\mathbb R is not complete. We next discuss a result that allows us to complete any measure space.

Theorem 11: Let (E,E,μ)(E, \mathscr{E}, \mu) be an arbitrary measure space. Define the collection Eμ\mathscr{E}_\mu to consist of all sets BEB \subseteq E for which there exist B1,B2EB_1, B_2 \in \mathscr E such that B1BB2 and μ(B2B1)=0.\begin{aligned} B_1 \subseteq B \subseteq B_2 \text{ and } \mu(B_2 \setminus B_1) = 0.\end{aligned}

Then Eμ\mathscr{E}_\mu is a σ\sigma-algebra on EE that includes E.\mathscr E. Define the function

μ ⁣:Eμ[0,]\begin{aligned} \overline{\mu} \colon \mathscr E_\mu \to [0, \infty]\end{aligned}
by letting μ(B)=μ(B1)\overline{\mu}(B) = \mu(B_1) in the notation above. Then μ\overline \mu is well-defined and is a measure on the σ\sigma-algebra Eμ\mathscr E_\mu whose restriction to E\mathscr E is simply μ.\mu. Finally, the measure space (E,Eμ,μ)(E, \mathscr E_\mu, \overline \mu) is complete, and is called the completion of (E,E,μ).(E, \mathscr{E}, \mu).

Let us start by showing that Eμ\mathscr{E}_\mu is a σ\sigma-algebra. That Eμ\mathscr{E}_\mu includes E\mathscr{E} is clear by taking B1=B2=BB_1 = B_2 = B for any BE.B \in \mathscr{E}. This, in particular, means that Eμ.\varnothing \in \mathscr{E}_\mu. Now let BEμB \in \mathscr{E}_\mu with B1,B2EB_1, B_2 \in \mathscr{E} as in the theorem statement. (1) implies

B2cBcB1c and μ(B1cB2c)=0,\begin{aligned} B_2^\mathsf{c} \subseteq B^\mathsf c \subseteq B_1^\mathsf{c} \text{ and } \mu(B_1^\mathsf{c} \setminus B_2^\mathsf{c}) = 0,\end{aligned}
showing that BcEμ.B^\mathsf{c} \in \mathscr{E}_\mu. Finally, suppose that {Bn}nN\{B^n\}_{n \in \mathbb N} is a sequence of sets in Eμ\mathscr{E}_\mu, such that for each nNn \in \mathbb N, B1n,B2nEB^n_1, B^n_2 \in \mathscr{E} be the sets satisfying (1). Then nNB1n\bigcup_{n \in \mathbb N} B^n_1 and nNB2n\bigcup_{n \in \mathbb N} B^n_2 both belong to E\mathscr E, and satisfy
nNB1nnNBnnNB2n, andμ(nNB2nnNB1n)μ(nN(B2nB1n))nNμ(B2nB1n)=0,\begin{aligned} \begin{gather*} \bigcup_{n \in \mathbb N} B^n_1 \subseteq \bigcup_{n \in \mathbb N} B^n \subseteq \bigcup_{n \in \mathbb N} B^n_2, \text{ and} \\ \mu \left( \bigcup_{n \in \mathbb N} B^n_2 \setminus \bigcup_{n \in \mathbb N} B^n_1 \right) \le \mu \left( \bigcup_{n \in \mathbb N} (B^n_2 \setminus B^n_1) \right) \le \sum_{n \in \mathbb N} \mu(B^n_2 \setminus B^n_1) = 0, \end{gather*}\end{aligned}
showing that nNBnEμ.\bigcup_{n \in \mathbb N} B^n \in \mathscr{E}_\mu.

Next, we show that μ\overline \mu is well-defined. Using the notation in the statement of the theorem, it follows immediately that μ(B1)=μ(B2).\mu(B_1) = \mu(B_2). Furthermore, if ABA \subseteq B such that AEA \in \mathscr{E}, then

μ(A)μ(B2)=μ(B1).\begin{aligned} \mu(A) \le \mu(B_2) = \mu(B_1).\end{aligned}
Hence
μ(B1)=sup{μ(A) : AE and AB},\begin{aligned} \mu(B_1) = \sup \{\mu(A) \,:\, A \in \mathscr{E} \text{ and } A \subseteq B\},\end{aligned}
and so the common value of μ(B1)\mu(B_1) and μ(B2)\mu(B_2) depends only on the set BB and not on the choice of B1B_1 and B2.B_2.

Next, we show that μ\overline \mu is a measure on Eμ.\mathscr E_\mu. μ\overline \mu is clearly an extension of μ\mu by letting B1=B2=BB_1 = B_2 = B for any BE.B \in \mathscr{E}. This, in particular, implies that μ()=0.\overline \mu(\varnothing) = 0. Non-negativity of the measure μ\mu implies the non-negativity of μ.\overline \mu. Finally, we check the countable additivity. Let {Bn}nN\{B^n\}_{n \in \mathbb N} be a sequence of disjoint sets in Eμ\mathscr E_\mu, such that for each nNn \in \mathbb N, B1n,B2nEB^n_1, B^n_2 \in \mathscr{E} be the sets satisfying (1). The disjointness of the sets {Bn}nN\{B^n\}_{n \in \mathbb N} implies the disjointness of the sets {B1n}nN\{B^n_1\}_{n \in \mathbb N}, and so we get

μ(nNBn)=μ(nNB1n)=nNμ(B1n)=nNμ(Bn).\begin{aligned} \overline{\mu}\left( \bigcup_{n \in \mathbb N} B^n \right) = \mu\left( \bigcup_{n \in \mathbb N} B^n_1 \right) = \sum_{n \in \mathbb N} \mu(B^n_1) = \sum_{n \in \mathbb N} \overline\mu(B^n).\end{aligned}
Lastly, we need to check that the measure space (E,Eμ,μ)(E, \mathscr E_\mu, \overline \mu) is complete. Suppose BEμB \in \mathscr E_\mu is such that μ(B)=0,\overline\mu(B) = 0, and suppose AB.A \subseteq B. We need to show that AEμ.A \in \mathscr E_\mu. Let B1,B2EB_1, B_2 \in \mathscr E be sets from (1). For A1=A_1 = \varnothing and A2=B2A_2 = B_2, the conditions (1) are satisfied for the set A,A, and we are done.

If we denote by N\mathscr{N} the collection of all μ\mu-negligible sets for the measure space (E,E,μ)(E, \mathscr{E}, \mu), then it is easy to see that Eμ={BN : BE and NN},\begin{aligned} \mathscr E_\mu = \{B \cup N \,:\, B \in \mathscr E \text{ and } N \in \mathscr N\},\end{aligned} and μ(BN)=μ(B).\overline \mu (B \cup N) = \mu(B). From this it follows that if the measure space (E,E,μ)(E, \mathscr{E}, \mu) is complete, then it is its own completion.

A very nice property that holds if we assume that the measure space (E,E,μ)(E, \mathscr{E}, \mu) is complete is that if ff and gg are real-valued functions defined EE such that ff is measurable and and f=gf = g holds μ\mu-a.e., then gg is also measurable. To see this, let BB(R)B \in \mathcal{B}(\mathbb R) be any Borel set. Then {fB}E\{f \in B\} \in \mathscr E and

A:={gB}{fg}{fg}NE\begin{aligned} A := \{g \in B\} \cap \{f \neq g\} \subseteq \{f \neq g\} \in \mathscr N \subseteq \mathscr E\end{aligned}
by our assumptions. Thus, {f=g}={fg}cE\{f = g\} = \{f \neq g \}^\mathsf{c} \in \mathscr E, and
{gB}=({fB}{f=g})AE\begin{aligned} \{g \in B\} = \left(\{f \in B\} \cap \{f = g\}\right) \cup A \in \mathscr E\end{aligned}
allows us to conclude that gg is measurable.

The above property can fail if the measure space is not complete. Similar problems can arise in the study of stochastic processes, and therefore completeness assumptions are made. For this whole blog post, we shall place ourselves on a complete probability space (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) endowed with a filtration F={Ft}0t<\mathbb{F} = \{\mathcal{F}_t\} _ {0 \le t < \infty}, i.e., a family of sub-σ\sigma-algebras of F\mathcal{F} which is increasing in the sense that

FsFtF:=σ(u[0,)Fu),0st<.\begin{aligned} \mathcal{F}_s \subseteq \mathcal{F}_t \subseteq \mathcal{F}_\infty := \sigma\left(\bigcup_{u \in [0, \infty)} \mathcal{F}_u\right), \quad 0 \le s \le t < \infty.\end{aligned}
We sometimes denote this setup conveniently as (Ω,F,F,P)(\Omega, \mathcal{F}, \mathbb{F}, \mathbb{P}) and call it a filtered probability space. We shall assume that this filtration satisfies the usual conditions:

  1. Right-continuity, i.e., Ft=Ft+:=s>tFs\mathcal{F}_t = \mathcal{F}_{t+} := \bigcap_{s > t} \mathcal{F}_s for all t[0,)t \in [0, \infty), and

  2. F0\mathcal{F}_0 contains all the P\mathbb{P}-negligible events in F\mathcal{F}.

The structure of the filtration implies that we can write Ft+\mathcal{F}_{t+} equivalently as nNFt+1/n.\bigcap_{n \in \mathbb N} \mathcal{F}_{t + 1/n}.

Often we will need to start with the natural filtration FtX:=σ(Xs; 0st),t0,\begin{aligned} \mathcal{F}_t^X := \sigma(X_s; \;0 \le s \le t), \quad t \ge 0,\end{aligned} generated by the process X={Xt}t0,X = \{X_t\}_{t \ge 0}, which is the smallest filtration with respect to which the process XX is adapted, i.e., XtX_t is FtX\mathcal{F}_t^X-measurable for each t0.t \ge 0. And then we will need to augment it to make it satisfy the usual conditions. To that end, we define the minimal augmented filtration generated by XX to be the smallest filtration that is right continuous and complete and with respect to which the process XX is adapted. This can be constructed in the following three steps:

  1. First, let {FtX}t0\left\{\mathcal{F}_t^{X}\right\}_{t \ge 0} be the natural filtration as in (3).

  2. Let N\mathscr{N} be the collection of all P\mathbb{P}-negligible sets for the complete probability space (Ω,F,P)(\Omega, \mathcal{F}, \mathbb{P}) (we can always complete it using Theorem 11 if it isn’t). For each t0,t \ge 0, let (FtX)P={FN : FFtX and NN}\begin{aligned} \left(\mathcal{F}^X_t\right)_{\mathbb{P}} = \left\{F \cup N \,:\, F \in \mathcal{F}_t^X \text{ and } N \in \mathscr N\right\}\end{aligned} be the completion of FtX\mathcal{F}_t^X just like in (2).

  3. Finally, for each t0,t \ge 0, let Ft:=((FtX)P)+=s>t(FsX)P\begin{aligned} \mathcal{F}_t := \left( \left(\mathcal{F}^X_t\right)_{\mathbb{P}} \right)_+ = \bigcap_{s > t} \left(\mathcal{F}^X_s\right)_{\mathbb{P}}\end{aligned} define the right-continuous filtration. To show right-continuity of {Ft}t0,\{\mathcal{F}_t\}_{t \ge 0}, we need to show that

    s>t(FsX)P=Ft=Ft+=r>ts>r(FsX)P.\begin{aligned} \bigcap_{s > t} \left(\mathcal{F}^X_s\right)_{\mathbb{P}} = \mathcal{F}_t = \mathcal{F}_{t+} = \bigcap_{r > t} \bigcap_{s > r} \left(\mathcal{F}^X_s\right)_{\mathbb{P}}.\end{aligned}
    But this is obvious since F(FsX)PF \in \left(\mathcal{F}^X_s\right)_{\mathbb{P}} for every s>r>ts > r > t is same as saying F(FsX)PF \in \left(\mathcal{F}^X_s\right)_{\mathbb{P}} for every s>t.s > t.

These steps give a filtration {Ft}t0,\{\mathcal{F}_t\}_{t \ge 0}, which is right continuous and complete and with respect to which the process XX is adapted. It is clear that it is also the smallest such filtration. Does it matter if we do step 3 before step 2, i.e., is it true that

((FtX)P)+=((FtX)+)P?\begin{aligned} \left( \left(\mathcal{F}^X_t\right)_{\mathbb{P}} \right)_+ = \left( \left(\mathcal{F}^X_t\right)_{+} \right)_{\mathbb{P}} ?\end{aligned}
It would be very disappointing if they weren’t equal. Let us start by showing
((FtX)+)P((FtX)P)+.\begin{aligned} \left( \left(\mathcal{F}^X_t\right)_{+} \right)_{\mathbb{P}} \subseteq \left( \left(\mathcal{F}^X_t\right)_{\mathbb{P}} \right)_+.\end{aligned}
From (2), if F((FtX)+)P,F \in \left( \left(\mathcal{F}^X_t\right)_{+} \right)_{\mathbb{P}}, then F=BNF = B \cup N for some B(FtX)+B \in \left(\mathcal{F}^X_t\right)_{+} and NN.N \in \mathscr N. Then BFsXB \in \mathcal{F}^X_s for every s>t,s > t, which means that BN(FsX)P,B \cup N \in \left(\mathcal{F}^X_s\right)_{\mathbb{P}}, showing F(FsX)PF \in \left(\mathcal{F}^X_s\right)_{\mathbb{P}} for every s>t.s > t.

For the other side,

((FtX)P)+((FtX)+)P,\begin{aligned} \left( \left(\mathcal{F}^X_t\right)_{\mathbb{P}} \right)_+ \subseteq \left( \left(\mathcal{F}^X_t\right)_{+} \right)_{\mathbb{P}},\end{aligned}
let F((FtX)P)+.F \in \left( \left(\mathcal{F}^X_t\right)_{\mathbb{P}} \right)_+. We need to construct B(FtX)+B \in \left(\mathcal{F}^X_t\right)_{+} and NNN \in \mathscr N such that F=BN.F = B \cup N. By the definition of right-continuity, F(Ft+1/nX)PF \in \left(\mathcal{F}^X_{t + 1/n}\right)_{\mathbb{P}} for every nN,n \in \mathbb N, which means there exist sequences {Bn}nN\{B_n\}_{n \in \mathbb N} and {Nn}nN\{N_n\}_{n \in \mathbb N} satisfying BnFt+1/nX,B_n \in \mathcal{F}^X_{t + 1/n}, NnN,N_n \in \mathscr N, and F=BnNnF = B_n \cup N_n for each nN.n \in \mathbb N. Define
B:=lim infnBn={Bn,eventually}=kNnkBn.\begin{aligned} B := \liminf_{n \to \infty} B_n = \{B_n, \text{eventually}\} = \bigcup_{k \in \mathbb N} \bigcap_{n \ge k} B_n.\end{aligned}
Then note that nkBnFt+1/kX,\bigcap_{n \ge k} B_n \in \mathcal{F}_{t + 1/k}^X, which implies that for each mN,m \in \mathbb N,
B=kNnkBn=kmnkBnFt+1/mX.\begin{aligned} B = \bigcup_{k \in \mathbb N} \bigcap_{n \ge k} B_n = \bigcup_{k \ge m} \bigcap_{n \ge k} B_n \in \mathcal{F}_{t + 1/m}^X.\end{aligned}
This says that B(FtX)+.B \in \left(\mathcal{F}^X_t\right)_{+}. Define N:=FB,N := F \setminus B, and note
FB=FkNnkBnc=kNnk(FBn)=kNnkNnN.\begin{aligned} F \setminus B = F \cap \bigcap_{k \in \mathbb N} \bigcup_{n \ge k} B_n^\mathsf{c} = \bigcap_{k \in \mathbb N} \bigcup_{n \ge k} (F \setminus B_n) = \bigcap_{k \in \mathbb N} \bigcup_{n \ge k} N_n \in \mathscr{N}.\end{aligned}
This shows the desired inclusion, and we are done.

Stochastic Processes

For our purposes, a stochastic process is simply a collection of real-valued random variables X={Xt}t0X = \{X_t\}_{t \ge 0} on (Ω,F).(\Omega, \mathcal{F}). The sample paths of XX are the mapping [0,)tXt(ω)R[0, \infty) \ni t \mapsto X_t( \omega) \in \mathbb R obtained when fixing ωΩ.\omega \in \Omega.

We need terminology to talk about when two stochastic processes are the “same”.

Definition 19: Consider two stochastic process X={Xt}t0X = \left\{ X_t\right\}_{t \ge 0} and Y={Yt}t0Y = \left\{ Y_t\right\}_{t \ge 0} defined on the same probability space (Ω,F,P).(\Omega, \mathcal{F}, \mathbb{P}). We say

  1. XX and YY have the same finite-dimensional distributions if, for any integer n1n \ge 1, real numbers 0t1<t2<<tn<,0 \le t_1 < t_2 < \cdots < t_n < \infty, and AB(Rn)A \in \mathcal{B}(\mathbb R^n), we have

P{(Xt1,,Xtn)A}=P{(Yt1,,Ytn)A};\begin{aligned} \mathbb{P}\{(X_{t_1}, \ldots, X_{t_n}) \in A\} = \mathbb{P}\{(Y_{t_1}, \ldots, Y_{t_n}) \in A\};\end{aligned}
  1. YY is a modification or a version of XX if, for every t[0,)t \in [0, \infty), we have P{Xt=Yt}=1;\mathbb{P}\{X_t = Y_t\} = 1;

  2. XX and YY are indistinguishable if almost all their sample paths agree, i.e.,

    P{Xt=Yt t[0,)}=1.\begin{aligned} \mathbb{P}\{X_t = Y_t \;\; \forall \, t \in [0, \infty)\} = 1.\end{aligned}
    The set {Xt=Yt t[0,)}\{X_t = Y_t \;\; \forall \, t \in [0, \infty)\} is measurable because of our completeness assumption.

We will need to make stronger measurability assumptions for the random variables {Xt}t0\{X_t\}_{t \ge 0} than just assuming that each XtX_t is measurable.

Definition 20: A stochastic process X={Xt}t0X = \{X_t\}_{t \ge 0} is called

  1. adapted, if XtX_t is Ft\mathcal{F}_t-measurable for every t[0,)t \in [0, \infty);

  2. measurable, if the mapping [0,)×Ω(t,ω)Xt(ω)R\begin{aligned} [0, \infty) \times \Omega \ni (t,\omega) \mapsto X_t(\omega) \in \mathbb R\end{aligned} is B([0,))F\mathcal{B}([0, \infty)) \otimes \mathcal{F}-measurable, when R\mathbb R is endowed with its Borel σ\sigma-algebra;

  3. progressively measurable, if for every t[0,)t \in [0, \infty) the mapping

    [0,t]×Ω(s,ω)Xs(ω)R\begin{aligned} [0,t] \times \Omega \ni (s,\omega) \mapsto X_s(\omega) \in \mathbb R\end{aligned}
    is B([0,t])Ft\mathcal{B}([0,t]) \otimes \mathcal{F}_t-measurable, when R\mathbb R is endowed with its Borel σ\sigma-algebra.

Recall that every progressively measurable process is both measurable and adapted; and every adapted process with right-continuous or with left-continuous paths, is progressively measurable (Karatzas and Shreve, 1998a), (Le Gall, 2016). On the other hand, we have the following result due to Meyer (1966).

Theorem 12: Every measurable and adapted process X={Xt}t0X = \{X_t\} _ {t \ge 0} has a progressively measurable modification Y={Yt}t0.Y = \{Y_t\} _ {t \ge 0}.

Consider the space L0\mathbb{L}^0 of equivalence classes of measurable functions f:ΩRf : \Omega \to \mathbb R, and endow it with the topology of convergence in probability, for example with the pseudo-metrics ρ(f,g)=E(1fg)\rho(f, g) = \mathbb E (1 \wedge |f-g|) or ρ(f,g)=E(fg1+fg).\rho(f,g) = \mathbb E \left( \frac{|f-g|}{1 + |f-g|} \right). Recall that if a sequence {fn}nNL0\{f_n\} _ {n \in \mathbb N} \subseteq \mathbb{L}^0 satisfies nNρ(fn,fn+1)<,\sum _ {n \in \mathbb N} \rho(f _ n, f _ {n+1}) < \infty, then it converges in probability “fast”, thus also almost surely.

Every process Y={Yt}t0Y = \{Y_t\} _ {t \ge 0} can be thought of as a mapping from [0,)[0, \infty) into L0\mathbb{L}^0, which associates to each element in [0,)[0, \infty) the equivalence class Yt\mathcal{Y} _ t of Yt.Y_t. Consider now the collection H\mathscr{H} of processes YY such that the mapping [0,)tYtL0[0, \infty) \ni t \mapsto \mathcal{Y}_t \in \mathbb{L}^0 satisfies

  1. it takes values in a separable subspace of L0\mathbb{L}^0;

  2. under it, the inverse image of every open ball of L0\mathbb{L}^0 is a Borel subset of [0,)[0, \infty); and

  3. it is the uniform limit of a sequence of simple measurable functions with values in the space L0.\mathbb{L}^0.

Then note that property 3 implies the other two, and that conversely, properties 1 and 2 together imply property 3, because a real-valued function is measurable if and only if it is the increasing limit of a sequence of simple measurable functions.

Next, note that H\mathscr{H} is a real vector space on account of property 3, and is closed under sequential pointwise convergence on account of properties 1 and 2. It is easy to see that H\mathscr{H} contains all processes of the form Y(t,ω)=K(t)ξ(ω)Y(t,\omega) = K(t) \xi(\omega), where KK is the indicator of an interval in [0,)[0, \infty) and ξ:ΩR\xi : \Omega \to \mathbb R is a bounded measurable function. Thus, monotone class theorem implies H\mathscr{H} contains all bounded measurable processes. Now using a standard slicing argument it is easily seen that every measurable process YY has properties 1, 2 and 3.

Consider now a measurable and adapted process X={Xt}t0.X = \{X_t\} _ {t \ge 0}. For each nN{n \in \mathbb N}, there exists a process X(n)={Xt(n)}t0X^{(n)} = \{X^{(n)}_t\} _ {t \ge 0} which is simple and measurable (in the sense that there exists a partition {Ak(n)}kN\left\{A^{(n)} _ k\right\} _ {k \in \mathbb N} of [0,)[0, \infty) into Borel sets, and a sequence of random variable {Hk(n)}kN\left\{H^{(n)} _ k\right\} _ {k \in \mathbb N}, such that Xt(n)=Hk(n)X^{(n)}_t = H^{(n)} _ k for tAk(n)t \in A^{(n)} _ k), and satisfies

ρ(Xt,Xt(n))2(n+1), t[0,).\begin{aligned} \rho(X_t, X^{(n)}_t) \le 2^{-(n+1)}, \quad \forall \; t \in [0, \infty).\end{aligned}

Next, we define a sequence of random variables {Gk(n)}kN\left\{G^{(n)} _ k\right\} _ {k \in \mathbb N} for each nN{n \in \mathbb N} using the sequence {Hk(n)}kN\left\{H^{(n)} _ k\right\} _ {k \in \mathbb N} to get “enhanced” measurability properties. To this end, define sk(n)=infAk(n)s^{(n)} _ k = \inf A^{(n)} _ k, and define Gk(n)=X(sk(n))G^{(n)} _ k = X(s^{(n)} _ k) if sk(n)Ak(n)s^{(n)} _ k \in A^{(n)} _ k; if, on the other hand sk(n)Ak(n)s^{(n)} _ k \notin A^{(n)} _ k, we pick any Fsk(n)\mathcal{F}_{s^{(n)} _ k}-measurable Gk(n)G^{(n)} _ k that satisfies

ρ(Gk(n),Hk(n))2(n+1).\begin{aligned} \rho\left(G^{(n)} _ k, H^{(n)} _ k\right) \le 2^{-(n+1)}.\end{aligned}
Such a Gk(n)G^{(n)} _ k always exists since we may take a decreasing sequence {θm}mNAk(n)\{\theta_m\} _ {m \in \mathbb N} \subseteq A^{(n)} _ k with θmsk(n)\theta_m \downarrow s^{(n)} _ k, and set Gk(n)=lim infmX(θm)G^{(n)} _ k = \liminf _ {m \to \infty} X(\theta_m); and the above inequality holds for all integers k.k.

We now define the process Y(n)={Yt(n)}t0Y^{(n)} = \{Y^{(n)}_t\} _ {t \ge 0} by Yt(n):=Gk(n), tAk(n)Y^{(n)}_t := G^{(n)} _ k, \, t \in A^{(n)} _ k, and check that it is progressively measurable and satisfies ρ(Xt,Yt(n))2(n+1)\rho\left(X_t, Y^{(n)}_t\right) \le 2^{-(n+1)} for all t[0,)t \in [0, \infty) and nN.{n \in \mathbb N}.

Finally, we construct a progressively measurable process YY by

Yt(ω):=limnYt(n)(ω), (t,ω)[0,)×Ω such that this limit exists,\begin{aligned} Y_t(\omega) := \lim_{n \to \infty} Y^{(n)}_t(\omega), \quad \forall \; (t,\omega) \in [0, \infty) \times \Omega \text{ such that this limit exists,}\end{aligned}
and Yt(ω):=0Y_t(\omega) := 0 otherwise. The properties of ρ\rho then implies P{Xt=Yt}=1\mathbb{P}\{X_t = Y_t\} = 1 holds for all t[0,)t \in [0, \infty), so YY is a modification of X.X.

Definition 21: The σ\sigma-algebra of progressively measurable sets, denoted by P\mathscr{P} _ \star, is the smallest σ\sigma-algebra on the product space [0,)×Ω[0, \infty) \times \Omega, with respect to which the mappings as in (6) are measurable for all progressively measurable processes X.X.

It can be verified that P\mathscr{P} _ \star is the collection of all sets AFB([0,))A \in \mathcal{F} \otimes \mathcal{B}([0, \infty)) such that the process Xt(ω)=1A(ω,t)X_t(\omega) = \mathbf{1}_A(\omega, t) is progressively measurable. This characterization gives the useful result that a subset AA of [0,)×Ω[0, \infty) \times \Omega belongs to P\mathscr{P} _ \star if and only if, for every t[0,)t \in [0, \infty), A([0,t]×Ω)B([0,t])Ft.A \cap ([0,t] \times \Omega) \in \mathcal{B}([0,t]) \otimes \mathcal{F}_t.

Stopping Times

Definition 22: A random variable T:Ω[0,]T : \Omega \to [0, \infty] is called a stopping time if {Tt}Ft\{T \le t\} \in \mathcal{F}_t for all t[0,).t \in [0, \infty).

Because the filtration F\mathbb{F} is right-continuous this is equivalent to {T<t}Ft\{T < t\} \in \mathcal{F}_t for all t[0,)t \in [0, \infty) as can be seen by writing

{Tt}=qQ,t<q<s{T<q}Fs.\begin{aligned} \{T \le t\} = \bigcap_{\substack{q \in \mathbb{Q}, \, t < q < s}} \{T < q\} \in \mathcal{F}_s.\end{aligned}

It is easily checked that if TT and SS are stopping times, then so are

TS, TS and T+S.\begin{aligned} T \wedge S,\, T \vee S \text{ and } T+S.\end{aligned}
Likewise, if {Tn}nN\{T_n\}_{n \in \mathbb N} is a sequence of stopping times, then
supnTn, infnTn, lim supnTn and lim infnTn\begin{aligned} \sup_n T_n,\, \inf_n T_n,\, \limsup_{n \to \infty} T_n \text{ and } \liminf_{n \to \infty} T_n\end{aligned}
are also stopping times. Stopping times go hand in glove with progressive measurability. Suppose X={Xt}t0X = \{X_t\}_{t \ge 0} is F\mathbb{F}-progressively measurable and TT an F\mathbb{F}-stopping . Then, the stopped process
{XTt}t0\begin{aligned} \left\{X_{T \wedge t}\right\}_{t \ge 0}\end{aligned}
is F\mathbb{F}-progressively measurable.

Definition 23: For a stopping time TT, we define

  1. σ\sigma-algebra of events prior to it by

F(T):={AF : A{Tt}F(t), t[0,)};\begin{aligned} \mathcal{F}(T) := \{A \in \mathcal{F}\,:\, A \cap \{T \le t\} \in \mathcal{F}(t), \quad \forall \; t \in [0, \infty)\};\end{aligned}
  1. σ\sigma-algebra of events strictly prior to it by

F(T):=σ(F(0){A{T>t} : t[0,), AF(t)}).\begin{aligned} \mathcal{F}(T-) := \sigma(\mathcal{F}(0) \cup \left\{A \cap \{T > t\} \,:\, t \in [0, \infty), \, A \in \mathcal{F}(t)\right\}).\end{aligned}

The following intuitive properties are useful and their proofs can be found in any standard textbook (Karatzas and Shreve, 1998a), (Le Gall, 2016).

Lemma 6: Let TT and SS be any two stopping times. Then

  1. F(T)F(T)\mathcal{F}(T-) \subseteq \mathcal{F}(T);

  2. TT is measurable with respect to both F(T)\mathcal{F}(T) and F(T)\mathcal{F}(T-);

  3. F(T)F(S)=F(TS)\mathcal{F}(T) \cap \mathcal{F}(S) = \mathcal{F}(T \wedge S);

  4. A{TS}F(S)A \cap \{T \le S\} \in \mathcal{F}(S), for all AF(T)A \in \mathcal{F}(T);

  5. A{T<S}F(S)A \cap \{T < S\} \in \mathcal{F}(S-), for all AF(T)A \in \mathcal{F}(T);

  6. If the stopping times satisfy TST \le S, then F(T)F(S)\mathcal{F}(T) \subseteq \mathcal{F}(S) and F(T)F(S)\mathcal{F}(T-) \subseteq \mathcal{F}(S-);

  7. If the stopping times satisfy T<ST < S on the set {0<S<}\{0 < S < \infty\}, then F(T)F(S)\mathcal{F}(T) \subseteq \mathcal{F}(S-);

  8. If {Tn}nN\{T_n\} _ {n \in \mathbb N} is an increasing sequence of stopping times and T=limnNTnT = \lim _ {n \in \mathbb N} \uparrow T_n, then

    F(T)=σ(nNF(Tn));as well as F(T)=σ(nNF(Tn))\begin{aligned} \mathcal{F}(T-) = \sigma\left(\bigcup_{n \in \mathbb N} \mathcal{F}(T_n-)\right); \quad \text{as well as } \mathcal{F}(T-) = \sigma\left(\bigcup_{n \in \mathbb N} \mathcal{F}(T_n)\right)\end{aligned}
    provided that Tn<TT_n < T holds on the set {0<T<}\{0 < T < \infty\} for every nN;{n \in \mathbb N};

  9. If ZZ is an integrable random variable, we have

E[ZF(T)]=E[ZF(ST)],P-a.e. on {TS},E[E[ZF(T)]F(S)]=E[ZF(ST)],P-a.e.;\begin{aligned} \mathbb{E}\left[ Z \mid \mathcal{F}(T) \right] &= \mathbb{E}\left[ Z \mid \mathcal{F}(S \wedge T)\right], \quad \mathbb{P}\text{-a.e. on } \{T \le S\} \text{,}\\ \mathbb{E}\left[ \mathbb{E}\left[ Z \mid \mathcal{F}(T) \right] \mid \mathcal{F}(S) \right] &= \mathbb{E}\left[ Z \mid \mathcal{F}(S \wedge T) \right], \quad \mathbb{P}\text{-a.e.};\end{aligned}
  1. If {Tn}nN\{T_n\}_{n \in \mathbb N} is a sequence of stopping times and T:=infnTnT:= \inf_n T_n, then

F(T)=nNF(Tn).\begin{aligned} \mathcal{F}(T) = \bigcap_{n \in \mathbb N} \mathcal{F}(T_n).\end{aligned}

Predictable Stopping Times

A very important result states that every stopping time is the decreasing limit of a sequence of discrete stopping time. More concretely, if TT is a stopping time, and we define

Tn(ω)={k2non {k12nT<k2n}T(ω)on {T=},\begin{aligned} T_n(\omega) = \begin{cases} k 2^{-n} &\text{on } \left\{\frac{k-1}{2^n} \le T < \frac{k}{2^n}\right\} \\ T(\omega) &\text{on } \{T = \infty\}, \end{cases}\end{aligned}

for n,kN,n,k \in \mathbb N, then each TnT_n is a discrete stopping time and we have TnT.T_n \downarrow T. When can we approximate a stopping with an increasing sequence of stopping times? Of course, being able to do that means that our stopping time is a very special one. We give a name to such stopping times.

Definition 24: A stopping time TT is called predictable, if there exists an increasing (“announcing”) sequence {Tn}nN\{T_n\} _ {n \in \mathbb N} of stopping times with T=limnNTnT = \lim _ {n \in \mathbb N} \uparrow T_n, and Tn<TT_n < T holds on the set {T>0}\{T > 0\} for all nN.{n \in \mathbb N}.

Note {Tt}=nN{Tnt}F(t)\{T \le t\} = \bigcap_{n \in \mathbb N} \{T_n \le t\} \in \mathcal{F}(t) and therefore if in the definition above we had not required TT to be a stopping time, it would still be a stopping time. A canonical example of a predictable stopping time is the first time Brownian motion WW hits or exceeds a certain level b.b. It is announced by the sequence of first times WW hits or exceeds the levels b1/nb - 1/n, for nN.{n \in \mathbb N}. In fact, we have (taken from (Almost Sure blog)):

Theorem 13: Let XX be a continuous adapted process and bb be a real number. Then
T=inf{t[0,) : X(t)b}\begin{aligned} T = \inf \{t \in [0, \infty)\,:\,X(t) \ge b\}\end{aligned}
is a predictable stopping time.
Let
Tn=inf{t[0,) : X(t)b1/n},\begin{aligned} T_n = \inf \{t \in [0, \infty)\,:\,X(t) \ge b - 1/n\},\end{aligned}
which, by the Debut Theorem below (Theorem 14) is a stopping time. This gives an increasing sequence {Tn}nN\{T_n\} _ {n \in \mathbb N} of stopping times bounded above by T.T. Also, X(Tn)b1/nX(T _ n) \ge b - 1/n whenever Tn<T _ n < \infty and, by left-continuity, setting τ=limnTn\tau = \lim _ n T _ n gives X(τ)bX(\tau) \ge b whenever τ<.\tau < \infty. So TτT \ge \tau, showing that T=limnNTn.T = \lim _ {n \in \mathbb N} \uparrow T_n. If 0<TnT<0 < T _ n \le T < \infty then, by continuity, X(Tn)=b1/n<b=X(T).X(T_n) = b - 1/n < b = X(T). So, Tn<TT_n < T on the set {0<T<}\{0 < T < \infty\} and the sequence {Tnn}nN\{T_n \wedge n\} _ {n \in \mathbb N} announces T.T.

Of course, in the above proof we cheated by assuming the Debut Theorem for stopping times, proving which is our next agenda. In the same flavor as Lemma-6 we have

Lemma 7: If TT is a predictable stopping time and SS an arbitrary stopping time, then

  1. A{TS}F(S)A \cap \{T \le S\} \in \mathcal{F}(S-) for all AF(T)A \in \mathcal{F}(T-);

  2. A{S=}F(S)A \cap \{S = \infty\} \in \mathcal{F}(S-) for all AF()A \in \mathcal{F}(\infty).

In particular, both the events {TS}\{T \le S\} and {T=S}\{T = S\} belong to F(S).\mathcal{F}(S-).

Debut of a Progressive Set

Definition 25: For any measurable mapping Z:Ω[0,]Z : \Omega \to [0, \infty] and any AFA \in \mathcal{F}, we define the restriction of ZZ on AA as
ZA:=Z1A+1Ac.\begin{aligned} Z_A := Z \mathbf{1}_{A} + \infty \mathbf{1}_{A^\mathsf{c}}.\end{aligned}

Notice that we can write {ZAZ}=A(Ac{Z=})\{Z_A \le Z\} = A \cup \left(A^\mathsf{c} \cap \{Z = \infty\}\right) as a disjoint union. If TT is a stopping time and AF(T)A \in \mathcal{F}(T), then since {TAt}={Tt}AF(t)\{T_A \le t\} = \{T \le t\} \cap A \in \mathcal{F}(t) for all t[0,)t \in [0, \infty), TAT_A is also a stopping time.

Theorem 14 [Debut of a Progressive Set]: Under the usual conditions, if AA is a progressively measurable set, i.e., APA \in \mathscr{P}_\star, then the debut DAD_A is a stopping time.

Since APB([0,))FA \in \mathscr{P}_\star \subseteq \mathcal{B}([0, \infty)) \otimes \mathcal{F}, Measurable Debut theorem (Theorem 9) implies DAD_A is a random variable.

Just like in the proof of Theorem 9, for any real number t>0t > 0, the set {DA<t}\{D_A < t\} is the projection onto Ω\Omega of the set At:=A([0,t)×Ω).A_t := A \cap ([0, t) \times \Omega). Recall that APA \in \mathscr{P}_\star implies AtB([0,t])F(t).A_t \in \mathcal{B}([0,t]) \otimes \mathcal{F}(t). Therefore, Measurable Projection theorem (Theorem 7) implies {DA<t}=π(At)F(t).\{D_A < t\} = \pi(A_t) \in \mathcal{F}(t). Right continuity of the filtration now implies DAD_A is an F\mathbb{F}-stopping time.

The following important theorem now becomes an easy corollary.

Theorem 15 [First Hitting Times]: In the context above, if XX is a progressively measurable process and ΓB(R)\Gamma \in \mathcal{B}(\mathbb R), the first hitting time
HΓ:=inf{t[0,) : X(t)Γ}\begin{aligned} H_\Gamma := \inf \{t \in [0, \infty)\,:\,X(t) \in \Gamma\}\end{aligned}
is a stopping time.
Since XX is a progressively measurable process, the set A={(t,ω)[0,)×Ω : X(t,ω)Γ}A = \{(t,\omega) \in [0, \infty) \times \Omega\,:\,X(t, \omega) \in \Gamma\} is a progressively measurable set. Theorem 14 then implies the debut DAD_A is a stopping time, but it is easy to see DA=HΓ.D_A = H_\Gamma.

Optional and Predictable Processes

Definition 26: The predictable σ\sigma-algebra, denoted by P\mathscr{P}, is the smallest σ\sigma-algebra on [0,)×Ω[0, \infty) \times \Omega with respect to which left-continuous F\mathbb{F}-adapted processes are measurable. A stochastic process is said to predictable if it is P\mathscr{P}-measurable.
Definition 27: The optional σ\sigma-algebra, denoted by O\mathscr{O}, is the smallest σ\sigma-algebra on [0,)×Ω[0, \infty) \times \Omega with respect to which right-continuous F\mathbb{F}-adapted processes are measurable. A stochastic process is said to optional if it is O\mathscr{O}-measurable.

Since all right-continuous or left-continuous adapted processes are progressively measurable, all optional or predictable processes are adapted. We need the concept of stochastic intervals to ease some notation and see some examples of optional and predictable sets.

Definition 28: If U,V:Ω[0,]U,V : \Omega \to [0, \infty] are two random variables, and UVU \le V, then define the stochastic intervals as
U,V:={(t,ω)[0,)×Ω : U(ω)tV(ω)}U,V:={(t,ω)[0,)×Ω : U(ω)t<V(ω)}U,V:={(t,ω)[0,)×Ω : U(ω)<tV(ω)}U,V:={(t,ω)[0,)×Ω : U(ω)<t<V(ω)}.\begin{aligned} \llbracket U,V \rrbracket &:= \{(t,\omega) \in [0, \infty) \times \Omega\,:\,U(\omega) \le t \le V(\omega)\} \\ \llbracket U,V \llbracket &:= \{(t,\omega) \in [0, \infty) \times \Omega\,:\,U(\omega) \le t < V(\omega)\} \\ \rrbracket U,V \rrbracket &:= \{(t,\omega) \in [0, \infty) \times \Omega\,:\,U(\omega) < t \le V(\omega)\} \\ \rrbracket U,V \llbracket &:= \{(t,\omega) \in [0, \infty) \times \Omega\,:\,U(\omega) < t < V(\omega)\}.\end{aligned}
Also, for AF(0)A \in \mathcal{F}(0), define
0A(ω)={0 if ωA if ωAc.\begin{aligned} 0_A(\omega) = \begin{cases} 0 &\text{ if } \omega \in A \\ \infty &\text{ if } \omega \in A^\mathsf{c}. \end{cases}\end{aligned}

Notice that U=U,U=0,U0,U\llbracket U \rrbracket = \llbracket U,U \rrbracket = \llbracket 0, U \rrbracket \setminus \llbracket 0, U \llbracket, and that 0A0_A is predictable with the announcing sequence {0An}nN.\{0_A \wedge n\} _ {n \in \mathbb N}. Denote by S\mathfrak{S} the collection of all stopping times and by S(P)\mathfrak{S}^{(\mathscr{P})} the collection of all predictable stopping times (S\mathfrak{S} is S in fraktur font).

Lemma 8: We collect some useful properties of stochastic intervals in relation to optional and predictable σ\sigma-algebras.

  1. If S,TSS,T \in \mathfrak{S} such that STS \le T, then the stochastic intervals S,T,S,T,S,T,S,T\llbracket S,T \rrbracket, \llbracket S,T \llbracket, \rrbracket S,T \rrbracket, \rrbracket S,T \llbracket and the graphs S,T\llbracket S \rrbracket, \llbracket T \rrbracket are optional.

  2. If S,TSS,T \in \mathfrak{S} such that STS \le T, then the stochastic interval S,T\rrbracket S,T \rrbracket is predictable. If SS(P)S \in \mathfrak{S}^{(\mathscr{P})}, then the stochastic interval S,\llbracket S, \infty \llbracket is predictable.

  3. The σ\sigma-algebra O\mathscr{O} of optional sets can be generated from stochastic intervals in multiple equivalent ways

O=σ({S, : SS})=σ({U,V : S,TS})=σ({S,T : S,TS}).\begin{aligned} \mathscr{O} = \sigma\left(\{\llbracket S,\infty \llbracket \,:\, S \in \mathfrak{S}\}\right) = \sigma\left(\{\llbracket U,V \llbracket \,:\, S, T \in \mathfrak{S}\}\right) = \sigma\left(\{\llbracket S,T \rrbracket \,:\, S,T \in \mathfrak{S}\}\right).\end{aligned}
  1. Similarly, σ\sigma-algebra P\mathscr{P} of predictable sets can be generated from stochastic intervals in multiple equivalent ways

P=σ({S,T,0A : S,TS and AF(0)})=σ({S, : SS(P)})=σ({S,T : S,TS(P)}).\begin{aligned} \mathscr{P} = \sigma\left(\{\rrbracket S,T \rrbracket, \llbracket 0_A \rrbracket \,:\, S,T \in \mathfrak{S} \text{ and } A \in \mathcal{F}(0)\}\right) = \sigma\left(\left\{\llbracket S,\infty \llbracket\,:\,S \in \mathfrak{S}^{(\mathscr{P})}\right\}\right) = \sigma\left(\left\{\llbracket S,T \rrbracket\,:\,S,T \in \mathfrak{S}^{(\mathscr{P})}\right\}\right).\end{aligned}

I will only be proving some results. For a complete picture see (He, Wang and Yan, 1992), (Almost Sure blog).

  1. It is easy to see that 1S,T\mathbf{1}_{\llbracket S,T \llbracket} is right-continuous and adapted, hence S,T\llbracket S,T \llbracket is optional. If Tn=S+1/nT_n = S + 1/n for nN{n \in \mathbb N}, then S=nNS,Tn\llbracket S \rrbracket = \bigcap _ {n \in \mathbb N} \llbracket S, T_n\llbracket, and thus S\llbracket S \rrbracket is also optional. Similarly, for T.\llbracket T \rrbracket. These facts along with simple manipulation of sets immediately imply S,T,S,T,S,T\llbracket S,T \rrbracket, \rrbracket S,T \rrbracket, \rrbracket S,T \llbracket are also optional.

  2. It is easy to see that 1S,T\mathbf{1}_{\rrbracket S,T \rrbracket} is left-continuous and adapted, hence S,T\rrbracket S,T \rrbracket is predictable.

It is a simple exercise to show

Lemma 9: Every predictable process is optional, and every optional process is progressively measurable. In other words
POPB([0,))F.\begin{aligned} \mathscr{P} \subseteq \mathscr{O} \subseteq \mathscr{P} _ \star \subseteq \mathcal{B}([0, \infty)) \otimes \mathcal{F}.\end{aligned}

Properties of Debuts

Recall the setting of section "Measurable Section" from part 1. We showed there that if AB([0,))F,A \in \mathcal{B}([0, \infty)) \otimes \mathcal{F}, then we can define a measurable mapping T:Ω[0,]T : \Omega \to [0, \infty] such that for all ωπ(A)\omega \in \pi(A) we have (T(ω),ω)A(T(\omega), \omega) \in A, or stated more succinctly, TA.\llbracket T \rrbracket \subseteq A. Now, what if we want something stronger, say, TT should be a stopping time? As we will see, we will have to make stronger assumptions on the set AA and even then if we want TT to be a stopping time we will have to let go of the nice property that {T<}=π(A).\{T < \infty\}= \pi(A).

But first, we need to prove a certain property of debuts which we will need in proving the aforementioned section theorem, and that's the goal of this section.

It isn't difficult to see that P\mathscr{P} coincides with the mosaic generated by the paving on [0,)×Ω[0, \infty) \times \Omega that consists of finite unions of stochastic intervals of the form S,T\llbracket S,T \rrbracket, where S,TS(P).S,T \in \mathfrak{S}^{(\mathscr{P})}. Generalizing this, let us consider a collection A\mathfrak{A} (A\mathfrak{A} is A in fraktur font) of stopping times that contains 00 and \infty, and which is closed under a.s. equality and under finitely many lattice operations (\wedge and \vee). We denote by J:={U,V : S,TA}\mathcal{J} := \{\llbracket U,V \llbracket\,:\,S,T \in \mathfrak{A}\}, and by T\mathcal{T} the collection of finite unions of elements of J.\mathcal{J}.

Simple observations like S,Tc=0,S T,\llbracket S,T \llbracket^\mathsf{c} = \llbracket 0, S \llbracket \, \cup \, \llbracket T, \infty \llbracket or S,T U,V=SU,(SU)(TV)\llbracket S,T \llbracket \, \cap \, \llbracket U,V \llbracket = \llbracket S \vee U, (S \vee U) \vee (T \wedge V) \llbracket can be used to show that J\mathcal{J} is a paving on [0,)×Ω[0, \infty) \times \Omega and T\mathcal{T} is an algebra on [0,)×Ω.[0, \infty) \times \Omega.

For more structure, let us impose the following properties on the collection A\mathfrak{A}:

  1. For any pair S,TAS,T \in \mathfrak{A}, the stopping time S{S<T}S _ {\{S < T\}} (recall the notation of Definition 25 and the observation right after it) belongs to A\mathfrak{A};

  2. For any increasing sequence {Sn}nNA\{S_n\} _ {n \in \mathbb N} \subseteq \mathfrak{A}, the limit limnNSn\lim _ {n \in \mathbb N} \uparrow S_n belongs to A.\mathfrak{A}.

Property 1 ensures that the debut of an element of T\mathcal{T} belongs to A.\mathfrak{A}. This is because S{S<T}=DS,TS_{\{S < T\}} = D_{\llbracket S,T \llbracket} and because the debut of union of a finite collection of elements from J\mathcal{J} is equal to to the minimum of the debuts of those elements.

Property 2 ensures that the debut of an element of Tδ\mathcal{T} _ \delta (which, recall, equals {nNAn : AnT for all nN}\left\{\bigcup _ {n \in \mathbb N} A_n \,:\, A_n \in \mathcal{T} \text{ for all } {n \in \mathbb N}\right\}) belongs to to A.\mathfrak{A}. This is the main result of this section, and is proved below as Theorem 16.

Coming out of the abstraction for a bit, we note that the collection S\mathfrak{S} of all stopping times satisfies all the conditions imposed above on A.\mathfrak{A}. Similarly for the collection S(P)\mathfrak{S}^{(\mathscr{P})} of all predictable stopping times. The first claim is trivial to see. To show the second claim, we first prove the following lemma.

Lemma 10: For any stopping time TT and set AF(T)A \in \mathcal{F}(T), the random variable TAT_A is a stopping time as we saw above. For TAT_A to be a predictable stopping time, it is necessary that AF(T)A \in \mathcal{F}(T-); this condition is also sufficient if TT is predictable.

If TAT_A is a predictable stopping time, then since A={TAT}(Ac{T=})A = \{T_A \le T\} \setminus \left(A^\mathsf{c} \cap \{T = \infty\}\right) and by Lemma 7 both {TAT},(Ac{T=})F(T)\{T_A \le T\}, \left(A^\mathsf{c} \cap \{T = \infty\}\right) \in \mathcal{F}(T-), we get AF(T).A \in \mathcal{F}(T-).

Conversely, suppose TT is a predictable stopping time. Consider the collection

A:={AF(T) : TA and TAc are both predictable}.\begin{aligned} \mathscr{A} := \{A \in \mathcal{F}(T-) \,:\, T_A \text{ and } T _ {A^\mathsf{c}} \text{ are both predictable}\}.\end{aligned}
It is a σ\sigma-algebra: ΩA\Omega \in \mathscr{A} clearly and it is closed under complements by definition; if A1,A2,AA_1, A_2, \ldots \in \mathscr{A} with A=nNAnA = \bigcup _ {n \in \mathbb N} A_n, note TA=infnNTAn.T_A = \inf _ {n \in \mathbb N} T _ {A _ n}. Let {Tnm}mN\{T _ n ^ m\} _ {m \in \mathbb N} be the announcing sequence for TAnT _ {A _ n} for each nN{n \in \mathbb N}, and define
τn:=min{Tjk : j,kn}.\begin{aligned} \tau _ n := \min \{T _ j ^ k \,:\, j,k \le n\}.\end{aligned}
Then {τn}nN\{\tau _ n\} _ {n \in \mathbb N} announces TAT_A and hence AA.A \in \mathscr{A}.

Thus, it suffices to show that TAT_A is predictable for any AA in a collection of sets that generates F(T).\mathcal{F}(T-). To this end, suppose {Tn}nN\{T^n\} _ {n \in \mathbb N} announces TT, and fix an arbitrary nNn \in \mathbb N and an arbitrary AF(Tn).A \in \mathcal{F}(T^n). For every integer mnm \ge n, the restriction TAmT _ A ^ m is a stopping time and so is Rm:=TAmm.R _ m := T ^ m _ A \wedge m. The sequence {Rm}mN\{R_m\} _ {m \in \mathbb N} then announces TAT_A, showing TAT_A is predictable; similarly for TAc.T _ {A^\mathsf{c}}. We conclude TA,TAcT_A, T _ {A^\mathsf{c}} are predictable for every AnNF(Tn).A \in \bigcup _ {n \in \mathbb N} \mathcal{F}(T^n). But recall

F(T)=σ(nNF(Tn)),\begin{aligned} \mathcal{F}(T-) = \sigma\left(\bigcup _ {n \in \mathbb N} \mathcal{F}(T^n)\right),\end{aligned}
since Tn<TT^n < T on {T>0}\{T > 0\}, so we obtain the property for all AF(T).A \in \mathcal{F}(T-).

Coming back to the second claim above, if S,TS(P)S,T \in \mathfrak{S}^{(\mathscr{P})}, then by Lemma 7, {S<T}F(S)\{S < T\} \in \mathcal{F}(S-), and thus by Lemma 10, S{S<T}S(P).S _ {\{S < T\}} \in \mathfrak{S}^{(\mathscr{P})}. The rest of the conditions are trivial to check.

Theorem 16: Suppose A\mathfrak{A} is a collection of stopping times with the properties postulated above, and let BTδ.B \in \mathcal{T}_\delta. Then the debut DBD_B of this set is a stopping time in A\mathfrak{A}, and its graph DBB.\llbracket D_B\rrbracket \subseteq B.

We first show that DBB.\llbracket D_B\rrbracket \subseteq B. If (s,ω)DB(s, \omega) \in\llbracket D_B\rrbracket, then DB(ω)=sD_B(\omega) = s which is same as saying

s=inf{t[0,) : (t,ω)B}.\begin{aligned} s = \inf \{t \in [0, \infty)\,:\,(t,\omega) \in B\}.\end{aligned}
But we constructed our collection Tδ\mathcal{T}_\delta using stochastic intervals of the form S,T\llbracket S,T \llbracket, and thus
B(ω)={t[0,) : (t,ω)B}\begin{aligned} B(\omega) = \{t \in [0, \infty)\,:\,(t,\omega) \in B\}\end{aligned}
is closed, implying (s,ω)B.(s, \omega) \in B.

To show that DBAD_B \in \mathfrak{A}, consider the collection

B:={SA : SDB},\begin{aligned} \mathfrak{B} := \{S \in \mathfrak{A}\,:\,S \le D_B\},\end{aligned}
and note that B\mathfrak{B} contains 00, is closed under pairwise maximization, i.e., S,TBS,T \in \mathfrak{B} implies STBS \vee T \in \mathfrak{B}, and is closed under countable increasing limits, i.e., if {Sn}nNB\{S _ n\} _ {n \in \mathbb N} \subseteq \mathfrak{B} is an increasing sequence then limnNSnB.\lim _ {n \in \mathbb N} \uparrow S_n \in \mathfrak{B}. Then B\mathfrak{B} contains a representative T:=ess supBT := \operatorname{ess\,sup} \mathfrak{B} of its essential supremum (see (He, Wang and Yan, 1992), (Karatzas and Shreve, 1998b)).

Since BTδB \in \mathcal{T}_\delta we can find a decreasing sequence {Bn}nNT\{B_n\} _ {n \in \mathbb N} \subseteq \mathcal{T} with nNBn=B.\bigcap _ {n \in \mathbb N} B_n = B. Define

Tn:=DBnT, for all nN,\begin{aligned} T_n := D_{B_n \cap \llbracket T, \infty \llbracket} \text{ for all } n \in \mathbb N,\end{aligned}
and note TnBn\llbracket T_n \rrbracket \subseteq B_n for each nN.n \in \mathbb N. As BnT, TB_n \cap \llbracket T, \infty \llbracket \, \in \mathcal{T} and thus can be expressed as a finite union k=1mUk,Vk\bigcup _ {k=1} ^ m \llbracket U_k, V_k \llbracket for Uk,VkTU_k, V_k \in \mathcal{T}, we have
Tn=min1kmDUk,Vk=min1km(Uk){Uk<Vk}T.\begin{aligned} T_n = \min_{1 \le k \le m} D_{\llbracket U_k, V_k \llbracket} = \min_{1 \le k \le m} (U_k)_{\{U_k < V_k\}} \in \mathcal{T}.\end{aligned}
Since BBnB \subseteq B_n, we have TTnDBT \le T_n \le D_B, and therefore TnB.T_n \in \mathfrak{B}. Thus, by definition of essential supremum Tn=TT_n = T a.s., and as
TnNBn=B,\begin{aligned} \llbracket T \rrbracket \subseteq \bigcap _ {n \in \mathbb N} B_n = B,\end{aligned}
we conclude that T=DB.T = D_B.

The next result is Proposition 1.2.26 from (Karatzas and Shreve, 1998a), and is left unproven there.

Theorem 17: Under the usual conditions, let XX be an adapted process with paths that are RCLL. The jumps of XX are exhausted by a sequence of stopping times {Tn}nN\{T_n\}_{n \in \mathbb N},
A:={(t,ω)(0,)×Ω : X(t,ω)X(t,ω)}nNTn.\begin{aligned} A := \{(t,\omega) \in (0, \infty) \times \Omega \,:\, X(t, \omega) \neq X(t-, \omega)\} \subseteq \bigcup_{n \in \mathbb N} \llbracket T_n \rrbracket.\end{aligned}

Write

A=nNAn for An:={(t,ω)(0,)×Ω : X(t,ω)X(t,ω)>1/n}.\begin{aligned} A = \bigcup_{n \in \mathbb N} A_n \text{ for }A_n := \{(t,\omega) \in (0, \infty) \times \Omega \,:\, |X(t, \omega) - X(t-, \omega)| > 1/n\}.\end{aligned}
For each nNn \in \mathbb N, the set AnA_n has no accumulation point in [0,)[0, \infty) because 1/n>0.1/n > 0. Therefore, it is easy to verify
An=pNDAnp for DAnp(ω):=inf{t[0,) : #([0,t]An(ω))p},\begin{aligned} A_n = \bigcup_{p \in \mathbb N} \llbracket D_{A_n}^p \rrbracket \text{ for } D_{A_n}^p(\omega) := \inf \{t \in [0, \infty) \,:\, \#([0,t] \cap A_n(\omega)) \ge p\},\end{aligned}
where the notation #B\#B denotes cardinaility of a set B.B.

We claim that all elements of the sequence {DAnp}p,nN\left\{D_{A_n}^p\right\}_{p,n \in \mathbb N} are stopping times. This is sufficient to prove our desired result.

Start by noting that the process XX is progressively measurable and the process (t,ω)X(t,ω)(t, \omega) \mapsto X(t-, \omega) is left-continuous and adapted, and thus predictable. Thus, the sets {An}nN\{A_n\}_{n \in \mathbb N} are progressively measurable. Fix any arbitrary nNn \in \mathbb N, and note that for p=1p=1, the mapping

ωDAn1(ω)=inf{t[0,) : X(t,ω)X(t,ω)>1/n}\begin{aligned} \omega \mapsto D_{A_n}^1(\omega) = \inf \{t \in [0, \infty) \,:\, |X(t, \omega) - X(t-, \omega)| > 1/n\}\end{aligned}
is a stopping time by Theorem 14 (debut of a progressive set). Arguing by induction, for each pNp \in \mathbb N the random variable DAnp+1D_{A_n}^{p+1} is also a stopping time, as it is the debut of the progressively measurable set
An DAnp,\begin{aligned} A_n \,\cap \; \rrbracket D_{A_n}^p, \infty \rrbracket\end{aligned}
by Lemma 8.

Section Theorems

We are now ready to come back to proving section theorems. We start by proving a general section theorem.

Theorem 18 [General Section Theorem]: In the context of Theorem 16, let H\mathscr{H} be the σ\sigma-algebra generated by the algebra T.\mathcal{T}. For every set AHA \in \mathscr{H} and ε>0\varepsilon > 0, there exists a stopping time TεAT_\varepsilon \in \mathfrak{A} with
TεA and P[π(A)]P(Tε<)+ε.\begin{aligned} \llbracket T_\varepsilon \rrbracket \subseteq A \text{ and } \mathbb{P}\left[\pi(A)\right] \le \mathbb{P}(T_\varepsilon < \infty) + \varepsilon.\end{aligned}

Fix a set AHA \in \mathscr H and ε>0.\varepsilon > 0. Measurable Section Theorem (Theorem 10 in part 1) implies there exists a random variable Z:Ω[0,]Z : \Omega \to [0, \infty] with ZA\llbracket Z \rrbracket \subseteq A and π(A)={Z<},\pi(A) = \{Z < \infty\}, where recall π ⁣:[0,)×ΩΩ\pi \colon [0, \infty) \times \Omega \to \Omega is the canonical projection map. We denote by ν\nu the measure on the measurable space ([0,)×Ω,H)\left([0, \infty) \times \Omega, \mathscr{H}\right) defined by

[0,)×Ωf(t,ω) ν(dt,dω)={Z<}f(Z(ω),ω) P(dω)\begin{aligned} \int_{[0, \infty) \times \Omega} f(t, \omega) \, \nu(\mathrm{d} t, \mathrm{d} \omega) = \int_{\{Z < \infty\}} f(Z(\omega), \omega) \, \mathbb{P}(\mathrm{d} \omega)\end{aligned}
for every H\mathscr{H}-measurable f:[0,)×Ω[0,).f : [0, \infty) \times \Omega \to [0, \infty). Notice that (Z(ω),ω)A(Z(\omega), \omega) \in A for all ω{Z<}.\omega \in \{Z < \infty\}.

By first taking ff to be the indicator 1A\mathbf{1}_{A} for AA and then taking ff to be the indicator 1[0,)×Ω\mathbf{1}_{[0, \infty) \times \Omega} for [0,)×Ω[0, \infty) \times \Omega in the equation above, we see that

ν(A)=P(Z<)P[π(A)]=ν([0,)×Ω).\begin{aligned} \nu(A) = \mathbb{P}(Z < \infty) \equiv \mathbb{P}\left[\pi(A)\right] = \nu([0, \infty) \times \Omega).\end{aligned}
Thus, the measure ν\nu is carried by the set A.A. Similarly, taking ff to be 1B\mathbf{1}_{B} for any BH,B \in \mathscr{H}, we get
ν(B)=π(A)1B(Z(ω),ω) P(dω)P[π(AB)].\begin{aligned} \nu(B) = \int_{\pi(A)} \mathbf{1}_{B}(Z(\omega), \omega) \, \mathbb{P}(\mathrm{d} \omega) \le \mathbb{P}\left[\pi(A \cap B)\right].\end{aligned}

Choquet’s capacitability theorem (Theorem 6 in part 1) applied to the paving T\mathcal{T} and to the capacity

ν(C):=infBH, CBν(B),C[0,)×Ω,\begin{aligned} \nu^*(C) := \inf _ {B \in \mathscr{H},\, C \subseteq B} \nu(B), \quad C \subseteq [0, \infty) \times \Omega,\end{aligned}
we obtain the existence of a subset BεTδB_\varepsilon \in \mathcal{T}_\delta of AA, such that
P[π(A)]=ν(A)ν(Bε)+εP[π(ABε)]+εP[π(Bε)]+ε.\begin{aligned} \mathbb{P}\left[\pi(A)\right] = \nu(A) \le \nu(B_\varepsilon) + \varepsilon \le \mathbb{P}\left[\pi(A \cap B_\varepsilon)\right] + \varepsilon \le \mathbb{P}\left[\pi(B_\varepsilon)\right] + \varepsilon.\end{aligned}
Now take TεT_\varepsilon to be the debut DBεD_{B_\varepsilon} which by Theorem 16 is a stopping time in A.\mathfrak{A}.

Theorem 19 [Optional and Predictable Sections]: Let AA be an optional (respectively, predictable) subset of the product space [0,)×Ω.[0, \infty) \times \Omega. For every ε>0\varepsilon > 0, there exists a stopping time (respectively, a predictable stopping time) TεT_\varepsilon with TεA and P[π(A)]P(Tε<)+ε.\begin{aligned} \llbracket T_\varepsilon \rrbracket \subseteq A \text{ and } \mathbb{P}\left[\pi(A)\right] \le \mathbb{P}(T_\varepsilon < \infty) + \varepsilon.\end{aligned}
Both statements follow from the General Section Theorem (Theorem 17) in conjunction with the observations before Lemma 10, by taking A\mathfrak{A} to be either S\mathfrak{S} or S(P).\mathfrak{S}^{(\mathscr{P})}.

As mentioned in the beginning of the section “Properties of Debuts”, we didn't get the nice equality TA and {T<}=π(A)\begin{aligned} \llbracket T \rrbracket \subseteq A \text{ and } \{T < \infty\} = \pi(A)\end{aligned} as we got in Measurable Section theorem, but instead obtained an approximation (7). The condition TεA\llbracket T_\varepsilon \rrbracket \subseteq A makes sure that {Tε<}π(A)\{T_\varepsilon < \infty\} \subseteq \pi(A), and the P[π(A)]P(Tε<)+ε\mathbb{P}\left[\pi(A)\right] \le \mathbb{P}(T_\varepsilon < \infty) + \varepsilon part ensures that the measure of the difference P[π(A){Tε<}]\mathbb{P}\left[ \pi(A) \setminus \{T_\varepsilon < \infty\} \right] of these two events can be made as small as desired.

To see that it is not always possible to choose a stopping time TT that satisfies (8) if AOA \in \mathscr{O}, we will construct a filtration {Ft}t0\{\mathcal{F}_t\}_{t \ge 0} and choose a set AOA \in \mathscr{O} that forces π(A){T<}\pi(A) \setminus \{T < \infty\} \neq \varnothing for every stopping time TT (the argument is from (Almost Sure blog)).

To this end, let τ:Ω(0,)\tau : \Omega \to (0, \infty) be a random variable such that P{τ<t}>0\mathbb{P}\{\tau < t\} > 0 for all t>0.t > 0. For example, let τ\tau be such that its distribution is uniform on (0,1).(0,1). Let {Ft}t0\{\mathcal{F}_t\}_{t \ge 0} be the completed filtration such that Ft\mathcal{F}_t is generated by {{τs} : st}\{\{\tau \le s\} \,:\, s \le t\}, and thus τ\tau becomes a stopping time. Let A=0,τ ,A = \rrbracket 0, \tau \llbracket\,, which we know from Lemma 8 belongs to O.\mathscr{O}. Note that P{π(A)}=1\mathbb{P}\{\pi(A)\} = 1 by our construction.

It is easy to see that Ft\mathcal{F}_t is trivial when restricted to {τ>t}\{\tau > t\}, i.e., contains only sets of measure 00 or 1.1. So, every Ft\mathcal{F}_t-measurable random variable is a.s. constant on the event {τ>t}.\{\tau > t\}. Therefore, any stopping time TT is deterministic on the event {T<τ}.\{T < \tau\}. So, if TA\llbracket T \rrbracket \subseteq A, we have

T={son {τ>s}on {τs}\begin{aligned} T = \begin{cases} s &\text{on } \{\tau > s\} \\ \infty &\text{on } \{\tau \le s\} \end{cases}\end{aligned}
a.s. for some fixed s>0.s > 0. But then
P{T<}=P{τ>s}<1=P{π(A)}.\begin{aligned} \mathbb{P}\{T < \infty\} = \mathbb{P}\{\tau > s\} < 1 = \mathbb{P}\{\pi(A)\}.\end{aligned}
A similar argument can be used to show that it is not possible to choose a predictable stopping time TT that satisfies (8) if AP.A \in \mathscr{P}.

Applications of Section Theorems

Recall the measurable graph theorem (Theorem 8 in part 1) which implies that a map T:Ω[0,]T : \Omega \to [0, \infty] is measurable if and only if its graph T\llbracket T \rrbracket is measurable. We also have the following neat result:

Theorem 20: A random variable T:Ω[0,]T : \Omega \to [0, \infty] is a stopping time (respectively, a predictable stopping time), if and only if its graph T\llbracket T \rrbracket is an optional (respectively, predictable) set.

The necessity follows from the characterization of O\mathscr{O} and P\mathscr{P} in Lemma 8.

In the optional case, sufficiency follows from Theorem 14 and the fact that OP.\mathscr{O} \subseteq \mathscr{P}_\star.

For the sufficiency in the predictable case, suppose T\llbracket T \rrbracket is predictable. Apply Predictable Section theorem (Theorem 18) for A=TA = \llbracket T \rrbracket to construct a sequence {Tn}nN\{T_n\} _ {n \in \mathbb N} of predictable stopping times such that

TnT and P(T<)P(Tn<)+2n, nN.\begin{aligned} \llbracket T_n \rrbracket \subseteq \llbracket T \rrbracket \text{ and } \mathbb{P}(T < \infty) \le \mathbb{P}(T_n < \infty) + 2^{-n}, \quad \forall \; {n \in \mathbb N}.\end{aligned}
Replacing TnT_n by T1TnT_1 \vee \cdots \vee T_n if necessary, we may assume that this sequence is increasing. It follows then that the limit
T=limnNTn=supnNTn\begin{aligned} T = \lim _ {n \in \mathbb N} \uparrow T_n = \sup _ {n \in \mathbb N} T_n\end{aligned}
is predictable: if {Tnm}mN\{T_n^m\} _ {m \in \mathbb N} announces TnT_n, then
τn=max{Tjk : j,kn}\begin{aligned} \tau_n = \max \{T_j^k \,: \, j,k \le n\}\end{aligned}
announces T.T.

Suppose XX and YY are measurable processes such that

P{XT1{T<}=YT1{T<}}=1\begin{aligned} \mathbb{P}\left\{X_T \mathbf{1}_{\{T < \infty\}} = Y_T \mathbf{1}_{\{T < \infty\}}\right\} = 1\end{aligned}
for each F\mathcal{F}-measurable T:Ω[0,].T : \Omega \to [0, \infty]. Then what can we say about XX and Y?Y? We will show that XX and YY are indistinguishable! To this end, let
F={(t,ω)[0,)×Ω : X(t,ω)Y(t,ω)}.\begin{aligned} F = \{(t,\omega) \in [0, \infty) \times \Omega \,:\, X(t,\omega) \neq Y(t, \omega)\}.\end{aligned}
It is measurable since XX and YY are measurable and due to our completeness assumption. We want to show that P{π(F)}=0\mathbb{P}\{\pi(F)\} = 0, so suppose to the contrary that P{π(F)}>0.\mathbb{P}\{\pi(F)\} > 0. Any F\mathcal{F}-measurable T:Ω[0,]T : \Omega \to [0, \infty] satisfying TF\llbracket T \rrbracket \subseteq F must also satisfy XT1{T<}YT1{T<}.X_T \mathbf{1}_{\{T < \infty\}} \neq Y_T \mathbf{1}_{\{T < \infty\}}. Now Measurable Section theorem (Theorem 10) says there exists a random variable T:Ω[0,]T : \Omega \to [0,\infty] such that TF\llbracket T \rrbracket \subseteq F and {T<}=π(F)\{T < \infty\} = \pi(F), and therefore
P{XT1{T<}YT1{T<}}P{T<}=P{π(F)}>0,\begin{aligned} \mathbb{P}\left\{X_T \mathbf{1}_{\{T < \infty\}} \neq Y_T \mathbf{1}_{\{T < \infty\}} \right\} \ge \mathbb{P}\{T < \infty\} = \mathbb{P}\{\pi(F)\} > 0,\end{aligned}
contradicting our initial assumption.

Remarkably, we have similar results for optional and predictable processes.

Theorem 21: Suppose two optional (respectively, predictable) processes XX and YY agree at finite stopping times (respectively, predictable stopping times), i.e.,
P{XT1{T<}=YT1{T<}}=1, TS (respectively, TS(P)).\begin{aligned} \mathbb{P}\left\{X_T \mathbf{1}_{\{T < \infty\}} = Y_T \mathbf{1}_{\{T < \infty\}}\right\} = 1, \quad \forall \; T \in \mathfrak{S} \; \left(\text{respectively, } T \in \mathfrak{S}^{(\mathscr{P})}\right).\end{aligned}
Then the two processes XX and YY are indistinguishable, i.e.,
P{Xt=Yt t[0,)}=1.\begin{aligned} \mathbb{P}\{X_t = Y_t \;\; \forall \, t \in [0, \infty)\} = 1.\end{aligned}
Consider the optional case—the predictable case is similar. Just like above we have to show that for the optional set
F={(t,ω)[0,)×Ω : X(t,ω)Y(t,ω)}, P{π(F)}=0.\begin{aligned} F = \{(t,\omega) \in [0, \infty) \times \Omega \,:\, X(t,\omega) \neq Y(t, \omega)\}, \; \mathbb{P}\{\pi(F)\} = 0.\end{aligned}
Suppose to the contrary that P{π(F)}=2ε>0\mathbb{P}\{\pi(F)\} = 2\varepsilon > 0 for some 0<ε1/2.0 < \varepsilon \le 1/2. Then the Optional Section theorem (Theorem 19) implies there exists a stopping time TεT_\varepsilon with the properties
TεF and 2ε=P{π(F)}P{Tε<}+ε.\begin{aligned} \llbracket T_\varepsilon \rrbracket \subseteq F \text{ and } 2\varepsilon = \mathbb{P}\{\pi(F)\} \le \mathbb{P}\{T_\varepsilon < \infty\} + \varepsilon.\end{aligned}
But then
P{XTε1{Tε<}YTε1{Tε<}}P{Tε<}ε>0,\begin{aligned} \mathbb{P}\left\{X_{T_\varepsilon} \mathbf{1}_{\{T_\varepsilon < \infty\}} \neq Y_{T_\varepsilon} \mathbf{1}_{\{T_\varepsilon < \infty\}}\right\} \ge \mathbb{P}\{T_\varepsilon < \infty\} \ge \varepsilon > 0,\end{aligned}
contradicting our initial assumption.

Projection Theorems

To motivate optional and predictable projections, let us start with a fundamental problem in filtering theory: Assume an underlying complete probability space (Ω,F,P).(\Omega, \mathcal{F}, \mathbb{P}). There is an underlying signal X={Xt}t0X = \{X_t\} _ {t \ge 0} which is modelled as a stochastic process and which we are interested in studying. Our observation process has noise and therefore instead of observing XX we observe a process Y={Yt}t0Y = \{Y_t\} _ {t \ge 0} such that

Yt=ft(Xt,Wt),\begin{aligned} Y_t = f_t(X_t, W_t),\end{aligned}
for Wiener process W={Wt}t0.W = \{W_t\}_{t \ge 0}. Therefore, it makes sense to define F={Ft}t0\mathbb{F} = \{\mathcal{F}_t\} _ {t \ge 0} to be the minimal augmented filtration generated by Y.Y. The problem, then, is to compute an estimate for XX based on the observable data at time tt, viz. Ft.\mathcal{F}_t.

We could look at Zt:=E(XtFt),\begin{aligned} Z_t := \mathbb{E} \left( X_t \mid \mathcal{F}_t \right),\end{aligned} as an estimate for XtX_t at each time t0.t \ge 0. The process Z={Zt}t0Z = \{Z_t\}_{t \ge 0} is, of course, adapted. However, since conditional expectation is defined only up to P\mathbb{P}-a.s., what version of ZZ should we choose? (9) does not fix the paths of the process ZZ, which requires specifying its values at the uncountable set of time in [0,).[0, \infty). We would be very lucky if it were possible for us to choose a version of ZZ such that (9) holds not only for all t[0,)t \in [0, \infty) but also for all finite stopping times. And indeed this is possible! This is part of the statement of the optional projection theorem.

On the other hand, if we want an estimate of XX based on observable data before time tt, then our estimate would be

Zt=E(XtFt)\begin{aligned} Z_t = \mathbb{E}(X_t \mid \mathcal{F}_{t-})\end{aligned}
where recall Ft:=σ(s<tFs).\mathcal{F}_{t-} := \sigma\left(\bigcup _ {s < t} \mathcal{F}_s\right). Again it is possible to choose a version of ZZ such that this equality holds not only for all t[0,)t \in [0, \infty) but also for all finite predictable stopping times, and this is part of the statement of the predictable projection theorem.

Theorem 22 [Optional and Predictable Projections]: Let XX be a bounded, measurable (though not necessarily adapted) process.

  1. There is a unique, modulo indistinguishability, optional process XoX^o, called optional projection of XX, that satisfies for all stopping times TST \in \mathfrak{S} the identity

E(XT1{T<}FT)=XTo1{T<}.\begin{aligned} \mathbb{E}\left( X_T \mathbf{1}_{\{T < \infty\}} \mid \mathcal{F}_T \right) = X^o_T \mathbf{1}_{\{T < \infty\}}.\end{aligned}
  1. There is a unique, modulo indistinguishability, predictable process XpX^p, called predictable projection of XX, that satisfies for all predictable stopping times TS(P)T \in \mathfrak{S}^{(\mathscr{P})} the identity

E(XT1{T<}FT)=XTp1{T<}.\begin{aligned} \mathbb{E}\left( X_T \mathbf{1}_{\{T < \infty\}} \mid \mathcal{F}_{T-} \right) = X^p_T \mathbf{1}_{\{T < \infty\}}.\end{aligned}

Remarks:

  1. Taking expectations in (10) we get E(XT1{T<})=E(XTo1{T<})\begin{aligned} \mathbb{E} \left(X_T \mathbf{1}_{\{T < \infty\}}\right) = \mathbb{E} \left(X_T^o \mathbf{1}_{\{T < \infty\}}\right)\end{aligned} for all stopping times TS.T \in \mathfrak{S}. Now suppose (12) holds for all stopping times TS.T \in \mathfrak{S}. Fix an arbitrary stopping time SSS \in \mathfrak{S} and a set AFSA \in \mathcal{F}_S, and write (12) for T=SAT = S_A to get

    E(XS1{S<}1A)=E(XSo1{S<}1A).\begin{aligned} \mathbb{E} \left(X_S \mathbf{1}_{\{S < \infty\}} \mathbf{1}_A\right) = \mathbb{E} \left(X_S^o \mathbf{1}_{\{S < \infty\}} \mathbf{1}_A\right).\end{aligned}
    But this immediately implies (10) for stopping time S.S. Therefore, requiring (10) to hold for all stopping times is equivalent to requiring (12) to hold for for all stopping times.

  2. Similarly, requiring (11) to hold for all predictable stopping times is equivalent to requiring E(XT1{T<})=E(XTp1{T<})\begin{aligned} \mathbb{E}\left( X_T \mathbf{1}_{\{T < \infty\}} \right) = \mathbb{E}\left(X^p_T \mathbf{1}_{\{T < \infty\}}\right)\end{aligned} to hold for all predictable stopping times.

  3. Operators o^o and p^p are linear operators, i.e., for bounded, measurable processes X,YX,Y and a,bRa,b \in \mathbb R,

    (aX+bY)o=aXo+bYo, and(aX+bY)p=aXp+bYp,\begin{aligned} (aX + bY)^o &= a X^o + bY^o, \text{ and} \\ (aX + bY)^p &= a X^p + bY^p,\end{aligned}
    as can be easily verified by using the properties of conditional expectation. The formulation of (12) and (13) makes it obvious that (Xo)o=Xo\left(X^o\right)^o = X^o and (Xp)p=Xp.\left(X^p\right)^p = X^p. This explains the terminology of “projection”.

[of Theorem 22] Uniqueness is immediate from Theorem 21, so let's focus on existence, first for optional projection and then for predictable projection.

  1. We will employ the monotone class theorem for functions. We need a simple class of processes for which finding an optional projection is easy to summon. To this end, consider the processes of the form

    Xt(ω)=1B(ω)1[u,v)(t),0u<v< and BF.\begin{aligned} X_t(\omega) = \mathbf{1}_{B}(\omega) \mathbf{1}_{[u,v)}(t), \quad 0 \le u < v < \infty \text{ and } B \in \mathcal{F}.\end{aligned}
    We claim that the candidate for the optional projection is
    Xto(ω):=Mt(ω)1[u,v)(t),\begin{aligned} X^o_t(\omega) := M_t(\omega) \mathbf{1}_{[u,v)}(t),\end{aligned}
    where MM is the right-continuous version of the bounded, thus also uniformly integrable, martingale E(1BFt)\mathbb{E}\left( \mathbf{1}_B \mid \mathcal{F}_t \right) (see (Karatzas and Shreve, 1998a), (Le Gall, 2016) for the proof of existence of such a martingale; it is here that the usual conditions on the filtration F\mathbb{F} become crucial).

    With these choices and an arbitrary stopping time TT, the left-hand side of (12) becomes P{B{uT<v}}\mathbb{P}\{B \cap \{u \le T < v\}\}, whereas optional stopping theorem (Karatzas and Shreve, 1998a), (Le Gall, 2016) shows that its right-hand side is

    E(E(1BFT)1[u,v)(T)).\begin{aligned} \mathbb{E} \left( \mathbb{E} \left(\mathbf{1}_B \mid \mathcal{F}_T \right) \mid \mathbf{1}_{[u,v)}(T) \right).\end{aligned}
    Recalling now that TT is FT\mathcal{F}_T-measurable shows that the two sides are equal.

    Finally, use linearity and monotone class arguments to establish existence for arbitrary bounded, measurable X.X.

  2. Similar ideas work in the predictable case. We first consider processes of the form

    Xt(ω)=1B(ω)1(u,v](t),0u<v< and BF.\begin{aligned} X_t(\omega) = \mathbf{1}_{B}(\omega) \mathbf{1}_{(u,v]}(t), \quad 0 \le u < v < \infty \text{ and } B \in \mathcal{F}.\end{aligned}
    We claim that the candidate for the predictable projection is
    Xtp(ω):=Mt(ω)1(u,v](t),\begin{aligned} X^p_t(\omega) := M_{t-}(\omega) \mathbf{1}_{(u,v]}(t),\end{aligned}
    where MtM_{t-} is the left-continuous version of the martingale MM defined previously. Again using FT\mathcal{F}_{T-}-measurability of TT we see the two sides of (13) are equal. Monotone class arguments now allow us to finish the proof.

We end our discussion with a result on time change. We call it a “time change” because given an adapted increasing process AA with right-continuous paths and A0=0,A_0 = 0, we imagine a clock which runs according to AA in the sense that at time tt this clock shows time At.A_t.

Theorem 23: Suppose AA is an adapted increasing process with right-continuous paths and A0=0.A_0 = 0.

  1. For any two bounded, measurable processes XX and YY that satisfy E(XT1{T<})=E(YT1{T<}), TS,\begin{aligned} \mathbb{E} \left(X_T \mathbf{1}_{\{T < \infty\}}\right) = \mathbb{E} \left(Y_T \mathbf{1}_{\{T < \infty\}}\right), \quad \forall \; T \in \mathfrak{S},\end{aligned} we have E0TXt dAt=E0TYt dAt, TS.\begin{aligned} \mathbb{E} \int_0^T X_t \,\mathrm{d}A_t = \mathbb{E} \int_0^T Y_t \,\mathrm{d}A_t, \quad \forall \; T \in \mathfrak{S}.\end{aligned}

  2. For any non-negative, RCLL and uniformly integrable martingale M,M, we have

E0TMt dAt=E(MTAT), TS.\begin{aligned} \mathbb{E} \int_0^T M_t \,\mathrm{d}A_t = \mathbb{E} \left( M_T A_T\right), \quad \forall \; T \in \mathfrak{S}.\end{aligned}

Only the first part requires any real effort.

  1. Introduce the time change

    C(s):=inf{t0 : At>s},\begin{aligned} C(s) := \inf \{t \ge 0 \,:\, A_t > s\},\end{aligned}
    and check that C(s)C(s) is a stopping time for each s0.s \ge 0. Recall that C(s)C(s) is increasing, right-continuous, and the “functional inverse” of AtA_t in the sense that
    At=inf{s0 : C(s)>t}.\begin{aligned} A_t = \inf \{s \ge0 \,:\, C(s) > t\}.\end{aligned}
    Therefore, CC gives us a way to read the correct time off our AA-clock—if the AA-clock shows time s,s, the actual time is C(s).C(s).

    We also check that the properties

    0g(At) dAt=0A()g(u) du, and 0g(t) dAt=0A()g(C(s)) ds=0g(C(s))1{C(s)<} ds\begin{aligned} \begin{gather*} \int_0^\infty g(A_t) \, \mathrm{d}A_t = \int_0^{A(\infty)} g(u) \, \mathrm{d}u, \text{ and } \\ \int_0^\infty g(t) \, \mathrm{d}A_t = \int_0^{A(\infty)} g(C(s)) \, \mathrm{d}s = \int_0^\infty g(C(s)) \mathbf{1}_{\{C(s) < \infty\}}\, \mathrm{d}s \end{gather*}\end{aligned}
    hold for any Borel-measurable g ⁣:[0,)[0,).g \colon [0, \infty) \to [0, \infty). It follows from these considerations and Fubini’s theorem, that
    E0TXt dAt=E0TX(C(s))1{C(s)<} ds=0E[X(C(s))1{C(s)<}] ds\begin{aligned} \mathbb{E} \int_0^T X_t \,\mathrm{d}A_t = \mathbb{E} \int_0^T X(C(s)) \mathbf{1}_{\{C(s) < \infty\}}\, \mathrm{d}s = \int_0^\infty \mathbb{E}\left[X(C(s)) \mathbf{1}_{\{C(s) < \infty\}} \right] \, \mathrm{d}s\end{aligned}
    holds, with a similar expression also holding for Y.Y.

    The claim (15) now follows directly from our assumptions, in the case T=.T = \infty. For general TST \in \mathfrak S we apply the above result to the increasing, right-continuous and adapted process

    A~t:=At10,T(t)+AT1T,(t),t0.\begin{aligned} \tilde{A}_t := A_t \mathbf{1}_{\llbracket 0, T\llbracket}(t) + A_T \mathbf{1}_{\llbracket T, \infty \llbracket}(t), \quad t \ge 0.\end{aligned}

  2. This follows immediately from part-1 by letting

    X:=M10,T and Y:=MT10,T,\begin{aligned} X := M \mathbf{1}_{\llbracket 0, T \rrbracket} \text{ and } Y := M_T \mathbf{1}_{\llbracket 0, T \rrbracket},\end{aligned}
    and using Doob’s optional sampling theorem (Karatzas and Shreve, 1998a), (Le Gall, 2016) to verify that these processes satisfy (14).

References