“Mathematics exists solely for the honour of the human mind.” — Jacobi in a letter to Legendre after the death of Fourier. Fourier had the opinion that the principal aim of mathematics was public utility and explanation of natural phenomena.
We previously discussed Choquet's theory of capacities and its applications in measure theory. The goal of this blog post is to discuss the debut, section and projection theorems in stochastic processes. These theorems form the core of the “general theory of processes”, developed primarily by the Strasbourg school of probability under the tutelage of Paul-André Meyer. At the risk of being overly simplistic, general theory of processes is the study of filtrations and stopping times—martingales will be absent in our discussion.
Unlike part 1 where there were no prerequisites other than basic measure theory, this part assumes a good amount of familiarity with stochastic processes, at the level of (Karatzas and Shreve, 1998a), (Le Gall, 2016). Without that the material here will feel unmotivated and difficult. On the other hand, the theorems proved here are skipped even in advanced courses in stochastic processes as they aren't the most useful for applications, but rather are necessary to fill in gaps if a rigorous treatment is warranted. Nevertheless the theory presented is extremely profound and beautiful, and reading it will be an honour for the mind, if nothing else.
Definition 18: A measure space (E,E,μ) is called complete if B∈E and μ(B)=0 implies that every subset of B is in E. We call A⊆E a μ-negligible or a μ-null set if there exists B∈E such that A⊆B and μ(B)=0. A given statement is said to hold μ-almost everywhere, or simply μ-a.e., if the set on which it fails is μ-negligible.
Thus the measure space (E,E,μ) is complete if and only if every μ-negligible subset of E belongs to E.
Recall that an outer measure on E is a function μ∗:P(E)→[0,∞] such that
μ∗(∅)=0,
if A⊆B⊆E, then μ∗(A)≤μ∗(B), and
if {An}n∈N is a sequence of subsets of E, then
μ∗(n∈N⋃An)≤n∈N∑μ∗(An).
Also recall that a subset B⊆E is called μ∗-measurable if
μ∗(A)=μ∗(A∩B)+μ∗(A∩Bc),∀A⊆E.
It is a simple exercise to show that if B⊆E is such that either μ∗(B)=0 or μ∗(Bc)=0, then B is μ∗-measurable. Finally, recall that if Mμ∗ denotes the collection of all μ∗-measurable subsets of E, then Mμ∗ is a σ-algebra, and the restriction of μ∗ to Mμ∗ is a measure on Mμ∗. Therefore, it follows that the measure space (E,Mμ∗,μ∗) is complete. In particular, the Lebesgue measure on the σ-algebra of Lebesgue subsets of R is complete. It can be shown that the restriction of Lebesgue measure to the σ-algebra of Borel subsets of R is not complete. We next discuss a result that allows us to complete any measure space.
Theorem 11: Let (E,E,μ) be an arbitrary measure space. Define the collection Eμ to consist of all sets B⊆E for which there exist B1,B2∈E such that B1⊆B⊆B2 and μ(B2∖B1)=0.
Then Eμ is a σ-algebra on E that includes E. Define the function
μ:Eμ→[0,∞]
by letting μ(B)=μ(B1) in the notation above. Then μ is well-defined and is a measure on the σ-algebra Eμ whose restriction to E is simply μ. Finally, the measure space (E,Eμ,μ) is complete, and is called the completion of (E,E,μ).
Let us start by showing that Eμ is a σ-algebra. That Eμ includes E is clear by taking B1=B2=B for any B∈E. This, in particular, means that ∅∈Eμ. Now let B∈Eμ with B1,B2∈E as in the theorem statement. (1) implies
B2c⊆Bc⊆B1c and μ(B1c∖B2c)=0,
showing that Bc∈Eμ. Finally, suppose that {Bn}n∈N is a sequence of sets in Eμ, such that for each n∈N, B1n,B2n∈E be the sets satisfying (1). Then ⋃n∈NB1n and ⋃n∈NB2n both belong to E, and satisfy
Next, we show that μ is well-defined. Using the notation in the statement of the theorem, it follows immediately that μ(B1)=μ(B2). Furthermore, if A⊆B such that A∈E, then
μ(A)≤μ(B2)=μ(B1).
Hence
μ(B1)=sup{μ(A):A∈E and A⊆B},
and so the common value of μ(B1) and μ(B2) depends only on the set B and not on the choice of B1 and B2.
Next, we show that μ is a measure on Eμ.μ is clearly an extension of μ by letting B1=B2=B for any B∈E. This, in particular, implies that μ(∅)=0. Non-negativity of the measure μ implies the non-negativity of μ. Finally, we check the countable additivity. Let {Bn}n∈N be a sequence of disjoint sets in Eμ, such that for each n∈N, B1n,B2n∈E be the sets satisfying (1). The disjointness of the sets {Bn}n∈N implies the disjointness of the sets {B1n}n∈N, and so we get
Lastly, we need to check that the measure space (E,Eμ,μ) is complete. Suppose B∈Eμ is such that μ(B)=0, and suppose A⊆B. We need to show that A∈Eμ. Let B1,B2∈E be sets from (1). For A1=∅ and A2=B2, the conditions (1) are satisfied for the set A, and we are done.
If we denote by N the collection of all μ-negligible sets for the measure space (E,E,μ), then it is easy to see that Eμ={B∪N:B∈E and N∈N}, and μ(B∪N)=μ(B). From this it follows that if the measure space (E,E,μ) is complete, then it is its own completion.
A very nice property that holds if we assume that the measure space (E,E,μ) is complete is that if f and g are real-valued functions defined E such that f is measurable and and f=g holds μ-a.e., then g is also measurable. To see this, let B∈B(R) be any Borel set. Then {f∈B}∈E and
A:={g∈B}∩{f=g}⊆{f=g}∈N⊆E
by our assumptions. Thus, {f=g}={f=g}c∈E, and
{g∈B}=({f∈B}∩{f=g})∪A∈E
allows us to conclude that g is measurable.
The above property can fail if the measure space is not complete. Similar problems can arise in the study of stochastic processes, and therefore completeness assumptions are made. For this whole blog post, we shall place ourselves on a complete probability space (Ω,F,P) endowed with a filtration F={Ft}0≤t<∞, i.e., a family of sub-σ-algebras of F which is increasing in the sense that
Fs⊆Ft⊆F∞:=σu∈[0,∞)⋃Fu,0≤s≤t<∞.
We sometimes denote this setup conveniently as (Ω,F,F,P) and call it a filtered probability space. We shall assume that this filtration satisfies the usual conditions:
Right-continuity, i.e., Ft=Ft+:=⋂s>tFs for all t∈[0,∞), and
F0 contains all the P-negligible events in F.
The structure of the filtration implies that we can write Ft+ equivalently as ⋂n∈NFt+1/n.
Often we will need to start with the natural filtration FtX:=σ(Xs;0≤s≤t),t≥0, generated by the process X={Xt}t≥0, which is the smallest filtration with respect to which the process X is adapted, i.e., Xt is FtX-measurable for each t≥0. And then we will need to augment it to make it satisfy the usual conditions. To that end, we define the minimal augmented filtration generated by X to be the smallest filtration that is right continuous and complete and with respect to which the process X is adapted. This can be constructed in the following three steps:
First, let {FtX}t≥0 be the natural filtration as in (3).
Let N be the collection of all P-negligible sets for the complete probability space (Ω,F,P) (we can always complete it using Theorem 11 if it isn’t). For each t≥0, let (FtX)P={F∪N:F∈FtX and N∈N} be the completion of FtX just like in (2).
Finally, for each t≥0, let Ft:=((FtX)P)+=s>t⋂(FsX)P define the right-continuous filtration. To show right-continuity of {Ft}t≥0, we need to show that
s>t⋂(FsX)P=Ft=Ft+=r>t⋂s>r⋂(FsX)P.
But this is obvious since F∈(FsX)P for every s>r>t is same as saying F∈(FsX)P for every s>t.
These steps give a filtration {Ft}t≥0, which is right continuous and complete and with respect to which the process X is adapted. It is clear that it is also the smallest such filtration. Does it matter if we do step 3 before step 2, i.e., is it true that
((FtX)P)+=((FtX)+)P?
It would be very disappointing if they weren’t equal. Let us start by showing
((FtX)+)P⊆((FtX)P)+.
From (2), if F∈((FtX)+)P, then F=B∪N for some B∈(FtX)+ and N∈N. Then B∈FsX for every s>t, which means that B∪N∈(FsX)P, showing F∈(FsX)P for every s>t.
For the other side,
((FtX)P)+⊆((FtX)+)P,
let F∈((FtX)P)+. We need to construct B∈(FtX)+ and N∈N such that F=B∪N. By the definition of right-continuity, F∈(Ft+1/nX)P for every n∈N, which means there exist sequences {Bn}n∈N and {Nn}n∈N satisfying Bn∈Ft+1/nX,Nn∈N, and F=Bn∪Nn for each n∈N. Define
B:=n→∞liminfBn={Bn,eventually}=k∈N⋃n≥k⋂Bn.
Then note that ⋂n≥kBn∈Ft+1/kX, which implies that for each m∈N,
B=k∈N⋃n≥k⋂Bn=k≥m⋃n≥k⋂Bn∈Ft+1/mX.
This says that B∈(FtX)+. Define N:=F∖B, and note
For our purposes, a stochastic process is simply a collection of real-valued random variables X={Xt}t≥0 on (Ω,F). The sample paths of X are the mapping [0,∞)∋t↦Xt(ω)∈R obtained when fixing ω∈Ω.
We need terminology to talk about when two stochastic processes are the “same”.
Definition 19: Consider two stochastic process X={Xt}t≥0 and Y={Yt}t≥0 defined on the same probability space (Ω,F,P). We say
X and Y have the same finite-dimensional distributions if, for any integer n≥1, real numbers 0≤t1<t2<⋯<tn<∞, and A∈B(Rn), we have
P{(Xt1,…,Xtn)∈A}=P{(Yt1,…,Ytn)∈A};
Y is a modification or a version of X if, for every t∈[0,∞), we have P{Xt=Yt}=1;
X and Y are indistinguishable if almost all their sample paths agree, i.e.,
P{Xt=Yt∀t∈[0,∞)}=1.
The set {Xt=Yt∀t∈[0,∞)} is measurable because of our completeness assumption.
We will need to make stronger measurability assumptions for the random variables {Xt}t≥0 than just assuming that each Xt is measurable.
Definition 20: A stochastic process X={Xt}t≥0 is called
adapted, if Xt is Ft-measurable for every t∈[0,∞);
measurable, if the mapping [0,∞)×Ω∋(t,ω)↦Xt(ω)∈R is B([0,∞))⊗F-measurable, when R is endowed with its Borel σ-algebra;
progressively measurable, if for every t∈[0,∞) the mapping
[0,t]×Ω∋(s,ω)↦Xs(ω)∈R
is B([0,t])⊗Ft-measurable, when R is endowed with its Borel σ-algebra.
Recall that every progressively measurable process is both measurable and adapted; and every adapted process with right-continuous or with left-continuous paths, is progressively measurable (Karatzas and Shreve, 1998a), (Le Gall, 2016). On the other hand, we have the following result due to Meyer (1966).
Theorem 12: Every measurable and adapted process X={Xt}t≥0 has a progressively measurable modification Y={Yt}t≥0.
Consider the space L0 of equivalence classes of measurable functions f:Ω→R, and endow it with the topology of convergence in probability, for example with the pseudo-metrics ρ(f,g)=E(1∧∣f−g∣) or ρ(f,g)=E(1+∣f−g∣∣f−g∣). Recall that if a sequence {fn}n∈N⊆L0 satisfies ∑n∈Nρ(fn,fn+1)<∞, then it converges in probability “fast”, thus also almost surely.
Every process Y={Yt}t≥0 can be thought of as a mapping from [0,∞) into L0, which associates to each element in [0,∞) the equivalence class Yt of Yt. Consider now the collection H of processes Y such that the mapping [0,∞)∋t↦Yt∈L0 satisfies
it takes values in a separable subspace of L0;
under it, the inverse image of every open ball of L0 is a Borel subset of [0,∞); and
it is the uniform limit of a sequence of simple measurable functions with values in the space L0.
Then note that property 3 implies the other two, and that conversely, properties 1 and 2 together imply property 3, because a real-valued function is measurable if and only if it is the increasing limit of a sequence of simple measurable functions.
Next, note that H is a real vector space on account of property 3, and is closed under sequential pointwise convergence on account of properties 1 and 2. It is easy to see that H contains all processes of the form Y(t,ω)=K(t)ξ(ω), where K is the indicator of an interval in [0,∞) and ξ:Ω→R is a bounded measurable function. Thus, monotone class theorem implies H contains all bounded measurable processes. Now using a standard slicing argument it is easily seen that every measurable process Y has properties 1, 2 and 3.
Consider now a measurable and adapted process X={Xt}t≥0. For each n∈N, there exists a process X(n)={Xt(n)}t≥0 which is simple and measurable (in the sense that there exists a partition {Ak(n)}k∈N of [0,∞) into Borel sets, and a sequence of random variable {Hk(n)}k∈N, such that Xt(n)=Hk(n) for t∈Ak(n)), and satisfies
ρ(Xt,Xt(n))≤2−(n+1),∀t∈[0,∞).
Next, we define a sequence of random variables {Gk(n)}k∈N for each n∈N using the sequence {Hk(n)}k∈N to get “enhanced” measurability properties. To this end, define sk(n)=infAk(n), and define Gk(n)=X(sk(n)) if sk(n)∈Ak(n); if, on the other hand sk(n)∈/Ak(n), we pick any Fsk(n)−measurable Gk(n) that satisfies
ρ(Gk(n),Hk(n))≤2−(n+1).
Such a Gk(n) always exists since we may take a decreasing sequence {θm}m∈N⊆Ak(n) with θm↓sk(n), and set Gk(n)=liminfm→∞X(θm); and the above inequality holds for all integers k.
We now define the process Y(n)={Yt(n)}t≥0 by Yt(n):=Gk(n),t∈Ak(n), and check that it is progressively measurable and satisfies ρ(Xt,Yt(n))≤2−(n+1) for all t∈[0,∞) and n∈N.
Finally, we construct a progressively measurable process Y by
Yt(ω):=n→∞limYt(n)(ω),∀(t,ω)∈[0,∞)×Ω such that this limit exists,
and Yt(ω):=0 otherwise. The properties of ρ then implies P{Xt=Yt}=1 holds for all t∈[0,∞), so Y is a modification of X.
Definition 21: The σ-algebra of progressively measurable sets, denoted by P⋆, is the smallest σ-algebra on the product space [0,∞)×Ω, with respect to which the mappings as in (6) are measurable for all progressively measurable processes X.
It can be verified that P⋆ is the collection of all sets A∈F⊗B([0,∞)) such that the process Xt(ω)=1A(ω,t) is progressively measurable. This characterization gives the useful result that a subset A of [0,∞)×Ω belongs to P⋆ if and only if, for every t∈[0,∞), A∩([0,t]×Ω)∈B([0,t])⊗Ft.
Definition 22: A random variable T:Ω→[0,∞] is called a stopping time if {T≤t}∈Ft for all t∈[0,∞).
Because the filtration F is right-continuous this is equivalent to {T<t}∈Ft for all t∈[0,∞) as can be seen by writing
{T≤t}=q∈Q,t<q<s⋂{T<q}∈Fs.
It is easily checked that if T and S are stopping times, then so are
T∧S,T∨S and T+S.
Likewise, if {Tn}n∈N is a sequence of stopping times, then
nsupTn,ninfTn,n→∞limsupTn and n→∞liminfTn
are also stopping times. Stopping times go hand in glove with progressive measurability. Suppose X={Xt}t≥0 is F-progressively measurable and T an F-stopping . Then, the stopped process
A very important result states that every stopping time is the decreasing limit of a sequence of discrete stopping time. More concretely, if T is a stopping time, and we define
Tn(ω)={k2−nT(ω)on {2nk−1≤T<2nk}on {T=∞},
for n,k∈N, then each Tn is a discrete stopping time and we have Tn↓T. When can we approximate a stopping with an increasing sequence of stopping times? Of course, being able to do that means that our stopping time is a very special one. We give a name to such stopping times.
Definition 24: A stopping time T is called predictable, if there exists an increasing (“announcing”) sequence {Tn}n∈N of stopping times with T=limn∈N↑Tn, and Tn<T holds on the set {T>0} for all n∈N.
Note {T≤t}=⋂n∈N{Tn≤t}∈F(t) and therefore if in the definition above we had not required T to be a stopping time, it would still be a stopping time. A canonical example of a predictable stopping time is the first time Brownian motion W hits or exceeds a certain level b. It is announced by the sequence of first times W hits or exceeds the levels b−1/n, for n∈N. In fact, we have (taken from (Almost Sure blog)):
Theorem 13: Let X be a continuous adapted process and b be a real number. Then
T=inf{t∈[0,∞):X(t)≥b}
is a predictable stopping time.
Let
Tn=inf{t∈[0,∞):X(t)≥b−1/n},
which, by the Debut Theorem below (Theorem 14) is a stopping time. This gives an increasing sequence {Tn}n∈N of stopping times bounded above by T. Also, X(Tn)≥b−1/n whenever Tn<∞ and, by left-continuity, setting τ=limnTn gives X(τ)≥b whenever τ<∞. So T≥τ, showing that T=limn∈N↑Tn. If 0<Tn≤T<∞ then, by continuity, X(Tn)=b−1/n<b=X(T). So, Tn<T on the set {0<T<∞} and the sequence {Tn∧n}n∈N announces T.
Of course, in the above proof we cheated by assuming the Debut Theorem for stopping times, proving which is our next agenda. In the same flavor as Lemma-6 we have
Lemma 7: If T is a predictable stopping time and S an arbitrary stopping time, then
A∩{T≤S}∈F(S−) for all A∈F(T−);
A∩{S=∞}∈F(S−) for all A∈F(∞).
In particular, both the events {T≤S} and {T=S} belong to F(S−).
Definition 25: For any measurable mapping Z:Ω→[0,∞] and any A∈F, we define the restriction of Z on A as
ZA:=Z1A+∞1Ac.
Notice that we can write {ZA≤Z}=A∪(Ac∩{Z=∞}) as a disjoint union. If T is a stopping time and A∈F(T), then since {TA≤t}={T≤t}∩A∈F(t) for all t∈[0,∞), TA is also a stopping time.
Theorem 14 [Debut of a Progressive Set]: Under the usual conditions, if A is a progressively measurable set, i.e., A∈P⋆, then the debut DA is a stopping time.
Just like in the proof of Theorem 9, for any real number t>0, the set {DA<t} is the projection onto Ω of the set At:=A∩([0,t)×Ω). Recall that A∈P⋆ implies At∈B([0,t])⊗F(t). Therefore, Measurable Projection theorem (Theorem 7) implies {DA<t}=π(At)∈F(t). Right continuity of the filtration now implies DA is an F-stopping time.
The following important theorem now becomes an easy corollary.
Theorem 15 [First Hitting Times]: In the context above, if X is a progressively measurable process and Γ∈B(R), the first hitting time
HΓ:=inf{t∈[0,∞):X(t)∈Γ}
is a stopping time.
Since X is a progressively measurable process, the set A={(t,ω)∈[0,∞)×Ω:X(t,ω)∈Γ} is a progressively measurable set. Theorem 14 then implies the debut DA is a stopping time, but it is easy to see DA=HΓ.
Definition 26: The predictable σ-algebra, denoted by P, is the smallest σ-algebra on [0,∞)×Ω with respect to which left-continuous F-adapted processes are measurable. A stochastic process is said to predictable if it is P-measurable.
Definition 27: The optional σ-algebra, denoted by O, is the smallest σ-algebra on [0,∞)×Ω with respect to which right-continuous F-adapted processes are measurable. A stochastic process is said to optional if it is O-measurable.
Since all right-continuous or left-continuous adapted processes are progressively measurable, all optional or predictable processes are adapted. We need the concept of stochastic intervals to ease some notation and see some examples of optional and predictable sets.
Definition 28: If U,V:Ω→[0,∞] are two random variables, and U≤V, then define the stochastic intervals as
Notice that [[U]]=[[U,U]]=[[0,U]]∖[[0,U[[, and that 0A is predictable with the announcing sequence {0A∧n}n∈N. Denote by S the collection of all stopping times and by S(P) the collection of all predictable stopping times (S is S in fraktur font).
Lemma 8: We collect some useful properties of stochastic intervals in relation to optional and predictable σ-algebras.
If S,T∈S such that S≤T, then the stochastic intervals [[S,T]],[[S,T[[,]]S,T]],]]S,T[[ and the graphs [[S]],[[T]] are optional.
If S,T∈S such that S≤T, then the stochastic interval ]]S,T]] is predictable. If S∈S(P), then the stochastic interval [[S,∞[[ is predictable.
The σ-algebra O of optional sets can be generated from stochastic intervals in multiple equivalent ways
It is easy to see that 1[[S,T[[ is right-continuous and adapted, hence [[S,T[[ is optional. If Tn=S+1/n for n∈N, then [[S]]=⋂n∈N[[S,Tn[[, and thus [[S]] is also optional. Similarly, for [[T]]. These facts along with simple manipulation of sets immediately imply [[S,T]],]]S,T]],]]S,T[[ are also optional.
It is easy to see that 1]]S,T]] is left-continuous and adapted, hence ]]S,T]] is predictable.
It is a simple exercise to show
Lemma 9: Every predictable process is optional, and every optional process is progressively measurable. In other words
Recall the setting of section "Measurable Section" from part 1. We showed there that if A∈B([0,∞))⊗F, then we can define a measurable mapping T:Ω→[0,∞] such that for all ω∈π(A) we have (T(ω),ω)∈A, or stated more succinctly, [[T]]⊆A. Now, what if we want something stronger, say, T should be a stopping time? As we will see, we will have to make stronger assumptions on the set A and even then if we want T to be a stopping time we will have to let go of the nice property that {T<∞}=π(A).
But first, we need to prove a certain property of debuts which we will need in proving the aforementioned section theorem, and that's the goal of this section.
It isn't difficult to see that P coincides with the mosaic generated by the paving on [0,∞)×Ω that consists of finite unions of stochastic intervals of the form [[S,T]], where S,T∈S(P). Generalizing this, let us consider a collection A (A is A in fraktur font) of stopping times that contains 0 and ∞, and which is closed under a.s. equality and under finitely many lattice operations (∧ and ∨). We denote by J:={[[U,V[[:S,T∈A}, and by T the collection of finite unions of elements of J.
Simple observations like [[S,T[[c=[[0,S[[∪[[T,∞[[ or [[S,T[[∩[[U,V[[=[[S∨U,(S∨U)∨(T∧V)[[ can be used to show that J is a paving on [0,∞)×Ω and T is an algebra on [0,∞)×Ω.
For more structure, let us impose the following properties on the collection A:
For any pair S,T∈A, the stopping time S{S<T} (recall the notation of Definition 25 and the observation right after it) belongs to A;
For any increasing sequence {Sn}n∈N⊆A, the limit limn∈N↑Sn belongs to A.
Property 1 ensures that the debut of an element of T belongs to A. This is because S{S<T}=D[[S,T[[ and because the debut of union of a finite collection of elements from J is equal to to the minimum of the debuts of those elements.
Property 2 ensures that the debut of an element of Tδ (which, recall, equals {⋃n∈NAn:An∈T for all n∈N}) belongs to to A. This is the main result of this section, and is proved below as Theorem 16.
Coming out of the abstraction for a bit, we note that the collection S of all stopping times satisfies all the conditions imposed above on A. Similarly for the collection S(P) of all predictable stopping times. The first claim is trivial to see. To show the second claim, we first prove the following lemma.
Lemma 10: For any stopping time T and set A∈F(T), the random variable TA is a stopping time as we saw above. For TA to be a predictable stopping time, it is necessary that A∈F(T−); this condition is also sufficient if T is predictable.
If TA is a predictable stopping time, then since A={TA≤T}∖(Ac∩{T=∞}) and by Lemma 7 both {TA≤T},(Ac∩{T=∞})∈F(T−), we get A∈F(T−).
Conversely, suppose T is a predictable stopping time. Consider the collection
A:={A∈F(T−):TA and TAc are both predictable}.
It is a σ-algebra: Ω∈A clearly and it is closed under complements by definition; if A1,A2,…∈A with A=⋃n∈NAn, note TA=infn∈NTAn. Let {Tnm}m∈N be the announcing sequence for TAn for each n∈N, and define
τn:=min{Tjk:j,k≤n}.
Then {τn}n∈N announces TA and hence A∈A.
Thus, it suffices to show that TA is predictable for any A in a collection of sets that generates F(T−). To this end, suppose {Tn}n∈N announces T, and fix an arbitrary n∈N and an arbitrary A∈F(Tn). For every integer m≥n, the restriction TAm is a stopping time and so is Rm:=TAm∧m. The sequence {Rm}m∈N then announces TA, showing TA is predictable; similarly for TAc. We conclude TA,TAc are predictable for every A∈⋃n∈NF(Tn). But recall
F(T−)=σ(n∈N⋃F(Tn)),
since Tn<T on {T>0}, so we obtain the property for all A∈F(T−).
Coming back to the second claim above, if S,T∈S(P), then by Lemma 7, {S<T}∈F(S−), and thus by Lemma 10, S{S<T}∈S(P). The rest of the conditions are trivial to check.
Theorem 16: Suppose A is a collection of stopping times with the properties postulated above, and let B∈Tδ. Then the debut DB of this set is a stopping time in A, and its graph [[DB]]⊆B.
We first show that [[DB]]⊆B. If (s,ω)∈[[DB]], then DB(ω)=s which is same as saying
s=inf{t∈[0,∞):(t,ω)∈B}.
But we constructed our collection Tδ using stochastic intervals of the form [[S,T[[, and thus
B(ω)={t∈[0,∞):(t,ω)∈B}
is closed, implying (s,ω)∈B.
To show that DB∈A, consider the collection
B:={S∈A:S≤DB},
and note that B contains 0, is closed under pairwise maximization, i.e., S,T∈B implies S∨T∈B, and is closed under countable increasing limits, i.e., if {Sn}n∈N⊆B is an increasing sequence then limn∈N↑Sn∈B. Then B contains a representative T:=esssupB of its essential supremum (see (He, Wang and Yan, 1992), (Karatzas and Shreve, 1998b)).
Since B∈Tδ we can find a decreasing sequence {Bn}n∈N⊆T with ⋂n∈NBn=B. Define
Tn:=DBn∩[[T,∞[[ for all n∈N,
and note [[Tn]]⊆Bn for each n∈N. As Bn∩[[T,∞[[∈T and thus can be expressed as a finite union ⋃k=1m[[Uk,Vk[[ for Uk,Vk∈T, we have
Theorem 17: Under the usual conditions, let X be an adapted process with paths that are RCLL. The jumps of X are exhausted by a sequence of stopping times {Tn}n∈N,
A:={(t,ω)∈(0,∞)×Ω:X(t,ω)=X(t−,ω)}⊆n∈N⋃[[Tn]].
Write
A=n∈N⋃An for An:={(t,ω)∈(0,∞)×Ω:∣X(t,ω)−X(t−,ω)∣>1/n}.
For each n∈N, the set An has no accumulation point in [0,∞) because 1/n>0. Therefore, it is easy to verify
An=p∈N⋃[[DAnp]] for DAnp(ω):=inf{t∈[0,∞):#([0,t]∩An(ω))≥p},
where the notation #B denotes cardinaility of a set B.
We claim that all elements of the sequence {DAnp}p,n∈N are stopping times. This is sufficient to prove our desired result.
Start by noting that the process X is progressively measurable and the process (t,ω)↦X(t−,ω) is left-continuous and adapted, and thus predictable. Thus, the sets {An}n∈N are progressively measurable. Fix any arbitrary n∈N, and note that for p=1, the mapping
ω↦DAn1(ω)=inf{t∈[0,∞):∣X(t,ω)−X(t−,ω)∣>1/n}
is a stopping time by Theorem 14 (debut of a progressive set). Arguing by induction, for each p∈N the random variable DAnp+1 is also a stopping time, as it is the debut of the progressively measurable set
We are now ready to come back to proving section theorems. We start by proving a general section theorem.
Theorem 18 [General Section Theorem]: In the context of Theorem 16, let H be the σ-algebra generated by the algebra T. For every set A∈H and ε>0, there exists a stopping time Tε∈A with
[[Tε]]⊆A and P[π(A)]≤P(Tε<∞)+ε.
Fix a set A∈H and ε>0. Measurable Section Theorem (Theorem 10 in part 1) implies there exists a random variable Z:Ω→[0,∞] with [[Z]]⊆A and π(A)={Z<∞}, where recall π:[0,∞)×Ω→Ω is the canonical projection map. We denote by ν the measure on the measurable space ([0,∞)×Ω,H) defined by
∫[0,∞)×Ωf(t,ω)ν(dt,dω)=∫{Z<∞}f(Z(ω),ω)P(dω)
for every H-measurable f:[0,∞)×Ω→[0,∞). Notice that (Z(ω),ω)∈A for all ω∈{Z<∞}.
By first taking f to be the indicator 1A for A and then taking f to be the indicator 1[0,∞)×Ω for [0,∞)×Ω in the equation above, we see that
ν(A)=P(Z<∞)≡P[π(A)]=ν([0,∞)×Ω).
Thus, the measure ν is carried by the set A. Similarly, taking f to be 1B for any B∈H, we get
ν(B)=∫π(A)1B(Z(ω),ω)P(dω)≤P[π(A∩B)].
Choquet’s capacitability theorem (Theorem 6 in part 1) applied to the paving T and to the capacity
ν∗(C):=B∈H,C⊆Binfν(B),C⊆[0,∞)×Ω,
we obtain the existence of a subset Bε∈Tδ of A, such that
P[π(A)]=ν(A)≤ν(Bε)+ε≤P[π(A∩Bε)]+ε≤P[π(Bε)]+ε.
Now take Tε to be the debut DBε which by Theorem 16 is a stopping time in A.
Theorem 19 [Optional and Predictable Sections]: Let A be an optional (respectively, predictable) subset of the product space [0,∞)×Ω. For every ε>0, there exists a stopping time (respectively, a predictable stopping time) Tε with [[Tε]]⊆A and P[π(A)]≤P(Tε<∞)+ε.
Both statements follow from the General Section Theorem (Theorem 17) in conjunction with the observations before Lemma 10, by taking A to be either S or S(P).
As mentioned in the beginning of the section “Properties of Debuts”, we didn't get the nice equality [[T]]⊆A and {T<∞}=π(A) as we got in Measurable Section theorem, but instead obtained an approximation (7). The condition [[Tε]]⊆A makes sure that {Tε<∞}⊆π(A), and the P[π(A)]≤P(Tε<∞)+ε part ensures that the measure of the difference P[π(A)∖{Tε<∞}] of these two events can be made as small as desired.
To see that it is not always possible to choose a stopping time T that satisfies (8) if A∈O, we will construct a filtration {Ft}t≥0 and choose a set A∈O that forces π(A)∖{T<∞}=∅ for every stopping time T (the argument is from (Almost Sure blog)).
To this end, let τ:Ω→(0,∞) be a random variable such that P{τ<t}>0 for all t>0. For example, let τ be such that its distribution is uniform on (0,1). Let {Ft}t≥0 be the completed filtration such that Ft is generated by {{τ≤s}:s≤t}, and thus τ becomes a stopping time. Let A=]]0,τ[[, which we know from Lemma 8 belongs to O. Note that P{π(A)}=1 by our construction.
It is easy to see that Ft is trivial when restricted to {τ>t}, i.e., contains only sets of measure 0 or 1. So, every Ft-measurable random variable is a.s. constant on the event {τ>t}. Therefore, any stopping time T is deterministic on the event {T<τ}. So, if [[T]]⊆A, we have
T={s∞on {τ>s}on {τ≤s}
a.s. for some fixed s>0. But then
P{T<∞}=P{τ>s}<1=P{π(A)}.
A similar argument can be used to show that it is not possible to choose a predictable stopping time T that satisfies (8) if A∈P.
Recall the measurable graph theorem (Theorem 8 in part 1) which implies that a map T:Ω→[0,∞] is measurable if and only if its graph [[T]] is measurable. We also have the following neat result:
Theorem 20: A random variable T:Ω→[0,∞] is a stopping time (respectively, a predictable stopping time), if and only if its graph [[T]] is an optional (respectively, predictable) set.
The necessity follows from the characterization of O and P in Lemma 8.
In the optional case, sufficiency follows from Theorem 14 and the fact that O⊆P⋆.
For the sufficiency in the predictable case, suppose [[T]] is predictable. Apply Predictable Section theorem (Theorem 18) for A=[[T]] to construct a sequence {Tn}n∈N of predictable stopping times such that
[[Tn]]⊆[[T]] and P(T<∞)≤P(Tn<∞)+2−n,∀n∈N.
Replacing Tn by T1∨⋯∨Tn if necessary, we may assume that this sequence is increasing. It follows then that the limit
T=n∈Nlim↑Tn=n∈NsupTn
is predictable: if {Tnm}m∈N announces Tn, then
τn=max{Tjk:j,k≤n}
announces T.
Suppose X and Y are measurable processes such that
P{XT1{T<∞}=YT1{T<∞}}=1
for each F-measurable T:Ω→[0,∞]. Then what can we say about X and Y? We will show that X and Y are indistinguishable! To this end, let
F={(t,ω)∈[0,∞)×Ω:X(t,ω)=Y(t,ω)}.
It is measurable since X and Y are measurable and due to our completeness assumption. We want to show that P{π(F)}=0, so suppose to the contrary that P{π(F)}>0. Any F-measurable T:Ω→[0,∞] satisfying [[T]]⊆F must also satisfy XT1{T<∞}=YT1{T<∞}. Now Measurable Section theorem (Theorem 10) says there exists a random variable T:Ω→[0,∞] such that [[T]]⊆F and {T<∞}=π(F), and therefore
P{XT1{T<∞}=YT1{T<∞}}≥P{T<∞}=P{π(F)}>0,
contradicting our initial assumption.
Remarkably, we have similar results for optional and predictable processes.
Theorem 21: Suppose two optional (respectively, predictable) processes X and Y agree at finite stopping times (respectively, predictable stopping times), i.e.,
Then the two processes X and Y are indistinguishable, i.e.,
P{Xt=Yt∀t∈[0,∞)}=1.
Consider the optional case—the predictable case is similar. Just like above we have to show that for the optional set
F={(t,ω)∈[0,∞)×Ω:X(t,ω)=Y(t,ω)},P{π(F)}=0.
Suppose to the contrary that P{π(F)}=2ε>0 for some 0<ε≤1/2. Then the Optional Section theorem (Theorem 19) implies there exists a stopping time Tε with the properties
To motivate optional and predictable projections, let us start with a fundamental problem in filtering theory: Assume an underlying complete probability space (Ω,F,P). There is an underlying signal X={Xt}t≥0 which is modelled as a stochastic process and which we are interested in studying. Our observation process has noise and therefore instead of observing X we observe a process Y={Yt}t≥0 such that
Yt=ft(Xt,Wt),
for Wiener process W={Wt}t≥0. Therefore, it makes sense to define F={Ft}t≥0 to be the minimal augmented filtration generated by Y. The problem, then, is to compute an estimate for X based on the observable data at time t, viz. Ft.
We could look at Zt:=E(Xt∣Ft), as an estimate for Xt at each time t≥0. The process Z={Zt}t≥0 is, of course, adapted. However, since conditional expectation is defined only up to P-a.s., what version of Z should we choose? (9) does not fix the paths of the process Z, which requires specifying its values at the uncountable set of time in [0,∞). We would be very lucky if it were possible for us to choose a version of Z such that (9) holds not only for all t∈[0,∞) but also for all finite stopping times. And indeed this is possible! This is part of the statement of the optional projection theorem.
On the other hand, if we want an estimate of X based on observable data before time t, then our estimate would be
Zt=E(Xt∣Ft−)
where recall Ft−:=σ(⋃s<tFs). Again it is possible to choose a version of Z such that this equality holds not only for all t∈[0,∞) but also for all finite predictable stopping times, and this is part of the statement of the predictable projection theorem.
Theorem 22 [Optional and Predictable Projections]: Let X be a bounded, measurable (though not necessarily adapted) process.
There is a unique, modulo indistinguishability, optional process Xo, called optional projection ofX, that satisfies for all stopping times T∈S the identity
E(XT1{T<∞}∣FT)=XTo1{T<∞}.
There is a unique, modulo indistinguishability, predictable process Xp, called predictable projection ofX, that satisfies for all predictable stopping times T∈S(P) the identity
E(XT1{T<∞}∣FT−)=XTp1{T<∞}.
Remarks:
Taking expectations in (10) we get E(XT1{T<∞})=E(XTo1{T<∞}) for all stopping times T∈S. Now suppose (12) holds for all stopping times T∈S. Fix an arbitrary stopping time S∈S and a set A∈FS, and write (12) for T=SA to get
E(XS1{S<∞}1A)=E(XSo1{S<∞}1A).
But this immediately implies (10) for stopping time S. Therefore, requiring (10) to hold for all stopping times is equivalent to requiring (12) to hold for for all stopping times.
Similarly, requiring (11) to hold for all predictable stopping times is equivalent to requiring E(XT1{T<∞})=E(XTp1{T<∞}) to hold for all predictable stopping times.
Operators o and p are linear operators, i.e., for bounded, measurable processes X,Y and a,b∈R,
(aX+bY)o(aX+bY)p=aXo+bYo, and=aXp+bYp,
as can be easily verified by using the properties of conditional expectation. The formulation of (12) and (13) makes it obvious that (Xo)o=Xo and (Xp)p=Xp. This explains the terminology of “projection”.
[of Theorem 22] Uniqueness is immediate from Theorem 21, so let's focus on existence, first for optional projection and then for predictable projection.
We will employ the monotone class theorem for functions. We need a simple class of processes for which finding an optional projection is easy to summon. To this end, consider the processes of the form
Xt(ω)=1B(ω)1[u,v)(t),0≤u<v<∞ and B∈F.
We claim that the candidate for the optional projection is
Xto(ω):=Mt(ω)1[u,v)(t),
where M is the right-continuous version of the bounded, thus also uniformly integrable, martingale E(1B∣Ft) (see (Karatzas and Shreve, 1998a), (Le Gall, 2016) for the proof of existence of such a martingale; it is here that the usual conditions on the filtration F become crucial).
With these choices and an arbitrary stopping time T, the left-hand side of (12) becomes P{B∩{u≤T<v}}, whereas optional stopping theorem (Karatzas and Shreve, 1998a), (Le Gall, 2016) shows that its right-hand side is
E(E(1B∣FT)∣1[u,v)(T)).
Recalling now that T is FT-measurable shows that the two sides are equal.
Finally, use linearity and monotone class arguments to establish existence for arbitrary bounded, measurable X.
Similar ideas work in the predictable case. We first consider processes of the form
Xt(ω)=1B(ω)1(u,v](t),0≤u<v<∞ and B∈F.
We claim that the candidate for the predictable projection is
Xtp(ω):=Mt−(ω)1(u,v](t),
where Mt− is the left-continuous version of the martingale M defined previously. Again using FT−-measurability of T we see the two sides of (13) are equal. Monotone class arguments now allow us to finish the proof.
We end our discussion with a result on time change. We call it a “time change” because given an adapted increasing process A with right-continuous paths and A0=0, we imagine a clock which runs according to A in the sense that at time t this clock shows time At.
Theorem 23: Suppose A is an adapted increasing process with right-continuous paths and A0=0.
For any two bounded, measurable processes X and Y that satisfy E(XT1{T<∞})=E(YT1{T<∞}),∀T∈S, we have E∫0TXtdAt=E∫0TYtdAt,∀T∈S.
For any non-negative, RCLL and uniformly integrable martingale M, we have
E∫0TMtdAt=E(MTAT),∀T∈S.
Only the first part requires any real effort.
Introduce the time change
C(s):=inf{t≥0:At>s},
and check that C(s) is a stopping time for each s≥0. Recall that C(s) is increasing, right-continuous, and the “functional inverse” of At in the sense that
At=inf{s≥0:C(s)>t}.
Therefore, C gives us a way to read the correct time off our A-clock—if the A-clock shows time s, the actual time is C(s).
We also check that the properties
∫0∞g(At)dAt=∫0A(∞)g(u)du, and ∫0∞g(t)dAt=∫0A(∞)g(C(s))ds=∫0∞g(C(s))1{C(s)<∞}ds
hold for any Borel-measurable g:[0,∞)→[0,∞). It follows from these considerations and Fubini’s theorem, that
holds, with a similar expression also holding for Y.
The claim (15) now follows directly from our assumptions, in the case T=∞. For general T∈S we apply the above result to the increasing, right-continuous and adapted process
Dellacherie, Claude. Capacités et processus stochastiques, Springer-Verlag, 1972.
Dellacherie, Claude and Meyer, Paul-André. Probabilities and Potential, North-Holland Publishing Company, 1978.
He, Sheng-wu and Wang, Jia-gang and Yan, Jia-an. Semimartingale Theory and Stochastic Calculus, CRC Press, 1992.
Karatzas, Ioannis and Shreve, Steven. Brownian Motion and Stochastic Calculus, Graduate Texts in Mathematics Volume 113, Springer-Verlag New York, 1998.
Karatzas, Ioannis and Shreve, Steven. Methods of Mathematical Finance, Probability Theory and Stochastic Modelling Volume 39, Springer-Verlag New York, 1998.
Le Gall, Jean-François. Brownian Motion, Martingales, and Stochastic Calculus, Graduate Texts in Mathematics Volume 274, Springer International Publishing, 2016.