Research
I am broadly interested in theoretical machine learning, nonparametric statistics and probability theory. The following are some of the topics I am working on right now:
Suppose X and Y are Rd−valued random vectors on a probability space (Ω,F,P) with distributions μ:=P∘X−1 and ν:=P∘Y−1 respectively. It is often of interest to find, if it exists, a measurable map T:Rd→Rd such that the pushforward T#μ:=μ∘T−1 of μ under T equals ν. It is possible that no such map exists, for example, if μ=δx0 for some x0∈Rd and ν does not equal a Dirac measure, because for any measurable T:Rd→Rd the pushforward T#μ is δT(x0).
On the other hand, if multiple such maps exist, then which one should we prefer? In the Monge's formulation of optimal transport problem, we have a measurable cost function c:Rd×Rd→[0,∞], and we want to find an optimal map
T0∈T#μ=νargmin∫Rdc(x,T(x))dμ(x), provided, of course, that it exists. It is not obvious when such a map exists, however, the following refinement by
(McCann, 1995) of the result discovered independently by
(Knott and Smith, 1984) and by
(Brenier, 1987) tells us when an optimal map exists, assuming quadratic cost. As a piece of terminology, we call a set
A⊆Rdsmall if it has a Hausdorff dimension at most
d−1. So, for example, Lebesgue-negligible sets are small.
Theorem: Let
μ and
ν be two Borel probability measures on
Rd, with
d a positive integer. Suppose that
μ does not give mass to small sets, and that the cost function
c:Rd×Rd→[0,∞] is the squared Euclidean norm
c(x,y)=∥x−y∥2. Then there is exactly one (upto
μ−a.e. of course) measurable map
T:Rd→Rd such that
T#μ=ν, and moreover
T=∇φ for some convex function
φ:Rd→R. In statistical optimal transport, the explicit form of the measures μ or ν is unknown, and instead only independent samples from them are observed. To that end, let X1,X2,… be independent copies of X, and let Y1,Y2,… be independent copies of Y. Let μn:=n1∑i=1nδXi denote the empirical measure based on the sample (X1,…,Xn) of the first n observations, and likewise for νn.
Brenier, Y.: Decomposition polaire et rearrangement monotone des champs de vecteurs. C. R. Acad. Sci. Paris Ser. I Math., 305:805–808, 1987.
Knott, M. and Smith, C. S.: On the optimal mapping of distributions. Journal of Optimization Theory and Applications, 43(1):39–49, 1984.
McCann, R.J.: Existence and uniqueness of monotone measure-preserving maps. Duke Mathematical Journal, 80(2):309 – 323, 1995.
Santambrogio, F.: Optimal Transport for Applied Mathematicians. Progress in Nonlinear Differential Equations and Their Applications, vol. 87. Birkhäuser/Springer, Cham (2015). Calculus of variations, PDEs, and modeling.
Villani, C.: Topics in Optimal Transportation. Graduate Studies in Mathematics, vol. 58. American Mathematical Society, Providence (2003).
NeurIPS 2022, AISTATS 2022, ICLR 2022, NeurIPS 2021, AAAI 2021, ICML 2020
Fall 2021, Spring 2021 — Masters level course on introduction to machine learning