2018年12月29日星期六

Levy’s equivalence theorem

We often says that the convergence almost surely, convergence in probability and convergence in law cannot be mixed and the first one implies the second one, which implies the third one. But in some situation, these three claims can be the same. One very famous situation is the Levy's equivalence theorem : We set $(X_m)_{m \geq 1}$ a family of independent random variables and we define $S_n = \sum_{m=1}^n X_m$, then convergence in law can imply the convergence almost surely !

This seems a little incredible and such an amazing theorem does not often appear in a general textbook. Maybe it is for the reason that such an easy theorem requires sometimes technical proof. Here, I give a personal argument using many big theorems.

The key idea is to use the Kolmogorov's three-series theorem. It is an equivalent criteria for that $S_n$ converges almost surely. The three conditions are to define that $Y_m = X_m \mathbf{1}_{|X_m| \leq A}$ and
$$
\sum_{m=1}^{\infty} \mathbb{P}[|X_m|>A] < \infty,  \quad \sum_{m=1}^{\infty}\mathbb{E}[Y_m]\text{ converges }, \quad \sum_{m=1}^{\infty}\mathbf{Var}[Y_m] < \infty.
$$

For the first one, we put the random variables in Skorohod representation theorem, and say that one is necessary to say the fluctuation is not so often.

For the third one, we argue by absurd and put it in the Lindeberg Feller central limit theorem, to say that if the variance goes to infinite, then
$$
\frac{\sum_{m}^n Y_m - \mathbb{E}[Y_m]}{\sqrt{\sum_{m=1}^{n}\mathbf{Var}[Y_m]}} \Rightarrow \mathcal{N}(0,1),
$$
and we know $\sum_{m}^{\infty} Y_m$ also converges in law to some random variable. Thus, the normalization seems too stronger and we obtain that  $\frac{\sum_{m}^n  - \mathbb{E}[Y_m]}{\sqrt{\sum_{m=1}^{n}\mathbf{Var}[Y_m]}} \Rightarrow \mathcal{N}(0,1),$ which is a contradiction.

Finally, to prove that the expectation is finite, we still argue by absurd. We go back to the very basic definition to test it by a function. However, we have a intuition that $\sum_{m=1}^{\infty}X_m \simeq \sum_{m=1}^{\infty}Y_m$ which will go to infinite, because that its drift is infinite while its variance is finite. This is the situation of vague convergence.

So we prove the three conditions and prove this theorem.

2018年9月12日星期三

Analysis and PDE : Basic fact about evolution equation

Equation of the type evolution is one important topic in mathematics. Today,  I read some very theoretical basis of this subject.
$$
\partial_t f = \Delta f(x) + a(x) \cdot \nabla f(x) + c(x) f(x) + \int_{\mathbb{R}^d} b(y,x) f(y) dy
$$
with the condition that $a \in W^{1,\infty}, c \in L^{\infty}, b \in L^2(\mathbb{R^d \times R^d})$.
Although there are many equations (maybe of types  degenerate), but this model describes the equation of diffusion, of branching, of transport and also of mean-field, so it is already very large. Of course, we will talk about the  existence, uniqueness of this equation.

The main frame to talk about the evolution equation  is to design a Hilbert space $H$ and a Banach subspace $V \subset H$, the  advantage  is to use the duality of Hilbert space to get
$$
V \subset H = H' \subset V' .
$$
so we can talk about the weak solution of the equation and we use $|\cdot|, \|\cdot\|$ to  represent respectively the norm in $H$ and $V$.

One propose a good condition, similar to that of Lax-Milgram. If we denote by $L$ the generator, then we requires that
(1)$L : V \rightarrow V'$ is bounded.
(2)$L$ is coercive  + dissipative that $\langle Lg, g \rangle \leq - \alpha \|g\|^2 + b|g|^2$ for every $g \in V$.
This condition gives an inequality
$$
|g(T)|^2_H + 2 \alpha \int_0^T \|g(s)\|_V^2 ds \leq e^{2bT}|g_0|^2
$$
at first correct for a $g$ more regular and then we can pass to general function by approximation. This is the energy inequality, it says a lot of things and the most important the uniqueness and make sure that  the solution stays always in the function space $g \in L^{\infty}(0, T; H) \cap L^2(0,T;V) \cap H^1(0, T; V')$.

(One remark : the dissipative part just says  the Cauchy-Liptchitz condition. )

Later, we use a theorem : $ L^2(0,T;V) \cap H^1(0, T; V') \subset C([0,T];H)$. This makes sense when we do formally $g'$ and then we can manipulate as if it is regular.

The existence theory is a variant theorem of Lax-Milgram theory. When we discretize the problem, it becomes a variant elliptic equation in small discrete steps. Then, we do a sub-sequence of weak limit to get the solution of the equation. .

2018年8月18日星期六

A very non-rigorous introduction of differential geometry

During the travelling time on train, I review the content of differential geometry - a topic that I have learned many times and long time ago, but handle so little in hand. I believe that I have learned it for at least three times : first time in Hong Kong, second time at Fudan and third time at Polytechnique. Some times later, I finally get understood its main idea, so in this short blog, I attempt to give a very non-rigorous introduction.

The first step is sometimes the most difficult step : to well define the object of manifold $\mathcal{M}$. Generally speaking, it is something similar to $\mathbb{R}^d$. But how it looks like? The standard definition of a manifold uses the language of local coordinate and atlas, I have to say that a very visualization is to imagine that a RPG guy runs in a game, and local we have to draw the map as a surface. We require that the $C^k$ condition to make locally the bijection is OK and we can always to approximation properly. Then, the compatibility condition makes all the local map together and define a good manifold.

The mapping between the manifold is used to define the equivalent class. The notation is complex but the idea is simple : when we talking about the property of squares, circles etc, we never identify a specific square of circle since they have the same property. For a manifold, we should also say the same story. That's why we invent this idea of the mapping between different manifold.

For $p \in \mathcal{M}$, its local increment has $d$ directions induced by the local map, so it has naturally a tangent plane $T_p\mathcal{M}$. When we remove the point $p$, it becomes a tangent fiber $T\mathcal{M}$. But until now, we say nothing about the distance and volume on the manifold. We have to add a metric $g$ to the manifold $(\mathcal{M}, g)$, which is a symmetric matrix, or a $(0,2)$-tensor. $g = g_{i,j}dx^i \otimes dx^j$. Then the curves can be written as
$$
L = \int_a^b \sqrt{g(\frac{d\gamma}{dt}, \frac{d\gamma}{dt})} dt,
$$
and the volume can be written as
$$
\int f d\mathcal{M} = \int_{\mathbb{R}^d} f \sqrt{det(g)} dx_1 dx_2 \cdots dx_d.
$$
We can check that all these definitions do not depend on the choice of local coordinate.

All these definitions make the manifold similar to $\mathbb{R}^d$ except some more profound for curvature. It starts from vector field $X$, which is generally not commutative, so we have the Lie crochet $[X, Y] = XY - YX$. And some more complicated idea is the covariant derivative (connection) written as $\nabla_{X}Y$. They develop the idea for curvature, Ricci curvature and scalar curvature. I skip these part since they may be more useful for the expert of geometry.

In the last par of the lecture book, people use these language to study the theory of relativity. For example, main idea is to find possible $g$ that makes the physical model work by some minimum action. Interested readers can check the lecture note of MAT568 by Jeremie Szeftel.

2018年7月20日星期五

Circle packing and its applications

One topic of this year's lecture of Saint-flour is about the circle packing, a very nice theory abridging complex analysis, probability and graph theory. I will record some main results in the two week's lectures.

Drawing circle packing : finite and infinite

A circle packing is a collection of disk $P$ to represent a planar graph.$G = (V, E)$, where the center of disk represents the vertex of graph and two disks are tangent if and only if the two vertex are connected.

For a circle  packing, we could draw easily a graph, but could we construct a circle packing from any planar graph ? This is the first question. The question could be answered by two steps : in finite case, we could have an algorithm to solve it by minimizing the energy, which depending only on the radius and for given radius we could draw the graph. Moreover, this drawing unique up to Mobius transform if we add an infinite point to make the out face a triangle.

In infinite case, the question is completed. However,  if we suppose that the bounded degree condition, this is a direct result of compactness.

Reversible Markov chain and electric network

The main question of the topic is to study the recurrence of a reversible Markov chain. In fact, we could always realize it by a electric network by designing conductance on the edges. The recurrence of an infinite simple random walk on this graph (Markov chain) is equivalent to say  $R_{eff}(x \leftrightarrow \infty) = \infty$. Very intuitively, this means that the effective resistance is infinite so we could not go to infinite freely.

Two useful lemma may be Dirichlet and Thomson : the voltage is a harmonic function which minimizes the energy and the current is unit flow which minimize the energy. 

By the electric network, the problem of recurrence is transformed to a problem of analysis of graph.

He-Schramn theorem : Classification of by circle packing

However, it's not easy to analyze the effective resistance. But He-Schramm theorem tells a classification of the infinite graph with only bounded degree : hyperbolic, circle packing in unit disk and transient; or parabolic, circle packing in plane and recurrent.

The main idea is that the circle packing provides a new view point to see the graph. The radius of disk gives something more than the distance on the graph, and this geometric information matches well with the analysis of resistance. If the circle packing is in plane, it could have a log increment of the effective resistance.

Benjamini-Schramm theorem

Our final object should be that of UIPT, which is a random object and the bounded degree condition is missing. To go pass it, a more powerful tool of Benjamini-Schramm theorem is called. This says that we could always start from some limit graph and if the object studied is a local limit of the finite, uniform rooted graph of bounded degree, it's current.

The proof of this theorem is tricky and depends mainly on a very nice estimate called "magic" lemma. Personally, I believe that this lemma has something to do with the maximal inequality and Whitney covering.

Recurrence of UIPT

The final step to prove UIPT requires all the preparation. Moreover, we apply some surgery to replacing some big degree vertex by some trees. This modification makes the Benjamini-Schramm theorem work since it reduces the degree of vertex. Moreover, some quantitative analysis is also needed to calibrate the increment of effective resistance, for example, we should know that the distribution of degree is exponential.


A final remark is that this  seems a nice proof, but how we collect all the necessary result of probability, graph theory and complex analysis ? This is a nice question. And the connection between these three area could be more profound in all senses.

2018年6月12日星期二

Analysis and PDE : Soblev space - integer and fractional type and inequality of Soblev

The Soblev space $W^{k,p}(\Omega)$ may be one of the most important tool in the domain of PDE. Its interest comes from two sides : from the view-point of mathematics, it engages on a lot of nice technique in functional and harmonic analysis. From the view-point of physic, some thing like 
$$\int_{\Omega} |\nabla u(x)|^2 dx < \infty$$
could be interpreted as the finite energy. Although the researchers use the embedding theorem everyday, sometimes when we would some more information from the constants, it's not easy, especially when we talk about the fractional order one $W^{s,p}(\Omega)$. The most direct way to illustrate its delicate usage may be a note. However, here we concentrate on the inequality of Soblev and Poincare type instead of the one of Holder type.


$W^{k,p}(\Omega), W_0^{k,p}(\Omega), W^{k,p}(\mathbb{R}^d)$

The definition of $W^{k,p}(\Omega)$ is just say that its all kind of weak derivative belongs to space $L^p(\Omega)$. However, since the function $C^{\infty}_0(\Omega)$ isn't always dense in this space, we have to define another guy as the closure under the norm of the former one. This case has no problem when we talk about the space $W^{k,p}(\mathbb{R}^d)$. The more tricky problem is that $W^{k,p}(\Omega)$ isn't the subspace of $W^{k,p}(\mathbb{R}^d)$ ! Luckily, we have one extension theorem to define $Ext(u)$ such that
$$\Vert Ext(u) \Vert_{W^{k,p}(\mathbb{R}^d)} \leq C(d,p,k, \Omega)\Vert u \Vert_{W^{k,p}(\Omega)}$$
But this theorem will use something about the property of $\Omega$. So this gives a very intuitive remark that the estimates about $W_0^{k,p}(\Omega), W^{k,p}(\mathbb{R}^d)$ don't use the property of $\Omega$ while the estimate of $W^{k,p}(\Omega)$ will use.


$W^{k,p}(\mathbb{R^d})$ and inequality of Soblev-Poincare

For $W_0^{k,p}(\Omega), W^{k,p}(\mathbb{R}^d)$, two useful inequalities are the Soblev inequality that
$$\Vert u\Vert_{L^{p*}(\mathbb{R^d})} \leq C(d,p,k)  \Vert \nabla^{k}u\Vert_{L^{p}(\mathbb{R^d})}$$
where $p*$ is some critical exponent that $p* = \frac{dp}{d - pk}$. Another important inequality is the one of Poincare that
$$\Vert u\Vert_{L^{p}(\mathbb{R^d})} \leq C(d,p,k) d(\Omega) \Vert \nabla u\Vert_{L^{p}(\mathbb{R^d})}$$
while this time it will gives one time of factor of $d(\Omega)$. The proof starts from the dense subspace $C^{\infty}_0(\Omega)$ and use the expression of integral. So it also works for the type of inequality of $u - (u)_{\Omega}$.

The interest of Poincare is that when we define the weighted norm
$$\Vert u\Vert_{\underline{W}^{k,p}(\Omega)} = \sum_{\beta \leq k}|\Omega|^{\frac{|\beta| - |k|}{d}}\Vert \partial^{\beta} u \Vert_{L^p(\Omega)}$$
we could just consider the last term that
$$\Vert u\Vert_{\underline{W}^{k,p}(\Omega)}  \simeq \Vert  \nabla^{k}u \Vert_{\underline{L}^{p}(\Omega)}$$
However, when we use the Soblev inequality, we get some supplementary power of the size.


$W^{s,p}(\Omega), W_0^{s,p}(\Omega), W^{s,p}(\mathbb{R}^d)$

Secondary, we think about another norm that
$$[u]^p_{W^{s,p}(\Omega)} = \int\int_{\Omega \times \Omega} \frac{|u(x) - u(y)|^p}{|x-y|^{d + sp}}dx dy$$
We draw an analogue when we define it in different context. However, this one isn't a norm since we observe that after a difference of constant, this semi-norm doesn't change. So when we talk about its norm, we have to add the one of $L^p$. One remark is that this semi-norm only take in account of the small size fluctuation. Two trivial inequalities are that
$$0 < s \leq s' < 1, [u]_{W^{s,p}(\Omega)} \leq (d(\Omega))^{s'-s}[u]_{W^{s',p}(\Omega)}$$


$$0 < s  < 1, [u]_{W^{s,p}(\Omega)} \leq C(d,p,s)(d(\Omega))^{1-s}\Vert \nabla u \Vert_{L^{p}(\Omega)}$$
They are correct for $u \in W_0^{s,p}(\Omega), W^{s,p}(\mathbb{R}^d)$, otherwise we have to add the constant of $C(\Omega)$.These two inequality share the same spirit of Poincare inequality that we use the diameter to replace the derivative. However, to get better description about the weighted norm, we have to think about other type of inequality. After all, we see that transform between the integer order and fractional order is always a little messy, although that we believe that they share the similar property. (As how we define them properly.)


$W^{s,p}(\mathbb{R}^d) \hookrightarrow  L^{p*}(\mathbb{R}^d)$

This result looks like the one of Soblev inequality : for $p* = \frac{dp}{d - ps}$ and $u \in W^{s,p}(\mathbb{R}^d)$, we have 
$$\Vert u\Vert_{L^{p*}(\mathbb{R^d})} \leq C(d,p,k)  [u]_{W^{s,p}(\mathbb{R^d})}$$
however, its proof is very tricky which uses many technique like decomposition dyadic. Once it's correct, we could also well take the highest derivative as the main part of the weighted norm that
$$\Vert u\Vert_{\underline{W}^{s,p}(\Omega)}  \simeq [u]_{\underline{W}^{s,p}(\Omega)}$$




$W^{1,p}(\mathbb{R}^d) \hookrightarrow  W^{s,p*}(\mathbb{R}^d)$

The final theorem, where we neglect the context since it's always same, is that for $0 < s < 1, p* = \frac{dp}{d + (s-1)p} \Leftrightarrow s - \frac{d}{p*} = 1 - \frac{d}{p}$, we have that 
$$ [u]_{W^{s,p*}(\mathbb{R^d})} \leq C(d,p,s)\Vert \nabla u\Vert_{L^{p}(\mathbb{R^d})}$$
To prove it, one important tool is Hardy-Littlewood-Soblev inequality. So this completes the relation between fractional and integer order Soblev space.

Conclusion

We got a lot of inequality. How to remember it ? Ok, we only have to take one ingredient in heart. The weighted norm eliminate the constant in the Poincare and Holder inequality, so we could use the highest derivative and any $p$ that we like. However, if we use the weighted norm of different order, the effect of scale always exists.  

2018年5月21日星期一

Analysis and PDE : Harmonic function

$\Delta u = 0$ may be one of the most important function in the PDE since it appears many times in different context and has nice properties : well, we have to say that its beautiful property implies the interesting result in physics and our natural. Here, we recall some basic property and proof strategy of this topic.

Classical solution

Although the existence should be considered as the first property necessary for study the object, maybe it's nicer to give some interesting property at first. We suppose that the solution is of class $C^2$, then one important formula is Stokes formula that 
$$\begin{eqnarray*}\int_{\partial \Omega} F \cdot \nu  d\sigma = \int_{\Omega} \nabla \cdot F(x) dx \end{eqnarray*}$$
$$\int_{\partial \Omega} u \nu  d\sigma = \int_{\Omega} \nabla u(x) dx $$
we apply this and obtain that in a domain of harmonic function we have 
$$\int_{\partial \Omega} \nabla u(x) \cdot \nu d\sigma = \int_{\Omega} \Delta u(x) dx = 0 $$
Furthermore, we obtain some useful formula as Green formula, we obtain the mean value principle that 
$$u(x) = \frac{1}{|\partial B_R(x)|}\int_{\partial B_R(x) }u(y) dy$$
One remarkable result is that this property improve the regularity as $C^{\infty}$ and it also implies the function is harmonic. (The proof is similar by the convolution below). Since if we do derivative, we get 
$$0 = \frac{d}{dr} \int_{\partial B_1(x) }u(x + ry) d\sigma = \int_{\partial B_1(x) }\nabla u(x + ry) \cdot \nu d\sigma = \frac{d}{dr} \int_{ B_r(x) } \Delta u(y) dy$$

A second important property is the Liouville proeprty. It says that a bounded harmonic function is trivial and is constant. The proof uses the fact that all the derivative are also harmonic 
$$|\partial_i u| \leq \frac{1}{\omega_d R^d} \int_{\partial B_R} |u| d\sigma \leq \frac{N}{R}\sup|u| \rightarrow 0$$
and then we use the mean value principle to analysis its size.

A third important property is the maximum principle. Idea is simple : the maximum and minimum of the function can only be attended at its boundary. Use the maximum principle, we prove that the uniqueness of the Dirichlet problem.

Weak solution is also classical solution

One very famous theorem Weyl states that all the weak solution that 
$$\int_{\Omega} u(x) \Delta \phi(x) dx = 0$$
for the test function $\phi \in C_c^{\infty}(\Omega)$ is also a strong solution of class $C^{\infty}$. The proof is very classical : we do convolution $u_{\epsilon} = u \ast \psi_{\epsilon}$. Then we prove that this function satisfies the harmonic by weak relation. Finally,  we pass this limit to the mean value principle. 
$$u_{\epsilon}(x) = \frac{1}{| B_R(x)|}\int_{ B_R(x) }u_{\epsilon}(y) dy \rightarrow u(x) = \frac{1}{| B_R(x)|}\int_{ B_R(x) }u(y) dy$$
Since this property doesn't require the regularity, we reprove that is is strong solution.

Existence

Finally, we come back to the problem of existence. There are two ways to prove the existence and it works in some more general framework. 
First one is the Lax-Milgram theorem. It treat the problem as find the inverse of operator in some function space 
$$a(u, v) = L(v)$$
The second one is the variational formation that we treat the solution  as a minimum of 
$$J(u) = \frac{1}{2}\int_{\Omega} |\nabla u|^2 dx - \int_{\Omega} f u dx$$
One can also deduce the characterization from one to another.

2018年5月16日星期三

Analysis and PDE : Covering and decomposition lemmas - Vitali, Calderon-Zygmund, Whitney

Some estimations a priori are important in PDE, while they come from some functional inequality, and some of them come from harmonic analysis and require some combinatoric covering lemma to prove it. This is a subject that I have learned several years ago from Prof. Hongquan LI when I was in Fudan. Today I spent a whole day to review them.

Vitali covering

When we study the Hardy-Littlewood maximal function 
$$\mathcal{M} f(x) = \sup_{r > 0} \frac{1}{|B_r(x)|} \int_{B_r(x)} f(y) dy$$
We would like to prove that this operator is of type strong $(p,p)$. The strategy is to prove that it is weak $(1,1)$ and strong $(\infty, \infty)$ and then uses the interpolation inequality. However, the weak $(1,1)$ isn't very clear. One tool used in the proof is so called Vitali covering. It says that for a covering $\{B_i\}_{i \geq 0}$ we could abstract a disjoint sub-covering $\{B_{kj}\}_{j \geq 0}$such that 
$$\bigcup_j^{\infty} B_{kj} \subset \bigcup_i^{\infty} B_i \subset \bigcup_j^{\infty} 3B_{kj}$$
This nice structure makes some sub-additive into additive and is very  useful in the proof.

Calderon-Zygmund decomposition

Calderon-Zygmund decomposition could be seen as a dyadic version of maximal inequality. We start by the system of dyadic cubes 
$$\mathcal{D} = \bigcup_{k \in \mathbb{Z}} \mathcal{D_k}$$ 
where $\mathcal{D_k}$ is the collection of cubes of  length $2^{-k}$. Then for any point $x$, it belongs to only one cube in the system $\mathcal{D_k}$. And any two cubes could be disjoint or one included in another - they could not have non-trivial intersection and difference at the same time. This decomposition is used to study singular integral, but at first we could use it to build a generalized cubic version of maximal function defined as 
$$\mathcal{M}_Q f(x) = \sup_{x \in Q \in \mathcal{D}} \frac{1}{|Q|} \int_{Q} f(y) dy$$

The first fact is that any open set in $\mathbb{R}^d$ could be built by the disjoint union of cubes. Idea is simple, we give one cube for each point $x$ such that $x \in Q \subset \Omega$. Then we mesh them if one belongs to another.

We apply this idea to the open set 
$$A_{\lambda} = \{x | \mathcal{M}_Q f(x)  \geq \lambda \}$$
and for each point, we choose the largest $Q$ admit. This works once we suppose $f \in L^1(\mathbb{R}^d)$ where an infinite cubes make it null. Then we know that this gives a perfect partition of the open set.
$$A_{\lambda} = \bigsqcup_{i=1}^{\infty}Q_i$$
Moreover, we know that the average at each cube is less than $2^d \lambda$ if $f$ is positive by the stopping time property, and $ \mathcal{M}_Q f(x) < \lambda \text{ a.e }$ for the point outside $A_{\lambda}$ by the Lebesgue differential theorem.  

Whitney decomposition

Finally, we give a stronger version of Whitney decomposition. It is also a dyadic decomposition but has two more properties
1, The distance of one cube from $\Omega^c$ is comparable to its length.
2, If two cubes are neighbors, their lengths are comparable.
So this decomposition looks more balanced.

Its construction is a little tricky : it applies a "level by level peeling" to cover the level by the distance to the boundary. Then we do mesh. This make the decomposition always comparable by some distance.