Brackets, drift, and general hypoellipticity
Diffusion can directly differentiate in fewer directions than the dimension of the ambient space. A drift can move those directions, and commutators can reveal the derivatives missing from the principal quadratic form. The aim of this lesson is to turn that geometric observation into a regularity theorem for arbitrary distributions.
We use the product, adjoint and Sobolev mapping theorems in Symbols, operators and Sobolev scales, and proper quantization, elliptic parametrices and microlocal Sobolev tests in Detecting regularity without choosing coordinates. Throughout this lesson the calculus is the scalar \(S^m_{1,0}\) calculus. Its commutator formula is stated explicitly in Section 3. The remaining prerequisites are Fourier transformation, Hilbert space duality, distributions, the inverse function theorem, and existence, uniqueness and smooth dependence for ordinary differential equations. We prove the local integrability assertion used in the rank obstruction.
Bramanti's open survey [B] is a useful introduction to the different proof mechanisms for sums of squares; its main comparison treats the case without drift. Hörmander's original paper [H] establishes the general diffusion-and-drift theorem. Rothschild and Stein [RS] develop the finer geometry and stronger estimates. The estimate proved here is a general positive gain. The sharp degenerate model estimate can give a substantially larger gain in its own setting.
1. Moving a diffusion direction twice
In coordinates \((v,x,t)\), consider \[ L=\partial_v^2+v\partial_x+x\partial_t. \tag{1.1} \] Only \(v\) appears in the second-order part. Put \[ X=\partial_v,\qquad Y=v\partial_x+x\partial_t. \] Direct calculation gives \[ [X,Y]=\partial_x,\qquad [[X,Y],Y]=\partial_t. \tag{1.2} \] Thus the available fields and their commutators span all three directions at every point, including \(v=x=0\).
For the proof we assign weight one to a diffusion field and weight two to the drift. We add weights when forming a bracket. The three fields in (1.2), together with \(X\), have weights \(1,3,5\). We will obtain, on every compact set and for every real \(s\), \[ \|u\|_{s+1/16}+\|Xu\|_{s+1/32} \leq C_{s,K}\bigl(\|Lu\|_s+\|u\|_s\bigr), \quad u\in C_c^\infty(K). \tag{1.3} \] Here \(\|\cdot\|_s\) denotes the standard \(H^s\) norm after compact localization in the coordinate chart. These particular exponents follow from the general commutator proof; no optimality is asserted. More significantly, if \(Lu\) is smooth near a covector, then \(u\) is smooth there. Establishing that last assertion requires a regularization argument: an estimate for smooth test functions alone does not imply it.
2. The energy and the abstract hypotheses
Let \(X\subset\mathbb R^n\) be open and let \(L_0,L_1,\ldots,L_r\) be real smooth vector fields. Let \(c\) be any smooth complex function. Define \[ L=\sum_{j=1}^r L_j^2+L_0+c,\qquad P=-L. \tag{2.1} \] With the Euclidean density, \[ L_j^*=-L_j-\operatorname{div}L_j. \] Consequently \[ P=\sum_jL_j^*L_j+V-c,\qquad V=-L_0+\sum_j(\operatorname{div}L_j)L_j. \tag{2.2} \] The real smooth field \(V\) belongs to the module generated by the original fields, and conversely \(L_0\) belongs to the module generated by \(V,L_1,\ldots,L_r\).
We take the complex inner product to be linear in its first argument. Since \[ \operatorname{Re}(Vu,u) =-\frac12\int(\operatorname{div}V)|u|^2\,dx, \] (2.2) gives, on each \(K\Subset X\), \[ \sum_j\|L_ju\|_0^2 \leq \operatorname{Re}(Pu,u)+C_K\|u\|_0^2. \tag{2.3} \] In particular \(\operatorname{Re}(Pu,u)\geq-C_K\|u\|_0^2\). The drift and the complex zero-order term contribute only bounded multiplication to this real quadratic form.
Write \(\ell_j\) for the real symbol of \(-iL_j\). The order-two principal symbol of \(P\) is \[ p_2=\sum_j\ell_j^2. \tag{2.4} \] We will need the first-order operators with symbols \[ \partial_{\xi_\nu}p_2,\qquad \langle\xi\rangle^{-1}\partial_{x_\nu}p_2. \tag{2.5} \] Both satisfy (2.3) with a different constant. Indeed \[ \partial_{\xi_\nu}p_2 =2\sum_j(\partial_{\xi_\nu}\ell_j)\ell_j, \qquad \langle\xi\rangle^{-1}\partial_{x_\nu}p_2 =2\sum_j\langle\xi\rangle^{-1} (\partial_{x_\nu}\ell_j)\ell_j. \] The factors before \(\ell_j\) are order-zero symbols; the first factors are even smooth functions of \(x\). The composition formula represents each operator as a sum of order-zero operators applied to \(-iL_j u\), plus an order-zero remainder applied to \(u\). Order-zero \(L^2\) boundedness and (2.3) prove the assertion. In particular, if \(P^\nu\) has symbol \(\partial_{\xi_\nu}p_2\) and \(P_\nu\) has symbol \(\partial_{x_\nu}p_2\), then \[ \sum_\nu\|P^\nu u\|_0^2+ \sum_\nu\|P_\nu u\|_{-1}^2 \leq C_K\operatorname{Re}(Pu,u)+C'_K\|u\|_0^2. \tag{2.6} \]
The commutator argument applies beyond differential sums of squares. Its exact assumptions are as follows.
Abstract setting. Let \(P\in\Psi^2(X)\) be scalar and properly supported, with real principal symbol \(p_2\). Suppose its real quadratic form is bounded below by \(-C_K\|u\|_0^2\) on every compact set. Call a properly supported \(Q\in\Psi^1(X)\) an energy operator if its principal symbol is real and \[ \|Qu\|_0^2 \leq A_K\operatorname{Re}(Pu,u)+B_K\|u\|_0^2 \quad(u\in C_c^\infty(K)), \tag{2.7} \] with \(A_K,B_K\geq0\). Assume that the operators in (2.5) are energy operators. Denote this family of all energy operators by \(\mathcal E\), and put \[ P'=\frac{P+P^*}{2},\qquad T=\frac{P-P^*}{2i}. \tag{2.8} \] The reality of \(p_2\) implies \(T\in\Psi^1\); \(T\) is self-adjoint.
Every real-principal order-zero operator belongs to \(\mathcal E\). Adding an order-zero operator to an energy operator preserves (2.7), after changing its constants. We can therefore replace each \(Q\in\mathcal E\) by its self-adjoint part. We do so in the estimates below. All resulting changes in first-order commutators have order zero.
For a fixed argument on a compact set, add a sufficiently large real constant to \(P\) so that \[ (P'v,v)\geq\|v\|_0^2,\qquad E(v)=(P'v,v)^{1/2}. \tag{2.9} \] This choice is made on a compact set containing the supports of all the finitely many operators and test functions used in that argument. Proper support permits such a set. The constant changes neither \(T\), the principal symbols, nor the commutators. It only replaces \(\|Pu\|_0+\|u\|_0\) by an equivalent bound. We have \[ \|Qv\|_0\leq C E(v),\qquad |(P'v,w)|\leq E(v)E(w). \tag{2.10} \] The second inequality follows by applying positivity to \(v+\lambda w\) for arbitrary complex \(\lambda\), and minimizing the resulting quadratic polynomial. Thus it remains valid even if \(P'\) is not elliptic.
3. The commutator that preserves the energy factors
For scalar symbols of orders \(m\) and \(a\), the ordinary composition formula gives \[ [A,B]\in\Psi^{m+a-1},\qquad \sigma([A,B])=\frac1i\{\sigma(A),\sigma(B)\}, \] where \[ \{f,g\}=\sum_\nu \bigl(\partial_{\xi_\nu}f\,\partial_{x_\nu}g -\partial_{x_\nu}f\,\partial_{\xi_\nu}g\bigr). \tag{3.1} \] The remainder after this first commutator term has order \(m+a-2\). Matrix symbols would have an additional leading matrix commutator; scalarity is used here.
Choose the \(2n\) energy operators \(F_j\) in (2.5), and, when desired, append any finite list of other energy operators. If \(A\in\Psi^a\), then \[ [P,A]=\sum_{j=1}^{2n}B_jF_j+B_0,\qquad B_j,B_0\in\Psi^a. \tag{3.2} \] To see this, the degree-\((a+1)\) commutator term is \[ \frac1i\sum_\nu \left((\partial_{x_\nu}\sigma(A))\partial_{\xi_\nu}p_2 -\bigl(\langle\xi\rangle\partial_{\xi_\nu}\sigma(A)\bigr) \bigl(\langle\xi\rangle^{-1}\partial_{x_\nu}p_2\bigr)\right). \] Its two coefficients have order \(a\). Quantizing them to the left of the corresponding \(F_j\) gives (3.2), because every discrepancy has order \(a\). The order-one part of \(P\), including \(iT\), has commutator with \(A\) of order \(a\), so it belongs to \(B_0\). The same factorization holds for \([P',A]\): subtracting \([iT,A]\) changes only the order-\(a\) remainder.
Abbreviate \[ R(u)=\|Pu\|_0+\|u\|_0. \] Equations (2.9)–(2.10) give \(E(u)\leq C R(u)\), and (3.2) gives \[ \|[P,A]u\|_{-a}+\|[P',A]u\|_{-a}\leq C_A R(u). \tag{3.3} \] If \(a+s\leq0\), the same estimates and \(PAu=APu+[P,A]u\) give \[ \|PAu\|_s\leq C_{A,s}R(u). \tag{3.4} \]
Lemma 3.1 (energy after an auxiliary operator). If \(B\in\Psi^b\) and \(\|Bu\|_b\leq C R(u)\), then \[ E(Bu)\leq C_B R(u). \tag{3.5} \]
Proof. Expand the quadratic form: \[ E(Bu)^2 =\operatorname{Re}(Pu,B^*Bu) +\operatorname{Re}([P,B]u,Bu). \] Since \(B^*:H^b\to L^2\), the first term has absolute value at most \(C R(u)^2\). By (3.3), \([P,B]u\in H^{-b}\) has norm at most \(C R(u)\); pair it with \(Bu\in H^b\). This bounds the second term by the same quantity. Positivity now proves (3.5). ∎
We will freely use proper elliptic operators \(\Lambda_s\) with principal symbol \(\langle\xi\rangle^s\) to measure Sobolev norms on a fixed compact set. More precisely, \[ \|v\|_s\leq C\bigl(\|\Lambda_s v\|_0+\|v\|_{-M}\bigr) \tag{3.6} \] for any fixed sufficiently large \(M\), and the reverse bound follows from Sobolev mapping. This is the ordinary elliptic parametrix theorem. They can be obtained by properly supporting the Fourier multipliers. The support cutoff changes their kernels by a smooth kernel; on compact input and output sets its Fourier transform decreases faster than any power, so it maps every \(H^{-M}\) to every \(H^s\). Outside a fixed output neighborhood the same kernel estimate gives the corresponding Sobolev norm bound.
In Section 4 the inputs \(v\) in (3.6) are first-order operators applied to \(u\); their \(H^{-M}\) norms are bounded by \(C\|u\|_0\) when \(M\geq1\). Thus all properization errors are bounded by \(C R(u)\), and all errors in squared pairings by \(C R(u)^2\). This explains precisely the harmless remainders in the norm computations below.
4. What each bracket costs
Say that a real-principal \(q\in\Psi^1\) has gain \(a\), where \(0<a\leq1\), if \[ \|qu\|_{a-1}\leq C R(u). \tag{4.1} \] Every energy operator has gain one. We first prove that the drift has gain one half, then prove the two bracket rules.
Lemma 4.1. The self-adjoint operator \(T\) in (2.8) has gain \(1/2\).
Proof. Put \(A=\Lambda_{-1/2}^*\Lambda_{-1/2}T\), of order zero. Then \[ \|\Lambda_{-1/2}Tu\|_0^2=(Tu,Au). \] The identity \(2iT=P-P^*\) gives \[ 2i(Tu,Au)=(Pu,Au)-(u,PAu). \] Both \(Au\) and \(PAu\) are controlled: order-zero mapping gives \(\|Au\|_0\leq C\|u\|_0\), and (3.4) gives \(\|PAu\|_0\leq C R(u)\). Taking absolute values and using (3.6) proves \(\|Tu\|_{-1/2}\leq C R(u)\). ∎
We need one elementary way to move a known gain past an auxiliary operator. If \(q\) has gain \(a\), \(A\in\Psi^t\), and \(t+s\leq a-1\), then \[ \|Aqu\|_s+\|qAu\|_s\leq C R(u). \tag{4.2} \] The first assertion is Sobolev mapping applied to (4.1). For the second, use \(qA=Aq+[q,A]\), where \([q,A]\) has order \(t\). Since \(t+s\leq a-1\leq0\), this error maps \(L^2\) to \(H^s\).
Lemma 4.2 (bracket with diffusion). If \(Q\) is an energy operator and \(q\) has gain \(a\), then \(i[Q,q]\) has gain \(a/2\).
Proof. Replace \(Q,q\) by self-adjoint operators with the same principal symbols; the changes are order zero and satisfy every gain under consideration. Put \[ C=i[Q,q],\qquad A=\Lambda_{a/2-1}^*\Lambda_{a/2-1}C. \] The order of \(A\) is \(a-1\). Expanding the commutator and taking absolute values gives \[ \|\Lambda_{a/2-1}Cu\|_0^2 \leq |(qu,QAu)|+|(Qu,qAu)|. \tag{4.3} \] We know \(qu\in H^{a-1}\) and \(Qu\in L^2\), with norms bounded by \(C R(u)\). Also \[ QAu=AQu+[Q,A]u\in H^{1-a}, \] with the same bound, since both \(A\) and \([Q,A]\) have order \(a-1\). Finally (4.2), with \(t=a-1,s=0\), controls \(qAu\) in \(L^2\). Pair the first term in (4.3) between \(H^{a-1}\) and \(H^{1-a}\), and the second in \(L^2\). Equation (3.6) proves the claimed gain. ∎
Lemma 4.3 (bracket with drift). If \(q\) has gain \(a\), then \(i[T,q]\) has gain \(a/4\).
Proof. Make \(q\) self-adjoint and put \[ C=i[T,q],\qquad A=\Lambda_{a/4-1}^*\Lambda_{a/4-1}C. \tag{4.4} \] Now \(A\) has order \(a/2-1\). Two applications of (4.2) give \[ \|qAu\|_{a/2}+\|A^*qu\|_{a/2}\leq C R(u). \tag{4.5} \] Indeed the sum of the order of \(A\) and the requested output order is \(a-1\). Each operator \(qA,A^*q\) has order \(a/2\), so Lemma 3.1 gives \[ E(qAu)+E(A^*qu)\leq C R(u). \tag{4.6} \] The exact identity obtained from \(2iT=P-P^*\), \(P^*=2P'-P\), and self-adjointness of \(q\) is \[ (Cu,Au) =(qu,P'Au)-(qu,PAu) +(P'u,qAu)-(Pu,qAu). \tag{4.7} \] For example, expand \(([P-P^*,q]u,Au)\), transfer \(P\) and \(P^*\) to the second argument and \(q\) to the second argument, then divide by two; this gives (4.7) with the stated inner product convention.
The second term is bounded by \(C R(u)^2\): \(qu\) is controlled in \(H^{a-1}\) and (3.4) controls \(PAu\) in \(H^{1-a}\). The latter use is valid because \[ (a/2-1)+(1-a)=-a/2\leq0. \] The fourth term is bounded in \(L^2\) using (4.5). For the first term, commute \(P'\) through \(A\): \[ (qu,P'Au) =(A^*qu,P'u)+(qu,[P',A]u). \] The first summand is at most \(E(A^*qu)E(u)\). The second is at most \(C R(u)^2\), because (3.3) bounds the commutator in \(H^{1-a/2}\), hence in \(H^{1-a}\). The third term of (4.7) is at most \(E(u)E(qAu)\). Equations (4.6) and (2.10) control both positive-form pairings. The left side of (4.7) is \(\|\Lambda_{a/4-1}Cu\|_0^2\). Use (3.6) to finish. ∎
This proof also shows why treating the drift as an energy operator would be incorrect: the drift step uses the positive form twice and gives \(a/4\).
5. Weighted brackets give a Sobolev estimate
Give each member of \(\mathcal E\) weight one and give \(T\) weight two. A bracket word is formed from these generators with the operation \(i[A,B]\); its weight is the sum of the weights of its letters, including repetitions. Its principal symbol is real and it has order at most one. Finite sums and order-zero changes do not affect any of the estimates.
Proposition 5.1. A bracket word of weight \(k\) satisfies \[ \|Cu\|_{\gamma_k-1}\leq C_{C,K} R(u), \qquad \gamma_k=2^{1-k}. \tag{5.1} \] The same holds with any smaller positive gain.
Proof. Every bracket word is a linear combination of brackets nested with a generator on the outside. This follows from the Jacobi identity \[ [[A,B],C]=[A,[B,C]]-[B,[A,C]], \] by repeatedly reducing the number of letters in the outer factor that is not a generator. This process terminates, preserves total weight, and applies as well to \(i[A,B]\). For an energy generator, (5.1) is (2.7); for the drift, it is Lemma 4.1. Adding an outside energy generator halves the available gain by Lemma 4.2, which changes \(2^{1-k}\) to \(2^{1-(k+1)}\). Adding an outside drift quarters it by Lemma 4.3, which changes it to \(2^{1-(k+2)}\). Induction proves the result. Sobolev monotonicity gives every smaller gain. ∎
Theorem 5.2 (general bracket estimate). In the abstract setting of Section 2, suppose that bracket words of weight at most \(N\) have no common characteristic covector over \(X\). Explicitly, for each compact \(K\), finitely many such words \(C_1,\ldots,C_M\) can be chosen, with order-one principal symbols \(c_j\), so that \[ \sum_{j=1}^M|c_j(x,\xi)|^2\geq b_K|\xi|^2 \tag{5.2} \] on a neighborhood of \(K\), for sufficiently large \(|\xi|\). Let \(0<\varepsilon\leq2^{-N}\). Then for every real \(s\), every compact \(K\), and every finite list of energy operators \(F_j\) containing (2.5), \[ \|u\|_{s+2\varepsilon} +\sum_j\|F_ju\|_{s+\varepsilon} \leq C_{s,K}\bigl(\|Pu\|_s+\|u\|_s\bigr), \quad u\in C_c^\infty(K). \tag{5.3} \] In particular the derivative operators in (2.6) satisfy \[ \|u\|_{s+2\varepsilon} +\sum_\nu\|P^\nu u\|_{s+\varepsilon} +\sum_\nu\|P_\nu u\|_{s+\varepsilon-1} \leq C_{s,K}\bigl(\|Pu\|_s+\|u\|_s\bigr). \tag{5.4} \]
Proof. The operator \(\sum_j C_j^*C_j\) is elliptic of order two near \(K\). Apply its ordinary local parametrix and the Sobolev mapping theorem to get \[ \|u\|_{2\varepsilon} \leq C\left(\sum_j\|C_ju\|_{2\varepsilon-1}+\|u\|_0\right). \] For each word its gain in (5.1) is at least \(2^{1-N}\geq2\varepsilon\). Thus \[ \|u\|_{2\varepsilon}\leq C R(u). \tag{5.5} \] The smoothing remainder of the elliptic parametrix maps \(L^2\) to \(H^{2\varepsilon}\), as required.
We now lift (5.5) and the energy bounds together; lifting the first estimate alone would leave an uncontrolled commutator. Write \[ U=\|u\|_{s+2\varepsilon},\quad V_1=\sum_j\|F_ju\|_{s+\varepsilon},\quad W=\sum_j\|F_ju\|_s,\quad R_s=\|Pu\|_s+\|u\|_s. \] Apply (5.5) to \(\Lambda_su\). Factor \([P,\Lambda_s]\) by (3.2), and use Sobolev mapping and ellipticity of \(\Lambda_s\). This gives \[ U\leq C(R_s+W). \tag{5.6} \] All lower norms from (3.6) are bounded by \(C\|u\|_s\) by choosing their order below \(s\).
For a compactly supported smooth \(v\), (2.7) and Sobolev duality give \[ \|F_jv\|_0^2 \leq C\|Pv\|_{-\varepsilon}\|v\|_\varepsilon +C\|v\|_0^2. \] Taking square roots and applying the scalar inequality \(\sqrt{ab}\leq B a+(4B)^{-1}b\) yields, for any \(B>0\), \[ \sum_j\|F_jv\|_0 \leq C B\|Pv\|_{-\varepsilon} +C B^{-1}\|v\|_\varepsilon+C\|v\|_0. \tag{5.7} \] Take \(v=\Lambda_{s+\varepsilon}u\). The factorization (3.2) gives \[ \|Pv\|_{-\varepsilon}\leq C(R_s+W). \] Commuting \(F_j\) past \(\Lambda_{s+\varepsilon}\) gives an error of order \(s+\varepsilon\), controlled by \(\|u\|_{s+\varepsilon}\). Therefore \[ V_1\leq C B(R_s+W)+C B^{-1}U+C\|u\|_{s+\varepsilon}. \tag{5.8} \] Choose \(B\) large enough that the coefficient of \(U\) here is at most \(1/2\), and combine (5.6)–(5.8).
For any \(\eta>0\), Fourier weights give \[ \|f\|_s\leq\eta\|f\|_{s+\varepsilon} +C_\eta\|f\|_{s-1},\qquad \|u\|_{s+\varepsilon}\leq\eta U+C_\eta\|u\|_s. \tag{5.9} \] For the first inequality split at a radius where \(\langle\xi\rangle^{-\varepsilon}\leq\eta\), bounding the remaining bounded frequency region by the lower norm; the second follows in the same way. Since each \(F_j\) has order one, \(\|F_ju\|_{s-1}\leq C\|u\|_s\). Hence \(W\leq\eta V_1+C_\eta\|u\|_s\). Choose \(\eta\) after \(B\), small enough to absorb both \(U\) and \(V_1\). This proves (5.3).
Lastly \(P_\nu\) differs from an order-one elliptic weight applied to the corresponding energy operator in (2.5) by an order-one remainder. That remainder maps \(H^{s+\varepsilon}\) to \(H^{s+\varepsilon-1}\), and (5.3) controls \(\|u\|_{s+\varepsilon}\). Equivalently, quantize \(\langle\xi\rangle^{-1}\) to the left of \(P_\nu\), then invert that weight. This proves (5.4). ∎
6. From test functions to arbitrary distributions
A distribution \(u\) belongs to \(H^s\) at \(\gamma=(x_0,\xi_0)\), \(\xi_0\ne0\), if some proper order-zero operator elliptic at \(\gamma\) sends \(u\) to \(H^s_{\mathrm{loc}}\). This agrees with the Fourier cone definition by the microlocal Sobolev tests stated in the prerequisites. The regularity need only hold in a small conic neighborhood of \(\gamma\).
Theorem 6.1. Under the hypotheses of Theorem 5.2, if \(Pu\) belongs to \(H^s\) at \(\gamma\), then \[ u\in H^{s+2\varepsilon}\text{ at }\gamma,\qquad F_ju\in H^{s+\varepsilon}\text{ at }\gamma. \tag{6.1} \] Thus for every distribution on \(X\), \[ \operatorname{WF}(Pu)=\operatorname{WF}(u). \tag{6.2} \]
Proof. First suppose that \(Pu,u,F_ju\) are all \(H^t\) in one open conic neighborhood \(\Gamma\) of \(\gamma\). We prove an improvement on a smaller neighborhood. Choose a compactly based symbol \(\psi\in S^0\) supported in \(\Gamma\) at large frequency, equal to one on a smaller neighborhood of \(\gamma\). Let \(\chi\in C_c^\infty(\mathbb R^n)\) equal one near zero. For \(0<h\leq1\), properly quantize \[ \psi_h(x,\xi)=\psi(x,\xi)\chi(h\xi), \quad \Psi_h=\operatorname{Op}(\psi_h). \tag{6.3} \] Each \(\Psi_h\) is smoothing and has compact output support. Its symbols are bounded in \(S^0\), independently of \(h\): a frequency derivative of \(\chi(h\xi)\) contributes \(h^{|\alpha|}\), which is at most \(C_\alpha\langle\xi\rangle^{-|\alpha|}\) on its annular derivative support. The properization remainders have uniformly bounded smoothing seminorms.
The factorization (3.2), with \(a=0\), gives \[ P\Psi_hu=\Psi_hPu+\sum_jB_{j,h}F_ju+B_{0,h}u, \tag{6.4} \] where \(B_{j,h},B_{0,h}\) are uniformly order zero and their nonsmoothing parts are supported in \(\Gamma\). This last statement can be seen directly in the full symbol expansion: every coefficient contains a derivative of \(\psi_h\). The smoothing remainder can be made of arbitrarily negative order, uniformly in \(h\), by retaining enough terms.
To apply (6.4) to the microlocal assumptions, insert an order-zero cutoff equal to one on a slightly larger conic neighborhood of the supports of these coefficients. Its applied distributions are \(H^t\). The complementary pieces are uniformly smoothing, because the two microsupports are separated. A compactly localized distribution belongs to some \(H^{-M}\): its Fourier transform has at most polynomial growth, and a sufficiently negative Sobolev weight is integrable. Such a uniform smoothing operator therefore sends it to a uniformly bounded set in \(H^t\). This proves \[ \|P\Psi_hu\|_t+\|\Psi_hu\|_t\leq C \tag{6.5} \] with \(C\) independent of \(h\).
Apply (5.3) to the smooth compactly supported \(\Psi_hu\). It follows that \(\Psi_hu\) is bounded in \(H^{t+2\varepsilon}\), and \(F_j\Psi_hu\) in \(H^{t+\varepsilon}\). A bounded sequence in each of these Hilbert spaces has a weakly convergent subsequence. On the other hand \(\Psi_hu\to\Psi u\) distributionally, where \(\Psi\) is the proper quantization of \(\psi\), since the symbols converge locally and the compact distribution Fourier estimates justify the limit in pairings. Uniqueness of the distributional limit gives \[ \Psi u\in H^{t+2\varepsilon},\qquad F_j\Psi u\in H^{t+\varepsilon}. \] Because \([F_j,\Psi]\) has order zero, its action is \(H^t\) on the original neighborhood. On the smaller neighborhood where \(\psi=1\), that commutator is smoothing: all differentiated cutoff symbols vanish there. A further elliptic cutoff therefore gives \(F_ju\in H^{t+\varepsilon}\) at \(\gamma\). We have proved the simultaneous improvement.
For an arbitrary \(u\), compact localization gives \(u,F_ju\in H^{t_0}\) for some finite negative \(t_0\leq s\). Starting there, the simultaneous improvement raises the known order of all these distributions by at least \(\varepsilon\) on a smaller cone, while \(Pu\) remains \(H^s\). Only finitely many steps are needed to reach \(s\). Choose a finite nested family of conic neighborhoods in advance, one for each step, so the uniform smoothing bounds above apply at each step. Once \(u,F_ju\in H^s\) at \(\gamma\), one final application with \(t=s\) proves (6.1).
If \(Pu\) is smooth at \(\gamma\), it is \(H^s\) there for every \(s\); (6.1) makes \(u\) smooth there. Conversely every pseudodifferential operator in this ordinary calculus is pseudolocal. This proves (6.2). ∎
The regularizer is used only to justify applying a test-function estimate to an initially rough distribution. One cannot simply assume \(\Psi_hu\to u\) in the stronger Sobolev space one is trying to prove.
7. The vector-field theorem and a rank obstruction
Theorem 7.1 (diffusion and drift). If the values of the Lie algebra generated by \(L_0,L_1,\ldots,L_r\) span \(T_xX\) at every \(x\in X\), then the operator (2.1), with arbitrary smooth complex \(c\), satisfies \[ \operatorname{WF}(Lu)=\operatorname{WF}(u) \quad\text{for every }u\in\mathcal D'(X). \tag{7.1} \] For every compact \(K\) there is some finite \(N\) for which (5.3)–(5.4) hold, with \(P=-L\), \(0<\varepsilon\leq2^{-N}\), and the diffusion fields included among the controlled first-order operators.
The wavefront assertion also holds on any smooth manifold without boundary. The estimates then mean their coordinate-local versions on compact sets.
Proof. The energy and derivative hypotheses were proved in Section 2. The self-adjoint parts of \(-iL_j\) are energy operators and have the same real principal symbols as \(-iL_j\). The principal symbol of \(T\) is that of \(-iV\), with \(V\) from (2.2). The module generated by all brackets of \(V,L_1,\ldots,L_r\) agrees with the one generated by all brackets of \(L_0,L_1,\ldots,L_r\). For completeness, multiplication by smooth coefficients causes only existing lower brackets: \[ [aA,bB]=ab[A,B]+a(A b)B-b(B a)A. \tag{7.2} \] Using (2.2), this identity proves both module containments by induction on the number of brackets.
At each point choose finitely many brackets whose values form a basis. They remain independent on a neighborhood. A finite covering of \(K\) supplies finitely many brackets and a largest weight \(N\). Their real linear symbols have no simultaneous zero on the unit cosphere, and compactness supplies the uniform lower bound (5.2). Theorems 5.2 and 6.1 apply. Changing \(P\) to \(-L\) changes neither Sobolev membership nor the wavefront set. ∎
For the manifold statement, choose local coordinates and a smooth positive density. Integration by parts has the same form, with divergence relative to that density, and multiplication by its smooth positive coordinate factor preserves all local Sobolev norms. Brackets and their ranks are intrinsic. Apply the coordinate proof in each chart; coordinate invariance of the Fourier wavefront set gives (7.1) globally.
Here is a larger family that illustrates how the rank condition is checked. Let \(f_1,\ldots,f_d\) be linearly independent real polynomials in one variable, of degree at most \(D\), and set \[ L=\partial_v^2+\sum_{j=1}^d f_j(v)\partial_{z_j}. \tag{7.3} \] The iterated brackets with \(\partial_v\) have coefficient vectors \((f_1^{(k)}(v),\ldots,f_d^{(k)}(v))\), \(0\leq k\leq D\). These span \(\mathbb R^d\) at every \(v\). Indeed a linear combination orthogonal to every such vector would define a polynomial \(f=\sum_j a_jf_j\) with all derivatives through degree \(D\) zero at that \(v\). Its Taylor formula would give \(f=0\), contradicting independence unless every \(a_j=0\). Together with \(\partial_v\), the bracket values have full rank. Their weights are at most \(D+2\), so the theorem gives gain \(2^{-D-1}\) for \(u\) from an \(L^2\) right-hand side. For example \(f_1(v)=1+v^2\) and \(f_2(v)=v^3-v\) give a three-dimensional operator regularizing in both \(z\) directions despite having only one diffusion field.
The rank condition has a local converse in the absence of the zero-order term.
Proposition 7.2. Suppose \(c=0\). If the Lie algebra has rank strictly less than \(n\) throughout a nonempty open set \(Y\), then \(L\) is not hypoelliptic on \(Y\).
Proof. The rank is integer-valued. Choose a point where it attains its maximum \(k<n\) on \(Y\). Some \(k\) brackets are independent there, and hence on a neighborhood; by maximality the rank is exactly \(k\) on that neighborhood. They give a smooth rank-\(k\) distribution \(\mathscr D\), and all original fields are tangent to it. Equation (7.2) and the Jacobi identity show that the distribution is involutive.
We explain why an involutive constant-rank distribution has local coordinates in which it is spanned by the first \(k\) coordinate directions. The assertion is trivial for \(k=0\). For \(k>0\), choose a nonvanishing field in \(\mathscr D\); its smooth flow and a transverse hypersurface, together with the inverse function theorem, give coordinates in which it is \(\partial_{x_1}\). Subtract multiples of \(\partial_{x_1}\) from the other frame fields, obtaining \(Y_2,\ldots,Y_k\) with no \(x_1\) component. Involutivity gives \[ [\partial_{x_1},Y_a]=\sum_{b=2}^k C_{ab}(x)Y_b. \tag{7.4} \] Let \(Y\) be the column of these fields. Solve the smooth matrix ordinary differential equation \[ \partial_{x_1}M=-MC,\qquad M(0,x')=I. \] Its solution is invertible on a smaller neighborhood: the determinant is nonzero initially and remains nonzero by uniqueness for the matrix equation and its inverse equation. Then \(Z=MY\) satisfies \(\partial_{x_1}Z=0\). The fields \(Z\) depend only on \(x'\) and span an involutive rank-\((k-1)\) distribution there; their brackets have no \(x_1\) component and remain in their span. Induction on \(k\) supplies coordinates on this transverse slice, independent of \(x_1\), which straighten that distribution. These, together with \(x_1\), are the required coordinates.
Thus every \(L_j\), including \(L_0\), differentiates only in the first \(k\) coordinates. The locally integrable function \[ u(x)=\mathbf1_{\{x_n>0\}} \] is independent of those coordinates. Distributionally \(L_ju=0\) for every \(j\): multiplication of its zero tangential derivatives by smooth coefficients is still zero. Hence \(Lu=0\), but \(u\) is not smooth along \(x_n=0\). This is a local failure of hypoellipticity inside \(Y\). ∎
This argument uses \(c=0\) exactly when concluding \(Lu=0\). It also requires a deficient-rank open set. Failure of the rank condition at one isolated point is not covered by Proposition 7.2.
8. Exercises with complete solutions
Exercise 1 — the drift sign, 6 points. For \(L=\partial_v^2+(v\partial_x+x\partial_t)+c(v,x,t)\), compute the field \(V\) in (2.2), the principal symbol of \(T\), and the real energy form of \(P=-L\).
Solution. The diffusion field \(\partial_v\) has zero divergence, so \(V=-v\partial_x-x\partial_t\). This field also has zero divergence. With frequency variables \((\nu,\xi,\tau)\), the symbol of \(-iV\), and thus the principal symbol of \(T\), is \(-v\xi-x\tau\). Equation (2.2) gives \[ \operatorname{Re}(Pu,u)=\|\partial_vu\|_0^2 -\int\operatorname{Re}c\,|u|^2. \] The imaginary part of \(c\) affects \(T\) only in order zero. On each compact set the last integral is bounded in absolute value by \(C_K\|u\|_0^2\).
Exercise 2 — a variable bracket coefficient, 8 points. Let \(X=\partial_x\), \(Y=x^3\partial_y+(1+x)\partial_z\), and \(L=X^2+Y\) on \(\mathbb R^3\). Check the rank condition at every \(x\). Give a uniform weight bound and the gain supplied by this lesson.
Solution. The brackets are \[ [X,Y]=3x^2\partial_y+\partial_z,\quad [X,[X,Y]]=6x\partial_y,\quad [X,[X,[X,Y]]]=6\partial_y. \] The last field supplies \(\partial_y\) everywhere. Subtract \(3x^2\partial_y\) from the first bracket to obtain \(\partial_z\); together with \(X\) these span all directions. The largest displayed weight is \(3+2=5\). Thus one may take \(N=5\), \(2\varepsilon=1/16\), and the diffusion estimate gains \(\varepsilon=1/32\). This bound holds on compact sets even at \(x=0\); its constants depend on those sets. It is a guaranteed gain, not a claim that \(1/16\) is optimal.
Exercise 3 — the two drift pairings, 10 points. In Lemma 4.3 verify the orders of \(A,qA,A^*q\), and explain why both positive-form pairings in (4.7) are bounded. Identify the hypothesis that would be missing if one tried to use just \(\|Tu\|_{-1/2}\).
Solution. Since \(C\) has order one, \[ \operatorname{ord}A=2(a/4-1)+1=a/2-1,\qquad \operatorname{ord}(qA)=\operatorname{ord}(A^*q)=a/2. \] Equation (4.2) gives \(qAu,A^*qu\in H^{a/2}\), with norms controlled by \(R(u)\). Apply Lemma 3.1 at order \(b=a/2\) to both operators. We then have \(E(qAu),E(A^*qu)\leq C R(u)\), and positivity gives \[ |(P'u,qAu)|\leq E(u)E(qAu),\quad |(A^*qu,P'u)|\leq E(A^*qu)E(u). \] After commuting \(P'\) through \(A\), these are precisely the two required form pairings. The isolated \(H^{-1/2}\) bound for \(Tu\) supplies neither energy bound for these auxiliary outputs. The energy factorization (3.2) is what provides them.
Exercise 4 — rough data, 10 points. Suppose \(N=3\), \(\varepsilon=1/8\), and initially \(u,F_ju\) are \(H^{-2}\) near a covector while \(Pu\) is \(H^0\) there. Give a finite bootstrapping schedule that proves the conclusion of Theorem 6.1. Explain why each step uses a smaller cone.
Solution. Start at \(t=-2\). Each simultaneous improvement gives \(u\in H^{t+1/4}\) and \(F_ju\in H^{t+1/8}\), so both are at least \(H^{t+1/8}\). Sixteen steps at orders \(t=-2+j/8\), \(0\leq j<16\), give both \(u,F_ju\in H^0\). The right-hand side is \(H^0\), hence belongs to every intermediate \(H^t\). One final step at \(t=0\) gives \(u\in H^{1/4}\) and \(F_ju\in H^{1/8}\). Choose seventeen nested cones with closures inside the previous cones before starting. Derivatives of the regularizing cutoff are supported away from the next inner cone, so their commutator errors are smoothing there; the larger cone carries the prior \(H^t\) control needed for (6.4). No convergence in \(H^{1/4}\) is assumed in advance.
Exercise 5 — what a deficient open set proves, 8 points. Consider \(L=\partial_x^2+x\partial_y\) and \(M=\partial_x^2\) on \(\mathbb R^2\). Compare their Lie ranks and construct the rank-obstruction solution for the one to which Proposition 7.2 applies. Does the vanishing of \(x\partial_y\) on \(x=0\) obstruct the first operator?
Solution. For \(L\), \([\partial_x,x\partial_y]=\partial_y\), so the Lie rank is two everywhere. The vanishing of the drift itself at \(x=0\) does not remove its commutator there; Theorem 7.1 makes \(L\) microlocally hypoelliptic. For \(M\), the generated distribution is spanned only by \(\partial_x\), with rank one throughout every open set. The function \(\mathbf1_{\{y>0\}}\) satisfies \(Mu=0\) distributionally and is not smooth on \(y=0\). This proves the failure of hypoellipticity for \(M\).
References
- [B] Marco Bramanti, On the proof of Hörmander's hypoellipticity theorem, Le Matematiche 75 (2020), 3–26. Open article. Sections 1–3 compare energy, commutator and flow methods for sums of squares.
- [H] Lars Hörmander, Hypoelliptic second order differential equations, Acta Mathematica 119 (1967), 147–171. doi:10.1007/BF02392081. The general theorem includes the real drift.
- [RS] Linda Preiss Rothschild and Elias M. Stein, Hypoelliptic differential operators and nilpotent groups, Acta Mathematica 137 (1976), 247–320. doi:10.1007/BF02392419.
Written by GPT-6.1 Sol (OpenAI), at Ultra reasoning effort, September 2026. Self-checked by the writing AI. Public domain (CC0).