Subellipticity and unique continuation · Self-checked by the writing AI

An adaptive scale for repeated brackets

A bracket can be large because one time derivative is large, or because several transverse derivatives interact. To distinguish these mechanisms, measure each bracket by a root that remembers its number of leaves. Where the first few transverse gradients are small, this measure stays comparable to a pure time-derivative scale throughout a useful neighborhood.

We will prove the derivative estimates, the comparison of the two scales, and the persistence of that comparison away from the center. The parameter used to strengthen a later estimate is chosen first; the frequency parameter is then made small. Keeping this order explicit matters when discarding transverse bracket terms.

Repeated brackets and their leaf counts are defined in Sign-constrained weighted Taylor geometry. The polynomial-jet translation argument also underlies A derivative scale for signed one-dimensional equations. Everything proved here concerns smooth scalar functions and finite derivatives.

1. A bracket remembers both its amplitude and its derivative scales

Use canonical coordinates \((x,\xi)\in\mathbb R^{2n}\), with

\[ \{f,g\} =\sum_{j=1}^n \bigl(\partial_{\xi_j}f\,\partial_{x_j}g -\partial_{x_j}f\,\partial_{\xi_j}g\bigr). \tag{1.1} \]

For two real smooth functions \(q_1,q_2\), define \(q_{(a)}=q_a\) and \(q_{(a,I)}=\{q_a,q_I\}\), where every letter is \(1\) or \(2\). The length \(|I|\) counts leaves, including the final leaf. All estimates are pointwise on an open set where the stated derivative bounds hold.

Lemma 1.1 (anisotropic bracket bound). Suppose \(P,K,A_j,B_j>0\) and

\[ |\partial_\xi^\alpha\partial_x^\beta q_a| \leq C_{\alpha\beta}PK A^\alpha B^\beta, \qquad a=1,2, \qquad PA_jB_j\leq1. \tag{1.2} \]

Then, for every word \(I\),

\[ |\partial_\xi^\alpha\partial_x^\beta q_I| \leq C_{I,\alpha\beta}PK^{|I|}A^\alpha B^\beta. \tag{1.3} \]

In the full product-rule expansion, the terms using a Poisson contraction in a specified coordinate \(j\) have the stronger bound

\[ C_{I,\alpha\beta} (PA_jB_j)\,PK^{|I|}A^\alpha B^\beta. \tag{1.4} \]

The constants depend on finitely many of the constants in (1.2), on the indices and on the dimension. They do not depend on the positive scale parameters.

Proof. The assertion for one leaf is (1.2). Suppose it holds for a word with \(\ell\) leaves. A term in \(\{q_a,q_I\}\) is a product of a derivative of \(q_a\) and a derivative of \(q_I\), with one extra \(x_j\) derivative and one extra \(\xi_j\) derivative. Applying \(\partial_\xi^\alpha\partial_x^\beta\) distributes its derivatives among the two factors. Each resulting term is bounded by a constant times

\[ P^2K^{\ell+1}A_jB_jA^\alpha B^\beta =(PA_jB_j)\,PK^{\ell+1}A^\alpha B^\beta. \tag{1.5} \]

The factor \(PA_jB_j\) is at most one. The finite sums in the bracket and the product rule prove the induction step for (1.3).

To retain (1.4), follow a specified contraction in each expanded term. If it is the new outer contraction, (1.5) retains its factor. If it occurs inside the previous word, use that word's marked bound and discard only the factor belonging to the outer contraction. Further derivatives change the finite constants and the powers \(A^\alpha B^\beta\); they do not remove the retained factor. Summing all terms that use the specified coordinate proves (1.4). ∎

The same argument applies to any bracket tree: bracketing expressions with \(\ell\) and \(r\) leaves gives \(P^2K^{\ell+r}A_jB_j\), and the same contraction factor reduces this to \(PK^{\ell+r}\). The definition above is sufficient for the scale we will use.

2. Normalize a finite bracket family

Write \(t=x_1\), \(\tau=\xi_1\), and \(w=(x',\xi')\in\mathbb R^{2n-2}\). Consider

\[ q_1=\tau,\qquad q_2=q(t,w), \tag{2.1} \]

where \(q\) is real and independent of \(\tau\). Fix an integer \(k\geq1\). Let \(0<\lambda<1\), and assume, with constants independent of \(\lambda\), that

\[ \begin{gathered} |\partial_t^a\partial_w^\beta q(t,w)| \leq C_{a\beta}\lambda^{|\beta|-2}, \qquad |t|<1,\quad |w|<\lambda^{-1},\\ \lambda^{-2} \leq C\sum_{1\leq |I|\leq k+1}|q_I(t,w)|, \qquad \tau=0. \end{gathered} \tag{2.2} \]

The first assumption gives bounded transverse second derivatives and a transverse first-derivative bound of order \(\lambda^{-1}\). The second assumption is a uniform finite-bracket condition.

Choose an amplification parameter \(\rho\geq1\), and define

\[ M_\rho(t,w) =\max_{1\leq |I|\leq k+1} \left(\frac{|q_I(t,0,w)|}{\rho}\right)^{1/|I|}. \tag{2.3} \]

Only the evaluation sets \(\tau=0\); brackets themselves are formed in all canonical coordinates. In particular, \(\partial_t^jq=\{\tau,\ldots,\{\tau,q\}\ldots\}\) has \(j+1\) leaves.

Lemma 2.1 (initial bounds). If \(\lambda^2\rho\leq1\), then

\[ c\left(\frac{\lambda^{-2}}{\rho}\right)^{1/(k+1)} \leq M_\rho(t,w) \leq C'\frac{\lambda^{-2}}{\rho}. \tag{2.4} \]

At the center, put \(M=M_\rho(0,0)\). Then

\[ \begin{aligned} |\partial_t^jq(0,0)|&\leq \rho M^{j+1}, &&0\leq j\leq k,\\ |\partial_t^jq(t,w)|&\leq C_j\rho M^{k+1}, &&j>k. \end{aligned} \tag{2.5} \]

Proof. First all the bracket functions in (2.2) are bounded by \(C_I\lambda^{-2}\). To check this with Lemma 1.1, use

\[ \begin{gathered} P=\lambda^{-2},\quad K=1,\\ A_1=\lambda^2,\quad B_1=1,\\ A_j=B_j=\lambda\quad (j>1). \end{gathered} \tag{2.6} \]

The leaf \(\tau\) satisfies (1.2) in \(|\tau|<\lambda^{-2}\). Its only nonzero derivative is \(\partial_\tau\tau=1=PKA_1\). The leaf \(q\) satisfies that bound by (2.2) and is independent of \(\tau\). Every product \(PA_jB_j\) equals one. Thus (1.3) gives the claimed bracket bound at \(\tau=0\).

There are finitely many words in (2.2), so at least one satisfies \(|q_I|\geq c_0\lambda^{-2}\). Put \(Z=\lambda^{-2}/\rho\geq1\). Taking the appropriate roots, and decreasing a fixed constant if necessary, gives the lower bound \(cZ^{1/(k+1)}\). Every upper bound \((C_IZ)^{1/|I|}\) is at most a fixed constant times \(Z\). This proves (2.4).

The first part of (2.5) follows directly from the pure time-derivative words in (2.3). Raising the lower bound for \(M\) to the power \(k+1\) yields \(\lambda^{-2}\leq C''\rho M^{k+1}\). The derivative bound in (2.2) then proves the second part of (2.5). ∎

Choose fixed exponents

\[ 0<\kappa_1<\kappa<\frac1{k+1}, \qquad R=\lambda^{-\kappa}, \qquad R_1=\lambda^{-\kappa_1}. \tag{2.7} \]

For \(\rho\geq1\) and small \(\lambda\), (2.4) gives

\[ M\geq1,\qquad R^2\leq\rho M,\qquad R<\lambda^{-1},\qquad \frac{R_1}{R}\longrightarrow0. \tag{2.8} \]

For example, \(R^2/(\rho M)\leq C\lambda^{2/(k+1)-2\kappa} \rho^{-k/(k+1)}\). We will also use

\[ \varepsilon=\frac{\rho}{R^2}\longrightarrow0. \tag{2.9} \]

This last limit takes \(\lambda\to0\) after fixing \(\rho\). Constants in the estimates below are independent of both parameters; the permitted smallness threshold for \(\lambda\) can depend on the chosen \(\rho\).

3. Small center gradients control every derivative on a cell

Assume at the center that

\[ |\nabla_w\partial_t^j q(0,0)| \leq \rho R^{-1}M^{j+1}, \qquad 0\leq j\leq\lfloor k/2\rfloor. \tag{3.1} \]

Proposition 3.1 (derivative bounds on the cell). Under (2.2) and (3.1), for sufficiently small \(\lambda\),

\[ |\partial_t^a\partial_w^\beta q(t,w)| \leq C_{a\beta}\rho M^{a+1}R^{-|\beta|}, \qquad |Mt|<1,\quad |w|<R. \tag{3.2} \]

Proof. Introduce dimensionless variables and a dimensionless function:

\[ \begin{gathered} s=Mt,\qquad y=w/R,\\ F(s,y)=\frac{q(s/M,Ry)}{\rho M}. \end{gathered} \tag{3.3} \]

We prove uniform bounds for every \(\partial_s^a\partial_y^\beta F\) when \(|s|<1,|y|<1\).

At the center, (2.5) gives

\[ |\partial_s^jF(0,0)|\leq1,\qquad 0\leq j\leq k. \tag{3.4} \]

The small-gradient hypothesis gives the same bound for \(|\nabla_y\partial_s^jF(0,0)|\) through \(\lfloor k/2\rfloor\). For the remaining center gradients, (2.2), the lower bound for \(M\), and \(R^2\leq\rho M\) give

\[ \begin{aligned} \frac{R|\nabla_w\partial_t^jq(0,0)|}{\rho M^{j+1}} &\leq C_j\frac{R\,\rho^{1/2}M^{(k+1)/2}}{\rho M^{j+1}}\\ &\leq C_jM^{k/2-j} \leq C_j,\qquad j>k/2. \end{aligned} \tag{3.5} \]

Here \(\lambda^{-1}\leq C\rho^{1/2}M^{(k+1)/2}\), by (2.4). Thus all center gradients needed in a time Taylor expansion are bounded.

For \(a\geq k+1\), (2.5) gives the stronger uniform bound

\[ |\partial_s^aF(s,y)| \leq C_aM^{k-a} \leq C_a/M. \tag{3.6} \]

For \(|\beta|\geq2\), the original derivative bounds yield

\[ \begin{gathered} |\partial_s^a\partial_y^\beta F|\\ \leq C_{a\beta} M^{-a}(\lambda R)^{|\beta|-2} \frac{R^2}{\rho M} \\ \leq C_{a\beta}. \end{gathered} \tag{3.7} \]

We used \(M\geq1\), \(\lambda R\leq1\), and (2.8).

Taylor's formula in \(s\), with (3.4) and the uniform bound for \(\partial_s^{k+1}F\), bounds \(\partial_s^aF(s,0)\) for \(a\leq k\). Larger \(a\) are already covered by (3.6).

For the gradients at \(y=0\), use Taylor's formula in \(s\) again, now with the center bounds from (3.1) and (3.5). Its remainder is uniformly bounded by

\[ |\nabla_y\partial_s^{k+1}F| \leq C\frac{R\lambda^{-1}}{\rho M^{k+2}} \leq CM^{-k/2-1}. \tag{3.8} \]

The same original derivative estimate bounds gradients with any time order \(a>k\), using the computation in (3.5).

Finally integrate the uniformly bounded \(y\)-Hessians from (3.7) along the segment from \(0\) to \(y\). This extends the first-derivative bounds to \(|y|<1\). Integrate those first derivatives once more to bound \(F\) and its time derivatives throughout that ball. We have obtained every required derivative bound. Returning to \((t,w)\) proves (3.2). ∎

This proof explains the cutoff \(k/2\) in (3.1): for higher time orders, the ordinary first-derivative bound and the inequality \(R^2\leq\rho M\) already supply the needed estimate.

4. The pure time scale is the full bracket scale on this cell

Define

\[ N_\rho(t,w) =\max_{0\leq j\leq k} \left(\frac{|\partial_t^jq(t,w)|}{\rho}\right)^{1/(j+1)}. \tag{4.1} \]

Theorem 4.1 (stable scale). Under the hypotheses of Proposition 3.1, choose \(\rho\geq1\) first and then \(\lambda\) sufficiently small. There are constants \(C,c,c_1>0\), independent of these parameters, such that

\[ \begin{gathered} M=N_\rho(0,0),\\ M_\rho(t,w)\leq CM,\\ \text{when }|Mt|<1,\quad |w|<R;\\[4pt] c_1M\leq N_\rho(t,w)\leq M_\rho(t,w),\\ M_\rho(t,w)\leq CM,\\ \text{when }|Mt|<1,\quad |w|<cR. \end{gathered} \tag{4.2} \]

The exact equality at the center is available for the chosen bracket-word normalization; comparability is the property needed in applications.

Proof. On a cylinder with \(|\tau|<\rho M\), apply Lemma 1.1 to (3.2), using

\[ \begin{gathered} P=\rho,\quad K=M,\\ A_1=(\rho M)^{-1},\quad B_1=M,\\ A_j=B_j=R^{-1}\quad(j>1). \end{gathered} \tag{4.3} \]

The leaf \(\tau\) again satisfies these bounds. Now

\[ PA_1B_1=1,\qquad PA_jB_j=\varepsilon\quad(j>1). \tag{4.4} \]

Thus \(|q_I|\leq C_I\rho M^{|I|}\), proving the upper bound in (4.2).

A term without a transverse contraction must be a pure time-derivative word, up to sign, or zero. Indeed, \(\tau\) differentiates a function independent of \(\tau\) in the \(t\) direction. A time contraction between two functions independent of \(\tau\) is zero. Therefore a nonzero time-only expansion contains exactly one \(q\) leaf and all its other leaves are \(\tau\). Every other term contains at least one transverse contraction. By the marked estimate (1.4), its bound includes the factor \(\varepsilon\).

For every word which is not a pure time-derivative word, we consequently have

\[ |q_I(0,0,0)| \leq C_I\varepsilon\rho M^{|I|}. \tag{4.5} \]

There are finitely many words under consideration. Take \(\lambda\) small enough that all their constants satisfy \(C_I\varepsilon<1\). Such a word cannot attain the maximum defining \(M\). The leaf \(\tau\) is zero at the center. A pure time-derivative word must therefore attain that maximum, proving \(M=N_\rho(0,0)\).

Use \(F\) from (3.3) and put

\[ P_0(s)=\sum_{j=0}^k \frac{\partial_s^jF(0,0)}{j!}\,s^j. \tag{4.6} \]

The equality of scales just proved implies

\[ \max_{0\leq j\leq k}|\partial_s^jF(0,0)|=1. \tag{4.7} \]

Polynomial translation does not destroy this normalization. For \(|s|\leq1\) and \(0\leq r\leq k\), the exact identity

\[ P_0^{(r)}(0) =\sum_{\ell=0}^{k-r} \frac{(-s)^\ell}{\ell!}\,P_0^{(r+\ell)}(s) \tag{4.8} \]

gives \(\max_{j\leq k}|P_0^{(j)}(s)|\geq e^{-1}\). By (3.6), Taylor's remainder gives

\[ \max_{j\leq k} |\partial_s^jF(s,0)-P_0^{(j)}(s)| \leq C/M. \tag{4.9} \]

Since \(M\to\infty\) for fixed \(\rho\), this error is at most \(1/(2e)\) for small \(\lambda\). Proposition 3.1 gives uniformly bounded \(\nabla_y\partial_s^jF\). Hence, when \(|y|<c\) for a sufficiently small fixed \(c\),

\[ \max_{j\leq k}|\partial_s^jF(s,y)| \geq \frac1{4e}, \qquad |s|<1. \tag{4.10} \]

Finally, rescaling (4.1) gives

\[ \frac{N_\rho(t,w)}{M} =\max_{j\leq k} |\partial_s^jF(s,y)|^{1/(j+1)}. \tag{4.11} \]

The lower bound (4.10) therefore proves \(N_\rho\geq c_1M\). Pure time words are included in the definition of \(M_\rho\), so \(N_\rho\leq M_\rho\). This completes the proof. ∎

In particular, the smaller neighborhood

\[ \begin{gathered} \mathcal U_0=\{(t,w):\\ |Mt|<1,\quad |w|<R_1\}. \end{gathered} \tag{4.12} \]

has \(M_\rho\) comparable to \(M\), because \(R_1/R\to0\). Choose a real \(\phi\in C_c^\infty(-1,1)\), equal to one on \([-3/4,3/4]\), with \(0\leq\phi\leq1\). A useful supported cutoff is

\[ \Phi_0(t,w) =\phi(Mt)\,\phi(|w|^2/R_1^2). \tag{4.13} \]

It satisfies

\[ |\partial_t^a\partial_w^\beta\Phi_0| \leq C_{a\beta}M^aR_1^{-|\beta|}. \tag{4.14} \]

To see this, rescale to \(s=Mt,z=w/R_1\). The function \(\phi(s)\phi(|z|^2)\) is fixed and smooth, and all its derivatives are bounded. Its support lies compactly inside \(\mathcal U_0\).

The derivative and finite-bracket hypotheses have supplied these cell estimates. In a general subelliptic argument, the orientation of sign changes is an additional hypothesis used to estimate the local models. The transverse-gradient alternative to (3.1), the covering by the resulting neighborhoods, and the analytic assembly remain separate steps.

5. Exercises with complete solutions

Exercise 1 — basic. In one canonical plane, let \(q_1=PKBx\) and \(q_2=PKA\xi\). Work where \(|Bx|<1,|A\xi|<1\), and assume \(PAB\leq1\). Compute the bracket and identify the retained factor in Lemma 1.1.

Solution. Both leaves have amplitude at most \(PK\), their first derivatives have the scales \(B,A\), and their higher derivatives vanish. Our sign convention gives

\[ \{q_1,q_2\} =-P^2K^2AB =-(PAB)\,PK^2. \tag{5.1} \]

Thus the coordinate contraction contributes exactly the marked factor \(PAB\). Discarding it proves the ordinary two-leaf bound; retaining it records the smaller bracket when \(PAB\) is small.

Exercise 2 — basic. Let \(q(t,w)=\lambda^{-2}t^k\), with \(k\geq1\). Compute \(M_\rho(0,0)\). Check (3.1), and determine when this function has no change from positive to negative as \(t\) increases.

Solution. There is no transverse dependence, so every nonzero bracket is a pure time derivative. At the center only the derivative of order \(k\) is nonzero. Therefore

\[ M_\rho(0,0) =\left(\frac{k!\lambda^{-2}}{\rho}\right)^{1/(k+1)}. \tag{5.2} \]

All the center gradients in (3.1) vanish. For even \(k\), the function is nonnegative and touches zero. For odd \(k\), it crosses from negative to positive. In both cases there is no positive-to-negative change. The calculation also shows why a \(k\)-th time derivative uses the root \(1/(k+1)\).

Exercise 3 — intermediate. In two canonical planes, take \(q(t,x_2,\xi_2)=\lambda^{-2}t^2+x_2\xi_2\), and \(k=2\). Verify the cell hypotheses at the origin and compute \(M\). Does the cell hypothesis ensure the required sign orientation at every nearby transverse point?

Solution. On \(|w|<\lambda^{-1}\), the value is bounded by \(C\lambda^{-2}\), first transverse derivatives by \(C\lambda^{-1}\), and the only nonzero transverse second derivative is \(\partial_{x_2}\partial_{\xi_2}q=1\). All higher derivatives satisfy (2.2). The pure second time derivative is \(2\lambda^{-2}\) everywhere, so the finite-bracket lower bound holds.

The center gradients of \(q\) and \(\partial_tq\) vanish. Since \(\{\tau,q\}=2\lambda^{-2}t\) and \(\{q,\{\tau,q\}\}=0\), the nonzero three-leaf value is \(\{\tau,\{\tau,q\}\}=2\lambda^{-2}\). Thus

\[ M=(2\lambda^{-2}/\rho)^{1/3}. \tag{5.3} \]

Choose a nearby point with \(x_2\xi_2<0\). As \(t\) increases through the negative root of \(\lambda^{-2}t^2+x_2\xi_2=0\), the function changes from positive to negative. The derivative and cell hypotheses hold, but this transverse point violates the needed sign orientation. The geometric scale estimate does not supply that additional hypothesis.

Exercise 4 — intermediate. Let \(P\) have degree at most \(k\) and \(\max_{j\leq k}|P^{(j)}(0)|=1\). Prove the lower bound used in (4.8). Then examine \(P(s)=(s-1)^k/k!\) at \(s=1\).

Solution. Taylor's formula for \(P^{(r)}\), centered at \(s\) and evaluated at zero, is exact:

\[ P^{(r)}(0) =\sum_{\ell=0}^{k-r} \frac{(-s)^\ell}{\ell!}P^{(r+\ell)}(s). \tag{5.4} \]

For \(|s|\leq1\), the sum of the absolute coefficient bounds is at most \(\sum_{\ell\geq0}1/\ell!=e\). Taking the maximum over \(r\) gives \(1\leq e\max_j|P^{(j)}(s)|\). For the specified polynomial, its derivatives at zero have absolute values \(1/(k-j)!\), so the normalization is exact. At \(s=1\), every derivative below order \(k\) vanishes, but \(P^{(k)}(1)=1\). The whole jet retains its size even when the lower derivatives vanish simultaneously.

Exercise 5 — advanced. In the model from Exercise 2, take \(k=3,\kappa=1/8,\kappa_1=1/16\), and fix \(\rho\). Compute the three ratios controlling the cell proof.

Solution. Here \(M=6^{1/4}\lambda^{-1/2}\rho^{-1/4}\). Hence

\[ \begin{aligned} \frac{R^2}{\rho M} &=6^{-1/4}\lambda^{1/4}\rho^{-3/4},\\ \frac{\rho}{R^2}&=\rho\lambda^{1/4},\\ \frac{R_1}{R}&=\lambda^{1/16}. \end{aligned} \tag{5.5} \]

All tend to zero for fixed \(\rho\) as \(\lambda\to0\). The first controls transverse derivatives, the second makes transverse bracket contributions negligible, and the third places the cutoff cell inside the region where the pure jet has a uniform lower bound.

Exercise 6 — advanced. Keep \(\kappa=1/8\), but choose \(\rho=\lambda^{-1}\). Does \(\lambda^2\rho\leq1\) guarantee that the transverse brackets can be discarded? More generally, what condition on \(\beta\) makes \(\rho/R^2\to0\) when \(\rho=\lambda^{-\beta}\)?

Solution. With \(\rho=\lambda^{-1}\), \(\lambda^2\rho=\lambda\leq1\), but

\[ \frac{\rho}{R^2}=\lambda^{-3/4}\longrightarrow\infty. \tag{5.6} \]

The initial bound for \(M\) is valid, while the small factor needed in (4.5) is absent. In general \(\rho/R^2=\lambda^{2\kappa-\beta}\), which tends to zero exactly when \(\beta<2\kappa\). The theorem uses the simpler, sufficient order of choices: fix \(\rho\), then decrease \(\lambda\). It asserts no uniform smallness threshold for an arbitrarily growing \(\rho\).

References

The source comparison for the anisotropic bracket estimate and the first stable neighborhood is Hörmander, The Analysis of Linear Partial Differential Operators IV, Chapter 27, Section 27.4, especially Lemma 27.4.1 and Proposition 27.4.2. The proofs and exercises above have been written independently. The subsequent transverse-gradient geometry and full sufficient estimate remain to be developed.

Written by GPT-6.1 Sol (OpenAI), at Ultra reasoning effort, September 2026. Self-checked by the writing AI. Public domain (CC0).