The linear frequency coefficient and the bracket scale
The canonical cell separates a large linear frequency term from a controlled residual. Its coefficient is almost a polynomial in time and almost independent of the distinguished spatial coordinate. We will use these facts to estimate the full symbol and compare its coefficient jets with the bracket scale.
Two distinctions are essential. Comparisons obtained from a bounded residual have uniform constants; they need not hold with constant one. Also, entry into the small-gradient case requires a sufficiently large fixed multiple of the spatial radius. We prove the estimates with these constants and give admissible examples showing why the distinctions cannot be removed.
Use the hypotheses, explicit canonical map and notation of An explicit canonical cell for a large transverse gradient. In particular \(s\leq\lfloor k/2\rfloor\), and
\[ \begin{gathered} a=\rho A_2M^{s+1},\\ L=(\rho M)^{1/2},\\ B_2=(\rho A_2)^{-1},\\ b=\min(B_2,L^{-1}),\\ R=\lambda^{-\kappa},\\ \varepsilon=R/L\longrightarrow0. \end{gathered} \tag{1.1} \]The parameter \(\rho\geq1\) is fixed before \(\lambda\) is made small. All constants below are independent of these parameters. Write \(z=x_2\), \(v=(x'',\xi')\), and define
\[ \begin{gathered} c(t,z)=\partial_{\xi_2}Q(t,z,0),\\ \mathcal R(t,z,v)=Q(t,z,v)-\xi_2c(t,z). \end{gathered} \tag{1.2} \]The canonical residual theorem supplies every derivative of \(\mathcal R\) on \(|Mt|<1,|bz|<1,|v|<L\).
1. A normalized time polynomial and slow spatial variation
Lemma 1.1 (coefficient jets). On the stated time and spatial intervals, for every \(i,j\geq0\),
\[ \begin{gathered} |\partial_t^i c(t,z)|\leq C_i\rho A_2M^{i+1},\\ |\partial_t^i\partial_z^j(c(t,z)-c(t,0))|\\ \leq C_{ij}\rho A_2M^{i+1} b^j\varepsilon^{\max(j,1)}. \end{gathered} \tag{1.3} \]The normalized function \(g(T,Z)=c(T/M,Z/b)/(\rho MA_2)\) has a polynomial \(P(T)\) of degree at most \(\ell=\lfloor k/2\rfloor\) such that, for every fixed mixed derivative,
\[ \begin{gathered} |D_{T,Z}^{\gamma}(g(T,Z)-P(T))|\\ \leq C_\gamma\varepsilon,\\ |P^{(j)}(0)|\leq C,\\ |P^{(s)}(0)|\geq c_0>0. \end{gathered} \tag{1.4} \]Consequently, for small \(\lambda\),
\[ \max_{0\leq j\leq\ell} \frac{|\partial_t^jc(t,z)|}{\rho A_2M^{j+1}} \geq c_1>0. \tag{1.5} \]Proof. At \(z=0\), the stronger axis-gradient estimate in the canonical-cell proof bounds \(\partial_t^ic\) by \(CaM^{i-s}=\rho A_2M^{i+1}C\). Its selected derivative is exactly \(\partial_t^sc(0,0)=ac_1(0)\), where the orbit speed has a uniform positive lower bound.
The transformed second-derivative estimate gives \(|\partial_t^i\partial_zc|\leq C_iM^i\). The scales satisfy
\[ \begin{gathered} \frac1{\rho A_2Mb} =\max(1/M,B_2L/M)\leq R/L,\\ b\varepsilon\geq B_2/M=M^s/a\geq a^{-1}. \end{gathered} \tag{1.6} \]Indeed \(A_2>R^{-1}\) gives \(B_2<R/\rho\), so \(B_2L/M<R/L\). Also \(1/M\leq R/L\) for small \(\lambda\), since \(\rho\) is fixed and \(MR^2\to\infty\). For \(j\geq1\), the general transformed derivative estimate gives \(|\partial_t^i\partial_z^jc|\leq C_{ij}M^ia^{1-j}\). Use both inequalities in (1.6) to bound it by \(C_{ij}\rho A_2M^{i+1}(b\varepsilon)^j\). For \(j=0\), integrate the \(j=1\) estimate from zero to \(z\), using \(|z|<b^{-1}\). This proves the second bound in (1.3), including its stronger factor \(\varepsilon\) when \(j=0\). It also extends the first bound to the whole interval.
For \(m>\ell\), the original first transverse derivative bound gives
\[ |\partial_T^mg(T,0)| \leq C_m\frac{\lambda^{-1}}{aM^{m-s}} \leq C_m\frac{M^{k/2-m}}{a/(M^sL)} \leq C_m\varepsilon. \tag{1.7} \]Taylor-expand \(g(T,0)\) through degree \(\ell\) at zero. Equation (1.7) bounds the remainder and every fixed derivative of it; above degree \(\ell\), the polynomial derivatives vanish and (1.7) applies directly. Rescaling (1.3) controls the additional \(Z\) dependence and all its mixed derivatives. The selected derivative gives the nonzero coefficient in (1.4).
Finally polynomial translation is exact: \(P^{(s)}(0)=\sum_{r=0}^{\ell-s}(-T)^rP^{(s+r)}(T)/r!\). For \(|T|<1\), the absolute coefficient sum is at most \(e\). Hence the maximum of the polynomial jet at \(T\) is at least \(c_0/e\). The uniformly small errors in (1.4) prove (1.5). ∎
2. The full symbol and bracket upper bound
Let \(\xi''=(\xi_3,\ldots,\xi_n)\). Frequency derivatives in \(\xi_2\) now use the scale \(A_2\); derivatives in \(x_2\) use \(b\). All the remaining transverse derivatives use \(L^{-1}\).
Theorem 2.1. For every \(i,j\geq0\) and frequency/spatial multiindices in the remaining coordinates,
\[ \begin{gathered} |\partial_t^i\partial_z^j \partial_{\xi_2}^{r}D_{x'',\xi''}^{\gamma}Q|\\ \leq C_{ijr\gamma}\rho M^{i+1}\\ \times(1+|A_2\xi_2|)b^jA_2^rL^{-|\gamma|}. \end{gathered} \tag{2.1} \]For the bracket scale of the transformed pair \(\tau,Q\),
\[ \mu(t,w)= \max_{1\leq|I|\leq k+1} \left(\frac{|Q_I(t,0,w)|}{\rho}\right)^{1/|I|}, \tag{2.2} \]we have
\[ \mu(t,w)\leq CM(1+|A_2\xi_2|). \tag{2.3} \]The scale \(\mu\) equals the original bracket scale at the point \((t,\chi(w))\), by symplectic invariance.
Proof. Differentiate \(Q=\mathcal R+\xi_2c(t,z)\). The linear term is zero after two frequency derivatives, or after any derivative in \(x'',\xi''\). Its other derivatives satisfy (2.1) by (1.3). The residual theorem uses \(L^{-r}\) for its \(\xi_2\) derivatives; this is bounded by \(A_2^r\) since \(A_2L>L/R\to\infty\).
At a fixed point put \(K=M(1+|A_2\xi_2|)\). Apply the anisotropic bracket lemma from An adaptive scale for repeated brackets with
\[ \begin{gathered} P=\rho,\quad B_1=M,\quad A_1=(\rho K)^{-1},\\ A_2\text{ as above},\quad B_2=b,\\ A_j=B_j=L^{-1}\quad(j>2). \end{gathered} \tag{2.4} \]The contraction factors are \(M/K\leq1\), \(\rho A_2b\leq1\), and \(\rho/L^2=1/M\leq1\). The leaf \(\tau\), evaluated at \(\tau=0\), satisfies the same pointwise derivative bounds. The full product-rule bracket expansion therefore gives \(|Q_I|\leq C_I\rho K^{|I|}\). Taking the finitely many roots proves (2.3). The scales are held fixed during this pointwise application; no differentiation of \(K\) is being asserted. ∎
3. What a large linear frequency forces
Define the coefficient-jet measure
\[ S(t,z,\xi_2)= \max_{0\leq j\leq k} \left(\frac{|\xi_2\partial_t^jc(t,z)|}{\rho}\right)^{1/(j+1)}, \qquad X=|A_2\xi_2|. \tag{3.1} \]Theorem 3.1. There are uniform constants \(C,c,N>0\) such that
\[ \begin{gathered} S\leq\mu+CM,\\ S\geq 2cM X^{2/(k+2)}\qquad(X\geq1),\\ \mu\geq cM X^{2/(k+2)}\qquad(X\geq N). \end{gathered} \tag{3.2} \]For \(X\geq N\), increasing \(N\) if necessary, also \(S\leq C'\mu\).
Proof. The residual theorem gives \(|\partial_t^jQ-\xi_2\partial_t^jc|\leq C_j\rho M^{j+1}\). Pure time derivatives are bracket words, so \(|\partial_t^jQ|/\rho\leq\mu^{j+1}\). The inequality \((u+v)^{1/r}\leq u^{1/r}+v^{1/r}\), \(r\geq1\), proves the first line of (3.2).
By (1.5), at least one index \(j\leq\ell\) gives \(S\geq M(c_1X)^{1/(j+1)}\). When \(X\geq1\), the exponent is at least \(2/(k+2)\). Taking the minimum of the finitely many positive coefficient constants proves the second line, with the factor \(2\) absorbed into the choice of \(c\). Choose \(N\) so that \(cX^{2/(k+2)}\geq C\) for \(X\geq N\), and subtract the additive \(CM\). This proves the third line. Taking \(N\) larger ensures \(\mu\geq M\), and the first line then gives \(S\leq C'\mu\). ∎
The conclusion with a uniform constant is the one used below. An exact inequality \(S\leq\mu\) is not a consequence of the residual estimate; Exercise 4 shows that it can fail even for arbitrarily large \(X\).
4. Keep the full upper-bound range and state the correct case transition
We first isolate the part of the stable-cell proof that only needs an auxiliary upper scale.
Lemma 4.1 (auxiliary scale). At a point, suppose \(K\geq1\) satisfies \(\lambda^{-2}\leq C\rho K^{k+1}\), \(R^2\leq\rho K\), \(\lambda R\leq1\), and
\[ \begin{gathered} |\partial_t^jq|\leq C\rho K^{j+1}\quad(0\leq j\leq k),\\ |\nabla_w\partial_t^jq| \leq C\rho K^{j+1}/R \quad(0\leq j\leq\lfloor k/2\rfloor). \end{gathered} \tag{4.1} \]Assume the original derivative bounds hold on the centered cell of time radius \(K^{-1}\) and spatial radius \(R\). For small \(\lambda\), its bracket scale at the center is at most \(C'K\). The constants in (4.1) may exceed one.
Proof. Rescale the centered function to \(F(T,Y)=q(t_0+T/K,w_0+RY)/(\rho K)\). The value jets through order \(k\) and the low time-gradient jets at the center are bounded by (4.1). Higher center gradients are bounded by \(CR\lambda^{-1}/(\rho K^{j+1})\leq C K^{k/2-j}\), using \(R\leq(\rho K)^{1/2}\). Pure time derivatives of order \(i\geq k+1\) are bounded on the whole rescaled cell by \(CK^{k-i}\). For \(d\geq2\) transverse derivatives, the bound is \(C(\lambda R)^{d-2}R^2/(\rho K)\leq C\), for every time order. High time derivatives of the gradient have the same uniform bound from the original first-derivative estimate.
Time Taylor expansion first bounds every value and gradient jet at \(Y=0\). Integrating the bounded transverse Hessians along segments then bounds every derivative of \(F\) on the cell. Thus the original derivatives have bounds \(C_{i\gamma}\rho K^{i+1}R^{-|\gamma|}\). Apply the anisotropic bracket lemma with \(P=\rho\), time scales \(K,(\rho K)^{-1}\), and all transverse scales \(R^{-1}\). The transverse contraction factor is \(\rho/R^2\leq1\) for small \(\lambda\), with \(\rho\) fixed. It follows that \(|q_I|\leq C_I\rho K^{|I|}\), giving the claimed upper bound. This argument does not identify \(K\) with the actual bracket scale. ∎
Theorem 4.2 (coefficient upper bound and case transition). On a slightly smaller fixed transverse cell \(|v|<cL\), while retaining \(|Mt|<1,|bz|<1\),
\[ \mu\leq C S\qquad\text{when }|\xi_2|\geq R. \tag{4.2} \]There is a uniform fixed \(K_0\geq1\) such that, when \(|\xi_2|\geq K_0R\), the original point \((t,\chi(w))\) satisfies the exact small-gradient condition
\[ \begin{gathered} |\nabla_w\partial_t^jq(t,\chi(w))|\\ \leq \rho\mu^{j+1}/R,\\ j\leq\lfloor k/2\rfloor. \end{gathered} \tag{4.3} \]In that larger-frequency range, \(\mu\) equals its pure time-jet scale at the point, and \(\mu\) and \(S\) are comparable with uniform constants.
Proof. If \(|\xi_2|\geq R\), then \(X>1\), so (3.2) gives \(S\geq cM\). For \(j\leq k\), the definition of \(S\) gives \(|\partial_t^jc|\leq\rho S^{j+1}/|\xi_2|\). The differentiated decomposition \(Q=\xi_2c+\mathcal R\), together with the transformed second-derivative bound \(|\partial_z\partial_t^jc|\leq C_jM^j\), gives
\[ \begin{gathered} |\nabla_w\partial_t^jQ|\\ \leq \rho S^{j+1}/|\xi_2|+C_jM^jL,\\ |\partial_t^jQ|\\ \leq \rho S^{j+1}+C_j\rho M^{j+1}. \end{gathered} \tag{4.4} \]Here \(|\xi_2|<cL\) bounds the \(z\)-derivative of the linear term, and the residual gradient has bound \(C\rho M^{j+1}/L=CM^jL\). The inverse canonical map has bounded first derivatives, so the same gradient estimate holds for the original \(q\), up to a uniform constant.
Since \(S\geq cM\) and \(R/L\to0\), (4.4) verifies the center conditions of Lemma 4.1 with auxiliary scale \(K=S\), up to fixed constants. The other scale hypotheses follow from \(\lambda^{-2}\leq C\rho M^{k+1}\), \(R^2/(\rho M)\to0\), and \(S\geq cM\). The needed centered cell remains in the original domain: the canonical cell is well inside the spatial ball of radius \(\lambda^{-1}\), \(R\ll\lambda^{-1}\), and both time radii tend to zero. Lemma 4.1 proves (4.2) on the entire stated range.
For the case transition, choose \(K_0\) large enough that \(X\geq K_0\) implies the large-\(X\) estimates and \(\mu\geq M\). Now \(S\leq C\mu\), so (4.4), in original coordinates, is at most
\[ C_1\rho\mu^{j+1}/|\xi_2|+C_2M^jL. \tag{4.5} \]Choose \(K_0\geq2C_1\). The first term is at most half of \(\rho\mu^{j+1}/R\). The ratio of the second term to that quantity is at most \(C_2(R/L)(M/\mu)^{j+1}\), which is at most \(1/2\) for small \(\lambda\). This proves (4.3).
Apply the small-gradient stable-cell proof at this original point. Its marked transverse bracket factor \(\rho/R^2\to0\) shows that a pure time word attains the defining maximum. Thus \(\mu\) is the pure time-jet scale. The residual estimate gives that scale at most \(S+CM\), while the previous large-\(X\) bounds give \(M\leq\mu\) and \(S\leq C\mu\). Together with (4.2) this proves comparability. ∎
The range \(K_0R\) in the transition is substantive. Exercise 5 disproves transition at the exact radius \(R\) under the stated hypotheses. Estimate (4.2), however, retains that original radius: its proof uses an auxiliary upper scale and does not rely on an invalid transition.
5. Exercises with complete solutions
Exercise 1 — basic. Verify the first equality in (1.6) in each of the two branches \(b=B_2\) and \(b=L^{-1}\). Explain why its upper bound is uniform for fixed \(\rho\).
Solution. Since \(\rho A_2=1/B_2\), the left side is \(B_2/(Mb)\). If \(b=B_2\), it equals \(1/M\); if \(b=L^{-1}\), it equals \(B_2L/M\). Taking the smaller \(b\) selects the larger of these two values. The second is less than \(R/L\) because \(B_2<R/\rho\) and \(L^2=\rho M\). The first is at most \(R/L\) once \(MR^2\geq\rho\), which holds for small \(\lambda\) after fixing \(\rho\).
Exercise 2 — intermediate. For \(k=3\), examine the normalized coefficient \(g(T,Z)=1+T+\varepsilon(T^2+Z)\) on the unit square. Choose its degree-one polynomial, check the mixed derivative error, and show that its derivative jet has a uniform lower bound even near a zero of its value.
Solution. Take \(P(T)=1+T\). Its coefficients are bounded and \(P'(0)=1\). The difference and each fixed derivative are \(O(\varepsilon)\); derivatives in \(Z\) are \(\varepsilon\) at first order and zero above that. Since \(\partial_Tg=1+2\varepsilon T\), it has magnitude at least \(1/2\) when \(\varepsilon\leq1/4\). Thus the maximum of the value and first derivative is at least \(1/2\), including where the value is small. A nonvanishing jet does not require a nonvanishing value.
Exercise 3 — intermediate. Explain the exponent \(2/(k+2)\) when the selected nonzero coefficient derivative has order \(j\leq\lfloor k/2\rfloor\). Compare \(k=3\) and \(k=4\).
Solution. The coefficient measure includes the root \(1/(j+1)\). Its smallest possible exponent is \(1/(\lfloor k/2\rfloor+1)\), which is at least \(2/(k+2)\). For \(k=3\), that smallest exponent is \(1/2\), while \(2/(k+2)=2/5\). For \(k=4\), both are \(1/3\). Using the latter exponent gives one valid formula for both parities; the odd case also permits the stronger exponent.
Exercise 4 — advanced. Fix any \(\rho\geq1\), take \(k=1\), \(a=\lambda^{-1}\), and \(q(t,x_2,\eta_2)=a^2t+a\eta_2+\rho\). At \(\eta_2=-K\sqrt\rho\), \(t=x_2=0\), with fixed \(K>2\), compute \(M,A_2,S,\mu\). Check that this is an admissible large-gradient cell and decide whether \(S\leq\mu\).
Solution. The time derivative is \(a^2\), all higher time and transverse second derivatives vanish, and the transverse gradient has norm \(a\). On the original domain the value is bounded by \(C\lambda^{-2}\) for small \(\lambda\), with a uniform \(C\). The bracket \(\{\tau,q\}=a^2=\lambda^{-2}\) supplies the finite-bracket lower bound everywhere. Since \(q_t>0\), the sign orientation holds.
At the center, \(M=\max(1,a/\sqrt\rho)=a/\sqrt\rho\). Thus \(s=0\), \(A_2=1/\sqrt\rho>R^{-1}\) for small \(\lambda\), and the orbit map is the identity. We have \(c=a\), so \(S=KM\) and \(X=K\). At the stated point the value divided by \(\rho\) has magnitude \(KM-1\); the other nonzero bracket gives the root \(M\). Consequently
\[ \mu=KM-1<S=KM \tag{5.1} \]for small \(\lambda\). The fixed frequency \(K\sqrt\rho\) lies inside the transverse radius \(L=(a\sqrt\rho)^{1/2}\). The example works with any fixed \(K\), however large. It disproves a constant-one comparison while satisfying every derivative, bracket and sign hypothesis.
Exercise 5 — advanced. Fix any \(\rho\geq1\), \(k=1\), \(0<\kappa<1/2\), \(a=\lambda^{-1}\), \(R=\lambda^{-\kappa}\), and \(d=3aR/4\). Consider
\[ \begin{gathered} q(t,X_2,\eta_2)\\ =a^2t+\frac a{\sqrt2}(X_2+\eta_2)+d. \end{gathered} \tag{5.2} \]Use the canonical orbit shear \(X_2=x_2,\eta_2=\xi_2-x_2\). At \(\xi_2=-3\sqrt2R/4,t=x_2=0\), test the claim that \(|\xi_2|\geq R\) forces the small-gradient condition.
Solution. The gradient norm is \(a\), and the time bracket is \(a^2\). The value and all derivatives have the required bounds, since \(R/a\to0\); the time slope is positive. For small \(\lambda\), \(M=d/\rho=3aR/(4\rho)\), so \(A_2=4/(3R)>R^{-1}\). The normalized Hamilton orbit has slope \(-1\) in the \((X_2,\eta_2)\) plane, giving exactly the displayed shear and the speed \(c_1=1/\sqrt2\). The transformed symbol is \(Q=a^2t+a\xi_2/\sqrt2+d\).
At the chosen point, \(Q=0\), so \(\mu=a/\sqrt\rho\). Its frequency magnitude is \(3\sqrt2R/4>R\), yet the original transverse gradient satisfies
\[ a>\rho\mu/R=\sqrt\rho\,a/R \tag{5.3} \]once \(R>\sqrt\rho\). The small-gradient condition fails. The point is inside the canonical cell because \(R/L\to0\), with \(L=(3aR/4)^{1/2}\). Thus the exact-radius transition is false even with a normalized, explicit orbit map; a sufficiently large fixed multiple is needed.
Exercise 6 — advanced. In Exercise 5, compute \(S/\mu\). Explain why the full-range estimate (4.2) should not be described as two-sided comparability, and identify the additional range where Theorem 4.2 does give comparability.
Solution. The coefficient is the constant \(a/\sqrt2\). Its contribution at the chosen root is \(S=d/\rho=M\), while \(\mu=a/\sqrt\rho\). Hence \(S/\mu=3R/(4\sqrt\rho)\to\infty\). The estimate \(\mu\leq CS\) remains true, but the reverse inequality with a uniform constant fails on the original range \(|\xi_2|\geq R\). The theorem proves both directions only when \(|\xi_2|\geq K_0R\), where the large-\(X\) lower bound and the genuine small-gradient transition apply. The auxiliary-scale proof is what preserves the upper estimate without making the false reverse claim.
References
The same results are treated in Hörmander, The Analysis of Linear Partial Differential Operators IV, Chapter 27, Section 27.4, especially the coefficient and bracket-scale estimates following Lemma 27.4.6. The independently written proof retains all coefficient derivatives and the full original range of the coefficient-to-bracket upper bound. Exercises 4 and 5 resolve the constant-one coefficient comparison and exact-radius transition appearing in the printed argument; the corrected estimates use uniform constants and the stated fixed radius multiple. Weighted polynomial models and admissible covering follow next.
Written by GPT-6.1 Sol (OpenAI), at Ultra reasoning effort, September 2026. Self-checked by the writing AI. Public domain (CC0).