An explicit canonical cell for a large transverse gradient
A large transverse gradient singles out a Hamiltonian direction. We can use one short Hamiltonian orbit to construct a canonical change of coordinates, without solving a general normal-form problem. In these coordinates the large gradient points in the \(\xi_2\) direction. Subtracting the corresponding linear term leaves a symbol whose derivatives are uniformly controlled.
The neighborhood has two possible shapes. Its transverse radius is \((\rho M)^{1/2}\), but its \(x_2\) radius can be longer. The distinction matters for mixed derivatives: \(x_2\) has its own scale, while every transverse frequency, including \(\xi_2\), uses the transverse radius.
We use the bracket normalization and initial bounds proved in An adaptive scale for repeated brackets, and the gradient alignment and projection estimates proved in How sign orientation aligns transverse gradients. The additional tools are the local existence theorem for smooth ordinary differential equations, Taylor's formula, the chain rule and finite-dimensional polynomial interpolation.
1. Select the gradient and its scales
Write \(t=x_1,\tau=\xi_1,w=(x',\xi')\), and use \(\{f,g\}=\sum_j(f_{\xi_j}g_{x_j}-f_{x_j}g_{\xi_j})\). Assume \(n\geq2\). Let \(q(t,w)\) be real, smooth and independent of \(\tau\). Fix \(k\geq1\), and suppose
\[ \begin{gathered} |\partial_t^i\partial_w^\gamma q| \leq C_{i\gamma}\lambda^{|\gamma|-2}, \qquad |t|<1,\quad |w|<\lambda^{-1},\\ \lambda^{-2}\leq C\sum_{1\leq |I|\leq k+1}|q_I(t,0,w)|,\\ q(t_-,w)>0\ \Longrightarrow\ q(t_+,w)\geq0, \qquad t_-<t_+. \end{gathered} \tag{1.1} \]Here \(q_1=\tau,q_2=q\), and the \(q_I\) are the repeated brackets defined in the preceding lesson. Constants in the derivative assumptions do not depend on \(\lambda\). The sign implication is required at each fixed transverse point throughout the stated domain.
Choose \(\rho\geq1\) first. Set
\[ \begin{gathered} M=\max_{1\leq |I|\leq k+1} \left(\frac{|q_I(0)|}{\rho}\right)^{1/|I|}, \qquad R=\lambda^{-\kappa},\\ 0<\kappa<\frac1{k+1}. \end{gathered} \tag{1.2} \]The earlier initial bounds give, for small \(\lambda\),
\[ \begin{gathered} M\geq c(\lambda^{-2}/\rho)^{1/(k+1)},\qquad M\geq1,\\ |\partial_t^i q(0)|\leq\rho M^{i+1}\quad(0\leq i\leq k),\\ \lambda^{-2}\leq C\rho M^{k+1},\qquad \frac{R}{(\rho M)^{1/2}}\longrightarrow0. \end{gathered} \tag{1.3} \]All limits below take \(\lambda\to0\) after fixing \(\rho\). The constants in the estimates are uniform; the allowed upper bound for \(\lambda\) may depend on \(\rho\).
Consider the alternative to the small-gradient cell:
\[ A_2=\max_{0\leq j\leq\lfloor k/2\rfloor} \frac{|\nabla_w\partial_t^j q(0)|}{\rho M^{j+1}} >R^{-1}. \tag{1.4} \]Choose an index \(s\) attaining this maximum and put
\[ \begin{gathered} a=|\nabla_w\partial_t^s q(0)| =\rho A_2M^{s+1},\\ L=(\rho M)^{1/2},\\ B_2=\frac{M^{s+1}}a=\frac1{\rho A_2},\\ b_2=\min(B_2,L^{-1}),\\ \mathcal A=\frac{a}{M^sL}. \end{gathered} \tag{1.5} \]Thus \(a\leq C\lambda^{-1}\), while \(\mathcal A>L/R\to\infty\). In particular \(L/a\to0\). The two \(x_2\) radii are combined by
\[ b_2^{-1}=\max(B_2^{-1},L). \tag{1.6} \]The maximum in (1.4) gives \(|\nabla_w\partial_t^j q(0)|\leq aM^{j-s}\) through \(\lfloor k/2\rfloor\). Higher time orders supply the same bound automatically, since
\[ \frac{\lambda^{-1}}{aM^{j-s}} \leq C\frac{M^{k/2-j}}{\mathcal A} \leq C,\qquad j>k/2. \tag{1.7} \]Taylor expansion in \(Mt\), with one higher time derivative for its remainder, therefore gives, for every fixed \(i\),
\[ |\nabla_w\partial_t^i q(t,0)| \leq C_i aM^{i-s},\qquad |Mt|<1. \tag{1.8} \]For completeness, when \(i\leq\lfloor k/2\rfloor\), expand this gradient through time order \(\lfloor k/2\rfloor\). Each center coefficient, after multiplication by \(t^{j-i}\), is bounded by \(CaM^{i-s}\). The next time order is bounded by \(\lambda^{-1}\), and (1.7), together with \(|t|<M^{-1}\), gives the same bound for the remainder. For larger \(i\), the original first-derivative estimate and (1.7) apply directly. No sign argument has been used yet.
2. Put a Hamiltonian orbit on the \(x_2\) axis
Define \(\psi(w)=\partial_t^s q(0,w)\). Rescale the transverse coordinates by \(a\):
\[ \Psi(y,\eta)=a^{-2}\psi(ay,a\eta). \tag{2.1} \]Then \(|\nabla\Psi(0)|=1\). Its derivatives of order at least two are uniformly bounded on a fixed small ball: for order \(r\geq2\), the bound is \(C_r(a\lambda)^{r-2}\), and \(a\lambda\leq C\). The value \(\Psi(0)\) need not be bounded; it does not enter the Hamilton equations.
Permute the canonical pairs and, if necessary, replace a pair \((y_j,\eta_j)\) by \((\eta_j,-y_j)\). Choose the sign so that \(\partial_{\eta_2}\Psi(0)>1/(2n)\). These linear maps preserve the symplectic form and the derivative hypotheses up to fixed constants. We retain the same notation after this preliminary change.
Solve the Hamilton equations from the origin,
\[ \dot y=\Psi_\eta,\qquad \dot\eta=-\Psi_y,\qquad (y,\eta)(0)=0. \tag{2.2} \]The Hessian bound implies \(|(\dot y,\dot\eta)|\leq1+C|(y,\eta)|\). Integration gives \(|(y,\eta)(u)|\leq(e^{C|u|}-1)/C\), with the usual value \(|u|\) when \(C=0\). Choose a uniform short time interval on which the orbit stays in the fixed ball and \(1/(4n)<\dot y_2<2\). We may then use \(y_2\) as its parameter. Smooth ODE estimates, and differentiation of the inverse parameter, give uniformly bounded smooth functions
\[ \begin{gathered} y_j=f_j(y_2)\quad(j\geq3),\\ \eta_j=g_j(y_2)\quad(j\geq2), \qquad f_j(0)=g_j(0)=0. \end{gathered} \tag{2.3} \]They are defined on a fixed interval about zero, with uniform bounds on every fixed derivative.
Use \((X',\eta')\) for the original transverse coordinates and \((x',\xi')\) for the new ones. Define \(\chi:(x',\xi')\mapsto(X',\eta')\) by
\[ \begin{gathered} X_2=x_2,\qquad X_j=x_j+a f_j(x_2/a)\quad(j\geq3),\\ \eta_j=\xi_j+a g_j(x_2/a)\quad(j\geq3),\\ \eta_2=\xi_2+a g_2(x_2/a)+\mathcal G(x',\xi'),\\ \mathcal G= \sum_{j\geq3} \bigl(g_j'(x_2/a)x_j-f_j'(x_2/a)\xi_j\bigr). \end{gathered} \tag{2.4} \]When \(n=2\), the sums and the coordinates with \(j\geq3\) are absent. This is still a valid construction.
Lemma 2.1 (explicit canonical map). On \(|w|<ca\), for a fixed sufficiently small \(c>0\), the map \(\chi\) is symplectic, fixes the origin, maps the \(x_2\) axis onto the selected orbit, and has an explicit smooth inverse. Both maps satisfy
\[ |D^\gamma\chi|\leq C_\gamma a^{1-|\gamma|}, \qquad |D^\gamma\chi^{-1}|\leq C_\gamma a^{1-|\gamma|} \tag{2.5} \]on their corresponding neighborhoods.
Proof. Let \(\theta=\sum_{j\geq2}\xi_j\,dx_j\). Expanding the pullback of the original canonical one-form gives
\[ \begin{gathered} \chi^*\left(\sum_{j\geq2}\eta_j\,dX_j\right)-\theta =dS,\\ S=a\sum_{j\geq3}g_j(x_2/a)x_j\\ \hspace{8mm} +a^2\int_0^{x_2/a} \left(g_2(u)+\sum_{j\geq3}g_j(u)f_j'(u)\right)\,du. \end{gathered} \tag{2.6} \]Indeed, the terms \(-f_j'\xi_j\,dx_2\) in \(\mathcal G\,dx_2\) cancel the terms \(f_j'\xi_j\,dx_2\) from \(\eta_jdX_j\). The remaining terms with \(x_jg_j'\) and \(ag_jdx_j\) are the differential of the first sum in \(S\); the one-variable integral supplies the rest. Taking an exterior derivative proves preservation of the symplectic form.
To invert the map, first recover \(x_2=X_2\). Then recover \(x_j=X_j-af_j(X_2/a)\) and \(\xi_j=\eta_j-ag_j(X_2/a)\) for \(j\geq3\). Finally recover \(\xi_2\) from the last two lines of (2.4). This proves invertibility directly. The map and this inverse each have the form \(a\) times a uniformly smooth function of the coordinates divided by \(a\), which proves (2.5).
On the new axis, (2.4) reduces to \((X_2,X_j,\eta_j)=(x_2,af_j(x_2/a),ag_j(x_2/a))\), the orbit in (2.3) rescaled to the original variables. ∎
Set
\[ Q(t,w)=q(t,\chi(w)). \tag{2.7} \]The canonical change leaves \(t,\tau\) fixed. Write \(Q^{(i)}=\partial_t^iQ\). Symplectic invariance of Hamilton fields and the parameter estimate in (2.2) give, on the \(x_2\) axis,
\[ \begin{gathered} H_{Q^{(s)}}(0,x_2,0)\\ =a c_1(x_2)\partial_{x_2} \pmod{\partial_\tau},\\ C^{-1}\leq c_1(x_2)\leq C,\\ |(a\partial_{x_2})^j c_1|\leq C_j. \end{gathered} \tag{2.8} \]The possible \(\partial_\tau\) component acts as zero on every function considered below. In particular, \(Q^{(s)}(0,x_2,0)=Q^{(s)}(0)\), and its transverse gradient on that axis is exactly \(a c_1(x_2)e_{\xi_2}\). The coefficient \(c_1(0)\) need not equal one; the uniform lower bound is what we need.
3. Retain the center jets and transformed derivative estimates
Lemma 3.1 (weighted center jets). For every \(i,j\geq0\),
\[ |\partial_{x_2}^j\partial_t^iQ(0)| \leq C_{ij}\rho M^{i+1}B_2^j. \tag{3.1} \]Proof. Put \(\psi=\partial_t^sq\) when forming brackets, and evaluate at \(t=0\). Applying \(H_\psi\) \(j\) times to \(\partial_t^iq\) produces a bracket tree with \(j(s+1)+i+1\) leaves. The Jacobi identity \(\{\{u,v\},h\}=\{u,\{v,h\}\}-\{v,\{u,h\}\}\) expands that tree into a finite linear combination of the bracket words used to define \(M\).
If the number of leaves is at most \(k+1\), its value at the center is bounded by a constant times \(\rho M^{j(s+1)+i+1}\). If it is larger, the initial anisotropic bracket estimate gives \(C_{ij}\lambda^{-2}\); (1.3) and \(M\geq1\) give the same claimed bound. This explains why the argument also covers jets beyond the defining finite family.
On the axis, \(a\partial_{x_2}=c_1^{-1}H_\psi\) on functions independent of \(\tau\). Repeated use of this identity expresses \((a\partial_{x_2})^j\partial_t^iQ(0)\) as a sum of \(H_\psi^\ell\partial_t^iq(0)\), \(0\leq\ell\leq j\), with uniformly bounded coefficients. The derivatives of \(c_1^{-1}\) needed here are bounded by (2.8). Since \(M\geq1\), the term with \(j\) gives an upper bound for all these powers. Dividing by \(a^j\) proves (3.1), because \(B_2=M^{s+1}/a\). ∎
Lemma 3.2 (all transformed derivatives). For every \(i\geq0\) and transverse multiindex \(\gamma\ne0\),
\[ |\partial_t^i\partial_w^\gamma Q| \leq C_{i\gamma}M^i a^{2-|\gamma|}. \tag{3.2} \]For all \(\gamma\), including zero, there is also the estimate
\[ |\partial_t^i\partial_w^\gamma Q| \leq C_{i\gamma}\rho M^{k+1}a^{-|\gamma|}. \tag{3.3} \]Both hold for \(|Mt|<1,|w|<ca\), after reducing the fixed \(c\).
Proof. In a chain-rule term with \(r\geq2\) derivatives on \(q^{(i)}\), the derivatives of \(\chi\) contribute \(a^{r-|\gamma|}\). Its bound is
\[ C\lambda^{r-2}a^{r-|\gamma|} =Ca^{2-|\gamma|}(a\lambda)^{r-2} \leq C'a^{2-|\gamma|}. \tag{3.4} \]For a term with only one derivative on \(q^{(i)}\), the Hessian bound and \(|\chi(w)|\leq Ca\) show
\[ |\nabla_w q^{(i)}(t,\chi(w))| \leq C_i(aM^{i-s}+a) \leq C_i'aM^i \tag{3.5} \]by (1.8). Its product with \(D^\gamma\chi\) is bounded by \(C_iM^ia^{2-|\gamma|}\). This proves (3.2).
For the second estimate, use the original bound \(C_i\lambda^{-1}\) for the first derivatives and \(C_i\lambda^{-2}\) for the value, instead of (3.5). Since \(a\leq C\lambda^{-1}\), a term with \(r\geq1\) derivatives on \(q^{(i)}\) is bounded by \(C\lambda^{-2}a^{-|\gamma|}\). The zero-derivative term has that bound as well. Now use (1.3). ∎
One smaller estimate will locate the gradient direction accurately. On the axis, the explicit map gives \(|\chi(x_2,0)|\leq C|x_2|\), and its first derivatives are bounded. Thus (1.8) and the Hessian bound give
\[ |\nabla_wQ(t,x_2,0)| \leq C(aM^{-s}+|x_2|). \tag{3.6} \]For all higher time orders, the original first-derivative estimate also gives \(|\nabla_w\partial_t^iQ(t,x_2,0)|\leq C_i\lambda^{-1}\). The same estimate holds on \(|w|<ca\). These estimates are stronger for this purpose than (3.2).
4. A derivative interpolation fact, including the boundary
We will need a uniform bound on lower derivatives of a bounded function whose sufficiently high derivatives are bounded. Here is a version that does not discard a smaller boundary neighborhood.
Lemma 4.1 (polynomial interpolation on a convex domain). Let \(\Omega\subset\mathbb R^d\) be a fixed bounded convex open set containing a ball about zero. If a smooth scalar or finite-dimensional vector function \(h\) satisfies
\[ \sup_\Omega |h|\leq C_0,\qquad \sup_\Omega |D^\gamma h|\leq C_r \quad\text{for }|\gamma|=r, \tag{4.1} \]then every derivative of order less than \(r\) is bounded throughout \(\Omega\) by a constant depending only on \(\Omega,d,r,C_0,C_r\).
Proof. Taylor-expand \(h\) at zero through degree \(r-1\). Convexity keeps each segment from zero to a point of \(\Omega\) inside \(\Omega\). The integral remainder is uniformly bounded because the domain is bounded and all order-\(r\) partial derivatives are bounded. Hence the Taylor polynomial \(P\) is uniformly bounded on a fixed ball contained in \(\Omega\).
Choose a finite interpolation grid in that ball with an invertible evaluation matrix on polynomials of degree at most \(r-1\). Start with a sufficiently small tensor grid with \(r\) distinct points in each coordinate: repeated one-dimensional interpolation determines every polynomial whose degree in each variable is at most \(r-1\), and hence determines this smaller polynomial space. Select a square full-rank subsystem of its evaluation matrix. Its inverse bounds all coefficients of \(P\) by the selected grid values. Therefore all the derivatives \(D^\gamma h(0)\), \(|\gamma|<r\), are bounded.
Taylor-expand \(D^\gamma h\) from zero to any point of \(\Omega\), using its already bounded center coefficients and the same order-\(r\) derivative bounds for the remainder. This bounds \(D^\gamma h\) on the whole domain. Apply the argument componentwise for a vector function. ∎
In particular, suppose \(g(T)\) is a vector function on \((-1,1)\), all derivatives of order \(m>\lfloor k/2\rfloor\) needed below are uniformly bounded, and \(|(I-\Pi_\omega)g(T)|\leq C\) for a fixed unit vector \(\omega\). Apply the lemma to \((I-\Pi_\omega)g\). It gives a uniform bound for \((I-\Pi_\omega)g^{(s)}(0)\) whenever \(s\leq\lfloor k/2\rfloor\). The constant vector \(\omega\) may depend on a parameter; the bound is uniform in that parameter.
5. The residual estimate with two radius branches
Write \(x''=(x_3,\ldots,x_n)\) and \(v=(x'',\xi')\). Thus \(\xi'\) includes \(\xi_2\), whereas \(x''\) excludes \(x_2\). Define
\[ \begin{gathered} \mathcal R(t,x_2,v)=Q(t,x_2,v)\\ -\xi_2\,\partial_{\xi_2}Q(t,x_2,0). \end{gathered} \tag{5.1} \]The cell \(\mathcal U\) is defined by \(|Mt|<1\), \(|b_2x_2|<1\) and \(|v|<L\).
Theorem 5.1 (canonical-cell residual). For all multiindices \(\alpha,\beta\), on \(\mathcal U\),
\[ \begin{gathered} |\partial_\xi^\alpha\partial_x^\beta\mathcal R|\\ \leq C_{\alpha\beta} \rho M^{\beta_1+1}b_2^{\beta_2}\\ \times L^{-(|\alpha'|+|\beta''|)}. \end{gathered} \tag{5.2} \]Here \(\alpha'=(\alpha_2,\ldots,\alpha_n)\) and \(\beta''=(\beta_3,\ldots,\beta_n)\). If \(\alpha_1>0\), the derivative is zero because the residual is independent of \(\tau\). Both the transverse-radius branch \(b_2=L^{-1}\) and the longer-\(x_2\) branch \(b_2=B_2<L^{-1}\) are included.
The entire cell is in the domain of \(\chi\) for sufficiently small \(\lambda\): its maximum coordinate radius divided by \(a\) is \(\max(L/a,M^{-(s+1)})\to0\). We now prove every derivative asserted in (5.2).
5.1. The transverse-radius branch
Suppose \(B_2\geq L^{-1}\). Rescale all transverse coordinates together:
\[ F(T,U)=\frac{\varepsilon_0}{\rho M} Q(T/M,LU),\qquad |T|<1,\quad |U|<2. \tag{5.3} \]The constant \(\varepsilon_0>0\) is fixed and sufficiently small. It does not depend on \(\lambda,\rho\). The enlarged fixed ball contains the product domain \(|x_2|<L,|v|<L\) after rescaling.
Taylor's formula in time, the center bounds in (1.3) and the high time bound in (3.3) give \(|Q(t,0)|\leq C\rho M\). Estimate (3.2) with two transverse derivatives gives \(\|F_{UU}\|\leq C\varepsilon_0\). Reduce \(\varepsilon_0\) so that the center height and curvature bounds needed for gradient alignment hold. The sign condition in (1.1) passes to \(F\).
Put \(g(T)=\nabla_UF(T,0)\). Equations (1.8), (2.8) and (3.6) give
\[ |g(T)|\leq C\mathcal A,\qquad g^{(s)}(0)=\varepsilon_0\mathcal A c_1(0)e_{\xi_2}. \tag{5.4} \]For every time order \(m>k/2\), the stronger first-derivative estimate after (3.6) gives
\[ |g^{(m)}(T)| \leq C_m\frac{\lambda^{-1}}{LM^m} \leq C_m'M^{k/2-m} \leq C_m'. \tag{5.5} \]The scaled gradient-alignment theorem supplies a fixed unit vector \(\omega\) with \(|(I-\Pi_\omega)g(T)|\leq C\) for all \(|T|<1\). Lemma 4.1 and (5.5) bound the corresponding projection of \(g^{(s)}(0)\). Its exact direction in (5.4) and \(c_1(0)\geq C^{-1}\) imply
\[ \|(I-\Pi_\omega)e_{\xi_2}\|\leq C/\mathcal A. \tag{5.6} \]The exact projector estimate from the preceding lesson says that \(\|\Pi_\omega-\Pi_{e_{\xi_2}}\|\) equals this sine of the angle. Combining it with \(|g(T)|\leq C\mathcal A\) yields
\[ |(I-\Pi_{e_{\xi_2}})g(T)|\leq C. \tag{5.7} \]Thus the derivative of the large gradient, rather than a possibly small gradient at a chosen time, identifies the direction. The argument does not assume that \(|g(T)|\) has a positive lower bound.
First subtract the coefficient at the origin:
\[ G_0(T,U)=F(T,U)-U_{\xi_2}\partial_{U_{\xi_2}}F(T,0). \tag{5.8} \]At \(U=0\), its gradient is the bounded vector in (5.7); its value is the bounded value \(F(T,0)\). Its Hessian equals that of \(F\). Integration along segments therefore bounds \(G_0\) on the full ball \(|U|<2\).
All derivatives of sufficiently high total order are bounded as well. If a derivative has \(d\geq2\) transverse differentiations, its bound from (3.2), after scaling, is \(C(L/a)^{d-2}\leq C\), for any time order, and the subtracted term has zero such derivative. With \(d\leq1\) and time order \(i\geq k+1\), (3.3) bounds the derivative of \(F\) by \(CM^{k-i}(L/a)^d\leq C\). The coefficient in (5.8) has the same bound by (5.5). Consequently every derivative of total order at least \(k+2\) is bounded. Apply Lemma 4.1 on \((-1,1)\times B_2(0)\), choosing an order larger than each derivative we want. Every derivative of \(G_0\) is uniformly bounded on that whole domain.
We must still use the coefficient at the \(x_2\) axis point, as (5.1) requires. Write \(U=(Z,Y)\), where \(Z=U_{x_2}\), and \(Y\) consists of the other transverse coordinates. The dimensionless requested residual is
\[ \begin{gathered} G(T,Z,Y)=F(T,Z,Y) -Y_{\xi_2}\partial_{Y_{\xi_2}}F(T,Z,0)\\ =G_0(T,Z,Y) -Y_{\xi_2}\partial_{Y_{\xi_2}}G_0(T,Z,0). \end{gathered} \tag{5.9} \]The equality follows by cancellation of the origin coefficient in (5.8). Both terms in the last line have uniformly bounded derivatives, including their time and \(Z\) derivatives, because they are values or traces of derivatives of \(G_0\). This proves uniform bounds for every derivative of \(G\) on \(|T|,|Z|<1,|Y|<1\). Rescaling proves (5.2) in this branch.
5.2. The longer-\(x_2\) branch
Suppose \(B_2<L^{-1}\). Now keep \(x_2\) on its own scale:
\[ \begin{gathered} F(T,Z,Y)\\ =\frac{\varepsilon_0}{\rho M} Q(T/M,Z/B_2,LY),\\ |T|,|Z|<1,\quad |Y|<1. \end{gathered} \tag{5.10} \]Here \(Y\) represents \(v=(x'',\xi')\). Notice that \(L<B_2^{-1}=a/M^{s+1}\); both radii still lie inside the canonical neighborhood.
We first bound \(F(T,Z,0)\). For every center derivative, (3.1) gives \(|\partial_T^i\partial_Z^jF(0,0,0)|\leq C_{ij}\). On the whole \((T,Z)\) square, (3.3) gives
\[ \begin{gathered} |\partial_T^i\partial_Z^jF(T,Z,0)|\\ \leq C_{ij}M^{k-i-(s+1)j}. \end{gathered} \tag{5.11} \]Taylor-expand through ordinary degree \(k\). Terms whose weight \(i+(s+1)j\) exceeds \(k\) have coefficients \(O(M^{-1})\), by (5.11). Every derivative of ordinary order \(k+1\) also has this bound. The remaining terms form a polynomial of weight at most \(k\), with uniformly bounded coefficients; the integral remainder is \(O(M^{-1})\). Thus \(F(T,Z,0)\) is uniformly bounded on the square. Reduce the fixed \(\varepsilon_0\) if needed so that its height is at most one.
For each \(Z\), regard \(F(T,Z,Y)\) as a function of \(T,Y\). The two-\(Y\)-derivative bound in (3.2) gives a uniformly bounded \(Y\)-Hessian, which we also normalize to at most one. The sign orientation is valid at each fixed \(Z,Y\), since the canonical map does not depend on time.
Let \(g_Z(T)=\nabla_YF(T,Z,0)\). Estimate (3.6) and \(|x_2|<a/M^{s+1}\) give \(|g_Z(T)|\leq C\mathcal A\), uniformly in \(Z\). Estimate (5.5) holds uniformly in \(Z\) as well. The exact orbit identity is now available at every \(Z\):
\[ g_Z^{(s)}(0) =\varepsilon_0\mathcal A c_1(Z/B_2)e_{\xi_2}. \tag{5.12} \]Apply gradient alignment for each fixed \(Z\), then Lemma 4.1 in the time variable and the projector calculation, exactly as in (5.4)–(5.7). This proves the uniform estimate \(|(I-\Pi_{e_{\xi_2}})g_Z(T)|\leq C\). The aligning vector may depend on \(Z\); we never differentiate that vector.
Define the requested residual directly:
\[ \begin{gathered} G(T,Z,Y)=F(T,Z,Y)\\ -Y_{\xi_2}\partial_{Y_{\xi_2}}F(T,Z,0). \end{gathered} \tag{5.13} \]Its center value and \(Y\)-gradient are bounded, and its \(Y\)-Hessian equals the bounded Hessian of \(F\). Hence \(G\) is uniformly bounded on the full product domain.
It remains to control mixed derivatives, including derivatives in \(Z\). For \(d\) differentiations in \(Y\), the two transformed estimates become
\[ \begin{gathered} |\partial_T^i\partial_Z^jD_Y^\gamma F|\\ \leq C_{ij\gamma} M^{-(s+1)j}(L/a)^{d-2}, \\j+d>0,\quad d=|\gamma|,\\ |\partial_T^i\partial_Z^jD_Y^\gamma F|\\ \leq C_{ij\gamma} M^{k-i-(s+1)j}(L/a)^d. \end{gathered} \tag{5.14} \]The first bound is used only for \(d\geq2\), where it is uniformly bounded. For \(d\leq1\), the second bound is uniformly bounded whenever \(i+(s+1)j\geq k\). These two alternatives bound every derivative of sufficiently high total order.
The same statement holds for \(G\). Derivatives with \(d\geq2\) annihilate its subtracted term. When \(d=1\), a surviving derivative of that term is a trace of \(\partial_T^i\partial_Z^j\partial_{Y_{\xi_2}}F\). When \(d=0\), it is \(Y_{\xi_2}\) times that trace. For total order at least \(k+2\), (5.14) bounds all such terms uniformly. This is a direct bound on mixed \(Z\) derivatives; pointwise gradient alignment alone would not give it.
Apply Lemma 4.1 to \(G\) on the fixed convex domain \((-1,1)^2\times B_1(0)\). Its bounded value and the bounded derivatives of arbitrarily large total order now bound every derivative on the entire domain. Rescaling (5.13) gives
\[ \begin{gathered} |\partial_t^i\partial_{x_2}^jD_v^\gamma\mathcal R|\\ \leq C_{ij\gamma}\rho M^{i+1}B_2^jL^{-|\gamma|}. \end{gathered} \tag{5.15} \]Since \(b_2=B_2\) in this branch, this is exactly (5.2). The equality case was included in the first branch. The proof of the theorem is complete. ∎
The estimate controls the residual after subtracting a linear frequency term. It does not assert that \(Q\) itself has bounded dimensionless first derivatives in the \(\xi_2\) direction. That coefficient can be large, and retaining it is essential for the later local model.
6. Exercises with complete solutions
Exercise 1 — basic. In three canonical pairs, take \(f_3(u)=u^2,g_3(u)=u,g_2(u)=0\) in (2.4). Write the map and its inverse, compute the primitive in (2.6), and show that it puts an orbit of
\[ \psi(X',\eta') =\frac{a\eta_2-aX_3+2X_2\eta_3-X_2^2}{\sqrt2} \tag{6.1} \]on the \(x_2\) axis. Check the gradient norm at the origin.
Solution. The map is
\[ \begin{gathered} X_2=x_2,\qquad X_3=x_3+x_2^2/a,\\ \eta_3=\xi_3+x_2,\qquad \eta_2=\xi_2+x_3-2x_2\xi_3/a. \end{gathered} \tag{6.2} \]Its inverse recovers \(x_2=X_2\), \(x_3=X_3-X_2^2/a\), \(\xi_3=\eta_3-X_2\), and \(\xi_2=\eta_2-X_3+2X_2\eta_3/a-X_2^2/a\). The primitive is \(S=x_2x_3+2x_2^3/(3a)\). Thus preservation of the symplectic form follows by the one-form calculation, with no omitted correction in \(\eta_2\).
Substitution gives \(\psi\circ\chi=a\xi_2/\sqrt2\). Its Hamilton field is \(a\partial_{x_2}/\sqrt2\), so the axis maps to a Hamiltonian orbit of \(\psi\). The two nonzero gradient components of \(\psi\) at zero are \(a/\sqrt2\) and \(-a/\sqrt2\); its norm is \(a\). This also gives the concrete coefficient \(c_1=1/\sqrt2\), illustrating why the proof retains its lower bound rather than setting it equal to one.
Exercise 2 — intermediate. Take \(s=1,i=0,j=2,k=3\) in the center-jet argument. Count the leaves of \(H_\psi^2q\), explain the bound when this exceeds the defining family, and verify the coefficient arising from two axis differentiations when \(H_\psi=ac_1(x_2)\partial_{x_2}\).
Solution. The leaf count is \(2(1+1)+0+1=5\), whereas the defining family has at most \(k+1=4\) leaves. The original anisotropic bracket bound gives \(C\lambda^{-2}\); (1.3) then gives \(C\lambda^{-2}\leq C'\rho M^4\leq C'\rho M^5\). The Jacobi expansion converts the tree to words before this bound is applied.
Write \(c=c_1\). Direct differentiation gives \(H_\psi^2q=a^2c^2q''+a^2cc'q'\), hence
\[ a^2q''=c^{-2}H_\psi^2q -ac'c^{-2}H_\psi q. \tag{6.3} \]The coefficient \(ac'\) is uniformly bounded. The second bracket has three leaves and contributes at most \(C\rho M^3\), which is at most \(C\rho M^5\). Dividing by \(a^2\) gives \(|\partial_{x_2}^2Q(0)|\leq C\rho MB_2^2\), as required.
Exercise 3 — intermediate. Compare the scale records \(\rho=1,k=3,M=\lambda^{-1/2},a=\lambda^{-7/8}\), with either \(s=1\) or \(s=0\). Choose \(\kappa=3/16\). Compute \(A_2,B_2,L,b_2^{-1}\) and \(\mathcal A\), and identify the two radius branches.
Solution. In both records \(L=\lambda^{-1/4}\) and \(a\lambda=\lambda^{1/8}\leq1\). For \(s=1\), \[ \begin{gathered} A_2=\lambda^{1/8}>\lambda^{3/16}=R^{-1},\\ B_2=\lambda^{-1/8}>L^{-1},\qquad b_2^{-1}=L=\lambda^{-1/4},\\ \mathcal A=\lambda^{-1/8}. \end{gathered} \tag{6.4} \] This is the transverse-radius branch. For \(s=0\), \[ \begin{gathered} A_2=\lambda^{-3/8}>R^{-1},\\ B_2=\lambda^{3/8}<L^{-1},\qquad b_2^{-1}=\lambda^{-3/8}>L,\\ \mathcal A=\lambda^{-5/8}. \end{gathered} \tag{6.5} \] This is the longer-\(x_2\) branch. The ratios \(L/a\) and \(b_2^{-1}/a\) tend to zero in both cases, placing either cell inside the canonical neighborhood. These are algebraic scale records; deciding whether a specified symbol realizes them additionally requires its gradient and bracket data.
Exercise 4 — intermediate. Suppose \(|h|\leq1\) and \(|h''|\leq K\) on \((-1,1)\). Prove a bound for \(|h'|\) valid arbitrarily near either endpoint. Explain what fails if only the value bound is given.
Solution. For \(x\geq0\), use the step \(\delta=-1/2\); for \(x<0\), use \(\delta=1/2\). The whole segment stays in the interval. Taylor's formula gives \(|\delta h'(x)|\leq|h(x+\delta)-h(x)|+K\delta^2/2\). Therefore \(|h'(x)|\leq4+K/4\), uniformly on the full interval. The functions \(h_N(x)=\sin(Nx)\) all have value bound one, but \(h_N'(0)=N\). Their second derivatives have size \(N^2\), so they do not satisfy a common high-derivative bound. A bounded residual alone is insufficient for the final derivative conclusion.
Exercise 5 — advanced. Let \(F(T,Z,\eta)=1+(A+Z)\eta+\eta^2\), where \(A\) is a constant. Compare subtraction of the coefficient at the origin with subtraction of the coefficient at the \(Z\) axis point. Identify the trace correction in (5.9).
Solution. The origin coefficient is \(A\), so \(G_0=1+Z\eta+\eta^2\). At the axis point, \(\partial_\eta F(T,Z,0)=A+Z\), so \(G=1+\eta^2\). Since \(\partial_\eta G_0(T,Z,0)=Z\), the trace correction is exactly \(Z\eta\): \(G=G_0-\eta\,\partial_\eta G_0(T,Z,0)\). It and every derivative are bounded on the fixed unit product domain, independently of \(A\). The subtraction removes the potentially large coefficient while preserving the distinct axis evaluation required in the theorem.
Exercise 6 — advanced. In the dimensionless variables \(T=Mt,Z=b_2x_2,Y=\xi_2/L\), take the residual model \(\mathcal R=\rho M T^2Z^3Y^4\). Compute \(\partial_t^2\partial_{x_2}^3\partial_{\xi_2}^4\mathcal R\), and identify its scale in (5.2).
Solution. Each differentiation in \(t,x_2,\xi_2\) contributes \(M,b_2,L^{-1}\), respectively. Hence the derivative is
\[ 2!\,3!\,4!\,\rho M\,M^2b_2^3L^{-4} =288\,\rho M^3b_2^3L^{-4}. \tag{6.6} \]Here \(\beta_1=2,\beta_2=3,|\alpha'|=4,|\beta''|=0\). The fourth frequency derivative uses \(L^{-4}\), even though it is in the distinguished frequency \(\xi_2\). Only the \(x_2\) derivatives use \(b_2^3\). Combining these indices into a single transverse scale would lose the precise statement.
References
The same results are treated in Hörmander, The Analysis of Linear Partial Differential Operators IV, Chapter 27, Section 27.4, especially the orbit construction and Lemma 27.4.6. The proof here supplies the explicit inverse and one-form check, the center-jet leaf count beyond the finite family, interpolation on the full domains, both radius branches and the axis-coefficient correction. The arguments and exercises have been written independently. The coefficient estimates, admissible covering, analytic local models and general sufficient-estimate assembly are subsequent steps.
Written by GPT-6.1 Sol (OpenAI), at Ultra reasoning effort, September 2026. Self-checked by the writing AI. Public domain (CC0).