Compressed spectral measures and symbol distributions

Written by GPT-6.1 Sol (OpenAI) and GPT-6 Astra (OpenAI). Self-checked by the writing AI. Original exposition: CC0.

Working question: Does an observable's mean determine its spectral distribution? The two-point distribution equally supported at −1-1 and 11 has mean zero, as does the point mass at zero, but their second moments differ. A compressed observable therefore needs more than a first trace. Comparing powers of the compressed matrix with compressed powers of the operator, and controlling leakage through the energy cutoff, recovers the full limiting law.

A spectral subspace contains all states below a chosen energy. Compressing an observable to that subspace gives a finite matrix. We will show that the eigenvalue distribution of this matrix approaches the distribution of the observable's principal symbol over a cotangent energy region.

Laptev and Safarov [LS] supply the finite-projection comparison underlying this topic. Guillemin and Sternberg's text [GS] explains the semiclassical interpretation, and the article of Duistermaat and Guillemin [DG] develops the wave-trace setting. We use Return times and spectral counting and Local spectral density and the subprincipal correction. Composition and Sobolev mapping come from Classical scalar symbols, summation and regularity. The complete finite-dimensional spectral, singular-decomposition, trace-duality, ideal and cyclicity proofs are in Finite-rank traces and the exact ideal bounds, starting from elementary Hilbert-space arguments. Every product requiring those bounds below has a finite-rank factor. Section 5 gives the complete Bernstein polynomial-approximation proof used for continuous tests, as in [Alt]. The bounded scalar spectral measures are supplied by Self-adjoint spectral calculus with the original domain.

As before, XX is compact, connected and without boundary, of dimension n≥2n\geq2. The scalar operator P∈Ψcl1(X;Ω1/2)P\in\Psi^1_{\mathrm{cl}}(X;\Omega^{1/2}) is positive elliptic and self-adjoint on H1H^1, with positive principal symbol pp. Let B∈Ψcl0B\in\Psi^0_{\mathrm{cl}} be self-adjoint, with real principal symbol bb.

Use the convention Πλ=1(−∞,λ](P),Hλ=ΠλL2,N(λ)=dim⁡Hλ.(1) \begin{gathered} \Pi_\lambda=\mathbf1_{(-\infty,\lambda]}(P),\\ H_\lambda=\Pi_\lambda L^2, \\ N(\lambda)=\dim H_\lambda. \end{gathered} \tag{1} Every argument also works with the open endpoint. Symplectic volume on T∗XT^*X is denoted by dzdz.

1. The finite matrix and its counting measure

Define the compression as an operator on its finite-dimensional space: Bλ=ΠλBΠλ∣Hλ.(2) B_\lambda=\Pi_\lambda B\Pi_\lambda\big|_{H_\lambda}. \tag{2} It is self-adjoint and has norm at most ∥B∥\|B\|. If its eigenvalues, repeated with multiplicity, are β1,λ,…,βN(λ),λ\beta_{1,\lambda},\ldots,\beta_{N(\lambda),\lambda}, set ρλ=∑ℓ=1N(λ)δβℓ,λ.(3) \rho_\lambda=\sum_{\ell=1}^{N(\lambda)} \delta_{\beta_{\ell,\lambda}}. \tag{3} Thus ρλ(R)=N(λ)\rho_\lambda(\mathbb R)=N(\lambda), and ρλ(f)=Tr⁡Hλf(Bλ).(4) \rho_\lambda(f)=\operatorname{Tr}_{H_\lambda}f(B_\lambda). \tag{4} Zero eigenvalues inside HλH_\lambda are included. The infinite-dimensional orthogonal complement of HλH_\lambda is not part of (3).

We will use two finite-rank consequences of the trace-ideal estimates. If TT has rank at most rr, then ∥T∥1≤r ∥T∥2,∥T∥1≤r∥T∥.(5) \|T\|_1\leq \sqrt r\,\|T\|_2,\qquad \|T\|_1\leq r\|T\|. \tag{5} Indeed the trace norm is the sum of the eigenvalues of ∣T∣|T|, while the squared Hilbert–Schmidt norm is the sum of their squares. There are at most rr nonzero terms; Cauchy–Schwarz gives the first bound and the operator norm bounds each term for the second. We also use ∥ATC∥1≤∥A∥∥T∥1∥C∥,∣Tr⁡T∣≤∥T∥1.(6) \begin{gathered} \|ATC\|_1\leq\|A\|\|T\|_1\|C\|, \\ |\operatorname{Tr}T|\leq\|T\|_1. \end{gathered} \tag{6} These hold for the indicated source and target Hilbert spaces. All products used below have a finite-rank factor, so are trace class. The trace can be taken either on HλH_\lambda or on L2L^2 after extension by zero.

2. A weighted first moment

Proposition 2.1. For every self-adjoint D∈Ψcl0D\in\Psi^0_{\mathrm{cl}}, with principal symbol dd, Tr⁡(ΠλD)=(2π)−nλn∫p<1d dz+O(λn−1).(7) \operatorname{Tr}(\Pi_\lambda D) =(2\pi)^{-n}\lambda^n\int_{p<1}d\,dz +O(\lambda^{n-1}). \tag{7}

Proof. First arrange positivity. Choose L>1+max⁡(∥D∥,sup⁡∣d∣). L>1+\max(\|D\|,\sup|d|). Then D+LID+LI is positive and its principal symbol d+Ld+L is strictly positive. The cumulative trace μ(λ)=Tr⁡(Πλ(D+LI))=∑λj≤λ⟨(D+LI)ϕj,ϕj⟩(8) \begin{gathered} \mu(\lambda)=\operatorname{Tr}(\Pi_\lambda(D+LI)) \\ =\sum_{\lambda_j\leq\lambda} \langle(D+LI)\phi_j,\phi_j\rangle \end{gathered} \tag{8} is increasing, vanishes at zero and is O(λn)O(\lambda^n).

Use the exact small-time model from the counting lesson, Section 3. Write KD(t)=Tr⁡(e−itP(D+LI))K_D(t)=\operatorname{Tr}(e^{-itP}(D+LI)) as a distribution, and choose an even real χ∈Cc∞\chi\in C_c^\infty, equal to one near zero and supported inside the common small-time construction. Set ν′(s)=12π⟨χ(t)KD(t),eits⟩,ν(λ)=∫0λν′(s) ds. \begin{aligned} \nu'(s)&=\frac1{2\pi}\langle\chi(t)K_D(t),e^{its}\rangle,\\ \nu(\lambda)&=\int_0^\lambda\nu'(s)\,ds. \end{aligned} The distributional pairing is well defined because its first argument has compact support. The tested kernel identity in the counting lesson, Section 2, proves dμ^=KD\widehat{d\mu}=K_D. The local coefficient construction, including its smooth time residual, shows that ν\nu is smooth and has ν(λ)=(2π)−nλn∫p<1(d+L) dz+O(λn−1).(9) \nu(\lambda)=(2\pi)^{-n}\lambda^n \int_{p<1}(d+L)\,dz +O(\lambda^{n-1}). \tag{9} It is real because KD(−t)=KD(t)‾K_D(-t)=\overline{K_D(t)} and χ\chi is even and real; it is normalized at zero by definition. Its derivative has leading coefficient M0=n(2π)−n∫p<1(d+L) dz>0M_0=n(2\pi)^{-n}\int_{p<1}(d+L)\,dz>0. The lower homogeneous terms and the negative-energy tail can be absorbed into ∣ν′(λ)∣≤M0(∣λ∣+a0)n−1(10) |\nu'(\lambda)|\leq M_0(|\lambda|+a_0)^{n-1} \tag{10} by increasing a0a_0, exactly as in the counting lesson.

There is a uniform deleted small-time interval without any base return. Choose a fixed smaller interval ∣t∣<T|t|<T. The Fourier transform of dμ−dνd\mu-d\nu is exactly (1−χ)KD(1-\chi)K_D. It vanishes near zero, and away from zero the wavefront inclusion in the counting lesson excludes returns. Multiplication by φ^(t/T)\widehat\varphi(t/T) therefore gives a smooth compactly supported function. Repeated integration by parts shows that its inverse transform is Schwartz. Thus the smoothing error is controlled for the actual model, including the smooth residual, rather than only for its homogeneous coefficient expansion.

Apply the bounded-integrable version of the quantitative Tauberian lemma with a=1/Ta=1/T. Its comparison function is the smooth, hence absolutely continuous, representative ν\nu. We obtain μ(λ)−ν(λ)=O(λn−1).(11) \mu(\lambda)-\nu(\lambda)=O(\lambda^{n-1}). \tag{11} The counting result also gives N(λ)=(2π)−nλn∫p<1dz+O(λn−1),(12) N(\lambda)=(2\pi)^{-n}\lambda^n\int_{p<1}dz +O(\lambda^{n-1}), \tag{12} since the inverse-period function is bounded. Subtract LN(λ)LN(\lambda) from (8)–(11), and use (12). This proves (7). Finite-rank trace cyclicity identifies this trace with that of ΠλDΠλ\Pi_\lambda D\Pi_\lambda on HλH_\lambda. ∎

In particular (7) applies to D=BjD=B^j for every fixed integer j≥1j\geq1: it is self-adjoint, classical of order zero and has principal symbol bjb^j. What remains is to compare this uncompressed power with the power of the compressed matrix.

3. How much can cross an energy cutoff?

Let Qλ=I−ΠλQ_\lambda=I-\Pi_\lambda. The map QλBΠλQ_\lambda B\Pi_\lambda measures the part of an observable that moves a low-energy vector out of the low-energy subspace.

Lemma 3.1 (trace norm of the leakage). ∥QλBΠλ∥1=O(λn−1/2).(13) \|Q_\lambda B\Pi_\lambda\|_1 =O(\lambda^{n-1/2}). \tag{13}

Proof. The scalar calculus gives C=[P,B]∈Ψcl0,(14) C=[P,B]\in\Psi^0_{\mathrm{cl}}, \tag{14} so CC is bounded on L2L^2. The possible order-one leading product cancels because the principal symbols are scalar. Also B:H1→H1B:H^1\to H^1, so the commutator identity is valid on each smooth eigenvector.

Choose Λ=λ+Δ\Lambda=\lambda+\Delta, with 1≤Δ≤λ1\leq\Delta\leq\lambda, and split QλBΠλ=(ΠΛ−Πλ)BΠλ+QΛBΠλ.(15) Q_\lambda B\Pi_\lambda =(\Pi_\Lambda-\Pi_\lambda)B\Pi_\lambda +Q_\Lambda B\Pi_\lambda. \tag{15} The near term has rank at most N(Λ)−N(λ)N(\Lambda)-N(\lambda). From (12), for λ≥1\lambda\geq1 and Λ≤2λ\Lambda\leq2\lambda, N(Λ)−N(λ)≤Cλn−1(1+Δ).(16) N(\Lambda)-N(\lambda) \leq C\lambda^{n-1}(1+\Delta). \tag{16} Indeed the difference of the leading volume terms is bounded by Cλn−1ΔC\lambda^{n-1}\Delta, and the two remainders are O(λn−1)O(\lambda^{n-1}). Equations (5) and (16) give ∥(ΠΛ−Πλ)BΠλ∥1≤C∥B∥λn−1(1+Δ).(17) \|(\Pi_\Lambda-\Pi_\lambda)B\Pi_\lambda\|_1 \leq C\|B\|\lambda^{n-1}(1+\Delta). \tag{17}

For the far term, expand in the orthonormal PP-eigenbasis. If λμ>Λ\lambda_\mu>\Lambda and λν≤λ\lambda_\nu\leq\lambda, then (λμ−λν)⟨Bϕν,ϕμ⟩=⟨Cϕν,ϕμ⟩.(18) (\lambda_\mu-\lambda_\nu) \langle B\phi_\nu,\phi_\mu\rangle =\langle C\phi_\nu,\phi_\mu\rangle. \tag{18} Thus Parseval yields ∥QΛBΠλ∥22=∑λν≤λ∑λμ>Λ∣⟨Bϕν,ϕμ⟩∣2≤Δ−2∑λν≤λ∥Cϕν∥2≤Δ−2∥C∥2N(λ).(19) \begin{aligned} \|Q_\Lambda B\Pi_\lambda\|_2^2 &=\sum_{\lambda_\nu\leq\lambda} \sum_{\lambda_\mu>\Lambda} |\langle B\phi_\nu,\phi_\mu\rangle|^2\\ &\leq\Delta^{-2}\sum_{\lambda_\nu\leq\lambda} \|C\phi_\nu\|^2\\ &\leq \Delta^{-2}\|C\|^2N(\lambda). \end{aligned} \tag{19} This map has rank at most N(λ)N(\lambda). Equations (5) and (19) therefore imply ∥QΛBΠλ∥1≤Δ−1∥C∥N(λ)≤CΔ−1λn.(20) \|Q_\Lambda B\Pi_\lambda\|_1 \leq\Delta^{-1}\|C\|N(\lambda) \leq C\Delta^{-1}\lambda^n. \tag{20} The inequalities concern a finite-dimensional input, even though the far output space may be infinite-dimensional.

Combine (17) and (20): ∥QλBΠλ∥1≤C[λn−1(1+Δ)+λnΔ−1].(21) \|Q_\lambda B\Pi_\lambda\|_1 \leq C\left[\lambda^{n-1}(1+\Delta) +\lambda^n\Delta^{-1}\right]. \tag{21} Take Δ=λ1/2\Delta=\lambda^{1/2}. Both varying terms have order λn−1/2\lambda^{n-1/2}, proving (13). This argument uses the full counting bound with O(λn−1)O(\lambda^{n-1}) remainder, but needs no measure-zero hypothesis on periodic covectors. ∎

4. Comparing the two powers

The Hilbert–Schmidt leakage has a sharper bound that controls both crossings of the cutoff in each moment.

Lemma 4.1. Put L=QλBΠλL=Q_\lambda B\Pi_\lambda, M=∥B∥M=\|B\|, and C=[P,B]C=[P,B]. Then ∥L∥22≤K(1+λ)n−1(M2+2∥C∥2)=O(λn−1).(21a) \begin{gathered} \|L\|_2^2 \\ \le K(1+\lambda)^{n-1}\bigl(M^2+2\|C\|^2\bigr) \\ =O(\lambda^{n-1}). \end{gathered} \tag{21a}

Proof. Divide the input into unit bands Er=1(λ−r−1,λ−r](P),r=0,1,….(21b) E_r=\mathbf1_{(\lambda-r-1,\lambda-r]}(P), \qquad r=0,1,\ldots. \tag{21b} Their ranks satisfy rank⁡Er≤K(1+λ)n−1\operatorname{rank}E_r\le K(1+\lambda)^{n-1}, uniformly in rr and λ≥2\lambda\ge2. For bands whose upper endpoint is at least two, subtract the two Weyl expansions (12): the leading volume term changes by at most C(1+λ)n−1C(1+\lambda)^{n-1}, and both remainders have that order. Bands meeting a bounded energy range have uniformly bounded rank by compact resolvent. Bands below the positive spectrum have rank zero. This also handles eigenvalues at the displayed endpoints.

The bands are mutually orthogonal and sum to Πλ\Pi_\lambda, with finitely many nonzero terms. At r=0r=0, the squared Hilbert–Schmidt norm of QλBE0Q_\lambda BE_0 is at most M2rank⁡E0M^2\operatorname{rank}E_0. For r≥1r\ge1, every input eigenvalue is at most λ−r\lambda-r, and every output eigenvalue is greater than λ\lambda. The gap is greater than rr. The exact commutator identity (18) and Parseval give ∥QλBEr∥22≤r−2∑ϕν∈ErL2∥Cϕν∥2≤r−2∥C∥2rank⁡Er.(21c) \begin{aligned} \|Q_\lambda BE_r\|_2^2 &\le r^{-2}\sum_{\phi_\nu\in E_rL^2} \|C\phi_\nu\|^2\\ &\le r^{-2}\|C\|^2\operatorname{rank}E_r. \end{aligned} \tag{21c} The output may be infinite-dimensional. Only the input rank enters the bound. Sum over rr, and use ∑r≥1r−2≤2\sum_{r\ge1}r^{-2}\le2, obtained by comparing the tail with ∫1∞t−2dt\int_1^\infty t^{-2}dt. This proves (21a). For the open spectral endpoint, change the complementary band endpoints consistently; the same gap and rank bounds apply. ∎

When powers of BλB_\lambda are extended by zero to L2L^2, we write them as (ΠλBΠλ)j(\Pi_\lambda B\Pi_\lambda)^j for j≥1j\geq1.

Lemma 4.2. For each fixed integer j≥1j\geq1, ∥ΠλBjΠλ−(ΠλBΠλ)j∥1≤(j−1)∥B∥j−1∥QλBΠλ∥1.(22) \begin{gathered} \left\|\Pi_\lambda B^j\Pi_\lambda -(\Pi_\lambda B\Pi_\lambda)^j\right\|_1 \\ \leq (j-1)\|B\|^{j-1} \|Q_\lambda B\Pi_\lambda\|_1. \end{gathered} \tag{22} At j=1j=1 the difference is zero.

Proof. Abbreviate Π=Πλ\Pi=\Pi_\lambda, Q=I−ΠQ=I-\Pi, and let the difference in (22) be FjF_j. Inserting I=Π+QI=\Pi+Q just before the final BB gives the exact recurrence Fj=Fj−1BΠ+ΠBj−1QBΠ,F1=0.(23) F_j=F_{j-1}B\Pi+\Pi B^{j-1}QB\Pi, \qquad F_1=0. \tag{23} All terms have finite rank. The ideal bound (6) gives ∥Fj∥1≤∥B∥∥Fj−1∥1+∥B∥j−1∥QBΠ∥1. \|F_j\|_1 \leq\|B\|\|F_{j-1}\|_1 +\|B\|^{j-1}\|QB\Pi\|_1. Induction proves (22), including B=0B=0, where all differences vanish. ∎

Lemma 4.3. For every integer j≥2j\ge2, ∥Fj∥1≤j(j−1)2Mj−2∥L∥22=Oj(λn−1).(23a) \|F_j\|_1 \le \frac{j(j-1)}2 M^{j-2}\|L\|_2^2 =O_j(\lambda^{n-1}). \tag{23a} At j=1j=1 the difference is zero. If B=0B=0, all differences vanish directly.

Proof. Set Ra=QBaΠR_a=QB^a\Pi. Inserting Π+Q=I\Pi+Q=I immediately before the last BB gives Ra=Ra−1BΠ+QBa−1Q L,∥Ra∥2≤aMa−1∥L∥2.(23b) \begin{gathered} R_a=R_{a-1}B\Pi+QB^{a-1}Q\,L,\\ \|R_a\|_2\le aM^{a-1}\|L\|_2. \end{gathered} \tag{23b} The second line follows by induction from the Hilbert–Schmidt ideal inequality and R1=LR_1=L. Self-adjointness identifies the second term in (23) as Rj−1∗LR_{j-1}^*L. Products of two Hilbert–Schmidt maps obey ∥Rj−1∗L∥1≤∥Rj−1∥2∥L∥2≤(j−1)Mj−2∥L∥22.(23c) \begin{gathered} \|R_{j-1}^*L\|_1 \le \|R_{j-1}\|_2\|L\|_2 \\ \le (j-1)M^{j-2}\|L\|_2^2. \end{gathered} \tag{23c} Here every input rank is finite. To see the first inequality directly, trace-norm duality tests the product against a contraction AA; the absolute trace is the Hilbert–Schmidt pairing of Rj−1A∗R_{j-1}A^* with LL, bounded by the product of their norms. Finite-dimensional singular-value decomposition proves this duality for these finite-rank products.

The recurrence therefore gives ∥Fj∥1≤M∥Fj−1∥1+(j−1)Mj−2∥L∥22\|F_j\|_1\le M\|F_{j-1}\|_1+(j-1)M^{j-2}\|L\|_2^2. Induction sums 1+⋯+(j−1)=j(j−1)/21+\cdots+(j-1)=j(j-1)/2, proving (23a). Both cutoff crossings contribute a Hilbert–Schmidt factor. ∎

Taking traces and using (21a) and (23a), ∣Tr⁡HλBλj−Tr⁡(ΠλBjΠλ)∣=Oj(λn−1).(24) \left| \operatorname{Tr}_{H_\lambda}B_\lambda^j -\operatorname{Tr}(\Pi_\lambda B^j\Pi_\lambda) \right| =O_j(\lambda^{n-1}). \tag{24} Together with (7), this gives every moment: λ−nρλ(sj)⟶(2π)−n∫p<1b(z)j dz,j=0,1,2,….(25) \begin{gathered} \lambda^{-n}\rho_\lambda(s^j) \\ \longrightarrow (2\pi)^{-n}\int_{p<1}b(z)^j\,dz, j=0,1,2,\ldots. \end{gathered} \tag{25} For j=0j=0 the left side is λ−nN(λ)\lambda^{-n}N(\lambda); the zeroth power is the identity on HλH_\lambda. Equation (12) proves that case.

5. The limiting distribution

Define the finite positive measure ρ=b∗ ⁣((2π)−n1{p<1} dz).(26) \rho=b_*\!\left((2\pi)^{-n}\mathbf1_{\{p<1\}}\,dz\right). \tag{26} The homogeneous symbol bb is defined off the zero section; the zero section has volume zero, so its value there is irrelevant. Equivalently, ρ(f)=(2π)−n∫p<1f(b(z)) dz.(27) \rho(f)=(2\pi)^{-n}\int_{p<1}f(b(z))\,dz. \tag{27}

Theorem 5.1 (the distribution of a compressed observable). For every f∈C(R)f\in C(\mathbb R), λ−nρλ(f)⟶ρ(f).(28) \lambda^{-n}\rho_\lambda(f)\longrightarrow\rho(f). \tag{28} For each fixed polynomial qq, the difference in (28) is Oq(λ−1)O_q(\lambda^{-1}). No assumption about the measure of periodic trajectories is required.

Proof. Choose R>max⁡(∥B∥,sup⁡∣b∣). R>\max(\|B\|,\sup|b|). Both measures are supported in the common compact interval [−R,R][-R,R]. Their masses are bounded after the required normalization: λ−nρλ(R)=λ−nN(λ)=O(1),ρ(R)=(2π)−n∫p<1dz<∞.(29) \begin{gathered} \lambda^{-n}\rho_\lambda(\mathbb R) =\lambda^{-n}N(\lambda)=O(1),\\ \rho(\mathbb R)=(2\pi)^{-n}\int_{p<1}dz<\infty. \end{gathered} \tag{29}

Here is the complete polynomial-approximation step. Put g(x)=f(2Rx−R)g(x)=f(2Rx-R) on [0,1][0,1]. A sequence in this interval has a convergent subsequence by nested bisection, as in the finite compactness proof in the trace provider. If gg were unbounded, a sequence with ∣g(xj)∣>j|g(x_j)|>j would have such a subsequence, contradicting continuity at its limit. If gg were not uniformly continuous, there would be ε0>0\varepsilon_0>0 and pairs xj,yjx_j,y_j with ∣xj−yj∣→0|x_j-y_j|\to0 but ∣g(xj)−g(yj)∣≥ε0|g(x_j)-g(y_j)|\ge\varepsilon_0. A convergent subsequence of xjx_j makes both points tend to the same limit, again contradicting continuity. Write Mg=sup⁡∣g∣M_g=\sup|g| and ωg(δ)=sup⁡∣x−y∣≤δ∣g(x)−g(y)∣,ωg(δ)⟶0. \omega_g(\delta)=\sup_{|x-y|\le\delta}|g(x)-g(y)|, \qquad\omega_g(\delta)\longrightarrow0. For an integer m≥1m\ge1, form the Bernstein polynomial wm,k(x)=(mk)xk(1−x)m−k,Bmg(x)=∑k=0mg(k/m)wm,k(x). \begin{gathered} w_{m,k}(x)=\binom mk x^k(1-x)^{m-k},\\ \mathcal B_mg(x)=\sum_{k=0}^m g(k/m)w_{m,k}(x). \end{gathered} The weights are nonnegative and sum to one by the binomial formula. The identities k(mk)=m(m−1k−1)k\binom mk=m\binom{m-1}{k-1} and k(k−1)(mk)=m(m−1)(m−2k−2)k(k-1)\binom mk=m(m-1)\binom{m-2}{k-2}, followed by the binomial formula again, give ∑k(k/m)wm,k(x)=x,∑k(k/m)2wm,k(x)=x2+x(1−x)/m,∑k(k/m−x)2wm,k(x)=x(1−x)/m≤1/(4m). \begin{aligned} \sum_k(k/m)w_{m,k}(x)&=x,\\ \sum_k(k/m)^2w_{m,k}(x)&=x^2+x(1-x)/m,\\ \sum_k(k/m-x)^2w_{m,k}(x)&=x(1-x)/m\le1/(4m). \end{aligned} For m=1m=1 the second factorial moment is zero and these formulas follow directly; the endpoint values are the continuous polynomial values. The total weight of indices with ∣k/m−x∣>δ|k/m-x|>\delta is at most 1/(4mδ2)1/(4m\delta^2). Splitting the approximation error into this set and its complement gives sup⁡0≤x≤1∣Bmg(x)−g(x)∣≤ωg(δ)+Mg2mδ2. \sup_{0\le x\le1}|\mathcal B_mg(x)-g(x)| \le\omega_g(\delta)+\frac{M_g}{2m\delta^2}. First choose δ>0\delta>0 small, then mm large. Composing with x=(s+R)/(2R)x=(s+R)/(2R) produces a polynomial q(s)q(s) with sup⁡[−R,R]∣f−q∣<ε\sup_{[-R,R]}|f-q|<\varepsilon. This proof works for complex-valued ff as well, and gives real coefficients when ff is real. This is the Bernstein construction [Alt, §3, equation (3.5) and Theorem 3.6], with an explicit error bound. Now ∣λ−nρλ(f)−ρ(f)∣≤ε[λ−nN(λ)+ρ(R)]+∣λ−nρλ(q)−ρ(q)∣.(30) \begin{aligned} |\lambda^{-n}\rho_\lambda(f)-\rho(f)| &\leq \varepsilon\left[\lambda^{-n}N(\lambda)+\rho(\mathbb R)\right]\\ &\quad+|\lambda^{-n}\rho_\lambda(q)-\rho(q)|. \end{aligned} \tag{30} The last term tends to zero by (25). Taking the upper limit and then ε↓0\varepsilon\downarrow0 proves (28). For a fixed polynomial, (7) and (24) give the stated rate, by summing its finitely many coefficients. Continuous approximation asserts convergence without a universal rate for arbitrary ff. ∎

There is also a probability version. The coefficient cP=(2π)−n∫p<1dz c_P=(2\pi)^{-n}\int_{p<1}dz is positive. Dividing (28) by N(λ)/λn→cPN(\lambda)/\lambda^n\to c_P gives ρλN(λ) ⟶ ρcP.(31) \frac{\rho_\lambda}{N(\lambda)} \ \longrightarrow\ \frac{\rho}{c_P}. \tag{31} The right side samples bb using normalized symplectic volume on the energy region. For a fixed polynomial, this probability-normalized convergence also has error Oq(λ−1)O_q(\lambda^{-1}): its numerator is λnρ(q)+Oq(λn−1)\lambda^n\rho(q)+O_q(\lambda^{n-1}), and (12) gives the denominator cPλn+O(λn−1)c_P\lambda^n+O(\lambda^{n-1}), with cP>0c_P>0. Continuous tests retain convergence without a universal rate.

All conclusions remain valid for a lower-bounded PP with the same positive principal symbol. To see this, choose cc such that P+c>0P+c>0. Its projection at λ+c\lambda+c is exactly Πλ\Pi_\lambda, and its commutator with BB is still [P,B][P,B]. Its principal symbol is unchanged. Since (λ+c)n/λn→1(\lambda+c)^n/\lambda^n\to1, applying the positive results at energy λ+c\lambda+c proves (28) and (31). The weighted first-moment error and leakage estimates retain their stated orders under this fixed translation.

5A. Smooth tests, compression error and stable distributions

Laptev and Safarov, Szegö type limit theorems, Section 1, Theorem 1.2, relates smooth functional calculus to leakage through an orthogonal projection. Here is a direct proof of its bounded-operator, finite-rank case, followed by a stability consequence. It keeps distinct the compression error and the asymptotic symbol calculation.

Proposition 5.2 (a smooth-test compression bound). Let B=B∗B=B^* be bounded, let Π\Pi be a finite-rank orthogonal projection, put Q=I−ΠQ=I-\Pi, and let A=ΠBΠ∣ran⁡ΠA=\Pi B\Pi|_{\operatorname{ran}\Pi}. If ff is real and C2C^2 on an interval containing [−∥B∥,∥B∥][-\|B\|,\|B\|], then

Df=Tr⁡ran⁡Π(Πf(B)Π−f(A)),∣Df∣≤12∥f′′∥∞∥QBΠ∥HS2. \begin{gathered} D_f=\operatorname{Tr}_{\operatorname{ran}\Pi} \bigl(\Pi f(B)\Pi-f(A)\bigr),\\ |D_f|\le \frac12\|f''\|_\infty\|QB\Pi\|_{\mathrm{HS}}^2. \end{gathered}

The trace on the left is nonnegative when ff is convex.

The same statement holds for f∈C1,1f\in C^{1,1}, with Lip⁡(f′)\operatorname{Lip}(f') in place of ∥f′′∥∞\|f''\|_\infty. This includes the continuously differentiable representative of a real W2,∞W^{2,\infty} function on the interval, which is the regularity used in [LS, Theorem 1.2].

Proof. Choose an orthonormal eigenbasis e1,…,ede_1,\ldots,e_d of AA, with eigenvalues μj\mu_j. The scalar spectral calculus for the bounded self-adjoint BB supplies probability measures σj\sigma_j with

⟨f(B)ej,ej⟩=∫f(t) dσj(t),∫t dσj(t)=μj,∫(t−μj)2dσj(t)=∥(B−μj)ej∥2. \begin{gathered} \langle f(B)e_j,e_j\rangle=\int f(t)\,d\sigma_j(t),\\ \int t\,d\sigma_j(t)=\mu_j,\\ \int(t-\mu_j)^2d\sigma_j(t)=\|(B-\mu_j)e_j\|^2. \end{gathered}

The scalar spectral measures come from Self-adjoint spectral calculus with the original domain. Only its bounded-operator restriction is needed here. Taylor's formula with integral remainder gives

∣f(t)−f(μj)−f′(μj)(t−μj)∣≤12∥f′′∥∞(t−μj)2. |f(t)-f(\mu_j)-f'(\mu_j)(t-\mu_j)| \le\frac12\|f''\|_\infty(t-\mu_j)^2.

For convex ff, apply its defining inequality at μj+θ(t−μj)\mu_j+\theta(t-\mu_j), subtract f(μj)f(\mu_j), divide by θ>0\theta>0, and let θ↓0\theta\downarrow0. Differentiability gives f(t)−f(μj)≥f′(μj)(t−μj)f(t)-f(\mu_j)\geq f'(\mu_j)(t-\mu_j), so the expression before the absolute value is nonnegative. Its linear term integrates to zero. Since ΠBej=μjej\Pi Be_j=\mu_je_j, one has (B−μj)ej=QBej(B-\mu_j)e_j=QBe_j. Sum over jj; the sum of these squared norms is exactly ∥QBΠ∥HS2\|QB\Pi\|_{\mathrm{HS}}^2. This proves both assertions. □\square

For the stated C1,1C^{1,1} extension, the fundamental theorem of calculus gives f(t)−f(μ)−f′(μ)(t−μ)=∫μt(f′(r)−f′(μ)) dr. f(t)-f(\mu)-f'(\mu)(t-\mu) =\int_\mu^t\bigl(f'(r)-f'(\mu)\bigr)\,dr. Its absolute value is at most Lip⁡(f′)∣t−μ∣2/2\operatorname{Lip}(f')|t-\mu|^2/2, for either order of the two endpoints. The preceding spectral-measure argument therefore applies unchanged. The constant 1/21/2 and the convex sign require no approximation of operators or differentiation of their spectral projections.

Here is the asserted weak-Sobolev representative. On a bounded interval write the distributional derivatives as g=f′∈L∞g=f'\in L^\infty and h=g′∈L∞h=g'\in L^\infty. Fix an interior point aa and put G(t)=∫ath(r) drG(t)=\int_a^t h(r)\,dr. The scalar primitive identity in Hilbert-valued integration, Section 3 gives G′=hG'=h distributionally, and ∣G(t)−G(s)∣≤∥h∥∞∣t−s∣|G(t)-G(s)|\leq\|h\|_\infty|t-s|. Thus g−Gg-G has zero distributional derivative. A locally integrable function with this property is a constant distribution: subtract a fixed integral-one test from any other test to make its integral zero, and use the compactly supported primitive of that difference. Consequently gg has a Lipschitz representative g~\widetilde g. Applying the same argument to f−∫atg~(r) drf-\int_a^t\widetilde g(r)\,dr gives a C1C^1 representative f~\widetilde f, with f~′=g~\widetilde f'=\widetilde g and Lip⁡(f~′)≤∥f′′∥∞\operatorname{Lip}(\widetilde f')\leq\|f''\|_\infty. Both representatives extend continuously to the endpoints. This proves the claimed W2,∞W^{2,\infty} case with the same constant.

For f(t)=t2f(t)=t^2 equality holds with the positive sign:

Tr⁡(ΠB2Π−A2)=∥QBΠ∥HS2. \operatorname{Tr}(\Pi B^2\Pi-A^2)=\|QB\Pi\|_{\mathrm{HS}}^2.

For the spectral cutoff of this lesson, the sharper squared Hilbert–Schmidt bound in Lemma 4.1 is ∥QλBΠλ∥HS2=O(λn−1)\|Q_\lambda B\Pi_\lambda\|_{\mathrm{HS}}^2=O(\lambda^{n-1}). Proposition 5.2 therefore gives a normalized compression defect Of(λ−1)O_f(\lambda^{-1}) for every fixed C2C^2 test. This controls the difference between the two operator traces. A rate for either trace against the classical symbol law still requires its separate symbol estimate; the defect bound does not create such an estimate.

Theorem 5.3 (a limit law survives small trace-norm changes). Let Ak,CkA_k,C_k be self-adjoint matrices of the same dimension dk≥1d_k\geq1, with ∥Ak∥,∥Ck∥≤M<∞\|A_k\|,\|C_k\|\le M<\infty, and suppose

εk=dk−1∥Ak−Ck∥1⟶0. \varepsilon_k=d_k^{-1}\|A_k-C_k\|_1\longrightarrow0.

If the probability eigenvalue measures of AkA_k converge to ν\nu, then those of CkC_k converge to the same ν\nu, against every continuous function. In particular, bounded changes of rank rk=o(dk)r_k=o(d_k) preserve the law. Counts in an interval also converge whenever ν\nu assigns zero mass to both endpoints.

Proof. For every integer m≥1m\ge1, telescoping in the given order gives

Akm−Ckm=∑j=0m−1Akm−1−j(Ak−Ck)Ckj. A_k^m-C_k^m=\sum_{j=0}^{m-1}A_k^{m-1-j}(A_k-C_k)C_k^j.

The trace-ideal inequality gives

dk−1∣Tr⁡(Akm−Ckm)∣≤mMm−1εk. d_k^{-1}|\operatorname{Tr}(A_k^m-C_k^m)| \le mM^{m-1}\varepsilon_k.

For a fixed polynomial p(t)=∑amtmp(t)=\sum a_mt^m, sum these bounds. Approximate a continuous ff uniformly on [−M,M][-M,M] by such a polynomial, using the Bernstein approximation already used in Theorem 5.1. The two probability masses are one, so the approximation contributes at most 2∥f−p∥∞2\|f-p\|_\infty. First take k→∞k\to\infty, then the approximation error to zero. This proves the common law without assigning a universal rate to a continuous test.

If Ak−CkA_k-C_k has rank at most rkr_k, then ∥Ak−Ck∥1≤2Mrk\|A_k-C_k\|_1\le2Mr_k, proving the rank assertion. For a closed interval [a,b][a,b], bound its indicator above and below by continuous piecewise linear functions which change from zero to one only within distance δ\delta of its endpoints. The difference of their limiting integrals is bounded by ν([a−δ,a+δ]∪[b−δ,b+δ])\nu([a-\delta,a+\delta]\cup[b-\delta,b+\delta]). This tends to zero when the endpoints have no atoms. The same argument covers open or half-open conventions. □\square

As a worked application, take Aλ=ΠλBΠλA_\lambda=\Pi_\lambda B\Pi_\lambda and add a uniformly bounded self-adjoint matrix EλE_\lambda with rank o(N(λ))o(N(\lambda)). Theorem 5.1 and Theorem 5.3 give the same symbol distribution for Aλ+EλA_\lambda+E_\lambda, even when the matrix modification has no pseudodifferential symbol. This is a statement about the fraction of states, not about the location of every eigenvalue: a rank-one change can leave a persistent outlier. For example Ak=0A_k=0 and Ck=diag⁡(1,0,…,0)C_k=\operatorname{diag}(1,0,\ldots,0) have a common limit δ0\delta_0, while their largest eigenvalues remain zero and one.

Use the conclusion

Check the cutoff-crossing estimate before taking higher moments. Identify the compact interval on which polynomial approximation is used, and then recover the distribution of the principal symbol rather than only its mean.

6. Exercises and complete solutions

Exercise 6.1 (what is being counted; introductory). Let B=cIB=cI, with real cc. Find ρλ\rho_\lambda and its limit. Explain why counting all eigenvalues of ΠλBΠλ\Pi_\lambda B\Pi_\lambda on ambient L2L^2 would give a different, unsuitable object.

Solution 6.1. On HλH_\lambda, the compression is cIHλcI_{H_\lambda}, so ρλ=N(λ)δc,λ−nρλ⟶cPδc. \rho_\lambda=N(\lambda)\delta_c,\qquad \lambda^{-n}\rho_\lambda\longrightarrow c_P\delta_c. Its symbol is the constant cc, and (26) gives exactly the same measure. On the orthogonal complement, the ambient operator is zero. That complement is infinite-dimensional, so counting its zero eigenvalue would add an infinite mass at zero. When c=0c=0, the intended measure still has just N(λ)N(\lambda) zeros, whereas the ambient zero operator has infinitely many. The finite spectral subspace in (2) is part of the definition.

Exercise 6.2 (balancing a strip and a gap; intermediate). In (21), choose Δ=λθ\Delta=\lambda^\theta, 0≤θ≤10\leq\theta\leq1. Determine the best exponent supplied by this family of choices.

Solution 6.2. The two varying exponents are n−1+θn-1+\theta and n−θn-\theta; the fixed term has exponent n−1n-1. Hence the estimate has exponent max⁡(n−1+θ,n−θ). \max(n-1+\theta,n-\theta). The first increases with θ\theta, the second decreases. Their intersection is θ=1/2\theta=1/2, with value n−1/2n-1/2. At either endpoint the maximum is nn. Thus the square-root gap is optimal for this particular trace-norm bound. If [P,B]=0[P,B]=0, leakage vanishes exactly and the gap argument is unnecessary.

Exercise 6.3 (a finite-rank change of observable; intermediate). Let K=K∗K=K^* have finite rank. Compare the moments of the compressions of B+KB+K and BB, after division by N(λ)N(\lambda). Prove directly that their normalized empirical distributions have the same limit whenever one of them converges.

Solution 6.3. The compressed difference is Kλ=ΠλKΠλ∣HλK_\lambda=\Pi_\lambda K\Pi_\lambda|_{H_\lambda}, with ∥Kλ∥1≤∥K∥1\|K_\lambda\|_1\leq\|K\|_1. Let Aλ=Bλ+KλA_\lambda=B_\lambda+K_\lambda and choose a constant MM bounding both operator norms. For j≥1j\geq1, the noncommutative identity Aλj−Bλj=∑r=0j−1Aλj−1−rKλBλr A_\lambda^j-B_\lambda^j =\sum_{r=0}^{j-1}A_\lambda^{j-1-r} K_\lambda B_\lambda^r gives trace norm at most jMj−1∥K∥1jM^{j-1}\|K\|_1, uniformly in λ\lambda. Since N(λ)∼cPλn→∞N(\lambda)\sim c_P\lambda^n\to\infty, the normalized moment differences tend to zero. The zeroth moments are both one. Both empirical probability measures have support in a fixed interval, so polynomial approximation and the argument of (30) give the claim for every continuous test function. This direct argument does not require KK to be pseudodifferential.

Exercise 6.4 (a stronger estimate for the second moment; advanced). Show Tr⁡(ΠλB2Πλ)−Tr⁡HλBλ2=∥QλBΠλ∥22≥0.(32) \operatorname{Tr}(\Pi_\lambda B^2\Pi_\lambda) -\operatorname{Tr}_{H_\lambda}B_\lambda^2 =\|Q_\lambda B\Pi_\lambda\|_2^2\geq0. \tag{32} First use the strip/gap decomposition to bound this difference by O(λn−2/3)O(\lambda^{n-2/3}). Then use the unit input bands to obtain O(λn−1)O(\lambda^{n-1}). Explain how the squared leakage controls every higher fixed power.

Solution 6.4. Self-adjointness gives ΠλB2Πλ−(ΠλBΠλ)2=ΠλBQλBΠλ=(QλBΠλ)∗(QλBΠλ). \begin{gathered} \Pi_\lambda B^2\Pi_\lambda -(\Pi_\lambda B\Pi_\lambda)^2 \\ =\Pi_\lambda BQ_\lambda B\Pi_\lambda \\ =(Q_\lambda B\Pi_\lambda)^*(Q_\lambda B\Pi_\lambda). \end{gathered} Taking the finite trace proves (32). The near and far terms in (15) have orthogonal output spaces, so their squared Hilbert–Schmidt norms add. The near term has squared norm at most ∥B∥2[N(Λ)−N(λ)]\|B\|^2[N(\Lambda)-N(\lambda)]; use a basis in its finite-dimensional output space and adjoint equality of the Hilbert–Schmidt norm to see this. Formula (19) bounds the far term. Thus ∥QλBΠλ∥22≤C[λn−1(1+Δ)+λnΔ−2]. \|Q_\lambda B\Pi_\lambda\|_2^2 \leq C\left[\lambda^{n-1}(1+\Delta) +\lambda^n\Delta^{-2}\right]. Choose Δ=λ1/3\Delta=\lambda^{1/3}. The two varying terms are both O(λn−2/3)O(\lambda^{n-2/3}), while the fixed term is smaller. The unit input bands of Lemma 4.1 instead bound the squared leakage by (21a), namely O(λn−1)O(\lambda^{n-1}). Each nonzero band uses its input rank and the inverse-square gap, whose sum is finite. Finally the second crossing in the recurrence (23) is Rj−1∗LR_{j-1}^*L; (23b)–(23c) bound its trace norm by (j−1)Mj−2∥L∥22(j-1)M^{j-2}\|L\|_2^2. Summing the recurrence gives (23a) for every fixed higher power. The one-sided trace-norm bound (13) remains a separate valid estimate.

Exercise 6.5 (an angular observable on a flat torus; advanced). On the two-dimensional torus (R/2πZ)2(\mathbb R/2\pi\mathbb Z)^2, let PP be a positive Fourier multiplier with high-frequency symbol ∣η∣|\eta|. Let BB be a self-adjoint Fourier multiplier with high-frequency symbol b(η)=η1/∣η∣b(\eta)=\eta_1/|\eta|. Find the limiting probability measure in (31), and explain why leakage is zero.

Solution 6.5. Section 6 of Local spectral density and the subprincipal correction proves the classical torus multiplier construction and completeness of the Fourier basis. It applies both to the smooth order-one symbol of PP and to the smooth order-zero extension of η1/∣η∣\eta_1/|\eta| away from the finitely many small modes. The two multipliers are diagonal in this same basis, so BB preserves every PP-spectral subspace. Thus QλBΠλ=0Q_\lambda B\Pi_\lambda=0, and compressed and uncompressed powers agree exactly.

The torus volume is (2π)2(2\pi)^2, canceling the Fourier factor in (27). In polar coordinates η=r(cos⁡θ,sin⁡θ)\eta=r(\cos\theta,\sin\theta), ρ(f)=∫01r dr∫02πf(cos⁡θ) dθ=∫−11f(s)1−s2 ds.(33) \begin{gathered} \rho(f)=\int_0^1r\,dr\int_0^{2\pi}f(\cos\theta)\,d\theta \\ =\int_{-1}^1\frac{f(s)}{\sqrt{1-s^2}}\,ds. \end{gathered} \tag{33} The last equality follows by splitting the angular integral into the two half-circles and substituting s=cos⁡θs=\cos\theta in each. The radial integral is 1/21/2. Consequently cP=πc_P=\pi and the limiting probability measure has density 1π1−s21{−1<s<1} ds.(34) \frac{1}{\pi\sqrt{1-s^2}}\mathbf1_{\{-1<s<1\}}\,ds. \tag{34} Its endpoint singularities are integrable and give no endpoint atoms. Finitely many choices of small Fourier modes do not affect the limit, as can also be seen from Exercise 6.3. The limiting measure comes from angular volume, even though each finite compression has a discrete spectrum.

References

[LS, §1, Theorems 1.2, 1.3 and 1.5–1.6; Appendix A] proves the compression inequality and the separated-band estimates. Its closed-manifold application [LS, §2, Theorem 2.2 and Lemma 2.3] uses an additional weighted spectral asymptotic and a smooth functional calculus. Here Proposition 2.1 supplies the weighted asymptotic from the preceding course proofs, while the moment argument supplies all continuous tests without invoking that additional smooth calculus. Zelditch [Z, §0] discusses the law of a single Zoll cluster; [Z, §4] computes its geometric band invariant. Those are different spectral subspaces from the cumulative cutoff in (1).