Trace inequalities for finite von Neumann algebras

Written by Claude Opus 5.5 (Anthropic), September 2026, revised October 2026. The September text was spot-checked by Claude Opus 5.5 in a separate session; the October revisions are self-checked by the writing AI. Public domain (CC0).

Introduction

In a von Neumann algebra with a trace, the trace turns operators into "noncommutative functions", and the spaces \(L^1\) and \(L^2\) of integrable and square-integrable operators behave in many ways like their classical namesakes. This lesson proves a handful of quantitative inequalities in these spaces. They are the analytic tools behind the classification of injective factors: they turn "almost commuting" information measured in the trace norm into projections, and into unitaries, that are close to what one wants.

The lesson has four themes.

  1. Joint distribution of two positive operators. Two positive operators \(h,k\) do not commute, so they have no joint spectral measure inside the algebra. Nevertheless there is a measure \(\mu\) on the plane which computes every \(L^2\) distance \(\|f(h)-g(k)\|_2\) exactly (Theorem 2.1). The reason is that on \(L^2\), left multiplication by \(h\) and right multiplication by \(k\) commute. The same measure does not compute \(L^1\) distances (Example 2.2).
  2. The Powers–Størmer inequality. For positive \(h,k\) in \(L^2\), \(\|h-k\|_2^2\le\|h^2-k^2\|_1\) (Theorem 3.1). So the square root is Hölder continuous of order \(\tfrac12\) from \(L^1\) to \(L^2\), and the modulus \(x\mapsto|x|\) is continuous on \(L^2\) with an explicit modulus of continuity.
  3. Integrating over the level. The spectral projection \(E_a(h)=\chi_{(a,\infty)}(h)\) depends very discontinuously on \(h\), but after integrating over the level \(a\) it becomes continuous in \(L^2\) (Propositions 4.2 and 4.3). The \(L^1\) analogue holds only in type I algebras of bounded degree (Theorem 5.1). From these integrated estimates we get the main result: for a finite family \(x_1,\dots,x_n\) of nearby vectors in \(L^2\) there is a common level \(a\) at which the truncated polar parts \(u_a(x_j)\) are all close to each other (Theorem 7.1). Applied to conjugates of a positive operator, this gives a projection that almost commutes with a given finite set of unitaries (Theorem 8.2).
  4. Two projections. If \(e\) and \(f\) are equivalent finite projections, there is a unitary \(W\) with \(WeW^*=f\) and \(|W-1|\le\sqrt2\,|e-f|\), and \(W\) commutes with \(|e-f|\) (Theorem 9.4). The constant \(\sqrt2\) is optimal.

Although the course uses these results for finite von Neumann algebras (II₁ factors and their ultrapowers), all proofs work for a faithful normal semifinite trace, and we state them in that generality. The finite case is the special case \(\tau(1)<\infty\).

Later lessons of this course use Theorem 3.1 (with Corollary 3.2), Theorem 7.1 (with Corollary 7.2), Theorem 8.2 (with Corollary 8.3) and Theorem 9.4 (with Corollary 9.5); among them are "Property \(\Gamma\) and the algebra generated by a factor and its commutant", "Approximately inner and centrally trivial automorphisms" and "Uniqueness of the injective II₁ factor".

What is assumed. Basic von Neumann algebra theory, as in the lessons "Projections and types of von Neumann algebras" and "Traces on von Neumann algebras"; noncommutative integration for a trace, as in "Measurable operators for a trace: examples, convergence and the commutant" and "Trace densities and noncommutative integration"; and the spectral theorem with joint spectral measures, as in "Spectral calculus with its domains retained" and "Joint spectral measures and commutator estimates in standard form". These are lessons of this programme. The facts used are listed precisely below.

Basic references are [Connes 1976], [Powers–Størmer 1970], [Haagerup 1975], [Fack–Kosaki 1986], [Connes 1994] and [Anantharaman–Popa].

Results used from other lessons

Throughout, \(N\) is a von Neumann algebra and \(\tau\) is a faithful normal semifinite trace on \(N\). All functions are Borel functions.

(B1) Measurable operators and \(L^p\). A closed densely defined operator \(x\) affiliated with \(N\) is \(\tau\)-measurable if \(\tau(\chi_{(\lambda,\infty)}(|x|))<\infty\) for some \(\lambda>0\). The \(\tau\)-measurable operators form a \(*\)-algebra under the closure of the sum and of the product, containing \(N\), and the usual algebraic rules hold. For a positive self-adjoint operator \(h\) affiliated with \(N\) put \[ \nu_h(B)=\tau(\chi_B(h))\qquad(B\subset(0,\infty)\ \text{Borel}), \] a positive measure on \((0,\infty)\) (by normality of \(\tau\)), and for a function \(f\ge0\) on \([0,\infty)\) with \(f(0)=0\) put \(\tau(f(h))=\int_{(0,\infty)}f\,d\nu_h\in[0,\infty]\). This extends \(\tau\) from \(N^+\), and is additive and monotone on positive \(\tau\)-measurable operators. For \(1\le p<\infty\), \(L^p(N,\tau)\) is the space of \(\tau\)-measurable \(x\) with \(\|x\|_p=\tau(|x|^p)^{1/p}<\infty\); it is a Banach space, \(N\cap L^p\) is dense in it, and \(\|x^*\|_p=\||x|\|_p=\|x\|_p\). The space \(L^2(N,\tau)\) is a Hilbert space with \(\langle x,y\rangle=\tau(y^*x)\). If \(h\ge0\) is affiliated with \(N\) and \(f(0)=0\), then \(f(h)\in L^p\) if and only if \(\int|f|^p\,d\nu_h<\infty\), and then \(\|f(h)\|_p^p=\int|f|^p\,d\nu_h\). The \(\tau\)-measurable operators and their algebra are constructed in Operators recovered from small trace defects (§§MT-05–MT-12), and the spaces \(L^p(N,\tau)\), their completeness, density of \(N\cap L^p\) and the Hilbert space \(L^2\) in Trace densities and noncommutative integration, §§TI-09 and TI-15. The formulas for \(\tau(f(h))\) are the spectral theorem for \(h\) together with normality of \(\tau\); see also Measurable operators for a trace, Section 9. See also [Fack–Kosaki 1986].

(B2) Hölder inequality and cyclicity. If \(1\le p,q,r<\infty\), \(1/p+1/q=1/r\), \(x\in L^p\), \(y\in L^q\), then \(xy\in L^r\) and \(\|xy\|_r\le\|x\|_p\|y\|_q\). If \(a,b\in N\) and \(x\in L^p\), then \(\|axb\|_p\le\|a\|\|x\|_p\|b\|\). The trace extends to a positive linear functional on \(L^1\) with \(|\tau(x)|\le\|x\|_1\); if \(1/p+1/q=1\) (allowing \(q=\infty\), with \(L^\infty=N\)), \(x\in L^p\) and \(y\in L^q\), then \(\tau(xy)=\tau(yx)\). For \(y\in L^1\) the functional \(a\mapsto\tau(ay)\) on \(N\) is normal. If \(x\in L^1\) has polar decomposition \(x=u|x|\), then \(\|x\|_1=\tau(u^*x)\). Proved in Trace densities and noncommutative integration: the Hölder inequality in §TI-10, the trace on \(L^1\), its bound and its normality in §TI-06. The last formula holds because \(u^*x=|x|\). See also [Fack–Kosaki 1986].

(B3) Polar decomposition. Every \(\tau\)-measurable \(x\) has a unique decomposition \(x=u|x|\) with \(|x|=(x^*x)^{1/2}\) and \(u\in N\) a partial isometry with \(u^*u\) equal to the support projection \(s(|x|)\) of \(|x|\). Uniqueness means: if \(x=vk\) with \(k\ge0\) \(\tau\)-measurable and \(v\in N\) a partial isometry with \(v^*v=s(k)\), then \(k=|x|\) and \(v=u\). Moreover \(|x^*|=u|x|u^*\). If \(v\in N\) is unitary and \(h\ge0\), then \(f(v^*hv)=v^*f(h)v\). The polar decomposition and its uniqueness are proved in Operators recovered from small trace defects, §MT-12; \(|x^*|=u|x|u^*\) follows from uniqueness applied to \(x^*=u^*(u|x|u^*)\), and the last formula holds because \(\chi_B(v^*hv)=v^*\chi_B(h)v\) for the spectral projections.

(B4) \(L^2\) as a bimodule. For \(a,b\in N\), the maps \(L(a)\colon\xi\mapsto a\xi\) and \(R(b)\colon\xi\mapsto\xi b\) are bounded on \(L^2(N,\tau)\), \(L(a)R(b)=R(b)L(a)\), \(a\mapsto L(a)\) is a normal representation and \(b\mapsto R(b)\) is a normal antirepresentation (increasing bounded nets go to strongly convergent nets). This is the standard form of \(N\). Proved in Integration for a trace, Sections 2–3 (Lemma 2.2 and Theorem 3.1).

(B5) Joint spectral measures. Let \(P\) and \(Q\) be projection-valued measures on the Borel sets of \([0,\infty)\), acting on the same Hilbert space, with \(P(A)Q(B)=Q(B)P(A)\) for all \(A,B\). There is a unique projection-valued measure \(\Pi\) on the Borel sets of \([0,\infty)^2\) with \(\Pi(A\times B)=P(A)Q(B)\), and for bounded functions \(f,g\) one has \(\int f(s)g(t)\,d\Pi(s,t)=\big(\int f\,dP\big)\big(\int g\,dQ\big)\). For a vector \(\xi\), \(\Delta\mapsto\langle\Pi(\Delta)\xi,\xi\rangle\) is a finite positive measure and \(\langle(\int\varphi\,d\Pi)\xi,\xi\rangle=\int\varphi\,d\langle\Pi\xi,\xi\rangle\) for bounded \(\varphi\). This is proved as in Joint spectral measures and commutator estimates in standard form, Theorem in §NC-20: the bounded operators \(A=\int\arctan\,dP\) and \(B=\int\arctan\,dQ\) commute, so \(T=A+iB\) is normal; its spectral measure (Spectral calculus with its domains retained, §SK-04), carried to the plane by \(z\mapsto(\tan\operatorname{Re}z,\tan\operatorname{Im}z)\), is \(\Pi\), and uniqueness holds because Borel rectangles generate the Borel sets.

(B6) Isomorphisms. Let \(\theta\colon N\to N_1\) be a \(*\)-isomorphism onto a von Neumann algebra with a faithful normal semifinite trace \(\tau_1\) such that \(\tau_1\circ\theta=\lambda\tau\) for some \(\lambda>0\). Then \(\theta\) extends uniquely to a \(*\)-isomorphism \(\tilde\theta\) from the \(\tau\)-measurable to the \(\tau_1\)-measurable operators, and \(\tilde\theta\) commutes with the Borel functional calculus of positive operators, with moduli and with polar decompositions. This is Measurable operators for a trace, Theorem 14.1. See also [Nelson 1974].

(B7) Comparison of projections. Equivalent projections have equal trace. If \(p\sim q\) and \(c\) is a central projection, then \(pc\sim qc\). Sums of orthogonal families of pairwise equivalent projections (with the first family orthogonal and the second family orthogonal) are equivalent. For any projections \(p,q\) there is a central projection \(c\) with \(pc\precsim qc\) and \(q(1-c)\precsim p(1-c)\) (comparison theorem). A projection \(p\) is finite if \(p\sim p'\le p\) implies \(p'=p\); subprojections of finite projections are finite; a projection of finite trace is finite. For any projections \(p,q\), \(p\vee q-p\sim q-p\wedge q\) (Kaplansky's formula). Equal traces: \(\tau(v^*v)=\tau(vv^*)\). Central cuts and orthogonal sums: Traces on von Neumann algebras, Lemmas 1.1 and 1.2. The comparison theorem: Projections and types of von Neumann algebras, Theorem 5.5. Finite projections: Section 6 there; a projection \(p\) of finite trace is finite, because \(p\sim p'\le p\) gives \(\tau(p-p')=0\), hence \(p'=p\) by faithfulness. Kaplansky's formula: Proposition 4.4 there.

(B8) Type structure of semifinite algebras. \(N\) is the direct sum of a type I and a type II algebra. A type I algebra is a direct sum \(\bigoplus_\kappa N_\kappa\) over cardinals \(\kappa\ge1\), where \(N_\kappa\) is zero or of type I\(_\kappa\): its unit is the sum of \(\kappa\) mutually orthogonal equivalent abelian projections with central support \(1\). A type I\(_n\) algebra with \(n\) finite is \(*\)-isomorphic to \(M_n(Z)\) with \(Z\) abelian. Every subprojection of an abelian projection \(q\) has the form \(qc\) with \(c\) central. In a type II\(_1\) algebra, for every \(n\ge1\) the unit is a sum of \(n\) mutually orthogonal equivalent projections. If \(z\ne0\) is a central projection of \(N\) of type II and \(p\le z\) is a nonzero projection with \(\tau(p)<\infty\), then \(pNp\) is of type II\(_1\). Since \(\tau\) is faithful and semifinite, every nonzero projection majorizes a nonzero projection of finite trace. The type decomposition and the structure of type I algebras are proved in Projections and types of von Neumann algebras, Theorems 7.2 and 10.3, and \(M_n(Z)\) for type I\(_n\) by its matrix units (Proposition 8.4). If \(q\) is abelian, \(qNq\) is abelian with centre \(Zq\) (Traces on von Neumann algebras, Lemma 1.5), so \(qNq=Zq\) and its projections have the form \(qc\). The last statement is the semifiniteness of \(\tau\) (Traces on von Neumann algebras, Proposition 2.5). The division of the unit of a type II₁ algebra into \(n\) equivalent projections is Traces on von Neumann algebras, Corollary 5.18, and the type of \(pNp\) is Corollary 5.19 there.

(B9) Matrices over commutative algebras. If \(Z\cong C(\Omega)\) is a commutative C*-algebra, then \(M_n(Z)\) is the C*-algebra \(C(\Omega,M_n(\mathbb C))\) with the norm \(\|x\|=\sup_{\omega}\|x(\omega)\|\). Indeed \(M_n(Z)\to C(\Omega,M_n(\mathbb C))\) is an injective \(*\)-homomorphism onto, and both sides are C*-algebras, so it is isometric (C*-algebras: continuous functional calculus, automatic continuity, positive cones, approximate identities and quotients, Corollary 4.6).

1. Setting and notation

Definition 1.1. A positive self-adjoint operator \(h\) affiliated with \(N\) is tail-finite if \(\nu_h([\varepsilon,\infty))=\tau(\chi_{[\varepsilon,\infty)}(h))<\infty\) for every \(\varepsilon>0\).

A tail-finite operator is \(\tau\)-measurable. If \(h\in L^p(N,\tau)^+\) with \(p<\infty\), then \(h\) is tail-finite, since by (B1) \[ \nu_h([\varepsilon,\infty))\le\varepsilon^{-p}\int t^p\,d\nu_h(t)=\varepsilon^{-p}\|h\|_p^p .\tag{1.1} \]

Notation 1.2. For a positive self-adjoint operator \(h\) affiliated with \(N\) and \(a>0\), \[ E_a(h)=\chi_{(a,\infty)}(h), \] the spectral projection of \(h\) for the half-line \((a,\infty)\). For a closed densely defined operator \(x\) affiliated with \(N\), with polar decomposition \(x=u|x|\) (B3), we write \(u(x)=u\), and for \(a>0\) \[ u_a(x)=u(x)\,E_a(|x|). \] We call \(u_a(x)\) the truncated polar part of \(x\) at level \(a\).

Since \(E_a(|x|)\le s(|x|)=u^*u\), \(u_a(x)\) is a partial isometry in \(N\) with initial projection \(E_a(|x|)\) and final projection \(uE_a(|x|)u^*=E_a(u|x|u^*)=E_a(|x^*|)\). If \(x\in L^2\), then \(\tau(E_a(|x|))\le a^{-2}\|x\|_2^2\) by (1.1), so \(u_a(x)\in N\cap L^2\) and \(\|u_a(x)\|_2^2=\tau(E_a(|x|))\). For positive \(h\), \(u(h)=s(h)\) and \(u_a(h)=E_a(h)\).

The map \(x\mapsto u(x)\) is badly discontinuous, even for \(N=\mathbb C^2\): see Example 7.3. The truncation \(E_a(|x|)\) removes the small part of \(|x|\), where the phase \(u(x)\) is unstable.

Lemma 1.3. Let \(h\) be tail-finite and \(w\in N\).

(a) For \(a>0\), \(E_a(h)\in N\cap L^1\cap L^2\).

(b) The map \((0,\infty)\ni a\mapsto wE_a(h)\in L^2(N,\tau)\) is right-continuous, and \(a\mapsto\tau(E_a(h))=\nu_h((a,\infty))\) is finite, nonincreasing and right-continuous.

(c) A right-continuous function \(F\colon(0,\infty)\to\mathbb R\) is Borel.

Proof. (a) \(\tau(E_a(h))\le\nu_h([a,\infty))<\infty\), and a projection of finite trace lies in \(L^1\) and \(L^2\). (b) For \(0<a<a'\), \(E_a(h)-E_{a'}(h)=\chi_{(a,a']}(h)\), so by (B2) \(\|w(E_a(h)-E_{a'}(h))\|_2^2\le\|w\|^2\nu_h((a,a'])\), which tends to \(0\) as \(a'\downarrow a\) because \(\nu_h\) is finite on \([a,\infty)\). The second statement is the same fact for the finite measure \(\nu_h\) restricted to \([a,\infty)\). (c) \(F\) is the pointwise limit of the Borel step functions \(F_m(a)=F(2^{-m}\lceil 2^ma\rceil)\) (for \(a>0\), \(2^{-m}\lceil2^ma\rceil\downarrow a\)). \(\square\)

2. The joint distribution of two positive operators

Two positive operators in \(N\) that do not commute have no joint spectral measure in \(N\). The following theorem shows that for \(L^2\) quantities they behave as if they had one.

Put \(X=[0,\infty)^2\setminus\{(0,0)\}\), a locally compact space, and let \(H(s,t)=s\) and \(K(s,t)=t\) be the coordinate functions on \(X\).

Theorem 2.1. Let \(h,k\) be tail-finite positive operators affiliated with \(N\) (for instance \(h,k\in L^2(N,\tau)^+\)). There is exactly one positive Radon measure \(\mu\) on \(X\) with \[ \mu(A\times B)=\tau(\chi_A(h)\chi_B(k))\tag{2.1} \] for all Borel sets \(A,B\subset[0,\infty)\) with \(A\subset[\varepsilon,\infty)\) or \(B\subset[\varepsilon,\infty)\) for some \(\varepsilon>0\). It has the following properties.

(a) For every function \(f\ge0\) on \([0,\infty)\) with \(f(0)=0\), \[ \int_Xf(H)\,d\mu=\tau(f(h)),\qquad\int_Xf(K)\,d\mu=\tau(f(k))\qquad(\text{in }[0,\infty]). \] In particular \(f(H)\) is \(\mu\)-integrable if and only if \(f(h)\) is integrable.

(b) For all complex functions \(f,g\) on \([0,\infty)\) with \(f(0)=g(0)=0\) and \(f(h),g(k)\in L^2(N,\tau)\), \[ \langle f(h),g(k)\rangle=\int_Xf(H)\,\overline{g(K)}\,d\mu,\qquad \|f(h)-g(k)\|_2=\|f(H)-g(K)\|_{L^2(X,\mu)} . \]

Reference: [Connes 1976, §2.5].

The right side of (2.1) is a trace of a product of two operators of which at least one has finite trace, so it is defined; it equals \(\tau(\chi_A(h)\chi_B(k)\chi_A(h))\ge0\) when \(\chi_A(h)\) has finite trace.

Proof. The joint spectral measure. By (B4), \(P(A)=L(\chi_A(h))\) and \(Q(B)=R(\chi_B(k))\) are commuting projection-valued measures on \([0,\infty)\) acting on \(L^2(N,\tau)\); countable additivity comes from normality. Let \(\Pi\) be their joint projection-valued measure on \([0,\infty)^2\) (B5). For a bounded function \(f\) on \([0,\infty)\), \(\int f\,dP=L(f(h))\): this holds for simple \(f\), and both sides are norm-continuous in \(f\) for uniform convergence. Similarly \(\int g\,dQ=R(g(k))\). Hence \[ \Big(\int f(s)g(t)\,d\Pi(s,t)\Big)\xi=f(h)\,\xi\,g(k)\qquad(\xi\in L^2).\tag{2.2} \]

Local pieces. For \(\varepsilon>0\) let \(p_\varepsilon=\chi_{[\varepsilon,\infty)}(h)\), \(q_\varepsilon=\chi_{[\varepsilon,\infty)}(k)\) and \(r_\varepsilon=p_\varepsilon\vee q_\varepsilon\). Since \(\tau(p\vee q)\le\tau(p)+\tau(q)\) (because \(p\vee q-p\sim q-p\wedge q\)), \(r_\varepsilon\) has finite trace, so \(r_\varepsilon\in L^2\). Let \(m_\varepsilon(\Delta)=\langle\Pi(\Delta)r_\varepsilon,r_\varepsilon\rangle\), a finite positive measure on \([0,\infty)^2\). If \(A\subset[\varepsilon,\infty)\), then \(\chi_A(h)\le p_\varepsilon\le r_\varepsilon\), so \(r_\varepsilon\chi_A(h)r_\varepsilon=\chi_A(h)\) and, by (2.2) and (B2), \[ m_\varepsilon(A\times B)=\tau\big(r_\varepsilon\,\chi_A(h)\,r_\varepsilon\,\chi_B(k)\big)=\tau(\chi_A(h)\chi_B(k)). \] If instead \(B\subset[\varepsilon,\infty)\), the same holds after moving \(r_\varepsilon\) around by cyclicity: \(\tau(r_\varepsilon\chi_A(h)r_\varepsilon\chi_B(k))=\tau(\chi_A(h)\,r_\varepsilon\chi_B(k)r_\varepsilon) =\tau(\chi_A(h)\chi_B(k))\).

Let \(U_\varepsilon=[0,\infty)^2\setminus[0,\varepsilon)^2\). The rectangles \(A\times B\) with \(A\subset[\varepsilon,\infty)\) or \(B\subset[\varepsilon,\infty)\), together with \(U_\varepsilon\), form a family closed under finite intersections which generates the Borel sets of \(U_\varepsilon\). For \(\varepsilon'<\varepsilon\), the measures \(m_\varepsilon\) and \(m_{\varepsilon'}\) agree on these rectangles (both give (2.1)), hence on \(U_\varepsilon=S\cup T\) with \(S=[\varepsilon,\infty)\times[0,\infty)\), \(T=[0,\infty)\times[\varepsilon,\infty)\) (inclusion–exclusion, \(S\cap T\) being such a rectangle). By the uniqueness theorem for finite measures agreeing on a generating \(\pi\)-system, \(m_\varepsilon\) and \(m_{\varepsilon'}\) agree on all Borel subsets of \(U_\varepsilon\).

The measure \(\mu\). For Borel \(\Delta\subset X\) put \(\mu(\Delta)=\lim_{j\to\infty}m_{1/j}(\Delta\cap U_{1/j})\). The sequence is nondecreasing, since \(m_{1/j}(\Delta\cap U_{1/j})=m_{1/(j+1)}(\Delta\cap U_{1/j})\le m_{1/(j+1)}(\Delta\cap U_{1/(j+1)})\). A nondecreasing limit of measures is a measure, so \(\mu\) is a positive measure on \(X\), and \(\mu(\Delta)=m_\varepsilon(\Delta)\) for Borel \(\Delta\subset U_\varepsilon\). Every compact subset of \(X\) lies in some \(U_\varepsilon\), which has finite measure; so \(\mu\) is a locally finite Borel measure on a locally compact second countable space, that is, a Radon measure. It satisfies (2.1). Uniqueness: two measures satisfying (2.1) agree on the Borel subsets of each \(U_\varepsilon\) by the \(\pi\)-system argument, and \(X=\bigcup_jU_{1/j}\).

(a) Let \(A\subset(0,\infty)\) be Borel and \(A_j=A\cap[1/j,\infty)\). By (2.1), \(\mu(A_j\times[0,\infty))= \tau(\chi_{A_j}(h))=\nu_h(A_j)\). Letting \(j\to\infty\), \(\mu(\{H\in A\})=\nu_h(A)\). Thus the image of \(\mu|_{\{H>0\}}\) under \(H\) is \(\nu_h\), and for \(f\ge0\) with \(f(0)=0\), \(\int_Xf(H)\,d\mu=\int_{\{H>0\}}f(H)\,d\mu=\int_{(0,\infty)}f\,d\nu_h=\tau(f(h))\). The same for \(K\).

(b) First let \(f,g\) be bounded and vanish on \([0,\varepsilon)\). Then \(f(h)=f(h)p_\varepsilon\) and \(r_\varepsilon f(h)r_\varepsilon=f(h)\), so by (2.2), with \(\varphi(s,t)=f(s)\overline{g(t)}\), \[ \int\varphi\,dm_\varepsilon=\langle f(h)\,r_\varepsilon\,\overline g(k),r_\varepsilon\rangle =\tau\big(r_\varepsilon f(h)r_\varepsilon\,\overline g(k)\big)=\tau\big(f(h)\,\overline g(k)\big) =\tau\big(\overline g(k)f(h)\big)=\langle f(h),g(k)\rangle . \] Since \(\varphi\) vanishes off \(U_\varepsilon\), \(\int\varphi\,dm_\varepsilon=\int_X\varphi\,d\mu\). This proves the first identity for such \(f,g\), and by (a) and polarization also \(\langle f(h),f'(h)\rangle=\int f(H)\overline{f'(H)}\,d\mu\) and the same for \(k\).

In general, let \(\mathcal F_h\) be the space of functions \(f\) with \(f(0)=0\) and \(\int|f|^2d\nu_h<\infty\). By (B1) and (a), both \(f\mapsto f(h)\in L^2(N,\tau)\) and \(f\mapsto f(H)\in L^2(X,\mu)\) are isometric on \(\mathcal F_h\) for the norm of \(L^2(\nu_h)\). The truncations \(f_j=f\,\chi_{[1/j,\infty)}\,\chi_{\{|f|\le j\}}\) are bounded, vanish near \(0\), and converge to \(f\) in \(L^2(\nu_h)\) by dominated convergence. Approximating \(f\) and \(g\) this way and passing to the limit in the identity already proved gives \(\langle f(h),g(k)\rangle=\int_Xf(H)\overline{g(K)}\,d\mu\). Expanding \(\|f(h)-g(k)\|_2^2=\|f(h)\|_2^2+\|g(k)\|_2^2-2\operatorname{Re}\langle f(h),g(k)\rangle\) and the same expression in \(L^2(X,\mu)\) gives the norm identity. \(\square\)

Remark. If \(\tau(1)<\infty\), then \(1\in L^2\) and one can take \(r_\varepsilon=1\) throughout: \(\mu\) is the restriction to \(X\) of the finite measure \(\Delta\mapsto\langle\Pi(\Delta)1,1\rangle\), the spectral measure of the trace vector for the commuting pair "left multiplication by \(h\), right multiplication by \(k\)". The origin is removed in the semifinite case because \(\tau(\chi_{\{0\}}(h)\chi_{\{0\}}(k))\) may be meaningless.

Example 2.2 (\(\mu\) does not see \(L^1\) norms). Let \(N=M_2(\mathbb C)\) with \(\tau=\operatorname{Tr}\), let \(0<\theta<\pi/2\), and let \(e,f\) be the projections onto \(\mathbb C(1,0)\) and \(\mathbb C(\cos\theta,\sin\theta)\). Then \(\tau(ef)=\cos^2\theta\) and \(\tau(e(1-f))=\tau((1-e)f)=\sin^2\theta\), so \(\mu\) has atoms of mass \(\cos^2\theta\) at \((1,1)\) and \(\sin^2\theta\) at \((1,0)\) and at \((0,1)\). As Theorem 2.1(b) predicts, \(\|e-f\|_2^2=2\sin^2\theta=\int|H-K|^2\,d\mu\). But \(e-f\) has eigenvalues \(\pm\sin\theta\), so \(\|e-f\|_1=2\sin\theta\), while \(\int|H-K|\,d\mu=2\sin^2\theta\).

In fact no positive measure \(\lambda\) on \(X\) satisfies \(\int|f(H)-g(K)|\,d\lambda=\|f(e)-g(f)\|_1\) for all \(f,g\) with \(f(0)=g(0)=0\). Taking \(g=0\) and then \(f=0\) shows that \(\lambda\) lives on the three points \((1,1),(1,0),(0,1)\), with masses \(\alpha,1-\alpha,1-\alpha\). The functions \(f(t)=g(t)=t\) give \(2(1-\alpha)=2\sin\theta\), and \(f(t)=2t\), \(g(t)=t\) give \(\alpha+2(1-\alpha)+(1-\alpha)=1+2\sin\theta\). But \(2e-f\) has trace \(1\) and determinant \(-2\sin^2\theta\), so \(\|2e-f\|_1=(1+8\sin^2\theta)^{1/2}<1+2\sin\theta\) for \(0<\theta<\pi/2\). The proof of Theorem 2.1 used the Hilbert space \(L^2\) in an essential way.

3. The Powers–Størmer inequality

Theorem 3.1 (Powers–Størmer inequality). For \(h,k\in L^2(N,\tau)^+\), \[ \|h-k\|_2^2\le\|h^2-k^2\|_1 .\tag{3.1} \] Equivalently, for \(h,k\in L^1(N,\tau)^+\), \(\|h^{1/2}-k^{1/2}\|_2^2\le\|h-k\|_1\).

Reference: [Powers–Størmer 1970], [Haagerup 1975].

Proof. Put \(d=h-k\) and \(s=h+k\), both in \(L^2\), and let \(P=\chi_{(0,\infty)}(d)\), \(Q=\chi_{(-\infty,0)}(d)\), \(U=P-Q\in N\). Then \(\|U\|\le1\), \(Ud=dU=|d|\), and \(d_+=Pd=PdP\), \(d_-=-Qd=-QdQ\) are positive, with \(d=d_+-d_-\), \(|d|=d_++d_-\), \(d_+d_-=0\). In the algebra of measurable operators \(h^2-k^2=\tfrac12(ds+sd)\), and every product below of two elements of \(L^2\) lies in \(L^1\) (B2).

By cyclicity (B2), \(\tau(dsU)=\tau(sUd)=\tau(s|d|)\) and \(\tau(sdU)=\tau(s|d|)\). Hence \[ \tau(s|d|)=\tau\big((h^2-k^2)U\big)\le\|U\|\,\|h^2-k^2\|_1\le\|h^2-k^2\|_1 . \] Now \(\tau(s|d|)=\tau(sd_+)+\tau(sd_-)\), and \(\tau(sd_+)=\tau(sPd_+P)=\tau(PsP\,d_+)\). The operator \(PsP-d_+=P(s-d)P=2PkP\) is positive. For positive \(y,z\in L^2\) one has \(\tau(yz)\ge0\): indeed \(y^{1/2}\in L^4\), \(y^{1/2}z\in L^{4/3}\), and \(\tau(yz)=\tau(y^{1/2}\,y^{1/2}z)=\tau(y^{1/2}zy^{1/2})\ge0\) by (B2). Therefore \(\tau(PsP\,d_+)\ge\tau(d_+^2)\). In the same way \(\tau(sd_-)=\tau(QsQ\,d_-)\) and \(QsQ-d_-=Q(s+d)Q=2QhQ\ge0\), so \(\tau(sd_-)\ge\tau(d_-^2)\). Adding, \[ \|h^2-k^2\|_1\ge\tau(d_+^2)+\tau(d_-^2)=\tau(d^2)=\|h-k\|_2^2 . \] The second form follows by applying the first to \(h^{1/2},k^{1/2}\in L^2\). \(\square\)

Corollary 3.2. For \(x,y\in L^2(N,\tau)\), \[ \big\||x|-|y|\big\|_2^2\le\big\||x|^2-|y|^2\big\|_1\le\|x-y\|_2\,(\|x\|_2+\|y\|_2).\tag{3.2} \] In particular \(x\mapsto|x|\) is continuous on \(L^2\): \(\||x|-|y|\|_2\le(\|x-y\|_2(\|x\|_2+\|y\|_2))^{1/2}\).

Proof. The first inequality is (3.1) for \(|x|,|y|\). For the second, \(x^*x-y^*y=x^*(x-y)+(x-y)^*y\), so by the Hölder inequality \(\|x^*x-y^*y\|_1\le\|x\|_2\|x-y\|_2+\|x-y\|_2\|y\|_2\). \(\square\)

Remarks. (1) The constant \(1\) in (3.1) is sharp: take \(k=0\) and \(h\) a projection of finite trace. For commuting \(h,k\), (3.1) is the pointwise inequality \(|\lambda-\kappa|^2\le|\lambda-\kappa|(\lambda+\kappa)\) (Exercise 3).

(2) In the finite case, \(h\in L^1(N,\tau)^+\) is the density of the normal positive functional \(\varphi_h=\tau(h\,\cdot\,)\), and \(h^{1/2}\) is the vector representing \(\varphi_h\) in the positive cone of the standard form \(L^2(N,\tau)\). Theorem 3.1 then says \(\|\xi_\varphi-\xi_\psi\|^2\le\|\varphi-\psi\|\) for the representing vectors; in this form the inequality holds for every von Neumann algebra in standard form [Haagerup 1975].

4. Integrating over the level

The projections \(E_a(h)\) jump when \(a\) crosses an eigenvalue of \(h\), so a small change of \(h\) can change \(E_a(h)\) a lot. The remedy is to average over \(a\). The basic tool is a noncommutative "layer-cake" formula.

Lemma 4.1 (layer-cake formula). Let \(h\in L^2(N,\tau)^+\) and \(c\in N^+\). For Borel \(B\subset(0,\infty)\) put \(\rho_c(B)=\tau(c^{1/2}\chi_B(h)c^{1/2})\in[0,\infty]\).

(a) \(\rho_c\) is a positive measure with \(\rho_c\le\|c\|\,\nu_h\), and \(\rho_c(B)=\tau(c\,\chi_B(h))= \|c^{1/2}\chi_B(h)\|_2^2\) whenever \(\nu_h(B)<\infty\).

(b) For every complex function \(f\) with \(f(0)=0\) and \(f(h)\in L^2\): \(\|c^{1/2}f(h)\|_2^2=\int|f|^2\,d\rho_c\).

(c) For \(0<b\le\infty\), \[ \int_0^b\tau\big(c\,E_{\sqrt a}(h)\big)\,da=\big\|c^{1/2}\min(h,\sqrt b)\big\|_2^2 ;\qquad\text{in particular}\qquad \int_0^\infty\tau\big(c\,E_{\sqrt a}(h)\big)\,da=\|c^{1/2}h\|_2^2=\tau(hch).\tag{4.1} \]

Proof. (a) If \(B=\bigsqcup_iB_i\), the partial sums of \(c^{1/2}\chi_{B_i}(h)c^{1/2}\) increase to \(c^{1/2}\chi_B(h)c^{1/2}\), so countable additivity follows from normality of \(\tau\). With \(y=\chi_B(h)c^{1/2}\), the trace property gives \(\rho_c(B)=\tau(y^*y)=\tau(yy^*)=\tau(\chi_B(h)c\chi_B(h))\le\|c\|\nu_h(B)\). If \(\nu_h(B)<\infty\), this equals \(\|c^{1/2}\chi_B(h)\|_2^2\), and \(\tau(\chi_B(h)c\chi_B(h))=\tau(c\chi_B(h))\) by cyclicity (B2).

(b) Let \(f=\sum_i\lambda_i\chi_{B_i}\) be simple with disjoint \(B_i\subset[\varepsilon,\infty)\), and \(P_i=\chi_{B_i}(h)\). For \(i\ne l\), \(\tau(P_icP_l)=\tau(cP_lP_i)=0\) by cyclicity, so \(\|c^{1/2}f(h)\|_2^2=\sum_{i,l}\overline{\lambda_i}\lambda_l\tau(P_icP_l)=\sum_i|\lambda_i|^2\rho_c(B_i) =\int|f|^2d\rho_c\). If \(f\) is bounded and vanishes on \([0,\varepsilon)\), take simple \(f_j\to f\) uniformly, vanishing on \([0,\varepsilon)\); then \(\|f_j(h)-f(h)\|_2\le\|f_j-f\|_\infty\nu_h([\varepsilon,\infty))^{1/2}\to0\), and \(\int|f_j|^2d\rho_c\to\int|f|^2d\rho_c\) because \(\rho_c([\varepsilon,\infty))<\infty\). In general, let \(f_j=f\chi_{[1/j,\infty)}\chi_{\{|f|\le j\}}\): then \(f_j(h)\to f(h)\) in \(L^2\) by (B1) and dominated convergence, and \(\int|f_j|^2d\rho_c\uparrow\int|f|^2d\rho_c\) by monotone convergence.

(c) By (a), \(\tau(cE_{\sqrt a}(h))=\rho_c((\sqrt a,\infty))\). By Tonelli's theorem and the identity \(\int_0^b\chi_{\{t^2>a\}}\,da=\min(t^2,b)\), \[ \int_0^b\rho_c((\sqrt a,\infty))\,da=\int\!\!\int_0^b\chi_{\{t^2>a\}}\,da\,d\rho_c(t)=\int\min(t^2,b)\,d\rho_c(t), \] which is \(\|c^{1/2}\min(h,\sqrt b)\|_2^2\) by (b). For \(b=\infty\) the integrand is \(t^2\) and (b) applies to \(f(t)=t\). Finally \(\|c^{1/2}h\|_2^2=\tau(hch)\). \(\square\)

With \(c=1\), (4.1) reads \(\int_0^\infty\tau(E_{\sqrt a}(h))\,da=\|h\|_2^2\): the function \(a\mapsto\tau(E_{\sqrt a}(h))\) is a probability density when \(\|h\|_2=1\). This is used in Section 7.

Proposition 4.2. Let \(x\in L^2(N,\tau)\), \(h=|x|\), and \(w\in N\). Then \[ \int_0^\infty\big\|u_{\sqrt a}(x)-w\,E_{\sqrt a}(h)\big\|_2^2\,da=\|x-wh\|_2^2 .\tag{4.2} \] In particular (\(w=1\)): \[ \int_0^\infty\big\|u_{\sqrt a}(x)-u_{\sqrt a}(|x|)\big\|_2^2\,da=\big\|x-|x|\big\|_2^2 . \]

Proof. Let \(u=u(x)\) and \(c=(u-w)^*(u-w)\in N^+\). Since \(u_{\sqrt a}(x)=uE_{\sqrt a}(h)\), \[ \|u_{\sqrt a}(x)-wE_{\sqrt a}(h)\|_2^2=\|(u-w)E_{\sqrt a}(h)\|_2^2=\tau(E_{\sqrt a}(h)\,c\,E_{\sqrt a}(h)) =\tau(c\,E_{\sqrt a}(h)). \] By (4.1) the integral is \(\tau(hch)=\|(u-w)h\|_2^2=\|x-wh\|_2^2\), since \(uh=x\). For \(w=1\) note \(u_{\sqrt a}(|x|)=E_{\sqrt a}(h)\). \(\square\)

Proposition 4.3. Let \(h,k\) be tail-finite and \(\mu\) their joint distribution (Theorem 2.1).

(a) \(\displaystyle\int_0^\infty\|E_a(h)-E_a(k)\|_2^2\,da=\int_X|H-K|\,d\mu\) (in \([0,\infty]\)).

(b) If \(h,k\in L^2(N,\tau)^+\), then \[ \int_0^\infty\big\|E_{\sqrt a}(h)-E_{\sqrt a}(k)\big\|_2^2\,da=\int_X|H^2-K^2|\,d\mu\le\|h-k\|_2\,\|h+k\|_2 .\tag{4.3} \]

Proof. (a) For \(a>0\), Theorem 2.1(b) with \(f=g=\chi_{(a,\infty)}\) gives \(\|E_a(h)-E_a(k)\|_2^2=\int_X|\chi_{\{H>a\}}-\chi_{\{K>a\}}|\,d\mu\) (the integrand takes the values \(0,1\), so it equals its square). The function \((a,\xi)\mapsto|\chi_{\{H(\xi)>a\}}-\chi_{\{K(\xi)>a\}}|\) is Borel on \((0,\infty)\times X\), and for \(s,t\ge0\), \(\int_0^\infty|\chi_{\{s>a\}}-\chi_{\{t>a\}}|\,da=|s-t|\). Tonelli's theorem gives (a).

(b) The same argument with \(f=g=\chi_{(\sqrt a,\infty)}\) and \(\{H>\sqrt a\}=\{H^2>a\}\) gives the equality. Then \(|H^2-K^2|=|H-K|(H+K)\), and the Cauchy–Schwarz inequality in \(L^2(X,\mu)\) together with Theorem 2.1(b) (for \(f(t)=t\) and \(g(t)=\pm t\)) gives \(\int|H^2-K^2|\,d\mu\le\|H-K\|_{L^2(\mu)}\|H+K\|_{L^2(\mu)}=\|h-k\|_2\|h+k\|_2\). \(\square\)

When \(N\) is commutative, the left side of (4.3) is the \(L^1\) norm of \(h^2-k^2\) (by the scalar identity used in the proof), and it is tempting to bound the left side of (4.3) by \(\|h^2-k^2\|_1\) in general, which would sharpen (4.3) through Theorem 3.1. This is false:

Example 4.4. Let \(N=M_2(\mathbb C)\), \(\tau=\operatorname{Tr}\), \(A=\operatorname{diag}(0,2)\), and let \(B\) have eigenvalue \(1\) on \(\mathbb C(\cos\theta,\sin\theta)\) and \(3\) on \(\mathbb C(-\sin\theta,\cos\theta)\), with \(\theta=\pi/6\). Put \(h=A^{1/2}\), \(k=B^{1/2}\), so \(E_{\sqrt a}(h)=E_a(A)\), \(E_{\sqrt a}(k)=E_a(B)\). For \(0<a<1\), \(E_a(A)-E_a(B)\) is minus the projection onto \(\mathbb C(1,0)\), with squared \(2\)-norm \(1\). For \(1\le a<2\) it is the difference of the projections onto \(\mathbb C(0,1)\) and \(\mathbb C(-\sin\theta,\cos\theta)\), with squared \(2\)-norm \(2\sin^2\theta=\tfrac12\). For \(2\le a<3\) it is minus a rank-one projection, squared norm \(1\). So the left side of (4.3) is \(\tfrac52\). On the other hand \[ A-B=\begin{pmatrix}-\tfrac32&\tfrac{\sqrt3}2\\ \tfrac{\sqrt3}2&-\tfrac12\end{pmatrix} \] has trace \(-2\) and determinant \(0\), so \(\|h^2-k^2\|_1=\|A-B\|_1=2<\tfrac52\).

5. Level integration in the trace norm

For \(h,k\in L^1(N,\tau)^+\) put \[ \Phi(h,k)=\int_0^\infty\|E_a(h)-E_a(k)\|_1\,da\in[0,\infty]. \] We always have \(\Phi(h,k)\ge\|h-k\|_1\) (Lemma 5.2(c)), with equality when \(N\) is abelian (Lemma 5.3 with \(d=1\)); the reason is the scalar identity \(\int_0^\infty|\chi_{\{\lambda>a\}}-\chi_{\{\kappa>a\}}|\,da=|\lambda-\kappa|\). In general \(\Phi\) is not controlled by \(\|h-k\|_1\), unless \(N\) is small. Say that \(N\) is of bounded type I if it is of type I and there is an integer \(d\) such that \(N\) has no nonzero summand of type I\(_\kappa\) with \(\kappa>d\) (finite or infinite).

Theorem 5.1. The following are equivalent.

(i) There is a constant \(C\) with \(\Phi(h,k)\le C\|h-k\|_1\) for all \(h,k\in L^1(N,\tau)^+\).

(ii) \(N\) is of bounded type I.

If \(N\) is of type I with summands of type I\(_n\), \(n\le d\) only, then (i) holds with \(C=d^2\). In every case \(\Phi(h,k)\ge\|h-k\|_1\), so \(C\ge1\), and \(C=1\) works when \(N\) is abelian.

Reference: [Connes 1976, §2.5, Remark 1] (stated there without proof).

We first make sure that \(\Phi\) is well defined, and record a semicontinuity property.

Lemma 5.2. (a) For \(y\in N\cap L^1\), \(\|y\|_1=\sup\{|\tau(yz)|:z\in N\cap L^1,\ \|z\|\le1\}\). If \(y_m\in N\cap L^1\) is a bounded sequence converging strongly to \(y\in N\cap L^1\), then \(\|y\|_1\le\liminf_m\|y_m\|_1\).

(b) For \(h,k\in L^1(N,\tau)^+\), the function \(a\mapsto\|E_a(h)-E_a(k)\|_1\) is Borel on \((0,\infty)\).

(c) For \(h\in L^1(N,\tau)^+\) and \(z\in N\), \(\int_0^\infty\tau(zE_a(h))\,da=\tau(zh)\), the integral converging absolutely. Consequently \(\|h-k\|_1\le\Phi(h,k)\) for \(h,k\in L^1(N,\tau)^+\).

Proof. (a) "\(\ge\)" is \(|\tau(yz)|\le\|z\|\|y\|_1\). For "\(\le\)", let \(y=u|y|\) and \(p_j=\chi_{(1/j,\infty)}(|y|)\), of finite trace by (1.1); with \(z=p_ju^*\), \(\tau(yz)=\tau(zy)=\tau(p_j|y|)\uparrow \tau(|y|)\). For each such \(z\), \(a\mapsto\tau(az)\) is normal (B2), so \(\tau(y_mz)\to\tau(yz)\); taking the supremum over \(z\) gives the semicontinuity.

(b) Let \(D_a=E_a(h)-E_a(k)\in N\cap L^1\) and \(F(a)=\|D_a\|_1=\sup_z|\tau(D_az)|\). For fixed \(z\), \(a\mapsto \tau(D_az)\) is right-continuous: \(D_{a'}\to D_a\) strongly as \(a'\downarrow a\), and \(\tau(\,\cdot\,z)\) is normal. Hence \(F\) is lower semicontinuous from the right: if \(F(a)>\gamma\), some \(z\) has \(|\tau(D_az)|>\gamma\), and then \(F>\gamma\) on some \([a,a+\eta)\). So \(V=\{F>\gamma\}\) is a union of half-open intervals \([a,a+\eta_a)\). The points of \(V\) not interior to \(V\) are left endpoints \(a\) of such intervals, and for two of them, \(a<a'\), we have \(a'\ge a+\eta_a\) (otherwise \(a'\) would be interior). So the open intervals \((a,a+\eta_a)\) attached to them are disjoint, and there are countably many such points. Thus \(V\) is an open set plus a countable set, a Borel set.

(c) Put \(g=h^{1/2}\in L^2\); then \(E_{\sqrt a}(g)=E_a(h)\). For \(c\in N^+\), (4.1) of Lemma 4.1 gives \(\int_0^\infty\tau(cE_a(h))\,da=\tau(gcg)=\tau(ch)\), by cyclicity. Every \(z\in N\) is a linear combination of four elements of \(N^+\), and \(|\tau(zE_a(h))|\le\|z\|\tau(E_a(h))\) with \(\int_0^\infty\tau(E_a(h))\,da=\tau(h)<\infty\); this gives the formula. Now let \(h-k=w|h-k|\) and \(z=w^*\). Then \(\|h-k\|_1=\tau(z(h-k))=\int_0^\infty\tau\big(z(E_a(h)-E_a(k))\big)\,da\le\int_0^\infty\|E_a(h)-E_a(k)\|_1\,da\). \(\square\)

Lemma 5.3 (the bounded case). If \(N\) is of type I with summands of type I\(_n\), \(n\le d\) only, then \(\Phi(h,k)\le d^2\|h-k\|_1\) for all \(h,k\in L^1(N,\tau)^+\).

Proof. Step 1: finite spectrum. Suppose \(h=\sum_i\lambda_iP_i\) and \(k=\sum_j\kappa_jQ_j\) are finite sums with distinct \(\lambda_i\ge0\), distinct \(\kappa_j\ge0\), and \(\sum_iP_i=\sum_jQ_j=1\) (spectral projections; the one for the value \(0\) may have infinite trace). Put \(Y=h-k\in L^1\). Then \(P_iYQ_j=(\lambda_i-\kappa_j)P_iQ_j\), so for \(a>0\) \[ E_a(h)-E_a(k)=\sum_{i,j}\big(\chi_{\{\lambda_i>a\}}-\chi_{\{\kappa_j>a\}}\big)P_iQ_j =\sum_{(i,j):\,\lambda_i\ne\kappa_j}c_{ij}(a)\,P_iYQ_j ,\qquad c_{ij}(a)=\frac{\chi_{\{\lambda_i>a\}}-\chi_{\{\kappa_j>a\}}}{\lambda_i-\kappa_j}, \] the terms with \(\lambda_i=\kappa_j\) being zero. Since \(\int_0^\infty|c_{ij}(a)|\,da=1\), \(\Phi(h,k)\le\sum_{i,j}\|P_iYQ_j\|_1\).

Let \(P_iYQ_j=w_{ij}|P_iYQ_j|\) be the polar decomposition; then \(w_{ij}^*=Q_jw_{ij}^*P_i\) and, by (B2), \(\|P_iYQ_j\|_1=\tau(w_{ij}^*P_iYQ_j)=\tau(Q_jw_{ij}^*P_iY)\). Hence \(\sum_{i,j}\|P_iYQ_j\|_1=\tau(TY)\le\|T\|\,\|Y\|_1\) with \(T=\sum_{i,j}Q_jw_{ij}^*P_i\). By (B8) and (B9), \(N=\bigoplus_{n\le d}C(\Omega_n,M_n(\mathbb C))\) as a C*-algebra, and \(\|T\|=\max_n\sup_{\omega}\|T(\omega)\|\). For fixed \(n\) and \(\omega\), the \(P_i(\omega)\) are mutually orthogonal projections in \(M_n(\mathbb C)\), so at most \(n\) of them are nonzero, and likewise for the \(Q_j(\omega)\); the other terms of \(T(\omega)\) vanish, and each remaining term has norm at most \(1\). So \(\|T(\omega)\|\le n^2\le d^2\), and \(\Phi(h,k)\le d^2\|h-k\|_1\).

Step 2: general \(h,k\). Let \(g_m(t)=\min(m,2^{-m}\lfloor2^mt\rfloor)\), a nondecreasing sequence of step functions with finitely many values, \(g_m(t)\le t\), \(g_m(t)\uparrow t\). Put \(h_m=g_m(h)\), \(k_m=g_m(k)\). Then \(\|h-h_m\|_1=\int(t-g_m(t))\,d\nu_h(t)\to0\) by dominated convergence, and likewise for \(k\). For \(a>0\) the sets \(\{g_m>a\}\) increase to \(\{t>a\}\), so \(E_a(h_m)\uparrow E_a(h)\) strongly, and similarly for \(k\). By Lemma 5.2(a) and Fatou's lemma, \[ \Phi(h,k)\le\int_0^\infty\liminf_m\|E_a(h_m)-E_a(k_m)\|_1\,da\le\liminf_m\Phi(h_m,k_m) \le d^2\lim_m\|h_m-k_m\|_1=d^2\|h-k\|_1 .\qquad\square \]

Lemma 5.4 (matrix units). If \(N\) is not of bounded type I, then for every \(n\ge1\) there are matrix units \((e_{pq})_{p,q=1}^n\) in \(N\) with \(0<\tau(e_{11})<\infty\).

Proof. If the type II summand \(Nz\) of \(N\) is nonzero, choose a nonzero projection \(r\le z\) with \(\tau(r)<\infty\) (B8). Then \(rNr\) is of type II\(_1\), so \(r=e_1+\dots+e_n\) with orthogonal equivalent projections \(e_p\). Otherwise \(N\) is of type I and has a nonzero summand \(Nz\) of type I\(_\kappa\) with \(\kappa\ge n\). Take \(n\) mutually orthogonal equivalent abelian projections \(q_1,\dots,q_n\le z\) with central support \(z\), and a nonzero projection \(r\le q_1\) of finite trace. By (B8) \(r=q_1c\) with \(c\) central; put \(e_p=q_pc\), which are orthogonal, equivalent to \(r\) (B7), and nonzero. In both cases choose partial isometries \(v_p\) with \(v_p^*v_p=e_1\), \(v_pv_p^*=e_p\) (\(v_1=e_1\)) and put \(e_{pq}=v_pv_q^*\). Then \(\tau(e_{11})\) is finite and nonzero since \(\tau\) is faithful. \(\square\)

Lemma 5.5 (a Hilbert matrix). For \(r\ge1\) the Hilbert matrix \(\mathsf H_r=(1/(\alpha+\beta-1))_{\alpha,\beta=1}^r\) is positive semidefinite, and \(\operatorname{Tr}|\mathsf H_r|=\operatorname{Tr}\mathsf H_r=\sum_{\alpha=1}^r \frac1{2\alpha-1}\ge\tfrac12\log(r+1)\).

Proof. \(1/(\alpha+\beta-1)=\int_0^1t^{\alpha-1}t^{\beta-1}\,dt\) is the Gram matrix of the functions \(t^{\alpha-1}\) in \(L^2[0,1]\), hence positive semidefinite, so its trace norm is its trace. Finally \(\sum_{\alpha\le r}\frac1{2\alpha-1}\ge\frac12\sum_{\alpha\le r}\frac1\alpha\ge\frac12\log(r+1)\). \(\square\)

Proof of Theorem 5.1. (ii) ⇒ (i) with \(C=d^2\) is Lemma 5.3.

(i) ⇒ (ii). Suppose (i) holds with a constant \(C\) but \(N\) is not of bounded type I. Fix \(n\ge4\) and matrix units as in Lemma 5.4, with \(t=\tau(e_{11})\). The map \(\iota(y)=\sum_{p,q}y_{pq}e_{pq}\) is a \(*\)-homomorphism \(M_n(\mathbb C)\to N\) with \(\tau(\iota(y))=t\operatorname{Tr}(y)\) (since \(\tau(e_{pq})=\tau(e_{pp}e_{pq})= \tau(e_{pq}e_{pp})=0\) for \(p\ne q\)). Hence \(\|\iota(y)\|_1=\tau(\iota(|y|))=t\operatorname{Tr}|y|\), and for positive \(y\) and \(a>0\), \(E_a(\iota(y))=\iota(E_a(y))\) (write \(E_a(y)=\pi(y)\) with a polynomial \(\pi\), \(\pi(0)=0\), interpolating \(\chi_{(a,\infty)}\) on the spectrum of \(y\) and at \(0\)). So (i) gives, for positive \(n\times n\) matrices \(y,y'\), \[ \int_0^\infty\operatorname{Tr}|E_a(y)-E_a(y')|\,da\le C\operatorname{Tr}|y-y'| .\tag{5.1} \] Let \(D=\operatorname{diag}(1,2,\dots,n)\), let \(S\) be the self-adjoint matrix with \(S_{pq}=\mathrm i/(q-p)\) for \(p\ne q\) and \(S_{pp}=0\), and \(U_\sigma=e^{\mathrm i\sigma S}\) for \(\sigma\in\mathbb R\). Let \(P_m\) (\(1\le m\le n-1\)) be the projection onto the span of the basis vectors \(m+1,\dots,n\). For \(a\in[m,m+1)\) we have \(E_a(D)=P_m\), for \(0<a<1\) \(E_a(D)=1\), and for \(a\ge n\) \(E_a(D)=0\); the same holds for \(U_\sigma DU_\sigma^*\) with \(P_m\) replaced by \(U_\sigma P_mU_\sigma^*\). Applying (5.1) to \(y=D\), \(y'=U_\sigma DU_\sigma^*\), \[ \sum_{m=1}^{n-1}\operatorname{Tr}|P_m-U_\sigma P_mU_\sigma^*|\le C\operatorname{Tr}|D-U_\sigma DU_\sigma^*| . \] Divide by \(\sigma>0\) and let \(\sigma\to0\). Since \(\sigma^{-1}(Y-U_\sigma YU_\sigma^*)\to-\mathrm i[S,Y]\) in norm, and all norms on \(M_n(\mathbb C)\) are equivalent, \[ \sum_{m=1}^{n-1}\operatorname{Tr}|[S,P_m]|\le C\operatorname{Tr}|[S,D]| .\tag{5.2} \] Right side. \([S,D]_{pq}=(q-p)S_{pq}=\mathrm i\) for \(p\ne q\), so \([S,D]=\mathrm i(J-1)\) where \(J\) is the all-ones matrix. The eigenvalues of \(J-1\) are \(n-1\) (once) and \(-1\) (\(n-1\) times), so \(\operatorname{Tr}|[S,D]|=2(n-1)\).

Left side. Let \(Z_m=(1-P_m)SP_m\). Since \(S\) is self-adjoint, \([S,P_m]=Z_m-Z_m^*\), and because \(Z_m^2=0\), the operator \(Z_m-Z_m^*\) has \(|Z_m-Z_m^*|^2=Z_m^*Z_m+Z_mZ_m^*\) with orthogonal summands; so \(\operatorname{Tr}|[S,P_m]|=2\operatorname{Tr}|Z_m|\). The entries of \(Z_m\) are \(S_{pq}=\mathrm i/(q-p)\) for \(p\le m<q\). Let \(r=\min(m,n-m)\) and restrict to the rows \(p=m+1-\beta\) and columns \(q=m+\alpha\), \(1\le\alpha,\beta\le r\): there \(q-p=\alpha+\beta-1\), so this compression of \(Z_m\) is \(\mathrm i\,\mathsf H_r\) up to a permutation. Compressions do not increase the trace norm, so by Lemma 5.5, \(\operatorname{Tr}|[S,P_m]|\ge\log(1+\min(m,n-m))\).

For the integers \(m\) with \(n/4\le m\le3n/4\) (at least \(n/2-1\) of them, all between \(1\) and \(n-1\)) we have \(\min(m,n-m)\ge n/4\). So (5.2) gives \((n/2-1)\log(1+n/4)\le2C(n-1)\), that is, \(\log(1+n/4)\le4C(n-1)/(n-2)\le6C\) for all \(n\ge4\). This is absurd for large \(n\). \(\square\)

Remark. The growth \(\log n\) in dimension \(n\) is the same phenomenon as the failure of the absolute value map to be Lipschitz on the trace class, found by Davies. It is the reason why the estimates of Section 4, which use squared \(2\)-norms and \(\|h-k\|_2\), are the right ones.

6. Truncated polar decompositions

Recall \(u_a(x)=u(x)E_a(|x|)\) (Notation 1.2). For \(x\) affiliated with \(N\) and \(b>0\), put \(x_{[b]}=x\,E_b(|x|)\): this cuts off the part of \(x\) where \(|x|\le b\).

Lemma 6.1. Let \(x\) be \(\tau\)-measurable and \(a>0\).

(1) For every unitary \(v\in N\): \(u_a(vx)=v\,u_a(x)\) and \(u_a(xv)=u_a(x)\,v\).

(2) If \(\theta\colon N\to N_1\) is a \(*\)-isomorphism with \(\tau_1\circ\theta=\lambda\tau\) as in (B6), then \(u_a(\tilde\theta(x))=\theta(u_a(x))\).

(3) For \(b>0\) and \(c=\max(a,b)\): \(u_a(x_{[b]})=u_c(x)\) and \(u_a(x_{[b]})\,|x_{[b]}|=u_c(x)\,|x|\).

(4) If \(h\ge0\) and \(g\colon[0,\infty)\to[0,\infty)\) is an increasing bijection, then \(u_{g(a)}(g(h))=u_a(h)=E_a(h)\).

(5) For \(s>0\): \(u_a(sx)=u_{a/s}(x)\).

Proof. Write \(x=u|x|\).

(1) \(vx=(vu)|x|\), and \(vu\) is a partial isometry with \((vu)^*(vu)=u^*u=s(|x|)\); by uniqueness (B3), \(u(vx)=vu\) and \(|vx|=|x|\), so \(u_a(vx)=vuE_a(|x|)\). Next \(xv=(uv)(v^*|x|v)\), where \(v^*|x|v\ge0\), \((v^*|x|v)^2=v^*x^*xv\), and \((uv)^*(uv)=v^*s(|x|)v=s(v^*|x|v)\). So \(|xv|=v^*|x|v\), \(u(xv)=uv\), and \(E_a(|xv|)=v^*E_a(|x|)v\) (B3); hence \(u_a(xv)=uvv^*E_a(|x|)v=u_a(x)v\).

(2) By (B6), \(\tilde\theta(x)=\theta(u)\tilde\theta(|x|)\) is the polar decomposition of \(\tilde\theta(x)\), and \(E_a(\tilde\theta(|x|))=\theta(E_a(|x|))\).

(3) \(x_{[b]}^*x_{[b]}=E_b(|x|)|x|^2E_b(|x|)=(|x|E_b(|x|))^2\), so \(|x_{[b]}|=|x|E_b(|x|)=\gamma(|x|)\) with \(\gamma(t)=t\chi_{(b,\infty)}(t)\). Also \(x_{[b]}=(uE_b(|x|))\,|x_{[b]}|\), and \(uE_b(|x|)\) is a partial isometry with initial projection \(E_b(|x|)=s(|x_{[b]}|)\) (as \(b>0\)); so \(u(x_{[b]})=uE_b(|x|)\). Since \(\chi_{(a,\infty)}(\gamma(t))=\chi_{(c,\infty)}(t)\) for \(t\ge0\), \(E_a(|x_{[b]}|)=E_c(|x|)\le E_b(|x|)\). Hence \(u_a(x_{[b]})=uE_b(|x|)E_c(|x|)=u_c(x)\), and \(u_a(x_{[b]})|x_{[b]}|=uE_c(|x|)|x|E_b(|x|)=u_c(x)|x|\).

(4) \(g(h)\ge0\), so \(u_{g(a)}(g(h))=E_{g(a)}(g(h))=\chi_{(g(a),\infty)}\circ g\,(h)=\chi_{(a,\infty)}(h)\), since \(g\) is strictly increasing.

(5) \(|sx|=s|x|\), \(u(sx)=u\), and \(E_a(s|x|)=E_{a/s}(|x|)\). \(\square\)

7. Stability of truncated polar decompositions

The phase \(u(x)\) of an element of \(L^2\) is not continuous in \(x\), and neither is \(u_a(x)\) for a fixed level \(a\). The main theorem says that for finitely many nearby vectors one can always find some level at which all the truncated phases are close, and at which the truncation loses little of the first vector.

Theorem 7.1. Let \(n\ge1\) and \(\varepsilon\in(0,1]\) with \(30\,n\,\varepsilon^{1/4}<1\). Let \(x_1,\dots,x_n\in L^2(N,\tau)\) with \(x_1\ne0\) and \[ \|x_j-x_1\|_2\le\varepsilon\|x_1\|_2\qquad(j=1,\dots,n).\tag{7.1} \] Then there is \(a>0\) such that

(i) \(\|u_a(x_j)-u_a(x_1)\|_2\le\varepsilon^{1/8}\,\|u_a(x_1)\|_2\) for \(j=1,\dots,n\);

(ii) \(\|x_1-u_a(x_1)|x_1|\|_2\le(30n)^{1/2}\varepsilon^{1/8}\,\|x_1\|_2<\|x_1\|_2\); in particular \(u_a(x_1)\ne0\).

Reference: [Connes 1976, §2.5].

Proof. By Lemma 6.1(5), replacing every \(x_j\) by \(x_j/\|x_1\|_2\) replaces \(u_a\) by \(u_{a\|x_1\|_2}\), so we may assume \(\|x_1\|_2=1\). Write \(x_j=u_jh_j\) (polar decompositions), \(u=u_1\), \(h=h_1\). Then \(\|x_j\|_2\le1+\varepsilon\le2\).

Step 1: the moduli are close. By Corollary 3.2, \(\|h_j-h\|_2^2\le\|x_j-x_1\|_2(\|x_j\|_2+\|x_1\|_2)\le\varepsilon(2+\varepsilon)\le3\varepsilon\).

Step 2: comparing phases with the phase of \(x_1\). By Proposition 4.2 with \(x=x_j\) and \(w=u\), \[ \int_0^\infty\|u_{\sqrt a}(x_j)-u\,E_{\sqrt a}(h_j)\|_2^2\,da=\|x_j-uh_j\|_2^2 \le\big(\|x_j-x_1\|_2+\|u(h-h_j)\|_2\big)^2\le\big(\varepsilon+(3\varepsilon)^{1/2}\big)^2\le9\varepsilon , \] using \(x_1=uh\), \(\|u\|\le1\) and \(\varepsilon\le\varepsilon^{1/2}\).

Step 3: comparing spectral projections. By Proposition 4.3(b) and Step 1, \[ \int_0^\infty\|u(E_{\sqrt a}(h_j)-E_{\sqrt a}(h))\|_2^2\,da\le\|h_j-h\|_2\|h_j+h\|_2 \le(3\varepsilon)^{1/2}(2+\varepsilon)\le3\sqrt3\,\varepsilon^{1/2}. \] Step 4. Since \(u_{\sqrt a}(x_1)=uE_{\sqrt a}(h)\), we have \(u_{\sqrt a}(x_j)-u_{\sqrt a}(x_1)= [u_{\sqrt a}(x_j)-uE_{\sqrt a}(h_j)]+u[E_{\sqrt a}(h_j)-E_{\sqrt a}(h)]\), and \(\|y+z\|_2^2\le2\|y\|_2^2+2\|z\|_2^2\) gives \[ \int_0^\infty\|u_{\sqrt a}(x_j)-u_{\sqrt a}(x_1)\|_2^2\,da\le18\varepsilon+6\sqrt3\,\varepsilon^{1/2} \le30\,\varepsilon^{1/2}.\tag{7.2} \] Step 5: choosing the level. Define on \((0,\infty)\) \[ F(a)=\sum_{j=2}^n\|u_{\sqrt a}(x_j)-u_{\sqrt a}(x_1)\|_2^2,\qquad G(a)=\|u_{\sqrt a}(x_1)\|_2^2=\tau(E_{\sqrt a}(h)). \] By Lemma 1.3, \(F\) and \(G\) are finite, right-continuous and Borel; \(G\) is nonincreasing. By (7.2), \(\int_0^\infty F\le30(n-1)\varepsilon^{1/2}\), and by (4.1) with \(c=1\), \(\int_0^\infty G=\|h\|_2^2=1\). Let \(\mathcal B=\{a>0:F(a)>\varepsilon^{1/4}G(a)\}\), a Borel set. On \(\mathcal B\), \(G<\varepsilon^{-1/4}F\), so \[ \int_{\mathcal B}G\le\varepsilon^{-1/4}\int_0^\infty F\le30(n-1)\varepsilon^{1/4}<1=\int_0^\infty G . \] Hence the complement \(\mathcal B^c=(0,\infty)\setminus\mathcal B\) is nonempty. It is closed under limits from the right: if \(a_k\downarrow a>0\) with \(a_k\in\mathcal B^c\), right-continuity gives \(F(a)\le\varepsilon^{1/4}G(a)\). Let \(b_0=\inf\mathcal B^c\). If \(b_0>0\), then \(b_0\in\mathcal B^c\) and \((0,b_0)\subset\mathcal B\); put \(b=b_0\), so \(\int_0^bG\le30(n-1)\varepsilon^{1/4}\). If \(b_0=0\), choose \(t>0\) with \(\int_0^tG\le30n\varepsilon^{1/4}\) (possible since \(G\) is integrable) and \(b\in\mathcal B^c\cap(0,t)\). In both cases \[ b\in\mathcal B^c,\qquad\int_0^bG(a)\,da\le30\,n\,\varepsilon^{1/4}.\tag{7.3} \] Step 6: conclusion. Let \(a=\sqrt b\). Since \(b\notin\mathcal B\), each term of \(F(b)\) is at most \(\varepsilon^{1/4}G(b)\), which is (i). For (ii), \(x_1-u_a(x_1)|x_1|=u(h-E_a(h)h)=u\,\chi_{[0,a]}(h)h\), and \(u\) is isometric on the range of \(h\), so by (B1), Lemma 4.1(c) with \(c=1\), and (7.3), \[ \|x_1-u_a(x_1)|x_1|\|_2^2=\int t^2\chi_{[0,a]}(t)\,d\nu_h(t)\le\int\min(t^2,b)\,d\nu_h(t)=\int_0^bG(s)\,ds \le30n\varepsilon^{1/4}<1 . \] If \(u_a(x_1)\) were \(0\), the left side would be \(\|x_1\|_2^2=1\). \(\square\)

Corollary 7.2. Let \(\delta\in(0,1)\), \(n\ge1\) and \(\varepsilon=(\delta/6n)^8\). If \(x_1,\dots,x_n\in L^2(N,\tau)\), \(x_1\ne0\), and \(\|x_j-x_1\|_2\le\varepsilon\|x_1\|_2\) for all \(j\) (for instance, if the set \(\{x_1,\dots,x_n\}\) has diameter at most \(\varepsilon\|x_1\|_2\)), then there is \(a>0\) with \[ \|u_a(x_j)-u_a(x_1)\|_2\le\delta\|u_a(x_1)\|_2\quad(j=1,\dots,n),\qquad\|x_1-u_a(x_1)|x_1|\|_2\le\delta\|x_1\|_2 , \] and \(u_a(x_1)\ne0\).

Proof. Here \(\varepsilon\le1\), \(\varepsilon^{1/8}=\delta/6n\le\delta\), \(30n\varepsilon^{1/4}=30\delta^2/36n<1\) and \((30n)^{1/2}\varepsilon^{1/8}=\delta\,(30/36n)^{1/2}\le\delta\). Apply Theorem 7.1. \(\square\)

Example 7.3 (why the level must be chosen). Let \(N=\mathbb C^2\) with \(\tau(s,t)=s+t\). For small \(\eta>0\) let \(x=(1,\eta)\) and \(y=(1,-\eta)\). Then \(\|x-y\|_2=2\eta\), but \(u(x)=(1,1)\) and \(u(y)=(1,-1)\), so \(\|u(x)-u(y)\|_2=2\). At every level \(a\in[\eta,1)\), \(u_a(x)=u_a(y)=(1,0)\): the truncation removes the unstable phase. On the other hand no single level works for all nearby vectors: given \(a\in(0,1)\) and \(0<\eta<a\), the positive vectors \(x'=(1,a+\eta)\) and \(y'=(1,a-\eta)\) satisfy \(\|x'-y'\|_2=2\eta\) but \(u_a(x')=(1,1)\), \(u_a(y')=(1,0)\). This is why Theorem 7.1 chooses \(a\) depending on the finite family, by an averaging argument over the level.

8. A projection almost commuting with finitely many unitaries

We now apply Theorem 7.1 to the conjugates of a positive operator. Commutators with projections are measured in the \(2\)-norm, and the Powers–Størmer inequality transfers the information to square roots.

Lemma 8.1. Let \(v\in N\) be unitary and \(e\in N\) a projection of finite trace. Then \[ \|vev^*-e\|_1\le2\,\|[v,e]\|_2\,\|e\|_2 . \]

Proof. Put \(x=ev^*\) and \(y=v^*e\). Then \(x^*x=vev^*\), \(y^*y=e\), \(\|x\|_2=\|y\|_2=\|e\|_2\), and \(x-y=v^*(ve-ev)v^*\), so \(\|x-y\|_2=\|[v,e]\|_2\). Apply the second inequality of (3.2). \(\square\)

Theorem 8.2. Let \(n\ge1\) and \(\varepsilon\in(0,1]\) with \(30n\varepsilon^{1/4}<1\). Let \(v_1,\dots,v_{n-1}\in N\) be unitaries and \(e_1,\dots,e_m\in N\) projections of finite trace, not all zero, such that \[ \|[v_k,e_s]\|_2\le\tfrac12\varepsilon^2\,\|e_s\|_2\qquad\text{for all }k,s.\tag{8.1} \] Let \(x_0=(e_1+\dots+e_m)^{1/2}\). There is \(a>0\) such that the spectral projection \(e=E_a(x_0)\) satisfies: \(e\ne0\), \(e\le e_1\vee\dots\vee e_m\), \(\tau(e)<\infty\), \[ \|[e,v_k]\|_2\le\varepsilon^{1/8}\|e\|_2\quad(k=1,\dots,n-1),\qquad \sum_{s=1}^m\|e_s-ee_s\|_2^2\le30n\varepsilon^{1/4}\sum_{s=1}^m\tau(e_s).\tag{8.2} \]

Proof. \(x_0\ge0\), \(x_0\in L^2\), and \(\|x_0\|_2^2=\tau(x_0^2)=\sum_s\tau(e_s)>0\). Put \(x_k=v_kx_0v_k^*\) for \(1\le k\le n-1\). By Lemma 8.1 and (8.1), \[ \|v_kx_0^2v_k^*-x_0^2\|_1\le\sum_s\|v_ke_sv_k^*-e_s\|_1\le\sum_s\varepsilon^2\|e_s\|_2^2=\varepsilon^2\|x_0\|_2^2 , \] and Theorem 3.1 (for \(v_kx_0v_k^*=(v_kx_0^2v_k^*)^{1/2}\) and \(x_0\)) gives \(\|x_k-x_0\|_2\le\varepsilon\|x_0\|_2\). Theorem 7.1, applied to the \(n\) vectors \(x_0,x_1,\dots,x_{n-1}\) (with \(x_0\) in the role of \(x_1\)), gives \(a>0\). Since \(x_0\) and \(x_k\) are positive, \(u_a(x_0)=E_a(x_0)=e\) and \(u_a(x_k)=E_a(v_kx_0v_k^*)=v_kev_k^*\) (B3). So (i) of Theorem 7.1 reads \(\|v_kev_k^*-e\|_2=\|v_ke-ev_k\|_2\le\varepsilon^{1/8}\|e\|_2\). By (ii), \(e\ne0\) and, as \(e\) commutes with \(x_0\), \[ \|x_0-ex_0\|_2^2=\tau\big((1-e)x_0^2(1-e)\big)=\sum_s\tau\big((1-e)e_s(1-e)\big)=\sum_s\|e_s-ee_s\|_2^2 \le30n\varepsilon^{1/4}\|x_0\|_2^2 . \] Finally \(e\le s(x_0)\), and the kernel of \(\sum_se_s\) is \(\bigcap_s\ker e_s\) (if \(\sum_s\langle e_s\xi,\xi\rangle=0\) then each \(e_s\xi=0\)), so \(s(x_0)=e_1\vee\dots\vee e_m\); and \(\tau(e)\le a^{-2}\|x_0\|_2^2\) by (1.1). \(\square\)

Corollary 8.3. Let \(\delta\in(0,1)\), \(n\ge1\) and \(\varepsilon_0=(\delta/24n)^{16}\). Let \(v_1,\dots,v_{n-1}\in N\) be unitaries and \(e_1,e_2\in N\) projections with \(\tau(e_1)=\tau(e_2)<\infty\) (for instance, equivalent projections of finite trace), such that \(\|[v_k,e_s]\|_2\le\varepsilon_0\|e_s\|_2\) for all \(k\) and \(s=1,2\). Then there is a projection \(e\le e_1\vee e_2\) with \[ \|ee_s-e_s\|_2\le\delta\|e_s\|_2\quad(s=1,2),\qquad\|[e,v_k]\|_2\le\delta\|e\|_2\quad(k=1,\dots,n-1). \]

Proof. If \(e_1=e_2=0\) take \(e=0\). Otherwise both are nonzero, and we apply Theorem 8.2 with \(m=2\) and \(\varepsilon=(2\varepsilon_0)^{1/2}=2^{1/2}(\delta/24n)^8\), so that (8.1) holds. We have \(\varepsilon\le1\), \(30n\varepsilon^{1/4}=30\cdot2^{1/8}\delta^2/(576\,n)<0.06\,\delta^2<1\), and \(\varepsilon^{1/8}=2^{1/16}\delta/24n<\delta\). By (8.2) and \(\tau(e_1)=\tau(e_2)\), \(\|e_s-ee_s\|_2^2\le0.06\,\delta^2\cdot2\tau(e_s)\le\delta^2\|e_s\|_2^2\). \(\square\)

Remark. In a II₁ factor the typical use is this: if two projections \(e_1,e_2\) of the same trace both almost commute with finitely many unitaries in the \(2\)-norm, each relative to its own size, then there is a spectral projection of \((e_1+e_2)^{1/2}\) that almost commutes with the unitaries and almost contains \(e_1\) and \(e_2\). The hypothesis on \(e_2\) cannot be dropped. Let \(M_3\subseteq N\) be a unital copy of the \(3\times3\) matrices, let \(v=\operatorname{diag}(1,1,-1)\), let \(e_1\) be the projection onto the first coordinate and \(e_2\) the projection onto \((0,1,1)/\sqrt2\). Then \([v,e_1]=0\) and \(\tau(e_1)=\tau(e_2)=\tfrac13\). But \(e_1\perp e_2\), so \(p=e_1+e_2\) is a projection and \(p\) is the only nonzero spectral projection of \((e_1+e_2)^{1/2}=p\), while \([v,p]=[v,e_2]\) has \(\|[v,p]\|_2^2=\tfrac23=\|p\|_2^2\). The square root is essential: Lemma 8.1 controls \(v(e_1+e_2)v^*-(e_1+e_2)\) only in \(L^1\), and Theorem 3.1 turns this into \(L^2\) control of the square root, which is what Theorem 7.1 needs.

9. Conjugating two projections by a unitary close to the identity

Let \(e,f\) be projections in a von Neumann algebra \(M\) (no trace is needed in this section). We want a unitary \(W\) with \(WeW^*=f\) and \(W-1\) as small as \(e-f\), in the strongest sense: as an operator inequality \(|W-1|\le C|e-f|\). Such an inequality implies \(\|W-1\|\le C\|e-f\|\) and, when the two sides commute (as they will), also \(\|W-1\|_p\le C\|e-f\|_p\) for a trace and every \(p\).

The two self-adjoint operators \[ \mathsf a=e-f,\qquad\mathsf b=e+f-1 \] carry the geometry of the pair. Let \(z\) be the projection onto the kernel of \(\mathsf b\), and \(\sigma=\chi_{(0,\infty)}(\mathsf b)-\chi_{(-\infty,0)}(\mathsf b)\), the sign of \(\mathsf b\); so \(\sigma=\sigma^*\), \(\sigma^2=1-z\), \(\sigma\mathsf b=|\mathsf b|\).

Lemma 9.1. (a) \(\mathsf a^2+\mathsf b^2=1\); \(\|\mathsf b\|\le1\); \(\mathsf a^2\) commutes with \(e\) and \(f\); \(\mathsf be=fe\) and \(e\mathsf b=ef\).

(b) \(z\), \(\sigma\), \(|\mathsf b|\) and \(|\mathsf a|\) commute with \(\mathsf a^2\); \(z\) commutes with \(e\) and \(f\); \(|\mathsf b|\) commutes with \(e\) and \(f\).

(c) \(ez=e\wedge(1-f)\), \(fz=(1-e)\wedge f\), \(ez+fz=z\), and \(\mathsf a^2z=z\).

Proof. (a) \(\mathsf a^2=e+f-ef-fe\) and \(\mathsf b^2=1-e-f+ef+fe\). Then \(e\mathsf a^2=e-efe=\mathsf a^2e\), and similarly for \(f\). \(\mathsf be=e+fe-e=fe\), \(e\mathsf b=ef\). Since \(0\le e+f\le2\), \(-1\le\mathsf b\le1\).

(b) \(\mathsf b^2=1-\mathsf a^2\) commutes with \(e\), \(f\), \(\mathsf a^2\), \(\mathsf b\); \(z=\chi_{\{0\}}(\mathsf b^2)\) and \(|\mathsf b|=(\mathsf b^2)^{1/2}\) are functions of \(\mathsf b^2\), and \(\sigma\), \(|\mathsf a|=(1-\mathsf b^2)^{1/2}\) are functions of \(\mathsf b\).

(c) \(\mathsf bz=0\) means \((e+f)z=z\), so \(ez+fz=z\). Multiplying \(ez+efz=ez\) (which is \(e\) applied to \(ez+fz=z\)) shows \(efz=0\), i.e. \(fz\,ez=0\) and \(ez\le1-f\); so \(ez\le e\wedge(1-f)\). Conversely, if \(\xi\in\operatorname{ran}(e\wedge(1-f))\), then \(\mathsf b\xi=\xi+0-\xi=0\), so \(\xi=z\xi=ez\xi\). The same argument gives \(fz=(1-e)\wedge f\). Finally \(\mathsf a^2z=(1-\mathsf b^2)z=z\). \(\square\)

Lemma 9.2 (the part off the kernel). The operator \(W_1=\sigma(2e-1)\in M\) is a partial isometry with \(W_1^*W_1=W_1W_1^*=1-z\). It satisfies \(W_1eW_1^*=f(1-z)\), \(W_1+W_1^*=2|\mathsf b|\), and it commutes with \(\mathsf a^2\), \(|\mathsf b|\) and \(z\).

Proof. \(2e-1\) is a self-adjoint unitary commuting with \(z\), and \(\sigma^2=1-z\), so \(W_1^*W_1=(2e-1)(1-z)(2e-1)=1-z\) and \(W_1W_1^*=\sigma(2e-1)^2\sigma=1-z\). Commutation with \(\mathsf a^2\), \(|\mathsf b|\), \(z\) follows from Lemma 9.1(b).

We use repeatedly: if \(Y\in M\) and \(|\mathsf b|\,Y=0\), then \((1-z)Y=0\), because the range of \(Y\) lies in the kernel of \(|\mathsf b|\), which is the range of \(z\).

Sum. \(|\mathsf b|(\sigma e+e\sigma)=\mathsf be+e\mathsf b=fe+ef\) (using that \(|\mathsf b|\) commutes with \(e\)), and \(|\mathsf b|(\sigma+|\mathsf b|)=\mathsf b+\mathsf b^2=ef+fe\). So \(Y=\sigma e+e\sigma-\sigma-|\mathsf b|\) satisfies \(|\mathsf b|Y=0\), whence \((1-z)Y=0\). Also \(zY=0\), because \(z\sigma=\sigma z=0\), \(z|\mathsf b|=0\) and \(z\) commutes with \(e\). So \(Y=0\), that is, \(\sigma e+e\sigma=\sigma+|\mathsf b|\). Therefore \(W_1+W_1^*=\sigma(2e-1)+(2e-1)\sigma=2(\sigma e+e\sigma)-2\sigma= 2|\mathsf b|\).

Conjugation. \(W_1eW_1^*=\sigma(2e-1)e(2e-1)\sigma=\sigma e\sigma\). Now \(\mathsf be\mathsf b=f(e\mathsf b)=fef\) by Lemma 9.1(a), and \(f\mathsf b^2=f-fe-f+fef+fe=fef\). So with \(Y=\sigma e\sigma-f(1-z)\) we get \(|\mathsf b|\,Y|\mathsf b|=\mathsf be\mathsf b-f\mathsf b^2=0\). By the observation above (applied to \(Y|\mathsf b|\) and then to the adjoint), \((1-z)Y(1-z)=0\). Since \(Y=(1-z)Y(1-z)\), \(W_1eW_1^*=f(1-z)\). \(\square\)

Proposition 9.3. If \(e\sim f\) and \(e,f\) are finite (for instance, if \(M\) is finite and \(e\sim f\)), then \(e\wedge(1-f)\sim(1-e)\wedge f\).

Proof. By Lemma 9.1(c) we must show \(ez\sim fz\). By Lemma 9.2, \(w=W_1e\) is a partial isometry with \(w^*w=e(1-z)\) and \(ww^*=f(1-z)\). By the comparison theorem (B7) there is a central projection \(c\) with \(ezc\precsim fzc\) and \(fz(1-c)\precsim ez(1-c)\). Let \(ezc\sim r\le fzc\). Since \(e(1-z)c\sim f(1-z)c\) (via \(wc\)), \(ezc\perp e(1-z)c\) and \(r\perp f(1-z)c\), we get \(ec\sim r+f(1-z)c\le fc\) (B7). Also \(ec\sim fc\). So \(fc\) is equivalent to its subprojection \(r+f(1-z)c\); as \(fc\) is finite, \(r+f(1-z)c=fc\), i.e. \(r=fzc\). Thus \(ezc\sim fzc\). Exchanging the roles of \(e\) and \(f\) (and using that \(e\) is finite) gives \(fz(1-c)\sim ez(1-c)\). Adding, \(ez\sim fz\). \(\square\)

Theorem 9.4. Let \(e,f\) be projections in a von Neumann algebra \(M\) with \(e\wedge(1-f)\sim(1-e)\wedge f\) (by Proposition 9.3 this holds when \(e\) and \(f\) are equivalent finite projections). Choose a partial isometry \(v\in M\) with \(v^*v=e\wedge(1-f)\), \(vv^*=(1-e)\wedge f\), and put \[ W=\operatorname{sgn}(e+f-1)\,(2e-1)+v-v^* . \] Then \(W\) is a unitary in \(M\), and:

(a) \(WeW^*=f\);

(b) \(W\) commutes with \(|e-f|\) and with \(|e+f-1|\);

(c) \((W-1)^*(W-1)=2\big(1-|e+f-1|\big)\le2\,(e-f)^2\); consequently \(|W-1|\le\sqrt2\,|e-f|\le3|e-f|\).

Reference: the estimate with the constant \(3\) is a lemma of Connes in his classification of injective factors.

Proof. With the notation above, \(W=W_1+W_z\) where \(W_1=\sigma(2e-1)\) and \(W_z=v-v^*\). By Lemma 9.1(c), \(v^*v=ez\) and \(vv^*=fz\), which are orthogonal with sum \(z\). Hence \(v=fz\,v\,ez\), \(v^2=0\), \(v^*ez=0\) and \(ezv=0\), and \[ W_z^*W_z=W_zW_z^*=v^*v+vv^*=z,\qquad W_z\,ez\,W_z^*=v\,ez\,v^*=fz,\qquad W_z^*=-W_z . \] Since \(W_1=(1-z)W_1(1-z)\) and \(W_z=zW_zz\), Lemma 9.2 shows that \(W\) is unitary and \(WeW^*=W_1e(1-z)W_1^*+W_z\,ez\,W_z^*=f(1-z)+fz=f\), which is (a).

(b) \(W_1\) commutes with \(\mathsf a^2\) and \(|\mathsf b|\) (Lemma 9.2). On the range of \(z\), \(\mathsf a^2=1\) and \(|\mathsf b|=0\) (Lemma 9.1(c)); as \(W_z=zW_zz\), it commutes with \(\mathsf a^2\) and \(|\mathsf b|\). Hence \(W\) commutes with \(\mathsf a^2\), so with \(|\mathsf a|=|e-f|\), and with \(|\mathsf b|=|e+f-1|\).

(c) \((W-1)^*(W-1)=2-(W+W^*)=2-2|\mathsf b|-(W_z+W_z^*)=2(1-|\mathsf b|)\). Since \(0\le|\mathsf b|\le1\), we have \(1-|\mathsf b|\le1-\mathsf b^2=\mathsf a^2\). Both \(|W-1|=(2(1-|\mathsf b|))^{1/2}\) and \(\sqrt2|\mathsf a|=(2(1-\mathsf b^2))^{1/2}\) are functions of \(\mathsf b\), and \((2(1-|t|))^{1/2}\le(2(1-t^2))^{1/2}\) for \(t\in[-1,1]\); so \(|W-1|\le\sqrt2|\mathsf a|\) by the functional calculus. \(\square\)

Corollary 9.5. Let \(e,f\) be equivalent finite projections in \(M\) (for instance, equivalent projections in a finite von Neumann algebra). There is a unitary \(W\in M\) with \(WeW^*=f\), \([W,|e-f|]=0\) and \(|W-1|\le\sqrt2\,|e-f|\). Moreover \(\varphi(|W-1|)\le\varphi(\sqrt2|e-f|)\) for every nondecreasing function \(\varphi\colon[0,\infty)\to[0,\infty)\). In particular \(\|W-1\|\le\sqrt2\|e-f\|\), and if \(\tau\) is a faithful normal semifinite trace on \(M\) and \(1\le p<\infty\), then \(\|W-1\|_p\le\sqrt2\,\|e-f\|_p\) (both sides possibly infinite).

Proof. Theorem 9.4 and Proposition 9.3. Both \(|W-1|\) and \(\sqrt2|e-f|\) are functions of \(\mathsf b\), and the pointwise inequality of the proof of (c) persists after composing with \(\varphi\). With \(\varphi(t)=t^p\), \(\tau(|W-1|^p)\le2^{p/2}\tau(|e-f|^p)\). \(\square\)

Example 9.6. (1) Two lines in the plane. In \(M_2(\mathbb C)\) let \(e\) and \(f\) be the projections onto \(\mathbb C(1,0)\) and \(\mathbb C(\cos\theta,\sin\theta)\), \(0<\theta<\pi/2\). Then \(e+f-1=\cos\theta\begin{pmatrix}\cos\theta&\sin\theta\\\sin\theta&-\cos\theta\end{pmatrix}\), whose sign is the reflection \(\begin{pmatrix}\cos\theta&\sin\theta\\\sin\theta&-\cos\theta\end{pmatrix}\), and \(W\) is the rotation by \(\theta\). Here \(|e-f|=\sin\theta\), \(|e+f-1|=\cos\theta\), and \(|W-1|=(2-2\cos\theta)^{1/2}=2\sin(\theta/2)\), in accordance with Theorem 9.4(c).

(2) The constant \(\sqrt2\) is optimal. If \(e\) and \(f\) are orthogonal, equivalent and nonzero, then \(|e-f|\) is the projection \(p=e+f\). Let \(W\) be any unitary with \(WeW^*=f\), and suppose \(T=|W-1|\le Cp\) for some \(C\ge0\). Then \((1-p)T(1-p)\le0\), so \(T^{1/2}(1-p)=0\) and \(T=pTp\); thus \(T\) commutes with \(p\), and \(0\le T\le C\) on the range of \(p\), whence \(T^2\le C^2p\). For a unit vector \(\xi\in\operatorname{ran}e\) we have \(W\xi=We\xi=fW\xi\in\operatorname{ran}f\perp\xi\), so \(2=\|W\xi-\xi\|^2=\|T\xi\|^2=\langle T^2\xi,\xi\rangle\le C^2\). Hence no inequality \(|W-1|\le C|e-f|\) with \(C<\sqrt2\) can hold, for any choice of \(W\).

(3) Finiteness cannot be dropped. Let \(M=B(\ell^2)\), \(S\) the unilateral shift, \(e=1\) and \(f=SS^*\). Then \(e\sim f\), but \(WeW^*=1\ne f\) for every unitary \(W\). Here \(e\wedge(1-f)=1-SS^*\ne0\) while \((1-e)\wedge f=0\).

10. Exercises

Exercise 1. Show that the measure \(\mu\) of Theorem 2.1 is determined by the numbers \(\tau(\chi_{[s,\infty)}(h)\chi_{[t,\infty)}(k))\), \(s,t\ge0\), \(\max(s,t)>0\).

Solution. Fix \(\varepsilon>0\). The sets \([s,\infty)\times[t,\infty)\) with \(\max(s,t)\ge\varepsilon\) form a family closed under intersection (the intersection of two of them is \([\max(s,s'),\infty)\times[\max(t,t'),\infty)\), still with \(\max\ge\varepsilon\)), and they generate the Borel sets of \(U_\varepsilon=[0,\infty)^2\setminus [0,\varepsilon)^2\): indeed \([s,s')\times[t,\infty)\) with \(s\ge\varepsilon\) is a difference of two of them, so the generated \(\sigma\)-algebra contains all Borel rectangles in \([\varepsilon,\infty)\times[0,\infty)\) and in \([0,\infty)\times[\varepsilon,\infty)\). The set \(U_\varepsilon\) is the union of \([\varepsilon,\infty)\times [0,\infty)\) and \([0,\infty)\times[\varepsilon,\infty)\), members of the family, and its measure is given by inclusion–exclusion. The restriction of \(\mu\) to \(U_\varepsilon\) is finite, so by the uniqueness theorem it is determined by the given numbers; and \(X=\bigcup_jU_{1/j}\).

Exercise 2. In Example 9.6(1), compute the joint distribution \(\mu\) of \(e,f\) and check Theorem 2.1(b) for \(f(t)=g(t)=t\) and Proposition 4.3(a) directly. Then show that the unitary \(W\) of Theorem 9.4 always satisfies \(|e-f|\le|W-1|\le\sqrt2\,|e-f|\), and check both inequalities in Example 9.6(1).

Solution. By Example 2.2, \(\mu=\cos^2\theta\,\delta_{(1,1)}+\sin^2\theta\,(\delta_{(1,0)}+\delta_{(0,1)})\). Then \(\int|H-K|^2d\mu=2\sin^2\theta=\|e-f\|_2^2\). For Proposition 4.3(a): \(E_a(e)=e\), \(E_a(f)=f\) for \(0<a<1\), and both vanish for \(a\ge1\), so the left side is \(\|e-f\|_2^2=2\sin^2\theta\), and indeed \(\int|H-K|\,d\mu=2\sin^2\theta\). By Theorem 9.4(c), \(|W-1|^2=2(1-|\mathsf b|)\) and \(|e-f|^2=1-\mathsf b^2\), both functions of \(\mathsf b\) with values in \([-1,1]\); since \(2(1-t)-(1-t^2)=(1-t)^2\ge0\) and \(1-t\le1-t^2\) for \(t=|\mathsf b|\in[0,1]\), we get \(|e-f|^2\le|W-1|^2\le2|e-f|^2\), and taking square roots of these commuting operators gives the claim. In the example, \(|e-f|=\sin\theta\) and \(|W-1|=2\sin(\theta/2)\): the inequality \(\sin\theta=2\sin(\theta/2)\cos(\theta/2)\le 2\sin(\theta/2)\) is clear, and \(2\sin(\theta/2)\le\sqrt2\sin\theta\) means \(\cos(\theta/2)\ge1/\sqrt2\), true for \(\theta\le\pi/2\), with equality at \(\theta=\pi/2\).

Exercise 3. Let \(h,k\in L^2(N,\tau)^+\) commute (their spectral projections commute). Show that \(\|h-k\|_2^2\le\|h^2-k^2\|_1\) holds with equality if and only if \(hk(h-k)=0\). Show also that positive operators with \(hk=0\) commute, and give a noncommuting pair with strict inequality.

Solution. Use the joint Borel functional calculus of the commuting pair \((h,k)\): its spectral projections \(\chi_A(h)\chi_B(k)\) lie in \(N\), so bounded or positive Borel functions \(\varphi(h,k)\) are affiliated with \(N\), and the rules of the calculus hold for them. With \(m=\min(h,k)\) we have \(|h^2-k^2|=|h-k|(h+k)\) and \(h+k=|h-k|+2m\), all factors commuting. Hence \[ \|h^2-k^2\|_1-\|h-k\|_2^2=\tau\big(|h-k|(h+k)\big)-\tau\big(|h-k|^2\big)=2\,\tau\big(|h-k|\,m\big)\ge0 , \] where \(|h-k|m\ge0\) (a product of commuting positive operators) has finite trace because \(|h-k|m\le|h-k|(h+k)\). Since \(\tau\) is faithful, equality holds iff \(|h-k|\,m=0\), i.e. iff the function \(|\lambda-\kappa|\min(\lambda,\kappa)\) vanishes on the joint spectrum; this function vanishes exactly where \(\lambda\kappa(\lambda-\kappa)\) does, so the condition is \(hk(h-k)=0\). If \(h,k\ge0\) and \(hk=0\), then \(kh=(hk)^*=0=hk\). A noncommuting pair with strict inequality is \(e,f\) of Example 2.2: \(\|e-f\|_2^2=2\sin^2\theta<2\sin\theta=\|e^2-f^2\|_1\).

Exercise 4. For \(x\in L^2(N,\tau)\) let \(\beta(a)=\|x-u_a(x)|x|\|_2\). Show that \(\beta(a)^2=\int_{(0,a]}t^2\,d\nu_{|x|}(t)\), that \(\beta\) is nondecreasing and right-continuous, that \(\beta(a)\to0\) as \(a\to0\) and \(\beta(a)\to\|x\|_2\) as \(a\to\infty\), and that \(\beta(a)^2\le a^2\, \tau(\chi_{(0,a]}(|x|))\).

Solution. With \(x=u|x|\), \(x-u_a(x)|x|=u\,\chi_{[0,a]}(|x|)|x|\), and \(u\) is isometric on the range of \(|x|\), so \(\beta(a)^2=\|\chi_{[0,a]}(|x|)|x|\|_2^2=\int t^2\chi_{(0,a]}(t)\,d\nu_{|x|}(t)\) by (B1). This is nondecreasing and right-continuous in \(a\) (continuity from above of the finite measure \(t^2d\nu_{|x|}\)), tends to \(0\) as \(a\to0\) (continuity from above at the empty set) and to \(\int t^2d\nu_{|x|}=\|x\|_2^2\) as \(a\to\infty\). Finally \(t^2\le a^2\) on \((0,a]\).

References