The entropy of a trace-preserving automorphism
Written by Claude Opus 5.5 (Anthropic), September 2026, revised October 2026. The September text was spot-checked by Claude Opus 5.5 in a separate session; the October revisions are self-checked by the writing AI. Public domain (CC0).
Introduction
In classical ergodic theory, the Kolmogorov–Sinai entropy of a measure-preserving transformation \(T\) of a probability space measures how much new information each application of \(T\) produces. One takes a finite partition \(\mathcal P\), refines it by its images \(T^{-1}\mathcal P,\dots,T^{-(k-1)}\mathcal P\), computes the Shannon entropy of the common refinement, divides by \(k\) and lets \(k\to\infty\). The supremum over all \(\mathcal P\) is the entropy \(h(T)\). The Bernoulli shift with probability vector \((p_1,\dots,p_n)\) has entropy \(-\sum_jp_j\log p_j\), so Bernoulli shifts with different entropies are not isomorphic.
This lesson builds the same invariant for an automorphism \(\theta\) of a finite von Neumann algebra \(R\) that preserves a faithful normal tracial state \(\tau\). The motivating problem is the following. Let \(M_n\) be the algebra of \(n\times n\) matrices with its normalized trace, and form the infinite tensor product \(R_n=\bigotimes_{\nu\in\mathbb Z}(M_n,\mathrm{tr})\). For every \(n\ge2\) this is the hyperfinite II₁ factor, so all the \(R_n\) are isomorphic. The \(n\)-shift \(S_n\) moves each tensor factor one place to the right. Are \(S_n\) and \(S_m\) conjugate for \(n\ne m\)? The answer is no, because \(S_n\) has entropy \(\log n\) (Corollary 6.3).
The difficulty is that the classical construction needs the common refinement \(\mathcal P\vee\mathcal Q\) of two partitions. Its noncommutative analogue would be the algebra generated by two finite-dimensional subalgebras, and that algebra is in general infinite-dimensional (Example 2.4). The lesson "Entropy of finite-dimensional subalgebras" circumvents this. It does not define a join; instead it defines directly a number \(H(N_1,\dots,N_k)\) that plays the role of the entropy of the join of \(N_1,\dots,N_k\), and proves that it has the properties one expects. With this number the definition of the entropy of \(\theta\) copies the classical one: \[ H(N,\theta)=\lim_{k\to\infty}\frac1kH(N,\theta(N),\dots,\theta^{k-1}(N)),\qquad H(\theta)=\sup_NH(N,\theta), \] the supremum being over finite-dimensional subalgebras \(N\).
What the lesson does.
- Section 3 defines \(H(\theta)\), proves that it is a conjugacy invariant, and shows that it agrees with the Kolmogorov–Sinai entropy when \(R\) is abelian.
- Section 4 proves the noncommutative Kolmogorov–Sinai theorem. If finite-dimensional subalgebras \(P_1\subset P_2\subset\cdots\) have weakly dense union in \(R\), then \(H(\theta)=\lim_qH(P_q,\theta)\) for every \(\tau\)-preserving \(\theta\). This is what makes \(H(\theta)\) computable.
- Section 5 constructs the infinite tensor product \(\bigotimes_{\mathbb Z}(B,\psi)\) of a finite-dimensional C*-algebra with a faithful state, and studies the centralizer of the product state. For \(B=M_n\) the centralizer is the hyperfinite II₁ factor and the shift restricts to an ergodic automorphism of it, the Bernoulli shift.
- Section 6 computes the entropy of these shifts: it is the von Neumann entropy \(-\sum_j\lambda_j\log\lambda_j\) of the state \(\psi\), where \(\lambda_1,\dots,\lambda_n\) are the eigenvalues of its density matrix. In particular \(H(S_n)=\log n\), and every positive real number is the entropy of some Bernoulli shift of the hyperfinite II₁ factor.
- Section 7 proves \(H(\theta^p)=|p|H(\theta)\) when the Kolmogorov–Sinai theorem applies, and the inequality \(H(\theta^p)\le|p|H(\theta)\) in general.
- Section 8 compares \(H\) with the entropy obtained from invariant abelian subalgebras, and explains why monotonicity in \(N\) is the property that drives the whole theory.
What is assumed. The lesson "Entropy of finite-dimensional subalgebras" of this course defines the joint entropy and the relative entropy of finite-dimensional subalgebras and proves their properties; Section 1 restates exactly what is used. From Course 2 we use one implication of Theorem 1.5 of "Uniqueness of the injective II₁ factor", to identify the Bernoulli shifts as automorphisms of the hyperfinite II₁ factor. Standard operator algebra theory is listed in "Results used from other lessons", with the place where each fact is proved. Classical ergodic theory is used only for motivation, and in Sections 3 and 8 to name what the abelian case recovers.
Basic references are [Connes–Størmer 1975] and [Connes 1994]. The lesson "Dynamical entropy of C*-algebras and von Neumann algebras" of this course extends the theory to state-preserving automorphisms [Connes–Narnhofer–Thirring 1987].
Conventions. Throughout, \(R\) is a von Neumann algebra with a faithful normal tracial state \(\tau\) (so \(R\) is finite). We call \((R,\tau)\) a finite tracial algebra. A subalgebra of \(R\) is a von Neumann subalgebra containing the unit of \(R\). We write \(\|x\|\) for the operator norm and \(\|x\|_2=\tau(x^*x)^{1/2}\); note \(\|x\|_2\le\|x\|\). An automorphism \(\theta\) of \(R\) is \(\tau\)-preserving if \(\tau\circ\theta=\tau\); these form a group \(\operatorname{Aut}(R,\tau)\). The function \(\eta\colon[0,\infty)\to\mathbb R\) is \(\eta(t)=-t\log t\), with \(\eta(0)=0\); logarithms are natural. We use the identity \[ \eta(st)=s\,\eta(t)+t\,\eta(s)\qquad(s,t\ge0).\tag{0.1} \] For \(x\in R\) with \(x\ge0\), \(\eta(x)\) is defined by functional calculus. The convention \(0\cdot\infty=0\) is used once, in Theorem 7.2.
Results used from other lessons
(B1) Finite-dimensional C*-algebras. Every finite-dimensional C*-algebra \(B\) is isomorphic to a direct sum \(\bigoplus_{k=1}^sM_{n_k}\) of full matrix algebras. It has a unique trace \(\operatorname{Tr}\) taking the value \(1\) on every minimal projection (the sum of the matrix traces). Every positive linear functional \(\psi\) on \(B\) is \(\psi=\operatorname{Tr}(\rho\,\cdot\,)\) for a unique \(\rho\ge0\) in \(B\), the density of \(\psi\), and \(\psi\) is faithful if and only if \(\rho\) is invertible. For every \(\rho\ge0\) in \(B\) there are pairwise orthogonal minimal projections \(e_1,\dots,e_d\) of \(B\), with \(d=\sum_kn_k\) and \(\sum_je_j=1\), and numbers \(\lambda_j\ge0\) with \(\rho=\sum_j\lambda_je_j\) (diagonalize \(\rho\) in each block). The algebraic tensor product of finitely many finite-dimensional C*-algebras is a C*-algebra of finite dimension, and its canonical trace is the tensor product of the canonical traces. A nonzero projection \(e\) is minimal in \(B\) if and only if \(eBe=\mathbb Ce\). The structure, the matrix units and the minimal projections are proved in AF-algebras, Section 2, and the traces of a multimatrix algebra in Lemma 10.2 there. The densities are the finite-dimensional case of the duality between trace-class and bounded operators (Compact and trace-class operators, Section 6), applied in each block, and \(\rho\) is diagonalized block by block. For the tensor product, \(M_m\otimes M_n\cong M_{mn}\) with \(\operatorname{Tr}\otimes\operatorname{Tr}=\operatorname{Tr}\). If \(eBe=\mathbb Ce\), every subprojection of \(e\) is \(0\) or \(e\); conversely, if \(e\) is minimal, the finite-dimensional algebra \(eBe\) has no projections other than \(0\) and \(e\), and it is spanned by its projections (Lemma 2.2(1) of the lesson on AF-algebras).
(B2) Conditional expectations. Let \((R,\tau)\) be a finite tracial algebra and \(Q\) a subalgebra. There is a unique map \(E_Q\colon R\to Q\) with \(\tau(E_Q(x)y)=\tau(xy)\) for all \(x\in R\), \(y\in Q\). It is a normal unital completely positive \(Q\)-bimodule map (\(E_Q(axb)=aE_Q(x)b\) for \(a,b\in Q\)), it satisfies \(\tau\circ E_Q=\tau\), and it is idempotent onto \(Q\). Proved in Integration for a trace, Theorem 9.1; complete positivity is Contractive retractions and the algebraic structure of expectations, §CE-006. See also Entropy of finite-dimensional subalgebras, (B1).
(B3) Density theorems. Let \(A_0\subset B(\mathcal H)\) be a unital \(*\)-algebra. Its weak closure is the bicommutant \(A_0''\) (the double commutant theorem), and the unit ball of \(A_0\) is strongly dense in the unit ball of \(A_0''\) (the Kaplansky density theorem). Proved in The double commutant theorem, Theorem 4.4 and Kaplansky's density theorem and its consequences, Theorem 7.1.
(B4) Topologies on bounded sets. Fix a normal positive functional \(\omega\) on a von Neumann algebra \(M\). On bounded subsets of \(M\) the map \(x\mapsto\omega(x^*x)^{1/2}\) is continuous for the strong operator topology. On bounded subsets the weak and the \(\sigma\)-weak topologies coincide, and the \(\sigma\)-weak topology of \(M\) does not depend on the faithful normal representation. A normal \(\omega\) is \(\sigma\)-strongly continuous (Compact and trace-class operators, Theorem 9.1(ii)), so \(\omega=\sum_n\langle\,\cdot\,\xi_n,\xi_n\rangle\) with \(\sum_n\|\xi_n\|^2<\infty\) (The double commutant theorem, Theorem 10.1), and \(\omega(x^*x)^{1/2}=(\sum_n\|x\xi_n\|^2)^{1/2}\) is a \(\sigma\)-strong seminorm. On bounded sets the strong and \(\sigma\)-strong, and the weak and \(\sigma\)-weak topologies agree (Compact and trace-class operators, Lemma 8.5). Independence of the representation is The universal enveloping von Neumann algebra of a C*-algebra, and W*-algebras, Corollary 11.4.
(B5) Finite factors. A factor with a faithful normal tracial state is either isomorphic to some \(M_n\) or of type II₁. A II₁ factor has exactly one normal tracial state, so every isomorphism between II₁ factors carries the trace to the trace. A von Neumann algebra acting on a separable Hilbert space has separable predual. The first statement is Projections and types of von Neumann algebras, Corollaries 7.3 and 10.4; uniqueness of the trace is Integration for a trace, Corollary 7.4. For the last statement, the predual is a quotient of the trace class (Compact and trace-class operators, Theorem 9.4(a)); finite-rank operators are dense in the trace class (Corollary 4.4(c) there), and a rank-one operator \(\theta_{\xi,\eta}\) depends continuously on \(\xi,\eta\) in trace norm, so the finite sums of \(\theta_{\xi,\eta}\) with \(\xi,\eta\) from a countable dense set are dense when the Hilbert space is separable.
1. What is used from "Entropy of finite-dimensional subalgebras"
In this section, every result cited by number belongs to the lesson "Entropy of finite-dimensional subalgebras", unless it is marked "of this lesson".
Let \((R,\tau)\) be a finite tracial algebra. For \(k\ge1\) let \(\mathcal S_k\) be the set of families \(x=(x_i)_{i\in\mathbb N^k}\) of positive elements of \(R\), only finitely many of them nonzero, with \(\sum_ix_i=1\). For \(l\in\{1,\dots,k\}\) and \(j\in\mathbb N\) the \(l\)-th marginal is \[ x^{(l)}_j=\sum_{i\,:\,i_l=j}x_i . \]
Definition 1.1. For finite-dimensional subalgebras \(N_1,\dots,N_k\) of \(R\) the joint entropy is \[ H(N_1,\dots,N_k)=\sup_{x\in\mathcal S_k}\Big(\sum_{i\in\mathbb N^k}\eta\big(\tau(x_i)\big) -\sum_{l=1}^k\sum_{j\in\mathbb N}\tau\big(\eta(E_{N_l}(x^{(l)}_j))\big)\Big),\tag{1.1} \] and for finite-dimensional subalgebras \(N,P\) the relative entropy is \[ H(N\mid P)=\sup_{x\in\mathcal S_1}\sum_{j\in\mathbb N}\Big(\tau\big(\eta(E_P(x_j))\big)- \tau\big(\eta(E_N(x_j))\big)\Big).\tag{1.2} \]
These are Definition 3.1 and Definition 5.1 (the latter restricted to finite-dimensional \(N,P\)). There the families are indexed by products \(I_1\times\dots\times I_k\) of finite sets; this gives the same suprema. A finitely supported family indexed by \(\mathbb N^k\) may be restricted to a finite product \(I_1\times\dots\times I_k\) containing its support, and a family indexed by finite sets \(I_l\) may be transported to \(\mathbb N^k\) by injections \(I_l\to\mathbb N\) and extended by \(0\). Zero entries contribute \(\eta(0)=0\), and zero marginals contribute \(\tau\eta(E_{N_l}(0))=0\).
That lesson proves the following facts. All algebras below are finite-dimensional subalgebras of \(R\).
Facts 1.2.
- (a) Range and symmetry (Remark 3.2(a),(b) and Corollary 4.2). \(0\le H(N_1,\dots,N_k)<\infty\), and \(H(N_1,\dots,N_k)\) does not change when the \(N_l\) are permuted.
- (b) Monotonicity (Proposition 3.3). If \(N_l\subset P_l\) for all \(l\), then \(H(N_1,\dots,N_k)\le H(P_1,\dots,P_k)\).
- (c) Subadditivity (Proposition 3.4). \(H(N_1,\dots,N_k,N_{k+1},\dots,N_m)\le H(N_1,\dots,N_k)+H(N_{k+1},\dots,N_m)\) for \(1\le k<m\).
- (d) Merging (Proposition 3.5; for \(r=0\) also Corollary 3.6(4)). If \(N_1,\dots,N_m\subset P\), then for all \(Q_1,\dots,Q_r\) (\(r\ge0\)) \[ H(N_1,\dots,N_m,Q_1,\dots,Q_r)\le H(P,Q_1,\dots,Q_r). \] By (a) the merged algebras may stand anywhere in the list.
- (e) Atoms (Theorem 4.1). If \((e_\alpha)_{\alpha}\) is a family of pairwise orthogonal minimal projections of \(N\) with \(\sum_\alpha e_\alpha=1\), then \(H(N)=\sum_\alpha\eta(\tau(e_\alpha))\).
- (f) Commuting algebras (Theorem 4.4). If \(N_1,\dots,N_k\) commute pairwise, then the algebra \((N_1\cup\dots\cup N_k)''\) is finite-dimensional and \(H(N_1,\dots,N_k)=H\big((N_1\cup\dots\cup N_k)''\big)\). We use this only for abelian \(N_l\).
- (g) Moving the arguments (Proposition 5.3). \(H(N_1,\dots,N_k)\le H(P_1,\dots,P_k)+\sum_{l=1}^kH(N_l\mid P_l)\).
From (e): a family as in (e) has at most \(\dim N\) members, and concavity of \(\eta\) gives \(\sum_{\alpha=1}^d\eta(t_\alpha)\le d\,\eta(1/d)=\log d\) whenever \(\sum_\alpha t_\alpha=1\). Hence (Corollary 4.2) \[ H(N)\le\log\dim N.\tag{1.3} \]
The relative entropy has further properties (Proposition 5.2): it vanishes when \(N\subset P\), increases in \(N\), decreases in \(P\), and satisfies a triangle inequality. For commuting \(N_1,N_2\), Theorem 5.5 computes \(H(N_2\mid N_1)=H\big((N_1\cup N_2)''\big)-H(N_1)\); the quantity \(H\big((N_1\cup N_2)''\mid N_1\big)\) agrees with this when \(N_1\) is abelian, but can be strictly larger otherwise (Example 5.6 there). None of these facts is needed below; the only relative-entropy results used are (g), the invariance proved in Lemma 2.1 of this lesson, and the continuity theorem.
For subalgebras \(N,P\) of \(R\) and \(\delta>0\) we write \(N\overset{\delta}{\subset}P\) if for every \(x\in N\) with \(\|x\|\le1\) there is \(y\in P\) with \(\|y\|\le1\) and \(\|x-y\|_2<\delta\) (Definition 6.1).
Theorem 1.3 (continuity of relative entropy; Theorem 7.3). For every \(n\in\mathbb N\) and \(\varepsilon>0\) there is \(\delta>0\) such that \(H(N\mid P)<\varepsilon\) for all subalgebras \(N,P\) of \(R\) with \(\dim N\le n\) and \(N\overset{\delta}{\subset}P\).
Lemma 4.1(b) of this lesson, which combines Theorem 1.3 with the Kaplansky density theorem, is Corollary 7.4 of that lesson; we include its short proof.
2. Invariance and two elementary inequalities
The joint entropy is built from \(\tau\) and the conditional expectations only, so it is invariant under trace-preserving isomorphisms.
Lemma 2.1 (invariance). Let \((R,\tau)\) and \((R',\tau')\) be finite tracial algebras and \(\gamma\colon R\to R'\) a \(*\)-isomorphism with \(\tau'\circ\gamma=\tau\).
- (a) For every subalgebra \(Q\) of \(R\), \(E_{\gamma(Q)}\circ\gamma=\gamma\circ E_Q\).
- (b) For finite-dimensional subalgebras of \(R\), \(H(\gamma(N_1),\dots,\gamma(N_k))=H(N_1,\dots,N_k)\) and \(H(\gamma(N)\mid\gamma(P))=H(N\mid P)\), the left sides computed in \((R',\tau')\).
In particular this holds for every \(\gamma\in\operatorname{Aut}(R,\tau)\); for automorphisms it is Remark 3.2(c) and Proposition 5.2(5) of "Entropy of finite-dimensional subalgebras". We need it also for isomorphisms between two algebras (Proposition 3.3(e)), so we give the proof.
Proof. (a) The map \(F=\gamma\circ E_Q\circ\gamma^{-1}\) takes values in \(\gamma(Q)\). For \(x'\in R'\) and \(y'=\gamma(y)\in\gamma(Q)\), \[ \tau'(F(x')y')=\tau'\big(\gamma(E_Q(\gamma^{-1}(x'))y)\big)=\tau\big(E_Q(\gamma^{-1}(x'))y\big) =\tau(\gamma^{-1}(x')y)=\tau'(x'y'). \] By the uniqueness in (B2), \(F=E_{\gamma(Q)}\).
(b) The map \(x\mapsto\gamma(x)=(\gamma(x_i))_i\) is a bijection of \(\mathcal S_k(R)\) onto \(\mathcal S_k(R')\). It commutes with taking marginals. Functional calculus commutes with \(*\)-isomorphisms, so \(\eta(E_{\gamma(N_l)}(\gamma(y)))=\gamma(\eta(E_{N_l}(y)))\) by (a), and \(\tau'\circ\gamma=\tau\). Hence each term of (1.1) for \(\gamma(x)\) and \(\gamma(N_1),\dots,\gamma(N_k)\) equals the corresponding term for \(x\) and \(N_1,\dots,N_k\). Taking suprema gives the first equality; (1.2) is treated in the same way. \(\square\)
Lemma 2.2 (adding an algebra). For finite-dimensional subalgebras \(N_1,\dots,N_{k+1}\), \[ H(N_1,\dots,N_k)\le H(N_1,\dots,N_k,N_{k+1}). \]
This is Corollary 3.6(1) of "Entropy of finite-dimensional subalgebras"; the proof is one line.
Proof. Let \(x\in\mathcal S_k\). Define \(y\in\mathcal S_{k+1}\) by \(y_{(i,1)}=x_i\) and \(y_{(i,j)}=0\) for \(j\ne1\). The terms \(\eta(\tau(y_{(i,j)}))\) sum to \(\sum_i\eta(\tau(x_i))\), since \(\eta(0)=0\). For \(l\le k\) the \(l\)-th marginals of \(y\) and \(x\) agree. The \((k+1)\)-th marginal of \(y\) is \(1\) at \(j=1\) and \(0\) elsewhere, and \(\eta(E_{N_{k+1}}(1))=\eta(1)=0\), \(\eta(0)=0\). So the expression in (1.1) for \(y\) equals the expression for \(x\). Taking the supremum over \(x\) proves the claim. \(\square\)
Corollary 2.3. If the list \(L\) of finite-dimensional subalgebras is obtained from the list \(L'\) by deleting some entries, then \(H(L)\le H(L')\).
Proof. By symmetry (Fact 1.2(a)) the deleted entries may be put at the end; apply Lemma 2.2 once per deleted entry. \(\square\)
Example 2.4 (there is no join). Let \(R=L^\infty([0,\pi/2])\otimes M_2\), the algebra of bounded measurable functions from \([0,\pi/2]\) to \(M_2\), with the faithful normal tracial state \(\tau(f)=\frac2\pi\int_0^{\pi/2}\mathrm{tr}(f(t))\,dt\). Let \(p\) be the constant function \(\left(\begin{smallmatrix}1&0\\0&0\end{smallmatrix}\right)\) and let \(q(t)\) be the projection onto the line spanned by \((\cos t,\sin t)\). Put \(N_1=\mathbb Cp+\mathbb C(1-p)\) and \(N_2=\mathbb Cq+\mathbb C(1-q)\), two two-dimensional abelian subalgebras. Then \(pqp\) is the function \(t\mapsto\cos^2t\left(\begin{smallmatrix}1&0\\0&0\end{smallmatrix}\right)\). Its spectrum is \([0,1]\), which is infinite, so the C*-algebra generated by \(pqp\), and with it \((N_1\cup N_2)''\), is infinite-dimensional. Thus there is no finite-dimensional algebra that could serve as \(N_1\vee N_2\), although \(H(N_1,N_2)\) is defined and, by Fact 1.2(c) and (1.3), at most \(2\log2\).
3. The entropy of an automorphism
Lemma 3.1 (Fekete's lemma). Let \(a_1,a_2,\dots\) be real numbers with \(a_k\ge0\) and \(a_{k+l}\le a_k+a_l\) for all \(k,l\). Then \(\lim_ka_k/k\) exists and equals \(\inf_ka_k/k\).
Proof. Put \(a_0=0\), so the inequality holds for \(k,l\ge0\). Fix \(m\ge1\). Every \(k\) can be written \(k=qm+r\) with \(0\le r<m\), and then \(a_k\le qa_m+a_r\le qa_m+\max_{r<m}a_r\). Dividing by \(k\) and letting \(k\to\infty\) (so \(q/k\to1/m\)) gives \(\limsup_ka_k/k\le a_m/m\). Hence \(\limsup_ka_k/k\le\inf_ma_m/m\le\liminf_ka_k/k\). \(\square\)
Definition 3.2. Let \((R,\tau)\) be a finite tracial algebra and \(\theta\in\operatorname{Aut}(R,\tau)\). For a finite-dimensional subalgebra \(N\) and \(k\ge1\) put \[ H_k(N,\theta)=H\big(N,\theta(N),\dots,\theta^{k-1}(N)\big). \] The entropy of \(\theta\) with respect to \(N\) and the entropy of \(\theta\) are \[ H(N,\theta)=\lim_{k\to\infty}\frac1kH_k(N,\theta),\qquad H(\theta)=\sup_NH(N,\theta)\in[0,\infty], \] the supremum being over all finite-dimensional subalgebras \(N\) of \(R\).
Using \(k+1\) algebras \(N,\dots,\theta^k(N)\) and dividing by \(k\) gives the same limit, since \(\frac{k+1}k\to1\).
Proposition 3.3. Let \(\theta\in\operatorname{Aut}(R,\tau)\) and let \(N,P\) be finite-dimensional subalgebras.
- (a) The sequence \(k\mapsto H_k(N,\theta)\) is subadditive, so the limit defining \(H(N,\theta)\) exists and equals \(\inf_kH_k(N,\theta)/k\). Moreover \(0\le H(N,\theta)\le H(N)\le\log\dim N\).
- (b) If \(N\subset P\), then \(H(N,\theta)\le H(P,\theta)\).
- (c) \(H(N,\theta)\le H(P,\theta)+H(N\mid P)\).
- (d) \(H(\theta(N),\theta)=H(N,\theta)\) and \(H(N,\theta^{-1})=H(N,\theta)\). Hence \(H(\theta^{-1})=H(\theta)\).
- (e) Let \(\gamma\colon(R,\tau)\to(R',\tau')\) be a trace-preserving isomorphism and \(\theta'=\gamma\theta\gamma^{-1}\). Then \(H(\gamma(N),\theta')=H(N,\theta)\) and \(H(\theta')=H(\theta)\).
Proof. (a) By Fact 1.2(c) and Lemma 2.1 applied to \(\theta^k\), \[ H_{k+l}(N,\theta)\le H_k(N,\theta)+H\big(\theta^k(N),\dots,\theta^{k+l-1}(N)\big)=H_k(N,\theta)+H_l(N,\theta). \] The numbers \(H_k(N,\theta)\) are finite and nonnegative by Fact 1.2(a), so Lemma 3.1 applies. The infimum is at most \(H_1(N,\theta)=H(N)\), and \(H(N)\le\log\dim N\) by (1.3).
(b) \(\theta^j(N)\subset\theta^j(P)\) for all \(j\); apply Fact 1.2(b), divide by \(k\) and let \(k\to\infty\).
(c) By Fact 1.2(g) and Lemma 2.1, \[ H_k(N,\theta)\le H_k(P,\theta)+\sum_{j=0}^{k-1}H\big(\theta^j(N)\mid\theta^j(P)\big)=H_k(P,\theta)+k\,H(N\mid P). \] Divide by \(k\) and let \(k\to\infty\).
(d) \(H_k(\theta(N),\theta)=H(\theta(N),\dots,\theta^k(N))=H_k(N,\theta)\) by Lemma 2.1. Next, applying \(\theta^{k-1}\) and then symmetry, \[ H_k(N,\theta^{-1})=H\big(N,\theta^{-1}(N),\dots,\theta^{-(k-1)}(N)\big) =H\big(\theta^{k-1}(N),\dots,\theta(N),N\big)=H_k(N,\theta). \] Taking suprema over \(N\) gives \(H(\theta^{-1})=H(\theta)\).
(e) Since \(\theta'^j(\gamma(N))=\gamma(\theta^j(N))\), Lemma 2.1 gives \(H_k(\gamma(N),\theta')=H_k(N,\theta)\). As \(N\mapsto\gamma(N)\) is a bijection between the finite-dimensional subalgebras of \(R\) and of \(R'\), the suprema agree. \(\square\)
Corollary 3.4. Let \(R\) and \(R'\) be II₁ factors with their unique normal tracial states. If \(\theta\in\operatorname{Aut}R\), \(\theta'\in\operatorname{Aut}R'\) and \(\gamma\colon R\to R'\) is an isomorphism with \(\gamma\theta\gamma^{-1}=\theta'\), then \(H(\theta)=H(\theta')\).
Proof. Every automorphism of a II₁ factor preserves its trace, and \(\gamma\) carries the trace of \(R\) to that of \(R'\) (B5). Apply Proposition 3.3(e). \(\square\)
So for a II₁ factor, \(H\) is defined on the whole automorphism group and is a conjugacy invariant.
Example 3.5 (zero entropy).
- (i) If \(R\) is finite-dimensional, then \(H(\theta)=0\) for every \(\theta\). Indeed all \(\theta^j(N)\) lie in \(R\), so Fact 1.2(d) with \(r=0\) gives \(H_k(N,\theta)\le H(R)\), a bound independent of \(k\).
- (ii) If \(\theta^p=\mathrm{id}\) for some \(p\ge1\), then \(H(\theta)=0\). The list \(N,\theta(N),\dots,\theta^{k-1}(N)\) contains only the \(p\) algebras \(N,\dots,\theta^{p-1}(N)\), repeated. Merging equal entries (Fact 1.2(d) with \(P=\theta^t(N)\), once for each \(t\)) gives \(H_k(N,\theta)\le H(N,\theta(N),\dots,\theta^{p-1}(N))\) for all \(k\ge p\). In particular \(H(\mathrm{id})=0\).
The commutative case recovers the classical theory. Let \(A\) be an abelian finite tracial algebra. For a finite-dimensional subalgebra \(Q\subset A\), the minimal projections of \(Q\) are its atoms; they are pairwise orthogonal with sum \(1\). Put \[ h(Q)=\sum_{e\text{ atom of }Q}\eta(\tau(e)), \] and for finite-dimensional \(Q_1,Q_2\subset A\) let \(Q_1\vee Q_2=(Q_1\cup Q_2)''\), again finite-dimensional. If \(A=L^\infty(X,\mu)\) with \(\tau=\int\cdot\,d\mu\), the finite-dimensional subalgebras of \(A\) correspond to the finite measurable partitions of \(X\) (up to null sets): the atoms are the indicator functions of the pieces. Under this correspondence \(\vee\) is the common refinement and \(h\) is the Shannon entropy of a partition. For a \(\tau\)-preserving automorphism \(\theta\) of \(A\), the Kolmogorov–Sinai entropy is \[ h(\theta\mid A)=\sup_{N}\lim_{k\to\infty}\frac1kh\big(N\vee\theta(N)\vee\dots\vee\theta^{k-1}(N)\big), \] over finite-dimensional \(N\subset A\). This is the measure-algebra form of the classical definition; see [Fremlin, Section 385].
Proposition 3.6 (the commutative case). Let \((R,\tau)\) be a finite tracial algebra, \(\theta\in\operatorname{Aut}(R,\tau)\), and \(A\subset R\) an abelian subalgebra with \(\theta(A)=A\). For every finite-dimensional \(N\subset A\) and \(k\ge1\), \[ H_k(N,\theta)=h\big(N\vee\theta(N)\vee\dots\vee\theta^{k-1}(N)\big). \] Consequently \(h(\theta|_A\mid A)=\sup_{N\subset A}H(N,\theta)\le H(\theta)\), with equality if \(R=A\).
Proof. The algebras \(\theta^j(N)\subset A\) are abelian and commute. By Fact 1.2(f), \(H_k(N,\theta)=H(Q)\) with \(Q=N\vee\dots\vee\theta^{k-1}(N)\). The atoms of \(Q\) are minimal projections of \(Q\) with sum \(1\), so \(H(Q)=h(Q)\) by Fact 1.2(e). Divide by \(k\), let \(k\to\infty\) and take the supremum over \(N\subset A\). If \(R=A\), every finite-dimensional subalgebra lies in \(A\). \(\square\)
Note that the right side of Fact 1.2(e) involves only \(N\) and \(\tau|_N\). More generally, although the supremum in (1.1) runs over families in all of \(R\), the joint entropy \(H(N_1,\dots,N_k)\) depends only on the algebra generated by \(N_1,\dots,N_k\) and on the restriction of \(\tau\) to it: replacing each \(x_i\) by its conditional expectation onto that algebra does not change the expression in (1.1) (Remark 3.2(d) of "Entropy of finite-dimensional subalgebras").
4. The Kolmogorov–Sinai theorem
The supremum in the definition of \(H(\theta)\) runs over all finite-dimensional subalgebras. The Kolmogorov–Sinai theorem reduces it to one increasing sequence. The two ingredients are Proposition 3.3(c), which bounds \(H(N,\theta)\) by \(H(P,\theta)\) up to \(H(N\mid P)\), and the continuity theorem, which makes \(H(N\mid P)\) small when \(N\) is almost contained in \(P\).
A reference for this section is [Connes–Størmer 1975].
Lemma 4.1 (approximation). Let \((R,\tau)\) be a finite tracial algebra and \(P_1\subset P_2\subset\cdots\) finite-dimensional subalgebras with \(\bigcup_qP_q\) weakly dense in \(R\), and fix a finite-dimensional subalgebra \(N\).
- (a) For every \(\delta>0\) there is \(q_0\) with \(N\overset{\delta}{\subset}P_q\) for all \(q\ge q_0\).
- (b) For every \(\varepsilon>0\) there is \(q_0\) with \(H(N\mid P_q)<\varepsilon\) for all \(q\ge q_0\).
Proof. (a) The set \(A_0=\bigcup_qP_q\) is a unital \(*\)-algebra with weak closure \(R\). By the Kaplansky density theorem (B3), every \(x\) in the unit ball of \(R\) is the strong limit of a net in the unit ball of \(A_0\); by (B4), applied to \(\omega=\tau\), the net converges to \(x\) in \(\|\cdot\|_2\). The unit ball of the finite-dimensional space \(N\) is norm-compact, so it contains a finite set \(x_1,\dots,x_m\) such that every \(x\) in the ball has \(\|x-x_s\|<\delta/2\) for some \(s\). Choose \(y_s\) in the unit ball of \(A_0\) with \(\|x_s-y_s\|_2<\delta/2\), say \(y_s\in P_{q_s}\), and let \(q_0=\max_sq_s\). For \(q\ge q_0\) all \(y_s\) lie in \(P_q\), and for \(x\) in the unit ball of \(N\), with \(s\) as above, \[ \|x-y_s\|_2\le\|x-x_s\|+\|x_s-y_s\|_2<\delta . \] (b) Let \(n=\dim N\), take \(\delta\) from Theorem 1.3 for \(n\) and \(\varepsilon\), and apply (a). \(\square\)
Theorem 4.2 (Kolmogorov–Sinai theorem). Let \((R,\tau)\) be a finite tracial algebra and \(P_1\subset P_2\subset\cdots\) finite-dimensional subalgebras of \(R\) whose union is weakly dense in \(R\). Then for every \(\theta\in\operatorname{Aut}(R,\tau)\) \[ H(\theta)=\lim_{q\to\infty}H(P_q,\theta)=\sup_qH(P_q,\theta). \]
Reference: [Connes–Størmer 1975, Section 4].
Proof. By Proposition 3.3(b) the sequence \(H(P_q,\theta)\) is increasing, so its limit is its supremum, and this is at most \(H(\theta)\). Conversely, let \(N\) be finite-dimensional and \(\varepsilon>0\). By Lemma 4.1(b) there is \(q\) with \(H(N\mid P_q)<\varepsilon\), and Proposition 3.3(c) gives \[ H(N,\theta)\le H(P_q,\theta)+\varepsilon\le\sup_qH(P_q,\theta)+\varepsilon . \] Since \(N\) and \(\varepsilon\) are arbitrary, \(H(\theta)\le\sup_qH(P_q,\theta)\). \(\square\)
Remark 4.3 (where the hypothesis holds). We call a finite tracial algebra with such a sequence \((P_q)\) approximately finite-dimensional. Examples:
- every finite-dimensional \(R\) (take \(P_q=R\));
- the hyperfinite II₁ factor, which is the weak closure of \(M_2\subset M_2\otimes M_2\subset\cdots\);
- every abelian \(A\) with separable predual: it is generated by a sequence of projections \(f_1,f_2,\dots\), and \(P_q\), the algebra generated by \(f_1,\dots,f_q\), is finite-dimensional;
- the centralizers of product states studied in Section 5 (Theorem 5.7), which include the infinite tensor products of finite-dimensional algebras with faithful tracial states.
Without such a sequence the statement has no content, and the theorem cannot even be formulated. Theorem 4.2 does not need \(R\) to be a factor, nor of type II₁.
Remark 4.4 (generators). In the classical theory one has more: if \(N\) generates under \(\theta\), that is, the algebras \(\theta^k(N)\), \(k\in\mathbb Z\), together generate everything, then \(h(\theta)=h(N,\theta)\). The proof uses the two-sided windows \(N_q=\theta^{-q}(N)\vee\dots\vee\theta^q(N)\), which increase to everything. One-sided joins \(N\vee\theta(N)\vee\dots\vee\theta^q(N)\) need not: for the bilateral Bernoulli shift on \(\{0,1\}^{\mathbb Z}\), the partition by the zeroth coordinate generates under all integer powers of the shift, but its joins with forward iterates involve only the coordinates on one side of \(0\). Two facts give \(h(N_q,\theta)=h(N,\theta)\). First, \(h(\theta^{-q}(P),\theta)=h(P,\theta)\), because \(\theta^{-q}\) commutes with \(\theta\) and preserves entropy. Second, \(h(P\vee\theta(P)\vee\dots\vee\theta^m(P),\theta)=h(P,\theta)\): the join of the \(\theta^j\)-images, \(0\le j<n\), of this window is \(P\vee\theta(P)\vee\dots\vee\theta^{n+m-1}(P)\), and its entropy divided by \(n\) has the same limit as when divided by \(n+m\). Since \(N_q=\theta^{-q}\big(N\vee\theta(N)\vee\dots\vee\theta^{2q}(N)\big)\), both facts give \(h(N_q,\theta)=h(N,\theta)\), and the Kolmogorov–Sinai theorem (the abelian case of Theorem 4.2) gives \(h(\theta)=\lim_qh(N_q,\theta)=h(N,\theta)\). In the noncommutative setting the joins need not exist (Example 2.4). When the algebras generated by \(\theta^{-q}(N),\dots,\theta^q(N)\) are finite-dimensional, Theorem 4.2 still applies to them, but computing \(H\) of these algebras requires an extra argument. Section 6 supplies it for shifts.
5. Infinite tensor products and centralizers
This section is pure operator algebra; entropy returns in Section 6.
5.1 The construction
Fix a C*-algebra \(B\) of finite dimension and a faithful state \(\psi\) on \(B\). Write \(\psi=\operatorname{Tr}(\rho\,\cdot\,)\) with \(\rho\) invertible (B1), and fix pairwise orthogonal minimal projections \(e_1,\dots,e_d\) of \(B\) with sum \(1\) and \(\rho=\sum_j\lambda_je_j\). Then \(\lambda_j=\operatorname{Tr}(\rho e_j)= \psi(e_j)>0\) and \(\sum_j\lambda_j=1\). The von Neumann entropy of \(\psi\) is \[ S(\psi)=\operatorname{Tr}\eta(\rho)=\sum_{j=1}^d\eta(\lambda_j).\tag{5.1} \] Let \(D=\operatorname{span}\{e_1,\dots,e_d\}\), an abelian subalgebra of \(B\).
For a finite interval \(I\subset\mathbb Z\) put \(B_I=\bigotimes_{\nu\in I}B\), with the product state \(\psi_I=\bigotimes_{\nu\in I}\psi\), the canonical trace \(\operatorname{Tr}_I\) and \(\rho_I=\bigotimes_{\nu\in I}\rho\). Then \(\psi_I=\operatorname{Tr}_I(\rho_I\,\cdot\,)\) and \(\rho_I\) is invertible. Let \(D_I=\bigotimes_{\nu\in I}D\subset B_I\). For intervals \(I\subset J\), identify \(B_I\) with \(B_I\otimes1\subset B_J=B_I\otimes B_{J\setminus I}\) (after reordering tensor factors). Then \[ \rho_J=\rho_I\otimes \rho_{J\setminus I},\qquad \psi_J(x\otimes1)=\psi_I(x),\qquad D_I\subset D_J . \] Let \(A=\bigcup_IB_I\), a unital \(*\)-algebra in which each \(B_I\) is a C*-algebra of finite dimension, and let \(\psi_\infty\) be the state on \(A\) that restricts to \(\psi_I\) on each \(B_I\).
Because each \(\psi_I\) is faithful, \(\langle a,b\rangle=\psi_\infty(b^*a)\) is an inner product on \(A\). Let \(\mathcal H\) be the completion and \(\Omega\in\mathcal H\) the vector given by \(1\in A\). For \(x,a\in B_J\) we have \(x^*x\le\|x\|^21\) in the C*-algebra \(B_J\), so \(\psi_J(a^*x^*xa)\le\|x\|^2\psi_J(a^*a)\). Hence left multiplication by \(x\) extends to a bounded operator \(\pi(x)\) on \(\mathcal H\) with \(\|\pi(x)\|\le\|x\|\), and \(\pi\) is a unital \(*\)-homomorphism of \(A\). It is injective, because \(\pi(x)\Omega=0\) forces \(\psi_J(x^*x)=0\). We usually write \(x\) for \(\pi(x)\). Define \[ M=\pi(A)'',\qquad \varphi(x)=\langle x\Omega,\Omega\rangle\ (x\in M),\qquad M_I=\pi(B_I). \] The pair \((M,\varphi)\) is the infinite tensor product \(\bigotimes_{\nu\in\mathbb Z}(B,\psi)\). The state \(\varphi\) is normal, \(\varphi\circ\pi=\psi_\infty\), and \(\mathcal H\) is separable because \(A\) has countable dimension.
Lemma 5.1.
- (a) \(\Omega\) is cyclic for \(M'\). Hence \(\Omega\) is separating for \(M\) and \(\varphi\) is faithful.
- (b) (Product property.) If \(I,J\) are disjoint finite intervals, \(x\in M_I\) and \(y\in M_J\), then \(\varphi(xy)=\varphi(x)\varphi(y)\).
- (c) (Eigenoperators.) Let \(y\in B_I\) and \(c>0\) with \(\rho_Iy\rho_I^{-1}=cy\). Then \(\varphi(yw)=c\,\varphi(wy)\) for all \(w\in M\).
Proof. (a) Let \(y\in B_I\) and put \(c_y=\|\rho_I^{-1/2}y\rho_I^{1/2}\|\). For an interval \(J\supset I\), \(\rho_J^{-1/2}y\rho_J^{1/2}=(\rho_I^{-1/2}y\rho_I^{1/2})\otimes1\) also has norm \(c_y\). Writing \(z=\rho_J^{-1/2}y\rho_J^{1/2}\) we get \(y\rho_Jy^*=\rho_J^{1/2}zz^*\rho_J^{1/2}\le c_y^2\rho_J\). Hence for \(w\in B_J\) \[ \psi_J(y^*w^*wy)=\operatorname{Tr}_J(wy\rho_Jy^*w^*)\le c_y^2\operatorname{Tr}_J(w\rho_Jw^*)=c_y^2\psi_J(w^*w). \] So \(w\Omega\mapsto wy\Omega\) extends to a bounded operator \(r(y)\) on \(\mathcal H\). It commutes with every \(\pi(x)\), \(x\in A\), because \(x(wy)=(xw)y\). Thus \(r(y)\in\pi(A)'=M'\), and \(r(y)\Omega=y\Omega\). Therefore \(M'\Omega\supset\pi(A)\Omega\) is dense. If \(x\in M\) and \(x\Omega=0\), then \(xr\Omega=rx\Omega=0\) for all \(r\in M'\), so \(x=0\). Finally \(\varphi(x^*x)=\|x\Omega\|^2\).
(b) Let \(K\) be an interval containing \(I\cup J\). In \(B_K\) the product \(xy\) is the elementary tensor \(x\otimes y\otimes1\), and \(\psi_K\) is a product state, so \(\psi_K(xy)=\psi_I(x)\psi_J(y)\).
(c) First let \(w\in B_J\) with \(J\supset I\). Then \(\rho_Jy\rho_J^{-1}=(\rho_Iy\rho_I^{-1})\otimes1=cy\), so \[ \psi_J(yw)=\operatorname{Tr}_J(\rho_Jyw)=c\operatorname{Tr}_J(y\rho_Jw)=c\operatorname{Tr}_J(\rho_Jwy)=c\,\psi_J(wy). \] Both \(w\mapsto\varphi(yw)=\langle w\Omega,y^*\Omega\rangle\) and \(w\mapsto\varphi(wy)=\langle wy\Omega,\Omega\rangle\) are weakly continuous, and \(\pi(A)\) is weakly dense in \(M\) (B3). \(\square\)
5.2 Automorphisms of the tensor product
Lemma 5.2 (the shift). There is an automorphism \(\alpha\) of \(M\) with \(\varphi\circ\alpha=\varphi\) and \(\alpha(\pi(x))=\pi(s(x))\) for \(x\in A\), where \(s\) is the automorphism of \(A\) that moves the tensor factor at site \(\nu\) to site \(\nu+1\). In particular \(\alpha(M_I)=M_{I+1}\) and \(\alpha(\pi(D_I))=\pi(D_{I+1})\).
Proof. Since all sites carry the same \((B,\psi)\), \(\psi_\infty\circ s=\psi_\infty\). So \(U(a\Omega)=s(a)\Omega\) defines an isometry with dense range, that is, a unitary on \(\mathcal H\), with \(U\Omega=\Omega\) and \(U\pi(x)U^*=\pi(s(x))\). The automorphism \(\operatorname{Ad}U\) of \(B(\mathcal H)\) maps \(\pi(A)\) onto \(\pi(A)\), hence its weak closure \(M\) onto \(M\). Let \(\alpha\) be its restriction. Then \(\varphi(\alpha(x))=\langle xU^*\Omega,U^*\Omega\rangle=\varphi(x)\). \(\square\)
Lemma 5.3 (permutations). Let \(B=M_n\), let \(K\) be a finite interval and \(g\) a bijection of \(K\). There is a unitary \(V_g\in B_K\) that commutes with \(\rho_K\) and satisfies \(V_g\big(\bigotimes_{\nu\in K}b_\nu\big)V_g^*=\bigotimes_{\nu\in K}b_{g^{-1}(\nu)}\). The automorphism \(\beta=\operatorname{Ad}\pi(V_g)\) of \(M\) satisfies \(\varphi\circ\beta=\varphi\), and \(\beta(M_I)=M_{g(I)}\) whenever \(I\subset K\) and \(g(I)\) are intervals.
Proof. Here \(B_K\) is the algebra of operators on \((\mathbb C^n)^{\otimes K}\). Let \(V_g\) permute tensor factors of vectors, \(V_g(\bigotimes_\nu\xi_\nu)=\bigotimes_\nu\xi_{g^{-1}(\nu)}\). Then \(V_g^*(\bigotimes_\nu\zeta_\nu)=\bigotimes_\nu\zeta_{g(\nu)}\), and a direct evaluation on elementary tensors gives the stated formula. Since all tensor factors of \(\rho_K=\rho^{\otimes K}\) are equal, \(V_g\rho_KV_g^*=\rho_K\). For an interval \(J\supset K\) and \(w\in B_J\), with \(V=V_g\otimes1\) commuting with \(\rho_J=\rho_K\otimes \rho_{J\setminus K}\), \[ \psi_J(VwV^*)=\operatorname{Tr}_J(V^*\rho_JVw)=\psi_J(w). \] So the normal functionals \(\varphi\circ\beta\) and \(\varphi\) agree on the weakly dense algebra \(\pi(A)\), hence on \(M\). The last statement follows from the formula for \(V_g\). \(\square\)
The next proposition is a zero–one law of Hewitt–Savage type: an element that is fixed by automorphisms moving any finite piece of the tensor product off itself is a scalar.
Proposition 5.4 (asymptotic independence). Let \(z\in M\). Suppose that for every finite interval \(I\) there are a finite interval \(I'\) disjoint from \(I\) and an automorphism \(\beta\) of \(M\) with \(\varphi\circ\beta=\varphi\), \(\beta(z)=z\) and \(\beta(M_I)\subset M_{I'}\). Then \(z\in\mathbb C1\).
Proof. Let \(\varepsilon>0\). By the Kaplansky density theorem there is \(a\in\pi(A)\) with \(\|a\|\le\|z\|\) and \(\|(z-a)\Omega\|<\varepsilon\); say \(a\in M_I\). Take \(I'\) and \(\beta\) for this \(I\) and put \(a'=\beta(a)\in M_{I'}\). For \(w\in M\), \(\|\beta(w)\Omega\|^2=\varphi(\beta(w^*w))=\|w\Omega\|^2\), so \(\|(z-a')\Omega\|=\|(z-a)\Omega\|<\varepsilon\). Now \[ \big|\varphi(z^*z)-\varphi(a^*a')\big|=\big|\langle z\Omega,z\Omega\rangle-\langle a'\Omega,a\Omega\rangle\big| \le|\langle(z-a')\Omega,z\Omega\rangle|+|\langle a'\Omega,(z-a)\Omega\rangle|\le2\varepsilon\|z\|. \] By Lemma 5.1(b) and \(\varphi\circ\beta=\varphi\), \(\varphi(a^*a')=\varphi(a^*)\varphi(\beta(a))=|\varphi(a)|^2\). Also \(|\varphi(a)-\varphi(z)|\le\|(a-z)\Omega\|<\varepsilon\), and \(|\varphi(a)|,|\varphi(z)|\le\|z\|\), so \(\big||\varphi(a)|^2-|\varphi(z)|^2\big|\le2\varepsilon\|z\|\). Hence \(|\varphi(z^*z)-|\varphi(z)|^2|\le4\varepsilon\|z\|\) for every \(\varepsilon\), and \[ \|(z-\varphi(z)1)\Omega\|^2=\varphi(z^*z)-|\varphi(z)|^2=0 . \] Since \(\Omega\) is separating (Lemma 5.1(a)), \(z=\varphi(z)1\). \(\square\)
5.3 The centralizer
Definition 5.5. The centralizer of \(\varphi\) is the set \(M_\varphi\) of all \(x\in M\) such that \(\varphi(xw)=\varphi(wx)\) for all \(w\in M\). For a finite interval \(I\), \(F_I\) denotes the set of those \(x\in B_I\) with \(x\rho_I=\rho_Ix\).
Since \(\operatorname{Tr}_I\) is a faithful trace, \(\psi_I(xy)=\psi_I(yx)\) for all \(y\in B_I\) if and only if \(\operatorname{Tr}_I((\rho_Ix-x\rho_I)y)=0\) for all \(y\), that is, \(x\in F_I\). So \(F_I\) is the centralizer of \(\psi_I\) in \(B_I\).
Lemma 5.6.
- (a) \(M_\varphi\) is a von Neumann algebra on \(\mathcal H\), and \(\tau=\varphi|_{M_\varphi}\) is a faithful normal tracial state on it.
- (b) \(D_I\subset F_I\), \(F_I\subset F_J\) for \(I\subset J\), and \(\pi(F_I)\subset M_\varphi\).
- (c) \(\alpha(M_\varphi)=M_\varphi\), \(\alpha(\pi(F_I))=\pi(F_{I+1})\) and \(\alpha(\pi(D_I))=\pi(D_{I+1})\).
Proof. (a) For fixed \(w\), the maps \(x\mapsto\varphi(xw)=\langle x(w\Omega),\Omega\rangle\) and \(x\mapsto\varphi(wx)=\langle x\Omega,w^*\Omega\rangle\) are weakly continuous on \(B(\mathcal H)\), so \(M_\varphi\) is weakly closed. It contains \(1\). If \(x,x'\in M_\varphi\), then \(\varphi(xx'w)=\varphi(x'wx)=\varphi(wxx')\), and \(\varphi(x^*w)=\overline{\varphi(w^*x)}=\overline{\varphi(xw^*)}=\varphi(wx^*)\). So \(M_\varphi\) is a von Neumann algebra; the state \(\tau\) is tracial by definition, and normal and faithful by Lemma 5.1(a).
(b) The elements of \(D_I\) are linear combinations of tensor products of the \(e_j\), which commute with \(\rho\), so they commute with \(\rho_I\). If \(x\in F_I\), then \(x\otimes1\) commutes with \(\rho_I\otimes \rho_{J\setminus I}=\rho_J\). For \(x\in F_I\), Lemma 5.1(c) with \(c=1\) gives \(\varphi(xw)=\varphi(wx)\) for all \(w\in M\).
(c) For \(x\in M_\varphi\) and \(w\in M\), \(\varphi(\alpha(x)w)=\varphi(x\alpha^{-1}(w))=\varphi(\alpha^{-1}(w)x)=\varphi(w\alpha(x))\), so \(\alpha(M_\varphi)\subset M_\varphi\); the same argument for \(\alpha^{-1}\) gives equality. The shift \(s\) carries \(\rho_I\) to \(\rho_{I+1}\) (both are tensor powers of \(\rho\)), hence \(F_I\) onto \(F_{I+1}\), and \(D_I\) onto \(D_{I+1}\). \(\square\)
Theorem 5.7 (the centralizer is approximately finite-dimensional). Put \(R=M_\varphi\), \(\tau=\varphi|_R\) and \(P_p=\pi(F_{[-p,p]})\). Then \(P_1\subset P_2\subset\cdots\) are finite-dimensional subalgebras of \(R\) and \[ R=\Big(\bigcup_pP_p\Big)''. \] Moreover, if \(E_p\) denotes the \(\tau\)-preserving conditional expectation of \(R\) onto \(P_p\), then \(\|E_p(x)-x\|_2\to0\) for every \(x\in R\).
Proof. The \(P_p\) are increasing subalgebras of \(R\) by Lemma 5.6(b). Let \(K\) be the closure of \(\bigcup_pP_p\Omega\) and \(Q\) the orthogonal projection onto \(K\).
Step 1: vectors orthogonal to \(R\Omega\). Let \(a\in B_{[-p,p]}\) and write \(\rho_{[-p,p]}=\sum_\mu\mu G_\mu\) with spectral projections \(G_\mu\in B_{[-p,p]}\) (all \(\mu>0\)). Then \[ a=\sum_\mu G_\mu aG_\mu+\sum_{\mu\ne\nu}G_\mu aG_\nu . \] The first sum commutes with \(\rho_{[-p,p]}\), so it lies in \(F_{[-p,p]}\). Each \(y=G_\mu aG_\nu\) with \(\mu\ne\nu\) satisfies \(\rho_{[-p,p]}y\rho_{[-p,p]}^{-1}=(\mu/\nu)y\). For such \(y\) and any \(x\in R\), Lemma 5.1(c) with \(w=x^*\) and the centralizer property give \[ \varphi(x^*y)=\varphi(yx^*)=\tfrac{\mu}{\nu}\,\varphi(x^*y), \] so \(\langle y\Omega,x\Omega\rangle=\varphi(x^*y)=0\). Hence every \(a\in\pi(A)\) can be written \(a=f+y\) with \(f\in\bigcup_pP_p\) and \(y\Omega\perp R\Omega\). Since \(K\subset\overline{R\Omega}\), this gives \(Qa\Omega=f\Omega\) and \((1-Q)a\Omega=y\Omega\perp R\Omega\).
Step 2: \(R\Omega\subset K\). Let \(x\in R\). As \(\pi(A)\Omega\) is dense in \(\mathcal H\), there are \(a_m\in\pi(A)\) with \(a_m\Omega\to x\Omega\). By Step 1, \(\langle(1-Q)a_m\Omega,x\Omega\rangle=0\) for all \(m\). In the limit, \(\|(1-Q)x\Omega\|^2=\langle(1-Q)x\Omega,x\Omega\rangle=0\), so \(x\Omega\in K\).
Step 3: \(R=R_0\), where \(R_0=(\bigcup_pP_p)''\). Since \(R\) is weakly closed, \(R_0\subset R\). Let \(E\) be the \(\tau\)-preserving conditional expectation of \(R\) onto \(R_0\) (B2). For \(x\in R\) and \(r\in R_0\), \[ \langle x\Omega-E(x)\Omega,r\Omega\rangle=\tau(r^*x)-\tau(r^*E(x))=\tau(r^*x)-\tau(E(r^*x))=0 . \] So \(E(x)\Omega\) is the orthogonal projection of \(x\Omega\) onto \(\overline{R_0\Omega}\supset K\). By Step 2, \(x\Omega\in K\), hence \(E(x)\Omega=x\Omega\), and \(x=E(x)\in R_0\) because \(\Omega\) is separating.
Step 4: convergence. The same computation, with \(P_p\) in place of \(R_0\), shows that \(E_p(x)\Omega=Q_px\Omega\), where \(Q_p\) is the projection onto the finite-dimensional space \(P_p\Omega\). The \(Q_p\) increase to \(Q\), and \(Qx\Omega=x\Omega\) by Step 2. Hence \(\|E_p(x)-x\|_2=\|Q_px\Omega-x\Omega\|\to0\). \(\square\)
If \(\psi\) is a trace, then \(\rho\) is central in \(B\), each \(\rho_I\) is central in \(B_I\), and \(F_I=B_I\). Theorem 5.7 then gives \(M_\varphi=M\): the product of traces is a trace on the whole tensor product.
Theorem 5.8 (ergodicity). The only elements of \(M\) fixed by \(\alpha\) are the scalars. In particular the restriction \(\theta=\alpha|_R\) is an ergodic automorphism of \(R=M_\varphi\) (its only fixed points in \(R\) are the scalars), and \(\theta\in\operatorname{Aut}(R,\tau)\).
Proof. Let \(\alpha(z)=z\). For \(I=[a,b]\) let \(\beta=\alpha^{b-a+1}\); it preserves \(\varphi\), fixes \(z\), and maps \(M_I\) onto \(M_{I'}\) with \(I'=I+(b-a+1)\) disjoint from \(I\). Proposition 5.4 gives \(z\in\mathbb C1\). By Lemma 5.6(c), \(\alpha\) restricts to an automorphism \(\theta\) of \(R\), and \(\tau\circ\theta=\tau\) because \(\varphi\circ\alpha=\varphi\). \(\square\)
Theorem 5.9 (the Bernoulli algebra is the hyperfinite II₁ factor). Let \(B=M_n\) with \(n\ge2\) and \(\psi\) a faithful state on \(M_n\). Then \(R=M_\varphi\) is isomorphic to the hyperfinite II₁ factor. The algebra \(M\) is a factor as well.
Proof. Factor. Let \(z\) be in the center of \(R\). Given \(I=[a,b]\), let \(I'=[b+1,2b-a+1]\), \(K=I\cup I'\), and let \(g\) be the bijection of \(K\) exchanging \(a+t\) and \(b+1+t\) for \(0\le t\le b-a\). By Lemma 5.3, \(V_g\in F_K\), so \(\pi(V_g)\in R\) by Lemma 5.6(b) and commutes with \(z\). Hence \(\beta=\operatorname{Ad}\pi(V_g)\) fixes \(z\), preserves \(\varphi\) and maps \(M_I\) onto \(M_{I'}\). By Proposition 5.4, \(z\in\mathbb C1\). The same argument applies to the center of \(M\), which also commutes with \(\pi(V_g)\in M\).
Type II₁. \(R\) contains \(\pi(D_I)\), of dimension \(n^{|I|}\) for every \(I\), so it is infinite-dimensional. It is a factor with the faithful normal tracial state \(\tau\), so it is of type II₁ (B5). It acts on the separable space \(\mathcal H\), so it has separable predual (B5).
Semidiscrete. The conditional expectations \(E_p\colon R\to P_p\) are normal unital completely positive maps into finite-dimensional C*-algebras (B2); compose with the inclusions \(P_p\to R\). For \(x\in R\), the sequence \(E_p(x)\) is bounded by \(\|x\|\) and \(\|(E_p(x)-x)\Omega\|\to0\) (Theorem 5.7). For \(r\in M'\subset R'\) and \(\zeta\in\mathcal H\), \[ \langle(E_p(x)-x)r\Omega,\zeta\rangle=\langle r(E_p(x)-x)\Omega,\zeta\rangle\to0 . \] Since \(M'\Omega\) is dense (Lemma 5.1(a)) and the sequence is bounded, \(E_p(x)\to x\) weakly, hence \(\sigma\)-weakly (B4). Thus \(R\) is semidiscrete in the sense of Definition 1.3 of "Uniqueness of the injective II₁ factor".
Conclusion. Semidiscreteness is a property of the von Neumann algebra and not of the representation: the maps \(E_p\) and the inclusions are normal, and \(\sigma\)-weak convergence does not depend on the faithful normal representation (B4). So \(R\), transported to its standard representation on \(L^2(R,\tau)\), is a semidiscrete II₁ factor with separable predual. Theorem 1.5 of "Uniqueness of the injective II₁ factor" states that for such a factor semidiscreteness (condition (c) there) implies that it is isomorphic to the hyperfinite II₁ factor (condition (a) there). \(\square\)
Definition 5.10 (Bernoulli shifts). Let \(n\ge2\), and let \(\psi\) be a faithful state on \(M_n\) whose density has eigenvalues \(\lambda_1,\dots,\lambda_n\) (all positive, with sum \(1\)). The automorphism \(\theta_\psi=\alpha|_R\) of the hyperfinite II₁ factor \(R=M_\varphi\) is the Bernoulli shift with weights \(\lambda_1,\dots,\lambda_n\). For \(\psi=\mathrm{tr}\), the normalized trace, \(R=M=\bigotimes_{\mathbb Z}(M_n,\mathrm{tr})\) and \(\theta_\psi\) is the \(n\)-shift \(S_n\).
Up to conjugacy, \(\theta_\psi\) depends only on the weights. If \(\psi'=\psi\circ\operatorname{Ad}u\) for a unitary \(u\in M_n\), the automorphism of \(A\) that applies \(\operatorname{Ad}u\) at every site carries \(\psi'_\infty\) to \(\psi_\infty\) and commutes with the shift. As in the proof of Lemma 5.2 it is implemented by a unitary between the two GNS spaces, which carries one centralizer onto the other and intertwines the two shifts.
Example 5.11 (edge cases).
- If \(n=1\), then \(M=R=\mathbb C\) and the shift is the identity. This is why \(n\ge2\) is required in Theorem 5.9.
- If \(B=\mathbb C^d\) is abelian with weights \(p_1,\dots,p_d\), then \(\psi\) is a trace, \(R=M\) is abelian, and \(\theta\) is the classical Bernoulli shift with probability vector \((p_1,\dots,p_d)\), written in the language of measure algebras. Theorem 5.8 is the classical ergodicity of Bernoulli shifts.
- If \(B\) is a direct sum of several matrix algebras and \(\psi\) is a trace, \(R=M\) is not a factor in general. For instance, the central projections \(z\otimes1\) with \(z\) central in \(B\) at site \(0\) lie in the center of \(M\).
6. The entropy of Bernoulli shifts
We keep the notation of Section 5: \(B\) finite-dimensional, \(\psi\) faithful, \(R=M_\varphi\), \(\tau=\varphi|_R\), \(\theta=\alpha|_R\). For a finite interval \(I\) and a multi-index \(j=(j_\nu)_{\nu\in I}\in\{1,\dots,d\}^I\), put \(e_j=\bigotimes_{\nu\in I}e_{j_\nu}\in D_I\).
Lemma 6.1. For every finite interval \(I\), \[ H(\pi(F_I))=H(\pi(D_I))=|I|\,S(\psi). \]
Proof. The projections \(e_j\), \(j\in\{1,\dots,d\}^I\), are pairwise orthogonal with sum \(1\). Each is minimal in \(B_I\), since \(e_jB_Ie_j=\bigotimes_\nu e_{j_\nu}Be_{j_\nu}=\mathbb Ce_j\) (B1). They lie in \(D_I\subset F_I\) (Lemma 5.6(b)), so they are minimal projections of \(D_I\) and of \(F_I\). By the product property, \(\tau(\pi(e_j))=\psi_I(e_j)=\prod_{\nu\in I}\lambda_{j_\nu}\). By Fact 1.2(e), both entropies equal \(\sum_j\eta\big(\prod_\nu\lambda_{j_\nu}\big)\). By induction on \(|I|\), this sum is \(|I|S(\psi)\): splitting off one site and using (0.1) and \(\sum_j\lambda_j=1\), \[ \sum_{j,j'}\eta(\lambda_j\mu_{j'})=\sum_{j,j'}\big(\lambda_j\eta(\mu_{j'})+\mu_{j'}\eta(\lambda_j)\big) =\sum_{j'}\eta(\mu_{j'})+\sum_j\eta(\lambda_j), \] where \((\mu_{j'})\) are the products over the remaining sites, which also sum to \(1\). \(\square\)
Theorem 6.2 (entropy of the tensor shift). With the notation above, \[ H(\theta)=S(\psi)=-\operatorname{Tr}(\rho\log \rho). \] More precisely, \(H(\pi(D_{\{0\}}),\theta)=H(P_p,\theta)=S(\psi)\) for every \(p\), where \(P_p=\pi(F_{[-p,p]})\).
Reference: [Connes–Størmer 1975, Section 4].
Proof. Lower bound. By Lemma 5.6(c), \(\theta^i(\pi(D_{\{0\}}))=\pi(D_{\{i\}})\). For \(0\le i\le k-1\) these algebras are abelian and commute pairwise, since they all lie in the abelian algebra \(\pi(D_{[0,k-1]})\). They generate it, because \(D_{[0,k-1]}\) is spanned by the products \(e_j\) of elements taken at single sites. By Fact 1.2(f) and Lemma 6.1, \[ H_k(\pi(D_{\{0\}}),\theta)=H(\pi(D_{[0,k-1]}))=k\,S(\psi), \] so \(H(\pi(D_{\{0\}}),\theta)=S(\psi)\) and \(H(\theta)\ge S(\psi)\).
Upper bound. For \(0\le i\le k-1\), \(\theta^i(P_p)=\pi(F_{[-p+i,p+i]})\subset\pi(F_{[-p,p+k-1]})\) by Lemma 5.6(b),(c). By Fact 1.2(d) with \(r=0\) and Lemma 6.1, \[ H_k(P_p,\theta)\le H(\pi(F_{[-p,p+k-1]}))=(2p+k)\,S(\psi). \] Dividing by \(k\) gives \(H(P_p,\theta)\le S(\psi)\). Since \(D_{\{0\}}\subset F_{[-p,p]}\), Proposition 3.3(b) and the lower bound give equality. By Theorem 5.7 the \(P_p\) increase and their union is weakly dense in \(R\), so Theorem 4.2 gives \(H(\theta)=\lim_pH(P_p,\theta)=S(\psi)\). \(\square\)
Corollary 6.3 (Bernoulli shifts and the \(n\)-shift).
- (a) The Bernoulli shift with weights \(\lambda_1,\dots,\lambda_n\) has entropy \(-\sum_{j=1}^n\lambda_j\log\lambda_j\).
- (b) The \(n\)-shift \(S_n\) of \(\bigotimes_{\mathbb Z}(M_n,\mathrm{tr})\), which is the hyperfinite II₁ factor, has \(H(S_n)=\log n\).
- (c) If \(n\ne m\), there is no isomorphism \(\gamma\colon\bigotimes_{\mathbb Z}(M_n,\mathrm{tr})\to \bigotimes_{\mathbb Z}(M_m,\mathrm{tr})\) with \(\gamma S_n\gamma^{-1}=S_m\). More generally, Bernoulli shifts with different values of \(-\sum_j\lambda_j\log\lambda_j\) are not conjugate.
Proof. (a) is Theorem 6.2 with \(B=M_n\), and Theorem 5.9 identifies \(R\). (b) For \(\psi=\mathrm{tr}\) the density is \(\rho=\frac1n1\), so \(S(\psi)=n\,\eta(1/n)=\log n\). (c) The algebras are II₁ factors, so Corollary 3.4 applies. \(\square\)
Corollary 6.4 (tracial tensor shifts). Let \(N_0\) be a C*-algebra of finite dimension with a faithful tracial state \(\tau_0\), let \((R,\tau)=\bigotimes_{\mathbb Z}(N_0,\tau_0)\) and \(S\) the shift. Then \(S\) is \(\tau\)-preserving and \(H(S)=H(N_0)\), the entropy of the copy of \(N_0\) at site \(0\).
Proof. Since \(\tau_0\) is a trace, \(M_\varphi=M=R\) (remark after Theorem 5.7). By Fact 1.2(e), applied to the minimal projections \(e_1,\dots,e_d\) at site \(0\), \(H(N_0)=\sum_j\eta(\tau_0(e_j))=S(\tau_0)\). Apply Theorem 6.2. \(\square\)
For \(N_0=\mathbb C^d\) this is the classical formula \(h=-\sum_jp_j\log p_j\) for Bernoulli shifts (with Proposition 3.6). For \(N_0=M_n\) it gives again \(H(S_n)=\log n\).
Corollary 6.5 (all entropies occur). For every \(t\in(0,\infty)\) there is a Bernoulli shift of the hyperfinite II₁ factor with entropy \(t\). Hence the hyperfinite II₁ factor has uncountably many pairwise non-conjugate ergodic automorphisms.
Proof. Choose \(n\ge2\) with \(\log n>t\). For \(s\in(0,1/n]\) let \(\lambda(s)=(1-(n-1)s,s,\dots,s)\); all entries are positive because \(1-(n-1)s\ge1/n\). The function \(f(s)=\eta(1-(n-1)s)+(n-1)\eta(s)\) is continuous on \((0,1/n]\), with \(f(1/n)=\log n>t\) and \(f(s)\to\eta(1)=0\) as \(s\to0\). By the intermediate value theorem \(f(s)=t\) for some \(s\), and the Bernoulli shift with weights \(\lambda(s)\) has entropy \(t\) by Corollary 6.3(a). It is ergodic by Theorem 5.8, and different entropies give non-conjugate shifts by Corollary 3.4. \(\square\)
7. Powers of an automorphism
Proposition 7.1. Let \((R,\tau)\) be a finite tracial algebra and \(\theta\in\operatorname{Aut}(R,\tau)\). For every finite-dimensional \(N\) and \(p\ge1\), \(H(N,\theta^p)\le p\,H(N,\theta)\). Consequently \(H(\theta^p)\le|p|\,H(\theta)\) for every \(p\in\mathbb Z\).
Proof. The list \(N,\theta^p(N),\dots,\theta^{(r-1)p}(N)\) is obtained from \(N,\theta(N),\dots,\theta^{rp-1}(N)\) by deleting entries, so by Corollary 2.3 \[ \frac1rH_r(N,\theta^p)\le p\cdot\frac1{rp}H_{rp}(N,\theta). \] Letting \(r\to\infty\) gives \(H(N,\theta^p)\le pH(N,\theta)\), and taking suprema gives \(H(\theta^p)\le pH(\theta)\) for \(p\ge1\). For \(p=0\), \(H(\mathrm{id})=0\) (Example 3.5). For \(p<0\), Proposition 3.3(d) gives \(H(\theta^p)=H(\theta^{-p})\le|p|H(\theta)\). \(\square\)
Theorem 7.2 (powers). Let \((R,\tau)\) be a finite tracial algebra with finite-dimensional subalgebras \(P_1\subset P_2\subset\cdots\) whose union is weakly dense. Then \(H(\theta^p)=|p|\,H(\theta)\) for every \(\theta\in\operatorname{Aut}(R,\tau)\) and \(p\in\mathbb Z\) (with \(0\cdot\infty=0\)).
Proof. By Proposition 7.1 and Proposition 3.3(d), it suffices to prove \(H(\theta^p)\ge pH(\theta)\) for \(p\ge1\). Let \(P_1\subset P_2\subset\cdots\) be the sequence, fix \(k\) and \(\varepsilon>0\). The algebras \(\theta^t(P_k)\), \(0\le t\le p-1\), are finite-dimensional, so by Lemma 4.1(b) there is \(q\) with \[ H(\theta^t(P_k)\mid P_q)<\varepsilon\qquad(0\le t\le p-1). \] Let \(r\ge1\). For \(0\le j\le rp-1\) write \(j=sp+t\) with \(0\le t\le p-1\), and put \(N_j=\theta^j(P_k)\) and \(Q_j=\theta^{sp}(P_q)\). By Fact 1.2(g) and Lemma 2.1 (applied to \(\theta^{sp}\)), \[ H_{rp}(P_k,\theta)=H(N_0,\dots,N_{rp-1})\le H(Q_0,\dots,Q_{rp-1})+\sum_{s=0}^{r-1}\sum_{t=0}^{p-1} H(\theta^t(P_k)\mid P_q). \] The double sum is less than \(rp\varepsilon\). In the list \(Q_0,\dots,Q_{rp-1}\) each algebra \(\theta^{sp}(P_q)\) occurs \(p\) times. Merging the copies (Fact 1.2(d), with symmetry) gives \(H(Q_0,\dots,Q_{rp-1})\le H_r(P_q,\theta^p)\). Therefore \[ \frac1rH_r(P_q,\theta^p)\ge p\cdot\frac1{rp}H_{rp}(P_k,\theta)-p\varepsilon . \] Letting \(r\to\infty\) gives \(H(\theta^p)\ge H(P_q,\theta^p)\ge pH(P_k,\theta)-p\varepsilon\). This holds for every \(k\) and \(\varepsilon\), and \(\sup_kH(P_k,\theta)=H(\theta)\) by Theorem 4.2. Hence \(H(\theta^p)\ge pH(\theta)\). \(\square\)
Classically the lower bound is immediate: the join \(Q=N\vee\theta(N)\vee\dots\vee\theta^{p-1}(N)\) satisfies \(h(Q,\theta^p)=p\,h(N,\theta)\). Noncommutatively \(Q\) need not be finite-dimensional (Example 2.4), and the proof above replaces it by a large \(P_q\) that almost contains \(N,\dots,\theta^{p-1}(N)\). This is where the continuity theorem is needed.
8. Abelian entropy and the role of monotonicity
Another natural candidate for the entropy of \(\theta\) uses only abelian subalgebras on which \(\theta\) acts.
Definition 8.1. The abelian entropy of \(\theta\in\operatorname{Aut}(R,\tau)\) is \[ H_a(\theta)=\sup\{h(\theta|_A\mid A):A\subset R\text{ an abelian subalgebra with }\theta(A)=A\}, \] where \(h(\theta|_A\mid A)\) is the Kolmogorov–Sinai entropy of Section 3. (The set is nonempty since \(A=\mathbb C1\) qualifies, with entropy \(0\).)
Proposition 8.2.
- (a) \(H_a(\theta)\le H(\theta)\) for every \(\theta\in\operatorname{Aut}(R,\tau)\).
- (b) For the tensor shifts of Section 6 (in particular for the Bernoulli shifts), \(H_a(\theta)=H(\theta)=S(\psi)\).
- (c) If \(A\) is abelian and \(\theta(A)=A\), then \(h(\theta^p|_A\mid A)=|p|\,h(\theta|_A\mid A)\) for every \(p\in\mathbb Z\). Consequently \(H_a(\theta^p)\ge|p|\,H_a(\theta)\).
Proof. (a) is Proposition 3.6.
(b) Let \(A=\big(\bigcup_I\pi(D_I)\big)''\), an abelian subalgebra of \(R\) (each \(\pi(D_I)\subset R\), and the \(D_I\) increase and commute). It is \(\theta\)-invariant by Lemma 5.6(c). By Proposition 3.6 and Theorem 6.2, \(h(\theta|_A\mid A)\ge H(\pi(D_{\{0\}}),\theta)=S(\psi)=H(\theta)\). Combine with (a).
(c) Let \(p\ge1\); \(A\) is also \(\theta^p\)-invariant. Proposition 3.6 (for \(\theta\) and for \(\theta^p\)) and Proposition 7.1 give \[ h(\theta^p|_A\mid A)=\sup_{N\subset A}H(N,\theta^p)\le p\sup_{N\subset A}H(N,\theta)=p\,h(\theta|_A\mid A). \] Conversely, for finite-dimensional \(N\subset A\) let \(Q=N\vee\theta(N)\vee\dots\vee\theta^{p-1}(N)\), which is finite-dimensional and contained in \(A\). The algebras \(\theta^{sp}(Q)\), \(0\le s\le r-1\), generate \(N\vee\dots\vee\theta^{rp-1}(N)\), so by Proposition 3.6 \[ H_r(Q,\theta^p)=h\big(N\vee\dots\vee\theta^{rp-1}(N)\big)=H_{rp}(N,\theta). \] Dividing by \(r\) and letting \(r\to\infty\) gives \(H(Q,\theta^p)=pH(N,\theta)\), hence \(h(\theta^p|_A\mid A)\ge p\,h(\theta|_A\mid A)\). For \(p<0\) use Proposition 3.3(d), and for \(p=0\) Example 3.5. Finally, a \(\theta\)-invariant abelian algebra is \(\theta^p\)-invariant, so \(H_a(\theta^p)\ge\sup_Ah(\theta^p|_A\mid A)=|p|H_a(\theta)\), the supremum over \(\theta\)-invariant \(A\). \(\square\)
Remark 8.3 (why \(H\) and not \(H_a\)). Two features make the abelian entropy unsatisfactory as a general invariant.
- It depends on the existence of large \(\theta\)-invariant abelian subalgebras. For a general \(\theta\) it is not clear that there are any beyond \(\mathbb C1\), and then \(H_a(\theta)=0\) says nothing.
- Its behaviour under powers is not controlled from above. An abelian algebra invariant under \(\theta^p\) need not be invariant under \(\theta\). The algebra generated by \(A,\theta(A),\dots,\theta^{p-1}(A)\) is \(\theta\)-invariant but in general not abelian. So the argument of Proposition 8.2(c) gives only \(H_a(\theta^p)\ge|p|H_a(\theta)\).
By contrast, \(H\) needs no invariant subalgebras, and \(H(\theta^p)=|p|H(\theta)\) whenever the Kolmogorov–Sinai theorem applies (Theorem 7.2).
Remark 8.4 (the role of monotonicity). The proof of Theorem 4.2 uses two properties of \(N\mapsto H(N,\theta)\): it is increasing (Proposition 3.3(b)), and \(H(N,\theta)\le H(P,\theta)+H(N\mid P)\) with \(H(N\mid P)\) small when \(N\) is almost contained in \(P\) (Proposition 3.3(c) and Theorem 1.3). Monotonicity is what lets the supremum over all \(N\) be computed along one increasing sequence; without it, the values \(H(P_q,\theta)\) need not approach the supremum. Other entropies of automorphisms of von Neumann algebras were proposed at the same time, for instance by Emch. For that definition the quantity playing the role of \(H(N,\theta)\) is not increasing in \(N\) [Connes–Størmer 1975], so the argument of Section 4 does not apply to it, and its values on shifts are not computed by the methods of this lesson.
9. Exercises
Exercise 1. Let \((R,\tau)\) be a finite tracial algebra and \(P_1\subset P_2\subset\cdots\) finite-dimensional subalgebras whose union is weakly dense. Let \(u\in\bigcup_qP_q\) be a unitary. Show that \(H(\operatorname{Ad}u)=0\).
Solution. Say \(u\in P_{q_0}\). For \(q\ge q_0\), \(u\in P_q\), so \(\operatorname{Ad}u\) maps \(P_q\) onto itself and all algebras in the list \(P_q,uP_qu^*,\dots\) equal \(P_q\). By Fact 1.2(d), \(H_k(P_q,\operatorname{Ad}u)\le H(P_q)\) for all \(k\), so \(H(P_q,\operatorname{Ad}u)=0\). Also \(\operatorname{Ad}u\) preserves \(\tau\). By Theorem 4.2, \(H(\operatorname{Ad}u)=\lim_qH(P_q,\operatorname{Ad}u)=0\). (Note that for \(q<q_0\) the algebras \(u^jP_qu^{-j}\) can all be different; the Kolmogorov–Sinai theorem lets us ignore them.)
Exercise 2. Let \(N_0=M_2\oplus\mathbb C\) with the faithful tracial state \(\tau_0(x\oplus c)=a\,\mathrm{tr}(x)+(1-a)c\), \(0<a<1\), where \(\mathrm{tr}\) is the normalized trace of \(M_2\). Compute the entropy of the shift on \(\bigotimes_{\mathbb Z}(N_0,\tau_0)\) and find its maximum over \(a\).
Solution. A maximal family of minimal projections is \(e_{11}\oplus0\), \(e_{22}\oplus0\), \(0\oplus1\), with traces \(a/2,a/2,1-a\). By Corollary 6.4, \[ H(S)=2\eta(a/2)+\eta(1-a)=a\log2+\eta(a)+\eta(1-a), \] using \(2\eta(a/2)=a\log2-a\log a\). The derivative in \(a\) is \(\log2-\log a+\log(1-a)=\log\frac{2(1-a)}a\), which vanishes at \(a=2/3\). There the value is \(\frac23\log2+\frac23\log\frac32+\frac13\log3=\log3\), the maximum. This agrees with (1.3) at the level of one site: three orthogonal minimal projections give at most \(\log3\), with equality when all three have trace \(1/3\).
Exercise 3. Let \(\theta_\psi\) be the Bernoulli shift of \(M_n\) with state \(\psi\), and \(p\ge1\). Show that \(\theta_\psi^p\) is the Bernoulli shift of \(M_{n^p}=M_n^{\otimes p}\) with the state \(\psi^{\otimes p}\), and check that Theorem 6.2 and Theorem 7.2 give the same value for \(H(\theta_\psi^p)\).
Solution. Group the sites into blocks \(\{mp,\dots,mp+p-1\}\), \(m\in\mathbb Z\). The algebra \(A=\bigcup_IB_I\) is also the union of the algebras built on unions of consecutive blocks. The state \(\psi_\infty\) is the product of the states \(\psi^{\otimes p}\) on the blocks. So the GNS space, \(M\), \(\varphi\) and \(M_\varphi=R\) are the same for \((M_n,\psi)\) and for \((M_n^{\otimes p},\psi^{\otimes p})\) with blocks as sites. The block shift is \(\alpha^p\), so the Bernoulli shift of \((M_{n^p},\psi^{\otimes p})\) is \(\theta_\psi^p\). By Theorem 6.2, \(H(\theta_\psi^p)=S(\psi^{\otimes p})\). The density of \(\psi^{\otimes p}\) is \(\rho^{\otimes p}\), with eigenvalues the products \(\lambda_{j_1}\cdots\lambda_{j_p}\), and the computation in Lemma 6.1 gives \(S(\psi^{\otimes p})=pS(\psi)\). Theorem 7.2 gives \(H(\theta_\psi^p)=pH(\theta_\psi)=pS(\psi)\) as well.
Exercise 4. Let \((R,\tau)\) be a finite tracial algebra and \(\theta\in\operatorname{Aut}(R,\tau)\) with \(\theta^p=\mathrm{id}\) for some \(p\ge1\). Explain why Proposition 7.1 does not give \(H(\theta)=0\), why Theorem 7.2 does when \(R\) has an approximating sequence, and why Example 3.5(ii) needs no such sequence.
Solution. Example 3.5(ii) gives \(H(\theta)=0\) directly. Proposition 7.1 gives \(0=H(\mathrm{id})=H(\theta^p)\le pH(\theta)\), which carries no information: the inequality goes the wrong way. If \(R\) has an approximating sequence, Theorem 7.2 gives \(pH(\theta)=H(\theta^p)=0\), another proof. Example 3.5(ii) holds for every finite tracial algebra, because merging needs no approximation.
Exercise 5. In Example 2.4, show that \(H(N_1,N_2)\ge H(N_1)\), and compute \(H(N_1)\).
Solution. By Lemma 2.2, \(H(N_1,N_2)\ge H(N_1)\). The minimal projections of \(N_1\) are \(p\) and \(1-p\), each of trace \(1/2\), so \(H(N_1)=2\eta(1/2)=\log2\) by Fact 1.2(e). Together with Example 2.4, \(\log2\le H(N_1,N_2)\le2\log2\).
References
- [Connes–Størmer 1975] A. Connes and E. Størmer, Entropy for automorphisms of II₁ von Neumann algebras, Acta Math. 134 (1975), no. 3–4, 289–306. Free at https://doi.org/10.1007/BF02392105
- [Connes–Narnhofer–Thirring 1987] A. Connes, H. Narnhofer and W. Thirring, Dynamical entropy of C* algebras and von Neumann algebras, Comm. Math. Phys. 112 (1987), no. 4, 691–719. Free at https://alainconnes.org/wp-content/uploads/entropy.pdf
- [Connes 1994] A. Connes, Noncommutative geometry, Academic Press, San Diego, CA, 1994. Free at https://alainconnes.org/wp-content/uploads/book94bigpdf.pdf
- [Fremlin] D. H. Fremlin, Measure Theory, Volume 3: Measure Algebras, Torres Fremlin. Free at https://www1.essex.ac.uk/maths/people/fremlin/mt.htm