arXiv:2607.19305v1 [cs.LG] 21 Jul 2026
Doctoral School in Information and Communication Technology
Riemannian Deep Learning: Modules, Networks, and Geometries Ziheng Chen
Advisor Prof. Nicu Sebe Università di Trento ELLIS Co-supervisor Prof. Bernhard Schölkopf Max Planck Institute for Intelligent Systems
September 2026
Publications († corresponding author, ‡ equal contribution, ♣ equal supervision)
The thesis is based on the following publications:
• Chapter 3: [1] Ziheng Chen, Yue Song, Yunmei Liu, and Nicu Sebe. “A Lie Group Approach to Riemannian Batch Normalization.” ICLR 2024. [2] Ziheng Chen, Yue Song, Xiao-Jun Wu, and Nicu Sebe. “Gyrogroup Batch Normalization.” ICLR 2025. • Chapter 4: [3] Ziheng Chen, Yue Song, Gaowen Liu, Ramana Rao Kompella, Xiao-Jun Wu, and Nicu Sebe. “Riemannian Multinomial Logistics Regression for SPD Neural Networks.” CVPR 2024. [4] Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, and Nicu Sebe. “RMLR: Extending Multinomial Logistic Regression into General Geometries.” NeurIPS 2024. • Chapter 5: [5] Ziheng Chen‡ , Zihan Su‡ , Bernhard Schölkopf, and Nicu Sebe. “Proper Velocity Neural Networks.” ICLR 2026. [6] Ziheng Chen, Bernhard Schölkopf, and Nicu Sebe. “Hyperbolic Busemann Neural Networks.” CVPR 2026. [7] Ziheng Chen, Xiao-Jun Wu, Bernhard Schölkopf, and Nicu Sebe. “Riemannian Networks over Full-Rank Correlation Matrices.” ICML 2026. • Chapter 6: [8] Ziheng Chen, Yue Song, Tianyang Xu, Zhiwu Huang, Xiao-Jun Wu, and Nicu Sebe. “Adaptive Log-Euclidean Metrics for SPD Matrix Learning.” IEEE TIP 2024. [9] Ziheng Chen, Yue Song, Xiao-Jun Wu, and Nicu Sebe. “Fast and Stable Riemannian Metrics on SPD Manifolds via Cholesky Product Geometry.” ICLR 2026. The following papers are published but are not included in this thesis: (10) Ziheng Chen, Yue Song, Xiao-Jun Wu, Gaowen Liu, and Nicu Sebe. “Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian Geometry.” ICLR 2025. (11) Shaocheng Jin, Tao Zhou, Rui Wang, Ziheng Chen, Xiaoqing Luo, Xiao-Jun Wu, and Josef Kittler. “Towards Robust EEG Decoding Based on Riemannian SelfAttention.” KDD 2026. (12) Rui Wang, Zihao Bi, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen† . “Riemannian Graph Convolutional Network for Skeleton-Based Two-Person Interaction Recognition.” IJCAI 2026. i
(13) Xianglong Shi‡ , Ziheng Chen†,‡ , Yunhan Jiang, and Nicu Sebe. “Intrinsic Lorentz Neural Network.” ICLR 2026. (14) Shanglin Li, Shiwen Chu, Okan Koç, Yi Ding, Qibin Zhao, Motoaki Kawanabe, and Ziheng Chen† . “HEEGNet: Hyperbolic Embeddings for EEG.” ICLR 2026. (15) Chen Hu‡ , Ziheng Chen‡ , Rui Wang, Yefeng Zheng, and Nicu Sebe. “Riemannian High-Order Pooling for Brain Foundation Models.” ICLR 2026. (16) Rui Wang, Yuting Jiang, Xiaoqing Luo, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen† . “Wasserstein-Aligned Hyperbolic Multi-View Clustering.” AAAI 2026 (Oral). (17) Rui Wang, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen† . “Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds.” NeurIPS 2025. (18) Rui Wang, Shaocheng Jin, Zhenyu Cai, Ziheng Chen† , Xiao-Jun Wu† , and Josef Kittler. “Learning a Better SPD Network for Signal Classification: A Riemannian Batch Normalization Method.” IEEE TNNLS 2025. (19) Chen Hu, Rui Wang♣ , Xiaoning Song, Tao Zhou, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen♣ . “A Correlation Manifold Self-Attention Network for EEG Decoding.” IJCAI 2025. (20) Rui Wang, Shaocheng Jin, Ziheng Chen† , Xiaoqing Luo, and Xiao-Jun Wu. “Learning to Normalize on the SPD Manifold under Bures-Wasserstein Geometry.” CVPR 2025. (21) Rui Wang, Jiayao Jin, Ziheng Chen† , Cong Wu† , Xiao-Jun Wu, and Nicu Sebe. “Structural Topology Refinement Network for Skeleton-Based Action Recognition.” IEEE TIM 2025. (22) Rui Wang, Chen Hu, Ziheng Chen† , Xiao-Jun Wu† , and Xiaoning Song. “A Grassmannian Manifold Self-Attention Network for Signal Classification.” IJCAI 2024. (23) Rui Wang, Xiao-Jun Wu, Ziheng Chen, Cong Hu, and Josef Kittler. “SPD Manifold Deep Metric Learning for Image Set Classification.” IEEE TNNLS 2024.
ii
Abstract
Recently, deep neural networks operating on manifold-valued representations have garnered significant attention across various machine learning applications. However, many basic neural components remain tied to particular manifolds or rely on Euclidean approximations, while the underlying geometry can make repeated computations costly or numerically unstable. This thesis addresses these limitations by developing reusable modules, exploiting manifold-specific structures when general constructions are insufficient, and introducing fast and stable geometries. We first generalize batch normalization and classification beyond individual manifolds. For normalization, we develop a framework on Lie groups, with theoretical control over Riemannian sample means and variances. To extend this principle beyond Lie groups, we introduce pseudo-reductive gyrogroups, which generalize classical gyrogroups and Lie groups, and build a normalization framework on this structure for a wider range of manifolds. For classification, we first extend Euclidean Multinomial Logistic Regression (MLR), which consists of a fully connected layer followed by softmax, to Symmetric Positive Definite (SPD) manifolds with flat metrics. We then use Riemannian trigonometry to extend MLR to general Riemannian manifolds. When a general formulation cannot exploit useful structure, we design networks for particular representations. For stable hyperbolic deep learning, we use Proper Velocity, an unconstrained representation of hyperbolic space, and develop its geometry and core neural layers. We also use Busemann functions to build intrinsic and efficient hyperbolic classification and fully connected layers. For full-rank correlation matrices, a normalized alternative to SPD matrices, we construct networks that operate directly on the manifold and derive accurate gradients for end-to-end training under two correlation geometries. Finally, we study how the geometry itself can improve learning on SPD manifolds. To move beyond fixed metrics, we make the metric learnable through parameterized matrix logarithms, allowing it to adapt to data and network dynamics with little additional computation. To improve efficiency and numerical stability, we exploit the product structure of Cholesky factors to construct SPD metrics with fast and stable closed-form operators. The proposed methods are supported by theoretical analysis and validated through numerical experiments and empirical applications in vision, signal processing, graph learning, and genomics. iii
Keywords Riemannian deep learning, geometric deep learning, Riemannian manifolds, matrix manifolds, Lie groups, constant-curvature manifolds
iv
Acknowledgements My Ph.D. journey would not have been possible without the guidance, collaboration, and encouragement of many people. First and foremost, I would like to express my deepest gratitude to my advisor, Prof. Nicu Sebe. He gave me the freedom to pursue the questions that genuinely interested me and created an open research environment in which I could explore new directions with confidence. At the same time, he was always generous with his time and consistently offered encouragement and guidance in our discussions. His support extended well beyond research. He advised me on career decisions, helped me engage with the research community, and supported my scholarship applications. This balance between intellectual freedom and dependable support has profoundly shaped both the work presented in this thesis and the researcher I have become. I am also deeply grateful to Prof. Bernhard Schölkopf. During my research stay with him, his generosity and intellectual openness allowed me to explore distributional geometry in depth and to develop new perspectives on its role in machine learning. I value his insight, trust, and support. I have greatly enjoyed discussing research questions with him and feel fortunate that our collaboration will continue during my postdoctoral research. I would also like to thank Prof. Xiaojun Wu, who was my master’s advisor. We continued to collaborate throughout my Ph.D., and I deeply appreciate the generous help, guidance, and encouragement he provided along the way. My sincere thanks also go to my collaborators Yue Song, Rui Wang, and Shanglin Li, as well as to the students I mentored: Zihan Su, Xianglong Shi, Youxing Li, Ying Zhang, Chen Hu, Shaocheng Jin, Zihao Bi, and Yuting Jiang. Our discussions and joint efforts have been an essential part of my Ph.D. experience. I have learned a great deal from each of them, and I am grateful for the ideas, dedication, and enthusiasm they brought to our work together. Finally, I would like to thank my family. My deepest and most personal thanks go to my girlfriend, Yunmei Liu. Throughout this journey, she has supported me with patience and warmth and has always believed in me. She has made the difficult moments easier and the joyful moments more meaningful. She has also brought more good fortune into my life than I could ever have imagined. I am profoundly grateful to have her by my side.
v
vi
Contents Notation and Conventions
xxi
1 Introduction 1.1 Riemannian Deep Learning . . . . . . . . . . . . . . . . . . . . . . . . . 1.2 Contributions and Outlines . . . . . . . . . . . . . . . . . . . . . . . . 1.3 Summary of Papers Excluded from the Thesis . . . . . . . . . . . . . .
1 1 3 5
2 Mathematical Background 2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.2 Topology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3 Differential Geometry . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.4 Riemannian Geometry . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.5 Metric Geometry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.6 Algebraic Structures on Manifolds . . . . . . . . . . . . . . . . . . . . . 2.7 Riemannian Optimization . . . . . . . . . . . . . . . . . . . . . . . . . 2.8 Matrix Functions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.8.1 Matrix Functions and Differentials . . . . . . . . . . . . . . . . 2.8.2 Backpropagation Through Matrix Functions . . . . . . . . . . . 2.9 Example Manifolds . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.9.1 Symmetric Positive Definite Manifolds . . . . . . . . . . . . . . 2.9.2 Full-Rank Correlation Manifolds . . . . . . . . . . . . . . . . . . 2.9.3 Grassmannian Manifolds . . . . . . . . . . . . . . . . . . . . . . 2.9.4 Special Orthogonal Groups . . . . . . . . . . . . . . . . . . . . . 2.9.5 Constant-Curvature Manifolds . . . . . . . . . . . . . . . . . . .
9 9 9 12 17 25 29 34 35 35 36 37 37 39 48 50 50
3 Riemannian Batch Normalization 3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2 Lie Group Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 3.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.3 Revisiting Normalization . . . . . . . . . . . . . . . . . . . . . . 3.2.4 LieBN . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.5 Manifestations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.6 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3 Gyrogroup Batch Normalization . . . . . . . . . . . . . . . . . . . . . .
55 55 56 56 58 59 61 65 71 76
vii
Contents
3.4
3.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.2 Pseudo-Reductive Gyrogroups . . . . . . . . . . . . . . . . . . . 3.3.3 GyroBN on Pseudo-Reductive Gyrogroups . . . . . . . . . . . . 3.3.4 Instantiations . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
76 79 84 86 99 107
4 Riemannian Multinomial Logistic Regression 109 4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 4.2 Multinomial Logistic Regression on SPD Manifolds . . . . . . . . . . . 110 4.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 4.2.2 SPD Multinomial Logistic Regression . . . . . . . . . . . . . . . 111 4.2.3 SPD MLRs under Deformed LEM and LCM . . . . . . . . . . . 115 4.2.4 Rethinking the Existing LogEig Classifier . . . . . . . . . . . . . 116 4.3 Extension to General Riemannian Manifolds . . . . . . . . . . . . . . . 117 4.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 4.3.2 Riemannian Multinomial Logistic Regression . . . . . . . . . . . 119 4.3.3 SPD Multinomial Logistic Regressions . . . . . . . . . . . . . . 123 4.3.4 Lie Multinomial Logistic Regression . . . . . . . . . . . . . . . . 126 4.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 4.4.1 Experiments on the Proposed SPD MLRs . . . . . . . . . . . . 128 4.4.2 Experiments on the Proposed Lie MLR . . . . . . . . . . . . . . 131 4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 5 Riemannian Neural Networks 133 5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133 5.2 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . . . . . 134 5.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134 5.2.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 5.2.3 Proper Velocity Geometry . . . . . . . . . . . . . . . . . . . . . 136 5.2.4 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . 139 5.2.5 Connections to the Hyperboloid . . . . . . . . . . . . . . . . . . 142 5.2.6 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 144 5.3 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . . . . . 149 5.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 5.3.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151 5.3.3 Busemann Multinomial Logistic Regression . . . . . . . . . . . . 152 5.3.4 Busemann Fully Connected Layer . . . . . . . . . . . . . . . . . 156 5.3.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159 5.4 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . . . . . 163 5.4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163 5.4.2 Log-Euclidean Correlation Layers . . . . . . . . . . . . . . . . . 165 5.4.3 Poly-Hyperbolic-Cholesky Layers . . . . . . . . . . . . . . . . . 169 5.4.4 Backpropagation over Correlation Geometries . . . . . . . . . . 171 viii
Contents
5.5
5.4.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
172 183
6 Fast and Stable Geometries on SPD Manifolds 185 6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 185 6.2 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . . . . . 186 6.2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 186 6.2.2 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . 187 6.2.3 Parameter Learning . . . . . . . . . . . . . . . . . . . . . . . . . 195 6.2.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 197 6.3 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . . . . . . 204 6.3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 204 6.3.2 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . 205 6.3.3 Product Geometries on the Cholesky . . . . . . . . . . . . . . . 206 6.3.4 Geometries on the SPD Manifold . . . . . . . . . . . . . . . . . 211 6.3.5 Applications to SPD Neural Networks . . . . . . . . . . . . . . 213 6.3.6 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 213 6.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220 7 Conclusion and Future Work 221 7.1 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 221 7.2 Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 222 Bibliography
223
A Experimental Details and Additional Discussions 245 A.1 Data Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 245 A.1.1 Skeleton-Based Action Recognition and Gesture Data Sets . . . 245 A.1.2 Radar and EEG Signal Data Sets . . . . . . . . . . . . . . . . . 246 A.1.3 Image Classification Data Sets . . . . . . . . . . . . . . . . . . . 246 A.1.4 Graph Data Sets . . . . . . . . . . . . . . . . . . . . . . . . . . 246 A.1.5 Genomic Sequence Data Sets . . . . . . . . . . . . . . . . . . . 247 A.2 Backbone Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 248 A.2.1 Backbone Networks on the SPD Manifold . . . . . . . . . . . . 248 A.2.2 Backbone Networks on the Grassmannian . . . . . . . . . . . . 249 A.2.3 Backbone Networks on Rotation Matrices . . . . . . . . . . . . 250 A.2.4 Backbone Networks on Hyperbolic Spaces . . . . . . . . . . . . 250 A.2.5 Riemannian Residual Network Backbones . . . . . . . . . . . . 252 A.3 Experimental Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253 A.3.1 Lie Group Batch Normalization . . . . . . . . . . . . . . . . . . 253 A.3.2 Gyrogroup Batch Normalization . . . . . . . . . . . . . . . . . . 257 A.3.3 Riemannian Multinomial Logistic Regression . . . . . . . . . . . 258 A.3.4 Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . 261 A.3.5 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . 263 A.3.6 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . 266 ix
Contents A.3.7 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . A.3.8 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . . A.4 Additional Discussions . . . . . . . . . . . . . . . . . . . . . . . . . . . A.4.1 Riemannian Multinomial Logistic Regression . . . . . . . . . . . A.4.2 Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . A.4.3 Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . A.4.4 Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . A.4.5 Product Cholesky Metrics . . . . . . . . . . . . . . . . . . . . .
269 271 273 273 282 286 289 290
B Proofs 291 B.1 Mathematical Background . . . . . . . . . . . . . . . . . . . . . . . . . 291 B.1.1 Proof of Thm. 57 . . . . . . . . . . . . . . . . . . . . . . . . . . 291 B.2 Lie Group Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 291 B.2.1 Proof of Thm. 61 . . . . . . . . . . . . . . . . . . . . . . . . . . 291 B.2.2 Proof of Thm. 62 . . . . . . . . . . . . . . . . . . . . . . . . . . 292 B.2.3 Proof of Thm. 65 . . . . . . . . . . . . . . . . . . . . . . . . . . 292 B.2.4 Proof of Thm. 66 . . . . . . . . . . . . . . . . . . . . . . . . . . 293 B.2.5 Proof of Thm. 67 . . . . . . . . . . . . . . . . . . . . . . . . . . 293 B.2.6 Proof of Thm. 68 . . . . . . . . . . . . . . . . . . . . . . . . . . 294 B.2.7 Proof of Thm. 69 . . . . . . . . . . . . . . . . . . . . . . . . . . 294 B.2.8 Proof of Thm. 70 . . . . . . . . . . . . . . . . . . . . . . . . . . 296 B.2.9 Proof of Thm. 72 . . . . . . . . . . . . . . . . . . . . . . . . . . 297 B.3 Gyrogroup Batch Normalization . . . . . . . . . . . . . . . . . . . . . . 298 B.3.1 Proof of Thm. 74 . . . . . . . . . . . . . . . . . . . . . . . . . . 298 B.3.2 Proof of Thm. 75 . . . . . . . . . . . . . . . . . . . . . . . . . . 299 B.3.3 Proof of Thm. 77 . . . . . . . . . . . . . . . . . . . . . . . . . . 300 B.3.4 Proof of Thm. 78 . . . . . . . . . . . . . . . . . . . . . . . . . . 300 B.3.5 Proof of Thm. 79 . . . . . . . . . . . . . . . . . . . . . . . . . . 302 B.3.6 Proof of Thm. 80 . . . . . . . . . . . . . . . . . . . . . . . . . . 303 B.3.7 Proof of Thm. 81 . . . . . . . . . . . . . . . . . . . . . . . . . . 304 B.3.8 Proof of Thm. 83 . . . . . . . . . . . . . . . . . . . . . . . . . . 305 B.3.9 Proof of Thm. 86 . . . . . . . . . . . . . . . . . . . . . . . . . . 306 B.3.10 Proof of Thm. 88 . . . . . . . . . . . . . . . . . . . . . . . . . . 309 B.3.11 Proof of Thm. 89 . . . . . . . . . . . . . . . . . . . . . . . . . . 309 B.3.12 Proof of Thm. 90 . . . . . . . . . . . . . . . . . . . . . . . . . . 312 B.3.13 Proof of Thm. 91 . . . . . . . . . . . . . . . . . . . . . . . . . . 312 B.3.14 Proof of Thm. 92 . . . . . . . . . . . . . . . . . . . . . . . . . . 314 B.3.15 Proof of Thm. 93 . . . . . . . . . . . . . . . . . . . . . . . . . . 318 B.3.16 Proof of Thm. 95 . . . . . . . . . . . . . . . . . . . . . . . . . . 318 B.3.17 Proof of Thm. 97 . . . . . . . . . . . . . . . . . . . . . . . . . . 318 B.3.18 Proof of Thm. 98 . . . . . . . . . . . . . . . . . . . . . . . . . . 321 B.3.19 Proof of Thm. 99 . . . . . . . . . . . . . . . . . . . . . . . . . . 322 B.4 SPD Multinomial Logistic Regression . . . . . . . . . . . . . . . . . . . 326 B.4.1 Proof of Thm. 104 . . . . . . . . . . . . . . . . . . . . . . . . . 326 x
Contents
B.5
B.6
B.7
B.8
B.9
B.4.2 Proof of Thm. 105 . . . . . . . . . . . . . . . . . . . . . . . . . B.4.3 Proof of Thm. 106 . . . . . . . . . . . . . . . . . . . . . . . . . B.4.4 Proof of Thm. 107 . . . . . . . . . . . . . . . . . . . . . . . . . B.4.5 Proof of Thm. 108 . . . . . . . . . . . . . . . . . . . . . . . . . B.4.6 Proof of Thm. 109 . . . . . . . . . . . . . . . . . . . . . . . . . B.4.7 Proof of Thm. 111 . . . . . . . . . . . . . . . . . . . . . . . . . Riemannian Multinomial Logistic Regression . . . . . . . . . . . . . . . B.5.1 Proof of Thm. 113 . . . . . . . . . . . . . . . . . . . . . . . . . B.5.2 Proof of Thm. 114 . . . . . . . . . . . . . . . . . . . . . . . . . B.5.3 Proof of Thm. 116 . . . . . . . . . . . . . . . . . . . . . . . . . B.5.4 Proof of Thm. 117 . . . . . . . . . . . . . . . . . . . . . . . . . B.5.5 Proof of Thm. 119 . . . . . . . . . . . . . . . . . . . . . . . . . B.5.6 Proof of Thm. 120 . . . . . . . . . . . . . . . . . . . . . . . . . Proper Velocity Neural Networks . . . . . . . . . . . . . . . . . . . . . B.6.1 Derivation of the Proper Velocity Metric . . . . . . . . . . . . . B.6.2 Proof of Thm. 121 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.3 Proof of Thm. 122 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.4 Proof of Thm. 123 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.5 Proof of Thm. 124 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.6 Proof of Thm. 125 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.7 Proof of Thm. 126 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.8 Proof of Thm. 127 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.9 Proof of Thm. 128 . . . . . . . . . . . . . . . . . . . . . . . . . B.6.10 Proof of Thm. 129 . . . . . . . . . . . . . . . . . . . . . . . . . Hyperbolic Busemann Neural Networks . . . . . . . . . . . . . . . . . . B.7.1 Proof of Thm. 130 . . . . . . . . . . . . . . . . . . . . . . . . . B.7.2 Proof of Thm. 132 . . . . . . . . . . . . . . . . . . . . . . . . . B.7.3 Proof of Thm. 135 . . . . . . . . . . . . . . . . . . . . . . . . . B.7.4 Proof of Thm. 136 . . . . . . . . . . . . . . . . . . . . . . . . . B.7.5 Proof of Thm. 137 . . . . . . . . . . . . . . . . . . . . . . . . . Full-Rank Correlation Networks . . . . . . . . . . . . . . . . . . . . . . B.8.1 Proof of Thm. 138 . . . . . . . . . . . . . . . . . . . . . . . . . B.8.2 Proof of Thm. 139 . . . . . . . . . . . . . . . . . . . . . . . . . B.8.3 Proof of Thm. 142 . . . . . . . . . . . . . . . . . . . . . . . . . B.8.4 Proof of Thm. 143 . . . . . . . . . . . . . . . . . . . . . . . . . B.8.5 Proof of Thm. 145 . . . . . . . . . . . . . . . . . . . . . . . . . B.8.6 Proof of Thm. 146 . . . . . . . . . . . . . . . . . . . . . . . . . B.8.7 Proof of Thm. 144 . . . . . . . . . . . . . . . . . . . . . . . . . Adaptive Log-Euclidean Metrics . . . . . . . . . . . . . . . . . . . . . . B.9.1 Proof of Thm. 147 . . . . . . . . . . . . . . . . . . . . . . . . . B.9.2 Proof of Thm. 148 . . . . . . . . . . . . . . . . . . . . . . . . . B.9.3 Proof of Thm. 149 . . . . . . . . . . . . . . . . . . . . . . . . . B.9.4 Proof of Thm. 150 . . . . . . . . . . . . . . . . . . . . . . . . . B.9.5 Proof of Thm. 152 . . . . . . . . . . . . . . . . . . . . . . . . . xi
327 327 327 328 328 329 332 332 333 333 333 338 338 338 338 339 341 342 350 350 354 359 359 360 363 363 364 366 367 368 370 370 372 373 376 377 379 380 380 380 381 381 381 382
Contents B.9.6 Proof of Thm. 154 B.9.7 Proof of Thm. 155 B.9.8 Proof of Thm. 156 B.9.9 Proof of Thm. 157 B.9.10 Proof of Thm. 158 B.9.11 Proof of Thm. 159 B.9.12 Proof of Thm. 160 B.9.13 Proof of Thm. 161 B.9.14 Proof of Thm. 163 B.9.15 Proof of Thm. 164 B.9.16 Proof of Thm. 167 B.9.17 Proof of Thm. 165 B.9.18 Proof of Thm. 166 B.10 Product Cholesky Metrics B.10.1 Proof of Thm. 169 B.10.2 Proof of Thm. 170 B.10.3 Proof of Thm. 172 B.10.4 Proof of Thm. 173 B.10.5 Proof of Thm. 174 B.10.6 Proof of Thm. 177 B.10.7 Proof of Thm. 179
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
xii
382 383 384 384 385 385 385 385 386 386 386 388 389 390 390 391 393 393 395 396 396
List of Tables 2.1 2.2 2.3 2.4 2.5 2.6 2.7
Euclidean prototypes for topological concepts. . . . . . . . . . . . . . . Euclidean prototypes for differential-geometric concepts. . . . . . . . . Euclidean prototypes for Riemannian-geometric concepts. . . . . . . . . Euclidean prototypes for metric-geometric concepts. . . . . . . . . . . . n Lie group structures and associated Riemannian operators on S++ . . . n Riemannian operators of (θ, α, β)-EM and BWM on S++ . . . . . . . . . Isometric prototype spaces and diffeomorphisms on the correlation manifold. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.8 Vector operations and Riemannian operators under ECM and LECM. . 2.9 Vector operations and Riemannian operators under OLM and LSM. . . 2.10 Riemannian operators on the Grassmannian under ONB and PP. . . . 2.11 Gyro operators on the Grassmannian under ONB and PP. . . . . . . . 2.12 Lie group structures and Riemannian operators on rotation matrices. . 2.13 Gyro operators on stereographic and Beltrami–Klein models. . . . . . . 2.14 Riemannian operator templates for stereographic and radius constantcurvature coordinates. . . . . . . . . . . . . . . . . . . . . . . . . . . .
12 18 24 29 39 39
3.1 Review of SPD Lie groups and invariant metrics. . . . . . . . . . . . . 3.2 Review of the rotation Lie group and invariant metric. . . . . . . . . . 3.3 Review of full-rank correlation Lie groups and invariant metrics. . . . . 3.4 Summary of some representative RBN methods. . . . . . . . . . . . . . 3.5 Summary of LieBN types. . . . . . . . . . . . . . . . . . . . . . . . . . 3.6 Key operators in calculating LieBN on SPD manifolds. . . . . . . . . . 3.7 Key operators in calculating LieBN on the rotation matrices. . . . . . . 3.8 Summary of LieBN on the correlation. . . . . . . . . . . . . . . . . . . 3.9 10-fold average results of SPDNet with and without SPDBN or LieBN. 3.10 Cross-validation results of TSMNet with SPDDSMBN and DSMLieBN. 3.11 Results of LieNet with or without rotation LieBN. . . . . . . . . . . . . 3.12 Results of SPDNet with or without correlation LieBN under different invariant metrics. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.13 Comparison of previous RBN methods with GyroBN, where M and V denote the sample mean and variance. . . . . . . . . . . . . . . . . . . 3.14 Summary of operators for GyroBN across representative manifolds. . .
59 59 60 60 65 68 68 70 72 73 75
xiii
46 46 46 49 49 50 52 54
76 77 99
List of Tables 3.15 Efficiency (in µs) of gyroaddition on the radius manifold: closed form versus Riemannian definition. Values in parentheses indicate the runtime of the closed-form implementation as a percentage of the corresponding Riemannian implementation. The best results are bold. . . . . . . . . 3.16 Comparison of GyroBN against other Grassmannian BN methods under the GyroGr backbone. Here, accuracy is reported as a percentage, fit time denotes the average training time per epoch (s/epoch), and #Params is reported in millions. Values in parentheses specify the dimension of the Grassmannian input to the BN layer. The largest number of parameters is marked in red. . . . . . . . . . . . . . . . . . . . . . . 3.17 Comparison of GyroBN against LRBN across five CCSs. ROC is the testing AUC reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. When LRBN degenerates the backbone network, the results are highlighted with red. . . . . . . . . . . . . . . . 3.18 Comparison of CorNet with or without RBN layers. Accuracy is reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
100
103
106 107
4.1 4.2
Several MLRs on different geometries are special cases of our MLR. . . 123 Properties of deformed metrics on SPD manifolds (θ ̸= 0 and min(α, α + nβ) > 0). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124 4.3 Comparison of SPDNet with LogEig against SPD MLRs on the Radar data set. The best results are bold. . . . . . . . . . . . . . . . . . . . . 127 4.4 Comparison of SPDNet with LogEig against SPD MLRs on the HDM05 data set. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127 4.5 Inter-session TSMNet results on Hinss2021. . . . . . . . . . . . . . . . . 127 4.6 Inter-subject TSMNet results on Hinss2021. . . . . . . . . . . . . . . . 128 4.7 Comparison of LogEig against SPD MLRs under the RResNet architecture.129 4.8 Comparison of LogEig against SPD MLRs under the SPDGCN architecture. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130 4.9 Comparison of LogEig against SPD MLRs for direct classification. . . . 130 4.10 Results of LogEig MLR against Lie MLR under the LieNet architecture. 131 5.1 5.2 5.3 5.4 5.5 5.6 5.7
Failure and violation rates (%) of r ⊗H x in FP32. . . . . . . . . . . . . ∥Log0 (Exp0 (v)) − v∥. . . . . . . . . . . . . . . . . . . . . . . . . . . . . Gradient magnitude ∥∇x fr (x)∥ across varying radii. . . . . . . . . . . . Top-1 image classification accuracy (%) of hyperbolic MLRs on ResNet18. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic), which comes from Bdeir et al. [15, Tab. 1]. . . . . . . Accuracies of hyperbolic networks on graph learning. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic). . Results of Tangent FC (TFC) vs PV FC, and Tangent BN (TBN) vs GyroBN. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Comparison of methods in calculating mean and variance in PV GyroBN. Time is measured in milliseconds per training epoch. . . . . . . . . . . xiv
144 145 145 146 146 147 147
List of Tables 5.8
Ablations on PVNN with or without exponential map for the input PV feature. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 5.9 Ablations on PV activations. . . . . . . . . . . . . . . . . . . . . . . . . 148 5.10 Comparison in MCC of hyperbolic and Euclidean convolutional networks, including PVCNN, on TEB data sets. . . . . . . . . . . . . . . . . . . . 149 5.11 Comparison of C-class MLR. In Dist, Real means the point-to-hyperplane distance is the real distance, obtained by inf y∈H d(x, y), where H is a hyperplane and d is the geodesic distance; Pseudo denotes a surrogate that coincides with the real distance only in Euclidean geometry. Compact params indicate whether each logit avoids an additional manifold-valued parameter. Batch efficiency indicates whether the MLR can avoid inefficient per-class loops in implementation (see Sec. A.4.2.1). In #Params, we highlight the heaviest in red. In FLOPs, we mark the slowest in red and the fastest in green. . . . . . . . . . . . . . . . . . . . . . . . . . . 155 5.12 Comparison of hyperbolic FC layers. For simplicity, BFC layers do not involve the gyroaddition and assume ϕ is the identity map, which is in line with the Möbius and Lorentz FC layers. . . . . . . . . . . . . . . . 158 5.13 Top-1 image classification accuracy (%) of MLR methods on the ResNet18 backbone. The best results within each hyperbolic model are bold. The slowest MLR and largest parameter count are shown in red. . . . 159 5.14 Genomic MCC of MLR methods under the CNN backbone. The best results within each hyperbolic model are bold. . . . . . . . . . . . . . . 160 5.15 Fit time (s/epoch) on genome sequence learning. The fastest times are bold and the slowest ones are red. . . . . . . . . . . . . . . . . . . . . 161 5.16 Comparison of hyperbolic FC layers on link prediction. The best results within each hyperbolic model are bold. . . . . . . . . . . . . . . . . . . 161 5.17 Node classification F1 scores of hyperbolic MLRs on the HGCN backbone, where δ denotes graph hyperbolicity (lower is more hyperbolic). The best results within each hyperbolic model are bold. . . . . . . . . 162 5.18 Efficiency comparison: fit time (s/epoch) and parameter count. Slowest results and largest parameter counts are in red. . . . . . . . . . . . . . 163 5.19 Correspondence between Euclidean and correlation-based layers. For convolution, kernel-based FC refers to applying a convolution kernel to a receptive field, which is an FC transformation. . . . . . . . . . . . . . 165 5.20 Five-fold results and training time per epoch on four data sets. The top 3 results are highlighted with red, blue, and cyan. ∗ denotes reproduced results due to missing official code. . . . . . . . . . . . . . . . . . . . . 173 5.21 Ablations on mixed geometries. Each row shows the metric used for Convolution (Conv), and each column is the metric for MLR. The diagonal entries indicate configurations where both layers use the same metric. The best result in each row is bold. . . . . . . . . . . . . . . . . . . . 174 5.22 SPDNet: SPD vs. correlation. . . . . . . . . . . . . . . . . . . . . . . . 175 xv
List of Tables 5.23 Comparison of SPDMLR-Trivlz on raw covariances against CorMLR on raw correlations on all three data sets. The input matrix dimensions are 93 × 93, 63 × 63, and 20 × 20, respectively. . . . . . . . . . . . . . . . . 5.24 SPD networks with or without normalized SPD inputs. . . . . . . . . . 5.25 Comparison of CorNet with or without activations. . . . . . . . . . . . 5.26 Average runtime (s) of a single forward pass in CorNet under different metrics and input dimensions. The best results are bold. . . . . . . . . 6.1 6.2 6.3 6.4 6.5 6.6 6.7 6.8 6.9
Parameter learning for the general matrix logarithm and exponential. . Results of ALog on the HDM05 data set. The best results are bold. . . Results of ALog on the FPHA data set. . . . . . . . . . . . . . . . . . . Results of ALog on the AFEW data set. . . . . . . . . . . . . . . . . . Results of fixed bases on the HDM05 and FPHA data sets. . . . . . . . Comparison of RBN methods on the HDM05 data set. . . . . . . . . . Experiments on RResNet under different geometries. . . . . . . . . . . Comparison of Gyro MLRs on the NTU60 data set. . . . . . . . . . . . Riemannian and gyro operators of different metrics on the Cholesky manifold. For the diagonal log metric, log(·) and exp(·) are diagonal logarithm and exponentiation. . . . . . . . . . . . . . . . . . . . . . . . . . 6.10 SPD MLRs under different metrics on the SPDNet backbone. The best results are bold. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.11 SPD MLRs on the GyroSPD backbone. . . . . . . . . . . . . . . . . . . 6.12 Results on residual blocks. . . . . . . . . . . . . . . . . . . . . . . . . . 6.13 Failure probabilities (%) of geodesics under different metrics with small eigenvalues in L ∈ Ln++ . An output matrix containing any Inf or NaN is considered a failure. Here, DLM denotes the diagonal log metric, while DPM and DBWM denote θ-DPM and θ-DBWM, respectively. . . . . . 6.14 Swelling effects of geodesic SPD interpolations. Deeper greens indicate greater swelling. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.15 Number of matrix functions required per sample for a C-class SPD MLR. Spectral matrix functions include matrix logarithm, matrix power, and the Lyapunov operator. . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.16 Asymptotic per-sample complexity of a C-class SPD MLR for an n × n input SPD matrix. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.17 Average runtime (in seconds) of one SPD MLR training step across different matrix dimensions. . . . . . . . . . . . . . . . . . . . . . . . . . . A.1 A.2 A.3 A.4 A.5 A.6 A.7 A.8
Summary statistics for the graph data sets. . . . . . . . . . . . . . . . . Summary statistics for TEB. . . . . . . . . . . . . . . . . . . . . . . . . Summary statistics for the adopted GUE data sets. . . . . . . . . . . . Matrix powers in LieBN-Cor under different metrics on each data set. . (θ, α, β) of SPD MLRs on the SPDGCN backbone. . . . . . . . . . . . Candidate values for hyperparameters in SPD MLRs. . . . . . . . . . . Hyperparameters for PVNN that vary across graph data sets. . . . . . Hyperparameters for PVNN that are shared across graph data sets. . . xvi
175 180 181 182 196 198 199 200 201 202 203 203 211 214 214 215
216 218 218 219 220 246 247 248 257 259 260 262 262
List of Tables A.9 Summary of the hyperbolic layers used in the graph node classification models. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.10 Hyperparameters for TEB. . . . . . . . . . . . . . . . . . . . . . . . . . A.11 Summary of hyperparameters used in the image classification task. . . . A.12 Hyperparameters for genome sequence learning. . . . . . . . . . . . . . A.13 Hyperparameters for node classification on Disease, Airport, PubMed, and Cora. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.14 Hyperparameters in CorNets. . . . . . . . . . . . . . . . . . . . . . . . A.15 Hyperparameter θ in θ-PCM and θ-BWCM. It is selected from the candidate values used for SPDNet and GyroSPD. . . . . . . . . . . . . . . A.17 Comparison of point-to-hyperplane distances. Real means the point-tohyperplane distance is the real distance, obtained by inf y∈H d(x, y) with H as a hyperplane and d as the geodesic distance. Instead, Pseudo means the point-to-hyperplane distance is a surrogate, which only equals the real distance in Euclidean geometry. . . . . . . . . . . . . . . . . . A.16 Comparison of hyperplanes. Compact params indicate whether the parameterization requires an additional manifold-valued point. . . . . . .
xvii
263 263 264 265 265 267 272
282 282
List of Tables
xviii
List of Figures 2.1
3.1 3.2 3.3 3.4 3.5 3.6 3.7 3.8
3.9 4.1
4.2
4.3 4.4
5.1 5.2
The black stars denote 2 × 2 correlation matrices, while the red, green, and blue dots denote corresponding SPD matrices. The black dots denote the boundary of the SPD cone. . . . . . . . . . . . . . . . . . . . . . . Illustration of LieBN on the SPD, rotation, and correlation Lie groups. Minimal examples of applying LieBN. . . . . . . . . . . . . . . . . . . . Visualization of input and output SPD matrices in LieBN. . . . . . . . Test accuracy curves of LieNet with rotation LieBN. . . . . . . . . . . . Illustration of GyroBN on manifold-valued data. Blue points, green points, and the red dashed curves indicate the input samples, normalized outputs, and data distributions, respectively. . . . . . . . . . . . . . . . Minimal examples of applying GyroBN. . . . . . . . . . . . . . . . . . . Comparison of derivation logic for gyrotranslation isometries. . . . . . . Visualization of GyroBN across different geometries. Blue and green points represent input and normalized data, respectively. Red and cyan points denote the input and output batch means. Black points mark the manifold boundary, and the gray surface depicts the manifold. . . . . . Training and testing curves of 1-block GyroGr on two NTU data sets.
40 56 58 74 75 77 79 80
101 104
Conceptual illustration of SPD hyperplanes induced by (α, β)-LEM and θ-LCM. In each subfigure, the black dots are SPSD matrices, denoting 2 the boundary of S++ , while the blue, red, and yellow dots denote three SPD hyperplanes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Illustration of the deformation (left) and Venn diagram (right) of metrics on SPD manifolds, where IEM, SREM, and 14 PAM denote Inverse Euclidean Metric, Square Root Euclidean Metric, and Polar Affine Metric scaled by 1/4, respectively. . . . . . . . . . . . . . . . . . . . . . . . Conceptual illustration of SPD hyperplanes induced by five families of 2 Riemannian metrics. The black dots denote the boundary of S++ . . . Conceptual illustration of a Lie hyperplane. Each pair of antipodal black dots corresponds to a rotation matrix with an Euler angle of π, while the green dots denote a Lie hyperplane. . . . . . . . . . . . . . . . . . . .
127
Illustration: red curves are different horospheres of B v . . . . . . . . . . Validation accuracy curves on ImageNet-1k. . . . . . . . . . . . . . . .
152 159
xix
117
123 124
List of Figures 5.3
5.4
5.5
5.6 5.7 5.8 5.9 6.1 6.2 6.3
Illustration of the Log-Euclidean 1D convolution with two kernels. The 3-channel input is first split into two receptive fields along the channel dimension. In each receptive field, two kernels are applied to the product space. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Illustration of the PHCM convolution and MLR. The multi-channel input correlation matrices are denoted as {C i }ci=1 . For the convolutional layer, the illustration focuses on the transformation within a receptive field and assumes a single-channel output. . . . . . . . . . . . . . . . . . . . . . Illustration of the decision hyperplanes in the correlation MLRs under five different geometries. The 3×3 correlation manifold can be embedded as an open elliptope in R3 , by visualizing the strictly lower triangular part of each C ∈ Cor+ (3). The black dots denote the boundary. The PHCM hyperplane is defined by the one in the β-concatenated Poincaré space. Distribution of per-sample coefficients of variation of diagonal variances on FPHA. Higher values indicate stronger diagonal variability, which could cause nuisance noise. . . . . . . . . . . . . . . . . . . . . . . . . . Distribution of per-sample coefficients of variation of diagonal variances on HDM05. Higher values indicate stronger diagonal variability, which could cause nuisance noise. . . . . . . . . . . . . . . . . . . . . . . . . . Distribution of ratios of diagonal to off-diagonal entries on FPHA. . . . Distribution of ratios of diagonal to off-diagonal entries on HDM05. . . Accuracy curves on the FPHA data set. . . . . . . . . . . . . . . . . . . Visualization of parameters in the ALog layer on the HDM05 and FPHA data sets. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Geodesic interpolation of SPD matrices under different Riemannian metrics. Each 3 × 3 SPD matrix can be visualized as an ellipsoid [9]. The two endpoints are fixed across all metrics. . . . . . . . . . . . . . . . .
B.1 Illustration of the Euclidean spaces LT0 (m), Hol(m) and Row0 (m), where ⋆ can be obtained by symmetry. . . . . . . . . . . . . . . . . . . . . .
xx
168
170
174 176 177 178 178 199 201 217 375
Notation and Conventions This chapter collects the notation used throughout the thesis.
General Sets and Linear Algebra R, Rn
The real numbers and the n-dimensional Euclidean space.
R+ , R++
The non-negative and positive real numbers.
In , 0n , 1n , 1, 0n×n
The n × n identity matrix, the n-dimensional zero vector, the n-dimensional all-one vector (also written 1 when its dimension is clear), and the n × n zero matrix. Dimension subscripts may be omitted when they are clear from context.
⟨x, y⟩, ∥x∥
The Euclidean inner product and the induced norm.
∥A∥, tr(A), rank(A)
The Frobenius norm, trace, and rank of a matrix.
diag(x), diag(A)
The diagonal matrix induced by a vector and the diagonal part of a matrix, with the intended meaning specified by context.
argmin, argmax
The minimizer and maximizer operators.
Manifold Geometry X, Y
Generic sets or topological spaces.
X
A generic metric space.
M, N
Smooth or Riemannian manifolds.
C ∞ (M)
The space of smooth real-valued functions on M.
Tp M
The tangent space of M at p.
gp
The Riemannian metric at p, viewed as an inner product on Tp M. xxi
Notation and Conventions gL, gR
Left- and right-invariant Riemannian metrics on a Lie group.
⟨v, w⟩p , ∥v∥p
The Riemannian inner product and norm in Tp M.
dM (x, y)
The geodesic distance between x, y ∈ M.
α, L(α)
A general curve on a space and its length.
γ
A geodesic or geodesic ray.
η
The momentum parameter used in LieBN and GyroBN running-statistics updates.
MKn , dK , DK
The n-dimensional model space of constant curvature K, its distance, and the diameter of MK2 .
CAT(K)
A metric space whose geodesic triangles satisfy the CAT(K) comparison inequality.
πC
The orthogonal projection onto a convex subset C of a CAT(0) space.
∂X
The boundary at infinity of a metric space, defined by equivalence classes of asymptotic geodesic rays when available.
B γ (x), HBτγ , Hτγ
The Busemann function associated with a geodesic ray γ, its horoball, and its horosphere at level τ .
Expp , Logp
The Riemannian exponential and logarithmic maps at p.
PTp→q
Parallel transport from Tp M to Tq M.
Tp→q
A vector transport from Tp M to Tq M.
Ha,p , H̃Ã,P
Euclidean and Riemannian hyperplanes used by multinomial logistic regression.
Pk , Ak , Ãk
Class-wise Riemannian prototype, fixed tangent-space parameter, and transported tangent parameter in RMLR.
zk , rk , vk (x)
The Euclidean direction parameter, scalar offset parameter, and class score used by the PV MLR and PV FC layers.
M , B, v 2 , s
The Riemannian mean, biasing parameter, variance, and scaling factor used by Riemannian normalization layers.
f∗,p , dp f (v)
The differential of a smooth map f at p, written in Chapter 2 as a linear map and, in later computations, equivalently as applied to a tangent vector v. xxii
Notation and Conventions ∇w f , gradw f
The Euclidean gradient of a scalar objective at a Euclidean parameter w and the Riemannian gradient at w ∈ M, respectively.
X(M), X(α)
Smooth vector fields on M and smooth vector fields along a curve α.
[X, Y ]
The Lie bracket of smooth vector fields X, Y ∈ X(M).
D , V′ D, dt
The Levi-Civita connection and the induced derivative V ′ = D (V ) along a curve. dt
FM({xi }), WFM({wi }, {xi })
The Fréchet mean and weighted Fréchet mean.
Algebraic and Gyrovector Operations G, e, E
A group-like algebraic structure and its identity element.
⊕
A binary operation for groups, Lie groups, and gyrogroups.
Lx , Rx
Left and right translations by x on a Lie group.
Lx
Left gyrotranslation by x, defined by Lx (y) = x ⊕ y.
⊖x
The inverse of x under an addition-like operation. In gyrogroups, this is called the gyro-inverse.
⊙
Scalar multiplication in gyrovector spaces.
gyr[x, y]
The gyration generated by x and y.
⟨x, y⟩gyr
The gyro inner product.
∥x∥gyr
The gyronorm.
dgyr (x, y)
The gyrodistance.
Barη
The binary barycenter operator with weight η, used for running mean updates on gyrogroups.
Matrix Manifolds Sn
The vector space of n × n symmetric matrices.
n S++
The manifold of n × n symmetric positive definite matrices.
Sn+
The set of n × n symmetric positive semidefinite matrices.
Chol(P )
The Cholesky factor of an SPD matrix P . xxiii
Notation and Conventions Ln++
The space of n × n lower triangular matrices with positive diagonal entries.
log(P ), exp(A), logα (P ), log−1 α (A)
The matrix logarithm and exponentiation, followed by the general matrix logarithm used by ALEM and its inverse.
⊕ALE , ⊙ALE
The ALEM-induced element addition and scalar multiplin cation on S++ .
f (S), f∗,S , Lf
A symmetric matrix function induced by a scalar function f , its differential at S, and the corresponding Loewner matrix.
A⊛B
The Hadamard product of matrices A and B with the same size.
log∗,P , Chol∗,P , Chol−1 ∗,L
The differentials of the matrix logarithm, the Cholesky decomposition, and the inverse Cholesky map at the indicated points.
Pθ (P ), (Pθ )∗,P
The matrix power map P 7→ P θ on SPD matrices and its differential at P .
n → S n , ⊙ϕ , g ϕ , ϕ : S++ dϕ
An isometry from SPD matrices to a Euclidean space and its induced abelian group operation, pullback Euclidean metric, and geodesic distance.
Dlog(D), ψLC (P )
The diagonal element-wise logarithm and the Log-Cholesky map ψLC (P ) = ⌊L⌋ + Dlog(D(L)), where L = Chol(P ).
LP [V ]
The Lyapunov operator defined by LP [V ]P + P LP [V ] = V .
⟨A, B⟩(α,β) , ∥A∥(α,β)
The two-parameter Euclidean inner product on S n and its induced norm.
ST
The admissible parameter set {(α, β) ∈ R2 | min(α, α + nβ) > 0}.
⌊A⌋, D(A)
The strictly lower triangular part and the diagonal matrix formed from the diagonal of A.
A1/2
The Cholesky half-diagonal operator A1/2 = ⌊A⌋ + 21 D(A).
AI
Affine-invariant geometry on SPD matrices.
LE
Log-Euclidean geometry on SPD matrices.
LC
Log-Cholesky geometry on SPD matrices.
PE
Power-Euclidean geometry on SPD matrices. xxiv
Notation and Conventions (α, β)-AIM
The two-parameter Affine-Invariant Metric (AIM) on SPD matrices.
(α, β)-LEM
The two-parameter Log-Euclidean Metric (LEM) on SPD matrices.
(θ, α, β)-AIM
The three-parameter Affine-Invariant Metric (AIM) on SPD matrices.
θ-LCM
The parameterized Log-Cholesky Metric (LCM) on SPD matrices.
(θ, α, β)-EM
The three-parameter Euclidean Metric (EM) on SPD matrices.
(a, b)-ALEM
The Adaptive Log-Euclidean Metric on SPD matrices.
θ-PCM
The deformed Power-Cholesky Metric on SPD matrices.
(θ, M)-BWCM
The deformed Bures–Wasserstein–Cholesky Metric on SPD matrices.
⊕AI , ⊙AI
AIM-induced SPD gyroaddition and scalar gyromultiplication.
⊕LieAI , ⊕LE , ⊕LC , ⊕θ-AI , ⊕θ-LC , ⊖LieAI
SPD Lie group operations under AIM, LEM, LCM, and their power-deformed AIM and LCM variants, with ⊖LieAI denoting the inverse under ⊕LieAI when operation-specific disambiguation is needed.
⊕LE , ⊙LE , ⊕LC , ⊙LC
LEM- and LCM-induced SPD gyrovector operations, which coincide with linear operations in the corresponding global charts.
BW
Bures–Wasserstein geometry on SPD matrices.
g AI , g LE , g LC , g θ-E , g BW , g CRI , g θ-CRI
Metric tensors associated with the corresponding SPD geometries.
GL(n), SL(n)
The general linear group and the special linear group.
O(n)
The orthogonal group.
St(p, n)
The Stiefel manifold represented by orthonormal frames.
f n) Gr(p, n), Gr(p,
The Grassmannian manifold represented by orthonormal bases and projection matrices. xxv
Notation and Conventions Ip,n , Iep,n
Identity points of the Grassmannian under the Orthonormal Basis (ONB) and Projector Perspective (PP) representations.
⊕Gr , ⊙Gr , ⊖Gr
Grassmannian gyroaddition, scalar multiplication, and inverse under the ONB representation.
e Gr , ⊙ e Gr , ⊖ e Gr ⊕
Grassmannian gyroaddition, scalar multiplication, and inverse under the PP representation.
f n) π : Gr(p, n) → Gr(p,
The Grassmannian isometry π(U ) = U U ⊤ from ONB to PP.
Matrix commutator
The operation [A, B] = AB − BA.
Cor+ (n)
The manifold of full-rank correlation matrices.
Cor, I, inv
The correlation normalization map, the cor-inversion operator on correlation matrices, and the matrix inversion operator on SPD matrices.
LTn
The Euclidean space of n × n lower triangular matrices.
LT0 (n)
The Euclidean space of n × n strictly lower triangular matrices.
LT1 (n)
The affine space of n×n lower triangular matrices with unit diagonal.
Ln
The manifold of Cholesky factors of full-rank correlation matrices, consisting of lower triangular matrices with positive diagonal entries and unit row norm.
Hol(n), Row0 (n)
The spaces of symmetric hollow matrices and symmetric matrices with null row sum, respectively.
Θ, ϕEC , Log◦ , Exp◦ , Log⋆ , Exp⋆
Diffeomorphisms and charts used by ECM, OLM, and LSM on correlation matrices.
D(H), D⋆ (C)
The diagonal correction operators used in OLM and LSM, respectively.
Diag(n), Diag+ (n)
The vector space of diagonal matrices and the positive diagonal matrix manifold.
Sn , S± (n)
The permutation group and the signed permutation group acting on correlation matrices.
(X)sym , (X)sym+
The averaged symmetrization (X + X ⊤ )/2 and the unnormalized symmetrization X + X ⊤ , respectively. xxvi
Notation and Conventions HSi , PHSn−1 , PPn−1
The open hemisphere, product hemisphere space, and product of unit Poincaré balls used by PHCM and GyroBN on correlation manifolds.
Φ : Cor+ (n) → PPn−1
The row-wise Cholesky identification from full-rank correlation matrices to the product of unit Poincaré balls.
SO(n), so(n)
The special orthogonal group and its Lie algebra.
Constant-Curvature Manifolds MnK
The K-radius model with curvature parameter K.
stnK , DnK
The K-stereographic model and its positive-curvature projected hypersphere case.
SnK
The spherical model space with curvature parameter K.
Sn , Dn
The unit sphere and unit projected hypersphere. Unit-space notation suppresses the fixed curvature subscript.
PnK
The Poincaré ball model of hyperbolic space with curvature K < 0.
Pn , K n , H n , L n
The unit Poincaré ball, unit Beltrami–Klein ball, unit hyperboloid, and unit Lorentz reference space. Unit-space notation suppresses the fixed curvature subscript.
⊕K , ⊖K , ⊙K
Gyroaddition, gyro-inverse, and scalar gyromultiplication on the K-stereographic model.
λK x
The conformal factor of the K-stereographic model at x.
⊕M , ⊖M , ⊙M
Möbius gyroaddition, gyro-inverse, and scalar multiplication on the Poincaré ball.
M M ⊕M K , ⊖K , ⊙K
Gyroaddition, gyro-inverse, and scalar gyromultiplication on the K-radius model.
LnK , HnK
Equivalent notation for the Lorentz, or hyperboloid, model of hyperbolic space with curvature K < 0.
⊕L , ⊖L , ⊙L
Lorentz gyroaddition, gyro-inverse, and scalar gyromultiplication on the Lorentz model.
⟨x, y⟩L , ∥x∥L
The Lorentzian inner product and its associated quantity. xxvii
Notation and Conventions ⟨x, y⟩K , ∥x∥K
The curvature-dependent ambient bilinear form on the Kradius model and its induced tangent-space norm. On the K < 0 branch, the ambient form is Lorentzian and becomes positive definite only after restriction to a tangent space.
KnK
The Beltrami–Klein model of hyperbolic space with curvature K < 0.
⊕E , ⊖ E , ⊙ E
Einstein gyroaddition, gyro-inverse, and scalar gyromultiplication on the Beltrami–Klein model.
γx
The Einstein gamma factor on the Beltrami–Klein model.
tanK , sinK , cosK
Curvature-aware trigonometric functions used for constantcurvature models.
PVnK
The Proper Velocity model of hyperbolic space with curvature K < 0.
⊕U , ⊖ U , ⊗ U
PV gyroaddition, gyro-inverse, and scalar gyromultiplication.
βx , γ y
The relativistic beta factor on PV space and the gamma factor used by the PV–Poincaré isometry.
πPVnK →PnK , πPnK →PVnK
The mutually inverse isometries between the PV model and the Poincaré ball.
πPVnK →HnK , πHnK →PVnK
The mutually inverse isometries between the PV model and the hyperboloid model.
xxviii
Chapter 1 Introduction 1.1
Riemannian Deep Learning
Over the past decade or so, Deep Neural Networks (DNNs) have achieved significant progress in machine learning [102, 127, 97, 202]. Traditionally, DNNs have been developed under the assumption that the latent geometry of the input data is Euclidean. However, many applications involve non-Euclidean structures, such as manifolds [30, 88]. Therefore, deep learning over Riemannian spaces, referred to as Riemannian deep learning, has shown great success in diverse applications, such as computer vision [107, 106, 108, 47, 120], natural language processing [76, 181, 141], graph and knowledge-graph learning [163, 38, 42, 136], multimodal learning [67, 166], recommendation systems [223], signal processing [31, 123], human neuroimaging [167, 134, 135, 233], medical imaging [37], astronomy [44], and genome sequence modeling [119]. Commonly encountered manifolds include special orthogonal groups [165], Symmetric Positive Definite (SPD) manifolds [9], Grassmannian manifolds [71, 19], spherical manifolds [165], and hyperbolic manifolds [173, 33]. Many of these manifolds admit computationally tractable Riemannian operators, including geodesics, exponential and logarithmic maps, and parallel transport. Building on these geometric tools, several fundamental Euclidean neural network components have been generalized to manifold-valued data, including normalization [31, 34, 142, 123], attention [167], residual blocks [201, 118], classification [76, 159, 161], and Fully Connected (FC) and convolutional layers [106, 107, 108, 76, 181, 45, 161]. However, most existing constructions either remain tied to particular geometries or carry out neural computations in intermediate flat spaces that approximate the intrinsic geometry. Existing normalization methods are either confined to particular SPD met1
1.1. Riemannian Deep Learning rics [31, 123], restricted to matrix Lie groups with a specific distance [34, Sec. 3.2], or applicable more generally but lack theoretical guarantees for controlling sample statistics, as in ManifoldNorm [34, Algs. 1–2] and RBN [142, Alg. 2]. Classification layers frequently map manifold-valued features to tangent Euclidean spaces [106, 31], ambient Euclidean spaces [107], or coordinate Euclidean spaces [36], while intrinsic alternatives may require the generalized law of sines [76] or gyrovector structures [159, 161]. FC and convolutional layers exhibit similar limitations. Early constructions are tailored to SPD, rotation, and Grassmannian manifolds [106, 107, 108], while hyperbolic counterparts rely on tangent-space mappings [76, 147], Lorentz spacetime [45], or Poincaré geometry [181]. The weighted-Fréchet-mean convolution [37] applies more broadly, but constrains the output manifold dimension to match the input dimension. These limitations motivate unified principles that formulate a module once at an appropriate geometric level and then instantiate it across manifolds carrying the required common structure. Beyond individual modules, deep architectures have been developed for matrixvalued manifolds, including SPD manifolds [106, 37], Grassmannian manifolds [108], and rotation manifolds [107], as well as vector-valued manifolds, including spherical manifolds [13, 182] and hyperbolic manifolds [76, 13, 182, 45, 15]. Nevertheless, existing networks remain concentrated on a limited set of geometric representations and construction tools. Hyperbolic networks, for example, predominantly use the Poincaré ball [76, 181] and Lorentz (hyperboloid) models [45, 15], while alternative models remain less explored [33]. Beyond standard Riemannian and gyrovector operators, Busemann functions and horospheres [29, Ch. II.8] have been incorporated into hyperbolic SVMs [73], hyperbolic PCA [39], Sliced-Wasserstein distances [24], and prototype-learning methods [81]. Despite these developments, their use in neural network design remains limited. Likewise, correlation matrices have received far less attention than SPD covariance representations despite being statistically compact alternatives to covariance matrices [7]. Their recently developed tractable Riemannian geometries [195, 191] indicate a broader design space that remains underexplored. Finally, both the module and network designs described above ultimately depend on the underlying Riemannian geometry. A Riemannian metric is not merely a means of measuring distance. It provides the theoretical foundation and concrete computational primitives of Riemannian learning algorithms. By assigning inner products to tangent spaces, the metric determines geodesic distances, exponential and logarithmic maps, parallel transport, and Fréchet statistics [165]. These operators enter directly into the design of deep-network components. For example, hyperbolic geometry de2
Chapter 1. Introduction fines Riemannian classifiers [76, 181], weighted Fréchet means support convolutional layers [37] and normalization [31, 34, 123], and exponential and logarithmic maps underpin attention [167], residual blocks [118], and pooling layers [205]. Consequently, changing the Riemannian metric changes not only the geometry, but also the formulas, parameterizations, computational cost, numerical stability, and ultimately the practical behavior of the associated deep networks. This role is especially evident on the SPD manifold, where a broad range of metrics has been developed [171, 9, 70, 137, 22, 93]. However, most existing metric tensors are fixed, which may limit the expressivity of the induced geometry and its ability to adapt to data or the dynamics of deep networks. Moreover, numerical stability is particularly important in deep-network training, where metric-induced operators are repeatedly evaluated. Developing Riemannian metrics that balance flexibility, tractability, efficiency, and stability is therefore a significant problem for Riemannian deep learning.
1.2
Contributions and Outlines
The contributions of this thesis are organized around three connected perspectives: unified Riemannian module design across manifolds, manifold-specific Riemannian network design, and the design of the underlying Riemannian geometries. The first perspective formulates principled network modules which can be applied to different geometries, including normalization and classification layers. The second perspective addresses cases in which a fully general construction is either intractable or unable to exploit useful manifold-specific structure. It therefore uses the additional structures of particular manifolds to develop manifold-specific modules and network architectures. The third perspective moves from network design under prescribed geometries to the design of the geometry itself, developing flexible, efficient, and numerically stable metrics on SPD manifolds. The contributions and organization of the remaining chapters are summarized below. Chapter 2 establishes the mathematical foundations used throughout the thesis. The chapter is organized into two parts. The first part develops the general theory required by the subsequent methods, including topology, differential and Riemannian geometry, metric geometry, algebraic structures on manifolds, Riemannian optimization, and matrix functions, and relates these constructions to their Euclidean counterparts. The second part turns to the concrete manifolds studied in later chapters: SPD, full-rank correlation, Grassmannian, and constant-curvature manifolds, together with special orthogonal groups. 3
1.2. Contributions and Outlines Chapter 3 develops unified frameworks for Batch Normalization (BN) across structured classes of manifolds, with the goal of controlling both the Riemannian mean and variance beyond a single manifold or metric. Sec. 3.2 first introduces Lie Group Batch Normalization (LieBN) on Lie groups under invariant metrics. LieBN uses group translations for centering and biasing and tangent-space scaling at the identity for variance control, thereby extending Euclidean BN while preserving the manifold structure. Concrete manifestations are developed on SPD, rotation, and full-rank correlation manifolds, together with efficient implementations and experimental validation. Sec. 3.3 first introduces pseudo-reductive gyrogroups, a new algebraic structure that generalizes classical gyrogroups and Lie groups, thereby providing a more general foundation for principled normalization. Building on this structure, it develops Gyrogroup Batch Normalization (GyroBN), which recovers LieBN as a special case and extends the same normalization principle to manifolds that need not possess a Lie group structure. Finally, GyroBN is instantiated on the Grassmannian, constant-curvature, and full-rank correlation manifolds, and experiments evaluate the resulting layers across matrix-manifold, constant-curvature, and graph learning tasks. Chapter 4 develops a unified framework for intrinsic classification across Riemannian manifolds. Sec. 4.2 first extends Euclidean Multinomial Logistic Regression (MLR) to SPD manifolds with flat pullback metrics. By formulating classification through the geodesic margin between an SPD input and a decision hyperplane, it derives closedform classifiers and provides an intrinsic explanation for the widely used LogEig MLR. Sec. 4.3 then replaces the potentially intractable point-to-hyperplane infimum with a Riemannian-trigonometric formulation, extending the classifier beyond flat SPD geometries. The resulting Riemannian Multinomial Logistic Regression (RMLR) requires only an explicit Riemannian logarithmic map, incorporates several existing manifold classifiers as special cases, and yields new instantiations on SPD manifolds under multiple metric families and on the special orthogonal group. Finally, Sec. 4.4 evaluates these classifiers within feedforward, residual, graph, and Lie group networks, as well as in direct manifold-valued classification. Chapter 5 turns to manifold-specific Riemannian network design, exploiting the additional structures of particular hyperbolic models and correlation manifolds. Sec. 5.2 first introduces the Proper Velocity (PV) model, an unconstrained model of hyperbolic space, and establishes its Riemannian geometry. Based on this geometry, Proper Velocity Neural Networks (PVNNs) develop MLR, FC, convolutional, activation, and normalization layers for stable hyperbolic deep learning. Sec. 5.3 uses Busemann functions and horospheres to develop intrinsic and efficient Busemann Multinomial Logistic 4
Chapter 1. Introduction Regression (BMLR) and Busemann Fully Connected (BFC) layers on both the Poincaré and Lorentz models. BMLR interprets its logits through point-to-horosphere distances, while BFC uses the same Busemann logits to construct feature transformations. Sec. 5.4 then develops Correlation Networks (CorNets) for full-rank correlation matrices. It constructs MLR, FC, and convolutional layers under five correlation geometries and derives accurate Riemannian backpropagation. Experiments across vision, graph, and genome learning tasks validate these manifold-specific designs. Chapter 6 develops flexible, fast, and numerically stable Riemannian metrics on SPD manifolds. Whereas the preceding chapters design modules and networks under prescribed geometries, this chapter designs the underlying metrics. Sec. 6.2 first develops Adaptive Log-Euclidean Metrics (ALEMs). By parameterizing general matrix logarithms within a pullback framework, ALEMs adapt the geometry to data while retaining closed-form Riemannian operators and a compatible abelian Lie group structure; the learned metrics are instantiated in SPD networks and other Riemannian building blocks. Sec. 6.3 then develops product Cholesky geometries, including the Power-Cholesky Metric (PCM) and Bures–Wasserstein–Cholesky Metric (BWCM). These metrics avoid the scalar logarithms and exponentials used by the Log-Cholesky Metric (LCM) and admit fast and stable closed-form Riemannian and algebraic operators. Their effectiveness is evaluated in SPD classification, residual learning, and tensor interpolation, while separate experiments assess computational efficiency, scalability, and numerical stability. Chapter 7 summarizes the thesis and discusses future directions.
1.3
Summary of Papers Excluded from the Thesis
The technical chapters focus on the core theoretical and methodological contributions. Fourteen additional publications on Riemannian deep learning belong to the same research theme and are summarized below. • Riemannian attention. We first developed self-attention on three specific manifolds: the Grassmannian [210], the full-rank correlation manifold [103], and the SPD manifold under the Bures–Wasserstein geometry [113]. Then, we generalized geometry-specific constructions into a unified attention framework for general matrix manifolds [212]. • Geometry-aware electroencephalography (EEG) representation learning. In Li et al. [135], we used hyperbolic embeddings to represent the hierarchical structure of EEG signals and improve cross-domain generalization. In Hu et al. 5
1.3. Summary of Papers Excluded from the Thesis [104], we introduced Riemannian high-order pooling on SPD manifolds to capture second-order correlations in EEG foundation models. • Geometry-specific Riemannian batch normalization. We developed SPD BN methods under specific Riemannian metrics [214, 215] for stable training. • Skeleton-based action recognition. We developed two SPD-based approaches to skeleton action recognition: a graph convolutional network that uses Gaussian embeddings of high-order skeletal statistics to model inter-subject interactions and global correlations for two-person interaction recognition [216], and an approach that partitions the skeleton into semantic body regions and models their longrange dependencies with SPD representations [213]. • Hyperbolic neural networks. In Shi et al. [179], we constructed a fully intrinsic hyperbolic Lorentz neural network whose FC, BN, concatenation, activation, and dropout modules operate within the Lorentz geometry. • Hyperbolic multi-view clustering. In Wang et al. [217], we developed hyperbolic multi-view clustering that aligns view-specific distributions through a hyperbolic sliced-Wasserstein distance while preserving hierarchical semantics in the Lorentz manifold. • Riemannian interpretation of global covariance pooling. In Chen et al. [53], we provided a unified Riemannian interpretation of matrix functions in global covariance pooling, showing that their effectiveness is explained by the Riemannian classifiers they implicitly respect. • SPD deep metric learning. In Wang et al. [211], we combined an SPD network encoder with a Riemannian decoder, local covariance regularization, and deep metric learning to improve image-set classification. These papers are excluded from this thesis because they extend or apply the core works developed in this thesis. For example, CorAtt [103] applies OLM/LSM correlation geometry to EEG attention, while Sec. 5.4 systematizes how to build neural networks and differentiation over the correlation matrices. GyroAtt [212] adopts the gyro paradigm, following Sec. 3.3. HEEGNet [135] applies the Lorentz GyroBN developed in Sec. 3.3 to domain-specific normalization. CBN [214] and GBWBN [215] specialize Sec. 3.2’s template to specific metrics. ILNN [179] accelerates the Lorentz 6
Chapter 1. Introduction GyroBN developed in Sec. 3.3 and follows Sec. 5.2’s point-to-hyperplane design. Finally, RiemGCP [53] applies the RMLR framework developed in Sec. 4.3 to interpret matrix-function normalization as implicit SPD Riemannian classifiers.
7
1.3. Summary of Papers Excluded from the Thesis
8
Chapter 2 Mathematical Background 2.1
Introduction
This chapter presents the mathematical foundations of the thesis in two parts. The first part develops the general theory and computational tools used throughout the subsequent chapters, beginning with topology and differential geometry [197], proceeding to Riemannian geometry [165], metric geometry [29], and algebraic structures on manifolds [128, 200], and then reviewing Riemannian optimization [1, 27], trivialization [132], and matrix functions and their differentials [21]. Although these concepts can be abstract, many generalize familiar Euclidean constructions, so Euclidean prototypes are provided whenever appropriate. The second part reviews the concrete spaces used later: Symmetric Positive Definite (SPD) and full-rank correlation manifolds, Grassmannian manifolds, special orthogonal groups, and constant-curvature manifolds. Their Riemannian and algebraic operators expose both the common primitives that support unified module design across manifolds and the additional structures later exploited by manifold-specific methods.
2.2
Topology
Topology provides the language for continuity and locality before coordinates are introduced [197, App. A]. This is essential because a manifold is usually not defined by a single global coordinate system as in Euclidean space, but by a family of local coordinate systems. Recall classical analysis on Rn , where many basic notions are formulated in terms of open sets. In the standard Euclidean setting Rn , openness is defined through open 9
2.2. Topology balls: a subset U ⊂ Rn is open if, for every x ∈ U , there exists r > 0 such that the open ball Br (x) = {y ∈ Rn | ∥y − x∥ < r} is contained in U . However, a general abstract space need not carry a norm or a distance, so it may not have a prior notion of open balls. Topology abstracts the essential closure properties of Euclidean open sets, leading to the following definition. Definition 1 (Topological space [197, Def. A.1]). A topology on a set X is a collection T of subsets of X such that (1) ∅ ∈ T and X ∈ T ,
(2) every union of elements of T belongs to T , i.e., for any {Uα }α∈A ⊂ T , [
α∈A
Uα ∈ T .
(2.1)
(3) every finite intersection of elements of T belongs to T , i.e., for any integer m ≥ 1 and any open sets U1 , . . . , Um ∈ T , m \
i=1
Ui ∈ T .
(2.2)
The pair (X, T ) is called a topological space. Elements of T are called open sets. A subset F ⊂ X is called closed if X \ F is open, where X \ F = {x ∈ X | x ∈ / F }. A basis records enough open sets to reconstruct the full topology, and second countability controls the size of this local description. Definition 2 (Basis and second countability [197, Defs. A.6 and A.12]). Let (X, T ) be a topological space. A collection B ⊂ T is a basis for T if every open set U ∈ T is a union of elements of B. Equivalently, for every x ∈ U with U ∈ T , there exists B ∈ B such that x ∈ B ⊂ U . A topological space is second countable if its topology has a countable basis. The next proposition gives the corresponding construction criterion when one starts from a candidate family of subsets rather than from an already specified topology. Proposition 3 (Criterion [197, Prop. A.8]). Let X be a set and let B be a collection of subsets of X. Then B is a basis for some topology T on X if and only if (1) X is the union of all sets in B, 10
Chapter 2. Mathematical Background
(2) whenever x ∈ B1 ∩ B2 with B1 , B2 ∈ B, there exists B3 ∈ B such that x ∈ B3 ⊂ B1 ∩ B2 . In Euclidean space Rn , the family of open balls generates the standard topology. Moreover, Rn is second countable because the open balls with centers in Qn and positive rational radii form a countable basis. Subspace topology is used whenever a geometric object is described as a subset of an ambient space. Definition 4 (Subspace topology [197, Sec. A.2]). Let (X, T ) be a topological space and let Y ⊂ X. The subspace topology on Y is TY = {Y ∩ U | U ∈ T } .
(2.3)
With this topology, Y is called a subspace of X. For Euclidean space, the unit sphere S n−1 = {x ∈ Rn | ∥x∥ = 1} carries the subspace topology inherited from Rn . Its open sets are exactly the sets S n−1 ∩ U , where U is open in Rn . Definition 5 (Continuous map and homeomorphism [197, Sec. A.7]). Let (X, TX ) and (Y, TY ) be topological spaces. A map f : X → Y is continuous if f −1 (V ) ∈ TX for every V ∈ TY . A map f : X → Y is a homeomorphism if it satisfies: (1) f is bijective,
(2) f and f −1 are continuous. In the Euclidean special case, this definition recovers the familiar notion of continuity from calculus. For a map f : Rn → Rm with the standard topologies, f is continuous in the sense of Thm. 5 if and only if, for every x ∈ Rn and every ε > 0, there exists δ > 0 such that ∥f (y) − f (x)∥ < ε whenever ∥y − x∥ < δ. Homeomorphic spaces are topologically identical. A simple homeomorphism is the translation x 7→ x + a on Rn , whose inverse is x 7→ x − a.
For manifolds, one also needs a separation condition that prevents distinct points from being topologically indistinguishable. Definition 6 (Hausdorff space [197, Def. A.16]). A topological space X is Hausdorff if, for any two distinct points x, y ∈ X, there exist disjoint open sets U, V ⊂ X such that x ∈ U and y ∈ V . 11
2.3. Differential Geometry Topological concept
Euclidean special case
Open and closed sets
Br (x) = {y ∈ Rn | ∥y − x∥ < r}, B r (x) = {y ∈ Rn | ∥y − x∥ ≤ r}. BQ = {Br (q) | q ∈ Qn , r ∈ Q>0 }. TS n−1 = {S n−1 ∩ U | U ⊂ Rn open}. Continuous map in calculus. x ̸= y ⇒ Br (x) ∩ Br (y) = ∅ for 0 < r < 1 ∥x − y∥. 2
Basis and second countability Subspace topology Continuous map Hausdorff space
Table 2.1: Euclidean prototypes for topological concepts. The Hausdorff condition ensures that points can be separated by neighborhoods. For Euclidean space, if x ̸= y, then the open balls Br (x) and Br (y) are disjoint whenever 0 < r < 21 ∥x − y∥. Thus Euclidean space is Hausdorff. Tab. 2.1 summarizes the Euclidean special cases of the above topological concepts.
2.3
Differential Geometry
In Rn , a basis provides a global coordinate system: every point is represented by a unique coordinate tuple. By contrast, a manifold usually does not admit a single coordinate map covering the entire space. Instead, each point has a neighborhood homeomorphic to an open subset of Rn , and this homeomorphism provides Euclidean coordinates on that neighborhood. This local Euclidean structure permits derivatives, tangent vectors, and smooth maps to be defined intrinsically. Definition 7 (Topological manifold [197, Sec. 5]). An n-dimensional topological manifold is a topological space M such that (1) M is Hausdorff,
(2) M is second countable, (3) every point p ∈ M has a neighborhood U homeomorphic to an open subset of Rn . The integer n is called the dimension of M. The Hausdorff and second-countability assumptions exclude pathological spaces and ensure that local coordinates behave like ordinary Euclidean neighborhoods. Charts are local coordinate systems, and an atlas is a collection of such coordinate systems covering the whole manifold. 12
Chapter 2. Mathematical Background Definition 8 (Chart and atlas [197, Sec. 5]). Let M be an n-dimensional topological manifold. A chart on M is a pair (U, φ) satisfying: (1) U ⊂ M is open,
(2) φ : U → φ(U ) ⊂ Rn is a homeomorphism onto an open subset of Rn .
Writing φ = (x1 , . . . , xn ), the functions xi : U → R are called the coordinate functions of the chart. Equivalently, if ri : Rn → R denotes the i-th standard coordinate projection, then xi = ri ◦ φ. An atlas is a collection of charts A = {(Uα , φα )}α∈A satisfying [ M= Uα . (2.4) α∈A
To do calculus consistently across charts, changes of coordinates must preserve smoothness. Definition 9 (Smooth atlas and smooth manifold [197, Sec. 5]). Two charts (U, φ) and (V, ψ) on M are smoothly compatible if the following transition maps are smooth maps between open subsets of Euclidean spaces: (1) ψ ◦ φ−1 : φ(U ∩ V ) → ψ(U ∩ V ), (2) φ ◦ ψ −1 : ψ(U ∩ V ) → φ(U ∩ V ).
A smooth atlas is an atlas whose charts are pairwise smoothly compatible. A smooth manifold is a topological manifold equipped with a maximal smooth atlas. Smooth compatibility means that changing coordinates does not destroy differentiability. This is the formal mechanism that allows a derivative computed in one coordinate chart to represent an intrinsic geometric object. Smooth maps between manifolds are defined by checking their coordinate representations. Definition 10 (Smooth map and diffeomorphism [197, Sec. 6]). Let M and N be smooth manifolds. A map f : M → N is smooth if, for every p ∈ M, every chart (U, φ) around p, and every chart (V, ψ) around f (p) with f (U ) ⊂ V , the coordinate representation ψ ◦ f ◦ φ−1 : φ(U ) → ψ(V ) (2.5) is smooth. A smooth map f : M → N is a diffeomorphism if it satisfies: (1) f is bijective,
13
2.3. Differential Geometry (2) f −1 : N → M is smooth. Diffeomorphisms refine homeomorphisms by preserving smooth structure, so diffeomorphic manifolds are identical from the viewpoint of smooth geometry. Before defining tangent spaces on manifolds, it is useful to recall what a tangent vector does in Euclidean calculus. Fix x ∈ Rn and a direction v ∈ Rn . The directional derivative at x in the direction v can be viewed as an operator on smooth functions, Dx,v : C ∞ (Rn ) → R,
Dx,v (f ) =
d f (x + tv). d t t=0
(2.6)
This operator is linear in f and, by the ordinary product rule, satisfies Dx,v (f h) = f (x)Dx,v (h) + h(x)Dx,v (f ),
f, h ∈ C ∞ (Rn ).
(2.7)
Thus a Euclidean tangent vector can be recognized not only as an arrow v, but also as a first-order operator that differentiates smooth functions at x. The derivation viewpoint keeps precisely this algebraic behavior and extends it to manifolds, where no global vector structure is available. Definition 11 (Tangent space [197, Sec. 8]). Let M be a smooth manifold and let p ∈ M. A derivation at p is a map v : C ∞ (M) → R satisfying:a (1) linearity,
a, b ∈ R,
v(af + bh) = av(f ) + bv(h),
f, h ∈ C ∞ (M),
(2.8)
(2) the Leibniz rule, v(f h) = f (p)v(h) + h(p)v(f ),
f, h ∈ C ∞ (M).
(2.9)
The set of all derivations at p is a vector space, called the tangent space of M at p and denoted by Tp M. a
Strictly speaking, the domain should be the algebra of germs of smooth functions at p, namely equivalence classes of smooth functions that agree on some neighborhood of p. For simplicity, we write C ∞ (M).
The tangent space Tp M is a vector space whose operations are pointwise addition and scalar multiplication of derivations. Let (U, φ) = (U, x1 , . . . , xn ) be a chart containing p. The chart induces the coordinate tangent vector ∂x∂ i p , which is the derivation at 14
Chapter 2. Mathematical Background p defined by ∂ ∂ (f ◦ φ−1 ) (f ) = (φ(p)) , ∂xi p ∂ri
f ∈ C ∞ (M),
i = 1, . . . , n.
(2.10)
Thus ∂x∂ i p differentiates f in the i-th Euclidean coordinate direction after f is expressed in the chart. The coordinate tangent vectors ∂x∂ 1 p , . . . , ∂x∂n p form a basis of Tp M [197, Prop. 8.9]. Consequently, every v ∈ Tp M has a unique coordinate representation v=
n X
vi
i=1
∂ , ∂xi p
v i ∈ R,
(2.11)
so this coordinate basis identifies Tp M with Rn . The differential of a smooth map generalizes the Jacobian matrix. Definition 12 (Differential [197, Sec. 8]). Let f : M → N be a smooth map and let p ∈ M. The differential of f at p is the linear map f∗,p : Tp M → Tf (p) N defined by (f∗,p (v)) (h) = v(h ◦ f ), v ∈ Tp M, h ∈ C ∞ (N ). (2.12) Throughout this chapter, we write the differential as f∗,p and its action on v as f∗,p (v). Later computations also use the notation dp f (v). The intrinsic differential recovers the ordinary Jacobian in local coordinates. Let (U, φ) = (U, x1 , . . . , xn ) be a chart around p ∈ M and let (V, ψ) = (V, y 1 , . . . , y m ) be a chart around f (p) ∈ N . The differential f∗,p is f∗,p
=
∂ ··· ∂x1 p
∂ ··· ∂y 1 f (p)
∂ ∂xn p
∂ (y 1 ◦ f ) ∂ (y 1 ◦ f ) (p) ∂x1 (p) · · · ∂xn ∂ . . .. . . . . . . ∂y m f (p) ∂ (y m ◦ f ) ∂ (y m ◦ f ) (p) · · · (p) ∂x1 ∂xn
(2.13)
The (i, j)-th entry of this matrix is ∂ (y i ◦ f ) /∂xj (p) for i = 1, . . . , m and j = 1, . . . , n. Hence, f∗,p is represented by the Jacobian matrix of the local coordinate representation ψ ◦ f ◦ φ−1 evaluated at φ(p) [197, Prop. 8.11]. The differential gives an intrinsic definition of the velocity of a curve. 15
2.3. Differential Geometry Definition 13 (Smooth curve and velocity [197, Sec. 8.6]). Let I ⊂ R be an open interval and let c : I → M be a smooth curve. The velocity vector of c at t0 ∈ I is c′ (t0 ) := c∗,t0
d d t t=t0
!
(2.14)
∈ Tc(t0 ) M.
Equivalently, for every h ∈ C ∞ (M), c′ (t0 )(h) =
d h(c(t)). d t t=t0
(2.15)
Proposition 14 (Velocity in local coordinates [197, Prop. 8.15]). Let c : I → M be a smooth curve and let (U, φ) = (U, x1 , . . . , xn ) be a chart containing c(t). Then the velocity is obtained by differentiating the coordinate functions: ′
c (t) =
n X d(xi ◦ c) i=1
dt
(t)
∂ . ∂xi c(t)
(2.16)
Thus the coefficients of c′ (t) in the chart-induced basis are precisely the ordinary derivatives of the coordinate representation φ ◦ c. Every smooth curve through p yields a tangent vector in Tp M, and conversely every tangent vector is the velocity of some smooth curve through p [197, Prop. 8.16]. Therefore, Tp M = {c′ (0) | ε > 0,
c : (−ε, ε) → M is smooth,
c(0) = p} .
(2.17)
For any curve c representing v ∈ Tp M in this way, the derivation v acts as the directional derivative v(h) = ddt t=0 h(c(t)) [197, Prop. 8.17]. Curves also provide a classical method for computing differentials. Proposition 15 (Differential via curves [197, Prop. 8.18]). Let F : M → N be a smooth map, let p ∈ M, and let v ∈ Tp M. If c : (−ε, ε) → M is any smooth curve satisfying c(0) = p and c′ (0) = v, then the differential maps the velocity of c to the velocity of its image curve: F∗,p (v) = (F ◦ c)′ (0) = 16
d F (c(t)). d t t=0
(2.18)
Chapter 2. Mathematical Background
In particular, this expression is independent of the chosen representative curve. The following examples illustrate the above proposition in the vector and matrix cases. Example 16 (Sphere). Let Sn = {y ∈ Rn+1 | ∥y∥ = 1} and consider the normalization map ν : Rn+1 \ {0} → Sn defined by ν(x) = x/ ∥x∥. For x ̸= 0 and v ∈ Rn+1 , the curve c(t) = x + tv has initial velocity v. Applying Eq. (2.18) gives d x + tv 1 ν∗,x (v) = = d t t=0 ∥x + tv∥ ∥x∥
⟨x, v⟩ x v− ∥x∥2
∈ Tν(x) Sn .
(2.19)
In particular, when x ∈ Sn , the differential is ν∗,x (v) = v − ⟨x, v⟩ x, the orthogonal projection of v onto Tx Sn . Example 17 (General linear group). Let GL(n) = {A ∈ Rn×n | det(A) ̸= 0} denote the manifold of invertible real n × n matrices. For G ∈ GL(n), define LG : GL(n) → GL(n) by LG (B) = GB. Since GL(n) is an open submanifold of Rn×n , its tangent spaces are identified with Rn×n . For X ∈ TIn GL(n) ∼ = Rn×n , the curve c(t) = In + tX remains in GL(n) for sufficiently small t and satisfies c′ (0) = X. Hence, (LG )∗,In (X) =
d G(In + tX) = GX, d t t=0
(2.20)
so the differential of left multiplication is again left multiplication [197, Ex. 8.19]. A vector field assigns a tangent vector to each point in a smooth way. Definition 18 (Vector field [197, Def. 12.7]). A vector field on a smooth manifold M is a map X : M → T M such that X(p) ∈ Tp M for every p ∈ M. It is smooth if, for every smooth function f ∈ C ∞ (M), the function p 7→ X(p)f is smooth. Tab. 2.2 summarizes the Euclidean special cases of the above differential-geometric concepts.
2.4
Riemannian Geometry
Riemannian geometry equips each tangent space with an inner product. This turns local tangent vectors into measurable directions and induces global geometric objects 17
2.4. Riemannian Geometry Differential-geometric concept
Euclidean special case
Smooth manifolds Charts Smooth map Tangent space
Rn . φ = idRn . Smooth map in calculus. Tx Rn ∼ = Rn .
Table 2.2: Euclidean prototypes for differential-geometric concepts.
such as lengths and distances. Definition 19 (Riemannian manifold [165, Defs. 3.1 and 3.2]). Let M be a smooth manifold. A Riemannian metric on M is a smooth assignment p 7→ gp : Tp M × Tp M → R
(2.21)
such that, for every p ∈ M, gp satisfies: (1) gp is bilinear,
(2) gp (v, w) = gp (w, v) for all v, w ∈ Tp M, (3) gp (v, v) > 0 for all nonzero v ∈ Tp M.
The pair (M, g) is called a Riemannian manifold. We write the value of the metric as ⟨v, w⟩p = gp (v, w) for v, w ∈ Tp M. The induced norm is q ∥v∥p = ⟨v, v⟩p ,
v ∈ Tp M.
(2.22)
The Riemannian metric extends the Euclidean inner product to curved spaces. Unless explicitly stated otherwise, a manifold means a Riemannian manifold, and the pair (M, g) is abbreviated as M. We also use ⟨v, w⟩p for the Riemannian metric. Once each tangent vector has a norm, one can measure the length of curves and define the induced shortest-path distance. Definition 20 (Length and geodesic distance [165, Defs. 5.11 and 5.15]). Let M be a manifold and let α : [a, b] → M be a piecewise smooth curve. The length of α is Z b
L(α) =
a
∥α̇(t)∥α(t) d t.
18
(2.23)
Chapter 2. Mathematical Background The geodesic distance between x, y ∈ M is dM (x, y) = inf L(α), α
(2.24)
where the infimum is taken over all piecewise smooth curves α : [a, b] → M satisfying α(a) = x and α(b) = y. The geodesic distance in Rn is exactly the straight-line distance. On a connected manifold M, the distance function dM : M × M → [0, ∞) is a metric, which makes (M, dM ) a metric space [165, Prop. 5.18]. The notion of a connection generalizes the directional derivative of tangent vectors to curved spaces. It is used to define geodesics, parallel transport, and curvature. Definition 21 (Connection and covariant derivative [165, Def. 3.9]). Let X(M) denote the space of smooth vector fields on M. A connection on M is a map D : X(M) × X(M) → X(M),
(X, Y ) 7→ DX Y,
(2.25)
that satisfies, for all X, Y, Z ∈ X(M), f1 , f2 , f ∈ C ∞ (M), and a, b ∈ R: (1) C ∞ (M)-linearity in the first argument,
Df1 X+f2 Y Z = f1 DX Z + f2 DY Z,
(2.26)
(2) R-linearity in the second argument, DX (aY + bZ) = aDX Y + bDX Z,
(2.27)
DX (f Y ) = X(f )Y + f DX Y.
(2.28)
(3) the Leibniz rule,
The vector field DX Y is called the covariant derivative of Y in the direction X. A Riemannian metric determines a canonical connection by requiring zero torsion and compatibility with the metric. Theorem 22 (Levi-Civita connection [165, Thm. 3.11]). Every Riemannian manifold (M, g) admits a unique connection D satisfying, for all smooth vector fields X, Y, Z ∈ X(M): 19
2.4. Riemannian Geometry (1) torsion-freeness, (2.29)
DX Y − DY X = [X, Y ], (2) compatibility with the metric, X (⟨Y, Z⟩) = ⟨DX Y, Z⟩ + ⟨Y, DX Z⟩ .
(2.30)
This connection is called the Levi-Civita connection. Example 23 (Euclidean Levi-Civita connection [165, Def. 3.8 and Lem. 3.14]). Let x1 , . . . , xn be the standard coordinates on Rn , and let n X
∂ X= X , ∂xi i=1 i
Y =
n X j=1
Yj
∂ ∂xj
(2.31)
be smooth vector fields. The Levi-Civita connection of the Euclidean metric is n X ∂Y j ∂ ∂ Xi i . DX Y = X(Y ) j = ∂x ∂x ∂xj i,j=1 j=1 n X
j
(2.32)
Thus, DX Y is the ordinary directional derivative of the component functions of Y along X. In particular, the standard coordinate vector fields are parallel: ∂ = 0, ∂x ∂xj
(2.33)
D ∂i for all i, j.
The Levi-Civita connection is the canonical connection of Riemannian geometry. It determines geodesics, parallel transport, and curvature. It also induces a covariant derivative for vector fields along a curve. Proposition 24 (Induced covariant derivative [165, Prop. 3.18]). Let M be a manifold with Levi-Civita connection D, let α : I → M be a smooth curve, and let X(α) denote the space of smooth vector fields along α. There is a unique map X(α) → X(α),
V 7→ V ′ =
D (V ), dt
called the induced covariant derivative along α, satisfying:
20
(2.34)
Chapter 2. Mathematical Background (1) (aV + bW )′ = aV ′ + bW ′ for V, W ∈ X(α) and a, b ∈ R. (2) (f V )′ = dd ft V + f V ′ for V ∈ X(α) and f ∈ C ∞ (I). (3) Let U ⊆ M be an open neighborhood of α(I), and let Y ∈ X(U ) be a smooth vector field on U . If V ∈ X(α) is the vector field along α obtained by restricting Y to the curve, namely V (t) = Yα(t) ∈ Tα(t) M, then V ′ (t) = Dα̇(t) Y,
(2.35)
where the right-hand side denotes the covariant derivative of the field Y . (4) For V, W ∈ X(α), d ⟨V, W ⟩α(t) = ⟨V ′ , W ⟩α(t) + ⟨V, W ′ ⟩α(t) . dt
(2.36)
Geodesics are the Riemannian counterparts of straight lines. Definition 25 (Geodesic [165, Ch. 3]). Let M be a manifold with Levi-Civita connection D. A smooth curve γ : I → M is a geodesic if Dγ̇ = Dγ̇ γ̇ = 0. dt
(2.37)
Consequently, every geodesic has constant speed : ∥γ̇(t)∥γ(t) is constant on I. Unlike straight lines in Euclidean space, geodesics on a general manifold need not be globally length-minimizing. They are locally length-minimizing curves [165, Lem. 5.14 and Prop. 5.16]. Geodesics also turn tangent vectors into manifold points through the exponential map, and locally turn nearby manifold points back into tangent vectors through the logarithmic map. Definition 26 (Exponential and logarithmic maps [165, Def. 3.29 and Prop. 3.30]). Let M be a manifold. For x ∈ M and v ∈ Tx M, let γx,v be the geodesic satisfying γx,v (0) = x and γ̇x,v (0) = v. The Riemannian exponential map at x is Expx (v) = γx,v (1),
(2.38)
where it is defined. As Expx is locally invertible around 0 ∈ Tx M, its local inverse 21
2.4. Riemannian Geometry is called the Riemannian logarithmic map at x and is denoted by Logx . This pair is used repeatedly in Riemannian neural networks to move between nonlinear manifold-valued data and linear tangent-space computations.
To compare or aggregate tangent vectors based at different manifold points, tangent vectors must be transported along curves. Definition 27 (Parallel transport [165, Prop. 3.19]). Let α : [a, b] → M be a smooth curve and let V (t) be a vector field along α. The field V is parallel along α if DV =0 (2.39) dt for all t ∈ [a, b]. Given v ∈ Tα(a) M, the endpoint V (b) ∈ Tα(b) M of the unique parallel vector field satisfying V (a) = v is called the parallel transport of v along α. When the connecting curve is the relevant geodesic from x to y, we denote parallel transport by PTx→y : Tx M → Ty M. Proposition 28 (Parallel transport is an isometry [165, Lem. 3.20]). Let V (t) and W (t) be parallel vector fields along a smooth curve α : [a, b] → M. Then ⟨V (t), W (t)⟩α(t)
(2.40)
is constant in t. Consequently, parallel transport along α defines a linear isometry between tangent spaces. The connection also measures the failure of second covariant derivatives to commute, which is encoded by the curvature tensor. Definition 29 (Curvature [165, Lem. 3.35]). Let M be a manifold with Levi-Civita connection D. The Riemannian curvature tensor is the map R : X(M) × X(M) × X(M) → X(M)
(2.41)
R(X, Y )Z = DX DY Z − DY DX Z − D[X,Y ] Z,
(2.42)
defined by
22
Chapter 2. Mathematical Background where [X, Y ] denotes the Lie bracket of vector fields: f ∈ C ∞ (M).
[X, Y ](f ) = X(Y (f )) − Y (X(f )),
(2.43)
Sectional curvature extracts from the curvature tensor the curvature of each twodimensional tangent plane. Definition 30 (Sectional curvature [165, Lem. 3.39]). Let M be a manifold with curvature tensor R. For a two-dimensional subspace σ ⊂ Tx M spanned by linearly independent vectors u, v ∈ Tx M, the sectional curvature of σ is Kx (σ) =
⟨R(u, v)v, u⟩x . ⟨u, u⟩x ⟨v, v⟩x − ⟨u, v⟩2x
(2.44)
It is independent of the choice of basis {u, v} for σ. Positive, zero, and negative curvature correspond to spherical, Euclidean, and hyperbolic behavior, which will be discussed later. A central consequence of non-positive curvature is that the exponential map can become globally invertible under suitable assumptions. Theorem 31 (Cartan–Hadamard theorem [165, Thm. 10.22]). Let M be a complete, connected, and simply connected manifold whose sectional curvature is everywhere non-positive. Then, for every x ∈ M, the exponential map Expx : Tx M → M is a global diffeomorphism. This theorem explains why several geometries in machine learning are algorithmically convenient, as their exponential maps allow tangent-space computations to be performed globally on the manifold. The distance also supports averaging manifold-valued data through Fréchet means. Definition 32 (Fréchet mean and variance [171, Sec. 2]). Let (X , d) be a metric PN space and let x1 , . . . , xN ∈ X with weights wi > 0 satisfying i=1 wi = 1. A weighted Fréchet mean is any minimizer WFM({wi }, {xi }) ∈ argminz∈X
N X
wi d(xi , z)2 .
(2.45)
i=1
When wi = N1 for all i, it is called the Fréchet mean, and we write FM({xi }) for 23
2.4. Riemannian Geometry Geometric concept
Euclidean special case
Riemannian metric Geodesic distance Geodesic Exponential map Logarithmic map Parallel transport Sectional curvature Weighted Fréchet mean Fréchet variance
⟨u, v⟩x = ⟨u, v⟩. d(x, y) = ∥x − y∥. γ(t) = (1 − t)x + ty. Expx (v) = x + v. Logx (y) = y − x. PTx→y = idRn . K ≡ 0 on Rn . P WFM({wi }, {xi }) = N i=1 wi xi . PN Var = minz i=1 wi ∥xi − z∥2 .
Table 2.3: Euclidean prototypes for Riemannian-geometric concepts.
the corresponding minimizer. The infimum of the objective is called the Fréchet variance. The objective above is defined on any metric space. On a manifold, the distance is usually the geodesic distance. If the data lie in a sufficiently small geodesic ball, the weighted Fréchet mean exists and is unique [2, Thm. 2.1]. Pullback metrics allow a complex manifold to inherit a computationally convenient geometry from a simpler prototype space. Definition 33 (Pullback and Riemannian isometry [165, Defs. 2.8 and 3.6]). Let M and N be smooth manifolds, let g be a Riemannian metric on N , and let f : M → N be smooth. The pullback of g by f is the symmetric tensor field f ∗ g on M defined by (f ∗ g)p (v, w) = gf (p) (f∗,p (v), f∗,p (w)) ,
p ∈ M,
v, w ∈ Tp M.
(2.46)
If f ∗ g is positive definite at every point, it is a Riemannian metric on M. In particular, when f is a diffeomorphism and M is equipped with the pullback metric f ∗ g, the map f : (M, f ∗ g) → (N , g) is a Riemannian isometry. A Riemannian isometry is the Riemannian version of a diffeomorphism: it preserves the metric, lengths, geodesic distances, geodesics, exponential maps, logarithmic maps, parallel transport, curvature, and distance-based quantities such as Fréchet means and variances [165, Ch. 3]. Tab. 2.3 summarizes the Euclidean special cases of the above geometric concepts. 24
Chapter 2. Mathematical Background
2.5
Metric Geometry
Metric geometry extends Riemannian geometry to the more general setting of metric spaces, where no differentiable structure is assumed.1 The theory develops geodesics and curvature without smooth structure, providing the metric tools used later for hyperbolic neural layers. We begin by recalling basic notions in metric spaces. Definition 34 (Metric space [29, Def. I.1.1]). A metric space is a pair (X , d) where X is a nonempty set and d : X × X → R satisfies, for all x, y, z ∈ X : (1) d(x, y) ≥ 0 and d(x, y) = 0 if and only if x = y, (2) d(x, y) = d(y, x), (3) d(x, z) ≤ d(x, y) + d(y, z). A connected Riemannian manifold becomes a metric space when equipped with its geodesic distance [165, Prop. 5.18]. The metric-space perspective abstracts away coordinates and focuses on distances, geodesics, and comparison geometry. Geodesics, rays, and lines generalize unit-speed minimizing geodesics to metric spaces. Definition 35 (Geodesic, geodesic ray, and geodesic line [29, Ch. I.1]). Let (X , d) be a metric space. A geodesic joining x to y is a continuous map γ : [0, l] → X with γ(0) = x and γ(l) = y such that d (γ(t), γ(t′ )) = |t − t′ |,
t, t′ ∈ [0, l].
(2.47)
A geodesic ray is a continuous map γ : [0, ∞) → X such that d (γ(t), γ(t′ )) = |t−t′ | for all t, t′ ≥ 0. A geodesic line is a continuous map γ : R → X such that d (γ(t), γ(t′ )) = |t − t′ | for all t, t′ ∈ R. Geodesic metric spaces abstract the requirement that every pair of points be joined by a distance-realizing geodesic. Definition 36 (Geodesic metric space [29, Ch. I.1]). The metric space (X , d) is a geodesic metric space, or more briefly a geodesic space, if every pair of points in X is joined by a geodesic. It is uniquely geodesic if there is exactly one geodesic 1
Here, the missing differentiable structure refers specifically to a smooth manifold structure. Metric analogues of geodesics and curvature can still be defined without it.
25
2.5. Metric Geometry joining x to y for all x, y ∈ X . Convexity is defined through geodesics, generalizing linear convexity in Euclidean space and geodesic convexity on manifolds. Definition 37 (Convex subset [29, Ch. I.1]). Let (X , d) be a metric space. A subset C ⊆ X is convex if every pair x, y ∈ C can be joined by a geodesic in X and the image of every such geodesic is contained in C. We next review concepts that extend curvature from manifolds to metric spaces. The reference spaces are the model spaces of constant curvature. Definition 38 (Model space [29, Ch. I.2]). For K ∈ R, the model space (MKn , dK ) is given by n √1 d , K > 0, S , K (2.48) (MKn , dK ) = (Rn , d) , K = 0, Ln , √ 1 d , K < 0, −K
where d is the geodesic distance in the corresponding manifold. The unit sphere Sn and unit Lorentz manifold Ln are reviewed in Sec. 2.9.5. The diameter of MK2 is denoted by √ π/ K, K > 0, (2.49) DK = ∞, K ≤ 0.
Definition 39 (Comparison triangle [29, Lem. I.2.14 and Sec. II.1]). Let (X , d) be a geodesic metric space and let △(x, y, z) be a geodesic triangle in X with side lengths a = d(y, z), b = d(x, z), and c = d(x, y). A comparison triangle for △(x, y, z) in MK2 is a triangle △(x̄, ȳ, z̄) such that dK (ȳ, z̄) = a, dK (x̄, z̄) = b, and dK (x̄, ȳ) = c. When a + b + c < 2DK , the comparison triangle exists. If p lies on the side from x to y, then a comparison point for p is the point p̄ on the side from x̄ to ȳ such that d(x, p) = dK (x̄, p̄) and d(y, p) = dK (ȳ, p̄). Comparison points on the other two sides are defined analogously.
CAT(K) spaces encode curvature through triangle comparison with the model plane MK2 . Intuitively, a CAT(K) space is a metric space whose triangles are thinner than the corresponding comparison triangles in MK2 .
26
Chapter 2. Mathematical Background Definition 40 (CAT(K) space [29, Def. II.1.1]). Let (X , d) be a metric space and let K ∈ R. Let ∆ be a geodesic triangle in X with perimeter less than 2DK , and ¯ ⊂ M 2 be a comparison triangle for ∆. The triangle ∆ satisfies the CAT(K) let ∆ K ¯ inequality if, for all p, q ∈ ∆ and all corresponding comparison points p̄, q̄ ∈ ∆, d(p, q) ≤ dK (p̄, q̄).
(2.50)
Then X is called a CAT(K) space as follows:
(1) If K ≤ 0, then X is a CAT(K) space if X is a geodesic space all of whose geodesic triangles satisfy the CAT(K) inequality. (2) If K > 0, then X is a CAT(K) space if X is DK -geodesic and all geodesic triangles in X of perimeter less than 2DK satisfy the CAT(K) inequality.
Here, DK -geodesic means that for every pair of points x, y ∈ X with d(x, y) < DK , there is a geodesic joining x to y.
Definition 41 (Hadamard space [29, p. 159]). A Hadamard space is a complete CAT(0) space. In particular, a Hadamard manifold, a complete, simply connected Riemannian manifold with non-positive sectional curvature, is a Hadamard space [29, Thm. II.4.1]. Proposition 42 (Orthogonal projection [29, Prop. II.2.4]). Let (X , d) be a CAT(0) space and let C ⊆ X be a convex subset that is complete in the induced metric. For every x ∈ X , there exists a unique point πC (x) ∈ C such that d (x, πC (x)) = inf d(x, y) = d(x, C). y∈C
(2.51)
If x′ belongs to the geodesic segment [x, πC (x)], then πC (x′ ) = πC (x).
(2.52)
The map πC : X → C is called the orthogonal projection, or simply the projection. With geodesics, one can define asymptotic rays, boundary points, Busemann functions, horoballs, and horospheres in metric spaces.
27
2.5. Metric Geometry Definition 43 (Asymptotic rays and boundary points [29, Def. II.8.1]). Let (X , d) be a metric space. Two geodesic rays γ, η : [0, ∞) → X are asymptotic if there exists C ≥ 0 such that d (γ(t), η(t)) ≤ C for all t ≥ 0.
(2.53)
The set ∂X of boundary points, also called points at infinity or ideal points, is the set of equivalence classes of geodesic rays, where two geodesic rays are equivalent if and only if they are asymptotic. Asymptotic rays generalize parallel lines. In hyperbolic geometry, they point toward the same ideal boundary point, which becomes a direction for defining Busemann functions. Definition 44 (Busemann function, horoball, and horosphere [29, Def. II.8.17]). Let (X , d) be a metric space and let γ : [0, ∞) → X be a geodesic ray. If the limit exists, the Busemann function associated with γ is B γ (x) = lim (d (x, γ(t)) − t) , t→∞
x ∈ X.
(2.54)
For τ ∈ R, the sublevel set HBτγ = {x ∈ X | B γ (x) ≤ τ }
(2.55)
is a horoball, and the level set Hτγ = {x ∈ X | B γ (x) = τ }
(2.56)
is a horosphere. In Hadamard spaces, the Busemann limit exists [29, Lem. II.8.18]. Moreover, Busemann functions associated with asymptotic rays agree up to an additive constant. Corollary 45 (Busemann functions of asymptotic rays [29, Cor. II.8.20]). If X is a Hadamard space, then the Busemann functions associated with asymptotic rays in X are equal up to addition of a constant. In Euclidean space, if γ(t) = tv with ∥v∥ = 1, then B γ (x) = − ⟨x, v⟩. Thus the Busemann function provides an intrinsic generalization of the Euclidean inner product, up to sign. Therefore, horospheres generalize Euclidean hyperplanes. Tab. 2.4
28
Chapter 2. Mathematical Background Metric-geometric concept
Euclidean special case
Geodesic, geodesic ray, and geodesic line Asymptotic geodesic rays Busemann function −B γ (x) Horosphere Hτγ Horospheres associated with asymptotic rays
Straight segment, ray, and line. Parallel rays. Inner product ⟨x, v⟩. Hyperplane. Parallel hyperplanes.
Table 2.4: Euclidean prototypes for metric-geometric concepts.
summarizes the Euclidean special cases of the metric-geometric concepts.
2.6
Algebraic Structures on Manifolds
Euclidean neural networks rely on vector addition and scalar multiplication to combine, translate, and rescale features. Manifold-valued representations generally lack a global linear structure, which motivates algebraic structures on manifolds that play analogous roles. This section reviews groups, Lie groups, gyrogroups and gyrovector spaces, which generalize addition, subtraction, and scalar multiplication to curved spaces. Definition 46 (Group [128, Ch. I.2]). A group is a nonempty set G equipped with a binary operationa ⊕ : G × G → G satisfying, for all x, y, z ∈ G: (1) associativity, x ⊕ (y ⊕ z) = (x ⊕ y) ⊕ z,
(2) identity, there exists e ∈ G, called the identity element or neutral element, such that e ⊕ x = x ⊕ e = x, (3) inverse, for every x ∈ G, there exists ⊖x ∈ G such that (⊖x)⊕x = x⊕(⊖x) = e. a
The group operation is usually written multiplicatively and called multiplication, as it is usually noncommutative. For a commutative group, it is often written additively and called addition. In this thesis, we use ⊕ for simplicity.
Definition 47 (Abelian group [128, Ch. I.2]). A group (G, ⊕) is an abelian group, or commutative group, if x ⊕ y = y ⊕ x for all x, y ∈ G. Lie groups are simultaneously algebraic and geometric objects. They add smooth structure to groups and require the group operations to be smooth. 29
2.6. Algebraic Structures on Manifolds Definition 48 (Lie group [197, Def. 6.20]). A Lie group is a smooth manifold G equipped with a binary operation ⊕ : G × G → G satisfying: (1) (G, ⊕) is a group,
(2) the multiplication map (x, y) 7→ x ⊕ y is smooth, (3) the inversion map x 7→ ⊖x is smooth. Definition 49 (Abelian Lie group). A Lie group (G, ⊕) is an abelian Lie group, or commutative Lie group, if its underlying group is abelian, namely x ⊕ y = y ⊕ x for all x, y ∈ G. Example 50 (Matrix Lie groups). The general linear group GL(n) introduced in Thm. 17 is a Lie group under matrix multiplication. Standard matrix Lie subgroups of GL(n) include the special linear group, the orthogonal group, and the special orthogonal group: SL(n) = {A ∈ GL(n) | det(A) = 1} , (2.57) O(n) = A ∈ GL(n) | A⊤ A = In , SO(n) = O(n) ∩ SL(n).
Gyrogroups relax the associativity axiom of groups. The deviation from associativity is controlled by gyrations through the gyroassociative law. Definition 51 (Gyrogroup [200, Def. 2.7]). Given a nonempty set G with a binary operation ⊕ : G × G → G, (G, ⊕) forms a gyrogroup if its binary operation satisfies the following axioms for any x, y, z ∈ G: (1) Left identity: there exists at least one element e ∈ G, called a left identity or neutral element, such that e ⊕ x = x. (2) Left inverse: there exists an element ⊖x ∈ G, called a left inverse of x, such that ⊖x ⊕ x = e. (3) Left gyroassociative law: there exists an automorphism gyr[x, y] : G → G for each x, y ∈ G such that x ⊕ (y ⊕ z) = (x ⊕ y) ⊕ gyr[x, y]z.
(2.58)
The automorphism gyr[x, y] is called the gyroautomorphism, or the gyration 30
Chapter 2. Mathematical Background of G generated by x, y. (4) Left reduction law: gyr[x, y] = gyr[x ⊕ y, y].
(2.59)
If every gyration is the identity map, the gyroassociative law reduces to ordinary associativity. Hence, gyrogroups naturally generalize groups. Definition 52 (Gyrocommutative gyrogroup [200, Def. 2.8]). A gyrogroup (G, ⊕) is gyrocommutative if it satisfies x ⊕ y = gyr[x, y](y ⊕ x) (gyrocommutative law).
(2.60)
Gyrocommutativity replaces commutativity. It says that exchanging two operands is possible after applying the appropriate gyration. Similarly, a gyrovector space generalizes a vector space. Definition 53 (Gyrovector space [54]). A gyrocommutative gyrogroup (G, ⊕) equipped with a scalar gyromultiplication ⊙ : R × G → G is called a gyrovector space if it satisfies the following axioms for s, t ∈ R and x, y, z ∈ G: (1) Identity scalar multiplication:
1 ⊙ x = x.
(2.61)
(s + t) ⊙ x = s ⊙ x ⊕ t ⊙ x.
(2.62)
(st) ⊙ x = s ⊙ (t ⊙ x).
(2.63)
gyr[x, y](t ⊙ z) = t ⊙ gyr[x, y]z.
(2.64)
(2) Scalar distributive law:
(3) Scalar associative law:
(4) Gyroautomorphism:
(5) Identity gyroautomorphism: gyr[s ⊙ x, t ⊙ x] = id, 31
(2.65)
2.6. Algebraic Structures on Manifolds where id is the identity map. Remark 54. Nguyen [157, Def. 2.3] presented a similar definition, except that the identity scalar multiplication axiom also includes 0 ⊙ x = t ⊙ e = e and (−1) ⊙ x = ⊖x. As implied by Ungar [200, Thm. 6.4], these conditions are redundant. Just as an inner product space augments a vector space with a compatible inner product, a real inner product gyrovector space augments a gyrovector space with an ambient inner product and corresponding compatibility axioms. Definition 55 (Real inner product gyrovector space [200, Def. 6.2]). Let (G, ⊕, ⊙) be a gyrovector space and let ⟨·, ·⟩ denote the Euclidean inner product on Rn with associated norm ∥·∥. We call (G, ⊕, ⊙, ⟨·, ·⟩) a real inner product gyrovector space if the following conditions hold. (1) G ⊆ Rn and inherits the inner product ⟨·, ·⟩ and norm ∥·∥. (2) Inner product gyroinvariance: ⟨gyr[x, y]u, gyr[x, y]v⟩ = ⟨u, v⟩ ,
∀x, y, u, v ∈ G.
(2.66)
∀s ∈ R \ {0}.
(2.67)
(3) Scaling property: x |s| ⊙ x = , ∥s ⊙ x∥ ∥x∥
∀x ∈ G \ {0},
(4) Let ∥G∥ = {± ∥x∥ | x ∈ G} ⊂ R. The set ∥G∥ forms a one-dimensional real vector space with respect to the vector addition and scalar multiplication induced by ⊕ and ⊙ on G. (5) Homogeneity property: ∥s ⊙ x∥ = |s| ⊙ ∥x∥ ,
∀x ∈ G,
∀s ∈ R.
(2.68)
(6) Gyrotriangle inequality: ∥x ⊕ y∥ ≤ ∥x∥ ⊕ ∥y∥ ,
32
∀x, y ∈ G.
(2.69)
Chapter 2. Mathematical Background Definition 56 (Gyrovector space isomorphisms [200, Def. 6.89]). Let (G1 , ⊕1 , ⊙1 ) and (G2 , ⊕2 , ⊙2 )
(2.70)
be real inner product gyrovector spaces. A map ϕ : G1 → G2 is a gyrovector space isomorphism if it is bijective and satisfies ∀x, y ∈ G1 ,
(2.71)
∀x ∈ G1 , ∀t ∈ R,
(2.72)
ϕ(x ⊕1 y) = ϕ(x) ⊕2 ϕ(y), ϕ(t ⊙1 x) = t ⊙2 ϕ(x),
and preserves the inner product of unit gyrovectors, ⟨x, y⟩ ⟨ϕ(x), ϕ(y)⟩ = , ∥ϕ(x)∥ ∥ϕ(y)∥ ∥x∥ ∥y∥
∀x, y ∈ G1 with x ̸= 0, y ̸= 0.
(2.73)
A useful property is that gyrovector space isomorphisms preserve the gyration, inverse, and identity. Proposition 57. [↓] Let (G1 , ⊕1 , ⊙1 ) and (G2 , ⊕2 , ⊙2 ) be real inner product gyrovector spaces with gyrations gyr1 and gyr2 , respectively. If ϕ : G1 → G2 is a gyrovector space isomorphism, then for all x, y, z ∈ G1 , ϕ (gyr1 [x, y]z) = gyr2 [ϕ(x), ϕ(y)]ϕ(z),
(2.74)
ϕ(e1 ) = e2 ,
(2.75)
ϕ(⊖1 x) = ⊖2 ϕ(x),
(2.76)
where e1 and e2 are the gyro identities in G1 and G2 , respectively. Intuitively, a gyrovector space generalizes a vector space to curved spaces. The gyrooperations associated with a manifold M can be defined as follows. Given a predefined origin e ∈ M, we assume that the relevant exponential maps, logarithmic maps, and parallel transports are well-defined. Following Nguyen and Yang [159, Eqs. (1)–(3)], for x, y, z ∈ M and t ∈ R, these operations are defined as x ⊕ y = Expx (PTe→x (Loge (y))) , t ⊙ x = Expe (t Loge (x)) , ⊖x = Expe (− Loge (x)) , 33
(2.77) (2.78) (2.79)
2.7. Riemannian Optimization gyr[x, y]z = (⊖(x ⊕ y)) ⊕ (x ⊕ (y ⊕ z)).
(2.80)
The corresponding gyro inner product, gyronorm, and gyrodistance are ⟨x, y⟩gyr = ⟨Loge (x), Loge (y)⟩e , q ∥x∥gyr = ⟨x, x⟩gyr ,
dgyr (x, y) = ∥⊖x ⊕ y∥gyr .
(2.81) (2.82) (2.83)
If the above operations satisfy the axioms in Thm. 53, then (M, ⊕, ⊙) is a gyrovector space. In Euclidean space, the above definitions reduce to the ordinary vector-space structure.
2.7
Riemannian Optimization
This section briefly reviews Riemannian optimization, which addresses the following manifold-constrained optimization problems: (2.84)
min f (w).
w∈M
Here, M is a Riemannian manifold and f : M → R is a smooth objective function.
Riemannian gradient [165, Def. 3.47]. The Riemannian gradient of f at w ∈ M, denoted by gradw f ∈ Tw M, is the unique tangent vector satisfying ⟨gradw f, v⟩w = f∗,w (v),
∀v ∈ Tw M.
(2.85)
Eq. (2.85) characterizes the Riemannian gradient through the directional derivative f∗,w (v). This is the same characterization as in Euclidean space, where the Euclidean gradient satisfies f∗,w (v) = ⟨∇w f, v⟩ for w, v ∈ Rn . As in Rn , the Cauchy–Schwarz inequality shows that, among all unit tangent directions, the directional derivative is minimized by − gradw f / ∥gradw f ∥w whenever gradw f ̸= 0. Therefore, − gradw f is the direction of steepest descent. Riemannian optimizers. To solve the objective in Eq. (2.84), we take Stochastic Gradient Descent (SGD) [175] as an example to show how to generalize a Euclidean optimizer to the Riemannian setting [1]. At iteration t, let ft denote a stochastic or mini-batch objective associated with f . When M = Rm , Euclidean SGD updates an 34
Chapter 2. Mathematical Background iterate w(t) ∈ Rm as
w(t+1) = w(t) − αt ∇w(t) ft ,
(2.86)
where αt > 0 is the learning rate and ∇w(t) ft is the Euclidean gradient of ft . On a general manifold, the descent direction must belong to the tangent space at the current iterate, and ordinary vector addition cannot return that direction to the manifold. The corresponding Riemannian Stochastic Gradient Descent (RSGD) [25, Sec. 2.3] update is w(t+1) = Expw(t) (−αt gradw(t) ft ) , (2.87) whenever the exponential map is defined for this step. In practical algorithms, a retraction can replace the exact exponential map by a first-order approximation [27, Def. 3.47 and Sec. 4.3], but a systematic treatment of retractions is beyond the scope of this chapter. A widely used package is the Geoopt [125], which integrates manifoldvalued parameters and Riemannian optimizers into PyTorch, supporting RSGD and adaptive optimizers such as Riemannian Adam [18]. In-depth discussions of Riemannian optimization can be found in Absil et al. [1], Boumal [27]. Trivialization. Apart from solving Eq. (2.84) directly on the manifold, another approach is to parameterize manifold-valued variables with Euclidean parameters and thereby use Euclidean optimization, which is called trivialization [132]. Let ϕ : Rm → M be a smooth surjective map, let z ∈ Rm be a trainable Euclidean parameter, and set w = ϕ(z) ∈ M. The constrained objective can then be written as min f (ϕ(z)) .
z∈Rm
(2.88)
In practice, the map ϕ could be the exponential map or retraction.
2.8
Matrix Functions
This section reviews some matrix functions widely used on matrix manifolds.
2.8.1
Matrix Functions and Differentials
We denote the Euclidean space of n × n real symmetric matrices by S n and the SPD n manifold of n × n SPD matrices by S++ . Let ˚ I be an open interval of R and let f :˚ I → R be a smooth function. For any symmetric matrix S whose eigenvalues lie in 35
2.8. Matrix Functions ˚ I, the associated symmetric matrix function is defined by f : S 7−→ U f (Σ)U ⊤ ∈ S n , with S = U ΣU ⊤ as the eigendecomposition.
(2.89)
Its differential is known as the Daleckii–Krein formula: f∗,S (V ) = U Lf ⊛ U ⊤ V U U ⊤ , ∀V ∈ S n , f (σi )−f (σj ) , if σ ̸= σ i j σi −σj [Lf ]i,j = f ′ (σ ), otherwise
(2.90) (2.91)
i
where Lf is called the Loewner matrix, its (i, j)-th entry is defined in Eq. (2.91), and ⊛ denotes the Hadamard product. Three special cases are the matrix logarithm n n , which is its inverse, and → S n , the matrix exponential exp : S n → S++ log : S++ n n defined by Pθ (P ) = P θ . See Bhatia [20, → S++ the matrix power map Pθ : S++ Eqs. (2.38)–(2.40)] or Bhatia [21, Thm. V.3.3] for more details. Let Ln++ be the Cholesky space of lower triangular matrices with positive diagonal n entries. The Cholesky map is Chol : S++ → Ln++ with inverse Chol−1 (L) = LL⊤ . Let n n P ∈ S++ , L = Chol(P ), V ∈ TP S++ , and X ∈ TL Ln++ . For any square matrix A, set A1/2 = ⌊A⌋ + 12 D(A), where ⌊A⌋ is the strictly lower triangular part and D(A) is the diagonal matrix formed from the diagonal of A. As shown by Lin [137, Prop. 4], the differentials of Chol and Chol−1 are Chol∗,P (V ) = L L−1 V L−⊤ 1/2 ,
2.8.2
⊤ ⊤ Chol−1 ∗,L (X) = XL + LX .
(2.92)
Backpropagation Through Matrix Functions
Symmetric matrix functions. The above symmetric matrix functions can be backpropagated using the Daleckii–Krein formula. Although PyTorch supports automatic differentiation through eigendecomposition [169], its backward pass requires computing (σi − σj )−1 [110, Prop. 1], which may trigger numerical instability when two eigenvalues approach each other. Following Brooks et al. [31, Eq. (13)], we instead use the Daleckii– Krein expression in Eqs. (2.90) and (2.91) for the backward pass [21, Thm. V.3.3]. As shown in Eq. (2.91), the divided difference converges to the derivative f ′ (σi ) when two eigenvalues approach each other, making the expression numerically more stable. Cholesky decomposition. The backpropagation of the Cholesky decomposition has been studied by Murray [154]. In the experiments, we use torch.linalg.cholesky 36
Chapter 2. Mathematical Background and its automatic differentiation.
2.9
Example Manifolds
This section collects the concrete model spaces that recur in later chapters. Each manifold is first described by its underlying set and representation, and then summarized through the Riemannian and algebraic operators.
2.9.1
Symmetric Positive Definite Manifolds
The SPD manifold has shown great success in diverse applications [106, 37, 205, 141, n and the vector space of n × n 48, 123]. We denote the set of n × n SPD matrices by S++ n n forms an open real symmetric matrices by S . As shown by Arsigny et al. [9], S++ submanifold of the Euclidean space S n . We review five common Riemannian metrics n : the affine-invariant metric (AIM) [171], the log-Euclidean metric (LEM) [9], on S++ the power-Euclidean metric (PEM) [70], the log-Cholesky metric (LCM) [137], and the Bures–Wasserstein metric (BWM) [22]. LEM, AIM, and PEM are represented by the parameterized families (α, β)-LEM, (α, β)-AIM, and (θ, α, β)-EM, respectively. Their common (α, β) parameters refer to the following O(n)-invariantinner product on S n [196]: ⟨A, B⟩(α,β) = α ⟨A, B⟩ + β tr(A) tr(B), (2.93) where A, B ∈ S n and (α, β) ∈ ST = {(α, β) ∈ R2 | min(α, α + nβ) > 0}. The standard LEM and AIM are recovered from (α, β)-LEM and (α, β)-AIM, respectively, at (α, β) = (1, 0). Likewise, the standard θ-PEM is recovered from (θ, α, β)-EM at (α, β) = (1, 0). The (θ, α, β)-EM formulas below assume θ ̸= 0; their limit as θ → 0 is (α, β)-LEM. Tabs. 2.5 and 2.6 summarize the associated operators under these five geometries with the following notation. n n (1) General notation. Let P, Q ∈ S++ be SPD matrices, let V, W ∈ TP S++ be N n tangent vectors, and let {Pi }i=1 ⊂ S++ be a dataset. The norms induced by ⟨·, ·⟩(α,β) and the standard Frobenius inner product ⟨·, ·⟩ are denoted by ∥·∥(α,β) and ∥·∥, respectively.2 2
We use ∥·∥ for both the Frobenius norm of matrices and the ℓ2 norm of vectors, since both norms are induced by the standard Euclidean inner product on the ambient Euclidean space.
37
2.9. Example Manifolds (2) Symmetric matrix functions. Symmetric matrix functions and their differentials are reviewed in Sec. 2.8.1. In addition to the matrix logarithm and exponential recorded there, the SPD operator tables use the matrix power map Pθ (P ) = P θ , whose differential is obtained by specializing Eq. (2.90). (3) LCM. The Cholesky map, its inverse, the half-diagonal operator, and their differentials are reviewed in Sec. 2.8.1. Let L = Chol(P ) and K = Chol(Q). In the LCM column, the corresponding tangent vectors are X = Chol∗,P (V ) and Y = Chol∗,P (W ). We define the Log-Cholesky map by ψLC (P ) = ⌊L⌋+Dlog(D(L)), where Dlog(·) is the diagonal element-wise logarithm. The diagonal matrices K, L, X, and Y contain the diagonal entries of K, L, X, and Y , respectively. (4) Gyrovector. The three geometries in Tab. 2.5 also admit gyrovector spaces. As shown by Nguyen [157, Sec. 3.1] and Nguyen and Yang [159, Sec. 2.4], AIM, LEM, and LCM each induce a gyro-structure via Eqs. (2.77) to (2.83). For LEM and LCM, the gyrovector spaces reduce to vector spaces: P ⊕LE Q = exp (log(P ) + log(Q)) ,
t ⊙LE P = exp (t log(P )) ,
−1 −1 P ⊕LC Q = ψLC (ψLC (P ) + ψLC (Q)), t ⊙LC P = ψLC (tψLC (P )).
(2.94)
The above two vector additions are the Lie group operations in Tab. 2.5. For AIM, the gyroaddition and scalar gyromultiplication are P ⊕AI Q = P 1/2 QP 1/2 ,
t ⊙AI P = P t .
(2.95)
In particular, the AIM gyroaddition above is not the AIM Lie group operation; the latter is listed in Tab. 2.5. (5) BWM. The Lyapunov operator LP [V ] is defined by LP [V ]P + P LP [V ] = V.
(2.96)
For BWM parallel transport, we only present the case where P and Q are commuting matrices, namely P = U ΣU ⊤ and Q = U ∆U ⊤ , with Σ = diag(σi ) and ∆ = diag(δi ). (6) Completeness. BWM and (θ, α, β)-EM with θ ̸= 0 are geodesically incomplete, meaning that their exponential maps are not defined globally. For BWM, the exponential map is locally defined only for tangent vectors satisfying LP [V ] + In ∈ 38
Chapter 2. Mathematical Background Metric
(α, β)-LEM
(α, β)-AIM
LCM
Q⊕P
exp (log(P ) + log(Q))
KP K ⊤
Chol−1 (⌊L + K⌋ + KL)
gP (V, W )
log∗,P (V ), log∗,P (W )
(α,β)
⟨P −1 V, W P −1 ⟩
(α,β)
log(Q−1/2 P Q−1/2 )
∥log(P ) − log(Q)∥ P exp N1 N i=1 log(Pi )
d(P, Q) FM({Pi })
(α,β)
γ(t; P, Q)
exp [log(P ) + t(log(Q) − log(P ))]
P 1/2 log(P −1/2 QP −1/2 )P 1/2 P 1/2 exp P −1/2 V P −1/2 P 1/2
References
[9, 196]; Thm. 147
[171, 196]
ExpP V
∥ψLC (P ) − ψLC (Q)∥ P N −1 1 ψLC i=1 ψLC (Pi ) N
Karcher flow
(log∗,P )−1 [log(Q) − log(P )] exp log(P ) + log∗,P (V )
LogP Q
⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨XL−1 , YL−1 ⟩
(α,β)
−1 Chol−1 ∗,L [⌊K⌋ − ⌊L⌋ + L Dlog(L K)]
Chol−1 (⌊L⌋ + ⌊X⌋ + L exp (XL−1 )) n o t Chol−1 ⌊L⌋ + t(⌊K⌋ − ⌊L⌋) + LKt−1
P 1/2 (P −1/2 QP −1/2 )t P 1/2
[137]; Thm. 147
n . Table 2.5: Lie group structures and associated Riemannian operators on S++
Metric
(θ, α, β)-EM
BWM
gP (V, W )
1 ⟨(Pθ )∗,P (V ), (Pθ )∗,P (W )⟩(α,β) θ2
1 ⟨LP [V ], W ⟩ 2
d(P, Q)
1 |θ|
Qθ − P θ
LogP Q
[(Pθ )∗,P ]−1 Qθ − P θ
PTP →Q (V )
γ(t; P, Q)
[(Pθ )∗,Q ]−1 ◦ (Pθ )∗,P (V ) 1/θ P θ + (Pθ )∗,P (V )
References
[70, 196]
ExpP V
1/2 tr(P ) + tr(Q) − 2 tr (P Q)1/2
(α,β)
(1 − t)P θ + tQθ
(P Q)1/2 + (QP )1/2 − 2P hq i δi +δj ⊤ U U V U ij U ⊤ σi +σj P + V + LP [V ]P LP [V ]
At A⊤ t 1/2 At = (1 − t)P 1/2 + tP −1/2 P 1/2 QP 1/2
1/θ
[22, 196]
n . Table 2.6: Riemannian operators of (θ, α, β)-EM and BWM on S++ n . For (θ, α, β)-EM, the exponential map is locally defined only for tangent S++ n vectors satisfying P θ + (Pθ )∗,P (V ) ∈ S++ .
(7) Fréchet mean. Karcher flow [116] refers to the standard iterative solver for the Fréchet mean objective in Thm. 32.
2.9.2
Full-Rank Correlation Manifolds
Given a covariance matrix Σ, its correlation matrix is defined as C = Cor(Σ) = D(Σ)− /2 ΣD(Σ)− /2 , 1
1
(2.97)
where D(·) extracts the diagonal part of Σ as a diagonal matrix. This diagonal normalization yields a scale-invariant representation: for any positive diagonal matrix D, Cor(DΣD) = Cor(Σ). Hence, correlation matrices remove marginal scales and empha39
2.9. Example Manifolds
Figure 2.1: The black stars denote 2 × 2 correlation matrices, while the red, green, and blue dots denote corresponding SPD matrices. The black dots denote the boundary of the SPD cone. size pairwise dependencies rather than raw variances. Only recently have Riemannian structures been developed for correlation matrices. The space of n × n full-rank correlation matrices, denoted by Cor+ (n), forms a Riemannian manifold and can be identified as a quotient manifold of the SPD manifold [62, Thm. 1]. As illustrated in Fig. 2.1, each correlation matrix corresponds to a surface in the SPD manifold. However, this quotient geometry does not guarantee uniqueness or closed forms of the Riemannian logarithm and Fréchet mean [195, Sec. 1.1]. To address this limitation, recent advances introduced five convenient Riemannian metrics on Cor+ (n): the Euclidean–Cholesky metric (ECM) [195], log-Euclidean–Cholesky metric (LECM) [195], poly-hyperbolic–Cholesky metric (PHCM) [195], off-log metric (OLM) [191], and log-scaled metric (LSM) [191]. These metrics are pullback metrics from simpler prototype spaces: ECM, LECM, OLM, and LSM are induced from Euclidean spaces, while PHCM is induced from a product of hyperbolic open hemispheres. We first review the associated prototype spaces. (1) LT1 (n) is the affine space of n × n lower triangular matrices with unit diagonal. (2) LT0 (n) is the Euclidean space of n × n lower triangular matrices with null diagonal. (3) Ln is the manifold of n × n lower triangular matrices with positive diagonals and unit row ℓ2 -norm. (4) Hol(n) is the Euclidean space of n × n symmetric matrices with null diagonals. The 40
Chapter 2. Mathematical Background tangent space TC Cor+ (n) at C ∈ Cor+ (n) can be identified with Hol(n). (5) Row0 (n) is the Euclidean space of n × n symmetric matrices with null row sum. ECM. It is derived from LT1 (n) by Θ=D(Chol(·))−1 Chol(·)
1 − ⇀ Cor+ (n) − ↽ −− −− −− −− −− −− −− −− −− −− −− − − LT (n), Θ−1 =Cor ◦ Chol−1
(2.98)
where Θ(C) = D(Chol(C))−1 Chol(C) for any C ∈ Cor+ (n). Here, Chol(C) is the Cholesky decomposition C = Chol(C) Chol(C)⊤ and D(·) returns a diagonal matrix consisting of the input diagonals. As LT1 (n) = In + LT0 (n), ECM is essentially induced from the Euclidean space of LT0 (n). Proposition 58 (ECM). Let ϕEC (C) = ⌊Θ(C)⌋, where ⌊·⌋ returns a strictly lower triangular matrix. ECM over Cor+ (n) is the pullback metric from the Euclidean space LT0 (n) by ϕEC . LECM. It is defined by further pulling back ECM: log ◦Θ
0 −− ⇀ Cor+ (n) ↽ −− −− −− −− −− −− −− −− −− −− −− −− −− −− −− − − LT (n), (log ◦Θ)−1 =Cor ◦ Chol−1 ◦ exp
(2.99)
where log(·) : LT1 (n) −→ LT0 (n) is the matrix logarithm with the matrix exponential exp(·) as its inverse. OLM. It is derived from a permutation-invariant inner product over Hol(n) by Log◦ =off◦log
− ⇀ Cor+ (n) − ↽ −− −− −− −− −− − − Hol(n). ◦ Exp
(2.100)
For any symmetric hollow matrix H ∈ Hol(n), the operator D(H) returns a unique diagonal matrix such that Exp◦ (·) : Hol(n) ∋ H 7−→ exp (D(H) + H) ∈ Cor+ (n) is a diffeomorphism. As shown by Archakov and Hansen [6, Cor. 1 and Sec. 5], D(H) can be computed by the following exponentially converging algorithm: Dk+1 = Dk − log (D (exp (Dk + H))), with D0 = 0n×n as the zero matrix. LSM. It is derived from a permutation-invariant inner product over Row0 (n) by Log⋆
−− ⇀ Cor+ (n) ↽ −− −− −− −− −− −− −− − − Row0 (n). ⋆ Exp =Cor ◦ exp
(2.101)
For any correlation matrix C ∈ Cor+ (n), there exists a unique positive diagonal matrix D⋆ (C) such that Log⋆ (·) : Cor+ (n) ∋ C 7−→ log(D⋆ (C)CD⋆ (C)) ∈ Row0 (n) is a diffeo41
2.9. Example Manifolds morphism. As shown by Thanwerdas [191, Sec. 3.5], D⋆ (C) corresponds to the unique 1 zero of f : x ∈ Rn++ 7−→ Cx − , where Rn++ denotes the set of n-dimensional posix tive vectors and x1 = x11 , . . . , x1n . This equation can be solved by damped Newton’s method. Tab. 2.7 summarizes the associated prototype spaces and diffeomorphisms. Tabs. 2.8 and 2.9 summarize the vector operations and Riemannian operators for ECM, LECM, OLM, and LSM. We use the following notation and record the following remarks. (1) General notation. Let C, C ′ ∈ Cor+ (n) be correlation matrices, let {Ci }N i=1 ⊂ + + ∼ Cor (n), and let V, W ∈ TC Cor (n) = Hol(n) be tangent vectors. Let L = Chol(C). (2) ECM and LECM. For any K ∈ LT1 (n) and X, ξ ∈ LT0 (n), the maps and differentials involved in ECM and LECM are (2.102)
Θ(C) = D(L)−1 L, − 12
Θ−1 (K) = D(KK ⊤ ) log(K) =
exp(ξ) =
KK ⊤ D(KK ⊤ )
n−1 X (−1)k−1 k=1 n−1 X
k
− 12
,
(K − In )k ,
1 k ξ , k! k=0
(2.103) (2.104) (2.105)
1 (2.106) Θ∗,C (V ) = Θ(C) L−1 V L−⊤ 1 − D L−1 V L−⊤ Θ(C), 2 2 (Θ∗,C )−1 (ξ) = Lξ ⊤ − CD Lξ ⊤ D(L) + D(L) ξL⊤ − D Lξ ⊤ C , (2.107) n−1 i X (−1)k−1 h log∗,K (ξ) = (K − In )k−1 ξ + · · · + ξ (K − In )k−1 , (2.108) k k=1 exp∗,X (ξ) =
n−1 X 1
k! k=1
X k−1 ξ + X k−2 ξX + · · · + ξX k−1 ,
(log ◦Θ)∗,C (V ) = log∗,Θ(C) (Θ∗,C (V )) .
(2.109) (2.110)
Due to the nilpotency of LT0 (n), the matrix logarithm over LT1 (n) and exponentiation over LT0 (n) are free from eigendecomposition. Although the Euclidean inner product in the ECM and LECM columns can be any inner product, we use the canonical one in this thesis. (3) OLM and LSM. Let H, W ∈ Hol(n), S = H + D(H), Σ = D⋆ (C)CD⋆ (C), X = Log⋆ (C) = log(Σ) ∈ Row0 (n), and Y ∈ Row0 (n). The involved maps and 42
Chapter 2. Mathematical Background differentials are [191, Thms. 2.4 and 4.1]: Log◦∗,C (V ) = off log∗,C (V ) ,
(2.111)
Exp◦∗,H (W ) = exp∗,S (W + D∗,H (W )) , 0 −1 D∗,H (W ) = − diag H D exp∗,S (W ) 1 , X n H 0 = [Hil0 ] ∈ S++ , Hil0 = Uij Uik Ulj Ulk [Lexp ]j,k ,
(2.112)
Log⋆∗,C (V ) = log∗,Σ
(2.115)
(2.113) (2.114)
j,k
1 0 0 V Σ + ΣV , ∆V ∆ + 2 1 ⋆ −1 exp∗,X (Y ) − ∆−2 D exp∗,X (Y ) Σ Exp∗,X (Y ) = ∆ 2 −2 −1 ∆ , +ΣD exp∗,X (Y ) ∆
(2.116) (2.117)
where S = U diag (λ1 , . . . , λn ) U ⊤ , Lexp is the Loewner matrix of exp∗,S , and 1 is the vector of all ones. Here, log∗ and exp∗ are given by the Daleckii–Krein formula in Eqs. (2.90) and (2.91), while diag(·) : Rn → Diag(n) returns a diagonal matrix from an input vector. The remaining auxiliary quantities are 1
∆ = D(Σ) 2 ,
V 0 = −2 diag (In + Σ)−1 ∆V ∆1 .
(2.118)
(4) Permutation. Let Sn be the group of permutation matrices Pσ = δi,σ(j) 1≤i,j≤n associated with permutations σ, and let D± (n) = {diag (ε1 , . . . , εn ) | ε ∈ {−1, 1}n } be the group of diagonal matrices with entries in {−1, 1}. Thanwerdas [191, Thm. 1.1] showed that the largest congruence action on full-rank correlation matrices is the action of signed permutation matrices: ⋆ : (A, C) ∈ S± (n) × Cor+ (n) 7−→ ACA⊤ ∈ Cor+ (n), S± (n) = D± (n)Sn .
(2.119)
As both Log⋆∗ and Log◦∗ are permutation-equivariant [191, Thms. 2.2(1) and 3.6(1)], permutation-invariant metrics over the correlation manifold can be induced by permutation-invariant inner products over Hol(n) and Row0 (n), respectively.
(5) Invariant inner products on Hol(n). For n ≥ 4, permutation-invariant inner 43
2.9. Example Manifolds products on Hol(n) are [190, Thm. 8.7] ⟨X1 , X2 ⟩(α,β,γ) = α tr(X1 X2 ) + β Sum (X1 X2 ) + γ Sum(X1 ) Sum(X2 ),
∀X1 , X2 ∈ Hol(n),
(2.120)
with α > 0, 2α + (n − 2)β > 0, and α + (n − 1)(β + nγ) > 0. For n = 3, permutation-invariant inner products have the same form with α = 0: ⟨X1 , X2 ⟩(α,β,γ) = β Sum(X1 X2 ) + γ Sum(X1 ) Sum(X2 ), with β > 0 and β + 3γ > 0.
(2.121)
For n = 2, they have the same form with α = β = 0: ⟨X1 , X2 ⟩(α,β,γ) = γ Sum(X1 ) Sum(X2 ),
with γ > 0.
(2.122)
(6) Invariant inner products on Row0 (n). For n ≥ 4, permutation-invariant inner products on Row0 (n) are [191, Thm. 4.2] ⟨Y1 , Y2 ⟩(α,δ,ζ) = α tr(Y1 Y2 ) + δ tr(D(Y1 )D(Y2 )) + ζ tr(Y1 ) tr(Y2 ),
∀Y1 , Y2 ∈ Row0 (n),
(2.123)
with α > 0, nα + (n − 2)δ > 0, and nα + (n − 1)(δ + nζ) > 0. For n = 3, the permutation-invariant inner products have the same form with α = 0. For n = 2, they have the same form with α = δ = 0. (7) OLM and LSM invariance. Combining the permutation-equivariant diffeomorphisms with the above invariant inner products gives permutation-invariant OLM and LSM. As shown by Thanwerdas [191, Thm. 2.7], OLM is further invariant under signed permutations when β = γ = 0, in which case the associated ⟨·, ·⟩(α,0,0) reduces to the scaled canonical Euclidean inner product: ⟨V, W ⟩(α,0,0) = α ⟨V, W ⟩ ,
∀V, W ∈ Hol(n).
(2.124)
In this thesis, we assume that ⟨·, ·⟩(α,β,γ) and ⟨·, ·⟩(α,δ,ζ) are the canonical Euclidean inner products. For n ≤ 3, these inner products remain permutation invariant; the dimension-specific forms above account for redundancies among the displayed trace and sum terms. 44
Chapter 2. Mathematical Background (8) Inverse consistency. We briefly review inverse-consistency, a property exclusive to LSM. The cor-inversion is defined as I : Cor+ (n) ∋ C 7−→ Cor (C −1 ) ∈ Cor+ (n) n [191, Def. 1.4]. It corresponds to the matrix inversion inv : S++ ∋ Σ 7−→ Σ−1 ∈ n , as represented by the following commuting diagram: S++ inv
n S++
n S++
Cor
Cor I
Cor+ (n)
(2.125)
Cor+ (n)
As shown by Thanwerdas [191, Thm. 1.7], LSM enjoys inverse-consistency: Log⋆ (I(C)) = − Log⋆ (C),
∀C ∈ Cor+ (n).
(2.126)
(9) Vector structure. ECM, LECM, OLM, and LSM are all pulled back from Euclidean vector spaces. It is therefore natural to inherit vector addition and scalar multiplication from their prototype spaces. For each corresponding isometry ϕ, these operations take the form C ⊕ C ′ = ϕ−1 (ϕ(C) + ϕ(C ′ )) , with t ∈ R.
45
t ⊙ C = ϕ−1 (tϕ(C)) ,
(2.127)
2.9. Example Manifolds
Metric
Prototype space
ECM [195]
LT (n) = LT (n) + In
LECM [195]
LT0 (n)
1
0
Diffeomorphisms
Properties
Θ : C ∈ Cor+ (n) 7−→ D(Chol(C))−1 Chol(C) ∈ LT1 (n) Θ−1 = Cor ◦ Chol−1 : LT1 (n) −→ Cor+ (n)
Null curvature
log ◦Θ : Cor+ (n) −→ LT0 (n) (log ◦Θ)−1 = Cor ◦ Chol−1 ◦ exp : LT0 (n) −→ Cor+ (n)
Null curvature
OLM [191]
Hol(n)
Log : C ∈ Cor (n) 7−→ (off ◦ log)(C) ∈ Hol(n) (Log◦ )−1 = Exp◦ : H ∈ Hol(n) 7−→ exp (D(H) + H) ∈ Cor+ (n)
LSM [191]
Row0 (n)
Log⋆ : C ∈ Cor+ (n) 7−→ log(D⋆ (C)CD⋆ (C)) ∈ Row0 (n) (Log⋆ )−1 = Exp⋆ : R ∈ Row0 (n) 7−→ Cor(exp(R)) ∈ Cor+ (n)
Permutation-invariance Inverse-consistency Null curvature
PHCM [195]
PHSn−1
Chol : Cor+ (n) −→ Ln ∼ = PHSn−1 Chol−1 : Ln ∼ = PHSn−1 −→ Cor+ (n)
Nonpositive sectional curvature
◦
+
Permutation-invariance Null curvature
Table 2.7: Isometric prototype spaces and diffeomorphisms on the correlation manifold.
Operation ′
C ⊕C t⊙C gC (V, W ) ExpC (V ) LogC (C ′ ) γ(t; C, C ′ ) d(C, C ′ ) Fréchet mean Curvature PTC→C ′ (V )
ECM
ϕ (C) + ϕ (C ) (ϕ ) (ϕEC )−1 tϕEC (C) ⟨Θ∗,C (V ), Θ∗,C (W )⟩ Θ−1 (Θ(C) + Θ∗,C (V )) ′ Θ−1 ∗,Θ(C) (Θ(C ) − Θ(C)) −1 Θ ((1 − t)Θ(C) + tΘ(C ′ )) ∥Θ(C) − Θ(C ′ )∥ P Θ−1 N1 N i=1 Θ(Ci ) 0 (Θ∗,C ′ )−1 (Θ∗,C (V )) EC −1
EC
EC
′
LECM −1
(log ◦Θ) (log ◦Θ(C) + log ◦Θ(C ′ )) (log ◦Θ)−1 (t log ◦Θ(C)) ⟨(log ◦Θ)∗,C (V ), (log ◦Θ)∗,C (W )⟩ (log ◦Θ)−1 (log ◦Θ(C) + (log ◦Θ)∗,C (V )) ′ (log ◦Θ)−1 ∗,log ◦Θ(C) (log ◦Θ(C ) − log ◦Θ(C)) −1 (log ◦Θ) ((1 − t) log ◦Θ(C) + t log ◦Θ(C ′ )) ∥log ◦Θ(C) − log ◦Θ(C ′ )∥ P (log ◦Θ)−1 N1 N i=1 (log ◦Θ)(Ci ) 0 ((log ◦Θ)∗,C ′ )−1 ((log ◦Θ)∗,C (V ))
Table 2.8: Vector operations and Riemannian operators under ECM and LECM.
Operation
OLM
LSM
C ⊕ C′ t⊙C gC (V, W ) ExpC (V ) LogC (C ′ ) γ(t; C, C ′ ) d(C, C ′ )
Exp◦ (Log◦ (C) + Log◦ (C ′ )) Exp◦ (t Log◦ (C)) (α,β,γ) Log◦∗,C (V ), Log◦∗,C (W ) Exp◦ Log◦ (C) + Log◦∗,C (V ) Exp◦∗,Log◦ (C) (Log◦ (C ′ ) − Log◦ (C)) Exp◦ ((1 − t) Log◦ (C) + t Log◦ (C ′ )) ∥Log◦ (C) Log◦ (C ′ )∥(α,β,γ) − P ◦ Exp◦ N1 N Log (C ) i i=1 0 (Log◦∗,C ′ )−1 Log◦∗,C (V )
Exp⋆ (Log⋆ (C) + Log⋆ (C ′ )) Exp⋆ (t Log⋆ (C)) (α,δ,ζ) Log⋆∗,C (V ), Log⋆∗,C (W ) Exp⋆ Log⋆ (C) + Log⋆∗,C (V ) Exp⋆∗,Log⋆ (C) (Log⋆ (C ′ ) − Log⋆ (C)) Exp⋆ ((1 − t) Log⋆ (C) + t Log⋆ (C ′ )) ∥Log⋆ (C) Log⋆ (C ′ )∥(α,δ,ζ) − P ⋆ Exp⋆ N1 N Log (C ) i i=1 0 (Log⋆∗,C ′ )−1 Log⋆∗,C (V )
Fréchet mean Curvature PTC→C ′ (V )
Table 2.9: Vector operations and Riemannian operators under OLM and LSM.
46
Chapter 2. Mathematical Background PHCM [195, Def. 4.3 and Thm. 4.4]. It is defined through the Cholesky decomposition and a product of open-hemisphere models of hyperbolic space. For a correlation matrix C ∈ Cor+ (n), let L = Chol(C). For k = 2, . . . , n, define the nonzero part of the k-th row of L by HSk−1 = x ∈ Rk | ∥x∥ = 1, xk > 0 .
ℓk = (Lk1 , . . . , Lkk ) ∈ HSk−1 ,
(2.128)
Thus, Ln is identified with the product of n−1 open hemispheres, denoted by PHSn−1 = Qn−1 i 0 i=1 HS . The first row, for which L11 = 1 and HS = {1}, is trivial and omitted from the product. PHCM is the pullback by the Cholesky decomposition of the product Q i i HSi metric on n−1 ), where each αi is a positive weight and g HS denotes the i=1 (HS , αi g metric tensor on HSi . In particular, PHCM with all weights equal to 1 is called the canonical PHCM, on which we focus below. Given C ∈ Cor+ (n) and L = Chol(C) ∈ Ln , define n−1 Y 1 n−1 n Ψ = ψ × ··· × ψ :L → HSi , (2.129) i=1
where
(2.130)
ψ i (L) = (Li+1,1 , . . . , Li+1,i+1 ) ∈ HSi . For Z ∈ TL Ln , its differential is the corresponding row extraction, i ψ∗,L (Z) = (Zi+1,1 , . . . , Zi+1,i+1 ) ∈ Tψi (L) HSi ,
n−1 1 Ψ∗,L = ψ∗,L × · · · × ψ∗,L .
(2.131)
The Riemannian operators under PHCM are obtained from the product geometry and the geometry of each HSi . Let C ′ ∈ Cor+ (n), L′ = Chol(C ′ ), and V, W ∈ TC Cor+ (n). i (Chol∗,C (V )). Then Write ai = ψ i (L), a′i = ψ i (L′ ), and ξi (V ) = ψ∗,L D
gC (V, W ) = D(L)−1 L L−1 V L−⊤ 1 , 2 E −1 −1 −⊤ D(L) L L W L , 1
(2.132)
2
ExpC (V ) = Chol
−1
LogC (C ′ ) = Chol−1 ∗,L γ(t; C, C ′ ) = Chol−1
−1
Ψ
!! 1 ExpHS a1 (ξ1 (V )) , . . . , n−1
ExpHS an−1 (ξn−1 (V ))
,
!! HS1 ′ Log (a ) , . . . , a1 1 (Ψ∗,L )−1 , HSn−1 Logan−1 a′n−1 !! HS1 ′ γ (t; a , a ) , . . . , 1 1 Ψ−1 , n−1 γ HS t; an−1 , a′n−1 47
(2.133) (2.134) (2.135)
2.9. Example Manifolds n−1 X
1 − ⟨ai , a′i ⟩ d(C, C ) = arccosh 1 + Li+1,i+1 L′i+1,i+1 i=1 ′ 2
i
i
2
.
(2.136)
Here, LogHS , ExpHS , and γ HS are the corresponding operators on HSi ; the distance and these closed forms follow from the hyperboloid–hemisphere isometry [195, Thm. 4.2].
2.9.3
i
Grassmannian Manifolds
The Grassmannian has been widely applied in machine learning, ranging from action recognition [108] to question answering [159], shape generation [224], image classification [204], and signal analysis [210]. The Grassmannian manifold is the set of p-dimensional subspaces of Rn [19]. It has two matrix representations: the projector perspective (PP) and the orthonormal-basis perspective (ONB): f n) = P ∈ S n | P 2 = P, rank(P ) = p , Gr(p, n n e ∈ St(p, n) | U e = U R, Gr(p, n) = [U ] | [U ] := U
oo R ∈ O(p) ,
(2.137)
where S n is the Euclidean space of symmetric matrices, St(p, n) is the Stiefel manifold, and O(p) is the orthogonal group. By abuse of notation, we use [U ] and U interchangeably for elements of Gr(p, n), where the n × p column-wise orthonormal matrix U is taken as a representative of an equivalence class. Helmke and Moore [100] show that the ONB perspective is diffeomorphic to the PP representation by f n). π : Gr(p, n) ∋ U 7→ U U ⊤ ∈ Gr(p,
(2.138)
As shown by Nguyen [157, Sec. 3.2] and Nguyen and Yang [159, Sec. 2.3.1], the Grassmannian admits gyro-structures defined by Eqs. (2.77) to (2.83). Under the ONB perspective, tangent vectors are represented by horizontal lifts. Given an orthogonal complement U⊥ ∈ St(n − p, n) of U , every tangent vector ∆ ∈ TU Gr(p, n) can be written as ∆ = U⊥ B,
B ∈ R(n−p)×p .
(2.139)
Under the PP perspective, every point can be written as P = OIep,n O⊤ with O ∈ O(n), 48
Chapter 2. Mathematical Background Operator
f n) PP: Gr(p,
ONB: Gr(p, n)
gU (∆, Ξ) or geP (A1 , A2 )
geP (A1 , A2 ) = 12 ⟨A1 , A2 ⟩
gU (∆, Ξ) = ⟨∆, Ξ⟩ ∥arccos(Σ)∥
d(U, V ) or e d(P, Q)
⊤
1 √ ∥log ((In − 2Q) (In − 2P ))∥ 2 2
SVD
U V := OΣR⊤
LogU V gP Q or Log
O arctan(Σ)R⊤ −1 SVD := OΣR⊤ In − U U ⊤ V U ⊤ V
γ(t; U, V ) or γ e(t; P, Q)
U R cos(tΣ)R⊤ + O sin(tΣ)R⊤
1 [log ((In − 2Q) (In − 2P )) , P ] 2
U R cos(Σ)R⊤ + O sin(Σ)R⊤
ExpU ∆ gPA or Exp
exp ([A, P ]) P exp (−[A, P ])
SVD
∆ := OΣR⊤ SVD
PTU →V (∆) ^ or PT P →Q (A)
LogU V := OΣR⊤ − sin(Σ) ⊤ UR O O + In − OO⊤ ∆ cos(Σ) SVD
g P Q, P ] A exp −[Log g P Q, P ] exp [Log
Karcher flow
Karcher flow
[71, 19]
[19]
LogU V := OΣR⊤
FM({Ui }) g or FM({P i }) References
g P Q, P ] P exp −t[Log g P Q, P ] exp t[Log
Table 2.10: Riemannian operators on the Grassmannian under ONB and PP. Operator
ONB: Gr(p, n)
Gyroaddition
U ⊕Gr V = exp(ΩU )V
Gyro identity
Ip,n
Scalar gyromultiplication
t ⊙Gr U = exp (tΩU ) Ip,n
Gyro-inverse
⊖Gr U = exp (−ΩU ) Ip,n
References
[159]
Gr
f n) PP: Gr(p,
e Q = exp(ΩP )Q exp (−ΩP ) P⊕ Iep,n
e Gr P = exp (tΩP ) Iep,n exp (−tΩP ) t⊙ e Gr P = exp (−ΩP ) Iep,n exp(ΩP ) ⊖ [157]
Table 2.11: Gyro operators on the Grassmannian under ONB and PP. f n) can be written as and every tangent vector A ∈ TP Gr(p, "
# 0 B⊤ A=O O⊤ , B 0
B ∈ R(n−p)×p .
(2.140)
Tab. 2.10 and Tab. 2.11 summarize the Riemannian and gyro operators, respectively; their formulas use the singular value decomposition (SVD). In these tables, U, V ∈ f n) and Gr(p, n) and ∆, Ξ ∈ TU Gr(p, n) refer to the ONB perspective, while P, Q ∈ Gr(p, f n) refer to the PP perspective. To distinguish the two perspectives, A, A1 , A2 ∈ TP Gr(p, g and Exp. g The differential PP Riemannian operators are marked by a tilde, such as Log g e (P ). Let of π is π∗,U (∆) = U ∆⊤ + ∆U ⊤ . For gyro operators, define P = Log Ip,n ⊤ ⊤ ⊤ e e e ΩU = [U U , Ip,n ] and ΩP = [P , Ip,n ]. The identities are Ip,n = [Ip , 0] and Ip,n = Ip,n Ip,n . 49
2.9. Example Manifolds Operator Expression References
R⊕S RS
gR (A1 , A2 )
d(R, S)
LogR S
⟨A1 , A2 ⟩
log(R⊤ S)
R log(R⊤ S) [146]
ExpR (A) R exp R⊤ A
γ(t; R, S) R exp t log(R⊤ S)
FM Karcher flow
Table 2.12: Lie group structures and Riemannian operators on rotation matrices.
2.9.4
Special Orthogonal Groups
The set of n × n rotation matrices forms a Lie group, known as the special orthogonal group and denoted by SO(n) [197]: SO(n) = R ∈ Rn×n | R⊤ R = In ,
det(R) = 1 .
(2.141)
Its group operation is the matrix product, with the identity matrix as the neutral element. Any tangent vector A ∈ TR SO(n) can be represented as A = RV , with V ∈ so(n). Here, so(n) is the Lie algebra of SO(n), which is the tangent space at the identity matrix, formed by the set of n × n skew-symmetric matrices: so(n) = Ω ∈ Rn×n | Ω⊤ = −Ω .
(2.142)
The Fréchet mean can be obtained by Karcher flow [146]. Furthermore, if all rotations lie in a closed ball of radius r < π/2, then Karcher flow converges to the unique mean [146, Thm. 5]. Given R, S ∈ SO(n) and tangent vectors A, A1 , A2 ∈ TR SO(n), Tab. 2.12 summarizes all the associated operators on SO(n). Across these matrix-manifold examples, the tables expose the Euclidean template behind the spaces used later: flat pullback metrics use ordinary addition and scaling in a chart, quotient metrics use horizontal representatives, product metrics act factor by factor, and Lie-group metrics use group translation.
2.9.5
Constant-Curvature Manifolds
Definition 59 (Constant-curvature space [165, Ch. 8]). A constant-curvature space (CCS) is a complete, simply connected, n-dimensional Riemannian manifold of constant curvature K. By O’Neill [165, Cor. 8.25], any two CCSs with the same dimension and the same curvature K are isometric. Thus, a CCS is determined up to isometry by K: Euclidean space corresponds to K = 0, spherical space to K > 0, and hyperbolic space to K < 0. In the following, we introduce several concrete models of these spaces that are used 50
Chapter 2. Mathematical Background later: the K-stereographic model, the K-radius model, and the Beltrami–Klein model. K-stereographic model [13]. This model has shown success in different applications, including computer vision [201], natural language processing [76, 180], graph learning [13, 85, 86], and astronomy [44]. It is defined as a model stnK with the conformal metric 2 K 2 λK , (2.143) ⟨u, v⟩st x = x = (λx ) ⟨u, v⟩ , 1 + K ∥x∥2 where K ∈ R is the constant curvature and λK x is a conformal factor. In particular, n n stK is the scaled R when K = 0. It unifies the spherical projected hypersphere DnK , Euclidean space Rn , and the hyperbolic Poincaré ball PnK : Dn = Rn , For K > 0, spherical geometry, K stnK = Rn , For K = 0, Euclidean geometry, Pn = x ∈ Rn | ∥x∥2 < − 1 , For K < 0, hyperbolic geometry. K K
(2.144)
Although DnK = Rn for K > 0, its metric is conformal to the Euclidean one. We abbreviate the K-stereographic model as the stereographic model. Bachmann et al. [13, Eqs. 2–3] show that this model admits a gyro-structure. K-radius model [182]. This model has been effective in various applications [38, 45, 15, 166, 99, 119]. It provides an extrinsic representation of the space with constant curvature K ∈ R, encompassing the sphere SnK , Euclidean space Rn , and the Lorentz, or hyperboloid, model LnK :
MnK =
SnK = x ∈ Rn+1 | ∥x∥2 = K1 ,
For K > 0, spherical geometry,
Rn , For K = 0, Euclidean geometry, Ln = x ∈ Rn+1 | ⟨x, x⟩ = 1 , x > 0 , For K < 0, hyperbolic geometry, t K L K
where ⟨x, x⟩L = ∥xs ∥2 − x2t is the Lorentzian quadratic form. Following the conventions ⊤ of the hyperboloid, we write x = (xt , x⊤ s ) , where xt ∈ R is the time component and xs ∈ Rn is the spatial component [173]. When K ̸= 0, the model can be written compactly as 1 n n+1 MK = x ∈ R | ⟨x, x⟩K = , K
⟨·, ·⟩K =
⟨·, ·⟩ ,
K > 0,
⟨·, ·⟩ , K < 0, L
(2.145)
with the hyperbolic branch restricted to xt > 0. We abbreviate the K-radius model 51
2.9. Example Manifolds Operator Gyroaddition Gyro identity Scalar gyromultiplication Gyro-inverse Gyration References
x ⊕K y =
Stereographic model stnK 1 − 2K ⟨x, y⟩ − K ∥y∥2 x + 1 + K ∥x∥2 y
1 − 2K ⟨x, y⟩ + K 2 ∥x∥2 ∥y∥2 0 p |K| ∥x∥ tanK t tan−1 K x p t ⊙K x = ∥x∥ |K| ⊖K x = −x 2 (Ast x + Bst y) gyr[x, y]z = z + Dst [200]
Beltrami–Klein model KnK 1 1 γxK x ⊕E y = x+ Ky −K ⟨x, y⟩ x 1 − K ⟨x, y⟩ γx 1 + γxK 0 √ −K ∥x∥ tanh t tanh−1 x √ t ⊙E x = ∥x∥ −K ⊖E x = −x AE x + B E y gyr[x, y]z = z + DE [200]
Table 2.13: Gyro operators on stereographic and Beltrami–Klein models. as the radius model. On its negative-curvature branch, we use MnK = LnK = HnK M M interchangeably; in hyperbolic-only contexts, the specializations of ⊕M K , ⊖K , and ⊙K are denoted by ⊕L , ⊖L , and ⊙L , respectively. Its gyro-structure is introduced later in Sec. 3.3. Beltrami–Klein model. There are five models of hyperbolic space [33]. Apart from the above Poincaré ball and hyperboloid models, we further study the Beltrami– Klein model: K ⟨x, v⟩ ⟨x, w⟩ 1 ⟨v, w⟩ 2 n n KK = x ∈ R | ∥x∥ < − , with gxK (v, w) = 2 , 2 − K 1 + K ∥x∥ 1 + K ∥x∥2 where K < 0 is the constant curvature and g K is its Riemannian metric. Although the Poincaré ball and Beltrami–Klein models share the same underlying set, their Riemannian metrics differ. This model admits an Einstein gyrovector space [200, Sec. 6.18].
The gyro-structures for the stereographic and Beltrami–Klein models are summarized in Tab. 2.13. In the table, x, y, z ∈ stnK for the stereographic model, x, y, z ∈ KnK −1/2 for the Beltrami–Klein model, t ∈ R, and γxK = 1 + K ∥x∥2 is the Einstein gamma factor. The stereographic convention is tanK = tanh for K < 0 and tanK = tan for K > 0. The stereographic gyration in Tab. 2.13, following Bachmann et al. [13, App. C.2.6], is gyr[x, y]z = z + 2 with
Ast x + Bst y , Dst
(2.146)
Ast = −K 2 ⟨x, z⟩ ∥y∥2 − K ⟨y, z⟩ + 2K 2 ⟨x, y⟩ ⟨y, z⟩ ,
Bst = −K 2 ⟨y, z⟩ ∥x∥2 + K ⟨x, z⟩ ,
Dst = 1 − 2K ⟨x, y⟩ + K 2 ∥x∥2 ∥y∥2
= (1 − K ⟨x, y⟩)2 + K 2 ∥x∥2 ∥y∥2 − ⟨x, y⟩2 ≥ 0. 52
(2.147)
Chapter 2. Mathematical Background The Cauchy–Schwarz inequality gives the last relation. The gyration formula applies whenever Dst > 0; singular positive-curvature configurations with Dst = 0 are excluded. The Einstein gyration coefficients for the Beltrami–Klein model are 2 γxK γyK − 1 ⟨x, z⟩ − KγxK γyK ⟨y, z⟩ AE = K K γx + 1 2 K 2 γxK γy 2 ⟨x, y⟩ ⟨y, z⟩ , + 2K (γxK + 1) γyK + 1 γyK γxK γyK + 1 ⟨x, z⟩ + γxK − 1 γyK ⟨y, z⟩ , BE = K K γy + 1
(2.148)
K DE = 1 + γxK γyK (1 − K ⟨x, y⟩) = 1 + γx⊕ . Ey
Tab. 2.14 summarizes the associated Riemannian operators for the stereographic and radius models when K ̸= 0; at K = 0, they reduce to the usual Euclidean operators. On the sphere, logarithmic maps and parallel transports are restricted away from antipodal pairs. The gyro operators of the radius model and the closed-form Riemannian operators of the Beltrami–Klein model are introduced later in Sec. 3.3. In Tab. 2.14, for K ̸= 0, the curvature-aware trigonometric functions are tanK (·) =
tan(·),
K > 0,
sinK (·) =
sin(·),
K > 0,
tanh(·), K < 0, sinh(·), K < 0, cos(·), K > 0, cosK (·) = cosh(·), K < 0,
All ratios in these tables are understood by continuous extension at removable zero denominators: in particular, t ⊙K 0 = 0, Expx (0) = x, and Logx (x) = 0. On the positive-curvature branch, logarithmic maps and parallel transports are restricted away from antipodal pairs and other stated singular configurations. For the radius model, if p x ∈ MnK and v ∈ Tx MnK , then ∥v∥K = ⟨v, v⟩K ; for K < 0, the Lorentzian form is positive definite only after restriction to this tangent space. These formulas show that constant-curvature neural layers can be implemented by choosing a model, applying the matching exponential, logarithmic, and transport maps, and reducing to Euclidean vector operations when K = 0.
53
2.9. Example Manifolds
Operator
Stereographic: PnK , DnK
Radius: LnK , SnK
Metric
⟨u, v⟩x= (λK )2 ⟨u, v⟩ p x −1 2 √ tanK |K| ∥−x ⊕K y∥ |K|
⟨u, v⟩x = ⟨u, v⟩K 1 √ cos−1 K (K ⟨x, y⟩K ) |K|
Geodesic distance Exponential map Logarithmic map Parallel transport
p λK v x ∥v∥ √ x ⊕K tanK |K| 2 |K|∥v∥ √ −1 |K|∥−x⊕K y∥ −x⊕ y 2 tanK K √ K ∥−x⊕ y∥ |K|λx
K
λK x gyr[y, −x]v λK y
√ p |K|∥v∥K sinK √ |K| ∥v∥K x + cosK v |K|∥v∥K
√
cos−1 K (β)
sign(K)(1−β 2 )
(y − βx), β = K ⟨x, y⟩K K⟨y,v⟩
Fréchet mean
K < 0: [142, Alg. 1] K > 0: Karcher flow
v − 1+K⟨x,y⟩K (x + y) K K < 0: [142, Alg. 3] K > 0: Karcher flow
References
[182]
[182]
Table 2.14: Riemannian operator templates for stereographic and radius constantcurvature coordinates.
54
Chapter 3 Riemannian Batch Normalization 3.1
Introduction
Motivated by the great success of normalization techniques [109, 12, 198, 219], researchers have sought to devise normalization layers tailored for manifold-valued data. Brooks et al. [31] introduced Riemannian Batch Normalization (RBN) designed specifically for the SPD manifold, with the ability to normalize the Riemannian mean. Kobler et al. [123] extended this approach to further control the Riemannian variance. However, the above methods are constrained to AIM on the SPD manifold, limiting their applicability. On the other hand, Chakraborty [34] proposed two distinct Riemannian normalization frameworks: one for Riemannian homogeneous spaces [34, Algs. 1–2] and another for matrix Lie groups [34, Algs. 3–4]. Nonetheless, the normalization designed for Riemannian homogeneous spaces can normalize neither the mean nor the variance, while the one for matrix Lie groups is confined to a specific type of distance [34, Sec. 3.2]. Meanwhile, Lou et al. [142, Alg. 2] proposed an RBN layer for general geometries. However, similar to Chakraborty [34, Algs. 1–2], it lacks theoretical guarantees for normalizing sample statistics. Therefore, a principled Riemannian normalization framework capable of controlling both Riemannian mean and variance remains unexplored. Given that Batch Normalization (BN) [109] serves as the foundational prototype for various types of normalization, this chapter focuses on RBN, with the potential to be extended to other normalization variants. We first present a general RBN framework for Lie groups, referred to as Lie Group Batch Normalization (LieBN), which can normalize both the Riemannian mean and variance under invariant metrics. To extend this normalization principle beyond Lie groups, we next introduce pseudo-reductive gyrogroups, 55
3.2. Lie Group Batch Normalization
Figure 3.1: Illustration of LieBN on the SPD, rotation, and correlation Lie groups. The 2×2 SPD, 3×3 rotation, and 3×3 correlation manifolds can be embedded into R3 as an open cone [220], a closed ball with antipodal points identified [95], and an open elliptope [195], respectively. LieBN is illustrated by (1) the left-invariant AIM and the proposed right-invariant CRIM geometry on the SPD manifold, (2) left or right translation under a bi-invariant metric on the rotation manifold, and (3) the bi-invariant ECM and LSM geometry on the correlation manifold. On the SPD and correlation manifolds, the batch mean and variance of the same input samples differ under different geometries. In all sub-figures, the black, blue, green, and red dots denote the boundary of the space, the input Lie group samples, the normalized samples, and the batch mean, respectively. As illustrated, our LieBN effectively normalizes the Lie group distribution. a relaxation of classical gyrogroups that also encompasses Lie groups as special cases. Building on this algebraic structure, we develop Gyrogroup Batch Normalization (GyroBN), which generalizes LieBN beyond group structures. We then instantiate LieBN and GyroBN on SPD manifolds, rotation matrices, correlation matrices, the Grassmannian, and different CCSs. Extensive experiments across multiple tasks demonstrate the effectiveness of this unified design.
3.2
Lie Group Batch Normalization
3.2.1
Introduction
Since several manifold-valued measurements form Lie groups, such as SPD manifolds [9, 137, 195], special orthogonal groups SO(n) [28], and full-rank correlation matrices 56
Chapter 3. Riemannian Batch Normalization [195, 191], we direct our attention to Lie groups. As each Lie group naturally admits leftand right-invariant metrics [69, Ch. 1.2], we propose a principled framework for RBN over Lie groups under invariant metrics, referred to as LieBN. Compared to previous work, our framework provides a theoretical guarantee for normalizing the Riemannian sample mean and variance. Empirically, we focus on the SPD, special orthogonal, and full-rank correlation manifolds. On SPD manifolds, we generalize three existing Lie group structures into parameterized ones by matrix power deformation. Additionally, we propose a novel rightinvariant metric, which, to the best of our knowledge, is the first non-trivial rightinvariant SPD metric1 , referred to as the Cholesky Right Invariant Metric (CRIM). We then instantiate our LieBN framework on SPD manifolds under these four Lie group structures. For rotation matrices, we adopt the popular bi-invariant metric [28], which will induce two types of LieBN: one w.r.t. left-invariance and another w.r.t. rightinvariance. On the correlation manifold, we manifest our LieBN under four recently developed correlation geometries [195, 191]. To facilitate usage, we provide a LieBN toolbox compatible with PyTorch, which can be used as a drop-in module. Fig. 3.1 illustrates our LieBN on different geometries, while Fig. 3.2 illustrates a minimal demo. Extensive experiments on SPD, rotation, and correlation manifolds involving radar recognition, human action recognition, and electroencephalography (EEG) classification demonstrate the effectiveness of our methods. We emphasize that our work is fundamentally distinct from Brooks et al. [31], Kobler et al. [123], Lou et al. [142] in theory and more general than Chakraborty [34]. Previous RBN methods are either designed for specific geometries [31, 123, 34] or fail to control both the mean and variance [142]. In contrast, our LieBN ensures the normalization of both the mean and variance across general Lie groups. In summary, our main contributions are: • A general LieBN framework with controllable first- and second-order moments; • A novel right-invariant metric on the SPD manifold, which is the first non-trivial right-invariant SPD metric; • Concrete instantiations of our LieBN framework on different geometries: four on SPD manifolds, one on rotation matrices, and four on correlation manifolds; 1
Although some metrics are bi-invariant, the associated group structures are commutative [9, 137]. Therefore, their bi-invariance is reduced to left-invariance.
57
3.2. Lie Group Batch Normalization from from from from
LieBN import LieBNSPD , LieBNRot , LieBNCor LieBN . Geometry . SPD import SPDMatrices LieBN . Geometry . Rotations import RotMatrices LieBN . Geometry . Correlation import Correlation
# ==== SPD matrices ==== P_spd = SPDMatrices ( n =5) . random (4 , 2 , 5 , 5) # Implemented metrics : LEM , ALEM , LCM , AIM , CRIM liebn_spd = LieBNSPD ([2 , 5 , 5] , metric = " LEM " , batchdim =[0]) output_spd = liebn_spd ( P_spd ) # ==== SO (3) matrices ==== P_so3 = RotMatrices () . random (4 , 2 , 3 , 3 , 3) # LieBN - Left if is_left else - Right liebn_so3 = LieBNRot ([3 , 3 , 3] , batchdim =[0 , 1] , is_left = False ) output_so3 = liebn_so3 ( P_so3 ) # ==== Correlation matrices ==== P_cor = Correlation ( n =5) . random (4 , 2 , 5 , 5) # Implemented metrics : ECM , LECM , OLM , LSM liebn_cor = LieBNCor ([2 , 5 , 5] , metric = " ECM " , batchdim =[0]) output_cor = liebn_cor ( P_cor )
Figure 3.2: Minimal examples of applying LieBN. • Validation of the effectiveness of our LieBN framework by extensive experiments on different geometries.2 Outline. Sec. 3.2.2 recalls the invariant metrics and Lie structures used by LieBN. Sec. 3.2.3 revisits Euclidean BN and RBN. Sec. 3.2.4 develops LieBN on Lie groups under left- and right-invariant metrics and establishes its statistical control. Sec. 3.2.5 instantiates LieBN on SPD, rotation, and full-rank correlation manifolds. Sec. 3.2.6 reports experiments that validate LieBN across these geometries. Proofs are deferred to Sec. B.2.
3.2.2
Preliminaries
An invariant metric can be understood as the Lie-group analogue of the Euclidean inner product being unaffected by translations. In Euclidean space, adding the same vector on the left or on the right preserves inner products, while on a Lie group the corresponding requirement is that left or right group translations preserve the Riemannian metric. 2
The code is available at https://github.com/GitZH-Chen/LieBN.git.
58
Chapter 3. Riemannian Batch Normalization Operator
(α, β)-AIM
(α, β)-LEM
LCM
Q⊕P ⊖P Identity WFM Invariance
KP K ⊤ Chol−1 (L−1 ) In Karcher Flow Left-invariance
exp (log(P ) + log(Q)) exp (− log(P )) P In exp ( i wi log(Pi )) Bi-invariance
Chol−1 (⌊L + K⌋ + KL) Chol−1 (−⌊L⌋ + L−1 ) P In −1 ψLC ( i wi ψLC (Pi )) Bi-invariance
Table 3.1: Review of SPD Lie groups and invariant metrics. Group
Q⊕P
SO(n)
QP
⊖P
Identity
WFM
Invariance
In
Karcher Flow
Bi-invariance
P −1 = P ⊤
Table 3.2: Review of the rotation Lie group and invariant metric.
Definition 60 (Invariance [69]). A Riemannian metric g L over a Lie group {M, ⊕} is left-invariant if, for any x, y ∈ M and V1 , V2 ∈ Ty M, it satisfies gyL (V1 , V2 ) = gLLx (y) (Lx∗,y (V1 ), Lx∗,y (V2 )), with Lx (y) = x⊕y as the left translation by x, and Lx∗,y as the differential map of Lx at y. Similarly, a right-invariant metric g R satisfies gyR (V1 , V2 ) = gRRx (y) (Rx∗,y (V1 ), Rx∗,y (V2 )), with Rx (y) = y⊕x as the right translation by x, and Rx∗,y as the differential map of Rx at y. Many popular matrix manifolds used in machine learning form Lie groups, including the SPD manifold, the full-rank correlation manifold, and the rotation group. Although the corresponding Riemannian geometries have been reviewed in Secs. 2.9.1, 2.9.2 and 2.9.4, Tabs. 3.1 to 3.3 provide a focused recap of their Lie structures. The notation follows Tabs. 2.5, 2.7 to 2.9 and 2.12.
3.2.3
Revisiting Normalization
3.2.3.1
Revisiting Euclidean Normalization
In Euclidean DNNs, normalization is a significant technique for accelerating network training by mitigating the issue of internal covariate shift [109]. While various normalization methods have been introduced [109, 12, 198, 219], they all share a common purpose: the normalization of the first and second moments. We focus on BN, the prototype of other normalization variants. Given a batch of activations {xi }N i=1 , the core operations in the standard Euclidean 59
3.2. Lie Group Batch Normalization Operator
ECM
LECM
C ⊕ C′ ⊖C Identity
OLM
LSM
ϕ−1 (ϕ(C) + ϕ(C ′ )) ϕ−1 (−ϕ(C)) −1 ϕP (0n×n ) N ϕ−1 w ϕ(C ) i i i=1 Bi-invariance
WFM Invariance
Table 3.3: Review of full-rank correlation Lie groups and invariant metrics. Here ϕ denotes Θ, log ◦Θ, Log◦ , and Log⋆ for ECM, LECM, OLM, and LSM, respectively. Methods
Involved Statistics
Controllable Mean
Controllable Variance
Geometries
SPDBN [31, Alg. 1] SPDBN [124, Alg. 1] SPDDSMBN [123] ManifoldNorm [34, Algs. 1–2] ManifoldNorm [34, Algs. 3–4] RBN [142, Alg. 2]
Mean Mean+Variance Mean+Variance Mean+Variance Mean+Variance Mean+Variance
✓ ✓ ✓ ✗ ✓ ✗
N/A ✓ ✓ ✗ ✓ ✗
SPD manifolds under AIM SPD manifolds under AIM SPD manifolds under AIM Riemannian homogeneous spaces A specific Lie group structure and distance Geodesically complete manifolds
LieBN (Ours)
Mean+Variance
✓
✓
Lie groups
Table 3.4: Summary of some representative RBN methods. BN can be expressed as:
xi − µ b +β ∀i ≤ N, xi ← γ p 2 vb + ϵ
(3.1)
where µb is the batch mean, vb2 is the batch variance, γ is the scaling parameter, β is the biasing parameter, and ϵ is a small scalar for stability. 3.2.3.2
Revisiting RBN
Although endeavors have been made to develop Riemannian normalization approaches tailored for manifolds, none of the existing methods effectively handle the first and second moments in a principled manner. Brooks et al. [31] introduced RBN over SPD manifolds under AIM. The core operations are defined as follows: 1
1
(3.2)
1 2
(3.3)
n Centering from mean M ∈ S++ : P̄i ← M − 2 Pi M − 2 , 1 2
n Biasing towards parameter B ∈ S++ : P̂i ← B P̄i B ,
where {Pi }N i=1 are SPD matrices, and M is their Fréchet mean under AIM. Let ΓP →Q (S) = ExpQ [PTP →Q (LogP (S))] , 60
(3.4)
Chapter 3. Riemannian Batch Normalization n where P, Q, S ∈ S++ . Under AIM, Eqs. (3.2) and (3.3) can be more generally expressed as ΓI→B [ΓM →I (Pi )]. (3.5)
However, Eqs. (3.2) and (3.3) only consider the Riemannian mean3 and do not consider the Riemannian variance. To remedy this limitation, Kobler et al. [123] further extended the RBN to involve the second-order statistics. The key operation is formulated as s
∀i ≤ N, P̄i ← ΓI→B [(ΓM →I (Pi )) v ],
(3.6)
where v 2 is the Fréchet variance, and s ∈ R is a scaling factor. However, this method is still limited to SPD manifolds under AIM. In parallel, Chakraborty [34, Algs. 1–2] proposed a general framework for Riemannian homogeneous spaces based on Eq. (3.5), which involves both first and second moments. However, Eq. (3.5) does not generally guarantee control over the Riemannian mean, resulting in agnostic Riemannian statistics [34, Sec. 3.1]. To mitigate this limitation, Chakraborty [34, Algs. 3–4] further proposed normalization over matrix Lie groups. However, the discussion is limited to a certain distance, limiting the applicability of their method. On the other hand, Lou et al. [142, Alg. 2] proposed an RBN based on a variant of Eq. (3.5). Similarly, their approach suffers from the same problem of agnostic Riemannian statistics on general manifolds. In summary, prevailing Riemannian normalization approaches lack a principled guarantee for controlling the first- and second-order statistics. In contrast, our method can normalize first- and second-order statistics over general Lie groups. We summarize the above RBN methods in Tab. 3.4.
3.2.4
LieBN
Since every Lie group naturally admits invariant metrics, we propose BN over Lie groups based on invariant metrics, referred to as LieBN. We first introduce the core operations under left-invariant metrics and then extend them to right-invariant metrics. Finally, we present the theoretical LieBN framework. In the following, we denote the neutral element in the Lie group M as E 4 . 3
Although not discussed in Brooks et al. [31], the congruent actions in Eqs. (3.2) and (3.3) can transfer the batch mean to a desired value under AIM. 4 The neutral element E is not necessarily the identity matrix.
61
3.2. Lie Group Batch Normalization 3.2.4.1
Ingredients under Left-invariant Metrics
In this subsection, we always assume that the Lie group M admits a left-invariant metric g L . Recalling the standard Euclidean BN [109] in Eq. (3.1), two key points are noteworthy: (a) the Euclidean BN implicitly assumes a Gaussian distribution and can effectively normalize the latent Gaussian distribution; (b) the centering and biasing operations control the mean, while the scaling controls the variance. Therefore, extending BN to Lie groups requires Lie-group counterparts of the Gaussian distribution, centering, biasing, and scaling. There are several notions of Gaussian distribution over manifolds [170, 218, 35, 14]. We adopt the intrinsic definition from Chakraborty and Vemuri [35], which characterizes a Gaussian distribution on the Lie group M with a mean parameter M ∈ M and variance σ 2 . This distribution is denoted as N (M, σ 2 ), and its probability density function (PDF) is d(X, M )2 2 , (3.7) p X | M, σ = k(σ) exp − 2σ 2
where k(σ) is the normalizing constant and d(·, ·) is the geodesic distance. When M is R with the standard Euclidean metric, Eq. (3.7) reduces to the Euclidean Gaussian.
On Lie groups, the natural counterparts of addition and subtraction in Eq. (3.1) are group operations. Therefore, centering and biasing on Lie groups can be defined by the left translation. Additionally, we define scaling via the tangent space. Specifically, for a batch of activations {Pi }N i=1 ⊂ M, we define the key operations of LieBN as follows: Centering from mean M ∈ M : P̄i ← L⊖M (Pi ), s Scaling: P̂i ← ExpE √ LogE (P̄i ) , v2 + ϵ Biasing towards parameter B ∈ M : P̃i ← LB P̂i ,
(3.8) (3.9) (3.10)
where M is the Fréchet mean, v 2 is the Fréchet variance, ⊖M ∈ M is the group inverse of M , L⊖M and LB are left translations (LB (Pi ) = B ⊕ Pi ), and s ∈ R \ {0} is a scaling parameter. The following two propositions demonstrate the above operations in normalizing mean and variance: one related to population statistics and the other related to sample statistics. Proposition 61 (Population). [↓] Given a random point X over {M, ⊕, g L }, and the Gaussian distribution N (M, v 2 ) defined in Eq. (3.7), we have the following for 62
Chapter 3. Riemannian Batch Normalization
the population statistics: 2 (1) (MLE of M ) Given {Pi }N i=1 ⊂ M i.i.d. sampled from N (M, v ), the maximum likelihood estimator (MLE) of M is the sample Fréchet mean.
(2) (Gaussian homogeneity) Given X ∼ N (M, v 2 ) and B ∈ M, we have LB (X) ∼ N (LB (M ), v 2 ).
(3.11)
Proposition 62 (Sample). [↓] Given N samples {Pi }N i=1 over the Lie group {M, ⊕, g L }, define ϕs (Pi ) = ExpE [s LogE (Pi )] . (3.12) We then have the following for the sample statistics. • Sample mean homogeneity: FM{LB (Pi )} = LB (FM{Pi }), ∀B ∈ M.
(3.13)
• Controllable dispersion from E: XN
i=1
wi d2 (ϕs (Pi ), E) = s2
XN
i=1
wi d2 (Pi , E),
(3.14)
where {wi }N i=1 are weights satisfying a convexity constraint, i.e., ∀i, wi > 0 and P i wi = 1.
Thm. 61 and Eq. (3.13) imply that our centering and biasing in Eqs. (3.8) and (3.10) can transfer the sample and population mean. As the post-centering mean is E, Eq. (3.14) implies that Eq. (3.9) can control the sample variance. More interestingly, the latent Gaussian distribution can be transferred under some geometries, such as SPD manifolds under LEM and LCM. Remark 63. The MLE of the mean of the Gaussian distribution has been examined in several previous works [176, 35, 34]. However, these studies primarily focus on particular manifolds or specific metrics. In contrast, our contribution lies in presenting a general result for Lie groups.
63
3.2. Lie Group Batch Normalization Remark 64. While Eq. (3.7) appeared in Kobler et al. [124], the authors only focus on SPD manifolds under AIM. The transformation of the population under their proposed RBN remains unexplored as well. Besides, while Chakraborty [34] analyzed the population properties for their RBN over matrix Lie groups, their results were confined within a specific distance. In contrast, our work provides a more extensive examination, encompassing both population and sample properties of our LieBN in a general manner. 3.2.4.2
Ingredients under Right-invariant Metrics
The key insight beneath Eqs. (3.8) and (3.10) and Thms. 61 and 62 is that left translation is an isometry under left-invariant metrics. Similarly, right translation is an isometry under right-invariant metrics. Therefore, it can be used for centering and biasing under right-invariant metrics. Following the previous notations, we define the centering and biasing under a right-invariant metric g R as centering to E: P̄i ← R⊖M (Pi ),
biasing towards B: P̃i ← RB (P̂i ).
(3.15) (3.16)
Similar to the case under left-invariant metrics, Thms. 61 and 62 can be easily extended to right-invariant metrics. Notably, the proofs for the MLE of M in Thm. 61 and controllable dispersion in Thm. 62 can be directly applied to the right-invariant metric. Therefore, we only show the homogeneity in the following proposition. Proposition 65. [↓] Given a random point X ∼ N (M, v 2 ) over {M, ⊕, g R }, B ∈ M, and N samples {Pi }N i=1 over M, we have: (1) Gaussian homogeneity: RB (X) ∼ N (RB (M ), v 2 ); (2) Sample homogeneity: FM{RB (Pi )} = RB (FM{Pi }). 3.2.4.3
LieBN under Invariant Metrics
With the above ingredients, Alg. 1 presents our theoretical LieBN framework. Similar to Ioffe and Szegedy [109], we use the moving average to update the running statistics. For a bi-invariant metric, LieBN can be implemented using either left or right translation. If the Lie group is commutative, LieBN under left and right translations are equivalent. Tab. 3.5 summarizes the LieBN types under different conditions. 64
Chapter 3. Riemannian Batch Normalization Commutativity
Non-commutative
Commutative
Invariance
Left
Right
Bi
Left = Right = Bi
LieBN Types
Left
Right
Left & Right
Left = Right
Table 3.5: Summary of LieBN types. The centering and biasing in Euclidean BN correspond to the group action of R. From a geometric perspective, the standard Euclidean metric is invariant under this group operation. Consequently, it is not surprising that our LieBN algorithm naturally generalizes the standard Euclidean BN. Proposition 66. [↓] The LieBN algorithm presented in Alg. 1 is equivalent to the standard Euclidean BN when M = Rn , both during the training and testing phases.
3.2.5
Manifestations
This section instantiates our LieBN in Alg. 1 on nine different Lie groups, including four on the SPD manifold, one on rotation matrices, and four on the correlation manifold. 3.2.5.1
LieBN on SPD Manifolds
We first extend the current Lie groups on SPD manifolds by the matrix power deformation, resulting in three families of parameterized Lie groups. Then, we propose a novel right-invariant metric on the SPD manifold, the first non-trivial right-invariant metric on this manifold. Finally, we construct LieBN layers based on these Lie structures. Deformed Lie structures on SPD manifolds. As shown in Tab. 3.1, there are three Lie groups on SPD manifolds, each with a left-invariant metric. These metrics include (α, β)-AIM, (α, β)-LEM, and LCM. For clarity, we denote the group operations w.r.t. (α, β)-AIM, (α, β)-LEM and LCM as ⊕LieAI , ⊕LE and ⊕LC , respectively.
Recently, Thanwerdas and Pennec [193] further extended (α, β)-AIM into a threeparameter family of metrics via the pullback of the matrix power function Pθ (·), scaled by θ12 and denoted by (θ, α, β)-AIM. The matrix power serves as a deformation, wherein (θ, α, β)-AIM encompasses (α, β)-AIM with θ = 1, and becomes (α, β)-LEM as θ approaches 0 [192]. Inspired by the deforming utility of the power function, we define the power-deformed metrics of (α, β)-LEM and LCM as the pullback metrics by Pθ and scaled by θ12 . We denote these two metrics as (θ, α, β)-LEM and θ-LCM, respectively. We have the following results with respect to the deformation. 65
3.2. Lie Group Batch Normalization Algorithm 1: Lie Group Batch Normalization (LieBN) Input : A batch of activations {Pi }N i=1 over Lie groups {M, ⊕, g}, a small positive constant ϵ, and momentum η ∈ [0, 1], running mean Mr = E, running variance vr2 = 1, biasing parameter B ∈ M, and scaling parameter s ∈ R \ {0}. Output : Normalized activations {P̃i }N i=1 . if training then Compute batch mean Mb and variance vb2 Update running statistics: Mr ← WFM({1 − η, η}, {Mr , Mb }) vr2 ← (1 − η)vr2 + ηvb2 Use the batch statistics, M ← Mb , v 2 ← vb2 else Use the running statistics, M ← Mr , v 2 ← vr2 for i ← 1 to N do Centering to the neutral element E: if g is left-invariant then P̄i ← L⊖M (Pi ) else P̄i ← R⊖M (Pi ) Scaling the variance: i h P̂i ← ExpE √vs2 +ϵ LogE (P̄i ) Biasing towards parameter B: if g is left-invariant then P̃i ← LB (P̂i ) else P̃i ← RB (P̂i )
Proposition 67 (Deformation). [↓] (θ, α, β)-LEM is equal to (α, β)-LEM. θ-LCM interpolates between g̃-LEM (as θ → 0) and LCM (θ = 1). Here, given any P ∈ n n S++ and tangent vectors V, W ∈ TP S++ , g̃-LEM is defined as ⟨V, W ⟩P = g̃(log∗,P (V ), log∗,P (W )),
(3.17)
where g̃(V1 , V2 ) = 21 ⟨V1 , V2 ⟩− 14 ⟨D(V1 ), D(V2 )⟩, D(Vi ) is a diagonal matrix consisting of the diagonal elements of Vi , and log∗,P is the differential map at P . As (θ, α, β)-LEM is equal to (α, β)-LEM, we focus on (α, β)-LEM, (θ, α, β)-AIM, and θ-LCM in the following. As a diffeomorphism, Pθ can also pull back the group operations ⊕LieAI and ⊕LC , denoted by ⊕θ-AI and ⊕θ-LC , respectively. We have the 66
Chapter 3. Riemannian Batch Normalization following proposition on the invariance. Proposition 68 (Invariance). [↓] (θ, α, β)-AIM is left-invariant w.r.t. ⊕θ-AI , while θ-LCM is bi-invariant w.r.t. ⊕θ-LC . SPD right-invariant metrics. AIM is left-invariant w.r.t. ⊕LieAI . We can also define a right-invariant metric w.r.t. ⊕LieAI by definition [69, Ch. 1.2]: gPCRI (V, W ) = ⟨(R⊖LieAI P )∗,P (V ), (R⊖LieAI P )∗,P (W )⟩I
(3.18)
where R(·) denotes Lie group right translation, ⊖LieAI P is the inverse of P under ⊕LieAI , n and ⟨·, ·⟩I denotes an arbitrary inner product on TI S++ . We set ⟨·, ·⟩I to be the same (α,β) as the AIM at I, i.e., ⟨·, ·⟩ . We call this metric CRIM, as the group operation is defined by the matrix product of Cholesky factors [195, Sec. 3.2]. n , Theorem 69. [↓] Given any SPD matrices P, Q and tangent vector V ∈ TP S++ n , g CRI } are the Riemannian operators on {S++
gPCRI (V, V ) =
L(L−1 V L−⊤ ) 1 L−1 2
(α,β) sym+
1 (α,β) e− 2 PeQ e− 12 d(P, Q) = log Q , ExpP (V ) = ⊖LieAI ExpAI − V̄ , e P ⊤ , LogP (Q) = − LL⊤ LVe L⊤ 1 2
!2
(3.19) (3.20) (3.21) (3.22)
sym+
where L is the Cholesky factor of P = LL⊤ , ⊖LieAI (·) is the group inverse, Pe e are the group inverses of P and Q, V̄ = L−1 V L−⊤ 1 L−1 L−⊤ and Q , and 2 sym+ ⊤ n×n e . Here, (X) Q denotes unnormalized Ve = LogAI sym+ = X + X , ∀X ∈ R Pe 1 symmetrization, and (X) 1 = ⌊X⌋ + 2 X. 2
Corollary 70. [↓] CRIM is geodesically complete, and the associated geodesic connecting SPD matrices P and Q is n o AI e e γ(P,Q) (t) = ⊖ γ (t; P , Q) 1 1 1 1 t LieAI − − e e e e =⊖ P 2 P 2 QP 2 Pe 2 , LieAI
67
(3.23)
3.2. Lie Group Batch Normalization Metric
(θ, α, β)-AIM
Invariance
Left-invariance
Bi-invariance LieBN-Left = LieBN-Right
θ-LCM
θ-CRIM Right-invariance
LieBN Type
LieBN-Left
Pullback Map
Pθ
log
Codomain
n , ⊕LieAI , θ12 g (α,β)-AI } {S++
{S n , ⟨·, ·⟩(α,β) } P +Q
P +Q
LQL⊤
L⊖Q (P ) or R⊖Q (P )
K −1 P K −⊤
P −Q
L−1 QL−⊤
ExpE [s LogE (P )]
Ps
P −Q sP
sP
FM
Karcher Flow
Arithmetic average
Arithmetic average
WFM({1 − η, η}, {P1 , P2 })
1 1 −1 −1 η P12 P1 2 P2 P1 2 P12
LQ (P ) or RQ (P ) Riemannian and Lie group operators in the codomain
(α, β)-LEM
KP K ⊤
Arithmetic weighted average
LieBN-Right
Pθ ◦ψLC
{LTn , θ12 ⟨·, ·⟩}
Arithmetic weighted average
Pθ n , ⊕LieAI , θ12 g CRI } {S++
⊖LieAI
⊖LieAI P
s
Karcher Flow
1 1 −1 −1 η ⊖LieAI Pe12 Pe1 2 Pe2 Pe1 2 Pe12
Table 3.6: Key operators in calculating LieBN on SPD manifolds. Invariance
LieBN Type
Bi-invariance
LieBN-Left & LieBN-Right
⊖R
R−1
LR (S)
RR (S)
ExpI [s LogI (R)]
FM
RS
SR
exp (s log (R))
[146, Alg. 1]
WFM({1 − η, η}, {R, S}) R exp(η log(R⊤ S))
Table 3.7: Key operators in calculating LieBN on the rotation matrices.
e = ⊖LieAI Q are group inverses, with γ AI as the geodesic where Pe = ⊖LieAI P and Q under AIM.
Similar to the discussion in Sec. 3.2.5.1, we define θ-CRIM as the deformed metric of CRIM by the pullback of matrix power function Pθ (·) and scaled by θ12 . As the pullback of CRIM, θ-CRIM is right-invariant w.r.t. ⊕θ-AI by definition. Proposition 71. θ-CRIM is right-invariant w.r.t. ⊕θ-AI . Manifestations on SPD manifolds. So far, there are four families of invariant metrics on the SPD Lie groups: (1) left-invariant (θ, α, β)-AIM w.r.t. ⊕LieAI ; (2) biinvariant (α, β)-LEM w.r.t. ⊕LE and θ-LCM w.r.t. ⊕LC ; (3) right-invariant θ-CRIM w.r.t. ⊕LieAI . Since all the above metrics are pullback metrics, the LieBN based on these metrics can be simplified and calculated in the codomain. We first show a general result on LieBN under the pullback metric. We denote Alg. 1 on the Lie group M as Pi ∈ {Pj }N j=1 ⊂ M.
LieBN(Pi ; B, s, ϵ, η),
(3.24)
Then we can obtain the following theorem. Theorem 72. [↓] Given a Lie group M1 , a Lie group M2 with an invariant metric g 2 , and a map f : M1 → M2 that is both a diffeomorphism and a Lie-group isomorphism, the map f induces an invariant metric g 1 on M1 , denoted as g 1 = 1 f ∗ g 2 . For a batch of activations {Pi }N i=1 in M1 , LieBN (Pi ; B, s, ϵ, η) in M1 can 68
Chapter 3. Riemannian Batch Normalization
be calculated in M2 by the following process: Mapping data into M2 : P̄i = f (Pi ), B̄ = f (B),
Performing LieBN in M2 : P̂i = LieBN2 (P̄i ; B̄, s, ϵ, η),
Mapping the resulting data back to M1 : P̃i = f −1 (P̂i ),
(3.25) (3.26) (3.27)
where LieBN2 is the LieBN on M2 . n Given a metric g on S++ , the power-deformed metric g̃ = θ12 P∗θ g is equal to P∗θ ( θ12 g). Thm. 72 indicates that the LieBN under g̃ can be calculated by the LieBN under 1 g. Besides, as the Christoffel symbols remain the same under constant scaling, the θ2 LieBNs under θ12 g and g only differ in the variance. We denote g (α,β)-AI and g (θ,α,β)-AI as the metric tensors of (α, β)-AIM and (θ, α, β)-AIM, respectively. Based on the above discussions, the computations of the LieBN under g (θ,α,β)-AI are reduced to the LieBN under θ12 g (α,β)-AI . Similarly, denoting g CRI and g θ-CRI as the metric tensors of CRIM and θ-CRIM, then the LieBN under θ-CRIM can be calculated by the one under 1 CRI g . Furthermore, as shown in Sec. 4.2.3, (α, β)-LEM is a pullback metric from the θ2 Euclidean space S n of symmetric matrices, while θ-LCM is a pullback metric from the Euclidean space LTn of lower triangular matrices. As shown in Thm. 66, the LieBN in the Euclidean space S n or LTn is simplified to the standard Euclidean BN. Therefore, the LieBNs under (α, β)-LEM and θ-LCM can be calculated by the Euclidean BN over S n and LTn , respectively. We denote the LieBN under left and right translations as LieBN-Left and LieBNRight, respectively. Then, the LieBNs under (θ, α, β)-AIM and θ-CRIM correspond to LieBN-Left and LieBN-Right, respectively. As ⊕LC and ⊕LE are commutative, the LieBN-Left and LieBN-Right under (α, β)-LEM and θ-LCM are equivalent. We denote n P, Q, P1 and P2 as points in the codomain, i.e., S++ with scaled CRIM for θ-CRIM, n n S++ with scaled (α, β)-AIM for (θ, α, β)-AIM, S for (α, β)-LEM, and LTn for θ-LCM, respectively. For CRIM, we denote Pei = ⊖LieAI Pi for i = 1, 2. We summarize all the necessary ingredients in Tab. 3.6 for calculating SPD LieBN. Note that for (θ, α, β)-AIM, our scaling operation defined in Eq. (3.9) encompasses the scaling operation in Kobler et al. [124, Eq. (9)] as a special case, when (θ, α, β) = (1, 1, 0).
3.2.5.2
LieBN on Rotation Matrices
As the Riemannian metric on the rotation matrices is bi-invariant, there are two instantiations of LieBN on this manifold, i.e., LieBN-Left based on the left translation 69
3.2. Lie Group Batch Normalization Metric
ECM
LECM
OLM
Invariance
Bi-invariance
LieBN Type
LieBN-Left = LieBN-Right
Pullback Map
Θ
Codomain
log ◦Θ
{LT1 (n), ⟨·, ·⟩}
{LT0 (n), ⟨·, ·⟩}
LSM
Log◦
Log⋆
{Hol(n), ⟨·, ·⟩(α,β,γ) }
{Row0 (n), ⟨·, ·⟩(α,δ,ζ) }
Table 3.8: Summary of LieBN on the correlation. permutation-invariant inner products [191].
⟨·, ·⟩(α,β,γ) and ⟨·, ·⟩(α,δ,ζ) are
and LieBN-Right based on the right translation. In particular, the scaling can be further simplified: ExpI (s LogI (R)) = exp (s log (R)). For the specific SO(3), the matrix exp and log can be efficiently calculated without matrix decomposition [95, Sec. 3.2]. Tab. 3.7 presents the expressions of the required operators in Alg. 1. 3.2.5.3
LieBN on Full-Rank Correlation Matrices
As summarized in Tab. 3.3, all four correlation metrics are bi-invariant, and their associated Lie groups are commutative. Consequently, LieBN-Left is identical to LieBNRight. Moreover, all four correlation metrics are pullback metrics from simpler Euclidean spaces. Therefore, LieBN over the correlation can be implemented according to Thm. 72: (1) map the correlation into the prototype Euclidean space, (2) apply Euclidean BN, and (3) map back to the correlation. Optimization. Finally, we discuss the optimization of the correlation-valued biasing parameter B ∈ Cor+ (n). As reviewed in Sec. 2.9.2, the correlation matrix can be identified by the product of hyperbolic spaces via the Cholesky decomposition. Given C ∈ Cor+ (n), the k-th row of the Cholesky factor L = Chol(C) is (Lk1 , . . . , Lk,k−1 , Lkk , 0, . . . , 0) with Lkk > 0, which belongs to the hyperbolic space of an open hemisphere HSk−1 = x ∈ Rk | ∥x∥ = 1, xk > 0 .
(3.28)
Besides, the open hemisphere HSn is isometric to the Poincaré ball Pn = {x ∈ Rn | ∥x∥ < 1} by πHSn →Pn ((x⊤ , xn+1 )⊤ ) =
x . 1 + xn+1
(3.29)
(3.30)
Therefore, each correlation can be parameterized with n − 1 Poincaré vectors. Each 70
Chapter 3. Riemannian Batch Normalization Poincaré vector can be optimized directly on its manifold using the Riemannian optimization strategy reviewed in Sec. 2.7. The above process can be expressed as
1 0 L21 L22 C7→ .. .. . . Ln1 Ln2
3.2.6
0 x1 ∈ P1 0 .. . → 7 . .. . xn−1 ∈ Pn−1 · · · Lnn
··· ··· .. .
(3.31)
Experiments
This section validates our LieBN on nine invariant metrics across the SPD, rotation, and correlation matrices. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.1. 3.2.6.1
Experiments of LieBN on the SPD Manifold
Note that our LieBN layers are architecture-agnostic and can be applied to any existing SPD neural network. Following the previous work [106, 31, 123], we focus on two network architectures: (1) SPDNet [106] for drone recognition on the Radar data set [31], and human action recognition on the HDM05 [153] and FPHA [80] data sets; (2) TSMNet [123] for EEG classification on the Hinss2021 data set [101]. In the EEG application, TSMNet is endowed with SPD Domain-Specific Momentum Batch Normalization (SPDDSMBN) (denoted TSMNet+SPDDSMBN) [123], which is a domain adaptation version of Kobler et al. [124]. For a fair comparison, we also implement a domain-specific momentum LieBN, referred to as DSMLieBN. The backbone network architectures are represented as {d0 , d1 , . . . , dL }, where the dimension of the parameter in the i-th BiMap layer is di × di−1 . As (α, β) only affects variance calculation throughout LieBN, we simply set (α, β) = (1, 0) and only tune the deformation factor θ. For each family of LieBN or DSMLieBN, we report two representatives: the standard one induced from the standard metric (θ = 1), and the one induced from the deformed metric with proper θ. If the standard one is already saturated, we only report the result of that standard variant. Application to SPDNet. As SPDNet is a classical SPD network, we apply our LieBN to SPDNet on the Radar, HDM05, and FPHA data sets. Additionally, we compare our method with SPDNetBN, which applies the SPDBN in Eqs. (3.2) and (3.3) to SPDNet. Following Brooks et al. [31], we use the architectures of {20, 16, 8}, {93, 30}, and {63, 33} for the Radar, HDM05, and FPHA data sets, respectively. The 10-fold av71
3.2. Lie Group Batch Normalization (a) Radar data set. SPDNetLieBN Acc Fit time (s) Mean±STD Max
SPDNet 0.98 93.25±1.10 94.4
SPDNetBN 1.56 94.85±0.99 96.13
Best θ
θ=1 AIM-(1)
LEM-(1)
LCM-(1)
CRIM-(1)
LCM-(-0.5)
1.62 95.47±0.90 96.27
1.28 94.89±1.04 96.8
1.11 93.52±1.07 95.2
1.82 94.35±0.68 95.6
1.43 94.80±0.71 95.73
(b) HDM05 data set. SPDNetLieBN Acc Fit time (s) Mean±STD Max
SPDNet 0.57 59.13±0.67 60.34
0.97 66.72±0.52 67.66
Best θ
θ=1
SPDNetBN AIM-(1)
LEM-(1)
LCM-(1)
CRIM-(1)
AIM-(1.5)
LCM-(0.5)
CRIM-(0.5)
1.14 67.79±0.65 68.75
0.87 65.05±0.63 66.05
0.66 66.68±0.71 68.52
1.37 63.25±0.88 64.94
1.46 68.16±0.68 69.25
1.01 70.84±0.92 72.27
1.74 65.76±0.54 66.96
(c) FPHA data set. SPDNetLieBN Acc Fit time (s) Mean±STD Max
SPDNet 0.32 85.59±0.72 86
0.62 89.33±0.49 90.17
Best θ
θ=1
SPDNetBN AIM-(1)
LEM-(1)
LCM-(1)
CRIM-(1)
AIM-(1.5)
LCM-(0.5)
CRIM-(-0.5)
0.8 89.70±0.51 90.5
0.55 86.56±0.79 87.83
0.39 77.64±1.00 79
0.92 84.65±1.20 86.67
1.03 90.39±0.66 92.17
0.65 86.33±0.43 87
1.21 86.40±0.57 87.17
Table 3.9: 10-fold average results of SPDNet with and without SPDBN or LieBN on the Radar, HDM05, and FPHA data sets. If the LieBN under the standard metric (θ = 1) is not saturated, the rightmost columns report the deformed LieBN. The best results are bold. erage results, including the average training time (s/epoch), are summarized in Tab. 3.9. We have three key observations regarding the choice of metrics, deformation, and training efficiency. • The choice of metrics. The metric that yields the most effective LieBN layer differs for each data set. Specifically, the optimal LieBN layers on these three data sets are the ones induced by AIM-(1), LCM-(0.5), and AIM-(1.5), respectively, which improve the performance of SPDNet by 2.22, 11.71, and 4.80 percentage points. Additionally, although the LCM-based LieBN performs worse than other LieBN variants on the Radar and FPHA data sets, it exhibits the best performance on the HDM05 data set. These observations demonstrate the value of a unified normalization framework that can be instantiated under multiple admissible SPD geometries. • The effect of deformation. Deformation patterns also vary across data sets. Firstly, the standard AIM and CRIM are already saturated on the Radar data set. Secondly, the appropriate deformation θ can further enhance the performance of LieBN. Notably, even though the LieBNs induced by LCM-(1) and CRIM-(1) 72
Chapter 3. Riemannian Batch Normalization (a) Inter-session classification
(b) Inter-subject classification
Method
Fit Time
Mean±STD
Method
Fit Time
Mean±STD
SPDDSMBN
0.16
54.12±9.87
SPDDSMBN
7.74
50.10±8.08
AIM-(1) LEM-(1) LCM-(1) CRIM-(1)
0.16 0.13 0.10 0.29
55.10±7.61 54.95±10.09 51.54±6.88 51.86±9.21
AIM-(1) LEM-(1) LCM-(1) CRIM-(1)
6.94 4.71 3.59 16.35
50.04±8.01 50.95±6.40 51.86±4.53 50.71±8.1
LCM-(0.5)
0.15
53.11±5.65
CRIM-(1.5) AIM-(-0.5)
19.51 8.71
51.34±5.82 53.97±8.78
DSMLieBN
DSMLieBN
Table 3.10: Cross-validation results of TSMNet with SPDDSMBN and DSMLieBN on the Hinss data set. If the DSMLieBN under the standard metric (θ = 1) is not saturated, the bottom rows report deformed DSMLieBN. impede the learning of SPDNet on the FPHA data set, they can improve the performance under an appropriate deformation θ. These findings highlight the efficacy of the deforming geometry on the SPD manifold. • Efficiency. Although our LieBN involves additional computations on variance compared with SPDNetBN, our LieBN achieves comparable or even better efficiency than SPDNetBN. In particular, the LieBN induced by standard LEM or LCM exhibits better efficiency than SPDNetBN. Even with deformation, the LCM-based LieBN is still comparable with SPDNetBN in terms of efficiency. This phenomenon could be attributed to the fast and simple computation of LCM and LEM. Application to EEG classification. We apply our method to TSMNet under two scenarios, inter-session and inter-subject. Following Kobler et al. [123], we adopt the architecture of {40, 20}. Compared to SPDDSMBN, DSMLieBN-AIM obtains the highest average scores of 55.10% and 53.97% in these two scenarios, outperforming SPDDSMBN by 0.98 and 3.87 percentage points, respectively. In the intersubject scenario, the efficiency advantage of our LieBN over SPDDSMBN is more evident. Specifically, both the LEM- and LCM-based DSMLieBN achieve similar or better performance compared to SPDDSMBN, while requiring considerably less training time. For example, DSMLieBN-LCM-(1) achieves better results with only half the training time of SPDDSMBN on inter-subject tasks. Interestingly, under the standard AIM, the sole difference between SPDDSMBN and our DSMLieBN is the way of centering and biasing. SPDDSMBN applies the matrix inverse square root and matrix square root to fulfill centering and biasing, while AIM-induced LieBN uses more efficient Cholesky decomposition. As such, the DSMLieBN induced by the standard AIM is more efficient than SPDDSMBN, particularly on the inter-subject task. On the other hand, 73
3.2. Lie Group Batch Normalization
Figure 3.3: Visualization of input and output 30 × 30 SPD matrices in LieBN using 2 × 2 Riemannian t-SNE embeddings. The first row shows the input and output under different metrics. Due to the significant difference in magnitude between the t-SNE embeddings of LieBN’s input and output, the second row separately visualizes the LieBN output (at a smaller scale). the CRIM-based LieBN shows less efficiency, due to the relatively complex Riemannian computation of this metric. Visualization. We randomly select 50 samples and visualize the input and output of LieBN on the HDM05 data set. Using Riemannian t-SNE [65], we map the 30 × 30 SPD matrices to 2 × 2 low-dimensional representations. As shown in Fig. 3.3, LieBN effectively normalizes the data distribution. Specifically, the input t-SNE embeddings are largely scattered and their elements can reach up to 400K, whereas those of the output embeddings are mostly constrained within 20. 3.2.6.2
Experiments of LieBN on Rotation Matrices
This subsection implements our LieBN on the special orthogonal groups, i.e., SO(n), also known as rotation matrices. As the Riemannian metric on SO(n) is bi-invariant, there are two instantiations of our LieBN on this group: LieBN-Left based on the left translation and LieBN-Right based on the right translation. We apply our LieBN to the classic LieNet backbone [107], where the latent space is the special orthogonal group. Following Huang et al. [107], we use three action recognition data sets, the G3D [23], HDM05 [153], and NTU60 [178] data sets. We denote the LieNet models 74
Chapter 3. Riemannian Batch Normalization HDM05
80 70 60 0
LieNet LieNetLieBN-Left LieNetLieBN-Right 50 100 150 Epochs
70
60 0
Acc.
80 Acc.
Acc.
90
65
LieNet LieNetLieBN-Left LieNetLieBN-Right 50 100 150 Epochs
NTU60-2Blocks
60
60
55
55
50 45 0
LieNet LieNetLieBN-Left LieNetLieBN-Right 25 50 Epochs
NTU60-3Blocks
65
Acc.
G3D
LieNet LieNetLieBN-Left LieNetLieBN-Right 25 50 Epochs
50 45 0
Figure 3.4: Test accuracy curves corresponding to Tab. 3.11. Method
G3D
HDM05
NTU60
Mean±STD
Max
Mean±STD
Max
2Blocks
3Blocks
LieNet
87.91±0.90
89.73
76.92±1.27
79.11
62.4
60.91
LieNetLieBN-Left LieNetLieBN-Right
88.88±1.62 88.12±1.12
90.67 90.3
78.89±1.07 79.39±1.13
80.88 80.67
63.51 63.6
62.62 62.72
Table 3.11: Results of LieNet with or without rotation LieBN. with our LieBN-Left and LieBN-Right as LieNetLieBN-Left and LieNetLieBN-Right, respectively. Results. We conduct 10-fold experiments on the G3D and HDM05 data sets under the suggested 3Blocks5 and 2Blocks architectures, respectively. On the NTU60 data set, we validate LieBN under the 2Blocks and 3Blocks settings. The results are presented in Tab. 3.11. Due to differences in software, our reimplemented LieNet (in PyTorch) performs slightly differently from the results reported by Huang et al. [107] (in MATLAB). However, we still observe a clear improvement when applying our LieBN to the vanilla LieNet backbone. Additionally, LieBN-Right performs slightly better than LieBN-Left. Although the effects of left and right translations on the sample statistics under the biinvariant metric are identical, their transformations on each sample differ, as illustrated in Fig. 3.1. This difference could slightly affect the network performance. The specific optimal choice of left or right translations depends on the data set’s characteristics. Training dynamics. Fig. 3.4 presents the test accuracy curves. We have the following additional observations, which can be attributed to the mitigated covariate shift by our LieBN, as our LieBN can effectively normalize the sample statistics. • Accelerated convergence. LieBN significantly accelerates the convergence of LieNet. Specifically, on the NTU60 data set, the largest data set involved, LieNet with LieBN converges by the 5th epoch, whereas the vanilla LieNet does not 5
Each block consists of a RotMap layer followed by a RotPooling layer. For more details, please refer to Huang et al. [107].
75
3.3. Gyrogroup Batch Normalization
Data Set
SPDNet
HDM05 FPHA
59.13±0.67 85.59±0.72
SPDNetLieBN-Cor ECM
LECM
OLM
LSM
65.37 ± 1.07 87.20 ± 0.12
61.35 ± 0.34 87.03 ± 0.32
60.33 ± 0.12 86.80 ± 0.12
60.00 ± 0.27 86.77 ± 0.29
Table 3.12: Results of SPDNet with or without correlation LieBN under different invariant metrics. converge until the 25th epoch. A similar phenomenon can also be observed on the HDM05 data set. • More stable performance. LieBN enhances the stability of network training. In particular, on the HDM05 and G3D data sets, the initial training fluctuations are greatly mitigated by our LieBN. 3.2.6.3
Experiments of LieBN on Correlation Matrices
We apply our correlation LieBN (LieBN-Cor) to SPD networks. Our experiments focus on the SPDNet backbone using the FPHA and HDM05 data sets. LieBN-Cor is applied before the final classification layer. Specifically, SPD features are first activated by the power function, then mapped into correlation matrices via Cor(·), and finally processed by LieBN-Cor. Results. The 5-fold average results are presented in Tab. 3.12. Although LieBNCor is not specifically designed for SPD networks, it still improves SPDNet’s performance, demonstrating its effectiveness. Among the four invariant metrics, ECM achieves the best performance. LieBN-SPD outperforms LieBN-Cor when applied to SPDNet, as expected because SPDNet is tailored to SPD matrices. However, this does not undermine the validity of LieBN-Cor. The consistent improvement over vanilla SPDNet highlights the potential of applying LieBN-Cor to correlation manifolds.
3.3
Gyrogroup Batch Normalization
3.3.1
Introduction
Although LieBN can normalize sample statistics, many important geometries in machine learning do not admit a Lie group structure. As a result, existing methods still lack a principled solution for Riemannian normalization. Recently, gyro-structures have emerged as effective tools for building Riemannian networks across various geometries, 76
Chapter 3. Riemannian Batch Normalization
M
GyroBN
M
Figure 3.5: Illustration of GyroBN on manifold-valued data. Blue points, green points, and the red dashed curves indicate the input samples, normalized outputs, and data distributions, respectively. Method
Controllable Statistics
SPDBN [31] SPDBN [124] SPDDSMBN [123] ManifoldNorm [34, Algs. 1–2]
M M+V M+V N/A
ManifoldNorm [34, Algs. 3–4]
M+V
RBN [142, Alg. 2]
N/A
LieBN (Sec. 3.2)
M+V
GyroBN
M+V
Applied Geometries
Incorporated by GyroBN
SPD manifolds under AIM SPD manifolds under AIM SPD manifolds under AIM Riemannian homogeneous spaces Matrix Lie groups under the distance d(X, Y ) = ∥log (X −1 Y )∥ Geodesically complete manifolds Lie groups under invariant metrics
✓ ✓ ✓ ✗
Pseudo-reductive gyrogroups with gyroisometric gyrations
✓ ✗ ✓ N/A
Table 3.13: Comparison of previous RBN methods with GyroBN, where M and V denote the sample mean and variance. including SPD [157], Grassmannian [157], hyperbolic [76], and spherical manifolds [182]. They naturally extend Euclidean vector structures while encompassing Lie groups and non-group geometries. For instance, the Grassmannian, hyperbolic, and spherical manifolds do not form Lie groups but instead form gyrogroups. Based on the analysis above, this part of the thesis first introduces the pseudoreductive gyrogroup, a relaxation of the classical gyrogroup that provides a broader algebraic foundation for Riemannian normalization. Building on this structure, we develop GyroBN, a general RBN framework on pseudo-reductive gyrogroups, as illustrated in Fig. 3.5. We employ gyrosubtraction, gyroaddition, and scalar gyromultiplication to generalize the centering (vector subtraction), biasing (vector addition), and scaling (scalar multiplication) in Euclidean BN to curved manifolds in a principled manner. We clarify why centering and biasing in GyroBN rely on left gyroaddition, rather than other candidates, such as right gyroaddition or gyrocoaddition [200, Def. 2.9]. We show that when gyrations are gyroisometries, GyroBN enjoys theoretical control over sample statistics. These conditions are satisfied by all known gyrogroups in machine learn77
3.3. Gyrogroup Batch Normalization ing, providing a principled and unified normalization mechanism. Moreover, several existing RBN methods arise as special cases of GyroBN, including LieBN and various SPD-based variants, as summarized in Tab. 3.13. Beyond the LieBN instantiations in Sec. 3.2.5, we instantiate GyroBN on seven representative geometries: the Grassmannian [19], five CCSs [76, 131, 13], and the full-rank correlation manifold [195]. For the Grassmannian, we propose an efficient implementation. For CCSs, we cover five models: Poincaré ball, hyperboloid, Beltrami–Klein, sphere, and projected hypersphere. To enable these instantiations, we refine the projected hypersphere structure [13], derive closed-form gyro-structures for the hyperboloid and sphere, and develop the Riemannian structure of the Beltrami–Klein model. For the correlation manifold, we demonstrate that its gyro-structure can be defined rowwise on its Cholesky factor. We also provide a PyTorch-compatible toolbox [169] with drop-in GyroBN layers, illustrated in Fig. 3.6. Experiments on networks over these seven geometries validate the effectiveness of our framework. In summary, the main contributions of this part are • Theoretical foundation: pseudo-reductive gyrogroups as a relaxation of classical gyrogroups; • General framework: GyroBN as a plug-and-play normalization mechanism, with pseudo-reduction and gyroisometric gyrations ensuring theoretical control of batch statistics; • Geometric insights: refined projected hypersphere gyro-structure, closed-form hyperboloid and sphere gyro-structures, Riemannian structure of the Beltrami–Klein model, and row-wise correlation manifold gyro-structure; • Practical instantiations: implementations on the Grassmannian, five CCSs, and the correlation manifold with extensive experiments.6 Outline. Sec. 3.3.2 introduces pseudo-reductive gyrogroups and analyzes their theoretical properties, while Sec. 3.3.3 develops the GyroBN framework. Sec. 3.3.4 shows that prior RBN methods are special cases and instantiates GyroBN on seven representative geometries. Sec. 3.3.5 reports experiments that validate GyroBN across these geometries. Proofs are deferred to Sec. B.3. 78
Chapter 3. Riemannian Batch Normalization from GyroBN import * from GyroBN . Geometry import * # ==== Grassmannian ==== manifold = GrassmannianGyro ( n =50 , p =10) X_gr = manifold . random_normal (30 , 50 , 10) gybn_gr = GyroBNGr ( shape =[50 , 10]) out_gr = gybn_gr ( X_gr ) # ==== Five CCSs ==== models = [ ( " Poincare " , Stereographic ( K = -1.0) ) , ( " Hyperboloid " , Hyperboloid ( K = -1.0) ) , ( " Klein " , Klein ( K = -1.0) ) , ( " Sphere " , Sphere ( K = 1.0) ) , ( " ProjSphere " , Stereographic ( K = 1.0) ) , ] for name , manifold in models : X_ccs = manifold . random_normal (30 , 16) gybn_ccs = GyroBNCCS ( shape =[16] , model = name , K = manifold . K ) out_ccs = gybn_ccs ( X_ccs ) # ==== Full - Rank Correlation ==== manifold = C o r P o l y H y p e r b o l i c C h o l e s k y M e t r i c ( n =10) X_cor = manifold . random (30 , 10 , 10) gybn_cor = GyroBNCor ( shape =[10 , 10]) out_cor = gybn_cor ( X_cor )
Figure 3.6: Minimal examples of applying GyroBN.
3.3.2
Pseudo-Reductive Gyrogroups
Given a gyrogroup (G, ⊕), the left gyrotranslation by x ∈ G is defined as Lx : G → G,
Lx (y) = x ⊕ y,
∀y ∈ G.
(3.32)
If any gyrotranslation is a gyroisometry, we can use gyrotranslation to center manifoldvalued samples for the normalization layer. Nguyen and Yang [159] shows that any left gyrotranslation on the SPD and Grassmannian manifolds is a gyroisometry. However, the proof relies on the left cancellation law of gyrogroups, which does not hold for non-reductive gyrogroups, such as the Grassmannian. Here, non-reductive gyrogroups refer to gyro-like groupoids that satisfy the first three gyrogroup axioms but fail the 6
The code is available at https://github.com/GitZH-Chen/GyroBN.git.
79
3.3. Gyrogroup Batch Normalization
Ours:
Axiom (G1-3) Pseudo-reduction
Left cancellation law
Or
Previous:
Axiom (G1-3) Left reduction (G4)
Left gyrotranslation law
Invariance of gyronorm under any gyration
Gyroisometries of any gyration and left gyrotranslation
Gyroisometry of the gyroinverse
Gyrocommutativity
Figure 3.7: A conceptual comparison of the derivation logic in our work with that in previous work [159], where the left gyrotranslation law is presented in Thm. 190. The previous work proves the results on the SPD and Grassmannian manifolds in a caseby-case manner. In contrast, we relax the left reduction into pseudo-reduction and give a general analysis. Our framework also corrects the proof for the Grassmannian cases. last axiom in Thm. 51, namely the left reduction law (G4), gyr[x, y] = gyr[x ⊕ y, y]. Therefore, the proof is not generally valid for the Grassmannian. We propose an intermediate structure, referred to as pseudo-reductive gyrogroups, which supports the left cancellation law and, therefore, the gyroisometry of gyrotranslation. This structure forms the algebraic foundation for building the normalization layer. As illustrated in Fig. 3.7, our derivation extends the case-by-case approach by Nguyen and Yang [159] to general pseudo-reductive gyrogroups. 3.3.2.1
From Gyrogroups to Pseudo-Reductive Gyrogroups
Definition 73 (Pseudo-reductive gyrogroups). A groupoid (G, ⊕) is a pseudoreductive gyrogroup if it satisfies the axioms (G1), (G2), (G3) and the following pseudo-reductive law: gyr[a, x] = id, for any left inverse a of x in G,
(3.33)
where id is the identity map. Eq. (3.33) can be intuitively viewed as an intermediate between reduction and nonreduction. For gyrogroups, Eq. (3.33) can be directly obtained from left gyroassociativity (G3) and reduction (G4) [200, Thm. 2.10, item 3]. However, there is no theoretical guarantee that Eq. (3.33) holds for non-reductive gyrogroups. Therefore, we name Eq. (3.33) pseudo-reduction. Nevertheless, for the specific non-reductive Grassmannian, it is indeed pseudo-reductive. 80
Chapter 3. Riemannian Batch Normalization
f n) are pseudo-reductive gyrocommutative Proposition 74. [↓] Gr(p, n) and Gr(p, gyrogroups.
Our pseudo-reductive gyrogroup naturally generalizes the vanilla gyrogroup, as it shares most of the basic properties of gyrogroups [200, Thms. 2.10–2.11]. Theorem 75 (First pseudo-reductive gyrogroup properties). [↓] Let (G, ⊕) be a pseudo-reductive gyrogroup. For any elements x, y, z, a ∈ G, we have: (1) If x ⊕ y = x ⊕ z, then y = z (General Left Cancellation law; see Case (8) below). (2) gyr[e, x] = id for any left identity e in G. (3) gyr[a, x] = id for any left inverse a of x in G. (4) There is a left identity that is a right identity. (5) There is only one left identity. (6) Every left inverse is a right inverse. (7) There is only one left inverse, ⊖x, of x, and ⊖(⊖x) = x. (8) The left cancellation law: ⊖x ⊕ (x ⊕ y) = y. (9) The gyrator identity: gyr[x, y]a = ⊖(x ⊕ y) ⊕ {x ⊕ (y ⊕ a)}. (10) gyr[x, y]e = e. (11) gyr[x, y](⊖a) = ⊖ gyr[x, y]a. (12) gyr[x, e] = id. (13) The gyrosum inversion law: ⊖(x ⊕ y) = gyr[x, y](⊖y ⊕ ⊖x). Remark 76. In non-reductive gyrogroups, the identities in Cases (2) and (3) are not guaranteed to hold. Consequently, any property relying on them, such as Case (4) and those from Case (6) to Case (10), is also not guaranteed to hold. The absence of these basic properties undermines the rationality of non-reductive gyrogroups. In contrast, our pseudo-reductive gyrogroups preserve most of the fundamental properties of gyrogroups.
81
3.3. Gyrogroup Batch Normalization 3.3.2.2
Isometries over Pseudo-Reductive Gyrogroups
The gyro-structure in the following is assumed to be defined as Eq. (2.77)–Eq. (2.83). We first clarify that the Riemannian distance agrees with the gyrodistance, and the Riemannian isometry agrees with the gyroisometry. These justify gyrodistance and gyroisometry for gyrospaces over manifolds. Lemma 77 (Distances). [↓] Given a pseudo-reductive gyrogroup (M, ⊕), we have d(x, y) = ∥Logx (y)∥x = ∥⊖x ⊕ y∥gyr = dgyr (x, y),
∀x, y ∈ M,
(3.34)
where d denotes the geodesic distance.7 f ⊕) e be two pseudo-reductive gyLemma 78 (Isometries). [↓] Let (M, ⊕) and (M, f respectively. If rogroups. Their gyro identity elements are e ∈ M and ee ∈ M, f is a Riemannian isometry with ee = ϕ(e), then the following hold. ϕ:M→M (1) The Riemannian isometry is a gyroisometry:
dgyr (x, y) = dg gyr (ϕ(x), ϕ(y)),
(3.35)
f where dgyr and dg gyr are the gyrodistances over M and M, respectively.
(2) If the gyroinverse, gyration, or left gyrotranslation over M is a gyroisometry, f is also a gyroisometry. its counterpart over M
Thm. 77 implies that, as long as the gyro-structure is defined by Eq. (2.77)–Eq. (2.83), the gyrodistance coincides with the geodesic distance. Unless otherwise specified, we shall not distinguish between the two and uniformly denote them by d(·, ·). Besides, the second result in Thm. 78 is particularly useful, as several geometries are isometric, such as the ONB and PP Grassmannian, as well as different models in hyperbolic geometry. Now, we analyze gyroisometries over pseudo-reductive gyrogroups. The most related property in Thm. 75 is the left cancellation law, one of the key prerequisites for a gyrotranslation to be a gyroisometry. Note that the left cancellation comes from left gyroassociativity and Eq. (3.33) [200, Thm. 2.10, item 9]. Therefore, left cancellation does not generally hold for non-reductive gyrogroups but exists in pseudo-reductive 7
On Cartan–Hadamard manifolds, the statement holds for all x, y ∈ M. More generally, the equality requires x, y to lie within a geodesic ball of convexity radius to ensure the well-definedness of the minimizing geodesic and logarithm. In this section, we implicitly assume these conditions are satisfied.
82
Chapter 3. Riemannian Batch Normalization gyrogroups. We first present an if-and-only-if statement about gyroisometry, which will be useful in the following. Theorem 79. [↓] Given a pseudo-reductive gyrogroup (G, ⊕), gyr[x, y] preserves the gyronorm for any x, y ∈ G if and only if gyr[x, y] is a gyroisometry for any x, y ∈ G. The gyroisometry of any gyration is a prerequisite for other operators to be gyroisometries. Theorem 80 (Gyroisometries). [↓] Given a pseudo-reductive gyrogroup (G, ⊕) with any gyr[·, ·] as a gyroisometry, we have the following. (1) The left gyrotranslation is a gyroisometry. (2) If (G, ⊕) is gyrocommutative, then any gyroinverse is a gyroisometry. Now, we discuss the gyroisometries for the gyro-structures reviewed in Tabs. 2.5, 2.11 and 2.13. Theorem 81. [↓] For the pseudo-reductive gyrogroups corresponding to the SPD manifold (under AIM, LEM, and LCM), the ONB and PP Grassmannian, and the stereographic model with K ≤ 0 (the Poincaré ball for K < 0 and Euclidean space for K = 0), the gyrodistance coincides with the geodesic distance. Moreover, the gyroinverse, any gyration, and any left gyrotranslation are gyroisometries.
Credit and sketch of the proof. As Thm. 77 already shows that the gyrodistance coincides with the geodesic distance, it remains to establish the isometries. For the Grassmannian and SPD manifolds, these were proved by Nguyen and Yang [159, Thms. 2.12– 2.14 and 2.16–2.18], although their arguments implicitly treated the non-reductive Grassmannian as a gyrogroup by using left cancellation. Our Thms. 74 and 75 confirms that the Grassmannian is pseudo-reductive and does satisfy left cancellation, thereby validating their results. However, we can directly establish these properties from Thms. 79 and 80. The complete proof is given in Sec. B.3.7. Remark 82. The remaining constant-curvature manifolds reviewed in Sec. 2.9.5, including the stereographic model with K > 0, radius model, and Beltrami– Klein model, also satisfy these properties. Their verifications will be presented in Sec. 3.3.4. 83
3.3. Gyrogroup Batch Normalization
3.3.3
GyroBN on Pseudo-Reductive Gyrogroups
Building on Thm. 81, which establishes that several geometries admit isometric gyrotranslations, we develop RBN in a principled way for general pseudo-reductive gyrogroups, referred to as GyroBN. Throughout, we assume (M, ⊕) is a pseudo-reductive gyrogroup with the gyro-structure defined by Eq. (2.77)–Eq. (2.83).8 As Thm. 77 establishes the equivalence between gyrodistance and geodesic distance, we use the terms “gyromean” and “gyrovariance” interchangeably with their Riemannian counterparts. 3.3.3.1
GyroBN
To generalize the Euclidean BN in Eq. (3.1) to gyrogroups, we first define sample mean, sample variance, centering, biasing, and scaling over gyrogroups. Then, we introduce the GyroBN framework with a theoretical analysis of the ability to normalize sample statistics. We define the gyromean as the Fréchet mean [74] under gyrodistance: µ = FM({xi ∈ M}N i=1 ) = argmin y∈M
1 XN 2 d (xi , y) . i=1 N
(3.36)
The gyrovariance is the corresponding Fréchet variance. By Thm. 77, the gyromean and gyrovariance coincide with the Riemannian mean and variance, i.e., the Fréchet mean and variance under geodesic distance. The existence and local uniqueness of the Fréchet mean are reviewed in Thm. 32. Easy computation shows that the Euclidean BN operations in Eq. (3.1) have direct gyrogroup counterparts in Rn . Centering corresponds to gyrosubtraction (Eqs. (2.77) and (2.79)), biasing to gyroaddition (Eq. (2.77)), and scaling to scalar gyromultiplication (Eq. (2.78)). Motivated by this, we define a normalization layer over gyrogroups via gyro operations. Given a batch of activations {xi }N i=1 ⊂ M, the core operations of GyroBN are Scaling Biasing Centering z }| { z}|{ z }| { s ∀i ≤ N, x̃i ← β⊕ √ ⊙ ⊖µ ⊕ xi , (3.37) v2 + ϵ
where µ ∈ M and v 2 are gyromean and gyrovariance, β ∈ M is the bias parameter, s ∈ R is the scaling parameter, and ϵ is a small value for numerical stability. The following theorem shows that Eq. (3.37) can normalize manifold-valued data. 8
In GyroBN, ⊙ is not required to satisfy the axioms of a gyrovector space (Thm. 53).
84
Chapter 3. Riemannian Batch Normalization Algorithm 2: Gyrogroup Batch Normalization (GyroBN) Require : batch of activations {xi }N i=1 ⊂ M, small positive constant ϵ, and momentum η ∈ [0, 1], running mean µr , running variance vr2 , bias parameter β ∈ M, scaling parameter s ∈ R. Return : normalized batch {x̃i }N i=1 ⊂ M if training then Compute batch mean µb and variance vb2 of {xi }N i=1 ; Update running statistics µr = Barη (µb , µr ), and vr2 = ηvb2 + (1 − η)vr2 ; (µ, v 2 ) = (µb , vb2 ) if training else (µr , vr2 ) ∀i ≤ N, x̃i = β ⊕ √vs2 +ϵ ⊙ (⊖µ ⊕ xi )
Theorem 83 (Homogeneity). [↓] Let (M, ⊕) be a pseudo-reductive gyrogroup in which every gyration gyr[·, ·] is a gyroisometry. For N samples {xi }N i=1 ⊂ M and any t ∈ R, we have N Homogeneity of gyromean: FM({β ⊕ xi }N i=1 ) = β ⊕ FM({xi }i=1 ),
Homogeneity of dispersion from e:
t 1 XN 2 d (t ⊙ xi , e) = i=1 N N
∀β ∈ M,
(3.38)
2 XN
i=1
d2 (xi , e), (3.39)
The most important property of the Euclidean BN [109] lies in its ability to normalize the sample mean and variance. Thm. 83 shows that the formulation in Eq. (3.37) enjoys the same property: homogeneity of the gyromean guarantees that centering and biasing shift the gyromean, while homogeneity of the dispersion from e ensures that scaling controls the sample variance. As a result, GyroBN provides a theoretical guarantee of normalization on any pseudo-reductive gyrogroup with isometric gyrations. Moreover, since the gyromean and gyrovariance coincide with their Riemannian counterparts, GyroBN also normalizes Riemannian statistics. To finalize GyroBN, we define the running mean updates over gyrogroups as the binary barycenter based on gyrodistance: Barη (x1 , x2 ) = argminy∈M η d2 (x1 , y) + (1 − η) d2 (x2 , y) ,
η ∈ [0, 1],
(3.40)
which can be calculated by the geodesic. With these ingredients, the general framework for GyroBN is presented in Alg. 2. In particular, it recovers the classic Euclidean 85
3.3. Gyrogroup Batch Normalization BN [109] when M = Rn . Remark 84. We make the following two remarks with respect to left gyrotranslation. • Other candidates. There are three alternatives to left gyrotranslation. However, they are not necessarily gyroisometries and therefore cannot support the general GyroBN construction without additional isometry assumptions. Specifically, analogous to the left gyrotranslation, the right gyrotranslation is Rx : G → G, Rx (y) = y ⊕ x, ∀y ∈ G. (3.41) Along with the gyroaddition, there is the gyrogroup coaddition [200, Def. 2.9]: x ⊞ y = x ⊕ gyr[x, ⊖y]y,
∀x, y ∈ G.
(3.42)
Coaddition is symmetric to gyroaddition in many ways; for instance, when gyroaddition is gyrocommutative, coaddition is commutative [200, Thm. 3.3]. However, the right gyrotranslation, as well as the left and right translations by coaddition, are not guaranteed to be gyroisometries. Numerical experiments confirm that these three translations on the hyperbolic Poincaré ball fail to preserve gyrodistance (see gyrocoadd.py). A theoretical reason is that they all lack a counterpart of the left gyrotranslation law, which is important for the translation to be a gyroisometry (see Fig. 3.7). • Special cases. For Lie groups with a right-invariant Riemannian metric, right translations are isometries, and GyroBN can then be formulated using them, with Thm. 83 extending directly. In general, however, only left gyrotranslations are guaranteed to be isometries.
3.3.4
Instantiations
As indicated by Thms. 81 and 83, GyroBN can be applied to different geometries with guaranteed normalization of the sample statistics. Once the required operators are specified, Alg. 2 can be used in a plug-and-play manner. We first show that existing RBN methods with control over sample statistics, such as LieBN on Lie groups and AIM-based SPDBNs, are special cases of the framework. We then instantiate GyroBN on seven representative geometries: the Grassmannian, five CCS models (Poincaré ball, projected hypersphere, hyperboloid, sphere, and Beltrami–Klein), and the full-rank 86
Chapter 3. Riemannian Batch Normalization correlation manifold. To support these instantiations, we simplify the Grassmannian operators for efficient computation, refine the gyro-structures on the Poincaré ball and projected hypersphere, establish new gyro-structures for the hyperboloid and sphere, characterize the Beltrami–Klein geometry, and formulate a row-wise realization for the correlation manifold. 3.3.4.1
LieBN as a Special Case
Chakraborty [34, Algs. 3–4] introduced the Riemannian normalization on matrix Lie groups under a specific distance. The extension to general Lie groups, yielding LieBN with theoretical normalization over the Riemannian mean and variance, is presented in Sec. 3.2. This subsection shows that LieBN is a special case of GyroBN. LieBN is formulated with an invariant metric on a Lie group. Centering and biasing are performed through group translations, while scaling is defined in the tangent space at the identity element. Since every Lie group is automatically a gyrogroup, gyrotranslation reduces exactly to the group translation. Consequently, the centering, biasing, and scaling in LieBN are identical to those in GyroBN. Moreover, as shown in Thm. 77, the mean, variance, and running mean update defined via the geodesic distance in LieBN are equivalent to their counterparts based on gyrodistance in GyroBN. Therefore, LieBN is a special case of GyroBN. The LieBN instantiations presented in Sec. 3.2.5 are therefore incorporated by GyroBN. 3.3.4.2
AIM-Based SPDBNs as Special Cases
As reviewed in Sec. 3.2.3.2, several RBNs on the SPD manifold were developed based on AIM [31, 124, 123]. These SPD normalization methods can be expressed in the following unified form: Normalization: ∀i ≤ N,
√ s 1 1 v 2 +ϵ − 12 − 12 e 2 B2, Pi ← B M Pi M
(3.43)
where M and v 2 are the Riemannian mean and variance. The running mean is updated by the binary barycenter under the geodesic distance. As gyrodistance is identical to the geodesic distance, the gyromean, gyrovariance, and running mean updates are identical to the Riemannian ones. Using the AIM gyrooperations in Eq. (2.95), Eq. (3.43) is exactly the specific implementation of Eq. (3.37) under the AIM-based gyrogroup on the SPD manifold. Therefore, the SPDBNs developed by Brooks et al. [31], Kobler et al. [124, 123] are also special cases of GyroBN. 87
3.3. Gyrogroup Batch Normalization Remark 85. Brooks et al. [31] only considers centering and biasing. Kobler et al. [124] uses a running mean for centering during training. Kobler et al. [123] uses different momentum parameters to update running statistics for training and testing, along with multi-channel mechanisms for domain adaptation. Nevertheless, all of them are based on Eq. (3.43). Therefore, tricks such as multi-channel and separate momentum can also be applied to GyroBN. This is what we mean by claiming that GyroBN incorporates their approaches. 3.3.4.3
Grassmannian Manifold
We focus on the ONB perspective. Let U, V ∈ Gr(p, n), t ∈ R, and ∆ ∈ TU Gr(p, n). Using the Grassmannian operators reviewed in Tabs. 2.10 and 2.11, we instantiate GyroBN on the ONB Grassmannian. Instantiation. Building on these operators, we now implement the ONB Grassmannian GyroBN. Given a batch of activations {U1···N }, the three core steps of GyroBN are Centering to the identity Ip,n : Ui1 = exp
−[M M ⊤ , Iep,n ]
Ui ,
s √ [Ui1 (Ui1 )⊤ , Iep,n ] Ip,n , v2 + ϵ Biasing towards B ∈ M: Ui3 = exp [B, Iep,n ] Ui2 .
Scaling the dispersion from Ip,n : Ui2 = exp
(3.44) (3.45) (3.46)
g e (·) is the Riemannian logarithm under the PP Grassmannian, M Here (·) = Log Ip,n 2 ⊤ (resp. v ) is the Riemannian batch mean (resp. variance), and Iep,n = Ip,n Ip,n is the PP identity. The mean M can be obtained by the Karcher flow [116], with the Riemannian logarithm computed by Bendokat et al. [19, Alg. 5.3]. Efficient computation. The commutators [M M ⊤ , Iep,n ] and [U 1 (U 1 )⊤ , Iep,n ] can be i
i
efficiently computed by the following result.
Proposition 86. [↓] Given U = (U1⊤ , U2⊤ )⊤ ∈ Gr(p, n) with U1 ∈ Rp×p and U2 ∈ R(n−p)×p , then ! e2⊤ 0 − U p×p [U U ⊤ , Iep,n ] = , (3.47) e2 0(n−p)×(n−p) U SVD
e2 = U2 Q arcsin(Ŝ) R⊤ and U1⊤ := QSR⊤ . Here S is in ascending order, Q where U Ŝ p and R are flipped column-wise, and Ŝ = Ip − S 2 . 88
Chapter 3. Riemannian Batch Normalization Remark 87. Two technical issues are worth noting. • Cut locus. The logarithm LogU (V ) exists only when U and V are not in each other’s cut locus [19]. Similarly, gyroaddition and gyromultiplication are not globally defined due to the cut locus [157, Sec. 3.2]. However, Bendokat et al. [19, Alg. 5.3] provides a numerical remedy. • PP Grassmannian. Although our derivation is based on the ONB Grassmannian, GyroBN under the PP Grassmannian can be obtained by mapping f n) → Gr(p, n), normalizing, and mapping back via π. data via π −1 : Gr(p, f n). This follows from the isometry π : Gr(p, n) → Gr(p,
3.3.4.4
Stereographic Model
We begin by analyzing its gyro-structure and then instantiate GyroBN. Stereographic gyrovector space. As reviewed in Sec. 2.9.5, the stereographic model stnK unifies CCS geometries: the hyperbolic Poincaré ball PnK for K < 0, Euclidean space Rn for K = 0, and the spherical projected hypersphere DnK for K > 0. For x, y, z ∈ stnK , t ∈ R, and v ∈ Tx stnK , its Riemannian and gyro operators are reviewed in Tabs. 2.13 and 2.14. For K < 0, the Poincaré ball (PnK , ⊕K , ⊙K ) forms a Möbius gyrovector space, as reviewed in Tab. 2.13. For K > 0, however, the situation is subtler [13]: y • gyroaddition is well-defined except when x = K∥y∥ 2;
√ • gyromultiplication is well-defined except when r tan−1 ( K ∥x∥) = π/2 + kπ for some k ∈ Z.
Even if assumed well-defined, it remains unclear whether (stnK , ⊕K , ⊙K ) with K > 0 satisfies the axioms of a gyrovector space, in contrast to the Poincaré ball. Closing this gap is part of our contribution: under the assumption of well-definedness, we show that the stereographic gyro operations coincide with those in Eqs. (2.77) and (2.78) and prove that the stereographic model satisfies all axioms of a gyrovector space. Proposition 88. [↓] The stereographic gyroaddition and gyromultiplication coincide with the Riemannian definitions: x ⊕K y = Expx (PT0→x (Log0 (y))) , t ⊙K x = Exp0 (t Log0 (x)) , 89
∀x, y ∈ stnK ,
∀t ∈ R, x ∈ stnK .
(3.48) (3.49)
3.3. Gyrogroup Batch Normalization Theorem 89. [↓] For any K ∈ R, the stereographic model (stnK , ⊕K ) satisfies all axioms of a gyrocommutative gyrogroup. When further endowed with gyromultiplication ⊙K , it satisfies all axioms of a gyrovector space. Stereographic GyroBN. Next, we extend Thm. 81 to the stereographic model with arbitrary curvature. Theorem 90. [↓] For the stereographic model, the gyrodistance coincides with the geodesic distance. Moreover, the gyroinverse, any gyration, and any left gyrotranslation are gyroisometries. Combining Thm. 83 and Thm. 90, GyroBN in the stereographic model is theoretically guaranteed to normalize sample statistics. Practically, implementation only requires substituting the stereographic operators reviewed in Tabs. 2.13 and 2.14 into Alg. 2. For efficient computation, the Poincaré Fréchet mean can be obtained using the algorithm of Lou et al. [142, Alg. 1], while the mean on the sphere is computed via the Karcher flow [116]. 3.3.4.5
Radius Model
As reviewed in Sec. 2.9.5, the radius model MnK is another representation of constantcurvature spaces, unifying the hyperboloid HnK for K < 0, the sphere SnK for K > 0, and the Euclidean space Rn for K = 0. The case K = 0 reduces trivially to Euclidean space, so the following analysis of the radius-model gyro-structure considers K ̸= 0. Although this model has been effective in various applications [38, 45, 15, 166, 99, 119], its gyro-structure has not been formalized. We first analyze the gyro-structure on MnK and then instantiate GyroBN. Radius gyrovector space. The radius model MnK is isometric to the stereographic model stnK via stereographic projection fixing the south pole [182]: "
# ξ ∈ R x p πMnK →stnK : MnK ∋ 7−→ ∈ stnK , n x∈R 1 + |K|ξ 2 √1 1−K∥y∥ 1+K∥y∥2 ∈ MnK . πstnK →MnK : stnK ∋ y 7−→ |K| 2y 1+K∥y∥2
(3.50) (3.51)
p The origin in MnK is defined as 0 = [ 1/|K|, 0, . . . , 0]⊤ , corresponding to 0 ∈ stnK . For K > 0, we have MnK = SnK and stnK = DnK . Since Eq. (3.50) is undefined at the 90
Chapter 3. Riemannian Batch Normalization south pole −0 when K > 0, we use the one-point compactification DnK ∪ {∞} with the identification πMnK →stnK (−0) = ∞ [182, Rmk. A.9]. For simplicity, we use DnK and DnK ∪ {∞} interchangeably.
Tab. 2.14 reviews the Riemannian operators. We adopt the following curvatureaware functions: cos sin if K > 0, if K > 0, cosK = sinK = cosh if K < 0, sinh if K < 0, (3.52) ⟨·, ·⟩ if K > 0, ⟨·, ·⟩K = ⟨·, ·⟩ if K < 0. L
p For v ∈ Tx MnK , ∥v∥K = ⟨v, v⟩K is the induced Riemannian norm. On the K < 0 branch, the ambient Lorentzian form is positive definite only after this tangent-space restriction. Moreover, (·)s denotes the space vector, and (·)t denotes the time scalar. Using the radius-model Riemannian operators reviewed in Tab. 2.14, we define the gyroaddition and gyromultiplication as Eqs. (2.77) and (2.78): x ⊕M K y = Expx (PT0→x (Log0 (y))) , t ⊙M K x = Exp0 (t Log0 (x)) ,
∀x, y ∈ MnK ,
∀t ∈ R, ∀x ∈ MnK .
(3.53) (3.54)
We give the following clarifications regarding the above two gyro operations. • For hyperbolic geometry (K < 0), Eq. (3.53) has been employed in prior work [38, 99]. However, there exists no closed-form expression, which could be more efficient than the composition of Riemannian operators. Besides, the underlying gyrostructure has not been formally discussed. • For spherical geometry (K > 0), the geodesic between antipodal points (x and −x) is not unique, making the logarithm and parallel transport along such a geodesic ill-defined. For the sphere, Eq. (3.53) assumes that x ̸= −0 and y ̸= −0, while Eq. (3.54) assumes that x ̸= −0. Like before, we always make these assumptions implicitly. In the following, we first give the closed-form expressions of Eqs. (3.53) and (3.54). Then, we show that Eqs. (3.53) and (3.54) conform to all the axioms of a gyrovector space. 91
3.3. Gyrogroup Batch Normalization ⊤ Proposition 91 (Gyromultiplication and gyroinverse). [↓] Let x = [xt , x⊤ s ] be a point in MnK , where xt ∈ R is the time scalar, and xs ∈ Rn is the spatial part. The gyromultiplication and inverse have closed-form expressions:
0,
p |K|x ) cosK t cos−1 ( t M K t ⊙K x = p 1 , −1 √ sin t cos ( |K|x ) K t |K| K xs ∥x ∥ " # s xt M ⊖M . K x = −1 ⊙K x = −xs
t = 0 ∨ x = 0, (3.55)
t ̸= 0,
(3.56)
In particular, the gyro identity is 0. Besides, the following shows that the sphere gyromultiplication is still valid for the singular cases in the projected hypersphere. Assume K > 0 and x ̸= ±0, and let stnK ∋ u = πMnK →stnK (x) be its stereographic √ Kxt ∈ (0, π). Then the following are equivaimage. For t ∈ R, set θ = cos−1 lent: (i) t tan−1
√
K ∥u∥ = π2 + kπ ⇐⇒ (ii) tθ = (2k + 1)π,
k ∈ Z.
(3.57)
In these singular cases, Eq. (3.55) is still valid, while the stereographic gyromultiplication returns infinity: n t ⊙M K x = −0 ∈ MK ,
t ⊙K u = ∞ ∈ stnK ,
(3.58)
where ∞ denotes the added point in the one-point compactification DnK ∪{∞} (stnK = DnK for K > 0) and πMnK →stnK (−0) = ∞. ⊤ ⊤ ⊤ Proposition 92 (Gyroaddition). [↓] Let x = [xt , x⊤ s ] and y = [yt , ys ] be points in MnK , where xt , yt ∈ R are the time scalars, and xs , ys ∈ Rn are the spatial parts.
92
Chapter 3. Riemannian Batch Normalization Then, the gyroaddition on MnK admits the closed form:
x, y = 0, y, x = 0, M x ⊕K y = 1 D−KN √ |K| D+KN , otherwise. 2(As xs +Ay ys )
(3.59)
D+KN
Here, As = ab2 − 2Kbsxy − Kany and Ay = b(a2 + Knx ), where the following quantities are defined by a=1+
p p |K|xt , b = 1 + |K|yt , nx = ∥xs ∥2 , ny = ∥ys ∥2 , sxy = ⟨xs , ys ⟩. (3.60) D = a2 b2 − 2Kabsxy + K 2 nx ny , N = a2 ny + 2absxy + b2 nx .
(3.61)
Besides, the following shows that the sphere gyroaddition is still valid under the singular cases in the projected hypersphere. Assume K > 0 and x, y ̸= ±0, and let u = πMnK →stnK (x) and v = πMnK →stnK (y) be the stereographic images. The following statements are equivalent: v (v ̸= 0); (1) u = K ∥v∥2 (2) xs = ys and xt = −yt (same meridian, mirrored across the equator); (3) D = 0. In such singular cases, we have N > 0 and Eq. (3.59) is still valid, while the stereographic gyroaddition returns infinity: x ⊕M K y = −0,
u ⊕K v = ∞,
(3.62)
where ∞ is the point added in the one-point compactification of DnK . The above two propositions immediately imply that πMnK →stnK preserves the gyro operations. Corollary 93 (Isomorphism). [↓] For the hyperbolic geometry (K < 0), the isom-
93
3.3. Gyrogroup Batch Normalization etry πMnK →stnK : HnK → PnK preserves the gyro operations: n x ⊕M πMnK →stnK (x) ⊕K πMnK →stnK (y) , ∀x, y ∈ MnK , K y = πstn K →MK n r ⊙M r ⊙K πMnK →stnK (x) , ∀r ∈ R, ∀x ∈ MnK . K x = πstn K →MK
(3.63)
For the spherical geometry (K > 0), the isometry πMnK →stnK : SnK → DnK ∪ {∞} preserves the gyro operations: n x ⊕M πMnK →stnK (x) ⊕K πMnK →stnK (y) , ∀x, y ∈ MnK /{−0}, K y = πstn K →MK (3.64) n n n n n /{−0}. r ⊙M x = π r ⊙ π (x) , ∀r ∈ R, ∀x ∈ M stK →MK K MK →stK K K Remark 94. We provide two clarifications regarding Thm. 93. • Compactified projected hypersphere. Although stereographic operations for K > 0 may be undefined in certain cases, they become well-defined on the one-point compactification DnK ∪ {∞}, where undefined cases correspond to ∞. • Sphere vs. projected hypersphere. Thm. 93 suggests a numerical advantage of the sphere: gyro-operations are well-defined on SnK at all points except the single south pole −0, including cases corresponding to singularities on the projected hypersphere. This broader domain can make computations on SnK more stable. M From the above corollary, it is expected that the operations ⊕M K and ⊙K also satisfy the axioms of a gyrovector space for both negative and positive curvature K.
Theorem 95 (Radius gyrovector spaces). [↓] (MnK , ⊕M K ) forms a gyrocommutative M M n gyrogroup, and (MK , ⊕K , ⊙K ) forms a gyrovector space.a a
For K > 0, we implicitly assume the addition and multiplication are well-defined; whenever they are, all corresponding axioms hold.
Due to the isometry, Thm. 78 implies gyroisometries over MnK . Theorem 96. On the radius model, the gyrodistance is identical to the geodesic distance, whereas the gyroinverse, gyration, and left gyrotranslation are gyroisometries. Radius GyroBN. Thm. 96 guarantees that GyroBN on the radius model normal-
94
Chapter 3. Riemannian Batch Normalization izes the sample mean and variance. Substituting the radius Riemannian operators from Tab. 2.14 and the gyro operators from Thms. 91 and 92 into Alg. 2 can directly yield the radius GyroBN. For K < 0 (hyperboloid), the Fréchet mean can be computed efficiently by Lou et al. [142, Alg. 3]; for K > 0 (sphere), we compute the Fréchet mean using the Karcher flow [116]. 3.3.4.6
Hyperbolic Beltrami–Klein
The Beltrami–Klein model has recently emerged as a promising alternative to the Poincaré ball for representing hyperbolic geometry [147]. As reviewed in Sec. 2.9.5, the Poincaré ball admits the Möbius gyrovector space and the Beltrami–Klein model admits the Einstein gyrovector space. While prior studies mainly focused on the case K = −1 [147], we develop the Beltrami–Klein Riemannian structure under arbitrary negative curvature and relate it to the Einstein gyrospace. This allows us to establish the equivalence between the gyro and Riemannian formulations and, ultimately, to instantiate GyroBN on this model. Beltrami–Klein Riemannian structure. The standard Einstein gyro operations on the Beltrami–Klein model are reviewed in Tab. 2.13. For K = −1, prior work related the Einstein operations to the Riemannian-form gyro operations in Eqs. (2.77) and (2.78) [147, Sec. 4.2]. We extend this equivalence to arbitrary K < 0 and derive the corresponding closed-form Beltrami–Klein Riemannian operators. To this end, we first establish the isometry between the Beltrami–Klein and Poincaré models. Proposition 97 (Beltrami–Klein isometries). [↓] The following maps are Riemannian isometries between the Beltrami–Klein and Poincaré ball models: 1 q x ∈ PnK , 2 1 + 1 + K ∥x∥ 2 πPnK →KnK : PnK ∋ x 7−→ x ∈ KnK . 1 − K ∥x∥2
πKnK →PnK : KnK ∋ x 7−→
(3.65) (3.66)
In particular, πPnK →KnK (0) = 0. Given x in the hyperbolic model H ∈ {KnK , PnK } and tangent vector v ∈ Tx H, the differential maps of πKnK →PnK and πPnK →KnK are (πKnK →PnK )∗,x (v) =
K ⟨x, v⟩ 1 q v− x, 2 q q 2 2 2 1 + 1 + K ∥x∥ 1 + 1 + K ∥x∥ 1 + K ∥x∥ 95
3.3. Gyrogroup Batch Normalization
(πPnK →KnK )∗,x (v) =
2 4K ⟨x, v⟩ 2 x. 2 v + 1 − K ∥x∥ 1 − K ∥x∥2
In particular, the differential maps at the zero vector are 1 (πKnK →PnK )∗,0 (v) = v, 2 n n (πPK →KK )∗,0 (v) = 2v.
(3.67) (3.68)
Moreover, these isometries preserve gyroaddition and gyromultiplication: πPnK →KnK (x ⊕M y) = πPnK →KnK (x) ⊕E πPnK →KnK (y), πPnK →KnK (t ⊙M x) = t ⊙E πPnK →KnK (x),
∀x, y ∈ PnK ,
∀t ∈ R, ∀x ∈ PnK ,
(3.69)
where ⊕M and ⊙M are Möbius operations, while ⊕E and ⊙E are the Einstein counterparts. Theorem 98 (Einstein by Beltrami–Klein). [↓] The Einstein gyro operations can be rewritten as Eqs. (2.77) and (2.78): x ⊕E y = Expx (PT0→x (Log0 (y))) , t ⊙E x = Exp0 (t Log0 (x)) ,
∀x, y ∈ KnK ,
∀x ∈ KnK , ∀t ∈ R.
(3.70) (3.71)
Thm. 98 demonstrates that the Einstein gyro-structure can be expressed by the Beltrami–Klein geometry. Conversely, the Beltrami–Klein geometry can also be formulated by the Einstein gyro-structure. Theorem 99 (Beltrami–Klein by Einstein). [↓] Given x, y ∈ KnK and v ∈ Tx KnK , the distance, exponential, and logarithmic operators under the Beltrami–Klein geometry are 2 d(x, y) = p |K|
p tanh−1 |K|
∥−x ⊕E y∥ , q 2 1 + 1 + K ∥−x ⊕E y∥
1 K ⟨x, v⟩ , q Expx (v) = x ⊕E Exp0 v − x q 2 2 2 1 + K ∥x∥ 1 + 1 + K ∥x∥ (1 + K ∥x∥ )
96
Chapter 3. Riemannian Batch Normalization
Logx (y) =
1 (πPnK →KnK )∗,ex (Log0 (−x ⊕E y)) , λxK e
where x e = πKnK →PnK (x). In particular, the exponential and logarithmic maps at the zero vector 0 are identical across the Beltrami–Klein and Poincaré ball models: p
v |K| ∥v∥) p , ∀v ∈ T0 H, |K| ∥v∥ p x , ∀x ∈ H, Log0 (x) = tanh−1 ( |K| ∥x∥) p |K| ∥x∥
Exp0 (v) = tanh(
(3.72) (3.73)
with H ∈ {KnK , PnK }.
Remark 100. Since the Beltrami–Klein and Poincaré ball models share the same Exp0 and Log0 , it naturally follows that the Einstein and Möbius gyromultiplication coincide. As the Beltrami–Klein model is isometric to the Poincaré ball, Thm. 78 implies gyroisometries. Theorem 101. On the Beltrami–Klein model, the gyrodistance is identical to the geodesic distance, whereas the gyroinverse, gyration, and left gyrotranslation are gyroisometries. Beltrami–Klein GyroBN. Thm. 101 ensures that GyroBN on the Beltrami–Klein model normalizes the sample mean and variance. Owing to the isometry between the Beltrami–Klein and Poincaré models, the Fréchet mean can be computed via the Poincaré ball: map the data to the Poincaré model using πKnK →PnK , compute the Poincaré Fréchet mean [142, Alg. 1], and map the result back using πPnK →KnK . Together with the gyro and Riemannian operators in Sec. 3.3.4.6, we have all the ingredients to implement Alg. 2. 3.3.4.7
Correlation Manifolds
As reviewed in Sec. 2.9.2, ECM, LECM, OLM, and LSM are correlation metrics induced from Euclidean, zero-curvature prototype spaces, and these four metrics have been instantiated for LieBN in Sec. 3.2.5.3. Here, we further handle the nonzerocurvature correlation metric, namely PHCM. Under PHCM, any correlation matrix can be identified with a product of hyperbolic spaces via its Cholesky decomposition. Given C ∈ Cor+ (n), let L = Chol(C) be its Cholesky factor. The k-th row of L has 97
3.3. Gyrogroup Batch Normalization the form (Lk1 , . . . , Lk,k−1 , Lkk , 0, . . . , 0) with Lkk > 0, which belongs to the hyperbolic open hemisphere9 HSk−1 = x ∈ Rk | ∥x∥ = 1, xk > 0 . (3.74) As detailed in Sec. 5.4, HSn is isometric to the unit Poincaré ball Pn = {x ∈ Rn | ∥x∥ < 1} by πHSn →Pn
"
x xn+1
#!
=
x . 1 + xn+1
(3.75)
(3.76)
Therefore, each correlation matrix can be identified with n − 1 Poincaré vectors:
1 0 L21 L22 Cor+ (n) ∋ C7→ .. .. . . Ln1 Ln2
0 1 x ∈ P 1 0 .. . → 7 . .. . n−1 xn−1 ∈ P · · · Lnn ··· ··· ...
(3.77)
⊤ Here, xi = πHSi →Pi L(i+1,1) , · · · , L(i+1,i+1) corresponds to the (i + 1)-th row of the Q i Cholesky factor. Let PPn−1 = n−1 i=1 P denote the product of unit Poincaré balls. We denote the identification in Eq. (3.77) by Φ : Cor+ (n) → PPn−1 . GyroBN on the correlation manifold can then be realized via the Poincaré GyroBN applied row-wise: first map C to PPn−1 via Φ, apply GyroBNi independently on each Pi , and finally map + back with Φ−1 . For a batch of activations {C i }N i=1 ⊂ Cor (n), the process can be expressed as
∀i ≤ N,
3.3.4.8
GyroBN 7−→ 1 x ei1 ∈ P1 xi1 ∈ P1 Φ−1 i Φ .. .. .. e. 7−→ C C i 7−→ . . . GyroBNn−1 x ein−1 ∈ Pn−1 xin−1 ∈ Pn−1 7−→
(3.78)
Summary
To conclude this section, Tab. 3.14 summarizes the key gyro operators needed to implement GyroBN on representative manifolds. The correlation manifold is excluded, since its GyroBN is realized row-wise through the Poincaré ball. 9
Also known as the Jemisphere model, where the “J” is pronounced as in Spanish [33, Sec. 7].
98
Chapter 3. Riemannian Batch Normalization U ⊕Gr V
⊖Gr U
t ⊙Gr U
Fréchet mean
exp(ΩU )V
exp (−ΩU ) Ip,n
exp (tΩU ) Ip,n
Karcher flow
(a) The ONB Grassmannian Operator
stnK
x⊕y
(1 − 2K⟨x, y⟩ − K∥y∥2 )x + (1 + K∥x∥2 )y 1 − 2K⟨x, y⟩ + K 2 ∥x∥2 ∥y∥2
⊖x
−x p tanK t tan−1 |K|∥x∥ K x p ∥x∥ |K| K < 0: [142, Alg. 1] K > 0: Karcher flow
t⊙x Fréchet mean
MnK
KnK
Eq. (3.59) " # xt −xs
Tab. 2.13
Eq. (3.55)
Tab. 2.13
K < 0: [142, Alg. 3] K > 0: Karcher flow
Via Poincaré ball (see Sec. 3.3.4.6)
−x
(b) Constant curvature spaces
Table 3.14: Summary of operators for GyroBN across representative manifolds.
3.3.5
Experiments
GyroBN layers are backbone-agnostic and can be integrated into different networks whenever the underlying manifold admits the required gyro operations. This subsection evaluates GyroBN on the Grassmannian, five CCSs, and the correlation manifold. The main findings are summarized as follows. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.2. • Numerical experiments (Sec. 3.3.5.1). The closed-form expression derived for radius gyroaddition in Eq. (3.59) significantly accelerates the computation compared with its definition-based counterpart in Eq. (3.53), achieving 2×–3× speedups. Furthermore, visualizations demonstrate that GyroBN effectively normalizes sample distributions across diverse geometries. • Performance (Secs. 3.3.5.2 to 3.3.5.4). On networks over the Grassmannian, CCSs, and the correlation manifold, GyroBN generally improves backbone networks, whereas existing RBN methods often degrade performance. Compared to these methods, GyroBN is generally faster or comparable in runtime, requires fewer or equal parameters, and improves robustness. 3.3.5.1
Numerical Experiments
Efficiency of the closed-form radius gyroaddition. Recalling Sec. 3.3.4.5, we derive closed-form expressions for the radius gyrooperations to improve computational 99
3.3. Gyrogroup Batch Normalization Geometry
Hyperboloid
Sphere
Dim
Riemannian
Closed form
Riemannian
Closed form
16 32 64 128 256 1024 2048
361.22 363.41 387.94 574.37 1149.58 3364.23 6479.95
121.68 (33.69%) 123.08 (33.87%) 181.50 (46.78%) 272.10 (47.37%) 534.09 (46.46%) 1414.67 (42.05%) 2497.47 (38.54%)
323.58 327.75 453.42 644.18 1137.62 3551.69 6930.69
122.02 (37.71%) 123.71 (37.74%) 181.05 (39.93%) 270.21 (41.95%) 538.58 (47.34%) 1421.67 (40.03%) 2449.56 (35.34%)
Table 3.15: Efficiency (in µs) of gyroaddition on the radius manifold: closed form versus Riemannian definition. Values in parentheses indicate the runtime of the closed-form implementation as a percentage of the corresponding Riemannian implementation. The best results are bold. efficiency. To assess this, we compare two variants of radius gyroaddition: (i) the definition-based operator Eq. (3.53), implemented via a composition of the Riemannian logarithm, parallel transport, and exponential map; and (ii) the closed-form operator Eq. (3.59). We report the mean wall-clock time (in µs), averaged over 100 runs with a batch size of 10,000, across varying dimensions. As shown in Tab. 3.15, the closed-form implementation consistently outperforms its definition-based counterpart, achieving speedups of roughly 2×–3× across the hyperboloid and sphere. Visualization of GyroBN on different geometries. To intuitively illustrate the effect of GyroBN, we visualize its behavior on the ONB Grassmannian Gr(1, 3), five CCSs with |K| = 1, and the correlation manifold Cor+ (3). For each geometry, we randomly generate a batch of points, compute their batch mean, apply GyroBN, and then plot the normalized batch along with the resulting mean. The visualizations are constructed using the following embeddings: • Grassmannian. Since Gr(1, 3) is homeomorphic to the real projective space RP2 , it is depicted as the unit hemisphere with antipodal points identified. • CCSs. The unit Poincaré ball P3 and unit Beltrami–Klein ball K3 are shown as the interior of the unit ball in R3 . The unit hyperboloid H2 is visualized as the upper sheet of a two-sheeted hyperboloid in R3 . The unit sphere S2 is embedded as a 2-sphere in R3 , while the projected hypersphere D3 coincides with R3 itself. • Correlation manifold. Cor+ (3) is embedded in R3 as an open elliptope using its strictly lower triangular part. For better visualization, we fix the bias parameter to the gyro identity element and set the scaling parameter to 0.4 for the Grassmannian, correlation manifold, and sphere; 0.7 for the Poincaré ball, Beltrami–Klein ball, and hyperboloid; and 1 for the projected 100
Chapter 3. Riemannian Batch Normalization
Figure 3.8: Visualization of GyroBN across different geometries. Blue and green points represent input and normalized data, respectively. Red and cyan points denote the input and output batch means. Black points mark the manifold boundary, and the gray surface depicts the manifold. hypersphere. As shown in Fig. 3.8, GyroBN consistently normalizes data distributions across these geometries. Notably, although the Poincaré and Beltrami–Klein inputs are identical, their GyroBN behavior and resulting sample distributions differ due to their distinct Riemannian metrics, underscoring that GyroBN faithfully respects the underlying geometry. 3.3.5.2
Experiments on Grassmannian Neural Networks
Data sets and preprocessing. In line with previous work [108, 159], we evaluate GyroBN on three skeleton-based action recognition tasks, including HDM05 [153], NTU60 [178], and NTU120 [138] data sets, focusing on mutual actions for NTU60 and NTU120. 101
3.3. Gyrogroup Batch Normalization Each sequence is represented as a Grassmannian matrix of size 93 × 10, 150 × 10, and 150 × 10 for HDM05, NTU60, and NTU120, respectively.
Comparative methods. Since no Grassmannian-specific BN methods exist, we adapt two previous approaches to the Grassmannian: ManifoldNorm [34, Algs. 1–2] and the RBN method of Lou et al. [142, Alg. 2], which we denote LRBN for clarity. Although neither was originally designed for the Grassmannian, they can be adapted by employing Riemannian operators such as geodesics, exponential/logarithmic maps, and parallel transport. The key difference is that GyroBN can normalize data distributions across different geometries, whereas the other two methods cannot. Backbone networks. We adopt the recently proposed GyroGr architecture [159] as the backbone, which is briefly reviewed in Sec. A.2.2. GyroGr replaces the non-intrinsic FRMap + ReOrth block in GrNet [108] with Grassmannian left gyrotranslation, thereby improving numerical stability and performance. It consists of three basic components: left gyrotranslation, pooling [108], and the Projection Map (ProjMap) [108], where ProjMap maps Grassmannian points to symmetric matrices for classification. We consider both the 1-block and L-block variants. The 1-block version is structured as: gyrotranslation → pooling → ProjMap → classification, where the classification is implemented as an FC layer with softmax. The L-block version stacks L blocks of gyrotranslation and pooling, followed by a final ProjMap and classification layer. Since each pooling step approximately halves the dimensionality, we omit the pooling operation in the last block when L > 1. Following Huang et al. [108], the number of channels is fixed to 8.
Implementation details. Following Nguyen and Yang [159], we use the Cayley map to approximate the matrix exponential of skew-symmetric matrices and apply the trivialization strategy reviewed in Sec. 2.7 to parameterize the Grassmannian variables in both the gyrotranslation and GyroBN layers. Specifically, each trainable Grassmannian parameter is represented by Euclidean coordinates through the exponential map at the identity. This allows direct use of PyTorch optimizers [169] and avoids direct Riemannian updates of these parameters. In contrast, we find that the Grassmannian LRBN benefits from Riemannian optimization. Thus, we employ Geoopt [125] to optimize its Grassmannian bias parameter. Similarly, we use Geoopt to update the orthogonal bias parameter in ManifoldNorm. For all models, the BN layer is inserted after the first pooling layer with a momentum of 0.1. Training uses SGD with a learning rate of 5e−2 , batch size 30, and 400, 200, and 200 epochs for HDM05, NTU60, and NTU120, respectively. All models are optimized with a standard cross-entropy loss. Following previous normalization methods on matrix manifolds [123, 215] and the LieBN implementation in Sec. A.3.1, we adopt a single Fréchet mean iteration. 102
Chapter 3. Riemannian Batch Normalization
Method
HDM05 (47 × 10)
NTU60 (75 × 10)
NTU120 (75 × 10)
Acc
Fit Time
#Params (M)
Acc
Fit Time
#Params (M)
Acc
Fit Time
#Params (M)
GyroGr GyroGr-ManifoldNorm GyroGr-LRBN
48.97±0.24 49.67±0.76 48.64±0.77
2.09 32.90 33.31
2.0744 2.0921 2.0781
70.13±0.16 68.56±0.43 67.77±0.52
28.16 232.60 238.53
0.5062 0.5512 0.5122
53.76±0.18 51.41±0.38 50.56±0.22
49.62 399.78 403.40
1.1812 1.2262 1.1872
GyroGr-GyroBN
51.89±0.37
3.04
2.0773
72.60±0.04
35.85
0.5114
55.47±0.10
67.37
1.1864
Table 3.16: Comparison of GyroBN against other Grassmannian BN methods under the GyroGr backbone. Here, accuracy is reported as a percentage, fit time denotes the average training time per epoch (s/epoch), and #Params is reported in millions. Values in parentheses specify the dimension of the Grassmannian input to the BN layer. The largest number of parameters is marked in red. Main results. We compare GyroBN with ManifoldNorm and LRBN under the 1-block GyroGr backbone. The 5-fold results are presented in Tab. 3.16. We have the following four findings, which highlight the effectiveness of GyroBN in facilitating network training. • Improved accuracy. Across all three data sets, GyroBN consistently improves performance, enhancing the accuracy of the vanilla GyroGr by 2.92, 2.47, and 1.71 percentage points on HDM05, NTU60, and NTU120, respectively. In contrast, both ManifoldNorm and LRBN often degrade performance, particularly on NTU60 and NTU120. This advantage comes from the theoretical guarantee of GyroBN for normalizing sample statistics, which is absent in the other two methods (see Tab. 3.13). • Enhanced efficiency. GyroBN is substantially more efficient than ManifoldNorm and LRBN. The efficiency gain is mainly attributed to: (i) replacing computationally expensive Riemannian operators (e.g., parallel transport, exponential/logarithmic maps) with simpler gyro operations, (ii) reducing matrix multiplications from n × p to (n − p) × p or p × p (see Thm. 86), and (iii) applying trivialization to avoid costly Riemannian optimization. • Improved parameter economy. GyroBN also requires fewer parameters than LRBN and ManifoldNorm. The key difference lies in the bias parameter. Due to the trivialization, GyroBN only needs an (n − p) × p Euclidean matrix, whereas LRBN requires an n × p Grassmannian matrix and ManifoldNorm requires an n × n orthogonal matrix. • Stronger generalization. As shown in Fig. 3.9, we observe that GyroBN can narrow the gap between training and testing accuracy, indicating a stronger generalization ability. 103
3.3. Gyrogroup Batch Normalization
Figure 3.9: Training and testing curves of 1-block GyroGr on two NTU data sets. 3.3.5.3
Experiments on Networks over CCSs
Data sets. Following Lou et al. [142], we focus on the link prediction task on four graph data sets: Cora [177], Disease [4], Airport [229], and PubMed [155]. Comparative methods. We compare GyroBN with LRBN [142, Alg. 2]. As the original LRBN is only implemented in the Poincaré ball, we extend it to the other four CCSs. In particular, the core difference between LRBN and GyroBN lies in normalization: GyroBN can normalize sample statistics across different geometries, while LRBN lacks this guarantee. Backbone networks. We use a Hyperbolic Neural Network (HNN) [76] for the Poincaré ball and Klein HNN (KNN) [147] for the Beltrami–Klein. For the sphere and projected hypersphere, we mimic the transformation and activation in HNN [76, Sec. 3.2] to build the corresponding layers. The layers on the four spaces above can be expressed as Transformation: xk = Expe M k Loge (xk−1 ) ⊕ bk , with bk ∈ N , and M k ∈ Rm×n , Activation: xk = Expe ϕ Loge (xk−1 ) , with ϕ as an activation,
where e is the origin, ⊕ is the gyroaddition, and N ∈ {PnK , KnK , SnK , DnK }. For the 104
Chapter 3. Riemannian Batch Normalization hyperboloid, we use the Lorentz fully-connected layer [45, Eq. 3] and Lorentz activation layer [15, Eq. 13], which are briefly reviewed in Sec. A.2.4. The above backbone network is referred to as HNN, KNN, SNN, PHNN, and LNN, respectively. We collectively call them Constant Curvature Neural Networks (CCNNs). Implementation details on CCNNs. We follow the official implementations of HGCN10 [38], LRBN11 [142], and HCNN12 [15] to conduct experiments, where we adopt the same training settings as Lou et al. [142, Sec. H.1]. Specifically, the baseline encoder is a CCNN with two transformation layers: the first maps the input feature dimension to 128, and the second maps 128 to 128. After each transformation layer, we use a ReLU activation, except for the Cora data set where activation is omitted. A BN layer, GyroBN or LRBN, is inserted after each transformation layer. The curvature is set as |K| = 1. For the manifold-valued bias parameter in GyroBN and LRBN, we apply the exponential map Expe (v) to trivialize it via a Euclidean parameter v. The optimization is performed with Adam [121], using a learning rate of 1e−2 and a weight decay of 1e−3 , except for the Cora data set, where weight decay is set to 0. The Fréchet mean iterations are performed until convergence. Main results. We compare GyroBN with LRBN across five CCSs under the CCNN backbone. Tab. 3.17 reports the 5-fold average testing AUC on four data sets. We highlight the following findings. • Improved performance. GyroBN consistently improves performance over the vanilla CCNNs across all data sets and geometries, whereas LRBN degrades performance in several cases (highlighted in red). The gains are especially pronounced on the sphere (SNN), where GyroBN achieves improvements of 17.65 (Disease), 7.49 (Airport), and 13.37 (PubMed) percentage points. This contrast underscores the advantage of GyroBN’s theoretical guarantee of normalizing sample statistics. • Efficiency. As shown in the “Fit Time” rows of Tab. 3.17, GyroBN is more efficient than LRBN on the Poincaré ball, Beltrami–Klein and projected hypersphere, due to the simplicity of gyro operations. On the hyperboloid and sphere, GyroBN and LRBN exhibit comparable efficiency. • Parameter equivalence. GyroBN and LRBN require the same number of parameters, only marginally more than those of the vanilla backbone. Thus, the 10
https://github.com/HazyResearch/hgcn https://github.com/CUAI/Differentiable-Frechet-Mean 12 https://github.com/kschwethelm/HyperbolicCV 11
105
3.3. Gyrogroup Batch Normalization Space
Poincaré Ball PnK
Method
Hyperboloid HnK
Beltrami–Klein KnK
HNN
HNN-LRBN
HNN-GyroBN
LNN
LNN-LRBN
LNN-GyroBN
KNN
KNN-LRBN
KNN-GyroBN
Disease
ROC Fit Time #Params (M)
79.21 ± 2.14 0.027 0.0180
76.58 ± 2.15 0.088 0.0183
81.18 ± 0.93 0.084 0.0183
87.71 ± 1.42 0.012 0.0184
85.31 ± 0.95 0.057 0.0187
88.87 ± 0.33 0.057 0.0187
81.31 ± 1.37 0.020 0.0180
78.98 ± 1.51 0.089 0.0183
81.56 ± 0.70 0.077 0.0183
Airport
ROC Fit Time #Params (M)
94.63 ± 0.19 0.054 0.0182
94.17 ± 0.40 0.122 0.0184
95.40 ± 0.17 0.119 0.0184
93.86 ± 0.21 0.047 0.0186
93.05 ± 1.00 0.089 0.0188
95.06 ± 0.12 0.092 0.0188
95.00 ± 0.05 0.056 0.0182
94.47 ± 0.33 0.136 0.0184
96.14 ± 0.05 0.126 0.0184
PubMed
ROC Fit Time #Params (M)
95.02± 0.42 0.125 0.0806
93.40 ± 0.20 0.342 0.0809
95.83 ± 0.11 0.335 0.0809
95.36 ± 0.10 0.111 0.0815
95.88 ± 0.09 0.242 0.0818
95.89 ± 0.11 0.249 0.0818
95.87 ± 0.11 0.125 0.0806
89.83 ± 0.27 0.353 0.0809
96.23 ± 0.12 0.337 0.0809
Cora
ROC Fit Time #Params (M)
89.96 ± 0.44 0.032 0.2001
93.47 ± 0.49 0.091 0.2003
94.32 ± 0.22 0.076 0.2003
91.84 ± 1.01 0.114 0.2019
92.62 ± 0.17 0.244 0.2021
93.66 ± 0.30 0.242 0.2021
90.03 ± 0.32 0.038 0.2001
93.38 ± 0.12 0.091 0.2003
93.48 ± 0.25 0.076 0.2003
(a) Results on three hyperbolic spaces. Space Method
Projected Hypersphere DnK
Sphere SnK
PHNN
PHNN-LRBN
PHNN-GyroBN
SNN
SNN-LRBN
SNN-GyroBN
Disease
ROC Fit Time #Params (M)
69.70 ± 2.01 0.022 0.0180
60.25 ± 1.25 0.066 0.0183
72.26 ± 0.61 0.063 0.0183
54.19 ± 2.21 0.029 0.0181
53.38 ± 4.07 0.042 0.0183
71.84 ± 0.89 0.043 0.0183
Airport
ROC Fit Time #Params (M)
89.60 ± 0.99 0.056 0.0182
87.06 ± 0.46 0.102 0.0184
90.44 ± 0.93 0.095 0.0184
83.63 ± 0.77 0.053 0.0182
86.14 ± 0.79 0.070 0.0184
91.12 ± 1.57 0.071 0.0184
PubMed
ROC Fit Time #Params (M)
89.86 ± 0.39 0.121 0.0806
90.06 ± 0.23 0.175 0.0809
92.06 ± 0.64 0.174 0.0809
79.94 ± 1.76 0.120 0.0806
90.10 ± 0.21 0.140 0.0809
93.31 ± 0.08 0.156 0.0809
Cora
ROC Fit Time #Params (M)
92.88 ± 0.26 0.026 0.2001
92.03 ± 0.42 0.067 0.2003
93.26 ± 0.42 0.063 0.2003
92.10 ± 0.40 0.025 0.2001
82.01 ± 0.71 0.044 0.2003
93.16 ± 0.32 0.045 0.2003
(b) Results on two spherical spaces.
Table 3.17: Comparison of GyroBN against LRBN across five CCSs. ROC is the testing AUC reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. When LRBN degenerates the backbone network, the results are highlighted with red. performance gains of GyroBN cannot be attributed to parameter size, but rather to its principled normalization mechanism. 3.3.5.4
Experiments on Correlation Neural Networks
Data sets and preprocessing. Following the CorNet protocol in Sec. 5.4, we use the Radar, HDM05, and FPHA data sets and model each input sequence as multichannel correlation matrices. Comparative methods. Similar to Sec. 3.3.5.3, we compare GyroBN against LRBN. Backbone networks. We adopt the PHCM-based CorNet detailed in Sec. 5.4 as the backbone network. CorNet identifies each correlation matrix C ∈ Cor+ (n) with a collection of Poincaré vectors and applies Poincaré layers. Specifically, C is mapped to the poly-Poincaré space PPn−1 via Eq. (3.77). The resulting multi-channel Poincaré vectors are then merged into a single Poincaré representation through β-concatenation [181, Sec. 3.3]. Then, a Poincaré FC layer followed by a Poincaré MLR layer constructs the network. Implementation details. Since CorNet operates in the Poincaré geometry, both 106
Chapter 3. Riemannian Batch Normalization Radar
Methods CorNet CorNet-LRBN CorNet-GyroBN
HDM05
FPHA
Acc
Fit Time
#Params (M)
Acc
Fit Time
#Params (M)
Acc
Fit Time
#Params (M)
96.56 ± 0.86
2.12
0.0438
82.26 ± 0.92
0.74
0.4091
0.70
1.1210
2.33
0.0439
N/A
1.64
0.4094
90.03 ± 0.63
1.12
1.1213
97.67 ± 0.36
2.19
0.0439
80.82 ± 0.86
1.33
0.4094
92.88 ± 0.20
1.07
1.1213
92.85 ± 2.46
81.53 ± 0.72
Table 3.18: Comparison of CorNet with or without RBN layers. Accuracy is reported as a percentage, fit time is measured in s/epoch, and #Params is reported in millions. GyroBN and LRBN are instantiated in the Poincaré model and applied after the Poincaré FC layer. To stabilize training, we scale the learning rate of the scaling parameter s by factors of 0.1 and 0.5 for Radar and respectively. On FPHA, HDM05, s we further regularize the scaling by clamping: min √v2 +ϵ , 4 . For better efficiency, the number of Fréchet mean iterations is set to 2. Main results. We summarize the comparison of CorNet with or without normalization layers in Tab. 3.18. Overall, GyroBN demonstrates clear benefits on Radar and FPHA with negligible parameter cost and modest efficiency trade-offs. On HDM05, however, neither GyroBN nor LRBN improves the baseline, with LRBN even diverging and failing to converge.
3.4
Conclusion
This chapter developed a unified approach to RBN across broad families of manifolds. We first presented LieBN, which operates under left-, right-, and bi-invariant metrics. LieBN uses the group translations for centering and biasing and performs scaling in the tangent space at the identity element. These operations provide theoretical control over Riemannian sample and population statistics. Their instantiations on SPD, rotation, and full-rank correlation manifolds further demonstrate how a choice of Lie group structure and invariant metric yields a concrete normalization layer, including the power-deformed SPD structures and CRIM developed in this chapter. The chapter then introduced pseudo-reductive gyrogroups as a relaxation of classical gyrogroups, providing a broader algebraic foundation for Riemannian normalization. Building on this structure, we developed GyroBN and showed that gyroisometric gyrations enable theoretical control over sample statistics. Since every Lie group is a gyrogroup with identity gyrations, GyroBN recovers LieBN as a special case. GyroBN replaces group subtraction, addition, and tangent-space scaling with gyrosubtraction, gyroaddition, and scalar gyromultiplication, thereby extending the normalization principle to broader geometries. Finally, we instantiated the framework on the Grassmannian, five CCS models, and the full-rank correlation manifold. Experiments across all 107
3.4. Conclusion considered geometries support the effectiveness of this extension. Together, LieBN and GyroBN extend normalization with theoretical control of intrinsic statistics from Lie groups to pseudo-reductive gyrogroups. Gyrogroups substantially broaden the scope beyond Lie groups but do not encompass all manifolds. A natural direction for future work is therefore to extend this normalization principle to manifolds that do not admit suitable gyro-structures. The next chapter considers intrinsic classification.
108
Chapter 4 Riemannian Multinomial Logistic Regression 4.1
Introduction
Although Riemannian networks demonstrated success in many applications, most approaches still rely on Euclidean spaces for classification, such as tangent spaces [106, 107, 31, 156, 207, 209, 157, 158, 123, 208, 47], ambient Euclidean spaces [205, 183, 184], or coordinate systems [36].1 However, these strategies distort the intrinsic geometry of the manifold, undermining the effectiveness of Riemannian networks. Researchers have recently started directly developing Riemannian Multinomial Logistic Regression (RMLR) on manifolds. Inspired by the idea of hyperplane margin [130], Ganea et al. [76] developed a hyperbolic MLR in the Poincaré ball for HNNs. Motivated by HNNs, Nguyen and Yang [159] developed three kinds of gyro SPD MLRs based on three distinct gyro-structures of the SPD manifold. Nguyen et al. [161] proposed gyro MLRs for the Symmetric Positive Semidefinite (SPSD) manifold based on the product of gyro spaces. However, these classifiers often rely on manifold-specific geometric structures, limiting their generalizability to other geometries. For instance, the hyperbolic MLR [76] relies on the generalized law of sines, while the gyro MLRs [159, 161] rely on gyro-structures. This chapter proceeds in two stages. We first study SPD manifolds endowed with pullback Euclidean metrics, whose flat geometry reduces the point-to-hyperplane infimum to a Euclidean problem and yields closed-form intrinsic MLRs. We then extend the classification principle to general Riemannian manifolds. Rather than evaluating 1
Notice that there are also works designing classifiers for grid-based manifold-valued data [37], but we focus on non-gridded data in line with many previous SPD networks.
109
4.2. Multinomial Logistic Regression on SPD Manifolds the potentially intractable point-to-hyperplane infimum on each manifold, the general RMLR adopts a Riemannian-trigonometric formulation that requires only a well-defined Riemannian logarithm. Since the SPD classifiers in Sec. 4.2 are special cases of the general RMLR framework in Sec. 4.3, and the two parts use the same experimental settings, all experiments2 are presented together in Sec. 4.4. Throughout this chapter and its accompanying appendix material, all parameters are assumed to satisfy the standing admissibility conditions.3
4.2
Multinomial Logistic Regression on SPD Manifolds
4.2.1
Introduction
SPD matrices are commonly encountered in a diverse range of scientific fields, such as medical imaging [36, 37], signal processing [8, 105, 32, 31], elasticity [152, 89], question answering [141, 157], graph classification [42], and computer vision [106, 94, 232, 34, 41, 230, 37, 48, 183, 156, 158, 185]. Despite their ubiquitous presence, traditional learning algorithms are ineffective in handling the non-Euclidean geometry of SPD matrices. To address this limitation, several Riemannian metrics [171, 9, 137] have been proposed. With these Riemannian metrics, various machine learning techniques can be generalized to SPD manifolds. Inspired by the great success of deep learning [102, 127, 97], several deep networks have been developed on SPD manifolds. Despite their promising performance, many approaches still rely on Euclidean spaces for classification, such as tangent spaces [106, 31, 156, 207, 157, 158, 123, 208, 47], ambient Euclidean spaces [205, 183, 184], and coordinate systems [36]. However, these strategies distort the intrinsic geometry of the SPD manifold, undermining the effectiveness of SPD neural networks. Notably, there are also some similarity-based classifiers originally designed for shallow learning methods [77, 94, 48]. Although these classifiers can be extended to deep SPD neural networks [209, 211], the calculation of pairwise distances might undermine training efficiency. Recently, motivated by HNNs [76], three kinds of SPD Multinomial Logistic Regression (MLR) classifiers based on the gyro-structures induced by LEM, LCM, and 2
The code is available at https://github.com/GitZH-Chen/RMLR. All normal vectors are nonzero. Whenever a displayed formula contains 1/θ or 1/θ2 , we assume θ ̸= 0, except when an explicit limit as θ → 0 is considered. The inner-product parameters satisfy (α, β) ∈ ST, as defined in Eq. (2.93). These conditions are not repeated below. 3
110
Chapter 4. Riemannian Multinomial Logistic Regression AIM were developed by Nguyen and Yang [159]. However, the proposed SPD MLRs rely on the gyro-structures, limiting their generality. Besides, Chakraborty et al. [37] also introduced an invariant layer for manifold-valued data mimicking the invariant FC layer in Convolutional Neural Networks (CNNs). However, it is designed for gridded manifold-valued data, which is not the primary data type encountered in many other SPD networks. Following the convention of most SPD networks, we only focus on non-gridded cases. In fact, SPD MLR can be directly derived under LEM and LCM without the assistance of gyro-structures. More generally, LEM and LCM are pullback Euclidean metrics, which are metrics pulled back from the Euclidean space. This section develops a unified construction of SPD MLR across pullback Euclidean metrics. On the empirical side, we focus on the parameterized LEM and LCM defined in Sec. 3.2.5.1, which generalize the standard LEM and LCM by the pullback of matrix power. Their deformation behavior is established in Thm. 67. We showcase our SPD MLRs under these parameterized metrics. Besides, our framework encompasses the gyro SPD MLRs induced by the standard LEM and LCM in Nguyen and Yang [159]. More importantly, our framework also provides an intrinsic explanation for the commonly used LogEig classifier on SPD manifolds, which consists of successive matrix logarithm, FC, and softmax layers. The main contributions are summarized as follows: (1) We develop a unified construction of SPD MLR under pullback Euclidean metrics and instantiate it under two parameterized metric families. (2) Our framework offers an intrinsic explanation of the most popular LogEig classifier, which stacks matrix logarithm, the FC layer, and softmax. Outline. Sec. 4.2.2 develops SPD MLRs under pullback Euclidean metrics. The deformed-metric instantiation follows in Sec. 4.2.3. Sec. 4.2.4 reinterprets the existing LogEig classifier intrinsically. The corresponding experiments are presented together with the general RMLR experiments in Sec. 4.4. Proofs are deferred to Sec. B.4.
4.2.2
SPD Multinomial Logistic Regression
This section first reformulates the Euclidean MLR. Then, we deal with SPD MLR under an arbitrary pullback Euclidean metric on SPD manifolds. We use the pullback metric in Thm. 33 with a Euclidean codomain. The induced abelian Lie-group structure, biinvariant metric, distance, and Riemannian operators required below are established later in Thm. 149. 111
4.2. Multinomial Logistic Regression on SPD Manifolds 4.2.2.1
Reformulation of Euclidean MLR
The Euclidean MLR was first reformulated by Lebanon and Lafferty [130] from the perspective of distances to margin hyperplanes. Hyperbolic MLR was designed based on this reformulation [76]. Nguyen and Yang [159] further proposed three gyro SPD MLRs based on the gyro-structures induced by AIM, LEM, and LCM. We now briefly review the reformulation of Euclidean MLR. Given C classes, MLR in Rn computes the following softmax probabilities: ∀k ∈ {1, . . . , C},
p(y = k | x) ∝ exp (⟨ak , x⟩ − bk ) ,
(4.1)
where bk ∈ R and x, ak ∈ Rn . As shown in Lebanon and Lafferty [130, Sec. 5] and Ganea et al. [76, Sec. 3.1], Eq. (4.1) can be reformulated as p(y = k | x) ∝ exp (sign (⟨ak , x − pk ⟩) ∥ak ∥ d (x, Hak ,pk )) ,
(4.2)
where ⟨ak , pk ⟩ = bk , and Hak ,pk is referred to as a hyperplane, defined as Hak ,pk = {x ∈ Rn | ⟨ak , x − pk ⟩ = 0} .
(4.3)
As reviewed in Tab. 2.3, Logp x is the natural generalization of the directional vector − → px = x − p starting at p and ending at x, while the Riemannian metric at p corresponds to the inner product. Therefore, the MLR in Eq. (4.2) and hyperplane in Eq. (4.3) can n be readily generalized to the SPD manifold (S++ , g). n n \{0}, we Definition 102 (SPD hyperplanes). Given P ∈ S++ and A ∈ TP S++ define the SPD hyperplane as
n H̃A,P = S ∈ S++ | gP (LogP S, A) = ⟨LogP S, A⟩P = 0 ,
(4.4)
where P and A are referred to as shift and normal matrices, respectively. Definition 103 (SPD MLR). The SPD MLR is defined as p(y = k | S) ∝ exp sign Ak , LogPk (S) P ∥Ak ∥Pk d S, H̃Ak ,Pk , k
(4.5)
n n n where Pk ∈ S++ , Ak ∈ TPk S++ \{0}, ⟨·, ·⟩Pk = gPk , ∥ · ∥Pk is the norm on TPk S++ n induced by g at Pk , and H̃Ak ,Pk is a margin hyperplane in S++ as defined in Eq. (4.4).
112
Chapter 4. Riemannian Multinomial Logistic Regression d S, H̃Ak ,Pk denotes the margin distance between S and the SPD hyperplane H̃Ak ,Pk , which is formulated as
d S, H̃Ak ,Pk =
inf
d(S, Q),
Q∈H̃Ak ,Pk
(4.6)
where d(S, Q) is the geodesic distance induced by g. In geometry, the hyperplane in Eq. (4.3) is actually a regular submanifold of the trivial manifold Rn . As for our definition of SPD hyperplanes, we have a similar result. Proposition 104 (Submanifolds). [↓] The SPD hyperplane, as defined in Eq. (4.4), n n is a global → TP S++ is a regular submanifold of the SPD manifold if LogP : S++ diffeomorphism. Thm. 104 rationalizes Thm. 102, as both the SPD hyperplane and Euclidean hyperplane are submanifolds. Nevertheless, we still follow the nomenclature of Ganea et al. [76], Lebanon and Lafferty [130] and call H̃A,P an SPD hyperplane.
4.2.2.2
SPD MLRs under Pullback Euclidean Metrics
For our SPD MLR in Thm. 103, under most Riemannian metrics on SPD manifolds, all the operators involved in Eq. (4.5) have closed-form expressions except the margin distance in Eq. (4.6). Therefore, the only difficulty lies in the calculation of the margin distance. This subsection proposes a general expression for SPD MLRs under pullback Euclidean metrics, which are defined by a pullback from Euclidean metric, as detailed in Thm. 149. We choose the identity matrix I as a fixed anchor. We choose pullback Euclidean metrics as our starting metrics mainly because of their extensive inclusion and easy computation. Several Riemannian metrics, including LEM, LCM, and their variants [196, 194], are pullback Euclidean metrics. Besides, due to their fast and simple calculation, the margin distance under a pullback Euclidean metric has a closed-form expression, while obtaining distances to hyperplanes under other metrics, such as AIM, would be complicated. We start by calculating the margin distance in Eq. (4.6) under a given pullback Euclidean metric. Lemma 105. [↓] Given a pullback Euclidean metric g, the margin distance defined 113
4.2. Multinomial Logistic Regression on SPD Manifolds in Eq. (4.6) has a closed-form solution: d S, H̃Ak ,Pk = d ϕ(S), Hϕ∗,Pk (Ak ),ϕ(Pk ) =
|⟨ϕ(S) − ϕ(Pk ), ϕ∗,Pk (Ak )⟩| , ∥Ak ∥Pk
(4.7) (4.8)
where | · | is the absolute value. Putting Eq. (4.8) into Eq. (4.5), we obtain our SPD MLR under a given pullback Euclidean metric: p(y = k | S) ∝ exp
Ak , LogPk (S) P
k
= exp (⟨ϕ(S) − ϕ(Pk ), ϕ∗,Pk (Ak )⟩) ,
(4.9) (4.10)
n n n where S, Pk ∈ S++ and Ak ∈ TPk S++ \{0}. When Pk is fixed, Ak ∈ TPk S++ indeed lies in a Euclidean space. However, Pk would vary during training, making Ak nonEuclidean. To remedy this issue, Ak can be generated from a Euclidean parameter in a fixed tangent space through one of three mechanisms: Riemannian parallel transport [69], vector transport4 [1], or the differential of a group translation, including Lie group translation [197, Sec. 20] and gyrogroup translation [199]. We focus on Riemannian parallel transport and the differential of a Lie group translation, for which we establish an equivalence under pullback Euclidean metrics. Under parallel transport, we write n Ak = PTQ→Pk (Ãk ) with Ãk ∈ TQ S++ as a Euclidean parameter. This is the solution also adopted by HNNs [76], where the tangent point is the zero vector. Since the Lie groups associated with pullback Euclidean metrics are abelian, we only consider the left translation. We have the following two lemmas to show the relation between parallel transport and the differential of left translation.
Lemma 106. [↓] Given a pullback Euclidean metric, any parallel transport is equivalent to the differential map of a left translation and vice versa. n Lemma 107. [↓] Given two fixed SPD matrices Q1 , Q2 ∈ S++ , we have the follow-
4
Vector transport is often used in optimization as a substitute for parallel transport because its expressions are typically simpler and cheaper to compute [27, Sec. 10.5].
114
Chapter 4. Riemannian Multinomial Logistic Regression
ing equivalence for parallel transports under a pullback Euclidean metric: n n ∀Ã1,k ∈ TQ1 S++ , ∃! Ã2,k ∈ TQ2 S++ , s.t. PTQ1 →Pk Ã1,k = PTQ2 →Pk Ã2,k .
(4.11)
Thm. 106 indicates that under pullback Euclidean metrics, the above two solutions are equivalent, while Thm. 107 implies that anchor points can be arbitrarily chosen. Therefore, without loss of generality, we generate Ak from the tangent space at the n ∼ n identity matrix I by parallel transport, i.e., Ak = PTI→Pk (Ãk ) with Ãk ∈ TI S++ =S . Together with Eq. (6.9), Eq. (4.10) can be further simplified. Theorem 108 (SPD MLR under a pullback Euclidean metric). [↓] Under any pullback Euclidean metric, SPD MLR and SPD hyperplane are D E p(y = k | S) ∝ exp ϕ(S) − ϕ(Pk ), ϕ∗,I (Ãk ) , n D E o n | ϕ(S) − ϕ(Pk ), ϕ∗,I (Ãk ) = 0 , H̃Ãk ,Pk = S ∈ S++
(4.12) (4.13)
n n is an SPD \{0} ∼ where Ãk ∈ TI S++ = S n \{0} is a symmetric matrix, and Pk ∈ S++ matrix.
4.2.3
SPD MLRs under Deformed LEM and LCM
In this section, we first review the deformed LEM and LCM, and then showcase our SPD MLR in Thm. 108 under these deformed metrics. As reviewed in Sec. 3.2.5.1, inspired by the deforming utility of the matrix power function [192, 194], we define (θ, α, β)-LEM and θ-LCM as the pullback metrics of (α, β)-LEM and LCM by the matrix power function (·)θ and scaled by θ12 for θ ̸= 0. As shown in Thm. 67, (θ, α, β)-LEM is equal to (α, β)-LEM and θ-LCM interpolates between the standard LCM for θ = 1 and an LEM-like metric as θ → 0. Besides, both (α, β)-LEM and θ-LCM are pullback Euclidean metrics, as summarized in Tab. 3.6. Therefore, the SPD MLRs under these two families of metrics can be directly obtained by Thm. 108. Corollary 109 (SPD MLRs under the deformed LEM and LCM). [↓] The SPD 115
4.2. Multinomial Logistic Regression on SPD Manifolds
MLR under (α, β)-LEM is p(y = k | S) ∝ exp
D
log(S) − log(Pk ), Ãk
E(α,β)
,
(4.14)
n ∼ n n where Ãk ∈ TI S++ . The SPD MLR under θ-LCM is = S and Pk ∈ S++
1 ⟨X, Y ⟩ , p(y = k | S) ∝ exp θ
(4.15)
h i X = ⌊K̃⌋ − ⌊L̃k ⌋ + Dlog D(K̃) − Dlog D(L̃k ) ,
(4.16)
with X and Y defined as
1 Y = ⌊Ãk ⌋ + D(Ãk ), 2
(4.17)
where K̃ = Chol(S θ ), L̃k = Chol(Pkθ ), and D(Ãk ) denotes a diagonal matrix with the diagonal elements of Ãk . !
x y ∈ y z S 2 is positive definite if and only if x, z > 0 ∧ xz > y 2 . Fig. 4.1 illustrates SPD hyperplanes induced by (α, β)-LEM and θ-LCM. 2 S++ can be visualized as an open cone in R3 by the condition that P =
Remark 110. This construction incorporates the results with respect to LEM and LCM presented by Nguyen and Yang [159]. For (α, β)-LEM, when (α, β) = (1, 0), (α, β)-LEM becomes the standard LEM. Our margin distance to the hyperplane in Thm. 105 becomes the pseudo-gyrodistance under LEM [159, Thm. 2.23]. For θ-LCM, when θ = 1, θ-LCM becomes the standard LCM. Thm. 105 then becomes the pseudo-gyrodistance induced by LCM [159, Thm. 2.24]. However, our framework does not require gyro-structures and directly obtains the margin distance and SPD MLR based on the Riemannian metric.
4.2.4
Rethinking the Existing LogEig Classifier
Many existing SPD neural networks [106, 31, 160, 207, 156, 208, 47] rely on a Euclidean MLR in the codomain of matrix logarithm, i.e., a matrix logarithm followed by an FC layer and a softmax layer. For simplicity, we call this classifier LogEig MLR. The existing explanation of LogEig MLR is that it approximates manifolds by a tangent 116
Chapter 4. Riemannian Multinomial Logistic Regression
Figure 4.1: Conceptual illustration of SPD hyperplanes induced by (α, β)-LEM and θ-LCM. In each subfigure, the black dots are SPSD matrices, denoting the boundary 2 , while the blue, red, and yellow dots denote three SPD hyperplanes. of S++ space. However, our framework can offer a novel intrinsic explanation for this widely used MLR. When (α, β) = (1, 0) for (α, β)-LEM, the SPD MLR in Eq. (4.14) is very similar to the LogEig MLR. However, due to the nonlinearity of log(·) and the non-Euclideanness of the SPD parameter Pk , SPD MLR cannot be hastily viewed as equivalent to LogEig MLR. Nevertheless, under special circumstances, Eq. (4.14) is indeed equivalent to a LogEig MLR. Proposition 111. [↓] The LEM-based SPD MLR is equivalent to a LogEig MLR with parameters in the FC layer optimized by Euclidean SGD when SPD manifolds are endowed with the standard LEM, the SPD parameter Pk in Eq. (4.14) is optimized by LEM-based RSGD, and the Euclidean parameter Ãk is optimized by Euclidean SGD. Thm. 111 implies that, when optimized by LEM-based RSGD, the LEM-based SPD MLR is equivalent to the Euclidean MLR in the codomain of matrix logarithm. Nevertheless, a substantial body of prior work underscores the theoretical and empirical superiority of AIM-based optimization over its LEM-based counterpart [186, 92]. Therefore, we adopt the AIM-based optimizer to update the involved SPD parameters.
4.3
Extension to General Riemannian Manifolds
4.3.1
Introduction
The SPD MLR framework developed in Sec. 4.2 constructs an intrinsic classifier from the geodesic distance between an SPD matrix and a margin hyperplane. Its central 117
4.3. Extension to General Riemannian Manifolds quantity is the point-to-hyperplane distance in Eq. (4.6), defined as the infimum of the geodesic distance over all points on the hyperplane. For the pullback Euclidean metrics considered in Sec. 4.2.2.2, this optimization admits a closed-form solution. On a general Riemannian manifold, however, directly evaluating this infimum can require solving a difficult, potentially non-convex optimization problem, which prevents an extension of the preceding framework. We circumvent this obstacle by reinterpreting the Euclidean margin distance from a trigonometric perspective rather than directly solving the infimum. This alternative characterization expresses the margin distance through the angle between geodesics and the geodesic distance from the input to the hyperplane anchor. Lifting this characterization to Riemannian manifolds yields a closed-form RMLR. Accordingly, the resulting framework only requires an explicit expression of the Riemannian logarithm, which is the minimal geometric requirement for extending Euclidean MLR to manifolds. Since this requirement is satisfied by many manifolds commonly encountered in machine learning, the framework applies broadly across different geometries. This reliance on an operator shared by many manifolds, rather than on additional manifold-specific structure, makes RMLR a unified classification module across geometries. We instantiate the framework on SPD manifolds and rotation matrices. On the SPD manifold, we systematically develop SPD MLRs under five families of power-deformed metrics and provide a complete theoretical discussion of their geometric properties. On the Lie group SO(n), we construct a Lie MLR under the widely used bi-invariant metric, providing the first extension of Euclidean MLR to Lie groups. Moreover, the framework incorporates several existing Riemannian MLRs, including gyro SPD MLRs [159], the SPD MLRs developed in Sec. 4.2, and gyro SPSD MLRs [161]. Our SPD MLRs are validated on four SPD backbone networks, including SPDNet [106] on the radar and human action recognition tasks and TSMNet [123] on the EEG classification tasks for the Riemannian feedforward network, RResNet [117] on the human action recognition task for the Riemannian residual network, and SPDGCN [231] on the node-classification task for the Riemannian graph neural network. Our Lie MLR is validated on the classic LieNet [107] backbone for the human action recognition task. Compared with previous non-intrinsic classifiers, our MLRs achieve consistent performance gains. In particular, our SPD MLRs outperform the previous classifiers by 14.23 percentage points on SPDNet and 13.72 percentage points on RResNet for human action recognition, and 4.46 percentage points on TSMNet for EEG inter-subject classification. Furthermore, our Lie MLR can improve both the training stability and performance. In summary, our main theoretical contributions are the 118
Chapter 4. Riemannian Multinomial Logistic Regression following: (1) We develop a unified RMLR framework for general Riemannian manifolds that requires only the Riemannian logarithm and incorporates several existing manifoldspecific RMLRs as special cases. (2) We systematically propose five families of SPD MLRs based on different geometries of the SPD manifold. (3) We propose a novel Lie MLR for deep neural networks on SO(n). Outline. Sec. 4.3.2 revisits the existing Riemannian MLRs and proposes the general RMLR framework. Sec. 4.3.3 instantiates the framework on SPD manifolds under five families of power-deformed metrics. Sec. 4.3.4 presents the Lie MLR on SO(n). Sec. 4.4 reports experiments on Riemannian feedforward, residual, and graph networks, direct SPD classification, and LieNet. Proofs are deferred to Sec. B.5.
4.3.2
Riemannian Multinomial Logistic Regression
Inspired by Lebanon and Lafferty [130], prior work extended the Euclidean MLR to hyperbolic, SPD, and SPSD manifolds [76, 159, 161]. Sec. 4.2 further develops SPD MLRs under pullback Euclidean metrics. However, these classifiers rely on specific Riemannian properties, such as the generalized law of sines, gyro-structures, and flat metrics, which limit their generality. In this section, we first revisit several existing MLRs and then propose our Riemannian classifiers with minimal geometric requirements.
4.3.2.1
Revisiting Existing Multinomial Logistic Regressions
The Euclidean MLR and its reformulation through margin distances to hyperplanes have already been reviewed in Sec. 4.2.2.1. By abstracting the expressions for the SPD MLR and SPD hyperplanes in Eqs. (4.4) and (4.5), we can readily extend this construction from the SPD manifold to a general Riemannian manifold M:
˜ p(y = k | S) ∝ exp sign(⟨Ãk , LogPk (S)⟩Pk )∥Ãk ∥Pk d(S, H̃Ãk ,Pk ) , n o H̃Ãk ,Pk = S ∈ M | gPk (LogPk S, Ãk ) = 0 , 119
(4.18) (4.19)
4.3. Extension to General Riemannian Manifolds where Pk ∈ M, Ãk ∈ TPk M\{0}, gPk is the Riemannian metric at Pk , and LogPk is the Riemannian logarithm at Pk . The margin distance is defined as an infimum: ˜ H̃ d(S, Ãk ,Pk ) =
inf Q∈H̃Ã ,P k
d(S, Q).
(4.20)
k
The MLRs in Lebanon and Lafferty [130], Ganea et al. [76], Nguyen and Yang [159] and Sec. 4.2 can be viewed as different implementations of Eqs. (4.18) to (4.20). To calculate the MLR in Eq. (4.18), one has to compute the associated Riemannian metrics, logarithmic maps, and margin distance. The associated Riemannian metrics and logarithmic maps often have closed-form expressions on the frequently encountered manifolds in machine learning. However, the computation of the margin distance can be challenging. On the Poincaré ball of hyperbolic manifolds, the generalized law of sines simplifies the calculation of Eq. (4.20) [76]. However, the generalized law of sines is not universally guaranteed on other manifolds. For pullback Euclidean metrics on SPD manifolds, Thm. 105 provides a closed-form solution for the margin distance. For curved manifolds, solving Eq. (4.20) would become a non-convex optimization problem. To address this challenge, Nguyen and Yang [159] defined gyro-structures on the SPD manifold and proposed a pseudo-gyrodistance to calculate the margin distance. Similarly, Nguyen et al. [161] proposed a pseudo-gyrodistance on the SPSD manifold based on the gyro product space. However, gyro-structures do not necessarily exist in general geometries. In summary, the aforementioned methods often rely on specific properties of their associated Riemannian metrics, which usually do not generalize to general geometries. 4.3.2.2
Riemannian Multinomial Logistic Regression
Recalling Eqs. (4.18) and (4.19), the minimum requirement for extending Euclidean MLR to manifolds is the well-definedness of LogPk (S) for each k. In this subsection, we will develop Riemannian MLR, which depends solely on the Riemannian logarithm, without additional requirements, such as gyro-structures and the generalized law of sines. In the following, we always assume the well-definedness of the Riemannian logarithm. We start by reformulating the Euclidean margin distance to the hyperplane from a trigonometric perspective and then present our Riemannian MLR. As we discussed before, obtaining the margin distance of Eq. (4.20) could be challenging. Inspired by Nguyen and Yang [159], we resort to the perspective of trigonometry to reinterpret Euclidean margin distance. In Euclidean space, the margin distance 120
Chapter 4. Riemannian Multinomial Logistic Regression is equivalent to d (x, Ha,p ) = sin (∠xpy ∗ ) d(x, p),
with y ∗ = argmax cos ∠xpy.
(4.21)
y∈Ha,p \{p}
We extend Eq. (4.21) to manifolds by the Riemannian trigonometry and geodesic distance, the counterparts of Euclidean trigonometry and distance. Definition 112 (Riemannian margin distance). Let H̃Ã,P be a Riemannian hyperplane defined in Eq. (4.19), and S ∈ M. The Riemannian margin distance from S to H̃Ã,P is defined as d S, H̃Ã,P = sin (∠SP Y ∗ ) d(S, P ),
(4.22)
Y ∗ = argmax cos ∠SP Y.
(4.23)
where d(S, P ) is the geodesic distance, and
Y ∈H̃Ã,P \{P }
The initial velocities of geodesics define cos ∠SP Y : cos ∠SP Y =
⟨LogP Y, LogP S⟩P , ∥ LogP Y ∥P ∥ LogP S∥P
(4.24)
where ⟨·, ·⟩P is the Riemannian metric at P , and ∥ · ∥P is the associated norm. The Riemannian margin distance in Thm. 112 has a closed-form expression. Theorem 113. [↓] The Riemannian margin distance defined in Thm. 112 is given by |⟨LogP S, Ã⟩P | d(S, H̃Ã,P ) = . (4.25) ∥Ã∥P Putting Eq. (4.25) into Eq. (4.18), we can obtain a closed-form expression for Riemannian MLR. Theorem 114 (RMLR). [↓] Given a Riemannian manifold (M, g), the Riemannian MLR induced by g is p(y = k | S ∈ M) ∝ exp ⟨LogPk S, Ãk ⟩Pk ,
121
(4.26)
4.3. Extension to General Riemannian Manifolds
where Pk ∈ M, Ãk ∈ TPk M\{0}, and Log is the Riemannian logarithm. As discussed for SPD MLRs in Sec. 4.2.2.2, a tangent vector attached to a trainable base point can be generated from a Euclidean parameter in a fixed tangent space through Riemannian parallel transport, vector transport, or the differential of a group translation. Following the hyperbolic and gyro MLRs [76, 159], we focus on parallel transport and Lie group translation: Ãk = ΓQ→Pk Ak , Ãk = LPk ⊙Q−1 ⊙
(4.27) ∗,Q
Ak ,
(4.28)
where Q ∈ M is a fixed point, Ak ∈ TQ M\{0}, Γ is the parallel transport along the geodesic connecting Q and Pk , and LPk ⊙Q−1 denotes the differential map at Q of ⊙ ∗,Q
left translation LPk ⊙Q−1 with Pk ⊙ Q−1 ⊙ denoting the Lie group product and inverse. ⊙ In this way, Ak lies in a fixed tangent space and, therefore, can be optimized by a Euclidean optimizer. Remark 115. We make the following remarks regarding our Riemannian MLR. (a) The reformulations of Eq. (4.21) in gyro MLR [159, 161] and in our work are different. Gyro MLR adopts gyro trigonometry and gyro distance to reformulate Eq. (4.21), while our method directly uses Riemannian trigonometry and geodesic distance. (b) Compared with the hyperbolic, gyro SPD, and gyro SPSD MLRs [76, 159, 161] and the SPD MLRs in Sec. 4.2, our framework enjoys broader applicability, as our framework only requires the Riemannian logarithm. This property is commonly satisfied by most manifolds encountered in machine learning, such as the five metrics on SPD manifolds mentioned in Sec. 2.9.1, the invariant metric on SO(n) [28], and the hyperbolic and spherical models in Sec. 2.9.5. Besides, several existing MLRs on different geometries are special cases of our Riemannian MLR, which are detailed in Tab. 4.1. (c) The well-definedness of the Riemannian logarithm is a much weaker requirement compared to the existence of the gyro-structure. The gyro-structure not only requires the Riemannian logarithm but also implicitly requires geodesic completeness [159, Eqs. (1)–(2)]. For instance, on SPD manifolds, EM and BWM [196] are incomplete, undermining the well-definedness of gyro operations.
122
Chapter 4. Riemannian Multinomial Logistic Regression MLR
Geometries
Requirements
Incorporated by Our MLR
Euclidean MLR (Eq. (4.1))
Euclidean geometry
N/A
✓(Sec. A.4.1.1)
Gyro SPD MLRs [159]
n AIM, LEM & LCM on S++
Gyro-structures
✓(Thm. 118)
SPSD product gyro spaces
Gyro-structures
✓(Sec. A.4.1.2)
Flat SPD MLRs (Sec. 4.2)
n (α, β)-LEM & (θ)-LCM on S++
Pullback metrics from the Euclidean space
✓(Thm. 118)
Ours
General geometries
Riemannian logarithm
N/A
Gyro SPSD MLRs [161]
Table 4.1: Several MLRs on different geometries are special cases of our MLR.
(", $, %)-EM (", $, %)-LEM
("-BWM
(", $, %)-AIM O ' -Invariant Metrics Riemannian Metrics on SPD Manifolds
"-LCM
Figure 4.2: Illustration of the deformation (left) and Venn diagram (right) of metrics on SPD manifolds, where IEM, SREM, and 41 PAM denote Inverse Euclidean Metric, Square Root Euclidean Metric, and Polar Affine Metric scaled by 1/4, respectively.
4.3.3
SPD Multinomial Logistic Regressions
This section showcases our RMLR framework on the SPD manifold. We first systematically discuss the power-deformed geometries of SPD manifolds. Based on these metrics, we will develop five families of deformed SPD MLRs. 4.3.3.1
Deformed Geometries of SPD Manifolds
As discussed in Sec. 2.9.1, there are five popular Riemannian metrics on SPD manifolds. n These metrics can all be extended to power-deformed metrics. For a metric g on S++ , the power-deformed metric is defined as g̃P (V, W ) =
1 n n gP θ ((ϕθ )∗,P (V ), (ϕθ )∗,P (W )) , ∀P ∈ S++ , V, W ∈ TP S++ , 2 θ
(4.29)
where ϕθ (P ) = P θ is the matrix power, and (ϕθ )∗,P is the differential map. The deformed metric g̃ can interpolate between a LEM-like metric (θ → 0) and g (θ = 1) [194]. Previous work power-deformed (α, β)-AIM and BWM into (θ, α, β)-AIM [192] and 2θ-BWM [194], respectively. Earlier, we introduced (θ, α, β)-LEM and θ-LCM in Sec. 3.2.5.1. They are the power deformations of (α, β)-LEM and LCM, respectively. By Thm. 67, (θ, α, β)-LEM is equal to (α, β)-LEM. We therefore use (α, β)-LEM for this family. Here, we further define the power deformation of (α, β)-EM through Eq. (4.29), 123
4.3. Extension to General Riemannian Manifolds Name
Properties
(θ, α, β)-LEM
Bi-invariance, O(n)-invariance, Geodesic Completeness
(θ, α, β)-AIM
Lie Group Left-Invariance, O(n)-invariance, Geodesic Completeness
(θ, α, β)-EM
O(n)-Invariance
θ-LCM
Lie Group Bi-Invariance, Geodesic Completeness
2θ-BWM
O(n)-Invariance
Table 4.2: Properties of deformed metrics on SPD manifolds (θ ̸= 0 and min(α, α + nβ) > 0).
Figure 4.3: Conceptual illustration of SPD hyperplanes induced by five families of 2 . Riemannian metrics. The black dots denote the boundary of S++ denoted by (θ, α, β)-EM. We have the following result for (θ, α, β)-EM. Proposition 116. [↓] (θ, α, β)-EM interpolates between (α, β)-LEM (θ → 0) and (α, β)-EM (θ = 1). So far, all five popular Riemannian metrics on SPD manifolds have been generalized to power-deformed families of metrics. We summarize their associated properties in Tab. 4.2 and present their theoretical relation in Fig. 4.2. We leave technical details in Sec. A.4.1.3. 4.3.3.2
Five Families of SPD Multinomial Logistic Regressions
This subsection presents five families of specific SPD MLRs using our general framework in Thm. 114 and the metrics discussed in Sec. 4.3.3.1. We focus on generating Ãk by parallel transport from the identity matrix, except for 2θ-BWM. Since the parallel transport under 2θ-BWM would undermine numerical stability (please refer to Sec. A.4.1.4 for more details), we resort to a newly developed Lie group operation [195]: S1 ⊙ S2 = L1 S2 L⊤ 1, where L1 = Chol(S1 ) is the Cholesky factor. 124
n ∀S1 , S2 ∈ S++ ,
(4.30)
Chapter 4. Riemannian Multinomial Logistic Regression Theorem 117 (SPD MLRs). [↓] By abuse of notation, we omit the subscripts k n of Ak and Pk . Given an SPD feature S, the SPD MLRs, p(y = k | S ∈ S++ ), are proportional to (α, β)-LEM : exp ⟨log(S) − log(P ), A⟩(α,β) , D E(α,β) 1 − θ2 θ − θ2 (θ, α, β)-AIM : exp log P S P ,A , θ 1 θ θ (α,β) , (θ, α, β)-EM : exp ⟨S − P , A⟩ θ h i * + ⌊ K̃⌋ − ⌊ L̃⌋ + Dlog(D( K̃)) − Dlog(D( L̃)) , 1 θ-LCM : exp , 1 θ ⌊A⌋ + D(A) 2 " * 2θ 2θ 1 +# 2θ 2θ 1 2θ 1 (P S ) 2 + (S P ) 2 − 2P , , 2θ-BWM : exp 4θ LP 2θ L̄AL̄⊤
(4.31) (4.32) (4.33) (4.34)
(4.35)
n \{0} is a symmetric matrix, log(·) is the matrix logarithm, LP [V ] where A ∈ TI S++ is the solution to the matrix linear system LP [V ]P + P LP [V ] = V , known as the Lyapunov operator, Dlog(·) is the diagonal element-wise logarithm, ⌊·⌋ is the strictly lower part of a square matrix, and D(·) is a diagonal matrix with diagonal elements of a square matrix. Besides, log∗,P is the differential map at P , K̃ = Chol(S θ ), L̃ = Chol(P θ ), and L̄ = Chol(P 2θ ).
The Lyapunov operator in Eq. (4.35) requires the eigendecomposition. However, the backpropagation of eigendecomposition involves 1/(σi −σj ) [111], undermining numerical stability. Therefore, we propose a numerically stable backpropagation for the Lyapunov operator, detailed in Sec. A.4.1.4. As 2 × 2 SPD matrices can be embedded into R3 as an open cone [220], we illustrate SPD hyperplanes induced by five families of metrics in Fig. 4.3. Remark 118. Our SPD MLRs extend the gyro SPD MLRs of Nguyen and Yang [159] and the flat SPD MLRs in Sec. 4.2. The pseudo-gyrodistance to an SPD hyperplane in Nguyen and Yang [159, Thms. 2.23–2.25] is incorporated by our Thm. 113, while the flat SPD MLRs under (α, β)-LEM and θ-LCM in Thm. 109 are special cases of our Thm. 117. Furthermore, our approach extends the scope of prior work because neither the framework in Sec. 4.2 nor the gyro SPD MLRs of Nguyen and Yang [159] cover SPD MLRs based on (θ, α, β)-EM and 2θ-BWM. The gyro operations in 125
4.3. Extension to General Riemannian Manifolds Nguyen and Yang [159, Eq. (1)] implicitly require geodesic completeness, whereas (θ, α, β)-EM and 2θ-BWM are incomplete. As neither (θ, α, β)-EM nor 2θ-BWM belong to pullback Euclidean metrics, the framework in Thm. 108 cannot be applied to these metrics. To the best of our knowledge, our work is the first to apply PEM and BWM to establish Riemannian neural networks, opening up new possibilities for utilizing these metrics in machine learning applications. Besides, neither the gyro SPD MLRs of Nguyen and Yang [159] nor the flat SPD MLRs in Sec. 4.2 cover the deformed metrics for building SPD MLRs.
4.3.4
Lie Multinomial Logistic Regression
This section introduces our Lie MLR on SO(n) based on the general RMLR framework in Thm. 114. The Riemannian metric on SO(n) is assumed to be the invariant metric in Tab. 2.12. We employ the vector transport on SO(n) given by Boumal and Absil [28, Tab. 1], which coincides with the differential of left translation in Eq. (4.28). Lemma 119. [↓] TQ→P (H) = (LP Q−1 )∗,Q (H) = P Q⊤ H,
∀P, Q ∈ SO(n),
H ∈ TQ SO(n). (4.36)
Similar to SPD MLRs, we set Q = I. The Lie MLR on SO(n) is presented in the following. Theorem 120. [↓] The Lie MLR on SO(n) is given by p(y = k | R ∈ SO(n)) ∝ exp where Pk ∈ SO(n) and Ak ∈ so(n).
log Pk⊤ R , Ak ,
(4.37)
We refer to the Riemannian hyperplanes (Eq. (4.19)) on SO(n) as Lie hyperplanes. As SO(3) is homeomorphic to 3-dimensional real projective space RP3 [95], Fig. 4.4 illustrates Lie hyperplanes in the closed ball in R3 of radius π. 126
Chapter 4. Riemannian Multinomial Logistic Regression
Figure 4.4: Conceptual illustration of a Lie hyperplane. Each pair of antipodal black dots corresponds to a rotation matrix with an Euler angle of π, while the green dots denote a Lie hyperplane.
4.4
Experiments (θ, α, β)-AIM
Architectures
LogEig MLR
2-Block 5-Block
92.88±1.05 93.47±0.45
(θ, α, β)-EM
(α, β)-LEM
2θ-BWM
θ-LCM
(1,1,0)
(1,1,0)
(1,1,1/8)
(1,1,0)
(1,1,1)
(0.5)
(0.25)
(1)
(0.5)
94.53±0.95 94.32±0.94
94.24±0.55 95.11±0.82
94.93±0.60 95.01±0.84
93.55±1.21 94.60±0.70
95.64±0.83 95.87±0.58
92.22±0.83 93.69±0.66
94.99±0.47 94.84±0.68
93.49±1.25 93.93±0.98
94.59±0.82 95.16±0.67
Table 4.3: Comparison of SPDNet with LogEig against SPD MLRs on the Radar data set. The best results are bold.
(θ, α, β)-AIM Architectures
LogEig MLR
1-Block 2-Block 3-Block
57.42±1.31 60.69±0.66 60.76±0.80
(α, β)-LEM
2θ-BWM
(1,1,0)
(1,1,0)
(θ, α, β)-EM (0.5,1.0,1/30)
(1,1,0)
(0.5)
(1)
θ-LCM (0.5)
58.07±0.64 60.72±0.62 61.14±0.94
66.32±0.63 66.40±0.87 66.70±1.26
71.65±0.88 70.56±0.39 70.22±0.81
56.97±0.61 60.69±1.02 60.28±0.91
70.24±0.92 70.46±0.71 70.20±0.91
63.84±1.31 62.61±1.46 62.33±2.15
65.66±0.73 65.79±0.63 65.71±0.75
Table 4.4: Comparison of SPDNet with LogEig against SPD MLRs on the HDM05 data set.
(θ, α, β)-AIM Classifiers
LogEig MLR
Balanced Acc.
53.83±9.77
(θ, α, β)-EM
(α, β)-LEM
2θ-BWM
(1,1,0)
(0.5,1,0.05)
(1,1,0)
(1,1,0)
(0.5)
(1)
θ-LCM (1.5)
53.36±9.92
55.27±8.68
54.48±9.21
53.51±10.02
55.54±7.45
55.71±8.57
56.43±8.79
Table 4.5: Inter-session experiments of TSMNet with different MLRs on the Hinss2021 data set. 127
4.4. Experiments (θ, α, β)-AIM Classifiers
LogEig MLR
Balanced Acc.
49.68±7.88
(θ, α, β)-EM
(α, β)-LEM
2θ-BWM
θ-LCM
(1,1,0)
(1.5,1,0)
(1,1,0)
(1.5,1,1/20)
(1,1,0)
(0.5)
(0.75)
(1)
(0.5)
50.65±8.13
51.15±7.83
50.02±5.81
51.38±5.77
51.41±7.98
50.26±7.23
51.67±8.73
52.93±7.76
54.14±8.36
Table 4.6: Inter-subject experiments of TSMNet with different MLRs on the Hinss2021 data set. We first instantiate our SPD MLRs in four SPD neural networks: SPDNet [106] and TSMNet [123] for Riemannian feedforward networks, RResNet [117] for Riemannian residual networks, and SPDGCN [231] for Riemannian graph neural networks. Then, we proceed with experiments of our Lie MLR under the classic LieNet architecture [107]. The classifier in all the above networks is the LogEig MLR (matrix logarithm + FC + softmax), a Euclidean MLR on the tangent space at the identity matrix. We substitute the original non-intrinsic LogEig MLR in each baseline model with our RMLRs. Notably, the gyro SPD MLRs [159] are special cases of our SPD MLRs under the standard AIM, LEM, and LCM ((θ, α, β) = (1, 1, 0)), while the flat SPD MLRs in Sec. 4.2 are incorporated by our SPD MLRs under (α, β)-LEM and θ-LCM. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.3.
4.4.1
Experiments on the Proposed SPD MLRs
In the following, we abbreviate SPD MLR-metric as metric. For instance, (θ, α, β)-AIM denotes the baseline endowed with the SPD MLR induced by (θ, α, β)-AIM, with (1, 1, 0) as the value of (θ, α, β). Experiments on the Riemannian feedforward network. We evaluate our SPD MLRs for Riemannian feedforward networks under the SPDNet and TSMNet backbones. Following Huang and Van Gool [106], Brooks et al. [31], on SPDNet, we use the Radar data set [31] for radar recognition and the HDM05 data set [153] for human action recognition. TSMNet [123] is one of the state-of-the-art methods for the EEG classification task. Following Kobler et al. [123], we use the Hinss2021 [101] data set. For each family of SPD MLRs, we report the SPD MLR induced by the standard metric (θ = 1, α = 1, β = 0) and the one induced by the deformed metric with the best (θ, α, β). Besides, if the standard SPD MLR is already saturated, we only report the results of the standard one. Under each metric, we highlight the results of our SPD MLR under the best hyperparameters in bold. (1) Radar. In line with Brooks et al. [31], we evaluate our classifiers under two network architectures: 2-Block and 5-Block configurations. The 10-fold results (mean±std) are presented in Tab. 4.3. Note that the SPD MLR induced by standard AIM is sat128
Chapter 4. Riemannian Multinomial Logistic Regression Data Sets
LogEig MLR
(θ, α, β)-AIM
(θ, α, β)-EM
(α, β)-LEM
2θ-BWM
θ-LCM
HDM05 NTU60
58.17 ± 2.07 45.22 ± 1.23
60.23 ± 1.26 48.94 ± 0.68
71.89 ± 0.60 (↑ 13.72) 52.24 ± 1.25
59.44 ± 0.87 46.99 ± 0.41
69.85 ± 0.23 50.56 ± 0.59
65.76 ± 0.96 53.63 ± 0.95 (↑ 8.41)
Table 4.7: Comparison of LogEig against SPD MLRs under the RResNet architecture. urated. Generally speaking, our SPD MLRs achieve superior performance against the vanilla LogEig MLR. Moreover, for most families of metrics, the associated SPD MLRs with proper (θ, α, β) outperform the standard SPD MLR, demonstrating the effectiveness of our parameterization. Besides, among all SPD MLRs, the ones induced by (α, β)-LEM achieve the best performance. (2) HDM05. Following Huang and Van Gool [106], three architectures are adopted: 1-Block, 2-Block and 3-Block configurations. The 10-fold results (mean±std) are presented in Tab. 4.4. Note that the standard SPD MLRs under AIM, LEM, and BWM are already saturated on this data set. As on the Radar data set, similar observations can be made on this data set. Our SPD MLRs can bring consistent performance gains for SPDNet, and properly selected hyperparameters can bring further improvement. Particularly, among all the SPD MLRs, the ones based on the 2θ-BWM and (θ, α, β)-EM achieve the best performance. Compared to the vanilla LogEig MLR, the highest performance improvement is 14.23 percentage points, highlighting our approach’s effectiveness. Notably, since 2θ-BWM and (θ, α, β)-EM are geodesically incomplete and not pulled back from a Euclidean space, the SPD MLR under these two metrics cannot be derived by the framework of gyro or flat MLR. This contrast confirms the applicability of our theoretical framework to a broader range of geometries. (3) Hinss2021. The results (mean±std) of leave-5%-out cross-validation are reported in Tabs. 4.5 and 4.6. Once again, our intrinsic classifiers demonstrate improved performance compared to the LogEig MLR in both inter-session and inter-subject scenarios. Besides, the SPD MLRs based on θ-LCM achieve the best performance, outperforming the vanilla classifier by 2.60 percentage points for intersession and by 4.46 percentage points for inter-subject. This finding highlights the versatility of our framework. Experiments on the Riemannian residual network. Following Katsman et al. [117], we use the HDM05 and NTU60 [178] data sets on the RResNet backbone. For the hyperparameter (θ, α, β) in our SPD MLRs, we borrow the best ones from Tab. 4.4. Tab. 4.7 reports the 10-fold and 5-fold results on the HDM05 and NTU60 data sets, 129
4.4. Experiments Disease
Classifiers LogEig MLR (θ, α, β)-AIM (θ, α, β)-EM (α, β)-LEM 2θ-BWM θ-LCM
Cora
Pubmed
Mean±STD
Max
Mean±STD
Max
Mean±STD
Max
90.55 ± 4.83
96.85
78.04 ± 1.27
79.6
70.99 ± 5.12
77.6
94.84 ± 2.27 90.87 ± 5.14 96.33 ± 2.19 91.93 ± 3.64 93.01 ± 2.14
98.43 98.03 98.82 96.85 98.43
79.79 ± 1.44 79.05 ± 1.23 79.89 ± 0.99 73.46 ± 2.18 77.59 ± 1.20
81.6 81 81.8 77.7 80.1
77.83 ± 1.08 78.16 ± 2.41 78.16 ± 2.41 73.22 ± 4.06 74.46 ± 5.81
80 79.5 79.5 78.1 78.9
Table 4.8: Comparison of LogEig against SPD MLRs under the SPDGCN architecture. Classifiers
Radar
HDM05
LogEig MLR
91.93 ± 1.30
48.43 ± 1.25
(θ, α, β)-AIM (θ, α, β)-EM (α, β)-LEM 2θ-BWM θ-LCM
95.21 ± 0.81 92.25 ± 1.20 95.09 ± 0.57 94.89 ± 0.41 95.67 ± 0.61 (↑ 3.74)
49.17 ± 1.08 61.60 ± 0.69 49.05 ± 0.91 66.77 ± 1.34 (↑ 18.34) 58.66 ± 0.51
Inter-session 39.76 ± 7.60
Hinss2021
41.14 ± 7.26 45.78 ± 8.51 (↑ 6.02) 40.88 ± 7.46 44.84 ± 8.00 43.17 ± 6.21
Inter-subject 44.66 ± 7.17
45.89 ± 6.52 45.84 ± 4.75 46.02 ± 5.96 (↑ 1.36) 45.21 ± 7.44 45.10 ± 6.20
Table 4.9: Comparison of LogEig against SPD MLRs for direct classification. respectively. The SPD MLRs still consistently outperform the vanilla LogEig MLR. Besides, similar to the SPD MLRs under the SPDNet backbone for action recognition (Tab. 4.4), the SPD MLR based on θ-LCM, 2θ-BWM, or (θ, α, β)-EM outperforms the vanilla LogEig MLR by a large margin. In particular, the highest performance improvements are 13.72 and 8.41 percentage points on these two data sets. Experiments on the Riemannian graph network. We use SPDGCN [231] as the backbone network for the Riemannian graph network. Following Zhao et al. [231], we use the Disease [4], Cora [177], and Pubmed [155] data sets for node classification. The 10-fold average and maximum results of the vanilla LogEig MLR against our SPD MLR with the best (θ, α, β) are reported in Tab. 4.8. Similar to the previous results, our SPD MLRs generally outperform the LogEig MLR. Besides, the SPD MLR based on (α, β)-LEM generally achieves the best performance for SPDGCN. Ablations of SPD MLRs on direct classification. For a more straightforward comparison, we compare LogEig against our SPD MLRs for direct classification. We adopt the Radar, HDM05, and Hinss2021 data sets. We follow the preprocessing of SPDNet and TSMNet to model features into the SPD manifold and directly use LogEig or our SPD MLRs for classification. The average results are presented in Tab. 4.9. The hyperparameters (θ, α, β) are borrowed from Tabs. 4.3 to 4.6. Our SPD MLRs consistently outperform the vanilla LogEig MLR. In particular, on the HDM05 data set, the highest performance improvement by our SPD MLRs is 18.34 percentage points, surpassing the non-intrinsic LogEig MLR by a large margin. 130
Chapter 4. Riemannian Multinomial Logistic Regression
Classifiers
G3D
HDM05
Mean±STD
Max
Mean±STD
Max
LogEig MLR
87.91±0.90
89.73
76.92±1.27
79.11
Lie MLR
89.13±1.7
92.12
78.24±1.03
80.25
Table 4.10: Results of LogEig MLR against Lie MLR under the LieNet architecture.
4.4.2
Experiments on the Proposed Lie MLR
We apply our Lie MLR to the classic SO(n) network, i.e., LieNet [107], where features are on the Lie group of SO(3) × · · · × SO(3). More precisely, this feature space is a product manifold of SO(3) factors, and Thm. 120 extends naturally under the product metric, with each class logit obtained by summing the factorwise inner products. Following LieNet [107], we use G3D [23] and HDM05 [153] data sets. We also extend the Riemannian optimization package Geoopt [125] to SO(3), allowing for the direct Riemannian optimization reviewed in Sec. 2.7. We find that RSGD performs best for LieNet. Tab. 4.10 presents the 10-fold average results of LieNet with or without Lie MLR. Note that on the HDM05 data set, LieNet might fail to converge, with the validation accuracy fluctuating between 70% and 75%. Therefore, we select the 10 best-performing folds out of 20 experimental folds. It can be observed that our Lie MLR can improve the performance of LieNet. Besides, our Lie MLR can also improve the training stability. On the HDM05 data set, LieNet fails to converge in 8 out of 20 folds. However, when endowed with our Lie MLR, LieNet+LieMLR only encounters convergence failures in 2 folds.
131
4.5. Conclusion
4.5
Conclusion
This chapter developed a unified approach to intrinsic classification in two stages, progressing from a structured family of flat SPD geometries to general Riemannian manifolds. The first part considered SPD manifolds endowed with pullback Euclidean metrics. Their flat geometry reduces the infimum defining the geodesic distance from an SPD point to a margin hyperplane to a Euclidean point-to-hyperplane problem. This yields a closed-form margin distance and, consequently, a unified construction of SPD MLR. We instantiated this construction under deformed LEM and LCM and showed that, under the corresponding optimization scheme, its LEM instance recovers the widely used LogEig classifier, thereby providing an intrinsic interpretation of the existing pipeline. The second part addressed the central obstacle to extending this construction beyond flat geometries. On a general Riemannian manifold, evaluating the point-to-hyperplane distance through its infimum can require solving a difficult, potentially non-convex optimization problem and may not admit a closed-form solution. Instead of solving this minimization problem, we replaced the infimum-based margin formulation with a Riemannian-trigonometric one that combines the geodesic distance from an input to the hyperplane anchor with the angle between the corresponding geodesics. This reformulation yields a closed-form RMLR that requires only a well-defined Riemannian logarithm, extending the classification principle from flat SPD geometries to a broad range of Riemannian manifolds. We instantiated the general framework as five families of SPD MLRs under powerdeformed metrics and as a Lie MLR on SO(n). Experiments across Riemannian feedforward, residual, graph, and Lie-group networks demonstrated the broad applicability of the framework.
132
Chapter 5 Riemannian Neural Networks 5.1
Introduction
The preceding chapters developed two fundamental network modules through unified geometric formulations that can be instantiated across different manifolds. Such unified constructions make essential modules reusable across manifold families, but not every neural component admits a sufficiently tractable or effective formulation based only on broadly shared Riemannian properties. The general RMLR in Sec. 4.3.2.2, for example, achieves broad applicability by replacing the potentially intractable point-tohyperplane infimum with a Riemannian-trigonometric formulation that requires only a well-defined Riemannian logarithm. In contrast, hyperbolic and flat correlation geometries permit exact evaluation of the corresponding point-to-hyperplane infima, while Busemann functions and horospheres provide an alternative hyperbolic decision principle. These examples illustrate how additional geometric or algebraic structure of a particular manifold can support more direct and better-tailored modules and architectures. This chapter therefore studies manifold-specific Riemannian network design through three complementary routes. In Sec. 5.2, we introduce the unconstrained Proper Velocity (PV) model, establish its Riemannian toolkit, and construct MLR, fully connected, convolutional, activation, and normalization layers. In Sec. 5.3, we exploit Busemann functions and horospheres to derive intrinsic and batch-efficient BMLR and BFC layers for the Poincaré and Lorentz models. Finally, Sec. 5.4 exploits the specific geometries of full-rank correlation manifolds to construct MLR, fully connected, and convolutional layers together with Riemannian backpropagation. 133
5.2. Proper Velocity Neural Networks
5.2
Proper Velocity Neural Networks
5.2.1
Introduction
Hyperbolic representations have recently delivered strong performance across different applications because the exponential volume growth of negatively curved manifolds enables low-distortion embeddings of tree-like and hierarchical structure [163]. These advantages have been validated in computer vision [78, 120, 72, 201, 79, 15, 99, 16, 189, 140, 217], graph learning [38, 13, 75, 188], multimodal learning [67, 166], recommendation systems [223], astronomy [44], genome sequence learning [119], natural language processing [163, 76, 164, 90, 98, 222], and brain signal decoding [135]. Recently, the focus has shifted from hyperbolic embeddings to building HNNs that operate entirely within hyperbolic space. As reviewed in Sec. 2.9.5, hyperbolic geometry admits multiple models, so the choice of representation is central to the design of hyperbolic networks. Most recent works rely on the Poincaré ball and hyperboloid models, which provide convenient Riemannian or gyrovector structures, thereby facilitating neural network construction. However, both models are constrained spaces, which can lead to numerical instabilities. In particular, as embeddings in the Poincaré ball approach the boundary, numerical computations become unstable and might cause gradients to vanish [91]. On the other hand, the Proper Velocity (PV) model originates from Einstein’s special relativity, where proper velocity provides a natural parameterization for relativistic velocity addition [200, Ch. 10]. Algebraically, PV admits a gyrovector space [200, Ch. 6], analogous to the Möbius gyrovector space of the Poincaré ball. Unlike the constrained Poincaré ball and hyperboloid models, PV offers an unconstrained representation that alleviates numerical instabilities. These properties have made the PV model successful in relativistic physics and motivate its exploration as a stable alternative geometry for HNNs. However, its Riemannian operators, including exponential and logarithmic maps and parallel transport, remain largely unexplored, despite being fundamental for constructing neural networks. Inspired by the above discussions, we propose Proper Velocity Neural Networks (PVNNs). To this end, we first establish the complete Riemannian geometry of PV by deriving closed-form expressions for the exponential map, logarithmic map, geodesic distance, and parallel transport. Building on this foundation, we extend several fundamental neural layers into PV space, including MLR classification, FC, convolutional, activation, and BN layers. Together, these layers form a complete PVNN framework 134
Chapter 5. Riemannian Neural Networks from which different network architectures can be constructed. We validate the framework through four sets of experiments, including numerical stability, image classification, graph learning, and genomic sequence learning, demonstrating both the stability of PV embeddings and effectiveness of PVNNs. To our knowledge, the PV model has remained largely unexplored in machine learning, and our work provides the first systematic study of its use for representation learning. In summary, our contributions are threefold: (1) We establish the complete Riemannian geometric toolkit of the PV manifold, deriving closed-form operators that enable its use as a new alternative to classical hyperbolic models. (2) We develop fundamental building blocks in PV space, including MLR, FC, convolutional, activation, and BN layers. (3) We validate the stability and effectiveness of PVNNs through experiments on four tasks: numerical stability, image classification, graph node classification, and genomic sequence learning.1 Outline. In Sec. 5.2.2, we introduce the PV model and its gyrovector operations. In Sec. 5.2.3, we develop the Riemannian geometry and closed-form operators of PV space. In Sec. 5.2.4, we construct the core layers of PVNNs. In Sec. 5.2.5, we connect PV constructions to hyperboloid neural layers, and in Sec. 5.2.6, we evaluate their numerical stability and effectiveness. Proofs are deferred to Sec. B.6.
5.2.2
Preliminaries
PV Space [200]. As shown in Sec. 2.9.5, hyperbolic space is a space with constant negative curvature K < 0 and admits several models one can work with. The popular models include the Poincaré ball and the hyperboloid (also known as the Lorentz model). The PV model PVnK = Rn is an alternative representation of hyperbolic geometry, which was initially named the Ungar gyrovector space and is used to describe algebraic structures of relativistic proper velocities [200]. Unlike the bounded Poincaré ball or the constrained hyperboloid, the PV model is an unconstrained space, offering better numerical stability. Its Riemannian metric is given by Sec. B.6.1: gx (u, v) = ⟨u, v⟩ + Kβx2 ⟨x, u⟩ ⟨x, v⟩ , 1
∀x ∈ PVnK , ∀u, v ∈ Tx PVnK .
The code is available at https://github.com/NickyoyoSu/PVNN.
135
(5.1)
5.2. Proper Velocity Neural Networks Here, βx = √
1 1−K∥x∥2
is the relativistic beta factor. In Ungar’s notation, the curvature
is parametrized by a positive constant s with s2 = −1/K, where s plays the role of the vacuum speed of light in special relativity [200, Sec. 3.8]. PV Gyrovector [200]. From an algebraic point of view, the PV space forms a gyrovector space [200, Def. 6.2], which extends the Euclidean vector space to manifolds. Given x, y, z ∈ PVnK and t ∈ R, PV gyroaddition ⊕U and scalar gyromultiplication ⊗U [200, Chs. 3.11 and 6.20] are defined as2
1 − βy βx x ⊕U y = x + y + −K ⟨x, y⟩ x, βy 1 + βx √ y t ⊗U y = sinh t sinh−1 −K ∥y∥ √ , −K ∥y∥
(5.2) (t ⊗U 0 = 0) .
(5.3)
In particular, the PV inverse is ⊖U x = −x, and the PV identity is the zero vector: 0 ⊕U x = x ⊕U 0 = x.
PV Gyration. As shown by Ungar [200, Eqs. 3.220 and 3.221], the PV gyration for any x, y, z ∈ PVnK is given by gyr[x, y]z = z +
Ax + By , D
(5.4)
where the coefficients are A = (1 − βy2 )K ⟨x, z⟩ − (1 + βx )(1 + βy )βx βy K ⟨y, z⟩ + 2βx2 βy2 K 2 ⟨x, y⟩ ⟨y, z⟩ ,
B = (1 − βx2 )βy2 K ⟨y, z⟩ + (1 + βx )(1 + βy )βx βy K ⟨x, z⟩ , D = (1 + βx )(1 + βy ) (1 − βx βy K ⟨x, y⟩ + βx βy ) . Here, βx = √
1 1−K∥x∥2
(5.5) (5.6) (5.7)
is the relativistic beta factor.
5.2.3
Proper Velocity Geometry
5.2.3.1
From Gyro Isomorphism to Riemannian Isometry
The Poincaré ball also admits a gyrovector space, named the Möbius gyrovector space, as reviewed in Sec. 2.9.5. Algebraically, the PV and Möbius gyrovector spaces are isomorphic. We further show that PV and the Poincaré ball are geometrically isometric. 2
The subscript U refers to the initial letter of Ungar.
136
Chapter 5. Riemannian Neural Networks The following bijections define the gyrovector space isomorphism [200, Tab. 6.1]: πPVnK →PnK : PVnK ∋ x 7→
βx x ∈ PnK , 1 + βx
where γy = √
is the gamma factor. The isomorphism preserves the gyro
operations:
1 1+K∥y∥2
(5.8)
πPnK →PVnK : PnK ∋ y 7→ 2γy2 y ∈ PVnK ,
πPVnK →PnK (x ⊕U y) = πPVnK →PnK (x) ⊕M πPVnK →PnK (y), πPVnK →PnK (r ⊗U x) = r ⊙M πPVnK →PnK (x),
∀x, y ∈ PVnK ,
∀x ∈ PVnK , ∀r ∈ R,
(5.9) (5.10)
where ⊙M and ⊕M are the Möbius gyro operations reviewed in Sec. 2.9.5. Lemma 121 (Differentials). [↓] The differentials of πPVnK →PnK and πPnK →PVnK are dx πPVnK →PnK (v) = K
βx βx3 ⟨x, v⟩ x + v, 2 (1 + βx ) 1 + βx
dy πPnK →PVnK (w) = −4Kγy4 ⟨y, w⟩ y + 2γy2 w,
∀x ∈ PVnK , ∀v ∈ Tx PVnK ,
∀y ∈ PnK , ∀w ∈ Ty PnK .
Let id be the identity map. The differentials at the origin 0 are d0 πPVnK →PnK = 12 id,
d0 πPnK →PVnK = 2 id .
(5.11)
Based on Thm. 121, we can prove that the above isomorphisms are isometries. Theorem 122 (Isometries). [↓] The mappings in Eq. (5.8) are Riemannian isometries. 5.2.3.2
Proper Velocity Riemannian Operators
The Poincaré ball admits the closed-form Riemannian operators reviewed in Sec. 2.9.5. By Thm. 122, we can readily obtain the counterparts on PV space via the properties of Riemannian isometries reviewed in Thm. 33.
137
5.2. Proper Velocity Neural Networks Theorem 123 (PV Riemannian operators). [↓] Let π = πPVnK →PnK . Given x, y ∈ PVnK and v ∈ Tx PVnK , the Riemannian operators on the PV space are Expx (v) = x ⊕U
1 √ sinh −K
√ dx π(v) −K(1 + βx ) ∥dx π(v)∥ , βx ∥dx π(v)∥
Logx (y) = σ(x, y)z + τ (x, y) ⟨x, z⟩ x,
1 + βx (1 + βx )βy ṽ − K ⟨y, ṽ⟩ y, βx (1 + βy )βx √ 2 tanh−1 −K ∥π(−x ⊕U y)∥ , d(x, y) = √ −K
PTx→y (v) =
(5.12) (5.13) (5.14) (5.15)
with z = (−x) ⊕U y. For the parallel transport, ṽ = gyrM [ȳ, −x̄] (dx π(v)) with gyrM βy βx x and ȳ = 1+β y. Here, the scalar as the Möbius gyration in Sec. 2.9.5, x̄ = 1+β x y coefficients in the logarithm are √ 2 tanh−1 −K ∥π(z)∥ σ(x, y) = √ , ∥z∥ −K √ √ −K tanh−1 −K ∥π(z)∥ 2βx τ (x, y) = . 1 + βx ∥z∥
(5.16)
At the identity 0, the above operators can be further simplified: √ v 1 , sinh −K ∥v∥ ∥v∥ −K √ y 1 Log0 (y) = √ sinh−1 −K ∥y∥ , ∥y∥ −K βy ⟨y, v⟩ y, PT0→y (v) = v − K 1 + βy βx2 PTx→0 (v) = v + K ⟨x, v⟩ x, 1 + βx √ 1 d(0, y) = √ sinh−1 −K ∥y∥ . −K Exp0 (v) = √
(5.17) (5.18) (5.19) (5.20) (5.21)
The expressions containing normalized vectors or ∥z∥−1 are understood by continuous extension in the zero cases. Thus, Expx (0) = x and Logx (x) = 0. In particular, Exp0 (0) = Log0 (0) = 0. This implies that PV gyro operations can be expressed via Riemannian operations.
138
Chapter 5. Riemannian Neural Networks Theorem 124 (Gyro by Riemannian). [↓] The PV gyro operations can be rewritten as x ⊕U y = Expx (PT0→x (Log0 (y))) , ∀x, y ∈ PVnK , (5.22) t ⊗U x = Exp0 (t Log0 (x)) , ∀x ∈ PVnK , ∀t ∈ R.
5.2.4
Proper Velocity Neural Networks
Building on the above gyrovector and Riemannian tools, we introduce fundamental building blocks for PV neural networks, including MLR, FC, convolutional, activation, and BN layers, thereby enabling the construction of concrete deep architectures in this space. 5.2.4.1
Proper Velocity Multinomial Logistic Regression
Following the point-to-hyperplane formulation in Sec. 4.2.2.1, we define the PV margin hyperplane and solve the corresponding point-to-hyperplane infimum under the PV geometry. We define the PV hyperplane as o n n Ha,p = x ∈ PVK | Logp (x), a p = 0 ,
p ∈ PVnK , a ∈ Tp PVnK ,
(5.23)
where p ∈ PVnK and a ∈ Tp PVnK are the hyperplane parameters. As the Poincaré hyperplane can be expressed by the Möbius gyro operations [76, Eq. (22)], the PV hyperplane can also be expressed by the PV gyro operations. In addition, building PV MLR requires the PV point-to-hyperplane distance. The following theorem provides these results. Theorem 125. [↓] Let π = πPVnK →PnK . Given x, p ∈ PVnK and a ∈ Tp PVnK , we have n o Ha,p = x ∈ PVnK | Logp (x), a p = 0 = {x ∈ PVnK | ⟨−p ⊕U x, dp π(a)⟩ = 0} , √ −K |⟨−p ⊕U y, dp π(a)⟩| 1 −1 sinh . d(y, Ha,p ) = inf d(y, w) = √ w∈Ha,p ∥dp π(a)∥ −K By Thm. 125, we define the C-class PV MLR as p(y = k | x) ∝ exp (vk (x)) , vk (x) = sign (⟨−pk ⊕U x, dpk π(ak )⟩) ∥ak ∥pk d (x, Hak ,pk ) ,
(5.24)
where pk ∈ PVnK and ak ∈ Tpk PVnK are the PV MLR parameters for class k. However, 139
5.2. Proper Velocity Neural Networks the above expression has three drawbacks: (i) the parameter pk is over-parameterized, as it corresponds to the scalar bias parameter in the Euclidean MLR; (ii) the gyroaddition in ⟨−pk ⊕U x, dpk π(ak )⟩ complicates the computation; and (iii) the parameters (pk , ak ) are constrained, making optimization costly. To address these drawbacks, we follow Shimizu et al. [181] and adopt the parameterization pk = Exp0 (rk zk /∥zk ∥), ak = PT0→pk (zk ) with zk ∈ T0 PVnK ∼ = Rn and rk ∈ R. This parameterization avoids Riemannian optimization in PV MLR and further simplifies the formulation. Theorem 126 (PV MLR). [↓] For x ∈ PVnK , the score vk (x) in Eq. (5.24) for each class k is √ p √ √ −K ∥zk ∥ −1 2 vk (x) = √ sinh cosh( −Krk ) ⟨x, zk ⟩ − sinh( −Krk ) 1 − K∥x∥ , ∥zk ∥ −K
(5.25)
where zk ∈ Rn and rk ∈ R are parameters for class k. In particular, as K → 0− we have vk (x) → ⟨x, zk ⟩ + bk with bk = −rk ∥zk ∥, which recovers the Euclidean MLR reviewed in Sec. 4.2.2.1. The parameterization (zk , rk ) is essential for efficiency. In the original form Eq. (5.24), computing vk (x) for a batch x ∈ Rb×n and C classes requires explicit gyroaddition −pk ⊕U x for each class, producing an intermediate tensor of size b × C × n that could cause out-of-memory errors in high dimensions. One could instead loop over classes, but this is computationally inefficient. In contrast, Eq. (5.25) depends on inner products ⟨x, zk ⟩, which can be implemented as a matrix multiplication. 5.2.4.2
Proper Velocity Fully Connected Layer
The Euclidean Fully Connected (FC) layer is defined as y = Ax + b with A ∈ Rm×n and b ∈ Rm . It can be expressed element-wise as yk = ⟨ak , x⟩ − bk = ⟨ak , x − pk ⟩ with ak , pk ∈ Rn and ⟨pk , ak ⟩ = bk . As shown by Shimizu et al. [181, Sec. 3.2] and Chen et al. [55, Sec. 3.1], the LHS yk is the signed distance from y to the hyperplane passing through the origin and orthogonal to the k-th axis of the output space, which can be formulated as sign (⟨ek , y − 0⟩) d(y, Hek ,0 ) = ⟨ak , x − pk ⟩ ,
∀1 ≤ k ≤ m,
(5.26)
where ek denotes the vector whose k-th element is 1 and all others are 0. For the PV model, the LHS of Eq. (5.26) can be formulated by the signed point-tohyperplane distance, while the RHS can be formulated by the vk in PV MLR. Specifi140
Chapter 5. Riemannian Neural Networks cally, the PV FC layer F : PVnK → PVm K from the n-dimensional to the m-dimensional PV spaces for the input x ∈ PVnK returns the output y ∈ PVm K by solving the m equations: sign (⟨d0 π(ek ), −0 ⊕U y⟩) d(y, Hek ,0 ) = vk (x),
∀1 ≤ k ≤ m,
(5.27)
where Hek ,0 and vk (x) are given by Thm. 125 and Eq. (5.25), respectively. This definition has an explicit solution. Theorem 127 (PV FC layer). [↓] The output y = F(x) ∈ PVm K has the closed form √ 1 yk = √ sinh( −Kvk (x)), 1 ≤ k ≤ m, (5.28) −K
where vk (x) is defined in Eq. (5.25) with zk ∈ Rn and rk ∈ R as the FC parameters. In particular, as K → 0− we have yk → ⟨x, zk ⟩ + bk with bk = −rk ∥zk ∥, which recovers the Euclidean FC layer.
Generalization. We can jointly express the Euclidean FC layer and activation σ, which yields the RHS of Eq. (5.26) with σ (⟨ak , x − pk ⟩). Accordingly, we extend the PV FC by applying the activation to vk (x) in Eq. (5.27). Then, Eq. (5.28) becomes √ 1 yk = √ sinh( −Kσ(vk (x))), −K 5.2.4.3
1 ≤ k ≤ m.
(5.29)
Proper Velocity Convolution and Activation
Convolution. As shown by Shimizu et al. [181], Bdeir et al. [15], Chen et al. [55], Euclidean convolution consists of linear maps between kernel weights and concatenated values in each receptive field. To define convolution on PV space, it therefore suffices to define PV concatenation, since we already have the PV FC layer. Because PV space is unconstrained, we define PV concatenation to coincide with Euclidean concatenation. For simplicity, we consider the 1D case. For PV inputs {xi ∈ PVnK }ki=1 in a 1D receptive field (where k is the kernel size), the PV convolution output y ∈ PVm K for this receptive field is y = F (Concat (x1 , . . . , xk )), where Concat(·) is standard Euclidean concatenation and F is the PV FC layer. Activation. A natural choice is to apply a Euclidean activation σ in the tangent space at the origin via the mapping x 7→ Exp0 (σ (Log0 (x))), which has been shown to be effective in Poincaré networks [76]. Alternatively, since PV space is unconstrained, we can apply the activation directly in PV space as x 7→ σ(x). This direct PV-space 141
5.2. Proper Velocity Neural Networks activation avoids exponential and logarithmic maps and is therefore more efficient. 5.2.4.4
Proper Velocity Normalization
We instantiate the GyroBN framework in Sec. 3.3.3 on the PV space and show that PV GyroBN can normalize sample statistics. Given activations {xi ∈ PVnK }N i=1 , the core operations of PV GyroBN are Centering Scaling }| { z z }| { z }| { s x̃i ← B⊕U √ ⊗U −M ⊕U xi , v2 + ϵ Biasing
∀i ≤ N,
(5.30)
where M and v 2 denote the Fréchet mean and variance, and B ∈ PVnK and s ∈ R are parameters. Owing to the isometry between the PV space and Poincaré ball, the PV Fréchet mean can be computed via the Poincaré ball: map the data to the Poincaré ball, compute the Poincaré mean [142, Alg. 1], and map the result back. The following theorem guarantees that PV GyroBN can normalize sample statistics. n Theorem 128 (Homogeneity). [↓] For N samples {xi }N i=1 ⊂ PVK , we have
N Homogeneity of mean: FM {B ⊕U xi }N ∀B ∈ PVnK , i=1 = B ⊕U FM {xi }i=1 , 1 XN 2 1 XN 2 Homogeneity of dispersion from 0: d (t ⊗U xi , 0) = t2 · d (xi , 0). i=1 i=1 N N
Thm. 128 directly explains the PV GyroBN in Eq. (5.30). After the centering, the batch mean is shifted to the identity 0. After the scaling, the variance becomes s2 . After the biasing, the batch mean is translated to B.
5.2.5
Connections to the Hyperboloid
This subsection discusses the connections between the PV model and the hyperboloid model. We first show the isometry between the two models. Then, we show that several current hyperboloid network layers can be rewritten as PV layers.
142
Chapter 5. Riemannian Neural Networks Proposition 129 (PV–hyperboloid isometries). [↓] The following maps are Riemannian isometries between the hyperboloid model HnK and the PV model PVnK : " # xt πHnK →PVnK :HnK ∋ 7→ xs ∈ PVnK , xs q 2 1 ∥x∥ − K πPVnK →HnK :PVnK ∋ x 7→ ∈ HnK . x
(5.31) (5.32)
The PV–hyperboloid isometries in Thm. 129 imply that several standard layers in hyperboloid networks can be rewritten as PV layers composed with πHnK →PVnK and πPVnK →HnK . The Lorentz activation [15, Eq. (13)], Lorentz FC layer [45, Sec. 3.1] and Lorentz concatenation [15, Eq. (32)] are " #! q 2 1 xt ∥σ(xs )∥ − K LAct = , xs σ(xs ) " #! q 2 1 xt ∥W xs + b∥ − K , LFC = xs W xs + b qP N N −1 2 i=1 xi,t + K x1,s N HCat({xi }i=1 ) = ∈ HnN K , . .. xN,s
(5.33)
(5.34)
(5.35)
⊤ ⊤ n ⊤ n where x = [xt , x⊤ s ] ∈ HK and xi = [xi,t , xi,s ] ∈ HK for 1 ≤ i ≤ N . Then Thm. 129 implies that the above Lorentz layers can be rewritten in terms of PV layers as follows:
LAct(x) = πPVnK →HnK (σ(πHnK →PVnK (x))),
(5.36)
LFC(x) = πPVnK →HnK (W πHnK →PVnK (x) + b),
(5.37)
n HCat({xi }N i=1 ) = πPVK →Hn K
Concat(πHnK →PVnK (x1 ), . . . , πHnK →PVnK (xN )) .
(5.38)
These identities show that many hyperboloid constructions effectively operate by mapping to PV space, applying Euclidean building blocks there, and mapping back through πPVnK →HnK . This perspective naturally motivates designing networks directly in PV space, instead of repeatedly switching between equivalent models. Moreover, even if one 143
5.2. Proper Velocity Neural Networks m follows the pattern HnK → PVnK → PVm K → HK to construct layers, the intermediate map should be the PV layers, such as Thm. 127, rather than Euclidean layers, since PV is a non-linear Riemannian manifold.
5.2.6
Experiments
We evaluate PV embeddings and PVNNs on four representative tasks: • Sec. 5.2.6.1 evaluates the numerical advantage of the PV model against Poincaré and hyperboloid. • Sec. 5.2.6.2 compares PV, Poincaré, and hyperboloid MLRs on image classification. • Sec. 5.2.6.3 evaluates our PV MLR, FC, and GyroBN layers on graph learning. • Sec. 5.2.6.4 compares fully PV convolutional networks with fully hyperboloid convolutional networks on genomic sequence learning. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.4. 5.2.6.1
Numerical Stability
We study three aspects: gyro operator, Riemannian operator, and gradient behavior. All experiments use curvature K = −1, dimension n = 16, and batch size 4096.
Failure rate
r 1 5 10 20 50 75 100 150 200 1000
Violation rate
PVnK
PnK
HnK
PVnK
PnK
HnK
0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0
0 0 0 4.23 64.42 79.63 88.26 96.43 100 100
N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A
0 0 0 0 0 0 0 0 0 0
32.50 92.36 99.76 100 100 100 100 100 100 100
Gyro Operator. We use scalar gyromultiplication r ⊗H x as a probe of numerical stability across hyperbolic models. Given random batches x and radii r, we evaluate two metrics. The failure rate is the fraction of outputs that contain Table 5.1: Failure and violation rates (%) NaN/Inf. The violation rate is defined of r ⊗H x in FP32. only for models with manifold constraints: Poincaré ball requires ∥x∥2 < −1/K, and hyperboloid requires ⟨x, x⟩L = ∥xs ∥2 −x2t = K1 ⊤ −8 for x = [xt , x⊤ s ] . The tolerance is set to 10 . As PV is unconstrained, its violation rate is reported as N/A. As shown in Tab. 5.1, PV maintains zero failures up to r = 1000 in 144
Chapter 5. Riemannian Neural Networks FP32. The Poincaré ball has zero failure and violation rates, whereas the hyperboloid model starts to fail around r = 20 and quickly accumulates both NaN/Inf outputs and off-manifold points under large scalar multipliers, revealing pronounced numerical instability. Riemannian Operator. We evaluate the exponential and logarithmic maps by measuring the Model FP32 FP64 round-trip error ∥Log0 (Exp0 (v)) − v∥ for tangent PnK 2.1 × 10−4 4.3 × 10−11 n 0 HK 1.0 × 10 1.0 × 100 vectors v with large norm ∥v∥ = 10. Since this n PVK 2.1 × 10−7 6.7 × 10−16 quantity is theoretically zero, any non-zero value reflects numerical instability. We sample a batch Table 5.2: ∥Log (Exp (v)) − v∥. 0 0 of such vectors and report the average error in Tab. 5.2. PV achieves stable behavior in both FP32 and FP64, whereas the Poincaré ball already exhibits noticeable errors in FP32 and the hyperboloid model remains unstable in both precisions. Gradient. To compare gradient behavior, we study the graModel ∥∇x fr (x)∥ Range Gradient behavior dient of fr (x) = ∥r ⊗H x − x∥ PnK [7.6 × 10−13 , 1.1 × 10−11 ] Vanishing gradients n HK [0, NaN] Exploding gradients with respect to x. SpecifiPVnK [2.1 × 10−6 , 1.1 × 10−4 ] Stable gradients cally, we sample 24 logarithmically spaced radii r ∈ [1, 1000] Table 5.3: Gradient magnitude ∥∇x fr (x)∥ across and, for each radius, measure the varying radii. ∥∇x fr (x)∥ on a random batch. The range of ∥∇x fr (x)∥ is summarized in Tab. 5.3. The Poincaré ball exhibits severe gradient vanishing near the boundary. In contrast, the hyperboloid model yields gradients that vary from 0 to NaN, reflecting gradient explosion. PV maintains gradients in a safer band. 5.2.6.2
Image Classification
We compare our PV MLR against previous Poincaré MLRs [76, 181] and Lorentz MLR [15]. Following Bdeir et al. [15], we train a ResNet-18 backbone [97] on CIFAR-10 and CIFAR-100 [126], replacing the final Euclidean MLR with a hyperbolic MLR. The backbone output is lifted to the target geometry via the exponential map at the identity. Since PV space is unconstrained, we also consider a direct variant that skips Exp0 and treats the backbone output as PV coordinates. We denote these two PV heads as PV MLR (with Exp0 ) and PV MLR (without Exp0 ). Tab. 5.4 reports the 5-fold results. 145
5.2. Proper Velocity Neural Networks Model
Method
CIFAR-10 (δ = 0.26)
CIFAR-100 (δ = 0.23)
PnK
Poincaré MLR [76, Eq. (25)] Unidirectional MLR [181, Eq. (6)]
HnK
Lorentz MLR [15, Thm. 2]
95.09 ± 1.51 95.12 ± 0.20
76.78 ± 0.67 77.19 ± 0.10
PVnK
PV MLR (with Exp0 ) PV MLR (without Exp0 )
95.27 ± 0.12 95.30 ± 0.18
78.19 ± 0.59 78.20 ± 0.37
95.02 ± 0.12
77.96 ± 0.09
Table 5.4: Top-1 image classification accuracy (%) of hyperbolic MLRs on ResNet-18. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic), which comes from Bdeir et al. [15, Tab. 1]. Model
Method
Disease (δ = 0)
Airport (δ = 1)
PubMed (δ = 3.5)
Cora (δ = 11)
KnK
KNN [147]
92.10 ± 0.97
69.36 ± 0.76
52.26 ± 1.99
PnK
HNN [76] HNN++ [181]
79.41 ± 0.55
HnK
LNN [15]
PVnK
PVNN
79.90 ± 0.01
75.20 ± 1.08
68.82 ± 0.88
79.90 ± 0.01 80.57 ± 0.23
81.15 ± 0.23
82.16 ± 2.95 88.40 ± 0.17
97.96 ± 0.42
69.28 ± 0.85 73.68 ± 0.39
74.33 ± 0.22
49.68 ± 1.25 52.06 ± 0.90
53.34 ± 1.65 51.42 ± 1.33
Table 5.5: Accuracies of hyperbolic networks on graph learning. The best results are bold. δ represents the δ-hyperbolicity (lower is more hyperbolic). PV MLR matches or outperforms prior hyperbolic baselines, with the largest gains on CIFAR-100 where the decision boundaries are more complex. Both PV variants, with and without Exp0 , achieve similar accuracies. 5.2.6.3
Graph Learning
Data and Setup. We study node classification on four standard graph data sets: Disease [4], Airport [229], Cora [177], and PubMed [155]. All models share the same architecture consisting of two FC layers with nonlinear activations followed by an MLR classifier; they differ only in the underlying hyperbolic model. Baselines include KNN [147] for the Klein ball, HNN/HNN++ [38, 181] for the Poincaré ball, and LNN [15] for the hyperboloid model. Our PVNN is built from PV FC, activation, and MLR layers. Main Results. For a fair comparison, we use a tangent activation in each model and set σ = id for the PV FC layer in Eq. (5.29). Tab. 5.5 summarizes the 5-fold results. On the three more hyperbolic data sets (Disease, Airport, and PubMed), PVNN consistently achieves the best performance, with especially large gains on Airport where it improves over the strongest baseline by 5.86%. On the weakly hyperbolic Cora data set, PVNN remains comparable to Poincaré- and Klein-based networks, and worse than 146
Chapter 5. Riemannian Neural Networks Method
Disease
Airport
PubMed
Cora
PVNN+TFC PVNN
80.86 ± 0.30 81.24 ± 0.36
86.99 ± 0.61 97.93 ± 0.29
PVNN+TBN PVNN+GyroBN
80.67 ± 0.38 81.24 ± 0.19
98.71 ± 0.36 99.03 ± 0.18
74.40 ± 0.43 74.16 ± 0.32
53.58 ± 0.81 52.26 ± 1.32
73.52 ± 0.12 74.34 ± 0.31
45.36 ± 2.44 46.64 ± 5.45
Table 5.6: Results of Tangent FC (TFC) vs PV FC, and Tangent BN (TBN) vs GyroBN. Method Tangent Euclidean Fréchet 1 iter Fréchet 2 iters Fréchet 5 iters Fréchet 10 iters Fréchet ∞
Disease
Airport
PubMed
Cora
Acc
Fit Time
Acc
Fit Time
Acc
Fit Time
Acc
Fit Time
81.15 ± 0.23 81.15 ± 0.23
26.08 25.80
98.56 ± 0.36 98.75 ± 0.31
55.48 55.19
61.50 ± 5.75 69.82 ± 3.58
3.10 2.99
33.10 ± 1.58 32.62 ± 0.65
7.12 7.29
81.05 ± 0.23 81.05 ± 0.23 81.24 ± 0.36 81.24 ± 0.19 80.86 ± 0.00
29.79 30.12 30.90 30.49 31.29
88.93 ± 1.17 94.11 ± 0.46 98.50 ± 0.16 99.03 ± 0.18 98.46 ± 0.15
65.19 67.37 82.28 105.79 122.37
62.52 ± 8.44 73.78 ± 0.20 73.92 ± 0.44 74.34 ± 0.31 71.16 ± 3.93
3.38 3.49 4.02 3.96 4.46
42.84 ± 6.15 45.68 ± 4.36 49.50 ± 1.83 46.64 ± 5.45 47.32 ± 4.73
7.67 8.21 9.15 9.77 9.27
Table 5.7: Comparison of methods in calculating mean and variance in PV GyroBN. Time is measured in milliseconds per training epoch. the hyperboloid-based one. Overall, these results suggest that PV geometry is more effective on strongly hyperbolic graphs. Tangent vs. Riemannian. A natural construction of hyperbolic layers is to work in the tangent space. To validate the benefits of our Riemannian PV layers, we compare our PV FC with TFC of the form Exp0 (A Log0 (x) + b), and our GyroBN with TBN given by Exp0 (BN(Log0 (x))) [109]. We denote these variants by PVNN+TFC and PVNN+TBN, respectively. As shown in Tab. 5.6, PVNN consistently outperforms PVNN+TFC on the more hyperbolic Disease and Airport data sets, while performance on the other two data sets is comparable and TFC can be slightly better. For normalization, PVNN+GyroBN improves over PVNN+TBN on all data sets. Overall, these ablations validate the effectiveness of our Riemannian PV constructions, especially in strongly hyperbolic settings. Ablations on Batch Statistics. PV GyroBN in Eq. (5.30) uses Fréchet mean and variance, which require iterative solvers. We also consider two efficient variants. A tangent variant computes batch statistics in the tangent space at the identity via M = Exp0
! N 1 X Log0 (xi ) , N i=1
1 X v2 = ∥Log0 (xi ) − Log0 (M )∥2 , N i=1 N
(5.39)
and a Euclidean variant computes standard Euclidean mean and variance directly in 147
5.2. Proper Velocity Neural Networks Disease
Airport
PubMed
Cora
81.05 ± 0.23
97.71 ± 0.34
74.22 ± 0.26
51.92 ± 2.01
Exp0 ✗ ✓
81.24 ± 0.36
97.93 ± 0.29
74.16 ± 0.32
52.26 ± 1.32
Table 5.8: Ablations on PVNN with or without exponential map for the input PV feature.
Method Tangent Act. FC σ FC σ + Tangent Act. Euc. Act.
Disease
Airport
PubMed
Cora
81.24 ± 0.36 81.34 ± 0.43 80.96 ± 0.19
97.93 ± 0.29 99.40 ± 0.15 99.15 ± 0.38
74.16 ± 0.32 74.02 ± 0.17 73.96 ± 0.22
52.26 ± 1.32 51.34 ± 0.46 51.30 ± 1.65
81.34 ± 0.43
98.87 ± 0.35
74.56 ± 0.59
38.10 ± 3.30
Table 5.9: Ablations on PV activations.
the unconstrained PV space. Tab. 5.7 shows that Tangent and Euclidean are up to 2× faster while achieving similar accuracies on Disease and Airport. Although Fréchetbased GyroBN attains the best accuracies, it is more computationally expensive. Ablations on PV Embedding. In the main experiments, the input features are first lifted to PV via Exp0 and then processed by PVNN. Since PV space is unconstrained, we also consider a variant that feeds the Euclidean features directly as PV coordinates. Tab. 5.8 compares these two settings. The two variants perform similarly, while using Exp0 provides small improvements on Disease, Airport, and Cora. This differs from image classification in Tab. 5.4, where the variant without Exp0 is marginally better. This slight discrepancy may stem from the different nature of the inputs. In vision, the ResNet encoder can adapt its learned representation to the chosen lifting, whereas in graphs the raw node features benefit slightly from the explicit exponential map. Ablations on Activation. We ablate three types of activations in PVNN: the internal nonlinearity σ in the PV FC layer (fixed to tanh), and explicit activations applied either directly in PV (Euc. Act.) or in the tangent space (Tangent Act.). Tab. 5.9 reports the results. First, when comparing these three choices individually, the differences are small on Disease and PubMed, while FC σ performs best on Airport, Tangent Act. performs best on Cora, and Euc. Act. degrades substantially on Cora. Second, when comparing the composite variant FC σ + Tangent Act. against Tangent Act., the composite does not yield consistent gains, suggesting redundancy. 148
Chapter 5. Riemannian Neural Networks Task
Data Set
Euclidean CNN
HCNN-S
PVCNN
Retrotransposons
LINEs SINEs hAT-Ac
76.12 ± 2.16 85.45 ± 1.16
81.83 ± 0.27 93.78 ± 0.54
DNA transposons
70.63 ± 1.24 85.15 ± 1.64
60.66 ± 0.82 51.94 ± 2.69
68.30 ± 0.93 56.10 ± 0.56
71.27 ± 0.78 62.31 ± 0.78
Pseudogenes
processed unprocessed
87.45 ± 0.90
89.61 ± 1.34
92.08 ± 0.80
Table 5.10: Comparison in MCC of hyperbolic and Euclidean convolutional networks, including PVCNN, on TEB data sets. 5.2.6.4
Genomic Sequence Learning
Khan et al. [119] recently proposed Hyperbolic Convolutional Neural Networks (HCNNs) on the hyperboloid for genomic sequence learning, demonstrating that HCNNs outperform Euclidean CNNs on this task. Following Khan et al. [119], we evaluate on the Transposable Elements Benchmark (TEB) data set for DNA transposable element prediction. To ensure a fair comparison, all models share the same backbone network architecture, which consists of two convolutional blocks followed by an FC layer and a final MLR classifier [119]. We use a single curvature shared for all layers. Tab. 5.10 reports 5-fold Matthews Correlation Coefficient (MCC). The PV Convolutional Network (PVCNN) achieves the best performance on all TEB tasks, with particularly strong gains on SINEs, where it improves over HCNN-S by about 9 MCC points. These results demonstrate the benefits of PV convolutional networks.
5.3
Hyperbolic Busemann Neural Networks
5.3.1
Introduction
The preceding section introduced the unconstrained PV representation and its neural layers, based on geodesic hyperplanes. We next use another powerful geometric tool, the Busemann function, to construct hyperbolic neural networks. For simplicity, we focus on the widely used Poincaré ball and Lorentz models. To support deep learning fully in hyperbolic spaces, several key building blocks in neural networks have recently been generalized to Poincaré or Lorentz spaces, including attention [90, 45, 221], BN [142, 15, 52, 54], linear feed-forward layers [76, 181, 45], activation [76, 15], residual blocks [201, 118, 99], MLR [181, 15, 162], and graph convolution [38, 139, 13, 60]. Among these components, MLR classification and FC layers play a fundamental role in final decision-making and feature transformation. 149
5.3. Hyperbolic Busemann Neural Networks Recently, hyperplanes and point-to-hyperplane distances, which have been explored in Chapter 4 and Sec. 5.2, have been adopted to construct hyperbolic MLR in both Poincaré [76, 181] and Lorentz [15] models. Ganea et al. [76, Sec. 3.1] introduced the first Poincaré MLR based on Poincaré hyperplanes, but the formulation suffers from over-parameterization and lacks batch efficiency. Shimizu et al. [181, Sec. 3.1] alleviated these issues through a re-parameterization strategy. Building on these ideas, Bdeir et al. [15, Sec. 4.3] proposed a Lorentz MLR. However, its hyperplane is defined by the ambient Minkowski space, which is model-specific and may distort Lorentzian geometry. For hyperbolic FC layers, three main formulations exist. Ganea et al. [76, Sec. 3.2] introduced Möbius matrix-vector multiplication through the tangent space on the Poincaré ball. Shimizu et al. [181, Sec. 3.2] further proposed the Poincaré FC layer, defined intrinsically but restricted to the Poincaré model. On the Lorentz model, Chen et al. [45, Sec. 2.2] constructed a Lorentz FC layer by applying linear transformations in the ambient Minkowski space followed by projection onto the Lorentz model. Thus, Möbius and Lorentz FC rely on flat-space (tangent or ambient) approximations that could distort intrinsic geometry, whereas Poincaré FC is intrinsic but model-specific. On the other hand, the Busemann function and its level sets, horospheres, have emerged as powerful intrinsic tools for hyperbolic learning. They enjoy convenient metric properties [29, Ch. II.8] and admit closed-form expressions on both the Poincaré and Lorentz models [24, Prop. 9]. These operators have supported several hyperbolic algorithms, including SVM [73], PCA [39], Sliced Wasserstein distances [24], and prototype learning [81]. We also note that Nguyen et al. [162, Cor. 4.3] proposed a Poincaré MLR based on the Busemann function. However, its induced point-to-hyperplane distance is pseudo, coincides with the true distance only in Euclidean geometry, remains over-parameterized, and is not batch efficient. These observations motivate intrinsic and batch-efficient formulations for MLR and FC layers that can operate on both the Poincaré ball and the Lorentz model. To address this need, we propose Busemann Multinomial Logistic Regression (BMLR) and Busemann Fully Connected (BFC) layers, two Busemann-based components for hyperbolic networks. Our contributions are summarized as follows: • We introduce BMLR, deriving intrinsic logits directly from Busemann functions with a point-to-horosphere distance interpretation. BMLR uses a compact perclass parameterization, eliminates manifold-valued parameters in prior MLRs, remains batch-efficient, and recovers Euclidean MLR as curvature tends to zero. 150
Chapter 5. Riemannian Neural Networks • We develop BFC layers by generalizing the FC and activation layers through the Busemann function, providing intrinsic constructions on both the Poincaré and Lorentz models. BFC preserves comparable complexity and parameter counts, and recovers Euclidean FC in the zero curvature limit. • We provide empirical validation across image classification, genome sequence learning, node classification, and link prediction. BMLR and BFC generally outperform existing hyperbolic layers. BMLR shows particularly large gains as the number of classes increases, and the Lorentz BMLR is the fastest among all hyperbolic MLRs.3 Outline. In Sec. 5.3.2, we recall Busemann functions and horospheres in hyperbolic space. In Sec. 5.3.3, we introduce BMLR and its point-to-horosphere interpretation. In Sec. 5.3.4, we develop BFC layers, and in Sec. 5.3.5, we evaluate both components. Proofs are deferred to Sec. B.7.
5.3.2
Preliminaries
The metric-geometric notions of geodesic rays, asymptotic rays, Busemann functions, horoballs, horospheres, and Hadamard spaces have been reviewed in Sec. 2.5. The Poincaré and Lorentz models and their Riemannian operators have been reviewed in Sec. 2.9.5. Their gyro operators are presented in Tab. 2.13 and Sec. 3.3.4.5, respectively. Their gyrovector spaces are denoted by {PnK , ⊕M , ⊙M } and {LnK , ⊕L , ⊙L }, respectively. In the following, we review the Busemann function on the hyperbolic space. In Euclidean space, the Busemann function associated with the geodesic γ(t) = tv that starts at 0 with unit direction v ∈ Sn−1 is B v (x) = − ⟨x, v⟩, which coincides, up to n a sign, with the inner product. Let HK ∈ {PnK , LnK } be a hyperbolic space. We write B v (x) for the Busemann function associated with the ray that emanates from the origin n n n e ∈ HK in the direction v ∈ Sn−1 ⊂ Te HK . For K < 0, v ∈ Sn−1 , and x ∈ HK , closed forms of the Poincaré and Lorentz Busemann functions [24, Prop. 9] are PnK : LnK :
3
! √ 2 v − −Kx 1 log , B v (x) = √ −K 1 + K ∥x∥2 √ 1 B v (x) = √ log −K (xt − ⟨xs , v⟩) . −K
The code is available at https://github.com/GitZH-Chen/HBNN.
151
(5.40) (5.41)
5.3. Hyperbolic Busemann Neural Networks
Poincaré
1
Lorentz
x2
xt
6 0
1
3
4 1
0 x1
1
(xs )01
0 )2 xs 4 4 (
1 4
Figure 5.1: Illustration: red curves are different horospheres of B v . The level sets of a Busemann function are horospheres, the hyperbolic counterpart of Euclidean hyperplanes. In Euclidean space, for a unit direction v, the hyperplanes Hτv = {x ∈ Rn | ⟨x, v⟩ = τ } with τ ∈ R are parallel. Analogously, for fixed v, the n horospheres Hτv = {x ∈ HK | B v (x) = τ } are equidistant, as established later in Thm. 132. Fig. 5.1 illustrates such horospheres, and Tab. 2.4 summarizes the above correspondence between Euclidean and hyperbolic notions.
5.3.3
Busemann Multinomial Logistic Regression
We begin by reformulating the Euclidean MLR, then lift it to hyperbolic space via the Busemann function, introducing BMLR. We also present a point-to-horosphere interpretation. Finally, we compare BMLR with existing hyperbolic MLRs, highlighting our advantages in geometric fidelity, parameterization, and computational efficiency. 5.3.3.1
Formulation
The Euclidean MLR softmax(Ax + b) computes the multinomial probability for each class k ∈ {1, . . . , C} given an input x ∈ Rn . It admits the inner product form: ∀k,
exp (⟨ak , x⟩ + bk ) p(y = k | x) = PC , exp (⟨a , x⟩ + b ) j j j=1
(5.42)
where ak ∈ Rn and bk ∈ R are the weight and bias for class k. We write p(y = k | x) ∝ exp (uk (x)) with uk (x) = ⟨ak , x⟩ + bk . Decomposing the weight vector into a magnitude 152
Chapter 5. Riemannian Neural Networks αk = ∥ak ∥ > 0 and a unit direction vk = ∥aakk ∥ ∈ Sn−1 , each logit is uk (x) = αk ⟨vk , x⟩ + bk .
(5.43)
As reviewed in Sec. 5.3.2, the Busemann function naturally generalizes the Euclidean inner product. Analogously to Eq. (5.43), we define the hyperbolic logits via the Busemann function, yielding BMLR: ∀k,
exp (uk (x)) p(y = k | x) = PC , j=1 exp (uj (x)) uk (x) = −αk B vk (x) + bk ,
(5.44) (5.45)
with αk > 0, vk ∈ Sn−1 , and bk ∈ R as parameters. The following result shows that, as K → 0− , both the Poincaré and Lorentz BMLRs reduce to the Euclidean MLR. Theorem 130 (Limits of BMLRs). [↓] As K → 0− , the hyperbolic Busemann functions converge to the Euclidean inner product: Poincaré: Lorentz:
B v (x) −−−−→ −2 ⟨v, x⟩ ,
K→0−
(5.46)
K→0−
B v (x) −−−−→ − ⟨v, xs ⟩ .
(5.47)
The hyperbolic BMLRs converge to the Euclidean MLR: uk (x) −−−−→ 2αk ⟨vk , x⟩ + bk ,
K→0−
(5.48)
Lorentz: uk (x) −−−−→ αk ⟨vk , xs ⟩ + bk .
K→0−
(5.49)
Poincaré:
Remark 131 (Intuition). On the Poincaré ball, letting K → 0− recovers Euclidean geometry [182, App. A.4.2]. For the Lorentz model, as K → 0− , the temporal coordinate diverges while the spatial component approaches Rn , making LnK converge to a Euclidean space. Consistently, the Poincaré and Lorentz Busemann functions and the associated BMLR logits reduce to their Euclidean counterparts, providing a natural generalization of Euclidean MLR. 153
5.3. Hyperbolic Busemann Neural Networks 5.3.3.2
Geometric Interpretation
The point-to-hyperplane strategy underlying MLR has been established in Sec. 4.2.2.1. We now show that BMLR admits the corresponding interpretation through point-tohorosphere distances. Theorem 132 (Hadamard horosphere distance). [↓] Let (X , d) be a geodesically complete Hadamard space, and let B γ : X → R be the Busemann function associated with a geodesic ray γ : [0, ∞) → X . For any τ1 , τ2 ∈ R, define the horospheres by i = 1, 2. (5.50) Hτγi = {x ∈ X | B γ (x) = τi } , The distance between these horospheres is constant: d Hτγ1 , Hτγ2 = d Hτγ2 , Hτγ1 = |τ2 − τ1 | .
(5.51)
In particular, the point-to-horosphere distance is d (x, Hτγ ) = |B γ (x) − τ | ,
∀x ∈ X .
(5.52)
Corollary 133 (Point-to-horosphere distance). In a hyperbolic space n HK ∈ {PnK , LnK }
(5.53)
with curvature K < 0, the point-to-horosphere distance is d (x, Hτv ) = |B v (x) − τ | ,
(5.54)
where Hτv = {x | B v (x) = τ } denotes the horosphere with respect to direction v ∈ Sn−1 . A Euclidean hyperplane can be parameterized by a unit direction, a positive magnitude, and a scalar bias. Similarly, we parameterize a hyperbolic horosphere as n Hv,α,b = {x ∈ HK | −αB v (x) + b = 0} ,
(5.55)
with v ∈ Sn−1 , α > 0, and b ∈ R. With this parameterization, the signed point-tohorosphere logit is uk (x) = signk αk d (x, Hvk ,αk ,bk ) , (5.56) 154
Chapter 5. Riemannian Neural Networks Method
Logit uk (x), ∀k ∈ {1, . . . , C}
Space
Dist
#Params
Compact params
FLOPs
Batch efficiency
Euclidean MLR
⟨ak , x⟩ + bk , with ak ∈ Rn , bk ∈ R ! √ λK ∥a ∥ 2 −K ⟨−pk ⊕M x, ak ⟩ k pk √ sinh−1 , 2 −K 1 + K ∥−pk ⊕M x∥ ∥ak ∥ with pk ∈ PnK , ak ∈ Tpk PnK
Rn
Real
C(n + 1)
✓
C(2n)
✓
PnK
Real
C(2n)
✗
C(19n + 29)
✗
PnK
Real
C(n + 2)
✓
C(4n + 52)
✓
PnK
Pseudo
C(2n)
✗
C(19n + 34)
✗
LnK
Real
C(n + 1)
✓
C(4n + 52)
✓
PnK LnK
Real
C(n + 2)
✓
PnK : C(6n + 12) LnK : C(2n + 12)
✓
Poincaré MLR [76, Eq. (25)]
Poincaré MLR [181, Eq. (6)]
Pseudo-Busemann MLR [162, Cor. 4.3]
Lorentz MLR [15, Eq. (12)]
BMLR
2 √ αk sinh−1 (α − β), √ √−K α = λK −K ⟨x,vk ⟩ cosh√ 2 −Kb x k , β = λK x − 1 sinh 2 −Kbk , with αk > 0, vk ∈ Sn−1 , bk ∈ R
B vk (−pk ⊕M x) , ∥−pk ⊕M x∥ with pk ∈ PnK , vk ∈ Sn−1 √ α 1 √ , sign(α)β sinh−1 −K β −K√ √ αq = cosh −Kbk ⟨zk , xs ⟩ − sinh −Kbk , √ √ 2 β= cosh( −Kbk )zk − (sinh( −Kbk ) ∥zk ∥)2 , with zk ∈ Rn , bk ∈ R − d(x, pk )
−αk B vk (x) + bk , with αk > 0, vk ∈ Sn−1 , bk ∈ R
Table 5.11: Comparison of C-class MLR. In Dist, Real means the point-to-hyperplane distance is the real distance, obtained by inf y∈H d(x, y), where H is a hyperplane and d is the geodesic distance; Pseudo denotes a surrogate that coincides with the real distance only in Euclidean geometry. Compact params indicate whether each logit avoids an additional manifold-valued parameter. Batch efficiency indicates whether the MLR can avoid inefficient per-class loops in implementation (see Sec. A.4.2.1). In #Params, we highlight the heaviest in red. In FLOPs, we mark the slowest in red and the fastest in green. where signk = sign (−αk B vk (x) + bk ). By Thm. 133, the exact point-to-horosphere distance is |−αB v (x) + b| d (x, Hv,α,b ) = . (5.57) α Consequently, Eq. (5.56) equals the BMLR logit in Eq. (5.45). Remark 134 (Generality). Since B v (x) = − ⟨v, x⟩ in Euclidean geometry, Eqs. (5.55) to (5.57) naturally generalize to their Euclidean counterparts. We also acknowledge Fan et al. [73, Eq. (2) and Prop. 3.1], who used horospheres and point-to-horosphere distances to construct a hyperbolic SVM. However, they considered only the unit Poincaré ball with curvature K = −1, which is a special case of Eqs. (5.55) and (5.57).
5.3.3.3
Comparison with Existing Hyperbolic MLRs
Based on the point-to-hyperplane reformulation, recent work extended MLR to the Poincaré [76, 181, 162] and Lorentz [15] models. Ganea et al. [76, Sec. 3.1] introduced the first Poincaré MLR by replacing the Euclidean point-to-hyperplane distance with its hyperbolic counterpart, where the hyperplane is defined by geodesics and the resulting 155
5.3. Hyperbolic Busemann Neural Networks distance is the real point-to-hyperplane distance, obtained as an infimum over the hyperplane. However, the formulation is not batch efficient (see Sec. A.4.2.1). It also requires per-class parameters ak ∈ Tpk PnK and pk ∈ PnK , which leads to over-parameterization. Shimizu et al. [181, Sec. 3.1] alleviated such issues via re-parameterization. Bdeir et al. [15, Sec. 4.3] further developed a Lorentz MLR, but its hyperplanes are defined by the ambient Minkowski space, which is tailored to the Lorentz model and does not fully respect intrinsic hyperbolic geometry. Moreover, Nguyen et al. [162, Cor. 4.3] proposed a Poincaré MLR based on the Busemann function. We refer to it as Pseudo-Busemann MLR, as the induced point-to-hyperplane distance is pseudo, coinciding with the real point-to-hyperplane distance only in Euclidean geometry. It also suffers from overparameterization and is not batch efficient. As summarized in Tab. 5.11,4 BMLR unifies advantages that prior hyperbolic MLRs offer only partially. In particular, BMLR respects the real point-to-horosphere distance, uses compact parameters without an additional manifold-valued point, attains the lowest FLOPs on LnK and a competitive cost on PnK , and supports batch-efficient computation. On LnK , its FLOPs are even close to those of the Euclidean MLR.
5.3.4
Busemann Fully Connected Layer
We first reformulate the Euclidean FC layer by the Busemann function, then present the manifestations in the Poincaré and Lorentz models. 5.3.4.1
Formulation
As discussed in Sec. 5.2.4.2, the Euclidean FC layer can be written as d̄ (y, Hek ,0 ) = ⟨ak , x⟩ + bk ,
∀k ∈ {1, . . . , m},
(5.58)
where d̄ (y, Hek ,0 ) = sign (⟨ek , y − 0⟩) d(y, Hek ,0 ) is the signed distance. To extend Eq. (5.58) into hyperbolic space, the right-hand side can be replaced by Eq. (5.45), as it generalizes ⟨ak , x⟩ + bk . For the left-hand side, a natural idea is to use the signed point-to-horosphere distance. However, as detailed in Sec. A.4.2.2, this may fail to admit a solution for y. We therefore follow the point-to-hyperplane distance in [76, Thm. 5] for the Poincaré model and the one in [15, Eq. (44)] for the Lorentz model. n n m Given x ∈ HK , the hyperbolic BFC layer F : HK ∋ x 7→ y ∈ HK is given by solving y 4
Relative to [162, Def. 4.2, Cor. 4.3, and App. B.1.2], the Pseudo-Busemann MLR written here includes an additional sign −; this is intentional and matches their official implementation.
156
Chapter 5. Riemannian Neural Networks via the following m equations: (5.59)
∀k ∈ {1, . . . , m},
d̄ (y, Hek ,e ) = uk (x),
where uk (x) = −αk B vk (x) + bk with {αk > 0, vk ∈ Sn−1 , bk ∈ R} as parameters. Here, d̄ (y, Hek ,e ) is the hyperbolic signed distance from y to the hyperplane passing through m the output-space origin e ∈ HK . Next, we show that the above implicit definition has an explicit solution for the output y. Theorem 135 (Poincaré BFC). [↓] Given an input x ∈ PnK , the Poincaré BFC layer F : PnK → Pm K is given by y= 1+
q
ω
,
ω=
2
1 − K ∥ω∥
"
sinh
√
−Kuk (x) √ −K
#m
,
(5.60)
k=1
where uk (x) = −αk B vk (x) + bk with {αk > 0, vk ∈ Sn−1 , bk ∈ R} as parameters for k = 1, . . . , m. Theorem 136 (Lorentz BFC). [↓] Given an input x ∈ LnK , the Lorentz BFC layer F : LnK → Lm K is given by q " # 2 1 + ∥y ∥ s yt −K y= = , √ 1 ys √ sinh −Ku(x)
(5.61)
−K
where u(x) = (u1 (x), . . . , um (x))⊤ with uk (x) = −αk B vk (x) + bk . Here, {αk > 0, vk ∈ Sn−1 , bk ∈ R} are parameters for k = 1, . . . , m. Analogously to Thm. 130, our BFC layers converge to their Euclidean counterparts as K → 0− . Theorem 137 (Limits of BFC layers). [↓] As K → 0− , the hyperbolic BFC layer n m HK ∋ x 7→ y ∈ HK reduces to a Euclidean FC layer: 1 K→0− Poincaré: yk −−−−→ αk ⟨vk , x⟩ + bk , 2 K→0−
Lorentz: (ys )k −−−−→ αk ⟨vk , xs ⟩ + bk .
157
(5.62) (5.63)
5.3. Hyperbolic Busemann Neural Networks Method Möbius [76, Eq. (27)] Poincaré FC [181, Eq. (7)]
Lorentz FC [45, Eq. (3)]
BFC
n m F : HK ∋ x 7→ y ∈ HK √ W x ∥W x∥ −1 tanh −K ∥x∥ ∥x∥ ∥W x∥ √ sinh −Kuk (x) ω q √ , ωk = , y= 2 −K 1 + 1 − K ∥ω∥ with uk (x) in Tab. 5.11 "q # ∥ψ(W x, v)∥2 − 1/K , y= ψ(W x, v) W ϕ(x) + b ψ(W x, v) = λσ v ⊤ x + b′ ∥W ϕ(x) + b∥ √ sinh −Ku(x) ω q √ ,ω= y= ; −K 1 + 1 − K ∥ω∥2 r √ 1 1 −Ku(x) , yt = + ∥ys ∥2 , ys = √ sinh −K −K with uk (x) = ϕ (−αk B vk (x) + bk )
1 √ tanh −K
Space
Methodology
PnK
Tangent
PnK
Poincaré geometry
LnK
PnK LnK
Parameters
#Params
FLOPs
mn
2nm + 2n +2m + 24
αk > 0, vk ∈ Sn−1 , bk ∈ R, k = 1, . . . , m
m(n + 2)
4nm + 71m + 4
Ambient Minkowski
W ∈ Rm×(n+1) , v ∈ Rn+1 , b ∈ Rm , b′ ∈ R, λ > 0
m(n + 1) + m +(n + 1) + 2
2nm + 8m +2n + 10
Busemann
αk > 0, vk ∈ Sn−1 , bk ∈ R, k = 1, . . . , m
m(n + 2)
6nm + 29m + 4 2nm + 30m + 2
W ∈R
m×n
Table 5.12: Comparison of hyperbolic FC layers. For simplicity, BFC layers do not involve the gyroaddition and assume ϕ is the identity map, which is in line with the Möbius and Lorentz FC layers. 5.3.4.2
Generalization
Following Sec. 5.2.4.2, we extend the hyperbolic BFC by inserting the activation into Eq. (5.59): d̄ (y, Hek ,e ) = ϕ (uk (x)) , ∀k ∈ {1, . . . , m}. (5.64) This is reflected in Thms. 135 and 136 by replacing every uk (x) with ϕ (−αk B vk (x) + bk ). Moreover, inspired by the Poincaré Möbius transformation [76, Sec. 3.2], a BFC transn m formation could be further followed by a gyroaddition ⊕H : HK ∋ x 7→ F(x) ⊕H b ∈ HK m as a gyro bias. For example, the Lorentz BFC layer is generalized as with b ∈ HK
LnK ∋ x 7→ y =
q
1 + ∥ys ∥2 −K
√ 1 sinh −K
√
m ⊕L b ∈ LK , −Ku(x)
(5.65)
where uk (x) = ϕ (−αk B vk (x) + bk ) with parameters {αk > 0, vk ∈ Sn−1 , bk ∈ R}m k=1 and b ∈ Lm K. 5.3.4.3
Comparison with Existing Hyperbolic FC Layers
Tab. 5.12 compares BFC with prior hyperbolic FC layers. BFC faithfully respects hyperbolic geometry, whereas the Möbius and Lorentz FC layers apply Euclidean transformations in the tangent or ambient Minkowski space, which can distort intrinsic geometry. BFC also offers flexibility across models, while Poincaré FC and Lorentz FC are tailored to their respective models. In addition, BFC uses a comparable parameterization and maintains O(nm) FLOPs. On LnK , its FLOPs are O(2mn), matching the fastest layers. 158
Chapter 5. Riemannian Neural Networks
Space
Method
Rn
CIFAR-10 (Num. classes: 10)
CIFAR-100 (Num. classes: 100)
Tiny-ImageNet (Num. classes: 200)
ImageNet-1k (Num. classes: 1000)
Acc
Fit Time
#Params
Acc
Fit Time
#Params
Acc
Fit Time
#Params
Acc
Fit Time
#Params
MLR
95.14 ± 0.12
10.66
5.13K
77.72 ± 0.15
10.60
51.30K
65.19 ± 0.12
69.17
102.60K
71.87
2263.12
513K
PnK
PMLR PBMLR-P BMLR-P
95.04 ± 0.13 95.23 ± 0.08 95.32 ± 0.14
11.94 21.92 12.01
5.14K 10.24K 5.14K
77.19 ± 0.50 77.78 ± 0.15 78.10 ± 0.35
12.11 76.84 12.13
51.40K 102.40K 51.40K
64.93 ± 0.38 65.43 ± 0.27 66.16 ± 0.19
71.90 336.58 71.98
102.80K 204.80K 102.80K
71.77 71.46 73.36
2300.11 3907.12 2300.77
514K 1024K 514K
LnK
LMLR BMLR-L
94.98 ± 0.12 95.25 ± 0.02
11.55 11.08
5.13K 5.14K
78.03 ± 0.21 78.07 ± 0.26
11.72 11.22
51.30K 51.40K
65.63 ± 0.10 65.99 ± 0.14
69.27 69.19
102.60K 102.80K
72.46 73.24
2277.17 2276.53
513K 514K
Table 5.13: Top-1 image classification accuracy (%) of MLR methods on the ResNet-18 backbone. The best results within each hyperbolic model are bold. The slowest MLR and largest parameter count are shown in red.
Poincaré
Lorentz 70 Top-1 Accuracy
Top-1 Accuracy
70 60 50 40
0
20
MLR PMLR PBMLR-P BMLR-P 40 60 80 100 Epoch
60 50 40
0
20
MLR LMLR BMLR-L 40 60 80 100 Epoch
Figure 5.2: Validation accuracy curves on ImageNet-1k.
5.3.5
Experiments
We first compare BMLRs with prior hyperbolic MLRs on three architectures: ResNet-18 (image classification), CNN (genome sequences), and HGCN (node classification). We then compare BFC with prior hyperbolic FC layers on link prediction. All experiments use both the Poincaré and Lorentz models. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.5. 5.3.5.1
Image Classification
Setup. Following Sec. 5.2.6.2, we use a hybrid architecture with a ResNet-18 [97] backbone and an MLR head. We compare Euclidean MLR with hyperbolic variants in both models. In Poincaré, we evaluate Poincaré MLR (PMLR) re-parameterized by Shimizu et al. [181], Pseudo-Busemann MLR (PBMLR-P) [162], and our BMLR-P. In Lorentz, we evaluate Lorentz MLR (LMLR) [15] and our BMLR-L. For hyperbolic MLRs, we 159
5.3. Hyperbolic Busemann Neural Networks Benchmark
TEB
GUE
Task
Data Set
Num.
PnK
LnK
classes
PMLR
PBMLR-P
BMLR-P
LMLR
BMLR-L
Retrotransposons
LTR Copia LINEs SINEs
2 2 2
75.34 ± 1.02 85.54 ± 0.61 95.30 ± 0.85
74.37 ± 1.48 85.92 ± 0.65 95.34 ± 1.58
76.73 ± 1.08 86.05 ± 1.08 95.99 ± 0.74
73.01 ± 1.07 83.14 ± 0.80 96.70 ± 0.87
75.86 ± 1.52 86.72 ± 0.58 96.29 ± 0.59
DNA transposons
CMC-EnSpm hAT-Ac
2 2
83.39 ± 0.56 89.38 ± 0.90
83.62 ± 1.00 89.86 ± 0.54
84.03 ± 0.71 89.62 ± 0.74
81.78 ± 1.05 88.94 ± 0.69
84.15 ± 1.00 90.70 ± 0.51
Pseudogenes
processed unprocessed
2 2
72.45 ± 1.49 75.37 ± 2.27
71.99 ± 2.04 71.99 ± 1.47
73.09 ± 1.66 75.71 ± 1.89
73.71 ± 1.76 74.54 ± 1.98
73.32 ± 1.65 76.15 ± 1.61
Core Promoter Detection
tata notata all
2 2 2
80.95 ± 1.47 70.02 ± 0.52 67.64 ± 0.77
79.32 ± 2.44 70.60 ± 0.75 68.02 ± 0.63
80.29 ± 1.63 70.48 ± 0.35 68.50 ± 0.61
80.90 ± 1.15 71.26 ± 0.56 67.63 ± 0.56
81.76 ± 1.16 70.43 ± 0.39 68.36 ± 1.07
Promoter Detection
tata notata all
2 2 2
80.30 ± 1.59 92.63 ± 0.36 90.53 ± 0.50
80.27 ± 2.71 93.05 ± 0.32 90.79 ± 0.77
82.83 ± 1.69 92.75 ± 0.51 90.20 ± 0.65
83.27 ± 1.95 91.74 ± 0.57 89.34 ± 0.40
82.55 ± 1.54 92.60 ± 0.49 89.82 ± 0.45
Covid Variant Classification
Covid
9
74.09 ± 0.25
70.84 ± 0.80
73.40 ± 0.30
64.07 ± 0.51
72.45 ± 0.21
Species Classification
Virus Fungi
20 25
67.24 ± 2.10 15.06 ± 1.32
59.17 ± 3.32 18.75 ± 1.77
77.12 ± 1.23 30.01 ± 0.76
71.34 ± 2.05 15.07 ± 1.88
77.21 ± 1.04 30.14 ± 2.48
Table 5.14: Genomic MCC of MLR methods under the CNN backbone. The best results within each hyperbolic model are bold. map the ResNet-18 features to the target hyperbolic space before classification. We evaluate on CIFAR-10 [126], CIFAR-100 [126], Tiny-ImageNet [129], and ImageNet-1k [66]. On the first three data sets, we conduct five-fold experiments. Results. Tab. 5.13 reports top-1 validation accuracy, fit time per epoch, and classifier-head parameters. Fig. 5.2 presents the ImageNet-1k accuracy curves. Overall, BMLR-P and BMLR-L consistently outperform prior hyperbolic MLRs with comparable parameters. Within each hyperbolic model, the accuracy margin over prior hyperbolic MLRs increases with the number of classes, from CIFAR-10 to CIFAR-100 and Tiny-ImageNet, with the largest gains on ImageNet-1k. This demonstrates the advantage of BMLR as task complexity increases. Besides, PBMLR-P uses approximately double the head parameters and is markedly slower due to complex batch-inefficient computation, whereas BMLR-L achieves the fastest fit time among all hyperbolic MLRs. 5.3.5.2
Genome Sequence Learning
Setup. Similar to Sec. 5.2.6.4, we evaluate hyperbolic MLRs on genome sequence learning. Following Khan et al. [119], we adopt a CNN backbone, which consists of three convolutional blocks and an MLR head. Similar to Sec. 5.3.5.1, we compare our BMLR against previous hyperbolic MLR heads by replacing the final Euclidean MLR with a hyperbolic MLR. We validate on two benchmarks: TEB [119] and Genome Understanding Evaluation (GUE) [234], covering a total of 16 data sets. Results. Tab. 5.14 summarizes five-fold average MCC across TEB and GUE. Compared with other hyperbolic MLRs, our BMLR-P and BMLR-L achieve higher MCC in most tasks. Similar to Sec. 5.3.5.1, the gains are more pronounced on complex data sets 160
Chapter 5. Riemannian Neural Networks
Data Set
PnK
LnK
PMLR
PBMLR-P
BMLR-P
LMLR
BMLR-L
LTR Copia LINEs SINEs
3.96 5.80 1.36
5.11 6.90 1.80
4.06 5.73 1.39
3.92 5.50 1.38
3.77 5.32 1.28
CMC-EnSpm hAT-Ac
3.11 4.37
4.38 5.37
3.04 4.38
2.89 4.13
2.82 3.96
processed unprocessed
5.02 3.30
5.94 3.90
4.89 3.29
4.68 3.05
4.58 2.95
CPD-tata CPD-notata CPD-all
0.72 6.13 6.85
1.35 12.40 14.35
0.70 6.00 6.59
0.67 5.73 6.35
0.62 5.71 6.33
PD-tata PD-notata PD-all
0.98 8.33 9.40
1.20 10.53 12.07
0.97 8.29 9.19
1.04 8.11 9.06
0.94 7.84 8.83
Covid
28.96
45.52
27.97
27.58
26.67
Virus Fungi
25.28 6.41
29.57 8.96
25.67 6.41
25.12 6.25
24.85 6.25
Table 5.15: Fit time (s/epoch) on genome sequence learning. The fastest times are bold and the slowest ones are red. Space
Method
Methodology
Disease δ=0
Airport δ=1
PubMed δ = 3.5
Cora δ = 11
PnK
Möbius Poincaré FC BFC-P
Tangent Poincaré geometry Busemann
76.35 ± 1.83 79.45 ± 1.01 80.45 ± 0.93
93.31 ± 0.41 94.31 ± 0.16 94.88 ± 0.39
94.93 ± 0.06 94.24 ± 0.25 94.85 ± 0.07
90.80 ± 0.56 88.21 ± 0.72 91.94 ± 0.32
LnK
LTFC Lorentz FC BFC-L
Tangent Ambient Minkowski Busemann
71.32 ± 5.36 72.78 ± 2.04 78.36 ± 0.51
92.68 ± 0.35 92.99 ± 0.33 95.37 ± 0.17
94.85 ± 0.17 94.20 ± 0.10 94.90 ± 0.04
89.37 ± 0.64 92.06 ± 0.62 92.28 ± 0.12
Table 5.16: Comparison of hyperbolic FC layers on link prediction. The best results within each hyperbolic model are bold. with more classes, e.g.,, Virus (20 classes) and Fungi (25 classes), demonstrating the effectiveness of our approach. Tab. 5.15 reports fit time per epoch, where PBMLR-P is consistently the slowest due to batch inefficiency, and BMLR-L is the fastest. 5.3.5.3
Node Classification
Setup. Following Nguyen et al. [162], we adopt the HGCN [38] backbone to evaluate our BMLR on graph data sets, including Disease [4], Airport [229], PubMed [155], and Cora [177]. The HGCN backbone consists of a hyperbolic Graph Convolutional Network (GCN) and an MLR as the final classification layer. Both the GCN and the MLR are built on the hyperbolic space. The vanilla HGCN uses a tangent MLR, which maps 161
5.3. Hyperbolic Busemann Neural Networks
Space
Method
Disease δ=0
Airport δ=1
PubMed δ = 3.5
Cora δ = 11
PnK
HGCN HGCN-PMLR HGCN-PBMLR-P HGCN-BMLR-P
86.87 ± 2.58 88.98 ± 1.96 89.05 ± 0.78 92.45 ± 0.96
85.34 ± 1.16 84.78 ± 1.48 85.04 ± 0.97 86.02 ± 0.53
76.29 ± 0.98 76.02 ± 1.09 75.89 ± 0.78 77.36 ± 0.73
76.56 ± 0.81 77.47 ± 1.15 77.90 ± 1.00 78.48 ± 1.52
LnK
HGCN HGCN-LMLR HGCN-BMLR-L
87.83 ± 0.77 89.72 ± 1.51 90.80 ± 1.15
84.94 ± 1.40 82.61 ± 1.01 85.27 ± 1.17
76.49 ± 0.88 75.44 ± 1.17 77.30 ± 0.41
77.37 ± 1.72 69.91 ± 3.61 77.65 ± 2.10
Table 5.17: Node classification F1 scores of hyperbolic MLRs on the HGCN backbone, where δ denotes graph hyperbolicity (lower is more hyperbolic). The best results within each hyperbolic model are bold. features into the tangent space via Loge and applies a Euclidean MLR. We replace this with different hyperbolic MLRs. Results. Tab. 5.17 reports average F1 scores. Our BMLRs consistently outperform prior hyperbolic MLRs within each hyperbolic model. As graphs become less hyperbolic, that is, for larger δ, existing hyperbolic heads could underperform the vanilla tangentbased MLR, for example, PBMLR-P on PubMed, and LMLR on Airport, PubMed, and Cora. Especially on Cora, which has the largest δ, LMLR lags the tangent baseline by a large margin (69.91 vs. 77.37). In contrast, BMLR remains the top performer across all δ values, indicating that Busemann-based decoding robustly strengthens HGCN over a broader range of graph hyperbolicity. 5.3.5.4
Link Prediction
Setup. We compare our BFC layers with prior hyperbolic FC layers, including the Möbius layer [76] that operates via the tangent space, the Lorentz FC layer [45] that operates through the ambient Minkowski space, and the Poincaré FC layer [181]. Mimicking the Möbius layer, we also implement a Lorentz tangent FC layer, Exp0 (M Log0 (x)), referred to as LTFC. Following Chami et al. [38], we evaluate on Disease, Airport, PubMed, and Cora. Following the HNN implementation [76, 38], all methods share the same backbone with two FC layers. For a fair comparison, all hyperbolic FC layers are followed by a gyroaddition biasing. For BFC, we use ϕ = tanh on Airport and Cora and the identity map on the other two data sets. Results. Tab. 5.16 reports five-fold test AUC. Our BFC layers generally outperform prior hyperbolic FC layers. The gains are most pronounced on Disease, which is the most hyperbolic (δ = 0), where Busemann-based decoding is markedly more effective than tangent or ambient methods, indicating better capture of intrinsic hyperbolic 162
Chapter 5. Riemannian Neural Networks Space
Method
Disease
Airport
PubMed
Cora
Fit Time
#Params
Fit Time
#Params
Fit Time
#Params
Fit Time
#Params
PnK
Möbius Poincaré FC BFC-P
0.0200 0.0198 0.0201
464 528 528
0.0535 0.0536 0.0512
480 544 544
0.1120 0.1176 0.1123
8288 8352 8352
0.0229 0.0248 0.0231
23216 23280 23280
LnK
LTFC Lorentz FC BFC-L
0.0343 0.0232 0.0244
464 563 528
0.0818 0.0715 0.0713
480 580 544
0.1633 0.1537 0.1525
8288 8876 8352
0.0370 0.0261 0.0280
23216 24737 23280
Table 5.18: Efficiency comparison: fit time (s/epoch) and parameter count. Slowest results and largest parameter counts are in red. geometry. This observation aligns with geometric intuition, since tangent space or ambient space approximations inherently struggle to represent curved manifolds in highly non-Euclidean cases. Training Time and Parameter Count. Tab. 5.18 summarizes fit time per epoch and parameter counts. Our BFC layers achieve training time and model size comparable to existing layers. In particular, LTFC is the slowest due to costly logarithmic and exponential maps, and LFC uses the largest number of parameters among Lorentz variants.
5.4
Full-Rank Correlation Networks
5.4.1
Introduction
The preceding two sections studied manifold-specific designs for hyperbolic learning. We now turn to neural networks on full-rank correlation manifolds. Covariance matrices in the SPD manifold have achieved success in various applications, with many deep network architectures adapted to leverage their Riemannian geometries [106, 31, 37, 59, 167, 123, 206, 47, 118, 134, 172, 215, 115, 104]. In contrast, correlation matrices, despite serving as statistically compact alternatives to covariance matrices [7], remain unexplored in deep learning. As discussed in Sec. 2.9.2, Riemannian structures for correlation matrices have only recently been developed. David and Gu [62] identified full-rank correlation matrices as a quotient manifold of the SPD manifold, referred to as the correlation manifold. However, this quotient geometry does not guarantee uniqueness or closed forms of the Riemannian logarithm and Fréchet mean [195, Sec. 1.1]. To close this gap, Thanwerdas and Pennec [195] proposed three theoretically and computationally convenient geometries: ECM, LECM, and PHCM. Thanwerdas [191] further introduced two efficient permutation-invariant metrics: OLM and LSM. These Riemannian structures provide 163
5.4. Full-Rank Correlation Networks promising foundations for extending Euclidean deep learning to the correlation manifold. On the other hand, several fundamental layers in Euclidean deep learning, such as MLR, FC, and convolutional layers, have been extended to different manifolds by leveraging their rich Riemannian or algebraic structures [106, 107, 108, 76, 37, 45, 181, 15, 51, 161]. For the SPD manifold, these layers have been constructed using bilinear mapping [106], weighted Fréchet means [37], gyrovector spaces [159, 161], and Riemannian geometry [49, 51]. Inspired by these advancements, we develop MLR, FC, and convolutional layers for correlation manifolds in a geometrically intrinsic manner. We begin by systematically introducing four types of correlation-based MLR, FC, and convolutional layers, corresponding to ECM, LECM, OLM, and LSM, respectively. Besides, we discuss backpropagation through Riemannian computations over the correlation manifold, with novel approaches for accurate backpropagation under OLM and LSM. As the above four metrics have zero curvature, our next focus is to build correlation layers under the geometry of non-zero curvature. We target PHCM, induced by the product of multiple hyperbolic spaces [195, Thm. 4.4]. By adapting existing Poincaré-based hyperbolic MLR, FC, and convolutional layers designed for a single Poincaré ball [76, 181], we construct their counterparts on the correlation manifold. Together with the corresponding backpropagation mechanisms, these layers constitute complete Correlation Networks (CorNets) under different geometries. The effectiveness is validated by experiments comparing our approach against existing SPD and Grassmannian baselines. Tab. 5.19 summarizes the correspondence between Euclidean and our correlation layers. In summary, our main contributions are as follows: (1) We systematically extend MLR, FC, and convolutional layers to the correlation manifold under five geometries: four with zero curvature and one with non-zero curvature. The developed layers enable flexible variation of the latent geometry under a consistent network architecture, allowing for straightforward comparisons across different correlation geometries. (2) We develop accurate backpropagation of Riemannian computations under OLM and LSM. (3) We conduct experiments against existing SPD and Grassmannian networks to demonstrate the effectiveness of correlation embeddings and networks. Outline. Sec. 5.4.2 constructs correlation MLR, FC, and convolutional layers under four flat geometries. Sec. 5.4.3 develops their counterparts under a non-zero-curvature 164
Chapter 5. Riemannian Neural Networks Space
Euclidean Rn
Correlation Cor+ (n)
C-class MLR FC layer Convolution Geometry
f : Rn ∋ x 7→ p = softmax(Ax + b) ∈ RC F : Rn ∋ x 7→ y = Ax + b ∈ Rm Kernel-based FC in each receptive field Euclidean
f : Cor+ (n) ∋ X 7→ p ∈ RC F : Cor+ (n) ∋ X 7→ Y ∈ Cor+ (m) Kernel-based correlation FC in each receptive field ECM, LECM, OLM, LSM and PHCM
Table 5.19: Correspondence between Euclidean and correlation-based layers. For convolution, kernel-based FC refers to applying a convolution kernel to a receptive field, which is an FC transformation. geometry and establishes the order-invariance of the associated β-operations. Sec. 5.4.4 presents backpropagation over the correlation geometries, and Sec. 5.4.5 evaluates CorNets under these five geometries. Proofs are deferred to Sec. B.8.
5.4.2
Log-Euclidean Correlation Layers
Since ECM, LECM, OLM, and LSM are derived via diffeomorphisms from Euclidean spaces, they are collectively termed Log-Euclidean metrics [191]. This motivates the principled development of MLR, FC, and convolutional layers [191]. 5.4.2.1
Log-Euclidean Correlation MLRs
As discussed in Sec. 4.3.2.1, the MLR can be rewritten by point-to-hyperplane formulations. We use the same margin-distance infimum, while the Euclidean isometries of ECM, LECM, OLM, and LSM make it possible to solve this infimum exactly under all four geometries. To avoid over-parameterization, we follow Sec. 5.2.4.1 and set Pk = ExpE (γk [Zk ]) and Ak = PTE→Pk (Zk ), with [Zk ] = ∥ZZkk∥ as the unit direction E vector of Zk . Here, E is the origin of M, while γk ∈ R and Zk ∈ TE M ∼ = Rm are the MLR parameters. This is a concrete instance of the trivialization strategy reviewed in Sec. 2.7. Under this trivialization, each hyperplane HAk ,Pk is denoted as HZk ,γk . As all Log-Euclidean metrics are isometric to Euclidean spaces, the corresponding MLRs admit principled closed forms. Theorem 138. [↓] Let M, g M be an m-dimensional manifold that is isometric to the standard Euclidean space Rm via the diffeomorphism ϕ : M → Rm . Denoting E = ϕ−1 (0) with 0 as the zero vector, each vk (X) and margin hyperplane HZk ,γk in the C-class Riemannian MLR are vk (X) = ⟨ϕ(X), ϕ∗,E (Zk )⟩ − γk ∥ϕ∗,E (Zk )∥ and HZk ,γk = {X ∈ M | vk (X) = 0}, respectively. Here, Zk ∈ TE M ∼ = Rm and γk ∈ R for 1 ≤ k ≤ C are MLR parameters, while ϕ∗ is the differential. 165
5.4. Full-Rank Correlation Networks Simple computations show that ECM: ϕEC (I) = 0, LECM: log ◦Θ(I) = 0,
OLM: Log◦ (I) = 0, LSM: Log⋆ (I) = 0.
(5.66)
Therefore, we define the origin of the correlation manifold under four Log-Euclidean metrics as the identity matrix. Besides, Thm. 138 suggests that Log-Euclidean MLRs can be obtained modulo the calculation of diffeomorphisms and their differentials at the identity matrix I. Proposition 139 (Differentials). [↓] For any tangent vector V ∈ TI Cor+ (n) ∼ = ◦ ⋆ EC Hol(n), the differentials of ϕ , log ◦Θ, Log , and Log at the identity matrix I are ϕEC (log ◦Θ)∗,I (V ) = ⌊V ⌋, ∗,I (V ) = ⌊V ⌋, (5.67) Log◦∗,I (V ) = V, Log⋆∗,I (V ) = V − diag(V 1), where diag : Rn → Diag(n) returns a diagonal matrix, and 1 = (1, · · · , 1)⊤ ∈ Rn . Putting Thm. 139 into Thm. 138, we obtain correlation MLRs under four LogEuclidean metrics. Theorem 140 (Log-Euclidean MLRs). Given C ∈ Cor+ (n), the logits vk (C) for the k-th class in the correlation MLRs under four Log-Euclidean metrics are vkEC (C) = ⟨⌊Θ(C)⌋, ⌊Zk ⌋⟩ − γk ∥⌊Zk ⌋∥ ,
vkLEC (C) = ⟨log ◦Θ(C), ⌊Zk ⌋⟩ − γk ∥⌊Zk ⌋∥ , vkOL (C) = ⟨Log◦ (C), Zk ⟩ − γk ∥Zk ∥ ,
(5.68)
vkLS (C) = Log⋆ (C), Log⋆∗,I (Zk ) − γk Log⋆∗,I (Zk ) ,
where Zk ∈ Hol(n) and γk ∈ R are parameters.
5.4.2.2
Log-Euclidean FC and Convolutional Layers
Following Sec. 5.2.4.2, we now generalize this point-to-hyperplane FC construction to the correlation manifold. Definition 141 (Correlation FC layers). Given a metric g, the correlation FC layer F : Cor+ (n) ∋ X 7→ Y ∈ Cor+ (m) returns the output Y by solving the following 166
Chapter 5. Riemannian Neural Networks d = m(m−1)/2 equations: sk d(Y, HOk ,I ) = vk (X; Zk , γk ),
1 ≤ k ≤ d,
(5.69)
where sk = sign (⟨LogI (Y ), Ok ⟩I ), I is the identity matrix, d is the dimension of Cor+ (m), {Ok }dk=1 is an orthonormal basis over TI Cor+ (m), d(·, ·) is the margin distance to the hyperplane HOk ,I , and vk is defined by Sec. 4.3.2.1 for Cor+ (n). The FC parameters are {Zk ∈ Hol(n)}dk=1 and {γk ∈ R}dk=1 . Sec. A.4.3.1 details how Thm. 141 extends the existing SPD, Poincaré, and Euclidean FC layers. Although Thm. 141 is implicitly defined by d equations, the FC layers under four Log-Euclidean geometries admit explicit expressions in a principled manner. Analogous to Thm. 138, a corresponding result for the FC layer is presented in Thm. 201, which yields the Log-Euclidean FC layers. Theorem 142 (Log-Euclidean FC layers). [↓] Given an input correlation C ∈ Cor+ (n), the correlation FC layers F(·) : Cor+ (n) → Cor+ (m) under different Log-Euclidean metrics are ECM: Y = Cor ◦ Chol−1 V EC + Im , LECM: Y = Cor ◦ Chol−1 ◦ exp V LEC , OLM: Y = Exp◦ V OL , LSM: Y = Cor ◦ exp V LS ,
(5.70) (5.71) (5.72) (5.73)
where the (i, j)-th elements in V EC ∈ LT0 (m), V LEC ∈ LT0 (m), V OL ∈ Hol(m), and V LS ∈ Row0 (m) are VijEC =
v EC (C), ij
if i > j
0, otherwise v LEC (C), if i > j ij VijLEC = 0, otherwise OL vij (C) √ , if i > j 2 OL Vij = VjiOL , if i < j 0, otherwise 167
(5.74)
(5.75)
(5.76)
5.4. Full-Rank Correlation Networks
1
+
C ∈ Cor (n) C 2 ∈ Cor+ (n) C 3 ∈ Cor+ (n)
Input 3-Channel Correlation Matrices
Split
! "2 {C 1 , C 2 } ∈ Cor+ (n)
! "2 {C 2 , C 3 } ∈ Cor+ (n)
F1 F2
F1 F2
C! 1 ∈ Cor+ (m) C! 2 ∈ Cor+ (m) C! 3 ∈ Cor+ (m) C! 4 ∈ Cor+ (m)
FC Transformation on Each Receptive Field
Figure 5.3: Illustration of the Log-Euclidean 1D convolution with two kernels. The 3-channel input is first split into two receptive fields along the channel dimension. In each receptive field, two kernels are applied to the product space.
LS (C) √ vij / 6, LS (C) √ vii / 3, VijLS = VjiLS , − Pm−1 V LS , kj k=1 P P m−1 m−1 V LS , lk l=1 k=1
if m > i > j ≥ 1 if m > i ≥ 1 if i < j
(5.77)
if i = m, 1 ≤ j < m if i = j = m
Each vijg with g ∈ {EC, LEC, OL, LS} is defined by Eq. (5.68) with parameters Zij ∈ Hol(n) and γij ∈ R. For vijEC , vijLEC , and vijOL , the indices satisfy i, j = 1, . . . , m and i > j. For vijLS , they satisfy i, j = 1, . . . , m − 1 and i ≥ j. Correlation Convolution. Following the convolution-as-FC construction in Sec. 5.2.4.3, we develop the correlation convolution. The c-channel correlation matrices {Ci ∈ Cor+ (n)}ci=1 within a receptive field are first concatenated into C ∈ (Cor+ (n))c . For each convolution kernel, C is then fed into a correlation FC layer.5 Fig. 5.3 illustrates the above process.
5
Thm. 142 naturally supports product geometries, which are detailed in Sec. A.4.3.2.
168
Chapter 5. Riemannian Neural Networks
5.4.3
Poly-Hyperbolic-Cholesky Layers
As detailed in Sec. 2.9.2, the space Ln , consisting of the Cholesky factors of Cor+ (n), can be identified with the product of n − 1 hyperbolic open hemispheres, PHSn−1 = Qn−1 i i=1 HS . We focus on the widely used hyperbolic Poincaré ball, whose MLR, FC, and β-concatenation components are reviewed in Sec. A.2.4. In the following, we focus on the canonical Poincaré ball (K = −1), namely the unit Poincaré ball Pn . We first Q i identify the correlation manifold with the poly-Poincaré space PPn−1 = n−1 i=1 P , the product of n−1 unit Poincaré balls. Then, we develop correlation layers from the layers on a single Poincaré space. 5.4.3.1
Correlation Geometry via Poincaré Balls
Proposition 143 (Isometries). [↓] The open hemisphere HSn is isometric to the unit Poincaré ball Pn by ψHSn →Pn ((x⊤ , xn+1 )⊤ ) =
x , 1 + xn+1
1 ψPn →HSn (y) = 1 + ∥y∥2
2y 1 − ∥y∥2
!
(5.78) ,
with (x⊤ , xn+1 )⊤ ∈ HSn ⊂ Rn × R+ and y ∈ Pn ⊂ Rn . Thm. 143 indicates that Cor+ (n) can be identified with PPn−1 = diffeomorphism Φ:
1 0 Chol L21 L22 C 7−→ .. .. . . Ln1 Ln2
··· ··· .. .
0 0 Qn−1 i=1 ψi .. 7−→ .
· · · Lnn
ψ1 (h1 ) .. .
Qn−1
i=1 P
i
via the
(5.79)
ψn−1 (hn−1 )
with C ∈ Cor+ (n), hi = (Li+1,1 , · · · , Li+1,i+1 )⊤ ∈ HSi , and ψi = ψHSi →Pi . This identification motivates us to construct the correlation layers using the corresponding layers over Poincaré spaces. 5.4.3.2
Revisiting Poincaré Layers
The Poincaré MLR and FC layers are reviewed in Sec. A.2.4 and follow the pointto-hyperplane logic discussed in Chapter 4 and Secs. 5.2 and 5.3. The convolutional 169
5.4. Full-Rank Correlation Networks Input 𝑐-Channel Correlation Matrices
Cholesky Factors
1 L121 . ..
L1n1
C 1 ∈ Cor+ (n)
··· ··· .. .
0 0 .. .
L1n2
···
L1nn
Poly-Poincaré Vectors xi !" #! $ L121 , L122 ∈ P1
Ψn−1
1 Lc21 .. . Lcn1
0 Lc22 .. .
··· ··· .. .
0 0 .. .
Lcn2
···
Lcnn
Lc = Chol(C c ) ∈ LT+ (n)
!"
.. .
L1n1 , · · · , L1nn
x1 ∈ PPn−1 =
!n−1
#! $
∈ Pn−1
i i=1 P ⊂ R
x ! ∈ PM M = m(m−1) 2
n(n−1) 2
1
$ Ψ−1 L21
β-Split
.. .
$ m1 L
0 $ 22 L .. . $ m2 L
··· ··· .. . ···
0 0 .. . $ mm L
C! ∈ Cor+ (m)
! ∈ LT+ (m) L
FC Transformation
β-Concate
…
Φ
…
…
Chol
Poincaré FC
Ψ1
L1 = Chol(C 1 ) ∈ LT+ (n)
C c ∈ Cor+ (n)
0 L122 .. .
" ! ! Ψ1 (Lc21 , Lc22 ) ∈ P1
... " ! ! Ψn−1 (Lcn1 , · · · , Lcnn ) ∈ Pn−1
x ∈ PP c
n−1
=
!n−1
i=1 P
i
⊂R
n(n−1) 2
Poincaré MLR x ∈ PN N = c n(n−1) 2
Identifying the Correlation Manifold with the Poly-Poincaré Space
Classification
MLR Classification
Figure 5.4: Illustration of the PHCM convolution and MLR. The multi-channel input correlation matrices are denoted as {C i }ci=1 . For the convolutional layer, the illustration focuses on the transformation within a receptive field and assumes a single-channel output. construction uses the Poincaré β-concatenation defined below. The Poincaré convolutional layer shares a logic similar to the correlation convolution, except it uses β-concatenation to concatenate the Poincaré vectors in each receptive field [181, Secs. 3.3–3.4], which can stabilize the norm of the Poincaré vector. The Poincaré β-concatenation generalizes the Euclidean concatenation via the scaled concatenation in the tangent space. Given inputs {xi ∈ Pni }N i=1 , it is defined as PN ⊤ n −1 ⊤ ⊤ ∈ P , where v = Log (x ) and n = v , · · · , β Exp0 βn βn−1 v i i 0 n 1 N i=1 ni . Here, 1 N α 1 βni and βn are defined by the beta function βα = B ( /2, /2). The inverse is called the Poincaré β-split. The Poincaré convolution is: (1) β-concatenating the multi-channel feature in a given receptive field; and (2) performing the Poincaré FC transformation. 5.4.3.3
Building Poly-Hyperbolic-Cholesky Layers
PHCM MLR. The input multi-channel correlation matrices, C = {C i ∈ Cor+ (n)}ci=1 , are first mapped into poly-Poincaré spaces as x = {xi = Φ(C i ) ∈ PPn−1 }ci=1 . The resulting Poincaré vectors are then β-concatenated into a single Poincaré vector x ∈ PN , where N = c n(n−1) . This concatenated vector is subsequently fed into the Poincaré MLR 2 for classification. PHCM Convolutional and FC Layer. The convolutional layer follows a logic similar to Log-Euclidean convolution. The multi-channel correlation matrices within a receptive field C = {C i ∈ Cor+ (n)}ci=1 are first mapped to a β-concatenated Poincaré vector x ∈ PN as in the PHCM MLR, which is then fed into the Poincaré FC layer for dimensionality transformation. This produces a vector x e ∈ PM , with M = k m(m−1) , 2 which is then split using β-split. Subsequently, applying Φ−1 reconstructs new k×m×m 170
Chapter 5. Riemannian Neural Networks correlation matrices. When the input is a single correlation matrix, it is reduced to the correlation FC. Fig. 5.4 illustrates the PHCM layers. However, there is an underlying ambiguity in the above discussion. To clarify, we write each xi ∈ PPn−1 in x as xi = {pi1 ∈ P1 , · · · , pin−1 ∈ Pn−1 }, which gives x = {pij ∈ Pj }i=c,j=n−1 i=1,j=1 . We can either concatenate twice by i → j or once along both i and j. A similar issue arises with β-split. The following theorem establishes this invariance. Theorem 144 (Order-invariance). [↓] Given multichannel data xi1 ,...,in ∈ Pnin with ij ∈ {1, . . . , Nj }, applying the β-concatenation sequentially n times in the order in → · · · → i1 is equivalent to a single β-concatenation along all indices simultaneously. Similarly, β-splittingx ∈ PN into multichannel data xi1 ,...,in ∈ Pnin Qn−1 PNn with ij ∈ {1, . . . , Nj } and N = j=1 Nj in =1 nin under the sequential order i1 → · · · → in is identical to the one under a single β-split to generate all indices simultaneously. Therefore, we always conduct the β-operation simultaneously along both i and j.
5.4.4
Backpropagation over Correlation Geometries
Except for D and D⋆ , all computations involved in the five metrics can be backpropagated using existing techniques or PyTorch’s auto-differentiation. The matrix logarithm, matrix exponentiation, and Cholesky decomposition, together with their differentials and backpropagation, are reviewed in Sec. 2.8. It therefore remains to discuss D and D⋆ .
D and D⋆ . Their gradients can be backpropagated either approximately through their iterative algorithms or accurately using the following two propositions. Proposition 145 (Gradients w.r.t. D). [↓] Let l(·) be the loss function and define F : Hol(n) → S n by F (H) = Y = D(H) + H for any symmetric hollow matrix H, where S n is the Euclidean space of n × n symmetric matrices. Let Y = U ∆U ⊤ be the eigendecomposition with (δ1 , · · · , δn ) as eigenvalues. Given the succeeding ∂l ∂l gradient ∂Y , the output gradient ∂H is ∂l = off ∂H
∂l − exp∗,Y ∂Y
∂l 0 −1 D (H ) Dv 1⊤ , ∂Y
n with H 0 ∈ S++ having entries Hil0 =
P
(5.80)
j,k Uij Uik Ulj Ulk [Lexp ]j,k , where Lexp is the
171
5.4. Full-Rank Correlation Networks Loewner matrix in Eq. (2.91) specialized to f = exp and σi = δi . Here, D(·) : Rn×n → Diag(n) extracts the diagonal matrix, while Dv(·) : Rn×n → Rn returns a vector of diagonal elements. Besides, off(·) subtracts the diagonal matrix from a matrix, and exp∗,Y is the differential of the symmetric matrix exponential given by Eq. (2.90). Proposition 146 (Gradients w.r.t. D⋆ ). [↓] Following the notation in Thm. 145, + ⋆ ⋆ define F : Cor+ (n) → Row+ 1 (n) by F (C) = Σ = D (C)CD (C), where Row1 (n) is the manifold of n × n SPD matrices with unit row sum. Given the succeeding ∂l ∂l gradient ∂Σ , the output gradient ∂C is ∂l =∆ ∂C
∂l −1 ⊤ − (I + Σ) ve1 sym ∆, ∂Σ
(5.81)
∂l ∂l where ∆ = D(Σ)1/2 , ve = Dv Σ ∂Σ + ∂Σ Σ , I is the identity matrix, and 1 ∈ Rn is ⊤ . the vector with all entries equal to 1. Here, (A)sym = A+A 2
5.4.5
Experiments
We construct Riemannian networks on the correlation manifold, termed CorNets, using the proposed convolutional and MLR layers. Following previous work [106, 31, 50], we evaluate our approach on the Radar data set [31] for radar signal classification, along with the HDM05 [153], FPHA [80] and NTU120 [138] data sets for human action recognition. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.6. Implementation. We denote CorNet-Metric as the CorNet composed of correlation convolution and MLR layers under a specified metric. In line with Nguyen et al. [161], each CorNet consists of one correlation convolutional layer followed by a correlation MLR layer, trained with cross-entropy loss. Following Wang et al. [210], Nguyen et al. [161], each raw feature is modeled as a multi-channel [c, n, n] SPD tensor. Since matrix power effectively activates SPD matrices by deforming their geometry, as detailed in Sec. 4.3.3.1 and prior work [194, 53], we first apply a matrix power, and then convert the result to correlation matrices as the input of CorNet. Due to trivialization, all trainable manifold-valued parameters are represented by Euclidean parameters and optimized by standard Euclidean optimizers. We compare CorNets against representative Grassmannian and SPD networks, including GrNet [108], GyroGr [159], 172
Chapter 5. Riemannian Neural Networks
Manifold
Method
Radar
HDM05
FPHA
NTU120
Mean±STD
Time
Mean±STD
Time
Mean±STD
Time
Mean±STD
Time
Grassmann
GrNet [108] GyroGr∗ [159] GyroGr-Scaling∗ [159]
90.48 ± 0.76 90.64 ± 0.57 88.88 ± 1.52
1.39 1.38 1.63
63.19 ± 0.70 58.32 ± 1.23 39.75 ± 0.93
1.64 2.48 3.52
85.31 ± 0.90 79.62 ± 0.49 58.62 ± 1.66
0.70 0.70 1.03
57.59 ± 0.22 53.76 ± 0.18 43.90 ± 0.23
50.97 136.96 338.01
SPD
SPDNet [106] SPDNetBN [31] SPDResNet-AIM [118] SPDResNet-LEM [118] SPDNetLieBN-AIM [50] SPDNetLieBN-LCM [50] SPDNetMLR [51] GyroLE∗ [159] GyroLC∗ [159] GyroAI∗ [159] GyroSPD++∗ [161]
93.25 ± 1.10 94.85 ± 0.99 95.71 ± 0.37 95.89 ± 0.86 95.47 ± 0.90 94.80 ± 0.71 94.59 ± 0.82 96.24 ± 0.24 93.60 ± 1.31 96.29 ± 0.48 95.20 ± 0.88
0.66 1.25 0.96 0.77 1.21 1.10 0.66 0.79 0.66 0.99 5.09
64.57 ± 0.61 71.28 ± 0.79 64.95 ± 0.82 70.12 ± 2.45 71.83 ± 0.69 71.78 ± 0.44 65.90 ± 0.93 73.17 ± 0.37 67.53 ± 0.85 72.34 ± 1.06 69.82 ± 1.79
0.50 0.94 1.23 0.55 1.15 1.11 5.46 2.86 1.49 22.80 103.57
85.59 ± 0.72 89.33 ± 0.49 86.63 ± 0.55 85.07 ± 0.99 90.39 ± 0.66 86.33 ± 0.43 85.60 ± 0.43 90.73 ± 0.92 76.10 ± 0.63 89.60 ± 0.37 89.50 ± 0.37
0.28 0.58 0.69 0.30 0.97 0.59 0.88 1.59 0.78 12.62 66.35
51.25 ± 0.36 54.35 ± 0.43 57.33 ± 0.35 61.34 ± 2.02 58.20 ± 0.46 57.96 ± 0.43 58.59 ± 0.13 59.29 ± 0.42 59.29 ± 0.42 62.21 ± 0.29 61.57 ± 0.30
12.77 19.78 23.84 13.00 31.10 22.06 22.48 22.08 14.14 98.31 216.46
Correlation
CorNet-ECM CorNet-LECM CorNet-OLM CorNet-LSM CorNet-PHCM
97.71 ± 0.61 98.40 ± 0.70 97.57 ± 0.76 96.24 ± 1.48 96.56 ± 0.86
1.01 1.12 1.35 1.50 2.37
81.35 ± 1.27 78.05 ± 1.14 81.46 ± 0.61 74.89 ± 1.07 82.26 ± 0.92
0.60 0.64 0.93 0.98 1.10
92.17 ± 0.49 91.17 ± 0.32 91.63 ± 0.12 83.43 ± 0.65 90.03 ± 0.63
0.50 0.54 0.79 0.83 0.77
65.04 ± 0.14 65.03 ± 0.10 64.41 ± 0.23 60.69 ± 0.85 60.01 ± 0.22
12.06 12.68 16.07 16.28 16.92
Table 5.20: Five-fold results and training time per epoch on four data sets. The top 3 results are highlighted with red, blue, and cyan. ∗ denotes reproduced results due to missing official code. GyroGr-Scaling [159], SPDNet [106], SPDNetBN [31], RResNet [118], LieBN [50], SPD MLR [51], Gyro [159], and GyroSPD++[161]. 5.4.5.1
Main Results
Tab. 5.20 reports the five-fold results comparing our CorNets against existing SPD and Grassmannian baselines. We summarize the key observations below. • Effectiveness. CorNets consistently outperform both SPD and Grassmannian networks. Specifically, CorNets surpass the classic SPDNet by 5.15%, 17.69%, 6.58%, and 13.79% on four data sets, respectively, and outperform the best Grassmannian networks by 7.76%, 19.07%, 6.86%, and 7.45%. Despite not using BN or residual blocks, CorNets achieve superior performance compared to SPDNetBN, SPDNetLieBN, and RResNet. Notably, although CorNets share the same high-level architecture as GyroSPD++ (one manifold convolutional layer followed by one manifold MLR layer), CorNets exhibit better performance. These results highlight the effectiveness of correlation embedding and our method for constructing correlation networks. • Optimal Metric. The optimal metric for CorNets varies across data sets, indicating that the choice of geometry is a critical hyperparameter in Riemannian networks. Our framework enables seamless switching among five correlation geometries in a consistent architecture, demonstrating the adaptability of our approach to different tasks. 173
5.4. Full-Rank Correlation Networks Data Set Conv
MLR
ECM LECM OLM LSM PHCM
HDM05
FPHA
ECM
LECM
OLM
LSM
PHCM
ECM
LECM
OLM
LSM
PHCM
81.35 ± 1.27 66.49 ± 1.13 77.82 ± 0.48 68.83 ± 1.19 81.16 ± 0.40
73.38 ± 0.34 78.05 ± 1.14 76.56 ± 0.89 70.41 ± 1.57 80.05 ± 0.45
80.11 ± 0.77 79.21 ± 1.23 81.46 ± 0.61 67.56 ± 1.52 81.96 ± 0.51
78.54 ± 0.43 73.61 ± 0.99 80.77 ± 0.81 74.89 ± 1.07 78.28 ± 0.64
80.80 ± 0.54 58.37 ± 2.24 77.39 ± 1.29 72.69 ± 3.56 82.26 ± 0.92
92.17 ± 0.49 87.90 ± 0.57 92.17 ± 0.58 78.97 ± 2.80 88.30 ± 0.81
91.50 ± 0.21 91.17 ± 0.32 92.27 ± 0.78 75.10 ± 1.15 79.80 ± 0.69
91.67 ± 0.28 90.25 ± 0.25 91.63 ± 0.12 82.25 ± 3.38 87.37 ± 0.72
87.37 ± 1.14 89.63 ± 0.31 89.90 ± 0.67 83.43 ± 0.65 86.63 ± 0.27
91.97 ± 0.24 86.09 ± 0.98 91.83 ± 0.15 78.97 ± 4.97 90.03 ± 0.63
Table 5.21: Ablations on mixed geometries. Each row shows the metric used for Convolution (Conv), and each column is the metric for MLR. The diagonal entries indicate configurations where both layers use the same metric. The best result in each row is bold. • Efficiency. CorNets achieve efficiency comparable to or better than several baseline methods. The most efficient CorNet variant is based on ECM, owing to the simplest computations of ECM. Although GyroSPD++ uses the same architecture, CorNets achieve significantly greater efficiency, attributed to the heavy computational cost of the AIM-based computations in GyroSPD++ and the lightweight Riemannian computations on the correlation manifold. Particularly, on the largest NTU120 data set, CorNet-ECM and CorNet-LECM are the top two most efficient ones. 5.4.5.2
Ablations on Mixed Geometries
Our main experiments use the same metric for convolution and MLR. To evaluate mixed geometries, we assign different metrics to the two layers. Tab. 5.21 reports fivefold results on HDM05 and FPHA. Overall, consistent metrics yield the best accuracy. 5.4.5.3
Visualization
Figure 5.5: Illustration of the decision hyperplanes in the correlation MLRs under five different geometries. The 3 × 3 correlation manifold can be embedded as an open elliptope in R3 , by visualizing the strictly lower triangular part of each C ∈ Cor+ (3). The black dots denote the boundary. The PHCM hyperplane is defined by the one in the β-concatenated Poincaré space. Fig. 5.5 shows that different metrics induce visibly distinct curved hyperplanes. 174
Chapter 5. Riemannian Neural Networks 5.4.5.4
Potential and Necessity
Although correlation matrices are still SPD, naively treating Input Radar HDM05 FPHA SPD 93.25 ± 1.10 64.57 ± 0.61 85.59 ± 0.72 them as SPD inputs and feeding Correlation 89.49 ± 0.67 66.81 ± 0.73 83.37 ± 0.40 them into existing SPD networks fails to leverage their intrinsic geTable 5.22: SPDNet: SPD vs. correlation. ometric structures. To illustrate this, we use the classic SPDNet [106] but replace its covariance inputs with correlation matrices. The five-fold average results in Tab. 5.22 reveal two key insights: (1) on the HDM05 data set, correlation inputs lead to improved performance, suggesting that correlation embeddings can serve as compact and effective alternatives to covariance representations; and (2) on the other two data sets, the performance degrades, indicating that ignoring the specific geometry of correlation matrices can be detrimental. These findings highlight both the promise and the necessity of designing networks respecting the unique geometry of the correlation manifold. 5.4.5.5
Ablations on Correlation Embeddings
Data Set
Measurement
SPDMLR-Trivlz
CorMLR
LEM
LCM
AIM
ECM
LECM
OLM
LSM
PHCM
Radar
Acc Fit Time (s/epoch)
95.47 ± 0.66 0.65
95.55 ± 0.35 0.63
94.87 ± 0.87 0.99
89.47 ± 0.93 0.56
87.41 ± 0.23 0.62
85.79 ± 0.83 0.78
91.63 ± 0.32 0.68
83.33 ± 1.29 0.74
HDM05
Acc Fit Time (s/epoch)
54.31 ± 1.65 3.24
45.12 ± 1.05 5.38
52.46 ± 2.44 260.67
65.57 ± 0.62 3.18
64.44 ± 0.63 3.87
62.86 ± 0.65 3.39
64.01 ± 0.92 3.57
62.78 ± 0.85 2.73
FPHA
Acc Fit Time (s/epoch)
84.13 ± 1.14 0.51
76.62 ± 0.43 0.52
83.25 ± 0.59 18.96
85.37 ± 0.16 0.51
85.24 ± 0.22 0.64
84.67 ± 0.27 0.8
80.17 ± 0.15 0.81
73.67 ± 0.32 0.45
Table 5.23: Comparison of SPDMLR-Trivlz on raw covariances against CorMLR on raw correlations on all three data sets. The input matrix dimensions are 93 × 93, 63 × 63, and 20 × 20, respectively. To further evaluate the effectiveness of correlation embeddings, we compare the performance of directly classifying raw covariance matrices using the SPD MLRs in Thm. 117 with that of classifying corresponding raw correlation matrices using correlation MLR (CorMLR). The original SPDMLR involves an SPD matrix parameter for each class, which causes heavy Riemannian computations. For a fair comparison, we also implement a similar trivialization as Sec. 5.4.2.1, denoted as SPDMLR-Trivlz. We implement SPDMLR-Trivlz under LEM, LCM, and AIM, respectively. Tab. 5.23 presents the 5-fold average results on all three data sets. CorMLR performs better than SPDMLR-Trivlz on HDM05 and FPHA. Although CorMLR performs worse on Radar, we emphasize that 175
5.4. Full-Rank Correlation Networks these comparisons are conducted on a single MLR layer, which fails to fully uncover the potential of correlation matrices. Besides, SPDMLR under AIM is much slower than others, especially on HDM05, due to its complex computations. In contrast, CorMLR, especially under ECM and PHCM, offers competitive or superior efficiency. 5.4.5.6
Analysis of Covariance versus Correlation
In this section, we analyze when and why correlation matrices provide stronger representations than covariance matrices. The coefficient of variation of diagonal variances quantifies the variability of diagonal variances via per-sample coefficients of variation, and the ratio of diagonal to off-diagonal entries compares the magnitudes of diagonal and off-diagonal entries via their ratios. These analyses lead to two insights: (1) large variability and magnitude of diagonal elements can act as nuisance noise for SPD networks by overshadowing informative off-diagonal correlations; (2) under such cases, correlation representations that normalize variances and emphasize pairwise correlations tend to be more effective, which is especially evident on HDM05. Coefficient of Variation of Diagonal Variances. Channel 0
Channel 1
Channel 2
60
30
Count
40
Count
Count
50
20 10 0
0.6
0.8
0.6
0.8
1.0
1.2
1.4
1.6
1.8
2.0
1.0
1.2
1.4
1.6
1.8
2.0
1.00
1.25
1.50
1.75
2.00
Coefficient of Variation Channel 3
0.75 1.00 1.25 1.50 1.75 2.00 2.25 2.50
1.0
Coefficient of Variation Channel 4
1.5
2.0
Coefficient of Variation Channel 5
2.5
60
30
Count
40
Count
Count
50
20 10 0
Coefficient of Variation Channel 6
0.8
1.0
1.2
1.4
1.6
0.75
1.00
1.25
1.50
1.75
1.8
Coefficient of Variation Channel 7
2.0
2.2
0.75 1.00 1.25 1.50 1.75 2.00 2.25 2.50
Coefficient of Variation Channel 8
60
30
Count
40
Count
Count
50
20 10 0
0.75
Coefficient of Variation
2.25
2.00
Coefficient of Variation
2.25
1.0
1.5
2.0
Coefficient of Variation
2.5
Figure 5.6: Distribution of per-sample coefficients of variation of diagonal variances on FPHA. Higher values indicate stronger diagonal variability, which could cause nuisance noise. 176
Chapter 5. Riemannian Neural Networks Channel 0
Channel 1
Channel 2
40
Count
60
Count
Count
80
20 0
1.0
1.5
2.0
2.5
Coefficient of Variation
3.0
1.0
1.5
2.0
2.5
Coefficient of Variation
3.0
1.0
1.5
2.0
2.5
Coefficient of Variation
3.0
Figure 5.7: Distribution of per-sample coefficients of variation of diagonal variances on HDM05. Higher values indicate stronger diagonal variability, which could cause nuisance noise. This section investigates why CorNets yield substantially larger gains over SPD networks on HDM05 compared to FPHA. n Setup. For each covariance matrix Σ ∈ S++ we extract the diagonal vector v = (Σ11 , . . . , Σnn ).
(5.82)
We compute the coefficient of variation of v as CV =
std(v) , mean(v) + ε
(5.83)
where ε = 10−8 ensures numerical stability. As shown in Sec. A.3.6.1, each sequence is modeled as a c-channel tensor of covariance matrices. The above procedure yields one coefficient of variation per channel for each sample. We visualize their empirical distributions per channel. Analysis. Figs. 5.6 and 5.7 show that the coefficients of variation w.r.t. diagonal variance are large on both data sets. On FPHA, most values fall between 0.8 and 2.0. On HDM05, they are even larger, typically between 1.0 and 3.0. Such large fluctuations indicate that diagonal variances change substantially and could bring nuisance noise for SPD networks. In contrast, correlation matrices allow CorNets to focus on pairwise relationships. This explains the consistent improvements over SPD networks and the larger gains on HDM05. Ratio of Diagonal to Off-Diagonal Entries in Covariance Features.
177
5.4. Full-Rank Correlation Networks Channel 1
Count
Count
60
60 40
1.5
2.0
2.5
3.0
3.5
0
4.0
Ratio of diagonal to off-diagonal entries Channel 3
70 60 50 40 30 20 10 0
1.5
2.0
2.5
3.0
3.5
0
4.0
Ratio of diagonal to off-diagonal entries Channel 4
80
Count
60 40 20 1.5
2.0
2.5
3.0
3.5
0
4.0
40 20
20
1.5
Ratio of diagonal to off-diagonal entries Channel 6
2.0
2.5
3.0
3.5
70 60 50 40 30 20 10 0
Ratio of diagonal to off-diagonal entries Channel 7 80
60
60
20
Count
80
60 40
40 20
0
1.5
2.0
2.5
3.0
3.5
4.0
2.0
1.5
2.0
2.5
3.0
3.5
2.5
3.0
3.5
40 20
0
4.5
1.5
Ratio of diagonal to off-diagonal entries Channel 5
Ratio of diagonal to off-diagonal entries Channel 8
80
Count
Count
Channel 2 80
80
Count
Count
Count
Channel 0 70 60 50 40 30 20 10 0
1.5
Ratio of diagonal to off-diagonal entries
2.0
2.5
3.0
3.5
0
4.0
Ratio of diagonal to off-diagonal entries
1.5
2.0
2.5
3.0
3.5
4.0
Ratio of diagonal to off-diagonal entries
Figure 5.8: Distribution of ratios of diagonal to off-diagonal entries on FPHA. Channel 0
150
Channel 1
Count
Count
100 75 50
Count
150
125
100 50
25 0
Channel 2
2
3
4
5
6
7
Ratio of diagonal to off-diagonal entries
0
2
3
4
5
6
Ratio of diagonal to off-diagonal entries
7
150 125 100 75 50 25 0
2
3
4
5
6
Ratio of diagonal to off-diagonal entries
Figure 5.9: Distribution of ratios of diagonal to off-diagonal entries on HDM05. This section further examines why CorNets achieve larger gains over SPD networks on HDM05 than on FPHA. We analyze the ratio of diagonal to off-diagonal entries in covariance matrices on FPHA and HDM05, to quantify how strongly variance terms overshadow pairwise correlations. n Setup. For each covariance matrix Σ ∈ S++ we compute the mean magnitude of diagonal entries n 1X D= |Σii |, (5.84) n i=1
178
Chapter 5. Riemannian Neural Networks and the mean magnitude of off-diagonal entries O=
X 1 |Σij |. n(n − 1) i̸=j
(5.85)
D , O
(5.86)
We then form the sample-wise ratio R=
which measures how much larger the diagonal amplitudes are compared to the offdiagonal correlations. Each sample yields one ratio per channel, and we visualize the empirical distributions of these ratios on FPHA and HDM05. Analysis. Figs. 5.8 and 5.9 show that both data sets have ratios well above one. On FPHA, most ratios lie between 1.7 and 3.0, indicating that diagonal amplitudes are noticeably larger than off-diagonal correlations. HDM05 exhibits even larger ratios, typically between 2.0 and 6.0, with many above 3.0. These statistics indicate that covariance representations on both data sets are strongly dominated by diagonal entries, with more pronounced dominance on HDM05. When diagonal terms dominate, SPD networks trained on covariance inputs tend to overemphasize variances and underexploit informative pairwise correlations. Correlation matrices normalize variances and highlight off-diagonal interactions, which explains why CorNets outperform SPD baselines on both data sets and why the improvement is substantially larger on HDM05. 5.4.5.7
Normalized Covariance vs. Correlation
Setup. We evaluate SPD-based baselines by covariance inputs normalized by their b= largest eigenvalue. Given a covariance matrix Σ, we get the normalized SPD input Σ Σ/λmax (Σ) and feed it into existing SPD networks. This variant is denoted by “-EigN”. We report results on the Radar, HDM05, and FPHA data sets for representative SPD models: SPDNet, SPDNetBN, SPDResNet, SPDNetLieBN, SPDNetMLR, GyroAI, and GyroSPD++. Here, SPDResNet is implemented under the LEM, while SPDNetLieBN follows the LCM.
179
5.4. Full-Rank Correlation Networks Manifold
n S++
Cor+ (n)
Method
Radar
HDM05
FPHA
SPDNet SPDNet-EigN
93.25 ± 1.10 86.91 ± 0.57
64.57 ± 0.61 66.62 ± 0.73
85.59 ± 0.72 84.90 ± 0.62
SPDNetBN SPDNetBN-EigN
94.85 ± 0.99 89.25 ± 1.19
71.28 ± 0.79 71.59 ± 0.68
89.33 ± 0.49 88.47 ± 0.39
SPDResNet SPDResNet-EigN
95.89 ± 0.86 92.61 ± 0.96
70.12 ± 2.45 71.02 ± 0.91
85.07 ± 0.99 84.53 ± 0.46
SPDNetLieBN SPDNetLieBN-EigN
94.80 ± 0.71 88.91 ± 1.21
71.78 ± 0.44 70.61 ± 1.04
86.33 ± 0.43 83.73 ± 0.65
SPDNetMLR SPDNetMLR-EigN
94.59 ± 0.82 89.41 ± 0.58
65.90 ± 0.93 66.89 ± 0.63
85.60 ± 0.43 83.63 ± 1.09
GyroAI GyroAI-EigN
96.29 ± 0.48 91.36 ± 0.80
72.34 ± 1.06 72.64 ± 0.70
89.60 ± 0.37 89.90 ± 0.31
GyroSPD++ GyroSPD++-EigN
95.20 ± 0.88 90.83 ± 1.09
69.82 ± 1.79 66.92 ± 0.28
89.50 ± 0.37 84.29 ± 0.14
CorNet-ECM CorNet-LECM CorNet-OLM CorNet-LSM CorNet-PHCM
97.71 ± 0.61 98.40 ± 0.70 97.57 ± 0.76 96.24 ± 1.48 96.56 ± 0.86
81.35 ± 1.27 78.05 ± 1.14 81.46 ± 0.61 74.89 ± 1.07 82.26 ± 0.92
92.17 ± 0.49 91.17 ± 0.32 91.63 ± 0.12 83.43 ± 0.65 90.03 ± 0.63
Table 5.24: SPD networks with or without normalized SPD inputs. Results. Tab. 5.24 summarizes the results. On HDM05, eigenvalue normalization has only a marginal effect and the normalized variants achieve accuracy comparable to their unnormalized counterparts. On FPHA and, in particular, on Radar, normalization usually reduces accuracy. The behavior of GyroSPD++ is especially informative. GyroSPD++ and CorNet share a similar architecture, consisting of one convolution followed by an MLR layer. However, GyroSPD++-EigN performs worse than GyroSPD++ on all three data sets, while CorNet with correlation inputs achieves clear improvements over GyroSPD++. These phenomena can be explained by two factors. (1) Redundancy. The raw samples on HDM05 and FPHA have already undergone centering, scaling, and normalization before covariance modeling. Dividing by λmax (Σ) therefore introduces little additional control over scale, which explains the marginal effect on HDM05. (2) Scaled Covariance versus Correlation. Since EigN is equivalent to uniformly rescaling the raw samples before covariance computation, the normalized covariance matrices remain covariances and do not encode new statistical information. 180
Chapter 5. Riemannian Neural Networks Moreover, forcing the largest eigenvalue to 1 can remove potentially informative differences in overall energy across samples, which aligns with the degradation observed for EigN variants, especially GyroSPD++-EigN. In contrast, correlation normalization uses a different scaling factor for each pair of variables, Σij Corij = p , Σii Σjj
(5.87)
producing standardized correlation coefficients. Therefore, global eigenvalue scaling is statistically distinct from correlation normalization and fails to capture the benefits of explicit correlation modeling. 5.4.5.8
Ablations on Activations
Metric
Activation
Radar
HDM05
FPHA
ECM
ReLU None
97.41 ± 0.25 97.71 ± 0.61
81.23 ± 0.46 81.35 ± 1.27
89.80 ± 0.58 92.17 ± 0.49
LECM
ReLU None
97.23 ± 0.67 98.40 ± 0.70
77.51 ± 1.02 78.05 ± 1.14
91.00 ± 0.15 91.17 ± 0.32
OLM
ReLU None
97.52 ± 0.47 97.57 ± 0.76
81.86 ± 0.65 81.46 ± 0.61
91.47 ± 0.19 91.63 ± 0.12
LSM
ReLU None
95.60 ± 0.97 96.24 ± 1.48
N/A 74.89 ± 1.07
N/A 83.43 ± 0.65
PHCM
ReLU None
96.40 ± 0.25 96.56 ± 0.86
77.32 ± 1.56 82.26 ± 0.92
88.63 ± 0.22 90.03 ± 0.63
Table 5.25: Comparison of CorNet with or without activations. In the main experiments, we follow HNN++ [181] and GyroSPD++ [161], and do not use explicit activations, as the manifold itself introduces nonlinearity. We further conduct an ablation on activations. Following Ganea et al. [76, Sec. 3.2], we define activations in the tangent space at the identity, i.e., ExpI ◦δ ◦ LogI for four Log-Euclidean metrics, and Exp0 ◦δ ◦ Log0 for PHCM in the β-concatenated Poincaré vector, where δ is ReLU [82]. Specifically, we insert a ReLU after the correlation convolution. As shown in Tab. 5.25, adding activations generally yields no benefits and can even degrade performance. The variant without activation consistently achieves higher or comparable accuracy, except CorNet-OLM for HDM05. Moreover, CorNet-LSM with activation 181
5.4. Full-Rank Correlation Networks fails to converge on HDM05 and FPHA. These results suggest that CorNet already provides sufficient nonlinearity, rendering additional activations redundant. 5.4.5.9
Scalability of Correlation Metrics
Dim
ECM
LECM
OLM
LSM
PHCM
30 50 100 150 200 250 300 400 500 600 700 800 900 1000
0.0004 0.0004 0.0008 0.0015 0.0025 0.0037 0.0053 0.0092 0.0143 0.0206 0.0289 0.039 0.0535 0.0706
0.0018 0.0027 0.0054 0.0100 0.0197 0.0345 0.0733 0.1796 0.3076 0.5983 1.0961 1.8689 2.9886 3.7259
0.0012 0.0318 0.0764 0.1247 0.1906 0.2352 0.3434 0.5163 0.6907 0.9331 1.2432 1.6658 2.2156 2.539
0.0019 0.0334 0.0781 0.1267 0.1938 0.2379 0.3454 0.5261 0.6961 0.9484 1.2575 1.6815 2.2303 2.5783
0.0131 0.0211 0.0413 0.2284 0.3320 0.4414 0.5732 0.4807 0.5693 0.7923 1.0417 1.3387 1.7324 1.229
Table 5.26: Average runtime (s) of a single forward pass in CorNet under different metrics and input dimensions. The best results are bold. We evaluate the computational efficiency of correlation metrics across increasing input dimensions using CorNet with one correlation FC layer followed by one correlation MLR layer. Each input correlation matrix of size [n, n] is mapped to [20, 20] by the FC layer and then classified into 10 classes by the MLR layer. For each listed dimension 30 ≤ n ≤ 1000, we randomly generate 30 correlation matrices and record the average runtime of a single forward pass. As implied by Tab. 2.7, the runtime is governed by two factors: the codomain computation (Euclidean or hyperbolic) and the complexity of the diffeomorphism. The results are summarized in Tab. 5.26. We have the following findings. • ECM is consistently the most efficient metric, benefiting from both a Euclidean codomain and the simplest diffeomorphism. 182
Chapter 5. Riemannian Neural Networks • At very low dimensions (n ≤ 100), the relative costs of the non-ECM metrics are not yet stable. At n = 50 and n = 100, the ordering is ECM < LECM < PHCM < OLM ≈ LSM.
(5.88)
At the smallest tested dimension n = 30, all runtimes remain small and their relative ordering differs. At intermediate dimensions (150 ≤ n ≤ 300), the ordering becomes ECM < LECM < OLM ≈ LSM < PHCM.
(5.89)
ECM < PHCM < OLM ≈ LSM < LECM.
(5.90)
Here, the cost of PHCM’s hyperbolic computations dominates, while the dimensiondependent cost of LECM’s matrix functions is not yet pronounced. • From n = 400, PHCM becomes faster than OLM and LSM, and at n = 700, it also becomes faster than LECM. At high dimensions (n ≥ 800), the ordering is Here, diffeomorphisms dominate: ECM and PHCM scale better thanks to relatively lightweight Cholesky decomposition, while OLM and LSM slow down due to matrix logarithm/exponentiation. LECM is the slowest, as its log ◦Θ requires two nested matrix functions.
5.5
Conclusion
This chapter developed manifold-specific neural components and architectures by exploiting the additional structures of particular hyperbolic models and correlation manifolds. The first part addressed representation choice through PVNN. The unconstrained PV model is connected to the Poincaré and hyperboloid models by Riemannian isometries, but it avoids their explicit constraints. By deriving closed-form core Riemannian operators and relating them to the PV gyrovector structure, we constructed PV MLR, FC, convolutional, activation, and normalization layers. Together, these layers constitute a complete PVNN framework for constructing model-specific hyperbolic architectures. Experiments demonstrate the superior numerical stability and competitive performance of the PV representation. The second part shifted attention from the representation itself to the geometric principle for building neural layers. Using Busemann functions and horospheres, we developed a common construction for the Poincaré and Lorentz models. BMLR ad183
5.5. Conclusion mits an exact point-to-horosphere interpretation, avoids an additional manifold-valued point parameter, supports batch-efficient evaluation, and recovers Euclidean MLR in the zero-curvature limit. The same Busemann logits yield explicit BFC layers with practical O(nm) complexity and corresponding Euclidean limits. Experiments on image classification, genome sequence learning, node classification, and link prediction showed that these layers generally improve upon existing hyperbolic alternatives while retaining comparable computational cost. The third part extended manifold-specific network design from hyperbolic vectors to full-rank correlation matrices. The four flat geometries, ECM, LECM, OLM, and LSM, admit Euclidean isometries that yield closed-form correlation MLR, FC, and convolutional layers. For the non-flat PHCM geometry, the Cholesky representation identified correlation matrices with a product of Poincaré balls, enabling analogous layers through β-concatenation and β-splitting. Analytic gradients for the Riemannian computations enabled accurate end-to-end backpropagation under OLM and LSM. Across the four evaluated data sets, the best CorNet geometry outperformed the considered SPD and Grassmannian baselines. All three designs nevertheless operate under prescribed Riemannian geometries. The next chapter therefore turns from the design of modules and architectures to that of the underlying metrics themselves.
184
Chapter 6 Fast and Stable Geometries on SPD Manifolds 6.1
Introduction
The preceding chapters developed Riemannian neural components and architectures under prescribed Riemannian metrics. Despite their different constructions, they share the underlying metric as a common starting point. As reviewed in Chapter 2, a Riemannian metric assigns inner products to tangent spaces and determines distance, geodesics, exponential and logarithmic maps, and parallel transport. The induced distance also defines the Fréchet statistics used to summarize manifold-valued features. When combined with compatible Lie group or gyrovector structures, a Riemannian metric further supports manifold analogues of addition and scalar multiplication. These geometric primitives offer powerful toolkits for building Riemannian neural networks. Based on this observation, we shift our attention from designing neural modules and architectures under prescribed geometries to designing the underlying geometry itself. Our goal is to balance geometric flexibility, theoretical convenience, computational efficiency, and numerical stability. Accordingly, we seek geometries whose Riemannian operators and compatible algebraic operations admit closed-form expressions and can be inserted directly into deep learning architectures. We focus on the SPD manifold, which encodes covariance or other second-order statistics and arises naturally in many applications [37, 31, 134, 61, 152, 141, 231, 106].1 1
Our focus is distinct from SPD metric-learning methods that learn distance functions induced by an existing Riemannian metric. Instead, we seek to design new Riemannian metrics themselves.
185
6.2. Adaptive Log-Euclidean Metrics We pursue two complementary routes. Sec. 6.2 starts from pullback Euclidean geometry and proposes Adaptive Log-Euclidean Metrics (ALEMs). The resulting adaptive geometry retains closed-form Riemannian operators and a convenient abelian Lie group structure. Sec. 6.3 instead exposes the product structure of the Cholesky manifold and transfers these geometries to the SPD manifold, yielding simple and stable closed-form Riemannian and gyro operators.
6.2
Adaptive Log-Euclidean Metrics
6.2.1
Introduction
As reviewed in Sec. 2.9.1, most popular Riemannian metrics on the SPD manifold are fixed, which can limit the expressive capacity of the associated geometry. A common way to construct SPD metrics is through pullbacks along diffeomorphisms, which transfer Riemannian structures from simpler source manifolds. For instance, Thanwerdas and Pennec [195] explained AIM as the pullback metric from a left-invariant metric on the Cholesky manifold. The matrix-power deformations in Secs. 3.2.5.1 and 4.3.3.1 are also representative examples. Inspired by the above observations, we leverage pullback techniques to introduce adaptive Riemannian metrics. In particular, we first show that several Riemannian metrics on SPD manifolds, including LEM, LCM, and their generalizations, can be explained as pullback metrics from the standard Euclidean space. We refer to these metrics as pullback Euclidean metrics. Then, we propose a general framework for characterizing the properties of pullback Euclidean metrics. Our framework can explain the widely used LEM [9] and LCM [137]. We focus on LEM on SPD manifolds and extend it into Adaptive Log-Euclidean Metrics (ALEMs). Besides, we present a complete theoretical discussion on the properties of ALEMs. Compared with the existing Riemannian metrics, our metrics are learnable, adapting to the characteristics of the data sets. The effectiveness of our metrics is demonstrated by experiments as well as the applications to recently developed Riemannian building blocks, including the RBN framework developed in Chapter 3, Riemannian residual blocks [117], and the Riemannian classifiers developed in Chapter 4. Drawing on this, our contributions are summarized as follows: (1) We reveal the connection of two popular Riemannian metrics (LEM and LCM) by the pullback technique and propose a general framework for pullback Euclidean metrics. 186
Chapter 6. Fast and Stable Geometries on SPD Manifolds (2) Based on our framework, we propose specific ALEMs on SPD manifolds and conduct comprehensive analyses in terms of the algebraic, analytic, and geometric properties. (3) Extensive experiments on widely used SPD learning benchmarks demonstrate that our metrics exhibit consistent performance gain across data sets.2 Outline. The necessary background on differential geometry, pullback metrics, and SPD geometry has been established in Chapter 2. Sec. 6.2.2 develops ALEM and its differentials, and Sec. 6.2.2.5 studies its geometric properties. Sec. 6.2.3 derives the gradients and parameter-update rules for the general matrix logarithm and exponential. Sec. 6.2.4 instantiates the general matrix logarithm as ALog in SPDNet and applies ALEM to other Riemannian building blocks. Proofs are deferred to Sec. B.9.
6.2.2
Adaptive Log-Euclidean Metrics
In this section, we show that both (α, β)-LEM and LCM are pullback metrics from the Euclidean space. Inspired by this observation, we present a general framework for characterizing pullback Euclidean metrics. Then, we focus on generalizing LEM. 6.2.2.1
Rethinking (α, β)-LEM and LCM
Among the existing Riemannian metrics on the SPD manifold, LEM is popular in many applications, given its closed form for the Fréchet mean and clear vector-space and Liegroup structures. In addition, the nascent LCM, which is gaining increasing attention, shares similar properties with LEM. LEM is derived from Lie group translation [9], while LCM is obtained as the pullback of a metric on Ln++ [137]. Besides, (α, β)-LEM is obtained as a pullback of LEM [196]. However, the same mathematical logic underlies their derivations. We denote the Euclidean space of n × n lower triangular matrices by n LTn . We define ψLC : S++ → LTn as ψLC (P ) = ⌊L⌋ + Dlog(D(L)),
(6.1)
where L is the Cholesky factor of the SPD matrix P , ⌊L⌋ is the strictly lower part of L, D(L) is a diagonal matrix with diagonal elements of L, and Dlog applies the natural logarithm element-wise to the diagonal. Then, we have the following theorem. 2
The code is available at https://github.com/GitZH-Chen/ALEM.
187
6.2. Adaptive Log-Euclidean Metrics Theorem 147. [↓] (α, β)-LEM is the pullback metric from the Euclidean space of S n with an O(n)-invariant inner product ⟨·, ·⟩(a,b) by matrix logarithm. Specifically, the standard LEM is the pullback metric from the Euclidean space of S n with the standard Frobenius inner product by matrix logarithm. LCM is the pullback metric from LTn with the Frobenius inner product by ψLC . As Euclidean spaces of the same dimension are naturally isometric, it follows that both (α, β)-LEM and LCM are pulled back from the standard Euclidean space S n . Corollary 148. [↓] (α, β)-LEM and LCM are pullback metrics from S n with the standard Frobenius inner product.
6.2.2.2
Pullback Euclidean Metrics on SPD Manifolds
Sec. 6.2.2.1 has shown how LEM is derived from matrix logarithm. Besides, as shown n are also induced in Arsigny et al. [9], operations in Lie group and linear space on S++ from matrix logarithm. Now, let us explain the underlying mechanism in detail. A matrix logarithm is a diffeomorphism (a smooth bijection with a smooth inverse). The property of bijection offers the possibility of transferring algebraic structures from S n n . The smoothness of the matrix logarithm and its inverse suggests that smooth into S++ structures, such as a Lie group structure and a Riemannian metric, can be transferred n n → S n , it suffices to . More generally, given an arbitrary diffeomorphism ϕ : S++ to S++ n pull various properties from the Euclidean space back to the SPD manifold S++ by ϕ as n well. Besides, the computation of the induced operators in S++ by ϕ is usually simple. n n Lemma 149. [↓] Let S1 , S2 ∈ S++ , V ∈ TS1 S++ , and k ∈ R, and let g E be the n Frobenius inner product in S n . Let ϕ : S++ → S n be a diffeomorphism, and denote n by ϕ∗,S . We define the following operations: its differential at S ∈ S++
Element Addition: S1 ⊙ϕ S2 = ϕ−1 (ϕ(S1 ) + ϕ(S2 )),
Scalar Multiplication: k ⊛ϕ S2 = ϕ−1 (kϕ(S2 )),
Inner Product: ⟨S1 , S2 ⟩ϕ = ⟨ϕ(S1 ), ϕ(S2 )⟩,
Riemannian Metric: g ϕ = ϕ∗ g E , Then, we have the following conclusions:
n (1) {S++ , ⊙ϕ , ⊛ϕ , ⟨·, ·⟩ϕ } is a Hilbert space over R.
188
(6.2) (6.3) (6.4) (6.5)
Chapter 6. Fast and Stable Geometries on SPD Manifolds
n n (2) {S++ , ⊙ϕ } is an abelian Lie group. {S++ , g ϕ } is a Riemannian manifold. The associated Riemannian operators are as follows:
dϕ (S1 , S2 ) = ∥ϕ(S1 ) − ϕ(S2 )∥F ,
(6.6)
ExpS1 V = ϕ−1 (ϕ(S1 ) + ϕ∗,S1 V ),
(6.7)
LogS1 S2 = ϕ−1 ∗,ϕ(S1 ) (ϕ(S2 ) − ϕ(S1 )),
(6.8)
PTS1 →S2 (V ) = ϕ−1 ∗,ϕ(S2 ) ◦ ϕ∗,S1 (V ),
(6.9)
n is a tangent vector, ExpS1 , where ∥ · ∥F is the Frobenius norm, V ∈ TS1 S++ LogS1 , and PTS1 →S2 are the Riemannian exponential map at S1 , logarithmic map at S1 , and parallel transport along the geodesic connecting S1 and S2 , −1 ϕ respectively, and ϕ−1 ∗ denotes the differential of ϕ . Then g is a bi-invariant metric, called a Pullback Euclidean Metric induced by ϕ.
(3) ϕ is an isomorphism: (a) a linear isomorphism preserving the inner product; (b) a Lie group isomorphism; (c) a Riemannian isometry. In fact, (α, β)-LEM and LCM are special cases of Thm. 149, as are the linear-space and Lie-group structures in Arsigny et al. [9] and the Lie-group structure in Lin [137]. In addition, neither Arsigny et al. [9] nor Lin [137] reveals the Hilbert space structures n . in S++
6.2.2.3
Adaptive Log-Euclidean Metrics
The key to Thm. 149 lies in the diffeomorphism ϕ. If we have a proper ϕ, Riemannian metrics on SPD manifolds can be induced. In the following, we will present our mappings and then discuss the induced metrics. As reviewed in Sec. 2.8.1, the matrix logarithm reduces to a scalar logarithm, which is a diffeomorphism between R++ and R. Following this hint, the eigenvalue-based n diffeomorphism between S++ and S n reduces to a scalar diffeomorphism between R++ and R. A very natural idea is to substitute the natural logarithm with logarithms with arbitrary proper bases. Throughout this part, log(·) without a subscript denotes the natural scalar or matrix logarithm. Every scalar or matrix logarithm with a general or adaptive base is written explicitly as logα (·), without omitting the subscript. The base parameter α is interpreted as either a scalar or a vector according to its argument. For 189
6.2. Adaptive Log-Euclidean Metrics a scalar base α ∈ R++ \ {1} and x ∈ R++ , we define logα (x) =
log(x) . log(α)
(6.10)
When α = e, this scalar logarithm reduces to the natural logarithm, which we write without a subscript as log(·). For a diagonal matrix X, the same notation is extended to a base vector as logα (X) = diag(loga1 (x11 ), loga2 (x22 ), · · · , logan (xnn )),
(6.11)
where α = (a1 , a2 , · · · , an ) ∈ (R++ \ {1})n is the base vector, diag(·) is the diagonalization operator, and X is an n × n diagonal matrix. When α is scalar in a matrix expression, it denotes the constant base vector (α, . . . , α). Together with eigendecomposition, a general matrix logarithm is defined by logα (S) = U logα (Σ)U ⊤ ,
(6.12)
where S = U ΣU ⊤ is the eigendecomposition. As a special case, α = (e, e, · · · , e)
=⇒
logα = log .
(6.13)
As with the scalar logarithm, we have the following proposition. Proposition 150 (Diffeomorphism). [↓] logα is a diffeomorphism, a smooth bijecn n tion with a smooth inverse log−1 α : S → S++ defined as Σ11 Σ22 Σnn ⊤ log−1 α (X) = U diag(a1 , a2 , · · · , an )U ,
(6.14)
where X = U ΣU ⊤ is the eigendecomposition. Remark 151. The general matrix logarithm logα is an arbitrary member of the following family {logα | α = (a1 , · · · , an ) ∈ (R++ \ {1})n } .
(6.15)
Besides, there could be some ambiguity in Eq. (6.12) under different arrangements of eigenvalues and eigenvectors. In fact, there is a correspondence between scalar
190
Chapter 6. Fast and Stable Geometries on SPD Manifolds logai and eigenvalues and eigenvectors. See Sec. A.4.4.1 for more details. n Since logα is a diffeomorphism from S++ onto S n , all the results in Thm. 149 hold true.
Theorem 152. [↓] Following the notation in Thm. 149, we define ⊕ALE and ⊙ALE as in Eqs. (6.2) and (6.3). We define ⟨·, ·⟩logα and g logα as in Eqs. (6.4) and (6.5). Then, we have the following conclusions: n (1) {S++ , ⊕ALE , ⊙ALE , ⟨·, ·⟩logα } is a Hilbert space over R. n n (2) {S++ , ⊕ALE } is an abelian Lie group. g logα is a Riemannian metric on S++ . We call this metric the Adaptive Log-Euclidean Metric (ALEM) and denote g logα by g ALE . The associated Riemannian operators are as follows:
dALE (S1 , S2 ) = ∥ logα (S1 ) − logα (S2 )∥F , log (S ) + (log ) ExpS1 V = log−1 V , 1 α α ∗,S1 α LogS1 S2 = log−1 α ∗,X1 (logα (S2 ) − logα (S1 )) , PTS1 →S2 (V ) = log−1 α ∗,X2 ◦ (logα )∗,S1 (V ),
(6.16) (6.17) (6.18) (6.19)
where Xi = logα (Si ) ∈ S n for i = 1, 2.
(3) logα is an isomorphism: (a) a linear isomorphism preserving the inner product; (b) a Lie group isomorphism; (c) a Riemannian isometry. Remark 153. Obviously, ALEM varies with different base parameters α in logα . We thus use the plural to describe our metrics. Besides, our metrics can be learned. This is why we call them adaptive metrics. Analogously to (α, β)-LEM, we can also define (a, b)-ALEM as the pullback of an O(n)-invariant inner product: g (a,b)-ALE = log∗α g (a,b)-E ,
(6.20)
where we denote the O(n)-invariant inner product ⟨·, ·⟩(a,b) by g (a,b)-E . g (a,b)-ALE also shares the properties presented in Thm. 152. Nevertheless, this part focuses on (a, b) = (1, 0).
191
6.2. Adaptive Log-Euclidean Metrics 6.2.2.4
Differentials of General Logarithms
Eqs. (6.17) to (6.19) require the differential maps of logα and log−1 α . This subsection introduces the concrete formulae of the associated differential maps. n Proposition 154 (Differentials). [↓] For a tangent vector V ∈ TS S++ , the differn n n ential (logα )∗,S : TS S++ → Tlogα (S) S of logα at S ∈ S++ is given by
(logα )∗,S (V ) = Q + Q⊤ + W,
(6.21)
where Q = DU logα (Σ)U ⊤ , DU = ( (σ1 In − S)+ V u1 · · · (σn In − S)+ V un ), ⊤ u⊤ u1 V u1 n V un ,··· , W = U diag U ⊤, σ1 log(a1 ) σn log(an ) (·)+ is the Moore–Penrose inverse, u1 , · · · , un are orthonormal eigenvectors of S, and the associated eigenvalues are σ1 , · · · , σn . Symmetrically, for a tangent vector Ve ∈ TX S n , the differential log−1 α ∗,X : n TX S n → Tlog−1 S n of log−1 α at X ∈ S is given by α (X) ++ e e e⊤ f log−1 α ∗,X (V ) = Q + Q + W ,
(6.22)
eΣ e e⊤ where X = U eigendecomposition. Here, DUe is defined similarly and U is the e ⊤ . Moreover, e = D e diag aσ1e1 , · · · , aσnen U Q U σ en ⊤ e f=U e diag log(a1 )aσ1e1 u eu e⊤ W e⊤ V e , · · · , log(a )a u e V u e 1 n n n U . 1 n
Arsigny et al. [9] write the differential of the matrix exponential as an infinite series. The differential of log−1 α can also be rewritten in this way. Proposition 155 (Differential as Infinite Series). [↓] Following the notation in Thm. 154, the differential of log−1 α can also be formulated as e log−1 α ∗,X (V )
∞ k−1 X 1 X e k−l−1 = ( (P X) (DPe X + PeVe )(PeX)l ), k! l=0 k=1
192
(6.23)
Chapter 6. Fast and Stable Geometries on SPD Manifolds
e BU e ⊤ , B = diag (log(a1 ), · · · , log(an )), D e = D e B U e⊤ + U e BD⊤ . where Pe = U e P U U
When log−1 α is reduced to the matrix exponential, Eq. (6.23) coincides with the expression in Arsigny et al. [9, Eq. (8)], and our ALEM becomes exactly LEM. 6.2.2.5
Properties of ALEM
Since our ALEMs are natural generalizations of LEM, they intuitively share many of its properties. This subsubsection introduces some useful properties of our ALEMs for machine learning. Fréchet means are important tools for SPD matrix learning [94, 36, 31, 34]. Like LEM, our ALEM also admits closed-form expressions for Fréchet means. We present a more general result, the weighted Fréchet mean. Proposition 156 (Weighted Fréchet Means). [↓] For m points S1 , . . . , Sm on the P SPD manifold with associated weights w1 , . . . , wm ∈ R+ satisfying m i=1 wi > 0, the n , dALE } has a closed form weighted Fréchet mean M over the metric space {S++ M = log−1 α
m X i=1
wi
Pm
j=1 wj
!
logα (Si ) .
(6.24)
Like LEM, although ALEM is not affine-invariant, it enjoys several other invariance properties. Proposition 157 (Bi-invariance). [↓] ALEM is a Lie group bi-invariant metric. Proposition 158 (Exponential Invariance). [↓] The Fréchet means under ALEM n are exponential-invariant. In other words, for S1 , . . . , Sm ∈ S++ and β ∈ R, β (FM(S1 , . . . , Sm ))β = FM(S1β , . . . , Sm ),
(6.25)
where FM(S1 , · · · , Sm ) denotes the Fréchet mean of S1 , · · · , Sm . In addition to exponential invariance, the Fréchet mean induced by our ALEM also satisfies various properties presented in Ando et al. [5]. Proposition 159. [↓] Let the following be SPD matrices: n A, B, C, A0 , B0 , C0 ∈ S++ .
193
(6.26)
6.2. Adaptive Log-Euclidean Metrics
Let FM(A, B, C) denote the Fréchet mean of A, B, C under ALEM. Then the Fréchet mean satisfies the following properties. (1) U1: permutation invariance. For any permutation π of {A, B, C}, FM π(A, B, C) = FM(A, B, C).
(6.27)
(2) U2. FM(A, A, A) = A. The following properties hold if A, B, C, A0 , B0 , C0 commute. (1) V1: joint homogeneity. FM(aA, bB, cC) = (abc)1/3 FM(A, B, C),
∀a, b, c > 0.
(6.28)
(2) V2: monotonicity. The map (A, B, C) 7→ FM(A, B, C) is monotone. Specifically, if A ≥ A0 , B ≥ B0 , and C ≥ C0 , then FM(A, B, C) ≥ FM(A0 , B0 , C0 ) in the positive semidefinite ordering. (3) V3: self-duality. FM(A, B, C) = FM(A−1 , B −1 , C −1 )−1 . (4) V4: determinant identity. det FM(A, B, C) = (det A · det B · det C)1/3 . In fact, Thm. 159 holds true for any finite number of SPD matrices. Besides, the geodesic distance induced by ALEMs has similarity invariance. Proposition 160 (Similarity Invariance). [↓] The geodesic distance under ALEM is similarity invariant. In other words, let R ∈ SO(n) be a rotation matrix and let s ∈ R++ be a scale factor. Given any two SPD matrices S1 and S2 , we have dALE (S1 , S2 ) = dALE (s2 RS1 R⊤ , s2 RS2 R⊤ ).
(6.29)
Let us explain a bit more about the above three kinds of invariance. First, among metrics on Lie groups, bi-invariant metrics are the most convenient ones [187, Ch. V]. Second, exponential invariance offers a fast computation for Fréchet means under exponential scaling. Finally, similarity invariance is significant for describing frequently encountered covariance matrices [9].
The above discussion focuses on the theoretical perspective. Now, let us reconsider Eq. (6.12) in a numerical way. 194
Chapter 6. Fast and Stable Geometries on SPD Manifolds Proposition 161. [↓] logα can be rewritten as logα (S) = U logα (Σ)U ⊤ ,
(6.30)
= U A log(Σ)U ⊤ ,
(6.31)
log(Σ) ⊤ U , B
(6.32)
=U
where X is the diagonal division, B = diag (log(a1 ), · · · , log(an )), and A = IBn . Y Based on the above proposition, more analyses could be carried out from a numerical point of view. First, logα (·) can balance the eigenvalues of an input SPD matrix S by exploiting different bases for different eigenvalues. In Riemannian algorithms, manifoldvalued features usually contain vibrant information. We expect that by the above adaptation, manifold-valued data could be better fitted and the learning ability of algorithms could be further promoted. Remark 162. Note that the discussion in Sec. 6.2.2.3 and Sec. 6.2.2.5 can also be readily transferred to LCM, generating an adaptive version of LCM.
6.2.3
Parameter Learning
As shown in Thm. 152, Riemannian computations under ALEM are built upon the general matrix logarithm logα and its inverse. Accordingly, this subsection studies the backpropagation and parameter optimization of these two maps. 6.2.3.1
Gradient Computation
For both the general matrix logarithm and exponential, gradients are required with respect to their parameters and inputs. Since logα involves a structured matrix decomposition, the following derivations rely heavily on structured-matrix backpropagation (BP) [110], whose key idea is invariance of the first-order differential form. The general matrix logarithm is a special case of eigenvalue functions. Based on the formula given by Bhatia [20] and the matrix BP techniques presented by Ionescu et al. [110], we can obtain all the gradients in the following propositions. d Proposition 163. [↓] Let us denote X = logα (S), where S ∈ S++ is the input SPD matrix. The input gradient ∇S L is obtained by specializing the Daleckii– Krein expression in Eqs. (2.90) and (2.91) with V = ∇X L and f (σi ) = Aii log(σi ).
195
6.2. Adaptive Log-Euclidean Metrics Name
Detail
Constraint
Method
RELU MUL DIV
Optimizing base vector α (Eq. (6.30)) Optimizing diagonal elements of A (Eq. (6.31)) Optimizing diagonal elements of B (Eq. (6.32))
Positive Unconstrained Unconstrained
shift-ReLU max(ϵ, α) Standard BP Standard BP
Table 6.1: Parameter learning for the general matrix logarithm and exponential.
The parameter gradient is ∇A L = [U ⊤ (∇X L)U ] ⊛ log(Σ),
(6.33)
where S = U ΣU ⊤ is the eigendecomposition of an SPD matrix and σ1 , . . . , σd are the diagonal entries of Σ. As the inverse map corresponding to Eq. (6.31), the general matrix exponential log−1 α can be rewritten as ⊤ Σ11 Σnn log−1 U α (X) = U diag a1 , · · · , an Σ U ⊤, = U exp A
(6.34)
where X = U ΣU ⊤ ∈ S n is an eigendecomposition of X. Following Thm. 163, we obtain the backpropagation of log−1 α . d Proposition 164. [↓] Let us denote X = log−1 α (S) with S ∈ S . The input gradient ∇S L is obtained by specializing the Daleckii–Krein expression in Eqs. (2.90) σi and (2.91) with V = ∇X L and f (σi ) = e Aii . The parameter gradient is
Σdd −Σ Σ11 ∇A L = [U (∇X L)U ] ⊛ diag a1 , · · · , ad , A2 ⊤
(6.35)
where S = U ΣU ⊤ is the eigendecomposition of a symmetric matrix and σ1 , . . . , σd are the diagonal entries of Σ. 6.2.3.2
Parameter Updates
The general matrix logarithm and exponential share the same base parameters. Let us focus on the former. Let the input SPD matrix S have dimension d × d. Recalling Eqs. (6.30) to (6.32), there are three ways to implement parameter learning. We could learn the base vector α in Eq. (6.30), diagonal matrix A in Eq. (6.31), or diagonal 196
Chapter 6. Fast and Stable Geometries on SPD Manifolds matrix B in Eq. (6.32), respectively. For learning A in Eq. (6.31) or B in Eq. (6.32), since the parameters (diagonal elements) lie in a Euclidean space Rd , the optimization can be easily integrated into the BP algorithm. We call learning A MUL and learning B DIV. For learning α in Eq. (6.30), each element a of α satisfies a > 0 and a ̸= 1. Since the equality case can be avoided by setting a = 1 + ϵ whenever a = 1, where ϵ ∈ R++ , it remains to enforce positivity during optimization. We consider two strategies for doing so. The first strategy applies the shift-ReLU max(ϵ, a) to an unconstrained parameter. We call this strategy RELU. Other transformations, such as squaring the parameter, are also feasible, but we focus on RELU. The second strategy takes a geometric approach by viewing a as a point on a one-dimensional SPD manifold and optimizing it using the Riemannian optimization strategy reviewed in Sec. 2.7. We call this strategy GEOM. Its RSGD update is given in the following proposition. Proposition 165. [↓] Viewing a positive scalar a as a point in a one-dimensional SPD manifold, we have the following RSGD update formula. a(t+1) = a(t) e−γ
(t) a(t) ∇
a(t)
L
,
(6.36)
where ∇a(t) L is the Euclidean gradient of L with respect to a at a(t) , γ (t) is the learning rate, and e(·) is the natural exponential function. Moreover, the following proposition shows that GEOM is equivalent to DIV. Proposition 166. [↓] For parameter learning in logα , optimizing the base vector α by RSGD is equivalent to optimizing the divisor matrix B by Euclidean stochastic gradient descent (ESGD). Consequently, the three distinct update schemes are RELU, DIV, and MUL, as summarized in Tab. 6.1.
6.2.4
Experiments
In this section, we validate the efficacy of our approaches on multiple data sets. Riemannian metrics are foundational to Riemannian neural networks. Therefore, our ALEM can redesign basic blocks in Riemannian neural networks. Beyond the main SPDNet experiments, we apply our ALEM to other Riemannian building blocks, including the LieBN framework in Chapter 3, Riemannian residual blocks [117], and Riemannian classifiers [159]. More details on data sets and experimental settings are provided in 197
6.2. Adaptive Log-Euclidean Metrics Learning rate
1e−2
5e−2
Architecture
{ 93, 30}
{ 93, 70, 30}
{ 93, 70, 50, 30}
{ 93, 30}
{ 93, 70, 30}
{ 93, 70, 50, 30}
SPDNet SPDNetBN ALog-MUL ALog-DIV ALog-RELU
62.92±0.81 63.03±0.75 63.52±0.75 63.60±0.79 63.02±0.79
62.87±0.60 58.27±1.7 63.86±0.58 63.93±0.52 63.94±0.64
63.03±0.67 52.02±2.34 63.94±0.44 63.81±0.7 63.14±0.65
63.89±0.73 63.75±0.69 64.4±0.68 64.81±0.64 63.97±0.75
64.00±0.65 48.78±5.15 64.60±0.69 64.84±0.65 64.10±0.63
63.72±0.61 37.84±6.10 64.36±0.49 64.80±0.36 63.78±0.46
Table 6.2: Results of ALog on the HDM05 data set. The best results are bold. Secs. A.1 and A.3.7. 6.2.4.1
Applications in SPDNet
In existing SPD neural networks, activation and classification layers commonly map SPD features into the logarithmic domain through the matrix logarithm [106, 232, 37, 156, 47]. This mapping is an isomorphism that identifies the SPD manifold under LEM with the Euclidean space S n . Replacing the natural matrix logarithm log with the learnable general logarithm logα allows the resulting layer to adapt the underlying geometry to the learned SPD features. We focus on the classic SPDNet [106], whose BiMap, ReEig, and LogEig layers are reviewed in Sec. A.2.1. Specifically, replacing the matrix logarithm in its LogEig layer with the learnable general matrix logarithm logα yields the Adaptive Logarithm (ALog) layer. We compare SPDNet, SPDNetBN, and the ALog-RELU/MUL/DIV variants on HDM05, FPHA, and AFEW. On the three data sets, the numbers of training epochs are set to 200, 500, and 100. We verify our ALog on SPDNet with various architectures. In addition, we test the robustness of the proposed layer against different learning rates on the HDM05 and FPHA data sets. Generally speaking, among all three implementations, ALog-MUL shows the most robust performance gain and achieves consistent improvement over the vanilla matrix logarithm. We also observe that ALog-MUL is comparable to or even better than SPDNetBN, which, however, introduces substantially greater complexity than our approach. The main reason for the superiority of our ALog against the vanilla matrix logarithm is that our ALog can adaptively respect the vibrant geometry of SPD manifolds, depending on the characteristics of data sets, while only LEM can be respected by the matrix logarithm. The following are detailed observations and analyses. Results on the HDM05 data set. The 10-fold results are presented in Tab. 6.2, where the data split and weight initialization are randomized. Following Huang and 198
Chapter 6. Fast and Stable Geometries on SPD Manifolds
SPDNet SPDNet-ALog-MUL
80
Acc
60 40 20 0
0
100
200 300 Training epoch
400
500
Figure 6.1: Accuracy curves on the FPHA data set. SPDNet
SPDNetBN
85.73±0.80
86.83±0.74
ALog MUL
DIV
RELU
87.8±0.71
88.07±1.13
86.65±0.68
Table 6.3: Results of ALog on the FPHA data set. Van Gool [106], three architectures are implemented on this data set, i.e., { 93, 30}, { 93, 70, 30}, and { 93, 70, 50, 30}. Generally speaking, endowed with ALog, SPDNet achieves consistent improvement. Among all three implementations, RELU only brings limited improvement. The reason might be that RELU fails to respect the innate geometry of the positive constraint. There is another interesting observation worth mentioning. In Brooks et al. [31], only the result of SPDNetBN under the architecture of {93, 30} is reported on this data set. Our experiments show that with the network going deeper, SPDNetBN tends to collapse, while our ALog layer performs robustly in all settings. Results on the FPHA data set. We validate our approach on this data set, with a learning rate of 1e−2 , over 10-fold cross-validation on random initialization. Since our experiments indicate that the vanilla SPDNet is already saturated with 1 BiMap layer, we just report the results on the architecture of {63, 33}, which are presented in Tab. 6.3. Although DIV performs best on this data set, it presents the largest variance. There is an underlying nonlinear scaling mechanism in the update of DIV, which might undermine its robustness. Without loss of generality, let us focus on a single scalar parameter b in Eq. (6.32). The ultimate factor multiplied by the plain logarithm is 1/b. 199
6.2. Adaptive Log-Euclidean Metrics Therefore, the change of the multiplier after the update would be 1/(b − ∆) − 1/b = ∆/[(b − ∆)b].
(6.37)
Eq. (6.37) will scale the original ∆ to some extent. This scaling mechanism might undermine the robustness of the ALog layer. However, ALog-MUL achieves robust improvement and even surpasses SPDNetBN. This again demonstrates the significance of our adaptive mechanism for Riemannian deep networks. Finally, in terms of convergence analysis, accuracy curves with and without ALog are also reported in Fig. 6.1. Results on the AFEW data set. On this data set, Depth 1 2 3 4 the learning rate is 5e−2 , and we SPDNet 48.53 46.89 48.24 47.22 validate our method under four SPDNetBN 46.89 46.65 47.62 48.35 ALog-MUL 48.57 48.13 49.45 50.62 network architectures, i.e., {512, ALog-DIV 48.42 48.02 48.13 49.89 100}, {512, 200, 100}, {512, 400, ALog-RELU 48.06 47.25 48.86 48.1 200, 100}, and {512, 400, 300, 200, 100}. Note that, on this Table 6.4: Results of ALog on the AFEW data set. data set, SPDNetBN tends to present relatively large fluctuations in performance, so we compute the median of the last ten epochs. On various architectures, consistent improvement can be observed when SPDNet is endowed with our ALog. In addition, MUL performs best among all three implementations. Another interesting observation is that SPDNetBN seems ineffective on these deep features, while our methods show consistently superior performance, most notably for ALog-MUL. This indicates that our adaptive layer maintains effectiveness when applied to covariance matrices from deep features. Model complexity. Our ALog manifests the same complexity, no matter how it is optimized. Without loss of generality, the discussion below focuses on ALog-MUL. The extra computation and memory costs caused by the ALog layer are minor. It only depends on the final dimension of the network. Let us take the deepest one on the AFEW data set as an example. Our ALog only brings 100 unconstrained scalar parameters, while SPDNetBN needs an SPD matrix parameter for each RBN layer. The total number of parameters in RBN layers is 4002 + 3002 + 2002 , which is much larger than ours. In addition, SPDNetBN needs to store the running mean of SPD matrices in every RBN layer, while our ALog only needs to store a vector. In terms of computation, the extra cost of our ALog is secondary as well. The forward and backward computation of our ALog is generally the same as the plain matrix logarithm, while computation in 200
Chapter 6. Fast and Stable Geometries on SPD Manifolds 1.25
SPDNet-ALog-MUL-[93, 30] SPDNet-ALog-MUL-[93, 70, 30] SPDNet-ALog-MUL-[93, 70, 50, 30]
1.20
1.2 1.1 1.0 Acc
Acc
1.15 1.10
0.8 0.7
1.05 1.00
0.9
0.6 0.5 0
5
10 15 20 Diagonal elements of A
25
30
0
(a) HDM05.
5
SPDNet-ALog-MUL-[63, 33] SPDNet-ALog-MUL-[63, 53, 33] SPDNet-ALog-MUL-[63, 53, 43, 33] 10 15 20 25 30 Diagonal elements of A
(b) FPHA.
Figure 6.2: Visualization of parameters in the ALog layer on the HDM05 and FPHA data sets. Data Set Architecture
{93, 30}
HDM05 {93, 70, 30}
{93, 70, 50, 30}
FPHA {63, 33}
SPDNet-Log2 SPDNet SPDNet-Log10 SPDNet-ALog-MUL
63.93±0.81 63.89±0.73 63.45±0.33 64.4±0.68
63.54±0.50 64.00±0.65 63.8±0.71 64.60±0.69
63.98±0.63 63.72±0.61 63.64±0.64 64.36±0.49
86.65±0.67 85.73±0.80 78.42±0.77 87.8±0.71
Table 6.5: Results of fixed bases on the HDM05 and FPHA data sets. the RBN layer is much more complex. All in all, our ALog can consistently improve the performance of SPDNet and achieve comparable or better results than SPDNetBN with much lower computation and memory costs. Visualization. We visualize the final learned parameters of the ALog layer. Since ALog-MUL is the most robust strategy, we visualize the parameters of ALog-MUL. Specifically, we plot the final values of the diagonal elements of A in Eq. (6.31) and visualize the results in Figs. 6.2a and 6.2b. We observe that the distribution of the parameters is consistent within the same data set but varies between data sets. This indicates that our approach can capture vibrant patterns in different data sets, respecting their specific geometry. Ablation studies. To further demonstrate the utility of the adaptive mechanisms in our approach, we validate the ALog layer with fixed bases. As decimal and binary are the two most common systems, we use log10 and log2 as examples of shrinking and expanding the natural logarithm log. Specifically, we set the scalar base α in logα to 10 and 2 in Eq. (6.30), respectively. We refer to the network with binary/decimal base as SPDNet-Log2/SPDNet-Log10. Note that when α = e, logα = log, and Eq. (6.30) 201
6.2. Adaptive Log-Euclidean Metrics Method
Geometry
[93, 30]
[93, 70, 30]
[93, 70, 50, 30]
None SPDNetBN SPDBN LieBN-LEM
N/A AIM AIM LEM
63.89±0.73 63.75±0.69 64.33±0.89 63.67±0.85
64.00±0.65 48.78±5.15 64.31±0.92 65.77±0.89
63.72±0.61 37.84±6.10 63.62±1.21 65.34±0.83
LieBN-ALEM
ALEM
65.24±0.71
70.11±0.96
68.86±0.72
Table 6.6: Comparison of RBN methods on the HDM05 data set. reduces to the vanilla matrix logarithm. The network is then our baseline, i.e., SPDNet. We conduct 10-fold experiments on the HDM05 and FPHA data sets and set the learning rate to 5e−2 and 1e−2 , respectively, while keeping the other settings consistent with previous experiments. The results are presented in Tab. 6.5. We observe that the fixed logarithms show similar or slightly worse results than the vanilla log, while our ALog shows consistent improvement. Besides, log10 does not converge on the FPHA data set. In fact, log10 could shrink the gradient, slowing down convergence, especially under a small learning rate. In contrast, our ALog maintains consistent effectiveness. In summary, our ALog can respect vibrant geometry induced by logα and thus benefit SPD network learning.
6.2.4.2
Riemannian Batch Normalization
The LieBN framework, its normalization guarantee, and its SPD manifestations are den veloped in Chapter 3. As shown in Thm. 152, {S++ , ⊕ALE } forms a Lie group. Besides, Thm. 157 demonstrates that ALEM is bi-invariant with respect to this group structure. Therefore, LieBN under ALEM can also normalize Riemannian sample statistics. Following the LieBN algorithm and its SPD specialization, we implement LieBN under ALEM, denoted LieBN-ALEM. We compare it against AIM-based SPDNetBN [31] and SPDBN [124], and against LieBN under LEM. Following previous work [124, 31], we adopt the SPDNet backbone. Tab. 6.6 presents the 10-fold average results on the HDM05 data set under different network architectures. Our LieBN-ALEM achieves the best performance among the RBN methods. In particular, the AIM-based SPDNetBN brings worse performance under deeper architectures. In contrast, our LieBN-ALEM can consistently improve the performance across different architectures. Besides, compared with LieBN-LEM, our LieBN-ALEM shows better performance, demonstrating the effectiveness of our ALEM. 202
Chapter 6. Fast and Stable Geometries on SPD Manifolds 6.2.4.3
Riemannian Residual Blocks
The general RResNet construction and SPD residual block are reviewed Method HDM05 NTU60 in Sec. A.2.5. Since its RiemanSPDNet 63.89±0.73 45.90±1.11 nian exponential is metric-dependent, RResNet-AIM 63.82±0.58 45.22 ± 1.23 RResNet-LEM 66.51±0.93 48.73±0.60 the Riemannian residual block under ALEM is obtained by substitutRResNet-ALEM 69.03±1.06 57.09±0.59 ing Eq. (6.17) into that construcTable 6.7: Experiments on RResNet under diftion. The required backpropagation ferent geometries. of log−1 α is derived in Thm. 164. Following Katsman et al. [117], we compare RResNet under different geometries on the HDM05 and NTU60 data sets. Tab. 6.7 reports the 10-fold and 5-fold average results on these data sets. Compared with the vanilla SPDNet, RResNet-AIM brings little improvement, while LEM and ALEM show much better performance. In particular, the ALEM-based RResNet can bring a clear performance improvement, underscoring the effectiveness of our ALEM. 6.2.4.4
Riemannian Classifiers
Euclidean MLR, which consists of FC and softmax, has become a stanLearning rate 1e−2 5e−2 dard classification block in Euclidean GyroMLR-AIM 54.28±0.47 41.41±0.71 neural networks. Inspired by this, GyroMLR-LCM 42.68±0.88 42.06±0.49 GyroMLR-LEM 53.22±0.47 39.62±1.30 Nguyen and Yang [159] extended EuGyroMLR-ALEM 56.21±0.39 51.65±0.44 clidean MLR to the SPD manifold using gyrostructures [157] for intrinTable 6.8: Comparison of Gyro MLRs on the sic classification, referred to as gyro NTU60 data set. MLR. Three gyro MLRs under LCM, AIM, and LEM were introduced by Nguyen and Yang [159]. Following the logic in Nguyen and Yang [159, Sec. 2.4.2], we can obtain the gyro MLR under ALEM. n and C classes, the Theorem 167 (Gyro MLR). [↓] Given an SPD feature S ∈ S++ SPD gyro MLR under ALEM computes the multinomial probability of each class:
p(y = k | S) ∝ exp
hD
logα (S) − logα (Pk ), (logα )∗,Pk (Ãk ) 203
Ei
,
(6.38)
6.3. Product Cholesky Metrics
n n where k ∈ {1, . . . , C}, Pk ∈ S++ , and Ãk ∈ TPk S++ . n Following Chapter 4, we set Ãk = PTIn →Pk (Ak ) with Ak ∈ TIn S++ . Therefore, the RHS of Eq. (6.38) becomes
hD Ei exp logα (S) − logα (Pk ), (logα )∗,In (Ak ) .
(6.39)
As (logα )∗,In (Ak ) ∈ T0 S n ∼ = S n , we view (logα )∗,In (Ak ) as the parameter.
We use SPDNet as the backbone. We compare Gyro MLR under our ALEM with those under LEM, LCM, and AIM on the NTU60 data set. Tab. 6.8 presents the 5-fold average results under different learning rates. Our ALEM outperforms the other metrics within the gyro MLR framework. When the learning rate is 5e−2 , our GyroMLRALEM shows a larger performance advantage, especially compared with GyroMLRLEM. These results demonstrate that Riemannian networks can benefit from the adaptivity of our ALEM.
6.3
Product Cholesky Metrics
6.3.1
Introduction
Whereas Sec. 6.2 introduces adaptive flexibility through pullback Euclidean geometry, this part pursues a complementary route centered on computational efficiency and numerical stability. LCM provides the natural bridge between these routes. It combines simple closed-form Riemannian operators with the fast and stable computation of the Cholesky decomposition [137]. LCM is induced by the Cholesky decomposition from a Riemannian metric on the Cholesky manifold, which is the space of lower triangular matrices with positive diagonal entries. We refer to this source metric as the diagonal log metric. Its interpretation as a pullback Euclidean metric was established in Thm. 147. We reveal a simple product structure underlying the diagonal log metric: a Euclidean metric on the strictly lower triangular part together with n copies of a Riemannian metric on R++ for the diagonal part. This observation opens up a principled design space, as any metric on R++ can induce a metric on the Cholesky manifold, and further yield a corresponding metric on the SPD manifold via the Cholesky decomposition. Building on this product structure, we introduce two Cholesky metrics, the diagonal power metric and the diagonal Bures–Wasserstein metric, which induce two SPD metrics via the Cholesky decomposition: the Power-Cholesky Metric (PCM) and 204
Chapter 6. Fast and Stable Geometries on SPD Manifolds Bures–Wasserstein–Cholesky Metric (BWCM). Unlike the diagonal logarithm in LCM, our metrics rely on diagonal powers, improving numerical stability by avoiding exponentiation and logarithms. We further define in the Cholesky factors of SPD matrices a diagonal power deformation, which continuously connects existing and new metrics. As power θ → 0, the deformed metric converges to LCM, while at θ = 1 it recovers our proposed metrics, thereby offering a tunable trade-off. All proposed SPD metrics admit closed-form Riemannian operators, including geodesics, logarithmic and exponential maps, parallel transport, as well as gyrovector operators [200], which extend vector addition and scalar multiplication into manifolds. These operators make our metrics directly applicable to SPD neural networks. In particular, by substituting these operators into the Riemannian MLR formulation developed in Chapter 4 and residual blocks [118], we directly obtain SPD MLR classifiers and residual blocks under our metrics. We validate our metrics with experiments on SPD neural networks, numerical stability analyses, and tensor interpolation, showing the effectiveness, efficiency, and robustness of our metrics. In summary, our main contributions are: (1) Revealing the underlying simple product structure in the Cholesky manifold; (2) Proposing two Cholesky metrics and their SPD counterparts, PCM and BWCM, which admit fast and stable closed-form operators; (3) Developing SPD classifiers and residual blocks based on our metrics for SPD neural networks.3 Outline. Sec. 6.3.2 recalls the Cholesky geometry used in this part. Sec. 6.3.3 develops the Cholesky product geometries, and Sec. 6.3.4 derives their SPD counterparts. Their applications to SPD neural networks and the experimental evaluation are presented in Secs. 6.3.5 and 6.3.6. Proofs are deferred to Sec. B.10.
6.3.2
Preliminaries
Pullback metrics are defined in Thm. 33, while the commonly used SPD geometries are summarized in Tabs. 2.5 and 2.6. We further recall the Cholesky geometry. The Euclidean space of n × n lower triangular matrices is denoted LTn . Its open subset, whose diagonal elements are all positive, is denoted by Ln++ . The Cholesky space Ln++ forms a submanifold of LTn [137]. For a Cholesky matrix L ∈ Ln++ and tangent vectors X, Y ∈ TL Ln++ , the Riemannian metric on the Cholesky manifold, referred to as the 3
The code is available at https://github.com/GitZH-Chen/PCM_BWCM.
205
6.3. Product Cholesky Metrics diagonal log metric, is gLDL (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨L−1 X, L−1 Y⟩.
(6.40)
Here, ⌊X⌋ and ⌊Y ⌋ are the strictly lower triangular parts of X and Y , while L, X, and Y are the diagonal matrices formed from their diagonal entries. LCM is the pullback metric of g DL by the Cholesky decomposition. As shown in Thm. 147, the diagonal log metric is the pullback metric, by the diagonal log map, of the Euclidean metric over LTn , which rationalizes our nomenclature.
6.3.3
Product Geometries on the Cholesky
We first unveil the product structure beneath the existing diagonal log metric on the Cholesky manifold. Based on this, we propose two novel Cholesky metrics. 6.3.3.1
Disentangling the Cholesky Geometry
We denote the space of n × n diagonal matrices with positive diagonal elements by Diag+ (n) and the space of n × n strictly lower triangular matrices by LT0 (n). Then, LT0 (n) is a Euclidean space, and Diag+ (n) ∼ = (R++ )n is an open submanifold of Rn . Recalling Eq. (6.40), it is defined separately on LT0 (n) and Diag+ (n). Besides, Diag+ (n) can be identified as the product of n copies of R++ . The above discussion implies a product structure. We denote the standard Euclidean metric over LT0 (n) by g E and define the Riemannian metric on R++ as gpR++ (v, w) = p−2 vw,
∀p ∈ R++ and v, w ∈ Tp R++ .
(6.41)
Then, Ln++ is the product manifold of LT0 (n) and n copies of R++ : n
}| { z 0 R++ R++ n DL E } × · · · × {R++ , g }. {L++ , g } = {LT (n), g } × {R++ , g 6.3.3.2
(6.42)
Product Geometries on the Cholesky
The following definition characterizes the underlying product structure in Eq. (6.42). 0
Definition 168 (Product geometries). Suppose g LT is a Euclidean inner product on LT0 (n) and {g i }ni=1 are Riemannian metrics on R++ . Then, the weighted product 206
Chapter 6. Fast and Stable Geometries on SPD Manifolds P 0 metric g on Ln++ is defined as gL (X, Y ) = g LT (⌊X⌋, ⌊Y ⌋) + ni=1 αi gLi ii (Xii , Yii ), with L ∈ Ln++ , X, Y ∈ TL Ln++ , and αi > 0. Here, ⌊X⌋ and ⌊Y ⌋ are the strictly lower triangular parts of X and Y , while Lii , Xii , and Yii are the i-th diagonal elements. 0
For simplicity, we focus on the case where g LT = g E is the standard Euclidean metric, all αi are equal to 1, and all g i are identical. Since R++ can be viewed as 1 a one-dimensional SPD manifold S++ , the Riemannian metrics reviewed in Sec. 2.9.1 can be immediately used to build Riemannian metrics on the Cholesky manifold. We additionally consider the Generalized Bures–Wasserstein Metric (GBWM), which is 1 1 n [93]. For clarity, we denote the pullback of BWM by S 7→ M − 2 SM − 2 for S, M ∈ S++ PEM and GBWM by θ-EM and M -BWM, respectively. Simple computations show 1 that AIM, LEM, and LCM coincide with Eq. (6.41) on S++ . Therefore, these metrics 1 reduce to three classes on S++ : (1) LEM, LCM, or AIM; (2) θ-EM; (3) M -BWM or BWM. When the metric on R++ is AIM (LEM or LCM), the resulting product metric on the Cholesky manifold is the diagonal log metric, and the pullback SPD metric via the Cholesky decomposition is exactly LCM. Inspired by the above analysis, we obtain two new metrics on the Cholesky manifold by setting each g i in Thm. 168 to θ-EM and M -BWM, termed the Diagonal Power Metric (θ-DPM) and Diagonal Bures–Wasserstein Metric (M-DBWM with M ∈ Diag+ (n)), respectively. By product geometries [131], we can obtain closed-form expressions for their Riemannian operators, such as the geodesic, logarithmic map, parallel transport, and weighted Fréchet mean. These operators are particularly important for building concrete learning algorithms [227, 132, 31, 141]. Theorem 169 (θ-DPM). [↓] Let L, K ∈ Ln++ and X, Y ∈ TL Ln++ , and let {Li ∈ PN N Ln++ }N i=1 have weights {wi }i=1 satisfying wi > 0 for all i and i=1 wi = 1. Then, the Riemannian operators under θ-DPM with θ ̸= 0 are gLθ-DE (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨Lθ−1 X, Lθ−1 Y⟩, 1 γ(L,X) (t) = ⌊L⌋ + t⌊X⌋ + L In + tθL−1 X θ , i θ 1 h LogL (K) = ⌊K⌋ − ⌊L⌋ + L L−1 K − In , θ 1−θ −1 PTL→K (X) = ⌊X⌋ + L K X, 1 d2 (L, K) = ∥⌊K⌋ − ⌊L⌋∥2F + 2 ∥Kθ − Lθ ∥2F , θ 207
(6.43) (6.44) (6.45) (6.46) (6.47)
6.3. Product Cholesky Metrics
WFM({wi }, {Li }) =
X
i
wi ⌊Li ⌋ +
X
i
wi Lθi
θ1
,
(6.48)
where ∥·∥F is the Frobenius norm. X, Y, L, K, and Li are diagonal matrices with diagonal elements from X, Y , L, K, and Li . γ(L,X) (t) denotes the geodesic starting at L with initial velocity X. PTL→K (·) is the parallel transport along the geodesic connecting L and K. Log, d, and WFM are the Riemannian logarithm, geodesic distance, and weighted Fréchet mean, respectively. Note that γ(L,X) (t) is locally defined in {t ∈ R|L + tθX ∈ Diag+ (n)}. Theorem 170 (M-DBWM). [↓] Following the notation in Thm. 169, the Riemannian operators under M-DBWM with M ∈ Diag+ (n) are 1 gLM-DBW (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨L−1 X, M−1 Y⟩, 4 2 1 −1 γ(L,X) (t) = ⌊L⌋ + t⌊X⌋ + L In + t L X , 2 h i 1 LogL (K) = ⌊K⌋ − ⌊L⌋ + 2L L−1 K 2 − In , 1 PTL→K (X) = ⌊X⌋ + L−1 K 2 X, 1 1 2 2 − 12 2 2 d (L, K) = ∥⌊K⌋ − ⌊L⌋∥F + ∥M K − L ∥2F , !2 X X 1 WFM({wi }, {Li }) = wi ⌊Li ⌋ + wi Li2 , i
(6.49) (6.50) (6.51) (6.52) (6.53) (6.54)
i
where the geodesic γ(L,X) (t) is locally defined in {t ∈ R | L + 2t X ∈ Diag+ (n)}. When M = In in M-DBWM, the resulting metric is denoted by DBWM. On the SPD manifold, GBWM is locally AIM [93]. Similarly, on the Cholesky manifold, our M-DBWM is locally the diagonal log metric at L ∈ Ln++ : gLL-DBW (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + 41 ⟨L−1 X, L−1 Y⟩. GBWM on the SPD manifold generally has no closedform expression for the Fréchet mean [22]. Moreover, the closed-form expression of parallel transport under BWM is known only if two SPD matrices commute [196]. In contrast, all these operators have closed-form expressions under M-DBWM on Ln++ . 6.3.3.3
Deformed Cholesky Metrics
On SPD manifolds, metrics deformed by the matrix power can interpolate between a given metric and an LEM-like metric [194, Sec. 3.1]. Inspired by this, we define 208
Chapter 6. Fast and Stable Geometries on SPD Manifolds a diagonal power deformation on the Cholesky manifold. For θ ̸= 0, we denote the diagonal power by DPowθ : Diag+ (n) ∋ P 7−→ Pθ ∈ Diag+ (n). We will show how our proposed metric is connected to the existing diagonal log metric by DPowθ . Definition 171. Let {Ln++ , g} = {LT0 (n), g E }×{Diag+ (n), g̃} be a product metric and θ ̸= 0. We define the diagonal-power-deformed metric of g as {Ln++ , g θ } = {LT0 (n), g E } × {Diag+ (n), θ12 DPow∗θ g̃}. The following lemma shows that g θ in Thm. 171 converges to a diagonal-log-like metric as θ → 0. Lemma 172. [↓] Given L ∈ Ln++ and X, Y ∈ TL Ln++ , g θ in Thm. 171 satisfies gLθ (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + g̃Lθ Lθ−1 X, Lθ−1 Y −→ ⟨⌊X⌋, ⌊Y ⌋⟩ + g̃In (L−1 X, L−1 Y). θ→0 (6.55) Now, we discuss the deformation of the diagonal log metric, θ-DPM, and M-DBWM. First, Eq. (6.55) indicates that the diagonal-power-deformed metric of the diagonal log metric is itself. Second, θ-DPM is the diagonal-power-deformed metric of the Euclidean metric on the Cholesky manifold. Besides, θ-DPM interpolates between the diagonal log metric (θ → 0) and the Euclidean metric (θ = 1). Third, the diagonal-power-deformed metric of M-DBWM, referred to as (θ, M)-DBWM, is (θ,M)-DBW
gL
1 (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨Lθ−2 X, M−1 Y⟩. 4
(6.56)
When M = In , the deformed metric of DBWM, i.e., θ-DBWM, tends to be a scaled diagonal log metric as θ → 0: 1 lim gLθ-DBW (X, Y ) = ⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨L−1 X, L−1 Y⟩. θ→0 4
(6.57)
As (θ, M)-DBWM is the pullback metric by diagonal power and scaled by a constant, the Riemannian operators also have closed-form expressions, which are discussed in Sec. A.4.5.1. 6.3.3.4
Algebraic Structures
As reviewed in Sec. 2.6, Eqs. (2.77) and (2.78) define gyroaddition and scalar gyromultiplication on a Riemannian manifold. The gyro operations under the diagonal log metric reduce to vector-space operations, as the metric is induced by the Euclidean 209
6.3. Product Cholesky Metrics metric over the lower triangular matrices. This subsection studies the gyro-structures over θ-DPM and (θ, M)-DBWM. Let the identity matrix be the origin and C be θ-DPM or (θ, M)-DBWM. We have the following. Lemma 173 (Gyro-structures). [↓] For L, K ∈ Ln++ and t ∈ R, the gyro operations are 1 L ⊕C K = ⌊L⌋ + ⌊K⌋ + Lβ + Kβ − In β , 1 t ⊙C L = t⌊L⌋ + tLβ + (1 − t)In β ,
(6.58) (6.59)
where β = θ for θ-DPM, and β = θ/2 for (θ, M)-DBWM. ⊕C requires L and K to satisfy Lβ + Kβ − In ∈ Diag+ (n), while ⊙C requires (1 − t)In + tLβ ∈ Diag+ (n). These gyro operations are defined only under the assumptions in Thm. 173, which arise from the locally defined Riemannian exponential map. Throughout, we impose these assumptions implicitly. Theorem 174. [↓] When the gyro operations are well defined under the conditions in Thm. 173, {Ln++ , ⊕C } satisfies all the axioms of gyrocommutative gyrogroups (Thms. 51 and 52), and {Ln++ , ⊕C , ⊙C } satisfies all the axioms of gyrovector spaces (Thm. 53). Corollary 175. The identity element of {Ln++ , ⊕C } is the identity matrix, i.e., 1 ∀L ∈ Ln++ , In ⊕C L = L. The inverses are ⊖C L = −1 ⊙C L = −⌊L⌋ + 2In − Lβ β , for L ∈ {L ∈ Ln++ | 2In − Lβ ∈ Diag+ (n)}. Remark 176. The gyrostructures on the Grassmannian have shown success in building Riemannian algorithms [157, 159]. Like our gyrostructure, the gyrostructures of the Grassmannian also require some assumptions for well-definedness [157, Sec. 3.2]. We therefore examine the well-definedness of the gyrostructures under our metrics. In practice, such positivity constraints can be remedied by numerical techniques. Taking 2In − Lβ ∈ Diag+ (n) as an example, one can use di ← max(di , ε) for each diagonal element di with a small constant ε > 0. 6.3.3.5
Numerical Advantages over Diagonal Log Metric
Tab. 6.9 summarizes all the Riemannian and gyro operators. The Riemannian operators under (θ, M)-DBWM and θ-DPM are mostly computed using the diagonal power 210
Chapter 6. Fast and Stable Geometries on SPD Manifolds Operators
Diagonal Log Metric
θ-DPM
(θ, M)-DBWM
gL (X, Y )
⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨L−1 X, L−1 Y⟩
⟨⌊X⌋, ⌊Y ⌋⟩ + ⟨Lθ−1 X, Lθ−1 Y⟩
γ(L,X) (t)
⌊L⌋ + t⌊X⌋ + L exp(tL−1 X)
LogL (K)
⌊K⌋ − ⌊L⌋ + L log(L−1 K)
PTL→K (X)
⌊X⌋ + (L−1 K)X
⌊L⌋ + t⌊X⌋ + L (In + tθL−1 X) θ h i θ ⌊K⌋ − ⌊L⌋ + 1θ L (L−1 K) − In
⟨⌊X⌋, ⌊Y ⌋⟩ + 14 ⟨Lθ−2 X, M−1 Y⟩ 2 ⌊L⌋ + t⌊X⌋ + L In + t 2θ L−1 X θ h i θ ⌊K⌋ − ⌊L⌋ + 2θ L (L−1 K) 2 − In
d2 (L, K)
∥⌊K⌋ − ⌊L⌋∥2F + ∥log(K) − log(L)∥2F
WFM({wi }, {Li }) L⊕K t⊙L
P
i wi ⌊Li ⌋ + exp (
P
i wi log(Li ))
⌊L⌋ + ⌊K⌋ + LK t⌊L⌋ + Lt
1−θ
⌊X⌋ + (L−1 K)
1
X
∥⌊K⌋ − ⌊L⌋∥2F + θ12 ∥Kθ − Lθ ∥2F P
i wi ⌊Li ⌋ +
P
θ i wi Li
θ1
1 ⌊L⌋ + ⌊K⌋ + Lθ + Kθ − In θ 1 t⌊L⌋ + tLθ + (1 − t)In θ
1− θ2
⌊X⌋ + (L−1 K)
X θ θ 1 ∥⌊K⌋ − ⌊L⌋∥2F + θ12 ∥M− 2 K 2 − L 2 ∥2F P
i wi ⌊Li ⌋ +
P
θ 2
i wi Li
θ2
θ θ2 θ ⌊L⌋ + ⌊K⌋ + L 2 + K 2 − In θ θ2 t⌊L⌋ + tL 2 + (1 − t)In
Table 6.9: Riemannian and gyro operators of different metrics on the Cholesky manifold. For the diagonal log metric, log(·) and exp(·) are diagonal logarithm and exponentiation. function, while those under the diagonal log metric are computed using the diagonal logarithm or exponentiation. This indicates that our θ-DPM and (θ, M)-DBWM may have better numerical stability than the existing diagonal log metric, as logarithm or exponentiation might overly stretch the diagonal elements compared with the power function. The gyro operations under our θ-DPM and (θ, M)-DBWM also have numerical advantages over those under the diagonal log metric. The former are based on linear operations combined with power and its inverse, causing relatively minor changes to the input magnitude, while the gyro operations under the diagonal log metric are based on products or powers, resulting in more noticeable alterations to the input magnitude.
6.3.4
Geometries on the SPD Manifold
This section discusses the Riemannian metrics on the SPD manifold via the Cholesky decomposition. We first review some basic properties of the Cholesky decomposition, followed by the SPD metrics. n The Cholesky decomposition, denoted by Chol(·) : S++ → Ln++ , is a diffeomorphism [137]. Therefore, it can pull back the Riemannian and gyro-structures from the n Cholesky manifold Ln++ to the SPD manifold S++ . We call the pullbacks of θ-DPM and (θ, M)-DBWM through the Cholesky decomposition the Power-Cholesky Metric (θ-PCM) and Bures–Wasserstein–Cholesky Metric ((θ, M)-BWCM), respectively. Then, the Riemannian operators under θ-PCM and (θ, M)-BWCM can be obtained from the properties of Riemannian isometries (Thm. 33).
Operators. Let C ∈ {θ-DPM, (θ, M)-DBWM} and S ∈ {θ-PCM, (θ, M)-BWCM}. We denote the Riemannian logarithm, exponential map, geodesic, parallel transport along the geodesic, geodesic distance, weighted Fréchet mean, gyroaddition, and scalar 211
6.3. Product Cholesky Metrics n gyromultiplication on {S++ , g S } by LogS , ExpS , γ S , PTS , dS (·, ·), WFMS , ⊕S , and ⊙S , respectively, while LogC , ExpC , γ C , PTC , dC (·, ·), WFMC , ⊕C , and ⊙C are their n n n counterparts on {Ln++ , g C }. For P, Q ∈ S++ , V, W ∈ TP S++ , and {Pi ∈ S++ }N i=1 with P N weights {wi }N i=1 satisfying wi > 0 for all i and i=1 wi = 1, we have the following Riemannian and gyro operators: −1 S γ(P,V ) (t) = Chol
C γ(L, (t) Ve )
,
LogSP (Q) = (Chol∗,P )−1 LogCL (K) , ExpSP (V ) = Chol−1 ExpCL Ve , PTSP →Q (V ) = (Chol∗,Q )−1 PTCL→K Ve , dS (P, Q) = dC (L, K),
WFMS ({wi }, {Pi }) = Chol−1 WFMC ({wi }, {Li }) , P ⊕S Q = Chol−1 (L ⊕C K), t ⊙S P = Chol−1 (t ⊙C L),
(6.60) (6.61) (6.62) (6.63) (6.64) (6.65) (6.66) (6.67)
where P = LL⊤ , Q = KK ⊤ , and Pi = Li L⊤ i are Cholesky decompositions. Here, Ve = Chol∗,P (V ) is the Cholesky differential as defined in Sec. 2.8.1. Besides, when C is the diagonal log metric, the above recovers the Riemannian and gyro-structures under LCM.
The above gyro operations, when well-defined, also satisfy the gyrovector-space axioms. n Theorem 177. [↓] {S++ , ⊕S , ⊙S } satisfies all the axioms of gyrovector spaces.
Remark 178. As discussed in Sec. 6.3.3.5, the Cholesky metrics θ-DPM and (θ, M)-DBWM are more numerically stable than the existing diagonal log metric. As pullback metrics through the Cholesky decomposition, our θ-PCM and (θ, M)-BWCM therefore preserve the advantage of numerical stability over the existing LCM. Besides, all the Riemannian operators have closed-form expressions and are easy to use, as the differential maps of the Cholesky decomposition can be easily calculated. 212
Chapter 6. Fast and Stable Geometries on SPD Manifolds
6.3.5
Applications to SPD Neural Networks
The closed-form operators derived in Sec. 6.3.4 make the proposed metrics directly applicable to SPD neural networks. In this section, we apply the proposed SPD metrics θ-PCM and (θ, M)-BWCM to build MLR classifiers and residual blocks on the SPD manifold. MLR. The Euclidean point-to-hyperplane formulation and its Riemannian extension are developed in Chapter 4. Substituting the operators of the proposed metrics into that formulation gives the following SPD MLRs. n , the C-class SPD MLRs Theorem 179. [↓] Given an input SPD matrix S ∈ S++ under θ-PCM and (θ, M)-BWCM are
⟨⌊K⌋ − ⌊Lk ⌋, ⌊Ak ⌋⟩ n , θ-PCM : p(y = k | S ∈ S++ ) ∝ exp 1 θ θ + ⟨K − Lk , Ak ⟩ 2θ ⟨⌊K⌋ − ⌊Lk ⌋, ⌊Ak ⌋⟩ n , (θ, M)-BWCM : p(y = k | S ∈ S++ ) ∝ exp θ θ 1 + ⟨K 2 − Lk2 , M−1 Ak ⟩ 4θ
(6.68)
(6.69)
where S = KK ⊤ and Pk = Lk L⊤ k are Cholesky decompositions. The parameters n n are Pk ∈ S++ and Ak ∈ LT for each class k = 1, · · · , C. Residual blocks. The general construction and the specialized SPD residual-block expression are reviewed in Sec. A.2.5. The only component that varies across metrics is the Riemannian exponential map; substituting the operators derived above therefore gives residual blocks under the proposed metrics.
6.3.6
Experiments
We first compare our metrics against the popular AIM, LEM, and LCM when building SPD MLR classifiers and residual blocks. Then, we evaluate the proposed metrics through numerical experiments. More details on data sets and experimental settings are provided in Secs. A.1 and A.3.8. 6.3.6.1
Riemannian Classifiers
Following the SPD learning setup in Sec. 4.4 and Huang and Van Gool [106], we adopt the Radar data set [31] for radar signal classification, and the HDM05 [153] and FPHA [80] data sets for human action recognition. We compare the SPD MLRs under our 213
6.3. Product Cholesky Metrics (a) Radar
(c) FPHA
(b) HDM05
Metric
Acc
Time
AIM LEM LCM
94.53 ± 0.95 93.55 ± 1.21 93.49 ± 1.25
0.80 0.76 0.72
θ-PCM θ-BWCM
95.79 ± 0.38 93.93 ± 0.79
0.72 0.71
Metric
1-Block
2-Block
3-Block
Metric
Acc
Time
85.57 ± 0.50 85.90 ± 0.47 86.37 ± 0.59
7.14 0.98 0.74
89.40 ± 0.13 86.27 ± 0.60
0.69 0.70
Acc
Time
Acc
Time
Acc
Time
AIM LEM LCM
58.07 ± 0.64 56.97 ± 0.61 60.69 ± 1.89
17.32 2.21 1.83
60.72 ± 0.62 60.69 ± 1.02 62.61 ± 1.46
18.75 2.92 2.40
61.14 ± 0.94 60.28 ± 0.91 62.33 ± 2.15
19.23 3.50 2.90
AIM LEM LCM
θ-PCM θ-BWCM
62.51 ± 1.65 62.71 ± 0.88
1.58 1.64
63.66 ± 1.30 64.52 ± 0.56
2.29 2.27
65.75 ± 2.86 67.40 ± 0.90
2.76 2.87
θ-PCM θ-BWCM
Table 6.10: SPD MLRs under different metrics on the SPDNet backbone. The best results are bold. metrics with those under AIM, LEM, and LCM in Thm. 117. Following Sec. 4.4 and Nguyen and Yang [159], we adopt SPDNet [106] and GyroSPD [159] as two backbones, both mimicking feedforward neural networks. For simplicity, we set M in (θ, M)-BWCM to the identity matrix. SPDNet. On HDM05, we further evaluate architectures with up to three transformation blocks. Tab. 6.10 reports the five-fold accuracy and training time per epoch, from which we draw the following observations. • Effectiveness. Our metrics generally yield higher accuracy than their counterparts. Notably, they outperform LCM, although both originate from the Cholesky product structure. This improvement is attributed to the fact that the diagonal logarithm and exponentiation in LCM tend to overly stretch the diagonal entries, i.e., eigenvalues of the Cholesky factors, whereas our diagonal power transformation achieves a more balanced scaling. • Efficiency. Our metrics substantially reduce computational cost compared to AIM, remain faster than LEM, and achieve efficiency comparable to LCM. Together with their superior accuracy, these results highlight the dual advantages of our approach in both effectiveness and efficiency. GyroSPD. Tab. 6.11 reports Radar HDM05 FPHA the results on the GyroSPD Metric Acc Time Acc Time Acc Time backbone, which consists of a AIM 96.80 ± 0.59 1.23 66.05 ± 1.80 21.65 85.77 ± 0.52 11.48 LEM 96.58 ± 0.27 1.18 66.42 ± 0.47 2.02 85.87 ± 0.79 1.22 single gyrotranslation layer folLCM 96.29 ± 0.53 1.12 68.37 ± 0.66 1.66 89.83 ± 0.28 0.98 θ-PCM 97.04 ± 0.64 1.18 71.93 ± 1.21 1.51 91.17 ± 0.30 1.00 lowed by an SPD MLR. Similar θ-BWCM 96.21 ± 0.25 1.05 72.74 ± 0.43 1.58 91.00 ± 0.11 0.96 to the SPDNet results, our metrics achieve comparable or supe- Table 6.11: SPD MLRs on the GyroSPD backbone. rior performance to LCM across all data sets while maintaining comparable efficiency. On HDM05 and FPHA, both proposed metrics deliver higher accuracy with lower runtime than AIM and LEM. On Radar, PCM achieves the highest accuracy, while BWCM has the lowest runtime. 214
Chapter 6. Fast and Stable Geometries on SPD Manifolds 6.3.6.2
Riemannian Residual Blocks
Riemannian ResNet (RResNet) was introRadar HDM05 FPHA duced by Katsman et al. [118]. Its backMetric Acc Time Acc Time Acc Time bone architecture largely follows SPDNet. AIM 96.4 1.02 57.01 1.14 87.33 0.72 LEM 97.07 0.81 67.52 0.52 86.17 0.32 The key difference lies in the head: while LCM 97.07 0.85 66.27 0.63 86.83 0.49 SPDNet directly applies a classification θ-PCM 97.87 0.85 68.05 0.63 88.33 0.48 layer, RResNet inserts a residual block beTable 6.12: Results on residual blocks. fore the classification head. Accordingly, we adopt the following classification head under each metric: LogIn +FC + softmax. Since the Riemannian exponential is similar for θ-PCM and (θ, M)-BWCM, we focus on θ-PCM and compare it against AIM, LEM, and LCM in constructing RResNet. Tab. 6.12 reports the best results across three trials, showing that our metric consistently achieves superior accuracy while maintaining comparable efficiency. 6.3.6.3
Numerical Stability
As discussed by Lin [137, p. 16], LCM is more stable than AIM and LEM owing to the numerical advantage of Cholesky decomposition over SVD. Moreover, as highlighted in Thm. 178, the essential distinction between our metrics and LCM lies in the diagonal operations: ours rely on diagonal power, while LCM employs diagonal exponentiation and logarithm. This structural difference grants our metrics stronger numerical stability and robustness compared with LCM, as well as AIM and LEM. To validate this, we evaluate geodesics on the Cholesky manifold. We generate 100,000 synthetic n × n Cholesky matrices L and tangent vectors X ∈ LTn , where each entry is uniformly sampled from [0, 1], and we set the smallest eigenvalue (diagonal entry) of L to ϵ. We test two representative sizes: 3 × 3 matrices, commonly used in diffusion tensor imaging [10], and 256 × 256 matrices, typical in computer vision [133, 206]. The deformation parameter θ is set to 1.5, 0.5, and 0.15. For (θ, M)-DBWM, we set M = In . As shown in Tab. 6.13, our θ-DPM and θ-DBWM remain highly stable across a wide range of ϵ. For 3 × 3 matrices, the diagonal log metric already deteriorates at ϵ = 1e−3 , with failure rates increasing rapidly as ϵ decreases. For 256 × 256 matrices, instability emerges even earlier at ϵ = 1e−1 . In contrast, our metrics remain stable down to ϵ = 1e−30 in most cases. The only exception occurs when θ = 0.15, where failures appear around ϵ = 1e−20 . This behavior is expected since both θ-DPM and θ-DBWM converge to the diagonal log metric as θ → 0, thereby inheriting its instability in this limit. Overall, 215
6.3. Product Cholesky Metrics 3 × 3 for small matrices ϵ
DLM
θ = 1.5
256 × 256 for large matrices
θ = 0.5
θ = 0.15
DPM
DBWM
DPM
DBWM
DPM
DBWM
DLM
θ = 1.5
θ = 0.5
θ = 0.15
DPM
DBWM
DPM
DBWM
DPM
DBWM
1e−1 1e−2 1e−3 1e−4 1e−5 1e−10 1e−15
0.62 5.70 51.32 94.34 99.39 100 100
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
14.29 18.48 58.35 95.02 99.47 100 100
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
1e−20 1e−21 1e−22 1e−23 1e−24 1e−25 1e−30
100 100 100 100 100 100 100
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0.002 0.03 0.25 2.26 22.98 86.34 100
100 100 100 100 100 100 100
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0 0 0 0 0 0 0
0.02 0.01 0.23 2.42 23.13 86.58 100
Table 6.13: Failure probabilities (%) of geodesics under different metrics with small eigenvalues in L ∈ Ln++ . An output matrix containing any Inf or NaN is considered a failure. Here, DLM denotes the diagonal log metric, while DPM and DBWM denote θ-DPM and θ-DBWM, respectively. these results demonstrate the superior numerical robustness of our metrics. 6.3.6.4
Tensor Interpolation
As shown by Arsigny et al. [9, 10], geodesic interpolation of SPD matrices is important in diffusion tensor imaging. This experiment illustrates geodesic interpolation under different SPD metrics. Let P = LL⊤ and Q = KK ⊤ be the Cholesky decompositions n . The geodesics connecting P and Q under θ-PCM and (θ, M)-BWCM of P, Q ∈ S++ are h 1 i θ-PCM: Chol−1 ⌊L⌋ + t(⌊K⌋ − ⌊L⌋) + Lθ + t(Kθ − Lθ ) θ , θ θ2 θ θ −1 . (θ, M)-BWCM: Chol ⌊L⌋ + t(⌊K⌋ − ⌊L⌋) + L 2 + t(K 2 − L 2 )
(6.70) (6.71)
As the geodesics under θ-PCM and (θ, M)-BWCM have similar expressions, we 3 focus on θ-PCM. Fig. 6.3 visualizes the geodesic interpolations on S++ under different metrics, including θ-EM, LEM, AIM, BWM, LCM, and θ-PCM. Tab. 6.14 presents the associated determinants of the interpolated SPD matrices. We can make the following observations. (1) The standard Euclidean metric (1-EM) exhibits a significant swelling effect, where the maximal determinant of interpolation is much larger than the determinants of the starting and end points. Although matrix power can mitigate the swelling effect, θ-EM still suffers from swelling. (2) We find that BWM also demonstrates a clear swelling effect. In contrast, LCM, 216
Chapter 6. Fast and Stable Geometries on SPD Manifolds
Figure 6.3: Geodesic interpolation of SPD matrices under different Riemannian metrics. Each 3 × 3 SPD matrix can be visualized as an ellipsoid [9]. The two endpoints are fixed across all metrics. AIM, and LEM show no swelling effect. (3) The trivial PCM (θ = 1) considerably mitigates the swelling effect compared to the Euclidean metric, but it still exhibits some level of swelling. However, by introducing Cholesky power deformation, θ-PCM effectively reduces the swelling effect. Notably, the swelling effect of θ-PCM is significantly weaker than that of θ-EM under the same θ. (4) Our PCM shows an interpolation visually similar to that under LCM. This suggests that our metric retains some practical potential of LCM but with better numerical stability (as demonstrated in Sec. 6.3.6.3). 6.3.6.5
Asymptotic Complexity
We investigate the scalability of our PCM and BWCM in comparison with five existing SPD metrics, namely AIM, LEM, LCM, PEM, and BWM. We use SPD MLR as a representative application. We first analyze the asymptotic complexity of each SPD MLR. We then complement this analysis with synthetic experiments that measure the actual runtime of a single SPD MLR training step across different matrix dimensions. 217
6.3. Product Cholesky Metrics
Metric
The determinant of the i-th interpolation 0
1
2
3
4
5
6
7
8
9
1.0-EM 0.5-EM 0.1-EM LEM AIM BWM LCM
3.07 3.07 3.07 3.07 3.07 3.07 3.07
104.86 18.67 4.25 3.1 3.1 15.32 3.1
182.09 39.93 5.42 3.14 3.14 32.04 3.14
234.38 59.53 6.38 3.17 3.17 48.14 3.17
261.35 71.96 6.96 3.2 3.2 59.22 3.2
262.64 73.79 7.05 3.24 3.24 62.09 3.24
237.86 64.14 6.62 3.27 3.27 55.33 3.27
186.64 45.01 5.75 3.31 3.31 39.93 3.31
108.61 21.73 4.6 3.34 3.34 19.98 3.34
3.38 3.38 3.38 3.38 3.38 3.38 3.38
0.1-PCM 0.5-PCM 1.0-PCM
3.07 3.07 3.07
3.15 3.35 3.6
3.23 3.59 4.07
3.29 3.79 4.46
3.34 3.91 4.72
3.37 3.97 4.83
3.39 3.94 4.76
3.4 3.83 4.49
3.4 3.64 4.03
3.38 3.38 3.38
Table 6.14: Swelling effects of geodesic SPD interpolations. Deeper greens indicate greater swelling.
Metric
Num. spectral matrix functions
Num. Cholesky decompositions
AIM LEM LCM PEM BWM PCM BWCM
1 + 2C 1+C 0 1+C 1 + 3C 0 0
0 0 1+C 0 C 1+C 1+C
Table 6.15: Number of matrix functions required per sample for a C-class SPD MLR. Spectral matrix functions include matrix logarithm, matrix power, and the Lyapunov operator.
As shown in Thms. 117 and 179, the C-class SPD MLRs under different metrics for n the input S ∈ S++ are AIM: p(y = k | S) ∝ exp
hD
log
−1 −1 Pk 2 SPk 2
, Ak
Ei
,
LEM: p(y = k | S) ∝ exp [⟨log(S) − log(Pk ), Ak ⟩] , "* +# ⌊K⌋ − ⌊Lk ⌋ + log(K) 1 LCM: p(y = k | S) ∝ exp , ⌊Ak ⌋ + Ak , 2 − log(Lk ) 1 θ θ PEM: p(y = k | S) ∝ exp S − Pk , Ak , θ D E 1 1 1 ⊤ (Pk S) 2 + (SPk ) 2 − 2Pk , LPk (Lk Ak Lk ) , BWM: p(y = k | S) ∝ exp 2 1 θ θ θ-PCM : p(y = k | S) ∝ exp ⟨⌊K⌋ − ⌊Lk ⌋, ⌊Ak ⌋⟩ + K − Lk , Ak , 2θ 218
Chapter 6. Fast and Stable Geometries on SPD Manifolds ⟨⌊K⌋ − ⌊Lk ⌋, ⌊Ak ⌋⟩ E , (θ, M)-BWCM : p(y = k | S) ∝ exp θ 1 D θ −1 2 2 K − Lk , M Ak + 4θ
n and Ak ∈ S n are MLR weights, log(·) is the matrix logarithm, and where Pk ∈ S++ LP [V ] is the solution to the matrix linear system LP [V ]P + P LP [V ] = V , known as the Lyapunov operator.
Analysis. Tab. 6.15 summarizes the number of spectral and Cholesky matrix functions required by each SPD MLR. Cholesky decomposition requires O(1/3n3 ) flops, while eigendecomposition costs O(9n3 ) flops [83, Algs. 4.2.3 and 8.3.3]. Combining these counts, Tab. 6.16 reports the resulting asymptotic per-sample complexity for each metric. Cholesky-based metrics (LCM, PCM, BWCM) are asymptotically more efficient than the eigen-based metrics (LEM, PEM, AIM, and BWM), with AIM and BWM being the slowest among the considered methods. In addition, PCM and BWCM can be practically more efficient than LCM, since diagonal powers are cheaper to compute than diagonal logarithms. Setup. To validate the asymptotic complexity in Tab. 6.16, we measure the average wallclock time of a single forward–backward training step of an SPD MLR classifier as the matrix dimension increases. The model consists of a single SPD MLR layer with 50 output classes followed by a cross-entropy loss. For each dimension n ∈ {32, 64, 128, 256, 512}, we randomly generate a batch of 30 n × n SPD matrices. In each run, we perform one forward and one backward pass and record the total runtime of this step. For PEM, we set the matrix power to 0.5.
Metric AIM LEM LCM PEM BWM PCM BWCM
Asymptotic complexity O 9(1 + 2C)n3 O 9(1 + C)n3 O 1+C n3 3 O 9(1 + C)n3 O (9(1 + 3C) + C3 )n3 O 1+C n3 3 O 1+C n3 3
Table 6.16: Asymptotic persample complexity of a C-class SPD MLR for an n × n input SPD matrix.
Results. As reported in Tab. 6.17, our PCM and BWCM are the fastest metrics across all tested dimensions, and the gap becomes particularly pronounced in the high-dimensional case. For small and medium scales (32 and 64), LCM, PCM, and BWCM have very similar runtimes and all are clearly faster than AIM, LEM, PEM, and BWM. When the dimension increases to 512, AIM and BWM require about 60 and 70 seconds per training step, whereas PCM and BWCM remain within roughly 1.7 seconds. In this setting, PCM and BWCM are even faster than LCM. 219
6.4. Conclusion Dim
AIM
LEM
LCM
PEM
BWM
PCM
BWCM
32 64 128 256 512
0.2380 1.0139 3.6256 14.5142 60.1918
0.0077 0.0395 0.1832 0.7793 3.2948
0.0046 0.0303 0.1490 0.5833 2.5030
0.0076 0.0473 0.1844 0.7853 3.4357
0.2377 1.1205 4.0674 16.5918 70.8647
0.0040 0.0251 0.1013 0.3848 1.7553
0.0040 0.0225 0.1019 0.4077 1.7526
Table 6.17: Average runtime (in seconds) of one SPD MLR training step across different matrix dimensions.
6.4
Conclusion
This chapter developed SPD metric design along two complementary routes. The first route proposed a pullback Euclidean framework and used a learnable general matrix logarithm to construct ALEM. Through this pullback construction, ALEM inherits compatible Hilbert-space, abelian Lie-group, and Riemannian structures, together with closed-form Riemannian operators and weighted Fréchet means. We further derived the differentials, gradients, and parameter-update schemes required to learn the logarithm bases. Experiments with applications to SPDNet, LieBN, Riemannian residual blocks, and gyro MLR demonstrate the effectiveness of our metrics. The second route revealed the product structure of the Cholesky manifold. This structure yielded diagonal power and diagonal Bures–Wasserstein geometries, whose deformed versions converge to the diagonal log metric as the power parameter approaches zero. Transferring these geometries through the Cholesky decomposition produced PCM and BWCM on the SPD manifold. The resulting metrics admit closed-form Riemannian and gyro operators, which directly yield SPD MLR classifiers and residual blocks. Experiments on SPD classifiers and residual networks, together with numerical experiments, supported their effectiveness, efficiency, and numerical robustness. Together, these routes show that the underlying Riemannian geometry can itself be designed to provide the flexibility, tractability, efficiency, and numerical stability required by Riemannian deep learning.
220
Chapter 7 Conclusion and Future Work 7.1
Conclusion
This thesis studied Riemannian deep learning from three connected perspectives: unified Riemannian module design across manifolds, manifold-specific Riemannian network design, and the design of the underlying Riemannian geometries. The main contributions are summarized as follows: • Chapters 3 and 4 developed unified network modules from geometric structures shared across different manifolds. The normalization chapter first developed LieBN on Lie groups, then introduced pseudo-reductive gyrogroups, a new algebraic structure that generalizes classical gyrogroups and Lie groups, and finally developed GyroBN on this foundation, with LieBN recovered as a special case. The classification development progressed from SPD MLRs obtained through exact evaluation of the point-to-hyperplane infimum under flat pullback Euclidean metrics to a general RMLR based on a Riemannian-trigonometric formulation requiring only a well-defined Riemannian logarithm. • Chapter 5 developed manifold-specific Riemannian network designs by exploiting additional structures. PVNN developed the geometry and core neural layers of the stable PV model, Hyperbolic Busemann Neural Networks (HBNN) derived intrinsic and efficient BMLR and BFC layers from Busemann functions and horospheres on the Poincaré and Lorentz models, and CorNet developed correlation MLR, FC, and convolutional layers under five geometries together with accurate Riemannian backpropagation under OLM and LSM. • Chapter 6 designed the underlying Riemannian geometries. ALEM learned pull221
7.2. Future Work back Euclidean metrics through general matrix logarithms, whereas PCM and BWCM used the product structure of the Cholesky manifold to obtain efficient and numerically stable SPD geometries. Taken together, these contributions address Riemannian deep learning at the levels of unified network modules, manifold-specific network designs, and underlying Riemannian geometries while balancing intrinsic structure, generality, computational tractability, and numerical stability.
7.2
Future Work
Two directions are particularly promising: representation learning with novel geometries and geometry-aware generative modeling. • Novel geometries for complex relational structure. Hyperbolic spaces provide an effective inductive bias for hierarchical data [164]. Although mixedcurvature products [87, 182] and matrix manifolds [59] have demonstrated the benefit of matching geometry to heterogeneous structure, real systems may also contain more complex relation types and structures. Future work could develop novel geometric structures to encode such complex relationships while retaining efficient and tractable Riemannian computations. • Geometry-aware generative modeling. Geometry arises both in latent representations and in spaces of probability distributions. At the latent level, Riemannian variational autoencoders [114], continuous normalizing flows [148], scorebased models [64], and flow matching [43] have shown how manifold geometry can shape generative dynamics, while other work has analyzed the geometry of learned latent spaces [11, 168]. At the distribution level, a growing body of work has explored the geometry of probability distributions through information geometry [56, 63, 57] and optimal transport [96, 58]. A central problem is therefore to balance the geometry of the latent space with the geometry of the distributions evolving on it.
222
Bibliography [1] P-A Absil, Robert Mahony, and Rodolphe Sepulchre. Optimization Algorithms on Matrix Manifolds. Princeton University Press, 2008. [2] Bijan Afsari. Riemannian Lp center of mass: Existence, uniqueness, and convexity. In Proceedings of the American Mathematical Society, 2011. [3] Shun-ichi Amari. Information geometry and its applications, volume 194. Springer, 2016. URL https://doi.org/10.1007/978-4-431-55978-8. [4] Roy M Anderson and Robert M May. Infectious Diseases of Humans: Dynamics and Control. Oxford University Press, 1991. [5] Tsuyoshi Ando, Chi-Kwong Li, and Roy Mathias. Geometric means. Linear algebra and its applications, 385:305–334, 2004. URL https://doi.org/10.1016/ j.laa.2003.11.019. [6] Ilya Archakov and Peter Reinhard Hansen. A new parametrization of correlation matrices. Econometrica, 89(4):1699–1715, 2021. [7] Ilya Archakov and Peter Reinhard Hansen. A canonical representation of block matrices with applications to covariance and correlation matrices. Review of Economics and Statistics, 106(4):1099–1113, 2024. [8] Marc Arnaudon, Frédéric Barbaresco, and Le Yang. Riemannian medians and means with applications to radar signal processing. IEEE Journal of Selected Topics in Signal Processing, 7(4):595–604, 2013. URL https://doi.org/10. 1109/JSTSP.2013.2261798. [9] Vincent Arsigny, Pierre Fillard, Xavier Pennec, and Nicholas Ayache. Fast and simple computations on tensors with log-Euclidean metrics. PhD thesis, INRIA, 2005. URL https://doi.org/10.1007/11566465_15. 223
Bibliography [10] Vincent Arsigny, Pierre Fillard, Xavier Pennec, and Nicholas Ayache. Geometric means in a novel vector space structure on symmetric positive-definite matrices. SIAM journal on matrix analysis and applications, 29(1):328–347, 2007. [11] Georgios Arvanitidis, Lars Kai Hansen, and Søren Hauberg. Latent space oddity: on the curvature of deep generative models. ICLR, 2018. [12] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. [13] Gregor Bachmann, Gary Bécigneul, and Octavian Ganea. Constant curvature graph convolutional networks. In ICML, 2020. [14] Frédéric Barbaresco. Gaussian distributions on the space of symmetric positive definite matrices from Souriau’s Gibbs state for Siegel domains by coadjoint orbit and moment map. In Geometric Science of Information: 5th International Conference, 2021. [15] Ahmad Bdeir, Kristian Schwethelm, and Niels Landwehr. Fully hyperbolic convolutional neural networks for computer vision. In ICLR, 2024. [16] Ahmad Bdeir, Johannes Burchert, Lars Schmidt-Thieme, and Niels Landwehr. Robust hyperbolic learning with curvature-aware optimization. In NeurIPS, 2025. [17] Gary Becigneul and Octavian-Eugen Ganea. Riemannian adaptive optimization methods. In ICLR, 2019. [18] Gary Bécigneul and Octavian-Eugen Ganea. Riemannian adaptive optimization methods. In ICLR, 2019. [19] Thomas Bendokat, Ralf Zimmermann, and P-A Absil. A Grassmann manifold handbook: Basic geometry and computational aspects. Advances in Computational Mathematics, 50(1):1–51, 2024. [20] Rajendra Bhatia. Positive Definite Matrices. Princeton University Press, 2007. [21] Rajendra Bhatia. Matrix analysis, volume 169 of Graduate Texts in Mathematics. Springer New York, 2013. doi: 10.1007/978-1-4612-0653-8. [22] Rajendra Bhatia, Tanvi Jain, and Yongdo Lim. On the Bures-Wasserstein distance between positive definite matrices. Expositiones Mathematicae, 37(2):165– 191, 2019. 224
Bibliography [23] Victoria Bloom, Dimitrios Makris, and Vasileios Argyriou. G3D: A gaming action dataset and real time action recognition evaluation framework. In CVPR Workshops, 2012. [24] Clément Bonet, Lucas Drumetz, and Nicolas Courty. Sliced-Wasserstein distances and flows on Cartan-Hadamard manifolds. JMLR, 2025. [25] Silvere Bonnabel. Stochastic gradient descent on Riemannian manifolds. IEEE Transactions on Automatic Control, 58(9):2217–2229, 2013. URL https:// arxiv.org/abs/1111.5280. [26] Silvere Bonnabel, Anne Collard, and Rodolphe Sepulchre. Rank-preserving geometric means of positive semi-definite matrices. Linear Algebra and its Applications, 438(8):3202–3216, 2013. [27] Nicolas Boumal. An introduction to optimization on smooth manifolds. Cambridge University Press, 2023. [28] Nicolas Boumal and P-A Absil. A discrete regression method on manifolds and its application to data on so (n). IFAC Proceedings Volumes, 44(1):2284–2289, 2011. [29] Martin R Bridson and André Haefliger. Metric spaces of non-positive curvature, volume 319. Springer Science & Business Media, 1999. [30] Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst. Geometric deep learning: going beyond Euclidean data. IEEE Signal Processing Magazine, 34(4):18–42, 2017. [31] Daniel Brooks, Olivier Schwander, Frédéric Barbaresco, Jean-Yves Schneider, and Matthieu Cord. Riemannian batch normalization for SPD neural networks. In NeurIPS, 2019. [32] Daniel A Brooks, Olivier Schwander, Frédéric Barbaresco, Jean-Yves Schneider, and Matthieu Cord. Exploring complex time-series representations for Riemannian machine learning of radar data. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3672– 3676. IEEE, 2019. URL https://doi.org/10.1109/ICASSP.2019.8683056. 225
Bibliography [33] James W Cannon, William J Floyd, Richard Kenyon, and Walter R Parry. Hyperbolic geometry. In Silvio Levy, editor, Flavors of Geometry, volume 31, pages 59–115. Cambridge University Press, 1997. [34] Rudrasis Chakraborty. ManifoldNorm: Extending normalizations on Riemannian manifolds. arXiv preprint arXiv:2003.13869, 2020. [35] Rudrasis Chakraborty and Baba C Vemuri. Statistics on the Stiefel manifold: theory and applications. The Annals of Statistics, 47(1):415–438, 2019. [36] Rudrasis Chakraborty, Chun-Hao Yang, Xingjian Zhen, Monami Banerjee, Derek Archer, David Vaillancourt, Vikas Singh, and Baba Vemuri. A statistical recurrent model on the manifold of symmetric positive definite matrices. In NeurIPS, 2018. [37] Rudrasis Chakraborty, Jose Bouza, Jonathan H Manton, and Baba C Vemuri. Manifoldnet: A deep neural network for manifold-valued data with applications. IEEE TPAMI, 2020. [38] Ines Chami, Zhitao Ying, Christopher Ré, and Jure Leskovec. Hyperbolic graph convolutional neural networks. NeurIPS, 2019. [39] Ines Chami, Albert Gu, Dat P Nguyen, and Christopher Ré. HoroPCA: Hyperbolic dimensionality reduction via horospherical projections. In ICML, 2021. [40] Woong-Gi Chang, Tackgeun You, Seonguk Seo, Suha Kwak, and Bohyung Han. Domain-specific batch normalization for unsupervised domain adaptation. In CVPR, 2019. [41] Kai-Xuan Chen, Jie-Yi Ren, Xiao-Jun Wu, and Josef Kittler. Covariance descriptors on a Gaussian manifold and their application to image set classification. Pattern Recognition, 107:107463, 2020. doi: 10.1016/j.patcog.2020.107463. [42] Kaixuan Chen, Jie Song, Shunyu Liu, Na Yu, Zunlei Feng, Gengshi Han, and Mingli Song. Distribution knowledge embedding for graph pooling. IEEE TKDE, 2023. [43] Ricky TQ Chen and Yaron Lipman. Flow matching on general geometries. In ICLR, 2024. [44] Tianyu Chen, Xingcheng Fu, Yisen Gao, Haodong Qian, Yuecen Wei, Kun Yan, Haoyi Zhou, and Jianxin Li. Galaxy walker: Geometry-aware VLMs for galaxyscale understanding. In CVPR, 2025. 226
Bibliography [45] Weize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. Fully hyperbolic neural networks. In ACL, 2022. [46] Yuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li, Ying Deng, and Weiming Hu. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In ICCV, 2021. [47] Ziheng Chen, Tianyang Xu, Xiao-Jun Wu, Rui Wang, Zhiwu Huang, and Josef Kittler. Riemannian local mechanism for SPD neural networks. In AAAI, 2023. [48] Ziheng Chen, Tianyang Xu, Xiao-Jun Wu, Rui Wang, and Josef Kittler. Hybrid Riemannian graph-embedding metric learning for image set classification. IEEE Transactions on Big Data, 9(1):75–92, 2023. doi: 10.1109/TBDATA.2021. 3113084. [49] Ziheng Chen, Yue Song, Gaowen Liu, Ramana Rao Kompella, Xiaojun Wu, and Nicu Sebe. Riemannian multinomial logistics regression for SPD neural networks. In CVPR, 2024. [50] Ziheng Chen, Yue Song, Yunmei Liu, and Nicu Sebe. A Lie group approach to Riemannian batch normalization. In ICLR, 2024. [51] Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, and Nicu Sebe. RMLR: Extending multinomial logistic regression into general geometries. In NeurIPS, 2024. [52] Ziheng Chen, Yue Song, Xiao-Jun Wu, and Nicu Sebe. Gyrogroup batch normalization. In ICLR, 2025. [53] Ziheng Chen, Yue Song, Xiaojun Wu, Gaowen Liu, and Nicu Sebe. Understanding matrix function normalizations in covariance pooling through the lens of Riemannian geometry. In ICLR, 2025. [54] Ziheng Chen, Xiao-Jun Wu, Bernhard Schölkopf, and Nicu Sebe. Riemannian batch normalization: A gyro approach. arXiv preprint arXiv:2509.07115, 2025. [55] Ziheng Chen, Xiaojun Wu, Bernhard Schölkopf, and Nicu Sebe. Building transformation layers for Riemannian neural networks, 2025. URL https://openreview. net/forum?id=1tJVBCpVD0. [56] Chaoran Cheng, Jiahan Li, Jian Peng, and Ge Liu. Categorical flow matching on statistical manifolds. In NeurIPS, 2024. 227
Bibliography [57] Chaoran Cheng, Jiahan Li, Jiajun Fan, and Ge Liu. α-flow: A unified framework for continuous-state discrete flow matching models, 2025. URL https://arxiv. org/abs/2504.10283. [58] Jaemoo Choi, Jaewoong Choi, and Myungjoo Kang. Scalable Wasserstein gradient flow for generative modeling through unbalanced optimal transport. In ICML, 2024. [59] Calin Cruceru, Gary Bécigneul, and Octavian-Eugen Ganea. Computationally tractable Riemannian manifolds for graph embeddings. In AAAI, 2021. [60] Jindou Dai, Yuwei Wu, Zhi Gao, and Yunde Jia. A hyperbolic-to-hyperbolic graph convolutional network. In CVPR, 2021. [61] Tingting Dan, Ziquan Wei, Won Hwa Kim, and Guorong Wu. Exploring the enigma of neural dynamics through a scattering-transform mixer landscape for Riemannian manifold. In ICML, 2024. [62] Paul David and Weiqing Gu. A Riemannian structure for correlation matrices. Operators and Matrices, 13(3):607–627, 2019. [63] Oscar Davis, Samuel Kessler, Mircea Petrache, Ismail Ilkan Ceylan, Michael Bronstein, and Avishek Joey Bose. Fisher flow matching for generative modeling over discrete data. In NeurIPS, 2024. [64] Valentin De Bortoli, Emile Mathieu, Michael Hutchinson, James Thornton, Yee Whye Teh, and Arnaud Doucet. Riemannian score-based generative modelling. In NeurIPS, 2022. [65] Thibault de Surrel, Sylvain Chevallier, Fabien Lotte, and Florian Yger. Geometryaware visualization of high dimensional symmetric positive definite matrices. TMLR, 2025. [66] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. [67] Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shanmukha Ramakrishna Vedantam. Hyperbolic image-text representations. In ICML, 2023. 228
Bibliography [68] Abhinav Dhall, Amanjot Kaur, Roland Goecke, and Tom Gedeon. Emotiw 2018: Audio-video, student engagement and group-level affect prediction. In Proceedings of the 20th ACM International Conference on Multimodal Interaction, pages 653– 656, 2018. URL https://doi.org/10.1145/3242969.3264993. [69] Manfredo Perdigao Do Carmo and J Flaherty Francis. Riemannian Geometry, volume 6. Springer, 1992. [70] Ian L Dryden, Xavier Pennec, and Jean-Marc Peyrat. Power Euclidean metrics for covariance matrices with application to diffusion tensor imaging. arXiv preprint arXiv:1009.3045, 2010. URL https://arxiv.org/abs/1009.3045. [71] Alan Edelman, Tomás A Arias, and Steven T Smith. The geometry of algorithms with orthogonality constraints. SIAM journal on Matrix Analysis and Applications, 20(2):303–353, 1998. [72] Aleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe, and Ivan Oseledets. Hyperbolic vision transformers: Combining improvements in metric learning. In CVPR, 2022. [73] Xiran Fan, Chun-Hao Yang, and Baba Vemuri. Horospherical decision boundaries for large margin classification in hyperbolic space. In NeurIPS, 2023. [74] Maurice Fréchet. Les éléments aléatoires de nature quelconque dans un espace distancié. Annales de l’institut Henri Poincaré, 10(4):215–310, 1948. [75] Xingcheng Fu, Yisen Gao, Yuecen Wei, Qingyun Sun, Hao Peng, Jianxin Li, and Xianxian Li. Hyperbolic geometric latent diffusion model for graph generation. In ICML, 2024. [76] Octavian Ganea, Gary Bécigneul, and Thomas Hofmann. Hyperbolic neural networks. NeurIPS, 2018. [77] Zhi Gao, Yuwei Wu, Mehrtash Harandi, and Yunde Jia. A robust distance measure for similarity-based classification on the spd manifold. IEEE TNNLS, 2019. [78] Zhi Gao, Yuwei Wu, Yunde Jia, and Mehrtash Harandi. Curvature generation in curved spaces for few-shot learning. In ICCV, 2021. [79] Zhi Gao, Chen Xu, Feng Li, Yunde Jia, Mehrtash Harandi, and Yuwei Wu. Exploring data geometry for continual learning. In CVPR, 2023. 229
Bibliography [80] Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. First-person hand action benchmark with RGB-D videos and 3D hand pose annotations. In CVPR, 2018. [81] Mina Ghadimi Atigh, Martin Keller-Ressel, and Pascal Mettes. Hyperbolic busemann learning with ideal prototypes. In NeurIPS, 2021. [82] Xavier Glorot, Antoine Bordes, and Yoshua Bengio. Deep sparse rectifier neural networks. In AISTATS, 2011. [83] Gene H. Golub and Charles F. Van Loan. Matrix Computations. JHU press, 2013. [84] Alexandre Gramfort. MEG and EEG data analysis with MNE-Python. Frontiers in Neuroscience, 7, 2013. [85] Karish Grover, Geoffrey J Gordon, and Christos Faloutsos. CurvGAD: Leveraging curvature for enhanced graph anomaly detection. In ICML, 2025. [86] Karish Grover, Haiyang Yu, Xiang Song, Qi Zhu, Han Xie, Vassilis N Ioannidis, and Christos Faloutsos. Spectro-Riemannian graph neural networks. In ICLR, 2025. [87] Albert Gu, Frederic Sala, Beliz Gunel, and Christopher Ré. Learning mixedcurvature representations in product spaces. In ICLR, 2019. [88] Nicolas Guigui, Nina Miolane, and Xavier Pennec. Introduction to Riemannian geometry and geometric statistics: From basic theory to implementation with Geomstats. Foundations and Trends in Machine Learning, 16(3):329–493, 2023. doi: 10.1561/2200000098. [89] Johann Guilleminot and Christian Soize. Generalized stochastic approach for constitutive equation in linear elasticity: a random matrix model. International Journal for Numerical Methods in Engineering, 90(5):613–635, 2012. URL https: //doi.org/10.1002/nme.3338. [90] Caglar Gulcehre, Misha Denil, Mateusz Malinowski, Ali Razavi, Razvan Pascanu, Karl Moritz Hermann, Peter Battaglia, Victor Bapst, David Raposo, Adam Santoro, et al. Hyperbolic attention networks. In ICLR, 2019. [91] Yunhui Guo, Xudong Wang, Yubei Chen, and Stella X Yu. Clipped hyperbolic classifiers are super-hyperbolic classifiers. In CVPR, 2022. 230
Bibliography [92] Andi Han, Bamdev Mishra, Pratik Kumar Jawanpuria, and Junbin Gao. On Riemannian optimization over positive definite matrices with the Bures-Wasserstein geometry. NeurIPS, 2021. [93] Andi Han, Bamdev Mishra, Pratik Jawanpuria, and Junbin Gao. Learning with symmetric positive definite matrices via generalized Bures-Wasserstein geometry. In International Conference on Geometric Science of Information, pages 405–415. Springer, 2023. [94] Mehrtash Harandi, Mathieu Salzmann, and Richard Hartley. Dimensionality reduction on SPD manifolds: The emergence of geometry-aware methods. IEEE TPAMI, 2018. [95] Richard Hartley, Jochen Trumpf, Yuchao Dai, and Hongdong Li. Rotation averaging. IJCV, 2013. [96] Doron Haviv, Aram-Alexandre Pooladian, Dana Pe’Er, and Brandon Amos. Wasserstein flow matching: Generative modeling over families of distributions. In ICML, 2025. [97] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. [98] Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Leandros Tassiulas, Menglin Yang, and Rex Ying. HELM: Hyperbolic large language models via mixture-of-curvature experts. In NeurIPS, 2025. [99] Neil He, Menglin Yang, and Rex Ying. Lorentzian residual neural networks. In KDD, 2025. [100] Uwe Helmke and John B Moore. Optimization and Dynamical Systems. Springer Science & Business Media, 2012. [101] Marcel F. Hinss, Ludovic Darmet, Bertille Somon, Emilie Jahanpour, Fabien Lotte, Simon Ladouce, and Raphaëlle N. Roy. An EEG dataset for cross-session mental workload estimation: Passive BCI competition of the Neuroergonomics Conference 2021, 2021. [102] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997. 231
Bibliography [103] Chen Hu, Rui Wang, Xiaoning Song, Tao Zhou, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. A correlation manifold self-attention network for EEG decoding. In IJCAI, 2025. [104] Chen Hu, Ziheng Chen, Rui Wang, Yefeng Zheng, and Nicu Sebe. Riemannian high-order pooling for brain foundation models. In ICLR, 2026. [105] Xiaoqiang Hua, Yongqiang Cheng, Hongqiang Wang, Yuliang Qin, Yubo Li, and Wenpeng Zhang. Matrix CFAR detectors based on symmetrized Kullback–Leibler and total Kullback–Leibler divergences. Digital Signal Processing, 69:106–116, 2017. URL https://doi.org/10.1016/j.dsp.2017.06.019. [106] Zhiwu Huang and Luc Van Gool. A Riemannian network for SPD matrix learning. In AAAI, 2017. [107] Zhiwu Huang, Chengde Wan, Thomas Probst, and Luc Van Gool. Deep learning on Lie groups for skeleton-based action recognition. In CVPR, 2017. [108] Zhiwu Huang, Jiqing Wu, and Luc Van Gool. Building deep networks on Grassmann manifolds. In AAAI, 2018. [109] Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. [110] Catalin Ionescu, Orestis Vantzos, and Cristian Sminchisescu. Matrix backpropagation for deep networks with structured layers. In ICCV, 2015. [111] Catalin Ionescu, Orestis Vantzos, and Cristian Sminchisescu. Training deep networks with structured layers by matrix backpropagation. arXiv preprint arXiv:1509.07838, 2015. [112] Vinay Jayaram and Alexandre Barachant. MOABB: trustworthy algorithm benchmarking for BCIs. Journal of Neural Engineering, 15(6):066011, 2018. [113] Shaocheng Jin, Tao Zhou, Rui Wang, Ziheng Chen, Xiaoqing Luo, Xiao-Jun Wu, and Josef Kittler. Towards robust EEG decoding based on Riemannian selfattention. In KDD, 2026. [114] Dimitris Kalatzis, David Eklund, Georgios Arvanitidis, and Søren Hauberg. Variational autoencoders with riemannian brownian motion priors. In ICML, 2020. 232
Bibliography [115] Huan Kang, Hui Li, Tianyang Xu, Xiao-Jun Wu, Rui Wang, Chunyang Cheng, and Josef Kittler. SMLNet: A SPD manifold learning network for infrared and visible image fusion. IJCV, 2025. [116] Hermann Karcher. Riemannian center of mass and mollifier smoothing. Communications on Pure and Applied Mathematics, 30(5):509–541, 1977. URL https://doi.org/10.1002/cpa.3160300502. [117] Isay Katsman, Eric Ming Chen, Sidhanth Holalkere, Anna Asch, Aaron Lou, SerNam Lim, and Christopher De Sa. Riemannian residual neural networks. In NeurIPS, 2023. [118] Isay Katsman, Eric Chen, Sidhanth Holalkere, Anna Asch, Aaron Lou, Ser Nam Lim, and Christopher M De Sa. Riemannian residual neural networks. NeurIPS, 2024. [119] Raiyan R Khan, Philippe Chlenski, and Itsik Pe’er. Hyperbolic genome embeddings. In ICLR, 2025. [120] Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Ustinova, Ivan Oseledets, and Victor Lempitsky. Hyperbolic image embeddings. In CVPR, 2020. [121] Diederik P Kingma. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. [122] Diederik P Kingma and Jimmy Lei Ba. Adam: A method for stochastic optimization. In ICLR, 2015. [123] Reinmar Kobler, Jun-ichiro Hirayama, Qibin Zhao, and Motoaki Kawanabe. SPD domain-specific batch normalization to crack interpretable unsupervised domain adaptation in EEG. In NeurIPS, 2022. [124] Reinmar J Kobler, Jun-ichiro Hirayama, and Motoaki Kawanabe. Controlling the Fréchet variance improves batch normalization on the symmetric positive definite manifold. In ICASSP, 2022. [125] Max Kochurov, Rasul Karimov, and Serge Kozlukov. Geoopt: Riemannian optimization in pytorch. arXiv preprint arXiv:2005.02819, 2020. [126] A Krizhevsky. Learning multiple layers of features from tiny images. Master’s thesis, University of Tront, 2009. 233
Bibliography [127] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, 2012. [128] Serge Lang. Algebra. Springer Science & Business Media, 2012. [129] Yann Le and Xuan Yang. Tiny ImageNet visual recognition challenge, 2015. [130] Guy Lebanon and John Lafferty. Hyperplane margin classifiers on the multinomial manifold. In ICML, 2004. [131] John M Lee. Introduction to Riemannian Manifolds, volume 2. Springer, 2018. [132] Mario Lezcano Casado. Trivializations for gradient-based optimization on manifolds. In NeurIPS, 2019. [133] Peihua Li, Jiangtao Xie, Qilong Wang, and Zilin Gao. Towards faster training of global covariance pooling networks by iterative matrix square root normalization. In CVPR, 2018. [134] Shanglin Li, Motoaki Kawanabe, and Reinmar J Kobler. SPDIM: Source-free unsupervised conditional and label shift adaptation in EEG. In ICLR, 2025. [135] Shanglin Li, Shiwen Chu, Okan Koç, Yi Ding, Qibin Zhao, Motoaki Kawanabe, and Ziheng Chen. HEEGNet: Hyperbolic embeddings for EEG. ICLR, 2026. [136] Yancong Li, Xiaoming Zhang, Ying Cui, and Shuai Ma. Hyperbolic graph neural network for temporal knowledge graph completion. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2024. [137] Zhenhua Lin. Riemannian geometry of symmetric positive definite matrices via Cholesky decomposition. SIAM Journal on Matrix Analysis and Applications, 40 (4):1353–1370, 2019. [138] Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. NTU RGB+D 120: A large-scale benchmark for 3D human activity understanding. IEEE T-PAMI, 2019. [139] Qi Liu, Maximilian Nickel, and Douwe Kiela. Hyperbolic graph neural networks. In NeurIPS, 2019. 234
Bibliography [140] Yuanpei Liu, Zhenqi He, and Kai Han. Hyperbolic category discovery. In CVPR, 2025. [141] Federico López, Beatrice Pozzetti, Steve Trettel, Michael Strube, and Anna Wienhard. Vector-valued distance and Gyrocalculus on the space of symmetric positive definite matrices. In NeurIPS, 2021. [142] Aaron Lou, Isay Katsman, Qingxuan Jiang, Serge Belongie, Ser-Nam Lim, and Christopher De Sa. Differentiating through the Fréchet mean. In ICML, 2020. [143] Miroslav Lovrić, Maung Min-Oo, and Ernst A Ruh. Multivariate normal distributions parametrized as a riemannian symmetric space. Journal of Multivariate Analysis, 74(1):36–48, 2000. [144] Jan R Magnus and Heinz Neudecker. Matrix differential calculus with applications in statistics and econometrics. John Wiley & Sons, 2019. URL https://doi. org/10.1002/9781119541219. [145] Luigi Malagò, Luigi Montrucchio, and Giovanni Pistone. Wasserstein Riemannian geometry of gaussian densities. Information Geometry, 1:137–179, 2018. [146] Jonathan H Manton. A globally convergent numerical algorithm for computing the centre of mass on compact Lie groups. In The 8th Control, Automation, Robotics and Vision Conference, 2004., volume 3, pages 2211–2216. IEEE, 2004. [147] Yidan Mao, Jing Gu, Marcus C Werner, and Dongmian Zou. Klein model for hyperbolic neural networks. arXiv preprint arXiv:2410.16813, 2024. [148] Emile Mathieu and Maximilian Nickel. Riemannian continuous normalizing flows. In NeurIPS, 2020. [149] Debin Meng, Xiaojiang Peng, Kai Wang, and Yu Qiao. Frame attention networks for facial expression recognition in videos. In 2019 IEEE International Conference on Image Processing (ICIP), pages 3866–3870. IEEE, 2019. URL https://doi. org/10.1109/ICIP.2019.8803603. [150] Aaron Meurer, Christopher P Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K Moore, Sartaj Singh, et al. Sympy: symbolic computing in python. PeerJ Computer Science, 3: e103, 2017. 235
Bibliography [151] Hà Quang Minh. Alpha Procrustes metrics between positive definite operators: a unifying formulation for the Bures-Wasserstein and Log-Euclidean/Log-HilbertSchmidt metrics. Linear Algebra and its Applications, 636:25–68, 2022. [152] Maher Moakher. On the averaging of symmetric positive-definite tensors. Journal of Elasticity, 82(3):273–296, 2006. URL https://doi.org/10.1007/ s10659-005-9035-z. [153] Meinard Müller, Tido Röder, Michael Clausen, Bernhard Eberhardt, Björn Krüger, and Andreas Weber. Documentation mocap database HDM05. Technical report, Universität Bonn, 2007. [154] Iain Murray. Differentiation of the Cholesky decomposition. arXiv preprint arXiv:1602.07527, 2016. [155] Galileo Namata, Ben London, Lise Getoor, Bert Huang, and U Edu. Query-driven active surveying for collective classification. In 10th International Workshop on Mining and Learning with Graphs, 2012. [156] Xuan Son Nguyen. Geomnet: A neural network based on Riemannian geometries of SPD matrix space and Cholesky space for 3D skeleton-based interaction recognition. In ICCV, 2021. [157] Xuan Son Nguyen. The Gyro-structure of some matrix manifolds. In NeurIPS, 2022. [158] Xuan Son Nguyen. A Gyrovector space approach for symmetric positive semidefinite matrix learning. In ECCV, 2022. [159] Xuan Son Nguyen and Shuo Yang. Building neural networks on matrix manifolds: A Gyrovector space approach. In ICML, 2023. [160] Xuan Son Nguyen, Luc Brun, Olivier Lézoray, and Sébastien Bougleux. A neural network based on SPD manifold learning for skeleton-based hand gesture recognition. In CVPR, 2019. [161] Xuan Son Nguyen, Shuo Yang, and Aymeric Histace. Matrix manifold neural networks++. In ICLR, 2024. [162] Xuan Son Nguyen, Shuo Yang, and Aymeric Histace. Neural networks on symmetric spaces of noncompact type. In ICLR, 2025. 236
Bibliography [163] Maximillian Nickel and Douwe Kiela. Poincaré embeddings for learning hierarchical representations. In NeurIPS, 2017. [164] Maximillian Nickel and Douwe Kiela. Learning continuous hierarchies in the Lorentz model of hyperbolic geometry. In ICML, 2018. [165] Barrett O’Neill. Semi-Riemannian Geometry with Applications to Relativity, volume 103 of Pure and Applied Mathematics. Academic Press, New York, 1983. [166] Avik Pal, Max van Spengler, Guido Maria D’Amely di Melendugno, Alessandro Flaborea, Fabio Galasso, and Pascal Mettes. Compositional entailment learning for hyperbolic vision-language models. In ICLR, 2025. [167] Yue-Ting Pan, Jing-Lun Chou, and Chun-Shu Wei. MAtt: A manifold attention network for EEG decoding. In NeurIPS, 2022. [168] Yong-Hyun Park, Mingi Kwon, Jaewoong Choi, Junghyo Jo, and Youngjung Uh. Understanding the latent space of diffusion models through the lens of Riemannian geometry. In NeurIPS, 2023. [169] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019. [170] Xavier Pennec. Probabilities and statistics on Riemannian manifolds: A geometric approach. PhD thesis, INRIA, 2004. [171] Xavier Pennec, Pierre Fillard, and Nicholas Ayache. A riemannian framework for tensor computing. IJCV, 2006. [172] Can Pouliquen, Mathurin Massias, and Titouan Vayer. Schur’s positive-definite network: Deep learning in the SPD cone with structure. In ICLR, 2025. [173] John G Ratcliffe. Foundations of Hyperbolic Manifolds. Springer, 2006. [174] Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. Accelerating 3D deep learning with PyTorch3D. arXiv preprint arXiv:2007.08501, 2020. [175] Herbert Robbins and Sutton Monro. A stochastic approximation method. The annals of mathematical statistics, pages 400–407, 1951. 237
Bibliography [176] Salem Said, Lionel Bombrun, Yannick Berthoumieu, and Jonathan H Manton. Riemannian Gaussian distributions on the space of symmetric positive definite matrices. IEEE TIT, 2017. [177] Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective classification in network data. AI magazine, 29 (3):93–93, 2008. [178] Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. NTU RGB+ D: A large scale dataset for 3D human activity analysis. In CVPR, 2016. [179] Xianglong Shi, Ziheng Chen, Yunhan Jiang, and Nicu Sebe. Intrinsic Lorentz neural network. In ICLR, 2026. [180] Ryohei Shimizu, Yusuke Mukuta, and Tatsuya Harada. Hyperbolic neural networks++. In ICLR, 2021. [181] Ryohei Shimizu, YUSUKE Mukuta, and Tatsuya Harada. Hyperbolic neural networks++. In ICLR, 2021. [182] Ondrej Skopek, Octavian-Eugen Ganea, and Gary Bécigneul. Mixed-curvature variational autoencoders. In ICLR, 2020. [183] Yue Song, Nicu Sebe, and Wei Wang. Why approximate matrix square root outperforms accurate svd in global covariance pooling? In ICCV, 2021. [184] Yue Song, Nicu Sebe, and Wei Wang. On the eigenvalues of global covariance pooling for fine-grained visual recognition. IEEE TPAMI, 2022. [185] Yue Song, Nicu Sebe, and Wei Wang. Fast differentiable matrix square root. In ICLR, 2022. [186] Suvrit Sra and Reshad Hosseini. Conic geometric optimization on the manifold of positive definite matrices. SIAM Journal on Optimization, 25(1):713–739, 2015. URL https://doi.org/10.1137/140978168. [187] Shlomo Sternberg. Lectures on differential geometry, volume 316. American Mathematical Soc., 1999. URL https://www.ams.org/journals/bull/1965-71-02/ S0002-9904-1965-11286-1/S0002-9904-1965-11286-1.pdf. 238
Bibliography [188] Yadong Sun, Xiaofeng Cao, Yu Wang, Wei Ye, Jingcai Guo, and Qing Guo. Geometry awakening: Cross-geometry learning exhibits superiority over individual structures. In NeurIPS, 2024. [189] Tanuj Sur, Samrat Mukherjee, Kaizer Rahaman, Subhasis Chaudhuri, Muhammad Haris Khan, and Biplab Banerjee. Hyperbolic uncertainty-aware few-shot incremental point cloud segmentation. In CVPR, 2025. [190] Yann Thanwerdas. Riemannian and stratified geometries on covariance and correlation matrices. PhD thesis, Université Côte d’Azur, 2022. [191] Yann Thanwerdas. Permutation-invariant log-Euclidean geometries on full-rank correlation matrices. SIAM Journal on Matrix Analysis and Applications, 45(2): 930–953, 2024. doi: 10.1137/22M1538144. [192] Yann Thanwerdas and Xavier Pennec. Is affine-invariance well defined on SPD matrices? a principled continuum of metrics. In Geometric Science of Information: 4th International Conference, GSI 2019, Toulouse, France, August 27–29, 2019, Proceedings 4, pages 502–510. Springer, 2019. [193] Yann Thanwerdas and Xavier Pennec. Exploration of balanced metrics on symmetric positive definite matrices. In Geometric Science of Information: 4th International Conference, GSI 2019, Toulouse, France, August 27–29, 2019, Proceedings 4, pages 484–493. Springer, 2019. [194] Yann Thanwerdas and Xavier Pennec. The geometry of mixed-Euclidean metrics on symmetric positive definite matrices. Differential Geometry and its Applications, 81:101867, 2022. [195] Yann Thanwerdas and Xavier Pennec. Theoretically and computationally convenient geometries on full-rank correlation matrices. SIAM Journal on Matrix Analysis and Applications, 43(4):1851–1872, 2022. doi: 10.1137/22M1471729. [196] Yann Thanwerdas and Xavier Pennec. O (n)-invariant Riemannian metrics on SPD matrices. Linear Algebra and its Applications, 661:163–201, 2023. [197] Loring W. Tu. An introduction to manifolds. Springer, 2011. [198] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: the missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016. 239
Bibliography [199] Abraham A Ungar. Analytic hyperbolic geometry: Mathematical foundations and applications. World Scientific, 2005. [200] Abraham Albert Ungar. Analytic Hyperbolic Geometry and Albert Einstein’s Special Theory of Relativity (Second Edition). World Scientific, 2022. [201] Max Van Spengler, Erwin Berkhout, and Pascal Mettes. Poincaré ResNet. In ICCV, 2023. [202] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. [203] Raviteja Vemulapalli, Felipe Arrate, and Rama Chellappa. Human action recognition by representing 3D skeletons as points in a Lie group. In CVPR, 2014. [204] Haoqi Wang, Zhizhong Li, and Wayne Zhang. Get the best of both worlds: Improving accuracy and transferability by Grassmann class representation. In ICCV, 2023. [205] Qilong Wang, Jiangtao Xie, Wangmeng Zuo, Lei Zhang, and Peihua Li. Deep cnns meet global covariance pooling: Better representation and generalization. IEEE TPAMI, 2020. [206] Qilong Wang, Zhaolin Zhang, Mingze Gao, Jiangtao Xie, Pengfei Zhu, Peihua Li, Wangmeng Zuo, and Qinghua Hu. Towards a deeper understanding of global covariance pooling in deep learning: An optimization perspective. IEEE TPAMI, 2023. [207] Rui Wang, Xiao-Jun Wu, and Josef Kittler. SymNet: A simple symmetric positive definite manifold deep learning method for image set classification. IEEE TNNLS, 2021. [208] Rui Wang, Xiao-Jun Wu, Ziheng Chen, Tianyang Xu, and Josef Kittler. DreamNet: A deep Riemannian manifold network for SPD matrix learning. In ACCV, 2022. [209] Rui Wang, Xiao-Jun Wu, Ziheng Chen, Tianyang Xu, and Josef Kittler. Learning a discriminative SPD manifold neural network for image set classification. Neural Networks, 151:94–110, 2022. 240
Bibliography [210] Rui Wang, Chen Hu, Ziheng Chen, Xiao-Jun Wu, and Xiaoning Song. A Grassmannian manifold self-attention network for signal classification. In IJCAI, 2024. [211] Rui Wang, Xiao-Jun Wu, Ziheng Chen, Cong Hu, and Josef Kittler. SPD manifold deep metric learning for image set classification. IEEE TNNLS, 2024. [212] Rui Wang, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. Towards a general attention framework on gyrovector spaces for matrix manifolds. In NeurIPS, 2025. [213] Rui Wang, Jiayao Jin, Ziheng Chen, Cong Wu, Xiao-Jun Wu, and Nicu Sebe. Structural topology refinement network for skeleton-based action recognition. IEEE TIM, 2025. [214] Rui Wang, Shaocheng Jin, Zhenyu Cai, Ziheng Chen, Xiao-Jun Wu, and Josef Kittler. Learning a better SPD network for signal classification: A Riemannian batch normalization method. IEEE TNNLS, 2025. [215] Rui Wang, Shaocheng Jin, Ziheng Chen, Xiaoqing Luo, and Xiao-Jun Wu. Learning to normalize on the SPD manifold under Bures-Wasserstein geometry. In CVPR, 2025. [216] Rui Wang, Zihao Bi, Chen Hu, Xiaoning Song, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. Riemannian graph convolutional network for skeleton-based twoperson interaction recognition. In IJCAI, 2026. [217] Rui Wang, Yuting Jiang, Xiaoqing Luo, Xiao-Jun Wu, Nicu Sebe, and Ziheng Chen. Wasserstein-aligned hyperbolic multi-view clustering. AAAI, 2026. [218] Yunfeng Wang and Gregory S Chirikjian. Error propagation on the euclidean group with applications to manipulator kinematics. IEEE Transactions on Robotics, 22(4):591–602, 2006. [219] Yuxin Wu and Kaiming He. Group normalization. In ECCV, 2018. [220] Or Yair, Mirela Ben-Chen, and Ronen Talmon. Parallel transport on the cone manifold of SPD matrices for domain adaptation. IEEE TIP, 2019. [221] Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, and Rex Ying. Hypformer: Exploring efficient transformer fully in hyperbolic space. In KDD, 2024. 241
Bibliography [222] Menglin Yang, Ram Samarth B B, Aosong Feng, Bo Xiong, Jiahong Liu, Irwin King, and Rex Ying. Hyperbolic fine-tuning for large language models. In NeurIPS, 2025. [223] Xin Yang, Xingrun Li, Heng Chang, Xihong Yang, Shengyu Tao, Maiko Shigeno, Ningkang Chang, Junfeng Wang, Dawei Yin, Erxue Min, et al. Hgformer: Hyperbolic graph transformer for collaborative filtering. In ICML, 2025. [224] Ryoma Yataka, Kazuki Hirashima, and Masashi Shiraishi. Grassmann manifold flows for stable shape generation. In NeurIPS, 2023. [225] Florian Yger. A review of kernels on covariance matrices for BCI applications. In 2013 IEEE International Workshop on Machine Learning for Signal Processing (MLSP), pages 1–6. IEEE, 2013. URL https://doi.org/10.1109/MLSP.2013. 6661972. [226] Hongwei Yong, Jianqiang Huang, Deyu Meng, Xiansheng Hua, and Lei Zhang. Momentum batch normalization for deep learning with small batch size. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part XII 16, pages 224–240. Springer, 2020. [227] Ying Yuan, Hongtu Zhu, Weili Lin, and James Stephen Marron. Local polynomial regression for symmetric positive definite matrices. Journal of the Royal Statistical Society Series B: Statistical Methodology, 74(4):697–719, 2012. [228] Ernesto Zacur, Matias Bossa, and Salvador Olmos. Left-invariant Riemannian geodesics on spatial transformation groups. SIMAX, 2014. [229] Muhan Zhang and Yixin Chen. Link prediction based on graph neural networks. In NeurIPS, 2018. [230] Tong Zhang, Wenming Zheng, Zhen Cui, Yuan Zong, Chaolong Li, Xiaoyan Zhou, and Jian Yang. Deep manifold-to-manifold transforming network for skeletonbased action recognition. IEEE TMM, 2020. [231] Wei Zhao, Federico Lopez, J Maxwell Riestenberg, Michael Strube, Diaaeldin Taha, and Steve Trettel. Modeling graphs beyond hyperbolic: Graph neural networks in symmetric positive definite matrices. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 122–139. Springer, 2023. 242
Bibliography [232] Xingjian Zhen, Rudrasis Chakraborty, Nicholas Vogt, Barbara B Bendlin, and Vikas Singh. Dilated convolutional neural networks for sequential manifold-valued data. In ICCV, 2019. [233] Runhe Zhou, Shanglin Li, Guanxiang Huang, Xinliang Zhou, Qibin Zhao, Motoaki Kawanabe, Yi Ding, and Cuntai Guan. EEG-based multimodal learning via hyperbolic mixture-of-curvature experts. In ICML, 2026. [234] Zhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta, Ramana Davuluri, and Han Liu. DNABERT-2: Efficient foundation model and benchmark for multi-species genome. In ICLR, 2024.
243
Bibliography
244
Appendix A Experimental Details and Additional Discussions A.1
Data Sets
This section collects descriptions of the data sets used in the experimental chapters.
A.1.1
Skeleton-Based Action Recognition and Gesture Data Sets
HDM05 [153]. The HDM05 data set1 consists of 2,273 skeleton-based motion capture sequences executed by different actors. Each frame records the 3D coordinates of 31 joints. Different experiments use the standard task-specific filtered versions adopted by their corresponding backbones, as detailed in the method-specific experimental settings. FPHA [80]. The FPHA data set2 includes 1,175 skeleton-based first-person hand gesture videos of 45 different categories, with 600 clips for training and 575 for testing. Each frame contains the 3D coordinates of 21 hand joints. G3D [23]. The G3D data set consists of 663 sequences of 20 different gaming actions. Each sequence records the 3D locations of 20 joints, i.e.,, 19 bones. NTU60 [178]. The NTU60 data set3 contains 56,880 skeleton sequences classified into 60 classes, where each frame includes 3D coordinates of 25 or 50 body joints. NTU120 [138]. The NTU120 data set4 contains 114,480 sequences in 120 action classes. 1
https://resources.mpi-inf.mpg.de/HDM05/ https://github.com/guiggh/hand_pose_action 3 https://github.com/shahroudy/NTURGB-D 4 https://github.com/shahroudy/NTURGB-D 2
245
A.1. Data Sets Data Set
#Nodes
#Edges
#Classes
#Features
Disease Airport PubMed Cora
1044 3188 19717 2708
1043 18631 44338 5429
2 4 3 7
1000 4 500 1433
Table A.1: Summary statistics for the graph data sets.
A.1.2
Radar and EEG Signal Data Sets
Radar [31]. The Radar data set5 contains 3,000 synthetic radar signals equally distributed in 3 classes. Hinss2021 [101]. The Hinss2021 data set6 is a competition data set containing EEG signals for mental workload estimation. The data set is employed for two tasks, inter-session and inter-subject, which are treated as domain adaptation problems. Geometry-aware methods [220, 123] have demonstrated promising performance in EEG classification.
A.1.3
Image Classification Data Sets
AFEW [68]. The Acted Facial Expressions in the Wild (AFEW) data set contains seven emotion categories, with 773 videos for training and 383 for validation. CIFAR-10 and CIFAR-100 [126]. The CIFAR-10 and CIFAR-100 data sets each contain 60,000 32 × 32 color images from 10 and 100 classes, respectively. We use the standard PyTorch splits: 50,000 training images and 10,000 test images. Tiny-ImageNet [129]. Tiny-ImageNet is a subset of ImageNet with 100,000 images from 200 classes, resized to 64 × 64. We use the official validation split for evaluation. ImageNet-1k [66]. ImageNet-1k contains 1.28M training images, 50K validation images, and 100K test images distributed across 1k classes.
A.1.4
Graph Data Sets
Cora [177]. Cora is a citation network where the nodes represent scientific papers in the area of machine learning, the edges are citations between them, and the labels of the nodes are academic subareas. 5 6
https://www.dropbox.com/s/dfnlx2bnyh3kjwy/data.zip?dl=0 https://zenodo.org/record/5055046
246
Appendix A. Experimental Details and Additional Discussions Task
Species
Data Sets
Num. classes
Max length
Train / Dev / Test
Retrotransposons
Plant
LTR Copia LINEs SINEs
2
500 1000 500
7666 / 682 / 568 22502 / 2030 / 1782 21152 / 1836 / 1784
DNA transposons
Plant
CMC-EnSpm hAT-Ac
2
200 1000
19912 / 1872 / 1808 17322 / 1822 / 1428
Pseudogenes
Human
processed unprocessed
2
1000
17956 / 1046 / 1740 12938 / 766 / 884
Table A.2: Summary statistics for TEB. Disease [4]. Disease represents a disease propagation tree that simulates the susceptible, infected, and recovered (SIR) disease transmission model, with each node representing either an infection or a non-infection state. Airport [229]. Airport is a transductive data set where nodes represent airports and edges represent airline routes from OpenFlights.org. PubMed [155]. PubMed is a standard benchmark that describes citation networks where nodes represent scientific papers in the area of medicine, the edges are citations between them, and the node labels are academic subareas. Tab. A.1 summarizes the above data sets.
A.1.5
Genomic Sequence Data Sets
Transposable Elements Benchmark [119]. The Transposable Elements Benchmark (TEB) comprises seven binary classification data sets that investigate transposable elements across plant and human genomes. The seven data sets are LTR Copia, LINEs, and SINEs for plant retrotransposons; CMC-EnSpm and hAT-Ac for plant DNA transposons; and processed and unprocessed pseudogenes for human pseudogenes. For each data set, positive examples are sequences spanning annotated elements of interest, and negatives are randomly sampled, non-overlapping genomic segments outside these regions. We adopt chromosome-level training, validation, and test splits, using chromosomes 8 and 9 for validation and test in plant genomes, and chromosomes 20–22 and 17–19 for validation and test in human genomes, respectively. Tab. A.2 provides the summary statistics. Genome Understanding Evaluation [234]. The Genome Understanding Evaluation (GUE) benchmark contains seven biologically significant genome analysis tasks that span 28 data sets. The data sets contain sequences ranging from 70 to 1000 base pairs in length and originate from yeast, mouse, human, and virus genomes. The HBNN experiments select Core promoter detection and Promoter detection, together with the 247
A.2. Backbone Networks Task
Species
Data Set
Num. classes
Length
Train / Dev / Test
Core promoter detection
Human
tata notata all
2
70
4904 / 613 / 613 42452 / 5307 / 5307 47356 / 5920 / 5920
Promoter detection
Human
tata notata all
2
300
4904 / 613 / 613 42452 / 5307 / 5307 47356 / 5920 / 5920
Covid variant classification
Virus
Covid
9
1000
77669 / 7000 / 7000
Species classification
Fungi Virus
Fungi Virus
25 20
5000 10000
8000 / 1000 / 1000 4000 / 500 / 500
Table A.3: Summary statistics for the adopted GUE data sets. multi-class tasks Covid variant classification and species classification, totaling nine data sets. Tab. A.3 provides their summary statistics.
A.2
Backbone Networks
This section collects the backbone network architectures used in the thesis.
A.2.1
Backbone Networks on the SPD Manifold
SPDNet. SPDNet [106] is a classical SPD neural network. It mimics conventional densely connected feedforward networks and consists of three basic building blocks: BiMap layer: S k = W k S k−1 W k⊤ , with W k semi-orthogonal,
(A.1)
ReEig layer: S k = U k−1 max(Σk−1 , ϵIn )U k−1⊤ , with S k−1 = U k−1 Σk−1 U k−1⊤ ,
(A.2)
LogEig layer: S k = log(S k−1 ).
(A.3)
where max(·) is element-wise maximization. BiMap and ReEig mimic transformation and non-linear activation, while LogEig maps SPD matrices into the tangent space at the identity matrix for classification. SPDNetBN. SPDNetBN [31] further proposed RBN based on AIM: 1
1
n : ∀i ≤ N, P̄i ← M − 2 Pi M − 2 , Centering from mean M ∈ S++ 1 2
1 2
n Biasing towards parameter B ∈ S++ : ∀i ≤ N, P̂i ← B P̄i B .
(A.4) (A.5)
TSMNet and SPDDSMBN. TSMNet [123] can be illustrated as ftc → fsc → fBiM ap → fReEig → fLogEig , where ftc and fsc denote temporal and spatial convolution, 248
Appendix A. Experimental Details and Additional Discussions respectively. SPDDSMBN [123] is an improved version of SPDNetBN. Apart from controlling the mean, it can also control variance. The key operation in SPDDSMBN for controlling the mean and variance is ∀i ≤ N,
s
P̄i ← ΓI→B [(ΓM →I (Pi )) v ],
(A.6)
n is the biasing where M is the Riemannian mean, v 2 is the Fréchet variance, B ∈ S++ parameter, and s ∈ R is the scaling factor. Inspired by Yong et al. [226], during training, SPDDSMBN generates running means and running variances for training and testing with distinct momentum parameters. It uses the training running statistics during training and the testing running statistics during testing. SPDDSMBN also applies domain-specific techniques [40], keeping multiple parallel BN layers and distributing observations according to the associated domains. To share cross-domain knowledge, s is uniformly learned across all domains, and B is set to the identity matrix. Kobler et al. [123] adopted SPDDSMBN for domain adaptation in EEG classification.
SPDGCN [231]. SPDGCN is used as a Riemannian graph neural network backbone. GyroSPD [159]. GyroSPD substitutes the BiMap layer in SPDNet with the AIM1 1 based gyrotranslation S k = W k ⊕AI S k−1 = W k 2 S k−1 W k 2 , where W k is an SPD matrix parameter.
A.2.2
Backbone Networks on the Grassmannian
GyroGr. GyroGr [159] mimics conventional densely connected feedforward networks and is composed of three basic building blocks. Given an ONB Grassmannian matrix U k−1 , the gyrotranslation and ProjMap layers are defined as Gyrotranslation: U k = W k ⊕Gr U k−1 , ProjMap: P k = U k−1 (U k−1 )⊤ .
W k ∈ Gr(p, n),
(A.7) (A.8)
In addition, the pooling is performed via the space of projection matrices [108]. Specifically, Grassmannian data are first mapped to the space of projection matrices via the ProjMap layer. A standard mean pooling operation is applied to the resulting projection matrices. Finally, the pooled matrices are projected back to the ONB 249
A.2. Backbone Networks Grassmannian by SVD. The entire procedure can be expressed as P k = fp U k−1 (U k−1 )⊤ ,
k U k = O1:p ,
SVD
where P k := Ok Σk (Ok )⊤ ,
(A.9)
where fp is a regular mean pooling.
A.2.3
Backbone Networks on Rotation Matrices
LieNet. LieNet [107] is a classical neural network on rotation matrices. Its latent space is the Lie group SON (3) = SO(3) × · · · × SO(3), i.e.,, R = (R1 , . . . , RN ) ∈ SON (3). The group and manifold structures on SON (3) are defined component-wise. For instance, 1 2 R1 ⊙ R2 = (R11 R12 , . . . , RN RN ). There are three basic layers in LieNet: RotMap layer: Rk = W k ⊙ Rk−1 , with W k ∈ SON (3), Rk−1 , if Θ Rk−1 > Θ Rk−1 , ni ,mi mi ,ni mi ,ni , RotPooling layer: Rik = Rk−1 , otherwise,
(A.10) (A.11)
ni ,mi
LogMap layer: Rk = log(Rk−1 ),
(A.12)
where Θ(·) is the Euler angle, and (ni , mi ) are two indices. The RotMap and RotPooling layers mimic the convolution and pooling layers, while the LogMap layer maps rotation matrices into the tangent space for classification. In LieNet, each rotation feature has shape [num, frame, 3, 3], where num and frame denote the spatial and temporal dimensions. The RotPooling layer is applied along either the spatial or temporal dimension, while the RotMap layer is applied along the spatial dimension, with W k of size [num, 3, 3].
A.2.4
Backbone Networks on Hyperbolic Spaces
We briefly recap the hyperboloid layers, Poincaré MLR, Poincaré FC layer, and Poincaré β-concatenation used in the experiments. Throughout this subsection, K < 0. Hyperboloid Neural Networks. Let x ∈ HnK be the input vector. The weight parameters are W ∈ Rm×(n+1) and v ∈ Rn+1 . The Lorentz FC layer [45, Eq. 3] and 250
Appendix A. Experimental Details and Additional Discussions activation layer [15, Eq. 13] are defined in a spacetime manner: Activation: y =
q 2 ∥ψ (xs )∥ − 1/K
(A.13)
,
ψ (xs )
"p # ∥ϕ(W x, v)∥2 − 1/K FC: y = , ϕ(W x, v)
(A.14)
with ϕ(W x, v) = λσ v ⊤ x + b′
W ψ(x) + b , ∥W ψ(x) + b∥
(A.15)
where b ∈ Rm and b′ ∈ R are biases, ψ is an activation function, σ is the sigmoid function, and λ > 0 is a learnable scaling parameter. Poincaré MLR. Lebanon and Lafferty [130] first reformulated the Euclidean MLR p(y=k | x) ∝ exp (⟨ak , x⟩ − bk ) via the point-to-hyperplane distance: p(y=k | x) ∝ exp (sign(⟨ak , x⟩ −bk ) ∥ak ∥ d(x, Hak ,bk )) , Ha,b = {x ∈ Rn | ⟨a, x⟩ − b = 0} ,
where a ∈ Rn \ {0} and b ∈ R.
(A.16) (A.17)
For x ∈ PnK , Ganea et al. [76, Eqs. 24–25] generalized this formulation via geometric reinterpretation, and Shimizu et al. [181, Sec. 3.1] further derived the resulting closed form: p p p 2∥zk ∥ K vk (x) = p |K|x, [z ]⟩ cosh(2 |K|r ) − λ − 1 sinh(2 |K|r ) , asinh λK ⟨ k k k x x |K|
2 −1 where λK is the conformal factor, p(y=k | x) ∝ exp (vk (x)), and x = 2(1 − |K|∥x∥ ) zk [zk ] = ∥zk ∥ . Here, zk ∈ Rn \ {0} and rk ∈ R are parameters. Under the identification ak = zk and bk = ∥zk ∥ rk , the Euclidean limit is limK→0 vk (x) = 4(⟨ak , x⟩ − bk ).
Poincaré FC Layer. Shimizu et al. [181] extended the Euclidean FC layer to the Poincaré ball via point-to-hyperplane distances. In Euclidean spaces, the FC can be written element-wise as yk = ⟨ak , x⟩ − bk with x, ak ∈ Rn and bk ∈ R. Thus, yk equals ∥ak ∥ times the signed distance from x to the hyperplane Hak ,bk . Under this interpretation, the Poincaré FC layer F : PnK → Pm K takes the closed form y=
w p , 1 + 1 + |K|∥w∥2
wk = |K|−1/2 sinh
p |K|vk (x) ,
(A.18)
m where |K| is the magnitude of the curvature. Here, Z = {zk }m k=1 and r = {rk }k=1 parameterize the orientations and biases, and vk (x) is the Poincaré MLR logit.
251
A.2. Backbone Networks Poincaré β-Concatenation. It generalizes the Euclidean concatenation into the hyperbolic Poincaré ball, stabilizing the norm of the Poincaré vector [181, Sec. 3.3]. Given inputs {xi ∈ PnKi }N i=1 , it is defined as Exp0 βn βn−1 v ⊤ , · · · , βn−1 v⊤ 1 1 N N where vi = Log0 (xi ), n = function.
A.2.5
PN
i=1 ni ,
⊤
∈ PnK ,
(A.19)
and βα = B (α/2, 1/2) is defined using the beta
Riemannian Residual Network Backbones
RResNet. Euclidean residual blocks can be written as x(i) = x(i−1) + ni x(i−1) ,
(A.20)
where ni is a network. Katsman et al. [118] generalized this to manifolds by replacing addition with the Riemannian exponential map: x(i) = Expx(i−1) ℓi (x(i−1) ) ,
(A.21)
where ℓi : M → T M outputs a vector field parameterized by the neural network. As Expx (v) = x + v for the Euclidean space, it can be immediately shown that Eq. (A.21) naturally extends Eq. (A.20) to manifolds. On the SPD manifold, the vector field is generated from the eigenvalues. Specifically, the SPD residual block [118, Eqs. 22–23] is Y = ExpX Q diag (f (spec(X))) Q⊤ , (A.22) n where X ∈ S++ , spec(X) contains all eigenvalues, f : Rn → Rn is a neural network, and Q is an orthogonal parameter. Here, diag(·) returns a diagonal matrix from the input vector.
252
Appendix A. Experimental Details and Additional Discussions
A.3
Experimental Details
A.3.1
Lie Group Batch Normalization
A.3.1.1
LieBN on the SPD Manifold
The SPDNet and TSMNet backbones, including SPDNetBN and SPDDSMBN, are reviewed in Sec. A.2.1. We next describe the domain-specific momentum LieBN and implementation details. SPD modeling and preprocessing. The data set descriptions are collected in Sec. A.1. Throughout this thesis, every sample covariance matrix is made strictly positive definite before subsequent manifold operations by adding a small diagonal perturbation, Σ ← Σ + ϵIn with ϵ > 0. Unless otherwise stated, this preprocessing convention is applied to all covariance-based SPD representations. Following the protocol of Brooks et al. [31], each Radar signal is divided into windows of length 20, whose series yields one 20 × 20 SPD covariance matrix. This produces 3,000 covariance matrices equally distributed across 3 classes. For HDM05, each frame consists of 3D coordinates of 31 joints, and each sequence is modeled by a 93 × 93 covariance matrix. Following the protocol of Brooks et al. [31], we trim the data set down to 2,086 sequences scattered throughout 117 classes by removing some under-represented classes. For the FPHA, we follow Wang et al. [207] to represent each sequence as a 63 × 63 covariance matrix. For Hinss2021, we choose the SOTA method, TSMNet [123], as our baseline model. We follow the Python implementation7 of Kobler et al. [123] to carry out preprocessing. In detail, the Python packages MOABB [112] and MNE [84] are used to preprocess the data set. The applied steps include resampling the EEG signals to 250/256 Hz, applying temporal filters to extract oscillatory EEG activity in the 4–36 Hz range, extracting short segments (≤ 3s) associated with a class label, and finally obtaining 40 × 40 SPD covariance matrices. Domain-specific momentum LieBN for EEG classification. Kobler et al. [123] proposed SPDDSMBN as a domain adaptation approach for EEG classification. SPDDSMBN, based on Eq. (3.6), performed normalization of mean and variance on SPD manifolds under the specific AIM. Additionally, SPDDSMBN utilized separate momentum parameters for updating training and testing running statistics, inspired by Yong et al. [226]. Following Kobler et al. [123, Alg. 1], we also present a momentum 7
https://github.com/rkobler/TSMNet
253
A.3. Experimental Details Algorithm 3: Momentum LieBN (MLieBN) Algorithm Input : A batch of activations {Pi }N i=1 over the Lie group {M, ⊕, g}, and a small positive constant ϵ running mean M̄r = E, running variance v̄r2 = 1 for training running mean M̃r = E, running variance ṽr2 = 1 for testing biasing parameter B ∈ M, scaling parameter s ∈ R \ {0}, momentum parameters for training and testing ηtrain , η ∈ [0, 1] Output : Normalized activations {P̃i }N i=1 if training then Compute batch mean Mb and variance vb2 of {Pi }N i=1 ; M̄r ← WFM({1 − ηtrain , ηtrain }, {M̄r , Mb }); v̄r2 ← (1 − ηtrain )v̄r2 + ηtrain vb2 ; M̃r ← WFM({1 − η, η}, {M̃r , Mb }); ṽr2 ← (1 − η)ṽr2 + ηvb2 ; if training then M ← M̄r , v 2 ← v̄r2 ; else M ← M̃r , v 2 ← ṽr2 ; for i ← 1 to N do Centering to the neutral element E: if g is left-invariant then P̄i ← L⊖M (Pi ); else P̄i ← R⊖M (Pi ); Scaling the dispersion: i h P̂i ← ExpE √vs2 +ϵ LogE (P̄i ) Biasing towards parameter B: if g is left-invariant then P̃i ← LB (P̂i ); else P̃i ← RB (P̂i );
LieBN (MLieBN) in Alg. 3. Here η is fixed and ηtrain is defined as 1
ηtrain = 1 − ρ T −1 max(T −t,0) + ρ,
where ρ =
1 , domains_per_batch
(A.23)
where T and t denote the total number of training epochs and the current epoch, respectively. Furthermore, following Kobler et al. [123], we adopt multi-channel mechanisms for domain-specific MLieBN (DSMLieBN), where each domain has its own MLieBN layer. Similar to Kobler et al. [123], we set the biasing parameter equal to the neutral element, and the scaling factor is shared across all domains. We denote Alg. 3 as MLieBN(Pj | B, s, ϵ, η, ηtrain ). Then our DSMLieBN follows DSMLieBN(Pj , i) = MLieBNi (Pj | E, s, ϵ, η, ηtrain ), 254
∀Pj ∈ {Pk }N k=1 ,
(A.24)
Appendix A. Experimental Details and Additional Discussions where i is the index of the domain. We follow the official code of SPDDSMBN8 to implement our DSMLieBN. Thus, the only difference between DSMLieBN and SPDDSMBN is how normalization is performed. Analogous to Thm. 72, computations for DSMLieBN under pullback metrics can also be performed by mapping, calculating, and then remapping. Implementation details. We use the official code of SPDNetBN9 [31] and TSMNet10 [123] to implement our experiments on the SPDNet and TSMNet backbones. For the SPDNet architecture, we compare our LieBN with SPDNetBN [31], which applies the SPDBN (Eqs. (3.2) and (3.3)) to SPDNet. Similar to SPDNetBN, we apply our LieBN after each transformation layer. In the EEG application, one of the state-ofthe-art methods is TSMNet+SPDDSMBN [123], which is a domain adaptation version of Kobler et al. [124]. For a fair comparison, we also implement a domain-specific momentum LieBN, referred to as DSMLieBN. Following Kobler et al. [123], we apply our DSMLieBN before the LogEig layer in TSMNet. We use the standard cross-entropy loss and optimize the parameters with the Riemannian AMSGrad optimizer [18]. The network architectures are represented as {d0 , d1 , . . . , dL }, where the dimension of the parameter in the i-th BiMap layer is di × di−1 . The experiments are conducted with a learning rate of 5e−3 , a batch size of 30, and 200 training epochs on the Radar, HDM05, and FPHA data sets. For the Hinss2021 data set, following Kobler et al. [123], we use a learning rate of 1e−3 with a weight decay of 1e−4 , a batch size of 50, and 50 training epochs. Scoring metrics. In line with the previous work [31, 123], we use accuracy as the scoring metric for the Radar, HDM05, and FPHA data sets, and balanced accuracy (i.e., the average recall across classes) for the Hinss2021 data set. Ten-fold experiments on the Radar, HDM05, and FPHA data sets are carried out with randomized initialization and split (the split is officially fixed for the FPHA data set), while on the Hinss2021 data set, models are fit and evaluated with randomized leave-5%-of-sessions-out (intersession) or leave-5%-of-subjects-out (inter-subject) cross-validation. A.3.1.2
LieBN on Rotation Matrices
Data sets and preprocessing. The data set descriptions are collected in Sec. A.1. Following Huang et al. [107], the G3D experiments use the cross-subject setting, with 8
https://github.com/rkobler/TSMNet https://proceedings.neurips.cc/paper_files/paper/2019/file/ 6e69ebbfad976d4637bb4b39de261bf7-Supplemental.zip 10 https://github.com/rkobler/TSMNet 9
255
A.3. Experimental Details half of the subjects used for training and the other half for testing, while the NTU60 experiments use the cross-view protocol [178]. We use the code11 of Vemulapalli et al. [203] to represent each skeleton sequence as a point on the Lie group SON ×T (3), where N and T denote spatial and temporal dimensions. As preprocessed in Huang et al. [107], we set T to 100, 16, and 64 on the G3D, HDM05, and NTU60 data sets, respectively. LieNet. The LieNet backbone is reviewed in Sec. A.2.3. Note that the official code of LieNet12 is implemented in MATLAB. We use the open-source PyTorch code13 to implement our experiments. To reproduce LieNet more faithfully, we made the following modifications to this PyTorch code. We reimplemented the LogMap and RotPooling layers to make them consistent with the official MATLAB implementation. In addition, we extended the Riemannian computations of Geoopt [125] to SO(3) to enable the direct Riemannian optimization, which is missing from the current package. We apply our LieBN before the LogMap layer. Note that the dimension of input features in LieNet is B × N × T × 3 × 3. We calculate Lie group statistics along the batch and temporal dimensions (B × T ). We denote the LieNet models with our LieBN-Left and LieBNRight as LieNetLieBN-Left and LieNetLieBN-Right, respectively. Implementation details. We find that SGD is the most effective optimizer for LieNet, and thus, we adopt it for our experiments. The learning rate is set to 1e−2 . The batch sizes are 30, 30, and 256 for the G3D, HDM05, and NTU60 data sets, respectively. On the NTU60 data set, the learning rate is reduced by a factor of 10 upon model convergence, specifically at the 5th and 25th epochs for LieNetLieBN and LieNet, respectively. For each model, we apply torch.nn.utils.clip_grad_norm_ with max_norm=5 to the transformation matrix in the final FC layer. A.3.1.3
LieBN on Correlation Matrices
We follow the same settings as the experiments on the SPD manifold with respect to the backbone architecture, batch size, number of training epochs, optimizer, and learning rate. The network architecture can be denoted as BiMap-[Power-Cov2CorLieBN-Cor]-LogEig, where Power denotes the matrix power and Cov2Cor is Cor(·) : n S++ → Cor+ (n). The matrix powers used for each data set are presented in Tab. A.4. A single iteration is sufficient to achieve saturated network performance with respect to calculating D and D⋆ in OLM and LSM, except for D⋆ on the HDM05 data set, which requires convergence with a maximum of 20 iterations. 11
https://ravitejav.weebly.com/kbac.html https://github.com/zhiwu-huang/LieNet 13 https://github.com/hjf1997/LieNet 12
256
Appendix A. Experimental Details and Additional Discussions Metric
Data Set
ECM
LECM
OLM
LSM
0.75 -0.5
0.5 -0.25
0.5 -0.25
-0.5 -0.25
HDM05 FPHA
Table A.4: Matrix powers in LieBN-Cor under different metrics on each data set.
A.3.2
Gyrogroup Batch Normalization
A.3.2.1
GyroBN on the Grassmannian
For HDM05, under-represented clips are removed, yielding 2,086 instances over 117 classes. For NTU60 and NTU120, we focus on mutual actions and adopt the cross-view and cross-setup protocols, respectively [178, 138]. Following Nguyen and Yang [159], each sequence is represented as a Grassmannian matrix of size 93 × 10, 150 × 10, and 150 × 10 for HDM05, NTU60, and NTU120, respectively. Because GyroBN is inserted after the first pooling layer, the corresponding inputs to GyroBN have sizes 47 × 10, 75 × 10, and 75 × 10, as reported in Tab. 3.16. The GyroGr backbone is reviewed in Sec. A.2.2. Trivialization. Following Nguyen and Yang [159], we apply the trivialization strategy reviewed in Sec. 2.7 to the Grassmannian parameters in the gyrotranslation and GyroBN layers. Each Grassmannian parameter U ∈ Gr(p, n) is parameterized by a matrix U ∈ R(n−p)×p such that " # i h 0 −U⊤ (A.25) = U U ⊤ , Iep,n , U 0 where (·) = LogIep,n (·). The parameter U can be retrieved by U = exp
h
U U ⊤ , Iep,n
i
Ip,n = exp
"
0 −U⊤ U 0
#!
Ip,n .
(A.26)
This reparameterization represents the trainable Grassmannian parameters by Euclidean coordinates, thereby allowing the direct use of PyTorch optimizers [169] and avoiding direct Riemannian updates of these parameters. 257
A.3. Experimental Details
A.3.3
Riemannian Multinomial Logistic Regression
A.3.3.1
RMLR on the SPD Manifold
This subsection offers additional details on the experiments on SPD MLRs. Backbone networks and LogEig MLR. The SPDNet, TSMNet, SPDNetBN, SPDDSMBN, and SPDGCN backbones are reviewed in Sec. A.2.1, while RResNet is reviewed in Sec. A.2.5. In the SPD baseline models considered here, the Euclidean MLR in the codomain of matrix logarithm (matrix logarithm + FC + softmax) is used for classification. Following the terminology introduced in Sec. 4.2.4, we call this classifier the LogEig MLR. The LogEig MLR is the Euclidean classifier in the tangent space at the identity, which might distort the innate geometry of the SPD manifold. SPD modeling and preprocessing. We use the same SPD modeling and preprocessing as the LieBN experiments in Sec. A.3.1.1. Implementation details. For SPDNet [106] and TSMNet [123], we follow the official PyTorch code of SPDNetBN14 and TSMNet15 to implement our experiments. To evaluate the performance of our intrinsic classifiers, we substitute the LogEig MLR in SPDNet and TSMNet with our SPD MLRs. We implement our SPD MLRs induced by five parameterized metrics. On the Radar and HDM05 data sets, the learning rate is 10−2 , and the batch size is 30. On the Hinss2021 data set, following Kobler et al. [123], the learning rate is 10−3 with a 10−4 weight decay, and the batch size is 50. The maximum numbers of training epochs are 200, 200, and 50, respectively. We use the standard cross-entropy loss as the training objective and optimize the parameters with the Riemannian AMSGrad optimizer [18]. RResNet [117]. We focus on the AIM-based RResNet and use the official code16 and suggested network settings to implement the experiments with RResNet. We conduct 10-fold and 5-fold experiments on the HDM05 and NTU60 data sets, respectively. Since RResNet is developed based on SPDNet, we use the same learning settings as SPDNet for the action recognition task and borrow the best (θ, α, β) from Tab. 4.4 for our SPD MLRs under the RResNet backbone. NTU60 SPD modeling. For NTU60 [178], each frame contains the 3D coordinates of 25 body joints and is therefore represented by a 25 × 3 = 75-dimensional coordinate vector. Following Katsman et al. [117], each sequence is modeled as a 75 × 75 temporal covariance matrix, and evaluation follows the cross-view protocol [178]. 14
https://proceedings.neurips.cc/paper_files/paper/2019/file/ 6e69ebbfad976d4637bb4b39de261bf7-Supplemental.zip 15 https://github.com/rkobler/TSMNet 16 https://github.com/CUAI/Riemannian-Residual-Neural-Networks
258
Appendix A. Experimental Details and Additional Discussions Data Sets
(θ, α, β)-AIM
(θ, α, β)-EM
(α, β)-LEM
2θ-BWM
θ-LCM
Disease Cora Pubmed
(0.25,1,0) (0.5,1,0) (0.5,1,0)
(0.25,1,0) (0.25,1,1/9) (0.5,1,0)
(1,1) (1,1/9) (1, −1/3)
0.25 0.25 0.25
0.5 0.5 0.5
Table A.5: (θ, α, β) of SPD MLRs on the SPDGCN backbone. SPDGCN [231]. We use the official code17 and the suggested network settings in Zhao et al. [231]. Note that SPDGCN with SPD MLR retains the same network settings as vanilla SPDGCN. Tab. A.5 presents the hyperparameters (θ, α, β) on different data sets. Network architectures. We denote the network architecture as [d0 , d1 , · · · , dL ], where the dimension of the parameter in the i-th BiMap layer (Sec. A.2.1) is di × di−1 . For SPDNet, we also validate our SPD MLRs under different network architectures on the Radar and HDM05 data sets. The network architectures on the Radar data set are [20, 16, 8] for the 2-block configuration and [20, 16, 14, 12, 10, 8] for the 5-block configuration, while on the HDM05 data set, the network architectures are [93, 30] for 1-block, [93, 70, 30] for 2-block, and [93, 70, 50, 30] for 3-block. For TSMNet, the 1-block architecture is [40, 20]. Scoring metrics and evaluation protocols. We use the same scoring metrics and evaluation protocols as Sec. A.3.1.1. For the graph-learning experiments, following Zhao et al. [231], we report the 10-fold average and maximum node-classification accuracy. Hyperparameters. We implement the SPD MLRs induced by not only five standard metrics, i.e., LEM, AIM, EM, LCM, and BWM, but also five families of parameterized metrics. Therefore, in our SPD MLRs, we have a maximum of three hyperparameters, i.e., θ, α, β, where (α, β) are associated with O(n)-invariance and θ controls deformation. For (α, β) in (θ, α, β)-LEM, (θ, α, β)-AIM, and (θ, α, β)-EM, recalling Eq. (2.93), α is a scaling factor, while β measures the relative significance of traces. As scaling is less important [192], we set α = 1. As for the value of β, we select it from a predefined set: {1, 1/n, 1/n2 , 0, −1/n + ϵ, −1/n2 }, where n is the dimension of the input SPD matrices in SPD MLRs. The purpose of including ϵ ∈ R+ is to ensure the positive definiteness of the inner product, i.e., α + nβ > 0 and hence (α, β) ∈ ST. These chosen values for β allow for amplifying, neutralizing, or suppressing the trace components, depending on the characteristics of the data sets. For the deformation factor θ, we roughly select its value around its deformation boundary, i.e., [0.25, 1.5] 17
https://github.com/andyweizhao/SPD4GNNs
259
A.3. Experimental Details Metric
(θ, α, β)-AIM
(θ, α, β)-EM
θ-LCM
2θ-BWM
Candidate Values
{0.25, 0.5, 0.75, 1, 1.25, 1.5}
{0.25, 0.5, 1, 1.5}
{0.5, 1, 1.5}
{0.25, 0.5, 0.75}
Table A.6: Candidate values for hyperparameters in SPD MLRs. for (θ, α, β)-AIM, [0.5, 1.5] for θ-LCM, [0.25, 1.5] for (θ, α, β)-EM, and [0.25, 0.75] for 2θ-BWM. The detailed values are listed in Tab. A.6. A.3.3.2
RMLR on Rotation Matrices
LieNet backbone. The LieNet backbone is reviewed in Sec. A.2.3. In the official MATLAB implementation, the LogMap layer uses the Euler axis–angle representation. Classification is performed using the Euler axis–angle representation, followed by an FC layer and a softmax layer. As the axis–angle representation is equivalent to the matrix logarithm, we call this classifier LogEig MLR as well. This classifier is, therefore, also non-intrinsic. Preprocessing. We use the shared rotation-matrix modeling, preprocessing, LieNet implementation, and SGD optimizer choice described in Sec. A.3.1.2. Following Huang et al. [107], the Lie MLR experiments use the G3D and HDM05 data sets. We trim HDM05 by removing under-represented sequences, resulting in 2,326 sequences across 122 classes. Lie MLR. We use our Lie MLR to replace the axis–angle classifier in LieNet and call the resulting network LieNet+LieMLR. To alleviate the computational burden, we set each Pk to have shape [num, 3, 3], where num is the spatial dimension of the input of the Lie MLR layer. In other words, Pk is shared in the temporal dimension. We adopt PyTorch3D [174] to calculate the matrix logarithm. Due to the instability of pytorch3d.transforms.so3_log_map, we first use pytorch3d.transforms.matrix_ to_axis_angle to calculate the rotation axis and angle and then convert this representation into the matrix logarithm18 . Training details. Following Huang et al. [107], we focus on the 3-block and 2block architectures for the G3D and HDM05 data sets, respectively, which are the suggested architectures for these two data sets. The learning rate is 10−2 on both data sets, and we further set the weight decay to 10−5 on the G3D data set. For LieNet and LieNet+LieMLR, we use torch.nn.utils.clip_grad_norm_ for gradient clipping with a clipping factor of 5. The clipping is imposed on the dimensionality reduction weight in the final FC linear layer of LieNet or, accordingly, A = {A1 , . . . , AC } in the 18
https://github.com/facebookresearch/pytorch3d/issues/188
260
Appendix A. Experimental Details and Additional Discussions Lie MLR layer of LieNet+LieMLR. Scoring metrics. For the G3D data set, following LieNet [107], we adopt a 10-fold cross-subject test setting, where half the subjects are used for training and the other half are employed for testing. For the HDM05 data set, following Huang et al. [107], we randomly select half of the sequences for training and the rest for testing. Due to the instability of LieNet, we conduct 20-fold experiments and select the best 10 folds to evaluate the performance.
A.3.4
Proper Velocity Neural Networks
Common Implementations. We use the trivialization strategy reviewed in Sec. 2.7 in our MLR, FC, and GyroBN layers. Consequently, all trainable manifold-valued parameters in PVNN are represented by Euclidean coordinates and optimized using standard Euclidean optimizers. A.3.4.1
Image Classification
The CIFAR-10 and CIFAR-100 data sets and their PyTorch splits are described in Sec. A.1.3. Following Bdeir et al. [15], we use data augmentation that includes random cropping with padding of 4 pixels and random horizontal flipping. Implementation Details. We implement the experiments using the official code19 of Bdeir et al. [15]. All models share a common backbone, which consists of a ResNet-18 encoder followed by a hyperbolic MLR classifier. Except for PV MLR without Exp0 , the output embedding of the ResNet-18 backbone is mapped to the target hyperbolic space via the exponential map at the identity e, that is, Expe (x). Here, e = 0 for the Poincaré and PV spaces, and e = 0 for the hyperboloid model. All models are trained from scratch. Optimization is performed using SGD [175] with an initial learning rate of 0.1, a momentum of 0.9, and a weight decay of 5 × 10−4 . Training is conducted with a batch size of 128 for 200 epochs. The learning rate is decayed by a factor of γ = 0.2 at epochs 60, 120, and 160. The curvature for the PV space is set as K = −0.5. A.3.4.2
Graph Learning
The descriptions and summary statistics of the Disease, Airport, PubMed, and Cora data sets are collected in Sec. A.1.4. Implementation Details. We adopt the official code of HGCN20 [38] to conduct 19 20
https://github.com/kschwethelm/HyperbolicCV https://github.com/HazyResearch/hgcn
261
A.3. Experimental Details Hyperparameter
Disease
Airport
PubMed
Cora
Learning rate Dropout Curvature
0.01 0.4 -0.3
0.01 0.4 -0.3
0.05 0.6 -1.0
0.05 0.6 -1.0
Table A.7: Hyperparameters for PVNN that vary across graph data sets. Setting
Epochs
Batch size
Weight decay
Value
2000
128
5 × 10−4
Table A.8: Hyperparameters for PVNN that are shared across graph data sets. experiments. The features of each node are embedded into the hyperbolic space via the exponential map at the identity. The hyperbolic network consists of two FC layers: the first maps the input feature dimension to a 16-dimensional hidden representation, and the second maps from 16 to 16. Each FC layer is followed by an activation function. An MLR layer is then used for classification. All models are trained using the Adam optimizer [122]. We evaluate performance every 10 epochs and employ early stopping with a patience of 200 evaluations, restoring the checkpoint with the best test accuracy. Tabs. A.7 and A.8 summarize the hyperparameters for PVNN. For KNN [147], HNN [76], HNN++ [181], and LNN [15], we follow their original papers to implement the experiments. Tab. A.9 summarizes the hyperbolic layers used in each model. A.3.4.3
Genomic Sequence Learning
The description and complete summary of TEB are collected in Sec. A.1.5. We focus on five of its seven data sets. Implementation Details. For the Euclidean CNN and the hyperbolic CNN baseline (HCNN-S), we directly use the results reported in the original paper [119, Tab. 2]. Our PVCNN architecture follows their implementation21 . Each DNA sequence is represented as a length-L sequence with 4 input channels. We first apply a PV convolution that maps the 4 input channels to 32 channels, followed by PV TBN and a tanh tangent activation. A second PV convolution layer is then applied. The final PV feature is concatenated and passed through an FC layer, and finally classified with a PV MLR head. The curvature is initialized at K = −0.5 and learned during training. We train for 100 epochs with a step learning-rate schedule, using milestones at epochs 60 and 21
https://github.com/rrkhan/HGE
262
Appendix A. Experimental Details and Additional Discussions Model
FC layer
Activation
MLR
PVNN KNN HNN HNN++ LNN
PV FC in Thm. 127 Log0 (W Exp0 (x)) Log0 (W Exp0 (x)) Poincaré FC [181] Lorentz FC [45]
Exp0 (σ (Log0 (x))) Exp0 (σ (Log0 (x))) Exp0 (σ (Log0 (x))) Exp0 (σ (Log0 (x))) Lorentz activation [15]
PV MLR in Thm. 126 Euclidean MLR after Exp0 Poincaré MLR [76] Poincaré MLR [181] Lorentz MLR [15]
Table A.9: Summary of the hyperbolic layers used in the graph node classification models. Setting
Value
Optimizer Learning rate Weight decay Batch size Dropout Adam (β1 , β2 )
Adam 1e−4 2e−2 100 0.1 (0.9, 0.999)
Table A.10: Hyperparameters for TEB. 85 with a decay factor of 0.1. For the PV FC layer, σ in Eq. (5.29) is set to tanh. All other hyperparameters are summarized in Tab. A.10.
A.3.5
Hyperbolic Busemann Neural Networks
A.3.5.1
Image Classification
Implementation Details. For CIFAR-10/100 and Tiny-ImageNet, we follow Bdeir et al. [15, App. C.1]. For ImageNet-1k, we follow Guo et al. [91, Sec. 4]. Tab. A.11 summarizes the data-set-specific hyperparameters. For the hyperbolic MLR, before mapping into the hyperbolic space, we clip the feature vector by r CLIP (x; r) = min 1, x ∥x∥
(A.27)
where r > 0 is a hyperparameter. The clipped Euclidean embedding is projected via the exponential map to the target hyperbolic space: Expe (CLIP(x; r)). For the Lorentz model, the clipping parameter is r = 1 on CIFAR-10/100 and r = 4 on Tiny-ImageNet and ImageNet-1k. On the Poincaré ball, r = 1 on all four data sets. All methods are implemented in PyTorch and trained with cross-entropy loss. The results of MLR, PMLR, and LMLR on CIFAR-10/100 and Tiny-ImageNet are copied 263
A.3. Experimental Details
Hyperparameter
CIFAR-10/100 Tiny-ImageNet
ImageNet-1k
Epochs Batch size Initial learning rate LR schedule Weight decay Optimizer Precision Curvature K
200 128 0.1 60, 120, 160; γ = 0.2 5e−4 SGD 32-bit −1
100 256 0.1 30, 60, 90; γ = 0.1 1e−4 SGD 32-bit −1
Table A.11: Summary of hyperparameters used in the image classification task. from Bdeir et al. [15, Tab. 1], while those of PBMLR-P on CIFAR-10/100 are copied from Nguyen et al. [162, Tab. 2]. The remaining results are obtained by our careful implementation. A.3.5.2
Genome Sequence Learning
Implementation Details. We mainly follow the official implementations of Khan et al. [119] for data processing, model architecture, and training. We adopt a simple CNN with three convolutional blocks followed by dense, ReLU-activated layers to extract features [119, Fig. 4]. Before classification, the features are clipped and mapped to the target hyperbolic space as in Eq. (A.27), then passed to either a prior hyperbolic MLR or our BMLR head. The clipping factor defaults to r = 1, with the following exceptions on GUE: on the Lorentz model, r = 2.0 for Covid variant classification and r = 5.0 for species classification; on the Poincaré ball, r = 2.0 for species classification. Following Khan et al. [119], we treat the curvature K as a learnable parameter initialized as −1. The remaining hyperparameters are listed in Tab. A.12. All models share these hyperparameters, except LMLR on Covid variant classification, for which we set the weight decay to 1e−3 to ensure convergence. All methods are implemented in PyTorch and trained with cross-entropy loss. Results are obtained from our reimplementation. A.3.5.3
Node Classification
Implementation Details. We follow the official implementations of HGCN [38] and PBMLR [162] and conduct experiments on both the Poincaré ball and the Lorentz model. We adhere to their experimental settings. The only changes are weight decay 264
Appendix A. Experimental Details and Additional Discussions Batch size Epochs Optimizer β1 , β2 Initial learning rate LR schedule Weight decay Initial curvature K
100 100 Adam 0.9, 0.999 1e−4 60, 85; γ = 0.1 0.1 −1
Table A.12: Hyperparameters for genome sequence learning. Space
Weight decay
Dropout
PnK LnK
1e − 4, 1e − 5, 1e − 3, 1e − 3 1e − 4, 5e − 5, 1e − 3, 1e − 3
0.3, 0, 0, 0.2 0, 0, 0, 0.3
Table A.13: Hyperparameters for node classification on Disease, Airport, PubMed, and Cora. and dropout. We train with cross-entropy loss and the Adam optimizer [122] for 5000 epochs with a learning rate of 1e−2 , curvature set to −1, embedding dimension 16, and three GCN layers. We tune weight decay and dropout and report the values in Tab. A.13. All methods are implemented in PyTorch. The results of HGCN-PMLR and HGCN-PBMLR-P on the Poincaré ball are taken from Nguyen et al. [162, Tab. 11]. Results for the remaining baselines are obtained from our reimplementation following the original settings.
A.3.5.4
Link Prediction
Implementation Details. We follow the official implementations of HNN [76], HNN++ [181], and HyboNet [45], and adopt the experimental protocol of Chami et al. [38] for link prediction. The encoder consists of two fully connected layers: the first maps the input features to 16, and the second maps 16 to 16. Each FC layer is instantiated as either our BFC or an existing hyperbolic FC layer. After each FC, we apply the activation Expe (ReLU(Loge (x))), where e denotes the model origin. Following Chami et al. [38], this activation is disabled on Cora. As in the Möbius layer, we apply a gyro bias after each FC, that is, x ⊕H b. We train with Adam [122] at a learning rate of 1e−2 and tune weight decay and FC dropout. For BFC, we use ϕ = tanh on Airport and Cora and the identity map on the other two data sets. 265
A.3. Experimental Details
A.3.6
Full-Rank Correlation Networks
The Radar, HDM05, FPHA, and NTU120 data sets are described in Sec. A.1. For HDM05 and FPHA, each sequence is normalized for body-part length, scale, and view using the preprocessing of Vemulapalli et al. [203]. For NTU120, we follow Chen et al. [46] to preprocess the data. A.3.6.1
Input Data
Correlation Input in CorNets. Following Wang et al. [210], Nguyen et al. [161], we first model each sample as a multi-channel SPD tensor. For HDM05 and FPHA, the skeleton preprocessing and multi-channel covariance construction follow the GyroGr pipeline of Nguyen and Yang [159]; CorNet subsequently converts every covariance matrix into a full-rank correlation matrix by 1
1
n Cor : S++ ∋ Σ 7−→ C = D(Σ)− 2 ΣD(Σ)− 2 ∈ Cor+ (n).
(A.28)
For Radar, we follow Wang et al. [210] and use temporal convolution followed by covariance pooling to obtain a multi-channel covariance tensor of shape [c, 20, 20]. After preprocessing, the input correlation tensor shapes are [7, 20, 20], [3, 28, 28], [9, 28, 28], and [6, 28, 28] on Radar, HDM05, FPHA, and NTU120, respectively. A.3.6.2
Implementation Details
SPD Baselines. We follow the official PyTorch code of SPDNetBN22 to implement SPDNet and SPDNetBN. For LieBN23 , we focus on the instantiation under LCM [137], while for RResNet24 , we implement the ones induced by AIM [171] and LEM [9]. For SPD MLR25 , we implement the one based on LCM. Due to the lack of official code, Gyro-based models are carefully reimplemented from their original papers. Following Nguyen et al. [161], GyroSPD++ combines an AIM-based convolution with an LEMbased MLR. Grassmannian Baselines. Since GrNet is officially implemented in MATLAB, we carefully re-implemented it using PyTorch. Additionally, as both GyroGr and GyroGrScaling do not release official code, we re-implemented them based on the original paper 22
https://proceedings.neurips.cc/paper_files/paper/2019/file/ 6e69ebbfad976d4637bb4b39de261bf7-Supplemental.zip 23 https://github.com/GitZH-Chen/LieBN 24 https://github.com/CUAI/Riemannian-Residual-Neural-Networks 25 https://github.com/GitZH-Chen/SPDMLR
266
Appendix A. Experimental Details and Additional Discussions Data Set
Model
Optimizer
lr
wd
Matrix Power
Converged Epoch
Radar
CorNet-ECM CorNet-LECM CorNet-OLM CorNet-LSM CorNet-PHCM
Adam Adam Adam Adam Adam
1e−2 1e−2 1e−2 1e−2 1e−2
N/A N/A N/A N/A N/A
1.5 -0.25 -0.25 0.75 0.75
50 50 50 50 50
HDM05
CorNet-ECM CorNet-LECM CorNet-OLM CorNet-LSM CorNet-PHCM
Adam Adam SGD Adam Adam
1e−3 1e−4 5e−2 1e−3 1e−2
1e−3 1e−3 1e−3 N/A N/A
0.125 0.5 0.25 -0.75 -0.25
100 150 200 50 50
FPHA
CorNet-ECM CorNet-LECM CorNet-OLM CorNet-LSM CorNet-PHCM
Adam Adam Adam Adam Adam
5e−3 5e−4 1e−4 1e−3 1e−3
N/A 1e−4 N/A N/A 1e−4
-0.25 -0.5 -1 -1 -0.5
150 150 50 50 150
NTU120
CorNet-ECM CorNet-LECM CorNet-OLM CorNet-LSM CorNet-PHCM
SGD SGD SGD SGD Adam
1e−2 1e−2 5e−3 1e−3 1e−3
N/A N/A N/A N/A N/A
0.25 0.25 0.25 0.25 0.25
50 50 50 50 50
Table A.14: Hyperparameters in CorNets. [159]. For all Grassmannian comparative methods, we use SGD [175] with a learning rate of 5e−2 . CorNets. On all four data sets, we employ a single convolutional kernel for global convolution, i.e., applying a global receptive field across the channel dimension. The output dimensions of the correlation convolutional layer are 8 × 8, 26 × 26, 26 × 26, and 11 × 11 for the Radar, HDM05, FPHA, and NTU120 data sets, respectively.
We primarily use the Adam [122] and SGD [175] optimizers. Inspired by the deformation effect on the latent SPD geometries by the matrix power over the SPD manifold illustrated in Fig. 4.2, we apply the matrix power before correlation modeling (Cor(·)) as activation. In particular, when the data are centered at zero and power is −1, Cor(Σ−1 ) corresponds to the partial correlation matrix of the covariance matrix Σ [191, Lem. 1.6]. The batch size is set to 30, and training is capped at 200 epochs, although most cases converge in fewer than 150 epochs. Due to the different correlation geometries, the hyperparameters vary for CorNets under different geometries. Tab. A.14 summarizes all the hyperparameters. Extra Computational Details for OLM and LSM Layers. For the MLR, FC, and convolutional layers induced by OLM and LSM, the key computations involve Exp◦ and Log⋆ , which depend on the calculations of D and D⋆ . In our experiments, we empirically observe that iterating until convergence is more effective for D, whereas a single step of Newton’s method generally performs best for D⋆ . Accordingly, we set D to iterate until convergence, leveraging Thm. 145 for accurate backpropagation. For 267
A.3. Experimental Details D⋆ , we adopt a single iteration in Newton’s method and use automatic differentiation (autograd) through this single step for backpropagation.
A.3.6.3
Additional Details on Visualization
We provide additional interpretations of the low-dimensional visualization and clarify how the decision hyperplanes in Fig. 5.5 are obtained. The correlation-in-SPD visualization is discussed in Sec. 2.9.2. Visualization of Low-Dimensional SPD and Correlation Matrices. Any 2 can be written as 2 × 2 covariance matrix in S++ Σ=
! a b , b d
a > 0, d > 0, ad − b2 > 0.
(A.29)
2 with the interior of the Embedding Σ into R3 via the map Σ 7→ (a, b, d) identifies S++ quadratic cone (a, b, d) ∈ R3 | a > 0, d > 0, ad − b2 > 0 , (A.30)
which is an open cone in R3 .
For 2 × 2 correlation matrices, any C ∈ Cor+ (2) has the form C=
! 1 r , r 1
r ∈ (−1, 1).
(A.31)
Thus, Cor+ (2) is one-dimensional. Embedding C into R3 as (1, r, 1) yields a line segment 2 inside the cone corresponding to S++ .
For 3 × 3 correlation matrices, any C ∈ Cor+ (3) is parameterized by its off-diagonal entries (r12 , r13 , r23 ): 1 r12 r13 C = r12 1 r23 . (A.32) r13 r23
1
Embedding C into R3 via C 7→ (r12 , r13 , r23 ) produces an open elliptope in R3 . This is the representation of Cor+ (3) used in Fig. 5.5, where each point in the elliptope corresponds to one 3 × 3 correlation matrix. Construction of Fig. 5.5. For ECM, LECM, OLM, and LSM, the decision hyperplane in the correlation MLR is the Riemannian hyperplane in Sec. 4.3.2.1 specialized 268
Appendix A. Experimental Details and Additional Discussions to M = Cor+ (n): HA,P = X ∈ Cor+ (n) | ⟨LogP (X), A⟩P = 0 ,
P ∈ Cor+ (n), A ∈ TP Cor+ (n). (A.33) + In Fig. 5.5, we focus on Cor (3) and visualize it as the open elliptope in R3 via the embedding C ∈ Cor+ (3) 7−→ (C21 , C31 , C32 ) ∈ R3 . (A.34) Given a Log-Euclidean metric and parameters (A, P ), each correlation matrix C is first mapped to the tangent space at P by LogP (C), and we evaluate the linear form ⟨LogP (C), A⟩P . The set of points in the elliptope where this scalar equals zero corresponds to the decision hyperplane HA,P and is plotted as the separating surface. For PHCM, the margin hyperplane is defined in the β-concatenated Poincaré embedding. Let Φ be the diffeomorphism in Eq. (5.79) that maps C ∈ Cor+ (n) to the polyPoincaré space PPn−1 , and let β-concatenation be the Poincaré operation in Sec. 5.4.3.2. We define n(n − 1) , (A.35) x̃(X) = β-concat (Φ(X)) ∈ PN , N= 2 and the PHCM hyperplane n o + Ha,p = X ∈ Cor (n) | Logp (x̃(X)), a p = 0 ,
p ∈ PN , a ∈ T p P N .
(A.36)
Here PN = x ∈ RN | ∥x∥2 < 1 is the N -dimensional unit Poincaré ball. In Fig. 5.5, we first map each correlation matrix C ∈ Cor+ (3) to x̃(C) ∈ PN , apply the Poincaré logarithm Logp at a reference point p, and then visualize the zero level set of the linear form Logp (x̃(C)), a p as the PHCM decision hyperplane.
A.3.7
Adaptive Log-Euclidean Metrics
A.3.7.1
Data Sets and Settings
Although the proposed ALog layers can be plugged into existing SPD networks, we focus on the SPDNet framework [106]. We follow the PyTorch code provided by SPDNetBN26 to reproduce SPDNet and SPDNetBN and implement our approaches. Following previous work [106, 31], we evaluate our methods on HDM05 [153], FPHA [80], and AFEW [68]. The data set descriptions are given in Sec. A.1. We use the HDM05 and FPHA temporal covariance representations and the filtered HDM05 setting 26
SPDNetBN supplementary code
269
A.3. Experimental Details specified in Sec. A.3.1.1. For AFEW, we use the released pre-trained FAN27 [149] to extract deep features and establish a 512 × 512 temporal covariance matrix for each video. We denote the dimensions of the transformation layers in the SPDNet backbone by {d0 , d1 , . . . , dL }. Following the settings in Brooks et al. [31], all networks are trained using the default RSGD [17] with a fixed learning rate γ and a batch size of 30. To make ALog start from the vanilla matrix logarithm, the parameters in MUL, DIV, and RELU are initialized as 1, 1, and e, respectively. By abuse of notation, SPDNet-ALog-MUL is abbreviated as ALog-MUL, denoting that we substitute the LogEig layer in SPDNet with the proposed ALog optimized by MUL.
A.3.7.2
Implementation Details of Additional Applications
NTU60 data set. For NTU60, we use the cross-view protocol and 75 × 75 covariance representation specified in Sec. A.3.3.1. As reported in Sec. 6.2.4, MUL shows the best performance. Therefore, we view A in Eq. (6.31) as the parameter for all experiments. In the following, we discuss in detail the specific implementation of each method. LieBN. We follow the official code28 to implement the experiments. The learning rate is 5e−2 . Since our LieBN-ALEM shows early convergence, we set the numbers of training epochs to 150, 50, and 30 for the [93, 30], [93, 70, 30], and [93, 70, 50, 30] architectures. Other settings are the same as Sec. 6.2.4. RResNet. We follow the official code29 to implement the experiments. For the HDM05 data set, we use RSGD [17] with a 5e−2 learning rate for 200 training epochs. For the NTU60 data set, we use Riemannian AMSGrad [17] with a 1e−2 learning rate for 50 training epochs. We adopt the architectures of [93, 30] and [75, 30] on these two data sets. Gyro MLR. Since the code of gyro MLR is not publicly available, we carefully re-implement the gyro MLR in Nguyen and Yang [159]. We adopt an architecture of [75, 30] under an SGD optimizer. The batch size and number of training epochs are 30 and 200, respectively. 27
https://github.com/Open-Debin/Emotion-FAN https://github.com/GitZH-Chen/LieBN 29 https://github.com/CUAI/Riemannian-Residual-Neural-Networks 28
270
Appendix A. Experimental Details and Additional Discussions
A.3.8
Product Cholesky Metrics
A.3.8.1
Implementation Details of SPD Neural Networks
Backbone networks. SPDNet and GyroSPD are reviewed in Sec. A.2.1, while RResNet is reviewed in Sec. A.2.5. Data sets and preprocessing. The Radar, HDM05, and FPHA data sets are described in Sec. A.1. The global covariance representations for SPDNet and RResNet follow Sec. A.3.1.1; the GyroSPD-specific multichannel construction is given below. SPD MLR. We follow the official PyTorch code30 to implement the SPD MLR developed in Sec. 4.2.2. Due to the lack of official code, the GyroSPD backbone is carefully reimplemented in PyTorch following the original paper [159]. For simplicity, we set M in (θ, M)-BWCM to the identity matrix. SPDNet. Following Huang and Van Gool [106], we replace the vanilla tangent classifier (LogEig + FC + softmax) in SPDNet with SPD MLRs induced by AIM, LEM, LCM, θ-PCM, and θ-BWCM. We use a Riemannian AMSGrad [17] with a learning rate of 1e−2 , a batch size of 30, and a maximum of 200 epochs. We denote the architectures by [d0 , d1 , . . . , dL ], where di is the output dimension of the i-th BiMap layer. Following prior work [106, 31], we adopt [20, 16, 8] on Radar, [63, 33] on FPHA, and [93, 30], [93, 70, 30], [93, 70, 50, 30] for 1-, 2-, and 3-block variants on HDM05. As shown in Tab. 3.9c, matrix power improves the performance of Cholesky-based metrics on FPHA. We therefore apply a power of −0.25 before SPD MLR layers under LCM, θ-PCM, and θ-BWCM. For θ-PCM on FPHA, we further adopt a weight decay of 1e−4 . The deformation factor θ is reported in Tab. A.15. GyroSPD. Following Nguyen and Yang [159], the backbone consists of one gyrotranslation layer followed by an SPD MLR. We compare metrics based on LEM, LCM, AIM, and our proposed geometries under the same settings: Riemannian AMSGrad with a learning rate of 1e−2 and a batch size of 30. The models are trained for up to 100, 100, and 50 epochs on Radar, HDM05, and FPHA, respectively. Similarly, a matrix power of 0.25 is applied on FPHA for LCM and our Cholesky-based metrics. The deformation factor θ is reported in Tab. A.15. SPD input of SPDNet. We use the global covariance representations specified in Sec. A.3.1.1. SPD input of GyroSPD. For Radar, the input is the same as SPDNet. For HDM05 and FPHA, we follow Nguyen and Yang [159] to model each sample into a multi-channel covariance tensor [c, n, n]. Specifically, we first identify the closest left 30
https://github.com/GitZH-Chen/SPDMLR
271
A.3. Experimental Details Backbone
Metric
Radar
HDM05
FPHA
SPDNet
θ-PCM θ-BWCM
-1.5 -0.75
-0.5 -1.5
0.75 1
GyroSPD
θ-PCM θ-BWCM
-0.75 -1.5
-0.75 -1.5
-0.75 -0.5
Table A.15: Hyperparameter θ in θ-PCM and θ-BWCM. It is selected from the candidate values used for SPDNet and GyroSPD. (right) neighbor of each joint based on its distance to the hip (wrist) joint, and then combine the 3D coordinates of each joint and those of its left (right) neighbor to create a feature vector for the joint. For a given frame t, we compute its Gaussian embedding [143]: " # ⊤ 1 Σ + µ (µ ) µ t t t t , (A.37) Yt = (det Σt )− n+1 (µt )⊤ 1 where µt and Σt are the mean vector and covariance matrix computed from the set of feature vectors within the frame. The lower part of the matrix log (Yt ) is flattened to obtain a vector ṽt . All vectors ṽt within a time window [t, t + c − 1], where c is determined from a temporal pyramid representation of the sequence (the number of temporal pyramids is set to 2 in our experiments), are used to compute a covariance matrix as t+c−1 1 X e Σt = (ṽi − v t ) (ṽi − v t )⊤ , (A.38) c i=t
P e t } are the covariance matrices that we need. where v t = 1c t+c−1 ṽi . The resulting {Σ i=t On FPHA, we generate the covariances based on three sets of neighbors: left, right, and vertical (bottom) neighbors. After preprocessing, the input covariance matrices are [3, 28, 28] and [9, 28, 28] on HDM05 and FPHA, respectively.
RResNet implementation. The backbone architectures of RResNet [118] are similar to those of SPDNet: [20, 16, 8] on Radar, [93, 30] on HDM05, and [63, 33] on FPHA, respectively. The only architectural difference lies in the head: SPDNet directly uses a classification layer, whereas RResNet attaches a residual block before the classification head. Following RResNet, we adopt LogIn + FC + softmax for classification under each metric. We use the official code31 to re-implement the AIM- and LEM-based RResNet. For the RResNets based on LCM and our θ-PCM, we adopt the following settings: a learning rate of 1e−2 , a batch size of 30, and a maximum of 200 epochs. 31
https://github.com/CUAI/Riemannian-Residual-Neural-Networks
272
Appendix A. Experimental Details and Additional Discussions We use the standard cross-entropy loss as the training objective and optimize the parameters with the Riemannian AMSGrad optimizer [17]. The Cholesky diagonal power is set to −1, 0.5, and −0.5 on the Radar, HDM05, and FPHA data sets, respectively. Additionally, on the FPHA data set, a matrix power of −0.25 is applied before the residual blocks to activate the latent geometry for both LCM and our metric. SPD input. The input SPD matrices are the same as those used in SPDNet.
A.4
Additional Discussions
A.4.1
Riemannian Multinomial Logistic Regression
A.4.1.1
RMLR as a Natural Extension of the Euclidean MLR
Proposition 180. When M = Rn is the standard Euclidean space, the RMLR defined in Thm. 114 becomes the Euclidean MLR in Eq. (4.1). Proof. On the standard Euclidean space Rn , Logy x = x − y, ∀x, y ∈ Rn . Besides, the differential maps of left translation and parallel transport are the identity maps. Therefore, given x, pk ∈ Rn and ak ∈ Rn \{0} ∼ = T0 Rn \{0}, we have p(y = k | x ∈ Rn ) ∝ exp ⟨Logpk x, ak ⟩pk , ∝ exp (⟨x − pk , ak ⟩) , ∝ exp (⟨x, ak ⟩ − bk ) ,
(A.39) (A.40) (A.41)
where bk = ⟨pk , ak ⟩. A.4.1.2
Gyro SPSD MLR as a Special Case of Our RMLR
Gyro SPSD MLR [161] is derived from the product of the Grassmannian and SPD gyro spaces. This section will show that the gyro SPSD MLR is a special case of our RMLR on the product geometry of the SPSD manifold. We first review some necessary results about gyro SPSD MLR and then show the equivalence. Following the notations in Nguyen et al. [161], we denote the Grassmannian with f n) and Gr(p, n), canonical metric under the projector and ONB perspective as Gr(p, + respectively. The space of n × n SPSD matrices with a fixed rank p, denoted as Sn,p , forms an SPSD manifold [26]. As shown in Bonnabel et al. [26], Nguyen et al. [161], the p + ∼ SPSD manifold is a product space, i.e., Sn,p = Gr(p, n) × S++ . In other words, every 273
A.4. Additional Discussions p + P ∈ Sn,p can be decomposed as P = UP SP UP⊤ with UP ∈ Gr(p, n) and SP ∈ S++ . We p,g further denote S++ as the SPD manifold with metric g, where g could be AIM, LEM, + or LCM. As shown in Nguyen et al. [161], the gyro space in Sn,p can be defined by the p,g product of gyro spaces of Gr(p, n) and S++ . By this product structure, Nguyen et al. [161] proposed the SPSD Pseudo-gyrodistance to a hyperplane. p,g Definition 181 (SPSD hypergyroplanes [161]). Let P, W ∈ Gr(p, n) × S++ . Then p,g hypergyroplanes in the structure space Gr(p, n) × S++ are defined as
n o psd,g p,g HW,P = Q ∈ Gr(p, n) × S++ | ⟨⊖psd,g P ⊕psd,g Q, W ⟩psd,g = 0 ,
(A.42)
where ⊕psd,g and ⟨·, ·⟩psd,g are the gyro addition and gyro inner product, respectively, as defined in Nguyen et al. [161]. Theorem 182 (SPSD pseudo-gyrodistance [161]). Let W = (UW , SW ) ,
p,g X = (UX , SX ) ∈ Gr(p, n) × S++ , (A.43)
P = (UP , SP ) ,
psd,g p,g and let HW,P be a hypergyroplane in the structure space Gr(p, n) × S++ . Then the psd,g pseudo-gyrodistance from X to HW,P is given by
λ psd,g ¯ d X, HW,P =
D
e gr UP ⊕ e gr UX ⊖
Egr ⊤ e gr UP ⊕ e gr UX ⊤ , UW UW ⊖
+ ⟨⊖g SP ⊕g SX , SW ⟩g q 2 ⊤ gr 2 + (∥SW ∥g ) λ UW UW
,
(A.44)
where ∥ · ∥gr and ∥ · ∥g are the gyro norms on the Grassmann and SPD manifolds e gr and ⊕g are gyro [161], and ⟨·, ·⟩gr and ⟨·, ·⟩g are gyro inner products [161]. ⊕ p,g additions on Gr(p, n) and S++ . Denoting g gr as the canonical metric on Gr(p, n) and g as AIM, LEM, or LCM, we can prove that Thm. 182 is a special case of Thm. 113. Theorem 183. Under the product metric g psd,g = λg gr × g, the Riemannian hyperplane in Eq. (4.19) on the SPSD manifold equals the SPSD hypergyroplane in Thm. 181. Similarly, the Riemannian margin distance in Thm. 113 on the SPSD manifold equals the SPSD Pseudo-gyrodistance in Thm. 182. Proof. Following the notations in Thms. 181 and 182, we further denote P = UP SP UP⊤ , 274
Appendix A. Experimental Details and Additional Discussions ⊤ Q = UQ SQ UQ⊤ , W = UW SW UW , and X = UX SX UX⊤ with UP , UQ , UW , UX ∈ Gr(p, n) p,g and SP , SQ , SW , SX ∈ S++ . Ip is the p × p identity matrix. Ip,n = (Ip , 0)⊤ is the egr , Log g gr , and ⟨·, ·⟩gr denote Riemannian parallel transport gyro identity on Gr(p, n). Γ UP
along a geodesic, the Riemannian logarithm, and the Riemannian metric at UP on Gr(p, n), respectively. Γg , Logg , and ⟨·, ·⟩gSP denote Riemannian parallel transport along p,g a geodesic, the Riemannian logarithm, and the Riemannian metric at SP on S++ , psd,g psd,g psd,g respectively. Γ , Log , and ⟨·, ·⟩X denote Riemannian parallel transport along + ∼ a geodesic, the Riemannian logarithm, and the Riemannian metric at X on Sn,p = p Gr(p, n) × S++ , respectively. First, we show that the SPSD hypergyroplane equals our Riemannian hyperplane in Eq. (4.19). We have the following: ⟨⊖psd,g P ⊕psd,g Q, W ⟩psd,g (1)
= λ⟨⊖gr UP ⊕gr UQ , UW ⟩gr + ⟨⊖g SP ⊕g SQ , SW ⟩g
(2)
gr g g = λ⟨Loggr UP UQ , ÃUW ⟩UP + ⟨LogSP SQ , ÃSW ⟩SP
(A.45)
(3)
= ⟨Logpsd,g Q, Ã⟩psd,g P P
where g gr (UW ) , egr Log ÃUW = Γ Ip,n Ip,n →UP ÃSW = ΓgIp →SP LoggIp (SW ) ,
(A.46) (A.47)
p,g + ∼ Ã = (ÃUW , ÃSW ) ∈ TP Sn,p = TUP Gr(p, n) × TSP S++ ,
(A.48)
where × denotes the Cartesian product. The above derivation comes from the following. (1) The definition of gyro addition, gyro inverse, and gyro inner product on the SPSD manifold [161, Sec. 3.3]. (2) The proof by Nguyen et al. [161, Prop. 3.2] indicates that similar results also hold on the Grassmannian. Combining that proposition with its Grassmannian counterparts yields the equation. (3) The Riemannian product geometry. By the product geometry of the SPSD manifold, we can immediately get psd,g à = (ÃUW , ÃSW ) = Γpsd,g Log (W ) , Ie →P Ie p,n
275
p,n
(A.49)
A.4. Additional Discussions ⊤ where Iep,n = Ip,n Ip,n is the gyro identity on the SPSD manifold.
Next, we show the equivalence between SPSD pseudo-gyrodistance and our Riemannian margin distance: d¯
psd,g X, HW,P
(1) ⟨⊖psd,g P ⊕psd,g X, W ⟩psd,g = ∥W ∥psd,g (2)
=
(3)
=
⟨Logpsd,g X, Ã⟩psd,g P P ∥W ∥psd,g
⟨Logpsd,g X, Ã⟩psd,g P P Logpsd,g (W ) Ie p,n
(4)
=
=
(A.50)
Iep,n
⟨Logpsd,g X, Ã⟩psd,g P P Γpsd,g Iep,n →P
(5)
psd,g
psd,g
Logpsd,g (W ) Iep,n
⟨Logpsd,g X, Ã⟩psd,g P P
P
psd,g
à P (6)
= d(X, H̃Ã,P )
(1) The definition of gyro addition, gyro inverse, gyro inner product, and gyro norm on the SPSD manifold. (2) Eq. (A.45). (3) The definition of SPSD gyro norm [161]. (4) Riemannian parallel transport is norm-preserving [69, Def. 3.1]. (5) Eq. (A.49). (6) Thm. 113.
Remark 184. We make the following remarks regarding the gyro MLR and our MLR on the SPSD manifold. (1) Eq. (A.49) indicates that when generating à in our RMLR by parallel trans+ porting a tangent vector A ∈ TIep,n Sn,p , à is the initial velocity of W in Eq. (A.44). 276
Appendix A. Experimental Details and Additional Discussions (2) Putting the pseudo-gyrodistance and Riemannian margin distance into Eq. (4.18) yields the gyro MLR and our Riemannian MLR. Therefore, Thm. 183 indicates the equivalence of the gyro MLR with our RMLR on the SPSD manifold. (3) As a metric g is required to induce gyro-structures, the metric g in gyro SPSD MLR is confined to AIM, LEM, and LCM. However, our SPSD MLR can be defined on the product space of the Grassmannian and SPD manifold under other metrics, such as BWM and PEM, as our framework does not require gyro-structures. A.4.1.3
Theories on the Deformed Metrics
Limiting cases of the deformed metrics. Thanwerdas and Pennec [192] generalized (α, β)-AIM to a three-parameter family of metrics by power deformation, i.e., (θ, α, β)-AIM. The family of (θ, α, β)-AIM comprises (α, β)-AIM for θ = 1 and approaches (α, β)-LEM with θ → 0 [192]. As reviewed in Sec. 3.2.5.1, LCM and (α, β)-LEM are extended to power-deformed metrics, denoted as (θ, α, β)-LEM and θ-LCM. Thm. 67 shows that (θ, α, β)-LEM is equal to (α, β)-LEM, and θ-LCM interpolates between g̃-LEM (θ → 0) and LCM (θ = 1), with g̃-LEM defined as 1 1 n , ⟨V1 , V2 ⟩P = ⟨Ve1 , Ve2 ⟩ − ⟨D(Ve1 ), D(Ve2 )⟩, ∀Vi ∈ TP S++ 2 4
(A.51)
where Vei = log∗,P (Vi ) with log∗,P as the differential map of matrix logarithm, and D(Vei ) is a diagonal matrix consisting of the diagonal elements of Vei .
Thanwerdas and Pennec [194] identified the Alpha-Procrustes metric [151] with power-deformed BWM, denoted as 2θ-BWM. Similarly, 2θ-BWM becomes BWM with θ = 0.5 [194]. We further show the limiting case of 2θ-BWM under θ → 0. Proposition 185. 2θ-BWM tends to ( 14 , 0)-LEM as θ → 0. Before starting the proof, we first recall a well-known property of deformed metrics [194]. Lemma 186. Let θ12 ϕ∗θ g be the deformed metric on SPD manifolds pulled back from n g by the matrix power ϕθ and scaled by θ12 . Then, as θ tends to 0, for all P ∈ S++ 277
A.4. Additional Discussions n , we have and all V ∈ TP S++
1 ( 2 ϕ∗θ g)P (V, V ) → gI (log∗,P (V ), log∗,P (V )). θ
(A.52)
Now, we present our proof for the limiting cases of deformed metrics. Proof of Thm. 185. First, we have 1 gIBWM (V, V ) = ⟨V, V ⟩. 4
(A.53)
By Thm. 186, we have the following: θ→0 gP2θ-BWM (V, V ) −−→ gIBWM log∗,P (V ), log∗,P (V ) 1 = ⟨log∗,P (V ), log∗,P (V )⟩ 4 ( 1 ,0)-LEM
= gP4
(A.54)
(V, V ) .
Proof of the properties of the deformed metrics (Tab. 4.2). In this subsection, we prove the properties presented in Tab. 4.2. We first present a useful lemma and then present our detailed proof. This lemma will be useful in the proof of our SPD MLRs as well. Lemma 187. Given a Riemannian manifold (M, g) and a positive real scalar a > 0, the scaled metric ag on M has the same Riemannian logarithmic maps, exponential maps, and parallel transport as g. Proof. Since the Christoffel symbols of ag are identical to those of g, the geodesics and parallel transport under both ag and g remain unchanged. The equivalence of geodesics implies that the Riemannian exponential maps are the same for ag and g. As the inverse of the Riemannian exponential maps, the Riemannian logarithmic maps under ag and g are also identical. According to Thm. 187, geodesic completeness is independent of the scaling factor a > 0. By the definition of O(n)-, left-, right-, and bi-invariance, these invariant properties are also independent of the scaling factor a > 0. Without loss of generality, we will omit the scaling factor in the following proof. 278
Appendix A. Experimental Details and Additional Discussions Proof. First, we prove O(n)-invariance of (θ, α, β)-LEM, (θ, α, β)-EM, (θ, α, β)-AIM, and 2θ-BWM. Since the differential of ϕθ is O(n)-equivariant, and (α, β)-LEM, (α, β)EM, (α, β)-AIM, and BWM are O(n)-invariant [196], O(n)-invariance is thus acquired. Next, we focus on geodesic completeness. It can be easily proven that Riemannian isometries preserve geodesic completeness. On the other hand, (α, β)-LEM, (α, β)AIM, and LCM are geodesically complete [196, 137]. As a direct corollary, geodesic completeness can be obtained since ϕθ is a Riemannian isometry. Finally, we deal with Lie group invariance. Similarly, it can be readily proved that Lie group invariance is preserved under isometries. LCM, LEM, and (α, β)-AIM are Lie group bi-invariant [137], bi-invariant [9], and left-invariant [195]. As an isometric pullback metric from the standard LEM [196], (α, β)-LEM is, therefore, Lie group biinvariant. As pullback metrics, (θ, α, β)-LEM, (θ, α, β)-AIM, and θ-LCM are therefore bi-invariant, left-invariant, and bi-invariant, respectively. A.4.1.4
Computational Details on the SPD MLR under Power-Deformed BWM
Matrix square roots in the SPD MLR under power-deformed BWM. In the 1 1 case of MLRs induced by 2θ-BWM, computing square roots like (BA) 2 and (AB) 2 n with B, A ∈ S++ poses a challenge. Eigendecomposition cannot be directly applied since BA and AB are no longer symmetric, let alone positive definite. Instead, we use the following formulas to compute these square roots [151]: 1
1
1
1
1
1
1
1
(BA) 2 = B 2 (B 2 AB 2 ) 2 B − 2 and (AB) 2 = [(BA) 2 ]⊤ ,
(A.55)
where the involved square roots can be computed using eigendecomposition or singular value decomposition (SVD). Numerical stability of the SPD MLR under power-deformed BWM. Let us first explain why we abandon parallel transport for the SPD MLR derived from 2θBWM. Then, we propose our numerically stable methods for computing the SPD MLR based on 2θ-BWM. Instability of parallel transport under power-deformed BWM. As discussed in Thm. 114, there are two ways to generate à in SPD MLR: parallel transport and Lie group translation. However, parallel transport under 2θ-BWM can cause numerical problems. Without loss of generality, we focus on the standard BWM because 2θ-BWM is isometric to the BWM. Although general parallel transport under BWM is obtained by solving an ODE, 279
A.4. Additional Discussions transport starting from the identity matrix has a closed-form expression [196]: ΓI→P (V ) = U
"r
# σi + σ j ⊤ U V U ij U ⊤ , 2
(A.56)
n where P = U ΣU ⊤ is the eigendecomposition of P ∈ S++ . The forward computation of Eq. (A.56) is numerically stable. However, its backpropagation requires differentiating through an eigendecomposition, involving the calculation of 1/(σi −σj ) [110, Prop. 2]. When σi is close to σj , this backpropagation can be problematic.
Numerically stable methods for the SPD MLR under power-deformed BWM. To bypass the instability of parallel transport under BWM, we use Lie group left translation to generate à in MLRs induced by 2θ-BWM. However, there is another problem that could cause instability. The computation of the Riemannian metric of 2θ-BWM requires solving the Lyapunov operator, i.e., LP [V ]P + P LP [V ] = V . For symmetric matrices, the Lyapunov operator can be obtained by eigendecomposition:
Vij′ LP [V ] = U σi + σj
U ⊤,
(A.57)
i,j
n . where V ∈ S n , U V ′ U ⊤ = V , and P = U ΣU ⊤ is the eigendecomposition of P ∈ S++ 1 Similar to Eq. (A.56), the backpropagation of Eq. (A.57) requires /(σi −σj ), undermining numerical stability.
To remedy this problem, we propose the following formula to stably compute the backpropagation of Eq. (A.57). n Proposition 188. For all P ∈ S++ and all V ∈ S n , we denote the Lyapunov equation as XP + P X = V, (A.58) ∂L where X = LP [V ]. Given the gradient ∂X of loss L with respect to X, the backpropagation of the Lyapunov operator can be computed by
∂L ∂L = LP [ ], ∂V ∂X ∂L ∂L ∂L = −XLP [ ] − LP [ ]X, ∂P ∂X ∂X where LP [·] can be computed by Eq. (A.57). 280
(A.59) (A.60)
Appendix A. Experimental Details and Additional Discussions Proof. Differentiating both sides of Eq. (A.58), we obtain d XP + X d P + d P X + P d X = d V,
(A.61)
=⇒ d XP + P d X = d V − X d P − d P X,
(A.62)
=⇒ d X = LP [d V − X d P − d P X].
(A.63)
Besides, easy computations show that LP [V ] : W = V : LP [W ], ∀W, V ∈ S n ,
(A.64)
where · : · denotes the standard Frobenius inner product. Then we have the following: ∂L ∂L : dX = : LP [d V − X d P − d P X], ∂X ∂X ∂L ∂L ∂L ∂L : d X = LP [ ] : d V + −XLP [ ] − LP [ ]X : d P. =⇒ ∂X ∂X ∂X ∂X
(A.65) (A.66)
Remark 189. Eq. (A.57) needs to be computed in the Lyapunov operator’s forward and backward processes. Therefore, in the process, we can save the interh forward i 1 mediate matrices U and K with Ki,j = σi +σj , and then use them to compute i,j
the backward process efficiently.
281
A.4. Additional Discussions
Method
Point-to-hyperplane distance
Euclidean MLR Poincaré MLR [76]
Pseudo-Busemann MLR [162]
Lorentz MLR [15]
√
1 sinh−1 −K
| ⟨a, x⟩ + b| ∥a∥ ! √ 2 −K |⟨−p ⊕M x, a⟩| 1 + K ∥−p ⊕M x∥2 ∥a∥ [76, Thm. 5]
B v (−p ⊕M x) ∥−p ⊕M x∥ [162, Cor. 4.3] √ ⟨v, x⟩L 1 −1 √ −K sinh ∥v∥L −K [15, Eq. (44)] d(x, p)
|−αB v (x) + b| α
BMLR
Applied manifolds
Dist
Rn
Real
PnK
Real
PnK
Pseudo
LnK
Real
PnK , LnK
Real
Table A.17: Comparison of point-to-hyperplane distances. Real means the point-tohyperplane distance is the real distance, obtained by inf y∈H d(x, y) with H as a hyperplane and d as the geodesic distance. Instead, Pseudo means the point-to-hyperplane distance is a surrogate, which only equals the real distance in Euclidean geometry.
A.4.2
Hyperbolic Busemann Neural Networks
A.4.2.1
Comparison with Existing Hyperbolic MLRs
Method
Hyperplane
Formulation
Applied manifolds
Compact params
Euclidean MLR
Euclidean
{x ∈ Rn | ⟨a, x⟩ + b = 0} a ∈ Rn , b ∈ R
Rn
✓
PnK
✗
o n x ∈ PnK | Logp (x) , a p = 0
Poincaré MLR [76]
Geodesic [76, Def. 3.1]
Pseudo-Busemann MLR [162]
Busemann & gyro [162, Def. 4.1]
{x ∈ PnK | B v (−p ⊕M x) = 0} v ∈ Sn−1 , p ∈ PnK
PnK
✗
Lorentz MLR [15]
Ambient Minkowski [15, Eq. (7)]
{x ∈ LnK | ⟨w, x⟩L = 0} p ∈ LnK , w ∈ Tp LnK
LnK
✗
BMLR
Horosphere
PnK , LnK
✓
p ∈ PnK , a ∈ Tp PnK
n {x ∈ HK | −αB v (x) + b = 0} α > 0, v ∈ Sn−1 , b ∈ R
Table A.16: Comparison of hyperplanes. Compact params indicate whether the parameterization requires an additional manifold-valued point. All the considered hyperbolic MLRs follow a point-to-hyperplane formulation. Therefore, the key difference lies in hyperplanes and point-to-hyperplane distances across methods. In addition to Tab. 5.11, Tabs. A.16 and A.17 further make this comparison. 282
Appendix A. Experimental Details and Additional Discussions We draw the following three conclusions. (1) Hyperplanes. Our BMLR uses Busemann-based horospheres that simultaneously satisfy three desiderata: (i) compact parameterization without a per-class manifoldvalued point, whereas other hyperbolic ones32 are over-parameterized; (ii) a natural generalization of Euclidean hyperplanes via horospheres, while the Lorentz MLR relies on the ambient Minkowski space, failing to fully respect the intrinsic geometry; and (iii) applicability across hyperbolic models, whereas the Lorentz MLR is tailored to the Lorentz model. (2) Point-to-Hyperplane Distances. Although pseudo-Busemann MLR also exploits Busemann functions, it relies on a pseudo point-to-hyperplane distance that only coincides with the real distance in Euclidean geometry. In contrast, our BMLR calculates the real point-to-horosphere distance, ensuring geometric fidelity across hyperbolic models. (3) Batch Efficiency. Recalling Tab. 5.11, the Poincaré MLR [76] computes logits using ⟨−pk ⊕M x, ak ⟩ and ∥−pk ⊕M x∥2 . For a batch X ∈ Rbs×n and C classes, evaluating −pk ⊕M X for every k yields an intermediate tensor of shape [bs, C, n], and materializing this tensor can cause GPU out of memory (OOM) when n or C is large. The same limitation holds for the pseudo-Busemann MLR. Consequently, their official implementations compute per class in a for-loop, which is batch inefficient. By contrast, BMLR uses logits −αk B vk (x) + bk , whose Busemann term reduces to class-wise inner products ⟨vk , x⟩ (or ⟨vk , xs ⟩). With X ∈ Rbs×n and V = [v1 , . . . , vC ] ∈ Rn×C , such inner products can be efficiently implemented as a single matrix multiplication XV without any [bs, C, n] intermediate, yielding high throughput and low memory usage. A.4.2.2
Busemann Fully Connected Layers and Point-to-Horosphere Distances
An apparently natural attempt to define a hyperbolic FC layer is to replace the LHS of Eq. (5.59) by the signed point-to-horosphere distance. The Euclidean hyperplane passing through the origin and orthogonal to ek ∈ Rm is Hek ,0 = {y ∈ Rm | ⟨ek , y⟩ = 0}. m The corresponding hyperbolic horosphere, following Eq. (5.55), is Hek ,1,0 = {y ∈ HK | ek −B (y) = 0}. By Eq. (5.57), the signed point-to-horosphere distance to Hek ,1,0 equals 32
We note that Shimizu et al. [181], Bdeir et al. [15] mitigate this issue through re-parameterization: z n a = PTe→p (z) and p = Expe b ∥z∥ , where e ∈ HK denotes the origin, z ∈ Rn , and b ∈ R. Nevertheless, the underlying definitions are over-parameterized.
283
A.4. Additional Discussions n m −B ek (y). Accordingly, we can define an alternative FC mapping F : HK ∋ x 7→ y ∈ HK via (A.67) B ek (y) = αk B vk (x) − bk , k = 1, . . . , m,
where uk (x) = −αk B vk (x) + bk , and αk > 0, vk ∈ Sn−1 , bk ∈ R are learnable. Although this only differs from Eq. (5.59) on the LHS, the following discussion shows that such a definition is infeasible in general and fails to deliver a valid hyperbolic FC layer.
Poincaré Model. Using Eq. (5.40) with v = ek , Eq. (A.67) becomes 1 √ log −K
ek −
√
−Ky
1 + K ∥y∥2
2
!
= −uk (x).
√ Define tk = exp − −Kuk (x) > 0. Exponentiating Eq. (A.68) gives √ tk 1 + K ∥y∥2 = 1 − 2 −Kyk − K ∥y∥2 ,
k = 1, . . . , m.
(A.68)
(A.69)
Writing R = ∥y∥2 , Eq. (A.69) yields an affine expression for each coordinate yk = ck + dk R, Imposing R =
Pm
2 k=1 yk =
where, denoting T = A2 =
Pm
1 − tk ck = √ , 2 −K
k=1 (ck + dk R)
2
dk =
√
−K (1 + tk ) . 2
gives a quadratic in R:
A2 R2 + (A1 − 1) R + A0 = 0,
Pm
k=1 tk and q =
−K (m + 2T + q) , 4
Pm
A1 =
(A.70)
(A.71)
2 k=1 tk ,
m−q , 2 284
A0 =
m − 2T + q ≥ 0. 4(−K)
(A.72)
Appendix A. Experimental Details and Additional Discussions Its discriminant is ∆P = (A1 − 1)2 − 4A2 A0 2 m−q −K m − 2T + q −1 −4 (m + 2T + q) = 2 4 4(−K) 2 m−2−q (m + 2T + q) (m − 2T + q) = − 2 4 1 = (m − 2 − q)2 − (m + q)2 − (2T )2 4 1 ((m − 2 − q) − (m + q)) ((m − 2 − q) + (m + q)) + 4T 2 = 4 1 = (−2 − 2q)(2m − 2) + 4T 2 4 = T 2 − (m − 1) (1 + q) .
(A.73)
A real solution R exists only if ∆P ≥ 0. In addition, feasibility requires 0 ≤ R < −1/K m so that y ∈ Pm K . Such conditions can fail for generic {uk (x)}, in which case no y ∈ PK satisfies Eq. (A.67). Lorentz Model. Using Eq. (5.41) with v = ek , Eq. (A.67) becomes √
√ 1 −K (yt − (ys )k ) = −uk (x). log −K
(A.74)
√ P Define tk = exp − −Kuk (x) > 0 and t = (t1 , . . . , tm )⊤ . Denoting T = m k=1 tk and Pm 2 q = k=1 tk , we obtain from Eq. (A.74), tk (ys )k = yt − √ , −K
k = 1, . . . , m,
=⇒
ys = yt 1 − √
1 t. −K
(A.75)
Enforcing the hyperboloid constraint ∥ys ∥2 − yt2 = 1/K yields a quadratic in yt : √ (m − 1)Kyt2 + 2 −KT yt − (1 + q) = 0.
(A.76)
The discriminant is √ 2 ∆L = 2 −KT − 4(m − 1)K (− (1 + q)) = 4(−K)T 2 + 4(m − 1)K (1 + q) = 4(−K) T 2 − (m − 1) (1 + q) .
(A.77)
Hence, a real yt exists only if T 2 − (m − 1) (1 + q) ≥ 0. In addition, feasibility re285
A.4. Additional Discussions quires yt > 0. Such conditions can fail for generic {uk (x)}, where no y ∈ Lm K satisfies Eq. (A.67). Summary. Equating Busemann coordinates as in Eq. (A.67) requires nontrivial inequalities on the responses {uk (x)}. These constraints are not guaranteed during learning, so the system can become infeasible and the output y undefined. This motivates our choice in Eq. (5.59) to use signed point-to-hyperplane distances on the LHS, which admit closed-form solutions that are feasible for all inputs and parameters in both the Poincaré and Lorentz models.
A.4.3
Full-Rank Correlation Networks
A.4.3.1
Connections among FC Layers: Correlation, SPD, Poincaré, and Euclidean
We clarify the correspondence between our FC formulation in Eq. (5.69) and previous FC layers. SPD Manifold. Nguyen et al. [161, Props. 3.4–3.6] introduced three SPD FC layers based on the gyrovector spaces under LEM, LCM, and AIM, respectively. These gyro SPD FC layers share the same definition as Eq. (5.69), except that their signed distance and vk are defined by gyrovector spaces. Poincaré Ball. We show that the Poincaré FC layer F(·) : PnK → Pm K reviewed in Sec. A.2.4 is also defined as our correlation FC layer in Thm. 141. We define the zero vector 0 ∈ PnK as the Poincaré origin, as it is the identity element of the Poincaré gyrovector space [76]. Obviously, {ek }m k=1 is the orthogonal basis over m m T0 PK , where ek = (δik )i=1 . Corresponding to Eq. (5.69), we have sign(⟨Log0 (y), ek ⟩) d(y, Hek ,0 ) = vk (x).
(A.78)
Compared with Shimizu et al. [181, Eq. (56)], we only need to show the LHS. The sign can be calculated as √ y (1) −1 sign(⟨Log0 (y), ek ⟩0 ) = sign 4 tanh −K∥y∥ √ , ek −K∥y∥ (A.79) (2) = sign (⟨y, ek ⟩) = sign (yk ) . The above follows from the following. 286
Appendix A. Experimental Details and Additional Discussions −1 2 (1) λK 0 = 1+K∥0∥2 = 2 and Log0 (y) = tanh
(2) tanh−1 (a) > 0, ⇐⇒ a > 0.
√
y −K∥y∥ √−K∥y∥ .
Therefore, the LHS of Eq. (A.78) is simplified as sign(⟨Log0 (y), ek ⟩) d(y, Hek ,0 ) = sign (yk ) d(y, Hek ,0 )
√ 2 −K|yk | 1 −1 = sign (yk ) √ sinh 1 + K∥y∥2 −K √ 1 2 −Kyk (2) −1 = √ . sinh 1 + K∥y∥2 −K (1)
(A.80)
The above comes from the following. (1) Ganea et al. [76, Thm. 5]. (2) 1 + K∥y∥2 > 0 by definition and sign(a) sinh−1 (|a|) = sinh−1 (a), ∀a ∈ R. The last equation in Eq. (A.80) is the LHS of Shimizu et al. [181, Eq. (56)], indicating the equality. Euclidean Space. We show that the Euclidean FC layer F(·) : Rn ∋ x → y = Ax + b ∈ Rm can also be defined as our correlation FC layer in Thm. 141. In Euclidean space Rn , the zero vector 0 ∈ Rn is the origin, and {ek }m k=1 is the orthogonal basis over T0 Rm ∼ = Rm . Then, the RHS of Eq. (5.69) becomes vk (x) = ⟨x − pk , ak ⟩ (1)
= ⟨x, zk ⟩ − γk ∥zk ∥ .
(A.81)
where (1) comes from Exp0 (γk [zk ]) = γk [zk ] and PT0→pk (zk ) = zk . The above takes the form of ⟨x, ak ⟩ + bk . On the other hand, the LHS of Eq. (5.69) becomes (1)
sign(⟨Log0 (y), ek ⟩) d(y, Hek ,0 ) = sign(yk ) d(y, Hek ,0 ) (2)
= yk .
The above comes from the following. (1) Log0 (y) = y and ⟨y, ek ⟩ = yk . k ⟩| (2) d(y, Hek ,0 ) = |⟨y,e = |yk |. ∥ek ∥
287
(A.82)
A.4. Additional Discussions A.4.3.2
Log-Euclidean Layers under Product Geometry
We first review some basic facts of the product geometry, and then discuss the LogEuclidean correlation MLR and FC layer under the product geometry. Product of Correlation Manifolds. Given a manifold (M, g), the n-fold product Q is (Mn , g) = ni=1 (M, g). Each point and tangent vector over Mn are (A.83)
Mn ∋ P = (P1 ∈ M, · · · , Pn ∈ M),
TP Mn ∋ V = (V1 ∈ TP1 M, · · · , Vn ∈ TPn M).
(A.84)
The product metric is ⟨V, W ⟩P =
n X i=1
⟨Vi , Wi ⟩Pi ,
∀V, W ∈ TP Mn .
(A.85)
Correlation MLR. Following Thm. 138, the logit for the k-th class and the input X = {Xl ∈ Cor+ (n)}cl=1 ∈ (Cor+ (n))c is vk (X) = =
c X
l=1 c X l=1
vkl (Xl ; Zkl , γkl ) (A.86) (⟨ϕn (Xl ), (ϕn )∗,In (Zkl )⟩ − γkl ∥(ϕn )∗,In (Zkl )∥) ,
where dn = n(n − 1)/2, ϕn : Cor+ (n) → Rdn is the diffeomorphism of the selected Log-Euclidean geometry, I = (In , · · · , In ), Zk = (Zk1 , · · · , Zkc ) ∈ TI (Cor+ (n))c , Zkl ∈ TIn Cor+ (n), and γkl ∈ R. Each vkl is the response of the l-th component given by Thm. 138, with the unit direction [Zkl ] = Zkl / ∥Zkl ∥In .
Correlation FC Layer. Following Thm. 201 and Eq. (A.86), the FC layer F(·) : (Cor+ (n))c → Cor+ (m) for the input X is Y = ϕ−1 m
dm X c X
vil (Xl ; Zil , γil )ei
i=1 l=1
!
,
(A.87)
where dm = m(m − 1)/2, ϕm : Cor+ (m) → Rdm is the output diffeomorphism, Zil ∈ TIn Cor+ (n) ∼ = Hol(n), and γil ∈ R.
Eq. (A.87) implies that F(·) : (Cor+ (n))c → Cor+ (m) differs from F(·) : Cor+ (n) → Cor+ (m) only in each scalar response vi , where the former is a summation over the c input components. For example, considering the FC layer F(·) : (Cor+ (n))c → Cor+ (m) 288
Appendix A. Experimental Details and Additional Discussions under ECM, its matrix-coordinate response vij for the input C = {Cl ∈ Cor+ (n)}cl=1 is vij (C) =
c X
EC (Cl ; Zijl , γijl ), vijl
(A.88)
l=1
where Zijl ∈ Hol(n) and γijl ∈ R, for i, j = 1, · · · , m with i > j, and 1 ≤ l ≤ c.
A.4.4
Adaptive Log-Euclidean Metrics
A.4.4.1
Well-definedness of General Matrix Logarithm
Eq. (6.12) does not explicitly specify the correspondence between eigenvalues and diagonal logarithms. Here, we present a detailed explanation. In implementations such as PyTorch or MATLAB, eigendecomposition routines return eigenpairs in a prescribed order. P We rewrite the eigendecomposition as S = σi Ei , where Ei = ui u⊤ i and ui is the corresponding eigenvector in U . Let S be an n × n SPD matrix, and let Pn be the permutation group on {1, . . . , n}. Changing the order of {1, . . . , n} can be viewed as a permutation, so we use π ∈ Pn to represent the corresponding changed order. Assume that the eigenvalues σi are sorted in ascending order, i.e., σ1 ≤ · · · ≤ σn . We use “the i-th eigenvalue” to refer to the eigenvalue in the i-th pair of the ordered eigenpair sequence (σ1 , u1 ), . . . , (σn , un ). Since each eigenvector ui is unique, it is safe to say the eigenvalues are ordered, and the i-th eigenvalue/eigenvector pair is unique. Let logα (Σ) denote applying the scalar logarithm logai to the i-th eigenvalue σi . P P Then logα (S) is rewritten as logα (S) = logai (σi )Ei , where S = σi Ei . In this way, logα is clearly well-defined. By definition, we can observe that the output of logα does not depend on the order in eigendecomposition. Suppose there are two eigendecompositions with different orders, i.e., S = U ΣU ⊤ = Ũ Σ̃Ũ ⊤ , where Ũ and Σ̃ are rearrangements of U and Σ. There exists a π ∈ Pn such that, for each j, there is a unique i satisfying ũj = uπ(i) and σ̃j = σ(i) . We then have P P logai (σi )Ei for S = U ΣU ⊤ and logaπ(i) (σπ(i) )Eπ(i) for S = Ũ Σ̃Ũ ⊤ , which indicates that the two eigendecompositions are equivalent. 289
A.4. Additional Discussions
A.4.5
Product Cholesky Metrics
A.4.5.1
Riemannian Operators under (θ, M)-DBWM
We first define a map ϕθ : Ln++ → Ln++ as ϕθ (L) = ⌊L⌋ + Lθ , ∀L ∈ Ln++ .
(A.89)
Its differential at L ∈ Ln++ is given by (ϕθ )∗,L (X) = ⌊X⌋ + θLθ−1 X,
∀X ∈ TL Ln++ .
(A.90)
Let g M-DBW and g (θ,M)-DBW be M-DBWM and (θ, M)-DBWM, respectively. Since constant scaling of a Riemannian metric preserves the Christoffel symbols, Riemannian operators such as the Riemannian logarithm, exponential map, and parallel transport under g (θ,M)-DBW are the same as those under the pullback metric ϕ∗θ g M-DBW . Following Thm. 33, these Riemannian operators under (θ, M)-DBWM can be obtained by ϕ∗θ g M-DBW . Besides, as constant scaling does not affect the weighted Fréchet mean (WFM), the WFM under (θ, M)-DBWM is the same as that under ϕ∗θ g M-DBW . The latter can be calculated by the properties of isometries presented in Thm. 33. Therefore, by the properties of isometry in Thm. 33 and Eqs. (A.89) and (A.90), we can obtain all the Riemannian operators.
290
Appendix B Proofs B.1
Mathematical Background
B.1.1
Proof of Thm. 57
Proof. The gyration-preservation property is shown by Ungar [200, Eq. (6.323)]. The proofs for the gyro identity and gyroinverse follow directly from the isomorphism and the uniqueness of inverse and identity [200, Thm. 2.10].
B.2
Lie Group Batch Normalization
B.2.1
Proof of Thm. 61
Proof. Case (1). The MLE of M is MMLE = argmax log(k(v)) − = argmin
N X
N X d(Pi , M )2 i=1
2v 2
(B.1)
d(Pi , M )2 .
i=1
Case (2). We denote Y = LB (X), and pX and pY as the density of X and Y , 291
B.2. Lie Group Batch Normalization respectively. The density of Y is (1)
pY (Q) = pX (L⊖B (Q)) d(L⊖B (Q), M )2 = k(σ) exp − 2σ 2 d(Q, LB (M ))2 (2) = k(σ) exp − . 2σ 2
(B.2)
The above comes from: (1) Pennec [170, Thm. 7]; (2) The isometry of the left translation.
B.2.2
Proof of Thm. 62
Proof. The isometry of LB directly implies the homogeneity of the sample mean. Now let us focus on Eq. (3.14). We have the following: XN
i=1
wi d2 (ϕs (Pi ), E) =
XN
= s2 = s2
where ∥ · ∥E is the norm on TE M.
B.2.3
wi ∥s LogE Pi ∥2E
i=1 XN
i=1 XN i=1
wi ∥ LogE Pi ∥2E
(B.3)
wi d2 (Pi , E),
Proof of Thm. 65
Proof. As the right-invariant metric has properties analogous to those of the leftinvariant metric, this proof follows the same logic as the two proofs above. Gaussian homogeneity. We denote Y = RB (X), and pX and pY as the density of X and Y , respectively. The density of Y is (1)
pY (Q) = pX (R⊖B (Q)) d(R⊖B (Q), M )2 = k(σ) exp − 2σ 2 d(Q, RB (M ))2 (2) = k(σ) exp − . 2σ 2 292
(B.4)
Appendix B. Proofs The above comes from: (1) Pennec [170, Thm. 13]; (2) The isometry of the right translation. Sample mean homogeneity. This is a direct corollary of the isometry of right translation.
B.2.4
Proof of Thm. 66
Proof. As Rn is an abelian group and the Euclidean inner product is bi-invariant, we focus on left translation in the following. The core of this proof lies in the fact that on Rn , (1) the Fréchet mean and variance are reduced to the familiar Euclidean statistics. (2) the calculation of the running mean becomes the weighted arithmetic mean. (3) Eqs. (3.8) to (3.10) become Eq. (3.1). We prove these three points one by one. As stated by Lou et al. [142, Prop. G.1 and Cor. G.2], from the view of the product manifold, the element-wise Fréchet mean and variance on Rn are equivalent to the vector-valued Euclidean mean and variance. Besides, by an argument similar to that of Lou et al. [142, Prop. G.1], the weighted Fréchet mean on Rn is simplified to the weighted arithmetic average. Therefore, on Rn , the calculation of running statistics in our Alg. 1 becomes the familiar moving average. Thirdly, on Rn , we know that Lx (y) = x + y, Expx v = x + v, Logx y = y − x, and the neutral element is 0. Since statistics, as well as the Euclidean BN, are calculated element-wise, it suffices to consider a single coordinate, i.e., R. For a batch of activations 2 {xi }N i=1 ⊂ R with batch mean µb and batch variance vb , Eqs. (3.8) to (3.10) are rewritten as follows: " #! γ xi − µ b Log0 (L−µb (xi )) = γp 2 + β. (B.5) Lβ Exp0 p 2 vb + ϵ vb + ϵ The above equation is the exact core computation of the standard Euclidean BN.
B.2.5
Proof of Thm. 67
Proof. We first prove the case of (θ, α, β)-LEM, and then proceed to the case of θ-LCM. (θ, α, β)-LEM. For clarity, we denote the metric tensor of (θ, α, β)-LEM as g (θ,α,β)-LE =
1 ∗ (α,β)-LE P g , θ2 θ
(B.6)
n n where g (α,β)-LE is the metric tensor of (α, β)-LEM. Let P ∈ S++ and V, W ∈ TP S++ .
293
B.2. Lie Group Batch Normalization Then we have (θ,α,β)-LE
gP
1 (α,β)-LE g ((Pθ )∗,P (V ), (Pθ )∗,P (W )) θ2 Pθ (P ) 1 = 2 ⟨(log ◦ Pθ )∗,P (V ), (log ◦ Pθ )∗,P (W )⟩(α,β) θ = ⟨log∗,P (V ), log∗,P (W )⟩(α,β)
(V, W ) =
(α,β)-LE
= gP
(B.7)
(V, W ).
θ-LCM. Let us first review a well-known fact of deformed metrics [194]. Let g̃ = θ12 P∗θ g be the power-deformed metric on SPD. Then when θ tends to 0, for all n n P ∈ S++ and all V ∈ TP S++ , we have g̃P (V, V ) → gI (log∗,P (V ), log∗,P (V )).
(B.8)
By Eq. (B.8), we can readily obtain the results.
B.2.6
Proof of Thm. 68
Proof. (α, β)-AIM is left-invariant [195]. As the pullback of (α, β)-AIM, (θ, α, β)-AIM is left-invariant as well. Besides, Thm. 147 shows that LCM is the pullback metric from the Euclidean space of LTn . Therefore, θ-LCM is bi-invariant.
B.2.7
Proof of Thm. 69
Proof. In the following, we denote the Riemannian operators under the left-invariant metric, i.e., AIM, by dL , LogL , and ExpL . Note that the Cholesky decomposition pulls back the group operation of matrix product from the Cholesky manifold Ln++ [195]. For simplicity, we abbreviate ⊕LieAI as ⊕.
The differential maps of the Cholesky decomposition and its inverse are reviewed in Eq. (2.92). Following the notation in this theorem, we further denote X ∈ TL Ln++ . Specifically, for the differential map at I, we have n . Chol∗,I (V ) = (V ) 1 , ∀V ∈ TI S++
(B.9)
(Chol−1 )∗,I (X) = (X)sym+ , ∀X ∈ TI Ln++ .
(B.10)
2
e and R e as the group translation on the Cholesky manifold Ln , we have the Denoting L ++ 294
Appendix B. Proofs following w.r.t. the differential maps of left and right translation: −1 eL )∗,L−1 ◦ Chol∗,⊖P Chol−1 ∗,I (L −1 eL )∗,L−1 = (Chol∗,⊖P )−1 (L ◦ Chol∗,I , (2) e L−1 )∗,L ◦ Chol∗,P . (R⊖P )∗,P = Chol−1 ∗,I ◦ (R (1)
((LP )∗,⊖P )−1 =
(B.11) (B.12) (B.13)
The above derivation comes from the following: eL ◦ Chol; (1) LP = Chol−1 ◦L
e L−1 ◦ Chol. (2) R⊖P = Chol−1 ◦R
Riemannian metric. For the differential of right translation, we have the following: e L−1 )∗,L ◦ Chol∗,P (V ) (R⊖P )∗,P (V ) = Chol−1 ∗,I ◦ (R −1 −1 −⊤ . = L(L V L ) 1 L 2
(B.14)
sym+
By Eq. (B.14), one can obtain the expression for the Riemannian metric tensor. Riemannian geodesic and exponential map. According to Zacur et al. [228], we have the following for the operators between left- and right-invariant metrics: ExpP (V ) = ⊖ ExpL⊖P − ((LP )∗,⊖P )−1 ◦ (R⊖P )∗,P (V ) , d(P, Q) = dL (⊖P, ⊖Q) .
(B.15) (B.16)
Putting the AIM-based geodesic distance into the RHS of Eq. (B.16), one can obtain the geodesic distance under CRIM. Now, we simplify Eq. (B.15). Putting Eqs. (B.12) and (B.13) into Eq. (B.15), we have the following: ExpP (V ) = ⊖
ExpL⊖P
−1 −1 e e − (Chol∗,⊖P ) ◦ (LL )∗,L−1 ◦ (RL−1 )∗,L ◦ Chol∗,P (V )
= ⊖ ExpL⊖P − (Chol∗,⊖P )−1 L−1 Chol∗,P (V )L−1 h io n (1) = ⊖ ExpL⊖P − (Chol∗,⊖P )−1 L−1 V L−⊤ 1 L−1 2 (2) L −1 −⊤ −1 −⊤ . = ⊖ Exp⊖P − L V L L 1 L 2
The above comes from the following: 295
sym+
(B.17)
B.2. Lie Group Batch Normalization (1) The first identity follows from Eq. (2.92); (2) The second identity follows from Eq. (2.92). Riemannian logarithm. From the second equality in Eq. (B.17), we have the following: L LogP (Q) = − Chol−1 ∗,P L Chol∗,⊖P Log⊖P (⊖Q) L (1) −1 ⊤ e LV L 1 L = − Chol∗,P (B.18) 2 ⊤ = − LL⊤ LVe L⊤ 1 2
sym+
The above comes from the following: n . (1) Chol∗,⊖P (V ) = L−1 LV L⊤ 1 , ∀V ∈ T⊖P S++ 2
B.2.8
Proof of Thm. 70
n . Proof. Completeness. Eq. (3.21) indicates that ExpI is defined over the whole TI S++ Besides, the SPD manifold is connected [171]. By Lee [131, Cor. 6.20], CRIM is complete.
Geodesic. For simplicity, we abbreviate ⊕LieAI as ⊕. The geodesic connecting P and Q can be obtained as follows: γ(t; P, Q) = ExpP (t LogP (Q)) (1) = ⊖ ExpL⊖P − (Chol∗,⊖P )−1 L−1 Chol∗,P (t LogP (Q)) L−1 (2) = ⊖ ExpL⊖P t LogL⊖P (⊖Q) n o e . = ⊖ γ AI (t; Pe, Q)
The above comes from the following:
(1) The second equality in Eq. (B.17); (2) The first equality in Eq. (B.18).
296
(B.19)
Appendix B. Proofs
B.2.9
Proof of Thm. 72
Proof. Without loss of generality, we focus on the case of the left-invariant metric. The results for the right-invariant metric can be proven similarly. We denote Eqs. (3.8) to (3.10) on Mi , i = 1, 2 as the mapping ξ i (·|M, v 2 , B, s). Let N B = {Pi }N i=1 and f (B) = {f (Pi )}i=1 . The core of this proof lies in three points:
(1) The Fréchet mean and variance of B in M1 correspond to the counterparts of f (B) in M2 . (2) ξ 1 (Pi |M, v 2 , B, s) in M1 is equal to f −1 (ξ 2 (f (Pi )|f (M ), v 2 , f (B), s)). (3) The updates of running statistics in M1 correspond to the counterparts in M2 .
We denote M as the Fréchet mean of B, and v 2 as the Fréchet variance of B. Then, by the isometry of f , the Fréchet mean and variance of f (B) are f (M ) and v 2 , respectively. On Mi , i = 1, 2, we denote Li , ⊕i , Expi , Logi as the Lie group and Riemannian operators, E i as the neutral element, and Eq. (3.9) as ϕis (·). With the isometry and Lie group isomorphism of f , we have the following equations: L1⊖1 M = f −1 ◦ L2⊖2 f (M ) ◦f, ϕ1s = Exp1E 1 s Log1E 1 (·) = f −1 Exp2E 2 s Log2E 2 (f (·)) = f −1 ◦ ϕ2s ◦ f,
L1B = f −1 ◦ L2f (B) ◦f.
(B.20) (B.21) (B.22) (B.23) (B.24)
Then we have ξ 1 (Pi |M, v 2 , B, s) = f −1 (ξ 2 (f (Pi )|f (M ), v 2 , f (B), s))
(B.25)
Lastly, we show the correspondence between running statistics. Since the Fréchet variance is the same for both B and f (B), we focus on the running mean. Let Mr and f (Mr ) denote the initial values of the running means in M1 and M2 , respectively, and let WFMi represent the weighted Fréchet mean in Mi . Then the updated running mean in M1 is WFM1 ({1 − η, η}, {Mr , M }) = f −1 (WFM2 ({1 − η, η}, {f (Mr ), f (M )})) 297
(B.26)
B.3. Gyrogroup Batch Normalization We can further simplify the above equation as (B.27)
WFM1 = f −1 ◦ WFM2 ◦f
Denoting LieBNi as the LieBN algorithm on Mi , Eq. (B.25) and Eq. (B.27) imply that LieBN1 (Pi |B, s, ϵ, η) = f −1 LieBN2 (f (Pi )|f (B), s, ϵ, η) . (B.28)
B.3
Gyrogroup Batch Normalization
B.3.1
Proof of Thm. 74
Proof. By Nguyen and Yang [159, Lem. 2.3], easy computations show that Eq. (3.33) f n). Without loss of generality, we holds for Gr(p, n) if and only if it holds for Gr(p, prove the case for the projector perspective. f n), Nguyen [157, Def. 3.18] gives the expression for gyration: Given any P, Q ∈ Gr(p, gyr[⊖P, P ]Q = F (⊖P, P )Q (F (⊖P, P ))−1 ,
(B.29)
with F (⊖P, P ) defined as h i h i h i F (⊖P, P ) = exp − ⊖P ⊕ P , Iep,n exp ⊖P , Iep,n exp P , Iep,n ,
(B.30)
where (·) = LogIep,n (·). This equation can be further simplified as
i h i h (1) F (⊖P, P ) = exp(0) exp ⊖P , Iep,n exp P , Iep,n i h i h (2) e e = exp ⊖P , Ip,n exp P , Ip,n (3)
= In .
The above derivation follows from (1) ⊖P ⊕ P = Iep,n = 0n×n ∈ Rn×n . (2) exp(0) = In .
(3) ⊖P = −P and exp
h
−P , Iep,n
i
= exp
h
298
P , Iep,n
i−1
.
(B.31)
Appendix B. Proofs Therefore, gyr[⊖P, P ] is the identity map.
B.3.2
Proof of Thm. 75
Proof. This theorem follows from Ungar [200, Thms. 2.10–2.11], which presents some useful properties for gyrogroups. We argue that all the properties except gyr[a, a] = id are independent of the left reduction law (G4), and are therefore satisfied on pseudoreductive gyrogroups. All the properties can be proven in the same way as in Ungar [200, Thms. 2.10–2.11]. We summarize the logic in the following: • left gyroassociativity ⇒ Case (1) • left gyroassociativity + Case (1) ⇒ Case (2) • definition ⇒ Case (3) • left gyroassociativity + Case (1) + Case (3) ⇒ Case (4) • definition ⇒ Case (5) • left gyroassociativity + (G2) + Case (1) + Case (3) + Case (4) + Case (5) ⇒ Case (6) • Case (1)+Case (6) ⇒ Case (7) • left gyroassociativity +Case (3) ⇒ Case (8) • left gyroassociativity + left cancellation in Case (8) ⇒ Case (9) • gyro identity in Case (9) ⇒ Case (10) • Case (10) ⇒ Case (11) • left cancellation in Case (8) + gyro identity in Case (9) ⇒ Case (12) • left cancellation in Case (8) + gyro identity in Case (9) ⇒ Case (13)
299
B.3. Gyrogroup Batch Normalization
B.3.3
Proof of Thm. 77
Proof. (1)
y = x ⊕ (⊖x ⊕ y) (2)
= Expx (PTe→x (Loge (⊖x ⊕ y)))
(B.32)
(3)
⇒ Logx (y) = PTe→x (Loge (⊖x ⊕ y)) .
The above comes from the following. (1) Left cancellation law. (2) Definition of gyroaddition. (3) Applying Logx (·) to both sides. By the last equation, we have
d(x, y) = ∥Logx (y)∥x = ∥PTe→x (Loge (⊖x ⊕ y))∥x (1)
= ∥Loge (⊖x ⊕ y)∥e
(B.33)
= dgyr (x, y), where (1) comes from • Parallel transport preserving the norm [69, Sec. 3.1]. • PTx→e ◦ PTe→x (v) = v, ∀v ∈ Te M.
B.3.4
Proof of Thm. 78
Proof. We denote the Riemannian logarithm, gyration, and gyronorm on {M, g} by g gyr, Log, gyr, and ∥·∥gyr , respectively, while Log, f and ∥·∥gyr f are their counterparts on f {M, ge}. We recall the following from Nguyen and Yang [159, Lems. 2.1–2.3]: e x ⊕ y = ϕ−1 ϕ(x)⊕ϕ(y) , e t ⊙ x = ϕ−1 t⊙ϕ(x) ,
where x, y, z ∈ M.
gyr[x, y]z = ϕ−1 (gyr[ϕ(x), f ϕ(y)]ϕ(z)) , 300
(B.34)
(B.35) (B.36)
Appendix B. Proofs Gyrodistance. e e dg ⊕ϕ(y) gyr (ϕ(x), ϕ(y)) = ⊖ϕ(x) gyr f (1)
= ∥ϕ(⊖x ⊕ y)∥gyr f rD E ]ee(ϕ(⊖x ⊕ y)), Log ]ee(ϕ(⊖x ⊕ y)) = Log ee q (2) = ⟨Loge (⊖x ⊕ y), Loge (⊖x ⊕ y)⟩e
(B.37)
= ∥⊖x ⊕ y∥gyr = dgyr (x, y).
The derivation above comes from the following. (1) By Eqs. (B.34) and (B.35): e ϕ(⊖x ⊕ y) = ϕ(⊖x)⊕ϕ(y).
(B.38)
(2) By the isometry: g ϕ(x) (ϕ(y)) , ∀x, y ∈ M, Logx (y) = (ϕ∗,x )−1 Log
⟨v, w⟩x = ⟨ϕ∗,x (v), ϕ∗,x (w)⟩ϕ(x) , ∀x ∈ M and ∀v, w ∈ Tx M,
(B.39) (B.40)
f where ϕ∗,x is the differential map. Here, the RHSs contain the operators over M, whereas the LHSs involve those over M.
ϕ.
Gyroisometry. Given any x, y, z, a ∈ M, we have the following by the isometry of For the gyroinverse: e e dgyr ⊖ϕ(y)) = dgyr f (⊖ϕ(x), f (ϕ(⊖x), ϕ(⊖y)) = dgyr (⊖x, ⊖y)
= dgyr (x, y)
= dgyr f (ϕ(x), ϕ(y)). 301
(B.41)
B.3. Gyrogroup Batch Normalization For the gyration: dgyr f ϕ(a)]ϕ(x), gyr[ϕ(z), f ϕ(a)]ϕ(y)) f (gyr[ϕ(z), = dgyr f (ϕ(gyr[z, a]x), ϕ(gyr[z, a]y)) = dgyr (gyr[z, a]x, gyr[z, a]y)
(B.42)
= dgyr (x, y) = dgyr f (ϕ(x), ϕ(y)). For the left gyrotranslation: e e dgyr ϕ(z)⊕ϕ(y)) f (ϕ(z)⊕ϕ(x), = dgyr f (ϕ(z ⊕ x), ϕ(z ⊕ y))
(B.43)
= dgyr (z ⊕ x, z ⊕ y) = dgyr (x, y) = dgyr f (ϕ(x), ϕ(y)).
B.3.5
Proof of Thm. 79
We first prove a useful lemma. Lemma 190 (Left gyrotranslation law). Every pseudo-reductive gyrogroup {G, ⊕} verifies the left gyrotranslation law: ⊖ (x ⊕ y) ⊕ (x ⊕ z) = gyr[x, y] (⊖y ⊕ z) ,
∀x, y, z ∈ G.
(B.44)
Proof. This lemma generalizes Nguyen and Yang [159, Lems. I.1 and L.1], which prove the left gyrotranslation law on the specific gyrogroups of the SPD and Grassmannian manifolds. Their proof only relies on the left cancellation and the basic axioms (G1– G3). Note that the original proof of left gyrotranslation on the Grassmannian [159, Lem. I.1] is questionable, as it relies on the left cancellation of gyrogroups, and the Grassmannian is not a gyrogroup but a non-reductive gyrogroup. Fortunately, as we show in Thm. 75, the general pseudo-reductive gyrogroups, including the Grassmannian, enjoy left cancellation. Therefore, the proof in Nguyen and Yang [159, Lem. I.1] can be readily generalized to pseudo-reductive gyrogroups. Proof of Thm. 79. ⇒: For any z, a ∈ G, the gyroautomorphism can be expressed by 302
Appendix B. Proofs the gyrator identity in Thm. 75: (B.45)
gyr[x, y]z = X ⊕ z̄,
(B.46)
gyr[x, y]a = X ⊕ ā,
where X = ⊖(x ⊕ y), z̄ = x ⊕ (y ⊕ z), and ā = x ⊕ (y ⊕ a). Then Nguyen and Yang [159, Eq. (30)], derived for the specific Grassmannian, can be directly extended to pseudoreductive gyrogroups, as it only relies on left gyrotranslation, invariance of the norm under gyroautomorphisms, and the axioms (G1–G3). ⇐: ∥gyr[x, y](z)∥gyr = ∥gyr[x, y](⊖e ⊕ z)∥gyr
(Case (7) in Thm. 75 indicates ⊖e = e)
= ∥⊖ gyr[x, y](e) ⊕ gyr[x, y](z)∥gyr
(automorphism)
= d (gyr[x, y](e), gyr[x, y](z)) = d (e, z) = ∥⊖e ⊕ z∥gyr = ∥z∥gyr .
B.3.6
(B.47)
Proof of Thm. 80
Proof. Given any x, y, z ∈ G, we prove the two claims as follows. Gyroisometry of the left gyrotranslation. This property generalizes Nguyen and Yang [159, Thms. 2.12 and 2.16], which deal with the gyrotranslation in the SPD and Grassmannian, respectively. We have the following: d(Lx (y), Lx (z)) = d(x ⊕ y, x ⊕ z) = ∥⊖(x ⊕ y) ⊕ (x ⊕ z)∥gyr = ∥gyr[x, y] (⊖y ⊕ z)∥gyr = ∥⊖y ⊕ z∥gyr
(left gyrotranslation law)
(gyroisometry of the automorphism)
= d(y, z). 303
(B.48)
B.3. Gyrogroup Batch Normalization Gyroisometry of the gyroinverse. d(⊖x, ⊖y) = ∥x ⊖ y∥gyr = ∥⊖y ⊕ x∥gyr
(gyrocommutativity and gyroisometry of the automorphism)
= d(y, x) = d(x, y)
B.3.7
(symmetry of the geodesic distance).
(B.49)
Proof of Thm. 81
We give a concise argument based on Thms. 77, 79 and 80. Proof. First, all these gyrospaces are characterized by Eqs. (2.77) and (2.78). By Thm. 77, their gyrodistances agree with the geodesic distances. It thus remains to establish the gyroisometries. As shown by Thms. 79 and 80, it suffices to show that gyrations in each space preserve the gyronorm. This argument on the SPD and ONB Grassmannian has already been proven [159, Lems. L.2 and I.2]. Since the ONB Grassmannian is isometric to the PP Grassmannian via Eq. (2.138), Thm. 78 implies that the same arguments apply to the PP. We therefore only need to treat stnK with K ≤ 0. As the Euclidean case Rn is trivial, we only need to show the Poincaré ball. In the following, a, b, x, y are arbitrary points in PnK . Norm invariance under gyrations. As the Poincaré ball forms a real innerproduct gyrovector space [200, Def. 6.2 and Thm. 6.85], any gyration preserves the Euclidean norm: ∥gyr[a, b](x)∥ = ∥x∥ , ∀x ∈ stnK . (B.50) For the gyronorm, we further have ∥gyr[a, b]x∥gyr = ∥ Log0 (gyr[a, b]x)∥0 p |K|∥ gyr[a, b]x∥ = √2 tanh−1 |K| p = √2 tanh−1 |K|∥x∥ |K|
= ∥x∥gyr .
304
(B.51)
Appendix B. Proofs
B.3.8
Proof of Thm. 83
Proof. According to Thm. 80, any left gyrotranslation is a gyroisometry. For any y ∈ M, we have the following: (1)
d (β ⊕ xi , y) = d (⊖β ⊕ (β ⊕ xi ), ⊖β ⊕ y) (2)
= d ((⊖β ⊕ β) ⊕ gyr[⊖β, β](xi ), ⊖β ⊕ y)
(B.52)
(3)
= d (xi , ⊖β ⊕ y) .
The above comes from the following. (1) Any left gyrotranslation is a gyroisometry. (2) Left gyroassociative law. (3) ⊖β ⊕ β = e and pseudo-reduction. Denoting the gyromean of {xi } and {β ⊕ xi } as µ and µ e, we have the following: (1)
β ⊕ µ = β ⊕ (⊖β ⊕ µ e) (2)
= gyr[β, ⊖β](e µ)
(3)
The above comes from the following.
=µ e.
(1) Eq. (B.52) indicates that µ = ⊖β ⊕ µ e. (2) Left gyroassociative law. (3) Pseudo-reduction. 305
(B.53)
B.3. Gyrogroup Batch Normalization Algorithm 4: ONB Grassmannian logarithm [19, Alg. 5.3] Input: U, Y ∈ Gr(p, n) are Stiefel representatives under the ONB perspective. SVD
QSR⊤ := Y ⊤ U with S in ascending order, and Q and R flipped column-wise p Ŝ) ⊤ accordingly Ŝ = Ip − S 2 ∆ = (In − U U ⊤ )Y Q arcsin( R Ŝ
Output: LogU (Y ) = ∆
Now, we proceed to deal with the second property. We have the following: (1)
d(t ⊙ xi , e) = ∥⊖e ⊕ (t ⊙ xi )∥gyr (2)
= ∥t ⊙ xi ∥gyr
= ∥t Loge (xi )∥e = |t| ∥Loge (xi )∥e = |t| ∥xi ∥gyr
(B.54)
(3)
= |t| ∥⊖e ⊕ xi ∥gyr
= |t| d(e, xi ) (4)
= |t| d(xi , e)
The above follows from the following. (1) Symmetry of gyrodistance (as geodesic distance). (2) ⊖e = e. (3) xi = ⊖e ⊕ xi . (4) Symmetry of gyrodistance (as geodesic distance). The last equation in Eq. (B.54) indicates the homogeneity of dispersion from e.
B.3.9
Proof of Thm. 86
We first review a fast and stable algorithm for the ONB Grassmannian logarithm [19, Alg. 5.3], and the calculation of Grassmannian logarithm under the projector perspective by the ONB Grassmannian logarithm [161, Prop. 3.12]. Alg. 4 reviews a fast and stable algorithm for the Grassmannian Riemannian logarithm under the ONB perspective Gr(p, n). The vanilla Riemannian logarithm in Tabs. 2.10 and 2.11 requires an n × p SVD and a p × p matrix inverse, while Alg. 4 only requires a p × p SVD. Therefore, Alg. 4 is more efficient than the vanilla logarithm. 306
Appendix B. Proofs Besides, Alg. 4 can also return a unique tangent vector when Y is in the cut locus of U . For more details, please refer to Bendokat et al. [19, Sec. 5.2]. As the projector perspective is isometric to the ONB perspective, the Grassmannian logarithm under the projector perspective can be calculated by the ONB Grassmannian logarithm [161, Prop. 3.12]. f n) with U = π −1 (P ) and V = Proposition 191 ([161]). Given any P, Q ∈ Gr(p, g P (Q) on Gr(p, f n) is given as π −1 (Q), the Riemannian logarithm Log g P (Q) = π∗,U (LogU V ) , Log
(B.55)
π∗,U (∆) = ∆U ⊤ + U ∆⊤ , ∀∆ ∈ TU Gr(p, n).
(B.56)
where Log is the Riemannian logarithm under the ONB perspective, π∗,U : f n) is the differential map of π at U , which is defined as TU Gr(p, n) → TP Gr(p, Now, we begin to present the proof. ge . Proof of Thm. 86. We first show the expression for LogIp,n and Log Ip,n First note the following:
0 0 0 In−p
⊤ (In − Ip,n Ip,n )=
⊤
U ⊤ Ip,n = U1⊤ , U2 = U1⊤ ,
Ip 0
!
,
!
(B.57)
(B.58)
By the above two equations, the ONB Grassmannian logarithm at Ip,n is LogIp,n (U ) = = = SVD
0 0 0 In−p 0
!
! U1 arcsin(Ŝ) ⊤ R (Alg. 4) Q U2 Ŝ !
Ŝ) ⊤ U2 Q arcsin( R Ŝ ! 0 , e U2
(B.59)
where QSR⊤ := U1⊤ with S in ascending order, and Q and R flipped column-wise 307
B.3. Gyrogroup Batch Normalization accordingly, and Ŝ =
p Ip − S 2 .
g e , we have For Log Ip,n
g e (U U ⊤ ) (1) Log = π Log (U ) ∗,Ip,n Ip,n Ip,n !! 0 (2) = π∗,Ip,n e2 U ! e⊤ 0 U (3) 2 = . e U2 0
(B.60)
The above derivation comes from the following. (1) Thm. 191 (2) Eq. (B.59)
⊤ ⊤ (3) For any ∆ = (∆⊤ 1 , ∆2 ) ∈ TIp,n Gr(p, n), where ∆1 is p × p, we have the following:
π∗,Ip,n
∆1 ∆2
!!
!
=
∆1 ∆2
=
∆1 0 ∆2 0
=
∆1 + ∆⊤ ∆⊤ 1 2 ∆2 0
Ip 0
(Ip , 0) + !
+
!
∆⊤ ∆⊤ 1 2 0 0 !
⊤ ∆⊤ 1 , ∆2
!
(B.61)
.
Combining all the above results together, we have the following: h i g e (U U ⊤ ), Iep,n [U U ⊤ , Iep,n ] = Log Ip,n " ! # e2⊤ 0 U = , Iep,n e2 0 U ! ! e2⊤ 0 U Ip 0 = − e2 0 U 0 0 ! e2⊤ 0 −U = . e2 U 0 308
Ip 0 0 0
!
e2⊤ 0 U e2 0 U
!
(B.62)
Appendix B. Proofs
B.3.10
Proof of Thm. 88
Proof. Given v ∈ T0 stnK , we have the following: (B.63)
λK 0 = 2,
Therefore, we have
p y −1 Log0 (y) = tanK , |K| ∥y∥ p |K| ∥y∥ λK 2 PT0→x (v) = 0K gyr[x, −0](v) = K v, λx λx p v . |K| ∥v∥ p Exp0 (v) = tanK |K| ∥v∥
Expx (PT0→x (Log0 (y))) = Expx
y p p tan−1 |K| ∥y∥ K ∥y∥ |K|λK x 2
(B.64) (B.65) (B.66)
!
(B.67)
= x ⊕K y, which implies
(B.68)
Exp0 (t Log0 (x)) = t ⊙K x.
B.3.11
Proof of Thm. 89
Proof. We only need to show the case of K > 0. Let x, y, z, w be any vectors in stnK , and s, t ∈ R be real scalars. We first notice the gyroaddition is a linear combination:
(1 − 2K ⟨x, y⟩ − K ∥y∥2 )x + (1 + K ∥x∥2 )y Ax + By x ⊕K y = = , 2 2 2 D 1 − 2K ⟨x, y⟩ + K ∥x∥ ∥y∥ where
A = 1 − 2K ⟨x, y⟩ − K ∥y∥2 ,
B = 1 + K ∥x∥2 ,
2
(B.69)
(B.70) 2
D = 1 − 2K ⟨x, y⟩ + K 2 ∥x∥ ∥y∥ . Recalling Eq. (2.146), the gyration gyr[x, y](z) is also a linear expression: gyr[x, y](z) = z + f1 · x + f2 · y. 309
(B.71)
B.3. Gyrogroup Batch Normalization Axioms G1–G3. Since e = 0 and ⊖K x = −x, G1 and G2 can be immediately verified. The left gyroassociative law follows from the definition of gyration [13, Eq. 36]: gyr[x, y](z) = − (x ⊕K y) ⊕K (x ⊕K (y ⊕K z)) .
(B.72)
To confirm that any gyration is an automorphism of (stnK , ⊕K ), we verify the following identity using symbolic computation: gyr[x, y](w ⊕K z) = gyr[x, y](w) ⊕K gyr[x, y](z).
(B.73)
Expanding all necessary inner products (e.g., ⟨x, w ⊕K z⟩, ⟨y, w ⊕K z⟩) and norms in terms of ⟨x, w⟩, ⟨x, z⟩, etc., we express both sides of Eq. (B.73) as linear combinations: gyr[x, y](w ⊕K z) = f1 x + f2 y + f3 w + f4 z,
gyr[x, y](w) ⊕K gyr[x, y](z) = f1′ x + f2′ y + f3′ w + f4′ z.
(B.74) (B.75)
We use SymPy to compare the coefficients of x, y, w, and z on both sides. The above is exposed in stereographic_gyr_automorphism.py. (G4) Left reduction law. Similarly, we expand gyr[x ⊕K z, y](w) and gyr[x, y](w) and compare the coefficients, as implemented in stereographic_left_reduction.py. Gyrocommutative law. This has been verified by Bachmann et al. [13, Lem. 11]. (V1) Identity scalar multiplication. This can be directly verified by definition: √ tan t tan−1 ( K ∥x∥) x √ , ∀x ̸= 0. t ⊙K x = ∥x∥ K
(B.76)
(V2) Scalar distributive law. We want to verify (s + t) ⊙K x = (s ⊙K x) ⊕K (t ⊙K x).
(B.77)
√ We only need to show the case of x ̸= 0. Let θ = tan−1 ( K ∥x∥). Then scalar multiplication is simplified as tan(tθ) t ⊙K x = √ x. K ∥x∥
(B.78)
Using symbolic computation (see stereographic_gyr_v2.py), we obtain the following 310
Appendix B. Proofs w.r.t. Eq. (B.77): tan ((s + t)θ) √ , K ∥x∥ tan(sθ) + tan(tθ) RHS = √ . K ∥x∥ (1 − tan(sθ) tan(tθ)) LHS =
(B.79) (B.80)
The identity follows from the tangent addition formula: tan((s + t)θ) =
tan(sθ) + tan(tθ) . 1 − tan(sθ) tan(tθ)
(B.81)
(V3) Scalar associative law. This can be directly verified by definition. (V4) Gyroautomorphism. We now verify
Let us denote
gyr[x, y](t ⊙K z) = t ⊙K gyr[x, y](z).
(B.82)
√ K ∥z∥ tan t tan−1 (z) √ . αt = K ∥z∥
(B.83)
By the linearity of gyration [13, Lem. 11], the left-hand side of Eq. (B.82) becomes (z)
LHS = αt gyr[x, y](z),
(B.84)
while the right-hand side reads gyr[x,y](z)
RHS = αt
gyr[x, y](z).
(B.85)
(·)
Since αt depends only on the norm of its argument, it suffices to show ∥z∥ = ∥gyr[x, y](z)∥ ,
(B.86)
which holds as proven by Bachmann et al. [13, Lem. 11, iv)]. (V5) Identity gyroautomorphism. In each of the three special cases when (i) x = 0, or (ii) y = 0, or (iii) x and y are parallel in V, x ∥ y, we have (1)
(B.87)
(2)
(B.88)
gyr[0, x](z) = z, gyr[x, 0](z) = z, 311
B.3. Gyrogroup Batch Normalization (3)
gyr[x, y](z) = z,
x ∥ y,
(B.89)
where (1–3) come from Ast x + Bst y = 0 in Eq. (2.146). Therefore, we have (x)
gyr[s ⊙K x, t ⊙K x] = gyr[αs(x) x, αt x] = id .
(B.90)
Note that (1–2) are also implied by the first gyrogroup theorem [200, Thm. 2.10].
B.3.12
Proof of Thm. 90
This proof largely follows the one for Thm. 81. Proof. Norm invariance under gyrations. By Bachmann et al. [13, Lem. 11], any gyration preserves the Euclidean norm: ∥ gyr[a, b](x)∥ = ∥x∥,
∀x ∈ stnK .
(B.91)
For the gyronorm, we further have ∥gyr[a, b]x∥gyr = ∥ Log0 (gyr[a, b]x)∥0 p = √2 tan−1 |K|∥ gyr[a, b]x∥ K |K| p = √2 tan−1 |K|∥x∥ K
(B.92)
|K|
= ∥x∥gyr ,
with tanK = tanh for K < 0 and tanK = tan for K > 0. Isometry of left gyrotranslation and gyroinverse. As shown in Thm. 89, forms a gyrocommutative gyrogroup. By Thm. 80, the left gyrotranslation and gyroinverse are gyroisometries. stnK
B.3.13
Proof of Thm. 91
Proof. Non-singular cases. Denoting ∥x∥s = ∥xs ∥ , ∀x ∈ MnK , we have p ( cos−1 |K|xt ) K (t Log0 x)s = t p xs , |K| ∥xs ∥ p cos−1 ( |K|xt ) ∥t Log0 x∥s = |t| Kp , |K| 312
(B.93) (B.94)
Appendix B. Proofs p p |K| ∥t Log0 x∥s = |t| cos−1 |K|xt ), K (
(B.95)
(t Log0 x)s xs = sgn(t) . ∥t Log0 x∥s ∥xs ∥
(B.96)
For t ̸= 0, the normalized spatial direction satisfies
Since cosK is even and sinK is odd, substituting the above expressions into Exp0 in p Tab. 2.14 recovers the signed argument t cos−1 K ( |K|xt ) in Eq. (3.55). The cases t = 0 and x = 0 are handled by the first branch of Eq. (3.55). For t = −1, we have
p −1 cos − cos ( |K|x ) t K 1 K √ −1 ⊙M −1 K x = p sinK − cosK ( |K|xt ) |K| xs ∥xs ∥ xt (1) √ −1 , = sinK cosK ( |K|xt ) √ − xs
(B.97)
|K|∥xs ∥
where (1) comes √ from cosK (−θ) = cosK (θ) and sinK (−θ) = − sinK (θ). The rest is to
show
sinK cos−1 K (
√
|K|xt )
= 1:
|K|∥xs ∥
sinK
p p |K|xt ) (1) sign(K)(1 − |K|x2 ) (2) t p p = = 1. |K| ∥xs ∥ |K| ∥xs ∥ cos−1 K (
(B.98)
The derivation above is based on the following: (1) cos2K (θ) + sign(K) sin2K (θ) = 1 implies sinK (θ) =
q sign(K)(1 − cos2K (θ)), ∀θ > 0.
(B.99)
−1 Considering cos−1 K (θ) ≥ 0 for any θ ∈ dom(cosK (·)), we have (1).
(2) ⟨x, x⟩K = sign(K)x2t + ∥xs ∥2 = ⇒ K ∥xs ∥2 = 1 − |K|x2t
1 K
⇒ |K| ∥xs ∥2 = sign(K)(1 − |K|x2t ). 313
(B.100)
B.3. Gyrogroup Batch Normalization Equivalence in the singular cases. By Eq. (3.50), √
K ∥u∥ =
√
K ∥xs ∥ √ . 1 + Kxt
(B.101)
√ √ √ Writing θ = cos−1 ( Kxt ) or Kxt = cos θ, the sphere constraint gives K ∥xs ∥ = sin θ. Hence √ K ∥u∥ =
sin θ θ = tan , 1 + cos θ 2
⇒ tan−1
√
θ K ∥u∥ = . 2
(B.102)
√ Therefore, (i) is equivalent to (ii) since t tan−1 ( K ∥u∥) = tθ2 .
Next, we use the closed form of gyromultiplication Eq. (3.58). For K > 0, the gyromultiplication reads cos (tθ) 1 t ⊙M (B.103) sin(tθ) . K x = √ xs K ∥xs ∥ √ ⊤ = If (ii) holds, then sin(tθ) = 0 and cos(tθ) = −1, whence t ⊙M K x = [−1/ K, 0] −0.
B.3.14
Proof of Thm. 92
⊤ ⊤ Proof. We denote x ⊕M K y = [zt , zs ] . As the results are trivial under x = 0 or y = 0, we assume x ̸= 0 and y ̸= 0 in the following.
Non-singular cases. We first consider the non-singular case: i) K < 0; ii) K > 0, v n n x, y ̸= ±0, and u ̸= K∥v∥ 2 . Since gyroadditions on MK and stK are both defined by Eq. (2.77), we have the following under isometries [159, Lem. 2.2]: n n →stn (x) ⊕K πMn →stn (y) . x ⊕M π →M M K y = πstn K K K K K K
(B.104)
Following Eq. (B.70), we rewrite the gyroaddition on the stereographic model as u ⊕K v =
(1 − 2K⟨u, v⟩ − K∥v∥2 )u + (1 + K∥u∥2 )v Au + Bv = , 1 − 2K⟨u, v⟩ + K 2 ∥u∥2 ∥v∥2 ∆
(B.105)
where A = 1 − 2K⟨u, v⟩ − K∥v∥2 ,
B = 1 + K∥u∥2 ,
314
∆ = 1 − 2K⟨u, v⟩ + K 2 ∥u∥2 ∥v∥2 . (B.106)
Appendix B. Proofs Using ⟨u, v⟩ = sxy /(ab), ∥u∥2 = nx /a2 , ∥v∥2 = ny /b2 , one obtains A= Hence
ab2 − 2Kbsxy − Kany , ab2
B=
a2 + Knx , a2
Au + Bv = stnK ∋ w = u ⊕K v = ∆
1 ∆
D . a2 b 2
(B.107)
A B xs + y s . a b
(B.108)
2w . 1 + K∥w∥2
(B.109)
∆=
Apply πstnK →MnK to w: 1 1 − K∥w∥2 zt = p , |K| 1 + K∥w∥2
zs =
A direct expansion (radius_gyroaddition.py) yields ∥w∥2 =
N A2 ∥u∥2 + B 2 ∥v∥2 + 2AB⟨u, v⟩ = . 2 ∆ D
(B.110)
Substituting the above into Eq. (B.109) gives 1 1 − KN/D 1 D − KN zt = p =p . |K| 1 + KN/D |K| D + KN
(B.111)
For zs , using Eq. (B.108), Eq. (B.107) and 1 + K∥w∥2 = (D + KN )/D, 2 1 A B zs = · xs + y s 1 + KN/D ∆ a b 2 = (Aab2 )xs + (Ba2 b)ys D + KN 2 (As xs + Ay ys ) . = D + KN
(B.112)
Gyroaddition in singular cases (K > 0). We show that the definition Eq. (3.59) indeed returns −0 in this case. √ Step 1: Log0 (y). Write R = 1/ K and choose the polar angle θ ∈ (0, π) so that "
# R cos θ x= , R sin θŝ
"
# −R cos θ y= , R sin θŝ 315
ŝ =
xs . ∥xs ∥
(B.113)
B.3. Gyrogroup Batch Normalization Using Log0 in Tab. 2.14, " # 0 √ 0 . Log0 (y) = cos−1 ( Kyt ) = √ ys (π − θ)Rŝ K∥ys ∥
Step 2: Parallel transport. With 1 + in Tab. 2.14, we have
√
(B.114)
Kxt = 1 + cos θ ̸= 0 (since x ̸= −0) and PT0→x
" # # 0 K⟨xs , (π − θ)Rŝ⟩ xt + √1K √ . PT0→x (Log0 (y)) = − (π − θ)Rŝ 1 + Kxt xs "
(B.115)
Using ⟨xs , ŝ⟩ = ∥xs ∥ = R sin θ, KR2 = 1, and xt = R cos θ, the scalar factor in Eq. (B.115) equals K(π − θ)R⟨xs , ŝ⟩ (π − θ) sin θ √ = . (B.116) 1 + cos θ 1 + Kxt A short simplification then yields the compact form "
# − sin θ w = PT0→x (Log0 (y)) = R(π − θ) . cos θŝ Note that ∥w∥x = R(π − θ), so with α =
√
(B.117)
K∥w∥x we have α = π − θ.
Step 3: Exponential at x. By Tab. 2.14, Expx (w) = cos(α)x +
sin(α) w, α
α=
√
K∥w∥x ,
(B.118)
and here cos(α) = cos(π − θ) = − cos θ, sin(α) = sin(π − θ) = sin θ. Substituting Eq. (B.113) and Eq. (B.117) into Eq. (B.118), "
"
# " ## " # cos θ − sin θ −1 Expx (w) = R − cos θ + sin θ =R = −0. sin θŝ cos θŝ 0
(B.119)
We conclude that x ⊕M K y = −0 in the singular configuration. Equivalence in singular cases (K > 0). Recalling the isometry between MnK and stnK , we write u = xas and v = ybs . 316
Appendix B. Proofs (i) ⇒ (ii). Write α = K ∥v∥2 > 0. From u = αv and Eq. (3.51) we get 1 1−α yt = √ , K1+α 1 1 − α1 1 1−α √ xt = √ = −yt , 1 = − K1+ α K1+α 2u 2v/α 2v xs = = ys , = = 2 2 2 1 + K∥u∥ 1 + K∥v∥ /α 1+α
(B.120)
Hence xs = ys and xt = −yt . (ii) ⇒ (iii). Under xs = ys and xt = −yt , set n = ∥xs ∥2 = ∥ys ∥2 and note sxy = n, nx = ny = n. Then, D = a2 b2 − 2Kabsxy + K 2 nx ny = (ab − Kn)2 .
(B.121)
On the sphere SnK , we have the constraint x2t + ∥xs ∥2 = 1/K . Since yt = −xt and ∥ys ∥2 = ∥xs ∥2 = n, √ √ ab = (1 + Kxt )(1 + Kyt ) √ √ = (1 + Kxt )(1 − Kxt ) (B.122) = 1 − Kx2t 1 − n = Kn. =1−K K
Therefore D = (ab − Kn)2 = 0.
(iii) ⇒ (i). This can be obtained using the Cauchy–Schwarz inequality, as in Bachmann et al. [13, App. C.2.1]. Consequences in singular cases (K > 0). Under (ii) we also have a + b = 2, so N = a2 n + 2abn + b2 n = (a + b)2 n = 4n > 0.
(B.123)
Using D = 0 and N > 0, Eq. (B.111) gives 1 1 −KN zt = √ = −√ . K KN K
(B.124)
Besides, since for xs = ys one checks As + Ay = (a + b)(ab − Kn) = 0, Eq. (B.112) yields zs = 0. On the other hand, u ⊕K v = ∞. 317
B.3. Gyrogroup Batch Normalization
B.3.15
Proof of Thm. 93
Proof. This is already implied by the proofs of Thms. 91 and 92.
B.3.16
Proof of Thm. 95
Proof. The gyrovector space over the stereographic model is also defined by Eqs. (2.77) and (2.78): x ⊕K y = Expx (PT0→x (Log0 (y))) , (B.125) r ⊙K x = Exp0 (r Log0 (x)) , with x, y ∈ stnK and r ∈ R. Since πMnK →stnK (0) = 0 and (stnK , ⊕K , ⊙K ) is a gyrovector M space, (MnK , ⊕M K , ⊙K ) is also a gyrovector space [159, Thm. 2.4].
B.3.17
Proof of Thm. 97
Proof. Isometries. The Beltrami–Klein model is isometric to the hyperboloid model by the following diffeomorphisms [131, Thm. 3.7]: πKnK →HnK : KnK ∋ x 7−→
x
1
p ,p √ −K 1 + K∥x∥2 1 + K∥x∥2
" # xt xs πHnK →KnK : HnK ∋ 7−→ √ ∈ KnK . −Kxt xs
!
∈ HnK ,
(B.126) (B.127)
Combining the isometries Eqs. (B.126), (B.127), (3.50) and (3.51), one can readily obtain the isometries between Beltrami–Klein and Poincaré ball models. Now, we turn to the differential maps. Given a curve over c(t) ∈ M with c(0) = x and c′ (0) = v, the 318
Appendix B. Proofs differential maps can be calculated by dπKnK →PnK (c(t)) d c(t) = p dt dt 1 + 1 + K∥c(t)∥2 t=0 t=0 p −1 1 + 1 + K∥x∥2 v − (1 + K∥x∥2 ) 2 K ⟨x, v⟩ x = 2 p 1 + 1 + K∥x∥2 =
1+
p
1
1 + K∥x∥2
v−
K ⟨x, v⟩ x, 2 p p 1 + 1 + K∥x∥2 1 + K∥x∥2
dπPnK →KnK (c(t)) 2c(t) d = dt dt 1 − K∥c(t)∥2 t=0 t=0 2v (1 − K∥x∥2 ) + 4K ⟨x, v⟩ x = (1 − K∥x∥2 )2 4K ⟨x, v⟩ 2 v+ = x. 2 (1 − K∥x∥ ) (1 − K∥x∥2 )2
(B.128)
Homomorphism. The homomorphism w.r.t. the scalar product can be readily obtained by the Riemannian isometry. Therefore, we first address it before proceeding to the addition. For simplicity, we denote ϕ = πPnK →KnK . Scalar product. As shown by Ungar [200], the geodesics under the Beltrami–Klein and Poincaré ball models are K γϕ(x)→ϕ(y) (t) = ϕ(x) ⊕E t ⊙E (−ϕ(x) ⊕E ϕ(y)) , P γx→y (t) = x ⊕M t ⊙M (−x ⊕M y) .
(B.129)
Here, we use the fact that the gyro inverses in the Möbius and Einstein gyrovector spaces are exactly the familiar vector inverse. The above geodesics satisfy ϕ(x ⊕M t ⊙M (−x ⊕M y)) = ϕ(x) ⊕E t ⊙E (−ϕ(x) ⊕E ϕ(y)) .
(B.130)
P K The above comes from ϕ(γx→y (t)) = γϕ(x)→ϕ(y) (t). In particular, the geodesic starting from the identity element yields (1)
P K (t) = t ⊙E ϕ(x), ϕ(t ⊙M x) = ϕ(γ0→x (t)) = γϕ(0)→ϕ(x)
(B.131)
where (1) comes from ϕ(0) = 0. The above holds for all x ∈ PnK and ∀t ∈ R, as every hyperbolic geometry is geodesically complete [131, p. 139]. 319
B.3. Gyrogroup Batch Normalization Addition. We first expand the LHS of Eq. (3.69). Inspired by Mao et al. [147, Eqs. 73–76], we express the Möbius addition as x ⊕M y = Bx+Cy by denoting A A = 1 − 2K⟨x, y⟩ + K 2 ∥x∥2 ∥y∥2 ,
B = 1 − 2K⟨x, y⟩ − K∥y∥2 ,
(B.132)
C = 1 + K∥x∥2 .
Then, the Möbius addition is ϕ(x ⊕M y) =
2 Bx+Cy A 2
1 − K Bx+Cy A 2ABx + 2ACy = A2 − K ∥Bx + Cy∥2 2ABx + 2ACy . = A2 − B 2 K ∥x∥2 − C 2 K ∥y∥2 − 2BCK ⟨x, y⟩
(B.133)
Inspired by Mao et al. [147, Eq. 77], we denote ∥x∥ = a, ∥y∥ = b, and ⟨x, y⟩ = ab cos(θ). In this way, we can resort to the symbolic computation package SymPy [150] for the heavy algebra computation, which brings ϕ(x ⊕M y) =
2 (−2Kab cos (θ) − Kb2 + 1) x K 2 a2 b2 − Ka2 − 4Kab cos (θ) − Kb2 + 1 2Ka2 + 2 + 2 2 2 y. K a b − Ka2 − 4Kab cos (θ) − Kb2 + 1
(B.134)
Now we turn to the RHS of Eq. (3.69). For any u, v ∈ KnK , the Einstein addition can be rewritten as 1 γu 1 u ⊕E v = u+ v−K ⟨u, v⟩ u 1 − K ⟨u, v⟩ γu 1 + γu (B.135) γu ⟨u, v⟩ 1 − K 1+γ 1 u = u+ v. 1 − K ⟨u, v⟩ γu (1 − K ⟨u, v⟩) 320
Appendix B. Proofs The gamma factor and inner product under isometry can be rewritten as 1 γϕ(x) = p 1 + K∥ϕ(x)∥2 1 =r 2 2 1 + K 1−K∥x∥2 ∥x∥2
1 − K∥x∥2 =q (1 − K∥x∥2 )2 + 4K∥x∥2
(1) 1 − K∥x∥
=
⟨ϕ(x), ϕ(y)⟩ =
(B.136)
2
1 + K∥x∥2
,
4 (1 − K∥x∥2 )(1 − K∥y∥2 )
⟨x, y⟩ .
Here, (1) comes from ∥x∥2 < − K1 ⇒ K∥x∥2 +1 > 0. Putting the above into Eq. (B.135) and following the same notation as Eq. (B.134), we can obtain the following by SymPy: ϕ(x) ⊕E ϕ(y) =
2 · (2Kab cos (θ) + Kb2 − 1) x 4Kab cos (θ) − (Ka2 − 1) (Kb2 − 1) −2Ka2 − 2 y, + 4Kab cos (θ) − (Ka2 − 1) (Kb2 − 1)
(B.137)
which is clearly equal to Eq. (B.134).
B.3.18
Proof of Thm. 98
Proof. For simplicity, we denote ϕ = πPnK →KnK . The homomorphism and bijection of ϕ imply the homomorphism of its inverse ϕ−1 . Also note that ϕ(0) = ϕ−1 (0) = 0. This yields x ⊕E y = ϕ ϕ−1 (x) ⊕M ϕ−1 (y) (1) P P P −1 = ϕ Expϕ−1 (x) PT0→ϕ−1 (x) Log0 (ϕ (y)) (B.138) (2) K K = ExpK x PT0→x Log0 (y) . The above comes from the following.
(1) For any Poincaré vectors u, w ∈ PnK , the following holds: PTP0→u (v) = LogPu (u ⊕M ExpP0 (v)) ⇒ u ⊕M w = ExpPu PTP0→u (LogP0 (w)) , (B.139)
where v = LogP0 (w) and the LHS comes from Ganea et al. [76, Thm. 4]; 321
B.3. Gyrogroup Batch Normalization (2) The second identity follows from the isometry: ExpPx (v) = ϕ−1 ExpK ϕ(x) (ϕ∗,x (v)) , K LogPx (y) = ϕ−1 ∗,x Logϕ(x) (ϕ(y)) ,
K PT (ϕ (v)) . PTPx→y (v) = ϕ−1 ∗,x ∗,y ϕ(x)→ϕ(y)
(B.140)
Similar to the gyro addition, the isometry of ϕ implies the following with respect to the gyro scalar product: t ⊙E x = ϕ t ⊙M ϕ−1 (x) (1) (B.141) = ϕ ExpP0 t LogP0 (ϕ−1 (x)) (2) K = ExpK 0 t Log0 (x) .
The above comes from the following. (1) Ganea et al. [76, Lem. 3];
(2) The isometry of ϕ.
B.3.19
Proof of Thm. 99
Proof. Following Thm. 98, we denote ϕ = πPnK →KnK and ψ = πKnK →PnK . The results can be obtained by the properties of isometry and isomorphism of ϕ. Riemannian exponential and logarithmic maps at the zero vector. First, we recall that the expressions of ExpP0 (v) and LogP0 (v) under the Poincaré ball model are exactly those in Eqs. (3.72) and (3.73) [76, Eq. 13]. 322
Appendix B. Proofs By Riemannian isometry, we have the following: P ExpK 0 (v) = ϕ Exp0 (ψ∗,0 (v)) √ −K v ∥v∥ √ = ϕ tanh 2 −K∥v∥ √ −K tanh ∥v∥ 2 2 √ v = √ 2 −K −K∥v∥ tanh ∥v∥ 2 √ 1 − K∥v∥2 −K∥v∥ √ −K 2 tanh ∥v∥ 2 v = 2 √ √ −K∥v∥ 1 + tanh −K ∥v∥
(B.142)
2
√ v = tanh( −K∥v∥) √ , −K∥v∥
(1)
2 tanh x where (1) comes from tanh(2x) = 1+tanh 2 . x
K As the inverse of ExpK 0 (v), Log0 (x), therefore, shares the expression with its counterpart under the Poincaré ball model. Here, we use the properties of isometry to further validate this result:
P LogK 0 (x) = ϕ∗,0 Log0 (ψ(x)) √ −K∥x∥ √ x q = 2 tanh−1 . −K∥x∥ 1 + 1 + K ∥x∥2
(B.143)
We only need to show 2 tanh−1 We denote a =
! √ √ −K∥x∥ p = tanh−1 ( −K∥x∥). 1 + 1 + K∥x∥2
√ −K∥x∥ < 1. Both sides are
(B.144)
! √ √a 1 + a 1 + 1 − a2 + a 2 1+ 1−a −1 √ √ LHS: 2 tanh = ln = ln , 1 − 1+√a1−a2 1 + 1 − a2 1 + 1 − a2 − a √ √ 1+a 1 − a2 −1 RHS: tanh (a) = ln √ = ln . 1−a 1−a (B.145)
323
B.3. Gyrogroup Batch Normalization Note that the below equation holds: √ √ 1 + 1 − a2 + a 1 − a2 √ = . 1−a 1 + 1 − a2 − a
(B.146)
Geodesic distances. (1)
dK (x, y) = dP (ψ(x), ψ(y)) √ 2 −1 =√ tanh −K ∥−ψ(x) ⊕M ψ(y)∥ −K √ 2 (2) tanh−1 = √ −K ∥ψ (−x ⊕E y)∥ −K √ 2 ∥−x ⊕E y∥ , q tanh−1 −K =√ 2 −K 1 + 1 + K ∥−x ⊕E y∥
(B.147)
where (1) comes from the isometry, while (2) comes from the isomorphism. Exponential maps. (1) −1 ExpK ExpPψ(x) (ψ∗,x (v)) x (v) = ψ
λK ψ(x)
(2)
= ψ −1 ψ(x) ⊕M Exp0
(3)
= x ⊕E ψ −1 Exp0
2
λK ψ(x) 2
!!
ψ∗,x (v) !!
(B.148)
ψ∗,x (v)
= x ⊕E Exp0 λK ψ(x) ψ∗,x (v) .
(4)
The above comes from the following. (1) The isometry of ψ; (2) ExpPx (v) = x ⊕M Exp0
K λx v ; 2
(3) The isomorphism of ψ; (4) ψ −1 ◦ Exp0 (v) = ψ −1 ◦ ExpP0 (v)
= ψ −1 ◦ ψ ◦ ExpK 0 (2v) = Exp0 (2v) .
324
(B.149)
Appendix B. Proofs It remains to calculate λK ψ(x) ψ∗,x (v): λK ψ(x) =
2 2
1 + K √ ∥x∥
1+K∥x∥2
1+
K λK ψ(x) ψ∗,x (v) = λψ(x) =p
1
2
2 p 2 1 + 1 + K∥x∥2 = 2 p 2 1 + 1 + K∥x∥ + K∥x∥2 2 p 2 2 1 + 1 + K∥x∥ p = 2 + 2 1 + K∥x∥2 + 2K∥x∥2 2 p 1 + 1 + K∥x∥2 p = 1 + 1 + K∥x∥2 + K∥x∥2 p 1 + 1 + K∥x∥2 p , = 1 + K∥x∥2
(B.150)
1 K ⟨x, v⟩ p v− x 2 p p 2 1 + 1 + K∥x∥ 1 + 1 + K∥x∥2 1 + K∥x∥2
1 + K∥x∥2
v−
Logarithmic maps.
K ⟨x, v⟩ x. p 1 + 1 + K∥x∥2 (1 + K∥x∥2 )
(1) P LogK x (y) = ϕ∗,ψ(x) Logψ(x) ψ(y) (2)
= ϕ∗,ψ(x)
(3)
=
2
2
λK ψ(x)
!
Log0 (−ψ(x) ⊕M ψ(y))
ϕ∗,ψ(x) (Log0 (−ψ(x) ⊕M ψ(y)))
λK ψ(x) 1 (4) = K ϕ∗,ψ(x) (Log0 (−x ⊕E y)) . λψ(x) The above comes from the following. (1) The isometry of ψ; (2) LogPx (y) = λ2K Log0 (−x ⊕M y); x
(3) The linearity of differential maps; 325
(B.151)
(B.152)
B.4. SPD Multinomial Logistic Regression (4) Log0 (−ψ(x) ⊕M ψ(y)) = LogP0 (−ψ(x) ⊕M ψ(y)) 1 −1 = LogK ψ (−ψ(x) ⊕ ψ(y)) M 0 2 1 = LogK 0 (−x ⊕E y) . 2
(B.153)
Remark 192. Mao et al. [147, Thm. 9] extended the Möbius matrix-vector multiplication [76, Lem. 6] to the Einstein space of the Beltrami–Klein model under K = −1, namely Exp0 (M Log0 (x))). Although their presented formulations are different from the Möbius one, this theorem indicates that matrix-vector multiplications over these two spaces are identical under any negative curvature. The equality can also be readily observed from Eqs. (B.145) and (B.146).
B.4
SPD Multinomial Logistic Regression
B.4.1
Proof of Thm. 104
This claim can be proven either by definition [197, Def. 9.1] or by the constant rank level set theorem [197, Thm. 11.2]. We focus on the latter.
n n \{0}. Define the function Proof. Consider any P ∈ S++ and A ∈ TP S++ n f : S++ → R,
S 7→ ⟨LogP S, A⟩P .
(B.154)
n For the SPD hyperplane H̃A,P , we have H̃A,P = f −1 (0). By assumption, LogP : S++ → n TP S++ is a global diffeomorphism, and f is therefore well-defined. We can rewrite f as a composition, i.e., f = h ◦ LogP , where h(·) = ⟨·, A⟩P is a linear map.
Since LogP is a diffeomorphism and h(·) is a nonzero linear map, the rank of f is globally constant. Therefore, there exists a neighborhood, e.g., the whole SPD manifold, of f −1 (0) where the rank of f is constant. According to the constant rank level set theorem [197, Thm. 11.2], we can obtain the claim. 326
Appendix B. Proofs
B.4.2
Proof of Thm. 105
Proof. By Thm. 149, we have D E ⟨LogP Q, A⟩P = ϕ∗,P ϕ−1 (ϕ(Q) − ϕ(P )) , ϕ A ∗,P ∗,ϕ(P ) = ⟨ϕ(Q) − ϕ(P ), ϕ∗,P A⟩ .
(B.155) (B.156)
Therefore, due to the isometry of ϕ, the SPD hyperplane H̃Ak ,Pk corresponds to the Euclidean hyperplane Hϕ∗,Pk (Ak ),ϕ(Pk ) . (B.157) Furthermore, the distance to the margin hyperplane is equivalent to inf ∥ϕ(S) − ϕ(Q)∥F
(B.158)
s.t. ⟨ϕ(Q) − ϕ(Pk ), ϕ∗,Pk Ak ⟩ = 0.
(B.159)
ϕ(Q)
The problem above is the familiar Euclidean distance from a point to a hyperplane. By simple computation, one can obtain the result.
B.4.3
Proof of Thm. 106
Proof. For simplicity, we abbreviate ⊙ϕ and g ϕ as ⊙ and g. By abuse of notation, we further denote Q ⊙ P⊙−1 as QP −1 , where P⊙−1 is the inverse of P under ⊙. According n to Thm. 149, (S++ , ⊙) is an abelian group and g is a bi-invariant Riemannian metric. By Lin [137, Lem. 6], any parallel transport can be expressed as a differential of left translation, n PTP →Q = (LQP −1 )∗,P , ∀P, Q ∈ S++ . (B.160)
B.4.4
Proof of Thm. 107
Proof. By Eq. (6.9), parallel transport under a pullback Euclidean metric is pathindependent and satisfies PTQ→P (V ) = ϕ−1 ∗,ϕ(P ) (ϕ∗,Q (V )) . 327
(B.161)
B.4. SPD Multinomial Logistic Regression n For any Ã1,k ∈ TQ1 S++ , define
Ã2,k = ϕ−1 ∗,ϕ(Q2 )
n ϕ∗,Q1 Ã1,k ∈ TQ2 S++ .
(B.162)
Since ϕ is a diffeomorphism, its differential at every point is a linear isomorphism. Hence, −1 PTQ2 →Pk Ã2,k = ϕ∗,ϕ(Pk ) ϕ∗,Q2 Ã2,k = ϕ−1 ϕ Ã (B.163) ∗,Q 1,k 1 ∗,ϕ(Pk ) = PTQ1 →Pk Ã1,k .
n satisfies the same equality, applying ϕ∗,Pk to both sides Moreover, if B̃2,k ∈ TQ2 S++ gives ϕ∗,Q2 B̃2,k = ϕ∗,Q1 Ã1,k . (B.164)
The invertibility of ϕ∗,Q2 then yields B̃2,k = Ã2,k , proving uniqueness.
B.4.5
Proof of Thm. 108
Proof. Ak = PTI→Pk (Ãk ) = ϕ−1 ϕ ( Ã ) . ∗,I k ∗,ϕ(Pk )
(B.165) (B.166)
One can obtain the result by putting Eq. (B.166) into Eq. (4.10).
B.4.6
Proof of Thm. 109
n n Proof. Denoting the matrix power as Pθ : S++ → S++ , we have
(B.167)
Pθ (I) = I, (Pθ )∗,I (A) = θA,
n ∀A ∈ TI S++ .
(B.168)
Next, we prove the two cases separately. (α, β)-LEM. We define the following map: ψ LEM = f ◦ log,
(B.169)
where f : S n → S n is the linear isometry between the standard Frobenius inner product 328
Appendix B. Proofs and the O(n)-invariant inner product ⟨·, ·⟩(α,β) . Then ψ LEM pulls back the standard n Euclidean metric on S n to (α, β)-LEM on S++ . Putting Eqs. (B.168) and (B.169) into Eq. (4.12), we have D
LEM
LEM
E
LEM (Pk ), ψ∗,I (Ãk )
ψ (S) − ψ hD Ei = exp f (log(S) − log(Pk )) , f (Ãk ) D E(α,β) = exp log(S) − log(Pk ), Ãk ,
exp
(B.170)
where the last equality follows from the linearity and isometry property of f , with f = f∗ . θ-LCM. We denote ψ LCM = Dlog ◦ Chol ◦ Pθ .
(B.171)
Then ψ LCM pulls back the Euclidean metric θ12 g E on the Euclidean space LTn of lower n triangular matrices to θ-LCM on S++ . The differential of Cholesky decomposition is presented by Lin [137, Prop. 4], while the differential of Dlog is obtained as the naturallogarithm specialization of Thm. 154. Then, simple computations show that LCM ψ∗,I (A) = θ
1 ⌊A⌋ + D(A) , 2
n ∀A ∈ TI S++ .
(B.172)
Putting Eqs. (B.171) and (B.172) into Eq. (4.12), we can obtain the result.
B.4.7
Proof of Thm. 111
To prove Thm. 111, we first present two lemmas about the general cases under pullback Euclidean metrics. One can observe that Eqs. (4.12) and (4.13) are very similar to those of a Euclidean MLR. However, since ϕ is normally nonlinear and Pk is an SPD parameter, Eq. (4.12) cannot hastily be identified with a Euclidean MLR. Under some special circumstances, SPD MLR can be reduced to the familiar Euclidean MLR. To show this result, we first specialize the RSGD strategy reviewed in Sec. 2.7 to pullback Euclidean metrics. In ambient-gradient notation, the required update is Wt+1 = ExpWt (−γt ΠWt (∇Wt f )) ,
(B.173)
where ΠWt denotes the projection mapping the Euclidean gradient ∇Wt f to the Rie329
B.4. SPD Multinomial Logistic Regression mannian gradient, and γt denotes the learning rate. We have already obtained the formula for the Riemannian exponential map in Eq. (6.7). We proceed to formulate Π. n n Lemma 193. For a smooth function f : S++ → R on S++ endowed with any kind n n n of pullback Euclidean metric, the projection map ΠP : S → TP S++ at P ∈ S++ is
−∗ ΠP (∇P f ) = ϕ−1 ϕ (∇ f ) , P ∗,ϕ(P ) ∗,ϕ(P )
(B.174)
D
(B.175)
∗ −∗ −1 n where ϕ∗,ϕ(P ) := ϕ∗,ϕ(P ) : TP S++ → Tϕ(P ) S n is the Frobenius adjoint of ϕ−1 ∗,ϕ(P ) : n n n n Tϕ(P ) S → TP S++ . Specifically, for all U ∈ TP S++ and Z ∈ Tϕ(P ) S , it satisfies E D E −∗ U, ϕ−1 (Z) = ϕ (U ), Z . ∗,ϕ(P ) ∗,ϕ(P )
n Proof. Given any smooth function f : S++ → R, denote its Riemannian gradient at P n as gradP f ∈ TP S++ .
⟨gradP f, V ⟩P = f∗,P (V ),
(B.176)
n ∀V ∈ TP S++ .
Let ∇P f denote the Euclidean gradient in the canonical chart. For any Z ∈ Tϕ(P ) S n , set V = ϕ−1 ∗,ϕ(P ) (Z). By Eq. (6.5) and the definition of the Frobenius adjoint, we have ⟨ϕ∗,P (gradP f ) , Z⟩ =
D
E
gradP f, ϕ−1 ∗,ϕ(P ) (Z)
ϕ−1 ∗,ϕ(P ) (Z)
P
= f∗,P D E −1 = ∇P f, ϕ∗,ϕ(P ) (Z) D E = ϕ−∗ (∇ f ) , Z . P ∗,ϕ(P )
(B.177)
Since Z is arbitrary, the nondegeneracy of the Frobenius inner product gives ϕ∗,P (gradP f ) = ϕ−∗ ∗,ϕ(P ) (∇P f ) .
(B.178)
Applying ϕ−1 ∗,ϕ(P ) to both sides yields −∗ gradP f = ϕ−1 ϕ (∇ f ) . P ∗,ϕ(P ) ∗,ϕ(P )
By definition, ΠP (∇P f ) = gradP f , which yields Eq. (B.174). We can describe the special case mentioned above with this lemma. 330
(B.179)
Appendix B. Proofs Lemma 194. Suppose the differential map ϕ∗,I is the identity map, and Pk in Eq. (4.12) is optimized by pullback-Euclidean-metric-based RSGD. Then Eq. (4.12) can be reduced to a Euclidean MLR in the codomain of ϕ updated by Euclidean SGD. Proof. Define a Euclidean MLR in the codomain of ϕ as p(y = k | S) ∝ exp
ϕ(S) − P̄k , Āk
where P̄k , Āk ∈ S n . We call this classifier ϕ-EMLR.
(B.180)
,
Define the SPD MLR under the pullback Euclidean metric induced by ϕ as p(y = k | S) ∝ exp
D
ϕ(S) − ϕ(Pk ), Ãk
E
,
(B.181)
n where Pk ∈ S++ and Ãk ∈ S n .
Suppose the SPD MLR and ϕ-EMLR satisfy P̄k = ϕ(Pk ). Other settings of the network are all the same, indicating that the Euclidean gradients satisfy ∂L ∂L = . ∂ϕ(Pk ) ∂ P̄k
(B.182)
The update of P̄k in the ϕ-EMLR is P̄k′ = P̄k − γ
∂L . ∂ P̄k
(B.183)
The update of Pk in the SPD MLR is ∂L −γΠPk ∂Pk −∗ ∂L −1 =ϕ ϕ(Pk ) − γϕ∗,Pk . ∂Pk
Pk′ = ExpPk
(B.184) (B.185)
Therefore, ϕ(Pk′ ) satisfies ϕ(Pk′ ) = ϕ(Pk ) − γϕ−∗ ∗,Pk
∂L ∂Pk
∗ = ϕ(Pk ) − γϕ−∗ ∗,Pk ϕ∗,Pk
= ϕ(Pk ) − γ
∂L ∂ϕ(Pk )
331
(B.186) ∂L ∂ϕ(Pk )
(B.187) (B.188)
B.5. Riemannian Multinomial Logistic Regression (B.189)
= P̄k′ .
Eq. (B.187) comes from the Euclidean chain rule of the differential. Let Y = ϕ(X). Then we have ∂L ∂L ∂L : dY = : ϕ∗,X d X = ϕ∗∗,X : d X, (B.190) ∂Y ∂Y ∂Y where : denotes the Frobenius inner product. The equivalence of Āk and Ãk is obvious. By mathematical induction, the claim can be proven. Proof. The claim follows directly from Thm. 194.
B.5
Riemannian Multinomial Logistic Regression
B.5.1
Proof of Thm. 113
Proof. Let us first solve Y ∗ in Eq. (4.22), which is the solution to the following constrained optimization problem: max Y
⟨LogP Y, LogP S⟩P ∥ LogP Y ∥P ∥ LogP S∥P
s.t.⟨LogP Y, Ã⟩P = 0.
(B.191)
Note that Eq. (B.191) is well-defined due to the existence of the Riemannian logarithm. Although Eq. (B.191) is normally non-convex, Eq. (B.191) and Eq. (4.22) can be reduced to a Euclidean problem: ⟨Ỹ , S̃⟩P ∥Ỹ ∥P ∥S̃∥P
s.t.⟨Ỹ , Ã⟩P = 0,
(B.192)
d(S, H̃Ã,P ) = sin(∠SP Y ∗ )∥S̃∥P ,
(B.193)
max Ỹ
where Ỹ = LogP Y and S̃ = LogP S. Let us first discuss Eq. (B.192). Denote the solution of Eq. (B.192) as Ỹ ∗ . Note that Ỹ ∗ is not necessarily unique. Note that ExpP is only well-defined locally. More precisely, ExpP is well-defined in an open ball Bϵ (0) centered at 0 ∈ TP M. Therefore, Ỹ ∗ might not be in Bϵ (0). In this case, we can scale Ỹ ∗ into Bϵ (0), and the scaled Ỹ ∗ is still the maximizer of Eq. (B.192). Therefore, without loss of generality, we assume Ỹ ∗ ∈ Bϵ (0). Putting Ỹ ∗ into Eq. (B.193), Eq. (B.193) is reduced to the distance to the hyperplane 332
Appendix B. Proofs ⟨Ỹ , Ã⟩P = 0 in the Euclidean space (TP M, ⟨·, ·⟩P ), which has a closed-form solution: |⟨S̃, Ã⟩P | ∥Ã∥P |⟨LogP S, Ã⟩P | = . ∥Ã∥P
(B.194)
d(S, H̃Ã,P ) =
B.5.2
(B.195)
Proof of Thm. 114
Proof. Putting the margin distance (Eq. (4.25)) into Eq. (4.18), we have the following: p(y = k | S) ∝ exp sign(⟨Ãk , LogPk (S)⟩Pk )∥Ãk ∥Pk d(S, H̃Ãk ,Pk ) = exp sign(⟨Ãk , LogPk (S)⟩Pk )∥Ãk ∥Pk
|⟨LogPk (S), Ãk ⟩Pk |
∥Ãk ∥Pk
!
(B.196)
= exp ⟨LogPk S, Ãk ⟩Pk .
B.5.3
Proof of Thm. 116
Proof. The Riemannian metric (α, β)-EM at I is (α,β)-EM
gI
(V, V ) = ⟨V, V ⟩(α,β) .
(B.197)
By Thm. 186, we have the following: (θ,α,β)-EM
gP
θ→0
(α,β)-EM
(V, V ) −−→ gI
log∗,P (V ), log∗,P (V )
= ⟨log∗,P (V ), log∗,P (V )⟩(α,β) (α,β)-LEM
= gP
B.5.4
(B.198)
(V, V ) .
Proof of Thm. 117
As the five families of metrics presented in Thm. 117 are pullback metrics, we first present a general result regarding Riemannian MLRs under pullback metrics. 333
B.5. Riemannian Multinomial Logistic Regression Lemma 195 (Riemannian MLRs under pullback metrics). Suppose (N , g) is a Riemannian manifold and ϕ : M → N is a diffeomorphism between manifolds. The Riemannian MLR by parallel transport, obtained by combining Eqs. (4.26) and (4.27), on M under g̃ = ϕ∗ g can be obtained using g: i h ˜ p(y = k | S ∈ M) ∝ exp ⟨LogPk S, Γ̃Q→Pk Ak ⟩Pk h i = exp ⟨Logϕ(Pk ) ϕ(S), Ãk ⟩ϕ(Pk ) ,
(B.199) (B.200)
˜ and Γ̃ are the Riemannian where Ãk = Γϕ(Q)→ϕ(Pk ) ϕ∗,Q (Ak ) with Ak ∈ TQ M; Log logarithm and parallel transport under g̃, respectively; and Log and Γ are their counterparts under g. Furthermore, if N has a Lie group operation ⊙, M could be endowed with a Lie ˜ by ϕ. The Riemannian MLR by left translation, obtained by group structure ⊙ ˜ can be calculated using g combining Eqs. (4.26) and (4.28), on M under g̃ and ⊙ and ⊙: " # ˜ P S, L̃ Ak (B.201) p(y = k | S ∈ M) ∝ exp Log k
R̃k
∗,Q
h
i
Pk
= exp ⟨Logϕ(Pk ) ϕ(S), Ãk ⟩ϕ(Pk ) ,
(B.202)
−1 ˜ −1 where Ãk = (LRk )∗,ϕ(Q) (ϕ∗,Q (Ak )), R̃k = Pk ⊙Q ˜ , Rk = ϕ(Pk ) ⊙ ϕ(Q)⊙ , and ⊙ ˜ L̃Pk ⊙Q ˜ −1 is the left translation under ⊙. ˜ ⊙
˜ and g̃ Proof. Before starting, we should point out that since ϕ is a diffeomorphism, ⊙ ˜ forms a are indeed well defined, and (M, g̃) forms a Riemannian manifold and (M, ⊙) Lie group. We first focus on the Riemannian MLR by parallel transport: p(y = k | S ∈ M) ˜ ∝ exp g̃Pk (LogPk S, Γ̃Q→Pk Ak ) h i −1 = exp gϕ(Pk ) ϕ∗,Pk ◦ ϕ−1 Log ϕ(S), ϕ ◦ ϕ Γ ϕ (A ) ∗,Pk k ϕ(Pk ) ∗,ϕ(Pk ) ∗,ϕ(Pk ) ϕ(Q)→ϕ(Pk ) ∗,Q = exp gϕ(Pk ) (Logϕ(Pk ) ϕ(S), Γϕ(Q)→ϕ(Pk ) ϕ∗,Q (Ak )) . (B.203)
In the case of the Riemannian MLR by left translation, we first note that L̃R̃k = ϕ−1 ◦ Lϕ(Pk )⊙ϕ(Q)−1 ◦ ϕ. ⊙ 334
(B.204)
Appendix B. Proofs Therefore, the associated differential is
L̃R̃k
∗,Q
= ϕ−1 ∗,ϕ(Pk ) ◦ (LRk )∗,ϕ(Q) ◦ ϕ∗,Q .
(B.205)
Putting Eq. (B.205) into Eq. (B.201), we can obtain the result. Now, we apply Thm. 195 to derive the expressions for our SPD MLRs presented in Thm. 117. In our cases of SPD MLRs, we set Q = I. For simplicity, we will omit the subscript k for Pk and Ak . We will first derive the expressions for SPD MLRs under (θ, α, β)-LEM, θ-LCM, (θ, α, β)-EM, and (θ, α, β)-AIM from Eq. (B.200). Then we will derive the expression for MLR under 2θ-BWM from Eq. (B.202). According to Thm. 187, the scaled metric ag shares the same Riemannian operators as g. We will use this fact throughout the following proof. Proof. For simplicity, we abbreviate ϕθ as ϕ during the proof. Note that for 2θ-BWM, ϕ should be understood as ϕ2θ . We first show ϕ(I) and the differential map ϕ∗,I , which will be frequently required in the following proof: (B.206)
ϕ(I) = I, n ϕ∗,I (A) = θA, ∀A ∈ TI S++ .
(B.207)
n n Let ϕ : (S++ , g̃) → (S++ , g). Then the SPD MLR under g̃ by parallel transport with Q = I is
n p(y = k | S ∈ S++ ) ∝ exp gϕ(P ) (Logϕ(P ) ϕ(S), ΓI→ϕ(P ) θA) .
(B.208)
Next, we begin to prove the five SPD MLRs one by one.
(α, β)-LEM. As shown in Thm. 147, the standard LEM is the pullback metric from the Euclidean space S n . Similarly, (α, β)-LEM is also a pullback metric: log
n (S++ , g (α,β)-LEM ) −→ (S n , g (α,β) ).
(B.209)
By Eq. (B.200), we have n p(y = k | S ∈ S++ ) ∝ exp ⟨log(S) − log(P ), log∗,I (A)⟩(α,β) = exp ⟨log(S) − log(P ), A⟩(α,β) .
(B.210) (B.211)
θ-LCM. Simple computations show that θ-LCM is the scaled pullback metric of the 335
B.5. Riemannian Multinomial Logistic Regression standard Euclidean metric in the Euclidean space of lower triangular matrices LTn : ϕ
Chol
Dlog
n n (S++ , θ2 g θ-LCM ) −→ (S++ , g LCM ) −→ (Ln++ , g CM ) −→ (LTn , g E ),
(B.212)
where g E is the standard Frobenius inner product, and g CM is the Cholesky metric on the Cholesky space Ln++ [137]. Denoting ζ = Dlog ◦ Chol ◦ϕ, we have 1 n . ζ∗,I (A) = θ ⌊A⌋ + D(A) , ∀A ∈ TI S++ 2
(B.213)
Similar to the case of (θ, α, β)-LEM, we have
1 ⟨ζ(S) − ζ(P ), ζ∗,I A⟩ (B.214) θ2 h i * + ⌊ K̃⌋ − ⌊ L̃⌋ + Dlog(D( K̃)) − Dlog(D( L̃)) , 1 = exp , 1 θ ⌊A⌋ + D(A) 2 (B.215)
n ) ∝ exp p(y = k | S ∈ S++
where K̃ = Chol(S θ ), L̃ = Chol(P θ ), D(K̃) is a diagonal matrix with diagonal elements from K̃, and ⌊K̃⌋ is a strictly lower triangular matrix from K̃.
1 (θ, α, β)-EM. Let η = |θ| ϕ. A simple computation shows that (θ, α, β)-EM is the pullback metric of (α, β)-EM: η
(B.216)
n η∗,I (A) = sgn(θ)A, ∀A ∈ TI S++ .
(B.217)
n n (S++ , g (θ,α,β)-EM ) −→ (S++ , g (α,β)-EM ).
Besides, we have the following for η:
According to Eq. (B.200), we have n p(y = k | S ∈ S++ ) ∝ exp ⟨η(S) − η(P ), sgn(θ)A⟩(α,β) 1 θ θ (α,β) . = exp ⟨S − P , A⟩ θ 336
(B.218) (B.219)
Appendix B. Proofs (θ, α, β)-AIM. Putting g (α,β)-AIM into Eq. (B.208), we have θ θ θ θ θ 1 (α,β)-AIM θ − θ − g P 2 log P 2 S P 2 P 2 , P 2 θAP 2 θ2 ϕ(P ) (B.220) D E (α,β) θ θ 1 = exp . (B.221) log P − 2 S θ P − 2 , A θ
n ) ∝ exp p(y = k | S ∈ S++
2θ-BWM. We first simplify Eq. (B.202) for SPD manifolds and then proceed to n n ˜ → (S++ focus on the case of g = g BWM . Denote ϕ : (S++ , g̃, ⊙) , g, ⊙), where the Lie group operation ⊙ [195] is defined as n S1 ⊙ S2 = L1 S2 L⊤ 1 , ∀S1 , S2 ∈ S++ , with L1 = Chol(S1 ).
(B.222)
n n Note that I is the identity element of (S++ , ⊙), and for any S ∈ S++ , the differential map of the left translation LS under ⊙ is n n (LS )∗,Q (V ) = LV L⊤ , ∀Q ∈ S++ , ∀V ∈ TQ S++ , with L = Chol(S).
(B.223)
n ˜ is ˜ the left translation L̃P ⊙I For the induced Lie group (S++ , ⊙), ˜ −1 under ⊙ ˜ ⊙
−1 L̃P ⊙I ◦ Lϕ(P )⊙ϕ(I)−1 ◦ ϕ, ˜ −1 = ϕ ⊙
(B.224)
= ϕ−1 ◦ LP 2θ ◦ ϕ,
(B.225)
˜ ⊙
2θ ϕ(P ) ⊙ ϕ(I)−1 ⊙ = P .
The associated differential at I is
L̃
˜ −1 P ⊙I ˜ ⊙
∗,I
(A) = ϕ−1 ∗,ϕ(P ) ◦ (LP 2θ )∗,ϕ(I) ◦ ϕ∗,I (A) ⊤ = 2θϕ−1 ∗,ϕ(P ) (L̄AL̄ ),
(B.226) (B.227)
˜ by left translation is where L̄ = Chol(P 2θ ). Then the SPD MLR under g̃ and ⊙ n p(y = k | S ∈ S++ ) ∝ exp 2θgϕ(P ) Logϕ(P ) ϕ(S), L̄AL̄⊤ .
(B.228)
Setting g = g BWM (we omit the scaling factor), we obtain the SPD MLR under 2θ-BWM: 1 BWM BWM n ⊤ p(y = k | S ∈ S++ ) ∝ exp 2θ · 2 gϕ(P ) Logϕ(P ) ϕ(S), L̄AL̄ (B.229) 4θ 337
B.6. Proper Velocity Neural Networks
1 2θ 2θ 12 2θ 2θ 12 2θ ⊤ = exp ⟨(P S ) + (S P ) − 2P , LP 2θ [L̄AL̄ ]⟩ . 4θ (B.230)
B.5.5
Proof of Thm. 119
Proof. During this proof, we use the ambient representation of tangent vectors. Given rotation matrices P, Q ∈ SO(n) and a tangent vector H ∈ TQ SO(n), let c(t) be a curve on SO(n) satisfying c(0) = Q and c′ (0) = H. The differential of the left translation LP Q−1 at Q is (LP Q−1 )∗,Q (H) =
dP Q−1 c(t) = P Q−1 H = P Q⊤ H, dt t=0
(B.231)
which is precisely the vector transport TQ→P (H) in Boumal and Absil [28, Tab. 1].
B.5.6
Proof of Thm. 120
Proof. Setting Q = I in Eq. (4.28) and using Thm. 119 give Ãk = Pk Ak . Combining this identity with the logarithmic map and metric in Tab. 2.12, we obtain D
LogPk R, Ãk
E
Pk
= Pk log Pk⊤ R , Pk Ak = log Pk⊤ R , Ak .
(B.232)
Substituting this expression into Eq. (4.26) yields the result.
B.6
Proper Velocity Neural Networks
B.6.1
Derivation of the Proper Velocity Metric
The PV line element at x ∈ PVnK can be written in terms of the curvature parameter K < 0 as Qx (u) = ∥u∥2 + Kβx2 ⟨x, u⟩2 , ∀u ∈ Tx PVnK ≃ Rn , (B.233) where βx = √
1
1−K∥x∥2 2
. This is equivalent to the expression in Ungar [200, Eq. (7.76)]
after substituting s = −1/K. Given u, v ∈ Tx PVnK , the bilinear form gx (u, v) is 338
Appendix B. Proofs obtained by the polarization identity: gx (u, v) = 14 (Qx (u + v) − Qx (u − v)) .
(B.234)
We first expand the two terms in the polarization identity: Qx (u + v) = ∥u + v∥2 + Kβx2 ⟨x, u + v⟩2
= ∥u∥2 + 2 ⟨u, v⟩ + ∥v∥2 + Kβx2 ⟨x, u⟩2 + 2 ⟨x, u⟩ ⟨x, v⟩ + ⟨x, v⟩2 ,
Qx (u − v) = ∥u − v∥2 + Kβx2 ⟨x, u − v⟩2
(B.235)
= ∥u∥2 − 2 ⟨u, v⟩ + ∥v∥2 + Kβx2 ⟨x, u⟩2 − 2 ⟨x, u⟩ ⟨x, v⟩ + ⟨x, v⟩2 .
Taking the difference yields
Qx (u + v) − Qx (u − v) = 4 ⟨u, v⟩ + 4Kβx2 ⟨x, u⟩ ⟨x, v⟩ .
(B.236)
Substituting this expression into the polarization identity, we obtain gx (u, v) = 41 (Qx (u + v) − Qx (u − v)) = ⟨u, v⟩ + Kβx2 ⟨x, u⟩ ⟨x, v⟩ ,
(B.237)
which coincides with the expression of the PV metric in Eq. (5.1).
B.6.2
Proof of Thm. 121
Proof. Differential of πPVnK →PnK . Consider the curve c : (−ε, ε) → PVnK which satisfies c(0) = x and c′ (0) = v. By definition of the differential, dx πPVnK →PnK (v) = βx Using πPVnK →PnK (x) = 1+β x with βx = √ x
d πPVnK →PnK (c(t)) . dt t=0 1 1−K∥x∥2
πPVnK →PnK (c(t)) = h(t)c(t),
(B.238)
, we write
h(t) :=
βc(t) . 1 + βc(t)
(B.239)
Let r(t) = ∥c(t)∥2 , so that βc(t) = (1 − Kr(t))−1/2 . Then r′ (0) = 2 ⟨x, v⟩ ,
′ βc(0) = 21 (1 − Kr(0))−3/2 Kr′ (0) = Kβx3 ⟨x, v⟩ .
339
(B.240)
B.6. Proper Velocity Neural Networks Differentiating h(t) at t = 0 gives ′
h (0) =
′ βc(0)
(1 + βx )
=K 2
βx3 ⟨x, v⟩ . (1 + βx )2
(B.241)
Finally, differentiating h(t)c(t) at t = 0 yields dx πPVnK →PnK (v) = h′ (0)x + h(0)v
(B.242)
βx3 βx v. =K ⟨x, v⟩ x + (1 + βx )2 1 + βx In particular, at x = 0 one has β0 = 1 and ⟨x, v⟩ = 0. Thus, we have d0 πPVnK →PnK (v) =
β0 1 v = v. 1 + β0 2
(B.243)
Differential of πPnK →PVnK . Consider the curve c : (−ε, ε) → PnK which satisfies c(0) = y and c′ (0) = w. By definition of the differential, dy πPnK →PVnK (w) =
d πPn →PVnK (c(t)). dt t=0 K
Using the explicit expression πPnK →PVnK (y) = 2γy2 y with γy = √
(B.244) 1
1+K∥y∥2
, we obtain
2 πPnK →PVnK (c(t)) = 2γc(t) c(t).
(B.245)
2 Let r(t) = ∥c(t)∥2 so that γc(t) = (1 + Kr(t))−1 . Then
r′ (0) = 2 ⟨y, w⟩ ,
Kr′ (0) d 2 γc(t) =− = −2Kγy4 ⟨y, w⟩ . dt t=0 (1 + Kr(0))2
(B.246)
2 c(t) at t = 0 yields Differentiating 2γc(t)
dy πPnK →PVnK (w) = 2
d γ 2 y + 2γy2 w dt t=0 c(t)
(B.247)
= −4Kγy4 ⟨y, w⟩ y + 2γy2 w. In particular, at y = 0 we have γ0 = 1 and ⟨y, w⟩ = 0. Thus, we have d0 πPnK →PVnK (w) = 2w. 340
(B.248)
Appendix B. Proofs
B.6.3
Proof of Thm. 122
Proof. It suffices to show that for any x ∈ PVnK and v, w ∈ Tx PVnK , gyP dx πPVnK →PnK (v), dx πPVnK →PnK (w) = gxPV (v, w),
(B.249)
where y = πPVnK →PnK (x).
We first recall the following equations from Eq. (5.1), Sec. 2.9.5, and Thm. 121: gxPV (v, w) = ⟨v, w⟩ + Kβx2 ⟨x, v⟩ ⟨x, w⟩ , ∀x ∈ PVnK , ∀v, w ∈ Tx PVnK , 2 gyP (u, z) = λK ⟨u, z⟩ , ∀y ∈ PnK , ∀u, z ∈ Ty PnK , y βx (B.250) πPVnK →PnK (x) = x, ∀x ∈ PVnK , 1 + βx βx βx3 dx πPVnK →PnK (v) = ⟨x, v⟩ x, ∀x ∈ PVnK , ∀v ∈ Tx PVnK . v+K 1 + βx (1 + βx )2 Let a=
βx , 1 + βx
b=K
βx3 . (1 + βx )2
(B.251)
Then dx πPVnK →PnK (v) = av + b ⟨x, v⟩ x and dx πPVnK →PnK (w) = aw + b ⟨x, w⟩ x. Thus, gyP dx πPVnK →PnK (v), dx πPVnK →PnK (w) 2 = λK ⟨av + b ⟨x, v⟩ x, aw + b ⟨x, w⟩ x⟩ y (B.252) 2 2 2 = λK a ⟨v, w⟩ + ab ⟨x, w⟩ ⟨v, x⟩ + ab ⟨x, v⟩ ⟨x, w⟩ + b ⟨x, v⟩ ⟨x, w⟩ ⟨x, x⟩ y 2 2 = λK a ⟨v, w⟩ + 2ab + b2 ∥x∥2 ⟨x, v⟩ ⟨x, w⟩ . y
Using y = πPVnK →PnK (x) and the relation between λK y , βx , and ∥x∥ from Sec. 2.9.5, we simplify the coefficients. First,
βx βx K ∥y∥ = K x, x 1 + βx 1 + βx βx2 =K ∥x∥2 (1 + βx )2 β2 − 1 = x , (1 + βx )2 2
341
(B.253)
B.6. Proper Velocity Neural Networks 1 where we use βx2 = 1−K∥x∥ 2 . Hence
1 + K ∥y∥2 = 1 +
βx2 − 1 2βx , = 2 (1 + βx ) 1 + βx
(B.254)
1+βx 2 which implies λK y = 1+K∥y∥2 = βx . Therefore
2 2 λK a = y
Next, we compute
1 + βx βx
2
βx 1 + βx
2
= 1.
βx βx3 2Kβx4 K = , 1 + βx (1 + βx )2 (1 + βx )3 βx4 (βx2 − 1) βx6 2 2 2 2 ∥x∥ = K , b ∥x∥ = K (1 + βx )4 (1 + βx )4
(B.255)
2ab = 2
(B.256)
which brings us to βx4 2 2(1 + β ) + β − 1 x x (1 + βx )4 βx4 βx4 (βx + 1)2 =K . =K (1 + βx )4 (1 + βx )2
2ab + b2 ∥x∥2 = K
Multiplying by λK y
2
2 λK y
=
1+βx βx
2
2
(B.257)
yields
2ab + b ∥x∥
2
=
1 + βx βx
2
K
βx4 = Kβx2 . (1 + βx )2
(B.258)
Substituting these identities into the expression for gP gives gyP dx πPVnK →PnK (v), dx πPVnK →PnK (w) = ⟨v, w⟩ + Kβx2 ⟨x, v⟩ ⟨x, w⟩ = gxPV (v, w). (B.259)
B.6.4
Proof of Thm. 123
Following the notation in the main theorem, we further denote: x̄ = π(x) ∈ PnK ,
ȳ = π(y) ∈ PnK , 342
v̄ = dx π(v) ∈ Tx̄ PnK ,
(B.260)
Appendix B. Proofs Recalling Eq. (5.8) and Thm. 121, we have the following: π(x) =
βx x, 1 + βx
π −1 (ȳ) = 2γȳ2 ȳ, dx π(v) = K
(B.261)
∀x ∈ PVnK ,
(B.262)
∀ȳ ∈ PnK ,
βx3 βx v, ⟨x, v⟩ x + 2 (1 + βx ) 1 + βx
∀x ∈ PVnK , ∀v ∈ Tx PVnK ,
(B.264)
∀ȳ ∈ PnK , ∀w ∈ Tȳ PnK .
dȳ π −1 (w) = −4Kγȳ4 ⟨ȳ, w⟩ ȳ + 2γȳ2 w,
(B.263)
Next, we derive the expressions for each PV operator.
B.6.4.1
PV Exponential Map
We recall from Tab. 2.14 that the Riemannian exponential on the Poincaré ball is ExpPx̄ (v̄) = x̄ ⊕M
1 √ tanh −K
√
−KλK x̄ ∥v̄∥ 2
v̄ ∥v̄∥
.
(B.265)
By the Riemannian isometry and the gyrovector isomorphism of π, for any x ∈ PVnK and v ∈ Tx PVnK we have (1) Expx (v) = π −1 ExpPx̄ (v̄) √ 1 −KλK v̄ (2) x̄ ∥v̄∥ −1 √ = x ⊕U π tanh , 2 ∥v̄∥ −K
(B.266)
The above equalities follow from the following facts. (1) Isometry. (2) Gyrovector isomorphism. Let
We have
1 u= √ tanh −K
√ −KλK v̄ x̄ ∥v̄∥ . 2 ∥v̄∥
1 ∥u∥ = √ tanh −K
√
343
−KλK x̄ ∥v̄∥ 2
.
(B.267)
(B.268)
B.6. Proper Velocity Neural Networks Let t =
√
√ K ∥v̄∥ −Kλx̄ so that −K ∥u∥ = tanh(t). Then 2 2 u 1 + K ∥u∥2 1 2 v̄ √ = tanh(t) ∥v̄∥ −K 1 + K ∥u∥2 1 v̄ 2 √ tanh(t) = 2 ∥v̄∥ 1 − tanh (t) −K v̄ 2 tanh(t) =√ 2 −K 1 − tanh (t) ∥v̄∥ 2 v̄ (1) = √ tanh(t) cosh2 (t) ∥v̄∥ −K v̄ 2 sinh(t) cosh2 (t) =√ ∥v̄∥ −K cosh(t) 2 v̄ =√ sinh(t) cosh(t) ∥v̄∥ −K 1 v̄ (2) = √ (2 sinh(t) cosh(t)) ∥v̄∥ −K 1 v̄ =√ sinh (2t) ∥v̄∥ −K √ v̄ 1 sinh −KλK ∥v̄∥ =√ x̄ ∥v̄∥ −K √ −K(1 + βx ) v̄ 1 (3) sinh ∥v̄∥ . = √ βx ∥v̄∥ −K
π −1 (u) =
(B.269)
The above equalities use: (1) 1 − tanh2 (t) = 1/ cosh2 (t); (2) sinh(2t) = 2 sinh(t) cosh(t); 1+βx (3) λK x̄ = βx .
B.6.4.2
PV Logarithmic Map
We recall from Tab. 2.14 that the Riemannian logarithm on the Poincaré ball is LogPx̄ (ȳ) = √
√ tanh−1 −K ∥z̄∥ 2 z̄, ∥z̄∥ −KλK x̄ 344
z̄ = (−x̄) ⊕M ȳ,
(B.270)
Appendix B. Proofs 2 where λK x̄ = 1+K∥x̄∥2 . We define
z̄ = π(z).
(B.271)
! √ tanh−1 −K ∥z̄∥ 2 √ z̄ ∥z̄∥ −KλK x̄
(B.272)
z = (−x) ⊕U y, By the Riemannian isometry of π, we have Logx (y) = dx̄ π −1 LogPx̄ (ȳ) = dx̄ π −1
= α(x, y)dx̄ π −1 (z̄) , where
√ tanh−1 −K ∥z̄∥ 2 α(x, y) = √ . ∥z̄∥ −KλK x̄
(B.273)
The differential of π −1 at x̄ = π(x) is dx̄ π −1 (h) = −4Kγx̄4 ⟨x̄, h⟩ x̄ + 2γx̄2 h, where γx̄ = √
1
1+K∥x̄∥2
(B.274)
∀h ∈ Tx̄ PnK ,
βx x and the relation 1 − K ∥x∥2 = β12 , we have . Using x̄ = 1+β x x
βx2 βx2 2 K ∥x̄∥ = K ∥x∥ = (1 + βx )2 (1 + βx )2 β2 − 1 2βx 1 + K ∥x̄∥2 = 1 + x = . (1 + βx )2 1 + βx 2
βx2 − 1 βx2
=
2
=
(1 + βx )2 . 4βx2
βx2 − 1 , (1 + βx )2
(B.275)
Hence γx̄2 =
1 1 + βx , 2 = 2βx 1 + K ∥x̄∥
γx̄4 =
1 + βx 2βx
(B.276)
Substituting these into dx̄ π −1 (h) yields (1 + βx )2 1 + βx dx̄ π (h) = −4K ⟨x̄, h⟩ x̄ + 2 h 2 4βx 2βx (1 + βx )2 1 + βx = −K ⟨x̄, h⟩ x̄ + h 2 βx βx 1 + βx = h − K ⟨x, h⟩ x, βx −1
(B.277)
βx x. Applying Eq. (B.277) with h = z̄ and where the last equality uses that x̄ = 1+β x
345
B.6. Proper Velocity Neural Networks using that z̄ = π(z) is collinear with z = (−x) ⊕U y, we obtain Logx (y) = α(x, y)(dx π)−1 (z̄) 1 + βx = α(x, y) z̄ − K ⟨x, z̄⟩ x . βx
(B.278)
Since z̄ = π(z) and π is given by Eq. (5.8), z and z̄ are collinear and z̄ = ρz,
ρ=
βz , 1 + βz
(B.279)
which also implies ⟨x, z̄⟩ = ρ ⟨x, z⟩. Substituting these into Eq. (B.278) yields
1 + βx Logx (y) = α(x, y) ρz − Kρ ⟨x, z⟩ x βx 1 + βx ρ z + (−Kα(x, y)ρ) ⟨x, z⟩ x. = α(x, y) | {z } βx | {z } τ (x, y) σ(x, y)
(B.280)
1+βx βz Using the definition of α(x, y) in Eq. (B.273) together with λK x̄ = βx and ρ = 1+βz , a straightforward simplification yields
√ −K ∥z̄∥ 2 tanh−1 , σ(x, y) = √ ∥z∥ −K
2βx τ (x, y) = 1 + βx
√
√ −K tanh−1 −K ∥z̄∥ . ∥z∥ (B.281)
Thus, Logx (y) = σ(x, y)z + τ (x, y) ⟨x, z⟩ x.
B.6.4.3
(B.282)
PV Parallel Transport
We recall from Tab. 2.14 that the parallel transport on the Poincaré ball is PTPx̄→ȳ (w) =
λK x̄ gyrM [ȳ, −x̄](w), λK ȳ 346
with w ∈ Tx̄ PnK .
(B.283)
Appendix B. Proofs We have PTx→y (v) = dȳ π −1 PTPx̄→ȳ (dx π(v)) K λx̄ −1 = dȳ π gyrM [ȳ, −x̄] (dx π(v)) λK ȳ K (2) λx̄ = K dȳ π −1 (gyrM [ȳ, −x̄] (dx π(v))) λȳ (B.284) K 1 + βy (3) λx̄ gyrM [ȳ, −x̄] (dx π(v)) − K ⟨y, gyrM [ȳ, −x̄] (dx π(v))⟩ y = K βy λȳ 1 + βy (4) (1 + βx )βy = gyrM [ȳ, −x̄] (dx π(v)) − K ⟨y, gyrM [ȳ, −x̄] (dx π(v))⟩ y (1 + βy )βx βy (1 + βx )βy 1 + βx gyrM [ȳ, −x̄] (dx π(v)) − K ⟨y, gyrM [ȳ, −x̄] (dx π(v))⟩ y. = βx (1 + βy )βx (1)
The above equalities use: (1) the isometry property of π; (2) linearity of dȳ π −1 ; (3) Eq. (B.277). (4) Using the relation between λK x̄ and βx in the proof of Thm. 122, λK x̄ = B.6.4.4
1 + βx βx
λK ȳ =
1 + βy . βy
(B.285)
PV Geodesic Distance
We recall from Tab. 2.14 that the geodesic distance on the Poincaré ball PnK is √ 2 −1 d (y1 , y2 ) = √ tanh −K ∥(−y1 ) ⊕M y2 ∥ , −K P
y1 , y2 ∈ PnK .
(B.286)
By isometry and isomorphism, the PV geodesic distance is d(x, y) = dP (π(x), π(y)) √ 2 =√ tanh−1 −K ∥(−π(x)) ⊕M π(y)∥ −K √ 2 −1 =√ tanh −K ∥π(−x ⊕U y)∥ . −K 347
(B.287)
B.6. Proper Velocity Neural Networks B.6.4.5
Special Cases at the Identity
Exponential Map at the Identity. √ v̄ 1 −K(1 + β0 ) ∥v̄∥ Exp0 (v) = √ sinh β0 ∥v̄∥ −K v √ 1 (2) = √ −K ∥v∥ . sinh ∥v∥ −K (1)
(B.288)
The above comes from the following.
(1) 0 is the gyro identity;
(2) β0 = 1 and d0 π(v) = 12 v. Logarithmic Map at the Identity. As z = −0 ⊕U y = y, we have Log0 (y) = σ(0, y)z + τ (0, y) ⟨0, z⟩ 0 = σ(0, y)y −1
=√
βy From π(y) = 1+β y and βy = √ y
√ Let t =
2 tanh −K
1 1−K∥y∥2
−K ∥π(y)∥ =
√ −K ∥π(y)∥ y. ∥y∥
, we obtain βy √ −K ∥y∥ . 1 + βy
√ √ −K ∥y∥ and s = 1 + t2 , so that βy = √ a=
(B.289)
1 1−K∥y∥2
βy t t= . 1 + βy s+1 348
(B.290)
= 1s . Define (B.291)
Appendix B. Proofs Using the hyperbolic double-angle identity, we have tanh 2 tanh−1 (a) =
2a 1 + a2 2t/(s + 1) = 1 + t2 /(s + 1)2 2t(s + 1) = (s + 1)2 + t2 2t(s + 1) = 2 s + 2s + 1 + t2 2t(s + 1) = 2(1 + t2 + s) t(s + 1) = 1 + t2 + s t(s + 1) = 2 s +s t t = =√ . s 1 + t2
(B.292)
q √ cosh(u) = 1 + sinh2 (u) = 1 + t2 .
(B.293)
Denoting u = sinh−1 (t), we have
Therefore,
sinh(u) t tanh sinh−1 (t) = tanh(u) = =√ . cosh(u) 1 + t2
(B.294)
Since tanh is strictly increasing on R, this implies that −1
2 tanh
βy t = sinh−1 (t). 1 + βy
(B.295)
Substituting this identity back gives
and therefore
√ 1 sinh−1 −K ∥y∥ σ(0, y) = √ , ∥y∥ −K
(B.296)
√ y 1 −1 Log0 (y) = √ sinh −K ∥y∥ . ∥y∥ −K
(B.297)
gyrM [0, ȳ] = gyrM [ȳ, 0] = id .
(B.298)
Parallel Transport from the Identity. The gyration satisfies
349
B.6. Proper Velocity Neural Networks Substituting this into Eq. (B.284) gives 1 + β0 (1 + β0 )βy d0 π(v) − K ⟨y, d0 π(v)⟩ y β0 (1 + βy )β0 1 2βy 1 =2· v−K · ⟨y, v⟩ y 2 1 + βy 2 βy =v−K ⟨y, v⟩ y. 1 + βy
PT0→y (v) =
(B.299)
Parallel Transport to the Identity. Taking y = 0 in Eq. (B.284) and using gyrM [0, −x̄] = id yields 1 + βx dx π(v) βx 1 + βx βx3 βx = K ⟨x, v⟩ x + v βx (1 + βx )2 1 + βx βx2 ⟨x, v⟩ x. =v+K 1 + βx
PTx→0 (v) =
(B.300)
Distance from the Identity. This can be directly obtained by the gyro identity. √ 2 −1 √ d(0, y) = −K ∥π(y)∥ . tanh −K Using the same identity as above with t =
√ −K ∥y∥ yields
√ 1 −1 d(0, y) = √ sinh −K ∥y∥ . −K
B.6.5
(B.301)
(B.302)
Proof of Thm. 124
Proof. As shown in Thm. 88, the Möbius gyroaddition and gyromultiplication can be written by the Riemannian operators. Besides, the isometry πPVnK →PnK preserves the identity: πPVnK →PnK (0) = 0. By Nguyen and Yang [159, Lems. 2.1–2.2], one can directly obtain the results.
B.6.6
Proof of Thm. 125
We first establish the PV hyperplane equivalence and then derive the distance formula. 350
Appendix B. Proofs B.6.6.1
Equivalent Characterization of the PV Hyperplane
We first review a useful lemma from Chen et al. [55, Lem. J.1]. Lemma 196. We assume that the manifold M admits a gyrogroup defined by x ⊕ y = Expx (PTe→x (Loge (y))) , ∀x, y ∈ M,
(B.303)
where e ∈ M is the origin of the manifold. Then, we have the following Logp (x), a p = ⟨Loge (⊖p ⊕ x), PTp→e (a)⟩e ,
∀x, p ∈ M and ∀a ∈ Tp M. (B.304)
Now, we are ready to prove Thm. 125. Proof of PV hyperplane. Thm. 124 indicates that the assumption of Thm. 196 holds with M = PVnK , ⊕ = ⊕U and e = 0. Then, the PV hyperplane
can be rewritten as
o n Ha,p = x ∈ PVnK | Logp (x), a p = 0
Ha,p = x ∈ PVnK | ⟨Log0 (−p ⊕U x), PTp→0 (a)⟩0 = 0 .
(B.305)
(B.306)
Using the explicit PV operators in Thm. 123 and the PV metric in Eq. (5.1), we have Log0 (−p ⊕U x) = α(−p ⊕U x), for some scalar α ≥ 0, for some scalar β > 0,
PTp→0 (a) = βdp π(a),
(B.307)
g0 (u, v) = ⟨u, v⟩ .
As α = 0 is trivial, we only consider the case α > 0: ⟨Log0 (−p ⊕U x), PTp→0 (a)⟩0 = 0
B.6.6.2
⇐⇒
⟨−p ⊕U x, dp π(a)⟩ = 0.
(B.308)
PV Point-to-Hyperplane Distance
We first prove a lemma on the isometry and point-to-hyperplane distance, which will be used to derive the PV point-to-hyperplane distance. 351
B.6. Proper Velocity Neural Networks Lemma 197 (Isometry and point-to-hyperplane distance). Let (M, g) and (M̄, ḡ) be Riemannian manifolds and let ϕ : M → M̄ be a Riemannian isometry. For p ∈ M and a ∈ Tp M, define the hyperplane Ha,p = x ∈ M | gp Logp (x), a = 0 .
(B.309)
Let p̄ = ϕ(p) and ā = dp ϕ(a) ∈ Tp̄ M̄, and define the corresponding hyperplane on M̄ by ¯ p̄ (x̄), ā = 0 . H̄ā,p̄ = x̄ ∈ M̄ | ḡp̄ Log (B.310) Then ϕ maps Ha,p onto H̄ā,p̄ , that is,
ϕ (Ha,p ) = H̄ā,p̄ .
(B.311)
Moreover, for every x ∈ M we have dM (x, Ha,p ) = dM̄ ϕ(x), H̄ā,p̄ ,
(B.312)
when the point-to-hyperplane distance exists. Here, dM and dM̄ denote the Riemannian distances on M and M̄, respectively.
Proof. Since ϕ is a Riemannian isometry, we have ¯ p̄ (ϕ(x)) , ā . gp Logp (x), a = ḡp̄ dp ϕ Logp (x) , dp ϕ(a) = ḡp̄ Log
Therefore,
gp Logp (x), a = 0
⇐⇒
¯ p̄ (ϕ(x)) , ā = 0, ḡp̄ Log
(B.313)
(B.314)
which shows that x ∈ Ha,p if and only if ϕ(x) ∈ H̄ā,p̄ , and hence ϕ (Ha,p ) = H̄ā,p̄ . For the point-to-hyperplane distance, recall that for a subset S ⊂ M the distance from x to S is dM (x, S) = inf dM (x, z). (B.315) z∈S
352
Appendix B. Proofs For the point-to-hyperplane distance, we have dM (x, Ha,p ) = inf dM (x, z) z∈Ha,p
= inf dM̄ (ϕ(x), ϕ(z)) z∈Ha,p
(B.316)
= inf dM̄ (ϕ(x), z̄) z̄∈H̄ā,p̄
= dM̄ ϕ(x), H̄ā,p̄ .
Next, we review the Poincaré hyperplane and point-to-hyperplane distance [76, Sec. 3.1]. Poincaré Point-to-Hyperplane Distance. For a point p ∈ PnK and a normal vector a ∈ Tp PnK , the Poincaré point-to-hyperplane distance is given by Ganea et al. [76, Thm. 5]: o n P n x ∈ PK | Logp (x), a p = 0 = {x ∈ PnK | ⟨−p ⊕M x, a⟩ = 0} , (B.317) ! √ 1 2 −K |⟨−p ⊕M y, a⟩| −1 P P d (y, Ha,p ) = √ sinh . (B.318) −K 1 + K ∥−p ⊕M y∥2 ∥a∥ P Ha,p =
Proof of the PV point-to-hyperplane distance. Let p̄ = π(p) ∈ PnK ,
ā = dp π(a) ∈ Tp̄ PnK ,
ȳ = π(y) ∈ PnK .
(B.319)
By Thm. 197, the point-to-hyperplane distances satisfy dPV (y, Ha,p ) = dP ȳ, H̄ā,p̄ .
(B.320)
Applying the Poincaré distance formula in Eq. (B.318) with p = p̄, a = ā, and y = ȳ gives ! √ 1 2 −K |⟨−p̄ ⊕M ȳ, ā⟩| P −1 d ȳ, H̄ā,p̄ = √ sinh . (B.321) −K 1 + K ∥−p̄ ⊕M ȳ∥2 ∥ā∥ The gyrovector isomorphism π implies
−p̄ ⊕M ȳ = π(−p ⊕U y). 353
(B.322)
B.6. Proper Velocity Neural Networks Denote z = −p ⊕U y. From Eq. (5.8), we have the explicit expression π(z) = with βz > 0. Since βz = √
1 1−K∥z∥2
βz z, 1 + βz
(B.323)
, we obtain 2
1 + K ∥π(z)∥ = 1 + K
βz 1 + βz
2
∥z∥2
Kβz2 ∥z∥2 (1 + βz )2 β 2 (1 − βz−2 ) =1+ z (1 + βz )2 β2 − 1 =1+ z (1 + βz )2 βz − 1 =1+ 1 + βz 2βz = 1 + βz =1+
The above yields
Therefore,
√ √ 2 −K |⟨π(z), ā⟩| −K |⟨z, ā⟩| = . 2 ∥ā∥ 1 + K ∥π(z)∥ ∥ā∥
1 d (y, Ha,p ) = √ sinh−1 −K
B.6.7
√
−K |⟨−p ⊕U y, dp π(a)⟩| ∥dp π(a)∥
(B.324)
(B.325)
.
(B.326)
Proof of Thm. 126
Proof of PV MLR. For clarity, we fix a class index k and omit k in the notation whenever possible. We denote π = πPVnK →PnK as in Thm. 125. Step 1: From Hyperplane Distance to a Signed Score. The PV MLR in 354
Appendix B. Proofs Eq. (5.24) associated with parameters (p, a) for x ∈ PVnK is vk (x) = sign (⟨−pk ⊕U x, dpk π(ak )⟩) ∥ak ∥pk d (x, Hak ,pk ) √ −K |⟨−pk ⊕U x, dpk π(ak )⟩| (1) ∥ak ∥pk −1 = √ sign (⟨−pk ⊕U x, dpk π(ak )⟩) sinh ∥dpk π(ak )∥ −K √ −K ⟨−pk ⊕U x, dpk π(ak )⟩ (2) ∥ak ∥pk . = √ sinh−1 ∥dpk π(ak )∥ −K (B.327) The above comes from the following. (1) Thm. 125; (2) sinh−1 is odd and strictly increasing. Step 2: Trivialization and Reduction to a Single Direction. We adopt the unidirectional parameterization in Sec. 5.2.4.1: pk = Exp0 (rk [zk ]) ,
ak = PT0→pk (zk ) ,
[zk ] =
zk , ∥zk ∥
(B.328)
with zk ∈ T0 PVnK ∼ = Rn and rk ∈ R. As parallel transport is an isometry, we have ∥ak ∥pk = ∥zk ∥0 = ∥zk ∥.
(B.329)
Moreover, pk and zk are collinear, because Exp0 in Thm. 123 preserves directions at the origin. Using the explicit expression of PT0→y at the origin in Thm. 123, we see that PT0→pk maps zk to a linear combination of zk and pk . Therefore, ak is also collinear with zk . The differential dpk π in Thm. 121 has the form dpk π(v) = αk v + βk ⟨pk , v⟩ pk ,
αk > 0, βk ∈ R,
(B.330)
so dpk π maps any vector in span{zk } into the same one-dimensional subspace. Consequently, there exists a scalar λk > 0 such that dpk π(ak ) = λk zk .
(B.331)
The sign of λk can be absorbed into zk by redefining zk ← −zk if necessary. Without loss of generality we may assume λk > 0. Putting Eq. (B.328), Eq. (B.329), Eq. (B.331) 355
B.6. Proper Velocity Neural Networks and ∥dpk π(ak )∥ = λk ∥zk ∥ into Eq. (B.327) yields ∥zk ∥ vk (x) = √ sinh−1 −K
√
−K ⟨−pk ⊕U x, zk ⟩ . ∥zk ∥
(B.332)
Step 3: Eliminating the Gyroaddition. The remaining task is to expand the gyro-additive term in Eq. (B.332). From Sec. 5.2.2, PV gyroaddition is given by u ⊕U v = u + v +
1 − βv βu −K ⟨u, v⟩ u, βv 1 + βu
Setting u = −pk and v = x yields −pk ⊕U x = −pk + x +
1 βw = p . 1 − K∥w∥2
βpk 1 − βx −K ⟨−pk , x⟩ (−pk ). βx 1 + β pk
(B.333)
(B.334)
Taking the inner product with zk gives βpk 1 − βx −K ⟨−pk , x⟩ ⟨−pk , zk ⟩ ⟨−pk ⊕U x, zk ⟩ = ⟨−pk , zk ⟩ + ⟨x, zk ⟩ + βx 1 + β pk 1 − βx βpk = ⟨x, zk ⟩ + 1 + −K ⟨−pk , x⟩ ⟨−pk , zk ⟩ . βx 1 + β pk (B.335)
Next, we rewrite the above expression using the unidirectional parameterization of pk . From Eq. (B.328) and the explicit PV exponential at the origin in Thm. 123, we have z √ 1 k sinh −Krk . (B.336) pk = Exp0 (rk [zk ]) = √ ∥z −K k∥ Thus,
√ 1 ⟨−pk , zk ⟩ = − √ −Krk ∥zk ∥. sinh −K
(B.337)
Moreover, since pk and zk share the same direction, any x admits the decomposition x = x∥ + x⊥ ,
x∥ =
⟨x, zk ⟩ zk , ∥zk ∥2
⟨x⊥ , zk ⟩ = 0,
(B.338)
which implies ⟨−pk , x⟩ = −pk , x∥
√ ⟨x, z ⟩ ⟨x, zk ⟩ 1 k = ⟨−pk , zk ⟩ = − √ sinh −Krk . 2 ∥zk ∥ ∥zk ∥ −K 356
(B.339)
Appendix B. Proofs The beta factor at pk is √ 1 βp k = p = sech −Krk , 1 − K∥pk ∥2
where we used ∥pk ∥2 = − K1 sinh2
√
(B.340)
−Krk and the identity 1 + sinh2 (t) = cosh2 (t).
Using ⟨−pk , zk ⟩, ⟨−pk , x⟩, and βpk , we have ⟨−pk ⊕U x, zk ⟩ 1 − βx βpk = ⟨x, zk ⟩ + 1 + −K ⟨−pk , x⟩ ⟨−pk , zk ⟩ βx 1 + β pk βp k 1 −K = ⟨x, zk ⟩ + ⟨−pk , x⟩ ⟨−pk , zk ⟩ βx 1 + β pk ! √ sinh −Krk 1 √ = ⟨x, zk ⟩ + − ∥zk ∥ βx −K ! ! √ √ sinh −Krk ⟨x, zk ⟩ sinh −Krk βpk √ √ − − ∥zk ∥ −K 1 + β pk ∥zk ∥ −K −K √ √ sinh −Krk ∥zk ∥ βpk sinh2 −Krk √ = ⟨x, zk ⟩ − + ⟨x, zk ⟩ βx 1 + β pk −K ! √ √ βpk sinh2 −Krk sinh −Krk ∥zk ∥ √ = 1+ ⟨x, zk ⟩ − . 1 + β pk βx −K Since βpk = sech have
√
(B.341)
√ √ −Krk and 1 + sinh2 −Krk = cosh2 −Krk = 1/βp2k , we
√ √ 1 + βpk + βpk sinh2 −Krk βpk sinh2 −Krk 1+ = 1 + β pk 1 + β pk √ 1 + βpk cosh2 −Krk = 1 + β pk √ 1 1 + 1/βpk = = = cosh −Krk , 1 + βpk βpk
(B.342)
which implies √ √ sinh −Krk ∥zk ∥ √ ⟨−pk ⊕U x, zk ⟩ = cosh −Krk ⟨x, zk ⟩ − . βx −K 357
(B.343)
B.6. Proper Velocity Neural Networks p Recalling that βx = 1/ 1 − K∥x∥2 , we obtain ⟨−pk ⊕U x, zk ⟩ = cosh
√
−Krk ⟨x, zk ⟩ −
sinh
√
√
−Krk
−K
∥zk ∥
Substituting Eq. (B.344) into Eq. (B.332), we arrive at
p 1 − K∥x∥2 . (B.344)
√−K √ p √ ∥zk ∥ −1 −Krk ⟨x, zk ⟩ − sinh −Krk 1 − K∥x∥2 . vk (x) = √ sinh cosh ∥zk ∥ −K (B.345) Proof of PV MLR limits. By Taylor expansions, we have √ Krk2 −Krk = 1 − + O(K 2 ), 2 √ √ sinh −Krk = −Krk + O (−K)3/2 ,
cosh
(B.346)
p K∥x∥2 1 − K∥x∥2 = 1 − + O(K 2 ). 2
The argument of sinh−1 (·) in Eq. (B.345) can be simplified as −1
√
√−K
√
p 2 −Krk 1 − K∥x∥
⟨x, zk ⟩ − sinh √ −K Krk2 −1 2 = sinh + O(K ) ⟨x, zk ⟩ 1− 2 ∥zk ∥ √ K∥x∥2 3/2 2 − −Krk + O (−K) + O(K ) 1− 2 √ ⟨x, zk ⟩ −1 3/2 = sinh −K − rk + O (−K) ∥zk ∥ √ ⟨x, zk ⟩ − rk + O (−K)3/2 . = −K ∥zk ∥
sinh
cosh
−Krk
∥zk ∥
(B.347)
Substituting this into Eq. (B.345) gives
∥zk ∥ √ ⟨x, zk ⟩ 3/2 vk (x) = √ −K − rk + O (−K) ∥zk ∥ −K ⟨x, zk ⟩ = ∥zk ∥ − rk + O(−K) ∥zk ∥ = ⟨x, zk ⟩ − rk ∥zk ∥ + O(−K), K→0−
−−−−→ ⟨x, zk ⟩ − rk ∥zk ∥. 358
(B.348)
Appendix B. Proofs
B.6.8
Proof of Thm. 127
Proof of PV FC layer. Specializing Thm. 125 to p = 0 and a = ek and using that −0 ⊕U y = y gives the LHS √ 1 −1 sign (⟨d0 π(ek ), −0 ⊕U y⟩) d (y, Hek ,0 ) = √ −Kyk , sinh −K
(B.349)
with yk = ⟨y, ek ⟩. Then, we obtain √ 1 −Kvk (x) , yk = √ sinh −K
k = 1, . . . , m.
(B.350)
Proof of PV FC limits. By Thm. 126, as K → 0− we have vk (x) → ⟨x, zk ⟩ + bk ,
bk = −rk ∥zk ∥.
(B.351)
For K < 0 and vk (x) ̸= 0, we can rewrite yk as
√ sinh −Kvk (x) √ yk = vk (x) , −Kvk (x)
(B.352)
√ and we define the fraction to be 1 when vk (x) = 0. Since −K → 0 and vk (x) converges √ to a finite limit, we have −Kvk (x) → 0. Using the standard limit limu→0 sinh(u)/u = 1, it follows that √ sinh −Kvk (x) √ → 1 as K → 0− . (B.353) −Kvk (x) Combining the above limits yields
lim yk = lim− vk (x) = ⟨x, zk ⟩ + bk .
K→0−
B.6.9
K→0
(B.354)
Proof of Thm. 128
Proof. The result is first established in the Poincaré ball model in Thms. 83 and 90. Since π = πPVnK →PnK : PVnK → PnK is a Riemannian isometry, for x, y ∈ PVnK , v ∈ Tx PVnK , 359
B.6. Proper Velocity Neural Networks and t ∈ R, it intertwines the key geometric operators used in the proof: P π ExpPV x (v) = Expπ(x) (dx π(v)) , dx π LogPV (y) = LogPπ(x) (π(y)) , x π (x ⊕U y) = π(x) ⊕M π(y), π (t ⊗U x) = t ⊙M π(x).
(B.355) (B.356) (B.357) (B.358)
Together with the preservation of geodesic distances and Fréchet means under the Riemannian isometry π, these identities imply that both homogeneity identities are preserved under π. Therefore the same theorem holds for the PV model by the isometry π.
B.6.10
Proof of Thm. 129
Proof. We first recall the isometries between the Poincaré ball and the hyperboloid [182, Sec. 2.1]: x ps , 1 + |K|xt 1 1 − K ∥y∥2 p|K| 1 + K ∥y∥2 . πPnK →HnK (y) = 2y
πHnK →PnK (x) =
(B.359)
(B.360)
1 + K ∥y∥2
Hence, the following are Riemannian isometries: πHnK →PVnK = πPnK →PVnK ◦ πHnK →PnK ,
πPVnK →HnK = πPnK →HnK ◦ πPVnK →PnK .
(B.361)
It remains to derive the explicit formulas. ⊤ n For x = [xt , x⊤ s ] ∈ HK , we first map to the Poincaré ball:
y = πHnK →PnK (x) = Applying πPnK →PVnK from Eq. (5.8) yields πHnK →PVnK (x) = πPnK →PVnK (y) = 2γy2 y, 360
x ps . 1 + |K|xt 1 γy = q . 2 1 + K ∥y∥
(B.362)
(B.363)
Appendix B. Proofs Using y = xs /(1 +
p |K|xt ), we compute 2
1 + K ∥y∥2 = 1 + K
∥xs ∥ 2 = p 1 + |K|xt
1+
Since x ∈ HnK satisfies ⟨x, x⟩L = 1/K, we have
2 p |K|xt + K ∥xs ∥2 . 2 p 1 + |K|xt
(B.364)
1 . K
(B.365)
2 2 p p 1 2 2 1 + |K|xt + K ∥xs ∥ = 1 + |K|xt + K xt + K 2 p = 1 + |K|xt + Kx2t + 1 p = 1 + 2 |K|xt + |K|x2t + Kx2t + 1 p = 2 1 + |K|xt .
(B.366)
1 K
⟨x, x⟩L = −x2t + ∥xs ∥2 =
⇒
∥xs ∥2 = x2t +
Substituting this into the numerator gives
Therefore,
and hence
Finally,
p 2 1 + |K|xt 2 p 1 + K ∥y∥2 = , 2 = p 1 + |K|x t 1 + |K|xt
(B.367)
p 1 + |K|xt 1 γy2 = . 2 = 2 1 + K ∥y∥
πHnK →PVnK (x) = 2γy2 y = 2 ·
1+
(B.368)
p |K|xt x ps · = xs . 2 1 + |K|xt
(B.369)
For πPVnK →HnK , take x ∈ PVnK and map to the Poincaré ball by Eq. (5.8): y = πPVnK →PnK (x) =
βx x, 1 + βx 361
βx = q
1
. 2
1 − K ∥x∥
(B.370)
B.6. Proper Velocity Neural Networks Applying πPnK →HnK , we obtain 1 1 − K ∥y∥2 p|K| 1 + K ∥y∥2 . πPVnK →HnK (x) = πPnK →HnK (y) = 2y 1 + K ∥y∥2
(B.371)
We now simplify the spatial and temporal components separately. We write βx y= x, 1 + βx
2
∥y∥ =
βx 1 + βx
2
∥x∥2 .
(B.372)
Using βx2 = 1/(1 − K ∥x∥2 ), we obtain βx2 βx2 − 1 βx − 1 = = 2 2 (1 + βx ) (1 + βx ) (1 + βx ) 2βx 2 ⇒ 1 + K ∥y∥2 = , 1 − K ∥y∥2 = . 1 + βx 1 + βx
K ∥y∥2 = K ∥x∥2
(B.373)
The spatial component of πPVnK →HnK (x) is βx
2 1+βx x 2y = x, 2 = 2βx 1 + K ∥y∥ 1+βx
(B.374)
and the temporal component is 2 1 1 − K ∥y∥ 1 1+βx p p = · 2 2βx |K| 1 + K ∥y∥ |K| 1+β x 2
Thus,
q 1 − K ∥x∥2 q 1 p =p = = ∥x∥2 − K1 . (B.375) |K|βx |K|
πPVnK →HnK (x) =
q 2 1 ∥x∥ − K
362
x
.
(B.376)
Appendix B. Proofs
B.7
Hyperbolic Busemann Neural Networks
B.7.1
Proof of Thm. 130
Proof. Denoting κ2 = −K with κ > 0, we rewrite the Busemann functions in Eqs. (5.40) and (5.41) as 1 (Poincaré) B v (x) = log κ (Lorentz) B v (x) =
∥v − κx∥2 1 − κ2 ∥x∥2
!
,
(B.377)
1 log (κxt − κ ⟨xs , v⟩) . κ
(B.378)
For the Poincaré case, ∥v − κx∥2 1 − 2κ ⟨v, x⟩ + κ2 ∥x∥2 −2κ ⟨v, x⟩ + 2κ2 ∥x∥2 = = 1 + . 1 − κ2 ∥x∥2 1 − κ2 ∥x∥2 1 − κ2 ∥x∥2 Using log (1 + u) = u + O (u2 ) as u → 0 and 1 − κ2 ∥x∥2
−1
= 1 + O (κ2 ), we obtain
1 −2κ ⟨v, x⟩ + 2κ2 ∥x∥2 1 + O κ2 + O κ2 κ = −2 ⟨v, x⟩ + 2κ ∥x∥2 + O (κ)
B v (x) =
(B.379)
(B.380)
= −2 ⟨v, x⟩ + O (κ) . K→0−
Therefore, B v (x) −−−−→ −2 ⟨v, x⟩. For the Lorentz case, the hyperboloid constraint −x2t + ∥xs ∥2 = −κ−2 and xt > 0 yield q 1 κxt = 1 + κ2 ∥xs ∥2 = 1 + κ2 ∥xs ∥2 + O κ4 . (B.381) 2
Set z = −κ ⟨xs , v⟩ + 12 κ2 ∥xs ∥2 + O (κ4 ). Then
log (κxt − κ ⟨xs , v⟩) = log (1 + z) = z + O z 2 = −κ ⟨xs , v⟩ + O κ2 ,
(B.382)
K→0−
which implies B v (x) = − ⟨xs , v⟩ + O (κ) and therefore B v (x) −−−−→ − ⟨xs , v⟩. Substituting the above two limits into uk (x) = −αk B vk (x) + bk gives the stated Euclidean limits of the logits. 363
B.7. Hyperbolic Busemann Neural Networks
B.7.2
Proof of Thm. 132
We first review two useful characterizations of Busemann functions in Hadamard spaces. The class B below is given by Bridson and Haefliger [29, p. 271]. The equivalence follows from Bridson and Haefliger [29, Prop. II.8.22], whose formulation in terms of Busemann functions uses the identification of horofunctions with Busemann functions up to additive constants in Bridson and Haefliger [29, Cor. II.8.20]. Definition 198. Let (X , d) be a Hadamard space, and let B be the set of functions h : X → R on (X , d) satisfying: (1) h is convex;
(2) 1-Lipschitz: |h(x) − h(y)| ≤ d(x, y) for all x, y ∈ X ; (3) for any x0 ∈ X and r > 0, the function h attains its minimum on the sphere Sr (x0 ) at a unique point y with h(y) = h(x0 ) − r. Proposition 199. For a function h : X → R, the following conditions are equivalent: (1) h is a Busemann function; (2) h ∈ B; (3) h is convex, and for every t ∈ R, the set h−1 (−∞, t] is nonempty; moreover, for each x ∈ X , the curve cx : [0, ∞) → X defined by t 7→ πh−1 (−∞,h(x)−t] (x) is a geodesic ray. Now, we are ready to prove Thm. 132.
Proof. As τ2 = τ1 is trivial, we only consider τ2 ̸= τ1 . We assume τ2 > τ1 , and discuss the other direction last. Step 1: Symmetry. By definition, d Hτγ1 , Hτγ2 =
inf
x∈Hτγ1 ,y∈Hτγ2
d(x, y) =
inf
y∈Hτγ2 ,x∈Hτγ1
d(y, x) = d Hτγ2 , Hτγ1 .
(B.383)
Step 2: Lower Bound. For any x ∈ Hτγ2 and y ∈ Hτγ1 , the 1-Lipschitz property of B γ gives |B γ (x) − B γ (y)| ≤ d(x, y). (B.384) 364
Appendix B. Proofs With B γ (x) = τ2 and B γ (y) = τ1 this yields τ2 − τ1 ≤ d(x, y)
∀y ∈ Hτγ1 .
(B.385)
Taking infimum in Eq. (B.385) over y ∈ Hτγ1 gives τ2 − τ1 ≤ d x, Hτγ1 .
(B.386)
Step 3: Upper Bound. For any x ∈ Hτγ2 , by property (3) in Thm. 199, the projection map cx (t) = π{B γ ≤B γ (x)−t} (x) = πHBτγ −t (x), 2
t ∈ [0, ∞),
(B.387)
is a unit-speed geodesic ray: d(x, cx (t)) = t. Let t = τ2 − τ1 > 0 and z = cx (t) ∈ HBτγ1 . If B γ (z) < τ1 , then z lies in the interior of the horoball HBτγ1 . We can move slightly from z toward x along the geodesic segment xz to obtain a point zε with d(x, zε ) < d(x, z), contradicting the minimality of the projection. Hence, the projected point z = cx (t) indeed lies on Hτγ1 : B γ (z) = τ1 . Then, we have the following: d(x, Hτγ1 ) ≤ d(x, HBτγ1 ) = d(x, z) = d(x, cx (t)) = t = τ2 − τ1 .
(B.388)
Step 4: Sandwich Closure. Combining Eqs. (B.386) and (B.388) gives d(x, Hτγ1 ) = τ2 − τ1 .
(B.389)
Combining Eq. (B.388) and Eq. (B.386), d x, Hτγ1 = τ2 − τ1 ,
for every x ∈ Hτγ2 .
(B.390)
The right-hand side does not depend on x, hence d Hτγ2 , Hτγ1 = τ2 − τ1 .
(B.391)
Step 5: Opposite Direction. If τ1 > τ2 , geodesic completeness of the Hadamard space allows the geodesic ray cx from Step 3 to extend to a complete geodesic e cx : R → γ γ X . For any x ∈ Hτ2 , let s = τ1 − τ2 and z = e cx (−s). Since B (cx (t)) = B γ (x) − t for t ≥ 0, convexity and the 1-Lipschitz property of B γ give B γ (z) = B γ (x) + s = τ1 . 365
B.7. Hyperbolic Busemann Neural Networks Thus, z ∈ Hτγ1 and d(x, z) = s. Combining this upper bound with the 1-Lipschitz lower bound from Step 2 gives d x, Hτγ1 = τ1 − τ2 ,
for every x ∈ Hτγ2 .
(B.392)
Therefore, in both directions,
d (x, Hτγ ) = |B γ (x) − τ | ,
∀x ∈ X ,
(B.393)
and the symmetry of the distance between horospheres gives d Hτγ2 , Hτγ1 = |τ2 − τ1 |.
B.7.3
(B.394)
Proof of Thm. 135
Proof. The hyperplane and point-to-hyperplane distance are presented in Tabs. A.16 and A.17. The origin of the Poincaré ball model is e = 0 ∈ Pm K . The specific ones w.r.t. the origin [181, Def. 1 and Eq. (56)] are Hek ,0 = {y ∈ Pm K | ⟨ek , y⟩ = 0}, √ 1 2 −Kyk −1 d̄ (y, Hek ,0 ) = √ sinh . −K 1 + K ∥y∥2
(B.395) (B.396)
Equating d̄ (y, Hek ,0 ) with uk (x) from Eq. (5.59) gives −1
sinh
√ √ 2 −Kyk = −Kuk (x), 2 1 + K ∥y∥
∀k ∈ {1, . . . , m}.
(B.397)
Note that Eq. (B.397) takes the same form as Shimizu et al. [181, Eq. (56)], except that their responses are different. The proof below is inspired by their derivation. Applying sinh(·) on both sides of Eq. (B.397) yields √ √ 2 −Kyk = sinh −Ku (x) . k 1 + K ∥y∥2 366
(B.398)
Appendix B. Proofs 1 Define ωk := √−K sinh
√
−Kuk (x) and ω = [ωk ]m k=1 . Then, we have 2yk = 1 + K ∥y∥2 ωk ,
∀k,
(B.399)
which is equivalent to the vector identity
2y = 1 + K ∥y∥2 ω.
(B.400)
Hence, y is collinear with ω. Write y = λω with λ ≥ 0. Substituting into Eq. (B.400) and taking norms gives a quadratic in λ: (B.401)
K ∥ω∥2 λ2 − 2λ + 1 = 0. Solving and selecting the branch that satisfies y → 0 as ω → 0 yields λ=
q 1 − 1 − K ∥ω∥2 K ∥ω∥
2
=
Therefore,
1 q . 2 1 + 1 − K ∥ω∥
(B.402)
−Kuk (x) √ , −K
(B.403)
m m Hw,p = {x ∈ Lm K | ⟨w, x⟩L = 0}, with p ∈ LK , w ∈ Tp LK .
(B.404)
y=
ω
q , 2 1 + 1 − K ∥ω∥
ωk =
sinh
√
which proves the claim. One can check that y ∈ Pm K.
B.7.4
Proof of Thm. 136
Proof. Recalling Tab. A.16, a Lorentz hyperplane is
The canonical origin is 0 ∈ Lm K . The tangent space at the origin is ⊤ ⊤ m T0 Lm K = {[0, v ] | v ∈ R },
(B.405)
where each tangent vector has a zero time component. Therefore, the coordinate hyperplane through the origin and orthogonal to the k-th axis is Hēk ,e = {y ∈ Lm K | ⟨ēk , y⟩L = 0}
= {y = (yt , ys ) ∈ Lm K | (ys )k = 0}, 367
(B.406)
B.7. Hyperbolic Busemann Neural Networks ⊤ m where ēk = [0, e⊤ k ] ∈ T0 LK . From Tab. A.17, the associated signed point-to-hyperplane distance is d̄ (y, Hēk ,e ) = sign (⟨ēk , y⟩L ) d (y, Hēk ,e ) √ (B.407) 1 −1 −K(ys )k . =√ sinh −K
Equating d̄ (y, Hēk ,e ) with uk (x) from Eq. (5.59) gives sinh−1
√
√ −K(ys )k = −Kuk (x),
1 ≤ k ≤ m.
(B.408)
√ 1 sinh −Kuk (x) , −K
1 ≤ k ≤ m.
(B.409)
Applying sinh(·) to both sides of Eq. (B.408) yields (ys )k = √
Stacking the coordinates gives ys = √
√ 1 sinh −Ku(x) , −K
u(x) = (u1 (x), . . . , um (x))⊤ .
(B.410)
2 2 Since y ∈ Lm K , the hyperboloid constraint ⟨y, y⟩L = 1/K implies −yt + ∥ys ∥ = 1/K. Taking the positive time component yields
yt =
r
1 + ∥ys ∥2 . −K
(B.411)
Combining the expressions for yt and ys proves the claim.
B.7.5
Proof of Thm. 137
Proof. Set K = −κ2 with κ > 0. Poincaré Case. Recall that y=
ω q , 2 2 1 + 1 + κ ∥ω∥
ωk =
sinh (κuk (x)) . κ
(B.412)
For any bounded scalar z, sinh (κz) = κz +
κ3 z 3 sinh (κz) κ2 z 3 + O κ5 ⇒ =z+ + O κ4 . 3! κ 6 368
(B.413)
Appendix B. Proofs Applying this to z = uk (x) yields κ2 ωk = uk (x) + uk (x)3 + O κ4 = uk (x) + O κ2 . 6
(B.414)
Thus, ω = u(x) + O (κ2 ) and ∥ω∥2 = ∥u(x)∥2 + O (κ2 ). For the denominator, q which gives
1 1 + κ2 ∥ω∥2 = 1 + κ2 ∥ω∥2 + O κ4 , 2
q 1 1 + κ2 ∥ω∥2 = 2 + κ2 ∥ω∥2 + O κ4 . 2 Taking the reciprocal produces 1+
1 1 q = + O κ2 . 2 1 + 1 + κ2 ∥ω∥2
(B.415)
(B.416)
(B.417)
Multiplying with ω = u(x) + O (κ2 ) gives 1 y = u(x) + O κ2 . 2
(B.418)
By Thm. 130, uk (x) → 2αk ⟨vk , x⟩ + bk , hence
1 yk → αk ⟨vk , x⟩ + bk . 2
(B.419)
r
(B.420)
Lorentz Case. Recall that 1 ys = sinh (κu(x)) , κ
yt =
1 + ∥ys ∥2 . κ2
Using the same expansion as above, 1 ys,k = κ
κ3 κuk (x) + uk (x)3 + O κ5 3!
= uk (x) + O κ2 .
Therefore, ys = u(x) + O (κ2 ) and ∥ys ∥2 = ∥u(x)∥2 + O (κ2 ). 369
(B.421)
B.8. Full-Rank Correlation Networks Factor out κ−1 and expand the square root: q 1 yt = 1 + κ2 ∥ys ∥2 κ 1 2 1 2 4 1 + κ ∥ys ∥ + O κ = κ 2 1 κ = + ∥ys ∥2 + O κ3 κ 2 1 = + O (κ) → ∞. κ
(B.422)
By Thm. 130, uk (x) → αk ⟨vk , xs ⟩ + bk . Using the spatial expansion yields (ys )k → αk ⟨vk , xs ⟩ + bk .
(B.423)
This completes the proof.
B.8
Full-Rank Correlation Networks
B.8.1
Proof of Thm. 138
We first prove a lemma for MLRs on general isometric manifolds, of which this theorem is a specific case. Notably, the result and proof can be readily extended to the case where Rm is endowed with an arbitrary inner product. Lemma 200 Riemannian MLRs). Given m-dimensional Riemannian (Isometric f M f → M, f manifolds M, g and M, g M with a Riemannian isometry ϕ : M f and ϕ(E) ∈ M. The Riemannian MLR over M f for the their origins are E ∈ M f of each class k = 1, · · · , C can be calculated by the one over M: input X ∈ M vkM (X; Zk , γk ) = vkM (ϕ(X); ϕ∗,E (Zk ), γk ), f
(B.424)
f∼ f → Tϕ(E) M as the differential map. with γk ∈ R, Zk ∈ TE M = Rm , and ϕ∗,E : TE M f Here, vkM and vkM are the specific realizations of the Riemannian MLR reviewed in f and M, respectively. Sec. 4.3.2.1 over M
e Log, g ⟨·, ·⟩ , Proof. We omit the subscript k in Ak and Pk for simplicity. We denote Γ, P e A,P ), H e A,P as the parallel transport along the geodesic, Riemannian log∥·∥P , e d(X, H arithm, Riemannian metric, the induced norm, margin distance and hyperplane over 370
Appendix B. Proofs f while the counterparts over M are denoted as Γ, Log, ⟨·, ·⟩ M, ϕ(P ) , ∥·∥ϕ(P ) , d, and H, respectively. From the isometry, we have D
∥A∥P = ∥ϕ∗,P (A)∥ϕ(P ) , E g P (X), A = Logϕ(P ) (ϕ(X)), ϕ∗,P (A) Log . ϕ(P ) P
(B.425) (B.426)
The above equations imply
e A,P = Hϕ (A),ϕ(P ) . ϕ H ∗,P
(B.427)
Denoting H = Hϕ∗,P (A),ϕ(P ) , we have the following for the margin distance e e A,P ) = d(X, H
inf e A,P Q∈H
(1)
=
inf
e d(X, Q)
d(ϕ(X), ϕ(Q))
e A,P Q∈H
(B.428)
(2)
= inf d(ϕ(X), R) R∈H
(3)
= d(ϕ(X), H).
The above comes from the following. (1) Isometry. (2) Eq. (B.427). (3) Definition of margin distance. Combining the above, we have v M (X; P, A) g P (X)⟩P )∥A∥P e e A,P ) = sign(⟨A, Log d(X, H (B.429) = sign Logϕ(P ) (ϕ(X)), ϕ∗,P (A) ϕ(P ) ∥ϕ∗,P (A)∥ϕ(P ) d(ϕ(X), Hϕ∗,P (A),ϕ(P ) ) f
= v M (ϕ(X); ϕ(P ), ϕ∗,P (A)) .
Finally, let us further consider trivialization. By isometry, we have the following: eE→P (Z) A=Γ
= ϕ−1 ∗,P PTϕ(E)→ϕ(P ) (ϕ∗,E (Z)) , 371
(B.430)
B.8. Full-Rank Correlation Networks
Then, we have
g E (γ[Z]) P = Exp
= ϕ−1 Expϕ(E) (γ[ϕ∗,E (Z)]) .
(B.431)
ϕ∗,P (A) = PTϕ(E)→ϕ(P ) (ϕ∗,E (Z)),
(B.432)
ϕ(P ) = Expϕ(E) (γ[ϕ∗,E (Z)]) .
(B.433)
Putting the above two equations into Eq. (B.429), we have v M (X; Z, γ) = v M (X; P, A) f
f
= v M (ϕ(X); ϕ(P ), ϕ∗,P (A))
= v M ϕ(X); Expϕ(E) (γ[ϕ∗,E (Z)]), PTϕ(E)→ϕ(P ) (ϕ∗,E (Z))
(B.434)
= vkM (ϕ(X); ϕ∗,E (Zk ), γk ).
Thm. 138 is a special case of Thm. 200 and can be readily proven accordingly. Proof of Thm. 138. MLR. In Euclidean space Rm , simple computations show that the Riemannian MLR reviewed in Sec. 4.3.2.1 becomes Eq. (4.1), where the latter is equal to ⟨ak , x − pk ⟩. Based on Thm. 200, we have vk (X; Zk , γk ) = vkR (ϕ(X); ϕ∗,E (Zk ), γk ), m
= ⟨ϕ(X) − γk [ϕ∗,E (Zk )], ϕ∗,E (Zk )⟩
(B.435)
= ⟨ϕ(X), ϕ∗,E (Zk )⟩ − γk ∥ϕ∗,E (Zk )∥ , Margin Hyperplane. In Euclidean space Rm , the Riemannian margin hyperplane becomes the Euclidean one, which is parameterized by ⟨ak , x − pk ⟩ = 0. Together with Eq. (B.435), the results can be easily obtained.
B.8.2
Proof of Thm. 139
Proof. First, we have the following: Θ(I) = I,
(B.436)
Chol(I) = I,
(B.437)
log∗,I (V ) = V,
∀V ∈ Hol(n),
372
(B.438)
Appendix B. Proofs log∗,I (V ) = V, D⋆ (I) = I.
(B.439)
∀V ∈ LT0 (n),
(B.440)
Putting the above into the differential formulas collected in Sec. 2.9.2, one can directly get the result w.r.t. ECM, LECM, and OLM. For LSM, based on Sec. 2.9.2, we have 1 0 0 V Σ + ΣV ∆V ∆ + 2 1 (1) =V + V0+V0 2
Log⋆∗,I (V ) = log∗,Σ
(B.441)
(2)
= V − diag (V 1) .
The above comes from the following. (1) Σ = ∆ = I (2)
V 0 = −2 diag (In + Σ)−1 ∆V ∆1
(B.442)
= − diag (V 1)
B.8.3
Proof of Thm. 142
Let dn = n(n−1) and dm = m(m−1) be the manifold dimensions of Cor+ (n) and Cor+ (m), 2 2 respectively. We have the following general results. Lemma 201. Let Cor+ (n), g n be isometric to Rdn by the diffeomorphism ϕn : Cor+ (n) → Rdn ,
and let Cor+ (m), g m be isometric to Rdm by the diffeomorphism ϕm : Cor+ (m) → Rdm .
(B.443)
(B.444)
The diffeomorphism satisfies In = ϕ−1 n (0dn ),
Im = ϕ−1 m (0dm ). 373
(B.445)
B.8. Full-Rank Correlation Networks
The correlation FC layer F : Cor+ (n) → Cor+ (m) for the input X ∈ Cor+ (n) is Y = ϕ−1 m
dm X
vi (X)ei
i=1
!
(B.446)
,
m m for where {ei }di=1 is the canonical orthonormal basis over Rdm with ei = (δik )dk=1 dm each i. Here, {vi (X)}i=1 is given by Thm. 138:
vi (X) = ⟨ϕn (X), (ϕn )∗,In (Zi )⟩ − γi ∥(ϕn )∗,In (Zi )∥ ,
(B.447)
with Zi ∈ TIn Cor+ (n) ∼ = Rdn and γi ∈ R as the FC parameters. dm dm Proof of Thm. 201. Let {Ok = (ϕm )−1 ∗,Im (ek )}k=1 . Then {Ok }k=1 is an orthonormal basis over TIm Cor+ (m).
The LHS of Eq. (5.69) is LogIm (Y ), Ok Im d(Y, HOk ,Im ) (1) −1 = sign (ϕm )∗,Im ϕm (Y ), Ok I d(Y, HOk ,Im )
sign
m
(2)
= sign (⟨ϕm (Y ), ek ⟩) d(Y, HOk ,Im )
(B.448)
(3)
= sign (⟨ϕm (Y ), ek ⟩) d(ϕm (Y ), Hek ,0dm )
= (ϕm (Y ))k , where (1)–(2) come from the isometry, and (3) comes from Eq. (B.428). The RHS of Eq. (5.69) can be implied by Thm. 138. Thm. 201 can be naturally extended to the cases where the inner products of Rdn and Rdm are not canonical. Lemma 202. Following all the notation in Thm. 201, we further assume that the inner products Qn (·, ·) over Rdn and Qm (·, ·) over Rdm are not necessarily canonical. In addition, f : (Rdm , Qm (·, ·)) → (Rdm , ⟨·, ·⟩) is a linear isometry to the canonical inner product. Then, we have −1 Y = ϕ−1 m ◦f
dm X i=1
vi (X)ei
!
(B.449)
, n
vi (X) = Q (ϕn (X), (ϕn )∗,In (Zi )) − γi ∥(ϕn )∗,In (Zi )∥Q , n
374
(B.450)
Appendix B. Proofs
Figure B.1: Illustration of the Euclidean spaces LT0 (m), Hol(m) and Row0 (m), where ⋆ can be obtained by symmetry.
n
where ∥·∥Q is the norm induced by Qn . Proof of Thm. 202. First, we denote ψ m = f ◦ ϕm : Cor+ (m), g m → (Rdm , ⟨·, ·⟩).
(B.451)
Note that the differential of any linear map between vector spaces is itself. The rest of the proof is identical to that of Thm. 201. Now, we present the proof of Thm. 142. Proof of Thm. 142. As ECM, LECM, OLM, and LSM are pullback metrics from Euclidean spaces, we resort to Thm. 201 and its extension Thm. 202. Denoting the zero matrix as 0, we have the following: ϕEC (In ) = log ◦Θ(In ) = 0 ∈ LT0 (n), Log◦ (In ) = 0 ∈ Hol(n),
Log⋆ (In ) = 0 ∈ Row0 (n).
(B.452) (B.453) (B.454)
Therefore, the identity matrix is indeed the origin defined in Thm. 201. Recalling Thm. 201, the prototype space is the vector space with the standard vector inner product. Obviously, LT0 (m), Hol(m), and Row0 (m) are linearly isomorphic to Rm(m−1)/2 . As shown in Fig. B.1, each L ∈ LT0 (m) can be identified with a vector of its lower triangular part. Besides, LT0 (m) with the canonical matrix inner product is m(m−1) identified with R 2 with standard vector inner product. Therefore, the basis over 0 LT (m) corresponding to the canonical orthonormal basis over Rm(m−1)/2 is LT0 (m)
(LT0 (m), ⟨·, ·⟩) : Uij
= Eij ,
375
1 ≤ j < i ≤ m,
(B.455)
B.8. Full-Rank Correlation Networks where Eij ∈ Rm×m is the standard basis matrix, with the (k, l)-th element defined as (Eij )kl =
1 0
if k = i and l = j,
(B.456)
otherwise.
Without loss of generality, we identify (LT0 (m), ⟨·, ·⟩) with (R {Eij }1≤j<i≤m as the canonical orthonormal basis.
m(m−1) 2
, ⟨·, ·⟩), and refer to
However, {Eij } is neither a canonical orthonormal basis nor even orthonormal for Hol(m) and Row0 (m) under the standard matrix inner product. According to Thm. 202, we only need to find the linear isometry that maps these two spaces into (LT0 (m), ⟨·, ·⟩). By Fig. B.1, we have the following linear isometries to pull back these two inner products to the standard ones over LT0 (m): fHol(m)→LT0 (m) :(Hol(m), ⟨·, ·⟩) → (LT0 (m), ⟨·, ·⟩), √ Hol(m) ∋ H 7−→ 2⌊H⌋ ∈ LT0 (m),
fRow0 (m)→LT0 (m) :(Row0 (m), ⟨·, ·⟩) → (LT0 (m), ⟨·, ·⟩), √ √ e + 3D(R) e ∈ LT0 (m), Row0 (m) ∋ R 7−→ 6⌊R⌋
(B.457)
e ∈ S m−1 is the leading principal submatrix of order m − 1 of R. The bases where R −1 −1 fHol(m)→LT 0 (m) ({Eij }) and fRow (m)→LT0 (m) ({Eij }) are as follows: 0 Eij + Eji √ , 1≤j<i≤m (B.458) 2 Eii −E√im −Emi , if 1 ≤ i < m Row0 (m) 3 (Row0 (m), ⟨·, ·⟩) : Uij = (B.459) −Emj −Ejm Eij +Eji −Emi −E √ im , if 1 ≤ j < i < m 6 Hol(m)
(Hol(m), ⟨·, ·⟩) : Uij
=
Putting the required diffeomorphisms and vijg in Thm. 140 into Thm. 202 for ECM, LECM, OLM, and LSM, the corresponding FC layers can be readily obtained.
B.8.4
Proof of Thm. 143
Proof. First, we review the isometries between the open hemisphere and hyperboloid [195, Eqs. (4.1)–(4.2)], and the one between Poincaré ball and hyperboloid [182, Sec. 2.1]: ψHSn →Hn : (x1 , . . . , xn+1 )⊤ ∈ HSn 7−→
1 xn+1 376
(1, x1 , . . . , xn )⊤ ∈ Hn ,
(B.460)
Appendix B. Proofs 1 ⊤ ⊤ ys , 1 ∈ HSn , yt xs ⊤ ⊤ n ∈ Pn , ψHn →Pn : xt , xs ∈ H 7−→ 1 + xt !⊤ 1 + ∥y∥2 2y ⊤ 1 n ψPn →Hn : y ∈ P 7−→ = 2, 2 1 − ∥y∥ 1 − ∥y∥ 1 − ∥y∥2
ψHn →HSn : yt , ys⊤
⊤
For any (x⊤ , xn+1 )⊤ ∈ HSn and y ∈ Pn , we have ψHSn →Pn
(B.461)
∈ Hn 7−→
x xn+1
!!
x
= ψHn →Pn ◦ ψHSn →Hn 1
= ψHn →Pn =
x
xn+1
1 x
xn+1 !!
(B.462) 1 + ∥y∥2 2y
!
∈ Hn . (B.463)
!! (B.464)
1
1 xn+1 1 + xn+1 x = 1 + xn+1
ψPn →HSn (y) = ψHn →HSn ◦ ψPn →Hn (y) = ψHn →HSn
1 1 − ∥y∥2
1 1 + ∥y∥2
2y 1 − ∥y∥2
=
B.8.5
1 + ∥y∥2 2y !
!!
(B.465)
Proof of Thm. 145
Proof. We denote D = D(H). By Sec. 2.9.2, we have dY = dD + dH −1 dD = − diag H 0 D exp∗,Y (dH) 1 .
(B.466)
Following Ionescu et al. [111], we denote the inner product ⟨·, ·⟩ as · : · for simplicity. By the invariance of differential and properties of trace [111, Eqs. 67–72], we have the 377
B.8. Full-Rank Correlation Networks following: ∂l ∂l ∂l : dY = : dD + : dH ∂Y ∂Y ∂Y ∂l ∂l = : − diag (H 0 )−1 D exp∗,Y (dH) 1 + : dH ∂Y ∂Y ! ⊤ ∂l ∂l (1) = tr −Dv : dH (H 0 )−1 D exp∗,Y (dH) 1 + ∂Y ∂Y " # ! ⊤ ∂l ∂l (2) = tr − 1Dv : dH (H 0 )−1 D exp∗,Y (dH) + ∂Y ∂Y ∂l ∂l 0 −1 = −(H ) Dv 1⊤ : D exp∗,Y (dH) + : dH ∂Y ∂Y ∂l ∂l 0 −1 1⊤ : exp∗,Y (dH) + : dH = −D (H ) Dv ∂Y ∂Y ∂l ∂l (3) 0 −1 = − exp∗,Y D (H ) Dv 1⊤ : dH ∂Y ∂Y ∂l ∂l (4) 0 −1 − exp∗,Y D (H ) Dv 1⊤ = off : dH ∂Y ∂Y
(B.467)
The above comes from the following. (1) A : diag(b) = Dv(A) : b, ∀A ∈ Rn×n , b ∈ Rn , a : b = a⊤ b = tr(a⊤ b), ∀a, b ∈ Rn .
(B.468) (B.469)
(2) Cyclic property of the trace for matrices A, B, and C of compatible dimensions: tr(ABC) = tr(CAB). (3) For any A ∈ S n and S ∈ S n , write Y = U ∆U ⊤ and let L = Lexp be the Loewner matrix in Eq. (2.91). By the Daleckii–Krein formula in Eq. (2.90) and the properties of trace, we have A : exp∗,Y (S) = A : U L ⊛ U ⊤ SU U ⊤ = U L ⊛ U ⊤ AU U ⊤ : S = exp∗,Y (A) : S.
(4) H has zero diagonal elements. 378
(B.470)
Appendix B. Proofs The invariance of the first-order differential gives ∂l ∂l : dY = : dH. ∂Y ∂H
(B.471)
∂l By the last equation in Eq. (B.467), we can obtain ∂H .
B.8.6
Proof of Thm. 146
⋆ ⋆ Proof. Denoting by f : Cor+ (n) → Row+ 1 (n) the map f (C) = D (C)CD (C) = Σ, we have Log⋆∗,C = log∗,Σ ◦f∗,C . (B.472)
Combining with the differential of Log⋆ shown in Sec. 2.9.2, we have the following differential equation: dΣ = ∆dC∆ − V 0 Σ + ΣV 0 , (B.473) with V 0 = diag (In + Σ)−1 ∆dC∆1 . Similarly to Thm. 145, we have the following: ∂l ∂l : dΣ = : ∆dC∆ − V 0 Σ + ΣV 0 ∂Σ ∂Σ ∂l ∂l = ∆ ∆ : dC − : V 0 Σ + ΣV 0 ∂Σ ∂Σ ∂l ∂l ∂l Σ+Σ = ∆ ∆ : dC − : diag (In + Σ)−1 ∆dC∆1 ∂Σ ∂Σ ∂Σ ∂l ∂l ∂l = ∆ ∆ : dC − Dv Σ+Σ : (In + Σ)−1 ∆dC∆1 ∂Σ ∂Σ ∂Σ ∂l = ∆ ∆ : dC − tr ve⊤ (In + Σ)−1 ∆dC∆1 ∂Σ ∂l = ∆ ∆ : dC − tr ∆1e v ⊤ (In + Σ)−1 ∆dC ∂Σ ∂l = ∆ ∆ : dC − ∆ (In + Σ)−1 ve1⊤ ∆ : dC ∂Σ ∂l −1 ⊤ = ∆ ∆ − ∆ (In + Σ) ve1 ∆ : dC ∂Σ ∂l −1 ⊤ = ∆ − (In + Σ) ve1 ∆ : dC. ∂Σ
By imposing symmetrization, we can obtain the results. 379
(B.474)
B.9. Adaptive Log-Euclidean Metrics
B.8.7
Proof of Thm. 144
As β-splitting is the inverse of β-concatenation [181], we only need to show the case w.r.t. β-concatenation. Besides, it suffices to prove the 2D case, which is shown in the following lemma. Lemma 203. Given xij ∈ Pnj with i ∈ {1, . . . , Ni } and j ∈ {1, . . . , Nj }, applying the β-concatenation sequentially 2 times in the order j → i is equivalent to a single β-concatenation along all indices simultaneously. Proof. Denoting d =
PNj
j=1 nj and vij = Log0 (xij ), we have the following
Nj −1 −1 Exp0 βNi ×d βd concatj=1 βd βnj vij i=Ni ,j=Nj βNi ×d βd−1 βd βn−1 v = Exp0 concati=1,j=1 ij j i=Ni ,j=Nj vij . = Exp0 concati=1,j=1 βNi ×d βn−1 j i concatN i=1
(B.475)
The last line implies the claim.
A special case of the above lemma is where all nj are identical. Corollary 204. Given xij ∈ Pn with i ∈ {1, . . . , Ni } and j ∈ {1, . . . , Nj }, applying the β-concatenation sequentially 2 times in the order j → i is equivalent to a single β-concatenation along all indices simultaneously. Thm. 144 can be obtained by Thms. 203 and 204.
B.9
Adaptive Log-Euclidean Metrics
B.9.1
Proof of Thm. 147
Proof of Thm. 147. Let us first deal with (α, β)-LEM. Substituting the differential of the matrix logarithm into Thm. 33 directly yields the result. Now, let us focus on LCM. Denote LCM, the standard Euclidean metric, and the metric on the Cholesky manifold [137] by g LC , g E , and g C , respectively. By Tab. 2.5, n {S++ , g LC } is isometric to {Ln++ , g C }, with the Cholesky decomposition Chol as an isometry. This is exactly how Lin [137] derived LCM. So, the key point lies in the Cholesky metric g C . Let us reveal why it is defined in this way. In fact, g C is derived 380
Appendix B. Proofs from g E by φln . Simple computations show that (φln )∗,L (V ) = ⌊V ⌋ + D(L)−1 D(V ),
(B.476)
where V ∈ TL Ln++ . By Eq. (B.476), Tab. 2.5 can be rewritten as gLC (X, Y ) = g E (φln )∗,L (X), (φln )∗,L (Y ) .
(B.477)
n Therefore, φln : Ln++ → LTn is an isometry. By transitivity, ψLC : S++ → LTn is also an isometry.
B.9.2
Proof of Thm. 148
Proof of Thm. 148. As Rn(n+1)/2 ∼ = LTn ∼ = S n , LCM is therefore a pullback metric from the standard Euclidean space S n . Second, any two Euclidean spaces of the same finite dimension are naturally isometric; hence, (α, β)-LEM is also a pullback metric from the standard Euclidean space S n .
B.9.3
Proof of Thm. 149
Proof of Thm. 149. By the definitions in Eqs. (6.2) to (6.5), the Hilbert-space and isomorphism claims follow directly. It remains to establish the geometric claims. As n every Euclidean space is an abelian Lie group, {S++ , ⊙ϕ } is an abelian Lie group. The geodesic distance in Eq. (6.6) also follows immediately because ϕ is a Riemannian isometry. We only need to prove Eqs. (6.7) to (6.9). Note that in the Euclidean space S n , for any x, y ∈ S n and tangent vector v ∈ Tx S n ∼ = S n , we have the following: Expx v = x + v,
(B.478)
Logx y = y − x,
(B.479)
PTx→y v = v.
(B.480)
By the isometry of ϕ, we can readily obtain Eqs. (6.7) to (6.9).
B.9.4
Proof of Thm. 150
Proof of Thm. 150. Obviously, log−1 α is the inverse of logα . What follows is to verify the smoothness of logα and its inverse. 381
B.9. Adaptive Log-Euclidean Metrics According to Magnus and Neudecker [144, Thm. 8.9], the map producing an eigenvalue or an eigenvector from a real symmetric matrix is C ∞ . Recalling logα and its inverse map log−1 α , it is obvious that they comprise arithmetic calculations or compositions of smooth maps. Therefore, logα is a diffeomorphism with inverse log−1 α .
B.9.5
Proof of Thm. 152
Proof of Thm. 152. This is a direct result of Thm. 149.
B.9.6
Proof of Thm. 154
Proof of Thm. 154. The differentials of log−1 α and logα can be derived similarly. In the following, we only present the process of deriving the differential of logα . First, let us recall the differentials of eigenvalues and eigenvectors. Magnus and Neudecker [144, Thm. 8.9] offers their Euclidean differentials, which are the exact formulations for differentials under the canonical base on SPD manifolds. Thus, we can readily obtain the differentials of eigenvalues and eigenvectors as follows: σ∗,S (V ) = u⊤ V u,
(B.481)
u∗,S (V ) = (σIn − S)+ V u,
(B.482)
where Su = σu, u⊤ u = 1, and (·)+ is the Moore–Penrose inverse. By the RHS of Eq. (6.30), the differential map of logα is (logα )∗,S (V ) = U∗,S (V ) logα (Σ)U ⊤ + U (logα )∗,Σ (Σ∗,S (V )) U ⊤ ⊤ + U logα (Σ)U∗,S (V )
(B.483)
= Q + Q⊤ + U (logα )∗,Σ (Σ∗,S (V )) U ⊤ , where Q = U∗,S (V ) logα (Σ)U ⊤ . For the differential of diagonal logarithm, it is 1 (logα )∗,Σ (Σ∗,S (V )) = A Σ∗,S (V ), Σ
(B.484)
where A is defined in Eq. (6.31). Denote the eigenvectors and eigenvalues of S = U ΣU ⊤ by U = (u1 , . . . , un ) and Σ = diag(σ1 , . . . , σn ). By Eqs. (B.481) to (B.484), the differential of logα can be 382
Appendix B. Proofs obtained.
B.9.7
Proof of Thm. 155
Proof of Thm. 155. Following the notation in the proposition, we prove the result as follows. By abuse of notation, in the following, we omit the wide tilde e. Now, we proceed to deal with the differential of log−1 α . We rewrite the formula of −1 logα as log−1 α (X)
(B.485)
= U diag (aσ1 1 , · · · , aσnn ) U ⊤ ,
(B.486)
= U diag elog(a1 )σ1 , · · · , elog(an )σn U ⊤ , = U diag
∞ X (log(a1 )σ1 )k k=0
∞ X (BΣ)k
=U
k=0
=
k!
k!
!
,··· ,
∞ X (log(an )σn )k k=0
k!
!
(B.487) U ⊤,
(B.489)
U ⊤,
∞ X (P X)k k=0
(B.488)
(B.490)
k!
where P = U BU ⊤ , U is obtained from the eigendecomposition X = U ΣU ⊤ , and B = diag (log(a1 ), · · · , log(an )) is diagonal. By the properties of normed vector algebras [197, Prop. 15.14], we can obtain the last equation. Then, we can compute the differential of n n log−1 α by curves. Given a curve c on S starting at X with initial velocity V ∈ TX S , write c(t) = U (t)Σ(t)U (t)⊤ and define P (t) = U (t)BU (t)⊤ , so that P (0) = P . We have d log−1 (c(t)) log−1 α ∗,X (V ) = dt t=0 α ∞ X d (P (t)c(t))k = . dt t=0 k=0 k!
(B.491)
Term-by-term differentiation gives log−1 α ∗,X (V )
k−1 ∞ X 1 X d = ( (P X)k−l−1 (P (t)c(t))(P X)l ). k! l=0 dt t=0 k=1
383
(B.492)
B.9. Adaptive Log-Euclidean Metrics By the chain rule, we have d (P (t)c(t)) = P ′ (0)X + P V. dt t=0 P ′ (0) is obtained by P ′ (0) =
d U (t)BU (t)⊤ dt t=0
= U ′ (0)BU ⊤ + U BU ′ (0)⊤
(B.493)
(B.494)
= DU BU ⊤ + U BDU⊤ , where DU is derived from the differential of eigenvectors, DU = ( (σ1 In − X)+ V u1 · · · (σn In − X)+ V un ).
(B.495)
Substituting Eqs. (B.493) to (B.495) into Eq. (B.492) yields the differential of log−1 α .
B.9.8
Proof of Thm. 156
n , dALE } is isometric to the space Proof of Thm. 156. Obviously, the metric space {S++ S n endowed with the standard Euclidean distance. Therefore, the weighted Fréchet n mean of {Si } in S++ corresponds to the weighted Fréchet mean of associated points {logα (Si )} in S n . The weighted Fréchet means in Euclidean spaces are clearly the familiar weighted means.
B.9.9
Proof of Thm. 157
Proof of Thm. 157. As logα is a Riemannian isometry and S n is bi-invariant, ALEM is therefore bi-invariant. 384
Appendix B. Proofs
B.9.10
Proof of Thm. 158
Proof of Thm. 158. Following the notation in this proposition, we proceed as follows. The right-hand side can be rewritten as m X 1
β (FM(S1β , · · · Sm )) = log−1 α
m
β logα (Si )
i=1 m X
β = log−1 α "
!
= log−1 α
!
1 logα (Si ) m i=1 !#β m X 1 logα (Si ) m i=1
(B.496)
= (FM(S1 , · · · Sm ))β .
B.9.11
Proof of Thm. 159
Proof of Thm. 159. Recalling Eq. (6.24), Properties U1 and U2 obviously hold. When the SPD matrices {Ai }i≤n commute, we have FM({Ai }) =
Y
Ai
i
! n1
.
(B.497)
With Eq. (B.497), Properties V1–V4 can be easily proved.
B.9.12
Proof of Thm. 160
Proof of Thm. 160. Obviously, for a given SPD matrix S, logα (RSR⊤ ) = R logα (S)R⊤ ,
logα (s2 S) = U logα (s2 In ) + logα (Σ) U ⊤ ,
where S = U ΣU ⊤ is the eigendecomposition. These identities yield the result.
B.9.13
Proof of Thm. 161
Proof of Thm. 161. The three equations can be directly obtained. 385
(B.498) (B.499)
B.9. Adaptive Log-Euclidean Metrics
B.9.14
Proof of Thm. 163
Proof of Thm. 163. The input gradient follows from the Daleckii–Krein formula reviewed in Eqs. (2.90) and (2.91). Now, let us focus on the gradient with respect to A. Differentiating both sides of Eq. (6.31) gives (B.500)
d X = (∗) + U (d A ⊛ log(Σ)) U ⊤ ,
where (∗) denotes other terms involving d U and d Σ. According to the invariance of the first-order differential form, we have (B.501)
∇X L : d X
(B.502)
= ∇S L : d S + ∇X L : U (d A ⊛ log(Σ)) U ⊤ = ∇S L : d S + [U ⊤ (∇X L)U ] ⊛ log(Σ) : d A,
(B.503)
where A : B = tr(A⊤ B) is the Euclidean Frobenius inner product. From the second term on the RHS of Eq. (B.503), we can obtain the gradient with respect to A.
B.9.15
Proof of Thm. 164
Proof of Thm. 164. The derivation follows the same logic as Thm. 163. We only need to show the derivation of Eq. (6.35). Similarly to Thm. 163, we have the following:
d X = (∗) + U d A ⊛ diag
∇X L : d X
Σnn 11 aΣ 1 , · · · , an
−Σ A2
Σnn 11 = ∇S L : d S + [U ⊤ (∇X L)U ] ⊛ diag aΣ 1 , · · · , an
B.9.16
−Σ A2
U ⊤,
(B.504)
(B.505) : d A.
(B.506)
Proof of Thm. 167
Proof of Thm. 167. Following Nguyen [157], Nguyen and Yang [159], we first define gyrostructures under ALEM: P ⊕ALE Q = ExpP PTIn →P LogIn (Q) , 386
(B.507)
Appendix B. Proofs gyr[P, Q]R = (⊖(P ⊕ALE Q)) ⊕ALE (P ⊕ALE (Q ⊕ALE R)), t ⊙ALE P = ExpIn t LogIn (P ) , ⊖P = −1 ⊙ALE P = ExpIn − LogIn (P ) , ⟨P, Q⟩gyr = LogIn (P ), LogIn (Q) In , q ∥P ∥gyr = ⟨P, P ⟩gyr ,
(B.508) (B.509) (B.510) (B.511) (B.512) (B.513)
dgyr (P, Q) = ⊖P ⊕ALE Q gyr ,
n where P, Q, R ∈ S++ , and In is the identity matrix. The above operations are called gyroaddition, gyroautomorphism, scalar gyromultiplication, gyroinverse, gyroinner product, gyronorm, and gyrodistance. Simple computations show that Eq. (B.507) and Eq. (B.509) are exactly ⊕ALE and ⊙ALE in Thm. 152. As indicated by Thm. 152, n {S++ , ⊕ALE , ⊙ALE } forms a gyrovector space [158, Def. 1]. In the following proof, we use ⊙ALE and ⊕ALE .
The gyro MLR [159] under ALEM is defined as p(y = k | S) ¯ ∝ exp sign(⟨Ãk , LogPk (S)⟩Pk )∥Ãk ∥Pk d(S, HÃk ,Pk ) ,
(B.514)
n n ¯ H where Pk ∈ S++ and Ãk ∈ TPk S++ . d(S, Ãk ,Pk ) is the margin distance to the SPD hyperplane HÃk ,Pk , which is defined as ∗ ¯ H d(S, Ãk ,Pk ) = sin(∠SPk Q )dgyr (S, Pk ),
Q∗ =
argmax
(B.515) (B.516)
(cos(∠SPk Q)) ,
Q∈HÃ ,P \{Pk } k
cos(∠SPk Q) =
k
⊖Pk ⊕ALE Q, ⊖Pk ⊕ALE S gyr
∥⊖Pk ⊕ALE Q∥gyr ∥⊖Pk ⊕ALE S∥gyr
n HÃk ,Pk = {S ∈ S++ | ⟨LogPk S, Ãk ⟩Pk = 0}.
,
(B.517) (B.518)
Eqs. (B.515), (B.517) and (B.518) are called the SPD pseudo-gyrodistance, SPD gyrocosine, and SPD gyrohyperplane.
For simplicity, we further omit the subscript k in Pk and Ãk . Eq. (B.518) can be 387
B.9. Adaptive Log-Euclidean Metrics simplified: ⟨LogP S, Ã⟩P E D (1) (log (S) − log (P )) , Ã = log−1 α α α ∗,logα (P ) P E D (2) (log (S) − log (P )) , (log ) ( Ã) = (logα )∗,P ◦ log−1 α α α ∗,P α ∗,logα (P ) D E = logα (S) − logα (P ), (logα )∗,P (Ã) .
(B.519)
The above derivation comes from the following. (1) Eq. (6.18). (2) The definition of ALEM.
Similarly, a simple computation shows that Eq. (B.517) can also be simplified as ⟨− logα (P ) + logα (Q), − logα (P ) + logα (S)⟩ . ∥− logα (P ) + logα (Q)∥F ∥− logα (P ) + logα (S)∥F
(B.520)
Together with Eqs. (B.519) and (B.520), Eq. (B.515) is equivalent to the distance to the hyperplane in the Euclidean space. Therefore, Eq. (B.515) has a closed-form solution: ¯ H ) d(S, Ã,P =
logα (S) − logα (P ), Ā Ā F
=
logα (S) − logα (P ), Ā
(B.521) ,
à P
where Ā = (logα )∗,P (Ã). Substituting Eq. (B.521) into Eq. (B.514) yields the claimed result.
B.9.17
Proof of Thm. 165
Proof of Thm. 165. To derive the ALEM-specific positive-scalar update, consider the general RSGD update reviewed in Sec. 2.7. For a minimization parameter w on an n-dimensional smooth connected Riemannian manifold M, we have w(t+1) = Expw(t) (−γ (t) πw(t) (∇w(t) L)), 388
(B.522)
Appendix B. Proofs where Expw (·) : Tw M → M is the Riemannian exponential map, which maps a tangent vector at w back into the manifold M, and πw (·) : Rn → Tw M is the projection operator, projecting an ambient Euclidean vector into the tangent space at w. In the n , X ∈ Rn×n , and V ∈ S n , the exponential case of the SPD manifold, for all S ∈ S++ map and projection operator are formulated as follows: X + X⊤ S, 2 ExpS (V ) = S 1/2 exp(S −1/2 V S −1/2 )S 1/2 , πS (X) = S
(B.523) (B.524)
where exp(·) is the matrix exponential. For more details about Eq. (B.523) and Eq. (B.524), see Yger [225] and Amari [3]. Substituting Eq. (B.523) and Eq. (B.524) into Eq. (B.522) immediately yields Eq. (6.36).
B.9.18
Proof of Thm. 166
Proof of Thm. 166. Without loss of generality, we focus on the equivalence between b = B11 and a = a1 . Note that b is essentially expressed as b = log(a). Suppose b(t) = log a(t) . Then, we have ∂ log(a) ∂a a(t) 1 = ∇b(t) L (t) . a
∇a(t) L = ∇b(t) L
(B.525)
By Eq. (6.36), log a(t+1) is
(t) (t) log a(t+1) = log a(t) e−γ a ∇a(t) L = log a(t) − γ (t) a(t) ∇a(t) L = log a(t) − γ (t) a(t) ∇b(t) L/a(t) = log a(t) − γ (t) ∇b(t) L
(B.526)
= b(t) − γ (t) ∇b(t) L.
The last row is the ESGD update formula for b. Therefore, if b(0) = log a(0) , the two optimization procedures yield equivalent iterates throughout training. 389
B.10. Product Cholesky Metrics
B.10
Product Cholesky Metrics
B.10.1
Proof of Thm. 169
Proof. Since θ-DPM is the product metric of {LT0 (n), g E } and n copies of {R++ , g θ-E }, we first show the Riemannian operators on {R++ , g θ-E }. We can then readily obtain the Riemannian operators on {Ln++ , g θ-DE } by the principles of product metrics.
As shown by Thanwerdas and Pennec [194], g θ-E is the pullback metric of g E by the power function Pθ (·) and scaled by θ12 , expressed as g θ-E = θ12 P∗θ g E . Besides, as constant scaling does not change the Christoffel symbols, the geodesic, Riemannian logarithm and exponential maps, and parallel transport along a geodesic remain the same under g θ-E and P∗θ g E . These Riemannian operators under P∗θ g E can be obtained by the properties of Riemannian isometries (Thm. 33). Specifically, given p, q ∈ R++ and w, v ∈ Tp R++ , we have the following: (Pθ )∗,p (v) = θpθ−1 v, 1 gpθ-E (v, w) = 2 g E (Pθ )∗,p (v), (Pθ )∗,p (w) = ⟨pθ−1 v, pθ−1 w⟩ = p2(θ−1) vw, θ γ(p,v) (t) = P−1 P (p) + t (P ) (v) θ θ θ ∗,p
(B.527) (B.528) (B.529)
= (pθ + tθpθ−1 v) θ
1
(B.530)
1 θ
(B.531)
= p(1 + tθp−1 v) , with t ∈ {t ∈ R | 1 + tθp−1 v ∈ R++ },
Logp (q) = (Pθ )−1 ∗,p (Pθ (q) − Pθ (p))
! θ 1 1−θ θ 1 q = p (q − pθ ) = p −1 , θ θ p q 1−θ −1 PTp→q (v) = (Pθ )∗,q (Pθ )∗,p (v) = v. p
(B.532) (B.533) (B.534)
The geodesic distance between p and q under g θ-E is given by d2 (p, q) = gpθ-E (Logp (q), Logp (q)) =
1 θ (q − pθ )2 . 2 θ
(B.535)
N The weighted Fréchet mean (WFM) of {pi ∈ R++ }N i=1 with weights {wi }i=1 satisfying P wi > 0 for all i and i wi = 1 under g θ-E is defined as
WFM({wi }, {pi }) = argmin p∈R++
390
N X i=1
wi d2 (p, pi ).
(B.536)
Appendix B. Proofs Obviously, the WFM of {pi } under g θ-E is the same as the one under P∗θ g E . Due to the isometry of P∗θ g E to g E , the WFM of {pi } under g θ-E can be calculated as 1 XN E θ θ WFM({wi }, {pi }) = P−1 WFM ({w }, {P (p )}) = , w p i θ i i i θ i=1
(B.537)
where WFME in Eq. (B.537) is the Euclidean WFM, which is the familiar weighted average. So far, we have obtained all the necessary Riemannian operators on {R++ , g θ-E }. Combining these results with the Euclidean space LT0 (n) yields the results in the theorem.
B.10.2
Proof of Thm. 170
Proof. As in the proof of Thm. 169, we only need to show the Riemannian operators on {R++ , g m-BW } with m ∈ R++ . Expressions for the Riemannian operators under GBWM can be found in Han et al. [93]. Here, we further simplify the associated expressions for the one-dimensional case. Specifically, given p, q ∈ R++ and w, v ∈ Tp R++ , we have the following: v , 2mp vw 1 , gpm-BW (v, w) = ⟨Lp,m (v), w⟩ = 2 4mp
(B.538)
Lp,m (v) =
(B.539)
(tv)2 tv 2 = p(1 + ), 4p 2p ! 12 1 1 q Logp (q) = 2 m(m−2 pq) 2 − p = 2 (pq) 2 − p = 2p −1 , p γ(p,v) (t) = p + tv + Lp,m (tv)2 m2 p = p + tv +
2 1 1 4 (pq) 2 − p 4mp 3 1 1 = pq − 2p 2 q 2 + p2 mp 1 1 −1 2 2 =m q − 2p q + p 1 2 1 1 . = m− 2 q 2 − p 2
(B.540) (B.541)
d2 (p, q) = gpm-BW (Logp (q), Logp (q)) =
(B.542)
n As shown by Han et al. [93], GBWM on S++ is the pullback metric of BWM by 1 1 n n 1 π(S) = M − 2 SM − 2 for all S ∈ S++ with M ∈ S++ . For the specific R++ ∼ , the = S++
391
B.10. Product Cholesky Metrics isometry is simplified as π(p) = m−1 p, ∀p ∈ R++ with m ∈ R++ .
(B.543)
The geodesic γ e(p,v) (t) under BWM on R++ exists in the interval {t o∈ R | 1+ n v tLp (v) ∈ R++ } [145], which can be simplified as t ∈ R | 1 + t 2p ∈ R++ . Therefore, m-BW the geodesic γ(p,v) (t) under g exists in the interval: v π∗,p (v) ∈ R++ = t ∈ R | 1 + t ∈ R++ . t∈R|1+t 2π(p) 2p
(B.544)
For BWM, the parallel transport is [196, Tab. 6] f p→q (v) = PT
12 q v. p
(B.545)
Therefore, the parallel transport on {R++ , g m-BW } is −1 PTp→q (v) = π∗,q
! 12 π(q) f π(p)→π(q) (π∗,p (v)) = π −1 PT π∗,p (v) ∗,q π(p) 12 q = v. p
(B.546)
Finally, we show the WFM on {R++ , g m-BW }. Given {pi ∈ R++ }N i=1 with weights P N m-BW } is {wi }N i=1 satisfying wi > 0 for all i and i=1 wi = 1, the WFM on {R++ , g WFM({wi }, {pi }) = argmin p∈R++
= argmin p∈R++
= argmin p∈R++
= argmin p∈R++
N X
wi d2 (p, pi )
i=1
N X
i=1 N X i=1 N X i=1
1 1 2 wi m−1 p 2 − pi2
1 1 2 wi p 2 − pi2
392
1 2
wi p − 2p pi
= argmin p − 2p p∈R++
1 2
1 2
N X i=1
(B.547)
1
wi pi2 .
Appendix B. Proofs 1
Let f (p) = p − 2p 2
PN
1
2 i=1 wi pi . Then, the first- and second-order derivatives are
X 1 1 df = 1 − p− 2 wi pi2 , i dp 2 X 1 d f 1 3 wi pi 2 > 0, = p− 2 2 i dp 2
(B.548) ∀p ∈ R++ .
(B.549)
Therefore, the optimal solution can be obtained by setting Eq. (B.548) equal to 0: X 1 2 df = 0 ⇒ p∗ = wi pi2 . i dp
(B.550)
Combining the above results with the Euclidean geometry on LT0 (n), one can readily obtain the results.
B.10.3
Proof of Thm. 172
Proof. The differential of DPowθ at L ∈ Diag+ (n) is (DPowθ )∗,L (V) = θLθ−1 V,
∀V ∈ TL Diag+ (n).
(B.551)
Substituting Eq. (B.551) into Thm. 171 yields the result.
B.10.4
Proof of Thm. 173
We first present a useful lemma. Lemma 205. The Riemannian exponential and logarithmic maps and parallel transport along the geodesic are the same under (θ, M)-DBWM and θ/2-DPM. Proof. This follows directly from Thms. 169 and 170. Now we begin to prove Thm. 173. Proof. We first derive the expressions for DPM with a generic nonzero parameter θ. According to Thm. 205, the expressions for (θ, M)-DBWM then follow by replacing θ with θ/2. We omit the superscript C for simplicity. 393
B.10. Product Cholesky Metrics For X ∈ TIn Ln++ , we have the following: 1 θ L − In , θ PTIn →L (X) = ⌊X⌋ + L1−θ X, LogIn (L) = ⌊L⌋ +
1 θ
ExpIn X = ⌊X⌋ + (In + θX) .
(B.552) (B.553) (B.554)
For the binary operation, substituting Eqs. (B.552) and (B.553) into Eq. (2.77) gives L ⊕ K = ExpL PTIn →L LogIn (K) 1 θ = ExpL PTIn →L ⌊K⌋ + K − In θ 1 1−θ θ K − In = ExpL ⌊K⌋ + L θ θ1 1 1−θ θ −1 = ⌊L⌋ + ⌊K⌋ + L In + θL L K − In θ 1 = ⌊L⌋ + ⌊K⌋ + Lθ + Kθ − In θ .
(B.555)
In the third row of Eq. (B.555), the well-definedness of the exponential map requires
1 1−θ θ L+θ L K − In ∈ Diag+ (n) ⇔ Lθ + Kθ − In ∈ Diag+ (n). θ
(B.556)
For the gyromultiplication, substituting Eqs. (B.552) and (B.554) into Eq. (2.78) gives t ⊙ L = ExpIn (t LogIn (L)) t θ L − In = ExpIn t⌊L⌋ + θ (B.557) θ1 θ = t⌊L⌋ + In + t L − In 1 = t⌊L⌋ + tLθ + (1 − t)In θ . In the second row of Eq. (B.557), the exponential map requires In + θ
t θ L − In θ
∈ Diag+ (n) ⇔ tLθ + (1 − t)In ∈ Diag+ (n).
394
(B.558)
Appendix B. Proofs
B.10.5
Proof of Thm. 174
Proof. In this proof, we assume L, K, J ∈ Ln++ and s, t ∈ R, with all gyro operations satisfying the conditions in Thm. 173. By Thm. 205, we only need to prove the case of θ-DPM. For simplicity, we omit the superscript C. Axiom (G1). Eq. (2.77) implies that the identity element is the identity matrix. Axiom (G2). We define the inverse element of L as 1 ⊖L = −1 ⊙ L = −⌊L⌋ + 2In − Lθ θ .
(B.559)
Simple computations show that ⊖L ⊕ L = In .
Axiom (G3). Gyroaddition in Thm. 173 indicates that
1 L ⊕ (K ⊕ J) = (L ⊕ K) ⊕ J = ⌊L⌋ + ⌊K⌋ + ⌊J⌋ + Lθ + Kθ + Jθ − 2In θ . (B.560)
Therefore the gyroautomorphism is the identity map, i.e., gyr[L, K] = id. Axiom (G4). This is a direct corollary of (G3). Gyrocommutative law. Gyroaddition in Thm. 173 indicates that L ⊕ K = K ⊕ L.
(B.561)
Axiom (V1). This follows from gyromultiplication in Thm. 173 and Eq. (B.559). Axiom (V2). 1 (s + t) ⊙ L = (s + t)⌊L⌋ + (s + t)Lθ + (1 − (s + t))In θ
1 = s⌊L⌋ + t⌊L⌋ + sLθ + (1 − s)In + tLθ + (1 − t)In − In θ
(B.562)
= (s ⊙ L) ⊕ (t ⊙ L).
Axiom (V3). 1 (st) ⊙ L = (st)⌊L⌋ + stLθ + (1 − st)In θ 1 = (st)⌊L⌋ + s tLθ + (1 − t)In + (1 − s)In θ h θ1 i θ = s ⊙ t⌊L⌋ + tL + (1 − t)In
(B.563)
= s ⊙ (t ⊙ L).
Axioms (V4) and (V5). These two axioms can be directly obtained, as gyroautomorphisms are all identity maps. 395
B.10. Product Cholesky Metrics
B.10.6
Proof of Thm. 177
Proof. According to Nguyen and Yang [159, Thm. 2.4], gyrovector operations are preserved under Riemannian isometries. Moreover, the Cholesky decomposition is a Riemannian isometry: n Chol : {S++ , g S } → {Ln++ , g C }. (B.564)
B.10.7
Proof of Thm. 179
Proof. Substituting the associated operators in Sec. 6.3.4 into Thm. 195 yields the results.
396