ConceptioArchivearXiv CS
arXiv CSopen access

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Pengzhen Chen 1 2 3 Yanwei Liu 1 3 Xiaoyan Gu 1 2 3 Antonios Argyriou 4 Wu Liu 5 Weiping Wang 1

1. Introduction

arXiv:2605.26702v1 [cs.CV] 26 May 2026

Abstract

The rapid proliferation of AI-generated content (AIGC) has democratized the synthesis of high-fidelity 360◦ panoramic imagery (Wang et al., 2025). Recent advances, including PanFusion (Zhang et al., 2024a) and text-to-360 models (Zhou et al., 2025b), enable the creation of immersive spherical environments from directly natural language prompts. These technologies are accelerating applications across virtual reality and the Metaverse (Shinde et al., 2023; Zhou et al., 2025a; Tukur et al., 2024), while critically, providing foundational data for World Models (Lu et al., 2025b; Li et al., 2025) and Embodied AI agents (Zheng et al., 2025; Wu et al., 2025). However, this ease of generation simultaneously exacerbates risks regarding copyright infringement and unauthorized content redistribution. Digital watermarking offers a principled mechanism for provenance tracking by embedding imperceptible identity signals into media. Yet, despite substantial progress in deep watermarking for planar images, extending these techniques to panoramic imagery remains fundamentally challenging due to its distinct geometric structure.

Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations. As panoramas are defined on the sphere, they naturally transform under the action of SO(3), rendering conventional planar representations and augmentation-based robustness strategies inadequate and devoid of theoretical guarantees. To address this, we formulate panoramas as spherical signals and leverage SO(3) representation theory to derive provably rotation-invariant descriptors. While spherical harmonic coefficients transform equivariantly under rotations, the natural invariant constructions are typically limited to zeroth-order statistics which eliminate directional information and severely constrain embedding capacity. In this work, we introduce a principled third-order invariant construction by coupling higher-order SO(3) irreducible representations via tensor products and projecting onto the trivial representation. This yields a spherical invariant bispectrum that preserves phase information while remaining strictly rotation-invariant. Leveraging this property, we embed watermarks into higher-order spherical harmonic coefficients and recover them from invariant bispectral scalars, enabling reliable extraction under arbitrary 3D rotations. We provide a theoretical proof of SO(3) invariance for it and demonstrate experimentally its near-perfect robustness to continuous rotations while maintaining high visual fidelity. Code is available here.

A panoramic image is not a signal defined on the Euclidean plane R2 , but rather a function residing on the unit sphere S2 . During consumption, users can freely alter their viewing direction via head-mounted displays, an interaction mathematically modeled by the action of the three-dimensional rotation group, SO(3) (Cohen et al., 2018), on the spherical signal. When represented via standard Equirectangular Projection (ERP), such rotations induce highly non-linear and latitude-dependent distortions, including severe polar stretching and large-scale texture displacement (Makadia & Daniilidis, 2003). Consequently, conventional watermark extraction methods, which rely on pixel-grid alignment or local convolutional consistency in Euclidean space, become inherently unstable under global rotations.

1 Institute of Information Engineering, Chinese Academy of Sciences 2 School of Cyber Security, University of Chinese Academy of Sciences 3 State Key Laboratory of Cyberspace Security Defense 4 University of Thessaly 5 University of Science and Technology of China. Correspondence to: Yanwei Liu <[email protected]>, Xiaoyan Gu <[email protected]>.

Current deep watermarking frameworks are(Lu et al., 2025a; Hu et al., 2024; Bui et al., 2023; Zhang et al., 2024b; Tancik et al., 2020; Chen et al., 2025a; 2026; Wu et al., 2023; Li et al., 2026) predominantly built upon Convolutional Neural Networks (CNNs) that exploit translational equivariance (Ben Jabra & Ben Farah, 2024). While effective against perturbations like Gaussian noise or JPEG compres-

Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).

1

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

sion, they lack intrinsic robustness to geometric transformations. Prior studies typically resort to data augmentation as a heuristic remedy by injecting distorted samples during training to force the network to memorize attacked variations. However, we argue that this strategy is insufficient for theoretically trusted traceability especially for spherical data. Specifically, since SO(3) is a continuous group describing infinite possible rotations (Esteves et al., 2018), each resulting in a distinct non-linear and complex pixel-space distortion. It is infeasible to exhaustively cover the transformation space via finite augmentation. Robustness obtained in this manner relies on memorization rather than geometric consistency, offering no theoretical guarantees and incurring significant training overhead. This fundamental geometric mismatch between the translational equivariance of planar CNNs and the rotational symmetry of spherical signals indicates that robust panoramic watermarking cannot be achieved through projection-space heuristics. Instead, it necessitates representations that explicitly respect the underlying SO(3) symmetry.

Signal 𝒇 ∈ 𝑺𝟐

Higher-Order Irreps

𝒘 ∈ {𝟎, 𝟏}𝒌

𝑽𝒍𝟏 ⨂𝑽𝒍𝟐 ⨂ 𝑽𝒍𝟑 Clebsch−Gordan

𝒘 ∈ {𝟎, 𝟏}𝒌

𝑽𝟎 Spherical Bispectrum Invariant

Figure 1. The grounding theory of TRIAD. By coupling higherorder spherical harmonics representations and projecting onto the trivial representation, we obtain a scalar invariant (the spherical bispectrum) that retains phase information while remaining invariant to rotations, enabling information embedding in sensitive equivariant coefficients with reliable invariant recovery.

tion theory rather than empirical heuristics.

To address these challenges, we propose TRIAD, a theoretically grounded framework for provably robust watermarking that delves into the natural spherical structure of panoramic images. As shown in Figure 1, we model images using a Spherical Harmonics (SH) expansion, which provides a compact and continuous parameterization of the sphere, avoiding the non-uniform sampling artifacts inherent to ERP. In the SH domain, the zeroth-order SH coefficient c0 is strictly rotation-invariant (Kondor, 2025). However, as c0 corresponds to the global average (DC component) of the signal (Sloan, 2008), embedding watermarks in this term would cause noticeable shifts in global luminosity and color, severely degrading perceptual quality. Therefore, we propose a novel embedding strategy based on the third-order spherical bispectrum. Specifically, we leverage higher-order spherical harmonic coefficients to carry the watermark information, which offers substantially greater capacity and improved imperceptibility. To recover this information robustly, we construct a third-order tensor product that couples three SO(3) irreducible representations. From a representation-theoretic perspective, decomposing this tensor product yields a trivial (l = 0) component corresponding to the bispectrum. This resulting scalar is mathematically guaranteed to be invariant under SO(3) rotations, enabling reliable recovery of watermark information embedded in rotation-sensitive higher-order coefficients while preserving strict rotation invariance at extraction time.

2. Spherical Bispectrum Invariant. We introduce a novel embedding-to-extraction mechanism based on the spherical bispectrum. By coupling higher-order spherical harmonic coefficients via tensor products, we derive a rotationinvariant scalar that carries messages embedded in higherorder coefficients. 3. TRIAD Framework. We propose TRIAD, an end-to-end framework that seamlessly integrates equivariant operations and invariant extraction. Extensive experiments demonstrate that TRIAD achieves superior robustness against arbitrary 360◦ rotations while maintaining high visual fidelity.

2. Related Work Robust Watermarking for Panoramic and 3D Data. Digital watermarking has evolved from traditional frequency-domain methods (DCT/DWT) (Al-Haj, 2007) to deep learning-based frameworks (Zhu et al., 2018; Luo et al., 2020). However, standard CNN-based methods fail to generalize to panoramic images due to the severe geometric distortions inherent in equirectangular projections. While several schemes for 360◦ images have been proposed (Liu et al., 2021), they typically operate in projection space and lack synchronization mechanisms to handle 3D rotations. Similarly, in the 3D data domain, recent works have explored watermarking for meshes (Narendra et al., 2024), point clouds (Zaman et al., 2025), and emerging 3D Gaussian Splatting representations (Chen et al., 2025b), employing techniques such as salient point learning, SVD-based embedding, and neural feature extraction. Despite differences in representation, these approaches share a common reliance on data augmentation during training, which only

Our main contributions are summarized as follows: 1. Provable Geometric Robustness. We identify the theoretical limitations of augmentation-based robustness for spherical data and propose a watermarking framework with certified SO(3) invariance, grounded in group representa-

2

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

3.1. Rotation Group SO(3) and Spherical Signals

provides empirical robustness, and often degrades under unseen transformations, offering no theoretical guarantees. In contrast, our approach exploits the algebraic structure of the rotation group SO(3), enabling mathematically strict rotation-invariance without exhaustive augmentation.

Three-dimensional rotations are described by the special orthogonal group SO(3), consisting of all 3 × 3 orthogonal matrices with determinant one. For a signal f (ω) defined on the unit sphere S2 , the action of a rotation R ∈ SO(3) is defined as: (RR f )(ω) = f (R−1 ω), (1)

Spherical CNNs and SO(3)-Equivariance. To handle spherical signals without projection-induced distortion, Cohen et al. (Cohen et al., 2018) introduced Spherical CNNs based on the Fourier theorem on SO(3). Subsequent studies generalized this via G-steerable convolutions (Weiler et al., 2018) and tensor field networks (Thomas et al., 2018), establishing the foundation for modern equivariant libraries like e3nn (Geiger & Smidt, 2022). These frameworks leverage Clebsch-Gordan (CG) coefficients (Kondor et al., 2018) to perform tensor products, ensuring that learned feature fields transform predictably under rotations, i.e., equivariantly. While highly effective for discriminative tasks including molecular modeling (Batzner et al., 2022; Kohler et al., 2025) and image processing (Ocampo et al., 2023; Esteves et al., 2018), their application to generative watermarking remains unexplored.

which corresponds to a rigid rotation of the underlying spherical domain. Unlike planar images, spherical signals do not admit a global notion of translation. Consequently, the translational equivariance exploited by conventional CNNs has no natural analogue on S2 . When spherical data is represented using planar projections such as equirectangular projection (ERP), rotations in SO(3) induce highly non-uniform, latitudedependent distortions. As a result, the locality and weightsharing assumptions underlying standard convolutional kernels are fundamentally violated, motivating the need for representations that explicitly respect rotational symmetry. 3.2. Spherical Harmonics

Higher-Order Invariants and Bispectrum. Invariant constructions are fundamental for handling geometric transformations, as they enable stable signal representations independent of group actions. A widely adopted approach is the power spectrum (Kazhdan et al., 2003), which achieves invariance by discarding phase information. While effective for recognition tasks (Poulenard et al., 2019), this phase elimination fundamentally limits its capacity of hiding messages. Addressing this requires higher-order statistics. In classical signal processing, the bispectrum (triple correlation) is known to retain phase information (Nikias & Mendel, 1993). Recently, higher-order invariant constructions have been revisited in geometric deep learning, where third-order tensor contractions are employed to capture complex relational structures (Mataigne et al., 2024; Iglesias Martı́nez et al., 2024; Sanborn & Miolane, 2023). Nevertheless, these studies primarily focus on characterizing fixed physical systems (e.g., atom potentials). Our work is the first to construct a learnable spherical bispectrum-based framework specifically for watermarking. By leveraging the phase-preserving property of third-order coupling, we enable high-capacity watermark embedding in rotation-sensitive subspace while extracting from strictly SO(3)-invariant bispectral scalars.

Spherical harmonics (SH) form an orthonormal basis for the space of square-integrable functions on the sphere, L2 (S2 ). Using spherical coordinates ω = (θ, ϕ), any spherical signal f (ω) can be expanded as: f (ω) =

lX max

l X

m cm l Yl (ω),

(2)

l=0 m=−l

where l denotes the frequency degree and m indexes angular variation within each degree. A key property of spherical harmonics is their structured response to rotations. Under a rotation R, the SH coefficients transform linearly as: c′l = Dl (R) cl ,

(3)

where cl ∈ C2l+1 collects the coefficients at degree l, and Dl (R) denotes the corresponding Wigner-D matrix. Importantly, coefficients of different degrees do not mix. This block-diagonal transformation law constitutes SO(3)equivariance and forms the algebraic backbone of spherical signal processing and equivariant neural networks. 3.3. Irreducible Representations and Tensor Products

3. Theoretical Background

In equivariant learning frameworks such as e3nn, features are organized as direct sums of irreducible representations (irreps) of SO(3). Each irrep is labeled by its degree l (and parity), with l = 0 corresponding to scalars and l ≥ 1 corresponding to vectors.

This section reviews the mathematical foundations underlying the proposed SO(3)-invariant watermarking framework for spherical signals. We briefly introduce rotation group actions, spherical harmonics, and higher-order equivariant constructions, focusing exclusively on the structures essential to our method.

The interaction between equivariant features is governed by tensor products. Given two irreducible representations Vl1 3

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

and Vl2 of degrees l1 and l2 , their tensor product decomposes into a direct sum of irreps: Vl1 ⊗ Vl2 ∼ =

lM 1 +l2

VJ ,

perturbations embedded in the higher-order spherical harmonic coefficients induce non-vanishing variations in I, satisfying the necessary condition for extraction from this rotation-invariant quantity.

(4)

J=|l1 −l2 |

Proof. See Appendix A.1 for the detailed derivation.

where the decomposition is mediated by Clebsch–Gordan (CG) coefficients. This operation enables controlled coupling across frequency degrees while preserving equivariance, and crucially allows the construction of invariant features by projecting onto the trivial irrep V0 .

A critical property of Theorem 4.1 is the spectral isolation provided by the SH basis. Due to the orthogonality of spherical harmonics, operations within specific subspaces Vl1 , Vl2 , Vl3 do not introduce interference in unrelated frequency bands. As shown in Figure 2, this enables targeted watermark injection in higher-order components while guaranteeing strict rotation invariance via bispectral projection.

4. Methodology 4.1. Design Principles

4.2. Watermark Embedding

We begin by formalizing the construction of a rotationinvariant watermark derived from a panoramic signal. Let f : S2 → RC denote a panoramic image represented in spherical harmonics (SH), with coefficients cm l ∈ Vl , where Vl denotes the degree-l irreducible representation of SO(3).

Given an input panorama x ∈ RH×2H and a watermark vector w ∈ {0, 1}k , the encoder embeds watermarks by operating directly in the spherical harmonic (SH) domain. We first lift the input panorama to its spectral representation by computing SH coefficients up to degree lmax :

The trivial representation V0 yields rotation-invariant statistics by construction. However it corresponds to global, low-frequency image characteristics, and any perturbation of this component induces perceptually severe artifacts. Intuitively, valid invariant watermarking requires embedding information beyond zeroth-order statistics while retaining the ability to extract it from an invariant quantity. To this end, we exploit third-order correlations among higher-order SH coefficients. Specifically, we consider the tensor product of three irreducible representations: M Vl1 ⊗ Vl2 ⊗ Vl3 = Vl . (5)

max c = {cl }ll=0 ,

cl ∈ Vl ,

(8)

where each irreducible component transforms equivariantly and independently under rotation. Watermark embedding is restricted to a selected set of higher-order subspaces: M Vembed = Vl , l > 0, (9) l∈Lembed

where Lembed denotes the selected embedding degrees. An SO(3)-equivariant backbone Φeq is applied to project the raw SH coefficients to selected structured spectral features

l

By projecting the representation onto the trivial subspace V0 , we obtain a scalar invariant derived exclusively from higher-order coefficients, thereby decoupling watermark embedding from invariant extraction. Concretely, the resulting bispectrum invariant I is defined as: X X 2 m3 I= Cl0,0 cm1 cm (6) l2 cl3 , 1 m1 l2 m2 l3 m3 l1

u = Φeq (c),

u ∈ Vembed .

(10)

Due to the orthogonality of spherical harmonic basis, operations within Vembed do not affect other frequency components. Watermark information is embedded into Vembed by conditioning the spectral features on the watermark vector from w. Specifically, watermark w is mapped to scalar features transforming under the trivial representation V0 , and injected through an SO(3)-equivariant interaction with the spectral features. Specifically, we employ a parameterized equivariant tensor product TPϑ1 (detailed in Appendix C.2) to produce a fused update:

l1 ,l2 ,l3 m1 ,m2 ,m3 0,0 where C... denotes the Clebsch–Gordan coupling coefficients projecting onto the trivial representation and can be computed via Wigner 3-j symbols: r (2l1 + 1)(2l2 + 1)(2l3 + 1) 0,0 Cl1 m1 l2 m2 l3 m3 =  4π  l l l l1 l2 l3 × 1 2 3 . 0 0 0 m1 m2 m3 (7)

TPϑ1 : Vembed ⊗V0 → Vembed . (11) By construction, ∆u lies in the same irreducible subspaces as u and therefore preserves the SO(3) transformation behavior of the spectral representations. ∆u = TPϑ1 (u, w)|Vembed ,

Theorem 4.1. The bispectrum-based scalar I defined above is invariant under arbitrary SO(3) rotations. Furthermore, 4

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

(a) Embedding: Higher-Order SH Injection Input Panorama 𝐱

SH Coefficients {𝒄𝒍 }

Watermark 𝒘 to 𝑽𝟎 transform

(b) Extraction: Mapping to Bispectrum Invariants Modified features

𝑢 + Δ𝑢

Third-Order Bispectrum Projection (𝑽𝒍𝟏 ⊗ 𝑽𝒍𝟐 ⊗ 𝑽𝒍𝟑 → 𝑽𝟎 )

Rotated Watermarked Input

𝑉𝑙𝑚𝑎𝑥

ෝ ∈ 𝑽𝒆𝒎𝒃𝒆𝒅 𝒖

SHT

𝑉2

Selected Higher-Order Subspaces 𝑽𝒆𝒎𝒃𝒆𝒅

SHT Equivariant Tensor Product Interaction TP𝝑𝟏

Watermarked SH Coeffs {෤𝒄𝒍 }

𝑉1

Watermarked Panorama ෥ 𝒙

𝑉0

SO(3)-Equivariant Backbone

SO(3)-Equivariant Backbone 𝜱𝒆𝒒

Spherical Signal 𝑳(𝜽, 𝝋)

… TP𝝑𝟐

… TP𝝑𝟑

Tensor Clebsch-Gorden Product Coupling & Projection to Coupling (𝑪𝟎,𝟎 … ) Scalar Rotation-Invariant Scalar 𝑰

MLP Decoder Perceptual ISHT Module

Recovered Watermark 𝒘 ෝ

Figure 2. Framework of TRIAD. (a) Given an input panorama, we first represent it in the spherical harmonics (SH) domain and process the resulting coefficients with an SO(3)-equivariant backbone Φeq . Watermark information is embedded by modifying selected higher-order SH subspaces, which preserves equivariance and perceptual fidelity under rotations. (b) During extraction, the watermarked SH coefficients are coupled through a third-order tensor product and projected onto the trivial representation using Clebsch–Gordan coefficients. This operation produces a rotation-invariant scalar (the spherical bispectrum), from which the embedded watermark can be reliably extracted regardless of the panorama’s orientation.

The modified spectral features are computed by u + ∆u. These features are projected back to the full SH coefficient space and transformed to the spatial domain via the inverse spherical harmonic transform (ISHT), yielding a residual image ∆x. To ensure imperceptibility, the residual is modulated by a perceptual module (Appendix C.1), which is composed of a learnable mask Mperc (x) together with a geometric prior from the structure characteristics of ERP Mgeo , and added to the original panorama to produce the final watermarked output: x̃ = x + Mperc (x) ⊙ Mgeo ⊙ ∆x.

equivariant tensor products: h = TPϑ2 (û, û)|Vembed ,

(12)

These invariant features are aggregated and mapped through a lightweight MLP network to produce the recovered watermark ŵ. Since extraction relies exclusively on projections onto the trivial representation V0 , the decoding process is provably invariant to arbitrary 3D rotations.

Given a watermarked panorama x̃, the decoder recovers the embedded watermark by extracting rotation-invariant thirdorder statistics from its spherical harmonic representation. The panorama is first lifted to the SH domain up to lmax : c̃l ∈ Vl .

(14)

where TPϑ3 constrains the output subspace to the trivial representation V0 . This two-stage construction is equivalent to forming a third-order tensor product TP(û, û, û) whose output is projected onto the trivial representation, yielding a learnable bispectrum invariant that preserves the sign of the embedded watermark.

4.3. Extraction

max c̃ = {c̃l }ll=0 ,

z = TPϑ3 (h, û)|V0 ,

4.4. Loss Functions The network is trained end-to-end using a weighted combination of image fidelity loss and watermark extraction loss:

(13)

These coefficients are projected onto the same embedding space Vembed and processed by an SO(3)-equivariant backbone, producing spectral features û ∈ Vembed that contain the watermark-induced perturbations. Subsequently, to construct rotation-invariant descriptors, the decoder computes third-order equivariant statistics of the spectral features. Specifically, we apply two successive parameterized

Ltotal = λm LM SE (x, x̃) + λbce LBCE (w, ŵ),

(15)

where LM SE is Mean Squared Error (MSE) loss to ensure visual quality, and LBCE is the binary cross-entropy loss for message recovery. λm and λbce are the weighted factors. 5

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

5. Experiments

methods exhibit severe performance degradation once the input is rotated, with bit accuracy collapsing to near-random guessing. This reflects the fundamental mismatch between spherical rotational symmetries and planar convolutional architectures, whose inductive bias is translational equivariance. In ERP representations, any SO(3) rotation can induce highly non-linear, latitude-dependent distortions that cannot be compensated for by local convolutional filters.

5.1. Experimental Setup Datasets. We utilize two publicly available panoramic datasets: panoContext (Zhang et al., 2014) and SUN360 (Xiao et al., 2012). We randomly select 10,000 panoramas for training and 2,000 for testing. All panoramas are resized to 512 × 1024 equirectangular format for evaluation. For baseline methods that do not support this resolution, we follow the resolution scaling strategy from TrustMark (Bui et al., 2023) to interpolate the watermark strength (see Appendix C.7), which has been shown to preserve watermarking performance.

While rotation-based data augmentation partially improves robustness at specific angles encountered during training, the resulting performance remains highly irregular and angle-dependent. Crucially, SO(3) is a continuous group with infinitely many possible elements, and any augmentation strategy can only cover a finite subset of this space. Consequently, models trained with augmentation remain vulnerable to unseen rotations. The oscillatory performance observed across rotation angles indicates that such robustness arises from empirical memorization rather than principled invariance, and does not generalize uniformly over SO(3).

Implementation Details. The SO(3)-equivariant backbone is implemented using e3nn (Geiger & Smidt, 2022), consisting of 2 layers of Gated Blocks operating on identical irreducible representations. Spherical harmonic transform is computed with a cutoff degree of lmax = 16. We set Lembed = {6, 8, 14}, corresponding to the subspace Vembed ∈ {128 × 6e, 128 × 8e, 64 × 14e}, unless otherwise specified. The network is trained using the Adam optimizer with a learning rate of 10−4 for 300 epochs. The watermark length is set to k = 32 bits. The experiments are conducted on a NVIDIA A100 GPU. The loss weights λBCE is set to 10 and λm is initially set to 1 and linearly increased to 20 over the second 100 epochs.

Moreover, aggressive augmentation introduces an additional trade-off between robustness and imperceptibility. To maintain watermark extractability under larger transformation uncertainty, baseline methods must increase embedding strength, which directly degrades perceptual quality, as shown in Appendix Table 7.

Baselines. We compare TRIAD against the following opensourced methods: StegaStamp (Tancik et al., 2020), SepMark (Wu et al., 2023), TrustMark (Bui et al., 2023), EditGuard (Zhang et al., 2024b), Robust-Wide (Hu et al., 2024), VINE (Lu et al., 2025a). All methods are evaluated using their released checkpoints unless specified.

In contrast, TRIAD maintains near-perfect bit accuracy across all rotation angles without any data augmentation. This is a direct consequence of the theoretically guaranteed SO(3) invariance of the proposed bispectral construction. The results demonstrate that for continuous and unbounded transformation groups such as SO(3), theoretical invariance is not merely advantageous but necessary for practical and reliable watermarking. Additional visualizations of rotated panoramas are provided in Appendix Figure 8.

Metrics. We evaluate the performance using Peak Signalto-Noise Ratio (PSNR), Structural Similarity (SSIM), and Bit Accuracy. PSNR and SSIM quantify the perceptual quality of watermarked images, and Bit Accuracy measures watermark extraction reliability.

5.3. General Comparison Comparative Robustness against Common Distortions. We evaluate the robustness across all methods under a range of common image distortions, with results reported in Table 1. Beyond robustness to arbitrary rotations, TRIAD demonstrates strong resilience to diverse perturbations without explicit data augmentation, including JPEG Compression, Gaussian Filter, Gaussian Noise, Median Filter, Resize, Brightness, and Contrast, achieving performance comparable to augmented baselines. In most cases, its performance is comparable to or exceeds that of augmented baselines. Detailed distortion parameters are provided in Appendix C.4.

5.2. Comparative Performance to Arbitrary SO(3) Rotations against Data Augmentation To evaluate robustness under unconstrained geometric transformations, we apply random 3D rotations sampled uniformly from SO(3) using random unit quaternions. After rotation, watermark extraction is performed directly, without alignment or inverse transformation. We compare TRIAD against all baseline methods under this setting. To validate the necessity of theoretical robustness over data augmentation, we additionally report baseline performance with extensive rotation-based augmentation. The augmentation training details are provided in Appendix C.5. As shown in Figure 3, without augmentation, all baseline

Notably, several of these robustness properties can be directly attributed to the representation-theoretic structure of the proposed framework. Operations such as resizing and 6

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Figure 3. Bit accuracy under random SO(3) rotations with increasing rotation angles. We parameterize rotations by their rotation angle, which measures the geodesic distance to the identity rotation on SO(3). For each angle, 1,000 rotation axes are sampled uniformly on S2 , and the average bit accuracy is reported.

Table 1. Comparison with baseline methods. The best and the second best results are highlighted in bold and underlined, respectively. Mixed denotes the averaged performance over combinations of three randomly selected distortions. M ETHOD S TEGA S TAMP (TANCIK ET AL ., 2020) S EP M ARK (W U ET AL ., 2023) T RUST M ARK (B UI ET AL ., 2023) E DIT G UARD (Z HANG ET AL ., 2024 B ) ROBUST-W IDE (H U ET AL ., 2024) VINE (L U ET AL ., 2025 A ) TRIAD (O URS )

C APACITY 100 30 100 64 64 100 32

PSNR↑ 27.96 35.68 40.83 36.58 41.65 36.33 39.22

G ENERAL D ISTORTIONS↑

SSIM↑ 0.8986 0.9799 0.9968 0.8865 0.9921 0.9865 0.9946

JPEG

R ESIZE

C ONTRAST

B RIGHTNESS

G AUSSIAN N OISE

G AUSSIAN B LUR

M EDIAN F ILTER

M IXED

0.973 0.985 0.993 0.957 0.997 1.000 0.978

0.812 0.864 1.000 0.634 0.998 1.000 1.000

0.987 0.988 0.982 0.966 0.973 0.994 1.000

0.986 0.984 0.955 0.941 0.990 0.976 0.988

0.961 0.978 0.986 0.935 0.989 1.000 0.975

0.879 0.987 0.973 0.513 0.999 0.951 1.000

0.894 0.969 0.984 0.546 1.000 0.965 1.000

0.978 0.977 0.979 0.679 0.992 0.986 0.984

Gaussian Filter correspond to isotropic low-pass perturbations in the spherical harmonic domain, inducing smooth, frequency-dependent attenuation of coefficients. Since the proposed watermark is recovered via third-order SO(3)invariant bispectral contractions, which multiplicatively couple multiple coefficients across frequency bands, such spectral attenuation results in a continuous scaling of the invariant response rather than structural destruction, thereby allowing stable extraction. Similarly, regarding additive Gaussian noise, while introducing a bias in higher-order statistics, this isotropic noise primarily induces a magnitude scaling proportional to the signal strength, which preserves the relative geometric configuration of the coupled coefficients. As a result, the bispectral invariant preserves sufficient structure for stable extraction under moderate noise levels. JPEG compression, while non-linear in the spatial domain, predominantly suppresses localized high-frequency content and does not introduce coherent global perturbations aligned with the invariant subspace. This allows reliable recovery of the globally aggregated bispectral features.

tional theoretical analysis is in Appendix A.2. Fidelity. TRIAD achieves high visual fidelity, with PSNR above 39.2 dB and SSIM exceeding 0.99, only marginally lower than TrustMark and Robust-Wide. Visualizations of watermarking patterns are in Appendix Figure 9. 5.4. Sensitivity Analysis Impact of Embedding Subspace Composition Vembed . We study the effects of spectral composition of the embedding subspace Vembed . A fundamental trade-off exists in spherical watermarking: embedding in lower-degree irreducible representations (e.g., l = 4) offers inherent geometric stability but induce perceptible low-frequency artifacts, whereas embedding in higher-degree representations (e.g., l = 16) ensures improved imperceptibility at the cost of increased vulnerability to high-frequency attenuation during signal processing. We conduct an ablation study on subspace configurations, ranging from single-degree targets to multi-scale combinations. As illustrated in Figure 4, restricting the embedding to low degree (V4 ) yields high robustness (> 99%) but compromises visual fidelity. Conversely, exclusively utilizing high degrees (V16 ) preserves quality but degrades extraction accuracy. Our proposed configuration, Vembed = V6 ⊕ V8 ⊕ V14 , exploits spectral diversity via direct sums. By distributing the watermark

In contrast to robustness obtained via data-driven augmentation, these properties arise intrinsically from the algebraic structure of the SO(3)-invariant representation. The observed robustness is therefore not incidental, but a direct consequence of embedding watermark information into globally invariant, higher-order spectral statistics. Addi7

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

signal across disjoint irreducible representations, the model effectively balances the objective: it anchor the invariant identity on the robust low-frequency modes while spreading the information payload into the perceptually masked mid-frequency regions. This spectral allocation enables our method to achieve near-perfect recovery while maintaining high perceptual quality (PSNR ≈ 39.2 dB), significantly outperforming single-degree strategies.

Table 2. Ablation over a wider range of cutoff degrees and embedding subspaces. Bit Accuracy is evaluated under 3D rotations.

Analysis of Cutoff Degree lmax . We investigate the tradeoff between computational cost, perceptual fidelity, and robustness under varying spherical harmonic cutoff degrees lmax . Although increasing lmax theoretically raises the complexity of the Spherical Harmonic Transform, our architecture demonstrates that its impact on inference is limited. Given the same number of selected embedding subspaces Vembed , the inference latency remains remarkably stable as lmax increases from 4 to 24. This indicates that the computational bottleneck is dominated by the channel-wise equivariant tensor product operations rather than the spatial-spectral transforms.

Table 3. Ablation study on invariant projection mechanism. The Power Spectrum fails to support high payload capacities due to phase blindness.

lmax

Lembed

PSNR ↑

Bit Accuracy ↑

16 20 24 28

{6, 8, 14} {6, 8, 14, 16} {6, 8, 14, 16, 20} {6, 8, 14, 16, 20, 22}

39.22 39.16 38.46 37.19

1.000 1.000 1.000 1.000

Projection Mechanism Power Spectrum Power Spectrum Bispectrum Bispectrum Bispectrum

Order

Capacity (bits)

Bit Acc (%)

2 2 3 3 3

16 32 16 32 64

92.4 61.3 100.0 100.0 100.0

We evaluate the efficacy of the bispectrum versus the power spectrum as projection mechanisms for mapping higherorder equivariant features to zeroth-order invariant scalars. While both serve as channels to distill rotation-invariant signatures from the embedding space, the bispectrum uniquely retains directional phase information via third-order coupling (Vl1 ⊗Vl2 ⊗Vl3 → V0 ), which is structurally discarded by the second-order Power Spectrum (Vl1 ⊗ Vl2 → V0 ). We identify this limitation as “Phase Blindness”: the Power Spectrum is a non-injective (many-to-one) mapping, meaning multiple distinct high-order signals can collapse to the same invariant scalar, creating ambiguity that limits watermark capacity. To validate this, we train a variant (PowerSpec) in which the bispectral extraction is replaced by power spectrum projection. As shown in Table 6, the Power-Spec model fails to converge when the payload exceeds 16 bits. In contrast, the bispectrum preserves phase coupling, allowing distinct signal configurations to be uniquely resolved in the invariant domain. This capability enables TRIAD to scale to higher capacity with near-perfect recovery, validating that third-order statistics are the optimal sufficient statistics for high-capacity invariant watermarking.

Regarding robustness and fidelity, higher cutoff degrees provide access to broader spectral subspaces and thus potentially larger embedding capacity. However, they also introduce a trade-off: higher-degree coefficients correspond to finer spatial details, which are more susceptible to attenuation under lossy compression and aliasing during ERP-tosphere projection. To examine whether this trend remains stable beyond the default setting, we further extend the ablation to a wider range of cutoff degrees and embedding subspaces. Specifically, we progressively increase lmax from 16 to 28 and enlarge Vembed by incorporating additional higherdegree irreducible subspaces. As summarized in Table 2, the bit accuracy under 3D rotations remains consistently at 100% across all tested configurations, indicating that no obvious numerical instability or abrupt robustness degradation is observed in this range. Meanwhile, PSNR decreases smoothly from 39.22 dB to 37.19 dB as the embedding subspace becomes broader, suggesting that the main effect of using wider spectral subspaces is a gradual reduction in visual fidelity rather than a failure of the rotation-invariant extraction mechanism. These results further support our choice of lmax = 16 as the default configuration. Although larger cutoffs remain robust to 3D rotations in the tested range, they bring limited robustness benefit while gradually sacrificing perceptual quality. Therefore, we adopt lmax = 16 together with Vembed = {6, 8, 14} as the fidelity–robustness sweet spot, which maximizes the usable embedding capacity within geometrically stable frequency subspaces while avoiding unnecessary fidelity degradation.

6. Discussion Despite the theoretical guarantees of rotation invariance, our framework faces an inherent trade-off between embedding capacity and spectral robustness. As indicated in our analysis (Section 5.4), spherical harmonic components of higher degrees (l > 16) are susceptible to attenuation from common distortions such as lossy compression and aliasing artifacts. To prioritize the reliable recovery of the watermark under strictly invariant geometric priors, we constrain the embedding to middle-frequency spectral bands. Conse-

Importance of Bispectrum over Power Spectrum. Detailed theoretical analysis is provided in the Appendix B.4. 8

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Figure 4. (a) Inference time and Bit Accuracy with different spherical harmonic cutoff degree (lmax ). (b) Performance with different settings of embedding subspace Vembed . Robustness is evaluated under contrast (0.7x) attack.

Impact Statement

quently, while strategies discussed in Appendix C.6 demonstrate potential for capacity enhancement, the effective payload is currently capped (e.g., 64 bits) to ensure stability. Scaling to high-capacity payloads without compromising the robustness of the invariant descriptors remains an open challenge for the spherical watermarking domain. Looking ahead, we will focus on extending the current global invariant framework to handle partial spherical signals by investigating local equivariant symmetries and enhancing the watermark capacity.

This work advances digital watermarking for spherical media by introducing a provably rotation-invariant framework grounded in group representation theory. By moving beyond the prevailing reliance on data augmentation, which offers only empirical and bounded robustness, we establish a framework for provably robust watermarking grounded in group representation theory. This theoretical guarantee is critical for the reliable provenance tracking of immersive content in increasingly complex pipelines, such as World Models and the Metaverse, where geometric transformations are intrinsic rather than adversarial. Furthermore, our introduction of the third-order spherical bispectrum as a carrier for information transmission provides a new direction for information hiding. By demonstrating that higher-order spectral invariants can serve as a strictly rotation-invariant domain for embedding and extraction, we provide a mathematically grounded alternative to spatial or frequency-based heuristics. This approach not only solves the immediate challenge of SO(3) robustness but also establishes a generalized methodology for signal processing on non-Euclidean manifolds. We anticipate this will inspire future research into equivariant information carriers for other geometric data types, fostering the development of trustworthy copyright protection mechanisms for the emerging applications in immersive media, 3D vision and embodied AI systems.

7. Conclusion In this work, we present TRIAD, a principled framework for rotation-invariant watermarking of panoramic imagery grounded in the representation theory of SO(3). By modeling panoramas as spherical signals and leveraging thirdorder representation coupling, we construct bispectral invariants that enable reliable watermark extraction under arbitrary 3D rotations. Unlike augmentation-based approaches, our method provides theoretical guarantees of invariance while preserving high perceptual quality by embedding information exclusively in higher-order spherical harmonic components. Extensive experiments demonstrate that TRIAD achieves near-perfect robustness to continuous SO(3) rotations and strong resilience to common signal distortions. We believe this work highlights the necessity of invariant representations for watermarking on non-Euclidean domains and opens new directions for information hiding based on higher-order group-theoretic invariants.

References Al-Haj, A. Combined dwt-dct digital image watermarking. Journal of computer science, 3(9):740–746, 2007.

Acknowledgements

Batzner, S., Musaelian, A., Sun, L., Geiger, M., Mailoa, J. P., Kornbluth, M., Molinari, N., Smidt, T. E., and Kozinsky, B. E (3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials. Nature communications, 13(1):2453, 2022.

This work was supported by the Strategic Priority Research Program of the Chinese Academy of Sciences (NO. XDB0690302), and the National Nature Science Foundation of China under Grant 62371450.

Ben Jabra, S. and Ben Farah, M. Deep learning-based water9

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

marking techniques challenges: a review of current and future trends. Circuits, Systems, and Signal Processing, 43(7):4339–4368, 2024.

Kondor, R. The principles behind equivariant neural networks for physics and chemistry. Proceedings of the National Academy of Sciences, 122(41):e2415656122, 2025.

Bui, T., Agarwal, S., and Collomosse, J. Trustmark: Universal watermarking for arbitrary resolution images. arXiv preprint arXiv:2311.18297, 2023.

Kondor, R., Lin, Z., and Trivedi, S. Clebsch–gordan nets: a fully fourier space spherical convolutional neural network. Advances in Neural Information Processing Systems, 31, 2018.

Chen, P., Liu, Y., Gu, X., Liu, E., Shang, Z., Ji, X., and Liu, W. Plugmark: A plug-in zero-watermarking framework for diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 17335– 17345, 2025a.

Li, J., Liu, Y., Chen, P., Shang, Z., and Gu, X. Gs-mark: Deep robust watermarking for graph signals. In ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 13922– 13926. IEEE, 2026.

Chen, P., Liu, Y., Gu, X., Chen, X., Liu, W., and Wang, W. Rel-zero: Harnessing patch-pair invariance for robust zero-watermarking against ai editing. arXiv preprint arXiv:2603.17531, 2026.

Li, X., He, X., Zhang, L., Wu, M., Li, X., and Liu, Y. A comprehensive survey on world models for embodied ai. arXiv preprint arXiv:2510.16732, 2025.

Chen, Z., Wang, G., Zhu, J., Lai, J., and Xie, X. Guardsplat: Efficient and robust watermarking for 3d gaussian splatting. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 16325–16335, 2025b.

Liu, Y., Liu, J., Argyriou, A., Ma, S., Wang, L., and Xu, Z. 360-degree vr video watermarking based on spherical wavelet transform. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 17(1):1–23, 2021.

Cohen, T. S., Geiger, M., Koehler, J., and Welling, M. Spherical cnns. In ICLR, 2018.

Lu, S., Zhou, Z., Lu, J., Zhu, Y., and Kong, A. Robust watermarking using generative priors against image editing: From benchmarking to advances. In International Conference on Learning Representations, volume 2025, pp. 83902–83936, 2025a.

Esteves, C., Allen-Blanchette, C., Makadia, A., and Daniilidis, K. Learning so (3) equivariant representations with spherical cnns. In Proceedings of the european conference on computer vision (ECCV), pp. 52–68, 2018. Geiger, M. and Smidt, T. e3nn: Euclidean neural networks. arXiv preprint arXiv:2207.09453, 2022.

Lu, T., Shu, T., Yuille, A., Khashabi, D., and Chen, J. Genex: Generating an explorable world. In International Conference on Learning Representations, volume 2025, pp. 52310–52335, 2025b.

Hu, R., Zhang, J., Xu, T., Li, J., and Zhang, T. Robust-wide: Robust watermarking against instruction-driven image editing. In European Conference on Computer Vision, pp. 20–37. Springer, 2024.

Luo, X., Zhan, R., Chang, H., Liu, F., and Milanfar, P. Distortion agnostic deep watermarking. CVPR, 2020.

Iglesias Martı́nez, M. E., Antonino-Daviu, J. A., Dunai, L., Conejero, J. A., and Fernández de Córdoba, P. Higherorder spectral analysis and artificial intelligence for diagnosing faults in electrical machines: An overview. Mathematics, 12(24):4032, 2024.

Makadia, A. and Daniilidis, K. Direct 3d-rotation estimation from spherical images via a generalized shift theorem. In 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003. Proceedings., volume 2, pp. II–217. IEEE, 2003.

Kazhdan, M., Funkhouser, T., and Rusinkiewicz, S. Rotation invariant spherical harmonic representation of 3d shape descriptors. In Symposium on geometry processing, volume 6, pp. 156–164, 2003.

Mataigne, S., Mathe, J., Sanborn, S., Hillar, C., and Miolane, N. The selective g-bispectrum and its inversion: Applications to g-invariant networks. Advances in Neural Information Processing Systems, 37:115682–115711, 2024.

Kohler, C., Patel, P., Vaska, N., Goodwin, J., Jones, M. C., Platt, R., Caceres, R. S., and Walters, R. Bridging equivariant gnns and spherical cnns for structured physical domains. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.

Narendra, M., Valarmathi, M., Anbarasi, L. J., and Gandomi, A. H. Levenberg–marquardt deep neural watermarking for 3d mesh using nearest centroid salient point learning. Scientific Reports, 14(1):6942, 2024. 10

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Nikias, C. L. and Mendel, J. M. Signal processing with higher-order spectra. IEEE Signal processing magazine, 10(3):10–37, 1993.

Wu, X., Liao, X., and Ou, B. Sepmark: Deep separable watermarking for unified source tracing and deepfake detection. In Proceedings of the 31st ACM International Conference on Multimedia, pp. 1190–1201, 2023.

Ocampo, J., Price, M. A., and McEwen, J. D. Scalable and equivariant spherical cnns by discrete-continuous (disco) convolutions. In International Conference on Learning Representations, 2023.

Xiao, J., Ehinger, K. A., Oliva, A., and Torralba, A. Recognizing scene viewpoint using panoramic place representation. In 2012 IEEE conference on computer vision and pattern recognition, pp. 2695–2702. IEEE, 2012.

Poulenard, A., Rakotosaona, M.-J., Ponty, Y., and Ovsjanikov, M. Effective rotation-invariant point cnn with spherical harmonics kernels. In 2019 International Conference on 3D Vision (3DV), pp. 47–56. IEEE, 2019.

Zaman, K. A. U., Alam, M. Z., Ali, M. N., and Miraz, M. H. Deep neural watermarking for robust copyright protection in 3d point clouds. arXiv preprint arXiv:2510.27533, 2025.

Sanborn, S. and Miolane, N. A general framework for robust g-invariance in g-equivariant networks. Advances in Neural Information Processing Systems, 36:67103– 67124, 2023.

Zhang, C., Wu, Q., Gambardella, C. C., Huang, X., Phung, D., Ouyang, W., and Cai, J. Taming stable diffusion for text to 360 panorama image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6347–6357, 2024a.

Shinde, Y., Lee, K., Kiper, B., Simpson, M., and Hasanzadeh, S. A systematic literature review on 360° panoramic applications in architecture, engineering, and construction (aec) industry. Journal of Information Technology in construction, 28, 2023.

Zhang, X., Li, R., Yu, J., Xu, Y., Li, W., and Zhang, J. Editguard: Versatile image watermarking for tamper localization and copyright protection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11964–11974, 2024b.

Sloan, P.-P. Stupid spherical harmonics (sh) tricks. In Game developers conference, volume 9, pp. 320–321, 2008.

Zhang, Y., Song, S., Tan, P., and Xiao, J. Panocontext: A whole-room 3d context model for panoramic scene understanding. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 612, 2014, Proceedings, Part VI 13, pp. 668–686. Springer, 2014.

Tancik, M., Mildenhall, B., and Ng, R. Stegastamp: Invisible hyperlinks in physical photographs. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2117–2126, 2020. Thomas, N., Smidt, T., Kearnes, S., Yang, L., Li, L., Kohlhoff, K., and Riley, P. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds. 2018.

Zheng, X., Liao, C., Weng, Z., Lei, K., Dongfang, Z., He, H., Lyu, Y., Jiang, L., Qi, L., Chen, L., et al. Panorama: The rise of omnidirectional vision in the embodied ai era. arXiv preprint arXiv:2509.12989, 2025.

Tukur, M., Schneider, J., Househ, M., Dokoro, A. H., Ismail, U. I., Dawaki, M., and Agus, M. The metaverse digital environments: A scoping review of the techniques, technologies, and applications. Journal of King Saud University-Computer and Information Sciences, 36(2): 101967, 2024.

Zhou, H., Chen, X., Li, J., Zhang, Z., Fu, Y., Liva, M. P., Greenbaum, D., and Hui, P. Generative artificial intelligence in the metaverse era: A review on models and applications. Research, 8:0804, 2025a.

Wang, H., Xiang, X., Xia, W., and Xue, J.-H. A survey on text-driven 360-degree panorama generation. IEEE Transactions on Circuits and Systems for Video Technology, 2025.

Zhou, S., Fan, Z., Xu, D., Chang, H., Chari, P., Bharadwaj, T., You, S., Wang, Z., and Kadambi, A. Dreamscene360: Unconstrained text-to-3d scene generation with panoramic gaussian splatting. In Computer Vision – ECCV 2024, pp. 324–342, Cham, 2025b. Springer Nature Switzerland.

Weiler, M., Hamprecht, F. A., and Storath, M. Learning steerable filters for rotation equivariant cnns. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 849–858, 2018.

Zhu, J., Kaplan, R., Johnson, J., and Fei-Fei, L. Hidden: Hiding data with deep networks. In ECCV, 2018.

Wu, S., Teng, F., Shi, H., Jiang, Q., Luo, K., Wang, K., and Yang, K. Quadreamer: Controllable panoramic video generation for quadruped robots. In Conference on Robot Learning, pp. 1777–1789. PMLR, 2025. 11

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

A. Theoretical Analysis A.1. Proof of Theorem 4.1 Proof. Fix arbitrary degrees (l1 , l2 , l3 ). Let R ∈ SO(3) be an arbitrary rotation. Under R, the spherical harmonics coefficients transform as l X ′ (l) cm → 7 Dmm′ (R) cm l l , m′ =−l

where D(l) (R) denotes the Wigner-D matrix corresponding to the irreducible representation Vl . Consider the corresponding bispectrum component X Il1 ,l2 ,l3 =

2 m3 Cl0,0 cm1 cm l2 cl3 . 1 m1 l2 m2 l3 m3 l1

m1 ,m2 ,m3

After applying the rotation R, the transformed quantity becomes    X X X ′ ′ m m (l ) (l )  Il1 ,l2 ,l3 (R) = Cl0,0 Dm11 m′ (R)cl1 1   Dm22 m′ (R)cl2 2  1 m1 l2 m2 l3 m3 1

m′1

m1 ,m2 ,m3

2

m′2

  X (l ) ′ m × D 3 ′ (R)c 3  . m′3

l3

m3 m3

Reordering the summations yields ! Il1 ,l2 ,l3 (R) =

X

X

m′1 ,m′2 ,m′3

m1 ,m2 ,m3

(l ) (l ) (l ) Cl0,0 Dm11 m′ (R)Dm22 m′ (R)Dm33 m′ (R) 1 m1 l2 m2 l3 m3 1 2 3

m′ m′ m′

cl1 1 cl2 2 cl3 3 .

By the invariance property of Clebsch–Gordan coefficients, the contraction of three Wigner-D matrices with C 0,0 satisfies X (l ) (l ) (l ) Cl0,0 Dm11 m′ (R)Dm22 m′ (R)Dm33 m′ (R) = Cl0,0 ′ ′ ′ . 1 m1 l2 m2 l3 m3 1 m l2 m l3 m 1

2

3

1

2

3

m1 ,m2 ,m3

Substituting this identity back, we obtain Il1 ,l2 ,l3 (R) = Il1 ,l2 ,l3 . Since the full bispectrum invariant I is obtained by summing Il1 ,l2 ,l3 over all (l1 , l2 , l3 ) and the above argument holds independently for each triple, the complete bispectrum invariant I is invariant under arbitrary SO(3) rotations. This establishes the rotation invariance of I. We next analyze its sensitivity to perturbations in higher-order spherical harmonics coefficients. Consider a small perturbation ∆cm l embedded in higher-order SH coefficients. Due to the linearity of the tensor product and the projection onto the trivial representation, the resulting change in I can be expressed as   X X m1 m2 m3 m1 m2 m3 m1 m2 m3 c ∆c + O(∆c2 ). ∆I = Cl0,0 ∆c c c + c ∆c c + c l1 l2 l3 l1 l2 l3 l1 l2 l3 1 m1 l2 m2 l3 m3 l1 ,l2 ,l3 m1 ,m2 ,m3

Hence, ∆I is non-zero for generic perturbations. This confirms that the embedded information is not lost during the projection to the invariant scalar I, but rather manifests as detectable structural changes, providing the discriminative basis for the learnable decoder to recover the watermark. Any information embedded in the higher-order SH coefficients that contributes to I manifests in the rotation-invariant scalar and can, in principle, be extracted given knowledge of the embedding scheme.

12

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

A.2. Theoretical Interpretation of Robustness against Common Distortions The robustness of the proposed watermarking framework to common image distortions arises naturally from the multilinearity, continuity, and global aggregation properties of third-order SO(3)-invariant bispectral representations. Unlike data augmentation, which empirically approximates robustness, these properties follow directly from the algebraic structure of the embedding space, providing principled stability guarantees beyond rotation invariance. Here, we analyze the theoretical stability of the spherical bispectrum operator. While the proposed TRIAD architecture employs a equivariant backbone prior to bispectrum computation, this theoretical analysis remains highly relevant. As demonstrated in Figure 6, the watermarking signal is basically embedded onto the targeted Spherical Harmonics subspace, which showcases that the backbone Φeq (·) effectively functions as a projection from full coefficients to those of targeted subspace. Therefore, the robustness inherent to the raw SH coefficients below directly extends to our method. A.2.1. P RELIMINARIES AND N OTATION Let f (ω) be a spherical signal defined on S2 with spherical harmonic expansion f (ω) =

lX max

l X

m cm l Yl (ω).

l=0 m=−l

Following the main paper, we define the third-order SO(3)-invariant bispectrum scalar as X X 2 m3 I= Cl0,0 cm1 cm l2 cl3 , 1 m1 l2 m2 l3 m3 l1

(16)

l1 ,l2 ,l3 m1 ,m2 ,m3

where Cl0,0 denotes the Clebsch–Gordan coefficients corresponding to projection onto the trivial representation. 1 m1 l2 m2 l3 m3 By construction, I is invariant under arbitrary SO(3) rotations. A.2.2. S TABILITY UNDER I SOTROPIC G AUSSIAN F ILTERING Isotropic Gaussian filtering on the sphere corresponds to convolution with the spherical heat kernel. In the spherical harmonic domain, this induces a frequency-dependent attenuation: m c̃m l = g(l) cl ,

Substituting c̃m l into Eq. (16) yields: X I˜ =

X

2

g(l) = e−σ l(l+1) .

1 m2 m3 Cl0,0 g(l1 )g(l2 )g(l3 ) cm l1 cl2 cl3 1 m1 l2 m2 l3 m3

(17)

l1 ,l2 ,l3 m1 ,m2 ,m3

=

X

g(l1 )g(l2 )g(l3 )

X

2 m3 Cl0,0 cm1 cm l2 cl3 . 1 m1 l2 m2 l3 m3 l1

(18)

m1 ,m2 ,m3

l1 ,l2 ,l3

Equivalently, if Il1 ,l2 ,l3 denotes the bispectral contribution of the triplet (l1 , l2 , l3 ), then X I˜ = g(l1 )g(l2 )g(l3 )Il1 ,l2 ,l3 .

(19)

l1 ,l2 ,l3

This shows that isotropic blur acts on the bispectrum through a structured and degree-dependent attenuation of frequencytriplet responses. Importantly, the coupling pattern induced by the Clebsch–Gordan projection is preserved: blur rescales each valid triplet contribution without mixing unrelated irreducible subspaces. For a fixed bandwidth lmax , the induced variation is bounded by X |I˜ − I| ≤ |g(l1 )g(l2 )g(l3 ) − 1| |Il1 ,l2 ,l3 |. (20) l1 ,l2 ,l3 ≤lmax

Therefore, under moderate blur, the bispectral descriptor undergoes controlled attenuation rather than unstructured distortion. This structured response explains why the learned extractor can retain reliable recovery when sufficient low- and midfrequency bispectral components remain. 13

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

A.2.3. E FFECT OF R ESIZING AND L OW-PASS P ERTURBATIONS Resizing with anti-aliasing can be approximated as a low-pass operation in the spherical harmonic domain: ( cm m l , l ≤ lc , c̃l = 0, l > lc . The corresponding bispectrum is then restricted to frequency triplets whose degrees remain below the cutoff: X X 2 m3 I˜ = Cl0,0 cm1 cm l2 cl3 . 1 m1 l2 m2 l3 m3 l1

(21)

(22)

l1 ,l2 ,l3 ≤lc m1 ,m2 ,m3

Thus, low-pass perturbations remove bispectral interactions involving truncated high-frequency components, while preserving the interactions among retained low- and mid-frequency irreducible subspaces. For natural images, spectral energy is typically concentrated in low- and mid-frequency bands, whereas very high-frequency components carry comparatively weaker energy and are more sensitive to sampling artifacts. Consequently, moderate resizing primarily suppresses high-frequency triplets while leaving a substantial portion of signal-dependent bispectral structure intact. To make this statement precise, let Tc = {(l1 , l2 , l3 ) : l1 , l2 , l3 ≤ lc }

(23)

denote the retained set of triplets. The amount of preserved bispectral information can be characterized by the retained bispectral energy ratio P 2 (l ,l ,l )∈T |Il1 ,l2 ,l3 | ρ(lc ) = P 1 2 3 c . (24) 2 l1 ,l2 ,l3 ≤lmax |Il1 ,l2 ,l3 | When ρ(lc ) remains sufficiently large, the invariant descriptor preserves discriminative, input-dependent structure after low-pass filtering. In this sense, the representation remains informative rather than degenerating into an input-independent or nearly constant descriptor. This formulation clarifies the role of low-/mid-frequency bispectral couplings in maintaining robust extraction under resizing. A.2.4. ROBUSTNESS TO A DDITIVE G AUSSIAN N OISE . We model additive noise in the spherical harmonic domain as: m m c̃m l = cl + ϵl ,

2 ϵm l ∼ N (0, σ ).

To evaluate the stability of the watermarking signal, we examine the expectation of the third-order bispectrum invariant under this perturbation. Substituting the noisy coefficients into the tensor contraction (Eq. (16)) and taking the expectation, we observe that while the first-order noise terms vanish (due to E[ϵ] = 0), the second-order interactions introduce a non-zero term: X 0,0 ˜ =I+ E[I] Cl1 l2 l3 (cl1 · E[ϵl2 ϵl3 ] + . . . ) + E[ϵ3 ] Since the noise is isotropic, the quadratic term E[ϵ2 ] is proportional to the noise variance σ 2 . Consequently, the expected value of the noisy invariant takes the form: ˜ ≈ I + β · c · σ 2 ≈ I(1 + λσ 2 ) E[I] where λ is a scalar factor derived from the coupling constants. This result indicates that while the estimator is mathematically biased (contradicting a zero-bias assumption), the bias is structure-preserving. Specifically, the noise acts primarily as a magnitude scaling factor proportional to the signal strength c, rather than an additive shift that disrupts the feature’s sign or relative orientation. This geometric property explains the experimental robustness without explicit noise augmentation. Since the perturbation scales the feature vector but preserves its directionality and relative ordering, it does not push the embedding across the decision boundary of the MLP decoder. Thus, the watermark remains recoverable even when the feature magnitude is modulated by noise interference. (Note: While the real-valued constraint of images imposes conjugate symmetry on ϵ, implying correlations between ϵm and ϵ−m , this strictly affects the magnitude of the scalar λ but does not alter the fundamental conclusion that the distortion manifests as a signal-dependent scaling.) 14

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling Table 4. Isolated analysis of bispectral descriptors under common distortions. Cosine similarity is computed between the bispectrum before and after applying each distortion.

Distortion

Strength

Cosine Similarity ↑

Gaussian Blur Resize Gaussian Noise Combined

σ=3 0.5× std. = 0.05 –

0.997849 0.999972 0.999965 0.997458

A.2.5. JPEG C OMPRESSION AND L OCALIZED D ISTORTIONS JPEG compression is non-linear in the spatial domain and difficult to model exactly in spherical harmonics. However, its primary effect consists of localized block-wise quantization and high-frequency suppression. Since bispectral invariants aggregate global spectral interactions through Clebsch–Gordan contractions, such localized distortions do not coherently align with the invariant subspace. Consequently, the global bispectral response remains stable, consistent with our empirical observations. A.2.6. I SOLATED Q UANTITATIVE A NALYSIS OF B ISPECTRAL D ESCRIPTORS To further validate the stability properties analyzed above, we evaluate the bispectral descriptor in isolation under representative distortions. This experiment directly measures the change of the invariant representation itself, before the learning-based watermark decoder. For each clean panorama x, we compute its bispectral descriptor B(x) from the spherical harmonic coefficients. Given a distorted image x̂, we compute B(x̂) using the same cutoff degree and frequency-triplet configuration. We then measure the cosine similarity between the two descriptors: Scos =

⟨B(x̂), B(x)⟩ . ∥B(x̂)∥2 ∥B(x)∥2

(25)

A higher cosine similarity indicates that the bispectral descriptor preserves its direction in the invariant feature space after distortion. As shown in Table 4, the bispectral descriptor remains highly consistent under isolated distortions, with cosine similarity above 0.997 even under the combined setting. These results quantitatively support the stability analysis above: common distortions perturb the bispectrum in a structured and smooth manner, while largely preserving the direction of the invariant descriptor.

B. More Analysis B.1. Overhead Evaluation We evaluate the encoding, decoding and total time cost and GPU memory usage of watermarking methods on an NVIDIA A100 GPU. The results are averaged over 1,000 images. As shown in Table 5, our method demonstrates a comparatively low cost both in inference time and in GPU usage. B.2. Spectral Precision and Orthogonal Embedding To verify the disentanglement capabilities of our encoder, we examine the distribution of watermark energy across the spherical harmonic degrees. We decompose the image signal into its spectral components and compute the power spectrum P (l) for each degree l: l X 2 P (l) = ∥cm l ∥ m=−l

where cm l denotes the spherical harmonic coefficients. Ideally, the embedding scheme should modulate only the specific degrees allocated for the payload, ensuring that the information is strictly confined to the target subspaces Vembed without leaking into or corrupting adjacent coefficients. 15

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling Table 5. Comparison of watermarking methods based on running time per single image and GPU memory usage. The results are averaged over 1,000 images.

Method

Encoding Time (s)

Decoding Time (s)

Total Time (s)

Memory (GB)

0.0354 0.0053 0.0197 0.1659 0.0111 0.0583 0.0081

0.0318 0.0055 0.0147 0.0773 0.0157 0.0103 0.0072

0.0672 0.0108 0.0344 0.2432 0.0268 0.0686 0.0153

2.0 0.9 0.6 1.7 3.0 4.9 0.8

StegaStamp (Tancik et al., 2020) SepMark (Wu et al., 2023) TrustMark (Bui et al., 2023) EditGuard (Zhang et al., 2024b) Robust-Wide (Hu et al., 2024) VINE (Lu et al., 2025a) TRIAD (Ours)

Figure 6 presents the spectral energy difference between the cover and watermarked images. These results demonstrate the high precision of our modulation mechanism, where energy variations are only observed exclusively at the pre-defined embedding degrees from Vembed . This confirms that the embedding strategy correctly map the latent watermark code onto the intended geometric features. While for the non-targeted degrees, the coefficient norms remain virtually identical to those of the original image (relative change ≈ 0). This indicates that our method achieves near-perfect orthogonality of spectral watermark injection. The absence of unintended perturbations in non-watermarked degrees validates the functionality of the extraction process and preserves the image quality by leaving the vast majority of the frequency components unchanged, which aligns perfectly with our design objective of localized and disentangled feature modulation. B.3. Verification of Rotational Invariance To empirically validate the theoretical invariance of our bispectrum-based architecture, we conduct a rigorous stability analysis on the latent feature space. Specifically, given an input x and a randomly rotated counterpart xrot = g · x (where g ∈ SO(3)), we extract their respective tensor product feature vectors z and zrot and analyze their correspondence. As illustrated in Fig 7, the Scatter Correlation analysis reveals a near-perfect linear relationship between z and zrot . All feature points densely cluster around the diagonal identity line (y = x), demonstrating that the feature magnitude remains consistent regardless of the input’s geometric orientation. The Pearson correlation coefficient is close to 1.0, indicating strong linear dependence. Furthermore, the overlay plot offers a microscopic view of the feature topology. We observe that the red dashed line (representing zrot ) explicitly tracks the blue solid line (representing z) across feature dimensions. The overlapping waveforms indicate that the bispectrum contraction mechanism successfully eliminates the perturbations induced by rotation. These observations align with our theoretical derivation. By computing the third-order invariants (bispectrum) via the contraction TP(TP(u, u), u) → V′ , the model effectively cancels out the Wigner-D matrices induced by rotation. This confirms that our decoder has learned a robust SO(3)-invariant representation, ensuring reliable watermark extraction even under extreme geometric distortions. B.4. From Power Spectrum to Bispectrum Achieving robustness to arbitrary 3D rotations requires extracting watermark information from representations that are strictly invariant under SO(3). Intuitively, we observe that it can be achieved by constructing invariants from spherical harmonic (SH) coefficients via tensor contractions that project onto the trivial representation. The simplest such invariant is the power spectrum, Pl =

l X

2 2 |cm l | = ∥cl ∥ ,

(26)

m=−l

which can be interpreted as a second-order contraction of f ⊗ f onto the scalar (0e) irreducible representation. Since Wigner-D matrices are unitary, Pl is exactly invariant under arbitrary rotations. However, this construction discards all relative phase information between SH coefficients. As a result, the mapping from signals to power spectra is highly non-injective: structurally distinct signals may share identical power spectra. For watermarking, this phase blindness fundamentally limits both embedding capacity and reconstructability. To overcome this limitation, we turn to higher-order invariant constructions. In particular, third-order tensor products 16

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

allow the preservation of phase coupling while maintaining exact SO(3) invariance. This leads naturally to the bispectrum invariant, which is obtained by contracting three SH coefficients via Clebsch–Gordan (CG) coefficients corresponding to projection onto the trivial representation: X X 2 m3 (27) I= Cl0,0 cm1 cm l2 cl3 . 1 m1 l2 m2 l3 m3 l1 l1 ,l2 ,l3 m1 ,m2 ,m3

By construction, I is a scalar invariant under SO(3). Unlike the power spectrum, the bispectrum retains closed-loop phase coupling across frequencies, resolving much of the non-injectivity inherent to second-order invariants. Classical results in harmonic analysis show that, under mild conditions, bispectral invariants uniquely characterize a signal up to a global rotation. Consequently, embedding watermark information into the bispectrum invariant space achieves exact rotational robustness while substantially increasing representational capacity. In our framework, this invariant arises naturally from the quadratic tensor product of equivariant features, followed by projection onto the 0e subspace. This representation-theoretic construction ensures that robustness to arbitrary rotations is achieved by design, rather than through data augmentation, while preserving the structural information necessary for reliable watermark decoding.

C. More Details C.1. Perceptual Module Detail The perceptual module modulates the watermark residual using two complementary components to ensure imperceptibility. First, a Learnable Content Mask is generated by a lightweight CNN (consisting of three convolutional layers with SiLU activations and a final Sigmoid) to adaptively hide information in high-texture regions. Second, we apply a fixed Geometric Prior Mgeo (θ) = sin(θ) to counteract the non-uniform sampling of Equirectangular Projection (ERP). Since the panoramic image is an Equirectangular Projection (ERP) of the sphere, pixels do not represent equal areas. The sampling density increases significantly near the poles (θ → 0, π), meaning modifications in these regions are spatially magnified when viewed on the sphere. To counteract this distortion, we introduce a fixed geometric prior Mgeo derived from the spherical area element dA = sin θdθdϕ. For an image of height H, the weight for row i (corresponding to polar angle θi ) is calculated as: i (28) Mgeo (i) = sin(θi ), where θi = π H This sinusoidal weighting suppresses watermark strength near the poles while allowing full strength at the equator, preventing polar artifacts.

Learnable Mask 𝑴𝒑𝒆𝒓𝒄

Original Panorama 𝒙

Watermarked Panorama ෥ 𝒙

Geometric Prior 𝑴𝒈𝒆𝒐

Residual ∆𝒙

Modulated Residual

CNN+SiLU

Figure 5. Perceptual module. The learnable mask Mperc is three blocks of Convolution networks followed by an activation.

C.2. Network Architecture Details We provide the detailed specifications of the core modules. 17

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

1. Equivariant Tensor Product Interaction TPϑ . To enable information mixing within the spherical harmonic domain, we employ a FullyConnectedTensorProduct from e3nn (Geiger & Smidt, 2022) library. We define the parameterized equivariant tensor product, denoted as TPϑ (·, ·)|Vout , which couples two input representations using fixed Clebsch-Gordan coefficients followed by learnable path weights to map the result to a specified output subspace Vout . 2. Watermark w Mapping. We apply 3 three layers of linear and activation block to project w from k to Dw = 512 dimensions. 3. SO(3)-Equivariant Backbone. The backbone consists of stacked Gated Blocks, which are essential for introducing non-linearity to geometric tensors without breaking equivariance. Each block performs the following operations: • Linear Projection: An equivariant linear layer. • Gating Mechanism: To activate higher-order tensors (which do not have a rotation-invariant sign), the features are separated into scalars and gated tensors. – Scalars (0e): Activated directly using the SiLU function. – Higher-order Tensors (l > 0): Modulated element-wise by a learned scalar gate (activated via Sigmoid to [0, 1]). This structure ensures that the magnitude of directional features is non-linearly transformed while their orientation remains equivariant. 4. MLP Decoder. The decoder maps the extracted rotation-invariant bispectral scalars to the binary watermark vector. It follows a standard Multi-Layer Perceptron structure of: A sequence of Linear layers (dim → 256 → 128) with SiLU activations and a final Linear projection maps the features to the target watermark dimension. C.3. More Ablation Analysis We validate the necessity of the proposed Perceptual Module by removing it from the embedding pipeline (w/o Perceptual Module). As shown in Table 6, disabling the module leads to a sharp decline in visual quality, with PSNR dropping from 39.22 dB to 34.15 dB. Without the learnable mask to suppress perturbations in smooth regions (e.g., sky or walls), the watermark becomes visually intrusive, confirming that the perceptual module helps to maintain high fidelity. We further analyze the optimal dimension Dw for the watermark projection MLP. Our method set Dw = 512. Reducing the dimension to 256 restricts the representation capacity, causing the Bit Accuracy to fall to 96.8%, as the bottleneck limits the effective encoding of the watermark signal. Meanwhile, increasing the dimension to 1024 introduces no gain in both bit accuracy and visual fidelity. This confirms that our design choice (Dw = 320), which confirms our setting is efficient yet effective. Table 6. Ablation study on network components and hyper-parameters. We investigate the contribution of the Perceptual Module and the impact of the watermark expansion dimension (Dw ). The default setting (Ours) uses Dw = 512 with the Perceptual Module enabled.

Method / Variant

PSNR (dB) ↑

Bit Acc (%) ↑

TRIAD (Ours)

39.22

100.0

Component Analysis w/o Perceptual Module

35.15

100.0

Watermark Expansion Dimension (Dw ) Dw = 256 39.45 Dw = 1024 38.60

96.8 100.0

C.4. Distortions In our method, we apply a set of commonly used noise perturbations and image transformations to evaluate robustness and performance under degraded visual conditions. The detailed parameter settings are summarized as follows: • JPEG Compression: quality factor = 60. 18

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

• Resize: scaling factor = 0.5. • Gaussian Blur: kernel size = 1, σ = 3. • Gaussian Noise: mean = 0, standard deviation = 0.05. • Median Filter: kernel size = 3. • Brightness Adjustment: range = [0.7, 1.3]. • Contrast Adjustment: range = [0.7, 1.3]. C.5. Additional Baseline Rotation Augmentation Training Details Unlike the continuous rotation invariance guaranteed by our method, baseline models lack intrinsic geometric priors and must rely on seeing specific transformations during training. To simulate this, we randomly select a fixed set of K = 10 discrete rotations from the continuous group SO(3) as anchor augmentations. We limit the set size to K = 10 because SO(3) rotations induce drastic, non-linear geometric deformations in the equirectangular domain. Attempting to train on a dense or continuous sampling of the rotation manifold introduces excessive distributional variance, which we found empirically to prevent the baseline models from converging. Thus, this limited set serves as a tractable representative sample of the global rotation space to ensure fair comparison. We fine-tune all baseline methods on the PanoContext and SUN360 datasets to evaluate their empirical robustness. Specifically, for each training step, a rotation is stochastically selected from this pre-defined set and applied to the image after watermark embedding. All models are initialized from their officially released checkpoints and trained for 100 epochs using the AdamW optimizer with a learning rate of 1 × 10−4 . We evaluate the performance of these augmented baselines in Table 7. As observed, visual fidelity drops significantly due to the inherent trade-off between imperceptibility and robustness in non-equivariant models. Crucially, even when trained and evaluated on this highly restricted subset—which is far simpler than real-world continuous rotation—their performance remains unsatisfactory. This failure to generalize even to a finite set further demonstrates the necessity of theoretical invariance over data augmentation for SO(3) robustness. Table 7. Performance of baseline methods on PSNR and Average Bit Accuracy on panoramas rotated from the pre-defined anchor rotation set. The results are averaged over 1,000 images and angles from the set. Method StegaStamp (Tancik et al., 2020) SepMark (Wu et al., 2023) TrustMark (Bui et al., 2023) EditGuard (Zhang et al., 2024b) Robust-Wide (Hu et al., 2024) VINE (Lu et al., 2025a)

PSNR (dB)

Average Bit Accuracy (%)

25.15 31.42 35.50 32.18 37.10 31.85

73.38% 76.65% 78.15% 77.14% 82.23% 79.40%

C.6. Group-wise Extension for Payload Enhancement In the main paper, watermark embedding is performed in the embedding space Vembed . To increase capacity, we introduce a group-wise extension that partitions Vembed into multiple independent subspaces. Concretely, the embedding space is composed of irreducible representations with multiplicity: Vembed =

M

(ml × Vl ) ,

(29)

l∈Lembed

where ml denotes the multiplicity of degree-l irreducible representations. We evenly partition the multiplicity dimension into G groups, yielding: G M  ml ml × Vl = ∀l ∈ Lembed . (30) G × Vl , g=1

19

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

The watermark vector w is partitioned into {w(g) }G g=1 . For each group g, watermark information is injected only into the corresponding multiplicity slice: (g)

∆u(g) = TPϑ1 (u(g) , w(g) )|( ml ×Vl ) , G

(g)

TPϑ1 :

ml G × Vl



⊗ V0 →

ml G × Vl



.

(31)

The full modified representation is obtained by concatenating all group-wise updates along the multiplicity dimension. During extraction, invariant statistics are computed independently within each group:   z (g) = TP û(g) , û(g) , û(g) ,

(32)

followed by concatenation of the group-wise decoded outputs. Importantly, groups are formed by splitting the multiplicity dimension of each irreducible representation, rather than by mixing different degrees l, which preserves orthogonality and SO(3)-equivariance within each group. C.7. Resolution Scaling Algorithm 1 is adapted from the scaling method proposed in TrustMark (Bui et al., 2023). It allows a watermark model trained at a fixed resolution to be applied to images of arbitrary resolutions without performance degradation. Algorithm 1 Resolution scaling - watermark embedding on arbitrary resolution images Input: Original image x, [binary watermark vector w] Output: Watermarked image y Data: Embedding network E trained on resolution m × n 1: H, W ← height(x), width(x) 2: x ← x/127.5 − 1 3: x̃ ← interpolate(x, (m, n)) 4: r ← E(x̃, w) − x̃ 5: r ← interpolate(r, (H, W )) 6: y ← clamp(x + r, −1, 1) 7: y ← y × 127.5 + 127.5

// Normalize to range [−1, 1] // Residual image

D. More Results D.1. Visualizations of SO(3) Rotations See Fig 8. D.2. Visualizations of Comparative Embedding See Fig 9.

20

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Figure 6. Analysis of Spectral Targeting Precision in Spherical Harmonics. We compare the aggregated energy spectrum (L2 norm of coefficients) across degrees l = 0 to lmax for the original (blue) and watermarked (orange) images. (Top) The spectral profiles exhibit high congruence, indicating that the watermarking process preserves the natural frequency distribution of the cover image. (Bottom) The relative change rate shows distinct, isolated modulations only at the specific degrees targeted for watermark embedding (e.g., l ∈ Ltarget ) and confirms that non-targeted degrees (l ∈ / Ltarget ) remain virtually perturbed (near-zero deviation). This validates the orthogonality of our embedding mechanism, demonstrating that the watermark is accurately confined to the designated subspaces without spectral leakage, thereby preserving the fidelity of unmodulated components.

21

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Figure 7. Analysis of Rotational Invariance in Feature Space. We evaluate the stability of the learned bispectrum features under arbitrary SO(3) rotations. (Left) The Scatter Correlation plot compares feature activations of the original image (x-axis) versus the rotated image (y-axis). The tight clustering along the identity line (y = x) indicates high invariance. (Right) Feature Vector Overlay: A dimension-wise comparison where the blue solid line (Original Features) and the red dashed line (Rotated Features) exhibit near-perfect alignment. This visual overlap confirms that the bispectrum quantity effectively marginalizes geometric pose information and presents as a SO(3) invariant.

22

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Original Image

Yaw:-156°,Pitch: -28°,Roll:36°

Yaw:133°,Pitch: -1°,Roll:-121°

Yaw:45°,Pitch: 81°,Roll:162°

Yaw:-79°,Pitch: 6°,Roll:-131°

Yaw:-57°,Pitch: 28°,Roll:-134°

Yaw:-174°,Pitch: -52°,Roll:-23°

Yaw:126°,Pitch: 37°,Roll:-103°

Yaw:109°,Pitch: 0°,Roll:-75°

Yaw:-22°,Pitch: -58°,Roll:-146°

Figure 8. Visualization of drastic geometric deformations induced by random SO(3) rotations in the Equirectangular Projection (ERP) domain. The top panel shows the original canonical view, while the subsequent panels display the same scene under randomly sampled 3D rotations (annotated with Yaw, Pitch, and Roll). Note that standard rigid 3D rotations manifest as complex, non-linear pixel displacements and distortions in the 2D ERP format. This visualization highlights the inherent challenge for baseline models to learn rotation robustness solely through data augmentation, motivating the need for our theoretically invariant architecture.

23

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

Watermarked

Residual

Original

Watermarked

Residual

TRIAD

VINE

Robust-Wide

EditGuard

TrustMark

SepMark

StegaStamp

Original

Figure 9. Qualitative evaluation of visual fidelity and imperceptibility. From left to right: the original cover panorama, the watermarked panorama, and the pixel-wise residual map.

24

Record · ID 229433 · SHA-256 8845bb1e50f2e7e4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.