Conceptio › Archive › arXiv CS
arXiv CSopen access

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

arXiv:2605.09279v1 [cs.GR] 10 May 2026

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting DAHENG YIN, Simon Fraser University, Canada and Jiangxing Intelligence Inc., China YILI JIN, McGill University, Canada and Simon Fraser University, Canada JIANXIN SHI, Nankai University, China and Simon Fraser University, Canada ISAAC DING, Simon Fraser University, Canada MIAO ZHANG, Simon Fraser University, Canada FANGXIN WANG, The Chinese University of Hong Kong, Shenzhen, China ZHAOWU HUANG, Fuzhou University, China and Southeast University, China CONG ZHANG, Simon Fraser University, Canada and Jiangxing Intelligence Inc., China JIANGCHUAN LIU∗ , Simon Fraser University, Canada FANG DONG, Southeast University, China Volumetric video (VV) streaming delivers truly immersive viewing experiences over the Internet, serving as a critical foundation for next-generation applications, including immersive telepresence in the metaverse, the surveillance of remote ecological systems, and robotic teleoperation for embodied AI, and beyond. Beyond immersive viewing, these applications turn VV streaming into a real-time interface to remote physical environments, imposing new system-level demands for photorealistic scene representation, low-latency interaction, and robust performance under heterogeneous network conditions. 3D Gaussian Splatting (3DGS) has been widely used for real-time photorealistic rendering, offering superior visual quality and rendering performance, but it faces challenges due to bandwidth consumption. Furthermore, as the foundation of adaptive VV streaming, existing Levels of Detail (LoD) methods based on density are not well-suited to Gaussian representations, leading to visible gaps and severe quality degradation. Recent studies have also explored attribute compression techniques to reduce bandwidth consumption. Our preliminary studies reveal that aggressive attribute compression primarily causes color distortion, which can be effectively corrected in the rendered image using a reference image. Motivated by these findings, we propose a novel Color-Adaptive scheme for adaptive VV streaming that uses vector quantization (VQ) to establish LoDs and correct color distortions with low-resolution reference images. We further present CAGS, an adaptive VV streaming system compatible with diverse Gaussian representations, which integrates the Color-Adaptive scheme by rendering reference images on the streaming server and performing color ∗ Corresponding authors.

Authors’ Contact Information: Daheng Yin, Simon Fraser University, Burnaby, Canada and Jiangxing Intelligence Inc., Shenzhen, China, [email protected]; Yili Jin, McGill University, Montreal, Canada and Simon Fraser University, Burnaby, Canada, yili.jin@mail. mcgill.ca; Jianxin Shi, Nankai University, Tianjin, China and Simon Fraser University, Burnaby, Canada, [email protected]; Isaac Ding, Simon Fraser University, Burnaby, Canada, [email protected]; Miao Zhang, Simon Fraser University, Burnaby, Canada, [email protected]; Fangxin Wang, The Chinese University of Hong Kong, Shenzhen, Shenzhen, China, [email protected]; Zhaowu Huang, Fuzhou University, Fuzhou, China and Southeast University, Nanjing, China, [email protected]; Cong Zhang, Simon Fraser University, Burnaby, Canada and Jiangxing Intelligence Inc., Shenzhen, China, [email protected]; Jiangchuan Liu, Simon Fraser University, Burnaby, Canada, [email protected]; Fang Dong, Southeast University, Nanjing, China, [email protected].

This work is licensed under a Creative Commons Attribution 4.0 International License. SIGGRAPH Conference Papers ’26, Los Angeles, CA, USA © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2554-8/2026/07 https://doi.org/10.1145/3799902.3811058

restoration on the client. Extensive experiments on our prototype system demonstrate that CAGS outperforms the existing adaptive streaming systems in PSNR by 5∼20 dB under fluctuating bandwidth, operates significantly faster than existing scalable Gaussian compression methods, and generalizes across different Gaussian representations. The code is available at https://github.com/yindaheng98/ColorAdaptiveGaussianSplatting. CCS Concepts: • Information systems → Multimedia streaming; • Networks → Application layer protocols. Additional Key Words and Phrases: Volumetric Video Streaming, 3D Gaussian Splatting, Color Restoration, Vector Quantization ACM Reference Format: Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong. 2026. CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers (SIGGRAPH Conference Papers ’26), July 19–23, 2026, Los Angeles, CA, USA. ACM, New York, NY, USA, 25 pages. https://doi.org/10.1145/3799902.3811058

1

Introduction

Modern Internet infrastructure and real-time streaming technologies are reshaping online media, evolving it from delivering static images and videos toward live, interactive, photorealistic spatial experiences. Volumetric Video (VV) streaming is central to this evolution, which captures dynamic scenes as evolving 3D content, allowing interactive exploration from arbitrary viewpoints with six degrees of freedom (6DoF). Built on this infrastructure, emerging applications such as immersive communication, live digital twins, and embodied intelligence are turning VV streaming from an immersive viewing medium into a real-time interface to the physical world. This transition introduces new system-level demands: VV streaming must deliver high visual fidelity for accurate interpretation, minimal latency for responsive interaction, and robust performance across heterogeneous network conditions for reliable deployment. Prior studies have established VV streaming as an effective 6DoF medium for real-world scenes, supporting more interactive experiences than conventional video in telepresence, collaboration, education, training, and immersive media [Guan et al. 2023; Han et al. 2020]. These capabilities naturally extend to a wide range of domains SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

Server Volumetric Video time

Internet

Client Receive & Buffer

Time-varying 3D data (e.g., Voxel, Mesh, 3DGS).

Compressed Chunks Selection

bandwidth

2

Decode

time

Feedback

Control, view, ...

Render

Fig. 1. Overview of a typical adaptive VV streaming pipeline. The server maintains time-varying 3D data, selects an appropriate quality-size version of each scene tile according to viewport and network feedback, and streams it to the client for decoding and rendering.

where remote perception, coordination, and action are critical, from healthcare [Gasques et al. 2021] and medical training [Rojas-Muñoz et al. 2019] to digital twins, robotics, autonomous systems [Tao et al. 2018; Zimmer et al. 2024], and environmental or industrial monitoring [de Koning et al. 2023; Hazeleger et al. 2024]. In this sense, VV streaming expands the Internet from a medium for transmitting images, videos, and messages into an infrastructure for delivering dynamic, interactive representations of the physical world. To support such applications, VV streaming systems must deliver interactive 6DoF content with stable geometry for navigation, accurate visuals for interpretation, low latency for responsiveness, and consistent performance across varying network bandwidths. Adaptive VV streaming systems address these requirements by encoding scene tiles at multiple quality-size levels and dynamically selecting optimal versions based on viewport feedback, user interactions, or network status (Figure 1) [Zhang et al. 2022]. As a spatially decomposable and rasterization-friendly technique, 3D Gaussian Splatting (3DGS) [Kerbl et al. 2023] is widely used for real-time photorealistic 3D rendering. It represents scenes using semi-transparent ellipsoids (“Gaussians”) with Spherical Harmonics (SH) for view-dependent photorealism, while enabling real-time rendering via rasterization. Despite these advantages, Gaussians carry complex attributes that produce large data volumes, posing significant challenges for streaming applications [Huang et al. 2026; Kim et al. 2025; Sun et al. 2025; Zhu et al. 2025]. Data reduction typically follows two orthogonal strategies: decreasing the number of Gaussians or compressing their attributes. Existing systems typically support network adaptation through density-based Level of Detail (LoD) [Liu et al. 2023; Shi et al. 2025], which provides progressive scene representations with different numbers of primitives. Since each Gaussian captures fine-grained details within a volumetric region, removing Gaussians can lead to insufficient scene coverage and introduce structural gaps (Figure 2a). Alternatively, attribute compression reduces data by compressing Gaussian attributes. A representative method is Vector Quantization (VQ), which maps high-dimensional Gaussian attributes to compact codebook indices and enables fast index-based decoding [Lee et al. 2024]. Nonetheless, existing attribute compression techniques primarily focus on static compression and lack the quality scalability required for adaptive VV streaming under fluctuating bandwidth.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

(a) LoD 0 (~21 MB)

(b) 40-bit VQ (~14 MB) (c) 100-bit VQ (~34 MB)

Fig. 2. Visual comparison of density-based LoD used in LTS [Sun et al. 2025] and attribute compression by Vector Quantization (VQ) (both compressed further by Draco [Google 2017]). Removing Gaussians causes structural gaps at low bitrates (a), while attribute compression better preserves structure with color distortion (b). Scenes are selected from the N3DV [Li et al. 2022b], Robo360 [Liang et al. 2023], AcinoSet [Joska et al. 2021] and MVSPS [Hein et al. 2025] datasets.

In this paper, we explore scalable attribute compression for adaptive, photorealistic VV streaming based on Gaussian representations. Our preliminary studies (Sec. 3) highlight that attribute compression largely avoids structural gaps but introduces color distortion as the dominant visual artifact. Accurate color representation is critical in streamed 3D scenes, as colors indicate object types and materials in robotics, liquid states and interactions in human contexts, tissue and instrument conditions in clinical settings, and species or environmental specifics in ecological observations (Figure 2b). Distorted color streams may obscure visual information required for accurate human interpretation and effective machine perception. Encouragingly, our preliminary studies further reveal that such color distortion caused by VQ can be effectively corrected in rendered images using a low-resolution reference image with accurate colors. Motivated by this finding, we advocate a Color Adaptation scheme that scalably compresses Gaussian attributes by VQ to establish LoDs, and restores colors using low-resolution reference images. In streaming systems, the server renders these reference images from the high-fidelity Gaussians and streams them to clients to enable color restoration. Implementing this concept in practical streaming systems introduces three challenges. First, adaptive streaming requires scalable

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

VQ rather than traditional flat VQ. We therefore design Scalable Vector Quantization (SVQ) that organizes Gaussian attributes into a base layer and multiple enhancement layers, establishing LoDs with different levels of quantization error. Second, latency constraints require the streaming server to predict the client’s viewport and render reference images in advance. Prediction errors inevitably result in mismatches between reference images and actual client viewports, severely degrading color restoration quality. We address this with Post-Render Perspective Alignment (PRPA), which realigns reference images on the client side using locally rendered depth. Third, prediction errors may leave visible regions uncovered by reference images, creating gaps during PRPA. We introduce an Adaptive Field of View strategy that uses server-side LSTM-based prediction to dynamically adjust the reference field of view (FoV), balancing visible-region coverage and reference quality. Based on these designs, we introduce Color-Adaptive Gaussian Streaming (CAGS), an adaptive VV streaming system compatible with Gaussian representations organized as differential volumetric frames (Sec. 4). This design facilitates integration with diverse Gaussian representations and can benefit from ongoing advances in 3D/4D Gaussian compression. The contributions of this paper are: • We analyze the interaction between VQ and color restoration, revealing an opportunity to build adaptive VV streaming around a novel Color Adaptation scheme. • We identify key system challenges of bringing Color Adaptation into adaptive VV streaming, and address them with SVQ, PRPA, and Adaptive FoV. • We implement and evaluate a CAGS prototype on volumetric video datasets and real-world network traces, demonstrating 5∼20 dB PSNR gains over existing LoD methods and broad generalizability across Gaussian representations.

2 Related Work 2.1 3D Gaussian Splatting for Volumetric Video Streaming As the most popular representation for VV streaming systems [Guan et al. 2023; Liu et al. 2023; Wang et al. 2024b; Zhang et al. 2024], point clouds require high-density data and significant bandwidth (300–800 Mbps) to ensure good visual quality [Han et al. 2020] due to their discrete nature [Han et al. 2020] and lack view-dependent effects that are critical for photorealism (e.g., specular highlights) [Wegen et al. 2024]. While Neural Radiance Fields (NeRF) [Mildenhall et al. 2021] offer superior fidelity with view-dependent effects, they suffer from severe rendering latency [Shi et al. 2024; Yin et al. 2024]. Alternatively, 3D Gaussian Splatting (3DGS) [Kerbl et al. 2023] enables real-time photorealistic rendering [Lee et al. 2024] and has been extended to dynamic scenes [Li et al. 2024; Luiten et al. 2024; Yan et al. 2024b]. Despite its advantages, the large data footprint of high-dimensional Gaussian attributes [Papantonakis et al. 2024] poses significant challenges for bandwidth-constrained streaming. Vector Quantization (VQ) has become popular in recent Gaussian compression and streaming systems [Girish et al. 2024; Li et al. 2025; Wang et al. 2024a], enhancing compression efficiency through advanced codebook designs [Wang et al. 2024a; Xu et al. 2024a] and clustering techniques [Xie et al. 2025; Xu et al. 2024b]. Critically, most existing VQ-based methods prioritize compression efficiency

•

3

for storage, overlooking the scalability indispensable for adaptive streaming [Guan et al. 2023; Han et al. 2020]. While other approaches explore scalable compression for 3DGS [Chen et al. 2025b; Liu et al. 2024], they typically involve complex structures that introduce high decoding latency (Table 4), thus hindering real-time streaming capabilities. To address this gap, we propose a scalable VQ framework that retains the fast decoding speed of VQ while enabling bitrate adaptation for robust VV streaming.

2.2

Level of Detail for Adaptive Streaming

Level of Detail (LoD) has long been used to reduce rendering cost [Cui et al. 2024; Kerbl et al. 2024; Yan et al. 2024a]. In adaptive streaming, LoD also facilitates bitrate adaptation by allowing clients to switch between quality levels according to real-time network conditions. To maximize visual quality without causing playback stalls, existing adaptive VV streaming systems typically combine visibility-aware transmission [Hladky et al. 2019; Zhu et al. 2025] with LoD selection [Han et al. 2020; Liu et al. 2023; Sun et al. 2025], prioritizing content that is visible or likely to become visible. For Gaussian-based representations, most LoD designs are density-based. Representative approaches include progressive training [Kerbl et al. 2024; Lu et al. 2024; Shi et al. 2025] and importance evaluation based on geometric properties [Fan et al. 2024; Papantonakis et al. 2024]. However, since each Gaussian occupies considerable space, low-density layers lack sufficient Gaussians to preserve fine details, resulting in significant quality degradation at lower LoDs. Alternatively, recent scalable neural codecs [Chen et al. 2025a; Liu et al. 2024] offer LoD-like quality-size scalability, but their complex decoding pipelines incur substantial latency (Table 4), limiting their suitability for real-time interaction. Other latency-resilient systems [Hladky et al. 2022; Lu and Rowe 2025] approximate views using geometry proxies, relying on server-side rendering, which increases computational load proportionally to client display resolution. In contrast, our method keeps Gaussian rendering on the client and uses only a fixed low-resolution reference image for color restoration, decoupling server-side rendering cost from the final display resolution.

2.3

Image Restoration

Learning-based image restoration has been widely studied for tasks such as super-resolution [Luo et al. 2022] and color restoration [Bozic et al. 2024]. While related techniques have been combined with neural scene representations to accelerate rendering [Huang et al. 2023; Wang et al. 2022], their potential for mitigating compression artifacts in 3DGS-based VV streaming remains largely unexplored. In this paper, we explore example-based color restoration [Cong et al. 2024; Zhang et al. 2019] to address the color distortion caused by VQ on Gaussians for VV streaming systems.

3

Measurement and Motivation

To motivate our proposed Color Adaptation scheme, we conduct preliminary experiments to evaluate the effectiveness of examplebased color restoration in addressing color distortions introduced by attribute compression. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

4

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

Offline 3DGS Sequence

Server

Scalable VQ (§4.1) Pose FoV LoD

Tiling

Compres s

Client

Compressed 3DGS Sequence 3D Tiles Frustum Culling Reference Rendering

Ref. Image

Distorted Color Restoration

3D Tile Buffer Render Reference Buffer

Depth

PRPA (§4.2) Ref. image

Viewport Prediction / Adaptive FoV (§4.3)

Upload viewport and bandwidth Fig. 3. Overview of CAGS. The server predicts the viewport, selects tiles and LoDs, renders a low-resolution reference image from the highest-quality layer of the compressed 3DGS, and streams it with Gaussian data. The client renders the tiles, aligns the reference image using PRPA, and restores colors for display.

3.2

Measurement Insights

Figure 4 shows that example-based color restoration consistently achieves better visual quality than single-image color restoration and super-resolution. The results indicate that aggressive VQ mainly damages color while preserving much of the scene structure. Hence, example-based color restoration leverages preserved structural details to achieve more accurate and visually appealing results. Moreover, when severe distortion makes the distorted image less trustworthy, example-based restoration relies more on the reference image and resembles super-resolution results. Thus, super-resolution defines the lower bound of example-based restoration. These findings highlight the potential benefit of integrating VQ and example-based color restoration into adaptive VV streaming systems.

4

System Design

Figure 3 provides an overview of the CAGS pipeline. CAGS is compatible with Gaussian representations organized as differential frames, where each frame stores only the Gaussians that differ from the previous frame. In the offline phase, CAGS establishes LoDs by Scalable Vector Quantization (Sec. 4.1). Each quantized frame is then tiled and compressed with Draco [Google 2017]. In the online streaming phase, the server predicts the client viewport from the uploaded viewport history and determines an Adaptive Field of View (Sec. 4.3). Based on the available bandwidth and predicted viewport, the server then renders low-resolution reference images and selects suitable tiles and LoDs. Low-resolution reference images are encoded as a video stream and streamed together with the selected tiles. For each frame, the client renders the received tiles to SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

video:coffee_martini

30

8

25

6

20 40 35 30 25 35 30 25 20

video:cook_spinach

4 6 4 2

video:cut_roasted_beef

6 4 2

Size (MB)

35

Color restoration with ref. Compressed 3DGS size 10

Size (MB)

PSNR (dB) PSNR (dB)

We select three videos from the N3DV dataset [Li et al. 2022b] and train a 3DGS scene on the first frame of each video. We render each scene from interpolated training viewports at two resolutions: 1600×1200 as ground truth and 400×300 as reference images. We compress each 3DGS scene using KMeans VQ at five quality levels, followed by lossless compression with Draco [Google 2017]. After decompression and dequantization, we render the dequantized scenes from the same viewports to obtain color-distorted images. We adapt a lightweight SRResNet [Ledig et al. 2017] for restoration by modifying its input layer to support three settings: 1) color restoration from the distorted image, 2) super-resolution from the reference image, and 3) example-based restoration from both images.

Color restoration Super-resolution

Size (MB)

Experimental Setup

PSNR (dB)

3.1

Config 1 Config 2 Config 3 Config 4 Config 5 Fig. 4. Comparison of compressed frame size and PSNR for super-resolution, single-image color restoration, and example-based color restoration. Configs 1–5 correspond to KMeans VQ settings where scales are quantized to 8/10/12/14/16 bits; rotation and SH are quantized to 4/7/10/13/16 bits; and opacity is fixed at 4 bits.

obtain a color-distorted image and a depth map, aligns the reference image using Post-Render Perspective Alignment (Sec. 4.2), and finally applies a lightweight network to restore colors.

4.1

Scalable Vector Quantization

While KMeans VQ effectively compresses Gaussian attributes [Fan et al. 2024; Lee et al. 2024], its flat clustering structure lacks the scalability required for adaptive streaming. Hierarchical vector quantization has been studied extensively for organizing codebooks across quality levels [Gersho and Shoham 1984]. A representative method is Agglomerative Hierarchical Clustering (AHC), which builds a cluster tree by repeatedly merging the closest clusters. Unfortunately, AHC involves computing pairwise distances in each iteration, which is computationally prohibitive for Gaussian representations that typically contain >100k Gaussians per frame. To bridge this gap, we propose Scalable Vector Quantization (SVQ) that integrates the efficiency of KMeans clustering with the hierarchical structure of AHC, with an index assignment strategy to build a scalable codebook.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

Leaf node

Non-leaf node

0

1 01

00 000

10

11

001 0010

00100

110 0011

00101

1100

11000

111

01 10 replace 00 000 11 111 C0 001 110 C1

1101

01 10 000 111 replace 001 0011 110 1101 0010 1100 C2

11001

(a) Cluster Tree

01 10 000 111 0011 1101 replace 0010 00101 1100 00100 11000 11001 C3

(b) Structure of Codebook

01 10 000 111 0011 1101 00101 00100 11000 11001

•

5

01concat01 01 01 10 10 10 10 00 0 000concat000 000 11 1 111 111 111 00 1 001 1 0011concat0011 11 0 110 1 1101 1101 00 1 001 0 0010 1 00101 00 1 001 0 0010 0 00100 ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ 11 0 110 0 1100 1 11001 L2 L3 L0 L1 (c) Structure of Quantized Data

Fig. 5. Illustration of Scalable Vector Quantization (L0: the base layer quantized data; L1–L3: enhancement layers quantized data; C0–C3: corresponding codebook layers; colors indicate cluster centers; each row in (c) represents a Gaussian).

4.2

 Find Color for Points

Reference

 Get Color for Pixels

Predicted View  Pixels to 3D Points Depth

Local View Fig. 6. PRPA aligns the reference image in three steps: ❶ unprojects clientview pixels using the depth map, ❷ reprojects them into the reference view, and ❸ samples the corresponding colors to produce an aligned image.

4.1.1 Hierarchical Codebook Construction. SVQ first runs KMeans on a random subset of Gaussians, then iteratively merges these clusters into a binary tree, similar to AHC. To preserve fidelity, we define a cluster distance metric based on the quantization error introduced by merging two clusters 𝑑 (𝐶 1, 𝐶 2 ): 𝑑 (𝐶 1, 𝐶 2 ) = mean({𝑐 − mean(𝐶 1 ∪ 𝐶 2 )|𝑐 ∈ 𝐶 1 ∪ 𝐶 2 }) At each iteration, SVQ merges the pair of clusters with the smallest distance, progressively forming a binary tree (Figure 5a), until the number of clusters matches the target codebook size. By starting from KMeans clusters rather than individual Gaussians, SVQ significantly reduces computational complexity compared to AHC. 4.1.2 Scalable Indexing. After building the tree, SVQ assigns indices to clusters to create a scalable codebook. Inspired by Huffman coding, indices are assigned based on cluster positions in the binary tree (Figure 5a). This design provides inherent scalability: truncating lower-order bits naturally maps to its parent cluster at a coarser LoD (Figure 5b). Consequently, the codebook is organized into a base layer to store higher-order index bits and subsequent enhancement layers to store lower-order bits (Figure 5b). 4.1.3 Dequantization. Dequantization involves only concatenating the received bits (Figure 5c) and indexing into the codebook (Figure 5b), and is therefore computationally efficient without becoming a performance bottleneck (Table 2).

Post-Render Perspective Alignment

To support real-time color restoration at high resolutions, CAGS relies on lightweight models such as SRResNet [Ledig et al. 2017]. Our evaluation shows that restoration quality depends on accurate alignment between the reference and distorted images (Sec. 5.4). Unfortunately, misalignment is inevitable in server-side rendering. Due to inherent latency on the Internet (typically 40–90 ms), uploading the client viewport and waiting for the corresponding reference image for every frame would exceed the latency budget for interactive VV streaming (33.3 ms per frame at 30 FPS) [Gül et al. 2022]. Han et al. [Han et al. 2020] have shown that practical systems must predict future viewports to satisfy this constraint. Accordingly, the CAGS server renders low-resolution reference images in advance from the highest-quality level of the compressed Gaussian sequence. The prediction errors inevitably introduce misalignment between the reference and distorted images. To mitigate this issue, we propose Post-Render Perspective Alignment (PRPA). PRPA takes the server-rendered reference image, its rendering viewport, and the client-side depth map produced by the 3DGS rasterizer. It aligns the reference image to the actual client viewport before color restoration. 4.2.1 Reference Image Alignment. As shown in Figure 6, PRPA reprojects pixels from the client-side depth map to the server-rendered reference and samples the corresponding colors to produce an aligned image. Note that PRPA fundamentally differs from Depth Image Based Rendering (DIBR) [Fehn 2004; Wu et al. 2023]. DIBR warps a reference image using its own depth to a target view. In contrast, PRPA aligns the reference image using the target depth. 4.2.2 Error Erosion on Occluded Regions. Naive alignment may map pixels to occluded regions in the reference view, causing visual artifacts (Figure 7b). PRPA reduces these artifacts through error erosion. During reprojection into the reference view, PRPA records the projected depth of each pixel. When multiple pixels map to the same reference pixel, only the pixel with the smallest depth is treated as visible, while the others are marked as occluded. PRPA then iteratively replaces each occluded pixel with the average color of its non-occluded neighbors, yielding a cleaner aligned reference image for color restoration (Figure 7e). Implemented with optimized GPU matrix operations, PRPA introduces negligible overhead that does not bottleneck real-time applications (Table 2). Pseudocode is provided in supplementary. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

6

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

(a) Reference image

(b) Aligned reference image

(a) Small FoV ref. image

(b) Aligned small FoV ref. image

(c) Large FoV ref. image

(d) Aligned large FoV ref. image

Background Occluded

Foreground (c) Before erosion

(d) Erosion Step 3

(e) Erosion Step 30

Fig. 7. Illustration of alignment errors and error erosion in PRPA. (b) Aligned reference image before error erosion, with red circles highlighting occlusion artifacts. (c)-(e) Iterative error erosion reduces these artifacts.

Fig. 8. Visualization of PRPA results with different reference FoVs. A small FoV (a) misses boundary content (b), whereas a large FoV (c) preserves coverage but reduces pixel density in the aligned reference image (d).

4.3

4.3.2 Fast Approximate Ground-truth FoV. Computing the exact ground-truth FoV requires projecting all pixels from the reference image into the client viewport and then finding the smallest FoV that covers them, which is computationally prohibitive. Instead, we approximate the ground-truth FoV by projecting only the four corner pixels of the reference image into the client viewport at a fixed depth (e.g., 10 m). Given that viewport prediction errors are usually small over short intervals [Han et al. 2020], this approximation can greatly reduce computation while maintaining sufficient accuracy. Since the server cannot access the actual client viewport in real-time, this approximate FoV is only used for offline supervision and to update 𝑠𝑖𝑎 once client viewport updates arrive. The FoV for reference rendering is predicted by the Adaptive FoV model.

Adaptive Field of View

Viewport prediction errors may cause the reference image to cover only part of the client viewport, leaving missing regions after PRPA (Figure 8b) and degrading restoration quality. While enlarging the reference FoV improves coverage (Figure 8c), it reduces pixel density in target regions (Figure 8d), which also hurts quality. To balance coverage and density, we propose an Adaptive FoV strategy driven by a lightweight LSTM model. We choose LSTM because it is efficient and effective for sequential, latency-sensitive prediction. Though complex backbones like Transformers could achieve similar accuracy, they usually require additional optimization for real-time use. Notably, our strategy is decoupled from the viewport prediction, providing a plug-and-play solution compatible with various prediction methods [Han et al. 2020; Liu et al. 2023]. In our evaluation, we use autoregression for viewport prediction. 4.3.1 LSTM-based Adaptive FoV Model. The LSTM predicts a scal𝑦 ing factor 𝑠𝑖 = (𝑠𝑖𝑥 , 𝑠𝑖 ) for frame 𝑖, which adjusts the reference FoV 𝑦 (𝐹 𝑥 , 𝐹 𝑦 ) relative to the fixed client viewport FoV (𝐹 0𝑥 , 𝐹 0 ), such that 𝑦 𝑦 𝑥 𝑥 𝑥 𝑦 𝐹 = (1 + 𝑠𝑖 )𝐹 0 and 𝐹 = (1 + 𝑠𝑖 )𝐹 0 . The model takes the client viewport (rotation 𝑞𝑖 and position 𝑝𝑖 ), their temporal changes, and historical states as inputs. Formally, the predicted scaling factor 𝑠ˆ𝑖 and hidden state ℎ𝑖 are computed as: (𝑠ˆ𝑖 , ℎ𝑖 ) = LSTM(𝑞𝑖 , 𝑞𝑖 − 𝑞𝑖 −1, 𝑝𝑖 , 𝑝𝑖 − 𝑝𝑖 −1, 𝑠𝑖𝑎−1, ℎ𝑖 −1 ) where 𝑠𝑖𝑎−1 denotes the approximate ground-truth FoV scaling factor from the previous frame (Sec. 4.3.2). In a real streaming system, viewport predictions typically span multiple future frames [Liu et al. 2023], where actual client viewports are unavailable. In this case, the model uses predicted viewports as 𝑞𝑖 and 𝑝𝑖 , and uses the previously predicted FoV scaling factor 𝑠ˆ𝑖 −1 as 𝑠𝑖𝑎−1 . To limit error accumulation, we refresh the LSTM hidden state whenever actual client viewport updates arrive. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

4.3.3 LSTM Model Training. We train the LSTM model offline using collected viewport datasets. First, we run the viewport prediction on ground-truth viewports to obtain predicted viewports. Then we compute the ground-truth FoV scaling factors 𝑠𝑖 and their approximations 𝑠𝑖𝑎 using the method above. The training loss L is computed between 𝑠ˆ𝑖 and 𝑠𝑖 over all 𝑁 frames: 𝑁

L=

1 ∑︁ |𝑠ˆ𝑖 − 𝑠𝑖 | 𝑁 𝑖=1

5 Evaluation 5.1 Prototype Implementation We implemented a CAGS prototype to evaluate its effectiveness, following prior VV streaming practice [Sun et al. 2025; Wang et al. 2024a]. We prepare volumetric videos using TrackerSplat [Yin et al. 2025] and apply importance-based pruning [Papantonakis et al. 2024; Zhou et al. 2024] to remove redundant Gaussians. To enable adaptive streaming, we build a linear LoD hierarchy by interleaving

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

SVQ layers according to the empirically measured visual impact of Gaussian attributes, prioritizing high-impact layers at lower LoDs. We tile Gaussians via Morton sorting [Jiang et al. 2025] to balance the number of Gaussians per tile. Spatial data and base LoD are compressed with Draco [Google 2017], while enhancement layers and the codebook are compressed with Gzip. During streaming, the server applies a bandwidth-aware adaptation strategy that prioritizes tiles covering visible Gaussians (identified via reference rendering) and progressively raises their LoDs until reaching the bandwidth limit. Our prototype is deployed on consumer-grade GPUs (RTX 3080) for both server and client, streaming at 30 FPS while rendering at 60 FPS on the client side. Additional engineering details and underlying rationales are provided in the supplementary.

5.2

Evaluation Setup

5.2.1 Dataset. We evaluate CAGS on four datasets with diverse motion patterns, scene scales, and capture resolutions: Neural 3D Video Synthesis (N3DV) [Li et al. 2022b], ST-NeRF [Zhang et al. 2021], Meeting Room [Li et al. 2022a], and Dynamic 3DGS [Luiten et al. 2024]. We adapt UnityGS [Pranckevičius 2023] and develop an application to record viewport traces using a Meta Quest 3. The traces used to train color restoration and FoV prediction are collected independently from those used for evaluation. 5.2.2 Metrics. We render the original uncompressed Gaussians at the corresponding viewports as ground truth. We report PSNR, SSIM, and LPIPS [Zhang et al. 2018].

Bandwidth (Mbps)

5.2.3 Network Settings. We evaluate both fixed and dynamic bandwidth settings. For fixed bandwidth, we set the bandwidth limit (Sec. B.6) to 30, 60, 90, 120, and 150 Mbps. For dynamic bandwidth, we use a representative segment from 5Gophers [Narayanan et al. 2020], as shown in Figure 9.

200

0 0

50

100

150 200 Frame Index

250

300

Fig. 9. Throughput of the selected network trace from the 5Gophers dataset (lines 3249 to 3549, 1.47–251.7 Mbps, average 133.7 Mbps).

5.2.4 Baselines. We compare CAGS with two state-of-the-art Gaussianbased VV streaming systems and two restoration variants, specifically focusing on methods that are compatible with the real-time client loop required for interactive streaming. LTS-F: LTS [Sun et al. 2025] is an adaptive VV streaming system based on density-based LoD. We use its released LoD component and implement its corresponding network adaptation strategy as the baseline LTS-F. For consistency with Sec. B, we encode the sequence into independently decodable Groups of Frames (GoF) and select the LoD for each frame during streaming. V3 -A: V3 [Wang et al. 2024a] is a VV streaming system based on VQ (hash grids trained with entropy loss). It does not support LoD and thus lacks a built-in network adaptation strategy. For the V3 -A

•

7

baseline, we integrate its released code as a compression module into our pipeline (including PRPA and adaptive FoV for fairness), and implement network adaptation in Sec. B.6, except that alreadytransmitted tiles are resent at higher quality until the bandwidth limit is reached. SR align ref.: Uses SRResNet [Ledig et al. 2017] for super-resolution on the reference image aligned by PRPA. CR w/o ref.: Performs color restoration (Sec. 3.1) directly on the distorted image without a reference.

5.3

Evaluation Results

5.3.1 Evaluation of Scalable Vector Quantization. Before evaluating the full system, we benchmark SVQ against state-of-the-art scalable compression methods: SPZ [Niantic Labs 2025], CompGS [Liu et al. 2024], HAC [Chen et al. 2025a], and HAC++ [Chen et al. 2025b]. As shown in Table 4, SVQ achieves comparable rate-distortion performance to these methods and outperforms them in decoding latency, which is uniquely suited for real-time streaming. 5.3.2 Evaluation under Fixed Bandwidth. Figure 10 reports the average PSNR and SSIM under fixed bandwidth constraints. CAGS achieves higher PSNR than the baselines in most cases, indicating better color fidelity. LTS-F is limited by density-based LoD, which provides a weaker rate–quality trade-off and therefore selects lowerquality LoDs under the same bandwidth. V3 -A lacks scalability and causes redundant transmissions, resulting in lower quality. In some cases such as the “walking” (Figure 10c), LTS-F achieves higher SSIM at high bandwidth. Analysis shows this sequence contains fewer Gaussians, allowing density-based LoD to approach uncompressed quality when bandwidth is sufficient, while CAGS and V3 -A remain bounded by lossy VQ compression. 5.3.3 Evaluation under Fluctuating Bandwidth. Figure 11 reports per-frame PSNR/SSIM under the fluctuating 5G trace. CAGS achieves higher and more stable quality over time, demonstrating robustness to bandwidth variations. Similar to fixed bandwidth results, the “walking” sequence favors the baseline due to its fewer Gaussians. Figure 11 also reveals occasional quality drops (e.g., frame 32 of “basketball” and frame 240 of “discussion”). These drops are caused by rapid head movements, which increase viewport prediction errors and lead to either large missing regions in the PRPA output or overly large FoV predictions. While noticeable in measurements, these degradations have minimal perceptual impact because: 1) they are short-lived, since humans cannot move rapidly for long, and once movement stabilizes, prediction errors drop and quality quickly recovers; and 2) rapid movements naturally limit human visual perception, making such temporary degradations less noticeable. 5.3.4 Visualization Results. Figure 12 shows a representative example of color restoration. The color distortion in Figure 12a is effectively corrected in Figure 12c, demonstrating the effectiveness of our restoration method. 5.3.5 System Performance. We profile key components on both server and client. Table 2 summarizes the profiling results, demonstrating that CAGS supports real-time streaming and rendering. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

8

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

Table 1. Performance and quality comparison of our SVQ method (with 66-bit and 76-bit initialization) against state-of-the-art scalable compression methods at the highest quality level. Best and second-best results are highlighted in bold and underline, respectively. Size is in MB and decoding time ("dec.t") in seconds. Full results including SSIM and LPIPS are provided in the supplementary material.

Table 2. Performance of system components.

Encoding (Offline) Server

Datasets Neural3DV Meet.Room Dyn.3DGS Methods psnr size dec.t psnr size dec.t psnr size dec.t SPZ low 23.3 3.33 0.076 21.8 2.54 0.062 21.1 2.26 0.052 SPZ high 23.1 4.53 0.075 21.8 3.41 0.057 21.0 3.32 0.057 CompGS 24.9 22.20 6.213 25.0 5.58 0.799 19.1 4.27 2.772 HAC 22.2 25.69 24.407 25.2 8.69 9.135 20.5 2.84 1.571 HAC++ 21.5 17.98 51.708 25.2 6.78 14.706 20.9 2.00 6.169 SVQ 66bit 24.5 2.00 0.015 26.0 1.53 0.018 25.3 2.25 0.028 SVQ 76bit 25.1 2.14 0.015 26.4 1.69 0.019 23.3 2.42 0.029

5.4

Ablation Studies

We conduct two ablation studies to quantify the contributions of key components: 1) w/o PRPA: feeds the misaligned reference image directly to restoration. 2) w/o Adaptive FoV: fixes the reference rendering FoV to 10% larger than the client viewport. Figures 10 and 11 show that removing PRPA leads to a significant quality drop, confirming that accurate reference–target alignment is essential for effective restoration. Adaptive FoV also provides meaningful improvements, indicating that sufficient boundary coverage is important for reducing the impact of viewport prediction errors. These components effectively handle viewport prediction errors: PRPA aligns reference images rendered from predicted viewpoints with the actual viewport, enabling accurate color restoration. Adaptive FoV expands the reference FoV according to head motion, ensuring complete coverage. Together, they ensure the quality stability observed in Sec. 5.3.3.

5.5

Generality to Other Representations

CAGS is designed for Gaussian representations where each frame stores the Gaussians that differ from the previous frame. To validate its generality, we integrate CAGS with three representative methods for preparing volumetric videos: Dynamic 3DGS [Luiten et al. 2024], 4DGS [Wu et al. 2024] and HiCoM [Gao et al. 2024], and evaluate them under the same bandwidth trace using the baselines (SR align ref., and CR w/o ref.) and ablations (w/o PRPA, w/o Adaptive FoV). Figures 13 and 14 show that CAGS consistently improves quality across different preparation methods. These results suggest that CAGS can extend beyond the tested methods, potentially supporting a wide range of 3D/4D Gaussian representations and benefiting from ongoing evolutions in 3D/4D Gaussian compression.

6

Limitation and Future Work

CAGS has certain limitations and leaves room for improvement. Error erosion reduces color distortion in occluded regions but still leaves artifacts in the PRPA output. To ensure real-time performance, we did not complicate our restoration model to specifically handle these artifacts. Although our lightweight restoration model can learn to suppress many of them during training, small artifacts may SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Client Decoding Client Rendering (1600x1200) Client Rendering (1600x1200, RTX 4090)

Component SVQ Codebook (per-video) SVQ & Draco Encoding Viewport Prediction Dynamic FoV Rendering (400x300) Draco Decoding SVQ Decoding Render Distorted+Depth PRPA Color Restoration Render Distorted+Depth PRPA Color Restoration

Time 36.8 s 370 ms 1 ms 1 ms 1.7 ms 9 ms 2.58 ms 9.32 ms 2.11 ms 6.5 ms 1.21 ms 1 ms 3.3 ms

still remain (can be observed in our provided video results). Future work can explore artifact handling without sacrificing speed. Although our VQ design is developed for Gaussian representations, VQ itself is more general and can also be applied to other 3D representations. Prior works have shown its effectiveness for NeRFs [Zhong et al. 2024] and also report color distortion as a side effect [Takikawa et al. 2022], suggesting that such distortion may be a common issue inherent to VQ across different 3D representations. This highlights the broader potential of the Color Adaptation scheme beyond Gaussian-based streaming: it could serve as a general solution for streaming 3D content across diverse scene representations. Future work can explore integrating the color-adaptive scheme with a wider range of 3D representations.

7

Conclusion

In this paper, we present a Color Adaptation scheme for volumetric video streaming based on dynamic 3D Gaussian Splatting. We first identify limitations of existing LoD methods for Gaussian representations. We reveal that vector quantization primarily causes color distortion, which can be effectively corrected using reference images. We also identify the challenges of implementing Color Adaptation in real systems and address them with the CAGS system. Extensive evaluation of our prototype demonstrates the effectiveness and efficiency of our design.

References Vukasin Bozic, Abdelaziz Djelouah, Yang Zhang, Radu Timofte, Markus Gross, and Christopher Schroers. 2024. Versatile Vision Foundation Model for Image and Video Colorization. In ACM SIGGRAPH 2024 Conference Papers (SIGGRAPH ’24). 1–11. Yihang Chen, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, and Jianfei Cai. 2025a. HAC: Hash-Grid Assisted Context for 3D Gaussian Splatting Compression. In Computer Vision – ECCV 2024, Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.). 422–438. Yihang Chen, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, and Jianfei Cai. 2025b. HAC++: Towards 100X Compression of 3D Gaussian Splatting. doi:10.48550/arXiv.2501.12255 Xiaoyan Cong, Yue Wu, Qifeng Chen, and Chenyang Lei. 2024. Automatic Controllable Colorization via Imagination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2609–2619. Jiadi Cui, Junming Cao, Fuqiang Zhao, Zhipeng He, Yifan Chen, Yuhui Zhong, Lan Xu, Yujiao Shi, Yingliang Zhang, and Jingyi Yu. 2024. LetsGo: Large-Scale Garage Modeling and Rendering via LiDAR-Assisted Gaussian Primitives. In SIGGRAPH Asia 2024 Conference Papers (SA ’24). Koen de Koning, Jeroen Broekhuijsen, Ingolf Kühn, Otso Ovaskainen, Franziska Taubert, Dag Endresen, Dmitry Schigel, and Volker Grimm. 2023. Digital twins: dynamic

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

model-data fusion for ecology. Trends in ecology & evolution 38, 10 (2023), 916–926. Zhiwen Fan, Kevin Wang, Kairun Wen, Zehao Zhu, Dejia Xu, and Zhangyang Wang. 2024. LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS. Advances in Neural Information Processing Systems 37 (2024), 140138– 140158. Christoph Fehn. 2004. Depth-image-based rendering (DIBR), compression, and transmission for a new approach on 3D-TV. In Stereoscopic Displays and Virtual Reality Systems XI, Vol. 5291. International Society for Optics and Photonics, SPIE, 93 – 104. Qiankun Gao, Jiarui Meng, Chengxiang Wen, Jie Chen, and Jian Zhang. 2024. HiCoM: Hierarchical Coherent Motion for Dynamic Streamable Scenes with 3D Gaussian Splatting. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. Danilo Gasques, Janet G Johnson, Tommy Sharkey, Yuanyuan Feng, Ru Wang, Zhuoqun Robin Xu, Enrique Zavala, Yifei Zhang, Wanze Xie, Xinming Zhang, et al. 2021. Artemis: A collaborative mixed-reality system for immersive surgical telementoring. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–14. A. Gersho and Y. Shoham. 1984. Hierarchical vector quantization of speech with dynamic codebook allocation. In ICASSP ’84. IEEE International Conference on Acoustics, Speech, and Signal Processing, Vol. 9. 416–419. Sharath Girish, Tianye Li, Amrita Mazumdar, Abhinav Shrivastava, David Luebke, and Shalini De Mello. 2024. QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. Google. 2017. Draco 3D Graphics Compression. https://github.com/google/draco. Yongjie Guan, Xueyu Hou, Nan Wu, Bo Han, and Tao Han. 2023. MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real Time. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking. Number 29. 1–15. Serhan Gül, Cornelius Hellge, and Peter Eisert. 2022. Latency Compensation Through Image Warping For Remote Rendering-Based Volumetric Video Streaming. In 2022 IEEE International Conference on Image Processing (ICIP). 2026–2030. Bo Han, Yu Liu, and Feng Qian. 2020. ViVo: Visibility-Aware Mobile Volumetric Video Streaming. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking. 1–13. W Hazeleger, JPM Aerts, Peter Bauer, MFP Bierkens, Gustau Camps-Valls, MM Dekker, FJ Doblas-Reyes, Veronika Eyring, C Finkenauer, Arthur Grundner, et al. 2024. Digital twins of the Earth with and for humans. Communications earth & environment 5, 1 (2024), 463. Jonas Hein, Nicola Cavalcanti, Daniel Suter, Lukas Zingg, Fabio Carrillo, Lilian Calvet, Mazda Farshad, Nassir Navab, Marc Pollefeys, and Philipp Fürnstahl. 2025. Nextgeneration surgical navigation: Marker-less multi-view 6DoF pose estimation of surgical instruments. Medical Image Analysis 103 (2025), 103613. Jozef Hladky, Hans-Peter Seidel, and Markus Steinberger. 2019. The camera offset space: real-time potentially visible set computations for streaming rendering. ACM Trans. Graph. 38, 6 (2019), 14 pages. Jozef Hladky, Michael Stengel, Nicholas Vining, Bernhard Kerbl, Hans-Peter Seidel, and Markus Steinberger. 2022. QuadStream: A Quad-Based Scene Streaming Architecture for Novel Viewpoint Reconstruction. ACM Trans. Graph. 41, 6 (2022), 13 pages. Xudong Huang, Wei Li, Jie Hu, Hanting Chen, and Yunhe Wang. 2023. RefSR-NeRF: Towards High Fidelity and Super Resolution View Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8244–8253. Zhaohui Huang, Cong Zhang, Jianxin Shi, Xiaoyi Fan, Laizhong Cui, and Jiangchuan Liu. 2026. ACPGS: Towards Bandwidth-Efficient Delivery of 3D Gaussian Splatting. In Proceedings of the 36th Workshop on Network and Operating System Support for Digital Audio and Video (NOSSDAV ’26). 127–133. Yuheng Jiang, Chengcheng Guo, Yize Wu, Yu Hong, Shengkun Zhu, Zhehao Shen, Yingliang Zhang, Shaohui Jiao, Zhuo Su, Lan Xu, Marc Habermann, and Christian Theobalt. 2025. Topology-Aware Optimization of Gaussian Primitives for HumanCentric Volumetric Videos. In Proceedings of the SIGGRAPH Asia 2025 Conference Papers (SA Conference Papers ’25). 1–12. Daniel Joska, Liam Clark, Naoya Muramatsu, Ricardo Jericevich, Fred Nicolls, Alexander Mathis, Mackenzie W. Mathis, and Amir Patel. 2021. AcinoSet: A 3D Pose Estimation Dataset and Baseline Models for Cheetahs in the Wild. arXiv:2103.13282 [cs.CV] Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 2023. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics 42, 4 (2023), 139:1–139:14. Bernhard Kerbl, Andreas Meuleman, Georgios Kopanas, Michael Wimmer, Alexandre Lanvin, and George Drettakis. 2024. A Hierarchical 3D Gaussian Representation for Real-Time Rendering of Very Large Datasets. ACM Transactions on Graphics 44, 3 (2024). Gunjoong Kim, Seonghoon Park, Jeho Lee, Chanyoung Jung, Hyungchol Jun, and Hojung Cha. 2025. Vega: Fully Immersive Mobile Volumetric Video Streaming with 3D Gaussian Splatting. In Proceedings of the 31st Annual International Conference on Mobile Computing and Networking. 1106–1120. Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang,

•

9

and Wenzhe Shi. 2017. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4681–4690. Joo Chan Lee, Daniel Rho, Xiangyu Sun, Jong Hwan Ko, and Eunbyung Park. 2024. Compact 3D Gaussian Representation for Radiance Field. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 21719–21728. Hao Li, Sicheng Li, Xiang Gao, Abudouaihati Batuer, Lu Yu, and Yiyi Liao. 2025. GIFStream: 4D Gaussian-based Immersive Video with Feature Stream. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jiyang Li, Lechao Cheng, Zhangye Wang, Tingting Mu, and Jingxuan He. 2024. LoopGaussian: Creating 3D Cinemagraph with Multi-view Images via Eulerian Motion Field. In Proceedings of the 32nd ACM International Conference on Multimedia (MM ’24). 476–485. Lingzhi Li, Zhen Shen, Zhongshu Wang, Li Shen, and Ping Tan. 2022a. Streaming Radiance Fields for 3D Video Synthesis. Advances in Neural Information Processing Systems 35 (2022), 13485–13498. Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, and Zhaoyang Lv. 2022b. Neural 3D Video Synthesis From Multi-View Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5521–5531. Litian Liang, Liuyu Bian, Caiwei Xiao, Jialin Zhang, Linghao Chen, Isabella Liu, Fanbo Xiang, Zhiao Huang, and Hao Su. 2023. Robo360: a 3D omnispective multi-material robotic manipulation dataset. arXiv preprint arXiv:2312.06686 (2023). Junhua Liu, Boxiang Zhu, Fangxin Wang, Yili Jin, Wenyi Zhang, Zihan Xu, and Shuguang Cui. 2023. CaV3: Cache-assisted Viewport Adaptive Volumetric Video Streaming. In 2023 IEEE Conference Virtual Reality and 3D User Interfaces (VR). 173– 183. Xiangrui Liu, Xinju Wu, Pingping Zhang, Shiqi Wang, Zhu Li, and Sam Kwong. 2024. CompGS: Efficient 3D Scene Representation via Compressed Gaussian Splatting. In Proceedings of the 32nd ACM International Conference on Multimedia (MM ’24). 2936–2944. Edward Lu and Anthony Rowe. 2025. QUASAR: Quad-based Adaptive Streaming And Rendering. ACM Trans. Graph. 44, 4 (2025), 18 pages. Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. 2024. Scaffold-Gs: Structured 3d Gaussians for View-Adaptive Rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20654–20664. Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. 2024. Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis. In 3DV. Zhengxiong Luo, Yan Huang, Shang Li, Liang Wang, and Tieniu Tan. 2022. Learning the Degradation Distribution for Blind Image Super-Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6063–6072. Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. 2021. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Commun. ACM 65, 1 (2021), 99–106. Arvind Narayanan, Eman Ramadan, Jason Carpenter, Qingxu Liu, Yu Liu, Feng Qian, and Zhi-Li Zhang. 2020. A First Look at Commercial 5G Performance on Smartphones. In Proceedings of The Web Conference 2020 (WWW ’20). 894–905. Niantic Labs. 2025. spz: File Format for 3D Gaussian Splats. https://github.com/nianticlabs/spz. Panagiotis Papantonakis, Georgios Kopanas, Bernhard Kerbl, Alexandre Lanvin, and George Drettakis. 2024. Reducing the Memory Footprint of 3D Gaussian Splatting. Proceedings of the ACM on Computer Graphics and Interactive Techniques 7, 1 (2024), 16:1–16:17. Aras Pranckevičius. 2023. Gaussian Splatting playground in Unity. https://github.com/ aras-p/UnityGaussianSplatting. Edgar Rojas-Muñoz, Maria Eugenia Cabrera, Daniel Andersen, Voicu Popescu, Sherri Marley, Brian Mullis, Ben Zarzaur, and Juan Wachs. 2019. Surgical telementoring without encumbrance: a comparative study of see-through augmented reality-based approaches. Annals of surgery 270, 2 (2019), 384–389. Jianxin Shi, Miao Zhang, Linfeng Shen, Jiangchuan Liu, Yuan Zhang, Lingjun Pu, and Jingdong Xu. 2024. Towards Full-scene Volumetric Video Streaming via Spatially Layered Representation and NeRF Generation. In Proceedings of the 34th Edition of the Workshop on Network and Operating System Support for Digital Audio and Video (NOSSDAV ’24). 22–28. Yuang Shi, Géraldine Morin, Simone Gasparini, and Wei Tsang Ooi. 2025. LapisGS: Layered Progressive 3D Gaussian Splatting for Adaptive Streaming. In International Conference on 3D Vision 2025. Yuan-Chun Sun, Yuang Shi, Cheng-Tse Lee, Mufeng Zhu, Wei Tsang Ooi, Yao Liu, Chun-Ying Huang, and Cheng-Hsin Hsu. 2025. LTS: A DASH Streaming System for Dynamic Multi-Layer 3D Gaussian Splatting Scenes. In Proceedings of the 16th ACM Multimedia Systems Conference (MMSys ’25). 136–147. Towaki Takikawa, Alex Evans, Jonathan Tremblay, Thomas Müller, Morgan McGuire, Alec Jacobson, and Sanja Fidler. 2022. Variable Bitrate Neural Fields. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Proceedings. 1–9.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

10

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

Fei Tao, He Zhang, Ang Liu, and Andrew YC Nee. 2018. Digital twin in industry: State-of-the-art. IEEE Transactions on industrial informatics 15, 4 (2018), 2405–2415. Chen Wang, Xian Wu, Yuan-Chen Guo, Song-Hai Zhang, Yu-Wing Tai, and Shi-Min Hu. 2022. NeRF-SR: High Quality Neural Radiance Fields Using Supersampling. In Proceedings of the 30th ACM International Conference on Multimedia (MM ’22). 6445–6454. Penghao Wang, Zhirui Zhang, Liao Wang, Kaixin Yao, Siyuan Xie, Jingyi Yu, Minye Wu, and Lan Xu. 2024a. V^3: Viewing Volumetric Videos on Mobiles via Streamable 2D Dynamic Gaussians. In SIGGRAPH Asia 2024 Conference Papers. Yizong Wang, Dong Zhao, Huanhuan Zhang, Teng Gao, Zixuan Guo, Chenghao Huang, and Huadong Ma. 2024b. Bandwidth-Efficient Mobile Volumetric Video Streaming by Exploiting Inter-Frame Correlation. IEEE Transactions on Mobile Computing 23, 10 (2024), 9410–9423. Ole Wegen, Willy Scheibel, Matthias Trapp, Rico Richter, and Jurgen Dollner. 2024. A Survey on Non-photorealistic Rendering Approaches for Point Cloud Visualization. IEEE Transactions on Visualization and Computer Graphics (2024), 1–20. Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 2024. 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20310–20320. Jiangkai Wu, Yu Guan, Qi Mao, Yong Cui, Zongming Guo, and Xinggong Zhang. 2023. ZGaming: Zero-Latency 3D Cloud Gaming by Image Prediction. In Proceedings of the ACM SIGCOMM 2023 Conference (ACM SIGCOMM ’23). 710–723. Shuzhao Xie, Jiahang Liu, Weixiang Zhang, Shijia Ge, Sicheng Pan, Chen Tang, Yunpeng Bai, Cong Zhang, Xiaoyi Fan, and Zhi Wang. 2025. SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming. In Proceedings of the 33rd ACM International Conference on Multimedia (MM ’25). Association for Computing Machinery, New York, NY, USA, 8214–8223. doi:10.1145/3746027.3755370 Jiawei Xu, Zexin Fan, Jian Yang, and Jij Xie. 2024a. Grid4D: 4D Decomposed Hash Encoding for High-Fidelity Dynamic Gaussian Splatting. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS ’24, Vol. 37). 123787–123811. Zhen Xu, Yinghao Xu, Zhiyuan Yu, Sida Peng, Jiaming Sun, Hujun Bao, and Xiaowei Zhou. 2024b. Representing Long Volumetric Video with Temporal Gaussian Hierarchy. ACM Trans. Graph. 43, 6 (2024), 171:1–171:18. Jinbo Yan, Rui Peng, Luyang Tang, and Ronggang Wang. 2024b. 4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes. In Proceedings of the 32nd ACM International Conference on Multimedia (MM ’24). 7871–7880.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Zhiwen Yan, Weng Fei Low, Yu Chen, and Gim Hee Lee. 2024a. Multi-Scale 3D Gaussian Splatting for Anti-Aliased Rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20923–20931. Daheng Yin, Isaac Ding, Yili Jin, Jianxin Shi, and Jiangchuan Liu. 2025. TrackerSplat: Exploiting Point Tracking for Fast and Robust Dynamic 3D Gaussians Reconstruction. In Proceedings of the SIGGRAPH Asia 2025 Conference Papers (SA Conference Papers ’25). 1–11. Daheng Yin, Jianxin Shi, Miao Zhang, Zhaowu Huang, Jiangchuan Liu, and Fang Dong. 2024. FSVFG: Towards Immersive Full-Scene Volumetric Video Streaming with Adaptive Feature Grid. In Proceedings of the 32nd ACM International Conference on Multimedia (MM ’24). 11089–11098. Anlan Zhang, Chendong Wang, Bo Han, and Feng Qian. 2022. YuZu: Neural-Enhanced Volumetric Video Streaming. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). 137–154. Anlan Zhang, Chendong Wang, Yuming Hu, Ahmad Hassan, Zejun Zhang, Bo Han, Feng Qian, and Shichang Xu. 2024. Habitus: Boosting Mobile Immersive Content Delivery through Full-body Pose Tracking and Multipath Networking. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 1677–1695. Bo Zhang, Mingming He, Jing Liao, Pedro V. Sander, Lu Yuan, Amine Bermak, and Dong Chen. 2019. Deep Exemplar-Based Video Colorization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8052–8061. Jiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao, Yanshun Zhang, Minye Wu, Yingliang Zhang, Lan Xu, and Jingyi Yu. 2021. Editable Free-Viewpoint Video Using a Layered Neural Representation. ACM Transactions on Graphics 40, 4 (2021), 149:1–149:18. Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR. Hongliang Zhong, Jingbo Zhang, and Jing Liao. 2024. VQ-NeRF: Neural Reflectance Decomposition and Editing With Vector Quantization. IEEE Transactions on Visualization and Computer Graphics 30, 9 (2024), 6247–6260. Zhi Zhou, Junke Zhu, and Zhangjin Huang. 2024. Gaussian Splatting with Neural Basis Extension. In Proceedings of the 32nd ACM International Conference on Multimedia (MM ’24). 6043–6052. Mufeng Zhu, Mingju Liu, Cunxi Yu, Cheng-Hsin Hsu, and Yao Liu. 2025. SGSS: Streaming 6-DoF Navigation of Gaussian Splat Scenes. In Proceedings of the 16th ACM Multimedia Systems Conference (MMSys ’25). 46–56. Walter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou, Rui Song, and Alois C Knoll. 2024. Tumtraf v2x cooperative perception dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 22668–22677.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

LPIPS

SSIM

PSNR

Ours

LTS-F

V3-A

w/o PRPA

w/o Adaptive FoV

CR w/o ref.

32.1

23.8

29

29.7

31.9

25.9

27.1

18.8

24

24.6

26.8

20.8

22.1

13.8

19

19.5

21.7

15.7

17.1

8.8

14

14.4

16.6

0.91

0.92

0.89

0.92

0.90

0.92

0.82

0.83

0.80

0.83

0.81

0.83

0.73

0.74

0.71

0.74

0.72

0.74

0.64

0.65

0.62

0.65

0.63

0.65

0.39

0.38

0.42

0.40

0.41

0.41

0.30

0.29

0.33

0.31

0.32

0.32

0.21

0.20

0.24

0.22

0.23

0.23

0.12

0.11

0.15

0.13

0.14

0.14

50 100 150 Bandwidth (Mbps)

(a) “cook spinach”

50 100 150 Bandwidth (Mbps)

(b) “cut roasted beef”

50 100 150 Bandwidth (Mbps)

(c) “walking”

50 100 150 Bandwidth (Mbps)

(d) “discussion”

11

SR align.ref.

31.0

50 100 150 Bandwidth (Mbps)

•

50 100 150 Bandwidth (Mbps)

(e) “stepin”

(f) “basketball”

Fig. 10. Comparison of visual quality under fixed bandwidth. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. Results for all videos are included in the supplementary material.

LPIPS

SSIM

PSNR

Ours

LTS-F

V3-A

w/o PRPA

w/o Adaptive FoV

CR w/o ref.

SR align.ref.

33.1

31.8

28.2

30.2

30.7

32.6

26.8

25.5

21.9

23.9

24.4

26.3

20.5

19.2

15.6

17.6

18.1

20.0

14.2

12.9

9.3

11.3

11.8

13.7

0.88

0.88

0.87

0.88

0.75

0.75

0.74

0.75

0.77

0.76

0.63

0.62

0.62

0.62

0.61

0.62

0.49

0.48

0.49

0.49

0.48

0.49

0.42

0.43

0.42

0.44

0.43

0.44

0.32

0.33

0.32

0.34

0.33

0.34

0.22

0.23

0.22

0.24

0.23

0.24

0.12

0.13

0.12

0.14

0.13

0.14

0

100 200 Frame Index

(a) “cook spinach”

300

0

100 200 Frame Index

300

(b) “cut roasted beef”

0

20 40 60 Frame Index

(c) “walking”

0

100 200 Frame Index

(d) “discussion”

300

0

100 200 Frame Index

(e) “stepin”

300

0

50 100 Frame Index

150

(f) “basketball”

Fig. 11. Comparison of visual quality under fluctuating bandwidth. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. Results for all videos are included in the supplementary material.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

12

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

(a) Color-distorted (1600x1200)

(b) Reference image (400x300)

(c) Color-restored (1600x1200)

(d) Ground truth (1600x1200)

Fig. 12. Visual results of the frame 64 in “coffee martini” under 30Mbps. Additional visual and video results are included in supplementary materials.

PSNR

Ours

w/o PRPA

w/o Adaptive FoV

CR w/o ref.

SR align.ref.

29.3

33.3

25.4

29.7

28.7

31.4

25.5

29.5

21.6

25.9

24.9

27.6

21.7

25.7

17.8

22.1

21.1

23.8

17.9

21.9

14.0

18.3

17.3

20.0

0.91

0.93

0.89

0.92

0.90

0.87

0.89

0.85

0.88

0.86

0.83

0.85

0.81

0.84

0.82

0.79

0.81

0.77

0.80

0.78

SSIM

0.95

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(b) “cut roasted beef”

0.85 0.80

50 100 150 Bandwidth (Mbps)

(c) “walking”

50 100 150 Bandwidth (Mbps)

(d) “discussion”

50 100 150 Bandwidth (Mbps)

(e) “stepin”

(f) “basketball”

LPIPS

(a) “cook spinach”

50 100 150 Bandwidth (Mbps)

0.90

Fig. 13. Comparison of visual quality under fixed bandwidth on volumetric videos prepared by HiCoM [Gao et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. Results for all videos (including volumetric videos prepared by Dynamic 3DGS [Luiten et al. 2024], 4DGS [Wu et al. 2024] and HiCoM [Gao et al. 2024]) under fixed bandwidth conditions are included in the supplementary material.

PSNR

Ours

w/o PRPA

w/o Adaptive FoV

SR align.ref.

32.5

32.8

28.9

30.7

30.3

27.7

27.8

26.1

24.9

25.9

29.8

22.9

22.8

23.3

19.1

21.5

24.6

18.1

17.8

20.5

13.3

17.1

19.4

0.93

0.92

0.92

0.91

0.94

0.86

0.84

0.86

0.86

0.90

0.80

0.81

0.86

0.74

0.76

0.82

0.92 SSIM

CR w/o ref.

0.89 0.76

0.79

0.86 0.68

0.72 0

100 200 Frame Index

0

100 200 Frame Index

300

(b) “cut roasted beef”

0

20 40 60 Frame Index

(c) “walking”

LPIPS

(a) “cook spinach”

300

0.83

0

100 200 Frame Index

(d) “discussion”

300

0

100 200 Frame Index

(e) “stepin”

300

0

50 100 Frame Index

150

(f) “basketball”

Fig. 14. Comparison of visual quality under fluctuating bandwidth on volumetric videos prepared by HiCoM [Gao et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. Results for all videos (including volumetric videos prepared by Dynamic 3DGS [Luiten et al. 2024], 4DGS [Wu et al. 2024] and HiCoM [Gao et al. 2024]) under fluctuating bandwidth conditions are included in the supplementary material.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

•

13

Supplementary materials for “CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting” A

Pseudocode of PRPA Algorithm

Algorithm 1 Post-Render Perspective Aligning (PRPA) Input: reference image 𝑨(𝑖, 𝑗), locally rendered depth map 𝑫 (𝑢, 𝑣), camera intrinsic matrix 𝑲𝑙 , 𝑲𝑟 , rotation matrix 𝑹𝑙 , 𝑹𝑟 and position 𝒕𝑙 , 𝒕𝑟 of locally rendered image and reference image, kernel size 𝑘 for average filtering, depth threshold 𝑇 for occlusion detection. Output: aligned reference image 𝑨′ (𝑢, 𝑣) where the alignment error has been eroded. 1: 𝑷𝑐 (𝑢, 𝑣) ← 𝑲𝑙 (𝑹𝑟 𝑹𝑙−1 (𝑲𝑙−1 [𝑢, 𝑣, 1] ⊤ 𝑫 (𝑢, 𝑣) − 𝒕𝑙 ) + 𝒕𝑟 ) 2: 𝑫𝑟 (𝑢, 𝑣) ← 𝑷𝑐 (𝑢, 𝑣)3 3: 𝑷 (𝑢, 𝑣) ← 𝑷𝑐 (𝑢, 𝑣)1,2 /𝑷𝑐 (𝑢, 𝑣)3 4: 𝑨′ (𝑢, 𝑣) ← 𝑨(𝑷 (𝑢, 𝑣)1 , 𝑷 (𝑢, 𝑣)2 ) ⊲ Align with error 5: 𝑪𝑟 (𝑖, 𝑗) ← |{(𝑢, 𝑣)|𝑷 (𝑢, 𝑣) = (𝑖, 𝑗)}| 6: 𝑪𝑙 (𝑢, 𝑣) ← 𝑪𝑟 (𝑷 (𝑢, 𝑣)1 , 𝑷 (𝑢, 𝑣)2 ) 7: 𝑫𝑟𝑚 (𝑖, 𝑗) ← min({𝑫𝑟 (𝑢, 𝑣)|𝑷 (𝑢, 𝑣) = (𝑖, 𝑗)}) 8: 𝑫𝑙𝑚 (𝑢, 𝑣) ← 𝑫𝑟𝑚 (𝑷 (𝑢, 𝑣)1 , 𝑷 (𝑢, 𝑣)2 ) 9: 𝑷𝑜 ← {(𝑢, 𝑣)|𝑪𝑙 (𝑢, 𝑣) > 0 ∧ |𝑫 (𝑢, 𝑣) − 𝑫𝑙𝑚 (𝑢, 𝑣)| ≤ 𝑇 } 10: 𝑷𝑑 ← {(𝑢, 𝑣)|𝑪𝑙 (𝑢, 𝑣) > 0 ∧ |𝑫 (𝑢, 𝑣) − 𝑫𝑙𝑚 (𝑢, 𝑣)| > 𝑇 } 11: while |𝑷𝑑 | > 0 do 12: 𝑬𝑑 ←MorphologyDilate(𝑷𝑑 )∩∁𝑷𝑑 13: for (𝑢𝐸 , 𝑣 𝐸 ) ∈ 𝑬𝑑 ∩ ∁𝑷𝑜 do 14: 𝑲 ← [𝑢𝐸 − 𝑘, 𝑢𝐸 + 𝑘] × [𝑣 𝐸 − 𝑘, 𝑣 𝐸 + 𝑘] 15: 𝑎 ← mean({𝑨′ (𝑢, 𝑣)|(𝑢, 𝑣) ∈ 𝑲 ∩ ∁𝑷𝑜 ∩ ∁𝑷𝑑 }) 16: for (𝑢, 𝑣) ∈ 𝑲 ∩ 𝑷𝑑 do 17: 𝑨′ (𝑢, 𝑣) ← 𝑎 ⊲ Erode error along edge 18: delete (𝑢, 𝑣) from 𝑷𝑑 19: end for 20: end for 21: end while 22: return 𝑨′ (𝑢, 𝑣)

B

Detailed Prototype Implementation

We implemented a prototype CAGS system to evaluate its effectiveness. In this section, we describe engineering details and underlying rationales of each component to facilitate reproducibility.

B.1

Preparing Volumetric Video

Following TrackerSplat [Yin et al. 2025], we represent each dynamic scene as a set of Gaussians whose positions and attributes are updated across frames. For each frame, we store only the Gaussians whose attributes change relative to the previous frame. Concretely, the first frame is treated as a keyframe containing all Gaussians, and each subsequent frame stores only the updated Gaussians. To further reduce the data footprint without sacrificing fidelity, we incorporate importance-based pruning from Reduced3DGS [Papantonakis et al. 2024] and GSNB [Zhou et al. 2024] into the keyframe training loop, removing low-contribution Gaussians while preserving rendering quality.

B.2

SVQ

We quantize each Gaussian attribute independently using SVQ. We use Table 4 to determine suitable bit widths for cluster initialization. As shown in Table 4, SVQ achieves compression ratios comparable to recent scalable compression methods, while avoiding their heavy decoding cost. Although a 76-bit initialization provides better quality than 66-bit, the improvement is not significant. Given that 66-bit initialization already provides competitive performance, we adopt it as the default setting in our system. In detail, the 66 bits are allocated as follows: scaling (10 bits), rotation (quaternion real: 8 bits, imaginary: 10 bits), opacity (8 bits), coefficients for level-0 SH (9 bits), and the remaining three SH levels (8, 7, and 6 bits). Finally, the initial clusters of each attribute are merged into 16 clusters (4 bits) respectively (Sec. 4.1.1).

B.3

LoD Assignment

Although each attribute can be compressed independently by SVQ, adaptive streaming requires a linear LoD structure, where levels are transmitted sequentially (LoD 0, 1, 2, ...) until reaching bandwidth limits [Sun et al. 2025]. This linearity avoids combinatorial decisions and supports real-time bitrate control by simply accumulating layer sizes. Since lower LoDs are always delivered earlier, layers with larger visual impact must be prioritized to maximize quality under limited bandwidth. We determine the layer ordering by evaluating the importance of enhancement layers for each attribute. For each attribute, we progressively decode it with increasing numbers of enhancement layers while keeping other attributes at full quality. For each decoded scene, we render the result, apply the restoration model in Sec. 3.1, and measure PSNR/SSIM. Results in Figure 16 show that scaling has the greatest impact, followed by rotation, spherical harmonics, and opacity. We also observe that higher layers provide diminishing returns, and the ranking is roughly stable across datasets. Based on these measurements, we interleave attribute layers according to their PSNR gains in Figure 16a to construct a unified linear LoD hierarchy. This design ensures early LoDs deliver the highest quality gains, while later LoDs provide progressively smaller refinements. Figure 15 plots quality versus frame size across LoDs under this hierarchy. It shows a steep quality increase in the first few LoDs, then levels off, demonstrating that critical SVQ layers are positioned at lower LoDs, as required for effective adaptive streaming.

B.4

Tiling

Following LTS [Sun et al. 2025], we partition Gaussians into spatial tiles for viewport-adaptive streaming that transmits only visible tiles. As our training integrates importance-based pruning, the resulting Gaussian distribution is notably sparser and less uniform than that in LTS. As a result, directly applying the bounding-box tiling strategy in LTS would lead to many under-populated tiles, reducing compression efficiency.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

14

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

"cook spinach" "walking" "boxes"

"cut roasted beef" "discussion" "football"

"flame salmon" "stepin" "softball"

"flame steak" "trimming" "tennis"

"sear steak" "basketball" "juggle"

40 0.98 38 0.96 36 0.94

0.92 SSIM

PSNR (dB)

34

32

0.90 30 0.88 28 0.86 26

0.84 1.0

1.5

2.0

2.5 3.0 Frame size (MB)

3.5

4.0

4.5

1.0

1.5

2.0

2.5 3.0 Frame size (MB)

3.5

4.0

4.5

Fig. 15. Visual quality versus frame size across LoDs constructed by interleaving SVQ layers, evaluated on the first frame of each video. Each point corresponds to one LoD. The x-axis shows the resulting frame size (excluding the codebook) after removing a given LoD and all higher LoDs, and the y-axis shows the resulting visual quality (PSNR/SSIM).

To address this, we leverage the Morton-code sorting mechanism from TaoGS [Jiang et al. 2025] to linearize Gaussians while preserving spatial locality. For each frame, we sort Gaussians by Morton code and partition them into tiles with approximately equal size, so that each tile contains a similar number of Gaussians. The tile size is set to 8,000 Gaussians for the first frame and 4,000 for subsequent frames.

based on rendered reference images. 2) Select tiles that cover all visible Gaussians from recent frames. 3) For tiles not yet transmitted, progressively increase their LoDs until their average LoD matches that of already-sent tiles or the bandwidth limit is reached. 4) For tiles already transmitted, further increase their LoDs in descending order according to the count of visible Gaussians, until the bandwidth limit is reached.

B.5

B.7

Compression

After tiling, Gaussian positions and the lowest-LoD attributes are compressed using enhanced Draco [Google 2017], following LTS [Sun et al. 2025]. Higher LoDs and the SVQ codebook are compressed separately using Gzip. In addition, we store unquantized frames (also compressed with Draco) on the server to support reference rendering.

B.6

Network Adaptation Strategy

For evaluation, we implement a straightforward network adaptation strategy based on bandwidth limit to handle fluctuating bandwidth conditions: 1) On the streaming server, identify visible Gaussians

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Streaming

Both the streaming server and client run on machines equipped with RTX 3080 GPUs. Before streaming starts, we allocate a short initialization period to collect initial viewport traces and download necessary components, including the codebook (81.9 KB), the color restoration model (219 KB), and the lowest LoD of the first frame (∼1.5 MB). The streaming frame rate is 30 FPS, and the client renders at 60 FPS. The client uses a fixed rendering FoV of 1.48 rad (horizontal) and 1.20 rad (vertical) at 1600×1200, while reference images are rendered at 400×300. Both the 3D Tile Buffer and the Reference Buffer are set to 6 frames, i.e., the server predicts the viewport and renders reference images 180 ms (6 frames) ahead of the actual viewing time.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

•

15

34.6

PSNR

33.3 32.0 30.7 0.97 0.96

SSIM

0.95 0.94 0.93 0.92

SH.coef opacity rotation scale −600 −300 0 Frame size reduction (KB)

PSNR

LPIPS

(a) All Average 34.4

34.6

33.1

34.5

34.7

33.1

33.3

31.8

33.2

33.4

31.8

32.0

30.5

31.9

32.1

30.5

30.7

29.2

30.6

30.8

0.96

0.96

0.95

0.95

0.97

SSIM

0.96

0.94 0.93 0.92

SH.coef opacity rotation scale

0.94 0.93 0.92

PSNR

LPIPS

−600 −300 0 Frame size reduction (KB)

SSIM

0.92 0.91

SH.coef opacity rotation scale

0.94 0.93 0.92

0.95 SH.coef opacity rotation scale

0.94 0.93 0.92

(b) “cook spinach”

(c) “cut roast beef”

(d) “flame salmon”

34.0

32.9

33.4

33.4

32.7

31.6

32.1

32.1

31.4

30.3

30.8

30.8

32.3

30.1

29.0

29.5

29.5

30.9

0.96

0.96

0.95

0.95

0.94

0.94

0.96 0.95 0.94

0.94 0.93 0.92

SH.coef opacity rotation scale −600 −300 0 Frame size reduction (KB)

(g) “walking” LPIPS

0.93

−600 −300 0 Frame size reduction (KB)

0.95

PSNR

SH.coef opacity rotation scale

0.96

0.95

0.94

−600 −300 0 Frame size reduction (KB)

0.96

SSIM

0.96

0.95

0.93 0.92 0.91

SH.coef opacity rotation scale

0.93 0.92

−600 −300 0 Frame size reduction (KB)

SH.coef opacity rotation scale

−600 −300 0 Frame size reduction (KB)

0.92

−600 −300 0 Frame size reduction (KB)

(h) “discussion”

−600 −300 0 Frame size reduction (KB)

(e) “flame steak”

0.93

(f) “sear steak”

33.7

0.96 0.95 SH.coef opacity rotation scale

0.94 0.93 0.92

−600 −300 0 Frame size reduction (KB)

(i) “stepin”

(j) “trimming”

(k) “basketball”

34.8

34.2

35.3

35.2

33.1

33.5

32.9

34.0

33.9

31.8

32.2

31.6

32.7

32.6

30.5

30.9

30.3

31.4

31.3

0.97

0.97

0.96

0.96

0.96

0.96

0.95

0.95

0.94 0.93 0.92

SH.coef opacity rotation scale −600 −300 0 Frame size reduction (KB)

0.93 0.92

0.95

0.94 SH.coef opacity rotation scale

0.93 0.92

SH.coef opacity rotation scale

0.94 0.93 0.92

0.95 SH.coef opacity rotation scale

0.94 0.93 0.92

SH.coef opacity rotation scale

−600 −300 0 Frame size reduction (KB)

−600 −300 0 Frame size reduction (KB)

−600 −300 0 Frame size reduction (KB)

−600 −300 0 Frame size reduction (KB)

(m) “football”

(n) “juggle”

(o) “softball”

(p) “tennis”

LPIPS

(l) “boxes”

0.94

0.95

SH.coef opacity rotation scale −600 −300 0 Frame size reduction (KB)

34.4

0.96

SH.coef opacity rotation scale

Fig. 16. Visual quality versus data size for individual SVQ enhancement layers. Each point corresponds to one SVQ layer. The x-axis shows data size reduction (excluding the codebook) after removing the given layer and all subsequent layers, while the y-axis reports the resulting rendering quality (PSNR/SSIM) for that video. SH.coef: coefficient for spherical harmonics. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

16

C

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

Additional Evaluation Results

Table 3. Performance and quality comparison of our SVQ method (with 66-bit and 76-bit initialization) against state-of-the-art scalable compression methods at the highest quality level. Best and second-best results are highlighted in bold and underline, respectively. Size is in MB and decoding time ("dec.t") in seconds.

Datasets Neural 3D Video dataset Meeting Room dataset Dynamic 3DGS dataset Methods psnr ssim lpips size dec.t psnr ssim lpips size dec.t psnr ssim lpips size dec.t SPZ low 23.3 .842 .112 3.33 0.076 21.8 .822 .070 2.54 0.062 21.1 .761 .130 2.26 0.052 SPZ high 23.1 .842 .110 4.53 0.075 21.8 .822 .068 3.41 0.057 21.0 .760 .112 3.32 0.057 CompGS 24.9 .898 .193 22.20 6.213 25.0 .880 .221 5.58 0.799 19.1 .721 .353 4.27 2.772 HAC 22.2 .889 .203 25.69 24.407 25.2 .873 .216 8.69 9.135 20.5 .758 .338 2.84 1.571 HAC++ 21.5 .885 .202 17.98 51.708 25.2 .868 .222 6.78 14.706 20.9 .780 .291 2.00 6.169 SVQ 66bit 24.5 .901 .108 2.00 0.015 26.0 .887 .093 1.53 0.018 25.3 .883 .153 2.25 0.028 SVQ 76bit 25.1 .907 .100 2.14 0.015 26.4 .894 .079 1.69 0.019 23.3 .874 .200 2.42 0.029 Table 4. Performance and quality comparison of our SVQ method against recent compression methods. Best and second-best results are shown in bold and underline, respectively. The "scal." column indicates whether a method is scalable. Size is in MB and decoding time ("dec.t") in seconds. Our method achieves comparable rate-distortion performance with significantly lower decoding latency, which is uniquely suited for real-time streaming.

Datasets Neural 3D Video dataset Meeting Room dataset Dynamic 3DGS dataset Methods scal. psnr ssim lpips size dec.t psnr ssim lpips size dec.t psnr ssim lpips size dec.t GIFStream × 23.6 .860 .130 0.58 0.050 24.3 .879 .093 0.48 0.042 13.6 .615 .375 0.13 0.015 QUEEN × 29.5 .933 .181 17.36 0.997 31.4 .954 .153 7.70 0.437 18.3 .788 .315 3.41 0.197 SVQ ✓ 23.2 .829 .108 0.75 0.015 21.8 .812 .093 0.96 0.018 21.6 .766 .153 1.08 0.028 Table 5. Viewport prediction and Adaptive FoV precision of the across datasets, measured by Position MAE (translation error), Angular MAE (degrees), and FoV Coverage (%).

Dataset Neural 3D Video dataset ST-NeRF dataset Meeting Room dataset Dynamic 3DGS dataset

Position MAE 0.1115 0.0804 0.1841 0.0677

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Angular MAE 1.54 1.58 1.90 2.38

FoV Coverage (%) 94.72 94.66 94.20 93.77

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

PSNR

Ours

LTS-F

27.7

34.7

31.1

26.3

22.1

29.1

25.5

19.4

20.7

16.5

23.5

19.9

13.8

15.1

10.9

17.9

14.3

0.90

0.87

0.89

0.72

0.75

0.72

0.75

0.75

0.58

0.60

0.57

0.61

0.60

0.44

0.45

0.42

0.47

0.45

0.47

0.46

0.47

0.43

0.44

0.36

0.35

0.36

0.32

0.33

0.25

0.24

0.25

0.21

0.22

0.14

0.13

0.14

0.10

0.11

SSIM PSNR SSIM LPIPS

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

(d) “flame steak”

(e) “sear steak”

28.8

29.3

25.6

31.5

18.0

23.2

23.7

20.0

25.9

12.4

17.6

18.1

14.4

20.3

6.8

12.0

12.5

8.8

14.7

0.86

0.90

0.85

0.84

0.87

0.71

0.75

0.71

0.70

0.73

0.56

0.60

0.57

0.56

0.59

0.41

0.45

0.43

0.42

0.45

0.50

0.48

0.49

0.50

0.49

0.39

0.37

0.38

0.39

0.38

0.28

0.26

0.27

0.28

0.27

0.17

0.15

0.16

0.17 50 100 150 Bandwidth (Mbps)

(g) “discussion”

(h) “stepin”

0.16

50 100 150 Bandwidth (Mbps)

(i) “trimming”

50 100 150 Bandwidth (Mbps)

(j) “basketball”

29.8

31.6

34.2

31.5

31.5

24.2

26.0

28.6

25.9

25.9

18.6

20.4

23.0

20.3

20.3

13.0

14.8

17.4

14.7

14.7

0.92

0.93 0.78

0.75

0.91

0.88

0.89

0.76

0.74

0.77

0.61

0.60

0.62

0.63

0.61

0.46

0.46

0.47

0.48

0.47

0.46

0.47

0.47

0.46

0.45

0.35

0.36

0.36

0.35

0.34

0.24

0.25

0.25

0.24

0.23

0.13

0.14

0.14

0.13

50 100 150 Bandwidth (Mbps)

(k) “boxes”

50 100 150 Bandwidth (Mbps)

(l) “football”

50 100 150 Bandwidth (Mbps)

(m) “juggle”

SR align.ref.

50 100 150 Bandwidth (Mbps)

23.6

50 100 150 Bandwidth (Mbps)

17

0.90

50 100 150 Bandwidth (Mbps)

(f) “walking”

PSNR

CR w/o ref.

31.9

50 100 150 Bandwidth (Mbps)

SSIM

w/o Adaptive FoV

25.0

50 100 150 Bandwidth (Mbps)

LPIPS

w/o PRPA

30.6

0.86

LPIPS

V3-A

•

0.12 50 100 150 Bandwidth (Mbps)

(n) “softball”

50 100 150 Bandwidth (Mbps)

(o) “tennis”

Fig. 17. Comparison of visual quality under fixed bandwidth. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

18

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

LPIPS

SSIM

PSNR

Ours

V3-A

LTS-F

w/o PRPA

w/o Adaptive FoV

33.1

31.8

28.1

34.7

33.1

26.8

25.5

21.8

28.4

26.8

20.5

19.2

15.5

22.1

20.5

14.2

12.9

9.2

15.8

14.2

0.84

0.88

0.88

0.77

0.76

0.71

0.75

0.75

0.63

0.62

0.58

0.62

0.62

0.49

0.48

0.45

0.49

0.49

0.42

0.43

0.44

0.41

0.41

0.32

0.33

0.34

0.31

0.31

0.22

0.23

0.24

0.21

0.21

0.12

0.13

0.14

0.11

0.11

0

100 200 Frame Index

300

0

100 200 Frame Index

300

0

100 200 Frame Index

300

LPIPS

SSIM

PSNR

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

PSNR

100 200 Frame Index

300

(d) “flame steak”

0

30.2

30.7

28.7

32.6

23.9

24.4

22.4

26.3

15.6

17.6

18.1

16.1

20.0

9.3

11.3

11.8

9.8

13.7

0.88

0.88

0.87

0.86

0.88

0.75

0.75

0.74

0.73

0.75

0.62

0.62

0.61

0.60

0.62

0.49

0.49

0.48

0.47

0.49

0.42

0.44

0.43

0.46

0.44

0.32

0.34

0.33

0.36

0.34

0.22

0.24

0.23

0.26

0.24

0.12

0.14

0.13

0.16

20 40 60 Frame Index

0

100 200 Frame Index

300

0

(g) “discussion”

100 200 Frame Index

300

(h) “stepin”

100 200 Frame Index

300

(i) “trimming”

0

31.6

33.0

33.9

32.2

25.3

26.7

27.6

25.9

18.1

19.0

20.4

21.3

19.6

11.8

12.7

14.1

15.0

13.3

0.87

0.86

0.87

0.74

0.73

0.74

0.61

0.60

0.61

0.64

0.61

0.48

0.47

0.48

0.50

0.48

0.43

0.46

0.44

0.43

0.43

0.33

0.36

0.34

0.33

0.33

0.23

0.26

0.24

0.23

0.23

0.13

0.16

0.14

0.13

0.13

(k) “boxes”

150

0

50 100 Frame Index

150

(l) “football”

300

50 100 Frame Index

150

(j) “basketball”

24.4

50 100 Frame Index

100 200 Frame Index

0.14

0

30.7

0

SR align.ref.

(e) “sear steak”

21.9

(f) “walking”

SSIM

0

28.2

0

LPIPS

CR w/o ref.

0.87 0.78

0

50 100 Frame Index

(m) “juggle”

150

0.74

0

50 100 Frame Index

(n) “softball”

150

0

50 100 Frame Index

150

(o) “tennis”

Fig. 18. Comparison of visual quality under fluctuating bandwidth. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

SSIM

PSNR

Ours

w/o PRPA

w/o Adaptive FoV

CR w/o ref.

•

19

SR align.ref.

32.8

34.5

27.3

36.6

30.3

29.0

30.7

23.5

32.8

26.5

25.2

26.9

19.7

29.0

22.7

21.4

23.1

15.9

25.2

18.9

0.92

0.93

0.88

0.94

0.92

0.88

0.89

0.84

0.90

0.88

0.84

0.85

0.80

0.86

0.84

0.80

0.81

0.76

0.82

0.80

0.24

0.18

0.20

0.18

0.12

0.14

0.12

0.06

0.08

0.22 LPIPS

0.21 0.17 0.15

0.12

0.09

0.07

LPIPS

SSIM

PSNR

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

(d) “flame steak”

(e) “sear steak”

26.5

29.8

30.2

30.3

32.8

22.7

26.0

26.4

26.5

29.0

18.9

22.2

22.6

22.7

25.2

15.1

18.4

18.8

18.9

21.4

0.90

0.91

0.90

0.91

0.94

0.86

0.87

0.86

0.87

0.90

0.82

0.83

0.82

0.83

0.86

0.78

0.79

0.78

0.79

0.82

0.26

0.26 0.21

0.23

0.21

0.21

0.21

0.16

0.16

0.15

0.17

0.15

0.11

(f) “walking”

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(g) “discussion”

0.09

0.11

0.09

0.11 50 100 150 Bandwidth (Mbps)

PSNR

50 100 150 Bandwidth (Mbps)

(h) “stepin”

50 100 150 Bandwidth (Mbps)

(i) “trimming”

50 100 150 Bandwidth (Mbps)

(j) “basketball”

29.7

31.1

35.8

32.4

31.5

25.9

27.3

32.0

28.6

27.7

22.1

23.5

28.2

24.8

23.9

24.4

21.0

20.1

18.3

19.7

0.93

0.92

0.89

0.88

0.85

0.84

0.81

0.80

SSIM

0.96 0.93

0.93 0.91

0.85

0.86

0.85

0.81

0.81

0.81

0.24

0.23

LPIPS

0.22

0.18 0.14

0.14

0.10

0.09 50 100 150 Bandwidth (Mbps)

(k) “boxes”

0.22

0.20

0.19 0.16

0.89

0.89

0.08 50 100 150 Bandwidth (Mbps)

(l) “football”

0.16

0.13

0.10

0.08 50 100 150 Bandwidth (Mbps)

(m) “juggle”

50 100 150 Bandwidth (Mbps)

(n) “softball”

50 100 150 Bandwidth (Mbps)

(o) “tennis”

Fig. 19. Comparison of visual quality under fixed bandwidth on volumetric videos prepared by Dynamic 3DGS [Luiten et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

20

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

PSNR

Ours

w/o PRPA 33.5

27.1

35.9

33.1

27.2

20.8

29.6

26.8

22.1

20.9

14.5

23.3

20.5

15.8

14.6

8.2

17.0

14.2

0.93

0.94

0.87 0.79

0.75

0.75

0.71

0.76

0.75

0.67

0.66

0.63

0.67

0.67

0.37

0.38

0.41

0.36

0.36

0.28

0.29

0.32

0.27

0.27

0.19

0.20

0.23

0.18

0.18

0.10

0.11

0.14

0.09

0.09

100 200 Frame Index

300

0

100 200 Frame Index

300

0

100 200 Frame Index

300

PSNR

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

SSIM LPIPS

0

100 200 Frame Index

300

(d) “flame steak”

0

30.5

29.9

32.7

36.0

24.2

23.6

26.4

29.7

14.7

17.9

17.3

20.1

23.4

8.4

11.6

11.0

13.8

17.1

0.89

0.90

0.92

0.81

0.82

0.84

0.92 0.83

0.73

0.74

0.73

0.74

0.76

0.65

0.65

0.65

0.66

0.68

0.40

0.40

0.38

0.39

0.38

0.31

0.31

0.29

0.30

0.29

0.22

0.22

0.20

0.21

0.20

0.13

0.13

0.11

0.12

0

20 40 60 Frame Index

0

(f) “walking”

100 200 Frame Index

300

0

(g) “discussion”

100 200 Frame Index

300

(h) “stepin”

100 200 Frame Index

300

0

31.3

35.6

35.6

34.8

24.3

25.0

29.3

29.3

28.5

18.0

18.7

23.0

23.0

22.2

11.7

12.4

16.7

16.7

15.9

0.91

0.90

0.85

0.82

0.75

0.74

0.76

0.76

0.76

0.67

0.66

0.67

0.68

0.67

0.39

0.38

0.39

0.38

0.40

0.29

0.30

0.29

0.31

0.20

0.21

0.20

0.22

0.11

0.12

0.11

0.13

0.21 0.12

0

50 100 Frame Index

(k) “boxes”

150

0

50 100 Frame Index

150

(l) “football”

0.84

0

50 100 Frame Index

(m) “juggle”

150

150

0.94

0.92

0.83

0.30

50 100 Frame Index

(j) “basketball”

30.6

0.85

300

0.11

0

(i) “trimming”

0.94

100 200 Frame Index

(e) “sear steak”

21.0

0.81

PSNR

0.83

27.3

0.89

SSIM

0.91

0.85

0.84

0

LPIPS

SR align.ref.

28.4

0.83 SSIM

CR w/o ref.

34.7

0.91

LPIPS

w/o Adaptive FoV

0

50 100 Frame Index

(n) “softball”

150

0

50 100 Frame Index

150

(o) “tennis”

Fig. 20. Comparison of visual quality under fluctuating bandwidth on volumetric videos prepared by Dynamic 3DGS [Luiten et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

SSIM

PSNR

Ours

w/o PRPA

w/o Adaptive FoV

CR w/o ref.

33.3

24.7

34.5

27.4

25.5

29.5

20.9

30.7

23.6

21.7

25.7

17.1

26.9

19.8

17.9

21.9

13.3

23.1

16.0

0.91

0.93

0.86

0.93

0.90

0.87

0.89

0.82

0.89

0.86

0.83

0.85

0.78

0.85

0.82

0.79

0.81

0.74

0.81

0.78

0.28

0.19

0.22

0.13

0.16

0.07

0.25

0.23 LPIPS

0.24

0.13 0.08

PSNR

50 100 150 Bandwidth (Mbps)

SSIM

0.20

0.18

0.12

21

SR align.ref.

29.3

0.18

•

50 100 150 Bandwidth (Mbps)

0.15 0.10

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

(d) “flame steak”

25.4

29.7

28.7

25.7

31.4

21.6

25.9

24.9

21.9

27.6

17.8

22.1

21.1

18.1

23.8

14.0

18.3

17.3

14.3

20.0

0.89

0.95

0.89

0.92

0.90 0.84

0.90

0.79

0.85

0.74

0.80

(e) “sear steak”

0.85

0.88

0.86

0.81

0.84

0.82

0.77

0.80

0.78

0.23

0.22

0.28

0.22

0.17

0.16

0.22

0.16

0.11

0.10

0.16

LPIPS

0.28 0.23 0.18 0.13 50 100 150 Bandwidth (Mbps)

PSNR

(f) “walking”

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(g) “discussion”

(h) “stepin”

0.10 50 100 150 Bandwidth (Mbps)

(i) “trimming”

50 100 150 Bandwidth (Mbps)

(j) “basketball”

29.1

31.1

34.6

31.3

30.5

25.3

27.3

30.8

27.5

26.7

21.5

23.5

27.0

23.7

22.9

17.7

19.7

23.2

19.9

19.1

0.92

0.92

0.93

0.93

0.88

0.88

0.89

0.89

0.84

0.84

0.85

0.85

0.85

0.80

0.80

0.80

0.81

0.81

SSIM

0.95 0.90

0.25

0.25 0.23

LPIPS

0.22

0.23 0.20

0.20 0.16

0.17

0.15

0.10

0.10 50 100 150 Bandwidth (Mbps)

(k) “boxes”

50 100 150 Bandwidth (Mbps)

(l) “football”

0.17

0.15

0.11

0.11

0.10 50 100 150 Bandwidth (Mbps)

(m) “juggle”

50 100 150 Bandwidth (Mbps)

(n) “softball”

50 100 150 Bandwidth (Mbps)

(o) “tennis”

Fig. 21. Comparison of visual quality under fixed bandwidth on volumetric videos prepared by HiCoM [Gao et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

22

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

PSNR

Ours

w/o PRPA 32.3

26.1

33.5

31.2

26.0

19.8

27.2

24.9

19.2

19.7

13.5

20.9

18.6

12.9

13.4

7.2

14.6

12.3

0.88

0.93

0.82

0.79

0.84

0.75

0.74

0.70

0.75

0.74

0.66

0.66

0.61

0.66

0.66

0.39

0.38

0.44

0.37

0.37

0.30

0.29

0.35

0.28

0.28

0.21

0.20

0.26

0.19

0.19

0.12

0.11

0.17

0.10

0.10

0.90

SSIM LPIPS LPIPS

SSIM

PSNR

0

100 200 Frame Index

300

0

100 200 Frame Index

300

0

100 200 Frame Index

300

0.90 0.82

0

100 200 Frame Index

300

0

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

(d) “flame steak”

27.2

30.5

29.3

27.7

34.3

20.9

24.2

23.0

21.4

28.0

14.6

17.9

16.7

15.1

21.7

8.3

11.6

10.4

8.8

15.4

0.89

0.90

0.91

0.90

0.94

0.81

0.82

0.82

0.81

0.85

0.73

0.74

0.73

0.72

0.76

0.65

0.66

0.64

0.63

0.67

0.42

0.40

0.39

0.41

0.39

0.33

0.31

0.30

0.32

0.30

0.24

0.22

0.21

0.23

0.21

0.15

0.13

0.12

0.14

0

20 40 60 Frame Index

0

(f) “walking”

PSNR

SR align.ref.

25.5

0.84

SSIM

CR w/o ref.

31.8

0.93

LPIPS

w/o Adaptive FoV

100 200 Frame Index

300

0

(g) “discussion”

100 200 Frame Index

300

(h) “stepin”

(e) “sear steak”

100 200 Frame Index

300

0

31.1

34.1

35.0

30.9

24.5

24.8

27.8

28.7

24.6

18.2

18.5

21.5

22.4

18.3

11.9

12.2

15.2

16.1

12.0

0.91

0.90

0.83

0.82

0.84

0.84

0.84

0.75

0.74

0.75

0.76

0.75

0.67

0.66

0.66

0.68

0.66

0.39

0.40

0.40

0.39

0.39

0.30

0.31

0.31

0.30

0.30

0.21

0.22

0.22

0.21

0.21

0.12

0.13

0.13

0.12

50 100 Frame Index

(k) “boxes”

150

0

50 100 Frame Index

150

(l) “football”

50 100 Frame Index

(m) “juggle”

150

150

0.93

0.92

0

50 100 Frame Index

(j) “basketball”

30.8

0

300

0.12

0

(i) “trimming”

0.93

100 200 Frame Index

0.12 0

50 100 Frame Index

(n) “softball”

150

0

50 100 Frame Index

150

(o) “tennis”

Fig. 22. Comparison of visual quality under fluctuating bandwidth on volumetric videos prepared by HiCoM [Gao et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

PSNR

Ours

w/o PRPA

w/o Adaptive FoV

CR w/o ref.

SSIM

23

SR align.ref.

33.2

34.3

28.3

36.0

30.4

29.4

30.5

24.5

32.2

26.6

25.6

26.7

20.7

28.4

22.8

21.8

22.9

16.9

24.6

19.0

0.92

0.93

0.91

•

0.94 0.94

0.86

0.88

0.89

0.84

0.85

0.81

0.80

0.81

0.76

0.20

0.20

0.23

0.14

0.14

0.17

0.08

0.08

0.11

0.89

0.90 0.86

0.84

0.82

0.79

LPIPS

0.21

PSNR

50 100 150 Bandwidth (Mbps)

SSIM

0.20 0.16

50 100 150 Bandwidth (Mbps)

0.14

0.11

0.08

0.06 50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

(d) “flame steak”

50 100 150 Bandwidth (Mbps)

(e) “sear steak”

26.6

29.4

30.0

29.5

33.1

22.8

25.6

26.2

25.7

29.3

19.0

21.8

22.4

21.9

25.5

15.2

18.0

18.6

18.1

21.7

0.90

0.90

0.90

0.90

0.93

0.86

0.86

0.86

0.86

0.89

0.82

0.82

0.82

0.82

0.85

0.78

0.78

0.78

0.78

0.81

0.26

0.27

0.24

0.26

0.21

0.22

0.19

0.21

0.16

0.17

0.14

0.16

0.11

0.12

0.09

0.11

LPIPS

0.23

50 100 150 Bandwidth (Mbps)

SSIM

PSNR

(f) “walking”

50 100 150 Bandwidth (Mbps)

50 100 150 Bandwidth (Mbps)

(g) “discussion”

(h) “stepin”

0.17 0.11

50 100 150 Bandwidth (Mbps)

(i) “trimming”

(j) “basketball”

29.9

31.2

35.5

33.0

30.4

26.1

27.4

31.7

29.2

26.6

22.3

23.6

27.9

25.4

22.8

18.5

19.8

24.1

21.6

19.0

0.93

0.92

0.93

0.94

0.93

0.89

0.88

0.89

0.90

0.89

0.85

0.84

0.85

0.86

0.85

0.81

0.80

0.81

0.82

0.81

0.26

0.26

0.21

0.21

0.26 0.22

LPIPS

50 100 150 Bandwidth (Mbps)

0.16

0.16

0.11

(k) “boxes”

50 100 150 Bandwidth (Mbps)

(l) “football”

0.18

0.16

0.10

0.11 50 100 150 Bandwidth (Mbps)

0.24 0.21

0.16

0.12

0.11 50 100 150 Bandwidth (Mbps)

(m) “juggle”

50 100 150 Bandwidth (Mbps)

(n) “softball”

50 100 150 Bandwidth (Mbps)

(o) “tennis”

Fig. 23. Comparison of visual quality under fixed bandwidth on volumetric videos prepared by 4DGS [Wu et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

24

•

Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, and Fang Dong

PSNR

Ours

w/o PRPA 33.0

28.3

35.1

33.2

26.7

22.0

28.8

26.9

21.2

20.4

15.7

22.5

20.6

14.9

14.1

9.4

16.2

14.3

0.93

0.90

0.91

0.91

0.84

0.81

0.83

0.83

0.75

0.75

0.72

0.75

0.75

0.67

0.66

0.63

0.67

0.67

0.37

0.38

0.40

0.36

0.37

0.28

0.29

0.31

0.27

0.28

0.19

0.20

0.22

0.18

0.19

0.10

0.11

0.13

0.09

0.10

0

50 Frame Index

100

0

100 200 Frame Index

0

100 200 Frame Index

300

LPIPS

SSIM

PSNR

(a) “cook spinach” (b) “cut roast beef” (c) “flame salmon”

PSNR

100 200 Frame Index

300

(d) “flame steak”

0

29.6

30.5

30.8

34.2

23.3

24.2

24.5

27.9

14.7

17.0

17.9

18.2

21.6

8.4

10.7

11.6

11.9

15.3

0.92

0.94

0.92

0.89

0.91

0.83

0.85

0.83

0.81

0.83

0.74

0.76

0.74

0.73

0.75

0.65

0.67

0.65

0.65

0.67

0.40

0.40

0.37

0.39

0.39

0.31

0.31

0.28

0.30

0.30

0.22

0.22

0.19

0.21

0.21

0.13

0.13

0.10

0.12

20 40 60 Frame Index

0

100 200 Frame Index

300

0

(g) “discussion”

100 200 Frame Index

300

(h) “stepin”

100 200 Frame Index

300

(i) “trimming”

0

31.5

35.6

36.6

33.4

25.2

29.3

30.3

27.1

18.0

18.9

23.0

24.0

20.8

11.7

12.6

16.7

17.7

14.5

0.93

0.94

0.84

0.85

0.75

0.75

0.76

0.76

0.76

0.67

0.66

0.67

0.68

0.67

0.40

0.40

0.39

0.39

0.41

0.31

0.31

0.30

0.30

0.32

0.22

0.22

0.21

0.21

0.23

0.13

0.13

0.12

0.12

0.14

0

50 100 Frame Index

(k) “boxes”

150

0

50 100 Frame Index

150

(l) “football”

0.85

0.84

50 100 Frame Index

(m) “juggle”

150

150

0.94

0.92

0

50 100 Frame Index

(j) “basketball”

24.3

0.83

300

0.12

0

30.6

0.91

100 200 Frame Index

(e) “sear steak”

21.0

(f) “walking”

SSIM

0

27.3

0

LPIPS

SR align.ref.

27.5

0.83 SSIM

CR w/o ref.

33.8

0.91

LPIPS

w/o Adaptive FoV

0

50 100 Frame Index

(n) “softball”

150

0

50 100 Frame Index

150

(o) “tennis”

Fig. 24. Comparison of visual quality under fluctuating bandwidth on volumetric videos prepared by 4DGS [Wu et al. 2024]. The y-axis ranges vary across subplots while maintaining equal scale spans for each metric across videos. SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting

(a) Color-distorted

(b) CR w/o ref.

(c) SR align ref.

(d) w/o PRPA

(e) w/o Adaptive FoV

(f) Ours

•

25

(g) Ground Truth

Fig. 25. Visual results of the frame 30 under 30Mbps.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Record · ID 175176 · SHA-256 deeb6251f6b55468
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.