ConceptioArchivearXiv CS
arXiv CSopen access

EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
databasesdatamanagementsqlstorage
databases, sql, data management, storage

EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database Kaiyuan Tang EVOLVE TTHRESH SZ3 AMGSRN++ fV-SRN

55

PSNR (dB)

50

, Maizhe Yang

, and Chaoli Wang TTHRESH CR: 6,334 X

Instant-NGP SIREN NeurComp ECNR ZFP

EVOLVE CR: 7,843 X

45

40

35 0

2,000

4,000

6,000

8,000

Compression Ratio

10,000

(a) rate-distortion curves

12,000

(b) encoding/decoding time

(c) GT

(d) TTHRESH vs. EVOLVE

arXiv:2607.18187v1 [cs.GR] 20 Jul 2026

Figure 1: After training, EVOLVE, as an offline compression method, can achieve a competitive compression ratio (CR) on the unseen ionization (H+) dataset (141 MB). Unlike prior learning-based methods that require storing one model per CR setting, a single trained EVOLVE model supports a wide range of variable-rate encoding, as shown in (a). The encoding/decoding time plot in (b) uses a log scale to improve the visibility of extremely small values. A more comprehensive comparison of encoding/decoding time is given in Tables 2 and 3. All evaluations are conducted on an NVIDIA 4090 GPU. Abstract—Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical. However, conventional compressors often struggle to preserve fine structural details at high compression ratios (CRs), and implicit neural representations (INRs) require costly per-volume optimization and produce models with fixed CRs. To respond, we present EVOLVE, an autoencoder (AE)-based volume-compression framework that targets high CRs for offline compression, with three key contributions. First, we construct a large-scale cross-domain database of 6,376 volumes from 21 scientific simulations, curated via perceptual hashing to ensure diversity, enabling the optimized model to extract features that generalize across volumes within the covered scientific simulation domains. Second, we reexamine the design space of AE-based compressors and incorporate several macro- and micro-designs into a vanilla AE to develop EVOLVE, which substantially improves the expressive power and compression capability. Third, we develop a learnable gain mechanism with a three-stage training strategy to enable variable-rate encoding, allowing a single model to support continuous CR adjustment at inference time. Experiments on multiple unseen scientific simulation datasets demonstrate that EVOLVE achieves substantially higher CRs than conventional compressors at comparable reconstruction quality, while delivering compression speeds that are orders of magnitude faster than INR-based methods, highlighting its promise as a strong alternative for compressing scientific data. The code, model weights, and results are available on our project page at https://evolve-vis.github.io. Index Terms—Volume compression, learning-based compressor, autoencoder, context model, database

1

I NTRODUCTION

Modern scientific simulations and instruments produce volumetric data at rates that far outpace the growth of storage and network capacity. Large-scale astrophysics, climate science, and turbulent combustion simulations routinely generate extreme-scale data, while the cost of storing and moving such data continues to grow. For this reason, data archival and transmission have become major bottlenecks in scientific workflows, limiting both the scale of simulations that can be preserved and the speed at which results can be shared and analyzed. As a viable solution, lossy compression addresses this challenge by significantly reducing data size while preserving essential features for visualization and analysis. Conventional error-bounded lossy compressors, such as ZFP [38], TTHRESH [6], and SZ3 [37], rely on predefined mathematical transforms or predictors. While effective at moderate compression ratios (CRs), say a few hundred, these methods suffer severe distortion at high CRs (over 1,000). Besides these general-purpose compressors, non-neural learning-based solutions, such as hierarchical vector quantization [56] and sparse dictionary learning [15,17,43], learn • The authors are with the Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN 46556, USA. • E-mail: {ktang2, myang9, chaoli.wang}@nd.edu. Manuscript received xx xxx. 202x; accepted xx xxx. 202x. Date of Publication xx xxx. 202x; date of current version xx xxx. 202x. For information on obtaining reprints of this article, please send e-mail to: [email protected]. Digital Object Identifier: xx.xxxx/TVCG.202x.xxxxxxx

codebooks or dictionaries directly from the data, which have mostly been integrated into real-time compression-domain rendering pipelines. Implicit neural representations (INRs) [59] treat compression as function approximation and support random access. They yield high CRs while preserving good data quality. However, INR-based approaches have significant drawbacks: fully connected INRs [22, 42, 59] may require several hours of training time per volume; grid-based INRs [77,79,81] accelerate training but achieve lower CRs; and a single trained model corresponds to only one specific CR. Autoencoder (AE) methods, such as AE-SZ [39] and IDLat [57], map volume blocks to compact latent representations, yet they exhibit notable limitations as compression tools. First, their CRs do not demonstrate significant advantages over state-of-the-art conventional compressors [6, 37]. Second, these methods are typically trained on small-scale datasets (e.g., a few volume blocks sampled from a single time-varying volumetric dataset). While they can generalize to other timesteps within the same dataset, cross-dataset generalization remains an open challenge. Third, as with INR methods, the CRs of AEs are determined by fixed latent-space dimensions, thereby requiring different model architectures for different target rates. Due to these limitations, prior research has predominantly leveraged AEs as visual analysis tools [20, 57] rather than as competitive compression tools. To address these challenges, we rethink the application of AEs to volumetric data compression and propose an AE-based framework, named EVOLVE (Efficient Learned VOLume Compression with Variable-

Rate Encoding). Unlike methods that incorporate compression-domain rendering, EVOLVE targets high CRs for offline purposes, where the full volume is reconstructed before visualization and analysis. Under this setting, EVOLVE aims to encode and decode various volumes with a single shared model (see Section 4.5). We introduce three key innovations to overcome the existing limitations. First, to achieve higher CRs than conventional compressors at equivalent reconstruction quality, we incorporate advanced context-aware entropy modeling that combines spatial-wise [12, 26, 44] and channel-wise context [45], fusing multi-source information to estimate latent-variable distributions accurately. Second, to enable cross-dataset generalization, we construct a large-scale training database comprising thousands of volumes spanning diverse scientific simulation domains, allowing the model to learn transferable feature representations and, after being optimized on sufficiently representative volumes from a given domain, it can support compress unseen data without per-volume optimization. Third, to support flexible CRs with a single model, we adopt a learnable gain mechanism that enables continuous rate variation, coupled with a three-stage training strategy to ensure consistent performance across the entire rate-distortion spectrum. We evaluate EVOLVE on multiple scientific simulation datasets against conventional compressors, INR methods, and prior AE approaches. At equivalent reconstruction quality, EVOLVE attains substantially higher CRs than both conventional and AE-based methods (see Figure 1), while compressing orders of magnitude faster than INR methods on encoding speed. Moreover, a single EVOLVE model spans a continuous rate-distortion range, eliminating the need to train a separate model for each target CR, and generalizes well to volumes unseen during training. That said, EVOLVE is validated primarily on curated scientific simulation data; extending it to other domains, such as medical or microscopy volumes, would require retraining on representative data. In summary, this paper makes the following contributions. • We propose EVOLVE, an AE-based volumetric data compression framework that achieves high compression efficiency through advanced context-aware entropy modeling and significantly outperforms existing methods at equivalent reconstruction quality, while supporting continuous single-model variable-rate compression without requiring multiple models per setting. • We construct the first large-scale volume database spanning diverse scientific domains, enabling a single trained compression model to generalize to unseen volumes from the covered scientific simulation domains without per-volume optimization. • Extensive experiments show that EVOLVE achieves state-ofthe-art compression performance across multiple datasets while demonstrating strong generalization and practical efficiency for unseen simulation data. 2 R ELATED W ORK 2.1

Lossy Volumetric Data Compression

Lossy compression has been critical for managing the massive storage requirements and addressing I/O challenges posed by large-scale scientific simulations. Many schemes embed compressed representations into the rendering pipeline, decoding only the portion needed per frame on demand. Treib et al. [70] developed a GPU-decodable wavelet codec for terascale turbulence, Nystad et al. [47] introduced the fixed-rate block-local texture format ASTC with constant-time random access, and Schneider and Westermann [56] proposed a hierarchical vectorquantization scheme. Gobbetti et al. [17] developed COVRA, which models octree data blocks at multiple resolutions as sparse combinations of learned dictionary atoms, and Marton et al. [43] and Díaz et al. [15] later scaled this direction to massive time-varying data with out-of-core GPU streaming of variable-rate sparse codes. Among them, works including [15, 17, 43, 56] are considered learning-based solutions that optimize dictionaries and sparse codes from the data. Such methods prioritize random access and decoding throughput at render time. For more details about compression-domain volume rendering, we refer readers to the surveys [8, 54]. A complementary line of work targets offline storage and transmission, decompressing the full volume before

any visualization or analysis. Lindstrom [38] developed ZFP, which applies customized block-level transforms with fast I/O access; Liang et al. [37] developed SZ3, an error-bounded framework built on an improved Lorenzo predictor; and Ballester-Ripoll et al. [6] proposed TTHRESH, using tensor decomposition via higher-order SVD. Yan et al. [82] further proposed TopoSZ to preserve topological features under user-specified error bounds. Like these methods, EVOLVE is designed for offline compression, reconstructing the full volume before use rather than for transient render-time decoding (for more discussion, refer to Section 4.5). 2.2 Deep Learning for Volume Visualization Deep learning solutions [2, 20, 23, 32, 62, 66, 68] have been increasingly applied to volume generation, scene representation and data compression [73]. Since SIREN [59] introduced periodic activation functions to capture high-frequency details, INRs have emerged as a promising paradigm for representing data as continuous functions parameterized by neural networks. NeurComp [42] was the first to apply SIREN-based architectures with residual connections to volumetric scalar field compression, utilizing network weight quantization to achieve high CRs. To accelerate training and inference, grid-based approaches replace a large part of the fully connected layers with trainable feature grids (e.g., the multiresolution hash encoding of Instant-NGP [46]) that are queried by a lightweight multilayer perceptron (MLP), trading larger model sizes for orders-of-magnitude faster optimization and rendering. fV-SRN [77] uses a coarse grid of learnable latent features, together with a small MLP, to enable interactive volume rendering via custom CUDA TensorCore kernels. APMGSRN [81] improves reconstruction quality through multiple spatially adaptive feature grids that dynamically allocate network parameters to high-error regions. For time-varying data, KD-INR [22] employs knowledge distillation to sequentially compress individual timesteps into a single coherent model, while ECNR [64] utilizes a Laplacian pyramid for multiscale spatiotemporal decomposition with parallel MLPs. Other examples include STSR-INR [65], Meta-INR [83], MC-INR [60], and Lossless-INR [63]. F-Hash [61] further extends multiresolution hash encoding to the spatiotemporal domain through a feature-based tesseract encoding, accelerating convergence on timevarying volumes. However, INR methods require per-volume training, limiting their practical applicability for general-purpose compression. AE-based methods offer an alternative by learning generalizable latent representations. Deep Fluids [33] employs a generative AE to learn compact latent representations of fluid flows from parameterized simulations. FlowNet [20] uses a sparse stacked AE to learn implicit feature descriptors from streamlines or stream surfaces for clustering and selection. AE-SZ [39] integrates a convolutional AE as a predictor within the SZ framework to improve prediction accuracy. IDLat [57] generates latent representations guided by spatial importance maps, ensuring higher reconstruction quality in user-specified regions. While these methods demonstrate the potential of AEs for scientific data, they are typically trained on small-scale datasets from single simulations and achieve CRs around a few hundred at comparable quality, which does not significantly outperform state-of-the-art conventional compressors. Our EVOLVE addresses these limitations by training on a database with thousands of volumes to achieve cross-dataset generalization, while substantially surpassing both prior AE methods and conventional compressors in CR at equivalent reconstruction quality. 2.3 Learned Neural Data Compression End-to-end learned compression has achieved remarkable success in image coding, surpassing traditional codecs such as JPEG [72] in rate-distortion performance. Unlike conventional approaches that rely on handcrafted transforms and entropy coders, learned methods jointly optimize the encoder, decoder, and entropy model through gradient descent, enabling the network to discover data-adaptive representations. Ballé et al. [4, 5] established the foundational variational AE framework with nonlinear transforms and introduced the scale hyperprior to capture spatial dependencies in latent representations. Subsequent work focused on improving entropy modeling through context models. Minnen et al. [44] combined the hyperprior with autoregressive context modeling to outperform traditional codecs, and subsequent research [25, 26, 45]

further improved the accuracy of entropy estimation and decoding efficiency. For variable-rate compression, Cui et al. [14] proposed asymmetric gain units that rescale latent magnitudes, enabling continuous bitrate adjustment within a single model. These techniques have been extended to video compression [41, 58] and Gaussian splatting [11, 50, 67, 75], demonstrating the broad applicability of learned entropy modeling. In this work, we extend 2D learned compression techniques to 3D volumetric data by adapting context models for spatial volumes and training on a large-scale volume database to achieve generalization across the diverse scientific simulation domains it covers. 3

EVOLVE

This section presents the three core contributions of EVOLVE. First, we describe how we collect and curate our high-quality database of thousands of volumes for training in Section 3.1. Second, in Section 3.2, we provide an ablation-based roadmap that progressively upgrades an initially underperforming vanilla AE into EVOLVE, ultimately outperforming existing state-of-the-art conventional compressors. Third, we describe how we achieve variable-rate encoding in EVOLVE, enabling multiple CRs within a single model for more flexible deployment, as detailed in Section 3.3. 3.1

subset of the training volumes, with its own 95%/5% train/validation split (see Section 3.2). For evaluation, we deliberately avoid splitting along the temporal axis, as holding out intermediate timesteps of highly correlated sequences would inflate test performance. All test volumes (see Table 1) are drawn from unseen simulations or variables and are disjoint from the 9,921 collected volumes. 3.2 Ablation-Based Roadmap In this section, we describe the design process of EVOLVE through a systematic architectural ablation roadmap. We conduct this ablation study using a small subset of the 6,376 training database to identify effective design choices. Specifically, for each dataset listed in Table 1, we randomly sample 5% of the selected volumes, resulting in a subset of 319 volumes. We then use 95% of them (303 volumes) for training and the remaining 5% (16 volumes) for validation. Experimenting first on such a small subset of volumes allows us to compare different architectural variants while avoiding the optimization noise introduced by long training cycles on the large volume database.

Vanilla AE

Database Curation

Training an effective learned volume compression model requires a largescale, diverse, and high-quality database: models trained on narrow data distributions tend to overfit dataset-specific spatial statistics rather than learn transferable representations. Therefore, we curate our database by collecting time-varying scientific simulation data from multiple established repositories, including the Open SciVis Datasets [34], IEEE SciVis Contest archives [31], ETH Zürich visualization datasets [13], and Well [48], as well as our in-house simulation data. The resulting database comprises 21 datasets totaling 9,921 volumes across various scientific domains, before further filtering. This breadth ensures that the training data covers a wide range of physical quantities and spatial structures. Its full composition is provided in Appendix A. Time-varying simulations often exhibit significant redundancy across consecutive frames, which can be suboptimal for training. For example, five-jet contains 2,000 timesteps, yet the dynamics evolve slowly, making adjacent frames nearly indistinguishable. Directly using all 9,921 volumes can lead the model to memorize over-represented local patterns, reducing training efficiency and harming out-of-distribution performance, as observed in prior large-scale database curation studies [1, 36]. To improve the quality and diversity of our database, we design a curation process to identify and remove redundant volumes before training. Specifically, we adopt pHash [85], a perceptual hash-based strategy, to quantify structural similarity and identify near-duplicate samples. Specifically, we apply a 3D discrete cosine transform (DCT) to each volume, binarize the low-frequency DCT coefficients into a perceptual hash, and measure pairwise similarity using the Hamming distance between hashes. Inspired by nearest neighbor search [49, 76], we define the nearest-neighbor similarity (NNS) for each volume as the similarity to its closest neighbor and use the 95th-percentile of NNS within each dataset as a diversity metric. In our corpus, datasets such as five-jet and supercurrent initially exhibit very high redundancy, with the 95th-percentile NNS exceeding 0.95, indicating that many frames are near-duplicates. We perform deduplication for datasets whose 95th-percentile NNS exceeds 0.85 by clustering highly similar volumes with a union–find algorithm and retaining one representative per cluster. We repeat this filtering until the dataset-level NNS falls below 0.85. After each filtering iteration, we further validate the results with a visual inspection step that operates only on the clusters identified by deduplication: we render each retained representative side-by-side with its flagged near-duplicates using random transfer functions. Since the goal is to detect structural near-identity between candidate duplicates rather than to produce perceptually meaningful visualizations, randomly sampled transfer functions are sufficient. After filtering, the final training set contains 6,376 volumes with substantially improved diversity in data distribution. Train/validation/test split. All 6,376 curated volumes are used exclusively for training, while design decisions are made on a small validation

Factorized

Macro Design

Hyperprior GDN → ResBlock Context Non-uniform Channels

Micro Design

PConv Non-uniform Slicing Attention Block 64 → 16

Block Size

64 → 32 64 → 128

EVOLVE TTHRESH ZFP SZ3

Figure 2: We modernize a vanilla AE towards EVOLVE through progressive improvements. All evaluations are conducted on an NVIDIA RTX 4090 workstation. A hatched bar indicates that the corresponding modification is not adopted. Note that although only a 303-volume subset is used for training in this roadmap study, the final EVOLVE model already achieves performance that surpasses conventional compression methods.

We start with a vanilla convolutional AE that has strided convolution layers for downsampling and deconvolution layers for upsampling, with generalized divisive normalization (GDN) [3] as the normalizing function, following common design choices in prior AE-based compressors [39, 57]. To avoid excessive memory consumption, we adopt the same block-based processing strategy [21, 23] in which the network operates on fixed-size blocks cropped from the target volume. We set the block resolution to 64×64×64 as our starting point. During inference, the decoded blocks are merged using a weighted blending scheme, which leverages spatial overlap to avoid boundary discontinuities. The vanilla AE is optimized using a mean squared error (MSE) loss computed between the ground truth (GT) and the decoded volumes. We then progressively modernize each component of this vanilla AE toward our final EVOLVE model, as summarized in Figure 2. 3.2.1 Training Techniques All model variants in the roadmap are trained for 500 epochs on a small subset of volumes, using a learning rate of 0.0001 and a batch size of 16. Because voxel statistics vary significantly across scientific domains, direct optimization on a large database suffers from instability and slow convergence, so we adopt several modern training techniques. In particular, we employ a linear warmup [18] during the first 20 epochs to stabilize early-stage optimization under heterogeneous data

y! <latexit sha1_base64="fc6MKbrBdMChd5BIPE9WXkOlHc0=">AAACC3icbVDLSsNAFJ3UV42vVJduBovgqiQi1WXBje4q2Ae0oUwmk3bozCTMTCwh5BP8BLf6Ae7ErR/h2h8xabOwrQcuHM65l3M5XsSo0rb9bVQ2Nre2d6q75t7+weGRVTvuqjCWmHRwyELZ95AijArS0VQz0o8kQdxjpOdNbwu/90SkoqF41ElEXI7GggYUI51LI6s2nFGfTJBOhx5PkywzR1bdbthzwHXilKQOSrRH1s/QD3HMidCYIaUGjh1pN0VSU8xIZg5jRSKEp2hMBjkViBPlpvPXM3ieKz4MQpmP0HCu/r1IEVcq4V6+yZGeqFWvEP/zBrEObtyUiijWROBFUBAzqENY9AB9KgnWLMkJwpLmv0I8QRJhnbe1lOLxzCxacVY7WCfdy4bTbDQfruqt+7KfKjgFZ+ACOOAatMAdaIMOwGAGXsAreDOejXfjw/hcrFaM8uYELMH4+gXJSpr9</latexit>

Channel Context Φch <latexit sha1_base64="knrwADg+D+2Nq4xnZ/DSjmR+rY4=">AAACD3icbVDLSsNAFJ34rPEVFVdugkVwVRKR6koKblxWsA9oQphMp83QmSTM3Igl5CP8BLf6Ae7ErZ/g2h9x2mZhWw9cOJxzL+dywpQzBY7zbaysrq1vbFa2zO2d3b196+CwrZJMEtoiCU9kN8SKchbTFjDgtJtKikXIaScc3U78ziOViiXxA4xT6gs8jNmAEQxaCqxjLxS514xYEeQe0CfISVQUgVV1as4U9jJxS1JFJZqB9eP1E5IJGgPhWKme66Tg51gCI5wWppcpmmIywkPa0zTGgio/n75f2Gda6duDROqJwZ6qfy9yLJQai1BvCgyRWvQm4n9eL4PBtZ+zOM2AxmQWNMi4DYk96cLuM0kJ8LEmmEimf7VJhCUmoBubSwlFYZq6FXexg2XSvqi59Vr9/rLauCn7qaATdIrOkYuuUAPdoSZqIYJy9IJe0ZvxbLwbH8bnbHXFKG+O0ByMr1+aip0V</latexit>

Ψ <latexit sha1_base64="9CR26dkQPruavPW80+PkAl6UnLQ=">AAACAXicbVDLSsNAFL2prxpfVZduBovgqiQi1ZUU3LisYB/ShDKZTtqhM0mYmQglZOUnuNUPcCdu/RLX/ojTNgvbeuDC4Zx7ufeeIOFMacf5tkpr6xubW+Vte2d3b/+gcnjUVnEqCW2RmMeyG2BFOYtoSzPNaTeRFIuA004wvp36nScqFYujBz1JqC/wMGIhI1gb6dELROY1Fcv7lapTc2ZAq8QtSBUKNPuVH28Qk1TQSBOOleq5TqL9DEvNCKe57aWKJpiM8ZD2DI2woMrPZgfn6MwoAxTG0lSk0Uz9O5FhodREBKZTYD1Sy95U/M/rpTq89jMWJammEZkvClOOdIym36MBk5RoPjEEE8nMrYiMsMREm4wWtgQit22TirucwSppX9Tceq1+f1lt3BT5lOEETuEcXLiCBtxBE1pAQMALvMKb9Wy9Wx/W57y1ZBUzx7AA6+sX5lWXPA==</latexit>

<latexit sha1_base64="riQTojURDDZDKG6xZFw53HPImco=">AAACGXicbVDLSsNAFJ3UV42vqEsXDhahgpSk+Oiy4MaVVLAPaGKYTCft0JkkzEyEErL0K/wEt/oB7sStK9f+iEmbha0euHA4517uvceLGJXKNL+00tLyyupaeV3f2Nza3jF29zoyjAUmbRyyUPQ8JAmjAWkrqhjpRYIg7jHS9cZXud99IELSMLhTk4g4HA0D6lOMVCa5xqHNkRphxJKbtGrz2I1OoS3pkCM3uq+f6K5RMWvmFPAvsQpSAQVarvFtD0IccxIozJCUfcuMlJMgoShmJNXtWJII4TEakn5GA8SJdJLpIyk8zpQB9EORVaDgVP09kSAu5YR7WWd+tlz0cvE/rx8rv+EkNIhiRQI8W+THDKoQ5qnAARUEKzbJCMKCZrdCPEICYZVlN7fF46mep2ItZvCXdOo166J2fntWaTaKfMrgAByBKrDAJWiCa9ACbYDBI3gGL+BVe9LetHftY9Za0oqZfTAH7fMHxwyfkw==</latexit>

N (µp , σp2 )

<latexit sha1_base64="xNeEikwST0zb6hpeGH3c1AK0HG8=">AAACEnicbVDLSsNAFJ34rPEVdSO4CRahbkoiUl0W3Oiugn1AG8tkMmmHzkzCzEQJIX6Fn+BWP8CduPUHXPsjTtssbOuBC4dz7uXee/yYEqkc59tYWl5ZXVsvbZibW9s7u9befktGiUC4iSIaiY4PJaaE46YiiuJOLDBkPsVtf3Q19tsPWEgS8TuVxthjcMBJSBBUWupbh71HEuAhVFnPZ1ma5/dZxT3NTbNvlZ2qM4G9SNyClEGBRt/66QURShjmClEoZdd1YuVlUCiCKM7NXiJxDNEIDnBXUw4Zll42+SC3T7QS2GEkdHFlT9S/ExlkUqbM150MqqGc98bif143UeGllxEeJwpzNF0UJtRWkT2Oww6IwEjRVBOIBNG32mgIBURKhzazxWfTVNz5DBZJ66zq1qq12/Ny/abIpwSOwDGoABdcgDq4Bg3QBAg8gRfwCt6MZ+Pd+DA+p61LRjFzAGZgfP0CAQ6dJQ==</latexit>

<latexit sha1_base64="ipBbxlb2oCQlIQ7GylpxW/I0mOQ=">AAACEHicbVDLSsNAFJ3UV42vqODGTbAIdVOSItVlwY3uKtgHNLFMJpN26EwSZiZKiPkJP8GtfoA7cesfuPZHnLZZ2NYDFw7n3Mu5HC+mREjL+tZKK6tr6xvlTX1re2d3z9g/6Igo4Qi3UUQj3vOgwJSEuC2JpLgXcwyZR3HXG19N/O4D5oJE4Z1MY+wyOAxJQBCUShoYR84j8fEIyszxWJbm+X1WrZ/lA6Ni1awpzGViF6QCCrQGxo/jRyhhOJSIQiH6thVLN4NcEkRxrjuJwDFEYzjEfUVDyLBws+n/uXmqFN8MIq4mlOZU/XuRQSZEyjy1yaAciUVvIv7n9RMZXLoZCeNE4hDNgoKEmjIyJ2WYPuEYSZoqAhEn6lcTjSCHSKrK5lI8luu6asVe7GCZdOo1u1Fr3J5XmjdFP2VwDE5AFdjgAjTBNWiBNkDgCbyAV/CmPWvv2of2OVstacXNIZiD9vULhw6c/g==</latexit>

<latexit sha1_base64="z7ayfaZMY3n2guI+NLWE6jAK0f0=">AAACGXicbVDLSsNAFJ3UV42vqEsXDhahgpSk+Oiy4MaVVLAPaGKYTCftkJkkzEyEErL0K/wEt/oB7sStK9f+iEnbha0euHA4517uvceLGZXKNL+00tLyyupaeV3f2Nza3jF29zoySgQmbRyxSPQ8JAmjIWkrqhjpxYIg7jHS9YKrwu8+ECFpFN6pcUwcjoYh9SlGKpdc49DmSI0wYulNVrV54gan0JZ0yJEb3NdPdNeomDVzAviXWDNSATO0XOPbHkQ44SRUmCEp+5YZKydFQlHMSKbbiSQxwgEakn5OQ8SJdNLJIxk8zpUB9CORV6jgRP09kSIu5Zh7eWdxtlz0CvE/r58ov+GkNIwTRUI8XeQnDKoIFqnAARUEKzbOCcKC5rdCPEICYZVnN7fF45lepGItZvCXdOo166J2fntWaTZm+ZTBATgCVWCBS9AE16AF2gCDR/AMXsCr9qS9ae/ax7S1pM1m9sEctM8ftuCfiQ==</latexit>

<latexit sha1_base64="dnphVHUsNrcUsTw+ExuN55oK3P4=">AAACEHicbVDLSsNAFJ3UV42vqODGTbAIdVMSleqy4EZ3FewDmlgmk0k7dCYJMxMlxPyEn+BWP8CduPUPXPsjTtssbOuBC4dz7uVcjhdTIqRlfWulpeWV1bXyur6xubW9Y+zutUWUcIRbKKIR73pQYEpC3JJEUtyNOYbMo7jjja7GfucBc0Gi8E6mMXYZHIQkIAhKJfWNA+eR+HgIZeZ4LEvz/D6rnp3kfaNi1awJzEViF6QCCjT7xo/jRyhhOJSIQiF6thVLN4NcEkRxrjuJwDFEIzjAPUVDyLBws8n/uXmsFN8MIq4mlOZE/XuRQSZEyjy1yaAcinlvLP7n9RIZXLoZCeNE4hBNg4KEmjIyx2WYPuEYSZoqAhEn6lcTDSGHSKrKZlI8luu6asWe72CRtE9rdr1Wvz2vNG6KfsrgEByBKrDBBWiAa9AELYDAE3gBr+BNe9betQ/tc7pa0oqbfTAD7esXiKWc/w==</latexit>

y!(1) y!(2)

N (µk , σk2 )

Ψ

Ψ

<latexit sha1_base64="6d4Fpicj+HSkoy8YJ34mSWew5Eg=">AAACD3icbVDLSsNAFJ34rPEVFVdugkVwVRKR6koKblxWsA9oQphMJ+3QmSTM3Igl5CP8BLf6Ae7ErZ/g2h9x2mZhWw9cOJxzL+dywpQzBY7zbaysrq1vbFa2zO2d3b196+CwrZJMEtoiCU9kN8SKchbTFjDgtJtKikXIaScc3U78ziOViiXxA4xT6gs8iFnECAYtBdaxF4rcaw5ZEeQe0CfIVVoUgVV1as4U9jJxS1JFJZqB9eP1E5IJGgPhWKme66Tg51gCI5wWppcpmmIywgPa0zTGgio/n75f2Gda6dtRIvXEYE/Vvxc5FkqNRag3BYahWvQm4n9eL4Po2s9ZnGZAYzILijJuQ2JPurD7TFICfKwJJpLpX20yxBIT0I3NpYSiME3dirvYwTJpX9Tceq1+f1lt3JT9VNAJOkXnyEVXqIHuUBO1EEE5ekGv6M14Nt6ND+NztrpilDdHaA7G1y/Awp0t</latexit>

<latexit sha1_base64="9CR26dkQPruavPW80+PkAl6UnLQ=">AAACAXicbVDLSsNAFL2prxpfVZduBovgqiQi1ZUU3LisYB/ShDKZTtqhM0mYmQglZOUnuNUPcCdu/RLX/ojTNgvbeuDC4Zx7ufeeIOFMacf5tkpr6xubW+Vte2d3b/+gcnjUVnEqCW2RmMeyG2BFOYtoSzPNaTeRFIuA004wvp36nScqFYujBz1JqC/wMGIhI1gb6dELROY1Fcv7lapTc2ZAq8QtSBUKNPuVH28Qk1TQSBOOleq5TqL9DEvNCKe57aWKJpiM8ZD2DI2woMrPZgfn6MwoAxTG0lSk0Uz9O5FhodREBKZTYD1Sy95U/M/rpTq89jMWJammEZkvClOOdIym36MBk5RoPjEEE8nMrYiMsMREm4wWtgQit22TirucwSppX9Tceq1+f1lt3BT5lOEETuEcXLiCBtxBE1pAQMALvMKb9Wy9Wx/W57y1ZBUzx7AA6+sX5lWXPA==</latexit>

<latexit sha1_base64="e/cffUJu+Kq72cDe7kaKYCDqtqg=">AAACGXicbVDLSsNAFJ3UV42vqEsXDhahgpSkSO2y4MaVVLAPaGKYTCfttDNJmJkIJWTpV/gJbvUD3IlbV679EZO2C1s9cOFwzr3ce48XMSqVaX5phZXVtfWN4qa+tb2zu2fsH7RlGAtMWjhkoeh6SBJGA9JSVDHSjQRB3GOk442vcr/zQISkYXCnJhFxOBoE1KcYqUxyjWObIzXEiCU3adnmsTs6h7akA47c0X31THeNklkxp4B/iTUnJTBH0zW+7X6IY04ChRmSsmeZkXISJBTFjKS6HUsSITxGA9LLaIA4kU4yfSSFp5nSh34osgoUnKq/JxLEpZxwL+vMz5bLXi7+5/Vi5dedhAZRrEiAZ4v8mEEVwjwV2KeCYMUmGUFY0OxWiIdIIKyy7Ba2eDzV81Ss5Qz+kna1YtUqtduLUqM+z6cIjsAJKAMLXIIGuAZN0AIYPIJn8AJetSftTXvXPmatBW0+cwgWoH3+ALP2n4g=</latexit>

<latexit sha1_base64="DSsFbhwYGDAyJM/RilLrpxkgOmY=">AAACGXicbVDLSsNAFJ3UV42vqEsXDhahgpSkSO2y4MaVVLAPaGKYTCft0JkkzEyEErL0K/wEt/oB7sStK9f+iEmbha0euHA4517uvceLGJXKNL+00srq2vpGeVPf2t7Z3TP2D7oyjAUmHRyyUPQ9JAmjAekoqhjpR4Ig7jHS8yZXud97IELSMLhT04g4HI0C6lOMVCa5xrHNkRpjxJKbtGrz2KXn0JZ0xJFL7+tnumtUzJo5A/xLrIJUQIG2a3zbwxDHnAQKMyTlwDIj5SRIKIoZSXU7liRCeIJGZJDRAHEinWT2SApPM2UI/VBkFSg4U39PJIhLOeVe1pmfLZe9XPzPG8TKbzoJDaJYkQDPF/kxgyqEeSpwSAXBik0zgrCg2a0Qj5FAWGXZLWzxeKrnqVjLGfwl3XrNatQatxeVVrPIpwyOwAmoAgtcgha4Bm3QARg8gmfwAl61J+1Ne9c+5q0lrZg5BAvQPn8AsLqfhg==</latexit>

<latexit sha1_base64="0PLGRf1df5raM5QeQprK9XT1LpQ=">AAACFXicbVDLSsNAFJ34rPEVddnNYBEqSEmK1C4LblxJBfuAJpbJdNoOnZmEmYlQQhZ+hZ/gVj/Anbh17dofMWmzsK0HLhzOuZd77/FDRpW27W9jbX1jc2u7sGPu7u0fHFpHx20VRBKTFg5YILs+UoRRQVqaaka6oSSI+4x0/Ml15nceiVQ0EPd6GhKPo5GgQ4qRTqW+VXQ50mOMWHyblF0eXUBX0RFHD9Vzs2+V7Io9A1wlTk5KIEezb/24gwBHnAiNGVKq59ih9mIkNcWMJKYbKRIiPEEj0kupQJwoL549kcCzVBnAYSDTEhrO1L8TMeJKTbmfdmYnq2UvE//zepEe1r2YijDSROD5omHEoA5glggcUEmwZtOUICxpeivEYyQR1mluC1t8nphZKs5yBqukXa04tUrt7rLUqOf5FEARnIIycMAVaIAb0AQtgMETeAGv4M14Nt6ND+Nz3rpm5DMnYAHG1y92yZ3O</latexit>

N (µ, ω 2 )

N (µi , σi2 )

(a) factorized model

N (µj , σj2 )

gch

DWConv x2 Blocks

<latexit sha1_base64="eeJvP0Pi/rnMWG5nAZBx+7Hh1Vw=">AAACDXicbVDLSsNAFJ34rPEV7dJNsAh1UxIp1WXBje4q2Ae0MUymk3bo5MHMjRhCvsFPcKsf4E7c+g2u/RGnbRa29cCFwzn3ci7HizmTYFnf2tr6xubWdmlH393bPzg0jo47MkoEoW0S8Uj0PCwpZyFtAwNOe7GgOPA47XqT66nffaRCsii8hzSmToBHIfMZwaAk1yiPHrJq/Tx3swHQJ8jIOM9do2LVrBnMVWIXpIIKtFzjZzCMSBLQEAjHUvZtKwYnwwIY4TTXB4mkMSYTPKJ9RUMcUOlks+dz80wpQ9OPhJoQzJn69yLDgZRp4KnNAMNYLntT8T+vn4B/5WQsjBOgIZkH+Qk3ITKnTZhDJigBniqCiWDqV5OMscAEVF8LKV6Q67pqxV7uYJV0Lmp2o9a4q1eat0U/JXSCTlEV2egSNdENaqE2IihFL+gVvWnP2rv2oX3OV9e04qaMFqB9/QISRJuu</latexit>

<latexit sha1_base64="UP2Kjsk9LC2JRAGYhAkFnDyheMY=">AAACEHicbVDLSsNAFJ3UV42vqODGTbAIdVMS0eqy4EZ3FewDmlgmk0k7dCYJMxMlxPyEn+BWP8CduPUPXPsjTtssbOuBC4dz7uVcjhdTIqRlfWulpeWV1bXyur6xubW9Y+zutUWUcIRbKKIR73pQYEpC3JJEUtyNOYbMo7jjja7GfucBc0Gi8E6mMXYZHIQkIAhKJfWNA+eR+HgIZeZ4LEvz/D6rnp/kfaNi1awJzEViF6QCCjT7xo/jRyhhOJSIQiF6thVLN4NcEkRxrjuJwDFEIzjAPUVDyLBws8n/uXmsFN8MIq4mlOZE/XuRQSZEyjy1yaAcinlvLP7n9RIZXLoZCeNE4hBNg4KEmjIyx2WYPuEYSZoqAhEn6lcTDSGHSKrKZlI8luu6asWe72CRtE9rdr1Wvz2rNG6KfsrgEByBKrDBBWiAa9AELYDAE3gBr+BNe9betQ/tc7pa0oqbfTAD7esXi9OdAQ==</latexit>

y!(4) y!(5) <latexit sha1_base64="dKtIDQiM1x55POOXXNxE2zLE9qc=">AAACDXicbVDLSsNAFJ34rPEV7dJNsAh1UxLR6rLgRncV7APaGCbTSTt08mDmRgwh3+AnuNUPcCdu/QbX/ojTNgvbeuDC4Zx7OZfjxZxJsKxvbWV1bX1js7Slb+/s7u0bB4dtGSWC0BaJeCS6HpaUs5C2gAGn3VhQHHicdrzx9cTvPFIhWRTeQxpTJ8DDkPmMYFCSa5SHD1n14jR3sz7QJ8jIKM9do2LVrCnMZWIXpIIKNF3jpz+ISBLQEAjHUvZsKwYnwwIY4TTX+4mkMSZjPKQ9RUMcUOlk0+dz80QpA9OPhJoQzKn69yLDgZRp4KnNAMNILnoT8T+vl4B/5WQsjBOgIZkF+Qk3ITInTZgDJigBniqCiWDqV5OMsMAEVF9zKV6Q67pqxV7sYJm0z2p2vVa/O680bot+SugIHaMqstElaqAb1EQtRFCKXtAretOetXftQ/ucra5oxU0ZzUH7+gUT55uv</latexit>

(3)

gch

(4)

(3)

Φch Φch

gch

(5)

gch

<latexit sha1_base64="9CR26dkQPruavPW80+PkAl6UnLQ=">AAACAXicbVDLSsNAFL2prxpfVZduBovgqiQi1ZUU3LisYB/ShDKZTtqhM0mYmQglZOUnuNUPcCdu/RLX/ojTNgvbeuDC4Zx7ufeeIOFMacf5tkpr6xubW+Vte2d3b/+gcnjUVnEqCW2RmMeyG2BFOYtoSzPNaTeRFIuA004wvp36nScqFYujBz1JqC/wMGIhI1gb6dELROY1Fcv7lapTc2ZAq8QtSBUKNPuVH28Qk1TQSBOOleq5TqL9DEvNCKe57aWKJpiM8ZD2DI2woMrPZgfn6MwoAxTG0lSk0Uz9O5FhodREBKZTYD1Sy95U/M/rpTq89jMWJammEZkvClOOdIym36MBk5RoPjEEE8nMrYiMsMREm4wWtgQit22TirucwSppX9Tceq1+f1lt3BT5lOEETuEcXLiCBtxBE1pAQMALvMKb9Wy9Wx/W57y1ZBUzx7AA6+sX5lWXPA==</latexit>

Conv 1x1 <latexit sha1_base64="81f/3Otw8jzoUbCf1+RGsx/B22c=">AAACFXicbVDLSsNAFJ3UV42vqstugkWom5IUqS4LbnRXwT6gqWUynTZDZ5IwcyOWkIVf4Se41Q9wJ25du/ZHnLZZ2NYDFw7n3Mu993gRZwps+9vIra1vbG7lt82d3b39g8LhUUuFsSS0SUIeyo6HFeUsoE1gwGknkhQLj9O2N76a+u0HKhULgzuYRLQn8ChgQ0YwaKlfKLqeSNyGz9L7pFw9S/uJC/QREuKnab9Qsiv2DNYqcTJSQhka/cKPOwhJLGgAhGOluo4dQS/BEhjhNDXdWNEIkzEe0a6mARZU9ZLZE6l1qpWBNQylrgCsmfp3IsFCqYnwdKfA4Ktlbyr+53VjGF72EhZEMdCAzBcNY25BaE0TsQZMUgJ8ogkmkulbLeJjiQno3Ba2eCI1TZ2Ks5zBKmlVK06tUrs9L9VvsnzyqIhOUBk56ALV0TVqoCYi6Am9oFf0Zjwb78aH8TlvzRnZzDFagPH1C5qbnzU=</latexit>

<latexit sha1_base64="JUeoVlToVm+nX58BtX7/EHa3zJ4=">AAACFXicbVC7TsNAEDyHVwivAGUaiwgpNJENKFBGooEuSOQhxSE6X87xKXe2dbdGRJYLvoJPoIUPoEO01NT8CJfEBUkYaaXRzK52d9yIMwWW9W3kVlbX1jfym4Wt7Z3dveL+QUuFsSS0SUIeyo6LFeUsoE1gwGknkhQLl9O2O7qa+O0HKhULgzsYR7Qn8DBgHiMYtNQvlhxXJE7DZ+l9Ujk7SfuJA/QREuKnab9YtqrWFOYysTNSRhka/eKPMwhJLGgAhGOlurYVQS/BEhjhNC04saIRJiM8pF1NAyyo6iXTJ1LzWCsD0wulrgDMqfp3IsFCqbFwdafA4KtFbyL+53Vj8C57CQuiGGhAZou8mJsQmpNEzAGTlAAfa4KJZPpWk/hYYgI6t7ktrkgLBZ2KvZjBMmmdVu1atXZ7Xq7fZPnkUQkdoQqy0QWqo2vUQE1E0BN6Qa/ozXg23o0P43PWmjOymUM0B+PrF5w+nzY=</latexit>

(2)

Φch

3D Checkerboard Spatial Context

(b) hyperprior model

<latexit sha1_base64="ddy8tcVwfnwd+/yQcri3o7RlUtE=">AAACDXicbVDLSsNAFJ34rPEV7dJNsAh1UxKV6rLgRncV7APaGCbTSTt08mDmRgwh3+AnuNUPcCdu/QbX/ojTNgvbeuDC4Zx7OZfjxZxJsKxvbWV1bX1js7Slb+/s7u0bB4dtGSWC0BaJeCS6HpaUs5C2gAGn3VhQHHicdrzx9cTvPFIhWRTeQxpTJ8DDkPmMYFCSa5SHD1n1/DR3sz7QJ8jIKM9do2LVrCnMZWIXpIIKNF3jpz+ISBLQEAjHUvZsKwYnwwIY4TTX+4mkMSZjPKQ9RUMcUOlk0+dz80QpA9OPhJoQzKn69yLDgZRp4KnNAMNILnoT8T+vl4B/5WQsjBOgIZkF+Qk3ITInTZgDJigBniqCiWDqV5OMsMAEVF9zKV6Q67pqxV7sYJm0z2p2vVa/u6g0bot+SugIHaMqstElaqAb1EQtRFCKXtAretOetXftQ/ucra5oxU0ZzUH7+gUQoZut</latexit>

(2)

Φsp

<latexit sha1_base64="6d4Fpicj+HSkoy8YJ34mSWew5Eg=">AAACD3icbVDLSsNAFJ34rPEVFVdugkVwVRKR6koKblxWsA9oQphMJ+3QmSTM3Igl5CP8BLf6Ae7ErZ/g2h9x2mZhWw9cOJxzL+dywpQzBY7zbaysrq1vbFa2zO2d3b196+CwrZJMEtoiCU9kN8SKchbTFjDgtJtKikXIaScc3U78ziOViiXxA4xT6gs8iFnECAYtBdaxF4rcaw5ZEeQe0CfIVVoUgVV1as4U9jJxS1JFJZqB9eP1E5IJGgPhWKme66Tg51gCI5wWppcpmmIywgPa0zTGgio/n75f2Gda6dtRIvXEYE/Vvxc5FkqNRag3BYahWvQm4n9eL4Po2s9ZnGZAYzILijJuQ2JPurD7TFICfKwJJpLpX20yxBIT0I3NpYSiME3dirvYwTJpX9Tceq1+f1lt3JT9VNAJOkXnyEVXqIHuUBO1EEE5ekGv6M14Nt6ND+NztrpilDdHaA7G1y/Awp0t</latexit>

<latexit sha1_base64="4uVp0pS8YVMKx0s6xb1sCZJjJY8=">AAACEHicbVDLSsNAFJ3UV42vqODGTbAIdVMSKdVlwY3uKtgHNLFMJpN26EwSZiZKiPkJP8GtfoA7cesfuPZHnLZZ2NYDFw7n3Mu5HC+mREjL+tZKK6tr6xvlTX1re2d3z9g/6Igo4Qi3UUQj3vOgwJSEuC2JpLgXcwyZR3HXG19N/O4D5oJE4Z1MY+wyOAxJQBCUShoYR84j8fEIyszxWJbm+X1WrZ/lA6Ni1awpzGViF6QCCrQGxo/jRyhhOJSIQiH6thVLN4NcEkRxrjuJwDFEYzjEfUVDyLBws+n/uXmqFN8MIq4mlOZU/XuRQSZEyjy1yaAciUVvIv7n9RMZXLoZCeNE4hDNgoKEmjIyJ2WYPuEYSZoqAhEn6lcTjSCHSKrK5lI8luu6asVe7GCZdM5rdqPWuK1XmjdFP2VwDE5AFdjgAjTBNWiBNkDgCbyAV/CmPWvv2of2OVstacXNIZiD9vULijydAA==</latexit>

y!(3)

<latexit sha1_base64="N07nBRf5y0m9mQjCFKXH+5j6feE=">AAACDXicbVDLSsNAFJ3UV42vaJdugkWom5IUqS4LbnRXwbZCG8NkOmmHTh7M3Igh5Bv8BLf6Ae7Erd/g2h9x2mZhWw9cOJxzL+dyvJgzCZb1rZXW1jc2t8rb+s7u3v6BcXjUlVEiCO2QiEfi3sOSchbSDjDg9D4WFAcepz1vcjX1e49USBaFd5DG1AnwKGQ+IxiU5BqV0UNWa5zlbjYA+gQZGee5a1StujWDuUrsglRRgbZr/AyGEUkCGgLhWMq+bcXgZFgAI5zm+iCRNMZkgke0r2iIAyqdbPZ8bp4qZWj6kVATgjlT/15kOJAyDTy1GWAYy2VvKv7n9RPwL52MhXECNCTzID/hJkTmtAlzyAQlwFNFMBFM/WqSMRaYgOprIcULcl1XrdjLHaySbqNuN+vN2/Nq66bop4yO0QmqIRtdoBa6Rm3UQQSl6AW9ojftWXvXPrTP+WpJK24qaAHa1y8O/pus</latexit>

Φsp

Slicing

(c) our context model

<latexit sha1_base64="4Jf80L6xuhy6Z1UToIKKb0tuARc=">AAACFXicbVDLSsNAFJ3UV42vqMtugkWom5JIqS4LbnRXwT6gqWEynbZDZ5IwcyOWkIVf4Se41Q9wJ25du/ZHnD4WtvXAhcM593LvPUHMmQLH+TZya+sbm1v5bXNnd2//wDo8aqookYQ2SMQj2Q6wopyFtAEMOG3HkmIRcNoKRlcTv/VApWJReAfjmHYFHoSszwgGLflWwQtE6tWHLLtPS5WzzE89oI+QkmGW+VbRKTtT2KvEnZMimqPuWz9eLyKJoCEQjpXquE4M3RRLYITTzPQSRWNMRnhAO5qGWFDVTadPZPapVnp2P5K6QrCn6t+JFAulxiLQnQLDUC17E/E/r5NA/7KbsjBOgIZktqifcBsie5KI3WOSEuBjTTCRTN9qkyGWmIDObWFLIDLT1Km4yxmskuZ52a2Wq7eVYu1mnk8eFdAJKiEXXaAaukZ11EAEPaEX9IrejGfj3fgwPmetOWM+c4wWYHz9Ap3hnzc=</latexit>

Φch

<latexit sha1_base64="P9y5YsGkqya0BH/k4NcL4DjUcW4=">AAACFXicbVC7TsNAEDyHVwivAGUaiwgpNJGNIFBGooEuSOQhxSE6X87xKXe2dbdGRJYLvoJPoIUPoEO01NT8CJfEBUkYaaXRzK52d9yIMwWW9W3kVlbX1jfym4Wt7Z3dveL+QUuFsSS0SUIeyo6LFeUsoE1gwGknkhQLl9O2O7qa+O0HKhULgzsYR7Qn8DBgHiMYtNQvlhxXJE7DZ+l9Ujk/SfuJA/QREuKnab9YtqrWFOYysTNSRhka/eKPMwhJLGgAhGOlurYVQS/BEhjhNC04saIRJiM8pF1NAyyo6iXTJ1LzWCsD0wulrgDMqfp3IsFCqbFwdafA4KtFbyL+53Vj8C57CQuiGGhAZou8mJsQmpNEzAGTlAAfa4KJZPpWk/hYYgI6t7ktrkgLBZ2KvZjBMmmdVu1atXZ7Vq7fZPnkUQkdoQqy0QWqo2vUQE1E0BN6Qa/ozXg23o0P43PWmjOymUM0B+PrF5+Enzg=</latexit>

(4)

(5)

(d) channel-context architecture

Figure 3: Comparison of different entropy models and our proposed context model. (a) Factorized model with independent Gaussian assumptions. (b) Hyperprior model that conditions latent distributions on hyperprior Ψ. (c) Our context model that further combines the 3D checkerboard spatial and channel context to achieve a more accurate probability estimation. (d) Channel-context architecture that decodes latent channels progressively, conditioning each slice on previously decoded ones.

<latexit sha1_base64="rKzUGeCCcjEhg7q2nW276SD2kGc=">AAACCnicbVDLSsNAFJ3UV42vqEs3g0VwVRKR6rLgRncVbCu0oUwmk3bozCTMTCoh5A/8BLf6Ae7ErT/h2h9x2mZhWw9cOJxzL+dygoRRpV3326qsrW9sblW37Z3dvf0D5/Coo+JUYtLGMYvlY4AUYVSQtqaakcdEEsQDRrrB+GbqdydEKhqLB50lxOdoKGhEMdJGGjhO/4mGZIR03g94nhXFwKm5dXcGuEq8ktRAidbA+emHMU45ERozpFTPcxPt50hqihkp7H6qSILwGA1Jz1CBOFF+Pvu8gGdGCWEUSzNCw5n69yJHXKmMB2aTIz1Sy95U/M/rpTq69nMqklQTgedBUcqgjuG0BhhSSbBmmSEIS2p+hXiEJMLalLWQEvDCtk0r3nIHq6RzUfca9cb9Za15V/ZTBSfgFJwDD1yBJrgFLdAGGEzAC3gFb9az9W59WJ/z1YpV3hyDBVhfv41wmuk=</latexit>

ED

Ψ <latexit sha1_base64="LFf0xgl1IRbrsgpGeHqK1gZu7lE=">AAACAXicbVDLSsNAFL2prxpfVZduBovgqiQi1WXBje4q2Ic0oUymk3boTBJmJkIJWfkJbvUD3Ilbv8S1P+K0zcK2HrhwOOde7r0nSDhT2nG+rdLa+sbmVnnb3tnd2z+oHB61VZxKQlsk5rHsBlhRziLa0kxz2k0kxSLgtBOMb6Z+54lKxeLoQU8S6gs8jFjICNZGevQCkXlNxfJ+perUnBnQKnELUoUCzX7lxxvEJBU00oRjpXquk2g/w1Izwmlue6miCSZjPKQ9QyMsqPKz2cE5OjPKAIWxNBVpNFP/TmRYKDURgekUWI/UsjcV//N6qQ6v/YxFSappROaLwpQjHaPp92jAJCWaTwzBRDJzKyIjLDHRJqOFLYHIbduk4i5nsEraFzW3XqvfX1Ybd0U+ZTiBUzgHF66gAbfQhBYQEPACr/BmPVvv1of1OW8tWcXMMSzA+voF6aSXRw==</latexit>

k /! y <k y!anc <latexit sha1_base64="WT05sFq+fGDWu+zAbb1Pyl4q7d4=">AAACM3icbVDLSgMxFM34tr6qLt0Ei+Cqzoj4ABeCG90p2Fbo1JJJb21oMjMkd9QS5lv8Cj/Bra7Fnbr1H0xrF2o9EDiccy4390SpFAZ9/8UbG5+YnJqemS3MzS8sLhWXV6omyTSHCk9koi8jZkCKGCooUMJlqoGpSEIt6h73/doNaCOS+AJ7KTQUu45FW3CGTmoWD8Jb0YIOQxtGyvby/Mp286YNEe7QspjnOd2io5lDFyqW/LI/AB0lwZCUyBBnzeJ72Ep4piBGLpkx9cBPsWGZRsEl5IUwM5Ay3mXXUHc0ZgpMww5OzOmGU1q0nWj3YqQD9eeEZcqYnopcUjHsmL9eX/zPq2fY3m9YEacZgjt3sKidSYoJ7fdFW0IDR9lzhHEt3F8p7zDNOLpWf22JVF4ouFaCvx2Mkup2Odgt757vlI5Oh/3MkDWyTjZJQPbIETkhZ6RCOLknj+SJPHsP3qv35n18R8e84cwq+QXv8wtEsa0o</latexit>

(a) EVOLVE Architecture

<latexit sha1_base64="0RYnAEficKBglXp4+2D7exiUj6U=">AAACC3icbVDLSsNAFJ34rPGV6tJNsAiuSiJSXRbc6K6CfUAbw2Q6aYfOJGHmRi0hn+AnuNUPcCdu/QjX/ojTNgvbeuDC4Zx7OZcTJJwpcJxvY2V1bX1js7Rlbu/s7u1b5YOWilNJaJPEPJadACvKWUSbwIDTTiIpFgGn7WB0NfHbD1QqFkd3ME6oJ/AgYiEjGLTkW+XBfTbK/awH9AkyMsxz36o4VWcKe5m4BamgAg3f+un1Y5IKGgHhWKmu6yTgZVgCI5zmZi9VNMFkhAe0q2mEBVVeNn09t0+00rfDWOqJwJ6qfy8yLJQai0BvCgxDtehNxP+8bgrhpZexKEmBRmQWFKbchtie9GD3maQE+FgTTCTTv9pkiCUmoNuaSwlEbpq6FXexg2XSOqu6tWrt9rxSvyn6KaEjdIxOkYsuUB1dowZqIoIe0Qt6RW/Gs/FufBifs9UVo7g5RHMwvn4BmlObgA==</latexit>

<latexit sha1_base64="yGCYlBWLSFaoBXNCdmHZAydXIkA=">AAACE3icbVDLSsNAFJ34rPUVdSVugkVwVRKR6rLgRncV7AOaGibTSTt0JgkzN2IZgl/hJ7jVD3Anbv0A1/6I08fCth64cDjnXu69J0w5U+C639bS8srq2npho7i5tb2za+/tN1SSSULrJOGJbIVYUc5iWgcGnLZSSbEIOW2Gg6uR33ygUrEkvoNhSjsC92IWMYLBSIF96IdC+7U+y+/1IA+0D/QRNOnneWCX3LI7hrNIvCkpoSlqgf3jdxOSCRoD4Viptuem0NFYAiOc5kU/UzTFZIB7tG1ojAVVHT1+IXdOjNJ1okSaisEZq38nNBZKDUVoOgWGvpr3RuJ/XjuD6LKjWZxmQGMyWRRl3IHEGeXhdJmkBPjQEEwkM7c6pI8lJmBSm9kSirxYNKl48xksksZZ2auUK7fnperNNJ8COkLH6BR56AJV0TWqoToi6Am9oFf0Zj1b79aH9TlpXbKmMwdoBtbXLx9tnwk=</latexit>

<latexit sha1_base64="BX44HuC+mVoIvHsVbh3YrEtUoJc=">AAACD3icbVDLSsNAFJ34rPEVFVdugkVwVRKR6sJFQRcuK9gHNLFMJpN26MwkzEyUEPIRfoJb/QB34tZPcO2POG2zsK0HLhzOuZdzOUFCiVSO820sLa+srq1XNszNre2dXWtvvy3jVCDcQjGNRTeAElPCcUsRRXE3ERiygOJOMLoe+51HLCSJ+b3KEuwzOOAkIggqLfWtQ++JhHgIVe4FLM+K4iG/GhV9q+rUnAnsReKWpApKNPvWjxfGKGWYK0ShlD3XSZSfQ6EIorgwvVTiBKIRHOCephwyLP188n5hn2gltKNY6OHKnqh/L3LIpMxYoDcZVEM5743F/7xeqqJLPyc8SRXmaBoUpdRWsT3uwg6JwEjRTBOIBNG/2mgIBURKNzaTErDCNHUr7nwHi6R9VnPrtfrdebVxU/ZTAUfgGJwCF1yABrgFTdACCOTgBbyCN+PZeDc+jM/p6pJR3hyAGRhfv5F0nRM=</latexit>

k 0/! yanc <latexit sha1_base64="nK16FA/EmKDimlz0E69U1e4FhT4=">AAACKHicbVDLSgMxFM34tr6qLt0Ei+CqzoioS0EXLhWsCp1aMukdG5rMDMkdtYT8hV/hJ7jVD3AnbvVHTGsXvg4EDufcy805SSGFwTB8C8bGJyanpmdmK3PzC4tL1eWVc5OXmkOD5zLXlwkzIEUGDRQo4bLQwFQi4SLpHQ78ixvQRuTZGfYLaCl2nYlUcIZealfrsWLYTVIbOrpF41vRgS5DGyfK9p27sj3XtjHCHVqWcefa1VpYD4egf0k0IjUywkm7+hF3cl4qyJBLZkwzCgtsWaZRcAmuEpcGCsZ77BqanmZMgWnZYS5HN7zSoWmu/cuQDtXvG5YpY/oq8ZODFOa3NxD/85olpvstK7KiRPCxhofSUlLM6aAk2hEaOMq+J4xr4f9KeZdpxtFX+eNKolyl4luJfnfwl5xv16Pd+u7pTu3gaNTPDFkj62STRGSPHJBjckIahJN78kieyHPwELwEr8Hb1+hYMNpZJT8QvH8C7cmn2A==</latexit>

<latexit sha1_base64="OZLT6fx6ojyXMWh/9MQ7aRp/hps=">AAACC3icbVDLSsNAFJ34rPGV6tJNsAiuSiJSXRbc6K6CfUAbw2Q6aYfOJGHmRi0hn+AnuNUPcCdu/QjX/ojTNgvbeuDC4Zx7OZcTJJwpcJxvY2V1bX1js7Rlbu/s7u1b5YOWilNJaJPEPJadACvKWUSbwIDTTiIpFgGn7WB0NfHbD1QqFkd3ME6oJ/AgYiEjGLTkW+XBfTbK/awH9AkyleS5b1WcqjOFvUzcglRQgYZv/fT6MUkFjYBwrFTXdRLwMiyBEU5zs5cqmmAywgPa1TTCgiovm76e2yda6dthLPVEYE/VvxcZFkqNRaA3BYahWvQm4n9eN4Xw0stYlKRAIzILClNuQ2xPerD7TFICfKwJJpLpX20yxBIT0G3NpQQiN03dirvYwTJpnVXdWrV2e16p3xT9lNAROkanyEUXqI6uUQM1EUGP6AW9ojfj2Xg3PozP2eqKUdwcojkYX7/Ai5uY</latexit>

k gsp

Φksp <latexit sha1_base64="M5Nt2zUr6w/P3pn1Cji2BUftLTk=">AAACE3icbVDLSsNAFJ34rPEVdSVugkVwVRKR6rLgRncV7AOaGCbTSTt0JgkzN2IJwa/wE9zqB7gTt36Aa3/E6WNhWw9cOJxzL/feE6acKXCcb2NpeWV1bb20YW5ube/sWnv7TZVkktAGSXgi2yFWlLOYNoABp+1UUixCTlvh4Grktx6oVCyJ72CYUl/gXswiRjBoKbAOvVDkXr3Pivt8UAS5B/QRcpUWRWCVnYozhr1I3CkpoynqgfXjdROSCRoD4Vipjuuk4OdYAiOcFqaXKZpiMsA92tE0xoIqPx+/UNgnWunaUSJ1xWCP1b8TORZKDUWoOwWGvpr3RuJ/XieD6NLPWZxmQGMyWRRl3IbEHuVhd5mkBPhQE0wk07fapI8lJqBTm9kSisI0dSrufAaLpHlWcauV6u15uXYzzaeEjtAxOkUuukA1dI3qqIEIekIv6BW9Gc/Gu/FhfE5al4zpzAGagfH1C0WlnyE=</latexit>

k y!anc /! yk

Conv 1x1

Spatial Channel

Conv 2x ↓

hs <latexit sha1_base64="BC6YJeSZ93WNTBuJyvj8lItMw+E=">AAAB+3icbVDLSsNAFL3xWeOr6tLNYBFclUSkuiy40V1F+4A2lMl00g6dmYSZiVBCPsGtfoA7cevHuPZHnLZZ2NYDFw7n3Mu994QJZ9p43reztr6xubVd2nF39/YPDstHxy0dp4rQJol5rDoh1pQzSZuGGU47iaJYhJy2w/Ht1G8/U6VZLJ/MJKGBwEPJIkawsdLjqK/75YpX9WZAq8QvSAUKNPrln94gJqmg0hCOte76XmKCDCvDCKe520s1TTAZ4yHtWiqxoDrIZqfm6NwqAxTFypY0aKb+nciw0HoiQtspsBnpZW8q/ud1UxPdBBmTSWqoJPNFUcqRidH0bzRgihLDJ5Zgopi9FZERVpgYm87CllDkrmtT8ZczWCWty6pfq9Yerir1+yKfEpzCGVyAD9dQhztoQBMIDOEFXuHNyZ1358P5nLeuOcXMCSzA+foFIayUmg==</latexit>

ED

z!

k gch Φkch

y!<k

DWConv Blocks x 2

<latexit sha1_base64="McoNqY2mBS2xHRsXmqRY61yaP88=">AAACAHicbVC7SgNBFL3rM8ZX1NJmMAhWYVc0ahewsbCIYB6QLGF2MpsMmZldZmaFsGzjJ9jqB9iJrX9i7Y84SbYwiQcuHM65l3vvCWLOtHHdb2dldW19Y7OwVdze2d3bLx0cNnWUKEIbJOKRagdYU84kbRhmOG3HimIRcNoKRrcTv/VElWaRfDTjmPoCDyQLGcHGSu1uINL73mXWK5XdijsFWiZeTsqQo94r/XT7EUkElYZwrHXHc2Pjp1gZRjjNit1E0xiTER7QjqUSC6r9dHpvhk6t0kdhpGxJg6bq34kUC63HIrCdApuhXvQm4n9eJzHhtZ8yGSeGSjJbFCYcmQhNnkd9pigxfGwJJorZWxEZYoWJsRHNbQlEVizaVLzFDJZJ87ziVSvVh4ty7SbPpwDHcAJn4MEV1OAO6tAAAhxe4BXenGfn3flwPmetK04+cwRzcL5+AbJ8loU=</latexit>

Bits

<latexit sha1_base64="ut/NpImId/uWDpaaFcM/Cdfkynk=">AAACAHicbVC7SgNBFL0bXzG+opY2g0GwCrsiUbuAjYVFBPOAZAmzk9lkyMzsMjMrhGUbP8FWP8BObP0Ta3/ESbKFSTxw4XDOvdx7TxBzpo3rfjuFtfWNza3idmlnd2//oHx41NJRoghtkohHqhNgTTmTtGmY4bQTK4pFwGk7GN9O/fYTVZpF8tFMYuoLPJQsZAQbK3V6gUjv+5dZv1xxq+4MaJV4OalAjka//NMbRCQRVBrCsdZdz42Nn2JlGOE0K/USTWNMxnhIu5ZKLKj209m9GTqzygCFkbIlDZqpfydSLLSeiMB2CmxGetmbiv953cSE137KZJwYKsl8UZhwZCI0fR4NmKLE8IklmChmb0VkhBUmxka0sCUQWalkU/GWM1glrYuqV6vWHi4r9Zs8nyKcwCmcgwdXUIc7aEATCHB4gVd4c56dd+fD+Zy3Fpx85hgW4Hz9ArDmloQ=</latexit>

Conv 2x ↑

y!

PConv Blocks x L5

! % & ! % & × × ×' × × ×' 64 64 64 32 32 32

Context Model

z! <latexit sha1_base64="TN6UsSTd7UFlg3rGToP0k58QPH4=">AAACCnicbVDLSsNAFJ3UV62vqEs3g0VwVRKR6rLgRncV7AOaUCaTSTt0ZhJmJpUa8gd+glv9AHfi1p9w7Y84bbOwrQcuHM65l3M5QcKo0o7zbZXW1jc2t8rblZ3dvf0D+/CoreJUYtLCMYtlN0CKMCpIS1PNSDeRBPGAkU4wupn6nTGRisbiQU8S4nM0EDSiGGkj9W3be6QhGSKdeQHPnvK8b1edmjMDXCVuQaqgQLNv/3hhjFNOhMYMKdVznUT7GZKaYkbyipcqkiA8QgPSM1QgTpSfzT7P4ZlRQhjF0ozQcKb+vcgQV2rCA7PJkR6qZW8q/uf1Uh1d+xkVSaqJwPOgKGVQx3BaAwypJFiziSEIS2p+hXiIJMLalLWQEvC8UjGtuMsdrJL2Rc2t1+r3l9XGXdFPGZyAU3AOXHAFGuAWNEELYDAGL+AVvFnP1rv1YX3OV0tWcXMMFmB9/QKPB5rq</latexit>

EE

<latexit sha1_base64="oDOeSeoIMF6/OFjGxTVilzq8PJ8=">AAACAHicbVC7SgNBFL3rM8ZX1NJmMAhWYVclahewsbCIYB6QLGF2MpsMmZldZmaFsGzjJ9jqB9iJrX9i7Y84SbYwiQcuHM65l3vvCWLOtHHdb2dldW19Y7OwVdze2d3bLx0cNnWUKEIbJOKRagdYU84kbRhmOG3HimIRcNoKRrcTv/VElWaRfDTjmPoCDyQLGcHGSu1uINL73kXWK5XdijsFWiZeTsqQo94r/XT7EUkElYZwrHXHc2Pjp1gZRjjNit1E0xiTER7QjqUSC6r9dHpvhk6t0kdhpGxJg6bq34kUC63HIrCdApuhXvQm4n9eJzHhtZ8yGSeGSjJbFCYcmQhNnkd9pigxfGwJJorZWxEZYoWJsRHNbQlEVizaVLzFDJZJ87ziVSvVh8ty7SbPpwDHcAJn4MEV1OAO6tAAAhxe4BXenGfn3flwPmetK04+cwRzcL5+Aa9QloM=</latexit>

Conv 2x ↑

<latexit sha1_base64="H4aDzmYjzxupntZSKFLDKrMN264=">AAACAHicbVC7SgNBFL3rM8ZX1NJmMAhWYTdI1C5gY2ERwTwgWcLsZDYZMjO7zMwKYdnGT7DVD7ATW//E2h9xkmxhEg9cOJxzL/feE8ScaeO6387a+sbm1nZhp7i7t39wWDo6bukoUYQ2ScQj1QmwppxJ2jTMcNqJFcUi4LQdjG+nfvuJKs0i+WgmMfUFHkoWMoKNlTq9QKT3/WrWL5XdijsDWiVeTsqQo9Ev/fQGEUkElYZwrHXXc2Pjp1gZRjjNir1E0xiTMR7SrqUSC6r9dHZvhs6tMkBhpGxJg2bq34kUC60nIrCdApuRXvam4n9eNzHhtZ8yGSeGSjJfFCYcmQhNn0cDpigxfGIJJorZWxEZYYWJsREtbAlEVizaVLzlDFZJq1rxapXaw2W5fpPnU4BTOIML8OAK6nAHDWgCAQ4v8ApvzrPz7nw4n/PWNSefOYEFOF+/rbqWgg==</latexit>

<latexit sha1_base64="/vGIawDcDkyrvdywcSZGi4Uo5DM=">AAAB+3icbVDLSsNAFL3xWeur6tLNYBFclUSkuiy40V1F+4A2lMl0kg6dmYSZiVBCPsGtfoA7cevHuPZHnLZZ2NYDFw7n3Mu99wQJZ9q47reztr6xubVd2inv7u0fHFaOjts6ThWhLRLzWHUDrClnkrYMM5x2E0WxCDjtBOPbqd95pkqzWD6ZSUJ9gSPJQkawsdJjNNCDStWtuTOgVeIVpAoFmoPKT38Yk1RQaQjHWvc8NzF+hpVhhNO83E81TTAZ44j2LJVYUO1ns1NzdG6VIQpjZUsaNFP/TmRYaD0Rge0U2Iz0sjcV//N6qQlv/IzJJDVUkvmiMOXIxGj6NxoyRYnhE0swUczeisgIK0yMTWdhSyDyctmm4i1nsEralzWvXqs/XFUb90U+JTiFM7gAD66hAXfQhBYQiOAFXuHNyZ1358P5nLeuOcXMCSzA+foFIBWUmQ==</latexit>

PConv Blocks x L3

gs

PConv Blocks x L2

<latexit sha1_base64="6xsCsFK8xePaYdV6G00pjnEipCI=">AAACCnicbVDLSsNAFJ3UV62vqEs3g0VwVRKR6rLgRncV7AOaUCaTSTt0ZhJmJtUS8gd+glv9AHfi1p9w7Y84bbOwrQcuHM65l3M5QcKo0o7zbZXW1jc2t8rblZ3dvf0D+/CoreJUYtLCMYtlN0CKMCpIS1PNSDeRBPGAkU4wupn6nTGRisbiQU8S4nM0EDSiGGkj9W3be6QhGSKdeQHPnvK8b1edmjMDXCVuQaqgQLNv/3hhjFNOhMYMKdVznUT7GZKaYkbyipcqkiA8QgPSM1QgTpSfzT7P4ZlRQhjF0ozQcKb+vcgQV2rCA7PJkR6qZW8q/uf1Uh1d+xkVSaqJwPOgKGVQx3BaAwypJFiziSEIS2p+hXiIJMLalLWQEvC8UjGtuMsdrJL2Rc2t1+r3l9XGXdFPGZyAU3AOXHAFGuAWNEELYDAGL+AVvFnP1rv1YX3OV0tWcXMMFmB9/QKL2Zro</latexit>

Conv 2x ↑

! x

Conv 2x ↑

<latexit sha1_base64="GQ56vxuo4mIh29SiKVXhu6I6Gx8=">AAACAHicbVC7SgNBFL0bXzG+opY2g0GwCrsiUbuAjYVFBPOAZAmzk9lkyMzsMjMrhGUbP8FWP8BObP0Ta3/ESbKFSTxw4XDOvdx7TxBzpo3rfjuFtfWNza3idmlnd2//oHx41NJRoghtkohHqhNgTTmTtGmY4bQTK4pFwGk7GN9O/fYTVZpF8tFMYuoLPJQsZAQbK3V6gUjv+17WL1fcqjsDWiVeTiqQo9Ev//QGEUkElYZwrHXXc2Pjp1gZRjjNSr1E0xiTMR7SrqUSC6r9dHZvhs6sMkBhpGxJg2bq34kUC60nIrCdApuRXvam4n9eNzHhtZ8yGSeGSjJfFCYcmQhNn0cDpigxfGIJJorZWxEZYYWJsREtbAlEVirZVLzlDFZJ66Lq1aq1h8tK/SbPpwgncArn4MEV1OEOGtAEAhxe4BXenGfn3flwPuetBSefOYYFOF+/rCSWgQ==</latexit>

PConv Blocks x L1

Bits

z

<latexit sha1_base64="3ULg6KyY4mkEXxobthYSZo8ImOE=">AAACEXicbZDLSsNAFIYnXmu9RV24cDNYhApSEpHqsqALlxXsBZpQJtNJO3RmEmYmQgl5Ch/BrT6AO3HrE7j2RZymWdjWHwY+/nMO58wfxIwq7Tjf1srq2vrGZmmrvL2zu7dvHxy2VZRITFo4YpHsBkgRRgVpaaoZ6caSIB4w0gnGt9N654lIRSPxqCcx8TkaChpSjLSx+vZx1Qt46vEku4A5KTrkKDvv2xWn5uSCy+AWUAGFmn37xxtEOOFEaMyQUj3XibWfIqkpZiQre4kiMcJjNCQ9gwJxovw0/0AGz4wzgGEkzRMa5u7fiRRxpSY8MJ0c6ZFarE3N/2q9RIc3fkpFnGgi8GxRmDCoIzhNAw6oJFiziQGEJTW3QjxCEmFtMpvbEvCsXDapuIsZLEP7subWa/WHq0rjrsinBE7AKagCF1yDBrgHTdACGGTgBbyCN+vZerc+rM9Z64pVzByBOVlfv3QBnOI=</latexit>

(µ, σ)

Bits

<latexit sha1_base64="mcKXprmFQIQiygeML1IuaUYjw9s=">AAACMnicbVDLSgMxFM34tr6qLt0Ei+CqzoioS0EXLivYVujUkklvbWgyMyR31BLmV/wKP8Gt7nUn3foRprULbT0QOJxzLjf3RKkUBn3/3ZuZnZtfWFxaLqysrq1vFDe3aibJNIcqT2SibyJmQIoYqihQwk2qgalIQj3qnQ/9+j1oI5L4GvspNBW7i0VHcIZOahVPwwfRhi5DG0bK9vP81vbylg0RHtGymOc5PaD/Zoolv+yPQKdJMCYlMkalVRyE7YRnCmLkkhnTCPwUm5ZpFFxCXggzAynjPXYHDUdjpsA07ejCnO45pU07iXYvRjpSf09Ypozpq8glFcOumfSG4n9eI8POadOKOM0Q3LWjRZ1MUkzosC7aFho4yr4jjGvh/kp5l2nG0ZX6Z0uk8kLBtRJMdjBNaofl4Lh8fHVUOrsY97NEdsgu2ScBOSFn5JJUSJVw8kReyCt58569D+/TG/xEZ7zxzDb5A+/rG6fkrN0=</latexit>

Factorized

! % & × × ×' 16 16 16

<latexit sha1_base64="LFf0xgl1IRbrsgpGeHqK1gZu7lE=">AAACAXicbVDLSsNAFL2prxpfVZduBovgqiQi1WXBje4q2Ic0oUymk3boTBJmJkIJWfkJbvUD3Ilbv8S1P+K0zcK2HrhwOOde7r0nSDhT2nG+rdLa+sbmVnnb3tnd2z+oHB61VZxKQlsk5rHsBlhRziLa0kxz2k0kxSLgtBOMb6Z+54lKxeLoQU8S6gs8jFjICNZGevQCkXlNxfJ+perUnBnQKnELUoUCzX7lxxvEJBU00oRjpXquk2g/w1Izwmlue6miCSZjPKQ9QyMsqPKz2cE5OjPKAIWxNBVpNFP/TmRYKDURgekUWI/UsjcV//N6qQ6v/YxFSappROaLwpQjHaPp92jAJCWaTwzBRDJzKyIjLDHRJqOFLYHIbduk4i5nsEraFzW3XqvfX1Ybd0U+ZTiBUzgHF66gAbfQhBYQEPACr/BmPVvv1of1OW8tWcXMMSzA+voF6aSXRw==</latexit>

<latexit sha1_base64="a+7hlFKQrwCffilPCtgTpccw69M=">AAAB/nicbVDLSgNBEOz1GddX1KOXwSB4Crsi0WPAi94imAckS5idzCZjZmaXmVkhLgE/wat+gDfx6q949kecJHswiQUNRVU33V1hwpk2nvftrKyurW9sFrbc7Z3dvf3iwWFDx6kitE5iHqtWiDXlTNK6YYbTVqIoFiGnzXB4PfGbj1RpFst7M0poIHBfsogRbKzU6IQiexp3iyWv7E2BlomfkxLkqHWLP51eTFJBpSEca932vcQEGVaGEU7HbifVNMFkiPu0banEguogm147RqdW6aEoVrakQVP170SGhdYjEdpOgc1AL3oT8T+vnZroKsiYTFJDJZktilKOTIwmr6MeU5QYPrIEE8XsrYgMsMLE2IDmtoRi7Lo2FX8xg2XSOC/7lXLl7qJUvc3zKcAxnMAZ+HAJVbiBGtSBwAO8wCu8Oc/Ou/PhfM5aV5x85gjm4Hz9Asc2lhs=</latexit>

Q

Conv 2x ↑

! $ % × × × &! 8 8 8

<latexit sha1_base64="rKzUGeCCcjEhg7q2nW276SD2kGc=">AAACCnicbVDLSsNAFJ3UV42vqEs3g0VwVRKR6rLgRncVbCu0oUwmk3bozCTMTCoh5A/8BLf6Ae7ErT/h2h9x2mZhWw9cOJxzL+dygoRRpV3326qsrW9sblW37Z3dvf0D5/Coo+JUYtLGMYvlY4AUYVSQtqaakcdEEsQDRrrB+GbqdydEKhqLB50lxOdoKGhEMdJGGjhO/4mGZIR03g94nhXFwKm5dXcGuEq8ktRAidbA+emHMU45ERozpFTPcxPt50hqihkp7H6qSILwGA1Jz1CBOFF+Pvu8gGdGCWEUSzNCw5n69yJHXKmMB2aTIz1Sy95U/M/rpTq69nMqklQTgedBUcqgjuG0BhhSSbBmmSEIS2p+hXiEJMLalLWQEvDCtk0r3nIHq6RzUfca9cb9Za15V/ZTBSfgFJwDD1yBJrgFLdAGGEzAC3gFb9az9W59WJ/z1YpV3hyDBVhfv41wmuk=</latexit>

Conv 2x ↑

! $ % × × × &! 4 4 4

y!

Ψ

Context

<latexit sha1_base64="McoNqY2mBS2xHRsXmqRY61yaP88=">AAACAHicbVC7SgNBFL3rM8ZX1NJmMAhWYVc0ahewsbCIYB6QLGF2MpsMmZldZmaFsGzjJ9jqB9iJrX9i7Y84SbYwiQcuHM65l3vvCWLOtHHdb2dldW19Y7OwVdze2d3bLx0cNnWUKEIbJOKRagdYU84kbRhmOG3HimIRcNoKRrcTv/VElWaRfDTjmPoCDyQLGcHGSu1uINL73mXWK5XdijsFWiZeTsqQo94r/XT7EUkElYZwrHXHc2Pjp1gZRjjNit1E0xiTER7QjqUSC6r9dHpvhk6t0kdhpGxJg6bq34kUC63HIrCdApuhXvQm4n9eJzHhtZ8yGSeGSjJbFCYcmQhNnkd9pigxfGwJJorZWxEZYoWJsRHNbQlEVizaVLzFDJZJ87ziVSvVh4ty7SbPpwDHcAJn4MEV1OAO6tAAAhxe4BXenGfn3flwPmetK04+cwRzcL5+AbJ8loU=</latexit>

<latexit sha1_base64="rKzUGeCCcjEhg7q2nW276SD2kGc=">AAACCnicbVDLSsNAFJ3UV42vqEs3g0VwVRKR6rLgRncVbCu0oUwmk3bozCTMTCoh5A/8BLf6Ae7ErT/h2h9x2mZhWw9cOJxzL+dygoRRpV3326qsrW9sblW37Z3dvf0D5/Coo+JUYtLGMYvlY4AUYVSQtqaakcdEEsQDRrrB+GbqdydEKhqLB50lxOdoKGhEMdJGGjhO/4mGZIR03g94nhXFwKm5dXcGuEq8ktRAidbA+emHMU45ERozpFTPcxPt50hqihkp7H6qSILwGA1Jz1CBOFF+Pvu8gGdGCWEUSzNCw5n69yJHXKmMB2aTIz1Sy95U/M/rpTq69nMqklQTgedBUcqgjuG0BhhSSbBmmSEIS2p+hXiEJMLalLWQEvDCtk0r3nIHq6RzUfca9cb9Za15V/ZTBSfgFJwDD1yBJrgFLdAGGEzAC3gFb9az9W59WJ/z1YpV3hyDBVhfv41wmuk=</latexit>

PConv Blocks x L5

y!

Conv 2x ↓

<latexit sha1_base64="ut/NpImId/uWDpaaFcM/Cdfkynk=">AAACAHicbVC7SgNBFL0bXzG+opY2g0GwCrsiUbuAjYVFBPOAZAmzk9lkyMzsMjMrhGUbP8FWP8BObP0Ta3/ESbKFSTxw4XDOvdx7TxBzpo3rfjuFtfWNza3idmlnd2//oHx41NJRoghtkohHqhNgTTmTtGmY4bQTK4pFwGk7GN9O/fYTVZpF8tFMYuoLPJQsZAQbK3V6gUjv+5dZv1xxq+4MaJV4OalAjka//NMbRCQRVBrCsdZdz42Nn2JlGOE0K/USTWNMxnhIu5ZKLKj209m9GTqzygCFkbIlDZqpfydSLLSeiMB2CmxGetmbiv953cSE137KZJwYKsl8UZhwZCI0fR4NmKLE8IklmChmb0VkhBUmxka0sCUQWalkU/GWM1glrYuqV6vWHi4r9Zs8nyKcwCmcgwdXUIc7aEATCHB4gVd4c56dd+fD+Zy3Fpx85hgW4Hz9ArDmloQ=</latexit>

Q

PConv Blocks x L4

Conv 2x ↓

y <latexit sha1_base64="Ki3lzXFVPbn46yY95v/wCRdDVew=">AAAB/nicbVDLSsNAFL2prxpfVZduBovgqiQi1WXBje4q2Ae0oUymk3bsTBJmJkIIAT/BrX6AO3Hrr7j2R5y2WdjWAxcO59zLvff4MWdKO863VVpb39jcKm/bO7t7+weVw6O2ihJJaItEPJJdHyvKWUhbmmlOu7GkWPicdvzJzdTvPFGpWBQ+6DSmnsCjkAWMYG2kdt8XWZoPKlWn5syAVolbkCoUaA4qP/1hRBJBQ004VqrnOrH2Miw1I5zmdj9RNMZkgke0Z2iIBVVeNrs2R2dGGaIgkqZCjWbq34kMC6VS4ZtOgfVYLXtT8T+vl+jg2stYGCeahmS+KEg40hGavo6GTFKieWoIJpKZWxEZY4mJNgEtbPFFbtsmFXc5g1XSvqi59Vr9/rLauCvyKcMJnMI5uHAFDbiFJrSAwCO8wCu8Wc/Wu/Vhfc5bS1YxcwwLsL5+AcWglho=</latexit>

EE ! $ % × × × &! 2 2 2

PConv Blocks x L4

<latexit sha1_base64="oDOeSeoIMF6/OFjGxTVilzq8PJ8=">AAACAHicbVC7SgNBFL3rM8ZX1NJmMAhWYVclahewsbCIYB6QLGF2MpsMmZldZmaFsGzjJ9jqB9iJrX9i7Y84SbYwiQcuHM65l3vvCWLOtHHdb2dldW19Y7OwVdze2d3bLx0cNnWUKEIbJOKRagdYU84kbRhmOG3HimIRcNoKRrcTv/VElWaRfDTjmPoCDyQLGcHGSu1uINL73kXWK5XdijsFWiZeTsqQo94r/XT7EUkElYZwrHXHc2Pjp1gZRjjNit1E0xiTER7QjqUSC6r9dHpvhk6t0kdhpGxJg6bq34kUC63HIrCdApuhXvQm4n9eJzHhtZ8yGSeGSjJbFCYcmQhNnkd9pigxfGwJJorZWxEZYoWJsRHNbQlEVizaVLzFDJZJ87ziVSvVh8ty7SbPpwDHcAJn4MEV1OAO6tAAAhxe4BXenGfn3flwPmetK04+cwRzcL5+Aa9QloM=</latexit>

PConv Blocks x L3

Conv 2x ↓

<latexit sha1_base64="H4aDzmYjzxupntZSKFLDKrMN264=">AAACAHicbVC7SgNBFL3rM8ZX1NJmMAhWYTdI1C5gY2ERwTwgWcLsZDYZMjO7zMwKYdnGT7DVD7ATW//E2h9xkmxhEg9cOJxzL/feE8ScaeO6387a+sbm1nZhp7i7t39wWDo6bukoUYQ2ScQj1QmwppxJ2jTMcNqJFcUi4LQdjG+nfvuJKs0i+WgmMfUFHkoWMoKNlTq9QKT3/WrWL5XdijsDWiVeTsqQo9Ev/fQGEUkElYZwrHXXc2Pjp1gZRjjNir1E0xiTMR7SrqUSC6r9dHZvhs6tMkBhpGxJg2bq34kUC60nIrCdApuRXvam4n9eNzHhtZ8yGSeGSjJfFCYcmQhNn0cDpigxfGIJJorZWxEZYYWJsREtbAlEVizaVLzlDFZJq1rxapXaw2W5fpPnU4BTOIML8OAK6nAHDWgCAQ4v8ApvzrPz7nw4n/PWNSefOYEFOF+/rbqWgg==</latexit>

PConv Blocks x L2

<latexit sha1_base64="y+zSLB6ZDYY/sCYjxtAYLUI3dvA=">AAAB/nicbVDLSgNBEOz1GddX1KOXwSB4Crsi0WPAi94imAckS5idzCZjZmaXmVkxLAE/wat+gDfx6q949kecJHswiQUNRVU33V1hwpk2nvftrKyurW9sFrbc7Z3dvf3iwWFDx6kitE5iHqtWiDXlTNK6YYbTVqIoFiGnzXB4PfGbj1RpFst7M0poIHBfsogRbKzU6IQiexp3iyWv7E2BlomfkxLkqHWLP51eTFJBpSEca932vcQEGVaGEU7HbifVNMFkiPu0banEguogm147RqdW6aEoVrakQVP170SGhdYjEdpOgc1AL3oT8T+vnZroKsiYTFJDJZktilKOTIwmr6MeU5QYPrIEE8XsrYgMsMLE2IDmtoRi7Lo2FX8xg2XSOC/7lXLl7qJUvc3zKcAxnMAZ+HAJVbiBGtSBwAO8wCu8Oc/Ou/PhfM5aV5x85gjm4Hz9AsQKlhk=</latexit>

Conv 2x ↓

Conv 2x ↓

x

(b) Context Model Architecture

ha <latexit sha1_base64="L7LHn9JSJxkZv6WWCgt3kGriiog=">AAAB+3icbVDLSsNAFL3xWeOr6tLNYBFclUSkuiy40V1F+4A2lMl00g6dmYSZiVBCPsGtfoA7cevHuPZHnLZZ2NYDFw7n3Mu994QJZ9p43reztr6xubVd2nF39/YPDstHxy0dp4rQJol5rDoh1pQzSZuGGU47iaJYhJy2w/Ht1G8/U6VZLJ/MJKGBwEPJIkawsdLjqI/75YpX9WZAq8QvSAUKNPrln94gJqmg0hCOte76XmKCDCvDCKe520s1TTAZ4yHtWiqxoDrIZqfm6NwqAxTFypY0aKb+nciw0HoiQtspsBnpZW8q/ud1UxPdBBmTSWqoJPNFUcqRidH0bzRgihLDJ5Zgopi9FZERVpgYm87CllDkrmtT8ZczWCWty6pfq9Yerir1+yKfEpzCGVyAD9dQhztoQBMIDOEFXuHNyZ1358P5nLeuOcXMCSzA+foFBTKUiA==</latexit>

<latexit sha1_base64="GQ56vxuo4mIh29SiKVXhu6I6Gx8=">AAACAHicbVC7SgNBFL0bXzG+opY2g0GwCrsiUbuAjYVFBPOAZAmzk9lkyMzsMjMrhGUbP8FWP8BObP0Ta3/ESbKFSTxw4XDOvdx7TxBzpo3rfjuFtfWNza3idmlnd2//oHx41NJRoghtkohHqhNgTTmTtGmY4bQTK4pFwGk7GN9O/fYTVZpF8tFMYuoLPJQsZAQbK3V6gUjv+17WL1fcqjsDWiVeTiqQo9Ev//QGEUkElYZwrHXXc2Pjp1gZRjjNSr1E0xiTMR7SrqUSC6r9dHZvhs6sMkBhpGxJg2bq34kUC60nIrCdApuRXvam4n9eNzHhtZ8yGSeGSjJfFCYcmQhNn0cDpigxfGIJJorZWxEZYYWJsREtbAlEVirZVLzlDFZJ66Lq1aq1h8tK/SbPpwgncArn4MEV1OEOGtAEAhxe4BXenGfn3flwPuetBSefOYYFOF+/rCSWgQ==</latexit>

ga

PConv Blocks x L1

<latexit sha1_base64="FvMftbpcMtNWxhyqV4Nzk9SgNBk=">AAAB+3icbVDLSsNAFL3xWeur6tLNYBFclUSkuiy40V1F+4A2lMl0kg6dmYSZiVBCPsGtfoA7cevHuPZHnLZZ2NYDFw7n3Mu99wQJZ9q47reztr6xubVd2inv7u0fHFaOjts6ThWhLRLzWHUDrClnkrYMM5x2E0WxCDjtBOPbqd95pkqzWD6ZSUJ9gSPJQkawsdJjNMCDStWtuTOgVeIVpAoFmoPKT38Yk1RQaQjHWvc8NzF+hpVhhNO83E81TTAZ44j2LJVYUO1ns1NzdG6VIQpjZUsaNFP/TmRYaD0Rge0U2Iz0sjcV//N6qQlv/IzJJDVUkvmiMOXIxGj6NxoyRYnhE0swUczeisgIK0yMTWdhSyDyctmm4i1nsEralzWvXqs/XFUb90U+JTiFM7gAD66hAXfQhBYQiOAFXuHNyZ1358P5nLeuOcXMCSzA+foFA5uUhw==</latexit>

PConv 3x3

ED

DWConv 3x3

Conv 1x1

Conv 1x1

LeakyReLU

LeakyReLU

Conv 1x1

Conv 1x1

(c) PConv Block

(d) DWConv Block

<latexit sha1_base64="TN6UsSTd7UFlg3rGToP0k58QPH4=">AAACCnicbVDLSsNAFJ3UV62vqEs3g0VwVRKR6rLgRncV7AOaUCaTSTt0ZhJmJpUa8gd+glv9AHfi1p9w7Y84bbOwrQcuHM65l3M5QcKo0o7zbZXW1jc2t8rblZ3dvf0D+/CoreJUYtLCMYtlN0CKMCpIS1PNSDeRBPGAkU4wupn6nTGRisbiQU8S4nM0EDSiGGkj9W3be6QhGSKdeQHPnvK8b1edmjMDXCVuQaqgQLNv/3hhjFNOhMYMKdVznUT7GZKaYkbyipcqkiA8QgPSM1QgTpSfzT7P4ZlRQhjF0ozQcKb+vcgQV2rCA7PJkR6qZW8q/uf1Uh1d+xkVSaqJwPOgKGVQx3BaAwypJFiziSEIS2p+hXiIJMLalLWQEvC8UjGtuMsdrJL2Rc2t1+r3l9XGXdFPGZyAU3AOXHAFGuAWNEELYDAGL+AVvFnP1rv1YX3OV0tWcXMMFmB9/QKPB5rq</latexit>

Figure 4: (a) EVOLVE architecture. Q, EE, and ED denote quantization, entropy encoding, and entropy decoding. Blue and red lines indicate the encoding and decoding data paths, while the purple ones are shared by both paths. Dashed arrows denote probability estimation. (b) Context model k fuses hyperprior Ψ, channel-wise context from previously decoded slices ŷ<k , and spatial context from decoded anchor positions ŷanc through an aggregation network to predict distribution parameters (µ, σ) for probability modeling. (c) and (d) Detailed structures of PConv and DWConv blocks.

distributions, and use the AdamW optimizer [40] for its robustness to scale variation and improved generalization through decoupled weight decay. In addition, common data augmentation techniques described in [35], such as random cropping, flipping, and rotation, are applied to enhance robustness to spatial variability and improve generalization across diverse volumetric structures. 3.2.2

Macro Design

We denote the input volume as x ∈ RH×W ×D . The encoder of the vanilla AE performs a non-linear analysis transform ga (·) to produce a latent representation y. The decoder then applies a non-linear synthesis transform gs (·) to generate the reconstructed volume x̂. Although the latent representation y is typically more compact than the original volume x, it still contains substantial redundancy that can be further exploited for compression. Motivated by this, we introduce a series of macro-level design choices to improve compression performance. Factorized entropy model. The most straightforward way to compress the latent representation y is quantization, which can be expressed as ŷ = Q(y), where Q(·) denotes the quantization operator. If a probability model pŷ (ŷ) is given, entropy coding techniques, such as arithmetic coding [53], can then be used to encode the quantized codes into a compact bitstream losslessly. In practice, under a factorized entropy model, pŷ (ŷ) is represented by learnable channel-wise Gaussian parameters (µ, σ) as illustrated in Figure 3(a) and optimized during training. Since arithmetic coding is near-optimal, the entropy of ŷ provides an accurate estimate of the coding rate. Consequently, we can formulate a rate-distortion loss LRD = R(ŷ) + λ D(x, x̂),

(1)

where R(ŷ) = E[− log2 pŷ (ŷ)] is the rate term that indicates the average number of bits used when saving ŷ as bitstreams. D(·) is the distortion term that is computed using the MSE loss between the GT and reconstructed data. λ is a Lagrange multiplier, a hyperparameter that controls the tradeoff between reconstruction quality and CR, and is set to 300 in our roadmap experiment. Note that the quantization operation is not differentiable. Therefore, we  approximate quantized latent ŷ by adding a uniform noise U − 12 , 12 to the extracted y during training. As shown in Figure 2, the factorized entropy model improves the CR, but leads to a degradation in reconstruction quality due to the introduction of the additional rate term R(ŷ) in Equation 1. Hyperprior entropy model. In a factorized entropy model, latent elements within the same channel are assumed to follow a shared distribution across spatial positions, ignoring spatially varying statistics. In contrast, the hyperprior entropy model [4] introduces additional side information to more accurately estimate the probability of latent elements at different spatial locations, as shown in Figure 3(b). Prior work, such as IDLat [57], falls into this category by adopting a hyperprior entropy model to improve CR. In particular, as illustrated in Figure 4(a), the hyperprior entropy model introduces an additional hyper-latent z extracted from y via a hyper-analysis transform ha . The quantized hyper-latent ẑ is entropy coded and transmitted as side information. By performing a hyper-synthesis transform hs , the decoded hyperprior Ψ is used to condition the probability model of ŷ, resulting in spatially varying probability modeling. Similar to ŷ, the quantized hyperlatent ẑ is also entropy coded and saved as bitstreams after encoding. Accordingly, an additional rate term R(ẑ) is introduced into Equation 1 to optimize the bit cost associated with this part. The hyperprior entropy model improves the CR from 71.94× of the factorized entropy model

to 644.0×, with a slight improvement in reconstruction quality. GDN → residual block. We replace the GDN [3] functions used in many AE-based compressors, such as AE-SZ [39] and IDLat [57], with residual blocks [27]. GDN was initially designed to provide local normalization together with a point-wise nonlinearity. However, we empirically observe that with a large number of model parameters, incorporating multiple GDNs slows convergence, whereas residual blocks with skip connections facilitate stronger gradient flow and greater nonlinear expressiveness during training. Therefore, we replace GDN layers with residual blocks in our architecture. This substitution leads to a marked improvement in reconstruction quality and CR, though computational cost increases due to the added residual layers. Context model. IDLat [57] adopted the hyperprior entropy model to achieve better compression performance than the factorized entropy model by enabling spatially varying probability modeling. However, the hyperprior model still assumes independence among latent elements given the side information and thus fails to fully exploit local dependencies between neighboring latent variables. In volumetric data, voxel values at nearby locations are often very close and highly correlated, indicating substantial local redundancy that can be removed to achieve more effective compression. To respond, we propose to design a context model [11,26,45,58,67,75] that explicitly captures such local dependencies in the latent space. As illustrated in Figure 3(c), we design our context model with two components: 3D checkerboard spatial context and channel context. The 3D checkerboard spatial context partitions the latent grid into anchor and non-anchor positions and predicts the probability distributions of nonanchor elements conditioned on already decoded anchor elements. This design enables local spatial dependency modeling while maintaining efficient parallel decoding. In addition to spatial dependencies, the context model also captures inter-channel correlations through a channelcontext architecture that slices the latent into different channel groups and models them progressively, conditioning each group on previously decoded ones, as illustrated in Figure 3(d). By jointly leveraging masked spatial and channel context, the proposed architecture enables more accurate probability estimation by conditioning on neighboring decoded information, thereby achieving better compression performance. The architecture of the context model is illustrated in Figure 4(b). When decoding the k-th channel group, the context module receives three types of conditioning information: the hyperprior Ψ, previously decoded channel groups ŷ <k , and spatial anchor elements from the current channel group. The spatial anchor elements are either a zero vector for decoding anchor positions or the decoded anchors k for decoding non-anchor positions. The spatial context feature ŷanc k Φksp is extracted by gsp , which is a single-layer masked convolution. k The channel context feature Φkch is extracted by gch , which consists of two depthwise convolution (DWConv) blocks followed by a 1×1 convolution. Since the dimensionality of ŷ <k varies with k, separate k k context transforms gsp and gch are employed for different channel groups. We adopt DWConv blocks to reduce computational cost, while the final 1×1 convolution enables effective channel-level feature fusion. Finally, the hyperprior, spatial context, and channel context features are then concatenated and fed into a lightweight aggregation network to estimate the Gaussian distribution parameters (µ, σ) for entropy coding. Our context model plays a critical role in improving compression performance and distinguishes EVOLVE from prior AE-based volume compression methods [39, 57]. While incorporating context modeling inevitably increases encoding and decoding latency due to its autoregressive nature, this design enables substantial improvements in CR while preserving reconstruction quality (see Figure 2), making it a key component of our framework. 3.2.3

Micro Design

Non-uniform channel allocation. Prior work [57] typically assigns the same number of channels to all network stages (i.e., C1 = C2 = C3 in Figure 4(a)). In contrast, we observe that allocating more channels to deeper layers—where spatial resolution is progressively reduced—significantly improves reconstruction quality. Increasing channel width in deeper layers enhances the network’s capacity to model

complex, high-level features in a more compact representation space, leading to more faithful reconstructions. To control computational cost, we reduce the channel width in shallow layers, where spatial resolution is higher, yielding larger savings in computation overhead. As shown in Figure 2, this design improves both CR and reconstruction quality. PConv in residual blocks. We replace standard convolutions with advanced partial convolutions (PConv) [10] in the original residual blocks. PConv applies a parametric filter only to a few input channels, leaving the rest untouched. This substitution significantly reduces computational complexity without sacrificing much of the representation capacity, leading to faster encoding and decoding. Non-uniform slicing in latent channels. In the channel-context architecture, a common practice [45] is to uniformly split latent channels into multiple groups. However, transform-based compression algorithms, such as TTHRESH [6], demonstrate that coefficients derived from high-dimensional data are not equally important and achieve competitive compression by truncating the long tail of minor coefficients. Inspired by this, we split ŷ into uneven groups along the channel dimension. Since a network with comparable capacity models each group, this allocation implies that smaller groups benefit from denser parametric representation. This treatment effectively encourages the framework to learn fine-grained dependencies among critical components while efficiently handling redundant components with coarse-grained modeling. Consequently, it also allows us to minimize the total number of channel groups required without compromising reconstruction quality, thereby improving both compression efficiency and processing speed. Attention block. Attention mechanisms [71] have demonstrated strong performance in large-scale models, and we also explore incorporating attention blocks into our architecture. While attention provides a modest improvement in reconstruction quality, it introduces a disproportionate increase in computational cost. As the marginal quality gain does not justify the added complexity, this modification is not adopted in the final model, as indicated by the hatched bar in Figure 2. 3.2.4 Block Size All preceding experiments use a 64×64×64 block size for training. We conclude the roadmap by examining the effect of block size on compression performance. Smaller blocks require more patches to cover the full volume, leading to redundant boundary regions and limited receptive fields, which degrade both CR and reconstruction quality. Increasing the block size to 128×128×128 allows the model to capture longer-range spatial dependencies within a single forward pass. This final adjustment yields the complete EVOLVE model, which achieves the best performance across all metrics. However, increasing the block size substantially elevates both training and inference memory consumption due to the cubic growth of feature map resolution and intermediate activations. Given our current computational resources, we restrict the maximum block size to 128×128×128, providing a favorable tradeoff between compression performance and training cost. 3.3 Variable-Rate Encoding A key limitation of existing deep-learning-based volume compressors [42, 57, 81] is that each trained model is tied to a single fixed CR. Supporting different rate-distortion tradeoffs, therefore, requires training, storing, and switching among multiple models, which is cumbersome in practical scientific workflows. To address this limitation, EVOLVE incorporates a learnable gain mechanism that enables a single model to span a continuous range of CRs during inference, without retraining or maintaining multiple models. Gain-based quantization regulation. The core idea is to modulate the quantization bin size via a learnable gain vector γ = [γ0 , γ1 , . . . , γA−1 ], where each γi is initialized with a substantially different magnitude to cover a wide range of CRs before optimization, and A denotes the number of discrete quality levels. During compression at quality level i, the latent representation y is scaled by the corresponding gain γi prior to quantization and rescaled by its inverse afterward ŷ = Q(y · γi )/γi ,

(2)

where Q(·) denotes the quantization operator. A large γi amplifies the latent values, equivalent to a smaller quantization bin size, resulting in

a lower CR and better preservation of fine details. Conversely, a small γi increases the effective bin size, leading to coarser quantization and higher CR. Following [14], we define a set of rate-distortion tradeoff parameters λ = [λ0 , λ1 , . . . , λA−1 ] in one-to-one correspondence with the gain vector γ during optimization. Before training starts, the p gain values in γ are initialized as λ/λ0 following [69]. After the model is sufficiently trained across the discrete quality levels, continuous variable-rate encoding can be achieved at inference time by interpolating between adjacent γ values, enabling a single EVOLVE model to support a continuous range of CRs. Three-stage training strategy. Training a variable-rate model endto-end is challenging because the gain values interact with all other network weights, potentially destabilizing the optimization. To ensure stable convergence across the full range of quality levels, we adopt a three-stage training strategy. (1) In the first stage, the model is trained at a single, fixed highest quality level while the gain parameters are kept frozen. This stage allows the encoder, decoder, and entropy model to converge to a strong baseline without being affected by rate variation. (2) In the second stage, the gain parameters are unfrozen. The model is trained jointly across all A quality levels: at each iteration, a quality level i is selected by deterministically cycling through the levels, and the corresponding λi replaces λ in Equation 1, with the network weights and gain parameters optimized simultaneously. (3) In the third stage, the model is further finetuned by replacing the uniform-noise quantization approximation with the straight-through estimator (STE) [84], ensuring that the learned gain parameters and network weights are well adapted to the discrete rounding operation used during inference. Table 1: Testing datasets used for compression performance comparison. dataset asteroids asteroids-T combustion (MF) gas half-cylinder (VLM, 6,400) ionization (H+) ionization-T (H+) isotropic magnetic Nyx-E

volume resolution (x × y × z) 1,000×1,000×1,000 500×500×500 480×720×120 512×512×512 640×240×80 600×248×248 600×248×248 512×512×512 512×512×512 512×512×512

4

R ESULTS AND D ISCUSSION

4.1

Implementation Details

# timesteps or ensembles 1 220 1 1 1 1 200 1 1 206

size 3.7 GB 102.5 GB 158 MB 512 MB 47 MB 141 MB 27.5 GB 512 MB 512 MB 103.0 GB

Training details. We trained EVOLVE on the 6,376 selected volumes listed in Table 1. For model architecture, we set [C1 , C2 , C3 ] to [96, 192, 384] and [L1 , L2 , L3 , L4 , L5 ] to [2, 2, 4, 1, 3]. We set the number of latent channels M to 320 and the number of hyperlatent channels N to 256. For channel-context slicing, we split the latent channels unevenly into five groups with channel numbers [16, 16, 32, 64, 192]. The resulting EVOLVE model contains 85.6M parameters (≈327 MB in FP32). Unlike INRs, where the network itself is the compressed representation of a single volume, the same shared EVOLVE model can compress both seen and unseen in-domain volumes without pervolume optimization, and the reported CRs account for the per-volume bitstream. We used the AdamW optimizer [40] with an initial learning rate of 10−4 and a weight decay of 0.01 for the encoder and decoder parameters. A separate optimizer with a learning rate of 10−3 was used for other parameters, including the entropy model, as they exhibit different gradient scales and require a higher learning rate for stable convergence. We optimized our model for 1,600 epochs in total, with 800, 400, and 400 epochs for the three stages, respectively, and applied a linear warmup during the first 20 epochs at the beginning of each stage. In the second stage, we set the number of discrete quality levels A to 8 and λ to [100, 200, 400, 800, 1,600, 3,200, 6,400, 12,800]. The training was conducted on an NVIDIA H200 GPU using a batch size of 24. For data augmentation, we randomly sampled 128×128×128 blocks from each dataset. If any dimension was smaller than 128, it was resized to allow block extraction. We additionally applied random axis-wise flipping and 90◦ rotations, each with 50% probability.

Baselines. For conventional lossy compressors, we chose ZFP [38], TTHRESH [6], and SZ3 [37], which represent state-of-the-art predictionbased and transform-based approaches. Since compressors are sensitive to their operating modes, we document the exact settings. TTHRESH uses target-PSNR mode (-p), and SZ3 uses an absolute error bound (-M ABS) derived from the target PSNR as ϵ = r · 10−PSNR/20 , with r being the data range. ZFP is run in fixed-accuracy mode (-a), which adapts the per-block bit budget to a global error tolerance and generally attains higher CRs at a given quality than the fixed-rate mode (-r); even though it forgoes the fixed-size blocks and random access that fixed-rate mode provides. All operating points are selected so that each compressor meets the high-fidelity threshold (PSNR≥40 dB) for each dataset. For deep-learning-based compressors, we considered fully-connected INRs (SIREN [59], NeurComp [42], and ECNR [64]) that use MLPs to map coordinates to scalar values, grid-based INRs (Instant-NGP [46], fVSRN [77], and AMGSRN++ [80]) that combine learnable feature grids with lightweight decoders for faster training and inference, and an AEbased method (IDLat [57]) that learns lightweight latent representations through convolutional encoding. Note that we did not compare with time-varying INR compression methods (e.g., KD-INR [22] and MoEINR [19]), as EVOLVE focuses on static volume compression. For IDLat, we followed the original implementation and trained it on the same curated database as EVOLVE (see Table 1), using identical latent and hyperlatent dimensions for a fair comparison. For INR-based methods, we control the CR by adjusting their network hyperparameters, such as the number of hidden neurons for fully connected INRs and the feature grid resolution for grid-based INRs. Testing datasets and metrics. All compression tests were conducted on a workstation with a single NVIDIA RTX 4090 GPU to ensure a fair comparison. We used the datasets listed in Table 1 for testing the compression performance of different methods. To prevent data leakage, the volumes in Table 1 have no overlap with the 9,921 volumes in Table 1, which are from unseen variables (combustion, half-cylinder, ionization) or new datasets (asteroids, gas, isotropic, magnetic). We assessed reconstruction quality using three complementary metrics: data-level peak signal-to-noise ratio (PSNR), image-level learned perceptual image patch similarity (LPIPS) [86], and surface-level Chamfer distance (CD) [7]. LPIPS quantifies perceptual differences by computing distances between deep features extracted from rendered images using AlexNet. CD evaluates geometric fidelity by measuring the average nearest-neighbor distance between isosurfaces extracted from decompressed and GT volumes using the L2 norm. The isovalues used for CD computation and isosurface rendering were manually selected per dataset to capture salient structures. Additional results with timevarying (asteroids-T, ionization-T) and ensemble (Nyx-E) datasets are presented in Appendices E and F. Table 2: PSNR (dB), LPIPS, and CD values, as well as encoding time (ET, in sec), decoding time (DT, in sec), and CR of EVOLVE and other conventional lossy compressors. v is the chosen isovalue used to compute CD. dataset asteroids (v = 0.2) combustion (MF) (v = 0.5) ionization (H+) (v = 0.1) isotropic (v = 0.7)

method ZFP TTHRESH SZ3 EVOLVE ZFP TTHRESH SZ3 EVOLVE ZFP TTHRESH SZ3 EVOLVE ZFP TTHRESH SZ3 EVOLVE

PSNR↑ 40.00 43.44 40.85 47.16 41.92 42.98 41.10 49.31 41.44 40.11 41.43 47.58 43.79 43.51 40.03 45.18

LPIPS↓ 0.330 0.093 0.096 0.073 0.436 0.352 0.338 0.146 0.372 0.379 0.300 0.203 0.332 0.220 0.301 0.151

CD↓ 0.68 0.46 0.36 0.33 0.68 0.73 0.92 0.44 0.58 0.92 0.37 0.36 0.49 0.50 0.79 0.42

ET↓ 2.08 159.59 10.09 47.71 0.25 5.50 0.57 2.25 0.24 4.15 0.40 1.96 1.27 17.63 2.43 6.63

DT↓ 3.18 31.91 4.24 48.33 0.20 1.07 0.26 2.30 0.12 0.86 0.20 1.60 0.55 3.89 1.55 6.14

CR↑ 1,164 2,401 2,482 9,803 129 2,714 2,948 6,047 107 6,334 1,051 7,843 64 870 985 2,846

4.2 Comparison with Conventional Lossy Compressors We compare EVOLVE with three representative conventional lossy compressors: ZFP, TTHRESH, and SZ3. All methods are evaluated

CR / PSNR

6,047× / 49.31 dB

2,948× / 41.10 dB

2,714× / 42.98 dB

129× / 41.92 dB

CR / PSNR

7,843× / 47.58 dB

1,051× / 41.43 dB

6,334× / 40.11 dB

107× / 41.44 dB

GT

EVOLVE

SZ3

TTHRESH

ZFP

Figure 5: Comparison of volume rendering results between EVOLVE and conventional lossy compressors. Top and Bottom: combustion (MF) and ionization (H+). The difference image shows noticeable pixel differences relative to the GT in the CIELUV color space. CR / PSNR

9,803× / 47.16 dB

2,482× / 40.85 dB

2,401× / 43.44 dB

1,164× / 40.00 dB

CR / PSNR

2,846× / 45.18 dB

985× / 40.03 dB

870× / 43.51 dB

64× / 43.79 dB

GT

EVOLVE

SZ3

TTHRESH

ZFP

Figure 6: Comparison of isosurface rendering results between EVOLVE and conventional lossy compressors. Top and bottom: asteroids and isotropic. The chosen isovalues are 0.2 and 0.7, respectively.

Table 3: PSNR (dB), LPIPS, and CD values, as well as ET (sec), DT (sec), and CR of EVOLVE and other deep-learning-based compressors. dataset

gas (v = 0.25)

half-cylinder (VLM, 6,400) (v = 0.6)

magnetic (v = 0.05)

method SIREN NeurComp ECNR Instant-NGP fV-SRN AMGSRN++ IDLat EVOLVE SIREN NeurComp ECNR Instant-NGP fV-SRN AMGSRN++ IDLat EVOLVE SIREN NeurComp ECNR Instant-NGP fV-SRN AMGSRN++ IDLat EVOLVE

PSNR↑ 41.37 41.56 43.35 42.22 40.24 40.78 41.28 45.62 40.44 42.29 41.50 41.35 42.23 40.90 42.17 44.91 42.09 43.48 43.29 43.91 42.16 44.23 38.84 45.20

LPIPS↓ 0.134 0.133 0.148 0.176 0.167 0.180 0.115 0.113 0.150 0.118 0.116 0.199 0.165 0.191 0.119 0.106 0.367 0.339 0.342 0.326 0.351 0.312 0.332 0.306

CD↓ 0.75 0.74 0.77 0.98 0.89 1.23 0.73 0.71 1.26 0.73 1.35 0.88 0.73 0.91 0.65 0.59 2.09 1.50 1.31 1.37 1.92 1.23 1.48 1.04

ET↓ 3,488 3,239 2,543 50.32 31.45 81.31 10.41 6.05 234.2 235.5 151.6 48.04 29.94 76.33 1.73 1.61 3,457 3,273 1,566 47.04 29.72 75.69 10.35 5.95

DT↓ 3.70 5.07 5.34 0.08 0.26 0.12 20.60 6.05 0.27 0.52 0.42 0.01 0.03 0.01 3.17 1.24 3.65 3.97 5.73 0.08 0.24 0.11 20.76 5.88

CR↑ 1,979 5,954 1,714 1,839 936 1,575 1,808 10,517 633 1,044 446 168 166 371 1,114 2,596 5,256 6,649 3,047 1,839 1,813 1,575 3,748 10,774

in a high-fidelity reconstruction setting (i.e., PSNR≥40 dB) to ensure a fair comparison of compression efficiency without sacrificing reconstruction quality. Table 2 summarizes the averaged quantitative results across all datasets. In terms of CR, EVOLVE consistently achieves substantially higher CR than conventional compressors on unseen volumes, while maintaining superior reconstruction quality as measured by PSNR, LPIPS, and CD. In particular, on the asteroids dataset, EVOLVE reaches a CR of 9,803×, which is 3.95× higher than the best conventional baseline, while simultaneously improving PSNR, LPIPS, and CD. Regarding runtime performance, EVOLVE achieves encoding speeds comparable to, and in some cases faster than, those of conventional compressors. The decoding stage is slower than the second-slowest baseline (TTHRESH) primarily due to additional operations, such as block merging. See Section 4.4 for comparisons at varying CRs. Figure 5 presents volume rendering comparisons on the combustion and ionization datasets. Although all compressors preserve the overall structure, the reconstructions from conventional methods exhibit noticeable artifacts and noise, particularly in the combustion zoom-in region. In contrast, EVOLVE yields high-fidelity renderings that better preserve fine-grained spatial details. Similarly, for isosurface renderings shown in Figure 6, ZFP and SZ3 exhibit noticeable roughness and jagged artifacts on the asteroids dataset, and TTHRESH introduces subtle distortions, whereas EVOLVE produces smoother, more geometrically consistent isosurfaces. 4.3

Comparison with Deep-Learning-Based Compressors

We compare EVOLVE with three categories of existing deep-learningbased compressors: fully connected INRs (SIREN, NeurComp, ECNR), grid-based INRs (Instant-NGP, fV-SRN, AMGSRN++), and the AEbased compressor IDLat. Following the same high-fidelity setting (PSNR≥40 dB), we tune the network capacity for INR-based methods to meet this quality threshold. For IDLat, since its CR is fixed once the model is trained on the same database as EVOLVE, we report its results directly without quality-based adjustment. Table 3 summarizes the quantitative comparisons across three test datasets. EVOLVE consistently achieves the highest CR while maintaining superior PSNR, LPIPS, and CD on all datasets. Moreover, thanks to its generalizability, EVOLVE performs compression via a feed-forward inference process without per-instance retraining, resulting in orders-of-magnitude faster encoding speeds than INR-based methods. Compared with IDLat, the improved efficiency and compression quality of EVOLVE stem primarily from the proposed context model and computationally efficient network designs, which enable faster, higher-quality compression. Specifically, on the magnetic dataset, EVOLVE achieves a CR of 10,774× at 45.20 dB, surpassing the best INR method (AMGSRN++, 1,575× at 44.23 dB) by nearly 7× in CR while also improving quality.

In terms of visual quality, Figure 7 compares the volume rendering results between EVOLVE and fully connected INRs. Under comparable visual fidelity, EVOLVE achieves nearly 2× higher CR than NeurComp while preserving sharper structural details. For more challenging datasets such as magnetic (Figure 8), both grid-based INRs and EVOLVE inevitably lose some high-frequency details under high CR. However, EVOLVE achieves a CR that is 5.9× higher than that of fV-SRN, while avoiding the excessive over-smoothing artifacts observed in fV-SRN. Against IDLat, EVOLVE achieves much higher CR while maintaining comparable visual quality (Figure 9). The scale maps further reveal the entropy model’s spatially adaptive behavior. Compared with IDLat, EVOLVE allocates fewer bits to the relatively homogeneous upper regions of the gas volume, while assigning a relatively high bitrate to structurally complex areas near the bottom. This indicates that the proposed context modeling can exploit spatial redundancy, enabling more efficient bit allocation and leading to high CR. 4.4 Performance Evaluation under Various CRs The comparisons in Sections 4.2 and 4.3 evaluate all methods at a single operating point under the high-fidelity setting (PSNR≥40 dB). Since scientific workflows often require flexible tradeoffs between quality and compression, we further compare how different methods perform across a range of CRs on the combustion (MF) and asteroids datasets. For EVOLVE, the variable-rate encoding mechanism (Section 3.3) enables a single trained model to produce a densely sampled rate-distortion curve by interpolating between learned gain values, without retraining or model switching. For conventional compressors (SZ3, TTHRESH), different CRs are obtained by varying their quality or error-bound parameters. For INR-based methods, each operating point requires a separately trained model with adjusted network capacity. ZFP is excluded from this comparison as its achievable CRs (refer to Table 2) fall well below the evaluated range, and IDLat is excluded because its fixed latent space dimensionality determines its CR. Figure 10 presents the resulting rate-distortion curves measured in both PSNR and LPIPS. First, at comparable CRs, EVOLVE consistently achieves the highest PSNR among all evaluated methods on both datasets, demonstrating that the learned representations generalize well to unseen data across a wide range of rate-distortion operating points. Second, a single EVOLVE model spans a wide, continuous CR range via the learnable gain mechanism. The densely and smoothly sampled points along the EVOLVE curves confirm that the gain interpolation provides stable, fine-grained rate control without abrupt quality jumps. Third, while conventional compressors can also operate across multiple CRs by tuning their parameters, they exhibit steeper quality degradation as the CR increases. For instance, on the asteroids dataset, SZ3 achieves high PSNRs at low CRs but drops sharply beyond 3,000×, whereas EVOLVE maintains competitive reconstruction quality even at CRs exceeding 10,000×. Similarly, INR-based methods are generally confined to narrower CR ranges, and the LPIPS curves further corroborate these findings, showing that EVOLVE preserves superior perceptual quality across the entire evaluated CR range. 4.5 Deployment Considerations Volume compression methods differ in where decompression is placed in the visualization pipeline. For some methods, decompression can be performed on demand at each frame during rendering (e.g., waveletbased GPU compression [70], HVQ [56], dictionary-based sparse representations [15, 17], GPU-accelerated time-varying volume exploration [43], and grid-based INRs [46, 77, 79, 81]). In contrast, other methods require decompressing the whole volume before rendering. EVOLVE belongs to the latter category. Its decoding latency, typically a few seconds per volume on a consumer GPU, is incurred only once per volume, not per frame. After decoding, all interactive operations, including transfer function editing, isovalue changes, and spatial navigation, proceed at native rendering rates. For time-varying data exploration, each timestep must likewise be decoded before visualization; future timesteps can be prefetched and decoded asynchronously while the current timestep is being inspected, hiding part of the switching latency. EVOLVE is therefore best suited when storage and transmission bandwidth savings outweigh the need for immediate random access to

CR / PSNR

2,596× / 44.91 dB

446× / 41.50 dB

1,044× / 42.29 dB

633× / 40.44 dB

GT

EVOLVE

ECNR

NeurComp

SIREN

Figure 7: Comparison of volume rendering results between EVOLVE and fully-connected INRs using the half-cylinder (VLM, 6,400) dataset. CR / PSNR

10,774× / 45.20 dB

1,575× / 44.23 dB

1,813× / 42.16 dB

1,839× / 43.91 dB

GT

EVOLVE

AMGSRN++

fV-SRN

Instant-NGP

Figure 8: Comparison of isosurface rendering results between EVOLVE and grid-based INRs using the magnetic dataset. The chosen isovalue is 0.05. CR / PSNR

10,517× / 45.62 dB

1,808× / 41.28 dB

50.0

55

AMGSRN++ fV-SRN Instant-NGP

ECNR SIREN NeurComp

10,000

15,000

20,000

10,000

15,000

20,000

50

45.0

PSNR (dB)

PSNR (dB)

47.5

EVOLVE SZ3 TTHRESH

42.5 40.0

45 40 35

37.5 35.0

30

32.5

GT

EVOLVE

0

IDLat

5,000

10,000

15,000

20,000

Compression Ratio

25,000

0

0.25

5,000

Compression Ratio

0.45 0.40

0.20

0.35

LPIPS

LPIPS

0.30 0.15

0.25 0.20

0.10

EVOLVE scale

IDLat scale

Figure 9: Top: Comparison of volume rendering results between EVOLVE and IDLat using the gas dataset. Bottom: Visualization of the predicted Gaussian scale parameter from the entropy model. Blue-to-red indicates increasing magnitude, corresponding to lower-to-higher bitrate allocation.

compressed timesteps. For transient rendering, strict timing constraints may limit the affordable decoding complexity and achievable CR; for offline storage, the autoregressive context model can trade decoding latency (refer to Tables 3) for substantially higher CRs. Deployment also incurs a one-time overhead from the shared model itself (85.6M parameters), which is amortized over many encoded and decoded volumes. 5

C ONCLUSIONS AND F UTURE W ORK

We have presented EVOLVE, an AE-based framework toward more generalizable learned volume compression for scientific data, shifting the focus from domain-specific optimization to unified representation learning. Its efficient, scalable design balances compression efficiency and reconstruction fidelity, and supports continuous variable-rate encoding within a single model, offering a practical path toward generalizable data-driven volumetric compression.

0.15 0.10

0.05

0.05 0

5,000

10,000

15,000

20,000

Compression Ratio

asteroids

25,000

0

5,000

Compression Ratio

combustion (MF)

Figure 10: Rate-distortion curves of different compression approaches.

Several limitations remain. First, while our approach differs from prior works [21, 28, 64] by training a generalizable model on a largescale database, the domains covered remain limited. In particular, our database contains only simulation data and no medical or microscopy volumes, whose noise characteristics and near-lossless requirements differ substantially; applying EVOLVE to such domains would require retraining with representative data. Second, training EVOLVE is computationally intensive, requiring a high-end GPU and a three-stage optimization schedule. Once trained on representative data from a domain, however, the model compresses both seen and unseen indomain volumes without per-volume optimization, and inference runs on consumer-level GPUs. Finally, the current framework treats each volume independently, and leveraging temporal redundancy in timevarying simulations could further improve CRs.

A

T RAINING DATABASE C OMPOSITION

Table 1 reports the full composition of the training database curated as described in Section 3.1. As exceptions, FIT and rotstrat are actually single static volumes with a resolution of 2,048×2,048×2,048, and we partition each volume into 512 subvolumes. Table 1: The composition of the training database. Dataset diversity is measured as the 95th percentile of the NNS score using pHash. dataset argon-bubble combustion earthquake explosion FIT five-jet gravity half-cylinder hurricane ionization mantle MHD neutron-star radiative-layer Rayleigh-Taylor rotstrat solar-plume supercurrent supernova Tangaroa vortex total

volume resolution (x × y × z) 640×256×256 480×720×120 256×256×96 128×128×128 256×256×256 128×128×128 128×128×128 640×240×80 500×500×100 600×248×248 360×201×180 256×256×256 192×128×66 256×128×128 128×128×128 256×256×256 128×128×512 256×128×32 432×432×432 300×180×120 128×128×128 —

# total volumes 165 400 599 174 512 2,000 833 600 192 600 251 297 400 600 100 512 28 908 60 600 90 9,921

# selected volumes 55 400 374 65 512 105 833 600 133 513 90 297 400 600 56 512 28 60 60 593 90 6,376

original NNS↓ 0.93 0.69 0.90 0.98 0.50 0.96 0.66 0.77 0.90 0.88 0.93 0.68 0.54 0.53 0.99 0.51 0.76 0.99 0.56 0.85 0.70 —

final NNS↓ 0.85 0.69 0.85 0.84 0.50 0.84 0.66 0.77 0.85 0.85 0.84 0.68 0.54 0.53 0.85 0.51 0.76 0.85 0.56 0.85 0.70 —

Nine datasets are multivariate or ensemble, each containing multiple physical variables or generated by different simulation parameters. Specifically, combustion [24] includes CHI, HR, VORTS, and YOH; explosion [30] includes RHO, P, and T; half-cylinder [55] provides VLM and VTM at three Reynolds numbers (160, 320, 640); hurricane [31] includes CLOUD, P, VAPOR, and WSMAG; ionization [78] includes GT, H, H2 , He, He+, and PD; MHD [9] comprises RHO, MFM, and VLM; and Tangaroa [51] includes ACC, DIV, VLM, and VTM. In addition, gravity [29] and radiative-layer [16] contribute 17 and six distinct parameter combinations, respectively. Each variable name denotes a physical quantity of the corresponding simulation. For combustion [24], CHI, HR, VORTS, and YOH refer to the scalar dissipation rate, heat release, vorticity magnitude, and OH mass fraction. For hurricane [31], a simulation of Hurricane Isabel, CLOUD, P, VAPOR, and WSMAG denote the total cloud moisture mixing ratio, atmospheric pressure, water vapor mixing ratio, and wind speed magnitude (computed from the U/V/W wind components), respectively. For ionization [78], which simulates radiation-driven ionization fronts in primordial gas with a nine-species chemistry network, GT and PD denote the gas temperature and particle density. In contrast, H, H2 , He, He+, and H+ denote the abundances of neutral atomic hydrogen, molecular hydrogen, neutral helium, singly ionized helium, and ionized hydrogen, respectively. The remaining abbreviations follow common conventions (e.g., P for pressure, T for temperature, RHO for density, VLM/VTM for velocity/vorticity magnitude). For detailed information on these datasets, we refer readers to the corresponding entries in the Open SciVis Datasets [34], IEEE SciVis Contest archives [31], ETH Zürich visualization datasets [13], and Well [48]. For ionization, we remove the early and late timesteps that are largely empty and retain 100 content-rich timesteps for each variable. B

C OMPRESSION AND D ECOMPRESSION API

EVOLVE provides a compression command and a decompression command. Tables 2 and 3 summarize the key parameters. Compression. To compress a volume, the user specifies the input data directory, the pretrained model checkpoint, and the desired quality level. EVOLVE uses a continuous gain factor (--factor, float in [0.0, 10.0]) for fine-grained bitrate control: larger values yield higher quality at lower CR. The volume is divided into overlapping blocks

with configurable size and stride, enabling memory-efficient processing of arbitrary-resolution volumes. Decompression. For decompression, only the compressed bitstream directory, the same model checkpoint, and the matching model configuration are needed. The volume resolution, block layout, and normalization ranges are automatically recovered from the stored meta info (refer to Appendix C), so no additional data-specific parameters are required. Table 2: Key parameters of the compression command.

parameter --data_path --checkpoint

type str str

--block_size --stride --factor --output_dir

int×3 int×3 float str

description input volume file path path to the pretrained EVOLVE model checkpoint spatial block resolution (H ×W ×D) stride for overlapping blocks gain factor for quality control (0.0–10.0) directory for saving compressed bitstreams

Table 3: Key parameters of the decompression command.

parameter --bitstream_path --checkpoint

type str str

--output_dir

str

description compressed bitstream file path path to the pretrained EVOLVE model checkpoint directory for saving decompressed volumes

C C OMPRESSED F ILE C OMPOSITION Each EVOLVE compressed bitstream consists of three components: latent ŷ, hyperlatent ẑ, and meta info. The latent ŷ is the entropy-coded output of the encoder, representing the primary signal content. It is partitioned into five channel slices, each further decomposed into anchor and non-anchor positions via a checkerboard pattern for contextbased entropy coding. Because ŷ has the same spatial resolution as 1 the downsampled input (i.e., 16 along each axis), it carries the vast majority of the information and dominates the file size. The hyperlatent ẑ serves as side information: the hyper-encoder compresses ŷ into a much coarser representation (an additional 8× spatial downsampling relative to ŷ) to estimate the mean and scale parameters of the Gaussian entropy model. As a result, ẑ is 8×8×8 = 512× smaller than ŷ in voxel count, and its contribution to the total bitstream is inherently small. The meta info is a fixed-cost overhead required for correct decoding. It includes: (1) a global header storing the volume resolution, number of blocks, and number of channel slices; (2) per-block headers recording spatial position, block size, normalization range (vmin , vmax ), and the byte lengths of ŷ and ẑ payloads; and (3) per-slice descriptors storing the byte lengths of anchor and non-anchor bitstreams within each slice. These fields are stored as fixed-width integers and floats, so meta info scales linearly with the number of patches but is independent of the data content. Table 4 reports the composition breakdown on the combustion (MF) and isotropic datasets at two different CRs. Across all configurations, ŷ dominates (75.5%–98.0% of the total file size). ẑ accounts for only 1.3%–14.5%, and meta info contributes 0.7%–10.0%. At higher CRs (smaller files), the fixed-cost meta info and ẑ occupy a proportionally larger share. At lower CRs (larger files), ŷ grows disproportionately, reaching 95–98% of the total, confirming that the overhead of ẑ and meta info is negligible. Table 4: Compressed bitstream composition of EVOLVE at different CRs. dataset combustion (MF) combustion (MF) isotropic isotropic

CR 7,595 1,432 3,812 607

latent ŷ (%) 75.5 95.3 87.5 98.0

hyperlatent ẑ (%) 14.5 2.8 8.1 1.3

meta info (%) 10.0 1.9 4.4 0.7

Table 5: Peak GPU and host (CPU) memory during block-based (128×128×128 block) compression on a single NVIDIA RTX 4090, for two test volumes at two different compression rates. Because blocks are streamed to the GPU individually, the GPU footprint remains constant regardless of volume size, whereas the host memory grows with volume size. dataset

resolution

asteroids

1,000×1,000×1,000

gas

512×512×512

CR 6,523.7 4,527.6 6,784.1 4,597.9

PSNR (dB) 49.72 50.87 49.14 50.86

GPU (GB) 2.7 2.7 2.7 2.7

CPU (GB) 21.2 21.2 4.0 4.0

D I NFERENCE M EMORY F OOTPRINT EVOLVE compresses a volume by partitioning it into 128×128×128 blocks, each encoded and decoded independently, streaming a single block to the GPU at a time. Table 5 reports the peak GPU and host (CPU) memory during compression on a single NVIDIA RTX 4090, measured with nvidia-smi (per-process GPU memory) and the peak resident set (VmHWM), for two test volumes, asteroids (1,000×1,000×1,000, 3.7 GB) and gas (512×512×512, 512 MB), at two different compression rates. Since only one block resides on the GPU at any time, the GPU memory holds just the network parameters and the activations of the current block, and is therefore constant (∼2.7 GB) regardless of the volume size. The host memory instead holds the full input volume, its reconstruction, and the accumulated bitstream, and thus grows with the volume size. For a given volume, the footprint is independent of the target rate. This constant and modest GPU footprint allows EVOLVE to compress even gigabyte-scale volumes on a single consumer GPU (see Section 4.5). For volumes whose reconstruction exceeds the host memory, one can partition them into smaller subvolumes and process them sequentially. E L ATENT S PACE E XPLORATION AND A NALYSIS Similar to other autoencoder-based methods [20, 52, 57], the latent vectors extracted from the pretrained EVOLVE model are interpretable. In Figure 1, we compare the representative timestep selection results for the ionization-T (H+) dataset using the importance-driven framework proposed by Wang et al. [74] and EVOLVE. For EVOLVE, following [52], we project the latent vectors extracted from all volumes into 2D using t-SNE, with each point corresponding to a timestep. We then connect neighboring timesteps to form a trajectory in the latent space and select representative timesteps based on the arclength and angle metrics computed along this trajectory. Compared with the information-theoretic selection shown in Figure 1(a), we can observe that EVOLVE produces a more balanced temporal coverage, as shown in Figure 1(b). Early formation stages, transitions, and stabilization phases remain consistently represented. This result highlights the potential of EVOLVE for volume analysis applications.

for higher CRs. asteroids-T is a time-varying dataset containing 220 timesteps; we select two timesteps (t=165 and t=219) for visualization. Nyx-E is an ensemble dataset with varying cosmological simulation parameters; we select two members (ΩM =0.126 and ΩM =0.155) for visualization. The volume rendering results are shown in Figure 2. We observe that our method faithfully reconstructs volumes across all three quality levels. Future work will focus on incorporating additional temporal or ensemble information into the context model to achieve higher CRs. Table 6: Average PSNR (dB), LPIPS, ET (sec), DT (sec), and CR per timestep or ensemble using EVOLVE with different quality settings. dataset asteroids-T Nyx-E

ET↓ 7.97 7.12 6.43 6.43 6.09 6.72

DT↓ 6.27 8.15 6.29 6.43 8.39 8.39

CR↑ 2,706 4,924 12,486 304 661 2,093

Table 7: PSNR (dB), LPIPS, ET (sec), DT (sec), and CR on scanned volumetric data. For foot and engine, EVOLVE reports its highest achievable quality.

engine

(b) EVOLVE

F C OMPRESSION OF T IME -VARYING AND E NSEMBLE DATA EVOLVE compresses time-varying and ensemble datasets by independently compressing each volume. Table 6 reports the average quantitative results of our method on the asteroids-T and Nyx-E datasets under three quality settings (high, medium, and low), which correspond to decreasing gain factors that trade reconstruction quality

LPIPS↓ 0.299 0.320 0.415 0.024 0.043 0.104

H E VALUATION OF S CANNED VOLUMETRIC DATA To assess how EVOLVE behaves outside its training distribution, we evaluate it on four scanned volumes from the Open SciVis Datasets [34]: chameleon (1,024×1,024×1,080), engine (256×256×128), foot (2563 ), and stag-beetle (832×832×494), all normalized to [0, 1] in FP32. We compare EVOLVE (the released model, without any retraining) against ZFP, TTHRESH, and SZ3. For each dataset, EVOLVE operates at the high-quality level. Table 7 reports the quantitative results, and Figure 7 presents the corresponding rendering results. Because these scanned volumes fall outside EVOLVE’s training distribution, its overall benefit here is limited: although EVOLVE still edges out the conventional compressors on chameleon, engine, and stag-beetle, the margin on engine is small. For foot, its quality saturates at 33.7 dB and cannot be improved further by increasing the rate budget. This demonstrates the necessity of training on broader, more diverse data to further improve EVOLVE’s performance and generalization across a wider range of volumes.

chameleon

(a) Wang et al. [74]

PSNR↑ 59.06 57.51 50.27 50.60 46.56 39.19

G A DDITIONAL R ENDERING R ESULTS Due to page limits, the main paper only presents a subset of rendering comparisons. In Figures 3, 4, 5, and 6, we provide the complete set of volume rendering and isosurface rendering results across all test datasets, comparing EVOLVE with both conventional lossy compressors and deep-learning-based compressors at their respective CRs and reconstruction quality.

dataset

Figure 1: Comparison of representative timestep selection results using the ionization-T (H+) dataset. Each method selects 20 timesteps (marked along the vertical timeline) from 200 timesteps with seven rendering snapshots highlighted. In (a), the horizontal axis represents the timesteps.

quality high medium low high medium low

foot

stag-beetle

method ZFP TTHRESH SZ3 EVOLVE ZFP TTHRESH SZ3 EVOLVE ZFP TTHRESH SZ3 EVOLVE ZFP TTHRESH SZ3 EVOLVE

PSNR↑ 43.72 43.46 46.22 49.64 43.13 43.45 43.50 44.08 44.11 45.68 45.59 33.71 42.92 44.94 44.67 50.54

LPIPS↓ 0.1685 0.1196 0.0831 0.0390 0.0663 0.0663 0.0623 0.0600 0.0063 0.0031 0.0036 0.0929 0.0801 0.0357 0.0292 0.0247

ET↓ 6.22 225.75 16.38 67.61 0.07 2.98 0.14 0.51 0.12 5.61 0.28 1.07 0.78 64.71 3.89 30.11

DT↓ 3.68 83.18 8.71 55.04 0.03 0.30 0.09 0.66 0.10 1.14 0.21 0.94 0.85 22.27 2.05 26.49

CR↑ 127.3 80.4 701.9 728.1 79.3 103.2 99.7 326.6 29.0 19.8 35.4 257.4 919.6 49.5 1,061.1 3,829.7

t=165 t=219

high (2,706×)

medium (4,924×)

low (12,486×)

GT

high (304×)

medium (661×)

low (2,093×)

ΩM =0.155

ΩM =0.126

GT

Figure 2: Volume rendering of EVOLVE decompression results on asteroids-T (top) and Nyx-E (bottom), showing selected timesteps and ensemble members at different quality settings.

CR / PSNR

9,803× / 47.16 dB

2,482× / 40.85 dB

2,401× / 43.44 dB

1,164× / 40.00 dB

CR / PSNR

2,846× / 45.18 dB

985× / 40.03 dB

870× / 43.51 dB

64× / 43.79 dB

GT

EVOLVE

SZ3

TTHRESH

ZFP

Figure 3: Comparison of volume rendering results between EVOLVE and conventional lossy compressors. Top and bottom: asteroids and isotropic.

CR / PSNR

10,517× / 45.62 dB

1,714× / 43.35 dB

5,954× / 41.56 dB

1,979× / 41.37 dB

1,575× / 40.78 dB

936× / 40.24 dB

1,839× / 42.22 dB

1,808× / 41.28 dB

CR / PSNR

2,596× / 44.91 dB

446× / 41.50 dB

1,044× / 42.29 dB

633× / 40.44 dB

371× / 40.90 dB

166× / 42.23 dB

168× / 41.35 dB

1,114× / 42.17 dB

CR / PSNR

10,774× / 45.20 dB

3,047× / 43.29 dB

6,649× / 43.48 dB

5,256× / 42.09 dB

1,575× / 44.23 dB

1,813× / 42.16 dB

1,839× / 43.91 dB

3,748× / 38.84 dB

GT

EVOLVE

ECNR

NeurComp

SIREN

AMGSRN++

fV-SRN

Instant-NGP

IDLat

Figure 4: Comparison of volume rendering results between EVOLVE and deep-learning-based compressors. From top to bottom: gas, half-cylinder (VLM, 6,400), and magnetic.

CR / PSNR

6,047× / 49.31 dB

2,948× / 41.10 dB

2,714× / 42.98 dB

129× / 41.92 dB

CR / PSNR

7,843× / 47.58 dB

1,051× / 41.43 dB

6,334× / 40.11 dB

107× / 41.44 dB

GT

EVOLVE

SZ3

TTHRESH

ZFP

Figure 5: Comparison of isosurface rendering results between EVOLVE and conventional lossy compressors. Top and bottom: combustion (MF) and ionization (H+). The chosen isovalues are 0.5 and 0.1, respectively.

CR / PSNR

10,517× / 45.62 dB

1,714× / 43.35 dB

5,954× / 41.56 dB

1,979× / 41.37 dB

1,575× / 40.78 dB

936× / 40.24 dB

1,839× / 42.22 dB

1,808× / 41.28 dB

CR / PSNR

2,596× / 44.91 dB

446× / 41.50 dB

1,044× / 42.29 dB

633× / 40.44 dB

371× / 40.90 dB

166× / 42.23 dB

168× / 41.35 dB

1,114× / 42.17 dB

CR / PSNR

10,774× / 45.20 dB

3,047× / 43.29 dB

6,649× / 43.48 dB

5,256× / 42.09 dB

1,575× / 44.23 dB

1,813× / 42.16 dB

1,839× / 43.91 dB

3,748× / 38.84 dB

GT

EVOLVE

ECNR

NeurComp

SIREN

AMGSRN++

fV-SRN

Instant-NGP

IDLat

Figure 6: Comparison of isosurface rendering results between EVOLVE and deep-learning-based compressors. From top to bottom: gas, half-cylinder (VLM, 6,400), and magnetic. The chosen isovalues are 0.25, 0.6, and 0.05, respectively.

CR / PSNR

728.1× / 49.64 dB

701.9× / 46.22 dB

80.4× / 43.46 dB

127.3× / 43.72 dB

CR / PSNR

326.6× / 44.08 dB

99.7× / 43.50 dB

103.2× / 43.45 dB

79.3× / 43.13 dB

CR / PSNR

257.4× / 33.71 dB

35.4× / 45.59 dB

19.8× / 45.68 dB

29.0× / 44.11 dB

CR / PSNR

3,829.7× / 50.54 dB

1,061.1× / 44.67 dB

49.5× / 44.94 dB

919.6× / 42.92 dB

GT

EVOLVE

SZ3

TTHRESH

ZFP

Figure 7: Volume rendering comparison on scanned data (top to bottom: chameleon, engine, foot, and stag-beetle). For each method, the bottom-left inset shows the pixel-wise difference from GT, and the bottom-right inset shows the zoom-in of the red box. For foot, EVOLVE (33.71 dB) visibly smooths the trabecular bone texture, reflecting its quality saturation on out-of-distribution CT data, while for the other datasets EVOLVE attains the highest CR at matched or higher quality without any retraining.

ACKNOWLEDGMENTS This research was supported in part by the U.S. National Science Foundation through grants IIS-2101696, OAC-2104158, IIS-2401144, and CCF-2550610, and the U.S. Department of Energy through grant DE-SC0023145. We thank the anonymous reviewers for their insightful comments. R EFERENCES [1] A. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos. SemDeDup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540, 2023. doi: 10.48550/arXiv.2303.09540 3 [2] K. Ai, K. Tang, and C. Wang. NLI4VolVis: Natural language interaction for volume visualization via LLM multi-agents and editable 3D Gaussian splatting. IEEE Trans. Vis. Comput. Graph., 32(1):46–56, 2026. doi: 10. 1109/TVCG.2025.3633888 2 [3] J. Ballé. Efficient nonlinear transforms for lossy image compression. In Proc. Picture Coding Symposium, pp. 248–252, 2018. 3, 5 [4] J. Ballé, V. Laparra, and E. P. Simoncelli. End-to-end optimized image compression. In Proc. ICLR, 2017. 2, 4 [5] J. Ballé, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston. Variational image compression with a scale hyperprior. In Proc. ICLR, 2018. 2 [6] R. Ballester-Ripoll, P. Lindstrom, and R. Pajarola. TTHRESH: Tensor compression for multidimensional visual data. IEEE Trans. Vis. Comput. Graph., 26(9):2891–2903, 2020. doi: 10.1109/TVCG.2019.2904063 1, 2, 5, 6 [7] H. G. Barrow, J. M. Tenenbaum, R. C. Bolles, and H. C. Wolf. Parametric correspondence and chamfer matching: Two new techniques for image matching. In Proc. IJCAI, pp. 659–663, 1977. 6 [8] J. Beyer, M. Hadwiger, and H. Pfister. State-of-the-art in GPU-based large-scale volume visualization. Comput. Graph. Forum, 34(8):13–37, 2015. doi: 10.1111/cgf.12605 2 [9] B. Burkhart, S. Appel, S. Bialy, J. Cho, A. Christensen, D. Collins et al. The catalogue for astrophysical turbulence simulations (CATS). Astrophys. J., 905(1):14, 2020. doi: 10.3847/1538-4357/abc484 10 [10] J. Chen, S.-H. Kao, H. He, W. Zhuo, S. Wen, C.-H. Lee et al. Run, don’t walk: Chasing higher FLOPS for faster neural networks. In Proc. IEEE/CVF CVPR, pp. 12021–12031, 2023. doi: 10.1109/CVPR52729.2023. 01157 5 [11] Y. Chen, Q. Wu, W. Lin, M. Harandi, and J. Cai. HAC: Hash-grid assisted context for 3D Gaussian splatting compression. In Proc. ECCV, pp. 422–438, 2024. doi: 10.1007/978-3-031-72667-5_24 3, 5 [12] Z. Cheng, H. Sun, M. Takeuchi, and J. Katto. Learned image compression with discretized Gaussian mixture likelihoods and attention modules. In Proc. IEEE/CVF CVPR, pp. 7936–7945, 2020. doi: 10.1109/CVPR42600. 2020.00796 2 [13] Computer Graphics Laboratory, ETH Zürich. Scientific visualization datasets. https://cgl.ethz.ch/research/visualization/data. php. 3, 10 [14] Z. Cui, J. Wang, S. Gao, T. Guo, Y. Feng, and B. Bai. Asymmetric gained deep image compression with continuous rate adaptation. In Proc. IEEE/CVF CVPR, pp. 10532–10541, 2021. doi: 10.1109/CVPR46437.2021. 01039 3, 6 [15] J. Díaz, F. Marton, and E. Gobbetti. Interactive spatio-temporal exploration of massive time-varying rectilinear scalar volumes based on a variable bit-rate sparse representation over learned dictionaries. Comput. Graph., 88:45–56, 2020. doi: 10.1016/j.cag.2020.03.002 1, 2, 8 [16] D. B. Fielding, E. C. Ostriker, G. L. Bryan, and A. S. Jermyn. Multiphase gas and the fractal nature of radiative turbulent mixing layers. Astrophys. J. Lett., 894(2):L24, 2020. doi: 10.3847/2041-8213/ab8d2c 10 [17] E. Gobbetti, J. A. I. Guitián, and F. Marton. COVRA: A compressiondomain output-sensitive volume rendering architecture based on a sparse representation of voxel blocks. Comput. Graph. Forum, 31(3pt4):1315– 1324, 2012. doi: 10.1111/J.1467-8659.2012.03124.X 1, 2, 8 [18] P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola et al. Accurate, large minibatch SGD: Training ImageNet in 1 hour. arXiv preprint arXiv:1706.02677, 2017. doi: 10.48550/arXiv.1706.02677 3 [19] J. Han, K. Tang, and C. Wang. MoE-INR: Implicit neural representation with mixture-of-experts for time-varying volumetric data compression. IEEE Trans. Vis. Comput. Graph., 32(1):254–264, 2026. doi: 10.1109/TVCG .2025.3633893 6 [20] J. Han, J. Tao, and C. Wang. FlowNet: A deep learning framework for clustering and selection of streamlines and stream surfaces. IEEE Trans. Vis. Comput. Graph., 26(4):1732–1744, 2020. doi: 10.1109/TVCG.2018.

2880207 1, 2, 11 [21] J. Han and C. Wang. SSR-TVD: Spatial super-resolution for timevarying data analysis and visualization. IEEE Trans. Vis. Comput. Graph., 28(6):2445–2456, 2022. doi: 10.1109/TVCG.2020.3032123 3, 9 [22] J. Han, H. Zheng, and C. Bi. KD-INR: Time-varying volumetric data compression via knowledge distillation-based implicit neural representation. IEEE Trans. Vis. Comput. Graph., 30(10):6826–6838, 2024. doi: 10.1109/ TVCG.2023.3345373 1, 2, 6 [23] J. Han, H. Zheng, D. Z. Chen, and C. Wang. STNet: An end-to-end generative framework for synthesizing spatiotemporal super-resolution volumes. IEEE Trans. Vis. Comput. Graph., 28(1):270–280, 2022. doi: 10. 1109/TVCG.2021.3114815 2, 3 [24] E. R. Hawkes, R. Sankaran, J. C. Sutherland, and J. H. Chen. Scalar mixing in direct numerical simulations of temporally evolving plane jet flames with skeletal CO/H2 kinetics. Proc. Combust. Inst., 31(1):1633–1640, 2007. doi: 10.1016/j.proci.2006.08.079 10 [25] D. He, Z. Yang, W. Peng, R. Ma, H. Qin, and Y. Wang. ELIC: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding. In Proc. IEEE/CVF CVPR, pp. 5708–5717, 2022. doi: 10.1109/CVPR52688.2022.00563 2 [26] D. He, Y. Zheng, B. Sun, Y. Wang, and H. Qin. Checkerboard context model for efficient learned image compression. In Proc. IEEE/CVF CVPR, pp. 14771–14780, 2021. doi: 10.1109/CVPR46437.2021.01453 2, 5 [27] K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proc. IEEE/CVF CVPR, pp. 770–778, 2016. doi: 10. 1109/CVPR.2016.90 5 [28] W. He, J. Wang, H. Guo, K.-C. Wang, H.-W. Shen, M. Raj et al. InSituNet: Deep image synthesis for parameter space exploration of ensemble simulations. IEEE Trans. Vis. Comput. Graph., 26(1):23–33, 2020. doi: 10. 1109/TVCG.2019.2934312 9 [29] K. Hirashima, K. Moriwaki, M. S. Fujii, Y. Hirai, T. R. Saitoh, and J. Makino. 3D-spatiotemporal forecasting the expansion of supernova shells using deep learning towards high-resolution galaxy simulations. Mon. Not. R. Astron. Soc., 526(3):4054–4066, 2023. doi: 10.48550/arXiv. 2302.00026 10 [30] K. Hirashima, K. Moriwaki, M. S. Fujii, Y. Hirai, T. R. Saitoh, J. Makino et al. Surrogate modeling for computationally expensive simulations of supernovae in high-resolution galaxy simulations. arXiv preprint arXiv:2311.08460, 2023. doi: 10.48550/arXiv.2311.08460 10 [31] IEEE VIS. IEEE SciVis contests. https://sciviscontest.ieeevis. org/. 3, 10 [32] S. Jeon, K. Tang, C. Wang, and W.-K. Jeong. Super-Gaussian: Interactive scene editing for 3D Gaussian splatting and NLI-based volume visualization in virtual reality. IEEE Trans. Vis. Comput. Graph., 33(1), 2027. Accepted. 2 [33] B. Kim, V. C. Azevedo, N. Thuerey, T. Kim, M. H. Gross, and B. Solenthaler. Deep Fluids: A generative network for parameterized fluid simulations. Comput. Graph. Forum, 38(2):59–70, 2019. doi: 10.1111/CGF.13619 2 [34] P. Klacansky. Open scientific visualization datasets. http://klacansky. com/open-scivis-datasets/, 2017. 3, 10, 11 [35] A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In Proc. NeurIPS, pp. 1106–1114, 2012. doi: 10.1145/3065386 4 [36] K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch et al. Deduplicating training data makes language models better. In Proc. ACL, pp. 8424–8445, 2022. doi: 10.18653/V1/2022.ACL-LONG.577 3 [37] X. Liang, K. Zhao, S. Di, S. Li, R. Underwood, A. M. Gok et al. SZ3: A modular framework for composing prediction-based error-bounded lossy compressors. IEEE Trans. Big Data, 9(2):485–498, 2023. doi: 10. 1109/TBDATA.2022.3201176 1, 2, 6 [38] P. Lindstrom. Fixed-rate compressed floating-point arrays. IEEE Trans. Vis. Comput. Graph., 20(12):2674–2683, 2014. doi: 10.1109/TVCG.2014. 2346458 1, 2, 6 [39] J. Liu, S. Di, K. Zhao, S. Jin, D. Tao, X. Liang et al. Exploring autoencoderbased error-bounded compression for scientific data. In Proc. IEEE CLUSTER, pp. 294–306, 2021. doi: 10.1109/CLUSTER48925.2021.00034 1, 2, 3, 5 [40] I. Loshchilov and F. Hutter. Decoupled weight decay regularization. In Proc. ICLR, 2019. 4, 6 [41] G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao. DVC: An end-toend deep video compression framework. In Proc. IEEE/CVF CVPR, pp. 11006–11015, 2019. doi: 10.1109/CVPR.2019.01126 3 [42] Y. Lu, K. Jiang, J. A. Levine, and M. Berger. Compressive neural represen-

tations of volumetric scalar fields. Comput. Graph. Forum, 40(3):135–146, 2021. doi: 10.1111/CGF.14295 1, 2, 5, 6 [43] F. Marton, M. Agus, and E. Gobbetti. A framework for GPU-accelerated exploration of massive time-varying rectilinear scalar volumes. Comput. Graph. Forum, 38(3):53–66, 2019. doi: 10.1111/cgf.13671 1, 2, 8 [44] D. Minnen, J. Ballé, and G. Toderici. Joint autoregressive and hierarchical priors for learned image compression. In Proc. NeurIPS, pp. 10794–10803, 2018. 2 [45] D. Minnen and S. Singh. Channel-wise autoregressive entropy models for learned image compression. In Proc. IEEE ICIP, pp. 3339–3343, 2020. doi: 10.1109/ICIP40778.2020.9190935 2, 5 [46] T. Müller, A. Evans, C. Schied, and A. Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, 2022. doi: 10.1145/3528223.3530127 2, 6, 8 [47] J. Nystad, A. Lassen, A. Pomianowski, S. Ellis, and T. Olson. Adaptive scalable texture compression. In Proc. ACM SIGGRAPH/Eurographics HPG, pp. 105–114, 2012. doi: 10.2312/EGGH/HPG12/105-114 2 [48] R. Ohana, M. McCabe, L. Meyer, R. Morel, F. Agocs, M. Beneitez et al. The Well: A large-scale collection of diverse physics simulations for machine learning. In Proc. NeurIPS, pp. 44989–45037, 2024. 3, 10 [49] M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov et al. DINOv2: Learning robust visual features without supervision. Trans. Mach. Learn. Res., 2024, 2024. 3 [50] Z. Pan, K. Tang, J. Xia, Y. Qin, L. Gu, C. Wang et al. SGI: Structured 2D Gaussians for efficient and compact large image representation. In Proc. IEEE/CVF CVPR, pp. 12162–12172, 2026. doi: 10.48550/arXiv.2603.07789 3 [51] S. Popinet, M. Smith, and C. Stevens. Experimental and numerical study of the turbulence characteristics of airflow around a research vessel. J. Atmos. Ocean. Technol., 21(10):1575–1589, 2004. doi: 10.1175/1520-0426 (2004)021<1575:EANSOT>2.0.CO;2 10 [52] W. P. Porter, Y. Xing, B. R. von Ohlen, J. Han, and C. Wang. A deep learning approach to selecting representative time steps for time-varying multivariate data. In Proc. IEEE VIS (Short Papers), pp. 131–135, 2019. doi: 10.1109/VISUAL.2019.8933759 11 [53] J. Rissanen and J. Langdon, Glen G. Universal modeling and coding. IEEE Trans. Inf. Theory, 27(1):12–22, 1981. doi: 10.1109/TIT.1981.1056282 4 [54] M. B. Rodríguez, E. Gobbetti, J. A. I. Guitián, M. Makhinya, F. Marton, R. Pajarola et al. State-of-the-art in compressed GPU-based direct volume rendering. Comput. Graph. Forum, 33(6):77–100, 2014. doi: 10.1111/cgf. 12280 2 [55] I. B. Rojo and T. Günther. Vector field topology of time-dependent flows in a steady reference frame. IEEE Trans. Vis. Comput. Graph., 26(1):280–290, 2020. doi: 10.1109/TVCG.2019.2934375 10 [56] J. Schneider and R. Westermann. Compression domain volume rendering. In Proc. IEEE VIS, pp. 293–300, 2003. doi: 10.1109/VISUAL.2003.1250385 1, 2, 8 [57] J. Shen, H. Li, J. Xu, A. Biswas, and H.-W. Shen. IDLat: An importancedriven latent generation method for scientific data. IEEE Trans. Vis. Comput. Graph., 29(1):679–689, 2023. doi: 10.1109/TVCG.2022.3209419 1, 2, 3, 4, 5, 6, 11 [58] X. Sheng, J. Li, B. Li, L. Li, D. Liu, and Y. Lu. Temporal context mining for learned video compression. IEEE Trans. Multim., 25:7311–7322, 2023. doi: 10.1109/TMM.2022.3220421 3, 5 [59] V. Sitzmann, J. N. P. Martel, A. W. Bergman, D. B. Lindell, and G. Wetzstein. Implicit neural representations with periodic activation functions. In Proc. NeurIPS, pp. 7462–7473, 2020. 1, 2, 6 [60] H. Son, J. Noh, S. Jeon, C. Wang, and W.-K. Jeong. MC-INR: Efficient encoding of multivariate scientific simulation data using meta-learning and clustered implicit neural representations. In Proc. IEEE VIS (Short Papers), pp. 206–210, 2025. doi: 10.1109/VIS60296.2025.00047 2 [61] J. Sun, D. Lenz, H. Yu, and T. Peterka. F-Hash: Feature-based hash design for time-varying volume visualization via multi-resolution tesseract encoding. IEEE Trans. Vis. Comput. Graph., 32(1):396–406, 2026. doi: 10.1109/TVCG.2025.3634812 2 [62] K. Tang, K. Ai, J. Han, and C. Wang. TexGS-VolVis: Expressive scene editing for volume visualization via textured Gaussian splatting. IEEE Trans. Vis. Comput. Graph., 32(1):933–943, 2026. doi: 10.1109/TVCG.2025. 3634643 2 [63] K. Tang, D. Burke, and C. Wang. Lossless-INR: Lossless volumetric implicit neural representations. In Proc. IEEE VIS (Short Papers), 2026. Accepted. 2 [64] K. Tang and C. Wang. ECNR: Efficient compressive neural representation

of time-varying volumetric datasets. In Proc. IEEE PacificVis, pp. 72–81, 2024. doi: 10.1109/PACIFICVIS60374.2024.00017 2, 6, 9 [65] K. Tang and C. Wang. STSR-INR: Spatiotemporal super-resolution for time-varying multivariate volumetric data via implicit neural representation. Comput. Graph., 119:103874, 2024. doi: 10.1016/j.cag.2024.01.001 2 [66] K. Tang and C. Wang. StyleRF-VolVis: Style transfer of neural radiance fields for expressive volume visualization. IEEE Trans. Vis. Comput. Graph., 31(1):613–623, 2025. doi: 10.1109/TVCG.2024.3456342 2 [67] K. Tang and C. Wang. ECoNGS: Efficient compressive neural Gaussian splats for volume visualization. IEEE Trans. Vis. Comput. Graph., 33(1), 2027. Accepted. 3, 5 [68] K. Tang, S. Yao, and C. Wang. iVR-GS: Inverse volume rendering for explorable visualization via editable 3D Gaussian splatting. IEEE Trans. Vis. Comput. Graph., 31(6):3783–3795, 2025. doi: 10.1109/TVCG.2025. 3567121 2 [69] K. Tong, Y. Wu, Y. Li, K. Zhang, L. Zhang, and X. Jin. QVRF: A quantization-error-aware variable rate framework for learned image compression. In Proc. IEEE ICIP, pp. 1310–1314, 2023. doi: 10.1109/ICIP49359 .2023.10222717 6 [70] M. Treib, K. Bürger, F. Reichl, C. Meneveau, A. S. Szalay, and R. Westermann. Turbulence visualization at the terascale on desktop PCs. IEEE Trans. Vis. Comput. Graph., 18(12):2169–2177, 2012. doi: 10.1109/TVCG. 2012.274 2, 8 [71] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez et al. Attention is all you need. In Proc. NeurIPS, pp. 5998–6008, 2017. doi: 10.48550/arXiv.1706.03762 5 [72] G. K. Wallace. The JPEG still picture compression standard. Commun. ACM, 34(4):30–44, 1991. doi: 10.1145/103085.103089 2 [73] C. Wang and J. Han. DL4SciVis: A state-of-the-art survey on deep learning for scientific visualization. IEEE Trans. Vis. Comput. Graph., 29(8):3714–3733, 2023. doi: 10.1109/TVCG.2022.3167896 2 [74] C. Wang, H. Yu, and K.-L. Ma. Importance-driven time-varying data visualization. IEEE Trans. Vis. Comput. Graph., 14(6):1547–1554, 2008. doi: 10.1109/TVCG.2008.140 11 [75] Y. Wang, Z. Li, L. Guo, W. Yang, A. C. Kot, and B. Wen. ContextGS: Compact 3D Gaussian splatting with anchor level context model. In Proc. NeurIPS, pp. 51532–51551, 2024. 3, 5 [76] R. Webster, J. Rabin, L. Simon, and F. Jurie. On the de-duplication of LAION-2B. arXiv preprint arXiv:2303.12733, 2023. doi: 10.48550/arXiv. 2303.12733 3 [77] S. Weiss, P. Hermüller, and R. Westermann. Fast neural representations for direct volume rendering. Comput. Graph. Forum, 41(6):196–211, 2022. doi: 10.1111/CGF.14578 1, 2, 6, 8 [78] D. Whalen and M. L. Norman. Ionization front instabilities in primordial H II regions. Astrophys. J., 673(2):664–675, 2008. doi: 10.1086/524400 10 [79] Q. Wu, D. Bauer, M. J. Doyle, and K.-L. Ma. Interactive volume visualization via multi-resolution hash encoding based neural representation. IEEE Trans. Vis. Comput. Graph., 30(8):5404–5418, 2024. doi: 10.1109/TVCG. 2023.3293121 1, 8 [80] S. W. Wurster and H.-W. Shen. AMGSRN++: Improved adaptive SRN for scientific visualization. In Proc. IEEE PacificVis, pp. 182–191, 2025. doi: 10.1109/PacificVis64226.2025.00024 6 [81] S. W. Wurster, T. Xiong, H.-W. Shen, H. Guo, and T. Peterka. Adaptively placed multi-grid scene representation networks for large-scale data visualization. IEEE Trans. Vis. Comput. Graph., 30(1):965–974, 2024. doi: 10. 1109/TVCG.2023.3327194 1, 2, 5, 8 [82] L. Yan, X. Liang, H. Guo, and B. Wang. TopoSZ: Preserving topology in error-bounded lossy compression. IEEE Trans. Vis. Comput. Graph., 30(1):1302–1312, 2024. doi: 10.1109/TVCG.2023.3326920 2 [83] M. Yang, K. Tang, and C. Wang. Meta-INR: Efficient encoding of volumetric data via meta-learning implicit neural representation. In Proc. IEEE PacificVis (Visualization Notes), pp. 246–251, 2025. doi: 10.1109/ PACIFICVIS64226.2025.00030 2 [84] P. Yin, J. Lyu, S. Zhang, S. J. Osher, Y. Qi, and J. Xin. Understanding straight-through estimator in training activation quantized neural nets. In Proc. ICLR, 2019. 6 [85] C. Zauner, M. Steinebach, and E. Hermann. Rihamark: Perceptual image hash benchmarking. In Proc. IS&T/SPIE MWSF, pp. 343–357, 2011. doi: 10.1117/12.876617 3 [86] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proc. IEEE/CVF CVPR, pp. 586–595, 2018. doi: 10.1109/CVPR.2018.00068 6

Related documents

Record · ID 386978 · SHA-256 c8a7814a806ed007
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.