Conceptio › Archive › arXiv CS
arXiv CSopen access

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs Thanapat Trachu, Samuele Cornell, William Chen, Shinji Watanabe

arXiv:2609.17509v1 [cs.SD] 15 Sep 2026

Language Technologies Institute, Carnegie Mellon University, Pittsburgh, USA {ttrachu, scornell, wc4, swatanab}@andrew.cmu.edu

Abstract—Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe. Index Terms—Audio Codec, Text-to-Speech, Dynamic Frame Rate Codec, Compression

I. I NTRODUCTION Neural audio codecs have emerged as an important component for speech language modeling. When used as speech tokenizers, they convert continuous speech into discrete codes, allowing speech generation [1]–[3] and speech language models [4]–[6] to be formulated as next-token prediction problems similar to those used in large language models. However, general neural audio codecs [7]–[9] typically operate at high frame rates, producing long sequences of discrete codes. Long sequences of discrete codes introduce two limitations. First, long sequences increase the computational cost, as the complexity of transformer-based models tends to scale with sequence length [10]. Second, the length mismatch between speech and text tokens can degrade the performance of speech language models, since substantially longer speech sequences make syntactic and semantic modeling more difficult, as shown in [11]. These limitations motivate dynamic frame rate codecs [12]– [15], which represent speech as discrete codes paired with durations. A compression model produces segmentation boundaries, which are frame positions where the frame-level embeddings are partitioned. These boundaries divide the embeddings into variable-length segments, and each segment is represented by

one code and its duration. Since each code now corresponds to a segment rather than a single frame, the frame rate is measured as the number of segments per second instead of the number of frames per second. We refer to this as the effective frame rate. For example, if one second of audio with 75 framelevel embeddings is compressed into 25 segments, the effective frame rate is 25 Hz, instead of 75 Hz. By assigning longer durations to regions with lower information density, dynamic frame rate codecs use fewer codes to represent speech than general codecs. Most existing dynamic frame rate codecs either operate on single-codebook codec models [13], [15], or apply a single compression step before multi-layer quantization [14], [16], [17]. In both cases, all quantization layers are constrained to share the same segmentation boundaries. This constraint may be suboptimal for multi-codebook codecs due to the hierarchical nature of their quantization process. In residual vector quantization (RVQ) [18], the first layer quantizes the input directly, while each subsequent layer quantizes the residual embeddings from the previous layer. Consequently, earlier layers capture the main signal structure, whereas deeper layers encode finer details. As a result, representations at different layers may change at different rates over time, suggesting that each layer requires its own segmentation boundaries. This intuition is supported by our analysis in Section V-C1, which shows that deeper-layer residual embeddings change more rapidly than earlier-layer residual embeddings. This is further supported by the findings of SNAC [19], which uses different frame rates across quantization layers, although it is not a dynamic frame rate codec because its frame rates are fixed. Motivated by these findings, we propose LACE (LayerAdaptive Codec Encoding), a dynamic frame rate codec that independently applies a compression step at each quantization layer. This allows each layer to choose its own segmentation boundaries. To adapt LACE to the common speech language model framework, we propose union alignment, which constructs shared segmentation boundaries by taking the union across layers. Since union alignment can increase the effective frame rate, we further propose boundary anchor, which constrains deeper layers to reuse boundaries introduced by earlier layers. Our contributions are summarized as follows: • We propose LACE, a dynamic frame rate codec that applies an independent compression step at each quantization layer. • We propose union alignment and boundary anchor to

resolve duration inconsistency across layers, enabling LACE tokens to be used in downstream TTS training. • We provide a theoretical analysis showing that LACE achieves a lower upper bound on the expected quantization error than single compression. • Experiments on LibriTTS [20] show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. II. BACKGROUND A. Neural Audio Codec Preliminaries A neural audio codec consists of an encoder, a quantization module, and a decoder. The encoder maps the raw waveform into frame-level embeddings H ∈ RT ×D , where T is the number of frames and D is the embedding dimension. The quantization module discretizes H into frame-level codes Z ∈ [1, V ]T ×C , where each frame is represented by C codes from a codebook of size V . The decoder then reconstructs the waveform from these codes. A commonly used quantization method is RVQ. RVQ discretizes H using C quantization layers, where each layer l contains a codebook Ql ∈ RV ×D . Here, qlk ∈ RD denotes the k-th codeword of Ql . RVQ performs quantization sequentially across quantization layers. At layer l, the nearest codeword is selected from the corresponding codebook Ql as zil = arg mink∈[1,V ] ∥rli − qlk ∥22 , where zil ∈ [1, V ] and rli ∈ RD denote a code and a residual embedding at frame i and quantization layer l, respectively. The residual embedding is updated recursively as rl+1 = rli − qlzl , with r1i = hi , i i where hi ∈ RD is the i-th frame-level embedding in H. We define the stacked frame-level residual embeddings of layer l as Rl ∈ RT ×D .

[1, V ]T ×C . This restores the frame rate expected by the decoder. The three compression models below share this pipeline and differ only in how they derive B. a) Dynamic Programming (DP)-Based Compression.: CodecSlime [15] determines B by minimizing the total reconstruction error between the original frame-level embeddings H and the segment-level embeddings H̄. For a segment of s consecutive frames ending at frame j, the segment-level embedding is definedPas the mean of the frames within the j segment: h̄j,s = 1s t=j−s+1 ht . The reconstruction loss Pj for this segment is ℓ(j, s) = t=j−s+1 ∥ht − h̄j,s ∥22 . Each segment is constrained to contain at most U frames. Let f [j, n] denote the minimum total reconstruction loss when the first j frames are compressed into n segments. The DP objective is defined recursively as: f [j, n] = min {f [j − s, n − 1] + ℓ(j, s)} , 1≤s≤U

(1)

with f [0, 0] = 0. At each state (j, n), we record the segment length s∗ [j, n] that achieves the minimum in Equation (1). Given the number of segments M , we start from the last boundary bM = T and recover each preceding boundary as bn−1 = bn − s∗ [bn , n] for n = M, M −1, . . . , 1, which terminates at b0 = 0 and yields the segmentation boundaries B. b) Cosine Similarity-Based Compression.: FlexiCodec [14] determines B by thresholding the similarity between adjacent frame-level embeddings. Let σt = cos(ht , ht+1 ) denote the similarity between frames t and t + 1. A boundary is placed wherever the similarity drops below a threshold τ , yielding B = {0, T } ∪ { t : σt < τ }. The number of segments M is therefore not fixed in advance but controlled implicitly by the threshold τ , unlike the DP-based compression model. c) Density Peak Clustering-Based Compression.: VARSTok [13] derives B using density peak clustering. For B. Compression in Dynamic Frame-Rate Codecs each frame t, a local density ρt is computed from its κ nearest The encoder of a neural audio codec produces frame-level neighbors, and a peak distance δt is the temporal distance to embeddings H at a fixed codec frame rate, yielding T frames. the nearest frame with higher density: A compression model groups these T frames into M variable1 X ϕ(ht , hj ), δt = min |j − t|, (2) ρt = length segments. Each segment is represented by C codes and j:ρj >ρt κ j∈KNN(t) an integer duration. Formally, the compression model determines segmentation where ϕ(hi , hj ) = (1 + ⟨hi , hj ⟩)/2 and KNN(t) denotes the boundaries B = {b0 , b1 , . . . , bM }, a sorted set with b0 = 0, set of κ nearest neighbors of frame t. Frames with a high bM = T , and bm ∈ [1, T − 1]. The m-th segment spans the peak score pt = ρt δt are selected as segment centers. Each frame interval (bm−1 , bm ]. Some methods control M explicitly center is greedily expanded to adjacent unassigned frames that using a target compression rate γ ∈ (0, 1], yielding M = ⌊γT ⌋, satisfy ϕ(hi∗ , ht ) − βpt > τ , up to a maximum of U frames. while others determine M implicitly through a threshold τ . The This process repeats until all frames are assigned to segments. frame-level embeddings within each segment are averaged to The segment endpoints are then sorted to form B. The number produce segment-level embeddings H̄ ∈ RM ×D . The segment- of segments M is controlled implicitly by τ . We use κ = 5, level embeddings are paired with the durations d ∈ NM , where β = 0.2, and U = 4, following the original paper. the duration of the m-th segment dm is bm − bm−1 . The All of these prior works force all quantization layers to segment-level embeddings H̄ are quantized using RVQ, as share the same segmentation boundaries. In contrast, LACE described in Section II-A, producing segment-level codes Z̄ ∈ applies an independent compression step at each quantization [1, V ]M ×C . This (code, duration) format resembles the classical layer, enabling layer-specific segmentation boundaries. Because run-length encoding scheme [21]. During reconstruction, Z̄ LACE is agnostic to the choice of compression model, any of is repeated according to d to produce frame-level codes Ẑ ∈ these methods can be used as the compression step within the

Residual Embedding at Layer l

Before Alignment

Quantized Layer 1 Compression Step

Dur: 2

Dur: 1

Dur: 2

Quantized Layer 2 Quantized Layer l

Dur: 3

Dur: 1

Dur: 1

After Alignment

Quantized Layer 1

Repeat

Dur: 2

Dur: 1

Dur: 1

Dur: 1

Dur: 2

Dur: 1

Dur: 1

Dur: 1

Quantized Layer 2

Fig. 1. Overview of the LACE codec. At each quantization layer, the residual embeddings are independently compressed before being passed to the quantization layer. Frames with the same color belong to the same segment.

proposed framework. III. M ETHOD We propose LACE, a dynamic frame rate codec that applies an independent compression step at each quantization layer, allowing each layer to choose its own segmentation boundaries. We then introduce two mechanisms, union alignment and boundary anchor, to make the duration consistent for downstream TTS training.

Fig. 2. An overview of union alignment. Before alignment (top), quantization layers 1 and 2 have different segmentation boundaries, resulting in inconsistent durations across layers. After alignment (bottom), the union of segmentation boundaries from both layers is applied to all layers.

segment, that segment is split into sub-segments. For example, the boundary b1 = 2 from the first layer splits the first segment in the second layer. The resulting sub-segments keep the same code and receive new corresponding durations. After alignment, all layers share the same segmentation boundaries and durations, at the cost of increasing the effective frame rate. C. Boundary Anchor

Union alignment can increase the effective frame rate because |B union | ≥ |B l | for all l. In the worst case, each layer introduces Unlike prior dynamic frame rate codecs that apply a single distinct boundaries and |B union | approaches T , negating the compression step and share segmentation boundaries across compression benefit. To mitigate this, we introduce a boundary quantization layers, LACE performs layer-wise compression on anchor. We choose an anchor layer l∗ , where layers l ≤ l∗ are the residual embeddings. At layer l, we apply the compression allowed to introduce new boundaries, while deeper layers l > l∗ step described in Section II-B to Rl . This produces segment- must reuse boundaries from earlier layers: B anchor = Sl∗ B l . l l l=1 level residual embeddings R̄l ∈ RM ×D , durations dl ∈ NM , By restricting the compression step at layers l > l∗ to place and layer-specific segmentation boundaries B l , where M l is boundaries only at positions in B anchor , we limit the growth of the number of segments at layer l. The segment-level residual |B union |. Here, l∗ is a tunable hyperparameter that trades off embeddings R̄l are then passed to quantization layer l, yielding between quality and the effective frame rate. l segment-level codes Z̄l ∈ [1, V ]M . For DP-based compression, we implement this restriction To compute the residual embeddings for the next layer, we with a constrained cost ℓ̃(j, s), which equals the segment loss expand Z̄l to frame-level codes Ẑl ∈ [1, V ]T by repeating ℓ(j, s) of Section II-B if both j and j−s lie in B anchor and +∞ each code according to dl . The next residual embedding is otherwise. We solve the same recursion as Equation (1) with ℓ then computed as rl+1 = rli − qẑl l . Unlike the standard RVQ replaced by ℓ̃, and use the unconstrained DP for layers l ≤ l∗ . i i update in Section II-A, qẑl l is obtained by quantizing a segment- While this formulation is specific to DP-based compression, i level residual embedding, whereas qlzl in the original update the boundary anchor itself is model-agnostic. It can be applied i is obtained from a frame-level residual embedding. We use the to other compression models in Section II-B by restricting its anchor l . frame-level residual embedding ri on the right-hand side so that candidate boundaries to B l+1 ri captures both the compression error and the quantization D. Theoretical Analysis error from layer l. Figure 1 illustrates layer-wise compression. Section III-C introduced the anchor layer l∗ as a hyperparamB. Union Alignment eter that trades off quality against the effective frame rate. We The layer-wise compression produces different segmentation now analyze the quality side of this tradeoff: how the choice boundaries B l for each quantization layer l. As a result, codes of l∗ affects the expected quantization error. In Sections II-B from different layers that cover the same time span can have and III-A, the compression step merges the frames before the different durations. This is problematic for TTS because the quantization layer, and the codes are repeated back to the model would need to predict durations for each layer while frame level after it. Since all repeated frames within a segment ensuring that the total duration after upsampling is identical are identical, repeating before or after quantization yields the across layers. We address this with union alignment, which same result. We therefore combine merging and repetition into constructs shared segmentation by taking the union a single matrix multiplication Al Rl , where Rl is the frameSC boundaries union l across layers: B = l=1 B . We then re-segment the level residual embedding at layer l. Concretely, Al ∈ RT ×T segment-level codes of every layer using B union . As shown is block-diagonal with one block per segment, and the block in Figure 2, when a boundary in B union falls inside an existing for a segment of length dm is the dm × dm matrix with every A. Layer-Wise Compression

entry equal to 1/dm . Since Al is symmetric and idempotent, the compression error Rl − Al Rl is orthogonal to any framelevel embeddings that use the same segmentation boundaries as layer l. Let Q̂l ∈ RT ×D denote the frame-level quantized embeddings, which stack the selected codewords qẑl l , so the i

residual update of Section III-A becomes Rl+1 = Rl − Q̂l . The quantized embeddings Q̂l repeat one codeword across each segment. Subtracting them in the residual update can therefore only change the part of Rl that is constant within each segment. The compression error, which varies within segments, passes to the next layer untouched. When all layers share segmentation boundaries, the quantized embeddings of every layer are constant within the same segments, so the compression error survives all C quantization layers. It becomes an error that quantization cannot remove. We formalize this intuition under three assumptions: (i) every layer merges at least two frames (M l < T ), so the compression error is strictly positive, (ii) each quantization layer is well-trained, approximately satisfying the centroid condition [22], so that E[∥Al Rl − Q̂l ∥2F ] ≤ ϵl E[∥Al Rl ∥2F ] with ϵmax = supl ϵl < 1, and (iii) the fraction of expected residual energy captured by the compression step, αl = E[∥Al Rl ∥2F ]/E[∥Rl ∥2F ], is bounded below by αmin = inf l αl > 0. Under these assumptions, the expected residual error after all C quantization layers with anchor layer l∗ satisfies1 ∗ −1 h i lY C+1 2 C−l∗ +1 E[∥R ∥F ] ≤ 1 − αl∗ 1 − ϵmax λl E[∥H∥2F ],

model additionally predicts a duration for that code. Duration prediction is formulated as a classification task. We add a duration head alongside the code prediction head and include a learned duration embedding in the input representation, which is added to the code embedding. The model is trained with cross-entropy loss for code prediction and focal loss for duration prediction to address class imbalance, with loss weights of 1 and 3, respectively. At inference, we use nucleus sampling with a top-p value of 0.8 and a temperature of 1.0. IV. E XPERIMENTAL S ETTING A. Dataset We use the LibriTTS dataset [20] at 24 kHz sampling rate. The training set consists of train-clean-100, train-clean-360, and train-other-500. Evaluation is performed on test-clean. The same dataset is used for both reconstruction and TTS experiments. For TTS, a reference utterance is randomly selected from a different utterance of the same speaker to condition the model on speaker identity. B. Model Architecture

1) Codec Models: We build our method on three neural audio codecs: SoundStream [8], EnCodec [9], and the Descript Audio Codec (DAC) [7], all pretrained on LibriTTS using ESPnet-Codec [24] recipes2 . We integrate the proposed layerwise compression into the quantization layers (Section III-A) l=1 and fine-tune each model end-to-end from the pretrained (3) weights for 80k iterations using the original codec objective. where RC+1 is the residual after the last layer and λl = We use Adam [25] with learning rate 10−4 , an exponential 1−αl (1−ϵl ) < 1. As C → ∞, Equation (3) reveals a hierarchy. decay schedule (rate 0.9998 per step, minimum 10−5 ), and A single compression, which shares the same segmentation shared optimization settings for the generator and discriminator. boundaries across layers (l∗ =1), converges to a constant error All experiments use 4 V100 32GB GPUs. floor (1−α1 ) E[∥H∥2F ], which is exactly the compression error 2) TTS Model: The TTS model is a 24-layer Transformer of the first layer. The boundary anchor reduces this floor by with 16 attention heads, model dimension 1024, feedforward diQl∗ −1 the factor l=1 λl < 1. Full layer-wise compression (l∗ =C) mension 4096, and dropout 0.1. Input transcripts are converted drives the bound to zero. Therefore, recomputing boundaries to phoneme sequences using the espeak-ng phonemizer [26]. A at each quantization layer removes the error floor inherent to randomly truncated 3-second reference utterance is prepended shared segmentation boundaries. The full derivation is provided to condition the model on speaker identity. The model is trained in the supplementary material. This bound concerns the codec for 40k iterations on 4 V100 32GB GPUs using Adam with quantization error only, not downstream TTS, where union learning rate 10−4 and 4000 warmup steps. alignment (Section III-B) also changes the token sequence. C. Baselines E. TTS Training with LACE Tokens We compare LACE with CodecSlime [15], FlexiCodec [14], Unlike standard TTS models [2], [23], a TTS model trained on LACE tokens must also predict the durations d. After union and VARSTok [13]. Because these methods use different codec alignment, all C quantization layers share the segmentation backbones and training setups, we re-implement them on boundaries B union and durations d. Therefore, the model LibriTTS with the same codec backbones and fine-tuning predicts durations only for the first quantization layer and configuration, changing only the compression model. These baselines use single compression, where the compression step applies the predicted durations to all subsequent layers. We use an autoregressive decoder-only Transformer that is applied once before the quantization layers so all layers predicts the segment-level codes Z̄ in a delay-pattern for- share the same segmentation boundaries. In contrast, LACE mat [23]. At each generation step for the first-layer code, the applies compression independently at each quantization layer. 1 The analysis assumes that layers l > l∗ reuse the segmentation boundaries of layer l∗ exactly, a special case of the boundary anchor in Section III-C.

2 https://huggingface.co/espnet/libritts soundstream24k, https://huggingface. co/espnet/libritts encodec 24k, https://huggingface.co/espnet/libritts dac 24k

TABLE I R ECONSTRUCTION PERFORMANCE ACROSS DIFFERENT CODECS AND COMPRESSION MODELS ON L IBRI TTS T E S T - C L E A N . A LL BASELINE RESULTS USE A SINGLE COMPRESSION STEP WITHOUT LACE. “+ LACE” DENOTES OUR PROPOSED LAYER - WISE COMPRESSION . T HE FRAME RATE COLUMN REPORTS THE EFFECTIVE FRAME RATE AND CODEC FRAME RATE FOR CODECS WITH AND WITHOUT COMPRESSION MODELS , RESPECTIVELY. F OR THRESHOLD - BASED COMPRESSION MODELS , WE REPORT THE AVERAGE EFFECTIVE FRAME RATE ACROSS LAYERS . BASELINE ROWS ARE RE - IMPLEMENTED USING THE SAME BACKBONES . T HE RESULTS REPORTED BY THE CITED SYSTEMS ARE NOT DIRECTLY COMPARABLE . Bitrate (kbps)

WER ↓ (%)

UTMOS ↑

-

-

2.02

4.06

-

-

-

75 75 75

24.00 24.00 24.00

2.16 2.09 2.41

3.95 3.93 3.61

3.59 3.25 2.56

0.97 0.96 0.93

0.99 0.98 0.97

27 22.5

8.69 8.64

4.92 2.22

2.49 3.81

1.61 3.08

0.86 0.95

0.94 0.98

Cosine [14] Cosine + LACE

27 22.5

8.69 8.64

6.90 2.57

1.84 3.55

1.29 2.73

0.81 0.94

0.90 0.97

Density Clustering [13] Density Clustering + LACE

27 22.5

8.69 8.64

4.25 2.62

2.19 2.89

1.45 2.08

0.85 0.95

0.94 0.96

Finetuned SoundStream

DP DP + LACE

27 22.5

8.69 8.64

7.37 2.78

2.54 3.60

1.52 2.45

0.83 0.92

0.92 0.97

Finetuned Encodec

DP DP + LACE

27 22.5

8.69 8.64

3.10 2.18

3.08 3.83

1.99 2.86

0.89 0.95

0.96 0.98

Codec

Compression

Ground Truth

–

DAC [7] Encodec [9] SoundStream [8]

– – – DP [15] DP + LACE

Finetuned DAC

Frame Rate (Hz)

D. Rate Computation For reconstruction, we report the effective frame rate before union alignment, since alignment is not required. For thresholdbased compression models that yield different effective frame rates across layers, we report the mean effective frame rate across layers. The bitrate accounts for both code and duration bits, where each duration costs log2 U bits per segment. For a fair comparison, we compare at matched bitrate (∼ 8.6 kbps). Note that because LACE adds one duration sequence per layer, it reaches this bitrate at a lower effective frame rate than single compression (22.5 vs. 27 Hz). E. Evaluation Metrics

PESQ ↑

STOI ↑

SpkSim ↑

TABLE II TTS EVALUATION ON L IBRI TTS T E S T - C L E A N . R ATE IS THE EFFECTIVE FRAME RATE AFTER UNION ALIGNMENT. W E REPORT MOS AND SMOS, WITH 95% CONFIDENCE INTERVALS . RTF IS MEASURED ON A SINGLE V100 GPU FOR ONE CANDIDATE . Method

γ l∗ Rate WER ↓ UTMOS ↑ SpkSim ↑ (Hz) (%)

MOS ↑

SMOS ↑

RTF ↓

No compression – – 75.0

2.31

4.18

0.91

3.94 ± 0.05 3.94 ± 0.06 1.95

Single LACE LACE

0.5 – 37.5 0.5 2 52.5 0.5 3 58.0

30.73 4.93 8.52

1.79 2.69 2.70

0.84 0.86 0.88

3.80 ± 0.06 3.80 ± 0.07 1.01 3.88 ± 0.06 3.86 ± 0.07 1.58 3.89 ± 0.06 3.84 ± 0.06 1.68

Single LACE

0.7 – 52.5 0.7 2 61.8

6.77 5.11

2.59 3.45

0.85 0.90

3.85 ± 0.06 3.82 ± 0.07 1.45 3.96 ± 0.05 3.90 ± 0.06 1.73

V. R ESULTS

A. Reconstruction Task a) Reconstruction.: We compute word error rate (WER) We evaluate reconstruction quality across three codec backwith Whisper-Large [27] by transcribing the reconstructed bones (DAC [7], SoundStream [8], EnCodec [9]) and three audio. UTMOS [28] assesses perceptual naturalness, while compression models (DP [15], Cosine Similarity [14], and PESQ [29] and STOI [30] measure signal quality and intelligiDensity Clustering [13]). All systems are matched at a similar bility. For speaker similarity (SpkSim), we compute the cosine bitrate (∼ 8.6 kbps), resulting in different effective frame rates. similarity between X-vectors from the generated and reference Table I shows that LACE surpasses the single-compression audio, extracted with WavLM-base [31] fine-tuned for speaker 3 baselines in every configuration, indicating that its benefit verification . All metrics use VERSA [32] to compute. is agnostic to both the compression model and the codec b) TTS.: Since TTS generation is non-deterministic, we backbone. Figure 3 provides a finer-grained view through the generate 10 utterances per input and select the one with the rate-quality curve. LACE outperforms single compression at lowest WER under Whisper-Small, following prior work [33]. every bitrates, even on the pretrained model without fine-tuning. The selected utterance is evaluated with Whisper-Large for This indicates that the improvement comes from LACE itself, WER, along with UTMOS and SpkSim. We also report while fine-tuning provides further gains. the real-time factor (RTF), the ratio of generation time to audio duration, to assess inference efficiency. To complement B. Text-to-Speech Task these objective metrics, we conduct a human evaluation using We train and evaluate the TTS model using DAC as the mean opinion score (MOS) and speaker similarity mean codec backbone with DP compression. We compare against opinion score (SMOS) to evaluate naturalness and speaker two baselines: DAC without compression and DAC with single similarity, respectively. Specifically, we randomly sample 25 DP compression. Table II shows that, at the same target test utterances and synthesize them with each TTS system. compression rate γ, LACE outperforms single compression on Then, we collect ratings from 35 evaluators. all quality metrics. This improvement comes at the cost of a 3 https://huggingface.co/microsoft/wavlm-base-plus-sv higher RTF because union alignment increases the effective

10 15 20 25 Bitrate (kbps) ( )

Finetuned + Single Finetuned + LACE 4.8 4.0 3.2 2.4 1.6 10 15 20 25 Bitrate (kbps) ( )

Fig. 3. Rate-distortion comparison between LACE and Single compression on LibriTTS test-clean using DAC as the codec backbone and DP compression. Layer 0 EnCodec

Layer 8

Layer 16

l * = 1 (Single) l * =2

VQ Error

4.0 3.5 3.0 2.5 2.0

WER(%) ( )

DP

UTMOS ( )

Pretrained + Single Pretrained + LACE

0.4 0.2

l * =8 l * = 32 (LACE)

0.3 0.2 0.1

0.0 0.1 0.3 0.5 0.7 0.9 1 4 8 16 32 Compression rate Number of codebooks used Fig. 5. Empirical validation of the theoretical analysis using the pretrained DAC model with DP compression. Left: quantization error at various compression rates γ using all 32 quantization layers. Right: quantization error as a function of the number of quantization layers, at γ=0.5.

Layer 31 SoundStream

variation of earlier layers suggests that segmentation decisions have a larger impact on these layers than on deeper layers. 4 We therefore let earlier layers choose its own segmentation 2 boundaries before deeper layers, justifying the boundary anchor 0 (Section III-C). 0 1 0 1 0 1 2) Empirical Result of VQ Error: While Figure 4 only Frame-to-Frame Cosine Similarity Fig. 4. Distribution of cosine similarity between consecutive frame-level motivates layer-wise compression, this analysis directly mearesidual embeddings across quantization layers. sures the error caused by shared segmentation boundaries. We frame rate. Although the MOS scores of single compression measure the quantization error E[∥RC+1 ∥2F ] on the pretrained with γ = 0.5 and no compression are close, the paired t-test DAC model with DP compression while varying the anchor shows that raters can still distinguish between them (p < 0.01). layer l∗ . Figure 5 shows that LACE yields consistently Compared to the no-compression baseline, LACE achieves a lower quantization error than single compression, matching lower RTF by reducing the effective frame rate. This comes at Equation (3). The right panel shows the gap widening as more a quality cost, partly because the compressed codec has lower layers are used. Single compression converges to a constant reconstruction quality (Table I). The target compression rate γ error floor, while LACE keeps reducing the error. The left panel directly controls the quality–efficiency tradeoff. Increasing γ shows the role of γ. As γ → 1, compression merges almost no improves synthesis quality while increasing RTF. In contrast, frames, so α1 → 1 and the floor (1 − α1 ) E[∥H∥2F ] vanishes, increasing the anchor layer l∗ beyond 2 does not improve explaining the narrow gap at high compression rates. At lower quality and degrades WER. We hypothesize that this behavior rates the floor grows and the benefit of LACE becomes more arises from the union alignment step. A larger l∗ causes long pronounced. segments to be split into shorter ones, introducing repeated VI. C ONCLUSION discrete code IDs after alignment. These repeated code IDs bias the TTS model toward repeatedly predicting the same code IDs We presented LACE, a dynamic frame rate codec that applies during generation. Incorporating repetition-aware sampling [2] an independent compression step at each quantization layer. to prevent repetitive code prediction is a promising direction Unlike prior methods that force all layers to share the same for future work. To ensure that the observed WER trends are segmentation boundaries, LACE allows each layer to choose robust to the reranking, we computed the mean WER over its own segmentation boundaries. To enable downstream TTS 10 generated candidates using Whisper-Small. The ranking training, we introduced union alignment, which constructs remains consistent. LACE achieves 14% WER compared with shared segmentation boundaries across layers, and boundary 61% for single compression at γ = 0.5, and 14% compared anchor, which limits the growth of the effective frame rate. with 21% at γ = 0.7. Experiments on LibriTTS show that LACE consistently outDensity

DAC

C. Analysis We perform two analyses to better understand LACE. The first examines how fast the residual embeddings change at each quantization layer. The second empirically validates the quantization error bound from Section III-D. 1) Histogram of Cosine Similarity: On the pretrained codecs, we compute the cosine similarity between consecutive framelevel residual embeddings, σtl = cos(rlt , rlt+1 ). Figure 4 shows that its distribution concentrates around 1 at earlier layers and gradually shifts toward 0 at deeper layers. This indicates that residual embeddings change more rapidly at deeper layers and confirms that different layers change at different rates, motivating layer-wise compression. Moreover, the slower

performs single-compression baselines across multiple codec backbones and compression models on the reconstruction task. Applied to TTS, LACE improves inference efficiency while maintaining competitive quality. ACKNOWLEDGMENT Experiments of this work used the Bridges2 system at PSC and Delta and DeltaAI system at NCSA through allocations CIS210014 and IRI120008P from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.

G ENERATIVE AI U SE D ISCLOSURE Generative AI tools were used only to edit and polish the manuscript language and to assist with code writing. All research ideas, experimental design, implementation, analysis, and reported results are the work of the authors. The conceptual framing, methodology, and scientific contributions are entirely human-generated. R EFERENCES [1] S. Chen, C. Wang, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei, “Neural codec language models are zero-shot text to speech synthesizers,” IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 705–718, 2025. [2] S. Chen, S. Liu, L. Zhou, E. Liu, X. Tan, J. Li, S. Zhao, Y. Qian, and F. Wei, “VALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers,” 2025. [Online]. Available: https://openreview.net/forum?id=0bcRCD7YUx [3] P. Peng, P.-Y. Huang, S.-W. Li, A. Mohamed, and D. Harwath, “VoiceCraft: Zero-shot speech editing and text-to-speech in the wild,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 12 442–12 462. [Online]. Available: https://aclanthology.org/2024.acl-long.673/ [4] S. Guo, S. Zhang, Q. Fang, Z. Ma, M. Zhang, and Y. Feng, “FastLongSpeech: Enhancing large speech-language models for efficient long-speech processing,” in Advances in Neural Information Processing Systems, 2026. [Online]. Available: https://openreview.net/forum?id= jaMPaFDAaZ [5] J. Tian, J. Shi, W. Chen, S. Arora, Y. Masuyama, T. Maekaku, Y. Wu, J. Peng, S. Bharadwaj, Y. Zhao et al., “ESPnet-SpeechLM: An open speech language model toolkit,” in Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL): System Demonstrations, 2025, pp. 116–124. [6] D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, H. Guo, X. Chang, J. Shi, S. Zhao, J. Bian, Z. Zhao, X. Wu, and H. M. Meng, “UniAudio: Towards universal audio generation with large language models,” in International Conference on Machine Learning, 2024. [Online]. Available: https://openreview.net/forum?id=SRmZw7nEGW [7] R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “Highfidelity audio compression with improved RVQGAN,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 27 980–27 993. [8] N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 495–507, 2022. [9] A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research, 2023, featured Certification, Reproducibility Certification. [Online]. Available: https://openreview.net/forum?id=ivCd8z8zR2 [10] Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour, “AudioLM: A language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2523–2533, 2023. [11] H. Wang, H. Wang, Y. Guo, Z. Li, C. Du, and K. Yu, “Why do speech language models fail to generate semantically coherent outputs? a modality evolving perspective,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 16 117– 16 121. [12] H. Zhang, Y. Guo, Z. Li, X. Hao, X. Chen, and K. Yu, “Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate,” in Interspeech 2025, 2025, pp. 5003–5007. [13] R.-C. Zheng, W. Liu, H.-P. Du, Q. Zhang, C. Deng, Q. Chen, W. Wang, Y. Ai, and Z.-H. Ling, “Say more with less: Variable-frame-rate speech tokenization via adaptive clustering and implicit duration coding,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 41, 2026, pp. 35 021–35 029.

[14] J. Li, Y. Qian, Y. Hu, L. Zhang, X. Wang, H. Lu, M. Thakker, J. Li, S. Zhao, and Z. Wu, “FlexiCodec: A dynamic neural audio codec for low frame rates,” in International Conference on Learning Representations, 2026. [Online]. Available: https://openreview.net/forum?id=kYkfCs4ZAH [15] H. Wang, Y. Guo, C. Shao, B. Li, and K. Yu, “CodecSlime: Temporal redundancy compression of neural speech codec via dynamic frame rate,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 17 017–17 021. [16] W. Zhang, Y. Qian, Y. Cao, C. He, S. Xu, and M. Wang, “CARVE: Content-adaptive rate-variable encoding for neural speech codecs,” IEEE Signal Processing Letters, vol. 33, pp. 2036–2040, 2026. [17] Y. Qian, W. Zhang, X. Zhuang, S. Xu, L. Zhou, and M. Wang, “Arbitrarily settable frame rate neural speech codec with content adaptive variable length segmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 14 452–14 456. [18] D. Lee, C. Kim, S. Kim, M. Cho, and W.-S. Han, “Autoregressive image generation using residual quantization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 11 523–11 532. [19] H. Siuzdak, F. Grötschla, and L. A. Lanzendörfer, “SNAC: Multi-scale neural audio codec,” in Audio Imagination: NeurIPS 2024 Workshop AI-Driven Speech, Music, and Sound Generation, 2024. [Online]. Available: https://openreview.net/forum?id=PFBF5ctj4X [20] H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in Interspeech 2019, 2019, pp. 1526–1530. [21] S. Golomb, “Run-length encodings,” IEEE Transactions on Information Theory, vol. 12, no. 3, pp. 399–401, 1966. [22] A. Gersho and R. M. Gray, Vector Quantization and Signal Compression. Kluwer Academic Publishers, 1992. [23] J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Defossez, “Simple and controllable music generation,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 47 704–47 720. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 2023/file/94b472a1842cd7c56dcb125fb2765fbd-Paper-Conference.pdf [24] J. Shi, J. Tian, Y. Wu, J.-W. Jung, J. Q. Yip, Y. Masuyama, W. Chen, Y. Wu, Y. Tang, M. Baali, D. Alharthi, D. Zhang, R. Deng, T. Srivastava, H. Wu, A. Liu, B. Raj, Q. Jin, R. Song, and S. Watanabe, “ESPnet-Codec: Comprehensive training and evaluation of neural codecs for audio, music, and speech,” in IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 562–569. [25] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations, 2015. [Online]. Available: https://arxiv.org/abs/1412.6980 [26] M. Bernard and H. Titeux, “Phonemizer: Text to phones transcription for multiple languages in Python,” Journal of Open Source Software, vol. 6, no. 68, p. 3958, 2021. [Online]. Available: https://doi.org/10.21105/joss.03958 [27] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, Jul. 2023, pp. 28 492–28 518. [Online]. Available: https://proceedings.mlr.press/v202/radford23a.html [28] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab system for VoiceMOS challenge 2022,” in Interspeech 2022, 2022, pp. 4521–4525. [29] A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ) — a new method for speech quality assessment of telephone networks and codecs,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), vol. 2, 2001, pp. 749–752. [30] C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A shorttime objective intelligibility measure for time-frequency weighted noisy speech,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2010, pp. 4214–4217. [31] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., “WavLM: Large-scale self-supervised pretraining for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022.

[32] J. Shi, H. jin Shim, J. Tian, S. Arora, H. Wu, D. Petermann, J. Q. Yip, Y. Zhang, Y. Tang, W. Zhang, D. S. Alharthi, Y. Huang, K. Saito, J. Han, Y. Zhao, C. Donahue, and S. Watanabe, “VERSA: A versatile evaluation toolkit for speech, audio, and music,” in Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL): System Demonstrations, 2025. [Online]. Available: https://openreview.net/forum?id=zU0hmbnyQm [33] P. Mousavi, G. Maimon, A. Moumen, D. Petermann, J. Shi, H. Wu, H. Yang, A. Kuznetsova, A. Ploujnikov, R. Marxer, B. Ramabhadran, B. Elizalde, L. Lugosch, J. Li, C. Subakan, P. Woodland, M. Kim, H. yi Lee, S. Watanabe, Y. Adi, and M. Ravanelli, “Discrete audio tokens: More than a survey!” Transactions on Machine Learning Research, 2025. [Online]. Available: https://openreview.net/forum?id=eqNchtvc6v

I. S UPPLEMENTARY M ATERIAL All sections, equation, and figure references below refer to the main paper unless stated otherwise. A. Upper Bound of the Expected Quantization Error a) Setup.: Let H ∈ RT ×D be the frame-level embeddings drawn from an in-domain speech distribution. At layer l, the compression step performs two operations: (1) averaging the frame-level embeddings within each segment, and (2) repeating each segment-level embedding according to its duration. In Sections II-B and III-A, the repetition happens after the quantization layer rather than before, since all repeated frames within a segment are identical, the two orders yield the same result. We therefore combine both operations into a single matrix multiplication Al Rl , where Al ∈ RT ×T is the compression matrix at layer l and Rl is the frame-level residual embedding from Section III-A, with R1 = H. The matrix Al is block-diagonal: the block of the m-th segment has every entry equal to 1/dm , where dm is the segment duration. From this structure, Al is symmetric ((Al )⊤ = Al ) and idempotent ((Al )2 = Al ). Let VQl (·) denote quantization at layer l followed by upsampling, so that the frame-level quantized embedding is Q̂l = VQl (Al Rl ) and the residual update of Section III-A becomes Rl+1 = Rl − Q̂l . Since Q̂l is constant within each segment of layer l, applying the compression matrix leaves it unchanged: Al Q̂l = Q̂l .

(1)

We analyze the boundary anchor setting of Section III-C: layers l ≤ l∗ compute new boundaries, while layers l > l∗ ∗ reuse the compression matrix of layer l∗ , denoted A⋆ = Al . b) Assumptions.: 1) Strict compression: every layer merges at least two frames (M l < T ), so the compression error is strictly positive: E[∥Rl − Al Rl ∥2F ] > 0. 2) Centroid condition: each quantization layer is welltrained. Its codewords satisfy the centroid condition [1], i.e., each codeword equals the conditional mean of the inputs assigned to it. We further assume that each layer captures positive energy, E[∥Q̂l ∥2F ] > 0, and the resulting distortion ratios (Lemma 2) are uniformly bounded by ϵmax = supl ϵl < 1. 3) Positive energy capture: we define αl =

E[∥Al Rl ∥2F ] E[∥Rl ∥2F ]

(2)

as the fraction of expected residual energy captured by the compression step at layer l, and assume it is uniformly bounded below by αmin = inf l αl > 0. Lemma 1 shows αl ≤ 1, together with Assumption 1 and Equation (11), αl < 1. Lemma 1 (Compression never increases energy). For any R ∈ RT ×D and any compression matrix A, ∥AR∥2F ≤ ∥R∥2F .

(3)

Proof. The compression matrix acts independently on each segment and each feature dimension, so it suffices to prove the claim for a single segment and a single dimension. Let x1 , . . . , xdm denote the scalar values within a segment of duration dm . Compression Pdm replaces every value with the xi , so the energy of the segment segment mean µ = d1m i=1 Pdm 2 Pdm 2 1 2 changes from i=1 xi to dm µ = dm i=1 xi . By the Cauchy–Schwarz inequality with the all-ones vector, dm X

xi · 1

2

≤

dm X

i=1

i=1

x2i

dm  X

dm  X 12 = dm x2i ,

i=1

(4)

i=1

Pdm 2 xi . Summing over all segments which gives dm µ2 ≤ i=1 and all feature dimensions yields the claim. ■ Lemma 2 (Expected rate-distortion bound). Under Assumption 2, the expected quantization error at layer l satisfies     (5) E ∥Al Rl − Q̂l ∥2F ≤ ϵl E ∥Al Rl ∥2F , where ϵl = 1 − E[∥Q̂l ∥2F ]/E[∥Al Rl ∥2F ] < 1. Proof. Write X = Al Rl for the quantization input and X̂ = Q̂l for its output. The quantization partitions the input space into regions {Vk } and maps every input in region Vk to the output ck . By the centroid condition, ck = E[X | X ∈ Vk ], so the expected quantization error within each region is zero: E[X − ck | X ∈ Vk ] = E[X | X ∈ Vk ] − ck = 0.

(6)

By the law of total expectation, the quantization error is therefore orthogonal to the quantization output in expectation:   X E ⟨X − X̂, X̂⟩F = P(X ∈ Vk ) E[X − ck | X ∈ Vk ], ck F k

= 0,

(7)

where ck is pulled out of the conditional expectation because it is constant within its region. The cross-term thus vanishes in the expected Pythagorean decomposition: E[∥X∥2F ] = E[∥X̂∥2F ] + E[∥X − X̂∥2F ].

(8)

Rearranging gives the exact equality  E[∥X̂∥2F ]  E[∥X∥2F ] = ϵl E[∥X∥2F ], E[∥X − X̂∥2F ] = 1 − E[∥X∥2F ] (9) and Assumption 2 (E[∥X̂∥2F ] > 0) gives ϵl < 1. ■ T ×D Lemma 3 (Orthogonality). For any R, S ∈ R and any compression matrix A, the compression error R − AR is orthogonal to any compressed embedding AS under the Frobenius inner product:  ⟨R − AR, AS⟩F = Tr R⊤ (A − A2 )S = 0, (10) using symmetry and idempotence of A. In particular, choosing S = R yields the Pythagorean identity: ∥R∥2F = ∥AR∥2F + ∥R − AR∥2F .

(11)

Lemma 4 (Per-layer energy contraction). For every layer l that computes new boundaries,

equals A⋆ Rl , so Lemma 2 (Equation (5)) applies with input A⋆ Rl and bounds it:

E[∥Rl+1 ∥2F ] ≤ λl E[∥Rl ∥2F ],

E[∥A⋆ Rl+1 ∥2F ] = E[∥A⋆ Rl − Q̂l ∥2F ] ≤ ϵmax E[∥A⋆ Rl ∥2F ]. (19)

(12)

where λl = 1 − αl (1 − ϵl ) < 1. Proof. Decompose the residual update as Rl+1 = (Rl − Al Rl ) + (Al Rl − Q̂l ). By Equation (1), the second term can be rewritten as a compressed embedding: Al Rl − Q̂l = Al Rl − Al Q̂l = Al (Rl − Q̂l ).

E[∥Rl+1 ∥2F ] = E[∥Rl − Al Rl ∥2F ] + E[∥Al Rl − Q̂l ∥2F ] ≤ (1 − αl ) E[∥Rl ∥2F ] + ϵl αl E[∥Rl ∥2F ] (14)

The first term follows from Equation (11) and the definition of αl in Equation (2), which together give E[∥Rl − Al Rl ∥2F ] = (1 − αl ) E[∥Rl ∥2F ]. The second term applies Equation (5) and then Equation (2): E[∥Al Rl − Q̂l ∥2F ] ≤ ϵl E[∥Al Rl ∥2F ] = ϵl αl E[∥Rl ∥2F ]. ■ Lemma 5 (Invariance of the compression error). If Al = A⋆ for all l ≥ l∗ , then for all such layers: Rl+1 − A⋆ Rl+1 = Rl − A⋆ Rl .

(15)

(16)

Subtracting Equation (16) from the residual update cancels Q̂l and yields the claim. Hence, the compression error introduced at layer l∗ passes through all subsequent layers unchanged: it cannot be reduced by any quantization that reuses A⋆ . ■ Theorem 1 (Hierarchy of error bounds). After all C quantization layers, the expected residual error satisfies: ∗

−1 h i lY C−l∗ +1 E[∥RC+1 ∥2F ] ≤ 1 − αl∗ 1 − ϵmax λl E[∥H∥2F ]. l=1

(17) Proof. Decompose the final residual as RC+1 = (RC+1 − A⋆ RC+1 ) + A⋆ RC+1 . The two parts are orthogonal by Lemma 3, and applying Lemma 5 repeatedly over layers l∗ , . . . , C replaces the first part with the compression error at layer l∗ :   ∗ ∗ E[∥RC+1 ∥2F ] = E ∥Rl − A⋆ Rl ∥2F   + E ∥A⋆ RC+1 ∥2F . (18) ∗ The first term equals (1 − αl∗ ) E[∥Rl ∥2F ] by Equation (11)

and the definition of αl∗ in Equation (2). For the second term, consider any layer l ≥ l∗ . These layers share the same compression matrix (Al = A⋆ ), so Equation (16) shows that A⋆ Rl+1 is exactly the quantization error of layer l. Since Al = A⋆ , the quantizer input Al Rl

∗

∗

∗

C−l +1 = ϵmax αl∗ E[∥Rl ∥2F ]. (20)  Summing the two terms gives 1 − αl ∗ 1 − l∗ 2 l∗ 2 C−l∗ +1 E[∥R ∥F ]. Bounding E[∥R ∥F ] by unrolling ϵmax Lemma 4 over layers 1, . . . , l∗ − 1 yields Equation (17). ■ Equation (17) reveals a hierarchy of bounds as C → ∞: ∗ • Shared segmentation boundary (l = 1). All layers reuse the boundaries of the first layer, corresponding to single compression. The second term of Equation (18) vanishes and the error converges to a constant floor:

lim E[∥RC+1 ∥2F ] = (1 − α1 ) E[∥H∥2F ],

C→∞

(21)

which is exactly the compression error of the first layer and is strictly positive by Assumption 1. ∗ • Boundary anchor (1 < l < C). The error floor is reduced by the contraction of the layer-wise phase:

Proof. Applying A⋆ to the residual update Rl+1 = Rl − Q̂l and using Equation (1): A⋆ Rl+1 = A⋆ Rl − Q̂l .

∗

C−l +1 E[∥A⋆ RC+1 ∥2F ] ≤ ϵmax E[∥A⋆ Rl ∥2F ]

(13)

Therefore, the two terms are orthogonal by Lemma 3 with S = Rl − Q̂l , and their squared norms add:

= λl E[∥Rl ∥2F ].

Applying this bound recursively over the C − l∗ + 1 layers from l∗ to C, and then converting the compressed energy at layer l∗ to residual energy via Equation (2):

lim

C→∞ •

E[∥RC+1 ∥2F ] ≤ (1 − αl∗ )

∗ lY −1

λl E[∥H∥2F ]. (22)

l=1

Layer-wise compression without anchor (LACE). Every layer computes new boundaries, so Lemma 4 applies at every layer: E[∥RC+1 ∥2F ] ≤

C Y

C λl E[∥H∥2F ] ≤ λmax E[∥H∥2F ],

l=1

(23) where λmax = 1 − αmin (1 − ϵmax ) < 1, so the bound converges to zero as C → ∞. Therefore, recomputing boundaries inside the residual loop removes the constant error floor inherent to a shared segmentation boundary, while the boundary anchor interpolates between the two regimes. Remark. This analysis assumes that layers l > l∗ reuse the segmentation boundary of layer l∗ exactly. In our implementation (Section III-C), these layers may instead select any subset of B anchor . The theorem corresponds to the special case where the full segmentation boundary of layer l∗ is reused. R EFERENCES [1] A. Gersho and R. M. Gray, Vector Quantization and Signal Compression. Kluwer Academic Publishers, 1992.

Record · ID 919425 · SHA-256 1927255634c01a09
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.