ConceptioArchivearXiv CS
arXiv CSopen access

SAME: A Semantically-Aligned Music Autoencoder

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

SAME: A Semantically-Aligned Music Autoencoder Julian D. Parker Zach Evans CJ Carr Zachary Zukowski Josiah Taylor Matthew Rice Jordi Pons

arXiv:2605.18613v1 [cs.SD] 18 May 2026

Stability AI {julian.parker, zach, cj, zachary.zukowski, josiah, matt.rice, jordi.pons}@stability.ai Abstract Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this work we introduce SAME (Semantically-Aligned Music autoEncoder), an autoencoder for stereo music and general audio that reaches a 4096× temporal compression ratio while maintaining reconstruction quality and downstream generative performance. We achieve this by combining a tranformer-based backbone with set of semantic regularisation approaches, phase-aware reconstruction losses and improved discriminator designs. The architecture delivers substantial computational cost benefits, through both its high compression ratio and its reliance on well-optimised transformer primitives. Two variants (a large SAME-L and a CPU-deployable SAME-S) are released in open-weights form.

1

Introduction

Most modern generative media models operate on latent distributions (continuous or discrete) rather than on raw data. The paradigm was established in the image domain [1] using continuous latents from a variational autoencoder (VAE). In the audio domain these models are called Neural Audio Codecs (NACs). The dominant NAC paradigm, established by SoundStream [2] and refined by EnCodec [3] and DAC [4], is a VQ-VAE [5] variant with convolutional encoder/decoder networks and a vector-quantised bottleneck, trained with STFT-based reconstruction and adversarial losses. The quantized tokens can then be modelled by an autoregressive generative model [6, 7]. Diffusion and flow-matching generative models require continuous latents, so a different approach is needed. A popular example is the continuous NAC in Stable Audio Open [8], which shares the discrete-NAC architecture and training recipe but replaces the quantized bottleneck with a VAE bottleneck. Recent continuous NACs for music and general audio have largely followed this recipe, adjusting the training objective for audio quality [9, 10]. Several also use diffusion-model variants for decoding [11– 13]. In the speech domain, NAC innovation has accelerated with transformer-based encoder/decoder backbones [14, 15], query-based resampling [16], and alignment with semantic representations [17, 18] or ground-truth features [19]. We present SAME (Semantically-Aligned Music autoEncoder), which synthesises and extends many of these innovations. SAME consists of: 1. A query-based transformer resampling block (the Transformer Resampling Block, TRB), enabling fast inference and scaling to large parameter counts. 2. A bottleneck regularised for generative tractability and alignment with specific semantic concepts, improving generative performance. 3. Improved multi-resolution STFT (MRSTFT) reconstruction losses and an improved discriminator design, improving audio quality. We train two variants. SAME-L is an 852M-parameter model that outperforms baselines on audio quality while being significantly faster at inference. SAME-S is a distilled 108M-parameter variant with extremely fast inference, intended for CPU use on edge devices. Weights for both are released.1 1 https://stability-ai.github.io/SAME

1

Auxillary Losses 2P × T P

Patch

d× T PS

Encoder TRB

d× T PS

Soft-Norm Bottleneck

2P × T P

Decoder TRB

Unpatch

Reconstruction Losses Adversarial Losses

Figure 1: SAME architecture and training losses. Total compression: P S×. Dashed boxes indicate loss components.

2

Architecture

SAME follows an encoder-bottleneck-decoder structure (Fig. 1), in which a parameter-free patching pretransform and a Transformer Resampling Block (TRB) jointly achieve the target compression ratio.

2.1

Patching Pretransform

Stereo audio waveforms of shape (B, 2, T ) are partitioned into non-overlapping patches of P samples per channel and reshaped to (B, 2P, T /P ), so each embedding is a 2P -dimensional vector concatenating the P left-channel and P right-channel samples. With P =256 this gives 256× temporal downsampling with no learned parameters. Decoding applies the inverse reshape. Gradients flow through the transform, so the encoder and decoder train end-to-end against the original waveform.

2.2

Transformer Resampling Blocks

The Transformer Resampling Block (TRB) performs temporal resampling through self-attention rather than strided convolution or pooling, an approach shown to work well in image and speech domains [16, 20] that builds on earlier cross-attention-based resampling [21]. A TRB operates in either encoder or decoder mode, depending on the direction of resampling. Both modes share the same structure: a sequence of input embeddings is interleaved with learnable output embeddings, processed by a stack of transformer layers, and the output embeddings are extracted as the resampled representation. A stride parameter S controls the resampling ratio. In encoder mode, the TRB downsamples by S. A weight-normalised linear projection maps patch embeddings to the transformer’s internal dimension. The result is partitioned into N = ⌈T /S⌉ nonoverlapping segments of S embeddings each. A single learnable output embedding (initialised near zero and perturbed with low-amplitude Gaussian noise) is appended to each segment, forming subsequences of length S+1. The interleaved sequence is processed by D transformer layers. The output embedding is then extracted from each subsequence, yielding an S-fold reduction in temporal resolution. A linear projection then maps to the desired latent dimension. Fig. 2 illustrates encoder-mode interleaving with S=2. In decoder mode, the TRB upsamples by S. Each input embedding is paired with S learnable output embeddings, each perturbed with Gaussian noise, forming length-(S+1) subsequences in which the roles are reversed: the input provides context and the transformer populates the S outputs. After the stack, the outputs are extracted and the latent embeddings are discarded, yielding an S-fold increase in temporal resolution. A weight-normalised linear projection maps from the transformer dimension back to the patch embedding dimension. Each transformer layer is a pre-norm residual block. Self-attention uses differential attention [22] with per-head QK-normalisation and rotary position embeddings (RoPE) [23]. Both normalisation sites use Dynamic Tanh (DyT) [24], a learnable tanh(α · x) plus affine transformation that replaces LayerNorm/RMSNorm; DyT avoids the per-batch-element-statistics issues caused by silence or low-level noise in the input [14]. The feed-forward network is a gated linear unit (GLU) with SiLU activation. In the decoder, the final K of D layers use a sinusoidal activation f (x)= sin(πx) instead, providing a periodic basis suited to reconstructing waveform-level detail [4]. All branch outputs are zero-initialised so each layer starts as an identity. An audio autoencoder must handle variable-length audio and is usally trained on very short sequences (5s or less), so standard attention (causal or not) is unsuitable. We use one of two strategies (Fig. 3). Sliding-window attention lets each embedding attend to a fixed number of neighbours on each side, giving 2

Seg. 1 x0

x1

Seg. 2 x2

q

Seg. 3

x3

x4

q

x5

Seg. 4 q

x6

x7

q

Transformer

T1

x0

x1

T2

y0

x2

T3

x3

···

T4

y1

x4

x5

y2

TD

x6

x7

y3

extract y embeddings (discard x)

Figure 2: Embedding interleaving in encoder-mode TRB (stride S=2). Chunked + Midpoint Shift

Sliding Window Seg. 1

Seg. 2

Seg. 3

Seg. 4

Seg. 1

Seg. 2

Seg. 3

Seg. 4

Figure 3: Attention masks for a 12-embedding interleaved sequence (4 segments of S+1=3, dashed lines). Left: sliding-window attention. Right: chunked attention with midpoint shift. Teal: standard chunk boundaries (layers 1. . .⌊D/2⌋). Orange: shifted boundaries (layers ⌊D/2⌋+1. . .D). linear complexity in sequence length and bounding the receptive field for length generalisation. This is the preferred option. However, sliding-window attention is not supported by current CPU-inference libraries (e.g. LiteRT [25]), making chunked attention necessary for CPU deployment. Chunked attention folds the sequence into fixed-size chunks processed independently. Hard chunk boundaries can produce audible artefacts at transitions. We mitigate this with a midpoint shift: the first ⌊D/2⌋ layers use standard chunk boundaries, then the sequence is padded by half a chunk on each side (by repeating the edge segments) and rechunked with offset boundaries for the remaining layers. The single mid-stack rechunking adds one extra chunk per shifted half at negligible cost.

2.3

Soft-Normalisation Bottleneck

Between encoder and decoder we use a lightly constrained bottleneck rather than the VAE formulation. The encoder output passes through a learnable per-channel affine transform (scale and bias), then is divided by a running standard deviation tracked by exponential moving average, adapting to data statistics during training and normalising latent magnitudes to a consistent range. A KL-like regularisation loss Lkl = E[µ2t + σt2 − log σt2 − 1] + 0.4 E[µ2c + σc2 − log σc2 − 1]

(1)

encourages zero-mean, unit-variance statistics along two axes independently, where (µt , σt2 ) are perchannel mean and variance over time and (µc , σc2 ) are per-timestep mean and variance over channels. The channel-axis term is downweighted to 0.4 to reflect the asymmetry between the two dimensions. The dual-axis penalty prevents both per-channel drift and per-timestep outliers. Decoding reverses the normalisation by multiplying by the running standard deviation. Gaussian noise scaled by the same standard deviation is added to the latent, at higher scale in training (5 × 10−2 ) than at inference (10−3 ). This smooths the latent manifold and makes the decoder robust to errors from downstream diffusion-based modelling, which has been shown crucial when modelling large semantic representations in the image domain [26, 27].

3

3

Training Objectives

3.1

Spectral Reconstruction Losses

We use a multi-resolution STFT loss [28] at seven resolutions (FFT sizes 32, 64, 128, 256, 512, 1024, 2048; 75% overlap each). At each resolution the loss combines a spectral contrast term, a modified log-magnitude L1 distance, and phase-derivative losses (below). A K-weighting pre-emphasis filter is applied before the STFT to focus the loss on perceptually relevant frequencies. For stereo we compute the loss independently on mid/side and left/right representations to preserve stereo image. The full multi-resolution spectral loss sums three components over all R resolutions: LMRSTFT =

R X

(r) (r) (r)  LSC + LLM + LIFGD ,

(2)

r=1

each of which is defined below. 3.1.1

Spectral Contrast

Let X (predicted) and Y (reference) denote magnitude spectrograms with X, Y ≥ 0. The standard spectral convergence loss ∥Y −X∥F /∥Y ∥F [28] normalises by the reference alone, making it asymmetric and unbounded when the prediction exceeds the reference. We replace it with a spectral contrast loss: LSC =

∥Y − X∥F , ∥X + Y ∥F + ϵ

(3)

with ϵ a small numerical-stability constant (used similarly throughout this section). Since ∥Y −X∥F ≤ ∥X+Y ∥F , the loss is bounded in [0, 1], symmetric, and scale-invariant. 3.1.2

Adaptive Log-Magnitude

The standard log-magnitude loss applies L1 between log(X + ϵ) and log(Y + ϵ) with a fixed ϵ. log(x + ϵ) transitions from approximately linear (≈ x/ϵ) for x ≪ ϵ to logarithmic for x ≫ ϵ, so ϵ sets the knee on an absolute scale. Small ϵ (e.g. 10−8 ) produces very large gradients 1/(x + ϵ) at low-amplitude bins, letting analysis-window leakage dominate [29]; the common remedy of ϵ=1 [29] suppresses leakage but fixes the knee at an arbitrary absolute magnitude unrelated to the signal. We replace the constant with an adaptive normalisation:   Y LLM = log X (4) σ + 1 − log σ + 1 1 , p where σ = std(X)2 + std(Y )2 is computed over frequency and time (gradient-detached). Bins well above σ (significant spectral content) receive logarithmic compression, while bins well below (noise floor, leakage) are linearised with reduced gradient. Since X/σ is dimensionless, the loss is globally scale-invariant while preserving relative weighting across bins. 3.1.3

Phase-derivative Losses

Magnitude-only spectral losses discard phase structure, yet phase coherence is critical for transient fidelity, pitch accuracy, and stereo imaging. Prior work uses L1 differences of the phase derivatives (instantaneous frequency and group delay, IFGD) [9], which capture perceptually-meaningful phase relationships but potentially suffer from discontinuities at the modulo-2π boundary. We instead operate on normalised complex phasors, avoiding phase unwrapping entirely. Henceforth let X, Y ∈ CF ×T be the complex STFTs of the predicted and reference signals, with F frequency bins and T time frames; · denotes complex conjugate. For Z ∈ {X, Y } we form crossframe and cross-bin products RtZ (f, t) = Z(f, t) Z(f, t−1) and RfZ (f, t) = Z(f, t) Z(f −1, t), whose arguments are the instantaneous-frequency and group-delay increments, and normalise them to unit phasors UtZ (f, t) = RtZ (f, t)/(|Z(f, t)| |Z(f, t−1)| + ϵ) and analogously UfZ . The IF and GD losses measure cosine distance between predicted and reference phasors, weighted by a detached, mean-normalised p geometric-mean magnitude factor wt ∝ |X(f, t)| |X(f, t−1)| |Y (f, t)| |Y (f, t−1)| (with wf analogous

4

over frequency-adjacent magnitudes) that focuses the loss on energetic time-frequency regions while keeping it scale-invariant: h   i LIF = E wt 1 − Re UtX UtY , (5) h   i LGD = E wf 1 − Re UfX UfY , (6)   2 Lcd = E log(|X − Y | /σsg + 1) . (7) The third term, a normalised complex-distance penalty, uses the stop-gradient standard deviation σsg of |X − Y |2 over time and frequency for self-normalising scaling. The combined phase-aware loss is LIFGD = LIF + LGD + Lcd .

3.2

Adversarial Training

We use a relativistic paired GAN objective [30]. For each discriminator k in a multi-view ensemble, the discriminator and generator losses take the softplus form: h  i (k) Ladv (D) = E log 1 + e−(Dk (x)−Dk (x̂)) , (8) h  i (k) Ladv (G) = E log 1 + eDk (x)−Dk (x̂) , (9) where x and x̂ are real and reconstructed audio. A feature-matching loss averages L1 distances between intermediate features across all discriminator layers: Lfm =

K

L

k=1

l=1

k 1 X 1 X ∥fk,l (x) − fk,l (x̂)∥1 , K Lk

(10)

where fk,l denotes the layer-l features of discriminator k. The total generator-side adversarial loss combines these with a feature-matching weight λfm : X (k) 1 Ladv (G) + λfm Lfm . (11) Ladv = K k

We use two multi-view discriminator architectures that share the same GAN and feature-matching objectives but differ in backbone and signal views, deployed at different stages of training (Section 4.2): the convolutional discriminator has sharper resolution but is prone to artefacts; the transformer-based discriminator avoids artefacts but can over-smooth. 3.2.1

Convolutional Discriminator

The first configuration extends the EnCodec multi-scale STFT discriminator [3] with two additional signal views, for a total of 7 discriminators. A multi-scale STFT component applies a 2D convolutional stack to the complex spectrogram at five resolutions (FFT sizes 128, 256, 512, 1024, 2048). A PQMF filter-bank component [10, 31] decomposes the waveform into subbands via pseudo-QMF analysis and applies weight-normalised 2D convolutions over the (subband, time) plane. A chroma component computes a 48-bin chromagram and discriminates on pitch-class distributions. 3.2.2

Transformer Discriminator

The second configuration replaces the convolutional STFT and chroma discriminators with TRB-based versions, and adds TRB-based patched-waveform discriminators. The PQMF filter-bank component is retained and remains convolutional, for a total of 10 discriminators. We use three STFT discriminators (FFT 128/1024/4096), three chroma discriminators (octave centres 1/5/9), and three patched waveform discriminators that reshape raw audio into non-overlapping patches at prime sizes (29, 443, 953) to avoid harmonic aliasing.

5

3.3

Auxiliary Losses

3.3.1

Generative Alignment Loss

For discrete tokenizers, it is common to jointly train a small auxiliary autoregressive model that backpropagates gradients into the encoder, shaping the latent space for downstream generation [16, 32]. The equivalent procedure for continuous latents under diffusion or flow-matching is rarer, with some recent precedent [27]. For SAME we train a small unconditional diffusion transformer (4 layers, 768-dim) jointly on the autoencoder’s latent space with a flow-matching objective [33]. At each training step, a timestep t is sampled from a truncated logistic-normal distribution and the latent z is noised via zt = (1 − t) z + t ε, ε ∼ N (0, I). The model predicts the velocity vθ (zt , t) ≈ ε − z:   Ldiff = Et,ε ∥vθ (zt , t) − (ε − z)∥22 . (12) During a warmup phase the diffusion model trains on detached latents. Gradients then flow through the encoder, shaping the latent geometry for diffusion-based generation. 3.3.2

Semantic Regression Losses

We train lightweight linear regressors (single 1×1 convolutions) to predict perceptually meaningful audio features directly from thePlatent representation. Each regressor gi maps latents to a target yi via a weighted L1 loss: Lsem = i λi ∥gi (z) − yi ∥1 . Chroma regression. Three chroma regressors target octave-band chromagrams centred at octaves 1, 5, and 9 (with octave widths of 1.0, 1.5, and 1.0 respectively), each projecting the latent to 128 chroma bins. The targets are computed from a high-resolution spectrogram of size 8192. Interaural level difference (ILD) regression. An additional regressor predicts the interaural level difference (the per-band log-magnitude difference between left and right channels, computed on a 32band mel spectrogram), explicitly encoding spatial information important for faithful stereo-image reconstruction. 3.3.3

Contrastive Latent Alignment

A transformer-based critic (4 layers, 1024-dim) is trained to decide whether a latent sequence, an audiofeature sequence, and a text embedding come from the same input. The audio features are an 8-level Cohen-Daubechies-Feauveau 9/7 biorthogonal wavelet decomposition [34]. The text embeddings are produced with T5Gemma [35]. The critic maps all three modalities into a shared space via learned linear projections, concatenates them along the sequence axis with a learnable critic token, and processes the result through transformer layers to produce a scalar score. A softplus margin loss compares positive (matched) triplets against negatives formed by independently rotating the audio and text components within the batch: h  i + − Lcon = E log 1 + em−(C(z,a,t) −C(z,a,t) ) , (13) where C is the critic score, +/− denote matched and mismatched inputs, and m is a margin hyperparameter. Sequence- and feature-level masking (dropping 40% and 35% of positions) and volume augmentation prevent the critic from relying on trivial cues. As with Ldiff , the critic trains on detached latents during a warmup phase before end-to-end gradients are enabled. This loss preserves audio-level and cross-modal semantics, complementing the geometric regularity targeted by Ldiff .

4

Model Configuration and Training

4.1

Model Configuration

4.1.1

Choice of downsampling-ratio and latent dim

Increasing temporal downsampling (Dt ) shortens the sequence length needed to represent a given duration of audio, making downstream modelling computationally cheaper (and potentially easier). At fixed latent dimension d, however, this raises the overall compression ratio and thus the reconstruction difficulty. Raising d mitigates this, though early latent-diffusion work argued that small d was needed to ease generative modelling [1]; more recent work shows larger d is tractable given good semantic structure [26].

6

We therefore target Dt =4096 (roughly twice the standard for audio autoencoders) and a relatively large d=256. Sec. 5.1 examines the interaction of these choices with our semantic regularisation. We train two configurations, both of which utilize a waveform patch size of 256 and a TRB stride of S = 16 to achieve the target downsampling ratio. 4.1.2

SAME-L (Large)

SAME-L uses a transformer dim of 1536 with 12 transformer blocks in both the encoder and decoder. The total parameter count is 852M. Sliding-window attention attending to S+1 positions on each side is used (Fig. 3). In the decoder, the last K=8 layers use sinusoidal activations. Training uses 4.46-second segments (196 608 samples per channel) at total batch size 192. 4.1.3

SAME-S (Small)

SAME-S reduces the transformer dimension to 768 and depth to 6 layers, and uses chunked attention with midpoint shift (Section 2.2) at a chunk size of 32. Differential attention and sinusoidal feed-forward layers are dropped to maximise CPU performance. Total parameter count is 108M. Training uses 0.56-second segments (24 576 samples per channel) at total batch size 1024. During pretraining (Stage 1, below) SAME-S is distilled from a frozen SAME-L teacher. A latent loss Ldistill = ∥zS − zT ∥1 aligns student and teacher encodings. LMRSTFT and Ladv are applied not only to the direct reconstruction DS (zS ) but also to three cross-decoded outputs: DT (zS ), DS (zT ), and DS (zS ) against DT (zT ). Each cross-term is weighted 0.25× the main reconstruction loss, ensuring bidirectional encoder–decoder compatibility.

4.2

Training Procedure

Both configurations follow a three-stage procedure on 32 H100 GPUs with Cautious AdamW [36] (β=(0.9, 0.95) and weight decay 10−4 for the autoencoder; β=(0.8, 0.99) for the discriminator), inversesquare-root learning-rate scheduling, and EMA weight averaging. Stage 1 – Pretraining (500k steps). The full autoencoder trains end-to-end with LMRSTFT , Lkl , the convolutional discriminator (Section 3.2.1), and model-specific auxiliary losses: Ldiff , Lsem , Lcon for SAME-L; the cross-model distillation objectives (Ldistill , cross-decoded LMRSTFT /Ladv ) and Lsem for SAME-S. Stage 2 – Decoder finetuning, convolutional discriminator (100k steps). The encoder is frozen and the convolutional discriminator is reset. Only LMRSTFT and Ladv remain active. Stage 3 – Decoder finetuning, transformer discriminator (100k steps). The convolutional discriminator is replaced by the transformer discriminator (Section 3.2.2), with the encoder still frozen. Synthetic linear chirps are appended to each batch to mitigate aliasing: frequencies sampled log-uniformly in [100 Hz, 22 kHz] over 2–6.5 octaves, amplitude uniform in [−24, −6] dBFS. All models are trained on Audiosparx2 production music, following the dataset and split of [37]: ≈19,500 h with a 66/25/9% mix of music, sound effects, and instrument stems.

5

Evaluation

All evaluation is performed on 446 track/caption pairs from the Song Describer Dataset (SDD) [38].

5.1

Ablation Studies

To isolate the contributions of the bottleneck and the auxiliary losses, we run a lightweight ablation: 50k autoencoder training steps at batch size 128 with the SAME-L backbone, without adversarial losses. This fixed-budget, spectral-loss-only setting removes GAN training dynamics as a confounder. VAE variants use a KL weight of 10−4 . For each configuration we then train an ≈1.4B-parameter DiT with a flow-matching objective, also for 50k steps at batch size 128. Tab. 2 reports one reconstruction metric on the autoencoder — MELlog1p , a multi-resolution log-mel error (see Section 5.3) — and two generation metrics on the DiT outputs conditioned on SDD captions: Fréchet audio distance in CLAP space (FAD-CLAP, via the fadtk [39] toolkit with the 630k-audioset-best checkpoint) and the reference-less MuQ-Eval [40] musical quality score. 2 https://www.audiosparx.com

7

ϵar-VAE

ACE-Step 1.5 SAO VAE CoDiCodec†

SAME-S

SAME-L

Dt d

1024 64

1920 64

2048 64

4096 64

4096 256

4096 256

RTF ↑

325

284

300

47

2069

561

7.0 ±3.3 0.084 ±0.051 0.069 ±0.034 93.2 ±4.7 76.5 ±20.0

6.2 ±3.3 0.092 ±0.055 0.079 ±0.039 92.2 ±5.2 73.3 ±19.5

−0.3 ±3.1 0.096 ±0.057 0.096 ±0.044 81.7 ±10.6 —

SI-SDR ↑ 12.0 ±3.9 STFTlog1p ↓ 0.080 ±0.053 MELlog1p ↓ 0.070 ±0.042 CCPC ↑ 97.2 ±2.2 MUSHRA ↑ 77.6 ±21.0

9.6 ±3.4 11.9 ±4.2 0.088 ±0.055 0.081 ±0.053 0.071 ±0.035 0.057 ±0.031 95.5 ±3.3 96.6 ±3.0 66.1 ±20.5 82.2 ±16.6

Table 1: Objective reconstruction quality and MUSHRA listening test. Bold: best. Underline: second best. Dt : temporal downsampling. d: latent dimension. RTF: audio duration / encode+decode wall-clock time; higher is faster. MUSHRA: 36 trials, 12 participants after filtering; 0–100 scale, mean ± 95% CI; reference scored 97.6±4.9 and the 64 kbps MP3 anchor 30.9±22.6 (excluded from ranking). † CoDiCodec in 2-step autoregressive mode; excluded from MUSHRA.

Dt d Bot. Ldiff Lsem , Lcon

E

A

B

C

D

1024 64 VAE — —

4096 256 VAE — —

4096 256 SN — —

4096 256 SN ✓ —

4096 256 SN ✓ ✓

MELlog1p ↓ 0.098 0.108 0.108 0.103 0.109 FAD-CLAP ↓ 0.724 0.651 1.061 0.593 0.576 MuQEval ↑ 3.194 3.252 2.783 3.340 3.870 Table 2: Ablation study. Dt : temporal downsampling; d: latent dimension; Bot. = bottleneck; SN = soft-normalisation; ✓/— indicate presence/absence of each auxiliary loss. Soft-normalisation alone (A → B) regresses both generation metrics: the simpler bottleneck requires the auxiliary losses it was designed to enable. The flow-matching alignment loss (B → C) recovers and surpasses the VAE baseline on all three metrics, and adding the semantic regressors and contrastive alignment (C → D) gives the best generation scores of any configuration at a small cost to reconstruction. The large MuQEval jump in particular suggests that modelling musical structure has become easier. The low-Dt reference E attains the best reconstruction but trails A, C, and D on generation, validating high Dt paired with larger d (A vs. E).

5.2

Baselines

Tab. 1 compares SAME against recent open-weights continuous-latent audio autoencoders. Stable Audio Open (SAO) [8], a convolutional VAE at 44.1 kHz stereo. ϵar-VAE [9], a hybrid convolutional/transformer VAE at 44.1 kHz stereo, trained with IF/GD phase losses. CoDiCodec [13], a STFT-domain consistency autoencoder at 44.1 kHz stereo, supporting both continuous and discrete modes (we use the continuous mode). ACE-Step 1.5 [41], a convolutional VAE at 48 kHz stereo.

5.3

Objective Evaluation

For objective evaluation we use four complementary metrics. SI-SDR measures waveform-level fidelity. STFTlog1p is a multi-resolution L1 distance on log(1+|X|) magnitude spectrograms at six FFT sizes (128, 256, 512, 1024, 2048, 4096; Hann, 75% overlap), and MELlog1p applies the same distance to 64-band mel projections at three FFT sizes (1024, 2048, 4096), rescaled by Nmel /F so their magnitude scale matches the unprojected STFT. CCPC (cross-channel phase coherence) [9] measures stereo-image fidelity as the energy-weighted mean phasor coherence between reference and reconstructed inter-channel phase differences, averaged across four STFT resolutions (FFT 512, 1024, 2048, 4096; Hann, 75%). We also report end-to-end inference speed (FP16, single H100, no chunking or torch.compile), averaged

8

over 50 two-minute SDD tracks. Results are in Tab. 1. Both SAME variants are faster than every baseline. SAME-S runs 6–7× faster than the convolutional VAE baselines (ϵar-VAE, SAO, ACE-Step), whilst SAME-L runs around 2× faster despite its substantially higher parameter count. On objective audio-quality metrics, SAME-L and ϵar-VAE are the strongest: ϵar-VAE is marginally ahead on SI-SDR, STFTlog1p , and CCPC, while SAME-L is significantly ahead on MELlog1p . SAME-S performs comparably to SAO at greatly reduced computational cost.

5.4

Subjective Evaluation

We evaluate perceptual quality of SAME-L and SAME-S with a MUSHRA listening test, using 64 kbps MP3 as anchor and a hidden reference. Stimuli are random excerpts from the Song Describer Dataset. Participants were filtered per-exampled based on their ability to correctly rate both the anchor and the reference, resulting in 36 valid trials from 12 unique participants. CoDiCodec is omitted to reduce test length as whilst it sounds perceptually plausible, it is often clearly different from the input. Results appear as the final row of Tab. 1. SAME-L is rated highest, with ϵar-VAE second. Audio examples are available online3 .

6

Conclusion

We introduced SAME, a stereo music and general-audio autoencoder that reaches a 4096× temporal compression ratio while maintaining sound quality (as judged by objective and subjective evaluation), generative tractability and fast inference speed. The architecture pairs a transformer resampling backbone and a bottleneck with three auxiliary losses (flow-matching alignment, semantic regression, and cross-modal contrastive alignment) that jointly shape the latent space for downstream use. We demonstrate that these auxiliary losses are a mechanism by which a simpler bottleneck can outperform a VAE at high compression. Model weights for SAME-L and SAME-S are released at https://stability-ai.github.io/SAME.

7

Acknowledgements

The authors would like to thank Yin-Jyun Luo and Boris Kuznetsov, who both participated in constructive discussions at the beginning of this project.

References [1] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2022. [2] N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Trans. Audio, Speech, Language Process., vol. 30, pp. 495–507, 2022. [3] A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Trans. Mach. Learning Res., 2023. [4] R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved RVQGAN,” in Advances in Neural Inform. Process. Syst., 2023. [5] A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Advances in Neural Inform. Process. Syst., 2017. [6] Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour, “AudioLM: A language modeling approach to audio generation,” IEEE/ACM Trans. Audio, Speech, Language Process., vol. 31, pp. 2523–2533, 2023. [7] J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” in Advances in Neural Inform. Process. Syst., 2023. 3 https://stability-ai.github.io/SAME

9

[8] Z. Evans, J. D. Parker, C. J. Carr, Z. Zukowski, J. Taylor, and J. Pons, “Stable Audio Open,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Process., 2025. [9] K. Wang, Z. Wu, D. Zhou, R. Lin, J. Dai, and T. Jiang, “Back to ear: Perceptually driven high fidelity music reconstruction,” arXiv preprint arXiv:2509.14912, 2025. [10] S. Ahn, B. J. Woo, M. H. Han, C. Moon, and N. S. Kim, “HILCodec: High-fidelity and lightweight neural audio codec,” IEEE J. Sel. Topics Signal Process., vol. 18, no. 8, pp. 1517–1530, 2024. [11] M. Pasini, S. Lattner, and G. Fazekas, “Music2Latent: Consistency autoencoders for latent audio compression,” in Proc. Int. Soc. Music Inform. Retrieval Conf., 2024. [12] ——, “Music2Latent2: Audio compression with summary embeddings and autoregressive decoding,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Process., 2025. [13] ——, “CoDiCodec: Unifying continuous and discrete compressed representations of audio,” in Proc. Int. Soc. Music Inform. Retrieval Conf., 2025. [14] J. D. Parker, A. Smirnov, J. Pons, C. J. Carr, Z. Zukowski, Z. Evans, and X. Liu, “Scaling transformers for low-bitrate high-quality speech coding,” in Proc. Int. Conf. Learning Representations, 2025. [15] H. Wu, N. Kanda, S. E. Eskimez, and J. Li, “TS3-Codec: Transformer-based simple streaming single codec,” in Proc. Interspeech, 2025. [16] D. Yang, S. Liu, H. Guo, J. Zhao, Y. Wang, H. Wang, Z. Ju, X. Liu, X. Chen, X. Tan, X. Wu, and H. Meng, “ALMTokenizer: A low-bitrate and semantic-rich audio codec tokenizer for audio language modeling,” in Proc. Int. Conf. Machine Learning, 2025. [17] X. Zhang, D. Zhang, S. Li, Y. Zhou, and X. Qiu, “SpeechTokenizer: Unified speech tokenizer for speech language models,” in Proc. Int. Conf. Learning Representations, 2024. [18] A. Défossez, L. Mazaré, M. Orsini, A. Royer, P. Pérez, H. Jégou, E. Grave, and N. Zeghidour, “Moshi: a speech-text foundation model for real-time dialogue,” arXiv preprint arXiv:2410.00037, 2024. [19] Z. Du, S. Zhang, K. Hu, and S. Zheng, “FunCodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Process., 2024. [20] Q. Yu, M. Weber, X. Deng, X. Shen, D. Cremers, and L.-C. Chen, “An image is worth 32 tokens for reconstruction and generation,” in Advances in Neural Inform. Process. Syst., 2024. [21] A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in Proc. Int. Conf. Machine Learning, 2021. [22] T. Ye, L. Dong, Y. Xia, Y. Sun, Y. Zhu, G. Huang, and F. Wei, “Differential Transformer,” in Proc. Int. Conf. Learning Representations, 2025. [23] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “RoFormer: Enhanced transformer with Rotary Position Embedding,” Neurocomputing, vol. 568, p. 127063, 2024. [24] J. Zhu, X. Chen, K. He, Y. LeCun, and Z. Liu, “Transformers without normalization,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition, 2025. [25] Google, “LiteRT: On-device runtime for cross-platform machine learning inference,” https://ai.google. dev/edge/litert, 2024, accessed 2026. [26] B. Zheng, N. Ma, S. Tong, and S. Xie, “Diffusion transformers with representation autoencoders,” arXiv preprint arXiv:2510.11690, 2025. [27] J. Heek, E. Hoogeboom, T. Mensink, and T. Salimans, “Unified latents (UL): How to train your latents,” arXiv preprint arXiv:2602.17270, 2026. [28] R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on Generative Adversarial Networks with multi-resolution spectrogram,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Process., 2020. 10

[29] S. Schwär and M. Müller, “Multi-scale spectral loss revisited,” IEEE Signal Process. Lett., vol. 30, pp. 1712–1716, 2023. [30] A. Jolicoeur-Martineau, “The relativistic discriminator: A key element missing from standard GAN,” in Proc. Int. Conf. Learning Representations, 2019. [31] T. Q. Nguyen, “Near-perfect-reconstruction pseudo-QMF banks,” IEEE Trans. Signal Process., vol. 42, no. 1, pp. 65–76, 1994. [32] H. Wang, S. Suri, Y. Ren, H. Chen, and A. Shrivastava, “LARP: Tokenizing videos with a learned autoregressive generative prior,” in Proc. Int. Conf. Learning Representations, 2025. [33] Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, and M. Nickel, “Flow Matching for generative modeling,” in Proc. Int. Conf. Learning Representations, 2023. [34] A. Cohen, I. Daubechies, and J.-C. Feauveau, “Biorthogonal bases of compactly supported wavelets,” Comm. Pure Appl. Math., vol. 45, no. 5, pp. 485–560, 1992. [35] B. Zhang, F. Moiseev, J. Ainslie, P. Suganthan, M. Ma, S. Bhupatiraju, F. Lebron, O. Firat, A. Joulin, and Z. Dong, “Encoder-decoder Gemma: Improving the quality-efficiency trade-off via adaptation,” arXiv preprint arXiv:2504.06225, 2025. [36] K. Liang, L. Chen, B. Liu, and Q. Liu, “Cautious optimizers: Improving training with one line of code,” arXiv preprint arXiv:2411.16085, 2024. [37] Z. Evans, J. D. Parker, C. J. Carr, Z. Zukowski, J. Taylor, and J. Pons, “Long-form music generation with latent diffusion,” in Proc. Int. Soc. Music Inform. Retrieval Conf., 2024. [38] I. Manco, B. Weck, S. Doh, M. Won, Y. Zhang, D. Bogdanov, Y. Wu, K. Chen, P. Tovstogan, E. Benetos, E. Quinton, G. Fazekas, and J. Nam, “The Song Describer Dataset: a corpus of audio captions for music-and-language evaluation,” in Machine Learning for Audio Workshop, NeurIPS, 2023. [39] A. Gui, H. Gamper, S. Braun, and D. Emmanouilidou, “Adapting Fréchet Audio Distance for generative music evaluation,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Process., 2024. [40] D. Zhu and Z. Li, “MuQ-Eval: An open-source per-sample quality metric for AI music generation evaluation,” arXiv preprint arXiv:2603.22677, 2026. [41] J. Gong, Y. Song, W. Zhao, S. Wang, S. Xu, J. Guo, and X. Yang, “ACE-Step 1.5: Pushing the boundaries of open-source music generation,” arXiv preprint arXiv:2602.00744, 2026.

11

Record · ID 200533 · SHA-256 f73a6c7b94e1b9e5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.