Conceptio › Archive › arXiv CS
arXiv CSopen access

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Haocheng Xi1,∗ , Yiming Xie2 , Hexu Zhao2 , Yiwen Zhang2 , Michael Liu1,2 , Thomas Creavin2 Kurt Keutzer1 , Xiuyu Li2 , Zhaoyang Lv2 , Chenfeng Xu3 , Haiwen Feng1,2 1

University of California, Berkeley

2

Impossible, Inc.

3

University of Texas at Austin

arXiv:2609.20744v1 [cs.LG] 17 Sep 2026

Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5× speedup over the 50-step dense H3 baseline on the same GPU count. GitHub: https://github.com/OpenVDN/vdn-minimax-h3 Weights: https://huggingface.co/OpenVDN/vdn-minimax-h3 Blog: https://openvdn.github.io/

First–last-frame-to-video · 14.3 s · 768p Dense H3 50 steps

VDN-H3 8 steps

Attention latency (s)

Dense H3

VDN

14.3 s video · 768p · DiT denoising Dense H3

30

1× B200 · 50 NFE

10

+ VDA

3.38× 3.38×

20

1.77× 1.77×

2.60× 2.60×

1× B200 · 50 NFE

2.57× 2.57×

6.25× 6.25×

+ few-step

5.9

8.0

10.1

12.3

6.70 s

8× B200 · 8 NFE

14.4

0

Video duration (s)

16.2× speedup · 1 GPU

7.36× 7.36×

+ distributed

0

307.9 s

49.3 s

1× B200 · 8 NFE

799.6 s

119.3× speedup · 8 GPUs 200

400

600

800

Denoising time (s)

(a) Speedup grows with video duration.

(b) Progressive acceleration of VDN-H3.

Figure 1 Quality and efficiency of VDN-H3. Top: matched video frames. Bottom: attention scaling and cumulative

denoising acceleration. * Part of the work done during an internship at Impossible, Inc.

1

Preprint

1

I NTRODUCTION

Attention is the dominant computational cost in long-sequence video generation. In the profiled MiniMax H3 workload, Softmax attention accounts for more than 85% of denoiser runtime. Dense Softmax gives every query direct access to the complete key sequence, but the resulting pairwise computation grows quadratically with sequence length. Frontier language models increasingly use recurrent linear attention to avoid this scaling bottleneck. Applying the same idea to video is appealing: distant context can be compressed into a fixed-size state, making its cost linear in the number of tokens. A direct replacement, however, meaningfully degrades generation quality because the compressed state cannot preserve all of the fine-grained interactions available to Softmax. Closing this quality gap requires addressing three mismatches. First, a fixed-size state must preserve global properties such as subject identity, scene layout, appearance, and long-range motion even though its capacity does not grow with sequence length, whereas Softmax retains an expanding set of keys and values. Second, delta-rule linear attention typically updates its recurrent state one token at a time, mirroring autoregressive language-model decoding. Video diffusion instead processes all spatial tokens in a frame together; imposing an arbitrary patch order is unnatural, while treating correlated writes independently can make them interfere. Third, adding a randomly initialized linear pathway to a Softmax-pretrained model changes both its information flow and residual-stream activation statistics. Without careful adaptation, it can disrupt capabilities learned during pretraining before the new branch becomes useful. We introduce Video DeltaNet (VDN), a hybrid attention architecture designed around these challenges. VDN retains Softmax for local video interactions and global boundary anchors, while bidirectional linear memory represents distant video context. The anchors keep the beginning and end of the clip directly accessible as global references. The linear branch uses Video Delta Attention (VDA), a video-native delta rule that updates memory once per frame by jointly incorporating its spatial tokens and key correlations. RMS normalization stabilizes the scale of the linear readout, while separate gates and output projections calibrate each branch before their outputs are combined. A staged adaptation recipe first aligns the new linear pathway with the pretrained teacher, then uses low-rank updates to co-adapt the hybrid layer while preserving the pretrained backbone. We instantiate the approach on MiniMax H3 to obtain VDN-H3. In collaboration with the SGLang team, we optimize its inference path. With eight-step distillation, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs. This corresponds to a 14.5× reduction relative to the 50-step dense H3 baseline on the same GPU count. On a fixed third-party benchmark, eight-step VDN-H3 matches or slightly exceeds 50-step Dense H3 across video-quality metrics, preserves comparable motion magnitude and first–last-frame conditioning fidelity, and maintains a clear margin over four-step FastH3. Our main contributions are: 1. A hybrid video attention architecture that combines local Softmax and global boundary anchors with bidirectional linear memory, together with branch-specific normalization, gates, and output projections. 2. A frame-wise delta operator that jointly incorporates a frame’s spatial tokens, with an analysis of the inherited-state transition and an efficient batched implementation. 3. A complete adaptation recipe for pretrained video models that combines staged teacher alignment, low-rank refinement, few-step distillation, and optimized inference without training a new foundation model from scratch.

2

V IDEO D ELTA N ET: H YBRID ATTENTION A RCHITECTURE

Video DeltaNet divides video-to-video attention by temporal role. Nearby frames retain explicit token-to-token Softmax attention, while distant video context is summarized by linear attention. Below, we introduce the two branches and explain how their outputs are combined. 2

Preprint

Text / text state Text

Local video

Video: 10 VAE chunks × 2

Boundary anchor Audio

T

Text

A

Distant video

Video: 10 VAE chunks × 2

T

Audio A

T

Query

Query

T

Audio

A

A

Key

Key

(a) Softmax

(b) Linear

Figure 2: Softmax and Linear attention allocation. The schematic uses one text token, twenty video tokens in ten two-frame chunks, and one audio token.

2.1

S LIDING -W INDOW S OFTMAX ATTENTION

Nearby frames contain local correspondences that determine texture, object boundaries, and shortterm motion. These interactions benefit from direct token-to-token matching across neighboring frames. VDN therefore retains exact Softmax attention within a bidirectional temporal window to preserve quality, while assigning distant interactions to linear attention. The window follows the video tokenizer. H3’s VAE decodes five consecutive latent frames as one temporal chunk, so each query chunk attends to itself and its immediately preceding and following chunks. This prevents the Softmax boundary from cutting through the tokenizer’s natural temporal unit, producing a 15-frame window except at sequence boundaries. VDN also adds two boundary anchors with four-way connectivity. Every video frame attends to all tokens in the first and last latent frames, and the first and last frames attend to the complete sequence. This pattern is particularly natural for full-clip video diffusion: the two boundary anchors provide explicit global references from opposite ends of the clip, while only two frame rows and columns receive dense connectivity. The same anchors are valuable for image-to-video and first– last-frame-to-video generation, where the provided visual conditions directly constrain the generated sequence. Anchor entries already inside a local window are included only once. 2.2

B IDIRECTIONAL L INEAR ATTENTION

Bidirectional linear attention handles the remaining video context through two temporal scans. For a query frame t, the forward state St→ summarizes frames before its Softmax window, and the reverse state St← summarizes frames after the window. Boundary anchors are excluded from both states because they are already available through Softmax. The two temporal regions are disjoint, so their query readouts can be added without counting any video frame twice. Each scan first builds frame states along its own direction. At readout time, VDN gathers the state immediately outside the corresponding window boundary, applies the accumulated channel-wise decay across the skipped local span, and evaluates the query against the resulting memory. The distant-context output is the sum of the two readouts: 3

Preprint

→ ← oL t = St qt + St qt .

(1a)

The linear memory is also text-aware. Before scanning the video sequence, VDN summarizes all text tokens into a state ST and initializes each directional scan with ST /2. Summing the two readouts therefore counts the prompt exactly once, providing the linear branch with global text conditioning while text remains directly visible to the Softmax branch: S0→ = S0← = 12 ST , 2.3

 S0→ + S0← q = ST q.

(1b)

C OMBINING O UTPUTS FROM T WO B RANCHES

Both branches receive their query, key, and value inputs from the pretrained QKV projections. The Softmax branch keeps H3’s QK normalization and rotary position processing. Following Gated DeltaNet (Yang et al., 2025b) and Kimi Delta Attention (Kimi Team, 2025), the linear branch then applies its own feature map: a separable short convolution to K and V, followed by SiLU, with L2 normalization on Q and K. The released K/V convolution consists of a depthwise 5 × 5 spatial filter and a five-tap temporal filter. No rotary embedding is added in the linear branch. The two branches have different output scales and therefore require calibration. Restricting Softmax to local windows and boundary anchors concentrates its probability mass over fewer keys, so VDN applies a content-dependent sigmoid gate to the Softmax readout. The linear readout is RMS-normalized and passed through its own sigmoid output gate. Following Gated DeltaNet and Kimi Delta Attention, separate decay and write gates control memory retention and frame updates inside the recurrence. Each branch also has its own output projection. Let the gated branch outputs be e S = GS ⊙ OS , O

e L = GL ⊙ RMSNorm(OL ). O

(1c)

The two branches are then projected independently and added: eS W S + O eL W L . Y =O O O

(1d)

Separate output projections allow the branches to contribute in different residual-stream directions. The combined output is then fed into the pretrained residual stream, while the original feed-forward sublayer is left unchanged.

3

V IDEO D ELTA ATTENTION : F RAME - WISE D ELTA RULE

Video Delta Attention (VDA) extends the delta rule from individual tokens to entire video frames. It jointly updates the recurrent state from all key-value pairs in a frame, allowing interactions among frame tokens to shape the write rather than accumulating independent corrections. We first review the standard token-wise rule and then derive the frame-wise update. 3.1

P RELIMINARIES OF L INEAR ATTENTION

Delta-rule memory reads a value associated with a key and writes a correction proportional to the prediction residual. Gated DeltaNet combines this mechanism with memory decay (Yang et al., 2025b), while Kimi Delta Attention introduces finer-grained decay (Kimi Team, 2025). At step t, for a state St−1 ∈ Rdv ×dk , a key kt ∈ Rdk , and a value vt ∈ Rdv , the familiar one-token update is S̄t = St−1 Diag(αt ), St = S̄t + βt (vt − S̄t kt )kt⊤ . 4

(2)

Preprint

Output

Output Text state / 2

Projection

Video Delta Attention

Projection

Forward scan + Reverse scan

Q

σ

K

V

Conv

Conv

Linear

Linear

L2

σ

Decay

RMSNorm

Softmax

α

Linear

Linear

β

L2

Input

σ

Input

(a) Hybrid attention

(b) Linear branch

Figure 3: Video DeltaNet architecture. Softmax and Linear outputs are independently gated and projected; the Linear branch uses bidirectional Video Delta Attention.

The decay gate αt controls how much inherited memory remains, while βt controls the strength of the erase-and-write correction. This token-wise recurrence is well matched to autoregressive decoding, where one new token arrives at each step. Video diffusion presents a different unit of computation: all spatial tokens of a latent frame are available together. Sequentially imposing a patch order is unnecessary, so a natural first adaptation is to compute their delta corrections in parallel from the same decayed state. SANA-WM follows this batched construction and adds frame-size key scaling to stabilize the resulting additive transition (Zhu et al., 2026). Let a video contain F latent frames with U = Hℓ Wℓ tokens each. For token u of frame t, denote its key, value, and write gate by kt,u ∈ Rdk , vt,u ∈ Rdv , and βt,u ≥ 0. A frozen-state update adds all token corrections evaluated at S̄t = St−1 Diag(αt ):

Stbatch = S̄t +

U X

⊤ βt,u (vt,u − S̄t kt,u )kt,u .

(3a)

u=1

Stack the keys and values as Kt ∈ RU ×dk and Vt ∈ RU ×dv , and write βt = (βt,1 , . . . , βt,U )⊤ . The two frame statistics are

At =

U X

⊤ βt,u kt,u kt,u = Kt⊤ Diag(βt )Kt ,

u=1

Bt =

U X

(3b) ⊤ βt,u vt,u kt,u = Vt⊤ Diag(βt )Kt .

u=1

Here At summarizes key correlations and Bt value–key writes, giving Stbatch = S̄t (I − At ) + Bt .

(3c)

Because every residual in Equation (3a) uses the same S̄t , overlapping keys can produce conflicting writes without accounting for one another (Appendix A.2). 5

Preprint 3.2

V IDEO D ELTA ATTENTION : A F RAME - WISE U PDATE

This independence becomes problematic when several patches address similar key directions. Their corrections can reinforce or conflict even though they are meant to describe one frame. VDA instead lets all spatial tokens determine a single new state together, so overlapping directions are resolved inside the update while previously accumulated memory remains a reference. The classical one-token delta rule can be viewed as one gradient step on its prediction error. Rather than taking one such step independently for every patch, VDA defines the frame-level state as the solution to a joint objective:

St = arg min S

U 1X 1 ∥S − S̄t ∥2F + βt,u ∥Skt,u − vt,u ∥22 . 2 2 u=1

(4)

The first term keeps the new memory close to the decayed state. The second asks that same state to fit all key–value associations in the frame simultaneously. Differentiating once gives a compact normal equation and closed-form update: St (I + At ) = S̄t + Bt ,

St = (S̄t + Bt )(I + At )−1 .

(5)

In contrast to Equation (3c), each residual in the joint solution is evaluated at the shared updated state St . The inverse couples the writes: when two patch keys overlap, their inner product affects both effective corrections. We compute this inverse in the dk × dk key-channel space; Appendix A.2 illustrates the effect of key correlations. 3.3

S TABILITY AND C ORRELATION AWARENESS OF V IDEO D ELTA ATTENTION

VDA has a stable inherited-state transition: the contribution carried from earlier frames cannot be amplified by a frame update when the prepared features and gates are fixed. Proposition 1 — Non-expansive inherited-state transition. For fixed prepared features and gates, with βt,u ≥ 0 and 0 ≤ αt,j ≤ 1, the transition Mt = Diag(αt )(I + At )−1 satisfies ∥Mt ∥2 ≤ 1. Thus changing the entering state by ∆S changes its carried contribution by at most ∥∆S∥F . Without frame-size scaling, additive frame-wise writes can amplify inherited state; SANA-WM √ therefore scales its keys by 1/ U (Zhu et al., 2026). VDA obtains non-expansiveness directly from (I + At )−1 . Appendix A.1 proves the proposition, and Appendix A.2 compares the two rules under different key correlations.

4

E ND - TO -E ND T RAINING P IPELINE

We first describe the three-stage adaptation that integrates the Linear branch into the pretrained Softmax backbone, followed by few-step distillation from 50 to eight denoising steps. 4.1

S TAGED A RCHITECTURE A DAPTATION

Architecture adaptation proceeds in three stages. Figure 4 shows the trainable components in each stage; the pretrained base weights remain frozen throughout. A1 - Per-layer alignment. Each Linear branch is initialized independently from frozen pretrained activations for 200 steps. This local objective avoids sending the initialization signal through a deep stack of simultaneously changing hybrid blocks. The backbone is frozen, the Softmax gate is fixed at its 0.99 initialization, and gradients are clipped independently for each layer at 0.1. A2 - End-to-end alignment. The calibrated branches are then installed together and optimized end to end for 500 steps. This stage corrects composition errors that are invisible when blocks are trained in isolation. The pretrained weights and Softmax gates remain frozen, and gradient clipping is applied globally at 1.0. 6

Preprint Frozen

Trainable

Starting point

Stage A1 per-layer / Stage A2 end-to-end

Stage B · Co-adaptation

Output head

MLP

Output head

MLP

Full Attention

MLP Softmax gate

Linear gate

Sliding-window

Video Delta Attention

Softmax

Hybrid Attention + QKVO LoRA

Embedding

Shared QKV

Embedding

(a) MiniMax H3

(b) Hybrid attention layer

(c) VDN-H3

Figure 4: Progressive adaptation of the pretrained H3 backbone. (a) The dense model is the frozen starting point. (b) Stage A1 aligns each Linear branch independently, then Stage A2 trains the assembled hybrid layers end to end while the pretrained weights and Softmax gates remain frozen. (c) Stage B jointly trains the Linear pathway, Softmax gates, and QKVO LoRA adapters.

Stage B - LoRA co-adaptation. We add LoRA adapters to the Q, K, V, and output projections and train them jointly with the Linear pathway and Softmax gates for 2,000 steps. 4.2

F EW-S TEP D ISTILLATION

After architecture adaptation, we distill the 50-step VDN-H3 into an eight-step sampler using a DMD2-style objective without the GAN term (Yin et al., 2024b). The student is initialized from the community MiniMax-H3-Turbo-LoRA (larryvrh, 2026) and trained against VDN-H3’s own 50step sampler, isolating step reduction from architecture conversion. Generator, real-score, and fakescore roles share one FSDP backbone through separate adapters, with three fake-score updates per generator update. The released checkpoint is trained for 250 generator steps.

5

E FFICIENT I NFERENCE

Fused VDA kernels. We organize VDA’s data preparation and readout into four fused Triton kernels. VDA-Prep combines temporal convolution, SiLU, L2 normalization, and the frame-major layout conversion for K and V. VDA-Stats constructs the per-frame statistics At and Bt in one pass, sharing input reads and reduction work. VDA-Gather collects the two directional states at local-window boundaries and applies the decay bridge. VDA-Epilogue combines RMS normalization, output gating, and the final layout conversion. QK normalization and rotary embedding are fused separately on the Softmax path. Chunk-wise scans. VDA defines an affine transition per frame, but attention reads memory only at VAE-chunk boundaries. We therefore compose each chunk into Sout = Sin Mchunk + Jchunk and scan the shorter chunk sequence in both directions with a single kernel launch. This preserves the required boundary states while reducing scan depth and launch overhead by roughly the chunk size. The prompt state is included as a leading virtual frame. Small-matrix inverse. Each frame requires computing (I + At )−1 . A batched Cholesky pipeline incurs several kernel launches and intermediate memory reads and writes, so we use one CUDA kernel that performs blocked Gauss–Jordan elimination in registers and directly emits the transition and injection terms. The inverse and recurrent state updates remain in FP32. Other optimizations. The released VDN-H3 model is served through SGLang. Window Softmax packs queries by visible-key pattern to use FlashAttention’s varlen API without a global mask. Following Ulysses (Jacobs et al., 2023), VDA is sharded by attention head and executed on a side stream 7

Preprint Dense H3 (50 NFE) 30

28.03

27.54

92

23.11 20

70

VQAA 2.20

Instruction Following

86

Perceptual Quality

VDN-H3 (8 NFE) 83.41

82.68

Q-Align

75

3.5

FAST-VQA

0.52

RAFTFlowMean

.88

28.85 1.77

8

27.59

0

.826

.183

.785

.104

1.65* 1.5

World Coherence

25

FL2VA PSNR

.70

DOVER++

.30

.833

28.67

3.28

3.22

11.71

11.55 9.19

30

1.84

13

78.02

2.0

4.45 4.22*

4.0

85

89.18

VQAT 4.40

2.25

FastH3 (4 NFE) 93.13

92.22

4.6

2.25

95

75.46

2.4

2.0

89.20

88.20

FL2VA SSIM

0

.116

FL2VA LPIPS

Figure 5: Video quality, motion, FIRM-Video dimensions, and FL2VA endpoint fidelity across Dense H3, FastH3, and VDN-H3. FL2VA metrics average the first and last conditioning frames; * denotes a reported significant difference from Dense H3. alongside window Softmax. MXFP8 accelerates the wide QKV, output, and feed-forward GEMMs, while recurrent states and small-matrix inverses remain in FP32. AdaLN modulation parameters are precomputed before the block loop to avoid repeating their projection inside each transformer block.

6

E XPERIMENTS

6.1

S ETTINGS

Data. We curated a training set of 10,015 video clips at 1344 × 768 resolution, each with 345 frames at 24 fps (14.375 seconds). Each sample is pre-encoded and cached as video latents of shape (24, 102, 48, 84), stereo audio latents of shape (2, 32, 575), and Qwen3-VL text embeddings of shape (L, 5120) with token-type tags (Bai et al., 2025). This removes the VAEs and text encoder from the training loop. For evaluation, we use 103 prompts from a fixed third-party set. All models render the same prompts at the same resolution and duration, without model-specific prompt selection. Baselines. Dense H3 with 50 neural function evaluations (NFEs) is the full-attention baseline, while FastH3 with four NFEs provides a fast-model baseline (FastVideo Team, 2026). Quality and qualitative comparisons use the final eight-step VDN-H3. We report 50-step VDN-H3 only in the efficiency study, where it isolates the speedup from hybrid attention before step distillation. Metrics. We report EvalCrafter VQAA and VQAT (Liu et al., 2024), Q-Align (Wu et al., 2024), FAST-VQA (Wu et al., 2022), and DOVER++ overall (Wu et al., 2023). Higher is better for all five; Q-Align, FAST-VQA, and DOVER++ use a ×100 scale. We additionally report RAFT mean flow magnitude in pixels as a motion diagnostic (Teed & Deng, 2020). FIRM-Video evaluates Instruction Following, Perceptual Quality (PQ), and World Coherence (WC) on a 1–5 scale (Zhang et al., 2026d). For first–last-frame-to-video (FL2VA), we measure PSNR, SSIM (Wang et al., 2004), and LPIPS (Zhang et al., 2018) at both conditioning frames and report their mean. Environment. We use PyTorch 2.13 and CUDA 12.9 on NVIDIA H200 and B200 clusters. Training uses FSDP2/HSDP, activation checkpointing, and pinned-memory activation offload. Parameters are gathered in bf16 and gradients reduced in FP32, except that the decay-α modules remain FP32. Additional optimization and training details are provided in Appendix A.4. 6.2

Q UALITY R ESULTS

Figure 5 summarizes overall video quality, motion, and conditional endpoint fidelity. Across the five no-reference quality metrics, eight-step VDN-H3 matches or exceeds 50-step Dense H3, with differences ranging from +0.06 to +1.00; FastH3 is 2.70–12.74 points lower than Dense H3. This agreement holds across aesthetic, technical, and learned perceptual metrics. VDN-H3 also preserves a similar RAFT motion magnitude (11.71 versus 11.55 pixels), while FastH3 falls to 9.19 8

Preprint Table 2: B200 efficiency. Gain is measured against the preceding row; speedup is cumulative against single-GPU Dense H3. The video-duration sweep uses one GPU and excludes distributed inference. Configuration

#GPU

NFE

s/NFE

End-to-end latency

Gain

Speedup

Latent frames

Dense H3 VDN-H3 + few-step + distributed

1 1 1 8

50 50 8 8

15.99 s 6.16 s 6.16 s 0.85 s

799.6 s 307.9 s 49.3 s 6.70 s

– 2.6× 6.3× 7.4×

1.0× 2.6× 16.2× 119.3×

Attention density 42.08% 26.82% 22.85% 19.98% Attention speedup 1.7× 2.3× 2.7× 3.0× End-to-end latency 19.0 s 34.1 s 41.6 s 49.3 s End-to-end speedup 10.3× 13.2× 14.7× 16.2×

42

72

87

102

Table 3: H200 efficiency. Gain is measured against the preceding row; speedup is cumulative against single-GPU Dense H3. The video-duration sweep uses one GPU and excludes distributed inference. Configuration

#GPU

NFE

s/NFE

End-to-end latency

Gain

Speedup

Latent frames

Dense H3 VDN-H3 + few-step + distributed

1 1 1 8

50 50 8 8

35.35 s 11.16 s 11.16 s 1.56 s

1767.3 s 557.9 s 89.3 s 12.5 s

– 3.2× 6.3× 7.1×

1.0× 3.2× 19.8× 141.6×

Attention density 42.08% 26.82% 22.85% 19.98% Attention speedup 2.0× 3.0× 3.5× 4.0× End-to-end latency 34.9 s 62.2 s 76.3 s 89.3 s End-to-end speedup 11.4× 15.6× 17.7× 19.8×

42

72

87

102

pixels. FIRM-Video shows the same pattern: VDN-H3 matches Dense H3 in Instruction Following (2.25), is slightly higher in Perceptual Quality (4.45 versus 4.40), and remains comparable in World Coherence (1.77 versus 1.84), whereas FastH3 is lower on all three dimensions. For FL2VA, VDN-H3 remains close to Dense H3, with gaps of 0.18 dB in PSNR (28.67 versus 28.85), 0.007 in SSIM (0.826 versus 0.833), and 0.011 in LPIPS (0.1156 versus 0.1044; lower is better). FastH3 shows a larger loss of endpoint fidelity: relative to Dense H3, its PSNR and SSIM decrease by 1.26 dB and 0.048, while LPIPS increases by 0.079. 6.3

E FFICIENCY R ESULTS

Backbone acceleration. We first compare one complete transformer evaluation at the same sequence length and NFE count. For the 14.3-second, 768p workload, the system-optimized VDN-H3 backbone reduces latency from 35.35 to 11.16 seconds on one H200 (3.2×), and from 16.0 to 6.2 seconds on one B200 (2.6×). Scaling with video length. Dense Softmax scores every video-token pair, so its video–video workload grows quadratically with the number of latent frames. VDN keeps only a fixed-width local window and the first- and last-frame anchors in Softmax, while VDA carries the remaining long-range context with linear scaling. As the sequence grows from 42 to 102 latent frames, the measured Softmax attention density falls from 42.1% to 20.0%. Over the same range, whole-backbone speedup increases from 1.8× to 3.2× on H200 and from 1.7× to 2.6× on B200. Sampling and distributed inference. Starting from the optimized backbone, eight-step distillation reduces one-B200 denoising from 307.9 to 49.3 seconds, and eight-GPU head-sharded inference brings the final latency to 6.7 seconds on B200 and 12.5 seconds on H200. Tables 2 and 3 report the complete deployment path and duration scaling on both accelerators. 6.4

A BLATION S TUDIES

We profile VDN’s core backbone operators at 102 latent frames. Figure 6 compares their direct and optimized implementations on H200 and B200, separating operator-level gains from few-step distillation and distributed inference. Fused VDA kernels. VDA-Prep reduces latency from 18.0 to 1.6 ms on H200 and from 17.1 to 3.4 ms on B200. VDA-Stats provides 2.1× and 2.9× speedups, VDA-Gather provides 7.2× and 7.5×, and VDA-Epilogue provides 7.3× and 9.1× on H200 and B200, respectively. Chunk-wise scans. Composing frame transitions adds overhead with all 56 heads on one GPU, but becomes effective after head sharding. With seven heads per rank in 8-GPU inference, scan latency falls from 4.6 to 1.1 ms on H200 and from 3.0 to 0.8 ms on B200. 9

Preprint Original

22

ms

16

18.02

8

H200

B200

H200

5.5

5.11

4.56 3.29

3.1

H200

0.28 0

B200

B200

9

4.58

(e) Chunk scan, 1 GPU

6.93

1.09

0.21

H200

B200

0

H200

0.76 B200

(d) VDA-Epilogue

ms

130

7.75

ms

112.6112.1

6.32 3.02

2.75

0

7.96

(c) VDA-Gather

ms

1.08 0

ms

5

(b) VDA-Stats

ms

4.58

1.2

3.99 0

10

2.01 1.57

6.45

(a) VDA-Prep 6.2

ms

11.39

3.37

1.58 0

2.4

13.36

17.07

11

Optimized

ms

H200

4.5

1.69

0.82 0

B200

(f) Chunk scan, 8-GPU

69.5

65

H200

56.1

1.26 B200

(g) Matrix inverse

0

H200

B200

(h) Window Softmax

Figure 6: Kernel-level optimization ablations on H200 and B200. Small-matrix inverse. Replacing the multi-kernel Cholesky path with the fused inverse reduces latency from 7.8 to 1.7 ms on H200 and from 6.3 to 1.3 ms on B200, a 4.6–5.0× speedup. Other optimizations. Figure 6 additionally isolates Window Softmax. It improves from 69.5 to 56.1 ms on B200, while H200 remains nearly unchanged at 112.6 versus 112.1 ms. 6.5

G ALLERY

In Figure 7, we present diverse examples comparing 50-step Dense H3 and eight-step VDN-H3 across visual styles, scene transitions, and motion patterns. The results demonstrate that VDN-H3 achieves visual quality on par with Dense H3 despite using substantially fewer denoising steps.

7

R ELATED W ORK

Efficient video attention algorithms. Recent video diffusion transformers process increasingly long spatiotemporal sequences (HaCohen et al., 2024; Jin et al., 2024; Kong et al., 2024; Polyak et al., 2024; Wan et al., 2025; Yang et al., 2025c; MiniMax, 2026). Attention acceleration spans optimized exact kernels (Dao et al., 2022; Dao, 2023; Shah et al., 2024) and low-precision kernels, including SageAttention (Zhang et al., 2025b), SageAttention2 (Zhang et al., 2024), and SageAttention3 (Zhang et al., 2025a). General-purpose sparse methods include SpargeAttn (Zhang et al., 2025d) and SpargeAttention2 (Zhang et al., 2026b); video-specific methods include Sparse VideoGen (Xi et al., 2025) and Sparse VideoGen2 (Yang et al., 2025a). SLA and SLA2 combine sparse and linear computation (Zhang et al., 2025c, 2026a). Quantized diffusion methods reduce the cost of DiT execution (Zhao et al., 2024; Li et al., 2025, 2026; Xue et al., 2026), while Quant VideoGen targets the recurrent cache of autoregressive video models (Xi et al., 2026). Linear attention and recurrent memory. Linear Transformers and Performers introduced influential linear-complexity alternatives to dense attention (Katharopoulos et al., 2020; Choromanski et al., 2021), while Mamba and Mamba-2 connected selective state-space models with efficient sequence processing (Gu & Dao, 2023; Dao & Gu, 2024). Gated Linear Attention (Yang et al., 2024b), DeltaNet (Yang et al., 2024a), Gated Delta Networks (Yang et al., 2025b), Kimi Linear (Kimi Team, 2025), and Gated DeltaNet-2 (Hatamizadeh et al., 2026) support data-dependent memory updates. In vision, VideoMamba applies state-space modeling to long-video understanding, and M4V develops a multimodal Mamba backbone for text-to-video generation (Li et al., 2024b; Huang et al., 2026). LoGeR similarly combines sliding-window attention with parametric long-term memory to preserve local geometry and global consistency over long video sequences (Zhang et al., 2026c). Linear and hybrid diffusion backbones include SANA (Xie et al., 2025), SANA-Video (Chen et al., 2025), SANA-Video 2.0 (Chen et al., 2026), and SANA-WM (Zhu et al., 2026). We also acknowl10

Preprint

Dense H3 VDNH3

(a) Cinematic Action

Dense H3 VDNH3

(b) Surreal Journey

Dense H3 VDNH3

(c) Scene Transition

Dense H3 VDNH3

(d) Graphic Animation

Dense H3 VDNH3

(e) Animated Sports

Figure 7: Qualitative gallery. Six synchronized frames from five prompts, with 50-step Dense H3 above eight-step VDN-H3 in each pair.

edge Reflections on Video DeltaNet as concurrent analysis of frame-level delta updates (Zhu, 2026). Few-step distillation and distributed inference. Progressive distillation, consistency models, latent consistency models, and adversarial distillation reduce the number of denoising evaluations (Salimans & Ho, 2022; Song et al., 2023; Luo et al., 2023; Sauer et al., 2023). Distribution Matching Distillation (Yin et al., 2024a) and DMD2 (Yin et al., 2024b) provide closely related objectives; AnimateLCM and T2V-Turbo extend few-step training to video (Wang et al., 2024; Li et al., 2024a). Multi-GPU inference is complementary: Ulysses and USP shard long attention sequences (Jacobs et al., 2023; Fang & Zhao, 2024), DistriFusion and PipeFusion exploit patch and pipeline parallelism across diffusion steps (Li et al., 2024c; Fang et al., 2024a), and xDiT composes multiple 11

Preprint parallel strategies in one inference engine (Fang et al., 2024b). StreamDiffusionV2 further targets low-latency distributed video streaming (Feng et al., 2026).

8

C ONCLUSION

Video DeltaNet accelerates video generation while largely preserving Dense H3 quality. On the 14.3-second, 768p workload, its optimized backbone is 2.6× faster on one B200 and 3.2× faster on one H200. With eight-step distillation and eight-GPU inference, DiT denoising takes 6.70 seconds on B200s and 12.5 seconds on H200s. Eight-step VDN-H3 remains close to 50-step Dense H3 across the reported quality evaluations.

R EFERENCES Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-VL technical report. arXiv preprint arXiv:2511.21631, 2025. Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, Junyu Chen, Shuai Yang, Xianbang Wang, Yicheng Pan, Daquan Zhou, Huan Ling, et al. SANA-Video: Efficient video generation with block linear diffusion transformer. arXiv preprint arXiv:2509.24695, 2025. Junsong Chen et al. SANA-Video 2.0: Hybrid linear attention with attention residuals for efficient video generation. arXiv preprint arXiv:2607.21553, 2026. Krzysztof Choromanski et al. Rethinking attention with performers. In International Conference on Learning Representations, 2021. Tri Dao. FlashAttention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691, 2023. Tri Dao and Albert Gu. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In International Conference on Machine Learning, 2024. Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memoryefficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems, 2022. Jiarui Fang and Shangchun Zhao. USP: A unified sequence parallelism approach for long context generative ai. arXiv preprint arXiv:2405.07719, 2024. Jiarui Fang, Jinzhe Pan, Aoyu Li, Xibo Sun, and Jiannan Wang. PipeFusion: Patch-level pipeline parallelism for diffusion transformers inference. arXiv preprint arXiv:2405.14430, 2024a. Jiarui Fang, Jinzhe Pan, Xibo Sun, Aoyu Li, and Jiannan Wang. xDiT: An inference engine for diffusion transformers with massive parallelism. arXiv preprint arXiv:2411.01738, 2024b. FastVideo Team. FastH3 Preview v1: Open-weight 4-step sparse-distilled MiniMax-H3. Hugging Face model repository, 2026. URL https://huggingface.co/FastVideo/ FastVideo-FastH3-4-step-Preview-v1-LoRA. Tianrui Feng, Zhi Li, Shuo Yang, Haocheng Xi, Muyang Li, Xiuyu Li, Lvmin Zhang, Keting Yang, Kelly Peng, Song Han, Maneesh Agrawala, Kurt Keutzer, Akio Kodaira, and Chenfeng Xu. StreamDiffusionV2: A streaming system for dynamic and interactive video generation. In Proceedings of Machine Learning and Systems, 2026. Chongjian Ge, Hanwen Jiang, Tianyu Wang, Jiuxiang Gu, Yiran Xu, Ziwen Chen, Shaoteng Liu, Jing Shi, Yicong Hong, Zefan Cai, Hailin Jin, and Hao Tan. Chimera: Designing and chinchilla-scaling hybrid visual diffusion transformers. arXiv preprint arXiv:2607.28611, 2026. Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. Yoav HaCohen et al. LTX-Video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103, 2024. Ali Hatamizadeh, Yejin Choi, and Jan Kautz. Gated DeltaNet-2: Decoupling erase and write in linear attention. arXiv preprint arXiv:2605.22791, 2026.

12

Preprint Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao, Yinlong Qian, Ling Chen, Yunchao Wei, and Lin Ma. M4V: Multimodal mamba for efficient text-to-video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. Sam Ade Jacobs et al. DeepSpeed Ulysses: System optimizations for enabling training of extreme long sequence transformer models. arXiv preprint arXiv:2309.14509, 2023. Yang Jin et al. Pyramidal flow matching for efficient video generative modeling. arXiv:2410.05954, 2024.

arXiv preprint

Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, 2020. Kimi Team. Kimi linear: An expressive, efficient attention architecture. arXiv preprint arXiv:2510.26692, 2025. Weijie Kong et al. HunyuanVideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. larryvrh. MiniMax-H3-Turbo-LoRA. Hugging Face model repository, 2026. huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora.

URL https://

Jiachen Li et al. T2V-Turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback. arXiv preprint arXiv:2404.08865, 2024a. Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. VideoMamba: State space model for efficient video understanding. arXiv preprint arXiv:2403.06977, 2024b. Muyang Li, Tianle Cai, Jiaxin Cao, Qinsheng Zhang, Han Cai, Junjie Bai, Yangqing Jia, Ming-Yu Liu, Kai Li, and Song Han. DistriFusion: Distributed parallel inference for high-resolution diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024c. Muyang Li et al. SVDQuant: Absorbing outliers by low-rank components for 4-bit diffusion models. In International Conference on Learning Representations, 2025. Xingyang Li et al. DeltaQuant: 4-bit video diffusion models with spatiotemporal delta smoothing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. EvalCrafter: Benchmarking and evaluating large video generation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. Simian Luo et al. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv preprint arXiv:2310.04378, 2023. MiniMax. MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities. Model release, 2026. URL https://www.minimax.io/blog/minimax-h3. Adam Polyak et al. Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720, 2024. Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2022. Axel Sauer et al. Adversarial diffusion distillation. arXiv preprint arXiv:2311.17042, 2023. Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. FlashAttention-3: Fast and accurate attention with asynchrony and low-precision. In Advances in Neural Information Processing Systems, 2024. Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In International Conference on Machine Learning, 2023. Zachary Teed and Jia Deng. RAFT: Recurrent all-pairs field transforms for optical flow. In European Conference on Computer Vision, 2020. Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025.

13

Preprint Fu-Yun Wang et al. AnimateLCM: Accelerating the animation of personalized diffusion models and adapters with decoupled consistency learning. arXiv preprint arXiv:2402.00769, 2024. Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004. Haoning Wu, Chaofeng Chen, Jingwen Hou, Liang Liao, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. FAST-VQA: Efficient end-to-end video quality assessment with fragment sampling. In European Conference on Computer Vision, 2022. Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, Qiong Yan, Xiongkuo Min, Guangtao Zhai, and Weisi Lin. Q-Align: Teaching LMMs for visual scoring via discrete text-defined levels. In International Conference on Machine Learning, 2024. Haocheng Xi, Shuo Yang, Yilong Zhao, Muyang Li, Han Cai, Xingyang Li, Yujun Lin, Zhuoyang Zhang, Jintao Zhang, Xiuyu Li, Zhiying Xu, Jun Wu, Chenfeng Xu, Ion Stoica, Song Han, and Kurt Keutzer. Quant VideoGen: Auto-regressive long video generation via 2-bit KV-cache quantization. arXiv preprint arXiv:2602.02958, 2026. Haocheng Xi et al. Sparse VideoGen: Accelerating video diffusion transformers with spatial-temporal sparsity. arXiv preprint arXiv:2502.01776, 2025. Enze Xie et al. SANA: Efficient high-resolution image synthesis with linear diffusion transformers. In International Conference on Learning Representations, 2025. Bowen Xue, Zihan Min, Xingyang Li, Zhekai Zhang, Haocheng Xi, Lvmin Zhang, Maneesh Agrawala, JunYan Zhu, Song Han, Yujun Lin, and Muyang Li. FourTune: Towards fully 4-bit efficient post-training for diffusion models. arXiv preprint arXiv:2607.05711, 2026. Shuo Yang et al. Sparse VideoGen2: Accelerate video generation with sparse attention via semantic-aware permutation. arXiv preprint arXiv:2505.18875, 2025a. Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving Mamba2 with delta rule. In International Conference on Learning Representations, 2025b. Songlin Yang et al. Parallelizing linear transformers with the delta rule over sequence length. In Advances in Neural Information Processing Systems, 2024a. Songlin Yang et al. Gated linear attention transformers with hardware-efficient training. In International Conference on Machine Learning, 2024b. Zhuoyi Yang et al. CogVideoX: Text-to-video diffusion models with an expert transformer. In International Conference on Learning Representations, 2025c. Tianwei Yin et al. One-step diffusion with distribution matching distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024a. Tianwei Yin et al. Improved distribution matching distillation for fast image synthesis. arXiv preprint arXiv:2405.14867, 2024b. Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, and Jianfei Chen. SageAttention2: Efficient attention with thorough outlier smoothing and per-thread INT4 quantization. arXiv preprint arXiv:2411.10958, 2024. Jintao Zhang, Jia Wei, Haoxu Wang, Pengle Zhang, Xiaoming Xu, Haofeng Huang, Kai Jiang, Jianfei Chen, and Jun Zhu. SageAttention3: Microscaling FP4 attention for inference and an exploration of 8-bit training. arXiv preprint arXiv:2505.11594, 2025a. Jintao Zhang et al. SageAttention: Accurate 8-bit attention for plug-and-play inference acceleration. In International Conference on Learning Representations, 2025b.

14

Preprint Jintao Zhang et al. SLA: Beyond sparsity in diffusion transformers via fine-tunable sparse-linear attention. arXiv preprint arXiv:2509.24006, 2025c. Jintao Zhang et al. SpargeAttn: Accurate sparse attention accelerating any model inference. In International Conference on Machine Learning, 2025d. Jintao Zhang et al. SLA2: Sparse-linear attention with learnable routing and QAT. arXiv:2602.12675, 2026a.

arXiv preprint

Jintao Zhang et al. SpargeAttention2: Trainable sparse attention via hybrid top-k+top-p masking and distillation fine-tuning. arXiv preprint arXiv:2602.13515, 2026b. Junyi Zhang, Charles Herrmann, Junhwa Hur, Chen Sun, Ming-Hsuan Yang, Forrester Cole, Trevor Darrell, and Deqing Sun. LoGeR: Long-context geometric reconstruction with hybrid memory. arXiv preprint arXiv:2603.03269, 2026c. Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, and Xue Yang. FIRM-Video: Check before you score for reliable text-to-video reward modeling. arXiv preprint arXiv:2608.21839, 2026d. Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. Tianchen Zhao et al. ViDiT-Q: Efficient and accurate quantization of diffusion transformers for image and video generation. arXiv preprint arXiv:2406.02540, 2024. Haoyi Zhu. Reflections on Video DeltaNet. Blog post, aug 2026. URL https://haoyizhu.site/blog/ video-delta-rule/. Haoyi Zhu et al. SANA-WM: Efficient minute-scale world modeling with hybrid linear diffusion transformer. arXiv preprint arXiv:2605.15178, 2026.

15

Preprint

A

D ERIVATION AND I MPLEMENTATION D ETAILS

A.1

P ROOF OF P ROPOSITION 1

Fix the prepared features and gates of frame t. Substituting S̄t = St−1 Diag(αt ) into Equation (5) separates the inherited state from the new write: St = St−1 Mt + Jt , Mt = Diag(αt )(I + At )−1 ,

Jt = Bt (I + At )−1 .

(A.1)

Mt . Since βt,u ≥ 0, the For two entering states with the same frame inputs, P∆St = ⊤∆St−1 2 Gram matrix is positive semidefinite: x⊤ At x = β (k x) ≥ 0 for every x. Writing t,u t,u u At = Q Diag(λi )Q⊤ , with λi ≥ 0, shows that (I + At )−1 has eigenvalues 1/(1 + λi ) ∈ (0, 1]. Together with 0 ≤ αt,j ≤ 1, this gives ∥∆St ∥F ≤ ∥∆St−1 ∥F ∥Mt ∥2 , ∥Mt ∥2 ≤ ∥ Diag(αt )∥2 ∥(I + At )−1 ∥2 ≤ 1.

(A.2)

This proves Proposition 1 without requiring the decay and Gram matrix to commute. The bound concerns inherited state at fixed frame inputs, not new writes or the full network. A.2

F RAME - SIZE SCALING AND CORRELATION AWARENESS

The difference between independent and joint frame writes can be seen by expanding Equation (3c): Stbatch = S̄t +

U X

⊤ βt,u (vt,u − S̄t kt,u )kt,u .

u=1

All corrections use the same pre-update state, so a patch’s residual ignores the other current-frame patches. By comparison, the first-order condition for VDA can be rearranged as St = S̄t +

U X

⊤ βt,u (vt,u − St kt,u )kt,u .

u=1

Here every residual depends on the shared St . Overlapping keys accumulate in At , so Equation (5) adjusts their joint update. With αt = 1, the additive inherited-state factor is I − At . Its eigenvalues 1 − λi (A Pt ) leave [−1, 1] if any λi (At ) > 2. For unit keys and sigmoid write gates, λmax (At ) ≤ tr(At ) = u βt,u ≤ U , and aligned keys can approach this bound. Thus per-token normalization alone does not ensure stability. √ SANA-WM scales each unit key by an additional 1/ U (Zhu et al., 2026), replacing At by At /U so that I − At /U is non-expansive. VDA instead uses (I + At )−1 and needs no frame-size scale. The inverse also responds to the observed key correlations. If all U positions repeat a unit key k P with a common gate β, writing s = S̄t k and v̄ = U −1 u vu gives St k =

1 Uβ s+ v̄. 1 + Uβ 1 + Uβ

(A.3)

Repeated keys thus receive a saturating U β/(1 + U β) weight, whereas orthogonal keys each receive β/(1 + β). SANA-WM’s scaled-key readout instead assigns β to repeated-key consensus and β/U to each orthogonal direction. VDA adapts to key geometry rather than applying the same U -based scale.

16

Preprint A.3

B OUNDARY GATHER AND DECAY BRIDGE

After excluding anchor indices 0 and F − 1, let Pj be the forward state through interior frame j and Rj the reverse state from the last interior frame down to j. Their virtual initial states are P0 = RF −1 = ST /2. For query t, let [ℓt , ht ] be its local window clipped to interior frame indices. The state read by its prepared queries is

Set = Pℓt −1 Diag

t Y

! αr

r=ℓt

+ Rht +1 Diag

ht Y

! αr

(A.4) ,

r=t

e oL t,u = St qt,u . The products are channel-wise. Missing distant video on one side selects the corresponding initial text state rather than a fabricated frame state. The bridge applies decay without local frame writes. If a window covers the complete clip, the implementation bypasses the linear pathway and uses the dense attention path. Clips consisting only of anchor frames also receive no linear video output. Algorithm 1. Bidirectional VDA inside one hybrid attention block 1. Compute shared QKV and retained multimodal Softmax with chunk and anchor masks. 2. Remove anchor video rows from the linear inputs; prepare Q/K/V features and frame gates. 3. Compute At , Bt , αt for all interior frames. Form the text state from the prompt. 4. Perform a batched factorization of I + At ; construct Ct , Mt , Jt for all frames and heads. 5. Scan S ← SMt + Jt in both temporal directions, each initialized with ST /2. 6. Gather outside-window states and apply the decay bridge in Equation (A.4). 7. Read each query; apply RMSNorm and the linear gate; zero anchor rows. 8. Gate and project Softmax; add the projected linear output to video rows. A.4

A RCHITECTURE ADAPTATION HYPERPARAMETERS

We use AdamW with linear warmup followed by cosine decay. Stages A1, A2, and B use a peak learning rate of 10−4 and decay to 5 × 10−6 ; Stage D uses 10−5 for the generator and 2 × 10−5 for the fake-score model, decaying to 10−6 and 2 × 10−6 , respectively. Following Chimera (Ge et al., 2026), parameters are divided into learning-rate groups by module scale: large fan-in matrices use the base schedule, while smaller vector parameters use larger multipliers. In Stage B, the Linear branch uses 0.4× the LoRA learning rate. The table below lists the remaining stage-specific settings.

17

Preprint Table 4: Hyperparameters for architecture adaptation. The table covers Stages A1, A2, and B; few-step distillation is described in Section 4.2. Setting

A1: per-layer

Parameter status

Train: one Linear branch Train: all Linear branches Train: branches, gates, and QKVO Freeze: backbone and Softmax Freeze: backbone and Softmax LoRA gate gates Freeze: backbone FFN, embeddings, and head

A2: end-to-end

500

B: LoRA co-adaptation

Steps

200

Optimizer / schedule

AdamW; linear warmup → co- AdamW; linear warmup → co- AdamW; linear warmup → cosine sine sine

2,000

β1 , β2

0.9, 0.999

0.9, 0.999

0.9, 0.999

Peak / minimum LR

10−4 / 5 × 10−6

10−4 / 5 × 10−6

10−4 / 5 × 10−6

Warmup steps

50

50

200

Weight decay

0

0

0

Gradient clip

0.1 per layer

1.0 global

1.0 global

LR multipliers

large 1.0 / small 5.0

large 1.0 / small 5.0

LoRA 1.0 / branch 0.4 / small 2.0

LoRA rank / scale

–

–

64 / 64

A.5

R ELEASE CONFIGURATION AND REPRODUCIBILITY

The architectural settings used in this report correspond to the released vdn_solve branch, chunk size 5 with radius 1, anchors in both row and column directions, K/V short convolution, text-state initialization, and the alpha decay bridge. Architecture metadata is created in A1 and inherited from the checkpoint in later stages. This is preferable to silently changing an attention mask or state rule when resuming a run. The release package will include exact weight hashes, dataset and evaluation-prompt manifests, random seeds, tokenizer and frame-alignment metadata, GPU topology, software versions, timing protocol, and quality-evaluation outputs. These artifacts connect each reported result to a reproducible model and execution configuration. A.6

Q UALITY ACROSS DENOISING BUDGETS

Figures 8–10 compare the same evaluation suite at 50, eight, and four neural function evaluations (NFEs). Each metric uses the same labeled, truncated vertical scale across the three figures, but different metrics have different scales. An asterisk reproduces the reported paired significance flag relative to 50-step Dense H3; it does not establish significance between two non-dense variants. RAFT mean flow is a motion-magnitude diagnostic rather than a quality score, and lower FL2VA LPIPS is better. Fifty steps. VDN-H3 after Stage B is broadly on par with Dense H3 before few-step distillation. Their aesthetic, technical, learned-quality, and FIRM scores are close, although two endpoint-fidelity measures show small degradations. This comparison isolates the hybrid architecture and its adaptation from any reduction in sampling steps.

18

Preprint Dense H3 (50 NFE) 30

92

27.69

27.54

20

88.20

VDN-H3 (50 NFE) 95

87.61

70

3.5

Q-Align 3.32

3.22

FAST-VQA

2.4

11.44

82.34

75

VQAT

13

82.68

92.14

86

VQAA

11.55

85

92.22

4.6

4.46

4.40

2.25 2.12

8

0

2

RAFT Mean Flow 2

30

1.84

4

Instruction Following

DOVER++

1.81

Perceptual quality

.88

28.85

28.83

.30

.833

.830* .111*

.104 1.5

25

.70

World coherence

FL2VA PSNR

0

FL2VA SSIM

FL2VA LPIPS

Figure 8: Quality at 50 NFEs before few-step distillation. VDN-H3 is the Stage-B checkpoint.

Eight steps. We compare distilled VDN-H3 with Dense H3 plus Larry’s eight-step adapter and FastH3 v2 at the same NFE count. VDN-H3 scores highest on all five no-reference quality measures and remains close to the dense-plus-Larry reference on FIRM and FL2VA. Its VQAT score is 10.78 points higher than FastH3 v2. FL2VA scores for FastH3 v2 were not available. Dense + Larry (8 NFE) 30

28.03

27.09

92

FastH3 v2 (8 NFE)

VDN-H3 (8 NFE)

95

89.20

87.25

92.37

26.00

70

13

3.5

11.71

Q-Align 3.28

2.94*

2.13

0

RAFT Mean Flow 2

4.6

4.38

2.25 0.61*

8

FAST-VQA

2.4

10.16*

Perceptual quality

.88

28.67*

4.45

4

Instruction Following

DOVER++

28.76*

4.38

2.11

2

30

1.83

75

VQAT

12.62*

83.41*

82.83 79.27*

86

VQAA

85

90.71*

78.42* 20

93.13*

.30

.827*

.826*

1.77 .116*

.110* 1.60* 1.5

25

World coherence

n/a

.70

FL2VA PSNR

n/a

FL2VA SSIM

0

n/a

FL2VA LPIPS

Figure 9: Quality at eight NFEs. FastH3 v2 has no reported FL2VA scores in this comparison.

Four steps. VDN-H3 remains compatible with Larry-style few-step adaptation at the lower sampling budget. It scores above both Dense H3 plus Larry and FastH3 v1 on all five no-reference quality measures; against FastH3 v1, its VQAT score is higher by 12.64 points, with better FL2VA endpoint fidelity. These are score differences: the supplied significance flags compare each variant with Dense 50, not VDN-H3 directly with FastH3.

19

Preprint Dense + Larry (4 NFE) 30

FastH3 v1 (4 NFE)

92

22.18*

23.11*

70

3.5

75

Q-Align

2.81*

9.19*

4.6

2.18

2.20

4.40

2.21

4.28*

0

RAFT Mean Flow 2

2

4

Instruction Following

DOVER++ 30

Perceptual quality

.88

28.56*

28.48*

1.81

.30

.823*

27.59*

.822*

.183*

.785*

.117*

1.65* 1.5

25

World coherence

4.22*

0.52*

8

1.74*

FAST-VQA

2.4

1.91*

10.71*

78.02*

86

VQAT

12.72*

82.96 81.19*

89.22* 89.18*

75.46*

VQAA 13

85

91.73

82.60*

24.59*

20

VDN-H3 (4 NFE)

95

88.10

.70

FL2VA PSNR

.122*

0

FL2VA SSIM

FL2VA LPIPS

Figure 10: Quality at four NFEs. VDN-H3 retains a clear margin over FastH3 v1 while using the same number of denoising steps.

20

Record · ID 978418 · SHA-256 a162ae18c95214cd
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.