Kwai Summary Attention Technical Report
arXiv:2604.24432v1 [cs.CL] 27 Apr 2026
OneRec Team Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agentic intelligence and recommendation system. However, the standard softmax attention exhibits quadratic time complexity with respect to sequence length. As the sequence length increases, this incurs substantial overhead in long-context settings, leading the training and inference costs of extremely long sequences deteriorate rapidly. Existing solutions mitigate this issue through two technique routings: i) Reducing the KV cache per layer, such as from the head-level compression GQA, and the embedding dimension-level compression MLA, but the KV cache remains linearly dependent on the sequence length at a 1:1 ratio. ii) Interleaving with KV Cache friendly architecture, such as local attention SWA, linear kernel GDN, but often involve trade-offs among KV Cache and long-context modeling effectiveness. Besides the two technique routings, we argue that there exists an intermediate path not well explored: Maintaining a linear relationship between the KV cache and sequence length, but performing semantic-level compression through a specific ratio 𝑘. This 𝑂 ( 𝑛/𝑘) path does not pursue a “minimum KV cache”, but rather trades acceptable memory costs for complete, referential, and interpretable retention of long distant dependency. Motivated by this, we propose Kwai Summary Attention (KSA), a novel attention mechanism that reduces sequence modeling cost by compressing historical contexts into learnable summary tokens. Specifically, summary tokens are injected into the input sequence at regular intervals, interleaved with text tokens, which partitions the sequence into distinct chunks, and enables text tokens and summary tokens to operate with different attention scopes. This design ensures both long-range dependencies via summary tokens and full expressiveness via dense attention with text tokens in adjacent chunks. Empirical results demonstrate that our hybrid-KSA significantly surpass other hybrid variants in long-context modeling, e.g., +3.69%/5.48% than hybrid-GDN at RULER-128K at from-scratch/CPT setting. In inference, KSA’s sequence-level compression is fully orthogonal with GQA and MLA. Composing KSA with these methods provides a reliable 8× further KV cache compression while preserving model performance.
Figure 1 | From scratch model variants performance on long-context and general benchmarks. Our hybrid-KSA (8× sequence compression, 3:1 KSA/Full mixture ratio) leads on multiple long-context retrieval tasks and achieves competitive general understanding ability with other Hybrid variants.
© 2026 Kuaishou. All rights reserved
Kwai Summary Attention Technical Report
1. Introduction In past years, the Transformer’s scaling-laws Kaplan et al. (2020) on data volume, model parameters and computation resources achieve great success to compress the knowledge of human society, pushing the AI intelligence frontier again and again. Excepted the above three well-explored scaling directions, the sequence modeling context length also provides a vital scaling perspective at same time. With parallel infrastructure improvement, valid context window gradually increases from earliest 1K to the latest 1M sequence Anthropic (2026); Kimi2.6 (2026); LongCat et al. (2026); Qwen3.5 (2026). As we know, the long-context modeling capabilities can significantly reduce model hallucination and incubate new Agent applications in solving intricate real-world tasks, such as the Code Agent Zhang et al. (2024), OpenClaw OpenClaw (2026) and so on. In the Agentic era, obtaining extensive memory across multi-turn conversations and accurately retrieving historical contexts are prerequisites for highly intelligent agents and long chain-of-thought reasoning. However, vanilla attention Vaswani et al. (2017) is subjected by its Achilles’ heel that imposes un-affordable storage and compute overhead as the sequence length grows: the KV cache scales linearly with sequence length, while the attention computation scales quadratically. To further unleash the long-context capability of LLMs and unlock the next tier of intelligence, inventing efficient attention variants and designing corresponding modeling architectures that alleviate the computation bottleneck, have become priorities research topic in the LLM field. Around the challenge of how to reduce KV cache and computational overhead in long-context setting, two technique paths have been well explored in recent years: • Reducing the KV cache per layer, evolving from MHA to GQA, e.g., the Qwen series Qwen2 et al. (2024); Qwen3 et al. (2025); Qwen3.5 (2026); Qwen3.6 (2026), MLA (DeepSeek series DeepSeek-V2 et al. (2024); DeepSeek-V3 et al. (2024); Guo et al. (2025); DeepSeek-AI (2026), Kimi K2 Kimi et al. (2025)), MLA-DSA (DeepSeek V3.2 Liu et al. (2025), GLM-5 Zeng et al. (2026) and DeepSeek NSA Yuan et al. (2025)), as refinement of the naive full attention. The core idea of this branch is to compress the per-token KV cache through head grouping and sharing, or low-rank KV projection. However, those efforts’ KV cache costs still maintain a strict 1:1 linear relationship with sequence length. • Interleaving with KV Cache friendly architecture, with varying mixing ratios, represented by linear Gated DeltaNet + GQA Qwen3.5 (2026) and SWA + GQA Agarwal et al. (2025); Xiao et al. (2026). This branch replaces the majority of layers with efficient attention variants that carry a much smaller KV cache, fundamentally decoupling the KV cache from sequence length in some layers. However, the drawback is also obvious: for linear attention, the fixed-size state is inherently lossy compression, in which long-range information becomes blurred and unattainable; local variants, on the other hand, completely discard any context outside the window, thus losing perception of the distant information. The above discussion reveals: i) KV cache compression methods, e.g., GQA, MLA still maintain strict linear relationship between sequence length and cache cost. ii) Architectures whose KV cache costs are independent of sequence length, e.g., Hybrid SWA, Hybrid Linear, face limited modeling expressivity and impaired long-context capacity. However, we argue that there should be have a intermediate technique roadmap which has previously been underestimated: Maintaining a linear relationship between the KV cache and sequence length, but performing semantic-level compression through a specific ratio 𝑘. This 𝑂 ( 𝑛/𝑘) path does not pursue a “minimum KV cache”, but rather trades acceptable memory costs for concrete and interpretable long distant dependency. Compared with SWA and Linear attention, it has the potential that preserving full fidelity over long-range dependency, more friendly for long-context reasoning, agent trajectories, and downstream RL training signals. In 2
Kwai Summary Attention Technical Report
t0
t1
t2
t3
S0
t4
t5
t6
t7
S1
t8
t9
t10
t11
S2
t12
t13
t14
t15
S3
t16
t17
t18
t19
S4
t20
t21
t22
t23
S5
t24
t25
t26
t27
S6
t28
t29
t30
0
1
2
3
3
4
5
6
7
7
8
9
10
11
11
12
13
14
15
15
16
17
18
19
19
20
21
22
23
23
24
25
26
27
27
28
29
30
⋯ pos
Chunk 0
Chunk 1
Chunk 2
Chunk 3
Chunk 4
Text Token
Chunk 5
Chunk 6
Summary Token
(a) Input sequence is sliced into chunks and summary tokens are injected by the boundary of each chunk. Chunk 0 Chunk 1 Chunk 2 Chunk 3 t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3 ⋯ t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3 ⋯
Chunk 0 Chunk 1 Chunk 2 Chunk 3 t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3 ⋯
⊕
FFN
1×
⊕
Full Attention
⊕
FFN
3×
⊕
Summary Attention
⋯ Summary Attention Mask
Chunk Summarization
Cross-chunk Summary
t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3 ⋯
Local Chunks Causal
Full Attention Mask
Masked
(b) Summary tokens only see within chunk, while text tokens consider both local texts and distant summaries.
Figure 2 | Our proposed Kwai Summary Attention (KSA), maintains sub-quadratic complexity while preserving the expressivity in modeling long-range dependencies. By interleaving KSA with full attention blocks, hybrid-KSA achieves satisfactory balance between performance and efficiency. other words, we believe that sequence-level token compression could provide a new perspective for reducing KV cache. At the same time, DeepSeek V4 DeepSeek-AI (2026) model series have been released recently, whose design follows the first-principle of sequence-level KV Cache compression too, further proves the engineering feasibility and long-context robustness of this technique roadmap. In light of the above, we propose Kwai Summary Attention (KSA), an efficient, adaptive, and scalable attention mechanism. KSA introduces a novel summary mechanism designed to distill historical contexts from lengthy sequences into a lightweight, learnable summary token. As illustrated in Figure 2a, our method operates by slicing the input sequence into multiple chunks of a specified size and injecting a learnable summary token at the end of each chunk. To achieve satisfactory trade-off between modeling capacity and efficiency in long-sequence processing, KSA enables text tokens and summary tokens to perform distinct computational patterns, as shown in Figure 2b: i) On the one hand, the summary tokens serve the purpose of distilling semantic information in the current chunk, thus considering only text tokens within the same chunk. As the sequence grows, these tokens act as concise priors of distant contexts for subsequent text tokens. ii) On the other hand, for the semantic fluency and stability, text tokens could interact with local neighboring text tokens via a sliding chunk mechanism, and could attend the long-range summary tokens of previous chunks. The local-text/global-summary visible tokens are able to provide favorable expressiveness and significantly reducing inference latency and KV-cache memory usage. To validate the effectiveness of KSA, we perform comprehensive from-scratch (Scratch) or continual pre-training (CPT) experiments on a wide range of downstream benchmarks and carry out detailed ablation studies with a series of architectural variants. Moreover, we open-source the KSA training recipes and computation kernel to facilitate the next generation LLM architecture research1 . 1Our KSA training scripts at https://github.com/Kuaishou-OneRec/KSA
3
Kwai Summary Attention Technical Report
2. Methodology 2.1. Rethinking Long-Context Modeling LLM long-context training and inference faces two main challenges: KV cache growth and attention computation cost. In retrospect, Full attention and its variants GQA/MLA preserves the complete history, but has KV cache grows linearly with sequence length, which becomes the primary bottleneck for long-context inference. In contrast, pure linear or local attention attains linear scaling through fixed-size recurrent states or window size, yet the limited state capacity often fails to retain finegrained semantic information callback over long contexts effectively. KSA strikes a trade-off between these two extremes: it continuously compresses long-context information into a growing set of summary states, enabling expressive modeling of distant dependencies without explicitly caching all historical tokens. Unlike pure linear attention, the summary state is not fixed-size but grows progressively as summary tokens; unlike sparse attention, long-range modeling no longer hinges on sparse token-to-token connections but routes distant information through summary tokens that act as compressed relays. Taken together, KSA can be regarded as a mixture mechanism that integrates local-global attention (exact local token-level modeling within a sliding window) with compressed global long-range attention (via a linearly growing set of summary tokens through a specific ratio), offering a better balance among modeling capacity, computational efficiency, and memory footprint. 2.2. Kwai Summary Attention This section presents the two key components of Kwai Summary Attention: summary token compression and sliding chunk attention. 2.2.1. Summary Token Compression Given an input sequence T = [𝑡0 , . . . , 𝑡𝑛−1 ] of 𝑛 text tokens and a chunk size 𝑘, we first partition it into 𝑛/𝑘 chunks (assume 𝑛 is divisible by 𝑘) and append a shared learnable summary embedding 𝑠 at the end of each chunk; we write 𝑠 𝑗 for the occurrence of 𝑠 placed in chunk 𝑗. Each summary token serves as a distilled representation of text tokens within the chunk. Formally, the augmented sequence T̂ is: T̂ = [chunk0 , chunk1 , . . . , chunk 𝑛𝑘 −1 ] , where chunk 𝑗 = [𝑡 𝑗𝑘 , 𝑡 𝑗𝑘+1 , . . . , 𝑡 𝑗𝑘+( 𝑘 −1) , 𝑠 𝑗 ];
(1)
where 𝑡 𝑗𝑘 is the first text token of chunk 𝑗, and 𝑠 𝑗 is its summary token (all 𝑠 𝑗 share the special learnable summary token 𝑠). Based on the two different token role, we impose structural constraints on the information flow to make the visibility of summary tokens and text tokens complementary: • Summary tokens can only see text tokens within their own chunk, and nothing else. Moreover, summary token position id is the same with its own chunk’s last text token’s. • Text tokens can see its short-term sliding neighbor text tokens and past long-term summary tokens, without direct access to the full text history. This design explicitly decouples token space from state space: summary tokens focus on short sequence semantic compression, while text tokens access long-range information through previous summary tokens.
4
Kwai Summary Attention Technical Report
Sliding Window Attention (Window Size = 4) Chunk 0
Chunk 1
Chunk 2
Sliding Chunk Attention (Sliding Chunk Num = 1) Chunk 0 Chunk 1 Chunk 2 Chunk 3
Chunk 3
t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3
Query position
Query Position
t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3 t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3
t₀ t₁ t₂ t₃ s₀ t₄ t₅ t₆ t₇ s₁ t₈ t₉ t₁₀ t₁₁ s₂ t₁₂ t13 t14 t15 s3
Key Position Chunk Summarization
Key Position Cross-chunk Summary
Local Chunks Causal
Masked
Info lost in SWA
Figure 3 | Sliding Window Attention (left) may cut through a chunk, resulting attention to overlook certain tokens at boundaries and causing information loss (red regions). Sliding Chunk Attention (right) aligns window boundaries with chunk boundaries, ensuring clean information routing. 2.2.2. Sliding Chunk Attention To ensure the text token information is orthogonal with the past summary tokens’, we further introduce a sliding chunk attention mechanism (SCA) that allows text tokens to see the latest several chunks’ text token. Formally, for a text token 𝑖 with chunk size 𝑘, a standard SCA visible scope is: 𝑡 ( ⌊ 𝑖 ⌋ −𝐶 )∗𝑘 , . . . , 𝑡𝑖 −1 , 𝑡𝑖 (2) 𝑘
where ⌊·⌋ is the floor division operator, and the ⌊ 𝑘𝑖 ⌋ − 𝐶 indicate the starting chunk index. In our KSA computation workflow, we combine the SCA and the summary tokens attention together, to inject the short-term neighboring and long-term summary information at the same time. Formally, for a text token 𝑖 with chunk size 𝑘, its text token visible scope as follows: 𝑠0 , 𝑠1 , . . . , 𝑠 ⌊ 𝑖 ⌋ −𝐶 −1 ∪ 𝑡 ( ⌊ 𝑖 ⌋ −𝐶 )∗𝑘 , . . . , 𝑡𝑖 −1 , 𝑡𝑖 (3) 𝑘 𝑘 | {z } | {z } distant summary tokens
sliding chunk text tokens
The left term covers distant summaries beyond the window; the right term covers text tokens within the window. Note that summary tokens within the sliding chunk are not considered by text tokens, since their information is already included, as shown in Figure 3 attention mask matrix. Why chunk-level sliding, not token-level? A natural alternative is standard token-level Sliding Window Attention (SWA), where text token could attends to a fix window size latest tokens and it have been already validated in many different LLMs (e.g., GPT-OSS). However, when combined with our summary tokens, its token-level sliding may omit information or introduce more information. As shown in Figure 3, if the window boundary cuts through a chunk, the text token sees only a subset of that chunk’s text tokens. Meanwhile, the chunk is not fully outside the window, so its summary token does not qualify as a distant summary either. As a result, the partially visible chunk is neither covered by complete text nor compensated by its summary, which results in loss of information. 5
Kwai Summary Attention Technical Report
(a) The KSA KV cache consists of contiguous tensors with an organized layout, enabling reading KV states with only one memory operation without concatenation during decoding.
(b) Inserting a new text token: the KV is written into the next available slot of the current chunk, and attention reads a single contiguous slice.
(c) Inserting the chunk summary token: the self-KV is written into the scratch slot left of the current chunk, so summary attention is also a single contiguous slice.
(d) Replacing the oldest text chunk: the just-finished chunk is copied to the ring’s write pointer, overwriting the oldest sliding chunk in place.
(e) Appending the summary token to the buffer: the new summary token is committed to the right end of the Summary Token Buffer.
Figure 4 | The KSA cache features organized layout of contiguous tensors to enable efficient memory access. The lifecycle involves writing new tokens, inserting summaries, and recycling chunks in place. By sliding at chunk granularity, we guarantee a clean partition: every past chunk is either fully within the window (all text tokens visible) or fully outside (accessed only through its summary token), with no intermediate state. This information routing ensures that no chunk’s information is lost. 2.3. Kernel Design The attention mask of KSA is a structured sparse mask defined jointly by the local sliding window of text tokens and the visibility of distant summary tokens. With augmented input of length 𝐿 = 𝑛 + 𝑛/𝑘 where 𝑛/𝑘 is the total number of summary tokens, building the full mask would cost a prohibitive 𝑂 ( 𝐿2 ) memory, which is infeasible for long sequences. To address this, we design a block-sparse attention kernel for training and prefill, and a summary KV cache for decoding, addressing compute efficiency and memory cost respectively. In the training / prefill kernel, Q/K/V are split into fixed-size blocks, and since KSA discard most Q-K interactions under sparse setting, the kernel only loads non-zero block pairs from HBM to SRAM for computation. 2.3.1. Efficient Decoding with Summary KV Cache The nature of auto-regressive decoding requires every generation step to take all previous KV entries into account, resulting in a memory-bandwidth bottleneck rather than a compute-bound one. A naive approach requires concatenating, discarding, and re-allocating KV segments at every step, and these scattered memory operations become the performance bottleneck in the tight decode loop. Composition of KSA KV cache (Figure 4a). We organize the KV cache as contiguous tensors, which can be logically divided into three distinct regions: i) Current Chunk, ii) Sliding Chunk Text, and iii) 6
Kwai Summary Attention Technical Report
Table 1 | Next per-token KV Cache comparison. Taking the 𝑘 = 8 chunk compression and ℎ = 128, 𝑑 = 128, 𝑔 = 8, 𝑑 𝑐 = 512, 𝑑𝑟 = 64 configurations as example. Mechanism
Effective Context
KV Cache Size (𝑛 → ∞)
Compression Rate
MHA
𝑛 (exactly)
2·𝑛·ℎ·𝑑
-
GQA MLA
𝑛 (exactly) 𝑛 (exactly)
2·𝑛·𝑔·𝑑 𝑛 · ( 𝑑 𝑐 + 𝑑𝑟 )
𝑔 /ℎ ≈ 6.25% ( 𝑑 𝑐 + 𝑑𝑟 )/(2 · ℎ · 𝑑 ) ≈ 1.76%
GDN SWA KSA
𝑛 (ambiguous) 𝑤 (exactly) 𝑛 (summarized)
2 · ℎ · 𝑑2 2·𝑤·𝑔·𝑑 2 · 𝑛/𝑘 · 𝑔 · 𝑑
𝑑 /𝑛 ≈ 0% 𝑤/𝑛 · 𝑔 /ℎ ≈ 0% 1/𝑘 ≈ 12.5%
KSA with GQA KSA with MLA
𝑛 (summarized) 𝑛 (summarized)
2 · 𝑛/𝑘 · 𝑔 · 𝑑 𝑛/ 𝑘 · ( 𝑑 𝑐 + 𝑑 𝑟 )
1/𝑘 · 𝑔 /ℎ ≈ 0.78% 1/𝑘 · ( 𝑑 𝑐 + 𝑑𝑟 )/(2 · ℎ · 𝑑 ) ≈ 0.22%
Summary Token Buffer. The Current Chunk region fills from right to left, with its right edge always aligned to the Sliding Chunk region. This alignment ensures the caches are physically contiguous, allowing text token attention to read all KV entries in a single operation without concatenation. The Summary Token Buffer grows to the right as new summaries are generated. Since RoPE is applied to each key before it enters the cache, every entry inherently carries its own position encoding, ensuring that the physical layout does not interfere with positional information. Updating the current text chunk (Figure 4b). The KV states of incoming tokens are written sequentially into the Current Chunk region at the next available slot. During attention computation, the model retrieves the KV cache as a unified, contiguous slice that spans the current chunk, the sliding chunk text, and the distant summaries. Chunk eviction and summary token insertion (Figures 4c, 4d, 4e). When the current chunk is filled, we execute the following update procedure: i) The KV of the new summary token is first written into the scratch slot to the left of the Current Chunk, so that it stays physically contiguous with the chunk text and summary attention can be computed via a single memory-slicing operation (Figure 4c). ii) The just-finished chunk text is then copied into the ring buffer at the write pointer, overwriting the oldest chunk in a modular fashion (Figure 4d). iii) Finally, the new summary KV is appended to the Summary Token Buffer (Figure 4e). Conclusively, every KV read during decoding corresponds to a contiguous slice within the buffer. This design eliminates the need for concatenation, gather operations, or explicit mask construction: the cache layout itself inherently encodes the visibility rule.
3. KV Cache Memory Analysis As we known, the key challenge for deploying LLM is the KV cache state amount bound during inference. To alleviate the cache amount issue, there are many comprehensive works were proposed, this section analyzes KSA and other methods KV cache amount in details. 3.1. Per-Token KV Cache Cost Table 1 summarizes the KV cache state amount for each mechanism in naive setting:
7
Kwai Summary Attention Technical Report
Naive Multi-Head Full Attention (MHA) Vaswani et al. (2017). Full Attention stores all 𝑛 past key-value pairs explicitly, resulting in a cache of size 2𝑛ℎ𝑑 that grows linearly with sequence length. Each prediction step requires attending over all 𝑛 cached KV results. While this provides access to full context, the linear KV cache growth becomes the notorious bottleneck for long-sequence inference. Grouped Query Attention (GQA) Ainslie et al. (2023). In standard MHA, each of the ℎ attention heads maintains its own KV states, the GQA proposes a simple-yet-effective technique to organize the total ℎ heads into 𝑔 groups, where same group of Q-heads shared the same cached KV-heads. Thus the KV cache reduces from 2𝑛ℎ𝑑 (MHA) to 2𝑛𝑔𝑑 (GQA), the KV compression ratio is about 𝑔/ℎ. Multi-head Latent Attention (MLA) DeepSeek-V2 et al. (2024). Compared with GQA, MLA takes a more aggressive approach by projecting all KV information into low-rank latent vectors. During inference: i) first introduce a 𝑛𝑑 𝑐 cached block to store the non-RoPE key and values; the per-head non-RoPE keys and values are reconstructed via per-head up-projection matrices. ii) additionally, MLA also introduce a small 𝑛𝑑𝑟 cached RoPE block to decoupled for key part. With the two different cache block, the total per-token cache amount is 𝑛 ( 𝑑 𝑐 + 𝑑𝑟 ). Gated Delta Net (GDN) Yang et al. (2025). Different with the MLA and GQA that follows the standard full attention mechanism, the linear GDN mechanism compresses the entire sequence history into a fixed 𝑑 × 𝑑 state matrix through diagonal gating. The KV cache is always a small constant value 2𝑑 2 , which is independent of sequence length. Nevertheless, the long-term information dependency is unobserved and unexplainable in inference process: as new tokens arrive, older information is progressively attenuated by the decay factors 𝛼𝑡 , making the effective context is hard to analyze. Sliding Window Attention (SWA) Beltagy et al. (2020). Following the standard full attention mechanism, SWA narrows the visual token window to reduce the KV cache amount. Specifically, the SWA maintains latest 𝑤 recent tokens in a FIFO buffer, with a fixed cache size 2𝑤𝑔𝑑 . Although effective, the limitation is also obviously: tokens relative index more than 𝑤 are completely lost, could not access the long-term information directly. Kwai Summary Attention (KSA) KSA extends SWA by augmenting the local window with compressed summary tokens. The cache consists of two parts: 2𝑤𝑑 for the local window (almost same with SWA) and the global 2𝑛/𝑘𝑑 for the summary tokens, where 𝑘 is the chunk size, resulting the total cache is about 2𝑤𝑔𝑑 + 2( 𝑛/𝑘) 𝑔𝑑 and about 2𝑛/𝑘𝑑 for a large 𝑛. 3.2. Orthogonality of KV Cache Compression Actually, the KV cache amount of one single layer can be represented as follows: KV Cache = PastToken × HeadNum × EmbeddingDim
(4)
Given the above discussion, GQA, MLA, and KSA are orthogonal to each other: GQA compresses the heads group, MLA shrinks the embedding dimension (which must be re-computed per head), and our KSA reduces the number of tokens to be attended. Since these three optimization directions are independent, the compression ratios multiply when combined. For KSA with GQA:
8
Kwai Summary Attention Technical Report
Figure 5 | KV cache size comparison across different mechanisms as sequence length grows.
𝑛/ 𝑘 |{z}
KV CacheKSA+GQA =
×
KSA compression
2·𝑔·𝑑 | {z }
(5)
GQA compression
In theory, KSA could be combined with MLA: 𝑛/ 𝑘 |{z}
KV CacheKSA+MLA =
KSA compression
×
( 𝑑 𝑐 + 𝑑𝑟 ) | {z }
(6)
MLA compression
Table 1 summarizes the combined compression for different configurations. Notably, KSA provides a qualitatively different scaling advantage: while GQA and MLA reduce the constant factor in 𝑂 ( 𝑛) cache growth, KSA reduces the growth rate itself to 𝑂 ( 𝑛/𝑘). As shown in Figure 5, for long sequences, this sub-linear scaling could further reach a considerable compression ratio (e.g., 0.78% in KSA with GQA and 0.22% in KSA with MLA).
4. Experiments 4.1. Experiments Setting 4.1.1. Model Hyper-parameters Configuration We evaluate different attention architectures under two experimental settings: train-from-scratch (Scratch, 400B Token, 128K Sequence) and continual-pretraining (CPT, 85B Token, 128K Sequence), covering both full training and adaptation scenarios. Baseline models. Our tested model variants are as follows: (1) Pure Full Attention, (2) hybridGDN Yang et al. (2025)/Ring-Linear Ling-Team et al. (2025), (3) Hybrid Sliding Window Attention (SWA), (4) Hybrid Sliding Chunk Attention (SCA), (5) Pure KSA, (6) hybrid-KSA, all of the hybrid model variants with a 3:1 mixing ratio. 9
Kwai Summary Attention Technical Report
Table 2 | Training model configurations of our KSA. Configuration Number of layers Hidden size Intermediate size Attention heads (Q/KV) Head dimension Hybrid architecture ratio (KSA:Full) Summary chunk size Sliding chunk number Tied embeddings
From Scratch
Continual Pretraining
24 2048 6144 16/16 128 3:1 8 128 False
36 2560 9728 32/8 128 3:1 8 128 True
Table 3 | Training hyperparameter settings. Configuration Sequence Length Stages Token Budget per Stage Max Learning Rate Min Learning Rate RoPE Theta Optimizer Weight Decay LR Schedule Gradient Clipping
From Scratch
Continual Pretraining
8K / 32K / 64K / 128K 32K / 64K / 128K 250B / 50B / 50B / 50B 25B / 35B / 25B −4 −5 8K: 4 × 10 ; ≥32K: 1 × 10 All stages: 1 × 10−4 −7 1 × 10 4 8K: 10 ; ≥32K: 106 106 AdamW ( 𝛽1 = 0.9, 𝛽2 = 0.95) 0.01 WSD Yu et al. (2025) 1.0
Train-from-Scratch. We train a 1.9B-parameter language model from scratch using the architecture specified in Table 2 and the training hyperparameters in Table 3. Training follows a progressive length-extension schedule: the model is first trained on 250B tokens at a sequence length of 8K, followed by three subsequent stages of 50B tokens each at 32K, 64K, and 128K, respectively, resulting in a total of 400B tokens across all stages for stable long-context training. Continual Pretraining. For continual pretraining, we initialize from Qwen3-4B-base Qwen3 et al. (2025) with the configuration summarized in Table 2, and further train on 85B tokens using the hyperparameters in Table 3. Training is conducted in three stages with progressively increasing sequence lengths, allocating 25B, 35B, and 25B tokens at 32K, 64K, and 128K, respectively. 4.1.2. Datasets Pipeline The pretraining corpus is composed of six categories: Common Knowledge, Math, Code, STEM, General Reasoning, and Long Context. A comprehensive breakdown of data formulation is in Figure 6. Long Context is a primary data class, as it directly serves our goal of reinforcing the model’s long-sequence modeling capacity. We construct a hybrid mixture tailored for context-length extension by combining naturally long documents, i.e., books, long-form QA across code, math, and STEM; with synthetic sequences that probe long-range information tracking. We further incorporate structured benchmark-style sequences from long-context evaluation tasks to ensure robustness. 10
Kwai Summary Attention Technical Report
General Reasoning
Common Knowledge
Subcategory Breakdown Common Knowledge 11.2% 7.32% 2.24% 1.19% 0.37% 0.13%
Code
11.3% 11.20% 0.06%
Long Context
55.0% 55.00%
Code
th
Ma
inhouse FineWeb-Edu DCLM Zh FineWeb-Edu v2 FineWeb-Pro
inhouse OpenCoder
EM ST
xt
onte
gC
Lon inhouse
Math
8.4% 6.91% 0.26% 0.18% 0.15% 0.15% 0.09% 0.08% 0.08% 0.53%
STEM
8.4% 8.14% 0.19% 0.09% 0.02%
General Reasoning
5.6% 4.57% 0.48% 0.18% 0.17% 0.15% 0.05% 0.03%
inhouse OpenMathInstruct AI2 Math OpenMathInstruct-2 AM-DS-R1-1.4M StackMathQA Gemma 2 Math Prompt AutoMathText other OSS inhouse FreeLaw DM Math other OSS inhouse MNBVC WanJuan UltraTextbooks OpenWebText PubMed other OSS
Figure 6 | Overall Distribution of Pretraining Data Proportions. Common Knowledge forms the foundational layer, sourced from diverse high-quality web text Penedo et al. (2024), educational content, and curated open-source corpora. The sources span filtered CommonCrawl snapshots Su et al. (2025), web-extracted educational material, and high-quality open-domain text, ensuring broad coverage of general world knowledge. Math and Code are constructed from both natural and synthetic sources. The math subset integrates web-crawled mathematical content, competition problems, Chinese-language resources, and QA pairs, augmented with synthetic problem, i.e., solution sequences to improve coverage of rare reasoning patterns. The code subset combines GitHub repositories, competition solutions, code-related QA pairs, and synthetic generation tasks spanning multiple programming languages. STEM and General Reasoning broaden the model’s domain expertise through academic papers, textbook-style content, encyclopedia entries, long-form books, and structured QA corpora. Domainspecific scientific documents and exam-style data further strengthen coverage of technical disciplines. The training pipeline follows the data mixture ratios shown in Figure 6 and progressively extends the context length across stages. In train-from-scratch experiments, the context length increases from 8K to 32K, 64K, and finally 128K tokens. In CPT experiments, training starts directly at 32K and advances to 128K, with all stages using the same ratios of data classes mixture. 4.1.3. Evaluation Benchmark We categorize the evaluation benchmarks into two groups, targeting long-context modeling ability and general language capability accordingly. Long-Context Benchmarks: RULER Hsieh et al. (2024) is a comprehensive long-context evaluation suite that assesses models across multiple dimensions including retrieval, multi-hop tracing, aggregation, and question answering, with configurable context lengths up to 128K tokens. General Capability Benchmarks: We evaluate general capabilities across three domains: general
11
Kwai Summary Attention Technical Report
FFN ❄
FFN 🔥
FFN 🔥
Layer Norm ❄
Layer Norm 🔥
Layer Norm 🔥
⋯ ⋯
⨂
⋯
Annealing
⋯
Summary Attn 🔥
Text Attn ❄
⋯
Summary Attn 🔥
Layer Norm ❄
Text Attn 🔥
Layer Norm 🔥
⋯
(a)Add independent summary attention parameter for stable CPT warmup.
⋯
⋯
Text Attn 🔥
Layer Norm 🔥
⋯
(b)Annealing the additional summary attention parameter for CPT warmup.
⋯
(c)Full parameter tuning for CPT/Scratch.
Figure 7 | The training recipes for CPT warmup: (a) insert additional Q/K/V learnable weights to adapt the summary token compression; (b) conduct the annealing strategy to drop the summary weights; (c) full parameter tuning on large data corpus. Residual connections are omitted for brevity. knowledge, mathematics, and code. For general knowledge, MMLU Hendrycks et al. (2021a) and its extensions, CMMLU Li et al. (2024), C-Eval Huang et al. (2023), and MMLU-Pro Wang et al. (2024), assess knowledge breadth and reasoning across diverse domains via multi-choice questions, where CMMLU and C-Eval provide Chinese benchmarks spanning multiple disciplines and examstyle tasks, and MMLU-Pro introduces more challenging, professionally curated questions with reduced data contamination. For mathematics, GSM8K Cobbe et al. (2021), CMATH Wei et al. (2023), and MATH Hendrycks et al. (2021b) cover problems from grade-school multi-step arithmetic to competition-level symbolic reasoning, including Chinese settings and diverse problem formats requiring multi-step derivations. For code, MBPP Austin et al. (2021) and HumanEval Chen et al. (2021) evaluate program synthesis from natural language, with HumanEval further emphasizing functional correctness through unit tests on more complex and realistic coding tasks. 4.1.4. Training Recipes To adapt summary attention with existing architecture seamlessly, we devise a three stage training stages for CPT: summary token adaptation, parameter annealing, and sequence length extension: Summary token adaptation for CPT. As described earlier, summary tokens are designed to distill chunk-wise semantic information and serve as compressed priors for distant contexts. To endow these tokens with dense, informative representations, we devise a multi-granularity distillation strategy that aligns the KSA student with a vanilla full-attention teacher at three levels: layer-wise, distribution-wise, and objective-wise. Concretely, we instantiate the summary token S as a new vocabulary entry and equip each summary layer with independent attention parameters 𝑊S𝑄 , 𝑊S𝐾 , and 𝑊S𝑉 . Layer-wise attention score alignment. Let 𝑊 𝑄 , 𝑊 𝐾 , 𝑊 𝑉 denote the pretrained attention projections, and let 𝑊S𝑄 , 𝑊S𝐾 , 𝑊S𝑉 denote the independent summary matrices introduced for KSA. Let T and T̂ denote the original and summary-augmented input sequences (Equation 1), S ⊂ {1, . . . , | T̂ |} index the summary positions. The teacher branch projects input tensors using the pretrained weights, 𝑋 = T (𝑊 𝑋 ) ⊤ ,
𝑋 ∈ { 𝑄, 𝐾, 𝑉 } .
(7)
The student processes its input by employing position-specific projections: text positions retain the pretrained weights, whereas summary tokens route through the newly introduced parameters, 12
Kwai Summary Attention Technical Report
( 𝑋ˆ𝑡 =
T̂𝑡 (𝑊 𝑋 ) ⊤ ,
𝑡 ∉ S,
𝑋 ⊤
𝑡 ∈ S,
T̂𝑡 (𝑊S ) ,
𝑋 ∈ { 𝑄, 𝐾, 𝑉 } .
(8)
Attention patterns also diverge between them. The teacher performs standard full attention:
√
𝑂 = softmax 𝑄𝐾 ⊤ / 𝑑 𝑉,
(9)
whereas the student adopts the KSA pattern MKSA , in which each text token attends to its local sliding-chunk window together with preceding summary tokens, and each summary token attends only to the text chunk it compresses; as detailed in Section 2.2: √ ˆ = softmax 𝑄 ˆ 𝐾ˆ ⊤ / 𝑑 + MKSA 𝑉ˆ . 𝑂 (10) ˆ (which have no To align the intermediate representations, we discard the summary rows of 𝑂 | T | × 𝑑 ˆ counterpart in T ), denoted 𝑂 T ∈ ℝ , and employ a mean squared error (MSE) loss: 𝐿 1 ∑︁ ˆℓ 2 , 𝑂ℓ − 𝑂 LMSE = T 2 𝐿 · |T | ℓ=1
(11)
where 𝐿 is the number of transformer blocks and |T | is the sequence length. Distribution-wise regularization. While LMSE constrains intermediate layer outputs, it does not directly enforce consistency in the final predictive distributions. We therefore introduce a KL regularizer on the output logits to provide a higher-level alignment between the KSA student and the full-attention teacher. Let 𝑊ℎ denote the shared LM head, and ℎ𝐿 and ℎˆ𝐿 be the final-layer hidden states of the teacher and student, respectively. The two predicted distributions along with KL regularizer are: 𝑝 = softmax ℎ 𝐿𝑊ℎ⊤ ,
ˆ𝑝 = softmax ℎˆ𝐿𝑊ℎ⊤ ,
LKL = KL( 𝑝 ∥ ˆ𝑝) =
∑︁ 𝑣
𝑝𝑣 𝑝𝑣 log
.
ˆ𝑝𝑣
(12)
Objective-wise training. The total distillation objective combines the language modeling loss on the student with the two alignment terms to ensure warmup eventually benefits model performance: L = LLM + 𝛼 LMSE + 𝛽 LKL ,
(13)
where 𝛼 and 𝛽 are hyperparameters tuned on a validation split. Analysis. All three loss terms, LMSE , LKL , and LLM ; are computed exclusively over text token positions. This design ensures dimensional consistency between the teacher and student outputs (as the teacher lacks summary positions) while preserving semantic fidelity. Notably, restricting the loss to text positions does not cut off gradient flow to the summary parameters: the distinctive attention pattern of text tokens, as defined in Equation (3), inherently captures the semantics of summary tokens, enabling gradients to propagate through the attention interactions. Through these three complementary, multi-granularity distillation objectives, the independent summary parameters learn to approximate the attention distribution of vanilla full attention as closely as possible, establishing a solid foundation for subsequent full-parameter training and scaling to longer sequences.
13
Kwai Summary Attention Technical Report
Parameter annealing for CPT. To avoid introducing additional parameters that would increase inference cost, we propose a parameter annealing strategy that gradually absorbs the independent summary parameters into the main LLM weights. Specifically, at each summary position, we perform two QKV projections on the same hidden state: one using the shared LLM weights, gaining ( 𝑞main , 𝑘main , 𝑣𝑠main ), and one through the independent summary weights, yielding ( 𝑞𝑠 , 𝑘𝑠 , 𝑣𝑠 ). The QKV 𝑠 𝑠 triplets fed into the single attention computation are then linearly interpolated: 𝑥˜𝑠 = 𝜆 𝑥 𝑠 + (1 − 𝜆 ) 𝑥 𝑠main ,
𝑥 ∈ {𝑞, 𝑘, 𝑣},
(14)
The interpolation coefficient follows an iteration-dependent schedule: 1, 𝑠 − 𝑠start , 𝜆 ( 𝑠) = 1 − 𝑠end − 𝑠start 0,
𝑠 ≤ 𝑠start , 𝑠start < 𝑠 < 𝑠end ,
(15)
𝑠 ≥ 𝑠end ,
where 𝑠 denotes the current training step, and 𝑠start , 𝑠end define the annealing window. When 𝑠 ≤ 𝑠start , the summary positions are governed entirely by the independent parameters; when 𝑠 ≥ 𝑠end , they rely exclusively on the main LLM weights, allowing the auxiliary parameters to be removed at inference time without any architectural modification. This strategy provides a smooth curriculum that facilitates the transition from a dedicated summary head to a fully shared representation. Sequence length extension for CPT/Scratch. For CPT settings, introduction of summary tokens stops at the context length of 32K, allowing the newly added parameters to learn stable representations. We then extend the context in two stages: 64K for 35B tokens, followed by 128K for 25B tokens. This staged schedule enables the summary mechanism to adapt to longer contexts incrementally, equipping KSA with the ability of handling extremely long sequences. For the Scratch setting, we directly share the attention weights of summary token and text token at the begin, and then training 250B/50B/50B/50B tokens with the sequence length 8K/32K/64K/128K. Note that we change the RoPE Theta from 104 to 106 after the 8K training finished. 4.2. Continual Pre-Training Performance Comparison We evaluate the effectiveness of KSA under the CPT setting, comparing full KSA and Hybrid-KSA against four representative baselines: Full attention, sliding window attention (Hybrid-SWA), sliding chunk attention (Hybrid-SCA), and linear attention (Hybrid-Linear). 4.2.1. Long-Context Benchamrk We first examine KSA on RULER Hsieh et al. (2024), a long-context stress test where it is designed to excel, against the previously introduced baselines. From Table 4, we can draw several conclusions: i) Hybrid-KSA demonstrates the strongest long-context retrieval capacity among all baselines. Specifically, it achieves the best results at RULER-4K (92.97), RULER-32K (86.65), and RULER-128K (71.67), outperforming the Full attention baseline while operating with substantially lower cost. ii) Summary tokens effectively compress and propagate distant context information when full attention becomes prohibitive. Notably, at the longest 128K context, Hybrid-KSA surpasses Full attention by +5.81 points and outperforms the strongest hybrid baseline (Hybrid-SWA) by +5.40 points.
14
Kwai Summary Attention Technical Report
Table 4 | For CPT setting, hybrid-KSA achieves superior long-context retrieval across all RULER lengths, while both KSA variants closely match or exceed Full attention on knowledge, math, and coding benchmarks, establishing the smallest performance gap among all sub-quadratic alternatives. Baselines
Benchmarks
Ours
Full Hybrid-SWA Hybrid-SCA Hybrid-Linear KSA Hybrid-KSA Long-Context Retrieval RULER-4K 92.88 RULER-8K 91.38 RULER-16K 89.12 RULER-32K 84.74 RULER-64K 78.16 RULER-128K 65.86
91.30 88.03 82.87 78.94 73.88 66.27
86.02 84.28 80.67 76.89 68.88 60.94
86.39 83.86 78.06 76.48 73.50 67.98
91.55 86.78 84.78 80.30 76.09 66.81
92.97 90.53 88.86 86.65 76.04 71.67
General Knowledge MMLU CMMLU C-Eval MMLU-Pro
71.83 75.00 73.66 46.36
70.57 73.69 72.36 45.23
69.83 72.59 71.66 45.11
64.33 68.41 67.42 38.83
70.73 73.29 72.14 45.70
70.50 72.63 72.66 45.39
Mathematics CMath GSM8K MATH
83.41 82.75 47.48
84.84 81.92 48.24
83.16 80.10 47.45
79.09 72.44 42.57
84.58 81.09 48.15
84.25 79.50 47.56
Code MBPP HumanEval
61.30 58.54
61.70 61.89
59.60 61.89
55.30 54.58
61.50 60.97
62.20 62.50
Average
73.50
72.12
69.94
67.28
72.30
73.59
iii) KSA’s way of information aggregation is a more faithful approximation of full attention than fixed-window or linearized alternatives. Compared to other hybrid variants (SWA, SCA, Linear), our KSA and Hybrid-KSA models consistently lead by clear margins across all RULER lengths. 4.2.2. General Benchmarks To test if our architectural design and training procedure preserve model’s general ability on real-world problems, we further validate KSA on multiple general benchmarks. The evaluation covers three knowledge domains: general knowledge (MMLU, CMMLU), mathematics (CMath, GSM8K), and code generation (MBPP). From Table 4, we observe the following: i) CPT of summary attention preserves the pretrained model’s general world knowledge. On MMLU and CMMLU, the full KSA model achieves 70.73 and 73.29, closely matching the Full attention upper bound (71.83 / 75.00), and substantially outperforms Hybrid-Linear (64.33 / 68.41), which suffers from the limited representation capacity of fixed-size memory updates. ii) KSA is capable of mathematical reasoning. The KSA model attains 84.58 on CMath, surpassing Full attention (83.41), and remains competitive on GSM8K (81.09 vs. 82.75). iii) Our method robustly scales to real-world coding problems. Hybrid-KSA achieves the best MBPP score (62.20) across all configurations, including Full attention. iv) The proposed approaches consistently yield smaller capability gaps relative to Full attention than other sub-quadratic baselines across all general-purpose benchmarks. 15
Kwai Summary Attention Technical Report
Table 5 | Hybrid-KSA consistently outperforms full and sub-quadratic baselines, achieving superior long-context scalability and strong general performance in from-scratch training. Baselines
Benchmarks
Ours
Full Hybrid-SWA Hybrid-SCA Hybrid-GDN KSA Hybrid-KSA Long-Context Retrieval RULER-4K 76.08 RULER-8K 72.85 RULER-16K 73.24 RULER-32K 69.06 RULER-64K 65.32 RULER-128K 48.75
74.54 71.69 69.54 67.86 63.03 56.64
77.72 75.22 72.55 67.74 63.54 58.01
79.83 76.01 74.04 70.41 69.39 59.87
70.44 65.91 66.74 62.54 57.13 39.29
80.65 73.35 74.07 72.30 69.95 65.35
General Knowledge MMLU CMMLU C-Eval MMLU-Pro
44.99 44.41 44.28 19.48
46.84 45.89 43.54 20.46
46.77 46.42 47.62 20.10
46.23 47.19 45.54 21.22
46.83 45.59 45.69 21.72
46.83 46.88 44.13 22.52
Mathematics CMath GSM8K MATH
55.33 48.29 23.38
54.83 47.46 31.46
62.33 52.39 28.82
58.00 50.95 33.30
58.50 54.81 30.04
61.83 59.14 36.92
Code MBPP HumanEval
30.60 25.61
30.00 28.05
31.60 26.83
34.80 27.44
35.80 29.88
36.40 31.71
Average
49.44
50.12
51.84
52.95
48.73
54.80
4.3. Train-from-scratch Performance Comparison We further evaluate KSA in a train-from-scratch setting, where all modules are optimized without pretrained initialization. This setting isolates the effect of the open-sourced model capabilities based on massive corpora, providing a stricter test of scalability and learning dynamics. We compare KSA and Hybrid-KSA against Full, Hybrid-SWA, Hybrid-SCA, and Hybrid-GDN. 4.3.1. Long-Context Benchmark Our methods showcase promising performance on the RULER benchmark as in Table 5: i) Hybrid-KSA achieves the best overall performance, surpassing even Full attention by a large margin. It consistently outperforms all baselines across nearly all context lengths, achieving the best scores at RULER-4K (80.65), 32K (72.30), 64K (69.95), and 128K (65.35). Notably, at 128K, it surpasses Full attention by a large margin (+16.60), demonstrating superior scalability. ii) KSA significantly improves robustness at extreme context lengths. While Full attention and other hybrid baselines degrade rapidly as sequence length increases: e.g., Full: 76.08 → 48.75, Hybrid-KSA maintains reliably performing, i.e., 80.65 → 65.35, indicating strong robustness for long-inputs. iii) Summary-based aggregation provides a stronger alternative to existing efficient attention designs. Compared to Hybrid-SWA and Hybrid-SCA, which are constrained by local windows or chunking, and Hybrid-GDN, which relies on compressed memory updates, Hybrid-KSA consistently achieves higher accuracy, especially in the long-context regime, e.g., +5.48 over GDN at 128K.
16
Kwai Summary Attention Technical Report
4.3.2. General Benchmarks We next evaluate general capabilities across knowledge, mathematics, and code benchmarks: i) Our method maintains competitive general knowledge performance. On MMLU and CMMLU, Hybrid-KSA achieves 46.83 and 46.88, matching or exceeding most baselines and remaining close to the best-performing methods, indicating no loss in general understanding ability. ii) KSA demonstrates strong gains in mathematical reasoning. Hybrid-KSA achieves the top on GSM8K (59.14) and MATH (36.92), outperforming Full attention by +10.85 and +13.54, respectively. This suggests our method can effectively support multi-step reasoning over long contexts. iii) Summary attention improves performance on code generation tasks. On MBPP and HumanEval, Hybrid-KSA achieves the highest scores (36.40 and 31.71), outperforming Full attention and all hybrid baselines, demonstrating strong capability in structured generation. iv) The proposed method provides a better efficiency–performance trade-off than prior sub-quadratic methods. Across all benchmarks, Hybrid-KSA consistently ranks among the top performers, while maintaining sub-quadratic complexity. Compared to other efficient attention mechanisms, it achieves a substantially smaller performance gap, or even improvements, relative to Full attention. 4.3.3. Training Loss and Evaluation Score Analysis We further analyze the training dynamics of Hybrid-KSA by examining both training loss and evaluation performance across multiple benchmarks under the train-from-scratch setting. Training Loss. Hybrid-KSA demonstrates best converge efficiency compared to all baselines: i) Hybrid-KSA achieves the lowest training loss throughout training. As shown in Figure 8a, at the end of training, Hybrid-KSA reaches the lowest loss (1.524), outperforming Hybrid-GDN (1.534), Hybrid-SWA (1.550), and Full attention (1.572). ii) The advantage becomes more pronounced in the long-tail regime. In the zoomed-in region (tokens ≥ 41.9B), Hybrid-KSA consistently stays below all baselines with a stable margin. Evaluation Scores. The improved optimization translates directly into stronger performance: i) Hybrid-KSA consistently matches or outperforms baselines across benchmarks. Across knowledge (MMLU, CMMLU, C-Eval), reasoning (GSM8K, MATH), and code (MBPP) tasks, Hybrid-KSA demonstrates competitive or superior performance throughout training. Notably, it achieves the best final scores on several tasks, including GSM8K, MATH, and MBPP. ii) Hybrid-KSA exhibits faster and more stable performance improvement. Compared to Full attention, which lags behind across most benchmarks, Hybrid-KSA shows consistently higher scores at intermediate training stages, suggesting improved learning efficiency. It also avoids the fluctuations observed in other hybrid baselines (e.g., Hybrid-SWA), leading to smoother and more reliable gains. iii) The advantage is particularly evident in reasoning-intensive tasks. On GSM8K, CMath, and MATH, Hybrid-KSA maintains a clear margin over Full attention and remains competitive with or better than other hybrid methods, indicating that the summary-based mechanism does not hinder, and may even enhance, multi-step reasoning ability.
17
Kwai Summary Attention Technical Report
(a) Hybrid-KSA has the best convergence efficiency during from-scratch training. Full
40
40
35 30 100 150 Tokens (B)
200
20
50
100 150 Tokens (B)
200
25
250
CMath
100 150 Tokens (B)
200
250
50
MATH
40 30
50
100 150 Tokens (B)
200
250
100 150 Tokens (B)
200
250
250
200
250
30
20
25 20 15
10 50
200
MBPP
15
20
100 150 Tokens (B)
35
25 Score
30
16
12 50
30 50 Score
Score
50
18
14
35
60
10
35 30
GSM8k
40
MMLU-Pro
22
40
35
250
Hybrid-SWA
45
30 50
C-Eval
Score
25
Hybrid-KSA
Score
45
Scratch Training Dynamics on 8k Context
Hybrid-GDN CMMLU
Score
45
Score
Score
MMLU
10 50
100 150 Tokens (B)
200
250
50
100 150 Tokens (B)
(b) Our method’s performance improves more rapidly than other models on multiple benchmarks.
Figure 8 | Hybrid-KSA delivers the best training dynamics, with superior convergence efficiency and faster performance gains across benchmarks in from-scratch training. 4.4. Needle-in-a-Haystack To delve deeper into KSA’s ability to retrieve specific information from long contexts, we conduct the Needle-in-a-Haystack benchmark Martin et al. (2023). In this task, a short factual statement, i.e., the needle, is embedded at varying depths within a long context, i.e., the haystack; and the model is asked to retrieve it via a targeted question. The retrieval success rate across different needle positions serves as a robust indicator of long-context information retrieval capability. 18
Kwai Summary Attention Technical Report
Table 6 | For CPT, Hybrid-KSA achieves the best overall results on RULER 128K subtasks, beating all other efficient baselines and Full attention on NIAH-Multivalue, VT, FWE, and SQuAD. Baselines
Subtasks Full NIAH-Single 100.00 NIAH-Multikey 75.00 NIAH-Multivalue 88.12 NIAH-Multiquery 95.62 VT 60.50 FWE 51.66 SQuAD 30.00
Ours
Hybrid-SWA Hybrid-SCA Hybrid-Linear 100.00 74.16 83.75 93.12 67.50 51.66 30.00
99.16 70.84 91.25 98.12 42.50 33.33 15.00
100.00 79.16 95.62 99.38 87.50 23.33 35.00
KSA
Hybrid-KSA
97.50 74.16 83.75 95.62 65.50 72.50 32.50
100.00 75.84 98.75 98.12 90.50 65.84 42.50
Figure 9 | NIAH-Single results after converting the model into a hybrid architecture and applying continual pre-training. Hybrid-KSA achieves near-perfect needle retrieval up to 128K context length, with only a minor drop from full accuracy at 128K. Figure 9 reports single-needle retrieval accuracy across context lengths from 4K to 128K and needle depths from 0% to 100%. Hybrid-KSA attention maintains near-perfect retrieval across all lengths and depths, with only a slight dip at 128K, suggesting that a small dose of full attention effectively compensates for the compression loss in extremely long sequences. Further quantitative analysis in Table 6 reveals that Hybrid-KSA consistently outperforms other methods with limited effective context across multiple subtasks: i) KSA achieves strong NIAH performance, even compared with Full attention. On NIAH-Multivalue, Hybrid-KSA attains 98.75, a substantial gain of +10.63 over Full attention (88.12) and the highest across all configurations. On NIAH-Multiquery, it scores 98.12, closely approaching the best result (99.38 by Hybrid-Linear) and surpassing Full attention (95.62). ii) Our method scales robustly to more challenging synthetic subtasks beyond simple retrieval. HybridKSA surpasses +30.0 over Full on VT and +14.18 over Full on FWE. iii) In summary, KSA serves as a high-fidelity compressed relay, enabling robust information retrieval and complex long-sequence reasoning across 128K tokens without the prohibitive overhead of full attention. 4.5. Kwai Summary Attention Design Ablation Analyses 4.5.1. Inference KV Cache and Speed Analysis To provide an intuitive understanding of KSA’s efficiency, we test the KV cache memory footprint across context lengths 16-128K and the decode throughput at 16K, and plot the results in Figure 10:
19
Kwai Summary Attention Technical Report
Full Attention(GQA)
Hybrid-SWA
Hybrid-Linear
(a) KV Cache Storage
Hybrid-KSA
(b) Decode Throughput 60 18.6
15 9.56
10
6.47
5
5.06
0
16k
40
49.8 46.1 42.2
41.3
39.5 35.9
35.6
1.37 1.29
32k
1.83
2.50 2.42
64k
3.38
27.5
20
Decode Step
0
25.4
19.9
11.0
10 128k
34.7
34.7 34.1
30.6
30
4.75 4.67
2.81 0.81 0.73 1.06
50.3
50
Tokens / s
KV Cache (GB)
20
16k
32k
64k
128k
Decode Step
Figure 10 | Hybrid-KSA reduces KV cache by up to 2.5× while maintaining competitive or higher decode throughput compared with other efficient baselines. i) Hybrid-KSA retains a substantially smaller cache footprint than Full attention At 128K, it consumes only 7.5 GB, which is 2.5× smaller than Full attention (18.6 GB). ii) Our method achieves higher decode throughput than other efficient baselines. At 16K, HybridKSA achieves a throughput of 1.06× relative to Full attention, surpassing both Hybrid-SWA and Hybrid-Ring-Linear which are 0.73× and 0.81× of Full attention respectively. iii) These results demonstrate that KSA offers a favorable trade-off: it compresses long-range context into a compact state, reclaiming memory without sacrificing decoding speed. 4.5.2. Hybrid-KSA Configuration Analysis We conduct ablations on key design choices of Hybrid-KSA, including chunk number N, chunk size S, and the hybrid ratio between summary and full attention; and put the results in Table 7. Chunk Number (N). Varying N controls the granularity of summarization. As shown in Table 7, N=128 strikes a balance between stable long-context ability and competitive general performance: i) Moderately increasing N improves long-context performance, but larger values yield diminishing or even negative returns. The RULER average improves from 82.80 to 83.16 as N increases from 32 to 64. However, further increasing N does not bring consistent gains and can degrade performance at longer ranges: e.g., RULER-64k: 76.35 at N=128 vs. 65.73 at N=256. ii) N has limited impact on general benchmarks. Performance remains largely stable across configurations, with slight improvements at larger N: e.g., reasoning avg: 75.89 at N=256). Chunk Size. Chunk size determines the local receptive field within each segment. We observe a clear trade-off between long-context modeling and general capability: i) Smaller chunks yield the stronger long-context performance. For example, S=8 achieves the best RULER average (82.97) and 64K performance (76.35). ii) Larger chunks improve general performance at the cost of long-context retrieval. For instance, S=32 attains the best reasoning average (75.23) and strong GSM8K performance (81.23), while its
20
Kwai Summary Attention Technical Report
Table 7 | The default setting achieves the best balance between long-context and general performance. Long-Context (RULER)
Configuration 4K
8K
16K
32K
64K
Knowledge
Reasoning & Code
Avg. MMLU CMMLU Avg. GSM8K CMath MBPP
(a) Chunk Num ( 𝑁 ) 𝑁 = 32 91.02 89.46 84.42 78.35 70.74 82.80 𝑁 = 64 93.36 88.65 87.21 76.91 69.69 83.16 𝑁 = 128 (Default) 88.69 88.01 83.62 78.19 76.35 82.97 𝑁 = 256 92.05 88.30 83.86 79.40 65.73 81.87
Avg.
69.93 70.43 70.18 69.94
72.68 72.27 72.16 72.53
71.30 71.35 71.17 71.23
81.23 80.13 80.75 81.16
83.59 83.00 82.48 84.92
60.50 75.11 61.20 74.78 60.20 74.48 61.60 75.89
88.69 88.01 83.62 78.19 76.35 82.97 88.77 83.78 82.20 77.34 71.97 80.82 90.21 84.66 80.26 77.31 75.09 81.50 86.50 81.61 78.12 72.63 70.09 77.79
70.18 69.75 69.91 69.83
72.16 72.50 71.99 72.28
71.17 71.13 70.95 71.05
80.75 80.81 81.23 80.48
82.48 83.59 83.25 81.84
60.20 74.48 61.20 75.20 61.20 75.23 62.10 74.80
(c) Hybrid Ratio (Summary : Full) 1:1 84.50 82.85 82.16 74.76 69.36 78.72 3:1 (Default) 88.69 88.01 83.62 78.19 76.35 82.97 5:1 93.75 89.12 87.54 78.47 70.31 83.84 8:1 90.80 85.88 78.94 71.83 66.19 78.73
69.39 70.18 69.47 68.31
72.61 72.16 72.55 71.06
71.00 71.17 71.01 69.69
81.09 80.75 79.83 79.98
84.50 82.48 83.41 82.91
60.40 75.33 60.20 74.48 59.20 74.15 58.10 73.67
(b) Chunk Size (𝑆) 𝑆 = 8 (Default) 𝑆 = 16 𝑆 = 32 𝑆 = 64
RULER average drops to 81.50. iii) The default choice S=8 provides a balanced anchor point, slightly favoring long-context modeling while retaining strong general capability. Hybrid Ratio. Hybrid ratio is crucial for model to leverage different architectures’ strengths: i) Increasing the proportion of summary attention improves long-context performance but weakens general capabilities. For example, moving from 3:1 to 5:1 improves the RULER average (82.97 → 83.84), but degrades reasoning (74.48 → 74.15) and code performance (MBPP: 60.20 → 59.20). ii) Increasing full attention enhances general performance at the cost of long-context retrieval. For instance, a 1:1 ratio improves reasoning (Cmath: 84.50) but significantly reduces long-context performance (RULER avg: 78.72). iii) The default choice 3:1 provides a balanced trade-off, achieving strong long-context performance while maintaining competitive general capability. 4.5.3. Per-Layer Attention Pattern Analysis To understand why hybrid architecture benefits long-context retrieval, we visualize the per-layer attention distributions of KSA and Hybrid-KSA on an out-of-window NIAH example in Figure 11. We compare the first block, i.e., shallow layers L0-3 and the last block i.e., deep layers L24-27; within each row, the y-axis is shared across the two models so that attention magnitudes are directly comparable. Two qualitative patterns emerge in Hybrid-KSA that are absent in KSA: i) Hybird SA attends to summary tokens more frequently. In shallow SA layers (most clearly in L2), the hybrid model develops a periodic “comb” pattern that attends to past summary tokens at every chunk boundary, while KSA exhibits only a single weak peak close to the needle position. ii) The interleaved full-attention layers (L3, L27, labeled “[F]”) serve as cross-chunk integrators. Their attention maps show a sharp spike at the needle chunk while KSA remains flat, providing an
21
Kwai Summary Attention Technical Report
Figure 11 | Per-layer attention patterns reveal that Hybrid-KSA enhances cross-chunk retrieval through “comb-like” summary attention and full-layer integration. explicit token-level retrieval indicator that KSA must instead approximate only summary tokens. iii) Both models focus heavily on the earliest summary tokens in the deep block with notably different shapes. Hybrid-KSA’s sink spreads across roughly six chunks, while the pure model’s sink collapses onto only two, suggesting that the hybrid distributes its “register” role over a wider sink basin. Taken together, the shallow comb pattern and the full-layer integrator provide two concerted retrieval indicators that KSA lacks, which we hypothesize is the key reason Hybrid-KSA retrieves out-of-window needles more reliably thus possessing robust long-context capacity.
5. Related Works In this section, we briefly review the latest efforts of long-context modeling. 5.1. Efficient Attention Variants Multi-head attention (MHA) Vaswani et al. (2017) has become a cornerstone of modern deep learning, serving as a key building block across language Radford et al. (2018); Devlin et al. (2019), vision Dosovitskiy (2020); Ramesh et al. (2021), and video generation Peebles and Xie (2023). Despite its remarkable success, MHA faces two major bottlenecks when scaling to long contexts: i) the quadratic cost of computing the full 𝑁 × 𝑁 attention score matrix during training and long-sequence prefill; and ii) the linearly growing KV cache during auto-regressive generation, which imposes 22
Kwai Summary Attention Technical Report
substantial memory and bandwidth pressure. To address these limitations, prior works have explored efficient attention mechanisms from several complementary directions. The first line of work reduces the memory footprint of KV cache during decoding. Multi-query attention (MQA) Shazeer (2019) lets all query heads share a single key-value head, while groupedquery attention (GQA) Ainslie et al. (2023) generalizes this idea by allowing groups of query heads to share key-value heads. Multi-head Latent Attention (MLA) DeepSeek-V2 et al. (2024) further compresses KV states by projecting them into a low-dimensional latent space. With matrix absorption during inference, MLA only stores compact latent vectors in the KV cache, substantially reducing memory overhead while preserving much of the expressiveness of MHA. These methods are effective for memory-efficient decoding, but they do not directly eliminate the quadratic computation during long-context prefill, and they still rely on token-level cache to access historical context. The second category reduces attention computation by limiting the attention scope. Local attention mechanisms, such as sliding-window and sliding-chunk approaches Child et al. (2019); Beltagy et al. (2020); Jiang et al. (2023); Gemma2 et al. (2024), restrict each query to attend only to KV states within a local window of size 𝑊 , reducing the complexity from O ( 𝑁 2 ) to O (𝑊 𝑁 ). These methods are simple and hardware-friendly, but weaken access to distant contexts that are crucial for long-context reasoning and retrieval. Sparse attention variants Liu et al. (2025); Yuan et al. (2025); Zhang et al. (2025); Yang et al. (2026) address this limitation by selecting a small subset of 𝐾 important KV tokens or blocks according to heuristic scores, learned routing, or attention statistics, and computing attention only over the selected subset. This reduces the attention cost to O ( 𝐾 𝑁 ) while potentially preserving selected long-range dependencies. However, sparse attention typically still requires maintaining a large KV cache, and dynamic token or block selection introduces additional routing, indexing, and hardware-efficiency challenges. The third direction replaces full softmax attention with linear, recurrent, or state-space sequencemixing mechanisms. Kernelized linear attention methods approximate or reformulate the softmax kernel with feature maps, avoiding explicit materialization of the full attention matrix. For example, Retentive Networks Sun et al. (2023) and Eagle Peng et al. (2024), maintain recurrent memory states with decay mechanisms, enabling efficient step-wise inference. In parallel, state-space and gated recurrent models, such as Mamba Gu and Dao (2024) and Gated Delta Net Yang et al. (2025), model long-range dependencies through selective fixed-dimensional states whose memory usage does not grow with the context length. These methods offer attractive inference efficiency, but compressing the entire history into a fixed-size state can lose fine-grained token-level information, often leading to performance degradation on tasks that require precise long-context retrieval or reasoning. Overall, existing efficient attention variants still face a fundamental performance-efficiency tradeoff. KV-cache compression reduces memory usage but does not fully address prefill cost; local and sparse attention reduce computation but may either lose distant information or retain large KV caches; and linear or recurrent variants improve efficiency but may sacrifice fine-grained retrieval ability. 5.2. Hybrid Attention Architecture To better balance the strengths of full attention and efficient attention, recent research has increasingly explored hybrid attention architectures. Rather than replacing all softmax attention layers with a single efficient mechanism, these models combine dense, sparse, local, recurrent, or state-space components within one architecture. A widely adopted practice is to allocate a small fraction of full-attention layers among more efficient layers, thereby preserving global recall and expressiveness while reducing KV-cache cost and computation. Hybrid-H3 Fu et al. (2023) builds a predominantly SSM-based backbone with only a small number 23
Kwai Summary Attention Technical Report
of attention layers, showing that a limited dose of softmax attention can mitigate the expressivity gap of pure SSMs in language modeling. RecurrentGemma Botev et al. (2024) combines gated linear recurrences with sliding-window attention at a fixed ratio, leveraging the complementary benefits of recurrent memory and local sparse attention. Mamba-2-Hybrid Waleffe et al. (2024) constructs an 8B-parameter model from a mixture of Mamba-2 layers, full-attention layers, and MLP layers, demonstrating that a modest attention budget can support strong long-context reasoning. Jamba Lieber et al. (2024); Team et al. (2024) further scales the hybrid paradigm by interleaving Mamba blocks with Transformer and Mixture-of-Experts (MoE) layers, enabling long-context modeling with improved throughput and memory efficiency. Beyond layer-level interleaving, several works introduce more structured hybridization strategies. Zamba Glorioso et al. (2024a); b shares a global attention block across multiple Mamba blocks, obtaining benefits from global attention with improved parameter efficiency. YOCO Sun et al. (2024) separates sequence encoding and generation through a self-decoder and a cross-decoder: the selfdecoder compresses the input sequence into compact KV states, while the cross-decoder repeatedly reuses them through cross-attention during generation, substantially reducing memory consumption across decoder layers. Hymba Dong et al. (2025) performs finer-grained hybridization by assigning different KV heads to softmax attention and SSM modules, allowing token-level information mixing across heterogeneous memory mechanisms. These hybrid architectures reveal an important trend: different efficient attention mechanisms are often complementary rather than mutually exclusive. Full attention provides high-fidelity global retrieval, local attention preserves nearby details with low cost, and recurrent or state-space modules offer compact memory for efficient generation. However, most existing hybrid designs combine these components at the layer, block, or head level, and rely on carefully engineered architectural schedules to balance efficiency and expressiveness. They usually lack an intrinsic interface that explicitly connects local fine-grained attention with compressed distant context. In contrast, our proposed KSA introduces summary tokens as an explicit bridge between local information and long-range memory. Instead of simply interleaving full and efficient layers, KSA performs hybridization inside the attention mechanism itself: local tokens preserve fine-grained short-range dependencies, while summary tokens provide compact access to distant contexts. This intrinsic mixing mechanism enables KSA to reduce long-context attention and KV-cache overhead while maintaining strong modeling capacity for both local details and global information.
6. Conclusion & Future Works In this work, we revisit the long-context efficiency and highlight the sequence-level KV cache compression: rather than pursuing either strict linear KV cache reduction (GQA/MLA), we advocate sequence-level semantic compression with a moderate ratio 𝑘, which maintains an 𝑂 ( 𝑛/𝑘) dependence on sequence length while preserving full addressability and interpretability over the distant history. Specifically, we propose Kwai Summary Attention (KSA), a drop-in efficient attention variant that i) periodically emits a learnable summary token for every text chunk, ii) routes local attention through a sliding chunk attention and distant attention through the accumulated summary tokens, and iii) guarantees a non-overlapping information conflict so that every past chunk is either fully visible as raw text or accessed exclusively through its summary, avoiding both information loss and double counting. Further, to support efficiency training and inference, we contribute: i) a block-sparse training and prefill kernel that exploits the structured sparsity of the KSA mask and avoids materializing the dense mask; ii) a contiguous, concatenation-free KV cache layout that encodes the visibility rule into memory addressing itself, such that every decoding step reads a single memory slice without
24
Kwai Summary Attention Technical Report
gather or dynamic reshape; and ii) a hybrid-KSA architectural recipe that interleaves full-attention and KSA layers to balance retrieval fidelity with memory and compute cost. On eleven benchmarks spanning long-context retrieval, knowledge, mathematics, and coding, our Hybrid-KSA configuration consistently establishes the smallest quality gap to full attention among all sub-quadratic alternatives, while surpassing full attention on the most context-intensive tasks such as RULER-128K. These results support our motivation: sequence-level token compression provides empirically favorable perspective for reducing the KV cache, and it is especially has the potential to support the complex agentic RL. In the future, we will explore the following directions: i) Sparse summary attention. Our current design almost enabling the distant summary tokens are fully visible. Replacing this with a learned sparse retriever that selects summaries conditioned on the query may further improve long-context capability, particularly at extreme context lengths. ii) KSA post-training. We currently only perform the pre-training process, next we will exploring how KSA interacts with post-training, including supervised fine-tuning, preference optimization, reasoning reinforcement learning and multi-teacher OPD. iii) Scaling laws of compression ratio. The relationship between chunk size , model capacity, and task difficulty is not yet fully explored; establishing scaling laws for sequence-level compression, analogous to those for parameters and data, would guide a principle to scale next level intelligence. iv) Unifying with OneRec. With the KV Cache friendly architecture, building a world knowledge enriched user KV cache memory is a promising iteration direction. Our next step is to build a unified generative recommendation foundation model on top of KSA, in which arbitrarily long user behavior trajectories are compressed into a hierarchy of summary tokens while recent interactions remain fully accessible through the local sliding window. We believe such an integration of KSA and OneRec could bridge the gap between language-model-style world knowledge and recommendation-style user modeling, and serve as a foundation to build more smart recommendation system.
References Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017). “Attention is all you need”. In: NeurIPS. Vol. 30. Radford, A., K. Narasimhan, T. Salimans, I. Sutskever, et al. (2018). “Improving language understanding by generative pre-training”. In. Child, R., S. Gray, A. Radford, and I. Sutskever (2019). “Generating long sequences with sparse transformers”. In: arXiv preprint arXiv:1904.10509. Devlin, J., M.-W. Chang, K. Lee, and K. Toutanova (2019). “Bert: Pre-training of deep bidirectional transformers for language understanding”. In: NAACL, pp. 4171–4186. Shazeer, N. (2019). “Fast transformer decoding: One write-head is all you need”. In: arXiv preprint arXiv:1911.02150. Beltagy, I., M. E. Peters, and A. Cohan (2020). “Longformer: The long-document transformer”. In: arXiv preprint arXiv:2004.05150. Dosovitskiy, A. (2020). “An image is worth 16x16 words: Transformers for image recognition at scale”. In: arXiv preprint arXiv:2010.11929. Kaplan, J., S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020). “Scaling laws for neural language models”. In: arXiv preprint arXiv:2001.08361. Austin, J., A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. (2021). “Program Synthesis with Large Language Models”. In: arXiv e-prints, arXiv–2108. Chen, M., J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. (2021). “Evaluating large language models trained on code”. In: arXiv preprint arXiv:2107.03374.
25
Kwai Summary Attention Technical Report
Cobbe, K., V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. (2021). “Training verifiers to solve math word problems”. In: arXiv preprint arXiv:2110.14168. Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021a). “Measuring Massive Multitask Language Understanding”. In: ICLR. Hendrycks, D., C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021b). “Measuring Mathematical Problem Solving With the MATH Dataset”. In: NeurIPS. Ramesh, A., M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever (2021). “Zero-shot text-to-image generation”. In: ICML. Pmlr, pp. 8821–8831. Ainslie, J., J. Lee-Thorp, M. De Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai (2023). “Gqa: Training generalized multi-query transformer models from multi-head checkpoints”. In: EMNLP, pp. 4895– 4901. Fu, D. Y., T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. Re (2023). “Hungry Hungry Hippos: Towards Language Modeling with State Space Models”. In: ICLR. Huang, Y., Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, Y. Fu, et al. (2023). “C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models”. In: Advances in neural information processing systems 36, pp. 62991–63010. Jiang, A. Q., A. Sablayrolles, A. Mensch, C. Bamford, and D. S. Chaplot (2023). “Mistral 7B”. In: Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed, pp. 50–72. Martin, L., K. Chandrayan, G. Kamradt, L. Hurtado, A. Arkhangorodsky, I. E. Ashimine, P. Arivalagan, and P. Král (2023). “Needle In A Haystack - Pressure Testing LLMs”. In: GitHub repository. url: https://github.com/gkamradt/LLMTest_NeedleInAHaystack. Peebles, W. and S. Xie (2023). “Scalable diffusion models with transformers”. In: ICCV. Sun, Y., L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei (2023). “Retentive network: A successor to transformer for large language models”. In: arXiv preprint arXiv:2307.08621. Wei, T., J. Luan, W. Liu, S. Dong, and B. Wang (2023). “CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?” In: arXiv e-prints, arXiv–2306. Botev, A., S. De, S. L. Smith, A. Fernando, G.-C. Muraru, R. Haroun, L. Berrada, R. Pascanu, P. G. Sessa, R. Dadashi, et al. (2024). “Recurrentgemma: Moving past transformers for efficient open language models”. In: arXiv preprint arXiv:2404.07839. DeepSeek-V2, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Dengr, C. Ruan, D. Dai, D. Guo, et al. (2024). “Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model”. In: arXiv preprint arXiv:2405.04434. DeepSeek-V3, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. (2024). “Deepseek-v3 technical report”. In: arXiv preprint arXiv:2412.19437. Gemma2, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al. (2024). “Gemma 2: Improving open language models at a practical size”. In: arXiv preprint arXiv:2408.00118. Glorioso, P., Q. Anthony, Y. Tokpanov, A. Golubeva, V. Shyam, J. Whittington, J. Pilault, and B. Millidge (2024a). “The zamba2 suite: Technical report”. In: arXiv preprint arXiv:2411.15242. Glorioso, P., Q. Anthony, Y. Tokpanov, J. Whittington, J. Pilault, A. Ibrahim, and B. Millidge (2024b). “Zamba: A compact 7b ssm hybrid model”. In: arXiv preprint arXiv:2405.16712. Gu, A. and T. Dao (2024). “Mamba: Linear-Time Sequence Modeling with Selective State Spaces”. In: CoLM. Hsieh, C.-P., S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, and B. Ginsburg (2024). “RULER: What’s the Real Context Size of Your Long-Context Language Models?” In: CoLM.
26
Kwai Summary Attention Technical Report
Li, H., Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin (2024). “Cmmlu: Measuring massive multitask language understanding in chinese”. In: ACL, pp. 11260–11285. Lieber, O., B. Lenz, H. Bata, G. Cohen, J. Osin, I. Dalmedigos, E. Safahi, S. Meirom, Y. Belinkov, S. Shalev-Shwartz, et al. (2024). “Jamba: A hybrid transformer-mamba language model”. In: arXiv preprint arXiv:2403.19887. Penedo, G., H. Kydlíček, A. Lozhkov, M. Mitchell, C. Raffel, L. Von Werra, T. Wolf, et al. (2024). “The fineweb datasets: Decanting the web for the finest text data at scale”. In: NeurIPS. Vol. 37, pp. 30811–30849. Peng, B., D. Goldstein, Q. G. Anthony, A. Albalak, E. Alcaide, S. Biderman, E. Cheah, T. Ferdinan, K. K. GV, H. Hou, S. Krishna, R. M. Jr., N. Muennighoff, F. Obeid, A. Saito, G. Song, H. Tu, R. Zhang, B. Zhao, Q. Zhao, J. Zhu, and R.-J. Zhu (2024). “Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence”. In: CoLM. Qwen2, A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, X. Liu, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, Z. Guo, and Z. Fan (2024). “Qwen2 Technical Report”. In: arXiv preprint arXiv:2407.10671. Sun, Y., L. Dong, Y. Zhu, S. Huang, W. Wang, S. Ma, Q. Zhang, J. Wang, and F. Wei (2024). “You Only Cache Once: Decoder-Decoder Architectures for Language Models”. In: NeurIPS. Team, J., B. Lenz, A. Arazi, A. Bergman, A. Manevich, B. Peleg, B. Aviram, C. Almagor, C. Fridman, D. Padnos, et al. (2024). “Jamba-1.5: Hybrid transformer-mamba models at scale”. In: arXiv preprint arXiv:2408.12570. Waleffe, R., W. Byeon, D. Riach, B. Norick, V. Korthikanti, T. Dao, A. Gu, A. Hatamizadeh, S. Singh, D. Narayanan, et al. (2024). “An empirical study of mamba-based language models”. In: arXiv preprint arXiv:2406.07887. Wang, Y., X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. (2024). “Mmlu-pro: A more robust and challenging multi-task language understanding benchmark”. In: Advances in Neural Information Processing Systems 37, pp. 95266–95290. Zhang, K., J. Li, G. Li, X. Shi, and Z. Jin (2024). “Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges”. In: ACL, pp. 13643– 13658. Agarwal, S., L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, et al. (2025). “gpt-oss-120b & gpt-oss-20b model card”. In: arXiv preprint arXiv:2508.10925. Dong, X., Y. Fu, S. Diao, W. Byeon, Z. CHEN, A. S. Mahabaleshwarkar, S.-Y. Liu, M. V. keirsbilck, M.-H. Chen, Y. Suhara, Y. C. Lin, J. Kautz, and P. Molchanov (2025). “Hymba: A Hybrid-head Architecture for Small Language Models”. In: ICLR. Guo, D., D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, et al. (2025). “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning”. In: arXiv preprint arXiv:2501.12948. Kimi, Y. Bai, Y. Bao, Y Charles, C. Chen, G. Chen, H. Chen, H. Chen, J. Chen, N. Chen, et al. (2025). “Kimi k2: Open agentic intelligence”. In: arXiv preprint arXiv:2507.20534. Ling-Team, B. Han, C. Tang, C. Liang, D. Zhang, F. Yuan, F. Zhu, J. Gao, J. Hu, L. Li, et al. (2025). “Every attention matters: An efficient hybrid architecture for long-context reasoning”. In: arXiv preprint arXiv:2510.19338. Liu, A., A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, et al. (2025). “Deepseekv3. 2: Pushing the frontier of open large language models”. In: arXiv preprint arXiv:2512.02556. Qwen3, A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang,
27
Kwai Summary Attention Technical Report
J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025). “Qwen3 Technical Report”. In: arXiv preprint arXiv:2505.09388. Su, D., K. Kong, Y. Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro (2025). “Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset”. In: ACL, pp. 2459–2475. Yang, S., J. Kautz, and A. Hatamizadeh (2025). “Gated Delta Networks: Improving Mamba2 with Delta Rule”. In: ICLR. Yu, T., Z. Wang, C. Wang, F. Huang, W. Ma, Z. He, T. Cai, W. Chen, Y. Huang, Y. Zhao, et al. (2025). “Minicpm-v 4.5: Cooking efficient mllms via architecture, data, and training recipe”. In: arXiv preprint arXiv:2509.18154. Yuan, J., H. Gao, D. Dai, J. Luo, L. Zhao, Z. Zhang, Z. Xie, Y. Wei, L. Wang, Z. Xiao, et al. (2025). “Native sparse attention: Hardware-aligned and natively trainable sparse attention”. In: ACL, pp. 23078–23097. Zhang, J., C. Xiang, H. Huang, J. Wei, H. Xi, J. Zhu, and J. Chen (2025). “Spargeattn: Accurate sparse attention accelerating any model inference”. In: ICML. Anthropic (2026). Introducing Claude Opus 4.6. url: https://www.anthropic.com/news/ claude-opus-4-6. DeepSeek-AI (2026). DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. Kimi2.6 (2026). Kimi K2.6: Advancing Open-Source Coding. url: https://www.kimi.com/blog/ kimi-k2-6. LongCat, M., A. Gui, B. Li, B. Tao, B. Zhou, B. Chen, C. Zhang, C. Gao, C. Zhang, C. Han, et al. (2026). “Longcat-flash-thinking-2601 technical report”. In: arXiv preprint arXiv:2601.16725. OpenClaw (2026). “OpenClaw — Personal AI Assistant”. In: GitHub repository. url: https:// github.com/openclaw/openclaw. Qwen3.5 (2026). Qwen3.5: Accelerating Productivity with Native Multimodal Agents. url: https: //qwen.ai/blog?id=qwen3.5. Qwen3.6 (2026). Qwen3.6-Plus: Towards Real World Agents. url: https://qwen.ai/blog?id= qwen3.6. Xiao, B., B. Xia, B. Yang, B. Gao, B. Shen, C. Zhang, C. He, C. Lou, F. Luo, G. Wang, et al. (2026). “Mimo-v2-flash technical report”. In: arXiv preprint arXiv:2601.02780. Yang, S., W. Huang, R. Chu, Y. Xiao, Y. Zhao, X. Wang, M. Li, E. Xie, Y.-C. Chen, Y. Lu, S. Han, and Y. Chen (2026). “LongLive: Real-time Interactive Long Video Generation”. In: ICLR. Zeng, A., X. Lv, Z. Hou, Z. Du, Q. Zheng, B. Chen, D. Yin, C. Ge, C. Huang, C. Xie, et al. (2026). “Glm-5: from vibe coding to agentic engineering”. In: arXiv preprint arXiv:2602.15763.
28
Kwai Summary Attention Technical Report
7. Author List and Acknowledgement Core Contributors Chenglong Chu, Guorui Zhou* , Guowang Zhang, Han Li, Hao Peng, Hongtao Cheng* , Jian Liang, Jiangxia Cao† , Kun Gai, Lingzhi Zhou, Lu Ren, Qi Zhang, Ruiming Tang† , Ruitao Wang* , Xinchen Luo* , Yi Su* , Zhiyuan Liang, Ziqi Wang Contributors Boyang Ding* , Chengru Song, Dunju Zang, Hui Wang, Jiao Ou, Jiaxin Deng, Jijun Shi, Jinghao Zhang, Junmin Chen, Lejian Ren* , Minxuan Lv, Qianqian Wang, Qigen Hu* , Shiyao Wang* , Siyang Mao, Tao Wang, Xingmei Wang, Zhixin Ling, Ziming Li, Zixing Zhang* All the authors listed alphabetically by first name. † Project Leaders.
* individuals who have departed from our team.
29