Conceptio › Archive › arXiv CS
arXiv CSopen access

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

arXiv:2609.08318v1 [cs.SE] 8 Sep 2026

ZHENGRAN ZENG∗ , Peking University, China YIXIN LI∗ , Peking University, China RUI XIE† , Peking University, China WEI YE† , Peking University, China SHIKUN ZHANG† , Peking University, China The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the resolution of complex real-world SE tasks. However, the trial-and-error nature of these agents generates lengthy interaction trajectories, creating severe bottlenecks in terms of context window limits and cost. While context compression offers a potential remedy, prior approaches suffer from static pruning strategies and granularity mismatches, often failing to preserve the semantic dependencies and syntactic details crucial for SE tasks. To strictly preserve critical task evidence while reducing context length, we introduce AttnCompress, a dynamic attention-guided trajectory compression framework. Unlike existing approaches, AttnCompress bridges the gap between semantic integrity and dynamic adaptability through three key mechanisms: (1) structure-aware segmentation via perplexity (PPL) spikes to preserve the syntactic structure of code and logs; (2) relevance estimation using proxy attention weights to quantify the precise relevance of historical blocks to the agent’s current reasoning; and (3) a dynamic rolling window to re-evaluate and recall historical context as the task evolves. Extensive evaluation on SWE-Bench-Verified and Multi-SWE-Bench demonstrates that AttnCompress achieves a pass rate of 53.17%, outperforming prior state-of-the-art baselines while reducing token consumption by 21.6% and total costs by 33.6%. The framework proves to be model-agnostic and generalizes effectively across diverse programming languages. CCS Concepts: • Software and its engineering → Automatic programming. Additional Key Words and Phrases: Software Engineering Agents, Large Language Models, Context Compression ACM Reference Format: Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang. 2026. AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents. Proc. ACM Softw. Eng. 3, ISSTA, Article ISSTA058 (October 2026), 23 pages. https://doi.org/10.1145/3832149

1

Introduction

The advent of Large Language Models (LLMs) has catalyzed a paradigm shift in software engineering (SE), transitioning from human-centric assistance to Autonomous Software Engineering (ASE) agents [22, 43]. Capable of perceiving, reasoning, and acting, these agents have demonstrated ∗ Both authors contributed equally to this research. † Those authors are the corresponding authors.

Authors’ Contact Information: Zhengran Zeng, Peking University, Beijing, China, [email protected]; Yixin Li, Peking University, Beijing, China, [email protected]; Rui Xie, Peking University, Beijing, China, [email protected]; Wei Ye, Peking University, Beijing, China, [email protected]; Shikun Zhang, Peking University, Beijing, China, [email protected].

This work is licensed under a Creative Commons Attribution 4.0 International License. © 2026 Copyright held by the owner/author(s). ACM 2994-970X/2026/10-ARTISSTA058 https://doi.org/10.1145/3832149 Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:2

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

remarkable potential in resolving complex real-world GitHub issues, as evidenced by benchmarks like SWE-Bench [15] and its extensions [2, 42, 45]. More broadly, SE agents have also been applied to a wide range of software engineering tasks that require sustained interaction with large codebases, including repository-level code understanding, feature implementation and test generation [16, 38]. Unlike simple Q&A tasks, SE agents operate through long-horizon interactions, engaging in iterative “trial-and-error” workflows involving codebase analysis, file editing, test execution, and debugging [5]. This process inevitably generates lengthy interaction trajectories. However, this extended context poses a severe efficiency bottleneck. As the trajectory grows, the accumulation of verbose logs, redundant file contents, and obsolete error stacks leads to rapidly growing token consumption. Such overhead results in increased latency, substantial API expenditures, and potentially surpassing the maximum context length of LLMs, which poses a significant barrier to industrial scalability. Furthermore, the “Lost-in-the-Middle” phenomenon suggests that feeding excessive noise to the model can degrade its reasoning performance [13, 23]. To mitigate this, context compression has become essential. Current research broadly falls into three categories: heuristic-based pruning, which removes historical messages based on rules (e.g., ObsMask [20]); summarization-based methods, which utilize LLMs to condense history into shorter text (e.g., AgentDiet [36]); and selection-based methods, which employ a small model to predict and filter out less important tokens (e.g., Lingua [14]). While these methods alleviate the context burden to some extent, they suffer from two critical limitations when applied to the dynamic nature of SE tasks: First, inability to handle dynamic context dependencies. Existing methods typically employ a “static, one-pass” strategy, where compression is performed once based on the agent’s current view of the task, producing a finalized reduced context. For instance, ObsMask mechanically removes observations older than a fixed window, while AgentDiet performs step-wise trajectory rewriting: at each step, it invokes a cost-efficient LLM to rewrite a previous step (typically a fixed lag behind the current step) into a shorter form. Such approaches are irrevocable and neglect the non-linear nature of debugging. In SE tasks, agents often exhibit focus shifting [27, 39], such as abandoning a hypothesis about a database error to investigate a network configuration mentioned twenty turns earlier. Static methods sever these long-range semantic dependencies because once a block of context is deemed “irrelevant” and removed, it cannot be recovered when the agent’s focus shifts, leading to context-induced hallucinations. Second, granularity mismatch and precision loss. There is a trade-off between semantic integrity and compression rate that prior works fail to balance. On one hand, coarse-grained methods like ObsMask operate at the message level, forcing the retention of entire verbose logs even if only a single error line is relevant. Conversely, summarization-based methods like AgentDiet employ an additional LLM to rewrite context. While this reduces length, the generative nature of summarization poses a severe risk to SE tasks which demand verbatim accuracy. Such methods often abstract away critical details (e.g., specific line numbers or variable names) and, more critically, are prone to hallucinating [25] non-existent code behaviors or incorrect error references, thereby misleading the agent. Furthermore, relying on an extra strong LLM for every step incurs prohibitive computational overhead. On the other hand, fine-grained token-level selection (e.g., Lingua [29]) disregards the syntactic structure of code. Arbitrarily dropping tokens can break JSON objects or function definitions, rendering the context incomprehensible for the LLM. To address these challenges, we introduce AttnCompress, a dynamic attention-guided trajectory compression framework designed specifically for the dynamic workflows of SE agents. Our approach bridges the gap between semantic integrity and dynamic adaptability through three key mechanisms:

Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:3

(1) Structure-Aware Segmentation via PPL Spikes: To resolve the granularity mismatch, we utilize a small proxy model to monitor the Perplexity (PPL) fluctuations [31] of the output stream. By identifying PPL spikes, which naturally occur at semantic boundaries (e.g., the switch from code to error logs), we segment the trajectory into semantically coherent blocks rather than arbitrary tokens. This strategy mitigates the risk of breaking syntactic structures like JSON objects and function bodies after compression. (2) Relevance Estimation via Proxy Attention: Instead of relying on heuristic rules or heavy summarization models, we leverage the attention distribution of a cost-effective small model (e.g., Qwen3-4B-Instruct [30]). We calculate the attention weights projected from the generation start token of the next response to historical blocks, enabling us to quantify the precise relevance of each block to the immediate reasoning step. (3) Dynamic Context Maintenance: To overcome the static limitation, we introduce a rolling window mechanism. Unlike static pruning, our framework maintains a dynamic buffer where recent history is continuously re-evaluated against the agent’s shifting intent. This allows previously suppressed information to be recalled if it becomes relevant again, effectively mitigating the risk of information loss during goal shifting. Our contributions are as follows: • We propose a dynamic attention-guided trajectory compression framework that leverages the attention distribution of a cost-effective proxy model to quantify the relevance between historical context and the agent’s current work. • We introduce a PPL-based block segmentation algorithm that preserves the syntactic structure of code and logs during compression. • Extensive evaluation on SWE-Bench-Verified [15] and Multi-SWE-Bench-Flash [45] demonstrates that AttnCompress achieves a pass rate of 53.17%, outperforming prior SOTA Agent Diet while reducing token consumption by 21.6% and total costs by 33.6%. 2 2.1

Background & Related Work LLM-based SE Agents & Workflow

In Autonomous Software Engineering (ASE), LLM-based agents are commonly deployed to address GitHub issues, implement feature requests, comprehend codebases, and generate tests in an endto-end manner [22, 38, 41]. As illustrated in the left panel of Figure 1, the operational workflow typically follows an iterative ReAct (Reasoning and Acting) pattern [44]. The process begins with 1) constructing task input, where the user query is combined with detailed descriptions of available tools (e.g., search, edit, and bash utilities). Subsequently, the agent enters a cyclic interaction loop: • LLM Processing: Based on the current history, the LLM generates a Thought for reasoning and an Action to interact with the environment. • Environment Execution: The environment executes the action using specific tools (e.g., searching the codebase or running a script) and yields an Observation, which contains the execution results such as file contents or error logs. This cycle repeats until the issue is resolved (“Done!”). While this mechanism enables complex problem-solving, it incurs a significant context overhead. As shown in the statistical breakdown in the right panel of Figure 1, (derived from the trajectories in RQ1), the distribution of tokens in interaction traces is highly imbalanced. Observation tokens account for the vast majority (62.6%) of the total context, significantly outweighing Action (24.8%), Thought (8.6%), and Task Input (4%). This data underscores that verbose environmental feedback is the primary contributor to context bloat, making the compression of observations the critical bottleneck for efficiency. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:4

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang Token Distribution of Different Traces Components

ReAct Agent

Task Input (4%)

Trajectory Manager User Query

① Construct Task Input

Task Input Thought, Action

Tools Description

LLM

Observation . . .

② LLM processes request and returns Action

Thought (8.6%)

Environment

Observation

� Search Tool

Thought, Action

� Edit Tool

Observation . . . Done!

Action (24.8%)

� Codebase

Observation (62.6%)

� Bash Tool ③ Environment runs Action and yields tool output

Task Output

Fig. 1. The typical workflow of a ReAct-based SE agent (left) and the token distribution across trace components (right). The statistical analysis reveals that Observation content (e.g., logs and code retrieval) dominates the context window (62.6%), posing a major challenge for long-horizon tasks.

2.2

The Context Challenge in SE

While extensive context provides necessary information, it introduces two fundamental challenges specific to software engineering tasks: Information Density Variance and Dynamic Relevance. An Example from Instance django-11133

[Assistant] Let me examine the HttpResponse implementation in the response.py file to investigate how response content is converted to bytes: Tool call: str_replace_editor {"path": "/testbed/django/http/response.py", "command": "view"}

1 2 3 4

[Tool] Here's the result of running `cat -n` on /testbed/django/http/response.py: import datetime import io import json import mimetypes

5-347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380-803 804 805 806 807 808

)

... [manually omitted 343 lines for better present] ... domain=domain, secure=secure, expires="Thu, 01 Jan 1970 00:00:00 GMT", samesite=samesite,

343 Low-Information Lines

While the agent focuses on auth_service.py, the content of utils.py (viewed in Turn 1) is deemed irrelevant by static policies and compressed.

# Common methods used by subclasses def make_bytes(self, value): """Turn a value into a bytestring encoded in the output charset.""" # Per PEP 3333, this response body must be bytes. To avoid returning # an instance of a subclass, this function returns `bytes(value)`. # This doesn't make a copy when `value` already contains bytes. # Handle string types -- we can't rely on force_bytes here because: # - Python attempts str conversion first # - when self._charset != 'utf-8' it re-encodes the content if isinstance(value, (bytes, memoryview)): return bytes(value) if isinstance(value, str): return bytes(value.encode(self.charset)) # Handle non-string types. return str(value).encode(self.charset)

15 Useful Lines

# These methods partially implement the file-like object interface. # See https://docs.python.org/library/io.html#io.IOBase # The WSGI server must call this method upon completion of the request. # See # http://blog.dscpl.com.au/2012/10/obligations-for-calling-close-on.html def close(self): for closer in self._resource_closers: ... [manually omitted 424 lines for better present] ... if json_dumps_params is None: json_dumps_params = {} kwargs.setdefault("content_type", "application/json") data = json.dumps(data, cls=encoder, **json_dumps_params) super().__init__(content=data, **kwargs)

utils.py has been pruned from the memory. The agent cannot check the regex details without re-retrieving the file.

[User Input] Please fix the bug in the authentication module. [Turn 1~10] Explored repository; opened utils.py, config.py, and auth_service.py. [Turn 11~20] Localized error in auth_service.py; focused on modifying exception handling logic within auth_service.py. [Turn 21] Fix failed; test feedback indicates the root cause is an invalid regex pattern in utils.py. [Turn 22] Discovered utils.py context is missing; forced to re-open utils.py to inspect the code. [Turn 30~40] Corrected the regex pattern in utils.py. [Turn 41] Fix succeeded

424 Low-Information Lines

(b) Focus Shifting Problem

(a) Granularity Mismatch Fig. 2. Challenges in SE context compression. (a) Granularity Mismatch: A case from django-11133 showing that within a large file output, only a small function (make_bytes) is useful. (b) Focus Shifting Problem: A debugging timeline where the agent shifts focus from auth_service.py to utils.py. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:5

Information Density Variance. The information density within an agent’s trajectory is highly uneven. A significant portion of the context often consists of verbose logs or irrelevant code snippets. As illustrated in Figure 2a, we present a real-world example from the SWE-Bench instance django-11133. The agent retrieves the content of response.py to inspect the HttpResponse implementation. This action generates a massive output where over 700 lines are essentially noise (highlighted in red as “Low-Information Lines”). The agent effectively needs only the make_bytes method (approximately 15 lines, highlighted in green) to proceed. This variance creates a practical granularity challenge for compression. If compression operates at a coarse unit such as a full message or a full tool output, the method tends to either keep large chunks of noise or discard large chunks that may still contain small but critical details. If compression operates at an overly fine unit such as individual tokens, it risks breaking code and structured logs. This motivates a segmentation strategy that identifies semantically coherent units within long observations so that the agent can retain complete and meaningful parts while discarding irrelevant parts. Dynamic Relevance. The relevance of historical information in SE tasks fluctuates as the agent’s debugging focus and hypotheses evolve. As illustrated in Figure 2b, the agent initially inspects the whole repository (Turns 1-10), then shifts its focus to auth_service.py (Turns 11-20) to address the error, believing the root cause lies therein. During this phase, static compression methods (e.g., ObsMask, LLMSummary, and AgentDiet) deem the previously viewed utils.py as irrelevant “old” context and permanently prune it. However, in Turn 21, execution feedback reveals that the root cause is actually an incorrect regex pattern defined in utils.py. Since the specific details of utils.py were discarded, the agent cannot “look back” to verify the pattern, forcing a redundant re-opening of the file (Turn 22). This inefficiency highlights the need for a dynamic mechanism that can recall previously suppressed blocks when the agent’s focus shifts back to them. To quantify how frequently this phenomenon occurs in practice, we randomly sampled 100 trajectories from LLMSummary runs and audited them using an LLM-assisted review followed by human verification. We found that 40 of these trajectories exhibited the focus shifting problem, indicating that this challenge is not merely illustrative but occurs in a substantial portion of real agent debugging workflows. 2.3

Related Work

Prior work on context reduction for LLMs and agents can be grouped into three categories. Each category addresses part of the long-context problem but exhibits limitations for SE agent trajectories. 2.3.1 Selection-based Compression. Selection-based compression aims to improve LLM efficiency by identifying and retaining only the most informative parts of the input. A significant body of work relies on training dedicated modules to score content relevance. For instance, RECOMP [37] trains an abstractive compressor to paraphrase documents, while CPC [21] employs a trained context-aware sentence encoder to judge similarity. Similarly, methods like Provence [7] and LLMLingua-2 [29] utilize distillation techniques to train classifier models that predict the necessity of individual tokens. Alternatively, non-training approaches such as FilCo [35] and Selective Context [19] rely on information-theoretic metrics, calculating the self-information of tokens to prune those with low density. However, these approaches encounter two fundamental limitations in the context of SE agents. First, the heavy reliance on specific training restricts their plug-and-play capability and generalizability across diverse SE scenarios. Second, most selection methods operate at a token granularity. Token-level selectors disregard syntactic boundaries, where removing ostensibly low-information tokens, such as brackets, colons, or indentation, which can corrupt the Abstract Syntax Tree (AST) of code or the structure of JSON logs, rendering the context incomprehensible to Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:6

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

the LLM. These shortcomings underscore the necessity for a structure-aware compression strategy that respects the semantic coherence of software artifacts. 2.3.2 Heuristic-based Compression. Another category relies on simple heuristics, such as keeping only the most recent turns, applying FIFO-style eviction, or masking older observations. ObsMask [20] is a representative approach that directly drops tool outputs from older dialogue history. Moreover, Pan et al. [28] propose reformatting code to remove tokens related to whitespaces and indentations. These methods are attractive due to simplicity and low overhead, and they can reduce context length substantially in practice. However, heuristic strategies typically lack semantic awareness. They assume that old content is less useful, and they do not distinguish between a verbose but unimportant log and a short but critical clue. Meanwhile, as discussed in Section 2.2, SE tasks often require revisiting earlier evidence after a hypothesis shift. Heuristic deletion can therefore discard essential information and cause irreversible information loss, reducing agent success rates on complex tasks. 2.3.3 Summarization-based Compression. A third family of methods compresses context by utilizing LLMs to summarize or rewrite historical information. This strategy is widely adopted in practical agent systems, often in an ad-hoc manner to handle context saturation. For instance, tools like Gemini-Cli [11] and Claude Code [4] trigger LLM-based compression only when the context window reaches a predefined threshold. Others rely on rigid heuristics: Trae Agent [34] truncates tool responses to a fixed size (e.g., 16KB), while SWE-agent [41] employs configurable regex patterns to remove specific text blocks. In the research domain, AgentDiet [36] proposes a more systematic approach. It utilizes an additional LLM to refine or shorten messages, aiming to remove irrelevant parts while keeping key content. These approaches can preserve high-level semantics, but they introduce new issues in SE settings. First, they strongly depend on the summarization model’s ability to retain precise details that are often indispensable for SE tasks, including exact file paths, line numbers, error codes, and variable names. If these details are dropped or altered by hallucinating, the agent may lose grounding and propose incorrect patches. Second, using an extra LLM for summarization increases runtime overhead and latency, and it can complicate deployment when the agent is expected to run at scale or under tight response constraints. In summary, existing techniques either prune too aggressively without adapting to dependency shifts, or compress by rewriting content in ways that may lose critical debugging details. These limitations motivate methods that can both preserve structural integrity and dynamically select context based on the agent’s current needs. Although chunking, relevance scoring, and rolling context maintenance are not new in isolation, AttnCompress differs by unifying PPL-based block segmentation, proxy-attention selection, and dynamic re-evaluation in one training-free middleware built for multi-turn SE agent trajectories, explicitly targeting the granularity mismatch and focusshifting challenges discussed above. 3 3.1

Approach Overview

We formalize AttnCompress as a plug-and-play middleware positioned between the SE Agent’s trajectory manager and the backend LLM. As illustrated in Figure 3, the workflow consists of three consecutive phases: Structure-Aware Segmentation, Attention-based Scoring, and Dynamic Context Maintenance. Given a raw trajectory consisting of multiple interaction turns 𝑇 = {(𝑎 1, 𝑜 1 ), . . . , (𝑎𝑡 , 𝑜𝑡 )}, where 𝑎𝑡 represents the agent’s action and 𝑜𝑡 represents the observation from the environment (typically Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:7

AttnCompress Input Construction . . . ai

Phase 1: Structure-Aware Segmentation . . . Perplexity ai

ReAct Agent Trajectory Manager Task Input

oi

a1

Small Proxy LLM

o1 . . . ai

oi . . . at ot

ai+1 oi+i

. . . ai score:0.9

block2

score:0.2 oi

block3

score:0.7

block4

score:0.1

ai+1

content

oi+i

. . . <|im_start|>

Greedy Selection

block1

oi

content

Phase 2: Importance Estimation

Attnation Score

ai block1 o’i

block3 ai+1 ai+2

ai+1

block1

. . .

. . .

oi+i score:0.1

. . .

. . .

Phase 3: Dynamic Context Maintenance a1 o’1 Agent LLM

...

al o’l

Long-term Memory

...

at-K o’t-K

Short-term Memory

...

at ot

Immediate Memory

Fig. 3. Overview of the AttnCompress Framework. The system intercepts the raw trajectory, segments observations into blocks, scores them via a proxy model’s attention, and dynamically maintains a compressed context using a rolling window mechanism.

the output of tool executions, often containing verbose code or logs), our goal is to maintain a compressed trajectory 𝑇 ′ = {(𝑎 1, 𝑜 1′ ), . . . , (𝑎𝑡 , 𝑜𝑡′ )} in real-time. The objective is to ensure |𝑇 ′ | ≪ |𝑇 | while maximizing the retention of semantic information relevant to the current reasoning step 𝑠𝑡 , thereby accommodating a longer interaction history within a limited context window. 3.2

Phase 1: Structure-Aware Segmentation via PPL Spikes

To address the “granularity mismatch” problem in compression, where token-level pruning destroys syntax and message-level pruning retains excessive noise, we introduced an adaptive segmentation algorithm based on Perplexity (PPL) spike detection [31]. This algorithm utilizes a small model (Proxy Model) as a probe to perceive semantic boundaries within text. Prior work has empirically shown that perplexity-based boundary detection can preserve the syntactic integrity of long code contexts more effectively than fixed-size or heuristic chunking [31]. PPL Calculation. For any given tool output text 𝑂 (e.g., file content read by cat or error traces from pytest), we first split it into lines 𝐿 = {𝑙 1, . . . , 𝑙𝑛 }. We use the proxy model to calculate the log probability of each token 𝑥 and compute the average perplexity 𝑃𝑃𝐿(𝑙𝑖 ) for each line 𝑙𝑖 : ! 1 ∑︁ 𝑃𝑃𝐿(𝑙𝑖 ) = exp − log 𝑃 (𝑥 | 𝑥 <𝑐𝑜𝑛𝑡𝑒𝑥𝑡 ) (1) |𝑙𝑖 | 𝑥 ∈𝑙𝑖

Typically, at boundaries where semantic content shifts drastically (e.g., from the end of a function definition to the start of a new one, or from a normal log stream to an error stack), the model exhibits higher “surprise,” manifested as local peaks in the PPL curve. Boundary Detection. We determine split boundaries by detecting “spikes” in the 𝑃𝑃𝐿 sequence. To adapt to the fluctuating PPL baselines of different text contents, we employ an adaptive thresholding strategy: Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:8

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

(1) Diff Calculation: Compute the PPL difference between adjacent lines 𝐷𝑖 = |𝑃𝑃𝐿(𝑙𝑖 ) − 𝑃𝑃𝐿(𝑙𝑖 −1 )|. (2) Adaptive Threshold: Calculate the mean 𝜇 and standard deviation 𝜎 of the difference sequence. Define the spike threshold 𝜏 = 𝜇 + ℎ · 𝜎, where ℎ is a sensitivity coefficient. (3) Segmentation: Mark all local maxima points where 𝐷𝑖 > 𝜏 as split boundaries. Additionally, to avoid fragmentation, minute blocks with length only one line are merged with the preceding block. Finally, the raw output 𝑂 is transformed into a series of semantic blocks 𝐵 = {𝑏 1, 𝑏 2, . . . , 𝑏𝑚 }, where each block acts as a relatively independent unit in terms of syntax or semantics. Besides, introducing overlapping regions between adjacent blocks is a feasible engineering improvement that could further improve cross-boundary coherence, but it would introduce additional hyperparameters (e.g., overlap size). We therefore adopt only the simplest non-overlapping segmentation scheme in this paper. 3.3

Phase 2: Importance Estimation via Proxy Attention

To evaluate the dynamic relevance of historical blocks to the current task, we leverage the attention weights of the proxy model as a filtering signal, which prior work has shown can identify tokens semantically salient to a query [24]. Input Construction. To calculate relevance, we construct an input sequence that incorporates the current context. The input sequence is formed by concatenating three components: (1) Context ′ History (using the compressed trajectory 𝑇1:𝑡 −1 or raw trajectory 𝑇1:𝑡 −1 ), (2) The New Observation 𝑂𝑛𝑒𝑤 , and (3) a single special Query Token 𝑞𝑔𝑒𝑛 appended at the very end (e.g., the generation start token in the chat template, such as <|im_start|>, which corresponds to 𝑇 [𝑡].𝑠𝑡𝑎𝑟𝑡_𝑡𝑜𝑘𝑒𝑛 in Algorithm 1). In a causal LLM, this token must attend to the preceding context before generating the agent’s next response. This construction allows the model to aggregate the attention from the entire preceding context onto this final token, representing the agent’s “current reasoning state." Scoring Metric. We feed the constructed sequence into the proxy model to perform a forward pass. Crucially, this process is computationally efficient: the logits required for PPL-based segmentation (Phase 1) and the attention maps required for scoring (Phase 2) are extracted simultaneously in a single inference pass. We utilize the attention map from a specific layer (the last layer is used in our experiments). For each candidate block 𝑏𝑖 , its importance score 𝑆𝑐𝑜𝑟𝑒 (𝑏𝑖 ) is defined as the average attention weight projected from the single query token 𝑞𝑔𝑒𝑛 to the tokens within the block: 𝑆𝑐𝑜𝑟𝑒 (𝑏𝑖 ) =

1 ∑︁ 𝐴(𝑞𝑔𝑒𝑛 , 𝑡) |𝑏𝑖 |

(2)

𝑡 ∈𝑏𝑖

Here, 𝐴(𝑞𝑔𝑒𝑛 , 𝑡) denotes the attention weight from the query token 𝑞𝑔𝑒𝑛 to a token 𝑡 inside block 𝑏𝑖 . This metric effectively captures the semantic alignment between the agent’s current state and the historical information. Selection Strategy. Based on the calculated scores, we apply a Greedy Selection strategy. Let 𝐵 = {𝑏 1, . . . , 𝑏𝑚 } be all candidate blocks segmented from the tool output context, and let |𝑏𝑖 | denote the number Í of tokens in block 𝑏𝑖 . Given a compression ratio 𝜌 ∈ (0, 1), we set a token budget 𝐿budget = 𝜌 · 𝑚 𝑖=1 |𝑏𝑖 |. We then sort all blocks in descending order of 𝑆𝑐𝑜𝑟𝑒 (𝑏𝑖 ) and greedily add blocks to the retention set until the accumulated retained tokens reach 𝐿budget . Unselected blocks are discarded. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

3.4

ISSTA058:9

Phase 3: Dynamic Rolling Maintenance

In SE tasks, the agent’s focus shifts continuously as the programming process evolves, leading to task focus drift. Traditional static compression (compress once, delete forever) results in the inability to retrieve old information. To address this, we design a Three-Tier Rolling Window mechanism to maintain context dynamically. The Context Buffer. As shown in Figure 3, at interaction turn 𝑡, we partition the trajectory into three regions: 𝑇 = {(𝑎 1, 𝑜 1 ), . . . , (𝑎𝑙 , 𝑜𝑙 ), . . . , (𝑎𝑡 −𝑘 , 𝑜𝑡 −𝑘 ), . . . , (𝑎𝑡 , 𝑜𝑡 ) } | {z }| {z } | {z } Long-term

Short-term

(3)

Immediate

(1) Immediate Memory (Raw): The most recent 𝑘 turns (e.g., Tail 2) are kept in their raw state without compression. This ensures the agent maintains coherent perception of the immediate interaction, preventing the loss of actionable details (e.g., filenames, line numbers). (2) Short-term Memory (Rolling Buffer): The intermediate region covering turns 𝐼𝑙𝑜𝑛𝑔 + 1 through 𝑡 − 𝑘. This serves as a dynamic buffer. In every interaction turn, all content within this region is re-evaluated against the current 𝑄𝑐𝑢𝑟𝑟 , and attention scores are re-calculated to re-compress the content. This ensures the agent can extract the most relevant information from recent history based on its latest intent. (3) Long-term Memory (Fixed Archive): The region from turn 1 to 𝐼𝑙𝑜𝑛𝑔 . This serves as the archived history. To minimize computational overhead, this region remains static during standard interaction steps, which allows the LLM to leverage prefix caching [18]. Consequently, we avoid reconstructing this long-term memory at every step, significantly reducing the inference cost. This archive is only reconstructed when a “Global Refresh” is triggered. Periodic Global Refresh. To balance computational cost with recall capability, we introduce a sliding-window-based Global Refresh mechanism (Algorithm 1). This mechanism operates in two distinct modes: Mode 1: Incremental Update (Standard). This corresponds to the else block (Lines 19-25). In most interaction steps where the buffer is not full, the long-term boundary 𝐼𝑙𝑜𝑛𝑔 remains unchanged. We perform selection and re-compression only on the short-term region 𝑇𝑠ℎ𝑜𝑟𝑡 (Lines 19-22), ′ generating 𝑇𝑠ℎ𝑜𝑟𝑡 . Crucially, as shown in Line 25, the final context is constructed by concatenating ′ ′ the pre-existing compressed archive 𝑇𝑙𝑜𝑛𝑔 with the newly updated short-term blocks. Since 𝑇𝑙𝑜𝑛𝑔 is textually invariant, this strategy maximizes the cache hit rate for the Agent LLM’s prefix cache. Mode 2: Global Refresh (Triggered). When the short-term buffer accumulates beyond threshold 𝑀 (condition at Line 9), a Global Refresh is triggered (Lines 10-17). To address the limitation of a frozen archive, we merge the existing long-term archive with the current short-term buffer (𝑇𝑙𝑜𝑛𝑔 + 𝑇𝑠ℎ𝑜𝑟𝑡 in Line 11) and perform a global selection over this combined history. This allows the agent to “resurrect” previously discarded blocks from the deep history if they become relevant to the current task. Finally, we advance the archive pointer 𝐼𝑙𝑜𝑛𝑔 to the current raw boundary 𝐼𝑟𝑎𝑤 (Line 16), effectively committing the current history to the long-term archive. While this step invalidates the prefix cache, it is performed infrequently to minimize overhead while ensuring semantic completeness. This mechanism effectively combines high-frequency updates for the short term (adapting to rapid focus changes) with low-frequency refreshes for the long term (adapting to major shifts in debugging direction), solving the focus drift problem while maintaining efficiency. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:10

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

Algorithm 1 Dynamic Context Maintenance with Rolling Window Require: Full Trajectory 𝑇 , Compressed Trajectory 𝑇 ′ , Current Turn 𝑡, Tail Size 𝑘, Rolling Window Size 𝑀 1: 𝐼𝑙𝑜𝑛𝑔 ← Index of the last long memory turn ⊲ Initially 0 2: 𝐼𝑟𝑎𝑤 ← 𝑡 − 𝑘 3: // Phase 3.1: Define Context Regions 4: 𝑇𝑙𝑜𝑛𝑔 ← 𝑇 [0 : 𝐼𝑙𝑜𝑛𝑔 ] ′ 5: 𝑇𝑙𝑜𝑛𝑔 ← 𝑇 ′ [0 : 𝐼𝑙𝑜𝑛𝑔 ] 6: 𝑇𝑠ℎ𝑜𝑟𝑡 ← 𝑇 [𝐼𝑙𝑜𝑛𝑔 : 𝐼𝑟𝑎𝑤 ] 7: 𝑇𝑟𝑎𝑤 ← 𝑇 [𝐼𝑟𝑎𝑤 : 𝑡] ⊲ Keep Raw 8: // Phase 3.2: Check Trigger for Global Refresh 9: if Length(𝑇𝑠ℎ𝑜𝑟𝑡 ) ≥ 𝑀 then 10: // Global Refresh: Re-evaluate EVERYTHING before Raw 11: 𝐶𝑎𝑛𝑑𝑖𝑑𝑎𝑡𝑒𝑠 ← Segment(𝑇𝑙𝑜𝑛𝑔 + 𝑇𝑠ℎ𝑜𝑟𝑡 ) 12: 𝑆𝑐𝑜𝑟𝑒𝑠 ← ProxyAttention(𝑇𝑙𝑜𝑛𝑔 ,𝑇𝑠ℎ𝑜𝑟𝑡 ,𝑇𝑟𝑎𝑤 , 𝑄𝑢𝑒𝑟𝑦 = 𝑇 [𝑡].𝑠𝑡𝑎𝑟𝑡_𝑡𝑜𝑘𝑒𝑛) 13: 𝑆𝑒𝑙𝑒𝑐𝑡𝑒𝑑 ← SelectTop(𝐶𝑎𝑛𝑑𝑖𝑑𝑎𝑡𝑒𝑠, 𝑆𝑐𝑜𝑟𝑒𝑠) ′ 14: 𝑇𝑐𝑜𝑚𝑝𝑟𝑒𝑠𝑠𝑒𝑑 ← Reconstruct(𝑆𝑒𝑙𝑒𝑐𝑡𝑒𝑑) 15: // Update Archive Pointer 16: 𝐼𝑙𝑜𝑛𝑔 ← 𝐼𝑟𝑎𝑤 ′ 17: Output Context: 𝑇𝑐𝑜𝑚𝑝𝑟𝑒𝑠𝑠𝑒𝑑 + 𝑇𝑟𝑎𝑤 18: else 19: // Local Rolling: Re-evaluate only Short-term 20: 𝐶𝑎𝑛𝑑𝑖𝑑𝑎𝑡𝑒𝑠 ← Segment(𝑇𝑠ℎ𝑜𝑟𝑡 ) ′ ,𝑇 21: 𝑆𝑐𝑜𝑟𝑒𝑠 ← ProxyAttention(𝑇𝑙𝑜𝑛𝑔 𝑠ℎ𝑜𝑟𝑡 ,𝑇𝑟𝑎𝑤 , 𝑄𝑢𝑒𝑟𝑦 = 𝑇 [𝑡].𝑠𝑡𝑎𝑟𝑡_𝑡𝑜𝑘𝑒𝑛) 22: 𝑆𝑒𝑙𝑒𝑐𝑡𝑒𝑑 ← SelectTop(𝐶𝑎𝑛𝑑𝑖𝑑𝑎𝑡𝑒𝑠, 𝑆𝑐𝑜𝑟𝑒𝑠) ′ 23: 𝑇𝑠ℎ𝑜𝑟𝑡 ← Reconstruct(𝑆𝑒𝑙𝑒𝑐𝑡𝑒𝑑) 24: // Merge with existing Long-term ′ ′ + 𝑇𝑟𝑎𝑤 + 𝑇𝑠ℎ𝑜𝑟𝑡 25: Output Context: 𝑇𝑙𝑜𝑛𝑔 26: end if

3.5

Implementation

AttnCompress is a generalized trajectory compression framework designed to be compatible with various ReAct-style LLM agents. To evaluate its effectiveness in a SOTA setting, we integrated AttnCompress into Trae-Agent [34], a leading open-source agent that has demonstrated top-tier performance on benchmarks like SWE-Bench Verified [3, 8]. For the proxy model (𝐿𝐿𝑀𝑝𝑟𝑜𝑥 𝑦 ), we selected Qwen3-4B-Instruct [30, 40]. We chose this model for two strategic reasons: first, despite its small size, it retains strong capabilities in code syntax understanding and context processing; second, its significantly lower parameter count incurs minimal computational overhead, satisfying the real-time requirements of the agent workflow. All experiments, including the inference of the proxy model and the execution of the agent loop, were conducted on a server equipped with 8 NVIDIA A100 GPUs. AttnCompress involves several key hyperparameters. To determine the optimal configuration, we conducted a preliminary ablation study on a subset of the SWE-Bench-Verified dataset (randomly select 100 instances). Based on the trade-offs between token cost and pass rate, we established the following default settings for our main evaluation: • Tail Size (𝑘): 2 (The 2 most recent turns are always kept raw). • Compression Ratio (𝜌): 0.2 (Retaining top 20% of information by attention score). • PPL Block Threshold: -2 (Used for adaptive boundary detection). Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:11

• Attention Layer: -1 (Using the attention map from the last layer of the proxy model). • Rolling Window Size (𝑀): 10 (Triggers a Global Refresh when the short-term buffer accumulates 10 turns). We omit the detailed analysis of how these parameters affect performance in this section. A comprehensive sensitivity analysis and the impact of different hyperparameter combinations are presented in Section 5.2. 4

Experimental Design

4.1

Research Questions

To systematically assess the proposed framework, we structure our evaluation around three key research questions: RQ1: Cost-Effectiveness Trade-off. Does AttnCompress achieve a superior balance between problem-solving effectiveness and computational efficiency compared to state-of-theart baselines? We investigate whether AttnCompress can maintain or improve the pass rate on SWE-Bench-Verified while reducing token consumption, monetary cost, and end-to-end latency compared to existing methods. RQ2: Component Contribution. How do the internal architectural components and hyperparameters impact the performance of AttnCompress? We perform an ablation study to quantify the individual contributions of the ppl-based segmentation, proxy-attention scoring, and rolling window mechanism, while also analyzing the sensitivity of key hyperparameters (e.g., tail size 𝑘 and compression ratio 𝜌). RQ3: Generalization Capabilities. Does the framework generalize well across different proxy models and datasets? We assess the robustness of AttnCompress by evaluating its performance with different proxy model (e.g., Qwen vs. Llama) and testing its adaptability on the multilingual Multi-SWE-Bench dataset. 4.2

Experimental Setup

4.2.1 Datasets. We evaluate AttnCompress and the baselines on two complementary benchmarks to assess both effectiveness and generalization, consistent with the setup in [36]. • SWE-Bench-Verified [8]: This is the primary dataset for our evaluation. It consists of 500 human-verified software engineering tasks derived from real-world GitHub issues. To maintain consistency and eliminate sampling bias, we do not sample a new subset; instead, we utilize the exact instance lists curated by AgentDiet [36]. Specifically, we use their defined validation set (100 instances) for the parameter tuning and ablation studies described in RQ2. The remaining test set (200 instances), as identified in their work, is strictly reserved for the main comparative evaluation in RQ1 and the proxy-model generalization study in RQ3. • Multi-SWE-Bench-Flash [6, 45]: To evaluate the generalization capabilities of our approach across different languages and environments (RQ3), we utilize Multi-SWE-Bench-Flash. This benchmark contains 300 instances covering seven programming languages (Rust, TypeScript, JavaScript, Java, Go, C, and C++). These tasks typically present higher complexity, often requiring the agent to troubleshoot environment build errors, providing a robust testbed for the agent’s adaptability. 4.2.2 Baselines. We compare AttnCompress against seven baselines. All baselines are integrated into the same Trae-Agent [34] framework to ensure a fair comparison. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:12

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

• Original: The unmodified Trae-Agent that retains the full interaction history. This serves as the upper bound for context completeness and the lower bound for efficiency. • Random: A sanity check that randomly drops 75% of the tokens from previous turns. • SlidingWindow: A strict token-budget sliding window baseline that retains only the most recent trajectory content under a fixed token budget. To keep its token cost comparable to the other compression methods, we set the window size to 8k tokens. • ObsMask [20]: A rule-based approach that masks the outputs of tools from older turns with placeholder, assuming that recent observations are more relevant. • Lingua [29]: A token compression method that uses a small BERT-based [9] model to classify and remove non-essential tokens. • LLMSummary [17]: A standard summarization approach where an additional LLM is prompted to summarize the history of multiple turns into a concise paragraph. • AgentDiet [36]: A recent SOTA method that employs a reflection module (i.e., a summarize LLM) to rewrite and condense the trajectory after each step. For all baseline methods, we generally adopt the default hyperparameter settings reported in their original works. However, to ensure a strictly fair comparison, we standardize the tail size parameter (the number of most recent turns retained in raw format) to 𝑘 = 2 across all applicable methods. This alignment follows the experimental protocol of AgentDiet [36], ensuring that observed performance differences stem from the compression strategy rather than the amount of immediate raw history. Meanwhile, for the summarization-based methods (LLMSummary and AgentDiet), we utilize the identical LLM backbone as the agent LLM to perform the summarization and reflection tasks. This ensures that these baselines operate at their optimal capability and are not bottlenecked by a weaker summary model. 4.2.3 Metrics. We utilize a comprehensive set of metrics to evaluate the trade-off between performance and cost. • Pass%: The ratio of successfully resolved instances in the benchmark. This is the primary indicator of whether compression harmed the agent’s reasoning capabilities. • Step: The average number of interaction turns required to solve a task. An increase in steps typically indicates that the agent lost critical information due to compression and had to perform redundant actions to recover it. • PStep: The average number of interaction turns required for only successfully resolved instances. Unlike Step, this metric isolates the efficiency of successful trajectories, filtering out noise from failed attempts that simply exhaust the maximum step limit. • Input (I) & Output (O): The accumulated number of input and output tokens used by the backend LLM. We first sum input and output tokens over all interaction turns within each instance, and then report the average of these per-instance totals across instances. • Agent Cost (𝐶𝑎𝑔𝑒𝑛𝑡 ): The monetary cost incurred solely by the backend SE agent during interaction steps. This reflects the direct expenditure on reasoning and generation based on the (potentially compressed) context. • Compression Cost (𝐶𝑐𝑜𝑚𝑝 ): The cost incurred specifically by the compression mechanism itself (e.g., the summarizer cost in AgentDiet). Note that while the proxy model in AttnCompress can be deployed locally to avoid API charges, we calculate 𝐶𝑐𝑜𝑚𝑝 based on its official API pricing [30] to ensure a fair comparison. • Total Cost (𝐶𝑡𝑜𝑡𝑎𝑙 ): The sum of the agent cost and compression cost (𝐶𝑡𝑜𝑡𝑎𝑙 = 𝐶𝑎𝑔𝑒𝑛𝑡 +𝐶𝑐𝑜𝑚𝑝 ). Crucially, all cost calculations account for the input token discounts provided by prefix caching to reflect realistic API pricing. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

5

ISSTA058:13

Results

In this section, we present the experimental results to answer the three research questions proposed in Section 4.1. 5.1 RQ1: Cost-Effectiveness Trade-off To demonstrate the superiority of AttnCompress, we evaluated its performance on SWE-bench Verified using three different backend agents: Qwen3-Coder-30B [33], Qwen3-235B-Instruct [32], and Gemini-3-Flash [12]. Table 1 summarizes the comprehensive results. Table 1. Main Results on SWE-bench Verified. We report Pass Rate (Pass%), Input/Output token usage (k), Agent Cost ($), Compression Overhead ($), Total Cost ($), Average Steps, and Steps for Passed instances (PStep). All token, cost, and step values are per-instance averages. Bold indicates the best performance among compression methods; Underline indicates the second best. Method

Original

Random

Lingua

LLMSummary

ObsMask

SlidingWindow

AgentDiet

AttnCompress

Agent LLM

Pass (%)

Input (k)

Output (k)

𝑪𝒂𝒈𝒆𝒏𝒕

𝑪𝒄𝒐𝒎𝒑

𝑪 𝒕 𝒐𝒕𝒂𝒍

Step

PStep

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

72.50 46.50 46.50

1150.97 1278.10 925.41

6.13 6.31 10.74

0.2156 0.0864 0.0546

/ / /

0.2156 0.0864 0.0546

50.93 45.62 41.24

47.31 35.84 34.51

Mean

55.17

1118.16

7.73

0.1189

/

0.1189

45.93

39.22

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

67.00 39.00 38.50

1075.24 875.20 762.83

6.69 7.62 13.27

0.1896 0.0679 0.0616

/ / /

0.1896 0.0679 0.0616

66.97 60.81 56.50

61.58 43.55 44.65

Mean

48.17

904.42

9.19

0.1064

/

0.1064

61.42

49.93

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

69.50 36.50 40.00

906.21 799.10 745.77

6.43 7.10 13.17

0.1469 0.0688 0.0578

/ / /

0.1469 0.0688 0.0578

62.84 55.57 56.53

58.03 39.07 42.70

Mean

48.67

817.03

8.90

0.0912

/

0.0912

58.31

46.60

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

67.00 42.50 43.00

896.45 836.14 698.96

5.76 7.56 10.98

0.1731 0.0801 0.0509

0.0185 0.0089 0.0063

0.1916 0.0890 0.0572

61.75 51.63 44.24

53.86 34.29 35.23

Mean

50.83

810.52

8.10

0.1014

0.0112

0.1126

52.54

41.13

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

67.50 31.50 42.50

775.20 634.67 476.91

6.29 8.82 10.38

0.0699 0.0350 0.0282

/ / /

0.0699 0.0350 0.0282

71.39 67.05 51.99

66.27 39.32 43.24

Mean

47.17

628.93

8.49

0.0443

/

0.0443

63.47

49.61

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

67.00 42.00 40.50

641.79 602.30 520.08

5.75 6.53 10.21

0.1572 0.0766 0.0567

/ / /

0.1572 0.0766 0.0567

57.63 52.12 43.90

51.50 36.20 36.49

Mean

49.83

588.06

7.50

0.0968

/

0.0968

51.21

41.40

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

68.00 43.00 42.50

713.40 1143.56 610.37

6.31 6.26 9.88

0.1400 0.0838 0.0424

0.0932 0.0423 0.0269

0.2332 0.1260 0.0693

51.93 49.96 42.49

47.15 35.14 35.00

Mean

51.17

822.44

7.49

0.0887

0.0541

0.1429

48.12

39.10

Gemini-3-Flash Qwen3-235B Qwen3-Coder-30B

71.50 43.50 44.50

780.17 619.19 534.62

5.45 6.21 10.20

0.1543 0.0706 0.0532

0.0027 0.0020 0.0019

0.1570 0.0726 0.0551

60.46 50.53 46.30

55.48 36.00 38.56

Mean

53.17

644.66

7.29

0.0927

0.0022

0.0949

52.43

43.35

Pass Rate Analysis. The primary challenge in context compression is balancing information retention with token reduction. As observed in Table 1, applying any compression strategy naturally results in a slight performance dip compared to the Original full-context baseline (Mean Pass Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:14

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

55.17%). However, relying on full context is often impractical or impossible. Most current LLMs are constrained by context windows, which are easily exceeded by the massive observation logs generated during long-horizon SE tasks. Consequently, compression strategies are necessary not merely for cost savings, but to enable the agent to function within these hard limits. Among all compression techniques, AttnCompress achieves the best balance between performance and overhead. It attains the highest mean Pass Rate of 53.17%, outperforming heuristic methods like ObsMask (47.17%) and SlidingWindow (49.83%), as well as selection-based methods like Lingua (48.67%). Moreover, AttnCompress consistently outperforms all baselines individually across all three agent LLMs (achieving 71.50%, 43.50%, and 44.50% on Gemini-3-Flash, Qwen3-235B, and Qwen3-Coder-30B, respectively). Compared to the previous SOTA summarization method AgentDiet (51.17%), AttnCompress achieves a 3.9% relative improvement in mean pass rate, demonstrating better information retention. To further assess statistical robustness, we repeated the SWE-Bench-Verified (200 case subset) evaluation five times using Qwen3-Coder-30B as the backend agent. As shown in Table 2, AttnCompress achieves a higher mean pass rate than both LLMSummary (43.70% vs. 42.70%) and AgentDiet (43.70% vs. 41.80%), suggesting that the improvement is not driven by a single lucky run. Table 2. Repeated-run robustness on SWE-Bench-Verified with Qwen3-Coder-30B. We report pass-rate and total-cost statistics over five runs. Pass (%)

Method Original LLMSummary ObsMask SlidingWindow AgentDiet AttnCompress

𝑪 𝒕 𝒐𝒕𝒂𝒍 ($)

Mean

Std

Min

Max

Mean

Std

Min

Max

44.10 42.70 40.80 39.20 41.80 43.70

2.38 2.22 2.08 2.41 2.99 2.17

41.00 40.00 38.00 36.00 38.50 40.50

46.50 46.00 43.00 42.00 45.50 46.50

0.0563 0.0585 0.0305 0.0567 0.0718 0.0566

0.0016 0.0012 0.0013 0.0012 0.0028 0.0013

0.0546 0.0572 0.0282 0.0556 0.0688 0.0551

0.0582 0.0599 0.0313 0.0584 0.0751 0.0584

To understand the remaining gap from the full-context baseline, we manually inspected the failed cases. The main failure mode is that AttnCompress can still remove information that later becomes important, such as an earlier file path, helper function, or error message. When the agent’s subsequent reasoning depends on such omitted evidence, it may make an incorrect decision and fail to recover within the remaining steps. Moreover, AttnCompress demonstrates superior cost efficiency compared to both the full-context and SOTA approaches. By avoiding the expensive process of using a large LLM to summarize every turn (as in AgentDiet), AttnCompress lowers the total cost to 0.0949$, representing a 33.6% reduction compared to AgentDiet (0.1429$) and a 20.2% reduction compared to the Original baseline (0.1189$). This cost efficiency is robust across all models; for example, on the expensive Qwen3-235B model, AttnCompress reduces total cost from 0.0864$ (Original) to 0.0726$. Furthermore, regarding token consumption, AttnCompress operates with a significantly leaner context (Mean Input: 644.66k). This corresponds to a 21.6% reduction against AgentDiet (822.44k) and a substantial 42.3% reduction against Original (1118.16k), validating that our attention-based selection mechanism efficiently filters noise in raw contexts. Efficiency and Latency. Beyond monetary cost, latency is a critical factor for real-time agents. We analyzed the time overhead using the breakdown illustrated in Figure 4. AttnCompress incurs a manageable computational overhead with an average total end-to-end time of 366.7s, where the compression (proxy model inference) accounts for only 66.6s. This Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:15

is significantly faster than the SOTA method AgentDiet, which requires an average of 498.5s due to substantial summarization overhead (248.1s). While heuristic methods like ObsMask and SlidingWindow are indeed faster (total time 303.6s and 284.7s, respectively) due to zero processing cost, they suffer from a significantly lower pass rate (47.17% and 49.83%, respectively). AttnCompress thus strikes a critical balance, offering a reasonable trade-off between the speed of heuristic pruning and the effectiveness of heavy summarization. Interestingly, our method also outperforms lighter baselines like Random and Lingua in terms of total efficiency, despite those methods having negligible compression overhead. As shown in the Step column of Table 1, those two methods often disrupt semantic continuity, causing the agent to get confused. This forces the agent to perform redundant actions to recover missing information, driving up the average step count (Random: 61.42 steps; Lingua: 58.31 steps). In contrast, AttnCompress preserves the semantic integrity of the trajectory, allowing the agent to solve problems in fewer steps (52.43 steps), thereby reducing the total cumulative inference time of the backend agent. Total Time (s) Compess Time (s)

498.5

Agent Execution Time (s)

366.7 304.2

284.7

284.7

308.6

344.3

369.5

52.7

35.3

Summary

Lingua

250.4

370.8 248.1

66.6 Original SlideWindow

404.8

308.6 300.1

304.2

397.0 370.8

ObsMask

AttnSelect

Random

AgentDiet

Fig. 4. End-to-end Latency Breakdown.

Conclusion 1 AttnCompress achieves the highest pass rate among studied compression methods (53.17%), improving upon prior SOTA methods by 3.9% while reducing total costs by over 33.6%. Furthermore, it lowers end-to-end latency by 26.4% compared to AgentDiet, offering more practical solution for long-horizon SE tasks. 5.2

RQ2: Component Contribution

To quantify the contribution of individual architectural components to the overall performance, we conducted an ablation study on a subset of SWE-Bench-Verified as detailed in Section 4.2.1 with Qwen3-Coder-30B as agent LLM. We subsequently analyzed the sensitivity of the framework to key hyperparameters to determine the optimal configuration. Analysis of PPL-based Segmentation. We first evaluate the necessity of our structure-aware segmentation by replacing PPL-based blocking with standard token-level pruning (while maintaining the same compression ratio). As shown in Table 3, removing the PPL module leads to a 3.0% decrease in Pass Rate (from 42.0% to 39.0%) and an increase in the average steps required. This degradation occurs because token-level pruning is agnostic to syntactic boundaries. In SE tasks, observations often consist of structured data such as JSON objects, stack traces, or code snippets. Randomly dropping tokens within these structures breaks the syntax, rendering the remaining context incomprehensible for the LLM. PPL-based segmentation ensures that we drop entire semantically coherent blocks (e.g., a redundant log line) rather than fragmenting critical code structures. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:16

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

Table 3. Ablation Study of Key Components. We compare the full AttnCompress framework against variants where specific mechanisms are removed or replaced. Method

Pass (%)

Input (k)

Output (k)

𝑪 𝒕 𝒐𝒕𝒂𝒍 ($)

Step

AttnCompress (Full)

42.0

547.27

10.74

0.0557

45.89

PStep 38.17

w/o PPL (Token-level) w/o Attention (Random) w/o Rolling (Fixed)

39.0 35.0 37.0

622.37 566.81 514.36

10.88 10.06 10.95

0.0635 0.0567 0.0619

50.48 47.38 48.44

40.13 35.89 39.51

Analysis of Proxy Attention. To assess the impact of our attention score module, we replaced the proxy model’s attention scoring with a random block selection strategy. This resulted in the most significant performance drop, with the pass rate plummeting to 35.0% (-7.0%). This result underscores the high noise ratio in SE trajectories. The majority of tool outputs are irrelevant to the specific bug at hand. The attention mechanism acts as a critical filter, aligning the historical context with the agent’s current reasoning. Without this guidance, the agent fails to locate the bug efficiently, forcing it to waste steps on unrelated files and often leading to task failure. Analysis of Dynamic Rolling Window. We investigated the importance of dynamic context maintenance by disabling the rolling window mechanism (using a fixed strategy where context is compressed once and never revisited). The pass rate dropped significantly to 37.0% (-5.0%). This validates the task focus drift hypothesis in debugging workflows. Information that appears irrelevant at step 𝑡 often becomes critical at step 𝑡 + 20 when the agent shifts its hypothesis. A static compression strategy permanently discards this information, preventing the agent from “looking back”. The rolling window mechanism is therefore essential for allowing the agent to recall previously suppressed blocks as its focus evolves. Conclusion 2 The ablation study confirms that all three components are essential. The attention mechanism is the primary driver of effectiveness (+7.0% pass rate) by filtering noise. The rolling window is crucial for handling non-linear debugging (+5.0% pass rate), and PPL segmentation ensures syntactic integrity (+3.0% pass rate) for code parsing. Hyperparameter Sensitivity. We further examined the impact of five key hyperparameters on the validation set. The results are detailed in Table 4. Tail Size (𝑘): There is a trade-off between performance and cost. Increasing the tail size from 2 to 10 improves the pass rate to 44.0%, as keeping more raw context helps the agent understand the immediate consequences of its actions. However, this comes at a significantly higher token cost (0.0821$ vs 0.0557$). To balance efficiency with effectiveness, and to maintain consistency with previous work like AgentDiet, we selected 𝑘 = 2 as the default. Compression Ratio (𝜌): As expected, a higher retention budget leads to better performance. Increasing 𝜌 to 0.3 yields a slight gain (43.0%). However, reducing it to 0.1 causes a drop (40.0%), indicating that essential information is being discarded. We selected 𝜌 = 0.2 as it offers a favorable cost-performance ratio, capturing the majority of relevant signals without inflating the context. Block Threshold: The results favor smaller thresholds, which correspond to finer-grained segmentation. A threshold of -2 (adaptive) or 0 allows the model to select precise lines of interest, whereas a coarser threshold of 2 (grouping larger chunks) degrades performance to 39.0%. Since the threshold has a negligible impact on computational overhead, we selected -2 to maximize segmentation flexibility. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:17

Table 4. Hyperparameter Sensitivity Analysis on the Validation Set. Default settings are marked with *. Parameter

Value

Pass (%)

Input (k)

Output (k)

𝑪 𝒕 𝒐𝒕𝒂𝒍 ($)

Step

PStep

Tail Size (𝑘)

2 (*) 5 10

42.0 42.0 44.0

547.27 563.19 603.26

10.74 10.69 10.74

0.0557 0.0657 0.0821

45.89 44.47 43.64

38.17 40.31 38.77

Comp. Ratio (𝜌)

0.1 0.2 (*) 0.3

40.0 42.0 43.0

465.49 547.27 600.65

10.17 10.74 10.72

0.0469 0.0557 0.0613

46.03 45.89 45.39

41.45 38.17 41.35

Block Threshold

-2 (*) 0 2

42.0 42.0 39.0

547.27 520.38 516.52

10.74 10.19 10.41

0.0557 0.0519 0.0486

45.89 45.23 45.39

38.17 35.69 37.08

Proxy Layer

Mean 0 (First) Middle -1 (Last) (*)

41.0 41.0 43.0 42.0

556.99 536.46 548.67 547.27

10.50 10.15 10.32 10.74

0.0541 0.0532 0.0539 0.0557

46.58 46.00 46.10 45.89

40.90 35.88 38.35 38.17

Window Size

2 5 10 (*)

35.0 38.0 42.0

541.79 488.52 547.27

10.79 10.00 10.74

0.0761 0.0529 0.0557

45.17 43.11 45.89

37.17 35.32 38.17

Proxy Layer selection: The choice of proxy layer has only a marginal impact on the final performance, as the results across different layers (First, Middle, and Last) are largely comparable. This suggests that the relevance signal derived from proxy attention is robust. Given this insensitivity, we selected the Last layer (-1) as it is computationally free to extract and performs robustly. Window Size: A larger rolling window is beneficial. Increasing the window size from 2 to 10 improves the pass rate from 35.0% to 42.0%. Meanwhile, increasing the Window Size has a minimal impact on token cost because the content within the window is already compressed. Therefore, we selected a larger window of 𝑀 = 10 to maximize the agent’s ability to recall historical context. In summary, while AttnCompress introduces several hyperparameters, Table 4 demonstrates that they primarily govern cost-performance trade-offs rather than acting as fragile triggers. Key variables like compression ratio (𝜌) and tail size (𝑘) smoothly scale performance up or down, whereas architectural choices (threshold and proxy layer) show robust plateaus (e.g., both threshold -2 and 0 achieve 42%, and middle/last layers achieve 43% and 42%). This indicates that the framework is resilient to parameter shifts and does not require exhaustive, task-specific fine-tuning. Conclusion 3 AttnCompress exhibits a clear trade-off between cost and performance. Notably, we adopted a conservative “balanced” configuration for our main evaluation rather than greedily maximizing the pass rate. This default setting already outperforms prior SOTA methods, suggesting that AttnCompress is highly effective even under cost constraints, while offering significant headroom for performance improvement if larger token budgets are permitted. 5.3

RQ3: Generalization Capabilities

Finally, we assess the generalization capabilities of AttnCompress. We investigate whether the framework maintains its effectiveness when utilizing different proxy model architectures and when applied to diverse programming languages beyond Python. Proxy Model Robustness. To verify that our approach is not dependent on a specific model family, we evaluated AttnCompress using five distinct small language models (SLMs) as the proxy scorer on the 200-instance SWE-Bench-Verified test set. This includes three models from the Qwen3 [30] Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:18

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

Table 5. Proxy–agent attention alignment (Qwen3-Coder-30B as reference scorer). Proxy Model

Spearman 𝜌

Qwen3-1.7B-Instruct Qwen3-4B-Instruct Qwen3-8B-Instruct Gemma3-4B-Instruct Llama3.2-3B-Instruct

0.64 0.63 0.61 0.60 0.56

Top-10% Ovlp. Top-20% Ovlp. 0.63 0.60 0.55 0.64 0.63

0.66 0.65 0.62 0.62 0.63

Table 6. Performance and Efficiency with different Proxy Models. Agent LLM

Proxy Model

Pass (%)

Input (k)

𝑪 𝒕 𝒐𝒕𝒂𝒍 ($)

Step

Ana Time (s)

Qwen3-Coder-30B

Qwen3-1.7B-Instruct Qwen3-4B-Instruct Qwen3-8B-Instruct Gemma3-4B-Instruct Llama3.2-3B-Instruct

43.50 44.50 46.50 45.00 45.50

551.63 534.62 558.34 694.98 507.71

0.0571 0.0551 0.0607 0.0588 0.0548

47.16 46.30 46.67 55.02 44.98

71.61 530.87 384.13 515.63 100.37

Gemini-3-Flash

Qwen3-1.7B-Instruct Qwen3-4B-Instruct Qwen3-8B-Instruct Gemma3-4B-Instruct Llama3.2-3B-Instruct

69.00 71.50 72.50 68.50 71.50

749.32 780.17 773.15 781.90 781.39

0.1940 0.1570 0.2077 0.1705 0.2072

64.90 60.46 65.78 71.48 68.41

131.49 252.47 682.06 790.62 239.08

family to analyze scaling laws (1.7B, 4B, 8B), as well as Gemma3-4B-Instruct [10] and Llama3.2-3BInstruct [26] to test cross-architecture generalization. We first verify whether proxy relevance scoring aligns with the backend agent’s needs. We use 200 uncompressed Original agent trajectories from the same 200-instance SWE-Bench-Verified subset evaluated below, and compare block-level attention rankings from each proxy model against Qwen3-Coder-30B on identical PPL-segmented blocks. Table 5 reports spearman rank correlation (𝜌), the fraction of shared blocks in the top 10% of each ranking, and the same for the top 20% (instance means). Our default Qwen3-4B-Instruct proxy achieves 𝜌=0.63 with 0.60 and 0.65 top10%/20% overlap, indicating substantial agreement on the blocks most likely to be retained under our compression budget. Cross-family proxies remain in a similar range, suggesting that imperfect global rankings still preserve the salient context the agent attends to. The end-to-end results are presented in Table 6, where we report both Qwen3-Coder-30B and Gemini-3-Flash as backend agents. Together with Table 5, this shows that small–big LLM attention agreement is sufficient for stable agent behavior: pass rates vary by only a few points across proxies. The data indicates a high degree of transferability across model families. On Qwen3-Coder30B, both Gemma3-4B-Instruct (45.0%) and Llama3.2-3B-Instruct (45.5%) achieve competitive pass rates comparable to Qwen3-4B-Instruct (44.5%). The same trend holds on Gemini-3-Flash, where Gemma3-4B-Instruct and Llama3.2-3B-Instruct reach 68.5% and 71.5%, respectively, close to the default Qwen3-4B-Instruct (71.5%). This is consistent with their strong top-10%/20% overlap with the backend agent in Table 5 (e.g., 0.63, 0.64 for Llama3.2-3B-Instruct), suggesting that the attention patterns used to distinguish signal from noise are universal features present in various LLM families, making AttnCompress model-agnostic. We further analyze the impact of model size within the Qwen3 series using Qwen3-Coder-30B as the backend agent. There is a nuanced trade-off between proxy model size, selection quality, and compression latency (Ana Time column). Specifically, Qwen3-1.7B-Instruct is the fastest (71.6s) but suffers from a slight performance drop (43.5%), suggesting that very small models may struggle to Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:19

accurately identify subtle long-term dependencies beyond coarse attention agreement. Qwen3-8BInstruct achieves the highest pass rate of 46.5%. However, this gain comes at a higher analysis time (384.1s). Therefore, Qwen3-4B-Instruct strikes the optimal balance. It improves pass rate over the 1.7B model (+1.0%) while incurring substantially lower analysis time than the 8B proxy. We also observe that the ranking of proxy models based on end-to-end agent performance (Table 6) does not strictly align with their attention alignment scores (Table 5). For instance, Qwen3-1.7B-Instruct achieves the highest top-20% overlap (0.66) but records the lowest pass rate on Qwen3-Coder-30B. We attribute this discrepancy primarily to two factors: (i) the differences in alignment across proxies are marginal (all top-20% overlaps fall within [0.62, 0.66]), meaning that minor variations in ranking do not necessarily translate to downstream success; and (ii) end-to-end evaluations inherently introduce variance due to agent stochasticity. While the exact mechanism by which attention alignment influences final agent outcomes warrants further investigation, our results yield two practical takeaways: proxy models from diverse families and scales share broadly similar attention distributions, and this shared structure enables effective compression with negligible variations in pass rates across proxies. Conclusion 4 The effectiveness of attention-based filtering is not tied to a single architecture; both Gemma and Llama models provide sufficient semantic signals to drive high agent performance. Multi-Language Generalization. We extended our evaluation to Multi-SWE-Bench-Flash to test adaptability across seven programming languages: C, C++, Go, Java, JavaScript, Rust, and TypeScript. Passed Instances (#) by Method Table 7 summarizes the overall results, and Figure 5 illustrates the specific pass countsLanguage per language. and Programming Java 17 TypeScript

Table 7. Main Results on Multi-SWE-Bench-Flash. Bold indicates the best performance among compression methods; Underline indicates the second best. Method

Pass (%)

Input (k)

Output (k)

𝑪 𝒕 𝒐𝒕𝒂𝒍 ($)

Step

Original

20.33

1629.38

11.69

0.1057

52.83

42.66

Random Lingua LLMSummary ObsMask SlidingWindow AgentDiet

17.00 16.00 17.39 15.72 17.67 18.33

1420.96 1360.88 1107.88 719.77 790.56 1014.70

14.54 16.50 11.54 12.35 12.02 11.56

0.1074 0.1046 0.0971 0.0376 0.0764 0.1132

79.51 78.24 59.07 72.05 61.76 58.45

62.24 58.69 45.62 55.85 52.36 45.04

AttnCompress

19.67

879.77

11.90

0.0826

62.38

49.32

C

12 8 4 0

Rust

C++

PStep JavaScript Original AttnCompress

Go AgentDiet LLM Summary

Fig. 5. Number of passed instances per language on Multi-SWE-Bench-Flash. AttnCompress matches the Original baseline in Java, Rust, and C++, outperforming other compression methods.

AttnCompress achieves the best performance among all compression methods, with a pass rate of 19.67%, which is close to the Original full-context baseline (20.33%). In contrast, other methods show significant degradation: ObsMask drops to 15.72%, SlidingWindow achieves 17.67%, and AgentDiet achieves 18.33% while incurring higher costs (0.1132$ vs 0.0826$). Language-specific analysis (Figure 5) reveals that AttnCompress maintains parity with the Original baseline in strictly typed and verbose languages. • C/C++ & Java: In Java, AttnCompress matches the Original baseline exactly (17 passed) and outperforms AgentDiet (13 passed). Similarly, in C, it successfully solves 5 instances Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:20

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

compared to 4 for the Original baseline. This indicates that our method remains effective in verbose, statically-typed environments. • Rust & Go: Performance is stable in modern systems languages. In Rust, AttnCompress matches the Original baseline (9 passed), and in Go, it shows a marginal difference (8 vs. 9). • Web Languages (TS/JS): For TypeScript and JavaScript, AttnCompress (7 passed) performs slightly below the Original (8 passed) but remains competitive with AgentDiet. Conclusion 5 AttnCompress generalizes across diverse programming languages, achieving near-parity with full-context agents (96.7% relative performance) while reducing token costs by 20%. 6

Threats to Validity

Threats to Internal Validity. The primary internal threat is Data Leakage, as proprietary LLMs might have seen the SWE-Bench issues during training. We mitigated this by including the more recent Multi-SWE-Bench in our evaluation. Furthermore, since all baselines utilize the same backend models, any leakage affects them equally, preserving the validity of relative comparisons. A second threat is hyperparameter overfitting. To address this, we strictly isolated a 100-instance validation set for ablation studies and parameter tuning, ensuring the reported performance on the 200-instance test set reflects genuine generalization rather than overfitting. Threats to External Validity. The main external threat is generalization across agent frameworks. Due to computational costs, we evaluated AttnCompress primarily on Trae-Agent. However, as most SE agents follow similar ReAct patterns, our middleware approach is theoretically transferable. We also addressed model and language generalization by validating our framework across diverse proxy models (Qwen, Llama, Gemma) and seven programming languages. The consistent results suggest our findings are not limited to a specific model architecture or language ecosystem. Threats to Construct Validity. A threat to construct validity is the reliance on test-based evaluation. A patch passing available tests (plausible) may not be semantically equivalent to the developer’s fix (correct). While this is a known limitation of benchmarks like SWE-Bench, it serves as a standard proxy for task success. Crucially, this metric limits all comparison methods equally; therefore, the observed improvements in pass rate reliably indicate that AttnCompress retains more critical task information than other compression baselines. 7

Conclusions

In this paper, we addressed the critical context scalability bottleneck in Autonomous Software Engineering (ASE) agents. We introduced AttnCompress, a dynamic compression framework that overcomes the limitations of static pruning and heuristic summarization through three key mechanisms: structure-aware segmentation via PPL spikes, proxy attention-guided relevance estimation, and a dynamic rolling window. This approach ensures the preservation of syntactic integrity and semantic dependencies essential for SE tasks. Extensive evaluation on SWE-Bench-Verified demonstrates that AttnCompress achieves a state-of-the-art pass rate of 53.17%, outperforming strong compression baselines while reducing token consumption by over 21.6% and total costs by 33.6%. Our results confirm that dynamic attention alignment offers a superior, model-agnostic solution for efficient, long-horizon software engineering tasks. Data Availability The replication package for our study, containing the necessary source code and scripts to reproduce our experiments, is available at the repository [1].

Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:21

References [1] [n. d.]. AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents. https: //github.com/ZZR0/AttnCompress. Accessed: 2026-07-16.. [2] 2025. SWE-bench Multilingual · Kabir Khandpur — kabirk.com. https://kabirk.com/multilingual. [Accessed 26-01-2026]. [3] 2026. SWE-bench Leaderboards — swebench.com. https://www.swebench.com/index.html. [Accessed 19-01-2026]. [4] Anthropic. 2025. Claude Code. https://claude.com/product/claude-code. Agentic AI coding tool for terminal and development workflows, accessed 2026-01-28. [5] Islem Bouzenia and Michael Pradel. 2025. Understanding Software Engineering Agents: A Study of Thought-ActionResult Trajectories. CoRR abs/2506.18824 (2025). arXiv:2506.18824 doi:10.48550/ARXIV.2506.18824 [6] ByteDance Seed Team. 2025. Multi-SWE-bench-flash. https://huggingface.co/datasets/ByteDance-Seed/Multi-SWEbench-flash. Subset of Multi-SWE-bench for rapid evaluation; accessed 2026-01-20. [7] Nadezhda Chirkova, Thibault Formal, Vassilina Nikoulina, and Stéphane Clinchant. 2025. Provence: efficient and robust context pruning for retrieval-augmented generation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https://openreview.net/forum?id=TDy5Ih78b4 [8] Neil Chowdhury, James Aung, Chan Jun Shern, Oliver Jaffe, Dane Sherburn, Giulio Starace, Evan Mays, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E. Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry. 2024. Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/ [9] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186. doi:10.18653/v1/N19-1423 [10] Google. 2024. Gemma-3-4B-IT. https://huggingface.co/google/gemma-3-4b-it. Model card, accessed 2026-01-28. [11] Google. 2025. Gemini CLI. https://geminicli.com/. Open-source AI command-line interface for Gemini models, accessed 2026-01-28. [12] Google DeepMind. 2025. Gemini 3 Flash. https://deepmind.google/models/gemini/flash/. Model card and official description, accessed 2026-01-28; proprietary large language model by Google DeepMind. [13] Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg. 2024. RULER: What’s the Real Context Size of Your Long-Context Language Models?. In First Conference on Language Modeling. https://openreview.net/forum?id=kIoBbc76Sy [14] Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 13358–13376. doi:10.18653/V1/2023.EMNLP-MAIN.825 [15] Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https://openreview.net/forum? id=VTF8yNQM66 [16] Haolin Jin, Linghan Huang, Haipeng Cai, Jun Yan, Bo Li, and Huaming Chen. 2024. From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future. CoRR abs/2408.02479 (2024). arXiv:2408.02479 doi:10.48550/ARXIV.2408.02479 [17] Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, and Saravan Rajmohan. 2025. Acon: Optimizing context compression for long-horizon llm agents. arXiv preprint arXiv:2510.00615 (2025). [18] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, Jason Flinn, Margo I. Seltzer, Peter Druschel, Antoine Kaufmann, and Jonathan Mace (Eds.). ACM, 611–626. doi:10.1145/3600006.3613165 [19] Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin. 2023. Compressing Context to Enhance Inference Efficiency of Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 6342–6353. doi:10.18653/V1/2023.EMNLP-MAIN.391 [20] Tobias Lindenbauer, Igor Slinko, Ludwig Felder, Egor Bogomolov, and Yaroslav Zharov. 2025. The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management. In NeurIPS 2025 Fourth Workshop on Deep Learning for Code. https://openreview.net/forum?id=OHVzruJl5k

Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

ISSTA058:22

Zhengran Zeng, Yixin Li, Rui Xie, Wei Ye, and Shikun Zhang

[21] Barys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov, Ali Etemad, and Shane K. Luke. 2025. Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference. In AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA, USA, Toby Walsh, Julie Shah, and Zico Kolter (Eds.). AAAI Press, 24595–24604. doi:10.1609/AAAI.V39I23.34639 [22] Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. 2024. Large Language Model-Based Agents for Software Engineering: A Survey. CoRR abs/2409.02977 (2024). arXiv:2409.02977 doi:10.48550/ARXIV.2409.02977 [23] Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguistics 12 (2024), 157–173. doi:10.1162/TACL_A_00638 [24] Lvzhou Luo, Yixuan Cao, and Ping Luo. 2025. AttnComp: Attention-Guided Adaptive Context Compression for Retrieval-Augmented Generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, 8456–8472. https://aclanthology.org/2025.findings-emnlp.449/ [25] Kishan Maharaj, Vitobha Munigala, Srikanth G. Tamilselvam, Prince Kumar, Sayandeep Sen, Palani Kodeswaran, Abhijit Mishra, and Pushpak Bhattacharyya. 2025. ETF: An Entity Tracing Framework for Hallucination Detection in Code Summaries. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 30639–30652. doi:10.18653/v1/2025.acl-long.1480 [26] Meta AI. 2024. Llama-3.2-3B-Instruct. https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct. Model card, accessed 2026-01-28. [27] Tianyue Ou, Wanyao Guo, Apurva Gandhi, Graham Neubig, and Xiang Yue. 2025. AgentDiagnose: An Open Toolkit for Diagnosing LLM Agent Trajectories. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Ivan Habernal, Peter Schulam, and Jörg Tiedemann (Eds.). Association for Computational Linguistics, Suzhou, China, 207–215. doi:10.18653/v1/2025.emnlp-demos.15 [28] Dangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo, and Xiaoning Du. 2025. The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget. CoRR abs/2508.13666 (2025). arXiv:2508.13666 doi:10.48550/ ARXIV.2508.13666 [29] Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Rühle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Dongmei Zhang. 2024. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression. In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, 963–981. doi:10.18653/V1/2024.FINDINGS-ACL.57 [30] Qwen Team. 2025. Qwen3-4B-Instruct-2507. https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507. Model card, accessed 2026-01-20. [31] Yuling Shi, Yichun Qian, Hongyu Zhang, Beijun Shen, and Xiaodong Gu. 2025. LongCodeZip: Compress Long Context for Code Language Models. CoRR abs/2510.00446 (2025). arXiv:2510.00446 doi:10.48550/ARXIV.2510.00446 [32] Qwen Team. 2025. Qwen3-235B-A22B-Instruct-2507. https://huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507. A 235B instruction-tuned causal Mixture-of-Experts model. Model card on Hugging Face.. [33] Qwen Team. 2025. Qwen3-Coder-30B-A3B-Instruct. https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct. A 30B instruction-tuned causal Mixture-of-Experts model. Model card on Hugging Face.. [34] Trae Research Team, Pengfei Gao, Zhao Tian, Xiangxin Meng, Xinchen Wang, Ruida Hu, Yuanan Xiao, Yizhou Liu, Zhao Zhang, Junjie Chen, Cuiyun Gao, Yun Lin, Yingfei Xiong, Chao Peng, and Xia Liu. 2025. Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling. (2025). arXiv:2507.23370 [cs.SE] https: //arxiv.org/abs/2507.23370 [35] Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md. Rizwan Parvez, and Graham Neubig. 2023. Learning to Filter Context for Retrieval-Augmented Generation. CoRR abs/2311.08377 (2023). arXiv:2311.08377 doi:10.48550/ARXIV.2311.08377 [36] Yuan-An Xiao, Pengfei Gao, Chao Peng, and Yingfei Xiong. 2025. Improving the Efficiency of LLM Agent Systems through Trajectory Reduction. CoRR abs/2509.23586 (2025). arXiv:2509.23586 doi:10.48550/ARXIV.2509.23586 [37] Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective Augmentation. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. https://openreview.net/forum?id=mlJLVigNHp [38] Jingxuan Xu, Ken Deng, Weihao Li, Songwei Yu, Huaixi Tang, Haoyang Huang, Zhiyi Lai, Zizheng Zhan, Yanan Wu, Chenchen Zhang, Kepeng Lei, Yifan Yao, Xinping Lei, Wenqiang Zhu, Zong-Xian Feng, Han Li, Junqi Xiong, Dailin Li, Zuchen Gao, Kun Wu, Wen Xiang, Ziqi Zhan, Yuanxing Zhang, Wuxuan Gong, Ziyuan Gao, Guanxiang Wang, Yirong Xue, Mengtong Li, Mengfei Xie, Xiaojiang Zhang, Jinghui Wang, Wenhao Zhuang, Zheng Lin, Huiming Wang, Zhaoxiang Zhang, Yuqun Zhang, Haotian Zhang, Bin Chen, and Jiaheng Liu. 2025. SWE-Compass: Towards Unified

Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents

ISSTA058:23

Evaluation of Agentic Coding Abilities for Large Language Models. CoRR abs/2511.05459 (2025). arXiv:2511.05459 doi:10.48550/ARXIV.2511.05459 [39] Qi Xuan, Aaron Okano, Premkumar T. Devanbu, and Vladimir Filkov. 2014. Focus-shifting patterns of OSS developers and their congruence with call graphs. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering, (FSE-22), Hong Kong, China, November 16 - 22, 2014, Shing-Chi Cheung, Alessandro Orso, and Margaret-Anne D. Storey (Eds.). ACM, 401–412. doi:10.1145/2635868.2635914 [40] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jian Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. 2025. Qwen3 Technical Report. CoRR abs/2505.09388 (2025). arXiv:2505.09388 doi:10.48550/ARXIV.2505.09388 [41] John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.). http://papers.nips.cc/paper_files/paper/2024/hash/ 5a7c947568c1b1328ccc5230172e1e7c-Abstract-Conference.html [42] John Yang, Carlos E Jimenez, Alex L Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik R Narasimhan, Diyi Yang, Sida Wang, and Ofir Press. 2025. SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?. In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=riTiq3i21b [43] Jian Yang, Xianglong Liu, Weifeng Lv, Ken Deng, Shawn Guo, Lin Jing, Yizhi Li, Shark Liu, Xianzhen Luo, Yuyu Luo, Changzai Pan, Ensheng Shi, Yingshui Tan, Renshuai Tao, Jiajun Wu, Xianjie Wu, Zhenhe Wu, Daoguang Zan, Chenchen Zhang, Wei Zhang, He Zhu, Terry Yue Zhuo, Kerui Cao, Xianfu Cheng, Jun Dong, Shengjie Fang, Zhiwei Fei, Xiangyuan Guan, Qipeng Guo, Zhiguang Han, Joseph James, Tianqi Luo, Renyuan Li, Yuhang Li, Yiming Liang, Congnan Liu, Jiaheng Liu, Qian Liu, Ruitong Liu, Tyler Loakman, Xiangxin Meng, Chuang Peng, Tianhao Peng, Jiajun Shi, Mingjie Tang, Boyang Wang, Haowen Wang, Yunli Wang, Fanglin Xu, Zihan Xu, Fei Yuan, Ge Zhang, Jiayi Zhang, Xinhao Zhang, Wangchunshu Zhou, Hualei Zhu, King Zhu, Bryan Dai, Aishan Liu, Zhoujun Li, Chenghua Lin, Tianyu Liu, Chao Peng, Kai Shen, Libo Qin, Shuangyong Song, Zizheng Zhan, Jiajun Zhang, Jie Zhang, Zhaoxiang Zhang, and Bo Zheng. 2025. From Code Foundation Models to Agents and Applications: A Comprehensive Survey and Practical Guide to Code Intelligence. CoRR abs/2511.18538 (2025). arXiv:2511.18538 doi:10.48550/ARXIV.2511.18538 [44] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https://openreview.net/forum?id=WE_vluYUL-X [45] Daoguang Zan, Zhirong Huang, Wei Liu, Hanwu Chen, Linhao Zhang, Shulin Xin, Lu Chen, Qi Liu, Xiaojian Zhong, Aoyan Li, Siyao Liu, Yongsheng Xiao, Liangqiang Chen, Yuyu Zhang, Jing Su, Tianyu Liu, Rui Long, Kai Shen, and Liang Xiang. 2025. Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving. CoRR abs/2504.02605 (2025). arXiv:2504.02605 doi:10.48550/ARXIV.2504.02605

Received 2026-01-30; accepted 2026-06-25

Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA058. Publication date: October 2026.

Record · ID 668110 · SHA-256 cb606f8cc7038028
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.