arXiv:2604.07767v1 [cs.DC] 9 Apr 2026
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation Senyao Li∗
Zhigang Zuo∗
Haozhao Wang†
[email protected] School of Computer Science and Technology, Huazhong University of Science and Technology China
School of Computer Science and Technology, Huazhong University of Science and Technology China
[email protected] School of Computer Science and Technology, Huazhong University of Science and Technology China
Junyu Chen
Zhanbo Jin
Ruixuan Li
School of Computer Science and Technology, Huazhong University of Science and Technology China
International School Beijing University of Posts and Telccommunication China
School of Computer Science and Technology, Huazhong University of Science and Technology China
Abstract Collaborative edge-cloud frameworks have emerged as the mainstream paradigm for mobile automation, mitigating the latency and privacy risks inherent to monolithic cloud agents. However, existing approaches centralize administration in the cloud while relegating the device to passive execution, inducing a cognitive lag regarding real-time UI dynamics. To tackle this, we introduce AdecPilot by applying the principle of administrative decentralization to the edge-cloud multi-agent framework, which redefines edge agency by decoupling high-level strategic designing from tactical grounding. AdecPilot integrates a UI-agnostic cloud designer generating abstract milestones with a bimodal edge team capable of autonomous tactical planning and self-correction without cloud intervention. Furthermore, AdecPilot employs a Hierarchical Implicit Termination protocol to enforce deterministic stops and prevent postcompletion hallucinations. Extensive experiments demonstrate proposed approach improves task success rate by 21.7% while reducing cloud token consumption by 37.5% against EcoAgent and decreasing end to end latency by 88.9% against CORE. The source code is available at https://anonymous.4open.science/r/Anonymous_codeB8AB.
1
Introduction
Mobile platforms have evolved into the primary interface for digital life, catalyzing the rise of LLM-driven autonomous agents for mobile automation [4, 7, 25, 31]. However, current architectures [16] face a fundamental dilemma: monolithic cloud controllers [4] incur prohibitive latency and privacy risks, while strict on-device [22, 30] constraints hinder the scalability of domain-specific models [10, 20]. To address this dilemma, the edge-cloud collaborative paradigm [21], which deploys a capable agent in the cloud and a lightweight agent at the edge to collaboratively implement the mobile automation, is emerging as the mainstream approach [11, 14]. However, edge-cloud collaboration suffers from a substantial challenge, i.e., reconciling the tension between two fundamental principles: • The scaling law of intelligence posits that reasoning capabilities ∗ Both authors contributed equally to this research. † Corresponding author.
1.Step-by-step Cloud Agent
2.One-shot Planner
3.AdecPilot (Our Approach)
Static Full Plan
Init
Action
…
Upload
Action
Init
Full Plan
Milestone Milestone Milestone
Strategic Milestones
Abstract
…
Status Only
Execute
Observe SLM
Pain Points
Latency
Privacy Risk (Raw Ul)
Pain Points
Edge AI Wasted
Fails on Dynamic Brittle Plan High Cost Ul(Delays) (Restart)
Advantages
Minimal Cloud RTT(Efficient)
Privacy
Robust(SelfCorrecting)
Figure 1: Evolution of Mobile Agent Paradigms. Left & Middle: Conventional methods struggle with high latency (1) or brittle static plans failing under perturbations (2). Right: AdecPilot (3) decouples Strategic Milestones from Tactical Grounding. Autonomous edge Planning ensures robust local resolution, minimal privacy exposure, and minimal cloud consumption.
scale with model parameter counts, favoring large cloud models for complex logic. • The law of observability states that planning efficacy correlates directly with the fidelity and immediacy of real-time data, favoring edge access for precise execution. To balance these two principles, existing works propose leveraging the cloud agent to perform strategic oversight while assigning the edge agent to handle execution [27], which falls into two main categories. The first category [7, 18, 29] involves the cloud agent generating a short-term plan and iteratively refining it into a complete plan throughout the process based on feedback from the edge agent’s execution. The second category [27] entails the cloud agent producing a full plan upfront and then progressively correcting it over time in response to execution feedback from the edge agent. Although these methods have achieved considerable success, both rely entirely on the cloud side to handle all planning-related tasks, relegating the edge agent to a purely mechanical executor, thus
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
resulting in an inability to plan against the visible real-time UI. Specifically, minor deviations such as unexpected icon relocations remain invisible to the remote planner, often causing the entire execution chain to fail before any error is even detected [3, 28]. Moreover, to avoid exposing sensitive user data, these approaches transmit only compressed summaries or cropped image patches, forcing the cloud-based planner to operate in a severely degraded visual environment [12, 21]. In summary, existing methods suffer from Remote Commander Paradox: the entity endowed with the highest intelligence has the poorest perception of the current interface, while the edge observer who possesses real-time visual access remains incapable of performing planning. To address the above challenges, we advocate for Administrative Decentralization within the computational system architecture. We reimagine the cloud as a strategic leader responsible solely for sparse top-level design while delegating concrete planning and execution to the edge. We propose AdecPilot. This framework redefines the collaborative boundary by separating strategic intent from tactical implementation. Adhering to the Scaling Law of Intelligence and Law of Observability, we preserve coarse-grained cloud supervision while empowering the edge with real-time UI observation and self-correction. To address the modality mismatch between high-latency visual diagnosis and high-speed text execution, our edge visual agent performs local planning and observation while utilizing text agent for execution and correction. The system transmits only specific actions rather than indiscriminate screen summaries strictly when the correction step limit is exceeded. To operationalize this decentralization, AdecPilot integrates a UI-agnostic Cloud Strategic Director for high-level decomposition with a Tactical Edge Team comprising a Vision Orchestrator and Textual Executor. Specifically, the cloud defines abstract milestones to guide the global trajectory while the edge autonomously resolves dynamic UI variances via local planning loops. This architecture confines heavy visual processing and atomic decision-making to the device, effectively reducing cloud token consumption and ensuring privacy exposure minimization. To safeguard execution, we further design the Hierarchical Implicit Termination protocol. By restricting validation to the final milestone, this mechanism enforces a deterministic stop upon logic exhaustion and prevents post-completion hallucinations common in lightweight models. Empirical results confirm that this architectural decoupling renders the system immune to network volatility and maintains baseline responsiveness even under severe bandwidth constraints where monolithic models fail. Compared to visual baselines like M3A [18], AdecPilot achieves 388.7× uplink data reduction and 43.8× computational efficiency gain via trajectory distillation. Our primary contributions are as follows:
• Redefining Edge Agency: We propose AdecPilot to decouple strategic cloud milestones from autonomous edge planning. This hierarchical separation ensures robustness against dynamic UI perturbations while minimizing cloud dependency.
Trovato et al.
• Bimodal Autonomy & HIT Protocol: We integrate a VLMbased Orchestrator and Text-based Executor for local planning without cloud intervention. Additionally, our Hierarchical Implicit Termination (HIT) protocol enforces deterministic zero latency exits, effectively preventing post-completion hallucinations. • SOTA Efficiency: Extensive evaluations demonstrate that AdecPilot significantly outperforms state-of-the-art baselines, validating its superiority in reducing cloud token consumption and overcoming transmission latency bottlenecks.
2 Related Work 2.1 Cloud Agents for UI Automation. Advent of Multimodal Large Language Models (MLLMs) transitioned UI automation from heuristic scripts to vision-driven agents [1, 2]. Pioneering frameworks, including AppAgent [29] and T3A [26], adopt stepwise execution paradigms. They utilize monolithic cloud MLLMs to process full-resolution screenshots for atomic action generation [26]. Recent efforts like PRISM incorporate video history to capture temporal execution context [28], attempting to resolve short-term memory deficits inherent to static screenshot analysis. While demonstrating competitive success rates on constrained benchmarks like AndroidWorld [20], these cloud-centric paradigms suffer from fundamental architectural flaws. Primarily, continuous transmission of raw pixels induces a prohibitive trilemma: excessive token consumption, unacceptable network latency, and severe visual privacy leakage [25]. Furthermore, cloud planners operate without real-time state perception, leading to inevitable semantic mismatch when confronting dynamic UI mutations.
2.2
Collaborative Edge-Cloud Multi-Agent
Mitigating resource constraints, recent research pivots toward hybrid collaborative architectures [9, 19, 25]. General-purpose frameworks including AdaSwitch [23] and Division-of-Thoughts [21] explore adaptive mechanisms dynamically distributing inference loads across heterogeneous models based on sample difficulty [6]. Within mobile domains, EcoAgent [27] minimizes cloud interaction via one-shot planning strategies, generating comprehensive action sequences upfront [5]. Conversely, CORE [7] addresses privacy concerns by restricting transmission to text-based UI representations. Despite reducing cloud dependency, these approaches exhibit fundamental architectural flaws in dynamic GUI environments. Such frameworks treat edge modules as passive actuators lacking autonomous tactical resolution. Furthermore, text-only transmissions in CORE [7] discard crucial spatial constraints, resulting in structural ambiguity during execution. Ultimately, existing collaborative paradigms fail to achieve robust Intent Grounding. They physically distribute computation but fail to implement administrative decentralization.
3
Problem Formulation
We formalize multimodal mobile agent workflow as hierarchical decision process parameterized by action space A, observation space O, and synchronization cost function C𝑠𝑦𝑛𝑐 . At step 𝑡, system operates based on defined inputs:
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
task instruction 𝐿𝑐𝑚𝑑 . Variable 𝐻 𝑓(𝑘𝑎𝑖𝑙) ∈ H defines textual diagnostic payload encapsulating failed milestone 𝑔𝑘 , expected invariant 𝐸𝑘 , alongside error execution trajectory, transmitted specifically upon failure 𝑓𝑘 . Transmitting solely textual payload 𝐻 𝑓(𝑘𝑎𝑖𝑙) provides cloud designer sufficient replanning context while mitigating visual data leakage and minimizing synchronization overhead. Term I(𝑓𝑘 ) denotes indicator function for failure event 𝑓𝑘 , formally defined as: ( 1, if Λ𝑒 fails to reach expected state 𝐸𝑘 of 𝑔𝑘 I(𝑓𝑘 ) = (4) 0, otherwise
Self-Correction Cloud Designer
Instruction
Milestones
Visual Orchestrator Agent
Textual Executor Agent
𝒔𝒕
Task Decomposition Expectation
Planning
Execute
Observation
Supplement Strategic Failure
Figure 2: Overview of AdecPilot. The UI-Agnostic Cloud Designer generates abstract milestones, while the Bimodal Edge Team autonomously executes them via local planning. This hierarchical loop enables real-time self-correction, ensuring robust and privacy-preserving automation. Task Instruction 𝐿𝑐𝑚𝑑 ∈ L: Natural language goal provided by user. Application Metadata 𝐶𝑚 : Static invariant functional schema resolving app specific structural priors. Strategic Decomposition: Unlike monolithic oracles 𝑓 : L × O → A depending upon continuous visual feedback, cloud designer Ψ𝑐 functions as open loop strategic designer. Initiating task, module maps user instruction 𝐿𝑐𝑚𝑑 , initial empty history ∅, and abstract app metadata 𝐶𝑚 directly to sequence of UI agnostic milestones 𝐺, strictly bypassing real time observation 𝑜𝑡 : 𝐺 = {(𝑔1, 𝐸 1 ), . . . , (𝑔𝐾 , 𝐸𝐾 )} ∼ 𝑃 Ψ𝑐 (· | 𝐿𝑐𝑚𝑑 , ∅, 𝐶𝑚 )
(1)
Tuple (𝑔𝑘 , 𝐸𝑘 ) encapsulates strategic subgoal 𝑔𝑘 and expected visual invariant 𝐸𝑘 , explicitly excluding low level UI directives. Formulation prioritizes stable business logic over transient rendering details. Bridging inference gap between abstract goals and concrete actions, bimodal edge pipeline assumes tactical autonomy. Process utilizes complementary sequential models: vision centric orchestrator Φ𝑚 aligns visual observation 𝑉𝑡 with expected state 𝐸𝑘 (𝑡 ) synthesizing meta instruction 𝑠𝑡 , while text centric executor Λ𝑒 grounds 𝑠𝑡 against textual hierarchy 𝑈𝑡 yielding atomic action 𝑎𝑡 . Execution pipeline is formally factorized exposing strict sequential dependency: 𝑠𝑡 = Φ𝑚 (𝑉𝑡 , 𝑔𝑘 (𝑡 ) , 𝐸𝑘 (𝑡 ) ),
𝑎𝑡 = Λ𝑒 (𝑠𝑡 , 𝑈𝑡 )
𝐾 ∑︁
I(𝑓𝑘 ) · |𝐻 𝑓(𝑘𝑎𝑖𝑙) |
Formulation mathematically decouples strategic design from tactical execution, confining cumulative errors to local limits.
4
(3)
𝑘=1
Here, operator | · | computes discrete token volume quantifying transmission bandwidth. Initial uplink payload comprises purely
Methodology
As shown in Fig. 2, we introduce AdecPilot. This framework decouples high level intent from low level grounding by assigning environment agnostic strategic design to cloud, while delegating environment specific tactical planning and execution to edge.
4.1
Cloud-Side Strategic Design
As shown in Fig. 3, the strategic cloud designer Ψ𝑐 functions as a high-level meta controller, operating strictly within a latent semantic space to direct the global trajectory of the task. Unlike conventional monolithic agents that entangle strategic reasoning with heavy pixel-level processing, we implement a UI-agnostic designing mechanism driven by a text-only LLM. Design choice is foundational to architecture: deliberately isolating designer from high dimensional raw visual stream 𝑉𝑡 and verbose view hierarchy 𝑈𝑡 compels model deriving milestones solely based upon logical reasoning and common sense knowledge regarding application workflow. Given instruction 𝐿𝑐𝑚𝑑 , failure context 𝐻 𝑓 𝑎𝑖𝑙 (empty initially), and static functional metadata 𝐶𝑚 , cloud generates 𝐾 coarse grained milestones directing global trajectory. Employing heuristic dispatcher, system routes instructions to task specific prompt templates via interrogative markers. Restricting 𝐶𝑚 to invariant functional schema rather than transient visual representation ensures designer remains UI Agnostic. Formal generative process: 𝐾 𝐺 = {(𝑔𝑘 , 𝐸𝑘 )}𝑘=1 ∼ 𝑃 Ψ𝑐 (· | 𝐿𝑐𝑚𝑑 , 𝐻 𝑓 𝑎𝑖𝑙 , 𝐶𝑚 )
(2)
where 𝑔𝑘 (𝑡 ) ∈ 𝐺 denotes currently active milestone at step 𝑡. System descriptive formulation quantifies total cloud communication cost structure over task lifespan. Formulation establishes rigorous information bottleneck, enforcing administrative decentralization principle. Unlike stepwise generation where synchronization cost scales linearly with trajectory length and observation size, decomposition approach limits cloud interaction to initial planning phase and sparse replanning moments triggered by failure. Total synchronization cost C𝑡𝑜𝑡𝑎𝑙 is formulated descriptively as sum of payload token volumes: C𝑡𝑜𝑡𝑎𝑙 = |𝐿𝑐𝑚𝑑 | +
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
(5)
Term 𝑃 Ψ𝑐 denotes strategic generation process executed by cloud designer Ψ𝑐 . Cloud designer processes strictly abstract app metadata 𝐶𝑚 alongside user instruction 𝐿𝑐𝑚𝑑 , decomposing global objective into sequence of UI agnostic milestones 𝐺. Initial planning enforces strictly empty history ∅. Triggering strategic redesign, cloud assimilates desensitized edge error trajectory 𝐻 𝑓 𝑎𝑖𝑙 redefining milestones. Tuple (𝑔𝑘 , 𝐸𝑘 ) encapsulates strategic subgoal 𝑔𝑘 and expected visual invariant 𝐸𝑘 dispatched directly to edge. By construction, cloud agent accesses zero real time rendering 𝑉𝑡 , prioritizing immutable business logic over transient visual details to significantly enhance robustness across heterogeneous device form factors.
4.2
Edge Side Collaborative Planning and Execution
Insight: Reasoning Gap and Modality Mismatch. Design stems from dual critical observations regarding edge intelligence. First,
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Sub-Goal 𝒈𝒌
“What is the first activity after {date} {time} in Simple Calendar Pro?” UpLink
Sample Sample Milestone: Searh for calendar rel….. Sample Milestone: Searh for calendar re……
software … Milestone: Search for … software Expection: Thecalendar calendarrelatsoftware … The ed software Expection: softw OS interface incalendar OS appears… Expectation: appears…The calendar software interface in OS appears…
Design
Cloud Designer
Planning Round ① Instruction s!
Milestone 𝑔!
Coordinates
Memory Retrieval
Planning Round ②
Node Selection
Current Screen
Next Screen
SWIPE(540,2016,540,288)
Execution ② Coordinates
Error log
Screenshot 𝑉!
ect ion f-C or r
Planning
Sel
Screenshot V!
Milestone: Clarify the name of the … software … calendar application in the system Expection: The calendar software Expectation: calendar name interface Display in OS appears… and icon…
Execution ①
Sample Milestone: Search for calendar related software … Expectation: The calendar software interface in OS appears…
Re-designing Sample Sample Milestone: Searh for calendar rel……
Instruction s! Replanning
Node Selection
Current Screen
Next Screen
SWIPE(248,2024,486,213)
Figure 3: Illustration of the AdecPilot workflow. The Cloud Designer orchestrates Strategic Projection, while the device-based Orchestrator and Executor collaborate to enable autonomous Self-Correction via local planning. Crucially, this closed loop is safeguarded by the HIT protocol, which enforces deterministic termination to prevent post-completion hallucinations. empirical evidence reveals reasoning gap within lightweight models. Failures originate not from capacity deficit but from impulsive tendency mapping pixels directly to actions bypassing intermediate analysis. Forcing VLM to generate reasoning trace before acting significantly improves decision quality. Second, system confronts modality mismatch. While VLMs excel diagnosing dynamic visual events, models suffer high latency and low coordinate precision. Conversely, text based models operating on UI trees offer rapid structural action grounding but lack visual context handling rendering anomalies. Resolving granularity mismatches driven by aforementioned observations, proposed bimodal collaborative architecture offloads visual processing entirely to edge. Decoupling strategic logic from implementation details minimizes latency and enhances robustness. Edge utilizes hierarchical pipeline where vision centric orchestrator conducts cognitive reasoning, guiding text centric executor through autonomous tactical planning. 4.2.1 Orchestrator Agent: Visual Reasoning and Planning. The execution cycle functions as a localized autonomous system driven by the Orchestrator Agent Φ𝑚 , parameterized by a quantized VLM. Unlike conventional edge agents confined to the passive execution of atomic commands, the Orchestrator functions as a Tactical Designer. It leverages the abstract expected state 𝐸𝑘 to perceive essential UI elements, bridging the gap between the raw screenshot 𝑉𝑡 and 𝐸𝑘 . Critically, the Orchestrator autonomously evaluates subtask completion; if the state remains unfulfilled, it synthesizes the
subsequent action based on this visual alignment analysis. Limited edge zero-shot capabilities necessitate explicit expected states 𝐸𝑘 for effective diagnosis. To bolster robustness in non-standard UIs, we implement dynamic context injection, overriding generic priors with local logic. This process is formalized as State Alignment Optimization. First, Orchestrator computes visual alignment score 𝑆𝑡 quantifying discrepancy between 𝑉𝑡 and 𝐸𝑘 . Bypassing heuristic vector similarities, system formulates alignment as visual question answering (VQA) verification task executing natively on VLM autoregressive head. Alignment score equals conditional probability of generating affirmative indicator token given visual context and interrogative query: 𝑆𝑡 = 𝑃Φ𝑚 (𝑦𝑡 = 𝑦 + | 𝑉𝑡 , Q (𝐸𝑘 ))
(6)
where 𝑦 + denotes affirmative vocabulary token identifying successful execution. Function Q (·) maps abstract expected state 𝐸𝑘 into deterministic verification query. Continuous confidence measure 𝑆𝑡 ∈ [0, 1] dictates execution continuity. Score falling below threshold 𝜏 (empirically set to 0.85) signifies critical trajectory deviation, instantaneously triggering local tactical re-planning. The selection of 𝜏 dictates the autonomy-cost trade-off: a higher 𝜏 ensures stricter visual alignment but increases local re-planning overhead, whereas a lower 𝜏 risks grounding errors. Instead of reporting failure to the cloud, the Orchestrator engages in planning to generate a corrective Meta-Instruction 𝑠𝑡 . This generates a local corrective trajectory without cloud intervention, achieving privacy exposure minimization and robust handling of dynamic UI elements.
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
4.2.2 Executor Agent: Atomic Structural Grounding. Upon the generation of the meta-instruction 𝑠𝑡 , control is transferred to the Executor Agent Λ𝑒 , instantiated by a lightweight text-only LLM. Designed to alleviate the computational burden on the vision-centric Orchestrator, the Executor exploits the system’s Computational Asymmetry: it operates exclusively on the textual View Hierarchy 𝑈𝑡 , enabling High-Velocity Execution without pixel-level processing. Executor treats task as dual problem integrating structural grounding and atomic actuation. Module must identify optimal DOM node 𝑢 ∗ and determine precise execution action. Replacing arbitrary heuristic matching, formulation frames grounding process as maximizing conditional semantic generation probability 𝑃Λ𝑒 . Variable 𝑃Λ𝑒 explicitly denotes normalized sequence probability aggregating autoregressive token likelihoods generating unique identifier 𝐼 𝐷𝑢 associated with candidate node 𝑢 given structural view hierarchy 𝑈𝑡 and meta instruction 𝑠𝑡 . Optimization objective is formalized as: 𝑢 ∗ = arg max log 𝑃Λ𝑒 (𝑢 | 𝑠𝑡 , 𝑈𝑡 ) − 𝛼R𝑠𝑡𝑟𝑢𝑐𝑡 (𝑢) (7)
Budget serves as explicit failure exploration boundary, fundamentally decoupled from success driven HIT protocol. If orchestrator Φ𝑚 fails to achieve expected state 𝐸𝑘 within 𝑇𝑟𝑒𝑝𝑙𝑎𝑛 , system synthesizes failure context 𝐻 𝑓 𝑎𝑖𝑙 . Empirical step exhaustion cases demonstrate circuit breaker operating correctly rather than algorithmic entrapment. Mechanism accommodates maximal local exploration against dynamic UI mutations, forcing edge agents to exhaust tactical possibilities before yielding. Structure completely prevents infinite cloud queries inherent to monolithic frameworks.
𝑢 ∈𝑈𝑡+
Objective balances semantic probability selecting candidate node 𝑢 against structural regularization term R𝑠𝑡𝑟𝑢𝑐𝑡 governed by scaling factor 𝛼. Search space restricts domain to 𝑈𝑡+ = {𝑢 ∈ 𝑈𝑡 | 𝑣𝑢 = 1}, strictly pruning non interactable elements governed by interactability indicator 𝑣𝑢 parsed natively from underlying Android structural metadata. Term R𝑠𝑡𝑟𝑢𝑐𝑡 mitigates visual hallucinations enforcing spatial layout constraints: R𝑠𝑡𝑟𝑢𝑐𝑡 (𝑢) = ∥p𝑢 − p𝑟𝑒 𝑓 ∥ 2
(8)
Vector p𝑢 denotes geometric centroid of node 𝑢. Spatial reference coordinate p𝑟𝑒 𝑓 resolves location ambiguity, extracted programmatically via regex from point coordinates embedded within orchestrator textual meta instruction output. Optimization explicitly grounds abstract meta instructions into deterministic XML elements, enforcing strict structural validation and mitigating visual hallucinations [10].
4.3
Hierarchical Error Recovery Mechanism
We distinguish locally resolvable Tactical Anomalies from Strategic Failures through hierarchical control loop. Local Self Correction (Inner Loop). Addressing granularity mismatch necessitates closed feedback circuit between edge agents. Textual executor acts as rapid filter: failing to identify node 𝑢 ∗ satisfying semantic constraints triggers tactical feedback signal F𝑡𝑎𝑐𝑡 instead of random actuation. Visual orchestrator integrates F𝑡𝑎𝑐𝑡 , visual context 𝑉𝑡 , and expected state 𝐸𝑘 performing direct prompt conditioning. Autoregressive generation directly synthesizes revised meta instruction 𝑠𝑡 +1 bypassing explicit probability marginalization. Paradigm proves feasibility regarding edge multi agent collaboration. Visual orchestrator and textual executor construct autonomous reasoning loop entirely on device. System resolves transient perturbations locally, strictly confining sensitive observation data to edge hardware. Design mitigates visual privacy leakage and eliminates redundant cloud token consumption. Strategic Redesigning (Outer Loop). To manage strategic deadends, system enforces tactical step budget 𝑇𝑟𝑒𝑝𝑙𝑎𝑛 per sub goal.
𝑇
−1
𝑟𝑒𝑝𝑙𝑎𝑛 𝐻 𝑓 𝑎𝑖𝑙 = ⟨(𝑔𝑘 , 𝐸𝑘 ), {(𝑄𝑡 , 𝑎𝑡 )}𝑡 =0 ⟩. | {z } | {z }
Cloud
(9)
Edge
Variable 𝑄𝑡 denotes textual execution trajectory of active milestone. Although 𝑄𝑡 incurs marginal privacy exposure, textual representation significantly mitigates leakage compared to frameworks transmitting raw visual trajectories or frame summaries. Receiving 𝐻 𝑓 𝑎𝑖𝑙 , cloud designer Ψ𝑐 transitions to Diagnostician, identifying root causes and regenerating corrected trajectory 𝐺 ′ . Architecture bounds error propagation, invoking expensive cloud intelligence strictly for genuine strategic failures. Maintaining optimal balance between local autonomy and global reasoning.
4.4
Adaptive Termination via Action Pruning
Standard benchmarks like AndroidWorld [26] impose a rigid termination tax by requiring explicit token generation. While feasible for cloud models, this protocol often induces Post-Completion Hallucination in lightweight edge models [13], where the agent invents destructive actions instead of stopping. To mitigate this, we propose the Hierarchical Implicit Termination (HIT) strategy, which enforces a "Fast Finish" via a multi-priority cascade restricted to the Final Milestone Phase. The execution flow is governed by three strictly ordered protocols. Priority 1 (System Level Real Time Detection) activates entering final sub goal. System intercepts structural view hierarchy 𝑈𝑡 post atomic action capturing deterministic OS callbacks including toast notifications, triggering environment terminate() before model hallucinates. Priority 2 (Designer Level Logic Exhaustion) signals definitive completion triggered upon sub goal queue depletion. For question answering tasks, orchestrator asserts ANSWER_READY state explicitly when visual alignment score 𝑆𝑡 > 𝜏𝑞𝑎 against expected text bounds, forcing immediate stop. Priority 3 (Budgetary Fallback) enforces strict global step limit 𝑇𝑚𝑎𝑥 resolving infinite loops. Upon triggering any priority, system wrapper executes environment terminate() function injecting either static success token or VLM extracted textual payload. This mechanism ensures evaluation metrics reflect agent capability rather than adherence to verbose syntax. However, HIT is calibrated for finite-horizon tasks; continuous orchestrating scenarios require adaptation to sliding-window triggers to prevent premature termination.
5
Experiments
This section describes experimental settings in Sec. 5.1, presents quantitative comparisons between AdecPilot and state-of-the-art baselines across success rate and efficiency in Sec. 5.2, analyzes privacy metrics in Sec. 5.3, evaluates boundary performance in Sec. 5.4, and concludes with ablation studies in Sec. 5.5.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Table 1: Main Performance on AndroidWorld. Evaluates task success rate SR, cloud token usage MT, relative cloud energy RCE. Metric RCE incorporates penalty factor 𝜇 = 1.2 for continuous image streaming. Symbol ∗ indicates latest Instruct version. Method
Architecture
Cloud Model
Edge Model
SR↑
MT↓
RCE↓
Privacy
ShowUI [15]
Single-Agent
InfiGUIAgent [17]
Single-Agent
–
ShowUI-2B
7.0%
0
0×
✓ Safe
–
InfiGUIAgent-2B
9.0%
0
0×
✓ Safe
Pure Device Baselines
Pure Cloud Baselines AppAgent [29]
Single-Agent
GPT-4o
–
11.2%
∼15k
9.0×
× High Risk
M3A [18]
Multi-Agent
GPT-4o×2
–
28.4%
∼87k
52.2×
× High Risk
Cloud-Device Collaborative UGround [8]
Open-loop
GPT-4o×2
UGround-2B
32.8%
∼45k
27.0×
× High Risk
EcoAgent [27]
Closed-loop
GPT-4o
OS-Atlas-4B+Qwen2-VL-2B
27.6%
∼3.2k
1.9×
! Text Summary
CORE [7]
Open-loop
GPT-4o
Gemma-2-9B-IT
26.7%
∼11.3k
6.78×
✓ Exposure Minimization
AdecPilot (Ours)
Hierarchical
GPT-4o
Qwen2.5-3B+Qwen3-VL-2B
33.6%
∼2k
1.0× (Baseline)
✓ Exposure Minimization
AdecPilot Pro (Ours)
Hierarchical
GPT-4o
Qwen3-4B+Qwen3-VL-2B∗
34.5%
∼1.9k
0.95×
✓ Exposure Minimization
(a) Performance-Efficiency Landscape Task Success Rate
(b) Domain-Specific Success Rate (AndroidWorld)
AdecPilot (Ours) CORE EcoAgent UGround M3A AppAgent
Contacts
Simple Calendar
(c) Domain-Specific Success Rate (AndroidLab)
AdecPilot (Ours) CORE EcoAgent AdecPilot Avg CORE Avg Eco Avg
Settings
AdecPilot (Ours) CORE EcoAgent GPT-4o Gemini-1.5-Pro LLaMA3.1-8B
Bluecoins Zoom
Calendar 80
Routine Tasks 60
60
Markor
Camera 40
20
Broccoli (Recipe)
Real-time Speed (1/Latency)
20
Settings
Complex Tasks
Cantook
Clock
Cost Efficiency (1/Tokens)
PiMusic
Clock
Audio Recorder
OpenTracks
Joplin
Simple SMS
Map
Contacts
Figure 4: Performance analysis. (a) Comprehensive comparison across task success rate and token cost. (b) Completion rate within AndroidWorld domain breakdown. (c) Completion rate within AndroidLab domain breakdown. Baseline selection strictly adapts to benchmark innate evaluation paradigms.
5.1
Implementation Details
Cloud designer Ψ𝑐 utilizes GPT-4o. Orchestrator agent Φ𝑚 deploys quantized Qwen3-VL-2B model on local server equipped with NVIDIA RTX 4070 TiS GPU. Replanning limit is set to 𝑅 = 1 round. Structural regularization factor 𝛼 within Eq. 7 is empirically set to 0.2, balancing semantic probability against spatial constraints bypassing exhaustive ablation. Datasets. Primary evaluation utilizes AndroidWorld [20] comprising 116 tasks across 20 applications. Programmatic verification ensures reproducibility over manual alternatives [24]. System operates on Pixel 6 emulator utilizing API 33. Accounting for environmental randomness, evaluations average across three independent runs utilizing distinct random seeds. Results exhibit minimal ±0.8% task success rate standard deviation, confirming objective stability and statistical significance. Although manual verification renders AndroidLab [26] suboptimal for automated scaling, system adjusts execution configurations benchmarking AdecPilot ensuring comprehensive cross benchmark comparison.
Metrics. System evaluation systematically investigates four critical dimensions: efficacy, operational efficiency, latency, and privacy preservation. These are measured by following metrics. Task Success Rate. Metric measures efficacy, defined as percentage of successfully completed tasks relative to total test corpus. Efficiency. Operational efficiency is quantified via Average Cloud Calls (MC) and Cloud Token Usage (MT). To normalize resource consumption, we report Relative Cloud Energy (RCE), estimating aggregate cloud-side burden compared against baseline. Reduction Rate. Metric evaluates privacy mitigation measuring reduction in UI elements uploaded to cloud compared against GPT-4o baseline. Enforcing strict fairness, evaluation exclusively considers rounds where comparative methods execute identical decisions on identical UI screens [7]. Let 𝐸𝐺𝑃𝑇 −4𝑜 and 𝐸𝑜𝑢𝑟𝑠 denote quantity of UI elements transmitted by GPT-4o baseline and AdecPilot respectively under identical conditions. Reduction rate calculation:
RR =
𝐸𝐺𝑃𝑇 −4𝑜 − 𝐸𝑜𝑢𝑟𝑠 𝐸𝐺𝑃𝑇 −4𝑜
(10)
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Minimizing 𝐸𝑜𝑢𝑟𝑠 inherently mitigates visual privacy leakage, objectively reducing the raw structural data exposure.
within DOM trees, decoupling mitigates image based privacy risks. AdecPilot Pro success rate (33.6%) trails Qwen2.5-Max (35.3%). Variance constitutes architectural trade off: trading 1.7% success rate mitigates 79.3% visual exposure while accelerating responsiveness. Strategy prioritizes operational efficiency over pure scaling [25]. Table 3 evaluates operational efficiency across uplink communication and local computation. Regarding data transmission, visual baselines including M3A [18] incur prohibitive costs reaching 5831 kB via continuous video transmission. EcoAgent lowers overhead to 120 kB. In contrast, AdecPilot transmits only textual logs minimizing load to 15 kB. This represents 8.0× reduction over EcoAgent and 388.7× reduction over M3A. Quantitative evidence validates text centric collaborative approach as optimal paradigm resolving cellular transmission bottlenecks. Regarding computational efficiency, monolithic baselines including M3A demand extreme computational resources reaching 529.16 TFLOPs per step. Collaborative baseline EcoAgent executes multimodal forward passes reaching 22.35 TFLOPs per atomic action. Conversely, proposed bimodal architecture enforces computational asymmetry. Orchestrator 2B handles sparse visual alignment utilizing 1024px downsampling while Executor 3B executes precise Intent Grounding operating as text only LLM. Decoupled design reduces total per step computation to 12.09 TFLOPs. Achieving 43.8× FLOPs reduction over M3A baseline, architecture fundamentally resolves power intensive multimodal redundancy inherent within visual frameworks validating local server execution feasibility. Communication Efficiency. Fig. 5(a) illustrates average cloud calls and token consumption. AdecPilot Pro minimizes metrics to 1.4 calls and 1.9k tokens per task. Conversely, monolithic baselines exhibit severe dependence upon continuous cloud synchronization. M3A requires 13.4 interactions consuming 87.0k tokens, while UGround consumes 45.0k tokens. Even optimized collaborative framework EcoAgent demands ∼ 3.2k tokens. Superior efficiency directly stems from administrative decentralization architecture. Delegating tactical planning entirely to edge orchestrator restricts cloud interaction strictly to initial strategic decomposition and sparse failure recovery. Design eliminates redundant step validation, fundamentally resolving transmission bottlenecks.
Table 2: AndroidWorld Privacy Performance. Table evaluates Success Rate(SR) and Reduction Rate(RR). Method
SR
RR
Qwen2.5-Max (Base) CORE [7](In Qwen2.5-Max)
35.3% (41/116) 27.6% (32/116)
0.0% 37.0%
AdecPilot (In Qwen2.5-Max) AdecPilot Pro (In Qwen2.5-Max)
31.9% (37/116) 33.6% (39/116)
75.7% 79.3%
5.2
Analysis of Cloud Tokens and Success Rate
Table 1 compares efficacy and operational cost. AdecPilot achieves a superior balance between execution capability and resource consumption. Notably, our method attains a success rate of 33.6%, surpassing EcoAgent [27] (27.6%). Crucially, we achieve this with reduced cloud reliance. EcoAgent consumes ∼3.2k tokens; proposed framework requires 2k. RCE evaluation, incorporating streaming penalties, achieves 1.9× overhead reduction over strongest collaborative baseline. Against monolithic M3A at ∼87k tokens, RCE reduction reaches factor 52.2. Metrics validate Tactical Planning: bimodal agents resolve granular actions without constant cloud queries. Minimal RCE 1.0× confirms offloading visual reasoning to edge, establishing standard for sustainable mobile automation. Figure 4 illustrates holistic performance efficiency trade offs and domain specific robustness. As shown in Fig. 4a, AdecPilot establishes optimal Pareto frontier among evaluated baselines. Monolithic cloud agents occupy low efficiency regions due to excessive token consumption. Proposed framework maximizes radar polygon area encompassing task success rate, real time speed, and cost efficiency. Cross Benchmark Robustness. Fig. 4b and Fig. 4c present success rates across AndroidWorld categories with successful baseline cases and AndroidLab respectively. As illustrated visually within complex tasks region, AdecPilot demonstrates superior generalization in logic heavy apps including Joplin and Broccoli. Furthermore, while customized path dependent sub goals within AndroidLab interfere with autonomous edge exploration, empirical results confirm consistent superiority over monolithic and collaborative baselines under strict path constraints. As illustrated within Fig. 4c, system achieves 56% success rate versus CORE [7] 37% evaluating Contacts. Unlike CORE exhibiting performance degradation in high fidelity spatial reasoning environments reaching only 17% evaluating PiMusic, proposed bimodal architecture ensures precise Intent Grounding achieving 27%. AndroidLab results confirm decoupling strategic intent from tactical execution effectively resolves domain specific brittle failures inherent in monolithic paradigms.
5.3
Analysis of Privacy and Communication Overhead
Table 2 reports visual privacy mitigation via Reduction Rate (RR). AdecPilot Pro reduces uploaded UI elements by 79.3% over monolithic baseline; CORE achieves 37.0%. Restricting uplink to abstract milestones confines visual data locally. Despite textual identifiers
5.4
Boundary Performance
Network Robustness. Fig. 5(b) depicts end to end latency across degrading bandwidth from high speed WiFi to severely constrained 2G networks. Monolithic visual streaming models including M3A and UGround suffer exponential latency surges under weak network conditions, exceeding 50 seconds under 2G bandwidth. In stark contrast, AdecPilot demonstrates exceptional environmental resilience, maintaining stable 3.03s responsiveness even under 50K/s constraints. Architecture confines visual pixel processing and autonomous self correction loops within device hardware. Transmitting solely lightweight text based abstract milestones immunizes system against unpredictable network volatility. Failure Attribution. Fig. 5(c) objectively details system failures. Primary bottlenecks concentrate in Vision Orchestrator (14 instances) and System Budget exhaustion (11 instances). While budget exhaustion confirms designed circuit breaker prevents infinite
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Table 3: Comprehensive Evaluation of System Overhead and Computational Efficiency. Table compares uplink data volume, data reduction ratios, device side TFLOPs per step, and computational efficiency gains across diverse agent architectures. Method
Edge Model Configuration
Uplink Data (kB) ↓
Data Reduction
TFLOPs / Step ↓
FLOPs Reduction
M3A [18]
–
5831
1.0×
529.16
1.00×
AppAgent [29]
–
2098
2.8×
190.39
2.8×
CORE [7]
Gemma 2 9B IT
741
7.86×
138.03
3.8×
EcoAgent [27]
ShowUI-2B + Qwen2-VL-2B
120
48.6×
22.35
23.7×
AdecPilot
Qwen2.5-3B + Qwen3-VL-2B
15
388.7×
12.09
43.8×
(a) Communication Overhead
(b) Latency vs. Bandwidth 16 14
12.2 13.4
60 45.0
6.5
8
6.5
40
6 4
20
15.0
1.9
11.3
3.2
0 A M3
r UG
d oun
t gen ppA
A
C
E OR
t gen
oA
Ec
Cloud Calls (MC)
12 10
16 M3A UGround AppAgent
60
End-to-End Latency (s)
Average MLLM Tokens(k) (MT)
87.0
80
(c) System Failure Attribution
70 CORE EcoAgent AdecPilot (Ours)
12
50 40 30 20 10
0
0
11
10 8 6 4
1.4
2
14
14
Failure Count
100
2
3 2
1.9
rs
Ou
Cloud Only
WiFi (10M/s)
4G (1M/s)
3G (200K/s)
2G (50K/s)
0
Cloud Design
Vision Orchestrator
Text Executor
System Budget
Figure 5: Holistic System Analysis (a) Average cloud calls (MC) and cloud token usage (MT) comparison within AndroidWorld. (b) End to end latency comparison across weak network environments. (c) Failure case distribution analysis within AndroidWorld. Table 4: Ablation Study on Architecture and Expected State. Validates textual executor Λ𝑒 , visual orchestrator Φ𝑚 , and explicit expected state 𝐸𝑘 within AndroidWorld dataset. Configurations utilizing Qwen2.5 3B and Qwen3 4B represent AdecPilot and AdecPilot Pro respectively. Method
SR ↑
Steps ↓
MT ↓
Replan ↓
w/o Executor Λ𝑒 w/o Orchestrator Φ𝑚 w/o Expectation 𝐸𝑘 AdecPilot(Qwen2.5-3B)
11.2% 6.0% 16.4% 33.6%
6.04 17.94 13.99 9.77
2492 6532 4410 2024
0.54 0.89 0.79 0.51
w/o Executor Λ𝑒 w/o Orchestrator Φ𝑚 w/o Expectation 𝐸𝑘 AdecPilot Pro(Qwen3-4B)
11.6% 7.8% 19.8% 34.5%
7.02 16.83 15.62 10.19
2006 6287 4534 1904
0.49 0.85 0.83 0.45
cloud queries, non trivial proportion (11/30) indicates severe dynamic UI mutations can still entrap edge agents in excessive local replanning overhead. Orchestrator anomalies indicate edge visual verification limits remain primary capability ceiling, establishing clear target for future trajectory distillation research.
5.5
Ablation Study
As detailed within Table 4, removing textual executor Λ𝑒 or visual orchestrator Φ𝑚 induces severe degradation. Omitting Φ𝑚 collapses SR to 6.0% while MT surges 3.2× (6532 tokens), confirming lacking visual verification triggers blind execution loops and redundant cloud redesigning. Expected state 𝐸𝑘 is foundational for tactical autonomy. Deleting 𝐸𝑘 forces unguided state alignment, dropping Qwen2.5 3B SR to 16.4% elevating replan rate to 0.79, and dropping Qwen3 4B SR to 19.8% elevating replan rate to 0.83. Table 4 indicates AdecPilot Pro (Qwen3 4B) achieves 34.5% SR while minimizing MT to 1904. Superior reasoning density reduces average replanning to 0.45, confirming efficient local self correction. As shown in Fig. 5b, AdecPilot sustains stable 3.03s latency under 2G constraints. Unlike visual streaming baselines exceeding 50s, proposed hierarchical decoupling ensures high speed local loops remain immune to bandwidth volatility. Hallucination Mitigation Analysis. Tab. 5 removing HIT protocol induces distinct hallucination paradigms across task categories including Question Answer and Operation. Operation tasks primarily suffer from post completion hallucination. Textual executor completes Intent Grounding but visual orchestrator lacks termination awareness. Conversely question answering tasks exhibit premature termination; agent observes target information visually but erroneously asserts completion before executing explicit textual answering. Results validate HIT protocol necessity
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Table 5: Ablation Study on Hallucination Mitigation. Total 116 AndroidWorld tasks partition into operation and question answering subsets evaluating success rate SR. Metric PCH signifies post completion hallucination rate.
[10] Xiao Han, Chen Zhu, Hengshu Zhu, and Xiangyu Zhao. 2025. Swarm intelligence in geo-localization: A multi-agent large vision-language model collaborative framework. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 814–825. [11] Senkang Hu, Yihang Tao, Guowen Xu, Yiqin Deng, Xianhao Chen, Yuguang Fang, and Sam Kwong. 2025. CP-Guard: Malicious Agent Detection and Defense in Collaborative Bird’s Eye View Perception. In AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA, USA. 23203–23211. [12] Guanjie Huang, Danny HK Tsang, Shan Yang, Guangzhi Lei, and Li Liu. 2025. Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition. In Proceedings of the 33rd ACM International Conference on Multimedia. 8313–8321. [13] Zaid Khan, Elias Stengel-Eskin, Jaemin Cho, and Mohit Bansal. 2025. DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. [14] Senyao Li, Haozhao Wang, Wenchao Xu, Rui Zhang, Song Guo, Jingling Yuan, Xian Zhong, Tianwei Zhang, and Ruixuan Li. 2025. Collaborative inference and learning between edge slms and cloud LLMs: A survey of algorithms, execution, and open challenges. arXiv preprint arXiv:2507.16731 (2025). [15] Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, and Mike Zheng Shou. 2025. Showui: One vision-language-action model for gui visual agent. In Proceedings of the Computer Vision and Pattern Recognition Conference. 19498–19508. [16] Wenjin Liu, Haoran Luo, Xueyuan Lin, Haoming Liu, Tiesunlong Shen, Jiapu Wang, Rui Mao, and Erik Cambria. 2025. Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning. arXiv preprint arXiv:2511.01016 (2025). [17] Yuhang Liu, Pengxiang Li, Zishu Wei, Congkai Xie, Xueyu Hu, Xinchen Xu, Shengyu Zhang, Xiaotian Han, Hongxia Yang, and Fei Wu. 2026. Infiguiagent: A multimodal generalist gui agent with native reasoning and reflection. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 1035–1051. [18] Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, Yuan Lin, Hang Li, Junbo Zhao, and Wei Li. 2025. Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory. arXiv preprint arXiv:2508.09736 (2025). [19] Jiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2025. Multi-agent system for comprehensive soccer understanding. In Proceedings of the 33rd ACM International Conference on Multimedia. 3654–3663. [20] Chris Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Toyama, Robert Berry, Divya Tyamagundlu, Timothy Lillicrap, and Oriana Riva. 2025. AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents. In International Conference on Representation Learning. 406–441. [21] Chenyang Shao, Xinyuan Hu, Yutang Lin, and Fengli Xu. 2025. Division-ofThoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents. In Proceedings of the ACM on Web Conference 2025, WWW 2025, Sydney, NSW, Australia, 28 April 2025- 2 May 2025. 1822–1833. [22] Hongda Sun, Hongzhan Lin, Haiyu Yan, Yang Song, Xin Gao, and Rui Yan. 2025. MockLLM: A Multi-Agent Behavior Collaboration Framework for Online Job Seeking and Recruiting. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, V.2, KDD 2025, Toronto ON, Canada, August 3-7, 2025. 2714–2724. [23] Hao Sun, Jiayi Wu, Hengyi Cai, Xiaochi Wei, Yue Feng, Bo Wang, Shuaiqiang Wang, Yan Zhang, and Dawei Yin. 2024. AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024. 8052–8062. [24] Luyuan Wang, Yongyu Deng, Yiwei Zha, Guodong Mao, Qinmin Wang, Tianchen Min, Wei Chen, and Shoufa Chen. 2024. Mobileagentbench: An efficient and user-friendly benchmark for mobile llm agents. arXiv preprint arXiv:2406.08184 (2024). [25] Quanmin Wei, Penglin Dai, Wei Li, Bingyi Liu, and Xiao Wu. 2025. CoPEFT: Fast Adaptation Framework for Multi-Agent Collaborative Perception with ParameterEfficient Fine-Tuning. In AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA, USA. 23351–23359. [26] Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, and Yuxiao Dong. 2025. Androidlab: Training and systematic benchmarking of android autonomous agents. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2144–2166. [27] Biao Yi, Xavier Hu, Yurun Chen, Shengyu Zhang, Hongxia Yang, Fan Wu, and Fei Wu. 2025. EcoAgent: An Efficient Edge-Cloud Collaborative Multi-Agent Framework for Mobile Automation. arXiv preprint arXiv:2505.05440 (2025). [28] Junfei Zhan, Haoxun Shen, Zheng Lin, and Tengjiao He. 2025. PRISM: PrivacyAware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch
Task type
Method
Steps ↓
SR ↑
PCH ↓
Operation
w/o HIT AdecPilot Pro
13.04 10.98
24/91 31/91
11/91 0/91
Question Answer
w/o HIT AdecPilot Pro
6.36 7.28
1/25 9/25
19/25 3/25
enforcing deterministic exits and resolving category specific evaluation anomalies.
6
Conclusion
Paper introduces AdecPilot mobile automation collaborative framework. Administrative decentralization decouples strategic milestones from autonomous tactical planning. Architecture endows edge teams with real time visual reasoning capability resolving edge cloud cognitive lag. Bimodal design ensures robust Intent Grounding under severe network deterioration. HIT protocol eliminates hallucination via deterministic exit. Empirical results confirm framework establishes leading superiority across multiple benchmarks.
References [1] Theodoros Aslanidis, Sokol Kosta, Spyros Lalis, and Dimitris Chatzopoulos. 2025. Cross-Domain DRL Agents for Efficient Job Placement in the Cloud-Edge Continuum. In Proceedings of the 5th Workshop on Machine Learning and Systems, EuroMLSys 2025, World Trade Center, Rotterdam, The Netherlands, 30 March 20253 April 2025. 276–285. [2] Hao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri, Alane Suhr, Sergey Levine, and Aviral Kumar. 2024. DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. [3] Ronshee Chawla, Daniel Vial, Sanjay Shakkottai, and R. Srikant. 2023. Collaborative Multi-Agent Heterogeneous Multi-Armed Bandits. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202). 4189–4217. [4] Xinran Chen, Yuchen Li, Hengyi Cai, Zhuoran Ma, Xuanang Chen, Haoyi Xiong, Shuaiqiang Wang, Ben He, Le Sun, and Dawei Yin. 2025. Multi-agent proactive information seeking with adaptive llm orchestration for non-factoid question answering. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 4341–4352. [5] Tong Cheng, Hang Dong, Lu Wang, Bo Qiao, Si Qin, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, and Thomas Moscibroda. 2023. Multi-Agent Reinforcement Learning with Shared Policy for Cloud Quota Management Problem. In Companion Proceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023. 391–395. [6] Alex Clinton, Yiding Chen, Jerry Zhu, and Kirthevasan Kandasamy. 2025. Collaborative Mean Estimation Among Heterogeneous Strategic Agents: Individual Rationality, Fairness, and Truthful Contribution. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025. [7] Gucongcong Fan, Chaoyue Niu, chengfei lv, Fan Wu, and Guihai Chen. 2025. CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [8] Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2025. Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents. In The Thirteenth International Conference on Learning Representations. [9] Xuehang Guo, Xingyao Wang, Yangyi Chen, Sha Li, Chi Han, Manling Li, and Heng Ji. 2025. SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Collaboration. arXiv preprint arXiv:2511.22788 (2025). [29] Chi Zhang, Zhao Yang, Jiaxuan Liu, Yanda Li, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2025. Appagent: Multimodal agents as smartphone users. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–20. [30] Guibin Zhang, Luyang Niu, Junfeng Fang, Kun Wang, Lei Bai, and Xiang Wang. 2025. Multi-agent Architecture Search via Agentic Supernet. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025. [31] Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang, Huazheng Wang, Yiran Chen, and Qingyun Wu. 2025. Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025.
Trovato et al.
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
A
Hyperparameter Sensitivity Analysis
Appendix investigates system robustness against critical scalar configurations. Primary focus evaluates visual alignment threshold 𝜏 governing local self correction and structural regularization factor 𝛼 dictating spatial grounding precision.
A.1
Table 6: Sensitivity Analysis on Visual Alignment Threshold 𝜏. Evaluates success rate SR, average steps, and replan frequency per task. Configuration 0.85 represents baseline AdecPilot Pro.
𝜏
SR ↑
Steps ↓
Replan Rate ↓
0.40
23.2%
6.88
0.12
0.60
30.2%
7.54
0.23
0.80
32.7%
9.35
0.38
0.85 (Ours)
34.5%
10.19
0.45
0.90
31.8%
13.45
0.76
0.95
29.3%
16.88
0.88
Structural Regularization Factor
Equation 7 introduces factor 𝛼 balancing textual semantic probability against geometric spatial constraints mitigating executor hallucination. Table 7 details spatial metric variations. Eliminating structural penalty 𝛼 = 0.0 induces severe spatial hallucination rate SHR 42.1% where executor selects semantically plausible but spatially incorrect elements. Extreme penalty 𝛼 = 1.0 overrides semantic logic forcing rigid coordinate selection dropping success rate to 24.1%. Configuration 𝛼 = 0.2 optimally suppresses SHR bounding value 5.4% sustaining maximal task completion.
A.3
Table 7: Sensitivity Analysis on Structural Regularization Factor 𝛼. Evaluates spatial hallucination rate SHR and overall success rate SR. Configuration 0.2 represents optimal baseline.
Visual Alignment Threshold
Threshold 𝜏 defines autonomy boundary within visual orchestrator Φ𝑚 . Metric dictates minimum acceptable confidence score 𝑆𝑡 required asserting expected state 𝐸𝑘 completion. Table 6 presents ablation across AndroidWorld dataset consistent with evaluation protocol in Section 5.1. Setting 𝜏 excessively low 0.40 permits premature milestone transition inducing cascading logic failures yielding marginal success rate 23.2%. Conversely excessively strict configurations 0.95 trigger redundant localized replanning exhausting step budget 𝑇𝑚𝑎𝑥 reducing success rate 29.3%. Empirical data validates configuration 𝜏 = 0.85 establishing optimal balance maximizing success rate 34.5% maintaining efficient average step count 10.19.
A.2
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
End-to-End Latency and Robustness Analysis
Table 8 highlights a critical divergence in network resilience. Visual baselines like M3A are 20× slower than our approach under 2G due to image transmission, whereas AdecPilot maintains stable performance, rising only from 2.88s to 3.03s. By limiting cloud interaction to abstract intent analysis which consumes 1.9k tokens
𝛼
SHR ↓
SR ↑
0.0
42.1%
22.4%
0.1
18.5%
29.3%
0.2 (Ours)
5.4%
34.5%
0.5
3.2%
30.2%
1.0
2.1%
24.1%
against 87k, we eliminate transmission bottlenecks. This validates that architectural decoupling renders the system immune to bandwidth volatility, ensuring responsiveness where monolithic agents fail.
A.4
Multi User Concurrent Scalability and Power Efficiency
Administrative decentralization fundamentally mitigates cloud computational bottlenecks during concurrent deployment. Table 9 evaluates system scalability metrics comparing 3 concurrent mobile agents against 5 concurrent mobile agents. Monolithic baseline M3A exhibits throughput scaling increasing from 0.76 to 1.25 Queries Per Second QPS alongside Resource Consumption Energy RCE spikes reaching 258.4. Concurrent visual processing resource saturation maintains query latency high around 54.11 to 55.32 seconds. Conversely AdecPilot confines high frequency visual verification locally. Architecture delegates strictly abstract strategic planning to cloud maintaining nearly perfect linear throughput scaling from 1.44 to 2.46 QPS. System suppresses RCE below 5.0 ensuring query latency remains perfectly stable around 3.11 seconds under 5 concurrent streams. Empirical evaluation restricts maximum concurrent instances to 5 concurrent emulated instances utilizing Pixel 6 emulator API 33 consistent with single agent evaluation protocol established within Section 5.1. Established scaling trajectories rigorously confirm edge cloud decoupling paradigm acts as prerequisite sustaining large scale mobile automation deployment. Isolated edge execution paradigm establishes optimal structural foundation deploying end cloud collaborative speculative decoding mechanisms theoretically enabling further exponential reduction regarding cloud side power consumption across massive multi user concurrent scenarios.
B
System Prompt Formulation
Section details meta instructions governing bimodal edge cloud multi agent system. Prompts enforce strict task boundaries ensuring administrative decentralization.
B.1
Cloud Designer Prompt
Cloud designer objective generates UI agnostic strategic milestones avoiding low level execution coordinates. Input comprises task instruction and application metadata. Output necessitates abstract milestone formatting.
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Table 8: End-to-End Latency Analysis. We report the total latency (inference + transmission) across varying network bandwidths (WiFi to 2G), where Cloud Inference time is estimated based on model modality.
MC↓
MT↓
Method
WiFi
4G
3G
2G
Gap(Slower)
(10M/s)
(1M/s)
(200K/s)
(50K/s)
(vs. 2G)
Cloud Latency (Calls)
(Tokens)
AppAgent [29]
6.46
15k
25.84s
25.85s
25.90s
26.14s
27.04s
8.9×
M3A [18]
13.39
87k
53.56s
53.60s
53.91s
55.30s
60.52s
20.0×
Pure Cloud Baselines
Cloud-Device Collaborative UGround [8]
12.21
45k
48.84s
48.86s
49.02s
49.74s
52.44s
17.3×
CORE [7]
6.46
11.3k
25.84s
25.84s
25.89s
26.07s
26.74s
8.8×
EcoAgent [27]
1.86
3.2K
5.72s
5.72s
5.73s
5.78s
5.98s
2.0×
AdecPilot Pro (Ours)
1.44
1.9k
2.88s
2.88s
2.89s
2.92s
3.03s
Baseline
Table 9: Cloud Infrastructure Scalability under Concurrent Loads. Evaluates throughput Query/Sec, Relative Cloud Energy RCE, and Average Query Latency Seconds comparing 3 concurrent mobile agents versus 5 concurrent mobile agents.
Throughput(Query/Sec) ↑
RCE ↓
Query Latency(Sec) ↓
Method
Baselines
Ours
3 Agent
5 Agents
3 Agent
5 Agents
3 Agent
5 Agents
M3A [18]
0.76
1.25
158.1
258.4
54.11
55.32
UGround [8]
0.72
1.23
80.8
133.4
50.14
51.89
EcoAgent [27]
1.01
1.62
5.9
9.1
5.75
5.96
AdecPilot
1.44
2.46
3.1
4.9
3.06
3.11
AdecPilot Pro
1.51
2.51
2.9
4.6
2.91
3.01
System Formulation: Role dictates strategic planner orchestrating mobile application workflows. Given user instruction and application metadata generate chronological milestone sequence. Strict constraint forbids coordinate generation. Milestone must encapsulate abstract functional goal. Expectation must encapsulate deterministic visual invariant indicating milestone completion. Format output utilizing tuple structure 𝑔𝑘 , 𝐸𝑘 . Cloud Designer Prompt role: You are a High-Level Strategic Planner for Android automation. goal: {task_instruction} task: Break this goal into a list of semantic milestones. for each milestone, define: • instruction: the high-level goal of this step (what to achieve, not how) • expectation: the visual state of the screen after this step is completed rules: • Do not specify specific UI elements such as “click the blue button” • Focus purely on the state transition good examples: {"instruction": "Open the Contacts app.", "expectation": "The Contacts app main list is visible."} {"instruction": "Fill in ‘Alice’.", "expectation": "The Name field shows ‘Alice’."}
Cloud Designer Prompt output format: [ {"instruction": "...", "expectation": "..."}, ... ]
B.2
Cloud Replan Prompt
Cloud replan objective revises the strategic milestone sequence when execution deviates from the original plan. Input comprises current task instruction, previous milestone sequence, and failed execution trajectory. Output necessitates a new milestone sequence aligned with the current environment state and remaining task objective. System Formulation: Role dictates strategic replanner for mobile application workflows. Given current user instruction, previous milestones, and failure trajectory, generate a revised chronological milestone sequence that recovers from execution failure while preserving task intent. Replanning must account for completed progress, discard invalid or obsolete milestones, and infer the most
Administrative Decentralization in Edge-Cloud Multi-Agent for Mobile Automation
plausible continuation from the current interface state. Strict constraint forbids low-level coordinates or primitive action descriptions. Each milestone must encode an abstract functional goal, and each expectation must specify a deterministic visual invariant indicating milestone completion. Format output utilizing tuple structure 𝑔𝑘 , 𝐸𝑘 . Cloud Replan Prompt role: You are a High-Level Strategic Planner for Android automation. goal: {goal} failed plan: {old_plan} execution trace: {trace} task: • analyze: read the [Orchestrator Thought] and [Action] in the trace • reflect: summarize why the previous plan failed and what should be changed • replan: generate a new list of high-level semantic milestones from the current screen analysis checklist: • Did the agent fail to find the app? • Did it get stuck on a specific screen? • Did the previous instruction mislead the agent? recovery strategies: • If stuck or lost, start with {"instruction": "Navigate Home", "expectation": "Home screen visible"} • If the app is not found, try a different entry point such as app drawer or app search • If a step is too complex, break it into smaller milestones output format: Reflection: brief analysis of failure and recovery strategy Plan: [ {"instruction": "...", "expectation": "..."}, ... ]
B.3
Edge Orchestrator Prompt
Vision orchestrator executes state alignment evaluating real time screenshot against cloud expectation 𝐸𝑘 . System Formulation: Role dictates visual diagnostician. Input comprises device screenshot 𝑉𝑡 and expected state description 𝐸𝑘 . Evaluate visual alignment generating structured JSON payload. Formulation maps probabilistic state alignment evaluating autoregressive logits corresponding token FINISHED within status field. Detecting alignment failure generate textual meta instruction 𝑠𝑡 detailing semantic target alongside approximate spatial centroid guiding executor regularization. Edge Orchestrator Prompt role: You are a strict Screen State Validator. current sub-goal: “{sub_goal}” success criteria: “{expectation}” recent action history: {history} task: • observe: look at the screenshot and decide whether it matches the success criteria • verify history: check whether an action was actually performed • judge: if the screen matches expectation and history confirms action, output FINISHED; otherwise output ONGOING capabilities note: • The Executor can open app, click, type, scroll, long press, swipe, and use system keys such as back and home
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Edge Orchestrator Prompt • Open app is a special action: if the goal is “Open X app” and the app is not visible, suggest open_app; do not suggest scrolling • On the home screen, scroll down opens the app drawer, while scroll up opens quick settings • By default, do not use swipe for navigation; use it only when the sub-task requires adjusting a slider output format: { "observation": "Brief description of current screen.", "status": "FINISHED" | "ONGOING", "reasoning": "Detailed analysis of why the screen matches or fails the expectation.", "suggestion": "If ONGOING, what is the exact next move?", "spatial_reference": "Approximate [x, y] centroid coordinate guiding textual executor." }
B.4
Edge Executor Prompt
Textual executor maps orchestrator meta instruction 𝑠𝑡 against structured view hierarchy 𝑈𝑡 extracting optimal execution node. System Formulation: Role dictates atomic executor. Input comprises structural view hierarchy and semantic meta instruction. Parse layout tree identifying specific interactable node 𝑢 ∗ fulfilling meta instruction intent utilizing spatial reference 𝑝𝑟𝑒 𝑓 enforcing structural regularization. Output requires strictly valid JSON encapsulating node ID alongside corresponding atomic action type CLICK SWIPE TYPE. Edge Executor Prompt role: You are a precise UI Operator. current sub-goal: “{sub_goal}” analysis from orchestrator agent: {hint_text} spatial reference: {p_ref} screen UI elements: {ui_tree} task: • read: read the orchestrator analysis to understand the target element • search: scan the screen UI element list and find the element whose text or description matches the hint • act: select the best action using the matched element index available actions: {"action_type": "click", "index": <index>} Tap a UI element. {"action_type": "input_text", "text": "...", "index": <index>} Type text into a field. {"action_type": "long_press", "index": <index>} Long press an element when necessary. {"action_type": "swipe", "direction": "<direction>", "index": <target_index>} Perform a continuous physical swipe on the target element. {"action_type": "open_app", "app_name": "..."} Directly launch an app by name. {"action_type": "scroll", "direction": "<direction>"} Use down to scroll down or open the app drawer, and up to scroll up or open quick settings. {"action_type": "navigate_back"} Use the system back button. {"action_type": "navigate_home"} Go to the home screen. Output strictly in JSON format.
C
Extended AndroidWorld Evaluation
Table 10 presents granular success rate decomposition across critical application domains within AndroidWorld extending radar
Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
Trovato et al.
Table 10: Granular Domain Success Rate Breakdown. Evaluates explicit success rate across top five hardest applications within AndroidWorld. Application Domain
AdecPilot Pro
CORE [7]
EcoAgent [27]
Calendar Tasks
41.2%
35.3%
29.4%
Broccoli-Recipe App
38.4%
34.6%
30.7%
Markor Editor
28.5%
25.0%
21.4%
Simple SMS
42.8%
35.7%
28.6%
Settings Configuration
53.3%
46.7%
43.3%
chart visualization. System demonstrates profound capability processing heavily structured data apps including Simple SMS and Settings achieving success rates 42.8% and 53.3% respectively. Superior performance attributes directly integrating bimodal text
executor retaining precise DOM parsing capabilities while visual baselines frequently misinterpret dense textual layouts.