ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Fei Tang∗ , Zhiqiong Lu∗ , Boxuan Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen† Zhejiang University
Abstract. GUI agents drive applications through their visual interfaces instead of programmatic APIs, interacting with arbitrary software via taps, swipes, and keystrokes, reaching a long tail of applications that CLI-based agents cannot. Yet progress in this area is bottlenecked less by modeling capacity than by the absence of a coherent full-stack infrastructure: online RL training suffers from environment instability and closed pipelines, evaluation protocols drift silently across works, and trained agents rarely reach real users on real devices. We present ClawGUI, an open-source framework addressing these three gaps within a single harness. ClawGUI-RL provides the first opensource GUI agent RL infrastructure with validated support for both parallel virtual environments and real physical devices, integrating GiGPO with a Process Reward Model for dense step-level supervision. ClawGUI-Eval enforces a fully standardized evaluation pipeline across 6 benchmarks and 11+ models, achieving 95.8% reproduction against official baselines. ClawGUI-Agent brings trained agents to Android, HarmonyOS, and iOS through 12+ chat platforms with hybrid CLI-GUI control and persistent personalized memory. Trained end to end within this pipeline, ClawGUI-2B achieves 17.1% Success Rate on MobileWorld GUI-Only, outperforming the same-scale MAI-UI-2B baseline by 6.0%. ClawGUI-RL RL Trainer
Multi-Env Parallel
Reward Manager
Real & Virtual Environment
ClawGUI-Agent
ClawGUI-Eval Infer
GRPO/GiGPO
Judge
Screenspot-Pro PRM / MLLM judger
MMBench-GUI
Build
…
Metric UI-Vision
Chat Apps
Agent Loop
Personalized Memory
Hybrid Operation
Deployment
Evaluate
…
Device Control
Real Devices
…
Multi-Model Support
AndroidControl
Message Virtual Devices
Result
Agent Loop
Perceive
OpenClaw Control
Rollout Manager
Rollout
arXiv:2604.11784v1 [cs.LG] 13 Apr 2026
{flysugar, syl}@zju.edu.cn § https://github.com/zju-real/ClawGUI https://zju-real.github.io/ClawGUI-Page
Reason
Reads screen
Plan / Reflex
Act TAP / TYPE / SCROLL
Result
Message sent · execution result returned
Figure 1 | Overview of ClawGUI, a unified open-source framework for GUI agent research and deployment. It integrates scalable online RL training (ClawGUI-RL) with outcome-based rewards and no human annotation, reproducible three-stage evaluation across 6 benchmarks and 11+ models (ClawGUI-Eval), and real-device deployment across Android, HarmonyOS, and iOS through 12+ chat platforms (ClawGUI-Agent).
∗ Core contribution;
† Corresponding authors;
Main contact: [email protected]
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Contents 1
Introduction
3
2
Related Work 2.1 GUI Agent Models: From Grounding to Navigation . . . . . . . . . . . . . . . . . . . . . . . . 2.2 Online Reinforcement Learning for GUI Agents . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3 Benchmarking and Reproducibility of GUI Agents . . . . . . . . . . . . . . . . . . . . . . . . . 2.4 Deploying GUI Agents to Real Users . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4 4 5 5 5
3
ClawGUI 3.1 System Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2 ClawGUI-RL: Scalable Online RL Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.1 Environment Manager . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.2 Reward Design: Binary Reward and Dense Reward . . . . . . . . . . . . . . . . . . . 3.2.3 RL Trainer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3 ClawGUI-Eval: Reproducible GUI Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.1 Benchmark and Model Coverage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.2 Pipeline Architecture: Infer, Judge, and Metric . . . . . . . . . . . . . . . . . . . . . . 3.4 ClawGUI-Agent: Personal GUI Assistant . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.1 Hybrid Device Control: Operating via CLI and GUI . . . . . . . . . . . . . . . . . . . 3.4.2 Personalized Memory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.3 Remote and Local Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.4 ClawGUI-Eval as a Deployable Skill . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6 6 7 7 7 8 8 9 9 10 10 11 11 11
4
Experiments 4.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.2 Main Results. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3 Every Step Counts: Dense Reward Unlocks Better GUI Policies . . . . . . . . . . . . . . . . . . 4.4 Benchmarking the Benchmarks: Can We Trust Published GUI Numbers? . . . . . . . . . . . .
11 11 12 13 13
5
Discussion
14
6
Conclusion
15
2
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
1. Introduction Graphical User Interfaces (GUIs) are the universal substrate through which humans interact with modern computing devices (Fu et al., 2023; Hong et al., 2024; Nakano et al., 2022; Shen et al., 2023; Tang et al., 2025b,c). An agent that can perceive screen state and execute low-level interface actions such as tapping, swiping, and typing is, in principle, capable of operating any application on any device without requiring dedicated APIs or backend access (Hu et al., 2024; Lai et al., 2024; Zhang et al., 2025). This generality has made GUI agents one of the most actively pursued directions toward end-to-end digital automation, with rapid progress over the past two years across grounding, navigation, and online reinforcement learning (RL) (Qin et al., 2025; Shi et al., 2025). Building a capable GUI agent, however, is not a single modeling problem but a full-stack engineering problem (Lu et al., 2025; Luo et al., 2025b; Tang et al., 2025a; Ye et al., 2025). A useful agent must be trained against realistic environments, evaluated under comparable conditions, and ultimately deployed to real devices where real users can benefit from it (Liu et al., 2024; Shi et al., 2025; Wang et al., 2025a). Existing research has made meaningful progress on each of these fronts in isolation: GUI grounding models have steadily improved element localization accuracy (Liu et al., 2025; Lu et al., 2025; Luo et al., 2025b; Tang et al., 2025a), navigation agents have extended task horizons (Gu et al., 2025; Qin et al., 2025), and online reinforcement learning has begun to push policy quality beyond what static supervision alone can achieve (Team et al., 2026). Yet when one attempts to assemble these pieces into a working pipeline, the cracks between them become apparent. The community still lacks a unified framework in which training, evaluation, and deployment operate as a coherent whole, and this absence, rather than any single missing technique, is what currently bottlenecks practical progress. We identify three concrete gaps that together define this bottleneck. The training ecosystem for GUI agents remains largely closed. Several recent systems report strong results from online RL training in virtual environments (Team et al., 2026; Wang et al., 2025a; Xu et al., 2026), but none release the underlying infrastructure, leaving outside researchers unable to reproduce the setup or build upon it. Even where code exists, it is tied exclusively to emulator-based sandboxes, and training directly on physical devices, which is ultimately where agents must perform, remains essentially unexplored in the open literature (Zhou et al., 2025). The central engineering difficulty here is not the RL algorithm itself but environment management: emulators drift out of healthy states during long runs, real devices cannot expose system-level verification signals, and reward signals in long-horizon GUI tasks are sparse almost by construction. Evaluation across GUI agent papers is badly misaligned. GUI benchmarks appear straightforward on the surface, yet reported numbers across papers are rarely directly comparable (Seed, 2026; Team et al., 2023). Prompt formatting, coordinate normalization conventions, image resolution, and sampling configuration each shift reported accuracy by several points, and these choices are often undocumented. The result is that the community has no shared baseline against which to measure true progress. A 2% improvement on ScreenSpot-Pro (Li et al., 2025) may reflect a genuine advance, a favorable prompt, or simply a different resolution, and there is currently no way for a reader to tell. The deployment loop from research to real users is broken. Agents trained in research pipelines almost never reach end users. A recent line of work has explored CLI-based agent harnesses (Anthropic, 2025; HKUDS, 2026; Steinberger and OpenClaw Contributors, 2026; Wener, 2026), which offer precise control through structured commands but cover only a narrow slice of real applications. Meanwhile, systems that connect a trained GUI policy to real hardware, expose it through interfaces users already use in daily life, and maintain persistent personalization over time remain largely absent from the open ecosystem (Agashe et al., 2024; Wang et al., 2024a; Zhang et al., 2023). Without this 3
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
final link, the real-world value of GUI agents goes largely unverified. Building on these insights, we present ClawGUI, an open-source framework designed to close all three gaps within a single coherent system. ClawGUI consists of three tightly integrated modules. ClawGUI-RL provides scalable online RL infrastructure with validated support for both Docker-based parallel Android emulators and real physical devices, integrating GiGPO (Feng et al., 2025b) together with a Process Reward Model that supplies dense step-level supervision to counteract the sparsity of outcome rewards in long-horizon GUI tasks. ClawGUI-Eval enforces a strict three-stage pipeline across 6 benchmarks and 11+ models, pinning every evaluation choice per model so that results become reproducible rather than nominally comparable. ClawGUI-Agent closes the loop from research to deployment, bringing trained agents to Android, HarmonyOS, and iOS through 12+ chat platforms, with a hybrid CLI-GUI control strategy that combines the precision of CLI with the universal coverage of GUI, and a persistent personalized memory system that allows the agent to adapt to individual users over time. To validate the framework end to end, we train ClawGUI-2B entirely within the ClawGUI-RL pipeline. On MobileWorld GUI-Only, ClawGUI-2B achieves a Success Rate (SR) of 17.1%, compared to 11.1% for the same-scale MAI-UI-2B baseline, and also surpasses substantially larger untrained models such as Qwen3-VL-32B (11.9%) and UI-Venus-72B (16.4%). Within the same pipeline, replacing episode-level GRPO with step-level GiGPO yields a 2.6% improvement (14.5% → 17.1%), directly confirming the value of dense credit assignment in GUI RL. On the evaluation side, ClawGUI-Eval achieves a 95.8% reproduction rate against published baselines across 6 benchmarks and 11+ models. Our main contributions are as follows: • We release ClawGUI, a unified open-source framework that integrates online RL training, standardized evaluation, and real-device deployment into a single pipeline for GUI agents. • We release ClawGUI-RL, the first open-source GUI agent RL infrastructure with validated support for both large-scale parallel virtual environments and real physical devices, integrating GiGPO with a Process Reward Model for dense step-level supervision. • We release ClawGUI-Eval together with all inference code and pre-computed predictions across 6 benchmarks and 11+ models, achieving a 95.8% reproduction rate against official baselines and enabling reliable cross-paper comparison. • We release ClawGUI-Agent, a production-ready deployment system that connects trained agents to real Android, HarmonyOS, and iOS devices through 12+ chat platforms with persistent personalized memory. • We release ClawGUI-2B, trained end to end within ClawGUI-RL, which reaches 17.1% on MobileWorld GUI-Only and outperforms the same-scale MAI-UI-2B baseline by 6.0 absolute points, validating the framework end to end.
2. Related Work 2.1. GUI Agent Models: From Grounding to Navigation Early GUI agents relied on cascaded pipelines that combined off-the-shelf perception modules such as OCR (Du et al., 2025; Feng et al., 2025a, 2026; Wang et al., 2024b), SAM (Kirillov et al., 2023; Ravi et al., 2024), and set-of-marks prompting (Lu et al., 2024; Yang et al., 2023) with a closedsource planner, a modular but error-accumulating design that precluded end-to-end optimization. As vision-language foundation models matured, end-to-end grounding became the dominant paradigm: SeeClick (Cheng et al., 2024), UI-TARS (Qin et al., 2025), Aguvis (Xu et al., 2024), and UGround (Gou et al., 2024) showed that localization accuracy scales with data and model capacity, and subsequent 4
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
work sharpened grounding further via RL-based coordinate rewards (Lu et al., 2025; Luo et al., 2025b; Tang et al., 2025a). On top of stronger grounding, a second wave targets long-horizon navigation, splitting into modular pipelines that pair a grounding model with a proprietary planner (Gou et al., 2024; Xu et al., 2024) and unified end-to-end policies that internalize perception and decision-making jointly (Qin et al., 2025; Team et al., 2026; Zhou et al., 2025). ClawGUI is orthogonal to this modeling axis: rather than proposing a new grounding or navigation model, it provides a shared harness in which both paradigms can be trained, evaluated, and deployed under consistent conditions. 2.2. Online Reinforcement Learning for GUI Agents Collecting large-scale trajectory data for long-horizon GUI tasks is expensive, as each demonstration requires step-by-step execution, precise action annotation, and faithful environment replay (Kong et al., 2025; Lin et al., 2025; Wang et al., 2026). Online reinforcement learning offers an attractive alternative, letting the agent generate its own experience through direct environment interaction and optimize toward task success via outcome rewards. A rapidly growing line of work including MobileGUI-RL (Shi et al., 2025), ComputerRL (Lai et al., 2025), MAI-UI (Zhou et al., 2025), UIVenus-1.5 (Team et al., 2026), and UI-TARS-2 (Wang et al., 2025a) has shown that sandbox-based online training yields consistent gains beyond SFT. Yet three difficulties persist: reward signals are sparse over long action sequences, multi-step credit assignment is nontrivial, and infrastructure cost is substantial, demanding parallel simulation and robust episode management across heterogeneous applications. More importantly for the community, none of these works open-source their training infrastructure, and all are validated solely in virtual sandboxes, leaving real-device training almost entirely unexplored. ClawGUI-RL directly targets this gap by releasing an open-source infrastructure that handles environment management, dense step-level reward supervision, and validated training on both parallel emulators and real physical devices. 2.3. Benchmarking and Reproducibility of GUI Agents A rich ecosystem of GUI benchmarks has emerged to measure grounding and navigation, including ScreenSpot-Pro (Li et al., 2025), ScreenSpot-V2, UI-Vision, MMBench-GUI, OSWorld-G, and AndroidControl, alongside interactive suites such as MobileWorld (Kong et al., 2025). These have become the de facto yardsticks for reporting GUI agent progress. In practice, however, reported numbers across papers are rarely directly comparable: prompt formatting, coordinate normalization, image resolution, sampling temperature, and post-processing rules interact in ways that shift reported accuracy by several points, and many of these choices are undocumented (Seed, 2026; Tang et al., 2025b; Team et al., 2023). As a result, the community has no reliable shared baseline, and small reported improvements are often indistinguishable from configuration drift. Prior standardization efforts have focused on a single benchmark, been bundled with a particular training recipe, or released evaluation scripts without the inference predictions, making independent re-judging infeasible. ClawGUI-Eval differs in both scope and philosophy: it decouples evaluation into standardized inference, judging, and metric computation, pins all configuration choices per model, and releases inference outputs across 6 benchmarks and 11+ models, enabling the community to reproduce, re-judge, and extend published results without re-running expensive inference. 2.4. Deploying GUI Agents to Real Users A capable GUI agent delivers value only when it reaches real users operating real devices. Driven by the success of OpenClaw (Steinberger and OpenClaw Contributors, 2026) and Hermes-Agent (organization, 2026), a growing body of work has turned toward CLI-based agent harnesses (Anthropic, 2025;
5
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
RL Infrastructure RL Trainer
Environment Manager
Multi-turn Online Rollout
Task 1
Real & Virtual Environment Real & Virtual API Server
Task 2
Task 3
manager_Logo.png
CPU Rollout Workers
Task N
Task Evaluation
GUI Agent Loop GUI Agent Loop
Get Reward
Exec Actions
Get Screenshot Policy Update
Eval Controller
ADB Controller
GUI Agent Loop
Allocate Parallel Envs
Text Answer Verification Database Verfication manager_Logo.png
GUI Agent Loop
Local Storage Inspection
Env-1 Application Callbacks
Env-2 Env-3
… Env-N
Environment Loop Ø Reset/restart environments
Real Device
Android Emulator
System Judge
Devices
Ø Health check and crash recovery Ø Spare Server Rotation Logic Ø Real Device Training Ø Multi-Environment Parallelism
MLLM Judge
APP Database
ry Que Emulator Judge
Real Device Judge
Figure 2 | Overview of ClawGUI-RL, consisting of an RL Infrastructure and a Real & Virtual Environment backend. The Environment Manager orchestrates multi-task parallel rollouts across real devices and Android emulators, with built-in health checking, crash recovery, and spare server rotation. Task evaluation combines system-level verification with MLLM-as-judge to provide robust reward signals for both virtual and real device training. HKUDS, 2026; organization, 2026; Rajasekaran and Engineering, 2026; Team, 2026; Wener, 2026) as a deployment substrate. CLI execution is efficient and precise, yet carries fundamental limitations: many applications expose no programmatic interface (Zhang et al., 2025), CLI operations are opaque to users who cannot observe or intervene in agent behavior (Zhao et al., 2025), and bypassing the visual layer forfeits the spatial grounding that makes agent actions interpretable (Fu et al., 2025; Tang et al., 2025c). GUI-based interaction addresses these limitations by operating directly on the screen and covering any application regardless of its underlying architecture, but introduces its own cost, since tasks resolved by CLI in a single call may require several sequential GUI actions (Jiang et al., 2025; Mozannar et al., 2025). Existing research deployments also tend to stop at demo notebooks or isolated Android controllers, leaving cross-platform coverage and persistent personalization largely unaddressed. ClawGUI-Agent closes this gap through a hybrid harness that leverages CLI efficiency where interfaces permit and falls back to GUI control where they do not, connects trained agents to Android, HarmonyOS, and iOS through 12+ chat platforms, and integrates a persistent personalized memory system that enables the agent to adapt to individual users over time.
3. ClawGUI 3.1. System Overview We introduce ClawGUI, a unified framework designed to cover the complete lifecycle of GUI agent development. As illustrated in Figure 1, ClawGUI consists of three tightly integrated modules: ClawGUI-RL for scalable online RL training, ClawGUI-Eval for standardized and reproducible evaluation, and ClawGUI-Agent for real-device deployment and human interaction.
6
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
3.2. ClawGUI-RL: Scalable Online RL Training GUI tasks are inherently sequential decision-making problems that require agents to learn through real environment interaction rather than static supervision alone. Despite growing interest in online RL for GUI agents, the community lacks an open-source infrastructure that is both scalable and validated on real physical devices. To this end, we build ClawGUI-RL to provide end-to-end support from environment management and reward design to policy optimization. 3.2.1. Environment Manager Stable and scalable environment management is a prerequisite for online RL training on GUI tasks. As shown in Figure 2, ClawGUI-RL abstracts all device backends behind a unified interface, allowing virtual environments and physical devices to be used interchangeably within the same training loop. Virtual Environment. ClawGUI-RL launches dozens of Docker-based Android emulators in parallel via MobileWorld (Kong et al., 2025), each exposing a backend URL that training workers interact with. Each environment follows a four-stage lifecycle: • Task Reset. At the beginning of each episode, the environment initializes the device state and loads a new task, ensuring a clean starting condition for every rollout. • Task Evaluation. Virtual environments expose system-level root access, enabling reliable task completion verification through direct inspection of app state and database records. This systemlevel signal is further complemented by an MLLM-as-judge that assesses the final screen state against the task instruction, providing a robust and comprehensive outcome reward. • Spare Server Rotation. Virtual sandbox environments are prone to becoming unhealthy during long training runs — a stalled or crashed container introduces training instability and can cause irrecoverable errors. To address this, ClawGUI-RL maintains a spare server queue. When a container is detected as unhealthy, the system automatically draws from the queue and rotates to a healthy replacement, allowing the affected task to resume without interrupting the training process. • Teardown. Containers are periodically restarted to prevent state accumulation and maintain environment fidelity across long training runs. Real Device Training. ClawGUI-RL supports training directly on physical Android devices or cloud phones through the same unified interface. Real-device training introduces two challenges that do not arise in virtual environments. • Task Source. Unlike virtual environments where tasks can be procedurally generated and automatically verified, real-device tasks must be manually authored to ensure they are both executable and verifiable on physical hardware. In ClawGUI-RL, we curate a set of humanauthored tasks covering representative real-world scenarios. • Task Evaluation. Physical devices do not expose system-level root access, making automated state verification infeasible. ClawGUI-RL therefore relies on MLLM-as-judge to assess task completion by evaluating the final screen state against the task instruction, providing a practical reward signal without requiring device-level privileges. 3.2.2. Reward Design: Binary Reward and Dense Reward Reward design is critical for online RL training on long-horizon GUI tasks. ClawGUI-RL adopts a two-level reward formulation combining a binary outcome reward with a dense process reward. 7
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Binary Outcome Reward. The primary reward signal is a binary score assigned at episode end: 1 for task success, 0 for failure. While straightforward, this signal suffers from a fundamental limitation in GUI environments — execution latency and multi-step interaction introduce significant delay between an action and its observable consequence, resulting in an extremely sparse reward signal that provides little guidance for intermediate steps. Dense Step-Level Reward via PRM. To complement the sparse outcome reward, ClawGUI-RL integrates a Process Reward Model (PRM). After each action, the PRM receives the previous screenshot, the current screenshot, and the full history of actions taken so far, and judges whether the current action meaningfully contributes to task completion. This produces a per-step score that is combined with the outcome reward: 𝑅 = 𝑅outcome + 𝑅step (1) By providing dense feedback at every step, the PRM substantially alleviates the sparsity problem and enables the optimizer to distinguish productive actions from dead ends throughout the episode. 3.2.3. RL Trainer ClawGUI-RL builds upon verl (Sheng et al., 2025) and verl-agent (Feng et al., 2025b), with out-of-thebox support for a suite of RL algorithms including Reinforce++ (Hu et al., 2025), PPO (Schulman et al., 2017), GSPO (Zheng et al., 2025), GRPO (Shao et al., 2024b), and GiGPO (Feng et al., 2025b). In our experiments, we integrate GRPO and GiGPO as the primary advantage estimation algorithms and analyze their impact on GUI agent training. GRPO (Shao et al., 2024a) estimates advantages by normalizing returns within a group of rollouts that share the same task. While straightforward and effective for single-turn tasks, GRPO assigns a uniform episode-level advantage to every step within a trajectory, which is too coarse for long-horizon GUI interaction. Consider two rollouts on the same task: rollout 𝐴 completes it in 4 steps, while rollout 𝐵 completes it in 8 steps. GRPO assigns both trajectories the same reward, providing no signal to distinguish the efficiency of individual steps or to credit productive actions over redundant ones. GiGPO (Feng et al., 2025b) addresses this limitation through a two-level hierarchical advantage estimation. At the episode level, GiGPO retains the macro relative advantage across complete trajectories, preserving global trajectory quality signals. At the step level, GiGPO introduces an anchor-state grouping mechanism: steps that encounter the same intermediate environment state across different rollouts are retroactively clustered into sub-groups, and micro relative advantages are estimated within each sub-group via discounted return normalization. This hierarchical structure yields fine-grained per-step credit assignment that captures both global trajectory quality and local step effectiveness, without requiring a learned value network or additional rollouts, making it particularly well-suited to the multi-step nature of GUI tasks. 3.3. ClawGUI-Eval: Reproducible GUI Evaluation Evaluation is the compass of research progress, yet GUI evaluation is harder to reproduce than it appears. Prompt ordering, coordinate normalization conventions, image resolution, and sampling temperature interact in ways that shift reported accuracy by several points across implementations. As a result, numbers across papers are rarely comparable, and the community has no reliable baseline against which to measure true progress. As shown in Figure 3, ClawGUI-Eval addresses this by pinning all evaluation choices per model and adopting a strict three-stage pipeline, achieving a 95.8% reproduction rate against official results across 6 benchmarks and 11+ models.
8
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Reproduced Results
ClawGUI-Eval: Official vs. Reproduced Results
results analysis
message to chat apps
Support Models Gemini
Qwen3-VL
…
auto evaluation UI-Venus
Claude
MAI-UI
GUI-G2
ClawGUI-Eval Benchmark
Infer
Judge
Metric
Screenspot-Pro
Transformers
Grounding
Success
Navigation
Pass@1
…
AndroidControl
…
API
…
…
Figure 3 | Overview of ClawGUI-Eval, featuring a standardized Infer → Judge → Metric pipeline across 6 benchmarks and 11+ models. Reproduced results are compared against official baselines across five benchmarks, achieving a 95.8% overall reproduction rate. The full pipeline can be triggered via a single natural language command through ClawGUI-Agent (OpenClaw-GUI). 3.3.1. Benchmark and Model Coverage ClawGUI-Eval covers 6 benchmarks spanning diverse GUI grounding and navigation scenarios: ScreenSpot-Pro (Li et al., 2025), ScreenSpot-V2 (Wu et al., 2024), UI-Vision (Nayak et al., 2025), MMBench-GUI (Wang et al., 2025b), OSWorld-G (Xie et al., 2025), and AndroidControl (Li et al., 2024). On the model side, it supports 11+ models including Qwen3-VL (Bai et al., 2025a), Qwen2.5VL (Bai et al., 2025b), UI-TARS (Qin et al., 2025), MAI-UI (Zhou et al., 2025), GUI-G2 (Tang et al., 2025a), UI-Venus (Gu et al., 2025), GUI-Owl (Ye et al., 2025), StepGUI (Yan et al., 2025), Gemini (Team et al., 2023), and Seed 1.8 (Seed, 2026). All inference results are publicly released alongside the evaluation code, enabling the community to reproduce, extend, and build upon our results directly. 3.3.2. Pipeline Architecture: Infer, Judge, and Metric ClawGUI-Eval decomposes evaluation into three decoupled stages, each with a clearly defined input and output. • Infer. Given a benchmark dataset and a target model, the inference stage generates raw predictions. ClawGUI-Eval supports two backends: local GPU inference via transformers, and remote API inference via any OpenAI-compatible endpoint. Multi-GPU parallel inference is handled automatically through Python multiprocessing, with each process pinned to a dedicated GPU. Shard-level checkpointing allows interrupted runs to resume from the last 9
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
User Side
Server Side - OpenClaw-GUI Framework PhoneAgent
Response
Message
Agent Loop Context LLM
Memory
Tools
Skills
Controlled Device Real Devices
Virtual Devices
… Phone
Web browser
Desktop
Figure 4 | Overview of ClawGUI-Agent, where users issue natural language instructions through 12+ chat platforms and the server side executes tasks via a message-driven agent loop with persistent memory and skills, controlling both virtual and real devices across phone, web browser, and desktop. completed shard without recomputation. • Judge. Raw model outputs are parsed and evaluated against ground truth. ClawGUI-Eval implements benchmark-specific judges: a point-in-box judge for standard GUI grounding benchmarks, a polygon and refusal-aware judge for OSWorld-G, and a multi-action judge for AndroidControl. Each judge produces a per-sample correctness label. • Metric. Per-sample labels are aggregated into final accuracy scores with breakdowns by platform, UI element type, and task category, enabling fine-grained analysis beyond top-line numbers. By decoupling these three stages, ClawGUI-Eval allows any single stage to be rerun independently — for instance, re-judging existing predictions with an updated parser without repeating expensive inference. 3.4. ClawGUI-Agent: Personal GUI Assistant As GUI agents grow more capable, the final challenge is closing the loop to real users. A trained agent that cannot be deployed, personalized, or integrated into daily workflows delivers no practical value. As shown in Figure 4, ClawGUI-Agent is designed to bridge this gap, providing a production-ready system that brings GUI agents into the hands of real users across real devices. 3.4.1. Hybrid Device Control: Operating via CLI and GUI Driven by the success of OpenClaw (Steinberger and OpenClaw Contributors, 2026), a growing body of work has focused on CLI-based agent control. CLI interaction is precise and efficient: a 10
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
single structured command can accomplish in one step what would otherwise require navigating multiple UI layers. However, CLI control carries fundamental limitations. Not all applications expose programmatic interfaces. CLI operations are opaque to users who cannot observe or intervene. It also bypasses the visual layer that makes agent behavior interpretable. GUI interaction addresses these limitations by operating directly on the screen, covering any application regardless of its underlying architecture, but introduces its own cost: tasks resolved by CLI in a single call may require multiple sequential GUI actions. We argue that neither paradigm alone is sufficient. ClawGUI-Agent adopts a hybrid approach that leverages CLI efficiency where interfaces permit, and falls back to GUI control where they do not. This combination preserves the speed of CLI for well-supported operations while ensuring broad coverage through GUI for everything else. 3.4.2. Personalized Memory ClawGUI-Agent incorporates a persistent personalized memory system. During task execution, the agent automatically extracts structured facts from interactions, including contact names and relationships, frequently used applications, and user habits and preferences, and stores them as vector embeddings in a persistent store. On subsequent tasks, the top-𝑘 most semantically similar memories are retrieved and injected into the system context, allowing the agent to recognize recurring entities and adapt to individual user patterns over time. Duplicate memories are detected and merged rather than accumulated, keeping the memory store lean and relevant. 3.4.3. Remote and Local Control ClawGUI-Agent supports two deployment modes. In remote control mode, the agent is accessed through 12+ chat platforms including Feishu, DingTalk, Telegram, Discord, Slack, and QQ, allowing users to issue tasks from a separate device to control the target phone remotely. In local control mode, users send instructions directly from a chat application running on the phone itself, upon which the agent takes over the local device in place, requiring no additional hardware or cloud relay. 3.4.4. ClawGUI-Eval as a Deployable Skill ClawGUI-Agent exposes ClawGUI-Eval as a built-in tool skill, enabling users to trigger a complete benchmark evaluation pipeline through a single natural-language command. Upon receiving an instruction such as “benchmark Qwen3-VL on ScreenSpot-Pro”, the agent automatically performs environment verification, launches multi-GPU parallel inference, runs the judge, computes metrics, and returns a structured result report with comparisons against official baselines, without writing a single script.
4. Experiments 4.1. Setting Training. We train ClawGUI-2B based on MAI-UI-2B (Zhou et al., 2025) using 64 parallel virtual environments on 8×A6000 (48GB) GPUs. We adopt the GiGPO algorithm with a rollout group size of 8, sampling temperature of 0.7, and learning rate of 1e-6 for 3 epochs with a training batch size of 8. For PRM-based step-level reward judgment, we employ Qwen3.5-72B as the judge model. Evaluation. We evaluate ClawGUI-2B on MobileWorld, an online interactive benchmark designed to assess the end-to-end task completion capability of GUI agents. MobileWorld comprises three task 11
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Model
MobileWorld SR (GUI-Only)
Agentic Framework Claude-4.5-Sonnet + UI-Ins-7B Gemini-3-Pro + UI-Ins-7B GPT-5 + UI-Ins-7B
47.8 55.6 54.0
End-to-End Model GUI-Owl-7B GUI-Owl-32B UI-Venus-7B UI-Venus-72B Qwen3-VL-8B Qwen3-VL-32B Qwen3-VL-235B-A22B Doubao-1.5-UI-TARS MAI-UI-2B MAI-UI-8B
7.7 8.5 8.5 16.4 9.4 11.9 12.8 26.3 11.1 19.7
Ours ClawGUI-2B
17.1
Table 1 | Comparison of models on GUI-Only (117 tasks) benchmark. categories: GUI-Only, MCP, and Call-User. We focus our evaluation on the GUI-Only split, which contains 117 tasks that require agents to complete real-world mobile interactions purely through visual GUI control, without any programmatic interface access. Task completion is measured by Success Rate, defined as whether the agent successfully accomplishes the task objective by the end of the episode. We set the maximum number of interaction steps to 50 for all evaluations. Method
Reward Type
SR (%)
GRPO GiGPO
Binary (episode-level) Dense (episode- & step-level)
14.5 17.1
Table 2 | Ablation on reward design on MobileWorld GUI-Only (117 tasks).
4.2. Main Results. Table 1 reports Success Rate on the MobileWorld GUI-Only benchmark. We highlight three key observations. (1) Infrastructure drives policy quality. ClawGUI-2B achieves 17.1% SR, surpassing the same-scale MAI-UI-2B baseline by a relative margin of 6.0%, which directly validates the effectiveness of our open-source RL infrastructure. Both models share the same base weights; the gain stems entirely from ClawGUI-RL’s scalable environment management and reward design. (2) Small well-trained models outperform larger untrained ones. ClawGUI-2B outperforms substantially larger end-to-end models, including Qwen3-VL-32B (11.9%) and UI-Venus-72B (16.4%), demonstrating that online RL training through real environment interaction contributes more to task completion capability than model scale alone. (3) Agentic frameworks remain a separate regime. Methods that couple proprietary frontier models with dedicated grounding modules achieve higher absolute numbers (e.g., Gemini-3-Pro + UI-Ins-7B at 55.6%), but rely on closed-source planners unavailable for end-to-end optimization. These systems are not directly comparable to compact trained agents and 12
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
represent a complementary rather than competing paradigm. Taken together, these results confirm that a well-engineered open-source training infrastructure can unlock strong GUI agent performance at modest model scale, closing a significant gap with much larger systems. 4.3. Every Step Counts: Dense Reward Unlocks Better GUI Policies GRPO assigns a single advantage score to an entire episode, which is too coarse for long-horizon GUI tasks where intermediate steps vary significantly in quality. A misclick early in a trajectory receives the same credit as a decisive final action, providing little signal for the optimizer to distinguish productive steps from dead ends. GiGPO addresses this through anchor-state grouping: steps that encounter the same intermediate state are clustered into sub-groups, and per-step advantages are estimated via discounted return normalization within each sub-group. This yields fine-grained step-level credit assignment without requiring a learned value network, making it particularly well-suited to the multi-step nature of GUI interaction. As shown in Table 2, replacing GRPO with GiGPO yields a 2.6% improvement in SR (14.5% → 17.1%) on the MobileWorld GUI-Only tasks, a relative gain of 17.9%. This consistent improvement demonstrates that dense step-level supervision provides a substantially richer learning signal than episode-level reward alone, and that accurate credit assignment is a critical factor in GUI agent training. Takeaway Dense step-level reward supervision via GiGPO yields a 17.9% relative gain over episode-level GRPO, confirming that fine-grained credit assignment is a critical factor in GUI agent RL training.
4.4. Benchmarking the Benchmarks: Can We Trust Published GUI Numbers? Reproducibility is a prerequisite for meaningful progress, yet GUI evaluation is notoriously difficult to reproduce. Prompt ordering, coordinate normalization conventions, image resolution, and sampling temperature interact in ways that can shift reported accuracy by several points across implementations. As a result, numbers across papers are rarely directly comparable, and the community has lacked a reliable common baseline against which to measure true progress. ClawGUI-Eval addresses this by pinning all evaluation choices per model and adopting a strict threestage Infer → Judge → Metric pipeline across 6 benchmarks and 11+ models. Table 3 reports our reproduced results against officially published numbers. A result is considered successfully reproduced (✓) if the reproduced value meets or exceeds the official number, or the absolute difference is ≤ 2%. As shown in Table 3, we highlight three observations. First, ClawGUI-Eval achieves an overall reproduction rate of 95.8% (46/48 cells with official baselines), with open-source models reaching 95.7% and frontier models reaching 100% on ScreenSpot-Pro. Second, the two failure cases (Qwen3VL-2B and UI-TARS 1.5-7B on SS-Pro) both involve models whose official evaluation configurations have not been publicly disclosed, suggesting that undisclosed prompt or resolution choices are the primary driver of irreproducibility in the field. Third, for closed-source frontier models where standard inference is infeasible (Liangyu Chen and Hanzhang Zhou and Quyu Kong and Xu Zhang and Wenxuan Wang and Qin Jin and Yue Wang, 2026), we adopt a Zoom paradigm, a two-stage crop-then-ground strategy that applies 25% crop tiles for Gemini and 50% crop tiles for Seed, which successfully
13
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Model
SS-Pro Off.
SS-Pro Ours
SS-V2 Off.
SS-V2 Ours
UIV Off.
UIV Ours
MMB Off.
MMB Ours
OSW-G Off.
OSW-G Ours
47.50 57.80 66.80 71.10 48.50 59.50 54.60 49.60 50.80 57.70 68.40 57.40 65.80 60.00
47.75 56.36 66.16 70.08 43.90 59.39 56.42 15.62 27.45 42.06 50.47 58.82 67.68 57.94 64.07 59.14
93.30 89.70 93.20 93.70 94.10 92.80 95.90 92.50 95.20 93.60
93.32 89.23 92.53 93.55 88.92 93.08 94.26 64.86 87.66 89.54 94.03 93.24 95.83 92.30 94.34 91.98
26.50 44.80 46.50 30.30 40.70 -
25.99 23.71 29.97 36.70 15.06 27.78 27.96 6.73 14.40 20.30 26.52 43.82 45.88 29.68 40.23 29.90
72.17 83.24 82.52 80.30 88.10 82.60 88.80 84.00
79.33 71.54 82.94 82.33 73.12 84.28 84.25 52.81 70.26 73.23 80.08 81.19 87.79 82.80 88.81 83.03
52.80 63.70 65.80 58.80 59.40 69.70 52.00 60.10 66.90
58.63 52.04 62.34 64.12 54.12 68.43 65.88 26.08 35.49 58.24 59.41 58.97 69.98 54.17 63.23 65.69
72.70 73.10 -
75.08 72.80 85.01
-
95.68 -
-
-
-
88.4 -
-
-
Open-Source Models GUI-G2 GUI-Owl 1.5-2B GUI-Owl 1.5-4B GUI-Owl 1.5-8B Qwen3-VL-2B Qwen3-VL-4B Qwen3-VL-8B Qwen2.5-VL-3B Qwen2.5-VL-7B UI-TARS 1.5-7B UI-Venus-7B UI-Venus 1.5-2B UI-Venus 1.5-8B MAI-UI-2B MAI-UI-8B StepGUI-4B Closed-Source Models Gemini 3.0 Pro (Zoom) Seed 1.8 (Zoom) Gemini 3.1 Pro (Zoom)
Table 3 | Reproduction results across GUI grounding benchmarks. Green bold indicates successful reproduction (| Δ | ≤ 2% or reproduced ≥ official); red indicates a gap exceeding the threshold; - indicates no official baseline. Benchmark abbreviations: SS-Pro = ScreenSpot-Pro, SS-V2 = ScreenSpot-V2, UIV = UIVision, MMB = MMBench-GUI, OSW-G = OSWorld-G. Closed-source models are evaluated on ScreenSpot-Pro only via a two-stage Zoom paradigm. recovers official performance without any access to model internals. Takeaway A standardized pipeline with pinned evaluation choices achieves 95.8% reproduction rate across 6 benchmarks and 11+ models, demonstrating that GUI evaluation discrepancies are an infrastructure problem, not a fundamental limitation.
5. Discussion Toward a Unified GUI-CLI Agentic Harness. The dominant lesson from the past year of agent engineering, from Claude Code (Anthropic, 2025; Rajasekaran and Engineering, 2026) and Hermes Agent (organization, 2026) to MiniMax M2.7’s self-evolving loop (Team, 2026), is that frontier capability comes as much from the surrounding harness as from the model itself. Permission pipelines, tool dispatch, context compaction, and multi-turn recovery determine whether a capable model becomes a reliable agent or spirals after several steps (Rajasekaran and Engineering, 2026). Yet CLI-based (organization, 2026; Steinberger and OpenClaw Contributors, 2026) and GUI-based (Team et al., 2026; Wang et al., 2025a; Zhou et al., 2025) agents have grown as two parallel ecosystems with almost no shared infrastructure, despite overlapping user goals. Recent hybrid systems such as Hermes Agent’s unified terminal-to-Android gateway (organization, 2026) and MiniMax’s shell-browser-MCP toolchains (Team, 2026) hint at convergence. We view ClawGUI as an early step toward a shared harness standard that treats CLI, GUI, and API calls as interchangeable actions and learns the routing policy itself from interaction data. Scaling Online RL Beyond Emulators. Current GUI agent RL training is confined almost entirely to emulator sandboxes (Kong et al., 2025; Shi et al., 2025; Team et al., 2026; Wang et al., 2025a), which drift from real app behavior and cannot cover the authenticated long tail of commercial applications. Two complementary directions are emerging. First, mock applications reconstructed by modern 14
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
code-generation models (Anthropic, 2025; Team, 2026) offer an authentication-free distribution that mirrors real interaction flows without real user credentials. Second, on-device RL with privacypreserving trajectory collection can tap the enormous pool of real user interaction without centralizing data. Both directions assume an infrastructure that handles environment instability at scale, which is precisely the role ClawGUI-RL is designed to fill. Scaling online RL is now as much a systems problem as an algorithmic one. Toward On-Device, Always-Present System Agents. As on-device inference becomes practical, the final shape of a GUI agent looks less like a remote service invoked on demand and more like a persistent system-level intelligence running locally. Recent efforts such as Hermes Agent’s Android device control (organization, 2026) and Google’s Gemma 4, an effective 2B mobile-first model (Farabet and Lacombe, 2026), together suggest that capable agentic models and ubiquitous device-level deployment are now converging. Such an agent would perceive full device state, retain persistent personalized memory, and execute multi-app workflows autonomously in the background. ClawGUI-Agent’s hybrid CLI-GUI control and personalized memory system are early instances of this pattern, but the full vision requires tighter operating-system integration, on-device policy training, and rigorous local-first privacy guarantees that the community has yet to establish. World Models for GUI Environments. Today’s GUI agents act reactively: observe a screenshot, predict an action, wait for environment feedback. What they lack is an internal model of how the screen will evolve in response to a candidate action, which is precisely what allows humans to plan several steps ahead before committing. Recent progress on general-purpose world models (Luo et al., 2025a; Parker-Holder et al., 2025; Zheng et al., 2026) suggests that learning UI dynamics as a predictive model is now tractable, with training signals drawn from the same screen-action trajectories already collected for imitation and RL. A GUI-specific world model would enable modelbased planning, counterfactual rollouts, and early dead-end detection, turning multi-step interaction from blind trial-and-error into deliberate search. We view ClawGUI-RL’s dense step-level trajectory logging as a natural substrate for training such models at scale.
6. Conclusion We presented ClawGUI, a unified open-source framework that integrates online RL training, standardized evaluation, and real-device deployment into a single coherent pipeline for GUI agent development. ClawGUI-RL provides the first open-source infrastructure with validated support for both large-scale parallel virtual environments and real physical device training, integrating GiGPO with a Process Reward Model for dense step-level reward supervision. ClawGUI-Eval establishes a reproducible evaluation standard across 6 benchmarks and 11+ models, achieving a 95.8% reproduction rate against official baselines. ClawGUI-Agent closes the loop from research to deployment, enabling natural-language-driven automation across Android, HarmonyOS, and iOS through 12+ chat platforms with a personalized memory system. Trained end-to-end within this pipeline, ClawGUI-2B achieves 17.1% MobileWorld SR, outperforming the same-scale baseline by a relative margin of 54% and surpassing models of substantially larger scale. We hope ClawGUI serves as a foundation for the community to build, evaluate, and deploy the next generation of GUI agents.
References S. Agashe, J. Han, S. Gan, J. Yang, A. Li, and X. E. Wang. Agent s: An open agentic framework that uses computers like a human, 2024. URL https://arxiv.org/abs/2410.08164. Anthropic. Claude code overview. https://code.claude.com/docs/en/overview, 2025. 15
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu. Qwen3-vl technical report, 2025a. URL https://arxiv.org/abs/2511.21631. S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin. Qwen2.5-vl technical report, 2025b. URL https://arxiv.org/abs/ 2502.13923. K. Cheng, Q. Sun, Y. Chu, F. Xu, Y. Li, J. Zhang, and Z. Wu. Seeclick: Harnessing gui grounding for advanced visual gui agents, 2024. URL https://arxiv.org/abs/2401.10935. Y. Du, Z. Chen, Y. Xie, W. Bai, H. Feng, W. Shi, Y. Su, C. Huang, and Y.-G. Jiang. Unirec-0.1b: Unified text and formula recognition with 0.1b parameters, 2025. URL https://arxiv.org/abs/2512 .21095. C. Farabet and O. Lacombe. Gemma 4: Byte for byte, the most capable open models. https: //blog.google/innovation-and-ai/technology/developers-tools/gemma-4/, Apr 2026. Google Innovation and AI blog post about the release of Gemma 4, a family of open-weight multimodal AI models under the Apache 2.0 license with advanced reasoning, multimodal capabilities, and support for agentic workflows. H. Feng, S. Wei, X. Fei, W. Shi, Y. Han, L. Liao, J. Lu, B. Wu, Q. Liu, C. Lin, J. Tang, H. Liu, and C. Huang. Dolphin: Document image parsing via heterogeneous anchor prompting. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 21919–21936, Vienna, Austria, July 2025a. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025.findings-acl.1130. URL https://aclantho logy.org/2025.findings-acl.1130/. H. Feng, W. Shi, K. Zhang, X. Fei, L. Liao, D. Yang, Y. Du, X. Wu, J. Tang, Y. Liu, H. Chen, and C. Huang. Dolphin-v2: Universal document parsing via scalable anchor prompting, 2026. URL https://arxiv.org/abs/2602.05384. L. Feng, Z. Xue, T. Liu, and B. An. Group-in-group policy optimization for llm agent training. arXiv preprint arXiv:2505.10978, 2025b. L. Fu, S. Li, Q. Li, L. Deng, F. Li, L. Fan, M. Chen, and X. He. Ufo2: A unified pre-training framework for online and offline speech recognition, 2023. URL https://arxiv.org/abs/2210.14515. T. Fu, A. Su, C. Zhao, H. Wang, M. Wu, Z. Yu, F. Hu, M. Shi, W. Dong, J. Wang, Y. Chen, R. Yu, S. Peng, M. Li, N. Huang, H. Wei, J. Yu, Y. Xin, X. Zhao, K. Gu, P. Jiang, S. Zhou, and S. Wang. Mano technical report, 2025. URL https://arxiv.org/abs/2509.17336. B. Gou, R. Wang, B. Zheng, Y. Xie, C. Chang, Y. Shu, H. Sun, and Y. Su. Navigating the digital world as humans do: Universal visual grounding for gui agents, 2024. URL https://arxiv.org/abs/ 2410.05243. Z. Gu, Z. Zeng, Z. Xu, X. Zhou, S. Shen, Y. Liu, B. Zhou, C. Meng, T. Xia, W. Chen, et al. Ui-venus technical report: Building high-performance ui agents with rft. arXiv preprint arXiv:2508.10833, 2025. 16
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
HKUDS. CLI-Anything: Making ALL software agent-native. https://github.com/HKUDS/CLI-A nything, 2026. W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Zhang, J. Li, B. Xu, Y. Dong, M. Ding, and J. Tang. Cogagent: A visual language model for gui agents, 2024. URL https: //arxiv.org/abs/2312.08914. J. Hu, J. K. Liu, H. Xu, and W. Shen. Reinforce++: Stabilizing critic-free policy optimization with global advantage normalization, 2025. URL https://arxiv.org/abs/2501.03262. S. Hu, M. Ouyang, D. Gao, and M. Z. Shou. The dawn of gui agent: A preliminary case study with claude 3.5 computer use, 2024. URL https://arxiv.org/abs/2411.10323. W. Jiang, Y. Zhuang, C. Song, X. Yang, J. T. Zhou, and C. Zhang. Appagentx: Evolving gui agents as proficient smartphone users. 2025. URL https://arxiv.org/abs/2503.02268. A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick. Segment anything, 2023. URL https://arxiv.org/abs/ 2304.02643. Q. Kong, X. Zhang, Z. Yang, N. Gao, C. Liu, P. Tong, C. Cai, H. Zhou, J. Zhang, L. Chen, et al. Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments. arXiv preprint arXiv:2512.19432, 2025. H. Lai, X. Liu, I. L. Iong, S. Yao, Y. Chen, P. Shen, H. Yu, H. Zhang, X. Zhang, Y. Dong, and J. Tang. Autowebglm: A large language model-based web navigating agent, 2024. URL https://arxiv. org/abs/2404.03648. H. Lai, X. Liu, Y. Zhao, H. Xu, H. Zhang, B. Jing, Y. Ren, S. Yao, Y. Dong, and J. Tang. Computerrl: Scaling end-to-end online reinforcement learning for computer use agents, 2025. URL https: //arxiv.org/abs/2508.14040. K. Li, Z. Meng, H. Lin, Z. Luo, Y. Tian, J. Ma, Z. Huang, and T.-S. Chua. Screenspot-pro: Gui grounding for professional high-resolution computer use, 2025. W. Li, W. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva. On the effects of data scale on ui control agents, 2024. URL https://arxiv.org/abs/2406.03679. Liangyu Chen and Hanzhang Zhou and Quyu Kong and Xu Zhang and Wenxuan Wang and Qin Jin and Yue Wang. Why your AI Agent keeps misclicking? A Practical Grounding Guide for Frontier Models. Blog Post, 2026. URL https://www.notion.so/Why-your-AI-Agent-keeps-mis
clicking-A-Practical-Grounding-Guide-for-Frontier-Models-32630d140ad8808 e895de98994dddb93. M. Lin, M. Liu, T. Lu, L. Yuan, Y. Liu, H. Xu, Y. Miao, Y. Chao, and Z. Li. Gui-rewalk: Massive data generation for gui agent via stochastic exploration and intent-aware reasoning, 2025. URL https://arxiv.org/abs/2509.15738. X. Liu, B. Qin, D. Liang, G. Dong, H. Lai, H. Zhang, H. Zhao, I. L. Iong, J. Sun, J. Wang, J. Gao, J. Shan, K. Liu, S. Zhang, S. Yao, S. Cheng, W. Yao, W. Zhao, X. Liu, X. Liu, X. Chen, X. Yang, Y. Yang, Y. Xu, Y. Yang, Y. Wang, Y. Xu, Z. Qi, Y. Dong, and J. Tang. Autoglm: Autonomous foundation agents for guis. 2024. URL https://arxiv.org/abs/2411.00820.
17
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
Y. Liu, P. Li, C. Xie, X. Hu, X. Han, S. Zhang, H. Yang, and F. Wu. Infigui-r1: Advancing multimodal gui agents from reactive actors to deliberative reasoners. 2025. URL https://arxiv.org/abs/ 2504.14239. Y. Lu, J. Yang, Y. Shen, and A. Awadallah. Omniparser for pure vision based gui agent, 2024. URL https://arxiv.org/abs/2408.00203. Z. Lu, Y. Chai, Y. Guo, X. Yin, L. Liu, H. Wang, H. Xiao, S. Ren, G. Xiong, and H. Li. Ui-r1: Enhancing efficient action prediction of gui agents by reinforcement learning. 2025. URL https://arxiv. org/abs/2503.21620. D. Luo, B. Tang, K. Li, G. Papoudakis, J. Song, S. Gong, J. Hao, J. Wang, and K. Shao. Vimo: A generative visual gui world model for app agents, 2025a. URL https://arxiv.org/abs/2504.13936. R. Luo, L. Wang, W. He, and X. Xia. Gui-r1 : A generalist r1-style vision-language action model for gui agents. 2025b. URL https://arxiv.org/abs/2504.10458. H. Mozannar, G. Bansal, C. Tan, A. Fourney, V. Dibia, J. Chen, J. Gerrits, T. Payne, M. K. Maldaner, M. Grunde-McLaughlin, E. Zhu, G. Bassman, J. Alber, P. Chang, R. Loynd, F. Niedtner, E. Kamar, M. Murad, R. Hosn, and S. Amershi. Magentic-ui: Towards human-in-the-loop agentic systems, 2025. URL https://arxiv.org/abs/2507.22358. R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman. Webgpt: Browser-assisted question-answering with human feedback, 2022. URL https://arxiv. org/abs/2112.09332. S. Nayak, X. Jian, K. Q. Lin, J. A. Rodriguez, M. Kalsi, R. Awal, N. Chapados, M. T. Özsu, A. Agrawal, D. Vazquez, C. Pal, P. Taslakian, S. Gella, and S. Rajeswar. Ui-vision: A desktop-centric gui benchmark for visual perception and interaction, 2025. URL https://arxiv.org/abs/2503.15661. N. organization. NousResearch/hermes-agent: The agent that grows with you, 2026. URL https: //github.com/NousResearch/hermes-agent. J. Parker-Holder, S. Fruchter, and G. DeepMind. Genie 3: A new frontier for world models. https: //deepmind.google/blog/genie-3-a-new-frontier-for-world-models/, Aug 2025. Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, W. Zhong, K. Li, J. Yang, Y. Miao, W. Lin, L. Liu, X. Jiang, Q. Ma, J. Li, X. Xiao, K. Cai, C. Li, Y. Zheng, C. Jin, C. Li, X. Zhou, M. Wang, H. Chen, Z. Li, H. Yang, H. Liu, F. Lin, T. Peng, X. Liu, and G. Shi. Ui-tars: Pioneering automated gui interaction with native agents, 2025. URL https://arxiv.org/abs/ 2501.12326. P. Rajasekaran and A. Engineering. Harness design for long-running applications. https://ww w.anthropic.com/engineering/harness-design-long-running-apps, March 2026. Engineering blog post on Anthropic’s multi-agent harness architecture for developing long-running autonomous applications, with separate planner, generator, and evaluator roles. N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. A. for Computer Use Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer. Sam 2: Segment anything in images and videos, 2024. URL https: //arxiv.org/abs/2408.00714. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms, 2017. URL https://arxiv.org/abs/1707.06347. 18
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
B. Seed. Seed1. 8 model card: Towards generalized real-world agency. arXiv:2603.20633, 2026.
arXiv preprint
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024a. Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024b. URL https://arxiv.org/abs/2402.03300. Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face, 2023. URL https://arxiv.org/abs/2303.17580. G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu. Hybridflow: A flexible and efficient rlhf framework. In Proceedings of the Twentieth European Conference on Computer Systems, EuroSys ’25, page 1279–1297. ACM, Mar. 2025. doi: 10.1145/3689031.3696075. URL http://dx.doi.org/10.1145/3689031.3696075. Y. Shi, W. Yu, Z. Li, Y. Wang, H. Zhang, N. Liu, H. Mi, and D. Yu. Mobilegui-rl: Advancing mobile gui agent through reinforcement learning in online environment. arXiv preprint arXiv:2507.05720, 2025. P. Steinberger and OpenClaw Contributors. OpenClaw: Your own personal AI assistant. https: //github.com/openclaw/openclaw, 2026. F. Tang, Z. Gu, Z. Lu, X. Liu, S. Shen, C. Meng, W. Wang, W. Zhang, Y. Shen, W. Lu, et al. Gui-g2 : Gaussian reward modeling for gui grounding. arXiv preprint arXiv:2507.15846, 2025a. F. Tang, Y. Shen, H. Zhang, S. Chen, G. Hou, W. Zhang, W. Zhang, K. Song, W. Lu, and Y. Zhuang. Think twice, click once: Enhancing gui grounding via fast and slow systems. 2025b. URL https: //arxiv.org/abs/2503.06470. F. Tang, H. Xu, H. Zhang, S. Chen, X. Wu, Y. Shen, W. Zhang, G. Hou, Z. Tan, Y. Yan, K. Song, J. Shao, W. Lu, J. Xiao, and Y. Zhuang. A survey on (m)llm-based gui agents. 2025c. URL https://arxiv.org/abs/2504.13865. G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. M. Team. Minimax m2.7: Early echoes of self-evolution. https://www.minimax.io/news/mini max-m27-en, Mar 2026. MiniMax official blog post announcing the release of the MiniMax M2.7 AI model and its self-evolution capabilities, including enhanced productivity tasks, benchmark performance, and agent-driven innovation. V. Team, C. Gao, Z. Gu, Y. Liu, X. Qiu, S. Shen, Y. Wen, T. Xia, Z. Xu, Z. Zeng, et al. Ui-venus-1.5 technical report. arXiv preprint arXiv:2602.09082, 2026. H. Wang, H. Zou, H. Song, J. Feng, J. Fang, J. Lu, L. Liu, Q. Luo, S. Liang, S. Huang, et al. Ui-tars-2 technical report: Advancing gui agent with multi-turn reinforcement learning. arXiv preprint arXiv:2509.02544, 2025a.
19
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
J. Wang, H. Xu, H. Jia, X. Zhang, M. Yan, W. Shen, J. Zhang, F. Huang, and J. Sang. Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration, 2024a. URL https://arxiv.org/abs/2406.01014. J. Wang, H. Xu, J. Ye, M. Yan, W. Shen, J. Zhang, F. Huang, and J. Sang. Mobile-agent: Autonomous multi-modal mobile device agent with visual perception, 2024b. URL https://arxiv.org/abs/ 2401.16158. X. Wang, Z. Wu, J. Xie, Z. Ding, B. Yang, Z. Li, Z. Liu, Q. Li, X. Dong, Z. Chen, W. Wang, X. Zhao, J. Chen, H. Duan, T. Xie, C. Yang, S. Su, Y. Yu, Y. Huang, Y. Liu, X. Zhang, Y. Zhang, X. Yue, W. Su, X. Zhu, W. Shen, J. Dai, and W. Wang. Mmbench-gui: Hierarchical multi-platform evaluation framework for gui agents, 2025b. URL https://arxiv.org/abs/2507.19478. Y. Wang, X. Chen, X. Jin, M. Wang, and L. Yang. Openclaw-rl: Train any agent simply by talking, 2026. URL https://arxiv.org/abs/2603.10165. J. Wener. OpenCLI: Make any website your CLI. https://github.com/jackwener/opencli, 2026. Z. Wu, Z. Wu, F. Xu, Y. Wang, Q. Sun, C. Jia, K. Cheng, Z. Ding, L. Chen, P. P. Liang, and Y. Qiao. Os-atlas: A foundation action model for generalist gui agents, 2024. URL https://arxiv.org/ abs/2410.23218. T. Xie, J. Deng, X. Li, J. Yang, H. Wu, J. Chen, W. Hu, X. Wang, Y. Xu, Z. Wang, Y. Xu, J. Wang, D. Sahoo, T. Yu, and C. Xiong. Scaling computer-use grounding via user interface decomposition and synthesis, 2025. URL https://arxiv.org/abs/2505.13227. H. Xu, X. Zhang, H. Liu, J. Wang, Z. Zhu, S. Zhou, X. Hu, F. Gao, J. Cao, Z. Wang, et al. Mobile-agent-v3. 5: Multi-platform fundamental gui agents. arXiv preprint arXiv:2602.16855, 2026. Y. Xu, Z. Wang, J. Wang, D. Lu, T. Xie, A. Saha, D. Sahoo, T. Yu, and C. Xiong. Aguvis: Unified pure vision agents for autonomous gui interaction. 2024. URL https://arxiv.org/abs/2412.04454. H. Yan, J. Wang, X. Huang, Y. Shen, Z. Meng, Z. Fan, K. Tan, J. Gao, L. Shi, M. Yang, S. Yang, Z. Wang, B. Li, K. An, C. Li, L. Lei, M. Duan, D. Liang, G. Liu, H. Cheng, H. Wu, J. Dong, J. Huang, M. Chen, R. Yu, S. Li, X. Zhou, Y. Dai, Y. Deng, Y. Liang, Z. Chen, W. Sun, C. Yan, C. Xu, D. Li, F. Xiao, G. Fan, G. Li, G. Peng, H. Li, H. Li, H. Chen, J. Xie, J. Li, J. Zhang, J. Ren, J. Yuan, J. Yin, K. Cao, L. Zhao, L. Tan, L. Shi, M. Ren, M. Xu, M. Liu, M. Luo, M. Wan, N. Wang, N. Wu, N. Wang, P. Ma, Q. Zhang, Q. Wang, Q. Zeng, Q. Gao, Q. Li, S. Zhong, S. Gao, S. Liu, S. Gao, S. Luo, X. Liu, X. Liu, X. Hou, X. Liu, X. Feng, X. Cai, X. Wen, X. Zhu, X. Liang, X. Liu, X. Zhou, Y. Sui, Y. Zhao, Y. Shi, Y. Xu, Y. Zeng, Y. Zhang, Z. Weng, Z. Yan, Z. Huang, Z. Wang, Z. Yan, Z. Ge, J. Li, Y. Zhu, B. Jiao, X. Zhang, and D. Jiang. Step-gui technical report, 2025. URL https://arxiv.org/abs/2512.15431. J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023. URL https://arxiv.org/abs/2310.11441. J. Ye, X. Zhang, H. Xu, H. Liu, J. Wang, Z. Zhu, Z. Zheng, F. Gao, J. Cao, Z. Lu, J. Liao, Q. Zheng, F. Huang, J. Zhou, and M. Yan. Mobile-agent-v3: Fundamental agents for gui automation, 2025. URL https://arxiv.org/abs/2508.15144. C. Zhang, Z. Yang, J. Liu, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu. Appagent: Multimodal agents as smartphone users, 2023. URL https://arxiv.org/abs/2312.13771. C. Zhang, S. He, L. Li, S. Qin, Y. Kang, Q. Lin, S. Rajmohan, and D. Zhang. Api agents vs. gui agents: Divergence and convergence, 2025. URL https://arxiv.org/abs/2503.11069. 20
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
P. Zhao, G. Liu, Y. Liang, W. He, Z. Lu, Y. Huang, Y. Guo, K. Zhang, H. Wang, L. Liu, and Y. Liu. Mas-bench: A unified benchmark for shortcut-augmented hybrid mobile gui agents, 2025. URL https://arxiv.org/abs/2509.06477. C. Zheng, S. Liu, M. Li, X.-H. Chen, B. Yu, C. Gao, K. Dang, Y. Liu, R. Men, A. Yang, J. Zhou, and J. Lin. Group sequence policy optimization, 2025. URL https://arxiv.org/abs/2507.18071. Y. Zheng, L. Zhong, Y. Wang, R. Dai, K. Liu, X. Chu, L. Lv, P. Torr, and K. Q. Lin. Code2world: A gui world model via renderable code generation, 2026. URL https://arxiv.org/abs/2602.09856. H. Zhou, X. Zhang, P. Tong, J. Zhang, L. Chen, Q. Kong, C. Cai, C. Liu, Y. Wang, J. Zhou, et al. Mai-ui technical report: Real-world centric foundation gui agents. arXiv preprint arXiv:2512.22047, 2025.
21