Conceptio › Archive › arXiv CS
arXiv CSopen access

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents Haoting Shi1 * , Wenhao Wang2 * † , Weicheng Fang2 , Yaozhong Liang2 , Tian Jin1 , Pengxiang Zhao2 , Guangyi Liu2 , Siheng Chen1† , Yanfeng Wang1 1

Shanghai Jiao Tong University

arXiv:2609.05374v1 [cs.AI] 4 Sep 2026

Abstract Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly through the GUI, often producing inefficient trajectories. Real-world computer work is hybrid, combining visual-state inspection with precise, high-throughput command-line operations, so capable agents must coordinate both modalities over shared application state. Yet scalable hybrid environments remain scarce because supporting both GUI and CLI over real applications typically requires substantial manual engineering for each application. Existing agents also struggle to use the two interfaces complementarily: CLI-native agents lack visual perception for tasks involving interface state or layout, while GUI-native agents are inefficient for operations better executed through commands. We introduce CUA-Universe, a scalable environment-todata pipeline that turns real desktop software into hybrid GUI+CLI environments. App-Forge adapts applications into reproducible VMs and command-line surfaces it discovers, wraps, or generates, scaling to 16 applications across diverse domains; Task-Weave synthesizes diverse hybrid tasks of controllable difficulty from reusable operations over seed files, turning each environment into a continuous task source; and Path-Steer rolls agents out along efficient hybrid paths and harvests verified trajectories for post-training. Training on this data shifts behavior from inefficient GUI interaction and brittle CLI scripting toward effective GUI+CLI orchestration, improving both success and efficiency for our 9B model on CUA-Verse (Score +39.3 pts; −37% steps, −60% tokens), OSWorld (SR +16.8 pts; −57% steps, −44% tokens), and OSWorld-MCP (Score +7.84 pts; −27% steps, −30% tokens). By converting real desktop software into hybrid environments and reusable training data, CUA-Universe provides a scalable path toward more capable and efficient computeruse agents. We will release our code and data.

1

Introduction

Computer-use agents (CUAs) have advanced rapidly, completing a broad range of real desktop and mobile tasks on benchmarks such as OSWorld and AndroidWorld (Xie et al. 2024; Rawles et al. 2025). Yet leading agents still interact predominantly through the graphical user interface (GUI), resulting in trajectories that are often unnecessarily * These authors contributed equally. †

Corresponding authors.

2

Zhejiang University

long (Abhyankar, Qi, and Zhang 2025) and increasingly brittle on extended workflows (Yuan et al. 2026). Real computer use, however, is inherently multi-modal: users rely on the GUI for visually grounded interaction, while using command-line tools, scripts, or APIs for precise and highthroughput operations. Capable CUAs should therefore operate in hybrid GUI+CLI environments, dynamically choosing the interface best suited to each subtask while maintaining a shared application state across modalities (Song et al. 2025; Yang et al. 2025; Jia et al. 2025). Realizing such hybrid computer-use agents faces two coupled bottlenecks: scalable real-world environments and cross-modality orchestration. On the environment side, GUI-centric environments are costly to build and often require substantial manual engineering, while CLI-centric environments are easier to automate but lack access to visual layout and interface state. Building hybrid environments that support both modalities over the same application state therefore remains expensive and difficult to scale across applications. On the agent side, existing agents are typically optimized for one modality and struggle to use the two interfaces complementarily. CLI-native agents often lack visual perception and therefore resort to brittle scripts for tasks that depend on interface state or visual layout, while GUI-native agents can be inefficient for batch operations that a single command could complete. The deeper challenge is therefore orchestration: deciding when to switch interfaces and carrying task state across them, for example, locating a target through the GUI and then processing it through the CLI. To address these bottlenecks, we introduce CUAUniverse, a real-software environment framework for synthesizing, evaluating, and training hybrid GUI+CLI computer-use agents. CUA-Universe organizes real software into a scalable environment-to-data pipeline with three components. App-Forge adapts a desktop application into a reproducible VM and exposes it through a command-line surface it discovers, wraps, or generates. Driven by a coding agent rather than per-application manual engineering, it scales CUA-Universe to 16 desktop applications across diverse domains. Task-Weave synthesizes GUI+CLI hybrid tasks of controllable difficulty from applications, tools, and seed states, turning each environment into a continuous source of tasks. Path-Steer guides agents toward efficient hybrid execution paths, using the CLI for batch and

Scale: 16 Real Desktop Applications

Closed-loop Environment-to-Data Pipeline 1. App-Forge Environments

Creative & Interactive

Impact: Better and Cheaper CUA Agents Interaction Behavior GUI-Only: Over-clicking

Blender

Godot

Draw.io

…

GIMP

Installer Agent

Office & Knowledge

Reproducible VM + Metadata

Tool Construction

CLI-Only: Over-scripting

…

2. Task-Weave Task Synthesis Zotero

LibreOffice Writer

LibreOffice Calc

GUI+CLI Hybrid: Efficient orchestration

LibreOffice Impress

…

Media & Playback Task Instantiation

Operation Abstraction Kdenlive

OBS Studio

Audacity

VLC

Agentic Task Refinement

Empirical Gains: Higher Success Rate, Lower Cost

3. Path-Steer Hybrid Rollout

Technical & Web

QGIS

VS Code

Thunder bird

Chrome

Hybrid Interface

Path Steering Rollouts

Traj Harvesting

Figure 1: Overview of CUA-Universe. (1) Scale: 16 real desktop applications spanning creative & interactive, office & knowledge, media & playback, and technical & web domains, each exposed through both a GUI and application-specific CLI interfaces. (2) Closed-loop environment-to-data pipeline: App-Forge adapts an application into a reproducible VM and the command-line surface it discovers, wraps, or generates (installer agent → reproducible VM + metadata → tool construction); Task-Weave synthesizes diverse hybrid GUI+CLI tasks at scale over seed files (operation abstraction → task instantiation → agentic refinement); and Path-Steer rolls agents out onto efficient hybrid paths and harvests high-quality, directly reusable trajectories for post-training. (3) Impact: training on this data shifts interaction behavior from inefficient GUI interaction and brittle CLI scripting toward effective GUI+CLI orchestration, producing more capable and efficient agents with consistent improvements across diverse benchmarks, including CUA-Verse (Score +39.3 pts; −37% steps, −60% tokens), OSWorld (SR +16.8 pts; −57% steps, −44% tokens), and OSWorld-MCP (Score +7.84 pts; −27% steps, −30% tokens). precise operations and the GUI for visually grounded ones, and harvests high-quality trajectories for post-training. Together, these components turn real software into scalable sources of hybrid tasks and training data, enabling agents to learn more effective cross-modality orchestration between GUI and CLI. Our evaluation spans three complementary axes. First, to directly evaluate a model’s ability to orchestrate GUI and CLI actions over a shared application state, we construct CUA-Verse, a held-out benchmark of 160 hybrid tasks built on the eight desktop applications introduced by CUA-Universe, with evaluation tasks disjoint from the training data. On CUA-Verse, our model improves success by roughly 3× over its identical base while using 60% fewer tokens, achieving the best open-source result. Second, we test whether this capability transfers beyond our benchmark to OSWorld, where our model gains +16.8 success-rate points over GUI-only execution while using fewer steps and tokens. Finally, on OSWorld-MCP (Jia et al. 2025), our model generalizes to a different tool-invocation interface unseen during training, improving Score by +7.84 points over its base. Together, these results show that CUA-Universe improves not only performance within the environments it constructs, but also hybrid interaction skills that transfer across tasks, benchmarks, and tool interfaces. We summarize our contributions as follows: • A scalable environment-to-data pipeline. CUAUniverse turns real desktop software into hybrid

GUI+CLI agent environments through three components, App-Forge, Task-Weave, and Path-Steer, which cover application adaptation, task synthesis, and trajectory generation. Agent-driven construction allows the framework to scale efficiently to 16 real applications across diverse domains. • Scalable hybrid task synthesis and efficient trajectory generation. The framework discovers, wraps, or generates application-specific command-line interfaces, synthesizes GUI+CLI tasks of controllable difficulty from reusable operations and seed states, and steers agents toward efficient hybrid execution paths. This turns each environment into a continuous source of tasks and training trajectories while encouraging more effective crossmodality orchestration. • A benchmark for hybrid GUI+CLI orchestration. CUA-Verse provides 160 held-out hybrid tasks across eight desktop applications introduced by CUA-Universe, directly evaluating an agent’s ability to coordinate GUI and CLI actions over a shared application state. • Strong performance and transfer. Our 9B model trained on CUA-Universe data achieves roughly 3× the success of its base model on CUA-Verse while using 60% fewer tokens. The learned hybrid interaction skills further transfer to OSWorld with a +16.8 point gain over GUI-only execution and generalize to the unseen tool interface of OSWorld-MCP.

2 2.1

Related Work

Environment Synthesis for Computer-Use Agents

A growing body of work automatically synthesizes environments and trajectories to avoid the cost of hand-curated benchmarks (Zhou et al. 2024; Deng et al. 2023; Xie et al. 2024; Rawles et al. 2025). InfiniteWeb (Zhang et al. 2026), GUI-Genesis (Cao et al. 2026), and AutoWebWorld (Wu et al. 2026) generate functional web environments for posttraining, but their targets are artificial web pages confined to the GUI modality, leaving a sim-to-real gap on actual software. Scaling further, CUA-Gym (Wang et al. 2026a) cogenerates environments, tasks, and verifiable rewards across desktop and mock web applications for RLVR, and GymAnything (Aggarwal, Neubig, and Welleck 2026) turns arbitrary applications into agent environments and distills successful trajectories into a model. These environments, however, are driven purely through the GUI, so the trajectories they yield are inherently single-modality and cannot exhibit when to leave the GUI. In contrast, CUA-Universe builds each application into a hybrid GUI+CLI environment by exposing it through both the GUI and application-specific CLI tools. The synthesized tasks therefore require coordinating the two modalities over a shared application state, while the harvested trajectories carry efficiency and orchestration signals that a GUI-only pipeline structurally cannot provide.

2.2

Single-Modality GUI and CLI Agents

Computer-use agents have progressed along two separate lines. On the GUI side, agents are increasingly evaluated across diverse desktop, mobile, and heterogeneous platform settings, including OSWorld (Xie et al. 2024), AndroidWorld (Rawles et al. 2025), and FedGUI (Wang et al. 2025, 2026b), but may remain inefficient, taking far longer trajectories than necessary (Abhyankar, Qi, and Zhang 2025) and degrading on long-horizon, cross-application workflows (Yuan et al. 2026). On the CLI side, the terminal has become a first-class target: Terminal-Bench (Merrill et al. 2026) and TerminalWorld (Chu et al. 2026) provide hard command-line tasks, CLI-Universe (Hua et al. 2026) synthesizes verifiable terminal tasks, and further work scales terminal training environments and recipes (Cheng et al. 2026; Ivison et al. 2026). Yet competence in one modality does not transfer to the other: terminal agents cannot handle operations that depend on visual state, while GUI agents fall back to slow, element-by-element manipulation for batch operations. CUA-Universe instead trains on GUI+CLI hybrid tasks over a shared application state, teaching agents when to switch modalities, a form of coordination that neither line develops in isolation and that improves both success rate and execution efficiency.

2.3

Toward Hybrid GUI+CLI Agents

A recent trend combines visual actions with CLI tool calls rather than relying on either alone: CoAct-1 (Song et al. 2025) pairs a GUI operator with a coding agent, and UltraCUA (Yang et al. 2025) trains a hybrid-action model whose tools are mined from generic documentation and code

repositories over a fixed task set. Hybrid benchmarks further quantify the GUI+CLI efficiency trade-off (Li et al. 2026; Zhou et al. 2026; Fu et al. 2026). Unlike these existing works, which train a hybrid agent over a fixed set of tools and tasks, CUA-Universe contributes a new way to construct environments and synthesize data. Rather than mining a generic tool set, it grounds the tool layer in each real desktop application. Inspired by CLI-Anything (Yang, Fan, and Huang 2026), we generate agent-native CLIs when needed while also exposing applications’ native commandline tools, turning each environment into a continuous source of tasks and trajectories. On top of this, Path-Steer steers rollouts onto efficient hybrid paths, raising success rates and yielding higher-quality trajectories for computeruse agent post-training.

3 3.1

Method

Overview and Problem Formulation

Overview. As illustrated in Figure 1, CUA-Universe turns real desktop applications into hybrid GUI+CLI environments and further converts them into scalable sources of tasks and training trajectories. The framework consists of three components. App-Forge (§3.2) scales environment construction by adapting applications into reproducible hybrid environments with programmatic tool surfaces. TaskWeave (§3.3) scales task generation by composing diverse hybrid tasks from reusable operations grounded in real application states. Path-Steer (§3.4) scales trajectory collection by steering agents toward efficient GUI+CLI execution paths and harvesting rollouts for post-training. Problem formulation. We model a hybrid task as a POMDP: at step t the agent observes ot (a screenshot plus optional textual CLI returns) and emits at ∈ Agui ∪ Acli , where both action spaces read and modify over a shared persistent application state st . A task is a tuple τ = (instr, s0 , V ) of an instruction, a seed initial state, and a verifier V (ζ) ∈ [0, 1] scoring a trajectory ζ (a VLM judge); the pipeline synthesizes such tasks at scale and collects trajectories that solve them along efficient hybrid paths.

3.2

App-Forge: Scalable Agentic Environment Construction

Scaling hybrid environments across real desktop applications faces two application-specific bottlenecks: environment adaptation, since applications differ substantially in installation, configuration, and runtime dependencies; and tool construction, since agents need usable interfaces across heterogeneous applications. App-Forge addresses both through a scalable agentic construction pipeline that produces reproducible application environments together with CLI surfaces aligned with their GUIs. Application adaptation. The first bottleneck is reproducibly adapting diverse desktop software, whose installation procedures, dependencies, and launch configurations vary substantially across applications. An installer agent, guided by an installation skill, operates a persistent VM harness to install and configure the target application and

its supporting command-line utilities, interactively diagnosing failures until successful launch is verified. The resulting setup is distilled into a reproducible configuration, while lightweight introspection extracts application metadata for later task grounding and verification. The same adaptation workflow is reused across applications, allowing CUAUniverse to scale to 16 desktop applications across diverse domains. Eight are inherited from OSWorld but were originally GUI-only, for which App-Forge adds a shared-state CLI layer, while the other eight are introduced by CUAUniverse (Appendix B). Tool construction. The second bottleneck is constructing an expressive CLI surface across heterogeneous applications. We draw on three sources: native command-line tools we discover (e.g., blender --python-expr, cvlc), scripting APIs we wrap (e.g., bpy, LibreOffice UNO, and GIMP Script-Fu), and agent-native CLIs we generate when existing interfaces are insufficient, inspired by CLI-Anything. This layered design supports diverse automation surfaces without requiring a manually designed CLI for every application. GUI and CLI operate over the same project state, with lightweight adapters resynchronizing stale GUI views after external CLI edits. Appendix C lists the complete tool inventory for all applications.

3.3

Task-Weave: Compositional Hybrid Task Synthesis

Scaling task generation across real applications requires synthesizing tasks that reflect executable capabilities, cover diverse compositions, and remain feasible in the live environment. Task-Weave addresses these requirements in three stages: it abstracts reusable operations from agent exploration, composes them into diverse seed-conditioned tasks, and validates the resulting tasks through real execution. Operation abstraction. We first build a reusable operation pool that captures what can be reliably performed in each application. For a family of exploration seeds, we run parallel exploration agents with diverse goals, each targeting a different facet of the application, such as structure and visibility, appearance, or export and organization. Each agent interacts with the GUI while recording screenshots and actions. We slide a window over each trajectory and prompt an LLM to abstract the interaction into a high-level reusable operation, such as export scene to gltf rather than click. Each operation is also annotated with its execution procedure, which can later inform rollout guidance. We filter trivial or non-reusable candidates, then deduplicate, cluster, and aggregate operations across runs into a global operation pool with supporting evidence. Task instantiation. We then compose operations into diverse tasks grounded in concrete application states. Each task is conditioned on a real seed project, such as a .blend scene or .odp deck, together with its metadata, which defines the initial state s0 and the objects available for manipulation. We sample and score candidate operation chains and retain only compositions that are meaningful for the target seed. Difficulty is controlled by chain length and com-

position, ranging from single-operation edits to multi-step hybrid workflows. Each selected chain is compiled with its seed, supporting evidence, and mapped CLI tools into a task package containing an instruction, initial state s0 , verifier, and guidance. The instruction is synthesized as a natural user goal rather than a sequence of low-level actions. Agentic task refinement. Finally, we ground synthesized tasks in real execution before retaining them. A ReAct-style review agent (Yao et al. 2022) launches the application and performs a short multimodal interaction to check whether the instruction is feasible, unambiguous, and not already satisfied by the seed. When needed, it revises the instruction and guidance based on execution feedback. Valid tasks are retained, fixable tasks are revised, and invalid tasks are discarded,reducing hallucinated or infeasible task specifications before rollout.

3.4

Path-Steer: Efficiency-Aware Hybrid Rollout

Path-Steer converts synthesized tasks into efficient, verified training trajectories through three stages: hybrid execution enables GUI and CLI actions within a shared trajectory, efficient-path steering guides agents toward appropriate modality choices, and trajectory harvesting retains highquality rollouts for post-training. Each task is attempted multiple times from fresh environment instances, with attempts executed in parallel for throughput. Hybrid execution interface. We first provide agents with a unified interface for flexibly interleaving GUI and CLI actions. At each step, the agent emits either a GUI action or a CLI action. CLI actions are parsed against the application’s tool registry, expanded into concrete commands with the current working-file path injected when needed, and executed in the VM. Their return code and truncated output are included in the next observation, allowing both modalities to operate seamlessly within one trajectory over the shared project state maintained by application adapters (§3.2). Efficient-path steering. We then guide rollouts toward more efficient and appropriate modality choices. From each task’s operation chain, we derive a hybrid execution prior indicating when an operation is better suited to the CLI, such as batch, precise, or high-throughput operations, or to the GUI, such as operations depending on visual layout or interface state. The prior provides lightweight modality-level guidance without specifying low-level actions, reducing inefficient GUI interaction and brittle CLI scripting while yielding shorter and less redundant trajectories. Scoring and trajectory harvesting. Finally, we score each completed rollout with the task verifier V (ζ) ∈ [0, 1], implemented as a VLM judge over the trajectory. Verified rollouts are serialized into step-level and trajectory-level records containing observations, reasoning, GUI and CLI actions with their returns, screenshots, and final scores. We retain high-scoring hybrid trajectories as directly reusable supervision for subsequent post-training.

Training-data Composition 2,526 episodes 51.3%

Episodes

2,397 episodes 48.7%

Blender 191 Episodes

GIMP 240 Episodes

Zotero 196 Episodes

VS Code 255 Episodes

Draw.io 216 Episodes

2,526 140,088 steps 59.5%

Training steps

0%

Episodes

95,320 steps 40.5%

25%

50%

CUA-Universe extensions

75%

100%

LibreOffice Impress 200 Episodes

2,397

QGIS 218 Episodes

VLC 290 Episodes

Episodes

Audacity 376 Episodes

LibreOffice Calc 339 Episodes

Godot 418 Episodes

LibreOffice Writer 354 Episodes

OBS 445 Episodes

Chrome 363 Episodes

Kdenlive 466 Episodes

Thunderbird 356 Episodes

OSWorld applications

(a) Overall distribution of episodes and training steps.

(b) Task distribution across CUA-Universe extension applications.

(c) Task distribution across OSWorld applications.

Figure 2: Training-data composition. Our pipeline scales to diverse applications and produces substantial training data, with 4,923 episodes and approximately 235K training steps across all 16 applications. The balanced coverage of CUA-Universe extensions and OSWorld applications further demonstrates the diversity and scalability of the generated data.

4

Experiments

Training setup. App installation and CLI-tool construction (App-Forge) are driven by a Codex coding agent (GPT5.6), while operation abstraction and hybrid task synthesis (Task-Weave) are performed by Kimi K2.5 (Team et al. 2026); task success throughout the pipeline is scored by a VLM judge (GPT-5.4 (OpenAI 2026)). Our training data is generated by rolling out the same Kimi K2.5 backbone under Path-Steer and keeping trajectories that pass the judge at a score threshold of 0.75, yielding 4,923 verified episodes (∼235K step-level records; Figure 2). We then fine-tune Qwen3.5-9B (Qwen Team 2026) with LoRA on these steplevel trajectories for 3 epochs using the ms-swift (Zhao et al. 2024) framework, on 8× A100 GPUs in roughly two days; full training details are in Appendix H.

4.1

∼3× with 37% fewer steps and 60% fewer tokens (255K vs. 643K per episode), yielding the best accuracy–cost tradeoff among open models. (3) Capability is structured: Ours is strongest on audio/video apps (Audacity 0.815, OBS 0.605) and weaker on 3D/spatial ones (Blender 0.398, Godot 0.460); among open 8–9B models it dominates EvoCUA-8B (0.330 avg.) everywhere except Blender (0.495), indicating headroom in 3D data coverage. Kimi K2.5

0.660

0.500

0.425

0.400

0.664

0.451

0.535

0.538

Seed2.1 Pro

0.483

0.608

0.430

0.414

0.421

0.820

0.868

0.749

1.0

GPT-5.5

0.748

0.573

0.682

0.708

0.730

0.928

0.890

0.882

EvoCUA-8B

0.495

0.160

0.293

0.285

0.319

0.405

0.394

0.285

Qwen3.5-9B

0.205

0.364

0.353

0.140

0.040

0.100

0.165

0.145

Ours

0.398

0.640

0.470

0.460

0.720

0.547

0.605

0.815

ro

ot

S GI

ve nli

dio

Evaluation on CUA-Verse

Setting. We evaluate on CUA-Verse, a held-out benchmark of 160 hybrid GUI+CLI tasks (eight professional desktop applications × 20 tasks) synthesized by the CUAUniverse pipeline. Its tasks are disjoint from all training data, though the eight applications themselves are indomain; it measures an agent’s ability to complete GUI+CLI hybrid tasks, exposing both a GUI and a CLI interface so the agent must coordinate the two over a shared application state to solve each task. The eight applications (Blender, Draw.io, Zotero, Godot, QGIS, Kdenlive, OBS, Audacity) are each scored by a unified VLM judge using GPT-5.4 (OpenAI 2026); we additionally report steps and per-episode token as efficiency metrics. We compare against three proprietary models (Kimi K2.5, Seed2.1 Pro, GPT-5.5) and three opensource models (Qwen3.5-9B, EvoCUA-8B, and Ours). Results. Three observations stand out (Table 1, Fig. 3). (1) Distillation lifts a 9B model to closed-source level: Ours reaches 0.582 Score, the best among open-source models— surpassing Kimi K2.5 (0.522) and trailing only the proprietary Seed2.1 Pro (0.599) and GPT-5.5 (0.768). Since it is trained only on the teacher’s successful trajectories, it learns the upper tail and exceeds the teacher’s mean on CUA-Verse, consistent with STaR (Zelikman et al. 2022) and ReST (Gulcehre et al. 2023). (2) The gains are clean and efficient: against the identical Qwen3.5-9B (0.189), Score improves

Score

d

er

en Bl

Dr

.io aw

Zo

te

d

Go

Q

e Kd

tu

OB

S S-

y cit da Au

0.5

0.0

Figure 3: Application-level performance on the eight CUA-Verse applications. Each cell is the score for one model–application pair; darker shading is better. The horizontal rule separates proprietary (top) from open-source (bottom) models. Ours surpasses its Qwen3.5-9B on all eight applications, and is strongest on audio/video apps (Audacity, OBS).

4.2

Transfer to OSWorld

Setting. We evaluate whether training on CUA-Universe improves general computer-use capability even under GUIonly execution, and whether providing CLI access yields further gains through learned hybrid orchestration. We evaluate on OSWorld under a controlled 244-task protocol that excludes the os and multi-app splits and scores every task with the official verifier; SR is the fraction of tasks with reward 1. We compare two action interfaces under identical task text, environment setup, and verifier: in the GUI setting the agent acts purely through screen coordinates, and in the GUI+CLI setting it additionally receives a native execute cli tool with per-application command definitions and a single injected sentence giving the task’s input-file path.

Model

Mode

Kimi K2.5 Seed2.1 Pro GPT-5.5 Qwen3.5-9B EvoCUA-8B Ours

GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI

Score

CUA-Verse Steps Token (K)

0.522

21.1

275

0.599

40.1

449

0.768

23.1

193

0.189

56.2

643

0.330

41.7

377

0.582

35.2

255

SR (%) 53.7 54.5 55.7 59.0 66.8 68.0 21.7 24.6 40.6 43.3 23.4 40.2

Steps ↓ 27.2 28.9 28.4 27.7 14.7 19.0 58.6 54.1 34.8 42.1 39.6 28.6

OSWorld Step Gain ↑ Token(K) ↓ 224.8 1.15 277.3 188.0 1.10 215.6 595.0 1.55 550.3 521.9 2.63 596.1 325.6 1.58 408.0 325.7 2.35 286.5

Token Gain ↑ 1.09 0.94 1.58 2.03 1.35 1.79

Table 1: Combined performance on CUA-Verse and OSWorld. CUA-Verse: 8 apps × 20 tasks; Score is the average VLM judge score, and Token is the average number of tokens per episode in thousands. OSWorld: 244-task controlled scope excluding os and multi-app tasks, comparing GUI and GUI+CLI execution. Token/task is reported in thousands. Step Gain and Token Gain are the mean per-task GUI-to-GUI+CLI cost ratios over jointly solved tasks. A gain of k× indicates that GUI+CLI uses k times fewer steps or tokens on the same successfully solved tasks. Kimi K2.5

Seed2.1 Pro

GPT-5.5

EvoCUA-8B

Qwen3.5-9B Family

GUI

CLI

GUI

CLI

GUI

CLI

GUI

CLI

GUI

CLI

Ours

Chrome

0.261

0.261

0.283

0.304

0.326

0.391

0.283

0.239

0.239

0.239

0.370

GIMP

0.692

0.808

0.462

0.615

0.538

0.423

0.692

0.615

0.192

0.269

0.423

LO-Calc

0.532

0.553

0.617

0.702

0.830

0.915

0.298

0.191

0.085

0.085

0.191

LO-Impress

0.574

0.638

0.681

0.638

0.745

0.723

0.404

0.340

0.255

0.298

0.319

LO-Writer

0.652

0.609

0.696

0.783

0.826

0.783

0.304

0.304

0.304

0.348

0.522

Thunderbird

0.800

0.667

0.533

0.467

0.800

0.800

0.733

0.467

0.267

0.333

0.733

VLC

0.471

0.412

0.706

0.765

0.706

0.765

0.294

0.353

0.235

0.294

0.471

VS-Code

0.609

0.565

0.609

0.565

0.739

0.739

0.522

0.435

0.261

0.261

0.435

Model

SR 1.0

Interface metrics. Because GUI and GUI+CLI are evaluated on the same 244 tasks, we report paired measures of both success and efficiency gains from adding CLI access. Success gain is the net number of tasks newly solved by GUI+CLI, computed as (SRGUI+CLI − SRGUI ) × 244, and is reported directly from the SR results. Step Gain and Token Gain are computed only on tasks solved by both interfaces, using the mean per-task ratio of GUI to GUI+CLI cost for steps and tokens, respectively. A value of k× means that GUI+CLI reaches the same successful outcome with k times fewer steps or tokens, while averaging per-task ratios prevents a few long trajectories from dominating the metric. Results and analysis. Table 1 and Figure 4 summarize the comparison, from which we draw five findings. (1) Task-

Tok↓

Kimi K2.5 35.25 30.33 27.52 Seed2.1 Pro† 26.23 22.54 40.62 GPT-5.5† 29.92 25.00 27.06

85.21 99.30 54.62

0.5

Qwen3.5-9B 20.90 10.66 37.22 125.68 Ours 28.69 23.36 27.25 87.95 0.0

Figure 4: Per-application OSWorld scores across models and interfaces (GUI vs. GUI+CLI). Cells are heat-shaded by score. Adding the CLI yields Ours (rightmost) its largest gains, converting +41 previously-failed GUI tasks into successes.

All agents share a 60-step budget; observation history is each baseline’s default, while our Qwen-family models use a 3frame history with at most 4 images per step. Steps is the mean number of model decision calls per trajectory and Token/task the mean input+output tokens per task (thousands). Full harness and hardware details are in Appendix G.

SR↑ TIR↑ ACS↓ †

Table 2: Generalization on OSWorldMCP (244-task subset, exclude os/multi apps). ↑ higher / ↓ lower is better; Tokens in M. † : unified prompt, reference value.

specific distillation lifts a 9B open model to near openSOTA at the lowest cost: with GUI+CLI our model reaches 40.2% SR—up +16.8 points from its GUI-only 23.4% and within 3 points of EvoCUA-8B (43.3%)—while spending the fewest tokens per task (286.5k) and the fewest steps (28.6) of any agent, and more than doubling the untuned Qwen3.5-9B (24.6%) under the same interface. (2) The CLI interface is where our model’s gain concentrates—by far the largest net solved-task gain of any agent: adding CLI converts +41 previously-failed GUI tasks into successes (23.4 → 40.2%), against only +3 to +8 for every other model, none of which gains more than eight tasks. (3) For strong closed backbones the interface alone helps only marginally: GPT-5.5 (+3), Seed2.1 (+8) and EvoCUA-8B (+7) improve modestly and Kimi K2.5 is essentially flat (+2)—exposing a CLI the model was not trained to exploit yields little without task-specific tuning. (4) When CLI helps, it is also cheaper: on jointly-solved tasks the CLI interface uses fewer steps for every model (1.10–2.63×) and fewer tokens for most (up to 2.03×); our model reaches the same successes with 2.35× fewer steps. The lone exception

Group

Accept Rate (≥ 0.75)

Mean score

Avg steps

Avg tokens

Avg cost

0.51 0.44 0.54 0.45

0.71 0.63 0.75 0.67

22.75 26.68 25.53 28.78

331,988 385,107 263,500 303,653

$0.26 $0.31 $0.29 $0.33

Kimi K2.5 w/ Path-Steer Kimi K2.5 w/o Path-Steer Seed2.1 Pro w/ Path-Steer Seed2.1 Pro w/o Path-Steer

Table 3: Rollout efficiency of Path-Steer (enabled vs. disabled), for the Kimi K2.5 data-generation backbone and Seed2.1 Pro (a cross-backbone check). Accept Rate is the fraction of the 320 tasks scoring ≥ 0.75; tokens are averaged per task. is Seed2.1, whose sub-1× token efficiency shows the CLI interface can trade extra tokens for its accuracy gain. (5) Overall ranking: among closed agents GPT-5.5 (68.0%) > Seed2.1 (59.0%) > Kimi K2.5 (54.5%); among open models EvoCUA-8B (43.3%) leads with our model (40.2%) a close second, both far above the base.

4.3

Generalization to OSWorld-MCP

Setting. We evaluate whether CUA-Universe teaches transferable cross-modality orchestration, rather than benchmark-specific interaction patterns, on the held-out OSWorld-MCP benchmark (Jia et al. 2025), which augments OSWorld with 158 MCP tools and allows agents to freely combine GUI actions and tool calls. We report Score (task accuracy), Strict SR (perfect-score rate), TIR (accuracy of tool-use decisions), and ACS (average completion steps, lower is better), together with tokens per run as an additional efficiency metric. Evaluation uses 244 tasks after excluding os and multi apps, including 159 tool-beneficial and 85 non-tool-beneficial tasks, with max steps = 50 and history n = 3. We compare against Kimi K2.5, Seed2.1 Pro, GPT-5.5, and Qwen3.5-9B. Results. Three observations stand out (Table 2). (1) Hybrid interaction skills transfer to an unseen MCP interface: our model is trained only with application-specific CLI tools in CUA-Universe and never sees the MCP action space, yet it improves SR from 20.90% to 28.69% (+7.79 points) and more than doubles TIR from 10.66% to 23.36% over the base, showing that the learned tooluse behavior transfers beyond the training format. (2) The transferred capability is competitive with substantially larger models: a single 9B model surpasses Seed2.1 Pro on Score (29.51% vs. 27.03%) and approaches GPT-5.5 and Kimi K2.5, narrowing the gap to far larger closed-source models. (3) Generalization remains efficient: Ours reaches an ACS of 27.25 with 87.95M tokens, reducing steps by 27% and tokens by 30% relative to its base. These results show that the orchestration capability learned in CUAUniverse transfers to MCP without sacrificing efficiency.

4.4

Rollout Efficiency

Setting. We evaluate Path-Steer during data generation to measure whether explicit modality guidance improves both rollout quality and efficiency. On a fixed set of 320 synthesized tasks, we compare rollouts with and without Path-Steer while keeping the hybrid GUI+CLI interface unchanged. We use Kimi K2.5, our data-generation backbone, and addition-

ally evaluate Seed2.1 Pro to test whether the effect generalizes across models. We report Accept Rate, the fraction of trajectories scoring at least 0.75, and Mean Score for trajectory quality, together with average steps, tokens, and estimated per-task cost for efficiency. Since both settings have identical CLI access, this comparison isolates the contribution of steering itself. Results. Three observations stand out (Table 3). (1) PathSteer improves trajectory quality and efficiency simultaneously: on Kimi K2.5, it raises Accept Rate from 0.44 to 0.51 and Mean Score from 0.63 to 0.71, while reducing steps from 26.68 to 22.75, tokens from 385K to 332K, and per-task cost from $0.31 to $0.26. This indicates that explicit modality guidance produces shorter and higher-quality hybrid trajectories by reducing inefficient GUI interaction and brittle CLI scripting. (2) The effect generalizes across backbones: Seed2.1 Pro shows the same pattern, with Accept Rate improving from 0.45 to 0.54 and Mean Score from 0.67 to 0.75, together with lower steps, tokens, and cost. (3) Steering, rather than CLI access alone, drives the gains: the w/o Path-Steer baseline already has access to the same CLI tools, so the consistent improvements isolate the contribution of modality steering itself. Step-by-step comparisons between the two settings are provided in Appendix E.2.

5

Conclusion

We presented CUA-Universe as a step toward a different way of scaling computer-use agents: scaling the environments and interaction spaces from which agents learn, rather than relying only on larger models or more GUI-only trajectories. By turning real desktop software into shared-state GUI+CLI environments, CUA-Universe provides a scalable source of hybrid tasks and trajectories that teach agents not only how to act, but how to orchestrate complementary interfaces efficiently. The resulting 9B model shows that this capability is learnable and transferable, with strong gains on CUA-Verse and OSWorld and further generalization to the unseen tool interface of OSWorld-MCP. More broadly, our results suggest that hybrid environment construction can become a new axis for training computer-use agents, where each additional application expands the space of tasks, tools, and interaction strategies available for learning. We hope this shifts CUA development from collecting increasingly large amounts of single-modality behavior toward building scalable environments that continuously generate richer supervision for more capable, efficient, and general computeruse agents.

References Abhyankar, R.; Qi, Q.; and Zhang, Y. 2025. Osworldhuman: Benchmarking the efficiency of computer-use agents. arXiv preprint arXiv:2506.16042. Aggarwal, P.; Neubig, G.; and Welleck, S. 2026. Gymanything: Turn any software into an agent environment. arXiv preprint arXiv:2604.06126. Cao, Y.; Ran, D.; Wu, M.; Guo, Y.; Chen, X.; Li, A.; Cao, G.; Zhi, G.; Yu, H.; Li, L.; et al. 2026. Gui-genesis: Automated synthesis of efficient environments with verifiable rewards for gui agent post-training. arXiv preprint arXiv:2602.14093. Cheng, Z.; Wang, H.; Liu, Z.; Wang, X.; Zhu, X.; Guo, Y.; Lin, W.; Pan, J. Z.; and Wang, Y. 2026. TerminalWorld: Scaling Terminal-Agent Environments via Agent Skills. arXiv preprint arXiv:2605.20876. Chu, Z.; Hu, J.; Jiang, X.; Zou, P.; Li, H.; Peng, C.; O’Hearn, P.; Barr, E. T.; Harman, M.; Sarro, F.; et al. 2026. TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks. arXiv preprint arXiv:2605.22535. Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; and Su, Y. 2023. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36: 28091–28114. Fu, Y.; Fu, B.; Wu, Z.; Cheng, S.; Sun, X.; Yang, B.; Li, Z.; Zhao, Y.; Ding, Z.; Liu, Z.; et al. 2026. MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop. arXiv preprint arXiv:2606.22557. Gulcehre, C.; Paine, T. L.; Srinivasan, S.; Konyushkova, K.; Weerts, L.; Sharma, A.; Siddhant, A.; Ahern, A.; Wang, M.; Gu, C.; et al. 2023. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998. Hua, Z.; Yao, Y.; Xie, W.; Zhao, Y.; Liu, M.; Qiu, R.; Huang, Z.; Wang, Z.; Ji, Y.; Ye, Y.; et al. 2026. CLIUniverse: Towards Verifiable Task Synthesis Engine for Terminal Agents. arXiv preprint arXiv:2606.22883. Ivison, H.; Yin, J. O.; Shao, R.; Xiao, T.; Lambert, N.; and Hajishirzi, H. 2026. Tmax: A simple recipe for terminal agents. arXiv preprint arXiv:2606.23321. Jia, H.; Liao, J.; Zhang, X.; Xu, H.; Xie, T.; Jiang, C.; Yan, M.; Liu, S.; Ye, W.; and Huang, F. 2025. Osworld-mcp: Benchmarking mcp tool invocation in computer-use agents. arXiv preprint arXiv:2510.24563. Li, W.; Zhou, B.; Yu, Y.; Xu, Z.; Yang, Y.; Li, D.; and Shan, C. 2026. WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces. arXiv preprint arXiv:2606.09426. Merrill, M. A.; Shaw, A. G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J. Y.; Walshe, T.; Buchanan, E. K.; et al. 2026. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868. OpenAI. 2026. Introducing GPT-5.4. https://openai.com/ index/introducing-gpt-5-4/. Accessed: 2026-07-22. Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents.

Rawles, C.; Clinckemaillie, S.; Chang, Y.; Waltz, J.; Lau, G.; Fair, M.; Li, A.; Bishop, W.; Li, W.; Campbell-Ajala, F.; et al. 2025. Androidworld: A dynamic benchmarking environment for autonomous agents. In International Conference on Learning Representations, volume 2025, 406–441. Song, L.; Dai, Y.; Prabhu, V.; Zhang, J.; Shi, T.; Li, L.; Li, J.; Savarese, S.; Chen, Z.; Zhao, J.; et al. 2025. Coact1: Computer-using agents with coding as actions. arXiv preprint arXiv:2508.03923. Team, K.; Bai, T.; Bai, Y.; Bao, Y.; Cai, S.; Cao, Y.; Charles, Y.; Che, H.; Chen, C.; Chen, G.; et al. 2026. Kimi K2. 5: Visual Agentic Intelligence. arXiv preprint arXiv:2602.02276. Wang, B.; Lu, D.; Wang, J.; Bai, T.; Liu, S.; Zhang, Z.; Wang, H.; Hu, H.; Xie, T.; Bai, S.; et al. 2026a. Cuagym: Scaling verifiable training environments and tasks for computer-use agents. arXiv preprint arXiv:2605.25624. Wang, W.; Shi, H.; Yuan, M.; Lin, Y.; Tong, P.; Zhou, H.; Liu, G.; Zhao, P.; Wang, Y.; and Chen, S. 2026b. FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems. In Findings of the Association for Computational Linguistics: ACL 2026, 28747–28767. Wang, W.; Yu, Z.; Ye, R.; Zhang, J.; Liu, G.; Liu, L.; Chen, S.; and Wang, Y. 2025. FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Data. In Christodoulopoulos, C.; Chakraborty, T.; Rose, C.; and Peng, V., eds., Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 26387–26408. Suzhou, China: Association for Computational Linguistics. ISBN 979-8-89176-332-6. Wu, Y.; Peng, Y.; Chen, Y.; Ruan, J.; Zhuang, Z.; Yang, C.; Zhang, J.; Chen, M.; Tseng, Y.; Yu, Z.; et al. 2026. Autowebworld: Synthesizing infinite verifiable web environments via finite state machines. arXiv preprint arXiv:2602.14296. Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T. J.; Cheng, Z.; Shin, D.; Lei, F.; et al. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37: 52040–52094. Yang, Y.; Fan, T.; and Huang, C. 2026. Cli-anything: Towards agent-native computer use. arXiv preprint arXiv:2606.03854. Yang, Y.; Yang, Z.; Dou, Z.-Y.; Nguyen, A.; You, K.; Attia, O.; Szot, A.; Feng, M.; Ramrakhya, R.; Toshev, A.; et al. 2025. Ultracua: A foundation model for computer use agents with hybrid action. arXiv preprint arXiv:2510.17790. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629. Yuan, M.; Zhou, Z.; Xiong, X.; Wu, W.; Sun, J.; Song, J.; Cui, K.; Wang, B.; Wu, H.; Li, Y.; et al. 2026. OSWorld2. 0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. arXiv preprint arXiv:2606.29537. Zelikman, E.; Wu, Y.; Mu, J.; and Goodman, N. 2022. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35: 15476–15488.

Zhang, Z.; Wang, Z.; Zhang, X.; Guo, Z.; Li, J.; Li, B.; and Lu, Y. 2026. InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training. arXiv preprint arXiv:2601.04126. Zhao, Y.; Huang, J.; Hu, J.; Wang, X.; Mao, Y.; Zhang, D.; Jiang, Z.; Wu, Z.; Ai, B.; Wang, A.; Zhou, W.; and Chen, Y. 2024. SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning. arXiv:2408.05517. Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; et al. 2024. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, volume 2024, 15585–15606. Zhou, X.; Zhang, S.; Zhao, Y.; Wei, J.; Song, T.; Cohan, A.; and Zhao, C. 2026. GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents. arXiv preprint arXiv:2606.24551.

A A.1

Additional Experiments

Action-Modality Behavior on CUA-Verse

Setting. Beyond overall performance, we examine how different models use the two modalities on CUA-Verse. CUA-Verse is explicitly constructed from hybrid tasks whose reference solutions span both GUI and CLI operations over a shared application state, including visually grounded interactions that cannot be reduced to command execution alone. From the stored trajectories, we classify each nonterminal action as CLI or GUI and report CLI %, the fraction of executed GUI+CLI actions that use the CLI. We report it alongside Steps and the CUA-Verse Score to characterize the relationship among modality choice, efficiency, and performance. Model

Steps

CLI %

Score

Kimi K2.5 Seed2.1 Pro GPT-5.5

21.1 40.1 23.1

30.3 49.2 100.0

0.522 0.599 0.768

Qwen3.5-9B EvoCUA-8B Ours

56.2 41.7 35.2

0.0 3.0 25.3

0.189 0.330 0.582

Table 4: Action-modality behavior on CUA-Verse, averaged over the eight applications. Steps is the mean executed action count, CLI % is the CLI share of executed GUI+CLI actions, and Score is the average VLM-judge score. Greater CLI use often accompanies higher efficiency and performance, but CLI usage alone is insufficient: even GPT-5.5, which executes entirely through CLI in these trajectories, does not fully solve the benchmark, reflecting the visually grounded demands of CUA-Verse.

Results. Table 4 shows that effective CLI use is important, but the key capability is cross-modality orchestration rather than maximizing CLI usage. (1) The untuned base remains GUI-bound. Qwen3.5-9B issues no CLI actions, uses nearly the full 60-step budget, and obtains the lowest Score (0.189), reflecting inefficient GUI-only execution. (2) Hybrid training changes interaction behavior. With the same backbone, our model raises CLI usage to 25.3%, reduces mean steps by 37%, and increases Score from 0.189 to 0.582, showing that CUA-Universe teaches the model to exploit CLI operations while retaining GUI interaction when needed. (3) More CLI is not itself sufficient. GPT5.5 uses CLI for all recorded actions and achieves the highest Score (0.768), yet still falls well short of perfect performance. Conversely, EvoCUA-8B remains almost entirely GUI-based despite having access to the CLI and obtains only 0.330. Together, these results suggest that CUA-Verse rewards the ability to select and coordinate modalities according to the task, rather than simply favoring either GUI or CLI in isolation.

App

Base

8-app LoRA (OOD-only)

16-app LoRA (+in-domain)

Chrome GIMP Calc Impress Writer Thunderbird VLC VS Code

23.9 26.9 8.5 27.7 34.8 33.3 29.4 26.1

26.1 30.8 10.6 25.8 39.1 46.7 31.4 34.8

37.0 42.3 18.8 31.8 52.2 73.3 37.4 43.5

All apps

24.2

27.4 (+3.2)

40.2 (+16.0)

Table 5: Out-of-domain training transfer on OSWorld under the controlled 244-task GUI+CLI scope. Values are success rates in percent. The 8-app OOD-only LoRA is trained on cua-universe extensions applications disjoint from the OSWorld applications, whereas the full 16-app LoRA is trained on both. Parenthesized values in the final row denote absolute percentage-point gains over the untuned Qwen3.5-9B.

A.2

Out-of-Domain Training Transfer

Setting. To test whether the LoRA gains reflect a transferable CLI-usage capability rather than memorization of the evaluation applications, we train an 8-app OOD-only LoRA exclusively on applications that are disjoint from the eight OSWorld evaluation applications and evaluate it zeroshot under the same controlled 244-task GUI+CLI scope. We compare it against the untuned Qwen3.5-9B and our full 16-app LoRA, whose training mixture additionally includes the eight OSWorld application domains. Results. Three observations stand out (Table 5). (1) Hybrid CLI-usage skills transfer across applications: despite being trained exclusively on applications disjoint from the OSWorld evaluation applications, the 8-app OOD-only LoRA improves over the untuned base on seven of the eight applications and raises the overall success rate from 24.2% to 27.4% (+3.2 points). This zero-shot improvement indicates that the LoRA learns a partially application-agnostic capability for selecting and coordinating GUI and CLI actions rather than only memorizing application-specific commands. (2) The transfer is heterogeneous across applications: the largest out-of-domain gains occur on Thunderbird (+13.4 points) and VS Code (+8.7), followed by Writer (+4.3) and GIMP (+3.9); Impress is the only application on which performance decreases, by 1.9 points. (3) In-domain coverage substantially compounds the transferable gain: the full 16-app LoRA improves over the base on all eight OSWorld applications and raises the overall success rate to 40.2% (+13.2 points), showing that transferable hybridinteraction skills and application-specific training data provide complementary benefits.

B

Adapted Applications

CUA-Universe currently provides adapters for the 16 desktop applications listed in Table 6, covering diverse software

domains. Eight applications are inherited from OSWorld, and eight additional applications are introduced by CUAUniverse. Each application implements the uniform adapter interface described in §3.2, exposing graphical, commandline, and programmatic interfaces that operate on the same underlying state.

OSWorld applications GIMP LibreOffice Writer Chrome LibreOffice Impress VLC VS Code LibreOffice Calc Thunderbird

46 30

15 14 13 13 10 10

CUA-Universe extensions

Application

Domain

Seed

Chrome

Web browsing

Profile

GIMP

Image editing

.xcf

LibreOffice Calc

Spreadsheets

.ods

LibreOffice Impress

Presentations

.odp

LibreOffice Writer

Word processing

.odt

Thunderbird

Email

Profile

VLC

Media playback

Media file

VS Code

Code editing

Workspace

Audacity

Audio editing

.aup3

Blender

3D modelling

.blend

Draw.io

Diagramming

.drawio

Godot

Game development

.tscn

Kdenlive

Video editing

.kdenlive

OBS Studio

Screen recording

.json

QGIS

Geospatial analysis

.qgz

Zotero

Reference management

Library

OSWorld applications

CUA-Universe extensions

Table 6: Desktop applications supported by CUAUniverse. Blue rows indicate applications inherited from OSWorld, whereas orange rows indicate applications added by CUA-Universe.

C

C.1

GIMP — 46 commands

The GIMP uses a live backend that drives the actual GUI application, while exposes the command interface and return

38 38 36 29 26 25 16 0

10

20

30

40

50

Agent-visible commands OSWorld applications

CUA-Universe extensions

Figure 5: Agent-visible command count per application, grouped by source: the eight applications reused from OSWorld and the eight CUA-Universe extensions (∼404 commands in total).

results in JSON format. PROJECT

project new project open project save project info project json project profiles

CANVAS

canvas resize canvas scale canvas crop canvas mode canvas dpi canvas info

LAYER

layer new layer add from file layer list layer remove layer duplicate layer move layer set layer flatten layer merge down

FILTER

filter add filter list filter list available filter info filter set filter remove

DRAW

draw text draw rect

MEDIA

media probe media histogram media list media check

EXPORT

export render export presets export preset info

LIVE

live status live canvas info live layer list live draw text live draw rect live export render

SESSION

session undo session redo session history session status

Tool Construction

The tool layer converts each application’s native automation surface—a scripting API, a project-file serializer, or an installed command-line binary—into a small, bounded vocabulary of subcommands that return JSON. A registry stores only metadata (tool ID, description, and the complete command template); the rollout runner executes the named command inside the benchmark VM and returns its exit code, stdout, and stderr to the next decision step. Where an application already ships a usable command-line interface, the registry wraps it directly; where it does not, a bounded harness operates on the native project artifact—or on the live application—so that every effect stays auditable in the same file the GUI opens. The deployed snapshot spans 16 applications and ∼404 agent-visible commands: 151 commands across the eight applications reused from OSWorld and 253 across the eight CUA-Universe extensions (Figure 5). Command counts are not padded to a target range—they track each application’s real capability surface. For brevity, the per-application sections below list tool IDs grouped by command family.

45

Blender QGIS Zotero OBS Studio Kdenlive Godot draw.io Audacity

C.2

Blender — 45 commands

A bounded bpy program runs in a single headless invocation, during which it has full read/write access to the real .blend artifact, allowing it to both inspect and alter the scene contents permanently. SCENE

scene new scene open scene save scene info scene profiles scene json

OBJECT

object add object remove object duplicate object transform object set object list object get

MATERIAL

material create material assign material set material list material get

MODIFIER

modifier list available modifier info modifier add modifier remove modifier set modifier list

item export item citation item bibliography item context item analyze item add to collection item move to collection note get note add

NOTE

camera add camera set camera set active camera list

SEARCH

search list search get search items

TAG / STYLE

tag list tag items style list

LIGHT

light add light set light list

IMPORT / SESSION

ANIMATION

animation keyframe animation remove keyframe animation frame range animation fps animation list keyframes

CAMERA

RENDER

render settings render info render presets render execute render script

SESSION

session status session undo session redo session history

C.3

OBS — 36 commands

A typed scene, source, filter, transition, and output model serialized to the native OBS scene collection, allowing the agent to compose and control live production setups. PROJECT

project info project json project save

SCENE

scene add scene remove scene duplicate scene set active scene list

SOURCE

source add source remove source duplicate source set source transform source list

FILTER

filter add filter remove filter set filter list filter available

AUDIO

audio add audio remove audio volume audio mute audio unmute audio monitor audio list

TRANSITION

transition add transition remove transition set active transition duration transition list

OUTPUT

output streaming output recording output settings output info output presets

QGIS — 38 commands

Bounded PyQGIS operations against a native .qgs project, with optional GUI sync, giving the agent full read and write access to layers, features, layouts, and processing pipelines. PROJECT

project new project open project info project set crs project save layer create vector layer list layer info layer remove layer style simple layer style graduated layer style categorized layer label simple layer label buffer layer category visibility layer filter attribute

LAYER

FEATURE

feature add feature add point feature list

LAYOUT

layout create layout list layout info layout remove layout add map layout add legend layout legend remove layer layout add label layout sync extent

PROCESS

process list process help process run

EXPORT

export presets export pdf export image

SESSION

session status session history

APP / REPL

app open repl

C.4

C.5

Zotero — 38 commands

import file import json session

C.6

LibreOffice Writer — 30 commands

ODF-backed Writer operations over the live UNO bridge, implementing the Writer slice of the shared cli-anythinglibreoffice harness for creating and editing rich text documents. DOCUMENT

document new document open document save document info document profiles document json

WRITER

writer add paragraph writer add heading writer add list writer add table writer table list writer table insert row writer table set cell writer table set row background writer add page break writer remove writer list writer set text

The seeded Zotero profile is accessed via its local database, browser connector, and Local API, enabling full management of collections, items, notes, and saved searches. APP

app status app version app launch app enable local api app ping

COLLECTION

collection list collection find collection tree collection get collection items collection use selected collection create collection export

STYLE

style create style modify style list style apply style remove

EXPORT

export presets export preset info export render

item list item find item get item children item notes item attachments item file

SESSION

session status session undo session redo session history

ITEM

C.7

Kdenlive — 29 commands

A project-file editor for Kdenlive’s MLT XML format, using melt for rendering and writing a native .kdenlive project with timeline clips, filters, transitions, and guides.

CONNECT

connect add connect remove connect label connect style connect list

EXPORT

export render

SESSION

session undo session redo session status

C.10

Audacity — 16 commands

PROJECT

project info project json project save project profiles

BIN

bin import bin list bin get

TIMELINE

timeline add track timeline add clip timeline remove clip timeline trim timeline split timeline move timeline list

IMPORT / TRACK

filter add filter set filter list filter available

SELECT

select all select time select tracks

EFFECTS

TRANSITION

transition add transition set transition list

amplify normalize fade in fade out reverse change pitch change speed

LABEL / EXPORT / RAW

GUIDE

guide add guide list

EXPORT

export xml export presets

SESSION

session status session undo session redo session history

FILTER

C.8

C.11

import track new get info

label add export run

Chrome — 15 commands

A restricted Chrome DOM surface (DOMShell), exposed through page, accessibility-tree, and action commands for navigation, file-system access, and element interaction.

Godot — 26 commands

Godot project assets, scenes, scripts, and export configuration, accessed via the engine’s conventions and CLI runtime for project creation, scene editing, and build export. ENGINE

engine version engine status

EDITOR

editor open

PROJECT

project create project info project scenes project scripts project resources project reimport project set setting project add input action project add input key scene create scene read scene add node scene set property scene set control layout

SCENE

PAGE

page open page reload page back page forward page info

FS

fs ls fs cd fs cat fs grep fs pwd

ACT

act click act type

SESSION

session status session daemon start session daemon stop

C.12

LibreOffice Impress — 14 commands

Native .odp slide, content, element, and export commands over ODF/UNO, implementing the Impress slice of the shared cli-anything-libreoffice harness. DOCUMENT

document save document info

IMPRESS

impress add slide impress remove slide impress set content impress list slides impress add element impress remove element impress move slide impress duplicate slide impress get slide

EXPORT

export presets export preset info export render

script run script inline script validate script read script write script append

SCRIPT

EXPORT

export presets export build

SESSION

session

C.9

A live client for Audacity’s Mod-Script-Pipe, treating the running GUI project as the source of truth for importing audio, making selections, and applying effects.

draw.io — 25 commands

The native .drawio XML/mxGraph model, manipulated through a constrained diagram API that exposes pages, shapes, connectors, and export operations. PROJECT

project new project open project save project info project xml

PAGE

page add page remove page rename page list

SHAPE

shape add shape remove shape list shape label shape move shape resize shape style

C.13

VS Code — 13 commands

The installed code binary, invoked to open paths, diff and merge files, manage extensions, and access settings and keybindings. OPEN

open open new window open user settings json open keybindings json

NAVIGATE

goto diff merge

WORKSPACE

add folder wait file closed

EXTENSIONS

install extension uninstall extension list extensions

STATUS

status

C.14

VLC — 13 commands

Direct vlc and cvlc invocations to launch media, open files at specific times or segments, capture snapshots, transcode audio and video, and manage configuration. APP

app open app help

OPEN

open media file open media at time open network stream open preferences file

PLAY

play fullscreen play paused segment

CAPTURE

snapshot frame

CONVERT

convert audio mp3 convert audio wav convert video mp4

CONFIG

reset user config

C.15

LibreOffice Calc — 10 commands

Bounded sheet and cell operations over ODF/UNO, implementing the Calc slice of the shared cli-anything-libreoffice harness for spreadsheet creation and manipulation. DOCUMENT

document save document info

SHEET

calc list sheets calc add sheet calc rename sheet

CELL

calc get cell calc set cell

EXPORT

export presets export preset info export render

C.16

Thunderbird — 10 commands

The installed executable’s profile-aware launch and compose capabilities, with every command naming the benchmark profile to open mail, address book, calendar, and compose emails. APP

app open app help

OPEN

open mail open addressbook open calendar

COMPOSE

compose email compose cc bcc compose attachment mailto email

DESKTOP HANDLER

xdg email

D D.1

Task-Weave

Seed Examples

A seed is a concrete, pre-loaded application state—an actual project, document, or media file opened in its native GUI app—that anchors task synthesis in real, verifiable context rather than a generic natural-language prompt. Each seed carries genuine content and metadata (e.g., a populated spreadsheet, a raster design with stable regions, or a

saved .drawio graph), so synthesized tasks reference elements that truly exist and produce app state that can be checked deterministically. Grounding generation in seeds reduces hallucinated targets, yields more executable and diverse instructions, and makes success criteria objective. Figure 6 shows six representative seeds spanning 3D modeling, spreadsheets, media playback, image editing, presentations, and diagramming; together they illustrate the breadth of GUI applications and file types our pipeline builds on. Seeds are not hand-authored. A coding agent (Codex) searches for and downloads real, diverse source files from public repositories and asset libraries, then uses formatconversion tools to normalize each one into an agentreadable, synthesis-friendly project state—so the seed pool scales automatically while staying grounded in authentic, heterogeneous content rather than templated fixtures.

E E.1

Steer-Path Rollouts

Rollout Agent Prompt

Figure 10 shows the message format of a data-generation rollout (the Kimi K2.5 backbone on a draw.io task). Two properties are worth noting. First, the rollout interface exposes GUI and CLI jointly: the system prompt lists the application’s CLI tool registry alongside pyautogui, wait, and terminate, and the baseline modality guidance is the generic instruction to “use CLI for precise operations and GUI for visual tasks.” Second, this is where Path-Steer enters: the task guidance field carries the efficiency-aware prior for this attempt (an ordered plan such as “First . . . ”), steering the backbone toward a shorter hybrid path during data generation. The agent then interleaves CLI edits with GUI waits, and each CLI call returns a structured JSON result (rc, output) that becomes the next observation. Path-Steer is used only during these rollouts; at evaluation the guidance field is empty (Appendix I.1). The task text is elided below; the point is the interface and the guided GUI+CLI interleaving, not the specific diagram.

E.2

Steer-Path Rollout Examples

In this VLC desktop task, the w/ Path-Steer trajectory overall outperformed the w/o Path-Steer trajectory. Although the w/ Path-Steer trajectory made several ineffective attempts with VLC CLI parameters in the early stage, it was able to switch to a feasible alternative solution in time, namely using ffmpeg to complete the video rotation, and eventually succeeded in both exporting the corrected video and opening the Audio Effects panel. In contrast, while the w/o Path-Steer trajectory followed a more intuitive GUI-based workflow, it became stuck in repeated searching and ineffective interactions during the filter configuration stage, and ultimately failed to complete the task. This case suggests that Path-Steer may not necessarily reduce local trial-and-error, but it can significantly improve task completion, error recovery, and goal convergence in complex desktop environments.

F

VLM Judge

We score task success with a VLM-as-judge, following the now-standard practice of using strong (vision-)language

Blender. This Objaverse product seed opens a concrete webcam-cover 3D asset, grounding modeling tasks in real mesh and material structure.

LibreOffice Calc. This budget workbook seed provides real sheets, headers, rows, and editable cells for grounded spreadsheet tasks.

VLC. This media seed opens a concrete video file, grounding playback and media-control tasks in a real duration and visual stream.

GIMP. This poster seed exposes a real raster design with stable regions for annotation, cropping, banner, and export tasks.

LibreOffice Impress. This photo-heavy deck seed grounds presentation tasks in existing slides, layouts, images, and editable text frames.

Draw.io. This flowchart seed grounds diagram-editing tasks in existing nodes, connectors, labels, and saved .drawio XML.

Figure 6: Six opened seed examples used to ground task synthesis across diverse GUI applications. Each seed is a concrete project or file with metadata and verifiable app state, rather than a generic natural-language prompt. models as automatic evaluators of agent trajectories; in the GUI and computer-use setting in particular, VLM judges are widely used both to score task completion from screenshots and to curate training trajectories for post-training. Independent of the task synthesizer and of any rollout reward, the judge takes the task instruction, a compact rendering of the trajectory, and a chronological sample of screenshots, and returns a score in [0, 1] with a short justification. Its rubric treats CLI output and exported-artifact evidence as authoritative for file-producing tasks while using the final screenshot as primary evidence otherwise. We run the judge (GPT5.4 (OpenAI 2026)) at temperature 0.1 with up to 15 screenshots; the full prompt is shown in Figure 11.

F.1

Agreement with Human Labels

Because Score is our headline CUA-Verse metric and the same judge filters training data, we validate it against human judgment. Three annotators independently re-scored a stratified sample of trajectories for all 16 applications, blind to the judge’s score. For each application we drew 30 trajectories the judge accepted (score ≥ 0.75) and 30 it rejected (score < 0.75), and labeled each as fully success, partial success, or failure. We take the binary decision positive = human “fully success” and negative = otherwise, matching the semantics of the 0.75 acceptance threshold used to filter training data. Because the sample is stratified rather than proportional, we reweight each application’s confusion cells

by its true acceptance rate p (Table 7, second column) before computing the aggregate agreement and κ. Two findings anchor the judge’s reliability. On the acceptance set (480 trajectories) the judge attains 99.0% precision (475/480), and crucially not one accepted trajectory was a human “failure”—the five imperfections are all partial successes—so the label-noise upper bound on training data is ≤ 1.0%. On the rejection set (480 trajectories) the falsenegative rate is 5.2% (25 fully-successful trajectories scored low); these cost data yield but do not inflate reported Score. Reweighted to the true distribution, judge–human agreement is 97.0% and Cohen’s κ = 0.94 (Landis–Koch “almost perfect”). Under a lenient criterion that counts partial successes as positive, the picture is unchanged where it matters: the judge accepts zero true failures (precision 100%; all 480 accepted trajectories are at least partial successes). Its rejection set is dominated by partial successes (406/480) with only 49 true failures, confirming that the 0.75 threshold deliberately screens out partially-completed work—the intended behavior for a training-data filter—rather than confusing success with failure.

G

Rollout Harness and Budget

Rollout Harness and hardware. Each task runs in an isolated OSWorld virtual machine provisioned with 4 vCPU and 4 GB of guest RAM on a copy-on-write overlay disk,

# Task Instruction:

The video I'm playing in VLC is upside-down. Please rotate it 180 degrees and save the corrected version as a new MP4 file named 'rotated_fixed.mp4' in my Documents folder. Then use VLC's Tools menu to open the Audio Effects panel so I can adjust the equalizer settings for playback.

# SteerPath Guidance:

First, use the CLI command `vlc_convert_video_mp4` with the transpose video filter to rotate 180 degrees: `cvlc -I dummy \"/home/user/Documents/rotated_orientation_sample.avi\" --video-filter=transpose --transpose-angle=180 mux=mp4,dst='/home/user/Documents/rotated_fixed.mp4'}\" vlc://quit`. After conversion completes, open the converted file in VLC with `vlc_open_media_file` to verify the rotation, then use ...

"__CLI__:cvlc -I dummy \"/home/user/Do cuments/rotated_orientation_sample.avi\" --video-filter=transpose ... "

pyautogui.click(714, 394)

"__CLI__:VLC_VERBOSE=-1 vlc --no-video-title-show \ "/home/user /Documents/rotated_fixed.mp4\""

pyautogui.click(1008, 394)

pyautogui.click(1044, 753)

pyautogui.scroll(-3, 936, 570)

FAIL

# Without SteerPath:

pyautogui.click(1131, 481)

Figure 7: VLC task trajectories w/ Path-Steer and w/o Path-Steer. The Path-Steer trajectory shows optimal solution using CLI tools and GUI actions, whereas the trajectory without sterr-path guidance becomes stuck during filter configuration and fails the task. with a 3 s post-action settle before each observation. All agents share a 60-step budget; observation history follows each baseline’s default, while our Qwen-family models use a 3-frame history with at most 4 images per step.

Rollout budget. Trajectory collection dominates the wallclock cost of building the fine-tuning dateset, because each rollout drives a full GUI environment rather than a single forward pass. A single rollout takes on average t̄ ≈ 5 minutes end-to-end (environment reset, per-step model calls, action execution, and the post-action settle), and a training run consumes on the order of N ≈ 104 rollouts. Because each environment needs only ∼4 vCPU and no GPU, this collection is CPU-bound and embarrassingly parallel. Concretely, a single commodity 128-core CPU server hosts P = ⌊128/4⌋ = 32 environments concurrently, so the full N t̄ ≈ 50,000 VM-minutes reduce to N t̄/P ≈ 1,560 minutes—just over one day of wall-clock time on that one machine. This is the key practicality of our recipe: the entire data-generation pipeline that produces our 9B model fits on a single ordinary CPU box, with no GPU cluster and no multimachine orchestration. Throughput in practice is slightly below this ceiling because of environment-reset overhead and occasional VM stalls (we observe ∼1.3–1.5 calendar days on one 128-core host), and the job is trivially shardable across additional machines when faster turnaround is needed.

H H.1

Training Details

Data

All records satisfy VLM score ≥ 0.75 and are step-level examples in Qwen XML format with a three-step history context. The CUA-Verse pool is balanced at 17,511 records per app across Audacity, Blender, draw.io, Godot, Kdenlive, OBS, QGIS, and Zotero. The OSWorld pool is balanced at 11,915 records per app across LibreOffice Writer, LibreOffice Calc, LibreOffice Impress, VS Code, VLC, Thunderbird, GIMP, and Chrome. The OSWorld build enforces exactly four images per record; the CUA-Verse build is VLMfiltered but not strictly four-image (22,590 records have an image count other than four). Figure 2 summarizes the overall split and the per-application episode counts for both pools, and Table 8 reports the final split sizes.

H.2

Base Model and Fine-tuning

We use Qwen3.5-9B as the base model in bfloat16. Finetuning is performed with LoRA while the base weights, vision tower, and multimodal aligner remain frozen; only the language-model linear modules are trainable. Table 10 lists the optimization and sequence settings used for fine-tuning.

I I.1

Experiments

Evaluation Agent Prompt

Each evaluated backbone uses its own actor prompt, and each is run in two modes across OSWorld and CUA-Verse:

# Task Instruction:

Please help me modify the setting of VS Code to keep my cursor focused on the debug console when debugging in VS Code, instead of automatically focusing back on the Editor.

GUI + CLI Mode: 5 CLI calls + 12 GUI actions -> reward = 1

pyautogui.click(492, 321)

pyautogui.typewrite ("debug focus console")

`CLI: cat > ~/.config/Code/User/settings.json << 'EOF' {"security.workspace.trust.startupPrompt": "never","editor.wordWrap": "on", "files.autoCreate" : true, "debug.focusEditorOnBreak":false} EOF

Done

GUI Mode: 60 GUI actions -> reward = 0

pyautogui.click(490, 321)

pyautogui.typewrite ("debug focus")

pyautogui.click(845, 388)

FAIL

Figure 8: An OSWorld VS Code task under GUI+CLI vs. GUI-only interfaces. GUI+CLI mode combines GUI navigation with direct settings-file editing: it locates the target setting, writes debug.focusEditorOnBreak: false to settings.json, and verifies persistence, completing the task in 5 CLI calls and 12 GUI actions (reward 1). GUI-only mode repeatedly searches and clicks through the settings UI but fails to commit the correct configuration, exhausting the 60-action budget (reward 0). competence on the software the pipeline covers. CLI abstract tools

GUI abstract tools

8.5

8.5

8

Abstract tools / task

a CLI mode that additionally exposes the execute cli function, and a GUI mode that follows the standard OSWorld screenshot-and-pyautogui protocol. Figure 12 shows the Qwen CLI-mode prompt as a representative example. Regardless of backbone or mode, the path-hint field defaults to none at evaluation, so no task hint or solution guidance is appended and the evaluated policy is prior-free—the PathSteer priors of Appendix E.1 act only during data-generation rollouts.

6.6

6.2

5.7

6

5.8 4.9

4.7 4

2

I.2

CUA-Verse Benchmark

CUA-Verse comprises 160 hybrid GUI+CLI tasks (eight applications × 20 tasks) synthesized by the same pipeline as the training data. To characterize what the tasks demand, we abstract each task’s reference solution into a deduplicated set of high-level tools and classify each as GUI (normalized keyboard/click/move/drag/scroll/vision) or CLI (application command groups such as godot scene * or qgis layer *). Figure 9 reports the resulting per-application composition. Tasks require 5.7 abstract tools on average, of which 59% are CLI, and every application mixes both modalities—confirming that CUA-Verse is genuinely hybrid rather than solvable by either modality alone. The two extremes are illustrative: Blender is GUI-dominant (spatial 3D manipulation), whereas Zotero and Kdenlive are CLI-heavy (structured library and timeline edits). The eight applications are in-domain by design, since we want to measure hybrid

0

y acit

Aud

der

w.io

Blen

Dra

ot

God

e io nliv Stud Kde OBS

QGIS

ro

Zote

Figure 9: CUA-Verse task composition. Mean number of abstract tools per task for each application, split into CLI (orange) and GUI (blue) operations. Every application requires both modalities.

J

OSWorld Example

Figure 8 contrasts the GUI+CLI and GUI-only interfaces on a single OSWorld task—modifying a VS Code setting so the cursor stays focused on the debug console during debugging rather than snapping back to the editor. In GUI+CLI mode the agent uses a few GUI actions to locate the relevant setting, then drops to the CLI

App

p

FP

FN

Prec. (%)

FNR (%)

Chrome GIMP LibreOffice Calc LibreOffice Impress LibreOffice Writer Thunderbird VLC VS Code Audacity Blender Draw.io Godot Kdenlive OBS QGIS Zotero

0.53 0.52 0.56 0.51 0.53 0.59 0.57 0.52 0.52 0.49 0.48 0.46 0.59 0.54 0.57 0.48

0 1 0 0 0 0 0 0 0 0 0 2 1 1 0 0

3 3 1 2 0 0 0 2 0 2 3 2 1 3 2 1

100.0 96.7 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 93.3 96.7 96.7 100.0 100.0

10.0 10.0 3.3 6.7 0.0 0.0 0.0 6.7 0.0 6.7 10.0 6.7 3.3 10.0 6.7 3.3

Overall

—

5

25

99.0

5.2

Table 7: VLM judge vs. human labels, all 16 applications. Three annotators, blind to the judge score, labeled 30 judgeaccepted and 30 judge-rejected trajectories per application. p is the true acceptance rate; FP counts accepted trajectories that were not a human “fully success” (out of 30, all of them partial successes—zero hard failures); FN counts rejected trajectories that were in fact fully successful (out of 30). Precision is on the acceptance set and FNR on the rejection set. Reweighting each application by p gives an overall agreement of 97.0% and Cohen’s κ = 0.94. Blue rows are OSWorld applications; orange rows are CUA-Universe extensions.

to write ”debug.focusEditorOnBreak”: false directly into settings.json and verify that it persists, resolving the task in 5 CLI calls and 12 GUI actions (reward 1). In GUI-only mode it repeatedly searches and clicks through the settings UI but never commits the correct configuration, exhausting the full 60-action budget (reward 0). The case makes the orchestration advantage concrete: the GUI locates the setting, while the CLI performs the precise, verifiable write that the GUI-only agent cannot reliably land.

Split

Episodes

Step records

Total (GB)

CUA-Verse OSWorld

2,526 2,397

140,088 95,320

172.97 113.67

Total

4,923

235,408

286.64

Table 8: Final training splits. “Total (GB)” is the on-disk size including all screenshots; step records reference images by path. Parameter Tuner type LoRA rank LoRA α LoRA dropout LoRA bias Target modules Vision tower Multimodal aligner

Limitations

Our study leaves several directions open. First, CUAUniverse synthesizes single-application tasks: each task is grounded in one application’s shared GUI+CLI state, which is what lets us build verifiable environments and steer efficient hybrid paths at scale. Cross-application workflows, where state is carried across several applications, are a natural extension of the same pipeline rather than a different design, and we leave them to future work. Second, we use the harvested trajectories for supervised fine-tuning only, so the learned policy is bounded by its data-generation backbone; because the pipeline already produces per-task verifiers, using them as rewards for reinforcement learning is a direct next step toward surpassing the teacher. Third, task success

LoRA 8 32 0.05 none all-linear frozen frozen

Table 9: LoRA configuration. is scored by a VLM judge rather than per-task programmatic checks, which may introduce label noise; the judge– human validation in Appendix F bounds this (99.0% acceptance precision, Cohen’s κ = 0.94), and we further mitigate it with a conservative acceptance threshold and by grounding the judge in CLI and exported-artifact evidence, but we do not eliminate it. Finally, application adaptation assumes software that is open-source or scriptable enough to expose a command-line surface, and our environments target desktop Linux; the coding-agent-driven construction is not tied to these choices in principle, but broadening to closed-source or non-desktop platforms remains open.

Parameter

K

Value

Optimizer Learning rate Warmup ratio Per-device train batch size Gradient accumulation Effective batch size (8 GPUs) Epoch argument Max length Max pixels Image max token budget Attention DeepSpeed

Value ms-swift default 1 × 10−4 0.05 2 1 16 3 18,000 602,112 1,024 Flash Attention 2 ZeRO-2

Table 10: Optimization and sequence settings.

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41

[SYSTEM] You are a computer use agent. You can operate the computer through both GUI (mouse/keyboard) and CLI (command-line) tools. You are given a task instruction, a screenshot of the current screen, and your previous interactions. Complete the task by issuing one action per step. Choose the most efficient tool for each step - use CLI for precise operations and GUI for visual tasks. For each step, respond in EXACTLY this format: –thought˝ ## Action: –action˝ ## Code: –code˝ In the code section, use ONE of: - A python block with pyautogui code (GUI action) - A cli block with a shell command (CLI action, only when CLI tools are listed) - A special function in a code block: - –”name”: ”computer.wait”, ”parameters”: –”time”: 3˝˝ - –”name”: ”computer.terminate”, ”parameters”: –”status”: ”success”˝˝ - –”name”: ”computer.terminate”, ”parameters”: –”status”: ”failure”˝˝ CLI Tool: Execute a CLI tool using a cli block. - CLI modifies the file on disk. The GUI may not auto-update. - Always check CLI output (rc and stdout). If rc != 0, diagnose and retry. - You can chain multiple CLI steps before switching to GUI. ## Available CLI Tools (draw.io registry, abridged) drawio˙project˙new — Create a new empty draw.io diagram file drawio˙project˙save — Save the current diagram to disk drawio˙shape˙add — Add a shape at a position (cylinder—rectangle—...) drawio˙shape˙style — Set a style property on a shape (fillColor, ...) drawio˙connect˙add — Add a labeled connector between two shapes drawio˙export˙render — Export the diagram to PNG/PDF/SVG ... (full registry of project/shape/connect/page/export/session commands) [Task guidance for this attempt: First ...] --- Interaction history (elided) --[USER] Instruction: –task˙instruction˝ Previous actions: –action˙history˝ Previous CLI results: –cli˙result˙history˝ –current˙screenshot˝ [ASSISTANT] –thought˝ ## Action: –next˙action˝ ## Code: –python˙or˙cli˙block˝ ... (this GUI/CLI interleaved turn repeats until terminate)

Figure 10: Message format of a data-generation rollout (Kimi K2.5 on a draw.io task). The system prompt exposes GUI and CLI jointly with generic modality guidance (“CLI for precise operations, GUI for visual tasks”); the task guidance field carries the Path-Steer efficiency prior for this attempt (an ordered plan, “First . . . ”), used only during rollouts and empty at evaluation. The interaction history is abstracted with placeholders: each turn supplies the instruction, running action/CLI-result histories, and the current screenshot, and the agent replies with a thought and one GUI or CLI block until it terminates. Task text and coordinates are elided.

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53

[SYSTEM] You are a task completion evaluator for a GUI automation agent operating a desktop application. You will be given: 1. A task instruction describing what the agent should accomplish 2. The execution trajectory showing actions taken and their results 3. Screenshots from the execution in chronological order Your job is to judge whether the task was completed successfully. How to read the screenshots: - The first screenshot is only the earliest available visual baseline. Do not treat it as the result. - Middle screenshots are context for how the state changed over time. - The final screenshot is the primary visual evidence for task completion. - If the trajectory says an edit/export succeeded but the final screenshot or final artifact evidence does not show it, score conservatively. - If the final screenshot is blank, stale, still in a menu/dialog, or appears unchanged from the baseline, do not give high credit for visual tasks. - For file-based editing tasks where the instruction says the exported file/project is the scored artifact, treat successful CLI output and exported artifact evidence as authoritative for file contents. Desktop application screenshots may be stale because many apps do not auto-refresh after external file edits; do not penalize missing GUI refresh when the trajectory provides structured CLI evidence that the exported artifact contains the requested edits. Scoring guidelines: - 1.0: Task fully completed, all requirements met, visual confirmation matches expectations - 0.7-0.9: Task mostly completed, minor issues (e.g. slightly wrong values, visual looks close) - 0.4-0.6: Task partially completed (some steps done, others missing or wrong) - 0.1-0.3: Task barely started or mostly failed - 0.0: Task not completed at all, or agent timed out without meaningful progress When evaluating, consider: - Did CLI commands succeed (rc=0) and produce expected output? - Does the final screenshot show the expected visual result? - Compared with the baseline screenshot, are the requested changes visible in the final screenshot? - Were all sub-tasks in the instruction addressed? - Did the agent reach a terminal state (DONE) or time out? Respond with ONLY a JSON object (no markdown, no extra text): –”score”: ¡float 0.0-1.0¿, ”reason”: ”¡brief explanation of what was/wasn’t completed¿”˝ [USER] # Task Instruction –instruction˝ # Execution Trajectory –trajectory˙evidence˝ # Screenshots The following –n˝ screenshot(s) are sampled from the trajectory and are ordered from earliest to latest. Use the earliest screenshot only as baseline/context. The final 3 screenshots, when present, show the end-state context. Use the FINAL screenshot as the main visual evidence for scoring. –image˙1˝ ... –image˙n˝ # base64 PNG image inputs, each labeled with its chronological role

Figure 11: Full VLM judge prompt. Braces denote per-attempt inputs: {instruction} the task instruction, {trajectory evidence} the compacted action/CLI trace, and {image 1..n} the sampled, role-labeled screenshots supplied as image inputs.

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52

You are a multi-purpose intelligent assistant. Based on my requests, you can use tools to help me complete various tasks. # Tools You have access to the following functions: ¡tools¿ –computer˙use˙function˙schema˝ –execute˙cli˙function˙schema˝ ¡/tools¿ If you choose to call a function ONLY reply in the following format with NO suffix: ¡tool˙call¿ ¡function=example˙function˙name¿ ¡parameter=example˙parameter˙1¿ value˙1 ¡/parameter¿ ¡parameter=example˙parameter˙2¿ This is the value for the second parameter that can span multiple lines ¡/parameter¿ ¡/function¿ ¡/tool˙call¿ ¡IMPORTANT¿ Reminder: - Function calls MUST follow the specified format: an inner ¡function=...¿¡/function¿ block must be nested within ¡tool˙call¿¡/tool˙call¿ XML tags - Required parameters MUST be specified - You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after - If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls - The current date is –runtime˙date˝. - Collapsed screenshots appear as text: This screenshot has been collapsed. ¡/IMPORTANT¿ # Response format Response format for every step: 1) Action: a short imperative describing what to do. 2) One or more ¡tool˙call¿...¡/tool˙call¿ blocks, each containing one tool call. Rules: - Output exactly in the order: Action, then the ¡tool˙call¿ block(s). - If multiple ¡tool˙call¿ blocks are needed, they will be executed sequentially in the order you output them. - Prefer one ¡tool˙call¿ for normal steps; use multiple only for tightly coupled low-level actions such as click-then-type or key sequences. - Use ¡function=execute˙cli¿ for CLI commands. Use ¡function=computer˙use¿ only for GUI actions. - Do not mix execute˙cli with other tool calls in the same step unless the command is immediately required by the same atomic action. - Be brief: one sentence for Action. - Do not output anything else outside those parts. - If finishing, use action=terminate in the tool call. # CLI mode rules - This evaluation is in CLI mode, meaning the separate execute˙cli function is available in addition to normal GUI actions. - For file-based inspection, edits, saves, exports, or verification covered by the listed CLI tools, call the separate ¡function=execute˙cli¿ tool before GUI interaction. - Use GUI actions for visual editing/selection when they are more natural or when no listed CLI tool covers the operation. - Do not use application menus, file pickers, or a different application for an operation that a listed CLI command can perform directly. - Do not run registry tool names as shell commands. Use the concrete command form shown in the command parameter description. - After a successful CLI edit/save/export/verification, use that CLI feedback to decide whether to terminate instead of repeating GUI confirmation loops. - Use wait when the environment needs time after either GUI or CLI actions. - Use terminate only after the requested result is complete. - Never output placeholder coordinates such as [0, 0].

Figure 12: Representative evaluation actor system prompt (Qwen, CLI+GUI); each backbone has its own prompt and a GUImode counterpart following the OSWorld protocol. It is application-agnostic; per-application capability enters only through the task-specific command registry injected into the execute cli schema.

Record · ID 660846 · SHA-256 5bc4e72a56a7c199
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.