CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents Haoting Shi1 * , Wenhao Wang2 * † , Weicheng Fang2 , Yaozhong Liang2 , Tian Jin1 , Pengxiang Zhao2 , Guangyi Liu2 , Siheng Chen1† , Yanfeng Wang1 1
Shanghai Jiao Tong University
arXiv:2609.05374v1 [cs.AI] 4 Sep 2026
Abstract Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly through the GUI, often producing inefficient trajectories. Real-world computer work is hybrid, combining visual-state inspection with precise, high-throughput command-line operations, so capable agents must coordinate both modalities over shared application state. Yet scalable hybrid environments remain scarce because supporting both GUI and CLI over real applications typically requires substantial manual engineering for each application. Existing agents also struggle to use the two interfaces complementarily: CLI-native agents lack visual perception for tasks involving interface state or layout, while GUI-native agents are inefficient for operations better executed through commands. We introduce CUA-Universe, a scalable environment-todata pipeline that turns real desktop software into hybrid GUI+CLI environments. App-Forge adapts applications into reproducible VMs and command-line surfaces it discovers, wraps, or generates, scaling to 16 applications across diverse domains; Task-Weave synthesizes diverse hybrid tasks of controllable difficulty from reusable operations over seed files, turning each environment into a continuous task source; and Path-Steer rolls agents out along efficient hybrid paths and harvests verified trajectories for post-training. Training on this data shifts behavior from inefficient GUI interaction and brittle CLI scripting toward effective GUI+CLI orchestration, improving both success and efficiency for our 9B model on CUA-Verse (Score +39.3 pts; −37% steps, −60% tokens), OSWorld (SR +16.8 pts; −57% steps, −44% tokens), and OSWorld-MCP (Score +7.84 pts; −27% steps, −30% tokens). By converting real desktop software into hybrid environments and reusable training data, CUA-Universe provides a scalable path toward more capable and efficient computeruse agents. We will release our code and data.
1
Introduction
Computer-use agents (CUAs) have advanced rapidly, completing a broad range of real desktop and mobile tasks on benchmarks such as OSWorld and AndroidWorld (Xie et al. 2024; Rawles et al. 2025). Yet leading agents still interact predominantly through the graphical user interface (GUI), resulting in trajectories that are often unnecessarily * These authors contributed equally. †
Corresponding authors.
2
Zhejiang University
long (Abhyankar, Qi, and Zhang 2025) and increasingly brittle on extended workflows (Yuan et al. 2026). Real computer use, however, is inherently multi-modal: users rely on the GUI for visually grounded interaction, while using command-line tools, scripts, or APIs for precise and highthroughput operations. Capable CUAs should therefore operate in hybrid GUI+CLI environments, dynamically choosing the interface best suited to each subtask while maintaining a shared application state across modalities (Song et al. 2025; Yang et al. 2025; Jia et al. 2025). Realizing such hybrid computer-use agents faces two coupled bottlenecks: scalable real-world environments and cross-modality orchestration. On the environment side, GUI-centric environments are costly to build and often require substantial manual engineering, while CLI-centric environments are easier to automate but lack access to visual layout and interface state. Building hybrid environments that support both modalities over the same application state therefore remains expensive and difficult to scale across applications. On the agent side, existing agents are typically optimized for one modality and struggle to use the two interfaces complementarily. CLI-native agents often lack visual perception and therefore resort to brittle scripts for tasks that depend on interface state or visual layout, while GUI-native agents can be inefficient for batch operations that a single command could complete. The deeper challenge is therefore orchestration: deciding when to switch interfaces and carrying task state across them, for example, locating a target through the GUI and then processing it through the CLI. To address these bottlenecks, we introduce CUAUniverse, a real-software environment framework for synthesizing, evaluating, and training hybrid GUI+CLI computer-use agents. CUA-Universe organizes real software into a scalable environment-to-data pipeline with three components. App-Forge adapts a desktop application into a reproducible VM and exposes it through a command-line surface it discovers, wraps, or generates. Driven by a coding agent rather than per-application manual engineering, it scales CUA-Universe to 16 desktop applications across diverse domains. Task-Weave synthesizes GUI+CLI hybrid tasks of controllable difficulty from applications, tools, and seed states, turning each environment into a continuous source of tasks. Path-Steer guides agents toward efficient hybrid execution paths, using the CLI for batch and
Scale: 16 Real Desktop Applications
Closed-loop Environment-to-Data Pipeline 1. App-Forge Environments
Creative & Interactive
Impact: Better and Cheaper CUA Agents Interaction Behavior GUI-Only: Over-clicking
Blender
Godot
Draw.io
…
GIMP
Installer Agent
Office & Knowledge
Reproducible VM + Metadata
Tool Construction
CLI-Only: Over-scripting
…
2. Task-Weave Task Synthesis Zotero
LibreOffice Writer
LibreOffice Calc
GUI+CLI Hybrid: Efficient orchestration
LibreOffice Impress
…
Media & Playback Task Instantiation
Operation Abstraction Kdenlive
OBS Studio
Audacity
VLC
Agentic Task Refinement
Empirical Gains: Higher Success Rate, Lower Cost
3. Path-Steer Hybrid Rollout
Technical & Web
QGIS
VS Code
Thunder bird
Chrome
Hybrid Interface
Path Steering Rollouts
Traj Harvesting
Figure 1: Overview of CUA-Universe. (1) Scale: 16 real desktop applications spanning creative & interactive, office & knowledge, media & playback, and technical & web domains, each exposed through both a GUI and application-specific CLI interfaces. (2) Closed-loop environment-to-data pipeline: App-Forge adapts an application into a reproducible VM and the command-line surface it discovers, wraps, or generates (installer agent → reproducible VM + metadata → tool construction); Task-Weave synthesizes diverse hybrid GUI+CLI tasks at scale over seed files (operation abstraction → task instantiation → agentic refinement); and Path-Steer rolls agents out onto efficient hybrid paths and harvests high-quality, directly reusable trajectories for post-training. (3) Impact: training on this data shifts interaction behavior from inefficient GUI interaction and brittle CLI scripting toward effective GUI+CLI orchestration, producing more capable and efficient agents with consistent improvements across diverse benchmarks, including CUA-Verse (Score +39.3 pts; −37% steps, −60% tokens), OSWorld (SR +16.8 pts; −57% steps, −44% tokens), and OSWorld-MCP (Score +7.84 pts; −27% steps, −30% tokens). precise operations and the GUI for visually grounded ones, and harvests high-quality trajectories for post-training. Together, these components turn real software into scalable sources of hybrid tasks and training data, enabling agents to learn more effective cross-modality orchestration between GUI and CLI. Our evaluation spans three complementary axes. First, to directly evaluate a model’s ability to orchestrate GUI and CLI actions over a shared application state, we construct CUA-Verse, a held-out benchmark of 160 hybrid tasks built on the eight desktop applications introduced by CUA-Universe, with evaluation tasks disjoint from the training data. On CUA-Verse, our model improves success by roughly 3× over its identical base while using 60% fewer tokens, achieving the best open-source result. Second, we test whether this capability transfers beyond our benchmark to OSWorld, where our model gains +16.8 success-rate points over GUI-only execution while using fewer steps and tokens. Finally, on OSWorld-MCP (Jia et al. 2025), our model generalizes to a different tool-invocation interface unseen during training, improving Score by +7.84 points over its base. Together, these results show that CUA-Universe improves not only performance within the environments it constructs, but also hybrid interaction skills that transfer across tasks, benchmarks, and tool interfaces. We summarize our contributions as follows: • A scalable environment-to-data pipeline. CUAUniverse turns real desktop software into hybrid
GUI+CLI agent environments through three components, App-Forge, Task-Weave, and Path-Steer, which cover application adaptation, task synthesis, and trajectory generation. Agent-driven construction allows the framework to scale efficiently to 16 real applications across diverse domains. • Scalable hybrid task synthesis and efficient trajectory generation. The framework discovers, wraps, or generates application-specific command-line interfaces, synthesizes GUI+CLI tasks of controllable difficulty from reusable operations and seed states, and steers agents toward efficient hybrid execution paths. This turns each environment into a continuous source of tasks and training trajectories while encouraging more effective crossmodality orchestration. • A benchmark for hybrid GUI+CLI orchestration. CUA-Verse provides 160 held-out hybrid tasks across eight desktop applications introduced by CUA-Universe, directly evaluating an agent’s ability to coordinate GUI and CLI actions over a shared application state. • Strong performance and transfer. Our 9B model trained on CUA-Universe data achieves roughly 3× the success of its base model on CUA-Verse while using 60% fewer tokens. The learned hybrid interaction skills further transfer to OSWorld with a +16.8 point gain over GUI-only execution and generalize to the unseen tool interface of OSWorld-MCP.
2 2.1
Related Work
Environment Synthesis for Computer-Use Agents
A growing body of work automatically synthesizes environments and trajectories to avoid the cost of hand-curated benchmarks (Zhou et al. 2024; Deng et al. 2023; Xie et al. 2024; Rawles et al. 2025). InfiniteWeb (Zhang et al. 2026), GUI-Genesis (Cao et al. 2026), and AutoWebWorld (Wu et al. 2026) generate functional web environments for posttraining, but their targets are artificial web pages confined to the GUI modality, leaving a sim-to-real gap on actual software. Scaling further, CUA-Gym (Wang et al. 2026a) cogenerates environments, tasks, and verifiable rewards across desktop and mock web applications for RLVR, and GymAnything (Aggarwal, Neubig, and Welleck 2026) turns arbitrary applications into agent environments and distills successful trajectories into a model. These environments, however, are driven purely through the GUI, so the trajectories they yield are inherently single-modality and cannot exhibit when to leave the GUI. In contrast, CUA-Universe builds each application into a hybrid GUI+CLI environment by exposing it through both the GUI and application-specific CLI tools. The synthesized tasks therefore require coordinating the two modalities over a shared application state, while the harvested trajectories carry efficiency and orchestration signals that a GUI-only pipeline structurally cannot provide.
2.2
Single-Modality GUI and CLI Agents
Computer-use agents have progressed along two separate lines. On the GUI side, agents are increasingly evaluated across diverse desktop, mobile, and heterogeneous platform settings, including OSWorld (Xie et al. 2024), AndroidWorld (Rawles et al. 2025), and FedGUI (Wang et al. 2025, 2026b), but may remain inefficient, taking far longer trajectories than necessary (Abhyankar, Qi, and Zhang 2025) and degrading on long-horizon, cross-application workflows (Yuan et al. 2026). On the CLI side, the terminal has become a first-class target: Terminal-Bench (Merrill et al. 2026) and TerminalWorld (Chu et al. 2026) provide hard command-line tasks, CLI-Universe (Hua et al. 2026) synthesizes verifiable terminal tasks, and further work scales terminal training environments and recipes (Cheng et al. 2026; Ivison et al. 2026). Yet competence in one modality does not transfer to the other: terminal agents cannot handle operations that depend on visual state, while GUI agents fall back to slow, element-by-element manipulation for batch operations. CUA-Universe instead trains on GUI+CLI hybrid tasks over a shared application state, teaching agents when to switch modalities, a form of coordination that neither line develops in isolation and that improves both success rate and execution efficiency.
2.3
Toward Hybrid GUI+CLI Agents
A recent trend combines visual actions with CLI tool calls rather than relying on either alone: CoAct-1 (Song et al. 2025) pairs a GUI operator with a coding agent, and UltraCUA (Yang et al. 2025) trains a hybrid-action model whose tools are mined from generic documentation and code
repositories over a fixed task set. Hybrid benchmarks further quantify the GUI+CLI efficiency trade-off (Li et al. 2026; Zhou et al. 2026; Fu et al. 2026). Unlike these existing works, which train a hybrid agent over a fixed set of tools and tasks, CUA-Universe contributes a new way to construct environments and synthesize data. Rather than mining a generic tool set, it grounds the tool layer in each real desktop application. Inspired by CLI-Anything (Yang, Fan, and Huang 2026), we generate agent-native CLIs when needed while also exposing applications’ native commandline tools, turning each environment into a continuous source of tasks and trajectories. On top of this, Path-Steer steers rollouts onto efficient hybrid paths, raising success rates and yielding higher-quality trajectories for computeruse agent post-training.
3 3.1
Method
Overview and Problem Formulation
Overview. As illustrated in Figure 1, CUA-Universe turns real desktop applications into hybrid GUI+CLI environments and further converts them into scalable sources of tasks and training trajectories. The framework consists of three components. App-Forge (§3.2) scales environment construction by adapting applications into reproducible hybrid environments with programmatic tool surfaces. TaskWeave (§3.3) scales task generation by composing diverse hybrid tasks from reusable operations grounded in real application states. Path-Steer (§3.4) scales trajectory collection by steering agents toward efficient GUI+CLI execution paths and harvesting rollouts for post-training. Problem formulation. We model a hybrid task as a POMDP: at step t the agent observes ot (a screenshot plus optional textual CLI returns) and emits at ∈ Agui ∪ Acli , where both action spaces read and modify over a shared persistent application state st . A task is a tuple τ = (instr, s0 , V ) of an instruction, a seed initial state, and a verifier V (ζ) ∈ [0, 1] scoring a trajectory ζ (a VLM judge); the pipeline synthesizes such tasks at scale and collects trajectories that solve them along efficient hybrid paths.
3.2
App-Forge: Scalable Agentic Environment Construction
Scaling hybrid environments across real desktop applications faces two application-specific bottlenecks: environment adaptation, since applications differ substantially in installation, configuration, and runtime dependencies; and tool construction, since agents need usable interfaces across heterogeneous applications. App-Forge addresses both through a scalable agentic construction pipeline that produces reproducible application environments together with CLI surfaces aligned with their GUIs. Application adaptation. The first bottleneck is reproducibly adapting diverse desktop software, whose installation procedures, dependencies, and launch configurations vary substantially across applications. An installer agent, guided by an installation skill, operates a persistent VM harness to install and configure the target application and
its supporting command-line utilities, interactively diagnosing failures until successful launch is verified. The resulting setup is distilled into a reproducible configuration, while lightweight introspection extracts application metadata for later task grounding and verification. The same adaptation workflow is reused across applications, allowing CUAUniverse to scale to 16 desktop applications across diverse domains. Eight are inherited from OSWorld but were originally GUI-only, for which App-Forge adds a shared-state CLI layer, while the other eight are introduced by CUAUniverse (Appendix B). Tool construction. The second bottleneck is constructing an expressive CLI surface across heterogeneous applications. We draw on three sources: native command-line tools we discover (e.g., blender --python-expr, cvlc), scripting APIs we wrap (e.g., bpy, LibreOffice UNO, and GIMP Script-Fu), and agent-native CLIs we generate when existing interfaces are insufficient, inspired by CLI-Anything. This layered design supports diverse automation surfaces without requiring a manually designed CLI for every application. GUI and CLI operate over the same project state, with lightweight adapters resynchronizing stale GUI views after external CLI edits. Appendix C lists the complete tool inventory for all applications.
3.3
Task-Weave: Compositional Hybrid Task Synthesis
Scaling task generation across real applications requires synthesizing tasks that reflect executable capabilities, cover diverse compositions, and remain feasible in the live environment. Task-Weave addresses these requirements in three stages: it abstracts reusable operations from agent exploration, composes them into diverse seed-conditioned tasks, and validates the resulting tasks through real execution. Operation abstraction. We first build a reusable operation pool that captures what can be reliably performed in each application. For a family of exploration seeds, we run parallel exploration agents with diverse goals, each targeting a different facet of the application, such as structure and visibility, appearance, or export and organization. Each agent interacts with the GUI while recording screenshots and actions. We slide a window over each trajectory and prompt an LLM to abstract the interaction into a high-level reusable operation, such as export scene to gltf rather than click. Each operation is also annotated with its execution procedure, which can later inform rollout guidance. We filter trivial or non-reusable candidates, then deduplicate, cluster, and aggregate operations across runs into a global operation pool with supporting evidence. Task instantiation. We then compose operations into diverse tasks grounded in concrete application states. Each task is conditioned on a real seed project, such as a .blend scene or .odp deck, together with its metadata, which defines the initial state s0 and the objects available for manipulation. We sample and score candidate operation chains and retain only compositions that are meaningful for the target seed. Difficulty is controlled by chain length and com-
position, ranging from single-operation edits to multi-step hybrid workflows. Each selected chain is compiled with its seed, supporting evidence, and mapped CLI tools into a task package containing an instruction, initial state s0 , verifier, and guidance. The instruction is synthesized as a natural user goal rather than a sequence of low-level actions. Agentic task refinement. Finally, we ground synthesized tasks in real execution before retaining them. A ReAct-style review agent (Yao et al. 2022) launches the application and performs a short multimodal interaction to check whether the instruction is feasible, unambiguous, and not already satisfied by the seed. When needed, it revises the instruction and guidance based on execution feedback. Valid tasks are retained, fixable tasks are revised, and invalid tasks are discarded,reducing hallucinated or infeasible task specifications before rollout.
3.4
Path-Steer: Efficiency-Aware Hybrid Rollout
Path-Steer converts synthesized tasks into efficient, verified training trajectories through three stages: hybrid execution enables GUI and CLI actions within a shared trajectory, efficient-path steering guides agents toward appropriate modality choices, and trajectory harvesting retains highquality rollouts for post-training. Each task is attempted multiple times from fresh environment instances, with attempts executed in parallel for throughput. Hybrid execution interface. We first provide agents with a unified interface for flexibly interleaving GUI and CLI actions. At each step, the agent emits either a GUI action or a CLI action. CLI actions are parsed against the application’s tool registry, expanded into concrete commands with the current working-file path injected when needed, and executed in the VM. Their return code and truncated output are included in the next observation, allowing both modalities to operate seamlessly within one trajectory over the shared project state maintained by application adapters (§3.2). Efficient-path steering. We then guide rollouts toward more efficient and appropriate modality choices. From each task’s operation chain, we derive a hybrid execution prior indicating when an operation is better suited to the CLI, such as batch, precise, or high-throughput operations, or to the GUI, such as operations depending on visual layout or interface state. The prior provides lightweight modality-level guidance without specifying low-level actions, reducing inefficient GUI interaction and brittle CLI scripting while yielding shorter and less redundant trajectories. Scoring and trajectory harvesting. Finally, we score each completed rollout with the task verifier V (ζ) ∈ [0, 1], implemented as a VLM judge over the trajectory. Verified rollouts are serialized into step-level and trajectory-level records containing observations, reasoning, GUI and CLI actions with their returns, screenshots, and final scores. We retain high-scoring hybrid trajectories as directly reusable supervision for subsequent post-training.
Training-data Composition 2,526 episodes 51.3%
Episodes
2,397 episodes 48.7%
Blender 191 Episodes
GIMP 240 Episodes
Zotero 196 Episodes
VS Code 255 Episodes
Draw.io 216 Episodes
2,526 140,088 steps 59.5%
Training steps
0%
Episodes
95,320 steps 40.5%
25%
50%
CUA-Universe extensions
75%
100%
LibreOffice Impress 200 Episodes
2,397
QGIS 218 Episodes
VLC 290 Episodes
Episodes
Audacity 376 Episodes
LibreOffice Calc 339 Episodes
Godot 418 Episodes
LibreOffice Writer 354 Episodes
OBS 445 Episodes
Chrome 363 Episodes
Kdenlive 466 Episodes
Thunderbird 356 Episodes
OSWorld applications
(a) Overall distribution of episodes and training steps.
(b) Task distribution across CUA-Universe extension applications.
(c) Task distribution across OSWorld applications.
Figure 2: Training-data composition. Our pipeline scales to diverse applications and produces substantial training data, with 4,923 episodes and approximately 235K training steps across all 16 applications. The balanced coverage of CUA-Universe extensions and OSWorld applications further demonstrates the diversity and scalability of the generated data.
4
Experiments
Training setup. App installation and CLI-tool construction (App-Forge) are driven by a Codex coding agent (GPT5.6), while operation abstraction and hybrid task synthesis (Task-Weave) are performed by Kimi K2.5 (Team et al. 2026); task success throughout the pipeline is scored by a VLM judge (GPT-5.4 (OpenAI 2026)). Our training data is generated by rolling out the same Kimi K2.5 backbone under Path-Steer and keeping trajectories that pass the judge at a score threshold of 0.75, yielding 4,923 verified episodes (∼235K step-level records; Figure 2). We then fine-tune Qwen3.5-9B (Qwen Team 2026) with LoRA on these steplevel trajectories for 3 epochs using the ms-swift (Zhao et al. 2024) framework, on 8× A100 GPUs in roughly two days; full training details are in Appendix H.
4.1
∼3× with 37% fewer steps and 60% fewer tokens (255K vs. 643K per episode), yielding the best accuracy–cost tradeoff among open models. (3) Capability is structured: Ours is strongest on audio/video apps (Audacity 0.815, OBS 0.605) and weaker on 3D/spatial ones (Blender 0.398, Godot 0.460); among open 8–9B models it dominates EvoCUA-8B (0.330 avg.) everywhere except Blender (0.495), indicating headroom in 3D data coverage. Kimi K2.5
0.660
0.500
0.425
0.400
0.664
0.451
0.535
0.538
Seed2.1 Pro
0.483
0.608
0.430
0.414
0.421
0.820
0.868
0.749
1.0
GPT-5.5
0.748
0.573
0.682
0.708
0.730
0.928
0.890
0.882
EvoCUA-8B
0.495
0.160
0.293
0.285
0.319
0.405
0.394
0.285
Qwen3.5-9B
0.205
0.364
0.353
0.140
0.040
0.100
0.165
0.145
Ours
0.398
0.640
0.470
0.460
0.720
0.547
0.605
0.815
ro
ot
S GI
ve nli
dio
Evaluation on CUA-Verse
Setting. We evaluate on CUA-Verse, a held-out benchmark of 160 hybrid GUI+CLI tasks (eight professional desktop applications × 20 tasks) synthesized by the CUAUniverse pipeline. Its tasks are disjoint from all training data, though the eight applications themselves are indomain; it measures an agent’s ability to complete GUI+CLI hybrid tasks, exposing both a GUI and a CLI interface so the agent must coordinate the two over a shared application state to solve each task. The eight applications (Blender, Draw.io, Zotero, Godot, QGIS, Kdenlive, OBS, Audacity) are each scored by a unified VLM judge using GPT-5.4 (OpenAI 2026); we additionally report steps and per-episode token as efficiency metrics. We compare against three proprietary models (Kimi K2.5, Seed2.1 Pro, GPT-5.5) and three opensource models (Qwen3.5-9B, EvoCUA-8B, and Ours). Results. Three observations stand out (Table 1, Fig. 3). (1) Distillation lifts a 9B model to closed-source level: Ours reaches 0.582 Score, the best among open-source models— surpassing Kimi K2.5 (0.522) and trailing only the proprietary Seed2.1 Pro (0.599) and GPT-5.5 (0.768). Since it is trained only on the teacher’s successful trajectories, it learns the upper tail and exceeds the teacher’s mean on CUA-Verse, consistent with STaR (Zelikman et al. 2022) and ReST (Gulcehre et al. 2023). (2) The gains are clean and efficient: against the identical Qwen3.5-9B (0.189), Score improves
Score
d
er
en Bl
Dr
.io aw
Zo
te
d
Go
Q
e Kd
tu
OB
S S-
y cit da Au
0.5
0.0
Figure 3: Application-level performance on the eight CUA-Verse applications. Each cell is the score for one model–application pair; darker shading is better. The horizontal rule separates proprietary (top) from open-source (bottom) models. Ours surpasses its Qwen3.5-9B on all eight applications, and is strongest on audio/video apps (Audacity, OBS).
4.2
Transfer to OSWorld
Setting. We evaluate whether training on CUA-Universe improves general computer-use capability even under GUIonly execution, and whether providing CLI access yields further gains through learned hybrid orchestration. We evaluate on OSWorld under a controlled 244-task protocol that excludes the os and multi-app splits and scores every task with the official verifier; SR is the fraction of tasks with reward 1. We compare two action interfaces under identical task text, environment setup, and verifier: in the GUI setting the agent acts purely through screen coordinates, and in the GUI+CLI setting it additionally receives a native execute cli tool with per-application command definitions and a single injected sentence giving the task’s input-file path.
Model
Mode
Kimi K2.5 Seed2.1 Pro GPT-5.5 Qwen3.5-9B EvoCUA-8B Ours
GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI GUI GUI+CLI
Score
CUA-Verse Steps Token (K)
0.522
21.1
275
0.599
40.1
449
0.768
23.1
193
0.189
56.2
643
0.330
41.7
377
0.582
35.2
255
SR (%) 53.7 54.5 55.7 59.0 66.8 68.0 21.7 24.6 40.6 43.3 23.4 40.2
Steps ↓ 27.2 28.9 28.4 27.7 14.7 19.0 58.6 54.1 34.8 42.1 39.6 28.6
OSWorld Step Gain ↑ Token(K) ↓ 224.8 1.15 277.3 188.0 1.10 215.6 595.0 1.55 550.3 521.9 2.63 596.1 325.6 1.58 408.0 325.7 2.35 286.5
Token Gain ↑ 1.09 0.94 1.58 2.03 1.35 1.79
Table 1: Combined performance on CUA-Verse and OSWorld. CUA-Verse: 8 apps × 20 tasks; Score is the average VLM judge score, and Token is the average number of tokens per episode in thousands. OSWorld: 244-task controlled scope excluding os and multi-app tasks, comparing GUI and GUI+CLI execution. Token/task is reported in thousands. Step Gain and Token Gain are the mean per-task GUI-to-GUI+CLI cost ratios over jointly solved tasks. A gain of k× indicates that GUI+CLI uses k times fewer steps or tokens on the same successfully solved tasks. Kimi K2.5
Seed2.1 Pro
GPT-5.5
EvoCUA-8B
Qwen3.5-9B Family
GUI
CLI
GUI
CLI
GUI
CLI
GUI
CLI
GUI
CLI
Ours
Chrome
0.261
0.261
0.283
0.304
0.326
0.391
0.283
0.239
0.239
0.239
0.370
GIMP
0.692
0.808
0.462
0.615
0.538
0.423
0.692
0.615
0.192
0.269
0.423
LO-Calc
0.532
0.553
0.617
0.702
0.830
0.915
0.298
0.191
0.085
0.085
0.191
LO-Impress
0.574
0.638
0.681
0.638
0.745
0.723
0.404
0.340
0.255
0.298
0.319
LO-Writer
0.652
0.609
0.696
0.783
0.826
0.783
0.304
0.304
0.304
0.348
0.522
Thunderbird
0.800
0.667
0.533
0.467
0.800
0.800
0.733
0.467
0.267
0.333
0.733
VLC
0.471
0.412
0.706
0.765
0.706
0.765
0.294
0.353
0.235
0.294
0.471
VS-Code
0.609
0.565
0.609
0.565
0.739
0.739
0.522
0.435
0.261
0.261
0.435
Model
SR 1.0
Interface metrics. Because GUI and GUI+CLI are evaluated on the same 244 tasks, we report paired measures of both success and efficiency gains from adding CLI access. Success gain is the net number of tasks newly solved by GUI+CLI, computed as (SRGUI+CLI − SRGUI ) × 244, and is reported directly from the SR results. Step Gain and Token Gain are computed only on tasks solved by both interfaces, using the mean per-task ratio of GUI to GUI+CLI cost for steps and tokens, respectively. A value of k× means that GUI+CLI reaches the same successful outcome with k times fewer steps or tokens, while averaging per-task ratios prevents a few long trajectories from dominating the metric. Results and analysis. Table 1 and Figure 4 summarize the comparison, from which we draw five findings. (1) Task-
Tok↓
Kimi K2.5 35.25 30.33 27.52 Seed2.1 Pro† 26.23 22.54 40.62 GPT-5.5† 29.92 25.00 27.06
85.21 99.30 54.62
0.5
Qwen3.5-9B 20.90 10.66 37.22 125.68 Ours 28.69 23.36 27.25 87.95 0.0
Figure 4: Per-application OSWorld scores across models and interfaces (GUI vs. GUI+CLI). Cells are heat-shaded by score. Adding the CLI yields Ours (rightmost) its largest gains, converting +41 previously-failed GUI tasks into successes.
All agents share a 60-step budget; observation history is each baseline’s default, while our Qwen-family models use a 3frame history with at most 4 images per step. Steps is the mean number of model decision calls per trajectory and Token/task the mean input+output tokens per task (thousands). Full harness and hardware details are in Appendix G.
SR↑ TIR↑ ACS↓ †
Table 2: Generalization on OSWorldMCP (244-task subset, exclude os/multi apps). ↑ higher / ↓ lower is better; Tokens in M. † : unified prompt, reference value.
specific distillation lifts a 9B open model to near openSOTA at the lowest cost: with GUI+CLI our model reaches 40.2% SR—up +16.8 points from its GUI-only 23.4% and within 3 points of EvoCUA-8B (43.3%)—while spending the fewest tokens per task (286.5k) and the fewest steps (28.6) of any agent, and more than doubling the untuned Qwen3.5-9B (24.6%) under the same interface. (2) The CLI interface is where our model’s gain concentrates—by far the largest net solved-task gain of any agent: adding CLI converts +41 previously-failed GUI tasks into successes (23.4 → 40.2%), against only +3 to +8 for every other model, none of which gains more than eight tasks. (3) For strong closed backbones the interface alone helps only marginally: GPT-5.5 (+3), Seed2.1 (+8) and EvoCUA-8B (+7) improve modestly and Kimi K2.5 is essentially flat (+2)—exposing a CLI the model was not trained to exploit yields little without task-specific tuning. (4) When CLI helps, it is also cheaper: on jointly-solved tasks the CLI interface uses fewer steps for every model (1.10–2.63×) and fewer tokens for most (up to 2.03×); our model reaches the same successes with 2.35× fewer steps. The lone exception
Group
Accept Rate (≥ 0.75)
Mean score
Avg steps
Avg tokens
Avg cost
0.51 0.44 0.54 0.45
0.71 0.63 0.75 0.67
22.75 26.68 25.53 28.78
331,988 385,107 263,500 303,653
$0.26 $0.31 $0.29 $0.33
Kimi K2.5 w/ Path-Steer Kimi K2.5 w/o Path-Steer Seed2.1 Pro w/ Path-Steer Seed2.1 Pro w/o Path-Steer
Table 3: Rollout efficiency of Path-Steer (enabled vs. disabled), for the Kimi K2.5 data-generation backbone and Seed2.1 Pro (a cross-backbone check). Accept Rate is the fraction of the 320 tasks scoring ≥ 0.75; tokens are averaged per task. is Seed2.1, whose sub-1× token efficiency shows the CLI interface can trade extra tokens for its accuracy gain. (5) Overall ranking: among closed agents GPT-5.5 (68.0%) > Seed2.1 (59.0%) > Kimi K2.5 (54.5%); among open models EvoCUA-8B (43.3%) leads with our model (40.2%) a close second, both far above the base.
4.3
Generalization to OSWorld-MCP
Setting. We evaluate whether CUA-Universe teaches transferable cross-modality orchestration, rather than benchmark-specific interaction patterns, on the held-out OSWorld-MCP benchmark (Jia et al. 2025), which augments OSWorld with 158 MCP tools and allows agents to freely combine GUI actions and tool calls. We report Score (task accuracy), Strict SR (perfect-score rate), TIR (accuracy of tool-use decisions), and ACS (average completion steps, lower is better), together with tokens per run as an additional efficiency metric. Evaluation uses 244 tasks after excluding os and multi apps, including 159 tool-beneficial and 85 non-tool-beneficial tasks, with max steps = 50 and history n = 3. We compare against Kimi K2.5, Seed2.1 Pro, GPT-5.5, and Qwen3.5-9B. Results. Three observations stand out (Table 2). (1) Hybrid interaction skills transfer to an unseen MCP interface: our model is trained only with application-specific CLI tools in CUA-Universe and never sees the MCP action space, yet it improves SR from 20.90% to 28.69% (+7.79 points) and more than doubles TIR from 10.66% to 23.36% over the base, showing that the learned tooluse behavior transfers beyond the training format. (2) The transferred capability is competitive with substantially larger models: a single 9B model surpasses Seed2.1 Pro on Score (29.51% vs. 27.03%) and approaches GPT-5.5 and Kimi K2.5, narrowing the gap to far larger closed-source models. (3) Generalization remains efficient: Ours reaches an ACS of 27.25 with 87.95M tokens, reducing steps by 27% and tokens by 30% relative to its base. These results show that the orchestration capability learned in CUAUniverse transfers to MCP without sacrificing efficiency.
4.4
Rollout Efficiency
Setting. We evaluate Path-Steer during data generation to measure whether explicit modality guidance improves both rollout quality and efficiency. On a fixed set of 320 synthesized tasks, we compare rollouts with and without Path-Steer while keeping the hybrid GUI+CLI interface unchanged. We use Kimi K2.5, our data-generation backbone, and addition-
ally evaluate Seed2.1 Pro to test whether the effect generalizes across models. We report Accept Rate, the fraction of trajectories scoring at least 0.75, and Mean Score for trajectory quality, together with average steps, tokens, and estimated per-task cost for efficiency. Since both settings have identical CLI access, this comparison isolates the contribution of steering itself. Results. Three observations stand out (Table 3). (1) PathSteer improves trajectory quality and efficiency simultaneously: on Kimi K2.5, it raises Accept Rate from 0.44 to 0.51 and Mean Score from 0.63 to 0.71, while reducing steps from 26.68 to 22.75, tokens from 385K to 332K, and per-task cost from $0.31 to $0.26. This indicates that explicit modality guidance produces shorter and higher-quality hybrid trajectories by reducing inefficient GUI interaction and brittle CLI scripting. (2) The effect generalizes across backbones: Seed2.1 Pro shows the same pattern, with Accept Rate improving from 0.45 to 0.54 and Mean Score from 0.67 to 0.75, together with lower steps, tokens, and cost. (3) Steering, rather than CLI access alone, drives the gains: the w/o Path-Steer baseline already has access to the same CLI tools, so the consistent improvements isolate the contribution of modality steering itself. Step-by-step comparisons between the two settings are provided in Appendix E.2.
5
Conclusion
We presented CUA-Universe as a step toward a different way of scaling computer-use agents: scaling the environments and interaction spaces from which agents learn, rather than relying only on larger models or more GUI-only trajectories. By turning real desktop software into shared-state GUI+CLI environments, CUA-Universe provides a scalable source of hybrid tasks and trajectories that teach agents not only how to act, but how to orchestrate complementary interfaces efficiently. The resulting 9B model shows that this capability is learnable and transferable, with strong gains on CUA-Verse and OSWorld and further generalization to the unseen tool interface of OSWorld-MCP. More broadly, our results suggest that hybrid environment construction can become a new axis for training computer-use agents, where each additional application expands the space of tasks, tools, and interaction strategies available for learning. We hope this shifts CUA development from collecting increasingly large amounts of single-modality behavior toward building scalable environments that continuously generate richer supervision for more capable, efficient, and general computeruse agents.
References Abhyankar, R.; Qi, Q.; and Zhang, Y. 2025. Osworldhuman: Benchmarking the efficiency of computer-use agents. arXiv preprint arXiv:2506.16042. Aggarwal, P.; Neubig, G.; and Welleck, S. 2026. Gymanything: Turn any software into an agent environment. arXiv preprint arXiv:2604.06126. Cao, Y.; Ran, D.; Wu, M.; Guo, Y.; Chen, X.; Li, A.; Cao, G.; Zhi, G.; Yu, H.; Li, L.; et al. 2026. Gui-genesis: Automated synthesis of efficient environments with verifiable rewards for gui agent post-training. arXiv preprint arXiv:2602.14093. Cheng, Z.; Wang, H.; Liu, Z.; Wang, X.; Zhu, X.; Guo, Y.; Lin, W.; Pan, J. Z.; and Wang, Y. 2026. TerminalWorld: Scaling Terminal-Agent Environments via Agent Skills. arXiv preprint arXiv:2605.20876. Chu, Z.; Hu, J.; Jiang, X.; Zou, P.; Li, H.; Peng, C.; O’Hearn, P.; Barr, E. T.; Harman, M.; Sarro, F.; et al. 2026. TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks. arXiv preprint arXiv:2605.22535. Deng, X.; Gu, Y.; Zheng, B.; Chen, S.; Stevens, S.; Wang, B.; Sun, H.; and Su, Y. 2023. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36: 28091–28114. Fu, Y.; Fu, B.; Wu, Z.; Cheng, S.; Sun, X.; Yang, B.; Li, Z.; Zhao, Y.; Ding, Z.; Liu, Z.; et al. 2026. MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop. arXiv preprint arXiv:2606.22557. Gulcehre, C.; Paine, T. L.; Srinivasan, S.; Konyushkova, K.; Weerts, L.; Sharma, A.; Siddhant, A.; Ahern, A.; Wang, M.; Gu, C.; et al. 2023. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998. Hua, Z.; Yao, Y.; Xie, W.; Zhao, Y.; Liu, M.; Qiu, R.; Huang, Z.; Wang, Z.; Ji, Y.; Ye, Y.; et al. 2026. CLIUniverse: Towards Verifiable Task Synthesis Engine for Terminal Agents. arXiv preprint arXiv:2606.22883. Ivison, H.; Yin, J. O.; Shao, R.; Xiao, T.; Lambert, N.; and Hajishirzi, H. 2026. Tmax: A simple recipe for terminal agents. arXiv preprint arXiv:2606.23321. Jia, H.; Liao, J.; Zhang, X.; Xu, H.; Xie, T.; Jiang, C.; Yan, M.; Liu, S.; Ye, W.; and Huang, F. 2025. Osworld-mcp: Benchmarking mcp tool invocation in computer-use agents. arXiv preprint arXiv:2510.24563. Li, W.; Zhou, B.; Yu, Y.; Xu, Z.; Yang, Y.; Li, D.; and Shan, C. 2026. WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces. arXiv preprint arXiv:2606.09426. Merrill, M. A.; Shaw, A. G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J. Y.; Walshe, T.; Buchanan, E. K.; et al. 2026. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868. OpenAI. 2026. Introducing GPT-5.4. https://openai.com/ index/introducing-gpt-5-4/. Accessed: 2026-07-22. Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents.
Rawles, C.; Clinckemaillie, S.; Chang, Y.; Waltz, J.; Lau, G.; Fair, M.; Li, A.; Bishop, W.; Li, W.; Campbell-Ajala, F.; et al. 2025. Androidworld: A dynamic benchmarking environment for autonomous agents. In International Conference on Learning Representations, volume 2025, 406–441. Song, L.; Dai, Y.; Prabhu, V.; Zhang, J.; Shi, T.; Li, L.; Li, J.; Savarese, S.; Chen, Z.; Zhao, J.; et al. 2025. Coact1: Computer-using agents with coding as actions. arXiv preprint arXiv:2508.03923. Team, K.; Bai, T.; Bai, Y.; Bao, Y.; Cai, S.; Cao, Y.; Charles, Y.; Che, H.; Chen, C.; Chen, G.; et al. 2026. Kimi K2. 5: Visual Agentic Intelligence. arXiv preprint arXiv:2602.02276. Wang, B.; Lu, D.; Wang, J.; Bai, T.; Liu, S.; Zhang, Z.; Wang, H.; Hu, H.; Xie, T.; Bai, S.; et al. 2026a. Cuagym: Scaling verifiable training environments and tasks for computer-use agents. arXiv preprint arXiv:2605.25624. Wang, W.; Shi, H.; Yuan, M.; Lin, Y.; Tong, P.; Zhou, H.; Liu, G.; Zhao, P.; Wang, Y.; and Chen, S. 2026b. FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems. In Findings of the Association for Computational Linguistics: ACL 2026, 28747–28767. Wang, W.; Yu, Z.; Ye, R.; Zhang, J.; Liu, G.; Liu, L.; Chen, S.; and Wang, Y. 2025. FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Data. In Christodoulopoulos, C.; Chakraborty, T.; Rose, C.; and Peng, V., eds., Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 26387–26408. Suzhou, China: Association for Computational Linguistics. ISBN 979-8-89176-332-6. Wu, Y.; Peng, Y.; Chen, Y.; Ruan, J.; Zhuang, Z.; Yang, C.; Zhang, J.; Chen, M.; Tseng, Y.; Yu, Z.; et al. 2026. Autowebworld: Synthesizing infinite verifiable web environments via finite state machines. arXiv preprint arXiv:2602.14296. Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; Hua, T. J.; Cheng, Z.; Shin, D.; Lei, F.; et al. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37: 52040–52094. Yang, Y.; Fan, T.; and Huang, C. 2026. Cli-anything: Towards agent-native computer use. arXiv preprint arXiv:2606.03854. Yang, Y.; Yang, Z.; Dou, Z.-Y.; Nguyen, A.; You, K.; Attia, O.; Szot, A.; Feng, M.; Ramrakhya, R.; Toshev, A.; et al. 2025. Ultracua: A foundation model for computer use agents with hybrid action. arXiv preprint arXiv:2510.17790. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629. Yuan, M.; Zhou, Z.; Xiong, X.; Wu, W.; Sun, J.; Song, J.; Cui, K.; Wang, B.; Wu, H.; Li, Y.; et al. 2026. OSWorld2. 0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. arXiv preprint arXiv:2606.29537. Zelikman, E.; Wu, Y.; Mu, J.; and Goodman, N. 2022. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35: 15476–15488.
Zhang, Z.; Wang, Z.; Zhang, X.; Guo, Z.; Li, J.; Li, B.; and Lu, Y. 2026. InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training. arXiv preprint arXiv:2601.04126. Zhao, Y.; Huang, J.; Hu, J.; Wang, X.; Mao, Y.; Zhang, D.; Jiang, Z.; Wu, Z.; Ai, B.; Wang, A.; Zhou, W.; and Chen, Y. 2024. SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning. arXiv:2408.05517. Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; et al. 2024. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, volume 2024, 15585–15606. Zhou, X.; Zhang, S.; Zhao, Y.; Wei, J.; Song, T.; Cohan, A.; and Zhao, C. 2026. GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents. arXiv preprint arXiv:2606.24551.
A A.1
Additional Experiments
Action-Modality Behavior on CUA-Verse
Setting. Beyond overall performance, we examine how different models use the two modalities on CUA-Verse. CUA-Verse is explicitly constructed from hybrid tasks whose reference solutions span both GUI and CLI operations over a shared application state, including visually grounded interactions that cannot be reduced to command execution alone. From the stored trajectories, we classify each nonterminal action as CLI or GUI and report CLI %, the fraction of executed GUI+CLI actions that use the CLI. We report it alongside Steps and the CUA-Verse Score to characterize the relationship among modality choice, efficiency, and performance. Model
Steps
CLI %
Score
Kimi K2.5 Seed2.1 Pro GPT-5.5
21.1 40.1 23.1
30.3 49.2 100.0
0.522 0.599 0.768
Qwen3.5-9B EvoCUA-8B Ours
56.2 41.7 35.2
0.0 3.0 25.3
0.189 0.330 0.582
Table 4: Action-modality behavior on CUA-Verse, averaged over the eight applications. Steps is the mean executed action count, CLI % is the CLI share of executed GUI+CLI actions, and Score is the average VLM-judge score. Greater CLI use often accompanies higher efficiency and performance, but CLI usage alone is insufficient: even GPT-5.5, which executes entirely through CLI in these trajectories, does not fully solve the benchmark, reflecting the visually grounded demands of CUA-Verse.
Results. Table 4 shows that effective CLI use is important, but the key capability is cross-modality orchestration rather than maximizing CLI usage. (1) The untuned base remains GUI-bound. Qwen3.5-9B issues no CLI actions, uses nearly the full 60-step budget, and obtains the lowest Score (0.189), reflecting inefficient GUI-only execution. (2) Hybrid training changes interaction behavior. With the same backbone, our model raises CLI usage to 25.3%, reduces mean steps by 37%, and increases Score from 0.189 to 0.582, showing that CUA-Universe teaches the model to exploit CLI operations while retaining GUI interaction when needed. (3) More CLI is not itself sufficient. GPT5.5 uses CLI for all recorded actions and achieves the highest Score (0.768), yet still falls well short of perfect performance. Conversely, EvoCUA-8B remains almost entirely GUI-based despite having access to the CLI and obtains only 0.330. Together, these results suggest that CUA-Verse rewards the ability to select and coordinate modalities according to the task, rather than simply favoring either GUI or CLI in isolation.
App
Base
8-app LoRA (OOD-only)
16-app LoRA (+in-domain)
Chrome GIMP Calc Impress Writer Thunderbird VLC VS Code
23.9 26.9 8.5 27.7 34.8 33.3 29.4 26.1
26.1 30.8 10.6 25.8 39.1 46.7 31.4 34.8
37.0 42.3 18.8 31.8 52.2 73.3 37.4 43.5
All apps
24.2
27.4 (+3.2)
40.2 (+16.0)
Table 5: Out-of-domain training transfer on OSWorld under the controlled 244-task GUI+CLI scope. Values are success rates in percent. The 8-app OOD-only LoRA is trained on cua-universe extensions applications disjoint from the OSWorld applications, whereas the full 16-app LoRA is trained on both. Parenthesized values in the final row denote absolute percentage-point gains over the untuned Qwen3.5-9B.
A.2
Out-of-Domain Training Transfer
Setting. To test whether the LoRA gains reflect a transferable CLI-usage capability rather than memorization of the evaluation applications, we train an 8-app OOD-only LoRA exclusively on applications that are disjoint from the eight OSWorld evaluation applications and evaluate it zeroshot under the same controlled 244-task GUI+CLI scope. We compare it against the untuned Qwen3.5-9B and our full 16-app LoRA, whose training mixture additionally includes the eight OSWorld application domains. Results. Three observations stand out (Table 5). (1) Hybrid CLI-usage skills transfer across applications: despite being trained exclusively on applications disjoint from the OSWorld evaluation applications, the 8-app OOD-only LoRA improves over the untuned base on seven of the eight applications and raises the overall success rate from 24.2% to 27.4% (+3.2 points). This zero-shot improvement indicates that the LoRA learns a partially application-agnostic capability for selecting and coordinating GUI and CLI actions rather than only memorizing application-specific commands. (2) The transfer is heterogeneous across applications: the largest out-of-domain gains occur on Thunderbird (+13.4 points) and VS Code (+8.7), followed by Writer (+4.3) and GIMP (+3.9); Impress is the only application on which performance decreases, by 1.9 points. (3) In-domain coverage substantially compounds the transferable gain: the full 16-app LoRA improves over the base on all eight OSWorld applications and raises the overall success rate to 40.2% (+13.2 points), showing that transferable hybridinteraction skills and application-specific training data provide complementary benefits.
B
Adapted Applications
CUA-Universe currently provides adapters for the 16 desktop applications listed in Table 6, covering diverse software
domains. Eight applications are inherited from OSWorld, and eight additional applications are introduced by CUAUniverse. Each application implements the uniform adapter interface described in §3.2, exposing graphical, commandline, and programmatic interfaces that operate on the same underlying state.
OSWorld applications GIMP LibreOffice Writer Chrome LibreOffice Impress VLC VS Code LibreOffice Calc Thunderbird
46 30
15 14 13 13 10 10
CUA-Universe extensions
Application
Domain
Seed
Chrome
Web browsing
Profile
GIMP
Image editing
.xcf
LibreOffice Calc
Spreadsheets
.ods
LibreOffice Impress
Presentations
.odp
LibreOffice Writer
Word processing
.odt
Thunderbird
Profile
VLC
Media playback
Media file
VS Code
Code editing
Workspace
Audacity
Audio editing
.aup3
Blender
3D modelling
.blend
Draw.io
Diagramming
.drawio
Godot
Game development
.tscn
Kdenlive
Video editing
.kdenlive
OBS Studio
Screen recording
.json
QGIS
Geospatial analysis
.qgz
Zotero
Reference management
Library
OSWorld applications
CUA-Universe extensions
Table 6: Desktop applications supported by CUAUniverse. Blue rows indicate applications inherited from OSWorld, whereas orange rows indicate applications added by CUA-Universe.
C
C.1
GIMP — 46 commands
The GIMP uses a live backend that drives the actual GUI application, while exposes the command interface and return
38 38 36 29 26 25 16 0
10
20
30
40
50
Agent-visible commands OSWorld applications
CUA-Universe extensions
Figure 5: Agent-visible command count per application, grouped by source: the eight applications reused from OSWorld and the eight CUA-Universe extensions (∼404 commands in total).
results in JSON format. PROJECT
project new project open project save project info project json project profiles
CANVAS
canvas resize canvas scale canvas crop canvas mode canvas dpi canvas info
LAYER
layer new layer add from file layer list layer remove layer duplicate layer move layer set layer flatten layer merge down
FILTER
filter add filter list filter list available filter info filter set filter remove
DRAW
draw text draw rect
MEDIA
media probe media histogram media list media check
EXPORT
export render export presets export preset info
LIVE
live status live canvas info live layer list live draw text live draw rect live export render
SESSION
session undo session redo session history session status
Tool Construction
The tool layer converts each application’s native automation surface—a scripting API, a project-file serializer, or an installed command-line binary—into a small, bounded vocabulary of subcommands that return JSON. A registry stores only metadata (tool ID, description, and the complete command template); the rollout runner executes the named command inside the benchmark VM and returns its exit code, stdout, and stderr to the next decision step. Where an application already ships a usable command-line interface, the registry wraps it directly; where it does not, a bounded harness operates on the native project artifact—or on the live application—so that every effect stays auditable in the same file the GUI opens. The deployed snapshot spans 16 applications and ∼404 agent-visible commands: 151 commands across the eight applications reused from OSWorld and 253 across the eight CUA-Universe extensions (Figure 5). Command counts are not padded to a target range—they track each application’s real capability surface. For brevity, the per-application sections below list tool IDs grouped by command family.
45
Blender QGIS Zotero OBS Studio Kdenlive Godot draw.io Audacity
C.2
Blender — 45 commands
A bounded bpy program runs in a single headless invocation, during which it has full read/write access to the real .blend artifact, allowing it to both inspect and alter the scene contents permanently. SCENE
scene new scene open scene save scene info scene profiles scene json
OBJECT
object add object remove object duplicate object transform object set object list object get
MATERIAL
material create material assign material set material list material get
MODIFIER
modifier list available modifier info modifier add modifier remove modifier set modifier list
item export item citation item bibliography item context item analyze item add to collection item move to collection note get note add
NOTE
camera add camera set camera set active camera list
SEARCH
search list search get search items
TAG / STYLE
tag list tag items style list
LIGHT
light add light set light list
IMPORT / SESSION
ANIMATION
animation keyframe animation remove keyframe animation frame range animation fps animation list keyframes
CAMERA
RENDER
render settings render info render presets render execute render script
SESSION
session status session undo session redo session history
C.3
OBS — 36 commands
A typed scene, source, filter, transition, and output model serialized to the native OBS scene collection, allowing the agent to compose and control live production setups. PROJECT
project info project json project save
SCENE
scene add scene remove scene duplicate scene set active scene list
SOURCE
source add source remove source duplicate source set source transform source list
FILTER
filter add filter remove filter set filter list filter available
AUDIO
audio add audio remove audio volume audio mute audio unmute audio monitor audio list
TRANSITION
transition add transition remove transition set active transition duration transition list
OUTPUT
output streaming output recording output settings output info output presets
QGIS — 38 commands
Bounded PyQGIS operations against a native .qgs project, with optional GUI sync, giving the agent full read and write access to layers, features, layouts, and processing pipelines. PROJECT
project new project open project info project set crs project save layer create vector layer list layer info layer remove layer style simple layer style graduated layer style categorized layer label simple layer label buffer layer category visibility layer filter attribute
LAYER
FEATURE
feature add feature add point feature list
LAYOUT
layout create layout list layout info layout remove layout add map layout add legend layout legend remove layer layout add label layout sync extent
PROCESS
process list process help process run
EXPORT
export presets export pdf export image
SESSION
session status session history
APP / REPL
app open repl
C.4
C.5
Zotero — 38 commands
import file import json session
C.6
LibreOffice Writer — 30 commands
ODF-backed Writer operations over the live UNO bridge, implementing the Writer slice of the shared cli-anythinglibreoffice harness for creating and editing rich text documents. DOCUMENT
document new document open document save document info document profiles document json
WRITER
writer add paragraph writer add heading writer add list writer add table writer table list writer table insert row writer table set cell writer table set row background writer add page break writer remove writer list writer set text
The seeded Zotero profile is accessed via its local database, browser connector, and Local API, enabling full management of collections, items, notes, and saved searches. APP
app status app version app launch app enable local api app ping
COLLECTION
collection list collection find collection tree collection get collection items collection use selected collection create collection export
STYLE
style create style modify style list style apply style remove
EXPORT
export presets export preset info export render
item list item find item get item children item notes item attachments item file
SESSION
session status session undo session redo session history
ITEM
C.7
Kdenlive — 29 commands
A project-file editor for Kdenlive’s MLT XML format, using melt for rendering and writing a native .kdenlive project with timeline clips, filters, transitions, and guides.
CONNECT
connect add connect remove connect label connect style connect list
EXPORT
export render
SESSION
session undo session redo session status
C.10
Audacity — 16 commands
PROJECT
project info project json project save project profiles
BIN
bin import bin list bin get
TIMELINE
timeline add track timeline add clip timeline remove clip timeline trim timeline split timeline move timeline list
IMPORT / TRACK
filter add filter set filter list filter available
SELECT
select all select time select tracks
EFFECTS
TRANSITION
transition add transition set transition list
amplify normalize fade in fade out reverse change pitch change speed
LABEL / EXPORT / RAW
GUIDE
guide add guide list
EXPORT
export xml export presets
SESSION
session status session undo session redo session history
FILTER
C.8
C.11
import track new get info
label add export run
Chrome — 15 commands
A restricted Chrome DOM surface (DOMShell), exposed through page, accessibility-tree, and action commands for navigation, file-system access, and element interaction.
Godot — 26 commands
Godot project assets, scenes, scripts, and export configuration, accessed via the engine’s conventions and CLI runtime for project creation, scene editing, and build export. ENGINE
engine version engine status
EDITOR
editor open
PROJECT
project create project info project scenes project scripts project resources project reimport project set setting project add input action project add input key scene create scene read scene add node scene set property scene set control layout
SCENE
PAGE
page open page reload page back page forward page info
FS
fs ls fs cd fs cat fs grep fs pwd
ACT
act click act type
SESSION
session status session daemon start session daemon stop
C.12
LibreOffice Impress — 14 commands
Native .odp slide, content, element, and export commands over ODF/UNO, implementing the Impress slice of the shared cli-anything-libreoffice harness. DOCUMENT
document save document info
IMPRESS
impress add slide impress remove slide impress set content impress list slides impress add element impress remove element impress move slide impress duplicate slide impress get slide
EXPORT
export presets export preset info export render
script run script inline script validate script read script write script append
SCRIPT
EXPORT
export presets export build
SESSION
session
C.9
A live client for Audacity’s Mod-Script-Pipe, treating the running GUI project as the source of truth for importing audio, making selections, and applying effects.
draw.io — 25 commands
The native .drawio XML/mxGraph model, manipulated through a constrained diagram API that exposes pages, shapes, connectors, and export operations. PROJECT
project new project open project save project info project xml
PAGE
page add page remove page rename page list
SHAPE
shape add shape remove shape list shape label shape move shape resize shape style
C.13
VS Code — 13 commands
The installed code binary, invoked to open paths, diff and merge files, manage extensions, and access settings and keybindings. OPEN
open open new window open user settings json open keybindings json
NAVIGATE
goto diff merge
WORKSPACE
add folder wait file closed
EXTENSIONS
install extension uninstall extension list extensions
STATUS
status
C.14
VLC — 13 commands
Direct vlc and cvlc invocations to launch media, open files at specific times or segments, capture snapshots, transcode audio and video, and manage configuration. APP
app open app help
OPEN
open media file open media at time open network stream open preferences file
PLAY
play fullscreen play paused segment
CAPTURE
snapshot frame
CONVERT
convert audio mp3 convert audio wav convert video mp4
CONFIG
reset user config
C.15
LibreOffice Calc — 10 commands
Bounded sheet and cell operations over ODF/UNO, implementing the Calc slice of the shared cli-anything-libreoffice harness for spreadsheet creation and manipulation. DOCUMENT
document save document info
SHEET
calc list sheets calc add sheet calc rename sheet
CELL
calc get cell calc set cell
EXPORT
export presets export preset info export render
C.16
Thunderbird — 10 commands
The installed executable’s profile-aware launch and compose capabilities, with every command naming the benchmark profile to open mail, address book, calendar, and compose emails. APP
app open app help
OPEN
open mail open addressbook open calendar
COMPOSE
compose email compose cc bcc compose attachment mailto email
DESKTOP HANDLER
xdg email
D D.1
Task-Weave
Seed Examples
A seed is a concrete, pre-loaded application state—an actual project, document, or media file opened in its native GUI app—that anchors task synthesis in real, verifiable context rather than a generic natural-language prompt. Each seed carries genuine content and metadata (e.g., a populated spreadsheet, a raster design with stable regions, or a
saved .drawio graph), so synthesized tasks reference elements that truly exist and produce app state that can be checked deterministically. Grounding generation in seeds reduces hallucinated targets, yields more executable and diverse instructions, and makes success criteria objective. Figure 6 shows six representative seeds spanning 3D modeling, spreadsheets, media playback, image editing, presentations, and diagramming; together they illustrate the breadth of GUI applications and file types our pipeline builds on. Seeds are not hand-authored. A coding agent (Codex) searches for and downloads real, diverse source files from public repositories and asset libraries, then uses formatconversion tools to normalize each one into an agentreadable, synthesis-friendly project state—so the seed pool scales automatically while staying grounded in authentic, heterogeneous content rather than templated fixtures.
E E.1
Steer-Path Rollouts
Rollout Agent Prompt
Figure 10 shows the message format of a data-generation rollout (the Kimi K2.5 backbone on a draw.io task). Two properties are worth noting. First, the rollout interface exposes GUI and CLI jointly: the system prompt lists the application’s CLI tool registry alongside pyautogui, wait, and terminate, and the baseline modality guidance is the generic instruction to “use CLI for precise operations and GUI for visual tasks.” Second, this is where Path-Steer enters: the task guidance field carries the efficiency-aware prior for this attempt (an ordered plan such as “First . . . ”), steering the backbone toward a shorter hybrid path during data generation. The agent then interleaves CLI edits with GUI waits, and each CLI call returns a structured JSON result (rc, output) that becomes the next observation. Path-Steer is used only during these rollouts; at evaluation the guidance field is empty (Appendix I.1). The task text is elided below; the point is the interface and the guided GUI+CLI interleaving, not the specific diagram.
E.2
Steer-Path Rollout Examples
In this VLC desktop task, the w/ Path-Steer trajectory overall outperformed the w/o Path-Steer trajectory. Although the w/ Path-Steer trajectory made several ineffective attempts with VLC CLI parameters in the early stage, it was able to switch to a feasible alternative solution in time, namely using ffmpeg to complete the video rotation, and eventually succeeded in both exporting the corrected video and opening the Audio Effects panel. In contrast, while the w/o Path-Steer trajectory followed a more intuitive GUI-based workflow, it became stuck in repeated searching and ineffective interactions during the filter configuration stage, and ultimately failed to complete the task. This case suggests that Path-Steer may not necessarily reduce local trial-and-error, but it can significantly improve task completion, error recovery, and goal convergence in complex desktop environments.
F
VLM Judge
We score task success with a VLM-as-judge, following the now-standard practice of using strong (vision-)language
Blender. This Objaverse product seed opens a concrete webcam-cover 3D asset, grounding modeling tasks in real mesh and material structure.
LibreOffice Calc. This budget workbook seed provides real sheets, headers, rows, and editable cells for grounded spreadsheet tasks.
VLC. This media seed opens a concrete video file, grounding playback and media-control tasks in a real duration and visual stream.
GIMP. This poster seed exposes a real raster design with stable regions for annotation, cropping, banner, and export tasks.
LibreOffice Impress. This photo-heavy deck seed grounds presentation tasks in existing slides, layouts, images, and editable text frames.
Draw.io. This flowchart seed grounds diagram-editing tasks in existing nodes, connectors, labels, and saved .drawio XML.
Figure 6: Six opened seed examples used to ground task synthesis across diverse GUI applications. Each seed is a concrete project or file with metadata and verifiable app state, rather than a generic natural-language prompt. models as automatic evaluators of agent trajectories; in the GUI and computer-use setting in particular, VLM judges are widely used both to score task completion from screenshots and to curate training trajectories for post-training. Independent of the task synthesizer and of any rollout reward, the judge takes the task instruction, a compact rendering of the trajectory, and a chronological sample of screenshots, and returns a score in [0, 1] with a short justification. Its rubric treats CLI output and exported-artifact evidence as authoritative for file-producing tasks while using the final screenshot as primary evidence otherwise. We run the judge (GPT5.4 (OpenAI 2026)) at temperature 0.1 with up to 15 screenshots; the full prompt is shown in Figure 11.
F.1
Agreement with Human Labels
Because Score is our headline CUA-Verse metric and the same judge filters training data, we validate it against human judgment. Three annotators independently re-scored a stratified sample of trajectories for all 16 applications, blind to the judge’s score. For each application we drew 30 trajectories the judge accepted (score ≥ 0.75) and 30 it rejected (score < 0.75), and labeled each as fully success, partial success, or failure. We take the binary decision positive = human “fully success” and negative = otherwise, matching the semantics of the 0.75 acceptance threshold used to filter training data. Because the sample is stratified rather than proportional, we reweight each application’s confusion cells
by its true acceptance rate p (Table 7, second column) before computing the aggregate agreement and κ. Two findings anchor the judge’s reliability. On the acceptance set (480 trajectories) the judge attains 99.0% precision (475/480), and crucially not one accepted trajectory was a human “failure”—the five imperfections are all partial successes—so the label-noise upper bound on training data is ≤ 1.0%. On the rejection set (480 trajectories) the falsenegative rate is 5.2% (25 fully-successful trajectories scored low); these cost data yield but do not inflate reported Score. Reweighted to the true distribution, judge–human agreement is 97.0% and Cohen’s κ = 0.94 (Landis–Koch “almost perfect”). Under a lenient criterion that counts partial successes as positive, the picture is unchanged where it matters: the judge accepts zero true failures (precision 100%; all 480 accepted trajectories are at least partial successes). Its rejection set is dominated by partial successes (406/480) with only 49 true failures, confirming that the 0.75 threshold deliberately screens out partially-completed work—the intended behavior for a training-data filter—rather than confusing success with failure.
G
Rollout Harness and Budget
Rollout Harness and hardware. Each task runs in an isolated OSWorld virtual machine provisioned with 4 vCPU and 4 GB of guest RAM on a copy-on-write overlay disk,
# Task Instruction:
The video I'm playing in VLC is upside-down. Please rotate it 180 degrees and save the corrected version as a new MP4 file named 'rotated_fixed.mp4' in my Documents folder. Then use VLC's Tools menu to open the Audio Effects panel so I can adjust the equalizer settings for playback.
# SteerPath Guidance:
First, use the CLI command `vlc_convert_video_mp4` with the transpose video filter to rotate 180 degrees: `cvlc -I dummy \"/home/user/Documents/rotated_orientation_sample.avi\" --video-filter=transpose --transpose-angle=180 mux=mp4,dst='/home/user/Documents/rotated_fixed.mp4'}\" vlc://quit`. After conversion completes, open the converted file in VLC with `vlc_open_media_file` to verify the rotation, then use ...
"__CLI__:cvlc -I dummy \"/home/user/Do cuments/rotated_orientation_sample.avi\" --video-filter=transpose ... "
pyautogui.click(714, 394)
"__CLI__:VLC_VERBOSE=-1 vlc --no-video-title-show \ "/home/user /Documents/rotated_fixed.mp4\""
pyautogui.click(1008, 394)
pyautogui.click(1044, 753)
pyautogui.scroll(-3, 936, 570)
FAIL
# Without SteerPath:
pyautogui.click(1131, 481)
Figure 7: VLC task trajectories w/ Path-Steer and w/o Path-Steer. The Path-Steer trajectory shows optimal solution using CLI tools and GUI actions, whereas the trajectory without sterr-path guidance becomes stuck during filter configuration and fails the task. with a 3 s post-action settle before each observation. All agents share a 60-step budget; observation history follows each baseline’s default, while our Qwen-family models use a 3-frame history with at most 4 images per step.
Rollout budget. Trajectory collection dominates the wallclock cost of building the fine-tuning dateset, because each rollout drives a full GUI environment rather than a single forward pass. A single rollout takes on average t̄ ≈ 5 minutes end-to-end (environment reset, per-step model calls, action execution, and the post-action settle), and a training run consumes on the order of N ≈ 104 rollouts. Because each environment needs only ∼4 vCPU and no GPU, this collection is CPU-bound and embarrassingly parallel. Concretely, a single commodity 128-core CPU server hosts P = ⌊128/4⌋ = 32 environments concurrently, so the full N t̄ ≈ 50,000 VM-minutes reduce to N t̄/P ≈ 1,560 minutes—just over one day of wall-clock time on that one machine. This is the key practicality of our recipe: the entire data-generation pipeline that produces our 9B model fits on a single ordinary CPU box, with no GPU cluster and no multimachine orchestration. Throughput in practice is slightly below this ceiling because of environment-reset overhead and occasional VM stalls (we observe ∼1.3–1.5 calendar days on one 128-core host), and the job is trivially shardable across additional machines when faster turnaround is needed.
H H.1
Training Details
Data
All records satisfy VLM score ≥ 0.75 and are step-level examples in Qwen XML format with a three-step history context. The CUA-Verse pool is balanced at 17,511 records per app across Audacity, Blender, draw.io, Godot, Kdenlive, OBS, QGIS, and Zotero. The OSWorld pool is balanced at 11,915 records per app across LibreOffice Writer, LibreOffice Calc, LibreOffice Impress, VS Code, VLC, Thunderbird, GIMP, and Chrome. The OSWorld build enforces exactly four images per record; the CUA-Verse build is VLMfiltered but not strictly four-image (22,590 records have an image count other than four). Figure 2 summarizes the overall split and the per-application episode counts for both pools, and Table 8 reports the final split sizes.
H.2
Base Model and Fine-tuning
We use Qwen3.5-9B as the base model in bfloat16. Finetuning is performed with LoRA while the base weights, vision tower, and multimodal aligner remain frozen; only the language-model linear modules are trainable. Table 10 lists the optimization and sequence settings used for fine-tuning.
I I.1
Experiments
Evaluation Agent Prompt
Each evaluated backbone uses its own actor prompt, and each is run in two modes across OSWorld and CUA-Verse:
# Task Instruction:
Please help me modify the setting of VS Code to keep my cursor focused on the debug console when debugging in VS Code, instead of automatically focusing back on the Editor.
GUI + CLI Mode: 5 CLI calls + 12 GUI actions -> reward = 1
pyautogui.click(492, 321)
pyautogui.typewrite ("debug focus console")
`CLI: cat > ~/.config/Code/User/settings.json << 'EOF' {"security.workspace.trust.startupPrompt": "never","editor.wordWrap": "on", "files.autoCreate" : true, "debug.focusEditorOnBreak":false} EOF
Done
GUI Mode: 60 GUI actions -> reward = 0
pyautogui.click(490, 321)
pyautogui.typewrite ("debug focus")
pyautogui.click(845, 388)
FAIL
Figure 8: An OSWorld VS Code task under GUI+CLI vs. GUI-only interfaces. GUI+CLI mode combines GUI navigation with direct settings-file editing: it locates the target setting, writes debug.focusEditorOnBreak: false to settings.json, and verifies persistence, completing the task in 5 CLI calls and 12 GUI actions (reward 1). GUI-only mode repeatedly searches and clicks through the settings UI but fails to commit the correct configuration, exhausting the 60-action budget (reward 0). competence on the software the pipeline covers. CLI abstract tools
GUI abstract tools
8.5
8.5
8
Abstract tools / task
a CLI mode that additionally exposes the execute cli function, and a GUI mode that follows the standard OSWorld screenshot-and-pyautogui protocol. Figure 12 shows the Qwen CLI-mode prompt as a representative example. Regardless of backbone or mode, the path-hint field defaults to none at evaluation, so no task hint or solution guidance is appended and the evaluated policy is prior-free—the PathSteer priors of Appendix E.1 act only during data-generation rollouts.
6.6
6.2
5.7
6
5.8 4.9
4.7 4
2
I.2
CUA-Verse Benchmark
CUA-Verse comprises 160 hybrid GUI+CLI tasks (eight applications × 20 tasks) synthesized by the same pipeline as the training data. To characterize what the tasks demand, we abstract each task’s reference solution into a deduplicated set of high-level tools and classify each as GUI (normalized keyboard/click/move/drag/scroll/vision) or CLI (application command groups such as godot scene * or qgis layer *). Figure 9 reports the resulting per-application composition. Tasks require 5.7 abstract tools on average, of which 59% are CLI, and every application mixes both modalities—confirming that CUA-Verse is genuinely hybrid rather than solvable by either modality alone. The two extremes are illustrative: Blender is GUI-dominant (spatial 3D manipulation), whereas Zotero and Kdenlive are CLI-heavy (structured library and timeline edits). The eight applications are in-domain by design, since we want to measure hybrid
0
y acit
Aud
der
w.io
Blen
Dra
ot
God
e io nliv Stud Kde OBS
QGIS
ro
Zote
Figure 9: CUA-Verse task composition. Mean number of abstract tools per task for each application, split into CLI (orange) and GUI (blue) operations. Every application requires both modalities.
J
OSWorld Example
Figure 8 contrasts the GUI+CLI and GUI-only interfaces on a single OSWorld task—modifying a VS Code setting so the cursor stays focused on the debug console during debugging rather than snapping back to the editor. In GUI+CLI mode the agent uses a few GUI actions to locate the relevant setting, then drops to the CLI
App
p
FP
FN
Prec. (%)
FNR (%)
Chrome GIMP LibreOffice Calc LibreOffice Impress LibreOffice Writer Thunderbird VLC VS Code Audacity Blender Draw.io Godot Kdenlive OBS QGIS Zotero
0.53 0.52 0.56 0.51 0.53 0.59 0.57 0.52 0.52 0.49 0.48 0.46 0.59 0.54 0.57 0.48
0 1 0 0 0 0 0 0 0 0 0 2 1 1 0 0
3 3 1 2 0 0 0 2 0 2 3 2 1 3 2 1
100.0 96.7 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 93.3 96.7 96.7 100.0 100.0
10.0 10.0 3.3 6.7 0.0 0.0 0.0 6.7 0.0 6.7 10.0 6.7 3.3 10.0 6.7 3.3
Overall
—
5
25
99.0
5.2
Table 7: VLM judge vs. human labels, all 16 applications. Three annotators, blind to the judge score, labeled 30 judgeaccepted and 30 judge-rejected trajectories per application. p is the true acceptance rate; FP counts accepted trajectories that were not a human “fully success” (out of 30, all of them partial successes—zero hard failures); FN counts rejected trajectories that were in fact fully successful (out of 30). Precision is on the acceptance set and FNR on the rejection set. Reweighting each application by p gives an overall agreement of 97.0% and Cohen’s κ = 0.94. Blue rows are OSWorld applications; orange rows are CUA-Universe extensions.
to write ”debug.focusEditorOnBreak”: false directly into settings.json and verify that it persists, resolving the task in 5 CLI calls and 12 GUI actions (reward 1). In GUI-only mode it repeatedly searches and clicks through the settings UI but never commits the correct configuration, exhausting the full 60-action budget (reward 0). The case makes the orchestration advantage concrete: the GUI locates the setting, while the CLI performs the precise, verifiable write that the GUI-only agent cannot reliably land.
Split
Episodes
Step records
Total (GB)
CUA-Verse OSWorld
2,526 2,397
140,088 95,320
172.97 113.67
Total
4,923
235,408
286.64
Table 8: Final training splits. “Total (GB)” is the on-disk size including all screenshots; step records reference images by path. Parameter Tuner type LoRA rank LoRA α LoRA dropout LoRA bias Target modules Vision tower Multimodal aligner
Limitations
Our study leaves several directions open. First, CUAUniverse synthesizes single-application tasks: each task is grounded in one application’s shared GUI+CLI state, which is what lets us build verifiable environments and steer efficient hybrid paths at scale. Cross-application workflows, where state is carried across several applications, are a natural extension of the same pipeline rather than a different design, and we leave them to future work. Second, we use the harvested trajectories for supervised fine-tuning only, so the learned policy is bounded by its data-generation backbone; because the pipeline already produces per-task verifiers, using them as rewards for reinforcement learning is a direct next step toward surpassing the teacher. Third, task success
LoRA 8 32 0.05 none all-linear frozen frozen
Table 9: LoRA configuration. is scored by a VLM judge rather than per-task programmatic checks, which may introduce label noise; the judge– human validation in Appendix F bounds this (99.0% acceptance precision, Cohen’s κ = 0.94), and we further mitigate it with a conservative acceptance threshold and by grounding the judge in CLI and exported-artifact evidence, but we do not eliminate it. Finally, application adaptation assumes software that is open-source or scriptable enough to expose a command-line surface, and our environments target desktop Linux; the coding-agent-driven construction is not tied to these choices in principle, but broadening to closed-source or non-desktop platforms remains open.
Parameter
K
Value
Optimizer Learning rate Warmup ratio Per-device train batch size Gradient accumulation Effective batch size (8 GPUs) Epoch argument Max length Max pixels Image max token budget Attention DeepSpeed
Value ms-swift default 1 × 10−4 0.05 2 1 16 3 18,000 602,112 1,024 Flash Attention 2 ZeRO-2
Table 10: Optimization and sequence settings.
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41
[SYSTEM] You are a computer use agent. You can operate the computer through both GUI (mouse/keyboard) and CLI (command-line) tools. You are given a task instruction, a screenshot of the current screen, and your previous interactions. Complete the task by issuing one action per step. Choose the most efficient tool for each step - use CLI for precise operations and GUI for visual tasks. For each step, respond in EXACTLY this format: –thought˝ ## Action: –action˝ ## Code: –code˝ In the code section, use ONE of: - A python block with pyautogui code (GUI action) - A cli block with a shell command (CLI action, only when CLI tools are listed) - A special function in a code block: - –”name”: ”computer.wait”, ”parameters”: –”time”: 3˝˝ - –”name”: ”computer.terminate”, ”parameters”: –”status”: ”success”˝˝ - –”name”: ”computer.terminate”, ”parameters”: –”status”: ”failure”˝˝ CLI Tool: Execute a CLI tool using a cli block. - CLI modifies the file on disk. The GUI may not auto-update. - Always check CLI output (rc and stdout). If rc != 0, diagnose and retry. - You can chain multiple CLI steps before switching to GUI. ## Available CLI Tools (draw.io registry, abridged) drawio˙project˙new — Create a new empty draw.io diagram file drawio˙project˙save — Save the current diagram to disk drawio˙shape˙add — Add a shape at a position (cylinder—rectangle—...) drawio˙shape˙style — Set a style property on a shape (fillColor, ...) drawio˙connect˙add — Add a labeled connector between two shapes drawio˙export˙render — Export the diagram to PNG/PDF/SVG ... (full registry of project/shape/connect/page/export/session commands) [Task guidance for this attempt: First ...] --- Interaction history (elided) --[USER] Instruction: –task˙instruction˝ Previous actions: –action˙history˝ Previous CLI results: –cli˙result˙history˝ –current˙screenshot˝ [ASSISTANT] –thought˝ ## Action: –next˙action˝ ## Code: –python˙or˙cli˙block˝ ... (this GUI/CLI interleaved turn repeats until terminate)
Figure 10: Message format of a data-generation rollout (Kimi K2.5 on a draw.io task). The system prompt exposes GUI and CLI jointly with generic modality guidance (“CLI for precise operations, GUI for visual tasks”); the task guidance field carries the Path-Steer efficiency prior for this attempt (an ordered plan, “First . . . ”), used only during rollouts and empty at evaluation. The interaction history is abstracted with placeholders: each turn supplies the instruction, running action/CLI-result histories, and the current screenshot, and the agent replies with a thought and one GUI or CLI block until it terminates. Task text and coordinates are elided.
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53
[SYSTEM] You are a task completion evaluator for a GUI automation agent operating a desktop application. You will be given: 1. A task instruction describing what the agent should accomplish 2. The execution trajectory showing actions taken and their results 3. Screenshots from the execution in chronological order Your job is to judge whether the task was completed successfully. How to read the screenshots: - The first screenshot is only the earliest available visual baseline. Do not treat it as the result. - Middle screenshots are context for how the state changed over time. - The final screenshot is the primary visual evidence for task completion. - If the trajectory says an edit/export succeeded but the final screenshot or final artifact evidence does not show it, score conservatively. - If the final screenshot is blank, stale, still in a menu/dialog, or appears unchanged from the baseline, do not give high credit for visual tasks. - For file-based editing tasks where the instruction says the exported file/project is the scored artifact, treat successful CLI output and exported artifact evidence as authoritative for file contents. Desktop application screenshots may be stale because many apps do not auto-refresh after external file edits; do not penalize missing GUI refresh when the trajectory provides structured CLI evidence that the exported artifact contains the requested edits. Scoring guidelines: - 1.0: Task fully completed, all requirements met, visual confirmation matches expectations - 0.7-0.9: Task mostly completed, minor issues (e.g. slightly wrong values, visual looks close) - 0.4-0.6: Task partially completed (some steps done, others missing or wrong) - 0.1-0.3: Task barely started or mostly failed - 0.0: Task not completed at all, or agent timed out without meaningful progress When evaluating, consider: - Did CLI commands succeed (rc=0) and produce expected output? - Does the final screenshot show the expected visual result? - Compared with the baseline screenshot, are the requested changes visible in the final screenshot? - Were all sub-tasks in the instruction addressed? - Did the agent reach a terminal state (DONE) or time out? Respond with ONLY a JSON object (no markdown, no extra text): –”score”: ¡float 0.0-1.0¿, ”reason”: ”¡brief explanation of what was/wasn’t completed¿”˝ [USER] # Task Instruction –instruction˝ # Execution Trajectory –trajectory˙evidence˝ # Screenshots The following –n˝ screenshot(s) are sampled from the trajectory and are ordered from earliest to latest. Use the earliest screenshot only as baseline/context. The final 3 screenshots, when present, show the end-state context. Use the FINAL screenshot as the main visual evidence for scoring. –image˙1˝ ... –image˙n˝ # base64 PNG image inputs, each labeled with its chronological role
Figure 11: Full VLM judge prompt. Braces denote per-attempt inputs: {instruction} the task instruction, {trajectory evidence} the compacted action/CLI trace, and {image 1..n} the sampled, role-labeled screenshots supplied as image inputs.
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52
You are a multi-purpose intelligent assistant. Based on my requests, you can use tools to help me complete various tasks. # Tools You have access to the following functions: ¡tools¿ –computer˙use˙function˙schema˝ –execute˙cli˙function˙schema˝ ¡/tools¿ If you choose to call a function ONLY reply in the following format with NO suffix: ¡tool˙call¿ ¡function=example˙function˙name¿ ¡parameter=example˙parameter˙1¿ value˙1 ¡/parameter¿ ¡parameter=example˙parameter˙2¿ This is the value for the second parameter that can span multiple lines ¡/parameter¿ ¡/function¿ ¡/tool˙call¿ ¡IMPORTANT¿ Reminder: - Function calls MUST follow the specified format: an inner ¡function=...¿¡/function¿ block must be nested within ¡tool˙call¿¡/tool˙call¿ XML tags - Required parameters MUST be specified - You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after - If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls - The current date is –runtime˙date˝. - Collapsed screenshots appear as text: This screenshot has been collapsed. ¡/IMPORTANT¿ # Response format Response format for every step: 1) Action: a short imperative describing what to do. 2) One or more ¡tool˙call¿...¡/tool˙call¿ blocks, each containing one tool call. Rules: - Output exactly in the order: Action, then the ¡tool˙call¿ block(s). - If multiple ¡tool˙call¿ blocks are needed, they will be executed sequentially in the order you output them. - Prefer one ¡tool˙call¿ for normal steps; use multiple only for tightly coupled low-level actions such as click-then-type or key sequences. - Use ¡function=execute˙cli¿ for CLI commands. Use ¡function=computer˙use¿ only for GUI actions. - Do not mix execute˙cli with other tool calls in the same step unless the command is immediately required by the same atomic action. - Be brief: one sentence for Action. - Do not output anything else outside those parts. - If finishing, use action=terminate in the tool call. # CLI mode rules - This evaluation is in CLI mode, meaning the separate execute˙cli function is available in addition to normal GUI actions. - For file-based inspection, edits, saves, exports, or verification covered by the listed CLI tools, call the separate ¡function=execute˙cli¿ tool before GUI interaction. - Use GUI actions for visual editing/selection when they are more natural or when no listed CLI tool covers the operation. - Do not use application menus, file pickers, or a different application for an operation that a listed CLI command can perform directly. - Do not run registry tool names as shell commands. Use the concrete command form shown in the command parameter description. - After a successful CLI edit/save/export/verification, use that CLI feedback to decide whether to terminate instead of repeating GUI confirmation loops. - Use wait when the environment needs time after either GUI or CLI actions. - Use terminate only after the requested result is complete. - Never output placeholder coordinates such as [0, 0].
Figure 12: Representative evaluation actor system prompt (Qwen, CLI+GUI); each backbone has its own prompt and a GUImode counterpart following the OSWorld protocol. It is application-agnostic; per-application capability enters only through the task-specific command registry injected into the execute cli schema.