Plover: Steering GUI Agents through Plan-Centric Interaction Madhumitha Venkatesan
Shicheng Wen
Jiajing Guo
University of California, Davis Davis, California, USA [email protected]
University of California, Davis Davis, California, USA [email protected]
Bosch Research North America Sunnyvale, California, USA [email protected]
Jorge Piazentin Ono
Liu Ren
Dongyu Liu
Bosch Research North America Sunnyvale, California, USA [email protected]
Bosch Research North America Sunnyvale, California, USA [email protected]
University of California, Davis Davis, California, USA [email protected]
arXiv:2607.15193v1 [cs.AI] 16 Jul 2026
Abstract Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent. Recent vision-based multimodal agents improve flexibility by operating directly over screenshots and natural language instructions, but planning and adaptation often remain internal, limiting users’ ability to inspect, supervise, or correct system behavior. We present Plover, a plan-centric vision-based GUI automation system that externalizes task plans and replanning as persistent, inspectable, and revisable artifacts. Through a planner–executor architecture, Plover supports explicit supervision of evolving execution, localized correction through editable plans, natural-language guidance, and screenshot-grounded interventions, while preserving prior progress during repair. A formative study with six participants informed the interaction design. We then evaluate Plover through benchmark failure-case repair and scenario-based workflow analyses. Our results show that many autonomous GUI-agent failures are structurally repairable when plans remain visible and interventions are localized, and that explicit replanning helps make GUI automation more transparent, controllable, and adaptable.
CCS Concepts • Human-centered computing → Interactive systems and tools; • Computing methodologies → Intelligent agents.
Keywords GUI agents, Human–AI Interaction, Mixed-initiative Systems, Interactive Task Automation ACM Reference Format: Madhumitha Venkatesan, Shicheng Wen, Jiajing Guo, Jorge Piazentin Ono, Liu Ren, and Dongyu Liu. 2026. Plover: Steering GUI Agents through Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. Conference’17, Washington, DC, USA © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM https://doi.org/XXXXXXX.XXXXXXX
Plan-Centric Interaction. In . ACM, New York, NY, USA, 24 pages. https: //doi.org/XXXXXXX.XXXXXXX
1
INTRODUCTION
Traditional rule-based scripts [20, 25, 52] and commercial Robotic Process Automation (RPA) tools [26, 74] are brittle because they depend on fixed selectors, pixel anchors, or other heuristics tied to interface structure. Even minor UI changes can invalidate these bindings and require repeated manual repair [51]. Vision-based agents instead treat rendered pixels as the primary representation of interface state, acting through screenshots and observable mouse and keyboard actions rather than hidden structural metadata. This makes them particularly useful in heterogeneous or constrained environments, such as legacy tools, proprietary software, non-web applications, and interfaces whose structure is unstable or inaccessible. Recent multimodal large language models (MLLMs) with computer-use capabilities [57, 66, 80, 82, 90] further extend this approach by integrating perception, planning, memory, and execution across diverse applications and workflows. Despite these advances, most systems remain autonomy-first. Once execution begins, planning and replanning typically remain internal. Agents may silently reinterpret intent, regenerate steps, or pursue plausible but incorrect actions without maintaining shared situational awareness with the user [56]. This creates a fundamental design tension: greater autonomy can improve flexibility, but it can also obscure system behavior and reduce user steerability when plans evolve under uncertainty. Accordingly, recent work has begun to explore collaboration and interaction as mechanisms for shared control in agentic systems [11, 19, 32, 64, 79]. Yet important interface and interaction design gaps remain. First, task plans are often presented as one-shot outputs rather than persistent, editable artifacts. Second, current interfaces remain largely text-centric despite the visual and spatial nature of GUI tasks. Third, replanning is typically treated as an internal recovery mechanism rather than an explicit, inspectable interaction. Together, these limitations make moments of instability harder for users to detect, interpret, and repair. As a result, even advanced vision-based agents can still obscure the very breakdowns that call for human guidance. Moreover, high-level natural language correction alone is often insufficient because GUI failures are frequently situated, spatial, and state-dependent. A user may know that the agent clicked the wrong dropdown, selected the wrong cell, or skipped a modal dialog, but describing the exact target verbally can be ambiguous when multiple visually similar elements are present. Similarly, a corrective
Conference’17, July 2017, Washington, DC, USA
prompt often gives the agent new intent without specifying which part of the existing workflow should be preserved or revised. This can lead to broad regeneration, duplicated work, or loss of already completed progress. Plover therefore treats correction not only as a conversational instruction, but as a localized update to an explicit plan grounded in the current screen state. We do not suggest that every GUI task should require continuous user supervision. For short, routine, or low-risk tasks, fully autonomous execution or simple prompt re-issuance may be sufficient. The challenge lies in long-horizon, high-friction, or failure-prone workflows where errors may silently propagate, partial progress is valuable, and users already need some ability to verify outcomes. In such settings, the interaction challenge is not to replace delegation with manual control, but to make the necessary moments of supervision precise, localized, and recoverable. Hence, we argue that robust GUI automation should be treated not only as a modeling problem, but also as an interaction problem. When execution unfolds over multiple steps in dynamic interfaces, users need visibility into how the agent plans, adapts, and recovers over time. We therefore study GUI automation as a supervised, repairable process rather than a fully autonomous one [38]. To address the design gaps above, we introduce Plover, a plan-centric GUI automation system that externalizes task plans as persistent, inspectable, and revisable artifacts. Built on a vision-based GUI agent running in OS-level virtual machines, Plover surfaces plan changes explicitly and supports localized correction through editable plans, natural-language guidance, and screenshot-based interventions. This design provides a mechanism for users to supervise evolving execution, repair pending workflow segments, and preserve prior progress as tasks drift over time, rather than relying on repeated prompt re-issuance or hidden replanning. We evaluate Plover through a formative study that informed the interaction design, followed by three complementary analyses of system behavior. In a benchmark failure-case repair analysis, we first ran 38 challenging tasks under autonomous execution, of which 26 resulted in non-successful outcomes. We then re-ran these 26 tasks in a collaborative setting and found that 23 improved, with 17 becoming complete successes and 6 becoming partial successes, requiring 2.04 interventions per task on average. We further conduct autonomy-only execution through scenario-based stability analysis using visual similarity and plan alignment metrics to examine replay stability and structural plan mismatch. Across both analyses, we identify recurring failure modes including spatial ambiguity in dense interfaces, execution drift in multi-step workflows, and misinterpreted intermediate states. These findings suggest that many GUI-agent failures are not simply terminal errors; they often become recoverable when plans are exposed, and intervention points are made explicit. In summary, we contribute the following: • A plan-centric interaction design for vision-based GUI automation that externalizes task plans as persistent, inspectable, and revisable artifacts, enabling explicit supervision of plan updates, localized repair, and visible replanning during execution. • Plover, a system that instantiates this design through editable plan artifacts, explicit Intelligent Replanning, and multimodal interventions that support targeted correction while preserving execution history and task continuity.
Venkatesan et al.
• An empirical characterization of expert-guided recoverability across benchmark GUI-agent failure cases and scenario-based workflows, showing where localized plan-centric repair can restore progress, where it fails, and which failure modes are most amenable to natural-language guidance, plan editing, multimodal annotation, or system-driven replanning.
2
RELATED WORK
Script-Based and Rule-Based GUI Automation. Traditional GUI automation systems, including commercial Robotic Process Automation (RPA) platforms, Selenium-style scripting frameworks, and record-and-replay tools, rely on low-level structural identifiers or fixed interaction traces to specify actions [15, 48, 61, 74]. Although effective in stable and well-structured environments, these approaches are brittle: interface updates, layout shifts, and rendering changes can easily break scripts and require manual repair [6, 13, 33, 84, 86]. As a result, they offer limited support for adaptation or recovery once execution deviates. Our work instead focuses on vision-based automation that operates over rendered interfaces while exposing task structure for user inspection and repair. Autonomous GUI Agents and Vision–Language Models. Recent advances in large language models (LLMs) and vision–language models (VLMs) have enabled GUI agents that operate directly over screenshots and natural language instructions, reframing automation as a perception–reasoning–action problem [53, 63, 88]. Early systems demonstrated language-driven planning for web and mobile interfaces [60, 72], while later work introduced stronger visual grounding and multimodal action generation that maps rendered pixels to executable actions [18, 24, 55]. More recent foundationlevel agents [37, 53, 77, 91] and Vision-Language-Action architectures [4, 30, 39, 81] have further expanded the scope of GUI automation and improved long-horizon execution across desktop and mobile environments. Many of these systems follow a ReAct-style paradigm [83], and recent work has begun to explore more explicit planning mechanisms for long-horizon GUI tasks [16, 42]. However, these systems remain largely autonomy-first. Planning and replanning are typically maintained as internal model state or exposed only as transient reasoning traces, giving users limited visibility into how actions are selected, how failures are interpreted, or how execution adapts over time [28, 78]. When execution drifts, recovery is usually handled by the agent alone, while user input remains confined to high-level natural language instructions rather than localized correction or collaborative revision of task structure. In contrast, our work externalizes evolving plans as persistent, editable, and inspectable artifacts, giving users direct leverage over how execution is reviewed, revised, and repaired over time. Interfaces for Human–Agent GUI Collaboration. Alongside increasingly capable GUI agents, recent work has begun to explore mixed-initiative and human-in-the-loop automation [2, 68, 92]. Early systems such as SUGILITE [38] showed that users can repair automation through multimodal programming-by-demonstration, and more recent work suggests that users often prefer “do-it-withme” interaction over fully autonomous execution [36]. Systems such as Cocoa [17], CowPilot [31], and Magentic-UI [45] further support shared control through procedural co-authoring, redirected navigation, or editable task plans. Related frameworks also model users
Plover: Steering GUI Agents through Plan-Centric Interaction
as active participants in agent workflows, emphasizing interactiondriven browsing and human-in-the-loop coordination [29, 65, 87]. These systems reintroduce user oversight into agent execution, but many still expose the agent’s reasoning, recovery logic, or task structure only partially, making it harder for users to inspect, revise, and negotiate how execution evolves over time [9, 54]. Recent work such as DoubleAgents [40] argues that plans should be treated as negotiable artifacts rather than fixed outputs. Rather than treating plan externalization as an end in itself, Plover uses the plan as the mechanism for repairing execution as it unfolds. Prior systems have shown that users can co-author, inspect, or edit agent plans, but Plover focuses on the moment when a GUI agent has already begun acting and its behavior starts to diverge from user intent. In this setting, the correction must specify what should change while preserving what has already been completed. Plover addresses this by making repairs operate on the remaining plan, grounded in the current screen state and recorded as part of the evolving workflow. This allows persistent plans to connect execution, correction, replanning, and provenance during GUI-agent repair, rather than serving only as previews, explanations, or editable task descriptions. A related challenge is how users communicate situated corrections during GUI tasks. Because GUI interaction is inherently spatial, natural-language-only feedback is often insufficient in visually dense interfaces. Systems such as Morae [49] and NaviPlus [12] address this challenge by surfacing ambiguity and asking clarifying questions during task execution. Other work explores deictic reference [44, 67] and multimodal grounding, including pointing, highlighting, freehand marks, and visual notes, as mechanisms for expressing intent or steering models [7, 10, 58, 73, 75, 85]. However, in most of these systems, such signals primarily serve as auxiliary inputs for disambiguation rather than as direct mechanisms for modifying automation. In contrast, our work treats visual annotations as actionable repair primitives over an explicit plan, enabling users not only to clarify intent but also to inspect, revise, and redirect execution as tasks drift. We also use annotation to resolve ambiguity, but the role of annotation in Plover is broader than disambiguating a target. Here, an annotation becomes an input to plan repair: it constrains which pending step or step sequence should be revised, anchors the revision to the current screen state, and produces an explicit proposal that the user can inspect before execution resumes. Thus, visual marks are not only grounding cues for action selection; they are repair primitives over a persistent task representation.
3 PLOVER DESIGN 3.1 Problem Framing and Formative Study Autonomy-first GUI agents can often execute short tasks from natural language alone, but they remain difficult to supervise when execution unfolds over many steps in dynamic interfaces. In highstakes settings, users must do more than specify a goal: they need to inspect evolving plans, detect breakdowns, and intervene without discarding prior progress. Consider a compliance analyst using a legacy reporting portal to transfer values from a spreadsheet and PDF report into a multi-page form. She begins with a conventional autonomy-first GUI agent, which moves directly from instruction to action without exposing
Conference’17, July 2017, Washington, DC, USA
a stable plan she can inspect or revise (L1). Midway through execution, an unexpected modal dialog appears; the agent adapts, but does not reveal what changed or why (L2). Later, it enters a value into the wrong numeric field, yet the user has no precise, visually grounded way to redirect behavior at the point of failure (L3). When she attempts to correct the task through another prompt, the system broadly regenerates the workflow rather than modifying only the problematic portion (L4). Even when the agent reports completion, it does not clearly surface uncertainty, ambiguity, or intermediate failure (L5). This scenario is representative of workflows where users cannot fully ignore the automation outcome: the task spans multiple UI states, mistakes may silently propagate, and redoing the entire workflow is costly. In such cases, the relevant interaction problem is not whether the user should supervise every low-level action, but how the system can make the few necessary moments of supervision precise, localized, and recoverable. The scenario further illustrates five recurring limitations of autonomy-first GUI agents: L1 Opaque Planning, L2 Silent Adaptation, L3 Weak Grounding for Correction, L4 All-or-Nothing Correction, and L5 Weak Verification. These limitations suggest that the challenge is not simply improving autonomous execution, but designing an interaction model that keeps task structure visible, supports targeted intervention, and preserves continuity across repair. To better understand these breakdowns, we conducted a formative study with six participants from our university community who had prior experience with agentic systems or GUI automation tools. Participants (P1–P6) completed a representative GUI task using an early prototype of Plover (Appendix Fig. 6) while thinking aloud. Sessions lasted 30–60 minutes and covered prompt authoring, plan inspection and editing, execution monitoring, and multimodal correction of a deliberately introduced interface error. We collected screen recordings, verbal feedback, and post-task discussion. The study was deemed exempt by our University Institutional Review Board (IRB), and all participants provided informed consent. Across participants, we observed four recurring challenges that motivated the design of Plover. Planning Challenges: users struggled to form accurate mental models of proposed plans and plan revisions before execution. Execution and Inspection Challenges: most participants struggled to judge whether execution was making meaningful progress, especially when repeated actions produced little visible change. Intervention Challenges: users wanted correction mechanisms that were precise, localized, and visually grounded; notably, all participants found screenshot annotation more precise than language alone for spatial correction, although several remained uncertain about how such inputs would affect the underlying plan. System-level Challenges: users expected the system to detect obvious non-progress and assist with recovery instead of placing the full burden of diagnosis on them.
3.2
Design Goals
Our formative findings suggest that the main challenge in collaborative GUI automation is not only executing tasks correctly, but also helping users stay aligned with evolving execution. Across planning, execution, and intervention, participants repeatedly struggled to understand what the agent was about to do, determine whether
Conference’17, July 2017, Washington, DC, USA
Figure 1: The Plover system architecture. it was making progress, and apply corrections without losing prior work. We synthesize these findings into three design questions: Q1) How can systems maintain alignment between evolving user intent and execution across multi-step workflows? Q2) How can interfaces support precise, visually grounded correction when execution deviates? Q3) How can replanning be surfaced as a transparent and inspectable process rather than a hidden internal update? We then translate these questions into the following design goals: DG1 Interpretable Plans. Users should be able to understand, inspect, and revise plans before execution, with clear differentiation between plan versions and structure (Q1). DG2 Execution Awareness. The system should clearly communicate progress, failure, and state transitions during execution so that users can monitor behavior and recognize breakdowns (Q1, Q3). DG3 Grounded Local Repair. Users should be able to intervene through both high-level language and precise spatial grounding, while constraining updates to the relevant portion of the plan (Q2). DG4 Failure-aware Recovery. The system should identify nonprogress and support recovery through explicit replanning rather than silent adaptation (Q3). DG5 Continuity across Repair. Intervention and replanning should preserve executed history and maintain a consistent sense of task progression (Q1, Q3).
4 PLOVER SYSTEM 4.1 Overview Plover is a mixed-initiative system for plan-centric GUI automation. Its central design choice is to treat task plans as a persistent shared state representation [27] that coordinates planning, execution, and repair, rather than as transient internal reasoning. This choice addresses a key limitation of autonomy-first GUI agents: when plans evolve silently during execution, users have little visibility into what changed, why it changed, or how to intervene without restarting. By externalizing plans as persistent, editable artifacts [17, 59], Plover is designed to allow users to Review
Venkatesan et al.
plan updates for alignment, Revise workflows through targeted intervention, and Repair failures without restarting execution. As shown in Fig. 1, Plover separates planning from execution across three coordinated components: the Agentic Interface, the Planner Service, and the Executor Service. A task begins when the user issues a prompt through the Agentic Interface 1 , which forwards a Plan Request to the Planner Service 2 . The planner performs Plan Generation to produce a versioned Structured Plan that makes the intended task structure explicit. The plan is then passed to the Executor Service 3 , which grounds each step into concrete keyboard, mouse, and observation actions within a Remote GUI Environment. As execution proceeds, the executor reports Execution State back to the interface 4 , which surfaces progress to the user through Execution Feedback 5 . This shared plan representation enables a coordination mechanism we call Intelligent Replanning (IR): a visible planadaptation process that revises pending steps while preserving executed history. IR has two modes: User-Driven IR, triggered by user corrections, and System-Driven IR, triggered by nonprogress detection. In User-Driven IR, users intervene upon observing an error through Plan Edits, Natural Language Guidance, or Multimodal Annotation, 1 triggering a Replan Request 2 that revises only pending steps while preserving executed history. Conversely, in System-Driven IR, the executor autonomously detects non-progress via Non-progress Detection 1 , prompting the planner to propose a localized recovery step surfaced explicitly on the interface 2 . In both modes, revisions update the shared plan, allowing users to inspect the rationale and scope of changes. This coordination loop transforms GUI automation from a brittle one-shot interaction into a collaborative, repairable process.
4.2
Agentic Interface
The Agentic Interface is the primary interaction surface through which users inspect plans, monitor execution, and steer adaptation. Fig. 2 shows a typical interaction: users author a task (b), review and optionally revise the generated plan (c), monitor execution through screen state and step-level feedback (d, e), and intervene through Natural Language Guidance (b) or Multimodal Annotation (e1) when execution drifts. Resulting IR events are versioned (f) so that users can inspect how the workflow changes over time while preserving completed progress. Rather than treating planning, execution, and replanning as separate stages, the interface presents them as a continuous workflow over a shared plan representation. It therefore emphasizes three primary functions: to Review plan alignment, Revise steps via intervention, and Repair failures. Shared Plan Workspace. To address opaque planning and silent adaptation ( L1 , L2 ), Plover externalizes task structure in a shared workspace centered on the Planner Chat (Fig. 2b) and Steps Panel (Fig. 2c). Plans are represented as editable Plan Artifacts that users can inspect, revise, and approve before and during execution in the Regenerated Plan Panel (Fig. 2c2) ( DG1 ). This workspace is not merely a display of model output; it serves as the primary coordination object between the user and the agent. After the user specifies a task in the chat box, the planner generates a structured step sequence in the Steps Panel. Before execution, users may regenerate
Plover: Steering GUI Agents through Plan-Centric Interaction
Conference’17, July 2017, Washington, DC, USA
Figure 2: Plover Agentic Interface. (a) System Status Bar shows execution phase and plan state. (b) Planner Chat supports natural language interaction and surfaces intervention updates as banners. (c) Steps Panel displays the interactive plan, Derived Constraints (c1), Regenerated Plan Proposals (c2), and System-Driven IR proposals (c3). (d) Live View provides visual grounding of the target environment. (e) Execution Panel includes annotation tools (e1), execution feed (e2), step timeline (e3), and expanded activity logs (e4). (f) Run Timeline provides an iconographic view of execution progress and access to expanded branching plan visualization (f1; Fig. 3). alternatives and approve a preferred version to enter a locked execution state. During execution, revisions introduced through Natural Language Guidance or Multimodal Annotation are surfaced as explicit proposals with localized changes ( L3 ) rather than silently overwriting the active plan. By separating the current plan from proposed revisions, Plover provides a mechanism for users to evaluate repair before committing to it, reducing the all-or-nothing correction behavior ( L4 ). To further support interpretability, each replanning event updates Derived Constraints (Fig. 2c1), shown as a concise summary of the agent’s inferred understanding of user intent. A dedicated proposal view supports side-by-side comparison between the current and proposed plans and includes a concise diff summary of updated steps (Fig. 4d). Together, these mechanisms are designed to make revision reviewable and support continuity across repair ( DG5 ). Execution Legibility. To support execution awareness without overwhelming users ( DG2 ), the Executor Panel (Fig. 2e) presents agent behavior through a step-centered execution view. It combines a live screenshot panel (e1), a concise semantic activity feed (e2), and a lightweight step timeline (e3). The screenshot panel shows the current interface state and supports direct Multimodal Annotation when correction is needed (Fig. 4b). The activity feed summarizes low-level operations such as clicking, typing, scrolling, and waiting into short human-readable descriptions that foreground agent intent rather than implementation detail. The step timeline shows which steps are completed, active, or failed. This design
reflects a deliberate tradeoff: users need enough visibility to judge whether execution is making progress, but raw tool traces alone would impose unnecessary cognitive load. Plover therefore uses progressive disclosure ( L2 , L5 ): high-level execution summaries are visible by default, while a secondary drawer (Fig. 2e4) supports deeper inspection through detailed logs and reasoning traces. System Status Bar. To reduce ambiguity about current system behavior ( L2 ), Plover provides a persistent System Status Bar (Fig. 2a) that communicates execution state and plan state at a glance. It indicates whether the system is planning, replanning, executing, paused, failed, or waiting for user input, while separately representing the plan as Draft, Proposal, or Locked. The bar also surfaces step progress, current grounding context, connectivity to the execution environment, elapsed time, and execution mode. This persistent overview gives users a stable frame of reference during mixed-initiative interaction and supports both execution awareness and failure recovery ( DG2 , DG4 ). Run Timeline and Provenance. To support mental model continuity and transparent plan evolution ( DG2 , DG5 ), Plover visualizes execution and revision history through a two-level iconographic Provenance Component (Fig. 2f). This design addresses formative study findings where participants struggled to track plan changes and understand how interventions affected ongoing execution. A persistent Run Timeline (Fig. 2f) represents the active plan as a sequence of status-based icon nodes corresponding to pending,
Conference’17, July 2017, Washington, DC, USA
Venkatesan et al.
Figure 3: Provenance Bar visualizing branching plan revisions and evolution. executing, completed, failed, and IR recovery states. Small adjacent histograms summarize execution metrics (e.g., latency and screenshot counts), while between-node intervention markers denote user interventions through Multimodal Annotation or Natural Language Guidance. Expanding the timeline reveals a branching version graph of approved plan revisions [8, 46] (Fig. 2f1; Fig. 3). Each row corresponds to a version labeled by revision cause (e.g., initial plan, manual regeneration, instruction update, annotation update, or Intelligent Replanning), with edges connecting revisions to their origin points (similar to a Git-style branching history). This provenance view serves a broader purpose than history logging. Since mixed-initiative automation unfolds through successive revisions, users need to understand not only the current plan, but also how the workflow arrived there. By making plan evolution explicit, Plover turns replanning from a hidden recovery mechanism into an auditable interaction process ( L2 , L5 ).
4.3
Intelligent Replanning
A central mechanism exposed through the interface is Intelligent Replanning (IR) (Fig. 1), which treats plan adaptation as a visible and revisable process ( L1 ). IR allows users to Review plan updates and Revise workflows through User-Driven IR, while Repairing failures through System-Driven IR, in which the system detects non-progress and proposes localized recovery steps ( DG4 ). Both operate over the same shared plan representation, ensuring that updates remain inspectable, versioned, and minimally disruptive. User-Driven IR. User-Driven IR allows users to intervene during execution by providing corrective signals grounded in either language or visual context. These interventions generate localized updates that modify only pending steps while preserving executed history ( DG3 , DG5 ). The key design principle is that user feedback should not trigger monolithic regeneration unless the task itself has fundamentally changed. (1) Natural Language Guidance. Users may provide free-form instructions to revise execution, such as “select the second option instead” or “skip this step and go to export.” These inputs are interpreted as high-level intent corrections and translated into localized edits of the relevant plan segment. This provides flexibility and low interaction overhead, but it can remain ambiguous in visually dense interfaces where multiple candidate targets are plausible.
Figure 4: Multimodal Annotation workflow for resolving spatial ambiguity through localized, plan-centric repair. (2) Multimodal Annotation. To resolve ambiguity, Plover supports screenshot-based annotation (Fig. 4b). When the execution feed reveals execution drift (Fig. 4a), users can mark the screenshot using strokes, shapes, or text overlays, producing a multimodal artifact anchored in pixel space. The UI maintains annotation primitives A = {𝑎𝑘 } and computes a bounding box 𝑏 (A) = (𝑥 min, 𝑦min, 𝑤, ℎ) over all primitives. On save, the artifact consisting of the annotated screenshot 𝐼𝑡 and geometry 𝑏 (A), is submitted to the Planner Service, with confirmation via status banners in the Planner Chat (Fig. 4c). In Plover, this annotation constrains repair scope by prioritizing spatial signals over language. The resulting schemaconstrained request triggers localized modifications that preserve executed history, surfaced as a revised plan proposal grounded in the annotated region (Fig. 4d). Once the user approves the proposal, the Executor acts on the specified target (Fig. 4e). By favoring the smallest consistent update, User-Driven IR ensures precise corrections and execution continuity, with the entire intervention recorded as a persistent marker in the Run Timeline (Fig. 4f). System-Driven IR. In addition to explicit user intervention, Plover triggers System-Driven IR when execution appears to stall. Autonomous GUI agents often fail through repeated actions that produce little meaningful interface change. Detecting such failures reliably is non-trivial: repeated actions alone may reflect benign retries, while a visually static interface may simply indicate legitimate waiting. Inspired by works like [23, 76], Plover therefore combines behavioral and visual signals to detect non-progress conservatively. (1) Behavioral Loop Detection. The executor maintains a canonicalized sequence of recent actions H𝑡 = (𝑎 1, 𝑎 2, . . . , 𝑎𝑡 ), where actions are abstracted to coarse semantic types (e.g., click, scroll, wait). Repeated subsequences beyond a threshold indicate potential behavioral thrashing.
Plover: Steering GUI Agents through Plan-Centric Interaction
(2) Visual Non-Progress Verification. To determine whether these repeated actions changed the environment, the system evaluates perceptual changes in the interface. For each screenshot 𝑠𝑡 , a perceptual hash ℎ𝑡 = dHash(𝑠𝑡 ) is computed, and similarity between consecutive states is measured using Hamming distance 𝑑𝑡 = Hamming(ℎ𝑡 , ℎ𝑡 −1 ). If 𝑑𝑡 < 𝜏 over multiple steps, the interface is considered visually stable, indicating a likely lack of progress. (3) Recovery Policy. Execution is classified as stuck only when both behavioral repetition and visual stability are observed, reducing false positives from ordinary retries or transient waits. When triggered, the system halts the current step, marks it as Failed, and inserts a corrective step as an IR recovery step. The proposed recovery is surfaced with a short rationale so users can understand why replanning occurred and what change is proposed.
4.4
Planner
The Planner Service synthesizes and revises structured plans that mediate between user intent and GUI execution. Rather than regenerating behavior monolithically, Plover treats plans as persistent, reviewable artifacts whose executed history is preserved while remaining steps can be selectively revised, enabling interpretable planning and localized repair ( DG1 , DG3 , DG5 ). Architecture. The planner mediates communication between the interface, the language model, and the executor. On each invocation, the interface submits a structured interaction context and receives a schema-constrained plan update. The planner does not directly control the execution environment. Instead, the environment state is supplied through a serialized interaction context, including user instructions, plan history, and optional screenshot metadata. Initial planning operates over task instructions and conversation history, while replanning with Multimodal Annotation additionally incorporates the latest screenshot and annotation geometry. This separation is deliberate. Decoupling planning from direct actuation keeps plan updates reviewable at the representation level rather than allowing hidden policy shifts inside the executor. To manage token constraints, the service retains full plan history while limiting visual context to the three most recent screenshots. This structured update mechanism enables visible adaptation rather than silent internal changes ( L1 , L2 ). Plan Representation and State Invariants. The planner outputs responses using a two-block schema. The <analysis> block summarizes task context, while the <steps> block contains ordered <completed> and <pending> sequences. Completed steps represent immutable execution history, while pending steps define the editable remainder. Each step is expressed as a deterministic UI instruction to support direct grounding by the executor. We formalize the planner state at time 𝑡 as 𝑃𝑡 = (𝐶𝑡 , 𝑈𝑡 ), where 𝐶𝑡 denotes executed steps and 𝑈𝑡 the editable suffix. Given input 𝑋𝑡 , the planner computes 𝑃𝑡 +1 = Update(𝑃𝑡 , 𝑋𝑡 ) subject to the invariant 𝐶𝑡 +1 = 𝐶𝑡 . This invariant captures a core design principle in Plover: once a step has been executed and committed to shared history, repair should operate over the remaining suffix rather than silently rewriting the entire workflow. In practice, this constraint makes revisions localized, preserves provenance, and prevents the broad regeneration behavior ( DG3 , L4 ) that users found difficult to interpret
Conference’17, July 2017, Washington, DC, USA
in autonomy-first systems. When replanning is triggered either through user intervention or system detection, the planner applies schema-constrained updates conditioned on the current plan 𝑃 and available context, such as annotation artifacts (𝐼𝑡 , 𝑏 (A)). These updates occur at three granularities: a local edit (single-step modification), a regional patch (short sequence update), or a full replan (workflow restructuring). This hierarchy enables the planner to adapt flexibly while preserving plan continuity and interpretability.
4.5
Executor
The Executor Service translates structured plan steps into concrete GUI interactions while maintaining a strict separation between reasoning and actuation ( DG2 , DG4 ). Each executor implements a shared RPC contract exposing deterministic primitives—pointer actions, keyboard input, scrolling, waits, and observation endpoints such as screenshots. This abstraction serves two purposes. First, it enforces a clear boundary between high-level planning decisions and low-level execution, ensuring that task logic remains visible in the plan representation rather than being hidden inside execution heuristics [1, 89]. Second, it supports portability across heterogeneous GUI environments by keeping the planner interface stable even when the underlying actuation layer changes (see Appendix D). To support mixed-initiative control, Plover also provides a parallel human-access channel to the execution environment. However, the system prioritizes lightweight, plan-centric corrections, such as Natural Language Guidance and Multimodal Annotation, over full manual override, because these interventions preserve explicit task structure and provenance. Currently, manual interventions can alter execution behavior but are not yet incorporated back into the plan representation; integrating explicit plan updates for human overrides remains future work.
5
EVALUATION
We evaluate Plover as a system for recoverable automation through three complementary analyses of failure and repair: (1) whether autonomous agent failures are structurally recoverable under plancentric mixed-initiative interaction, (2) how stable autonomous execution and plan reconstruction remain in realistic multi-step workflows, and (3) which classes of failures are most amenable to repair. Rather than focusing only on end-to-end success rates, our evaluation tests Plover’s core claims: that many GUI-agent failures are structurally recoverable, exposing plans makes deviations easier to diagnose, and lightweight interventions can restore progress. The benchmark repair study (Section 5.1) measures recoverability under mixed-initiative interaction. The scenario-based stability study (Section 5.2) examines autonomous execution and plan alignment in realistic workflows. We then synthesize both through a failure taxonomy identifying when different repair mechanisms are most effective (Section 5.3). Together, these analyses show whether Plover improves outcomes, and how plancentric interaction reshapes GUI automation dynamics.
5.1
Benchmark Failure-Case Repair Analysis
We first evaluate the central claim of Plover: that many autonomous GUI-agent failures are recoverable when plans are externalized and repair is localized. To do so, we curated 38 tasks from the
Conference’17, July 2017, Washington, DC, USA
Auto State (a)
Venkatesan et al.
What Went Wrong
Auto
Agent performs
✗
Failure Type
Intervention
How It Was Repaired
MI
System-Driven IR
✓
detects
repeated actions Execution Drift
without meaningful
non-progress and proposes a recovery
progress Repeated actions
step with its without progress reasoning
(b)
Agent fails to detect
✗
User performs
✓
Multimodal
the correct Perception Error
dropdown element
Annotation to ground the agent
Fails to detect UI element
(c)
Plan generates
✗
User directly
incorrect
✓
corrects the error in Planning Error
intermediate step
the editable plan prior to execution
Incorrect step sequence
(d)
Agent misinterprets
✗
User clarifies via
✓
brief Natural
that the correct cell State Misinterpretation was clicked and
Language Guidance
claims success Incorrect understanding of state
(e)
Compound errors
✗
Multiple
propagate through
✗
interventions Compound
dependent steps
unable to recover accumulated drift
Compound errors accumulate
Figure 5: Representative failure archetypes and recovery pathways in Plover. Each row shows the autonomous failure state (left) and the corresponding mixed-initiative outcome (right). Rows (a–d) illustrate distinct failure categories and their targeted repair mechanisms, while row (e) shows a compound failure where accumulated errors prevented recovery. OSWorld-Verified benchmark that were previously reported as failures for Claude 4.5 Sonnet using native computer-use capabilities [78]. OSWorld-Verified is a curated subset of OSWorld with verified task definitions, initial states, and evaluation procedures. These tasks, detailed in Appendix E, span heterogeneous environments, including web browsers and desktop applications such as LibreOffice, and emphasize multi-step workflows known to challenge vision-based automation [5, 14, 47]. We first ran all tasks in a fully autonomous setting to establish a baseline. Of the 38 tasks, 26
resulted in non-success outcomes (partial or failure); these cases formed the basis for our repair analysis. We then re-ran the 26 autonomous non-success cases in a mixedinitiative setting, where an expert (first author) provided targeted interventions—richer than mere prompt re-issuance—using Natural Language Guidance, Multimodal Annotation, and Plan Edits, while Plover triggered System-Driven IR when drift was detected. Interventions were lightweight and localized: naturallanguage guidance typically used 1–2 sentences clarifying spatial
Plover: Steering GUI Agents through Plan-Centric Interaction
Conference’17, July 2017, Washington, DC, USA
Table 1: Summary of mixed-initiative improvements over autonomous non-success trials. S = Success, P = Partial Success, F = Failure. Autonomous
Mixed-Initiative
P/F
#Trials
S/P/F
Improv. Rate
Browser Calc Writer Multi-App
2P / 2F 1P / 6F 2P / 3F 5P / 5F
4 7 5 10
4S / 0P / 0F 3S / 3P / 1F 4S / 1P / 0F 6S / 2P / 2F
100% 86% 100% 80%
1.75 1.86 1.60 2.40
Total
10P / 16F
26
17S / 6P / 3F
–
–
Average
–
–
–
88%
2.04
App
Avg. Int.
references or correcting element selection (e.g., “use the supplier dropdown on the left, not the archived list”)(<15s), multimodal annotations marked specific screen regions and sometimes added a short text label (<30s), and plan edits involved editing, adding, removing, or reordering steps in the initial or pending plan(<20s). No intervention required code, system internals, or domain knowledge beyond what was visible in the interface. This setup establishes an upper bound on plan-centric recoverability rather than typical user performance. That choice is intentional: our goal is to test whether failures are structurally recoverable through plancentric interaction, not how quickly naive users can identify and apply corrections. Establishing recoverability as a property of the failure class is a prerequisite for later user-facing studies [50]. Mixed-initiative interaction improved 23 of 26 autonomous nonsuccess cases, converting 17 to complete success and 6 to partial success (completing subtasks but not the goal), with only 3 remaining failures (Table 1). Recovery required an average of 2.04 interventions per task. Notably, all 10 autonomous partial successes were converted to complete success, and no regressions were observed. Improvements were consistent across application types, with 100% recovery in browser and writer tasks, and substantial gains in Calc (86%) and multi-application workflows (80%). Together, these findings suggest that many GUI-agent failures are not irrecoverable planning errors, but localized breakdowns in grounding, state interpretation, or execution continuity that become structurally recoverable when plans are externalized and repair is localized.
5.2
Scenario-Based Stability Analysis
We next examine the complementary setting in which no user intervenes. Whereas the repair study tests recoverability, this analysis probes how well Plover maintains execution stability and plan alignment in realistic multi-step workflows. To do so, we constructed five practical scenarios involving dense interaction patterns such as constrained typing, multi-step filtering, and application navigation. Rather than manually authoring prompts, we generated realistic task instructions from sampled interaction trajectories and replayed them through Plover without human intervention. This setup tests whether Plover can reconstruct plausible task structure and maintain execution alignment in autonomy-only workflows. Scenario details are provided in Appendix E. Trajectory Sampling and Prompt Reconstruction. To generate diverse task instances, we built a constrained exploration pipeline
sampling plausible interaction trajectories within each scenario, following established methods for automated GUI data collection [22]. For each environment, we specify initialization commands, curated clickable regions, restricted regions, and typing regions paired with text inputs. These constraints define valid interaction spaces over screenshots. During exploration, the system captures screenshots, queries a vision-based inference endpoint with Omniparser [41] for candidate actions, and executes only actions allowed by the scenario specification. We ran 100 exploration trials per scenario with randomized trajectory lengths of 5–15 steps, then filtered out redundant traces and retained 5 diverse trajectories per scenario after manual verification, yielding 25 trajectories overall [35]. We convert each retained trajectory into a natural-language task description using Claude Sonnet 4.5, employing a reverse task synthesis approach to infer user intent [62]. We then replay each generated prompt through the full Plover pipeline without human intervention. This setup tests whether Plover can reconstruct a plausible plan and reach a comparable end state from prompts inferred from behavior rather than manually authored instructions. Visual Fidelity. Autonomous replay proved substantially more stable in browser workflows than in desktop workflows. To assess replay fidelity, we compare the final interface state produced during replay against the reference trajectory using structural similarity (SSIM) [70], mean squared error (MSE), and perceptual hashing (dHash) [43]. As shown in Table 2, Firefox-based scenarios achieve consistently high similarity (SSIM > 0.98), whereas LibreOffice scenarios exhibit substantially lower similarity (SSIM 0.61–0.69) and higher variation (Appendix E). We treat image similarity as a proxy for execution stability rather than task success. These results suggest that browser workflows are relatively stable and reproducible, while desktop workflows are more sensitive to intermediate state differences, layout changes, and modal interactions. This gap highlights when autonomy breaks down and why plan-centric repair becomes necessary, particularly in desktop settings. Plan Alignment. Autonomously generated plans were often executable but structurally imperfect. We compare generated plans against reference trajectories using phase coverage, order alignment, redundancy, and actionability, following the Goal-Plan-Action (GPA) alignment framework [34]. To enable comparison, both plans and trajectories are mapped to a shared interface-agnostic phase taxonomy (e.g., navigation, filtering, selection), collapsing consecutive duplicates. Coverage measures recovered phases, order measures sequence agreement, redundancy captures repeated phases [21], and actionability measures executable steps. As shown in Table 2, actionability remains high (0.97), while coverage (0.62) and order alignment (0.41) are only moderate. This gap between step-level executability and workflow-level alignment [3] is exactly the space that Plover targets: plans may remain locally plausible even when they are globally misaligned, creating an opportunity for users to inspect and correct them before deviations propagate. Non-trivial redundancy (0.33) further suggests that predicted plans often reach subgoals through reordered or repetitive paths, which in autonomy-only settings can accumulate into drift or partial completion [21, 71]. We also observe that higher redundancy correlates with lower replay fidelity, suggesting that
Conference’17, July 2017, Washington, DC, USA
Venkatesan et al.
Table 2: Replay-based image similarity and plan comparison metrics across scenarios. SSIM, Coverage, Order, Redundancy, and Actionability range from 0–1; dHash ranges from 0–64; MSE is unbounded. Higher SSIM, Coverage, Order, and Actionability indicate better alignment, while lower MSE, dHash, and Redundancy indicate better performance. Visual Fidelity
Plan Alignment
Scenario
SSIM ↑
MSE ↓
dHash ↓
Coverage ↑
Order ↑
Redundancy ↓
Actionability ↑
Firefox Fillable Form Firefox Machine Dashboard Firefox Visualization Dashboard LibreOffice Incident Sheet LibreOffice Sensor Logs Sheet
0.9817 0.9809 0.9386 0.6094 0.6855
122.43 105.58 448.10 2226.31 1795.41
2.0 1.2 7.8 14.2 10.0
0.72 0.42 0.65 0.66 0.67
0.42 0.42 0.44 0.35 0.42
0.39 0.33 0.21 0.27 0.44
1.00 0.94 0.94 1.00 0.97
Overall Average
0.8392
939.57
7.04
0.62
0.41
0.33
0.97
repeated phases often reflect unsuccessful recovery attempts. Trial-level results are reported in Appendix E.
5.3
Characterizing Repairable Failures
To understand when mixed-initiative interaction is most effective, we analyzed the 26 autonomous non-success cases and examined how different interventions restored progress. Fig. 5 summarizes representative failure archetypes and their corresponding repair pathways. A key observation emerges: repairability is strongly tied to the locality of the failure. Breakdowns such as execution drift, perception errors, planning errors, and state misinterpretation (Fig.5a-d) remain recoverable through targeted, localized interventions. In contrast, compound failures (Fig. 5e), where multiple errors propagate across dependent steps, proved significantly harder to repair. Here, failure was no longer localized to a single action or state, but distributed across the plan, execution history, and current interface context. Under these conditions, each repair intervention provided only partial corrective leverage, since none could fully reconstruct the lost structural alignment of the task. This suggests that plan-centric interaction is most effective for localized breakdowns, but becomes fundamentally constrained when early errors cascade into globally misaligned workflows. Planning Errors and Execution Drift. Execution Drift emerged as the primary failure mode, appearing in 46% (𝑛 = 12/26) of cases. One-third of these instances (𝑛 = 4) were Compound Failures, where drift stemmed from underlying Planning Errors. These compound cases were the most difficult, accounting for two of the only three unrecovered tasks (𝑛 = 3/26). Despite this complexity, Plover effectively arrested divergent behavior in the remaining instances. Natural Language Guidance was the most frequent recovery channel (resolving 11 cases, 5 exclusively), while System-Driven IR and Multimodal Annotation each successfully resolved 9 cases (3 each exclusively). These findings suggest that while compound errors represent the current boundary of agent autonomy, Plover’s key contribution is that explicit, localized replanning can restore progress before drift becomes terminal. Perception and Grounding Breakdowns. Perception Errors (𝑛 = 7/26) and Action Grounding Failures (𝑛 = 5) occurred when agents selected incorrect UI elements despite correct intent. These failures were particularly common in visually dense interfaces such as LibreOffice and multi-application workflows. Most were recoverable
through Natural Language Guidance or Multimodal Annotation, which provided contextual clarification or spatial grounding. Replay analysis further showed lower visual fidelity in LibreOffice scenarios (SSIM as low as 0.61), reflecting increased visual ambiguity and sensitivity to intermediate state differences. State Misinterpretation and Missing Context. State Misinterpretation (𝑛 = 5/26) and Missing Context (𝑛 = 2) occurred when agents misinterpreted modal dialogs, document states, or required domainspecific inputs. These failures were mostly resolved through brief Natural Language Guidance, suggesting that many breakdowns stemmed from incomplete local context rather than fundamentally incorrect task decomposition. Scenario-Based Patterns. Failure patterns also varied across domains. Browser tasks achieved full recovery across all non-success cases, while multi-application workflows were the most challenging, accounting for 2 of the 3 remaining failures and requiring the highest number of interventions. In these harder cases, the initial planning error produced a fundamentally misaligned decomposition that propagated through dependent steps. Even after user- and system-driven interventions, the accumulated divergence was too large to recover through localized edits alone, suggesting that some failure modes require earlier intervention during planning rather than mid-execution repair. Replay metrics support the same pattern, with browser scenarios achieving higher visual similarity (SSIM ≈ 0.98) compared to LibreOffice scenarios (SSIM 0.61–0.69). Despite these differences, high actionability (0.97) suggests that plans were technically sound, though valid planning alone did not guarantee successful execution grounding when alignment was imperfect. Overall, failures were typically incremental and locally repairable. Mixed-initiative interaction improved 88% of non-success cases with an average of 2.04 interventions per task, and most improvements occurred after lightweight, localized corrections. Ultimately, these results indicate that GUI agent reliability depends less on perfect autonomy and more on providing the plan-centric visibility required to transform latent drift into a collaborative and repairable interaction.
6
Discussion
From One-Shot Delegation to Ongoing Alignment. Our findings suggest that one promising direction for reliable GUI automation is to support ongoing alignment between user intent and agent
Plover: Steering GUI Agents through Plan-Centric Interaction
behavior, rather than relying entirely on one-shot delegation. Here, automation is not only about handing a task to an agent and waiting for a final result, but also enabling users to inspect, guide, and repair execution as it unfolds. Success, therefore, depends not only on the agent’s raw capability but also on whether its understanding can be made visible and revised when needed. By treating the plan as a coordination interface rather than only an explanation of reasoning, plan-centric interaction creates a shared workspace where users and agents can easily realign when execution begins to drift. Reducing the Burden of Supervision. A related implication is that reliable GUI automation requires reducing supervision effort during execution, sometimes described as intervention fatigue [69]. In opaque automation, users often need to reconstruct system state to determine what went wrong and how to respond. Plan-centric interaction may reduce this burden by making the agent’s intermediate intent and execution state explicit throughout the task. When plans are persistent and revisable, users can focus on the workflow segment that has drifted rather than re-checking the entire task from the beginning. This suggests that future agent interfaces may aim not only to reduce how often users intervene, but also to make each intervention simpler and less mentally demanding. Matching Repair to the Level of Failure. A further implication is that effective repair depends on where breakdowns occur and what kind of precise correction they require. No single repair modality is sufficient across all failures. Natural language is useful for clarifying high-level intent, but can be too underspecified for visually ambiguous interfaces. Spatial or visual grounding can resolve such ambiguities more directly, but may not fully convey the broader rationale behind a correction. Multimodal repair is therefore valuable because it can help users to intervene at the level most appropriate to the failure, combining semantic guidance with spatial specificity when needed. More broadly, this suggests that repair mechanisms should support multiple and flexible forms of input while remaining sensitive to when each is most useful. Limitations and Future Work. Our results also clarify the scope of this approach. Plan-centric interaction appears especially useful for long-horizon, multi-step, or drift-prone tasks, where preserving partial progress and supporting targeted repair are valuable. For shorter or more routine workflows, however, the overhead of inspecting and revising plans may outweigh the benefits. In addition, our design assumes users generally know the intended path well enough to correct the agent when needed. An important direction for future work is supporting settings where the task structure is itself uncertain, allowing users and agents to co-construct plans dynamically rather than only revising existing ones. Another open question is how visible plans affect user trust and verification behavior. While externalized plans can improve inspectability, they may also encourage users to accept progress too readily when the system appears confident. Future interfaces may therefore need stronger support for uncertainty communication, such as highlighting fragile steps, signaling low-confidence decisions, or prompting verification at points where errors are likely.
Conference’17, July 2017, Washington, DC, USA
7
CONCLUSION
We present Plover, a plan-centric framework for mixed-initiative GUI automation that externalizes plans as persistent, editable artifacts and supports intervention through plan revision, multimodal interaction, and intelligent replanning. Our evaluation shows that 88% of autonomous failures are recoverable through lightweight, localized interventions, while scenario-based analysis reveals domaindependent stability gaps that motivate human–agent collaboration. Together, these findings suggest that many GUI-agent failures are structurally repairable, and that plan-centric interaction can support more transparent, adaptable, and reliable supervision and repair in GUI automation systems.
Acknowledgments This work was supported in part by the U.S. National Science Foundation under Grant No. IIS-2427770 and by a Bosch Research Award.
References [1] Mohamed Aghzal, Gregory J Stein, and Ziyu Yao. 2026. Why do LLM-based web agents fail? A hierarchical planning perspective. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 32157–32180. [2] Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13. [3] Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, and Jordan Lee Boyd-Graber. 2025. A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, 11568–11595. [4] Tanvir Bhathal and Asanshay Gupta. 2025. Websight: A vision-first architecture for robust web agents. arXiv preprint arXiv:2508.16987 (2025). [5] Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault L De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks. Advances in Neural Information Processing Systems 37 (2024), 5996–6051. [6] Sacha Brisset, Romain Rouvoy, Lionel Seinturier, and Renaud Pawlak. 2022. Erratum: Leveraging Flexible Tree Matching to repair broken locators in web automation scripts. Inf. Softw. Technol. 144 (2022), 106754. doi:10.1016/J.INFSOF. 2021.106754 [7] Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. 2024. ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024. IEEE, 12914–12923. doi:10.1109/CVPR52733.2024.01227 [8] Tathagata Chakraborti, Kshitij P. Fadnis, Kartik Talamadupula, Mishal Dholakia, Biplav Srivastava, Jeffrey O. Kephart, and Rachel K. E. Bellamy. 2018. Visualizations for an Explainable Planning Agent. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden, Jérôme Lang (Ed.). ijcai.org, 5820–5822. [9] Tathagata Chakraborti, Sarath Sreedharan, Sachin Grover, and Subbarao Kambhampati. 2019. Plan explanations as model reconciliation–an empirical study. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). Ieee, 258–266. [10] Juntong Chen, Jiang Wu, Jiajing Guo, Vikram Mohanty, Xueming Li, Jorge Piazentin Ono, Wenbin He, Liu Ren, and Dongyu Liu. 2025. InterChat: Enhancing generative visual analytics using multimodal interactions. In Computer Graphics Forum, Vol. 44. Wiley Online Library, e70112. [11] Weihao Chen, Chun Yu, Huadong Wang, Zheng Wang, Lichen Yang, Yukun Wang, Weinan Shi, and Yuanchun Shi. 2023. From gap to synergy: Enhancing contextual understanding through human-machine collaboration in personalized systems. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–15. [12] Ziming Cheng, Zhiyuan Huang, Junting Pan, Zhaohui Hou, and Mingjie Zhan. 2025. Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions. arXiv preprint arXiv:2503.24180 (2025).
Conference’17, July 2017, Washington, DC, USA
[13] Riccardo Coppola, Luca Ardito, and Marco Torchiano. 2019. Fragility of layout-based and visual GUI test scripts: an assessment study on a hybrid mobile application. In Proceedings of the 10th ACM SIGSOFT International Workshop on Automating TEST Case Design, Selection, and Evaluation. ACM, 28–34. doi:10.1145/3340433.3342824 [14] Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vázquez, Nicolas Chapados, and Alexandre Lacoste. 2024. WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks? Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 235 (2024), 11642–11662. [15] Anurag Dwarakanath, Neville Dubash, and Sanjay Podder. 2018. Machines that test Software like Humans. CoRR abs/1809.09455 (2018). arXiv:1809.09455 http://arxiv.org/abs/1809.09455 [16] Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. In Proceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu (Eds.). PMLR, 15419–15462. [17] KJ Kevin Feng, Kevin Pu, Matt Latzke, Tal August, Pao Siangliulue, Jonathan Bragg, Daniel S Weld, Amy X Zhang, and Joseph Chee Chang. 2026. Cocoa: Co-planning and co-execution with ai agents. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, 1–23. [18] Boyu Gou, Demi Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2025. Navigating the digital world as humans do: Universal visual grounding for gui agents. In International Conference on Learning Representations, Vol. 2025. 30851–30883. [19] Nitesh Goyal, Minsuk Chang, and Michael Terry. 2024. Designing for HumanAgent Alignment: Understanding what humans want from their agents. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–6. [20] Maria Fernanda Granda, Otto Parra, and Bryan Alba-Sarango. 2021. Towards a Model-Driven Testing Framework for GUI Test Cases Generation from User Stories.. In ENASE. 453–460. [21] Danil S. Grigorev, Alexey K. Kovalev, and Aleksandr I. Panov. 2025. VerifyLLM: LLM-Based Pre-Execution Task Plan Verification for Robots. In IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2025, Hangzhou, China, October 19-25, 2025. IEEE, 18489–18496. [22] Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent. CoRR abs/2511.06417 (2025). arXiv:2511.06417 doi:10.48550/ARXIV.2511.06417 [23] Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-aware gui agent: Adaptive perception through component recommendation and human-in-theloop refinement. arXiv preprint arXiv:2508.04025 (2025). [24] Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to-end web agent with large multimodal models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 6864–6890. [25] Theodore D Hellmann and Frank Maurer. 2011. Rule-based exploratory testing of graphical user interfaces. In 2011 Agile Conference. IEEE, 107–116. [26] Peter Hofmann, Caroline Samp, and Nils Urbach. 2020. Robotic process automation. Electronic markets 30, 1 (2020), 99–106. [27] Eric Horvitz. 1999. Principles of Mixed-Initiative User Interfaces. In Proceeding of the CHI ’99 Conference on Human Factors in Computing Systems: The CHI is the Limit, Pittsburgh, PA, USA, May 15-20, 1999, Marian G. Williams and Mark W. Altom (Eds.). ACM, 159–166. doi:10.1145/302979.303030 [28] Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The dawn of gui agent: A preliminary case study with claude 3.5 computer use. arXiv preprint arXiv:2411.10323 (2024). [29] Wenyue Hua, Mengting Wan, Jagannath Vadrevu, Ryan Nadel, Yongfeng Zhang, and Chi Wang. 2025. Interactive speculative planning: Enhance agent efficiency through co-design of system and user interface. In International Conference on Learning Representations, Vol. 2025. 14256–14283. [30] Zhiyuan Huang, Ziming Cheng, Junting Pan, Zhaohui Hou, and Mingjie Zhan. 2025. Spiritsight agent: Advanced gui agent with one look. In Proceedings of the Computer Vision and Pattern Recognition Conference. 29490–29500. [31] Faria Huq, Zora Zhiruo Wang, Frank F Xu, Tianyue Ou, Shuyan Zhou, Jeffrey P Bigham, and Graham Neubig. 2025. Cowpilot: a framework for autonomous and human-agent collaborative web navigation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations). 163–172. [32] Alayt Issak, Jeba Rezwana, and Casper Harteveld. 2025. MOSAAIC: Managing Optimization towards Shared Autonomy, Authority, and Initiative in Co-creation. arXiv:2505.11481 [cs.AI] https://arxiv.org/abs/2505.11481 [33] Arushi Jain, Shubham Paliwal, Monika Sharma, Lovekesh Vig, and Gautam Shroff. 2024. SmartFlow: Robotic Process Automation using LLMs. arXiv:2405.12842 [cs.RO] https://arxiv.org/abs/2405.12842
Venkatesan et al.
[34] Allison Sihan Jia, Daniel Huang, Nikhil Vytla, Nirvika Choudhury, Shayak Sen, John C. Mitchell, and Anupam Datta. 2025. What Is Your Agent’s GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment. CoRR abs/2510.08847 (2025). [35] Linjia Kang, Zhimin Wang, Yongkang Zhang, Duo Wu, Jinghe Wang, Ming Ma, Haopeng Yan, and Zhi Wang. 2026. Learning with Challenges: Adaptive Difficulty-Aware Data Generation for Mobile GUI Agent Training. arXiv preprint arXiv:2601.22781 (2026). [36] Anjali Khurana, Xiaotian Su, April Yi Wang, and Parmit K. Chilana. 2025. Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26 April 2025- 1 May 2025, Naomi Yamashita, Vanessa Evers, Koji Yatani, Sharon Xianghua Ding, Bongshin Lee, Marshini Chetty, and Phoebe O. Toups Dugas (Eds.). ACM, 880:1–880:18. doi:10.1145/3706598.3713431 [37] Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 881–905. [38] Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017. SUGILITE: creating multimodal smartphone automation by demonstration. In Proceedings of the 2017 CHI conference on human factors in computing systems. 6038–6049. [39] Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, and Mike Zheng Shou. 2025. Showui: One vision-language-action model for gui visual agent. In Proceedings of the Computer Vision and Pattern Recognition Conference. 19498–19508. [40] Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu, and Lydia B Chilton. 2025. DoubleAgents: Interactive Simulations for Alignment in Agentic AI. arXiv preprint arXiv:2509.12626 (2025). [41] Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Hassan Awadallah. 2025. OmniParser for Pure Vision Based GUI Agent. https://openreview.net/forum? id=C6hUK6Q1Pi [42] Shang Ma, Xusheng Xiao, and Yanfang Ye. 2025. Agent+ P: Guiding UI Agents via Symbolic Planning. arXiv preprint arXiv:2510.06042 (2025). [43] Jordan Madden, Moxanki Bhavsar, Lhamo Dorje, and Xiaohua Li. 2024. Robustness of Practical Perceptual Hashing Algorithms to Hash-Evasion and HashInversion Attacks. In The Third Workshop on New Frontiers in Adversarial Machine Learning. https://openreview.net/forum?id=hraOxsleRl [44] Valérie Maquil, Dimitra Anastasiou, Hoorieh Afkari, Adrien Coppens, Johannes Hermen, and Lou Schwartz. 2023. Establishing Awareness through Pointing Gestures during Collaborative Decision-Making in a Wall-Display Environment. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems, CHI EA 2023, Hamburg, Germany, April 23-28, 2023, Albrecht Schmidt, Kaisa Väänänen, Tesh Goyal, Per Ola Kristensson, and Anicia Peters (Eds.). ACM, 104:1–104:7. [45] Hussein Mozannar, Gagan Bansal, Cheng Tan, Adam Fourney, Victor Dibia, Jingya Chen, Jack Gerrits, Tyler Payne, Matheus Kunzler Maldaner, Madeleine Grunde-McLaughlin, et al. 2025. Magentic-ui: Towards human-in-the-loop agentic systems. arXiv preprint arXiv:2507.22358 (2025). [46] Arpit Narechania, Shunan Guo, Eunyee Koh, Alex Endert, and Jane Hoffswell. 2025. Utilizing Provenance as an Attribute for Visual Data Analysis: A Design Probe With ProvenanceLens. IEEE Trans. Vis. Comput. Graph. 31, 10 (2025), 8452–8465. [47] Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UIVision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction. In Proceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267), Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu (Eds.). PMLR, 45817–45851. [48] Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Md. Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A. Rossi, and Franck Dernoncourt. 2025. GUI Agents: A Survey. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025 (Findings of ACL, Vol. ACL 2025), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, 22522–22538. https://aclanthology.org/2025.findings-acl.1158/ [49] Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, and Amy Pavel. 2025. Morae: Proactively Pausing UI Agents for User Choices. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST 2025, Busan, Korea, 28 September 2025 - 1 October 2025, Andrea Bianchi, Elena L. Glassman, Wendy E. Mackay, Shengdong Zhao, Jeeeun Kim, and Ian Oakley (Eds.). ACM, 198:1–198:14. doi:10.1145/3746059.3747797
Plover: Steering GUI Agents through Plan-Centric Interaction
[50] Christopher Potts and Moritz Sudhof. 2026. Invisible failures in human-AI interactions. arXiv preprint arXiv:2603.15423 (2026). [51] Petr Průcha, Michaela Matoušková, and Jan Strnad. 2025. Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows. arXiv preprint arXiv:2509.04198 (2025). [52] Ju Qian, Zhengyu Shang, Shuoyan Yan, Yan Wang, and Lin Chen. 2020. Roscript: a visual script driven truly non-intrusive robotic testing system for touch screen applications. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 297–308. [53] Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian Li, Haihua Yang, Haifeng Liu, Feng Lin, Tao Peng, Xin Liu, and Guang Shi. 2025. UI-TARS: Pioneering Automated GUI Interaction with Native Agents. CoRR abs/2501.12326 (2025). arXiv:2501.12326 doi:10.48550/ARXIV.2501.12326 [54] Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, and Arvind Narayanan. 2026. Towards a Science of AI Agent Reliability. arXiv preprint arXiv:2602.16666 (2026). [55] Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From pixels to ui actions: Learning to follow instructions via graphical user interfaces. Advances in Neural Information Processing Systems 36 (2023), 34354– 34370. [56] Erfan Shayegani, Keegan Hines, Yue Dong, Nael Abu-Ghazaleh, Roman Lutz, Spencer Whitehead, Vidhisha Balachandran, Besmira Nushi, and Vibhav Vineet. 2026. Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness. In The Fourteenth International Conference on Learning Representations. https: //openreview.net/forum?id=9W4bPRsEIT [57] Minjie Shen, Yanshu Li, Lulu Chen, and Qikai Yang. 2025. From mind to machine: The rise of manus ai as a fully autonomous digital agent. arXiv preprint arXiv:2505.02024 (2025). [58] Aleksandar Shtedritski, Christian Rupprecht, and Andrea Vedaldi. 2023. What does CLIP know about a red circle? Visual prompt engineering for VLMs. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023. IEEE, 11953–11963. doi:10.1109/ICCV51070.2023.01101 [59] Kihoon Son, Hyewon Lee, DaEun Choi, Yoonsu Kim, Tae Soo Kim, Yoonjoo Lee, John Joon Young Chung, HyunJoon Jung, and Juho Kim. 2026. " When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction. arXiv preprint arXiv:2603.02050 (2026). [60] Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. Visiontasker: Mobile task automation using vision based ui understanding and llm task planning. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17. [61] Zihe Song, S. M. Hasan Mansur, Ravishka Rathnasuriya, Yumna Fatima, Wei Yang, Kevin Moran, and Wing Lam. 2025. Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging. In 12th IEEE/ACM International Conference on Mobile Software Engineering and Systems, MOBILESoft@ICSE 2025, Ottawa, ON, Canada, April 27-28, 2025. IEEE, 32–43. doi:10.1109/ MOBILESOFT66462.2025.00011 [62] Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, 5555–5579. [63] Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, et al. 2025. A survey on (m) llm-based gui agents. arXiv preprint arXiv:2504.13865 (2025). [64] Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, et al. 2026. Dark patterns meet gui agents: Llm agent susceptibility to manipulative interfaces and the role of human oversight. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (2026), 1–26. [65] Yuanrong Tang, Huiling Peng, Bingxi Zhao, Hengyang Ding, Hanchao Song, Tianhong Wang, Chen Zhong, and Jiangtao Gong. 2026. Human Tool: An MCPStyle Framework for Human-Agent Collaboration. arXiv preprint arXiv:2602.12953 (2026). [66] Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al. 2026. Kimi K2. 5: Visual Agentic Intelligence. arXiv preprint arXiv:2602.02276 (2026). [67] Ching-Yi Tsai, Nicole Tacconi, Andrew D Wilson, and Parastoo Abtahi. 2026. Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target Selection. arXiv preprint arXiv:2602.13433 (2026).
Conference’17, July 2017, Washington, DC, USA
[68] Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, et al. 2024. A Survey on Human-AI Collaboration with Large Foundation Models. arXiv preprint arXiv:2403.04931 (2024). [69] Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu. 2026. Computer Agent Arena: Toward Human-Centric Evaluation and Analysis of Computer-Use Agents. In The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=3x4SDbXbgl [70] Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. 13, 4 (2004), 600–612. [71] Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu. 2025. PlanGenLLMs: A Modern Survey of LLM Planning Capabilities. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, 19497–19521. [72] Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. Autodroid: Llm-powered task automation in android. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking. 543–557. [73] Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, and Wei Chen. 2025. Exploring Multimodal Prompt for Visualization Authoring with Large Language Models. CoRR abs/2504.13700 (2025). doi:10.48550/ARXIV.2504.13700 [74] Judith Wewerka and Manfred Reichert. 2020. Robotic Process Automation - A Systematic Literature Review and Assessment Framework. CoRR abs/2012.11951 (2020). arXiv:2012.11951 https://arxiv.org/abs/2012.11951 [75] Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Ruiyi Zhang, Subrata Mitra, Dimitris N. Metaxas, Lina Yao, Jingbo Shang, and Julian J. McAuley. 2024. Visual Prompting in Multimodal Large Language Models: A Survey. CoRR abs/2409.15310 (2024). arXiv:2409.15310 doi:10.48550/ARXIV.2409.15310 [76] Penghao Wu, Shengnan Ma, Bo Wang, Jiaheng Yu, Lewei Lu, and Ziwei Liu. 2026. Gui-reflection: Empowering multimodal gui models with self-reflection behavior. Advances in Neural Information Processing Systems 38 (2026), 101861–101896. [77] Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. 2025. OSATLAS: Foundation Action Model for Generalist GUI Agents. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https://openreview.net/forum?id=n9PDaFNi8t [78] Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.). [79] Yuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun, and Xin Tong. 2026. DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI 2026, Barcelona, Spain, April 13-17, 2026 (2026), 305:1–305:22. [80] John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems 37 (2024), 50528–50652. [81] Yuhao Yang, Zhen Yang, Zi-Yi Dou, Anh Nguyen, Keen You, Omar Attia, Andrew Szot, Michael Feng, Ram Ramrakhya, Alexander Toshev, et al. 2025. Ultracua: A foundation model for computer use agents with hybrid action. arXiv preprint arXiv:2510.17790 (2025). [82] Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems 35 (2022), 20744–20757. [83] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations. [84] Tom Yeh, Tsung-Hsiang Chang, and Robert C. Miller. 2009. Sikuli: using GUI screenshots for search and automation. In Proceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology, Victoria, BC, Canada, October 4-7, 2009, Andrew D. Wilson and François Guimbretière (Eds.). ACM, 183–192. doi:10.1145/1622176.1622213 [85] Ryan Yen, Jian Zhao, and Daniel Vogel. 2025. Code Shaping: Iterative Code Editing with Free-form AI-Interpreted Sketching. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26
Conference’17, July 2017, Washington, DC, USA
April 2025- 1 May 2025, Naomi Yamashita, Vanessa Evers, Koji Yatani, Sharon Xianghua Ding, Bongshin Lee, Marshini Chetty, and Phoebe O. Toups Dugas (Eds.). ACM, 872:1–872:17. doi:10.1145/3706598.3713822 [86] Shengcheng Yu, Chunrong Fang, Mingzhe Du, Yuchen Ling, Zhenyu Chen, and Zhendong Su. 2024. Practical Non-Intrusive GUI Exploration Testing with Visualbased Robotic Arms. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, Portugal, April 14-20, 2024. ACM, 130:1–130:13. doi:10.1145/3597503.3639161 [87] Hyeonggeun Yun and Jinkyu Jang. 2025. Interaction-Driven Browsing: A Humanin-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents. CoRR abs/2509.12049 (2025). doi:10.48550/ARXIV.2509. 12049 [88] Shaojie Zhang, Ruoceng Zhang, Pei Fu, Shaokang Wang, Jiahui Yang, Xin Du, Bin Qin, Ying Huang, Zhenbo Luo, and Jian Luan. 2026. Btl-ui: Blink-think-link reasoning model for gui agent. Advances in Neural Information Processing Systems 38 (2026), 56035–56056. [89] Yichun Zhang, Xiangwu Guo, Yauhong Goh, Jessica Hu, Zhiheng Chen, Xin Wang, Difei Gao, and Mike Zheng Shou. 2026. ShowUI-Aloha: Human-Taught GUI Agent. CoRR abs/2601.07181 (2026). [90] Zhou Zhao, Shengyu Zhang, Liang Wang, Xiangxin Zhou, Zhaokai Wang, Kun Kuang, Fei Wu, Wangchunshu Zhou, Shuofei Qiao, Jiwei Li, Guoyin Wang, Ziyu Zhao, Hongxia Yang, Fan Wu, Jiasheng Ye, Shenzhi Wang, Ruixuan Xiao, Tieyong Zeng, Yuhuai Li, Yuchen Eleanor Jiang, Meiling Tao, Xueyu Hu, Tao Xiong, Biao Yi, Keting Yin, Yurun Chen, Zishu Wei, Xinchen Xu, and Shengze Xu. 2025. OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use. arXiv:2508.04482 [cs.AI] https://arxiv.org/abs/2508.04482 [91] Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. 2024. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, Vol. 2024. 15585–15606. [92] Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Jizhou Guo, Yankai Chen, Chunyu Miao, Hoang H Nguyen, Yue Zhou, Weizhi Zhang, Liancheng Fang, et al. 2026. Llm-based human-agent collaboration and interaction systems: A survey. Findings of the Association for Computational Linguistics: ACL 2026 (2026), 36335–36364.
Venkatesan et al.
Appendix Table of Contents Appendix A: Early Prototype of Plover . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 Appendix B: Formative Study User Experience Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 Appendix C: System-Driven Intelligent Replanning Details . . . . . . . . . . . . . . . . . . . . . . . 18 Appendix D: Plover Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 Appendix E: Evaluation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19
Conference’17, July 2017, Washington, DC, USA
A
Venkatesan et al.
Early Prototype of Plover
Figure 6: Early Plover prototype used in Formative Study. The interface supported prompt authoring, plan inspection/editing, and execution monitoring. Observations from this prototype informed the redesign of Plover (DG1–DG5).
Figure 7: Early Annotation Panel. Users could provide visual annotations during execution; feedback from this component design informed improved intervention and replanning support in Plover.
Plover: Steering GUI Agents through Plan-Centric Interaction
B
Conference’17, July 2017, Washington, DC, USA
Formative Study User Experience Analysis
Table 3: Summary of observations from 6 participants (P1–P6) from the formative user study. Feedback is grouped by thematic categories and includes representative participant observations and corresponding design implications. Category
Participant P1
Understanding & Mental Models
Felt reviewing each step was mentally demanding and suggested previewing the expected outcome instead of inspecting every step individually.
P1
Felt execution was slow and said they would prefer to take over manually because it would be faster. Wanted to intervene during execution when simple actions generated many intermediate steps or screenshots.
Provide controls that allow users to pause execution and intervene when needed. Allow targeted intervention during execution to prevent unnecessary automated steps.
Trusted the system most during execution but least when it generated plans. Trusted the plan panel but found the execution view overwhelming due to excessive low-level steps.
Improve trust by making planning decisions visible and easier to inspect before execution. Present execution feedback at an appropriate level of abstraction rather than exposing raw system traces.
Felt the execution panel contained too many detailed steps and summaries that were difficult to read. Execution traces resembled debugging logs and used terminology such as “tool” that felt too technical for users. Wanted clearer links between the prompt, execution logs, and system state to better understand what the agent was doing.
Replace verbose logs with concise summaries of system actions. Translate low-level tool operations into human-readable descriptions of actions. Visually connect execution progress to the task description and corresponding plan steps.
P1
Drawing annotations were helpful for spatial references but were limited when relevant UI elements were outside the screenshot.
P3
Said text instructions would be used most of the time while annotations would only be necessary for spatial references. Suggested referencing annotated regions via natural language. Felt drawing was useful for locating regions but still required accompanying text to explain the intended action.
Improve annotation support by grounding user feedback directly in the visible interface context. Support combined text and annotation input so users can describe intent while referencing visual regions. Integrate textual instructions with spatial annotations to support clearer correction workflows.
P5
P4 P1 P2
Visibility of System State & Feedback
P2 P3 P6
Intervention & Repair
P5
Interaction Flow & Transitions
Found the step panel too narrow and disliked clicking the edit icon repeatedly to read full steps. Also misunderstood the proposed-plan panel as additions to the current plan. Felt diff labels were sometimes inaccurate and step numbering differences made comparison between plans difficult.
P6
P2
P4
Trust & Comfort
Design Implication Expose full step descriptions and clearly distinguish regenerated plans from existing plans to support accurate mental models. Provide clearer summaries of plan revisions to help users understand how regenerated plans differ from previous ones. Support localized plan editing so users can modify individual steps without regenerating entire plans. Improve plan readability by emphasizing key action words and organizing steps for easier scanning. Provide concise summaries of plan changes and visually distinguish planning and execution views. Surface high-level summaries of regenerated plans to reduce cognitive load during inspection.
P3
Control & Agency
Observation
P2
P3 P6
Wanted the ability to add or modify a single step after replanning instead of accepting the entire regenerated plan. Also noted that the boundary between “Added” and “Changed” diff labels was unclear. Found the planner intuitive but opening each step individually was mentally taxing. Suggested highlighting keywords and grouping related steps. Preferred a concise summary rather than inspecting every step change and felt the planner and execution panels looked too similar.
Execution panel was overwhelming due to many detailed steps and suggested highlighting action words to make steps easier to interpret. Felt diff labels were not very useful and preferred a more meaningful way to understand what changed in the plan. Felt the interface exposed too much low-level detail and suggested showing higher-level explanations with visual cues on screenshots.
Improve visual hierarchy to help users quickly identify key actions and system progress. Provide structured explanations of plan updates rather than relying solely on diff labels. Present execution feedback through higherlevel summaries supported by visual cues in the interface.
Conference’17, July 2017, Washington, DC, USA
C
System-Driven IR Details
Algorithm 1 System-Driven IR Runtime Policy Require: Interaction history 𝑀, session 𝑠, current step index 𝑖 Ensure: Next execution state or replanning event 1: 𝑇 ← LatestToolActions(𝑀) 2: if 𝑇 = ∅ then return normal execution 3: end if 4: 𝑎 ← Canonicalize(𝑇 [−1]) 5: Append 𝑎 to rolling action sequence 𝑆𝑠 6: 𝑓 ← DetectRepetition(𝑆𝑠 ) 7: if 𝑓 = ∅ then return normal execution 8: end if 9: 𝑉 ← LastScreenshots(𝑀, 𝑘 = 3) 10: if |𝑉 | < 3 then return normal execution 11: end if 12: 𝐻 ← {dHash(𝑣) | 𝑣 ∈ 𝑉 } 13: 𝐷 ← {Hamming(𝐻 𝑗 , 𝐻 𝑗+1 )} 14: if ∃𝑑 ∈ 𝐷 such that 𝑑 > 40 then 15: return normal execution 16: end if 17: 𝜏 ← Tail(𝑆𝑠 ) 18: Clear sequence 𝑆𝑠 19: Inject <failure_detected> message with type 𝑓 into 𝑀 20: Append PROPOSAL_MODE_SUFFIX to system prompt 21: 𝑅 ← InvokeModel(𝑀, proposal_only=True) 22: 𝜎 ← ExtractSummary(𝑅) 23: 𝜌 ← ExtractRationale(𝑅) 24: return INTELLIGENT_REPLAN (𝑖, 𝑓 , 𝜏, 𝜎, 𝜌, 𝑅.tool_uses)
Prompt Artifacts. During System-Driven IR, the system augments the model context with two prompt artifacts. First, it injects a structured failure message that signals repeated non-progress and requests a change in tactic. Second, it appends a proposal-mode prompt suffix that constrains the model to generate a recovery proposal in a strict format consisting of a short imperative next step and a brief rationale. Fig.8 shows the structured failure message injected into the conversation context when the runtime watchdog detects repeated non-progress. Fig.9 shows the proposal-mode prompt suffix used to constrain the model during Intelligent Replanning.
<failure_detected type=’REPEAT_SEQ_L3_R3’> Stuck/repetition detected. Stop. Change tactic that can solve this problem. </failure_detected>
Figure 8: Structured failure message injected into the conversation context when the watchdog confirms a stuck state. The type field encodes the detected repeated-action pattern and conditions the model to propose an alternative tactic rather than continuing the same behavior.
Venkatesan et al.
<PROPOSAL_MODE> You are proposing actions only. STRICT FORMAT: - Output EXACTLY ONE text block. - That text block must contain ONLY TWO LINES in this exact order: 1) SUMMARY: <one short imperative step sentence> 2) RATIONALE: <1–2 sentences explaining the detected failure and why the proposed next action helps> - The SUMMARY must: • start with a strong action verb • be written as a standalone executable step • not contain “I will”, “Let’s”, or future tense • not mention internal tool names - Do NOT restate the user’s original request. - Do NOT comment on the failure message directly. - Then output tool_use blocks if needed. - No other prose, analysis, or explanation. Do NOT assume actions will execute. </PROPOSAL_MODE>
Figure 9: Proposal-mode prompt suffix used during System-Driven IR. When execution drift is detected, this suffix is appended to the system prompt to constrain the model to produce a structured recovery proposal consisting of a short next-step summary and rationale.
D
Implementation Details
This section provides additional technical details regarding the prompt engineering, technical implementation and cross-platform execution architecture of Plover.
D.1
Planner Prompt Design
We design the planner prompt to enforce deterministic task decomposition while separating reasoning from actuation. The prompt uses a sectioned Markdown guide with bolded constraints to ensure the model functions as a high-level decomposer for a separate executor. To ensure controllability, the planner follows a tag-based output schema (<analysis>, <steps>), separating intent interpretation from executable UI actions. To support incremental updates, the prompt enforces a plan representation with two ordered blocks: (1) <completed>: An immutable execution history of previously finished steps. (2) <pending>: An editable suffix for remaining or revised actions. This invariant ensures that interventions lead to localized updates rather than monolithic regeneration, preserving execution continuity. Planning is further guided by four principles: (1) Deterministic step decomposition into mechanical UI actions; (2) Contextual grouping of actions to reduce fragmentation; (3) Reasoning-before-action to improve reliability; and (4) History preservation by updating only the pending suffix. To ensure bounded execution, the prompt enforces strict stopping conditions: the planner must terminate and request user guidance if it encounters subjective ambiguity or sensitive data (e.g., credentials). The final step must always return a visible outcome or screen state for verification. Finally, few-shot examples demonstrate valid formatting and termination, guiding the model toward consistent, schema-adherent plan generation.
Plover: Steering GUI Agents through Plan-Centric Interaction
D.2
Executor System Prompt Design
The agent operates under a structured System Capability and UI Summary Protocol designed to ensure deterministic GUI control and human-readable provenance. The prompt design emphasizes four technical pillars: • Environment Contextualization: The planner is grounded in the source environment using DISPLAY exports for GUI subshells. It is instructed to use curl for network requests and pdftotext for document parsing to bypass the visual layout limitations of complex PDFs. • Efficiency via Chaining: To mitigate the latency of highfidelity computer function calls, the prompt encourages action chaining, requiring the agent to batch multiple operations into a single tool request where feasible. • State Verification: The design enforces strict verification loops. The agent must use progressive zooming and multiple PageDown/PageUp sequences to ensure full visibility, followed by immediate screenshots to confirm the successful launch of GUI applications. • Standardized Action Tagging: To support execution legibility, the prompt enforces a UI Summary Protocol. Before each tool invocation, the agent must generate a <ui_summary> tag using a present-progressive verb (e.g., Typing, Clicking, Capturing). This creates a human-readable activity feed that avoids hallucinated labels by defaulting to generic descriptors if UI text is not clearly legible. This prompt structure ensures that the agent’s low-level mechanical actions remain synchronized with the high-level plan while providing the necessary logs for the system’s mechanisms.
D.3
planner from environmental primitives through this abstraction layer, switching between Ubuntu and Windows requires zero modification to planning logic, history preservation, or mixed-initiative policies. This architecture demonstrates that Plover ’s core reasoning and recovery mechanisms are environmentally portable and extensible to any GUI-driven platform.
E Evaluation Details E.1 Failure Analysis Table 4: Failure type taxonomy used to categorize breakdowns during task execution. Failure Type
Description
Perception Error
The model fails to detect or correctly interpret a UI element.
Action Grounding Failure
The model identifies the correct action but performs it incorrectly (e.g., clicking the wrong location or failing to drag correctly).
Planning Error
The system generates an incorrect step in the plan or missequences task actions.
State Misinterpretation
The system misinterprets the interface state (e.g., assuming an action succeeded when it did not).
Execution Drift
The system continues performing actions without making progress toward task completion, typically triggering Intelligent Replanning.
Missing Context
The system lacks required knowledge or context (e.g., domain-specific information such as spreadsheet formulas).
Plover Technical Implementation
Plover’s implementation consists of a distributed architecture coordinated via a centralized orchestration layer. The frontend is a React web application styled with Tailwind CSS, utilizing an HTML5 Canvas overlay to capture high-fidelity spatial annotations for multimodal intervention. The core Planner and Executor services are built with Python 3.10 using the FastAPI framework, ensuring asynchronous handling of long-horizon planning tasks.To enable robust, high-reasoning capabilities, the system leverages the Anthropic Claude 4.5 Sonnet model via the computer-use-2025-01-24 beta. Execution occurs across heterogeneous environments, including a Dockerized Ubuntu container and a native Windows machine, both of which stream visual feedback to the UI via a VNC connection at a fixed resolution of 1024 × 768.
D.4
Conference’17, July 2017, Washington, DC, USA
Cross-Platform Executor Architecture Details
We implement the executor as a modular service under a unified RPC interface, ensuring that Plover remains agnostic to the underlying operating system. The Ubuntu executor synthesizes input events via xdotool and native capture utilities, while the Windows executor implements the same contract using pyautogui for automation and screen capture. Despite these platform-specific diversities, both implementations expose identical gRPC services and return standardized success/failure signals. By isolating the
Conference’17, July 2017, Washington, DC, USA
Venkatesan et al.
Table 5: Results on 38 OSWorld-Verified failure cases originally reported as failed for Claude Sonnet 4.5 with computer-use capabilities. We compare autonomous execution and mixed-initiative execution. For mixed-initiative runs, we report the number of user interventions and the primary recovery mechanism. S = success, P = partial success, F = failure. Task
App
Auto
HAI
# Int.
Failure Cause
Recovery Type
Spreadsheet merge via CLI XLSX to HTML conversion ODS to CSV conversion Author extraction to Excel Spreadsheet to Word Paper Citation Search
Multi-app Multi-app Multi-app Multi-app Multi-app Multi-app
F P P P P F
S S S S S P
1 1 1 2 2 4
State Misinterpretation Perception Error Action Grounding Failure Planning Error, Perception Error State Misinterpretation Planning Error, Perception Error
S F F P F
2 4 3 4 2
F
P
2
LibreOffice Calc LibreOffice Calc LibreOffice Calc LibreOffice Calc LibreOffice Calc LibreOffice Calc
P S F F F S
S – P P S –
1 – 2 2 1 –
Planning Error, Perception Error Planning Error, Execution Drift Planning Error, Execution Drift Execution Drift, State Misinterpretation Action Grounding Failure, Execution Drift Action Grounding Failure, State Misinterpretation Perception Error – Missing Context Missing Context Planning Error –
Annotation NL Guidance NL Guidance Plan Edit NL Guidance Plan Edit, NL Guidance, Annotation Plan Edit, NL Guidance System-driven IR, Annotation System-driven IR, Annotation System-driven IR, NL Guidance Annotation, System-driven IR
Contact Info Collection Finding Research Papers Restaurant Info PDF Form Filling Revenue + pivot table
Multi-app Multi-app Multi-app Multi-app LibreOffice Calc
P F F F F
Branch lookup table
LibreOffice Calc
Column chart creation Name splitting Calculation + Pivot Table Sparkline Charts Filling missing data Creating new sheet + calculation Filtering + calculation Hiding rows with missing data Text alignment Text color coding Duplicates removal Line Spacing Text to Table Font Formatting 1 Font Formatting 2 Citation reference Password settings navigation Shirt filtering >=50% discount Coffee maker filtering Rental Car Booking Flight Booking Appointment Booking Locating FAQ page Locating discussion thread Hotel booking with filters Apparel shopping with filters
Annotation – NL Guidance NL Guidance Plan Edit –
LibreOffice Calc LibreOffice Calc LibreOffice Writer LibreOffice Writer LibreOffice Writer LibreOffice Writer LibreOffice Writer LibreOffice Writer LibreOffice Writer LibreOffice Writer Browser Browser Browser Browser Browser Browser Browser Browser Browser Browser
F S P F P S S F F S S S P S S P S F F S
S – S P S – – S S – – – S – – S – S S –
2 – 2 2 2 – – 2 2 – 0 0 1 0 0 2 0 2 2 0
Execution Drift – Execution Drift Planning Error, Execution Drift Planning Error, Execution Drift – – Planning Error, Action Grounding Execution Drift, State Misinterpretation – – – Action Grounding Failure – – Execution Drift – Execution Drift, Perception Error Execution Drift, Perception Error –
System-driven IR – System-driven IR Plan Edit, System-driven IR NL Guidance, System-driven IR – – Plan Edit, Annotation System-driven IR, NL Guidance – – – Annotation – – System-driven IR – System-driven IR, NL Guidance System-driven IR, Annotation –
Annotation, NL Guidance
Plover: Steering GUI Agents through Plan-Centric Interaction
Conference’17, July 2017, Washington, DC, USA
Table 6: Results on 26 OSWorld-Verified autonomous non-success cases. Rows are grouped by the autonomous outcome first (P then F). For each group, we report the mixed-initiative outcome, number of interventions, failure cause, and recovery type. Summary rows report the number of tasks, total interventions, and average interventions. S = success, P = partial success, F = failure.
Task
App
Auto
HAI
# Int.
XLSX to HTML conversion ODS to CSV conversion Author extraction to Excel Spreadsheet to Word Contact Info Collection Column chart creation Text alignment Duplicates removal Coffee maker filtering Appointment Booking
Multi-app Multi-app Multi-app Multi-app Multi-app LibreOffice Calc LibreOffice Writer LibreOffice Writer
P P P P P P P P
S S S S S S S S
Browser Browser
P P
Total (Auto=P) Average
Failure Cause
Recovery Type
1 1 2 2 2 1 2 2
Perception Error Action Grounding Failure Planning Error, Perception Error State Misinterpretation Planning Error, Perception Error Perception Error Execution Drift Planning Error, Execution Drift
S S
1 2
Action Grounding Failure Execution Drift
NL Guidance NL Guidance Plan Edit NL Guidance Plan Edit, NL Guidance Annotation System-driven IR NL Guidance, System-driven IR Annotation System-driven IR
10 P
10 S / 0 P / 0 F
16 1.60
Spreadsheet merge via CLI Paper Citation Search
Multi-app Multi-app
F F
S P
1 4
Finding Research Papers
Multi-app
F
F
4
Restaurant Info
Multi-app
F
F
3
PDF Form Filling
Multi-app
F
P
4
Revenue + pivot table
LibreOffice Calc
F
F
2
Branch lookup table
LibreOffice Calc
F
P
2
Calculation + Pivot Table Sparkline Charts Filling missing data Filtering + calculation Text color coding Font Formatting 1 Font Formatting 2
LibreOffice Calc LibreOffice Calc LibreOffice Calc LibreOffice Calc LibreOffice Writer LibreOffice Writer LibreOffice Writer
F F F F F F F
P P S S P S S
2 2 1 2 2 2 2
Locating discussion thread
Browser
F
S
2
Hotel booking with filters
Browser
F
S
2
Total (Auto=F) Average
16 F
7S/6P/3F
37 2.31
Overall Total Overall Average
26
17 S / 6 P / 3 F
53 2.04
State Misinterpretation Planning Error, Perception Error
Annotation Plan Edit, NL Guidance, Annotation Planning Error, Execution Drift System-driven IR, Annotation Planning Error, Execution Drift System-driven IR, Annotation Execution Drift, State Misinterpreta- System-driven IR, NL Guidtion ance Action Grounding Failure, Execution Annotation, System-driven Drift IR Action Grounding Failure, State Mis- Annotation, NL Guidance interpretation Missing Context NL Guidance Missing Context NL Guidance Planning Error Plan Edit Execution Drift System-driven IR Planning Error, Execution Drift Plan Edit, System-driven IR Planning Error, Action Grounding Plan Edit, Annotation Execution Drift, State Misinterpreta- System-driven IR, NL Guidtion ance Execution Drift, Perception Error System-driven IR, NL Guidance Execution Drift, Perception Error System-driven IR, Annotation
Conference’17, July 2017, Washington, DC, USA
Venkatesan et al.
Table 7: Failure types, recovery mechanisms, and outcomes across 26 non-success trials. Category
Count
Success
Partial
Failure
Failure Types Execution Drift Planning Error Perception Error Action Grounding State Misinterpretation Missing Context
12 9 7 5 5 2
8 6 5 3 3 1
3 2 2 1 2 1
1 1 0 1 0 0
Recovery Mechanisms System-driven IR NL Guidance Annotation Plan Edit
12 12 10 6
7 8 7 4
3 4 2 1
2 0 1 1
Plover: Steering GUI Agents through Plan-Centric Interaction
E.2
Conference’17, July 2017, Washington, DC, USA
Auto-evaluation Details Table 8: Image similarity–based auto-evaluation results for Plover across five scenarios. Scenario
Firefox Fillable Form
Firefox Machine Dashboard
Firefox Vis Dashboard
LibreOffice Incident Sheet
LibreOffice Sensor Logs Sheet
Overall Average
Thresholds
Trial
SSIM ↑
MSE ↓
dHash ↓
Match
High: SSIM ≥ 0.98, MSE ≤ 100, dHash ≤ 2 Partial: SSIM ≥ 0.97, MSE ≤ 250, dHash ≤ 5
8 10 11 41 60 Avg
0.9720 0.9899 0.9842 0.9770 0.9856 0.9817
205.37 56.260 87.060 186.16 77.320 122.43
5 0 0 4 1 2.0
Partial High High Partial High –
High: SSIM ≥ 0.978, MSE ≤ 140, dHash ≤ 2 Partial: SSIM ≥ 0.95, MSE ≤ 250, dHash ≤ 5
5 15 16 43 60 Avg
0.9776 0.9825 0.9796 0.9763 0.9887 0.9809
138.73 111.55 104.82 113.06 59.74 105.58
2 1 1 1 1 1.2
Partial High High Partial High –
High: SSIM ≥ 0.97, MSE ≤ 200, dHash ≤ 3 Partial: SSIM ≥ 0.90, MSE ≤ 700, dHash ≤ 10
3 13 37 64 74 Avg
0.9168 0.9831 0.9319 0.9449 0.9164 0.9386
566.26 106.38 582.76 404.99 580.11 448.10
10 1 8 10 10 7.8
Partial High Partial Partial Partial –
High: SSIM ≥ 0.95, MSE ≤ 2000, dHash ≤ 5 Partial: SSIM ≥ 0.80, MSE ≤ 8000, dHash ≤ 12
36 52 66 83 99 Avg
0.6448 0.5134 0.7143 0.5720 0.6025 0.6094
2404.98 2788.43 1461.53 2580.62 1895.99 2226.31
2 19 18 24 8 14.2
Low Low Low Low Low –
High: SSIM ≥ 0.95, MSE ≤ 2000, dHash ≤ 5 Partial: SSIM ≥ 0.80, MSE ≤ 8000, dHash ≤ 12
6 18 32 57 84 Avg
0.6164 0.6383 0.6117 0.9318 0.6292 0.6855
2347.07 1766.4 2327.08 428.05 2108.47 1795.41
14 9 17 2 8 10.0
Low Low Low Partial Low –
0.8392
939.57
7.04
–
Conference’17, July 2017, Washington, DC, USA
Venkatesan et al.
Table 9: Plan comparison metrics across all trials. Coverage measures phase overlap, Order measures phase sequence alignment, Redundancy measures consecutive duplicate phases, and Actionability measures the actionable step ratio. #Ref refers to the number of reference phases in the random trajectories, while #Pred refers to the number of reference phases in the predicted trajectories. Task
Trial
Coverage
Order
Redundancy
Actionability
#Ref
#Pred
Firefox Fillable Form
7 10 11 41 60 Avg
1.00 0.50 0.60 1.00 0.50 0.72
0.71 0.33 0.38 0.43 0.25 0.42
0.22 0.33 0.38 0.54 0.50 0.39
1.00 1.00 1.00 1.00 1.00 1.00
7 9 8 7 4 7.0
7 6 5 6 3 5.4
Firefox Machine Dashboard
5 15 16 43 54 Avg
0.25 0.33 0.75 0.25 0.50 0.42
0.18 0.50 0.50 0.43 0.50 0.42
0.45 0.22 0.36 0.38 0.21 0.33
1.00 1.00 1.00 0.69 1.00 0.94
11 4 10 7 8 8.0
6 14 9 8 11 9.6
Firefox Vis Dashboard
3 13 37 64 74 Avg
0.75 0.67 0.50 0.67 0.67 0.65
0.38 0.38 0.42 0.64 0.40 0.44
0.25 0.13 0.25 0.23 0.19 0.21
1.00 1.00 0.83 1.00 0.88 0.94
8 8 12 11 10 9.8
6 7 9 10 21 10.6
LibreOffice Incident Sheet
36 52 66 83 99 Avg
0.67 0.40 0.50 1.00 0.75 0.66
0.25 0.20 0.25 0.43 0.60 0.35
0.38 0.00 0.38 0.38 0.200 0.27
1.00 1.00 1.00 1.00 1.00 1.00
8 10 4 7 5 6.8
5 3 5 5 4 4.4
LibreOffice Sensor Logs
6 18 32 57 84 Avg
1.00 0.67 0.50 0.50 0.67 0.67
0.50 0.57 0.33 0.20 0.50 0.42
0.43 0.57 0.50 0.40 0.33 0.44
1.00 1.00 1.00 1.00 0.83 0.97
4 7 3 5 4 4.6
4 6 3 3 8 4.8
Overall Avg
–
0.62
0.41
0.33
0.97
7.24
6.96