L AB OSB ENCH : B ENCHMARKING C OMPUTER U SE AGENTS TOWARDS S CIENTIFIC I NSTRUMENT C ONTROL
arXiv:2606.16802v1 [cs.AI] 15 Jun 2026
A P REPRINT Anqi Zou1,2 , Han Deng1,3 , Chengyu Zhang1 , Junquan Hu1,2 , Yu Wang1,2 , Yuxiang Xing1,2 , Aokai Zhang1,2 , Hanling Zhang1,3 , Zhaoyang Liu†,4 , Ben Fei†,3 , Zhihui Wang†,2 , Wanli Ouyang1,3 1 Shenzhen Loop Area Institute 2 Dalian University of Technology 3 The Chinese University of Hong Kong 4 The Hong Kong University of Science and Technology [email protected], [email protected], [email protected]
A BSTRACT Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scenarios require coordinated control over complex interfaces, and feedback-driven parameter adjustment. However, directly evaluating agents on physical highprecision instruments is impractical due to high cost, safety risks, limited accessibility, and difficulty in ensuring reproducible evaluation. This motivates the need for a simulated yet realistic testbed that preserves the operational challenges of scientific instruments while enabling scalable and safe benchmarking. To this end, we introduce LabOSBench, a challenging benchmark for multimodal GUI agents built on a suite of web-based scientific-instrument simulators. Operating directly via a browser, LabOSBench avoids resource-heavy OS virtualization while supporting flexible task configuration and execution-based evaluation. Specifically, LabOSBench constructs 96 subtasks across eight instrument simulators, covering workflows from sample loading, alignment, parameter tuning, and data acquisition to result inspection. We evaluate general-purpose vision-language models, specialized GUI agent models, and advanced agentic frameworks at both subtask and end-to-end levels. Our experiments reveal that while existing agents can complete many structured GUI subtasks, they still struggle with feedback-driven operations and long-horizon workflow execution. Overall, LabOSBench provides a reproducible, low-cost testbed for advancing computer-using agents toward scientific-instrument control.
1
Introduction
Humans perform many computational tasks through graphical user interfaces (GUIs) and command-line interfaces (CLIs), including web browsing, file management, data analysis, and software development. Recent advances in large vision-language models (VLMs) have enabled autonomous digital agents that follow natural-language instructions and interact with GUIs through reasoning-and-acting loops [Yao et al., 2022, Qin et al., 2025], offering a promising path toward simplifying human–computer interaction. Despite this progress, developing multimodal agents for domain-specific professional settings remains challenging. Existing benchmarks such as OSWorld provide full operating-system environments [Xie et al., 2024, Bonatti et al., 2024, Rawles et al., 2025], but rely on heavy virtualization frameworks such as VMware or Docker, making them costly to deploy and difficult to scale. Web-based benchmarks such as WebArena and Mind2Web focus on general web navigation [Zhou et al., 2024, Deng et al., 2023], and thus do not capture scientific-instrument interfaces with intricate layouts, long-horizon procedures, and continuous parameter tuning. Moreover, many prior benchmarks rely on static demonstrations or offline evaluation, limiting their ability to assess interactive learning, exploration, and alternative valid action sequences.
LabOSWorld
A P REPRINT
Scientific instrument interfaces pose unique challenges for multimodal agents. Benchmarking this setting is difficult because it requires realistic procedural workflows, dense professional interfaces, and feedback-driven parameter tuning, while avoiding the cost, safety risks, and limited accessibility of physical instruments. Although recent work has explored scientific agents, autonomous laboratories, and intelligent instrumentation systems [Tom et al., 2024, Szymanski et al., 2023, Boiko et al., 2023, Wang et al., 2022, Jansen et al., 2024, Deng et al., 2026], existing benchmarks provide limited support for evaluating GUI-based instrument control. Prior work mainly focuses on scientific reasoning, experiment planning, autonomous discovery, or instrument-specific intelligence [Tom et al., 2024, Szymanski et al., 2023, Boiko et al., 2023, Wang et al., 2022, Jansen et al., 2024], whereas our work evaluates existing computer-use agents on scientific-instrument GUIs when direct device APIs are unavailable. To address this gap, we introduce LabOSBench, a lightweight, web-based, executable benchmark for evaluating multimodal agents on scientific instrument simulators. LabOSBench runs entirely in a standard browser environment, eliminating the need for OS-level virtualization while preserving realistic GUI interactions. It supports native mouse and keyboard control, configurable initial states, execution-based evaluation. LabOSBench includes 96 subtasks across eight instrument simulators, covering workflows such as X-ray diffraction scanning, focused ion beam milling, light/fluorescence microscopy imaging, and scanning probe microscopy operation. We evaluate state-of-the-art LLM/VLM-based agents, specialized GUI models [Qin et al., 2025], and agentic frameworks [Han et al., 2026]. Results show that strong multimodal models can complete many structured subtasks, but still struggle with feedback-driven adjustment, scientific-state interpretation, and long-horizon execution. Our analysis further reveals failures in visual grounding, action localization, recovery strategy, and instrument-specific GUI understanding. Our contributions are threefold: (i) we introduce LabOSBench, a lightweight executable benchmark for scientificinstrument GUI control; (ii) we design 96 subtasks across eight instrument simulators, covering procedural operation, visual feedback, and scientific-state adjustment; and (iii) we systematically evaluate general-purpose multimodal models, specialized GUI models, and agentic frameworks, revealing substantial gaps in feedback-driven and end-to-end scientific workflows.
2
Related Work
Computer-Using Agents (CUAs). CUAs have evolved from DOM-based web agents to vision-based computer-use agents that operate directly on screenshots [Deng et al., 2023, Zhou et al., 2024, He et al., 2024, Zheng et al., 2024, Xie et al., 2024, Davydova et al., 2025, Rawles et al., 2025, Qin et al., 2025, Xu et al., 2025, Liu et al., 2025, Wu et al., 2025, Li et al., 2026]. Screen parsing methods [Lu et al., 2024a, Li et al., 2025] improve agents’ ability to identify actionable elements in complex visual observations. However, current agents still struggle with long-horizon planning, precise localization, and state-dependent interaction dynamics. Agents for Science. Scientific-agent systems have studied reasoning, experiment planning, chemistry tool use, machine-learning experimentation, and autonomous discovery [Tom et al., 2024, Szymanski et al., 2023, Boiko et al., 2023, Wang et al., 2022, Jansen et al., 2024, Bran et al., 2023, Huang et al., 2023, Lu et al., 2024b]. MyScope1 provides browser-based microscope simulators for human education. In contrast, LabOSBench adds executable evaluation for computer-use agents, including natural-language task specifications, subtask-level success checkers, episode logging, browser execution, and instrument-specific metrics. Computer-Using Benchmarks. Existing benchmarks evaluate agents in web, desktop, mobile, and scientific workflow environments [Deng et al., 2023, Zhou et al., 2024, Koh et al., 2024, Xie et al., 2024, Xu et al., 2024, Davydova et al., 2025, Rawles et al., 2025, Yang et al., 2026, Sun et al., 2025]. Web-based benchmarks are lightweight but focus mainly on general navigation, while OS-level benchmarks improve realism at the cost of heavy virtualization. More recently, ScienceBoard [Sun et al., 2025] evaluates multimodal autonomous agents in realistic scientific workflows with professional software. However, existing benchmarks still provide limited coverage of scientific-instrument GUI operation, where agents must manipulate dense control panels, track instrument states, perform continuous parameter adjustment, and respond to state-dependent visual or numerical feedback. In contrast, LabOSBench focuses on executable, browser-based scientific instrument simulators to evaluate GUI-based instrument control when device APIs are unavailable. 1
https://myscope.training/
2
LabOSWorld
A P REPRINT
Figure 1: Overview of LabOSBench. The benchmark evaluates computer use agents on eight types of scientific instrument web interfaces, where agents receive natural-language instructions and screenshots, execute GUI actions, and are scored using episode-level and subtask-level metrics across 96 subtasks.
3
The LabOSWorld Benchmark
We introduce LabOSBench, a benchmark suite for evaluating computer-use agents on scientific-instrument interfaces. Unlike general web or desktop tasks, scientific-instrument control requires procedural dependencies, domain-specific controls, state transitions, and ordered acquisition or export operations. LabOSBench complements general computeruse benchmarks [Xie et al., 2024, Zhou et al., 2024, Deng et al., 2023, Rawles et al., 2025] with domain-specific web interfaces and laboratory-style workflows. Each run is launched from a lightweight specification consisting of command-line arguments, environment variables, and a natural-language instruction. As shown in Fig. 2, a browser-backed coordinator executes agent actions, exports episode logs through in-page benchmark hooks, and aggregates instrument-specific metrics. 3.1
Scientific Instrument Tasks
LabOSBench covers eight simulated scientific instruments that span three foundation experimental paradigms: (i) advanced microscopy imaging, including Scanning Electron Microscopy (SEM), Transmission Electron Microscopy (TEM), and Light/Fluorescence Microscopy (LFM); (ii) Diffraction and Spectroscopic Analysis, including X-Ray Diffraction (XRD), Energy-Dispersive Spectroscopy (EDS) and Atom Probe Tomography (APT); (iii) precision micronanofacrication and manipulation, including Focused Ion Beam (FIB) and Scanning Probe Microscopy (SPM). Each instrument is implemented as a high-fidelity, interactive web environment that meticulously preserves domain-specific 3
LabOSWorld
A P REPRINT
Multimodal GUI Agent Coordinator
Run specification CLI flags instrument/url/max_steps /subtask/agent/model
Environment
Observations (screenshot)
Playwright Driver
Action
(click drag type)
Session 2 (SEM)
page.evaluate DOM
Instruction
Task manager
NL prompt
schema.json JSON Schema for exported episode_*.json
Session 1 (XRD)
Browser-backed simulator
API_KEY API_URL
Episode log contract
Execution platform
Setup interpreter
Postprocess
Getter
Metrics
(navigate optional fastforward)
(navigate optional fastforward)
(export episode JSON)
(success step_effi subtasks)
Static web simulator + benchmark.js
Episode : JSON logs + screenshots
Figure 2: Evaluation infrastructure of LabOSBench. Each run is specified by command-line arguments, environment variables, and a natural-language instruction. A browser-backed coordinator executes agent actions in scientific instrument web simulators, exports episode logs through in-page benchmark hooks, and aggregates metrics according to instrument-specific JSON Schemas.
control modalities, such as multi-dimensional sliders, discrete toggles, real-time dynamic plots, and multi-state status indicators. Each task within our benchmark is operationalized via a natural-language instruction describing the target scientific operation. Rather than navigating standardized, linear web forms, agents in LabOSWorld must master highly heterogeneous, instrument-specific operational logic. For example, XRD tasks require agents to select a specimen, configure scan parameters, run the scan, and save the diffraction result, while SEM tasks involve chamber control, sample selection, vacuum preparation, high-voltage activation, imaging adjustment, scanning, and image saving. Consequently, successfully solving these tasks demands an integration of fine-grained cross-modal visual-grounding, non-linear procedural reasoning under physical dependencies, and robust state tracking across extended interaction horizons.
3.2
Workflow and Subtask Decomposition
To facilitate structured evaluation, we structure the workflow of each instrument as a sequence of subtasks. Each subtask corresponds to a meaningful operation stage detectable from simulator states or interactions with specific GUI controls. This decomposition follows actual scientific operating procedures rather than arbitrary or heuristic screen-region segmentations. Across instruments, subtasks fall into recurring structural categories: (i) Sample preparation (e.g., specimen selection and mounting); (ii) Environment conditioning (e.g., vacuum evacuation); (iii) Power or beam activation (e.g., powerup or beam alignment); (iv) Parameter configuration (e.g., focus adjustment and astigmatism correction); (v) Data acquisition (e.g., sample scanning and signal capturing); This design enables interpretable evaluation beyond binary task success and supports targeted diagnosis of failures at specific workflow stages. 4
LabOSWorld
A P REPRINT
Inst.
#Sub.
Goal
XRD SEM TEM FIB SPM LFM EDS APT
8 12 10 20 14 12 8 12
Diffractogram Imaging Imaging Preparation Scanning Brightfield Spectrum Reconstruction
Total
96
–
Figure 3: Overview of scientific-instrument workflows and GUI task taxonomy in LabOSBench. The left panel summarizes the distribution of 96 subtasks across high-level and fine-grained GUI operation categories, while the right panel lists the eight instrument simulators, their subtask counts, and workflow goals.
3.3
Browser-Based Evaluation Infrastructure
Unlike traditional environments that depend on heavy virtual machines, LabOSBench runs natively in a browser. We use Playwright to open simulator pages, capture screenshots, and execute agent actions. At each step, the agent receives the instruction and current screenshot, then outputs an executable GUI action. The action space covers clicking, double-clicking, dragging, typing, selecting dropdown options, pressing keys, scrolling, and waiting for simulator state transitions. The browser driver parses and executes these actions, using DOM-level event dispatch when possible and mouse events otherwise for JavaScript-based controls. Each simulator page is instrumented with an in-page benchmark script, such as benchmark_xrd.js or benchmark_sem.js. These scripts record subtask completion, control attempts, episode metadata, and step-level traces. Finally, a Python-based coordinator exports the in-page episode object and step screenshots, following instrumentspecific JSON Schemas with fields such as success, summary, subtasks, and steps.
3.4
Full-Episode and Subtask-Level Evaluation
We evaluate agents in two complementary modes. In the full-episode setting, the agent starts from the initial simulator state and must complete the entire workflow within a pre-defined step budget. This setting measures long-horizon task completion and captures compounding failures caused by early mistakes, missed state transitions, or incorrect parameter settings. In the subtask-level setting, the simulator is initialized directly to the canonical state immediately preceding a target subtask, and the agent is evaluated solely on that individual subtask. This mode isolates local capabilities such as locating a domain-specific widget, selecting a menu item, dragging a parameter slider, or recognizing operation completion. For simulators that support programmatic state initialization, such as XRD, we implement fast-forward functions exposed on the browser window, e.g., XRD_fast_forward_to_subtask(Sk ). These functions mark preceding subtasks as completed and position the UI at the beginning of the target subtask. Such fast-forwarding is utilized exclusively for diagnostic subtask evaluation and is disabled during full-episode evaluation. Together, the two settings distinguish local GUI grounding failures from long-horizon workflow failures. An agent that succeeds on subtasks but fails in full episodes may suffer from limitations in memory, sequential ordering, or error recovery, whereas failure on individual subtasks indicates weak grounding of domain-specific scientific controls. 5
LabOSWorld
3.5
A P REPRINT
Metrics
We report episode-level and subtask-level metrics. An episode is successful if all required subtasks are completed within the step budget. For each subtask, we record whether the corresponding operation is completed, revealing where an agent fails in full episodes and enabling localized success-rate analysis. Let Sd denote the set of subtasks associated with instrument d, and let Rd,s denote the number of repeated runs available for subtask s. In our main subtask-level evaluation, Rd,s = 2. For each subtask s ∈ Sd , we compute: Rd,s 1 X 1 [successd,s,r ] , SRd,s = Rd,s r=1
where successd,s,r indicates whether the agent successfully completes subtask s of instrument d in run r. The instrument-level score is then computed by averaging over all subtasks: Scored =
1 X SRd,s . |Sd | s∈Sd
The reported results in Table 1 correspond to Scored for each instrument. We also log interaction attempts and grounding-oriented statistics for diagnostic analysis. For image-producing tasks, we compute pixel-level quality metrics when a reference image or canonical view is available. In SEM, the saved or current micrograph is compared against a reference view after display-region normalization. We report peak signal-to-noise ratio (PSNR) to measure the visual fidelity of the resulting micrograph. This metric assesses whether parameter adjustments produce a scientifically usable visual state beyond merely triggering the correct GUI action.
4
Experiments
In this section, we present the experimental settings and main results for representative GUI agent baselines on LabOSBench, including general-purpose multimodal models, specialized GUI models, and agentic frameworks. 4.1
Baseline Models and Agentic Frameworks
We evaluate three groups of baselines on LabOSBench. The first group includes general-purpose multimodal models, including Qwen3VL-32B [Bai et al., 2025], EvoCUA-8B [Xue et al., 2026], Claude Sonnet-4.5 [Anthropic, 2025a], Kimi-K2.5 [Kimi Team et al., 2026], Seed-1.6 [ByteDance Seed Team, 2025], GPT-5.5 [OpenAI, 2026], and Claude Opus-4.5 [Anthropic, 2025b]. The second group consists of specialized GUI models, including UI-TARS-1.5-7B [Qin et al., 2025] and GUI-Owl-7B [Deng et al., 2026]. The third group contains agentic frameworks built on strong foundation models, including GTA1 w/ GPT-5.5, VLAA-GUI w/ Opus-4.5 [Han et al., 2026], and Hippo Agent w/ Opus-4.5. For each task, the agent receives a natural-language instruction and the current screen observation, and then outputs an executable GUI action, such as clicking, typing, scrolling, selecting a widget, or triggering an instrument-specific operation. We use a unified interaction protocol across all instruments and models. For subtask-level evaluation, each model is given at most 50 interaction steps per subtask. Unless otherwise specified, we run each model–subtask pair twice (runs=2) and report subtask-normalized success rates. For each instrument d, we first compute the success rate of each subtask over repeated runs and then average across all subtasks, ensuring that each subtask contributes equally. The instrument-level scores reported in Table 1 correspond to this subtask-normalized score. Benchmark-level difficulty. Before comparing individual models, we first analyze the structural difficulty of LabOSBench at the instrument level. Figure 4 visualizes each simulator according to its average non-human model success rate, workflow length, and human–best model gap. This view shows that the difficulty of scientific-instrument GUI control is multi-dimensional. FIB is challenging because it combines the longest workflow with low average model performance, indicating strong long-horizon dependency and error-accumulation pressure. LFM has a shorter workflow than FIB but a larger human–model gap, suggesting that feedback-driven optical adjustment and visual-state interpretation remain difficult for current agents. In contrast, EDS and XRD are located in the short-workflow and high-performance region, reflecting their more structured operation sequences. 6
LabOSWorld
A P REPRINT
Table 1: Main results of average subtask-level scores on the LabOSBench benchmark across eight scientific instrument simulators. Scores are reported in the range [0, 1]. The Avg. column reports the macro-average over the eight instrument simulators. Bold values indicate the best non-human method for each column. Category
Model
SEM SPM TEM XRD LFM FIB
General Models
Qwen3VL-32B EvoCUA-8B Claude Sonnet-4.5 Kimi-K2.5 Seed-1.6 GPT-5.5 Claude Opus-4.5
0.833 0.643 0.500 0.750 0.500 0.275 0.667 0.750 0.615 0.833 0.571 0.500 0.750 0.500 0.350 0.708 0.562 0.597 0.833 0.714 0.400 0.625 0.500 0.250 0.667 1.000 0.624 0.625 0.357 0.500 0.875 0.667 0.525 0.667 0.562 0.597 0.792 0.786 0.650 0.875 0.667 0.650 0.875 0.812 0.763 0.833 0.786 0.500 0.625 0.500 0.650 0.917 1.000 0.726 0.833 0.643 0.400 0.750 0.500 0.250 0.583 1.000 0.620
APT EDS
Avg.
Specialized Models
UI-TARS-1.5-7B GUI-Owl-7B
0.500 0.750 0.600 0.812 0.542 0.650 0.667 0.750 0.659 0.833 0.714 0.500 0.750 0.333 0.325 0.667 1.000 0.640
GTA1 w/ GPT-5.5 0.875 0.821 0.850 0.875 0.708 0.675 0.833 0.875 0.814 Agentic VLAA-GUI w/ Opus-4.5 0.333 0.429 0.400 0.625 0.333 0.150 0.500 0.875 0.456 Frameworks Hippo Agent w/ Opus-4.5 0.833 0.714 0.500 0.750 0.583 0.300 0.583 1.000 0.658 Human
Human
0.917 0.857 0.900 1.000 0.917 0.800 1.000 1.000 0.924
Figure 4: Instrument difficulty map of the eight scientific-instrument simulators in LabOSBench. The x-axis shows the average non-human success rate, the y-axis shows the number of subtasks, and bubble size denotes the human–model gap. Dashed lines indicate instrument-level means.
4.2
Results and Analysis
Category-level analysis. Table 2 reveals distinct strengths and limitations across method families. General-purpose multimodal models benefit from strong visual perception and instruction following: Seed-1.6 performs well on experimental execution and post-processing, while GPT-5.5 is strong in measurement configuration and execution. This suggests that general VLMs can handle instruction understanding, visual feedback, and parameter selection, but still struggle with long-horizon tracking and fine-grained closed-loop adjustment. Specialized GUI models show advantages in conventional interface grounding. UI-TARS-1.5-7B and GUI-Owl-7B are competitive on widget selection, panel operation, and structured GUI procedures. However, their gains are less consistent for tasks requiring scientific-state interpretation, such as alignment, focusing, or visually guided adjustment, indicating that GUI localization alone is insufficient for scientific-instrument control. Agentic frameworks achieve the strongest overall results, but their benefits depend on reliable grounding and domainvalid recovery. GTA1 w/ GPT-5.5 obtains the best overall average and leads on Activation & Alignment and Measurement Configuration, suggesting that planning, verification, and execution monitoring can improve complex workflows. In contrast, the weaker performance of VLAA-GUI w/ Opus-4.5 shows that simply adding an agent loop is not enough: wrong clicks, incorrect control adjustments, and invalid recovery actions can still derail scientific procedures. 7
LabOSWorld
A P REPRINT
Table 2: Success rates across five scientific-instrument task categories. Scores are reported as percentages. Each category-level score is computed by averaging subtask-level success rates over all operation-level subtasks assigned to that category. Model Qwen3VL-32B Hippo Agent w/ Opus-4.5 UI-TARS-1.5-7B GUI-Owl-7B EvoCUA-8B Kimi-K2.5 Seed-1.6 VLAA-GUI w/ Opus-4.5 Claude Sonnet-4.5 Claude Opus-4.5 GTA1 w/ GPT-5.5 GPT-5.5
Preparation & Sample Handling
Activation & Alignment
Measurement Configuration
Experimental Execution
Post-processing & Completion
82.4 88.2 91.2 85.3 85.3 76.5 97.1 82.4 76.5 82.4 94.1 82.4
38.6 43.2 40.9 34.1 40.9 45.5 59.1 18.2 45.5 40.9 63.6 50.0
61.1 66.7 64.8 64.8 64.8 53.7 66.7 29.6 63.0 63.0 87.0 81.5
46.4 42.9 64.3 50.0 50.0 57.1 82.1 28.6 35.7 35.7 78.6 71.4
62.5 65.6 71.9 68.8 43.8 62.5 81.2 56.2 68.8 62.5 75.0 75.0
Figure 5: Qualitative comparison of final outputs on the SEM focus-adjustment task under the same task setting.
Scientific-state understanding. Table 2 suggests that scientific-instrument GUI control cannot be reduced to discrete widget grounding. Tasks involving alignment, focusing, and image adjustment require agents to interpret scientific visual states and perform feedback-driven control. We therefore analyze the SEM focus-adjustment task, where the agent must iteratively adjust focus and judge image clarity. As shown in Fig. 6, GPT-5.5 achieves both the highest focus success rate and the highest best PSNR, showing that LabOSBench can evaluate intermediate scientific-state quality rather than only whether a task is eventually completed. End-to-end workflow difficulty. Although our main evaluation is subtask-level, we further perform an end-to-end pilot evaluation using GPT-5.5 with 10 runs per workflow. GPT-5.5 achieves non-zero success on EDS, APT, and XRD, with success rates of 70.0%, 80.0%, and 60.0%, respectively, but fails on SEM, SPM, TEM, LFM, and FIB. The average end-to-end success rate is only 26.3%, much lower than its subtask-level performance, indicating severe error accumulation in complete workflows. Instrument-level difficulty. Performance varies substantially across instruments. EDS and XRD are relatively easier due to structured operations, clear feedback, and short action sequences. In contrast, SEM, FIB, and LFM are more challenging because they involve feedback-driven focusing, image-quality adjustment, long procedures, fine-grained operations, and visual target localization. These differences show that LabOSBench evaluates not only generic GUI navigation, but also visual grounding, scientific-domain understanding, long-horizon reasoning, and closed-loop control. 8
LabOSWorld
A P REPRINT
Figure 6: Scientific-state quality on the SEM focus-adjustment task. The two lines show best PSNR and focus success rate (SR) across models. GPT-5.5 achieves the highest values on both metrics. Summary. The results highlight four findings. First, subtask-level evaluation reveals limitations of current GUI agents, while end-to-end evaluation exposes severe error accumulation. Second, general-purpose multimodal models outperform specialized GUI models overall, but still struggle with feedback-intensive and long-horizon tasks. Third, agentic frameworks help only when their recovery strategies align with scientific-instrument operation semantics. Fourth, LabOSBench challenges agents on GUI grounding, scientific-state understanding, sequential operation, and continuous visual adjustment.
5
Conclusion
We presented LabOSBench, a lightweight executable benchmark for scientific-instrument GUI control built on webbased simulators. By avoiding full OS virtualization while preserving dense interfaces, procedural dependencies, continuous parameter tuning, and feedback-driven visual adjustment, LabOSBench provides a practical testbed for evaluating multimodal GUI agents in scientific workflows. Our evaluation shows that current agents remain far from reliable scientific-instrument operation: strong multimodal models can solve many discrete subtasks, and agentic frameworks can improve robustness, but failures persist in scientific-state interpretation, precise spatial grounding, closed-loop adjustment, and long-horizon workflows. The SEM focus analysis and end-to-end study further demonstrate that successful agents must not only complete GUI actions, but also maintain valid intermediate scientific states and avoid error accumulation. Overall, LabOSBench highlights scientific-instrument control as a challenging and underexplored setting that connects computer-use agent evaluation [Xie et al., 2024, Zhou et al., 2024, Qin et al., 2025] with autonomous laboratory systems [Tom et al., 2024, Szymanski et al., 2023, Boiko et al., 2023], and we hope it encourages future work on domain-aware GUI agents with accurate grounding, state-aware planning, and scientific feedback interpretation.
Limitations Although LabOSBench provides a lightweight and executable testbed for scientific-instrument GUI control, it still has several limitations. First, it is built on web-based simulators rather than physical instruments, so it cannot fully capture hardware latency, calibration uncertainty, safety constraints, or real laboratory failure modes. Second, the current benchmark covers a selected set of instruments and workflows, including microscopy, spectroscopy, diffraction, and tomography, but does not yet include broader laboratory scenarios such as wet-lab protocols, robotic manipulation, chemical synthesis, or multi-instrument experimental planning. Finally, the evaluation mainly relies on screenshots and logged simulator states, while some scientific decisions may require richer observations, domain knowledge, or multimodal sensor signals beyond the current environment. LabOSBench is designed for evaluating agents in simulated scientific-instrument environments and does not provide direct control over real laboratory hardware. Nevertheless, scientific-instrument automation may involve safety-critical operations if deployed on physical devices. Future systems should include human supervision, permission control, operation logging, and safety interlocks before being connected to real instruments. The benchmark is intended to support reproducible research on GUI agents and should not be interpreted as a recommendation to deploy autonomous agents on physical laboratory equipment without appropriate safety validation. 9
LabOSWorld
A P REPRINT
Table 3: XRD subtasks. Subtask
Level
Description
1. Select Specimen 2. Load Sample 3. Power Up 4. Set Angles 5. Set Step Size 6. Set Scan Rate 7. Run Scan 8. Save Result
E ASY M EDIUM M EDIUM M EDIUM M EDIUM M EDIUM E ASY E ASY
Select a specimen from the dropdown menu. Click DOORS to open the chamber, load the sample, and then close the chamber. Click STANDBY to power up the instrument and wait until it is ready. Use the angle adjustment buttons to set the start and end angles. Select the step size from the dropdown menu. Select the scan rate from the dropdown menu. Click START SCAN to begin scanning. Click SAVE DIFFRACTOGRAM to save the scan result.
Subtask
Level
Description
1. Vent Chamber 2. Pump Down 3. Select Sample 4. E-beam On 5. E-beam Live Focus 6. WD 7 mm 7. Tilt 10◦ 8. Stage Z Center 9. Ion Beam Live Center 10. First Rect Start 11. Delete Pattern 12. Second Rect Start 13. Beam Current 10 pA 14. Pt Needle In 15. Pt Deposition Start 16. Ion Snapshot 5000× 17. Cross Section Cut Start 18. Cleaning Section Start
E ASY E ASY E ASY E ASY M EDIUM M EDIUM M EDIUM M EDIUM M EDIUM E ASY E ASY M EDIUM M EDIUM E ASY H ARD M EDIUM H ARD H ARD
19. Tilt 0◦ 20. Task Complete
E ASY E ASY
Click VENT once and wait for the chamber animation to complete. Click PUMP once and wait for the vacuum process to complete. Select Si Wafer from the SAMPLE dropdown menu. Click the electron-beam HT switch. Focus the electron-beam image. Adjust the WD slider to approximately 7 mm. Set the stage tilt to approximately 10◦ . Move along the Z axis to a suitable height. Center the target region in the ion-beam field of view. Click START once on the correctly selected pattern. Click DELETE PATTERN once to remove the current pattern. Select the second rectangle and click START. Set the beam current and use DELETE PATTERN. Click to extend the Pt needle. Select the Pt Deposition pattern, place it at the correct position, and click START. Set the magnification to approximately 5000× and capture an ion-beam snapshot. Select the cross-section cutting pattern and click in the correct region. Select the cutting pattern, set the beam current to 0.1 nA, drag the pattern to the target region, and click START. Check Cross Section for imaging and set the stage tilt to 0◦ . Click SAVE to store the cut image.
Table 4: FIB subtasks.
Resources Project Page: https://su-ise-2001.github.io/LABOSBENCH/ The project website contains benchmark documentation, simulator demonstrations, task definitions, evaluation results, and future updates.
A
Benchmark Task Details
This appendix provides the complete list of 96 subtasks used in LabOSBench. Each scientific-instrument simulator is decomposed into a sequence of operation-level subtasks. For each subtask, we report its name, difficulty level, and natural-language description. The difficulty levels are assigned according to the type of GUI interaction, the degree of scientific-state interpretation required, and the amount of feedback-driven adjustment involved. A.1
Task Construction Principles
The subtasks in LabOSBench are constructed according to scientific workflow stages rather than arbitrary GUI regions. Each subtask corresponds to an operation-level unit that changes the simulator state, configures an instrument parameter, performs a scientific action, or finalizes an output. This design preserves the scientific meaning of the original workflow while enabling diagnostic subtask-level evaluation. A.2
Complete Subtask List
The following eight tables summarize the complete subtask lists for all scientific-instrument simulators. These subtasks are used for the subtask-level evaluation reported in the main paper. 10
LabOSWorld
A P REPRINT
Table 5: LFM subtasks. Subtask
Level
Description
1. Drag Sample 2. Select Brightfield 3. Select Halogen 4. Select 10× Focus 5. Field Diaphragm Close 6. Field Diaphragm Focus 7. Field Diaphragm Center 8. Field Diaphragm Open 9. Aperture Diaphragm 10. White Balance 11. Exposure Adjust 12. BF Capture
E ASY E ASY E ASY M EDIUM E ASY H ARD H ARD E ASY H ARD H ARD M EDIUM E ASY
Drag the H&E-stained kidney specimen onto the microscope stage. Check the SETUP & BRIGHTFIELD mode. Click HALOGEN LAMP to select the halogen lamp. Select the 10× objective and use the FOCUS slider. Push the slider fully right to close the field diaphragm. Adjust the slider so that the diaphragm edge is sharply focused. Use the up, down, left, and right buttons to center the diaphragm. Push the slider left to open the diaphragm until the blades just exceed the field of view. Click to remove the right eyepiece, adjust the aperture diaphragm, and reinstall the eyepiece. Click the target, move to a blank area, click to set white balance, and then move back. Adjust the exposure slider for optimal exposure and click LUT. Click LUT to check the dynamic range and then click CAPTURE.
Subtask
Level
Description
1. Vent Chamber 2. Open Chamber 3. Close Chamber 4. Evacuate Chamber 5. Select Sample 6. Turn On HT 7. Set Acc Voltage 8. Set Contrast 9. Adjust Clarity 10. Start Scan 11. Save Image 12. Mosaic ROI Region Save
E ASY E ASY M EDIUM E ASY E ASY M EDIUM H ARD H ARD H ARD E ASY E ASY H ARD
Click VENT to release vacuum and allow chamber access. Open the chamber door to load the sample. Close the chamber door after loading. Click PUMP to evacuate the chamber. Select the specimen from the SAMPLE dropdown menu. Click the HT button to turn on the electron beam. Set the acceleration voltage from the dropdown menu. Adjust the contrast slider for optimal image quality. Use the focus control to achieve a sharp image. Click START SCAN to begin acquisition. Click SAVE to store the acquired image. In the 2×2 mosaic view, pan or double-click to adjust the field of view, use the MAGNIFICATION slider to zoom into vertically aligned parallel tubular or stripe-like structures, center and sharpen the region, and click SAVE IMAGE.
Table 6: SEM subtasks.
A.3
Difficulty Level Definitions
We annotate each subtask with a coarse difficulty level to characterize the interaction complexity. E ASY subtasks typically involve a single explicit GUI action, such as clicking a clearly labeled button or selecting an item from a dropdown menu. M EDIUM subtasks require a short sequence of actions, parameter adjustment, or waiting for a simulator state transition. H ARD subtasks require feedback-driven control, precise spatial grounding, scientific-state interpretation, or multi-step manipulation under visual feedback. The difficulty labels are intended for analysis and interpretation rather than for changing the evaluation protocol. All subtasks are evaluated using the same action interface and success-checking mechanism. A.4
Mapping to High-Level Task Categories
For category-level analysis, each subtask is further mapped to one of five high-level operation categories: Preparation & Sample Handling, Activation & Alignment, Measurement Configuration, Experimental Execution, and Post-processing & Completion. The mapping is based on the primary scientific operation required by the subtask rather than the low-level GUI action. For example, a click on a button may be categorized differently depending on whether it loads a sample, activates a beam, starts a scan, or saves a result. Table 11 defines the five high-level categories and lists the operation-level subtasks assigned to each category.
B
Evaluation Protocol and Implementation Details
This appendix provides additional details about the evaluation protocol and implementation of LabOSBench. The goal is to make the browser-based execution process, action interface, subtask initialization, step budget, and logging format explicit. All models and agentic frameworks are evaluated through the same interaction protocol unless otherwise specified. 11
LabOSWorld
A P REPRINT
Table 7: EDS subtasks. Subtask
Level
Description
1. Select Point Mode 2. Pick Sample Point 3. Open Label 4. Open Semi Quant 5. Open Maps 6. Toggle Si Mapping 7. Open Composite 8. Return to Main
E ASY E ASY E ASY M EDIUM E ASY E ASY E ASY E ASY
Click Point to select point-analysis mode. Mark the sampling point for subsequent analysis. Open the label of the sampling point or spectrum. Open the semi-quantitative analysis window. Open the element distribution map. Turn on or off the silicon element distribution map layer. Open the composite diagram. Return to the main control panel.
Subtask
Level
Description
1. Select Tapping Mode 2. Laser Alignment 3. Photodiode Alignment 4. Set Target Amplitude 5. Set Frequency 6. Auto Tune 7. Set Scan Size 8. Set Integral Gain 9. Set Scan Rate 10. Set Set Point 11. Motor Approach 12. Engage 13. Scan 14. Save
E ASY H ARD H ARD E ASY M EDIUM E ASY E ASY M EDIUM E ASY M EDIUM M EDIUM E ASY E ASY E ASY
Select TAPPING from the MODE dropdown menu. Align the laser spot onto the cantilever using the alignment controls. Center the laser spot on the photodiode using the photodiode controls. Set the target amplitude to 500 mV. Set the frequency range to 200–500 kHz. Click the AUTO TUNE button to initiate automated tuning. Select the desired scan size from the SCAN SIZE menu. Set the integral gain value for the feedback loop. Select the scan rate from the SCAN RATE menu. Adjust the set point slider to control the imaging force. Use the MOTOR slider to bring the probe toward the sample surface. Click the ENGAGE button to begin the automated approach. Click the SCAN button to initiate the topographic scan. Click the SAVE button to export the acquired image.
Table 8: SPM subtasks.
B.1
Observation and Action Space
At each interaction step, the agent receives a natural-language task instruction and the current screenshot of the simulator interface. The agent then outputs one executable GUI action. We use a unified action interface across all simulators so that the same evaluation runner can be applied to different scientific-instrument workflows. The action space includes clicking, double-clicking, dragging, typing, selecting dropdown options, pressing keys, scrolling, and waiting for simulator state transitions. Click and double-click actions are used for buttons, panels, and visual target regions; drag actions support sliders, stage movement, and pattern placement; type and select actions handle numeric inputs and menus; wait actions allow instrument animations, scans, reconstructions, or state transitions to complete. For coordinate-based actions, the runner maps the agent-produced coordinate to the current browser viewport. When an action targets a standard HTML element, the runner uses DOM-level event dispatch when possible. For canvas-like or JavaScript-rendered controls, it falls back to mouse events. This hybrid execution strategy allows LabOSBench to support both conventional web widgets and simulator-specific interactive controls.
B.2
Interaction Loop
Each evaluation episode follows a closed-loop observe–act–execute protocol. At the beginning of an episode, the runner launches the target simulator page, applies the required initial state, and provides the agent with the task instruction and the initial screenshot. At each step, the agent predicts the next GUI action based on the instruction and the current observation. The runner parses the action, executes it in the browser, waits for the simulator state to update, captures a new screenshot, and queries the in-page benchmark state. The loop terminates when the target subtask or workflow is successfully completed, when the step budget is exhausted, or when the runner detects an unrecoverable execution error. Invalid or unparsable actions are recorded in the episode log and counted as failed interaction attempts. This protocol allows the benchmark to capture both successful task completion and intermediate failure behavior, such as repeated clicks, wrong coordinate selection, ineffective waiting, or incorrect parameter manipulation. 12
LabOSWorld
A P REPRINT
Table 9: APT subtasks. Subtask
Level
Description
1. Select Sample 2. Align Sample 3. Set Temperature 4. Set Detection 5. Set Frequency 6. Set Pulse 7. Start Experiment 8. Wait Completion 9. Start Reconstruction 10. Set ICF 11. Set K-Factor 12. Finish Experiment
E ASY H ARD M EDIUM M EDIUM M EDIUM M EDIUM H ARD H ARD H ARD E ASY E ASY E ASY
Select a sample from the Sample menu. Click ALIGN SAMPLE and wait for the alignment process to finish. Set the specimen temperature. Set the detection rate. Set the voltage pulse frequency. Set the voltage pulse fraction. Click START to run the experiment. Wait for the experiment to finish and click again if needed. Click RECONSTRUCT to enter the reconstruction view. Choose an ICF parameter. Choose a K-factor parameter. Click Finish and then close the popup window.
Subtask
Level
Description
1. Select Sample 2. Remove Holder 3. Pump Airlock 4. Insert Specimen 5. Set Voltage 6. Enable Beam 7. Adjust Magnification 8. Move Stage 9. Focus Objective 10. Acquire Image
E ASY E ASY E ASY E ASY E ASY H ARD H ARD H ARD H ARD E ASY
Select a sample from the SAMPLE dropdown menu. Click REMOVE to take out the empty holder. Click PUMP in the AIRLOCK panel to evacuate the airlock. Click INSERT to insert the sample. Choose an accelerating voltage from the ACC menu. Turn the beam on and adjust the beam intensity. Select a magnification and adjust image brightness. Move the stage in X, Y, and Z to center the region of interest. Adjust the OBJECTIVE LENS FOCUS to sharpen the image. Insert the camera and click ACQUIRE to capture a TEM image.
Table 10: TEM subtasks.
B.3
Subtask Initialization
For subtask-level evaluation, each simulator is initialized to a canonical state immediately before the target subtask. This initialization isolates the local capability required by the target operation and avoids conflating subtask performance with failures accumulated in earlier workflow stages. For example, when evaluating a parameter-setting subtask, the simulator is first placed into the state where the relevant control is available and all prerequisite subtasks have already been completed. When a simulator supports programmatic state control, we use instrument-specific fast-forward functions exposed through the browser window. These functions mark prerequisite subtasks as completed and restore the GUI to the canonical pre-subtask state. Fast-forwarding is used only for diagnostic subtask-level evaluation. It is not used in full-episode evaluation, where the agent must complete the entire workflow from the initial simulator state. This design allows LabOSBench to distinguish local GUI-control ability from long-horizon workflow execution. An agent that succeeds under subtask initialization but fails in full-episode evaluation may suffer from planning, ordering, memory, or recovery problems. In contrast, failure under subtask initialization indicates that the agent cannot reliably complete the local scientific-instrument operation itself. B.4
Step Budget and Termination
For subtask-level evaluation, each model is allowed up to 50 interaction steps per subtask. A subtask episode is marked as successful if the corresponding success condition is satisfied within the step budget. Otherwise, the episode is marked as failed after the budget is exhausted. For full-episode evaluation, the agent starts from the initial simulator state and must complete all required subtasks in order. A full episode is successful only if all required workflow stages are completed within the specified workflow-level step budget. This setting evaluates long-horizon execution and captures error accumulation across the entire scientific workflow. An episode may terminate under three conditions: (i) all target success conditions are satisfied, (ii) the maximum number of allowed interaction steps is reached, or (iii) the runner encounters an unrecoverable execution error. All termination causes are recorded in the episode log. 13
LabOSWorld
A P REPRINT
Table 11: Definitions of the five high-level task categories and their assigned operation-level subtasks. Task Category
Definition
Assigned Subtasks
Preparation & Sample Handling
Operations that prepare the simulator, select the target object, or load the required sample before measurement or manipulation.
SEM-S1 VentChamber; SEM-S2 OpenChamber; SEM-S3 CloseChamber; SEM-S4 EvacuateChamber; SEM-S5 SelectSample; TEM-T1 Select sample; TEM-T2 Remove empty holder; TEM-T3 Airlock pump; TEM-T4 Insert specimen; XRD-S1 SelectSpecimen; XRD-S2 Doors; LFM-L1 DragSample; FIB-F1 VentChamber; FIB-F2 PumpDown; FIB-F3 SelectSample; APT-S1 SelectSample; EDS-S2 PickSamplePoint.
Activation & Alignment
Operations that activate instrument components or align the system into a valid operating state.
SEM-S6 TurnOnHT; SPM-S1 SelectTappingMode; SPM-S2 LaserAlignment; SPM-S3 PhotodiodeAlignment; SPM-S11 MotorApproach; SPM-S12 Engage; TEM-T5 Select acceleration voltage; TEM-T6 Beam on and filament current; XRD-S3 PowerUp; LFM-L2 SelectBrightfield; LFM-L3 SelectHalogen; LFM-L4 Select10xFocus; LFM-L5 FieldDiaphragmClose; LFM-L6 FieldDiaphragmFocus; LFM-L7 FieldDiaphragmCenter; LFM-L8 FieldDiaphragmOpen; LFM-L9 ApertureDiaphragm; FIB-F4 EbeamOn; FIB-F5 EbeamLiveFocus; FIB-F8 StageZCenter; FIB-F9 IonBeamLiveCenter; APT-S2 AlignSample.
Measurement Configuration
Operations that configure parameters SEM-S7 SetAccVoltage; SEM-S8 SetContrast; SEM-S9 AdjustClarity; SPM-S4 before acquisition, scanning, fabrication, SetTargetAmplitude; SPM-S5 SetFrequency; SPM-S6 AutoTune; SPM-S7 or reconstruction. SetScanSize; SPM-S8 SetIntegralGain; SPM-S9 SetScanRate; SPM-S10 SetSetPoint; TEM-T7 Magnification and brightness; TEM-T8 Stage position centering; TEM-T9 Objective lens focus; XRD-S4 SetAngles; XRD-S5 SetStepSize; XRD-S6 SetScanRate; LFM-L10 WhiteBalance; LFM-L11 ExposureAdjust; FIB-F6 WD7mm; FIB-F7 Tilt10deg; FIB-F13 BeamCurrent10pA; FIB-F16 IonSnapshot5000x; APT-S3 SelectTemperature; APT-S4 SelectDetectionRate; APT-S5 SelectPulseFreq; APT-S6 SelectPulseEnergy; EDS-S1 SelectPointMode.
Experimental Execution
Operations that start, monitor, or complete the main scientific action after the setup is ready.
SEM-S10 StartScan; SPM-S13 Scan; TEM-T10 Camera insert and acquire; XRD-S7 RunScan; LFM-L12 BFCapture; FIB-F10 FirstRectStart; FIB-F11 DeletePattern; FIB-F12 SecondRectStart; FIB-F14 PtNeedleIn; FIB-F15 PtDepositionStart; FIB-F17 CrossSectionCutStart; FIB-F18 CleaningCrossSectionStart; APT-S7 StartExperiment; APT-S8 WaitCompletion.
Post-processing & Completion
Operations that inspect, save, export, reconstruct, or finalize experimental outputs.
SEM-S11 SaveImage; SEM-S12 MosaicRoiRegionSave; SPM-S14 Save; XRD-S8 SaveResult; FIB-F19 Tilt0deg; FIB-F20 TaskComplete; APT-S9 Reconstruct; APT-S10 SetICF; APT-S11 SetKFactor; APT-S12 Finish; EDS-S3 OpenLabel; EDS-S4 OpenSemiQuant; EDS-S5 OpenMaps; EDS-S6 ToggleSiMapping; EDS-S7 OpenComposite; EDS-S8 ReturnToMain.
B.5
Logging Format
After each evaluation, the runner exports a structured JSON log for each simulator and model. The log records the number of runs per subtask, subtask-level outcomes, run-level episode metadata, diagnostic grounding metrics, and overall aggregated statistics. This format allows us to compute the reported subtask success rates while preserving additional diagnostic information for failure analysis. At the top level, each log contains num_runs_per_subtask, a list of subtasks, and an overall_metrics_summary. Each subtask entry stores the subtask identifier, subtask name, number of runs, success count, success rate, averaged diagnostic metrics, and the detailed records of individual runs. Each run record stores whether the run succeeds, the directory of the corresponding episode trace, and run-level metrics such as actual interaction steps, grounding accuracy, target-subtask attempts, and target-subtask completion. The official subtask-level score is computed from success_count and success_rate. Other fields, such as actual_steps, target_subtask_attempts, widget_grounding_accuracy, text_grounding_accuracy, state_grounding_accuracy, and target_subtask_success, are used as diagnostic signals rather than as the primary score. This distinction allows the benchmark to report final task success while still analyzing intermediate grounding behavior and failure modes. For anonymity, machine-specific absolute paths in episode_dir are removed or anonymized before release. B.6
Implementation Notes
The evaluation runner is implemented as a lightweight browser-based pipeline. It launches local simulator pages, captures screenshots, dispatches GUI actions, and exports structured logs after each episode. This design avoids full operating-system virtualization while retaining executable interaction with realistic scientific-instrument interfaces. Because the simulator state is logged at each step, the same traces can be used for quantitative scoring, failure localization, and qualitative case analysis. 14
LabOSWorld
A P REPRINT
Table 12: Main fields in the exported evaluation logs. Level
Field
Description
Global Global Global
num_runs_per_subtask subtasks overall_metrics_summary
Number of repeated runs used for each subtask. List of subtask-level evaluation records. Aggregated diagnostic metrics across all subtasks in the simulator.
Subtask Subtask Subtask Subtask Subtask Subtask
subtask_id subtask_name num_runs success_count success_rate metrics_summary
Subtask
runs
Identifier of the target subtask, such as S1. Name of the target subtask. Number of runs available for the subtask. Number of successful runs for the subtask. Success rate of the subtask over repeated runs. Averaged diagnostic metrics over runs, such as actual steps, grounding accuracy, and target-subtask attempts. Detailed records for individual runs.
Run Run Run
run success episode_dir
Run
metrics
Metric Metric Metric Metric Metric
actual_steps target_subtask_attempts widget_grounding_accuracy text_grounding_accuracy state_grounding_accuracy
Metric
target_subtask_success
Index of the repeated run. Whether the run is judged successful by the evaluation script. Directory storing the episode trace and screenshots. Paths are anonymized in released logs. Run-level diagnostic metrics. Number of interaction steps used by the agent. Number of detected attempts related to the target subtask. Whether the agent correctly grounds the target GUI widget when applicable. Whether the agent correctly grounds the relevant text or label when applicable. Whether the agent correctly reaches or recognizes the required simulator state when applicable. Whether the target-subtask event is detected in the in-page log.
Table 13: Subtask-level success rates on SEM. ID Subtask S1 VentChamber S2 OpenChamber S3 CloseChamber S4 EvacuateChamber S5 SelectSample S6 TurnOnHT S7 SetAccVoltage S8 SetContrast S9 AdjustClarity S10 StartScan S11 SaveImage S12 MosaicRoiRegionSave
C
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo 1.00 1.00 1.00 1.00 0.50 1.00 1.00 0.50 1.00 1.00 1.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00
0.50 1.00 1.00 1.00 0.00 1.00 0.50 0.00 0.50 1.00 1.00 0.00
1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00 0.50 1.00 1.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00
1.00 1.00 0.00 1.00 1.00 1.00 0.00 0.00 0.00 1.00 0.00 0.00
1.00 1.00 1.00 1.00 0.50 1.00 1.00 0.50 1.00 1.00 1.00 0.00
1.00 1.00 1.00 1.00 0.50 1.00 1.00 1.00 1.00 1.00 1.00 0.00
1.00 1.00 1.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
1.00 1.00 1.00 1.00 0.50 1.00 1.00 0.50 1.00 1.00 1.00 0.00
Subtask-level Results
This appendix reports the subtask-level success rates of all evaluated models on the LabOSBench benchmark. Tables 13– 20 provide the complete results grouped by scientific-instrument simulator. Each row represents an operation-level subtask, and all scores are reported in the range [0, 1].
15
LabOSWorld
A P REPRINT
Table 14: Subtask-level success rates on SPM. ID Subtask
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo
S1 SelectTappingMode S2 LaserAlignment S3 PhotodiodeAlignment S4 SetTargetAmplitude S5 SetFrequency S6 AutoTune S7 SetScanSize S8 SetIntegralGain S9 SetScanRate S10 SetSetPoint S11 MotorApproach S12 Engage S13 Scan S14 Save
1.00 0.00 0.00 0.50 0.50 1.00 1.00 0.50 1.00 0.50 0.00 1.00 1.00 1.00
1.00 0.00 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 0.00 1.00 1.00 1.00
1.00 0.00 0.00 1.00 1.00 1.00 1.00 0.00 1.00 0.00 1.00 1.00 1.00 1.00
1.00 0.00 0.00 0.00 0.00 1.00 1.00 1.00 0.50 0.00 0.00 0.00 0.00 0.50
1.00 0.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00
1.00 0.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00
1.00 0.00 0.00 1.00 1.00 1.00 1.00 0.00 1.00 0.00 0.00 1.00 1.00 1.00
1.00 0.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.50 0.00 1.00 1.00 1.00
1.00 0.00 0.00 1.00 0.50 1.00 1.00 0.50 1.00 0.50 0.50 1.00 1.00 1.00
1.00 0.50 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.50 0.50 1.00 1.00 1.00
1.00 0.00 0.00 0.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00 1.00
1.00 0.00 0.00 1.00 0.50 1.00 1.00 0.50 1.00 0.50 0.50 1.00 1.00 1.00
ID Subtask
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo
Table 15: Subtask-level success rates on TEM. T1 Select sample T2 Remove empty holder T3 Airlock pump T4 Insert specimen T5 Select acceleration voltage T6 Beam on and filament current T7 Magnification and brightness T8 Stage position centering T9 Objective lens focus T10 Camera insert and acquire
0.50 1.00 0.50 1.00 0.50 0.00 1.00 0.00 0.00 0.50
1.00 1.00 0.00 1.00 1.00 0.00 1.00 0.00 0.00 0.00
0.00 1.00 0.00 1.00 1.00 0.00 1.00 0.00 0.00 0.00
1.00 1.00 0.00 0.00 1.00 1.00 0.00 0.00 0.00 1.00
1.00 1.00 1.00 1.00 1.00 0.50 0.00 0.00 0.00 1.00
0.00 1.00 1.00 1.00 0.00 1.00 1.00 0.00 0.00 0.00
0.00 1.00 0.00 1.00 1.00 0.00 1.00 0.00 0.00 0.00
1.00 1.00 1.00 1.00 1.00 0.00 0.00 0.00 0.00 1.00
0.50 1.00 0.50 1.00 0.50 0.00 1.00 0.00 0.00 0.50
1.00 1.00 1.00 1.00 1.00 0.50 1.00 0.50 0.50 1.00
1.00 1.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 0.00
0.50 1.00 0.50 1.00 0.50 0.00 1.00 0.00 0.00 0.50
Table 16: Subtask-level success rates on XRD. ID Subtask
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo
S1 SelectSpecimen S2 Doors S3 PowerUp S4 SetAngles S5 SetStepSize S6 SetScanRate S7 RunScan S8 SaveResult
0.50 1.00 1.00 0.50 1.00 0.50 0.50 1.00
1.00 1.00 1.00 0.00 1.00 1.00 0.00 1.00
1.00 0.00 1.00 0.00 1.00 1.00 0.00 1.00
1.00 1.00 1.00 1.00 1.00 0.00 1.00 1.00
1.00 1.00 1.00 1.00 1.00 0.00 1.00 1.00
0.00 1.00 1.00 0.00 1.00 1.00 0.00 1.00
1.00 1.00 1.00 0.00 1.00 1.00 0.00 1.00
1.00 1.00 0.50 1.00 1.00 0.00 1.00 1.00
0.50 1.00 1.00 0.50 1.00 0.50 0.50 1.00
1.00 1.00 1.00 0.50 1.00 1.00 0.50 1.00
1.00 1.00 1.00 0.00 0.00 1.00 0.00 1.00
0.50 1.00 1.00 0.50 1.00 0.50 0.50 1.00
ID Subtask
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo
Table 17: Subtask-level success rates on LFM. L1 DragSample L2 SelectBrightfield L3 SelectHalogen L4 Select10xFocus L5 FieldDiaphragmClose L6 FieldDiaphragmFocus L7 FieldDiaphragmCenter L8 FieldDiaphragmOpen L9 ApertureDiaphragm L10 WhiteBalance L11 ExposureAdjust L12 BFCapture
1.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
1.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
1.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
1.00 1.00 1.00 0.00 1.00 0.00 0.00 1.00 0.00 1.00 1.00 1.00
16
1.00 1.00 1.00 1.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
1.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
1.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
1.00 1.00 0.50 0.00 1.00 0.00 0.00 1.00 0.00 0.00 1.00 1.00
0.50 0.50 0.00 0.00 0.50 0.00 0.00 0.50 0.00 0.00 1.00 1.00
1.00 1.00 0.50 0.50 1.00 0.00 0.50 1.00 0.50 0.50 1.00 1.00
1.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.00 1.00
1.00 1.00 0.50 0.00 1.00 0.00 0.00 1.00 0.00 0.50 1.00 1.00
LabOSWorld
A P REPRINT
Table 18: Subtask-level success rates on FIB. ID Subtask
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo
F1 VentChamber F2 PumpDown F3 SelectSample F4 EbeamOn F5 EbeamLiveFocus F6 WD7mm F7 Tilt10deg F8 StageZCenter F9 IonBeamLiveCenter F10 FirstRectStart F11 DeletePattern F12 SecondRectStart F13 BeamCurrent10pA F14 PtNeedleIn F15 PtDepositionStart F16 IonSnapshot5000x F17 CrossSectionCutStart F18 CleaningCrossSectionStart F19 Tilt0deg F20 TaskComplete
0.50 1.00 1.00 0.50 0.00 0.50 0.00 0.00 0.00 0.00 0.00 0.00 0.50 0.50 0.00 0.50 0.00 0.50 0.00 0.00
1.00 1.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.00 1.00 0.00 1.00 0.00 1.00 0.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
1.00 0.50 1.00 0.50 0.50 0.50 0.00 0.00 0.00 1.00 0.50 0.00 0.50 1.00 0.00 1.00 0.50 1.00 0.50 0.50
1.00 1.00 1.00 1.00 0.00 0.50 0.00 0.00 0.00 1.00 0.50 0.00 1.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00
1.00 1.00 1.00 1.00 1.00 1.00 0.00 0.00 0.00 0.00 1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
1.00 1.00 1.00 1.00 0.00 1.00 0.00 0.00 0.00 1.00 1.00 0.00 1.00 1.00 0.00 1.00 0.00 1.00 1.00 1.00
1.00 1.00 1.00 0.50 0.00 0.50 0.00 0.00 0.00 0.00 0.50 0.00 0.50 0.50 0.00 0.50 0.00 0.50 0.00 0.00
1.00 1.00 1.00 1.00 0.50 1.00 0.00 0.00 0.50 0.50 0.50 0.50 1.00 1.00 0.50 1.00 0.50 1.00 0.50 0.50
1.00 1.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
1.00 1.00 1.00 0.50 0.00 0.50 0.00 0.00 0.00 0.00 0.00 0.00 0.50 0.50 0.00 0.50 0.00 0.50 0.00 0.00
Table 19: Subtask-level success rates on APT. ID Subtask
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo
S1 SelectSample S2 AlignSample S3 SelectTemperature S4 SelectDetectionRate S5 SelectPulseFreq S6 SelectPulseEnergy S7 StartExperiment S8 StopExperiment S9 Reconstruct S10 SetICF S11 SetKFactor S12 Finish
1.00 0.50 0.50 1.00 0.50 1.00 0.50 1.00 0.50 0.50 0.50 0.50
0.50 1.00 1.00 1.00 0.50 1.00 1.00 1.00 1.00 0.50 0.00 0.00
1.00 0.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00 0.00 0.00 1.00
1.00 0.00 1.00 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00
1.00 0.50 1.00 1.00 1.00 1.00 1.00 1.00 0.00 1.00 1.00 1.00
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.50 0.50
1.00 0.00 0.00 1.00 1.00 1.00 1.00 1.00 0.00 0.00 0.00 1.00
1.00 0.00 1.00 1.00 1.00 1.00 0.00 0.00 0.00 1.00 1.00 1.00
1.00 0.50 0.50 1.00 0.50 1.00 0.50 1.00 0.50 0.50 0.50 0.50
1.00 0.50 1.00 1.00 1.00 1.00 1.00 1.00 0.50 0.50 0.50 1.00
1.00 0.00 0.00 1.00 0.00 1.00 1.00 1.00 0.00 0.00 0.00 1.00
1.00 0.50 0.50 1.00 0.50 1.00 0.50 0.50 0.50 0.50 0.00 0.50
Table 20: Subtask-level success rates on EDS. ID Subtask S1 SelectPointMode S2 PickSamplePoint S3 OpenLabel S4 OpenSemiQuant S5 OpenMaps S6 ToggleSiMapping S7 OpenComposite S8 ReturnToMain
Qwen3VL-32B EvoCUA Sonnet-4.5 Kimi Seed GPT-5.5 Opus-4.5 UI-TARS GUI-Owl GTA1 VLAA Hippo 0.50 0.50 0.50 0.50 1.00 1.00 1.00 1.00
1.00 1.00 1.00 0.50 0.00 1.00 0.00 0.00
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
0.00 1.00 1.00 0.00 1.00 0.50 1.00 0.00
1.00 0.50 1.00 0.50 1.00 1.00 0.50 1.00
17
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
1.00 0.50 1.00 1.00 0.00 1.00 1.00 0.50
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
1.00 0.50 1.00 0.50 1.00 1.00 1.00 1.00
1.00 0.00 1.00 1.00 1.00 1.00 1.00 1.00
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
LabOSWorld
A P REPRINT
References Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022. Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al. UI-TARS: Pioneering automated GUI interaction with native agents. arXiv preprint arXiv:2501.12326, 2025. Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems, 37:52040–52094, 2024. Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, et al. Windows agent arena: Evaluating multi-modal os agents at scale. arXiv preprint arXiv:2409.08264, 2024. Chris Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, et al. Androidworld: A dynamic benchmarking environment for autonomous agents. In International Conference on Learning Representations, volume 2025, pages 406–441, 2025. Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. WebArena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, volume 2024, pages 15585–15606, 2024. Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2Web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36:28091–28114, 2023. Gary Tom, Stefan P Schmid, Sterling G Baird, Yang Cao, Kourosh Darvish, Han Hao, Stanley Lo, Sergio Pablo-García, Ella M Rajaonson, Marta Skreta, et al. Self-driving laboratories for chemistry and materials science. Chemical Reviews, 124(16):9633–9732, 2024. Nathan J Szymanski, Bernardus Rendy, Yuxing Fei, Rishi E Kumar, Tanjin He, David Milsted, Matthew J McDermott, Max Gallant, Ekin Dogus Cubuk, Amil Merchant, et al. An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature, 624(7990):86, 2023. Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023. Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. ScienceWorld: Is your agent smarter than a 5th grader? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11279–11298, 2022. Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. DiscoveryWorld: A virtual environment for developing and evaluating automated scientific discovery agents. Advances in Neural Information Processing Systems, 37:10088–10116, 2024. Han Deng, Anqi Zou, Hanling Zhang, Ben Fei, Chengyu Zhang, Haobo Wang, Xinru Guo, Zhenyu Li, Xuzhu Wang, Peng Yang, et al. Owl-auraid 1.0: An intelligent system for autonomous scientific instrumentation and scientific data analysis. arXiv preprint arXiv:2603.29828, 2026. Qijun Han, Haoqin Tu, Zijun Wang, Haoyue Dai, Yiyang Zhou, Nancy Lau, Alvaro A Cardenas, Yuhui Xu, Ran Xu, Caiming Xiong, et al. VLAA-GUI: Knowing when to stop, recover, and search, a modular framework for GUI automation. arXiv preprint arXiv:2604.21375, 2026. Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. WebVoyager: Building an end-to-end web agent with large multimodal models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6864–6890, 2024. Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. GPT-4V(ision) is a generalist web agent, if grounded. arXiv preprint arXiv:2401.01614, 2024. Mariya Davydova, Daniel Jeffries, Patrick Barker, Arturo Márquez Flores, and Sinéad Ryan. Osuniverse: Benchmark for multimodal gui-navigation ai agents. arXiv preprint arXiv:2505.03570, 2025. Yifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu, Hanchen Zhang, Bohao Jing, Shudan Zhang, Yuting Wang, Wenyi Zhao, and Yuxiao Dong. Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025. URL https://arxiv.org/abs/2509.18119. Zhaoyang Liu, JingJing Xie, Zichen Ding, Zehao Li, Bowen Yang, Zhenyu Wu, Xuehui Wang, Qiushi Sun, Shi Liu, Weiyun Wang, et al. Scalecua: Scaling open-source computer use agents with cross-platform data. arXiv preprint arXiv:2509.15221, 2025. 18
LabOSWorld
A P REPRINT
Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Qiushi Sun, Zhaoyang Liu, Zhoumianze Liu, Yu Qiao, Xiangyu Yue, Zun Wang, et al. Os-oracle: A comprehensive framework for cross-platform gui critic models. arXiv preprint arXiv:2512.16295, 2025. Zehao Li, Zhenyu Wu, Yibo Zhao, Bowen Yang, Jingjing Xie, Zhaoyang Liu, Zhoumianze Liu, Kaiming Jin, Jianze Liang, Zonglin Li, et al. Os-themis: A scalable critic framework for generalist gui rewards. arXiv preprint arXiv:2603.19191, 2026. Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed Awadallah. OmniParser for pure vision based GUI agent. arXiv preprint arXiv:2408.00203, 2024a. Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua. ScreenSpot-Pro: GUI grounding for professional high-resolution computer use. In Proceedings of the 33rd ACM International Conference on Multimedia, pages 8778–8786, 2025. Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. ChemCrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376, 2023. Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. MLAgentBench: Evaluating language agents on machine learning experimentation. arXiv preprint arXiv:2310.03302, 2023. Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The AI scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292, 2024b. Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 881–905, 2024. Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, and Yuxiao Dong. Androidlab: Training and systematic benchmarking of android autonomous agents, 2024. URL https://arxiv.org/abs/2410.24024. Pei Yang, Hai Ci, and Mike Zheng Shou. macOSWorld: A multilingual interactive benchmark for GUI agents. Advances in Neural Information Processing Systems, 38:134014–134056, 2026. Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, et al. Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflows. arXiv preprint arXiv:2505.19897, 2025. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-VL technical report. arXiv preprint arXiv:2511.21631, 2025. Taofeng Xue, Chong Peng, Mianqiu Huang, Linsen Guo, Tiancheng Han, Haozhe Wang, Jianing Wang, Xiaocheng Zhang, Xin Yang, Dengchang Zhao, et al. EvoCUA: Evolving computer use agents via learning from scalable synthetic experience. arXiv preprint arXiv:2601.15876, 2026. Anthropic. Claude sonnet 4.5 system card. https://www.anthropic.com/claude-sonnet-4-5-system-card, 2025a. Official system card. Accessed: 2026-05-22. Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al. Kimi K2.5: Visual agentic intelligence. arXiv preprint arXiv:2602.02276, 2026. ByteDance Seed Team. Seed1.6: Bytedance seed general-purpose model series. https://seed.bytedance.com/ en/seed1_6, 2025. Official technical page. Accessed: 2026-05-22. OpenAI. Gpt-5.5 system card. https://openai.com/index/gpt-5-5-system-card/, 2026. Official system card. Accessed: 2026-05-22. Anthropic. Claude opus 4.5 system card. https://www.anthropic.com/claude-opus-4-5-system-card, 2025b. Official system card. Accessed: 2026-05-22.
19