Conceptio › Archive › arXiv CS
arXiv CSopen access

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

arXiv:2609.22067v1 [cs.HC] 18 Sep 2026

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw Renkai Ma∗

Ruyuan Wan∗

Xuan Lu

[email protected] School of Information Technology University of Cincinnati Cincinnati, Ohio, United States

College of Information Sciences and Technology The Pennsylvania State University University Park, Pennsylvania, United States

College of Information Science University of Arizona Tucson, Arizona, United States

Fan Yang

Chen Chen

Lingyao Li

University of South Carolina Columbia, South Carolina, United States

[email protected] Knight Foundation School of Computing and Information Sciences Florida International University Miami, Florida, United States

[email protected] College of Information Science University of Arizona Tucson, Arizona, United States

Abstract

AI agent that runs on a user’s own computer and performs tasks through chat apps such as WhatsApp, exemplifies this class of system. Released in November 2025, it had drawn more than 370,000 GitHub stars by August 2026 [40]. Such an agent differs from a typical conversational assistant in where execution occurs and when a user can review it [21, 68]. A typical conversational assistant asked to fix a bug returns a patch for the user to inspect, whereas an agent reads the repository, edits the files, runs the tests, and commits the result under the user’s credentials, with few checkpoints in between [12, 38]. We call one such stretch of delegated work, from the user’s instruction to the outcome the user reads afterward, a run. Consequently, a chatbot that misreads a request returns a poor answer, whereas an agent that misreads one can read files the user never exposed [50], exhaust an allowance the user cannot replenish, or delete work the user cannot recover [67]. An agent can do all three even though human–AI interaction design guidelines have long asked that AI systems make their limits legible and support recovery when they fail [1]. How AI agents are evaluated follows how they are built. Technical research decomposes them into modules that separate memory from planning, and both from the actions that call external tools [57, 62]. Agent evaluations follow this decomposition, scoring whether a workflow completes a task [28, 34, 37, 63, 64, 69], while human–AI alignment often scores moral judgments on prompts [22]. However, a benchmark score reflects the construct its evaluation was built to measure [25, 43], and few evaluations measure users’ experience of these modules. HCI research has begun to examine such experience, showing that how and when a user is involved changes what the user and the agent achieve together [21, 29] and how far the user overrelies on the agent [7, 21]. Less is known about how users meet these modules together, in their own everyday work. We investigate this experience through Friedman et al.’s Value Sensitive Design (VSD), which holds that design should account for human values throughout development and defines a value as what a person or group considers important in life [16, 17]. VSD argues that the properties designers build into a technology more readily support some values and hinder others, yet whether a value is realized depends on the goals of the people interacting with it [17].

Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect’s corpus share, values clustered not at the agent’s outputs but at the operating conditions users set around a run. Values were usually met where users described what the agent delivered, in five of six groups, and mostly unmet where users described supervising it, in all six groups. We conceptualize this pattern as value-sensitive delegation. Supporting human values requires attention not only to what an agent accomplishes, but to the conditions users set around delegation, including cost, access, and oversight.

CCS Concepts • Human-centered computing → Empirical studies in HCI.

Keywords Personal Autonomous Agent, Value Sensitive Design, OpenClaw, Social Media Analysis

1

Introduction

An autonomous agent is an AI system that independently executes tasks and manages workflows on behalf of a user. Users increasingly delegate real work to these agents1 . OpenClaw [8],2 an open-source ∗ Both authors contributed equally to this research. 1We use “AI agents” throughout to refer to autonomous, large language model (LLM)-

based agents. Prior work uses several labels for the same class of system, including “GUI agent,” “LLM-based agent,” and “multi-agent system” [11, 38, 65], as do the posts we analyze, and we preserve the posts’ wording when quoting them. 2 https://github.com/openclaw/openclaw Conference’17, Washington, DC, USA 2026. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM https://doi.org/10.1145/nnnnnnn.nnnnnnn 1

Conference’17, July 2017, Washington, DC, USA

Ma et al.

An autonomous agent changes where that interaction happens, because it acts between the user’s instruction and the outcome the user receives, so a value can be realized while the user is away. We therefore adopt VSD as our analytical lens. We ask which values users invoke when an agent acts for them, and which agent aspects3 those values attach to. Prior work applies VSD to AI primarily through design guidelines [13, 44, 54] and interviews with AI practitioners [45], while separate studies of responsible-AI practice examine how organizations translate values into work [55, 58]. That evidence comes from the developers who build AI systems, not from the users who delegate work to them, and as agents take on longer chains of action with fewer checkpoints, who answers for what an agent does becomes harder to settle from the developer’s side alone. We thus ask two research questions: • RQ1. What human values surface in users’ first-person experience of using an autonomous AI agent like OpenClaw and which agent aspects do those values attach to? • RQ2. What outcomes do users attribute to agent use when those values are met and when they are not? To answer these questions, we collected Reddit posts about OpenClaw, screened them for first-person accounts, and coded each for its human value, agent aspect, value fulfillment, and user outcome. The schema’s 21 values comprise 12 of the 13 values in Friedman et al.’s list [17] and nine study-specific values, such as affordability and meaningful human control, that we developed from the corpus for delegating work to an agent [31]. We mapped all 21 values onto six value groups. RQ1 analyzes the 73,093 coded posts and RQ2 the 44,767 of them that carry an attributable user outcome. Five of the 21 values accounted for 70.8% of posts. The agent aspects most overrepresented relative to their corpus share were the operating conditions users set around a run, namely what it cost, what it could reach, what it recorded, when it had to ask, and how it was installed. The model core, the underlying LLM that produces the agent’s output, was not among them (Section 4.2). Fulfillment varied from 77.3% met in Autonomous Operation to 42.8% in Equitable Access, and resource accounting had the lowest met rate of all 18 agent aspects at 35.3% (Section 4.3). Value fulfillment split what an agent delivered from what delegating to it cost and risked, with the post’s value met in 67.7% of task-effectiveness posts but in 34.7% of resource-burden and 10.7% of risk-exposure posts (Section 5.7). Value fulfillment and violation were therefore registered in these posts at agent aspects that the evaluations we reviewed do not cover [25, 43], and posts about what an agent could do on its own carried different fulfillment rates from posts about whether it could be relied on, a separation prior work on trust and reliance draws but agent evaluation does not [7, 56]. Because these values sat in what users had set before a run rather than in what the run returned, we discuss them as value-sensitive delegation, a relationship between a human value and the condition that carries it. Our contributions are threefold: (1) An empirical characterization of the human values that surface in 73,093 first-person posts of OpenClaw use, and of where in an agent each of those human values attaches. (2) The conceptualization of value-sensitive delegation,

showing that values fell short at the operating conditions governing cost, reach, and setup at higher rates than at the model core of the OpenClaw agent, and that posts about an agent’s reach, its reliability, its costs, and what it delivered carried systematically different fulfillment rates. That difference shifts the target of valuealignment work from the model to the boundaries users set around it. (3) Design implications keyed to the agent aspects where the shortfalls concentrate.

2

Related Work

We situate this study in two bodies of work. Section 2.1 reviews how autonomous agents are evaluated and where such evaluations stop short of user experience. Section 2.2 reviews how VSD has been applied to AI and why that work has not reached the experience of using an agent.

2.1

Evaluating Autonomous AI Agents

Using an autonomous agent is an act of delegation, where a user hands a task to a system, decides how much autonomy to grant, and rejoins the work where they reserved a say. Early work on software agents and automation defined delegation as relying on another’s action to reach a goal and distinguished levels of delegation by how much of the task the delegate decides [9]. Human-factors research described the handover along two axes, which functions a machine takes over and how far it automates each of them [42], and proposed delegation interfaces for supervisors to set the scope of automated systems [36]. More recent work asked which tasks people are willing to delegate to AI and proposed that motivation, difficulty, risk, and trust determine that willingness [35]. LLM-based agents extend this handover, because a single instruction can start a run that edits files, spends money, and contacts external services before the user sees the results [10]. Yet, the evaluation of these agents still centers on the technical question of whether an agent can complete a task in a given test environment. For example, benchmarks like AgentBench [34], WebArena [69], and OSWorld [63] test agents across interactive environments, websites, and desktop applications. Others apply this execution-based logic to domain-specific behavior, such as resolving GitHub issues (e.g., SWE-bench [28]) or reaching a correct database state during customer service interactions (e.g., 𝜏-bench [64]). While observable execution results verify task completion and enable technical capability comparisons, benchmark scholarship cautions against treating specific task performance as evidence of broader capability when relevant contexts are omitted [25, 43]. For AI agents, this omitted context is the user’s experience. A run scored as complete can still have taken supervision the score does not record, hidden further agent actions from the user, and left effects the user had to undo. Security research evaluates these agents by architectural layer to identify where failures originate. For example, an analysis of 470 OpenClaw advisories found the dominant weakness was per-layer trust enforcement rather than a unified policy boundary [52]. Other work has separated cognitive, execution, and information-system risks [66], traced attacks that propagate from prompt processing through tool invocation [60], and demonstrated that poisoning one dimension of an OpenClaw deployment’s persistent state increases attack success [61]. While this work shows that failures concentrate

3We use “agent aspect” to refer to the part of an agent system a user is dealing with when a value comes into play, which may be the model that produces an output, the account that meters a run, or the credentials that let an agent reach a file.

2

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

3

in identifiable parts of a system, it identifies them through advisory analysis and adversarial probing and scores them as attack success or exploitability rather than as issues end users encounter during daily use. Complementing these technical evaluations, HCI evaluates autonomous AI agents as interactive systems. Human–AI interaction design guidelines set out what such systems owe a user, including legible capabilities and support for error recovery [1], protection against overreliance [7], and trust built through human actors rather than system features alone [56]. Research has begun mapping these requirements onto human–agent interaction by establishing design spaces for computer-use agents [12], studying early adopters balancing autonomy with human oversight [38], and analyzing the failures and workarounds of GUI web-browsing agents [67]. However, to our knowledge, no shared framework has yet emerged for comparing how specific aspects of an AI agent shape user-reported experiences and outcomes. Our study addresses this gap by analyzing first-person posts of everyday OpenClaw use on Reddit.

2.2

Methods

To answer our research questions, we conducted a value-centered analysis of first-person Reddit posts about the use of OpenClaw. This section describes our data preparation (Section 3.1), coding schema development and application (Section 3.2), human validation (Section 3.3), statistical and qualitative analysis (Section 3.4), and ethics (Section 3.5).

3.1

Data Preparation

We collected Reddit posts and comments4 created between January 31, 2026 and April 26, 2026. We set the start date at the cutoff for use of the OpenClaw name and the end date at our final data retrieval. We first collected candidate posts and comments with Brandwatch5 , a social-media listening platform whose full-archive search returned Reddit IDs. We then used the Reddit PRAW API6 to retrieve the full content of each matched post and comment. The search used the case-insensitive keywords “openclaw,” “clawdbot,” and “moltbot,” the agent’s current and earlier names, yielding a corpus of 1,100,308 posts, comprising 38,753 original posts and 1,061,555 comments.

VSD and the Lived Experience of Agent Use

3.2

Autonomous AI agents do more than complete tasks. Because a run reaches a user’s files, credentials, and accounts, using an agent raises the question of whether end users retain meaningful control over delegated work, and a completion score does not answer it. VSD was built for questions of that kind, and pursues them through conceptual, empirical, and technical investigations that run throughout design [17]. It asks what a technology supports or undermines for its stakeholders, not only whether the technology works. Applying VSD to AI development remains difficult. Studies of responsible-AI practice show that unclear roles and organizational structures impede translating human values into practice [55, 58], while responsible-AI toolkits incorporate values through collaborative and educational features [44]. Responsible-AI and participatoryAI research continues to ask whose values a system represents [18, 27, 45] and how stakeholders should influence design [4, 15, 30, 49]. Because this evidence comes from design processes, practitioner intentions, and values elicited outside everyday use, it does not reveal which values end users experience as supported or undermined after an agent acts. Work extending VSD helps distinguish design intentions from end-user experience, recognizing that predefined value classifications can obscure local expressions of value, whereas engagement with lived experience supports value discovery [31]. Computational research offers a complementary route by evaluating value-relevant model behavior, encoding moral judgments about text scenarios [22], testing safety consistency [59], or linking diverse participant profiles to live LLM conversations [30]. However, its units of analysis remain elicited scenarios, test responses, and model conversations rather than end users’ accounts of agents in everyday settings. Neither route yields an empirical account linking a concrete agent aspect to the value a user experienced, to whether that value was supported or undermined, and to the outcome reported in naturally occurring use. Our study provides that account.

Coding Schema Development & Application: Agent Aspect, Human Value (Group), and User Outcome

Stage 1: Relevance screening for first-person experience. In Stage 1, we applied a screening-only prompt that classified each post as a first-person experience of using or attempting to use OpenClaw, a secondhand observation, or too thin to judge, keeping only the first-person posts (Appendix B defines the three categories). Stage 1 retained 187,479 posts, and Section 3.3 reports its validation. Stage 2: Value-centered coding. We call this stage value-centered because it anchored the coding on the human value. We first identified the primary human value at stake in an excerpt, then mapped the mentioned agent aspect and any resulting user outcome to that same excerpt. Every label was grounded in a single excerpt. Value fulfillment was coded as met when the excerpt described the value as supported, and not met when it described the value as undermined, threatened, or unavailable. A post was retained only if it passed both stages and its three required fields, the human value, the agent aspect, and the value fulfillment status, were all valid; the user outcome field was optional. Appendix B summarizes the coding instructions. We grounded the schema’s three coding dimensions in existing literature to give a standardized vocabulary across a large dataset. (1) Agent aspects. We defined 18 agent-aspect categories (Table 1 and Appendix B) identifying the system component or configurable boundary implicated in the value-centered excerpt. The categories drew on prior technical literature, such as decompositions of agent architecture [51, 57, 62] and automation [42], along with HCI guidance on human oversight and recovery [1], a GUIagent evaluation framework [11], and AI agent security risks [50]. 4 For simplicity, the rest of this paper uses “posts” to refer to both Reddit posts and

comments. 5 https://www.brandwatch.com 6 https://praw.readthedocs.io 3

Conference’17, July 2017, Washington, DC, USA

Ma et al.

We refined these definitions during codebook development to keep agent aspects distinct from human values and user outcomes. (2) Human values. We developed the value taxonomy through a deductive–inductive process [23]. We used 12 of the 13 values in Friedman et al.’s list [17] as sensitizing concepts, omitting courtesy because it concerns interpersonal politeness rather than how users delegate work to an agent. Inductively, we read a separate 50-post sample to identify human–agent interaction concerns absent from the deductive list [31]. We consolidated these into nine additional values, including meaningful human control, dependability, and contextual integrity, drawing from prior work [3, 24, 39, 46]. Coder review of a 400-post sample (Section 3.3) clarified definitions before full coding, and those review discussions acted as the kind of practice that surfaces values inside a team [48]. Appendix C defines all 21 values. (3) User outcomes. We defined nine user outcomes (Table 1) separately from the human values and agent aspects. We drew initial constructs from prior work on usability [26], user experience [20], workload [19], trust [32], and automation [41]. We then refined the boundaries between those constructs using user examples from the 400-post codebook-development sample, so that the categories captured the consequences users actually attributed to agent use. A value captured what mattered in the excerpt, whereas an outcome recorded the consequence the author reported. Using a large language model to apply a codebook at this scale trades coder time against a labeling error the researchers do not observe directly, so the procedure needs its own validation [47]. We applied GPT-5 mini with the Stage 2 prompt (Appendix B) to the retained posts, keeping 73,797 of them (Section 3.3 reports agreement with human coders). RQ1 analyzed the 73,093 of those posts assigned to one of the 21 predefined values; we excluded posts with open-coded values because they lacked a documented mapping to the taxonomy. RQ2 analyzed the 44,767 posts within that RQ1 sample that carried an attributable user outcome. A missing outcome indicates no attributable consequence was coded, not a neutral outcome. Appendix A summarizes sample construction. For group-level analysis, we used an affinity diagram to organize the 21 value definitions into six value groups, each covering a distinct facet of delegated use (Table 1). For example, Autonomous Operation clusters autonomy, human welfare, and calmness to capture a user’s goals and state during agent execution. Because this grouping was constructed for this analysis and not validated as a measurement model, we treat all group-level comparisons as exploratory summaries of cross-category patterns rather than tests of a latent construct.

3.3

not estimate label error across the full corpus. Six coders across three pairs reviewed the source text and initial LLM labels, entering decisions independently before discussion (an LLM-assisted, rather than blinded, procedure). Appendix E reports exact agreement, Cohen’s 𝜅 [2, 14], accuracy, macro-precision, macro-recall, macro-𝐹 1 , and weighted-𝐹 1 . Overall, human–human agreement was consistently high (exact agreement = 82.0%–98.8%; 𝜅 = .797–.979), whereas LLM–human correspondence showed high accuracy (.900– .994) but more variable macro-𝐹 1 (.604–.924) across coding targets.

3.4

Statistical and Qualitative Analysis

RQ1 analyses. We summarized the 21 values and six value groups, then cross-tabulated the six groups with the 18 agent aspects. Pearson’s 𝜒 2 summarized departure from independence and Cramér’s 𝑉 its magnitude, while observed-to-expected (O/E) ratios described single cells, each comparing a value group’s share at an agent aspect with that aspect’s share of the whole corpus. We applied the same two measures to value group and value fulfillment. A value group’s met rate reflected both its values and its agent aspects, because posts in different value groups raised different agent aspects, and because agent aspects differed in how often the values raised at them were met. We therefore used indirect standardization, comparing each value group’s observed met rate with the rate expected if its posts had been met at the corpus-wide rate for the agent aspects they raised. Appendix F describes the equations, the bootstrap procedure behind the intervals we report, and what those intervals leave out. These standardized differences are descriptive. RQ2 analyses. We cross-tabulated the six value groups with the nine user outcomes across the 44,767 outcome-coded posts, and applied the same two measures to value fulfillment and user outcome. Every expected count exceeded five in the tables we report (Appendix F). Each cell of the 6 × 9 table reports its post count and its met rate, the share of that cell’s posts whose value was coded met. We displayed, but did not interpret, cells holding fewer than 20 posts, because in a cell that small a single post moves the met rate by more than five percentage points. Thematic analysis of value-centered excerpts. The crosstabulations above report which values, agent aspects, and user outcomes co-occurred, but not how users described them, so we returned to the verbatim excerpt that grounded each coded post. We read these excerpts through an inductive thematic analysis [6], where codes came from the excerpts rather than from a prior list, while reading stayed inside the dimensions the Stage 2 coding had already established. We treated each value group’s two most frequent user outcomes as the unit of reading, because those twelve cells hold 69.9% of the outcome-coded posts, and read a stratified random sample of 20 excerpts per cell, 240 excerpts in total. One researcher labeled what each excerpt claimed, grouped the labels into sub-themes, and grouped the sub-themes into themes, discussing the developing codebook with the research team throughout and revising labels and boundaries after each discussion. For example, the researcher labeled separately the posts reporting that a run ended without acting, that a fallback model answered without raising an error, and that finished-looking output had never been tested. Those labels were then collected into the sub-theme silent or partial failure and paired with verification taken back by the user to form

Human Validation

Stage 1 validation. We evaluated the relevance screen on a 50-post relevance pilot, distinct from the inductive sample in Section 3.2, against human consensus labels, achieving a precision of .900, recall of .783, and 𝐹 1 = .837. Because these posts were not randomly sampled from the full corpus, these estimates do not establish fullcorpus recall. Stage 2 validation. We evaluated the Stage 2 codebook and initial LLM labels on the same 400-post sample used to develop the codebook (201 original posts, 199 comments), so these metrics do 4

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

Table 1: The coding schema. The table lists six value groups constructed for this analysis, the 18 agent aspects in corpusfrequency order, and the nine user outcomes. Appendix B gives the operational definitions and boundary rules used in coding; Appendix C adds a definition and an example excerpt for each of the 21 values. Value group

Definition

Human values (Friedman et al.’s list [17]; † study-specific, developed for this corpus following [31])

Dependable Operation Autonomous Operation Affordable Operation

Whether the agent’s service holds and can be repaired The user’s own goals and state while the agent acts for them The money and resources a run consumes

Bounded Reach

What the agent may access and who controls that access

Reviewability

Whether users can see, approve, and answer for what the agent does

Equitable Access

Who can use the agent at all

dependability†, trust, repairability† autonomy, human welfare, calmness affordability†, resource stewardship†, environmental sustainability security†, privacy, identity, property ownership, data sovereignty†, contextual integrity† transparency†, meaningful human control†, accountability, informed consent universal usability, freedom from bias

Agent aspect (adapted from [1, 11, 42, 50, 51, 57, 62])

Definition

Agent aspect

Definition

Model core Resource accounting System access

Multi-agent orchestration Observability Human oversight

Coordination among agents and delegated workers Traces, logs, and visibility into what the agent did Approval gates and points where a user intervenes

Environment access

The underlying LLM and how it is configured What a run consumes and what it is billed for Getting the agent installed, authenticated, and ready to use What the agent may reach once access is configured

Error handling

Tool execution

Invoking tools, commands, and external operations

Other

Action effects

User-visible changes the agent makes to files and data The goal, scope, and stopping rules set for a run Context, retrieval, and what the agent keeps or forgets Latency, responsiveness, and stability under load

Tool selection

Detecting, reporting, and recovering from failed actions A substantive agent-design concern outside the taxonomy Which tool or command the agent chooses

Task specification Memory Runtime performance

Planning Agent profile Reasoning

User outcome (adapted from [19, 20, 26, 32, 41])

Definition

Task effectiveness Time efficiency Resource burden Supervision workload Affective response Trust calibration Adoption behavior Risk exposure Recovery behavior

Whether the agent completed the task and how correct or usable the result was Whether the agent saved the user time or cost them time What a run cost the user in money and metered resources The effort of watching, checking, and managing what the agent did How the user felt about the experience How far the user relied on the agent, and whether that reliance was warranted Whether the user continued with the agent, intended to, or moved away from it Harm to the user’s data, systems, or compliance position that a post reported or clearly anticipated What the user did to repair or work around what the agent did

the theme finishing a run stopped meaning the work was done. The procedure yielded 63 initial codes, 24 sub-themes, and 12 themes, two per value group, which Sections 5.1–5.6 report as the bolded claims. Appendix D gives the codebook with a representative excerpt for each sub-theme. Because this step interpreted posts that the pipeline had already coded, it changed no count reported in this paper, and we report no interrater statistic for it. Unit of analysis. The user post, not the user, was our unit of analysis. Because the dataset may contain multiple posts from a single user, our association statistics and bootstrap intervals do not adjust for within-user dependence and should be read as corpusspecific rather than population-level estimates. Every value and outcome we record also comes from what a user wrote about their own use rather than from an independent measure of the agent’s behavior.

3.5

Decomposing a task and forming or revising a plan The system prompt, persona, and framing instructions Reasoning the agent shows or the user reports

publicly released research materials further excluded the verbatim source text. Use of AI tools. We used a large language model as a coding instrument, as Section 3.2 describes, and AI-based assistants for copy-editing and for checking references while preparing this manuscript. Following the ACM Policy on Authorship,7 no generative AI tool is listed as an author, every passage such a tool touched was reviewed by the authors, and the authors take full responsibility for the entire text.

4

RQ1: What Human Values Surface in First-Person Posts of OpenClaw Use, and Where Do They Attach?

The RQ1 sample contained 73,093 posts. We first reported the distribution of human values (Section 4.1), then compared the agentaspect distributions of the six value groups (Section 4.2). Finally, we examined value fulfillment by comparing each group’s observed met rate, the actual percentage of posts where the value was coded as successfully supported, with the rate expected given its distribution of agent aspects (Section 4.3).

Ethics

Research ethics. Our university’s institutional review board determined the study exempt because it analyzed publicly available posts without interacting with their authors. We reported aggregate patterns from posts and quoted only short excerpts reviewed for identifying details. To reduce reidentification risk, we removed usernames, links, timestamps, and community identifiers, and the

7 https://www.acm.org/publications/policy-on-authorship

5

Conference’17, July 2017, Washington, DC, USA

4.1

Ma et al.

The Most Common Human Values Across Value Groups Were Autonomy, Dependability, Affordability, Resource Stewardship, and Universal Usability

(O/E 1.5 and 1.2). The two value groups separated below it. Dependable Operation reached O/E ratios of 3.7 for error handling and 2.0 for runtime performance, whereas Autonomous Operation reached 2.3 for action effects and 1.9 for both tool execution and task specification. The agent aspect a value group discussed most was therefore not always the agent aspect that distinguished it. These ratios describe cells within the overall association, not separate cellwise tests.

Five of the 21 values accounted for 51,784 of the 73,093 posts (70.8%). Figure 1 shows all 21 values within their six groups. Specifically, autonomy was the most frequent value (19.9%), followed by dependability (19.8%), affordability (14.0%), resource stewardship (8.9%), and universal usability (8.3%). Four of the remaining 16 values appeared in fewer than 250 posts each. We therefore reported the full value distribution but used the six value groups for cross-category comparisons, which would otherwise include value-level cells with few posts. The three largest value groups, Dependable Operation, Autonomous Operation, and Affordable Operation, contained 52,542 posts (71.9%); Bounded Reach, Reviewability, and Equitable Access contained the remaining 20,551 posts. User posts thus focused more on whether the OpenClaw agent worked, what it could do on its own, and what it cost, and less often on what it could reach, whether users could review it, or who could use it. By provenance, the 12 VSD-derived values accounted for 29,260 posts (40.0%) and the nine study-specific values for 43,833 (60.0%).

4.2

4.3

Value Fulfillment Rate Varied Across Value Groups, but Only Autonomous and Affordable Operation Exceeded Expected Rates

Across the RQ1 sample, the value at stake was coded met in 39,878 of 73,093 posts (54.6%). Met rates ranged from 77.3% for Autonomous Operation to 42.8% for Equitable Access, with 53.9% for Reviewability, 47.9% for Affordable Operation, 46.9% for Dependable Operation, and 44.4% for Bounded Reach (Figure 3(a)). Also, value group and fulfillment were associated at the post level, 𝜒 2 (5, 𝑁 = 73,093) = 5,056.4, with Cramér’s 𝑉 = .263. Dependable Operation and Affordable Operation together accounted for 18,374 of the 33,215 not-met posts (55.3%). A high group met rate also did not mean uniform fulfillment inside a value group, since autonomy, the most frequent value, still had 2,208 of its 14,554 posts (15.2%) coded not met, 6.6% of all not-met posts. Across the 18 agent aspects, met rates ranged from 81.3% for planning (401 of 493 posts) to 35.3% for resource accounting (3,347 of 9,483). Figure 3(b) shows the eight agent aspects with the most posts. Value fulfillment thus varied more widely across agent aspects than across value groups, so part of the spread between value groups could reflect which agent aspects their posts raised (Section 4.2) rather than the values themselves. We therefore applied the indirect standardization described in Section 3.4 (Figure 4). Observed minus expected was positive for two value groups, +17.0 pp for Autonomous Operation (conditional 95% percentile bootstrap interval [+16.4, +17.6]) and +3.5 for Affordable Operation ([+2.7, +4.2]). It was negative for the other four, −6.5 for Reviewability ([−7.7, −5.3]), −8.8 for Equitable Access ([−10.0, −7.6]), −9.8 for Dependable Operation ([−10.6, −9.1]), and −10.8 for Bounded Reach ([−11.9, −9.7]). These intervals resample posts within each value group with corpus-wide agent-aspect-specific met rates held fixed, and describe the coded corpus rather than causal or user-level effects. A positive difference means a value group met its values more often than the agent aspects it discussed would predict. Affordable Operation ranked third on observed met rate yet exceeded its expectation, because 54.4% of its posts concerned resource accounting, the least-met agent aspect in the corpus at 35.3%. Affordable Operation’s shortfall therefore lay in the agent aspect its posts raised rather than in the values at stake. Reviewability moved the other way, its second-highest observed met rate falling short of an expectation set by two agent aspects met more often than the corpus-wide 54.6%, observability at 56.4% and human oversight at 67.7%. Autonomous Operation’s advantage survived the adjustment, at 77.3% observed against 60.3% expected, so it did not follow from the agent aspects its posts raised.

Four Value Groups Were Each Associated with One or Two Agent Aspects, Whereas the Two Largest Shared One Aspect, the Model Core

Value group and agent aspect were associated at the post level, 𝜒 2 (85, 𝑁 = 73,093) = 94,536.8, with Cramér’s 𝑉 = .509. Figure 2 reports each value group’s agent-aspect distribution beside the corresponding corpus-wide share. Three value groups each concentrated on a single agent aspect. Resource accounting comprised 54.4% of Affordable Operation posts against 13.0% corpus-wide (O/E 4.2), environment access 46.5% of Bounded Reach posts against 8.2% (O/E 5.7), and system access 44.7% of Equitable Access posts against 11.6% (O/E 3.9). Read the other way, those three value groups differ sharply. Affordable Operation held 96.1% of all resource-accounting posts, Bounded Reach 55.8% of environment-access posts, and Equitable Access only 32.8% of system-access posts. Cost was therefore discussed almost exclusively within Affordable Operation, whereas user setup and authentication arose across groups rather than belonging to Equitable Access. Reviewability rested on two agent aspects rather than one, observability at 25.4% of its posts (O/E 8.6) and human oversight at 24.6% (O/E 9.2). Reviewability held 83.9% of all observability posts and 89.9% of all human-oversight posts. Seeing what the OpenClaw agent had done and being asked to approve it therefore surfaced together, and almost only where users raised Reviewability. The two largest value groups shared an agent aspect. Both discussed the model core most often, at 39.6% of Dependable Operation posts and 32.8% of Autonomous Operation posts. The model core was also the most common agent aspect corpus-wide at 26.4%, so neither value group was more than modestly overrepresented there 6

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA 18,169 (24.9%)

Dependable Operation

14,483 (19.8%)

dependability trust repairability

2,101 (2.9%) 1,585 (2.2%)

17,614 (24.1%)

Autonomous Operation

14,554 (19.9%)

autonomy human welfare calmness

1,989 (2.7%) 1,071 (1.5%)

16,759 (22.9%)

Affordable Operation

10,220 (14.0%)

affordability resource stewardship environmental sustainability 63 (0.1%)

Bounded Reach

security privacy identity property ownership data sovereignty contextual integrity

6,476 (8.9%)

7,199 (9.8%) 3,841 (5.3%) 1,244 (1.7%) 1,145 (1.6%) 455 (0.6%) 285 (0.4%) 229 (0.3%)

7,142 (9.8%)

Reviewability

transparency meaningful human control accountability informed consent

3,411 (4.7%) 3,303 (4.5%) 365 (0.5%) 63 (0.1%)

6,210 (8.5%)

Equitable Access

universal usability freedom from bias

6,051 (8.3%) 159 (0.2%)

Figure 1: Distribution of values and their six value groups in the RQ1 sample (𝑁 = 73,093). Solid bars show the number of posts of each value. The shaded group bars show the number of posts in each value group; the six group counts therefore sum to the RQ1 sample. n

tio

ra

agent aspect

All

ts os rp

e

us

pe eO

l ab

d en

p

De

m no

to

Au

ion

at

r pe

sO ou

rd

fo

Af

le ab

n

tio

ra

e Op

Bo

ch

ea

dR

de un

ity bil

a iew

v

Re

le

b ita

ss ce

Ac

u

Eq

model core

26.4

39.6

32.8

19.4

13.7

10.0

21.7

resource accounting

13.0

0.7

0.4

54.4

0.1

1.4

1.2

system access

11.6

9.7

6.8

10.3

8.5

5.5

44.7

environment access

8.2

3.9

5.2

2.1

46.5

3.6

6.8

tool execution

5.9

7.3

11.2

1.2

5.8

2.6

3.5

action effects

5.5

4.3

12.8

1.0

5.9

3.3

2.4

task specification

5.2

3.7

10.0

1.2

3.5

10.7

2.9

memory

4.2

6.7

3.7

2.5

6.1

3.4

1.9

runtime performance

3.6

7.3

2.9

2.3

0.1

0.2

6.7

multi agent orchestration

3.6

4.1

5.9

2.2

1.5

3.7

1.6

observability

3.0

0.7

0.3

0.2

1.3

25.4

0.3

human oversight

2.7

0.3

0.5

0.0

0.6

24.6

0.1

error handling

2.2

7.9

0.2

0.2

0.3

0.5

0.2

other

1.6

1.0

2.2

0.3

3.3

1.3

4.0

tool selection

1.6

1.1

2.2

2.2

0.7

0.9

1.4

planning

0.7

0.5

1.8

0.1

0.1

0.6

0.2

agent profile

0.7

0.6

0.6

0.2

1.8

1.2

0.4

reasoning

0.4

0.6

0.5

0.1

0.1

1.2

0.0

larger share inside the group than in the corpus smaller share inside the group than in the corpus

Figure 2: Distribution of agent aspects across the six value groups in the RQ1 sample (𝑁 = 73,093). Each cell shows the percentage of posts in a given value group coded for that agent aspect, and each group column sums to 100%. The first column gives each aspect’s share of the whole corpus; orange marks a larger share inside the group than in the corpus (O/E ratio above 1) and blue a smaller share, and Appendix F defines the ratio.

5

RQ2: What Outcomes Do Users Attribute to Agent Use Across Value Groups?

the six value-group summaries comparable, each subsection below reports the two most frequent user outcomes and describes how users characterized them. A closing subsection then reads the same posts across value groups, describing how often values were met in each user outcome and user outcomes when a value was met and not (Section 5.7).

The RQ2 subset contained 44,767 of the 73,093 posts (61.2%) that carried an attributable user outcome. Among the nine user-outcome categories (Table 1), task effectiveness was the most frequent (18,646 posts, 41.7%), followed by resource burden (6,955, 15.5%) and time efficiency (5,332, 11.9%). Value group and user outcome were associated at the post level, 𝜒 2 (40, 𝑁 = 44,767) = 41,092.5, with Cramér’s 𝑉 = .428. Figure 5 reports that 6 × 9 cross-tabulation. To make 7

Conference’17, July 2017, Washington, DC, USA

Ma et al. (a) By value group user posts

54.6%

All 73,093 user posts Autonomous Operation Reviewability Affordable Operation Dependable Operation Bounded Reach Equitable Access

45.4% 77.3%

22.7%

53.9% 47.9% 46.9% 44.4% 42.8%

46.1% 52.1% 53.1% 55.6% 57.2%

73,093 17,614 7,142 16,759 18,169 7,199 6,210

(b) By agent aspect (the eight aspects carrying the most user posts) user posts

All 73,093 user posts action effects tool execution task specification memory model core environment access system access resource accounting

54.6%

45.4%

72.4% 69.0% 63.1% 60.3% 55.8% 51.2% 47.0% 35.3%

27.6% 31.0% 36.9% 39.7% 44.2% 48.8% 53.0% 64.7%

value coded met

73,093 4,029 4,324 3,818 3,082 19,278 6,002 8,464 9,483

value coded not met

Figure 3: Value fulfillment in the RQ1 sample (𝑁 = 73,093). Panel (a) reports met and not-met shares for the six value groups. Panel (b) reports the same shares for the eight most frequent agent aspects. expected

47.9%

44.5%

53.9%

Reviewability Equitable Access 42.8% Dependable Operation

77.3%

60.3%

Autonomous Operation Affordable Operation

observed

51.7%

46.9%

56.7%

Bounded Reach 44.4% 40%

60.4%

55.2% 50%

60%

70%

observed − expected

95% CI

+17.0 pp

[+16.4, +17.6]

+3.5 pp

[+2.7, +4.2]

−6.5 pp

[−7.7, −5.3]

−8.8 pp

[−10.0, −7.6]

−9.8 pp

[−10.6, −9.1]

−10.8 pp

[−11.9, −9.7]

80%

user posts coded met (%)

Figure 4: Observed value fulfillment and fulfillment expected from the distribution of agent aspects in each group in the RQ1 sample (𝑁 = 73,093). Differences are observed minus expected percentage points; intervals are conditional 95% percentile bootstrap intervals from 2,000 post-level resamples within each value group.

5.1

Autonomous Operation: Reach Users Widened Rather Than Took Back

decreasing by the day. What’s the point of everything if no one uses it? Building is easy now, but who will purchase, given that building is easy? How will supply and demand match?” This user questioned what their own effort was worth once task execution became easy.

Task effectiveness was the most frequent user outcome for Autonomous Operation (5,285 of the 11,754 outcome-coded posts in this group, 45.0%), followed by time efficiency (2,823 posts, 24.0%). Users kept widening what the agent could touch. One user with no coding knowledge asked how to make OpenClaw do what they had seen others do, describing “letting it open programs, searching the web, actually doing work autonomously.” Another described what they were adding next: “I’m working on getting mine a phone number, and access to a debit card with access to capital.” In both cases, users widened what OpenClaw could reach on their devices and accounts rather than restricting it. Users got time back, not better output. When users said what they had gained from using OpenClaw, they often described being freed from doing the work themselves rather than getting a better result. One user wrote, “Giving OpenClaw a pair of hands has freed up my hands!” Time efficiency here recorded the hours the user no longer had to spend, not a task finished faster. Another user set the promise of returned time against the wider situation it arrived in: “OpenClaw was supposed to give us more time, but it feels like time is

5.2

Dependable Operation: Delivery That Did Not Hold Across Runs

Dependable Operation’s posts led with task effectiveness (8,993 of its 14,256 outcome-coded posts, 63.1%), followed by adoption behavior (1,249 posts, 8.8%). The remaining user outcomes were sparse. Dependable Operation posts stood apart because users described the same OpenClaw setup as delivering at one time and failing to keep delivering at another, which left them asking whether the agent would carry out a task again. Finishing a run stopped meaning the work was done. Users expected an instruction to become an OpenClaw action, and posts about failure described the agent stopping short of that step rather than producing a wrong result. One user described an agent that would not act on a plan it had itself proposed: “No matter how many times I approve, it never actually transitions from ‘planning’ to ‘doing.’ ” Failure was also not always visible at the moment it occurred, 8

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

since free models, as one user put it, “don’t always fail loudly” and would “quietly ship you a stub and move on.” Because a finished run and an empty run looked alike, task completion stopped serving as evidence that the work had been done, and users took verification back. One user, describing repeated breakage, wrote: “I kept breaking them in production almost daily because I treated OpenClaw changes like scratch work instead of production deployments.” This user treated OpenClaw configuration as a deployment problem rather than a settings change, and keeping that configuration working was itself infrastructure work, the kind of task OpenClaw was meant to remove. Where a setup did hold, the posts were correspondingly plain, as one user who had helped others debug their OpenClaw configurations observed: “the setups that survive are the ones where the person can explain what their agent does in one sentence.” Users kept the agent but asked less of it. Repeated failure rarely ended agent use. One user planned to “uninstall Qwen 9.5 9b today and give Gemma a shot,” and to “switch to Ollama cloud” if the replacement also fell short. Another judged that “the openclaw product became undesirable at least compared to other platform options” and chose to downgrade and wait until it improved, rather than to stop. After a failure, these users changed the model, the host, or the task’s ambition, not the decision to delegate. Adoption behavior in this value group therefore recorded lowered expectations; the OpenClaw agent stayed, but users stopped counting on it to behave the same way twice.

5.3

like a running cost, visible after the fact but not available to plan against. Users moved the spending rather than stopping it. Posts that described cost as settled rarely described a cheaper agent. They described moving the spending somewhere the user could predict it. One user compared plans directly and got “more usage out of GitHub Copilot Pro+ for $40/month using Claude models than I do from 5x Claude Team accounts for $100.” Others changed the agent rather than the plan, adopting a minimal agent that consumed fewer tokens under default prompts, which one user reported had become “the only coding agent I use now.” Where the accounting did settle, users measured the agent against the tools it displaced. One described applications an agent had built that “I previously would have had to pay anywhere from 20 to 200 dollars a month for.” Affordability in this value group was accordingly not a property of the agent but an arrangement users assembled around it, and that arrangement lasted only while they kept adjusting it.

5.4

Bounded Reach: Risk and Usefulness Both Settled at OpenClaw Setup

Risk exposure was the most frequent user outcome for Bounded Reach (719 of the 2,425 outcome-coded posts in this group, 29.6%), followed by task effectiveness (585 posts, 24.1%). Users described the two in different terms. The risk users described was access, not damage. Posts with this user outcome described the access the user had granted rather than harm the user had suffered. That access was a property of an OpenClaw configuration, meaning which files, credentials, and networks the agent could reach, and users fixed it when they installed the agent rather than during a task. One user weighing a business deployment was “afraid to connect it to my company’s data and tools.” Another listed “whether it can touch anything outside the workspace indirectly.” Users understood the permission model as an instruction set, and as one put it, the configuration was “a set of instructions your agent follows with full permissions.” Another user described a default that shipped open: “OpenClaw has had 14 CVEs [Common Vulnerabilities and Exposures] since launch, 8 of them critical. The default config binds the gateway to 0.0.0.0, which means anyone who finds your IP has full access to your OpenClaw admin. You need to bind to localhost, set up sandbox mode, configure tool deny lists, and add systemd isolation.” This user separated the discovered defects from a setting that shipped open and listed four things the operator had to add. The gateway was reachable not because a run went wrong but because that was how the software shipped, so the risk lay in how OpenClaw was set up rather than in what a run produced. Useful runs happened inside limits set beforehand. The posts described not an agent given more room but a setup whose reach the user had already set in advance. One user reported a placement, having “launched this [OpenClaw] on LightSail to avoid it having access to my files.” Another stated a data boundary, “I’ll use my local inference models for internal stuff only.” Users described tasks that went well after such configuration. One noticed that “it keeps memory separated by project” and found this “helpful when organizing a lot of scattered notes and materials.”

Affordable Operation: Cost That Accrued Where Users Could Not See It

Resource burden was the most frequent user outcome for Affordable Operation (6,776 of the 10,549 outcome-coded posts, 64.2%), followed by task effectiveness (1,359 posts, 12.9%). These two user outcomes did not offset each other. A run that completed a task could still generate a bill, and the bill arrived without explaining which agent actions had cost money, so posts in this value group registered OpenClaw’s usefulness and its resource consumption as separate primary outcomes. The bill came from work users never asked for. What these posts objected to was rarely that a price was too high, but that consumption happened where the user was not looking. One user found they had been “paying Sonnet rates for 48 heartbeats a day to silently check if anything was scheduled,” a charge produced by the agent’s readiness rather than by any task they had asked for. Another traced the same problem to what each message loaded, proposing that “if the user says ‘hello’, don’t inject the full system prompt and skill definitions.” A third user added: “OpenClaw is not being proactive in monitoring token usage and preventing input tokens from reaching higher than 100K per message, even when I’m only writing ‘test’ and with things in place to alert at 40K input tokens to compact and reduce back down to sub-20K.” This user had already built the accounting that the agent lacked, including an alert threshold and scheduled sessions to compact context, and still could not keep consumption down. What they described was a spending decision made inside a run rather than at its boundary. Cost in this value group, therefore, behaved less like a price than 9

Conference’17, July 2017, Washington, DC, USA

5.5

Ma et al.

Reviewability: Checkpoints That Worked at a Run’s Edges but Not Inside It

Users weighed the install before they weighed the agent. Posts about adoption behavior often recorded a decision made before or during setup rather than after extended use. One user packaging a standalone build reasoned that “not everyone is comfortable using command line or Docker on Linux,” and another asked, before installing anything, “Openclaw for non-techie.” Completing the install did not settle the question, as one user who had finished it called the process “still annoying as hell to setup.” A third user added: “Accessibility is what made OpenClaw popular. Not the accessibility of OpenClaw itself, but what it allows non/moderately technical people to achieve, and what real-world problems they could solve with it.” This user separated reaching the agent from what the agent then made achievable, and credited OpenClaw only with the second.

Task effectiveness was the most frequent user outcome for Reviewability (982 of the 2,590 outcome-coded posts in this group, 37.9%), followed by affective response (492 posts, 19.0%). Checking before or after a run cost nothing. Where users described OpenClaw’s work as useful, they had set the point of control before the run or reviewed the outcome after it. One user who had delegated guest-post outreach reported that the agent “didn’t just say yes to everything” and “pushed back” on a topic outside their expertise. That refusal came from instructions the user had given in advance, not from a question the user answered during the run. Another user kept one standing gate on the outbound run, “nothing sends without a human,” and reduced their involvement to that single decision. In these posts, what users counted as success was the outcome itself, such as the outreach that closed. Reviewing the agent asked little of these users because none of it happened while the run was executing. The cost was in answering, not in looking. What users objected to was not that the agent asked for approval, but that it asked about a step they had already approved or asked somewhere they could not answer. One user asked, “How can I make it so that my Mac Mini doesn’t continue to request permissions to do things I am asking it to do? It continues to interrupt the workflow, and I feel like I haven’t set things up correctly.” This user had authorized the steps, and the agent asked again. They read OpenClaw’s repeated request not as oversight but as evidence that their own setup was wrong. Another user, running OpenClaw on a headless node, reported that “the approval dialogs never actually appear,” and kept a browser window open to authorize each step by hand. The gate existed, but it was positioned out of reach. A third user, after switching to a different frontend, could “actually see the tool calls and reasoning steps in a clean workspace” and called the view “a huge sanity saver.” Such a view let the user monitor the agent without interruption. What separated a workable checkpoint from an interruption in this value group was not whether users could watch the agent, but whether watching required an answer during the run.

5.6

5.7

Values Were Met More Often Where Users Described Delivery Than Where They Described Cost

Among the 44,767 outcome-coded posts, the value was met in 25,172 (56.2%) and not met in 19,595 (43.8%). Figure 5 breaks that split down by value group and user outcome, giving the post count and the met rate in each of its 54 cells. These cell rates describe the outcomecoded subset and are therefore not directly comparable with the group met rates in Section 4.3, which cover all 73,093 coded posts. The share of a value group’s posts carrying an attributable user outcome also varied, from 33.7% in Bounded Reach to 78.5% in Dependable Operation, so the value groups are not equally represented in this subset. Whether a value held depended more on the user outcome a post reported than on its value group. The value was met in 67.7% of the 18,646 task-effectiveness posts and in 76.1% of the 5,332 time-efficiency posts. The value was met in only 34.7% of the 6,955 resource-burden posts, 29.4% of the 1,914 supervision-workload posts, and 10.7% of the 1,026 risk-exposure posts. Task effectiveness was majority-met in five of the six value groups, from 93.0% in Autonomous Operation down to 53.6% in Dependable Operation; only in Equitable Access, at 48.4%, did its posts fall close to even without clearing half. Supervision workload ran the other way in all six, from 46.9% met in Autonomous Operation down to 10.3% in Equitable Access. Risk exposure ran the same way but fell below the 20-post display threshold in two value groups, so supervision workload was the only majority-not-met user outcome whose cells cleared that threshold in every value group. Three value groups paired one majority-met and one majoritynot-met leading user outcome. Bounded Reach showed the widest gap, with the value met in 76.4% of its task-effectiveness and in 11.4% of its risk-exposure posts. Reviewability followed, with task effectiveness at 72.6% and affective response at 30.1%, then Affordable Operation, with task effectiveness at 75.1% and resource burden at 34.7%. Each post carried one user outcome, so these opposite directions came from different sets of posts within the same value group rather than from single users. In the other three value groups, the two leading user outcomes sat close together, high for Autonomous Operation (93.0% and 93.6%) and nearly even for Dependable Operation (53.6% and 48.0%) and Equitable Access (48.4% and 45.6%). Particular user outcomes also concentrated in particular

Equitable Access: Useful Work Reported From the Far Side of Setup

Task effectiveness (1,442 of the 3,193 outcome-coded posts, 45.2%) and adoption behavior (585 posts, 18.3%) were the two most frequent user outcomes for Equitable Access. These posts came from two kinds of users, those who had already reached a working OpenClaw setup and those still describing what it cost to reach one. Posts of useful work came from users already past setup. One user had cleared setup and moved into a threaded chat client, reporting that “I can chat with my agent in different threads at the same time, and everything stays nice and neat.” When users explained how they reached a working setup, they credited help that someone else had prepared. One user switching providers followed a documented onboarding command and reported that “the change was pretty easy.” Another user, running OpenClaw on a hosting provider, credited the provider rather than the agent, writing that “Contabo has made it really easy to get something like OpenClaw up, I’ve started messing with it.” 10

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

value groups, with Affordable Operation holding 6,776 of the 6,955 resource-burden posts (97.4%) and Bounded Reach holding 719 of the 1,026 risk-exposure posts (70.1%). Setting the value groups aside and cross-tabulating value fulfillment against user outcome shows which user outcomes users reported on each side of fulfillment, 𝜒 2 (8, 𝑁 = 44,767) = 5,033.9, with Cramér’s 𝑉 = .335 (Table 2). Among met posts, task effectiveness and time efficiency together accounted for 66.2%, and both name what a run gave back rather than what it cost. Among notmet posts those two user outcomes accounted for 37.3%. Four user outcomes ran the other way, with resource burden, affective response, supervision workload, and risk exposure accounting for 46.9% of not-met posts against 19.5% of met posts, and each of those four recorded what running the agent cost. Adoption behavior did not follow that contrast, standing at 2,802 met and 2,171 not-met posts (11.1% of each), though the six value groups differed on it. These were post-level associations among one coded label per field and did not establish causal effects, estimates for unique users, or co-occurring user outcomes within a post.

6

and a condition is what the user can set at that place. Most agent aspects carry no condition at all, because the agent exercises them independently. A user can choose which model runs, but not how it reasons, and can set a stopping rule, but not the plan the agent forms to meet it. Where a condition does exist, its timing varies, since some are settled at install while others can be reset before each later run. We name this pattern value-sensitive delegation, the relationship between a human value and the operating condition that carries it, together with the moment at which a user can still change that condition. Two claims follow. First, in four of the six value groups the value attached to those conditions rather than to the finished work (Section 4.2). Second, whether a user could protect a human value depended on when its condition could be set. In four value groups, the condition can be fixed before a run starts. In Affordable Operation and Reviewability, by contrast, the condition moves while the run is underway, and users in those two groups reported values at the conditions they had fixed in advance but reported costs at the conditions that moved (Sections 5.3 and 5.5). The relationship runs in both directions, since the conditions a user sets bound what the agent may do, and what the agent does then sends the user back to adjust those conditions for the next run (Step 4 in Figure 6). In five of the six value groups, the most discussed agent aspect and the condition users acted on align. However, in Autonomous Operation, users most often discussed the model core, yet they acted on the agent’s reach (the files, credentials, and services the agent was permitted to touch). What these users discussed and what they could change therefore came apart (Sections 4.2 and 5.1). AI engineering research decomposes agents into the same modules of memory, planning, and tool actions [51, 57, 62], while security research locates risks across architectural layers [52, 66]. Both vocabularies identify a property or risk inside the agent. Neither indicates which condition a user would change or when that change could occur, which is what Figure 6 provides. Because these conditions can be reset between runs, users adjust their delegation strategy without abandoning OpenClaw. Prior work examines user reliance on automation [41], and when such reliance is appropriate [32]. Recent agent research maps the design space for user control [12] and balances autonomy with human oversight [38]. These accounts do not record what a user changed before the next run. Following failures, Dependable Operation users swapped models or hosts rather than discarding OpenClaw, while Autonomous Operation users widened the agent’s reach even without failures (Sections 5.2 and 5.1). Value-sensitive delegation thus refines these experiences of reliance [32, 41] by separating continued agent use from the conditions users changed before the next run. The Reviewability value group highlights this timing difference by mixing pre-run checks with mid-run interruptions. Prior work studies how user involvement impacts trust, performance, and cognitive load [7, 21, 53, 68], and distinguishes interventions that halt an agent from those that steer it [29]. However, our findings separate three mechanisms often collapsed into “human oversight”: the gate, the record, and the request. A gate is a boundary set once that holds without active watching. A record supports verification after the run. A request, unlike either, requires the user inside a run that is still going. While requests support human judgment,

Discussion

While prior work evaluates agents by task completion [28, 34, 64] or scores value-relevant judgments [30, 59], our study is among the first to trace, using OpenClaw as an exemplar agent, which human values users invoke in everyday agent use and which agent aspects those values attach to. Below, we describe this relationship as value-sensitive delegation and ground it in VSD’s interactional position (Section 6.1), examine the gains and costs users attributed to running OpenClaw (Section 6.2), and translate both into design implications (Section 6.3).

6.1

Value-Sensitive Delegation: Locating Values in the Conditions of Agent Use

Value Sensitive Design (VSD) treats human values as interactional, meaning that a value takes shape in the relationship between a technology’s properties, the people affected by it, and the context where it is used [17]. A delegated AI agent like OpenClaw stretches that relationship over time, because the user configures the agent, the task runs autonomously, and the user checks the outcomes afterward. Section 4.2 reports that in four of the six value groups, user posts concentrated on agent aspects the user configures rather than on agent aspects the agent exercises without the user. These included what a run could spend (Affordable Operation), what it could reach (Bounded Reach), when it had to ask (Reviewability), and what had to be installed (Equitable Access). While Dependable Operation and Autonomous Operation frequently discussed the model core, the model core was the most discussed agent aspect corpus-wide and did not uniquely set those two value groups apart. We define an operating condition, or condition for short, as the configurable boundary a user establishes around an agent’s run, together with the time point at which they can act on it (Figure 6). Reading the six value groups together, we identify five conditions in these posts, namely the model and host, the resource arrangement, the granted reach, the approval gate and record, and the initial install (Sections 5.1–5.6). A condition is not a further coded category. An agent aspect names the place where a value is at stake, 11

Primary value group

Conference’17, July 2017, Washington, DC, USA

Ma et al.

Dependable Operation (n=14,256)

n=8,993 54% met

n=73 29% met

n=1,025 49% met

n=1,249 48% met

n=660 33% met

n=747 12% met

n=816 61% met

n=135 10% met

n=558 29% met

Autonomous Operation (n=11,754)

n=5,285 93% met

n=24 50% met

n=2,823 94% met

n=1,218 79% met

n=1,804 53% met

n=514 47% met

n=26 50% met

n=47 2% met

n=13* 46% met

Affordable Operation (n=10,549)

n=1,359 75% met

n=6,776 35% met

n=614 68% met

n=1,348 55% met

n=328 52% met

n=75 20% met

n=39 36% met

n=5* 40% met

n=5* 40% met

Bounded Reach (n=2,425)

n=585 76% met

n=25 56% met

n=87 69% met

n=374 38% met

n=498 39% met

n=69 32% met

n=36 31% met

n=719 11% met

n=32 47% met

Reviewability (n=2,590)

n=982 73% met

n=38 32% met

n=230 77% met

n=199 47% met

n=492 30% met

n=392 46% met

n=75 53% met

n=111 11% met

n=71 41% met

Equitable Access (n=3,193)

n=1,442 48% met

n=19* 5% met

n=553 46% met

n=585 46% met

n=422 31% met

n=117 10% met

n=45 7% met

n=9* 0% met

n=1* 0% met

Task Resource effectiveness burden

Time efficiency

Adoption behavior

Affective response

Supervision workload

Recovery behavior

Risk exposure

Trust calibration

Primary user outcome

0

20

40

60

80

100

Share of the cell's posts whose primary value was coded met (%)

Figure 5: Value group by user outcome among the 44,767 outcome-coded posts. Each cell describes the post count and the met rate, the share of its posts whose human value was coded as met. Blue is majority met, red is majority not met, and cells at 50% are unshaded. Asterisks mark cells with fewer than 20 posts, shown but not interpreted. Table 2: User outcomes among the 44,767 outcome-coded posts, split by whether the post’s value was coded met. Rows sum to 100% before rounding, and the final row gives the met-minus-not-met difference in percentage points.

Value met Value not met Difference (pp)

1

Task eff.

Resource burden

Time eff.

Adoption behavior

Affective response

Supervision workload

Recovery behavior

Risk exposure

Trust calibration

𝑛

50.1 30.8

9.6 23.2

16.1 6.5

11.1 11.1

7.2 12.2

2.2 6.9

2.3 2.3

0.4 4.7

0.9 2.4

25,172 19,595

+19.3

−13.6

+9.6

0.0

−5.0

−4.7

0.0

−4.3

−1.5

Before an OpenClaw

Agent Run

the conditions a user set in advance

Value Group

Most Discussed Agent Aspect

The Condition

When Users Could Change It

Autonomous Operation Dependable Operation Affordable Operation

model core

the reach the model and host the resource arrangement

Bounded Reach Reviewability Equitable Access

environment access

before a later run after a failure, before a later run before a run; spent inside it at setup, before any run at a run’s edges; asked inside it at install, before the first run

model core resource accounting observability, human oversight

the reach the gate, the record

system access

the install

§5.1 §5.2 §5.3 §5.4 §5.5 §5.6

In Autonomous Operation alone, the aspect most discussed and the condition its users changed differ. a mid-run request §5.5

2

The Change Available

While the Run Is Under Way

3

the agent acts; the user is reached only by a request

4

After the Run Finished

the bill, the record, and verification

Figure 6: Value-sensitive delegation across the six value groups. Each row gives a group’s most discussed agent aspect (Section 4.2), the condition its users changed, and when posts described them changing it. The third and fourth columns are synthesized across Sections 5.1–5.6 rather than paired within user posts. Step 4 marks the return path, in which a user changes a condition for a later run rather than for the run that raised the value. they can also re-ask about a decision the user already approved, or arrive where the user cannot answer them. Separating these three explains why supervision was reported as both costless and

costly within the same value group, since gates and records keep the user’s authority in force without occupying their time, whereas a request has to be answered during the run (Sections 5.5 and 5.7). 12

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

The conditions also explain how users could name a value as threatened even when no harm occurred, as seen in Bounded Reach posts concerning granted access (Section 5.4). Value-sensitive delegation thus contributes a concrete unit of analysis to VSD. VSD identifies a human value between what a designer built and what a user brings, and a delegated run holds the two apart in time. A condition is what carries the user’s intent across that interval. The user fixes the condition beforehand and it stays in force while the agent works, so it is the user’s contribution that is present when the value is realized. A study of a deployed agent can therefore ask which condition a reported value attached to, and when the user could have changed it. A generic list of values supports neither question, because it records what mattered to someone without recording what they could change. Value-sensitive delegation therefore locates a user’s remaining agency in the conditions rather than in the run itself.

6.2

One cost held its direction everywhere. Supervision workload was mostly not met in all six value groups, and it was the only majority-not-met user outcome whose cells cleared the display threshold everywhere (Section 5.7). Its direction therefore did not depend on what a value group cared about or on which condition its users acted. Human workload research has long kept result and cost apart, rating how a task went separately from what it cost the user who performed it [19], and agent studies have priced human oversight by evaluating specific designer-chosen interventions, such as checkpoints or interruptions [7, 21, 53, 68]. The charge in these user posts has no such author, since it attaches to letting a run proceed at all and survives across six value groups that share relatively little else. Automation research offers the user one remedy, adjusting how much work they hand over [32, 41], but that remedy does not address a charge whose direction persists regardless of what the user delegates. We therefore read supervision workload in these posts as a charge of delegation itself rather than as the price of a particular human oversight mechanism. What did vary was the gain. Where users described what the agent delivered, how often the value held moved with the value group. In three of the six value groups, a leading user outcome that mostly held sat beside one that mostly did not, most widely in Bounded Reach (Section 5.7). Value alignment work argues for human values built for a context [33] and for reaching the people whose values are at stake [4, 30, 45]. Our findings support that argument for the gains and complicate it for the supervision cost. The gains moved with the value group, so a context-specific method would find them, whereas the supervision cost pointed in one direction in every group, even as its magnitude varied, leaving such a method nothing to catch. Holistic evaluation asks for many measures reported together rather than one headline score [5], and recent surveys of agent evaluation [37, 65] and a human-centered evaluation framework [11] describe a field where task completion dominates while cost and safety draw less measurement. User posts in our study thus point to a sharper rule than reporting more numbers. A measure that changes direction with what users care about and a measure that holds one direction should be reported separately, because averaging them hides the one that never turns.

Gains and Costs Users Attributed to Running OpenClaw Autonomously

In Autonomous Operation, the gain users described was time they no longer had to spend rather than a better result (Section 5.1). Human–AI collaboration measures what a user and an AI system achieve together [21] and what taking part costs the user [7], but both measurements require the user to be present during the run. User experience investigation [20], like our study, focuses on the experience in a user’s own state and surroundings rather than in an AI system’s properties. A delegated agent like OpenClaw carries that argument to its limit, because the hours a user gets back are hours spent entirely away from the agent. Recent agent studies have moved the same way, looking past the run to what people manage around it [12, 38, 67]. Our findings give that move an empirical footing, since the benefit these users reported does not appears in what a run produced. One post in Autonomous Operation ran the other way, describing time as shrinking rather than expanding and asking what a user’s own effort is worth once the agent does the work (Section 5.1). A single post settles nothing, but it marks a limit on the gain, because an account that counts the hours a user no longer spends, without asking what the user does with them, would read that post as a success. The costs users mentioned similarly materialized outside the finished task before anything went wrong. In Affordable Operation, the charge came from the agent’s readiness rather than from a requested task. In Bounded Reach, the risk users described was access they had granted, and in both value groups these were the posts where the value was mostly not met (Sections 5.3, 5.4, and 5.7). Two lines of work describe this situation, and neither treats a granted permission as a cost the user is already paying. Dependable computing counts such a condition as an active fault only once it produces an error, and treats it as dormant until then [3], while security research on these agents judges a deployment by what an adversary could achieve from it [50, 52, 60, 61, 66]. The user posts in our study part company with both, because for their authors the granted access was already the cost, written down while the deployment was working as intended. The accounts differ in what counts as harm, an event in the literature and an ongoing state in these user posts.

6.3

Design Implications

The design question that follows from Sections 6.1 and 6.2 is not how to make an agent perform better, but what an AI agent should let a user set, and when. Each implication below takes one of the five conditions and answers that question for it, from the narrowest change to the one that decides whether a user can make any of the others. From spending decided inside a run → Design a ceiling the agent cannot exceed. Affordable Operation concentrated on resource accounting, and that agent aspect carries the lowest met rate in the user posts (Sections 4.2 and 4.3). The resource arrangement is therefore where design should start. One user had already assembled the accounting the agent lacked, an alert threshold and scheduled context compaction, and still could not hold consumption down (Section 5.3). An alert posts spending that has already happened, so the user is told about a decision the run has already 13

Conference’17, July 2017, Washington, DC, USA

Ma et al.

made. Human–AI design guidelines ask a system to let a user customize what it monitors and how it behaves [1], thereby giving the user control. However, a global control that declares what an agent will do cannot stop what a run spends once the run begins. AI agents like OpenClaw should therefore accept a ceiling they must run within, not an alert the user must watch. A bill should also name the trigger behind each charge, so that a charge produced by the agent’s readiness reads differently from requested work. From reach that ships open → Design a narrow default and let users widen it. Bounded Reach concentrated on environment access, and risk exposure led its user outcomes (Sections 4.2 and 5.4). These users documented the reach they had granted before rather than harm they had suffered. One user separated the defects found in the software from a default that shipped open. A default of that kind is not a failure produced by a run, so no record of a run surfaces it. The harnesses that evaluate these agents already narrow reach, since a harness fixes what an agent may touch [28, 34, 63, 64, 69]. That narrowing stays inside the evaluation and does not reach the person who installs the agent. The narrow default should therefore ship in place, leaving widening as the user’s action. An agent should also show what a run may reach at the moment a user grants it, so opening reach becomes a decision. Nothing here shows that a narrower default would have prevented an incident. From requests that repeat or go unseen → Design an approval that persists and reaches the user. Reviewability was the one group whose posts concentrated on two agent aspects, observability and human oversight (Section 4.2). That pair separates looking from answering. Within the same value group, one leading user outcome mostly held and the other mostly did not (Section 5.7). Conditions set before a run and records read afterward asked nothing of the user while the run was underway. A request that arrived mid-run cost something, either because the user had already approved that step or because it appeared somewhere they could not see it (Section 5.5). Work on agent oversight models how much confirmation a run should carry [68] and what form an intervention should take [29]. These user posts point to a different variable. What made a request expensive was not its frequency but whether it reopened a settled decision and whether it landed where the user could answer. An approval should therefore persist within the scope the user granted it, so a run stops re-asking about a step that the user authorized. A request should also travel the channel the agent already uses to reach that user, not a dialog on a machine no one watches. From swapping the LLM model as the available repair → Design a completion test the user writes. The two largest value groups both discussed the model core more than any other agent aspect, without being set apart by it (Section 4.2). The condition available there is substitution, and these posts record users changing which LLM model ran rather than how it behaved. After a failure, they swapped the model or lowered what they asked of it (Section 5.2). A design that offers a better model therefore offers them another substitution rather than a new kind of control. The fix these value groups needed is not in the model or host but in a condition set before a run, because completion no longer evidenced that the work was done and these users had already taken verification back themselves. Benchmark harnesses already solve that problem, declaring what counts as done before a run and having

something other than the agent check it [28, 64]. That mechanism transfers to a deployment, with one change. The user writes the test, not a benchmark author, so it becomes a condition set around the agent rather than a fixture inside a harness. OpenClaw should therefore accept a completion test that its user writes before a run. A retry and verification policy should sit alongside it as a second condition set in advance. From conditions only a user who got in can set → Design the install as part of the agent. Equitable Access concentrated on system access and carried the lowest met rate of the six value groups (Sections 4.2 and 4.3). What these users credit for a working agent is the path rather than the agent at the end of it (Section 5.6). A documented onboarding command made a provider switch easy, a host provider earned credit the agent did not, and other users weighed the install itself, asking whether someone uncomfortable with a command line could run it at all. Every implication above assumes a user who got far enough to set the condition it names. An agent should therefore ship with the spending ceiling, the narrow reach default, the persistent approval, and the completion test already in place and adjustable without a command line, rather than leaving the user to assemble them.

7

Limitations and Future Work

Our study is observational and post-level. Because we analyzed Reddit posts from OpenClaw’s first few months of public use, the people it records are early adopters whose posts may overrepresent experiences worth writing about. Without stable author identifiers, we cannot separate many posts by one user from a single post by each of many users, so every percentage in this paper describes posts rather than users. Our associations also do not establish causal effects. As detailed in Section 3.3, the LLM-assisted coding showed variable reliability across targets, so estimates for infrequent categories carry the most classification uncertainty. Two user outcomes carry a further limit. Assessing whether a user’s trust is appropriately calibrated requires knowing both how much they relied on the agent [41] and whether the agent’s actual performance justified that reliance [32]. Because a single outcome label applied to a Reddit post cannot capture the agent’s unobserved, ground-truth performance, we draw no claims from trust calibration or recovery behavior in the value-group narratives in Section 5, although their cells appear in Figure 5 and Table 2. Finally, the six value groups we constructed for this analysis were not validated as a measurement model. These limits mark out our next studies. The coding schema can be applied to posts about other agents and platforms to test which of the patterns here belong to OpenClaw and which belong to delegation itself. User-level designs, such as interviews or diary studies that follow the same user across runs, could test whether the operating conditions that carried the values in these posts also govern individual experience, and could recover the outcomes that public posts undercount. Our findings also suggest a change to agent evaluation. Recording the conditions an agent ran under alongside its task score would let benchmark results speak to the surface where these users registered value success and failure. 14

Value-Sensitive Delegation in Everyday AI Agent Use

8

Conference’17, July 2017, Washington, DC, USA

Conclusion

[14] Jacob Cohen. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20, 1 (1960), 37–46. doi:10.1177/ 001316446002000104 [15] Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. Association for Computing Machinery, Article 37, 23 pages. doi:10.1145/3617694.3623261 [16] Batya Friedman. 1996. Value-Sensitive Design. Interactions 3, 6 (1996), 16–23. doi:10.1145/242485.242493 [17] Batya Friedman, Peter H. Kahn, Jr., and Alan Borning. 2006. Value Sensitive Design and Information Systems. In Human-Computer Interaction and Management Information Systems: Foundations, Ping Zhang and Dennis Galletta (Eds.). M.E. Sharpe, Armonk, NY, 348–372. [18] Iason Gabriel. 2020. Artificial Intelligence, Values, and Alignment. Minds and Machines 30, 3 (2020), 411–437. doi:10.1007/s11023-020-09539-2 [19] Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. In Human Mental Workload, Peter A. Hancock and Najmedin Meshkati (Eds.). Advances in Psychology, Vol. 52. North-Holland, Amsterdam, The Netherlands, 139–183. doi:10.1016/S0166-4115(08)62386-9 [20] Marc Hassenzahl and Noam Tractinsky. 2006. User Experience—A Research Agenda. Behaviour & Information Technology 25, 2 (2006), 91–97. doi:10.1080/ 01449290500330331 [21] Gaole He, Gianluca Demartini, and Ujwal Gadiraju. 2025. Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 414, 22 pages. doi:10.1145/3706598.3713218 [22] Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI With Shared Human Values. In International Conference on Learning Representations. https://openreview.net/ forum?id=dNy_RKzJacY [23] Hsiu-Fang Hsieh and Sarah E. Shannon. 2005. Three Approaches to Qualitative Content Analysis. Qualitative Health Research 15, 9 (2005), 1277–1288. doi:10. 1177/1049732305276687 [24] Patrik Hummel, Matthias Braun, Max Tretter, and Peter Dabrock. 2021. Data Sovereignty: A Review. Big Data & Society 8, 1, Article 2053951720982012 (2021), 17 pages. doi:10.1177/2053951720982012 [25] Ben Hutchinson, Negar Rostamzadeh, Christina Greer, Katherine Heller, and Vinodkumar Prabhakaran. 2022. Evaluation Gaps in Machine Learning Practice. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, 1859–1876. doi:10.1145/3531146. 3533233 [26] International Organization for Standardization. 2018. Ergonomics of HumanSystem Interaction—Part 11: Usability: Definitions and Concepts (2 ed.). International Standard ISO 9241-11:2018. International Organization for Standardization, Geneva, Switzerland. https://www.iso.org/standard/63500.html [27] Maurice Jakesch, Zana Buçinca, Saleema Amershi, and Alexandra Olteanu. 2022. How Different Groups Prioritize Ethical Values for Responsible AI. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, 310–323. doi:10.1145/3531146.3533097 [28] Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve RealWorld GitHub Issues?. In International Conference on Learning Representations. https://openreview.net/forum?id=VTF8yNQM66 [29] Joohee Kim, Sungbeom Cho, Duc M. Nguyen, Jaehyeong Jeon, Minjeong Shin, and Sungahn Ko. 2026. “Here, Let Me Help”: An Empirical Study of User Interventions in Human–Web Agent Collaboration. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 1541, 18 pages. doi:10.1145/3772318.3791536 [30] Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale. 2024. The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models. In Advances in Neural Information Processing Systems, Vol. 37. Curran Associates, Inc., 105236–105344. doi:10.52202/079017-3342 [31] Christopher A. Le Dantec, Erika Shehan Poole, and Susan P. Wyche. 2009. Values as Lived Experience: Evolving Value Sensitive Design in Support of Value Discovery. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, 1141–1150. doi:10.1145/1518701. 1518875 [32] John D. Lee and Katrina A. See. 2004. Trust in Automation: Designing for Appropriate Reliance. Human Factors 46, 1 (2004), 50–80. doi:10.1518/hfes.46.1. 50_30392 [33] Enrico Liscio, Michiel van der Meer, Luciano C. Siebert, Catholijn M. Jonker, and Pradeep K. Murukannaiah. 2022. What Values Should an Agent Align With?

This study presented an LLM-assisted, VSD-grounded content analysis of 73,093 first-person Reddit posts of OpenClaw use. We identified an autonomy that usually held against a dependability that did not, resource accounting as the agent aspect that least often met the value raised against it, and an oversight cost that stayed low at a run’s edges but rose once the agent stopped to ask inside one. Our findings call for locating the human values at stake in agent use within the conditions a user sets around a run, namely what it may spend, what it may reach, and when it must ask, rather than within what a run returns, a relationship we name value-sensitive delegation. An agent that finishes the task can still fail the user who handed it over.

References [1] Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. 2019. Guidelines for HumanAI Interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 3, 13 pages. doi:10.1145/3290605.3300233 [2] Ron Artstein and Massimo Poesio. 2008. Survey Article: Inter-Coder Agreement for Computational Linguistics. Computational Linguistics 34, 4 (2008), 555–596. doi:10.1162/coli.07-034-R2 [3] Algirdas Avizienis, Jean-Claude Laprie, Brian Randell, and Carl Landwehr. 2004. Basic Concepts and Taxonomy of Dependable and Secure Computing. IEEE Transactions on Dependable and Secure Computing 1, 1 (2004), 11–33. doi:10.1109/ TDSC.2004.2 [4] Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Díaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022. Power to the People? Opportunities and Challenges for Participatory AI. In Proceedings of the 2022 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. Association for Computing Machinery, Article 6, 8 pages. doi:10.1145/3551624.3555290 [5] Rishi Bommasani, Percy Liang, and Tony Lee. 2023. Holistic Evaluation of Language Models. Annals of the New York Academy of Sciences 1525, 1 (2023), 140–146. doi:10.1111/nyas.15007 [6] Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101. doi:10.1191/ 1478088706qp063oa [7] Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1, Article 188 (2021), 21 pages. doi:10.1145/3449287 [8] Dylan Butts and Matthew Chin. 2026. From Clawdbot to Moltbot to OpenClaw: Meet the AI Agent Generating Buzz and Fear Globally. https://www.cnbc.com/2026/02/02/openclaw-open-source-ai-agent-risecontroversy-clawdbot-moltbot-moltbook.html. CNBC. Published 2 February 2026, updated 12 March 2026. Accessed 13 August 2026. [9] Cristiano Castelfranchi and Rino Falcone. 1998. Towards a Theory of Delegation for Agent-Based Systems. Robotics and Autonomous Systems 24, 3–4 (1998), 141–157. doi:10.1016/S0921-8890(98)00028-1 [10] Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstantinos Voudouris, Umang Bhatt, Adrian Weller, David Krueger, and Tegan Maharaj. 2023. Harms from Increasingly Agentic Algorithmic Systems. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’23). Association for Computing Machinery, 651–666. doi:10.1145/3593013.3594033 [11] Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia-Jun Li. 2025. Toward a Human-Centered Evaluation Framework for Trustworthy LLMPowered GUI Agents. arXiv:2504.17934 [cs.HC] https://arxiv.org/abs/2504.17934 [12] Ruijia Cheng, Jenny T. Liang, Eldon Schoop, and Jeffrey Nichols. 2026. Mapping the Design Space of User Experience for Computer Use Agents. In Proceedings of the 31st International Conference on Intelligent User Interfaces. Association for Computing Machinery, 646–662. doi:10.1145/3742413.3789132 [13] Christina Cociancig and Hendrik Heuer. 2026. Toward a Clearer Process for Value Sensitive Artificial Intelligence. Science and Engineering Ethics 32, 2, Article 13 (2026). doi:10.1007/s11948-026-00583-2 15

Conference’17, July 2017, Washington, DC, USA

Ma et al.

An Empirical Comparison of General and Context-Specific Values. Autonomous Agents and Multi-Agent Systems 36, 1, Article 23 (2022). doi:10.1007/s10458-02209550-0 [34] Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. 2024. AgentBench: Evaluating LLMs as Agents. In International Conference on Learning Representations. https: //openreview.net/forum?id=zAdUB0aCTQ [35] Brian Lubars and Chenhao Tan. 2019. Ask Not What AI Can Do, But What AI Should Do: Towards a Framework of Task Delegability. In Advances in Neural Information Processing Systems, Vol. 32. Curran Associates, Inc. https://proceedings. neurips.cc/paper/2019/hash/d67d8ab4f4c10bf22aa353e27879133c-Abstract.html [36] Christopher A. Miller and Raja Parasuraman. 2007. Designing for Flexible Interaction Between Humans and Automation: Delegation Interfaces for Supervisory Control. Human Factors 49, 1 (2007), 57–75. doi:10.1518/001872007779598037 [37] Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. 2025. Evaluation and Benchmarking of LLM Agents: A Survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. Association for Computing Machinery, 6129–6139. doi:10.1145/3711896.3736570 [38] Suchismita Naik, Amanda Snellinger, Austin L. Toombs, Scott Saponas, and Amanda K. Hall. 2025. Exploring Early Adopters’ Use of AI Driven Multi-Agent Systems to Inform Human-Agent Interaction Design: Insights from Industry Practice. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 677, 8 pages. doi:10.1145/3706599.3706693 [39] Helen Nissenbaum. 2004. Privacy as Contextual Integrity. Washington Law Review 79, 1 (2004), 119–157. https://digitalcommons.law.uw.edu/wlr/vol79/iss1/10/ [40] OpenClaw Foundation. 2026. OpenClaw: Your Own Personal AI Assistant. Any OS. Any Platform. GitHub repository. https://github.com/openclaw/openclaw MIT License. Accessed 13 August 2026. [41] Raja Parasuraman and Victor Riley. 1997. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors 39, 2 (1997), 230–253. doi:10.1518/ 001872097778543886 [42] Raja Parasuraman, Thomas B. Sheridan, and Christopher D. Wickens. 2000. A Model for Types and Levels of Human Interaction with Automation. IEEE Transactions on Systems, Man, and Cybernetics—Part A: Systems and Humans 30, 3 (2000), 286–297. doi:10.1109/3468.844354 [43] Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada. 2021. AI and the Everything in the Whole Wide World Benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Vol. 1. https://datasets-benchmarks-proceedings.neurips.cc/ paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html [44] Malak Sadek, Marios Constantinides, Daniele Quercia, and Céline Mougenot. 2024. Guidelines for Integrating Value Sensitive Design in Responsible AI Toolkits. In Proceedings of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 472, 20 pages. doi:10.1145/3613904. 3642810 [45] Malak Sadek and Céline Mougenot. 2025. Challenges in Value-Sensitive AI Design: Insights from AI Practitioner Interviews. International Journal of Human– Computer Interaction 41, 17 (2025), 10877–10894. doi:10.1080/10447318.2024. 2439021 [46] Filippo Santoni de Sio and Jeroen van den Hoven. 2018. Meaningful Human Control over Autonomous Systems: A Philosophical Account. Frontiers in Robotics and AI 5, Article 15 (2018), 14 pages. doi:10.3389/frobt.2018.00015 [47] Hope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David Mimno, and Sarita Schoenebeck. 2025. Large Language Models in Qualitative Research: Uses, Tensions, and Intentions. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 481, 17 pages. doi:10.1145/3706598.3713120 [48] Katie Shilton. 2013. Values Levers: Building Ethics into Design. Science, Technology, & Human Values 38, 3 (2013), 374–397. doi:10.1177/0162243912436985 [49] Mona Sloane, Emanuel Moss, Olaitan Awomolo, and Laura Forlano. 2022. Participation Is not a Design Fix for Machine Learning. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. Association for Computing Machinery, Article 1, 6 pages. doi:10.1145/3551624. 3555285 [50] Hang Su, Jun Luo, Chang Liu, Xiao Yang, Yichi Zhang, Yinpeng Dong, and Jun Zhu. 2025. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents. arXiv:2506.23844 [cs.AI] https://arxiv.org/abs/2506.23844 [51] Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L. Griffiths. 2024. Cognitive Architectures for Language Agents. Transactions on Machine Learning Research (2024). https://openreview.net/forum?id=1i6ZCvflQJ [52] Surada Suwansathit, Yuxuan Zhang, and Guofei Gu. 2026. A Security Analysis of the OpenClaw AI Agent Framework. arXiv:2603.27517 [cs.CR] https://arxiv. org/abs/2603.27517 [53] Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang

Ye, Tianshi Li, Ziang Xiao, Yaxing Yao, and Toby Jia-Jun Li. 2026. Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 403, 26 pages. doi:10.1145/3772318.3791568 [54] Steven Umbrello and Ibo van de Poel. 2021. Mapping Value Sensitive Design onto AI for Social Good Principles. AI and Ethics 1, 3 (2021), 283–296. doi:10. 1007/s43681-021-00038-3 [55] Rama Adithya Varanasi and Nitesh Goyal. 2023. “It is currently hodgepodge”: Examining AI/ML Practitioners’ Challenges during Co-production of Responsible AI Values. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 251, 17 pages. doi:10.1145/3544548.3580903 [56] Oleksandra Vereschak, Fatemeh Alizadeh, Gilles Bailly, and Baptiste Caramiaux. 2024. Trust in AI-Assisted Decision Making: Perspectives from Those Behind the System and Those for Whom the Decision Is Made. In Proceedings of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 28, 14 pages. doi:10.1145/3613904.3642018 [57] Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024. A Survey on Large Language Model Based Autonomous Agents. Frontiers of Computer Science 18, 6, Article 186345 (2024). doi:10.1007/ s11704-024-40231-1 [58] Qiaosi Wang, Michael Madaio, Shaun Kane, Shivani Kapania, Michael Terry, and Lauren Wilcox. 2023. Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 249, 16 pages. doi:10.1145/3544548.3581278 [59] Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, and Yingchun Wang. 2024. Fake Alignment: Are LLMs Really Aligned Well?. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, 4696–4712. doi:10.18653/v1/2024.naacl-long.263 [60] Yuhang Wang, Feiming Xu, Zheng Lin, Guangyu He, Yuzhe Huang, Haichang Gao, Zhenxing Niu, Shiguo Lian, and Zhaoxiang Liu. 2026. From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent. arXiv:2602.08412 [cs.AI] https://arxiv.org/abs/2602.08412 [61] Zijun Wang, Haoqin Tu, Letian Zhang, Hardy Chen, Juncheng Wu, Xiangyan Liu, Zhenlong Yuan, Tianyu Pang, Michael Qizhe Shieh, Fengze Liu, Zeyu Zheng, Huaxiu Yao, Yuyin Zhou, and Cihang Xie. 2026. Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw. arXiv:2604.04759 [cs.CR] https://arxiv. org/abs/2604.04759 [62] Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang, Qi Zhang, and Tao Gui. 2025. The Rise and Potential of Large Language Model Based Agents: A Survey. Science China Information Sciences 68, 2, Article 121101 (2025). doi:10.1007/s11432-0244222-0 [63] Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Jing Hua Toh, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. In Advances in Neural Information Processing Systems, Vol. 37. Curran Associates, Inc., 52040–52094. doi:10.52202/079017-1650 [64] Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. 2025. 𝜏 bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. In The Thirteenth International Conference on Learning Representations. https: //openreview.net/forum?id=roNSXZpUDN [65] Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. 2026. A Survey on Evaluation of LLM-based Agents. In Findings of the Association for Computational Linguistics: ACL 2026. Association for Computational Linguistics, San Diego, California, United States, 26690–26714. doi:10.18653/v1/2026.findings-acl.1330 [66] Zonghao Ying, Xiao Yang, Siyang Wu, Yumeng Song, Yang Qu, Hainan Li, Tianlin Li, Jiakai Wang, Aishan Liu, and Xianglong Liu. 2026. Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw. arXiv:2603.12644 [cs.CR] https://arxiv.org/abs/2603.12644 [67] Shuning Zhang, Jingruo Chen, Zhiqi Gao, Jiajing Gao, Xin Yi, and Hewu Li. 2026. Characterizing Unintended Consequences of GUI Agents For Web Browsing. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, Article 247, 19 pages. doi:10.1145/3772318. 3790696 [68] Jieyu Zhou, Aryan Roy, Sneh Gupta, Daniel Weitekamp, and Christopher J. MacLellan. 2026. When Should Users Check? Modeling Confirmation Frequency in Multi-Step Agentic AI Tasks. In Proceedings of the 2026 CHI Conference on 16

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

Human Factors in Computing Systems. Association for Computing Machinery, Article 1649, 20 pages. doi:10.1145/3772318.3790655 [69] Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A Realistic Web Environment for Building Autonomous Agents. In International Conference on Learning Representations. https://openreview.net/forum?id=oKn9c6ytLx

A

Agent Aspect Taxonomy: - system_access: installation, onboarding, authentication, provider configuration, subscription state, account readiness, billing access required before use, or workspace setup. - model_core: the underlying LLM, including model choice, provider, version, tier, decoding parameters, backend, or runtime model. - agent_profile: the system prompt, persona, role definition, identity, or framing instructions that shape agent behavior. - task_specification: the task goal, scope, boundary, constraint, authority level, approval rule, stopping condition, or success criterion. - reasoning: visible or user-reported reasoning, reflection, self-critique, deliberation, decision rationale, or intermediate reasoning before acting. Do not infer hidden reasoning. - planning: task decomposition, goal sequencing, plan formation, or plan revision. - memory: context-window management, short-term task state, long-term memory, retrieval, summarization, persistence, forgetting, or conversation history. - tool_selection: the choice of which tool, function, command, or action to invoke. - tool_execution: invocation of tools, function calls, API calls, browser actions, terminal commands, app actions, or external operations. - action_effects: downstream changes to user-visible artifacts, including creation, editing, deletion, movement, or modification of files, documents, code, repositories, or data. - environment_access: sandbox boundaries, file permissions, credential access, local or cloud resources, data-access scope, external accounts, or workspace permissions. - observability: traces, logs, action history, current-state surfacing, decision rationale, explanation, or progress visibility. - error_handling: error detection, reporting, debugging, retry, rollback, repair, or recovery after a failed action. - human_oversight: approval gates, confirmation prompts, intervention points, supervision affordances, or human-in-the-loop controls. - multi_agent_orchestration: coordination among agents or subagents, including role assignment , agent-to-agent communication, and delegated workers. - resource_accounting: tokens, compute, quota, rate limits, budget pressure, billing impact, or resource use during or after operation. - runtime_performance: latency, speed, responsiveness, throughput, stability, or performance under task load. - other: a substantive agent-design concern outside the taxonomy when no listed aspect fits.

Corpus Construction and Analytic Samples

The analytic samples were constructed as follows: (1) The source corpus contained 1,100,308 posts. (2) Stage 1 retained 187,479 candidate posts; the rest were not first-person, were too thin or unclear, or did not yield a usable classification. (3) Among the candidates included in the final coding run, Stage 2 retained 73,797 first-person posts with the required coding fields and schema-valid output; the rest were not first-person or were too thin or unclear, lacked required coding fields, or failed output validation. (4) RQ1 uses the 73,093 posts assigned to a predefined value; the other retained posts had open-coded values without a documented mapping to the predefined taxonomy. RQ2 uses the 44,767 posts within the RQ1 sample that contain an attributable user outcome.

B

Agent Aspect Boundary Rules: - system_access concerns getting the system ready for use; environment_access concerns what resources the agent may reach after access is configured. - tool_execution concerns invoking an operation; action_effects concerns the user-visible change produced by that operation. - A failed task is not automatically error_handling. Use error_handling only when detection, reporting, retry, rollback, repair, or recovery is central. - resource_accounting concerns measured or constrained resource use; runtime_performance concerns latency, responsiveness, throughput, or stability. - Use task_specification only when the excerpt concerns a task goal, scope, constraint, delegated authority, approval rule, stopping condition, or success criterion. - Select only the component linked to the dominant value-centered excerpt, even when a product description mentions several components.

Prompt Design for Value-Centered Coding

The Stage 2 prompt below also asks for an outcome sentiment label. We coded that field but do not analyze or report it in this paper, and no result in the paper depends on it. The specification below summarizes the substantive Stage 2 instructions used to generate the labels analyzed in this study. It presents the coding task, evidence requirements, category definitions and decision boundaries, and primary output fields. Stage 1 used a screening-only prompt that assigned the same three relevance categories and produced no other labels. The “HAAI refinements” in the listing are the nine study-specific values described in Section 3.2.

Human Value Taxonomy: The initial VSD values are sensitizing concepts. The human-autonomous-agent interaction (HAAI) refinements extend them for recurring conditions in interaction with autonomous agents. Both sets describe stakeholder-relevant conditions rather than agent components, safeguards, or implementation mechanisms. Initial VSD Values: - human welfare: the physical, material, and psychological well-being of people directly or indirectly affected by the technology. - property ownership: people's rights to possess, use, manage, benefit from, transfer, or dispose of information and other assets. - privacy: a person's claim or entitlement to determine how information about them is accessed , collected, used, and communicated. - freedom from bias: protection from systematic unfairness produced by pre-existing, technical , or emergent bias. - universal usability: the ability of people with diverse skills, abilities, resources, platforms, and contexts to use the technology successfully. - trust: a person's willingness to rely on an agent or its responsible operators while accepting vulnerability to their actions, failures, or betrayal. - autonomy: a person's ability to decide, plan, and act in pursuit of their own goals. - informed consent: voluntary agreement made after adequate disclosure and comprehension, with competence and a genuine ability to choose. - accountability: the state in which actions and outcomes are traceable to people or institutions that can explain and answer for them. - identity: a person's understanding of who they are over time, including continuity, authorship, expertise, and social or professional role. - calmness: a peaceful and composed state that is not disrupted by unnecessary interruption, anxiety, or cognitive overload. - environmental sustainability: the preservation of ecosystems and resources for present needs without compromising future generations.

Listing 1: Summary of the Stage 2 value-centered coding instructions. Task: Analyze one Reddit text unit about OpenClaw use. Assign a relevance category and, when the text supports value-relevant evidence, identify one primary human value, the linked agent aspect, value fulfillment, and any attributable user outcome and sentiment. Evidence and Gating Rules: 1. Code only what the text states or strongly implies. Do not invent values, agent behavior, or outcomes. 2. A value may be explicit or strongly implied by an evaluation, concern, expectation, barrier , tradeoff, control decision, workaround, failure, benefit, or desired design condition. First-person language, a capability statement, or a product description alone does not establish a value. 3. For each first-person unit, re-read the full text once for implicit value evidence before leaving value_mentioned blank. 4. Select one verbatim phrase, sentence, or short multi-sentence excerpt that best supports the primary value. The value, linked agent aspect, fulfillment status, and optional outcome must all be grounded in that excerpt; do not combine unrelated evidence. 5. Assign one primary label per categorical field, keep the taxonomies conceptually distinct, and leave a field blank when evidence is insufficient. 6. If relevance_category is too_thin_or_unclear or no value is supported, leave all substantive fields blank and set outcome_sentiment to null. 7. Code a user outcome only when the shared excerpt states or strongly implies a concrete consequence for the author. A value concern alone is not an outcome.

HAAI Value Refinements: - meaningful human control: the state in which an agent remains responsive to relevant human reasons and its actions and outcomes remain traceable to people who understand and can assume responsibility. - transparency: the availability of relevant and intelligible information about an agent's goals, capabilities, limits, state, reasoning, actions, and outcomes. - repairability: the ability to diagnose, reverse, correct, or recover from errors and unwanted changes. - contextual integrity: the preservation of context-specific norms governing who sends or receives what information, for which purpose, and under what conditions. - data sovereignty: the legitimate authority of individuals or collectives over data access, storage, processing, transfer, retention, and deletion. - affordability: the state in which monetary costs do not unreasonably exclude or burden intended users. - resource stewardship: the responsible and proportionate use of tokens, compute, quota, energy, network capacity, and related service resources. - dependability: a justified expectation that the agent will provide correct and consistent service under stated conditions, including reliability, availability, integrity, safety, and maintainability. - security: protection from unauthorized access, disclosure, manipulation, disruption, or control.

Relevance Categories: - first_person_experience: the author's own use, attempted use, concrete setup, experienced problem, help-seeking about that setup, or directly observed agent behavior. - secondhand_observation: a substantive evaluation or value concern not grounded in the author 's own experience. - too_thin_or_unclear: insufficient context, a passing mention, generic promotion or question, off-topic text, a deleted or truncated fragment, or no substantive observation. - Product introductions, promotions, link shares, and lists of possible uses remain too_thin_or_unclear unless they contain a specific evaluation or value concern.

17

Conference’17, July 2017, Washington, DC, USA

Ma et al.

added for how each value applies to a delegated agent. The conceptual boundaries of the study-specific categories follow prior accounts of meaningful human control [46], contextual integrity [39], data sovereignty [24], and dependable and secure computing [3]. Example excerpts are drawn verbatim from the 400-post codebookdevelopment sample described in Section 3.3, except for calmness, environmental sustainability, identity, accountability, informed consent, and freedom from bias, which that sample contained no agreed instance of; those six come from the coded corpus. All were selected for brevity and clarity after review for identifying details.

Value Source: Value source records which part of the taxonomy supplied the primary label; it is not a separate value. - existing_vsd: value_mentioned matches one of the twelve initial VSD labels. - haai_refinement: value_mentioned matches one of the nine HAAI refinement labels. - open_code: neither list represents the human value. Do not open-code agent components, implementation mechanisms, safeguards, emotions, task outcomes, workload, debugging actions, or recovery actions as values. Value Boundary Rules: - accountability concerns who must explain or answer for actions; dependability concerns whether the agent works correctly and consistently. - autonomy concerns a person's capacity to pursue their own goals; meaningful human control concerns governance of delegated agent action. - data sovereignty concerns authority over data boundaries and lifecycle; resource stewardship concerns proportionate use of computational or service resources. - transparency concerns intelligible information and visibility; accountability concerns answerability and responsibility. - affordability concerns monetary access or burden; resource stewardship concerns consumption of computational or service resources. - trust concerns willingness to rely while vulnerable; dependability concerns functional correctness and consistency. - privacy concerns a person's claim over information about them; contextual integrity concerns appropriate information flow within a social context; data sovereignty concerns authority over data access and lifecycle.

D

Table 5 gives the codebook produced by the thematic analysis described in Section 3.4. Coding was bottom-up within the dimensions the Stage 2 coding had already established, so each value group was read through its two most frequent user outcomes. One researcher labeled the sampled excerpts and grouped the labels into sub-themes, discussing the developing codebook with the research team throughout and revising labels and boundaries after each discussion. The two sub-themes above and the two sub-themes below each rule within a value group form that group’s two themes, which Sections 5.1–5.6 report as the bolded claims. The third column gives one excerpt from the sampled posts for each sub-theme, quoted as the poster wrote it.

Value Fulfillment: - met: the value is supported, enabled, protected, respected, or fulfilled. - not_met: the value is undermined, violated, threatened, frustrated, absent, requested, or expected but unavailable. - blank: the excerpt identifies a value but does not establish whether it is met. - Do not assign fulfillment when value_mentioned is blank. User Outcome Taxonomy: - task_effectiveness: success, failure, output quality, correctness, completion, or partial completion. - time_efficiency: time saved or wasted, faster completion, delay, or interruption. - resource_burden: cost, tokens, compute, quota, rate-limit pressure, or billing burden experienced as a consequence. - supervision_workload: monitoring effort, cognitive load, checking, babysitting, or contextmanagement burden. - affective_response: frustration, confusion, anxiety, calmness, delight, annoyance, or relief. - trust_calibration: confidence, trust repair, distrust, appropriate reliance, overreliance, or reduced reliance. - adoption_behavior: continued use, adoption intention, abandonment, disuse, churn, or tool switching. - risk_exposure: experienced, observed, or clearly anticipated data loss, unwanted change, security exposure, privacy exposure, or compliance consequence. - recovery_behavior: debugging, rollback, repair, prompt workaround, workflow workaround, or recovery-strategy switching. - Outcome labels describe consequences, not values. A value concern alone is insufficient for risk_exposure.

E

Human Validation

The separate 50-post Stage 1 relevance-screening pilot reported precision of .900, recall of .783, and 𝐹 1 = .837. Because the pilot was not a corpus-random audit of excluded posts, these metrics do not establish full-corpus recall. The 400-post codebook-development sample supports two comparisons. A coder decision retained the initial LLM label when the review cell was blank and used the coder-entered replacement when the cell was nonblank. Human–human reliability compares the two coder decisions using exact agreement and pooled Cohen’s 𝜅. LLM–human correspondence treats the initial LLM label as the prediction and each eligible coder decision as a separate reference. For targets other than relevance, the human–human metrics use the 328 posts for which neither coder’s relevance decision was too thin or unclear. The LLM–human metrics use 671 individual coder decisions for which that coder’s relevance decision was not too thin or unclear. Accuracy, macro-precision, macro-recall, macro-𝐹 1 , and weighted-𝐹 1 therefore compare the LLM labels with coder decisions rather than measure accuracy against an objective gold standard. Metric calculations trimmed outer whitespace but otherwise treated each recorded replacement string as a distinct label; no post hoc class normalization was applied. Table 6 reports both comparisons. Coders viewed the initial LLM labels before entering their decisions, so these results measure assisted human–human agreement and agreement between the LLM and coders rather than blind independent coding. The sample was used for codebook development and does not estimate label error across the full corpus. LLM–human correspondence varied across coding targets: macro-𝐹 1 was .635 for human value and .604 for user outcome despite higher accuracy and weighted-𝐹 1 . We therefore interpret estimates for infrequent

Outcome Sentiment: - -1: negative. - 0: neutral. - 1: positive. - null: no coded outcome; missing is not the same as neutral. Selected Contrastive Examples: - Beginner onboarding: "I am an absolute beginner and do not know where to start" can support universal usability:not_met and system_access; leave the outcome blank unless the excerpt states a consequence. - First-person activity without value evidence: "I am developing custom skills for clients" is first_person_experience, but the statement alone does not establish a value. - Tool versus effect: "The browser call never runs" supports tool_execution; "the agent deleted my file" supports action_effects. - Dependability and outcome: repeated incorrect results support dependability:not_met and may support task_effectiveness with negative sentiment. Output: Return one JSON object with the core evidence and primary coding fields below. { "relevance_category": "<first_person_experience | secondhand_observation | too_thin_or_unclear>", "coding_quote_excerpt": "<verbatim value-centered excerpt, or empty string>", "agent_aspect": "<one taxonomy label, or empty string>", "value_mentioned": "<one primary value, or empty string>", "value_source": "<existing_vsd | haai_refinement | open_code | empty string>", "value_fulfillment": "<met | not_met | empty string>", "user_outcome": "<one outcome-family label, or empty string>", "outcome_sentiment": <-1 | 0 | 1 | null> }

C

Thematic Analysis Codebook

Coding Schema Reference

Table 4 lists the 21 human values with their value groups, and Table 3 lists the nine user-outcome categories. Values marked with a dagger are the nine study-specific categories developed for this corpus, and their definitions condense the codebook wording used in Stage 2. The remaining 12 come from Friedman et al.’s heuristic list of values with ethical import and carry that list’s own definitions [17], with the leading refers to dropped and a parenthesis 18

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

Table 3: The nine user-outcome categories, defined as in Table 1. Each names what a reported consequence concerned; value fulfillment carries its direction. Example excerpts come from the 400-post codebook-development sample. User outcome

Definition

Example post excerpt

Task effectiveness

Whether the agent completed the task and how correct or usable the result was. Whether the agent saved the user time or cost them time. What a run cost the user in money and metered resources. The effort of watching, checking, and managing what the agent did. How the user felt about the experience. How far the user relied on the agent, and whether that reliance was warranted. Whether the user continued with the agent, intended to, or moved away from it. Harm to the user’s data, systems, or compliance position that a post reported or clearly anticipated. What the user did to repair or work around what the agent did.

“I’ve been using my agent to build a temperature trading bot with mixed results.”

Time efficiency Resource burden Supervision workload Affective response Trust calibration Adoption behavior Risk exposure Recovery behavior

“Im using an old surface pro 6 and it pretty slow with OpenClaw” “OpenClaw made openAI and Anthropic $100K from my credit card.” “I prefer cowork but it’s too hard to keep a remote instance running smoothly for me.” “It’s a total mess at my side as everyone is using everything.” “If an AI can read your emails, track your behavior, and make decisions for you... At what point do you stop being the one in control?” “I am loving openclaw and I plan to do much more.” “Just don’t give it access to anything important.” “I stopped updating open claw because every new version critically broke something and ended up costing me several hours just to get it back on track”

categories and cells cautiously; the exact-string calculation treats free-text replacement variants as distinct labels.

F

Equations 4 and 5 define this expected count and the corresponding observed-minus-expected difference. A positive value indicates that the observed met rate exceeds the rate expected from the group’s distribution of agent aspects; it does not identify a causal group effect. For fulfillment standardization, we calculated 95% percentile bootstrap intervals from 2,000 within-group post resamples using seed 7. We held the corpus-wide agent-aspect met rates 𝑝𝑎 , estimated from the 73,093-post RQ1 sample, fixed across resamples. Thus, the intervals omit uncertainty in those estimates and dependence among posts written by the same author. For the outcome heatmap, let 𝑛𝑔,𝑜 denote the number of posts in met denote the number value group 𝑔 with user outcome 𝑜, and let 𝑛𝑔,𝑜 of those posts whose primary value was coded met. We display the cell met rate as ! met 𝑛𝑔,𝑜 𝑀𝑔,𝑜 = 100 . (6) 𝑛𝑔,𝑜

Analysis Measures

Let 𝑂𝑖 𝑗 be the number of posts whose coded label in one field falls in row 𝑖 and whose coded label in a second field falls in column 𝑗, with row total 𝑛𝑖 · , column total 𝑛 · 𝑗 , and grand total 𝑁 . Under independence, the expected count is 𝑛𝑖 · 𝑛 · 𝑗 . (1) 𝐸𝑖 𝑗 = 𝑁 The Pearson statistic and Cramér’s 𝑉 are √︄ ∑︁ ∑︁ (𝑂𝑖 𝑗 − 𝐸𝑖 𝑗 ) 2 𝜒2 𝜒2 = , (2) , 𝑉 = 𝐸𝑖 𝑗 𝑁 min(𝑟 − 1, 𝑐 − 1) 𝑖 𝑗 where 𝑟 and 𝑐 are the table dimensions. Equations 1 and 2 define the expected counts and post-level association summaries. Every expected count exceeded five in the tables we report, with a minimum of 36.8 in the value-group-by-user-outcome table and 297.6 in the value-fulfillment-by-user-outcome table, so the 𝜒 2 approximation holds in both. The cell-level observed-to-expected (O/E) ratio is 𝐿𝑖 𝑗 =

𝑂𝑖 𝑗 𝑃 (column 𝑗 | row 𝑖) = . 𝐸𝑖 𝑗 𝑃 (column 𝑗)

Equation 6 defines the heatmap statistic. Values above 50 indicate that met posts outnumber not-met posts within the cell, and values below 50 indicate the reverse. Posts without a coded user outcome do not enter the heatmap, and cells with fewer than 20 posts are marked with an asterisk but are not interpreted.

(3)

Equation 3 shows that the O/E ratio compares a within-row share with its corresponding corpus-wide share. The within-row percentage 𝑂𝑖 𝑗 /𝑛𝑖 · , the O/E ratio 𝐿𝑖 𝑗 , and the reverse conditional share 𝑂𝑖 𝑗 /𝑛 · 𝑗 answer different questions and are reported separately. For value fulfillment, let 𝑝𝑎 be the corpus-wide met rate at agent aspect 𝑎, and let 𝑛𝑔,𝑎 be the number of posts in value group 𝑔 coded at aspect 𝑎. The expected number of met posts based on the distribution of agent aspects is ∑︁ 𝐸𝑔met = 𝑛𝑔,𝑎 𝑝𝑎 . (4) 𝑎

With 𝑂𝑔met denoting the observed number of met posts and 𝑛𝑔 the group total, the standardized difference in pp is   100 𝑂𝑔met − 𝐸𝑔met Δ𝑔met = . 𝑛𝑔

(5) 19

Conference’17, July 2017, Washington, DC, USA

Ma et al.

Table 4: The 21 human values and their six value groups; † marks the nine study-specific categories. Example excerpts come from the 400-post codebook-development sample. Value group

Value

Definition

Example post excerpt

Dependable Operation

Dependability†

Justified expectation that the agent delivers its intended service correctly and consistently under stated conditions. Expectations that exist between people who can experience good will, extend good will toward others, feel vulnerable, and experience betrayal [17] (agent: the user’s reliance on the agent and its operators, and the vulnerability that reliance creates). Ability to diagnose, reverse, correct, and recover from agent errors or unwanted changes.

“OpenClaw feels dumber than Claude... even though it uses Claude?” “Trust grows once people see how consistently it performs.”

People’s ability to decide, plan, and act in ways that they believe will help them to achieve their goals [17] (agent: the goals a user pursues by delegating a run rather than performing the work). People’s physical, material, and psychological well-being [17] (agent: the well-being of the user who delegates work to the agent). A peaceful and composed psychological state [17] (agent: a state the agent’s interruptions during and between runs do not disturb).

“I want my Mac back!”

Monetary costs of obtaining and using the agent do not unreasonably exclude or burden users. Responsible and proportionate use of finite resources such as tokens, compute, quota, and energy. Sustaining ecosystems such that they meet the needs of the present without compromising future generations [17] (agent: the energy and water the infrastructure a run depends on consumes).

“I have literally paid money that I cannot use”

Protection of people, systems, data, and credentials against unauthorized access, manipulation, or control. A claim, an entitlement, or a right of an individual to determine what information about himself or herself can be communicated to others [17] (agent: what the agent may read about its user and pass on). People’s understanding of who they are over time, embracing both continuity and discontinuity over time [17] (agent: continuity of authorship and role once work is delegated). A right to possess an object (or information), use it, manage it, derive income from it, and bequeath it [17] (agent: the assets and outputs a run consumes or produces). Legitimate authority to determine how one’s data are accessed, stored, processed, transferred, and deleted. Preservation of context-specific norms governing who may send or receive what information.

“Giving an agent system-level access is a security minefield.”

Availability of intelligible information about the agent’s goals, state, reasoning, actions, and outcomes. The agent remains responsive to human reasons, and its outcomes remain traceable to responsible people. The properties that ensure that the actions of a person, people, or institution may be traced uniquely to the person, people, or institution [17] (agent: tracing a run’s actions to the person answerable for them). Garnering people’s agreement, encompassing criteria of disclosure and comprehension (for informed) and voluntariness, competence, and agreement (for consent) [17] (agent: agreement to what the agent does and to what it does with a user’s data).

“probably staring at logs quite a lot if issues pop up.”

Making all people successful users of information technology [17] (agent: reaching a working agent setup regardless of skill, platform, or resources). Freedom from systematic unfairness perpetrated on individuals or groups, including pre-existing social bias, technical bias, and emergent social bias [17] (agent: bias in what the agent’s underlying model produces or filters).

“OpenClaw orchestration should be a tool at everyone’s disposal.”

Trust

Repairability† Autonomous Operation

Autonomy

Human welfare

Calmness

Affordable Operation

Affordability† Resource stewardship† Environmental sustainability

Bounded Reach

Security† Privacy

Identity

Property ownership

Data sovereignty† Contextual integrity† Reviewability

Transparency† Meaningful human control† Accountability

Informed consent

Equitable Access

Universal usability

Freedom from bias

20

“I had this issue after an openclaw update got fixed with another update”

“switched from self-hosting openclaw to managed hosting and my weekends came back” “OpenClaw always sends me a Telegram message after each run, even when the script finds nothing new.”

“I’m trying to figure out the smartest way to assign models based on what the task actually needs.” “AI data centers consuming offensive amounts of electricity and water is a driving factor in this for me.”

“OpenClaw is the worst privacy nightmare.”

“I’m a professional writer and have been resisting the temptation of using AI to tweak my work.” “I’m paying for X amount of tokens, I should be able to burn X amount of tokens via whatever appropriate means I can.” “Local-first control (your data stays yours)” “it keeps memory separated by project, so different topics don’t get mixed together”

“risky actions require human approval” “the llm doesn’t make a choice and I effectively need to be the final decision maker AKA fall guy if something bad were to occur.”

“why were the AI features ON by default, so opt-out, instead of OFF by default, so opt-in?”

“All the AI sourcing tools we’ve tried spit out the same people.”

Value-Sensitive Delegation in Everyday AI Agent Use

Conference’17, July 2017, Washington, DC, USA

Table 5: Thematic analysis codebook. Each row pairs a sub-theme with one representative excerpt from the sampled posts. The two sub-themes on either side of a rule within a value group make up one theme, and the twelve themes are the bolded claims in Sections 5.1–5.6. Value group

Sub-theme

Autonomous Operation

Extending reach to new surfaces

“The biggest unlock was letting the agent improve its own environment.”

Delegating whole workflows rather than steps

“So this morning I got OpenClaw to build an automated script that pays for parking at the times I used to get tickets.”

Hours returned rather than quality gained Returned time questioned

“My micro SaaS ops went from 8 hours a week to 45 minutes. Run Lobster (OpenClaw) runs the rest.”

Silent or partial failure

“OpenClaw sees empty content and silently falls through to the next model in your fallback chain — no error, just the wrong model answering.” “‘Verify before asserting’ exists because the agent told me a feature was working when it had never been tested.”

Dependable Operation

Verification taken back by the user Substitution as the available repair Lowered expectations rather than exit Affordable Operation

Consumption outside a requested task Spending visible only after the fact Re-arranging where the money goes Cost measured against what it displaced

Bounded Reach

Granted access as the exposure Anticipated rather than experienced harm Placement and isolation set at setup Boundaries that made a run usable

Reviewability

Control set in advance Reading the record afterward Requests that reopen settled decisions Requests that arrive out of reach

Equitable Access

Representative quotation

“Good point. When using n8n, the incremental value of openclaw and the time sink just wasn’t worth it.”

“Chat performance of GLM was really good, but I did not get tools to work stably with Ollama. That is why I switched to Qwen. But the performance drop is huge.” “GPT-oss 20b was kinda useless i found, but I still use it for direct question / answer stuff rather than reasoning.” “Mine just ate 2.88M tokens in half an hour for simple task like playing a song in youtube, volume change and search for Gemini API key usage.” “Been running openclaw agents for a while and kept hitting the same wall. No visibility into cost until the bill arrived.” “I was getting tired of how fast OpenClaw burns through money with API tokens, so I switched to running everything locally with Ollama.” “I replaced my $3900/year sales stack using Claude Code and OpenClaw in 4 days. It now costs me $40/mo to run.” “The OpenClaw config file was world-readable (644 permissions). Anyone else on the system could see tokens and settings.” “I’ve been running OpenClaw for a few months and I’m increasingly worried about prompt injection. Content filtering didn’t work — is anyone else thinking about this?” “That’s exactly what I do and it works really well. I have proxmox on my homelab, and create a Linux VM for each agent so there’s clear separation” “Because of the precautions I put in place, the bot refused, checked the Telegram ID, and shut the conversation down.” “Biggest win was having a simple ‘supervisor’ agent that enforces a checklist (sources, confidence, next actions) and a hard stop when tools fail.” “Now I can chat with either agent, check memory, browse skills, and inspect tool calls from the same app.” “It is funny that you think the job is done, but when you open your claude app you found a random approval stall everything and it is already in your allowed list. Just straight frustration.” “when I asking OpenClaw to do stuff, it starts and then suddenly stops giving feedback, I have no idea if it’s done, or there were issues.”

Working setup as the precondition Credit given to the path in

“I’ve been running OpenClaw as my personal AI assistant that works for my whole immediate family.” “Had my openclaw set up and running within the day and im positive I probably would have given up halfway through if I didnt have the guide.”

Install as the decision point Setup ability as the gate on who benefits

“failed at installing openclaw 3 times. gave up and moved to Run Lobster (OpenClaw).” “My problem wasn’t OpenClaw itself. It was my co-founder / non-technical members of my team who couldn’t deploy without a lot of hand holding.”

21

Conference’17, July 2017, Washington, DC, USA

Ma et al.

Table 6: Human validation in the 400-post codebook-development sample. Blank review cells retain the initial LLM label and nonblank cells use the coder’s replacement. Human–human metrics compare the two coder decisions; LLM–human metrics treat coder decisions as references and initial LLM labels as predictions. Macro-P and Macro-R are macro-averaged precision and recall. Panel A: Human–human exact agreement and pooled Cohen’s 𝜅 Coding target

𝑛 posts

Exact agreement

𝜅

Relevance Agent aspect Human value Value fulfillment User outcome

400 328 328 328 328

94.75% 82.01% 90.85% 98.78% 92.99%

.906 .797 .897 .979 .918

Panel B: Initial LLM label versus coder decisions Coding target

𝑛 decisions

Accuracy

Macro-P

Macro-R

Macro-𝐹 1

Weighted-𝐹 1

Relevance Agent aspect Human value Value fulfillment User outcome

800 671 671 671 671

.948 .900 .943 .994 .961

.948 .895 .622 .747 .599

.911 .870 .661 .746 .609

.924 .863 .635 .747 .604

.946 .894 .937 .993 .955

22

Record · ID 1006881 · SHA-256 5360b34f98a70911
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.