ConceptioArchivearXiv CS
arXiv CSopen access

Gricea: An Open Science Platform for Conversational AI Research

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

arXiv:2609.22039v1 [cs.HC] 18 Sep 2026

Gricea: An Open Science Platform for Conversational AI Research Nikhil Sharma

Yunlin Gong

[email protected] Johns Hopkins University Baltimore, MD, USA

[email protected] Johns Hopkins University Baltimore, MD, USA

Xinyang Cheng

Ziang Xiao

[email protected] Johns Hopkins University Baltimore, MD, USA

[email protected] Johns Hopkins University Baltimore, MD, USA

A Represent and preserve

B Inspect, change, redeploy

C Build collective knowledge

Researcher

Another Researcher

Research Community

RQ: How does sycophantic AI change trust?

RQ: What happens if I change the context to controversial topics?

Study 1 · full study artifact

Shared executable research artifacts for researchers and agents to build upon.

Studies build on studies

Inspect Study 1

Study Flow One condition per participant

Read the design and experience the participant task.

Study 4

Study 2

Build on Study 1 → Study 2 Change only the RAG database

Chat task · Task Flow

Study 1

Study 7

Study 5

Study artifact

General topics

Controversial topics

Study 3

Retain everything else

Study 6

Configured values + preserved defaults

Deploy and share Study 2 A comprehensive representation of studies based on the CAI design space

A new study, with the rest of its design held constant.

+ Researchers + AI agents New findings accumulate into collective knowledge.

Configure: prompts, models, RAG, interface, flow, assignment, timing, measures.

A shared representation for researchers and AI agents · design, inspect, execute, preserve, reuse

Figure 1: A shared research foundation for conversational AI. A: Gricea represents a study through connected Study Flow and Task Flow definitions, preserving both configured choices and unchanged defaults in an inspectable, executable artifact. B: Other researchers can inspect and extend the artifact by changing its grounding context while retaining the remaining study design; The updated study can be deployed and shared. C: The research community and AI agents can build on shared executable artifacts across successive studies, accumulating new findings into collective knowledge

Abstract

conversational task behavior in. In a replication study using Gricea, we replicated configurations 93% of eligible CUI 2026 papers; while also flagging missing information in 96% of papers that hinder faithful replication — further motivating Gricea’s need. In a user study, researchers and practitioners from diverse backgrounds successfully constructed runnable studies addressing various open-ended research questions. Together, these findings demonstrate Gricea’s support for constructing, reproducing, and extending CAI studies

We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facing systems, and 1

Sharma et al.

through shared research artifacts, enabling cumulative knowledge building through open science.

Furthermore, all of these conditions can interplay with each other creating a vast and diverse design space. A shared representation must make the design choices both at the study level and at the task configuration level explicit; preserving a comprehensive representation of based on the CAI design space (Section 3). Even without the overhead of a shared representation, conducting these studies has high costs beyond just recruiting and compensating participants. Researchers must coordinate study procedures, task interfaces, agent behavior, and data collection, often by integrating survey tools, custom interfaces, model services, and deployment infrastructure. Commercial conversational systems offer limited experimental control, while changes to their interfaces, retrieval policies, or models can alter the conditions under investigation. Building and maintaining custom systems therefore demands time and engineering expertise that can constrain both who conducts conversational AI research and which questions they pursue. Therefore, a successful shared representation artifact must not only find a scalable way to represent the vast and diverse design space but also reduce the costs of building these custom systems while making the artifact a by-product of the process rather than an additional cost to the researchers. In this paper we present Gricea, a platform that represents conversational AI study designs as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Researchers visually author study procedures and configure complex interactive task conditions within a platform that supports participant-facing execution and instrumentation. The representation researchers inspect is also the definition the runtime executes, connecting experimental assignment and task behavior to the participant experience. Gricea reduces the technical overhead of constructing controlled studies and preserves immutable study versions as reusable templates. Other researchers can inspect the original design, reproduce its conditions, and extend it to new questions, making individual studies reusable resources enabling cumulative knowledge (Figure 1). Gricea’s design was informed by a formative analysis of 57 papers on conversational AI (Section 3). The resulting desiderata guided two complementary representations. Study Flow represents the experimental procedure, connecting surveys, tasks, and betweenand within-subject structures. Task Flow represents behavior within each task, including agent configuration and the participant-facing interface. Researchers can combine these elements to implement existing designs and complex configurations involving multiple agents, multiple users, and customized interfaces. We evaluated Gricea with 𝑁 = 10 researchers and practitioners from diverse disciplinary backgrounds, who completed an assisted authoring walkthrough before independently designing studies around their own research questions. Participants implemented valid, runnable studies that varied across procedures, interfaces, models, surveys, and outcomes, and the no-code interface reduced barriers for researchers without systems backgrounds. We also examined CUI 2026 full papers, none of which informed Gricea’s design. Of 29 eligible papers, we replicated designs from 27 as executable artifacts: 10 completely and 17 partially. Together, these evaluations demonstrate support for authoring diverse studies and reconstructing published designs for inspection and reuse.

CCS Concepts • Human-centered computing → HCI design and evaluation methods; User studies; Empirical studies in HCI; Collaborative and social computing; Field studies; Usability testing; User models; Natural language interfaces; Web-based interaction; Collaborative interaction; User interface management systems; Interaction design process and methods; • Information systems;

Keywords Conversational AI, Controlled Studies, Human Subject Studies, Research Platform Platform

1

Å Tutorial

Introduction

Conversational AI systems built on large language models are becoming a default interface for information access, communication, and everyday work [15]. As these systems become everyday infrastructure, they shape how people learn, create, collaborate, make decisions, and form relationships with artificial agents [48, 51]. Understanding these changes requires empirical research on human behavior and experience alongside evaluation of the systems people use. Such evidence is essential to the design, governance, and deployment of conversational AI. Prior human–AI interaction research shows that outcomes depend on various factors such as model performance, how systems communicate uncertainty, present evidence, structure initiative, support verification, and distribute control [5, 51, 85, 98]. Misinformation exposure, overreliance, persuasion, and uneven information access emerge through the interplay of model behavior, interface design, and user context [77, 86, 87, 89, 100]. Studying these effects requires examining how people interact with conversational AI while pursuing goals in specific tasks and contexts; researchers use Controlled human-subject studies to isolate effects of different design choices [60]. Building cumulative knowledge using controlled studies about conversational AI requires a shared frame of reference for what participants experienced, how conditions were configured, and how outcomes were measured. When those configurations and materials are difficult to recover, subsequent researchers must reconstruct the interaction before they can reproduce or extend a study. Differences in interface behavior, agent configuration, or procedure can then obscure what was preserved and what changed. Open science therefore requires preserving these methodological choices in an inspectable, reusable form alongside the findings [3, 37]. A shared representation gives researchers a common starting point for reproducing conditions, systematically varying design choices, and relating new findings to prior work. However, representing these conditions is challenging because design choices span several connected parts of a study: the interface and actions available to participants, agent behavior context, task, modality, participant population, and experimental procedure. 2

Gricea: An Open Science Platform for Conversational AI Research

This paper makes three contributions. First, we contribute a shared, executable representation of conversational AI studies that connects study procedure, task behavior, participant-facing conditions, and instrumentation, enabling researchers to inspect, reproduce, and extend experimental designs. Second, we present Gricea, a configurable no-code research platform that operationalizes this representation through visual authoring, participant-facing execution, immutable study versions, shareable configurations, and reusable community templates. These mechanisms lower technical barriers while making individual studies available as resources for subsequent research. Third, we contribute empirical findings from an authoring study with researchers and practitioners from diverse disciplinary backgrounds and reproductions of CUI 2026 study configurations, demonstrating Gricea’s support for diverse research questions and published study designs while identifying remaining needs for guidance, validation, and workflow support.

Human–AI research platforms bring agent behavior into this experimental infrastructure. Deliberate Lab combines no-code experimental stages, human and LLM participants, agent mediators, and cohort management for studying human–AI group dynamics [76]. For conversational AI studies, the experimental condition depends on more than the sequence of study stages or an agent’s configuration. To represent a broad set of CAI studies, we conduct a formative study to uncover the design space, allowing Gricea to extend existing efforts to a more general reusable research infrastructure through a coupled representation of study procedure and conversational task behavior.

2.2 2

Related Work

Gricea builds on research that broadens participation in online studies, preserves methods as reusable research resources, and makes conversational systems configurable through higher-level representations. These efforts address complementary requirements for conducting research and building on its findings. Gricea brings these requirements together through a shared representation of study procedure and conversational task behavior. The representation used to design a study governs its deployment and data collection, preserving an inspectable research artifact as part of building and running the study.

2.1

Infrastructure for Open Science

Open-science infrastructure supports the preservation and exchange of research materials across teams. Foster and Deardorff [37] describe infrastructure for project organization, collaboration, file versioning, and registration, making materials easier to preserve and share. However, accessible materials must also be sufficiently specified and connected for others to use them. Iarygina et al. [46] identified obstacles to computational reproduction among CHI papers that shared data and analysis code, illustrating the difference between making resources available and enabling others to reproduce the work. For conversational AI studies, researchers need to understand how the procedure, interface, agent behavior, and materials jointly determined what participants experienced. Executable research representations connect methodological specification to implementation. Aguilar et al. [3] represent experiment components through automation code and digital documentation, including infrastructure, data collection, analysis, and management. Nobre et al. [69] support inspecting participant behavior through interaction provenance and replay, while Cutler et al. [25] connect study specification, execution, analysis, and dissemination within a browser-based framework. Subsequent LLM integration preserves conversation history and supports replay of chatbot interactions [44]. Gricea builds on this connection between executable methods and inspectable interactions through a shared representation of conversational task logic and the surrounding experimental procedure which is also the same representation that the runtime executes. Reporting frameworks and agent-native research artifacts further clarify what must survive publication. Feuerriegel et al. [35] call for explicit documentation of LLM use, including model versions, prompts, and configurations, while Liu et al. [63] connect scientific logic, executable code, exploration traces, and evidence so that humans and AI agents can understand and build on research. Gricea integrates artifact preservation into the development and execution of participant studies. The configured procedure, prompts, materials, and interaction logic constitute the study that is deployed, so researchers do not need to reconstruct a separate artifact after implementation. Sharing that representation makes the implemented method available for inspection, reconfiguration, and reuse within the same environment.

Infrastructure for Crowdsourced and Online Studies

Online research infrastructure has expanded where studies can be conducted and who can participate. Kittur et al. [52] examined how task design and quality checks influence crowdsourced judgments, while Reinecke and Gajos [79] used personalized feedback to attract uncompensated participants and evaluated online replications of laboratory studies. Subsequent comparisons examined differences in participant diversity and data quality across recruitment platforms [73, 74]. These efforts demonstrate the influence of recruitment in the quality of the online studies. Conducting experiments with these participant populations also requires infrastructure for implementing tasks, assigning conditions, coordinating interactions, and collecting responses. Reusable experiment frameworks address these requirements by providing components that researchers can adapt across studies. oTree supports browser-based experiments through Python and HTML, including a library of reusable game templates [16]. Empirica supports configurable experimental designs and reusable protocols for realtime group experiments [4], while jsPsych enables researchers to construct behavioral experiments from reusable plugins and contribute new tasks to a community ecosystem [28]. Across these systems, reusable components allow the implementation work behind one study to support subsequent studies, reducing the effort required to develop and extend experimental designs. 3

Sharma et al.

2.3

Infrastructure for visual programming of CAI studies

these codes to identify recurring outcome areas, manipulation dimensions, procedural structures, and infrastructural demands. As part of this analysis, we also examined how papers visually represented their study designs. Papers used staged diagrams, branching structures, and flowcharts to communicate condition assignment, task sequences, and follow-up measures. These representations make explicit how study logic structures the activities participants experience; motivating Gricea’s support for executable visual representations: researchers should be able to design, inspect, communicate, and run a study through the same representation, without reconstructing its logic manually in code.

Visual and declarative systems make computational choices accessible through representations that users can inspect and modify. Wu et al. [95] support composing and debugging multi-step LLM chains, while Arawjo et al. [8] support systematic comparison of prompt and model variations through a visual dataflow environment. Cai et al. [14] allow users to edit a proposed workflow before an LLM executes it, and Feng et al. [34] support structured specification and testing of model behavior within interface design work. Conversational application platforms also provide deployment environments, and live-traffic experiments [39–41]. Gricea brings this control over computational behavior into the representation of a human-subject study, connecting experimental assignment, participant interaction, and measurement. Research-oriented representations bring methodological choices into these abstractions. Jun et al. [50] allow researchers to declare study designs, assumptions, and hypotheses for statistical analysis. Yao et al. [97] provide an experiment configuration language and controls over collaborative environments, agent perception and action, and synchronized interaction logs, while Zhang et al. [99] support configurable human–AI teaming environments and feedback collection. Gricea separates and couples study procedure and conversational task logic through the same graphs that drive execution. Researchers can examine how a procedural decision changes the participant-facing condition and preserve that relationship when a study is shared, reproduced, or extended. The formative analysis that follows identifies the recurring study requirements that informed this design.

3

3.1

What studies on conversational AI investigate

Studies on conversational AI investigate how configured assistants shape human behavior, judgment, and experience within particular task settings. In our corpus, these settings included co-writing, conversational search, learning, dietary recommendation, and daily planning and reflection. Participants composed text with generated suggestions[48], explored information through dialogue, received personalized recommendations[59], or revisited plans across sessions[1]. Each task establishes what participants are trying to accomplish and the role the assistant plays in that activity. Within these settings, the outcomes of interest are similarly broad. Prior work examines trust, reliance, persuasion, misinformation response, privacy behavior, writing quality, learning, and collaboration [48, 85–87, 89, 100]. What links these studies is not a single application domain, but a common methodological concern: how a conversational system condition shapes what users believe, do, and produce over time. A platform for this area must therefore support both configuring the participant-facing interaction and collecting the evidence needed to examine its outcomes, including self-reports, behavioral traces, and task outputs [10, 57, 61].

Formative Analysis: The Science of Conversational AI Studies

To scope the infrastructural requirements for Gricea, we conducted a formative design space analysis of papers on conversational AI systems. Our goal was to identify recurring patterns across prior work: what studies on conversational AI investigate, what they manipulate, how they are typically conducted, what technical demands those choices create, and what forms of infrastructure existing systems already provide. From these recurring patterns, we identified what researchers need to specify and control, and which details must remain inspectable for others to reproduce and build on a study. These requirements motivate five design desiderata for Gricea’s study representation and authoring environment.

3.2

The manipulation space of conversational AI studies

Our formative analysis shows that studies on conversational AI manipulate far more than prompts or underlying models. The true experimental object is a configured interaction condition: the combination of agent behavior, interface, context, and procedure that defines what participants experience. For example, a study of chatbot relationship framing varied both the agent’s self-description and the visibility of conversation history across sessions [23]. Therefore, representation of studies requires specifying both what researchers manipulate and the surrounding configuration they hold constant. We synthesize recurring configurations into six interacting dimensions: Interface Condition, Agent Condition, Context & Grounding, Task & Modality, Domain & Audience, and Study Procedure. Table 1 summarizes their configurations and infrastructural implications. These dimensions connect what researchers configure, what participants experience, and how the study is conducted. Across these dimensions, understanding a study requires inspecting how its procedure, interface, agent behavior, and contextual information jointly produce the participant experience [23, 48, 85]. Researchers need to distinguish the choices that define a condition from those held constant, and to understand how those choices are

Analysis Procedure. We began by collecting 100 candidate papers using keyword combinations around conversational AI, agent, or LLM, together with terms related to users, humans, and studies, across venues and repositories such as CHI, UIST, CUI. We then filtered this set to 57 papers that centered participant-facing conversational or agentic AI systems and provided sufficient detail about the study design, system configuration, or evaluated interaction condition. For each paper, the research team coded the study type, focus area, participant count, independent and dependent variables, between- and within-subject structure, procedural stages, system or pipeline components, analysis methods, and the overall structure of the study procedure. The research team reviewed and clustered 4

Gricea: An Open Science Platform for Conversational AI Research

Dimension

Examples of study configuration

Implications for study infrastructure

Interface Condition

Layout and workspace organization; citation and provenance display; highlighting; progress guides; verification prompts; structural overlays; editability of shared artifacts; time-synchronized guidance and conversation navigation [17, 19, 20, 49, 68].

Represent what participants can see, inspect, and change. Preserve presentation and available actions as part of the study condition, including how evidence, guidance, and task progress are displayed.

Agent Condition

System instructions; persona or role framing; stance; response style; initiative; clarification behavior; trust-adaptive interventions; tool-use policy; response delays and waiting cues; coordination of model components [30, 48, 56, 58, 65, 66].

Make agent behavior, permitted actions, timing, and coordination explicit. Distinguish the roles participants encounter from the model components that implement those roles.

Context & Grounding

Retrieved documents; retrieval policy; context filtering; citation constraints; external evidence injection; opposing or balanced context; task context; interaction history; persistent memory and its inspection or editing [20, 55, 85, 86, 96].

Specify what information the agent receives, what participants can inspect or change, and how context persists across turns or sessions. Distinguish information available to the agent from history displayed to participants and records retained for research.

Task & Modality

Search, writing, tutoring, coding, annotation, voice, multimodal interaction, direct artifact manipulation, and conversational control of haptic feedback [12, 36, 67, 76, 97].

Connect different participant activities and modality-specific inputs and outputs to a shared study procedure. Support taskspecific actions and outcome collection without rebuilding the surrounding study infrastructure.

Domain & Audience

Topic domain; expertise; prior beliefs and experience; participant population; targeted community; high-stakes versus low-stakes settings; language and dialect; accessibility and care needs [26, 31, 36, 38, 71, 77].

Document the task setting, study materials, and relevant participant characteristics separately from assigned conditions. Preserve the context needed to interpret findings and adapt a study for another population or setting.

Study Procedure

Pre-task elicitation; condition assignment; branching; withinor between-subject structure; counterbalancing; post-task measures; longitudinal follow-up; scheduled check-ins and cross-session continuity [1, 13, 23, 48, 75, 76].

Make assignment, task order, branching, measurement timing, and cross-session continuity explicit and authorable. Preserve the procedure as executable logic that can be inspected, reused, and extended.

Table 1: Recurring dimensions of conversational AI study configuration, synthesized from the formative analysis. Examples include experimental manipulations, fixed settings, and participant characteristics. The infrastructure implications identify what must remain explicit to design, inspect, reproduce, and extend a study.

implemented. A shared study representation should preserve these relationships so that both the original research team and subsequent researchers can inspect the design, reproduce its conditions, and make deliberate changes when extending it.

3.3

The study procedure determines how participants move through the study: the instructions they receive, the condition they encounter, when branching occurs, and when measurements are collected. The interactive runtime determines what participants experience within a task: what they can see and do, what information the system receives, how it responds, how the interface adapts, and which actions and responses are logged. Together, these levels specify both how participants encounter a condition and how that condition operates during the interaction. Across these studies, condition assignment and information collected before the task can configure the agent’s behavior and available context, while participants’ actions within the task can determine subsequent interaction paths and when post-task measures are collected [1, 49, 59]. A research platform must therefore represent these dependencies explicitly, connecting the assigned condition, the interaction participants experience, and the evidence collected about its outcomes. This requirement motivates Gricea’s coupled support for Study Flow and Task Flow, allowing researchers to specify the surrounding procedure and within-task behavior as connected parts of the same study.

The Procedural Anatomy of Conversational AI Studies

The analysis also shows that studies on conversational AI combine system configurations with structured research procedures. A study may begin with consent, instructions, and pre-task elicitation, proceed through condition assignment and an interactive task, and conclude with post-task measures, reflection, or interviews. Prior work combines these elements in co-writing tasks with pre- and post-task measures [30, 48], conversational search with turn-level behavior and post-task attitudes [9, 85], and daily planning and reflection that link repeated conversations to daily surveys and an exit interview [1]. A study on conversational AI therefore couples a study procedure with an interactive runtime. 5

Sharma et al.

3.4

Synthesis of Infrastructural Challenges

behavior, and measures correspond to the intended experimental design before publication.

Executing studies across this broad manipulation space is challenging because the experimental condition is distributed across many components that must be built and controlled in tandem. Researchers often need to stitch together survey tools, custom interfaces, backend orchestration, model and retrieval pipelines, assignment logic, deployment infrastructure, and fine-grained behavioral logging [76, 97]. Even when the intended manipulation is conceptually straightforward, implementing it as a controlled, reproducible participant experience requires substantial engineering effort, creating barriers for researchers without the technical expertise or resources to build and maintain this infrastructure. Maintaining experimental control also requires researchers to specify how interface affordances, retrieval behavior, model configuration, and system defaults jointly produce the participant experience [76, 97]. Reliance on commercial systems adds dependencies whose behavior may not be fully exposed or preserved across versions, with documented changes in model behavior showing why the same prompt and model name do not establish an equivalent condition [18]. Study infrastructure must therefore make the configuration and its dependencies inspectable alongside the record of what participants actually encountered. When instructions, prompts, task materials, interface behavior, and procedural logic remain embedded in one-off implementations, publishing the findings does not necessarily preserve the condition needed to reproduce or extend the study. Subsequent researchers must reconstruct how these elements were connected before they can determine whether a new implementation reproduces the original condition or introduces consequential differences. Preserving the configured study as an inspectable, executable artifact provides a shared frame of reference for comparing implementations, adapting procedures, and building on prior work [76, 97]. Hence, such infrastructures must reduce the effort of constructing studies along with preserving the configuration necessary for collective knowledge to accumulate.

4

D2: Coupled Support for Study Flow and Task Flow. The platform must support both the overall study procedure and the behavior of the interactive task, while keeping their roles distinct. Study Flow specifies how participants move through instructions, condition assignment, tasks, and measures, whereas Task Flow specifies how participant actions, agent responses, and branching shape progression within a task. Prior studies connect information collected before an interaction to the agent’s behavior and assess the resulting experience through subsequent measures [59]. Researchers must therefore be able to author and modify each flow separately while specifying how information enters a task, when the task ends, and how its outputs connect to subsequent study stages.

D3: Lower Technical Barriers to Controlled Study Authoring. The platform must reduce the amount of bespoke engineering required to build and deploy controlled studies on conversational AI [76]. Our formative analysis identified staged diagrams, branches, and condition flows as recurring ways of representing study logic. The authoring model should build on these representations, enabling researchers without systems backgrounds to specify and connect study components through a no-code interface. Reusable support for participant interfaces, model integration, deployment, and data collection should allow researchers to move from a study concept to an executable artifact without reconstructing the surrounding infrastructure from scratch; Lowering the barriers for conducting these studies allowing for a broader range of researchers to contribute.

D4: Reproducibility Through Inspectable and Reusable Artifacts.

Design Desiderata

The platform must preserve authored studies as explicit research artifacts rather than leaving critical details embedded in transient setup steps or one-off implementations [25]. Each study version must retain its procedure, task logic, prompts, model and context settings, and participant-facing materials in an executable form. Recorded interactions must remain linked to the corresponding version so that researchers can inspect both the authored condition and the experience participants encountered. Another researcher should be able to inspect, redeploy, adapt, and build on the artifact, with changes made explicit across versions. Preserving this continuity provides a shared frame of reference for reproducing studies and building cumulative knowledge.

Building on the formative analysis (Section 3), we derive five design desiderata for infrastructure that supports the design, execution, and reuse of conversational AI studies.

D1: Explicit Representation of Study Conditions. The platform must represent the configured interaction condition shown to participants rather than only isolated prompts, screens, or model calls. Prior studies manipulate agent behavior and participant-facing interaction support while also specifying shared interfaces, controls, and contextual information across conditions [20, 48]. The infrastructure must therefore make the relationships among procedure, interface, agent behavior, and context explicit, allowing researchers to distinguish what is manipulated from what is held constant. Researchers must be able to inspect how these choices shape what participants see, what actions they can take, and how the system responds, without reconstructing the condition from separate implementation details. This representation must support checking whether the implemented assignment, task

D5: Extensibility Across Tasks, Modalities, and Study Settings. The platform must remain extensible as the design space of conversational AI continues to expand. Studies may involve text chat, coding, multimodal interaction, voice, longer-running procedures, or community-facing workflows. New task types, modalities, and 6

Gricea: An Open Science Platform for Conversational AI Research

study settings should be supported through extensions that reuse the platform’s mechanisms for study execution and data collection, rather than requiring reimplementation of the surrounding system [97]. These extensions must remain configurable within the study representation, allowing researchers to inspect, version, and reuse the resulting studies through the same authoring environment.

5

Logic and Condition Control. Condition control is represented through graph structure, node configurations, and variables. Branching, variable assignment, weighted randomization, and within-subject routing determine which conditions participants encounter and in what order. Assignments and condition orders are retained in the participant’s execution state. Variables connect assigned conditions to task configurations, making explicit how experimental assignment shapes the participant-facing interaction. Researchers can inspect the parameters deliberately varied across conditions alongside the settings held constant within the same authored study.

Gricea: A Platform for Configurable Conversational AI Studies

Gricea represents controlled conversational AI studies as configurable, executable research artifacts. Its architecture consists of three decoupled layers that operate on this shared representation: a researcher-facing visual authoring environment, a participantfacing execution runtime, and a publication layer for versioning, reuse, and community sharing. Each artifact specifies the experimental procedure and the interactive conditions participants encounter within it. This representation allows complex study conditions to be authored, executed, preserved, and reused within a unified framework rather than reconstructed through ad hoc infrastructure. Gricea operationalizes study designs as directed graphs. A Study Flow graph governs participant progression through the experiment, while nested Task Flow graphs specify the behavior of each participant-facing interactive condition. Publication serializes these graphs, their configurations, and their assets into an immutable study version that the runtime engine executes directly.

5.1

5.2

Researcher Authoring Environment

Researchers author studies visually through no-code, node-based canvas editors for both Study Flow and Task Flow (Figure 2). The authoring model builds on the staged procedures, condition branches, and task entry points identified in our formative analysis (Section 3), allowing researchers to express experimental logic in the same graphs that drive execution. Researchers can construct the procedure and interactive condition within one environment while retaining separate control over each. At the study level, researchers add and connect instructions, elicitation stages, branching logic, randomization, within-subject blocks, and task entry points. At the task level, they configure each task’s Task Flow and participant interface, including system prompts, retrieved context, model parameters, supported tool settings, interface layout, scaffolds, and modality-specific settings. Study-level variables connect procedural decisions to the task configuration that participants encounter. The platform separates iterative authoring from participant deployment. Researchers can run a participant-facing preview and revise an editable draft before publication validates and locks the authored graphs, prompts, configuration states, and runtime bindings into a published study version. The participant runtime executes that version while further development proceeds in new drafts, preserving the exact authored condition deployed to participants independently of subsequent changes. Researchers can inspect assignment and routing in Study Flow, examine their effects on interaction in Task Flow, and preview the participant experience. Publication and execution use this same definition, avoiding a separate translation into bespoke software. Preview complements structural validation by letting researchers check the represented study against their intended design.

Core System Abstractions

Gricea represents authored studies through two connected levels of executable graphs. Study Nodes form the procedural Study Flow, while Task Nodes form Task Flows embedded in its interactive stages. Logic Nodes provide branching, variable assignment, and condition control within these flows. Publication validates the authored definition and makes a specific Study Version available as a Published Study. Templates and Community Document Collections provide reusable components and grounding materials, while Participant Analytics summarizes the evidence produced during execution. Study Flow. The Study Flow graph encodes the procedural structure through which participants navigate. Its nodes represent participant-facing stages, including instructions, surveys, task entry, annotations, and completion, alongside control operations for branching, randomization, and within-subject ordering. During execution, the engine follows the authored transitions using assignment results and participant-specific state, advancing through control operations until it reaches a stage requiring participant interaction or study completion.

5.3

Participant Runtime and Instrumentation

The participant runtime combines browser-based rendering with server-side execution of the published study. It resolves the Study Flow for procedural progression, invokes the associated Task Flow upon task entry, and binds participant interactions to the authored graph logic. The browser renders the interface defined for the current stage or task, submits participant actions to the runtime, and applies the resulting updates. The published artifact therefore governs both how participants progress through the study and how the interactive condition responds to their actions. The runtime supports chat, web-based search tasks, voice interaction, human–AI coding, annotation workflows, image generation,

Task Flow. The Task Flow graph encodes the runtime behavior of the interactive condition. Its nodes represent operations such as LLM inference calls, retrieval and context injection, real-time voice interaction, loops, participant input, and updates to interface components. The engine executes these operations according to the graph, coordinating model calls and interface changes with participant actions. When the task completes, control returns to the surrounding Study Flow, where information produced during the interaction can inform subsequent stages and measures. 7

Sharma et al.

Figure 2: The Gricea authoring interface, illustrating how researchers can assemble multi-stage study procedures (left) and independently configure agent pipelines and interface scaffolds for specific conversational tasks (right). and split-view configurations. These tasks use shared mechanisms for execution, state management, and event collection, allowing researchers to vary the participant-facing experience while preserving a consistent connection between the study configuration and the resulting data. The instrumentation layer records timestamps, clicks, focus events, scroll depth, input timing, text-edit counts, and taskspecific interaction traces according to the study’s collection settings. Where enabled, additional capture includes voice recordings, masked browser-session replay, screen recordings, file-upload traces, and eye-tracking data. These records are associated with the participant session and its immutable Study Version, allowing researchers to interpret behavior in relation to the exact configured condition participants encountered.

5.4

The publication layer supports reuse of both complete studies and their constituent resources. Researchers can publish Community Templates containing whole studies or selected nodes and flow segments, share Community Document Collections used for RAG grounding, and release de-identified dataset assets. Sharing these resources preserves not only the study’s outputs, but also its executable design and the materials used to construct the conversational condition. A peer researcher can therefore inspect the underlying graph logic and configurations, experience the study live as a participant, and fork the artifact to run an independent replication. The fork provides an editable draft for changing conditions, tasks, or measures without altering the published source. A researcher extending the study can change a selected parameter while retaining the surrounding procedure and materials, then publish and deploy the revised artifact. The source and revised artifacts provide a shared frame of reference for examining what was preserved and what changed, allowing subsequent research to build directly on existing experimental designs.

Publication, Reuse, and Community Workflows

Gricea treats reproducibility as an intrinsic property of the study artifact itself. Each published study is preserved as an immutable version containing its prompts, task parameters, graph structures, assignment logic, and participant-facing materials. These configurations remain inspectable after deployment, and researchers can distribute published studies through shareable links, making the executable study available alongside its written description.

5.5

Architecture and Extensibility

Gricea separates its visual authoring client, execution runtime, and data and publication services so that new capabilities can be added without rebuilding the surrounding study infrastructure. 8

Gricea: An Open Science Platform for Conversational AI Research

CONVERSATIONAL AI + SURVEYS

REALTIME VOICE AI + DRIVING

CONVERSATIONAL CODING

MULTI USER AND MULTI AGENT

Figure 3: Participant-facing interfaces configured in Gricea: conversational AI with integrated surveys (top left), real-time voice interaction alongside a driving task (top right), conversational coding with an editable code workspace (bottom left), and a participant interacting with two conversational agents (bottom right). Developers extend task behavior through new node types, each with configuration definitions and an execution handler. Participantfacing components similarly declare the properties researchers can configure and the interaction events they produce. These additions integrate with the existing Study Flow and Task Flow model. The architectural objective is to preserve a common study representation as the supported tasks and modalities expand. Extensions for richer multimodal interactions, longitudinal deployments, and community-facing workflows can reuse the same mechanisms for procedural control, publication, and instrumentation, keeping the resulting studies configurable, inspectable, and reusable.

construct their own studies. Combining replication and usability studies follows prior work on toolkit evaluation, which examines what systems enables and it’s usability [54, 70]. Our evaluation addresses four questions. RQ1: Can Gricea replicate conversational AI study designs beyond those included in our formative analysis? RQ2: Can researchers from diverse backgrounds implement their intended study designs with Gricea? RQ3: Can Gricea represent and execute a diverse range of conversational AI studies? RQ4: What frictions remain in translating research intent into an executable study? We first conducted a replication study of CUI 2026 papers outside our formative corpus to examine whether Gricea could represent and execute full study configurations or their conversational portions, addressing RQ1 and providing evidence of expressivity for RQ3 (Section 6.2). We then conducted a usability study in which researchers and AI practitioners authored studies addressing their own research questions, examining their ability to implement intended designs (RQ2), the range of studies they constructed (RQ3), and the frictions they encountered during authoring (RQ4).

6 Evaluation 6.1 Evaluation Rationale and Questions Gricea is intended to support the replication of existing conversational AI studies and the construction of new studies through a shared, executable representation. We evaluated whether the platform provides the features needed to implement published study designs beyond those included in our formative analysis, and whether researchers from diverse backgrounds can use those features to 9

Sharma et al.

6.2

6.3

RQ1: Replication of Published CUI 2026 Studies

Participants

We recruited 𝑁 = 10 participants from diverse disciplinary backgrounds and roles, including researchers, students, faculty, and industry practitioners. The participant summary is provided in the Table 3. Of the 10 participants 5 self-identified as Male and 5 self-identified as female. The median age of participants was in the range of 25-34. There were 4 PhDs, 2 Master students, 2 professionals, 1 professor and 1 undergraduate student in our sample. Participants varied in their prior experience with controlled studies and in the kinds of conversational AI questions they had previously explored or hoped to explore. This diversity was intentional since Gricea is meant to onboard researchers from diverse backgrounds by lowering the barriers to conduct user studies.

To evaluate Gricea’s support for published conversational AI studies, we examined 37 CUI 2026 full papers, of which 29 reported participant-facing human-subject studies. The full corpus is listed in Appendix A.2. These included live conversations, prerecorded conversational stimuli, expert annotation and rating, multimodal interfaces, group interaction, and embodied agents [11, 42, 81, 92, 94]. The research team extracted study details from the papers and supplementary materials, attempted reconstruction in Gricea, and recorded replication failures. We assessed replication success by whether Gricea provided the features needed to represent and execute the study procedures, participant-facing interfaces, and conversational task behavior. Complete replication covered the full study configuration, whereas partial replication covered the conversational portions that could be implemented when missing source materials or external dependencies prevented reproduction of the full study. Studies with full replication can still have missing details from studies such as missing participant facing instructions but they do not block representation of the study configurations and CAI tasks. We recorded unavailable information separately from requirements for external hardware or systems to distinguish reporting gaps and scope boundaries from the study features supported by Gricea.

6.4

Study Design and Procedure

Each session was designed to assess both onboarding and openended study authoring. The procedure drew on a common pattern in toolkit and platform evaluation: a structured task that helps participants develop the system’s basic mental model, followed by an open-ended task that reveals how they apply the platform to questions that matter to them [8, 98]. Sessions consisted of four phases: pre-task interview, assisted authoring walkthrough, openended think-aloud authoring, and post-task reflection. The total study lasted 90 minutes and participants were paid 30$ through amazon gift cards. Throughout, we recorded the participant screen and audio after obtaining their informed consent. We began each session with a brief introduction followed by a pre-task interview on their background, prior experience with human subject studies, challenges encountered in running human subject studies, and research questions they were interested in the area of Conversational AI. In the assisted authoring walkthrough, participants implemented a fixed research question: How does the stance of an AI assistant on controversial issues affect users’ perceived trust? During this phase, the facilitator provided guidance on the platform features, while participants retained control of the interface and performed the authoring themselves. In the open-ended phase, participants were asked to design a conversational AI study of their own choosing using Gricea. They were instructed to think aloud as they worked. After the authoring tasks, participants completed a post-task survey and took part in a semi-structured interview. These instruments focused on perceived usability, expressivity, reproducibility, likely time savings relative to current workflows, and adoption potential. The interview further probed where participants felt confident, where they felt stuck, and what forms of support would make the platform more useful in their own work.

Replication outcome: Using Gricea, we successfully replicated full study configurations or their conversational portions from 27 of the 29 eligible papers: 10 completely and 17 partially. Gricea provided the features needed to represent and execute the replicated procedures, interfaces, and conversational tasks, covering prompt-based manipulations, model and stance comparisons, roleplay conversations, fixed-media evaluations, repeated sessions, and speech-versus-typing tasks. Missing source materials and external dependencies limited complete replication, while unavailable interview and survey materials prevented replication of the remaining two studies. These results demonstrate that Gricea supports executable study designs beyond those used to inform its representation. Reporting Gaps and Their Impact on Replication: We found incomplete reporting or unavailable materials in 28 of the 29 eligible papers, although not all omissions prevented replication. The most common gaps were questionnaire items, revisions or participantfacing interview wording (21 papers), complete agent prompts, grounding inputs, or configuration rules (15), and stimulus or exercise materials (11), with overlap across categories. When unavailable information was necessary to implement the study configuration, it limited which portions could be replicated.

6.5

Analysis Approach

We analyzed the survey, transcripts and screen recordings using a combination of descriptive summaries and qualitative analysis. Completion, timing, hints, and breakdowns were summarized descriptively across participants. Think-aloud transcripts, facilitator notes, and interview responses were analyzed thematically to identify recurring patterns related to learnability, expressivity, reproducibility, and unmet support needs.

Gricea External Dependencies and Scope Boundaries. Fourteen papers involved external requirements in at least one phase: robots or wearables, voice-cloning pipelines, specialized applications, or in-person group coordination. These should generally remain external systems connected to Gricea and hence were only partially replicated. 10

Gricea: An Open Science Platform for Conversational AI Research

6.6

Findings

expressive capacity. Participants successfully generated and operationalized a highly diverse set of research questions, confirming they were able to express the exact study designs they had in mind (𝑀 = 5.5, 𝑆𝐸 = 0.40). Implemented studies spanned the breadth of the conversational AI manipulation space. P7 investigated how AI assistance affects subsequent non-AI creativity. P5 explored the “AI penalty” by testing evaluations of text with and without AI-disclosure. P2 operationalized a study on serendipity across standard web search, generative search, and RAG architectures. Other participants configured studies on proactive versus reactive coding agents (P3), debates involving two voice agents and a human moderator (P8), conditions prompting users to utilize agentic tool-calls (P1), human factchecking behaviors with generative AI (P9), and how different tiers of models exacerbate the digital divide (P10). Table 2 lists the research questions. Participants explicitly confirmed that Gricea easily accommodated these designs. P3 stated, “All of it was able to be done, and I have been trying to think of studies that I cannot run but I cannot think of one yet.” P6 reported Gricea supported “even more than I thought was possible,” and P8 summarized the platform’s capacity as “all of it and more.” Even for highly domain-specific protocols—such as P6’s mechanical design collaboration or P8’s clinical protocols—participants found the platform broadly expressive enough to capture their required experimental manipulations and compliant with medical standards for PII data.

6.6.1 Overcoming the Prototyping and Reproducibility Bottleneck. In the pre-task interviews, participants consistently described current conversational AI workflows as bespoke, costly, and highly unstable. P7 noted, “There’s no standardized way of conducting these studies... these models keep changing over time.” P6 described building custom interfaces from scratch, noting, “I manually coded all of that,” which took “a few weeks.” For participants with limited programming experience, this overhead was prohibitive. P9 stated, “I cannot code so I never thought I would be able to do this on my own, I would have to just pay someone to build it out for me.” Furthermore, reproducibility emerged as a primary concern. P7 articulated this explicitly: “If I want to release my system, it is not clear to me how my systems can stand over time and how will my study be replicated and built upon.” Following the authoring tasks, participants overwhelmingly viewed Gricea as a structural solution to these bottlenecks. P8 estimated that utilizing Gricea would save “at least 3 to 6 months of effort” and roughly “$25,000 worth of money,” noting that building their protocol manually “would have been a nightmare.” This qualitative enthusiasm was supported by the post-task survey, where participants indicated a strong likelihood to recommend Gricea to colleagues (𝑀 = 6.8, 𝑆𝐸 = 0.13) and expressed high confidence that studies authored in Gricea could be reliably reproduced by other researchers (𝑀 = 5.9, 𝑆𝐸 = 0.37). 6.6.2 Participants from diverse backgrounds were able to author studies. Participants across disciplines successfully mapped their conceptual study designs onto Gricea’s dual-flow abstractions, reporting high overall ease of use (𝑀 = 6.1, 𝑆𝐸 = 0.31). Several explicitly stated that the node-based architecture mirrored their internal cognitive models of experimental design. P1 noted that “all those component building blocks make sense to me” and could be used “intuitively” to build procedures. P3 found that “the plug, play, click and edit pipeline is really intuitive,” preserving their “chain of thought,” while P4 highlighted that “the most intuitive part is when you connect those together.” Three platform strengths consistently emerged as critical for supporting researchers from such diverse backgrounds. First, the visual canvas provided necessary architectural clarity; P8 noted, “The visual interface is quite easy to develop the protocol... The visual overlay made a lot of difference.” Second, participants valued the unified consolidation of procedure and runtime. P6 appreciated having surveys, AI tasks, and deployment “contained into one system,” contrasting it with fragmented legacy workflows. Third, the localized validation and preview mechanisms were highly praised. P7 appreciated “being able to preview it... and seeing if what I thought is what is actually happening,” while P3 praised “the ease with which a study could be validated,” calling the feature “super cool.” These accounts show why inspection matters beyond usability: participants used the representation and preview to connect their intended design to the procedure and interaction the platform would execute.

6.6.4 High Degrees of Freedom and the Need for additional scaffolding. While Gricea’s flexibility enabled diverse RQs, the high degrees of freedom introduced new methodological frictions. Breakdowns occurred most often for variables and branching: P4 found the randomization logic “confusing for me for like the branching and the randomized [nodes],” while P1 struggled with “how to dynamically insert a prompt” using variables. On a similar note, several participants experienced a conceptual boundary when transitioning from the macro Study Flow to the micro Task Flow. P6 noted that entering the Task Builder “brought us to another workflow that was a little confusing.” P2 found the task canvas overly granular, stating that “it offers too much detail” and “for social scientists who don’t need complex configurations they might to abstract away the complexities.” These frictions suggest that while Gricea’s core abstractions are powerful, researchers, especially those from non-tehcnical background, require stronger scaffolding to manage the flow of variables across nested nodes. P1 and P4 also noted a “high-learning curve”. To resolve these issues, participants requested additional support for methodological mapping. P7 desired an AI feature to provide “a basic starting template” to help structure the flow. P1 articulated an important boundary condition for the platform: Gricea is highly effective when the experimental design is already concrete, but early-stage ideation still requires “collaborative brainstorming with AI agents.” All participants however echoed that once they were used to the platform and ran a few studies they could see themselves getting over these barriers.

6.6.3 Gricea supported RQs across the Conversational AI Design Space. The open-ended phase provided robust evidence of Gricea’s 11

Sharma et al.

Participant Open-ended research question P01 P02 P03 P04 P05 P06 P07 P08 P09 P10

Under what conditions and tasks do users rely on different agentic tool calls, such as web search? How does serendipity differ across web search, generative-AI search, and retrieval-augmented generation? How do proactive versus reactive coding agents shape participants’ coding practices? How does AI-assisted creative writing compare to human-assisted creative writing? If people cannot distinguish AI-written text from human-written text, where does the perceived “AI penalty” in evaluation come from? How do different personas of AI chatbots affect engineers in mechanical design tasks? How does AI assistance in creativity tasks affect subsequent creativity without AI? How does an LLM-augmented debate with a human moderator affect discussion and decision-making? How do people fact-check information when using generative AI? How do differences in model tier exacerbate the digital divide?

Table 2: Open-ended research questions participants implemented as valid, runnable Gricea studies. These studies varied across procedures, interfaces, models, surveys, and outcomes, illustrating the range of research designs expressible through the shared study representation.

7 Discussion 7.1 Conversational AI Studies as Research Artifacts

7.2

Related Research Questions Require Inspectable Study Configurations

Our CUI reconstructions show that studies can address a shared research question while implementing substantially different experimental conditions. Studies examining uncertainty and reliance implemented that relationship through different tasks and procedures: one compared model and prompting conditions during travelplanning conversations, whereas another manipulated linguistic hedging and decision friction in financial advice with explicit allocation decisions [2, 90]. Beyond the underlying model, the studies differ in condition assignment, conversational progression, available actions, and the operationalization of reliance. Preserving the configured conditions in a shared representation allows subsequent researchers to inspect an earlier study and identify which parts to retain and which to vary when investigating a related hypothesis. Studies of conversational contact with an outgroup likewise address related questions through different interaction designs. Prior work has compared individual dialogue with a chatbot expressing a vegan perspective against a static essay, and facilitator-mediated group dialogue with an agent grounded in outgroup members’ discussions against a document containing that material [32, 42]. Although both investigate the consequences of interactive exposure to an outgroup perspective, they differ in who participates, whose perspective grounds the agent, how messages are composed, and how attitude change is measured. Preserving the study configurations in a shared artifact allows researchers to compare how related questions have been operationalized and design subsequent studies that retain or vary specific aspects of the interaction.

Studies on conversational AI should remain useful beyond the team that created them, but a publication alone cannot preserve the participant experience when procedures, interfaces, prompts, grounding, and API orchestration remain fragmented across code and configuration. Gricea makes the executable study available for inspection, reproduction, and extension, giving the research community a shared basis for building on an individual study. Participants valued versioning, inspectability, templates, and shareable links as ways to understand what another researcher built and ran. Preserving the authored condition makes methods easier to audit and interpret, extending the value of these features beyond individual authoring convenience. Our replication study of CUI 2026 papers revealed a gap between reporting a study for publication and preserving the information needed to reproduce its implementation (Section 6.2). Although these papers had passed peer review, questionnaire wording, agent prompts and grounding inputs, and stimulus or exercise materials were frequently incomplete or unavailable. When those details were necessary to implement a study, their absence prevented replication of the corresponding portions. A published account can therefore communicate a study’s rationale, procedure, and findings while leaving another research team unable to reconstruct the participant experience. Reporting and reproducibility frameworks call for making methodological choices explicit [3, 33, 35]. Our findings motivate preserving executable study artifacts as part of the methodological record alongside the paper, so that the community can inspect and reuse the implementation underlying the reported method. Gricea connects study materials and configurations to the executable study version, preserving them during authoring and execution rather than requiring researchers to reconstruct those connections after publication.

7.3

Study Flow and Task Flow as a Methodological Abstraction

In Gricea, the distinction between Study Flow and Task Flow clarifies the methodological structure of conversational AI studies as well as organizing the platform. Researchers design procedures that assign participants to conditions, collect pre- and post-task measures, and manage staged progression, while also configuring an interactive runtime in which participants encounter a conversational 12

Gricea: An Open Science Platform for Conversational AI Research

7.4

condition that unfolds turn by turn. The participant experience therefore depends on both the study procedure and the behavior of the interactive task. Across the papers reviewed in our formative analysis, researchers often communicated study procedures through staged diagrams, branching depictions, and flowchart-like structures. The diagrams externalized how researchers conceptualized and communicated experimental logic, beyond summarizing a paper after the fact. Participants in our evaluation repeatedly described Gricea’s visual canvas as matching how they thought about study design, and several highlighted that connecting stages and previewing the resulting flow helped them reason about the experiment more concretely. Visual authoring therefore matters both for usability and for its alignment with how researchers already represent study logic in practice. The process of developing Gricea suggests that research platforms should organize their abstractions around the choices researchers need to distinguish and control. Study Flow separates assignment, ordering, and measurement from the operations that implement the interactive task, while Task Flow makes those operations inspectable without hiding their dependencies on the surrounding procedure. Separating study procedure from interactive task behavior allows a researcher to change conversational behavior while retaining the study sequence, or to change the experimental design while retaining the task. Because Study Flow and Task Flow also drive execution, the relationships a researcher inspects are the relationships used to implement the study. Block-based authoring makes the methodological choices explicit and reusable, extending work on reusable experimental components and methodological specifications [28, 50]. Separating study procedure from interactive task behavior also provides a basis for extending research platforms without fragmenting their methods. A new task component should expose its configurable properties, required inputs, completion conditions, and recorded outputs so that it can participate in the existing procedure and data collection mechanisms. In Gricea, a modality-specific extension can reuse the surrounding study infrastructure while making its methodological consequences visible to the researcher. The component and its connections become part of the executable study representation, preserving the relationship between a configuration change and the participant experience it produces when another research team reuses or extends the study. The distinction between procedure and interactive task also applies to online studies of decision aids, interactive visualizations, educational interfaces, and collaborative tools, which combine assignment and measurement with a task that responds to participant actions [4, 16, 25]. The representation could extend to these settings by replacing or adding task components while retaining the surrounding Study Flow and its connection to execution records. Different participant-facing activities could therefore share the same mechanisms for specifying, running, and preserving a study.

Lowering Technical Barriers Requires Methodological Scaffolding

The cost of building controlled conversational AI studies can constrain who conducts them and which questions they explore. Lowering technical barriers broadens participation in producing research, not only in using a platform. Researchers from different disciplinary backgrounds implemented valid, runnable studies across the research questions in Table 2. Gricea supported participants’ study designs through a common representation while reducing barriers for researchers without systems backgrounds. Sharing the resulting artifacts also allows other teams to inspect and adapt the studies without rebuilding their infrastructure. Our evaluation also showed where researchers need methodological scaffolding: participants requested help translating research intent into variables, branching, randomization, and validation, and previewing how those choices shape the participant experience. Related systems likewise show the need for guidance alongside greater authoring control [8, 97]. In Gricea, methodological scaffolding should help researchers connect a research question to executable study logic while retaining visibility and control over the resulting design.

7.5

Making Design-to-Implementation Fidelity Inspectable

With a shared study representation, researchers can examine assignment branches and task configurations, preview the participant experience, and relate execution records to the published Study Version to check design-to-implementation fidelity. P7 described the check as asking whether “what I thought is what is actually happening”. Inspecting the configured study and its execution lets researchers examine the experimental choices used during execution rather than infer them from a separate implementation. Even a structurally valid study can diverge from its intended design: a valid within-subject flow may still misrepresent an intended between-subject comparison. Researcher review checks whether the implemented study matches the intended design, while publication preserves the inspected definition and version-linked records expose its execution. Connecting researcher review, the published study definition, and execution records gives the original team and subsequent researchers a common basis for examining fidelity.

7.6

Open-Science Workflows for Cumulative Conversational-AI Research

When we scale across a sequence of studies, the community value of Gricea becomes obvious: one team may publish a conversational task and its experimental conditions; another may retain that procedure while changing an agent behavior, interface feature, or participant population to examine a related hypothesis. Preserving the source and revised artifacts makes the methodological relationship between studies inspectable, giving researchers a basis for explaining which findings concern the same configured interaction and which concern a deliberate extension, rather than treating each implementation as an unrelated starting point. 13

Sharma et al.

Connected artifacts could help a research community trace evolving hypotheses, identify unresolved comparisons, and synthesize findings across related studies (Figure 1). Synthesizing findings across studies requires examining differences in populations, measures, and contexts as well as shared configurations. Gricea provides the mechanisms for publishing, inspecting, and reusing artifacts, while the contribution of shared artifacts to cumulative knowledge will depend on sustained community use.

7.7

The published-study evaluation is also limited to CUI 2026. The corpus provides varied conversational interfaces and study procedures, but does not cover the full range of social-science experiments or the complex combinations of task behavior, modalities, and study logic that Gricea permits. Future work: Our evaluation highlights three directions for future work. First, to help researchers navigate Gricea’s high degrees of freedom, we plan to integrate an AI copilot that proposes initial study configurations from high-level research questions and modifications to existing artifacts. Proposals would remain within the shared study representation and subject to researcher inspection, validation, and preview before publication. Second, addressing feedback that Gricea is currently “best used when there is [a] clear RQ” (P1, P10), we will add methodological scaffolding to help researchers brainstorm and identify potential confounders during early-stage ideation. Third, Gricea allows support for running AI participant simulations that could support pre-deployment checks of study paths, interaction behavior and pilot studies, further lowering the cost of conducting human subject studies. [72].

AI-Assisted Research Through a Shared Study Representation

AI-assisted research can build on shared study representations: Liu et al. [63] argue that agents need executable research artifacts to understand, reproduce, and extend scientific work. In Gricea, authoring agents could propose or modify Study Flow and Task Flow while researchers inspect changes to assignment, conversational behavior, measurement, and the resulting participant experience. Integrating coding assistance with Study Flow and Task Flow would preserve a common method for reviewing and reusing agent-authored studies. Supporting AI-native research requires distinguishing agents that help conduct research from agents whose behavior is part of the experiment. An authoring agent might retrieve a prior study, propose a controlled variation, or explain a change to its configuration. An agent participating within a study instead operates under the roles, information access, and interaction rules specified by that study, as human–AI research platforms have begun to support [76, 97]. A shared representation can make both authoring changes and within-study agent behavior inspectable, while keeping changes to the research design separate from actions taken within an experimental condition. Agent-proposed changes could retain a link to the source study and expose which conditions they alter, allowing researchers to review the proposed revisions before accepted changes enter the representation used for deployment and data collection. Human and AI contributions could then be inspected and reused through a common research workflow.

8

References [1] Adnan Abbas, Caleb Wohn, Donghan Hu, Eugenia H Rho, and Sang Won Lee. 2025. PITCH: Designing Agentic Conversational Support for Planning and Self-reflection. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machinery, New York, NY, USA, Article 62, 22 pages. doi:10.1145/3719160.3736634 [2] Yasmeen Abdrabou, Yomna Abdelrahman, Efe Bozkir, Youssef Mazen, Florian Alt, and Enkelejda Kasneci. 2026. Beyond Benchmarks: A User-Centric Framework for Evaluating Large Language Models. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–17. doi:10.1145/3816046.3816227 [3] Leonel Aguilar, Michal Gath-Morad, Jascha Grübel, Jasper Ermatinger, Hantao Zhao, Stefan Wehrli, Robert W. Sumner, Ce Zhang, Dirk Helbing, and Christoph Hölscher. 2024. Experiments as Code and its application to VR studies in humanbuilding interaction. Scientific Reports 14, 1 (30 Apr 2024), 9883. doi:10.1038/ s41598-024-60791-3 [4] Abdullah Almaatouq, Joshua Becker, James P. Houghton, Nicolas Paton, Duncan J. Watts, and Mark E. Whiting. 2021. Empirica: a virtual lab for highthroughput macro-level experiments. Behavior Research Methods 53, 5 (Mar 2021), 2158–2171. doi:10.3758/s13428-020-01535-9 [5] Saleema Amershi, Daniel S. Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi T. Iqbal, Paul N. Bennett, Kori Inkpen Quinn, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. 2019. Guidelines for Human-AI Interaction. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (2019). https://api.semanticscholar.org/CorpusID: 86866942 [6] Mehrasa Amiri Besheli and Mahmood Jasim. 2026. SPARC: Exploring Interaction, Sensemaking, and Engagement in AI-Augmented News Reading. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–15. doi:10.1145/3816046.3816213 [7] Shutaro Aoyama, Kiyoshi Suganuma, He Jiang, and Shunichi Kasahara. 2026. Designing a Feedback Loop Between a Human and Their AI Clones for Science Communication in Museums. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–18. doi:10.1145/3816046.3816205 [8] Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena L. Glassman. 2024. ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). ACM, 1–18. doi:10.1145/3613904.3642016 [9] Sandeep Avula, Bogeum Choi, and Jaime Arguello. 2022. The Effects of System Initiative during Conversational Collaborative Search. Proc. ACM Hum.-Comput. Interact. 6, CSCW1, Article 66 (apr 2022), 30 pages. doi:10.1145/3512913 [10] Dünya Baradari, Nataliya Kosmyna, Oscar Petrov, Rebecah Kaplun, and Pattie Maes. 2025. NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machinery, New York, NY, USA, Article 57, 21 pages. doi:10.1145/3719160.3736623 [11] Arndt Bieberstein, Benjamin Lukas Schnitzer, Stefano Gampe, and Oliver Korn. 2026. Do Prompt-Level Empathy Instructions Influence User Experience? Evidence From A Controlled Chatbot Study. In Proceedings of the 8th

Limitations and Future Work

Gricea is intended for a broad research community. The authoring study demonstrates how researchers from varied backgrounds implemented conversational AI studies, but its sample of ten participants does not capture the full range of prospective researchers and research practices. The evaluation covers a subspace of the methodologies and configurations that Gricea is intended to support. The reconstructed online components do not capture the facilitated group activities of participatory co-design workshops, situated observations of people using their own devices and assistive technologies, or the physical behavior of embodied agents [21, 80, 81]. These methods involve participant activities, physical settings, and researcher involvement beyond the procedures and interactions represented in the reconstructed components. 14

Gricea: An Open Science Platform for Conversational AI Research

ACM Conference on Conversational User Interfaces (Bremen, Germany) (CUI ’26). Association for Computing Machinery, New York, NY, USA, 3:1–3:15. doi:10.1145/3816046.3816221 [12] Arthur Caetano, Kavya Verma, Atieh Taheri, Radha Kumaran, Zichen Chen, Jiaao Chen, Tobias Höllerer, and Misha Sra. 2025. Agentic Workflows for Conversational Human-AI Interaction Design. arXiv:2501.18002 [cs.HC] https: //arxiv.org/abs/2501.18002 [13] Wanling Cai, Yucheng Jin, Xianglin Zhao, and Li Chen. 2023. “Listen to Music, Listen to Yourself”: Design of a Conversational Agent to Support Self-Awareness While Listening to Music. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 119, 19 pages. doi:10.1145/ 3544548.3581427 [14] Yuzhe Cai, Shaoguang Mao, Wenshan Wu, Zehua Wang, Yaobo Liang, Tao Ge, Chenfei Wu, Wang You, Ting Song, Yan Xia, Jonathan Tien, Nan Duan, and Furu Wei. 2024. Low-code LLM: Graphical User Interface over Large Language Models. arXiv:2304.08103 [cs.CL] https://arxiv.org/abs/2304.08103 [15] Aaron Chatterji, Thomas Cunningham, David J Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025. How people use chatgpt. Technical Report. National Bureau of Economic Research. [16] Daniel L. Chen, Martin Schonger, and Chris Wickens. 2016. oTree—An opensource platform for laboratory, online, and field experiments. Journal of Behavioral and Experimental Finance 9 (Mar 2016), 88–97. doi:10.1016/j.jbef.2015.12. 001 [17] Jiaqi Chen, Yanzhe Zhang, Yutong Zhang, Yijia Shao, and Diyi Yang. 2025. Generative Interfaces for Language Models. arXiv:2508.19227 [cs.CL] https: //arxiv.org/abs/2508.19227 [18] Lingjiao Chen, Matei Zaharia, and James Zou. 2024. How Is ChatGPT’s Behavior Changing Over Time? Harvard Data Science Review 6, 2 (Mar 2024). doi:10. 1162/99608f92.5317da47 [19] Ruijia Cheng, Titus Barik, Alan Leung, Fred Hohman, and Jeffrey Nichols. 2024. BISCUIT: Scaffolding LLM-Generated Code with Ephemeral UIs in Computational Notebooks. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC). https://arxiv.org/abs/2404.07387 [20] Peggy Chi, Senpo Hu, Lei Shi, Tanya Kraljic, Justin Secor, Tao Dong, Irfan Essa, and Mike Cleron. 2025. WatchWithMe: LLM-Based Interactive Guided Watching of Review Videos. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–15. doi:10.1145/3719160.3736624 [21] Sena Choi and Joel E Fischer. 2026. “I’M BLIND, ChatGPT”: Interactional Breakdown and Repair in LLM-based Conversational AI for Visually Impaired Users. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–14. doi:10.1145/3816046.3816209 [22] Orla Cooney, Sophie J. Leonard, Paola R. Peña, Jaimy Hannah, and Benjamin R. Cowan. 2026. Exploring Perceptions of Robo-advisors for Personal Investing. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–12. doi:10.1145/3816046.3816211 [23] Samuel Rhys Cox, Rune Møberg Jacobsen, and Niels van Berkel. 2025. The Impact of a Chatbot’s Ephemerality-Framing on Self-Disclosure Perceptions. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machinery, New York, NY, USA, Article 60, 17 pages. doi:10.1145/3719160.3736617 [24] Everlyne Kimani Cross, Luiza Santos, Laurent Denoue, Scott Carter, Nayeli Suseth Bravo, and Kate Sieck. 2026. ConvoDojo: Structured LLM-based Sparring Partners for Difficult Workplace Conversations.. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–15. doi:10.1145/3816046.3816230 [25] Zach Cutler, Jack Wilburn, Hilson Shrestha, Yiren Ding, Brian Bollen, Khandaker Abrar Nadib, Tingying He, Andrew McNutt, Lane Harrison, and Alexander Lex. 2026. ReVISit 2: A Full Experiment Life Cycle User Study Framework. IEEE Transactions on Visualization and Computer Graphics 32, 1 (Jan 2026), 13–23. doi:10.1109/tvcg.2025.3633896 [26] Elaine Czech, Ewan Soubutts, Ian Craddock, and Aisling Ann O’Kane. 2025. Understanding the Multimodal Voice Assistant as an Informal and Social Care Support Tool in the UK. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–14. doi:10.1145/3719160.3736605 [27] Nhi Dam and Leonhard Glomann. 2026. AI Interviews the Interviewers: Practitioner Experience and Evaluation of Conversational AI Interviewing. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–14. doi:10.1145/3816046.3816218 [28] Joshua R. de Leeuw, Rebecca A. Gilbert, and Björn Luchterhandt. 2023. jsPsych: Enabling an Open-Source Collaborative Ecosystem of Behavioral Experiments. Journal of Open Source Software 8, 85 (May 2023), 5351. doi:10.21105/joss.05351 [29] Smit Desai, Jessie Chin, Dakuo Wang, Benjamin R. Cowan, and Michael Twidale. 2026. Toward Metaphor-Fluid Conversation Design for Voice User Interfaces. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–21. doi:10.1145/3816046.3816223 [30] Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel Peter Robert. 2024. Shaping Human-AI Collaboration: Varied

Scaffolding Levels in Co-writing with Language Models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 1044, 18 pages. doi:10.1145/3613904.3642134 [31] Sarah Elwahsh, Nora Stern, Aneesha Singh, and Amid Ayobi. 2025. Linguistic Diversity and Mental Well-Being: Co-Designing Custom AI Chatbots with Multilingual Mothers. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–17. doi:10.1145/3719160.3736615 [32] Jana Falkner and Yvonne Kammerer. 2026. More Interactivity, More OpenMindedness? The Influence of AI-based Chatbot Interactivity on Attitudes Toward Vegans. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–11. doi:10.1145/3816046.3816208 [33] Sebastian S Feger, Sünje Dallmeier-Tiessen, Paweł W Woźniak, and Albrecht Schmidt. 2019. The role of hci in reproducible science: Understanding, supporting and motivating core practices. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems. 1–6. [34] K. J. Kevin Feng, Q. Vera Liao, Ziang Xiao, Jennifer Wortman Vaughan, Amy X. Zhang, and David W. McDonald. 2025. Canvil: Designerly Adaptation for LLM-Powered User Experiences. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM, 1–22. doi:10.1145/3706598. 3713139 [35] Stefan Feuerriegel, Christopher Barrie, M. J. Crockett, Laura K. Globig, Killian L. McLoughlin, Dan-Mircea Mirea, Arthur Spirling, Diyi Yang, Tim Althoff, Maria Antoniak, Lisa P. Argyle, Ashwini Ashokkumar, Mohammad Atari, Hannah Bailey, Kevin Bauer, Umang Bhatt, Yidong Chai, Tanmoy Chakraborty, Yanto Chandra, Huimin Chen, Hal Daumé III, Gianmarco De Francisci Morales, Morteza Dehghani, Danica Dillion, Johannes C. Eichstaedt, Kerstin Forster, Dominique Geissler, Kurt Gray, Thomas L. Griffiths, Jochen Hartmann, Oliver P. Hauser, James K. He, Rahul Hemrajani, Felix Holzmeister, Angel Hsing-Chi Hwang, Tiancheng Hu, Anna A. Ivanova, Nils Köbis, Yara Kyrychenko, Himabindu Lakkaraju, Jia Liu, Abdurahman Maarouf, Sebastian Maier, Lennart Meincke, Rada Mihalcea, Brent Mittelstadt, Saif M. Mohammad, Mor Naaman, Oded Netzer, Alice Oh, Desmond C. Ong, Francesco Pierri, Barbara Plank, Iyad Rahwan, Talal Rahwan, Pooja S. B. Rao, Claire E. Robertson, David M. Rothschild, Matthew J. Salganik, Eric Schulz, Chirag Shah, Yash Raj Shrestha, Ekaterina Shutova, Alexandra A. Siegel, Almog Simchon, Huan Sun, Malte Toetzke, Jay J. Van Bavel, Michelle Vaccaro, Jennifer Wortman Vaughan, Effy Vayena, Pedro O. S. Vaz-de Melo, Briana Vecchione, Angelina Wang, Robert West, Robb Willer, Dirk U. Wulff, Renwen Zhang, Simone Zhang, Steve Rathje, and Manoel Horta Ribeiro. 2026. A reporting checklist for large language models in behavioural science. Nature Human Behaviour 10, 7 (June 2026), 1182–1186. doi:10.1038/s41562-026-02492-7 [36] Claudia Flores-Saviaga, Benjamin V. Hanrahan, Kashif Imteyaz, Steven Clarke, and Saiph Savage*. 2025. The Impact of Generative AI Coding Assistants on Developers Who Are Visually Impaired. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM, 1–17. doi:10.1145/3706598.3714008 [37] Erin D. Foster and Ariel Deardorff. 2017. Open Science Framework (OSF). Journal of the Medical Library Association 105, 2 (Apr 2017). doi:10.5195/jmla. 2017.88 [38] Xiao Ge, Chunchen Xu, Daigo Misaki, Hazel Rose Markus, and Jeanne L Tsai. 2024. How Culture Shapes What People Want From AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). ACM, 1–15. doi:10.1145/3613904.3642660 [39] Google Cloud. 2026. Dialogflow CX: Experiments. https://docs.cloud.google. com/dialogflow/cx/docs/concept/experiments Official documentation; updated September 3, 2026; accessed September 10, 2026. [40] Google Cloud. 2026. Dialogflow CX: Playbooks. https://docs.cloud.google. com/dialogflow/cx/docs/concept/playbook Official documentation; updated September 3, 2026; accessed September 10, 2026. [41] Google Cloud. 2026. Dialogflow CX: Versions and Environments. https://docs. cloud.google.com/dialogflow/cx/docs/concept/version Official documentation; updated September 3, 2026; accessed September 10, 2026. [42] Koken Hata, Rintaro Chujo, Reina Takamatsu, Wenzhen Xu, and Yukino Baba. 2026. GroupEnvoy: A Conversational Agent Speaking for the Outgroup to Foster Intergroup Relations. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (Bremen, Germany) (CUI ’26). Association for Computing Machinery, New York, NY, USA, 31:1–31:18. doi:10.1145/3816046.3816204 [43] Ari Hautasaari, Yuta Kibayashi, Rintaro Chujo, Yuji Hatada, and Takeshi Naemura. 2026. Butlerliezer: Context- and Receiver-Aware Appropriately Deceptive Auto-Reply System based on Egocentric Video. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–15. doi:10.1145/3816046.3816234 [44] Tingying He and Alexander Lex. 2026. Integrating an LLM-Based Chatbot into a reVISit Study. reVISit project tutorial. https://revisit.dev/blog/2026/04/30/llmin-revisit/ Published April 30, 2026; accessed September 7, 2026.

15

Sharma et al.

[45] Wanqing Psyche He and Susan R. Fussell. 2026. XPLAIN: A Proactive Scaffold Across Speech Processing Stages—Supporting Non-Native Speakers in RealTime AI-Mediated Turn-Taking. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–17. doi:10.1145/3816046.3816235 [46] Olga Iarygina, Kasper Hornbæk, and Aske Mottelson. 2026. On the Computational Reproducibility of Human-Computer Interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). ACM, 1–16. doi:10.1145/3772318.3791129 [47] Habiba Imam and Jaisie Sin. 2026. “More Like a Person’s Voice”: Exploring the Design of Empathetic Virtual Agents for Older Adults. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–11. doi:10.1145/3816046.3816212 [48] Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023. Co-Writing with Opinionated Language Models Affects Users’ Views. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 111, 15 pages. doi:10.1145/3544548.3581196 [49] Daeun Jeong, Sungbok Shin, and Jongwook Jeong. 2025. Conversation Progress Guide: UI System for Enhancing Self-Efficacy in Conversational AI. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 180, 11 pages. doi:10.1145/3706598.3714222 [50] Eunice Jun, Maureen Daum, Jared Roesch, Sarah Chasins, Emery Berger, Rene Just, and Katharina Reinecke. 2019. Tea: A High-level Language and Runtime System for Automating Statistical Analysis. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology (UIST ’19). ACM, 591–603. doi:10.1145/3332165.3347940 [51] Hannah Rose Kirk, Iason Gabriel, Chris Summerfield, Bertie Vidgen, and Scott A. Hale. 2025. Why human–AI relationships need socioaffective alignment. Humanities and Social Sciences Communications 12 (2025), 728. doi:10.1057/s41599025-04532-5 [52] Aniket Kittur, Ed H. Chi, and Bongwon Suh. 2008. Crowdsourcing user studies with Mechanical Turk. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’08). ACM, 453–456. doi:10.1145/1357054.1357127 [53] Abhishek M Kulkarni, Sarah Anne Brown, Neha Rani, and Sharon Lynn Chu. 2026. Towards Pedagogy-Grounded Conversational AI Tutors for InterestBased Learning. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–16. doi:10.1145/3816046.3816215 [54] David Ledo, Steven Houben, Jo Vermeulen, Nicolai Marquardt, Lora Oehlberg, and Saul Greenberg. 2018. Evaluation strategies for HCI toolkit research. In Proceedings of the 2018 CHI conference on human factors in computing systems. 1–17. [55] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. nips 33 (2020), 9459–9474. https://arxiv.org/abs/2005.11401 [56] Yubo Li, Xiaobin Shen, Xinyu Yao, Xueying Ding, Yidi Miao, Ramayya Krishnan, and Rema Padman. 2025. Beyond Single-Turn: A Survey on MultiTurn Interactions with Large Language Models. arXiv:2504.04717 [cs.CL] https://arxiv.org/abs/2504.04717 [57] Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin. 2024. The Value, Benefits, and Concerns of Generative AI-Powered Assistance in Writing. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 1048, 25 pages. doi:10.1145/3613904.3642625 [58] Minhui Liang and Yuhan Luo. 2025. Exploring Multi-LLM Collaboration to Power Conversational Recommender System: A Case Study of Dietary Recommendation. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–6. doi:10.1145/3719160.3737635 [59] Minhui Liang, Jinping Wang, and Yuhan Luo. 2025. SmartEats: Investigating the Effects of Customizable Conversational Agent in Dietary Recommendations. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). Association for Computing Machinery, New York, NY, USA, Article 63, 16 pages. doi:10.1145/3719160.3736635 [60] Q. Vera Liao and Ziang Xiao. 2025. Rethinking Model Evaluation as Narrowing the Socio-Technical Gap. arXiv:2306.03100 [cs.HC] https://arxiv.org/abs/2306. 03100 [61] Chang Liu, Qinyi Zhou, Xinjie Shen, Xingyu Bruce Liu, Tongshuang Wu, and Xiang’Anthony’ Chen. 2026. Behavioral Indicators of Overreliance During Interaction with Conversational Language Models. arXiv preprint arXiv:2602.11567 (2026). [62] Jinyu Liu and Effie L-C Law. 2026. Beyond Emotional Mirroring: Understanding Affective Alignment in Artist-in-the-Loop Art Chatbots. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–14. doi:10.1145/3816046.3816206 [63] Jiachen Liu, Jiaxin Pei, Jintao Huang, Chenglei Si, Ao Qu, Xiangru Tang, Runyu Lu, Lichang Chen, Xiaoyan Bai, Haizhong Zheng, Carl Chen, Zhiyang Chen, Haojie Ye, Yujuan Fu, Zexue He, Zijian Jin, Zhenyu Zhang, Shangquan Sun,

Maestro Harmon, John Dianzhuo Wang, Jianqiao Zeng, Jiachen Sun, Mingyuan Wu, Baoyu Zhou, Chenyu You, Shijian Lu, Yiming Qiu, Fan Lai, Yuan Yuan, Yao Li, Junyuan Hong, Ruihao Zhu, Beidi Chen, Alex Pentland, Ang Chen, Mosharaf Chowdhury, and Zechen Zhang. 2026. The Last Human-Written Paper: Agent-Native Research Artifacts. arXiv:2604.24658 [cs.LG] https://arxiv. org/abs/2604.24658 [64] Joshua Mu-En Liu, Clarence Chi-Hong Weng, Yu-Hsuan Lin, Pang-Yu Hsiao, and Yoyo Tsung-Yu Hou. 2026. AI Echoing in the Backstage: Private AI Consulting May Strengthens Confidence and Limits Depolarization. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–17. doi:10.1145/3816046.3816199 [65] Xingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu, Takeo Igarashi, and Xiang ’Anthony’ Chen. 2025. Proactive Conversational Agents with Inner Thoughts. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 184, 19 pages. doi:10.1145/3706598.3713760 [66] Mykola Maslych, Mohammadreza Katebi, Christopher Lee, Yahya Hmaiti, Amirpouya Ghasemaghaei, Christian Pumarada, Janneese Palmer, Esteban Segarra Martinez, Marco Emporio, Warren Snipes, Ryan P. McMahan, and Joseph J. LaViola Jr. 2025. Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–15. doi:10.1145/3719160.3736636 [67] Anchit Mishra and Oliver Schneider. 2025. TacTalk: Personalizing Haptics Through Conversation. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–19. doi:10.1145/3719160.3736638 [68] Hannah Vy Nguyen, Grace Yu-Chun Yen, Omar Shakir, Hang Huynh, Sebastian Gutierrez, June A Smith, Sheila Jimenez, Salma E Abdelgelil, and Stephen MacNeil. 2025. Feedstack: Layering Structured Representations Over Unstructured Feedback to Scaffold Human–AI Conversation. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–6. doi:10.1145/3719160.3737636 [69] Carolina Nobre, Dylan Wootton, Zach Cutler, Lane Harrison, Hanspeter Pfister, and Alexander Lex. 2021. reVISit: Looking Under the Hood of Interactive Visualization Studies. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). ACM, 1–13. doi:10.1145/3411764.3445382 [70] Dan R Olsen Jr. 2007. Evaluating user interface systems research. In Proceedings of the 20th annual ACM symposium on User interface software and technology. 251–258. [71] Rock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas, Ziang Xiao, Emily Tseng, and Danielle Bragg. 2025. Understanding the LLMification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review. arXiv:2501.12557 (jan 2025). doi:10.48550/arXiv.2501.12557 arXiv:2501.12557 [cs]. [72] Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442 [cs.HC] https://arxiv.org/abs/2304.03442 [73] Eyal Peer, Laura Brandimarte, Sonam Samat, and Alessandro Acquisti. 2017. Beyond the Turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology 70 (May 2017), 153–163. doi:10.1016/ j.jesp.2017.01.006 [74] Eyal Peer, David Rothschild, Andrew Gordon, Zak Evernden, and Ekaterina Damer. 2022. Data quality of platforms and panels for online behavioral research. Behavior Research Methods 54, 4 (01 Aug 2022), 1643–1662. doi:10.3758/s13428021-01694-3 [75] Jacob Penney, Pawan Acharya, Peter Hilbert, Priyanka Parekh, Anita Sarma, Igor Steinmacher, and Marco Aurelio Gerosa. 2025. Outcomes, Perceptions, and Interaction Strategies of Novice Programmers Studying with ChatGPT. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–15. doi:10.1145/3719160.3736625 [76] Crystal Qian, Vivian Tsai, Michael Behr, Nada Hussein, Léo Laugier, Nithum Thain, and Lucas Dixon. 2025. Deliberate Lab: A Platform for Real-Time HumanAI Social Experiments. arXiv:2510.13011 [cs.HC] https://arxiv.org/abs/2510. 13011 [77] Girish Rathod M.S. 2024. Information Diet: Understanding its Role in a Digital Age with Statistical Insights. International Journal of Research in Library Science 10 (11 2024), 67–77. doi:10.26761/ijrls.10.4.2024.1797 [78] Nishant Rathore, Tushar Billakanti, Jessy Ceha, and Jay Henderson. 2026. A Comparison of Speech and Typing Input for Creative Generative AI Tasks. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–10. doi:10.1145/3816046.3816224 [79] Katharina Reinecke and Krzysztof Z. Gajos. 2015. LabintheWild: Conducting Large-Scale Online Experiments With Uncompensated Samples. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (CSCW ’15). ACM, 1364–1378. doi:10.1145/2675133.2675246 [80] Penelope Rekkas, Katrien Verbert, and Vero Vanden Abeele. 2026. Design and Evaluation of ChatBlend: A Framework and Card Deck for the HumanCentered Design of Mental Health Chatbots in Blended Care. In Proceedings of 16

Gricea: An Open Science Platform for Conversational AI Research

[96] Zhicheng Yan, Joel E Fischer, and Jeremie Clos. 2025. Remembering Things Makes Chatbots Sound Smarter, but Less Trustworthy - A Pilot Study. In Proceedings of the 7th ACM Conference on Conversational User Interfaces (CUI ’25). ACM, 1–8. doi:10.1145/3719160.3737617 [97] Bingsheng Yao, Jiaju Chen, Chaoran Chen, April Wang, Toby Jia jun Li, and Dakuo Wang. 2026. Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). ACM, 1–30. doi:10.1145/3772318.3790879 [98] J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023. Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 437, 21 pages. doi:10.1145/3544548. 3581388 [99] Lingyu Zhang, Zhengran Ji, and Boyuan Chen. 2024. CREW: Facilitating HumanAI Teaming Research. Transactions on Machine Learning Research (2024). https: //openreview.net/forum?id=ZRXwHRXm8i [100] Zhehao Zhang, Ryan A. Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, Ruiyi Zhang, Jiuxiang Gu, Tyler Derr, Hongjie Chen, Junda Wu, Xiang Chen, Zichao Wang, Subrata Mitra, Nedim Lipka, Nesreen Ahmed, and Yu Wang. 2025. Personalization of Large Language Models: A Survey. arXiv:2411.00027 [cs.CL] https://arxiv.org/abs/2411.00027

the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–21. doi:10.1145/3816046.3816210 [81] Gisela Reyes-Cruz, Marta Romeo, Dominic James Price, Matthew Peter Aylett, Chris Greenhalgh, and Joel E. Fischer. 2026. Embodied Active Listening: How Non-Verbal Backchannel Behaviour Influences Trust and Perception of a Social Robot. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (Bremen, Germany) (CUI ’26). Association for Computing Machinery, New York, NY, USA, 29:1–29:11. doi:10.1145/3816046.3816228 [82] Mazen Salous, Mikołaj P. Woźniak, Timo von Reeken, Matthias Kramer, Wilko Heuten, Susanne Boll, and Larbi Abdenebaoui. 2026. Beyond Captions: Shaping Imagery Models for Blind or Visually Impaired People During Interacting with AI-powered Conversational Visual Assistant. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–15. doi:10.1145/ 3816046.3816207 [83] Ariadna Sanchez, Jinzuomu Zhong, Artemis Deligianni, Alice Ross, and Simon King. 2026. When Text-to-Speech Speaks in Your Voice: A Study on Public Perception. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–17. doi:10.1145/3816046.3816200 [84] William Seymour, Adam D G Jenkins, Mark Cote, and Jose Such. 2026. Beliefs and Misconceptions around Integrated Conversational AI. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–10. doi:10.1145/3816046.3816203 [85] Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024. Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information Seeking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 1033, 17 pages. doi:10.1145/3613904.3642459 [86] Nikhil Sharma, Kenton Murray, and Ziang Xiao. 2025. Faux Polyglot: A Study on Information Disparity in Multilingual Large Language Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, 8090–8107. doi:10.18653/ v1/2025.naacl-long.411 [87] Nikhil Sharma, Zheng Zhang, Daniel Lee, Namita Krishnan, Guang-Jie Ren, Ziang Xiao, and Yunyao Li. 2026. Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents. arXiv preprint arXiv:2602.01405 (2026). [88] Srividya Sheshadri, Manju Balakrishnan Pillai, Renuka Kumar, Pavithra P M Nair, Apoorv Gupta, Megha Tm, Sooraj S Nair, Balu M Menon, Unnikrishnan Radhakrishnan, and Bhavani Rao R. 2026. Sruthi: What Becomes Sayable Through Peer-Mode Reflective Scaffolding in Fieldwork Communication. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–15. doi:10.1145/3816046.3816219 [89] Yike Shi, Qing Xiao, Qing Hu, Hong Shen, and Hua Shen. 2026. The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models. arXiv:2509.10830 [cs.HC] https://arxiv.org/abs/2509.10830 [90] Laura Spillner, Johanna Rockstroh, Nina Wenig, Nima Zargham, Leonard Kiefner, Robert Porzel, and Rainer Malaka. 2026. Linguistic Uncertainty Markers for Trust Calibration in AI-Assisted Decision-Making. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–15. doi:10.1145/3816046.3816231 [91] Madeleine Steeds, Katie Seaborn, Takao Fujii, Orla Cooney, and Iona Gessinger. 2026. Is ChatGPT Gender-Neutral? Implicit Stereotyping Persists Over Time and With Experience. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–16. doi:10.1145/3816046.3816201 [92] Marisa Victoria Tschopp, Cheng Chen, Magdalena Wischnewski, Miriam Gieselmann, Johannes Schöning, S. Shyam Sundar, and Kai Sassenberg. 2026. Does Humanizing Chatbots Promote or Hinder Voice Shopping?. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (Bremen, Germany) (CUI ’26). Association for Computing Machinery, New York, NY, USA, 28:1–28:14. doi:10.1145/3816046.3816202 [93] Giulia Valcamonica, Giovanni Caleffi, Francesco Piferi, Marta Tagliani, Maria Vender, Denis D. Delfitto, and Franca Garzotto. 2026. LLM-based conversational agents for dyslexia treatment: From chatbots to structured tutors. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (CUI ’26). ACM, 1–12. doi:10.1145/3816046.3816226 [94] Albatool Wazzan, Iman Qaiser, Stephen MacNeil, and Richard Souvenir. 2026. Comparing Intent Communication Modes for Instruction-based Image Editing. In Proceedings of the 8th ACM Conference on Conversational User Interfaces (Bremen, Germany) (CUI ’26). Association for Computing Machinery, New York, NY, USA, 17:1–17:15. doi:10.1145/3816046.3816216 [95] Tongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J Cai. 2022. PromptChainer: Chaining Large Language Model Prompts through Visual Programming. In CHI Conference on Human Factors in Computing Systems Extended Abstracts (CHI ’22). ACM, 1–10. doi:10. 1145/3491101.3519729

17

Sharma et al.

A Appendix A.1 Participant Details The participant summary is provided in Table 3. ID

Role

Conducted studies?

Study count

Primary research interests / intended study

P01

Researcher

Yes

5–10

P02

Master Student

No

N/A

P03

Master Student

No

N.A.

P04

PhD

Yes

1–5

P05

PhD

Yes

1–5

P06

PhD

Yes

1–5

P07

PhD

Yes

1–5

P08

Professor

Yes

10+

P09 P10

Professional Undergraduate Student

Yes No

10+ N/A

Capability boundary of conversational AI; what tasks can and cannot be automated; human–human / human–agent boundaries. Serendipity in search; comparing web search, GenAI search, and RAG search; detailed search-session logging. Legal AI and hallucinated precedent; coding agents; mental-health advice and human influence. Social consequences of conversational agents; tutor-like interaction; transfer from human–AI to human–human interaction. AI penalty: how judgments change when people learn a text was AI-generated versus human-generated. Utility of AI systems in design / mechanical collaboration; AI personas in design tasks. Long-term creativity after AI assistance; comparing no-AI, AI-guidance, and answer-like AI support. Surgeon knowledge assessment; using LLMs to identify knowledge deficits and support dialogue around experiential expertise. How users use physical devices related to healthcare; N/A

Table 3: Completed participant records used in the analysis (𝑛 = 10 usable completed records in the attached JSON export). Participants spanned multiple roles and brought diverse conversational-AI study ideas to the authoring task.

A.2

CUI 2026 Replication Corpus

Table 4 lists all 29 eligible papers included in the published-study replication evaluation (Section 6.2).

18

Gricea: An Open Science Platform for Conversational AI Research

Table 4: The complete corpus of 29 eligible CUI 2026 papers that we attempted to recreate. Reference

Paper

[11] [64] [2] [90] [84] [80]

Do Prompt-Level Empathy Instructions Influence User Experience? Evidence From A Controlled Chatbot Study AI Echoing in the Backstage: Private AI Consulting May Strengthens Confidence and Limits Depolarization Beyond Benchmarks: A User-Centric Framework for Evaluating Large Language Models Linguistic Uncertainty Markers for Trust Calibration in AI-Assisted Decision-Making Beliefs and Misconceptions around Integrated Conversational AI Design and Evaluation of ChatBlend: A Framework and Card Deck for the Human-Centered Design of Mental Health Chatbots in Blended Care “More Like a Person’s Voice”: Exploring the Design of Empathetic Virtual Agents for Older Adults Beyond Captions: Shaping Imagery Models for Blind or Visually Impaired People During Interacting with AI-powered Conversational Visual Assistant “I’M BLIND, ChatGPT”: Interactional Breakdown and Repair in LLM-based Conversational AI for Visually Impaired Users LLM-based conversational agents for dyslexia treatment: From chatbots to structured tutors Comparing Intent Communication Modes for Instruction-based Image Editing Butlerliezer: Context- and Receiver-Aware Appropriately Deceptive Auto-Reply System based on Egocentric Video Toward Metaphor-Fluid Conversation Design for Voice User Interfaces ConvoDojo: Structured LLM-based Sparring Partners for Difficult Workplace Conversations AI Interviews the Interviewers: Practitioner Experience and Evaluation of Conversational AI Interviewing Sruthi: What Becomes Sayable Through Peer-Mode Reflective Scaffolding in Fieldwork Communication SPARC: Exploring Interaction, Sensemaking, and Engagement in AI-Augmented News Reading A Comparison of Speech and Typing Input for Creative Generative AI Tasks XPLAIN: A Proactive Scaffold Across Speech Processing Stages—Supporting Non-Native Speakers in Real-Time AI-Mediated Turn-Taking Does Humanizing Chatbots Promote or Hinder Voice Shopping? Embodied Active Listening: How Non-Verbal Backchannel Behaviour Influences Trust and Perception of a Social Robot Exploring Perceptions of Robo-advisors for Personal Investing GroupEnvoy: A Conversational Agent Speaking for the Outgroup to Foster Intergroup Relations Beyond Emotional Mirroring: Understanding Affective Alignment in Artist-in-the-Loop Art Chatbots Towards Pedagogy-Grounded Conversational AI Tutors for Interest-Based Learning More Interactivity, More Open-Mindedness? The Influence of AI-based Chatbot Interactivity on Attitudes Toward Vegans Is ChatGPT Gender-Neutral? Implicit Stereotyping Persists Over Time and With Experience When Text-to-Speech Speaks in Your Voice: A Study on Public Perception Designing a Feedback Loop Between a Human and Their AI Clones for Science Communication in Museums

[47] [82] [21] [93] [94] [43] [29] [24] [27] [88] [6] [78] [45] [92] [81] [22] [42] [62] [53] [32] [91] [83] [7]

19

Record · ID 1006882 · SHA-256 dbfad21124aed92c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.