A Taxonomy of Human-Robot Teamwork Requirements
arXiv:2607.27302v1 [cs.SE] 29 Jul 2026
Anastasia Mavridou1∗ , Hazel M. Taylor2∗ , Sandy Lozito3 , Louise A. Dennis2 , Michael Fisher2 , Marie Farrell2 1 KBR Inc. at NASA Ames Research Center, Moffett Field, USA 2 University of Manchester, Manchester, UK 3 NASA Ames Research Center, Moffett Field, USA Abstract—Autonomous systems are increasingly deployed in safety- and mission-critical domains where humans and robots must operate as a team to complete complex tasks. Existing requirements for Human-Robot teamwork remain fragmented across disparate sources, with no unified framework that addresses complexities of collaborative Human-Robot tasks. We address this gap by presenting a taxonomy of Human-Robot Teamwork (HRT) requirements derived from analysis of (academic and industrial) literature, standards and regulatory guidance. We extracted a construction corpus of 361 requirements from 14 cross-domain sources. Through iterative classification and refinement, we develop a two-level hierarchical taxonomy comprising 6 high-level categories and 21 low-level subcategories that distinguish information provision, relational control, decision support, safety mechanisms, performance monitoring, and foundational system capabilities. We validate the taxonomy through expert evaluation with 5 domain specialists and a utility demonstration on an independently assembled corpus of 448 requirements drawn from 19 sources spanning six HRT domains. The taxonomy classifies 412 of these requirements, with distributional patterns that converge with the construction corpus while revealing domain-specific patterns and gaps. We examine how requirements distribute across Human-Led, RobotLed and Shared operational perspectives, revealing responsibility boundaries that shape safe collaboration. This work provides a structured foundation for Requirements Engineering (RE) practitioners to systematically elicit and specify HRT capabilities.
I. I NTRODUCTION Autonomous systems are increasing in complexity and capability, enabling new possibilities for collaboration between humans and robots in critical settings. Human-Robot teams are increasingly being tasked with completing complex operations in a plethora of domains, including autonomous robotic systems that (1) reduce human exposure to hazardous environments (e.g., nuclear applications [15]), and (2) enable operations that are simply not possible for humans in isolation due to hazardous and hostile environments (e.g., space missions [13]). In these critical settings, humans must work alongside AIdriven robots to complete tasks that require complementary capabilities, shared decision-making, and coordinated action. Requirements are central to software development in safetycritical domains. Yet, for HRT, they remain fragmented across disparate standards, industry technical reports, and academic literature. This hinders the systematic deployment of HumanRobot teams in critical areas such as aerospace, nuclear op* These authors contributed equally to this work.
erations, and advanced manufacturing, where comprehensive requirements are mandatory. This work fills this gap with a taxonomy and corpora for HRT requirements that organize distributed and disparate sources into a coherent hierarchy. It supports traceability and comparison across systems and team structures, providing requirements engineers with a reference against which specifications can be assessed for coverage and clarity. We seek to answer two research questions: RQ1: What are suitable categories of HRT requirements? RQ2: To what extent can these categories capture diverse patterns of HRT? For RQ1, we consider categories as suitable if they are coherent and well-defined, cover Human-Robot teamwork requirement types without excessive overlap, and are judged useful by domain experts for elicitation and specification tasks. We address RQ1 by deriving a structured taxonomy from the sources listed in Table I. For RQ2, we employ two complementary validation methods: 1) expert evaluation with 5 domain experts; and 2) a utility demonstration [59] on an independently assembled validation corpus. Contributions. In this paper, we contribute: • A construction corpus of 361 requirements for HumanRobot teams extracted from standards, regulatory documents and peer-reviewed literature (§III). • A two-level hierarchical taxonomy of requirements for Human-Robot teamwork comprising 6 high-level and 21 low-level categories, each characterized by distinct teamwork patterns (§IV and §V). • An expert evaluation validates the taxonomy’s conceptual clarity and explores Human-Led, Robot-Led and Shared operational perspectives (§VI). • A validation corpus of 448 requirements across 6 HRT domains, used to validate the taxonomy with a utility demonstration that classifies 412 out-of-sample requirements and highlights domain differences (§VII). II. BACKGROUND AND R ELATED W ORK Human-Robot teams collaborate through supervisory control (humans oversee and intervene), shared decision-making (combining strengths to achieve goals), and independent collaboration (performing distinct tasks while sharing information). Their growing use in critical domains demands verifiable
requirements (e.g., DO-178C [49]), yet current requirements focus on technical correctness while not adequately capturing teamwork aspects critical to reliability and safety. A recent NASA technical standard [46] defines humansystem interaction requirement classes spanning crew health, performance, training, and hardware/software design. We draw inspiration from [46] and reuse its nomenclature where applicable. We also refine categories from [46] to more accurately reflect HRT and introduce new categories not present in [46]. The Brahms language provides semi-formal descriptions of HRT for space applications [4], [62]. Mission requirement patterns for autonomous mobile robots are proposed in [60], though they do not capture the human. Explainability requirements and patterns are studied in [8], [56], [57]. Our taxonomy includes these but focuses more broadly on HRT, complementing existing work. [24] formalizes HRI properties using tockCSP and RoboScene, providing a UML-style notation for HRI scenarios. The conditions that they include in their models are likely instantiations of some of the categories of our taxonomy. A systematic literature review identifies 17 service robot design features under 3 headings: technical design, HRI design, and service system design [69]. Several of their feature categories are mirrored in our taxonomy, including FAILURE R ECOVERY (Failure Impact Feature and Recovery Mechanism Feature in [69]) and S YSTEM C APABILITIES (Functional Capability Feature in [69]). Others are higher-level: e.g., Trust and Emotional Feature in [69] has no direct analog, though our I NFORMATION S HARING AND AWARENESS supports trust calibration. Related work on recommender systems in social robots could help to create trustworthy systems but we did not find concrete examples in our search [25]. We also did not find case studies where ethical requirements were present but this is an important aspect to be considered in future work [45]. Another systematic review defines a taxonomy of robot autonomy for HRI [35]. They identify 6 distinct forms of autonomy, grouped into 3 categories: robot and human involvement at runtime, human involvement before runtime, and expressions of autonomy at runtime. Our taxonomy does not distinguish between autonomy types (out of scope for this work) but the requirements in our corpora likely span them. III. C ONSTRUCTION C ORPUS M ETHODOLOGY We follow four stages: 1) literature and standards review to identify requirement sources, 2) requirement extraction and corpus assembly, 3) iterative taxonomy development, and 4) validation through expert evaluation and utility demonstration. Search Strategy: We employed a hybrid, multi-stage approach. Initial sources were identified through expert recommendations and targeted searches of organization websites (e.g., NASA, FAA, and ISO), then expanded via snowballing [65] through references in relevant technical standards, reports and publications. We also searched IEEE Xplore, ACM Digital Library, and Springer using combinations of core terms, e.g., “human-robot team”, “human-robot collaboration”, “human-robot interaction”, combined with “require-
Venue Advances in Industrial and Manufacturing Engineering Conference on Human Factors in Computing Systems Towards Autonomous Robotic Systems IEEE Intelligent Systems IEEE/AIAA Digital Avionics Systems Conference International Journal of Robotics Research IEEE Access MITRE International Requirements Engineering Conference ISO
Type Journal Conference Conference Journal Conference Journal Journal Tech Report Conference Technical Standard
Federal Aviation Administration (FAA) NASA Code of Federal Regulations (CFR) National Transportation Systems Center
Technical Standard Technical Standard Technical Standard Tech Report
Ref. [50] [1] [19] [37] [58] [63] [6] [44] [34] [29], [28], [30] [14] [46] [5] [71]
TABLE I S OURCES INCLUDED IN THE CONSTRUCTION CORPUS
ment”, “specification”, “standard”, “safety”. We did not apply date restrictions, as foundational standards remain relevant. Inclusion Criteria: Sources were included if they: 1) contained explicit requirements, design guidance or specification patterns; 2) addressed HRI in collaborative or team-based settings; 3) focused on operational, safety, and/or teamwork aspects; 4) provided sufficient detail to extract requirements. Exclusion Criteria: Sources were excluded if they: 1) focused solely on algorithmic performance without human interaction context; 2) addressed single-agent autonomous systems without human-robot collaboration; 3) provided highlevel conceptual frameworks without specific requirements; 4) focused solely on shared goals or joint intentions in classical multi-agent systems literature (e.g., [32]). Source Selection Process: We began with 72 initial candidate sources. Following title and abstract screening, 42 documents were retained for full-text review. Application of the inclusion and exclusion criteria resulted in a construction corpus of 14 sources used for requirement extraction in Table I. Extraction Process: Two researchers independently reviewed each source to identify requirement statements. The researchers compared their extractions and discussed discrepancies. For ambiguous requirement statements, a third researcher reviewed to determine if the requirement statement followed the inclusion/exclusion criteria. Initial agreement between the primary extractors was 87%. Disagreements primarily concerned: 1) Boundary between teamwork and general HRI; 2) whether statements were requirements versus contextual information. For extraction, we distinguish teamwork requirements from general HRI requirements based on interdependent collaboration. We retained only requirements that explicitly support coordinated roles, shared responsibility, or distributed decision-making toward a common objective. Disagreements were resolved through discussion and third-party review, resulting in the final construction corpus of 361 requirements. Taxonomy Development: Once we assembled our construction corpus, we examined and classified each requirement. Two existing frameworks were directly relevant to our work: the McDermott et al. Human-Machine Teaming (HMT) framework [44] and NASA’s Human-System Standard [46]. McDermott et al. organize HMT into 10 themes across 4
higher-level categories (Transparency, Augmenting Cognition, Coordination, Design Specifics), providing generic requirements spanning autonomous vehicles to cognitive assistants that can be tailored to specific systems. NASA’s standard defines human-system interaction requirement types spanning crew health, performance, training, and hardware/software design. These were developed through extensive expert consultation and operational use in safety-critical systems. We also considered broader academic frameworks of HRI and CPS teaming (e.g., levels-of-autonomy [51] and interaction-mode classifications [37]), but these describe conceptual collaboration modes rather than concrete requirement types. We adopted NASA’s Human-System Standard [46] as our initial resource because its categories are organized around concrete interaction types (e.g., automation system status provision, mode change notification) and written as specific requirements consistent with regulated specification practice [31]. By contrast, McDermott et al. [44] organize categories around high-level HMT properties (e.g., Observability, Calibrated Trust, Common Ground), and the requirements within them are intentionally generic. NASA’s Standard thus more closely matched the granularity of our corpus requirements and the RE practices our taxonomy is meant to support. Preliminary analysis revealed that NASA’s framework was not enough for complete coverage of our construction corpus, necessitating extension and refinement. We followed iterative refinement toward two ending conditions: an objective condition (all objects are classifiable without a residual category) and a subjective condition (categories are concise, robust, comprehensive, and meaningful to domain experts) [47]. Categories from [46] were 1) broadened when few requirements matched a narrow category, and 2) narrowed when requirements fitting a category had clear differences between them. Taxonomy Structure: The taxonomy derived from our construction corpus defines a two-level hierarchy, with highlevel categories capturing abstract teamwork concerns, while low-level categories represent more concrete and actionable requirement groupings. The two-level hierarchical structure is grounded in classification theory [39], where hierarchical organization is appropriate when concepts can be meaningfully distinguished at different levels of abstraction. Each high-level category may comprise many low-level categories, while every low-level category maps exactly to one high-level category. IV. TAXONOMY: H IGH -L EVEL C ATEGORIES Our categorization provides a high-level classification of HRT requirements (Fig. 1, Table II). Each high-level category reflects a functional concern within teamwork and differs in terms of what is being exchanged (e.g., data, commands, or safety measures), the purpose of the exchange (e.g., situational awareness, control, or safety assurance), and the type of interaction (one-way, two-way, proactive, or reactive). I NFORMATION S HARING AND AWARENESS addresses state visibility without prescribing action; C ONTROL AND C OORDINA TION concerns allocation and authority; C OMMUNICATION AND D ECISION S UPPORT addresses influence on choice; S AFETY
AND R ISK M ITIGATION focuses on protection and containment;
T EAM P ERFORMANCE M ONITORING concerns assessment of agent state; and T EAM C APABILITIES AND AUTHORITY defines foundational properties. The first 3 categories primarily focus on real-time teamwork and decision-making, ensuring shared understanding and synchronized action, whereas S AFETY AND R ISK M ITIGATION and T EAM P ERFORMANCE M ONITORING focus on ensuring safe, reliable, and effective collaboration. In addition, some categories (e.g., I NFORMATION S HARING AND AWARENESS) are predominantly proactive, while others (e.g., S AFETY AND R ISK M ITIGATION) define primarily reactive interactions. These distinctions ensure analytical separability while acknowledging operational interdependence - especially between I NFORMATION S HARING AND AWARENESS, which provides state visibility without directing behavior, and C OMMUNI CATION AND D ECISION S UPPORT , which actively shapes action. For each category we provide a description of the (1) rationale; (2) type of exchanged information; (3) type of interaction and (4) key differentiator(s) between it and the other categories. Each category is decomposed into several subcategories, we describe these with examples in §V.
A. I NFORMATION S HARING AND AWARENESS Rationale: This focuses on the availability and accessibility of critical data for human and robotic teammates. It emphasizes maintaining situational awareness and ensuring that necessary data is provided to both parties for effective collaboration. Type of information: Raw data and timely status updates. Purpose: Ensuring both human(s) and robot(s) have adequate data to understand the system’s state and operate effectively. Type of interaction: Primarily one-way. For example, robot providing data to human, or vice versa. Key Differentiator: This category is about passive awareness. It ensures that operators are informed but it does not consider direct decision-making or control actions. Example Scenario: The robot provides telemetry and automation status to ensure that the human has the required data. B. C ONTROL AND C OORDINATION Rationale: This category encapsulates the ability of human operators to configure, control, and delineate responsibilities within the Human-Robot team. It ensures that the robot is used and performs in a way that aligns with human intent while maintaining safety and operational efficiency. Type of information: Commands, permissions, and role delineation between human and robot. Purpose: Enabling the human to define, adjust, and/or authorize the robot’s automated/autonomous actions while ensuring that the system operates within the defined framework. Type of interaction: Two-way negotiation where humans set automation parameters and robots operate within them. Key Differentiator: Unlike I NFORMATION S HARING AND AWARENESS, this category is about actively shaping how the robot functions within the team. Example Scenario: The human sets robot automation levels, defines responsibilities and authorizes restarts after failures.
Category
I NFORMATION S HARING AND AWARENESS
Focus Providing system status and data
Interaction Type Primarily one-way (status updates)
C ONTROL AND C OORDINATION
Managing tasks and system behavior
C OMMUNICATION AND D ECISION S UPPORT
Aiding decision making
S AFETY AND R ISK M ITIGATION
Preventing/responding to hazards
T EAM P ERFORMANCE M ONITORING
Assessing team effectiveness
T EAM C APABILITIES AND AUTHORITY
Defining system capabilities and ensuring human authority
Two-way negotiation (commands and permissions) Interactive (recommendations and alerts) Proactive and reactive (safety features, failure responses) Continuous assessment (monitoring behavior) Static system design and emergency control
Key Differentiator Ensures situational awareness but does not shape actions Ensures the human defines automation’s role and actions Provides guidance rather than just raw data Focuses on preventing harm and handling failures Tracks ongoing performance rather than providing static data Establishes fundamental system abilities and human control mechanisms
TABLE II H IGH - LEVEL CATEGORIZATION OF REQUIREMENTS
C. C OMMUNICATION AND D ECISION S UPPORT Rationale: Requirements in this category center on the provision and clarity of information. This category highlights the role of communication in facilitating decision-making and understanding between human and robot teammates. It includes requirements that are related to decision aids, interface clarity and notifications to support effective HRT. Type of Information: Includes processed information, recommendations, and alerts from the robot to the human. Purpose: Helping the human to make better decisions faster while maintaining situational awareness and trust in the robot. Type of Interaction: Interactive guidance, i.e., robot provides recommendations that the human interprets and acts upon. Key Differentiator: Unlike I NFORMATION S HARING AND AWARENESS, which provides data for passive awareness, this captures data provision that actively shapes human decisionmaking. This includes on-demand decision aids and proactively volunteered data whose nature compels or directs human responses, differing from data updating situational awareness. Example Scenario: A decision aid provides suggested actions, explains its reasoning, and notifies the human when it lacks sufficient data to provide recommendations. D. S AFETY AND R ISK M ITIGATION Rationale: This comprises requirements that ensure safe operation within Human-Robot collaborative environments. It includes preventative safety measures, emergency responses, and failure recovery to protect human operators and robots. Type of information: Safety, failure, and environmental safeguards between human, robot, and workspace. Purpose: Ensuring safe HRT, minimizing the risks to both human and robotic agents. Type of interaction: Proactive safety measures like workspace design. Reactive safety measures like failure recovery. Key Differentiator: Unlike C ONTROL AND C OORDINATION, which focuses on who does what, this category ensures that no matter who is in control, safety is prioritized at all times. Example Scenario: If the robot fails, it enters a safe mode and/or provides alerts to the human. Safety zones ensure that humans and robots operate without interference. E. T EAM P ERFORMANCE M ONITORING Rationale: This category covers the continuous assessment of both human and robot performance to maintain situational
awareness and ensure optimal teamwork. It includes monitoring human behavior as well as robotic system performance. Type of information: Observations of human and robot behavior, performance metrics, and real-time monitoring data. Purpose: Ensuring humans and robots perform optimally by adjusting operations based on observed performance. Type of interaction: Continuous assessment where the robot monitors human behavior and vice versa. Key Differentiator: While I NFORMATION S HARING AND AWARENESS provides static data, this focuses on continuous tracking and evaluation of human and robot performance. Example: The system detects operator fatigue or degraded robot performance and adjusts operations accordingly. F. T EAM C APABILITIES AND AUTHORITY Rationale: Defines fundamental capabilities of robotic and human teammates, ensuring that robots are designed with characteristics to support effective collaboration, while guaranteeing that humans maintain ultimate system authority. Type of information: System design considerations, humanfocused requirements, and override mechanisms. Purpose: Establishing the fundamental capabilities of both human and robot teammates while keeping human authority. Type of interaction: Static design with emergency control: robot has built-in capabilities, while humans retain override. Key Differentiator: Unlike C ONTROL AND C OORDINATION, which manages real-time interactions, this defines the fundamental capabilities and limits of human and robot teammates. C ONTROL AND C OORDINATION governs the runtime allocation of authority, whereas this category establishes the structural constraints and override boundaries for such allocation. Example Scenario: The robot’s designed capabilities (e.g., sensor accuracy, movement speed) determine how well it can collaborate. Meanwhile, the human retains the ability to override or shut down the robot at any time. V. TAXONOMY: L OW- LEVEL CATEGORIES Fig. 1 decomposes each high-level category into subcategories. A. I NFORMATION S HARING AND AWARENESS 1) Robot Data Availability Rationale: Data, such as telemetry, states, modes, etc., should be saved and made available to operators. Example Requirement:
Fig. 1. Hierarchy of categories in our taxonomy. The high-level categories are shown in the blue boxes and the low-level categories are in the orange boxes.
“Automated/robotic systems shall record and make available operational and performance data to both crew and ground support personnel.” [46]
2) System Status Visibility Rationale: Human must maintain situational awareness to calibrate system trust and avoid errors. They need visibility of system health and projected state to assess robot performance. Example Requirement: “Automated/robotic systems shall provide the human operator with the following information: system state and projection of future state, including failure or decrements in performance (e.g., battery power versus traverse distance and assessment of uncertainty in projection of future state).” [46]
3) Robot Mode Change Notification Rationale: Clear indication of the current mode helps prevent errors, where operators take inappropriate actions or fail to act due to misunderstanding the system’s state. Appropriate notification allows the operator to prepare for a mode change. Example Requirement: “The automated/robotic system shall notify the human operator of mode changes of any safety-critical operations.” [46]
4) Human Status Visibility Rationale: As well as passing information from robot to humans (e.g., System Status Visibility), it is crucial that there is a way of transferring information from human to robot. This subcategory encapsulates the requirements relating to robots that must interpret and receive information from humans. Example Requirement: “The automation/autonomy shall enable users to input informal priorities and changes in those priorities.” [44]
B. C ONTROL AND C OORDINATION 1) System Initiation Rationale: The human operator is to remain in control and must authorize a system restart after an emergency stop. Example Requirement:
“If a detected failure occurs, autonomous operation resumes after a deliberate restart from outside the collaborative workspace.” [28]
2) Robot System Configuration Rationale: The human operator must retain control and be able to modify automation configuration parameters, such as setup inputs, initial conditions, and termination conditions. However, certain configurations must remain restricted due to system-specific performance or safety constraints. Example Requirement: “Automated/robotic systems shall provide the human operator the ability to modify system configuration within the safety and performance limits of the system.” [46]
3) Responsibility Delineation Rationale: A clear delineation of the responsibilities of the human and robot is vital to ensure that the operator performs appropriate and timely actions. Transitions between robotic and human control must execute under well-defined procedures that specify the system state before and after transitions. Example Requirement: “Automated/robotic systems shall indicate whether a human operator or system is expected to perform a particular operation at a specific time.” [14]
4) Multi-Agent Coordination Rationale: Requirements specifically relating to the interaction between more than one agent in a team, e.g., the connection between and sharing of information amongst systems. This includes multi-robot and multi-human coordination. Example Requirement: “The automated/robotic system shall make connections between requests from different agents, both human and automated.” [44]
5) Range of Control (Levels of Automation) Rationale: Robotic systems are expected to be operated in different manners; e.g., controlled directly (i.e., manual control) or commanded remotely (i.e., supervisory control). The multiple options associated with these systems must be
operable via the range of controls available to the human operator. Example Requirement:
take measures to remain in a safe state, including alerting the operator of the degraded state and transfer of control. Example Requirement:
“Adaptive automation should be implemented at the point at which the user ignores a critical amount of information.” [14]
“Automated/robotic systems not selected for operation remain in a safe state in a multi-robot cell.” [28]
C. C OMMUNICATION AND D ECISION S UPPORT 1) Decision Support Rationale: Decision aids provide pertinent information, analysis, and/or suggested solutions for continued operations. The system ultimately needs to enable the operator to make those decisions, whether or not it is the operator that acts on them. Example Requirement: “The automated/robotic system shall provide enough information about the developing event so that the operator can intervene effectively.” [44]
2) Decision Transparency Rationale: The human must understand why the robot recommends actions, and the action consequences, to make informed decisions without significantly disrupting the operator’s task, maintaining situational awareness, and calibrating trust. Example Requirement: “Decision aid systems shall provide explanations and rationales, and consequences of potential actions.” [46]
3) Decision Aid Limitation Rationale: Human operators must be alerted when a decision aid cannot assist due to lack of data or design limitations so that data can be provided, or other avenues explored promptly. Example Requirement: “Decision aids shall notify the human operator when a problem or situation is beyond the aid’s capability.” [46]
4) Interface Design Rationale: Communication through displays and interfaces, including how and when warnings should be presented. Example Requirement: “Warning and caution alerts must provide timely attention-getting cues through at least two different senses by a combination of aural, visual, or tactile indications.” [5]
D. S AFETY AND R ISK M ITIGATION 1) Environmental Factors Rationale: Collaborative workspaces must prioritize operator safety and efficiency. Environment considerations such as clearly defined collaborative zones, perimeter safeguarding, and protective measures ensure that robots operate safely within shared spaces, allowing for controlled entry and human oversight. These safeguards create a structured and adaptive environment where humans and robots work together securely. Example Requirement: “The design of the collaborative workspace shall be such that the operator can easily perform all tasks and the location of equipment and machinery shall not introduce additional hazards.” [29]
2) Safety State Preservation Rationale: In a failed transfer of control from the robot to the human, the robot experiencing degraded performance must
3) Robot Safe Mode Rationale: Protective actions include avoidance maneuvers, and protective stops if the human operator can no longer command the system. The transition to operator control needs to occur safely and minimize harm to the operator and vehicle. Example Requirement: “The automated/robotic system shall take protective action (e.g., avoidance maneuver, protective stop) or request that the operator safely take control if the system’s operational safety threshold is exceeded.” [46]
4) Failure Recovery Rationale: If a failure occurs, the data provided by the system must enable the human operator to collect confirming/exclusion evidence to determine a safe course of action. Example Requirement: “Early warning notification of pending automation failure or performance decrements should use estimates of time needed for the user to adjust to task load changes due to automation failure.” [14]
5) Override and Shut-Down Capabilities Rationale: The system allows the human to override/shut down robotic systems if they present a risk, or if redirection is needed. It is vital that the override or shut down is performed safely, i.e., avoids inadvertent harm to crew and vehicle. Example Requirement: “The automated/robotic system shall provide the operator the capability to override the automation and assume partial or full manual control of the system to achieve operational goal states.” [44]
E. T EAM P ERFORMANCE M ONITORING 1) Human Performance Monitoring Rationale: The robotic agent in the team must monitor aspects of human behavior, performance and abilities to maintain understanding of the situation/tasks. The robot must monitor the situational awareness of its human teammate(s) so that they can take actions to improve this awareness if needed [2]. Example Requirement: “The autonomous teammate shall monitor the actions and performance of the human teammate(s).” [58]
F. T EAM C APABILITIES AND AUTHORITY 1) System Capabilities Rationale: Robot capabilities and characteristics are critical in HRT to determine the system’s ability to operate reliably, predictably, and collaboratively with human teammates. These requirements represent capabilities and characteristics that are essential for collaboration in teams. Example Requirement: “Automation should be designed to adapt by providing the most help during times of highest user workload, and somewhat less help during times of lowest workload.” [14]
Category
Total
I NFORMATION S HARING AND AWARENESS Robot Data Availability System Status Visibility Robot Mode Change Notification Human Status Visibility C ONTROL AND C OORDINATION System Initiation Robot System Configuration Responsibility Delineation Multi-Agent Coordination Range of Control (Levels of Automation) C OMMUNICATION AND D ECISION S UPPORT Decision Support Decision Transparency Decision Aid Limitation Interface Design S AFETY AND R ISK M ITIGATION Environmental Factors Safety State Preservation Robot Safe Mode Failure Recovery Override and Shut-Down Capabilities T EAM P ERFORMANCE M ONITORING Human Performance Monitoring T EAM C APABILITIES AND AUTHORITY System Capabilities Human Operational Constraints
56 6 35 10 5 59 4 14 13 5 23 100 49 23 3 25 61 11 10 18 11 11 9 9 76 64 12
Human- Robot- Shared Led Led N Y Y Y Y Y Y Y Y N Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y Y -
TABLE III T HE CATEGORY DISTRIBUTION OF OUR CONSTRUCTION CORPUS . W E CLASSIFY CATEGORIES AS H UMAN -L ED , ROBOT-L ED OR S HARED . T HE TOTAL REQUIREMENTS FOR EACH CATEGORY IS THE SUM OF ITS SUBCATEGORIES . G RAY SHADING INDICATES AN ANALOG IN [46].
2) Human Operational Constraints Rationale: Requirements that focus on the human counterpart rather than the robotic system. This includes requirements relating to the actions and behaviors that a human should withhold, and system requirements that aid the human teammate. Example Requirement: “Users should be given an adequate understanding of how the automated system works in order to monitor effectively.” [14]
VI. M ETRICS AND E XPERT E VALUATION The full list of collected requirements, the slides and interview protocol used for expert interviews are available in [43]. A. Category and Subcategory Metrics Table III collates the 361 requirements in our construction corpus on a per taxonomy category (§IV) and subcategory (§V) basis. Most requirements (100/361) fall into C OMMU NICATION AND D ECISION S UPPORT , mainly in the Decision Support subcategory. The next largest category was T EAM C A PABILITIES AND AUTHORITY , primarily comprising the System Capabilities subcategory. Many of these addressed handover tasks between humans and robots, e.g., [63]. Several requirements in S AFETY AND R ISK M ITIGATION concerned safety mechanisms for robots operating in close proximity to humans, including provisions specific to handover tasks. In Environmental Factors, we found several requirements that referred to the design of the collaborative workspace and physical protective measures. These requirements are also relevant in multi-robot teaming situations.
Within C ONTROL AND C OORDINATION, many requirements fell into the Range of Control (Levels of Automation) subcategory. Many of these were concerned with the robot moving through different degrees of automation/autonomous functionality ranging from tele-operated to fully autonomous. The largest subcategory within I NFORMATION S HARING AND AWARENESS was System Status Visibility. This is unsurprising as the human must remain aware of the robot’s current state, to make informed decisions using their knowledge and trust in the robot. Several requirements specified that the robot maintain awareness of the human teammate’s status, which we classified as Human Status Visibility. We found 9 requirements on Human Performance Monitoring. We did not find any on monitoring robot performance in the construction corpus. Table III indicates (shaded rows) taxonomy categories that have an analogous category in [46], which contributed 14 requirement categories and multiple requirement examples to our taxonomy. We critically refined several of these, updating terminology based on evidence from our search and expert evaluation. For example, we replaced Automation System Status Provision [46] with the System Status Visibility subcategory. We replaced ‘Automation’ with ‘Robot’ across several categories to better reflect the scope of requirements that we identified, which extend beyond automation including broader system components. We adopt ‘Visibility’ in place of ‘Provision’ to more accurately capture the intent of these requirements and to address terminology that consistently caused confusion during expert review. Our taxonomy extends beyond [46] by introducing 7 new categories derived from the wider literature. These categories have no counterpart in existing standards and represent a substantive original contribution to the systematic characterization of HRT requirements. B. Expert Evaluation To validate the taxonomy, we interviewed 5 experts with backgrounds in HRT, Human Factors and RE (2 academics, 2 from large robotics engineering companies and 1 from a space agency). We examined conceptual validity (RQ1) and practical applicability (RQ2). Participants were guided through the 6 high-level categories and asked about conceptual clarity and distinctness, subcategory classification rationale, and whether each category captured the intended class of HRT requirements (RQ1). For RQ2, participants assessed coverage and potential gaps by relating categories to real HRT systems and reflected on the taxonomy’s usefulness for requirements elicitation and specification analysis. We thematically analyzed [41] the discussions into 5 topics: 1) Awareness vs Action: Experts recognized the separation between awareness and action as a useful operational distinction (I NFORMATION S HARING AND AWARENESS, and C OMMUNI CATION AND D ECISION S UPPORT , respectively). However, they noted that information initially intended to support passive understanding can, depending on the context, become relevant or urgent to decisions. As such, systems may need mechanisms to signal changing criticality while still preserving the conceptual boundary between awareness/informing and acting/directing.
Several experts emphasized the importance of shared Situational Awareness (SA) [11] relating to team members’ (human and robot) internal models of external world [9]. SA helps to ensure that both humans and robots maintain a shared and accurate understanding of the environment, tasks, and each other’s states [46]. This enables safe coordination, decisionmaking, and effective responses to unexpected changes. SA was not included as a stand-alone taxonomy category as it is a cross-cutting cognitive outcome rather than a discrete functional concern. Confining it to a single category would obscure the distinct requirement types that collectively support it across the taxonomy. We decompose SA into operational requirement classes [10]. I NFORMATION S HARING AND AWARE NESS supports level 1 (perception) through visibility of system state, C OMMUNICATION AND D ECISION S UPPORT supports level 2 (comprehension) by providing structured rationale and interpretations, and T EAM P ERFORMANCE M ONITORING supports aspects of level 3 (projection) by anticipating performance decrements. However, projection in Human-Robot teams extends beyond these decrements and requires anticipation of authority transition and shifts in responsibility and control, which is not fully represented in our taxonomy. Requirements that enable structured degradation and explicit authority delineation therefore operationalize Level 3 SA as a coordination property rather than a cognitive one, suggesting that SA in HumanRobot teams is closely connected with authority management. 2) Reduced Human to Robot Information and Negotiated Control: While the taxonomy and extracted requirements provide structure for information flow from robot to human, there is less detail of what data the robot should have about its human teammate. Experts highlighted the importance of representing human priorities, condition, and behavior to enable adaptation and coordination. Further, they challenged the implicit assumption that control flows primarily from human to robot, noting situations where the robot must guide, influence, or temporarily supersede human action, especially during error, uncertainty, or time pressure. These observations reframe control as a dynamic, bidirectional exchange grounded in reciprocal awareness rather than a fixed allocation of authority. 3) Ergonomics and Human Variability: Safety cannot be reduced to collision avoidance or fail-safe mechanisms; workspace design, human variability, and fatigue considerations must also be incorporated within ergonomic requirements. Experts further emphasized that effective teamwork requirements must accommodate variability in human behavior and attention, rather than relying on assumptions of consistent compliance. Together, these observations highlight the need to account for physical, cognitive, and behavioral variability in the design of safe and effective human-robot collaboration. 4) Human Performance Monitoring Challenges: All experts regarded this category as essential but challenging. Measuring physical performance is feasible but assessing cognitive state, intention, or understanding is more difficult. One expert extended this further by suggesting future systems may need explainability of human behavior. 5) Additional Categories: Within T EAM P ERFORMANCE M ON -
ITORING , several experts suggested further refinement into subcategories such as Human Performance, Robot Performance, Quality of Interaction, and Monitoring of Operational Parameters. These suggestions reflect a desire to distinguish between physical performance metrics, cognitive or behavioral assessment, and interaction-level dynamics. The utility demonstration in §VII provides empirical support for the experts’ intuition, with requirements involving robot monitoring identified as falling outside the existing subcategory, corroborating it as a candidate for future extension. The experts viewed the taxonomy as conceptually coherent and practical for requirement elicitation. They also identified areas for refinement. As mentioned in §VI-A, we made modifications to several taxonomy categories in light of the feedback from the expert reviewers. We detail the changes in [43].
C. Human-Led, Robot-Led and Shared Classification To answer RQ2, we analyze the classification from three operational perspectives: (1) Human-Led, (2) Robot-Led and (3) Shared (Table III) to expose patterns of authority and interdependence often implicit in existing standards. This framing clarifies how authority boundaries shape safe collaboration. 1) Structural Boundaries of Authority: Our ternary classification aligns with established models of autonomy such as Sheridan’s 10 Levels of Automation (LoA) [51]. We map Human-Led to Low LoA (1–4), Robot-Led to High LoA (7– 10), and Shared to Mid LoA (5–6). This operational analysis reinforces the distinction between runtime authority allocation (C ONTROL AND C OORDINATION) and structural authority constraints (T EAM C APABILITIES AND AUTHORITY). Failure handling exemplifies the boundary. Detection is Robot-Led (the robot autonomously enters safe mode or executes failure recovery), while choice of response is typically Human-Led, and execution is Shared. For effective handover, Failure Recovery requires the robot to convey error severity and type rather than generic failures, so that the human can classify the failure progression and respond accordingly. Human-Led subcategories include the authority to override or shutdown the robot in unsafe situations (Override and ShutDown Capabilities) which is essential to meet legal and ethical requirements, especially in regulated domains. Most subcategories within C ONTROL AND C OORDINATION are Human-Led, reflecting the need for retained human supervisory authority. This override authority interacts with the Robot-Led safety functions in S AFETY AND R ISK M ITIGATION (e.g., Robot Safe Mode). While the robot can autonomously stabilize or enter a safe state, human override remains the final control layer when automated safeguards fail. This supports work on shared control in HRI that emphasizes dynamic arbitration of control authority through intent detection and feedback [42]. These authority boundaries are especially critical during failure, when control transitions under time and safety constraints. 2) Graceful Degradation: A critical concern in Robot-Led teamwork is graceful degradation [27], i.e., the ability of the system to fail safely and support a structured, predictable transition of control to the human under failure or performance
decrement. As robot autonomy increases, the human operator becomes progressively removed from active control, making transition quality increasingly consequential. Requirements addressing graceful degradation are captured primarily within S AFETY AND R ISK M ITIGATION. Effective safety degradation requires explicit specification of how authority transitions during failure, who stabilizes, who diagnoses, and who determines subsequent action. Shared safety protocols and structured responsibility shifts ensure that control transfer occurs predictably and without introducing additional risk, whether that be during Human-Led, Robot-Led or Shared tasks. 3) Shared Awareness: Beyond authority transitions, effective teamwork requires reciprocated modeling between agents. The lack of requirements concerning human cognitive state, adaptability and reciprocal awareness suggests that existing standards prioritize system functionality over modeling the human as an active teammate. For example, if the robot detects elevated cognitive strain in the human, (within Human Performance Monitoring), it may adjust its communication frequency or physical behavior accordingly. Expert feedback highlighted both the limited representation of human cognitive modeling in current requirements and the difficulty of specifying observable indicators of intention, understanding, or trust. Recent work [7] explores user indicators that may prompt explanations, suggesting a pathway toward specifying observable measures of mutual understanding. A related challenge is vigilance decrement [18], [23], i.e., prolonged low workload or monotonous activity reduces alertness and increases susceptibility to error. Requirements for human performance monitoring must capture both overload and under-stimulation. C OMMUNICATION AND D ECISION S UPPORT is Shared, and its effectiveness depends on the human’s ability to understand the rationale behind robot recommendations. 4) Interdependence: Shared functions regulate interdependence by coordinating communication, role delineation and risk management, to support efficient and effective teamwork. These requirements manage the complexities of interdependence, demanding that both (human and robot) agents engage in mutual communication and coordination to mitigate shared risk and ensure role compatibility. For example, C ONTROL AND C OORDINATION is shared when role delineation and transition procedures are in place, and I NFORMATION S HARING AND AWARENESS is shared when status updates are bidirectional. When functions are shared between human and robotic agents, clear delineation of authority is essential. Ambiguity over who is ultimately in charge during shared tasks contributes to coordination failures and safety incidents [17]. Many of the categories and subcategories are Shared (Table III), which underscores the need for a comprehensive delineation of responsibilities in Human-Robot teamwork scenarios. This undoubtedly extends to teams that involve multiple human and robot agents working to achieve a shared goal. These would be explicitly considered in the Multi-Agent Coordination and Responsibility Delineation categories. These observations suggest teamwork depends on authority structures, fault management, and mutual coordination rather
than on hardware setup or the number or type of agents. VII. U TILITY D EMONSTRATION We conducted a utility demonstration, per Usman et al.’s guidance on taxonomy validation [59], as a second measure to validate our taxonomy. A. Construction of an Independent Validation Corpus We assembled a second corpus (available in [43]) of requirements from 19 sources, spanning 6 HRT domains: Aerospace & Aviation [16], [64], Emergency & First Response [54], [48], [26], [22], Healthcare [61], [38], [55], Industrial Manufacturing [20], [21], [66], [3], [68], Defense & Security [52], [70], and Social & Service Robotics [53], [36], [67]. These were not used to construct the taxonomy in §III. Sources were selected via purposive sampling [12], to evaluate applicability across diverse contexts rather than achieving exhaustive coverage. Source identification used two strategies. First, we conducted targeted searches of Google Scholar and NASA’s Technical Reports Server (NTRS) using the search terms (individually and combined): CONOPS, requirements, technical document, rover, robot, ARTEMIS requirements, human-machine interaction. Second, we manually reviewed proceedings from the International Conference on Human-Robot Interaction (HRI) (2016–2026), extracting requirement statements from 9 relevant papers. We chose HRI as a top peer-reviewed venue to ensure the validation corpus reflects representative humanrobot interaction research. To ensure that any inability to classify a requirement could be attributed to gaps in the taxonomy rather than methodological drift between the corpora, we reused the inclusion and exclusion criteria from §III. B. Extraction and Classification Procedure We followed the same extraction protocol used for the construction corpus in §III. Two researchers independently extracted candidate requirements and applied the inclusion criteria, with an initial inclusion agreement of 91.09%, compared to 87% for the construction corpus. From an initial 823 candidate statements, 448 were retained as in-scope after filtering, 375 were excluded as they described single-agent autonomy or contexts different than teamwork requirements. Requirements were independently categorized by two researchers, achieving 94% initial agreement. Disagreements clustered at boundaries between adjacent categories rather than being distributed randomly, indicating that classification was straightforward for most requirements and ambiguity arose where two categories had overlapping concerns. These were resolved through discussion against the category definitions. C. Coverage results Table IV shows the distribution of 448 requirements across taxonomy subcategories and domains. Classification was achieved for 91.9% (412/448) of in-scope requirements using the existing taxonomy, demonstrating that the 6 highlevel categories and 21 subcategories provide broad coverage across domains not in the construction corpus. The remaining
requirements could each be allocated a high-level category but did not fit an existing subcategory (beige rows, Table IV) and are classed as candidates for taxonomy extension. Convergence with construction corpus: The distribution of requirements closely mirrors that of the construction corpus (Table III). C OMMUNICATION AND D ECISION S UPPORT remains the largest category (29.2% here and 27.7% in the construction corpus), followed by T EAM C APABILITIES AND AUTHORITY (23.2% here and 21.1% in the construction corpus). The T EAM P ERFORMANCE M ONITORING is again the sparsest (4.7% here and 2.5% in the construction corpus). The three middle-ranked categories reorder slightly but each remains within 13–17% in both corpora. The convergence of these patterns across two independently assembled corpora suggests that they reflect the structure of the field rather than artifacts of source selection. Domain-specific characteristics: Table IV shows that each domain focuses on different categories. Industrial Manufacturing is led by S AFETY AND R ISK M ITIGATION, where most requirements (39/43) belong in the Environmental Factors subcategory, reflecting shared workspace design standards [29]. Healthcare is led by C OMMUNICATION AND D ECISION S UPPORT, as these requirements (42%) often define how data is shown to clinicians/patients. This is mirrored in Social & Service Robotics with 38.7% of requirements in this category. In contrast, Emergency & First Response had a low incidence of requirements in this category (5%) with T EAM C APABILITIES AND AUTHORITY being its most populous category (49%). Defense & Security focuses on C ONTROL AND C OORDINATION reflecting the importance of adjustable range of automation in mission-critical settings. These patterns show that the taxonomy generalizes across domains without flattening their differences, supporting its use as a shared reference framework. Sparsely populated categories: Several subcategories remain sparsely populated in both corpora. T EAM P ERFORMANCE M ONITORING contains only 21 requirements in the validation corpus, being completely absent from Social & Service Robotics. This persistent under-representation, observed across both corpora and corroborated by expert evaluation (§VI-B), suggests that human-side monitoring is under-specified in current practice. This illustrates how the taxonomy can act as a diagnostic instrument that surfaces systematic absences and directs elicitation effort toward the concerns most likely to be missing from a specification. A similar pattern holds for Failure Recovery (1/448), System Initiation (3/448), Override and Shutdown Capabilities (4/448), and Decision Aid Limitation (no representation in the validation corpus). These subcategories collectively describe how a Human-Robot team behaves at the boundary of nominal operation, meaning at failure, restart and handover. Their consistent under-representation across different sources indicates a genuine specification gap. Candidates for taxonomy extension: Some validation-corpus requirements (36/448) fit a high-level category but not any existing subcategory. These are candidates for extension rather than signs of a structural weakness since all requirements were classifiable at the high-level category. From preliminary
Category I NFORMATION S HARING AND AWARENESS Robot Data Availability System Status Visibility Robot Mode Change Notification Human Status Visibility C ONTROL AND C OORDINATION System Initiation Robot System Configuration Responsibility Delineation Multi-Agent Coordination Range of Control (Levels of Automation) C OMMUNICATION AND D ECISION S UPPORT Decision Support Decision Transparency Decision Aid Limitation Interface Design Other S AFETY AND R ISK M ITIGATION Environmental Factors Safety State Preservation Robot Safe Mode Failure Recovery Override and Shut-Down Capabilities T EAM P ERFORMANCE M ONITORING Human Performance Monitoring Other T EAM C APABILITIES AND AUTHORITY System Capabilities Human Operational Constraints Other Total
Total 70 10 40 6 14 62 3 15 16 8 19 131 22 23 0 64 21 60 40 7 8 1 4 21 15 6 104 78 17 9 448
AA EF 10 10 2 3 6 7 2 6 9 1 1 1 4 3 4 1 24 3 9 1 8 7 1 3 8 1 1 4 1 3 1 5 1 5 1 13 30 12 17 1 13 61 61
HC IM DS 11 20 7 2 1 2 5 14 8 1 2 4 4 11 17 17 2 1 6 4 1 4 3 4 5 3 11 32 40 8 7 4 1 4 4 18 17 4 6 15 1 43 4 39 2 3 1 1 1 1 6 6 3 6 3 1 3 2 14 19 14 13 19 11 1 1 2 76 164 58
SS 12 3 4 2 2 24 1 6 17 1 1 14 6 1 7 62
TABLE IV D ISTRIBUTION OF 448 UTILITY ANALYSIS REQUIREMENTS ACROSS TAXONOMY CATEGORIES AND DOMAINS . Other SUBCATEGORY ENTRIES INDICATE CANDIDATES FOR TAXONOMY EXTENSIONS . AA = Aerospace & Aviation, EF = Emergency & First Response, HC = Healthcare, IM = Industrial Manufacturing, DS = Defense & Security, SS = Social & Service.
analysis, these fall into candidate subcategories: requirements for assessing robot performance as a symmetric counterpart to Human Performance Monitoring (n=6 in T EAM P ERFORMANCE M ONITORING); requirements concerning the robot’s social presentation and sustaining user engagement over time, rather than any specific information exchange or decision (n=21 in C OMMUNICATION AND D ECISION S UPPORT); and requirements relating to explainability and human acceptance that sit below the interaction level addressed by C OMMUNICATION AND D ECI SION S UPPORT , concerning instead structural prerequisites for explainability and conditions for accepting robotic teammates (n=9 in T EAM C APABILITIES AND AUTHORITY). Full definitions and empirical validation of these subcategories is future work.
VIII. U SING THE TAXONOMY IN P RACTICE The taxonomy is intended for both RE and HRT researchers working on autonomous and AI-enabled systems, and systems engineering practitioners in safety-critical domains who must produce HRT specifications. For both audiences, we position the taxonomy with respect to the RE process described in standards such as ISO/IEC/IEEE 29148 [31]. The taxonomy targets the early stages of this process, i.e., requirements elicitation and specification review, where structural coverage of teamwork concerns is most likely to be missed. It is not intended to replace later stages such as
detailed requirement formalization, verification, or compliance demonstration, but to provide a structured input to them. We see the following primary uses: Elicitation: Without a structured reference, requirement types are easily missed. The 21 subcategories provide a guidance list that helps practitioners cover teamwork concerns, with the examples in §V acting as starting points. This could be paired with the interview guide proposed by [44] which provides leading questions to help with the elicitation process. Coverage assessment and gap diagnosis: A draft requirement specification can be mapped onto the taxonomy to identify which categories are well-covered and which are sparse. As shown in §VII, several subcategories are consistently underrepresented across two independently assembled corpora. Practitioners can use the taxonomy as a diagnostic instrument to flag likely blind spots in their own specifications. Cross-domain comparison: The taxonomy allows requirement profiles to be compared across domains, which is relevant when the same robot platform is deployed across multiple contexts. For example, Boston Dynamics’ Spot has been used in Industrial Manufacturing, Defense & Security, and Healthcare [40]. Each of these domains has a different requirement profile: Industrial Manufacturing concentrates on S AFETY AND R ISK M ITIGATION , Healthcare on C OMMUNICATION AND D ECISION S UPPORT while Defense & Security focuses primarily on C ONTROL AND C OORDINATION and T EAM C APABILITIES AND AUTHORITY . The taxonomy helps practitioners identify which requirements transfer across these settings, which need domain-specific revision, and where gaps are likely to appear. IX. T HREATS TO VALIDITY We endeavored to be comprehensive throughout the development of our taxonomy but recognize several threats to validity and limitations. To begin, our construction corpus was obtained via a structured literature review (§III). Literature reviews are subject to coverage error, meaning that we may have missed relevant literature that could have been revealed if we had used different search terms. We mitigated this by starting our search with well-established literature and applying snowballing, in addition to systematic keyword searches. Further, although we found “Human-Machine Teaming” resources (e.g., [44]), we did not include this as a search term as we wanted to focus on robotics, rather than broader team structures such as cognitive assistants. Including it might have provided additional resources for developing the taxonomy, but it could have reduced its specificity to robotics, which was a key objective. Source selection bias could be present since we focused on English-language sources and Western standards and our validation corpus drew primarily from the HRI conference. Different venues or search keywords may have surfaced additional requirement types. While the expert evaluation serves as a key mechanism for mitigating coverage-related threats and refining the taxonomy, it is itself subject to limitations. Our panel of experts, though deliberately recruited from diverse backgrounds, may not be
fully representative of the breadth of HRT expertise. Nevertheless, convergent feedback between all experts provides confidence in the taxonomy’s broader applicability which is further examined by our utility demonstration. There is a wealth of standards in the aerospace domain, since it is global and highly-regulated, a dominance of aerospace-related results in our construction corpus threatens domain coverage. However, rather than undermining the taxonomy, this reflects the current state of the field, i.e., aerospace and industrial robotics represent some of the most mature and documented areas of HRT, making them an appropriate foundation for an initial taxonomy. We explicitly recognize that emerging domains may surface additional requirement types, and our utility demonstration broadened this to include requirements from diverse domains (see Table IV). Finally, the process of categorizing requirements into a taxonomy involves interpretive judgment. To mitigate construct validity threats, the taxonomy was developed iteratively across multiple researchers and refined through expert review, ensuring that classification decisions were not made unilaterally. All requirements in the construction corpus were successfully classified without needing a residual “other” category, which is an indicator of structural completeness with respect to the construction corpus. Our utility demonstration identified a small number of requirements that fell into an ’other’ category. The taxonomy will thus be extended to incorporate and evaluate these new subcategories in future work. X. C ONCLUSIONS AND F UTURE W ORK This paper presents a two-level taxonomy of HRT requirements based on 361 requirements from standards and the literature. The taxonomy was developed with the aim to support requirement elicitation, coverage assessment, and cross-domain traceability. We evaluate our taxonomy through an expert evaluation, and a utility demonstration using an independently-constructed corpus of 448 requirements. Future work will comprise broader validation across diverse human-robot team configurations, including large-scale industrial case studies. We will also examine whether taxonomy categories exhibit domain-specific precedence. For example, safety may take priority even when information sharing is limited such as in [33]. Finally, we will investigate the extent to which taxonomy categories can be formalized to support verifiable requirement specification patterns (e.g., [60], [13]). Data Availability: Our corpora are publicly available in [43]. Acknowledgments: We are grateful to Amber Drinkwater, Laura Hoang, Lukman Irshad, Caroline Jay, Federico Tavella, and Hannah Walsh for insightful discussions and helpful feedback. We also thank the anonymous reviewers, whose constructive comments substantially strengthened the paper. This work was supported by VeTSS and the Royal Academy of Engineering through a Research Fellowship and a Chair in Emerging Technologies. This work was also supported by the Centre for Robotic Autonomy in Demanding and Long Lasting Environments (CRADLE) under EPSRC grant EP/X02489X/1. Anastasia was supported by NASA contract 80ARC020D0010.
R EFERENCES [1] A. Agrawal, S. J. Abraham, B. Burger, C. Christine, L. Fraser, J. M. Hoeksema, S. Hwang, E. Travnik, S. Kumar, W. Scheirer, et al. The Next Generation of Human-Drone Partnerships: Co-Designing an Emergency Response System. In CHI Conference on Human Factors in Computing Systems, pages 1–13, 2020. [2] A. Ali, L. P. Robert, and D. M. Tilbury. Estimating situation awareness for human-robot teaming. In 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), pages 1216–1221. IEEE, 2025. [3] C. Boonyard, S. S. Johansen, C. Jouffrais, A. M. Brock, and T. Merritt. Speaking of safety: Drone-based voice communication for industry. In Proceedings of the 21st ACM/IEEE International Conference on HumanRobot Interaction, pages 785–795, 2026. [4] W. J. Clancey, M. Sierhuis, C. Kaskiris, and R. van Hoof. Advantages of Brahms for Specifying and Implementing a Multiagent Human-Robotic Exploration System. In International Florida Artificial Intelligence Research Society Conference, pages 7–11. AAAI Press, 2003. [5] Code of Federal Regulations. Title 14 - Aeronautics and Space. Technical report, 2019. [6] P. Damacharla, A. Y. Javaid, J. J. Gallimore, and V. K. Devabhaktuni. Common Metrics to Benchmark Human-Machine Teams (HMT): A Review. IEEE Access, 6:38637–38655, 2018. [7] H. Deters, L. Reinhardt, J. Droste, M. Obaidi, and K. Schneider. Identifying Explanation Needs: Towards a Catalog of User-Based Indicators. In International Requirements Engineering Conference, pages 31–42. IEEE, 2025. [8] J. Droste, H. Deters, M. Obaidi, and K. Schneider. Explanations in Everyday Software Systems: Towards a Taxonomy for Explainability Needs. In International Requirements Engineering Conference, pages 55–66. IEEE, 2024. [9] M. R. Endsley. Design and Evaluation for Situation Awareness Enhancement. In Proceedings of the Human Factors Society Annual Meeting, volume 32, pages 97–101. Sage Publications, 1988. [10] M. R. Endsley. A Taxonomy of Situation Awareness Errors. Human Factors in Aviation Operations, 3(2):287–292, 1995. [11] M. R. Endsley. Toward a Theory of Situation Awareness in Dynamic Systems. In Situational Awareness, pages 9–42. Routledge, 2017. [12] I. Etikan, S. A. Musa, R. S. Alkassim, et al. Comparison of convenience sampling and purposive sampling. American journal of theoretical and applied statistics, 5(1):1–4, 2016. [13] M. Etumi, H. M. Taylor, and M. Farrell. Towards A Catalogue of Requirement Patterns for Space Robotic Missions. In International Workshop on Formal Methods for Autonomous Systems, pages 136–166. EPTCS, 2025. [14] FAA Federal Aviation Administration. Human Factors Design Standard (HF-STD-001). Technical report, William J. Hughes Technical Center, 2016. [15] M. Fisher, R. C. Cardoso, E. C. Collins, C. Dadswell, L. A. Dennis, C. Dixon, M. Farrell, A. Ferrando, X. Huang, M. Jump, et al. An Overview of Verification and Validation Challenges for Inspection Robots. Robotics, 10(2):67, 2021. [16] A. Fuchs, C. Bejarano, A. Metge, S. Ruano, J. M. Cordero, A. Perillo, P. Vaiopoulos, G. Fedrizzi, and A. G. Vicario. Safe return-to-land operations in future cockpits: An analysis of cases and mitigation technologies. In International Conference on Human-Computer Interaction, pages 54–73. Springer, 2025. [17] A. Fuchs, A. Passarella, and M. Conti. Optimizing Delegation in Collaborative Human-AI Hybrid Teams. ACM Transactions on Autonomous and Adaptive Systems, 19(4), Nov. 2024. [18] E. T. Greenlee, P. R. DeLucia, and D. C. Newton. Driver Vigilance in Automated Vehicles: Effects of Demands on Hazard Detection Performance. Human Factors, 61(3):474–487, 2019. [19] E. C. Grigore, K. Eder, A. Lenz, S. Skachek, A. G. Pipe, and C. Melhuish. Towards Safe Human-Robot Interaction. In Towards Autonomous Robotic Systems, pages 323–335. Springer, 2011. [20] L. Gualtieri, F. Fraboni, H. Brendel, L. Pietrantoni, R. Vidoni, and P. Dallasega. Updating design guidelines for cognitive ergonomics in human-centred collaborative robotics applications: An expert survey. Applied ergonomics, 117:104246, 2024. [21] L. Gualtieri, E. Rauch, R. Vidoni, and D. T. Matt. Safety, ergonomics and efficiency in human-robot collaborative assembly: design guidelines and requirements. Procedia CIRP, 91:367–372, 2020.
[22] M. Hagenow, E. Senft, R. Radwin, M. Gleicher, M. Zinn, and B. Mutlu. A system for human-robot teaming through end-user programming and shared autonomy. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, pages 231–239, 2024. [23] P. Hancock. Human Vigilance in the Age of Intelligent Machines: Challenges and Prospects. Ergonomics, pages 1–17, 2026. [24] H. Hendry, A. Cavalcanti, C. McCall, and M. Chattington. RoboScene: Notation for Formal Verification of Human-Robot Interaction. In International Conference on Fundamental Approaches to Software Engineering, pages 166–187. Springer, 2025. [25] J. Huang, F. I. Doğan, and H. Gunes. Reimagining social robots as recommender systems: Foundations, framework, and applications. In ACM/IEEE International Conference on Human-Robot Interaction, pages 406–416, 2026. [26] C. M. Humphrey and J. A. Adams. Human roles for robot augmented first response. In 2015 IEEE International Symposium on Safety, Security, and Rescue Robotics (SSRR), pages 1–6. IEEE, 2015. [27] T. Ishigooka, S. Otsuka, K. Serizawa, R. Tsuchiya, and F. Narisawa. Graceful Degradation Design Process for Autonomous Driving System. In International Conference on Computer Safety, Reliability, and Security, pages 19–34. Springer, 2019. [28] ISO. ISO 10218-2:2025 Technical Standard - Robotics — Safety Requirements —Part 2: Industrial Robot Applications and Robot Cells. Technical report, 2011. [29] ISO. ISO-10218 Technical Standard - Robots and Robotic Devices — Safety Requirements for Industrial Robots —Part 1: Robots. Technical report, 2011. [30] ISO. ISO/TS 15066 Robots and robotic devices—Collaborative robots. Technical report, 2016. [31] Systems and software engineering — life cycle processes — requirements engineering. Standard ISO/IEC/IEEE 29148:2018, International Organization for Standardization / Institute of Electrical and Electronics Engineers, Geneva, Switzerland, 11 2018. [32] B. Jarvis, D. Jarvis, and L. Jain. Teams in Multi-Agent Systems. In International Conference on Intelligent Information Processing, pages 1–10. Springer, 2006. [33] L. Jia, B. K. Styler, and N. Du. More than automation: User insights into the functionality and interface of wheelchair-mounted robotic arms. In ACM/IEEE International Conference on Human-Robot Interaction (HRI), pages 677–685. IEEE, 2025. [34] T. E. Jost, P. Grünbacher, and C. Stary. Digital Process Twins for Interleaving Requirements Elicitation and Design of Cyber-Physical Systems. In International Requirements Engineering Conference, pages 345–353. IEEE, 2024. [35] S. Kim, J. R. Anthis, and S. Sebo. A taxonomy of robot autonomy for human-robot interaction. In ACM/IEEE International Conference on Human-Robot Interaction, pages 381–393, 2024. [36] Y. Kim, C. P. Lee, and B. Mutlu. Speaking with screens: Design space and guidelines for informational robot screens. In Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction, pages 385–394, 2026. [37] G. Klien, D. D. Woods, J. M. Bradshaw, R. R. Hoffman, and P. J. Feltovich. Ten Challenges for Making Automation a “Team Player”’ in Joint Human-Agent Activity. IEEE Intelligent Systems, 19(6):91–95, 2005. [38] A. Kubota, D. Cruz-Sandoval, S. Kim, E. W. Twamley, and L. D. Riek. Cognitively assistive robots at home: Hri design patterns for translational science. In 2022 17th ACM/IEEE International Conference on HumanRobot Interaction (HRI), pages 53–62. IEEE, 2022. [39] B. H. Kwasnik. The role of classification in knowledge representation and discovery. 1999. [40] W. Li, G. Yang, J. Wu, C. Pan, L. Sheng, and Q. Zhang. A comprehensive review on humanoid robots: perspectives from academia and industry. ENGINEERING Information Technology & Electronic Engineering, 27(2):1–37, 2026. [41] C. R. Lochmiller. Conducting Thematic Analysis with Qualitative Data. The Qualitative Report, 26(6):2029–2044, 2021. [42] D. P. Losey, C. G. McDonald, E. Battaglia, and M. K. O’Malley. A Review of Intent Detection, Arbitration, and Communication Aspects of Shared Control for Physical Human–Robot Interaction. Applied Mechanics Reviews, 70(1):010804, 2018. [43] A. Mavridou, H. M. Taylor, S. Lozito, L. A. Dennis, M. Fisher, and M. Farrell. Teamwork Taxonomy Repository. https://github.com/
anmavrid/Teamwork Requirements Taxonomy, 2026. Accessed: 202606-15. [44] P. McDermott, C. Dominguez, N. Kasdaglis, M. Ryan, I. Trhan, and A. Nelson. Human-Machine Teaming Systems Engineering Guide. Technical report, MITRE, 2018. [45] K. Nam, J. Yee, R. A. Jinnat, K. K. Rigual, A. Thylane, and R. Korpan. Not the intended user: Queer perspectives on identity, risk, and trust in robot companions. In ACM/IEEE International Conference on HumanRobot Interaction, pages 417–426, 2026. [46] NASA National Aeronautics and Space Administration. NASA Spaceflight Human-System Standard Vol 2: Human Factors, Habitability, and Environmental Health. Technical report, 2023. [47] R. C. Nickerson, U. Varshney, and J. Muntermann. A method for taxonomy development and its application in information systems. European journal of information systems, 22(3):336–359, 2013. [48] K. S. Pratt, R. Murphy, S. Stover, and C. Griffin. Conops and autonomy recommendations for vtol mavs based on observations of hurricane katrina uav operations. J. Field Robot, 26:636–650, 2006. [49] L. Rierson. Developing Safety-Critical Software: A Practical Guide for Aviation Software and DO-178C Compliance. CRC Press, 2017. [50] P. Segura, O. Lobato-Calleros, A. Ramı́rez-Serrano, and I. Soria. Human-Robot Collaborative Systems: Structural Components for Current Manufacturing Applications. Advances in Industrial and Manufacturing Engineering, 3:100060, 2021. [51] T. B. Sheridan and W. L. Verplank. Human and Computer Control of Undersea Teleoperators. Technical report, Massachusetts Institute of Technology, 1978. [52] J. Shively. Human-autonomy teaming: Supporting dynamically adjustable collaboration. Technical report, 2017. [53] S. Stange, T. Hassan, F. Schröder, J. Konkol, and S. Kopp. Selfexplaining social robots: An explainable behavior generation architecture for human-robot interaction. Frontiers in Artificial Intelligence, 5:866920, 2022. [54] B. Streiffert and C. B. S. JPL. Autonomous Rover for Ground-based Optical Surveillance (ARGOS). [55] A. Taylor, T. Tanjim, M. J. Sack, M. Hirsch, K. Cheng, K. Ching, J. S. George, T. Roumen, M. F. Jung, and H. R. Lee. Rapidly built medical crash cart! lessons learned and impacts on high-stakes team collaboration in the emergency room. In 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI), pages 501–510. IEEE, 2025. [56] H. M. Taylor, M. Luckuck, M. Farrell, C. Jay, A. Cangelosi, and L. Dennis. Eliciting Explainability Requirements for Safety-Critical Systems: A Nuclear Case Study. In International Working Conference on Requirements Engineering: Foundation for Software Quality, 2025. [57] H. M. Taylor, A. Mavridou, M. Farrell, and L. A. Dennis. Explainability Pattern Specifications for Human-Robot Teamwork. In International Conference on Engineering Reliable Autonomous Systems. IEEE, 2025. [58] G. Tokadlı and M. C. Dorneich. Characteristics of Human-Autonomy Teaming for Future Aerospace Operations. In IEEE/AIAA Digital Avionics Systems Conference, pages 1–7. IEEE, 2022. [59] M. Usman, R. Britto, J. Börstler, and E. Mendes. Taxonomies in software engineering: A systematic mapping study and a revised taxonomy development method. Information and Software Technology, 85:43–59, 2017. [60] G. Vázquez, A. Mavridou, M. Farrell, T. Pressburger, and R. Calinescu. Robotics: A New Mission for FRET Requirements. In NASA Formal Methods Symposium, pages 359–376. Springer, 2024. [61] J. Wang, C. Wang, R. Soltani Zarrin, and Z. Erickson. Bidirectional human-robot communication for physical human-robot interaction. In Proceedings of the 21st ACM/IEEE International Conference on HumanRobot Interaction, pages 572–580, 2026. [62] M. Webster, L. A. Dennis, C. Dixon, M. Fisher, R. Stocker, and M. Sierhuis. Formal Verification of Astronaut-Rover Teams for Planetary Surface Operations. In IEEE Aerospace Conference, pages 1–8, 2020. [63] M. Webster, D. Western, D. Araiza-Illan, C. Dixon, K. Eder, M. Fisher, and A. G. Pipe. A Corroborative Approach to Verification and Validation of Human–Robot Teams. The International Journal of Robotics Research, 39(1):73–99, 2020. [64] C. Westin, K. J. Klang, J. Basjuka, G. Söderholm, J. Lundberg, M. Bång, K. Lundin Palmerius, S. Boonsong, Å. Taraldsson, and G. Fylkner. Human-ai teaming in the urban air mobility coordinator work position: A proof-of-concept design. In International Conference on HumanComputer Interaction, pages 256–275. Springer, 2025.
[65] C. Wohlin. Guidelines for Snowballing in Systematic Literature Studies and a Replication in Software Engineering. In International Conference on Evaluation and Assessment in Software Engineering. ACM, 2014. [66] J. Wu, J. Ren, O. Ravn, and L. Nalpantidis. A risk-informed design framework for functional safety system design of human–robot collaboration applications. Safety, 11(1):24, 2025. [67] M. F. Xu, E. Zhao, Y. Zhang, J. E. Michaelis, S. Sebo, and B. Mutlu. Designing robots for families: In-situ prototyping for contextual reminders on family routines. In Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction, pages 356–365, 2026. [68] E. Yang and C. Mavrogiannis. Implicit communication in humanrobot collaborative transport. In 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI), pages 23–33. IEEE, 2025. [69] S. Yang, J. Park, S. You, and Y. Chang. Why do service robots fail? a systematic literature review from a service design perspective. In ACM/IEEE International Conference on Human-Robot Interaction, pages 108–118, 2026. [70] X. Ye, W. Jo, A. Ali, S. C. Bhatti, C. Esterwood, H. A. Kassie, and L. P. Robert. Autonomy acceptance model (aam): The role of autonomy and risk in security robot acceptance. In Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, pages 840–849, 2024. [71] M. Yeh, C. Swider, Y. J. Jo, C. Donovan, et al. Human Factors Considerations in the Design and Evaluation of Flight Deck Displays and Controls: Version 2.0. Technical report, John A. Volpe National Transportation Systems Center, 2016.