2026-5-4
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving Ashwin George1,2,*,+ , Lucas Elbert Suryana1,3,* , Lorenzo Flipse1,2 , Bart van Arem1,3 , David A. Abbink1,2 , Simeon Craig Calvert1,3 , Luciano Cavalcante Siebert1,4 and Arkady Zgonnikov1,2 1 Centre for Meaningful Human Control, 2 Department of Cognitive Robotics, 3 Department of Transport & Planning, 4 Department of
arXiv:2605.00556v1 [cs.HC] 1 May 2026
Intelligent Systems; Delft University of Technology, * Authors contributed equally and listed in alphabetical order, + Corresponding author: [email protected], Mekelweg 2, 2628 CD, Delft, Zuid Holland, The Netherlands.
Partial driving automation creates a tension: drivers remain legally responsible for vehicle behaviour, yet their active control is significantly reduced. This reduction undermines the engagement and sense of agency needed to intervene safely. Meaningful human control (MHC) has been proposed as a normative framework to address this tension. However, empirical methods for evaluating whether existing systems actually provide MHC remain underdeveloped. In this study, we investigated the extent to which drivers experience MHC when interacting with partially automated driving systems. Twenty-four drivers completed a simulator study involving silent automation failures under two modes haptic shared control (HSC) and traded control (TC). We derived behavioural metrics from telemetry data, subjective perception scores from post-trial surveys and used them to test hypothesised relations between them derived from the properties of systems under MHC. The confirmatory analysis showed a significant negative correlation between the perception of the automated vehicle (AV) understanding the driver and conflict in steering torques. An exploratory analysis also revealed a surprising positive correlation between reaction times and the perception of sufficient control. Qualitative feedback from open-ended post-experiment questionnaires revealed that mismatches in intentions between the driver and automation, lack of safety, and resistance to driver inputs contribute to the reduction of perceived MHC, while subtle haptic guidance aligned with driver intent had a positive effect. These findings suggest that future designs should prioritise effortless driver interventions, transparent communication of automation intent, and context-sensitive authority allocation to strengthen meaningful human control in partially automated driving.
1. Introduction Responsibility goes hand-in-hand with the level of control [1]. For example, the driver of a car has a much higher degree of responsibility and control over the vehicle as compared to a passenger in a train. This balance of responsibility and control can be disrupted when automation systems take up part of the control while the drivers of such vehicles are still held responsible for accidents [2]. Therefore, understanding driver’s perception of responsibility and control in partially automated vehicles is essential for designing vehicle automation. Vehicle automation is widely seen as a promising way to improve road safety and traffic efficiency [3]. Early deployments of fully automated systems have demonstrated benefits, such as reductions in specific crash types [4]. In practice, however, most vehicles today operate at SAE levels 2–3 [5], where automation handles some driving tasks while human drivers must supervise the system and intervene when necessary. These levels create a paradox: drivers are expected to remain constantly attentive even though their control role is significantly reduced [6]. One manifestation of this paradox is driver complacency, an over-reliance on automation in which drivers’ reactions slow when manual intervention is required due to reduced vigilance. This decline occurs when automation shifts drivers’ role from active control to passive supervision, reducing engagement [7]. In human factors terms, this reflects a vigilance decrement caused by cognitive underload during passive monitoring
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
Subjective Perception
Perceive
Act
Human reacts Human intention AV intention
Human - AV Interaction
Control
Act
Behavioural Metrics
Perceive
Haptic Shared Control
Control Modes
Both human and automation in control
Traded Control
Either human or automation in control
Figure 1 | Understanding the relationship between the subjective perception of human drivers and behavioural metrics observable to machines is vital for designing automated systems under meaningful human control. This study explores how different control modes — haptic share control (HSC) and traded control (TC) affect perceptions of control and responsibility when the human and automated system have conflicting reasons. [6, 8]. Such disengagement is critical in scenarios where the driver must retake control unexpectedly: if unprepared, the driver may respond too slowly, increasing accident risk. At the same time, some commercial systems are marketed as being capable of “autonomous driving”, even though manufacturers specify that drivers must remain ready to intervene when needed. These conflicting expectations reinforce the above paradox of automation: while the systems are presented as relieving the driver of control, safe operation still depends on sustained human vigilance raising questions about whether drivers can reasonably be assigned full moral and legal responsibility under these conditions. Beyond the paradox of automation, partial automation can also erode drivers’ felt responsibility for vehicle behaviour, even though they remain legally accountable. This reduced sense of responsibility is often explained through the sense of agency (SoA), the subjective experience of initiating and controlling an action. Prior work shows that SoA tends to decline as automation increases, with higher autonomy reducing awareness and engagement [9, 10, 11]. When individuals feel they did not cause an action, they are also less likely to accept moral responsibility for its outcomes [12]. Taken together, complacency and declining SoA create a layered challenge: systems expect drivers to remain attentive and accountable, yet their design undermines the very capacities required to do so. This mismatch produces misattribution of responsibility [13, 14], where it is unclear who should be responsible (and held accountable) for the actions of an automated system. Recent evaluations of commercial driving systems illustrate these gaps: drivers often lack understanding of how well the system can react (or how well they themselves can react) while manufacturers frequently distance themselves from responsibility through disclaimers [15]. To mitigate such responsibility gaps, meaningful human control (MHC) has been proposed as a normative stance that humans should maintain some form of control and responsibility over the behaviour of artificial systems, even if these systems perform tasks autonomously [16]. Automated systems (including driving automation) must meet two conditions to be under MHC — 1) the tracking condition which asserts that the behaviour of automated systems must track the reasons of relevant humans and 2) the tracing condition which asserts that it should be possible to trace the responsibility for the behaviours of the automated system to at least one human [16]. Based on these two fundamental conditions, subsequent work proposed a range of practice-oriented frameworks to operationalize tracking and tracing, such as actionable properties of Cavalcante Siebert et al. [17] or MHC operationalisation framework of Calvert [18].
2
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
While MHC has often been discussed as a guiding principle for system design [16, 19, 17, 18], a necessary first step is to evaluate whether an existing system satisfies its conditions in practice. Without such evaluation, MHC risks being practically under-validated. Previous studies have proposed several ways to evaluate MHC of specific systems [15, 20] or general classes of such systems [21], but assess it only indirectly. For example, MHC has been assessed indirectly through post-crash analyses [15] or user interviews [20]. However, evaluations based on post-crash analyses do not capture the nuances of real-time interaction between the driver and the automated system, while those based solely on user interviews rely on questions not specifically designed to measure MHC, even though partial alignment with MHC principles can be inferred. More generally, the current literature lacks methods that directly evaluate whether the interaction between human drivers and automated driving systems provides for meaningful human control. In this work, we address this gap by combining subjective and behavioural measures to evaluate meaningful human control in driver interactions with automated driving systems. Specifically, in our evaluation methodology (Fig. 1), we integrate detailed driving interaction data with a questionnaire explicitly designed to assess MHC based on the established properties [17]. This approach is inspired by Verhagen et al. [22], who combined subjective reports and interaction metrics to evaluate MHC in human-robot interaction for firefighting systems. To demonstrate our evaluation methodology, we focus on interaction strategies currently implemented in automated vehicles. We use two common interaction strategies, haptic shared control (HSC) and traded control (TC), as representative approaches to explore how such systems align with MHC in practice. In HSC, both the driver and the automation act simultaneously through force feedback on the steering wheel, supporting smoother collaboration and reducing conflict [23, 24, 25]. In contrast, TC relies on explicit handovers, which can be simpler to implement but risk disorientation if poorly timed [26]. With these two interaction strategies in mind, we address the following research questions (RQ): RQ1 To what extent do stereotypical strategies of human-automation interaction enable drivers to experience meaningful human control? RQ2 How are objective metrics associated with driver’s subjective experience of meaningful human control? We investigate these questions in a controlled driving simulator study where participants interact with an automated vehicle under both HSC and TC conditions in safety-critical situations. The main contributions of this paper are: • A methodology that integrates behavioural metrics, subjective questionnaires, and qualitative insights to evaluate how meaningful the control of human drivers is when interacting with driving automation • An experiment-based comparison of two established human-automation interaction strategies in terms of the driver’s perception of control and responsibility and the relation of those perceptions to observable behaviour.
2. Methods We conducted a driving simulator experiment (Section 2.1) and used surveys to gauge drivers’ subjective perception of interacting with driving automation through different control modes (Section 2.2). These subjective measurements were complemented with objective behavioural metrics derived from the telemetry data (Section 2.3), which were further analysed to study the influence of control modes, and relations between subjective perceptions and behavioural metrics (Section 2.4). The data and analysis scripts are publicly available along with supplementary materials1 . Twenty-four participants (13 male, 11 female) 23 to 36 years old (with average and SD 29.2 ± 3.75 years) were recruited from the student and research community at Delft University of Technology between December 2023 and April 2024 via flyers distributed through personal contacts and snowball sampling, where enrolled participants referred others. Eligible candidates must (i) have held a valid driving licence for at least one year, (ii) have normal or corrected-to-normal vision without spectacles (contact lenses were permitted), and (iii) have no history of epilepsy or other conditions that could be aggravated by virtual reality. The study protocol was approved by the Human Research Ethics Committee (HREC) of Delft University of Technology (ID: 111053). Written informed consent was obtained from all participants prior to the experiment. 1 Link to supplementary materials: https://osf.io/.. (will be updated after acceptance).
3
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
“Undetected” rider (collision target)
Ego vehicle Reference Trajectory Other vehicles
Figure 2 | Driving simulator experiment: Experimental setup: the fixed-base driving simulator with a virtual reality headset and haptic steering wheel, and the repeated overtaking scenario simulating a silent automation failure on a two-lane, two-way road. The ego vehicle (white car) overtakes other vehicles and motorcyclists in the presence oncoming traffic (yellow). To simulate silent automation failure (i.e., the system failing to detect a motorcyclist), the reference trajectory was used that crossed the path of a motorcyclist, leading to a collision in case the participant did not intervene. Only one motorcyclist served as the collision target, while the other travelled without conflict. The figure illustrates one representative configuration; the position of the vehicles and motorcycles, including the target motorcycle, were counterbalanced across trials using a Latin square design across trials. At the beginning of the experiment, participants were reminded of their right to withdraw at any time without penalty. A €10 voucher was provided to each participant upon completion of the experiment. All data were anonymised and stored using unique participant identification codes. 2.1. Driving simulator experiment The participants were asked to drive through a section of a rural road supported by driving automation; this was done repeatedly over a sequence of trials. In each trial, participants had to complete a sequence of overtaking manoeuvres in a bidirectional road, requiring them to repeatedly swerve into the opposite lane while managing oncoming traffic, while interacting with an automated driving system through a haptic steering wheel. Participants were presented a first-person view of the driving scene using Varjo VR-3 virtual reality headset while seated in front of a SensoWheel SD-LC steering wheel capable of providing high-fidelity haptic feedback (Fig. 2). The experiment was configured in JOAN [27] – the framework for running experiments in the CARLA environment [28]. Sony WH-1000XM3 noise-cancelling headphones were used to reduce auditory distractions. As the participants drove through a bidirectional rural road with straight and winding sections, the speed of the ego vehicle was fixed at 50 km/h under cruise control. Participants had no control over the accelerator or brake pedals, controlling only the steering wheel. This was done to isolate the lateral control task and ensure that participants’ behaviour primarily reflected their interaction with the steering-based haptic guidance, rather than confounding influences from longitudinal control. While driving, participants encountered right-hand traffic travelling in both directions at a constant speed of 40 km/h: nine vehicles travelling in the same direction, and five in the opposite direction (Fig. 2). Due to the speed difference, participants had to overtake the vehicles travelling in the same direction by swerving into the opposite lane. The vehicles were spaced in such a way that, with appropriate steering, it was possible for the participants to overtake all vehicles while maintaining
4
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
sufficient safety margins to oncoming vehicles. The experiment included three conditions: fully manual control (baseline; participants had complete control over steering), haptic shared control (HSC), and traded control (TC). In the two latter conditions, the steering was controlled by a simple automated driving system which used a pure pursuit controller to follow a pre-recorded reference trajectory. In HSC, both the automation and the driver could apply torques on the steering wheel at the same time and thus steer the vehicle together. In TC, if the driver applied a torque above a threshold (0.2 N m), the automation’s authority decreased to zero, giving the driver complete authority, and was only restored after the driver’s torque dropped below the threshold for more than one second. If the torque applied by the human remained below the threshold for more than one second, automation torque gradually increased until it regained full authority or was intervened upon by the driver. Both HSC and TC were implemented based on the four design choices architecture [29]. To familiarise themselves with the control modes and driving task, each participant completed four familiarisation trials: (i) manual driving without traffic, (ii) manual driving with traffic, (iii) driving with traffic under HSC, and (iv) driving with traffic under TC. The main experiment then consisted of nine trials: Trial 1 involved manual driving, while Trials 2–9 featured automated driving in either TC or HSC, with four trials per mode. The order of automated trials was randomized independently for each participant. To investigate participants’ interaction with automation in safety-critical scenarios, we simulated a silent automation failure in every automated driving trial. Specifically, the ego vehicle was programmed to follow a trajectory leading toward a potential side collision with a motorcycle (Fig. 2). Only one motorcycle served as the collision target per trial; its path intersected with the automation’s reference trajectory, while the other motorcycle travelled without conflict. The target motorcycle’s was counterbalanced with those of other vehicles and motorcycles across trials using a Latin square design, with position distributions balanced across participants and control modes to avoid learning effects. Participants were instructed to remain in their lane (unless overtaking is needed), avoid collisions, and stay on the road. They were also told: “Please note that the automation system is still in development and may falsely detect objects, potentially leading to collisions. You need to intervene when it behaves unsafely ”. Participant’s subjective perceptions were measured after each trial using a post-trial survey (Section 2.2.1). After completing all trials, a descriptive post-experiment questionnaire was used to qualitatively assess the subjective perceptions about the interaction (Section 2.2.2). 2.2. Subjective perception questionnaires Post-trial surveys with a rating scale were used to quantify the perception of the participants (Section 2.2.1) and open ended post-experiment surveys were used to get a deeper understanding of the cues that influenced to their perception (Section 2.2.2). 2.2.1. Post-trial subjective scores Seven Likert-scale (1 to 10) questions (Table 1) were asked after each HSC or TC trial to gauge how participants experienced the interaction with the driving automation (referred to as the automated vehicle (AV) in the questionnaires) and to ultimately assess the extent to which the driving automation operated under MHC. The questions were derived from the previously proposed actionable properties of systems under MHC [17]. Based on tracking and tracing conditions [16], these properties require that: P1 a moral operational design domain (moral ODD; extending the traditional notion of operational design domain towards relational and moral aspects) is clearly defined and mechanisms are in place for keeping the system in that domain; P2 representations which humans and automation have of each other, the environment, and the moral ODD, are shared, i.e., mutually compatible; P3 humans have sufficient ability and authority to control the outcome of the automation; P4 automation actions can be explicitly linked to at least one human who is aware of their responsibility. Since here we focused on the interaction between the human and the automation, rather than on the design of the automation itself, we designed the control modes HSC and TC to have identical moral operational design domains. Hence, in our comparison of HSC and TC we did not include any questions for Property 1. Property 5
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
Post-trial questions – automated driving trials (HSC and TC) A1: I felt that I had sufficient control over the automated vehicle A2: I felt that the automated vehicle understood my intentions during the driving task A3: I felt that I had sufficient understanding about the behaviours of automated vehicle A4: I felt that the automated vehicle and I were working together towards the same goal A5: I felt responsible for the driving task when I was using the automated vehicle A6: I felt safe in the automated vehicle during the driving task A7: I trusted the automated vehicle during the driving task
Property Sufficient control (P3) AV understood me (P2) I understood AV (P2) Working together (P2) Responsible (P4) Safe Trust
Post-trial questions – manual driving trials M1. I felt that I had sufficient control over the vehicle. M2. I felt responsible for the driving task when I was using the vehicle. M3. I felt safe in the vehicle during the driving task.
Sufficient control Responsible Safe
Table 1 | Post-trial questions that participants had to answer on a Likert scale (1 to 10) after interacting with the automated driving system through haptic shared control or traded control, and their relation to the properties of systems under meaningful human control [17]. Questions A6 and A7 were designed to assess perceived safety and trust, concepts that are related to MHC but do not map directly onto the four MHC properties. The questions for manual driving were redacted since there was no automation involved. 2 (shared representations) was divided into three components: AV understood me, I understood AV, and working together. This division reflects Cavalcante Siebert et al.’s [17] description that shared representations between human and AI systems depend on how both agents understand each other and are able to update their representations in response to changing reasons. The remaining properties were each represented by one questionnaire item: sufficient control for Property 3 (ability and authority to control) and responsible for Property 4 (actions are linked to a responsible human). In addition to items directly related to the MHC properties [17], we included questions on perceived safety and trust; as previous research has demonstrated, they can strongly influence evaluations of MHC based on how drivers experience interactions with driving automation [30]. As the manual driving trials did not include driving automation, only three of the above questions were applicable to these trials; different formulations were used to account for the absence of automation. Thus, answers to questions M1, M2 and M3 were used to evaluate participants’ baseline perception of sufficient control, responsibility and safety, respectively (Table 1). 2.2.2. Post-experiment open-ended questions After all the driving trials, an open-ended post-experiment questionnaire was administered to get a deeper understanding of the perceptions of the participant and the factors that might have influenced these perceptions. The questions related to the same concepts mentioned in post-trial questions (sufficient control, AV understood me, I understood AV, working together responsibility, safety, and trust). Two questions per concept asked the participant to recall positive and negative experience related to that concept. For example, the questions related to sufficient control were: D1: “Were there any situations where you felt you had sufficient control over the automated vehicle operation? Please describe.” and, D2: “Were there any situations where you felt you did not have sufficient control over the automated vehicle operation? Please describe.” An exhaustive list of the questions is included in the supplementary material. These items were designed to record in greater detail the overall impressions the participants had of HSC and TC, to elicit scenarios that positively or negatively affected their perception of the qualities being measured in the post-trial questionnaire, and to compare their experiences across conditions. In particular, they addressed aspects that could not be evaluated meaningfully on a trial-by-trial basis, such as general system preference, system characteristics that affect workload, and qualitative reflections on system behaviour. This approach
6
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
follows [22] in employing post-experiment reflections to evaluate MHC. In our study, such qualitative insights provide valuable context for understanding the Likert-scale scores recorded after each trial. 2.3. Behavioural metrics and hypotheses To study how subjective perceptions are related to behavioural aspects of driver-automation interaction, we formulated trial-level behavioural metrics that quantify interaction dynamics between the driver and the automation, driver’s performance, and the characteristics of vehicle trajectories. We then formulated hypotheses regarding a) correlation of these metrics to certain subjective perception scores, and b) differences in these metrics between HSC and TC (Fig. 3). Brief descriptions of metrics and the hypotheses linked to them are mentioned below; all metrics except reaction time and time to collision were calculated by aggregating the data over the whole duration of each trial. Detailed definitions of each metric are provided in supplementary materials. Reaction time: In experiment trials where a simulated silent automation failure triggered an incursion manoeuvre towards a motorcyclist, the reaction time was quantified as the time between the initiation of the automation manoeuvre and the first detectable human response. The automation initiation was identified as the first moment when the gradient of the automation torque exceeded a threshold (0.18 Nm/s), while the human response was determined based on conflict onset in HSC and authority in TC. We hypothesised that reaction times would be negatively correlated with subjective scores for item A1 (‘sufficient control’). Because of the continuous nature of the interaction in HSC, we hypothesised that reaction times would be shorter for HSC than for TC. Average conflict in steering torques: In the experiment, the human driver interacted with the automation through the steering wheel. We used the conflict in steering torques — defined as the negative signed product of human and automation applied torques that exceed a threshold (0.2 Nm) — to measure the disagreement between the human and the automation. In line with prior research [31], we hypothesised that conflict is negatively correlated with subjective scores A2, A3, and A4 (‘AV understood me’, ‘I understood AV’, and ‘working together’, respectively). Based on the assumption that interactions with HSC will be smoother, we hypothesised that conflict will be lower for HSC than for TC. Maximum steering torque: When interacting with the automation, the steering torque exerted by the participant reflects their control effort. Thus, we used maximum value of the human-applied torque during the trial as a metric of effort. Earlier studies have shown that drivers exert maximum torque when resisting lane-keeping or lane-departure assistance [32]. Thus, we hypothesised that maximum steering torques are negatively correlated with scores of ‘sufficient control’ (A1) and ‘working together’ (A4). Due to the smoothness of the interaction in HSC, we hypothesised that maximum steering torques will be lower for HSC than in TC. Steering reversal rate: Another established method for quantifying the control effort of a driver considers steering reversals: the greater the rate of steering reversals, the greater the control effort of the driver [33]. Steering reversal rate was computed as the number of steering reversals per second across a full trial, where a reversal is counted when the rate of change of steering angle is zero and the difference between two extremes exceeds 1° [34]. We hypothesised that the steering reversal rate is negatively correlated with the scores A1, A3, and A4 (‘sufficient control’, ‘I understood AV’ and ‘working together’, respectively). Based on the continuous nature of HSC, we hypothesised that the steering reversal rate would be lower for HSC than for TC. Jerk of vehicle trajectories: The smoothness of the trajectories of the ego vehicle was quantified as the rootmean-squared jerk (longitudinal and lateral jerks combined) computed over the whole trial. We hypothesised that trajectory jerk would be negatively correlated with the score A1 (‘sufficient control’). Also, assuming that HSC would result in smoother trajectories, we hypothesised that jerk would be smaller for HSC than TC. Time to collision (TTC): In each trial where a simulated silent automation failure could lead to a collision with a motorcyclist, the minimum TTC between the ego vehicle and the motorcyclist can be used to measure the criticality of the interaction: the lower the TTC, the more critical the interaction is. The TTC was calculated using a two-dimensional TTC algorithm [35], with the minimum taken over the duration of the failure event. We hypothesised that low min-TTC values (higher interaction criticality) would be associated with reduced perception of ‘sufficient control’; that is, we expected a positive correlation between TTC and A1. Furthermore, as the continuous engagement of the driver can reduce the criticality of the interaction in HSC, we hypothesised that TTC would be larger for HSC than for TC . 7
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
Property 1
Property 2
Moral operational design domain
MHC Properties
Shared representations
Property 3
Property 4
Ability and authority to control
Actions are linked to a responsible human
[17] Questions based on properties
Subjective perception
Sufficient control
AV understood me
I understood AV
Working together
Responsible
Hypothesised relations between subjective perception and objective metrics
+
-
Behavioural Metrics
Reaction time
Conflict
Maximum steering torque
Steering reversal rate
Jerk
Time to Collision
Take overs
Trajectory Deviation
HSC < TC
HSC < TC
HSC < TC
HSC < TC
HSC < TC
HSC > TC
HSC > TC
HSC > TC
Hypothesised relations between objective metrics and control modes
Figure 3 | Deriving hypotheses about survey questions and behavioural metrics from the properties of systems under meaningful human control (MHC): Survey questions related to MHC properties were used to evaluate subjective perception of MHC. The figure also shows the hypothesised dependence of subjective perception of MHC on behavioural metrics and control modes. Number of takeovers: When a driver is interacting with the automation through TC, a takeover happens if the torque applied by the driver exceeds the pre-defined threshold which would disengage the automation. For HSC, a takeover was counted each time the conflict signal transitioned from zero to positive. The total number of such events was summed per trial. We assumed that takeovers would be triggered when the behaviour of the automation was not aligned with the intentions of the driver. Hence, we hypothesised that the number of takeovers would be negatively correlated with scores A2 (‘AV understood me’) and A4 (‘working together’). At the same time, we hypothesised that the number of takeovers would be positively correlated with the score A5 (‘responsible’), as the driver would have more influence on the trajectory of the ego vehicle with more takeovers. We hypothesised that participants would be more engaged in HSC and that the number of takeovers would be greater in the case of HSC than for TC. Trajectory deviation: The deviation of the trajectory of the ego vehicle from the reference trajectory can be used to capture the extent to which the driver contributed to the behaviour of the vehicle. As an example, if the trajectory doesn’t deviate at all from the reference, that would mean that the driver fully relied on automation. This was quantified as the root-mean-squared error between the actual and reference trajectories, computed over the whole duration of a trial. The reference point for each time step was taken as the nearest point on the automation’s reference trajectory. We hypothesised that root-mean-squared trajectory deviation is positively correlated with the score A5 (‘responsible’). Also, assuming that drivers are more engaged in HSC, we hypothesised that trajectory deviation is higher for HSC than TC. Overtaking time: To quantify how driver’s preferences might deviate from the reference trajectory, we also included the overtaking time defined as the total time spent by the ego vehicle in the opposite lane during a trial. Since the preferences of individual drivers might lead to longer or shorter overtaking times, we did not associate any hypotheses with the this metric; it was only used for exploratory analyses. By providing objective insight into the behaviour of drivers, these metrics complement the data collected about the subjective perception of various concepts related to MHC. Detailed definitions of these metrics can be found in supplementary materials.
8
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
2.4. Analysis Behavioural data and subjective scores collected during the experiment were analysed quantitatively to test hypotheses (Section 2.4.1) and the answers to post-experiment questionnaires were analysed qualitatively to garner deeper insights into the factors affecting them (Section 2.4.2). 2.4.1. Quantitative analysis To analyse the effect of the control modes (HSC and TC) on the behaviour of participants, we fit linear mixed-effect models (LMMs) for each behavioural metric (Section 2.3): behavioural-metric ∼ control-mode + (1|participant) . Here, behavioural metrics (z-scored) were the dependent variables, control mode was the independent variable with a fixed effect, and participant ID was a random intercept. To examine how participant’s perceptions during the experiment were related to their behaviour and control modes, we analysed the post-trial subjective scores (Section 2.2.1) as a function of control mode and behavioural metrics. For each subjective perception item, one LMM was fit with the subjective score as the dependent variable, control modes and all behavioural metrics as fixed-effect independent variables, and participant ID as a random intercept: subjective-score ∼ all-behavioural-metrics + control-mode + (1|participant) . Control mode is included as a fixed effect to account for the effect of interaction type on subjective scores. The results of these LMMs were used in the confirmatory analyses of our hypotheses about the relationships between behavioural metrics and subjective scores (Fig. 3). For the confirmatory analyses, we used the statistical significance level 𝛼 = 0.05. In addition to testing our hypotheses, we conducted an exploratory analysis (based on the above LMMs) to identify potential relations between subjective scores and behavioural metrics that might have been overlooked in the hypotheses. The purpose of the exploratory analysis was not to make further claims, but to highlight possible relationships to investigate in future experiments. In our initial hypotheses, we had assumed that the relationships between subjective scores and behavioural metrics would be the same across control modes. However, some responses to the post-experiment questionnaire hinted that the nature of the interaction, as determined by the control mode, might also influence the relationships between subjective scores and behavioural metrics. To study whether control modes influence these fixed effects, we split the data and fit separate LMMs for each control mode. For the exploratory analysis, statistical significance of slopes was assessed at 𝑝 < 0.05. No correction for multiple comparisons was applied as we prioritised minimizing Type-II errors (false negatives) over Type-I errors (false positives). We did not formulate hypotheses about the effect of control mode on subjective scores; these relationships were therefore explored as part of the exploratory analysis only. 2.4.2. Qualitative analysis To analyse participants’ responses to the post-experiment questionnaire, we categorised their free-text answers into thematic topics using a two-coder approach. Both coders reviewed participants’ responses in their entirety to ensure familiarity with the dataset. An initial coding pass was conducted by one coder to identify all possible topics mentioned by participants, without restricting the scope to pre-defined categories. This exploratory, data-driven approach allowed for the inclusion of topics beyond those anticipated by the MHC framework. Following the initial coding, similar topics were consolidated into broader thematic categories. Within each category, related ideas were organised into subtopics that together captured the range of participant experiences. This hierarchical structure enabled the preservation of nuanced perspectives while reducing redundancy and complexity in the dataset. Once the preliminary categorisation was complete, a second coder independently reviewed the grouped topics. Both coders then developed labels and classifications for each factor, determining the most appropriate and precise naming. This independent labelling stage ensured that classification decisions were made without bias from prior discussion. 9
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
Reaction time b = 1.33, p < 0.001
ns
3 Z-scored metrics
Maximum Steering Torque b = 0.57, p < 0.001
Conflict b = -0.10, p = 0.42
Steering reversal rate b = -0.00, p = 0.99
Take overs b = -1.37, p < 0.001
TTC b = -0.35, p = 0.032
ns
***
***
Jerk b = -0.10, p = 0.36
Trajectory deviation b = -0.10, p = 0.48
Overtaking Time b = 0.46, p < 0.001
ns
***
*** ns
2
*
1 0 1 2 3
HSC
TC
HSC
TC
HSC
TC
HSC
TC
HSC
TC
HSC
TC
HSC
TC
HSC
TC
HSC
TC
(a) Control Modes Control Modes Control Modes Control Modes Control Modes Control Modes Control Modes Control Modes Control Modes Sufficient control b = -1.03, p < 0.001
Subjective Score
10
ns
I understood AV b = -0.85, p < 0.001
***
Working together b = -0.79, p < 0.01
**
Responsible b = -0.28, p = 0.21
Safe b = -1.17, p < 0.001
Manual HSC TC
Manual HSC TC
ns
***
Trust b = -0.68, p = 0.015
*
8 6 4 2
(b)
***
AV understood me b = -0.38, p = 0.15
Manual HSC TC
Control Modes
HSC
TC
Control Modes
HSC
TC
Control Modes
HSC
TC
Control Modes
Control Modes
Control Modes
HSC
TC
Control Modes
Figure 4 | The effects of control mode on behavioral metrics (a) and subjective perception scores (b). Statistical comparisons in panel (a) are based on confirmatory analyses using linear mixed-effect models (Section 2.4.1; Table 2). Comparisons in panel (b) are exploratory and indicated for reference only. Statistical tests are based on LMMs used for testing the relationships between behavioural metrics and subjective scores (Section 2.4.1) which also included control mode as the fixed effect; significance codes: ∗∗∗ 0.001 ∗∗ 0.01 ∗ 0.05. For visualization purposes, outliers (z-scores > 3) were excluded in panel (a). To assess coding reliability, both coders independently evaluated the presence or absence of each identified factor within participant responses. Each factor–response pair constituted one coded item for the reliability analysis. Inter-coder agreement was quantified using Cohen’s kappa [36] and Gwet’s AC1 [37]. See supplementary materials on osf.io for more details about the calculation of these metrics. Discrepancies between coders were resolved through discussion until full consensus was achieved. This consensus-based, investigator-triangulation approach follows established best practices for directed content analysis [38, 39, 40].
3. Results 3.1. Quantitative findings The quantitative results are reported in Tables 2 and 3 and Figs. 4 and 5. The confirmatory analysis tested two sets of hypotheses: the effect of control mode on each behavioural metric (Fig. 4 and Table 2), and the relationships between behavioural metrics and subjective scores (Table 3). The exploratory analysis examined additional relationships between behavioural metrics and subjective scores that were not part of the original hypotheses (Fig. 5). For each behavioural metric below, we first report the confirmatory relationship with control mode, then with subjective score, and finally any additional findings.
10
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
Behavioural Metric Reaction time Conflict Maximum steering torque Steering reversal rate Jerk TTC Takeovers Trajectory deviation
Hypothesised relation
Observed relation
HSC < TC HSC < TC HSC < TC HSC < TC HSC < TC HSC > TC HSC > TC HSC > TC
< <
>
Slope 𝛽
p-value
1.33 -0.10 0.57 -0.00 -0.10 -0.35 -1.37 -0.10
< 0.001∗∗∗ 0.42 < 0.001∗∗∗ 0.99 0.36 0.03 ∗ < 0.001∗∗∗ 0.48
Hypothesis accepted ✓ ✓
✓ ✓
Table 2 | Confirmatory analysis – behavioural metrics vs control modes: Hypothesised relations and results of statistical tests based on LMMs: behavioural-metric ∼ control-mode + (1|participant) (Section 2.4.1); significance codes: ∗∗∗ 0.001 ∗∗ 0.01 ∗ 0.05. Reaction time: In accordance with our hypothesis, reaction times in simulated automation failures were significantly lower for HSC than for TC ( 𝛽 = 1.33, 𝑝 < 0.001). However, contrary to our hypothesis, reaction times were positively correlated with the score for ‘sufficient control’ ( 𝛽 = 0.24, 𝑝 = 0.01). The exploratory analysis also revealed a positive correlation between reaction time and the score of ‘responsible’ for HSC ( 𝛽 = 0.18, 𝑝 = 0.04). Conflict in steering torques: There was no evidence for the hypothesized difference in conflict in steering torques between HSC and TC ( 𝛽 = −0.1, 𝑝 = 0.42). At the same time, the hypothesis that conflict in steering is negatively correlated with the score for ‘AV understood me’ was confirmed ( 𝛽 = −0.34, 𝑝 = 0.001). The exploratory analysis suggested that conflict might be positively correlated with the score of ‘sufficient control’ for HSC ( 𝛽 = 0.39, 𝑝 = 0.01) and negatively correlated with the score of ‘responsible’ for TC ( 𝛽 = −0.27, 𝑝 = 0.03). Maximum steering torque: Consistent with our hypothesis, maximum steering torque was lower for HSC than for TC ( 𝛽 = 0.57, 𝑝 < 0.001). The exploratory analysis indicated a possible positive correlation between maximum steering torque and the score of ‘responsible’ for TC ( 𝛽 = 0.235, 𝑝 = 0.04). Steering reversal rate: There was no evidence for relationships between steering reversal rate and control mode or subjective scores. Jerk of vehicle trajectories: The data showed no evidence for differences in the jerk of vehicle trajectories between HSC and TC. None of the hypothesised correlations of jerk with subjective scores were substantiated in the confirmatory analysis. The exploratory analysis showed that trajectory jerk could be negatively correlated with the score of ‘safe’ ( 𝛽 = −0.22, 𝑝 = 0.01), especially for TC ( 𝛽 = −0.314, 𝑝 = 0.01). Time to collision: The data did not show a significant difference in minimum TTC during silent automation failure between the two control modes. There was no evidence for the hypothesised correlation between TTC and the score of ‘sufficient control’ either. However, the exploratory analysis suggested a positive correlation between minimum TTC and the score for ‘AV understood me’ for TC ( 𝛽 = −0.253, 𝑝 = 0.03). Number of takeovers: As per our hypothesis, the number of takeovers was larger for HSC than for TC ( 𝛽 = −1.37, 𝑝 < 0.001). No evidence was however found for correlations between the number of takeovers and any of the subjective scores. Trajectory deviation: No evidence was found for difference in deviation of the vehicle trajectory from prerecorded reference between HSC and TC. The confirmatory analysis did not yield evidence for any of the hypothesised correlations of trajectory deviation with subjective scores. The exploratory analysis indicated that trajectory deviation can be negatively correlated with the subjective score for ‘sufficient control’ ( 𝛽 = −0.19, 𝑝 = 0.02) and positively correlated with the score for ‘trust’ ( 𝛽 = 0.24, 𝑝 = 0.01), especially for TC ( 𝛽 = 0.29, 𝑝 = 0.05). Overtaking time: Overtaking time was not part of our hypotheses; exploratory analyses suggested that overtaking time was larger for TC than for HSC ( 𝛽 = 0.46, 𝑝 < 0.001). Overtaking time might be positively
11
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
Hypothesised Slope 𝛽 relation
p-value
Subjective score
Behavioural metric
Sufficient control
Reaction time Maximum steering torque Steering reversal rate Jerk TTC
+
0.238 0.038 -0.178 -0.148 0.033
0.014 ∗ 0.666 0.086 0.083 0.621
AV understood me
Conflict Takeovers
-
-0.344 -0.019
0.001∗∗∗ 0.880
I understood AV
Conflict Steering reversal rate
-
-0.130 -0.038
0.203 0.720
Working together
Conflict Maximum steering torque Steering reversal rate Takeovers
-
-0.110 -0.080 0.061 -0.176
0.307 0.402 0.565 0.167
Responsible
Takeovers Trajectory deviation
+ +
0.182 -0.022
0.900 0.769
Hypothesis accepted
✓
Table 3 | Confirmatory analysis — subjective scores vs behavioural metrics: Hypothesised relations and results of statistical tests based on LMMs subjective-score ∼ behavioural-metric+control_mode+ (1|participant) (Section 2.4.1); significance codes: ∗∗∗ 0.001 ∗∗ 0.01 ∗ 0.05. ‘+’ indicates that the perception score was hypothesized to be positively correlated with the corresponding behavioural metric; ‘-’ indicates hypothesised negative correlation. correlated with subjective score for ’safety’ ( 𝛽 = 0.22, 𝑝 = 0.04) and ’trust’ ( 𝛽 = 0.24, 𝑝 = 0.05). Additionally, for HSC, we found potential positive correlation between overtaking time and the score for ‘sufficient control’ ( 𝛽 = 0.35, 𝑝 = 0.01). 3.2. Qualitative findings We performed a topic analysis of participants’ responses to the post-experiment questionnaire to gain deeper insight into the factors that influenced their subjective experience by categorising those experiences into thematic topics. As part of the qualitative analysis, inter-coder agreement was assessed on a set of 57 coded factors. The observed agreement between coders was 66.67%, with Cohen’s kappa [36] of –0.1993 and Gwet’s AC1 [37] of 0.5385. The negative kappa value was attributable to the absence of cases where both coders labeled a factor as absent (see contingency table in supplementary materials for details), which can produce prevalence bias in kappa calculations [41]. The agreement between both coders was on the presence of a factor in 38 cases, with disagreement in 19 cases (nine coded as present by Coder 1 only, ten coded as present by Coder 2 only), and no instances where both coded the factor as absent. Given these conditions, Gwet’s AC1 was taken as the more robust measure of agreement. Full consensus was reached after discussions between coders consolidated overlapping factors and resolved disagreements by jointly evaluating the arguments presented by each coder. This process resulted in 47 final factors that either positively or negatively affected perceptions related to MHC. These factors along with remarks for control modes are summarised in Table 4. We classified the factors into three categories: (1) attitudinal factors ○ relating to mental states of participants, (2) interaction factors ■ relating to human-AV interaction through the steering wheel, and (3) trajectory factors ▶ relating to the motion of vehicles. We follow this categorisation when describing the findings for each type of perception. Participant quotes are referenced by anonymised IDs (e.g., ID61, ID90); these are randomised identifiers and do not reflect the sequential numbering of the 24 participants.
12
Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving