1
The LAIA Dataset: Labelled Attention for Intelligent Automobiles
arXiv:2607.25570v2 [cs.CV] 29 Jul 2026
A. Contreras, D. Porres, R. Abad, P. Cano, A. Levy, G. Villalonga, A. M. López and A. Hernández-Sabaté
Abstract—The development of autonomous vehicles (AVs) usually relies heavily on data-driven Artificial Intelligence (AI) models that require large volumes of sensor data with groundtruth annotations. While modular architectures are widely used, end-to-end driving paradigms offer a promising alternative by directly mapping sensor inputs to control actions. However, their adoption is limited due to challenges in interpretability and explainability. To address this, we present LAIA (Labelled Attention for Intelligent Automobiles), a novel synthetic dataset designed to enrich end-to-end driving research with human attention data. Collected using the CARLA simulator in closedloop environments, LAIA comprises over 15 hours of driving from 44 participants across carefully crafted scenarios designed to evoke natural responses. Each sequence includes RGB images in six weather conditions, semantic and instance segmentation, depth, optical flow, CAN bus signals, and synchronized eyetracking data. LAIA opens the doors to the use of human gaze driving for different applications such as training attention-aware end-to-end AI drivers, predicting driver behavior, or developing methods to detect anomalous driver-attention patterns, as well as to improve the explainability of the models. Within the scope of this work, we use LAIA to compare human attention with the perceptual attention emerging in our end-to-end driving models, thereby providing interpretability regarding their behavior. Index Terms—Autonomous vehicles, driver attention dataset, end-to-end driving, explainability
I. I NTRODUCTION Autonomous vehicles could save hundreds of thousands of lives over the coming decades while reducing congestion, pollution, and travel time [1], making autonomous driving one of the most active research areas in artificial intelligence, computer vision, and robotics. Among the paradigms proposed for vehicle automation, end-to-end driving has attracted growing attention for its direct mapping from sensory input to control actions, offering a scalable alternative to modular pipelines in which perception, prediction, planning, and control are treated as separate components. However, despite competitive driving performance, end-to-end models remain difficult to interpret [2], limiting their adoption relative to modular approaches whose explicit intermediate representations naturally support explainability. In safety-critical settings, predictive accuracy alone is insufficient: it is equally important to verify which scene elements a model attends to and whether they are the same elements that a human driver would consider relevant, such as pedestrians, traffic lights, or nearby vehicles. This challenge is rooted in how end-to-end models are trained — through behavior cloning from sensor-action pairs collected All the authors are with the Computer Vision Center. A.M. López and A. Hernández-Sabaté are also with the Computer Science Department of Universitat Autònoma de Barcelona.
in human-driven vehicles — which captures what the driver did but not what the driver looked at, creating a gap between observable behavior and the underlying cognitive process. Human attention offers a promising way to narrow this gap: gaze data can serve both as an auxiliary supervisory signal during training, improving robustness and generalization, and as a post-hoc interpretability tool for evaluating whether model attention aligns with human visual decision-making. We introduce LAIA (Labelled Attention for Intelligent Automobiles), a public collection that couples eye gaze, control actions and full simulator annotations for large-scale autonomous-driving research, enabling reproducible largescale studies of vision-guided driving. LAIA is designed to support the training, testing, and validation of end-to-end driving models while also enabling attention-based explainability studies. The dataset combines synchronized driving data with human attention information, allowing direct comparison between model attention and human gaze. By providing both the signals required for end-to-end learning and the information needed to analyze visual decision-making, the dataset aims to facilitate research toward autonomous driving systems that are not only accurate, but also more interpretable and cognitively grounded. In particular, LAIA contains comprehensive driving scene data that reflect the inputs used by end-to-end AI driving models, including on-board sensor streams and vehicle control signals, all synchronized with human driver attention data captured via eye-tracking technology. The data were collected in closed-loop driving simulation environments, where 44 participants drove through a series of carefully designed routes featuring sudden and unexpected driving events. The simulation has been conducted using the CARLA autonomous driving simulator [3], complemented by a monitoring system that tracks human behavior while navigating within CARLA’s virtual towns. A total of 8 different CARLA routes with 18 different events carefully distributed to cause human reactions during their driving were implemented. RGB images in six environmental variations, along with ground-truth annotations for semantic segmentation, scene depth, panoptic instance segmentation, and synchronized human eye-tracking data are publicly released1 [4]. Unlike existing driving datasets, LAIA uniquely integrates dynamic human attention signals into the driving loop, providing a foundation for more cognitively grounded development of AI drivers. The dataset allows researchers to assess not only if an AI driver makes the correct driving decisions, 1 https://cloningdcb.org/
2
Fig. 1. Platform Setup. Left: Dynamic driving platform. Right: Static driving platform. In both cases the simulation software is CARLA and the simulations used in human driving were based on the scenarios in Figure 2.
but also how similar its perception is with regard to human attention, an important step toward building interpretable and reliable autonomous systems. Additionally, the attention maps extracted from human gaze data can serve as weak supervision or as an auxiliary task in multi-objective learning, supporting training strategies that improve both accuracy and robustness. Beyond learning, LAIA facilitates systematic benchmarking of attention-aware models and allows direct comparisons between human and model behavior under controlled, repeatable conditions. The remainder of the paper is as follows. Section II describes all the elements necessary for creating the dataset, including the data recording protocol. Section III is devoted to explaining the synchronization and mapping of the gaze data with the images generated by the simulator. Section IV describes the format and structure of the data while, Section V shows different examples of the use of the dataset. Finally, conclusions are summarized in Section VI. II. DATA ACQUISITION D ESCRIPTION Data were collected in two different centres, the Institute of Biomechanics of Valencia2 (IBV) and the Computer Vision Center3 (CVC), but the only differences among the subdatasets lie in the physical simulator environments. During the human driving sessions, eye-tracker data, CARLA images and additional data were captured to replay the episodes offline with the aim of generating additional data without preventing real-time performance during human driving sessions. Adapted autopilots accessing CARLA insider information were also used to automatically generate data associated with driving episodes. In all cases, erratic data were removed. In this Section, the main characteristics of the simulation environments, experimental scenarios (design, implementation, and experiment structure), and participants used to collect the database are described. A. Physical Simulation Environments The setup of the simulator environments consists of a system emulating a car, containing Hardware and Software Interfaces. 2 www.ibv.org 3 www.cvc.uab.cat
The simulator at IBV is mounted on a dynamic platform (Motion Systems PS-6TM-550) with a regulable seat that simulates the kinematic experience of a real vehicle (acceleration, braking, cornering, crashes and terrain irregularities) in response to manoeuvres executed by the human driver (Figure 1, left). The driving scene is displayed across three 55” 4K screens arranged at 46-degree angles between them. The hardware simulated environment mounted at CVC is a static platform composed of a foldable aluminium frame, a Z-1 sports seat, the Thrustmaster T300 RS haptic simulation hardware including a GP-4 carbon offset steering wheel together with compatible turn signals and pedals (Figure 1, right). In this case, three articulated monitors of 28” and a resolution of 1920x1080, arranged at 46-degree angles between them, represent the front view of the car plus the two mirrors. In both simulators, a tablet displays the speed. CARLA simulator, version 0.9.14 was used to design and run the scenarios containing sudden events. Experiments were run on a PC (ProArt X670E-creator wifi, Ryzen-9 7900X, 2 x 32GB PC5-5600U) with Windows 10 and NVIDIA GeForce 4090 Liquid plus NVIDIA GeForce 4090 INNO3D.
B. Simulated scenarios We define an episode as the human driving experience from the time the subject sits on the platform until the end of driving in a predefined scenario. We define a scenario, named by TownXX RYY, as the combination of a town and a route comprising several events properly triggered as each subject drives along the pre-defined route. The instructions to follow the route were given by an automated voice (options: English/Spanish/Catalan). To design the scenarios, four different CARLA maps, namely towns, were selected and, for each map, between one and three routes were designed. Different events that may occur both in a city and on a highway were taken into account in a total of four towns with Single Lane (SL), Multilane (ML), or Highway (HW). In particular, the events implemented are the following and can be grouped in: •
Traffic Negotiation: Unprotected left turn at intersection, Turn at intersection with crossing traffic, Crossing traffic running a red light at an intersection and Crossing with oncoming bicycles.
3
E2. Unprotected left turn at intersection
E3. Turn at intersection with crossing traffic
E5. Crossing traffic running a red light at an intersection
E6. Crossing with oncoming bicycles
E7. Highway merge on ramp
E8. Highway cut-in from on-ramp
E9. Static cut-in
E10. Highway exit
E12. Obstacle in lane
E14. Slow moving hazard at lane edge
E14a. Slow moving hazard at lane edge
E16. Longitudinal control after leading vehicle’s brake
E17. Obstacle avoidance without prior action
E18. Pedestrian emerging from behind parked vehicle
E19. Obstacle avoidance with prior action - pedestrian or bicycle
E19a. Obstacle avoidance with prior action - vehicle
E20. Parking cut-in
E21. Parking exit
Fig. 2. Implemented events.
Highway: Highway merge on ramp, Highway cut-in from on-ramp, Static cut-in, Highway exit. • Obstacle Avoidance: Obstacle in lane, Slow moving hazard at lane edge, Slow moving hazard at lane edge.
•
•
Braking and Lane Changing: Longitudinal control after leading vehicle’s brake, Obstacle avoidance without prior action, Pedestrian emerging from behind parked vehicle, Obstacle avoidance with prior action - pedestrian or
4
TABLE I E VENT DISTRIBUTION ACROSS SCENARIOS ( SEE THE EVENT ID IN FIG . 2)
Town01 R01 Town04 R01 Town04 R02 Town04 R03 Town06 R01 Town06 R02 Town06 R03 Town10 R01
Traffic Negotiation E2 E3 E5 E6 x x x x x x x
E7
Highway E8 E9
E10
Obstacle Avoidance E12 E14 E14a x
x
x
x
x
Braking and Lane Changing E16 E17 E18 E19 E19a x x x x x x x
x
Parking E20 E21
x x
x x x
bicycle, Obstacle avoidance with prior action - vehicle. Parking: Parking cut-in, Parking exit. Figure 2 illustrates all the events implemented while Table I summarizes the distribution of the different events across the different scenarios, taking into account that SL is present in Towns 1 and 4, ML is present in towns 4 and 6 and HW can be found in towns 4, 6 and 10. •
x x x
x x x
x
Normal or corrected to normal vision (without glasses). Very little tendency to suffer motion sickness. • Non-professional drivers. • Driving at least once a week. Professional drivers were explicitly excluded to avoid overrepresentation of highly automated or task-optimized driving behaviors that may not reflect typical human attention patterns. The average age of participants was 36.59±9.31 and the average of years of license was 16.23±8.76. The characteristics of all the participants regarding age, years of license and number of scenarios run are detailed in table II: • •
TABLE II C HARACTERISTICS OF PARTICIPANTS
IBV CVC
Men Women Men Women
# Drivers 11 13 10 10
Age 40 ± 5.99 39 ± 9.02 31 ± 12.2 32.4 ± 9.5
License 20 ± 6.82 19 ± 9.21 12.4 ± 10.7 11.9 ± 8.6
Scenarios 3 4
Fig. 3. Eyetracker device.
For a better illustration Figures 4 and 5 show the different CARLA maps with the routes the human drivers have to follow, the simulated vehicles trajectories and the points where the events are implemented.
The details of recordings per scenario are detailed in table III. Notice that the number of frames for each modality is the same because they are synchronized. TABLE III R ECORDING DETAILS
C. Eyetracker To capture driver attention a wearable eye-tracker, as the one shown in Figure 3, was used (Tobii glasses 3 [5]). These glasses are equipped with four eye cameras and a wideangle scene camera, allowing for precise tracking of eye movements and visual attention. The glasses are lightweight and unobtrusive, ensuring natural behavior from participants during studies. Gaze data was recorded at 100Hz and the scene camera recorded at 25 fps with a resolution of 1920 x 1080 pixels. D. Participants A total of 44 participants (23 women and 21 men) were asked to drive during 20-30 minutes across 3-4 scenarios. To ensure a baseline level of driving competence and familiarity with common traffic situations, participants were healthy people without any condition that might have caused an imbalance in the data recorded: • At least two years of driving license.
Town01 R01 Town04 R01 Town04 R02 Town04 R03 Town06 R01 Town06 R02 Town06 R03 Town10 R01
# drivers 38 10 14 13 12 11 9 27
# frames 667211 58397 132884 88725 37119 89349 66601 209468
Total time (minutes) 445.19 38.96 88.65 59.19 24.76 59.61 44.43 139.75
E. Data Recording Protocol The main steps of the recording protocol are the following: 1) The participant receives the information sheet about the project under which LAIA is developed and the informed consent to be signed. 2) As soon as the participant agrees, (s)he sits in the driving simulator and puts on the eye-tracker glasses. 3) The participant is asked to drive in the simulated environment, relying on the driving simulator and following the audio instructions for approximately half an hour.
5
Fig. 4. Scenarios for CARLA’s towns 01, 04 and 10. Blue lines show the route the human drivers have to follow, from the starting point in dark green to the ending red one. Red lines show the automatic vehicles trajectories, from the starting bright green points to the smallest red ones. Events are marked with a blue box.
Fig. 5. Scenarios for town maps 06. Blue lines show the route the human drivers have to follow, from the starting point in dark green to the ending red one. Red lines show the automatic vehicles trajectories, from the starting bright green points to the smallest red ones. Events are marked with a blue box.
These instructions are used to guide the driver along the specified route, in the same way as commercial navigation assistants do. 4) During human driving, the following data is collected: a) Driver’s gaze captured by the eye-tracking glasses.
b) The images projected by CARLA to provide the driving experience. c) Log of the driving simulation to allow offline playback of the driving episodes. In particular, CARLA data is captured using adapted autopilots
6
that access all the CARLA privileged information needed for expert-level driving. That is, the autopilots know where the road is, lane lines, cars, pedestrians, traffic lights, etc. without the need to interpret images or any other sensory information. The methodology of the experimentation procedure was approved by the ”Comitè d’Ètica en la Recerca (CERec) de la Universitat Autònoma de Barcelona” (research ethics committee of the university, reference number: CEEAH 6936). For all the experiments, a written informed consent was obtained from each participant. The consent form explains the goal of the experiment and describes what kind of data are collected and the terms of privacy in the use of personal data. Additionally, it emphasizes that the data released to the general public does not contain information that can directly identify the subject and that any data and research results already shared with other investigators or the general public cannot be destroyed, withdrawn, or recalled. Each consent was hand-signed by each subject on the day of the first experiment. The eye-tracking glasses were manually calibrated for each driver prior to the start of every scenario. Both human driving data and data captured by autopilots were reviewed to eliminate episodes of erratic driving. III. DATA P REPARATION AND VALIDATION To synchronize the driver attention data with the visual driving context, the eye-tracking videos captured during driving sessions are mapped to high-fidelity panoramic imagery generated using the CARLA simulator’s playback functionality. It is worth noting that this process revealed a bug4 in CARLA’s playback system, where, in some cases, spawning new actors caused ID conflicts, leading to incorrect aliasing of IDs during playback. We fixed the bug before producing the playbacks contained in LAIA. As figure 6 shows, the attention mapping pipeline proceeds through the following stages: 1) Visual Data Recording, 2) Intermediate Panoramic Construction, 3) Gaze Mapping and Fixation Modeling, and 4) Final Panoramic Projection.
Fig. 6. Attention mapping pipeline.
Fig. 7. Original view of the eye tracker worn by the driver.
A. Visual Data Recording The raw visual data is recorded using eye-tracking glasses worn by the driver. As Figure 7 shows, the eyetracker’s scene camera records all simulator displays visible to the participant, including the rearview mirrors, which are positioned to replicate their real-world positions within the cockpit. Rendered square reference markers, with a unique color for each screen, are embedded into the borders of the simulator’s physical displays. These markers serve as fixed reference points to geometrically align the eyetracker video frames with the composite CARLA-rendered panorama. Their small size and simple shape were intentionally selected to reduce visual saliency and minimize the likelihood of interfering with the driver’s natural attention. Narrow vertical strips at the junctions between displays are minimally visible in the eyetracker video and are considered negligible for alignment purposes. 4 https://github.com/carla-simulator/carla/pull/8960
B. Intermediate Panoramic Construction Using projective geometry and the aforementioned markers, an intermediate panoramic image is constructed offline to establish a common geometric frame between the eyetracker footage and the simulator imagery. First, each raw eyetracker image is rotation-corrected using head-tilt information from the glasses’ Inertial Measurement Unit (IMU), aligning the camera view with the horizontal axis of the display setup. Next, color normalization is applied to compensate for illumination differences between the eyetracker video and the CARLA rendering. The colored reference markers are then detected with a color-based filtering procedure, and their configuration within the screen layout of simulator display are identified by a geometric analysis. Using these correspondences, a projective mapping (homography) is estimated to warp gaze coordinates from each eyetracker view into an
7
Fig. 8. Intermediate Panoramic Construction.
intermediate panoramic coordinate system extracted from the CARLA playback image. Figure 8 illustrates the mapping of an eyetracker frame into the intermediate panoramic coordinate system. For consistency with this panoramic playback image, the rearview mirrors are ultimately represented as the leftmost and rightmost sections of the resulting representation. Notice that this image spatially aligns the eyetracker’s visual field with the simulated environment but it contains only regions observed by the eyetracker and therefore excludes areas outside the eyetracker’s field of view. It is used solely as a staging space for robust, frameaccurate gaze projection. C. Gaze Mapping and Fixation Modeling To project driver attention, gaze data are mapped from the eyetracker frame onto this intermediate panorama. Following the Tobii I-VT Attention Filter, which Tobii explicitly recommends for Tobii Pro Glasses 3 recordings in mobile environments, a fixation is defined as a set of gaze points where the maximum angular velocity does not exceed 100 degrees per second and the duration ranges between 0.1 and 0.5 seconds. This higher threshold, relative to the 30°/s used in stationary lab studies, is necessary because head movements during driving inflate point-to-point velocities, causing lower thresholds to misclassify fixations as saccades [6]. As a known trade-off, this threshold includes smooth pursuits and a small fraction (10–15%) of short saccades, which slightly overestimates attention dwell time [7]. Fixations lasting longer than 0.5 seconds are divided into shorter fixations. Thus, attention is modeled as a two-dimensional Gaussian distribution centered at the gaze point, with the standard deviation σ accounting for multiple sources of uncertainty: • The intrinsic fixation error of the eye-tracking device (typically < 5 pixels in the eyetracker image). • Downsampling from the 100 Hz gaze stream to the 25 fps scene camera. • Temporal misalignment between eyetracker and simulator video playback. • Reduction of human attention to a single gaze coordinate per frame.
Some regions remain without information in this intermediate image, as they fall outside the eyetracker’s field of view. These areas are recovered in the next stage. D. Final Panoramic Projection The complete and accurate panoramic image is reconstructed offline using CARLA’s playback system, which renders all camera views (forward, lateral, rear, and mirrors) and can provide additional channels such as semantic labels, depth, optical flow, and CAN bus states. Because playback is driven by ground-truth CARLA simulator logs rather than the participant’s head pose, it produces a consistent, full-coverage panorama that includes regions not visible in the eyetracker footage. The attention maps computed in the intermediate space are then projected into this final panoramic coordinate system, ensuring that gaze and attention are precisely aligned with the full simulator scene. The Gaussian attention fields retain their modeled uncertainty, thereby propagating spatial and temporal confidence through the entire pipeline. This rigorous multi-step process ensures that human attention is accurately reflected in the original simulator-based visual scenes. The final result is suitable for training, evaluation, and interpretability analyses of attention-aware driving models. Figure 9 illustrates the final process of the panoramic image reconstruction together with the correct projection of gaze data. Notice that missing information in the eyetracker video data is recovered in the final panoramic view. This figure also shows samples of the depth channel and the semantic segmentation channel in the Frontal-Left and the Frontal-Right views, respectively. E. Qualitative Validation To qualitatively support the correctness of our Gaussian modeling of gaze attention, Figure 10 compares two Gaussianbased representations of the driver gaze: the native distribution provided by the eye-tracker (colored heatmap), together with the scanpath (red lines), and the distribution computed by our pipeline (red-outlined circle). The inner crosses mark
8
Fig. 9. Final Panoramic Projection. On the top, the intermediate panoramic image using projective geometry. On the bottom the final panoramic image generated using CARLA’s playback feature with an additional channel of semantic segmentation. Human attention is projected and expanded with uncertainty modeled by a Gaussian distribution.
IV. DATA F ORMAT AND S TRUCTURE
Fig. 10. Visualization of gaze uncertainty modeling and comparison to the original Gaussian provided by the eye-tracking device. The broader heatmap-like distribution corresponds to the native distribution provided by the eye-tracker device, while the red-outlined circles represent the attention distribution modeled by our pipeline.
the instantaneous gaze point recorded by the device, and the centers of both Gaussians are nearly perfectly concentric. Figure 11 shows different examples of qualitative validation of the whole mapping. The widths of our Gaussians vary depending on the duration of the fixation.
The dataset described has been made publicly available in a federated and multidisciplinary data repository [4] and in https://cloningdcb.org/. No registration is required and users can select and download specific data through a customizable download configurator. This interface allows users to tailor their dataset selection based on their needs, ensuring flexible and efficient access to the data. Below, we list the publicly available data, while Figure 12 presents representative examples of each data modality included in the dataset. For each driver and each scenario the data available are as follows: 1) Playback RGB images, obtained from central, right and left cameras, as well as left and right mirrors (Fig 12(a)). These images are generated in 6 different weather conditions: clear noon, clear sunset, Hard rain noon, wet noon, cloudy noon, mid rain sunset. 2) Panoptic Segmentation: instance (Fig 12(b)) and semantic segmentation (Fig 12(c)). 3) Depth (pixel distance obtained per camera using a CARLA depth camera) (Fig 12(d)). 4) Optical flow data (Fig 12(e)). 5) CAN bus information at every timestamp, at a frequency of 25 Hz. It contains a .json file with the information of speed, acceleration, brake, throttle, steer, ego position, gyroscope, accelerometer, hand brake, rear gear, blinker, and direction. All of these signals are synchronized with all the data. 6) Attention map generated from eye-tracker mapped on the aligned RGB images (Fig 12(f)).
9
Fig. 11. Qualitative validation of the eye-tracking mapping onto the original simulator-grounded visual scenes. Each grouped image contains the original image captured by the eye-tracking on the bottom, the intermediate panoramic image with a heat map visualizing the projected human attention, together with the instant gaze produced by the eye-tracker device, and the final panoramic image with the gaze data accurately projected.
Internally, data are structured as follows: driver map route data_modality camera weather frame In addition to the dataset itself, a separate directory named miscellaneous contains the scripts used to generate the driving routes, including the corresponding voice instructions in three languages. This material supports the reproducibility of the experiments and also enables researchers to go beyond the current dataset by extending the scenarios or introducing their own agents.
semantically meaningful regions such as pedestrians, vehicles, or traffic signs, or making it possible to assess whether the models focus on the same scene elements that human drivers consider relevant, ultimately fostering the development of more explainable AI systems. The deliberate design of specific driving events along the routes also makes it possible to study how gaze patterns correlate with driving maneuvers in response to particular situations, such as a pedestrian emerging unexpectedly or a vehicle cutting in, thus supporting research on driving behavior modeling and prediction. Finally, by combining human attention with rich ground-truth annotations, the dataset provides a valuable benchmark for evaluating perception models under diverse weather conditions, viewpoints, and occlusion levels. These applications demonstrate the versatility of LAIA as a resource for improving autonomous driving systems.
V. E XAMPLES OF A PPLICATION The LAIA dataset enables a wide range of applications aimed at advancing both the performance and interpretability of AI-based driving systems. In particular, the availability of synchronized gaze data, control signals, and rich scene annotations opens new directions for cognitive modeling and human-centered AI development. Next, we outline several representative use cases. From the perspective of gaze data alone, the dataset supports the training and evaluation of gaze prediction models. In addition, LAIA enables systematic comparison between the attention of AI drivers and that of human drivers, which can help assess whether models focus on
VI. C ONCLUSION We presented LAIA, a novel dataset designed to support the development of attention-aware end-to-end autonomous driving systems. By combining raw sensory input, control actions, and high-resolution human gaze data in a wide variety of simulated scenarios, LAIA offers a unique resource for studying the interplay between perception, decision-making, and visual attention. A key strength of LAIA lies in its deliberate design choices regarding data density and repeatability. Rather than maximizing the number of distinct environments, the dataset
10
(a)
(b)
(c)
(d)
(e)
(f)
Fig. 12. Illustration of available data. a) RGB , b) Panoptic Instance Segmentation, c) Panoptic Semantic Segmentation, d) Depth, e) Optical flow, and f) Gaze overlaid on an RGB image.
prioritizes depth over breadth: 44 participants drove through the same carefully crafted routes across a reduced set of towns, ensuring that each scenario was experienced by a large and diverse group of drivers. This design yields a high degree of within-scenario variability in human attention and driving behavior, since individual differences in gaze patterns, reaction times, and control responses can be directly compared under identical environmental conditions. The result is a statistically robust foundation for learning-based models, where the same scene is observed through the eyes of dozens of different drivers spanning a range of ages, experience levels, and driving styles. With over 15 hours of recorded driving synchronized with eye-tracking data, control signals, and rich ground-truth annotations, LAIA provides the repetition density that allows
a robust comparison between AI and human drivers. Furthermore, the six weather variations applied uniformly across all sequences add an orthogonal axis of robustness, ensuring that attention patterns are available not only across participants but also across lighting and meteorological conditions for the same underlying scene geometry. Combined with the controlled placement of driving events — designed specifically to elicit natural human reactions — this structure allows researchers to study not just where drivers look on average, but how attention shifts and varies in response to specific, reproducible stimuli. The dataset enables new research directions focused on improving the transparency, interpretability, and cognitive alignment of AI drivers. By grounding end-to-end model
11
development in the rich, repeated, and diverse human attention signals that LAIA provides, the autonomous driving community gains a tool designed not only for training, but for rigorous evaluation and meaningful comparison of attentionaware systems under controlled and repeatable conditions. R EFERENCES [1] N. Kalra and D. G. Groves, “The enemy of good: Estimating the cost of waiting for nearly perfect automated vehicles,” RAND Corporation, Tech. Rep. RR-2150-RC, 2017. [Online]. Available: https://doi.org/10.7249/RR2150 [2] L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li, “End-to-end autonomous driving: Challenges and frontiers,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 10 164– 10 183, 2024. [Online]. Available: https://doi.org/10.1109/TPAMI.2024. 3435937 [3] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “Carla: An open urban driving simulator,” in Conference on robot learning. PMLR, 2017, pp. 1–16. [4] A. Contreras, D. Porres, R. Abad, P. Cano, A. Levy, G. Villalonga, A. M. López, and A. Hernández-Sabaté, “LAIA: Labelled Attention for Intelligent Automobiles,” 2025. [Online]. Available: https://doi.org/10. 34810/DATA2570 [5] “Tobii pro glasses 3 — latest in wearable eye tracking,” accessed: 202607-28. [Online]. Available: https://www.tobii.com/products/eye-trackers/ wearables/tobii-pro-glasses-3 [6] A. Hossain and E. Miléus, “Eye movement event detection for wearable eye trackers.(2016),” Linköpings Universitet thesis LiTH-MAT-EX2016/02-SE, 2016, uRN: urn:nbn:se:liu:diva-129616. [7] R. Andersson, L. Larsson, K. Holmqvist, M. Stridh, and M. Nyström, “One algorithm to rule them all? an evaluation and discussion of ten eye movement event-detection algorithms,” Behavior research methods, vol. 49, no. 2, pp. 616–637, 2017. [Online]. Available: https://doi.org/10.3758/s13428-016-0738-9