Conceptio › Archive › arXiv CS
arXiv CSopen access

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

Hwanjo Heo et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation Hwanjo Heo∗

Junhee Lee∗

Jinwoo Kim†

[email protected] ETRI Daejeon, Republic of Korea

[email protected] Chungbuk National University Cheongju, Republic of Korea

[email protected] Chungbuk National University Cheongju, Republic of Korea

arXiv:2609.35040v1 [cs.CR] 28 Sep 2026

Abstract Mixed-reality headsets such as Apple Vision Pro replace the touchscreen with gaze-as-pointer interaction: the wearer looks at a target and confirms with an air pinch. Because the display is inside the headset and the eye tracker is walled off from third-party software, such input is widely assumed to be unobservable to bystanders—a built-in defense against the shoulder-surfing that plagues phones and laptops. We present GAZEleak, a side-channel attack that recovers gaze-driven input from external video of head motion alone, under a strictly local, physical-observer threat model: the adversary only films the wearer from across the room and installs no software on the device. The attack exploits the centrally coupled eye–head motor program—gaze shifts recruit small, targetdependent head reorientations—which projects into sub-degree pose changes recoverable from commodity video. GAZEleak implements a measurement-based, sparse-optical-flow inference pipeline for users’ 6-digit device passcode. On a preliminary front-view dataset from three author-subjects who were aware of the attack hypothesis, GAZEleak places the true code within the top ten guesses for 56% (10 of 18) of test codes under a cross-person protocol with no labeled victim data, and for every code of the most exposed subject. The performance is subject-dependent: no passcode from the least exposed subject reaches the top ten, although its median guessed passcode rank is 12,786 rather than 500,000 expected from an uninformative ordering. These results provide preliminary evidence that gaze-coupled head motion can expose passcode information under controlled conditions, while motivating broader evaluation across users, behaviors, and capture settings.

CCS Concepts • Security and privacy → Side-channel analysis and countermeasures; Privacy protections; • Human-centered computing → Mixed / augmented reality; Virtual reality.

Keywords XR security, eye tracking, side-channel attack, passcode inference, head motion ∗ Co-first authors. † Corresponding author.

This work is licensed under a Creative Commons Attribution 4.0 International License. WPES ’26, The Hague, Netherlands © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-3026-9/2026/11 https://doi.org/10.1145/3847192.3847361

ACM Reference Format: Hwanjo Heo, Junhee Lee, and Jinwoo Kim. 2026. GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation. In 25th Workshop on Privacy in the Electronic Society (WPES ’26), November 15–19, 2026, The Hague, Netherlands. ACM, New York, NY, USA, 13 pages. https://doi.org/10.1145/3847192.3847361

1

Introduction

Personal computing is moving beyond the smartphone toward wearable, head-mounted devices that promise always-on, handsfree interaction. The first commercial wave is already deployed: Apple Vision Pro [3] and Meta Quest 3 [15] ship as mixed-reality headsets, while smart glasses such as Meta Ray-Ban [22] and the recently previewed Meta Orion [17] target an ambient-computing form factor. Apple Vision Pro exemplifies a shift in how users provide input: the wearer gazes at an interface element to point and performs an air pinch to confirm the selection. Unlike a touchscreen tap or a handheld-controller movement, the pinch reveals when a selection occurs but not which virtual target was selected. Even with this confirmation gesture visible, the interaction appears resistant to shoulder surfing for two reasons. First, the display is enclosed by the headset and is not directly visible from an ordinary external viewpoint. Second, the operating system restricts access to raw gaze data, preventing ordinary third-party applications from observing where the user looks. A nearby observer can therefore see the wearer and the confirmation gesture while remaining unable to see either the rendered user interface or the selected input. These protections conceal the digital input channel, but they do not necessarily conceal the wearer’s physical response. Gaze shifts can recruit small, target-dependent head movements, and the headset makes those movements visible as image-plane translation. The selected target may consequently leave a physical side channel even when the display and eye-tracking stream remain inaccessible. The security question here is whether an adversary can infer sensitive information with some confidence by analyzing the victim’s head motions recorded in external videos (see Figure 1). To answer this question, we present GAZEleak, a side-channel attack that infers gaze-entered passcodes using external video of an XR-headset wearer’s head motion. The attack exploits the centrallycoupled eye–head motor program: gaze shifts can recruit small, target-dependent head movements [25], and the resulting subdegree pose changes are recoverable from commodity video using classical sparse feature tracking. While some XR keylogging attacks have been proposed so far, they do not address this physical adversary. TyPose [27] inferred virtual-keyboard input from device-reported head-motion telemetry, while GAZEploit [32] recovered typed input from the animated eyes of a Vision Pro Persona

WPES ’26, November 15–19, 2026, The Hague, Netherlands

Hwanjo Heo, Junhee Lee, and Jinwoo Kim

the performance with various factors. Section 7 discusses results and security implications. Section 8 reviews related work, and Section 9 concludes the paper. 1

3

4

5

6

3

Adversary

Victim

2 Inferring Passcode via Head-motion Analysis

Recording Tools

Figure 1: Our threat model: an adversary records the head motions of a victim wearing an XR device to infer their 6digit passcode entered via air-pinch gestures.

This section provides the background and motivation for GAZEleak. We first review eye tracking as an emerging input modality in consumer XR devices (Section 2.1), then describe the eye–head coordination that makes gaze shifts externally observable (Section 2.2), and finally explain how this physical coupling motivates our attack (Section 2.3).

2.1 transmitted over FaceTime. These attacks exploit signals exposed through the device or a software-rendered representation. We instead study a local and physical observer who records the victim directly and receives no cooperation from the headset, operating system, or third-party software. Device-side access control and avatar suppression do not remove the underlying motion signal, although interface changes can reduce its exploitability. Scope of this paper. To evaluate GAZEleak, we perform a controlled, preliminary feasibility study of six-digit gaze-and-pinch passcode entry only from frontal camera recordings of the author subjects. The results therefore establish signal availability under controlled conditions rather than a population-level estimate of attack success. A larger study with unaware participants, additional camera viewpoints, and input tasks beyond numeric passcodes is planned subject to institutional ethics approval. We discuss these limitations in Section 7.5. Contributions. We make the following contributions. • To our knowledge, we present the first passcode-inference attack against gaze-as-pointer XR interaction that relies solely on external video. The adversary requires no software on the target device and no access to its display, operating-system APIs, eye-tracking stream, or head-motion telemetry. • We design and implement GAZEleak, an end-to-end inference pipeline that detects visible air-pinch events, segments a passcode entry into inter-key transitions, measures sub-degree head motion using classical sparse feature tracking, predicts keypad displacements, and accumulates transition evidence to rank all 106 six-digit passcode candidates. • We conduct a controlled three-subject feasibility study under both within-person and cross-person protocols, using securityrelevant outcomes that include top-10 accuracy and the true passcode’s rank. Under the cross-person protocol, GAZEleak places the true code within ten guesses for 56% (10 of 18) of all test codes and for every code of the most exposed subject. • We characterize the conditions that determine attack effectiveness through experiments on repeated observations, spatial and temporal capture degradation, transition-label granularity, scoring methods, descriptor selection, and vision model alternatives. The rest of the paper is organized as follows. Section 2 provides necessary background and our motivation. Section 3 defines our adversary assumption. Section 4 and Section 5 present the design and implementation of GAZEleak, respectively. Section 6 evaluates the attack under within- and cross-person protocols and examines

Background

Eye-Tracking Capability in XR Devices

Eye tracking has shifted from a laboratory peripheral to a built-in sensor in mainstream consumer headsets. The Apple Vision Pro is the most notable example: rather than treating gaze as an auxiliary signal for foveated rendering or social-presence avatars, Vision Pro elevates it to the primary pointing modality, with the operating system inferring the user’s target object from where the eye is looking and confirming selection through an air pinch. Comparable hardware exists on the Meta Quest Pro [16], which integrates Tobiibased gaze trackers behind the lenses, and HTC’s Vive Pro Eye and Vive XR Elite [13], although these primarily use gaze for rendering optimization or social-VR animation rather than as a UI pointer. Several smaller platforms (e.g., Pico Neo 3 Pro Eye, Pico 4 Pro, PlayStation VR2) ship with eye-tracking subsystems as well. Eye tracking has not yet become standard in lighter AI glasses form factor. Meta’s commercial Ray-Ban smart glasses include cameras and audio but no eye-tracking sensor. The recently-previewed Meta Orion prototype and the research-oriented Meta Aria Gen 2 platform [23] both incorporate inward-facing eye cameras, signaling that mainstream AI glasses are converging toward the same gaze-as-input paradigm established by current XR headsets. We therefore treat gaze-as-pointer as the de facto interaction modality across the XR device class our threat model targets.

2.2

Synchronization of Eye and Head Movement

A long-standing finding in motor neuroscience is that eye and head movements are not independently controlled but driven by a single coordinated motor program. Sidenmark and Gellersen [25] establish this directly in a VR setting: across 7,600 gaze shifts collected from 20 participants, they report that natural gaze re-orientation systematically recruits all three of the eyes, head, and torso, and that the relative contribution of each component varies as a function of the target’s angular position. Earlier work in non-immersive settings reaches the same conclusion — gaze shifts decompose into an eye component and a head component driven by a centrally coordinated command [8]. In the meantime, it is also reported that the head barely contributes when targets fall within a comfortable oculomotor range, but as the angular demand grows the head increasingly compensates [28]. Crucially, the eye–head coupling is largely involuntary and cannot simply be suppressed at will. Corneil et al. [10] record neckmuscle EMG in head-restrained primates and observe that the appearance of a visual target evokes a sharply time-locked, lateralized activation of the neck muscles within roughly 55–95 ms of stimulus

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

onset, even though the head itself is physically prevented from moving. The central motor command for head rotation is therefore issued in parallel with, and often preceding, the eye saccade; mechanical restraint hides the overt movement but does not prevent its generation. In an XR setting, where the wearer is unrestrained, the small but consistent head motion accompanying each gaze shift may therefore be observable to an external camera that can resolve sub-degree head pose — the side channel that underlies the attack we present in this paper.

2.3

Motivation

Our study is motivated by direct observation of how people use gaze-as-pointer XR devices in practice. While working alongside colleagues who routinely use headsets such as Apple Vision Pro, one of the authors noticed that most users do not hold the head still and move only the eyes. Instead, they reorient the head noticeably toward on-screen targets, to the point that the movement is plainly visible to a bystander. This is the everyday manifestation of the eye–head coupling reviewed in Section 2.2: gaze-recruited head motion is not a corner case but the common mode of interaction, and it is overt enough to be captured by an ordinary camera. The tendency appears especially pronounced during authentication processes. When users unlock the headset with a passcode or type a password on a virtual keyboard, they tend to enter each character deliberately, orienting the head toward the target key and exaggerating the movement relative to casual browsing.1 An informal conversation with several users suggested why: during authentication they are anxious about input errors, because a mistaken digit or character must be undone with awkward gaze-driven backspacing, so they slow down and aim carefully at each key. The care that improves their entry accuracy also produces larger, cleaner, more target-specific head movements. This observation is what makes passcode entry a compelling target rather than a hypothetical one. The setting in which the user is most careful—security-sensitive authentication—is precisely the setting in which the gaze-coupled head-motion side channel is strongest. We note that this accuracy-driven exaggeration is distinct from, and works against, the behavior of an attack-aware user who might consciously suppress head motion (which we treat as a limitation in Section 7.5): a typical user, unaware of the attack and focused only on entering the code correctly, is inclined to move more. These observations are informal and motivate our evaluation with real measurements and also the systematic user study we plan as future work.

3

Threat Model

Figure 1 illustrates our threat model. We assume that a victim is a user of a commercial XR headset or AI-glasses whose primary or co-primary pointing modality is eye gaze. Our running example is Apple Vision Pro, in which the OS recognizes the target object from the user’s gaze and the user confirms with an air pinch. The model generalizes to other eye-tracked devices such as Meta Quest Pro, HTC Vive Pro Eye, and the emerging class of AI glasses with inward-facing eye cameras. 1 A similar observation motivates a VR keylogging attack that requires malicious

software to be installed on the headset [27].

WPES ’26, November 15–19, 2026, The Hague, Netherlands

The adversary is strictly external. They do not install or compromise software on the device, do not query any OS-level API, and do not have access to the inward micro-display, the eye-tracker stream, or the on-device gaze coordinates. They cannot mount any software-level side-channel attack on the headset. The adversary can capture video of the victim’s head and upper body from a passive vantage point: a smartphone held discreetly in the same room, a fixed surveillance camera, a webcam on a nearby laptop, or a body-worn camera. Recordings are short (seconds to a minute) at standard consumer resolution and frame rate (e.g., 4K/30fps), with no calibration to the victim or the device. The adversary aims to exfiltrate sensitive gaze-driven input by the victim. Specifically, in this paper, we aim to infer the six-digit passcode that is entered with gaze selection and air-pinch confirmation. The goal of the adversary is to infer the correct six-digit passcode from 106 possible candidates. In many practical passcode settings, several retries may be possible before lockout or escalation, and an attacker who can reduce the search space to a top-𝑘 list has substantially increased risk. We choose 𝑘 = 10 as the Apple Vision Pro allows up to 10 failed passcode attempts [5]. Among a number of prospective attack scenarios by an adversary with gaze-inferring capability, we consider the passcode inference by an adversary with opportunistic passcode input session capture. The victim wakes their XR headset after sleep, a power cycle, or an inactivity timeout and is prompted for a 6-digit device passcode.2 The adversary is co-located (a coffee shop, an office, a lobby) and records the brief (< 10 s) entry from a smartphone in hand or pocket, or from a CCTV camera overhead. A single capture may be noisy, but the adversary may be able to accumulate multiple recordings of the same passcode over time.

4

GAZEleak Design

This section describes the design of GAZEleak. We first give an overview of the whole attack pipeline and formalize the passcode as a sequence of keypad transitions (Section 4.2). The following subsections then detail each inference stage, from transition segmentation (Section 4.3) and motion description (Section 4.4) to label scoring (Section 4.5) and passcode ranking (Section 4.6).

4.1

Overview

Figure 2 illustrates the passcode inference pipeline of GAZEleak, which converts an external passcode-entry video into a ranked list of six-digit candidates. End-to-end, GAZEleak 1 detects six air-pinch events from the victim recording, 2 segments the five intervening head-motion transitions (Section 4.3), 3 encodes each transition into a fixed-length motion descriptor (Section 4.4), 4 maps each descriptor to a keypad displacement and scores it against every candidate transition label (Section 4.5), and 5 pools the perposition evidence to rank all 106 candidate passcodes (Section 4.6). The unknown passcode is not used during feature extraction, scoring, or ranking. In addition, GAZEleak operates in two modes: (i) offline training where it learns transitions from a victim or others, and (ii) attack time that infers the passcode of a victim. 2 Apple Vision Pro requires the user to input the passcode, even when Optic ID (Apple’s

iris-based biometric authentication system) is enabled, after a reboot or an extended period of inactivity for security purposes [4].

WPES ’26, November 15–19, 2026, The Hague, Netherlands

1

INPUT RECORDING

2

TRANSITION SEGMENTATION

3

Hwanjo Heo, Junhee Lee, and Jinwoo Kim

4

TRANSITION MOTION DESCRIPTION

5

SCORING

RANKED CANDIDATES

OFFLINE TRAINING

1

2

3

4

5

6

7

8

9

0

Reference video

labeled x→y transitions 100 pairs

T1

T2

T3

T4

T5

REFERENCE MODEL

SHARED TRANSITION ENCODER

Known press boundaries

each Tj is uniformly sampled to 12 frames

CANDIDATE D

ridge regression

෡ reference 𝒅

or

centroid five inter-key clips

3

5

6

7

8

9

0 D = (d0, …, d5)

trained model

4

S(D) = ∑ Aj j=0

Air-pinch timing normalized pinch distance

head ROI + Shi-Tomasi + elliptical mask Lucas-Kanade

8 motion channels

1

f1

f2

f3

f4

f5

f6

select 14 of 100 dims + z-score

[f1, f2], …, [f5, f6]

target <𝒅𝒋 >

ℓμ(Δ̂) = −‖Δ̂ − Δμ‖2 / 2σ|Δμ|2 T1 T2 T3 T4 T5

transition label μ

five transition clips

(κ(dj, dj + 1))

POSITION EVIDENCE

100

12 stats/channel + 4 path features

8 x 12 + 4 = 100-d descriptor d

unknown 6 digits

up to several recordings

2

4

μ: movement label (Δcol, Δrow)

ATTACK TIME

Victim Recording

1

enumerate 106 codes 1 2 3 ⋮ k

D(1) D(2) D(3)

D(k)

Ranked top-k guesses

Figure 2: Overview of GAZEleak attack pipeline for 6-digit passcode inference.

4.2

Transition-Based Passcode Model

To formalize the gaze-inference problem, we model the passcode as a sequence of transitions rather than as six independent digit observations. Let 𝐷 = (𝐷 0, . . . , 𝐷 5 ) be a six-digit passcode. Here, we define three transition labels at different granularities. The exactpair label at position 𝑗 is 𝜋 (𝐷 𝑗 , 𝐷 𝑗+1 ) = 10 · 𝐷 𝑗 + 𝐷 𝑗+1,

𝑗 ∈ {0, . . . , 4}.

Thus, the numeric keypad induces 100 exact transition labels. For robustness analysis, we also consider coarser label spaces. Let each digit have a keypad coordinate (col(𝐷 𝑗 ), row(𝐷 𝑗 )). The movement label is:  𝜇 (𝐷 𝑗 , 𝐷 𝑗+1 ) = Δcol, Δrow , which groups exact pairs that share the same keypad displacement where Δcol = col(𝐷 𝑗+1 ) − col(𝐷 𝑗 ) and Δrow = row(𝐷 𝑗+1 ) − row(𝐷 𝑗 ). The direction label further coarsens this to:  𝛿 (𝐷 𝑗 , 𝐷 𝑗+1 ) = sgn(Δcol), sgn(Δrow) . At the two extremes, exact-pair labels retain absolute source and destination information, whereas direction labels retain only the signs of keypad displacement. Consequently, direction scoring constrains the passcode path but generally leaves multiple exact passcodes indistinguishable unless additional absolute-position evidence is available. GAZEleak adopts the movement label 𝜇 as its readout (Section 4.5); we treat exact-pair and direction as alternatives evaluated in the ablation (Section 6.6).

4.3

frames 𝑓0 < · · · < 𝑓5 . The five transition clips are the intervals [𝑓0, 𝑓1 ], [𝑓1, 𝑓2 ], . . . , [𝑓4, 𝑓5 ]. For labeled transition video clips used at the training phase (Section 6), CSV recorded key input boundaries are utilized for convenience.

Transition Segmentation

As each digit selection is confirmed by an air pinch motion, the six presses partition the recording into five transitions. Because the keypad is invisible to the camera, we recover the press timing from the wearer’s hand alone. For every frame, we run a single-hand landmark detector [34] and compute the thumb-tip–to–index-tip distance, normalized by an intrinsic hand scale (wrist to index knuckle) so that it is invariant to hand size and camera distance. A pinch is a sharp local minimum of this distance (the fingertips touch); we smooth the signal, extract local minima, and take the six deepest, subject to a minimum temporal separation, as the press

4.4

Transition Motion Description

The signal is measured inside a per-view head/headset region of interest (ROI) 𝐵 = (𝑥 0, 𝑦0, 𝑥 1, 𝑦1 ), chosen to contain stable headset, visor, face, and head-boundary texture while excluding the hands and as much background as possible. We localize 𝐵 automatically with a promptable segmentation model (e.g., SAM 2 [21]) and confirm it manually. The resulting ROI is fixed within a recording, but separate recordings may use different boxes. Within the ROI, we detect feature points only inside an inscribed elliptical mask. The head is approximately elliptical, so the rectangle corners are the pixels most likely to hold static background; the ellipse suppresses those background and ROI-corner points while retaining the high-gradient head boundary. The mask constrains only the initial point detection—its boundary is never used as a motion cue—which avoids segmentation jitter while preserving stable texture for optical flow. We encode each transition clip with a fixed-length descriptor obtained from sparse optical flow, following the classical detectthen-track paradigm. Each clip is temporally resampled to 𝑁 = 12 frames (uniform indices over [𝑓𝑘 , 𝑓𝑘+1 ]) so descriptors are invariant to transition duration. For each transition clip, we crop the ROI from the sampled frames and convert the crops to grayscale. On the first frame we detect Shi–Tomasi corners [24]. The selected points are then tracked across the sampled frames with pyramidal Lucas–Kanade optical flow [9]. Points that fail the tracker status check are dropped. Let p (0) = (𝑥 𝑗(0) , 𝑦 𝑗(0) ) ⊤ be the location of tracked corner 𝑗 in 𝑗 the first sampled frame 𝑓0 and let p (𝑖𝑗 ) = (𝑥 𝑗(𝑖 ) , 𝑦 𝑗(𝑖 ) ) ⊤ be its Lucas– Kanade location in the 𝑖-th sampled frame 𝑓𝑖 . From these correspondences, we estimate the dominant 2-D similarity transform with

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

RANSAC [11]:  𝑎 b p (𝑖𝑗 ) = 𝑖 𝑏𝑖

−𝑏𝑖 𝑎𝑖

 𝑡𝑥,𝑖 h (0) 𝑥𝑗 𝑡 𝑦,𝑖

𝑦 𝑗(0)

1

i⊤ .

Here, (𝑡𝑥,𝑖 , 𝑡 𝑦,𝑖 ) is translation in ROI pixels, while 𝑎𝑖 and 𝑏𝑖 en√︃ code isotropic scale 𝜌𝑖 = 𝑎𝑖2 + 𝑏𝑖2 and in-plane rotation 𝜃 𝑖 = atan2(𝑏𝑖 , 𝑎𝑖 ). The constrained model excludes shear and nonuniform scale, which are not expected from short head movements, and RANSAC suppresses static-background features and tracking outliers that remain inside the ROI. We normalize translation as (𝑡𝑥,𝑖 /𝑊 , 𝑡 𝑦,𝑖 /𝐻 ) using the ROI width 𝑊 and height 𝐻 . Because a single fitted transform can be unstable when few tracks agree, we also retain robust statistics of the raw point flow u 𝑗,𝑖 = p (𝑖𝑗 ) − p (0) 𝑗 . The per-frame motion vector is   med 𝑗 𝑢𝑥,𝑗,𝑖 med 𝑗 𝑢 𝑦,𝑗,𝑖 IQR 𝑗 𝑢𝑥,𝑗,𝑖 IQR 𝑗 𝑢 𝑦,𝑗,𝑖 𝑡𝑥,𝑖 𝑡 𝑦,𝑖 m𝑖 = 𝑊 , 𝐻 , 𝜃 𝑖 , 𝜌𝑖 − 1, , , , , 𝑊 𝐻 𝑊 𝐻 comprising transform translation, rotation, scale change, median point flow, and point-flow dispersion. The sequence M = (m1, . . . , m𝑁 −1 ) therefore contains eight motion channels over the 𝑁 − 1 frame-toreference estimates. To obtain a duration-invariant descriptor we summarize each channel over the series with twelve statistics capturing both its endpoint and its dynamics (final value, mean, standard deviation, etc). We append four path-geometry scalars computed from the translation and rotation channels—net displacement magnitude, total translation path length, final rotation angle, and total absolute angular variation—that describe the trajectory shape. This gives a 8 × 12 + 4 = 100-dimensional transition descriptor d ∈ R100 .

4.5

Label Scoring

We extract a fixed 14-dimensional subset of the 100-dimensional descriptor, chosen for stable head-motion features, and normalize each kept dimension with a z-score 𝑥ˆ 𝑗 = (𝑥 𝑗 − 𝑚 𝑗 )/𝑠 𝑗 .3 The mean 𝑚 𝑗 and the standard deviation 𝑠 𝑗 come from the reference set when the target provides labeled transitions, and from the target’s own transitions otherwise. In the cross-person protocol, the adversary has no labeled data from the victim but does have the recordings of the attacked passcode (five entries; Section 6.3). We therefore compute 𝑚 𝑗 and 𝑠 𝑗 from the transitions of those recordings. This removes the victim’s constant head-pose bias before the regressor is applied. For each transition, a ridge regressor predicts the keypad moved Δrow) š from the standardized descriptor. It is ment 𝜇ˆ = ( Δcol, trained on the reference transitions, whose movements are known. In the cross-person protocol (Section 6.1) the target provides no labels, so we refine the pooled regressor by unsupervised selfcalibration: we repeatedly snap each predicted movement to the nearest valid keypad move, refit the map to those assignments, and iterate. This aligns the reference map to the target’s own motion scale without using any target label. We score each candidate movement 𝜇 by how close it is to ˆ with a Gaussian on their gap, written ℓ (𝜇) = the prediction 𝜇, 2 ), where the variance grows with the movement −∥ 𝜇ˆ − 𝜇 ∥ 2 /(2𝜎 |𝜇 | 3 The descriptor-subset choice, i.e., retaining 14 of the total of 100 dimensions, is

validated by our ablation study in Section 6.6.

WPES ’26, November 15–19, 2026, The Hague, Netherlands

magnitude so that large movements are judged with a wider tolerance. A softmax over labels turns these scores into a per-position log-posterior, ∑︁ ′ log 𝑃 (𝜇 | x̂) = ℓ (𝜇) − log 𝑒 ℓ (𝜇 ) , 𝜇′

which the pooling step below accumulates.

4.6

Passcode Ranking by Position Pooling

Each target transition belongs to one of the five transition indices 𝑗 ∈ {0, . . . , 4} within the passcode sequence. We accumulate the perposition label evidence by summing the log-posteriors of the tranÍ sitions observed at that position, 𝐴 𝑗 (𝜇) = 𝑖: pos(𝑖 )=𝑗 log 𝑃 (𝜇 | x̂𝑖 ), and floor labels unseen in the reference to a small log-probability. A candidate passcode 𝐷 = (𝐷 0, . . . , 𝐷 5 ) is then scored by summing the position-appropriate label evidence over its five transitions, 𝑆 (𝐷) =

4 ∑︁

 𝐴 𝑗 𝜇 (𝐷 𝑗 , 𝐷 𝑗+1 ) .

𝑗=0

We enumerate all 106 candidate codes, sort by 𝑆 (𝐷) in descending order, and report the rank of the true code—equivalently, the attacker’s ordered guess list, of which the top 𝑘 matters when the device permits 𝑘 attempts (Section 3).

5

Implementation

We implemented the full pipeline of GAZEleak and report the concrete implementation details here for reproducibility. GAZEleak VisionOS App. To collect labeled data with precise ground truth, we built a data-collection application4 for Apple Vision Pro that reproduces the system numeric passcode keypad—a 3 × 4 grid whose key size and placement match the native passcode screen—rendered at a fixed position in the immersive scene (Figure 3). The user selects keys with the device’s native gaze pointer and air-pinch confirmation, so the captured behavior matches ordinary gaze-as-pointer input. The app was built with Swift 5.1 for visionOS 2.5. The app runs two session types. Training sessions prompt the user with target digits and log each accepted press: a short Phase 1 warm-up that visits the digits 0–9 in order, followed by Phase 2, which traverses the ordered digit-to-digit transitions in a Eulerian sequence and repeats that traversal until the session reaches roughly 300 transitions, so every transition is collected the same number of times. Test sessions prompt the user to enter six six-digit passcodes—three Shared codes that are identical across profiles and three Personal codes unique to each profile—and each press is logged with its profile, passcode index, position within the code, and entered digit. The visionOS app streams these events over the local network to a companion macOS app, which stores them in CSV format. GAZEleak Training/Inference Pipeline. All training and inference stages run offline on the recorded video; no component runs on or communicates with the headset. Recordings are made with a single commodity smartphone camera placed in front of the seated 4 Our app is a measurement instrument for constructing our dataset; it is not part of

the attack. The adversary in our threat model (Section 3) installs no software on the victim’s device and uses only external video.

WPES ’26, November 15–19, 2026, The Hague, Netherlands

Figure 3: VisionOS app’s keypad layout mimics the system passcode interface. Participants are instructed to input the highlighted (green) digit for training sessions (left) or assigned 6-digit passcode for test sessions (right).

wearer and encoded at 4K/30fps. Frames are processed at native resolution; no per-wearer or per-device calibration is applied. The attack pipeline is implemented offline in Python with OpenCV and it runs on Intel Xeon Gold 6526Y Linux server with NVIDIA H100 and no component executes on the headset. Each transition clip is resampled to 𝑁 = 12 frames. Inside the elliptical ROI mask we detect up to 450 Shi–Tomasi corners (quality level 0.004, minimum spacing 7 px) and track them with pyramidal Lucas–Kanade, then fit a partial-affine transform with RANSAC to suppress background and outlier tracks. Each transition yields a 100-dimensional descriptor, of which a fixed 14-dimensional subset is standardized and mapped to a keypad displacement by a ridge regressor with regularization 𝜆 = 1.0.

6

Evaluation

This section evaluates GAZEleak. In particular, we aim to answer the following research questions (RQ): • RQ1. Does GAZEleak accurately infer the victim’s transitions and passcode? (Section 6.2) • RQ2. Does observing multiple entries of the same passcode improve inference? (Section 6.3) • RQ3. Does passcode difficulty affect inference performance? (Section 6.4) • RQ4. Does the quality of the captured video affect inference performance? (Section 6.5) • RQ5. Which design components of GAZEleak contribute to the inference performance? (Section 6.6)

6.1

Experiment Settings

Data Collection. We collected a controlled dataset of passcodestyle gaze input on an Apple Vision Pro, as a representative XR platform. A participant wore the headset and entered digits on the numeric keypad, whose size and position were matched to the Apple Vision Pro system layout (Figure 3), while a single RGB camera recorded the wearer head-on in 4K resolution. Each session

Hwanjo Heo, Junhee Lee, and Jinwoo Kim

follows an Eulerian sequence that visits every ordered digit-todigit transition uniformly, so that all pairwise head movements between the ten keys are represented, and none is over-sampled; this reference session supplies the labeled transitions used to fit the scorer. Separate held-out sessions provide six-digit test passcodes. Participants. All three authors served as subjects. As no external participants were involved, this study did not require institutional review (Section 7.6). We report each subject separately rather than averaging, because the amount of head motion a wearer produces differs from person to person and that difference turns out to drive the result. Metrics. To measure how GAZEleak infers the victim’s transitions, we measure transition accuracy, the fraction of the five transitions per code whose predicted keypad displacement lands on the correct destination key, for which chance is 0.10. To see how GAZEleak infers the victim’s passcode, we use top-10 accuracy, the fraction of test codes whose true code appears within the attacker’s first ten ranked guesses, which matches the ten failed attempts that Apple Vision Pro permits before lockout (Section 3). Note that a uniform guesser reaches 10/106 . Additionally, we use median rank, the median position of the true code when all 106 candidates are sorted by the attacker’s predicted likelihood, where a uniform guesser sits at 500,000. Passcodes with identical movement sequences receive identical scores; none of the 18 test passcodes shares its movement sequence with another valid passcode, so ties did not affect the reported ranks. Protocols. We frame each evaluation as an attack on one victim, and report two protocols that differ in where the attacker’s reference transitions come from. In the within-person protocol, the training session and the test codes come from the same person, so the victim supplies their own labeled transitions. In the cross-person protocol, the reference transitions come from the other two subjects and the victim contributes no labels; an unsupervised self-calibration step aligns the pooled map to the victim’s own transition statistics. With only two reference subjects, this protocol measures transfer from a small reference pool rather than population-level generalization.

6.2

Passcode Recovery Performance

Table 1 reports transition accuracy and passcode recovery results. Performance varies substantially across the three subjects. For Subject 1, every test passcode appears within the top ten under both protocols, with median rank 1 within-person and 3 cross-person. Subject 2 reaches top-10 accuracy of 0.50 within-person and 0.67 cross-person, with median ranks 9 and 6, respectively. Subject 3 has no top-10 recovery: its transition accuracy is 0.30 within-person and 0.27 cross-person, and its median passcode ranks are 9,528 and 12,786. Across the three subjects, the cross-person protocol recovers 10 of the 18 test passcodes (56%) within ten guesses, and the within-person protocol recovers 9 of 18 (50%). Although Subject 3 does not meet the operational top-10 criterion (Section 3), its ranking remains more informative than an unstructured search. We discuss the operational implications in Section 7.1. The central pattern is that stronger passcode recovery coincides with more clearly separated recovered keypad geometry. Figure 4 visualizes recovered keypad geometry and helps explain the observed pattern. Subjects 1 and 2 produce separated landing clusters

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

Subject 1 code 572398

Table 1: Per-subject inference performance under the withinand cross-person protocols. Transition accuracy is the fraction of the five inter-digit transitions in each six-digit passcode whose predicted displacement reaches the correct destination key (chance: 0.10). Top-10 accuracy is the fraction of test passcodes whose true code is ranked among the first ten of 106 candidates, with the number of recovered codes out of the six test passcodes per subject in parentheses. Median rank is the median position of the true code in the ranked list (lower is better; an uninformative ordering has median rank 500,000). Every result pools the five recordings of each test passcode (𝑁 = 5). Section 6.3 varies the observation budget. Within-person Transition Subject

Acc.

Subject 1 Subject 2 Subject 3

0.93 0.73 0.30

1.00 (6/6) 0.50 (3/6) 0.00 (0/6)

1 9 9,528

Transition Acc.

Med. Rank

0.80 0.67 0.27

1.00 (6/6) 0.67 (4/6) 0.00 (0/6)

3 6 12,786

Subject 3

Passcode

2

3

4

5

6

1

2

3

4

5

6

4

5

6

7

8

9

7

8

9

0 predicted transition (ours)

start key

Figure 5: Estimated trajectories of the best case (Subject 1) and worst case (Subject 3) when recovering a sample passcode.

6.3

1

3

Figure 5 provides a complementary transition-level diagnostic for one example passcode for the best (Subject 1) and the worst case (Subject 3). For visualization, both trajectories are anchored at the ground-truth initial digit and formed by integrating the five continuous displacement estimates. Subject 1 preserves the approximate path geometry, whereas the first two predicted displacements of Subject 3 have notably smaller amplitudes than the corresponding true movements, after which the accumulated trajectory drifts away from the keypad path. This contraction could reflect Subject 3’s greater eye–in–head contribution to those gaze shifts, but we cannot attribute it to eye compensation without an independent eye-motion measurement. Note that this diagnostic visualization shows only transition-level error and does not depict the exhaustive passcode-ranking procedure.

target digit

0

2

Passcode Top-10 Acc.

Subject 2

1

true path (ground truth)

Transition

Subject 1

Med. Rank

Subject 3 code 572398

0

Cross-person

Passcode Top-10 Acc.

WPES ’26, November 15–19, 2026, The Hague, Netherlands

7

8

9

Figure 4: Recovered keypad geometry from (top) labeled transition and (bottom) passcode entry sessions. Each small point is a predicted destination position, colored by its groundtruth destination digit for evaluation; the larger colored markers show each digit’s recovered keypad position, taken as the centroid of its predicted landing points, and the gray circles mark the nominal keypad location.

that preserve the keypad’s spatial organization, whereas Subject 3’s clusters are compressed toward the center and overlap. This pattern is consistent with lower externally observable head-motion amplitude for Subject 3, but does not necessarily establish that the subject relied primarily on eye rotation because eye motion was not measured independently.

Effect of Repeated Observations

Our threat model assumes that the adversary can record the same passcode more than once (Section 3). To see if repeated observations help an adversary recover the true passcode, we recorded the test passcode five times. To avoid privileging an arbitrary record ing order, we evaluate all 𝑁5 subsets for each observation budget 𝑁 ∈ {1, . . . , 5}. For a given subset, the selected entries are pooled position by position at decoding time; this operation does not retrain or otherwise change the transition scorer. Pooling combines only the five recordings of the same passcode. The other five test passcodes are not used, because a victim has only one device passcode. We evaluate each of the six test passcodes separately in this way. We then compute top-10 accuracy and median rank across the six test passcodes. Figure 6 plots, for each subject and 𝑁 , the median  of each metric over the 𝑁5 subsets, with the shaded band spanning the minimum and maximum. The bands therefore measure sensitivity to which recordings are available. The benefit of repetition differs across subjects. For Subject 1, the median 𝑁 = 1 subset already achieves top-10 accuracy 1.00 and median rank 2, so additional observations provide little improvement. Subject 2 benefits more: its median rank decreases from 174 at 𝑁 = 1 to 19 at 𝑁 = 2 and 9 at 𝑁 = 4, while top-10 accuracy

WPES ’26, November 15–19, 2026, The Hague, Netherlands

Hwanjo Heo, Junhee Lee, and Jinwoo Kim

Table 2: Exploratory comparison of entry-combination strategies under the within-person protocol. Entry selection strategy

(a) Top-10 accuracy

Subject

𝐸𝑐𝑣

All entries

Best entry

Selected strategy

Subject 1 Subject 2 Subject 3

0.10 0.12 0.49

1 9 9,528

2 84 128

All entries All entries Best entry

(b) Median rank

Figure 6: Effect of the number of observed entries 𝑁 on (a) top-10 accuracy and (b) median rank. Lines show the median over all combinations of choosing 𝑁 entries from the five available recordings, and shaded bands span the corresponding minimum and maximum numbers.

increases from 0.17 to 0.50. Subject 3 does not reach the top-10 criterion, but its median rank decreases from 251,507 at 𝑁 = 1 to 9,528 when all five entries are pooled. Repeated observations can therefore reinforce consistent evidence and substantially reduce the search space even when they do not produce immediate top10 recovery. The min–max bands also show that the gain is not determined by 𝑁 alone: different subsets at the same observation budget can produce substantially different rankings. For Subject 3, the best single recording yields a median true-passcode rank of 128, whereas the worst performs near the random-ranking baseline. Exploratory Entry Selection. The subset variability in Figure 6 raises a separate question: can the adversary identify an informative recording without knowing the true passcode? To this end, let 𝐸𝑟 be the mean predicted motion magnitude in entry 𝑟 . We measure how much these values differ across the five entries with the labelfree coefficient of variation 𝐸 cv = sd(𝐸 1, . . . , 𝐸 5 )/mean(𝐸 1, . . . , 𝐸 5 ). We then consider two different entry selection strategies: (i) all entries: an adversary averages all entries; (ii) best entry: an adversary chooses the entry with the largest predicted head-motion magnitude. Note that selecting recordings with large predicted head-motion magnitude is a heuristic rather than a validated optimal strategy. Establishing its generalizability requires evaluation on a substantially larger cohort with the selection rule fixed in advance. Table 2 shows that 𝐸 cv predicts which strategy wins. When the entries are alike (𝐸 cv = 0.10 and 0.12 for Subjects 1 and 2), averaging cancels per-entry noise and pooling all five is best; keeping a single entry instead carries its noise and can hurt sharply, raising Subject 2’s median rank from 9 to 84. When the entries differ (𝐸 cv = 0.49 for Subject 3, whose quietest recording carries far less predicted motion than its strongest), pooling dilutes the one informative entry, so retaining the largest-motion entry is far better and cuts the median rank from 9,528 to 128. A label-free rule that pools all entries at low 𝐸 cv and keeps only the largest-motion entry at high 𝐸 cv therefore matches the better fixed strategy for every subject (last column of Table 2). Because the selection uses only predicted displacement magnitudes, it requires no labels and remains available in the cross-person setting.

Figure 7: Passcode recovery for the three passcodes shared by all subjects under the within-person protocol. Each line shows one subject’s median true-passcode rank as total vertical keypad travel increases from 4 to 10 steps.

6.4

Effect of Passcode Structure

To examine whether keypad-path geometry affects recovery, we focus on the three passcodes shared by all subjects. This matched subset holds the passcode constant across subjects and avoids comparing different personal codes. For a passcode 𝐷, we summarize Í its path by total horizontal and vertical travel, 𝑇ℎ = 4𝑗=0 |Δcol 𝑗 | Í4 and 𝑇𝑣 = 𝑗=0 |Δrow 𝑗 |, measured in keypad steps. The three shared codes have the same horizontal travel, 𝑇ℎ = 4, but increasing vertical travel: 𝑇𝑣 = 4 for 378569, 𝑇𝑣 = 5 for 572398, and 𝑇𝑣 = 10 for 293068. Figure 7 shows an association between greater vertical travel and poorer recovery. The subject-level pattern is not uniform, however. The ranks for Subjects 2 and 3 worsen as vertical travel increases, whereas Subject 1 remains at rank 1 or 2 for all three codes. This subject is already near the best possible rank, leaving little room for the code structure to produce a visible monotonic effect. Unfortunately, our preliminary experiment is not designed to isolate passcode-structure effects. Figure 7 only provides exploratory evidence of passcode-dependent difficulty, not an estimate of the effect of vertical travel. A future study should use a larger, balanced set of passcodes that independently varies horizontal and vertical travel, step lengths, and direction changes with randomized passcode assignment.

6.5

Robustness to Video Capture Quality

The recordings used in our main experiments (Section 6.2) rely on 4K resolution at 30 frames per second, a setting available on commodity smartphones. In a surveillance setting, however, the wearer’s head may occupy only a small part of a wide-field recording, so cropping the head produces a substantially lower-resolution

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

(a) Transition acc. per resolution

(b) Med. rank per resolution

WPES ’26, November 15–19, 2026, The Hague, Netherlands

(c) Transition acc. per FPS

(d) Med. rank per FPS

Figure 8: Effect of inference performance per captured video quality. (a, b) Transition accuracy and median rank across capture resolutions. (c, d) Transition accuracy and median rank across frame rates. Recovery remains effective down to 540p and 6 fps. input to the tracker. Stored CCTV footage may also have a lower frame rate [29]. These conditions motivate evaluating both spatial and temporal degradation. To isolate the effect of each sampling axis, we degrade the same source videos along one axis at a time and rerun the pipeline. Varying Video Resolution. Figure 8a and 8b show the transition accuracy and median ranks when varying the video resolutions. As video resolution gets lower, we observe that the inference performance is preserved to 540p, but degrades at 270p where the head region spans roughly one hundred and sixty pixels. Considering that 270p is an extreme condition, our attack is robust at standard video resolutions. Additionally, the degradation is not uniform across subjects. Subject 2 falls from a transition accuracy of 0.47 to 0.00, and its median rank rises from 26 to 436,481, whereas Subject 1 still reaches 0.87 and Subject 3 is already low enough that little changes. Varying Video Frame Rates. Figure 8c and 8d show the transition accuracy and median ranks when varying the video frame rates. As our pipeline relies on 30 fps, we evaluate lower frame rates, where temporal information is progressively lost. Inference performance remains stable at most lower frame rates before failing abruptly. At 6 fps, the transition accuracy remains 0.80 for Subject 1 and 0.43 for Subject 2, with no change in the order of magnitude of the median rank. At 2 fps, however, all three subjects fall to a transition accuracy between 0.07 and 0.13, effectively reaching chance performance. This suggests that the descriptor remains robust to moderate temporal degradation but fails once too little temporal information is retained. In contrast, reducing the spatial resolution degrades performance more steadily, with failure occurring at 270p. Overall, these findings indicate that the attack does not require specialized capture equipment. A standard surveillance camera or smartphone operating at 1080p provides sufficient detail, supporting the practical viability of the threat model described in Section 3.

6.6

Pipeline Ablations

We now justify the main design choices in GAZEleak through a sequence of ablations. We (i) replace the whole descriptor-and-readout stack with a frozen learned end model, (ii) vary the transition label granularity and the readout while holding the rest of the pipeline fixed, and (iii) vary the descriptor subset that the scorer consumes. Unless stated otherwise, all numbers are reported within-person and averaged over the three subjects

Pretrained Visual Representations. We first replace our motion descriptor and regressor with a frozen pretrained model that predicts the destination key of each transition through a linear classifier trained on its frozen features. The image models receive the final frame of each transition, whereas the video models receive the complete transition clip. Table 3: Evaluating frozen learned models as replacements for our pipeline. Image and video models take the final frame and the full clip as input, respectively. Transition

Passcode

Model

Temporal Input

Acc.

Top-10 Acc.

Med. Rank

VideoMAE-B V-JEPA2 ViT-L DINOv2 ViT-L R(2+1)D-18 CLIP ViT-L

Transition clip Transition clip Final transition frame Transition clip Final transition frame

0.36 0.33 0.26 0.20 0.16

0.11 0.00 0.00 0.00 0.00

12,526 14,536 101,520 161,066 393,790

Ours (LK)

Transition clip

0.66

0.50

9

Table 3 compares VideoMAE [30], V-JEPA 2 [6], DINOv2 [19], R(2+1)D [31], and CLIP [20]. The two strongest learned alternatives are the video models: VideoMAE reaches 0.36 transition accuracy and V-JEPA 2 reaches 0.33. However, R(2+1)D performs below the image-based DINOv2 model, indicating that temporal input alone does not explain performance. The best learned alternative, VideoMAE, yields a median true-passcode rank of 12,526, compared with 9 for GAZEleak. This experiment establishes a substantial performance gap for the tested configurations, but it does not by itself identify the cause; Section 7.2 discusses reduced feature resolution as one plausible explanation. Label Granularity and Readout. Table 4 crosses three transitionlabel spaces (Section 4.2) with the applicable readouts. Exact-pair labels 𝜋 retain the source and destination digits as 100 categorical labels. Movement labels 𝜇 merge transitions with the same two-dimensional keypad displacement into 31 labels, and direction labels 𝛿 retain only the signs of the horizontal and vertical displacements, producing 9 labels. The nearest-centroid and multinomial classifier readouts apply to all three spaces. Ridge regression applies to the vector-valued movement and direction spaces, but not to arbitrary exact-pair identifiers, which have no meaningful Euclidean ordering. Among the three label granularities, movement 𝜇 gives the best passcode recovery, and it matches the mechanics of gaze-coupled

WPES ’26, November 15–19, 2026, The Hague, Netherlands

Hwanjo Heo, Junhee Lee, and Jinwoo Kim

Table 4: Joint effect of transition-label granularity and readout under the within-person protocol with all five entries pooled. Granularity

Readout

Top-10 Acc.

Median Rank

Exact-pair 𝜋

nearest centroid multinomial classifier ridge regression

0.22 0.17

4,086 1,751

nearest centroid multinomial classifier ridge regression (ours)

0.17 0.28 0.50

846 356 9

nearest centroid multinomial classifier ridge regression

0.06 0.22 0.11

3,730 165 258

Movement 𝜇

Direction 𝛿

N/A

directly, i.e., the summary statistics of the translation and in-planerotation channels together with the path-geometry terms. Unlike a PCA projection, which is competitive (rank 12 at 14 dimensions) but mixes all channels into opaque components, this subset preserves the physical reading of every axis, which is what lets the recovered displacements be interpreted as keypad geometry (Section 6.2). The subset is chosen once and applied unchanged to every subject and to both protocols.

7

Discussion

This section interprets the residual ranking beyond the top-10 criterion (Section 7.1), explains why modern visual models underperform our classical pipeline (Section 7.2), and then discusses a behavioral confound, countermeasures, limitations, and ethical considerations.

Table 5: Descriptor-subset selection comparison.

7.1 Transition

Passcode

# dims

Selection

Acc.

Top-10 Acc.

Median Rank

100 20

All entries PCA

0.48 0.59

0.33 0.50

635 26

14 14 14

Random PCA Custom (ours)

0.43 0.62 0.66

0.28 0.50 0.50

618 12 9

8

PCA

0.59

0.44

23

head motion, which depends chiefly on how far the gaze travels rather than on the absolute source and destination keys. Because displacement lives in a continuous, ordered space, we read it out with a ridge regressor rather than a classifier, which improves the median rank from 356 to 9; a regression readout is either inapplicable to the unordered exact-pair or inferior in the case of the most coarse granularity of direction. Descriptor Subset. Table 5 reports the inference performance with selection strategies of features that are fed into the descriptor (Section 4.4). Note that our 100-dimensional descriptor is deliberately over-complete: it collects many summary statistics per motion channel so that no informative cue is missed at extraction time. Feeding all 100 entries to the scorer, however, is counter-productive—each subject supplies only several hundred reference transitions, which cannot constrain a 100-dimensional ridge map, and the resulting overfit inflates the median rank to 635. The scorer therefore consumes a reduced subset, and two decisions make that subset principled rather than arbitrary: how many entries to keep, and which ones. Comparing the PCA (Principal Component Analysis) baseline, the median rank peaks at 14 (rank 12) and is worse at both 8 (rank 23) and 20 (rank 26). Thus, 14-dimension is the empirical knee that balances descriptive capacity against the variance budget of a several-hundred-transition reference set. At a fixed dimension, which entries are kept matters more than the count: a random 14entry subset reaches only median rank 618, because most descriptor statistics are redundant or weakly coupled to keypad displacement. We instead retain the entries that carry the displacement signal most

Security Impact Beyond Top-10 Recovery

Top-10 recovery captures an immediately actionable online attack under the guess budget used in our evaluation, but a top-10 miss does not imply an absence of leakage. For Subject 3, the median true-passcode rank is 9,528 within-person and 12,786 cross-person (Section 6.2). Thus, at the median, the true passcode lies within the top 0.95% and 1.28% of the 106 -code space, respectively, rather than near the median rank of 500,000. This reduction is insufficient to unlock an unmodified device within ten attempts, but it shows that Subject 3 still leaks information that can materially prioritize a larger guessing campaign. The operational significance of this reduction depends on a second capability: the adversary must obtain more verification attempts than the normal interface permits. Prior work such as bypassing the retry counter with NAND mirroring [26] and parallelized brute-forcing with UID key extraction [14] have demonstrated such capabilities on earlier Apple devices. They do not, however, demonstrate such a compromise against Apple Vision Pro. Apple’s current design binds passcode derivation to a per-device UID and enforces delays and attempt counters through the Secure Enclave [2]. Our claim is therefore conditional: if a well-resourced adversary can separately obtain an offline verifier, reset verification state, or instantiate independent verifier copies, GAZEleak’s ranked output substantially reduces the work of the ensuing search. The residual rank also identifies opportunities to strengthen the inference stage. Training on a larger and more diverse participant cohort could better cover variation in eye–head coordination and reduce cross-person domain shift. Multiple synchronized viewpoints could provide complementary projections of the same small rotation and remain informative when one view is occluded or nearly insensitive to a motion axis. Repeated observations of the same passcode could likewise be pooled, with motion-aware weighting or robust aggregation so that weak entries do not dilute stronger ones.

7.2

Why Modern Visual Models Fall Short

GAZEleak utilizes the classical vision-based tracking method to quantify the user’s head movement (Section 4) instead of utilizing recent advanced vision or video models. Our design choice is based on our benchmark results (see Table 3) showing that frozen learned

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

representations underperform with a significant margin. Even the strongest learned alternative, VideoMAE-B, records several orders of magnitude higher median rank. A plausible explanation is a mismatch between the representations and the signal of interest. GAZEleak must distinguish subdegree head rotations that may produce only subpixel or few-pixel displacements in the captured video. Modern image and video encoders commonly reduce spatial resolution through strided feature maps [31], patch tokens [19, 20], or spatio-temporal tubelets [6, 30]. Although these representations support strong semantic and long-range correspondence performance, they may attenuate the subpixel-to-few-pixel displacement that carries our side-channel signal.

7.3

Prompt Consultation as a Behavioral Confound

The performance difference across subjects may partly reflect how they performed the experimental task, rather than only an inherent difference in eye–head coordination. During post hoc review, Subject 3, who is also an author, reported repeatedly consulting the assigned passcode displayed above the keypad (Figure 3, right) instead of memorizing the complete six-digit sequence before entry. Such prompt consultation violates an assumption of our transition model: the motion between two consecutive pinch events may include a gaze shift toward the prompt and back, rather than primarily the direct movement from one keypad digit to the next. The resulting composite trajectory can distort the direction and amplitude assigned to a digit transition and may have contributed to Subject 3’s lower accuracy. This explanation is retrospective, however, and cannot be established from the present data because prompt-directed glances were not annotated or experimentally controlled. This behavior differs from our intended device-unlock scenario, in which a user normally enters a familiar passcode without consulting an on-screen copy. Future studies should therefore separate passcode familiarity from inference difficulty. Before each entry, participants should be shown an assigned synthetic passcode, required to memorize all six digits, and allowed to begin only after the prompt has been removed and recall has been verified. To better approximate habitual entry without collecting real credentials, participants could practice synthetic passcodes to a predefined accuracy criterion before recording. A controlled comparison among visible-prompt, newly memorized, and practiced-passcode conditions would then quantify whether prompt consultation explains part of the per-subject performance gap. Until such an experiment is performed, Subject 3’s result should not be attributed solely to a stable low-head-motion profile.

7.4

Countermeasure

Physical Leakage Beyond Software Control. Unlike previous keystroke or gaze-inference attacks that can be prevented by a software patch (Section 8), GAZEleak exploits a biomechanical signal—the head reorientation that the central motor program couples to every gaze shift (Section 2.2)—from an external camera the device does not control. A vendor cannot completely suppress it without redesigning gaze interaction itself.

WPES ’26, November 15–19, 2026, The Hague, Netherlands

Mitigations. One direct defense is to shrink the spatial extent of sensitive UI elements. If all passcode buttons lie within a small angular region near the current fixation, gaze shifts can be completed mostly by eye rotation and are less likely to recruit measurable head motion. This micro-UI mode could be reserved for short secrets such as passcodes or OTP entry, rather than used throughout the interface. The cost is usability: smaller targets increase visual crowding, selection errors, dwell time, and accessibility burden, especially for users with reduced acuity, tremor, or less stable eye tracking. A secure design would therefore need to quantify the privacy gain against error rate and entry time. Another simple yet effective defense is to randomize the keypad layout during sensitive input. Randomization weakens the attack because a fixed head-motion direction no longer maps consistently to the same digit transition. It is particularly useful against an adversary who pools multiple recordings of the same passcode: if the layout changes across entries, repeated observations do not reinforce the same ordered-pair classes. The drawback is again interaction cost. Randomized keypads slow entry, increase cognitive load, and make motor memory unusable. They may also increase legitimate lockouts if users enter familiar passcodes under an unfamiliar layout. A practical system might therefore use partial randomization, row/column shuffling, or randomization only after risk triggers such as public-location use, recent failed attempts, or external-camera warnings. We view robust mitigation as an open problem and a motivation for future research.

7.5

Limitation and Future Direction

There are several limitations. First, the current data come from author subjects rather than recruited unaware participants. The knowledge of attack hypothesis can bias head behavior in either direction: a subject may consciously hold the head still and rely on eye movement alone — suppressing the very signal we measure — or may over-rotate the head toward each key, inflating it. Either way, the head motion recorded here is unlikely to match that of an unaware user entering a passcode out of habit. The experimental protocol for recruited participants should be carefully designed to differentiate the effect of the subject’s awareness of attack hypothesis that can be expressed by suppressing their head motion. Second, the front perspective is favorable because the face and headset are visible; real surveillance may observe the user from the side, behind, or under partial occlusion. Third, the six-digit passcode setting is only one gaze-input task. Passwords, OTPs, text entry, app focus, and document selection may have different leakage profiles. Finally, we only evaluate our attack against a single device, i.e., Apple Vision Pro. We plan to address these limitations with an IRB-approved user study, more participants, multiple camera perspectives, carefully designed experiment protocol and additional tasks beyond numeric passcodes.

7.6

Ethical Concerns

All recordings were collected from the authors themselves. Because the study involves no third-party participants, it does not constitute human-subjects research and does not require institutional review board (IRB) approval. Nor are any real secrets exposed: the

WPES ’26, November 15–19, 2026, The Hague, Netherlands

Hwanjo Heo, Junhee Lee, and Jinwoo Kim

Table 6: Qualitative comparison of GAZEleak with prior keystroke and gaze-inference attacks on XR devices. “External observ.” indicates that the exploited signal can be captured by sensing the victim from outside the headset; “No SW install” indicates that the attack requires no attacker-controlled software on the victim’s device; and “Resists SW patch” indicates that the leakage comes from an externally measurable physical signal rather than a software-exposed data or rendering channel.

Attack

Exploited signal

Input modality

External observ.

GAZEploit [32] TyPose [27] Hidden Reality [12] VR-Spy [1] VRecKey [18] Yang et al. [33]

Avatar eye animation Head motion (IMU) Hand-gesture video Wi-Fi CSI Controller IR LEDs Avatar hand motion

Gaze-as-pointer Controller/hand Controller/hand Controller/hand Controller/hand Phys. keyboard

✗ ✗ ✓ ✓ ✓ ✗

✓ ✗ ✓ ✓ ✓ ✓

✗ ✗ ✓ ✓ ✓ ✗

GAZEleak (ours)

Head motion (video)

Gaze-as-pointer

✓

✓

✓

entered codes are experimental stimuli generated for the study, not genuine device passcodes or account credentials, so neither our data nor its release reveals a real passcode. The larger study with recruited, unaware participants that we outline as future work will be conducted only under IRB approval, with informed consent and no collection of participants’ real credentials. Although we evaluate GAZEleak on Apple Vision Pro, the observed leakage should not be interpreted as a defect specific to the vendor or device. It arises from gaze-coupled head motion observable outside the headset and may therefore affect other systems that use gaze to select spatially separated interface elements. Nevertheless, because Apple Vision Pro is our evaluated platform and these findings may inform the design of its security-sensitive interfaces, we have shared our preliminary findings with Apple by filing a security report. We also present this issue as a broader design-level concern for the gaze-as-pointer research community.

No SW Resists SW install patch

GAZEleak removes this assumption entirely: head motion is observed externally, from a commodity camera, with no software running on the victim’s device. Table 6 summarizes related work. GAZEleak is the first inference attack against gaze-as-pointer XR devices from local and external observation by an adversary. A second line of work studies the use of eye-tracking data itself as a biometric signal. Aziz and Komogortsev [7] show that gaze streams exposed to third-party applications — in combination with head and hand motion streams — enable re-identification of XR users with high accuracy even when state-of-the-art privacy mechanisms are applied to each stream independently. These attacks differ from GAZEleak in that they require the gaze stream to be available to the attacker (i.e., they target the output of the eye tracker), whereas GAZEleak reconstructs gaze direction without observing it directly.

Acknowledgements 8

Related Work

The privacy implications of eye-tracking sensors integrated into consumer XR headsets have only recently begun to attract systematic attention. Most closely related to our work is GAZEploit [32], which infers the gaze direction of an Apple Vision Pro user by analyzing the eye animation of the user’s Persona avatar broadcast through a video call. While GAZEploit and GAZEleak share the underlying goal of recovering gaze without direct access to the device’s eye-tracker, the two attacks differ fundamentally in threat model: GAZEploit assumes the victim is actively sharing their Persona via a video stream while GAZEleak operates against any wearer visible to an external camera, regardless of avatar broadcast or OS version, and exploits a biomechanical leak (eye–head coordination) that vendors cannot fully eliminate without redesigning the gaze-based interaction. On VR headsets that predate gaze-as-pointer interaction, TyPose [27] showed that virtual-keyboard input on VR headsets can be reconstructed from the wearer’s head-pose telemetry alone. The attack relies on the same biomechanical signal we exploit — head motion synchronized with gaze — but operates from a fundamentally different vantage point: it assumes the attacker has installed a background application on the headset itself, where Inertial Measurement Unit (IMU) data are accessible without any special permission.

This work was partially supported by Institute of Information & communications Technology Planning & Evaluation (IITP) [RS2023-00215700, Trustworthy Metaverse: blockchain-enabled convergence research, 70%] and the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) [RS2026-25495427, 30%].

9

Conclusion

We presented GAZEleak, a preliminary external-observation attack on gaze-based passcode entry in XR. The attack exploits gazecoupled head motion that can be observed by a nearby adversary without access to the display or eye-tracking data. Although preliminary, our evaluation suggests that protecting gaze input may require more than concealing visual output and eye-tracking streams; it may also require mitigating information leakage through users’ physical movements. We hope this work motivates the development of practical defenses and further investigation of gaze-input privacy under a broader range of users, adversaries, and real-world attack scenarios.

References [1] Abdullah Al Arafat, Zhishan Guo, and Amro Awad. 2021. Vr-spy: A side-channel attack on virtual key-logging in vr headsets. In 2021 IEEE Virtual Reality and 3D User Interfaces (VR). IEEE, 564–572.

GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation

[2] Apple Inc. 2022. The Secure Enclave, Apple Platform Security. https://support. apple.com/guide/security/the-secure-enclave-sec59b0b31ff. Accessed: 2026-0522. [3] Apple Inc. 2024. Apple Vision Pro. https://www.apple.com/apple-vision-pro/. Accessed: 2026-05-22. [4] Apple Inc. 2024. How to set up Optic ID. https://support.apple.com/en-us/120145. Accessed: 2026-05-21. [5] Apple Inc. 2026. Set a passcode and use Optic ID to unlock Apple Vision Pro. https://support.apple.com/guide/apple-vision-pro/set-a-passcode-and-useoptic-id-tanf79745605/visionos. Accessed: 2026-07-09. [6] Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. 2025. V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985 (2025). [7] Samantha Aziz and Oleg Komogortsev. 2025. Exploring the uncoordinated privacy protections of eye tracking and vr motion data for unauthorized user identification. In 2025 IEEE Conference Virtual Reality and 3D User Interfaces (VR). IEEE, 217–227. [8] Emilio Bizzi, Ronald E Kalil, and Vincenzo Tagliasco. 1971. Eye-head coordination in monkeys: evidence for centrally patterned organization. Science 173, 3995 (1971), 452–454. [9] Jean-Yves Bouguet et al. 2001. Pyramidal implementation of the affine lucas kanade feature tracker description of the algorithm. Intel corporation 5, 1-10 (2001), 4. [10] Brian D Corneil, Etienne Olivier, and Douglas P Munoz. 2004. Visual responses on neck muscles reveal selective gating that prevents express saccades. Neuron 42, 5 (2004), 831–841. [11] Martin A Fischler and Robert C Bolles. 1981. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 24, 6 (1981), 381–395. [12] Sindhu Reddy Kalathur Gopal, Diksha Shukla, James David Wheelock, and Nitesh Saxena. 2023. Hidden reality: Caution, your hand gesture inputs in the immersive virtual world are visible to all!. In 32nd USENIX security symposium (USENIX Security 23). 859–876. [13] HTC Corporation. 2019. VIVE products. https://www.vive.com/us/product. Accessed: 2026-05-22. [14] Oleksiy Lisovets, David Knichel, Thorben Moos, and Amir Moradi. 2021. Let’s take it offline: Boosting brute-force attacks on iPhone’s user authentication through SCA. IACR Transactions on Cryptographic Hardware and Embedded Systems (2021), 496–519. [15] Meta Platforms, Inc. 2023. Meta Quest 3. https://www.meta.com/quest/quest-3/. Accessed: 2026-05-22. [16] Meta Platforms, Inc. 2024. Learn about Eye Tracking on Meta Quest Pro. https: //www.meta.com/help/quest/8107387169303764/. Accessed: 2026-07-25. [17] Meta Platforms, Inc. 2024. Orion. https://about.fb.com/realitylabs/orion/. Accessed: 2026-05-22. [18] Tao Ni, Yuefeng Du, Qingchuan Zhao, and Cong Wang. 2024. Non-intrusive and unconstrained keystroke inference in vr platforms via infrared side channel. arXiv preprint arXiv:2412.14815 (2024).

WPES ’26, November 15–19, 2026, The Hague, Netherlands

[19] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin ElNouby, et al. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023). [20] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748–8763. [21] Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. 2025. Sam 2: Segment anything in images and videos. In International Conference on Learning Representations, Vol. 2025. 28085–28128. [22] Ray-Ban and Meta Platforms, Inc. 2023. Ray-Ban Meta AI Glasses. https://www. ray-ban.com/usa/ray-ban-meta-ai-glasses. Accessed: 2026-05-22. [23] Michael Rice, Lorenz Krause, and Waqar Shahid Qureshi. 2025. EgoDrive: Egocentric Multimodal Driver Behavior Recognition using Project Aria. In Proceedings of the First International Workshop on Gaze Data and Natural Language Processing. 18–25. [24] Jianbo Shi et al. 1994. Good features to track. In 1994 Proceedings of IEEE conference on computer vision and pattern recognition. IEEE, 593–600. [25] Ludwig Sidenmark and Hans Gellersen. 2019. Eye, head and torso coordination during gaze shifts in virtual reality. ACM Transactions on Computer-Human Interaction (TOCHI) 27, 1 (2019), 1–40. [26] Sergei Skorobogatov. 2016. The bumpy road towards iPhone 5c NAND mirroring. arXiv preprint arXiv:1609.04327 (2016). [27] Carter Slocum, Yicheng Zhang, Nael Abu-Ghazaleh, and Jiasi Chen. 2023. Going through the motions: { AR/VR } keylogging from user head motions. In 32nd USENIX Security Symposium (USENIX Security 23). 159–174. [28] John S Stahl. 1999. Amplitude of human head movements associated with horizontal saccades. Experimental brain research 126, 1 (1999), 41–54. [29] Surveillance Camera Commissioner. 2018. Surveillance Camera Commissioner’s Buyers Toolkit. UK Home Office, https://assets.publishing.service.gov.uk/ government/uploads/system/uploads/attachment_data/file/711368/MVP_v5.1_ for_SCC_sign_off.pdf. Accessed: 2026-07-24. [30] Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35 (2022), 10078–10093. [31] Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 6450–6459. [32] Hanqiu Wang, Zihao Zhan, Haoqi Shan, Siqi Dai, Maximilian Panoff, and Shuo Wang. 2024. Gazeploit: Remote keystroke inference attack by gaze estimation from avatar views in vr/mr devices. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 1731–1745. [33] Zhuolin Yang, Zain Sarwar, Iris Hwang, Ronik Bhaskar, Ben Y Zhao, and Haitao Zheng. 2024. Can virtual reality protect users from keystroke inference attacks?. In 33rd USENIX Security Symposium (USENIX Security 24). 2725–2742. [34] Fan Zhang, Valentin Bazarevsky, Andrey Vakunov, Andrei Tkachenka, George Sung, Chuo-Ling Chang, and Matthias Grundmann. 2020. Mediapipe hands: On-device real-time hand tracking. arXiv preprint arXiv:2006.10214 (2020).

Record · ID 1108623 · SHA-256 1535e0cbcce581a4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.