Conceptio › Archive › arXiv CS
arXiv CSopen access

X-SPUR: Explainable Surprisal-Based Protocol-Aware Unsupervised Reasoning for Automotive Ethernet Intrusion Detection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

1

X-SPUR: Explainable Surprisal-Based Protocol-Aware Unsupervised Reasoning for Automotive Ethernet Intrusion Detection

arXiv:2609.21217v1 [cs.CR] 18 Sep 2026

Jisoo Kim, Student Member, IEEE, and Seonghoon Jeong, Member, IEEE

Abstract— Automotive Ethernet carries heterogeneous multi-protocol traffic in modern in-vehicle networks, where labeled attack data are rarely available and the strongest prior unsupervised detector still relies on handcrafted traffic features. This article presents X-SPUR, an explainable, surprisal-based, protocol-aware unsupervised reasoning framework that instead represents raw packet fields as token sequences, learns benign traffic patterns through causal language modeling, and detects anomalies from per-token cross-entropy surprisal. To incorporate temporal context, we introduce a bimodal fusion architecture that combines payload-token embeddings with inter-packet timing through additive fusion and a Hadamard interaction. To handle the heterogeneous score distributions of different protocol families, we further propose a dual top-k% perprotocol Z-score calibration that jointly captures moderately distributed and sparse anomaly signatures. On the TOW-IDS dataset, X-SPUR achieves an AUC of 0.9987. This is marginally higher than the 0.9969 reported for AERO. X-SPUR also eliminates handcrafted feature engineering. We train a separate CarDS model using the same architecture and training hyperparameters. This model retains strong performance on the second automotive Ethernet dataset. Beyond detection, per-token surprisal provides fine-grained explainability by attributing anomaly scores to specific protocol fields, supporting interpretable security analysis in heterogeneous in-vehicle networks. Index Terms— Anomaly Detection, Automotive Ethernet, Explainability, Mamba, Unsupervised Learning

I. I NTRODUCTION

A

UTOMOTIVE Ethernet is rapidly replacing the legacy Controller Area Network (CAN) as the backbone of

This article has been accepted for publication in IEEE Transactions on Industrial Informatics. This is the author’s accepted manuscript. Manuscript received July 13, 2026; revised September 6, 2026; accepted September 11, 2026. This research was supported by Sookmyung Women’s University Research Grants (1-2403-2028). (Corresponding author: Seonghoon Jeong.) Jisoo Kim is with the Department of Data Science, Sookmyung Women’s University, Seoul 04310, Republic of Korea (e-mail: [email protected]). Seonghoon Jeong is with the Division of Artificial Intelligence Engineering, Sookmyung Women’s University, Seoul 04310, Republic of Korea (e-mail: [email protected]). © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

modern in-vehicle networks (IVNs), supporting heterogeneous traffic such as Audio/Video Transport Protocol (AVTP, IEEE 1722) streams, generalized Precision Time Protocol (gPTP, IEEE 802.1AS) synchronization, and service-oriented communication over UDP. As this transition increases bandwidth and protocol diversity, it also expands the attack surface of connected and autonomous vehicles. Intrusion detection systems (IDSs) for automotive Ethernet must therefore operate under heterogeneous protocol semantics while remaining effective even when labeled attack data are unavailable during deployment. Among recent unsupervised approaches, AERO [1] has demonstrated strong detection performance by engineering three multimodal features (an abstract protocol sequence, raw payloads, and timestamp statistics) and training a lightweight neural outlier detector on benign traffic. Such feature engineering, however, introduces two limitations. First, handcrafted features may restrict generalizability across protocol variants and traffic conditions. Second, even when detection performance is strong, an outlier score computed on engineered features does not reveal which packet fields triggered an alert, limiting interpretability for security analysts. To address these limitations, we propose X-SPUR, an explainable unsupervised IDS for automotive Ethernet based on causal next-token prediction. X-SPUR treats parsed packet fields as a structured token sequence and learns benign traffic patterns through language modeling. Under this formulation, anomalous traffic manifests as tokens with unexpectedly high cross-entropy, or surprisal, relative to the benign distribution learned by the model. Because anomaly evidence is computed at the token level, it can be traced back to individual protocol fields, providing an intrinsic basis for explainability. X-SPUR is built on a Mamba2 state-space language model, whose sequence-processing cost grows linearly with input length. Raw protocol fields extracted with tshark are converted into text-like packet representations and tokenized using byte-level byte-pair encoding (BBPE), preserving fine-grained protocol content without requiring manual feature design. To incorporate temporal information, we introduce a bimodal fusion mechanism that combines payload-token embeddings with inter-packet timing embeddings through additive fusion and an element-wise Hadamard interaction. A central challenge is that benign score distributions differ sharply across protocol families. This makes a global threshold on raw scores difficult to apply across heterogeneous

2

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

traffic. X-SPUR therefore computes window-level anomaly scores from the top-k% token surprisal values, smooths them within each flow, and normalizes them with per-protocol Zscore calibration. Different top-k granularities favor different anomaly morphologies. The final hybrid score therefore fuses two calibrated granularities (k = 5% and 3%). We evaluate X-SPUR on the TOW-IDS automotive Ethernet dataset. It achieves an area under the receiver operating characteristic curve (AUC) of 0.9987. This is marginally higher than the 0.9969 reported for AERO [1]. X-SPUR operates on raw tokenized packets and also provides per-field anomaly attribution. We separately train a CarDS model using the same architecture and training hyperparameters. It achieves an AUC of 0.9923, whereas a faithful AERO reimplementation achieves an AUC of 0.6239. The results also show that protocol-aware calibration suppresses protocol-specific false positives. Token-level surprisal further yields meaningful attack signatures across different intrusion types. The main contributions of this article are threefold: • We propose X-SPUR, an explainable unsupervised automotive Ethernet IDS that replaces handcrafted multimodal feature engineering with end-to-end causal nexttoken prediction over tokenized packet fields, realized by a bimodal Mamba2 architecture that fuses payload content and inter-packet timing through additive fusion and a Hadamard interaction. On the TOW-IDS dataset, X-SPUR achieves an AUC of 0.9987. AERO [1] reports 0.9969. • We propose a dual top-k% per-protocol Z-score calibration strategy that accounts for heterogeneous protocol score distributions and differing attack morphologies. We apply the same model architecture and training hyperparameters to a separately trained CarDS model. We do not conduct a CarDS-specific search over those training hyperparameters. • We show that per-token cross-entropy surprisal provides intrinsic, field-level explainability, exposing distinct attack signatures and tracing both detections and residual false positives to the same protocol fields. The remainder of this article is organized as follows: §II reviews related work; §III describes the detection task, dataset, and packet-to-token pipeline; §IV presents the X-SPUR framework; §V reports the experimental setup, results, and analysis; and §VI concludes this article. II. B ACKGROUND A. Intrusion Detection in Automotive Ethernet IVN security research has traditionally targeted the CAN bus [2], [3]. However, the bandwidth demands of connected and autonomous vehicles have established automotive Ethernet as the de facto standard for next-generation IVNs, shifting the defensive problem from a single broadcast bus to switched, multi-protocol traffic that CAN-specific detectors cannot cover. One line of automotive Ethernet IDSs relies on explicit protocol knowledge or engineered features. Rule-based designs pair protocol-conformance checks with gated recurrent unit (GRU) analysis or decision trees [4], [5]. Weighted-histogram

statistics enable lightweight per-packet screening [6]. Featurebased detectors operate on handcrafted Scalable serviceOriented MiddlewarE over IP (SOME/IP), AVTP, or wavelet features [7]–[9] (Table I). A second line employs deep learning (DL) end to end. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) learn directly from stacked payload bytes or sequential SOME/IP traffic [10]–[12]. Supervised training presumes labeled attack traffic that zero-day threats by definition lack. Unsupervised and semi-supervised alternatives include reconstruction-based autoencoders and sequence-to-sequence long short-term memory (LSTM) models [13]–[15]. Among these, AERO [1] attains real-time unsupervised detection with a lightweight neural outlier detector over three multimodal features. Table I summarizes these studies. Two gaps persist across them. First, detection quality still hinges on engineered inputs: even the strongest unsupervised design depends on handcrafted multimodal features, which can tie a detector to the protocol mix it was engineered for. Second, intrinsic explainability remains confined to shallow rule-based designs: DL-based detectors localize anomalies no finer than the packet level and leave their decisions opaque to analysts. These gaps motivate the two pillars of X-SPUR—end-to-end causal modeling of raw packet tokens on a linear-time sequence backbone, and anomaly evidence attributable to individual protocol fields by construction. B. Sequence Modeling: From Transformers to State-Space Models Transformer-based architectures have achieved state-of-theart performance in network IDSs by capturing long-range dependencies through self-attention [17]. Dense self-attention has O(L2 ) computational complexity in the sequence length L. This growth can increase the computational burden of long traffic contexts on automotive electronic control units (ECUs). State-space models (SSMs), most prominently Mamba [18] and its successor Mamba2 [19], alleviate this bottleneck by replacing attention with a selective state-space recurrence that runs in O(L) time. In network security, Mamba-based detectors model long traffic sequences with Transformer-level accuracy at markedly lower computational cost [20], and NetMamba [21] showed that a unidirectional Mamba encoder suffices for efficient traffic classification. These results, however, concern conventional IP networks; to our knowledge, SSMs remain largely unexplored for automotive Ethernet. X-SPUR closes this gap with a Mamba2based causal language model as its detection backbone. C. Explainability in IVN Security To mitigate the opacity of DL-based IDSs, recent studies attach post-hoc explainable artificial intelligence (XAI) methods such as SHAP and LIME [22]. An operator who cannot see which inputs drove an alert has little basis to trust it, debug detector mistakes, or reduce false positives [23], and such post-hoc explanations often resolve only to whole flows or graph edges rather than individual fields [24]. These methods

X-SPUR: EXPLAINABLE SURPRISAL-BASED REASONING FOR AUTOMOTIVE ETHERNET IDS

3

TABLE I C OMPARISON OF REPRESENTATIVE STUDIES ON AUTOMOTIVE E THERNET IDS WITH OUR STUDY. S EQUENCE MODELING : ✓ EXPLICIT SEQUENCE MODEL , ● PARTIAL ( TEMPORAL CONTEXT ONLY, e.g., WINDOWING OR PAYLOAD STACKING ), ✗ NONE . Study

Input / Modality

Method

Paradigm

[4] (2023) [6] (2025)

SOME/IP AVTP-byte

Rule-based / GRU layer Weighted histogram

[5] (2023) [7] (2026) [8] (2022) [9] (2023)

SOME/IP SOME/IP AVTP-byte AVTP, gPTP, CAN/UDP AVTP, gPTP, UDP, SOME/IP AVTP-byte SOME/IP AVTP-byte

Decision-tree XGBoost XGBoost Wavelet + Deep CNN

[12] (2025) [10] (2021) [11] (2021) [13] (2022) [14] (2024) [1] (2024) [15] (2025) [16] (2024) Ours

Packet fields AVTP, gPTP, CAN/UDP (multimodal) AVTP, gPTP, CAN/UDP AVTP, gPTP, CAN/UDP Packet-field tokens

Sequence Modeling

Localization Granularity

Explainability

Rule-based / Supervised Unsupervised / Statistical Supervised Supervised Supervised Supervised

✓ ●

Sequence level Packet level

Limited Limited

✗ ● ● ●

Field level Packet level Window level Window level

Intrinsic Limited None None

Supervised

✓

Window level

None

CNN, Residual attention 2D-CNN RNN Convolutional Autoencoder Autoencoder Neural outlier detector

Supervised Supervised Unsupervised

● ✓ ●

Packet level Sequence level Window level

None Limited None

Semi-supervised Unsupervised

● ●

Window level Window level

None None

Seq2seq / LSTM

Unsupervised

✓

Sequence level

None

Random Forest / Pruned CNN Mamba-based

Supervised

●

Window level

None

Unsupervised

✓

Field level

Intrinsic

probe an already-trained detector, and the added per-alert computation can bottleneck real-time in-vehicle deployment. Foundation models for traffic understanding (e.g., Lens [25]) face a parallel barrier: full-scale generative inference exceeds the memory and latency budgets of automotive ECUs. X-SPUR instead makes explanation intrinsic to detection: because its causal model scores every token by surprisal against learned benign behavior, the quantity that raises an alert also localizes it to a specific protocol field and timestep, without auxiliary explainer modules. Deriving explanations from the detector’s own anomaly evidence is established in general network and log anomaly detection [26]–[28]. Within IVNs, X-CANIDS ranks CAN signals by per-signal reconstruction error [3]. X-SPUR extends this principle to protocolfield attribution in automotive Ethernet. III. P ROBLEM S ETUP AND DATA R EPRESENTATION

TABLE II ATTACK TYPES IN TOW-IDS Attack Name CAN DoS CAN Replay AVTP Frame Injection MAC Flooding PTP Sync Injection

# Windows Description 126,651 Flood of CAN-over-UDP packets 161,370 Replayed CAN-over-UDP packets 61,894 Injected AVTP video frames 9,000 Random MAC address flooding 90,006 Spoofed gPTP synchronization

extracted from a flow. For evaluation, a test window is labeled anomalous if any of its packets is an attack packet, and its attack type is inherited from the packet-level annotation. Protocol-wise calibration statistics and candidate thresholds are computed from benign validation data (sections IV-E and IV-F). For benchmark reporting, we use test labels to select the best-F1 operating percentile. This selection does not update model weights or calibration statistics.

A. Detection Task and Threat Setting We consider unsupervised intrusion detection on automotive Ethernet under a benign-only setting. The training and validation splits contain only benign traffic. The test split mixes benign traffic with multiple attack types. Here, unsupervised means that no attack samples or attack labels are used to optimize the model. This is also a one-class anomaly-detection setting. Unlike the semi-supervised IDS in [29], X-SPUR does not train on labeled attacks. X-SPUR learns through next-token prediction (6). It differs from AERO [1] in its scoring mechanism, not its supervision regime. AERO scores distance to a learned prototype. X-SPUR instead uses tokenlevel surprisal. Detection is performed at the window level, and each sample is a sliding window of at most w = 10 consecutive packets

B. Datasets and Packet-to-Text Representation TOW-IDS [9] comprises two packet-dump captures from an automotive Ethernet testbed plus an extended normaldriving capture. We follow the dataset split strategy discussed in AERO [1] (cf. Table II therein), with the addition of a large benign capture containing ≈7M packets (the TrainHuge split) for pretraining. The test split contains five attack scenarios (CAN DoS, CAN Replay, AVTP Frame Injection, MAC Flooding, and PTP Sync Injection), summarized in Table II. The benign traffic spans three protocol families with distinct statistical characteristics—UDP (CAN-over-UDP), AVTP (IEEE 1722 audio/video), and gPTP (IEEE 802.1AS)—whose markedly different score distributions later motivate protocol-

4

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

TABLE III C AR DS [30] AUTOMOTIVE E THERNET DATASET COMPOSITION USED FOR CROSS - DATASET EVALUATION . Split

Class

# Traces

# Packets

Train-Huge Benign Train Benign Validation Benign

38 27 6

14,276,579 7,239,662 4,120,212

5 18 5 2 5 1

2,916,705 6,002,445 1,774,426 543,372 2,762,142 133,813

Test

Benign Nmap SYN scan (-sS) Nmap ACK scan (-sA) Nmap FIN scan (-sF) RTP camera replacement IVI REST API crash

aware score calibration. Minor protocol types (ARP, 0x88f5, and SRP) are excluded owing to their limited sample sizes. To assess the transferability of the proposed methodology, we further evaluate X-SPUR on the automotive Ethernet (AE) portion of CarDS [30]. CarDS was captured from a 2020 commercial electric vehicle and contains 181M AE messages over 258 traces. Unlike TOW-IDS, its AE network is IPv6-based and VLAN-segmented. It carries general-purpose IP protocols such as TCP, UDP, ICMPv6, the Real-time Transport Protocol (RTP), and HTTP. We use the same model architecture and training procedure for its AE traffic. Benign traces form the train and validation splits. The test set contains AE pentesting traces, including Nmap scans with MAC/IP/VLAN spoofing, rear-view-camera RTP replacement, and an in-vehicle infotainment (IVI) REST-API crash (Table III). The preprocessing pipeline (Fig. 1) separates raw PCAP files into per-flow captures keyed by source/destination MAC address, source/destination IP address, and transport-layer ports. Each flow is then parsed with tshark into a text representation spanning up to 180 protocol fields across the link, VLAN, and network/transport layers, together with protocolspecific fields from the IEEE 1722 (AVTP), IEC 61883, MPEG-TS, H.264, MPEG-PES, and PTP/gPTP layers. C. BBPE Tokenization and Window Construction To convert packet text into model input tokens, we train a BBPE tokenizer with vocabulary size |V| = 16,000 on the Train-Huge split. The vocabulary reserves special tokens for padding (<pad>), unknown bytes, source/destination IP markers, packet boundaries (<sep>), and end of sequence. The resulting subword tokens preserve protocol-specific patterns while remaining flexible across heterogeneous field values. For example, the string eth:ethertype:ip:udp:data is tokenized as [eth, :, ethertype, :, ip, :, udp, :, data]. Tokenized flows are then segmented into sliding windows (w = 10 packets, stride 1), and the packet token sequences within a window are concatenated with <sep> tokens at packet boundaries. Sequences are truncated to their first Lmax tokens and padded if shorter.

inter-packet time deltas ∆ti = ti − ti−1 and log-scaled as si = log10 (∆ti + 10−7 ).

(1)

The values are clipped to [−7, 7] and stored using 1,000 uniform quantization bins. Their bin midpoints are fed as continuous scalars to the time-projection multilayer perceptron (MLP) in (5), rather than through a learned per-bin embedding. Finally, each test window is assigned its binary label by mapping window indices back to the per-packet annotations, following the rule in §III-A. IV. P ROPOSED X-SPUR F RAMEWORK As illustrated in Fig. 2, the X-SPUR framework operates in three stages. During training, a bimodal causal language model learns to predict the next token in benign traffic. During calibration, validation-set scores are computed at two top-k levels, smoothed within each flow, and normalized per protocol. During detection, a test window is scored by token-level cross-entropy, converted into top-k window scores, smoothed, calibrated, and thresholded. The same cross-entropy values provide field-level attribution. A. Mamba2-Based Causal Language Model The backbone of X-SPUR is a causal language model built on the Mamba2 state-space architecture [19]. Its O(L) sequence processing limits the growth in computational cost as the input context increases (§II). The base architecture embeds each token and its position through a token embedding etok : V → Rd and a learnable absolute positional embedding epos , applies dropout to their sum, and processes the sequence with a stack of N pre-norm residual Mamba2 blocks,    x(ℓ+1) = x(ℓ) + Mamba2 LN x(ℓ) , (2) followed by a final layer normalization and a language-model head whose output weights are tied to the token embedding matrix. Each Mamba2 block can be viewed as a selective statespace model: given an input sequence x(t), the hidden state evolves as h′ (t) = A h(t) + B(t) x(t),

y(t) = C(t) h(t),

(3)

where A is a structured (diagonal) state-transition matrix and the projections B(t) and C(t) are input-dependent. This selectivity lets the model modulate its state by the current token context, which suits detecting unusual protocol-field transitions. For discrete sequences, the continuous model is discretized with a learned step size ∆, yielding recurrent updates computed efficiently by a parallel scan. B. Bimodal Fusion of Payload and Inter-Packet Time

D. Time Representation and Window Labels In addition to packet content, each window carries aligned inter-packet timing. Raw packet timestamps are converted into

To incorporate temporal context beyond packet content, XSPUR extends the base model with a bimodal fusion module that combines token, positional, and time embeddings through

X-SPUR: EXPLAINABLE SURPRISAL-BASED REASONING FOR AUTOMOTIVE ETHERNET IDS

❶ Flow Separation

Flow 1

[gPTP example] eth:ethertype:ptp

60

<latexit sha1_base64="ClQUR/NWpxbtsD4oIwV6YWCPKow=">AAAB83icbVDLSsNAFL2pjz58VV3qYrAIrkrionVn0Y3LCvYBTSiT6aQdOpmEmYlQQsGvcONCEbdu/BR3foZ/4KTtQlsPDBzOuZd75vgxZ0rb9peVW1vf2MwXiqWt7Z3dvfL+QVtFiSS0RSIeya6PFeVM0JZmmtNuLCkOfU47/vg68zv3VCoWiTs9iakX4qFgASNYG8l1Q6xHfpD6077TL1fsqj0DWiXOglQuv93iw8dVvtkvf7qDiCQhFZpwrFTPsWPtpVhqRjidltxE0RiTMR7SnqECh1R56SzzFJ0aZYCCSJonNJqpvzdSHCo1CX0zmWVUy14m/uf1Eh1ceCkTcaKpIPNDQcKRjlBWABowSYnmE0MwkcxkRWSEJSba1FQyJTjLX14l7fOqU6vWbp1K4xjmKMARnMAZOFCHBtxAE1pAIIZHeIYXK7GerFfrbT6asxY7h/AH1vsPIPuUnA==</latexit>

60

dc:a6:32:5e:48:47

01:80:c2:00:00:0e

[AVTP example] eth:ethertype:ieee1722:ieee17221

Flow 2

❺ Time Bin Construction

❷ Tshark Feature Extraction [UDP example] eth:ethertype:ip:udp:data

Raw PCAP

5

82

dc:a6:32:5d:ce:d2

00:fc:70:00:00:02

91:e0:f0:01:00:00

0x88f7

0x0800

192.168.10.2 …

0x00fc70fffe000002

00:fc:70:00:00:02

0x22f0

b1 t1 t2 t3 t4 t5

Binning

…

0xfa

Log-scaling

…

Packet timestamps !tj = tj → tj→1 <latexit sha1_base64="e+8mIIdywuu2wy49FlczG6+B9Z8=">AAACA3icbVC7SgNBFJ2Nj8T4itppMxgFm4Rdi2gjBEyhlQmYByTLMjuZJJPMPpi5K4QlYON/WNlYKGLrT9jZ+yFOHoUmHpjL4Zx7uXOPGwquwDS/jMTS8spqMrWWXt/Y3NrO7OzWVBBJyqo0EIFsuEQxwX1WBQ6CNULJiOcKVncHl2O/fsek4oF/C8OQ2R7p+rzDKQEtOZn9VokJIBicPr6Y1JyucT9njZxM1sybE+BFYs1Itpi8/r55rJTKTuaz1Q5o5DEfqCBKNS0zBDsmEjgVbJRuRYqFhA5IlzU19YnHlB1PbhjhY620cSeQ+vmAJ+rviZh4Sg09V3d6BHpq3huL/3nNCDrndsz9MALm0+miTiQwBHgcCG5zySiIoSaESq7/immPSEJBx5bWIVjzJy+S2mneKuQLFStbPEJTpNABOkQnyEJnqIiuUBlVEUX36Am9oFfjwXg23oz3aWvCmM3soT8wPn4AVfyY7g==</latexit>

…

…

BBPE merges

Flow N

❸ BBPE Tokenization

Same row key

<pad> id=0 <sep> id=4

<latexit sha1_base64="+X/4Qo89bjnokYchP5IJcjOAfts=">AAACDnicbVBNSwMxEM1q1Vq/Vj16CdZCBSm7HqrHghePFewHtMuSTbNtaDZZkqxSlv4CL/4VLx4U8erZm//G7LaCtj4IefNmhpl5Qcyo0o7zZa2sFtbWN4qbpa3tnd09e/+grUQiMWlhwYTsBkgRRjlpaaoZ6caSoChgpBOMr7J8545IRQW/1ZOYeBEachpSjLSRfLsy9mk/QrHSomp+PQrC9H7q07OfIDDBqW+XnZqTAy4Td07KjUKYo+nbn/2BwElEuMYMKdVznVh7KZKaYkampX6iSIzwGA1Jz1COIqK8ND9nCitGGcBQSPO4hrn6uyNFkVKTKDCV2ZJqMZeJ/+V6iQ4vvZTyONGE49mgMGFQC5h5AwdUEqzZxBCEJTW7QjxCEmFtHCwZE9zFk5dJ+7zm1mv1G7fcOAEzFMEROAZV4IIL0ADXoAlaAIMH8ARewKv1aD1bb9b7rHTFmvccgj+wPr4Bj7mfRA==</latexit>

ki →↑ (wi , bi )

Token sequence

eth + er → ether ty + pe → type ether + type → ethertype 08 + 00 → 0800 …

eth

:

ethertype

:

❹ Window Generation

ieee 1722 … <latexit sha1_base64="lSFtnTs9QEpaKdgQBsZQwU/WvN0=">AAAB83icbVDLSsNAFL2pjz58VV3qYrAIrkriorqz6MZlBfuAJpTJdNIOnUzCzEQpoeBXuHGhiFs3foo7P8M/cNJ2oa0HBg7n3Ms9c/yYM6Vt+8vKrayurecLxdLG5tb2Tnl3r6WiRBLaJBGPZMfHinImaFMzzWknlhSHPqdtf3SV+e07KhWLxK0ex9QL8UCwgBGsjeS6IdZDP0jvJz2nV67YVXsKtEycOalcfLvFh4/LfKNX/nT7EUlCKjThWKmuY8faS7HUjHA6KbmJojEmIzygXUMFDqny0mnmCTo2Sh8FkTRPaDRVf2+kOFRqHPpmMsuoFr1M/M/rJjo491Im4kRTQWaHgoQjHaGsANRnkhLNx4ZgIpnJisgQS0y0qalkSnAWv7xMWqdVp1at3TiV+iHMUIADOIITcOAM6nANDWgCgRge4RlerMR6sl6tt9lozprv7MMfWO8/QQ6UsQ==</latexit>

148 119

168

119

229

228

… <latexit sha1_base64="K+oeIOSjNaa6fOtUQ4j0eKTZfMQ=">AAAB83icbVDLSsNAFL3x1YevqktdDBbBVUm6qO4sunFZwT6gCWUynbRDJ5MwM1FKKPgVblwo4taNn+LOz/APnLRdaOuBgcM593LPHD/mTGnb/rJWVtfWN3L5QnFza3tnt7S331JRIgltkohHsuNjRTkTtKmZ5rQTS4pDn9O2P7rK/PYdlYpF4laPY+qFeCBYwAjWRnLdEOuhH6T3k161VyrbFXsKtEycOSlffLuFh4/LXKNX+nT7EUlCKjThWKmuY8faS7HUjHA6KbqJojEmIzygXUMFDqny0mnmCToxSh8FkTRPaDRVf2+kOFRqHPpmMsuoFr1M/M/rJjo491Im4kRTQWaHgoQjHaGsANRnkhLNx4ZgIpnJisgQS0y0qaloSnAWv7xMWtWKU6vUbpxy/QhmyMMhHMMpOHAGdbiGBjSBQAyP8AwvVmI9Wa/W22x0xZrvHMAfWO8/QpKUsg==</latexit>

w1 P1 <sep> P2 <sep> P3 <sep> P4 <sep> P5 w2 P2 <sep> P3 <sep> P4 <sep> P5 <sep> P6

Fig. 1. Overview of the X-SPUR preprocessing pipeline. Raw PCAP traffic is separated into flows, parsed with tshark, converted to BBPE token sequences, grouped into sliding windows, and aligned with log-scaled inter-packet time features.

Per-flow smoothing (K=63)

LM Head

Top-k% Mamba2 Block

htok

htime

Threshold Determinator

Hybrid scoring

Explainability Analysis

Token-level Surprisal

ti+1 ti+2 ti+3 ti+4

…

CE Loss 0.42 0.31 2.87 3.95

…

Token <latexit sha1_base64="AeYe9tUWlAcTnqF2WGWJEr76CSI=">AAAB73icbVDLSgNBEJz1GeMr6tHLYBA8hV2R6DHgxWME84BkCbOTTjJkZnad6RXCkp/w4kERr/6ON//GSbIHTSxoKKq66e6KEiks+v63t7a+sbm1Xdgp7u7tHxyWjo6bNk4NhwaPZWzaEbMghYYGCpTQTgwwFUloRePbmd96AmNFrB9wkkCo2FCLgeAMndQe9TIUCqa9Utmv+HPQVRLkpExy1Hulr24/5qkCjVwyazuBn2CYMYOCS5gWu6mFhPExG0LHUc0U2DCb3zul507p00FsXGmkc/X3RMaUtRMVuU7FcGSXvZn4n9dJcXATZkInKYLmi0WDVFKM6ex52hcGOMqJI4wb4W6lfMQM4+giKroQguWXV0nzshJUK9X7q3LNz+MokFNyRi5IQK5JjdyROmkQTiR5Jq/kzXv0Xrx372PRuublMyfkD7zPH2+8kDE=</latexit>

Per-protocol |Z|-score calibration

5%

Additive + Hadamard Fusion

<latexit sha1_base64="ZCJJKqKLg851YnuR5rZ9ewAlLfA=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseCF48V7Ae0oWy2m3bJZhN2J0IJ/RFePCji1d/jzX/jts1BWx8MPN6bYWZekEph0HW/ndLG5tb2Tnm3srd/cHhUPT7pmCTTjLdZIhPdC6jhUijeRoGS91LNaRxI3g2iu7nffeLaiEQ94jTlfkzHSoSCUbRSdzLMMYlmw2rNrbsLkHXiFaQGBVrD6tdglLAs5gqZpMb0PTdFP6caBZN8VhlkhqeURXTM+5YqGnPj54tzZ+TCKiMSJtqWQrJQf0/kNDZmGge2M6Y4MaveXPzP62cY3vq5UGmGXLHlojCTBBMy/52MhOYM5dQSyrSwtxI2oZoytAlVbAje6svrpHNV9xr1xsN1rekWcZThDM7hEjy4gSbcQwvawCCCZ3iFNyd1Xpx352PZWnKKmVP4A+fzB7TSj8Y=</latexit>

Scoring & Calibration

3%

Bi-Modal Fusion Model

<latexit sha1_base64="2Sp3tqwPzNoYr/CcMUVR9JksHIA=">AAAB7nicbZC7SgNBFIbPxluMGqOWNoNBEISwaxHFKmBjGcFcIAlhdjKbDJmdXWbOCmFJ4SPYWChim0fwOezsfRAnl0ITfxj4+P9zmHOOH0th0HW/nMza+sbmVnY7t7O7l98vHBzWTZRoxmsskpFu+tRwKRSvoUDJm7HmNPQlb/jDm2neeODaiEjd4yjmnZD2lQgEo2itBnZTce6Nu4WiW3JnIqvgLaBYyT9+BN/Xk2q38NnuRSwJuUImqTEtz42xk1KNgkk+zrUTw2PKhrTPWxYVDbnppLNxx+TUOj0SRNo+hWTm/u5IaWjMKPRtZUhxYJazqflf1kowuOqkQsUJcsXmHwWJJBiR6e6kJzRnKEcWKNPCzkrYgGrK0F4oZ4/gLa+8CvWLklcule+8YsWFubJwDCdwBh5cQgVuoQo1YDCEJ3iBVyd2np03531emnEWPUfwR87kB+4rkvk=</latexit>

<latexit sha1_base64="v7FbpkKDxp6JG8F+k2VBO7SI0i0=">AAAB7nicbZC7SgNBFIbPxluMt6ilzWAQBCHspohiFbCxjGAukIQwO5lNhszOLjNnhbCk8BFsLBSxzSP4HHb2PoiTS6GJPwx8/P85zDnHj6Uw6LpfTmZtfWNzK7ud29nd2z/IHx7VTZRoxmsskpFu+tRwKRSvoUDJm7HmNPQlb/jDm2neeODaiEjd4yjmnZD2lQgEo2itBnZTcVEad/MFt+jORFbBW0Chsv/4EXxfT6rd/Ge7F7Ek5AqZpMa0PDfGTko1Cib5ONdODI8pG9I+b1lUNOSmk87GHZMz6/RIEGn7FJKZ+7sjpaExo9C3lSHFgVnOpuZ/WSvB4KqTChUnyBWbfxQkkmBEpruTntCcoRxZoEwLOythA6opQ3uhnD2Ct7zyKtRLRa9cLN95hYoLc2XhBE7hHDy4hArcQhVqwGAIT/ACr07sPDtvzvu8NOMseo7hj5zJD++wkvo=</latexit>

<latexit sha1_base64="2Suu9kkcrvaJlviPC3cp4lD/JkY=">AAAB7nicbZC7SgNBFIbPeo1RY9TSZjAIghB2FaJYBWwsI5gLJCHMTmaTIbOzy8xZISwpfAQbC0Vs8wg+h529D+LkUmjiDwMf/38Oc87xYykMuu6Xs7K6tr6xmdnKbu/s5vby+wc1EyWa8SqLZKQbPjVcCsWrKFDyRqw5DX3J6/7gZpLXH7g2IlL3OIx5O6Q9JQLBKFqrjp1UnF2MOvmCW3SnIsvgzaFQzj1+BN/X40on/9nqRiwJuUImqTFNz42xnVKNgkk+yrYSw2PKBrTHmxYVDblpp9NxR+TEOl0SRNo+hWTq/u5IaWjMMPRtZUixbxaziflf1kwwuGqnQsUJcsVmHwWJJBiRye6kKzRnKIcWKNPCzkpYn2rK0F4oa4/gLa68DLXzolcqlu68QtmFmTJwBMdwCh5cQhluoQJVYDCAJ3iBVyd2np03531WuuLMew7hj5zxD/E1kvs=</latexit>

<latexit sha1_base64="moZeF8xEeZE+iZpLnaab7e/wm9U=">AAAB7nicbZC7SgNBFIbPeo1RY9TSZjAIghB2RaJYBWwsI5gLJCHMTmaTIbOzy8xZISwpfAQbC0Vs8wg+h529D+LkUmjiDwMf/38Oc87xYykMuu6Xs7K6tr6xmdnKbu/s5vby+wc1EyWa8SqLZKQbPjVcCsWrKFDyRqw5DX3J6/7gZpLXH7g2IlL3OIx5O6Q9JQLBKFqrjp1UnF2MOvmCW3SnIsvgzaFQzj1+BN/X40on/9nqRiwJuUImqTFNz42xnVKNgkk+yrYSw2PKBrTHmxYVDblpp9NxR+TEOl0SRNo+hWTq/u5IaWjMMPRtZUixbxaziflf1kwwuGqnQsUJcsVmHwWJJBiRye6kKzRnKIcWKNPCzkpYn2rK0F4oa4/gLa68DLXzolcqlu68QtmFmTJwBMdwCh5cQhluoQJVYDCAJ3iBVyd2np03531WuuLMew7hj5zxD/K6kvw=</latexit>

BBPE offset mapper

Field-level attribution

Detailed in Fig. 4

Fig. 2. Overall framework of X-SPUR, from preprocessing and bimodal causal modeling to calibration and explainability.

both additive fusion and an element-wise Hadamard interaction: xi = etok (xi ) + epos (i) + etime (si ) (4) + α (etok+pos ⊙ etime (si )) , where α = 1.0 is a fixed interaction strength, ⊙ denotes element-wise multiplication, etime (si ) ∈ Rd is a learned embedding of the log-scaled inter-packet delta si from (1), and etok+pos abbreviates the sum etok (xi ) + epos (i). The time embedding itself is produced by a two-layer MLP: etime (s) = W2 GELU(LN(W1 s + b1 )) + b2 ,

(5)

where W1 ∈ Rdproj ×1 , W2 ∈ Rd×dproj , dproj = 64, and dropout with p = 0.1 is applied between the two linear layers. Because BBPE yields a variable number of subword tokens per packet while time deltas are per packet, sj is replicated across all token positions of packet j. The model uses dimension d = 768 with N = 24 layers and maximum sequence length Lmax = 256. We use head dimension 128, state dimension dstate = 16, convolution width dconv = 4, and dropout 0.1. The model has ≈98.5M parameters.

C. Benign-Only Training Objective On TOW-IDS, training is benign-only and two-phase: pretraining on the Train-Huge split (≈7M benign windows), then fine-tuning on the smaller Train split (≈228k windows) to stabilize learned patterns and match the target distribution. We pretrain for 10 epochs and fine-tune for 3 epochs with batch size 32. Both phases use the same causal language-modeling objective. For a model input window X = (x1:L , s1:L ), comprising a token sequence and its aligned time features, let IX = {j ∈ {2, . . . , L} : xj ̸= <pad>} denote the valid targettoken positions. The training loss is the standard next-token cross-entropy: LCLM = −

1 X log P (xj | x<j , s<j ; Θ), |IX |

(6)

j∈IX

where Θ denotes the model parameters and <pad> positions are excluded from the sum. Algorithm 1 summarizes the offline training and calibration procedure.

6

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

D. Window-Level Anomaly Scoring via Token Surprisal Algorithm 2 summarizes detection and attribution. At inference time, a test window X = (x1:L , s1:L ) is scored using token-level causal language-model cross-entropy. For each j ∈ IX , CEj = − log P (xj | x<j , s<j ; Θ),

(7)

where positions corresponding to <pad> are masked out. The value CEj measures the surprisal of target token xj under the benign model. Unusually high values indicate deviations from learned protocol behavior. Rather than averaging cross-entropy across the full window, X-SPUR computes a top-k% anomaly score: 1 X CEj , (8) Sk (X) = |Tk | j∈Tk

where Tk contains the max{1, ⌊k|IX |⌋} positions in IX with the highest cross-entropy values. Many attacks affect only a small subset of packet fields, producing sparse surprisal spikes that whole-window averaging would dilute. Top-k scoring preserves their contribution. We use two settings, k = 5% and k = 3%.

Since neighboring windows within the same flow overlap in nine out of ten packets, their raw anomaly scores are highly correlated. To reduce transient noise, we apply moving-average smoothing to both top-k scores within each flow, using kernel size K = 63 (selected in §V): 1 K

i+⌊K/2⌋

X

Sj .

where z0.05 and z0.03 denote the calibrated top-5% and top3% scores, respectively. Whichever granularity produces the more extreme deviation drives detection. F. Decision Rule and Threshold Selection For a fixed operating percentile ρ, the decision threshold τ is computed from the calibrated benign-only validation scores. A window is flagged as anomalous if and only if Shybrid ≥ τ . In practice, a domain expert should select the operating percentile based on the deployment’s tolerance for false alarms. Following the evaluation protocol of AERO [1], we sweep ρ from the 90th to the 99.99th percentile. For benchmark reporting, we select the percentile with the highest test-set F1score. The best-F1 operating point is attained at the 99.94th percentile (τ = 3.9122 on the hybrid score). It yields a test precision of 0.9785 and recall of 0.9775. This reporting choice does not update model weights or calibration statistics. G. Intrinsic Explainability via Field Attribution

E. Flow-Level Smoothing and Protocol-Aware Calibration

S̃i =

sparse-spike attacks whose evidence concentrates in a few fields. The final hybrid calibrated score is therefore defined as Shybrid = max (|z0.05 |, |z0.03 |) , (11)

(9)

j=i−⌊K/2⌋

At flow boundaries, the nearest endpoint score is replicated. Flows with at most K windows are left unsmoothed. Because benign score levels differ sharply across protocol families (gPTP windows sit far above UDP), X-SPUR normalizes per protocol. For each protocol p, identified from the flow’s frame.protocols field, and each top-k fraction k, the mean µp,k and standard deviation σp,k of the smoothed validation scores are computed. A test score is then normalized as S̃k − µp,k zk = , (10) max(σp,k , ϵ) where ϵ = 10−6 ensures numerical stability. If protocol p is absent from the benign validation split, the corresponding global validation mean and standard deviation are used as a fallback. On CarDS, this fallback also applies to protocols with fewer than 200 benign validation windows. Since both unusually high and unusually low scores may indicate abnormal behavior, the absolute value |zk | enters the final calibration stage. Empirically, the two granularities are complementary: the top-5% view is essential for PTP Sync Injection—its moderately distributed evidence all but vanishes from the narrower top-3% token set—whereas the top-3% view slightly sharpens

Because the window-level score is derived from token-level cross-entropy, field-level attribution requires no additional post-hoc explanation model. It proceeds directly from the scored window. The BBPE tokenizer’s offset mapping locates each token’s character span in the original packet text. The span is mapped to its tab-separated field index and field name in the tshark schema. Token-level cross-entropy values are then aggregated within each field: 1 X CEj , (12) CEf = |Tf | j∈Tf

where Tf is the set of valid target-token positions j whose token xj belongs to field f . This token-to-field mapping traces anomaly scores back to concrete protocol fields and reveals distinct attack signatures localized to the specific fields each intrusion disturbs. Two properties delimit what such an explanation means. First, it is directly grounded in the model score. Equation (12) aggregates field-associated values from the same token-level cross-entropies used to form the window score. These values precede the top-k pooling, smoothing, and calibration of §IVD and §IV-E. Unlike a separately fitted surrogate, it therefore reports local model evidence rather than an approximation from another model. Second, surprisal marks deviation from learned benign regularity, not maliciousness itself: attribution answers which fields deviate, and §V-G shows the same readout diagnosing residual false positives. V. E XPERIMENTAL R ESULTS A. Experimental Setup Unless stated otherwise, experiments use benign-only pretraining and fine-tuning. They use the TOW-IDS test split and

X-SPUR: EXPLAINABLE SURPRISAL-BASED REASONING FOR AUTOMOTIVE ETHERNET IDS

7

Algorithm 1: Offline Training and Calibration

Algorithm 2: Detection and Attribution

Input: Benign PCAP splits Phuge , Ptr , and Pv ; fixed tshark field schema Top-k fractions R = {0.03, 0.05}; fixed operating percentile ρ Output: Frozen BBPE tokenizer B; trained parameters Θ Statistics {(µp,k , σp,k )}p,k and global fallback {(µ∗,k , σ∗,k )}k Threshold τ

Input: Raw test flow F of protocol p; frozen B and field schema Trained Θ; calibration and fallback statistics from Algorithm 1; R = {0.03, 0.05}; threshold τ Output: Window decisions yi ∈ {0, 1}, one per resulting window For every yi = 1, a ranked field attribution 1 Apply the frozen preprocessing of Algorithm 1 to F , retaining text and BBPE offsets → (X1 , . . . , Xn ). 2 for i ← 1 to n do 3 Forward Xi once; retain {CEi,j }j (7). 4 for k ∈ R do 5 Compute Si,k from the same retained surprisals (8).

// Step 1—Data preparation (§III-C and §III-D)

For each d ∈ {huge, tr, v}, split Pd into ordered flows and serialize the fixed packet fields with tshark → Rd . 2 Train B (|V| = 16,000) on Rhuge only. 3 For each d ∈ {huge, tr, v}, BBPE-encode Rd , insert <sep>, align log-scaled time deltas, and form stride-1 windows (w = 10, Lmax = 256) → Sd . 1

// Step 2—Benign-only model fitting (§IV-C)

Initialize Θ. 5 foreach (S, E) ∈ [(Shuge , 10), (Str , 3)] in order do 6 for epoch ← 1 to E do 7 for mini-batch X ⊂ S do 8 Fuse the token, position, and time embeddings; forward once through the Mamba2 model. 9 Compute LCLM (6) and update Θ with AdamW. 4

// Step 3—Benign validation calibration (§IV-E)

6 7

for k ∈ R do Apply the centered K = 63 smoother within F (9) → {S̃i,k }n i=1 .

For each k ∈ R, set (µ̄k , σ̄k ) to (µp,k , σp,k ) if available, otherwise to (µ∗,k , σ∗,k ). 9 for i ← 1 to n do 10 zi,k ← (S̃i,k − µ̄k )/ max(σ̄k , ϵ) for each k ∈ R. 11 Si,hybrid ← max(|zi,0.03 |, |zi,0.05 |); yi ← 1 if Si,hybrid ≥ τ , else yi ← 0. 12 if yi = 1 then 13 Map the retained surprisals to tshark fields via BBPE offsets; rank fields by CEf (12). 8

for each ordered validation window Xi ∈ Sv do Forward Xi once and retain its valid target-token surprisals {CEi,j }j . 12 for k ∈ R do 13 Compute Si,k from the largest k fraction of {CEi,j }j (8).

10

11

// Step 4—Threshold (§IV-F)

For every Xi ∈ Sv , compute zi,k with its protocol statistics or the global fallback (10) and set Si,hybrid = maxk∈R |zi,k |. 20 τ ← Percentileρ ({Si,hybrid }i ). 19

the same scoring procedure. The ablations in Sections V-B– V-D vary only the settings under study around the default configuration (w = 10, sequence length 256, K = 63, k ∈ {3%, 5%}). We report threshold-independent AUC. For threshold-dependent metrics, we report the best F1-score over the stated sweep of validation-derived percentile thresholds. Precision and recall are reported at that operating point. We also report per-attack false-negative rates (FNRs) and perprotocol false-positive rates (FPRs). B. Context and Smoothing-Width Selection Fig. 3 reports the context sweep over window size w and maximum sequence length. Shorter windows generally give

768

1024

★ best AUC 0.9987

.952

0.96

.957

.961

.961

.972

0.98

0.92

.936

0.94

.933

.944

ROC--AUC

.982

.997

.996

.997

.997

.994

.997

.998

512

.934

(µ∗,k , σ∗,k ) ← global mean and standard deviation of S̃i,k .

★

384

//

18

.998

1.00

14

.991

for k ∈ R do 15 Smooth {Si,k }i separately within each flow (9) → {S̃i,k }i . 16 for each protocol p with sufficient validation windows (§IV-E) do 17 (µp,k , σp,k ) ← mean and standard deviation of S̃i,k for windows of protocol p.

Max sequence length 256

w=5

w = 10

w = 15

w = 20

Window size (packets)

Fig. 3. Context sweep over window size w and maximum sequence length (AUC), grouped by w and shaded by sequence length; the dashed line marks the best AUC and the starred bar (w = 10, sequence length 256) the selected operating point.

higher AUC. Larger windows show no consistent gain and greater variation across sequence lengths. We therefore adopt w = 10 with sequence length 256, the best operating point (AUC = 0.9987) and the default for all subsequent analyses. To select the smoothing width, we sweep K ∈ {7, 15, 31, 63, 127} under this context setting. AUC rises steadily with the smoothing window, from 0.9901 at K = 7 to its maximum 0.9987 at K = 63, then falls to 0.9924 at K = 127. We therefore fix K = 63. Smoothing is decisive for the hardest attack: at the same validation-percentile threshold, unsmoothed scores miss 97% of AVTP Frame Injection

8

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

TABLE IV C ALIBRATION ABLATION ACROSS THE TWO TOP -k VIEWS ; “P ROT.- CAL .” DENOTES PER - PROTOCOL Z- SCORE NORMALIZATION .

TABLE V P ER - ATTACK FALSE - NEGATIVE RATE (FNR) VERSUS AERO [1], WITH MEAN HYBRID ANOMALY SCORE .

Top-k Stage

AUC F1-score

Attack

X-SPUR

AERO

Mean Shybrid

3% 3% 3%

Raw Smooth. Prot.-cal.

0.9838 0.9942 0.9958

0.9296 0.9518 0.9604

5% 5% 5%

Raw Smooth. Prot.-cal.

0.9763 0.9907 0.9986

0.9087 0.9411 0.9749

CAN DoS CAN Replay AVTP Frame Inj. MAC Flooding PTP Sync Inj.

0.00% 0.14% 15.77% 0.00% 0.13%

0.86% 4.25% 2.43% 3.30% 0.04%

8.621 9.607 5.315 21.157 5.495

Hybrid max(|z0.05 |, |z0.03 |) 0.9987

0.9780

F. Field-Level Explainability Analysis windows versus 16% after smoothing, because the weak perwindow evidence of injected frames becomes detectable only when accumulated across a contiguous burst. Run-to-run results. We repeated the TOW-IDS experiment in five additional runs with different random seeds. The mean AUC was 0.9939, with a sample standard deviation of 0.0041. The AUC ranged from 0.9880 to 0.9985. The AUC of 0.9987 in Tables VIII and IX is from the original experiment (seed 42). These results indicate consistently high detection performance across repeated runs.

C. Time-Fusion Ablation To isolate the contribution of inter-packet timing, we retrain the model with the time channel removed (payload only) using the same context settings. Discarding the time modality lowers the AUC from 0.9987 to 0.9962 and the F1-score from 0.9780 to 0.9663, with a consistent drop across precision and recall. This is expected in automotive Ethernet, where replayed, flooded, and spoofed-synchronization attacks perturb inter-packet timing as much as payload—deviations the timeaware embedding surfaces directly but a payload-only model can only infer from content.

D. Calibration Ablation Table IV shows a consistent progression: raw top-k scores separate benign and malicious traffic reasonably well, flowlevel smoothing pushes both views above 0.99 AUC, perprotocol Z-score normalization further stabilizes the distributions, and the hybrid score fuses the two calibrated views for the best operating point (AUC 0.9987).

E. Per-Attack Performance Table V reports the attack-wise FNR of X-SPUR versus AERO [1], together with each attack’s mean hybrid anomaly score ((11)). CAN DoS, CAN Replay, and MAC Flooding are detected essentially perfectly. PTP Sync Injection has well above 99% recall, only marginally behind AERO. AVTP Frame Injection remains hardest by a wide margin. Its evidence is confined to a few IEC 61883 stream fields already unpredictable in benign traffic (§V-F).

Fig. 4 shows X-SPUR’s intrinsic field-level attribution (§IVG) in five panels, (a)–(e), using TOW-IDS test examples. All panels use the model and settings in §V-A, with w = 10 and Lmax = 256. Each panel has columns for the tshark fields. Each token is shaded by its cross-entropy (7) averaged over the sliding windows scoring it. Attribution reuses the token scores computed for detection. Reading a field’s column topto-bottom exposes deviations from learned benign behavior, and the most-discriminative field is highlighted. The panels reveal how differently each attack manifests. Some ignite only one or two fields while the rest stay cold. Others warm nearly every header field at once. Still others surface only as a broken cadence in a single structural field. Because attribution operates on the decoded tshark schema, the panels name the responsible fields directly. The field-level patterns are consistent with the per-attack results in Table V. CAN DoS and CAN Replay show clear field-level changes and near-zero false-negative rates. AVTP Frame Injection shows weaker contrast and the highest falsenegative rate. CAN-over-UDP (Fig. 4(a,b)). Both CAN attacks ignite only transport and payload fields and leave every Ethernet/IP header cold. CAN DoS (a) surprises the model at udp.dstport— the spoofed attacker port 60000, never seen in benign traffic, reaches a cross-entropy near 12—together with the injected payload and udp.srcport. CAN Replay (b) instead lights only payload: the replayed frames reuse valid headers and ports, so the sole anomaly is the stale byte sequence. Fewer than 1% of tokens are extreme, and this tight localization yields near-zero false negatives. AVTP Frame Injection (Fig. 4(c)). The injected frames are structurally valid MPEG-TS content, so the video payload stays cold; the injection instead resets the IEC 61883 stream counters (iec61883.seqnum, iec61883.dbc) and the media timestamp iec61883.spht. Yet spht carries an arbitrary media-clock timestamp already highly surprising in benign traffic, so the injection lifts it only marginally and no field turns cleanly from cold to warm. This absence of a sharp field-level signal is the visible reason AVTP Frame Injection keeps the highest false-negative rate (Table V); detection instead rests on the aggregate window score. PTP Sync Injection (Fig. 4(d)). Benign gPTP alternates Sync and Follow-Up messages of 60 and 90 bytes; the Sync flood collapses this cadence into a run of 60-byte Syncs, so frame.len ignites, while the sequence-id counter stays high in both benign and attack and is not the discriminator.

X-SPUR: EXPLAINABLE SURPRISAL-BASED REASONING FOR AUTOMOTIVE ETHERNET IDS

9

Surprisal Loss Scale

<0.5

.5–1.5

1.5–3

3–6

6–12

>12

(a) CAN-over-UDP · CAN DoS | C_D eth:ethertype:ip:udp:data …

eth.src

eth.type

ip.src

ip.dst

udp.srcport

udp.dstport

udp.length

payload

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08082e2ed500000039

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08163c3ce200000070

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08254a4af0000000a9

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08335857fe000000e0

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

084166650d00000019

▶ C_D

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

41480

60000

17

080000000000000000

▶ C_D

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

41480

60000

17

080000000000000000

▶ C_D

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

41480

60000

17

080000000000000000

(b) CAN-over-UDP · CAN Replay | C_R eth:ethertype:ip:udp:data …

eth.src

eth.type

ip.src

ip.dst

udp.srcport

udp.dstport

udp.length

payload

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08c1e6db8300000005

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08cff5e9910000003e

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08dd04f79f00000077

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08eb1206ad000000b0

benign

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

08fa2014ba000000e8

▶ C_R

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

089db66a840000f00f

▶ C_R

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

089eb66b840000f00f

▶ C_R

…

dc:a6:32:5d:ce:d2

0x0800

192.168.10.2

192.168.10.3

57746

60021

17

089fb76b850000f00f

(c) AVTP / MPEG-TS · Frame Injection | F_I eth:…:iec61883:mp2t …

channel

dbc

dbs

…

seqnum

sid

sph

spht

stream_data_len

…

benign

…

31

0xa0

0x06

…

0x57

63

True

0x6c86420b,0x6c86420

392

…

benign

…

31

0xb0

0x06

…

0x58

63

True

0x6c86420b,0x6c86420

392

…

benign

…

31

0xc0

0x06

…

0x59

63

True

0x6c86420b,0x6c86420

392

…

benign

…

31

0xd0

0x06

…

0x5a

63

True

0x6c86420b,0x6c86420

392

…

benign

…

31

0xe0

0x06

…

0x5b

63

True

0x6c86420b,0x6c86420

392

…

▶ F_I

…

31

0x00

0x06

…

0x00

63

True

0x442bd9b3,0x442bd9b

392

…

▶ F_I

…

31

0x10

0x06

…

0x01

63

True

0x442bd9b3,0x442bd9b

392

…

▶ F_I

…

31

0x20

0x06

…

0x02

63

True

0x442bd9b3,0x442bd9b

392

…

(d) gPTP · PTP Sync Injection | P_I eth:ethertype:ptp frame.protocols

frame.len

eth.dst

eth.src

eth.type

…

sequenceid

sourceportid

…

benign

eth:ethertype:ptp

90

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

34035

1

…

benign

eth:ethertype:ptp

60

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

34036

1

…

benign

eth:ethertype:ptp

90

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

34036

1

…

benign

eth:ethertype:ptp

60

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

34037

1

…

benign

eth:ethertype:ptp

90

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

34037

1

…

▶ P_I

eth:ethertype:ptp

60

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

18895

1

…

▶ P_I

eth:ethertype:ptp

60

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

18895

1

…

▶ P_I

eth:ethertype:ptp

60

01:80:c2:00:00:0e

00:fc:70:00:00:02

0x88f7

…

18895

1

…

(e) Malformed eth/ip · MAC Flooding | M_F eth:ethertype:ip attack-only (flow has no benign packets) frame.protocols

frame.len

eth.dst

eth.src

eth.type

ip.src

ip.dst

▶ M_F

eth:ethertype:ip

60

a4:3a:57:1a:4b:c6

5c:d1:98:50:23:a2

0x0800

69.124.20.25

184.161.182.41

▶ M_F

eth:ethertype:ip

60

a4:3a:57:1a:4b:c6

5c:d1:98:50:23:a2

0x0800

69.124.20.25

184.161.182.41

▶ M_F

eth:ethertype:ip

60

a4:3a:57:1a:4b:c6

5c:d1:98:50:23:a2

0x0800

69.124.20.25

184.161.182.41

▶ M_F

eth:ethertype:ip

60

a4:3a:57:1a:4b:c6

5c:d1:98:50:23:a2

0x0800

69.124.20.25

184.161.182.41

Fig. 4. Intrinsic field-level explainability of X-SPUR on TOW-IDS. Panels (a)–(e) show representative test examples for the five attack classes. Columns are tshark fields. Cells hold the raw tokens shaded by mean per-token cross-entropy (surprisal). The most-discriminative field is highlighted. Structurally fixed fields stay cold while the attacked field ignites: udp.dstport for CAN DoS (a) and the replayed payload for CAN Replay (b). PTP Sync Injection (d) shows the collapsed Sync/Follow-Up cadence in frame.len. MAC Flooding (e) shows a uniform elevation across every header. AVTP Frame Injection (c) is hardest: its valid frame only resets the already-warm iec61883.spht to an equally unpredictable value, so surprisal barely rises. Ellipses (. . . ) denote cold, low-surprisal fields omitted to fit the figure width.

10

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

TABLE VI P ER - PROTOCOL FALSE - POSITIVE RATES . Protocol

FP

Total benign

FPR

AVTP UDP gPTP

5,932 3,663 51

343,991 842,345 11,448

1.72% 0.43% 0.45%

Total

9,646

1,197,784

0.81%

TABLE VII C ROSS - DATASET DETECTION ON C AR DS; AERO F1 AT A MATCHED 0.5% BENIGN -FP BUDGET. T HE GRU AND AERO AUC S COINCIDE AT THIS PRECISION . Metric

X-SPUR

CNN

GRU

AERO [1]

AUC F1

0.9923 0.9842

0.6030 0.3577

0.6239 0.3167

0.6239 0.1074

Surprisal reflects spoofed synchronization semantics rather than corrupted payload, and detection stays at 99.9%. MAC Flooding (Fig. 4(e)). At the opposite extreme, malformed bare-IP frames whose random addresses never occur in benign traffic warm the entire header stack at once—the flood violates the learned packet structure as a whole, and recall is perfect. Overall, apart from MAC Flooding the evidence stays confined to the few fields each attack actually perturbs rather than spreading across generic Ethernet or IP headers, and surprisal magnitude and locality together predict difficulty. Recall is 99.6% at the first and last anomalous windows of each flow, compared with 97.7% for interior windows. G. Per-Protocol False-Positive Analysis Table VI splits the false positives by protocol (aggregate 0.81%). The gPTP figure is the clearest sign that per-protocol Z-score calibration removes protocol-level bias rather than relocating a global threshold: gPTP carries one of the highest benign baseline surprisals yet ends at an FPR comparable to UDP. The residual errors are localized, not diffuse: two AVTP MP2T streams alone produce 3,977 and 1,892 falsepositive windows, with a few UDP flows adding most of the remainder. Crucially, the fields that trigger them are the same IEC 61883/MPEG-TS stream fields, led by iec61883.spht, that carry the AVTP Frame Injection signal in Fig. 4(c). The false positives and the hardest attack thus share one interpretable cause—the model’s heightened sensitivity to the IEC 61883/MPEG-TS region—which tokento-field attribution makes directly visible where a windowlevel score could not. The coupling is quantitative: missed AVTP Frame Injection windows are near-misses (median z = 3.4 against τ = 3.91), but roughly twice as many benign AVTP windows occupy the same band, so recovering them by lowering τ would cost about two false positives per miss. H. Cross-Dataset Evaluation on CarDS To test the transferability of the proposed methodology beyond TOW-IDS, we train a separate X-SPUR model from

scratch on CarDS [30]. We use the 38-trace Train-Huge corpus to broaden benign-traffic coverage during pretraining. Finetuning uses the 27-trace Train split (Table III). We retain the model architecture, pretrain-to-finetune procedure, and training hyperparameters used for TOW-IDS (§IV-C). We also fit the BBPE tokenizer on the Train-Huge split listed in Table III. Calibration statistics and candidate thresholds are derived from its benign validation split. Despite a markedly different protocol stack (IPv6-based TCP/UDP with RTP and HTTP services rather than AVTP/gPTP) and attack taxonomy (network-reconnaissance scans, RTP stream replacement, and an infotainment service crash), X-SPUR retains strong performance (AUC 0.9923, F1-score 0.9842). The baselines do not: CNN and GRU causal language models trained on the same tokenized input collapse to AUC 0.6030 and 0.6239 (Table VII), and an architecture-verified reimplementation of AERO, trained on the same CarDS benign splits and scored on the same test captures, peaks at AUC 0.6239 across four trained configurations—essentially blind to the network scans (recall ≤8% at a 0.5% false-positive budget) and only partially sensitive to RTP replacement. Without changing the model architecture or training hyperparameters, X-SPUR achieves superior performance in this new environment. Per attack family, RTP camera replacement is detected almost perfectly (FNR 0.03%). The dominant Nmap SYN/ACK scans yield FNRs of 1.4–2.6%. The smallest classes remain the hardest. Their FNRs are 7.3% for FIN scans and 9.0% for the single IVI REST-API-crash trace. The same per-token attribution pattern appears on CarDS. Surprisal concentrates on a small set of protocol fields. These include TCP ports and flags for the scans, frame length for RTP replacement, and crafted TCP connection fields for the REST crash. Surprisal magnitudes are roughly 40× lower than on TOW-IDS. Perprotocol calibration accommodates this score-scale difference. I. Comparison with Prior Art Table VIII compares X-SPUR with representative automotive Ethernet baselines on TOW-IDS. These include reconstruction-based autoencoders, a semi-supervised classifier [29], and AERO [1]. X-SPUR achieves an AUC of 0.9987. AERO reports 0.9969. This difference is marginal and does not establish statistical superiority. X-SPUR also uses a larger parameter budget (98.5M versus 289k). AERO remains better on two attack families (Table V). The first distinction is intrinsic field-level attribution. AERO’s prototype-distance score does not natively map back to raw protocol fields. Its outlier score is the squared distance between a compressed 16-dimensional representation and a single learned prototype. Obtaining comparable field-level explanations would therefore require an additional attribution mechanism outside the detector. X-SPUR’s field attribution instead re-aggregates the same per-token surprisal used for detection (§IV-G). It averages the same token-level evidence within each field. This coupling is also reflected empirically in Fig. 4 and Table V. AVTP Frame Injection shows weak fieldlevel contrast between benign and attack traffic. It remains the hardest class (§V-F).

X-SPUR: EXPLAINABLE SURPRISAL-BASED REASONING FOR AUTOMOTIVE ETHERNET IDS

TABLE VIII A NOMALY DETECTION PERFORMANCE COMPARISON ON TOW-IDS ( BASELINE RESULTS AS REPORTED IN AERO [1]).

Method

Feature

# Params

Conv. autoencoder Payload (60 × 60) Conv. autoencoder Payload (60 × 60), byte changes Conv. autoencoder Payload (434 × 434) Vanilla autoencoder Time interval (mean, std., skew.) Shibly et al. [29] Payload (10% labeled) AERO [1] Three multimodal features X-SPUR Raw packet tokens (BBPE)

AUC

2,661k 0.7904 2,661k 0.8103 4,122k 0.9560 1,336k 0.9648 112k 0.9730 289k 0.9969 98.5M 0.9987

TABLE IX B ACKBONE COMPARISON ON TOW-IDS. PARAMETER COUNTS ARE FOR THE COMPLETE MODELS . Backbone Mamba2 (X-SPUR) Transformer LSTM Dilated TCN

# Params

AUC

98.5M 97.6M 97.6M 99.4M

0.9987 0.9944 0.9818 0.9960

The second distinction is the absence of protocol-specific handcrafted feature design. AERO’s own feature extractor requires an expert to hand-select payload byte-offset parameters (j, n) for each IVN [1]. Its AERO-FG2 ablation, using (j, n)=(0, 434), shows that this choice materially affects detection performance. X-SPUR instead operates on a byte-level token representation that requires no protocolspecific feature engineering, avoiding the deployment-specific redesign required by AERO’s feature extractor. The CarDS evaluation provides an empirical test of this property. We use the same model architecture and training hyperparameters on its markedly different protocol stack (§V-A). A separate XSPUR model is trained from scratch using CarDS benign data, without dataset-specific feature redesign or a search over training hyperparameters (§V-H). J. Backbone Comparison Table IX compares Mamba2 with Transformer, LSTM, and dilated TCN variants. All models use the same TOW-IDS splits and token–time fusion, with w = 10 and Lmax = 256. They share the dual top-k scoring, K = 63 smoothing, and per-protocol calibration. The total parameter count of each alternative differs from Mamba2’s by less than 1%. The Transformer also achieves a high AUC of 0.9944. Mamba2 achieves the highest AUC of 0.9987, supporting its use as the X-SPUR backbone in this comparison. K. Computational and Deployment Feasibility We examine sequence-length scaling using FP32 forwardpass timings on a single RTX 5090. The parameter counts of Mamba2 and the Transformer differ by less than 1%. We use randomly initialized weights and synthetic inputs. At batch 32, Mamba2 takes 40.57 ms at L = 256 and 167.75 ms at L = 1024. The Transformer takes 37.98 ms and 214.55 ms, respectively. At batch 32, Mamba2 is faster at L = 1024 but

11

not at L = 256. The Transformer is faster at both lengths with batch 1. We also profile the trained 98.5M-parameter model on Jetson AGX Orin 64 GB in MAXN mode. We use FP16, CUDA Graphs, and synthetic inputs at L = 256. Timing covers input transfer, model inference, and dual top-k scoring. Mean latency is 6.19 ms at batch 1 and 140.51 ms at batch 32. The latter yields 227.74 windows/s with 1.48 GB of peak PyTorchreserved GPU memory. The TOW-IDS captures generate about 1,973 windows/s on average at w = 10 and stride 1. To reduce this load, we shorten the input to L = 192 without retraining. We retain w = 10 but evaluate one window every 12 packets within each flow. At this stride, processing the TOW-IDS traffic requires about 164 window evaluations/s on average. On Orin, the measured batch-32 throughput is 298.71 windows/s. This exceeds the average processing requirement. We also assess detection performance under this sampling policy. We reduce the smoothing width to K = 7 for the wider sampling interval. We recalibrate using benign validation data. For windows not evaluated by the model, we carry forward the most recent score. This yields an offline TOW-IDS AUC of 0.9962 over all test windows. Thus, this setting can meet the average processing demand while retaining high detection performance. VI. C ONCLUSION This article presented X-SPUR, an explainable unsupervised intrusion detection framework for automotive Ethernet that replaces handcrafted multimodal feature engineering with causal next-token prediction over tokenized packet fields. Learning benign sequential structure with a Mamba2 statespace language model, X-SPUR detects intrusions without attack labels while keeping anomaly evidence at the token level. On the TOW-IDS dataset, X-SPUR achieves an AUC of 0.9987. AERO [1] reports 0.9969. Our ablations show that bimodal payload–timing fusion adds discriminative signal beyond packet content. Flow-level smoothing with dual topk per-protocol Z-score calibration also suppresses protocolspecific false positives to 0.81% of benign windows. Tokenlevel surprisal yields intrinsic field-level attribution. It exposes distinct attack signatures and traces detections and residual false positives to the same IEC 61883/MPEG-TS stream fields. We also apply the same model architecture and training hyperparameters to a separately trained CarDS model. It preserves strong detection with an AUC of 0.9923. A faithful AERO reimplementation achieves an AUC of 0.6239. These results support the transferability of the proposed methodology across datasets and protocol stacks. Limitations and Future Work. AVTP Frame Injection remains the hardest attack to detect. Its counter and timestamp fields are difficult to predict even in benign traffic. This leaves a weak surprisal contrast after injection. The centered K = 63 smoother uses up to 31 subsequent windows from the same flow. At stride 1, the median waiting time is 466 ms on TOW-IDS and 10.7 ms on CarDS. The corresponding 95thpercentile values are 959 ms and 25.2 ms. Causal smoothing is a future direction for latency-sensitive deployment.

12

IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS — ACCEPTED MANUSCRIPT

The 98.5M-parameter model also remains computationally demanding for production ECUs. The single-Orin profile does not sustain the observed aggregate stride-1 load. Distillation and quantization remain priorities for embedded deployment. Field-level surprisal indicates deviation from learned benign behavior. It supports analyst interpretation but does not establish maliciousness. R EFERENCES [1] S. Jeong, H. K. Kim, M. L. Han, and B. I. Kwak, “AERO: Automotive Ethernet real-time observer for anomaly detection in in-vehicle networks,” IEEE Trans. Ind. Informat., vol. 20, no. 3, pp. 4651–4662, 2024. [2] S. Tariq, S. Lee, H. K. Kim, and S. S. Woo, “CAN-ADF: The controller area network attack detection framework,” Comput. Secur., vol. 94, p. 101857, 2020. [3] S. Jeong, S. Lee, H. Lee, and H. K. Kim, “X-CANIDS: Signal-aware explainable intrusion detection system for controller area network-based in-vehicle network,” IEEE Trans. Veh. Technol., vol. 73, no. 3, pp. 3230– 3246, 2024. [4] F. Luo, Z. Yang, Z. Zhang, Z. Wang, B. Wang, and M. Wu, “A multilayer intrusion detection system for SOME/IP-based in-vehicle network,” Sensors, vol. 23, no. 9, p. 4376, 2023. [5] F. Gail, R. Rieke, F. Fenzl, and C. Krauß, “Evaluation of decision treebased rule derivation for intrusion detection in automotive Ethernet,” in Proc. IEEE 22nd Int. Conf. Trust, Secur. Privacy Comput. Commun. (TrustCom), 2023, pp. 1392–1399. [6] Y. Wang, Y. Wu, Y. Xu, K. Zhang, and Y. Xu, “Research on network intrusion detection based on weighted histogram algorithm for in-vehicle Ethernet,” Sensors, vol. 25, no. 11, p. 3541, 2025. [7] T. Kim, H. Park, I. You, and B. I. Kwak, “XGBoost-based anomaly detection framework for SOME/IP in in-vehicle networks,” Systems, vol. 14, no. 2, p. 196, 2026. [8] P. R. X. Carmo, P. Freitas de Araujo-Filho, D. R. Campelo, E. Freitas, A. T. de Oliveira Filho, and D. F. H. Sadok, “Machine learning-based intrusion detection system for automotive Ethernet: Detecting cyberattacks with a low-cost platform,” in Proc. 40th Braz. Symp. Comput. Netw. Distrib. Syst. (SBRC), 2022, pp. 196–209. [9] M. L. Han, B. I. Kwak, and H. K. Kim, “TOW-IDS: Intrusion detection system based on three overlapped wavelets for automotive Ethernet,” IEEE Trans. Inf. Forensics Security, vol. 18, pp. 411–422, 2023. [10] S. Jeong, B. Jeon, B. Chung, and H. K. Kim, “Convolutional neural network-based intrusion detection system for AVTP streams in automotive Ethernet-based networks,” Veh. Commun., vol. 29, p. 100338, 2021. [11] N. Alkhatib, H. Ghauch, and J.-L. Danger, “SOME/IP intrusion detection using deep learning-based sequential models in automotive Ethernet networks,” in Proc. IEEE 12th Annu. Inf. Technol. Electron. Mobile Commun. Conf. (IEMCON), 2021, pp. 0954–0962. [12] Y. Zhang and J. Xiu, “Real-time automotive Ethernet intrusion detection using sliding window-based temporal convolutional residual attention networks,” J. Inf. Secur. Appl., vol. 94, p. 104263, 2025. [13] N. Alkhatib, M. Mushtaq, H. Ghauch, and J.-L. Danger, “Unsupervised network intrusion detection system for AVTP in automotive Ethernet networks,” in Proc. IEEE Intell. Veh. Symp. (IV), 2022, pp. 1731–1738. [14] J. Liu, W. Fan, Y. Dai, E. Lim, Z. Pan, and A. Lisitsa, “Leveraging semisupervised learning for enhancing anomaly-based IDS in automotive Ethernet,” in Proc. IEEE 23rd Int. Conf. Trust, Secur. Privacy Comput. Commun. (TrustCom), 2024, pp. 1563–1571. [15] M. S. G. A. Leandro, P. Freitas de Araujo-Filho, D. R. Campelo, and L. F. M. da Luz, “SeqWatch: Unsupervised sequence-based intrusion detection system for automotive Ethernet,” in Proc. 43rd Braz. Symp. Comput. Netw. Distrib. Syst. (SBRC), 2025, pp. 378–391. [16] L. F. M. da Luz, P. Freitas de Araujo-Filho, and D. R. Campelo, “Multistage deep learning-based intrusion detection system for automotive Ethernet networks,” Ad Hoc Netw., vol. 162, p. 103548, 2024. [17] Z. Wu, H. Zhang, P. Wang, and Z. Sun, “RTIDS: A robust transformerbased approach for intrusion detection system,” IEEE Access, vol. 10, pp. 64 375–64 387, 2022. [18] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in Proc. 1st Conf. Lang. Model. (COLM), 2024. [19] T. Dao and A. Gu, “Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality,” in Proc. 41st Int. Conf. Mach. Learn. (ICML), 2024, pp. 10 041–10 071.

[20] Y.-C. Yu, Y.-C. Ouyang, and C.-A. Lin, “CBMAD: Anomaly detection in IoT network traffic via consistent bidirectional Mamba autoencoder,” in Proc. IEEE 26th Int. Conf. High Perform. Switching Routing (HPSR), 2025, pp. 1–6. [21] T. Wang, X. Xie, W. Wang, C. Wang, Y. Zhao, and Y. Cui, “NetMamba: Efficient network traffic classification via pre-training unidirectional Mamba,” in Proc. IEEE 32nd Int. Conf. Netw. Protocols (ICNP), 2024, pp. 1–11. [22] S. Neupane, J. Ables, W. Anderson, S. Mittal, S. Rahimi, I. Banicescu, and M. Seale, “Explainable intrusion detection systems (X-IDS): A survey of current methods, challenges, and opportunities,” IEEE Access, vol. 10, pp. 112 392–112 415, 2022. [23] D. Han, Z. Wang, W. Chen, Y. Zhong, S. Wang, H. Zhang, J. Yang, X. Shi, and X. Yin, “DeepAID: Interpreting and improving deep learning-based anomaly detection in security applications,” in Proc. ACM SIGSAC Conf. Comput. Commun. Secur. (CCS), 2021, pp. 3197– 3217. [24] K. Kaya, E. Ak, S. Bas, B. Canberk, and S. G. Oguducu, “X-CBA: Explainability aided CatBoosted Anomal-E for intrusion detection system,” in Proc. IEEE Int. Conf. Commun. (ICC), 2024, pp. 2288–2293. [25] X. Li, C. Qian, Q. Wang, J. Kong, Y. Wang, Z. Yao, B. Ji, L. Cheng, G. Zhou, and H. Shao, “Lens: A knowledge-guided foundation model for network traffic,” 2026, arXiv:2402.03646. [26] Q. P. Nguyen, K. W. Lim, D. M. Divakaran, K. H. Low, and M. C. Chan, “GEE: A gradient-based explainable variational autoencoder for network anomaly detection,” in Proc. IEEE Conf. Commun. Netw. Secur. (CNS), 2019, pp. 91–99. [27] M. Du, F. Li, G. Zheng, and V. Srikumar, “DeepLog: Anomaly detection and diagnosis from system logs through deep learning,” in Proc. ACM SIGSAC Conf. Comput. Commun. Secur. (CCS), 2017, pp. 1285–1298. [28] H. Guo, S. Yuan, and X. Wu, “LogBERT: Log anomaly detection via BERT,” in Proc. Int. Joint Conf. Neural Netw. (IJCNN), 2021, pp. 1–8. [29] K. H. Shibly, M. D. Hossain, H. Inoue, Y. Taenaka, and Y. Kadobayashi, “A feature-aware semi-supervised learning approach for automotive Ethernet,” in Proc. IEEE Int. Conf. Cyber Secur. Resilience (CSR), 2023, pp. 426–431. [30] W. Hellemans, J. Hamborg, T. Lauser, M. M. Rabbani, B. Preneel, C. Krauß, and N. Mentens, “CarDS - controller area network and automotive Ethernet realistic data set,” in Proc. IEEE Annu. Comput. Secur. Appl. Conf. (ACSAC), 2025, pp. 798–814.

Jisoo Kim (Student Member, IEEE) is currently pursuing the B.S. degree in data science with Sookmyung Women’s University, Seoul, Republic of Korea. She is an Undergraduate Researcher with the System and Network Security Laboratory (SNSec Lab), Sookmyung Women’s University. Her research focuses on data-driven cybersecurity, particularly explainable and unsupervised intrusion detection for automotive Ethernet. Her research interests include network security, automotive cybersecurity, anomaly detection, and explainable machine learning for cybersecurity.

Seonghoon Jeong (Member, IEEE) received the Ph.D. degree in information security from Korea University, Seoul, Republic of Korea, in 2023. He is currently an Assistant Professor with the Division of Artificial Intelligence Engineering, Sookmyung Women’s University, Seoul, Republic of Korea. He leads the System and Network Security Laboratory (SNSec Lab). His research focuses on system and network security, particularly explainable and unsupervised intrusion detection for connected vehicles. His current research interests also include learning-based binary analysis, microarchitectural security, and foundation models for cybersecurity.

Record · ID 1006802 · SHA-256 f3ab9716d38a1057
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.