Conceptio › Archive › arXiv CS
arXiv CSopen access

TSNBench: Benchmarking LLM Proficiency in Time-Sensitive Networking

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

arXiv:2605.09481v1 [cs.NI] 10 May 2026

TSNBench: Benchmarking LLM Proficiency in Time-Sensitive Networking

Rubi Debnath1∗ Daniel Bujosa Mateu2 Luxi Zhao3 Silviu S. Craciunas2,4 Paul Pop2 Sebastian Steinhorst1 1 Technical University of Munich, Munich, Germany 2 Technical University of Denmark, Kongens Lyngby, Denmark 3 Beihang University, Beijing, China 4 NXP Semiconductors, Vienna, Austria

Abstract We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.1 standards for deterministic communication with bounded latency in safety-critical domains such as autonomous vehicles, aviation, defense, and industrial automation. While LLMs have been extensively evaluated on general knowledge tasks, their capabilities in safety-critical networking domains remain largely unexplored. TSNBench comprises 939 expert-validated multiple-choice questions (MCQs) covering diverse TSN mechanisms, along with 100 open-ended Worst-Case Delay (WCD) computation tasks for Credit-Based Shaper (CBS) and Cyclic Queuing and Forwarding (CQF) across varying network topologies and traffic conditions. MCQ answers are validated by domain experts, and open-ended ground truth WCD values are computed using a verified Network Calculus (NC) solver for CBS and closed-form mathematical upper bounds for CQF. We evaluate 16 LLMs and find that although models achieve 67 to 95% accuracy on MCQs, they fail substantially on open-ended WCD computation. For CBS, only GPT-5 achieves a Mean Absolute Percentage Error (MAPE) of 36.2%, meaning its predicted WCD deviates by 36.2% of the actual TSN flow delay on average, while most models exceed 80%. For CQF, the best model achieves 41.8% MAPE, with most models clustering between 80% and 100%. Such errors are large relative to TSN latency budgets and can lead to violations of real-time constraints and unsafe configurations. TSNBench demonstrates that MCQ benchmarks may overestimate LLM capabilities in safety-critical networking domains.

1

Introduction

Recent advances in large language models (LLMs) across different domains such as engineering [Jackson et al., 2025, Guo et al., 2025], medicine [Xie et al., 2025, Liu et al., 2023, Li et al., 2024], clinical practice [Kweon et al., 2024], computer networking [Sharma and Yegneswaran, 2023], telecommunications [Maatouk et al., 2026, Ferrag et al., 2026, Oluwaseyi et al., 2025, Gajjar et al., 2025], and automation [Shen et al., 2024] have shown groundbreaking performance in assisting engineers, practitioners, researchers [Huang et al., 2023, Sun et al., 2024], and doctors in solving real-world problems. System engineers are increasingly using LLMs to design and configure networks [Wang et al., 2024a], generate code, and analyze network logs. With this, they are entering new territory: safety-critical application domains such as autonomous vehicles, aerospace [Fiori et al., 2024, Sanchez-Garrido et al., 2021], defense [Elliott, 2023], and industrial communication [Zhang et al., ∗ Corresponding Author.

Preprint.

2024]. In these contexts, the accuracy, reliability, and consistency of LLMs become far more than leaderboard metrics, as they become engineering requirements. Time-Sensitive Networking (TSN) [802, 2018], standardized by the IEEE 802.1 Working Group (WG), is a layer-2 Ethernet technology that provides deterministic communication guarantees for safety-critical applications. TSN deployments typically separate traffic based on timing criticality. Safety-critical periodic communication with guaranteed latency and bounded jitter is categorized as time-triggered (TT) [Ademaj et al., 2019] traffic and is served using the IEEE 802.1Qbv timed-gate mechanism. TT transmissions are controlled by a Gate-Control List (GCL), computed offline using exact methods such as SMT-based synthesis [Craciunas et al., 2016] or heuristic approaches [Pop et al., 2016, Gavriluţ et al., 2018, Bujosa et al., 2022]. In contrast, periodic or sporadic communication requiring bounded end-to-end latency but less stringent jitter control is classified as Audio Video Bridging (AVB) stream traffic [Böhm and Wermser, 2021, Bruckner et al., 2019]. Consequently, Worst-Case Delay (WCD) estimation errors of tens or hundreds of microseconds are significant, as they can consume timing margins, violate deadlines, or lead to infeasible TSN configurations. In mission-critical deployments, such errors can have severe consequences. A misconfigured TSN network can cause, for example, a robotic arm to miss a critical assembly step, a brake system to fail on a highway, an aircraft control system to respond incorrectly, a defense mechanism to collapse, or a spacecraft to miss a vital signal. These failures may result from sub-millisecond timing violations caused by a single misconfiguration. These risks highlight the importance of accurate analysis and configuration in TSN systems, especially as LLMs are increasingly integrated into network management workflows. Therefore, their domain proficiency must be rigorously evaluated. However, to the best of our knowledge, no existing benchmark evaluates LLM proficiency in TSN. To fill this gap, we introduce TSNBench, the first benchmark for evaluating LLM proficiency in TSN, comprising two complementary evaluation components. The first is a 939-question expert-validated multiple-choice question and answer (MCQA) dataset, generated from 83 peer-reviewed research papers using three LLMs from distinct model families and rigorously reviewed by five domain experts, each with over eight years of TSN research experience. The second is a set of open-ended questions requiring multi-step WCD computation for two widely deployed TSN mechanisms, namely Credit-Based Shaper (CBS) [802, 2010] and Cyclic Queuing and Forwarding (CQF) [802, 2017, Yan et al., 2020], across varying network topologies and traffic flows, with ground truth computed using a verified Network Calculus (NC) solver [Zhao et al., 2018] for CBS and closed-form mathematical upper bounds for CQF [Wang et al., 2023]. These open-ended WCD questions are intended as a closed-book stress test of standalone model capability, rather than as a deployment workflow for free-text LLM timing outputs. Detailed background on TSN, NC, CBS, and CQF is provided in Appendix 7, 8, 9, and 10, respectively. While general-purpose benchmarks such as MMLU [Hendrycks et al., 2021] and MMLU-Pro [Wang et al., 2024b] evaluate broad subject knowledge spanning elementary mathematics, history, and law, they are fundamentally unsuited for safety-critical domain-specific evaluation. Answering a multiple-choice question about elementary school history is categorically different from answering TSN terminology questions and correctly computing a WCD under NC constraints for a given network topology. Without a benchmark that captures this distinction, there is no principled way to measure LLM progress in deterministic networking domains. TSNBench is designed precisely to expose this gap. We evaluate 16 LLMs comprising open-source and closed-source models, as well as general-purpose and reasoning-specialized architectures. Our results reveal a striking dissociation, where models achieve 67 to 95% accuracy on MCQA yet fail substantially on open-ended WCD computation. The best-performing model, GPT-5, achieves a Mean Absolute Percentage Error (MAPE) of 36.2% on CBS, while most models exceed 80%. This is concerning in a domain where timing violations of tens of microseconds, even 1% of a 1000 µs deadline, may cause system failures. Our key contributions are: 1. First expert-validated TSN benchmark: TSNBench evaluates LLM knowledge of TSN mechanisms through 939 expert-validated MCQs derived from peer-reviewed TSN literature. 2. Open-ended timing-analysis tasks: TSNBench includes open-ended WCD computation tasks for CBS and CQF with ground truth computed using a verified NC solver for CBS and closed-form mathematical bounds for CQF. 2

3. Evaluation across 16 LLMs: We evaluate both open-source and closed-source models, including general-purpose and reasoning-specialized models, and show that high MCQA accuracy does not reliably predict accurate WCD computation. In summary, TSNBench provides the research community with the first rigorous evaluation resource for LLM proficiency in TSN, offering valuable insights to both the real-time networking community exploring LLM-assisted TSN management and the machine learning community seeking to understand the limits of LLMs in safety-critical, computationally demanding domains.

2

Related Work

General LLM Benchmarks: Benchmarking and datasets are essential for measuring LLM progress and identifying key gaps and limitations [Hendrycks et al., 2021, Wang et al., 2024b]. General knowledge benchmarks such as MMLU [Hendrycks et al., 2021] and MMLU-Pro [Wang et al., 2024b] evaluate broad subject knowledge including elementary mathematics, history, computer science, and law, using multiple-choice questions. Domain-specific benchmarks have extended this paradigm to medicine [Xie et al., 2025, Liu et al., 2023, Li et al., 2024], clinical practice [Kweon et al., 2024], law [Guha et al., 2023], code generation [Hua et al., 2025, Huang et al., 2024], and scientific research [Sun et al., 2024]. While these benchmarks have driven significant progress, they are not designed to evaluate safety-critical networking tasks. Most rely on multiple-choice evaluation, and none assess whether a model can perform the multi-step computational reasoning required in safety-critical networking domains. TSNBench addresses this gap by introducing MCQA and open-ended WCD computation questions with ground truth verified by state-of-the-art NC solvers, providing an evaluation of TSN that no existing general benchmark captures. Networking and Telecommunications Benchmarks: In the last few years, several benchmarks have evaluated LLM proficiency in networking and telecommunications domains. TeleQnA [Maatouk et al., 2026] presents an MCQ dataset for telecommunications, generated from research documents and 3GPP standards and validated by domain experts. 6G-Bench [Ferrag et al., 2026] presents an MCQ-based dataset for 6G networks containing 3,722 difficult questions validated through automated filtering and expert human review. Beyond question-answering benchmarks, NetConfEval [Wang et al., 2024a] evaluates LLMs on network configuration tasks and demonstrates that LLMs can simplify and automate complex network management tasks. LLMs for TSN and Real-Time Networks The application of LLMs to TSN management and orchestration is still at a very early stage, with only limited initial studies available. Windmann et al. [2025] explored the use of LLMs for configuring hybrid 5G/TSN networks by assisting users with manual configuration tasks and suggesting configurations in a 5G-TSN network. However, this work remains preliminary and does not provide experimental results. Overall, prior work does not provide a systematic benchmark or rigorous evaluation of LLM proficiency across TSN mechanisms, nor does it assess computational reasoning capabilities for WCD analysis. TSNBench fills this gap by providing the first structured benchmark covering both declarative TSN knowledge through MCQA and computational reasoning through open-ended WCD evaluation.

3

TSNBench

Unlike established domains such as medicine [Xie et al., 2025], 5G [Oluwaseyi et al., 2025, Maatouk et al., 2026], general human knowledge [Phan et al., 2026, Hendrycks et al., 2021, Wang et al., 2024b], coding [Hua et al., 2025, Huang et al., 2024], and law [Guha et al., 2023], no open-source TSN dataset exists for LLM evaluation [Zhang et al., 2024, Peng et al., 2023, Zanbouri et al., 2025, Adil et al., 2026]. As highlighted in [Liu et al., 2023], the data source determines the reliability of a dataset, and generating a high-quality dataset is a crucial prerequisite for meaningful benchmarking. We describe the TSNBench construction pipeline below, with full details provided in Appendix 11. 3.1

Dataset Source Selection

Published research papers and standards are among the most reliable sources for building domainspecific datasets [Liu et al., 2023]. Since TSN knowledge originates primarily from peer-reviewed 3

Generator LLM Model

Research Documents

Keyword Generator

TSN Keyword

Expert Review

Final TSN Keyword

Figure 1: TSNBench keyword-generation pipeline. TSN keywords are extracted from research documents using an LLM, expert-verified, and used for MCQA generation as described in Section 3.3. Table 1: Models used in the TSNBench keyword extraction and question generation pipeline. All models are used with default settings and last accessed in April 2026. Claude Sonnet 4 serves two distinct roles: keyword extraction and question generation. These roles use identical model configurations but operate on different inputs and prompts. Full dataset generation details are provided in Appendix 11. Model

API

Model ID

Organization

Usage

Claude Sonnet 4

Anthropic API

claude-sonnet-4-20250514

Anthropic

Keyword extractor

Claude Sonnet 4 GPT-4o mini Llama 3.1 70B

Anthropic API OpenAI API HF Router

claude-sonnet-4-20250514 gpt-4o-mini Llama-3.1-70B-Instruct

Anthropic OpenAI Meta

Generator Generator Generator

research and IEEE 802.1 TSN standards, we curate a collection of open-access research documents as our source corpus. To avoid copyright issues and exclude papers with incorrect results or flawed methodologies, we include only published open-access papers. For papers not available in openaccess form, we use arXiv versions that have been published or accepted, excluding unpublished preprints with unverified results. Where possible, we also collect author manuscript versions with proper attribution. To ensure quality, we prioritize highly cited papers from reputable venues while accounting for publication timeline, as recent papers naturally have fewer citations. In total, we collect 83 research papers covering a broad range of TSN mechanisms, including Time-Aware Shaper (TAS), CBS, CQF, NC-based schedulability analysis, performance evaluation, hardware experiments, combined shapers such as TAS+CBS [Zhao et al., 2022], and Multi-CQF [Alexandris et al., 2022]. Detailed background on TSN, related work, and its mechanisms is given in Appendix 7. 3.2

Keyword and Acronym Extraction

TSN employs specialized vocabulary, similar to other communication domains [Andrews et al., 2014, Saad et al., 2020, Ma et al., 2019]. A successful LLM that understands TSN should be able to reason correctly about TSN terminology. A model that cannot differentiate between TAS and CBS, or cannot correctly expand TSN-specific acronyms, cannot be considered proficient in TSN. To capture this dimension, we extract keywords and acronyms widely used in TSN literature and use them to guide MCQA generation. All terms are extracted from the 83 research documents using Claude Sonnet 4, as shown in Table 1, and stored in JSON format. Each document is preprocessed to remove non-relevant content, including author names, affiliations, figures, tables, URLs, and pseudocode. The model is instructed to extract only terms defined within the document, without relying on pretrained knowledge, and to provide each term’s acronym, full form, and one-to-two-sentence definition from the source. The extracted set is then reviewed by domain experts to resolve duplicates, retaining the longer definition in cases of conflict. Figure 1 illustrates this pipeline. 3.3

MCQA Generation, Post-Processing, and Expert Review

Raw MCQA Generation: To optimize time and reduce manual effort, we use an LLM-based approach to generate MCQAs from research documents. The keyword file is provided alongside the research documents as additional input, serving as an independent source to complement research paper content during generation. We use three models from distinct families, namely Claude Sonnet 4, GPT-4o mini, and Llama 3.1 70B, as shown in Table 1. These models are deliberately selected to ensure diverse styles and reasoning capabilities, thereby reducing generative bias. The same 4

system prompt is used for all models, and each research paper is assigned to exactly one model in a round-robin manner. Non-relevant sections, such as author information, affiliations, references, URLs, figures, tables, and pseudocode, are removed from each document before generation. Post-Processing: LLM-generated MCQAs cannot be used directly for benchmarking, as they may contain incorrectly formulated questions, incomplete options, or vague and incorrect answer choices. To address positional bias introduced by the generating model, answer options are shuffled randomly prior to human expert review, with the correct answer label updated to reflect the new ordering. Human-Based Domain Expert Review: Given the safety-critical nature of TSN, rigorous human validation is essential. We engage five TSN domain experts: three senior professors with more than 15 years of research experience and two postdoctoral researchers with more than 8 years of expertise. Each question is independently evaluated with four outcomes: (i) accept - correct and clear; (ii) revise - requires modification for clarity or correctness; (iii) reject - the question is incorrect, misleading, or irrelevant; or (iv) doubtful - the expert is uncertain and passes it to remaining reviewers for consensus. Questions without consensus are discarded. Full review criteria are provided in Appendix 11.1 and Table 5. Table 2 summarizes the dataset statistics and Figure 2 illustrates the full pipeline. Table 2: TSNBench dataset construction statistics. Full generation details are in Appendix 11. Type

Category

Count

Total raw questions generated by models 1326 Questions removed after expert review 387 Questions revised by domain experts 185 Questions in the final dataset (used for benchmarking) 939

MCQA

Credit-Based Shaper (CBS) Open-ended questions Cyclic Queuing and Forwarding (CQF)

100 100

Reject Reject LLM Generator Research Documents

Doubtful MCQA Generator

Raw Question Set

Second Expert Review

Revise

TSN MCQA dataset

Post Processing Expert Review Revise

TSN Keywords

Accept

Accept

Figure 2: Pipeline of our TSNBench MCQA dataset generator, showing all steps from raw generation to the final validated dataset. 3.4

Open-Ended Question Formulation

While MCQA evaluates declarative TSN knowledge, open-ended questions assess whether LLMs can perform the multi-step mathematical reasoning required in real TSN deployment. We evaluate WCD computation, as WCD is a central key performance indicator (KPI) in TSN network design and directly determines whether a network meets its stringent timing requirements. We select two TSN mechanisms for this evaluation: CBS and CQF. CBS is widely deployed for audio-video traffic and requires NC-based analysis, making it mathematically demanding. CQF is a more recently standardized TSN mechanism whose WCD can be computed from a closed-form equation given routing and cycle duration (T ), providing a complementary evaluation that isolates formula application from NC complexity. Together, these two mechanisms span a meaningful range of WCD computation difficulty. Ground truth WCD values are computed using a verified state-of-the-art NC tool [Zhao et al., 2018] for CBS and closed-form mathematical upper bound for CQF. We release all ground truth WCD values alongside the questions to support future open-source community evaluations. Each open-ended question is formulated by domain experts, as shown in Figure 3, and comprises three 5

+ Network topology

Human expert

+ Flow information

Routing of the flows

Question formulation

Open-ended TSN questions

Figure 3: Pipeline for TSNBench open-ended question formulation by domain experts. Each question comprises three components: network topology, flow information, and flow routing. components: network topology, flow information, and flow routing. In TSNBench, three topologies are used to cover a broad range of scenarios: (i) one-switch topology (Figure 15), (ii) medium-mesh topology (Figure 16), and (iii) ring topology, representing industrial networks (Figure 17). Each topology consists of end nodes and switches connected via Ethernet links, with unicast traffic flows transmitted from a sender to a single receiver. Flows consist of Ethernet frames whose maximum payload is bounded by the Maximum Transmission Unit (MTU). Further topology, flow, and routing details are provided in Appendix 11.4. 3.5

Prompt Design

For both MCQA and open-ended evaluations, each prompt defines the model’s role as a TSN expert. For MCQA, we use zero-shot prompting with no in-context examples, representing a conservative approach that measures inherent TSN proficiency, ensuring that the output performance reflects the model’s domain knowledge rather than in-context pattern matching. For open-ended questions, we also use a zero-shot setting, providing no example WCD calculations or NC or CQF equations, ensuring the model independently recalls and applies the correct computational methodology. For both question types, the model is asked to provide a confidence score alongside its answer. The open-ended prompt comprises three variable components: network topology, flow parameters, and pre-computed shortest path routes. The same prompt template is used across all 100 open-ended evaluation instances per mechanism, with only these three components varying. Fixed network constants are maintained throughout to ensure comparability across models and instances. A detailed discussion of the open-ended prompt design is provided in Appendix 11.3. 3.6

Model Scoring and Ground Truth

For the MCQA dataset, performance is measured as the percentage of questions answered correctly, reported as accuracy. For the open-ended questions, we evaluate the computational reasoning capability of each model by comparing its predicted WCD values against ground truth values. For CBS, ground truth WCD values are derived using NC-based Total Flow Analysis (TFA). Specifically, h the worst-case delay upper bound Dfh for flow f ∈ FM at h equals the worst-case delay upper bound i h DMi for all flows with the same priority Mi aggregating at h,   h h h h h Dfh = DM = hDev(αM , βM ) = sup inf τ ≥ 0 | αM (t) ≤ βM (t+τ ) , (1) i i i i i t≥0

h where αM (t) represents the arrival curve of aggregate flows of priority Mi passing through h, and i h βMi (t) represents the service curve for these corresponding flows. The end-to-end WCD for a flow is

obtained by summing per-port delay bounds along its route. Full NC methodology and proofs are provided in Appendix 8. For CQF, the worst-case end-to-end delay is given by the closed-form expression WCD = fi .ϕ + (SWnum + 1) · T + ξ, 6

(2)

Model Grok 4.1 Fast† Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-thinking) GPT-4o GPT-4o mini Llama 3.3 Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3† GPT-5† DeepSeek-V3.2 (Thinking)† Gemini 2.5 Flash Llama 3.2 1B Qwen3 8B Ministral 3 8B †

Accuracy (%)

Avg. Consistency

Avg. Latency (ms)

Avg. Conf.

ECE↓

Brier↓

CW Rate↓

93.2 91.7 94.0 91.8 88.3 88.9 92.1 92.8 95.3 94.7 95.0 94.7 90.1 67.4 83.7 86.9

0.9858 0.9986 0.9993 0.9957 0.9950 0.9950 0.9965 0.9975 0.9993 0.9840 0.99 0.9819 0.9847 1.0 0.9897 0.9954

6673 515 804 729 799 365 653 5498 1842 3845 5630 4400 6744 669 15103 345

0.9509 0.9760 0.9312 0.8782 0.9004 0.9082 0.9779 0.9476 0.9374 0.7524 0.8773 0.9202 0.9674 0.8529 0.8616 0.9649

0.0151 0.0328 0.0105 0.0354 0.0538 0.0450 0.0295 0.0214 0.0181 0.1874 0.0569 0.0224 0.0539 0.1859 0.0351 0.0822

0.0599 0.0764 0.0526 0.0765 0.0974 0.0918 0.0750 0.0646 0.0429 0.0852 0.0475 0.0487 0.0942 0.2544 0.1322 0.1230

99.0 100.0 96.4 99.2 77.8 100.0 100.0 100.0 86.6 3.4 51.7 78.1 95.4 99.0 100.0 100.0

Temperature parameter not supported. Evaluated with default settings. ↓ lower is better.

o3 (Default Temp.)

Accuracy

ECE=0.195 CW=3.4%

1.0 0.5 0.0

0.2

0.4 0.6 0.8 1.0 Confidence Grok 4.1 Fast (NR) (Temp=0)

ECE=0.059 CW=100.0%

Accuracy

Figure 4: TSNBench MCQA results across 16 models. Accuracy is the percentage of correct answers out of 939 questions, and consistency measures whether the model gives the same response across three runs. All models are evaluated at temperature 0.0 for deterministic performance; models without temperature support use their default setting and are marked with † . Full model details are given in Table 6, and extended results with temperature comparisons are provided in Table 7 in Appendix 12.

1.0 0.5 0.0 0.80 0.85 0.90 0.95 1.00 Confidence

Figure 5: Reliability plot for o3 and Grok 4.1 Fast (NR). Full reliability analysis are in Figure 6.

where fi .ϕ is the flow offset at the source node in µs, SWnum is the number of switches along the flow route, T is the cycle duration in µs, and ξ denotes the network specific delays including processing delay, propagation delay, switching delay, and time synchronization error. The derivation and proof of this bound are provided in Appendix 10.

4

Experiments

We evaluate 16 state-of-the-art LLMs spanning open-source and closed-source models across generalpurpose and reasoning-specialized architectures. Table 6 in Appendix 12 provides the full list of models with their model IDs and organizations. All models are accessed via their respective official vendor APIs with no fine-tuning applied: GPT (OpenAI API), DeepSeek (DeepSeek API), Mistral (Mistral AI API), Claude (Anthropic API), Gemini (Google AI API), Grok (xAI API), and Llama and Qwen (Hugging Face inference router). All client-side operations, including prompt construction, API handling, response parsing, and metric computation, are performed on a standard workstation. To assess repeatability and stochasticity, each MCQA and open-ended question is evaluated three times under two temperature settings: deterministic (T = 0.0) and stochastic (T = 0.7). Since TSN is widely used in safety-critical domains, deterministic responses are essential, as non-determinism would undermine the reliability of LLM-based TSN reasoning. For models that do not expose a temperature parameter, evaluations use the vendor default configuration, as noted in Table 4. Full cost and latency details are provided in Appendix 12, Table 8. 4.1

MCQA Evaluation

Contamination Analysis: Since the MCQs were generated using models from families included in the evaluation, as shown in Table 1, contamination is a potential concern. We therefore separate the evaluated models into generator families (Claude, GPT, Llama) and non-generator families (all remaining models) and compare their average MCQA accuracy. Generator-family models achieve an average accuracy of 88.8%, whereas non-generator-family models achieve 91.0%. The generatorfamily models do not perform better than the non-generator-family models, so we do not observe evidence of a systematic advantage. This analysis does not rule out all possible contamination pathways, but it addresses this specific concern. The open-ended timing tasks are less likely to be affected because their topology, flow, and routing inputs were constructed specifically for TSNBench. Evaluation Metrics: Model performance on the MCQA dataset is measured using accuracy, defined as the percentage of correctly answered questions out of 939, averaged across three runs. We additionally report Expected Calibration Error (ECE) [Pavlovic, 2025] and Brier score [Hoessly, 7

2026] to evaluate the alignment between the model’s expressed confidence and its actual correctness. Calibration is particularly critical in safety-critical domains such as TSN, where high-confidence incorrect answers may lead to misleading configuration decisions, deadline violations, or network instability in industrial and automotive systems. We therefore also evaluate the Confidently Wrong (CW) rate to determine the fraction of incorrect answers where the model expresses high confidence (≥0.8). All calibration metrics are computed on the full 939-MCQA dataset across three runs per model. Results and Discussion: Table 4 reports accuracy, average (avg.) consistency, calibration, and average latency for all 16 models. The top performers are Claude Sonnet 4.5 (95.3%) and GPT-5 (95.0%), with Claude Sonnet 4.5 also achieving the lowest Brier score (0.0429), indicating strong accuracy and calibration. Llama 3.2 1B achieves the lowest accuracy (67.4%), consistent with its substantially smaller parameter count compared with the other models. A notable finding emerges from the reasoning models. Despite their stronger general reasoning capabilities, o3, GPT-5, and DeepSeek-V3.2 (Thinking) do not outperform the best non-reasoning models on MCQA, all scoring below Claude Sonnet 4.5. This suggests that TSN MCQA performance is primarily driven by domain knowledge rather than general reasoning, and that reasoning-specialized architectures offer limited advantage on declarative knowledge retrieval tasks. The calibration results reveal key differences across models. While most models are well-calibrated (ECE < 0.06), o3 has the highest ECE (0.1874) despite 94.7% accuracy, yet achieves the lowest CW rate (3.4%), rarely assigning high confidence to incorrect answers (refer to Figure 5). In contrast, many non-reasoning models have CW rates of 100%, assigning high confidence to incorrect answers. Mistral Medium 3.1 has the highest average confidence (0.9779) while maintaining 92.1% accuracy. All models have zero refusal rate, indicating that the MCQA dataset does not trigger response refusals. 4.2

Reliability Analysis

Figure 6 presents the reliability plot for all 16 evaluated models on the MCQA dataset. Each diagram shows the observed accuracy against the model’s expressed confidence, binned across the confidence range. A perfectly calibrated model would fall on the gray dashed diagonal line. This means the model’s confidence would perfectly align with its actual accuracy. The red shaded region indicates overconfidence, meaning the model’s confidence exceeds its actual accuracy. The green shaded region indicates underconfidence, meaning the model is more accurate than its expressed confidence suggests. In safety-critical TSN deployments, overconfidence is significantly more dangerous than underconfidence. A model that is incorrect but expresses high confidence may mislead a network engineer with an erroneous WCD estimate or misconfigured scheduling parameters. By contrast, an underconfident model that expresses uncertainty on correct answers prompts additional verification. The majority of the evaluated models sit in the high-confidence region (0.8 to 1.0) regardless of their actual accuracy. This indicates that the models tend to exhibit overconfidence. Grok 4.1 Fast (NR), Mistral Medium 3.1, Mistral Large 3, and Ministral 3 8B achieve CW rates of 100%, meaning all incorrect answers fall in the high-confidence range. This represents the most critical calibration behavior for TSN deployment. GPT-4o, Gemini 2.5 Flash, Llama 3.2 1B, and Qwen3 8B similarly exhibit CW rates exceeding 95%. A notable exception is o3, which is the only model that falls predominantly in the green underconfident zone, with a CW rate of just 3.4%. Despite having the highest ECE (0.1874) among all evaluated models, o3 is the safest among the evaluated models from a calibration perspective, as it rarely expresses high confidence on incorrect MCQA answers. This highlights an important distinction between aggregate calibration metrics and safety-relevant calibration behavior. DeepSeek-V3.2 (NT) achieves the lowest ECE (0.0105), suggesting strong overall calibration, yet maintains a CW rate of 96.4%, demonstrating that a low ECE does not guarantee safe and realistic confidence behavior. 4.3

Open-Ended Question Evaluation

Evaluation Metrics: For the open-ended questions, we report two widely used metrics: Mean Absolute Error (MAE) and Mean Absolute Percentage Error (MAPE), computed per test case (TC). 8

Accuracy

Accuracy

Accuracy

Accuracy

Gray Dashed diagonal = perfect calibration | Red shaded portion = overconfident | Green shaded portion = underconfident Grok 4.1 Fast (Default Temp.) 1.2 ECE=0.020 CW=99.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.7 0.8 0.9 1.0

Grok 4.1 Fast (NR) (Temp=0) 1.2 ECE=0.059 CW=100.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.80 0.85 0.90 0.95 1.00

DeepSeek-V3.2 (NT) (Temp=0) 1.2 ECE=0.010 CW=96.4% 1.0 0.8 0.6 0.4 0.2 0.0 0.7 0.8 0.9 1.0

GPT-4o (Temp=0) 1.2 ECE=0.038 CW=99.2% 1.0 0.8 0.6 0.4 0.2 0.0 0.7 0.8 0.9

GPT-4o mini (Temp=0) 1.2 ECE=0.017 CW=77.8% 1.0 0.8 0.6 0.4 0.2 0.0 0.7 0.8 0.9

Llama 3.3 70B (Temp=0) 1.2 ECE=0.035 CW=100.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.75 0.80 0.85 0.90 0.95 1.00

Mistral Medium 3.1 (Temp=0) 1.2 ECE=0.057 CW=100.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.90 0.95 1.00

Mistral Large 3 (Temp=0) 1.2 ECE=0.019 CW=100.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.75 0.80 0.85 0.90 0.95 1.00

Claude Sonnet 4.5 (Temp=0) 1.2 ECE=0.018 CW=86.6% 1.0 0.8 0.6 0.4 0.2 0.0 0.4 0.6 0.8 1.0

o3 (Default Temp.) 1.2 ECE=0.195 CW=3.4% 1.0 0.8 0.6 0.4 0.2 0.0 0.2 0.4 0.6 0.8

GPT-5 (Default Temp.) 1.2 ECE=0.072 CW=51.7% 1.0 0.8 0.6 0.4 0.2 0.0 0.2 0.4 0.6 0.8

DeepSeek-V3.2 (T) (Default Temp.) 1.2 ECE=0.024 CW=78.1% 1.0 0.8 0.6 0.4 0.2 0.0 0.4 0.6 0.8 1.0

Gemini 2.5 Flash (Temp=0) 1.2 ECE=0.079 CW=95.4% 1.0 0.8 0.6 0.4 0.2 0.0 0.0 0.2 0.4 0.6 0.8 1.0 Confidence

Llama 3.2 1B (Temp=0) 1.2 ECE=0.184 CW=99.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.2 0.4 0.6 0.8 Confidence

1.0

1.0

Qwen3 8B (Temp=0) 1.2 ECE=0.036 CW=100.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.7 0.8 0.9 Confidence

1.0

1.0

1.0

1.0

Ministral 3 8B (Temp=0) 1.2 ECE=0.097 CW=100.0% 1.0 0.8 0.6 0.4 0.2 0.0 0.7 0.8 0.9 1.0 Confidence

Figure 6: Reliability diagram representing the performance of all 16 state-of-the-art models evaluated on the MCQA dataset in TSNBench. The gray dashed line represents the perfect calibration where confidence is equal to the accuracy. A model which is 100% confident and has 100% accuracy will fall on this gray dashed line. The red shaded region represents the over-confidence of the model (model confidence exceeds the actual accuracy of the model), and the green shaded region represents the under-confidence of the model (actual accuracy of the model exceeds it’s given confidence). Each TC consists of n flows, denoted fi where i = 1 · · · n. For each flow fi , ŷT C x ,fi denotes the WCD predicted by the model for T C x and yT C x ,fi denotes the ground truth WCD of flow fi for TC number x, computed using a verified NC solver for CBS and using Eq. 22 for CQF. The MAE for each TC is defined as: n 1X MAET C x = |ŷT C x ,fi − yT C x ,fi |, (3) n i=1 where x denotes the TC index and x ∈ {1, . . . , 100}. The MAPE for each TC is defined as: n

MAPET C x =

1 X |ŷT C x ,fi − yT C x ,fi | × 100 n i=1 yT C x ,fi

(4)

The overall MAE and MAPE for a model are obtained by averaging across all 100 TCs: 100

100

1 X MAE = MAET C x , 100 x=1

1 X MAPE = MAPET C x 100 x=1 9

(5)

Table 3: Open-ended WCD estimation results for CBS and CQF across 100 test cases (TCs). MAE and MAPE are reported as mean ± standard deviation across all TCs. Median MAE is a robust measure against outlier TCs. A model is excluded (“–”) if: (i) it responded to fewer than 50 TCs, (ii) fewer than 80% of flows per TC received a WCD estimate, or (iii) all predicted WCD values were zero (trivial failure). CBS WCD Accuracy

Model Grok 4.1 Fast† Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-thinking) GPT-4o⋆ GPT-4o mini Llama 3.3 70B Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3† GPT-5† DeepSeek-V3.2 (Thinking)† Gemini 2.5 Flash Llama 3.2 1B Qwen3 8B§ Ministral 3 8B

CQF WCD Accuracy

MAE (µs) ↓

MAPE (%) ↓

Median (µs) ↓

MAE (µs) ↓

MAPE (%) ↓

Median (µs) ↓

174.6 ± 314.5 3246.3 ± 3762.8 – – 378.5 ± 189.7 313.3 ± 174.0 337.4 ± 225.7 240.1 ± 152.3 292.8 ± 173.9 262.5 ± 319.4 150.2 ± 198.2 – 552.7 ± 1821.0 – – 70287.8 ± 403636.4

127.9 ± 514.1 1102.5 ± 1112.9 – – 97.2 ± 13.9 84.2 ± 39.0 102.7 ± 93.8 62.7 ± 27.5 71.7 ± 15.3 84.4 ± 106.0 36.2 ± 36.4 – 277.8 ± 1417.7 – – 25498.1 ± 164932.0

107.0 2185.4 – – 337.7 273.3 258.0 205.2 264.4 142.4 92.4 – 225.6 – – 879.1

139.6 ± 90.0 168.3 ± 69.0 172.3 ± 72.8 82.2 ± 264.7 180.5 ± 82.0 160.9 ± 83.7 166.8 ± 92.6 59.5 ± 27.2 211.5 ± 1057.3 102.2 ± 76.0 107.0 ± 69.0 – 112.0 ± 89.9 – – 2918.5 ± 4017.6

83.2 ± 56.6 90.6 ± 23.2 94.1 ± 40.4 61.9 ± 193.6 99.2 ± 44.8 99.0 ± 77.2 96.2 ± 82.6 41.8 ± 27.1 116.2 ± 607.6 60.4 ± 46.0 62.4 ± 42.1 – 60.6 ± 46.5 – – 1705.5 ± 2382.1

137.7 167.7 178.2 1.2 175.3 147.0 141.0 50.0 60.7 81.1 107.0 – 92.5 – – 1046.0

† Temperature parameter not supported. Evaluated with default settings. ⋆ GPT-4o returned all-zero WCD values for all CBS test cases (trivial failure) but produced efficient WCD response for CQF. § Qwen3 8B evaluation failed due to repeated API timeout errors. No valid responses recorded for any TC. Llama 3.2 1B

provided WCDs for fewer than 5 TCs and furthermore provided insufficient valid response for both CBS and CQF. DeepSeek-V3.2 (Thinking) provided empty response for all TCs. lower ↓ is better for MAE, MAPE, and Median.

We additionally report the median MAE across TCs as a robust measure against outlier TCs. Further example and details on the evaluation metrics are provided in Appendix 13 and Table 9.

MCQA Accuracy Ministral 3 8B Qwen3 8B Llama 3.2 1B Gemini 2.5 Flash DeepSeek-V3.2 (T) GPT-5 o3 Claude Sonnet 4.5 Mistral Large 3 Mistral Medium 3.1 Llama 3.3 GPT-4o mini GPT-4o DeepSeek-V3.2 (NT) Grok 4.1 Fast (NR) Grok 4.1 Fast (R)

CBS vs CQF - Per-Flow MAE (µs) | One-Switch Topology

87% 84% 67% 90% 95% 95% 95% 95% 93% 92% 89% 88% 92% 94% 92% 93%

0 20 40 MCQA (%)

60

80 100 120 0

higher is better

CBS 1000

lower is better

2000

MAE (µs)

3000

4000

CQF 5000

Figure 7: Performance comparison across MCQA and open-ended WCD computation for all 16 evaluated models in TSNBench. (Left) MCQA accuracy (%) per model. (Right) Per-TC MAE distribution (in µs) for CBS and CQF open-ended questions, shown as box plots over One-Switch topology test cases. Results and Discussion: Table 3 presents the WCD computation results for both CBS and CQF across all 100 TCs. The central finding is a striking dissociation between MCQA accuracy and computational reasoning performance. Models that achieve above 90% accuracy on MCQA still fail substantially on open-ended WCD computation, with the best-performing model, GPT-5, achieving a median MAE of 92.4 µs on CBS, which is concerning because industrial TSN traffic can have strict timing requirements [Ekrad et al., 2025]. Detailed per-TC results are provided in Appendix 13. For CBS, most models produce large errors, with many exceeding 200 µs MAE and 70% MAPE. Several models exhibit distinct failure modes. On CBS, Llama 3.2 1B responds to fewer than 50 evaluated TCs, returning all-zero WCD values for few TCs and partially incorrect values for some 10

TCs, with incomplete flow coverage in all responses. Grok 4.1 Fast (Reasoning) returns truncated JSON, providing flow profile metadata but no WCD values, suggesting that the model hit an output length limit. DeepSeek-V3.2 (Thinking) returns empty responses for more than 70 TCs across both mechanisms. The NC-based computation required for CBS is mathematically demanding and complex, and the zero-shot setting reveals that most models cannot independently recall or correctly apply the full NC methodology. Among models that produce valid CBS responses, GPT-5 achieves the best performance (MAE 150.2 µs, MAPE 36.2%). Notably, OpenAI reasoning models and Grok 4.1 Fast perform better on CBS than non-reasoning models, with GPT-5 achieving substantially lower MAE than all non-reasoning models, suggesting that multi-step mathematical reasoning capability provides an advantage for NC-based WCD computation even when it does not improve MCQA accuracy. For CQF, performance is more varied, with median MAE ranging from 1.2 µs (GPT-4o) to 1,046 µs (Ministral 3 8B), and MAPE ranging from 41.8% (Mistral Large 3) to 1705.5% (Ministral 3 8B). GPT-4o achieves the lowest median MAE on CQF (1.2 µs, MAPE 61.9%) despite failing completely on CBS, suggesting it can correctly apply the CQF closed-form equation. Mistral Large 3 achieves the lowest MAPE on CQF (41.8%), indicating the most accurate relative WCD estimation across all evaluated models. Llama 3.2 1B exhibits the most severe hallucination failure, fabricating up to 1,013 flows (flow 0–1012) instead of predicting WCD for the actual flows (fewer than 30 flows per TC), and returning WCD = 0 for all. Qwen3 8B fails to produce any response for either CBS or CQF due to repeated API timeouts. Ministral 3 8B, despite being a small model, produces valid responses for both CBS and CQF but with large errors (MAPE 25498.1% for CBS and 1705.5% for CQF), demonstrating that context handling is necessary but not sufficient for correct WCD computation. Comparison across MCQA and open-ended questions: Figure 7 illustrates the performance differences between models across two evaluation types, MCQA and open-ended questions. The right-hand figure shows the MAE for a one-switch topology across different models, while the left-hand figure presents the MCQA accuracy. The MCQA accuracy remains high, above 80%, for all models except Llama 3.2 1B. However, the MAE is still significant for TSN flows with deadlines in the range of 1000 to 5000 µs. Figure 18 further presents the performance differences between models for MCQA and open-ended questions in a ring topology.

5

Conclusion

We present TSNBench, the first benchmark for evaluating LLM proficiency in Time-Sensitive Networking (TSN), comprising 939 expert-validated multiple-choice questions (MCQs) and 100 open-ended questions per mechanism for Credit-Based Shaper (CBS) and Cyclic Queuing and Forwarding (CQF). The ground truth WCD values are computed using a verified Network Calculus (NC) solver for CBS and closed-form mathematical upper bounds for CQF. We evaluate 16 LLMs and find that models achieve 67-95% MCQA accuracy yet fail substantially on open-ended WCD computation, with the best model (GPT-5) still achieving a Mean Absolute Percentage Error (MAPE) of 36.2% on CBS. Despite CBS being extensively researched and an older mechanism, models cannot correctly apply NC, whereas CQF, with its simpler closed-form equation, is handled more successfully, confirming that WCD computation performance is governed by mathematical complexity rather than mechanism maturity. TSNBench demonstrates that MCQ benchmarks substantially overestimate LLM capability in safety-critical domains. Limitations and Future Directions: TSNBench has three primary limitations. First, the MCQA dataset is generated from open-access research papers, limiting coverage of certain mechanisms. Second, the open-ended evaluation covers only CBS and CQF. Extending to TAS is a natural next step, though its NP-hard gate control list (GCL) synthesis problem poses additional challenges beyond CBS and CQF. Third, the open-ended tasks evaluate standalone zero-shot model behavior under a closed-book prompt and should not be interpreted as a recommended deployment workflow for safety-critical TSN systems. Evaluating LLMs in settings where they produce checkable artifacts verified by deterministic analysis tools, and evaluating whether providing NC equations in the prompt improves WCD computation accuracy, are important directions for future versions of TSNBench.

References IEEE Standard for Local and Metropolitan Area Networks - Virtual Bridged Local Area Networks Amendment 12:

11

Forwarding and Queuing Enhancements for Time-Sensitive Streams. IEEE Std 802.1Qav-2009 (Amendment to IEEE Std 802.1Q-2005), pages C1–72, 2010. doi: 10.1109/IEEESTD.2009.5375704. IEEE Standard for Local and metropolitan area networks–Bridges and Bridged Networks–Amendment 29: Cyclic Queuing and Forwarding. IEEE 802.1Qch-2017 (Amendment to IEEE Std 802.1Q-2014 as amended by IEEE Std 802.1Qca-2015, IEEE Std 802.1Qcd(TM)-2015, IEEE Std 802.1Q-2014/Cor 1-2015, IEEE Std 802.1Qbv-2015, IEEE Std 802.1Qbu-2016, IEEE Std 802.1Qbz-2016, and IEEE Std 802.1Qci-2017), pages 1–30, 2017. doi: 10.1109/IEEESTD.2017.7961303. IEEE Standard for Local and Metropolitan Area Network–Bridges and Bridged Networks. IEEE Std 802.1Q-2018 (Revision of IEEE Std 802.1Q-2014), pages 1–1993, 2018. doi: 10.1109/IEEESTD.2018.8403927. A Ademaj, D Puffer, D Bruckner, G Ditzel, L Leurs, MP Stanica, P Didier, R Hummen, R Blair, and T Enzinger. Industrial automation traffic types and their mapping to QoS/TSN mechanisms. TSN mechanisms, 3, 2019. Muhammad Adil, Tie Qiu, Xiaobo Zhou, Danish Javeed, Zhenrui Cao, and Dapeng Oliver Wu. Integrated 5G and Time Sensitive Networking for Emerging Applications: A Survey of Advancements, Challenges, and Future Directions. IEEE Communications Surveys & Tutorials, 28:4016–4050, 2026. doi: 10.1109/COMST. 2025.3632286. Konstantinos Alexandris, Paul Pop, and Tongtong Wang. Configuration and Evaluation of Multi-CQF Shapers in IEEE 802.1 Time-Sensitive Networking (TSN). IEEE Access, 10:109068–109081, 2022. doi: 10.1109/ ACCESS.2022.3214007. Jeffrey G. Andrews, Stefano Buzzi, Wan Choi, Stephen V. Hanly, Angel Lozano, Anthony C. K. Soong, and Jianzhong Charlie Zhang. What will 5g be? IEEE Journal on Selected Areas in Communications, 32(6): 1065–1082, 2014. doi: 10.1109/JSAC.2014.2328098. Dietmar Bruckner, Marius-Petru Stănică, Richard Blair, Sebastian Schriegel, Stephan Kehrer, Maik Seewald, and Thilo Sauter. An introduction to OPC UA TSN for industrial communication systems. Proceedings of the IEEE, 107(6):1121–1131, 2019. doi: 10.1109/JPROC.2018.2888703. Daniel Bujosa, Mohammad Ashjaei, Alessandro V Papadopoulos, Thomas Nolte, and Julián Proenza. HERMES: Heuristic multi-queue scheduler for TSN time-triggered traffic with zero reception jitter capabilities. In Proc. RTNS, 2022. doi: 10.1145/3534879.3534906. Daniel Bujosa Mateu. Improved Configuration and Analysis Solutions for Time-Sensitive Networks with Support for Legacy Systems. Malardalen University (Sweden), 2024. Martin Böhm and Diederich Wermser. Multi-domain time-sensitive networks—control plane mechanisms for dynamic inter-domain stream configuration. Electronics, 10(20), 2021. ISSN 2079-9292. doi: 10.3390/ electronics10202477. Silviu S. Craciunas, Ramon Serna Oliver, Martin Chmelík, and Wilfried Steiner. Scheduling Real-Time Communication in IEEE 802.1Qbv Time Sensitive Networks. In Proceedings of the 24th International Conference on Real-Time Networks and Systems, RTNS ’16, page 183–192, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 9781450347877. doi: 10.1145/2997465.2997470. URL https://doi.org/10.1145/2997465.2997470. Rubi Debnath, Mustafa Selman Akinci, Devika Ajith, and Sebastian Steinhorst. 5GTQ: QoS-Aware 5G-TSN Simulation Framework. In 2023 IEEE 98th Vehicular Technology Conference (VTC2023-Fall), pages 1–7, 2023a. doi: 10.1109/VTC2023-Fall60731.2023.10333533. Rubi Debnath, Philipp Hortig, Luxi Zhao, and Sebastian Steinhorst. Advanced Modeling and Analysis of Individual and Combined TSN Shapers in OMNeT++. In 2023 IEEE 29th International Conference on Embedded and Real-Time Computing Systems and Applications (RTCSA), pages 176–185, 2023b. doi: 10.1109/RTCSA58653.2023.00029. Rubi Debnath, Philipp Hortig, Luxi Zhao, and Sebastian Steinhorst. Quantifying the Impact of Frame Preemption on Combined TSN Shapers. In NOMS 2024-2024 IEEE Network Operations and Management Symposium, pages 1–9, 2024. doi: 10.1109/NOMS59830.2024.10575564. Rubi Debnath, Mohammadreza Barzegaran, and Sebastian Steinhorst. Toward an optimized multi-cyclic queuing and forwarding in time-sensitive networking with time injection. IEEE Internet of Things Journal, 12(20): 43034–43051, 2025a. doi: 10.1109/JIOT.2025.3597560. Rubi Debnath, Luxi Zhao, Mohammadreza Barzegaran, and Sebastian Steinhorst. CyclicSim: Comprehensive Evaluation of Cyclic Shapers in Time-Sensitive Networking. In 2025 IEEE 22nd Consumer Communications & Networking Conference (CCNC), pages 01–09, 2025b. doi: 10.1109/CCNC54725.2025.10975975.

12

Rubi Debnath, Luxi Zhao, and Sebastian Steinhorst. Learning-Based Traffic Classification for Mixed-Critical Flows in Time-Sensitive Networking. In ICC 2025 - IEEE International Conference on Communications, pages 5926–5932, 2025c. doi: 10.1109/ICC52391.2025.11161468. Kasra Ekrad, Inés Alvarez Vadillo, Bjarne Johansson, Saad Mubeen, and Mohammad Ashjaei. A Methodology to Map Industrial Automation Traffic to TSN Traffic Classes. In 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA), pages 1–8, 2025. doi: 10.1109/ETFA65518.2025. 11205571. Leonard Elliott. Time-sensitive networking (tsn) in military ground vehicle architectures. In 2024 NDIA Michigan Chapter Ground Vehicle Systems Engineering and Technology Symposium. National Defense Industrial Association Michigan Chapter, August 2023. doi: https://doi.org/10.4271/2024-01-4122. URL https://doi.org/10.4271/2024-01-4122. Mohamed Amine Ferrag, Abderrahmane Lakas, and Mérouane Debbah. 6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning With Foundation Models in AI-Native 6G Networks. IEEE Open Journal of the Communications Society, 7:3305–3330, 2026. doi: 10.1109/OJCOMS.2026. 3680457. Norman Finn. Introduction to Time-Sensitive Networking. IEEE Communications Standards Magazine, 2(2): 22–28, 2018. doi: 10.1109/MCOMSTD.2018.1700076. Tiziana Fiori, Francesco Giacinto Lavacca, Francesco Valente, and Vincenzo Eramo. Proposal and Investigation of a Lite Time Sensitive Networking Solution for the Support of Real Time Services in Space Launcher Networks. IEEE Access, 12:10664–10680, 2024. doi: 10.1109/ACCESS.2024.3353466. Pranshav Gajjar, Cong Shen, and Vijay K Shah. Tele-LLM-hub: Building context-aware multi-agent LLM systems for telecom networks. In NeurIPS 2025 Workshop: AI and ML for Next-Generation Wireless Communications and Networking, 2025. URL https://openreview.net/forum?id=AencYkmJtl. Voica Gavriluţ and Paul Pop. Traffic-type Assignment for TSN-based Mixed-criticality Cyber-physical Systems. 4(2), January 2020. ISSN 2378-962X. doi: 10.1145/3371708. URL https://doi.org/10.1145/3371708. Voica Gavriluţ, Luxi Zhao, Michael L. Raagaard, and Paul Pop. Avb-aware routing and scheduling of timetriggered traffic for tsn. IEEE Access, 6:75229–75243, 2018. doi: 10.1109/ACCESS.2018.2883644. Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Aditya K, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan Choi, Kevin Tobia, Margaret Hagan, Megan Ma, Michael Livermore, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Sharad Goel, Shang Gao, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, and Zehua Li. LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 44123–44279. Curran Associates, Inc., 2023. URL https://proceedings.neurips.cc/paper_files/paper/2023/file/ 89e44582fd28ddfea1ea4dcb0ebbf4b0-Paper-Datasets_and_Benchmarks.pdf. Xingang Guo, Yaxin Li, XiangYi Kong, YILAN JIANG, Xiayu Zhao, Zhihua Gong, Yufan Zhang, Daixuan Li, Tianle Sang, Beixiao Zhu, Gregory Jun, Yingbing Huang, Yiqi Liu, Yuqi Xue, Rahul Dev Kundu, Qi Jian Lim, Yizhou Zhao, Luke Alexander Granger, Mohamed Badr Younis, Darioush Keivan, Nippun Sabharwal, Shreyanka Sinha, Prakhar Agarwal, Kojo Vandyck, Hanlin Mai, Zichen Wang, Aditya Venkatesh, Ayush Barik, Jiankun Yang, Chongying Yue, Jingjie He, Libin Wang, Licheng Xu, Hao Chen, Jinwen Wang, Liujun Xu, Rushabh Shetty, Ziheng Guo, Dahui Song, Manvi Jha, Weijie Liang, Weiman Yan, Bryan Zhang, Sahil Bhandary Karnoor, Jialiang Zhang, Rutva Pandya, Xinyi Gong, Mithesh Ballae Ganesh, Feize Shi, Ruiling Xu, Yifan Zhang, Yanfeng Ouyang, Lianhui Qin, Elyse Rosenbaum, Corey Snyder, Peter Seiler, Geir Dullerud, Xiaojia Shelly Zhang, Zuofu Cheng, Pavan Kumar Hanumolu, Jian Huang, Mayank Kulkarni, Mahdi Namazifar, Huan Zhang, and Bin Hu. Toward engineering AGI: Benchmarking the engineering design capabilities of LLMs. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URL https://openreview.net/forum?id=Wmsnx7EPel. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=d7KBjmI3GmQ. Linard Hoessly. On misconceptions about the brier score in binary prediction models. Glob. Epidemiol., 11 (100242):100242, June 2026.

13

Tianyu Hua, Harper Hua, Violet Xiang, Benjamin Klieger, Sang T. Truong, Weixin Liang, Fan-Yun Sun, and Nick Haber. Researchcodebench: Benchmarking LLMs on implementing novel machine learning research code. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URL https://openreview.net/forum?id=3k70Vt0YFS. Dong Huang, Yuhao Qing, Weiyi Shang, Heming Cui, and Jie M. Zhang. EffiBench: Benchmarking the Efficiency of Automatically Generated Code. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 11506–11544. Curran Associates, Inc., 2024. doi: 10. 52202/079017-0367. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ 15807b6e09d691fe5e96cdecde6d7b80-Paper-Datasets_and_Benchmarks_Track.pdf. Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. Benchmarking large language models as AI research agents. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. URL https:// openreview.net/forum?id=kXlTY0BmK3. Jason J Jackson, Terry Huang, Henry Velasquez, Kevin Zhu, and Sunishchal Dev. Predicting Emergent Software Engineering Capabilities by Fine-tuning. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025. URL https://openreview.net/forum? id=EwchHtwavV. Sunjun Kweon, Jiyoun Kim, Heeyoung Kwak, Dongchul Cha, Hangyul Yoon, Kwanghyun Kim, Jeewon Yang, Seunghyun Won, and Edward Choi. EHRNoteQA: An LLM Benchmark for RealWorld Clinical Practice Using Discharge Summaries. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 124575–124611. Curran Associates, Inc., 2024. doi: 10. 52202/079017-3958. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ e15c4afff22f12c4986c1fcb4e941e03-Paper-Datasets_and_Benchmarks_Track.pdf. Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S. Ilgen, Emma Pierson, Pang Wei Koh, and Yulia Tsvetkov. MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 28858–28888. Curran Associates, Inc., 2024. doi: 10.52202/079017-0908. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ 32b80425554e081204e5988ab1c97e9a-Paper-Conference.pdf. Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, LEI ZHU, and Michael Lingzhi Li. Benchmarking Large Language Models on CMExam - A comprehensive Chinese Medical Exam Dataset. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 52430–52452. Curran Associates, Inc., 2023. URL https://proceedings.neurips.cc/paper_files/paper/2023/ file/a48ad12d588c597f4725a8b84af647b5-Paper-Datasets_and_Benchmarks.pdf. Yongsen Ma, Gang Zhou, and Shuangquan Wang. WiFi Sensing with Channel State Information: A Survey. ACM Comput. Surv., 52(3), June 2019. ISSN 0360-0300. doi: 10.1145/3310194. URL https://doi.org/ 10.1145/3310194. Ali Maatouk, Fadhel Ayed, Nicola Piovesan, Antonio De Domenico, Merouane Debbah, and Zhi-Quan Luo. Teleqna: A benchmark dataset to assess large language models telecommunications knowledge. IEEE Network, 40(2):253–260, 2026. doi: 10.1109/MNET.2025.3576035. Ahmed Nasrallah, Akhilesh S. Thyagaturu, Ziyad Alharbi, Cuixiang Wang, Xing Shao, Martin Reisslein, and Hesham Elbakoury. Performance Comparison of IEEE 802.1 TSN Time Aware Shaper (TAS) and Asynchronous Traffic Shaper (ATS). IEEE Access, 7:44165–44181, 2019. doi: 10.1109/ACCESS.2019. 2908613. Giwa Oluwaseyi, Michael Adewole, Tobi Awodumila, and Pelumi Aderinto. The LLM as a network operator: A vision for generative AI in the 6g radio access network. In NeurIPS 2025 Workshop: AI and ML for NextGeneration Wireless Communications and Networking, 2025. URL https://openreview.net/forum? id=81mgAfsFJv. Maja Pavlovic. Understanding Model Calibration - A gentle introduction and visual exploration of calibration and the expected calibration error (ECE). In The Fourth Blogpost Track at ICLR 2025, 2025. URL https://openreview.net/forum?id=BxBeCjQd2y. Yifei Peng, Boxin Shi, Tigang Jiang, Xiaodong Tu, Du Xu, and Kun Hua. A Survey on In-Vehicle Time-Sensitive Networking. IEEE Internet of Things Journal, 10(16):14375–14396, 2023. doi: 10.1109/JIOT.2023.3264909.

14

Long Phan, Alice Gatti, Nathaniel Li, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Dan Hendrycks, Ziwen Han, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Aakaash Nattanmai, Gordon McKellips, Anish Cheraku, Asim Suhail, Ethan Luo, Marvin Deng, Jason Luo, Ashley Zhang, Kavin Jindel, Jay Paek, Kasper Halevy, Allen Baranov, Michael Liu, Advaith Avadhanam, David Zhang, Vincent Cheng, Brad Ma, Evan Fu, Liam Do, Joshua Lass, Hubert Yang, Surya Sunkari, Vishruth Bharath, Violet Ai, James Leung, Rishit Agrawal, Alan Zhou, Kevin Chen, Tejas Kalpathi, Ziqi Xu, Gavin Wang, Tyler Xiao, Erik Maung, Sam Lee, Ryan Yang, Roy Yue, Ben Zhao, Julia Yoon, Xiangwan Sun, Aryan Singh, Clark Peng, Tyler Osbey, Taozhi Wang, Daryl Echeazu, Timothy Wu, Spandan Patel, Vidhi Kulkarni, Vijaykaarti Sundarapandiyan, Andrew Le, Zafir Nasim, Srikar Yalam, Ritesh Kasamsetty, Soham Samal, David Sun, Nihar Shah, Abhijeet Saha, Alex Zhang, Leon Nguyen, Laasya Nagumalli, Kaixin Wang, Aidan Wu, Anwith Telluri, Summer Yue, Alexandr Wang, Dmitry Dodonov, Tung Nguyen, Jaeho Lee, Daron Anderson, Mikhail Doroshenko, Alun Cennyth Stokes, Mobeen Mahmood, Oleksandr Pokutnyi, Oleg Iskra, Jessica P. Wang, John-Clark Levin, Mstyslav Kazakov, Fiona Feng, Steven Y. Feng, Haoran Zhao, Michael Yu, Varun Gangal, Chelsea Zou, Zihan Wang, Serguei Popov, Robert Gerbicz, Geoff Galgon, Johannes Schmitt, Will Yeadon, Yongki Lee, Scott Sauers, Alvaro Sanchez, Fabian Giska, Marc Roth, Søren Riis, Saiteja Utpala, Noah Burns, Gashaw M. Goshu, Mohinder Maheshbhai Naiya, Chidozie Agu, Zachary Giboney, Antrell Cheatom, Francesco Fournier-Facio, Sarah-Jane Crowson, Lennart Finke, Zerui Cheng, Jennifer Zampese, Ryan G. Hoerr, Mark Nandor, Hyunwoo Park, Tim Gehrunger, Jiaqi Cai, Ben McCarty, Alexis C. Garretson, Edwin Taylor, Damien Sileo, Qiuyu Ren, Usman Qazi, Lianghui Li, Jungbae Nam, John B. Wydallis, Pavel Arkhipov, Jack Wei Lun Shi, Aras Bacho, Chris G. Willcocks, Hangrui Cao, Sumeet Motwani, Emily de Oliveira Santos, Johannes Veith, Edward Vendrow, Doru Cojoc, Kengo Zenitani, Joshua Robinson, Longke Tang, Yuqi Li, Joshua Vendrow, Natanael Wildner Fraga, Vladyslav Kuchkin, Andrey Pupasov Maksimov, Pierre Marion, Denis Efremov, Jayson Lynch, Kaiqu Liang, Aleksandar Mikov, Andrew Gritsevskiy, Julien Guillod, Gözdenur Demir, Dakotah Martinez, Ben Pageler, Kevin Zhou, Saeed Soori, Ori Press, Henry Tang, Paolo Rissone, Sean R. Green, Lina Brüssel, Moon Twayana, Aymeric Dieuleveut, Joseph Marvin Imperial, Ameya Prabhu, Jinzhou Yang, Nick Crispino, Arun Rao, Dimitri Zvonkine, Gabriel Loiseau, Mikhail Kalinin, Marco Lukas, Ciprian Manolescu, Nate Stambaugh, Subrata Mishra, Tad Hogg, Carlo Bosio, Brian P. Coppola, Julian Salazar, Jaehyeok Jin, Rafael Sayous, Stefan Ivanov, Philippe Schwaller, Shaipranesh Senthilkumar, Andres M. Bran, Andres Algaba, Kelsey Van den Houte, Lynn Van Der Sypt, Brecht Verbeken, David Noever, Alexei Kopylov, Benjamin Myklebust, Bikun Li, Lisa Schut, Evgenii Zheltonozhskii, Qiaochu Yuan, Derek Lim, Richard Stanley, Tong Yang, John Maar, Julian Wykowski, Mart Oller, Anmol Sahu, Cesare Giulio Ardito, Yuzheng Hu, Ariel Ghislain Kemogne Kamdoum, Alvin Jin, Tobias Garcia Vilchis, Yuexuan Zu, Martin Lackner, James Koppel, Gongbo Sun, Daniil S. Antonenko, Steffi Chern, Bingchen Zhao, Pierrot Arsene, Joseph M. Cavanagh, Daofeng Li, Jiawei Shen, Donato Crisostomi, Wenjin Zhang, Ali Dehghan, Sergey Ivanov, David Perrella, Nurdin Kaparov, Allen Zang, Ilia Sucholutsky, Arina Kharlamova, Daniil Orel, Vladislav Poritski, Shalev BenDavid, Zachary Berger, Parker Whitfill, Michael Foster, Daniel Munro, Linh Ho, Shankar Sivarajan, Dan Bar Hava, Aleksey Kuchkin, David Holmes, Alexandra Rodriguez-Romero, Frank Sommerhage, Anji Zhang, Richard Moat, Keith Schneider, Zakayo Kazibwe, Don Clarke, Dae Hyun Kim, Felipe Meneguitti Dias, Sara Fish, Veit Elser, Tobias Kreiman, Victor Efren Guadarrama Vilchis, Immo Klose, Ujjwala Anantheswaran, Adam Zweiger, Kaivalya Rawal, Jeffery Li, Jeremy Nguyen, Nicolas Daans, Haline Heidinger, Maksim Radionov, Václav Rozhoň, Vincent Ginis, Christian Stump, Niv Cohen, Rafał Poświata, Josef Tkadlec, Alan Goldfarb, Chenguang Wang, Piotr Padlewski, Stanislaw Barzowski, Kyle Montgomery, Ryan Stendall, Jamie Tucker-Foltz, Jack Stade, T. Ryan Rogers, Tom Goertzen, Declan Grabb, Abhishek Shukla, Alan Givré, John Arnold Ambay, Archan Sen, Muhammad Fayez Aziz, Mark H. Inlow, Hao He, Ling Zhang, Younesse Kaddar, Ivar Ängquist, Yanxu Chen, Harrison K. Wang, Kalyan Ramakrishnan, Elliott Thornley, Antonio Terpin, Hailey Schoelkopf, Eric Zheng, Avishy Carmi, Ethan D. L. Brown, Kelin Zhu, Max Bartolo, Richard Wheeler, Martin Stehberger, Peter Bradshaw, JP Heimonen, Kaustubh Sridhar, Ido Akov, Jennifer Sandlin, Yury Makarychev, Joanna Tam, Hieu Hoang, David M. Cunningham, Vladimir Goryachev, Demosthenes Patramanis, Michael Krause, Andrew Redenti, David Aldous, Jesyin Lai, Shannon Coleman, Jiangnan Xu, Sangwon Lee, Ilias Magoulas, Sandy Zhao, Ning Tang, Michael K. Cohen, Orr Paradise, Jan Hendrik Kirchner, Maksym Ovchynnikov, Jason O. Matos, Adithya Shenoy, Michael Wang, Yuzhou Nie, Anna Sztyber-Betley, Paolo Faraboschi, Robin Riblet, Jonathan Crozier, Shiv Halasyamani, Shreyas Verma, Prashant Joshi, Eli Meril, Ziqiao Ma, Jérémy Andréoletti, Raghav Singhal, Jacob Platnick, Volodymyr Nevirkovets, Luke Basler, Alexander Ivanov, Seri Khoury, Nils Gustafsson, Marco Piccardo, Hamid Mostaghimi, Qijia Chen, Virendra Singh, Tran Quoc Khánh, Paul Rosu, Hannah Szlyk, Zachary Brown, Himanshu Narayan, Aline Menezes, Jonathan Roberts, William Alley, Kunyang Sun, Arkil Patel, Max Lamparth, Anka Reuel, Linwei Xin, Hanmeng Xu, Jacob Loader, Freddie Martin, Zixuan Wang, Andrea Achilleos, Thomas Preu, Tomek Korbak, Ida Bosio, Fereshteh Kazemi, Ziye Chen, Biró Bálint, Eve J. Y. Lo, Jiaqi Wang, Maria Inês S. Nunes, Jeremiah Milbauer, M. Saiful Bari, Zihao Wang, Behzad Ansarinejad, Yewen Sun, Stephane Durand, Hossam Elgnainy, Guillaume Douville, Daniel Tordera, George Balabanian, Hew Wolff, Lynna Kvistad, Hsiaoyun Milliron, Ahmad Sakor, Murat Eron, Andrew Favre, Shailesh Shah, Xiaoxiang Zhou, Firuz Kamalov, Sherwin Abdoli, Tim Santens, Shaul Barkan, Allison Tee, Robin Zhang, Alessandro Tomasiello, G. Bruno De Luca, Shi-Zhuo Looi, Vinh-Kha Le, Noam Kolt, Jiayi Pan, Emma Rodman, Jacob Drori, Carl J. Fossum, Niklas Muennighoff, Milind Jagota, Ronak Pradeep, Honglu Fan, Jonathan Eicher, Michael

15

Chen, Kushal Thaman, William Merrill, Moritz Firsching, Carter Harris, Stefan Ciobâcă, Jason Gross, Rohan Pandey, Ilya Gusev, Adam Jones, Shashank Agnihotri, Pavel Zhelnov, Mohammadreza Mofayezi, Alexander Piperski, David K. Zhang, Kostiantyn Dobarskyi, Roman Leventov, Ignat Soroko, Joshua Duersch, Vage Taamazyan, Andrew Ho, Wenjie Ma, William Held, Ruicheng Xian, Armel Randy Zebaze, Mohanad Mohamed, Julian Noah Leser, Michelle X. Yuan, Laila Yacar, Johannes Lengler, Katarzyna Olszewska, Claudio Di Fratta, Edson Oliveira, Joseph W. Jackson, Andy Zou, Muthu Chidambaram, Timothy Manik, Hector Haffenden, Dashiell Stander, Ali Dasouqi, Alexander Shen, Bita Golshani, David Stap, Egor Kretov, Mikalai Uzhou, Alina Borisovna Zhidkovskaya, Nick Winter, Miguel Orbegozo Rodriguez, Robert Lauff, Dustin Wehr, Colin Tang, Zaki Hossain, Shaun Phillips, Fortuna Samuele, Fredrik Ekström, Angela Hammon, Oam Patel, Faraz Farhidi, George Medley, Forough Mohammadzadeh, Madellene Peñaflor, Haile Kassahun, Alena Friedrich, Rayner Hernandez Perez, Daniel Pyda, Taom Sakal, Omkar Dhamane, Ali Khajegili Mirabadi, Eric Hallman, Kenchi Okutsu, Mike Battaglia, Mohammad Maghsoudimehrabani, Alon Amit, Dave Hulbert, Roberto Pereira, Simon Weber, Handoko, Anton Peristyy, Stephen Malina, Mustafa Mehkary, Rami Aly, Frank Reidegeld, Anna-Katharina Dick, Cary Friday, Mukhwinder Singh, Hassan Shapourian, Wanyoung Kim, Mariana Costa, Hubeyb Gurdogan, Harsh Kumar, Chiara Ceconello, Chao Zhuang, Haon Park, Micah Carroll, Andrew R. Tawfeek, Stefan Steinerberger, Daattavya Aggarwal, Michael Kirchhof, Linjie Dai, Evan Kim, Johan Ferret, Jainam Shah, Yuzhou Wang, Minghao Yan, Krzysztof Burdzy, Lixin Zhang, Antonio Franca, Diana T. Pham, Kang Yong Loh, Joshua Robinson, Abram Jackson, Paolo Giordano, Philipp Petersen, Adrian Cosma, Jesus Colino, Colin White, Jacob Votava, Vladimir Vinnikov, Ethan Delaney, Petr Spelda, Vit Stritecky, Syed M. Shahid, Jean-Christophe Mourrat, Lavr Vetoshkin, Koen Sponselee, Renas Bacho, Zheng-Xin Yong, Florencia de la Rosa, Nathan Cho, Xiuyu Li, Guillaume Malod, Orion Weller, Guglielmo Albani, Leon Lang, Julien Laurendeau, Dmitry Kazakov, Fatimah Adesanya, Julien Portier, Lawrence Hollom, Victor Souza, Yuchen Anna Zhou, Julien Degorre, Yiğit Yaln, Gbenga Daniel Obikoya, Rai Michael Pokorny, Filippo Bigi, M. C. Boscá, Oleg Shumar, Kaniuar Bacho, Gabriel Recchia, Mara Popescu, Nikita Shulga, Ngefor Mildred Tanwie, Thomas C. H. Lux, Ben Rank, Colin Ni, Matthew Brooks, Alesia Yakimchyk, Huanxu Quinn Liu, Stefano Cavalleri, Olle Häggström, Emil Verkama, Joshua Newbould, Hans Gundlach, Leonor Brito-Santana, Brian Amaro, Vivek Vajipey, Rynaa Grover, Ting Wang, Yosi Kratish, Wen-Ding Li, Sivakanth Gopi, Andrea Caciolai, Christian Schroeder de Witt, Pablo Hernández-Cámara, Emanuele Rodolà, Jules Robins, Dominic Williamson, Brad Raynor, Hao Qi, Ben Segev, Jingxuan Fan, Sarah Martinson, Erik Y. Wang, Kaylie Hausknecht, Michael P. Brenner, Mao Mao, Christoph Demian, Peyman Kassani, Xinyu Zhang, David Avagian, Eshawn Jessica Scipio, Alon Ragoler, Justin Tan, Blake Sims, Rebeka Plecnik, Aaron Kirtland, Omer Faruk Bodur, D. P. Shinde, Yan Carlos Leyva Labrador, Zahra Adoul, Mohamed Zekry, Ali Karakoc, Tania C. B. Santos, Samir Shamseldeen, Loukmane Karim, Anna Liakhovitskaia, Nate Resman, Nicholas Farina, Juan Carlos Gonzalez, Gabe Maayan, Earth Anderson, Rodrigo De Oliveira Pena, Elizabeth Kelley, Hodjat Mariji, Rasoul Pouriamanesh, Wentao Wu, Ross Finocchio, Ismail Alarab, Joshua Cole, Danyelle Ferreira, Bryan Johnson, Mohammad Safdari, Liangti Dai, Siriphan Arthornthurasuk, Isaac C. McAlister, Alejandro José Moyano, Alexey Pronin, Jing Fan, Angel Ramirez-Trinidad, Yana Malysheva, Daphiny Pottmaier, Omid Taheri, Stanley Stepanic, Samuel Perry, Luke Askew, Raúl Adrián Huerta Rodrguez, Ali M. R. Minissi, Ricardo Lorena, Krishnamurthy Iyer, Arshad Anil Fasiludeen, Ronald Clark, Josh Ducey, Matheus Piza, Maja Somrak, Eric Vergo, Juehang Qin, Benjámin Borbás, Eric Chu, Jack Lindsey, Antoine Jallon, I. M. J. McInnis, Evan Chen, Avi Semler, Luk Gloor, Tej Shah, Marc Carauleanu, Pascal Lauer, Tran Duc Huy, Hossein Shahrtash, Emilien Duc, Lukas Lewark, Assaf Brown, Samuel Albanie, Brian Weber, Warren S. Vaz, Pierre Clavier, Yiyang Fan, Gabriel Poesia Reis e Silva, Long Tony Lian, Marcus Abramovitch, Xi Jiang, Sandra Mendoza, Murat Islam, Juan Gonzalez, Vasilios Mavroudis, Justin Xu, Pawan Kumar, Laxman Prasad Goswami, Daniel Bugas, Nasser Heydari, Ferenc Jeanplong, Thorben Jansen, Antonella Pinto, Archimedes Apronti, Abdallah Galal, Ng Ze-An, Ankit Singh, Tong Jiang, Joan of Arc Xavier, Kanu Priya Agarwal, Mohammed Berkani, Gang Zhang, Zhehang Du, Benedito Alves de Oliveira Junior, Dmitry Malishev, Nicolas Remy, Taylor D. Hartman, Tim Tarver, Stephen Mensah, Gautier Abou Loume, Wiktor Morak, Farzad Habibi, Sarah Hoback, Will Cai, Javier Gimenez, Roselynn Grace Montecillo, Jakub Łucki, Russell Campbell, Asankhaya Sharma, Khalida Meer, Shreen Gul, Daniel Espinosa Gonzalez, Xavier Alapont, Alex Hoover, Gunjan Chhablani, Freddie Vargus, Arunim Agarwal, Yibo Jiang, Deepakkumar Patil, David Outevsky, Kevin Joseph Scaria, Rajat Maheshwari, Abdelkader Dendane, Priti Shukla, Ashley Cartwright, Sergei Bogdanov, Niels Mündler, Sören Möller, Luca Arnaboldi, Kunvar Thaman, Muhammad Rehan Siddiqi, Prajvi Saxena, Himanshu Gupta, Tony Fruhauff, Glen Sherman, Mátyás Vincze, Siranut Usawasutsakorn, Dylan Ler, Anil Radhakrishnan, Innocent Enyekwe, Sk Md Salauddin, Jiang Muzhen, Aleksandr Maksapetyan, Vivien Rossbach, Chris Harjadi, Mohsen Bahaloohoreh, Claire Sparrow, Jasdeep Sidhu, Sam Ali, Song Bian, John Lai, Eric Singer, Justine Leon Uro, Greg Bateman, Mohamed Sayed, Ahmed Menshawy, Darling Duclosel, Dario Bezzi, Yashaswini Jain, Ashley Aaron, Murat Tiryakioglu, Sheeshram Siddh, Keith Krenek, Imad Ali Shah, Jun Jin, Scott Creighton, Denis Peskoff, Zienab EL-Wasif, Ragavendran P, Michael Richmond, Joseph McGowan, Tejal Patwardhan, Hao-Yu Sun, Ting Sun, Nikola Zubić, Samuele Sala, Stephen Ebert, Jean Kaddour, Manuel Schottdorf, Dianzhuo Wang, Gerol Petruzella, Alex Meiburg, Tilen Medved, Ali ElSheikh, S. Ashwin Hebbar, Lorenzo Vaquero, Xianjun Yang, Jason Poulos, Vilém Zouhar, Sergey Bogdanik, Mingfang Zhang, Jorge Sanz-Ros, David Anugraha, Yinwei Dai, Anh N. Nhu, Xue Wang, Ali Anil Demircali, Zhibai Jia, Yuyin Zhou, Juncheng Wu, Mike He, Nitin Chandok, Aarush Sinha, Gaoxiang Luo, Long Le, Mickaël Noyé, Michał Perełkiewicz, Ioannis Pantidis, Tianbo Qi, Soham Sachin Purohit, Letitia Parcalabescu,

16

Thai-Hoa Nguyen, Genta Indra Winata, Edoardo M. Ponti, Hanchen Li, Kaustubh Dhole, Jongee Park, Dario Abbondanza, Yuanli Wang, Anupam Nayak, Diogo M. Caetano, Antonio A. W. L. Wong, Maria del RioChanona, Dániel Kondor, Pieter Francois, Ed Chalstrey, Jakob Zsambok, Dan Hoyer, Jenny Reddish, Jakob Hauser, Francisco-Javier Rodrigo-Ginés, Suchandra Datta, Maxwell Shepherd, Thom Kamphuis, Qizheng Zhang, Hyunjun Kim, Ruiji Sun, Jianzhu Yao, Franck Dernoncourt, Satyapriya Krishna, Sina Rismanchian, Bonan Pu, Francesco Pinto, Yingheng Wang, Kumar Shridhar, Kalon J. Overholt, Glib Briia, Hieu Nguyen, David Quod Soler Bartomeu, Tony CY Pang, Adam Wecker, Yifan Xiong, Fanfei Li, Lukas S. Huber, Joshua Jaeger, Romano De Maddalena, Xing Han Lù, Yuhui Zhang, Claas Beger, Patrick Tser Jern Kon, Sean Li, Vivek Sanker, Ming Yin, Yihao Liang, Xinlu Zhang, Ankit Agrawal, Li S. Yifei, Zechen Zhang, Mu Cai, Yasin Sonmez, Costin Cozianu, Changhao Li, Alex Slen, Shoubin Yu, Hyun Kyu Park, Gabriele Sarti, Marcin Briański, Alessandro Stolfo, Truong An Nguyen, Mike Zhang, Yotam Perlitz, Jose Hernandez-Orallo, Runjia Li, Amin Shabani, Felix Juefei-Xu, Shikhar Dhingra, Orr Zohar, My Chiffon Nguyen, Alexander Pondaven, Abdurrahim Yilmaz, Xuandong Zhao, Chuanyang Jin, Muyan Jiang, Stefan Todoran, Xinyao Han, Jules Kreuer, Brian Rabern, Anna Plassart, Martino Maggetti, Luther Yap, Robert Geirhos, Jonathon Kean, Dingsu Wang, Sina Mollaei, Chenkai Sun, Yifan Yin, Shiqi Wang, Rui Li, Yaowen Chang, Anjiang Wei, Alice Bizeul, Xiaohan Wang, Alexandre Oliveira Arrais, Kushin Mukherjee, Jorge Chamorro-Padial, Jiachen Liu, Xingyu Qu, Junyi Guan, Adam Bouyamourn, Shuyu Wu, Martyna Plomecka, Junda Chen, Mengze Tang, Jiaqi Deng, Shreyas Subramanian, Haocheng Xi, Haoxuan Chen, Weizhi Zhang, Yinuo Ren, Haoqin Tu, Sejong Kim, Yushun Chen, Sara Vera Marjanović, Junwoo Ha, Grzegorz Luczyna, Jeff J. Ma, Zewen Shen, Dawn Song, Cedegao E. Zhang, Zhun Wang, Gaël Gendron, Yunze Xiao, Leo Smucker, Erica Weng, Kwok Hao Lee, Zhe Ye, Stefano Ermon, Ignacio D. Lopez-Miguel, Theo Knights, Anthony Gitter, Namkyu Park, Boyi Wei, Hongzheng Chen, Kunal Pai, Ahmed Elkhanany, Han Lin, Philipp D. Siedler, Jichao Fang, Ritwik Mishra, Károly Zsolnai-Fehér, Xilin Jiang, Shadab Khan, Jun Yuan, Rishab Kumar Jain, Xi Lin, Mike Peterson, Zhe Wang, Aditya Malusare, Maosen Tang, Isha Gupta, Ivan Fosin, Timothy Kang, Barbara Dworakowska, Kazuki Matsumoto, Guangyao Zheng, Gerben Sewuster, Jorge Pretel Villanueva, Ivan Rannev, Igor Chernyavsky, Jiale Chen, Deepayan Banik, Ben Racz, Wenchao Dong, Jianxin Wang, Laila Bashmal, Duarte V. Gonçalves, Wei Hu, Kaushik Bar, Ondrej Bohdal, Atharv Singh Patlan, Shehzaad Dhuliawala, Caroline Geirhos, Julien Wist, Yuval Kansal, Bingsen Chen, Kutay Tire, Atak Talay Yücel, Brandon Christof, Veerupaksh Singla, Zijian Song, Sanxing Chen, Jiaxin Ge, Kaustubh Ponkshe, Isaac Park, Tianneng Shi, Martin Q. Ma, Joshua Mak, Sherwin Lai, Antoine Moulin, Zhuo Cheng, Zhanda Zhu, Ziyi Zhang, Vaidehi Patil, Ketan Jha, Qiutong Men, Jiaxuan Wu, Tianchi Zhang, Bruno Hebling Vieira, Alham Fikri Aji, Jae-Won Chung, Mohammed Mahfoud, Ha Thi Hoang, Marc Sperzel, Wei Hao, Kristof Meding, Sihan Xu, Vassilis Kostakos, Davide Manini, Yueying Liu, Christopher Toukmaji, Eunmi Yu, Arif Engin Demircali, Zhiyi Sun, Ivan Dewerpe, Hongsen Qin, Roman Pflugfelder, James Bailey, Johnathan Morris, Ville Heilala, Sybille Rosset, Zishun Yu, Peter E. Chen, Woongyeong Yeo, Eeshaan Jain, Sreekar Chigurupati, Julia Chernyavsky, Sai Prajwal Reddy, Subhashini Venugopalan, Hunar Batra, Core Francisco Park, Hieu Tran, Guilherme Maximiano, Genghan Zhang, Yizhuo Liang, Hu Shiyu, Rongwu Xu, Rui Pan, Siddharth Suresh, Ziqi Liu, Samaksh Gulati, Songyang Zhang, Peter Turchin, Christopher W. Bartlett, Christopher R. Scotese, Phuong M. Cao, Ben Wu, Jacek Karwowski, and Davide Scaramuzza. A benchmark of expert-level academic questions to assess ai capabilities. Nature, 649(8099):1139–1146, January 2026. ISSN 1476-4687. doi: 10.1038/s41586-025-09962-4. URL http://dx.doi.org/10.1038/s41586-025-09962-4. Paul Pop, Michael Lander Raagaard, Silviu S. Craciunas, and Wilfried Steiner. Design optimization of cyberphysical distributed systems using IEEE time-sensitive networks (TSN). IET-CPS, 1(1):86–94, 2016. doi: https://doi.org/10.1049/iet-cps.2016.0021. Laria Reynolds and Kyle McDonell. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, CHI EA ’21, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450380959. doi: 10.1145/3411763.3451760. URL https://doi.org/10.1145/3411763.3451760. Walid Saad, Mehdi Bennis, and Mingzhe Chen. A Vision of 6G Wireless Systems: Applications, Trends, Technologies, and Open Research Problems. IEEE Network, 34(3):134–142, 2020. doi: 10.1109/MNET.001. 1900287. Jorge Sanchez-Garrido, Beatriz Aparicio, José Gabriel Ramírez, Rafael Rodriguez, Mariasole Melara, Lorenzo Cercós, Eduardo Ros, and Javier Diaz. Implementation of a Time-Sensitive Networking (TSN) Ethernet Bus for Microlaunchers. IEEE Transactions on Aerospace and Electronic Systems, 57(5):2743–2758, 2021. doi: 10.1109/TAES.2021.3061806. Ramon Serna Oliver, Silviu S. Craciunas, and Wilfried Steiner. IEEE 802.1Qbv Gate Control List Synthesis Using Array Theory Encoding. In 2018 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS), pages 13–24, 2018. doi: 10.1109/RTAS.2018.00008. Prakhar Sharma and Vinod Yegneswaran. PROSPER: Extracting Protocol Specifications Using Large Language Models. In Proceedings of the 22nd ACM Workshop on Hot Topics in Networks, HotNets ’23, page 41–47,

17

New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9798400704154. doi: 10.1145/ 3626111.3628205. URL https://doi.org/10.1145/3626111.3628205. Yongliang Shen, Kaitao Song, Xu Tan, Wenqi Zhang, Kan Ren, Siyu Yuan, Weiming Lu, Dongsheng Li, and Yueting Zhuang. TaskBench: Benchmarking Large Language Models for Task Automation. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 4540–4574. Curran Associates, Inc., 2024. doi: 10.52202/079017-0148. URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ 085185ea97db31ae6dcac7497616fd3e-Paper-Datasets_and_Benchmarks_Track.pdf. Johannes Specht and Soheil Samii. Urgency-Based Scheduler for Time-Sensitive Switched Ethernet Networks. In 2016 28th Euromicro Conference on Real-Time Systems (ECRTS), pages 75–85, 2016. doi: 10.1109/ ECRTS.2016.27. Liangtai Sun, Yang Han, Zihan Zhao, Da Ma, Zhennan Shen, Baocai Chen, Lu Chen, and Kai Yu. SciEval: a multi-level large language model evaluation benchmark for scientific research. In Proceedings of the ThirtyEighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’24/IAAI’24/EAAI’24. AAAI Press, 2024. ISBN 978-1-57735-887-9. doi: 10.1609/aaai.v38i17.29872. URL https://doi.org/10.1609/aaai.v38i17.29872. Changjie Wang, Mariano Scazzariello, Alireza Farshin, Simone Ferlin, Dejan Kostić, and Marco Chiesa. NetConfEval: Can LLMs Facilitate Network Configuration? Proc. ACM Netw., 2(CoNEXT2), June 2024a. doi: 10.1145/3656296. URL https://doi.org/10.1145/3656296. Xiaolong Wang, Haipeng Yao, Tianle Mai, Zehui Xiong, Fu Wang, and Yunjie Liu. Joint Routing and Scheduling With Cyclic Queuing and Forwarding for Time-Sensitive Networks. IEEE Transactions on Vehicular Technology, 72(3):3793–3804, 2023. doi: 10.1109/TVT.2022.3216958. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024b. URL https://openreview.net/forum?id=y10DM6R2r3. Stefan Windmann, Janis Albrecht, Maxim Friesen, and Jürgen Jasperneite. NetPilot - Towards LLM-Assisted Configuration of Hybrid TSN/5G Networks. In 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA), pages 1–4, 2025. doi: 10.1109/ETFA65518.2025.11205555. Jiacheng Xie, Yang Yu, Ziyang Zhang, Shuai Zeng, Jiaxuan He, Ayush Vasireddy, Xiaoting tang, Congyu Guo, Lening Zhao, Congcong Jing, Guanghui An, and Dong Xu. TCM-ladder: A benchmark for multimodal question answering on traditional chinese medicine. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URL https://openreview.net/ forum?id=ZDrT1eG54T. Jinli Yan, Wei Quan, Xuyan Jiang, and Zhigang Sun. Injection time planning: Making cqf practical in timesensitive networking. In IEEE INFOCOM 2020 - IEEE Conference on Computer Communications, pages 616–625, 2020. doi: 10.1109/INFOCOM41043.2020.9155434. Jialin Yang, Dongfu Jiang, Tony He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Zhiheng Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, and Wenhu Chen. Structeval: Benchmarking LLMs’ capabilities to generate structural outputs. Transactions on Machine Learning Research, 2026. ISSN 2835-8856. URL https://openreview.net/forum?id=buDwV7LUA7. J2C Certification. Kouros Zanbouri, Md. Noor-A-Rahim, Jobish John, Cormac J. Sreenan, H. Vincent Poor, and Dirk Pesch. A Comprehensive Survey of Wireless Time-Sensitive Networking (TSN): Architecture, Technologies, Applications, and Open Issues. IEEE Communications Surveys & Tutorials, 27(4):2129–2155, 2025. doi: 10.1109/COMST.2024.3486618. Tianyu Zhang, Gang Wang, Chuanyu Xue, Jiachen Wang, Mark Nixon, and Song Han. Time-Sensitive Networking (TSN) for Industrial Automation: Current Advances and Future Directions. ACM Comput. Surv., 57(2), October 2024. ISSN 0360-0300. doi: 10.1145/3695248. URL https://doi.org/10.1145/ 3695248. Luxi Zhao, Paul Pop, Zhong Zheng, and Qiao Li. Timing Analysis of AVB Traffic in TSN Networks Using Network Calculus. In 2018 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS), pages 25–36, 2018. doi: 10.1109/RTAS.2018.00009.

18

Luxi Zhao, Paul Pop, Zhong Zheng, Hugo Daigmorte, and Marc Boyer. Latency Analysis of Multiple Classes of AVB Traffic in TSN With Standard Credit Behavior Using Network Calculus. IEEE Transactions on Industrial Electronics, 68(10):10291–10302, 2021. doi: 10.1109/TIE.2020.3021638. Luxi Zhao, Paul Pop, and Sebastian Steinhorst. Quantitative Performance Comparison of Various Traffic Shapers in Time-Sensitive Networking. IEEE Transactions on Network and Service Management, 19(3):2899–2928, 2022. doi: 10.1109/TNSM.2022.3180160. Luxi Zhao, Yida Yan, and Xuan Zhou. Minimum Bandwidth Reservation for CBS in TSN With Real-Time QoS Guarantees. IEEE Transactions on Industrial Informatics, 20(4):6187–6198, 2024. doi: 10.1109/TII.2023. 3342466.

19

6

Limitations and Broader Impact

6.1

Limitations

While TSNBench fills a significant research gap and proposes a step forward towards evaluating TSN capabilities in LLMs, it has several limitations: Dataset scope: TSNBench currently only covers CBS and CQF in open-ended questions. Evaluating other TSN mechanisms is necessary to fully cover the entire TSN mechanism. Prompt Design: TSNBench does not provide any mathematical equation to the model as input for NC WCD calculation for CBS or the upper bound delay calculation of CQF. MCQA Scope: MCQs are solely developed using published research papers and the IEEE standards are not used to generate the MCQs. Solving the license issue and utilizing standards to include MCQs using IEEE 802.1 standard will enhance the entire MCQA dataset. Topology coverage: TSNBench open-ended question currently covers three different topologies: one-switch, medium-mesh, and ring topology. Covering diverse topologies and flow parameters will present a comprehensive evaluation. 6.2

Improvement Strategies

To address the limitations of TSNBench, we propose the following additions and improvements in the future version of TSNBench. 1. Larger and more diverse dataset: Our current TSNBench dataset covers 100 TCs across three topology types. In future versions, we will include larger and more complex topologies with higher flow counts. As model performance improves, more complex open-ended evaluations should be integrated with complex topologies and combined TSN mechanisms. 2. Additional scheduling mechanisms: TSNBench currently evaluates CBS and CQF. Future versions should extend to TAS and ATS to cover a broader range of the TSN standard suite. 3. Updated MCQA: Our MCQA dataset was developed using open-source research documents. In future work, we will update the dataset with MCQAs formulated directly from TSN standards. 4. Fine-tuned and domain-adapted models. TSNBench currently evaluates general-purpose LLMs without any TSN-specific fine-tuning. Future versions should benchmark domainadapted models trained on TSN standards and network calculus literature. 6.3

Broader Impact

TSNBench enables the real-time systems community and the machine learning community to objectively measure LLM performance and readiness for management and deployment assistance in safety-critical deterministic networks. By highlighting the critical aspects of TSN and the performance gap of the models between MCQA and computational reasoning, TSNBench alerts the incompetence of the models which may lead to misconfigurations and safety-critical issues. This benchmark provides a concrete direction to improve LLMs for deterministic networking. TSNBench further highlights the potential benefits of using LLMs thereby automating the management and deployment of TSN networks. Moreover, open-sourced ground truth WCD values computed by NC solvers provide a reliable resource for the entire community to further evaluate different benchmarking datasets. 6.4

Negative Impacts

While TSNBench is intended to advance research on LLM proficiency in TSN, we acknowledge the following potential negative impacts. Overreliance on model outputs: Models trained on the open-access dataset provided by TSNBench may achieve high accuracy on WCD analysis tasks, which could lead practitioners to deploy such models directly in real-world deployments without independent verification. Any WCD values or 20

network configuration decisions produced by an LLM should be verified using formally verified solvers and NC tools before real-world deployment. False confidence from MCQA performance: Our results demonstrate that strong MCQA performance does not transfer to open-ended WCD estimation. A practitioner or system engineer who evaluates an LLM solely on MCQA benchmarks may incorrectly conclude that the model is suitable for TSN configuration tasks, leading to unsafe deployments in systems where timing guarantees are required. Data contamination and benchmark overfitting: As TSNBench is released as an open-access dataset, future models may be trained directly on the benchmark questions, leading to inflated performance that does not reflect genuine TSN reasoning capability. We recommend that researchers introduce randomization in the test cases to prevent bias in results. Researchers should be cautious when interpreting results from models whose training data may overlap with the TSNBench dataset. Misuse of the dataset: The dataset can be used to train models to configure TSN networks. Owing to the safety-critical nature of TSN applications, such models could potentially be exploited by attackers to manipulate network configurations, introduce timing violations, or deliberately cause deadline misses in industrial and automotive systems.

7

Time-Sensitive Networking

Time-Sensitive Networking (TSN) [Finn, 2018] is a set of amendments and additions to the IEEE 802.1 standards that, since its inception in 2012, has become one of the most relevant technologies for enabling deterministic and real-time communications over Ethernet networks. TSN extends standard Ethernet by introducing mechanisms for bounded latency, low jitter, and high reliability, making it suitable for applications such as industrial automation, automotive systems, and professional audio-video networks. Figure 8 showcases a simple TSN network with flows.

safety-critical traffic non-critical traffic

TSN Sender

TSN Sender

TSN Receiver

TSN Receiver

TSN Sender

TSN Receiver TSN Switch

TSN Switch

TSN Sender

TSN Receiver

TSN Sender

TSN Receiver

Figure 8: A sample TSN network with TSN senders, receivers, and TSN switches in the network. TSN senders are sending mixed-critical including safety-critical and non-critical traffic to the TSN receivers. In TSN, communication between end-stations is based on the transmission of Ethernet frames across a network of interconnected Ethernet links and TSN switches. These switches, as well as the output ports of end-stations, implement a queuing architecture with up to eight First-In-First-Out (FIFO) queues, each associated with one of the eight traffic priorities defined in IEEE 802.1Q [802, 2018]. TSN is not just limited to wired domain. The growing necessity of deterministic communication has extended to wireless domain gaining a significant interest in wireless-TSN networks. Although TSN is fundamentally an IEEE 802.1 bridged Ethernet technology, wireless and 5G-TSN [Debnath et al., 2023a] integration requires additional adaptation or translation functions, together with timesynchronization mechanisms that preserve deterministic latency guarantees across heterogeneous network segments. We showcase a 5G-TSN system in Figure 9, where TSN senders are sending mixed criticality traffic types to wireless receiver nodes over a TSN switch and 5G system in the network. Some of the most commonly used abbreviations in TSN are given in Table 4. 21

safety-critical traffic non-critical traffic

TSN Sender

Robotic Arm

TSN Sender

Robot

TSN Sender

AGV TSN Switch

5G System

TSN Sender

Robot

TSN Sender

Robotic Arm

Figure 9: A sample wireless-TSN network with TSN senders, wireless receivers (such as robotic arm and automated guided vehicles (AGVs)), TSN switches, and 5G system in the network. TSN senders are sending mixed-critical including safety-critical and non-critical traffic to the wireless receivers. Frames are classified into traffic classes and assigned to egress queues based on their priority, with transmission selection typically governed by strict priority. Industrial TSN traffic is commonly categorized into traffic types such as isochronous traffic, cyclic-synchronous traffic, cyclic-asynchronous traffic, network-control traffic, alarms and events, configuration and diagnostics, and best-effort traffic [Ademaj et al., 2019]. These traffic types require different timing guarantees: safety-critical isochronous traffic is typically mapped to time-triggered (TT) traffic, requiring guaranteed latency and bounded jitter, and is commonly handled by time-triggered mechanisms such as the Time-Aware Shaper (TAS) [Craciunas et al., 2016, Serna Oliver et al., 2018]. In contrast, cyclic-synchronous or cyclic-asynchronous traffic that requires bounded end-to-end latency but less stringent jitter control is commonly mapped to AVB stream traffic and is often supported by the Credit-Based Shaper (CBS) [Zhao et al., 2018]. TSN also defines mechanisms such as Asynchronous Traffic Shaping (ATS) [Specht and Samii, 2016, Debnath et al., 2023b, Nasrallah et al., 2019], Frame preemption (FP) [Debnath et al., 2024], and Cyclic Queuing and Forwarding (CQF) [Wang et al., 2023, Debnath et al., 2025a, Yan et al., 2020] to provide deterministic communication under different traffic and deployment assumptions. These mechanisms regulate when and how frames are transmitted, allowing the network to provide guarantees such as bounded delay, jitter, and controlled bandwidth allocation. In the MCQA dataset of TSNBench, we covered the basics of different TSN mechanisms, including TAS, CBS, ATS, CQF, and CBS. The MCQAs are theoretical in nature and cover the basic understanding of the mechanisms without going into their mathematical or analytical details. In contrast, for the openended mechanisms, we evaluate the capability of the models to perform numerical analysis, formulate mathematical equations, and find the WCD values for the flows in the network. For this, we selected two TSN mechanisms: CBS and CQF. The WCD values of the flows using the CBS mechanism are calculated using NC analysis, which is mathematically complex. Therefore, we also evaluate the CQF mechanism as a simpler mechanism. The WCD values of the flows using the CQF mechanism can be directly calculated using the routing of the flow and the cycle duration. The detailed working mechanism and architecture of CQF and CBS are described in detail in the Appendix 9 and 10. The theory of NC is further explained along with the mathematical equations in Appendix 8.

8

Network Calculus Theory

Network Calculus (NC) is a theory for calculating worst-case bounds in communication networks based on min-plus algebra. Its basic paradigm involves two operators: convolution ⊗ (f ⊗g)(t) = inf {f (t−s)+g(s)},

(6)

(f ⊘g)(t) = sup{f (t+s)−g(s)}.

(7)

0≤s≤t

and deconvolution ⊘,

s≥0

22

Table 4: Abbreviations and mechanisms used in TSNBench. Keyword TSN TAS CBS ATS CQF NC WCD AVB TT

Abbreviations Time-Sensitive Networking Time Aware Shaper Credit-Based Shaper Asynchronous Traffic Shaper Cyclic Queuing and Forwarding Network Calculus Worst-Case Delay Audio Video Bridging Time-Triggered

Based on this algebra, the arrival curve and the service curve are constructed to describe the maximum arrival traffic data and the minimum service capability over any time interval, respectively. In the hybrid TSN/TAS+CBS architecture, the service for ET traffic is constrained not only by the bandwidth reservation, but also by high-priority TT traffic. We adopt the state-of-the-art network calculus model [Zhao et al., 2021, 2024] to ensure deadline guarantees for ET flows with an arbitrary number of SR classes in the TSN/TAS+CBS architecture. Since, in our open-end CBS questions, we do not have any TAS mechanism, we use the TSN/TAS+CBS architecture without the TAS mechanism in it with only CBS mechanism for the AVB flows in the network. As described in [Zhao et al., 2024], the service curve β(t) is for constraining the minimum service capabilities, satisfying R∗ (t) ≥ (R ⊗ β) (t). (8) ∗ The function R(t) (resp. R (t)) is the input (resp. output) cumulative function counting the total data bits of the flow that arrive at (resp. departure from) the server up to time t. A typical example of a service curve is the rate-latency form, βR,T (t) = R[t − T ]+

(9)

+

with the service rate R and latency T . The notation [x] equals x if x ≥ 0, and 0 otherwise. In the hybrid TSN/TAS+CBS architecture, the CBS service curve [Zhao et al., 2021] for the arbitrary SR Class Mi (i ∈ [1, NSR ]) with the impact of TT traffic at the output port h is, " # h,max + h c α (t) M h h i βM (t) = idSlM t − T AS − , (10) i i h C idSlM i ↑

where ch,max is the credit upper bound for SR Class Mi , Mi Pi−1 h ch,max = idSlM · Mi i

h,min h,max − l>i j=1 cMj , Pi−1 h j=1 idSlMj − C

(11)

h,max h,max h,max where l>i = maxj>i {lM , lBE } is the maximum frame size with priority lower than Class j h,max Mi at h, lM is the maximum frame size of Class Mi at h, and ch,min is the lower credit bound of Mi j SR Class Mi , h,max lM h i ch,min = sdSlM · . (12) Mi i C αTh AS (t) in Eq. (10) is the arrival curve of TT traffic scheduled by GCL.

The arrival curve α(t) is for constraining the arrival process of the flow, satisfying R(t) ≤ (R ⊗ α) (t).

(13)

A typical example of an arrival curve is the burst-rate form, α(t) = b + ρ · t, 23

(14)

for t > 0 and 0 otherwise, with the parameters b as the maximum burst tolerance and ρ as the long-term rate of the flow. For each ET flow f at its source ES h0 , the arrival curve can be modeled as, αfh0 (t) = bhf 0 + ρhf 0 t,

(15)

where bhf 0 = lf , and ρhf 0 = lf /Pf . The arrival curve of flow f at intermediate node h is the output − arrival curve of f departing from the server h ,

−

αfh (t) = αfh ⊘ δDh− (t),

(16)

f

−

where Dfh is the latency upper bound of flow f queuing at server h− , and δD (t) is the pure-delay function. The aggregate arrival curve for ET flows of SR Class Mi at h is obtained by summing the arrival curves of individual flows. It also incorporates the link shaping curve and the CBS shaping curve to improve the tightness of the analysis results. X X h−,h h−,h h αM (t) = αfh (t)∧σlink (t)∧σM (t), (17) i i h− ∈H f ∈F h−,h Mi

−

h ,h where x ∧ y = min{x, y}, σlink (t) is the link shaping curve from the preceding output h− to the current output port h: h−,h h−,h,max σlink (t) = Ct + lM , (18) i −

h ,h,max considering the packetization impact of the maximum frame size lM of flows with Class Mi i −

h ,h from h− to h. σM (t) is the CBS shaping curve of Class Mi from h− to h: i " # − − h−,max −chMi,min βTh AS (t) cMi h−,h h−,h,max h− σMi (t) = idSlMi t− + +lM , − i h C idSlMi

(19)

βTh AS (t) represents the minimum service supplied to TT traffic on the output port h. h With NC-based Total Flow Analysis (TFA), the worst-case delay upper bound Dfh for flow f ∈ FM i h at h equals the worst-case delay upper bound DM for all flows with the same priority Mi aggregating i at h,   h h h h h Dfh = DM = hDev(αM , βM ) = sup inf τ ≥ 0 | αM (t) ≤ βM (t+τ ) (20) i i i i i t≥0

h h where αM (t) is the arrival curve of aggregate flows of Class Mi from Eq. (17), and βM (t) is the i i service curve for Class Mi from Eq. (10). The upper bound of the worst-case end-to-end delay for the flow f is then obtained by summing the per-port latency bounds along its route.

9

Credit-Based Shaper

Credit-Based Shaper (CBS) is a TSN mechanism designed to prevent starvation of lower-priority traffic while guaranteeing a reserved portion of bandwidth for higher-priority queues, thereby providing reliability through bounded end-to-end delays. Traffic assigned to queues using CBS is typically referred to as Audio Video Bridging (AVB) traffic. Here, we build on the description from [Bujosa Mateu, 2024]. In CBS, each AVB queue is associated with a credit value. This credit increases over time when a frame is waiting to be transmitted or when the credit is negative, and decreases while a frame is being transmitted. Moreover, if the credit is positive and there are no AVB frames waiting to be transmitted, the credit is immediately reset to 0. The rates at which credit is increased and decreased are defined by the parameters idleSlope and sendSlope, respectively. Each queue implementing CBS is configured with its own idleSlope 24

Priority Filter

AVB Class N3

Queue 6

CBS

Switching Fabric

Queue 7

CBS

AVB Class N2

. . .

Queue 1

CBS

AVB Class N8

Transmission Selection Algorithm

Queue 8

CBS

AVB Class N1

Figure 10: A simple CBS mechanism with eight queues in the egress port of the switch with different AVB class mapped to different queues. and sendSlope values, which determine its allocated bandwidth share. In particular, the bandwidth reserved for a queue is expressed as Eq. (21). A queue is eligible for transmission only when its credit is zero or positive. Reserved BW =

idleSlope · BW idleSlope + sendSlope

(21)

Consider the example illustrated in Figure 11, which includes two AVB queues and one Best Effort (BE) queue. Frames 1 and 4 are assigned to the higher-priority AVB queue, while frames 2 and 3 belong to the lower-priority AVB queue and the BE queue, respectively. At time T0, both AVB queues are eligible for transmission. Due to strict priority scheduling, the higher-priority AVB queue (priority 2) is selected, and frame 1 is transmitted. During this transmission, its credit decreases, while the credit of the lower-priority AVB queue increases because it is waiting. At time T1, the higher-priority AVB queue has accumulated negative credit and is therefore no longer eligible for transmission. As a result, the lower-priority AVB queue is selected, and frame 2 is transmitted, even though a higher-priority frame (frame 4) is waiting. During this time, the lower-priority queue’s credit decreases, while the higher-priority queue’s credit recovers. By time T2, both AVB queues have negative credit, making them ineligible for transmission. Consequently, the BE queue is selected, and frame 3 is transmitted, despite the presence of a higher-priority AVB frame waiting. Finally, at time T3, the credit of the higher-priority AVB queue has recovered to zero, making it eligible again. Therefore, frame 4 is transmitted.

10

Cyclic Queuing and Forwarding (CQF)

Cyclic Queuing and Forwarding (CQF) [Debnath et al., 2025b] is a TSN shaping mechanism which uses a single cycle duration, denoted as T , across the entire network. T is the minimum scheduling unit where we put the TSN flows. Furthermore, T defines the granularity of the end-to-end delay of the flows in the network. The unit of T is in µs in TSNBench. In a TSN switch, every egress port in the network has eight queues. TSN flows are stored in the queues depending on its priority. In CQF, for each egress port, two queues are used: an even queue and an odd queue. Figure 14 shows the basic working diagram of CQF with two queues (even and odd). As shown in Figure 14, CQF works by employing two queues, let’s say, Q8 and Q7 for TT flows by operating them in a ping-pong manner where Q7 receives and Q8 transmits at the first cycle slot (T1 ). During the second cycle slot (T2 ), Q8 receives and Q7 transmits. Selecting or allotting a cycle slot for a flow means selecting the cycle slot number (within the hyperperiod H) and the queue for the flow. 25

4

1

CBS

AVB queue Priority 1

2

CBS

BE queue Priority 0

3

Strict Priority

AVB queue Priority 2

1 T0

2 T1

3 T2

4 T3

T4

Priority 2 Credit

Priority 1 Credit

Figure 11: TSN output port with two AVB queues employing CBS and one BE queue.

Transmission Selection Algorithm

Queue 8

Priority Filter

Switching Fabric

Queue 7 Queue 6

Queue 5 Queue 4

Queue 3 Queue 2

Queue 1 GCL

Figure 12: A simple CQF mechanism with eight queues in the egress port of the switch with two queues (Queue 8 and 7) operating as even and odd queue as shown in red.

In the CQF evaluation of TSNBench, we provide the cycle duration (T) and network-specific delays to the model as input through the prompt. WCD CQF: The worst case end-to-end delay of the TT flows in the CQF network is quantified as follows: Max Delay = fi .ϕ + (SWnum + 1) · T + ξ,

(22)

where fi · ϕ is the offset of the flow fi in µs, SWnum is the total number of switches in the route of the TT flow, T is the cycle duration in µs, and ξ denotes the network specific delays: processing delay, propagation delay, and time synchronization error (syncerror ).

T = 50 µs

1 4

T 3

3 4

4 ... 3 4 Hyperperiod = 400 µs

3

4

m 3

Figure 13: The Hypercyle also known as the scheduling cycle of the CQF (400 µs) with cycle duration (T) of 50µs. The different cycle slots are numbered as 1,2 · · · m in red. 26

Transmission Selection Algorithm

Switching Fabric

Even Queue Queue 8 CQF Odd Queue Queue 7 PCP 5 Queue 6 . . . PCP 0 Queue 1 GCL

Figure 14: In this figure, we showcase the even and the odd queue in CQF architecture and during one cycle duration one queue receives the flows and another queue transmits the flows received in the previous cycle duration.

11

More on TSNBench

11.1

Human evaluation decision mechanism

To maintain the same standards across all human reviewers, we use the following rules to evaluate the MCQA dataset. There are four possible options for every question in the MCQA dataset. 1. Accept: i. Technically correct. ii. Clearly worded and self-contained. iii. Unambiguous options. iv. Accurate and sufficient explanation. v. The correct answer is actually the correct answer. 2. Reject: i. Incorrect or misleading. ii. Poorly constructed beyond revision. iii. Irrelevant to TSN. iv. Incomplete information. v. Too paper-dependent. vi. Duplicate questions. 3. Revise: i. Minor issues in grammar, clarity, or wording. ii. Options need improvement. iii. Explanation needs refinement. 4. Doubtful: i. Paper-specific or uncertain about the correctness of the question. ii. Explanation seems questionable. iii. Needs further clarification. For a doubtful multiple-choice question, we read the research paper and re-evaluate the question. Afterward, the decision can be accept, reject, or revise; if it is still doubtful, we send it to another expert reviewer for a consensus-based group decision. Key principles followed while reviewing the dataset: We ensured that the MCQAs were technically accurate and aligned with TSN fundamentals. We avoided tricky questions and preferred clarity over 27

complexity. The same set of rules was given to all expert reviewers who worked on this dataset and served as human judges. After the review, 185 questions were revised by the domain experts, as shown below in Table 5. Table 5: Statistical data of the MCQA dataset after domain expert human review.

11.2

Category

Count

Total questions revised by domain experts

185

Sample Questions

We present three representative sample questions from our MCQA dataset below. Q1

TSN Keyword

What does TAS stand for in TSN traffic management? A. B. C. D.

Transmission Access Scheduler Traffic Analysis System Time-Aware Shaper Traffic Admission Service

Correct Answer: C

Q2

Research Paper

In a Cyclic Queuing and Forwarding (CQF) network what fundamental limitation would prevent effective fault tolerance using Frame Replication and Elimination for Reliability (FRER) in a linear topology where each switch has maximum transmission unit (MTU) sized frames frequently queued? A. B. C. D.

CQF’s ping-pong queue switching would create timing conflicts with FRER’s frame elimination mechanism. EMI interference would corrupt both original and replicated frames equally, making spatial redundancy ineffective. FRER cannot detect bit errors caused by EMI since it lacks Cyclic Redundancy Check (CRC) verification capabilities. Linear topologies cannot provide the disjoint paths required for FRER’s spatial redundancy approach, forcing expensive hardware additions.

Correct Answer: D

Q3

Research Paper

What fundamental challenge makes the Time Aware Shaper (TAS) implementation complex despite its ability to provide guaranteed end-to-end delays? A. B. C. D.

The requirement to synchronize all network devices to a common time reference. The need to maintain separate queues for each traffic class simultaneously. The difficulty in estimating worst-case transmission times for variable-length frames. The synthesis of the gate control list, which is an NP-complete problem.

Correct Answer: D

11.3

Prompt design of open-ended questions

For the CBS and CQF mechanisms, two different approaches are used for WCD calculation. NC is used to calculate the CBS WCD, whereas an analytical mathematical calculation is used to find the 28

WCD for the CQF mechanism. Since these two mechanisms work differently, we design prompts tailored to each mechanism. Role: We start by defining the role of the model: “You are an expert Time-Sensitive Networking (TSN) orchestrator.” We inject three network inputs: (i) network topology, (ii) TSN flow information, and (iii) the routes of the flows. We use the prompt-as-program [Reynolds and McDonell, 2021] approach to separate the network topology, flow information, and flow routes. All of these are provided in text format. However, to evaluate different topologies, flows, and routes, we separate them from the prompt logic. This ensures that the prompt remains the same across different network topologies and parameters. Constants: To correctly calculate the WCD, information about the network parameters is required. To prevent the model from assuming these values and to keep the constant values consistent across all models, we provide this information in the prompt. Constants for CBS open-ended questions: Bandwidth = 100 Mbps, P ropagation delay = 1 µs, Switching delay = 1 µs, T ime synchronization error = 1 µs, The switches of the network are cut-through switches, IdleSlope = 75%

By controlling these network parameters, we directly mitigate hallucinations and assumptions about numerical values. Architecture Restriction: TSN supports multiple architectures that affect the Quality of Service (QoS) and the WCD of the flows. The prompt restricts the model to using only one TSN mechanism through the following directive. For the CBS mechanism, we use: TSN Mechanism: Only Credit-Based Shaper (CBS, IEEE 802.1Qav) is allowed; All flows are AVB Class A, PCP = 6, using queue 6 only.

For the CQF mechanism, we use: TSN Mechanism: Only Cyclic Queuing and Forwarding (CQF, IEEE 802.1Qch) is allowed; All flows are TT, PCP = 7, using queue 7 (odd) and 6 (even) only.

Our reasoning is that letting the model select the TSN architecture or mechanism is a separate benchmarking problem, where the model is evaluated on architecture design performance. In TSNBench, our goal is to benchmark LLMs in TSN. Without an explicit restriction, the model may select an incorrect or inappropriate mechanism, producing a hallucinated architecture that does not satisfy the QoS requirements of the flows. This restriction forces the model to use a single solution space. It further ensures that the WCDs provided by different models are not caused by architectural faults or mechanism selection ambiguity, but rather by calculation and implementation errors within the specified mechanism. Structured Output: We instruct the model through the prompt to provide the output strictly in JSON format [Yang et al., 2026].

29

11.4

TSNBench Open-Ended Question Details

For the open-ended questions, there are three variable entries: network topology, flow information, and flow routing. We use the K-shortest path algorithm to determine the routes of the flows. The routes are then directly provided to the models as input for further evaluation. Network Topologies Used: For the open-ended questions, we selected three different topologies to evaluate the models: a one-switch topology, a medium-mesh topology, and an industrial ring topology. Figures 15, 16, and 17 represent the one-switch, medium-mesh, and ring topologies used in TSNBench, respectively.

SW1

Figure 15: One-switch topology used to evaluate open-ended questions in TSNBench.

ES3

ES4

ES1

SW1

SW2

ES6

ES12

SW4

SW3

ES7

ES10

ES9

ES2

ES11

ES5

ES8

Figure 16: Medium-mesh topology used to evaluate open-ended questions in TSNBench.

SW1

SW2

SW9

SW10

SW8

SW3

SW16

SW11

SW7

SW4

SW15

SW12

SW6

SW5

SW14

SW13

Figure 17: Ring topology representing the industrial ring network used to evaluate open-ended questions in TSNBench.

Flow parameters: We show the flow information used in TSNBench as follows. 30

Flow Information

TC1_flows.txt

0,node2_1,node5_2,2500,709,965 1,node5_4,node3_2,2500,610,825 2,node0_4,node0_1,1000,786,887 3,node2_3,node4_3,2500,1088,1233 4,node0_4,node3_3,1000,1015,488 5,node0_4,node0_1,2500,926,501 ...

Ground Truth WCD Values The ground-truth WCD values of the flows for all open-ended test cases for the CBS mechanism are calculated using a verified NC tool [Zhao et al., 2018, Debnath et al., 2025c, Gavriluţ and Pop, 2020]. For the WCD of the CQF mechanism, we use the mathematical equation given in Eq. 22.

12

More on TSNBench MCQA Evaluation

We evaluate both open-source and closed-source state-of-the-art LLMs on TSNBench. A detailed list of the models, along with their model numbers and snapshots, is given in Table 6. This ensures that the results are reproducible by the community. Table 6: Details of the models used for the benchmarking on TSN. Both MCQA and open-end questions are evaluated on these models. We provide the specific model number and snapshot for reproducibility. Chat Models Model

Family

Model ID

Organization

Country

Grok 4.1 Fast Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-thinking Mode) GPT-4o GPT-4o mini Llama 3.3 Mistral Medium 3.1 Mistral Large 3

Grok Grok DeepSeek GPT GPT Llama Mistral Mistral

grok-4-1-fast-reasoning grok-4-1-fast-non-reasoning deepseek-chat gpt-4o-2024-08-06 gpt-4o-mini-2024-07-18 Llama-3.3-70B-Instruct mistral-medium-2508 mistral-large-2512

xAI xAI DeepSeek AI OpenAI OpenAI Meta (via HF) Mistral AI Mistral AI

USA USA China USA USA USA France France

Claude Sonnet 4.5 o3 GPT-5 DeepSeek-V3.2 (Thinking Mode) Gemini 2.5 Flash

Claude GPT GPT DeepSeek Gemini

Anthropic OpenAI OpenAI DeepSeek AI Google

USA USA USA China USA

Meta (via HF) Alibaba Cloud Mistral AI

USA China France

Reasoning/Thinking Models claude-sonnet-4-5-20250929 o3-2025-04-16 gpt-5-2025-08-07 deepseek-reasoner gemini-2.5-flash

Small Models Llama 3.2 1B Qwen3 8B Ministral 3 8B

12.1

Llama QwenLM Ministral

llama-3.2-1B Qwen3-8B ministral-8b-2512

Extended Experimental Evaluation

We evaluate the models under two different configurations: (i) default temperature settings (0.7) and (ii) temperature set to 0.0, for both MCQA and open-ended questions. As in safety-critical networks, we want to ensure deterministic results. Therefore, we evaluate whether LLMs can provide consistent results when the temperature is set to 0.0. For models that do not support the temperature parameter, we use the default temperature for evaluation. Table 7 provides the accuracy and average consistency of the models for the MCQA dataset under the default temperature and temperature set to 0.0. Average consistency represents the ability of the model to provide the same results across three runs. 31

Table 7: Extended evaluation results of TSNBench MCQA dataset across different state-of-theart models across different families. We provide the accuracy in percentage under two different temperature setting (default and set to 0.0). The consistency shows the performance of the model in providing the same response across three runs. For those models which do not support temperature = 0.0, we use their default temperature and this is marked next to the model in the table. MCQA Accuracy (%)

Model Grok 4.1 Fast† Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-thinking) GPT-4o GPT-4o mini Llama 3.3 Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3† GPT-5† DeepSeek-V3.2 (Thinking)† Gemini 2.5 Flash Llama 3.2 1B Qwen3 8B Ministral 3 8B

Average Consistency (%)

Default Temp.

Temp=0.0

Temp=0.7

Default Temp.

Temp=0.0

Temp=0.7

93.2 – – – – – – – – 94.7 95.0 94.7 – – – –

– 91.7 94.0 91.8 88.3 88.9 92.1 92.8 95.3 – – – 90.1 67.4 83.7 86.9

– 91.6 93.4 92.1 88.2 89.1 92.3 92.9 95.3 – – – 90.8 67.0 82.8 86.5

0.99 – – – – – – – – 0.98 0.99 0.98 – – – –

– 1.00 1.00 1.00 0.99 0.99 1.00 1.00 1.00 – – – 0.98 1.00 0.99 1.00

– 1.00 0.98 0.98 0.98 0.99 0.99 1.00 1.00 – – – 0.97 0.93 0.97 0.97

† Temperature parameter not supported. Evaluated with default settings.

12.2

Cost and Latency

The cost and latency of a model are important evaluation parameters for the research community. Spending a large amount of money on benchmark evaluation is a real bottleneck for research groups. Moreover, not all models can be evaluated locally. Table 8 presents the cost and latency of the TSNBench MCQA and open-ended questions. Evaluating MCQA is relatively much cheaper than evaluating open-ended questions. Table 8: Extended results of cost and latency comparison for MCQA and open-ended evaluation in TSNBench. “–” indicates that cost and latency are not reported for this model, as it successfully evaluated fewer than 50 out of 100 TCs, where a TC is considered successfully evaluated only if the model provided WCD estimates for at least 80% of the flows within that TC. CBS Open-ended questions

CQF Open-ended questions

Cost (USD)

MCQA Latency (ms)

Cost (USD)

Latency (ms)

Cost (USD)

Latency (ms)

0.2490 0.2612 0.0420 2.3661 0.1417 0.5224 0.4432 0.4868 3.5967 7.6293 12.8766 0.4069 0.4164 0.0864 0.3736 0.1237

18,769,322 1,450,175 2,264,129 2,053,601 2,251,786 1,028,334 1,839,587 15,487,495 5,190,222 10,831,712 15,860,682 12,385,365 18,732,689 1,883,306 42,272,121 972,292

0.2241 0.3256 – – 0.3438 0.7399 1.9918 1.5989 21.1719 10.9954 57.5232 – 2.6471 – – 0.2425

43,788,625 3,200,483 – – 7,083,613 1,264,847 6,335,803 11,149,853 14,342,470 10,127,407 74,473,434 – 24,160,941 – – 10,122,749

0.3047 0.3251 0.4816 4.3866 0.3642 0.7137 2.2442 1.2298 18.3093 10.0134 42.4470 – 2.7690 – – 0.2434

49,367,058 3,383,209 13,211,941 2,122,167 7,092,033 1,136,920 6,658,056 8,226,564 12,583,398 8,591,382 58,404,966 – 17,688,716 – – 9,603,062

Model Grok 4.1 Fast† Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-thinking) GPT-4o GPT-4o mini Llama 3.3 Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3† GPT-5† DeepSeek-V3.2 (Thinking)† Gemini 2.5 Flash Llama 3.2 1B Qwen3 8B Ministral 3 8B

† Temperature parameter not supported. Evaluated with default settings.

13

More on TSNBench Open-Ended Questions Evaluation

We provide MAE and MAPE evaluations for the open-ended questions. A sample calculation is given as follows: 32

MAE and MAPE calculation example: Consider a model evaluated on three test cases (TCs). These three TCs may have different topologies, different flows and flow parameters, and different routes. For each TC, we have the ground-truth and predicted WCD values shown in Table 9. The ground truth is calculated using an NC solver for CBS and a mathematical equation for CQF. Table 9: Sample example of test cases (TC) with ground truth, predicted and absolute error values. TC

Flow

Ground Truth (µs)

Predicted (µs)

Abs. Error (µs)

TC1 TC1 TC1

F0 F1 F2

200 150 500

212 180 490

12 30 10

TC2 TC2

F0 F1

100 300

108 255

8 45

TC3 TC3 TC3

F0 F1 F2

400 250 600

420 265 600

20 15 0

Per-TC MAE: Suppose TC1, TC2, and TC3 contain three, two, and three flows, respectively. {f1 , f2 , f3 } ∈ T C1; {f1 , f2 } ∈ T C2; {f1 , f2 , f3 } ∈ T C3; Let Γ(f0 ) denote the absolute error of flow f0 in TC1, β(f0 ) denote the predicted WCD of flow f0 given by the LLM model, and Ω(f0 ) denote the ground truth of flow f0 . We calculate Γ(f0 ) as follows: Γ(f0 ) = |β(f0 ) − Ω(f0 )| In the given example, let Γ(f0 ) = 12, Γ(f1 ) = 30, and Γ(f2 ) = 10 for TC1. Similarly, for TC2, Γ(f0 ) = 8 and Γ(f1 ) = 45 and for TC3, Γ(f0 ) = 20, Γ(f1 ) = 15, and Γ(f2 ) = 0. We calculate the MAE for TC1, TC2, and TC3 represented as MAETC1 , MAETC2 , and MAETC3 as follows: MAETC1 = (12 + 30 + 10)/3 = 17.3 µs MAETC2 = (8 + 45)/2 = 26.5 µs MAETC3 = (20 + 15 + 0)/3 = 11.7 µs For every model, we have 100 test cases, and the final MAE is averaged across all test cases (in this example 3 test cases) and is represented as: MAE = (17.3 + 26.5 + 11.7)/3 = 18.5 µs The per-flow MAPE denoted as α(f0 ) is calculated as follows: α(f0 ) =

|β(f0 ) − Ω(f0 )| × 100 Ω(f0 )

For TC1, we calculate the MAPE as follows: MAPETC1 =

α(f0 ) + α(f1 ) + α(f2 ) = 8.7% 3

Similarly, the MAPE for TC2 and TC3 is given as follows: MAPETC2 = 11.5% MAPETC3 = 3.7% The final MAPE for each model is averaged across the 3 test cases: MAPE = (8.7 + 11.5 + 3.7)/3 = 8.0% 33

Table 10: MAE (µs) for CBS open-ended evaluation across One-Switch topology test cases. “–” denotes invalid, missing, or partial response (model predicted fewer than 80% of flows in the one TC). “0” denotes trivial failure (model returned all-zero WCD values). Best result per TC shown in bold. The MAE (µs) reported in this table is based on average across three runs per TC. Model

TC1

TC2

TC3

TC4

TC5

TC6

TC7

TC8

TC9

TC10

TC11

Grok 4.1 Fast Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-Thinking) GPT-4o GPT-4o mini Llama 3.3 70B Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3 GPT-5 DeepSeek-V3.2 (Thinking) Gemini 2.5 Flash

51.93 – 147.46 505.94 – 516.89 509.17 397.0 499.76 61.05 178.03 – –

91.48 – 698.76 486.53 – 322.23 488.48 381.23 484.85 121.41 86.44 – 225.61

175.11 – 291.06 294.39 293.06 301.96 294.34 172.06 286.34 10.57 30.66 – 220.86

31.03 – – 85.41 133.41 140.34 130.85 20.79 109.67 17.5 15.86 – 61.78

16.04 – 130.71 124.37 – 43.2 35.19 16.35 122.86 105.48 50.59 – 57.02

209.24 – – 265.02 301.19 300.79 295.19 183.52 267.11 84.57 40.84 – 135.3

81.05 – – 204.64 210.14 213.14 205.13 76.58 202.77 143.37 14.55 – 84.65

– – – 63.72 – 58.87 70.37 181.61 61.65 19.64 23.66 – 39.35

136.0 – 237.33 237.22 242.67 244.67 240.64 127.67 216.11 125.73 136.91 – 164.93

172.41 – 432.63 510.61 – 391.42 510.08 403.18 499.51 171.22 225.17 – 334.01

167.5 – – 415.26 – 427.25 418.62 302.1 413.98

– – 778156.77

– – 1467.11

– – 652.98

– – 1360.14

– – 1183.86

– – 10.08

– – 882.67

– – 979.38

– – 905.23

– 158.82

Small Models Llama 3.2 1B Qwen3 8B Ministral 3 8B

– – 858.59

– – 863.29

Table 11: MAE (µs) for CQF open-ended evaluation across One-Switch topology test cases. “–” denotes invalid, missing, or partial response (model predicted fewer than 80% of flows in the one TC). “0” denotes trivial failure (model returned all-zero WCD values). Best result per TC shown in bold. The MAE (µs) reported in this table is based on average across three runs per TC. Model

TC1

TC2

TC3

TC4

TC5

TC6

TC7

TC8

TC9

TC10

TC11

Grok 4.1 Fast Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-Thinking) GPT-4o GPT-4o mini Llama 3.3 70B Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3 GPT-5 DeepSeek-V3.2 (Thinking) Gemini 2.5 Flash

177.2 95.0 38.38 316.33 88.33 67.72 801.11 100.0 87.43 34.33 169.3 – 80.68

112.99 93.67 95.0 283.33 85.67 97.72 93.0 84.0 44.48 159.11 162.19 – 91.67

54.6 71.09 198.24 633.0 93.0 29.26 91.0 101.0 11.68 146.89 85.82 – 79.79

46.64 91.0 45.78 0.0 91.0 66.82 90.33 82.67 46.29 68.99 74.85 – 5.34

28.29 36.11 41.0 1.0 97.0 99.33 91.0 101.0 47.44 32.78 27.11 – 15.68

66.95 94.5 95.0 32.33 98.33 37.74 93.0 100.0 43.38 69.33 76.64 – 21.28

– 91.0 95.0 313.67 97.0 57.94 91.17 101.0 30.28 32.4 54.51 – 49.73

52.89 95.0 57.0 15.67 29.67 90.25 58.94 101.67 48.89 70.37 25.58 – 8.42

150.67 93.67 89.0 949.0 93.0 717.31 93.0 84.0 46.54 47.0 108.33 – 44.36

288.03 – – 317.0 91.0 171.29 92.67 50.0 13.69 126.11 160.43 – 157.85

128.18 91.0 149.89 283.67 95.44 109.71 502.0 78.11 31.28 81.15 115.66 – 80.49

Llama 3.2 1B Qwen3 8B Ministral 3 8B

– – 1216.33

– – 2279.78

– – 362.33

– – 765.22

– – 80.56

– – 593.44

– – 722.67

– – –

– – 2119.22

Small Models – – 1053.89

– – 1177.67

Table 12: MAE (µs) for CBS open-ended evaluation across Ring topology test cases (TC1-TC20), taken from 100 total test cases spanning three topologies. “–” denotes invalid, missing, or partial response (model predicted fewer than 80% of flows in the one TC). “0” denotes trivial failure (model returned all-zero WCD values). Best result per TC shown in bold. The MAE (µs) reported in this table is based on average across three runs per TC. Model

TC1

TC2

TC3

TC4

TC5

TC6

TC7

TC8

TC9

TC10

TC11

TC12

TC13

TC14

TC15

TC16

TC17

TC18

TC19

TC20

Grok 4.1 Fast Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-Thinking) GPT-4o GPT-4o mini Llama 3.3 70B Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3 GPT-5 DeepSeek-V3.2 (Thinking) Gemini 2.5 Flash

– 353.37 0 0 575.21 537.49 222.92 511.74 530.12 250.72 110.06 – –

– 2107.44 0 0 1021.42 0 894.74 – 941.39 474.97 132.15 – –

– 1346.37 0 0 1010.98 834.68 236.19 919.4 938.51 859.62 754.56 – 750.1

– 947.7 0 0 974.13 940.46 477.27 – 880.56 380.65 541.75 – 448.42

– 1088.22 0 0 435.05 347.3 394.88 375.99 279.59 225.1 260.18 – 444.25

– 8750.11 0 0 602.95 568.77 449.45 – 400.01 172.92 113.28 – 536.47

– 3031.05 0 0 680.03 531.01 231.57 – 595.62 225.86 254.26 – 898.0

0 4794.74 0 0 763.06 688.05 283.69 693.06 393.4 434.22 245.89 – 608.64

– 212.79 – – 409.39 409.44 183.42 312.21 315.46 73.84 60.51 – –

– 2154.81 – 0 330.35 292.29 295.81 – 251.47 124.12 303.68 – 237.37

– 449.66 0 0 538.01 282.04 248.21 456.66 400.11 260.18 – – 362.35

– 10680.41 0 0 0 271.79 808.44 271.54 317.76 92.34 164.84 – 84.77

– 317.1 0 0 339.8 0 910.92 – 266.55 91.49 156.55 – 350.92

– 1815.54 0 0 366.84 – 653.12 – 295.4 194.26 91.19 – 276.76

0 215.52 0 0 343.9 – 889.7 – 289.41 152.06 107.45 – 347.04

– 264.96 0 0 216.14 – 709.57 – 171.74 172.15 167.67 – 255.09

– 1491.62 0 0 513.14 399.24 498.84 404.25 115.43 91.48 257.56 – 234.62

– 1992.75 0 0 509.49 503.22 781.82 – 399.78 111.08 172.03 – 91.7

– 686.01 – – 713.7 – 690.99 – 616.89 429.72 464.36 – 238.5

– 559.0 0 0 828.38 448.82 229.13 – 733.79 784.25 305.86 – 679.01

Llama 3.2 1B Qwen3 8B Ministral 3 8B

0 – 3571.31

0 – 413.05

0 – 318.9

0 – –

0 – 826.43

0 – 486.78

0 – –

0 – 460.25

0 – 323.18

0 – 274.1

0 – 976.95

0 – –

0 – 898.99

0 – –

– – –

0 – 787.18

0 – 740.15

0 – 293.11

– – –

Small Models 0 – 581.21

34

Table 13: MAE (µs) for CQF open-ended evaluation across Ring topology test cases (TC1-TC20), selected from 100 total test cases spanning three topologies. “–” denotes invalid, missing, or partial response (model predicted fewer than 80% of flows in the one TC). “0” denotes trivial failure (model returned all-zero WCD values). Best result per TC shown in bold. The MAE (µs) reported in this table is based on average across three runs per TC. Model

TC1

TC2

TC3

TC4

TC5

TC6

TC7

TC8

TC9

TC10

TC11

TC12

TC13

TC14

TC15

TC16

TC17

TC18

TC19

TC20

Grok 4.1 Fast Grok 4.1 Fast (Non-Reasoning) DeepSeek-V3.2 (Non-Thinking) GPT-4o GPT-4o mini Llama 3.3 70B Mistral Medium 3.1 Mistral Large 3 Claude Sonnet 4.5 o3 GPT-5 DeepSeek-V3.2 (Thinking) Gemini 2.5 Flash

137.1 141.15 237.15 8.08 147.23 148.92 175.03 26.15 5.71 115.13 162.28 – 122.97

179.01 166.57 173.33 0.57 175.3 – 173.92 – 135.05 148.83 110.45 – 51.27

55.82 172.65 183.33 2.57 173.17 173.41 158.68 43.14 72.06 107.3 198.84 – 73.69

27.44 152.08 164.15 0.45 160.0 – 157.65 – 122.75 58.18 98.48 – 28.2

20.53 173.0 7.0 0.78 195.33 196.0 149.24 45.0 55.69 52.29 126.95 – 32.12

– 177.0 199.16 29.18 190.38 – 143.71 – 3.12 207.42 114.18 – 15.17

– 166.62 183.67 1.51 175.24 – 132.14 – 123.14 105.08 93.79 – 93.29

47.88 178.33 0 0.41 185.53 203.84 124.24 44.84 166.63 101.71 92.27 – 78.73

57.49 209.56 218.6 223.13 222.29 211.58 203.6 16.27 96.75 99.13 210.55 – 165.45

13.89 204.81 212.31 0.52 200.1 – 195.5 – 181.58 80.73 8.97 – 16.18

169.51 202.16 213.13 144.23 211.15 208.08 202.84 1.0 102.01 81.12 154.81 – 15.06

71.4 185.33 197.46 5.14 186.95 187.3 138.42 68.67 44.4 77.46 153.82 – 26.55

169.18 204.27 219.53 1.0 208.8 – 116.09 – 164.4 48.74 27.24 – 69.58

141.5 198.7 205.5 1.75 208.87 – 113.37 – 157.96 98.05 121.55 – 76.32

151.1 220.07 237.87 1.4 226.62 – 161.92 – 209.06 192.89 40.93 – 123.55

– 185.43 197.75 5.4 198.73 – 156.62 – 16.1 74.34 27.79 – 38.22

– 216.04 233.42 5.07 223.73 227.18 139.56 44.27 17.23 82.77 117.35 – 126.44

175.44 184.53 193.73 1.0 190.13 – 117.54 – 3.69 166.86 107.33 – 136.07

201.77 230.49 237.82 66.8 239.33 – 223.69 – 127.51 50.62 166.49 – 231.22

301.2 220.4 228.17 14.04 228.17 – 191.39 – 7.45 198.67 129.34 – 215.35

Llama 3.2 1B Qwen3 8B Ministral 3 8B

0 – 4348.12

0 – 1196.35

0 – 6907.14

0 – 714.58

0 – 5215.69

0 – 905.0

0 – 768.1

0 – 9065.2

0 – 206.04

0 – 785.87

0 – 811.05

0 – –

0 – 1641.37

0 – 2440.85

0 – 852.55

0 – 6831.6

0 – 1151.91

0 – 237.98

– – 4420.84

Small Models 0 – 121.13

MCQA Accuracy Ministral 3 8B Qwen3 8B Llama 3.2 1B Gemini 2.5 Flash DeepSeek-V3.2 (T) GPT-5 o3 Claude Sonnet 4.5 Mistral Large 3 Mistral Medium 3.1 Llama 3.3 GPT-4o mini GPT-4o DeepSeek-V3.2 (NT) Grok 4.1 Fast (NR) Grok 4.1 Fast (R)

CBS vs CQF - Per-Flow MAE (µs) | 100 Test Cases (TCs)

87% 84% 67% 90% 95% 95% 95% 95% 93% 92% 89% 88% 92% 94% 92% 93%

0 20 40 MCQA (%)

60

80 100 120 0

higher is better

CBS 1000

lower is better

2000

MAE (µs)

3000

4000

CQF 5000

Figure 18: Performance comparison across MCQA and open-ended WCD computation for all 16 evaluated models in TSNBench, illustrating the dissociation between declarative knowledge and computational reasoning. (Left) MCQA accuracy (%) per model. (Right) Per-TC MAE distribution (in µs) for CBS and CQF open-ended questions, shown as box plots over 100 total evaluated test cases, aggregated across three independent runs. Models achieving above 90% MCQA accuracy exhibit substantially high MAE on open-ended WCD computation. In TSNBench, all test cases contributes equally towards the model performance irrespective of the number of flows in the network. As per the network architecture, all flows are equally critical and needs the same preference. This ensures that for each network scenario all the flows are weighted equally.

35

Table 14: CBS Error Analysis Case 1: Lack of Specific Knowledge. Test Case: TC1 TSN mechanism: CBS

You are an expert Time-Sensitive Networking (TSN) orchestrator. Your task is to calculate the worst case delay (WCD) for each TSN flow. Input: Network Topology (T C1_topo.txt) Flow Information (T C1_f lows.txt) Routing of the Flow (T C1_route.txt) Constant: Bandwidth link = 100 Mbps; Propagation delay = 1 µs; Switching delay = 1 µs; Time synchronization error = 1 µs; The switches of the network are cut-through switches. TSN Mechanism: Only Credit-Based Shaper (CBS, IEEE 802.1Qav) is allowed; All flows are AVB Class A, PCP = 6, using queue 6 only. Task: 1. Map each egress port’s queues and collect the set of flows traversing from that port, using the given topology, flows, and route of the flow. 2. For each egress port, use the given IdleSlope and then compute the SendSlope. 3. For each flow, construct an arrival curve from its frame size and periodicity. 4. For each port, derive a lower-bounded CBS service curve. 5. Calculate the worst case delay (WCD) in microseconds (µs) for each flow using Network Calculus method. 6. Provide the confidence score between 0.0 and 1.0 from your answers. 1.0 means mathematically or procedurally provable from given info with zero ambiguity. 0.0 means zero confidence.

36

Table 14: CBS Error Analysis Case 1: Lack of Specific Knowledge. (continued)

Grok 4.1 Fast (Non-Reasoning) output: F0: 1452.0, F1: 1124.0, F2: 678.0, F3: 1234.0, F4: 1567.0, F5: 689.0, F6: 456.0, F7: 789.0, F8: 567.0, F9: 589.0, F10: 678.0, F11: 645.0, F12: 1123.0, F13: 789.0, F14: 456.0, F15: 678.0, F16: 789.0, F17: 1567.0, F18: 456.0, F19: 1234.0

Model Reasoning and Explanation Output: Parameters used: send_slope: 75% × 100 Mbps = 75 Mbps = 75000 bits/µs frame_overhead: 24 bytes (Ethernet header + FCS), mtu: 1522 bytes (max frame size), queue: 6 (PCP=6, AVB Class A), cbs_credit_max: MTU × 8 = 12176 bits WCD equation given by the model: α(t) =

Lmax t + Period Period

(arrival curve)

β(t) = Rsend × t − C (CBS service curve) W CD = h × (prop_delay + switch_delay + sync_error) + max(α ∩ β) SendSlope = IdleSlope × Bandwidth = 0.75 × 100 Mbps Rsend = SendSlope = 75 Mbps C = creditmax = MTU × 8 bits h = hop count from path Per-flow WCD: W CDFi =

X

 dhprop + dhswitch + dhsync + max (αFi (t) ∩ βport (t)) t≥0

h∈path(Fi )

37

Table 14: CBS Error Analysis Case 1: Lack of Specific Knowledge. (continued)

Run 2 of the model (Grok 4.1 Fast (Non-Reasoning)): WCD equation given by the model during run 2: Frame size: Li = payload_i + 24 bytes, 8Li Transmission time: Txi = , R Li Arrival curve: αi (t) = · t, pi CBS service curve: β(t) = min {αhi (t), sendSlope · t + creditmax } , X Worst-case delay: W CDi = max (Tprop,h + Tswitch,h + Tqueue,h ) + Tsync , h∈path

Li + burst_interference, αC Send slope: sendSlope = idleSlope · R − idleSlope · R = αR(1 − 2α);

Queueing delay:

Tqueue,h =

Per-flow WCD: W CDi = hop_counti · (1 + 1) + max_queue_delay + 1; Expert Explanation: The model uses different equations across different runs even with temperature set to 0.0. When analyzing the first run, the model makes several fundamental errors in evaluating the interference/blocking experienced by the queue under analysis, i.e., both the arrival curve and CBS service curve. First, it assumes the wrong maximum burst size and does not consider the variant output arrival curve at subsequent hops along the path. Second, it assumes a non-zero maximum credit, which is possible in this scenario due to the non-preemption frame of lower priority of non-CBS traffic. However, the model applies the wrong maximum credit to the CBS service curve, incorrectly understanding the relation of the corresponding rate (75%). Furthermore, in the final WCD calculation, the model considers only the arrival curve of the frame under analysis. This is incorrect: the proper approach requires using the aggregate arrival curve of all frames transmitted through the queue. It is also not clear whether the max(α ∩ β) operation represents the maximum horizontal deviation between α and β. The output arrival curve is not considered in the following hops’ calculation. On the other hand, the second run exhibits even more severe hallucinations, producing calculations of the wrong arrival curve without considering any burst size, and the wrong CBS service curve considering the wrong service rate (sendSlope) and wrong relation between the corresponding rate and the maximum credit, as well as without considering the output arrival curve either etc. GPT-5 output: F0: 714.65, F1: 821.79, F2: 171.48, F3: 687.17, F4: 907.26, F5: 171.48, F6: 183.9, F7: 453.91, F8: 611.08, F9: 181.3, F10: 611.08, F11: 321.35, F12: 1124.12, F13: 483.67, F14: 172.43, F15: 489.26, F16: 521.33, F17: 918.73, F18: 181.3, F19: 624.12

38

Table 14: CBS Error Analysis Case 1: Lack of Specific Knowledge. (continued) Model Reasoning and Explanation Output: sendSlope = −25 Mbps, X Arrival curve: b =

Li bits,

∀Fi | share output port

X

r=

∀Fi | share output port

Li bits/µs, pi

α(t) = b + rt, Ttech,total = Tprop + Tswitching + Tsync = hops × 1 µs + (hops − 1) × 1 µs + (hops − 1) × 1 µs = 13 µs, W CDtotal = W CDsum + Ttech,total = 701.65 µs + 13 µs = 714.65 µs. Expert Explanation: GPT uses the correct arrival and service curves. Its main error is that it fails to include the output port (output arrival curves) in the analysis. Ministral 3 8B output: F0: 1250.0, F1: 1100.0, F2: 500.0, F3: 1500.0, F4: 600.0, F5: 1100.0, F6: 2500.0, F7: 1200.0, F8: 500.0, F9: 500.0, F10: 2000.0, F11: 500.0, F12: 1000.0, F13: 1000.0, F14: 2000.0, F15: 500.0, F16: 1500.0, F17: 1000.0, F18: 500.0, F19: 2000.0 Model Reasoning and Explanation Output: Bandwidth:

R = 100 Mbps = 12.5 MB/s,

Propagation delay:

Tprop = 1 µs

per hop,

Switching delay:

Tswitch = 1 µs

per switch,

Time synchronization error:

Tsync = 1 µs,

IdleSlope:

IdleSlope = 0.75R = 75 Mbps = 9.375 MB/s,

Cut-through switch: Payload overhead:

Loverhead = 18 B + 4 B = 22 B,

Maximum frame size: Maximum transmission time:

Tbuffer = 0,

Lmax = 1433 B + 22 B = 1455 B, Ttx,max =

39

1455 × 8 = 9.24 µs. 100 Mbps

Table 14: CBS Error Analysis Case 1: Lack of Specific Knowledge. (continued)

SendSlope = IdleSlope (1 − IdleSlope) = 9.375 MB/s × 0.25 = 2.34375 MB/s, Bburst,i = SendSlope · pi ,  CBS service curve: S(t) = SendSlope t − Burst size:

Bburst SendSlope

 ,

t≥

Bburst . SendSlope

WCD Calculation: Arrival curve for flow Fx :

R(t) = min {Bx , Sx t} , Bx = Lpayload,x + Loverhead ,

8Lpayload,x , px  Service curve for each port: S(t) = SendSlope t − Sx =



Bburst , SendSlope   Bburst − SendSlope · P eriod W CD = max 0, SendSlope + Tprop + Tswitch + Tsync . Bburst SendSlope

,

t≥

Expert Explanation: Even though it is a small model, this agent delivers a poor result. Similar to what was observed with Grok 4.1 Fast (Non-Reasoning), it fails to account for the arrival curves of flows and the wrong CBS service curve. In addition, it analyzes only a single queue, instead of computing the delay at the output port of each device along the end-to-end path. Moreover, it introduces unit inconsistencies such as dividing the frame size expressed in bytes by a link speed given in bits/µs, which leads to incorrect results.

40

Table 15: CQF Error Analysis Case 1: Lack of Specific Knowledge. Test Case: TC1 TSN mechanism: CQF

You are an expert Time-Sensitive Networking (TSN) orchestrator. Your task is to calculate the worst case delay (WCD) for each TSN flow. Input: Network Topology (T C1_topo.txt) Flow Information (T C1_f lows.txt) Routing of the Flow (T C1_route.txt) Constant: Bandwidth link = 100 Mbps; Propagation delay = 1 µs; Switching delay = 1 µs; Time synchronization error = 1 µs; The switches of the network are cut-through switches. TSN Mechanism: Only Cyclic Queuing and Forwarding (CQF, IEEE 802.1Qch) is allowed; All flows are TT, PCP = 7, using queue 7 (odd) and 6 (even) only. Task: 1. Map each egress port’s queues and collect the set of flows traversing that port, using the given topology, flows, and route of the flow. 2. For the entire network, use the given cycle duration and compute the Hypercycle. 3. For each flow, set the offset or the start time of the flow from the sending node as 0. 4. Calculate the worst case delay (WCD) in microseconds (µs) for each flow. 5. Provide the confidence score between 0.0 and 1.0 from your answers. 1.0 means mathematically or procedurally provable from given info with zero ambiguity. 0.0 means zero confidence.

41

Table 15: CQF Error Analysis Case 1: Lack of Specific Knowledge. (continued)

Claude Sonnet’s output: F0: 257.72, F1: 206.8, F2: 105.096, F3: 218.704, F4: 253.904, F5: 104.008, F6: 103.808, F7: 153.52, F8: 119.408, F9: 119.752, F10: 122.896, F11: 118.464, F12: 222.04, F13: 156.552, F14: 103.808, F15: 156.576, F16: 155.68, F17: 256.616, F18: 101.696, F19: 208.328 Model Reasoning and Explanation Output:

LCM of all periods and cycle: WCD equation:

W CD =

X

Cycle duration: = 50 µs, LCM(1000, 2500, 5000, 50) = 5000 µs, Hypercycle: = 5000 µs. (Ttrans + Tprop + Tsw + Tsync ) + Tbuffering .

hops

Ttrans =

(Payload + Ethernet_overhead) · 8 , R

Tprop = 1 µs

per link,

Tsw = 1 µs

per switch,

Tsync = 1 µs

per hop,

Tbuffering = (hop_count − 1) · Tcycle In CQF, frames may wait up to one cycle duration at each intermediate switch if they arrive just after the transmission window closes. For cut-through with CQF, worst case adds (N − 1) ∗ cycle_duration where N is hop count. Hop count = 5, Payload_bytes = 965 bytes, Frame_size_bytes = 1003 bytes, Transmission_delay = 80.24 µs, Total propagation delay = 5 µs, Total switching delay = 4 µs, Total synchronization error = 5 µs, CQF buffering = 200 µs, WCD = 257.72 µs. Ground truth WCD:: Route hop count of f1 : 4, T = 50 µs, W CDf1 = f1 · ϕ + (SWnum + 1)T + ξ = (3 + 1) × 50 + 5 = 205 µs. Expert Explanation: The model knows the working mechanism of CQF. However, the model calculated the wrong Hypercycle by including cycle duration along with the periods. Furthermore, the CQF WCD explanation given by the model is wrong.

42

Table 15: CQF Error Analysis Case 1: Lack of Specific Knowledge. (continued)

GPT-4o output: F0: 260.0, F1: 210.0, F2: 60.0, F3: 210.0, F4: 260.0, F5: 60.0, F6: 60.0, F7: 110.0, F8: 60.0, F9: 60.0, F10: 60.0, F11: 60.0, F12: 210.0, F13: 110.0, F14: 60.0, F15: 110.0, F16: 110.0, F17: 260.0, F18: 60.0, F19: 210.0 Model Reasoning and Explanation Output: Hypercycle = 100 µs. WCD = (hop_count × (cycle_duration + Tswitch + Tprop + Tsync )) +

Payload Bandwidth

Each flow’s path, cycle duration, hypercycle, and CQF parameters are detailed, along with the equation used for WCD calculation. No missing inputs are identified, but assumptions are made based on typical network configurations. Expert Explanation: The Hypercycle calculated by the model is wrong, moreover the WCD equation used by the Payload model has an extra wrong component Bandwidth . One of the most interesting finding is all the models provide very different WCD equations across different runs and different test cases. Ministral 3 8B output: F0: 1160.0, F1: 1020.0, F2: 100.0, F3: 1160.0, F4: 1260.0, F5: 100.0, F6: 100.0, F7: 100.0, F8: 100.0, F9: 100.0, F10: 100.0, F11: 100.0, F12: 1160.0, F13: 100.0, F14: 100.0, F15: 100.0, F16: 100.0, F17: 1260.0, F18: 100.0, F19: 1160.0 Model Reasoning and Explanation Output: Hypercycle:

H = 10000 µs,

Max frame size = 1526 bytes = 1500 + 26 bytes, Max transmission time = 12.208 µs, Max queuing delay = 48 µs, W CD = h (Tprop + Tswitch + Tsync )   payload + Ttx,max max_frame_size + Tqueue,max (h − 1),

43

Table 15: CQF Error Analysis Case 1: Lack of Specific Knowledge. (continued)

Hop count:

h = 5,

Propagation + switching + sync = 5 × (1 + 1 + 1) = 15 µs,   2500 Transmission time = 12.208 × = 12.208 × 2 = 24.416 µs, 1526 Queueing delay = 48 × (5 − 1) = 192 µs, Total =15 + 24.416 + 192 = 231.416 µs, Adjusted worst-case = 1160 µs.

Expert Explanation: The WCD equation provided by the model is wrong. Even though the model takes into consideration the number of hops present in the route, the delays accumulated across each hop and also calculates the hop count. However, the model misses the most crucial part of the WCD equation which Furthermore, the two components of the l is the cycle duration. m payload WCD equation (Ttx,max max_frame_size ) and (Tqueue,max (h − 1)) considered by the model is entirely hallucinated. These two components are mainly contributing to the large WCD values of this model.

14

Failure Mode Analysis

To understand the nature of WCD computation failures, we identify five distinct failure modes observed across models and mechanisms. Trivial Zero Failure: The model returns WCD = 0 for all flows, producing a structurally valid JSON response but with no computational content. This failure mode affects GPT-4o and DeepSeekV3.2 (Non-thinking) on CBS, and Llama 3.2 1B across all test cases for CBS and CQF. This suggests these models recognize the output format requirement but cannot engage with the underlying NC computation or any reasoning behind the WCD calculation. Partial Prediction Failure: The model produces valid WCD values for fewer than 80% of flows in a given TC, resulting in incomplete coverage. This affects Mistral Large 3 on CBS and Llama 3.3 on CBS, suggesting these models lose track of flow indexing in large topologies. Timeout and Context Failure. The model cannot process the full open-ended prompt due to context window limitations or API timeout. This affects Qwen3 8B (API timeout across all TCs) and Llama 3.2 1B (context limit exceeded), confirming that small models are structurally unsuited for TSN open-end evaluation. Empty Response: The model returns an empty response for all open-ended test cases, regardless of network topology or flow count. This failure mode exclusively affects DeepSeek-V3.2 (Thinking), which produces no output, neither WCD values nor intermediate reasoning, across all evaluated topologies, including one-switch, medium-mesh, and ring configurations, and across all flows, for both CBS and CQF mechanisms.

44

Record · ID 175172 · SHA-256 c194625db21a9a41
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.