Conceptio › Archive › arXiv CS
arXiv CSopen access

Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix

arXiv:2605.16112v1 [cs.LG] 15 May 2026

Jinhao Zhang1

Kangfei Zhao1 Qiuhao Zeng2 Long-Kai Huang3† 1 Beijing Institute of Technology 2 University of Toronto 3 Hong Kong Baptist University

Abstract Transformer-based architectures have become the dominant paradigm for Continuous-Time Dynamic Graph (CTDG) learning, yet their performance remains limited on temporally shifted datasets. In this work, we identify attention dispersion as a shared failure mode of dynamic graph Transformers under temporal distribution shift. Through controlled ablation contrasting structurally and temporally distinguished historical neighbors against random ones, we show that prediction depends on a class of critical nodes that carry consistently more predictive signal than arbitrary neighbors. However, existing Transformers fail to focus on these nodes even when they are present in the input, as temporal shift weakens attention contrast and produces overly dispersed attention distributions. This diagnosis suggests a simple and transferable fix: replace standard attention with differential attention, which suppresses common-mode attention and amplifies distinctive token-level signals. When added to three representative CTDG Transformer baselines, differential attention consistently improves performance, with gains concentrated on high-shift datasets. Attention-level measurements further confirm the mechanism, showing reduced attention entropy and increased attention mass on critical nodes. Building on these findings, we introduce DiffDyG, a reference implementation combining differential attention with standard input encodings. Across 9 benchmarks and three negative sampling protocols, DiffDyG achieves SOTA performance, with especially large gains on the most shifted datasets.

1

Introduction

Continuous-Time Dynamic Graphs (CTDGs) provide a natural framework for modeling temporal interactions in evolving systems, including social networks, e-commerce platforms, communication infrastructures, and biological processes [10, 11, 9]. Recent CTDG models increasingly adopt Transformer-based architectures [23], which use self-attention to process variable-length historical interaction sequences. A line of recent work [24, 31, 15] has improved how dynamic graph structures are converted into Transformer-compatible sequences, for example through neighbor co-occurrence encoding and temporal patching. These designs allow Transformers to model longer histories and have made them a dominant paradigm on standard CTDG benchmarks. Despite these advances, existing Transformer-based models plateau on the same several datasets. Across the 9 standard CTDG benchmarks we examine, the strongest existing transformer reaches only 71.1% AP on US Legis., 66.5% on UN Trade, and 69.0% on UN Vote, with little improvement despite years of architectural innovation in sequence construction and feature design. The pattern is consistent across architectures rather than specific to any one method, and the affected datasets correlate with high temporal distribution shift between training and test windows. This suggests † Correspondence to: Long-Kai Huang <[email protected]>

Preprint.

a shared failure mode under temporal distribution shift. Existing approaches, all of which target sequence construction or feature design, do not fix it. We trace this limitation to attention allocation. Through controlled ablation, we identify a class of historical neighbors that we call critical nodes. These nodes either occupy central positions in the local interaction structure or participate in temporally stable relationships. Masking critical nodes causes a much larger performance drop on shifted datasets than masking the same number of randomly selected neighbors, showing that these nodes carry disproportionately important predictive information. However, existing Transformer-based models still perform poorly on shifted datasets even when all critical nodes are present in the input. Thus, the issue is not simply information availability. It is also unlikely to be a pure capacity issue, since the same architectures perform well on less-shifted datasets. Instead, the failure lies in how attention distributes probability mass over historical tokens. Under temporal shift, the attention score provides weaker relative separation among tokens, and the subsequent softmax produces a flatter attention distribution. As a result, attention becomes more dispersed, and the model fails to concentrate on the critical historical signals that remain present in the input. This diagnosis leads to two testable predictions. First, if attention allocation is the shared bottleneck, then replacing only the attention mechanism in existing Transformers should improve their robustness under temporal shift while leaving their sequence construction and other components unchanged. Second, the improved attention mechanism should produce less dispersed attention distributions and assign more mass to critical nodes. We test these predictions using differential attention [29], originally proposed for language modeling, which subtracts two softmax attention maps computed over the same input, thereby suppressing dispersed common-mode attention and amplifying distinctive token-level signals. Both predictions are supported empirically. When added as a plug-in module to three architecturally different Transformer baselines, namely DyGFormer, TIDFormer, and TCL, differential attention improves their average AP by +8.3, +5.7, and +2.9, respectively. The gains are concentrated on highshift datasets, where improvements range from +11 to +26 AP. Direct attention-level measurements further support the proposed mechanism. Differential attention consistently reduces attention entropy and increases the proportion of attention mass assigned to critical nodes, with the strongest effects appearing on the datasets where temporal distribution shift is most severe. Together, the failure-mode analysis, the cross-architecture performance recovery, and the attention-level measurements point to the same conclusion: attention allocation is a shared bottleneck of CTDG Transformers under temporal distribution shift, and differential attention provides a targeted remedy. Building on this finding, we introduce DiffDyG, a CTDG Transformer that combines differential attention with standard input-construction components, including RoPE, neighbor co-occurrence encoding, and spatial distance encoding. Across 9 benchmarks, DiffDyG achieves the strongest overall performance. On the three most challenging shifted datasets, it improves AP from 71.1 to 87.5 on US Legis., from 66.5 to 99.0 on UN Trade, and from 69.0 to 88.8 on UN Vote. These results show that a targeted change to attention allocation can close performance gaps that prior sequence-construction improvements have not resolved. Our contribution is to identify where Transformer-based CTDG models fail under temporal distribution shift, show that this failure is shared across architectures, validate a targeted attention-level fix, and verify the mechanism through direct measurements of attention behavior. Contributions are summarized as: 1) We diagnose attention dispersion as a shared failure mode of Transformer-based CTDG models under temporal distribution shift. Through controlled ablations on critical and random historical neighbors, we show that key predictive signals are present in the input but are not properly attended to under shift. 2) We validate this diagnosis across architectures and mechanisms. Adding differential attention to three representative Transformer baselines yields consistent gains, especially on high-shift datasets, while direct attention-level measurements show reduced attention entropy and increased attention mass on critical nodes. 3) We introduce DiffDyG, a reference implementation that combines differential attention with standard input-construction components and achieves new SOTA performance on all 9 benchmarks under three negative sampling protocols.

2

Table 1: Temporal distribution gap measured by Maximum Mean Discrepancy (MMD) across datasets. Metric

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

MMD

0.271

0.257

0.340

0.457

0.355

0.372

0.697

0.476

0.752

2

Related works

Dynamic graph learning. Existing dynamic graph modeling approaches broadly fall into two categories: discrete-time modeling [3, 18, 30] and continuous-time modeling [4, 19, 26]. Discretetime dynamic graph (DTDG) methods represent a dynamic graph as a sequence of snapshots. Static graph encoders are applied to each snapshot and sequential modules capture the time-series dynamics [8, 14]. In contrast, continuous-time dynamic graph (CTDG) methods represent a graph as a stream of timestamped interactions, enabling fine-grained temporal reasoning. CTDG models include temporal Point Process (TPP)-based models [7, 22, 32], random-walk based models [12, 13, 26], memory-based models [17, 25, 21], and specialized temporal message-passing networks [11, 31, 27], where Transformer-based [23] architectures have achieved strong performance in CTDG learning. Specifically, TCL [24] combines a graph Transformer with contrastive learning to master temporal and topological information in dynamic graphs. DyGFormer [31] takes a neighbor co-occurrence encoding as the input of vanilla Transformers. SimpleDyG [27] proposes an ego-graph tokenization scheme along the temporal dimension, and TIDFormer [15] designs calendar-based temporal encodings for mixed-granularity interaction modeling. Despite these research efforts, most existing methods neglect the temporal shift problem underlying dynamic graphs.

3

Diagnosing the Failure Mode of Dynamic Graph Transformers

In this section, we investigate why transformer-based CTDG models systematically degrade under temporal distribution shift. We show that performance degradation is strongly associated with temporal distribution shift and that the same datasets are difficult for several Transformer architectures (§3.1). To localize the cause, we identify a class of historical neighbors carrying additional predictive signal, which we define as critical nodes, and show through controlled ablation that the failure on shifted data is not one of information availability or model capacity but of attention allocation (§3.2). We first formalize the problem and setup. Continuous-Time Dynamic Graph (CTDG). A CTDG is a chronologically ordered sequence of temporal interactions G = {(u1 , v1 , t1 ), · · · , (un , vn , tn )}, where 0 ≤ t1 ≤ · · · ≤ tn . Each tuple (ui , vi , ti ) represents a directed interaction at time ti between nodes ui , vi ∈ N . Nodes and interactions may have associated features u ∈ RdN and e ∈ RdE .. Problem definition. Given a source u, destination v, timestamp t, and the history {(u′ , v ′ , t′ ) | t′ < t}, dynamic link prediction estimates the likelihood of an interaction between u and v at time t. Datasets and evaluation. We use 9 standard CTDG benchmarks [16]: Wikipedia, Reddit, MOOC, Enron, UCI, Can. Parl., US Legis., UN Trade, and UN Vote. Following [31], we use chronological 70%/15%/15% train/validation/test splits. Unless otherwise specified, we use random negative sampling strategy and report Average Precision (AP) in transductive setting. For diagnosis, we analyze 3 Transformer-based methods with different sequence construction designs: DyGFormer [31], TIDFormer [15], and TCL [24]. Experimental details are in Appendix A. 3.1

Temporal Shift Exposes a Shared Failure Pattern

We quantify temporal distribution shift between the training and test windows using Maximum Mean Discrepancy (MMD) [6]. MMD is computed on features extracted from the two windows using trained DyGFormer [31]. Tab. 1 shows that the benchmarks differ substantially in temporal shift. Wikipedia and Reddit have low shift, with MMD below 0.30, whereas US Legis. and UN Vote exceed 0.69. UN Trade has the third-largest MMD. Fig. 1 shows a strong negative correlation between MMD and AP for all 3 Transformer baselines in both transductive and inductive settings, with Pearson’s R ranging from −0.81 to −0.93. 3

MMD

(a) DyGFormer

transductive (Pearson R=-0.83) inductive (Pearson R=-0.82)

90 80 70

MMD

gis . Le US

UN

E UN nron Tra de

50

Vo te

60

U CaMOOCI n. C Pa rl.

Vo te

Le

UN

50

gis .

60

Average Precision (%)

70

100

WikRe ipde dit dia

80

US

gis .

Vo te

Le

UN

US

WikRe ipde dit dia

50

E UN nron Tra de

60

90

E UN nron Tra de

70

transductive (Pearson R=-0.81) inductive (Pearson R=-0.86)

U CaMOOCI n. C Pa rl.

80

Average Precision (%)

90

100

WikRe ipde dit dia

transductive (Pearson R=-0.91) inductive (Pearson R=-0.93)

U CaMOOCI n. C Pa rl.

Average Precision (%)

100

MMD

(b) TIDFormer

(c) TCL

Figure 1: Performance on 9 datasets with varying MMD. Legend reports Pearson’s R between AP and MMD. The consistency across architectures is important. DyGFormer, TIDFormer, and TCL construct historical sequences in different ways, yet they degrade on the same high-shift datasets. In particular, the best existing Transformer reaches only 71.1% AP on US Legis., 66.5% on UN Trade, and 69.0% on UN Vote. This pattern suggests that the bottleneck is unlikely to be tied only to a specific sequence construction strategy. Instead, it points to a component shared by these models. 3.2

Critical neighbors as a Diagnostic Probe

To localize the cause of the shared failure, we examine which historical interactions carry the predictive signal needed for link prediction, whether existing models exploit them, and what changes between low-shift and high-shift conditions. (a) Structural centrality

Critical nodes. For a test query (u, v, t), let C(u, v, t) = (H(u, t) ∪ H(v, t)) \ {u, v} denotes the combined historical neighborhood of the source and destination before time t, where H(u, t) denotes the historical 1-hop neighborhoods of u. Not all neighbors in C are equally informative. We define a node w ∈ C(u, v, t) as a critical node if it satisfies at least one of two properties: 1) Structural centrality: w is connected to multiple other nodes in C(u, v, t) and therefore reflects local community structure; 2) Temporal stability: w has interacted more than once with another node in C(u, v, t), and both nodes have multiple historical interactions with the source or destination, thereby reflecting a stable local relationship anchored to the query pair. The critical nodes are illustrated in Fig. 2.

t₄ t₁ t₅ t₂

t₅

t₁ t₂

t₃

t₁ t₅

t₃, t₄, t₅

t₁, t₄, t₅ t₃

t

t₄ t₂

(b) Temporal stability

t₂ t₅ t₃

t₂

t

t₃

t₄

t₅

t₁, t₄

t₄

t₄

t₃

t₅

t₄, t₅

t₅

t₁, t₂

t₂, t₃

t₁

t₄ t₅ t₅ t₃ t₂

0 ≤ t₁ ≤ t₂ ≤ t₃ ≤ t₄ ≤ t₅ ≤ t ● source node

● destination node

● historical node

● critical node

Figure 2: Illustration of the two criteria for critical nodes. The purple and brown nodes denote the source and destination, gray nodes denote ordinary historical neighbors, and green nodes denote critical nodes. The dashed edge is the query interaction at time t, and solid edges are historical interactions before t.

We denote the resulting set by K(u, v, t). This definition is not intended as the only possible definition of important neighbors. It is used as a diagnostic probe: if these nodes matter more than randomly selected neighbors of equal count, then they provide a useful way to test whether attention focuses on predictive historical signals. Critical-node ablation. For each test query (u, v, t), we mask tokens corresponding to K(u, v, t) at different retention levels and evaluate AP. At a retention ratio of r%, only r% of critical nodes are kept, while the remaining critical nodes are masked. All non-critical historical tokens are preserved. Fig. 3 reports the results across all nine datasets. The results show that critical nodes carry substantial predictive signal. On low-shift datasets such as Wikipedia and Reddit, performance remains relatively high even when many critical nodes are removed, suggesting that useful signals are more redundant. On medium-shift datasets such as UCI, Enron, MOOC, and Can. Parl., models perform well when critical nodes are retained, but performance drops sharply as these nodes are removed. On high-shift datasets such as US Legis., UN Trade, and UN Vote, existing Transformer baselines already perform poorly even with all critical nodes available, and their performance further decreases after critical-node masking. This indicates that critical nodes are important, but their presence alone is insufficient under severe temporal shift. Random-neighbor control. To test whether the above degradation is simply caused by removing historical tokens, we repeat the experiment with random masking. For each retention level, we mask the same number of randomly selected historical nodes as in the critical-node ablation. The random set is sampled from the full historical neighborhood and may include critical nodes by chance. 4

MOOC 86.48

82.12

66.57

58.46

82.38

78.58

63.56

54.62

93.42

86.21

72.2

65.18

97.01

88.74

74.59

66.09

Enron 91.33

83.08

79.22

63.25

79.7

76.33

66.26

60.74

93.17

85.01

79.82

64.63

98.77

85.41

80.54

64.77

UCI 95.51

93.15

76.54

69.14

89.57

79.07

73.62

65.33

97.44

93.8

76.8

68.4

98.54

94.07

77.25

69.68

Reddit 98.82 Wikipedia 98.81

95.87

93.36

89.11

97.53

93.51

90.44

87.95

99.38

97.06

94.05

89.72

99.64

97.16

94.62

90.25

95

91.25

85.2

96.47

91.2

90.36

82.64

99.3

95.79

91.23

85.14

99.57

96.67

91.43

86.17

100% 90%

50%

0%

100% 90%

50%

0%

100% 90%

50%

0%

100% 90%

50%

0%

DyGFormer

DiffDyG

TIDFormer

TCL

UN Vote 55.55

51.79

51.04

50.13

51.9

50.13

48.37

46.8

69.01

52.36

51.95

50.16

88.75

54.82

52.02

50.22

UN Trade 66.46 US Legis. 71.11

63.17

62.6

60.13

62.21

57.86

56.66

54.68

60.84

56.19

54.6

53.72

98.97

71.63

67.79

59.73

66.87

64.12

57.44

69.59

66.22

62.14

54.98

66.51

63.26

60.44

53.61

87.52

69.82

64.27

58.68

Can. Parl. 97.36

71.44

68.04

50.82

68.67

58.92

52.45

41.86

99.84

75.12

71.58

52.17

99.63

74.09

70.56

51.27

MOOC 86.48

65.87

61.87

56.21

82.38

62.57

58.63

52.13

93.42

65.94

61.93

56.62

97.01

66.87

62.07

54.07

Enron 91.33

73.84

68.32

53.8

79.7

64.94

55.13

51.11

93.17

74.28

69.21

53.85

98.77

74.33

71.05

53.98

UCI 95.51

76.36

71.79

62.92

89.57

74.07

69.96

57.82

97.44

77.34

71.85

63.05

98.54

78.87

72.03

63.24

Reddit 98.82 Wikipedia 98.81

93.92

93.15

88.89

97.53

91.7

89.57

87.49

99.38

94.15

93.34

88.76

99.64

94.23

93.38

88.66

86.95

82.91

79.44

96.47

85.81

80.3

77.4

99.3

87.56

83.05

80.74

99.57

87.92

83.44

80.85

100% 90%

50%

0%

100% 90%

50%

0%

100% 90%

50%

0%

100% 90%

50%

0%

60

40

100

80

60

40

Figure 3: Critical-node ablation. The x-axis shows the retention ratio of critical nodes: 100% keeps all critical nodes, while 0% masks all critical nodes. Non-critical nodes are kept unchanged. DyGFormer

DiffDyG

TIDFormer

TCL

UN Vote 55.55

52.67

51.62

50.39

51.9

50.84

48.69

47.73

69.01

57.61

52.48

50.51

88.75

73.54

64.82

58.67

UN Trade US Legis. 71.11

64.38

63.18

61.45

62.21

59.34

57.67

55.28

60.84

58.42

55.16

54.38

98.97

83.56

77.14

70.05

69.23

66.69

61.57

69.59

67.68

63.2

59.65

66.51

64.5

62.32

57.84

87.52

71.17

67.15

62.19

Can. Parl. 97.36

88.24

81.23

72.16

68.67

63.95

55.36

48.39

99.84

90.01

82.87

74.29

99.63

89.65

82.63

74.53

MOOC 86.48

82.12

66.57

58.46

82.38

78.58

63.56

54.62

93.42

86.21

72.2

65.18

97.01

88.74

74.59

66.09

Enron 91.33

83.08

79.22

63.25

79.7

76.33

66.26

60.74

93.17

85.01

79.82

64.63

98.77

85.41

80.54

64.77

UCI 95.51

93.15

76.54

69.14

89.57

79.07

73.62

65.33

97.44

93.8

76.8

68.4

98.54

94.07

77.25

69.68

Reddit 98.82 Wikipedia 98.81

95.87

93.36

89.11

97.53

93.51

90.44

87.95

99.38

97.06

94.05

89.72

99.64

97.16

94.62

90.25

95

91.25

85.2

96.47

91.2

90.36

82.64

99.3

95.79

91.23

85.14

99.57

96.67

91.43

86.17

100% 90%

50%

0%

100% 90%

50%

0%

100% 90%

50%

0%

100% 90%

50%

0%

66.46

100

80

60

40

Figure 4: Random-node ablation. For each retention ratio, we mask the same number of randomly selected historical nodes as in the corresponding critical-node ablation. Thus, 0% masks a random DiffDyG set with the same DyGFormer TIDFormer TCL 100 cardinality as the full critical-node set, rather than masking all historical nodes.88.75 54.82 52.02 50.22 69.01 52.36 51.95 50.16 51.9 50.13 48.37 46.8 UN Vote 55.55 51.79 51.04 50.13 UN Trade 66.46 US Legis. 71.11

63.17

62.6

60.13

62.21

57.86

56.66

54.68

60.84

56.19

54.6

53.72

98.97

71.63

67.79

59.73

66.87

64.12

57.44

69.59

66.22

62.14

54.98

66.51

63.26

60.44

53.61

87.52

69.82

64.27

58.68

Enron 91.33

73.84

68.32

53.8

79.7

64.94

55.13

51.11

93.17

74.28

69.21

53.85

98.77

74.33

71.05

53.98

Therefore, this control asks whether the structurally99.84 and75.12 temporally defined nodes in K(u, 80v, t) are 99.63 74.09 70.56 51.27 68.67 58.92 52.45 41.86 71.58 52.17 Can. Parl. 97.36 71.44 68.04 50.82 more informative than an equally sized random subset. The comparison in Fig. confirms this. 62.07 97.01 66.87 4 54.07 82.38 62.57 58.63 52.13 93.42 65.94 61.93 56.62 MOOC 86.48 65.87 61.87 56.21 60 Across models and datasets, masking critical nodes97.44 is generally more 98.54 damaging than masking the 78.87 72.03 63.24 89.57 74.07 69.96 57.82 77.34 71.85 63.05 UCI 95.51 76.36 71.79 62.92 same number of random nodes. For example, on US Legis. with DyGFormer, masking all critical 99.64 94.23 93.38 88.66 99.38 94.15 93.34 88.76 97.53 91.7 89.57 87.49 Reddit 98.82 93.92 93.15 88.89 80.85 83.44cardinality 99.3 87.56 83.05 80.74 80.3 86.95 71.1% 96.47 85.81 whereas 98.81from 82.91 79.44to 57.4%, 77.4 random nodes Wikipedia drops AP masking with99.57 the87.92 same drops 40 100%the 90%pattern 50% 0% 100%and 90% UN 50%Vote, 0% and 90% on 50%UN 0%Trade 100% 90% 50% 0% AP only to 61.6%. Similar gaps100% appear is also visible for DiffDyG, where critical-node masking causes larger degradation than random masking under the same mask size. Since random masking can remove critical nodes by chance, this gap gives a conservative estimate of the additional predictive signal carried by critical nodes.

Implication. These ablations separate information availability from information use. Critical nodes are present in the input across all datasets. On medium-shift datasets, the same architectures can use them effectively. On high-shift datasets, the models still fail even when the critical nodes remain in the input. Thus, the main bottleneck is not simply missing information or insufficient capacity. Rather, it is the model’s ability to assign enough relative weight to predictive historical tokens among many less informative ones. Since the same failure appears across different Transformer architectures, the evidence points to the shared attention mechanism. 3.3

Hypothesis: Attention Dispersion under Shift

We hypothesize that temporal distribution shift weakens√the contrast produced by standard softmax attention. Standard attention computes softmax(QK ⊤ / d)V, where the softmax distributes probability mass over all historical tokens. When test-time interaction patterns differ from training-time patterns, the relative gaps in QK ⊤ can become smaller. After softmax, smaller score gaps lead to flatter attention distributions. The model may still assign slightly higher mass to useful tokens, but the contrast is too weak for downstream prediction to reliably extract their signal. We refer to this failure mode as attention dispersion. The hypothesis makes two testable predictions. First, if attention allocation is the shared bottleneck, then replacing only the attention module in existing CTDG Transformers should improve performance on shifted datasets while leaving their sequence construction unchanged. Second, a successful replacement should produce less dispersed attention distributions and assign more mass to critical nodes. §4 validates the diagnosis by testing both predictions directly. 5

Table 2: Effect of adding Differential Attention (DA) to Transformer-based baselines the transductive setting. ∆ denotes the absolute AP improvement after adding DA. Model

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

Avg. ∆

TIDFormer 99.30±0.03 99.38±0.03 97.44±0.10 93.17±0.06 93.42±0.15 99.84±0.09 66.51±2.56 60.84±1.09 69.01±1.24 – TIDFormer +DA 99.37±0.03 99.46±0.02 98.00±0.11 95.18±0.11 95.37±0.63 99.87±0.31 76.23±0.62 84.13±1.02 82.36±0.54 – ∆ +0.07 +0.08 +0.56 +2.01 +1.95 +0.03 +9.72 +23.29 +13.35 +5.67 DyGFormer 98.81±0.02 98.82±0.06 95.51±0.20 91.33±0.16 86.48±0.06 97.36±0.45 71.11±0.59 66.46±1.29 55.55±0.42 – DyGFormer +DA 98.89±0.09 99.17±0.05 96.18±0.01 94.12±0.15 95.21±0.26 97.41±0.03 82.15±0.34 90.73±0.17 81.94±0.77 – ∆ +0.08 +0.35 +0.67 +2.79 +8.73 +0.05 +11.04 +24.27 +26.39 +8.26 TCL TCL +DA ∆

4

96.47±0.16 97.53±0.02 89.57±1.63 79.70±0.71 82.38±0.24 68.67±2.67 69.59±0.48 62.21±0.03 51.90±0.30 – 97.52±0.37 98.32±0.08 91.53±0.54 85.72±0.23 85.41±0.40 71.28±0.95 73.26±0.69 66.04±0.12 54.80±0.08 – +1.05 +0.79 +1.96 +6.02 +3.03 +2.61 +3.67 +3.83 +2.90 +2.87

Validating the Diagnosis

Differential attention as an intervention. We test both predictions in §3.3 using differential attention as a targeted intervention. Differential attention [29] is suitable for this test because it changes the attention computation while leaving the surrounding architecture unchanged. Given an Q input representation Z ∈ RN ×din , we form two query-key pairs   and a shared value: [Q1 ; Q2 ] = ZW , ⊤ Q K m [K1 ; K2 ] = ZWK , V = ZWV . Let Am = softmax √d m , m ∈ {1, 2}. Diff attention computes attn

DiffAttn(Z) = (A1 − λA2 )V,

(1)

where λ is a learnable scalar. The subtraction suppresses attention components that appear similarly in both maps and preserves token-level signals that are more distinctive. This matches the diagnosis in §3.3: if standard attention becomes too diffuse under temporal shift, subtracting a common-mode attention component should increase contrast among historical tokens. Importantly, this intervention does not require changing the feature and sequence construction. 4.1

Cross-architecture transferability

We first test whether this attention-level change transfers across existing CTDG Transformers. We replace standard multi-head self-attention with differential attention in three representative baselines: DyGFormer [31], TIDFormer [15], and TCL [24]. These models use different sequence construction strategies, including neighbor co-occurrence encoding, calendar-based temporal partitioning, and graph-topology-aware temporal encoding. For each model, all other components are kept unchanged, including the input sequence, feature channels, position or time encoding, and training procedure. We denote the resulting variants as DyGFormer +DA, TIDFormer +DA, and TCL +DA. Tab. 2 reports the transductive AP results. Adding differential attention improves all three baselines, with average gains of +8.26 for DyGFormer, +5.67 for TIDFormer, and +2.87 for TCL. The gains are concentrated on the high-shift datasets identified in §3.1. For example, DyGFormer +DA improves over DyGFormer by +11.04, +24.27, and +26.39 AP on US Legis., UN Trade, and UN Vote, respectively. TIDFormer +DA shows the same pattern, with gains of +9.72, +23.29, and +13.35 on the same datasets. In contrast, on low-shift datasets such as Wikipedia and Reddit, where the original baselines are already near saturation, the gains are small. These results support the diagnosis in two ways. First, the improvement appears across architectures with different sequence construction designs, while the only modified component is attention. Second, the gains are largest exactly where temporal shift is strongest. This pattern would be difficult to explain if the main limitation were a model-specific input construction issue. It is more consistent with the hypothesis that attention allocation is a shared bottleneck under shift. 4.2

Attention-level mechanism verification

We next test whether differential attention changes attention behavior in the way predicted by §3.3. Since differential attention forms a signed map B = A1 − λA2 , entropy and attention-mass measurements require a nonnegative normalized map. We therefore compute attention statistics using [B ]+ b DA , where [·]+ = max(·, 0) and ϵ is a small the normalized positive component A = P [Bij ij iℓ ]+ +ϵ ℓ constant for numerical stability. For standard attention, we use the usual softmax map. This gives a comparable probability distribution over historical tokens for both attention mechanisms. 6

Table 3: Comparison of attention entropy between standard attention and differential attention in the last layer. The reported values are averaged over attention heads. Attention

Wikipedia

Reddit

UCI

Enron

Mooc

CanParl

USLegis

UNTrade

UNVote

Differential Attention Standard Attention

2.8075 3.2894

2.9179 3.8371

3.1967 3.8586

3.0203 3.3862

2.5780 3.9390

2.7699 3.6248

2.3099 2.9130

2.8705 3.9199

2.3461 3.2337

Table 4: Proportion of attention scores assigned to critical nodes in the last layer. Attention

Wikipedia

Reddit

UCI

Enron

Mooc

CanParl

USLegis

UNTrade

UNVote

Differential Attention Standard Attention

0.8006 0.7692

0.7866 0.7777

0.7951 0.7315

0.9873 0.9836

0.7958 0.7840

0.8187 0.7865

0.7800 0.7768

0.9706 0.7445

0.9801 0.8255

Table 5: Proportion of critical nodes among the top-5%-ranked attention tokens in the last layer. Attention

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl. US Legis. UN Trade UN Vote

Avg.

Differential Attention Standard Attention

0.9395 0.7850

0.9055 0.8945

0.8990 0.6785

0.9880 0.9825

0.8985 0.8850

0.8840 0.8220

0.9168 0.8403

0.7845 0.7810

0.9710 0.8470

0.9810 0.8875

P Attention entropy. For each test query, we compute the Shannon entropy H(a) = − i ai log ai over historical tokens, averaged across heads and test examples. Lower entropy indicates more concentrated attention. Tab. 3 compares entropy at each layer when using differential vs. standard attention, with all other components identical. The pattern matches the prediction. Differential attention produces lower entropy than standard attention on most datasets, with especially large reductions on shifted datasets. For example, entropy decreases from 2.91 to 2.31 on US Legis., from 3.23 to 2.35 on UN Vote, and from 3.91 to 2.87 on UN Trade. On low-shift datasets where standard attention is already relatively concentrated, the gap is generally smaller. This supports the first mechanism prediction: differential attention reduces attention dispersion, and the reduction is strongest where the diagnosis predicts dispersion to be most harmful. Attention mass on critical nodes. Lower entropy alone is not sufficient, since attention could become sharper on the wrong tokens. We therefore measure whether the recovered concentration falls on the critical nodes defined in §3.2. We use two statistics. The first is the total attention mass assigned to critical nodes. The second is the proportion of critical nodes among the top-ranked attention tokens, with the top-5% region reflecting the model’s most confident attention assignments. Tab. 4 shows that differential attention assigns more total mass to critical nodes than standard attention on all 9 datasets, with large gains on shifted datasets such as UN Trade and UN Vote. Tab. 5 further shows that critical nodes are more frequent among the top-5% attended tokens under differential attention, increasing the average ratio from 0.84 to 0.92. Thus, differential attention does not merely sharpen attention arbitrarily. It shifts attention toward the same structurally and temporally important nodes whose predictive role was established by the ablation study in §3.2. Together, the plug-in experiments and attention-level measurements validate the diagnostic chain. Replacing only attention improves three different CTDG Transformers, with the largest gains on high-shift datasets. The same replacement also reduces attention entropy and increases focus on critical nodes. These results support the central conclusion of §3: the shared failure mode is not simply missing historical information or insufficient sequence construction, but dispersed attention allocation under temporal distribution shift.

5

DiffDyG: A Reference Implementation

The validation in §4 shows that replacing standard attention with differential attention improves several CTDG Transformers under temporal distribution shift. We now instantiate this finding in a single model, DiffDyG, which serves as the configuration used for end-to-end comparison with state-of-the-art baselines in §6. DiffDyG keeps the overall structure of a standard Transformer-based CTDG model, while replacing self-attention with differential attention and using standard temporal and structural encodings for dynamic graph inputs. Specifically, differential attention provides stronger temporal dispersion ability, allowing DiffDyG to use a smaller embedding dimension than existing Transformer-based baselines. Architecture overview. DiffDyG consists of three components. First, it constructs a historical token sequence for each query node from its recent interactions. Second, each token combines node, edge, temporal, co-occurrence, and structural information. Third, the resulting sequence is processed by 7

Table 6: Performance (Average Precision) comparison in the transductive setting under random negative sampling strategies. Average Rank (AR) is computed across the 9 datasets. The best and second best results are shown in bold and underlined. Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 99.57±0.03 99.64±0.14 98.54±0.03 98.77±0.46 97.01±0.22 99.63±0.09 87.52±0.47 98.97±0.01 88.75±1.45 1.11 RepeatMixer 99.16±0.02 99.22±0.01 96.74±0.08 92.66±0.07 92.76±0.10 81.30±0.76 68.36±0.28 64.00±0.06 55.33±0.07 4.89 TIDFormer 99.30±0.03 99.38±0.03 97.44±0.10 93.17±0.06 93.42±0.15 99.84±0.09 66.51±2.56 60.84±1.09 69.01±1.24 3.78 DyGFormer 98.81±0.02 98.82±0.06 95.51±0.20 91.33±0.16 86.48±0.06 97.36±0.45 71.11±0.59 66.46±1.29 55.55±0.42 4.78 99.09±0.04 99.22±0.01 96.85±0.08 92.48±0.10 94.18±0.07 86.09±0.14 78.69±0.12 76.92±0.06 66.40±0.12 3.00 CNEN TGAT 96.94±0.06 98.52±0.02 79.63±0.70 71.12±0.97 85.84±0.15 70.73±0.72 68.52±3.16 61.47±0.18 52.21±0.98 8.78 TCL 96.47±0.16 97.53±0.02 89.57±1.63 79.70±0.71 82.38±0.24 68.67±2.67 69.59±0.48 62.21±0.03 51.90±0.30 9.22 TGN 98.45±0.06 98.63±0.06 92.34±1.04 86.53±1.11 89.15±1.60 70.88±2.34 75.99±0.58 65.03±1.37 65.72±2.17 5.89 98.76±0.03 99.11±0.01 95.18±0.06 89.56±0.09 80.15±0.25 69.82±2.34 70.58±0.48 65.39±0.12 52.84±0.10 6.67 CAWN EdgeBank 90.37±0.00 94.86±0.00 76.20±0.00 83.53±0.00 57.97±0.00 64.55±0.00 58.39±0.00 60.41±0.00 58.49±0.00 10.00 GraphMixer 97.25±0.03 97.31±0.01 93.25±0.57 82.25±0.16 82.78±0.15 77.04±0.46 70.74±1.02 62.61±0.27 52.11±0.16 7.78

stacked Transformer layers whose self-attention modules are replaced with differential attention. Source and destination nodes are encoded separately, then combined for link prediction. 5.1

Input construction

We describe the input construction for the source node u at time t. The destination node v is processed in the same way. Token sequence. For source node u, we form a sequence of K+1 tokens. The first token represents u itself, and the remaining K tokens represent its most recent interaction neighbors before time t, ordered chronologically. Unless otherwise stated, we use K = 20 first-hop neighbors. For three smaller benchmarks (Wikipedia, UCI, Can. Parl.) we additionally include 5 second-hop neighbors. Details and motivation are in Appendix A.3.1. Feature channels. Each token aggregates five channels: (i) the node feature of u (or its neighbor vj′ for j > 0); (ii) the edge feature of the corresponding interaction; (iii) a temporal encoding of the elapsed time ∆tj = t − tj via cosine basis with learnable frequencies; (iv) the co-occurrence frequency of vj′ in the histories of both u and the destination v, following [31]; (v) the spatial distance (hop count) from u. Each channel is linearly projected to a common dimension d and the five projections are concatenated into a 5d-dimensional token representation. Stacking the tokens yields the input matrix Ztu ∈ R(K+1)×5d . Specifically, for the source node u (the 0-th token), its edge feature and temporal encoding feature are 0, and its spatial distance is 0. Positional encoding. We apply Rotary Positional Embeddings (RoPE) [20] to the token sequence before attention. RoPE provides relative position information within the historical sequence, while the elapsed-time encoding captures the actual temporal interval between the query time and each historical interaction. 5.2

Differential attention encoder

Source and destination input matrices Ztu and Ztv are processed independently through L = 2 stacked differential attention transformer layers. Each layer contains a differential attention sublayer followed by a SwiGLU feed-forward network, with RMSNorm before each sublayer and residual connections around each. The output sequences are average-pooled along the token dimension to produce nodelevel embeddings Yut , Yvt ∈ Rdout , which are concatenated and passed through an MLP classifier to produce the link probability ptu,v . The model is trained end-to-end with binary cross-entropy loss using one negative edge per positive edge. The full training procedure is summarized in Algorithm 1 in the appendix.

6

Experimental Evaluation of DiffDyG

We evaluate whether DiffDyG translates the attention-level fix validated in §4 into end-to-end performance gains. The main text focuses on transductive link prediction under random negative sampling, which is the standard setting used for direct comparison. Full results for AUC-ROC, inductive evaluation, historical negative sampling, inductive negative sampling, efficiency, and additional analyses and extension to dynamic node prediction task are reported in Appendix B. 8

Table 7: Ablation study on 9 datasets in transductive settings. Results are reported in Average Precision (AP) AP

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

DiffDyG w/o RoPE w/o DA w/o SE

99.57±0.03 99.02±0.04 98.98±0.36 99.27±0.35

99.64±0.14 99.16±0.02 99.02±0.15 99.34±0.32

98.54±0.03 97.14±2.00 96.94±0.29 98.06±0.01

98.77±0.46 97.81±3.43 92.83±0.24 98.15±0.65

97.01±0.22 96.37±2.09 88.61±0.38 95.87±0.06

99.63±0.09 97.56±0.30 97.42±0.17 98.87±0.01

87.52±0.47 85.09±2.41 74.51±0.08 85.89±0.37

98.97±0.01 95.50±4.00 68.11±0.32 96.23±0.58

88.75±1.45 84.98±2.72 57.21±0.27 86.43±0.01

Setup. We use the 9 CTDG benchmarks introduced in §3, with chronological 70%/15%/15% train/validation/test splits. We report Average Precision (AP) averaged over 5 random seeds. Baselines include memory-based models (TGN [17], EdgeBank [16]), random-walk-based models (CAWN [26]), temporal message-passing transformers (TGAT [28], TCL [24], DyGFormer [31], TIDFormer [15]), sequence-based models (GraphMixer [4], RepeatMixer [33]), and structureencoding models (CNEN [2]). Full baseline descriptions and implementation details are in Appendix A.2 and Appendix A.3. 6.1

Performance comparison with baselines

Tab. 6 reports the transductive AP results under random negative sampling. DiffDyG achieves the best average rank of 1.11, ranking first on 7 out of 9 datasets and second on the remaining 2. The strongest gains appear on the high-shift datasets identified in §3.1. On US Legis., DiffDyG outperforms the second-best baseline from 78.69% to 87.52%. On UN Trade, it improves from 76.92% to 98.97%. On UN Vote, it improves from 69.01% to 88.75%. These are precisely the datasets where existing CTDG Transformers plateaued despite changes to sequence construction and feature design. On low-shift datasets such as Wikipedia and Reddit, where existing methods are already near saturation, DiffDyG preserves performance, reaching 99.57% and 99.64% AP. On medium-shift datasets such as UCI and Enron, it further improves AP to 98.54% and 98.77%. Thus, the attentionlevel fix does not trade off performance on easier benchmarks for gains on shifted ones. The gain pattern matches the diagnosis in §3. Improvements are small on saturated low-shift datasets, moderate on medium-shift datasets, and largest on high-shift datasets. This is consistent with the view that attention dispersion becomes most harmful under temporal distribution shift, and that differential attention primarily helps by restoring sharper allocation over informative historical tokens. The full results in Appendix B.2 show the same trend. DiffDyG achieves the best average ranking across all three negative sampling strategies for both AP and AUC-ROC in both transductive and inductive settings, confirming that its advantage is not tied to a specific evaluation protocol. 6.2

Component ablation

We next examine which components drive DiffDyG’s gains. We compare the full model against three variants: w/o DA, which replaces differential attention with standard multi-head attention; w/o RoPE, which removes rotary positional embeddings; and w/o SE, which removes spatial distance encoding. The full dataset-wise ablation is reported in Appendix B.5; the main text summarizes the key effects. Removing differential attention causes the largest drop, with an average AP decrease of 10.53 points across the 9 datasets. The drop is especially severe on high-shift datasets: −13.01 on US Legis., −30.86 on UN Trade, and −31.54 on UN Vote. In contrast, removing RoPE causes an average drop of 1.75 points, and removing spatial encoding causes an average drop of 1.14 points. This confirms within DiffDyG what §4.1 showed across architectures: differential attention is the dominant source of improvement, while the supporting encodings provide smaller refinements.

7

Discussion and Limitations

This work studied why Transformer-based CTDG models underperform on benchmarks with large temporal distribution shift. Instead of focusing on sequence construction or feature design, we examined whether the bottleneck lies in the attention mechanism itself. Our diagnosis shows that shifted datasets still contain predictive historical signals, captured by structurally and temporally important critical nodes, but standard attention does not allocate sufficient mass to these nodes. The failure is therefore not simply missing information or insufficient model capacity. It is an attention9

allocation problem: under shift, attention becomes too dispersed to reliably separate predictive historical neighbors from less informative ones. We validated this diagnosis by using differential attention as a targeted intervention. Replacing only the attention module in three existing CTDG Transformers consistently improved performance, with the largest gains on the high-shift datasets identified by the diagnosis. Direct attention-level measurements further showed that differential attention reduces entropy and increases attention mass on critical nodes. Building on these findings, DiffDyG provides a reference implementation of this attention-level fix and achieves SOTA results across the standard CTDG benchmarks. These results suggest that further progress on difficult CTDG datasets may require not only better historical sequence construction, but also better mechanisms for allocating attention over those sequences. Limitations. Our analysis is primarily empirical. Although the ablation studies, plug-in experiments, and attention measurements provide consistent evidence for attention dispersion, we do not provide a formal characterization of when softmax attention becomes dispersed under temporal shift or when differential attention is guaranteed to help. Developing such theory is an important direction for future work. In addition, our definition of critical nodes is a diagnostic probe based on structural centrality and temporal stability, rather than a unique or optimal definition of predictive historical neighbors. Other definitions, including learned or task-adaptive ones, may reveal additional signals. Finally, our evaluation focuses on dynamic link prediction on standard CTDG benchmarks. Extending the diagnosis to other dynamic graph tasks and to substantially larger industrial graphs remains future work.

References [1] J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Proc. EMNLP, pages 4895–4901. Association for Computational Linguistics, 2023. [2] K. Cheng, L. Peng, J. Ye, L. Sun, and B. Du. Co-neighbor encoding schema: A light-cost structure encoding method for dynamic link prediction. In Proc. KDD, pages 421–432. ACM, 2024. [3] W. Cong, Y. Wu, Y. Tian, M. Gu, Y. Xia, M. Mahdavi, and C. J. Chen. Dynamic graph representation learning via graph transformer networks. CoRR, abs/2111.10447, 2021. [4] W. Cong, S. Zhang, J. Kang, B. Yuan, H. Wu, X. Zhou, H. Tong, and M. Mahdavi. Do we really need complicated model architectures for temporal networks? In Proc. ICLR. OpenReview.net, 2023. [5] T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. In Proc. NeurIPS, 2022. [6] A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola. A kernel two-sample test. The journal of machine learning research, 13(1):723–773, 2012. [7] Z. Han, J. Jiang, Y. Wang, Y. Ma, and V. Tresp. The graph hawkes network for reasoning on temporal knowledge graphs. In Learning with Temporal Point Processes Workshop at the 33rd Conference on Neural Information Processing Systems (TPP@ NeurIPS 2019)), 2019. [8] W. Hu, Y. Yang, Z. Cheng, C. Yang, and X. Ren. Time-series event prediction with evolutionary state graph. In Proc. WSDM, pages 580–588. ACM, 2021. [9] S. Huang, Y. Hitti, G. Rabusseau, and R. Rabbany. Laplacian change point detection for dynamic graphs. In Proc. KDD, pages 349–358. ACM, 2020. [10] X. Huang, Y. Yang, Y. Wang, C. Wang, Z. Zhang, J. Xu, L. Chen, and M. Vazirgiannis. Dgraph: A large-scale financial dataset for graph anomaly detection. In Advanced in NeurIPS, 2022. [11] S. Kumar, X. Zhang, and J. Leskovec. Predicting dynamic embedding trajectory in temporal interaction networks. In Proc. KDD, pages 1269–1278. ACM, 2019. [12] Y. Li, Y. Shen, L. Chen, and M. Yuan. Zebra: When temporal graph neural networks meet temporal personalized pagerank. Proc. VLDB Endow., 16(6):1332–1345, 2023. 10

[13] X. Lu, L. Sun, T. Zhu, and W. Lv. Improving temporal link prediction via temporal walk matrix projection. In Advanced in NeurIPS, 2024. [14] A. Pareja, G. Domeniconi, J. Chen, T. Ma, T. Suzumura, H. Kanezashi, T. Kaler, T. B. Schardl, and C. E. Leiserson. Evolvegcn: Evolving graph convolutional networks for dynamic graphs. In Proc. AAAI, pages 5363–5370. AAAI Press, 2020. [15] J. Peng, Z. Wei, and Y. Ye. Tidformer: Exploiting temporal and interactive dynamics makes A great dynamic graph transformer. In Proc. KDD, pages 2245–2256. ACM, 2025. [16] F. Poursafaei, S. Huang, K. Pelrine, and R. Rabbany. Towards better evaluation for dynamic link prediction. In Advanced in NeurIPS, 2022. [17] E. Rossi, B. Chamberlain, F. Frasca, D. Eynard, F. Monti, and M. M. Bronstein. Temporal graph networks for deep learning on dynamic graphs. CoRR, abs/2006.10637, 2020. [18] A. Sankar, Y. Wu, L. Gou, W. Zhang, and H. Yang. Dysat: Deep neural representation learning on dynamic graphs via self-attention networks. In Proc. WSDM, pages 519–527. ACM, 2020. [19] A. H. Souza, D. Mesquita, S. Kaski, and V. Garg. Provably expressive temporal graph networks. In Advanced in NeurIPS, 2022. [20] J. Su, M. H. M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. [21] J. Su, D. Zou, and C. Wu. PRES: toward scalable memory-based dynamic graph neural networks. In ICLR. OpenReview.net, 2024. [22] R. Trivedi, H. Dai, Y. Wang, and L. Song. Know-evolve: Deep temporal reasoning for dynamic knowledge graphs. In Proc. ICML, volume 70, pages 3462–3471. PMLR, 2017. [23] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In Advanced in NIPS, pages 5998–6008, 2017. [24] L. Wang, X. Chang, S. Li, Y. Chu, H. Li, W. Zhang, X. He, L. Song, J. Zhou, and H. Yang. TCL: transformer-based dynamic graph modelling via contrastive learning. CoRR, abs/2105.07944, 2021. [25] X. Wang, D. Lyu, M. Li, Y. Xia, Q. Yang, X. Wang, X. Wang, P. Cui, Y. Yang, B. Sun, and Z. Guo. APAN: asynchronous propagation attention network for real-time temporal graph embedding. In Proc. SIGMOD, pages 2628–2638. ACM, 2021. [26] Y. Wang, Y. Chang, Y. Liu, J. Leskovec, and P. Li. Inductive representation learning in temporal networks via causal anonymous walks. In ICLR. OpenReview.net, 2021. [27] Y. Wu, Y. Fang, and L. Liao. On the feasibility of simple transformer for dynamic graph modeling. In Proc. Web Conference, pages 870–880. ACM, 2024. [28] D. Xu, C. Ruan, E. Körpeoglu, S. Kumar, and K. Achan. Inductive representation learning on temporal graphs. In ICLR. OpenReview.net, 2020. [29] T. Ye, L. Dong, Y. Xia, Y. Sun, Y. Zhu, G. Huang, and F. Wei. Differential transformer. In Proc. ICLR. OpenReview.net, 2025. [30] J. You, T. Du, and J. Leskovec. ROLAND: graph learning framework for dynamic graphs. In Proc. KDD, pages 2358–2366. ACM, 2022. [31] L. Yu, L. Sun, B. Du, and W. Lv. Towards better dynamic graph learning: New architecture and unified library. In Advanced in NeurIPS, 2023. [32] Z. Zhao, X. Zhu, T. Xu, A. Lizhiyu, Y. Yu, X. Li, Z. Yin, and E. Chen. Time-interval aware share recommendation via bi-directional continuous time dynamic graphs. In Proc. SIGIR, pages 822–831. ACM, 2023. [33] T. Zou, Y. Mao, J. Ye, and B. Du. Repeat-aware neighbor sampling for dynamic graph learning. In Proc. KDD, pages 4722–4733. ACM, 2024. 11

A

Additional experimental details

A.1

Details of datasets Table 8: Summary of dynamic graph datasets

Datasets

Domains

#Nodes

#Links

#Node & Link Feat.

Bipartite

Wikipedia Reddit MOOC Enron UCI Can. Parl. US Legis. UN Trade UN Vote

Social Social Interaction Social Social Politics Politics Economics Politics

9,227 10,984 7,144 184 1,899 734 225 255 201

157,474 672,447 411,749 125,235 59,835 74,478 60,396 507,497 1,035,742

– & 172 – & 172 –&4 –&– –&– –&1 –&1 –&1 –&1

✓ ✓ ✓ ✗ ✗ ✗ ✗ ✗ ✗

Duration

Time Granularity

# Steps

1 month 1 month 17 months 3 years 196 days 14 years 12 congresses 32 years 72 years

Unix timestamps Unix timestamps Unix timestamps Unix timestamps Unix timestamps years congresses years years

152,757 669,065 345,600 22,632 58,911 14 12 32 72

We conduct experiments on 9 widely used dynamic graph benchmarks. These datasets are collected Poursafaei et al. [16]. Tab. 8 provides a summary of their key statistics. Below we describe each dataset. Wikipedia is a bipartite dataset that captures editing activities on Wikipedia during one month. The nodes represent users and pages. Each edge corresponds to an edit action at a specific timestamp and is associated with a 172-dimensional LIWC feature. The dataset also includes dynamic labels that indicate whether a user was temporarily banned from editing. Reddit is another bipartite graph from the social domain. It captures posting activity of users on different subreddits during one month. The nodes represent users and subreddits. Each edge denotes a posting event at a specific timestamp and is associated with a 172-dimensional LIWC feature vector. Dynamic labels indicate whether a user was banned from posting. MOOC is a bipartite graph constructed from interactions on an online education platform. The nodes represent students and course units such as videos or problem sets. Each edge represents a student accessing a course unit at a specific timestamp and includes a 4-dimensional feature vector. Enron is a non-bipartite graph that captures email communication among Enron employees over three years. The nodes represent employees. Each edge corresponds to an email exchange at a recorded timestamp. UCI is a non-bipartite graph that records message exchanges among university students over approximately 196 days. The nodes represent students. Each edge represents a directed message sent at a specific timestamp. Can. Parl. is a political graph that tracks voting behavior among Canadian Members of Parliament (MPs) from 2006 to 2019. The nodes represent Members of Parliament. An edge exists between two MPs in a given year when both vote “yes” on the same bill. The edge weight equals the number of such shared “yes” votes in that year. US Legis. is a co-sponsorship network of the U.S. Senate, spanning eight years. The nodes represent senators. An edge exists when two legislators co-sponsor the same bill. The weight of an edge reflects the number of cosponsorships within one congressional session. UN Trade is a long-term trade network that captures agricultural and food exchanges among 181 countries over 32 years. The edge weight between two countries reflects the total normalized value of imports and exports between them in this domain. UN Vote is a voting network of the United Nations General Assembly over 72 years. For each resolution, an edge weight between two countries increases by one when both cast a “yes” vote on that resolution. A.2

Details of Baselines

TGAT [28] learns node embeddings by attending over each node’s temporal–topological neighbors via a self-attention framework, coupled with a time-encoding function that captures temporal patterns in dynamic graphs. 12

TCL [24] processes dynamic graphs by first utilizing a graph-topology-aware transformer to capture temporal and topological information. It then employs a two-stream encoder to separately extract representations from the temporal neighborhoods of two interacting nodes. A co-attentional transformer models the inter-dependencies between these nodes at a semantic level, enhanced by contrastive learning. TGN [17] maintains an evolving memory for every node and updates it whenever a new interaction arrives, using a message function, a message aggregator, and a memory updater; an embedding module then produces time-aware node representations. CAWN [26] first samples multiple causal anonymous walks per node to capture the dynamics and relative identities, encodes each walk with recurrent networks, and aggregates the walk encodings to form the final node representation for downstream tasks. EdgeBank [16] is a parameter-free, memory-based method tailored to transductive dynamic link prediction. It predicts an interaction as positive if it has been retained in memory and negative otherwise, operating purely over stored historical edges. GraphMixer [4] shows that a fixed (non-trainable) time encoding can outperform a learned one. It embeds temporal links with an MLP-Mixer–style link encoder and summarizes node features via neighbor mean-pooling. DyGFormer [31] is Transformer-based and emphasizes source–target correlation modeling through a neighbor co-occurrence encoding over historical sequences. It further introduces a patching trick that splits long histories into patches for efficient, effective learning. RepeatMixer [33] learns temporal interaction patterns with an MLP encoder driven by an evolving repeat-behavior sampling scheme (covering first- and higher-order repeats), and adopts a time-aware aggregation that adaptively weighs representations from different orders according to their temporal salience. CNEN [2] proposes a lightweight dynamic graph model that stores structural signals in a hashtablebased co-neighbor memory with short- and long-term components; it fuses node, edge, temporal, and structural cues via MLPs to produce time-aware node embeddings for efficient link prediction. TIDFormer [15] is a Transformer architecture that leverages calendar-based time partitioning and derives informative interaction embeddings using only sampled first-order neighbors on both bipartite and non-bipartite graphs; a simple decomposition module tracks shifts in historical interaction patterns to model temporal and interactive factors. A.3

Details of Model Configurations, Implementation, and Evaluation

For DiffDyG, we set the number of Transformer layers to 2, the number of attention heads to 2, and the dropout rate to 0.2. We set the dimension d to 36 and the per-head attention dimension dattn to 45. By default, we sample 20 first-order neighbors for both the source and destination nodes to construct the input sequence. The hyperparameters of all baseline methods follow the default settings reported in their original papers. To ensure a fair comparison, all models are trained for up to 100 epochs with early stopping based on a patience of 5 epochs. We use the Adam optimizer with a learning rate of 1 × 10−4 , a batch size of 200, and a weight decay of 1 × 10−4 . All experiments are conducted on a single NVIDIA GeForce RTX 4090 GPU. All experiments use 5 random seeds and we report mean and standard deviation. The default hyperparameter settings of our method are summarized in Tab. 9. For negative sampling, we adopt random negative sampling by default. A.3.1

Multi-Hop Neighborhood Extension

The input construction described in §5.1 uses only the direct (1-hop) interaction neighbors of each node. In some smaller-scale datasets (namely, Wikipedia, UCI, and Can. Parl.), the 1-hop interaction is relatively sparse. Therefore, we apply a straightforward extension to 2-hop neighborhoods for these 3 datasets, where the additional computational cost is manageable. Extended token sequence. Starting from the 1-hop neighbors {(v1′ , t1 ), . . . , (vK , tK )} of source node u, we recursively expand: for each hop level h = 2, . . . , H, we sample Kh most recent 13

interaction neighbors of each (h−1)-hop neighbor, excluding nodes already in the sequence. The full token sequence then becomes  (0)  Ztu = |{z} z ; z(1) , . . . , z(K) ; z(K+1) , . . . ; · · · ∈ R(1+K+K2 +···+KH )×5d . | {z } | {z } source

1-hop

(2)

2-hop

Per-token features. The five feature channels are computed identically to the 1-hop case, with one (j) natural change: the spatial distance embedding kS now encodes hop distance h ∈ {0, 1, . . . , H} rather than being uniformly 1 for all neighbors. This allows the attention mechanism to distinguish neighbors at different structural distances from the source node. Ordering. Within each hop level, tokens are ordered chronologically by interaction time. Hop levels are concatenated in increasing order (1-hop before 2-hop, etc.), so the sequence reflects both temporal recency and structural proximity. Computational considerations. Obtaining multi-hop samples can be expensive for large and dense graphs. Therefore, we only apply 2-hop sampling on the three datasets. Additionally, multihop sampling increases the sequence length and, consequently, the cost of attention, which scales quadratically. In our experiments, we sample K=20 first-hop and K2 =5 second-hop neighbors (H=2), extending the sequence from 21 to 26 tokens per node. This overhead is modest. For the remaining six larger datasets, we use 1-hop sampling only, and the performance difference is minor (see Tab. 15 and 16 for the hop ablation). Table 9: Default Hyperparameters. Hyperparameters

Values

Number of Transformer Layers L Number of Attention Heads H Dimension of Time Intervals Encoding dT Dimension of Interaction Frequency dC Dimension of Structural Distance Encoding dS d dattn Number of Sampled 1-hop Neighbors K Dropout Rate Learning Rate Batch Size Weight Decay

2 2 100 36 1 36 45 20 0.2 0.0001 200 0.0001

B

Additional experimental results

B.1

Additional results on dynamic node classification

In addition to dynamic link prediction, we further evaluate DiffDyG on dynamic node classification to examine whether the proposed model remains effective for node-level temporal prediction tasks. Following the experimental setup of DyGFormer [31], we conduct experiments on the standard Wikipedia and Reddit benchmarks and report AUC-ROC. As shown in Tab. 14, DiffDyG achieves the best result on Reddit and the second-best result on Wikipedia. Although CNEN obtains a slightly higher AUC-ROC on Wikipedia, its performance drops on Reddit, while DiffDyG remains consistently strong on both datasets. As a result, DiffDyG achieves the best overall AR of 1.5 among all compared methods. These results indicate that the proposed model is not limited to dynamic link prediction, but also generalizes well to node-level temporal prediction tasks. 14

Table 10: Performance (Average Precision) comparison in the transductive setting. Average Rank (AR) is computed across the 9 datasets. The best and second best results are shown in bold and underlined. Random Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 99.57±0.03 99.64±0.14 98.54±0.03 98.77±0.46 97.01±0.22 99.63±0.09 87.52±0.47 98.97±0.01 88.75±1.45 1.11 RepeatMixer 99.16±0.02 99.22±0.01 96.74±0.08 92.66±0.07 92.76±0.10 81.30±0.76 68.36±0.28 64.00±0.06 55.33±0.07 4.89 TIDFormer 99.30±0.03 99.38±0.03 97.44±0.10 93.17±0.06 93.42±0.15 99.84±0.09 66.51±2.56 60.84±1.09 69.01±1.24 3.78 DyGFormer 98.81±0.02 98.82±0.06 95.51±0.20 91.33±0.16 86.48±0.06 97.36±0.45 71.11±0.59 66.46±1.29 55.55±0.42 4.78 CNEN 99.09±0.04 99.22±0.01 96.85±0.08 92.48±0.10 94.18±0.07 86.09±0.14 78.69±0.12 76.92±0.06 66.40±0.12 3.00 TGAT 96.94±0.06 98.52±0.02 79.63±0.70 71.12±0.97 85.84±0.15 70.73±0.72 68.52±3.16 61.47±0.18 52.21±0.98 8.78 TCL 96.47±0.16 97.53±0.02 89.57±1.63 79.70±0.71 82.38±0.24 68.67±2.67 69.59±0.48 62.21±0.03 51.90±0.30 9.22 98.45±0.06 98.63±0.06 92.34±1.04 86.53±1.11 89.15±1.60 70.88±2.34 75.99±0.58 65.03±1.37 65.72±2.17 5.89 TGN CAWN 98.76±0.03 99.11±0.01 95.18±0.06 89.56±0.09 80.15±0.25 69.82±2.34 70.58±0.48 65.39±0.12 52.84±0.10 6.67 EdgeBank 90.37±0.00 94.86±0.00 76.20±0.00 83.53±0.00 57.97±0.00 64.55±0.00 58.39±0.00 60.41±0.00 58.49±0.00 10.00 GraphMixer 97.25±0.03 97.31±0.01 93.25±0.57 82.25±0.16 82.78±0.15 77.04±0.46 70.74±1.02 62.61±0.27 52.11±0.16 7.78

Historical Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 90.13±0.14 84.01±1.14 89.21±0.16 86.27±0.69 92.80±0.49 97.15±0.37 86.80±0.03 78.50±4.13 82.69±0.99 1.89 RepeatMixer 90.20±1.04 83.02±1.20 87.23±0.23 87.38±0.18 92.19±0.58 75.65±1.99 66.35±5.64 58.57±2.45 55.32±0.02 4.33 TIDFormer 91.21±0.78 83.27±1.01 89.57±1.17 81.59±1.17 95.37±1.09 66.35±9.00 69.65±0.30 60.06±1.17 60.64±0.90 3.89 DyGFormer 82.23±2.54 81.57±0.67 82.17±0.82 75.63±0.73 85.85±0.66 97.00±0.31 85.30±3.88 64.41±1.40 60.84±1.58 4.89 CNEN 78.19±2.28 82.76±0.83 82.10±0.95 77.66±0.15 87.68±0.13 76.74±1.11 77.07±7.77 74.78±0.21 67.55±0.86 4.78 TGAT 87.38±0.22 79.55±0.20 68.27±1.37 64.07±1.05 82.19±0.62 67.13±0.84 62.14±6.60 55.74±0.91 52.96±2.14 8.56 89.05±0.39 77.14±0.16 80.25±2.74 70.66±0.39 77.06±0.41 65.93±3.00 80.53±3.95 55.90±1.17 52.30±2.35 8.11 TCL TGN 86.86±0.33 81.22±0.61 80.43±2.12 73.91±1.76 87.06±1.93 68.42±3.07 74.00±7.57 58.44±5.51 69.37±3.93 6.11 71.21±1.67 80.82±0.45 65.30±0.43 64.73±0.36 74.05±0.95 66.53±2.77 68.82±8.23 55.71±0.38 51.26±0.04 9.56 CAWN EdgeBank 73.35±0.00 73.59±0.00 65.50±0.00 76.53±0.00 60.71±0.00 63.84±0.00 63.22±0.00 81.32±0.00 84.89±0.00 7.89 GraphMixer 90.90±0.10 78.44±0.18 84.11±1.35 77.98±0.92 77.77±0.92 74.34±0.87 81.65±1.02 57.05±1.22 51.20±1.60 6.00

Inductive Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 88.89±2.13 91.87±0.02 86.07±0.01 84.15±0.03 85.27±1.75 96.89±0.04 86.64±0.65 76.47±0.57 81.93±0.08 1.11 RepeatMixer 88.86±0.97 91.11±0.73 84.20±0.34 83.17±0.50 83.11±1.28 76.16±2.12 69.01±0.60 58.57±2.45 55.32±0.02 4.56 TIDFormer 91.15±1.17 91.57±0.50 85.81±1.39 79.67±0.66 84.53±1.11 64.00±0.77 69.87±1.72 60.32±2.13 59.87±0.19 4.44 DyGFormer 78.29±5.38 91.11±0.40 72.25±1.71 77.41±0.89 81.24±0.69 95.44±0.57 81.25±3.62 55.79±1.02 51.91±0.84 5.67 CNEN 74.99±3.07 90.76±0.29 69.19±1.61 74.88±0.23 80.42±0.23 75.87±1.23 72.27±8.11 72.28±0.46 66.09±0.68 5.78 87.00±0.16 89.59±0.24 68.67±0.84 63.94±1.36 75.95±0.64 68.82±1.21 61.91±5.82 60.61±1.24 52.89±1.61 7.78 TGAT TCL 86.76±0.72 87.45±0.29 76.01±1.11 71.29±0.32 74.65±0.54 65.85±1.75 78.15±3.34 61.06±1.74 50.62±0.82 7.22 TGN 85.62±0.44 88.10±0.24 70.94±0.71 70.89±2.72 77.50±2.91 65.34±2.87 67.57±6.47 61.04±6.01 67.63±2.67 7.00 CAWN 74.06±2.62 91.67±0.24 64.61±0.48 75.15±0.58 73.51±0.94 67.75±1.00 65.81±8.52 62.54±0.67 52.19±0.34 7.33 EdgeBank 80.63±0.00 85.48±0.00 57.43±0.00 73.89±0.00 49.43±0.00 62.16±0.00 64.74±0.00 72.97±0.00 66.30±0.00 8.22 GraphMixer 88.59±0.17 85.26±0.11 80.10±0.51 75.01±0.79 74.27±0.92 69.48±0.63 79.63±0.84 60.15±1.29 51.60±0.73 6.78

B.2

Comparison result with different metric in both transductive and inductive settings

In the main paper, Tab. 6 reports the AP results in the transductive setting under random negative sampling strategies. Here, Tabs. 10, 11, 12, and 13 provide the results under all 3 negative sampling strategies for transductive AP, AUC-ROC, inductive AP, and inductive AUC-ROC, respectively. The results show that DiffDyG consistently achieves the best overall ranking across different evaluation settings. For transductive AUC-ROC, DiffDyG obtains the best AR under random, historical, and inductive negative sampling, with AR values of 1.56, 2.11, and 1.67, respectively. This indicates that the advantage of DiffDyG is not limited to AP, but also holds under a threshold-independent ranking metric. In the inductive setting, DiffDyG remains the strongest overall method. For AP, it achieves the best AR under all three negative sampling strategies, with AR values of 1.00, 1.56, and 1.56 under random, historical, and inductive negative sampling, respectively. For AUC-ROC, DiffDyG again obtains the best AR under all three strategies, with AR values of 1.33, 1.78, and 2.00. These results are notable because the inductive setting is more challenging, as the model must generalize to unseen nodes rather than only predicting future links among previously observed nodes. Overall, the consistent top rankings across AP and AUC-ROC, across transductive and inductive settings, and across three negative sampling strategies demonstrate that the gains of DiffDyG are robust. This supports the central claim that sharpening temporal attention helps the model identify informative historical interactions under different evaluation protocols, rather than overfitting to a particular choice of negative samples. 15

Table 11: Performance (AUC-ROC) comparison in the transductive setting. Average Rank (AR) is computed across the 9 datasets. The best and second best results are shown in bold and underlined. Random Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 99.37±0.38 99.07±1.24 98.41±0.05 98.23±0.69 97.10±0.27 99.59±0.49 81.16±0.14 98.97±0.34 87.49±1.29 1.56 RepeatMixer 99.04±0.01 99.15±0.01 95.36±0.49 93.47±0.17 93.60±0.39 86.15±0.31 75.07±0.21 69.92±0.11 56.58±0.12 4.22 TIDFormer 99.19±0.09 98.93±0.16 96.64±0.14 92.72±0.15 92.79±0.58 99.87±0.05 71.81±2.91 61.25±2.80 67.43±1.42 4.78 DyGFormer 98.64±0.07 98.63±0.01 94.01±0.67 91.11±0.85 86.29±0.24 97.76±0.41 77.90±0.58 70.20±1.44 57.12±0.62 4.89 99.01±0.05 99.15±0.01 96.03±0.08 93.12±0.09 95.69±0.07 89.71±0.07 84.14±0.63 78.57±0.07 69.52±0.34 2.56 CNEN TGAT 96.67±0.07 98.47±0.02 78.53±0.74 68.89±1.10 87.11±0.19 75.69±0.78 75.84±1.99 64.01±0.12 52.83±1.12 8.89 TCL 95.84±0.18 97.42±0.02 87.82±1.36 75.74±0.72 83.12±0.18 72.46±3.23 76.27±0.63 64.72±0.05 51.88±0.36 9.33 TGN 98.37±0.07 98.60±0.06 92.03±1.13 88.32±0.99 91.21±1.15 76.99±1.80 83.34±0.43 69.10±1.67 69.71±2.65 5.44 CAWN 98.54±0.04 99.01±0.01 93.87±0.08 90.45±0.14 80.38±0.26 75.70±3.27 77.16±0.39 68.54±0.18 53.09±0.22 6.56 90.78±0.00 95.37±0.00 77.30±0.00 87.05±0.00 60.86±0.00 64.14±0.00 62.57±0.00 66.75±0.00 62.97±0.00 9.56 EdgeBank GraphMixer 96.92±0.03 97.17±0.02 91.81±0.67 84.38±0.21 84.01±0.17 83.17±0.53 76.96±0.79 65.52±0.51 52.46±0.27 8.11

Historical Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 86.34±0.02 82.16±0.07 84.19±0.63 82.67±0.05 92.02±0.46 97.83±0.06 80.21±0.17 79.10±1.78 81.58±1.21 2.11 RepeatMixer 85.32±0.70 81.95±0.56 78.85±0.43 84.33±0.13 92.09±0.32 81.83±1.66 71.89±9.62 63.69±3.02 56.64±0.04 4.22 TIDFormer 83.14±0.23 81.99±0.21 84.19±0.77 78.78±2.94 88.55±0.31 68.96±1.28 68.23±0.58 59.90±2.30 61.16±1.66 5.78 DyGFormer 78.80±1.95 80.54±0.29 76.97±0.24 76.55±0.52 87.04±0.35 97.61±0.40 90.77±1.96 73.86±1.13 64.27±1.78 5.11 CNEN 76.90±1.34 81.46±0.50 78.98±0.36 78.05±0.08 88.25±0.09 79.20±0.60 83.93±3.99 76.49±0.25 70.45±0.85 4.56 82.87±0.22 79.33±0.16 58.89±1.57 61.85±1.43 80.81±0.67 70.86±0.94 73.47±5.25 60.37±0.68 53.95±3.15 8.44 TGAT TCL 85.76±0.46 76.49±0.16 72.25±3.46 67.95±0.88 72.09±0.56 69.95±3.70 83.97±3.71 61.43±1.04 52.29±2.39 7.89 TGN 82.74±0.32 81.11±0.19 77.25±2.68 77.09±2.22 88.00±1.80 73.23±3.08 83.53±4.53 63.93±5.41 73.40±5.20 5.33 67.84±0.64 80.27±0.30 57.86±0.15 65.10±0.34 71.57±1.07 72.06±3.94 78.62±7.46 63.09±0.74 51.27±0.33 9.11 CAWN EdgeBank 77.27±0.00 78.58±0.00 69.56±0.00 79.59±0.00 61.90±0.00 63.04±0.00 67.41±0.00 86.61±0.00 89.62±0.00 7.22 GraphMixer 87.68±0.17 77.80±0.12 77.54±2.02 75.27±1.14 76.68±1.40 79.03±1.01 85.17±0.70 63.20±1.54 52.61±1.44 6.11

Inductive Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 84.53±0.07 88.47±3.15 83.13±0.28 80.82±0.03 85.31±0.22 96.73±0.14 79.97±3.02 77.80±0.51 80.62±0.14 1.67 RepeatMixer 83.67±0.89 86.45±1.03 76.56±0.30 80.05±0.53 81.94±0.95 82.46±1.62 76.01±0.61 63.69±3.02 56.64±0.04 4.78 TIDFormer 86.13±1.14 88.87±0.65 82.95±0.38 78.30±2.08 83.43±0.69 64.76±1.38 68.89±2.08 60.22±0.98 59.54±0.94 5.00 DyGFormer 75.09±3.70 86.23±0.51 65.96±1.18 74.07±0.64 80.76±0.76 96.70±0.59 87.96±1.80 62.56±1.51 53.37±1.26 5.78 CNEN 70.22±1.47 85.66±0.34 65.63±0.77 74.34±0.18 79.79±0.26 78.29±1.03 80.13±4.77 73.77±0.26 68.38±0.68 5.78 81.93±0.22 87.13±0.20 60.80±1.01 60.45±2.12 73.18±0.33 72.47±1.18 71.62±5.42 66.13±0.78 53.04±2.58 7.56 TGAT 82.19±0.48 84.67±0.29 70.05±1.86 67.64±0.86 70.36±0.37 69.47±2.12 82.54±3.91 67.80±1.21 52.02±1.64 7.22 TCL TGN 80.97±0.31 84.56±0.24 64.11±1.04 71.34±2.46 77.44±2.86 69.57±2.81 78.12±4.46 66.37±5.39 72.69±3.72 7.22 CAWN 70.95±0.95 88.04±0.29 58.06±0.26 75.17±0.50 70.32±1.43 72.93±1.78 76.45±7.02 71.73±0.74 52.75±0.90 6.89 EdgeBank 81.73±0.00 85.93±0.00 58.03±0.00 75.00±0.00 48.18±0.00 61.41±0.00 68.66±0.00 74.20±0.00 72.85±0.00 7.44 GraphMixer 84.28±0.30 82.21±0.13 74.59±0.74 71.53±0.85 72.45±0.72 70.52±0.94 84.22±0.91 66.53±1.22 51.89±0.74 6.67

(a) Wikipedia

(b) Reddit

Figure 5: Comparison of models’ AP, parameter size (MB), and training time (seconds per epoch).

B.3

Efficiency and model size

Although differential attention introduces an additional softmax computation compared with standard attention, DiffDyG remains efficient in practice. Its stronger temporal denoising allows the use of a smaller embedding dimension—36 in DiffDyG versus 50 in DyGFormer —and consequently fewer trainable parameters: 0.58M for DiffDyG, compared with 0.85M for TCL, 1.05M for TIDFormer, 16

Table 12: Performance (Average Precision) comparison in the inductive setting. Average Rank (AR) is computed across the 9 datasets. The best and second best results are shown in bold and underlined. Random Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 98.97±0.25 99.04±0.01 98.06±0.37 94.48±1.95 95.16±0.70 98.12±0.34 73.78±0.39 97.82±1.00 74.46±1.82 1.00 RepeatMixer 98.70±0.05 98.85±0.01 95.04±0.12 87.97±0.29 93.05±0.12 59.77±0.70 51.70±0.97 61.36±0.21 54.47±0.08 4.78 TIDFormer 98.93±0.02 98.98±0.03 95.95±0.12 89.92±0.45 93.29±0.39 93.97±2.93 60.83±1.25 60.31±1.87 66.02±1.34 2.89 DyGFormer 98.36±0.03 98.35±0.02 94.14±0.43 88.63±0.18 85.27±0.09 87.74±0.71 54.28±2.87 64.55±0.62 55.93±0.39 4.89 98.37±0.03 98.78±0.01 95.03±0.16 89.66±0.22 91.89±0.31 68.31±0.59 59.44±0.44 66.58±0.27 69.71±0.48 3.33 CNEN 96.22±0.07 97.09±0.04 79.54±0.48 67.05±1.51 85.50±0.19 55.18±0.79 51.00±3.11 61.03±0.18 52.24±1.46 8.33 TGAT TCL 96.22±0.17 94.09±0.07 87.36±2.03 76.14±0.79 80.60±0.22 54.30±0.66 52.59±0.97 62.21±0.12 51.60±0.97 8.44 TGN 97.83±0.04 97.50±0.07 88.12±2.05 77.94±1.02 89.04±1.17 54.10±0.93 58.63±0.37 58.31±3.15 58.85±2.51 6.89 CAWN 98.24±0.03 98.62±0.01 92.73±0.06 86.35±0.51 81.42±0.24 55.80±0.69 53.17±1.20 65.24±0.21 49.94±0.45 6.33 GraphMixer 96.65±0.02 95.26±0.02 91.19±0.42 75.88±0.48 81.41±0.21 55.91±0.82 50.71±0.76 62.17±0.31 50.68±0.44 8.11

Historical Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 85.67±1.12 66.08±0.05 86.27±2.14 74.32±0.13 84.37±0.06 88.84±0.31 69.88±0.81 67.24±4.14 67.35±2.95 1.56 RepeatMixer 84.11±1.88 66.14±1.40 85.52±0.16 82.86±0.47 83.19±1.72 56.14±1.34 53.90±2.90 57.48±1.95 54.53±0.13 4.00 TIDFormer 84.99±1.89 65.62±0.92 86.11±1.81 76.99±1.21 84.61±1.17 63.84±0.56 69.56±2.11 59.13±1.41 60.15±0.26 2.67 DyGFormer 71.42±4.43 65.37±0.60 72.13±1.87 67.07±0.62 80.82±0.30 87.40±0.85 56.31±3.46 53.20±1.07 52.63±1.26 5.89 70.70±2.75 64.09±0.92 70.23±1.75 70.90±0.39 80.59±0.10 66.14±0.77 63.48±1.31 60.48±0.15 64.39±0.28 5.00 CNEN TGAT 84.17±0.22 63.47±0.36 70.52±0.93 61.40±1.31 76.73±0.29 56.72±0.47 51.83±3.95 55.28±0.71 53.05±3.10 7.33 82.20±2.18 60.83±0.25 76.71±1.00 67.11±0.62 74.27±0.53 55.71±0.74 53.87±1.41 55.76±1.03 54.19±2.17 7.11 TCL TGN 81.76±0.32 64.85±0.85 70.78±0.78 62.91±1.16 77.07±3.41 54.42±0.77 61.18±1.10 52.80±3.19 63.74±3.00 6.67 67.27±1.63 63.67±0.41 64.54±0.47 60.70±0.36 74.68±0.68 57.14±0.07 55.56±1.71 55.00±0.38 47.98±0.84 8.22 CAWN GraphMixer 87.60±0.30 64.50±0.26 81.66±0.49 72.37±1.37 74.00±0.97 55.84±0.73 52.03±1.02 54.94±0.97 48.09±0.43 6.56

Inductive Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 85.67±1.12 66.07±0.05 86.27±2.14 74.33±0.13 84.36±0.06 88.52±0.31 69.88±0.81 61.25±0.93 67.36±2.95 1.56 RepeatMixer 84.10±1.88 66.13±1.40 85.53±0.16 82.87±0.47 83.19±1.72 56.92±1.36 52.18±1.53 57.48±1.95 54.53±0.13 4.00 TIDFormer 84.99±1.89 65.62±0.92 86.11±1.81 76.98±1.21 84.61±1.17 63.27±0.32 69.56±2.11 59.13±1.41 59.09±0.19 2.67 DyGFormer 71.42±4.43 65.35±0.60 72.13±1.86 67.07±0.62 80.82±0.30 87.22±0.82 56.31±3.46 52.56±1.70 52.61±1.25 6.00 70.73±2.76 64.08±0.91 70.24±1.74 71.03±0.47 80.58±0.11 65.83±0.62 63.94±0.97 60.53±0.17 64.72±0.35 5.00 CNEN TGAT 84.17±0.22 63.40±0.36 70.49±0.93 61.40±1.30 76.72±0.30 56.46±0.50 51.83±3.95 55.58±0.68 53.08±3.10 7.44 TCL 82.20±2.18 60.81±0.26 76.65±0.99 67.11±0.62 74.28±0.53 55.46±0.69 53.87±1.41 55.66±0.98 54.13±2.16 7.00 81.77±0.32 64.84±0.84 70.73±0.79 62.90±1.16 77.07±3.40 54.18±0.73 61.18±1.10 52.80±3.24 63.71±2.97 6.56 TGN CAWN 67.24±1.63 63.65±0.41 64.54±0.47 60.72±0.36 74.69±0.68 57.06±0.08 55.56±1.71 54.97±0.38 48.01±0.82 8.22 GraphMixer 87.60±0.29 64.49±0.25 81.64±0.49 72.37±1.38 73.99±0.97 55.76±0.65 52.03±1.02 54.88±1.01 48.10±0.40 6.56

and 1.15M for DyGFormer. The gains in §6.1 are therefore not attributable to additional model capacity. Combined with FlashAttention [5] and Grouped Query Attention [1], DiffDyG trains 1.5× to 2.5× faster than DyGFormer and TIDFormer (Fig. 5), making it practical for large-scale dynamic graph learning while improving on both effectiveness and parameter efficiency. B.4

Hyperparameter sensitivity

We evaluate the sensitivity of DiffDyG to five hyperparameters: the number of Transformer layers, attention heads, sampling hops, and the number of neighbors for both first and second hops. As shown in Tabs. 15 and 16, DiffDyG demonstrates strong stability across different configurations. Performance remains largely consistent as the number of layers and attention heads increases, suggesting that the model is robust to variations in architectural depth and breadth. When it comes to the neighbor sampling strategy, increasing the interaction range from 1-hop to 2-hop provides a noticeable performance improvement, as the 2-hop neighborhood introduces more topological context for representation learning. However, it is important to note that the performance drop from reducing to 1-hop is relatively minor compared to the impact of removing key architectural components, such as Differential Attention, RoPE, or Spatial Encoding, which were analyzed in the ablation studies in §6.2. B.5

Additional ablation results

Tab. 17 provides the ablation studies in transductive setting with AUC-ROC and in inductive settings. 17

Table 13: Performance (AUC-ROC) comparison in the inductive setting. Average Rank (AR) is computed across the 9 datasets. The best and second best results are shown in bold and underlined. Random Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 98.97±0.13 98.79±0.62 97.32±0.54 92.59±0.47 95.40±0.90 98.05±0.64 59.38±1.64 98.81±0.03 70.99±1.16 1.33 RepeatMixer 98.63±0.02 98.69±0.02 93.54±0.14 89.04±0.26 94.10±0.07 60.03±1.19 49.29±1.90 66.17±0.17 54.32±0.25 4.00 TIDFormer 98.76±0.09 98.23±0.28 93.35±0.13 85.62±0.31 92.55±0.58 94.52±2.69 59.55±2.87 61.14±2.84 65.13±1.00 4.11 DyGFormer 98.16±0.32 98.05±0.03 91.97±1.68 88.56±1.07 85.00±0.10 89.33±0.48 53.21±3.04 67.25±1.05 56.73±0.69 4.78 CNEN 98.23±0.01 98.62±0.01 93.34±0.17 90.18±0.15 92.76±0.29 70.22±0.77 61.52±0.52 67.26±0.10 69.17±0.53 2.89 TGAT 95.90±0.09 96.98±0.04 77.64±0.38 64.63±1.74 86.84±0.17 56.51±0.75 48.27±3.50 62.72±0.12 51.83±1.35 8.33 95.57±0.20 93.80±0.07 84.49±1.82 72.33±0.99 81.43±0.19 55.83±1.07 50.43±1.48 63.76±0.07 50.51±1.05 8.78 TCL TGN 97.72±0.03 97.39±0.07 86.68±2.29 78.83±1.11 91.24±0.99 55.86±0.75 62.38±0.48 59.99±3.50 61.23±2.71 6.44 CAWN 98.03±0.04 98.42±0.02 90.40±0.11 87.02±0.50 81.86±0.25 58.83±1.13 51.49±1.13 67.05±0.21 48.34±0.76 6.22 GraphMixer 96.30±0.04 94.97±0.05 89.30±0.57 76.51±0.71 82.77±0.24 58.32±1.08 47.20±0.89 63.48±0.37 50.04±0.86 8.11

Historical Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 80.19±0.93 65.29±1.58 78.02±0.34 74.14±0.01 82.79±1.51 89.18±1.78 55.16±0.37 67.87±2.31 64.00±3.15 1.78 RepeatMixer 77.58±1.00 65.15±0.36 76.91±0.23 78.84±0.29 82.68±1.28 54.27±2.50 53.26±3.85 61.67±2.81 54.62±0.18 4.22 TIDFormer 73.81±2.16 64.87±1.13 77.35±0.57 72.86±2.45 79.63±1.29 64.82±0.74 67.90±0.91 59.03±2.74 59.05±0.33 4.22 DyGFormer 68.33±2.82 64.81±0.25 65.55±1.01 65.78±0.42 80.77±0.63 88.68±0.74 56.57±3.22 58.46±1.65 53.85±2.02 5.89 CNEN 69.88±1.81 63.18±0.42 67.78±0.66 70.57±0.25 80.83±0.14 67.25±0.99 64.26±1.02 60.95±0.23 63.86±0.68 4.78 TGAT 78.38±0.20 64.43±0.27 62.32±1.18 57.84±2.18 74.08±0.27 58.30±0.61 49.99±4.88 59.74±0.59 51.73±4.12 7.44 79.79±0.96 61.43±0.26 70.46±1.94 64.06±1.02 69.82±0.32 57.30±1.03 52.12±2.13 61.12±0.97 54.66±2.11 6.44 TCL TGN 75.75±0.29 64.55±0.50 62.69±0.90 62.68±1.09 77.69±3.55 55.64±0.54 64.87±1.65 55.61±3.54 68.59±3.11 6.22 CAWN 62.04±0.65 64.94±0.21 56.39±0.10 62.25±0.40 71.68±0.94 60.11±0.48 54.41±1.31 60.95±0.80 48.01±1.77 7.22 GraphMixer 82.87±0.21 64.27±0.13 75.98±0.84 68.20±1.62 72.53±0.84 56.68±1.20 49.28±0.86 59.88±1.17 45.49±0.42 6.67

Inductive Negative Sampling Baseline

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

AR

DiffDyG 80.19±0.93 65.29±1.58 78.02±0.34 74.14±0.01 82.79±1.54 88.76±1.78 55.16±0.37 64.07±4.17 64.02±3.15 2.00 RepeatMixer 77.58±1.00 65.15±0.36 76.92±0.23 78.85±0.29 82.68±1.28 55.68±2.54 51.96±1.70 61.67±2.81 54.62±0.18 4.33 TIDFormer 80.38±0.10 64.87±1.15 77.35±0.57 72.86±2.45 79.63±1.29 64.24±0.50 67.90±0.16 59.03±2.74 59.05±0.33 3.67 DyGFormer 68.33±2.82 64.80±0.25 65.58±1.00 65.79±0.42 80.77±0.63 88.51±0.73 56.57±3.22 57.28±3.06 53.87±2.01 5.89 CNEN 69.89±1.86 63.17±0.48 67.80±0.64 70.63±0.27 80.82±0.15 66.91±1.13 64.85±0.81 61.01±0.16 64.36±0.45 4.56 TGAT 78.38±0.20 64.39±0.27 62.29±1.17 57.83±2.18 74.07±0.27 58.15±0.62 49.99±4.88 59.98±0.59 51.78±4.14 7.44 TCL 79.79±0.96 61.36±0.26 70.42±1.93 64.05±1.02 69.83±0.32 56.88±0.93 52.12±2.13 61.01±0.93 54.65±2.20 6.44 TGN 75.76±0.29 64.55±0.50 62.66±0.91 62.68±1.09 77.68±3.55 55.43±0.42 64.87±1.65 55.62±3.59 68.58±3.08 6.44 62.02±0.65 64.91±0.21 56.39±0.11 62.27±0.40 71.69±0.94 60.01±0.47 54.41±1.31 60.88±0.79 48.04±1.76 7.33 CAWN GraphMixer 82.88±0.21 64.27±0.13 75.97±0.85 68.19±1.63 72.52±0.84 56.63±1.09 49.28±0.86 59.71±1.17 45.57±0.41 6.78

Table 14: Performance comparison on dynamic node classification. Results are reported in AUC-ROC. The best and second best results are shown in bold and underlined.

B.6

Model

Wikipedia

Reddit

AR

DiffDyG DyGFormer TIDFormer TCL CNEN RepeatMixer TGAT TGN CAWN GraphMixer

88.12±0.73 87.44±1.08 87.53±1.12 77.83±2.13 88.37±0.28 80.94±0.19 84.09±1.27 86.38±2.34 84.88±1.33 86.80±0.79

70.94±0.65 68.00±1.74 69.59±1.70 68.87±2.15 66.84±1.19 68.12±0.64 70.04±1.09 63.27±0.90 66.34±1.78 64.22±3.32

1.5 5.0 3.0 7.0 4.0 7.0 5.0 8.0 7.5 7.0

Efficiency and model size

Fig. 5 compares training efficiency and parameter size on Wikipedia and Reddit datasets, where bubble size represents the number of parameters. DiffDyG achieves the best accuracy–efficiency trade-off among all evaluated methods. By leveraging FlashAttention [5] and Grouped Query Attention [1], DiffDyG trains 1.5–2.5× faster than other Transformer-based models such as DyGFormer and TIDFormer, while maintaining a comparable or smaller model size. This demonstrates that differential attention is not only more effective but also practically efficient for large-scale dynamic graph learning.

C

The pseudo-code of D IFF DY G

The algorithm of the proposed DiffDyG is summarized in Algo. 1. 18

Table 15: Hyperparameter analysis in transductive setting Hyperparameter & Value AP

Wikipedia AUC-ROC

UCI AP

AUC-ROC

AP

Can. Parl. AUC-ROC

# layers

2 4 6

99.57±0.03 99.61±0.07 99.67±0.12

99.37±0.38 99.42±0.05 99.53±0.28

98.54±0.03 98.46±0.52 98.82±0.05

98.41±0.05 97.56±0.47 98.52±0.07

99.63±0.09 99.74±0.06 99.85±0.02

99.59±0.49 99.63±0.08 99.72±0.06

# heads

2 4 8

99.57±0.03 99.58±0.14 99.60±0.01

99.37±0.38 99.42±0.07 99.45±0.03

98.54±0.03 98.58±0.02 98.63±0.28

98.41±0.05 98.44±0.11 98.47±0.02

99.63±0.09 99.69±0.01 99.76±0.02

99.59±0.49 99.63±0.03 99.66±0.05

# hops

1 2

99.03±0.01 99.57±0.03

98.79±0.62 99.37±0.38

97.65±0.54 98.54±0.03

97.06±0.37 98.41±0.05

98.97±0.05 99.63±0.09

99.02±0.23 99.59±0.49

# 1-hop neighbors

10 20 30

99.06±0.14 99.57±0.03 99.23±0.06

98.89±0.02 99.37±0.38 99.01±0.35

98.49±0.02 98.54±0.03 98.57±0.03

98.13±0.07 98.41±0.05 98.36±0.02

99.34±0.21 99.63±0.09 99.72±0.04

99.12±0.04 99.59±0.49 99.63±0.05

# 2-hop neighbors

5 10 15

99.57±0.03 99.65±0.06 99.61±0.01

99.37±0.38 99.46±0.07 99.42±0.05

98.54±0.03 98.76±0.01 98.89±0.61

98.41±0.05 98.95±0.03 99.01±0.14

99.63±0.09 99.71±0.03 99.65±0.04

99.59±0.49 99.65±0.02 99.63±0.06

Table 16: Hyperparameter analysis in inductive setting Hyperparameter & Value AP

Wikipedia AUC-ROC

UCI AP

AUC-ROC

AP

Can. Parl. AUC-ROC

# layers

2 4 6

98.97±0.25 98.84±0.13 98.87±0.17

98.97±0.13 98.74±0.02 98.83±0.29

98.06±0.37 97.43±0.81 98.17±0.19

97.32±0.54 96.83±0.59 97.38±0.35

98.12±0.34 98.67±0.24 98.81±0.04

98.05±0.64 98.73±0.01 98.96±0.27

# heads

2 4 8

98.97±0.25 99.01±0.16 99.03±0.06

98.97±0.13 98.99±0.24 99.02±0.01

98.06±0.37 98.13±0.45 98.39±0.05

97.32±0.54 97.46±0.19 97.59±0.86

98.12±0.34 98.36±0.74 98.54±0.17

98.05±0.64 98.23±0.59 98.33±0.32

# hops

1 2

98.45±0.26 98.97±0.25

98.24±0.05 98.97±0.13

96.95±0.70 98.06±0.37

96.37±0.47 97.32±0.54

97.75±0.46 98.12±0.34

97.24±0.54 98.05±0.64

# 1-hop neighbors

10 20 30

98.67±0.25 98.97±0.25 98.71±0.46

98.46±0.13 98.97±0.13 98.52±0.14

97.52±0.15 98.06±0.37 98.52±0.16

96.79±0.31 97.32±0.54 97.06±0.27

98.04±0.48 98.12±0.34 98.65±0.07

97.79±0.12 98.05±0.64 98.34±0.32

# 2-hop neighbors

5 10 15

98.97±0.25 98.99±0.05 98.93±0.52

98.97±0.13 99.01±0.24 98.95±0.18

98.06±0.37 97.17±0.32 97.19±0.38

97.32±0.54 95.99±0.22 96.02±0.37

98.12±0.34 98.75±0.39 98.64±0.14

98.05±0.64 98.42±0.09 98.27±0.59

Algorithm 1 Training pipeline for D IFF DY G 1: Input: Training set G = (u, v, t, y), number of hop K, model with encoder g and classifier f 2: Initialize model parameters 3: for epoch = 1, 2, 3, . . . do 4: for (u, v, t, y) ∈ G do 5: for k = 1, . . . , K do 6: Obtain the k-hop neighbor sets of u and v as Hk (u, t), Hk (v, t) 7: Obtain five feature channels Xtk∗,N , Xtk∗,E , Xtk∗,T , Xtk∗,C , Xtk∗,S with each row corresponding to a node in the k-hop neighbor sets. ∗ can be u or v. 8: Xtk∗ ← [Xtk∗,N , Xtk∗,E , Xtk∗,T , Xtk∗,C , Xtk∗,S ] 9: end for 10: Obtain multi-hop encoding representations: Zt∗ ← [Xt1∗ , · · · , XtK∗ ] 11: Obtain final representations Y∗t = g(Zt∗ ) t , Y t ]) 12: Compute the link probability p ← f ([Yu v 13: Compute the loss LBCE (p, y) 14: Update the model 15: end for 16: end for

19

Table 17: Ablation study on 9 datasets in transductive & inductive settings. Transductive AUC-ROC

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

DiffDyG w/o RoPE w/o DA w/o SE

99.37±0.38 99.03±1.62 98.82±0.07 99.20±0.30

99.07±1.24 99.03±0.85 98.96±0.34 99.05±0.74

98.41±0.05 96.96±0.01 96.22±0.31 97.98±0.01

98.23±0.69 97.19±4.62 93.83±0.21 98.19±0.61

97.10±0.27 95.07±0.56 88.74±0.81 95.26±0.13

99.59±0.49 97.69±0.26 97.78±0.53 98.64±0.36

81.16±0.14 78.03±0.09 77.98±0.34 78.84±0.68

98.97±0.34 95.41±0.02 71.50±0.23 96.05±0.42

87.49±1.29 83.89±2.14 59.17±0.25 85.78±0.03

AP

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

DiffDyG w/o RoPE w/o DA w/o SE

98.97±0.25 98.75±0.26 98.45±0.33 98.81±0.04

99.04±0.01 98.73±0.31 98.59±0.47 98.89±0.56

98.06±0.37 96.01±0.03 94.88±0.51 97.73±1.53

94.48±1.95 94.25±2.79 90.74±0.34 93.55±2.75

95.16±0.70 94.66±0.75 87.97±0.95 95.10±0.25

98.12±0.34 96.98±0.01 92.15±0.24 97.57±0.01

73.78±0.39 70.99±3.51 63.09±0.13 70.36±1.57

97.82±1.00 92.11±0.10 65.25±0.25 93.65±0.13

74.46±1.82 66.60±2.55 56.76±0.54 68.21±1.00

AUC-ROC

Wikipedia

Reddit

UCI

Enron

MOOC

Can. Parl.

US Legis.

UN Trade

UN Vote

DiffDyG w/o RoPE w/o DA w/o SE

98.97±0.13 98.51±0.03 98.48±0.19 98.73±0.65

98.79±0.62 98.47±2.19 98.32±0.61 98.65±1.34

97.32±0.54 95.19±0.01 93.32±0.26 95.60±1.38

92.59±0.47 90.99±0.33 90.47±0.57 91.49±3.70

95.40±0.90 94.27±1.31 87.58±0.86 94.39±0.93

98.05±0.64 96.98±0.01 93.05±0.67 97.37±0.01

59.38±1.64 57.05±0.60 55.26±0.94 56.12±2.53

98.81±0.03 92.08±1.03 68.36±0.19 93.31±0.52

70.99±1.16 66.44±1.99 58.72±0.49 68.69±0.05

Inductive

20

Record · ID 192399 · SHA-256 071544d4daca9721
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.