OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation Yuze Dai* , Zhihan Zhang* , Yan Zhao, Ruoyu Wu, Xunkai Li, Zekai Chen, Qiangqiang Dai, Hongchao Qin, Ronghua Li Beijing Institute of Technology {1120240577, 3220241443, 1120243657}@bit.edu.cn, [email protected] [email protected], [email protected], [email protected] [email protected], [email protected]
arXiv:2607.19108v1 [cs.AI] 21 Jul 2026
Abstract Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often imperfect, with quality issues arising from text, structure, and labels, and typically manifesting as sparsity, noise, and imbalance. These dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Although prior studies have explored specific mitigation strategies, existing evidence remains fragmented across degradation types, datasets, tasks, and model families, leaving TAG robustness insufficiently understood. To address this gap, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified 3 × 3 taxonomy and supports standardized evaluation across nine TAG datasets and three downstream tasks. It systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, investigates the effectiveness, efficiency, and robustness of scenario-matched baselines, and further examines model behavior under composite degradation scenarios. OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings.
Introduction Text-attributed graphs (TAGs) have become an important data form for graph learning because they couple relational structure with rich node-associated text (Yan et al. 2023; Zhang et al. 2024a), supporting applications such as academic networks, social platforms, e-commerce systems, and knowledge-intensive information networks. However, realworld TAGs are often imperfect: their quality issues arise from three coupled modalities, namely text, structure, and labels, and typically manifest as sparsity, noise, and imbalance. Together, these dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Specifically, text sparsity or corruption weakens semantic features, structural sparsity or noise distorts neighborhood context and graph-text alignment, and label sparsity, noise, or imbalance degrades supervision signals. Such challenges affect not only traditional GNNs (Kipf and Welling 2017; Veličković et al. 2018; Hamilton, Ying, and * These authors contributed equally.
Leskovec 2017) and LLM-GNNs (He et al. 2023; Zhu et al. 2024b; Zhang et al. 2025b; Ren et al. 2024), but also recent graph foundation models (GFMs) (Chen et al. 2024; Liu et al. 2023), whose robustness under incomplete, noisy, and skewed TAG settings remains insufficiently understood. Motivated by this gap, we study TAG robustness through a benchmark that systematically covers all nine degradation scenarios. A number of methods have been proposed to mitigate specific data-quality problems, including text denoising and completion (Sun and Jiang 2019; Zhang et al. 2025b), graph structure learning (Zhu et al. 2021; Zhang et al. 2025a; Li et al. 2022; Liu et al. 2022), label correction, imbalanceaware training (Qin et al. 2024; Wang et al. 2024; Qin et al. 2025), and LLM-assisted graph modeling (Zhang et al. 2024b; Li et al. 2024a). Yet the current evidence remains fragmented. Existing studies are often tied to a specific degradation type, a small set of datasets, a particular backbone, or a single downstream task; as a result, they do not provide a unified experimental basis for understanding robustness in TAG learning. More importantly, they leave several benchmark-level questions insufficiently answered: ❶ Are the constructed degradation scenarios valid, meaningful, non-collapsing, and practical to instantiate? ❷ How sensitive are GNNs, LLM-GNNs, and GFMs to different degradation modalities? ❸ Which text-, structure-, and label-oriented baseline methods are truly effective under matched lowquality scenarios? ❹ And how well do existing methods perform under composite degradation scenarios? To answer these questions, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified taxonomy over three modalities and three degradation types, yielding nine representative scenarios. It systematically evaluates scenario validity and model sensitivity across multiple TAG datasets and three downstream tasks, including node classification, node clustering, and link prediction, while comparing traditional GNNs, LLM-GNNs, and a representative GFM. It also investigates the effectiveness, efficiency, and robustness of scenario-matched baselines for text, structure, and label degradation, and further examines model behavior under composite degradation scenarios. Rather than serving as another clean-data leaderboard,
OpenRTAG provides a standardized testbed for understanding where TAG learning systems fail, which repair strategies help, and how robustness conclusions vary across scenarios, datasets, tasks, and model paradigms. Our Contributions. (1) Unified degradation taxonomy. We define a 3 × 3 benchmark space for TAG quality issues, spanning text, structure, and label modalities under sparsity, noise, and imbalance. (2) Standardized benchmark framework. We build OpenRTAG, a unified framework for scenario construction, data preparation, model and baseline evaluation, and analysis across multiple datasets and three downstream tasks. (3) Systematic robustness study. We evaluate scenario validity and model sensitivity, compare representative GNNs, LLM-GNNs, and GFMs, study the effectiveness, efficiency, and robustness of scenario-matched baselines, and examine model behavior under composite degradation scenarios.
Structural degradation. Let A be the adjacency matrix, di be node degree, and sim(i, j) be a semantic or labelaware similarity proxy. Structure sparsity removes valid edges, structure noise injects spurious edges, and structure imbalance creates uneven structural support: S-Spa :
E′ = E \ E−,
|E − |/|E| ≈ α,
S-Noi :
E′ = E ∪ E+,
sim(i, j) low for (i, j) ∈ E + ,
S-Imb :
d′i or Vari (d′i ) is large. max d′i / min ′ i
i:di >0
(4) Label degradation. Let L ⊆ V be the labeled training nodes, yi∗ be the latent clean label, and nc be the number of training labels in class c. Label sparsity reduces supervision, label noise corrupts observed labels, and label imbalance skews class support: L-Spa :
L′ ⊂ L,
Problem Setting and Degradation Taxonomy
L-Noi :
Pr(yi = c | yi∗ = c′ ) = ηc′ c ,
Text-Attributed Graphs
L-Imb :
max n′c / min n′c is large. ′
We consider a text-attributed graph (TAG) as a graph in which each node is associated with node-level text (Yan et al. 2023). Formally, a TAG is denoted as G = (V, E, T, Y ), where V is the node set, E is the edge set, T = {ti }i∈V denotes node-associated texts, and Y denotes available supervision signals when labels exist. Thus, a TAG contains three coupled information sources: textual semantics T , relational structure (V, E), and supervision Y .
Nine Quality-Degradation Scenarios OpenRTAG focuses on data-quality failures from three information sources in TAG learning: text, structure, and labels. Text degradation affects node semantics, structural degradation affects relational context, and label degradation affects supervision. Each modality is organized by three degradation types: sparsity, noise, and imbalance, corresponding to missing information, corrupted information, and uneven quality or support, respectively. Let the modality and degradation-type sets be: M = {text, structure, label}, D = {sparsity, noise, imbalance}.
(1)
The benchmark scenario space is then defined as: S = M × D, which yields nine representative degradation scenarios. Given a clean TAG G, a scenario s = (m, d) ∈ S is instantiated by a scenario constructor As,α with perturbation strength α: Gs,α = As,α (G). (2) Text degradation. Let ℓi = |ti | denote the length of node text and let ϕ(ti ) denote its semantic content. Text sparsity removes or truncates node text, text noise corrupts tokens or inserts irrelevant content, and text imbalance makes text quality uneven across nodes or classes: t′i = ∅ or ℓ′i ≪ ℓ̄, |Wnoise (t′i )| T-Noi : ≥ α, |W (t′i )| T-Imb : Vari (ℓ′i ) or Vari (ϕ(t′i )) is large. T-Spa :
(3)
c
|L′ |/|L| ≈ 1 − α, (5)
c:nc >0
These definitions are intentionally compact: they specify the target modality and the expected direction of degradation while leaving implementation details, such as exact sampling rules and compatibility constraints, to the benchmark constructor. The taxonomy therefore covers incomplete, corrupted, and uneven node text; missing, noisy, and imbalanced graph topology; and scarce, noisy, and long-tailed supervision under a shared controlled scenario space.
OpenRTAG Benchmark Design OpenRTAG is organized around benchmark assets, scenario construction, evaluation coverage, and evidence reporting. Figure 1 summarizes the resulting framework.
Benchmark Data and Scenario Space OpenRTAG treats each clean TAG as a reusable source with node text, graph topology, labels, and fixed task splits. To characterize the current TAG ecosystem, we survey a broader candidate pool comprising citation and academic graphs (Cora, CiteSeer, PubMed, and Arxiv) (Yang, Cohen, and Salakhutdinov 2016; Hu et al. 2020), a wiki graph (WikiCS) (Mernyei and Cangea 2020), two social graphs (Instagram and Reddit) (Huang et al. 2024; Li et al. 2024c), and e-commerce graphs (Children, Ratings, History, Photo, and Products) (McAuley et al. 2015; Yan et al. 2023). These datasets span small citation networks to a large academic graph with over one hundred thousand nodes and reflect recent graph-text benchmark resources (Yan et al. 2023; Li et al. 2024c; Feng et al. 2024). The standardized main evaluation in this paper uses nine datasets: Cora, CiteSeer, Instagram, WikiCS, PubMed, Children, Photo, History, and Arxiv. The degradation taxonomy and evaluation protocol can subsequently extend to the remaining surveyed TAG collections. The core scenario space is the 3 × 3 taxonomy introduced in Section 2. Each scenario is instantiated by a configurable constructor that takes a clean TAG, a scenario type,
OpenRTAG Benchmark Framework A unified benchmark for degradation scenarios, model robustness, and scenario-matched baselines
C. Q1 Model-Paradigm Robustness
B. Core Scenario Space
A. Benchmark Data Interface
E. Evidence Outputs
Clean vs. degraded TAGs across model families
Clean TAG inputs
Unified 3 x 3 Quality-Degradation Taxonomy
Text
Sparsity
Noise
Imbalance
Text sparsity
Text noise
Text imbalance
Q1 Scenario Validity scenario stats + model drops Clean vs degraded TAGs
Multi-domain TAG datasets academic · social · product · web
Structure
Structure noise
Structure sparsity
Structure imbalance
Traditional GNNs
Text attributes documents / titles / descriptions
Labels
Label noise
Label sparsity
Q2 Text Evaluation text degradation drops + repair gains
Label imbalance
Q3 Structure Evaluation topology perturbation + repair gains
LLM-GNNs
GCN · GAT · GraphSAGE
ENGINE · GIANT· GFM-style Models OpenGraph … GraphAdapter …
Q4 Label Evaluation supervision quality + long-tail effects
Task Evaluation
Graph topology nodes / edges / neighborhoods Scenario Generator Labels supervision / class signals
Clean TAGs
Clean / degraded TAG pairs
Scenario Constructor
3 x 3 Scenario type
Task Splits train / validation / test indices
Node classification Node clustering NMI / ARI Acc / Macro-F1
Degraded TAGs
Q1 Robustness Drop clean score - degraded score absolute / relative drop harmful ratio · variance
report: validity stats / logs
ratio / strength / seed
Link prediction AUC / AP
Q5 Composite & Transfer paired degradation + downstream gains
Q1 robustness evidence Q2-Q5 repair evidence
matched degraded scenarios
Repaired TAGs new text / graph / labels
D. Q2-Q5 Scenario-Matched Baseline Evaluation Structure Repair Track
Text Repair Track
noisy / sparse / imbalanced text Denoising
denoise / generate
Label Repair Track
`
` clean text Generation
missing / noisy / imbalanced topology
edge refine
Graph Structure Learning
repaired graph Graph imbalance
scarce / wrong / long-tailed labels
Label-efficient
propagate / filter / reweight
robust supervision
Noisy-label
Graph imbalance
Representative methods are grouped by repair mechanism; Q5 further tests composite degradation and downstream transfer.
GNN Backbone shared downstream backbone NC / Clu / LP Acc-F1 / NMI-ARI / AUC-AP Repair Gain repaired - degraded
Figure 1: Overview of the OpenRTAG benchmark framework, covering scenario construction, model-paradigm robustness, scenario-matched baselines, and evidence reporting. and perturbation parameters such as ratio, strength, or random seed. The output is a degraded TAG with diagnostic records such as validity statistics and generation logs. Thus, scenario semantics are separated from low-level implementation details, while all generated variants remain comparable through aligned splits and task protocols.
Evaluation Protocol and Coverage As summarized in Table 1, OpenRTAG organizes evaluation around five research questions to assess method robustness and scenario validity under controlled quality degradation. In particular, Q1 focuses on model-paradigm robustness under controlled degradation. By comparing clean and degraded TAG pairs across traditional GNNs, LLM-GNNs, and representative GFMs, Q1 serves two purposes: it evaluates the robustness of different backbone paradigms, and it verifies whether the constructed scenarios induce meaningful, non-trivial, and consistent performance changes across model families. Q2–Q4 evaluate scenario-matched methods under text, structure, and label degradation, respectively. In each case, methods are matched to the degraded modality. These questions are designed to assess three aspects of method behavior: effectiveness, namely whether a method can improve performance under its matched low-quality scenario; robustness, namely whether its performance remains stable as the perturbation level increases; and efficiency, including runtime, preprocessing cost, memory usage when available, and
run stability. This design separates broad backbone robustness from targeted repair and mitigation performance. Q5 further extends the evaluation beyond the main controlled setting. It examines whether the better-performing methods identified in Q2–Q4 remain effective on different downstream tasks, and how they behave when the benchmark is extended from single-scenario degradation to composite degradation settings formed by combining two degradation scenarios. In this way, Q5 tests both the task-level generalizability of promising methods and their robustness under more realistic multi-factor degradation conditions.
Experiments We organize the experiments around five research questions: Q1. Are the degradation scenarios valid and broadly harmful across TAG backbones? Q2. How effective, robust, and efficient are text-oriented baselines under text degradation? Q3. How effective, robust, and efficient are structure-oriented baselines under structure degradation? Q4. How effective, robust, and efficient are label-oriented baselines under label degradation? Q5. Do strong baselines generalize to other downstream tasks and composite degradation scenarios? Together, Q1 validates the scenario space, Q2–Q4 evaluate scenario-matched methods, and Q5 extends the analysis to downstream-task and composite-degradation settings.
Table 1: Overview of model and baseline coverage in OpenRTAG. The table summarizes what is evaluated in each track; representative papers are cited in the corresponding setup paragraphs. Coverage axis
Benchmark target
Representative methods
Track
Evidence produced
Model paradigms
Clean-to-degraded robustness across model families
Traditional GNNs; LLM-GNN / TAG backbones; GFM-style representatives
Q1
Scenario validity, robustness drop, harmful-case ratio
Text repair
Sparse, noisy, or imbalanced node text
Denoising; completion; generation; rule-based or LM-based implementation baselines
Q2
Repair gain under text degradation
Structure repair
Missing, noisy, or imbalanced topology
Graph structure learning; graph purification; graph imbalance learning
Q3
Repair gain under structural degradation
Label repair
Scarce, noisy, or long-tailed supervision
Label-efficient learning; noisy-label learning; imbalance-aware graph learning
Q4
Repair gain under label degradation
Task and composite extension
Generalization to downstream tasks and composite scenarios
Better-performing methods from Q2–Q4
Q5
Task-level effectiveness and performance under composite degradation
Degradation scenario
Cora
WikiCS
Text-Spar
2.4
2.3
4.9
2.6
5.4
3.0
Text-Noise
4.0
2.2
4.6
2.6
4.6
4.1
Text-Imb
5.8
2.8
5.3
4.6
5.9
3.0
Struct-Spar
5.6
8.7
3.8
5.1
6.9
5.8
Struct-Noise
6.1
5.6
4.0
1.9
5.5
3.1
Struct-Imb
2.7
8.6
5.0
3.0
4.7
4.2
Label-Spar
2.5
2.6
2.7
4.7
4.1
3.0
Label-Noise
5.5
3.5
4.5
8.9
10.9
7.9
Label-Imb
10.1
13.6
9.5
6.6
14.4
10.6
Photo
2.8
2.6
2.5
2.5
--
0.6
0.2
0.3
-0.1
--
0.8
0.3
0.6
0.6
--
4.5
5.2
6.0
5.4
16.5
4.3
4.1
3.7
4.0
6.2
6
1.5
1.8
1.8
1.6
15.3
4
2.7
3.3
3.2
2.7
-1.9
5.5
6.4
5.9
4.3
19.1
9.9
11.0
11.9
10.2
-2.3
14 12
20
15
History
10
2.3
0.9
1.9
1.2
1.8
1.5
4.6
2.0
3.5
2.3
3.1
2.2
3.3
0.8
1.0
-1.1
1.0
0.9
6.2
8.0
5.3
5.3
5.8
4.9
8.4
7.7
5.0
4.9
6.1
5.7
2.7
4.3
3.0
2.9
3.4
2.6
0.9
1.1
0.7
1.7
2.0
1.4
2.0
0.9
0.5
1.7
2.9
2.2
8.0
8.2
5.6
0.0
0.4
0.0
1.6
1.8
1.6
1.3
--
25
2.4
3.0
2.7
2.8
--
20
6
4.3
5.0
5.0
4.2
--
15
4
1.8
2.1
2.2
2.1
20.4
10
1.8
1.8
2.1
2.0
-12.6
5
0.6
1.4
1.1
1.2
12.5
0
1.3
2.0
1.8
1.8
5.7
−5
0.9
1.1
0.9
1.2
-3.5
−10
24.0
24.1
21.3
23.1
0.6
−15
8
10 8
10
5
2 0
N GC
T GA
h NE raph T raph t ap Gr GE NGI G P G Tex G SA E
Model backbone
2 0
T ph AN Gra ter GI ap Ad
A OF
G ero
Z
0
−5
en Op aph Gr
T GA
N GC
−2 −4
h NE raph T raph t ap Gr GE NGI G P G Tex G SA E
h ap Gr pter a Ad
T AN
GI
Model backbone
Model backbone
A OF
G ero
Z
en Op aph Gr
Model backbone
Figure 2: Q1 model-paradigm robustness. Heatmaps report clean-to-degraded node-classification accuracy drops under the nine degradation scenarios, with color scales normalized per dataset. Cora
Q1: Scenario Validity and TAG-Backbone Robustness Accuracy (%)
Sparsity
WikiCS
Imbalance
Photo
Noise
Sparsity
Imbalance
90 85 80 75 70 rtC Be
se NG ram
rtC
Ba
Be
Ba se NG ram
g
rtD Be
se
Re
rtC
Ba
Be
se NG ram
rtC
ratio=0.2
Ba
Be
Ba se NG ram
g
rtD Be
Re
se
65 Ba
Setup. Q1 aims to verify whether the proposed 3 × 3 degradation space provides valid and meaningful stress-test scenarios for TAG learning. We use node classification as the anchor task and evaluate clean and degraded TAGs over nine datasets and nine degradation scenarios. The evaluation covers three standard GNNs (Kipf and Welling 2017; Veličković et al. 2018; Hamilton, Ying, and Leskovec 2017), seven LLM-related or broad TAG backbones (Zhu et al. 2024b; Chien et al. 2022; Tang et al. 2024; Huang et al. 2024; Zhao et al. 2023; Liu et al. 2024; Li et al. 2024b), and a separate GFM-style compatibility analysis with OpenGraph (Xia, Kao, and Huang 2024). Results. The results (Fig. 2) confirm that the constructed scenarios are valid stress tests. The internal quality statistics move in the intended directions, and the GCN anchor degrades under all nine scenarios, showing that the perturbations are both effective and non-collapsing. Label imbalance is the strongest stress case, with an average accuracy drop of 21.8 percentage points. Structure noise and sparsity are also consistently harmful, while text degradation causes milder but observable drops. Cross-backbone results further show that these degradation effects persist across traditional GNNs, LLM-GNNs, and GFM-style representatives.
Noise
95
Cite.
ratio=0.8
Figure 3: Q2 text-repair methods’ robustness under low and high perturbation ratios.
Insight. Overall, Q1 shows that the nine degradation scenarios are reasonable, effective, and non-trivial benchmark conditions. More importantly, the degradation effects are not limited to a single backbone: traditional GNNs, LLMGNNs, and GFM-style representatives all exhibit different degrees of sensitivity to low-quality TAG inputs. This motivates the need for TAG robustness evaluation beyond cleandata leaderboards.
Table 2: Node-classification accuracy (%) under text degradation with GCN as the shared backbone. Missing or failed runs are marked as OOM; the best available method for each dataset and scenario is highlighted. Scenario
Method
Cora
Citeseer
WikiCS
PubMed
Children
Text noise
Text sparsity
BaseModel CTD MLM BertDenoise RegexDenoise
84.44±0.46 84.75±0.38 83.70±0.53 83.39±0.98
78.06±0.57 77.59±0.31 77.22±0.71 77.64±0.59
65.96±0.28 65.17±0.58 65.09±0.29 64.86±0.62
81.70±0.59 81.76±0.28 81.76±0.31 82.16±0.69
85.49±0.36 85.26±0.19 85.44±0.32 84.95±0.38
44.58±0.29 44.24±0.12 44.28±0.18 44.16±0.90
78.45±0.11 79.12±0.16 67.21±0.19 78.43±0.06 79.32±0.27 67.16±0.17 78.53±0.11 OOM OOM 78.54±0.30 OOM OOM
BaseModel BertComplete NgramComplete PoDA UltraTAG S
85.98±0.49 85.42±1.11 85.67±0.38 84.81±1.19 86.29±0.56
74.29±0.41 73.77±0.45 73.88±0.24 74.92±0.41 73.46±0.09
66.46±0.32 65.40±0.50 65.30±0.73 65.70±0.50 65.80±0.83
79.61±0.34 80.14±0.69 80.28±0.56 79.89±0.90 80.07±0.73
85.73±0.41 85.86±0.38 86.00±0.29 85.90±0.10 86.05±0.08
45.65±0.37 45.81±0.78 45.71±0.65 45.77±0.48 45.78±0.35
80.79±0.20 80.53±0.03 80.75±0.26 81.05±0.30 80.88±0.21
Text imbalance BaseModel BertComplete NgramComplete PoDA UltraTAG S
82.53±1.59 85.49±0.75 84.26±1.97 84.19±0.95 82.35±0.11
77.85±0.39 77.27±0.83 77.22±0.71 77.74±0.54 76.02±1.18
75.19±0.32 75.65±0.49 73.93±0.44 78.56±0.23 77.06±0.18
81.56±0.69 81.12±0.70 81.33±0.27 82.17±0.47 80.28±0.22
87.32±0.13 87.41±0.13 87.15±0.08 87.52±0.35 87.76±0.14
46.62±0.39 46.77±0.40 46.34±0.50 47.34±0.07 46.54±0.43
79.83±0.22 76.71±0.19 66.74±0.08 79.75±0.13 76.78±0.31 OOM 79.67±0.08 OOM OOM 80.32±0.22 OOM OOM 79.91±0.37 OOM OOM
Text sparsity
Text noise
Cora
2
10
1
10
0
Cite.
Sparsity
Noise
95
Accuracy (%)
10
History
80.41±0.14 67.10±0.13 80.57±0.27 OOM 80.71±0.39 67.03±0.07 OOM OOM 80.63±0.15 67.43±0.14
WikiCS
Imbalance
Arxiv
Photo Sparsity
Noise
Imbalance
85 75 65
−1
ratio=0.2
Figure 4: Runtime of representative Q2 text-oriented baselines. Bars report runtime in log-scale minutes, with different colors denoting datasets ordered as Cora, CiteSeer, WikiCS, and Photo within each method group.
Q2: Effectiveness, Robustness, and Efficiency under Text Degradation Q2 evaluates text-oriented baselines with GCN. Three denoising methods target text noise (Sun and Jiang 2019; Flint et al. 2017), whereas four completion/generation methods target text sparsity and imbalance (Langkilde and Knight 1998; Wang et al. 2019; Zhang et al. 2025b). Results and insights. Table 2, Figure 3, and Figure 4 show that text-oriented repair is highly scenario- and dataset-dependent. Under text noise, denoising methods provide gains on some datasets, but the BaseModel remains competitive or best in many cases. Under text sparsity and imbalance, completion and generation better match insufficient semantic information, yet their gains remain inconsistent across datasets. Beyond average accuracy, some methods are effective only under mild degradation and become less stable as the perturbation ratio increases. Runtime further exposes substantial preprocessing cost and OOM failures for generationor LM-based repair. Future text-robust TAG methods should decide when and how much to intervene using degradation type, graph context, and computational budget.
4G
TA M
LT E
B
se Ba
SU
se
SL -G SE
se
Ba
Ba
GA ug ST AB LE
4G
55 TA M
DA
LT E
Po
B
-C RT
BE
se
se
Ba
Ba
G TA
tra
SU
Ul
se
ram
SL
NG
Ba
se
Ba
-G
x
ge
Re
SE
D
E
CT
se
se
Ba
Ba
ST AB L
10
Text imbalance
3
GA ug
Runtime (min, log)
10
Photo
ratio=0.8
Figure 5: Q3 structure-repair methods’ robustness under low and high perturbation ratios.
Q3: Effectiveness, Robustness, and Efficiency under Structural Degradation Q3 compares structure-oriented baselines under the same GCN backbone. We select scenario-matched subsets rather than applying the full GSL library to every topology failure: four relation-recovery methods for structure sparsity, four edge-purification methods for structure noise, and imbalance-oriented graph methods for structure imbalance (Guo et al. 2024; Zhang et al. 2024b; Zou et al. 2023; Liu et al. 2022; Li et al. 2022; Fang et al. 2024; Ju et al. 2023; Yun et al. 2022; Liu, Nguyen, and Fang 2021; Song, Park, and Yang 2022; Chen et al. 2022). Results and insights. Table 3, Figure 5, and Figure 6 show that topology repair is strongly failure- and datasetdependent. Graph editing methods are often effective under structure noise, where harmful edges need to be removed or corrected. In contrast, structure sparsity and imbalance show more mixed results: the BaseModel remains competitive on several datasets, and some repair methods even reduce performance, suggesting possible negative transfer from unreliable rewiring or reweighting. Perturbation-strength results further show that some methods lose stability under stronger degradation, while runtime results reveal sub-
Table 3: Node-classification accuracy (%) under structural degradation with GCN as the shared backbone. For structure sparsity and structure noise, we report scenario-matched GSL subsets rather than all structure-learning methods. Missing or failed runs are marked as OOM; the best available method for each dataset and scenario is highlighted. Scenario
Method
Cora
Citeseer
WikiCS
PubMed
Children
Structure noise
Structure sparsity
BaseModel GAugLLM GraphEdit LLM4RGNN STABLE
82.29±0.74 74.85±1.57 83.46±0.53 78.78±0.55 79.34±0.80
71.21±0.45 68.55±0.79 75.08±0.72 69.07±0.63 72.62±0.86
65.53±0.39 65.34±0.48 65.80±0.45 65.86±0.54 65.02±0.23
69.16±0.72 72.97±0.36 80.01±0.30 76.06±0.58 73.37±0.38
81.82±0.23 84.36±0.17 87.17±0.11 84.81±0.70 83.28±0.22
43.70±0.43 43.34±0.80 47.92±0.79 44.23±0.24 46.41±0.65
BaseModel GraphEdit LLM4RGNN SEGSL SUBLIME
82.78±1.38 82.04±0.75 77.80±0.56 82.23±0.77 80.38±1.13
74.40±0.45 75.50±0.39 69.12±0.57 74.19±1.11 76.12±0.77
66.20±0.37 65.06±0.46 63.46±0.08 66.61±0.44 65.40±0.29
74.63±0.41 81.59±0.33 75.85±0.78 76.99±1.08 77.62±0.36
85.44±0.15 86.87±0.12 84.01±0.61 OOM 85.39±0.41
43.53±0.72 76.87±0.11 79.80±0.15 65.64±0.11 49.40±0.27 80.21±0.38 OOM OOM 42.91±0.59 73.36±0.28 80.02±0.35 OOM OOM OOM OOM OOM 46.62±0.12 OOM OOM OOM
Structure imbalance BaseModel GraphPatcher LTE4G TailGNN TAM
85.67±0.38 78.91±1.23 85.42±0.49 81.30±7.34 86.78±0.87
75.60±1.07 69.59±1.65 74.29±0.31 73.72±0.74 76.70±1.02
66.11±0.62 59.96±0.77 60.21±0.29 64.12±2.20 65.80±0.51
79.82±0.28 72.23±1.55 79.47±0.16 79.54±0.67 80.86±0.68
87.36±0.16 88.87±0.00 85.44±0.05 87.67±1.45 87.03±0.24
47.34±0.50 48.13±1.09 41.09±0.33 43.96±0.61 48.39±1.47
Struct. sparsity
80.37±0.05 76.42±0.66 71.08±0.21 79.16±0.65 81.27±1.05
Cite.
81.58±0.20 80.13±0.79 74.71±0.74 80.36±0.55 81.54±0.45
WikiCS
Imbalance
68.75±0.07 69.72±0.31 62.22±0.22 68.09±0.52 69.79±1.49
Photo Sparsity
Noise
Imbalance
Figure 6: Runtime of representative Q3 structure-oriented baselines. Bars report runtime in log-scale minutes, with different colors denoting datasets ordered as Cora, CiteSeer, WikiCS, and Photo within each method group.
stantial overhead and OOM failures for several structurelearning pipelines. These results suggest that future structure-robust TAG methods should treat topology repair as selective intervention rather than global graph rewiring, jointly estimating which edges are trustworthy, which missing links are worth recovering, and which nodes are most vulnerable to structural defects under the downstream task.
Q4: Effectiveness, Robustness, and Efficiency under Label Degradation Q4 evaluates label-oriented baselines under scarce, corrupted, and long-tailed supervision. We use two labelsparsity methods, three noisy-label methods, and four imbalance-oriented methods (Lee et al. 2022; Xie, Wang, and Kuo 2021; Du et al. 2023; Zhu et al. 2024a; Qian et al. 2023; Yun et al. 2022; Liu, Nguyen, and Fang 2021; Song, Park, and Yang 2022; Chen et al. 2022). Results and insights. Table 4, Figure 7, and Figure 8
ratio=0.2
se
4G
Ba
LT E
se Ba
TA M
N
GH op
NN
Gr aF
RT G
se
GN N
30 PI
P
se
M TA
Ba
se
Ba
4G
SL
Ba
-G
SE
LT E
4R
M
LL
se
se
Ba
Ba
ug
TA M
GA
N
hE ap Gr
40 GH op
se
r he atc
NN
Ba
dit
50
Gr aF
−1
60
RT G
10
74.75±0.34 74.48±0.46 60.43±0.42 75.45±0.52 OOM OOM 78.64±0.03 81.17±0.18 OOM 71.70±0.29 80.57±0.16 OOM OOM OOM OOM
70
se
10
0
Arxiv
80
Ba
1
GN N
10
History
90
PI
10
2
Cora Sparsity
Noise
Accuracy (%)
10
Struct. imbalance
OOM
Runtime (min, log)
Struct. noise 3
Photo
ratio=0.8
Figure 7: Q4 label-repair methods’ robustness under low and high perturbation ratios.
show that supervision failures differ substantially across failure modes. Label imbalance causes severe drops on several datasets, and imbalance-aware methods such as TAM and LTE4G can bring large gains when long-tailed supervision dominates. Label sparsity also benefits from label-efficient methods on many datasets, although the best method varies between GraFN and GraphHop. In contrast, label noise is more difficult to repair: the BaseModel remains competitive in many cases, suggesting that explicit label correction may yield unstable gains when corrupted labels are hard to identify. The perturbation-strength results further show that supervision-oriented methods are sensitive to degradation severity, while runtime results indicate that they are generally more lightweight than text or structure repair methods, despite some costly or failed runs. These results suggest that future supervision-robust TAG methods should move beyond fixed label correction, propagation, or reweighting, and instead learn uncertainty-aware supervision policies that decide when to trust, correct, propagate, or rebalance labels under different forms of supervision failure.
Table 4: Node-classification accuracy (%) under label degradation with GCN as the shared backbone. Missing or failed runs are marked as OOM; the best available method for each dataset and scenario is highlighted. Scenario
Method
Label noise
BaseModel PIGNN RNCGLN RTGNN
Label sparsity
BaseModel 85.85±1.11 76.70±0.65 65.20±0.41 81.16±0.15 87.65±0.33 44.09±0.55 82.15±0.41 81.26±0.20 69.35±0.07 GraFN 88.50±0.56 76.18±0.41 64.99±0.34 84.69±0.25 88.70±0.27 48.41±0.40 86.63±0.28 OOM 74.39±0.44 OOM GraphHop 87.58±0.11 75.44±0.09 64.56±0.11 84.17±0.05 89.16±0.03 47.19±0.01 84.46±0.16 82.93±0.04
Label imbalance BaseModel LTE4G TAM TOPOAUC
Cora
Citeseer
WikiCS
PubMed
82.90±0.59 84.62±0.28 73.86±0.21 78.66±1.67
74.92±1.10 74.82±0.09 73.04±0.16 75.03±1.11
64.29±0.46 62.45±0.23 63.74±0.09 64.01±0.20
78.91±0.09 76.76±1.46 76.11±0.22 72.80±1.34
85.83±0.42 85.73±0.29 86.67±0.19 80.73±0.56
78.28±2.59 84.46±0.24 82.27±1.13 81.79±1.84
Runtime (min, log)
Label sparsity 10
1
10
0
10
Label noise
History
Arxiv
46.62±0.37 81.07±0.24 81.77±0.25 68.56±0.04 46.86±0.17 76.83±1.18 80.99±0.35 62.60±0.76 48.44±0.09 OOM OOM OOM 37.29±0.76 68.70±5.74 79.53±0.44 OOM
Cora
Label imbalance T-Spa+S-Spa
T-Noi+S-Noi
T-Spa+L-Imb
WikiCS
12.3
12.3
21.4
Base
NG
GE
13.4
12.5
15.7
Base
CTD
GAug
18.9
17.1
Base
NG
--
10.2
10.3
8.3
Base
NG
GE
--
12.0
12.1
10.1
Base
CTD
GAug
4.9
23.6
23.9
TAM
Base
NG
15.1
20.3
17.1
GAug
RNC
Base
--
--
--
0.3
--
TAM
−1
Ba
se
se
p
N
o hH rap
aF Gr
Ba
G PI
G
se
NN
NN
M TA
Ba
G RT
S-Noi+L-Noi
G
E4 LT
Cora
Cite.
WikiCS
0.7
TS
TI
se
Link prediction (AUC)
SN
SS
TN
1.0
TS
TI
SN
SS
0.9
0.5
0.8 B
se
SU
Ba
se
dit
Ba
GE
se
ram
NG
ram
Ba
NG
g
se Ba
se
Re
Ba
B
se
SU
Ba
se
dit
Ba
GE
se
ram
NG
ram
Ba
g
se Ba
NG
Re
se
0.3 Ba
Base
Lower is better
--
t Tex
Str
0
.
bel
uct
5
se
t Tex
15
20
Ba
La
10
--
12.3
5.8
GAug
RNC
.
uct
Str
bel
La
Acc. drop (pp)
Figure 10: Q5 Performance under composite scenarios.
PubMed
Node clustering (NMI) TN
12.8
Ba
Figure 8: Runtime of representative Q4 label-oriented baselines. Bars report runtime in log-scale minutes, with different colors denoting datasets ordered as Cora, CiteSeer, WikiCS, and Photo within each method group.
Score
Photo
74.24±1.35 13.40±2.06 66.59±1.48 79.21±0.40 25.62±2.22 75.05±0.40 39.69±1.20 39.79±0.27 76.70±0.47 63.80±0.05 80.23±0.08 84.23±0.11 39.89±0.47 71.82±0.47 76.05±0.15 61.44±0.19 OOM 78.82±2.13 76.57±0.37 82.43±0.67 86.61±0.06 50.74±2.25 84.34±1.00 82.89±0.37 79.44±0.19 OOM OOM OOM OOM OOM OOM OOM
2
10
Children
TN/TS/TI: text noise/sparsity/imbalance; SN/SS: structure noise/sparsity. Hollow markers denote Base.
Figure 9: Q5 Performance under different downstream tasks.
Q5: Extension to Composite Scenarios and Downstream Tasks Q5 extends the main node-classification analysis to harder and broader settings. We first construct four two-factor composite scenarios by combining text, structure, and label degradation, and compare the GCN with the best matched baselines selected from Q2–Q4. We further evaluate representative text and structure scenarios on node clustering and link prediction to examine whether the observed robustness patterns generalize beyond node classification.
Results and insights. Figures 9 and 10 show that the proposed degradation settings remain informative beyond node classification. Similar performance gaps appear in node clustering and link prediction, indicating that OpenRTAG captures quality issues that affect TAG learning across tasks. Under composite degradation, existing methods show less stable gains than in single-scenario settings, suggesting that repairing one modality alone may be insufficient when text, structure, and label quality issues coexist. These results suggest that future robust TAG methods should move beyond one-defect-at-a-time repair and jointly optimize text, topology, and supervision signals under multi-factor degradation.
Conclusion We present OpenRTAG, a robustness benchmark for textattributed graph learning under low-quality data. OpenRTAG defines a unified 3 × 3 degradation space across text, structure, and labels, covering sparsity, noise, and imbalance, and evaluates representative models and baselines across multiple datasets and tasks. The results show that TAG robustness is scenario-dependent and that existing methods remain limited under strong or composite degradation, motivating future methods that jointly address text, topology, and supervision quality. A limitation is that Open-
RTAG focuses on representative degradation settings, leaving more dynamic and domain-specific quality issues for further investigation in future work.
References Chen, J.; Xu, Q.; Yang, Z.; Cao, X.; and Huang, Q. 2022. A Unified Framework against Topology and Class Imbalance. In Proceedings of the 30th ACM International Conference on Multimedia, 180–188. Chen, Z.; Mao, H.; Liu, J.; Song, Y.; Li, B.; Jin, W.; Fatemi, B.; Tsitsulin, A.; Perozzi, B.; Liu, H.; et al. 2024. TextSpace Graph Foundation Models: Comprehensive Benchmarks and New Insights. arXiv preprint arXiv:2406.10727. Chien, E.; Liu, W.; Wang, P.; Wang, W. Y.; Yu, P. S.; and Milenkovic, O. 2022. Node Feature Extraction by SelfSupervised Multi-scale Neighborhood Prediction. In International Conference on Learning Representations. Du, X.; Bian, T.; Rong, Y.; Han, B.; Liu, T.; Xu, T.; Huang, W.; Li, Y.; and Huang, J. 2023. Noise-robust Graph Learning by Estimating and Leveraging Pairwise Interactions. Transactions on Machine Learning Research. Fang, Y.; Fan, D.; Zha, D.; and Tan, Q. 2024. GAugLLM: Improving Graph Contrastive Learning for Text-Attributed Graphs with Large Language Models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 747–758. Feng, J.; Liu, H.; Kong, L.; Chen, Y.; and Zhang, M. 2024. TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models. arXiv:2406.14683. Flint, E.; Ford, E.; Thomas, O.; Caines, A.; and Buttery, P. 2017. A Text Normalisation System for Non-Standard English Words. In Proceedings of the 3rd Workshop on Noisy User-generated Text, 107–115. Guo, Z.; Xia, L.; Yu, Y.; Wang, Y.; Yang, Z.; Wei, W.; Pang, L.; Chua, T.-S.; and Huang, C. 2024. GraphEdit: Large Language Models for Graph Structure Learning. arXiv:2402.15183. Hamilton, W. L.; Ying, Z.; and Leskovec, J. 2017. Inductive Representation Learning on Large Graphs. In Advances in Neural Information Processing Systems. He, X.; Bresson, X.; Laurent, T.; Perold, A.; LeCun, Y.; and Hooi, B. 2023. Harnessing Explanations: LLM-to-LM Interpreter for Enhanced Text-Attributed Graph Representation Learning. arXiv preprint arXiv:2305.19523. Hu, W.; Fey, M.; Zitnik, M.; Dong, Y.; Ren, H.; Liu, B.; Catasta, M.; and Leskovec, J. 2020. Open Graph Benchmark: Datasets for Machine Learning on Graphs. Advances in Neural Information Processing Systems. Huang, X.; et al. 2024. Can GNN be Good Adapter for LLMs? In Proceedings of the ACM Web Conference 2024. Ju, M.; Zhao, T.; Yu, W.; Shah, N.; and Ye, Y. 2023. Mitigating Degree Bias for Graph Neural Networks via Test-Time Augmentation. In Advances in Neural Information Processing Systems.
Kipf, T. N.; and Welling, M. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Representations. Langkilde, I.; and Knight, K. 1998. The Practical Value of N-Grams Is in Generation. In Natural Language Generation. Lee, J.; Oh, Y.; In, Y.; Lee, N.; Hyun, D.; and Park, C. 2022. GraFN: Semi-Supervised Node Classification on Graph with Few Labels via Non-Parametric Distribution Assignment. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. Li, K.; Liu, Y.; Ao, X.; Chi, J.; Feng, J.; Yang, H.; and He, Q. 2022. Reliable Representations Make A Stronger Defender: Unsupervised Structure Refinement for Robust GNN. arXiv:2207.00012. Li, X.; Wu, Z.; Wu, J.; Cui, H.; Jia, J.; Li, R.-H.; and Wang, G. 2024a. Graph Learning in the Era of LLMs: A Survey from the Perspective of Data, Models, and Tasks. arXiv preprint arXiv:2412.12456. Li, Y.; Wang, P.; Li, Z.; Yu, J. X.; and Li, J. 2024b. ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 1725– 1735. Li, Y.; Wang, P.; Zhu, X.; Chen, A.; Jiang, H.; Cai, D.; Chan, V. W. K.; and Li, J. 2024c. GLBench: A Comprehensive Benchmark for Graph with Large Language Models. In Advances in Neural Information Processing Systems Datasets and Benchmarks Track. Liu, H.; Feng, J.; Kong, L.; Liang, N.; Tao, D.; Chen, Y.; and Zhang, M. 2024. One for All: Towards Training One Graph Model for All Classification Tasks. In International Conference on Learning Representations. Liu, J.; Yang, C.; Lu, Z.; Chen, J.; Li, Y.; Zhang, M.; Bai, T.; Fang, Y.; Sun, L.; Yu, P. S.; and Shi, C. 2023. Towards Graph Foundation Models: A Survey and Beyond. arXiv:2310.11829. Liu, Y.; Zheng, Y.; Zhang, D.; Chen, H.; Peng, H.; and Pan, S. 2022. Towards Unsupervised Deep Graph Structure Learning. In Proceedings of the ACM Web Conference 2022. Liu, Z.; Nguyen, T.-K.; and Fang, Y. 2021. Tail-GNN: TailNode Graph Neural Networks. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. McAuley, J.; Targett, C.; Shi, Q.; and van den Hengel, A. 2015. Image-based Recommendations on Styles and Substitutes. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval. Mernyei, P.; and Cangea, C. 2020. Wiki-CS: A WikipediaBased Benchmark for Graph Neural Networks. arXiv preprint arXiv:2007.02901. Qian, S.; Ying, H.; Hu, R.; Zhou, J.; Chen, J.; Chen, D. Z.; and Wu, J. 2023. Robust Training of Graph Neural Networks via Noise Governance. In Proceedings of the 16th ACM International Conference on Web Search and Data Mining.
Qin, J.; Yuan, H.; Sun, Q.; Xu, L.; Yuan, J.; Huang, P.; Wang, Z.; Fu, X.; Peng, H.; Li, J.; and Yu, P. S. 2025. IGL-Bench: Establishing the Comprehensive Benchmark for Imbalanced Graph Learning. In International Conference on Learning Representations. Qin, J.; Yuan, H.; Sun, Q.; Xu, L.; Yuan, J.; Huang, P.; Wang, Z.; Fu, X.; Peng, H.; Li, J.; et al. 2024. IGL-Bench: Establishing the Comprehensive Benchmark for Imbalanced Graph Learning. arXiv preprint arXiv:2406.09870. Ren, X.; Tang, J.; Yin, D.; Chawla, N. V.; and Huang, C. 2024. A Survey of Large Language Models for Graphs. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Song, J.; Park, J.; and Yang, E. 2022. TAM: TopologyAware Margin Loss for Class-Imbalanced Node Classification. In International Conference on Machine Learning. Sun, Y.; and Jiang, H. 2019. Contextual Text Denoising with Masked Language Model. In Proceedings of the 5th Workshop on Noisy User-generated Text. Tang, J.; Yang, Y.; Wei, W.; Shi, L.; Su, L.; Cheng, S.; Yin, D.; and Huang, C. 2024. GraphGPT: Graph Instruction Tuning for Large Language Models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; and Bengio, Y. 2018. Graph Attention Networks. In International Conference on Learning Representations. Wang, L.; Zhao, W.; Jia, R.; Li, S.; and Liu, J. 2019. Denoising based Sequence-to-Sequence Pre-training for Text Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. Wang, Z.; Sun, D.; Zhou, S.; Wang, H.; Fan, J.; Huang, L.; and Bu, J. 2024. NoisyGL: A Comprehensive Benchmark for Graph Neural Networks under Label Noise. In Advances in Neural Information Processing Systems Datasets and Benchmarks Track. Xia, L.; Kao, B.; and Huang, C. 2024. OpenGraph: Towards Open Graph Foundation Models. In Findings of the Association for Computational Linguistics: EMNLP 2024, 2365– 2379. Xie, T.; Wang, B.; and Kuo, C.-C. J. 2021. GraphHop: An Enhanced Label Propagation Method for Node Classification. arXiv:2101.02326. Yan, H.; Li, C.; Long, R.; Yan, C.; Zhao, J.; Zhuang, W.; Yin, J.; Zhang, P.; Han, W.; Sun, H.; et al. 2023. A Comprehensive Study on Text-Attributed Graphs: Benchmarking and Rethinking. Advances in Neural Information Processing Systems, 36: 17238–17264. Yang, Z.; Cohen, W. W.; and Salakhutdinov, R. R. 2016. Revisiting Semi-Supervised Learning with Graph Embeddings. In International Conference on Machine Learning. Yun, S.; Kim, K.; Yoon, K.; and Park, C. 2022. LTE4G: Long-Tail Experts for Graph Neural Networks. In Proceedings of the 31st ACM International Conference on Information and Knowledge Management.
Zhang, D. C.; Yang, M.; Ying, R.; and Lauw, H. W. 2024a. Text-Attributed Graph Representation Learning. In Companion Proceedings of the ACM Web Conference 2024. Zhang, Z.; Li, X.; Lei, Z.; Zeng, G.; Li, R.; and Wang, G. 2025a. Rethinking Graph Structure Learning in the Era of LLMs. arXiv preprint arXiv:2503.21223. Zhang, Z.; Li, X.; Li, R.-H.; Zhou, B.; Li, Z.; and Wang, G. 2025b. Toward General and Robust LLM-enhanced Textattributed Graph Learning. arXiv:2504.02343. Zhang, Z.; Wang, X.; Zhou, H.; Yu, Y.; Zhang, M.; Yang, C.; and Shi, C. 2024b. Can Large Language Models Improve the Adversarial Robustness of Graph Neural Networks? arXiv preprint arXiv:2408.08685. Zhao, J.; Zhuo, L.; Shen, Y.; Qu, M.; Liu, K.; Bronstein, M.; Zhu, Z.; and Tang, J. 2023. GraphText: Graph Reasoning in Text Space. Uncertain mapping for method name GraphText, arXiv:2310.01089. Zhu, Y.; Feng, L.; Deng, Z.; Chen, Y.; Amor, R.; and Witbrock, M. 2024a. Robust Node Classification on Graph Data with Graph and Label Noise. In Proceedings of the AAAI Conference on Artificial Intelligence. Zhu, Y.; Wang, Y.; Shi, H.; and Tang, S. 2024b. Efficient Tuning and Inference for Large Language Models on Textual Graphs. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. Zhu, Y.; Xu, W.; Zhang, J.; Liu, Q.; Wu, S.; and Wang, L. 2021. Deep Graph Structure Learning for Robust Representations: A Survey. arXiv preprint arXiv:2103.03036. Zou, D.; Peng, H.; Huang, X.; Yang, R.; Li, J.; Wu, J.; Liu, C.; and Yu, P. S. 2023. SE-GSL: A General and Effective Graph Structure Learning Framework through Structural Entropy Optimization. In Proceedings of the ACM Web Conference 2023.