arXiv:2606.07492v1 [cs.IR] 5 Jun 2026
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies Ekaterina Grishina∗
Stepan Kuznetsov∗
Askar Tsyganov∗
[email protected] HSE University Moscow, Russian Federation
[email protected] HSE University Moscow, Russian Federation
[email protected] HSE University Moscow, Russian Federation
Ilya Ivanov∗
Daria Korovaitceva∗
Margarita Rusanova∗
[email protected] HSE University Moscow, Russian Federation
[email protected] HSE University Moscow, Russian Federation
[email protected] HSE University Moscow, Russian Federation
Uliana Parkina∗
Alexander Derevyagin∗
Evgeny Frolov
[email protected] HSE University Moscow, Russian Federation
[email protected] AXXX HSE University Moscow, Russian Federation
[email protected] AXXX HSE University Moscow, Russian Federation
Sergey Samsonov
Anton Lysenko
[email protected] HSE University Moscow, Russian Federation
[email protected] HSE University Moscow, Russian Federation
Abstract
Keywords
The ranking of recommendation algorithms is a challenging problem since model performance is sensitive to dataset characteristics such as sparsity, sequential structure, and scale. This drives a demand for a proper methodology for fair comparison between algorithms. Naive aggregation of performance metrics (e.g., averaging NDCG over benchmarks) can yield misleading rankings, undermining practical selection. To address this problem, we introduce a novel, data-driven ranking methodology based on Bradley-Terry (BT) model. We demonstrate that the obtained ranking depends on key dataset statistics. Additionally, we propose a novel metric for evaluating ranking consistency and demonstrate robustness of our ranking to incomplete data. Finally, we introduce a dataset-specific methodology for ranking algorithms on unseen datasets without running the models, relying on extensions of the Bradley–Terry framework, including BT trees and BT models with covariates.
Bradley-Terry Model, Pairwise Comparison, Recommender Systems, Algorithm Ranking
CCS Concepts • Information systems → Recommender systems; • Mathematics of computing → Probabilistic algorithms; • Computing methodologies → Learning to rank. ∗ Equal contribution
This work is licensed under a Creative Commons Attribution 4.0 International License. KDD ’26, Jeju Island, Republic of Korea © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2259-2/2026/08 https://doi.org/10.1145/3770855.3817890
ACM Reference Format: Ekaterina Grishina, Stepan Kuznetsov, Askar Tsyganov, Ilya Ivanov, Daria Korovaitceva, Margarita Rusanova, Uliana Parkina, Alexander Derevyagin, Evgeny Frolov, Sergey Samsonov, and Anton Lysenko. 2026. Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD ’26), August 09–13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3770855.3817890 Resource Availability: The source code of this paper has been made publicly available at https://doi. org/10.5281/zenodo.20383718 or https://github.com/fallnlove/btl_recsys.
1
Introduction
In recent years, the field of recommendation systems has been characterized by a wide variety of tasks, datasets, and algorithmic approaches. Almost every dataset has its own specifics: the types of interactions, sparsity, scale and time structure can vary. As a result, the same algorithm can demonstrate high performance on some subset of datasets and significantly lose on others. This makes it challenging for practitioners to choose the most suitable method for a specific task without costly direct comparison. From an academic research perspective, the variety of existing algorithms poses a challenge of robust and scalable evaluation. Direct comparison by individual metrics and datasets often gives a fragmentary picture and scales poorly with an increasing number of methods and datasets. These issues highlight the need for a principled methodology for fair algorithm comparison and dataset-aware selection.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Grishina et al.
In this work, we propose to evaluate the recommendation algorithms within the probabilistic ranking framework based on the Bradley-Terry model [3] and its modifications, including covariateadjusted [11, 33] and tree-based variants [14, 38, 46, 47]. We conduct an extensive benchmark of 14 recommender algorithms across 89 datasets, using a framework that aggregates disparate experimental results into a single quality scale to provide more informed recommendations for algorithm selection based on dataset taxonomy. Our main contributions are: • We introduce a theoretically grounded methodology for aggregating repeated evaluations across diverse datasets into a probabilistic ranking of recommendation algorithms based on the Bradley–Terry model. This yields rankings tailored to key dataset classes, such as sequential or sparse datasets, supporting baseline selection (Section 6.3). • We introduce the "transitive triplets" metric for evaluating ranking consistency under incomplete benchmark results (Section 4). Using this metric, we show that Bradley–Terry-based rankings are substantially more stable than rankings obtained by simple metric aggregation (Section 6.1). • We extend the framework with dataset-aware baseline selection using BT trees and covariate-adjusted BT [33, 38] enabling ranking prediction on a target dataset without running candidate models (Section 7.3).
existence of certain true ranking or suppose the possibility of comparing items dynamically. Thus, they are either inapplicable in real-life scenarios or the resulting rankings may depend on the comparison schedule and lack uncertainty quantification. Regarding the comparison of algorithms over multiple datasets, traditional methods rely on frequentist statistical tests [10] or aggregated performance metrics [35]. Although these approaches can provide global rankings, they may be unstable or misleading due to the weak theoretical foundation. Bayesian alternatives have been proposed to address these shortcomings by enabling direct pairwise comparisons and probabilistic statements about algorithm performance [2]. Recent works apply Bradley-Terry-type models to algorithm benchmarking, demonstrating their suitability for deriving consistent and interpretable rankings across datasets [44]. However, to the best of our knowledge, such approaches have not yet been applied to RecSys algorithms.
2
3.1
Related work
Pairwise comparison models have a long history as a principled approach to ranking, originating with the Bradley-Terry (BT) model [3], first formulated by Zermelo [48]. Due to their interpretability properties, BT-based models have been widely adopted in various domains, including sports, chess, and other competitions, ranking of AI models [6]. Several extensions of the Bradley-Terry model have been proposed to address different comparison settings, such as the Plackett-Luce model [28], which builds rankings based on listwise comparisons instead of pairwise. Other extensions include the Thurstone-Mosteller [26, 40] and Rao-Kupper [30] models, which account for ties or latent score variability. Recently, Bayesian formulations of the Bradley-Terry model have been proposed, allowing for uncertainty-aware inference and more robust pairwise comparisons. These formulations have been applied to evaluation of classification algorithms across multiple datasets [44]. A notable extension of the Bradley-Terry model is the covariate-adjusted BT [11, 33], that allows incorporating information about the varying context of each singular comparison (e.g. dataset characteristics). A special case of this approach is BT trees [38], which recursively partition the subjects and identify groups of subjects with homogeneous preference scalings in a data-driven way. For instance, in AutoML benchmark [14], BT-trees have been used to discover subsets of tasks where the relative AutoML framework rankings differ. Beyond statistical models, a parallel line of work focuses on algorithmic methods for reconstructing rankings from pairwise comparisons. These approaches aim to efficiently recover approximate or exact rankings via dynamic or static algorithms [19, 45], attempting to make as little pairwise comparisons as possible. It is important to note that such methods essentially presume the
3
Bradley-Terry model
Our approach is based on probabilistic ranking. In this section, we first describe the pairwise comparison model; then we move on to the methods for computing the ranking from pairwise comparisons. Finally, we describe the Plackett-Luce model, which produces rankings based on multiple (listwise) comparisons.
Pairwise comparisons properties
The Bradley-Terry model [3, 48] is a probability model for ranking "players" in "tournaments" based on pairwise comparisons or "matches". Consider the setting of a tournament between 𝑛 players. Within this model, the probability that the player 𝑖 wins over player 𝑗 is modeled as Pr(𝑖 ≻ 𝑗) =
𝑝𝑖 , 𝑝𝑖 + 𝑝 𝑗
where 𝑝𝑖 ≥ 0 are weights assigned to each of 𝑛 players, representing their strengths or abilities. The final ranking of the players is defined by the rank of their weights 𝑝𝑖 . The weights are invariant to a multiplicative constant, so in order to get a unique set, an additional Í Î constraint is usually imposed, e.g., 𝑛𝑖=1 𝑝𝑖 = 1 or 𝑛𝑖=1 𝑝𝑖 = 1. In our setup, instead of "players", we want to rank recommender algorithms based on pairwise comparisons between them. We say that the algorithm 𝑖 beats the algorithm 𝑗 on a dataset, if it achieves a higher metric on this dataset. This is how we get the tournament table 𝑊 with 𝑊𝑖 𝑗 being the number of datasets, where algorithm 𝑖 beats algorithm 𝑗. We assume that 𝑊𝑖𝑖 = 0. If the algorithms have achieved almost the same metric value, we can classify it as a “tie”. There are many approaches to handling ties, e.g [8] modifies original model by introducing additional parameters, while [44] leaves the original model unchanged, but modifies 𝑊 . Based on the work [44], we take the ties into account by adding 0.5 to both 𝑊𝑖 𝑗 and 𝑊 𝑗𝑖 in case if the achieved metric intervals metrici ± stdi and metricj ± stdj overlap (see definition of intervals in Section 5). After that, we run any standard (unaware of ties) Bradley-Terry algorithm on the obtained matrix 𝑊 .
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
3.2
Estimating parameters in the Bradley-Terry model
Within the classic Bradley-Terry (BT) model [3, 48], the likelihood Î of the tournaments’ outcomes is 1≤𝑖,𝑗 ≤𝑛 [𝑃 (𝑖 ≻ 𝑗)]𝑊𝑖 𝑗 and the loglikelihood is Ö ℓ (𝑝) = ln [Pr(𝑖 ≻ 𝑗)]𝑊𝑖 𝑗 =
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
and the approach of rank centrality estimates it using the stationary measure 𝜋ˆ of 𝑃: 𝜋ˆ ⊤ 𝑃 = 𝜋ˆ ⊤ . Í ∗ Imposing the condition 𝑖 𝜃 𝑖 = 0, we get 𝑛
𝜃 𝑖∗ = log 𝜋ˆ𝑖 −
1 ∑︁ 𝜋ˆ𝑘 . 𝑛 𝑘=1
1≤𝑖,𝑗 ≤𝑛
=
∑︁
[𝑊𝑖 𝑗 (ln(𝑝𝑖 ) − ln(𝑝𝑖 + 𝑝 𝑗 ))].
1≤𝑖,𝑗 ≤𝑛
Zermelo [48] showed that this expression has a single maximum and suggested to find it by a simple iteration: Í𝑛 𝑗=1 𝑊𝑖 𝑗 ′ 𝑝𝑖 ← Í𝑛 , (𝑊 + 𝑊 𝑗𝑖 )/(𝑝𝑖 + 𝑝 𝑗 ) 𝑖𝑗 𝑗=1 ! 1/𝑛 𝑛 Ö ′ . 𝑝𝑖 ← 𝑝𝑖 / 𝑝𝑗
3.3
(𝑦𝑖 1 ≻ 𝑦𝑖 2 ≻ · · · ≻ 𝑦𝑖𝑇𝑖 ) of 𝑇𝑖 ≤ 𝑛 objects as input, where 𝑛 is the total number of objects, and models probabilities
𝑗=1
Theoretical properties of the Bradley-Terry model estimates are fairly well understood, see recent works [13], [36], and references therein. However, the classic Bradley-Terry model has a limitation: it provides only a point estimate of parameters 𝑝𝑖 without confidence intervals. To address it, the authors of [44] proposed the Bayesian Bradley-Terry model. This model assumes that there are 𝑁𝑖 𝑗 = 𝑁 𝑗𝑖 comparisons between algorithms (𝑁𝑖 𝑗 = 𝑊𝑖 𝑗 + 𝑊 𝑗𝑖 if there are no ties). Using natural logarithms of parameters 𝛽𝑖 = ln 𝑝𝑖 , the authors introduce the following model: 𝑒 𝛽𝑖 𝑊𝑖 𝑗 ∼ Binomial 𝑁𝑖 𝑗 , , 𝑒 𝛽𝑖 + 𝑒 𝛽 𝑗 (1) ¯ 𝛽𝑖 ∼ Normal(0, 𝜎), 𝜎¯ ∼ LogNormal(0, 0.5). The parameters of this model are estimated using Metropolis-Hastings. Metropolis-Hastings outputs a set of samples for the parameters ¯ and the final ranking is determined by the mean 𝛽𝑖 across 𝛽𝑖 , 𝜎, all samples. The samples from MCMC allow to build confidence intervals for the weights 𝑝𝑖 and probabilities 𝑃 (𝑖 ≻ 𝑗). Another approach called rank centrality or spectral method [13, 27] models the comparisons with a random walk on the underlying Erdős-Rényi graph with adjacency matrix 𝐴. It models the Markov chain with 𝑛 states, given the sample transition matrix 𝑃: ( 1 𝐴𝑖 𝑗 𝑤¯ 𝑗𝑖 , 𝑖 ≠ 𝑗, 𝑃𝑖 𝑗 ≡ 2𝑛𝑑 1 Í 1 − 2𝑛𝑑 𝑘:𝑘≠𝑖 𝐴𝑖𝑘 𝑤¯ 𝑘𝑖 , 𝑖 = 𝑗 . where 𝑤¯ 𝑖 𝑗 = 𝑊𝑖 𝑗 /(𝑊𝑖 𝑗 + 𝑊 𝑗𝑖 ) is the ratio of wins and 𝑑 is the edge density of the graph with adjacency matrix 𝐴. This transition matrix is modeled with ( 1 𝐴𝑖 𝑗 𝜎 (𝜃 ∗ − 𝜃 𝑖∗ ), 𝑖 ≠ 𝑗, ∗ 𝑃𝑖 𝑗 ≡ E(𝑃𝑖 𝑗 |𝐴) = 2𝑛𝑑 1 Í 𝑗 1 − 2𝑛𝑑 𝑘:𝑘≠𝑖 𝐴𝑖𝑘 𝜎 (𝜃 𝑘∗ − 𝜃 𝑖∗ ), 𝑖 = 𝑗, where 𝜎 is the sigmoid function and 𝜃 𝑖∗ are the parameters (BradleyTerry weights). This transition matrix admits the stationary measure ! ∗ ∗ 𝑒𝜃1 𝑒 𝜃𝑛 ∗ 𝜋 ≡ Í𝑛 𝜃 ∗ , . . . , Í𝑛 𝜃 ∗ , 𝑘 𝑘 𝑘=1 𝑒 𝑘=1 𝑒
Multiple comparisons
Plackett-Luce (PL) model [9, 28] is an extension of Bradley-Terry model to listwise rankings, that is, simultaneous comparisons of more than two objects. Plackett-Luce model takes 𝑅 rankings
Pr(𝑦𝑖 1 ≻ 𝑦𝑖 2 ≻ · · · ≻ 𝑦𝑖𝑇𝑖 ) =
𝑇𝑖 Ö
𝑝 𝑖𝑘 Í𝑇𝑖
𝑘=1
𝑗=𝑘
. 𝑝𝑖 𝑗
The authors of [4, Section 4] proposed to introduce for 𝑖 = 1, . . . , 𝑅 and 𝑗 = 1, . . . 𝑇𝑖 − 1 latent variables 𝑍 = {𝑧𝑖 𝑗 }: 𝑓 (𝑧|𝐷𝑎𝑡𝑎, 𝑝) =
𝑅 𝑇Ö 𝑖 −1 Ö
Exp(𝑧𝑖 𝑗 ;
𝑖=1 𝑗=1
𝑇𝑖 ∑︁
𝑝𝑖𝑘 ),
𝑘=𝑗
leading to loglikelihood 𝑅 𝑇∑︁ 𝑖 −1 ∑︁
𝑇𝑖
©∑︁ ª log 𝑝 𝑦𝑖 𝑗 − 𝑝 𝑦𝑖𝑘 ® 𝑧𝑖 𝑗 . 𝑖=1 𝑗=1 « 𝑘=𝑗 ¬ This loglikelihood can be maximized using EM algorithm. An alternative approach proposed by the authors of [4] is to sample from 𝑓 (𝑝, 𝑧|𝐷𝑎𝑡𝑎) using Gibbs sampler: ℓ (𝑝, 𝑧) =
𝑇𝑖
©∑︁ ª 𝑍𝑖(𝑡𝑗 ) | 𝐷𝑎𝑡𝑎, 𝜆 (𝑡 −1) ∼ E 𝜆𝜌(𝑡𝑖𝑘−1) ® , « 𝑘=𝑗 ¬ 𝜆𝑘(𝑡 ) | 𝐷𝑎𝑡𝑎, 𝑍 (𝑡 ) ∼ G
𝑎 + 𝑤𝑘 , 𝑏 +
𝑅 𝑇∑︁ 𝑖 −1 ∑︁
! 𝛿𝑖 𝑗𝑘 𝑍𝑖(𝑡𝑗 )
,
𝑖=1 𝑗=1
where 𝑤𝑘 is the number of rankings where the 𝑘-th individual is not in the last ranking position and 𝛿𝑖 𝑗𝑘 is the indicator of the event that individual 𝑘 receives a rank no better than 𝑗 in the 𝑖-th ranking. In our setup, we obtain comparisons (𝑦𝑖 1 ≻ 𝑦𝑖 2 ≻ · · · ≻ 𝑦𝑖𝑇𝑖 ) by sorting recommender algorithms on each dataset according to their values of a metric.
4
Validation of the rankings
A common metric for assessing the consistency of rankings is rank correlation, such as Kendall’s 𝜏 or Spearman’s 𝜌. To evaluate the validity of a given ranking, one may take the mean rank correlation between that ranking and the per-dataset rankings. However, when a metric value is missing for a particular algorithm on a given dataset, that algorithm must be excluded from both rankings before computing the correlation. Consequently, standard rank correlation measures do not capture the robustness of a ranking with respect
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
to missing data. To address both validity and robustness simultaneously, we propose a metric based on the number of transitive triplets in the ranking. The ranking 𝑖 𝑑1 ≻ 𝑖 𝑑2 ≻ 𝑖 𝑑3 is called transitive on dataset 𝑑 if 𝑖 𝑑1 wins 𝑖 𝑑2 , 𝑖 𝑑2 wins 𝑖 𝑑3 and 𝑖 𝑑1 wins 𝑖 𝑑3 . In an ideal ranking all triplets are transitive, i.e., all duels between any three algorithms are consistently ordered. Hence, we can use the ratio of transitive triplets across datasets as a quality metric. We count only distinct triplets, so the total number of ordered triplets is (3𝑛 ) · 𝐷, where 𝑛 is the number of algorithms, and 𝐷 is the number of datasets. If there are missing comparisons in data, then the total number of ordered triplets decreases.. A higher ratio indicates a better ranking:
Grishina et al.
5.2
Evaluation protocol
In addition, this metric can take ties into account. If the metrics achieved by two algorithms 𝑖 and 𝑗 are very close (intervals metrici ± stdi and metricj ± stdj overlap), we consider the comparison a tie. Taking ties into account, we say that the triplet is transitive if 𝑖 1 ⪰ 𝑖 2 ⪰ 𝑖 3 . For all ranking methods, we rely solely on the weights 𝑝𝑖 to rank algorithms and do not treat overlapping confidence intervals produced by statistical models as ties in a ranking.
We follow the experimental setup described in [16], specifically the global temporal split (GTS) protocol. Each dataset is divided into training, validation, and test periods based on global timestamps with a 90/5/5 split ratio. For datasets without timestamps, we assigned random timestamps to conform to this setup. For validation, we use the last-item strategy, selecting the last interaction during the validation period as the target and for testing, the random-item as a target item was used. For robustness, we average the metrics over 10 different random seeds. Using these results, we obtain metric intervals mean_metrici ± stdi , which we use to identify ties between algorithms. We used the Optuna [1] framework with tree- structured Parzen Estimator (TPE) sampler to perform a hyperparameter optimization. For all models, hyperparameters are selected by optimizing NDCG@10 on the validation set. We use a single search space for all datasets to ensure fair comparison and set a budget of 200 trials (with 20 startup trials) to balance search depth and cost. Due to the prohibitive computational cost of certain algorithms on larger datasets, the optimization process was occasionally terminated early due to time constraints. In such cases, we selected the best hyperparameters identified up to that point. After selecting the optimal hyperparameters, the model was retrained on the combined training and validation sets, and test metrics were collected.
5 Experimental Setup 5.1 Datasets and recommendation algorithms
6 Experiments 6.1 Comparison with aggregation baselines
To ensure a broad and representative evaluation, we selected a diverse collection of recommendation datasets covering different domains and structural properties. Most come from the APS benchmark set [43], which provides a collection of standard datasets popular in recent recommender systems studies. To further increase the diversity, we included a few additional datasets. Following common practice [21] in recommender system evaluation, all datasets were preprocessed using a 5-core filtering scheme. The final statistics of the datasets after preprocessing are reported in Table 10 in Appendix. For further analysis, we grouped datasets along four dimensions: density, user-item ratio, mean interactions per user, and sequentiality. The assessment of sequentiality follows the methodology proposed by Klenitskiy et al. [22] based on 2-gram statistics. In experiments comparing performance across these dimensions, we evaluated the models on the top-20 datasets with the highest and lowest values for each metric (e.g., the 20 most and the 20 least sequential). Datasets without timestamp information were considered non-sequential. We utilize a set of standard and widely adopted baselines to ensure a comprehensive evaluation. Our selection includes trivial non-personalized heuristics (Random, PopRandom) and three neighborhood-based algorithms (User-KNN, Item-KNN, Seq-KNN). We also examine five matrix factorization and linear models (ALS [20, 39], BPR [32], SGD MF [31], PureSVD [7], EASEr [37]). Finally, to represent modern developments in the field, we include two graph-based models (LightGCN [18], UltraGCN [24]) and two advanced sequential architectures: the transformer-based SASRec [21, 25] and the tensor-based GASATF [12]. Implementation details are described in Table 7 in Appendix.
In recommender systems and other domains, it is common to rank algorithms by the mean or sum of metrics across datasets [29, 34]. However, these aggregation methods might yield unstable or contradictory rankings, as the comparison of metrics is usually valid only for a particular problem. Approaches based on the BT model (described in Section 3) offer a different perspective by simulating a "competition" between algorithms. In this section, we validate the BT model, while in section 7 we assess its predictive power and compare with covariate-adjusted extensions.
𝐷
Triplet ratio =
𝑛
∑︁ ∑︁ 1 𝐼 {𝑖 𝑑1 ≻ 𝑖 𝑑2 ≻ 𝑖 𝑑3 }. 𝑛 (3 ) · 𝐷 − missing 𝑑 𝑑 𝑑 𝑑=1 𝑖 ≠𝑖 ≠𝑖 1
2
3
Quality of rankings. To compare the quality of different rankings, we consider the ratio of transitive triplets in a ranking and mean Kendall’s 𝜏 between given ranking and per-dataset rankings, as described in Section 4. A high-quality ranking is expected to be largely transitive: if algorithm 𝐴 outperforms 𝐵, and 𝐵 outperforms 𝐶, then 𝐴 should also outperform 𝐶. As shown in Table 1, BradleyTerry achieves a higher ratio of transitive triplets for all datasets and especially for subsets of datasets with long user history and sparse datasets. Notice that the PL model demonstrates results very similar to the ones of the BT model. We also observe, that the triplet ratio is consistent with Kendall’s 𝜏, and by both metrics BT model shows the best results (for more Kendall’s 𝜏 results see table 8 in Appendix). However, Kendall’s 𝜏 can not be used to assess stability of the rankings when there are many missing pairwise comparisons. Instead, we apply our proposed “triplet ratio“ metric as a proxy to compare the robustness of the rankings. Stability of rankings. Since in practice running all compared algorithms on all available datasets may be computationally expensive and time consuming, the results of some pairwise comparisons
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
Table 1: Ranking quality. Algorithms are compared on all datasets and subsets of sparse and long-history-datasets. Comparison is based on NDCG@10. All w/o ties w ties
Long history w/o ties w ties
Sparse w/o ties w ties
Ratio of transitive triplets. Mean Sum BT PL
0.488 0.484 0.510 0.503
0.613 0.613 0.631 0.631
0.433 0.445 0.471 0.450
0.556 0.569 0.588 0.576
0.495 0.458 0.505 0.486
0.615 0.575 0.624 0.614
0.542 0.499 0.565 0.547
0.542 0.499 0.558 0.547
Mean Kendall’s 𝜏. 0.542 0.541 0.567 0.558
Ratio of transitive triples
Mean Sum BT PL
0.542 0.541 0.564 0.558
0.488 0.5 0.533 0.502
0.488 0.5 0.524 0.502
Mean Sum Bradley-Terry Plackett-Luce
0.6 0.5 0.0
0.2 0.4 0.6 0.8 Ratio of missing comparisons
Figure 1: Ratio of transitive triples in rankings of different methods (by NDCG@10) depending on the ratio of missing comparisons in the data.
0.05
2
0.00
0
Plackett-Luce 0.2
S -KN LigeqhtG N CN GASAT SASRecF Easer UltraGBPR CN Item SGD-KNMN Pure SAVDF LS U se rK PopRan NN om Randdom
Bradley-Terry
S -KN LigeqhtG N S CN GAASSRATec F UltraEGaser CN Item-KBPR NN Pure SALS User-KVD GD NMNF PopRSan om Randdom
Sum
S -KN LigeqhtG N SASRCecN E G SasATer ItemA-K F UltraGNN SGD CMN ALSF Pure SBPR V PopURseanr-KNDN om Randdom
4
7.5 5.0 2.5 0.0
S E-KasNer LigeqhtG N S CN GASSRATec ItemA-K F UltraGNN CN SGDBMPR F Pure SALS V PopURseanr-KNDN om Randdom
weight
Mean 0.10
0.0
Gaps ratio 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9
Figure 2: Weights of rankings by NDCG@10 for data with varying ratio of missing values.
might be missing. Therefore, it is important to evaluate robustness of difference ranking methods to incomplete data. To assess stability, we randomly omit a certain fraction of entries from the table with algorithm’s metrics, and recompute the ranking based on the remaining data. Figure 1 shows that the ratio of transitive triples in the rankings by mean and sum of values rapidly decreases as the proportion of missing data increases. In contrast, the ratio of transitive triples remains almost constant for the BT and PL models. While the weights of BT and PL appear stable in Figure 2, the weights of the "sum" and "mean" rankings fluctuate with growing gap ratio. This experiment indicates that BT and PL rankings are substantially more robust to missing data.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
6.2
Comparison of statistical models
We evaluated three Bradley-Terry estimators from Section 3.2: the classic Zermelo algorithm (MLE), the Bayesian variant, and the Rank Centrality method. On our benchmark data, all three converge to identical weights and rankings. This convergence occurs because our win matrix 𝑊 is complete and all the models are aimed at maximizing similar loglikelihood. We further use the Bayesian Bradley-Terry model approach (see (1)) due to its principal advantage – the availability of confidence intervals for weights, which allows for uncertainty-aware comparisons. In contrast, the PlackettLuce model often produces different weights and rankings. Figure 3 visualizes this divergence, comparing the Bayesian BT weights (with and without tie handling) with PL weights. We observe that when the ties are taken into account the weights of consecutive algorithms in the ranking become closer and the quantile boxes overlap more, showing more uncertainty. However, the rankings of BT with and without ties are very similar, indicating that taking ties into account has minimal impact on our benchmark. Moreover, BT is highly concordant with PL for all and sequential datasets (subfigures (a) and (b)). Both methods clearly identify the same cluster of top-performing algorithms (SASRec, GASATF, etc.). Minor rank variations within this cluster fall within the statistical uncertainty of the estimates. The models disagree more noticeably on datasets with long user histories, where the weights are more smooth. For example, PL moves SASRec from 4th to 6th place and Item-KNN from 9th to 5th. This suggests that while both approaches are valid, when there is more uncertainty in the data, they can give different interpretations of algorithmic performance.
6.3
Ranking on different dataset classes
We analyze how various dataset characteristics influence the performance of recommender systems by using the Bradley-Terry model. Table 2 shows that model rankings are not static and change significantly depending on the nature of the data. We observe that no single algorithm is the best across all possible scenarios. The most dramatic changes occur in the Sequentiality category. For example, SASRec and GASATF are the clear leaders on sequential datasets, where they hold the first and second ranks. However, their performance drops sharply on non-sequential datasets, where they fall to the 10th and 11th positions. In these non-sequential environments, traditional models like ALS and UltraGCN show significant improvements, with ALS rising from the 11th rank to become the second-best model. The amount of available user history also plays a critical role in determining which model performs best. On datasets with long history, Seq-KNN and GASATF take the top spots. However, when the interaction history is short, LightGCN becomes the most effective model. In this short history scenario, GASATF drops to 5th place and UltraGCN drops to 8th. This suggests that while some models are very powerful when given a lot of data, they are less robust when information is limited. On sparse datasets, GASATF and User-KNN show improved relative performance compared to their rankings on dense datasets. Furthermore, UltraGCN and ALS improve their rankings on datasets with a low user-to-item ratio, suggesting that these methods are better suited to settings where the item catalog is large relative to the user base.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Grishina et al.
Table 2: NDCG@10-based BT rankings across opposing dataset characteristics. Significant changes (magnitude of at least 3) are highlighted: improvements in green ↑ and drops in red ↓.
Density Rank
Large
Small
30
Bayesian BT w ties
20 10 0
0.4 0.3 0.2 0.1 0.0
Plackett-Luce Gibbs
Bayesian BT w ties
20 10
GASAT SASReF c Seq-K LightGNN EasCeNr Pure Item-KSVD BNN SG PR UltraDGMF ACN Us LS PopRear-n KNN d Randoom m
0
0.4 0.3 0.2 0.1 0.0
Plackett-Luce Gibbs
GASAT SASReF c Seq-K NN LightG EasCeNr Item-K BNN Pure S PR SG VD UltraDGMF ACN Us LS PopRear-n KNN Randdoom m
30
SAS GASRATec F LightG Seq-KNCN Ea N Pure Sser Item-KVD BNN UltraG PR SGD CMN AF Us LS PopRear-n KNN d Randoom m
weight
Bayesian BT w/o ties
(b) Sequential datasets.
2
Seq-KN GASATN F LightG SA CN UltraSGRec EasCeNr A Pure LS Item-KSVD SGD NMN B F User-K PR PopRan NN Randdoom m
0
Bayesian BT w ties
Plackett-Luce Gibbs
4 3 2 1 0
0.2
Seq-KN N LightG GA CN UltraSGATF Item-KCN S AS NN SGDRMec EaseFr BPR A Pure LS User-KSVD PopRan NN Randdoom m
4
GASAT F Seq-K SASRNeN c LightG C N UltraG EasCeNr A Pure LS Item-KSVD SGD NMN B F User-K PR PopRan NN Randdoom m
Bayesian BT w/o ties weight
Long history Short history LightGCN Seq-KNN SASRec EASEr GASATF ↓ ALS BPR ↑ UltraGCN ↓ User-KNN ↑ Item-KNN SGD MF PureSVD ↓ PopRandom Random
0.1 0.0
(c) Top 20 datasets with long user history.
Figure 3: Comparison of weights of different models. Metric for rankings is NDCG@10.
The heatmaps in Figure 4 provide a visualization of the performance clusters through pairwise win probabilities. In Figure 4b a
Sequentiality Sequential
Non-Sequential
SASRec LightGCN GASATF ALS ↑ LightGCN Seq-KNN Seq-KNN EASEr EASEr UltraGCN ↑ PureSVD Item-KNN Item-KNN SGD MF ↑ BPR BPR UltraGCN User-KNN ↑ SGD MF SASRec ↓ ALS GASATF ↓ User-KNN PureSVD ↓ PopRandom PopRandom Random Random
Table 3: Comparison of rankings across all datasets based on NDCG@10. Superscripts indicate rank deviation from BT.
Algorithm
GASAT SASReF c Seq-K NN LightG EasCeNr Item-K BNN Pure S PR SG VD UltraDGMF ACN Us LS PopRear-n KNN Randdoom m
Bayesian BT w/o ties
(a) All datasets. 40 30 20 10 0
Mean Interaction per User
Seq-KNN Seq-KNN LightGCN LightGCN Seq-KNN LightGCN LightGCN SASRec Seq-KNN GASATF ALS GASATF ↑ Seq-KNN GASATF LightGCN UltraGCN SASRec GASATF SASRec SASRec EASEr UltraGCN EASEr EASEr UltraGCN SASRec User-KNN ↑ Item-KNN UltraGCN ↑ EASEr SGD MF EASEr PureSVD Item-KNN ALS Item-KNN Item-KNN BPR ALS ↑ PureSVD GASATF BPR UltraGCN PureSVD Item-KNN BPR PureSVD SGD MF SGD MF SGD MF PureSVD ALS ↓ ALS BPR ↓ BPR User-KNN SGD MF ↓ User-KNN User-KNN User-KNN PopRandom PopRandom PopRandom PopRandom PopRandom Random Random Random Random Random
GASAT SASReF c Seq-K LightGNN EasCeNr Pure Item-KSVD BNN SG PR UltraDGMF ACN Us LS PopRear-n KNN d Randoom m
40 30 20 10 0
Sparse
SAS GASRATec F LightG Seq-KNCN Ea N Pure Sser Item-KVD BNN UltraG PR SGD CMN AF Us LS PopRear-n KNN d Randoom m
weight
1 2 3 4 5 6 7 8 9 10 11 12 13 14
Dense
User-Item Ratio
Seq-KNN LightGCN SASRec GASATF EASEr BPR Item-KNN UltraGCN PureSVD ALS SGD MF User-KNN PopRandom Random
BT 1 2 3 4 5 6 7 8 9 10 11 12 13 14
Sum
Mean
PL
1 2 3 5+1 4-1 10+4 6-1 7-1 11+2 9-1 8-3 12 13 14
2+1
1 2 4+1 3-1 5 6 8+1 7-1 10+1 11+1 9-2 12 13 14
3+1 4+1 5+1 1-4 8+2 6-1 7-1 11+2 10 9-2 12 13 14
very dark red block in the top-right corner is clearly observable. This represents a cluster of models including SASRec, GASATF, and LightGCN that have an extremely high probability of beating almost any other algorithm in a sequential context. When we move to Figure 4c for non-sequential data, the cluster changes completely. The dark red areas shift toward LightGCN and ALS, while the previous sequential leaders lose their dominance. This visualization confirms that the competition between models is context-dependent.
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
SASRec GASATF LightGCN Seq-KNN EASEr PureSVD Item-KNN BPR UltraGCN SGD MF ALS User-KNN PopRandom Random
(a) All
1.00
LightGCN ALS Seq-KNN EASEr UltraGCN Item-KNN SGD MF BPR User-KNN SASRec GASATF PureSVD PopRandom Random
Win Probability
0.75
0.50
Seq
tGC ALNS -KN EAS N Ultr Er aG Item CN -K SG NN DM BP F Use R r-KN SAS N R GA ec S Pur ATF Pop eSVD Ran Random dom
0.25
Ligh
SAS R GA ec S Ligh ATF tGC Seq N -KN EAS N Pur Er eSV Item D -KN BPN Ultr R aGC SG N DM AL F Use S Pop r-KNN Ran Random dom
Seq -K Ligh NN tG SAS CN R GA ec SA EASTF Er BP Item R -KN Ultr N aG Pur CN eSV ALDS SG D Use MF Pop r-KNN Ran Random dom
Seq-KNN LightGCN SASRec GASATF EASEr BPR Item-KNN UltraGCN PureSVD ALS SGD MF User-KNN PopRandom Random
(b) Sequential
0.00
(c) Non-sequential
Figure 4: Pairwise performance comparison of recommender models. Each cell (𝑖, 𝑗) shows the probability that the model in row 𝑖 outperforms model in column 𝑗 across all datasets, sequential datasets, and non-sequential datasets. Bayesian BT w/o ties
2+1 4+2 5+2 1-3 3-2 7+1 6-1 9+1 8-1 11+1 12+1 10-2 13 14
1 3+1 4+1 5+1 2-3 8+2 6-1 7-1 9 11+1 10-1 12 13 14
weight
0.2
10
0.1 0.0
0
Bayesian BT w ties
Plackett-Luce Gibbs
Bayesian BT w/o ties 20
30
15
20
10
10
5
0
0
0.4 0.2 0.0
Figure 5: Comparison of weights of different models on sequential datasets. The metric for ranking is NDCG@10. Top row - before shuffling timestamps, bottom row - after. Sequential models are highlighted in green.
6.4
Finally, we compare the BT rankings with aggregation methods (Sum and Mean) and discuss why a more sophisticated model is necessary. As shown in Tables 3 and 4, simple aggregation methods often produce different results because they do not account for the strength of the opponents. In Table 3, the Sum method ranks BPR four positions lower than the BT model. In Table 4, both the Sum and Mean methods rank EASEr three positions higher. These rank deviations happen because simple methods treat every win in a dataset as equal. The BT model provides a more accurate hierarchy by considering what opponent an algorithm beats and how strong that opponent is. This makes the BT framework a much more reliable tool for understanding the true strengths of different recommendation algorithms across diverse data environments.
0.3
GASA SASRTF ec Seq-K LightGNN EASCEN Item-K r BNN PureS PR SGD VD Ultra MF GCN User- ALS PopR KNN Raannddom om
2+1 4+2 5+2 1-3 3-2 7+1 6-1 9+1 8-1 11+1 12+1 10-2 13 14
Plackett-Luce Gibbs
LightG EA CN Seq-KSEr GASANN B TF Item-K PR SG NN UltraD MF PureGSCN SASRVD ec User- ALS PopR KNN Raannddom om
1 2 3 4 5 6 7 8 9 10 11 12 13 14
0.4
GASA SASRTF ec Se -K NN LighqtG EASCEN r Pu Itemre-KSVD BNN Ultra PR SGDGCN MF User- ALS PopR KNN a Rannddom om
LightGCN ALS Seq-KNN EASEr UltraGCN Item-KNN SGD MF BPR User-KNN SASRec GASATF PureSVD PopRandom Random
Bayesian BT w ties
LightG EA CN GASASEr TF Seq-K PureSNN BVPDR Item-K SA NN UltraSRec SGDGCN MF Us ALS PopRer-KNN a Rannddom om
PL
30 20
SAS GASRAec TF LightG Seq-KCN EASNEN r Pu Itemre-KSVD BNN Ultra PR SGDGCN MF Us ALS PopRer-KNN a n Randdom om
Mean
50 40 30 20 10 0
LightG EA CN Seq-KSEr GASANN PureS TF SASRVD B ec Item-K PR Ultra NN SGDGCN MF Us ALS PopRer-KNN a n Randdom om
Sum
weight
BT
After Shuffle
Algorithm
Before Shuffle
Table 4: Rankings on non-sequential datasets based on NDCG@10. Superscripts indicate rank deviation from BT.
Separation strength
We check the separation strength of the ranking by providing an ablation study on sequential datasets. We shuffle timestamps in the datasets and compare the ranking before and after shuffling. Contrary to the setup in [22], we apply timestamp shuffling to the training sequences and retrain the models. Figure 5 presents the results. As expected, sequential models dominate sequential datasets before timestamp shuffling, but lose this advantage and become less separated afterward. The ranking among sequential models also changes: SASRec is among the top-2 sequential models for all rankings before shuffling, but after shuffling, Seq-KNN and GASATF become ranked higher than SASRec. These models are more stable under different setups, while the rank of SASRec substantially degrades after shuffling. The weights of BT model decrease after shuffling, since the obvious leaders – sequential models – lose their advantage from the sequential data structure.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
7
Bradley-Terry model with covariates
A huge disadvantage of a simple BT model is its inability to model different rankings on different datasets and, hence, predict a specific ranking on a previously unseen dataset (which may drastically differ from the global BT ranking obtained on the training datasets). This can be amended by introducing dataset characteristics as covariates of the pairwise comparisons and exploit a covariate-adjusted BT model. We experiment with different covariate-based approaches (described in Sections 7.1, 7.2) in order to predict algorithms ranking on new datasets and evaluate their prediction accuracy in Section 7.3, providing the guidelines for strong baseline selection for the practitioners.
7.1
use numeric characteristics as numeric covariates and sequentiality as a categorical one, following the algorithm of [46]. We tried two variants of model: first with an enforced root split by sequentiality (acknowledged as an important characteristic, see e.g. [35]), and then without any constraints. Figure 6 shows the resulting trees: left with forced sequentiality split, right without adjustments. We observe that for sequential datasets rankings remain largely stable, whereas for non-sequential ones they depend mostly on the number of users and the number of interactions. Finally, we can use the constructed trees to predict ranking on a new dataset by simply traversing through them with its characteristics with zero computational cost.
Covariate-adjusted Bradley-Terry model
In the standard BT model 𝑃 (𝑖 ≻ 𝑗) can be rewritten as 𝑃 (𝑖 ≻ 𝑗) = 𝜎 (𝜃 𝑖 − 𝜃 𝑗 ), where 1 ≤ 𝑖, 𝑗 ≤ 𝑛 are players and 𝜃 𝑖 , 𝜃 𝑗 ∈ R are their corresponding strengths, 𝜎 (𝑧) = 1+𝑒1−𝑧 . Now suppose that for each pairwise comparison exists a covariate vector 𝑥 ∈ R𝑑 , influencing players’ strengths. Covariate-adjusted BT models (and other pairwise Learning-to-Rank models) typically integrate it via a linear combination, what renders a huge space for various approaches and remains convenient for theoretical analysis and interpretation (see [5, 11, 33] ). Namely, player 𝑖 is now characterized by a strength parameter 𝛽𝑖0 ∈ R and covariate weights 𝛽𝑖 = (𝛽𝑖1, . . . , 𝛽𝑖𝑑 ) ⊤ ∈ R𝑑 , quantifying different covariates’ influence on 𝑖’s advantage. Thus, we can define 𝑃 (𝑖 ≻ 𝑗 | 𝑥) = 𝜎 (𝛽𝑖0 − 𝛽 𝑗0 + ⟨𝑥, 𝛽𝑖 ⟩ − ⟨𝑥, 𝛽 𝑗 ⟩). This model has B = ∪𝑖 {𝛽𝑖0, 𝛽𝑖 }, |B| = 𝑛(𝑑 + 1) parameters and requires training data 𝑅 = {(𝑖𝑘 , 𝑗𝑘 , 𝑦𝑘 , 𝑥𝑘 )}𝑘 , where 1 ≤ 𝑖𝑘 , 𝑗𝑘 ≤ 𝑛 are players, 𝑦𝑘 ∈ {0, 1} is a player 𝑖𝑘 win indicator and 𝑥𝑘 ∈ R𝑑 is a covariate vector. It can be fit by maximizing log-likelihood ℓ (B) = Í 𝑘 (𝑦𝑘 log 𝑃 (𝑖𝑘 ≻ 𝑗𝑘 | 𝑥𝑘 ) + (1 − 𝑦𝑘 ) log(1 − 𝑃 (𝑖𝑘 ≻ 𝑗𝑘 | 𝑥𝑘 )) via any appropriate optimization method under identifiability constraints Í ∀𝑗 𝑛𝑖=1 𝛽𝑖 𝑗 = 0. However, this simple model is prone to overfitting to the global BT ranking with static and clearly separable strengths and little to no covariate effect. This can be amended by introducing Í √︁ L1 penalty following [33] as 𝐿(B) = 𝜆0 𝑖< 𝑗 (𝛽𝑖0 − 𝛽 𝑗0 ) 2 + 𝜀 + Í Í √︁ 𝜆1 𝑑𝑘=1 𝑖< 𝑗 (𝛽𝑖𝑘 − 𝛽 𝑗𝑘 ) 2 + 𝜀, where 𝜀 > 0, the first term regulates the closeness of overall strengths and the second term – the closeness of covariate effects weights. The final target function is ℓ (B) − 𝐿(B) → max. We optimize the penalized objecÍ tive with L-BFGS-B in the subspace ∀𝑗 𝑛𝑖=1 𝛽𝑖 𝑗 = 0 by assuming Í𝑛−1 𝛽𝑛 𝑗 = − 𝑖=1 𝛽𝑖 𝑗 . In our setup, we consider log-normalized numeric and binary categorical dataset characteristics as covariates and predict ranking on a new dataset with covariates 𝑥 by sorting algorithms’ strengths on this particular dataset, precisely 𝛽𝑖0 + ⟨𝑥, 𝛽𝑖 ⟩.
7.2
Grishina et al.
Bradley-Terry trees
A non-parametric niche alternative to the parametric approach presented above is the recursive covariate-dependent partitioning via Bradley-Terry trees. It can identify groups of datasets with distinct algorithm rankings in a data-driven way [46], thus yielding interpretable schemes that can predict algorithms rankings based on dataset characteristics. In order to build the trees for our data, we
Figure 6: BT trees for datasets. Each node displays the split covariate and its 𝑝-value, with split conditions on the edges; leaf sample sizes are denoted by 𝑛.
7.3
Ranking on unseen dataset
In order to evaluate different models’ ability to select strong baselines for different previously unobserved datasets, we use the following pipeline. We randomly choose 10 holdout datasets, fit models on the remaining training datasets, and predict rankings for each holdout dataset accordingly. Evaluation metrics are: (i) hit rate of the top-1 predicted algorithm being among the top-1/2/3 ground-truth algorithms on the holdout dataset (denoted as top-1/2/3 hits), and (ii) the overlap of top-2/3/5 predicted algorithms and the top-2/3/5 ground-truth algorithms correspondingly (denoted as top-2/3/5 overlap). Results are averaged over 5 runs. We compare trivial mean aggregation (obtaining the same ranking for all holdout datasets by averaging the metric on training datasets, denoted as Mean), simple Bradley-Terry model (obtaining the same ranking for all holdout datasets by fitting a BT model on training datasets, denoted as BT), Bradley-Terry trees (as described in Section 7.2, denoted as BT tree) and covariate-adjusted BradleyTerry model with fusion regularization (as described in Section 7.1, denoted as Cov. BT). As shown in Table 5, with sufficient training data (Train (79) / Holdout (10)) all of BT variations consistently outperform the trivial Mean aggregation and demonstrate a decent performance overall, predicting top-1 algorithm with 0.24-0.28 accuracy and approximately 3.3 out of top-5 algorithms on average. It is also important to note that with the decreasing amount of data for fitting a simple BT model (i.e., less training and more holdout datasets), its ability to accurately predict top-1 algorithm decreases in comparison with
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
covariate models, while the ability to predict strong baselines remains on par with others. Such results show that covariate-adjusted models are more accurate, especially with fewer comparisons, but for our purposes BT model is a strong and reliable method. Moreover, this highlights the importance of using the large and diverse set of datasets we have gathered. Table 5: Prediction accuracy metrics for holdout sizes of 10 and 30 datasets. Model Metric
Mean
BT
Cov. BT
BT tree
Train (79) / Holdout (10) top-1 hits top-2 hits top-3 hits top-2 overlap top-3 overlap top-5 overlap
0.02 0.12 0.16 0.62 1.38 3.32
0.24 0.58 0.78 0.86 1.64 3.32
0.28 0.62 0.78 0.96 1.72 3.28
0.26 0.50 0.64 0.86 1.64 2.88
Train (59) / Holdout (30) top-1 hits top-2 hits top-3 hits top-2 overlap top-3 overlap top-5 overlap
0.08 0.17 0.25 0.63 1.41 3.40
0.15 0.47 0.63 0.87 1.64 3.40
0.24 0.61 0.71 0.97 1.64 3.28
0.13 0.40 0.56 0.77 1.42 3.00
Finally, we observe the following phenomenon: although covariateadjusted Bradley-Terry models provide an overall more accurate ranking, simple Bradley-Terry ranking (providing the same top-5 algorithms for all new datasets) is actually sufficient for strong baseline prediction. To show this, we ran our pipeline calculating average Kendall correlation between rankings obtained by different BT models and ground-truth ranking by the metic value on the dataset as well as MAP@5, NDCG@5 and top-5 overlap (as metrics reflecting specifically the ability to obtain strong baselines). The results presented in Table 6 confirm our hypothesis. Table 6: Ranking accuracy metrics. Metric
Mean
BT
Cov. BT
BT tree
Kendall’s 𝜏 MAP@5 NDCG@5 top-5 overlap
0.438 0.689 0.575 3.10
0.489 0.872 0.709 3.32
0.572 0.873 0.732 3.42
0.501 0.869 0.693 3.18
Overall, these results show that in order to obtain an adequate set of strong baselines for a new dataset, the global BT ranking from Table 3 can be used. However, if a more accurate ranking prediction is needed, the covariate-augmented BT model is preferable.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
8
Discussion
A major limitation of the Bradley-Terry model is that it compares algorithms based on the number of wins, but does not take the possibility of insignificant difference between algorithm performance (i.e. a “tie”) into account. This can be done via simply introducing the parameter regulating the probability of a tie (see [8]), as well as through more complex methods taking into account the magnitude of the strengths (see [15]). The considerate shortcomings of pointwise estimates have also been addressed in the literature [17] with a true-pairwise learning-to-rank algorithm, training a bivariate MLP and deploying BT on pairwise scores. Our work has a significant potential in a variety of applications in both research and practice. We provide an open repository with 14 RecSys algorithms implementations and ready-to-use framework for hyperparameter optimization, evaluation on 89 preprocessed datasets, results aggregation and analysis via Bradley-Terry model. Our work can also serve as a practical guide for selecting strong algorithms for a given dataset based on its features.
9
Conclusion
The paper proposes a methodology for ranking recommendation algorithms based on the Bradley–Terry model. The method is supplemented by the transitive triplets metric for assessing ranking consistency and robustness to missing data. Experiments reveal that rankings vary significantly with dataset characteristics such as sequentiality and sparsity. Building on this observation, BT trees and covariate-adjusted BT enable dataset-aware baseline selection by predicting algorithm rankings for a target dataset from its characteristics, without running the candidate models. Our approach provides a robust and interpretable comparison that is substantially more stable than simple metric aggregation methods, even in the presence of missing data. The resulting open benchmark and analysis tools provide a practical foundation for reproducible and more reliable algorithm comparison in recommender systems.
Acknowledgments The work was supported by the grant for research centers in the field of AI provided by the Ministry of Economic Development of the Russian Federation in accordance with the agreement 000000C313925P4E0002 and the agreement with HSE University № 139-15-2025-009. This research was supported in part through computational resources of HPC facilities at HSE University [23].
References [1] Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorage, AK, USA) (KDD ’19). Association for Computing Machinery, New York, NY, USA, 2623–2631. doi:10.1145/3292500.3330701 [2] Alessio Benavoli, Giorgio Corani, Janez Demšar, and Marco Zaffalon. 2017. Time for a Change: a Tutorial for Comparing Multiple Classifiers Through Bayesian Analysis. Journal of Machine Learning Research 18, 77 (2017), 1–36. http://jmlr. org/papers/v18/16-305.html [3] Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika 39, 3/4 (1952), 324–345. [4] Francois Caron and Arnaud Doucet. 2012. Efficient Bayesian inference for generalized Bradley–Terry models. Journal of Computational and Graphical Statistics 21, 1 (2012), 174–196.
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
[5] Giuseppe Casalicchio, Gerhard Tutz, and Gunther Schauberger. 2015. Subjectspecific Bradley–Terry–Luce models with implicit variable selection. Statistical Modelling 15, 6 (2015), 526–547. [6] Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. 2024. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132 (2024). [7] Paolo Cremonesi, Yehuda Koren, and Roberto Turrin. 2010. Performance of recommender algorithms on top-n recommendation tasks. In Proceedings of the Fourth ACM Conference on Recommender Systems (Barcelona, Spain) (RecSys ’10). Association for Computing Machinery, New York, NY, USA, 39–46. doi:10.1145/ 1864708.1864721 [8] Roger R Davidson. 1970. On extending the Bradley-Terry model to accommodate ties in paired comparison experiments. J. Amer. Statist. Assoc. 65, 329 (1970), 317–328. [9] Gerard Debreu. 1960. Individual choice behavior: A theoretical analysis. [10] Janez Demšar. 2006. Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research 7, 1 (2006), 1–30. http://jmlr.org/ papers/v7/demsar06a.html [11] Jianqing Fan, Jikai Hou, and Mengxin Yu. 2024. Uncertainty quantification of MLE for entity ranking with covariates. Journal of Machine Learning Research 25, 358 (2024), 1–83. [12] Evgeny Frolov and Ivan Oseledets. 2023. Tensor-Based Sequential Learning via Hankel Matrix Representation for Next Item Recommendations. IEEE Access 11 (2023), 6357–6371. doi:10.1109/ACCESS.2023.3234863 [13] Chao Gao, Yandi Shen, and Anderson Y Zhang. 2023. Uncertainty quantification in the Bradley–Terry–Luce model. Information and Inference: A Journal of the IMA 12, 2 (2023), 1073–1140. [14] Pieter Gijsbers, Marcos LP Bueno, Stefan Coors, Erin LeDell, Sébastien Poirier, Janek Thomas, Bernd Bischl, and Joaquin Vanschoren. 2024. Amlb: an automl benchmark. Journal of Machine Learning Research 25, 101 (2024), 1–65. [15] Mark E Glickman. [n. d.]. Paired comparison models with strength-dependent ties and order effects. Statistical Modelling ([n. d.]), 1471082X251400474. [16] Danil Gusak, Anna Volodkevich, Anton Klenitskiy, Alexey Vasilev, and Evgeny Frolov. 2025. Time to Split: Exploring Data Splitting Strategies for Offline Evaluation of Sequential Recommenders. In Proceedings of the Nineteenth ACM Conference on Recommender Systems (RecSys ’25). Association for Computing Machinery, New York, NY, USA, 874–883. doi:10.1145/3705328.3748164 [17] Malay Haldar, Daochen Zha, Huiji Gao, Liwei He, and Sanjeev Katariya. 2025. Beyond Pairwise Learning-To-Rank At Airbnb. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (Seoul, Republic of Korea) (CIKM ’25). ACM, New York, NY, USA, 8 pages. doi:10.1145/ 3746252.3761521 [18] Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, YongDong Zhang, and Meng Wang. 2020. LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation (SIGIR ’20). Association for Computing Machinery, New York, NY, USA, 639–648. doi:10.1145/3397271.3401063 [19] Reinhard Heckel, Max Simchowitz, Kannan Ramchandran, and Martin Wainwright. 2018. Approximate Ranking from Pairwise Comparisons. In Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 84), Amos Storkey and Fernando Perez-Cruz (Eds.). PMLR, 1057–1066. https://proceedings.mlr.press/ v84/heckel18a.html [20] Yifan Hu, Yehuda Koren, and Chris Volinsky. 2008. Collaborative filtering for implicit feedback datasets. In 2008 Eighth IEEE international conference on data mining. Ieee, 263–272. [21] Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Recommendation. In Proceedings of the IEEE International Conference on Data Mining (ICDM). IEEE, 197–206. [22] Anton Klenitskiy, Anna Volodkevich, Anton Pembek, and Alexey Vasilev. 2024. Does it look sequential? an analysis of datasets for evaluation of sequential recommendations. In Proceedings of the 18th ACM Conference on Recommender Systems. 1067–1072. [23] PS Kostenetskiy, RA Chulkevich, and VI Kozyrev. 2021. HPC resources of the higher school of economics. In Journal of Physics: Conference Series, Vol. 1740. IOP Publishing, 012050. [24] Kelong Mao, Jieming Zhu, Xi Xiao, Biao Lu, Zhaowei Wang, and Xiuqiang He. 2021. UltraGCN: ultra simplification of graph convolutional networks for recommendation. In Proceedings of the 30th ACM international conference on information & knowledge management. 1253–1262. [25] Gleb Mezentsev, Danil Gusak, Ivan V. Oseledets, and Evgeny Frolov. 2024. Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item Catalogs. In Proceedings of the 18th ACM Conference on Recommender Systems (RecSys). ACM, 475–485. [26] Frederick Mosteller. 1951. Remarks on the Method of Paired Comparisons: I. The Least Squares Solution Assuming Equal Standard Deviations and Equal Correlations. Psychometrika 16, 1 (1951), 3–9.
Grishina et al.
[27] Sewoong Oh. 2017. Rank centrality: Ranking from pairwise comparisons. Operations research (2017). [28] Robin L Plackett. 1975. The analysis of permutations. Journal of the Royal Statistical Society Series C: Applied Statistics 24, 2 (1975), 193–202. [29] Wang Qinsi, Jinghan Ke, Masayoshi Tomizuka, Kurt Keutzer, and Chenfeng Xu. 2025. Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives. In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=kws76i5XB8 [30] P. V. Rao and Lawrence L. Kupper. 1967. Ties in Paired-Comparison Experiments: A Generalization of the Bradley–Terry Model. J. Amer. Statist. Assoc. 62, 317 (1967), 194–204. [31] Steffen Rendle and Christoph Freudenthaler. 2014. Improving pairwise learning for item recommendation from implicit feedback. In Proceedings of the 7th ACM International Conference on Web Search and Data Mining (New York, New York, USA) (WSDM ’14). Association for Computing Machinery, New York, NY, USA, 273–282. doi:10.1145/2556195.2556248 [32] Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian Personalized Ranking from Implicit Feedback. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence (UAI). Montreal, QC, Canada, 452–461. [33] Gunther Schauberger and Gerhard Tutz. 2019. BTLLasso: a common framework and software package for the inclusion and selection of covariates in BradleyTerry models. Journal of Statistical Software 88 (2019), 1–29. [34] Mete Sertkan, Sophia Althammer, Sebastian Hofstätter, Peter Knees, and Julia Neidhardt. 2023. Exploring Effect-Size-Based Meta-Analysis for Multi-Dataset Evaluation. In Proceedings of the 3rd Workshop Perspectives on the Evaluation of Recommender Systems 2023 co-located with the 17th ACM Conference on Recommender Systems (RecSys 2023) (CEUR Workshop Proceedings, Vol. 3476). CEUR-WS.org. https://ceur-ws.org/Vol-3476/paper2.pdf [35] Valeriy Shevchenko, Nikita Belousov, Alexey Vasilev, Vladimir Zholobov, Artyom Sosedka, Natalia Semenova, Anna Volodkevich, Andrey Savchenko, and Alexey Zaytsev. 2024. From Variability to Stability: Advancing RecSys Benchmarking Practices. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Barcelona, Spain) (KDD ’24). Association for Computing Machinery, New York, NY, USA, 5701–5712. doi:10.1145/3637528.3671655 [36] Vladimir Spokoiny. 2025. Semiparametric plug-in estimation, sup-norm risk bounds, marginal optimization, and inference in BTL model. arXiv preprint arXiv:2503.15045 (2025). [37] Harald Steck. 2019. Embarrassingly Shallow Autoencoders for Sparse Data. In The World Wide Web Conference (San Francisco, CA, USA) (WWW ’19). Association for Computing Machinery, New York, NY, USA, 3251–3257. doi:10.1145/3308558. 3313710 [38] Carolin Strobl, Florian Wickelmaier, and Achim Zeileis. 2011. Accounting for individual differences in Bradley-Terry models by means of recursive partitioning. Journal of Educational and Behavioral Statistics 36, 2 (2011), 135–153. [39] Gábor Takács, István Pilászy, and Domonkos Tikk. 2011. Applications of the conjugate gradient method for implicit feedback collaborative filtering. In Proceedings of the fifth ACM conference on Recommender systems. 297–300. [40] Louis L. Thurstone. 1927. A Law of Comparative Judgment. Psychological Review 34 (1927), 273–286. [41] Askar Tsyganov, Evgeny Frolov, Sergey Samsonov, and Maxim Rakhuba. 2026. Matrix-Free Two-to-Infinity and One-to-Two Norms Estimation. Proceedings of the AAAI Conference on Artificial Intelligence 40, 31 (2026), 26010–26018. [42] Askar Tsyganov, Uliana Parkina, Ekaterina Grishina, Sergey Samsonov, and Maxim Rakhuba. 2026. Faster SVD via Accelerated Newton-Schulz Iteration. In ICLR Blogposts 2026 (April 27, 2026). https://iclr-blogposts.github.io/2026/blog/ 2026/polar-svd/ [43] Tobias Vente, Michael Heep, Abdullah Abbas, Theodor Sperle, Joeran Beel, and Bart Goethals. 2025. APS Explorer: Navigating Algorithm Performance Spaces for Informed Dataset Selection. In Proceedings of the Nineteenth ACM Conference on Recommender Systems. 1322–1324. [44] Jacques Wainer. 2023. A Bayesian Bradley–Terry Model to Compare Multiple Machine Learning Algorithms on Multiple Data Sets. Journal of Machine Learning Research 24 (10 2023), 1–34. [45] Fabian Wauthier, Michael I. Jordan, and Nebojsa Jojic. 2013. Efficient Ranking from Pairwise Comparisons. In Proceedings of the 30th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 28). 109–117. [46] Achim Zeileis, Torsten Hothorn, and Kurt Hornik. 2008. Model-based Recursive Partitioning. Journal of Computational and Graphical Statistics 17, 2 (2008), 492–514. doi:10.1198/10618600SX319331 [47] A Zeileis, C Strobl, F Wickelmaier, and J Kopf. 2011. psychotree: Recursive partitioning based on psychometric models. R package version 0.12-1, URL http://CRAN. R-project. org/package= psychotree (2011). [48] Ernst Zermelo. 1929. Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung. Mathematische Zeitschrift 29, 1 (1929), 436–460.
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
A
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Supplementary tables Table 7: Recommendation algorithms used in the study
Algorithm
Description
Implementation
Random PopRandom User-KNN Item-KNN Seq-KNN SGD MF BPR ALS PureSVD EASEr GASATF LightGCN UltraGCN SASRec
Recommends items by sampling uniformly at random Recommends items by sampling in proportion to their frequency Recommends items from similar users Recommends items similar to the user’s history Recommends items from similar sequential user histories One batch SGD with adaptive negative sampling [31] and folding-in procedure Based on [32] with uniform negative sampling and folding-in procedure Based on [20, 39] Based on [7], using SVD routine from [42] Based on [37] Based on [12] Original LightGCN [18] model with BPR objective Original UltraGCN [24] model with folding-in procedure. We also tried regularization as in [41] SASRec [21] with Scalable Cross-Entropy [25] and alpha-beta reparametrization of bucket sizes
self implemented self implemented self implemented self implemented self implemented self implemented self-implemented “Implicit” package self implemented self-implemented self implemented Pytorch-geometric self implemented. from official repository
Table 8: Kendall’s 𝜏 between BT rankings and baseline aggregation methods. The highest correlation for each group is bold. NDCG@10
HitRate@10
Coverage@10
Dataset group
Mean
Sum
PL
Mean
Sum
PL
Mean
Sum
PL
All Sequential Non-sequential Dense Sparse Mostly users Mostly items Long history Short history
0.802 0.890 0.802 0.824 0.692 0.692 0.780 0.736 0.626
0.824 0.912 0.802 0.824 0.736 0.714 0.824 0.780 0.648
0.912 0.890 0.868 0.802 0.912 0.912 0.758 0.758 0.890
0.824 0.846 0.868 0.824 0.736 0.692 0.758 0.758 0.648
0.846 0.868 0.868 0.824 0.758 0.692 0.802 0.802 0.670
0.868 0.912 0.912 0.824 0.956 0.912 0.868 0.758 0.846
0.934 0.912 0.846 0.846 0.978 0.934 0.890 0.802 0.912
0.890 0.934 0.846 0.846 0.758 0.846 0.846 0.824 0.890
0.956 0.978 0.868 0.890 0.956 0.912 0.890 0.934 0.956
Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009
Table 9: Recall@10-based BT rankings across opposing datasets. Significant improvements are green ↑, drops are red ↓. Density
User-Item Ratio
Mean Interaction per User
Sequentiality
Rank
Dense
Sparse
Large
Small
Long history
Short history
Sequential
Non-Sequential
1 2 3 4 5 6 7 8 9 10 11 12 13 14
Seq-KNN LightGCN ALS UltraGCN Easer SGD MF SASRec GASATF Item-KNN BPR User-KNN Pure SVD PopRandom Random
Seq-KNN LightGCN SASRec ↑ GASATF ↑ Easer UltraGCN BPR Item-KNN User-KNN Pure SVD ALS ↓ SGD MF ↓ PopRandom Random
LightGCN Seq-KNN SASRec GASATF Easer Item-KNN BPR Pure SVD UltraGCN SGD MF ALS User-KNN PopRandom Random
LightGCN GASATF Seq-KNN SASRec UltraGCN ↑ Easer Item-KNN ALS SGD MF BPR ↓ Pure SVD User-KNN PopRandom Random
Seq-KNN GASATF LightGCN SASRec UltraGCN ALS Easer SGD MF Item-KNN Pure SVD BPR User-KNN PopRandom Random
Seq-KNN LightGCN SASRec GASATF Easer BPR ↑ ALS UltraGCN User-KNN SGD MF Item-KNN Pure SVD PopRandom Random
GASATF SASRec LightGCN Seq-KNN Easer Pure SVD BPR Item-KNN SGD MF UltraGCN ALS User-KNN PopRandom Random
LightGCN ALS ↑ UltraGCN ↑ Easer Seq-KNN SGD MF ↑ Item-KNN BPR SASRec ↓ User-KNN GASATF ↓ Pure SVD ↓ PopRandom Random
KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea
Grishina et al.
Table 10: Left: statistics of the datasets used in the study. Right: number of missing trials for the specified algorithm.
Dataset Amazon2014-Instant-Video Amazon2014-Apps-For-Android Amazon2014-Automotive Amazon2014-Baby Amazon2014-Beauty Amazon2014-Books Amazon2014-CDs-Vinyl Amazon2014-Cell-Phones-And-Accessories Amazon2014-Clothing-Shoes-And-Jewelry Amazon2014-Digital-Music Amazon2014-Electronics Amazon2014-Grocery-And-Gourmet Amazon2014-Health-And-Personal-Care Amazon2014-Home-And-Kitchen Amazon2014-Kindle-Store Amazon2014-Movies-And-TV Amazon2014-Musical-Instruments Amazon2014-Office-Products Amazon2014-Patio-Lawn-And-Garden Amazon2014-Pet-Supplies Amazon2014-Sports-And-Outdoors Amazon2014-Tools-And-Home-Improvement Amazon2014-Toys-And-Games Amazon2014-Video-Games Amazon2018-Arts-Crafts-And-Sewing Amazon2018-Automotive Amazon2018-CDs-And-Vinyl Amazon2018-Cell-Phones-And-Accessories Amazon2018-Digital-Music Amazon2018-Electronics Amazon2018-Gift-Cards Amazon2018-Grocery-And-Gourmet-Food Amazon2018-Home-And-Kitchen Amazon2018-Industrial-And-Scientific Amazon2018-Kindle-Store Amazon2018-Luxury-Beauty Amazon2018-Magazine-Subscriptions Amazon2018-Movies-And-TV Amazon2018-Musical-Instruments Amazon2018-Office-Products Amazon2018-Patio-Lawn-And-Garden Amazon2018-Pet-Supplies Amazon2018-Prime-Pantry Amazon2018-Software Amazon2018-Sports-And-Outdoors Amazon2018-Tools-And-Home-Improvement Amazon2018-Toys-And-Games Amazon2018-Video-Games Anime-Recommendations-Database Behance BookCrossing CiaoDVD CiteULike-A CiteULike-T DeliveryHero-SE FilmTrust FoodComRecipes Foursquare-NYC1 Foursquare-NYC2 Foursquare-Tokyo Frappe GoogleLocal2018 GoogleLocal2021-Alaska GoogleLocal2021-Delaware GoogleLocal2021-District-Of-Columbia GoogleLocal2021-New-Hampshire GoogleLocal2021-North-Dakota GoogleLocal2021-Rhode-Island GoogleLocal2021-South-Dakota GoogleLocal2021-Vermont GoogleLocal2021-Wyoming Hetrec-LastFM Jester-4 KGRec-Music LearnFromSets Librarything MarketBias-ModCloth ModCloth-Clothing-Fit MovieLens-100K MovieLens-1M MovieLens-Latest-Small MovieTweetings Myket-Android Personality RentTheRunway Retailrocket Steam-Australian-Reviews WikiLens Yoochoose
Mean Mean Number Number Number of User-Item Density Interaction Interaction BPR EASEr GASATF LightGCN SGD MF SASRec UltraGCN of Users of Items Interactions Ratio (%) per User per Item 5130 87271 2928 19445 22363 603668 75258 27879 39387 5541 192403 14681 38609 66519 68223 123960 1429 4905 1686 19856 35598 16638 19412 24303 55969 193328 112134 157037 16252 728489 456 127278 776839 10715 139768 3589 326 297377 27403 101133 103102 236897 14169 1779 331844 240464 207725 55144 60970 23724 15798 1822 5543 4447 35992 1208 13392 1261 1083 2293 694 65080 42808 71478 74236 103256 44070 67440 58439 33406 44743 1859 4101 5199 854 27858 2628 2238 943 6040 610 23421 10000 1819 4931 60112 2414 241 43135
1685 13209 1835 7050 12101 367982 64443 10429 23033 3568 63001 8713 18534 28237 61934 50052 900 2420 962 8510 18357 10217 11924 10672 22611 78977 73303 47983 11269 159729 147 40994 188643 5013 98752 1366 138 59925 10449 27500 32472 42403 4962 729 103911 73153 78098 17286 8027 29794 38093 2069 15450 6811 18279 406 32983 1594 9989 15177 1435 81193 8787 11132 7747 18289 8682 12134 10326 7641 8287 2823 136 8640 8870 55542 763 301 1349 3416 3650 11888 7863 14885 3171 34185 576 1594 4659
37126 752937 20473 160792 198502 8898041 1097592 194439 278677 64706 1689188 151254 346355 551682 982619 1697533 10261 53258 13272 157836 296337 134476 167597 231780 438698 1637345 1399872 1119546 142820 6526514 2953 1063722 6627987 70053 2217545 26784 2171 3281108 218236 738270 744439 1948008 131058 11598 2675299 1957540 1754038 472857 6314631 687070 585579 28144 205793 87191 296647 31668 438576 9996 52785 146998 15253 1021304 680269 1150348 875411 1660900 717137 1120757 893541 461690 609316 71355 98559 751531 450123 990968 39652 15799 99287 999611 90274 802784 545817 984567 42481 387914 15307 20539 245210
3.04 6.61 1.6 2.76 1.85 1.64 1.17 2.67 1.71 1.55 3.05 1.68 2.08 2.36 1.1 2.48 1.59 2.03 1.75 2.33 1.94 1.63 1.63 2.28 2.48 2.45 1.53 3.27 1.44 4.56 3.1 3.1 4.12 2.14 1.42 2.63 2.36 4.96 2.62 3.68 3.18 5.59 2.86 2.44 3.19 3.29 2.66 3.19 7.6 0.8 0.41 0.88 0.36 0.65 1.97 2.98 0.41 0.79 0.11 0.15 0.48 0.8 4.87 6.42 9.58 5.65 5.08 5.56 5.66 4.37 5.4 0.66 30.15 0.6 0.1 0.5 3.44 7.44 0.7 1.77 0.17 1.97 1.27 0.12 1.56 1.76 4.19 0.15 9.26
0.43 0.07 0.38 0.12 0.07 0.00 0.02 0.07 0.03 0.33 0.01 0.12 0.05 0.03 0.02 0.03 0.8 0.45 0.82 0.09 0.05 0.08 0.07 0.09 0.03 0.01 0.02 0.01 0.08 0.01 4.41 0.02 0.00 0.13 0.02 0.55 4.83 0.02 0.08 0.03 0.02 0.02 0.19 0.89 0.01 0.01 0.01 0.05 1.29 0.1 0.1 0.75 0.24 0.29 0.05 6.46 0.1 0.5 0.49 0.42 1.53 0.02 0.18 0.14 0.15 0.09 0.19 0.14 0.15 0.18 0.16 1.36 17.67 1.67 5.94 0.06 1.98 2.35 7.8 4.84 4.05 0.29 0.69 3.64 0.27 0.02 1.1 5.35 0.12
7.24 8.63 6.99 8.27 8.88 14.74 14.58 6.97 7.08 11.68 8.78 10.3 8.97 8.29 14.4 13.69 7.18 10.86 7.87 7.95 8.32 8.08 8.63 9.54 7.84 8.47 12.48 7.13 8.79 8.96 6.48 8.36 8.53 6.54 15.87 7.46 6.66 11.03 7.96 7.3 7.22 8.22 9.25 6.52 8.06 8.14 8.44 8.57 103.57 28.96 37.07 15.45 37.13 19.61 8.24 26.22 32.75 7.93 48.74 64.11 21.98 15.69 15.89 16.09 11.79 16.09 16.27 16.62 15.29 13.82 13.62 38.38 24.03 144.55 527.08 35.57 15.09 7.06 105.29 165.5 147.99 34.28 54.58 541.27 8.62 6.45 6.34 85.22 5.68
22.03 57.0 11.16 22.81 16.4 24.18 17.03 18.64 12.1 18.14 26.81 17.36 18.69 19.54 15.87 33.92 11.4 22.01 13.8 18.55 16.14 13.16 14.06 21.72 19.4 20.73 19.1 23.33 12.67 40.86 20.09 25.95 35.14 13.97 22.46 19.61 15.73 54.75 20.89 26.85 22.93 45.94 26.41 15.91 25.75 26.76 22.46 27.35 786.67 23.06 15.37 13.6 13.32 12.8 16.23 78.0 13.3 6.27 5.28 9.69 10.63 12.58 77.42 103.34 113.0 90.81 82.6 92.37 86.53 60.42 73.53 25.28 724.7 86.98 50.75 17.84 51.97 52.49 73.6 292.63 24.73 67.53 69.42 66.14 13.4 11.35 26.57 12.89 52.63
127 127 127 127 -
190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 190 -
147 147 147 147 147 -
192 192 192 192 192 192 192 192 192 192 192 -
165 165 165 165 -
200 200 200 200 200 200 200 200 200 200 200 200 200 200
190 190 190 190 190 -