ConceptioArchivearXiv CS
arXiv CSopen access

STEB: Style Text Embedding Benchmark

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

STEB: Style Text Embedding Benchmark Rafael Rivera Soto1 , Anna Wegmann2 , and Cristina Aggazzotti1 1

Johns Hopkins University 2 Utrecht University 1 [email protected], [email protected], [email protected]

arXiv:2606.31741v1 [cs.CL] 30 Jun 2026

Abstract While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains fragmented, with each work relying on their own set of tasks and datasets. To bridge this gap, we introduce the Style Text Embedding Benchmark, a comprehensive open-source benchmark intended to standardize the evaluation of style embeddings. STEB encompasses 96 datasets across 7 languages, spanning applications such as authorship verification, authorship retrieval, AI-text detection, probing of linguistic features, and others. We find that semantic embeddings consistently fail in stylistic tasks, and that there is no style embedding that is universally superior across all tasks evaluated. We open-source the STEB code base at: https://github.com/rrivera1849/STEB.

1

Introduction

Representation learning has driven progress in NLP in the last decade, including on information retrieval, clustering, classification, and semantic search. While such models like SBERTmodels (Reimers and Gurevych, 2019), E5 (Wang et al., 2024a), LLM2Vec (BehnamGhader et al., 2024), and EmbeddingGemma (Vera et al., 2025) focus on capturing semantic meaning, a more disregarded parallel field focuses on the style of text.1 Style representations have proven valuable in applications such as authorship attribution (AA), style transfer, and AI-generated text detection (Horvitz et al., 2024a,b; Khan et al., 2024; Kim et al., 2025a; Rivera Soto et al., 2021, 2023; Wegmann et al., 2022), and might support the development of more style-aware LLMs (Wegmann et al., 2026). Currently, evaluation approaches for style embeddings are inconsistent. Works vary considerably in the tasks and datasets they evaluate on and in 1

Semantic meaning (or content) and style are hard to disentangle and might not be fully disjoint (Wegmann et al., 2026).

Figure 1: STEB Score by model category Style embeddings (blue) score above general-purpose semantic models (orange) on stylistic tasks; dashed lines mark the best model in each category. Qwen3-Embedding-8B, a top-5 MTEB model, ranks poorly on STEB.

evaluation protocol decisions like preprocessing, encoded text length, and documents per embedding. For example, LUAR (Rivera Soto et al., 2021) evaluates only on authorship retrieval, STAR adds author clustering (Huertas-Tato et al., 2024), and LISA (Patel et al., 2023) uses the STEL evaluation framework (Wegmann and Nguyen, 2021; Wegmann et al., 2022). The Massive Text Embedding Benchmark (MTEB) (Muennighoff et al., 2023), positioned as a general-purpose text embedding benchmark, offers no alternative as it includes no style-specific tasks, and many of its leading models (Bai et al., 2023; Lee et al., 2025; Li et al., 2023; Wang et al., 2024a; Xiao et al., 2023) are trained on objectives such as semantic similarity, information retrieval, and natural language inference (NLI) that benefit less from stylistic features (see Fig. 1). As a

result, cross-work comparisons are unreliable, and progress on style embeddings is hard to quantify. We introduce the open-source Style Text Embedding Benchmark (STEB). STEB consists of 96 datasets across 7 languages, organized into 5 evaluation tasks (clustering, pair classification, order alignment, retrieval, and probing) under fixed evaluation protocols (e.g., metrics, long document handling) covering applications like AI-text detection (ATD), authorship verification (AV) and authorship retrieval (AR). STEB reports two complementary scores. The operational score reflects the datasets and tasks that the field has built to study style. It inherits the emphasis of the field—most prominently AA work—without committing to an a priori definition. The definitional score instead reweighs results according to the style definition proposed in Wegmann et al. (2026). We evaluate 40 models spanning style embeddings, general-purpose semantic embeddings, masked language models (MLMs), causal language models (CLMs) and nonneural baselines. We find that no single model dominates across STEB. Instead, the best representation depends on whether the goal is strong performance on a specific downstream application or broad coverage of stylistic attributes, as well as on the degree to which the task is entangled with semantics. Notably, recent semantic embeddings that lead MTEB (e.g., Qwen-Embedding-8B) perform substantially worse than specialized style representations, highlighting a gap in MTEB’s evaluation of stylistic tasks. Furthermore, re-evaluating a prior multilingual setup under a different protocol for fairly handling long texts produces substantially different rankings, reinforcing the importance of standardized evaluation procedures. We also find that offthe-shelf MLMs perform surprisingly well at capturing linguistic features and remain competitive on AV and AR, suggesting promising directions for future research. Overall, we hope that STEB serves as a “yardstick” for measuring progress toward better style representations. Why not just use LLM prompting? While generative models can solve an increasing number of problems, text representations remain highly relevant and widely used in practice (Enevoldsen et al., 2025; Warner et al., 2025) for several reasons: (i) Text representations are more efficient: state-of-the-art encoder models are smaller (millions vs. billions of parameters) and only use a single forward pass, whereas generative models in-

cur cost proportional to the output length; (ii) they are competitive on discriminative tasks, matching or outperforming much larger generative models (Warner et al., 2025); and (iii) they scale much better to larger document pools for applications like RAG (Ram et al., 2023) or clustering, and might reduce hallucinations (Gao et al., 2024). We demonstrate that LUAR-CRUD (Rivera Soto et al., 2021) clearly beats GPT-5.2 on a small AR task, while using ≫750× fewer FLOPs (Tab. 1; full setup in App. E). Model

R@1 R@8 MRR TFLOPs USD

LUAR-CRUD GPT-5.2

83.0 59.0

95.0 69.0

87.8 63.4

≈ 32 ≫ 24k

0 22

Table 1: Embeddings outperform LLMs in AR at a fraction of the compute. The GPT-5.2 FLOP estimate uses 1 billion active parameters as a conservative lower bound.

2

Related works

Benchmarks for style representations Evaluation of style representations has mostly focused on authorship tasks (Wegmann et al., 2026). The longest-running effort in this space is the PAN shared task series at CLEF, which has run continuously since the late 2000s and describes itself as covering “digital text forensics and stylometry”.2 Early editions established cross-domain AA and AV as standard tasks (Juola and Stamatatos, 2013; Kestemont et al.; Stamatatos et al., 2014, 2015, 2018a), while later editions broadened to author profiling, style change detection, multi-author writing style analysis, hate-speech spreader profiling, and ATD (Ayele et al., 2024; Bevendorff et al., 2020, 2021, 2025a, 2026; Stamatatos et al., 2023). Most of the style embeddings we evaluate, evaluate on at least one PAN split. For AR, LUAR (Rivera Soto et al., 2021) introduces its own evaluation splits, which have since continued to be used (e.g., Man et al. (2026a)). Few evaluation approaches have used more theoretical definitionbased understandings. For example, STEL (Wegmann and Nguyen, 2021) introduces a parallel-text ranking evaluation over a small set of style dimensions, and STEL-or-Content (Wegmann et al., 2022) extends it with a content-controlled variant that penalizes models reliant upon semantics. Patel et al. (2025) further extends STEL by prompting 2

https://pan.webis.de/

LLMs to generate synthetic instances across 40 stylistic features.

3

The Style Text Embedding Benchmark

STEB consists of 96 datasets across 7 languages, organized into 5 evaluation tasks. In Section 3.1, we outline these tasks, define the metrics used for scoring each, and discuss how each serves to evaluate the quality of style embeddings. In §3.2, we provide information about the included datasets and their properties. In §3.3, we explain the operational and the definitional STEB score. In §3.4, we mention principles that we considered, and should generally be considered, when extending STEB. 3.1

Tasks and evaluation

See Tab. 12 for examples of all task types in STEB. Pair classification The goal of pair classification is to determine whether two input texts share the same label (e.g., same author in AV). To perform this, we embed both inputs and compute their cosine similarity, expecting high similarity between same-label pairs and low similarity otherwise. We implement two variants: (1) All-to-All, where every sample in the set is compared against every other sample providing an exhaustive evaluation, and (2) Predefined, where the pairs are restricted to a predefined list of pairs. This latter setting is critical for controlling for confounding variables (e.g., enforcing cross-topic pairs). We report the area under the receiver operating curve (AUROC). Clustering While semantic embeddings are optimized to cluster documents by topic, style embeddings should induce clusters corresponding to authorship (Andrews and Bishop, 2019), register (e.g., formality vs. informality) (Patel et al., 2025), and LM provenance (Rivera Soto et al., 2023). We evaluate whether style embeddings are capable of clustering text samples across style-relevant labels. Following the MTEB (Muennighoff et al., 2023) protocol, we use a mini-batch k-means algorithm with a batch size of 32 and k equal to the number of ground-truth labels. We report the V-measure (Rosenberg and Hirschberg, 2007). Authorship retrieval AR has emerged as a primary application of style embeddings (Agarwal et al., 2025; Fincke and Boschee, 2024; Kim et al., 2025a; Man et al., 2026b; Rivera Soto et al., 2021). Unlike traditional information retrieval, which optimizes for semantic relevance between a query and

a document, AR aims to identify other texts written by a specific author from within a large candidate pool. We use standard retrieval benchmarks where topic and authorship may correlate but are not the sole signal. We report the mean reciprocal rank (MRR). Order alignment Order alignment evaluates whether a style embedding ranks a set of texts in the same stylistic order as a reference set. In its simplest form, this corresponds to the STEL task introduced by Wegmann and Nguyen (2021), which uses parallel texts. A pair of texts has to be aligned to the order of another base pair of texts using stylistic information (see App. Fig. 2). Every order alignment task also has a distractor variant. Here, the unordered set includes one additional text that shares the same topic as the ordered texts but stylistically unrelated to any of them, a model that relies on semantic features more than style will misalign it. In a setup with a one element ordered set, and with a two element unordered set including one distractor, this is equivalent to the STEL-orContent task (Wegmann et al., 2022). We report two accuracies, one for the original and one for the distractor variant. To achieve a high score in this task, an embedding must capture features other than semantics, and in the distractor variant it must ignore semantic similarity altogether. Probing The probing task assesses which linguistic features are encoded in an embedding by training a linear classifier on top of frozen representations (Conneau et al., 2018). We extract foundation-level linguistic features using the LFTK toolkit (Lee and Lee, 2023a), excluding features that may correlate with semantics (e.g., entity counts). Each continuous feature is discretized into five quantile bins, and the train/validation/test splits are balanced across the discretized labels. We then train a logistic regression probe on the frozen embeddings, selecting the L2 regularization strength from {10−5 , 10−4 , 10−3 , 10−2 } via a validation set, and report test accuracy averaged across all features. 3.2

STEB datasets

An overview of all datasets can be found in §C. The added datasets can generally be grouped into 6 categories: those that target capabilities with respect to single linguistic features (7 datasets), authorial styles (58), dialectal styles (3), registers (12), genres (4), historical styles (3) and demographics (9).

3.3

STEB Scores

With STEB, we evaluate models on the most common datasets and tasks that the field has built to study style. We aggregate the results differently for the operational score and the definitional score. Operational STEB Score While we do not align the operational score with a definition, we take care to not overweigh groups of correlated datasets. We automatically discover redundancies following Olmo et al. (2026) using the Spearman correlation between the model rankings on each dataset and apply hierarchical clustering (Ward linkage with a 0.5 distance threshold, cf. Ward Jr., 1963). Scores are macro-averaged within each and then across clusters. Definitional STEB Score To contrast to the operational view of style (cf. § 1), we also evaluate embeddings based on an a priori definition of style, namely the definition introduced in Wegmann et al. (2026). We evaluate whether embeddings (i) encode linguistic features, (ii) capture the stylistic patterns of different objects of study (e.g., dialect, genre, idiolect), and (iii) are more sensitive to stylistic than content information. See App. Tab. 13 for an overview of the dataset-to-cluster assignment. 3.4

Design principles

We provide code and datasets with STEB that are meant to be extended by the community. We emphasize the principles on which we designed STEB and that we recommend the community to follow: Tasks where style matters Fundamentally, STEB targets scenarios where the semantics of a text is neither the sole nor strongest signal. Where possible, we enforce topic-controlled settings that penalize models that rely on any semantic shortcuts (e.g., Tab. 4). As such, each STEB dataset should benefit from representations that capture stylistic information and should be more difficult for models that heavily rely on semantic information. Length control To preserve ecological validity (de Vries et al., 2020), we evaluate models on the full text length inherent to each dataset. This approach preserves the natural length distributions between domains, for example, retaining the verbosity of blog posts versus the brevity of Reddit comments. For models with fixed context windows (e.g., 512 tokens), we employ a chunk-and-pool strategy where inputs exceeding the limit are seg-

mented in sentence-boundary-aware chunks. We embed each chunk independently and apply mean pooling to derive a single document representation (see Fig. 4 for a visual of the approach). Reproducibility and extensibility For reproducibility, all stochastic steps in evaluation (e.g., k-means initialization, probe training) use fixed random seeds. To facilitate community adoption, STEB is open-source and easy to extend. We welcome contributions from the community in adding new tasks and datasets.

4

Models

We test non-neural (§4.1) as well as automaticallylearned representations (§4.2). 4.1

Non-neural representations

We test various representations using predefined, linguistically motivated features. Note that the strongest non-neural representation generally combines several of these features and tailors the set to the specific dataset or task (Wegmann et al., 2026). Function word frequencies Function words (i.e., stop words) like determiners (the) and pronouns (you) are useful across text analyses, such as AA (Kestemont, 2014; Mosteller and Wallace, 1963), genre/text type classification (Venglařová and Matlach, 2024), and deception detection (Liu et al., 2012). We count the frequencies of 390 function words and 69 function phrases (instead of ), saving these counts as a vector per text (Aggazzotti and Smith, 2025). TF-IDF-weighted n-grams N-grams—sequences of n consecutive units (i.e., characters, tokens, part-of-speech (POS) tags)—have long been used for text analysis, especially for AA (Houvardas and Stamatatos, 2006; Peng et al., 2003). Following others (e.g., Weerasinghe and Greenstadt, 2020), we weight the n-grams by their Term FrequencyInverse Document Frequency (TF-IDF) and fit the vectorizer to a 10 billion token sample of FineWeb (Penedo et al., 2024); see §D for more details.3 Stylometric features There is no agreed upon fixed set of best stylometric features (Juola, 2006; Nini, 2023). There are perhaps equally as many feature extraction tools for various use cases. Due to its popularity in the NLP community, we use 3

Note that n-grams can capture topic as well as style signals, depending on the n-gram unit and their frequency.

LFTK (Lee and Lee, 2023b) and combine its surface and POS features, which we found to perform best across datasets while remaining fast. NeuroBiber NeuroBiber (Alkiek et al., 2025) is a transformer-based system that extracts 96 syntactic and lexical “Biber features”, such as function words, subordination types, and tenses (Biber, 1988; Biber and Conrad, 2019). 4.2

Masked language models (MLMs) We include encoder-only transformers pre-trained with masked language modeling and without embeddingspecific fine-tuning. Specifically, we include RoBERTa (Liu et al., 2019), DeBERTa-v3 (He et al., 2023), and ModernBERT (Warner et al., 2025). Since these models do not produce a single vector, we obtain embeddings by mean-pooling the token-level hidden states from the last layer, weighted by the attention mask.

Neural models

We benchmark models with state-of-the-art results on various style and semantic embedding tasks, grouped into four families: style embeddings, semantic embeddings, MLMs, and CLMs. Style embeddings Most style embedding models we evaluate are trained contrastively, differing primarily in how they construct positive and negative pairs. LUAR (Rivera Soto et al., 2021) and STAR (Huertas-Tato et al., 2024) treat texts by the same author as positives and different authors as negatives, learning authorship representations from Reddit and various social media sources, respectively. CISR (Wegmann et al., 2022) further constrains positives to be same-author but different-topic pairs, encouraging content-independent style representations. StyleDistance (Patel et al., 2025) constructs synthetic hard-positives and hard-negatives designed to isolate stylistic correlations from topical ones. mStyleDistance (Qiu et al., 2025) extends this approach to the multilingual setting. MSR (Kim et al., 2025b) learns multilingual authorship embeddings using probabilistic content masking to suppress content features and language-aware batching to reduce cross-lingual easy-negatives. LISA (Patel et al., 2023) departs from purely contrastive training, learning interpretable style embeddings via a linear projection over an EncT5 encoder. Semantic embeddings These models are trained for general-purpose text similarity, primarily on semantic retrieval and NLI data. We include models like all-mpnet-base-v2 (Reimers and Gurevych, 2019) that perform well according to the semantic textual similarity evaluations and various models that perform(ed) well according to MTEB (Muennighoff et al., 2023): Qwen3-Embedding8B (Zhang et al., 2025), GTE (Li et al., 2023), E5 (Wang et al., 2024a), BGE (Xiao et al., 2023), and Jina Embeddings v3 (Sturua et al., 2025).

Causal language models (CLMs) Autoregressive models are pretrained with next-token prediction and have no embedding-specific training. We include them to assess whether large-scale causal language modeling implicitly captures stylistic features. We include GPT-2 XL (Radford et al., 2019), OPT-1.3B (Zhang et al., 2022) and the Qwen family at multiple scales. We use the hidden state of the last non-padding token as the embedding, as it is the only position with full sequence context.

5

Results

We discuss operational and definitional STEB results (§ 5.1) and provide a per-application (§ 5.2) and a multilingual (§5.3) analysis; all but the last are restricted to English. 5.1

Overall Results

STEB (operational) results are in Tab. 2 and selected STEB (definitional) results in Tab. 3; full results for all 40 models are in App. Tab. 6 and Tab. 9. No single winner across task categories and style definitions There is no consistent winner across categories, for neither STEB (definitional) nor STEB (operational). While StyleDistance wins for the definitional, STAR, LUAR-CRUD, and LUAR-MUD score similarly for the operational score. STAR’s advantage is in clustering and predefined pair classification, while it falls behind on retrieval, where LUAR-CRUD leads by over 10 points, and order alignment, where StyleDistance leads by nearly 19 points. Models like STAR and LUAR might profit from content-entangled features for different objects of study that are punished more heavily for the definitional STEB score. These patterns confirm previous observations that style is conceptualized differently across the literature, and that different training and evaluation objectives— motivated by different style definitions—result in

Stylespecific

LUAR-MUD LUAR-CRUD StyleDistance mStyleDistance MSR CISR STAR LISA

20.03 19.65 18.46 7.35 16.67 15.28 25.92 12.74

62.33 61.00 59.82 53.83 64.34 59.71 64.56 60.91

72.33 73.12 66.21 53.71 71.01 66.43 73.28 62.40

20.81 19.80 44.02 31.38 21.85 33.81 23.93 16.48

75.67 77.64 49.61 20.04 63.49 45.87 66.96 42.96

53.04 53.75 57.87 45.96 51.01 52.54 49.86 46.38

50.70 50.82 49.33 35.38 48.06 45.61 50.75 40.31

Semantic

all-mpnet-base-v2 E5-large-v2 Qwen3-Embedding-8B

8.48 11.06 2.67

60.41 60.20 51.61

63.87 64.70 52.79

14.62 15.82 23.11

48.99 49.91 35.17

45.64 49.03 42.96

40.34 41.79 34.72

MLM

BERT-large-cased RoBERTa-large DeBERTa-v3-large ModernBERT-large

12.02 13.21 21.44 16.58

60.99 60.97 59.37 60.19

67.11 68.79 70.45 65.38

21.57 20.71 24.23 20.56

58.88 61.07 63.24 52.17

55.16 57.92 53.41 54.06

45.96 47.11 48.69 44.82

CLM

Clustering All-to-All Predefined Order Align. Retrieval Probing STEB score 53 24 26 11 4 4 122

OPT-1.3B Qwen3.5-4B-Base

3.83 4.23

52.40 52.80

55.59 57.41

24.08 24.88

39.07 37.15

45.82 43.02

36.80 36.58

Nonneural

Model Num. Datasets

Function words TFIDF n-grams Stylometric NeuroBiber

7.17 6.78 9.67 8.24

55.17 57.18 56.88 55.10

60.88 64.45 60.48 57.27

14.51 13.72 25.87 14.13

43.14 50.95 39.52 9.06

37.98 55.74 53.30 51.64

36.47 41.47 40.95 32.57

Table 2: Overall STEB (operational) results (×100). Bold = best, underline = 2nd best per column. Metrics: V-measure (Clustering), AUC (All-to-All and Predefined Pair Classification), distractor accuracy (Order Alignment), MRR (Retrieval), and average accuracy (Probing). Scores are macro-averaged across automatically discovered dataset clusters within each task, then averaged across tasks to produce the overall STEB score. Num. Datasets = datasets contributing to each task; a dataset may contribute to more than one task. The full table is in App. Tab. 6.

Style emb.

Top 5

Object of Study Model Num. Datasets

Genre Register Time Demo. Dialect Idiolect 4 7 3 3 3 29

StyleDistance StyleDist. Synth. CISR DeBERTa large RoBERTa-base

62.84 55.45 62.43 63.26 63.79

40.01 38.81 36.09 38.25 42.52

46.21 43.27 41.52 46.94 45.94

33.67 33.13 33.03 35.37 32.84

56.91 55.64 54.76 43.37 40.16

STAR LUAR-MUD mStyleDistance LUAR-CRUD MSR LISA

68.68 66.67 49.41 65.80 66.39 56.89

39.79 38.72 32.83 38.14 39.08 37.95

47.25 42.14 35.04 41.30 41.98 37.38

40.99 35.53 32.25 35.32 37.05 32.95

49.39 44.30 42.83 41.52 41.52 27.19

Ling. Feat. Content-Ind.

Avg.

7

Order Al. 11

67

49.94 45.26 47.76 49.24 47.51

74.02 70.87 65.85 69.56 77.95

44.23 46.10 32.31 18.72 10.27

56.06 54.08 48.64 45.84 45.24

53.06 50.58 38.27 49.85 49.12 41.06

67.66 70.50 56.17 70.37 66.52 64.02

14.08 11.29 38.32 10.57 12.35 7.10

44.93 44.12 44.25 43.59 42.66 37.39

Avg. 49

60.03 45.26 58.72 68.26 59.83 72.24 76.11 37.23 76.99 68.69 54.00

Table 3: Top 5 models for STEB Score (definitional). Attributes are grouped into three clusters after Wegmann et al. (2026) and averaged: Object of Study (Genre, Style, Time, Demographics, Dialect, Idiolect), Linguistic Features, and Content-Independence. See App. Tab. 13 for selected datasets and App. Tab. 9 for full results.

different strengths and weaknesses across downstream tasks (Wegmann et al., 2026). Style embeddings consistently outperform semantic embeddings. Every style-specific model except LISA and mStyleDistance outperform the best general-purpose semantic embedding on both STEB scores. This disparity is especially pronounced in order alignment, AR, and predefined pair classification. Semantic embeddings are opti-

mized on objectives that favor matching texts primarily by semantic content, and the gap on STEB suggests style features are not well captured by semantic objectives alone. These findings underscore the need for a dedicated style benchmark, as semantic-oriented evaluation fails to capture the capabilities of style-specific models. MLM pre-training captures stylistic information. All MLMs we evaluate outperform every

Model

ATD

ATD (Adv.)

Authorship Verification

AR

Avg.

Top 5

deberta-v3-large STAR LUAR-MUD deberta-v3-base LUAR-CRUD

36.03 22.79 28.20 29.55 25.20

100.00 100.00 75.01 99.24 72.61

73.29 77.52 76.54 69.40 76.35

77.68 88.37 86.84 73.46 86.60

62.92 76.02 73.62 59.80 75.92

55.87 58.26 61.10 54.75 61.23

63.24 66.96 75.67 55.17 77.64

68.14 66.82 63.85 63.34 62.95

Style Embeddings

Overall Easy Medium Hard

StyleDistance mStyleDistance StyleDistance (Synthetic) MSR CISR LISA

30.10 15.85 22.76 17.38 22.78 11.61

70.08 5.48 44.03 0.12 64.33 6.06

70.44 54.41 59.00 73.88 71.57 65.03

75.52 60.49 64.82 88.92 79.39 84.41

69.17 49.63 55.28 73.43 70.74 66.71

62.18 50.53 52.62 55.99 65.56 47.86

49.61 20.04 31.53 63.49 45.87 42.96

55.06 23.95 39.33 38.72 51.14 31.41

Table 4: Top 5 models across application clusters, plus the remaining style embeddings. ATD = AI-Text Detection, AR = Authorship Retrieval. Authorship Verification: Overall aggregates a superset of datasets (PAN13–15, PAN20–21, Enron, and the PAN22–26 style-change benchmarks), while Easy/Medium/Hard are computed over the PAN22–26 style-change subset only—they do not average to the Overall column. To avoid double-counting, only Overall contributes to Avg. Bold = best, underline = 2nd best per column (across all models). Each column reports the macro-averaged score (×100) within the corresponding manual cluster.

semantic embedding and every causal LM on both STEB scores, despite having no embedding-specific fine-tuning. DeBERTa-v3-large outperforms some style embeddings on retrieval and matches stylespecific models on authorship verification, without any author-specific training (Tab. 2). RoBERTa leads over style embeddings in sensitivity to linguistic features (probing column of Tab. 2; linguisticfeature column of Tab. 3), suggesting that stylespecific fine-tuning can come at the cost of linguistic feature sensitivity. Joint training on MLM and style objectives may help recover this sensitivity in style-specific models. The newer ModernBERT variants (only shown in App. Tab. 6 & Tab. 9) underperform, suggesting that more recent MLM recipes have shifted away from preserving linguistic feature information. Together, these results suggest that new MLM objectives are a promising direction for future research. Non-neural models are competitive on certain tasks. TFIDF n-grams is the strongest non-neural model, followed closely by Stylometric. Both generally outperform causal LMs and perform similarly to general-purpose sentence embedders, but lag behind style-specific and masked models. They perform best on tasks like probing and predefined classification, but degrade on tasks like clustering and retrieval that often require richer representations. Despite TFIDF n-grams being effectively a bag-of-words model, it nearly matches sophisticated sentence embedders. Also, although the literature has shown no general-purpose best set

of stylometric features, using surface and POS tag features does fairly well across datasets and tasks.4 CISR and StyleDistance models perform best at content independence. StyleDistance, CISR, and mStyleDistance perform the best at contentindependence (i.e., the distractor variant of the order alignment tasks). This is expected as they were all trained similarly with hard negatives to improve content independence. Across models, though, there is still a lot of room for improvement on content-independence. STAR is most sensitive to style across stylistics attributes. STAR performs the best at recovering distinctive patterns across “objects of study” (e.g., genre, time, registers) beyond authorial styles. This might be influenced by its being one of few models that was trained on several different domains (i.e., social media, blogs, and books). Including more diverse training datasets (and potentially training tasks, cf. Wegmann et al., 2026) might thus be a promising direction for more generalizability. 5.2

Per-application analysis

To answer questions such as “Which model is best for AI-text detection?”, we manually group datasets into four application-oriented clusters (Tab. 4): ATD, ATD where text has been modified to evade detection (Adv.), AV, and AR. For AV we additionally report three sub-columns over the PAN 4 Tailoring stylometric features to each dataset/task would likely produce even higher results, and we encourage testing specialized feature sets for more representative performance.

Style embeddings

Top 5

Model

Pair Class. Retrieval

Avg.

deberta-v3-large StyleDistance roberta-large CISR deberta-v3-base

74.03 72.40 71.91 70.08 71.85

47.77 44.45 44.02 37.93 35.42

60.90 58.42 57.96 54.01 53.63

LUAR-MUD LUAR-CRUD mStyleDistance StyleDist. Synth. MSR STAR LISA

69.50 72.37 59.94 64.03 69.43 71.32 62.67

32.38 30.07 14.69 19.46 28.38 35.80 26.70

50.94 51.22 37.32 41.74 48.90 53.56 44.68

Table 5: Top 5 multilingual STEB results (×100). Pair Class. averages AUC across 13 PAN13/14/15 AV datasets (Dutch, Greek, Spanish); Retrieval averages MRR across 4 PAN18 cross-domain AA datasets (French, Italian, Polish, Spanish). Bold = best, underline = 2nd best per column (across all models).

style-change benchmarks ordered by how contententangled the pairs are: Easy, Medium, and Hard, where Hard pairs do not share any semantic content. The top five comprise three style-specific embeddings (STAR, LUAR-MUD, LUAR-CRUD) and two DeBERTa-v3 variants. In ATD, DeBERTa-v3large5 leads in the standard setting, while in the adversarial setting, STAR and DeBERTa-v3-large perform best. In AV and AR, the style-specific embedding models take the lead, suggesting that they are better suited for capturing the fine-grained nuances of everyday authors. The Easy/Medium/Hard AV columns reveal a finer pattern. On the Easy and Medium sets (where topic can be exploited), both STAR and LUAR perform well. However, in the Hard setting (where topic correlations cannot be exploited), CISR and StyleDistance perform best, suggesting that disentangling style from content matters for such applications. 5.3

Multilingual analysis

STEB includes non-English datasets across six languages (Dutch, French, Greek, Italian, Polish, Spanish) drawn from the PAN13/14/15 AV shared tasks and the PAN18 cross-domain AA shared task. These datasets are excluded from the main STEB scores (Tab. 2 and Tab. 3) as most evaluated models are monolingual English encoders. Here, we 5

While striking, this result is explained by the fact that DeBERTav3 adds a replaced token detection objective, where a discriminator is trained to detect which tokens in a sequence were replaced by the MLM (He et al., 2023). This objective is none other than ATD in a pre-ChatGPT era (first DeBERTa version was published in 2021)!

investigate whether multilingual style embeddings perform better than monolingual English style embeddings on these multilingual applications. We derive a single score by averaging across all dataset metrics and across the independent tasks (pair classification, retrieval). Multilingual style-representations underperform their English counterparts Tab. 5 shows that the top-5 by multilingual average comprises two DeBERTa-v3 variants, RoBERTa-large, and two style embeddings (StyleDistance, CISR). The explicitly multilingual style embeddings (MSR, mStyleDistance) do not appear in the top five. These results suggest that current multilingual training strategies are not yet effective, and as such, they remain a challenge for future work.6 The two style embeddings that do appear are notable for having been trained so as to separate semantic content from style, suggesting that this separation might be a useful strategy for the generalizability of style embeddings across languages.

6

Conclusion

We introduced STEB, a benchmark for evaluating embeddings across style-related NLP tasks, spanning 5 tasks, 96 datasets, and 7 languages. Across 40 models, we find that no single embedding is best. Style embeddings that retain semantic signal lead on authorship-related applications where author and topic correlate, while contentdisentangled style embeddings lead on contentcontrolled applications where topic shortcuts are removed. On ATD, MLMs outperform style-specific embeddings, and on predicting second-order style features (e.g., genre, dialect, register, time), the authorship-focused STAR model leads. Multilingual style embeddings underperform their English counterparts on the multilingual subset of STEB, highlighting cross-lingual style as an open problem. We release STEB to enable consistent, reproducible comparisons and to lower the cost of measuring progress on style. We welcome contributions from the community in improving the benchmark. 6

To stress-test this conclusion, we reproduce Kim et al. (2025b)’s multilingual evaluation setup on the same PAN13/14/15 multilingual AV datasets and style embeddings. We find that STEB’s default chunk-and-pool strategy for handling large texts is the main driver. Under per-document fixed truncation, MSR achieves the top score, while with the chunk-and-pool strategy, StyleDistance and CISR perform best. These results underscore the necessity of shared evaluation protocols, since without them our conclusions would differ materially. See §F for the full comparison.

Limitations Leakage of datasets / contamination Style embeddings are commonly trained and tested on similar datasets, often using datasets for training and testing with differing splits between publications. Representations might be tested on some data that they have seen during training. For example, the StyleDistance models were trained on one of the seven “linguistic features” datasets and on two of the eleven “Content Independence” datasets in Tab. 3. Excluding these shrinks the difference to CISR to ≈2 percentage points on content independence and make the synthetic version and CISR have almost the same overall performance. However, the overall ranking of the top models is preserved. Bounded by dataset coverage Both the empirical clustering in §5.1 and the attribute clustering in §5.1 are bounded by what existing datasets expose. In particular, most of the field’s emphasis has been on authorship-related tasks, and as such other axes of analysis remain underexplored. We call on the community to create datasets that explore aspects other than authorship and to expand STEB with these additions. Languages other than English remains underrepresented It is a fact of the field that not as many non-English datasets are available. As such, we inherit this limitation. Moreover, only recently has there been an effort to extend style embedding models to multilingual settings (Kim et al., 2025a; Qiu et al., 2025).

References Shantanu Agarwal, Joel Barry, Steven Fincke, and Scott Miller. 2025. Cross-Genre Authorship Attribution via LLM-Based Retrieve-and-Rerank. arXiv preprint. ArXiv:2510.16819 [cs]. Cristina Aggazzotti and Elizabeth Allyn Smith. 2025. A stylometric analysis of speaker attribution from speech transcripts. Preprint, arXiv:2512.13667. Kenan Alkiek, Anna Wegmann, Jian Zhu, and David Jurgens. 2025. Neurobiber: Fast and interpretable stylistic feature extraction. arXiv preprint. ArXiv:2502.18590 [cs]. Tiago A. Almeida, José María G. Hidalgo, and Akebo Yamakami. 2011. Contributions to the study of SMS spam filtering: new collection and results. In Proceedings of the 11th ACM symposium on Document engineering, DocEng ’11, pages 259–262, New York, NY, USA. Association for Computing Machinery.

Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, and Lucia Specia. 2020. ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting Transformations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4668–4679, Online. Association for Computational Linguistics. Nicholas Andrews and Marcus Bishop. 2019. Learning invariant representations of social media users. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1684– 1695, Hong Kong, China. Association for Computational Linguistics. Abinew Ali Ayele, Nikolay Babakov, Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashaf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korenčić, Maximilian Mayerl, Daniil Moskovskiy, Animesh Mukherjee, Alexander Panchenko, Martin Potthast, Francisco Rangel, Naquee Rizwan, Paolo Rosso, Florian Schneider, Alisa Smirnova, Efstathios Stamatatos, Elisei Stakovskii, Benno Stein, Mariona Taulé, Dmitry Ustalov, Xintong Wang, Matti Wiegmann, Seid Muhie Yimam, and Eva Zangerle. 2024. Overview of PAN 2024: Multi-author writing style analysis, multilingual text detoxification, oppositional thinking analysis, and generative ai authorship verification condensed lab overview. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 15th International Conference of the CLEF Association, CLEF 2024, Grenoble, France, September 9–12, 2024, Proceedings, Part II, page 231–259, Berlin, Heidelberg. Springer-Verlag. Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report. Preprint, arXiv:2309.16609. Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1644–1650, Online. Association for Computational Linguistics. Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2vec: Large language mod-

els are secretly powerful text encoders. In First Conference on Language Modeling. Janek Bevendorff, Ian Borrego-Obrador, Mara ChineaRíos, Marc Franco-Salvador, Maik Fröbe, Annina Heini, Krzysztof Kredens, Maximilian Mayerl, Piotr P˛ezik, Martin Potthast, Francisco Rangel, Paolo Rosso, Efstathios Stamatatos, Benno Stein, Matti Wiegmann, Magdalena Wolska, and Eva Zangerle. 2023. Overview of PAN 2023: Authorship Verification, Multi-Author Writing Style Analysis, Profiling Cryptocurrency Influencers, and Trigger Detection. In Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 459–481. Springer, Cham. Janek Bevendorff, Berta Chulvi, Gretel Liz De La Peña Sarracén, Mike Kestemont, Enrique Manjavacas, Ilia Markov, Maximilian Mayerl, Martin Potthast, Francisco Rangel, Paolo Rosso, Efstathios Stamatatos, Benno Stein, Matti Wiegmann, Magdalena Wolska, and Eva Zangerle. 2021. Overview of PAN 2021: Authorship verification, profiling hate speech spreaders on twitter, and style change detection. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 12th International Conference of the CLEF Association, CLEF 2021, Virtual Event, September 21–24, 2021, Proceedings, page 419–431, Berlin, Heidelberg. Springer-Verlag. Janek Bevendorff, Berta Chulvi, Elisabetta Fersini, Annina Heini, Mike Kestemont, Krzysztof Kredens, Maximilian Mayerl, Reynier Ortega-Bueno, Piotr P˛ezik, Martin Potthast, Francisco Rangel, Paolo Rosso, Efstathios Stamatatos, Benno Stein, Matti Wiegmann, Magdalena Wolska, and Eva Zangerle. 2022. Overview of PAN 2022: Authorship Verification, Profiling Irony and Stereotype Spreaders, and Style Change Detection. In Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 382–394. Springer, Cham. Janek Bevendorff, Daryna Dementieva, Maik Fröbe, Bela Gipp, André Greiner-Petter, Jussi Karlgren, Maximilian Mayerl, Preslav Nakov, Alexander Panchenko, Martin Potthast, Artem Shelmanov, Efstathios Stamatatos, Benno Stein, Yuxia Wang, Matti Wiegmann, and Eva Zangerle. 2025a. Overview of PAN 2025: Voight-kampff generative ai detection, multilingual text detoxification, multi-author writing style analysis, and generative plagiarism detection. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 16th International Conference of the CLEF Association, CLEF 2025, Madrid, Spain, September 9–12, 2025, Proceedings, page 388–411, Berlin, Heidelberg. Springer-Verlag. Janek Bevendorff, Daryna Dementieva, Maik Fröbe, Bela Gipp, André Greiner-Petter, Jussi Karlgren, Maximilian Mayerl, Preslav Nakov, Alexander Panchenko, Martin Potthast, Artem Shelmanov, Efstathios Stamatatos, Benno Stein, Yuxia Wang, Matti Wiegmann, and Eva Zangerle. 2025b. Overview of PAN 2025: Generative AI detection, multilingual

text detoxification, multi-author writing style analysis, and generative plagiarism detection. In Advances in Information Retrieval, pages 434–441. Springer, Cham. Janek Bevendorff, Maik Fröbe, André GreinerPetter, Andreas Jakoby, Maximilian Mayerl, Preslav Nakov, Henry Plutz, Martin Potthast, Benno Stein, Minh Ngoc Ta, Yuxia Wang, and Eva Zangerle. 2026. Overview of PAN 2026: Voight-kampff generative ai detection, text watermarking, multi-author writing style analysis, generative plagiarism detection, and reasoning trajectory detection. Preprint, arXiv:2602.09147. Janek Bevendorff, Bilal Ghanem, Anastasia Giachanou, Mike Kestemont, Enrique Manjavacas, Ilia Markov, Maximilian Mayerl, Martin Potthast, Francisco Rangel, Paolo Rosso, Günther Specht, Efstathios Stamatatos, Benno Stein, Matti Wiegmann, and Eva Zangerle. 2020. Overview of PAN 2020: Authorship verification, celebrity profiling, profiling fake news spreaders on twitter, and style change detection. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 11th International Conference of the CLEF Association, CLEF 2020, Thessaloniki, Greece, September 22–25, 2020, Proceedings, page 372–383, Berlin, Heidelberg. Springer-Verlag. Douglas Biber. 1988. Variation across Speech and writing. New York, NY: Cambridge University Press. Douglas Biber and Susan Conrad. 2019. Register, Genre, and Style, 2nd edition. Cambridge University Press, Cambridge, UK. Keith Carlson, Allen Riddell, and Daniel Rockmore. 2018. Evaluating prose style transfer with the Bible. Royal Society Open Science, 5(10):171920. Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136, Melbourne, Australia. Association for Computational Linguistics. Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Dan Jurafsky, Jure Leskovec, and Christopher Potts. 2013. A computational approach to politeness with application to social factors. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 250–259, Sofia, Bulgaria. Association for Computational Linguistics. Harm de Vries, Dzmitry Bahdanau, and Christopher D. Manning. 2020. Towards ecologically valid research on language user interfaces. ArXiv, abs/2007.14435. Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra Winata, Saba Sturua, Saiteja Utpala,

Mathieu Ciancone, Marion Schaeffer, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Veysel Çağatan, Akash Kundu, Martin Bernstorff, Shitao Xiao, Akshita Sukhlecha, Bhavish Pahwa, Rafał Poświata, Kranthi Kiran Gv, Shawon Ashraf, Daniel Auras, Björn Plüster, Jan Philipp Harries, Loïc Magne, Isabelle Mohr, Dawei Zhu, Hippolyte Gisserot-Boukhlef, Tom Aarsen, Jan Kostkan, Konrad Wojtasik, Taemin Lee, Marek Suppa, Crystina Zhang, Roberta Rocca, Mohammed Hamdy, Andrianos Michail, John Yang, Manuel Faysse, Aleksei Vatolin, Nandan Thakur, Manan Dey, Dipam Vasani, Pranjal A. Chitale, Simone Tedeschi, Nguyen Tai, Artem Snegirev, Mariya Hendriksen, Michael Günther, Mengzhou Xia, Weijia Shi, Xing Han Lù, Jordan Clive, Gayatri K, Maksimova Anna, Silvan Wehrli, Maria Tikhonova, Henil Shalin Panchal, Aleksandr Abramov, Malte Ostendorff, Zheng Liu, Simon Clematide, Lester James Validad Miranda, Alena Fenogenova, Guangyu Song, Ruqiya Bin Safi, Wen-Ding Li, Alessia Borghini, Federico Cassano, Lasse Hansen, Sara Hooker, Chenghao Xiao, Vaibhav Adlakha, Orion Weller, Siva Reddy, and Niklas Muennighoff. 2025. MMTEB: Massive Multilingual Text Embedding Benchmark. In The Thirteenth International Conference on Learning Representations (ICLR).

Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023. DeBERTav3: Improving deBERTa using ELECTRAstyle pre-training with gradient-disentangled embedding sharing. In The Eleventh International Conference on Learning Representations. Zachary Horvitz, Ajay Patel, Chris Callison-Burch, Zhou Yu, and Kathleen McKeown. 2024a. ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16):18216–18224. Zachary Horvitz, Ajay Patel, Kanishk Singh, Chris Callison-Burch, Kathleen McKeown, and Zhou Yu. 2024b. TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 13376–13390, Miami, Florida, USA. Association for Computational Linguistics. John Houvardas and Efstathios Stamatatos. 2006. Ngram feature selection for authorship identification. In Proceedings of the 12th International Conference on Artificial Intelligence: Methodology, Systems, and Applications, AIMSA’06, page 77–86, Berlin, Germany. Springer.

Steven Fincke and Elizabeth Boschee. 2024. Separating Style from Substance: Enhancing Cross-Genre Authorship Attribution through Data Selection and Presentation. arXiv preprint. ArXiv:2408.05192 [cs].

Baixiang Huang, Canyu Chen, and Kai Shu. 2024. Can Large Language Models Identify Authorship? In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 445–460, Miami, Florida, USA. Association for Computational Linguistics.

Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv preprint. ArXiv:2312.10997 [cs].

Javier Huertas-Tato, Alejandro Martín, and David Camacho. 2024. Understanding writing style in social media with a supervised contrastively pre-trained transformer. Know.-Based Syst., 296(C).

Lukas Gehring and Benjamin Paaßen. 2025. Assessing LLM text detection in educational contexts: Does human contribution affect detection? Preprint, arXiv:2508.08096. Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020. Investigating AfricanAmerican Vernacular English in Transformer-Based Text Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5877–5883, Online. Association for Computational Linguistics. Abhay Gupta, Jacob Cheung, Philip Meng, Shayan Sayyed, Kevin Zhu, Austen Liao, and Sean O’Brien. 2025. EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 16830–16855, Suzhou, China. Association for Computational Linguistics. Oren Halvani. 2017. Enron authorship verification corpus. https://data.mendeley.com/datasets/n 77w7mygwg/1.

Nancy Ide, Collin Baker, Christiane Fellbaum, Charles Fillmore, and Rebecca Passonneau. 2008. MASC: the Manually Annotated Sub-Corpus of American English. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), Marrakech, Morocco. European Language Resources Association (ELRA). Nancy Ide, Keith Suderman, Collin Baker, Rebecca Passonneau, and Christiane Fellbaum. 2013. Manually annotated sub-corpus third release. Web Download. Patrick Juola. 2006. Authorship attribution. Foundations and Trends in Information Retrieval, 1(3):233–334. Patrick Juola and Efstathios Stamatatos. 2013. Overview of the author identification task at PAN 2013. In Conference and Labs of the Evaluation Forum. Dongyeop Kang, Varun Gangal, and Eduard Hovy. 2019. (Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple Personas. In Proceedings of the 2019 Conference on Empirical Methods

in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1696–1706, Hong Kong, China. Association for Computational Linguistics. Mike Kestemont. 2014. Function words in authorship attribution. From black magic to theory? In Proceedings of the 3rd Workshop on Computational Linguistics for Literature (CLFL), pages 59–66, Gothenburg, Sweden. Association for Computational Linguistics. Mike Kestemont, Michael Tschuggnall, Efstathios Stamatatos, Walter Daelemans, Günther Specht, Benno Stein, and Martin Potthast. Overview of the Author Identification Task at PAN-2018. Aleem Khan, Elizabeth Fleming, Noah Schofield, Marcus Bishop, and Nicholas Andrews. 2021. A Deep Metric Learning Approach to Account Linking. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5275–5287, Online. Association for Computational Linguistics. Aleem Khan, Andrew Wang, Sophia Hager, and Nicholas Andrews. 2024. Learning to generate text in arbitrary writing styles. Preprint, arXiv:2312.17242. Junghwan Kim, Haotian Zhang, and David Jurgens. 2025a. Leveraging multilingual training for authorship representation: Enhancing generalization across languages and domains. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 34855–34880, Suzhou, China. Association for Computational Linguistics. Junghwan Kim, Haotian Zhang, and David Jurgens. 2025b. Leveraging multilingual training for authorship representation: Enhancing generalization across languages and domains. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 34867–34892, Suzhou, China. Association for Computational Linguistics. Bernd Kortmann, Kerstin Lunkenheimer, and Katharina Ehret. 2020. The electronic world atlas of varieties of English. Available online at https://ewave-atl as.org/. Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020. Reformulating Unsupervised Style Transfer as Paraphrase Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 737–762, Online. Association for Computational Linguistics. Taja Kuzman, Igor Mozetič, and Nikola Ljubešić. 2023. Automatic Genre Identification for Robust Enrichment of Massive Text Collections: Investigation of Classification Methods in the Era of Large Language Models. Machine Learning and Knowledge Extraction, 5(3):1149–1175.

Veronika Laippala, Samuel Rönnqvist, Miika Oinonen, Aki-Juhani Kyröläinen, Anna Salmela, Douglas Biber, Jesse Egbert, and Sampo Pyysalo. 2023. Register identification from the unrestricted open Web using the Corpus of Online Registers of English. Language Resources and Evaluation, 57(3):1045–1079. Bruce W. Lee and Jason Lee. 2023a. LFTK: Handcrafted features in computational linguistics. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), pages 1–19, Toronto, Canada. Association for Computational Linguistics. Bruce W. Lee and Jason Lee. 2023b. LFTK: Handcrafted features in computational linguistics. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), pages 1–19, Toronto, Canada. Association for Computational Linguistics. Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2025. NV-embed: Improved techniques for training LLMs as generalist embedding models. In The Thirteenth International Conference on Learning Representations. Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, and Yue Zhang. 2024. MAGE: Machine-Generated text detection in the wild. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 36–53. Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards general text embeddings with multi-stage contrastive learning. Preprint, arXiv:2308.03281. Xiong Liu, Jeffrey Hancock, Guangfan Zhang, Roger Xu, David Markowitz, and Natalya Bazarova. 2012. Exploring linguistic features for deception detection in unstructured text. In Hawaii International Conference on System Sciences. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach. Preprint, arXiv:1907.11692. Yiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, and LouisPhilippe Morency. 2021. StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2116–2138, Online. Association for Computational Linguistics. Hieu Man, Van-Cuong Pham, Nghia Trung Ngo, Franck Dernoncourt, and Thien Huu Nguyen. 2026a. Explainable disentangled representation learning for generalizable authorship attribution in the era of generative ai. Preprint, arXiv:2604.21300.

Hieu Man, Van-Cuong Pham, Nghia Trung Ngo, Franck Dernoncourt, and Thien Huu Nguyen. 2026b. Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AI. arXiv preprint. ArXiv:2604.21300 [cs]. Frederick Mosteller and David L. Wallace. 1963. Inference in an authorship problem: A comparative study of discrimination methods applied to the authorship of the disputed Federalist Papers. Journal of the American Statistical Association, 58(302):275–309. Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. MTEB: Massive text embedding benchmark. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2014–2037, Dubrovnik, Croatia. Association for Computational Linguistics. Terttu Nevalainen, Helena Raumolin-Brunberg, Samuli Kaislaniemi, Mikko Laitinen, Minna Nevala, Arja Nurmi, Minna Palander-Collin, Tanja Säily, and Anni Sairio. 2021. CEECES 1 = Corpus of Early English Correspondence Extension Sampler part 1. Version 2. XML conversion and encoding by Lassi Saario. Terttu Nevalainen, Helena Raumolin-Brunberg, Samuli Kaislaniemi, Mikko Laitinen, Minna Nevala, Arja Nurmi, Minna Palander-Collin, Tanja Säily, and Anni Sairio. 2022. CEECES 2 = Corpus of Early English Correspondence Extension Sampler part 2. XML conversion and encoding by Lassi Saario. Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 188–197, Hong Kong, China. Association for Computational Linguistics. Andrea Nini. 2023. A Theory of Linguistic Individuality for Authorship Analysis. Elements in Forensic Linguistics. Cambridge University Press. Team Olmo et al. 2026. arXiv:2512.13961.

Olmo 3.

Preprint,

Ajay Patel, Delip Rao, Ansh Kothary, Kathleen McKeown, and Chris Callison-Burch. 2023. Learning interpretable style embeddings via prompting LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 15270–15290, Singapore. Association for Computational Linguistics. Ajay Patel, Jiacheng Zhu, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathleen McKeown, and Chris Callison-Burch. 2025. StyleDistance: Stronger content-independent style embeddings with synthetic parallel examples. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers),

pages 8662–8685, Albuquerque, New Mexico. Association for Computational Linguistics. Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. The FineWeb Datasets: Decanting the web for the finest text data at scale. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Fuchun Peng, Dale Schuurmans, Shaojun Wang, and Vlado Keselj. 2003. Language independent authorship attribution using character level language models. In Proceedings of the Tenth Conference on European Chapter of the Association for Computational Linguistics - Volume 1, EACL ’03, page 267–274, USA. Association for Computational Linguistics. Justin Qiu, Jiacheng Zhu, Ajay Patel, Marianna Apidianaki, and Chris Callison-Burch. 2025. mStyleDistance: Multilingual Style Embeddings and their Evaluation. In Findings of the Association for Computational Linguistics: ACL 2025, pages 16917–16931, Vienna, Austria. Association for Computational Linguistics. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-Context Retrieval-Augmented Language Models. Transactions of the Association for Computational Linguistics, 11:1316–1331. Nils Reimers and Iryna Gurevych. 2019. SentenceBERT: Sentence Embeddings using Siamese BERTNetworks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics. Rafael Rivera Soto, Kailin Koch, Aleem Khan, Barry Y. Chen, Marcus Bishop, and Nicholas Andrews. 2023. Few-shot detection of machine-generated text using style representations. In The Twelfth International Conference on Learning Representations (ICLR). Rafael Rivera Soto, Olivia Elizabeth Miano, Juanita Ordonez, Barry Y. Chen, Aleem Khan, Marcus Bishop, and Nicholas Andrews. 2021. Learning universal authorship representations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 913–919, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Andrew Rosenberg and Julia Hirschberg. 2007. Vmeasure: A conditional entropy-based external cluster evaluation measure. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural

Language Learning (EMNLP-CoNLL), pages 410– 420, Prague, Czech Republic. Association for Computational Linguistics. Jonathan Schler, Moshe Koppel, Shlomo Argamon, and James Pennebaker. 2006. Effects of Age and Gender on Blogging. pages 199–205, Stanford, USA. Serge Sharoff. 2018. Functional Text Dimensions for the annotation of web corpora. Corpora, 13(1):65– 95. Efstathios Stamatatos, Walter Daelemans, Ben Verhoeven, Martin Potthast, Benno Stein, Patrick Juola, Miguel A Sanchez-Perez, Alberto Barrón-Cedeño, et al. 2014. Overview of the author identification task at PAN 2014. In CEUR workshop proceedings, volume 1180, pages 877–897. CEUR-WS. Efstathios Stamatatos, Krzysztof Kredens, Piotr Pezik, Annina Heini, Janek Bevendorff, Benno Stein, and Martin Potthast. 2023. Overview of the authorship verification task at PAN 2023. In CLEF 2023: Conference and Labs of the Evaluation Forum. Efstathios Stamatatos, Martin Potthast, Francisco Rangel, Paolo Rosso, and Benno Stein. 2015. Overview of the PAN/CLEF 2015 Evaluation Lab. In Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 518–538. Springer, Cham. Efstathios Stamatatos, Francisco Rangel, Michael Tschuggnall, Benno Stein, Mike Kestemont, Paolo Rosso, and Martin Potthast. 2018a. Overview of PAN 2018. In Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 267–285. Springer, Cham. Efstathios Stamatatos, Francisco Rangel, Michael Tschuggnall, Benno Stein, Mike Kestemont, Paolo Rosso, and Martin Potthast. 2018b. Overview of pan 2018: Author identification, author profiling, and author obfuscation. In International conference of the cross-language evaluation forum for european languages, pages 267–285. Springer. Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, and Han Xiao. 2025. Jina embeddings v3: Multilingual text encoder with low-rank adaptations. In Advances in Information Retrieval: 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy, April 6–10, 2025, Proceedings, Part V, page 123–129, Berlin, Heidelberg. Springer-Verlag. Sowmya Vajjala and Ivana Lučić. 2018. OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification. In Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications, pages 297–304, New Orleans, Louisiana. Association for Computational Linguistics. Klára Venglařová and Vladimír Matlach. 2024. Beyond content: discriminatory power of function words in

text type classification. Digital Scholarship in the Humanities, 39(2):765–789. Henrique Schechter Vera, Sahil Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, Weiyi Wang, Zhe Li, Gus Martins, Jinhyuk Lee, Mark Sherwood, Juyeong Ji, Renjie Wu, Jingxiao Zheng, Jyotinder Singh, Abheesht Sharma, Divyashree Sreepathihalli, Aashi Jain, Adham Elarabawy, AJ Co, Andreas Doumanoglou, Babak Samari, Ben Hora, Brian Potetz, Dahun Kim, Enrique Alfonseca, Fedor Moiseev, Feng Han, Frank Palma Gomez, Gustavo Hernández Ábrego, Hesen Zhang, Hui Hui, Jay Han, Karan Gill, Ke Chen, Koert Chen, Madhuri Shanbhogue, Michael Boratko, Paul Suganthan, Sai Meher Karthik Duddu, Sandeep Mariserla, Setareh Ariafar, Shanfeng Zhang, Shijie Zhang, Simon Baumgartner, Sonam Goenka, Steve Qiu, Tanmaya Dabral, Trevor Walker, Vikram Rao, Waleed Khawaja, Wenlei Zhou, Xiaoqi Ren, Ye Xia, Yichang Chen, YiTing Chen, Zhe Dong, Zhongli Ding, Francesco Visin, Gaël Liu, Jiageng Zhang, Kathleen Kenealy, Michelle Casbon, Ravin Kumar, Thomas Mesnard, Zach Gleicher, Cormac Brick, Olivier Lacombe, Adam Roberts, Qin Yin, Yunhsuan Sung, Raphael Hoffmann, Tris Warkentin, Armand Joulin, Tom Duerig, and Mojtaba Seyedhosseini. 2025. Embeddinggemma: Powerful and lightweight text representations. Preprint, arXiv:2509.20354. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2024a. Text embeddings by weakly-supervised contrastive pre-training. Preprint, arXiv:2212.03533. Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, Thomas Arnold, Alham Aji, Nizar Habash, Iryna Gurevych, and Preslav Nakov. 2024b. M4: Multi-generator, multi-domain, and multilingual black-box machine-generated text detection. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1369– 1407, St. Julian’s, Malta. Association for Computational Linguistics. Joe H. Ward Jr. 1963. Hierarchical grouping to optimize an objective function. Journal of the American Statistical Association, 58(301):236–244. Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference. In Proceedings of the 63rd Annual Meeting of the

Association for Computational Linguistics (Volume 1: Long Papers), pages 2526–2547, Vienna, Austria. Association for Computational Linguistics.

Task: Align the order of set B to that of set A w.r.t. style: Set A

Janith Weerasinghe and Rachel Greenstadt. 2020. Feature vector difference based neural network and logistic regression models for authorship verification. Notebook for PAN at CLEF.

Set B

Anna Wegmann, Cristina Aggazzotti, Rafael Rivera Soto, and Dong Nguyen. 2026. A Survey on Representing Linguistic Style: Challenges and Opportunities. Anna Wegmann and Dong Nguyen. 2021. Does it capture STEL? a modular, similarity-based linguistic style evaluation framework. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7109–7130, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Anna Wegmann, Marijn Schraagen, and Dong Nguyen. 2022. Same author or just same topic? towards content-independent style representations. In Proceedings of the 7th Workshop on Representation Learning for NLP, pages 249–268, Dublin, Ireland. Association for Computational Linguistics. Junchao Wu, Runzhe Zhan, Derek F Wong, Shu Yang, Xinyi Yang, Yulin Yuan, and Lidia S Chao. 2024. Detectrl: Benchmarking llm-generated text detection in real-world scenarios. Advances in Neural Information Processing Systems, 37:100369–100401. Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-pack: Packaged resources to advance general chinese embedding. Preprint, arXiv:2309.07597. Wei Xu, Alan Ritter, Bill Dolan, Ralph Grishman, and Colin Cherry. 2012. Paraphrasing for Style. In Proceedings of COLING 2012, pages 2899–2914, Mumbai, India. The COLING 2012 Organizing Committee. Helen Yannakoudakis, Ted Briscoe, and Ben Medlock. 2011. A New Dataset and Method for Automatically Grading ESOL Texts. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 180–189, Portland, Oregon, USA. Association for Computational Linguistics. Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: Open Pre-trained Transformer language models. Preprint, arXiv:2205.01068. Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren

1. r u a fan of them or something?

2. Are you one of their fans?

Oh, and also that young physician got an unflattering haircut

Oh yea and that young dr got a bad haircut

Solution: B2, B1

Figure 2: Order Alignment Example. Set A is written in an informal and a formal style, respectively. Set B is written in the reverse stylistic order. The task is to reorder B to match A’s style sequence. We aim for sets to only include sentences showing the same content. The “distractor” variant is signified by the modifications in grey and red. Figure was taken and slightly modified from Wegmann et al. (2022).

Zhou. 2025. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176.

Appendix A

Full aggregate results

The full tables of results for the experiments in §5.1, §5.2, and §5.3 are shown in Tab. 6, Tab. 7, and Tab. 8, respectively.

B

Task examples

We display an example of the order alignment task in Figure 2.

Model Num. Datasets

Clustering All-to-All Predefined Order Align. Retrieval Probing STEB score 53 24 26 11 4 4 122

Style specific LUAR-MUD LUAR-CRUD StyleDistance StyleDistance (Synthetic) mStyleDistance MSR CISR STAR LISA

20.03 19.65 18.46 13.46 7.35 16.67 15.28 25.92 12.74

62.33 61.00 59.82 57.02 53.83 64.34 59.71 64.56 60.91

72.33 73.12 66.21 57.67 53.71 71.01 66.43 73.28 62.40

20.81 19.80 44.02 42.76 31.38 21.85 33.81 23.93 16.48

75.67 77.64 49.61 31.53 20.04 63.49 45.87 66.96 42.96

53.04 53.75 57.87 56.69 45.96 51.01 52.54 49.86 46.38

50.70 50.82 49.33 43.19 35.38 48.06 45.61 50.75 40.31

Semantic embedders all-mpnet-base-v2 GTE-base-en-v1.5 GTE-large-en-v1.5 E5-base-v2 E5-large-v2 BGE-base-en-v1.5 BGE-large-en-v1.5 Jina Embeddings v3 Qwen3-Embedding-8B

8.48 10.76 11.48 9.28 11.06 7.93 7.15 7.90 2.67

60.41 60.04 60.99 60.40 60.20 59.40 59.84 60.58 51.61

63.87 60.33 62.39 63.42 64.70 62.97 63.81 64.60 52.79

14.62 15.62 16.20 15.26 15.82 15.82 16.23 17.11 23.11

48.99 49.36 49.60 49.62 49.91 51.44 48.79 47.35 35.17

45.64 50.64 48.87 52.13 49.03 49.45 47.23 45.62 42.96

40.34 41.13 41.59 41.69 41.79 41.17 40.51 40.53 34.72

Masked language models BERT-large-cased BERT-large-uncased RoBERTa-base RoBERTa-large DeBERTa-v3-base DeBERTa-v3-large ModernBERT-base ModernBERT-large

12.02 11.18 12.59 13.21 19.42 21.44 14.05 16.58

60.99 60.44 60.97 60.97 58.35 59.37 59.28 60.19

67.11 66.92 66.49 68.79 66.82 70.45 64.13 65.38

21.57 21.18 22.47 20.71 23.41 24.23 19.20 20.56

58.88 60.23 50.63 61.07 55.17 63.24 51.25 52.17

55.16 52.48 61.32 57.92 51.73 53.41 53.91 54.06

45.96 45.41 45.74 47.11 45.81 48.69 43.63 44.82

Causal language models GPT-2 XL OPT-1.3B Qwen2-0.5B Qwen3-0.6B-Base Qwen3.5-0.8B-Base Qwen3.5-2B-Base Qwen3.5-4B-Base

4.42 3.83 3.77 3.73 4.21 3.97 4.23

51.68 52.40 51.59 51.43 52.73 52.46 52.80

53.49 55.59 52.78 53.14 57.23 57.09 57.41

26.32 24.08 25.51 26.01 25.07 24.61 24.88

28.82 39.07 25.70 23.67 32.83 35.13 37.15

48.66 45.82 42.34 41.78 43.30 42.61 43.02

35.57 36.80 33.62 33.29 35.89 35.98 36.58

Pre-defined features Function words TFIDF FineWeb 1-2-grams TFIDF FineWeb 1-3-grams TFIDF Reddit 1-2-grams TFIDF Reddit 1-3-grams Stylometric NeuroBiber

7.17 6.78 6.37 6.79 6.52 9.67 8.24

55.17 57.18 56.97 56.93 56.68 56.88 55.10

60.88 64.45 64.45 64.45 64.25 60.48 57.27

14.51 13.72 13.67 13.84 13.84 25.87 14.13

43.14 50.95 49.00 49.77 50.26 39.52 9.06

37.98 55.74 55.51 55.54 55.34 53.30 51.64

36.47 41.47 40.99 41.22 41.15 40.95 32.57

Table 6: Full STEB (operational) results (×100). Bold = best, underline = 2nd best per column. Metrics: V-measure (Clustering), AUC (All-to-All and Predefined Pair Classification), distractor accuracy (Order Alignment), MRR (Retrieval), and average accuracy (Probing). Scores are macro-averaged across automatically discovered dataset clusters within each task, then averaged across tasks to produce the overall STEB score. Num. Datasets counts the datasets contributing to each task; a dataset may contribute to more than one task.

Model

ATD

ATD (Adv.)

Authorship Verification

AR

Avg.

63.24 66.96 75.67 55.17 77.64 49.61 52.17 45.87 61.07 60.23 58.88 51.25 31.53 63.49 50.63 50.26 49.77 50.95 49.00 49.36 39.52 49.91 43.14 42.96 49.62 49.60 51.44 48.99 47.35 48.79 37.15 39.07 35.13 32.83 20.04 35.17 28.82 9.06 25.70 23.67

68.14 66.82 63.85 63.34 62.95 55.06 51.84 51.14 47.71 44.34 43.11 42.69 39.33 38.72 38.19 33.91 33.13 32.74 32.47 31.89 31.84 31.69 31.42 31.41 31.26 30.77 30.66 30.32 30.27 30.17 26.39 25.50 25.42 25.30 23.95 23.74 23.57 22.90 21.83 21.39

Overall Easy Medium Hard Top 5 deberta-v3-large STAR LUAR-MUD deberta-v3-base LUAR-CRUD StyleDistance ModernBERT-large CISR roberta-large bert-large-uncased bert-large-cased ModernBERT-base StyleDistance (Synthetic) MSR roberta-base TFIDF Reddit 1-3-grams TFIDF Reddit 1-2-grams TFIDF FineWeb 1-2-grams TFIDF FineWeb 1-3-grams gte-base-en-v1.5 Stylometric e5-large-v2 Function words LISA e5-base-v2 gte-large-en-v1.5 bge-base-en-v1.5 all-mpnet-base-v2 jina-embeddings-v3 bge-large-en-v1.5 Qwen3.5-4B-Base opt-1.3b Qwen3.5-2B-Base Qwen3.5-0.8B-Base mStyleDistance Qwen3-Embedding-8B gpt2-xl NeuroBiber Qwen2-0.5B Qwen3-0.6B-Base

36.03 22.79 28.20 29.55 25.20 30.10 26.39 22.78 23.13 17.53 16.61 24.85 22.76 17.38 24.83 12.37 11.93 12.77 13.36 14.30 24.10 9.48 14.25 11.61 9.33 8.36 6.13 6.25 6.56 5.97 9.01 8.13 8.27 9.25 15.85 6.95 12.52 20.34 8.29 8.29

100.00 100.00 75.01 99.24 72.61 70.08 60.69 64.33 35.31 30.22 27.57 27.94 44.03 0.12 8.29 7.17 4.64 0.81 1.25 0.16 0.87 0.20 5.20 6.06 0.11 0.20 0.01 0.04 0.01 0.01 1.86 0.03 1.27 1.80 5.48 0.11 0.09 3.18 0.38 0.34

73.29 77.52 76.54 69.40 76.35 70.44 68.10 71.57 71.34 69.39 69.36 66.71 59.00 73.88 69.02 65.83 66.20 66.44 66.28 63.76 62.87 67.16 63.10 65.03 65.97 64.93 65.05 66.00 67.15 65.90 57.55 54.78 57.02 57.33 54.41 52.75 52.84 59.02 52.94 53.29

77.68 88.37 86.84 73.46 86.60 75.52 85.02 79.39 85.49 82.49 85.59 83.86 64.82 88.92 84.16 73.90 75.72 77.03 75.46 87.36 77.32 88.54 68.40 84.41 87.83 86.67 86.64 87.36 88.83 87.12 62.97 55.99 61.82 62.35 60.49 62.78 53.91 71.14 61.94 61.53

62.92 76.02 73.62 59.80 75.92 69.17 64.09 70.74 66.43 65.10 64.93 61.03 55.28 73.43 64.37 60.50 60.34 60.69 61.09 67.91 50.71 68.08 58.25 66.71 67.96 67.65 65.66 66.33 68.93 66.79 51.11 51.32 50.61 51.71 49.63 52.41 51.43 55.94 50.58 50.27

55.87 58.26 61.10 54.75 61.23 62.18 50.75 65.56 52.61 50.94 49.53 49.89 52.62 55.99 50.31 50.98 50.79 51.03 51.12 43.27 46.58 48.21 52.30 47.86 46.38 44.15 43.76 44.30 44.80 44.15 47.61 50.35 47.77 47.86 50.53 47.54 49.49 52.30 48.08 47.91

Table 7: Scores across application clusters, sorted in descending order. ATD = AI-Text Detection, AR = Authorship Retrieval. Authorship Verification: Overall aggregates a superset of datasets (PAN13–15, PAN20–21, Enron, and the PAN22–26 style-change benchmarks), while Easy/Medium/Hard are computed over the PAN22–26 style-change subset only—they do not average to the Overall column. To avoid double-counting, only Overall contributes to Avg. Bold = best, underline = 2nd best per column (across all models). Each column reports the macro-averaged score (×100) within the corresponding manual cluster.

Model deberta-v3-large StyleDistance roberta-large bert-large-cased CISR deberta-v3-base STAR roberta-base bert-large-uncased LUAR-CRUD LUAR-MUD TFIDF Reddit 1-2-grams TFIDF Reddit 1-3-grams MSR TFIDF FineWeb 1-3-grams TFIDF FineWeb 1-2-grams e5-large-v2 e5-base-v2 ModernBERT-base LISA Stylometric StyleDistance (Synthetic) bge-base-en-v1.5 ModernBERT-large jina-embeddings-v3 all-mpnet-base-v2 bge-large-en-v1.5 mStyleDistance gte-large-en-v1.5 Qwen3.5-0.8B-Base Qwen3.5-2B-Base Qwen3.5-4B-Base Function words gte-base-en-v1.5 opt-1.3b NeuroBiber Qwen3-Embedding-8B Qwen3-0.6B-Base Qwen2-0.5B gpt2-xl

Pair Class. Retrieval 74.03 72.40 71.91 71.33 70.08 71.85 71.32 66.01 69.36 72.37 69.50 63.02 62.61 69.43 61.94 62.42 67.05 66.20 69.09 62.67 67.38 64.03 62.86 66.95 64.02 60.04 62.93 59.94 62.78 56.83 55.72 55.93 54.25 58.53 55.30 54.98 54.56 54.64 53.43 48.75

47.77 44.45 44.02 39.11 37.93 35.42 35.80 39.70 34.92 30.07 32.38 38.52 37.34 28.38 34.53 31.98 24.60 24.99 20.55 26.70 19.37 19.46 18.90 13.80 16.23 19.17 14.61 14.69 10.16 15.68 16.41 14.52 15.16 10.41 11.55 10.59 10.51 10.41 9.41 12.55

Avg. 60.90 58.42 57.96 55.22 54.01 53.63 53.56 52.86 52.14 51.22 50.94 50.77 49.98 48.90 48.24 47.20 45.83 45.59 44.82 44.68 43.37 41.74 40.88 40.37 40.12 39.61 38.77 37.32 36.47 36.25 36.07 35.23 34.71 34.47 33.43 32.79 32.54 32.53 31.42 30.65

Table 8: Full Multilingual STEB results (×100), sorted in descending order. Pair Class. averages AUC across 13 PAN13/14/15 authorship verification datasets (Dutch, Greek, Spanish); Retrieval averages MRR across 4 PAN18 cross-domain AA datasets (French, Italian, Polish, Spanish). Bold = best, underline = 2nd best per column (across all models).

Object of Study

Ling. Feat. Content-Ind. Avg. 7

Order Al. † 11

67

49.94 45.26 47.76 53.06 38.27 50.58 49.85 49.12 41.06

74.02 70.87 65.85 67.66 56.17 70.50 70.37 66.52 64.02

44.23 46.10 32.31 14.08 38.32 11.29 10.57 12.35 7.10

56.06 54.08 48.64 44.93 44.25 44.12 43.59 42.66 37.39

43.96 57.80 58.53 56.56 58.25 57.35 57.25 57.26 57.50

34.77 40.67 41.36 40.17 40.07 40.20 40.94 41.27 40.82

57.68 69.27 67.54 65.77 65.79 65.54 63.80 63.40 62.18

24.10 5.11 5.22 5.31 5.22 5.28 5.53 5.45 5.06

38.85 38.35 38.04 37.08 37.03 37.01 36.75 36.71 36.02

43.37 40.16 42.07 46.08 41.23 42.28 37.71 44.03

68.26 59.83 66.21 62.29 64.12 60.13 58.98 64.81

49.24 47.51 48.94 47.01 48.91 47.20 45.58 48.17

69.56 77.95 75.48 68.28 73.58 73.38 72.90 68.55

18.72 10.27 9.86 18.49 9.22 10.35 9.64 10.89

45.84 45.24 44.76 44.59 43.90 43.64 42.71 42.54

28.29 28.24 27.29 28.33 28.43 28.26 28.76

32.34 34.66 34.00 33.39 34.70 35.77 30.68

45.08 39.32 40.83 46.08 47.35 38.48 46.92

38.08 37.26 35.31 37.89 38.07 36.30 36.92

60.61 60.46 58.54 59.49 58.63 57.99 57.85

26.03 26.76 29.14 25.49 25.72 28.12 24.29

41.57 41.49 41.00 40.96 40.81 40.80 39.69

32.44 31.56 31.67 31.70 31.84 31.75 31.46

35.49 27.06 27.71 26.31 25.33 24.81 28.07

51.19 53.12 57.98 58.70 58.05 57.64 34.04

41.12 40.31 39.57 39.32 39.04 38.76 34.62

67.36 63.21 66.90 67.13 66.55 66.68 64.11

24.89 9.52 5.66 5.55 5.62 5.47 7.99

44.46 37.68 37.38 37.33 37.07 36.97 35.58

Model Num. Datasets

Genre Register Time Demo. Dialect Idiolect Avg. 4 7 3 3 3 29 49

Style specific StyleDistance StyleDistance (Synth.) CISR STAR mStyleDistance LUAR-MUD LUAR-CRUD MSR LISA

62.84 55.45 62.43 68.68 49.41 66.67 65.80 66.39 56.89

40.01 38.81 36.09 39.79 32.83 38.72 38.14 39.08 37.95

46.21 43.27 41.52 47.25 35.04 42.14 41.30 41.98 37.38

33.67 33.13 33.03 40.99 32.25 35.53 35.32 37.05 32.95

56.91 55.64 54.76 49.39 42.83 44.30 41.52 41.52 27.19

60.03 45.26 58.72 72.24 37.23 76.11 76.99 68.69 54.00

Semantic embedders Qwen3-Embedding-8B E5-base-v2 E5-large-v2 GTE-base-en-v1.5 BGE-base-en-v1.5 BGE-large-en-v1.5 Jina Embeddings v3 GTE-large-en-v1.5 all-mpnet-base-v2

39.57 55.91 57.58 53.62 54.69 56.15 56.24 57.09 56.07

33.88 37.26 37.70 36.51 36.67 37.31 37.51 36.67 37.55

37.26 39.72 40.81 39.44 36.96 38.23 39.77 41.20 39.03

26.92 30.50 31.38 31.76 30.75 29.71 31.10 31.38 31.47

27.05 22.85 22.16 23.10 23.10 22.47 23.74 24.01 23.31

Masked language models DeBERTa-v3-large RoBERTa-base RoBERTa-large DeBERTa-v3-base BERT-large-cased ModernBERT-large ModernBERT-base BERT-large-uncased

63.26 63.79 65.84 58.34 66.24 63.99 62.23 62.24

38.25 42.52 40.01 36.63 41.99 39.77 38.76 39.90

46.94 45.94 46.43 44.97 45.11 44.29 43.14 45.27

35.37 32.84 33.07 33.72 34.79 32.75 32.67 32.80

Causal language models Qwen3.5-0.8B-Base Qwen2-0.5B GPT-2 XL Qwen3.5-2B-Base Qwen3.5-4B-Base Qwen3-0.6B-Base OPT-1.3B

46.66 44.76 41.36 45.71 44.42 42.81 43.20

35.59 36.01 32.22 34.81 34.62 34.26 33.19

40.51 40.57 36.15 39.04 38.93 38.22 38.79

Pre-defined features Stylometric Function words TFIDF Reddit 1-2-grams TFIDF FineWeb 1-2-grams TFIDF Reddit 1-3-grams TFIDF FineWeb 1-3-grams NeuroBiber

54.60 51.00 49.91 49.19 48.90 48.49 47.31

38.59 35.58 33.96 33.79 33.75 33.55 32.38

34.43 43.56 36.19 36.20 36.37 36.30 34.49

Table 9: Full STEB (definitional) results. Results are grouped into three clusters after Wegmann et al. (2026) and averaged: Object of Study (Genre, Register, Time, Demographics, Dialect, Idiolect), Linguistic Features, and Content-Independence. The Avg. column under Object of Study reports the mean of the six objects of study. Models are sorted by overall Avg. (descending). The datasets used for each column are given in App. Tab. 13. Order Alignment† uses the distractor variant of the order-alignment task for the respective datasts and no other task for that dataset. The other columns use all available tasks for the included dataset and the acc variant of order alignment. Bold = best, underline = 2nd best per column (across all models).

C

Dataset details

Table 10 summarizes which stylistic attribute each STEB dataset is intended to probe, using the seven categories around which this section is organized. A check mark indicates the category the dataset is grouped under in our analyses; some datasets could plausibly fit additional categories (e.g. the Corpus of Diverse Styles spans dialect, time, and genre) but we keep a single assignment for clarity.

Table 10: STEB datasets by intended stylistic attribute. Feat. = linguistic feature, Dial. = dialect, Reg. = register, Gen. = genre, Demo. = demographics, Auth. = authorship. Each row receives a single check mark for the bucket the dataset is grouped under in this appendix; a few datasets span multiple attributes but are assigned to a single bucket for clarity. Dataset

Feat. Dial. Reg. Gen. Time Demo. Auth.

STEL_feature SynthSTEL_feature StylePTB probing_amazon probing_blog probing_reddit probing_stackexchange

✓ ✓ ✓ ✓ ✓ ✓ ✓

twitter_aave_sae endive eWAVE STEL_register SynthSTEL_register OneStopEnglishCorpus graded_formality ASSET wikipedia_politeness stackexchange_politeness enron_spam sms_spam telegram-spam-ham hate_speech hate_speech_and_offensive_language corpus-of-diverse-styles core x_genre masc_text_genre parallel_shakespeare ceeces bible_versions fce_l1 blog_age blog_gender pastel_age pastel_education pastel_gender pastel_ethnic pastel_politics pastel_tod pan13_authorship_verification_english_test pan14_authorship_verification_corpus1_english_essays_test pan14_authorship_verification_corpus1_english_novels_test pan14_authorship_verification_corpus2_english_essays_test pan14_authorship_verification_corpus2_english_novels_test pan15_authorship_verification_english_test pan20_authorship_verification_test pan21_authorship_verification_test enron_authorship_corpus pan18_style_change pan22_style_change_basic pan22_style_change_advanced pan22_style_change_sentence pan23_style_change_easy pan23_style_change_medium pan23_style_change_hard pan24_style_change_easy pan24_style_change_medium pan24_style_change_hard pan25_style_change_easy pan25_style_change_medium pan25_style_change_hard pan26_style_change_easy pan26_style_change_medium pan26_style_change_hard

✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

Dataset pan18_cross_domain_authorship_attribution_english amazon fanfiction stackexchange_retrieval gede_essay_detection machine_text_detection_DetectRL_arxiv machine_text_detection_DetectRL_direct_prompt machine_text_detection_DetectRL_writing_prompt machine_text_detection_DetectRL_xsum machine_text_detection_DetectRL_yelp_review machine_text_detection_DetectRL_paraphrase_attacks machine_text_detection_DetectRL_perturbation_attacks machine_text_detection_DetectRL_prompt_attacks machine_text_detection_M4_arxiv machine_text_detection_M4_peerread machine_text_detection_M4_reddit machine_text_detection_M4_wikihow machine_text_detection_M4_wikipedia machine_text_detection_MAGE_cmv machine_text_detection_MAGE_eli5 machine_text_detection_MAGE_hswag machine_text_detection_MAGE_roct machine_text_detection_MAGE_sci_gen machine_text_detection_MAGE_squad machine_text_detection_MAGE_tldr machine_text_detection_MAGE_wp machine_text_detection_MAGE_xsum machine_text_detection_MAGE_yelp machine_text_detection_PAN24_news machine_text_detection_PAN25_26_essays machine_text_detection_PAN25_26_fiction machine_text_detection_PAN25_26_news machine_text_detection_PAN25_collaborative

Feat. Dial. Reg. Gen. Time Demo. Auth. ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

C.1

Feature datasets

STEL_feature added as order alignment task. We split the publicly available STEL version (Wegmann and Nguyen, 2021; Wegmann et al., 2022) into registers (cf. Section C.3) and linguistic features (i.e., number substitution, contraction or emoji use characteristic). STEL_feature consist of ≈ 400 instances total. The STEL characteristics are publicly released. Note that in order alignment tasks we pair each style transfer pair with every other pair from the same style category—thus, we do not use the predefined pairings from STEL or SynthSTEL. SynthSTEL_feature added as order alignment task. We split the SynthSTEL dataset (Patel et al., 2025) into 8 register (cf. Section C.3) and 32 linguistic features (i.e., everything else, including features like All Lower Case / Proper Capitalization). We combine the train and test split, leading to 100 pairs per feature. SynthSTEL is released under the MIT license per its HuggingFace dataset card at https://huggingface.co/datasets/StyleDis tance/synthstel. StylePTB added as order alignment task. StylePTB (Lyu et al., 2021) is a style-transfer benchmark with ≈ 60k parallel pairs over 21 individual styles. We keep 15 and remove six out of the 21 dimensions as some change meaning (e.g., antonym replacement) or have no consistent style change within the same dimension (e.g., synonym replacement uses WordNet without clear stylistic commonalities in the selected synonyms like market replaced with marketplace and years with twelvemonths). Note that the distribution is heavily skewed: to_future has ≈ 7k pairs while pp_back_to_front has only 7. To keep as many categories as possible, we combine the train/dev/test partitions. The text is provided in tokenized form: we do not detokenise. The StylePTB transformations are released under CC BY 4.0 (https://github.com/lvyiwei1/StylePTB /blob/master/LICENSE). Probing added as probing task across four domains (probing_amazon, probing_blog, probing_reddit, probing_stackexchange). We extract 46 foundation-level linguistic features per document using the LFTK toolkit (Lee and Lee, 2023b): 12 token-level counts (words, stop-words, punctuation, syllables, characters, plus token frequency norms) and 34 POS-tag

counts (total and unique counts for each of 17 Universal Dependencies POS tags). Each feature is independently quantile-binned into discrete labels and the resulting per-feature classification task is solved with a small MLP probe on top of frozen embeddings. The reported scores are macroaveraged across the 46 features. The underlying source texts are sampled from publicly available Amazon product reviews (Ni et al., 2019), blog posts (via Kaggle https://www.kaggle.com/d atasets/rtatman/blog-authorship-corpus), Reddit comments (via HuggingFace dataset AnnaWegmann/StyleEmbeddingData), and Stack Exchange posts, yielding 27, 584, 39, 111, 16, 447, and 9, 253 records respectively after binning and balancing. The scripts used to create the probing datasets will be released publicly on our GitHub repository. C.2

Dialectal datasets

Twitter AAE added as order alignment task. We use a set of parallel constructed tweets written originally in AAE variety and rewritten in SAE by annotators(Groenwold et al., 2020). The dataset consists of ≈ 2k instances. It can be accessed via the EMNLP 2020 supplementary archive at https://aclanthology.org/attachments/202 0.emnlp-main.473.OptionalSupplementaryM aterial.zip. There is no license file; we treat it as research-use only consistent with its release as paper supplementary material. EnDive added as clustering and all-to-all pair classification task. We add parallel sentences from the EnDive benchmark (Gupta et al., 2025) for Jamaican English, African American English, Colloquial Singaporean English, Indian English and Chicano English across 12 NLU benchmarks7 at https://huggingface.co/datasets/abhaygup ta1266/. The dataset was created using few-shot prompting with examples from native speakers. License information is not provided, but we assume it to be permissive based on the fact that it is meant to be a public benchmarks freely shared on HuggingFace. eWAVE added as clustering and all-to-all pair classification task. We add dialect example sentences from the Electronic World Atlas of Varieties of English (Kortmann et al., 2020) Cross7 specifically logic_bench_yn, logic_bench_mcq svamp mbpp humaneval gsm8k, folio, boolq, copa, multirc, sst-2, wsc

Linguistic Data Formats (CLDF) release at https: //zenodo.org/records/17433568 and https: //github.com/cldf-datasets/ewave. We include 53 varieties ranging from ≈ 20 to ≈ 200 instances. Varieties include Hong Kong English, Minx English, Bahamian Creole and Indian English. The CLDF release is licensed CC BY 4.0. C.3

Register datasets

STEL_register added as order alignment tasks. We split the publicly available8 STEL version (Wegmann and Nguyen, 2021; Wegmann et al., 2022) into registers (i.e., formality and complexity dimension) and linguistic features (cf. C.1). The registers consist of ≈ 200 instances each. SynthSTEL_register added as order alignment task. We split the SynthSTEL dataset (Patel et al., 2025) into 8 register (e.g., formal tone, offensive language, sarcasm) and 32 linguistic features (cf. Section C.1). Note that depending on ones definition of “register”, some categories like positive sentiment expression might not be considered style (Wegmann et al., 2026). We combine the train and test split, leading to 100 pairs per feature. SynthSTEL is released under the MIT license per its HuggingFace dataset card at https://huggingf ace.co/datasets/StyleDistance/synthstel. OneStopEnglish added as order alignment task. We use the OneStopEnglish corpus (Vajjala and Lučić, 2018) at https://github.com/nishk alavallabhi/OneStopEnglishCorpus/. It provides 189 parallel texts at three reading levels (beginner, intermediate, advanced) compiled from news articles rewritten by English teachers to suit the levels of the learners. It is shared with Creative Commons Attribution Share Alike 4.0 International, see https://github.com/nishkal avallabhi/OneStopEnglishCorpus/blob/mast er/LICENSE.markdown. Graded Formality Dataset added as order alignment task. From TweetEval (Barbieri et al., 2020) we take the ≈ 4k Tweets classified as non-offensive at https://huggingface.co/datasets/ca rdiffnlp/tweet_eval/. From SMS Spam (Almeida et al., 2011), we take the ≈ 4k SMS classified not spam at https://huggingface. 8

The full formality dataset of 815 items is not publicly available, so we use the public version of STEL with 100 instances of two pair seach, see https://github.com/nlp soc/STEL/

Level 1 2 3 4 5

Example Lol your always so convincing. Lol, you’re always so convincing. You’re always so convincing. You are always very convincing. You are consistently persuasive.

Table 11: Graded formality example.

co/datasets/ucirvine/sms_spam. We use gpt-5-mini-2025-08-07 rewrite the utterances more formally 4 times, see Figure 3. Both datasets are publicly shared on HuggingFace. An example of the resulting dataset can found in Table 11. An example of the resulting sentences are: “Lol your always so convincing.”, “Lol, you’re always so convincing.”, “You’re always so convincing.”, “You are always very convincing.”, “You are consistently persuasive.”. ASSET added as order alignment task. We use the ASSET corpus (Alva-Manchego et al., 2020), a multi-reference sentence-simplification benchmark covering ≈ 2k English Wikipedia source sentences with ten crowdsourced simplifications per sentence. To turn ASSET into an order-alignment task on the simple axis, we form one length-two ordered pair per source sentence by sampling one of its ten reference simplifications uniformly at random as the simpler element and pairing it with the original Wikipedia sentence as the more-complex element. We combine the validation and test splits on HuggingFace https://huggingface.co/datas ets/facebook/asset. The HuggingFace card shows the license as CC BY SA 4.0. wikipedia_politeness added as clustering and all-to-all pair classification task. We add the Wikipedia Politeness Corpus (Danescu-NiculescuMizil et al., 2013) shared at https://zissou .infosci.cornell.edu/convokit/datasets /wikipedia-politeness-corpus/. The corpus consists of ≈ 4k polite, neutral and impolite texts of Wikipedia editor requests. The corpus is shared with a CC BY license v4.0, see https://convokit.cornell.edu/documentati on/wiki_politeness.html. stackexchange_politeness added as clustering and all-to-all pair classification task. We add the Stack Exchange Politeness Corpus of DanescuNiculescu-Mizil et al. (2013) shared at https://

zissou.infosci.cornell.edu/convokit/data sets/stack-exchange-politeness-corpus/. We use the same labels as wikipedia_politeness. It consists of ≈ 6k Stack Exchange requests. The corpus is shared with a CC BY license v4.0, see https://convokit.cornell.edu/documentati on/stack_politeness.html C.4

Genre datasets

Corpus of Diverse Styles added as clustering, all-to-all pair classification and order alignment task. We use the Corpus of Diverse Styles (Krishna et al., 2020) taken from HuggingFace at https://huggingface.co/datasets/billra y110/corpus-of-diverse-styles. The original dataset totals ≈ 400k records across 11 “styles” (i.e., lyrics, three decade slices of the Corpus of Historical American English, Tweets, African American English Tweets, the writing of James Joyce, the Bible, the Switchboard telephone-speech transcripts, English poetry, and Shakespeare). Note that while we could distribute the labels between dialect, time and genre, we assign everything to genre for simplicity. The code is MIT-licensed (https://github.com/martiansideofthemoo n/style-transfer-paraphrase/blob/master/ LICENSE). CORE added as clustering and all-to-all pair classification task. We add the Corpus of Online Registers of English (CORE) (Laippala et al., 2023) as found at https://github.com/TurkuNLP/CO RE-corpus. We use the sub-labels as the relevant classes for pair classification, i.e., everything but the main labels {IN, NA, HI, LY, SP, IP, ID, OP}. If a document falls in several categories, we have it show up in several. We use documents from any file across the train, dev and test sets. This dataset was licensed under CC BY-SA 4.0. x_genre added as clustering and all-to-all pair classification task. We add the English slice of the X-GENRE genre-classification dataset (Kuzman et al., 2023) shared at https://huggingf ace.co/datasets/TajaKuzmanPungersek/X-G ENRE-text-genre-dataset. X-GENRE merges FTD (Sharoff, 2018) and CORE (Laippala et al., 2023) into a joint 9-label schema using classification models. Compared to CORE, 15 sub-labels and their texts are discarded. The added dataset consists of ≈ 2k English documents. We merge the train/dev/test partitions. X-GENRE is released under CC BY-SA 4.0 on HuggingFace.

MASC added as clustering and all-to-all pair classification task. We add the Manually Annotated Sub-Corpus of American English (MASC) 3.0.0 (Ide et al., 2008, 2013) at http://www.an c.org/data/masc/downloads. MASC is a corpus including texts from 19 genres in American English (e.g., debate transcripts, e-mail, jokes). We create 150 samples per genre. MASC is distributed without license or other restrictions, see https://anc.org/data/masc/about/. C.5

Time datasets

Shakespeare added as order alignment task. We add the parallel Shakespeare corpus of Xu et al. (2012) at https://github.com/cocoxu/Shak espeare. It includes labels for historical (i.e., Shakespearean) and contemporary English, providing ≈20k line-aligned pairs of originals and their modern paraphrases for 17 plays. The license file only includes a request for citation at https://github.com/cocoxu/Shakespeare/bl ob/master/README_LICENSE, thus it is likely to be shared at least for research purposes. CEECES added as clustering and all-to-all pair classification. We add the Corpus of Early English Correspondence Extension Sampler (CEECES), parts 1 and 2 (Nevalainen et al., 2021, 2022). It consists of ≈ 15k paragraphs from ≈ 2.5k historical letters labeled by six 20-year periods from 16801800. Both Zenodo records are licensed under Creative Commons Attribution Non Commercial 4.0 International and Creative Commons Attribution Non Commercial No Derivatives 4.0 International at https://zenodo.org/records/6411789 and https://zenodo.org/records/5887101. Bible versions added as clustering and all-toall pair classification task. We add the parallel Bible corpus of Carlson et al. (2018) shared at https://github.com/keithecarlson/Sty leTransferBibleData. The dataset comprises eight verse-aligned English Bible translations including, for example, the American Standard Version (from 1901) and the Kings James Version (from 1611). We remove verses that are identical (ignoring whitespacing and casing). The dataset includes no specific license, but all included Bibles are considered public domain and the associated paper was released with the Creative Commons Attribution License.

C.6

Demographic datasets

fce_l1 added as clustering and all-to-all pair classification task for distinguishing different native language speakers. We add the the publicly released subset of the Cambridge Learner Corpus’s First Certificate in English exam corpus (CLC FCE) (Yannakoudakis et al., 2011) at https://ilex ir.co.uk/datasets/index.html. It consists of ≈ 2.5k responses for ≈ 1.2k English Exams by students of 16 different native languages. The dataset is released with a non-commercial research and educational license. blog_age, blog_gender added as clustering and all-to-all pair classification tasks. We add the Blog Authorship Corpus (Schler et al., 2006) via HuggingFace at https://huggingface.co/dataset s/barilan/blog_authorship_corpus. We use the validation split to construct age (i.e., age buckets of 10s, 20s and 30s) and gender labeled texts. The corpus is meant for non-commercial research use, see https://u.cs.biu.ac.il/~koppel/Bl ogCorpus.htm. PASTEL_age, PASTEL_education, PASTEL_gender, PASTEL_ethnic, PASTEL_politics added as clustering and all-to-all pair classification tasks. We add PASTEL (Kang et al., 2019) corpus at https://github.com/dykang/PASTEL. It consists of ≈ 8k five-sentence stories written by crowd annotators. Texts are labeled with demographic information of annotators, including six age, eight education, three gender and, nine ethnic and three politics labels. There is no explicit license information provided, but since the dataset is released as a benchmark, we assume at least a intended research-purpose license. C.7

Authorship datasets

PAN Authorship Verification added as predefined pair classification task. We use the official test sets of the PAN authorship-verification shared tasks: PAN13 (Juola and Stamatatos, 2013), PAN14 (Stamatatos et al., 2014), PAN15 (Stamatatos et al., 2015), PAN20 (Bevendorff et al., 2020), and PAN21 (Bevendorff et al., 2021). PAN13, PAN14, and PAN15 contain multiple languages, including Greek, and Spanish in PAN13, Dutch, Greek, and Spanish in PAN14 and PAN15. PAN20 and PAN21 are large-scale cross-domain fanfiction verification benchmarks, we subsample

each to a balanced 500-pair subset (250 sameauthor, 250 different-author) with a fixed seed for reproducibility. All corpora are available at Zenodo, and the download links can be found in the PAN website: https://pan.webis.de/data.html. All PAN datasets are available for research use. PAN Style Change added as predefined pair classification task. We use the publicly datasets of the PAN Style Change Detection shared tasks: PAN18 (Stamatatos et al., 2018a), PAN22 (Bevendorff et al., 2022), PAN23 (Bevendorff et al., 2023), PAN24 (Ayele et al., 2024), PAN2025 (Bevendorff et al., 2025a), and PAN26 (Bevendorff et al., 2026). For each of the aforementioned style cahnge datasets, consecutive paragraphs or sentence pairs (depending on the datasets granularity) are taken to be same- or different-author trials. PAN23–26 ship three difficulty levels (easy, medium, hard) where each subsequent difficulty level has more topic overlap. The PAN26 raw pairs are heavily imbalanced (up to ≈96% same-author in the medium split). We apply per-document stratified down-sampling that keeps all different-author pairs and matches the same-author count. All corpora are available at Zenodo, and the download links can be found in the PAN website: https: //pan.webis.de/data.html. All PAN datasets are available for research use under the CC BY 4.0 license. Enron Authorship Corpus (Halvani, 2017) added as predefined pair classification task. We add the 80-author Enron e-mail authorshipverification corpus at https://prod-dcd-dat asets-public-files-eu-west-1.s3.eu-wes t-1.amazonaws.com/f0527105-2774-423e-8 0ca-cc692b70b6cb. Each verification case is comprised of 5 documents, where 4 documents come from one author and 1 comes from another. We concatenate the 4 documents together separated by a double newline so as to create two pairs of text for the verification trail. The corpora is available for research use under the CC BY 4.0 licence. C.8

AI-text detection datasets

Standard ATD variants use per-generator labels (human plus one label per LLM). Adversarial variants use binary human-vs-LLM labels because the question is whether embeddings still separate human from machine text after evasion techniques have been applied.

Task Type

Example

Solution

Order Alignment

Task: Align the order of the set B to that of set A w.r.t. style: Set A Set B santa was to fat, and the woman was Santa was excessively overweight driving. while the woman was driving. They cannot see anything in the be- Well, they probably can’t see anyginning. thing at first.

B2 B1

Pair Classification

Clustering

Authorship trieval

Re-

Task: Are the two texts written in the same style? Text A Text B Santa was excessively overweight santa was to fat, and the woman was while the woman was driving. driving.

different style

Task: Cluster the texts by style: Text T1. Round him peered Lenehan. T2. O lead me onward to the loneliest shade, T3. All Star Classic Game 1 Orlando 09 Game 1 - West Coast vs East T4. ( Goes up and comes down to center, shrieking and laughing.

T1 → joyce T2 → poetry T3 → tweets T4 → coha_1890

Task: Retrieve the target written by the same author as the query: Query Q. Bought as a Christmas gift but the case is nice

T1 (author 100066)

Targets T1. Very well made and very loud! Feelnmich safer having it in my hunting bag! T2. Perfect for any animal lover - it says it all. The decal came shipped really nicely and shipped fast Probing

Task: Predict each stylistic feature label from the text: Text Want to get one of these: But don’t have enough money. Alas. $249 for the smallest one seems a little steep.. . . Features (which quintile does the feature fall into?) n_adj=1, n_verb=0, t_word=0, t_syll=0, . . . (42 more features)

n_adj → 1 n_verb → 0 t_word → 0 t_syll → 0

Table 12: Examples of each STEB Tasks.

GEDE added as a clustering task for distinguishing human and the various LLMs. The Generative Essay Detection in Education (GEDE) dataset (Gehring and Paaßen, 2025) is a collection of three academic essay datasets (Argument Annotated Essays (AAE), PERSUADE 2.0, and British Academic Written English (BAWE)) plus machinegenerated versions of those essays at various levels of contribution using GPT-4o-mini and Llama-3.370b-Instruct. There are a total of 916 human essays and 12,703 LLM essays. The dataset is provided under a CC BY-NC-SA 4.0 DEED AttributionNonCommercial-ShareAlike 4.0 International license at https://github.com/lukasgehring/ Assessing-LLM-Text-Detection-in-Educati onal-Contexts.

DetectRL (Wu et al., 2024) added as five standard machine-textdetection datasets (DetectRL_{arxiv, direct_prompt, writing_prompt, xsum, yelp_review} and three adversarial datasets (DetectRL_{paraphrase, perturbation, prompt}_attacks, binary human-vs-LLM). Source: https://github.com/NLP2CT/Detect RL.

M4 (Wang et al., 2024b) added as five clustering datasets over English-only domains: arxiv, peerread, reddit, wikihow, and wikipedia. Each domain contains data from 5 different LLMs (Cohere, davinci, ChatGPT, Dolly-v2, and BloomZ), but we skip BloomZ due to the problematic nature of the data where parts of the prompt are in the “generated" text. Source: https: //github.com/mbzuai-nlp/M4. The data is available for research use and has previously been used in SemEval 2024 (https://github.com/m bzuai-nlp/SemEval2024-task8/). MAGE Li et al. (2024) added as ten clustering datasets over the HuggingFace dataset yaful/MAGE: cmv, eli5, hswag, roct, sci_gen, squad, tldr, wp, xsum, and yelp. Each domain’s records are labelled by generator model only, collapsing across generation methods (continuation, specified, topical), giving 27 unique machine labels plus human per domain. The data is available for public use in HuggingFace (yaful/MAGE). PAN24 Generative Authorship added as clustering and all-to-all pair classification task. We use the test partition of the PAN24 Generative Authorship Detection (Ayele et al., 2024) task. All PAN datasets are available for research use under

the CC BY 4.0 license. PAN25/26 Generative AI Detection (Task 1) added as three clustering datasets (PAN25_26_{essays, fiction, news}). The PAN25 and PAN26 editions of Task 1 (Bevendorff et al., 2025b, 2026) share the same underlying validation data. We split it by the genre field. Each record is labelled by its model field, yielding a multi-class clustering task with human plus all participating LLMs per genre. All PAN datasets are available for research use under the CC BY 4.0 license. PAN25 Human-AI Collaborative Text Classification added as clustering task. The PAN25 Task 2 development set (Bevendorff et al., 2025b) labels each text with one of six human-AI collaboration categories (e.g. fully human-written, human-initiated, then machine-continued, machineinitiated, then human-continued). We perform the clustering across these different labels. All corpora are available at Zenodo, and the download links can be found in the PAN website: https: //pan.webis.de/data.html. All PAN datasets are available for research use under the CC BY 4.0 license. C.9

ange. It is licensed by CC BY-SA 3.0 which allows research use. We collect a small sample of StackExchange data spanning 500 authors who contribute to any community at random. The retrieval setup is similar to Rivera Soto et al. (2021), where multiple documents serve as a query and multiple documents serve as a target.

D

TF-IDFngrams model details

While many approaches compare the frequencies of particular n-grams, simply counting their presence in documents does not reveal broader contextual information, such as how common the n-gram is in general. Therefore, n-grams can be weighted using their Term Frequency-Inverse Document Frequency (TF-IDF), which considers how often an ngram appears in a particular document compared to how rare it is in a corpus overall. This measurement increases the weight of rare words, which might be distinguishing of an author, and reduces the weight of common words. The authorship verification baseline for the PAN competition (Stamatatos et al., 2023), for instance, uses TF-IDF-weighted character 4-grams.

Authorship retrieval datasets

PAN18 Cross-Domain Authorship Attribution (Stamatatos et al., 2018b) added as retrieval task. Each problem provides a set of candidate authors with one or more known texts and a set of unknown query texts to be attributed. We concatenate each candidate’s known texts into a single target document and add each unknown text as a query with the corresponding ground-truth label. The PAN18 dataset also contains data in French, Italian, Polish, and Spanish. All PAN datasets are available for research use under the CC BY 4.0 license. Amazon and Reddit added as retrieval task. We use the same retrieval test splits as Rivera Soto et al. (2021), subsampling each dataset to 1000 query and target pairs to make the evaluation of all models tractable within our compute and timing constraints (the original datasets contain over 100, 000 queries and targets). We re-distribute the subsampled test splits for research purposes only. StackExchange added as retrieval task. StackExchange data is readily available for use from: https://archive.org/download/stackexch

We experimented with fitting the vectorizer to a few diverse datasets—Corpus of Diverse Styles (Krishna et al., 2020), Reddit Million User Dataset (Khan et al., 2021), and FineWeb (Penedo et al., 2024)—but found performance to be similar for all. After experimenting with various n values for each of these models, we found that character 3-5grams, token 1-2-grams, and POS tag 1-2-grams worked best overall. Based on these findings, for the main paper, we chose a TFIDFngrams model fit to a 10 billion token sample of FineWeb (Penedo et al., 2024), a large collection of cleaned English web data from CommonCrawl, with these n-gram settings. Tab. 6 through Tab. 9 show results for a few variations of TF-IDF n-gram models based on the dataset it was trained on (Reddit, FineWeb) and the chosen n (1-2, 1-3) for the token n-grams and POS tag n-grams. The character n-gram values were not decreased because very short character n-grams occur extremely often but do not carry much discriminative information; this makes the feature matrix larger and sparser, increases computational cost, and can introduce noise that hurts generalization.

Object of Study

Ling. Feat. Content Independence

Dataset

Gen. Register Time Demo. Dialect Idiolect

CORE x_genre MASC Corpus of Diverse Styles

✓ ✓ ✓ ✓

STEL_register SynthSTEL_register Graded Formality OneStopEnglish ASSET wikipedia_politeness stackexchange_politeness

Order Alignment†

✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

Bible versions CEECES Shakespeare Blog FCE L1 PASTEL

✓ ✓ ✓ ✓ ✓

✓ ✓ ✓

✓ ✓ ✓ ✓

EnDive Twitter AAE/SAE eWAVE Enron authorship corpus PAN13 AV PAN14 AV PAN15 AV PAN20 AV PAN21 AV PAN18 SC PAN22 SC PAN23 SC PAN24 SC PAN25 SC PAN26 SC PAN18 cross-domain AA amazon fanfiction stackexchange_retrieval STEL_feature SynthSTEL_feature StylePTB probing_amazon probing_blog probing_reddit probing_stackexchange

✓ ✓ ✓

✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

✓ ✓ ✓

Table 13: Datasets used for each cluster of the STEB Score (definitional). This mirrors the three clusters of Tab. 3 and Tab. 9 after Wegmann et al. (2026): Object of Study (Genre, Register, Time, Demographics, Dialect, Idiolect), Linguistic Features (Ling. Feat.), and Content-Independence (Order Alignment task). Idiolect includes authorship verification (PAN style-change and PAN authorship-verification editions, plus Enron) and authorship retrieval (PAN18 cross-domain AA English, Amazon, fanfiction, stackexchange_retrieval). PAN rows are collapsed across tasks and difficulty levels (AV = authorship verification, SC = style change, AA = authorship attribution); Order Alignment† uses the distractor variant of the order-alignment task for the respective datasts and no other task for that dataset. The other columns use all available tasks for the included dataset and the acc variant of order alignment.

E

Embeddings vs. prompting

In practice: Why not just use prompting? To demonstrate this concretely, we compare prompting GPT-5.2 against LUAR-CRUD on authorship retrieval, using a relatively small corpus of 100 query and 100 target collections of 16 sentences

each, created with Amazon Reviews data (§ E). The dataset is a subsample of the STEB Amazon retrieval dataset described in § C.9. As no established prompts exist for retrieval (likely due to scaling issues, cf. §E), we adapt the linguistically informed prompts created for authorship verification by Huang et al. (2024), with minor grammat-

ical corrections. We provide the full prompt in App. Fig. 3. Table 1 reports the same performance measures on retrieval as STEB (Section 3). We additionally report tera floating point operations (TFLOPs) and API costs in US dollars (GPT-5.2 only). LUAR-CRUD, a model from 2021, outperforms GPT-5.2 released in 2025 while being at least 500× more efficent than causal models with 1B+ parameters. We display the prompt used with GPT-5.2 in Figure 3. The lower efficiency of LLM models is unsurprising: authorship retrieval requires comparing a query against every target collection in the pool, which scales poorly with prompting approaches. As the target pool grows, it may even exceed the model’s context window or lead to large costs in US dollars or FLOPs.9 To illustrate, a common Reddit authorship retrieval dataset (Khan et al., 2021) contains 120k authors with 16 texts each, requiring approximately 5.1M tokens per retrieval query, which is considerably larger than GPT-5.2’s context window of 400k.10 Even if the context window was not restricted, retrieval on the 120k dataset would require about 20 billion TFLOPs for a 1 billion parameter causal model and 40k TFLOPS for models like LUAR.

F

Multilingual truncation analysis

The main-text multilingual ranking (Tab. 5) is computed under STEB’s default protocol, which segments long inputs into sentence-boundary-aware chunks and mean-pools the chunk embeddings (§G). To isolate the effect of this evaluation choice, we reproduce the multilingual evaluation setup of Kim et al. (2025b): the same PAN13/14/15 authorship verification datasets (Dutch, Greek, Spanish), the same pre-defined pair classification protocol, the same style embeddings, and the same AUC metric. The only change is that each document is truncated to the model’s maximum length rather than chunked and pooled as in STEB. Tab. 14 compares the two protocols. Under truncation, MSR moves from rank 4 to rank 1, while without truncation it ranks 4th, underperforming the style embeddings trained on English data. These results underscore the necessity of shared evaluation protocols, as without them our conclusions would differ materially. 9

This parallels problems faced in information retrieval and RAG (Gao et al., 2024). 10 See https://developers.openai.com/api/docs/m odels/gpt-5.2

Chunk & Pool Model

AUC

Rank

StyleDistance CISR LUAR-MUD MSR mStyleDistance

72.40 70.08 69.50 69.43 59.94

1 2 3 4 5

Truncate AUC Rank 58.40 60.31 58.52 67.51 56.88

4 2 3 1 5

Table 14: Mean AUC (×100) on the PAN13/14/15 multilingual AV datasets used by Kim et al. (2025b), under STEB’s default chunk-and-pool protocol versus perdocument truncation. Bold = best per protocol.

G

Chunking large documents

We demonstrate STEBs strategy for embedding large inputs into a single representation in Fig. 4. Note that this approach is fair to all encoders, as each one is able to observe the same amount of tokens regardless of what their token limits are.

H

Use of AI Assistants

We used AI-assistants (Claude Code, Cursor) to help us develop the STEB code-base. That said, all the code underwent human-review, and we rigorously tested the code-base for correctness. We also used LLMs to give us ideas for rephrasing and shortening excerpts of the paper.

I

Compute Requirements

We ran each model on a single H100 or A100 40Gb GPU. No single run took longer than 24hrs, and the largest model evaluated has 8B parameters. We upper-bound our GPU expenditure at 960 GPU hours (24 hours times 40 models). Please note that this is a very large over-estimate, most models took less than 8 hours to run.

System Message Respond with a JSON object with two elements: { "analysis": Reasoning behind your answer. "answer": A list of all candidate IDs (integers 0 to N-1) sorted from most to least likely to be written by the same author as the query text. The list must contain every candidate ID exactly once. } User Prompt You are given a set of texts written by one unknown person (the query author) and several sets of candidate texts written by several known authors (the candidate authors). Rank all candidate sets of texts by how likely each was written by the query author. Analyze the writing style only, disregarding the differences in topic and content. Base your reasoning on linguistic features such as phrasal verbs, modal verbs, punctuation, rare words, affixes, quantities, humor, sarcasm, typographical errors, and misspellings. Query texts: [16 QUERY DOCUMENTS UNKNOWN AUTHOR] Candidates (rank all N by likelihood of matching the query author): Candidate 0: [16 TARGET DOCUMENTS AUTHOR 0] Candidate 1: [16 TARGET DOCUMENTS AUTHOR 1] .. . Candidate N −1: [16 TARGET DOCUMENTS AUTHOR N − 1]

Figure 3: Prompt template for AA via LLM-based stylistic analysis.

Segmentation Chunk 1

Encodings

Mean Pooling

Encoder

[0-512)

Long Input Document Chunk 2 (e.g. 2048 tokens)

1 N

Encoder

[512-1024)

Chunk N

P Final Vector

Encoder

(Remainder)

Embeddings

Figure 4: Example of STEBs chunk-and-pool strategy. In the example, it is assumed that the model’s maximum context-length is 512 tokens. The long input document is chunked up into segments of 512 tokens that respect sentence boundaries. Each chunk is then embedded individually be the encoder, and finally we mean-pool across the chunks derive our final embedding.

J

Sample Counts per Dataset

We specify the number of classes, number of samples per class, and total number of texts embedding in Tab. 15. Table 15: Per-dataset sample counts in STEB. “# Classes” refers to classes for clustering and pair-classification tasks, trials for pre-defined pair classification, style groups for order alignment, unique author/item ids for retrieval, and probing features for probing. “Samples/Class” is the per-class evaluation budget. “Total” is the total number of texts the embedder is asked to encode.

Dataset

# Classes

Samples/Class

Total

8 3 2 6 38 11 53 6 2 15 4 2 3 5 5 2 2 2 5 5 5 6 6 6 5 5 28 28 28 28 28 28 28 28 28 28 14 13 12

200 200 200 200 25 200 29 200 200 25 200 200 200 200 200 200 200 200 200 200 200 200 200 200 200 200 60 73 66 76 36 55 57 70 86 42 200 25 25

1,600 600 400 1,200 950 2,200 1,537 1,200 400 375 800 400 600 1,000 1,000 400 400 400 1,000 1,000 1,000 1,200 1,200 1,200 1,000 1,000 1,680 2,044 1,848 2,128 1,008 1,540 1,596 1,960 2,408 1,176 2,800 325 300

Clustering (53 datasets) bible_versions blog_age blog_gender ceeces core corpus-of-diverse-styles eWAVE endive enron_spam fce_l1 gede_essay_detection hate_speech hate_speech_and_offensive_language machine_text_detection_DetectRL_arxiv machine_text_detection_DetectRL_direct_prompt machine_text_detection_DetectRL_paraphrase_attacks machine_text_detection_DetectRL_perturbation_attacks machine_text_detection_DetectRL_prompt_attacks machine_text_detection_DetectRL_writing_prompt machine_text_detection_DetectRL_xsum machine_text_detection_DetectRL_yelp_review machine_text_detection_M4_arxiv machine_text_detection_M4_peerread machine_text_detection_M4_reddit machine_text_detection_M4_wikihow machine_text_detection_M4_wikipedia machine_text_detection_MAGE_cmv machine_text_detection_MAGE_eli5 machine_text_detection_MAGE_hswag machine_text_detection_MAGE_roct machine_text_detection_MAGE_sci_gen machine_text_detection_MAGE_squad machine_text_detection_MAGE_tldr machine_text_detection_MAGE_wp machine_text_detection_MAGE_xsum machine_text_detection_MAGE_yelp machine_text_detection_PAN24_news machine_text_detection_PAN25_26_essays machine_text_detection_PAN25_26_fiction

Continued on next page

Table 15 – continued Dataset machine_text_detection_PAN25_26_news machine_text_detection_PAN25_collaborative masc_text_genre pastel_age pastel_education pastel_ethnic pastel_gender pastel_politics pastel_tod sms_spam stackexchange_politeness telegram-spam-ham wikipedia_politeness x_genre

# Classes

Samples/Class

Total

14 6 19 6 8 9 3 3 5 2 3 2 3 9

59 200 150 25 25 25 55 200 200 200 200 200 200 46

826 1,200 2,850 150 200 225 165 600 1,000 400 600 400 600 414

8 3 2 6 38 11 53 6 2 15 2 3 19 6 8 9 3 3 5 2 3 2 3 9

200 200 200 200 25 200 29 200 200 25 200 200 150 25 25 25 55 200 200 200 200 200 200 46

1,600 600 400 1,200 950 2,200 1,537 1,200 400 375 400 600 2,850 150 200 225 165 600 1,000 400 600 400 600 414

80 30 100 100 200 200

2 2 2 2 2 2

160 60 200 200 400 400

All-to-All Pair Classification (24 datasets) bible_versions blog_age blog_gender ceeces core corpus-of-diverse-styles eWAVE endive enron_spam fce_l1 hate_speech hate_speech_and_offensive_language masc_text_genre pastel_age pastel_education pastel_ethnic pastel_gender pastel_politics pastel_tod sms_spam stackexchange_politeness telegram-spam-ham wikipedia_politeness x_genre Pre-Defined Pair Classification (25 datasets) enron_authorship_corpus pan13_authorship_verification_english_test pan14_authorship_verification_corpus1_english_essays_test pan14_authorship_verification_corpus1_english_novels_test pan14_authorship_verification_corpus2_english_essays_test pan14_authorship_verification_corpus2_english_novels_test

Continued on next page

Table 15 – continued Dataset pan15_authorship_verification_english_test pan18_style_change pan20_authorship_verification_test pan21_authorship_verification_test pan22_style_change_advanced pan22_style_change_basic pan22_style_change_sentence pan23_style_change_easy pan23_style_change_hard pan23_style_change_medium pan24_style_change_easy pan24_style_change_hard pan24_style_change_medium pan25_style_change_easy pan25_style_change_hard pan25_style_change_medium pan26_style_change_easy pan26_style_change_hard pan26_style_change_medium

# Classes

Samples/Class

Total

500 1,492 500 500 9,537 2,141 22,105 2,826 4,112 7,013 2,471 4,131 4,592 10,247 10,648 12,759 71,108 54,141 24,406

2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2

1,000 2,984 1,000 1,000 19,074 4,282 44,210 5,652 8,224 14,026 4,942 8,262 9,184 20,494 21,296 25,518 142,216 108,282 48,812

1 1 3 2 14 32 8 11 2 1 1

400 600 200 392 50 200 200 400 1,000 400 400

400 600 600 784 700 6,400 1,600 4,400 2,000 400 400

500 500 50 500

195 86 5 125

97,787 43,411 259 62,980

46 46 46 46

27,584 39,111 16,447 9,253

27,584 39,111 16,447 9,253

Order Alignment (11 datasets) ASSET OneStopEnglishCorpus STEL_feature STEL_register StylePTB SynthSTEL_feature SynthSTEL_register corpus-of-diverse-styles graded_formality parallel_shakespeare twitter_aave_sae Retrieval (4 datasets) amazon fanfiction pan18_cross_domain_authorship_attribution_english stackexchange_retrieval Probing (4 datasets) probing_amazon probing_blog probing_reddit probing_stackexchange

Record · ID 324902 · SHA-256 3d7524aef7561cf6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.