BERT- AS - A -J UDGE : A R OBUST A LTERNATIVE TO L EXICAL M ETHODS FOR E FFICIENT R EFERENCE -B ASED LLM E VALUATION Hippolyte Gisserot-Boukhlef1,4 Nicolas Boizard2,4 Emmanuel Malherbe1
Céline Hudelot4
1 Artefact Research Center
3 Cohere
2 Diabolocom
Pierre Colombo3
arXiv:2604.09497v1 [cs.CL] 10 Apr 2026
4 MICS, CentraleSupélec, Université Paris-Saclay
Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases. In practice, however, evaluating generative outputs typically relies on rigid lexical methods to extract and assess answers, which can conflate a model’s true problem-solving ability with its compliance with predefined formatting guidelines. While recent LLM-as-a-Judge approaches mitigate this issue by assessing semantic correctness rather than strict structural conformity, they also introduce substantial computational overhead, making evaluation costly. In this work, we first systematically investigate the limitations of lexical evaluation through a large-scale empirical study spanning 36 models and 15 downstream tasks, demonstrating that such methods correlate poorly with human judgments. To address this limitation, we introduce BERT-as-a-Judge, an encoder-driven approach for assessing answer correctness in reference-based generative settings, robust to variations in output phrasing, and requiring only lightweight training on synthetically annotated question-candidate-reference triplets. We show that it consistently outperforms the lexical baseline while matching the performance of much larger LLM judges, providing a compelling tradeoff between the two and enabling reliable, scalable evaluation. Finally, through extensive experimentation, we provide detailed insights into BERT-as-a-Judge’s performance to offer practical guidance for practitioners, and release all project artifacts to foster downstream adoption.
Correspondence: [email protected] Code: https://github.com/artefactory/BERT-as-a-Judge Models & Data: https://hf.co/collections/artefactory/bert-as-a-judge Date: April 2, 2026
1
Introduction
Evaluation lies at the core of the large language model (LLM) ecosystem. In recent years, considerable effort has been devoted to rigorously and fairly assessing model performance across a wide range of tasks, to guide model selection and downstream adoption (Liang et al., 2022; Bommasani et al., 2021). For instruction-tuned models (optimized for human interaction and question answering), evaluation is typically conducted in zero-shot generative settings (Wei et al., 2022; Ouyang et al., 2022), in which models are prompted to directly generate an answer without access to task-specific examples. While conceptually straightforward, this setup poses two challenges for evaluation: reliably extracting the model’s predicted answer for comparison with a reference, and performing the comparison itself. The former arises from answer formatting variations, such as “The answer is X” versus “Answer: X”, the latter occurs when comparing outputs like “2.00” versus “2$”, both of which should be treated as equivalent. A common mitigation strategy is to enforce constrained output formats via prompting, enabling answers to be extracted with regular expressions (regex) (Liang et al., 2023; Gao et al., 2024), and then rely on metrics beyond exact match, such as ROUGE (Lin, 2004), BERTScore (Zhang et al., 2019), or MathVerify (Hugging Face, 2024), thereby avoiding errors caused by formatting inconsistencies or lexical variations. Although more flexible than strict exact match, these metrics can still fail to accurately capture answer correctness, especially since models often do not strictly follow prescribed output formats, making reliable answer parsing difficult. Such deviations may stem from differences in model scale, instruction-tuning data mixtures, or alignment strategies, and can artificially deflate measured downstream performance. While
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
REGEX-BASED EVALUATION
Prompt: What is 2+2? Format your answer as “Answer: <answer>”
BERT-AS-A-JUDGE
Answer Generation
Regex Parsing
Lexical Match
Output 1: Answer: 4
Extraction 1: →4
Score 1: →1
Output 2: The answer is 4
Extraction 2: → N/A
Score 2: →0
Answer Generation
Prompt: What is 2+2? Format your answer as “Answer: <answer>”
Reference: 4
Fine-tuned Encoder
Answer Assessment
Output 1: Answer: 4
Score 1: →1
Output 2: The answer is 4
Score 2: →1
Reference: 4
True Ranking Gemma-3 12B
Qwen-3 14B
Qwen-3 14B
Qwen-3 14B
Ministral-3 14B
Ministral-3 14B
Ministral-3 14B
Phi-4 14B
Phi-4 14B
Phi-4 14B
Gemma-3 12B
Gemma-3 12B
Figure 1: Comparison between regex-based (lexical) evaluation and BERT-as-a-Judge. Top: illustration of both approaches with simple examples. Bottom: model rankings for four similarly sized models from different families, computed via task-wise Borda count. formatting adherence is itself an important capability, particularly for instruction following and structured generation (Ouyang et al., 2022), it should not confound the evaluation of orthogonal competencies such as factual knowledge, mathematical reasoning, or reading comprehension (Hendrycks et al., 2021a; Cobbe et al., 2021). Recently, LLM-as-a-Judge frameworks have emerged as a compelling alternative (Zheng et al., 2023; Wang et al., 2023). By delegating answer comparison to a separate language model, these approaches reduce dependence on rigid formatting constraints and can correctly credit semantically valid but structurally unconventional responses. However, they introduce substantial computational overhead and additional sources of variance, including sensitivity to the choice of judge model and prompt design (Boizard et al., 2025b). Question. How can we measure a model’s core problem-solving ability without relying on output formatting or expensive inference? Contributions.
In this work, we make the following three contributions:
• Through a comprehensive empirical study across a diverse set of models and tasks, we show that lexical evaluation exhibits weak correlation with human judgments (§3). • To address this limitation, we introduce BERT-as-a-Judge, an encoder-driven approach for evaluating generative models in reference-based settings, leveraging the strength of bidirectional attention for text classification (Figure 1). We show that BERT-as-a-Judge consistently outperforms lexical evaluation and even surpasses LLM-as-a-Judge under comparable inference conditions (§4). • We provide detailed insights into BERT-as-a-Judge’s performance through an extensive set of experiments, offering practical guidance for downstream applications (§5). Additionally, we release the packaged code1 along with the full set of generated and annotated data, covering outputs from 36 models across 15 tasks, and open-source all post fine-tuning checkpoints used in our experiments.2 1 https://github.com/artefactory/Bert-as-a-Judge
2 https://hf.co/collections/artefactory/bert-as-a-judge
2
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
2
Experimental Protocol
2.1
Answer Generation
Tasks. The backbone of LLM evaluation consists of tasks whose outputs can be unambiguously judged as correct or incorrect, providing an objective basis for model assessment (Grattafiori et al., 2024; Yang et al., 2025; Olmo et al., 2025; Ramos et al., 2026; Apertus et al., 2025). In this work, we focus on three families of widely used benchmarks: • Multiple-choice, in which models are given a question along with a set of options: MMLU (Hendrycks et al., 2021a), MMLU-Pro (Wang et al., 2024), TruthfulQA (Lin et al., 2021), ARC-Easy/Challenge (Clark et al., 2018), and GPQA (Rein et al., 2024). • Context extraction, where models must provide answers grounded in a given passage by citing relevant evidence: SQuAD-v2 (Rajpurkar et al., 2018), HotpotQA (Yang et al., 2018), DROP (Dua et al., 2019), and CoQA (Reddy et al., 2019). • Open-form mathematics, in which models generate a final closed-form answer in free text: GSM8K (Cobbe et al., 2021), MATH (Hendrycks et al., 2021b), AsDiv (Miao et al., 2020), AIME 24 (Zhang & Math-AI, 2024), and AIME 25 (Zhang & Math-AI, 2025). Models. We perform inference across a broad range of recent open-weight instructiontuned model families, spanning from 135M to 70B parameters. Our study includes 36 models in total: Llama-3 (1B, 3B, 8B, 70B) (Grattafiori et al., 2024), Qwen-3 (600M, 4B, 8B, 14B, 32B) (Yang et al., 2025), Gemma-3 (1B, 4B, 12B, 27B) (Team et al., 2025), Falcon-3 (1B, 3B, 7B) (Team, 2024), Phi-4 (3.8B, 14B) (Abdin et al., 2024; Abouelenin et al., 2025), SmolLM-2 and 3 (135M, 360M, 1.7B, 3B) (Allal et al., 2025; Bakouch et al., 2025), OLMo-3 (7B, 32B) (Olmo et al., 2025), Ministral-3 (3B, 8B, 14B) (Liu et al., 2026), LFM-2 (350M, 700M, 1.2B, 2.6B) (Liu et al., 2026), EuroLLM (1.7B, 9B, 22B) (Martins et al., 2025b;a; Ramos et al., 2026), and Apertus (8B, 70B) (Apertus et al., 2025). Generation parameters. For each task-model pair, responses are produced in a zero-shot setting using greedy decoding, with a maximum generation length of 2048 tokens. For experimental purposes, models are prompted to conclude their outputs in the format “Final answer: [answer]” to facilitate downstream regex parsing and ensure fair comparison between model- and regex-based assessment methods. 2.2
Labeling
Synthetic labeling. For annotation, we employ Nemotron-Super-v1.5 (Bercovich et al., 2025) as an automatic evaluator. The model is provided with the question, the candidate answer, and the reference answer, and is asked to determine whether the candidate response is correct given the available information.3 Inference is conducted using greedy decoding in non-reasoning mode. Human labeling. To validate the reliability of the synthetic labeling approach, we perform human annotation on a subset of the data. Specifically, we randomly sample instances from the generated dataset and have them independently labeled by a pool of 11 human evaluators, totaling 3,212 annotations, and resulting in an overall average agreement of 97.5% with the synthetic labels.4 2.3
Evaluation Methods
BERT-as-a-Judge. We propose to train a BERT-like encoder model on labeled questioncandidate-reference triplets constructed as described in §2.1 and §2.2, leveraging its bidirectional attention mechanism well suited for structured text classification (Zhang et al., 2025). We construct the training mixture from the tasks described in §2.1 that provide an explicit 3 The full evaluation prompt is provided in Appendix A. 4 Further details are provided in Appendix C.
3
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
training split, namely MMLU, ARC-Easy, ARC-Challenge, SQuAD-v2, HotpotQA, GSM8K, and Math. The training dataset is constructed to balance the number of samples across task categories and models, resulting in approximately 1M synthetically labeled samples in total. We initialize the encoder from EuroBERT 210M (Boizard et al., 2025a) and fine-tune it for one epoch using binary cross-entropy. We employ a learning rate of 2 × 10−5 , following the authors’ recommendations for sequence classification, along with a 5% warmup ratio and a linear decay schedule. Training is conducted on 8 MI250x GPUs, yielding an effective batch size of 32, taking approximately 20 GPU hours per run. Baselines.
We compare BERT-as-a-Judge to the following baselines:
• Regex: Extracts answers using a regular expression based on the pattern “Final answer: [answer]” and evaluates multiple-choice tasks with exact match, context extraction with ROUGE-L (Lin, 2004), and open-form math with Math-Verify (Hugging Face, 2024). For answer parsing, we build on the regex rules provided by the lm-evaluation-harness framework (Gao et al., 2024), adapting them to our prompting format and the range of evaluated models. • LLM-as-a-Judge: Uses a generative model to determine whether a candidate response matches the reference answer for a given question, following the procedure described in §2.1. To keep inference costs comparable to the encoder, we use a model of similar scale by default (Qwen-3 0.6B), prompting it to respond directly with “True” or “False”. Larger LLM judges and more flexible prompting strategies are also evaluated in §5. 2.4
Assessment of Evaluation Quality
Metric. We assess each method by its accuracy against synthetic labels from NemotronSuper-v1.5,5 reflecting how well it predicts whether a given answer is correct or incorrect.6 Benchmarks. We assess all evaluation methods on the full set of tasks introduced in §2.1, including both the test splits of tasks used during encoder training (§ 2.3) and the tasks reserved exclusively for out-of-domain evaluation.
3
Limitations of Regex-Based Evaluation
This section analyzes the impact of regex-based evaluation on measured downstream performance, noting discrepancies from both formatting-related parsing failures and postparsing matching errors. Specifically, we quantify parsing failure rates across a range of models in Figure 2 and assess performance deltas relative to ground-truth labels in Table 1. Model scale, family, and task type have a high impact on output formatting. Figure 2 shows that larger models tend to produce fewer formatting errors, as illustrated by the Llama-3 models on context extraction and Qwen-3 on open-form math tasks. Model family also plays a significant role: Qwen-3 and Gemma-3 consistently achieve near-perfect formatting compliance on context extraction, whereas smaller Llama-3 models exhibit substantially higher failure rates. Task type further impacts formatting accuracy. Open-form math proves the most challenging, with Llama-3 70B generating incorrectly formatted outputs over 60% of the time and Qwen-3 32B around 20%, while multiple-choice and context extraction tasks are much easier, with mid- to large-scale models often achieving near-zero failure rates. Regex-based evaluation distorts performance measurements. Table 1 illustrates the risks of relying on regex-based evaluation, showing substantial negative deltas in measured performance across a broad set of models. Notably, even models with high formatting compliance (e.g., Gemma-3 family on context extraction tasks) suffer from substantial 5 For methods producing scores between 0 and 1 (e.g., encoder models with soft probabilities), we use a default threshold of 0.5. 6 To estimate performance with respect to human annotations, we apply a correction based on the observed agreement between human and synthetic labels (§2.2); results are reported in Appendix C.
4
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Context Extraction
Open-Form Math
15
40
5
10
20
0
0
0
Model Size (B)
1 3 8 70 1 4 12 27 0.6 4 8 14 32
20
10
Llama-3 Gemma-3 Qwen-3
60
30
Model Size (B)
1 3 8 70 1 4 12 27 0.6 4 8 14 32
20
1 3 8 70 1 4 12 27 0.6 4 8 14 32
Parsing Failures (%)
Multiple-Choice
Model Size (B)
Figure 2: Quantification of regex parsing failures. Values represent the failure rate, defined as the percentage of instances with unparsable outputs. Results are shown for the Llama-3, Gemma-3, and Qwen-3 model families and aggregated by task category. Multiple-Choice
Context Extraction
Open-Form Math
Family
Size
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
Llama-3
1B 3B 8B 70B
-0.9 (-2.6/+1.7) -13.1 (-8.5/-4.6) -23.3 (-0.9/-22.3) -1.1 (-1.1/-0.1)
↑4.8 ↓0.5 ↓6.4 ↑1.3
-23.9 (-18.1/-5.8) -27.5 (-15.6/-11.9) -21.0 (-4.0/-17.0) -18.4 (-0.1/-18.3)
↓1.2 ↓0.4 ↑3.5 ↑6.5
-11.9 (-11.1/-0.8) -7.0 (-2.6/-4.4) -5.4 (-0.8/-4.6) -30.4 (-29.0/-1.4)
↓0.8 ↑1.7 ↑4.2 ↓13.3
Gemma-3
1B 4B 12B 27B
-0.1 (-0.7/+0.5) +0.3 (-0.2/+0.5) -0.1 (-0.2/+0.1) +0.1 (-0.1/+0.2)
↑5.7 ↑6.2 ↑5.3 ↑4.7
-12.3 (-0.1/-12.2) -30.8 (-0.0/-30.8) -27.6 (-0.1/-27.4) -28.7 (-0.0/-28.7)
↑4.6 ↓1.5 ↓4.2 ↓2.1
-11.3 (-6.8/-4.5) -11.8 (-1.2/-10.6) -10.3 (-1.4/-8.9) -10.8 (-1.1/-9.7)
↓1.2 ↓0.1 ↑1.6 ↑0.4
Qwen-3
0.6B 4B 8B 14B 32B
-20.8 (-0.0/-20.8) -3.2 (-0.5/-2.7) -7.0 (-0.6/-6.3) -19.3 (-0.7/-18.6) -23.8 (-0.5/-23.3)
↓2.8 ↑3.5 ↑0.7 ↓13.2 ↓17.6
-13.8 (-0.0/-13.8) -20.2 (-0.0/-20.2) -29.7 (-0.0/-29.7) -20.9 (-0.0/-20.9) -25.1 (-0.0/-25.1)
↑2.8 ↑7.0 ↓7.2 ↑0.5 ↓1.1
-10.8 (-5.8/-5.0) -12.0 (-5.3/-6.7) -15.2 (-7.9/-7.3) -15.0 (-7.8/-7.2) -10.3 (-2.6/-7.7)
↓0.5 ↑1.0 ↓0.2 ↑0.8 ↑2.8
Table 1: Impact of regex-based evaluation on measured model performance. The “∆Accuracy” columns report the difference between regex-based and ground-truth (synthetic-label) accuracy, with values in parentheses indicating the contributions from parsing failures and post-parsing matching, respectively. The “∆Rank” columns show the corresponding average change in ranking across the 36 evaluated models, where greener values indicate upward rank changes and redder values indicate downward rank changes.
underestimation, often due to overly verbose outputs that technically follow formatting rules but fail lexical matching.7 Overall, regex-based assessment significantly distorts leaderboard rankings; for instance, Qwen-3 32B drops 18 positions while Gemma-3 4B climbs 6 on multiple-choice tasks, with these shifts often representing mere artifacts of formatting quirks and rigid lexical matching rather than true differences in capability.
4
Encoder-Based Evaluation
In this section, motivated by the observation that regex-based evaluation fails to accurately reflect true model performance across a wide range of benchmarks (§3), we train a BERT-asa-Judge encoder model to assess answer correctness following the methodology described in §2.3. We then assess the proposed approach using the setup detailed in §2.4 and report the results in Table 2. To further support our analysis, Table 3 shows performance for models whose outputs were excluded from the training mixture. 7 More details are provided in Appendix E.
5
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
LLM-Judge
MultipleChoice
Context Extraction
Open-Form Math
ID
OOD
ID
OOD
ID
OOD
3B 8B 14B
97.0 97.9 98.2
Ministral-3 96.9 83.7 82.5 97.8 87.9 87.1 98.3 89.1 88.6
81.4 83.5 83.5
86.5 84.8 82.6
0.35B 0.7B 1.2B 2.6B
94.8 96.8 97.5 97.9
94.1 96.7 97.1 97.8
LFM-2 90.5 87.4 91.2 87.1
88.6 84.4 90.9 86.0
97.1 97.4 94.6 93.9
96.9 97.0 94.4 93.7
98.8 93.7
1.7B 9B 22B
93.4 98.6 98.2
93.1 98.5 98.1
EuroLLM 91.5 91.2 90.7 90.2 90.6 90.6
98.5 94.5 91.5
98.4 94.1 91.1
90.0 91.4 95.3
8B 70B
97.6 98.1
97.4 98.0
Apertus 90.3 89.9 89.5 89.5
97.5 97.2
97.4 97.1
Task
Regex
ARC-Challenge ARC-Easy MMLU
Multiple-Choice 89.0 50.2 88.2 54.0 88.1 50.3
BERT-Judge 99.4 99.7 98.5
GPQA MMLU-Pro TruthfulQA
86.5 88.8 92.5
93.5 96.5 98.6
HotpotQA SQuAD-v2 CoQA DROP
Context Extraction 75.6 70.0 72.3 62.5 67.0 75.2 77.0 69.3
GSM8K Math
Open-Form Math 94.4 71.3 73.4 58.9
AIME24 AIME25 ASDiv
87.8 91.8 89.2
66.2 57.1 54.5
77.9 83.0 75.5
Size
90.9 89.3 88.1 88.6
Table 2: Accuracies of evaluation methods against ground-truth labels across tasks, averaged over models. Dashed lines separate test-only tasks from those with a training split. Bold indicates the highest accuracy per task.
Table 3: Assessment accuracy on out-ofdomain models. “ID” denotes training on all model outputs, while “OOD” excludes specific models from the training mixture. Results are aggregated by task category.
BERT-as-a-Judge shows the strongest alignment with human judgments. As shown in Table 2, our trained encoder achieves the highest accuracy against ground-truth labels across all benchmarks. This advantage holds across task types, reaching near-perfect alignment on multiple-choice datasets (e.g., 99.7% on ARC-Easy and 98.5% on MMLU) and remaining high on complex-output tasks (98.8% on GSM8K and 93.9% on MATH). It also consistently outperforms the regex-based method by substantial margins (e.g., +21.1% on CoQA, +20.3% on MATH, +10.4% on ARC-Challenge), demonstrating that a dedicated encoder more reliably captures answer correctness than rigid lexical heuristics, while remaining computationally efficient (≈200 ms per sample on an Apple M1 CPU). BERT-as-a-Judge is robust to out-of-domain tasks. Beyond its strong overall accuracy, the encoder-based method maintains high accuracy even on tasks excluded from the training mixture (e.g., 98.6% on TruthfulQA, 88.1% on CoQA, and 95.3% on ASDiv), highlighting its strong generalization ability across the three task categories considered (Table 2). BERT-as-a-Judge generalizes to unseen models. Table 3 shows that removing generations from specific models in the training mixture has minimal impact on downstream assessment quality for those excluded instances. This demonstrates the strong generalization ability of our approach, indicating it can be safely extended to additional model families outside the training mixture and supporting broader adoption. Typical encoder scales appear insufficient for LLM-as-a-Judge. Even at three times the size of our encoder, the baseline LLM judge (Qwen-3 0.6B) consistently produces the weakest results, substantially underperforming the regex baseline across all task categories (Table 2). For instance, on ARC-Challenge, it achieves only 50.2% accuracy versus 89.0% for regex, with similar gaps on context extraction tasks (62.5% vs. 72.3% on SQuAD-v2). These results suggest that the evaluation capabilities of LLM-as-a-Judge do not hold for generative models under 1B parameters, highlighting the critical role of scale for such methods. We examine this limitation further in §5.
5
Experimental Analysis
In this section, we further investigate the properties of our encoder-based evaluation method through a series of complementary analyses. 6
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Assessment Accuracy
Multiple-Choice
Context Extraction
Open-Form Math
100 80 BERT-J. Qw.-3 (S) Qw.-3 (L) Gem.-3 (S) Gem.-3 (L)
60 40 10
9
10
10
10
11
Inference FLOPs
10
12
10
9
10
10
10
11
10
12
10
Inference FLOPs
9
10
10
10
11
10
12
Inference FLOPs
10
13
Figure 3: Comparison between encoder-based evaluation and LLM judges from the Qwen-3 and Gemma-3 families across different model sizes and inference budgets. “S” (“short”) denotes the default generation setup, in which the model answers directly with “True” or “False”, while “L” (“long”) allows the generation of intermediate chain-of-thought tokens before the final judgment.
BERT-as-a-Judge consistently outperforms LLM-as-a-Judge across a wide range of inference budgets. As shown in Table 2, generative evaluation performs poorly at small scales (0.6B parameters). To complement these findings, we conduct a more extensive comparison by varying the judge family (Qwen-3, Gemma-3), model size (0.6B to 32B for Qwen-3 and 1B to 27B for Gemma-3), and inference budget (allowing or not intermediate chain-of-thought tokens before producing the final assessment). Figure 3 shows that our encoder matches the performance of the top-performing LLM judges in the defined setup, while remaining drastically less computationally expensive in terms of inference FLOPs.8
Assessment Accuracy
BERT-as-a-Judge is training-efficient. By default, we 98 train encoder models on 1M question-candidate-reference triplets. In this experiment, we evaluate lighter configura96 tions: 500K, 200K, and 100K samples (Figure 4). Remark94 ably, 100K training samples are sufficient to accurately Multiple-Choice 92 Context Extraction evaluate multiple-choice and open-form math tasks, with Open-Form Math 90 no significant improvement observed beyond this point. Gains are more noticeable for context extraction, as ex88 pected, since this task category requires more than simple 100K 200K 500K 1M candidate-reference matching and often demands underTraining Samples standing of the context provided by the question. Overall, with just 2 GPU hours of training (corresponding to the Figure 4: BERT-as-a Judge eval100K-sample configuration), our encoder achieves a high uation quality across different assessment accuracy, making it well-suited for settings training budgets. with limited data and computational resources. Task Category
Regex
BERT-J.
Regex+ BERT-J.
Multiple-Choice Context Extraction Open-Form Math
88.8 73.0 87.3
97.7 89.2 93.9
90.5 75.2 89.9
Task Category Multiple-Choice Context Extraction Open-Form Math
Table 4: Comparison of hybrid answer evaluation (Regex+BERT-J.) with standalone regex and BERT-as-a-Judge. Bold values indicate the highest accuracy in each row.
Regex 88.8 73.0 87.3
BERT-Judge w/ Q.
w/o Q.
97.7 89.2 93.9
97.3 84.2 93.9
Table 5: Impact of including (w/ Q.) or excluding (w/o Q.) the question in the encoder training prompt. Bold values denote the highest accuracy for each task category.
8 Inference FLOPs are estimated using the formula from Kaplan et al. (2020): FLOPs = 2 × model
size (in parameters) × number of generated tokens.
7
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Combining BERT-as-a-Judge with regex offers an efficient compromise. In this experiment, we use BERT-as-a-Judge as a fallback when regex parsing fails. While it does not reach the performance of the standalone encoder, this hybrid approach substantially improves over regex alone. The results demonstrate that selectively applying the encoder can recover a significant portion of assessment accuracy while keeping computational overhead low (for example, reducing total compute by a factor of five for a model with 20% regex failures). Removing the question from the prompt yields a controlled performance decrease. As shown in Table 5, omitting the question during encoder training (leaving only the candidate and reference) reduces overall assessment accuracy by removing some contextual information. However, this decrease is well-controlled across tasks. The question-free encoder still outperforms regex and remains close to the full-prompt setup, particularly on multiplechoice and open-form math tasks, while also reducing runtime due to shorter prompts and enabling application to any task with fully textual outputs, including multimodal tasks. The performance gap is slightly larger for context extraction, where the question provides critical information, underscoring the encoder’s context-aware capabilities. Test Set (Form.)
Task Multiple-Choice Context Extraction Open-Form Math
Test Set (Free)
Regex
BERT-J. (Form.)
BERT-J. (Free)
Regex
BERT-J. (Form.)
BERT-J. (Free)
88.8 73.0 87.3
97.7 89.2 93.9
97.4 85.8 93.7
– – –
94.0 84.3 93.1
97.6 91.6 93.5
Table 6: Encoder’s robustness to answer formatting. “Test Set (Form.)” denotes test sets with formatted answers, while “Test Set (Free)” contains unformatted answers. “Regex” columns show regex-based results, and “BERT-J. (Form.)” and “BERT-J. (Free)” report accuracies for BERT-as-a-Judge encoders trained on formatted and unformatted answers, respectively. Regex results are omitted for the free-format test set, as answers cannot be reliably parsed.
Assessment Accuracy
BERT-as-a-Judge is robust to variations in answer formatting guidelines. In our core experiments, and to ensure a fair comparison with the regex baseline, we train and evaluate the encoder on answers formatted to facilitate lexical parsing (§2.1). In practice, however, users may follow custom guidelines or allow free-form responses.9 Table 6 shows that under cross-formatting evaluation, meaning free-to-formatted (training on free-form, evaluating on formatted answers) and formatted-to-free (the reverse), we observe a slight performance drop compared to aligned settings. Nevertheless, the encoder still substantially outperforms regex, demonstrating strong robustness to variations in answer formatting. As expected, the free-to-formatted encoder consistently outperforms the formatted-to-free variant, benefiting from exposure to a wider range of formats during training and making it the preferred choice for downstream applications.
Multiple-Choice
100
Context Extraction
Open-Form Math
80 60
ARC-Challenge ARC-Easy GPQA MMLU MMLU-Pro TruthfulQA
40 20 0.00
0.25
0.50
Threshold
0.75
AIME24 AIME25 ASDiv GSM8K Math
CoQA DROP HotpotQA SQuAD-v2
1.00 0.00
0.25
0.50
Threshold
0.75
1.00 0.00
0.25
0.50
Threshold
0.75
1.00
Figure 5: Effect of score thresholding on BERT-as-a-Judge downstream assessment accuracy across the three task categories, averaged over all models. 9 Further details are provided in Appendix A.
8
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
BERT-as-a-Judge is robust to decision threshold variations. Encoder classifiers output continuous sigmoid probabilities, requiring a decision threshold for discrete evaluation. While our main experiments (§4) use a standard 0.5 threshold, Figure 5 demonstrates that accuracies remain remarkably stable across a broad spectrum of threshold values for all task categories. This invariance indicates strong separation between classes, enabling reliable off-the-shelf deployment without the need for task-specific threshold tuning.
6
Related Work
Traditional LLM evaluation. Pretrained language models have traditionally been evaluated using log-likelihood (Radford et al., 2019) or few-shot generation (Brown et al., 2020; Rae et al., 2021; Chowdhery et al., 2023; Touvron et al., 2023; Bai et al., 2023). With the rise of instruction-tuned models, zero-shot generative evaluation has become standard (Wei et al., 2022; Chung et al., 2024; Yang et al., 2025; Ramos et al., 2026). This paradigm typically enforces structured outputs via prompting (Liang et al., 2023; Gao et al., 2024), followed by rule-based comparison to references using deterministic metrics such as exact match, ROUGE (Lin, 2004), Math-Verify (Hugging Face, 2024), or Code-Eval (Chen et al., 2021), making evaluation highly sensitive to surface-level formatting. Model-based evaluation. Both lexical parsing and matching introduce limitations in capturing semantic correctness and robustness. Lexical overlap does not guarantee semantic equivalence, motivating neural metrics such as BERTScore (Zhang et al., 2019) and InfoLM (Colombo et al., 2022) for general text generation, as well as task-specific evaluators like COMET (Rei et al., 2022a;b; Guerreiro et al., 2024), MetricX (Juraska et al., 2023; 2025), and BLEURT (Sellam et al., 2020). Additionally, reliably extracting model outputs is challenging when formatting is inconsistent (Zhou et al., 2023; Pyatkin et al., 2025). To mitigate these issues, LLM-as-a-Judge approaches (Zheng et al., 2023; Wang et al., 2023; Bavaresco et al., 2025; Kim et al., 2023; 2024) directly assess candidate-reference equivalence across tasks, offering greater robustness to formatting artifacts, albeit at a substantial computational cost.
7
Conclusion
In this work, we show that standard evaluation protocols often conflate a model’s underlying problem-solving ability with its compliance to formatting constraints. Across extensive experiments spanning diverse models and tasks, we demonstrate that regex-based evaluation can substantially underestimate true performance. To address this, we propose BERT-as-a-Judge, a lightweight encoder-based framework that better captures semantic correctness, aligns more closely with human judgment, and avoids the high computational cost of LLM-as-a-Judge methods, enabling efficient and more reliable evaluation.
8
Limitations and Future Work
While BERT-as-a-Judge shows strong alignment with human judgments and effectively mitigates the limitations of lexical assessment, our study focuses on a specific subset of evaluation settings, namely, English benchmarks with objectively verifiable answers, where correctness can be clearly defined. Building on these results and the existing literature, a natural next step is to broaden the scope of encoder-based evaluation toward more general-purpose settings. This includes expanding beyond fact-based and structured tasks to open-ended generation scenarios such as summarization, machine translation, code generation, and instruction following. Additionally, adapting the framework to multilingual contexts would further improve its applicability across diverse use cases. As foundation models continue to evolve toward multimodal capabilities, extending this approach to handle vision and speech inputs also presents a promising avenue. Exploring such cross-modal evaluation settings (e.g., visual question answering, image captioning, or speech-based reasoning) could help move toward a unified and efficient evaluation paradigm applicable across tasks and modalities. 9
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Ethics Statement In conducting this research, we recognize the critical importance of fair and reliable evaluation in the LLM ecosystem. Evaluation metrics that are closely aligned with human judgments are essential to ensure that model comparisons accurately reflect real-world capabilities across the widest possible range of tasks. At the same time, the increasing scale of model evaluation, driven by more models, longer outputs, and a growing number of benchmark tasks, can impose substantial computational costs, raising concerns about accessibility and environmental impact. Our work emphasizes the development of lightweight, encoder-based evaluation methods that maintain high correlation with human judgments while minimizing compute requirements. By prioritizing both fairness and efficiency, we aim to support responsible, reproducible, and scalable evaluation practices in the broader LLM research community.
Acknowledgments We gratefully acknowledge the ADASTRA supercomputer at CINES for its technical support and access to HPC resources (grants C1615122 and GDA2401). This work was also supported by the French government under the France 2030 program (ArGiMi project). We further thank the 11 human annotators from the Artefact Research Center, whose careful efforts were instrumental in validating the correlation between our synthetic labeling strategy and human judgments, ensuring the reliability and rigor of our evaluation results.
10
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
References Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. Phi-4 technical report. arXiv preprint arXiv:2412.08905, 2024. Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, et al. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras. arXiv preprint arXiv:2503.01743, 2025. Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martı́n Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlı́ček, Agustı́n Piqueres Lajarı́n, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, and Thomas Wolf. Smollm2: When smol goes big – datacentric training of a small language model, 2025. Project Apertus, Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pasztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, et al. Apertus: Democratizing open and compliant llms for global language environments. arXiv preprint arXiv:2509.14233, 2025. Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. Elie Bakouch, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Lewis Tunstall, Carlos Miguel Patiño, Edward Beeching, Aymeric Roucher, Aksel Joonas Reedi, Quentin Gallouédec, Kashif Rasul, Nathan Habib, Clémentine Fourrier, Hynek Kydlicek, Guilherme Penedo, Hugo Larcher, Mathieu Morlon, Vaibhav Srivastav, Joshua Lochner, XuanSon Nguyen, Colin Raffel, Leandro von Werra, and Thomas Wolf. SmolLM3: smol, multilingual, long-context reasoner, 2025. Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 238–255, 2025. Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah, Ran El-Yaniv, Omri Puny, Ido Galil, Zach Moshe, Tomer Ronen, Najeeb Nabwani, Ido Shahaf, Oren Tropp, Ehud Karpas, Ran Zilberstein, Jiaqi Zeng, Soumye Singhal, Alexander Bukharin, Yian Zhang, Tugrul Konuk, Gerald Shen, Ameya Sunil Mahabaleshwarkar, Bilal Kartal, Yoshi Suhara, Olivier Delalleau, Zijia Chen, Zhilin Wang, David Mosallanezhad, Adi Renduchintala, Haifeng Qian, Dima Rekesh, Fei Jia, Somshubra Majumdar, Vahid Noroozi, Wasi Uddin Ahmad, Sean Narenthiran, Aleksander Ficek, Mehrzad Samadi, Jocelyn Huang, Siddhartha Jain, Igor Gitman, Ivan Moshkov, Wei Du, Shubham Toshniwal, George Armstrong, Branislav Kisacanin, Matvei Novikov, Daria Gitman, Evelina Bakhturina, Jane Polak Scowcroft, John Kamalu, Dan Su, Kezhi Kong, Markus Kliegl, Rabeeh Karimi, Ying Lin, Sanjeev Satheesh, Jupinder Parmar, Pritam Gundecha, Brandon Norick, Joseph Jennings, Shrimai Prabhumoye, Syeda Nahida Akter, Mostofa Patwary, Abhinav Khattar, Deepak Narayanan, Roger Waleffe, Jimmy Zhang, Bor-Yiing Su, Guyue Huang, Terry Kong, Parth Chadha, Sahil Jain, Christine Harvey, Elad Segal, Jining Huang, Sergey Kashirsky, Robert McQueen, Izzy Putterman, George Lam, Arun Venkatesan, Sherry Wu, Vinh Nguyen, Manoj Kilaru, Andrew Wang, Anna Warno, Abhilash Somasamudramath, Sandip Bhaskar, Maka Dong, Nave Assaf, Shahar Mor, Omer Ullman Argov, Scot Junkin, Oleksandr Romanenko, Pedro Larroy, Monika Katariya, Marco Rovinelli, Viji Balas, Nicholas Edelman, Anahita Bhiwandiwalla, Muthu Subramaniam, Smita Ithape, Karthik Ramamoorthy, Yuting Wu, Suguna Varshini Velury, Omri Almog, Joyjit Daw, Denys Fridman, Erick Galinkin, 11
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Michael Evans, Katherine Luna, Leon Derczynski, Nikki Pope, Eileen Long, Seth Schneider, Guillermo Siman, Tomasz Grzegorzek, Pablo Ribalta, Monika Katariya, Joey Conway, Trisha Saar, Ann Guan, Krzysztof Pawelec, Shyamala Prayaga, Oleksii Kuchaiev, Boris Ginsburg, Oluwatobi Olabiyi, Kari Briski, Jonathan Cohen, Bryan Catanzaro, Jonah Alben, Yonatan Geifman, Eric Chung, and Chris Alexiuk. Llama-nemotron: Efficient reasoning models, 2025. Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, and Pierre Colombo. Eurobert: Scaling multilingual encoders for european languages, 2025a. Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El-Haddad, Céline Hudelot, and Pierre Colombo. When does reasoning matter? a controlled study of reasoning’s contribution to model performance. arXiv preprint arXiv:2509.22193, 2025b. Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of machine learning research, 24(240):1–113, 2023. Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instructionfinetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. Pierre Jean A Colombo, Chloé Clavel, and Pablo Piantanida. Infolm: A new metric to evaluate summarization & data2text generation. In Proceedings of the AAAI conference on artificial intelligence, volume 36, pp. 10554–10562, 2022. Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proc. of NAACL, 2019. Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. The language model evaluation harness, 07 2024. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. 12
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Nuno M Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André FT Martins. xcomet: Transparent machine translation evaluation through finegrained error detection. Transactions of the Association for Computational Linguistics, 12: 979–995, 2024. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021a. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021b. Hugging Face. Introducing the math verify leaderboard, 2024. Accessed: 2026-03-05. Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. Metricx-23: The google submission to the wmt 2023 metrics shared task. In Proceedings of the Eighth Conference on Machine Translation, pp. 756–767, 2023. Juraj Juraska, Tobias Domhan, Mara Finkelstein, Tetsuji Nakagawa, Geza Kovacs, Daniel Deutsch, Pidong Wang, and Markus Freitag. Metricx-25 and gemspaneval: Google translate submissions to the wmt25 evaluation shared task. In Proceedings of the Tenth Conference on Machine Translation, pp. 957–968, 2025. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. Prometheus: Inducing finegrained evaluation capability in language models. In The Twelfth International Conference on Learning Representations, 2023. Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. Prometheus 2: An open source language model specialized in evaluating other language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 4334–4353, 2024. Percy Liang, Rishi Bommasani, Tony Lee, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022. Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. Holistic evaluation of language models. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. Featured Certification, Expert Certification. Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004. Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods, 2021. Alexander H Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, et al. Ministral 3. arXiv preprint arXiv:2601.08584, 2026. 13
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Pedro Henrique Martins, João Alves, Patrick Fernandes, Nuno M Guerreiro, Ricardo Rei, Amin Farajian, Mateusz Klimaszewski, Duarte M Alves, José Pombal, Nicolas Boizard, et al. Eurollm-9b: Technical report. arXiv preprint arXiv:2506.04079, 2025a. Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M Guerreiro, Ricardo Rei, Duarte M Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, et al. Eurollm: Multilingual language models for europe. Procedia Computer Science, 255: 53–62, 2025b. Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th annual meeting of the Association for Computational Linguistics, pp. 975–984, 2020. Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, et al. Olmo 3. arXiv preprint arXiv:2512.13961, 2025. Long Ouyang, Jeffrey Wu, Xu Jiang, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 2022. Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert, and Hannaneh Hajishirzi. Generalizing verifiable instruction following, 2025. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021. Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don’t know: Unanswerable questions for SQuAD. In Iryna Gurevych and Yusuke Miyao (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 784–789, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-2124. Miguel Moura Ramos, Duarte M Alves, Hippolyte Gisserot-Boukhlef, João Alves, Pedro Henrique Martins, Patrick Fernandes, José Pombal, Nuno M Guerreiro, Ricardo Rei, Nicolas Boizard, et al. Eurollm-22b: Technical report. arXiv preprint arXiv:2602.05879, 2026. Siva Reddy, Danqi Chen, and Christopher D. Manning. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266, 2019. doi: 10.1162/tacl a 00266. Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. COMET-22: Unbabel-IST 2022 submission for the metrics shared task. In Philipp Koehn, Loı̈c Barrault, Ondřej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Tom Kocmi, André Martins, Makoto Morishita, Christof Monz, Masaaki Nagata, Toshiaki Nakazawa, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Marco Turchi, and Marcos Zampieri (eds.), Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 578–585, Abu Dhabi, United Arab Emirates (Hybrid), December 2022a. Association for Computational Linguistics. doi: 10.18653/v1/2022.wmt-1.52. Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins. CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task. In Philipp Koehn, Loı̈c Barrault, Ondřej Bojar, Fethi Bougares, Rajen 14
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Chatterjee, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Tom Kocmi, André Martins, Makoto Morishita, Christof Monz, Masaaki Nagata, Toshiaki Nakazawa, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Marco Turchi, and Marcos Zampieri (eds.), Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 634–645, Abu Dhabi, United Arab Emirates (Hybrid), December 2022b. Association for Computational Linguistics. doi: 10.18653/v1/2022.wmt-1.60. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level googleproof q&a benchmark. In First conference on language modeling, 2024. Thibault Sellam, Dipanjan Das, and Ankur Parikh. BLEURT: Learning robust metrics for text generation. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7881–7892, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/ v1/2020.acl-main.704. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Beyer, Xiaohai Zhai, Anton Tsitsulin, Robert Busa-Fekete, Alex Feng, Noveen Sachdeva, Benjamin Coleman, Yi Gao, Basil Mustafa, Iain Barr, Emilio Parisotto, David Tian, Matan Eyal, Colin Cherry, Jan-Thorsten Peter, Danila Sinopalnikov, Surya Bhupatiraju, Rishabh Agarwal, Mehran Kazemi, Dan Malkin, Ravin Kumar, David Vilar, Idan Brusilovsky, Jiaming Luo, Andreas Steiner, Abe Friesen, Abhanshu Sharma, Abheesht Sharma, Adi Mayrav Gilady, Adrian Goedeckemeyer, Alaa Saade, Alex Feng, Alexander Kolesnikov, Alexei Bendebury, Alvin Abdagic, Amit Vadi, András György, André Susano Pinto, Anil Das, Ankur Bapna, Antoine Miech, Antoine Yang, Antonia Paterson, Ashish Shenoy, Ayan Chakrabarti, Bilal Piot, Bo Wu, Bobak Shahriari, Bryce Petrini, Charlie Chen, Charline Le Lan, Christopher A. Choquette-Choo, CJ Carey, Cormac Brick, Daniel Deutsch, Danielle Eisenbud, Dee Cattle, Derek Cheng, Dimitris Paparas, Divyashree Shivakumar Sreepathihalli, Doug Reid, Dustin Tran, Dustin Zelle, Eric Noland, Erwin Huizenga, Eugene Kharitonov, Frederick Liu, Gagik Amirkhanyan, Glenn Cameron, Hadi Hashemi, Hanna Klimczak-Plucińska, Harman Singh, Harsh Mehta, Harshal Tushar Lehri, Hussein Hazimeh, Ian Ballantyne, Idan Szpektor, Ivan Nardini, Jean Pouget-Abadie, Jetha Chan, Joe Stanton, John Wieting, Jonathan Lai, Jordi Orbay, Joseph Fernandez, Josh Newlan, Ju yeong Ji, Jyotinder Singh, Kat Black, Kathy Yu, Kevin Hui, Kiran Vodrahalli, Klaus Greff, Linhai Qiu, Marcella Valentine, Marina Coelho, Marvin Ritter, Matt Hoffman, Matthew Watson, Mayank Chaturvedi, Michael Moynihan, Min Ma, Nabila Babar, Natasha Noy, Nathan Byrd, Nick Roy, Nikola Momchev, Nilay Chauhan, Noveen Sachdeva, Oskar Bunyan, Pankil Botarda, Paul Caron, Paul Kishan Rubenstein, Phil Culliton, Philipp Schmid, Pier Giuseppe Sessa, Pingmei Xu, Piotr Stanczyk, Pouya Tafti, Rakesh Shivanna, Renjie Wu, Renke Pan, Reza Rokni, Rob Willoughby, Rohith Vallu, Ryan Mullins, Sammy Jerome, Sara Smoot, Sertan Girgin, Shariq Iqbal, Shashir Reddy, Shruti Sheth, Siim Põder, Sijal Bhatnagar, Sindhu Raghuram Panyam, Sivan Eiger, Susan Zhang, Tianqi Liu, Trevor Yacovone, Tyler Liechty, Uday Kalra, Utku Evci, Vedant Misra, Vincent Roseberry, Vlad Feinberg, Vlad Kolesnikov, Woohyun Han, Woosuk Kwon, Xi Chen, Yinlam Chow, Yuvein Zhu, Zichuan Wei, Zoltan Egyed, Victor Cotruta, Minh Giang, Phoebe Kirk, Anand Rao, Kat Black, Nabila Babar, Jessica Lo, Erica Moreira, Luiz Gustavo Martins, Omar Sanseviero, Lucas Gonzalez, Zach Gleicher, Tris Warkentin, Vahab Mirrokni, Evan Senter, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, Yossi Matias, D. Sculley, Slav Petrov, Noah Fiedel, Noam Shazeer, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Jean-Baptiste Alayrac, Rohan Anil, Dmitry, Lepikhin, Sebastian Borgeaud, Olivier Bachem, Armand Joulin, Alek Andreev, Cassidy Hardin, Robert Dadashi, and Léonard Hussenot. Gemma 3 technical report, 2025. TII Team. The falcon 3 family of open models, December 2024. 15
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. Xuechen Wang et al. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631, 2023. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024. Jason Wei, Maarten Bosma, Vincent Zhao, et al. Finetuned language models are zero-shot learners. International Conference on Learning Representations, 2022. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2018. Junyan Zhang, Yiming Huang, Shuliang Liu, Yubo Gao, and Xuming Hu. Do bert-like bidirectional models still perform better on text classification in the era of llms. arXiv preprint arXiv:2505.18215, 1, 2025. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019. Yifan Zhang and Team Math-AI. American invitational mathematics examination (aime) 2024, 2024. Yifan Zhang and Team Math-AI. American invitational mathematics examination (aime) 2025, 2025. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023. Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models, 2023.
16
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
A
Prompting Details
In this section, we detail how model outputs are generated for both answer generation (Table 7) and answer assessment (Table 9 , Table 10). For answer generation, we also describe the suffixes used to impose different formatting instructions (Table 8). Throughout the main text, we use the soft configuration by default, as it allows both regex-based answer parsing and the inclusion of intermediate chain-of-thought tokens, which improve answer quality (Appendix B). For regex parsing under both the soft and strict formatting constraints, we use the pattern “Final answer:\\s*(.+)”, which provides a general and flexible mechanism for extracting the predicted answer. Task Category
Generation Prompt Answer the following multiple-choice question. Question: {question}
Multiple-Choice
Choices: A) {choice_1} B) {choice_2} C) {choice_3} D) {choice_4} [...] Answer the question based on the provided context.
Context Extraction
Context: {context}
Open-Form Math
{question}
Question: {question}
Table 7: Base generation prompts for each task category. Task Category
Multiple-Choice
Formatting Instruction
Generation Suffix
Free
None
Soft Strict Free
Context Extraction
Open-Form Math
Conclude your response with "Final answer: X", where X is the letter of the correct choice. Respond only with the exact format "Final answer: X", where X is the letter of the correct choice. None
Soft
Conclude your response with "Final answer: X", where X is the exact span from the context that answers the question.
Strict
Respond only with the exact format "Final answer: X", where X is the exact span from the context that answers the question.
Free
None
Soft
Conclude your response with "Final answer: X", where X is the computed solution.
Strict
Respond only with the exact format "Final answer: X", where X is the computed solution.
Table 8: Generation suffixes used across task categories for all formatting strategies.
17
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Assessment Prompt You are an expert evaluator. Your task is to determine whether the CANDIDATE response correctly answers the QUESTION. Judge the CANDIDATE as correct only if its final answer, disregarding any intermediate reasoning or explanation, is semantically equivalent to the REFERENCE with respect to the QUESTION. Base your judgment solely on the information given. Do not rely on external knowledge. [QUESTION starts here] {question} [QUESTION ends here] [REFERENCE starts here] {reference} [REFERENCE ends here] [CANDIDATE starts here] {candidate} [CANDIDATE ends here] Conclude your response with exactly one of the following: - "Final answer: True" if the CANDIDATE is correct - "Final answer: False" if the CANDIDATE is incorrect
Table 9: Answer assessment prompt for LLM judges, allowing intermediate token generation before the final judgment. This involves Nemotron-Super-v1.5 for label generation (§2.2) and generative judges evaluated under inference budget L (Figure 3).
Assessment Prompt You are an expert evaluator. Your task is to determine whether the CANDIDATE response correctly answers the QUESTION. Judge the CANDIDATE as correct only if its final answer, disregarding any intermediate reasoning or explanation, is semantically equivalent to the REFERENCE with respect to the QUESTION. Base your judgment solely on the information given. Do not rely on external knowledge. [QUESTION starts here] {question} [QUESTION ends here] [REFERENCE starts here] {reference} [REFERENCE ends here] [CANDIDATE starts here] {candidate} [CANDIDATE ends here] Respond with exactly one of the following strings (add no additional text): - "Final answer: True" if the CANDIDATE is correct - "Final answer: False" if the CANDIDATE is incorrect
Table 10: Prompt used for direct assessment by LLM judges, applied to generative judges evaluated under inference budget S (Figure 3).
18
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
B
Effect of Generation Mode on Downstream Performance
In this section, we compare different answer production modes to assess their impact on model performance. Specifically, answers are generated under three formatting regimes:10 • Log-likelihood: Candidate answers are iteratively appended to the prompt, and the model’s prediction is derived from the sequence with the highest log-likelihood.11 • Strict: The model is prompted to respond exactly with “Final answer: [answer]”. • Soft: The model is prompted to conclude its response with “Final answer: [answer]” but may reason before answering. • Free: The model may answer in any format. The results are reported in Table 11.
Task
Log-lik.
Generative Strict
Soft
Free
Multiple-Choice ARC-Challenge 46.3 74.4 ARC-Easy 65.0 84.6 GPQA 25.9 32.3 MMLU 39.9 59.4 MMLU-Pro 21.1 38.0 TruthfulQA 33.2 51.3
76.2 84.9 31.8 62.0 44.3 51.6
75.0 84.4 28.8 60.5 42.6 51.3
CoQA DROP HotpotQA SQuAD-v2
Context Extraction – 71.3 – 48.4 – 66.0 – 44.7
76.5 60.2 71.0 53.5
86.7 64.5 82.4 62.0
AIME24 AIME25 ASDiv GSM8K Math
Open-Form Math – 18.1 – 13.2 – 68.2 – 43.1 – 49.3
19.8 14.4 81.5 73.6 60.2
18.1 15.7 82.1 73.6 57.3
Models demonstrate greater capacity in gen- Table 11: Comparison of evaluation erative mode. We first examine multiple- modes across all benchmarks. Results choice tasks by comparing generative evalua- are averaged over all models, and bold tion against the log-likelihood approach. Our values indicate the best performance for results indicate that the likelihood-based setup each task. consistently and severely impairs performance across all evaluated multiple-choice benchmarks (e.g., -22.1% on MMLU and -29.9% on ARC-Challenge). This suggests that while likelihood evaluation offers a convenient, regex-free parsing mechanism, it significantly bottlenecks the model’s inherent problem-solving capabilities compared to generative inference. Strict formatting constraints impair performance. Setting aside likelihood-based evaluation, we compare strict and soft generative prompting strategies. We observe that strict prompting yields the lowest overall performance. While it performs comparably to the soft method on multiple-choice tasks, it significantly degrades performance on tasks requiring more complex outputs (e.g., -11.8% on DROP and -30.5% on GSM8K). This pronounced drop highlights the importance of allowing intermediate chain-of-thought generation to fully leverage the model’s problem-solving capacity.
10 Full prompting details are provided in Appendix A. 11Applies only to multiple-choice tasks.
19
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
C
Human-Synthetic Label Agreement
This section complements § 2.2 in the main text by presenting detailed results of the human annotations. We report human-synthetic average agreement per task category (Table 12), showing consistently high agreement across categories, and analyze how this agreement impacts downstream performance measurement.
Task Category
Accuracy (%)
Context Extraction Multiple-Choice Open-Form Math
96.83 96.81 98.70
Average
97.45
Table 12: Accuracy between human and As described in § 2.4, the reported accuracies in synthetic labels across task categories. the main text are computed using synthetic labels generated by Nemotron-Super-v1.5. To estimate performance with respect to human annotations, we can apply a correction based on the observed agreement between human and synthetic labels. Let A H denote the accuracy with respect to human labels (unknown), AS the accuracy against synthetic labels, and ρ the agreement rate between synthetic and human judgments. Let YH , YS , and Ŷ be random variables representing, for a given example, the human label, the synthetic label, and the predicted label, respectively. A H = P Ŷ = YH
(1)
= P Ŷ = YH |YH = YS P (YH = YS ) + P Ŷ = YH |YH ̸= YS P (YH ̸= YS ) = P Ŷ = YS |YH = YS P (YH = YS ) + P Ŷ ̸= YS |YH ̸= YS P (YH ̸= YS ) = P Ŷ = YS P (YH = YS ) + P Ŷ ̸= YS P (YH ̸= YS ) = AS ρ + (1 − AS )(1 − ρ) = (2ρ − 1) AS + 1 − ρ
(2) (3) (4) (5) (6)
The final expression (Equation 6)12 can be interpreted as pulling the estimated accuracy A H toward random guessing. If human and synthetic labels are uncorrelated (ρ = 0.5), the estimated accuracy drops to a random guess, regardless of AS (see Figure 6 for illustration). 1.0
= 1.0 = 0.975 = 0.95
AH
0.8
= 0.9 = 0.75 = 0.5
AS = 1.0 AS = 0.8 AS = 0.6
AS = 0.4 AS = 0.2 AS = 0.0
0.6 0.4 0.2 0.0
0.0
0.2
0.4
AS
0.6
0.8
1.0
0.5
0.6
0.7
0.8
0.9
1.0
Figure 6: Sensitivity of the A H estimate to variations in AS and ρ.
12 In Equation 4, we assume (Ŷ, Y ) ⊥ (Y = Y ), in line with our empirical observations. Intuitively, H S S
this means that agreement between the predicted and synthetic labels, Ŷ and YS , is independent of whether the synthetic label YS matches the human label YH .
20
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
D
Detailed Results
D.1
Regex Parsing Failures
In this section, we extend Figure 2 from the main text by presenting disaggregated regex parsing failure rates across models and tasks. Family
Size
ARC Challenge
ARC Easy
GPQA
MMLU
MMLU Pro
Truthful QA
Apertus
8B 70B
0.2 0.1
0.1 0.1
4.9 21.2
0.8 2.6
4.5 12.3
0.0 1.5
EuroLLM
1.7B 9B 22B
97.6 0.0 0.1
98.7 0.0 0.0
83.0 11.8 7.4
94.1 1.2 0.4
69.9 11.8 3.0
66.3 0.0 0.0
Falcon-3
1B 3B 7B
5.1 0.0 0.0
4.0 0.2 0.0
11.2 2.0 1.1
5.5 0.2 0.1
31.9 2.5 1.3
3.7 0.1 0.1
Gemma-3
1B 4B 12B 27B
0.1 0.0 0.0 0.0
0.0 0.0 0.0 0.0
12.5 6.2 4.9 3.6
2.3 0.3 0.1 0.1
14.5 3.3 3.0 1.3
0.0 0.0 0.0 0.0
LFM-2
0.35B 0.7B 1.2B 2.6B
0.4 2.9 0.5 0.0
0.4 2.8 0.5 0.2
7.1 18.1 6.7 10.0
1.2 6.1 1.0 1.4
11.3 15.5 5.4 9.9
4.0 1.6 0.4 0.6
Llama-3
1B 3B 8B 70B
5.0 8.1 0.1 0.0
3.6 5.2 0.3 0.0
36.4 32.4 25.2 7.4
9.0 21.8 2.6 0.9
26.0 47.4 16.1 6.0
5.4 16.9 2.3 0.0
Ministral-3
3B 8B 14B
0.1 0.0 0.2
0.0 0.0 0.1
23.9 25.0 16.5
1.2 1.4 1.3
15.1 10.3 8.6
0.0 0.0 0.0
OLMo-3
7B 32B
0.7 0.3
0.4 0.0
22.3 51.1
1.9 1.6
15.8 17.4
0.0 0.0
Phi-4
3.6B 14B
0.0 0.0
0.0 0.0
2.5 0.9
0.2 0.0
3.0 0.6
0.0 0.0
Qwen-3
0.6B 4B 8B 14B 32B
0.0 0.0 0.1 0.1 0.0
0.0 0.0 0.0 0.0 0.0
0.0 9.6 10.9 11.4 7.8
0.1 0.4 0.7 0.3 0.1
0.0 4.6 5.9 4.7 4.3
0.0 0.0 0.0 0.0 0.0
SmolLM-2/3
0.135B 0.36B 1.7B 3B
98.4 6.7 0.0 0.7
98.4 5.2 0.0 0.5
97.8 43.1 0.2 19.4
97.6 12.2 0.2 5.1
97.8 37.0 0.5 16.9
98.3 7.1 0.2 1.6
Table 13: Parsing failure rates on multiple-choice benchmarks.
21
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
CoQA
DROP
HotpotQA
SQuAD-v2
Apertus
8B 70B
0.0 0.0
0.2 0.3
0.1 0.0
0.1 0.1
EuroLLM
1.7B 9B 22B
41.4 0.0 0.2
39.7 0.0 0.0
16.3 0.0 0.3
38.6 0.0 0.0
Falcon-3
1B 3B 7B
20.6 2.0 0.6
6.2 0.8 0.5
3.9 1.2 0.2
5.2 3.6 0.2
Gemma-3
1B 4B 12B 27B
0.0 0.0 0.2 0.0
0.8 0.0 0.0 0.0
0.1 0.1 0.3 0.0
0.3 0.0 0.1 0.0
LFM-2
0.35B 0.7B 1.2B 2.6B
1.8 20.8 0.2 1.8
2.2 1.3 0.5 2.8
4.1 0.5 1.5 1.6
3.5 1.1 0.2 3.2
Llama-3
1B 3B 8B 70B
31.2 34.6 6.4 0.0
53.0 5.3 7.7 0.0
28.9 17.1 3.5 0.0
33.9 27.9 9.8 0.3
Ministral-3
3B 8B 14B
0.0 0.0 0.0
0.1 0.0 0.0
0.1 0.0 0.0
0.1 0.0 0.3
OLMo-3
7B 32B
0.0 0.0
2.2 0.1
0.5 0.1
1.8 2.3
Phi-4
3.6B 14B
0.0 0.0
0.0 0.0
0.0 0.0
0.0 0.0
Qwen-3
0.6B 4B 8B 14B 32B
0.0 0.0 0.0 0.0 0.0
0.0 0.0 0.1 0.0 0.0
0.0 0.0 0.0 0.0 0.0
0.0 0.0 0.0 0.0 0.0
SmolLM-2/3
0.135B 0.36B 1.7B 3B
47.8 54.2 1.2 0.4
46.1 44.2 0.1 0.7
26.0 24.8 0.4 0.1
50.8 48.1 3.3 0.3
Table 14: Parsing failure rates on context extraction benchmarks.
22
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
AIME24
AIME25
ASDiv
GSM8K
Math
Apertus
8B 70B
50.0 26.7
46.7 36.7
0.9 1.3
0.8 0.4
13.3 15.5
EuroLLM
1.7B 9B 22B
70.0 63.3 33.3
76.7 50.0 30.0
5.3 16.9 2.5
6.3 16.2 6.8
54.1 43.7 11.7
Falcon-3
1B 3B 7B
83.3 46.7 13.3
90.0 40.0 16.7
69.0 0.9 0.0
59.9 0.5 0.2
77.7 16.2 2.9
Gemma-3
1B 4B 12B 27B
93.3 50.0 33.3 33.3
83.3 16.7 33.3 23.3
10.3 0.3 0.1 0.2
18.9 0.7 0.2 0.2
47.5 8.9 9.4 5.2
LFM-2
0.35B 0.7B 1.2B 2.6B
73.3 60.0 33.3 33.3
66.7 66.7 33.3 43.3
3.2 1.3 5.1 0.3
4.5 2.8 12.2 0.4
56.2 29.3 15.1 13.3
Llama-3
1B 3B 8B 70B
100.0 53.3 53.3 100.0
100.0 70.0 56.7 100.0
25.9 1.8 3.1 6.7
28.1 1.3 1.2 25.4
84.9 20.2 18.5 92.1
Ministral-3
3B 8B 14B
96.7 86.7 93.3
86.7 90.0 86.7
0.8 0.7 0.4
1.6 1.4 1.1
26.6 19.1 23.4
OLMo-3
7B 32B
60.0 80.0
46.7 83.3
0.3 0.3
1.0 1.3
11.1 15.8
Phi-4
3.6B 14B
40.0 16.7
23.3 16.7
0.1 0.0
0.1 0.0
18.7 3.2
Qwen-3
0.6B 4B 8B 14B 32B
83.3 56.7 56.7 50.0 50.0
73.3 63.3 60.0 46.7 43.3
0.5 0.2 0.3 0.1 0.1
2.0 0.6 0.1 0.0 0.6
34.5 8.7 7.9 6.1 6.3
SmolLM-2/3
0.135B 0.36B 1.7B 3B
70.0 70.0 63.3 50.0
66.7 63.3 80.0 43.3
36.4 15.8 2.4 0.9
40.6 22.6 5.8 1.4
68.2 52.3 38.7 20.4
Table 15: Parsing failure rates on open-form math benchmarks.
23
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
D.2
Impact of Regex-Based Evaluation on Performance Measurement
This section extends Table 1 from the main text by providing a breakdown of how regexbased evaluation affects downstream measured performance across every model and task. Family
Size
Apertus
CoQA
DROP
HotpotQA
SQuAD-v2
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
8B 70B
-30.2 (-0.0/-30.2) -41.4 (-0.0/-41.4)
↓3.0 ↓7.5
-21.8 (-0.0/-21.8) -21.9 (-0.1/-21.9)
↑1.5 ↑3.0
-20.6 (-0.1/-20.5) -22.7 (-0.0/-22.7)
↑2.0 ↓3.0
-19.2 (-0.0/-19.2) -25.3 (-0.0/-25.3)
↑10.0 ↓1.0
EuroLLM
1.7B 9B 22B
-49.6 (-30.6/-19.0) -27.6 (-0.0/-27.6) -24.8 (-0.2/-24.6)
↓8.0 ↓1.0 ↑1.0
-19.5 (-11.6/-7.9) -18.9 (-0.0/-18.9) -25.4 (-0.0/-25.4)
↓1.0 ↑1.0 0.0
-37.2 (-9.9/-27.3) -15.9 (-0.0/-15.9) -16.3 (-0.2/-16.1)
↓6.0 ↑5.0 ↑1.0
-30.0 (-16.6/-13.4) -22.6 (-0.0/-22.6) -29.2 (-0.0/-29.2)
↓7.0 ↑6.0 0.0
Falcon-3
1B 3B 7B
-42.6 (-16.2/-26.4) -34.0 (-2.0/-32.0) -19.8 (-0.4/-19.4)
↓7.5 ↓1.0 ↑4.0
-11.5 (-1.7/-9.8) -18.6 (-0.5/-18.1) -14.6 (-0.3/-14.3)
0.0 0.0 ↑6.0
-17.4 (-2.2/-15.3) -23.4 (-1.0/-22.4) -12.4 (-0.2/-12.2)
0.0 ↓1.0 ↑4.0
-19.3 (-2.1/-17.2) -21.7 (-1.7/-20.0) -18.0 (-0.1/-17.9)
0.0 ↑2.0 ↑14.0
Gemma-3
1B 4B 12B 27B
-19.0 (-0.0/-19.0) -43.0 (-0.0/-43.0) -28.8 (-0.2/-28.6) -28.8 (-0.0/-28.8)
↑10.5 ↓5.0 ↓5.0 ↓1.0
-6.0 (-0.1/-5.9) -27.4 (-0.0/-27.4) -31.8 (-0.0/-31.8) -31.1 (-0.0/-31.1)
↑2.0 ↑2.0 ↓5.0 ↓4.0
-9.9 (-0.1/-9.8) -23.8 (-0.1/-23.8) -20.1 (-0.3/-19.9) -25.4 (-0.0/-25.4)
↑2.0 0.0 ↓7.0 ↓2.5
-14.1 (-0.1/-14.0) -29.0 (-0.0/-28.9) -29.5 (-0.1/-29.4) -29.4 (-0.0/-29.4)
↑4.0 ↓3.0 0.0 ↓1.0
LFM-2
0.35B 0.7B 1.2B 2.6B
-20.0 (-1.4/-18.6) -33.6 (-15.8/-17.8) -14.8 (-0.0/-14.8) -29.6 (-1.6/-28.0)
↑8.5 ↑1.0 ↑14.5 ↑2.5
-3.7 (-0.6/-3.1) -5.5 (-0.4/-5.1) -8.9 (-0.1/-8.8) -28.6 (-1.2/-27.4)
↑3.0 ↑2.0 ↑2.0 ↓2.0
-8.2 (-2.0/-6.1) -10.3 (-0.3/-10.0) -9.9 (-1.2/-8.7) -22.6 (-1.1/-21.5)
↑3.0 ↑2.0 ↑7.0 0.0
-10.3 (-1.5/-8.8) -11.4 (-0.6/-10.7) -13.2 (-0.0/-13.1) -33.5 (-1.2/-32.4)
↑5.0 ↑3.0 ↑13.0 ↓9.0
Llama-3
1B 3B 8B 70B
-30.0 (-21.0/-9.0) -40.4 (-29.2/-11.2) -23.6 (-4.6/-19.0) -17.2 (-0.0/-17.2)
↑3.0 ↓4.5 ↓1.0 ↑7.0
-24.1 (-20.4/-3.7) -17.9 (-3.0/-14.9) -21.5 (-3.7/-17.8) -17.0 (-0.0/-17.0)
↓6.0 ↑8.0 ↑3.0 ↑3.0
-22.0 (-17.2/-4.8) -22.6 (-13.4/-9.3) -17.1 (-2.0/-15.1) -11.5 (-0.0/-11.5)
0.0 0.0 ↑3.0 ↑10.0
-19.7 (-13.9/-5.8) -29.1 (-17.0/-12.1) -21.9 (-5.9/-16.0) -27.8 (-0.3/-27.5)
↓2.0 ↓5.0 ↑9.0 ↑6.0
Ministral
3B 8B 14B
-36.8 (-0.0/-36.8) -30.0 (-0.0/-30.0) -27.4 (-0.0/-27.4)
↓1.0 ↑2.0 ↓2.0
-32.9 (-0.0/-32.9) -35.0 (-0.0/-35.0) -30.7 (-0.0/-30.7)
↓6.0 ↓9.0 ↓5.0
-29.8 (-0.0/-29.7) -24.9 (-0.0/-24.9) -23.4 (-0.0/-23.4)
↓5.0 ↓6.0 ↓9.0
-30.6 (-0.0/-30.6) -34.0 (-0.0/-34.0) -36.3 (-0.0/-36.3)
↓8.0 ↓6.0 ↓7.0
OLMo-3
7B 32B
-17.6 (-0.0/-17.6) -24.8 (-0.0/-24.8)
↑5.0 ↓3.5
-22.5 (-0.8/-21.7) -27.3 (-0.0/-27.2)
↑4.0 ↓2.0
-11.7 (-0.0/-11.6) -16.1 (-0.0/-16.0)
↑9.5 0.0
-26.4 (-0.2/-26.2) -35.6 (-2.2/-33.4)
↑7.0 ↓4.0
Phi-4
3.6B 14B
-42.0 (-0.0/-42.0) -36.0 (-0.0/-36.0)
↓8.5 ↓0.5
-23.2 (-0.0/-23.2) -32.0 (-0.0/-32.0)
↑1.0 ↓4.0
-29.4 (-0.0/-29.4) -27.8 (-0.0/-27.8)
↓3.0 ↓8.0
-35.1 (-0.0/-35.1) -45.5 (-0.0/-45.5)
↓12.0 ↓18.0
Qwen-3
0.6B 4B 8B 14B 32B
-27.2 (-0.0/-27.2) -23.0 (-0.0/-23.0) -30.6 (-0.0/-30.6) -19.8 (-0.0/-19.8) -23.0 (-0.0/-23.0)
↑4.0 ↑6.0 ↓8.0 0.0 ↓2.5
-9.4 (-0.0/-9.4) -20.0 (-0.0/-20.0) -34.5 (-0.0/-34.4) -19.9 (-0.0/-19.9) -25.3 (-0.0/-25.3)
↑1.0 ↑6.0 ↓9.0 ↓1.0 ↑2.0
-7.3 (-0.0/-7.3) -12.9 (-0.0/-12.9) -19.7 (-0.0/-19.7) -12.8 (-0.0/-12.8) -17.1 (-0.0/-17.1)
↑3.0 ↑7.0 ↓4.0 0.0 ↓2.0
-11.1 (-0.0/-11.1) -25.0 (-0.0/-25.0) -34.2 (-0.0/-34.2) -31.1 (-0.0/-31.1) -35.1 (-0.0/-35.1)
↑3.0 ↑9.0 ↓8.0 ↑3.0 ↓2.0
SmolLM-2/3
0.135B 0.36B 1.7B 3B
-38.0 (-24.2/-13.8) -55.0 (-35.2/-19.8) -25.2 (-1.0/-24.2) -35.2 (-0.4/-34.8)
0.0 ↓6.0 ↑6.5 ↑1.0
-11.5 (-6.8/-4.7) -17.6 (-11.4/-6.1) -8.4 (-0.1/-8.3) -20.5 (-0.1/-20.3)
↑1.0 ↓2.0 0.0 ↑4.5
-31.7 (-10.1/-21.6) -36.5 (-12.7/-23.8) -14.7 (-0.3/-14.4) -22.3 (-0.0/-22.2)
↓1.0 ↓4.0 ↑3.0 0.0
-23.9 (-12.5/-11.4) -27.9 (-16.0/-11.9) -18.3 (-1.5/-16.8) -25.4 (-0.0/-25.3)
0.0 ↓3.0 0.0 ↑2.0
Table 16: Effect of regex-based evaluation on performance measurement for context extraction benchmarks.
24
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
Apertus
ARC-Challenge
ARC-Easy
GPQA
MMLU
MMLU-Pro
TruthfulQA
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
8B 70B
0.0 (-0.1/+0.1) -1.8 (-0.1/-1.7)
↑3.0 ↑3.0
0.0 (-0.0/0.0) -1.5 (-0.0/-1.5)
↑6.0 ↑2.0
+5.8 (-0.0/+5.8) -1.3 (-1.3/0.0)
↑16.0 ↑8.5
-0.9 (-0.1/-0.8) -3.5 (-0.4/-3.1)
↑4.0 ↑3.0
+0.9 (-0.5/+1.4) -2.9 (-1.2/-1.7)
↑4.0 ↑1.0
-0.5 (-0.0/-0.5) -3.3 (-0.5/-2.8)
0.0 ↓2.0
EuroLLM
1.7B 9B 22B
-29.8 (-28.2/-1.5) -0.3 (-0.0/-0.3) -8.6 (-0.1/-8.5)
↓1.0 ↑4.0 ↓6.0
-29.2 (-28.4/-0.8) -0.0 (-0.0/-0.0) -7.7 (-0.0/-7.7)
↓1.0 ↑6.0 ↓8.0
-28.3 (-19.2/-9.2) -8.9 (-2.7/-6.3) -9.4 (-1.1/-8.3)
↓17.5 ↑0.5 ↓2.5
-27.4 (-24.6/-2.9) -1.6 (-0.5/-1.2) -12.7 (-0.1/-12.6)
↓1.0 ↑2.0 ↓8.0
-22.4 (-9.3/-13.1) -8.2 (-2.7/-5.5) -10.0 (-0.7/-9.3)
↓7.0 0.0 ↓2.0
-29.0 (-16.3/-12.7) +0.2 (-0.0/+0.2) -6.2 (-0.0/-6.2)
↓5.0 ↑3.0 ↓1.0
Falcon-3
1B 3B 7B
-26.7 (-2.6/-24.1) -3.2 (-0.0/-3.2) -5.2 (-0.0/-5.2)
0.0 ↑2.0 ↓2.0
-33.1 (-3.0/-30.1) -2.2 (-0.1/-2.1) -6.6 (-0.0/-6.6)
0.0 ↑2.0 ↓4.0
-14.7 (-2.2/-12.5) -4.5 (-0.2/-4.2) -6.0 (-0.7/-5.4)
↓10.5 ↑7.0 ↑2.0
-15.4 (-1.7/-13.7) -3.9 (-0.1/-3.9) -6.7 (-0.0/-6.7)
↑2.0 ↑2.0 ↓1.0
-13.8 (-6.3/-7.5) -7.5 (-0.4/-7.1) -7.1 (-0.4/-6.7)
0.0 ↓1.0 ↑1.0
-6.6 (-0.6/-6.0) -1.1 (-0.1/-1.0) -2.0 (-0.0/-2.0)
↑3.0 ↑2.0 ↑1.0
Gemma-3
1B 4B 12B 27B
+0.1 (-0.1/+0.2) +0.2 (-0.0/+0.2) +0.3 (-0.0/+0.3) +0.3 (-0.0/+0.3)
↑5.0 ↑3.0 ↑6.0 ↑5.0
-0.6 (-0.0/-0.6) 0.0 (-0.0/0.0) 0.0 (-0.0/0.0) +0.2 (-0.0/+0.2)
↑5.0 ↑6.0 ↑4.0 ↑4.0
+2.5 (-1.6/+4.0) +1.1 (-0.4/+1.6) -0.4 (-0.4/0.0) +0.2 (-0.2/+0.4)
↑11.0 ↑12.0 ↑7.0 ↑6.0
-1.0 (-0.6/-0.4) +0.9 (-0.1/+0.9) 0.0 (-0.0/0.0) +0.3 (-0.0/+0.3)
↑4.0 ↑7.0 ↑6.0 ↑4.0
-2.2 (-1.8/-0.5) -0.4 (-0.5/+0.1) -1.1 (-0.5/-0.6) -0.9 (-0.4/-0.5)
↑5.0 ↑6.0 ↑5.0 ↑5.0
+0.5 (-0.0/+0.5) +0.2 (-0.0/+0.2) +0.6 (-0.0/+0.6) +0.2 (-0.0/+0.2)
↑4.0 ↑3.0 ↑4.0 ↑4.0
LFM-2
0.35B 0.7B 1.2B 2.6B
-36.1 (-0.2/-35.9) -16.6 (-1.5/-15.1) -5.7 (-0.3/-5.5) -0.1 (-0.0/-0.1)
↑1.0 ↓2.0 ↑2.0 ↑5.5
-50.0 (-0.4/-49.6) -21.6 (-2.4/-19.2) -4.6 (-0.4/-4.2) -0.2 (-0.0/-0.2)
↑1.0 ↓1.0 ↑2.0 ↑6.5
-9.8 (-0.4/-9.4) -11.8 (-3.6/-8.3) -4.9 (-0.7/-4.2) -5.4 (-1.1/-4.2)
↑3.5 ↑1.0 ↑3.0 ↑3.5
-29.9 (-0.5/-29.4) -23.5 (-2.9/-20.6) -10.1 (-0.4/-9.8) -1.3 (-0.3/-1.0)
↑1.0 ↓3.0 ↓1.0 ↑6.0
-15.0 (-1.8/-13.2) -14.4 (-3.1/-11.3) -13.5 (-0.9/-12.6) -4.5 (-1.7/-2.8)
0.0 ↓1.0 ↓2.0 ↑3.0
-20.3 (-2.0/-18.4) -6.6 (-1.1/-5.5) -9.9 (-0.0/-9.9) -2.1 (-0.1/-2.0)
↓1.0 ↑2.0 ↓2.0 ↑1.0
Llama-3
1B 3B 8B 70B
-1.9 (-2.4/+0.5) -14.8 (-6.5/-8.3) -38.1 (-0.0/-38.1) +0.2 (-0.0/+0.2)
↑5.0 ↓1.0 ↓11.0 ↑3.0
-1.9 (-2.1/+0.2) -13.0 (-4.3/-8.8) -32.5 (-0.2/-32.4) +0.1 (-0.0/+0.1)
↑5.0 ↓1.0 ↓10.5 0.0
+2.0 (-3.8/+5.8) -2.2 (-5.8/+3.6) -5.6 (-2.0/-3.6) -3.3 (-2.7/-0.7)
↑9.5 ↑8.0 ↑4.0 0.0
-1.2 (-2.5/+1.4) -17.9 (-11.7/-6.2) -26.4 (-0.5/-25.9) -0.6 (-0.7/+0.1)
↑5.0 ↓4.0 ↓9.0 0.0
-1.7 (-3.2/+1.5) -14.7 (-14.0/-0.8) -14.3 (-2.0/-12.3) -3.5 (-2.9/-0.6)
↑4.0 ↓2.0 ↓4.0 ↑3.0
-1.0 (-1.6/+0.6) -15.9 (-8.7/-7.2) -22.8 (-1.0/-21.8) +0.4 (-0.0/+0.4)
0.0 ↓3.0 ↓8.0 ↑2.0
Ministral
3B 8B 14B
-2.0 (-0.0/-2.0) -4.4 (-0.0/-4.4) -1.8 (-0.0/-1.8)
↑3.0 0.0 0.0
-0.7 (-0.0/-0.7) -2.6 (-0.0/-2.6) -0.9 (-0.1/-0.8)
↑4.0 ↑2.0 0.0
-11.4 (-6.7/-4.7) -8.7 (-5.6/-3.1) -10.3 (-3.8/-6.5)
↑0.5 ↑2.0 ↓1.0
-2.1 (-0.3/-1.7) -8.3 (-0.6/-7.7) -5.7 (-0.5/-5.1)
↑2.0 ↑1.0 0.0
-8.0 (-5.1/-2.9) -12.3 (-3.1/-9.2) -11.3 (-3.0/-8.3)
↑3.0 ↑1.0 ↑1.0
-1.5 (-0.0/-1.5) -9.1 (-0.0/-9.1) -8.1 (-0.0/-8.1)
↑0.5 ↓5.0 ↓2.0
OLMo-3
7B 32B
-0.3 (-0.2/-0.1) +0.2 (-0.1/+0.3)
↑5.0 ↑3.0
-0.0 (-0.0/-0.0) 0.0 (-0.0/0.0)
↑7.0 ↑4.0
-8.0 (-6.7/-1.3) -25.9 (-22.3/-3.6)
↑1.5 ↓11.0
-0.9 (-0.5/-0.4) -1.6 (-0.6/-1.1)
↑6.0 ↑4.0
-8.4 (-5.8/-2.6) -13.0 (-9.5/-3.5)
↑1.0 ↓2.0
+0.4 (-0.0/+0.4) 0.0 (-0.0/0.0)
↑3.0 0.0
Phi-4
3.6B 14B
-0.4 (-0.0/-0.4) -6.7 (-0.0/-6.7)
↑5.0 ↓7.5
-0.7 (-0.0/-0.7) -8.0 (-0.0/-8.0)
↑4.0 ↓13.0
+3.1 (-0.2/+3.3) -10.5 (-0.2/-10.3)
↑14.0 ↓1.0
-3.0 (-0.1/-2.9) -12.3 (-0.0/-12.3)
↑4.0 ↓4.0
-1.9 (-0.4/-1.6) -6.5 (-0.1/-6.4)
↑4.0 ↓1.0
-0.1 (-0.0/-0.1) -10.0 (-0.0/-10.0)
↑1.5 ↓5.0
Qwen-3
0.6B 4B 8B 14B 32B
-31.2 (-0.0/-31.2) -0.2 (-0.0/-0.2) -2.2 (-0.0/-2.2) -16.0 (-0.0/-16.0) -19.5 (-0.0/-19.5)
↓2.0 ↑4.0 ↑3.0 ↓18.0 ↓20.5
-43.3 (-0.0/-43.3) +0.1 (-0.0/+0.1) -0.9 (-0.0/-0.9) -11.2 (-0.0/-11.2) -13.0 (-0.0/-13.0)
↓2.0 ↑5.0 ↑2.0 ↓18.0 ↓20.0
-14.7 (-0.0/-14.7) -7.1 (-2.0/-5.1) -11.6 (-2.0/-9.6) -26.1 (-2.9/-23.2) -38.6 (-1.3/-37.3)
↓11.0 ↑4.0 ↓2.0 ↓15.0 ↓26.0
-19.3 (-0.0/-19.3) -3.1 (-0.1/-3.0) -7.8 (-0.3/-7.5) -21.6 (-0.0/-21.5) -27.9 (-0.0/-27.9)
0.0 ↑3.0 0.0 ↓14.0 ↓19.0
-8.0 (-0.0/-8.0) -8.5 (-0.9/-7.6) -15.1 (-1.5/-13.7) -30.4 (-1.4/-29.0) -37.7 (-1.3/-36.4)
0.0 ↑2.0 ↓1.0 ↓10.0 ↓19.0
-8.4 (-0.0/-8.4) -0.1 (-0.0/-0.1) -4.0 (-0.0/-4.0) -10.3 (-0.0/-10.3) -5.8 (-0.0/-5.8)
↓2.0 ↑3.0 ↑2.0 ↓4.0 ↓1.0
SmolLM-2/3
0.135B 0.36B 1.7B 3B
-23.2 (-22.9/-0.3) -23.9 (-1.7/-22.2) -57.2 (-0.0/-57.2) -0.3 (-0.1/-0.3)
↑1.0 ↑2.0 ↓7.0 ↑4.5
-24.0 (-23.7/-0.3) -26.5 (-1.3/-25.2) -79.2 (-0.0/-79.2) -0.4 (-0.1/-0.3)
↑2.5 ↑1.5 ↓9.0 ↑6.0
-20.8 (-20.5/-0.2) -21.2 (-10.0/-11.2) -27.7 (-0.2/-27.5) -6.9 (-3.3/-3.6)
↓6.0 ↓6.0 ↓15.5 ↑0.5
-22.7 (-22.0/-0.7) -24.1 (-2.8/-21.4) -44.9 (-0.1/-44.8) -3.0 (-1.7/-1.3)
0.0 ↑2.0 ↓7.0 ↑3.0
-11.9 (-11.6/-0.3) -9.2 (-4.0/-5.2) -15.9 (-0.1/-15.8) -4.9 (-3.1/-1.8)
↓1.0 ↑3.0 ↓1.0 ↑2.0
-18.2 (-17.9/-0.4) -22.3 (-0.9/-21.4) -18.4 (-0.1/-18.2) -1.7 (-0.4/-1.3)
↑1.0 0.0 0.0 ↑1.0
Table 17: Effect of regex-based evaluation on performance measurement for multiple-choice benchmarks.
Family
Size
Apertus
AIME24
AIME25
ASDiv
GSM8K
Math
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
∆Accuracy
∆Rank
8B 70B
0.0 (-0.0/0.0) 0.0 (-0.0/0.0)
↑4.0 ↑4.0
0.0 (-0.0/0.0) 0.0 (-0.0/0.0)
↑5.0 ↑5.0
-5.1 (-0.1/-5.0) -5.8 (-0.7/-5.1)
↑3.0 ↑1.0
-0.4 (-0.2/-0.2) -0.3 (-0.0/-0.3)
↑3.0 ↑2.0
-10.2 (-1.2/-9.0) -13.3 (-1.3/-12.0)
↑6.0 ↑7.0
EuroLLM
1.7B 9B 22B
0.0 (-0.0/0.0) -10.0 (-6.7/-3.3) -20.0 (-3.3/-16.7)
↑4.0 ↓7.5 ↓3.0
0.0 (-0.0/0.0) -3.3 (-3.3/0.0) -6.7 (-0.0/-6.7)
↑5.0 ↓3.0 ↓7.5
-1.7 (-0.9/-0.8) -29.5 (-11.9/-17.6) -32.3 (-1.9/-30.4)
↑1.0 ↓4.0 ↓8.0
-0.8 (-0.2/-0.7) -27.5 (-10.7/-16.8) -29.7 (-5.4/-24.3)
↑1.0 ↓6.0 ↓8.0
-2.7 (-2.0/-0.7) -30.2 (-15.5/-14.7) -38.6 (-4.2/-34.4)
0.0 ↓1.0 ↓7.5
Falcon-3
1B 3B 7B
-3.3 (-3.3/0.0) -6.7 (-0.0/-6.7) -6.7 (-0.0/-6.7)
↓1.0 ↓4.0 ↑4.5
-3.3 (-0.0/-3.3) -6.7 (-0.0/-6.7) -6.7 (-0.0/-6.7)
↓3.0 ↓7.5 ↑3.5
-51.1 (-45.0/-6.1) -6.7 (-0.6/-6.1) -5.8 (-0.0/-5.8)
↓4.0 0.0 ↑5.0
-36.7 (-31.1/-5.6) -1.9 (-0.2/-1.7) -0.2 (-0.1/-0.1)
↓4.0 ↑1.0 ↑2.0
-22.6 (-18.8/-3.8) -24.7 (-4.1/-20.5) -24.9 (-0.8/-24.1)
↓1.0 ↑1.0 ↑3.0
Gemma-3
1B 4B 12B 27B
-6.7 (-6.7/0.0) -6.7 (-0.0/-6.7) -6.7 (-3.3/-3.3) -10.0 (-3.3/-6.7)
↓4.0 ↑3.5 ↑9.0 ↑6.5
-3.3 (-3.3/0.0) -16.7 (-3.3/-13.3) -6.7 (-0.0/-6.7) -6.7 (-0.0/-6.7)
↓3.0 ↓3.5 ↑5.5 ↑4.5
-10.8 (-5.2/-5.7) -7.2 (-0.0/-7.1) -7.2 (-0.0/-7.2) -7.5 (-0.1/-7.4)
↓1.0 ↓4.5 ↓6.5 ↓10.0
-6.7 (-5.5/-1.3) -1.3 (-0.3/-1.0) -1.1 (-0.1/-1.0) -0.8 (-0.2/-0.6)
↑2.0 ↑1.0 ↑2.0 0.0
-28.7 (-13.3/-15.5) -27.3 (-2.5/-24.8) -29.9 (-3.6/-26.3) -29.0 (-2.0/-27.0)
0.0 ↑3.0 ↓2.0 ↑1.0
LFM-2
0.35B 0.7B 1.2B 2.6B
0.0 (-0.0/0.0) 0.0 (-0.0/0.0) -10.0 (-3.3/-6.7) -13.3 (-3.3/-10.0)
↑4.0 ↑4.0 ↓7.5 ↑5.0
0.0 (-0.0/0.0) 0.0 (-0.0/0.0) -3.3 (-0.0/-3.3) -13.3 (-10.0/-3.3)
↑5.0 ↑5.0 ↓3.0 ↓3.0
-19.5 (-1.7/-17.7) -5.3 (-0.8/-4.5) -8.5 (-3.6/-4.9) -4.7 (-0.0/-4.6)
0.0 ↑3.0 ↑1.0 ↑12.0
-22.8 (-1.2/-21.6) -1.4 (-1.1/-0.3) -7.1 (-6.7/-0.3) -0.6 (-0.1/-0.5)
↓1.0 ↑1.0 ↑1.0 ↑1.0
-24.0 (-14.0/-10.0) -16.8 (-7.5/-9.3) -20.6 (-5.0/-15.6) -30.6 (-6.9/-23.7)
↓1.0 ↑3.0 ↑3.0 ↑2.0
Llama-3
1B 3B 8B 70B
-3.3 (-3.3/0.0) -6.7 (-6.7/0.0) 0.0 (-0.0/0.0) -33.3 (-33.3/0.0)
↓1.0 ↑4.5 ↑9.5 ↓17.0
-3.3 (-3.3/0.0) -3.3 (-0.0/-3.3) 0.0 (-0.0/0.0) -10.0 (-10.0/0.0)
↓3.0 ↓3.0 ↑5.0 ↓9.5
-15.8 (-14.5/-1.3) -4.7 (-1.2/-3.5) -6.7 (-1.6/-5.1) -11.4 (-5.9/-5.5)
0.0 ↑1.0 ↑2.0 ↓6.0
-8.7 (-8.7/-0.0) -0.5 (-0.4/-0.1) -0.5 (-0.2/-0.2) -24.1 (-24.2/+0.1)
↑2.0 ↑3.0 ↑3.0 ↓18.0
-28.1 (-25.5/-2.7) -19.6 (-4.5/-15.1) -19.8 (-2.4/-17.5) -73.1 (-71.5/-1.6)
↓2.0 ↑3.0 ↑1.5 ↓16.0
Ministral
3B 8B 14B
-36.7 (-36.7/-0.0) -40.0 (-40.0/0.0) -50.0 (-50.0/0.0)
↓11.5 ↓9.0 ↓16.5
-20.0 (-20.0/0.0) -36.7 (-36.7/-0.0) -33.3 (-33.3/-0.0)
↓5.0 ↓9.5 ↓8.5
-6.3 (-0.4/-5.9) -6.1 (-0.4/-5.7) -5.9 (-0.2/-5.7)
↑1.0 ↑5.0 ↑3.0
-0.6 (-0.3/-0.3) -1.9 (-0.9/-1.0) -1.1 (-0.7/-0.5)
↑2.0 0.0 ↑0.5
-34.9 (-14.4/-20.5) -33.7 (-10.7/-23.0) -37.9 (-15.9/-22.0)
↓3.0 ↓3.0 ↓6.5
OLMo-3
7B 32B
-30.0 (-20.0/-10.0) -36.7 (-36.7/-0.0)
↓5.5 ↓6.0
-13.3 (-13.3/0.0) -23.3 (-23.3/-0.0)
↑3.0 ↓2.0
-6.4 (-0.1/-6.3) -5.9 (-0.0/-5.9)
↑0.5 ↑2.0
-0.7 (-0.2/-0.5) -0.8 (-0.5/-0.3)
↑1.0 ↓1.0
-30.3 (-5.5/-24.8) -31.7 (-9.5/-22.1)
↓0.5 ↓3.0
Phi-4
3.6B 14B
-6.7 (-6.7/0.0) -16.7 (-10.0/-6.7)
↓4.0 ↑4.5
-3.3 (-3.3/-0.0) -6.7 (-6.7/0.0)
↑2.5 ↑3.5
-5.6 (-0.0/-5.6) -6.7 (-0.0/-6.7)
↑6.0 ↓3.5
-1.5 (-0.0/-1.5) -1.3 (-0.0/-1.3)
0.0 ↑1.0
-28.7 (-9.8/-18.8) -28.3 (-1.5/-26.8)
↑1.0 ↑2.0
Qwen-3
0.6B 4B 8B 14B 32B
-10.0 (-10.0/0.0) -13.3 (-13.3/-0.0) -16.7 (-16.7/0.0) -20.0 (-16.7/-3.3) -10.0 (-6.7/-3.3)
↑0.5 ↑3.5 ↑3.5 ↓1.0 ↑5.5
-3.3 (-3.3/0.0) -10.0 (-10.0/0.0) -23.3 (-20.0/-3.3) -20.0 (-20.0/0.0) -6.7 (-3.3/-3.3)
↓3.0 ↑4.5 ↓5.0 ↓1.0 ↑5.5
-4.3 (-0.1/-4.2) -7.3 (-0.1/-7.2) -7.0 (-0.2/-6.8) -6.5 (-0.0/-6.5) -5.8 (-0.0/-5.8)
↑3.0 ↓5.0 ↓3.0 ↑1.0 0.0
-3.4 (-0.5/-3.0) -1.1 (-0.4/-0.7) -1.4 (-0.0/-1.4) -0.2 (-0.0/-0.2) -0.8 (-0.4/-0.5)
↑1.0 0.0 ↑1.5 ↑3.0 ↓1.0
-33.0 (-15.4/-17.7) -28.4 (-2.8/-25.6) -27.7 (-2.7/-25.0) -28.3 (-2.2/-26.1) -28.3 (-2.5/-25.8)
↓4.0 ↑2.0 ↑2.0 ↑2.0 ↑4.0
SmolLM-2/3
0.135B 0.36B 1.7B 3B
0.0 (-0.0/0.0) 0.0 (-0.0/0.0) 0.0 (-0.0/0.0) -10.0 (-3.3/-6.7)
↑4.0 ↑4.0 ↑10.0 ↑0.5
0.0 (-0.0/0.0) 0.0 (-0.0/0.0) 0.0 (-0.0/0.0) -3.3 (-3.3/0.0)
↑5.0 ↑5.0 ↑5.0 ↑5.5
-3.1 (-3.0/-0.1) -3.3 (-2.4/-1.0) -4.7 (-1.8/-2.9) -5.7 (-0.3/-5.4)
0.0 ↑1.0 ↑2.0 ↑2.0
-0.1 (-0.6/+0.5) -1.7 (-1.1/-0.6) -2.4 (-1.7/-0.6) -0.4 (-0.2/-0.2)
0.0 0.0 ↑2.0 ↑2.0
-1.7 (-1.7/-0.0) -4.0 (-3.2/-0.8) -14.2 (-7.9/-6.3) -33.2 (-12.6/-20.6)
0.0 ↑1.0 ↑3.0 0.0
Table 18: Effect of regex-based evaluation on performance measurement for open-form math benchmarks.
25
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
D.3
BERT-as-a-Judge vs. Regex
In this section, we provide a detailed, unaggregated comparison of regex-based evaluation and BERT-as-a-Judge, reported by model and task. The results are based on the default BERT-as-a-Judge configuration, trained on 1M question-candidate-reference triplets, with candidates generated under the soft-constraint instruction (Appendix A). Family
Size
ARC Challenge
ARC Easy
GPQA
MMLU
MMLU Pro
Truthful QA
Apertus
8B 70B
99.4 99.5
99.6 99.7
92.4 95.1
98.2 98.4
96.7 96.6
99.1 99.0
EuroLLM
1.7B 9B 22B
99.5 99.9 99.7
99.7 100.0 99.9
92.4 94.9 94.4
97.0 99.3 98.4
88.5 97.5 97.2
83.0 99.8 99.9
Falcon-3
1B 3B 7B
98.4 99.7 99.9
98.9 99.7 99.9
92.4 94.6 96.2
97.9 99.2 99.5
95.0 97.1 97.7
97.6 99.8 99.6
Gemma-3
1B 4B 12B 27B
99.4 99.8 99.6 99.7
99.8 99.9 99.9 99.8
90.6 96.0 94.6 95.5
97.5 98.3 99.1 98.9
96.5 97.0 97.6 98.0
99.5 99.8 99.4 99.8
LFM-2
0.35B 0.7B 1.2B 2.6B
98.8 99.1 98.7 99.5
98.6 99.2 99.4 99.7
89.3 91.5 94.4 94.9
95.2 96.8 97.5 98.5
93.2 96.1 96.6 96.1
93.5 98.2 98.2 99.0
Llama-3
1B 3B 8B 70B
99.1 99.5 99.9 99.7
99.5 99.7 99.8 99.8
91.5 93.5 94.0 94.4
97.7 98.0 98.9 99.0
96.5 96.5 97.2 97.6
99.1 99.1 99.0 98.4
Ministral-3
3B 8B 14B
99.5 99.7 99.7
99.8 99.8 99.9
90.4 93.5 94.2
98.7 99.1 99.2
94.0 96.3 97.2
99.3 99.3 99.0
OLMo-3
7B 32B
99.4 99.5
99.8 99.9
92.6 77.7
98.7 99.2
94.1 91.4
99.6 99.8
Phi-4
3.6B 14B
99.7 99.9
99.9 100.0
94.2 96.0
99.4 99.2
97.8 98.0
99.6 99.0
Qwen-3
0.6B 4B 8B 14B 32B
99.8 99.4 99.0 99.7 99.4
99.9 99.7 99.5 99.9 99.7
95.8 95.5 94.2 94.0 96.7
99.1 98.8 98.8 99.2 99.1
97.4 97.5 97.2 97.8 97.9
100.0 99.6 99.4 99.6 99.5
SmolLM-2/3
0.135B 0.36B 1.7B 3B
97.6 99.5 99.9 99.5
98.9 99.4 99.9 99.7
88.6 97.3 97.5 94.4
96.7 98.7 99.5 98.6
97.0 98.1 98.7 97.0
96.7 98.7 99.9 98.8
Table 19: BERT-as-a-Judge assessment accuracy on multiple-choice benchmarks.
26
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
ARC Challenge
ARC Easy
GPQA
MMLU
MMLU Pro
Truthful QA
Apertus
8B 70B
99.0 97.7
99.2 97.9
91.5 93.8
96.4 94.9
96.2 94.5
97.8 96.2
EuroLLM
1.7B 9B 22B
70.2 99.5 91.2
70.8 99.9 92.1
71.7 89.7 85.3
72.6 98.0 85.8
77.6 91.3 87.3
71.0 99.8 93.5
Falcon-3
1B 3B 7B
73.3 96.6 94.6
66.8 97.7 93.3
84.8 93.3 90.4
84.3 95.5 92.8
85.8 91.1 91.8
93.1 98.7 97.8
Gemma-3
1B 4B 12B 27B
98.7 99.8 99.6 99.7
98.9 99.9 99.9 99.8
86.4 95.8 94.2 95.8
95.3 98.2 99.1 98.9
94.1 96.9 97.6 98.1
99.5 99.8 99.4 99.8
LFM-2
0.35B 0.7B 1.2B 2.6B
63.9 82.5 92.1 99.2
49.8 78.0 94.2 99.4
87.9 86.4 89.7 91.1
69.6 75.4 87.5 97.5
84.7 84.8 85.1 94.5
79.4 92.4 87.6 96.9
Llama-3
1B 3B 8B 70B
96.9 84.2 61.8 99.7
97.6 86.4 67.3 99.8
90.4 89.7 90.4 92.6
95.9 79.9 72.7 98.0
94.7 83.4 84.3 95.6
97.8 83.8 76.7 97.4
Ministral-3
3B 8B 14B
98.0 95.5 98.2
99.1 97.3 98.9
88.6 89.5 89.3
97.3 91.5 94.2
91.8 87.5 88.6
97.3 90.5 90.5
OLMo-3
7B 32B
99.2 99.5
99.8 99.9
91.1 74.1
98.2 98.0
91.4 86.9
99.6 99.5
Phi-4
3.6B 14B
99.2 93.3
99.1 92.0
94.2 86.4
96.2 87.0
95.9 92.4
99.1 88.5
Qwen-3
0.6B 4B 8B 14B 32B
68.8 98.6 96.1 83.8 80.0
56.7 99.5 98.1 88.7 86.5
83.9 90.6 85.7 72.5 60.9
80.0 95.2 91.2 78.1 71.7
89.7 90.3 84.1 69.3 62.0
91.6 99.6 94.7 89.2 93.5
SmolLM-2/3
0.135B 0.36B 1.7B 3B
76.8 76.1 42.8 99.0
76.0 73.5 20.8 99.2
79.2 78.3 71.4 88.6
77.3 75.8 55.1 95.6
88.1 90.6 84.1 93.7
81.8 77.7 81.6 96.3
Table 20: Regex assessment accuracy on multiple-choice benchmarks.
27
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
CoQA
DROP
HotpotQA
SQuAD-v2
Apertus
8B 70B
90.2 87.4
89.1 89.8
91.9 91.8
89.9 89.0
EuroLLM
1.7B 9B 22B
91.4 90.0 90.4
91.3 89.6 89.2
92.1 93.5 93.3
91.2 89.6 89.4
Falcon-3
1B 3B 7B
90.4 88.6 92.8
90.2 89.3 92.2
90.4 89.0 93.8
91.5 89.0 89.1
Gemma-3
1B 4B 12B 27B
84.8 75.4 90.6 87.2
90.1 83.5 85.2 85.4
91.0 88.9 91.8 88.7
91.1 87.8 85.2 83.9
LFM-2
0.35B 0.7B 1.2B 2.6B
88.8 83.4 91.4 85.6
90.4 89.0 90.5 85.2
90.7 85.2 91.5 88.8
92.1 92.1 91.5 88.9
Llama-3
1B 3B 8B 70B
91.6 88.4 88.4 89.8
89.4 88.8 87.8 89.0
94.5 91.1 92.8 90.5
92.7 89.7 88.3 87.3
Ministral-3
3B 8B 14B
75.6 87.2 88.6
85.0 87.1 89.3
86.4 89.9 91.0
87.7 87.5 87.7
OLMo-3
7B 32B
92.2 92.0
85.0 81.7
92.1 91.5
88.2 87.3
Phi-4
3.6B 14B
90.0 82.2
90.5 89.2
90.8 89.4
89.8 88.2
Qwen-3
0.6B 4B 8B 14B 32B
86.0 85.4 91.6 93.0 91.2
90.5 88.8 84.6 90.6 89.2
92.9 91.4 90.5 93.5 92.2
91.5 90.2 89.0 87.3 88.4
SmolLM-2/3
0.135B 0.36B 1.7B 3B
86.8 85.2 91.8 84.8
92.0 89.7 93.5 88.7
88.3 90.1 93.8 88.1
91.1 91.2 92.0 89.9
Table 21: BERT-as-a-Judge assessment accuracy on context extraction benchmarks.
28
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
CoQA
DROP
HotpotQA
SQuAD-v2
Apertus
8B 70B
67.0 56.6
75.7 76.3
73.7 72.5
78.3 72.8
EuroLLM
1.7B 9B 22B
49.2 70.8 72.4
80.0 79.9 73.0
60.5 80.5 80.2
69.0 76.1 69.4
Falcon-3
1B 3B 7B
55.4 61.6 77.4
85.1 79.2 83.9
77.5 71.3 83.8
78.8 76.1 80.1
Gemma-3
1B 4B 12B 27B
75.0 54.6 68.8 68.4
88.8 70.3 66.9 67.4
83.4 71.2 76.0 69.6
82.9 69.5 69.0 68.3
LFM-2
0.35B 0.7B 1.2B 2.6B
76.0 62.0 80.4 66.4
90.2 86.6 86.3 69.7
84.0 78.1 82.9 72.2
86.0 84.2 84.3 64.7
Llama-3
1B 3B 8B 70B
66.4 56.8 74.0 78.4
75.1 79.4 77.1 81.1
74.8 72.8 79.9 82.3
79.0 69.6 76.5 70.2
Ministral-3
3B 8B 14B
62.8 67.2 70.6
65.5 63.8 68.2
66.6 71.0 73.0
67.5 64.5 62.3
OLMo-3
7B 32B
78.4 72.0
74.8 70.3
83.3 78.9
71.8 62.5
Phi-4
3.6B 14B
56.8 63.2
75.0 66.9
66.5 68.7
63.4 53.0
Qwen-3
0.6B 4B 8B 14B 32B
70.4 71.4 67.8 77.8 75.4
88.9 76.8 62.5 79.2 73.0
86.6 81.0 75.9 83.3 79.1
86.7 73.2 64.2 67.5 63.2
SmolLM-2/3
0.135B 0.36B 1.7B 3B
60.8 45.0 72.4 62.8
87.4 81.8 90.5 76.9
67.0 61.4 81.5 71.5
75.6 71.3 80.1 73.0
Table 22: Regex assessment accuracy on context extraction benchmarks.
29
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
AIME24
AIME25
ASDiv
GSM8K
Math
Apertus
8B 70B
100.0 100.0
100.0 100.0
95.4 94.5
98.6 99.1
93.3 92.6
EuroLLM
1.7B 9B 22B
100.0 90.0 80.0
100.0 96.7 93.3
97.8 95.4 95.1
97.6 98.6 98.9
97.0 91.8 90.2
Falcon-3
1B 3B 7B
96.7 96.7 93.3
96.7 93.3 93.3
95.6 95.4 95.2
97.9 99.4 99.8
90.6 93.3 93.9
Gemma-3
1B 4B 12B 27B
93.3 90.0 96.7 90.0
100.0 83.3 93.3 93.3
93.5 95.2 94.8 95.1
96.5 98.6 98.9 99.2
90.4 93.6 94.9 95.4
LFM-2
0.35B 0.7B 1.2B 2.6B
100.0 100.0 90.0 90.0
100.0 100.0 96.7 90.0
96.2 94.8 94.6 95.2
97.2 98.9 98.9 99.2
92.3 93.1 92.9 94.9
Llama-3
1B 3B 8B 70B
96.7 96.7 100.0 93.3
96.7 96.7 100.0 93.3
93.7 94.9 94.0 95.3
97.6 99.2 99.0 99.6
92.9 93.0 92.5 94.5
Ministral-3
3B 8B 14B
66.7 66.7 70.0
56.7 63.3 60.0
94.4 94.9 95.3
99.2 99.8 99.6
90.3 92.8 92.5
OLMo-3
7B 32B
80.0 60.0
80.0 66.7
95.5 95.6
98.9 99.3
94.7 93.6
Phi-4
3.6B 14B
96.7 83.3
100.0 100.0
95.4 95.2
99.5 99.7
94.6 95.9
Qwen-3
0.6B 4B 8B 14B 32B
86.7 86.7 83.3 80.0 93.3
100.0 93.3 83.3 80.0 93.3
95.7 94.9 95.0 95.4 95.3
97.0 99.2 99.4 99.7 99.6
92.4 95.3 95.9 95.4 95.8
SmolLM-2/3
0.135B 0.36B 1.7B 3B
100.0 100.0 100.0 93.3
100.0 100.0 100.0 96.7
98.3 98.0 95.8 95.2
98.6 98.9 98.0 99.0
97.7 95.7 93.1 94.6
Table 23: BERT-as-a-Judge assessment accuracy on open-form math benchmarks.
30
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Family
Size
AIME24
AIME25
ASDiv
GSM8K
Math
Apertus
8B 70B
100.0 100.0
100.0 100.0
94.0 92.9
99.2 99.7
88.4 85.0
EuroLLM
1.7B 9B 22B
100.0 90.0 80.0
100.0 96.7 93.3
96.4 70.2 66.9
97.5 72.5 70.3
96.5 69.4 61.1
Falcon-3
1B 3B 7B
96.7 93.3 93.3
96.7 93.3 93.3
48.6 92.5 92.5
63.3 97.8 99.8
77.0 74.7 74.5
Gemma-3
1B 4B 12B 27B
93.3 93.3 93.3 90.0
96.7 83.3 93.3 93.3
88.7 91.5 91.9 91.5
93.1 98.6 98.9 99.2
71.1 72.3 69.8 70.6
LFM-2
0.35B 0.7B 1.2B 2.6B
100.0 100.0 90.0 86.7
100.0 100.0 96.7 86.7
79.8 93.1 90.5 92.8
76.6 97.9 92.6 99.4
75.9 81.8 78.1 69.2
Llama-3
1B 3B 8B 70B
96.7 93.3 100.0 66.7
96.7 96.7 100.0 90.0
83.7 93.8 91.9 87.3
90.8 99.4 99.4 75.7
71.8 79.9 79.9 26.9
Ministral-3
3B 8B 14B
63.3 60.0 50.0
80.0 63.3 66.7
91.4 91.7 91.8
99.4 98.1 98.9
64.9 66.0 62.0
OLMo-3
7B 32B
70.0 63.3
86.7 76.7
92.2 92.4
99.3 99.2
69.3 68.3
Phi-4
3.6B 14B
93.3 83.3
96.7 93.3
92.2 92.1
98.2 98.7
70.7 71.3
Qwen-3
0.6B 4B 8B 14B 32B
90.0 86.7 83.3 80.0 90.0
96.7 90.0 76.7 80.0 93.3
93.9 91.3 91.7 91.9 91.8
96.4 98.9 98.5 99.6 99.2
66.8 71.3 71.8 71.3 71.3
SmolLM-2/3
0.135B 0.36B 1.7B 3B
100.0 100.0 100.0 90.0
100.0 100.0 100.0 96.7
95.9 95.5 94.0 92.6
98.7 98.0 96.9 99.6
97.5 94.6 85.1 66.3
Table 24: Regex assessment accuracy on open-form math benchmarks.
31
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
E
Illustrative Examples
In this section, we present examples of common failure cases in regex-based evaluation. Table 25 illustrates a case where parsing fails despite the model producing a correct answer, while Table 26 shows a case where parsing succeeds but additional formatting introduced by the model prevents correct assessment against the reference. Question
Answer the question based on the provided context. Context: A psychological identity relates to self-image (one's mental model of oneself), self-esteem, and individuality. Consequently, Weinreich gives the definition "A person's identity is defined as the totality of one's self-construal, in which how one construes oneself in the present expresses the continuity between how one construes oneself as one was in the past and how one construes oneself as one aspires to be in the future"; this allows for definitions of aspects of identity, such as: "One's ethnic identity is defined as that part of the totality of one's self-construal made up of those dimensions that express the continuity between one's construal of past ancestry and one's future aspirations in relation to ethnicity" (Weinreich, 1986a). The description or representation of individual and group identity is a central task for psychologists, sociologists and anthropologists and those of other disciplines where "identity" needs to be mapped and defined. How should one describe the identity of another, in ways which encompass both their idiosyncratic qualities and their group memberships or identifications, both of which can shift according to circumstance? Following on from the work of Kelly, Erikson, Tajfel and others Weinreich's Identity Structure Analysis (ISA), is "a structural representation of the individual's existential experience, in which the relationships between self and other agents are organised in relatively stable structures over time . . . with the emphasis on the socio-cultural milieu in which self relates to other agents and institutions" (Weinreich and Saunderson, (eds) 2003, p1). Using constructs drawn from the salient discourses of the individual, the group and cultural norms, the practical operationalisation of ISA provides a methodology that maps how these are used by the individual, applied across time and milieus by the "situated self" to appraise self and other agents and institutions (for example, resulting in the individual's evaluation of self and significant others and institutions).[citation needed] Question: What is the ISA? Conclude your response with "Final answer: X", where X is the exact span from the context that answers the question.
Reference
Identity Structure Analysis (ISA)
Candidate
The ISA, or Identity Structure Analysis, is a methodology developed by Weinreich that maps how individuals perceive themselves and others in relation to their social environment, including their own group memberships and identifications, as well as the cultural norms and discourses that shape these perceptions.
Table 25: Example 488 from the CoQA dataset, along with the answer from Llama-3 1B
32
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Question
We have a triangle $\triangle ABC$ such that $AB = 6,$ $BC = 8,$ and $CA = 10.$ If $AD$ is an angle bisector such that $D$ is on $BC,$ then find the value of $ADˆ2.$ Conclude your response with "Final answer: X", where X is the computed solution.
Reference
45
Candidate
To solve for \( ADˆ2 \) in \(\triangle ABC\) where \( AB = 6 \), \( BC = 8 \), and \( CA = 10 \), and \( AD \) is the angle bisector of \(\angle BAC\) with \( D \) on \( BC \), we can use the Angle Bisector Theorem and the Angle Bisector Length Formula. ### Step-by-Step Solution: 1. **Angle Bisector Theorem:** The Angle Bisector Theorem states that the angle bisector divides the opposite side into segments proportional to the adjacent sides. Therefore, if \( D \) divides \( BC \) into segments \( BD \) and \( DC \), then: \[ \frac{BD}{DC} = \frac{AB}{AC} = \frac{6}{10} = \frac{3}{5} \] [...] 4. **Calculate \( ADˆ2 \):** \[ ADˆ2 = 6 \cdot 10 \left( 1 - \frac{8ˆ2}{(6 + 10)ˆ2} \right) \] \[ ADˆ2 = 60 \left( 1 - \frac{64}{16ˆ2} \right) \] \[ ADˆ2 = 60 \left( 1 - \frac{64}{256} \right) \] \[ ADˆ2 = 60 \left( 1 - \frac{1}{4} \right) \] \[ ADˆ2 = 60 \left( \frac{3}{4} \right) \] \[ ADˆ2 = 60 \cdot 0.75 \] \[ ADˆ2 = 45 \] Final answer: \( \boxed{45} \)
Table 26: Example 2070 from the MATH dataset, along with the answer from Falcon-3 7B
33