ConceptioArchivearXiv CS
arXiv CSopen access

Cost-Effective Model Evaluation with Meta-Learning

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Cost-Effective Model Evaluation with Meta-Learning Trinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen, Thanh Tam Nguyen

arXiv:2605.23595v1 [cs.LG] 22 May 2026

Abstract The rapid growth of machine learning has produced an ever-expanding ecosystem of models, making it increasingly challenging to verify the reliability of newly released models on unseen, unlabeled data. Conventional evaluation pipelines depend on expensive annotation, repeated fine-tuning, or narrow assumptions that fail to transfer across model families. We present MetaEvaluator, a cost-effective, model-agnostic framework for rapid, label-free assessment of unseen models spanning diverse architectures and modalities. MetaEvaluator leverages meta-learning over a pool of reference models to obtain a transferable initialization, enabling accurate evaluation of new models while amortizing cost across the pool and removing the need for per-model retraining. To the best of our knowledge, this is the first model-agnostic framework capable of evaluating new models on entirely unlabeled datasets. Extensive experiments show that MetaEvaluator produces stable and accurate performance estimates at substantially reduced cost compared to conventional approaches, making scalable benchmarking of emerging models on unlabeled data practical.

Keywords model evaluation, meta-learning, unseen models, unlabeled data

1

Introduction

Recent progress in machine learning is driven by large pretrained model families—spanning domains from recommendation [40, 51, 70, 77, 78] to graph learning [8, 9, 23, 50, 75]—and rapidly growing data collections, most of which remain unlabeled [18, 37]. This trend introduces a core deployment problem for organizations: choosing among many newly released and unseen models for an unlabeled workload. Consider an organization that deploys a new Text-to-SQL model to query an internal enterprise database. No labeled question–SQL pairs exist, and manual annotation would require domain experts and weeks of effort [21, 44–46], making rapid model selection impractical. This setting exposes a double challenge: the model is unseen, and the target dataset is entirely unlabeled. In practice, such model-selection decisions are increasingly intertwined with trustworthiness concerns—including robustness to poisoning, federated and privacy-preserving deployment, and explainability [25, 26, 41, 42, 48, 52–54, 58, 81]—which makes reliable, label-free assessment before deployment all the more valuable. Most existing evaluation pipelines nevertheless operate on one model at a time. Some require human or pseudo-labeling [1, 2, 13, 79]. Others rely on repeated fine-tuning [3, 7, 27, 72, 86]. Several approaches quantify train–test shifts for an individual model [7, 16, 92]. Judgebased systems using large language models (LLMs) have also been proposed [15, 38, 91]. However, these techniques are designed for fixed architectures and model-specific behaviors, and they incur substantial computational cost [71] and human labor overhead [22, 57, 80, 82, 89] when applied to each new model. Inspired by findings [73, 88] that pretrained systems exhibit structured and predictable performance trends across architectures and domains

Figure 1: Unlike existing methods that struggle to assess unseen models without labels, MetaEvaluator leverages metalearning to efficiently estimate accuracy for unseen models. rather than arbitrary variation, we pose the central question of this work: Can we learn to evaluate unseen models on unlabeled data by transferring knowledge from previously evaluated models? As illustrated in Fig. 1, we answer this question by introducing MetaEvaluator, a model-agnostic framework designed to generalize performance estimation to newly arriving, unseen models on unlabeled target workloads without relying on extensive human annotation or repeated per-model training. MetaEvaluator reframes evaluation as a meta-learning problem: it learns transferable performance patterns from a shared pool of reference models that have been systematically evaluated across diverse datasets, architectures, and distribution shifts. By distilling these patterns into compact context representations, MetaEvaluator can rapidly adapt to a new model on an unlabeled dataset. Across all settings, MetaEvaluator produces predictions that closely track ground-truth performance while substantially reducing evaluation overhead. This design amortizes cost across reference models and enables scalable assessment in rapidly evolving model ecosystems. Our main contributions are: • Formulation: We formulate the double challenge of estimating performance for unseen models on unlabeled target data and design a method that generalizes across heterogeneous architectures and modalities, including Text2SQL and Image Classification. • Method: We propose MetaEvaluator, a model-agnostic metalearning framework that learns how performance varies across models and shifts by transferring knowledge from a pool of reference models, enabling rapid adaptation to newly released architectures on unlabeled workloads without permodel retraining. • Dataset: We introduce MetaDataset, a large-scale and systematically constructed corpus of model–shift pairs that

Preprint, 2026,

spans Text2SQL and Image Classification, providing diverse unlabeled deployment scenarios and accurate performance supervision for MetaEvaluator. • Benchmarking: MetaEvaluator enables a lightweight, fast, and automated benchmarking framework that can ingest newly released models and promptly return accurate performance estimates on unlabeled workloads, supporting rapid deployment cycles and model-selection feedback in real organizational settings.

2

Related Work

Most existing label-free evaluation methods assess a single fixed model and do not handle the harder setting in which both the model and the target dataset are unseen at deployment time. Model ecosystems expand rapidly and pipelines that rely on per-model fine-tuning quickly become impractical. To the best of our knowledge, no prior work addresses this double challenge: evaluating unseen models on unlabeled data across modalities. Therefore, we study both Text2SQL and Image Classification to pursue a unified solution that remains effective as architectures evolve. In Image Classification, AutoEval [7] and DoC [16] estimate accuracy from representation-level distribution distances or confidence shifts, but both are trained for a specific backbone and must be retrained for each new architecture. SelfTrainEns [3] estimates accuracy through agreement patterns among auxiliary ensembles trained on the same task, but this substantially increases computational cost because multiple models must be trained solely for evaluation. ATC [14] further reduces overhead by selecting a confidence threshold and transferring it to unlabeled workloads. However, the threshold is still model-specific and must be re-tuned for every new system. A separate line of work, including AGD [27] and PseudoAutoEval [2], relies on retraining the target model or generating pseudo-labels, which introduces substantial computational overhead, additional inference passes, and human annotation. However, these approaches are largely developed for image. In Text2SQL, where schemas and query distributions evolve rapidly with scarce labels [19, 32], label-free evaluation of unseen models remains largely unexplored; the closest efforts estimate Text2SQL accuracy on unlabeled data [61, 64]. NL2SQL-BUGS [38] targets fine-grained debugging with both automated detectors and humanin-the-loop judgment, thus having low-throughput analysis at the query level and requiring substantial human involvement. Overall, as summarized in Fig. 1, prior methods are model-specific and costly, preventing scalable evaluation under the double challenge of unseen models and unlabeled data. We close this gap with a meta-learning framework that learns evaluators across reference models and shifts, enabling rapid, model-agnostic performance estimation for both Image Classification and Text2SQL.

3

Problem Formulation

Motivated by the limitations of prior work, we formalize a setting in which both the evaluated model and the target dataset are unseen and unlabeled at deployment time. To address this double challenge, we start from classical supervised learning and extend it to label-free evaluation and meta-learning. In the single-model case,

Pham et al.

supervised learning is used to train an evaluator which is a regressor over dataset-level signals that summarize: train–test mismatch and the model’s true accuracy. We then lift this formulation to meta-learning by aggregating such supervised evaluation problems across many reference models and shifts so that the evaluator can adapt to new architectures on unlabeled workloads.

3.1

Model

We begin with standard supervised learning, where one is given a labeled dataset D = {(𝑥𝑖 , 𝑦𝑖 )}𝑛𝑖=1 and a predictive model 𝑓𝜃 with paÍ rameters 𝜃 , and training minimizes the empirical risk 𝑛𝑖=1 L (𝑓𝜃 (𝑥𝑖 ), 𝑦𝑖 ), where L denotes a loss such as mean-squared error (MSE). At evalutest 𝑛 ′ ation time, the model is tested on a labeled set Dtest = {(𝑥 test 𝑗 , 𝑦 𝑗 )} 𝑗=1 , producing predictions 𝑦ˆ 𝑗 = 𝑓𝜃 (𝑥 test 𝑗 ), and robustness to distribution shift is quantified through an aggregate error such as mean absolute Í ′ test error MAE = 𝑛1′ 𝑛𝑗=1 |𝑓𝜃 (𝑥 test 𝑗 ) − 𝑦 𝑗 |. Label-free Model Evaluation. We now turn from learning predictive models to learning evaluators. For a trained model 𝑓 (· | 𝜓 ) with training set D𝑆 and a labeled evaluation workload D𝑇 , prior works [7, 16] construct a shift descriptor SD summarizing distributional or confidence-level shifts between D𝑆 and D𝑇 , and pair it with the model’s true accuracy 𝑎. This yields a training set: Dtrain = {(SD 𝑗 , 𝑎 𝑗 )}𝑁𝑗=1,

(1)

from which an evaluator 𝑔𝜃 is trained  by minimizing an empirical Í regression loss 𝑁𝑗=1 L (𝑔𝜃 (SD 𝑗 ), 𝑎 𝑗 . At deployment, for a new unlabeled target workload D𝑇new , one computes SDnew between D𝑆 and D𝑇new , predicts 𝑎ˆ = 𝑔𝜃 (SDnew ), and later measures generalization by comparing 𝑎ˆ with the unknown true accuracy 𝑎 new once labels become available offline. Scaling to Unseen Models. When evaluation must extend to a set of unseen models {𝑚𝑘 }, a naive strategy repeatedly acquires la(𝑚 ) beled target workloads D𝑇 𝑘 and forms the evaluator training set (𝑚𝑘 ) (𝑚𝑘 ) with new pairs (SD ,𝑎 ), followed by supervised fine-tuning of 𝑔𝜃 . This procedure amounts to optimizing a single parameter Í Í (𝑚 ) (𝑚 )  vector 𝜃 ★ = arg min𝜃 𝑘 𝑗 L (𝑔𝜃 (SD 𝑗 𝑘 ), 𝑎 𝑗 𝑘 , which aggregates losses across all previously observed models. Consequently, predictions for a novel architecture 𝑚 new are forced through this global solution 𝜃 ★, effectively using its shift descriptors with an evaluator that lacks task-specific adaptation. Meta-Learning for Model Evaluation. To overcome this limitation, we cast evaluation itself as a meta-learning problem. We construct a reference pool of models M and treat each 𝑚 ∈ M as a separate task. For every 𝑚, we form a task-specific dataset: 𝑁𝑚 D (𝑚) = {(SD𝑖(𝑚) , 𝑎𝑖(𝑚) )}𝑖=1 ,

(2)

where SD𝑖(𝑚) captures the shift between the training set of 𝑚 and a labeled workload which is drawn from MetaDataset (later introduced in §4), and 𝑎𝑖(𝑚) is the corresponding accuracy. The collection of all such tasks defines the meta-training set S = {D (𝑚) : 𝑚 ∈ M}. We meta-learn parameters 𝜃 for an evaluator 𝑔𝜃 so that, after a small number of adaptation steps on a new labeled workload D𝑛𝑒𝑤 , the adapted evaluator, called 𝑔𝜃𝑚 , accurately predicts accuracies for that model 𝑚. In this way, 𝜃 is optimized not to evaluate any single architecture, but to encode a transferable strategy for evaluating unseen ones.

Cost-Effective Model Evaluation with Meta-Learning

Preprint, 2026,

Figure 2: MetaEvaluator applies meta-learning over a pool of reference models, using data from MetaDataset to learn how to map shift descriptors to estimates by adapting to each reference model and minimizing error against known performance. At test time, we are given an unseen model, formulated by 𝑓new (· | 𝜓 new ), and an unlabeled target workload: D𝑇 = {𝑥𝑖𝑇 }𝑛𝑖=1,

(3)

on which the model produces predictions 𝑦ˆ𝑖 = 𝑓new (𝑥𝑖𝑇 | 𝜓 new ). Let 𝑦𝑇𝑖 ★ denote the unknown ground-truth outputs and define the true dataset-level metric: 𝑛 1 ∑︁ 𝑀★ = 𝑚(𝑦ˆ𝑖 , 𝑦𝑇𝑖 ★), (4) 𝑛 𝑖=1 where 𝑚(·, ·) specifies performance measures, e.g., exact match. Our goal is to estimate 𝑀 ★ without observing 𝑦𝑇𝑖 ★. We compute shift descriptors between the training data of 𝑓new and D𝑇 , adapt the meta-learned evaluator 𝑔𝜃 using a small set of reference signals ˆ Performance is measured if available, and output a prediction 𝑀. by the mean absolute error | 𝑀ˆ − 𝑀 ★ | across many unseen models and deployment shifts, reflecting whether the system has learned to evaluate rather than to memorize any specific architecture.

3.2

workload D𝑇 , without access to the ground-truth labels {𝑦𝑇𝑖 ★ } or modifying the model itself, thereby directly addressing 𝜉1 and 𝜉4. The estimator must further satisfy efficiency constraints in deployment scenarios (𝜉5). Let SDsrc denote descriptors computed from 𝑓new on its training set D𝑆 , and let SDtgt denote descriptors computed on the unlabeled target set D𝑇 . We summarize train–test differences through:  SD = ℎ SDtgt, SDsrc , (5) and seek an evaluator 𝑔, parameterized by 𝜃 , that maps these descriptors to a performance estimate: b = 𝑔𝜃 (SD). 𝑀

The evaluator is designed to achieve the following objectives: b ★ | across target workloads (1) Accuracy (𝜉1–𝜉3): minimize |𝑀−𝑀 and unseen models. (2) Uncertainty (𝜉2): provide calibrated uncertainty estimates, b −𝛿𝛼 , 𝑀 b +𝛿𝛼 ] such that P(𝑀 ★ ∈ e.g., a prediction interval [𝑀 b −𝛿𝛼 , 𝑀 b +𝛿𝛼 ]) ≥ 1−𝛼, where 𝛿𝛼 is the interval half-width [𝑀 at miscoverage level 𝛼. (3) Generality (𝜉1, 𝜉2, 𝜉4): require no access to ground-truth labels and no changes to the unseen model or its parameters. (4) Efficiency (𝜉5): operate with low runtime and resource consumption to enable scalable benchmarking.

Challenges

Building on the formulation above, we study the double challenge of estimating the true performance 𝑀 ★ of an unseen model 𝑓new (· | 𝜓 new ) on an unlabeled workload D𝑇 . This deployment scenario introduces several fundamental difficulties: 𝜉1 Absence of ground truth: the labels {𝑦𝑇𝑖 ★ } for D𝑇 are unavailable, so 𝑀 ★ in Eq. 4 cannot be computed directly. 𝜉2 Cross-model generalization: the evaluator 𝑔𝜃 must remain reliable when confronted with a novel model 𝑓new (· | 𝜓 new ) whose behavior lies outside the reference pool M used during training. 𝜉3 Distribution shift: the target distribution D𝑇 may differ substantially from the source data D𝑆 in domain or input statistics, inducing unpredictable performance changes. 𝜉4 Limited access: the evaluator must operate without modifying the unseen model or its parameters 𝜓 new . 𝜉5 Efficiency constraints: evaluation must remain lightweight in computation to enable scalable benchmarking and practical pre-deployment model selection.

3.3

Objective

Our objective is to estimate the true dataset-level performance 𝑀 ★ of an unseen model 𝑓new (· | 𝜓 new ) on the unlabeled target

(6)

4

Methodology

Motivated by Objectives 1–3 in §3, we construct MetaDataset, a unified multimodal corpus for meta-learning label-free evaluation. This exposes MetaEvaluator to varied shifts and model behaviors, enabling generalization to unseen architectures and supporting calibration through repeated observation across conditions. Our framework proceeds in two stages: (1) MetaDataset construction and (2) MetaEvaluator learning.

4.1

MetaDataset

Unlike conventional datasets that benchmark a fixed model on a single test distribution, MetaDataset is organized around model– shift pairs, providing multiple target workloads and true accuracies for each reference model and thus forming diverse meta-learning tasks. It must satisfy two general principles across modalities: (i) it should expose models to diverse and controllable distribution shifts [43, 47, 76] so evaluators can learn realistic performance

Preprint, 2026,

degradation, and (ii) it should be scalable and low-barrier, allowing practitioners to synthesize large volumes of shifted data without expert annotation. Practitioners are free to generate task-specific environments (e.g., graph [49, 55], voice [68], medical [67], and financial [24] settings), and control the shifts. In this work, we instantiate this dataset for two representative domains: Text2SQL and Image Classification. Text2SQL. We design relational environments that expose evaluators to schema evolution, SQL structural variation, and linguistic shift, reflecting deployment conditions in prior benchmarks [60, 63]. Database diversity. Starting from real-world tables (e.g., TabLib [10] and KaggleDBQA [29]), we apply acquisition, refinement, and synthesis steps, using GPT-5 as the backbone model, to remove noise, cluster compatible schemas, infer foreign keys, standardize columns, and construct multi-table databases with realistic connectivity. SQL diversity. We generate queries ranging from simple projections to nested analytics. To overcome template bias, we augment large-scale generators with SQLForge [17], PARSQL [5], and semantics-preserving rewriting [4], producing diverse join patterns, subqueries, and dialectal variants (SQLite, PostgreSQL, Snowflake). Question diversity. Each SQL query is paired with multiple naturallanguage realizations (e.g., formal, colloquial, conversational, vague, etc.) inspired by SynSQL-2.5M [33], SParC [85], and CoSQL [83]. We further inject realistic noise (distractors and modifiers) observed in KaggleDBQA and BIRD [34], while preserving the underlying SQL intent. Image Classification. We construct a multi-stage pipeline for label-preserving image generation under realistic distribution shift, combining vision–language guidance with diffusion models. Dataset preparation. For each class in several vision benchmarks (§5.1), we collect diverse seed images and apply CLIP-based filtering [65] to remove ambiguous or low-alignment samples. Semantic edit. Following EvolveDirector [90] and recent advances in instruction-guided image editing [56, 62], a vision–language controller proposes edits that are executed by a text-conditioned diffusion model [69]. We instantiate five families of shifts: illumination, material/surface properties, camera perturbations, background relocation, and contextual changes (e.g., weather, occlusion, motion blur). Shift severity is controlled by activating one family (mild), two to three (moderate), or at least three including material plus background or context (strong). Validation and filtering. A second CLIP-based alignment check removes label drift, ensuring semantic consistency under shift. Summary. MetaDataset is a unified multimodal dataset for MetaEvaluator to learn distribution-aware, label-free performance estimation, thus generalizing across models and deployment conditions.

Pham et al.

Algorithm 1 Meta-learning MetaEvaluator 𝑔𝜃 with Mtrain . 1: Input: Strain , Sval , pool models M train , initial parameters 𝜃 , ini-

tial context vectors {𝑐𝑡𝑥𝑚 }𝑚∈ Mtrain , learning rates 𝛼 inner, 𝛼 outer , epochs 𝐸. 2: for 𝑒 = 1 to 𝐸 do 3: for 𝑚 ∈ Mtrain do 4: // Inner loop on Strain : 5: Ltrain (𝑚) ← 0. 6: for each (𝑆𝐷𝑖 , 𝑀𝑖★) in Strain do 2 7: Ltrain (𝑚) ← Ltrain (𝑚) + 𝑔𝜃 (𝑆𝐷𝑖 , 𝑐𝑡𝑥𝑚 ) − 𝑀𝑖★ . 8: end for 9: // RMSE loss: √︁ 10: Ltrain (𝑚) ← Ltrain (𝑚)/|Strain |. 11: // Update 𝑐𝑡𝑥𝑚 (𝜃 fixed): 12: 𝑐𝑡𝑥𝑚 ← 𝑐𝑡𝑥𝑚 − 𝛼 inner ∇𝑐𝑡𝑥𝑚 Ltrain (𝑚). 13: 14: 15: 16:

// Outer loop on Sval : Lval ← 0. for each (𝑆𝐷 𝑗 , 𝑀 ★ 𝑗 ) in Sval do

2 Lval ← Lval + 𝑔𝜃 (𝑆𝐷 𝑗 , 𝑐𝑡𝑥𝑚 ) − 𝑀 ★ . 𝑗 18: end for 19: end for √︁ 20: Lval ← Lval /|Sval |. 21: // Update 𝜃 (all 𝑐𝑡𝑥𝑚 fixed): 22: 𝜃 ← 𝜃 − 𝛼 outer ∇𝜃 Lval . 23: end for ★} 24: Output: 𝜃 ★ = 𝜃 , {𝑐𝑡𝑥𝑚 𝑚∈ Mtrain . 17:

Algorithm 2 Evaluation for 𝑚𝑛𝑒𝑤 with MetaEvaluator 𝑔𝜃 ★ . 1: Input: an unseen model 𝑚𝑛𝑒𝑤 , initial context 𝑐𝑡𝑥 new , meta-set

Strain , unseen and unlabeled data DT , optimal parameters 𝜃 ★, learning rate 𝛼, adaptation steps 𝐾. 2: for 𝑘 =1 to 𝐾 do 3: L ← 0. 4: // Inner loop on Strain : 5: for each (𝑆𝐷𝑖 , 𝑀𝑖★) in Strain do 2 6: L ← L + 𝑔𝜃 ★ (𝑆𝐷𝑖 , 𝑐𝑡𝑥 new ) − 𝑀𝑖★ . 7: end for √︁ 8: L ← L/|Strain |. 9: // Update 𝑐𝑡𝑥 new (𝜃 ★ fixed): 10: 𝑐𝑡𝑥 new ← 𝑐𝑡𝑥 new − 𝛼∇𝑐𝑡𝑥 new L. 11: end for 12: 13: // Compute 𝑆𝐷 T and Estimate:

4.2

MetaEvaluator

To satisfy the objectives in §3, we meta-learn an evaluator over the broad coverage provided by MetaDataset (§4.1) across diverse model architectures and distribution shifts. Meta-Learning. As shown in Alg. 1, we partition MetaDataset into two splits, Dtrain and Dval , and optimize MetaEvaluator over a reference model pool Mtrain . In each episode, for a model 𝑚 ∈ Mtrain and a collection of sample sets drawn from either split, we compute shift descriptors SD that summarize differences between

b = 𝑔𝜃 ★ (𝑆𝐷 T, 𝑐𝑡𝑥 new ). 14: 𝑀 b 15: Output: 𝑀.

the outputs of 𝑚 on its original training data and on each sample set. Together with the corresponding true dataset-level performance 𝑀 ★, we collect many such (SD, 𝑀 ★) pairs, forming a sample set. In contrast to MAML [12], we treat each model as a task rather than each dataset. Sample sets drawn from Dtrain are grouped into

Cost-Effective Model Evaluation with Meta-Learning

Preprint, 2026,

Text2SQL Dataset

Source

WikiSQL [93] Spider [84] SParC [85] CoSQL [83] BIRD [34] ScienceBenchmark [87] EHRSQL [30] KaggleDBQA [29] SynSQL-2.5M [33] MetaDataset (ours)

Human+Template Human Human Human Human Hybrid Human+Template Human LLM-Gen Synthetic

# Examples 80,654 10,181 12,726 10,000+ 12,751 5,031 20,108 272 2,544,390 3,373,204

Handwritten Real-world (scanned) Real-world (street-view) Human-annotated Human-annotated Curated Web Synthetic

70,000 9,298 99,289 123,287 11,540 1,331,167 2,487,936

(a) Text2SQL

Image Classification MNIST [28] USPS [20] SVHN [39] COCO 2017 [36] PASCAL VOC 2012 [11] ImageNet ILSVRC12 [6] MetaDataset (ours)

(b) Image Classification

Table 1: Dataset sizes and source categories used in the coverage analysis for Text2SQL and Image Classification. a meta-set of size 𝑠 train : 𝑠

train Strain = {(SD𝑖 , 𝑀𝑖★)}𝑖=1 ,

(7)

while sample sets drawn from Dval form a meta-set of size 𝑠 val : 𝑠

val Sval = {(SD 𝑗 , 𝑀 ★ 𝑗 )} 𝑗=1 ,

(8)

Learning proceeds in two stages: (i) an adaptation step that updates only the model-specific context vector 𝑐𝑡𝑥𝑚 , which has dimension 512, using Strain while keeping the global parameters 𝜃 fixed; and (ii) a generalization step that updates 𝜃 using Sval while holding all {𝑐𝑡𝑥𝑚 }𝑚∈ Mtrain fixed. The overall procedure is depicted in Fig. 2. Evaluation. During evaluation for an unseen model 𝑚 new on D𝑇 , we adapt only the newly initialized context vector 𝑐𝑡𝑥 new of 𝑚 new , while keeping the globally meta-trained parameters 𝜃 ★ fixed. We compute shift descriptors SD as distributional differences between the outputs of 𝑚 new on its original training data and on each sample 𝑠 train set in the meta-set Strain , yielding {(SD𝑖 , 𝑀𝑖★)}𝑖=1 . We then perform a small number of lightweight adaptation steps on 𝑐𝑡𝑥 new by minimizing the loss over Strain while holding 𝜃 ★ fixed, as detailed in Alg. 2. Finally, given an unseen and unlabeled target workload D𝑇 , we compute SD𝑇 as the distributional difference between the outputs of 𝑚 new on its original training data and D𝑇 , and estimate:   b = 𝑔𝜃 ★ SD𝑇 , 𝑐𝑡𝑥 new . 𝑀 (9)

5

Experiment

Building a label-free evaluation method that remains reliable for unseen models under distribution shift raises four fundamental research questions: • RQ1 (Data Coverage) (§5.1): How well does MetaDataset capture diverse, realistic deployment scenarios and distribution shifts across tasks and domains? • RQ2 (Evaluator Learning) (§5.2): How accurately does MetaEvaluator learn and generalize to unseen models and distribution shifts across modalities?

Figure 3: t-SNE of semantic coverage across modalities. • RQ3 (Benchmarking Capability) (§5.3): Does MetaEvaluator enable lightweight benchmarking as the number of unseen models and the reference model pool grow? • RQ4 (Component Utility) (§5.4): How do the individual components of MetaEvaluator contribute to its overall estimation accuracy and robustness across unseen models and distribution-shifted workloads? Environment. We run all experiments on a workstation with four NVIDIA GeForce RTX 4090 GPUs (24GB each) and an Intel Core i7-14700 CPU (20 cores, 2.1,GHz base). To maximize throughput during evaluator meta-learning, we use mixed-precision computation (bfloat16) and parallel data loading. Base Models. We evaluate MetaEvaluator on two learning domains: Text2SQL and Image Classification, using modality-specific pools of base models. In each domain, every model is fine-tuned on its training data until achieving its best validation performance. The resulting models are then applied to generate predictions and latent representations, which are aggregated into fixed-length vectors that summarize the train–test shift, called shift descriptors. These shift descriptors serve as inputs to MetaEvaluator for estimating performance on unlabeled target workloads. Shift Descriptors. We construct a shift descriptor SD by concatenating three complementary hidden-space summaries: a Gaussian Fr’echet term SD𝐹 , a Mahalanobis term SD𝑀 , and a sliced Wasserstein term SD𝑆𝑊 , forming SD = [SD𝐹 , SD𝑀 , SD𝑆𝑊 ] as input to MetaEvaluator. SD𝐹 captures global changes in embedding statistics, SD𝑀 emphasizes rare or low-density examples such as uncommon SQL constructs or corrupted images, and SD𝑆𝑊 models directional geometric shifts caused by systematic changes in query structure or visual conditions. Together, these components provide a compact and expressive summary of train–test mismatch across modalities. Training Settings. MetaEvaluator is implemented as a three-layer MLP with hidden dimensions {256, 128, 64}, ReLU activations, and layer normalization, and is trained to regress the true accuracy of a model from shift descriptors. Training uses a batch size of 64 and the AdamW optimizer with learning rate 1×10−4 (cosine decay), 𝛽 1 =0.9, 𝛽 2 =0.999, and weight decay 1×10−3 . MetaEvaluator performs meta-learning up to 100 epochs with early stopping based

Preprint, 2026,

Pham et al.

on validation MAE, applies dropout of 0.2 between hidden layers, and optimizes mean squared error (MSE) between predicted and true accuracies. Evaluation Metrics. We evaluate MetaEvaluator by measuring how accurately it predicts dataset-level performance on each target data. For Image Classification, the task metric is classification accuracy (Acc). For Text2SQL, we report both exact match (EM) ★ 𝑛 and execution accuracy  (EX).  Given image data D = {(𝑥𝑖 , 𝑦𝑖 )}★𝑖=1 : 1 Í𝑛 ★ Acc(D) = 𝑛 𝑖=1 I 𝑦ˆ𝑖 = 𝑦𝑖 , where 𝑦ˆ𝑖 is the prediction, 𝑦𝑖 is the ground-truth label, and I[·] is the indicator function.  Given  Í Text2SQL data D = {(𝑥𝑖 , 𝑞𝑖★)}𝑛𝑖=1 : EM(D) = 𝑛1 𝑛𝑖=1 I 𝑞ˆ𝑖 = 𝑞𝑖★ , ★ where 𝑞ˆ𝑖 is the predicted SQL, and 𝑞𝑖 is the ground-truth SQL. Let Exec(·) on the database: EX(D) =  denote query execution  1 Í𝑛 ★ ) . Across 𝑛 samples in target data, ˆ I Exec( 𝑞 ) = Exec(𝑞 𝑖 𝑖=1 𝑖 𝑛 we quantify accuracy estimation quality using Mean Absolute Error Í𝑁 b 𝑀𝑖 − 𝑀𝑖★ . We report MAE separately for (MAE): MAE = 𝑁1 𝑖=1 Acc, EM, and EX, where smaller values indicate a more accurate and reliable evaluator. Budget Constraints. MetaDataset is generated under a fixed budget 𝐵=1,000 (USD). We decompose the total cost into three operations: (i) generation, (ii) filtering/validation, and (iii) execution (SQL only). For each modality 𝑡 ∈ {sql, img}, we partition the corpus into sample units U𝑡 (e.g., schema–workload units for Text2SQL gen and dataset–class units for images). For each unit 𝑢 ∈ U𝑡 , 𝑛𝑢 val denotes the number of generated candidates, 𝑛𝑢 denotes the number of candidates that are filtered/validated, and 𝑛𝑢exec denotes the number of SQL executions used for verification (with 𝑛𝑢exec =0 for gen img). With per-operation unit costs 𝑐𝑡 , 𝑐𝑡val , and 𝑐𝑡exec , the total cost is:

𝐶=

∑︁

 ∑︁  gen gen 𝑛𝑢 𝑐𝑡 + 𝑛𝑢val𝑐𝑡val + 𝑛𝑢exec𝑐𝑡exec ≤ 𝐵.

(10)

𝑡 ∈ {sql,img} 𝑢 ∈ U𝑡

For Text2SQL (𝑁 ≈3.4M), unit costs {𝑐 gen, 𝑐 val, 𝑐 exec }={8, 2, 40}×10−5 yield a projected total 𝐶 sql ≈489.6. For Images (𝑁 ≈2.5M), costs {𝑐 gen, 𝑐 val }={15, 3}×10−5 yield 𝐶 img ≈456.8. Thus, the total estimated gen cost is 946.4≤𝐵. We enforce strict per-sample caps 𝑛¯𝑡 and 𝑛¯𝑡exec gen gen exec (e.g., 𝑛¯sql =160, 𝑛¯sql =40, 𝑛¯img =300) to guarantee the worst-case. Training Data Formation. From MetaDataset, we construct training workloads for meta-learning by forming a meta-set: S = S𝑡𝑟𝑎𝑖𝑛 + S𝑣𝑎𝑙 (Eq. 7 and Eq. 8). Meta-set contains multiple sample sets 𝑠𝑖 across both modalities. Based on this, we form instances (Dtrain, 𝑠𝑖 ) for meta-learning (§4.2), where Dtrain is a fixed training set and 𝑠𝑖 is a sample set drawn from 𝑆. Each 𝑠𝑖 captures a distinct distribution shift scenario. For Text2SQL, 𝑠𝑖 varies database schemas, SQL operators, and linguistic forms. For Image Classification, it varies class subsets, background domains, and acquisition styles. We generate a meta-set with |S|=30K such sample sets 𝑠𝑖 , where |𝑠𝑖 |=10K. This balances coverage and computational cost while spanning both small-scale human-curated regimes and the large synthetic corpora, as summarized in Fig. 3 and Tab. 1.

Figure 4: Calibration of accuracy estimation across transfers.

Figure 5: Latency–MAE trade-offs on unseen models.

5.1

Data Coverage

To answer RQ1 (Data Coverage), we assess whether MetaDataset provides comprehensive and realistic coverage of deployment scenarios across Text2SQL and Image Classification, capturing diversity in schema structure, SQL composition, natural-language usage, as well as visual domains, styles, and semantic categories. Data Size. Tab. 1 summarizes the scale of the datasets used in our coverage analysis across Text2SQL and Image Classification. Human-curated Text2SQL benchmarks are typically limited to at most tens of thousands of examples, with several datasets containing fewer than 15K queries, whereas recent LLM-generated corpora such as SynSQL-2.5M [33] expand to millions of instances. Our MetaDataset further increases scale to over 3.3M Text2SQL queries, exceeding all existing benchmarks. A similar pattern appears in Image Classification: classic handwritten and real-world digit datasets remain under 100K examples, while large curated collections such as ImageNet ILSVRC12 [6] exceed one million images. MetaDataset reaches 2.49M images, placing it among the largest resources used in this study and enabling systematic analysis of large-scale deployment regimes. Semantic Coverage. We visualize embedding-space geometry with t-SNE in Fig. 3 to assess whether our construction spans realistic deployment scenarios across Text2SQL and Image Classification. In Text2SQL (Fig. 3a), WikiSQL [93] forms a distinct single-table cluster, Spider [84], SParC [85], and CoSQL [83] group together under multi-table and conversational settings, and BIRD [34] occupies a separate region reflecting higher schema complexity, while Spider 2.0 [31] and SynSQL-2.5M [33] spread across multiple clusters, indicating broad schema and SQL coverage. Similar patterns appear in question space, where MetaDataset forms the widest envelope and fills gaps between benchmarks. For image classification (Fig. 3b),

Cost-Effective Model Evaluation with Meta-Learning

Preprint, 2026,

Table 2: MAE (↓) of dataset-level accuracy estimation on unseen models across Text2SQL and Image Classification. Each cell reports mean ± 95% CI (percentage points). Best in bold, second best underlined. Tasks

Methods

Text2SQL

DoC [16] ATC [14] AGD [27] PseudoAutoEval [2] AutoEval [7] NL2SQL-BUGS [38] MetaEvaluator (Ours)

Image Classification

DoC [16] ATC [14] AGD [27] PseudoAutoEval [2] AutoEval [7] SelfTrainEns [3] MetaEvaluator (Ours)

Meta-Llama-3-70B

Qwen2.5-32B

XiYanSQL-14B

Ministral-3-14B

gemma-2-2b

Avg.

15.42 ± 2.31 17.21 ± 2.18 14.77 ± 2.26 13.84 ± 2.05 11.62 ± 1.94 9.31 ± 1.42 3.41 ± 0.78

15.97 ± 2.45 17.88 ± 2.34 15.18 ± 2.39 14.31 ± 2.22 12.05 ± 2.08 9.68 ± 1.55 3.76 ± 0.84

16.18 ± 2.28 18.06 ± 2.11 14.96 ± 2.21 14.58 ± 2.09 12.33 ± 1.91 9.84 ± 1.37 3.55 ± 0.71

15.66 ± 2.39 17.52 ± 2.27 15.04 ± 2.30 14.07 ± 2.16 11.89 ± 2.01 9.47 ± 1.49 3.69 ± 0.80

16.05 ± 2.51 17.95 ± 2.41 15.25 ± 2.44 14.42 ± 2.28 12.21 ± 2.13 9.76 ± 1.61 3.88 ± 0.89

15.86 ± 2.39 17.72 ± 2.26 15.04 ± 2.32 14.24 ± 2.16 12.02 ± 2.01 9.61 ± 1.49 3.66 ± 0.80

ResNeXt-50-32x4d

RegNetY-8GF

ConvNeXt-Tiny

ViT-Tiny

DeiT-Small

Avg.

16.03 ± 2.21 17.02 ± 2.34 15.11 ± 2.08 13.67 ± 2.01 11.44 ± 1.86 10.02 ± 1.31 3.58 ± 0.73

16.54 ± 2.37 17.46 ± 2.51 15.59 ± 2.22 14.02 ± 2.15 11.79 ± 1.97 9.28 ± 1.43 3.74 ± 0.81

16.27 ± 2.18 17.19 ± 2.28 15.32 ± 2.05 13.81 ± 1.98 11.58 ± 1.82 10.11 ± 1.26 3.61 ± 0.69

16.88 ± 2.46 17.83 ± 2.59 15.74 ± 2.31 14.19 ± 2.24 12.01 ± 2.05 10.46 ± 1.51 3.89 ± 0.88

17.02 ± 2.55 17.96 ± 2.67 15.97 ± 2.40 14.37 ± 2.33 12.18 ± 2.12 10.63 ± 1.60 3.97 ± 0.94

16.55 ± 2.35 17.49 ± 2.48 15.55 ± 2.21 14.01 ± 2.14 11.80 ± 1.96 11.30 ± 1.42 3.76 ± 0.81

Figure 6: Total training and evaluation latency as the number of unseen models increases.

Figure 8: MLP attains the lowest error and benefits most from larger meta-sets, while costs rise sharply beyond 30K with marginal gains. Table 3: Impact of meta-learning algorithm on evaluation accuracy and cost. Each MAE cell reports mean ± 95% CI; best is in bold and second best is underlined. MAML [12] FO-MAML [12] Reptile [59] Meta-SGD [35] ANIL [66] ProtoNet [74] MetaEvaluator (Ours)

Figure 7: Meta-learning improves with pool size. Inset: Hessian spectra remain stable as pool size increases. digit datasets (MNIST [28], USPS [20], SVHN [39]) separate from natural-image datasets (COCO [36], PASCAL [11], ImageNet [6]) in background space, while class space reflects shared object semantics, and in both views MetaDataset overlaps all groups, capturing within-family and cross-family shifts. Overall, these visualizations qualitatively confirm broad semantic coverage across established benchmarks in both modalities.

5.2

Evaluator Learning

Using MetaDataset from §5.1, we perform meta-learning and evaluate MetaEvaluator for label-free accuracy estimation, addressing RQ2 (Evaluator Learning). The reference model pool is specified

MAE

# Steps

# Extra params (M)

11.63 ± 1.21 8.88 ± 1.29 9.12 ± 1.34 8.21 ± 1.18 8.47 ± 1.25 12.45 ± 1.41 3.26 ± 0.96

12 10 9 9 8 0 3

0.00 0.00 0.00 0.38 0.00 0.00 0.12

in our codebase, and none of these models overlap with the unseen models evaluated in subsequent experiments. Evaluator Benchmark. Results in Tab. 2 average over multiple source–target transfers unseen during meta-learning, where each model is trained on a source dataset and evaluated on an unseen and unlabeled target set, with MAE measuring error against the true target accuracy. Text2SQL transfers include Spider–BIRD, WikiSQL– Spider, SParC–CoSQL, SynSQL-2.5M–Spider, and WikiSQL–Spider 2.0, while Image Classification uses MNIST–USPS, MNIST–SVHN, COCO– PASCAL, and COCO–ImageNet. MetaEvaluator attains the lowest MAE (≈ 3–4), substantially outperforming all baselines: DoC and ATC degrade sharply under shift, AGD, PseudoAutoEval, and AutoEval remain far from deployment-level accuracy, NL2SQL-BUGS is strongest for Text2SQL yet struggles on highly compositional

Preprint, 2026,

Pham et al.

accuracy improves, making it suitable for scalable benchmarking in rapidly evolving model ecosystems. Given a target MAE, practitioners can further choose the size of the reference pool to balance estimation quality against computational budget.

5.4

Figure 9: MetaEvaluator consistently reduces both MAE and latency compared to other meta-learning algorithms. targets (e.g., Spider 2.0, BIRD), and SelfTrainEns leads prior vision methods but still lags behind while requiring costly auxiliary training. These results highlight the benefit of directly meta-learning evaluation across models and shifts rather than relying on conventional heuristics. Detailed results are provided in our code. Estimation Calibration. Fig. 4 compares predictions with ground truth across transfers. MetaEvaluator tracks GT most closely on both tasks, with small deviations on easier shifts and controlled bias on harder ones, whereas ATC and DoC systematically overestimate and other baselines fluctuate widely. This stable calibration explains the low MAE observed in Tab. 2.

5.3

Benchmarking Capability

To answer RQ3 (Benchmarking Capability), we evaluate whether MetaEvaluator supports fast and lightweight benchmarking when an organization must screen many newly released models against unlabeled workloads, and when the reference model pool continues to expand. Evaluation Latency. The best 3 methods from Tab. 2 are selected to compare the evaluation latency. Fig. 5 shows that Text2SQL (Fig. 5 above) incurs substantially higher cost than Image Classification due to LLM decoding and schema-conditioned reasoning. Trainingbased baselines (AGD, PseudoAutoEval, AutoEval) remain slow because each new model triggers retraining. In contrast, MetaEvaluator stays both the fastest (about 1–2 minutes per model) and the most accurate. Moreover, Fig. 6 confirms its much flatter growth as the number of unseen models increases. By amortizing training latency across the reference pool and requiring only a few lightweight adaptation forward passes per model, MetaEvaluator enables practitioners to benchmark large streams of candidate Image Classification or Text2SQL systems quickly on unlabeled data, placing it on a strictly better accuracy–latency Pareto frontier. Model Pool. Fig. 7 shows that MAE consistently decreases as the reference model pool expands, indicating that a larger coverage of the model improves generalization by exposing MetaEvaluator to additional failure modes and shift regimes. The inset Hessian spectra largely overlap across pool sizes, suggesting stable curvature and no growing optimization difficulty as the pool increases. Together, these results demonstrate that MetaEvaluator remains stable while

Ablation Study

To answer RQ4 (Component Utility), we examine how design choices in MetaEvaluator support fast and lightweight benchmarking, focusing on the meta-learning algorithm and the meta-set size used during training. Meta-set Size. Each meta instance (Dtrain, 𝑠𝑖 ) encodes a train–test shift through 𝑆𝐷 train = ℎ(𝜙 Dtrain , 𝜙𝑠𝑖 ) (Eq. 5). As shown in Fig. 8, we compare MetaEvaluator’s MLP with classical regressors while varying 𝑛. Estimation error decreases for all methods as 𝑛 grows, but the MLP continues improving up to 𝑛=30K while simpler models saturate earlier, and training cost rises sharply beyond this point with only marginal accuracy gains. This enables practitioners to select 𝑛 that meets target MAE levels while keeping training cost compatible with rapid and automated benchmarking pipelines. Meta-learning Algorithm. We study how different meta-learning algorithms shape the accuracy–efficiency trade-off. As shown in Tab. 3, Meta-SGD and ANIL approach the accuracy of full MAML with fewer adaptation steps, while Reptile and ProtoNet fall behind, suggesting that explicit adaptation mechanisms are better suited for scalable performance estimation under distribution shift. These observations motivate the design of MetaEvaluator, which attains the lowest MAE with only a few adaptation steps. Fig. 9 further confirms that MetaEvaluator consistently outperforms the average of alternative meta-learning baselines in both MAE and latency across Text2SQL and Image Classification. Together, these results demonstrate that MetaEvaluator’s meta-learning strategy directly supports fast and reliable benchmarking of many unseen models on unlabeled workloads.

6

Conclusion and Future Work

We introduce MetaEvaluator, a label-free framework for estimating the performance of unseen models on unseen workloads across Text2SQL and Image Classification. By combining meta-learning with compact shift descriptors, MetaEvaluator amortizes evaluation cost across reference models and achieves significantly lower error and latency than conventional approaches. Extensive experiments demonstrate that MetaEvaluator achieves a strong accuracy– efficiency trade-off and scales as reference and new models grow. These results position MetaEvaluator as a practical tool for fast, automated model benchmarking, without requiring additional annotation or per-model retraining. Future directions include extending MetaEvaluator to new modalities and studying joint evaluation of multiple candidate models to further reduce deployment overhead in large model pools.

References [1] Anastasios N Angelopoulos, Stephen Bates, Clara Fannjiang, Michael I Jordan, and Tijana Zrnic. 2023. Prediction-powered inference. Science (2023). [2] Pierre Boyeau, Anastasios Nikolas Angelopoulos, Tianle Li, Nir Yosef, Jitendra Malik, and Michael I. Jordan. 2025. AutoEval Done Right: Using Synthetic Data for Model Evaluation. In ICML.

Cost-Effective Model Evaluation with Meta-Learning

[3] Jiefeng Chen, Frederick Liu, Besim Avci, Xi Wu, Yingyu Liang, and Somesh Jha. 2021. Detecting errors and estimating accuracy on unlabeled data with self-training ensembles. NeurIPS (2021). [4] Shaoguo Cui, Keying Wen, Binbin Sang, Tiansong Li, Yi Zhang, and Huan Gao. 2025. LLM-Based Data Synthesis and Distillation for High-Quality Text-to-SQL Training. In ICIC. [5] Yaxun Dai, Haiqin Yang, Mou Hao, and Pingfu Chao. 2025. PARSQL: Enhancing Text-to-SQL through SQL Parsing and Reasoning. In ACL. [6] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR. [7] Weijian Deng and Liang Zheng. 2021. Are labels always necessary for classifier accuracy evaluation?. In CVPR. [8] Chi Thang Duong, Thanh Tam Nguyen, Trung-Dung Hoang, Hongzhi Yin, Matthias Weidlich, and Quoc Viet Hung Nguyen. 2022. Deep MinCut: Learning Node Embeddings from Detecting Communities. Pattern Recognition (2022), 109126. [9] Chi Thang Duong, Thanh Tam Nguyen, Hongzhi Yin, Matthias Weidlich, Thai Son Mai, Karl Aberer, and Quoc Viet Hung Nguyen. 2022. Efficient and Effective MultiModal Queries Through Heterogeneous Network Embedding. IEEE Transactions on Knowledge and Data Engineering 34, 11 (2022), 5307–5320. [10] Gus Eggert, Kevin Huo, Max Biven, Jeff Waugh, et al. 2023. TabLib: A Dataset of 627M Tables with Context. arXiv:2310.07875 (2023). [11] Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. 2010. The Pascal Visual Object Classes (VOC) Challenge. IJCV (2010). [12] Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017. Model-Agnostic MetaLearning for Fast Adaptation of Deep Networks. In ICML. [13] Adam Fisch, Joshua Maynez, R. Alex Hofer, Bhuwan Dhingra, Amir Globerson, and William W. Cohen. 2024. Stratified Prediction-Powered Inference for Effective Hybrid Evaluation of Language Models. In NeurIPS. [14] Saurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur, and Hanie Sedghi. 2022. Leveraging unlabeled data to predict out-ofdistribution performance. (2022). [15] Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. 2024. A survey on llm-as-a-judge. The Innovation (2024). [16] Devin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell, and Ludwig Schmidt. 2021. Predicting with confidence on unseen distributions. In ICCV. [17] Yu Guo, Dong Jin, Shenghao Ye, Shuangwu Chen, Jian Yang, and Xiaobin Tan. 2025. SQLForge: Synthesizing Reliable and Diverse Data to Enhance Text-to-SQL Reasoning in LLMs. In ACL. [18] Rundong He, Yicong Dong, Lan-Zhe Guo, Yilong Yin, and Tailin Wu. 2025. ReEvaluating the Impact of Unseen-Class Unlabeled Data on Semi-Supervised Learning Model. In ICLR. [19] Thanh Dat Hoang, Thanh Trung Huynh, Matthias Weidlich, Thanh Tam Nguyen, Tong Chen, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2026. Boosting Small Language Models for Text-to-SQL with Fine-Grained Execution Feedback and Cost-Efficient Rewards. In ICDE. IEEE. [20] Jonathan J. Hull. 2002. A database for handwritten text recognition research. TPAMI (2002). [21] Nguyen Quoc Viet Hung, Duong Chi Thang, Nguyen Thanh Tam, Matthias Weidlich, Karl Aberer, Hongzhi Yin, and Xiaofang Zhou. 2017. Answer validation for generic crowdsourcing tasks with minimal efforts. The VLDB Journal 26 (2017), 855–880. [22] Nguyen Quoc Viet Hung, Matthias Weidlich, Nguyen Thanh Tam, Zoltán Miklós, Karl Aberer, Avigdor Gal, and Bela Stantic. 2019. Handling probabilistic integrity constraints in pay-as-you-go reconciliation of data models. Information Systems 83 (2019), 166–180. [23] Thanh Trung Huynh, Chi Thang Duong, Thanh Tam Nguyen, Vinh Tong Van, Abdul Sattar, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2021. Network alignment with holistic embeddings. TKDE 35, 2 (2021), 1881–1894. [24] Thanh Trung Huynh, Minh Hieu Nguyen, Thanh Tam Nguyen, Phi Le Nguyen, Matthias Weidlich, Quoc Viet Hung Nguyen, and Karl Aberer. 2023. Efficient integration of multi-order dynamics and internal dynamics in stock movement prediction. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining. 850–858. [25] Thanh Trung Huynh, Trong Bang Nguyen, Phi Le Nguyen, Thanh Tam Nguyen, Matthias Weidlich, Quoc Viet Hung Nguyen, and Karl Aberer. 2024. Fast-fedul: A training-free federated unlearning with provable skew resilience. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 55–72. [26] Thanh Trung Huynh, Trong Bang Nguyen, Thanh Toan Nguyen, Phi Le Nguyen, Hongzhi Yin, Quoc Viet Hung Nguyen, and Thanh Tam Nguyen. 2025. Certified Unlearning for Federated Recommendation. ACM Transactions on Information Systems (2025). [27] Yiding Jiang, Vaishnavh Nagarajan, Christina Baek, and J Zico Kolter. 2022. Assessing Generalization of SGD via Disagreement. In ICLR.

Preprint, 2026,

[28] Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 2002. Gradientbased learning applied to document recognition. Proc. IEEE (2002). [29] Chia-Hsuan Lee, Hao Cheng, Jacob Devlin, Kristina Toutanova, and Jianfeng Gao. 2021. KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers. In ACL. [30] Gyubok Lee, Hyeonji Hwang, Seongsu Bae, Yeonsu Kwon, Woncheol Shin, Seongjun Yang, Minjoon Seo, Jong-Yeup Kim, and Edward Choi. 2022. EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records. In NeurIPS. [31] Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin SU, ZHAOQING SUO, Hongcheng Gao, Wenjing Hu, Pengcheng Yin, Victor Zhong, Caiming Xiong, Ruoxi Sun, Qian Liu, Sida Wang, and Tao Yu. 2025. Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows. In ICLR. [32] Boyan Li, Yuyu Luo, Chengliang Chai, Guoliang Li, and Nan Tang. 2024. The Dawn of Natural Language to SQL: Are We Fully Ready? VLDB (2024). [33] Haoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang, Jing Zhang, Fuxin Jiang, Shuai Wang, Tieying Zhang, Jianjun Chen, Rui Shi, Hong Chen, and Cuiping Li. 2025. OmniSQL: Synthesizing High-Quality Text-to-SQL Data at Scale. VLDB (2025). [34] Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, et al. 2023. Can LLM Already Serve as a Database Interface? A BIg Bench for Large-Scale Database Grounded Text-toSQLs. In NeurIPS. [35] Zhenguo Li, Fengwei Zhou, Fei Chen, and Hang Li. 2017. Meta-SGD: Learning to Learn Quickly for Few-Shot Learning. arXiv:1707.09835 (2017). [36] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV. [37] Renpu Liu and Jing Yang. 2025. Unlabeled Data Can Provably Enhance In-Context Learning of Transformers. In NeurIPS. [38] Xinyu Liu, Shuyu Shen, Boyan Li, Nan Tang, and Yuyu Luo. 2025. NL2SQL-BUGs: A Benchmark for Detecting Semantic Errors in NL2SQL Translation. In SIGKDD. [39] Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Baolin Wu, Andrew Y Ng, et al. 2011. Reading digits in natural images with unsupervised feature learning. In NeurIPS. [40] Dong Duc Anh Nguyen, Minh Hieu Nguyen, Phi Le Nguyen, Jun Jo, Hongzhi Yin, and Thanh Tam Nguyen. 2024. Multi-task Learning of Heterogeneous Hypergraph Representations in LBSNs. In International Conference on Advanced Data Mining and Applications. Springer, 161–177. [41] Minh Hieu Nguyen, Thanh Trung Huynh, Thanh Toan Nguyen, Phi Le Nguyen, Hien Thu Pham, Jun Jo, and Thanh Tam Nguyen. 2025. On-device diagnostic recommendation with heterogeneous federated BlockNets. Science China Information Sciences 68, 4 (2025), 140102. [42] Minh Hieu Nguyen, Thanh Tam Nguyen, Jun Jo, Duc Anh Nguyen, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2026. Handling Data Sparsity and Model Poisoning Attacks in Federated Sequential Recommender Systems. KnowledgeBased Systems (2026), 115545. [43] Quoc Viet Hung Nguyen, Son Thanh Do, Thanh Tam Nguyen, and Karl Aberer. 2015. Tag-based paper retrieval: minimizing user effort with diversity awareness. In International Conference on Database Systems for Advanced Applications. 510– 528. [44] Quoc Viet Hung Nguyen, Chi Thang Duong, Thanh Tam Nguyen, Matthias Weidlich, Karl Aberer, Hongzhi Yin, and Xiaofang Zhou. 2017. Argument discovery via crowdsourcing. The VLDB Journal 26, 4 (2017), 511–535. [45] Quoc Viet Hung Nguyen, Thanh Tam Nguyen, Vinh Tuan Chau, Tri Kurniawan Wijaya, Zoltán Miklós, Karl Aberer, Avigdor Gal, and Matthias Weidlich. 2015. SMART: A tool for analyzing and reconciling schema matching networks. In ICDE. 1488–1491. [46] Quoc Viet Hung Nguyen, Tam Nguyen Thanh, Zoltán Miklós, and Karl Aberer. 2014. Reconciling schema matching networks through crowdsourcing. EAI Endorsed Transactions on Collaborative Computing 1, 2 (2014), e2. [47] Quoc Viet Hung Nguyen, Kai Zheng, Matthias Weidlich, Bolong Zheng, Hongzhi Yin, Thanh Tam Nguyen, and Bela Stantic. 2018. What-if analysis with conflicting goals: Recommending data ranges for exploration. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). IEEE, 89–100. [48] Thanh Tam Nguyen, Thanh Trung Huynh, Zhao Ren, Thanh Toan Nguyen, Phi Le Nguyen, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2025. Privacy-preserving explainable AI: a survey. Science China Information Sciences 68, 1 (2025), 111101. [49] Thanh Tam Nguyen, Thanh Trung Huynh, Hongzhi Yin, Matthias Weidlich, Thanh Thi Nguyen, Thai Son Mai, and Quoc Viet Hung Nguyen. 2023. Detecting rumours with latency guarantees using massive streaming data. The VLDB Journal 32, 2 (2023), 369–387. [50] Thanh Toan Nguyen, Thanh Tam Nguyen, Thanh Hung Nguyen, Hongzhi Yin, Thanh Thi Nguyen, Jun Jo, and Quoc Viet Hung Nguyen. 2023. Isomorphic Graph Embedding for Progressive Maximal Frequent Subgraph Mining. ACM Transactions on Intelligent Systems and Technology 15, 1 (2023), 1–26. [51] Thanh Tam Nguyen, Thanh Toan Nguyen, Matthias Weidlich, Jun Jo, Quoc Viet Hung Nguyen, Hongzhi Yin, and Alan Wee-Chung Liew. 2024. Handling Low Homophily in Recommender Systems with Partitioned Graph Transformer.

Preprint, 2026,

IEEE Transactions on Knowledge and Data Engineering (2024). [52] Thanh Tam Nguyen, Thanh Cong Phan, Minh Hieu Nguyen, Matthias Weidlich, Hongzhi Yin, Jun Jo, and Quoc Viet Hung Nguyen. 2022. Model-agnostic and diverse explanations for streaming rumour graphs. Knowledge-Based Systems 253 (2022), 109438. [53] Thanh Tam Nguyen, Thanh Cong Phan, Hien Thu Pham, Thanh Thi Nguyen, Jun Jo, and Quoc Viet Hung Nguyen. 2023. Example-based explanations for streaming fraud detection on graphs. Information Sciences 621 (2023), 319–340. [54] Thanh Toan Nguyen, Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Thanh Trung Huynh, Thanh Thi Nguyen, Matthias Weidlich, and Hongzhi Yin. 2024. Manipulating recommender systems: A survey of poisoning attacks and countermeasures. Comput. Surveys 57, 1 (2024), 1–39. [55] Thanh Tam Nguyen, Zhao Ren, Thanh Toan Nguyen, Jun Jo, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2024. Portable graph-based rumour detection against multi-modal heterophily. Knowledge-Based Systems 284 (2024), 111310. [56] Thanh Tam Nguyen, Zhao Ren, Trinh Pham, Phi Le Nguyen, Quoc Viet Hung Nguyen, and Hongzhi Yin. 2026. A review of instruction-guided image editing. EAAI (2026). [57] Thanh Tam Nguyen, Matthias Weidlich, Hongzhi Yin, Bolong Zheng, Quang Huy Nguyen, and Quoc Viet Hung Nguyen. 2020. Factcatch: Incremental pay-as-yougo fact checking with minimal user effort. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 2165–2168. [58] Toan Nguyen Thanh, Nguyen Duc Khang Quach, Thanh Tam Nguyen, Thanh Trung Huynh, Viet Hung Vu, Phi Le Nguyen, Jun Jo, and Quoc Viet Hung Nguyen. 2023. Poisoning GNN-based recommender systems with generative surrogate-based attacks. ACM Transactions on Information Systems 41, 3 (2023), 1–24. [59] Alex Nichol, Joshua Achiam, and John Schulman. 2018. On First-Order MetaLearning Algorithms. arXiv:1803.02999 (2018). [60] Khanh Trinh Pham, Thu Huong Nguyen, Jun Jo, Quoc Viet Hung Nguyen, and Thanh Tam Nguyen. 2025. Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents. In Australasian Database Conference. Springer, 108–123. [61] Khanh Trinh Pham, Thanh Tam Nguyen, Viet Huynh, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2026. An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data. In 2026 IEEE 42nd International Conference on Data Engineering (ICDE). IEEE. [62] Minh Tam Pham, Thanh Trung Huynh, Thanh Tam Nguyen, Thanh Toan Nguyen, Thanh Thi Nguyen, Jun Jo, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2024. A dual benchmarking study of facial forgery and facial forensics. CAAI Transactions on Intelligence Technology 9, 6 (2024), 1377–1397. [63] Minh Tam Pham, Quoc Viet Hung Nguyen, Jun Jo, and Thanh Tam Nguyen. 2025. An Extensible Benchmark for Value Ambiguity Resolution in Text-to-SQL. In Australasian Database Conference. Springer, 124–138. [64] Trinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen, and Thanh Tam Nguyen. 2026. Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning. In KDD. [65] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML. [66] Anirudh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals. 2020. Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML. In ICLR. [67] Zhao Ren, Yi Chang, Thanh Tam Nguyen, Yang Tan, Kun Qian, and Björn W Schuller. 2024. A comprehensive survey on heart sound analysis in the deep learning era. IEEE Computational Intelligence Magazine 19, 3 (2024), 42–57. [68] Zhao Ren, Thanh Tam Nguyen, and Wolfgang Nejdl. 2022. Prototype learning for interpretable respiratory sound analysis. In Proc. ICASSP. 9087–9091. [69] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR. [70] Darnbi Sakong, Viet Hung Vu, Thanh Trung Huynh, Phi Le Nguyen, Hongzhi Yin, Quoc Viet Hung Nguyen, and Thanh Tam Nguyen. 2024. Higher-order knowledge-enhanced recommendation with heterogeneous hypergraph multiattention. Information Sciences 680 (2024), 121165. [71] David Salinas, Omar Swelam, and Frank Hutter. 2025. Tuning LLM Judge Design Decisions for 1/1000 of the Cost. In ICML. [72] Sebastian Schelter, Tammo Rukat, and Felix Biessmann. 2020. Learning to Validate the Predictions of Black Box Classifiers on Unseen Data. In SIGMOD. [73] Konstantin Schürholt, Diyar Taskiran, Boris Knyazev, Xavier Giró-i Nieto, and Damian Borth. 2022. Model zoos: A dataset of diverse populations of neural network models. NeurIPS (2022). [74] Jake Snell, Kevin Swersky, and Richard Zemel. 2017. Prototypical Networks for Few-shot Learning. In NeurIPS. [75] Duong Chi Thang, Hoang Thanh Dat, Nguyen Thanh Tam, Jun Jo, Nguyen Quoc Viet Hung, and Karl Aberer. 2022. Nature vs. nurture: Feature vs. structure

Pham et al.

for graph neural networks. PRL 159 (2022), 46–53. [76] Duong Chi Thang, Nguyen Thanh Tam, Nguyen Quoc Viet Hung, and Karl Aberer. 2015. An evaluation of diversification techniques. In International Conference on Database and Expert Systems Applications. 215–231. [77] Nguyen Thanh Toan, Phan Thanh Cong, Nguyen Thanh Tam, Nguyen Quoc Viet Hung, and Bela Stantic. 2018. Diversifying group recommendation. IEEE Access 6 (2018), 17776–17786. [78] Huynh Thanh Trung, Tong Van Vinh, Nguyen Thanh Tam, Jun Jo, Hongzhi Yin, and Nguyen Quoc Viet Hung. 2022. Learning holistic interactions in LBSNs with high-order, dynamic, and multi-role contexts. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 5002–5016. [79] Boris van Breugel, Nabeel Seedat, Fergus Imrie, and Mihaela van der Schaar. 2023. Can You Rely on Your Model Evaluation? Improving Model Evaluation with Synthetic Test Data. In NeurIPS. [80] Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021. Want To Reduce Labeling Cost? GPT-3 Can Help. In EMNLP Findings 2021. [81] Chaoqun Yang, Wei Yuan, Liang Qu, and Thanh Tam Nguyen. 2024. PDC-FRS: Privacy-Preserving Data Contribution for Federated Recommender System. In International Conference on Advanced Data Mining and Applications. Springer, 65–79. [82] Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li, Tongyu Liu, and Xiaoyong Du. 2018. Cost-effective data annotation using game-based crowdsourcing. (2018). [83] Tao Yu, Rui Zhang, Alexander Er, Suyi Li, Eric Xue, Bo Zhang, Shreya Pang, Xi Victoria Lin, Yuwen Li, et al. 2019. CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases. In EMNLP. [84] Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task. In EMNLP. [85] Tao Yu, Rui Zhang, Michihiro Yasunaga, Heyang Tan, Xi Victoria Lin, Suyi Li, Alexander Er, Irene Li, Shreya Pang, Tao Chen, et al. 2019. SParC: Cross-Domain Semantic Parsing in Context. In ACL. [86] Yaodong Yu, Zitong Yang, Alexander Wei, Yi Ma, and Jacob Steinhardt. 2022. Predicting out-of-distribution error with the projection norm. In ICML. [87] Yi Zhang, Jan Deriu, George Katsogiannis-Meimarakis, Catherine Kosten, Georgia Koutrika, and Kurt Stockinger. 2024. ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL Systems. VLDB (2024). [88] Yi-Kai Zhang, Ting-Ji Huang, Yao-Xiang Ding, De-Chuan Zhan, and Han-Jia Ye. 2023. Model spider: Learning to rank pre-trained models efficiently. NeurIPS (2023). [89] Bo Zhao, Han van der Aa, Thanh Tam Nguyen, Quoc Viet Hung Nguyen, and Matthias Weidlich. 2021. Eires: Efficient integration of remote data in event stream processing. In Proceedings of the 2021 International Conference on Management of Data. 2128–2141. [90] Rui Zhao, Yuxuan Li, Songhua Zhang, Zuxuan Wu, et al. 2024. EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models. In NeurIPS. [91] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. NeurIPS (2023). [92] Xin Zheng, Miao Zhang, Chunyang Chen, Soheila Molaei, Chuan Zhou, and Shirui Pan. 2023. GNNEvaluator: Evaluating GNN Performance On Unseen Graphs Without Labels. In NeurIPS. [93] Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. arXiv:1709.00103 (2017).

Record · ID 222608 · SHA-256 781aa2a8faeb7fbb
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.