ConceptioArchivearXiv CS
arXiv CSopen access

Building informative materials datasets beyond targeted objectives

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

arXiv:2605.05104v1 [cond-mat.mtrl-sci] 6 May 2026

Building informative materials datasets beyond targeted objectives Rafael Espinosa Castañeda1,4 , Ashley Dale1,4 , Hongchen Wang1 ,Yonatan Kurniawan1 , Hao Wan1 , Runze Zhang1 ,Adji Bousso Dieng3 , Kangming Li2 , Jason Hattrick-Simpers1,5,∗ 1 Department of Materials Science and Engineering, University of Toronto, Canada. 2 Department of Materials Science and Applied Physics,

King Abdullah University of Science and Technology, Saudi Arabia. 3 Department of Computer Science, Princeton University 4 Vector Institute for Artificial Intelligence, Toronto, Canada 5 Acceleration Consortium, Toronto, Canada ∗ Corresponding author:[email protected]

Abstract Materials science data collection can be expensive, making the reuse and long-term utility of datasets critical important for future discovery campaigns. In practice, researchers prioritize a subset of properties due to research interests. However, ignoring a subset of outcomes in data collection campaigns potentially generate datasets poorly suited for future learning tasks. Here, we present a framework for dataset construction that maximizes informativeness for target properties of interest while preserving performance on untargeted ones. Our approach uses diversity-aware selection to ensure broad coverage of the materials space. In noisy experimental dataset construction, we find that without our diversityaware framework, prediction performance on untargeted properties can degrade by up to ∼ 40% relative to random sampling, whereas applying our framework yields improvements of up to ∼ 10% . For targeted properties, performance can degrade with respect to random sampling by up to ∼ 12.5% without diversity, while our framework achieves gains of up to ∼ 25%. Incorporating diversity into dataset construction not only preserves informativeness for the targeted properties, but also improves materials coverage for potential future objectives. As a result, the constructed datasets remain broadly informative across considered and unconsidered outcomes, ensuring unbiased quality entries and mitigating cold-start limitations in subsequent modeling and discovery campaigns.

1

1

Introduction

narrow regions of materials space that are most informative for the targeted properties. As a consequence, datasets constructed in this manner may be poorly suited for future modeling tasks or broader scientific reuse. Recent evidence further suggests that simply increasing dataset size does not guarantee improved scientific utility. Li et al. [40] showed that widely used repositories such as Materials Project, JARVIS and OQMD contain substantial redundancy: between 70–95% of the data can be removed from training sets without exceeding a 10% degradation in out-of-distribution predictive performance for single-objective tasks. These findings indicate that many datasets oversample redundant regions of materials space while underrepresenting the broader diversity of materials and properties. As a result, data acquisition strategies that focus solely on expanding dataset size may fail to improve the informativeness or general utility of the resulting datasets. These observations raise a fundamental question: how should materials datasets be constructed so that they remain informative not only for immediate research objectives, but also for future scientific tasks that were not anticipated at the time of data collection? Here we present a framework for diversity-aware dataset construction that addresses this challenge. Our approach balances two competing objectives during data acquisition: maximizing informativeness for targeted properties while preserving broad coverage of the underlying materials feature space. We hypothesize that by explicitly incorporating diversity into the sampling process, the resulting datasets remain informative for the properties that motivated the data collection while also retaining predictive utility for outcomes that were not explicitly optimized. To evaluate this framework, we design a poolbased active learning workflow. Acquisition policies are assessed according to the performance of surrogate models trained on the selected samples. The accuracy of the models is measured on both targeted and untargeted outcomes. We compare diversity-aware policies against random sampling and uncertainty-driven policies that do not ex-

Data that can be reused across different scientific objectives is essential for accelerating discovery while reducing the time and cost of research. In materials science, high-quality experimental and computational data are often expensive to generate, requiring specialized instrumentation [1], complex synthesis [2], difficult measurements [3], or large-scale simulations [4–6]. Consequently, the long-term value of a dataset depends not only on the quality of the measurements, but also on its ability to support future scientific questions beyond those originally envisioned. Large community repositories have played a transformative role in enabling such reuse. Databases such as the Materials Project [7, 8], JARVIS [9], OQMD [10–12], AFLOW [13, 14], and ICSD [15–17] have supported a broad range of data-driven materials discovery workflows. These repositories have enabled materials characterization [18, 19], predictive modeling of materials properties [20–24], generative and inverse-design models based on crystal structures [25–28], and a wide spectrum of computational and experimental discovery pipelines [29–33]. By providing shared resources for the scientific community, these datasets allow existing measurements to be repurposed for new research questions and help mitigate the cold-start problem faced by emerging applications with limited task-specific data. Despite their success, comparatively little attention has been given to how such datasets should be constructed in the first place. In practice, data acquisition campaigns are typically guided by specific research objectives and therefore prioritize a limited set of target properties. For instance, many active learning and Bayesian optimization workflows in materials discovery focus on optimizing a single property at a time [30, 34–38]. Extensions to bi-objective optimization and involving three [39] outcomes have also been implemented. As more outcomes are included, these implementations become more uncommon in materials science. Although such strategies are effective for achieving immediate objectives, they implicitly concentrate sampling in 2

Figure 1: Overview of data selection pipeline for each iteration when explicitly incorporating diversity in feature space. The image shows how 1% of the pool data is selected. For each iteration, performance of the two models-Random Forest and XGBoost on hold-out test dataset is recorded. The pipeline is repeated until the complete pool data has been selected.

plicitly account for diversity. The analysis is performed under single, two, and three-objective acquisition settings in order to evaluate how increasing the number of optimized outcomes influences dataset quality.

paigns. By promoting broad coverage of the materials feature manifold, diversity-aware dataset construction mitigates distributional bias and reduces the risk of distribution shift in future data acquisition. Consequently, this framework provides a principled strategy for designing materials datasets that remain broadly informative across scientific objectives and across successive generations of data-driven discovery.

We test these strategies across five datasets comprising both computational and experimental materials data: four density functional theory (DFT) repositories and one experimental thermoelectric dataset composed. Across these settings, we show that diversity-aware sampling enables datasets to remain highly informative for targeted properties while preserving predictive performance for outcomes that were not considered during dataset construction. Moreover, diversityaware acquisition promotes efficient coverage of materials space, enabling the construction of compact datasets that retain strong predictive utility.

2

Methods

2.1

Datasets

We selected five datasets that each contain at least three physically distinct and weakly linearly correlated properties. These datasets enable us to evaluate our framework in different data distribution settings. Furthermore, guarantying different uncorrelated outcomes ensures that performance on unconsidered properties in dataset construction cannot arise from linear inter-property correlations and instead reflects cross-objective generalization. Four of the datasets are density functional theory (DFT) repositories: the 2018.06.01 (MP18)

More broadly, many data acquisition efforts build sequentially upon existing repositories. For example, slab structures in the Open Catalyst dataset were generated from bulk materials derived from the Materials Project [41, 42]. In such sequential discovery pipelines, the statistical structure of the initial dataset strongly influences the efficiency of subsequent data collection cam3

and 2021.11.10 (MP21) releases of the Materials Project, and the 2018.07.07 (JARVIS18) and 2022.12.12 (JARVIS22) releases of JARVIS. These DFT-based datasets provide three outcomes: electronic bandgap, elastic bulk modulus, and formation energy. Together, these properties probe different aspects of materials behavior — electronic structure, mechanical response, and thermodynamic stability. Furthermore, they are linearly weakly correlated (Pearson correlation values shown in the Appendix), hence we avoid that sampling for one outcome ensures immediate informativeness for other targets. Our analysis restricted to materials for which bandgap, bulk modulus, and formation energy were jointly available, discarding all other entries. For JARVIS22 and MP21, we further refined the datasets by removing materials with formation energies exceeding 5 eV atom−1 , as such values correspond to highly unstable or nonphysical structures and are often associated with calculation artifacts rather than meaningful chemical behavior. Across all DFT datasets, we additionally verified that bulk modulus values were physically meaningful and retained only entries with bulk modulus greater than 0 GPa. Structural and compositional descriptors (273 features) were then extracted using the Voronoi tessellation[43] featurization as implemented in matminer[44] .

tive for experimentally noise and heterogeneous measurements. Hence, our results demonstrates its robustness applicability for building informative datasets on experimental and computational campaigns. Mixed or composite formulas were first parsed into weighted sub-compositions and combined into a single effective composition, ensuring correct stoichiometric treatment of multicomponent systems. Each resulting composition was then featurized using a comprehensive suite of composition-based descriptors from matminer. Following featurization, we removed redundant feature columns containing exclusively missing values, eliminated duplicated rows, and retained only four target outcome properties of: the thermoelectric figure of merit (zT ), thermal conductivity, electrical conductivity, and the Seebeck coefficient. Additional properties, electrical thermal conductivity, lattice thermal conductivity and power factor that exhibit strong linear correlation with Thermal conductivity and zT were excluded. This curation step ensures that performance on non considered outcomes cannot be attributed to strong linear correlations among targets. The explicit correlation matrix of the included features can be found in the appendix. The final number of curated entries used for the five data sets is summarized in Table 1. By evaluating performance jointly across DFT repositories and experimental data, we probe data sampling under two fundamentally different data regimes, with controlled and uncontrolled noise. Furthermore, this complementary design supports that our findings are independent of specific dataset structure or data distribution.

For the databases JARVIS18 and MP18, we did not directly use the originally reported structures or property values. Instead, we used the material identifiers from these older releases to query their corresponding entries in the most recent database versions, ensuring that all structures and labels reflected up-to-date calculations. This procedure mitigates inconsistencies introduced by database revisions. The fifth dataset used in this work is the Systematically Verified Thermoelectric (sysTEm) dataset [45]. It contains experimentally measured thermoelectric materials. The dataset was curated with doped and undoped materials. Evaluating our framework on this dataset allows to assess whether acquisition strategies remain effec-

Dataset

Number of Materials

JARVIS18 JARVIS22 MP18 MP21 sysTEm

18,964 23,472 6,979 7,112 7,771

Table 1: Number of materials used in each data set after data curation.

4

2.2

Active Learning Policies and Models

tify local epistemic uncertainty through prediction disagreement between two surrogate model families. For each outcome i ∈ O, we train independent single-output predictors: one RF model and one XGB model per outcome. Thus, cross-outcome dependencies are not imposed at the model level, and disagreement is evaluated outcome-wise. For a candidate sample s in the pool P , the disagreement score is defined as X |ŷRF,i (s) − ŷXGB,i (s)| , (1) Dis(s) = ∗ ∗ |ŷ RF,i (si ) − ŷXGB,i (si )| i∈O

Prior to initiating a loop for data selection, we randomly allocated 10% of the curated data—using a fixed random seed—to serve as an independent test set T ′ . The remaining 90% of the data was used for the active learning data selection. This last mentioned dataset is partitioned into two disjoint subsets: the training set T and the candidate pool P dataset. As initial training data, we randomly assign 1% of the entries to T . To assess the sensitivity of the learning trajectory to the initial seed, this initialization is repeated for five different random seeds. At each acquisition iteration, before selecting new samples from the pool P , we train three XGBoost regressors and three Random Forest regressors on the current training set T , one model per target property. The specific hyper-parameters of the models are shown in the appendix. Using these trained models, we then generate predictions for all samples in the test set T ′ and record the performance metrics (R2 , RMSE, and MAE) for each property at every iteration. Random Forest and XGBoost regressors were selected due to their low computational cost compared to other complex models and strong predictive performance on tabular data [46]. These make them well-suited for iterative data acquisition , where models must be retrained repeatedly. For uncertainty quantification, we use prediction disagreement as the uncertainty measure. The corresponding predictions for XGBoost are provided in the Appendix. Three data-acquisition policies are compared in this study: one that does not account for diversity and other two that incorporates diversity into the selection process. One diversity policy incorporates it on the outcome space and the second incorporates it explicitly on feature space and output space. This comparison enables us to quantify the benefits of diversity-aware selection and to determine the data volume required for model performance to saturate for each outcome. At each iteration, 1% of the pool data P is added to the training data, depending on the policy. Explicitly, the policies implemented in this work are: 1) Committee Disagreement. We quan-

where ŷRF,i (s) and ŷXGB,i (s) denote the predictions of the RF and XGB models trained specifically for outcome i. The reference sample s∗i = arg max |ŷRF,i (s) − ŷXGB,i (s)| s∈P

is the pool element exhibiting the maximum model disagreement for that outcome. The denominator normalizes the disagreement magnitude per outcome by its maximum observed scale in the pool, preventing outcomes with inherently larger numerical ranges from dominating the score. Consequently, Dis(s) measures relative cross-model model disagreement aggregated across outcomes. The acquisition policy selects samples that maximize Dis(s), thereby prioritizing regions of the design space where independently trained model families disagree most strongly, i.e. areas of elevated epistemic uncertainty. 2) Explicitly Diversity-Weighted Selection. Batch selection is formulated as a multiobjective optimization problem solved using the genetic evolutionary algorithm NSGA-II procedure [47]. Chromosomes correspond to indices of pool samples; therefore, each individual sample in the pool represents a candidate batch B ⊂ P . The evolutionary search identifies batches that lie on the Pareto front defined by (i) predictive disagreement and (ii) diversity. The (normalized) average disagreement for outcome i within a batch is defined as 1 X |ŷRF,i (s) − ŷXGB,i (s)| AvgDisi (B) = |B| s∈B |ŷRF,i (s∗i ) − ŷXGB,i (s∗i )| 5

Moreover, we define internal batch diversity as divwithin (B) = 1 −

where objectives are min–max normalized over the Pareto frontier P as :

X 1 cos(xi , xj ), |B|(|B| − 1) i,j∈B

Div(B) − minDiv(B) g Div(B) =

i̸=j

(2) where |B|(|B| − 1) is the number of ordered sample pairs. This term measures mean dissimilarity within the batch. Furthermore, we define batch diversity relative to the training data as

B∈P

maxDiv(B) − minDiv(B) B∈P

,

(6)

B∈P

AvgDisi (B) − minAvgDisi (B) ^ i (B) = AvgDis

B∈P

maxAvgDisi (B) − minAvgDisi (B) B∈P

B∈P

(7) These normalizations place all objectives on a common [0, 1] scale. The selected batch therefore lies on the Pareto front while maximizing a balanced trade-off between predictive disagreement and diversity. This described procedure when including diversity is shown in figure 1. 3) Committee Disagreement with NSGAII (without explicit diversity): To compare the effect of incorporating diversity only in outcome space into NSGA-II–based dataset construction (as introduced in Policy 2), we additionally consider a variant in which dataset selection is driven solely by committee disagreement optimized via NSGA-II, without an explicit diversity term. This allows us to assess the extent to which NSGA-II alone can produce broadly useful datasets when diversity is not directly enforced on feature space. Accordingly, the policy defined in Eq. 5 is modified as

! 1 X 1 X 1− cos(xi , xj ) , divTrain (B) = |B| i∈B |T | j∈T (3) where the inner weighted sum represents the average similarity between a sample from the batch and the training set T . Thus, Eq. 3 measures the mean dissimilarity of B from T . At each iteration, we ℓ2 -normalize all feature vectors in both the training set T and the pool P , yielding unit-norm representations. Under this normalization, cosine similarity reduces to the dot product between samples. Finally, the overall diversity objective is

Div(B) = wTrain divTrain (B) + wwithin divwithin (B), (4) with wTrain +wwithin = 1. In this work we use equal 1 X ^ B ∗ = arg max AvgDisi (B), (8) weights; sensitivity analysis is left for future study. |O| i∈O B∈P The NSGA-II algorithm performs selection, crossover, mutation, non-dominated sorting, and where the normalized average disagreement crowding-distance ranking to evolve candidate ^ i (B) is computed as described in Eq. 7. AvgDis batches. Unlike the committee-disagreement polAt each iteration, the evolutionary algorithm exicy, disagreement is not aggregated across outplores candidate batches and identifies solutions comes. Instead, each AvgDisi (B) constitutes a that lie on the Pareto frontier P, yielding a balseparate objective, along with Div(B). anced trade-off among predictive disagreements After the maximum number of generations of across the outcomes of interest. sample batches has been reached, we select the Although this policy does not explicitly infinal batch from the Pareto set P using a normalcorporate diversity in the input space, NSGAized scalarization: II enforces crowding distance in the objective ! (disagreement) space. As a result, the selected X 1 g ^ i (B) , batches exhibit implicit diversity in the output Div(B) + AvgDis B ∗ = arg max |O| + 1 B∈P space, enabling comparison with diversity-aware i∈O (5) variants and clarifying the role of explicit diver6

.

Figure 2: RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool JARVIS18.

sity mechanisms in constructing robust, broadly informative datasets.

3

The mechanism underlying this behavior is revealed by the coverage of the chemical data manifold. As shown in Fig. 3, when diversity is included, the sampled set preserves the overall distributional envelope of the full pool dataset. In contrast, single objective focused policy produces a skewed training distribution, even when samples have already been collected in certain regions of the principal component (PC) space. For example, Fig. 3 left panel shows that the region with P C1 < −5 remains largely unexplored despite the selection of multiple samples with P C1 > −5. This imbalance persists even when 60% of the data has been sampled, although one might expect model disagreement between XGBoost and Random Forest to be larger in such unexplored regions. By contrast, when diversity is incorporated, samples are consistently distributed across the entire data manifold. As a result, the sampled dataset maintains the global envelope of the underlying distribution. On the other hand, introducing diversity can moderately reduce the ability to select the most informative samples for predicting the targeted property, as shown in Fig. 2a). Across all datasets and single-objective policies, the largest RMSE degradation at the saturation point—defined as the stage where the non-diversity policy exhibits less than 5% degradation relative to the full

Results

In this section, we present representative results using the DFT and the sysTEM (experimental) datasets as material pools. These results illustrate the behavior of our framework in both experimental and computational settings. Additional results are provided in the Appendix. 3.1

DFT Datasets construction

Samples selection without explicit diversity introduces pronounced sampling biases. For example, when building datasets by targeting bandgap on JARVIS18 pool, predictive performance on unseen properties—bulk modulus and formation energy—is severely degraded relative to random sampling, as shown in Figs. 2b) and 2c). Introducing diversity into the acquisition strategy removes this behavior. When diversity is included, the RMSE for bulk modulus becomes comparable to that of random sampling Fig. 2b), while the RMSE for formation energy is indistinguishable from random sampling during early iterations and slightly lower after ∼ 35% of data collection Fig. 2c). 7

Figure 3: Data Manifold coverage comparison without and with diversity for single target DFT dataset building. Notoriously, skewed distributions are obtained even with 50% of sampled data when unconsidering diversity.

dataset (here ∼ 3.41%)—occurs for QBC bandgap sampling (Fig. 2a)). In this worst case scenario, RMSE saturation for bandgap prediction is achieved when approximately 40% of the pool is sampled using QBC, whereas saturation with diversity inclusion occurs near 70% of the data. Despite this difference in sampling efficiency, both approaches achieve similar predictive accuracy. When 40% of the data has been sampled, the RMSE differs by only ∼ 0.04, eV when diversity is included respect to non-diversity. This corresponds to a ∼ 10.39% degradation relative to the full dataset and ∼ 6.76% relative to the nondiversity.

tion decreases from ∼ 31.38% without diversity to ∼ 18.41% when diversity is included. Across DFT datasets, this pattern also holds: on not considered outcomes, diversity-aware sampling achieves better performance than non diversity aware. This can be contextualized in terms of   AUCRMSE-Policy Improvement= 1 − × 100%, AUCRMSE-Random where the RMSE area under R the curve (AUC) is calculated as AUCRMSE = RMSE(p) dp where p denotes the percentage of data collected. The only exceptions are JARVIS18 under QBC when targeting formation energy, and when jointly targeting formation energy and bulk modulus while using datasets constructed for bandgap prediction. Nevertheless, the performance can be considered statistically the same given that for these cases standard deviation bars overlap (see Figs. 4, 7). As can be noticed in Fig. 4, diversity inclusion improvement respect to random sampling is always on average above zero for single target datasets construction when using Random Forest as predictor. Using XGBoost, improvements can be negative, but the worst average degradation remains below ∼ 1% relative to random sampling. Also, standard deviation bars overlap with the random average, indicating statistically similar behavior (see Appendix). In contrast, without diversity awareness, not targeted properties can

Importantly, the moderate worst case reduction in targeted performance is offset by improvements in untargeted properties. For example, when 40% of the data has been sampled, the RMSE for bulk modulus prediction decreases from 30.325GPa without diversity to 27.957GPa with diversity, representing a ∼ 7.81% improvement relative to the non-diversity policy. Moreover, the bulk modulus RMSE degradation relative to the full dataset is substantially smaller when diversity is included (∼ 6.4%) than without diversity (∼ 15.37%). A similar trend is observed for formation energy, where the RMSE improves from 0.314 eV atom−1 to 0.283 eV atom−1 , corresponding to a 9.87% improvement. Relative to the full dataset, the formation energy degrada8

Figure 4: Random Forest improvement of all policies respect to random sampling with single outcome targeting in DFT datasets construction.

lead up to ∼ 13% performance degradation respect to random sampling with Random Forests and up to ∼ 17% with XGBoost. Moreover, only three diversity-aware cases, evaluated with the Random Forest regressor, show standard deviation bars overlapping the random sampling mean: MP21 targeting bulk modulus and evaluated on bandgap; JARVIS18 targeting formation energy and evaluated on bulk modulus; and JARVIS18 targeting bulk modulus and evaluated on bandgap 4. In all other cases, the AUC improvement bars lie above the random mean. This indicates that diversity increases the informativeness of the constructed datasets for properties not used during sampling. Overall, these results show that targeted outcome–driven acquisition introduces strong biases in DFT built datasets built. This reduces generalization to untargeted properties. Adding diversity mitigates this effect by improving coverage of the chemical data manifold. As a result, the constructed datasets better preserve the global distribution while maintaining predictive performance. Although diversity can slightly reduce accuracy on the targeted property, this loss is small compared to the gains on untargeted properties. In some cases, diversity also improves targeted performance on average, such as bulk modulus in the two-objective setting (see Fig. 7 for Random

Forest and Appendix for XGBoost). In general, diversity-aware sampling yields datasets that are informative for the targets and representative of the materials space, enabling more accurate predictions across multiple properties. 3.2

Experimental Datasets construction

The bias introduced by outcome-focused acquisition policies is also observed in experimental dataset construction. In contrast, when considering diversity, predictive performance on not considered outcomes behaves comparably to, or better than, random sampling. The reduction in biased behaviour for untargated properties applies for all cases in our study. For single target case can be noticed in Fig. 5, where all cases improvement AUC is increased when including diversity. As shown in Fig. 5, when a dataset is built targeting zT and used to predict the Seebeck coefficient, performance degrades by ∼ 40% relative to random sampling. In contrast, incorporating diversity yields performance comparable to random sampling. For untargeted properties, diversity improves performance respect to random sampling by up to ∼ 10%. For example, this gain is observed when jointly targeting zT and thermal conductivity with diversity and predicting electrical conductivity (see appendix). 9

Figure 5: Random Forest improvement of all policies respect to random sampling with single outcome targeting in experimental datasets construction. Using as pool the sysTEm dataset.

This indicates that diversity plays a decisive role for experimental data where there is noise. In contrast to DFT datasets where targeted focus sampling can have a moderate degradation impact on the targets themselves, in experimental cases diversity ensures better performance than random sampling and greater accuracy than solely targeted policies.

On the other hand, the diversity aware performance on untargated properties can be below than random sampling (Fig. 5). However, in nearly all cases for degraded improvement, standard deviation bars overlap with random sampling mean, indicating statistically comparable performance. The only exception occurs when datasets are constructed with implicit diversity by jointly targeting electrical conductivity, thermal conductivity and seebeck coefficient and evaluated on zT prediction. The RMSE learning curve can be seen in Fig. 9 a). However, even in this worst-case deviation, as noticed in the RMSE curve, the zT untargeted performance remains effectively close to that of random sampling.

Furthermore, in contrast to DFT dataset construction, experimental datasets can not exhibit a clear global skewness in the sampled chemical manifold when targeted QBC is used (Fig. 6). Nevertheless, although target aware policy broadly spans the chemical space, sampling within clusters remains limited. For example, as shown in the PC1–PC2 projection when 30% of points have been sampled building datasets targeting Thermal Conductivity (Fig. 6), in clusters C2, C3 and C4 can be noticed how some regions have been under-sampled when using QBC while QBC with diversity evenly spans within the clusters. Meanwhile, in cluster C1, more points have been sampled when using QBC than when including diversity. This behavior highlights the role of the global diversity objective in Eq. 4, which evaluates diversity not only relative to the previously collected training set but also within each proposed batch. While disagreement promotes broad exploration of the chemical manifold, the diversity objective ensures coverage within clus-

In all cases, diversity-aware sampling outperforms non-diversity methods on the targeted properties. For the Seebeck coefficient, QBC targeting alone yields a mean degradation of ∼ 10% relative to random sampling with Random Forest, whereas QBC with diversity yields a ∼ 12.5% mean improvement (see Fig. 5). This trend is model independent. With XGBoost, targeting alone leads to degradations of ∼ 5% for zT and ∼ 10% for the Seebeck coefficient, while adding diversity yields improvements of ∼ 5% and ∼ 10%, respectively. For targeted properties, performance can degrade respect to random sampling by up to ∼ 12.5% without diversity, while our framework achieves gains of up to ∼ 25% (see appendix). 10

Figure 6: Data Manifold coverage comparison without and with diversity for single target with sysTEm dataset as pool. Global diversity and general coverage is noted without and with diversity. Nevertheless, within clusters without diversity tends to focus in specific regions while diversity aware spans evenly within clusters.

ters. This mechanism is particularly important for the experimental dataset, where many materials are highly similar and often differ only through dopant substitutions. In summary, these results indicate that the incorporation of diversity is particularly critical for the construction of experimental datasets. The diverse selected data is more informative for targeted and not targeted properties. Moreover, diversity awareness as defined by Eq. 4 ensures diversity between clusters and within clusters in the chemical data manifold. Additional experimental results can be found in the appendix. 3.3

that it constrains not only the outcome space but also the feature space. Meanwhile, 2QBC will be used to show data building without diversity. Including feature diversity generally yields slightly lower prediction performance on average than NSGA-II. For instance, three exceptions are observed when using Random Forest as predictor model. First, when datasets are built jointly targeting bandgap and bulk modulus on the MP21 pool and then used to predict bulk modulus (Fig. 6- blue bars in the MP21 bulk modulus targeted cases). Second, when datasets are built jointly targeting bandgap and formation energy and evaluated for bulk modulus prediction in MP18 (Fig. 6- green bars in the MP18 bulk modulus untargeted cases). Third, under the same conditions in MP21, diversity-aware sampling considerably improves performance compared to NSGA-II (Fig. 6- green bars in the MP21 bulk modulus untargeted cases). In most DFT cases, NSGA-II and NSGA-II with feature diversity perform within one standard deviation, indicating similar statistical performance (see Fig. 7). For all constructed experimental datasets with NSGA-II and NSGAII with feature diversity, trained models on these are statistically equivalent, with results within one standard deviation bars. Additional results are provided in the Appendix.

Including Feature-Space Diversity versus Implicit Outcome-Space Diversity in Dataset Construction

When using one outcome, we used NSGA-II to incorporate explicitly feature diversity as objective to find a pareto front with disagreement. When extending to two outcomes, we can compare the difference between building datasets enforcing outcome diversity and jointly diversity of feature space and outcome space. By construction, NSGA-II enforces diversity through the crowding distance mechanism when proposing new batches, which promotes spread in outcome space. On the other hand, explicitly enforcing diversity differs in 11

Figure 7: Improvement of all policies respect to random sampling with two outcomes targeting in DFT datasets construction.

Figure 8: Data Manifold coverage NSGA-II QBC vs NSGA-II QBC with feature diversity. Globally, both create diverse datasets. However, NSGA-II without diversity still creates skewed distributions even with 50% of sampled data.

In all cases, feature diversity-aware sampling performs better than or comparably to random sampling while exploring the full chemical data manifold without producing skewed sampling distributions. This behavior is notable even for the bandgap–bulk modulus case, where the relatively uniform distribution of bulk modulus values would suggest naturally broad sampling across the data manifold. However, as shown in Fig. 8, NSGAII still produces a skewed sampling distribution along PC1. In contrast, enforcing diversity in feature space leads to a more uniform coverage of the

data manifold. 3.4

Datasets for Untargeted Properties with Nonlinear Correlation

Directly sampling three outcomes that are nonlinearly related to a fourth—namely zT =

S 2 σT , κ

where zT is the figure of merit, S the Seebeck coefficient, σ electrical conductivity, T tempera12

Figure 9: RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the targetd used for data construction are electrical conductivity, seebeck coefficient and thermal conductivity.

ture, and κ thermal conductivity—does not improve predictive performance for zT in experimental data.

outcome-driven optimization when the target of interest is not explicitly included. Collectively, in materials science, many properties lack well-established or explicit relationships with other measurable quantities, and it is often unclear which observables are most relevant for modeling a given target. As exploration proceeds across the vast and countably infinite chemical space, our understanding of materials properties continues to evolve. Consequently, if datasets are to remain informative for not well understood targets, diversity-aware sampling strategies are essential. In experimental settings with noise, selecting properties that we know that are nonlinearly related to a quantity does not guarantee that machine learning models will capture the associated relationships more effectively than random sampling. In the most severe cases, biased acquisition can actively obscure these relationships, ultimately impeding—rather than accelerating—materials discovery.

This trend is clearly observed in Fig. 9a). For NSGA-II alone and NSGA-II augmented with diversity (dotted and solid orange curves, respectively), the RMSE for zT remains marginally worse than that of random sampling, despite substantially improved predictive accuracy on the construction targets—Thermal Conductivity, Seebeck Coefficient, and Electrical Conductivity (Figs. 9b), 9c), and 9d). The most adverse behavior emerges in the absence of diversity awareness. In such cases, the acquisition process can induce systematic sampling biases for learning the fourth non-linear related property. Notably, even after the RMSE for the construction targets saturates at approximately 60% of the available data, predictive performance for the figure of merit zT remains inferior to that achieved through random selection. This highlights the limitations of 13

4

Discussion

of feature-space structure. On the other hand, several aspects of our methodology deserve discussion. First, we enforce diversity in feature space using average cosine dissimilarity. It has the potential issue of not considering features correlation and potentially become less sensitive as the number of modes increases [64]. However, cosine dissimilarity is simple and computationally efficient. By normalizing features, the cosine similarity reduces to simple dot product. Therefore, our approach for diversity calculation makes it practical for active learning on large datasets. Robust computation of diversity for our framework while making the computation reasonable time consuming is left as future work. Second, our feature-diversity calculation may suffer from the curse of dimensionality in very high-dimensional spaces [65–67]. However, we use all features to avoid losing information and to keep diversity defined only in feature space. Assigning weights or removing features would require assumptions about how features affect the outcome space, making the diversity measure dependent on current targets or collected samples. Furthermore, we do not aim to construct datasets whose diversity is outcome- or task-dependent, as this would mainly benefit tasks of current interest [68]. In contrast, our goal is to keep diversity independent of any specific target, so that the resulting dataset remains useful for future objectives. We tested the method on 273-dimensional feature space for DFT datasets, where the diversity measure selected samples broadly across the chemical space. This strongly suggests its applicability up to this dimensionality. Testing in higher-dimensional settings, and methods to address possible curse of dimensionality effects, are left for future work. Third, we evaluate our framework on five datasets. Dataset choice for pool samples can introduce bias into the dataset construction behavior. However, these datasets span experimental and computational data distributions. This suggests that our findings are broadly generalizable. Our framework also highlight an important point: informative, and reusable datasets can be

To the best of our knowledge, this is the first of its kind framework for materials dataset construction that accounts for current targets while preserving informativeness for future targets. Our approach is fundamentally different from prior work. Many diversity-aware Bayesian optimization methods focus on selecting diverse solutions within high-performing candidates.[48–53]. Also, different Batch Bayesian Optimization methods have been developed [54–56] and applied in materials science [57–61] where samples to be selected for the design targets are spread in the chemical space. These strategies are conceptually similar to NSGA-II applied to target disagreements in outcome space. Nevertheless, as shown in figure 8, these approaches can lead to skewed distributions. In contrast, our method does not restrict diversity to high-performing candidates nor enforce spread only in outcome space. Instead, it identifies candidates on the Pareto front P defined by both informativeness and diversity, and selects batches that balance these objectives. This leads to datasets that remain useful beyond the initial targets while maintaining broad coverage of the chemical space. Optimizing all objectives during data collection becomes ineffective as the number of objectives m increases. The probability that one sample dominates another decays exponentially as 2−(m−1) [62],[63], rapidly weakening effective selection. Even for m = 4, this probability is only 0.125, making it unlikely to identify samples that meaningfully improve the targets. In contrast, our framework focuses on the objectives of interest, enabling efficient identification of true Pareto-optimal samples, while performance on untargeted objectives remains better or comparable to random sampling. Our framework effectively selects informative samples for targeted objectives across different dataset settings. The experimental dataset contains clustered samples in feature space, as many compounds differ only by dopants. In contrast, the DFT datasets are broadly distributed. Hence, our results strongly suggests that our diversitybased approach performs consistently regardless 14

achieved. Fourth, we use NSGA-II to find the pareto front P of diversity and informativeness. We selected this method because it is multi-objective computationally efficient and promotes diversity through crowding distance among Pareto-optimal candidates. This makes this particular algorithm ideal for our diversity purposes. Other multi-objective optimization methods may also be effective and should be explored in future work. In summary, this work introduces the first materials framework for dataset construction for immediate and future objectives. The results support the central hypothesis that combining informativeness and diversity yields datasets that serve current goals while reducing bias for future tasks. More broadly, this work shows that informative and reusable materials datasets can be built systematically. Several directions remain to be explored for improvement. We hope this work encourages deeper understanding and improved methods for efficiently creating reusable datasets.

5

erty. However, incorporating diversity mitigates this limitation and improves performance to being comparable to random sampling. We further show that with implicit outcome diversity strategies as NSGA-II algorithm, data collection can produce skewed distributions that do not cover the chemical data manifold. Our framework explicitly enforces diversity in chemical space and therefore improves manifold coverage. Also, for all experimental cases the improvement AUC for NSGA-II and NSGA-II with explicit feature diversity overlap their uncertainty bands, showing no statistically difference in prediction performance. Overall, our results show that our diversity framework allows future modeling needs to be anticipated. This approach mitigates cold-start failures, reduces worst-case performance degradation on unmeasured outcomes, enables broad chemical space exploration while maintaining performance and preserves transferability across objectives. As a result, our framework enables data collection campaigns that are not only efficient in the short term but also robust and reusable in the long term.

Conclusions

We show that data building strategies that include diversity improves predictive performance on properties not targeted during dataset construction. For DFT datasets construction, including diversity, on average can slightly degrade performance on the targeted properties compared with policies that focus only the targets. Nevertheless, the degradation on targets is generally within uncertainty bands of only targeted dataset building policies. Moreover, diversity mitigates biases for not targeted outcomes, making prediction performance equally or better than random sampling. On the other hand, for noisy experimental datasets building, incorporating diversity improves performance on the targeted and not targeted outcomes. We also find that in noisy experimental settings, collecting informative samples that are nonlinearly related to another property does not guaranty good predictive performance for that prop-

Author contributions R.E.C. and J.H.S. conceived and designed the project. J.H.S. supervised the project. R.E.C and A.D. conducted the experiments. R.E.C. drafted the manuscript. Hongchen W., K.L. and A.B.D. provided technical support for the diversity policy construction. R.Z, Y.K. and Hao W. helped to conduct experiments in the Digital Research Alliance of Canada supercomputer. All authors discussed the results and reviewed the manuscript.

Conflicts of interest The authors declare no conflicts of interest.

Data and code availability Data and code will be made available on GitHub. 15

Acknowledgments

A. Davydov, J. Jiang, R. Pachter, G. Cheon, E. Reed, A. Agrawal, X. Qian, V. Sharma, H. Zhuang, S. V. Kalinin, B. G. Sumpter, G. Pilania, P. Acar, S. Mandal, K. Haule, D. Vanderbilt, K. Rabe and F. Tavazza, npj Computational Materials, 2020, 6, 173.

The authors acknowledge financial support from the Natural Sciences and Engineering Research Council of Canada (NSERC) Alliance grants (ALLRP 601812-24), the National Research Council of Canada’s Critical Battery Materials Initiative (CBMI-002-1). The research was also, in part, made possible thanks to funding provided to the University of Toronto’s Acceleration Consortium by the Canada First Research Excellence Fund (CFREF-2022-00042).

[10] S. Kirklin, J. E. Saal, B. Meredig, A. Thompson, J. W. Doak, M. Aykol, S. Rühl and C. Wolverton, npj Computational Materials, 2015, 1, 15010. [11] J. Shen, S. D. Griesemer, A. Gopakumar, B. Baldassarri, J. E. Saal, M. Aykol, V. I. Hegde and C. Wolverton, Journal of Physics: Materials, 2022, 5, 031001. [12] J. E. Saal, S. Kirklin, M. Aykol, B. Meredig and C. Wolverton, JOM, 2013, 65, 1501–1509.

References

[13] S. Divilov, H. Eckert, S. D. Thiel, S. D. Griesemer, R. Friedrich, N. H. Anderson, M. J. Mehl, D. Hicks, M. Esters, N. Hotz, X. Campilongo, A. Calzolari and S. Curtarolo, High Entropy Alloys & Materials, 2025, 3, 178–187.

[1] J. Liu, Z. Wang, J. Kou and K. Chen, Photon Science, 2026, 1, 91–103. [2] J. S. Solomon, N. Mrkyvkova, V. Kliner, T. SotoMontero, I. Fernandez-Guillen, M. Ledinský, P. P. Boix, P. Siffalovic and M. Morales-Masis, npj 2D Materials and Applications, 2025, 9, 50.

[14] C. Oses, M. Esters, D. Hicks, S. Divilov, H. Eckert, R. Friedrich, M. J. Mehl, A. Smolyanyuk, X. Campilongo, A. van de Walle, J. Schroers, A. G. Kusne, I. Takeuchi, E. Zurek, M. B. Nardelli, M. Fornari, Y. Lederer, O. Levy, C. Toher and S. Curtarolo, Computational Materials Science, 2023, 217, 111889.

[3] G. W. Lee, A. K. Gangopadhyay, K. F. Kelton, R. W. Hyers, T. J. Rathz, J. R. Rogers and D. S. Robinson, Physical review letters, 2004, 93, 037802.

[15] D. Zagorac, H. Müller, S. Ruehl, J. Zagorac and S. Rehme, Journal of Applied Crystallography, 2019, 52, 918–925.

[4] M. Mascagni, N. Tchipev, S. Seckler, M. Heinen, J. Vrabec, F. Gratl, M. Horsch, M. Bernreuther, C. W. Glass, C. Niethammer, N. Hammer, B. Krischok, M. Resch, D. Kranzlmüller, H. Hasse, H.-J. Bungartz and P. Neumann, Int. J. High Perform. Comput. Appl., 2019, 33, 838–854.

[16] R. Allmann and R. Hinek, Acta Crystallographica Section A, 2007, 63, 412–417. [17] A. Belsky, M. Hellenbrandt, V. L. Karen and P. Luksch, Acta Crystallographica Section B, 2002, 58, 364–369.

[5] K. KADAU, T. C. GERMANN and P. S. LOMDAHL, International Journal of Modern Physics C, 2006, 17, 1755–1761.

[18] J.-W. Lee, W. B. Park, J. H. Lee, S. P. Singh and K.-S. Sohn, Nature Communications, 2020, 11, 86.

[6] D. Wines and K. Choudhary, Materials Futures, 2024, 3, 025602.

[19] B. Cao, Z. Zheng, Y. Liu, L. Zhang, L. W.-Y. Wong, L.-T. Weng, J. Li, H. Li and T.-Y. Zhang, National Science Review, 2025, 12, nwaf421.

[7] A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder and K. A. Persson, APL Materials, 2013, 1, 011002.

[20] T. Xie and J. C. Grossman, Phys. Rev. Lett., 2018, 120, 145301. [21] K. Choudhary and B. DeCost, npj Computational Materials, 2021, 7, 185.

[8] A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder and K. A. Persson, APL Materials, 2013, 1, 011002.

[22] R. E. A. Goodall and A. A. Lee, Nature Communications, 2020, 11, 6280.

[9] K. Choudhary, K. F. Garrity, A. C. E. Reid, B. DeCost, A. J. Biacchi, A. R. Hight Walker, Z. Trautt, J. Hattrick-Simpers, A. G. Kusne, A. Centrone,

[23] D. Jha, L. Ward, A. Paul, W.-k. Liao, A. Choudhary, C. Wolverton and A. Agrawal, Scientific Reports, 2018, 8, 17593.

16

[24] O. Isayev, C. Oses, C. Toher, E. Gossett, S. Curtarolo and A. Tropsha, Nature Communications, 2017, 8, 15679.

[37] H. Wang, R. E. Castañeda, J. R. Werber, Y. Fehlis, E. Kim and J. Hattrick-Simpers, Training-Free Active Learning Framework in Materials Science with Large Language Models, 2025, https://arxiv.org/ abs/2511.19730.

[25] Y. Dan, Y. Zhao, X. Li, S. Li, M. Hu and J. Hu, npj Computational Materials, 2020, 6, 84.

[38] R. Xin, E. M. D. Siriwardane, Y. Song, Y. Zhao, S.-Y. Louis, A. Nasiri and J. Hu, The Journal of Physical Chemistry C, 2021, 125, 16118–16128.

[26] J. Noh, J. Kim, H. S. Stein, B. Sanchez-Lengeling, J. M. Gregoire, A. Aspuru-Guzik and Y. Jung, Matter, 2019, 1, 1370–1384.

[39] L. Rebuffi, S. Kandel, X. Shi, R. Zhang, R. J. Harder, W. Cha, M. J. Highland, M. G. Frith, L. Assoufid and M. J. Cherukara, Opt. Express, 2023, 31, 39514– 39527.

[27] H. Xiao, R. Li, X. Shi, Y. Chen, L. Zhu, X. Chen and L. Wang, Nature Communications, 2023, 14, 7027. [28] C. Zeni, R. Pinsler, D. Zügner, A. Fowler, M. Horton, X. Fu, Z. Wang, A. Shysheya, J. Crabbé, S. Ueda, R. Sordillo, L. Sun, J. Smith, B. Nguyen, H. Schulz, S. Lewis, C.-W. Huang, Z. Lu, Y. Zhou, H. Yang, H. Hao, J. Li, C. Yang, W. Li, R. Tomioka and T. Xie, Nature, 2025, 639, 624–632.

[40] K. Li, D. Persaud, K. Choudhary, B. DeCost, M. Greenwood and J. Hattrick-Simpers, Nature Communications, 2023, 14, 7283. [41] L. Chanussot, A. Das, S. Goyal, T. Lavril, M. Shuaibi, M. Riviere, K. Tran, J. Heras-Domingo, C. Ho, W. Hu, A. Palizhati, A. Sriram, B. Wood, J. Yoon, D. Parikh, C. L. Zitnick and Z. Ulissi, ACS Catalysis, 2021, 11, 13062–13065.

[29] L. Ward, A. Agrawal, A. Choudhary and C. Wolverton, npj Computational Materials, 2016, 2, 16028. [30] A. G. Kusne, H. Yu, C. Wu, H. Zhang, J. HattrickSimpers, B. DeCost, S. Sarker, C. Oses, C. Toher, S. Curtarolo, A. V. Davydov, R. Agarwal, L. A. Bendersky, M. Li, A. Mehta and I. Takeuchi, Nature Communications, 2020, 11, 5966.

[42] R. Tran, J. Lan, M. Shuaibi, B. M. Wood, S. Goyal, A. Das, J. Heras-Domingo, A. Kolluru, A. Rizvi, N. Shoghi, A. Sriram, F. Therrien, J. Abed, O. Voznyy, E. H. Sargent, Z. Ulissi and C. L. Zitnick, ACS Catalysis, 2023, 13, 3066–3084.

[31] N. J. Szymanski, B. Rendy, Y. Fei, R. E. Kumar, T. He, D. Milsted, M. J. McDermott, M. Gallant, E. D. Cubuk, A. Merchant, H. Kim, A. Jain, C. J. Bartel, K. Persson, Y. Zeng and G. Ceder, Nature, 2023, 624, 86–91.

[43] L. Ward, R. Liu, A. Krishna, V. I. Hegde, A. Agrawal, A. Choudhary and C. Wolverton, Phys. Rev. B, 2017, 96, 024104. [44] L. Ward, A. Dunn, A. Faghaninia, N. E. Zimmermann, S. Bajaj, Q. Wang, J. Montoya, J. Chen, K. Bystrom, M. Dylla, K. Chard, M. Asta, K. A. Persson, G. J. Snyder, I. Foster and A. Jain, Computational Materials Science, 2018, 152, 60–69.

[32] E. A. Pogue, A. New, K. McElroy, N. Q. Le, M. J. Pekala, I. McCue, E. Gienger, J. Domenico, E. Hedrick, T. M. McQueen, B. Wilfong, C. D. Piatko, C. R. Ratto, A. Lennon, C. Chung, T. Montalbano, G. Bassen and C. D. Stiles, npj Computational Materials, 2023, 9, 181.

[45] L. Tang, L. Purdy, T. Mohanty, L. Ng and T. Sparks, Systematically Verified Experimental Thermoelectric Dataset For Data-driven Approaches, 2025.

[33] A. Palizhati, S. B. Torrisi, M. Aykol, S. K. Suram, J. S. Hummelshøj and J. H. Montoya, Scientific Reports, 2022, 12, 4694. [34] D. Xue, D. Xue, R. Yuan, Y. Zhou, P. V. Balachandran, X. Ding, J. Sun and T. Lookman, Acta Materialia, 2017, 125, 532–541.

[46] L. Grinsztajn, E. Oyallon and G. Varoquaux, Proceedings of the 36th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2022.

[35] D. Xue, P. V. Balachandran, J. Hogden, J. Theiler, D. Xue and T. Lookman, Nature Communications, 2016, 7, 11241.

[47] K. Deb, A. Pratap, S. Agarwal and T. Meyarivan, IEEE Transactions on Evolutionary Computation, 2002, 6, 182–197.

[36] D. Xue, P. V. Balachandran, R. Yuan, T. Hu, X. Qian, E. R. Dougherty and T. Lookman, Proceedings of the National Academy of Sciences, 2016, 113, 13301– 13306.

[48] D. Eriksson, M. Pearce, J. R. Gardner, R. Turner and M. Poloczek, Scalable Global Optimization via Local Bayesian Optimization, 2020, https://arxiv.org/ abs/1910.01739.

17

[49] N. Maus, K. Wu, D. Eriksson and J. Gardner, Discovering Many Diverse Solutions with Bayesian Optimization, 2023, https://arxiv.org/abs/2210. 10953.

[62] L.-s. Wei and E.-c. Li, Journal of Computational Design and Engineering, 2023, 10, 1988–2018. [63] M. Li, S. Yang and X. Liu, Artificial Intelligence, 2015, 228, 45–65.

[50] H. P. Vanchinathan, A. Marfurt, C.-A. Robelin, D. Kossmann and A. Krause, Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 2015, p. 1195–1204.

[64] D. Friedman and A. B. Dieng, The Vendi Score: A Diversity Evaluation Metric for Machine Learning, 2023, https://arxiv.org/abs/2210.02410. [65] K. S. Beyer, J. Goldstein, R. Ramakrishnan and U. Shaft, Proceedings of the 7th International Conference on Database Theory, Berlin, Heidelberg, 1999, p. 217–235.

[51] Q. Nguyen and A. B. Dieng, Quality-Weighted Vendi Scores And Their Application To Diverse Experimental Design, 2024, https://arxiv.org/abs/2405. 02449. [52] G. Malkomes, B. Cheng, E. H. Lee and M. Mccourt, Proceedings of the 38th International Conference on Machine Learning, 2021, pp. 7423–7434.

[66] C. C. Aggarwal, A. Hinneburg and D. A. Keim, Proceedings of the 8th International Conference on Database Theory, Berlin, Heidelberg, 2001, p. 420–434.

[53] E. Nava, M. Mutný and A. Krause, Diversified Sampling for Batched Bayesian Optimization with Determinantal Point Processes, 2022, https://arxiv. org/abs/2110.11665.

[67] M. Radovanović, A. Nanopoulos and M. Ivanović, Proceedings of the 26th Annual International Conference on Machine Learning, New York, NY, USA, 2009, p. 865–872. [68] B. T. Mamillapalli, M. R. Kochi and S. M. Moosavi, AI for Accelerated Materials Design - ICLR 2026, 2026.

[54] C. Gong, J. Peng and Q. Liu, Proceedings of the 36th International Conference on Machine Learning, 2019, pp. 2347–2356. [55] E. Nava, M. Mutny and A. Krause, Proceedings of The 25th International Conference on Artificial Intelligence and Statistics, 2022, pp. 7031–7054. [56] T. Kathuria, A. Deshpande and P. Kohli, Proceedings of the 30th International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2016, p. 4213–4221. [57] R. Shibukawa, S. Matsuda, K. Nakamura, R. Tamura and K. Tsuda, npj Computational Materials, 2026. [58] R. Couperthwaite, A. Molkeri, D. Khatamsaz, A. Srivastava, D. Allaire and R. Arròyave, JOM, 2020, 72, 4431–4443. [59] N. Wilson, D. Willhelm, X. Qian, R. Arróyave and X. Qian, Computational Materials Science, 2022, 208, 111330. [60] T. Hastings, M. Mulukutla, D. Khatamsaz, D. Salas, W. Xu, D. Lewis, N. Person, M. Skokan, B. Miller, J. Paramore, B. Butler, D. Allaire, V. Attari, I. Karaman, G. Pharr, A. Srivastava and R. Arróyave, Acta Materialia, 2025, 297, 121173. [61] S. M. A. A. Alvi, B. Vela, V. Attari, J. Janssen, D. Perez, D. Allaire and R. Arróyave, npj Computational Materials, 2026, 12, 105.

18

A

Appendix

Figure 10: Pearson correlation of outcomes variables of DFT datasets.

Figure 11: Pearson correlation of outcomes variables of the sysTEm experimental dataset.

19

Figure 12: Outcomes distribution shown in the first two PCs of feature space.

Figure 13: Outcomes distribution shown in the first two PCs of feature space.

Figure 14: Outcomes distribution shown in the first two PCs of feature space.

20

Figure 15: Outcomes distribution shown in the first two PCs of feature space.

Figure 16: Outcomes distribution shown in the first two PCs of feature space in the sysTEm dataset.

21

A.1

Thermoelectric Dataset-sysTEm

A.2

Improvement metrics

Figure 17: XGBoost improvement of all policies respect to random sampling with single outcome targeting in experimental datasets construction.

Figure 18: Random Forest Improvement of all policies respect to random sampling with two outcomes targeting in experimental datasets construction. Using as pool the sysTEm dataset.

22

Figure 19: XGBoost improvement of all policies respect to random sampling with two outcomes targeting in experimental datasets construction.

Figure 20: Random Forest improvement of all policies respect to random sampling with three outcomes targeting in experimental datasets construction.

23

Figure 21: XGBoost improvement of all policies respect to random sampling with three outcomes targeting in experimental datasets construction.

24

A.2.1

Single Target Dataset Construction (performance metrics-Random Forest)

Figure 22: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is Thermal Conductivity.

Figure 23: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is zT.

25

Figure 24: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is electrical conductivity.

Figure 25: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is Seebeck coefficient.

A.2.2

Two targets Dataset Construction (performance metrics-Random Forest)

26

Figure 26: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity and Thermal Conductivity.

Figure 27: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity and zT.

27

Figure 28: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Seebeck Coefficient and Thermal Conductivity.

Figure 29: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Seebeck Coefficient and zT.

28

Figure 30: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are zT and Thermal Conductivity.

29

A.2.3

Three targets Dataset Construction (performance metrics-Random Forest)

Figure 31: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity, Seebeck Coefficient and zT.

30

Figure 32: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity, zT and Thermal Conductivity.

Figure 33: Random Forest RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Seebeck Coefficient, zT and Thermal Conductivity.

31

A.2.4

Single Target Dataset Construction (performance metrics-XGBoost)

Figure 34: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is Thermal Conductivity.

Figure 35: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is zT.

32

Figure 36: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is Electrical Conductivity.

Figure 37: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction is Seebeck Coefficient.

A.2.5

Two targets Dataset Construction (performance metrics-XGBoost)

33

Figure 38: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity and Thermal Conductivity.

Figure 39: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity and Thermal Conductivity.

34

Figure 40: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity and zT.

Figure 41: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Seebeck Coefficient and Thermal Conductivity.

35

Figure 42: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Seebeck Coefficient and zT.

Figure 43: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are zT and Thermal Conductivity.

A.2.6

Three targets Dataset Construction (performance metrics-XGBoost)

36

Figure 44: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity, Seebeck Coefficient and Thermal Conductivity.

Figure 45: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity, Seebeck Coefficient and zT.

37

Figure 46: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity, Thermal Conductivity and zT.

Figure 47: XGBoost RMSE curves on hold out test data for Thermal Conductivity, zT, Seebeck Coefficient and Electrical conductivity when the target used for data construction are Electrical Conductivity, Thermal Conductivity and zT.

38

B

DFT Improvement metrics with XGBoost

Figure 48: XGBoost improvement of all policies respect to random sampling with single outcome targeting in DFT datasets construction.

Figure 49: XGBoost improvement of all policies respect to random sampling with two outcomes targeting in DFT datasets construction.

C

JARVIS 18

C.1

Single Target Dataset Construction (performance metrics-Random Forest)

39

Figure 50: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool JARVIS18.

Figure 51: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool JARVIS18.

40

C.2

Two Targets Dataset Construction (performance metrics-Random Forest)

Figure 52: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and bulk modulus using as pool JARVIS18.

41

Figure 53: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and formation energy using as pool JARVIS18.

Figure 54: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are formation energy and bulk modulus using as pool JARVIS18.

42

C.3

Single target Dataset Construction (performance metrics-XGBoost)

Figure 55: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool MP21.

Figure 56: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool MP21.

43

Figure 57: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool MP21.

C.4

Two Targets Dataset Construction (performance metrics-XGBoost)

Figure 58: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and bulk modulus using as pool JARVIS18.

44

Figure 59: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and formation energy using as pool JARVIS18.

Figure 60: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are formation energy and bulk modulus using as pool JARVIS18.

45

C.5

Data-Manifold Coverage

46

47 Figure 61

48 Figure 62

49 Figure 63

D

JARVIS 22

D.1

Single Target Dataset Construction (performance metrics-Random Forest)

Figure 64: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool JARVIS22.

Figure 65: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool JARVIS22.

50

Figure 66: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool JARVIS22.

D.2

Two targets Dataset Construction (performance metrics-Random Forest)

Figure 67: RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and bulk modulus using as pool JARVIS22.

51

Figure 68: RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and formation energy using as pool JARVIS22.

Figure 69: RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are formation energy and bulk modulus using as pool JARVIS22.

52

D.3

Single Target Dataset Construction (performance metrics-XGBoost)

Figure 70: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool JARVIS22.

Figure 71: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool JARVIS22.

53

Figure 72: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool JARVIS22.

D.4

Two targets Dataset Construction (performance metrics-XGBoost)

Figure 73: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and bulk modulus using as pool JARVIS22.

54

Figure 74: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and formation energy using as pool JARVIS22.

Figure 75: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are formation energy and bulk modulus using as pool JARVIS22.

55

D.5

Data-Manifold Coverage

Figure 76

Figure 77

56

Figure 78

E

MP 18

E.1

Single Target Dataset Construction (performance metrics-Random Forest)

Figure 79: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool MP18.

57

Figure 80: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool MP18.

Figure 81: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool MP18.

58

E.2

Two Targets Dataset Construction (performance metrics-Random Forest)

Figure 82: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and bulk modulus using as pool MP18.

59

Figure 83: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and formation energy using as pool MP18.

Figure 84: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are formation energy and bulk modulus using as pool MP18.

60

E.3

Single Target Dataset Construction (performance metrics-XGBoost)

Figure 85: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool MP18.

Figure 86: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool MP18.

61

Figure 87: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool MP18.

E.4

Two Targets Dataset Construction (performance metrics-XGBoost)

Figure 88: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and bulk modulus using as pool MP18.

62

Figure 89: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are bandgap and formation energy using as pool MP18.

Figure 90: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the targets used for data construction are formation energy and bulk modulus using as pool MP18.

63

E.5

Data-Manifold Coverage

Figure 91

Figure 92

64

Figure 93

F

MP 21

F.1

Single target Dataset Construction (performance metrics-Random Forest)

Figure 94: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool MP21.

65

Figure 95: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool MP21.

Figure 96: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool MP21.

66

F.2

Two targets Dataset Construction (performance metrics-Random Forest)

Figure 97: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction are bandgap and bulk modulus using as pool MP21.

67

Figure 98: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction are bandgap and formation energy using as pool MP21.

Figure 99: Random Forest RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction are formation energy and bulk modulus using as pool MP21.

68

F.3

Single target Dataset Construction (performance metrics-XGBoost)

Figure 100: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bandgap using as pool MP21.

Figure 101: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is bulk modulus using as pool MP21.

69

Figure 102: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction is formation energy using as pool MP21.

F.4

Two targets Dataset Construction (performance metrics-XGBoost)

Figure 103: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction are bandgap and bulk modulus using as pool MP21.

70

Figure 104: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction are bandgap and formation energy using as pool MP21.

Figure 105: XGBoost RMSE curves on hold out test data for bandgap, bulk modulus and formation energy when the target used for data construction are formation energy and bulk modulus using as pool MP21.

71

F.5

Data-Manifold Coverage

Figure 106

Figure 107

72

Figure 108

73

Record · ID 158547 · SHA-256 4a60e83362fc01e5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.