Conceptio › Archive › arXiv CS
arXiv CSopen access

Bayesian classification of astronomical spectra with class uncertainties

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Astronomy & Astrophysics manuscript no. aa56350-25 September 21, 2026

©ESO 2026

Bayesian classification of astronomical spectra with class uncertainties Simon Barton1 , Martin Sahlén1 , Andreas Korn2 , and Christian Glaser3, 4

arXiv:2609.21694v1 [astro-ph.IM] 18 Sep 2026

1

Theoretical Astrophysics, Division of Astronomy and Space Physics, Department of Physics and Astronomy, Uppsala University, Box 524, 751 20 Uppsala, Sweden. Corresponding author e-mail: [email protected] 2 Observational Astrophysics, Division of Astronomy and Space Physics, Department of Physics and Astronomy, Uppsala University, Box 524, 751 20 Uppsala, Sweden 3 Division of High Energy Physics, Department of Physics and Astronomy, Uppsala University, Box 524, 751 20 Uppsala, Sweden 4 Dept. of Physics, TU Dortmund University, D-44221 Dortmund, Germany September 21, 2026 ABSTRACT Context. We developed a probabilistic machine learning method with the aim of performing the O(10)-way classification of low- and

high-resolution spectra of stellar and extragalactic targets for the upcoming 4MOST survey. In fulfilment of the survey requirements, this method should be able to express uncertainty in the input data as well as uncertainty introduced in its prediction. Aims. Four different methods are explored: (1) convolutional neural networks (CNNs), (2) the Dirichlet distribution, (3) Monte Carlo dropout (MCD), (4) Bayesian neural Networks (BNNs) + variational inference (VI). Training and validation was performed using labelled spectra from the SDSS database and a custom 4MOST mock dataset. All the methods were compared in terms of the same metrics: accuracy, area under the curve (AUC), expected calibration error (ECE), Shannon entropy, negative log-likelihood (NLL), Brier score, training time, and inference time. Methods. A CNN with simple architecture and ∼ 2 × 105 parameters was trained to achieve classification accuracies of 91.5% on SDSS data and 92.8% on 4MOST mock data. The direct Dirichlet prediction and VI models tested provide uncertainties on class membership probabilities, but they confuse classes more often. The MCD on a CNN is found to be the most suitable; it boosts the point-estimate accuracies to 92.6% and 93.9%, while still providing fast training and sufficiently fast inference. Compared to a standard CNN, the method additionally provides well-calibrated uncertainties at marginal extra cost. Results. Conclusions. Key words. Methods: data analysis – Methods: statistical – Techniques: spectroscopic – Surveys

1. Introduction

The next-generation 4-metre Multi-Object Spectroscopic Telescope (4MOST), due to start full science operations in 2026, will operate with low- and high-resolution spectrographs in parallel to observe up to 2436 celestial targets simultaneously (De Jong et al. 2019).1 Over its first five years, the instrument will acquire up to 40 million spectra, necessitating scalable data processing methods beyond the capabilities of template-based pipelines. Processing the large amounts of data that come with modern astronomical spectroscopic surveys includes the classification of targets, both for confirming catalogue labels and for part of a follow-up after a first photometric detection. Next to conventionally used line-fitting algorithms, which are slow and require physical modelling, machine learning (ML) methods have been successfully used in this processing (Zhong et al. 2024). However, such methods can lead to overconfident predictions. Hence, the fully automated 4MOST classification pipeline calls for probabilistic models that estimate uncertainty on class membership. Understanding and quantifying the uncertainty in classification tasks has become increasingly important, particularly in high-stakes applications where model confidence must 1

https://www.4most.eu

be both interpretable and well calibrated. The foundational notion of calibration was introduced by Guo et al. (2017) and was refined by the agreement between predicted probabilities and observed frequencies (Nixon et al. 2019), by inspired recalibration strategies (Kranzlein et al. 2021), and by improved evaluation techniques for multi-class settings (Silva Filho et al. 2023). In parallel, the literature has expanded on how to represent and decompose predictive uncertainty. The methods range from geometric and entropy-based measures (Chlaily et al. 2023) to Bayesian approaches that consider the distribution over performance metrics derived from confusion matrices rather than predictive distributions directly (Caelen 2017). Bayesian neural networks (BNNs) have emerged as promising approaches for modelling uncertainty in deep learning by placing probability distributions over model parameters. Unlike conventional neural networks that yield point estimates, BNNs capture epistemic uncertainty by marginalizing over these distributions, thus resulting in predictive probabilities that reflect both model confidence and data ambiguity. This makes them particularly well suited for applications where calibrated uncertainty estimates are critical, such as classification under limited or noisy data (Gawlikowski et al. 2023; Goan & Fookes 2020). The decomposition of uncertainty into aleatoric and epistemic components was implemented in practical deep learning Article number, page 1

A&A proofs: manuscript no. aa56350-25

Flux

Flux

bels in their coverage of a wide range of both galactic and extragalactic objects. In total, we used 260,000 spectra in splits, as described above. Some example spectra are displayed in Fig. 1.

0.2

Flux

0.25

0.25

0.00 0.50

QSO (S/N=48)

QSO-BROADLINE (S/N=63) 5000

6000 7000 Wavelength [Å]

8000

0.3

9000

1 - Gini impurity

0.2 0.1 0.8

STAR-K5 STAR-K3 STAR-K1 STAR-G2 STAR-F9 STAR-F5 STAR-A0

0.7

Density

0.6 0.5

QSO QSO-BROADLINE GALAXY GALAXY-STARFORMING GALAXY-STARBURST GALAXY-AGN

0.4

The spectra of our first dataset were obtained from eBOSS and are available as part of the SDSS Data Release 16 (Jönsson et al. 2020). Unnormalized fluxes are given for 3,600 wavelengths between 4000Å and 9000Å. The labels correspond to the 13 SDSS (sub)classes with more than 20 000 classified objects: seven galactic classes of different main sequence stars (distinguished by their surface temperature according to the Harvard spectral classification system, where each letter is further subdivided numerically, e.g. G2 star) and six extragalactic classes (four galaxies of different star-forming activity and nucleus activity, plus quasars with and without broad lines). Although the selected classes only constitute a small subset of the total 181 SDSS subclasses, they are similar to the planned 4MOST classification la-

0.2

Article number, page 2

GALAXY-STARFORMING (S/N=85)

Fig. 1: Selected example spectra of the SDSS dataset with the highest S/N of the respective class.

0.3

https://codeberg.org/simonbarton/bnn

GALAXY (S/N=55)

Flux

0.00 0.8 0.4 0.2 0.0 4000

2.1. SDSS

2

STAR-G2 (S/N=127)

0.0 0.50

2. Spectroscopic datasets To test our classification algorithms we used two custom datasets of low-resolution optical spectra. Both datasets were created for a spectrum classification task in Zhong et al. (2024). To remove distance-dependent variations and bring the data into a form suitable for deep learning, the unnormalized fluxes were rescaled such that the sum of all flux values was equal to 1000 for each spectrum. We note that the first BatchNorm layer normalizes the fluxes again. In total, we used 70% spectra for training and 15% for validation and testing, respectively. The validation set was used to select the hyperparameters and perform early stopping, while the test set was used to compare the results across different models. All splits were artificially class-balanced to ensure stable training and fair comparison across classes. This allowed us to better assess the model’s ability to distinguish minority classes and also to ensure comparability to other models. However, when applied within 4MOST, the true class distribution will be skewed and using a test set that reflects the actual astronomical object abundances may provide a more realistic estimate of deployment performance. In particular, the models may overestimate performance on rare classes and underestimate overall uncertainty under natural class priors. An estimate of the final accuracy for the SDSS dataset is given in Sect. 5, assuming the target class balance matches the SDSS catalogue abundance.

0.4 0.2 0.0

Flux

settings using Bayesian deep learning models (Kendall & Gal 2017; Mancarella et al. 2022). It has been shown how proper scoring rules can be used to quantify and distinguish between these two types of uncertainty in a principled manner (Hofman et al. 2024). BNNs have been applied in many fields, such as gravitational waves (Lin & Wu 2021) and the cosmic microwave background, (Hortúa et al. 2020) and used to estimate redshifts of galaxies from photometric DESI data (Zhou et al. 2025). Specifically for classification, they have been compared to other Bayesian deep learning classifiers for supernova signals (Möller & de Boissière 2020) and to assess modified theories of gravity (Thummel et al. 2024). The comparisons have been made between different methods for uncertainty estimation and their robustness against out-of-distribution data (Vranken et al. 2021; Mancarella et al. 2022). All methods were implemented using PyTorch (Paszke et al. 2019). The entire machine learning code is available online.2 Throughout this work the notation |X| is used for the number of elements in X.

0.1 0.0

0

10

20

30 S/N

40

50

60

Fig. 2: S/N distribution of each class for training and validation sets (bottom). Galactic (blue) and extragalactic (red) targets dominate the low- and high-noise regimes, respectively. Expected random-guessing accuracy based on the class ratio (top). For low S/N, the labels become easier to predict. In addition, class balance varies with the signal-to-noise ratio (S/N; see Fig. 2). Even though the validation set is a good rep-

Simon Barton et al.: Bayesian classification of astronomical spectra with class uncertainties

resentation of the total training data, this correlation can affect classification behaviour across the S/N range. quantify the PK To effect, we calculated the Gini purity G = i=1 p2i with the number of classes K = 13. In the context of classification, G provides an estimate for the expected accuracy of a classifier that guesses labels based solely on the class distribution, without considering the input spectrum. For highly imbalanced datasets, where one class dominates, G ≈ 1 leads to a high baseline accuracy. Conversely, balanced datasets where class probabilities are uniform lead to G = 1/K, corresponding to the expected accuracy of random guessing.

(Guo et al. 2017) P(y∗ = y | p∗iy∗ = p) = p

As a second dataset we used around 200,000 spectra in the same splits as above, simulated using the 4MOST Exposure Time Calculator based on templates for selected 4MOST surveys. The optical wavelength range from 4000Å and 9000Å and the wavelength resolution of 5001 are comparable to the products of the 4MOST low-resolution spectrographs. The five galactic classes from the 4MOST surveys S1–S4 and S9 are metal-poor stars and other dynamics tracers (Dyn), Cepheids in the Magellanic Clouds (GalHR), white dwarfs (ESN), and stars of the galactic disc (GalDiskLR) and of the Magellanic Clouds (MCsn). The five extragalactic classes represent the targets of the extragalactic 4MOST surveys S5–S8 and S10. In particular, we have S6 active galactic nuclei (AGN; COSMO_AGN), S5 bright cluster galaxies (ClusB), S7 galaxies (WAVES), S8 red galaxies (RedGAL), and S10 supernova host galaxies (tides_host). Example spectra for some classes and the S/N distribution of this dataset are shown in Figs. A.1 and A.2, respectively.

3. Uncertainty in classification Given that perfect classification of these spectra is generally unachievable, for example due to noise or imprecise labels, providing uncertainties of the predictions is essential. To formalN ize the problem, we denote the data as D = {xi , yi }i=1 where each xi is a spectrum and each yi ∈ {1, . . . , K} the true label. A non-probabilistic classifier m is then a function m : xi 7→ p∗i· , where p∗ik is the confidence that the i-th sample belongs to the k-th class and · acts as a placeholder. The predicted class y∗i = argmaxk ({p∗ik }) is simply the one with the highest confidence assigned. The canonical metric to evaluate such a model is the accuracy, N

accuracy =

1 X δy∗i =yi N i

,

(1)

with δa=b = 1 if a = b and 0 otherwise. Complementary insight is given by the receiver operating characteristic (ROC) as the relation the true positive rate (efficiency) against the false positive rate (contamination), or in summary its integral (AUC). An ideal classifier achieves an AUC of 1, whereas a model making random predictions yields an AUC of 0.5. While widely used for evaluating binary classification models, we use it in a oneversus-rest fashion for this multi-class problem.

∀i ∈ {1, . . . , N}

.

(2)

For example, the model’s class membership predictions with a confidence of p∗ = 30% should coincide with the true label in 30% of the cases. The discrepancy between the predicted confidence and the empirical accuracy is summarized in the expected calibration error (ECE), which in the limit of infinitely many bins3 reads K

ECE = 2.2. 4MOST mock

∀p ∈ [0, 1]

1 X ECEk , K k

N

ECEk =

1 X δy∗i =yi − p∗ik N i

.

(3)

The ECE (or related error definitions) is essential for any classifier. It is a low calibration error and not the construction via a softmax function, which justifies identifying the output confidence values p∗ with probabilities. 3.2. Entropy

One possibility to quantify the uncertainty of a single point estimate is the Shannon entropy H(p∗i · ) := −

K X

p∗ik log p∗ik

with

0 log 0 := 0

,

(4)

k

which measures the distribution of confidence across classes. Model m remains a point-estimate predictor, in that it is deterministic and it provides a single categorical distribution p∗ . Probabilistic models instead predict a second-order distribution p∗i (y) over these p∗i· (Gawlikowski et al. 2023). The set of all point estimates can be identified with the (K − 1)-standard simplex S k−1 , whose K corners correspond to the classes, and each convex combination thereof represents one categorical distribution with normalized probability mass (see Fig. 3). This allows the calculation of the differential entropy of the joint simplex distribution: Z ∗ h(pi ) = − p∗i (y) log(p∗i (y)) dy . (5) S k−1

Unfortunately, in high dimensions it is practically unfeasible to calculate this entropy from samples without assumptions about the form of the underlying distribution. We tried various methods, such as kernel density estimation, without satisfying results, and therefore do not use the joint entropy as an uncertainty estimate in our discussion. The uncertainty of a prediction y∗i comprises two distinct components. Aleatoric uncertainty (AU) arises from inherent random noise in the data-generating process and is irreducible, regardless of the amount of data available. In contrast, epistemic uncertainty (EU) reflects the model’s ignorance due to limited data or limited model capacity, and can be reduced by incorporating additional information, either in the form of model architecture or increasing the training data. The total predicted uncertainty of a single output p∗ (y) can be decomposed into these two components (Depeweg et al. 2018) as h(p(y | x)) = AU + EU with

3.1. Calibration

AU = Eω [h(p∗ (y | x, ω))] ,

EU = H(Eω [p∗ (y | x, ω)]) ,

To measure the quality of the vector p∗·· directly, we compute the calibration. A model m is said to be calibrated exactly if

3 The ECE can significantly increase with the chosen number of bins, which is why we report only the asymptotic limit.

(6)

Article number, page 3

A&A proofs: manuscript no. aa56350-25 Galaxy

0.0

0.0

1.0

Galaxy

0.0

Galaxy

1.0

2 0.2

4

r Sta

1.0

0.6

0.4

0.0

0.2

0.0

r Sta

1.0

0.8

0.6

0.4

0.2

0.0

r Sta

1.0

0.8

0.6

0.4

QSO

QSO

QSO

0.2

0.4

1.0

1.0

1.0 0.0

0

Log probability density

4 0.0

2 0.6

0.6

0.2

0.4

2

0.8

0.8

4 0.0

0.4

0.8

0.8

0.2

0

Log probability density

2

2 0.6

0.6

0.6

0.4

0.8

0.4

0.4

0

Log probability density

2 0.6

4

0.2

0.2

0.2

0.8

1.0

4

0.8

4

Fig. 3: Demonstration of the entropies of a predictive probability distribution in the three-simplex. Left: Large Shannon entropy H and large differential entropy h. The model assigns a high aleatoric ambiguity between galaxy and QSO to the input spectrum. Centre: Large H and small h. The model assigns low aleatoric uncertainty to the input spectrum, but is epistemically uncertain about its own prediction. Right: Small H and h. The model is certain about the spectrum and its prediction. where p∗ denotes the posterior, h the entropy of the predictive distribution, and ω the latent model parameters. Visually, this separation corresponds to the fact that an epistemic uncertainty translates to the flatness of the predicted simplex distribution, whereas aleatoric uncertainty is represented by a sharp but centred distribution (Malinin & Gales 2018; Gawlikowski et al. 2023). From this it follows that the Shannon entropy specifies aleatoric uncertainty, and any point estimator is by design unaware of any epistemic uncertainty because it assumes precise knowledge of the predictive distribution (Gawlikowski et al. 2023) (see Fig. 3 for the different behaviours in a three-class case).

The reference convolutional neural network (CNN) and MCD are trained with cross-entropy loss, the Dirichlet model uses negative log-likelihood loss, and the VI models maximize the evidence lower bound (ELBO). For all models, we used the Adam optimizer with weight-decay regularization (Loshchilov & Hutter 2017) and custom annealing learning rate scheduling plus early stopping based on validation loss. If prior knowledge about the true class distribution should be reflected in the predictions, the predictive posterior can be adjusted during post-processing using a second level of Bayesian inference to incorporate the desired class prior. 4.1. Baseline CNN

3.3. Proper scoring rules

Another rigorous formulation of calibrated uncertainties is provided by proper scoring rules (see Lakshminarayanan et al. (2017) and Gal (2016) for a definition and discussion). They are already used during the training in the form of the cross-entropy, and guarantee there that the point estimates are calibrated. The same principle can be applied to the continuous second-order distribution of probabilistic models. Two proper scores are the negative log-likelihood of the label in the joint distribution, and the mean Brier score: K 2 1 X NLLi = − log p∗iyi , Brieri = δyi =k − p∗ik . (7) K k Theoretically founded on the proper scoring rules, these two quantities are good measures for coherent and well-calibrated uncertainty.

4. Methodology In this section we present one point-estimate neural network and three probabilistic methods for the classification of astronomical spectra: (1) by fixing a closed-form (Dirichlet) distribution over the label space, (2) by using Monte Carlo dropout (MCD) on a neural network to produce stochastic outputs, and (3) by learning the distribution over model weights that propagate into a predictive distribution using Bayesian variational inference. As discussed in Sect. 3, we calculated different quantities to measure the prediction uncertainty. An overview of these methods and the flow of data to obtain uncertainties from their different output types is shown in Fig. 4. Article number, page 4

As a reference model, we present a CNN for the point-estimate classification of spectra. One of the goals was to keep the model small in terms of parameters, also in view of the possibility of building more computationally demanding models based on it. This CNN uses four blocks, where in each block a convolutional layer is followed by a batch normalization layer and a max-pooling layer to reduce the input signal (see Fig. 5). No dropout is used for the CNN. This model will serve as a baseline and the building block for other methods. It has ∼ 122, 000 trainable parameters, most of which are in the dense layer head. 4.2. Dirichlet distribution

Because the Dirichlet distribution is the conjugate prior of the categorical distribution, it is a natural distribution over point estimates on the (K − 1) simplex. Our Dirichlet model uses the baseline CNN, followed by a softplus activation to predict the concentration parameters of a K-dimensional Dirichlet distribution. The loss function during training is then the negative loglikelihood of the training labels in the distribution. For numerical reasons we offset these labels from the simplex corners, yi 7→ pik =

δk=yi + s , 1 + sK

(8)

and we chose the softening factor s = 10−4 . The point-estimate class membership probabilities y∗ can be retrieved using a Monte Carlo integration, which reduces to a simple sampling from the PDF. It is handy to embed the standard simplex into the K-dimensional Euclidean space such that it is spanned between the points {(IdK )i }, thus making the vector

Simon Barton et al.: Bayesian classification of astronomical spectra with class uncertainties

Model

Data

Data

Uncertainty

Output

CNN

Dirichlet

mean

Categorical

Total

Per class Calibration

Shannon entropy

N/A

ECE

Marginals

Total N/A

MC Dropout

marginalize

Per class Std Quantile Entropy

Variational Inference

Simplex Joint

Total

Per class

Joint entropy

Contours

Less information

More information

Fig. 4: Overview scheme for this study. The arrows indicate flow of data. components correspond to the label probabilities and ensuring the independence of classes. All other quantities can be computed similarly from these samples.

4.4. Variational inference

This is the only model where we were able to compute the joint entropy analytically, to represent epistemic uncertainty. However, we found this entropy not in good correlation with prediction success.

mω : xi 7→ p∗i (y)

Extending the notation above, we now consider a probabilistic model ,

(9)

where ω denotes the internal trainable weights of m. We are interested in finding the weights that explain the data best, or formally, the posterior p(ω | D). Directly applying Bayes’ rule, p(D | ω) p(ω) p(D)

,

4.3. Monte Carlo dropout

p(ω | D) =

As shown by Gal & Ghahramani (2016), the Bayesian predictive posterior p(y | x, ω) of a neural network can be numerically approximated by using a stochastic parameter dropout. This technique, known as Monte Carlo dropout (MCD), provides an easyto-implement and computationally inexpensive way to combine Bayesian inference with neural networks.

one finds the marginal likelihood p(D) to be intractable when ω is high-dimensional. Instead, we resort to finding a surrogate qθ (ω), parametrized by θ to approximate the posterior. In other words, we are looking for the value of θ that minimizes the Kullback-Leibler (KL) divergence to the true posterior, n o q∗ (ω) = argminθ KL(qθ (ω) || p(ω | D)) . (11)

We start with the same CNN architecture as the baseline model, insert dropout layers after each batch norm operation as shown in Fig. 5, and perform the training as usual. During inference we then keep sampling a new dropout mask for each layer and each x to obtain a distribution of predictions, represented by a collection of samples. From there the further processing is done as for the Dirichlet method. The success of the method can be sensitive to the choice of hyperparameters, notably the dropout rate rd and the parameter decay rate dw of the optimizer (Gal & Ghahramani 2016). In a 2D grid search (see Fig. 6) we observed a slight influence of the decay rate and a clear maximum of the accuracy for dropout rates between 8% and 18%. For further analysis, we selected the combination rd = 12.5%, dw = 0.01, which maximizes accuracy.

(10)

From here one can rewrite the unknown posterior in terms of the known joint, # " qθ (ω) p(D) KL (qθ (ω) || p(ω | D)) = Eω∼qθ log . (12) p(ω, D) In this form, the intractability has been moved into p(D), but because this term is not dependent on θ, the remaining expressions can be computed to form a loss function that is used in a gradient descent algorithm on θ. Once the model is trained, one can easily compute the predictive posterior from the parameter surrogate by applying Bayes’ rule again, p∗ (y | x, ω) =

p(y | ω) p(x, y | ω) p(x, ω)

.

(13) Article number, page 5

A&A proofs: manuscript no. aa56350-25

Linear(128→K) MC Dropout ReLU Linear(512→128) Flatten AdaptiveMaxPool(4) Block(c,p)

MC Dropout ReLU Linear(32→16)

(Adaptive*)MaxPool(p) MC Dropout

Block(128, *32)

ReLU BatchNorm

Block(64, 2)

Conv1D out channels: c kernel size: 5

Block(32, 5)

Block(16, 5)

Input

Fig. 5: Baseline CNN architecture, underlying all models. The hatched layers are enabled depending on the concrete model. The last block uses adaptive pooling to make the model compatible with different input spectra sizes.

This method for approximating the posterior by a surrogate is called variational inference (VI). In its stochastic version (SVI), it scales well and is therefore particularly useful in large-scale problems (Kingma et al. 2017; Goan & Fookes 2020; Gawlikowski et al. 2023). In practice, to design a BNN one chooses a functional model mω , a stochastic model that contains both priors p(ω) and p(y | x, ω) and, in the case of VI, the form of the parameter posterior p(ω | D). This choice is the equivalent of a loss function in pointestimate ML (Jospin et al. 2022). For the functional model we reused our baseline CNN. As a fixed form of qθ we chose the multivariate normal distribution with low-rank covariance. As done in Ong et al. (2018) we factorize the covariance matrix Σ = BBT + C 2

(14)

into a factor loading matrix B of size |ω| × f and a diagonal term C = diag(ci ). While the ci are the standard deviations of an independent multinormal surrogate, B allows a covariance structure of rank f , while keeping the total number of variational parameters low at |θ| = |ω|( f + 2). Although f = 1 would allow a closed form of the natural gradients (Tran et al. 2020), we used traditional backpropagation to experiment with f ∈ {0, 1, 2}, without computational limitations. The underlying neural network uses batch normalization (BN; Ioffe & Szegedy 2015), which maintains a running average of means and variances that is updated with each mini-batch during training. Although Mukhoti et al. (2020) derived that BN does not interfere with VI techniques, we found deteriorated performance and an increase in noise when using them in combination, and therefore we disabled all BN layers for the VI methods. As a parameter prior, we set an isotropic normal distribution with dimension-independent variance σ2 and trained the model both with fixed and variable σ. As pointed out in Murphy (2023), this procedure breaks the Bayesian idea, but it can be useful in real applications. Because the prior acts as a regularization, this construction allows the model to self-regularize. In fact, no overfitting is observed. To stay data-agnostic, we did not set a predictive prior p(y) for this work, but choices based on astronomical object spatial density, distance and signal strength, or catalogue labels are conceivable. For the update we used a stochastic scheme (SVI; Hoffman et al. 2013), where in each gradient descent step only one mini-batch is consumed for the likelihood approximation. To maintain balance within the ELBO, this requires the rescaling of the likelihood because it scales with the number of data points, while the KL terms scale with the number of model parameters. For the computation of the expectation values in Eq. 12 we used three MC steps per gradient descent step. We find this choice to have little effect on the final training state, but more effect on the noise level of the ELBO value. In each update we therefore sample three times from the parameter distribution and evaluate the ELBO on a random batch of size 128. We define an epoch as the number of updates equivalent to using the full set of training spectra: updates_per_epoch = ntrain /(n samples · batch_size).

5. Results

Fig. 6: Influence of the hyperparameters dropout rate and Adam weight decay on MCD accuracy (SDSS dataset).

Article number, page 6

In this section we compare the four methods in terms of their discriminative performance, as well as their ability to provide meaningful uncertainty estimates. For each model we performed three independent and unbiased training experiments that started with different randomly initialized weights. We did not handpick the presented training runs. All runs of all models used the

Simon Barton et al.: Bayesian classification of astronomical spectra with class uncertainties

Table 1: Results of all models used for the SDSS dataset. Model Baseline CNN Dirichlet MCD100 (p = 0.125) MCD1000 (p = 0.125) VI ( f = 0,σ = 0.3) VI ( f = 1,σ = 0.3) VI ( f = 2,σ = 0.3) VI ( f = 0,σ = variable) VI ( f = 1,σ = variable) VI ( f = 2,σ = variable)

Accuracy [%] 91.52 ± 0.02 87.9 ± 0.4 92.6 ± 0.1 92.58±0.07 91.35 ± 0.07 91.30 ± 0.10 91.15 ± 0.15 89.22 ± 0.22 89.47 ± 0.75 89.48 ± 0.26

Results (SDSS dataset) AUC ×1000 ECE [%] NLL 997.4 ± 0.2 988 ± 1 998.07 ± 0.01 997.98 ± 0.01 997.35 ± 0.04 997.27 ± 0.06 997.22 ± 0.03 995.83 ± 0.14 996.00 ± 0.44 996.07 ± 0.19

1.78±0.05 5.47±0.53 1.80±0.02 1.80±0.02 2.21±0.01 2.24±0.02 2.25±0.02 2.86±0.05 2.76±0.12 2.71±0.05

N/A 0.76 ± 0.07 0.210 ± 0.003 0.206 ± 0.009 0.230 ± 0.001 0.233 ± 0.003 0.236 ± 0.002 0.296 ± 0.005 0.288 ± 0.015 0.284 ± 0.006

Brier score ×1000 N/A 31 ± 3 10.34 ± 0.03 9.79 ± 0.02 11.34 ± 0.03 11.44 ± 0.11 11.50 ± 0.08 13.97 ± 0.22 13.54 ± 0.64 13.31 ± 0.25

Training time [s] 95 223 436 436 1365 2175 2135 1759 2318 2537

Inference time [ms] 0.012 0.029 0.852 8.479 0.609 0.617 0.616 0.620 0.620 0.617

Notes. Shown are the means and standard deviations of three independent training experiments with the same data splits, but different initialization weights. The highlighted values are the best in the column, within their uncertainty. p, f , and σ denote MC dropout rate, the rank of the VI covariance, and the standard deviation of its parameter prior, respectively. Training times are given until early stopping; inference times are given per spectrum. Both were taken on an Nvidia A5000 GPU.

0.2

0.4

0.6

0.8

1.0

24 58 108 2 4 G-AGN 2804 93% 1% 2% 4% 0% 0% 22 2886 76 13 3 G-SB 1% 96% 3% 0% 0% 220 2523 86 2 5 1 G-SF 163 5% 7% 84% 3% 0% 0% 0% 89 26 113 2727 1 37 3 1 2 1 G 3% 1% 4% 91% 0% 1% 0% 0% 0% 0% 11 1 2749 239 Q-BL 0% 0% 92% 8% 21 4 59 220 2686 2 2 1 2 1 2 Q 1% 0% 2% 7% 90% 0% 0% 0% 0% 0% 0% 1 3 2878 117 1 A0 0% 0% 96% 4% 0% 153 2826 1 20 F5 5% 94% 0% 1% 2763 142 89 6 F9 92% 5% 3% 0% 2 25 117 2855 1 G2 0% 1% 4% 95% 0% 1 82 2782 135 K1 0% 3% 93% 5% 8 177 2737 78 K3 0% 6% 91% 3% 1 1 119 2879 K5 0% 0% 4% 96%

G-AGN G-SB G-SF G Q-BL Q A0 F5 F9 G2 K1 K3 K5

Actual

0.0

Prediction

Fig. 7: Confusion matrix for a Monte Carlo dropout model on the SDSS dataset. The class names were abbreviated. The colours indicate percentages. The black lines separate the classes of different 4MOST coarse labels: star, quasar, galaxy. The percentages may not sum to exactly 100% due to rounding.

same split of data into training, validation, and test sets. All the calculated metrics are tabulated in Tables 1 and 2. All the following figures are based on evaluating the models on the unseen test set. The training histories of all four models on the 4MOST dataset are shown in Fig. C.1. While the CNN exhibits clear overfitting beyond epoch 30, the probabilistic models show significantly reduced overfitting; the gap between training and validation accuracy remains below 0.5 percentage points. Partic-

Fig. 8: Examples for some of the most common confusion pairs on the SDSS dataset. Left: Two highest S/N spectra. The label format is (actual → prediction). Right: 100 samples from the MCD predictive posterior, where all other classes were marginalized into ‘other’. The crosses show the expectation values. ularly, the MC dropout makes the validation accuracy even slightly exceed the training accuracy. All three runs converge in a similar manner for all models, which shows that the training is stable and converges reliably to the same posterior. The Dirichlet model displays noticeably larger accuracy fluctuations across epochs, despite the use of low learning rates. The training curves for the SDSS dataset are qualitatively identical, and are therefore omitted for brevity. The reference CNN achieves a 13-way classification accuracy of 91.5% on the SDSS dataset and a 10-way accuracy of 92.8% on the mock dataset. For comparison, the much larger Article number, page 7

A&A proofs: manuscript no. aa56350-25

0.8 0.6 0.4 0.2

1.0

ECE = 1.78% Mean confidence Output Fraction of samples

0.8 Accuracy

Accuracy

1.0

ECE = 6.03% Mean confidence Output Fraction of samples

0.8

0.6

Accuracy

1.0

0.4 0.2

0.0 0.0

0.2

0.4 0.6 Confidence

0.8

1.0

ECE = 2.26% Mean confidence Output Fraction of samples

0.6 0.4 0.2

0.0 0.0

0.2

0.4 0.6 Confidence

0.8

1.0

0.0 0.0

0.2

0.4 0.6 Confidence

0.8

1.0

Fig. 9: Calibration diagrams (also known as reliability diagrams) for three models: Dirichlet (left), MCD (centre), and VI with rank f = 1 and fixed prior (right). For each input (x, y) and each class c, the blue bars indicate how often c coincides with y, given the binned predicted probability p(y|x). The perfect values are shown in red, and take the binning effect into account. The ECE is the expected calibration error and measures the deviation between the blue bars and the red line, weighted by the respective prediction rate (black). A lower ECE is better. Blue bars above and below the red line correspond to under- and overconfidence of the model, respectively. These plots average all the probabilities and the error over all c. ResNet-like network from Zhong et al. (2024) achieves respective accuracies of 92.4% and 93.4% on the same datasets. The Dirichlet model underperforms in all metrics, which is probably due to its artificially enforced form of the posterior. The VI models show worse accuracy than the baseline CNN; allowing a dependence between the posteriors of individual weights (k > 0) shows no improvement, thus indicating that these models are not limited by their fixed form parameter posterior qθ . Instead, the limiting factor might be the choice of the prior. The MCD model with best-fitting hyperparameters outperforms the baseline CNN in terms of accuracy, even though probabilistic models are generally slightly inferior in this regard (Gal 2016; Gawlikowski et al. 2023). For all samplers we chose a validation MC sample size of 100. No performance gain is observed for higher sample numbers, but the inference time scales linearly with the number of samples. For reference, an MCD inference run with a sampling size of 1000 is included in Table 1.

0.975 0.950

Accuracy

0.925 0.900 0.875 0.850 0.825

CNN Dirichlet MCD VI

0.800 0.775

0

10

20

30

S/N

40

50

60

70

Fig. 10: Effect of S/N on accuracy for all models. The error bars indicate standard deviations from three independent training experiments (SDSS dataset). Article number, page 8

The confusion matrix for the SDSS test set (Fig. 7) summarizes the counts of the true versus the predicted labels across all classes. Overall, the model distinguishes well between galactic and extragalactic objects, with almost no confusion observed between these coarse classes. Within the stellar classes, about 80% of the misclassifications occur between adjacent spectral types. This is expected, as the Morgan–Keenan spectral classification scheme underlying the labels is based on a continuous temperature sequence with boundaries that do not correspond to distinct spectral features. This smooth gradation in stellar spectra naturally leads to overlap in feature space, particularly between neighbouring types. In contrast, the separation between stars and quasi-stellar objects (QSOs) is remarkably clean despite the historical terminology. Within the extragalactic classes, QSOs are well separated from other galaxies, with minimal cross-class confusion. However, their broad-line nature is misclassified in approximately seven percent of the cases, indicating some ambiguity in the identification of spectral line widths. Starburst (SB) galaxies are primarily confused with star-forming (SF) galaxies, which is expected given their similar emission-line features and overlapping star formation indicators. In contrast, the SF class exhibits substantial confusion with all other extragalactic classes, reflecting its broad spectral diversity and overlap with both AGN-hosting and quiescent galaxies. The confusion between AGN and normal galaxies lacks a clear physical interpretation. This may point to limitations in the spectral resolution or feature representation used by the model, or possibly to an intrinsic ambiguity in the labelling of weak or composite AGN spectra. A closer inspection reveals that many actual–prediction pairs are in fact confused due to low S/N (Fig. D.1). Example spectra of commonly misclassified extragalactic sources are shown in Fig. 8. The ternary distribution plots on the right give an impression of the models certainty. While there is high confidence in the wrong AGN predictions (rows two and four), labelled AGN are classified with significant uncertainty (rows one and three). We note that AGN and galaxies with starforming activity are often confused due to their physical mechanisms and overlapping definitions (Teimoorinia et al. 2024). In our evaluation, no model outperformed the baseline CNN in terms of probability calibration. The expected calibration er-

Simon Barton et al.: Bayesian classification of astronomical spectra with class uncertainties

Table 2: Results for the 4MOST mock dataset. Model

Accuracy [%] 92.80 ± 0.08 87.37 ± 0.09 93.94±0.07 93.95±0.08 93.95±0.05 91.71 ± 0.14 91.73 ± 0.38 91.43 ± 0.17 89.14 ± 0.77 88.74 ± 0.55 89.55 ± 0.85

Baseline CNN Dirichlet MCD50 (p = 0.125) MCD100 (p = 0.125) MCD200 (p = 0.125) VI ( f = 0,σ = 0.3) VI ( f = 1,σ = 0.3) VI ( f = 2,σ = 0.3) VI ( f = 0,σ = variable) VI ( f = 1,σ = variable) VI ( f = 2,σ = variable)

Results (4MOST mock dataset) AUC ×1000 ECE [%] NLL 995.15 ± 0.04 986.95 ± 0.72 996.25 ± 0.02 996.25 ± 0.02 996.25 ± 0.02 994.71 ± 0.17 994.72 ± 0.12 994.50 ± 0.11 992.37 ± 0.81 992.02 ± 0.82 992.45 ± 0.62

1.82±0.01 6.01±0.16 1.87±0.01 1.87±0.01 1.87±0.01 2.60±0.05 2.59±0.07 2.69±0.01 3.48±0.30 3.67±0.34 3.46±0.26

N/A 0.620 ± 0.013 0.142 ± 0.001 0.142 ± 0.001 0.142 ± 0.001 0.192 ± 0.004 0.191 ± 0.005 0.199 ± 0.002 0.260 ± 0.020 0.272 ± 0.024 0.258 ± 0.018

Brier score ×1000 N/A 38.71 ± 0.99 9.59 ± 0.05 9.60 ± 0.05 9.60 ± 0.05 13.26 ± 0.21 13.15 ± 0.36 13.67 ± 0.11 16.95 ± 1.23 17.86 ± 1.51 16.81 ± 0.98

Training time [s] 229 968 576 576 576 1124 1422 1436 665 988 959

Inference time [ms] 0.015 0.037 0.581 1.168 2.336 0.799 0.807 0.804 0.810 0.809 0.805

Notes. All values, as in Table 1.

0.6

CNN MCD VI

S/N

1.0 -0.21 -0.13 0.1

Shannon -0.21

1.0

0.62 -0.47

0.4

Brier -0.13 0.62

1.0 -0.89

0.3

Acc.

0.1 0.0

0

20

40

S/N

60

80

Acc.

Brier

0.2

Shannon

0.1 -0.47 -0.89 1.0

S/N

Shannon Entropy

0.5

100

Fig. 11: Shannon entropy vs S/N on the mock dataset. The Dirichlet model predictions have much higher entropy. The correlation matrix shows the relation of S/N (cause), accuracy (effect), Brier score (measure), and Shannon entropy (indicating uncertainty) of MCD predictions.

ror (ECE) remains low (≤ 3%) across all models, with the exception of the Dirichlet model, which exhibits noticeably worse calibration. Figure 9 displays three representative reliability diagrams. The subpar calibration performance of the Dirichlet model is likely attributable to its limited degrees of freedom, which constrain its ability to capture complex predictive uncertainty. This limitation could potentially be addressed by replacing the Dirichlet distribution with a more expressive alternative, such as a normalizing flow (Rezende & Mohamed 2015). The performance of classifiers is expected to depend on the signal quality. Figure 10 shows how the accuracy of all models correlates positively with S/N and levels out above S/N ≈ 40. For these low-noise spectra, MCD reaches accuracies of over 96%. The lowest performance is obtained at S/N ≈ 5 and increases again for high-noise signals S/N ≤ 5, due to the reduced data variety in accordance with the Gini impurity (cf. Fig. 2). This behaviour is observed for all four models, while their ordering

in accuracy is consistent across signal strengths. Comparison to the distribution of S/N in the dataset suggests that the total accuracy is limited by the noise in the extragalactic targets. Application within 4MOST may therefore yield higher (or lower) values, while MCD is expected to consistently perform best on other data. To provide an estimate of how well the algorithms work on an unbiased sample of the sky (rather than on a class-balanced test set), we first estimated the ratio of galactic to extragalactic sources from the photometric SDSS table (PhotoOb jAll) to be 1.37, and we then used the subclass counts from the spectroscopic table (S pecOb j) within each coarse class. The resulting class ratios are dominated by galaxies (27%) and F stars (31%), while AGN and starburst galaxies are least abundant (together 1.2%). Using these values as weights to the per-class accuracy of the trained MCD models results in an adjusted total accuracy of 92.0% ± 0.3%. This value is slightly lower than on the balanced test set; galactic sources have a positive impact on the difference, while extragalactic contribute negatively. The NLL and Brier score are not themselves measures of uncertainty, but are indicators of the calibration of uncertainty, given that the true label is known (Lakshminarayanan et al. 2017; Gal 2016). Both metrics reward high-confidence correct predictions and heavily penalize confident errors. Their strong threshold-like separation (Fig. E.1) between correct and incorrect predictions for the Bayesian methods reflects how the two metrics respond to the confidence encoded in the predictive distribution, and indicates that the models’ predicted probabilities are well aligned with the actual outcomes, such that confidence is a reliable proxy for correctness. The strong similarity of the distributions of MCD and VI suggests that both models are learning the same posterior, while MCD is better able to resolve an offGaussian PDF. The fact that the distribution looks so different for the Dirichlet model may be another hint of the lack of the model’s flexibility, as mentioned above. A statistical statement on the epistemic uncertainties, delivered by the probabilistic models is made in Fig. 11. While the drops in prediction entropy at S /N ≈ 14 and S /N ≈ 65 can again be attributed to the low data entropy (Fig. A.2), we verified that the uncertainty correlates positively with the Brier score, but negatively with accuracy and signal quality, as confidence in incorrect predictions and predictions on high-noise signals is expected to decrease. Article number, page 9

A&A proofs: manuscript no. aa56350-25

0.0 2948 96%

0.8

1.0

118 4% 3022 100% 1 2958 0% 99%

ESN GalDiskLR 118 4%

20 1%

23 1%

1 0%

2855 96%

1 0% 2 0%

WAVES RedGAL

MCsn

GalDiskLR

ESN

GalHR

Dyn

tides_host

tides_host

ClusB

RedGAL

3 0% 2927 15 78 97% 1% 3% 1 27 2711 49 150 0% 1% 92% 2% 5% 908 232 1830 1 30% 8% 62% 0% 17 2936 1% 99%

WAVES

2436 100%

COSMO_AGN

ClusB

Actual

0.6

3008 100%

GalHR

MCsn

0.4

COSMO_AGN

Dyn

0.2

rors that come with using misclassified objects. This way, a classifier with uncertainties could even improve the effective completeness of a 4MOST survey catalogue by correcting its labels and enabling the controlled inclusion of objects that would otherwise be excluded by conservative selection cuts, while keeping the class contamination quantifiable. The related chosen uncertainty thresholds are survey and task specific, and are dependent on the object abundance, the desired number of usable objects, and the confused classes themselves. On the contrary, using a neural network to predict a Dirichlet distribution to describe the posterior is clearly affected by the rigid form of the distribution. Although easy to adapt, this approach yields a poor accuracy and a poorly calibrated uncertainty. Despite their more principled Bayesian approach, the models trained with variational inference were found to be inferior in all metrics. Allowing a low-rank covariance had little effect on these results. This may be due to a bad choice of prior. Learning a diagonal normal prior during the training could not overcome these problems. Acknowledgements. We thank Thorsten Glüsenkamp and Anish Amarsi for useful comments and fruitful discussions, and Fucheng Zhong for providing training data. We would like to acknowledge financial support from the eSSENCE graduate school for data intensive science. A.K. and M.S. were supported by the Swedish National Space Agency (SNSA).

Prediction Fig. 12: Class confusion of the MCD model for the 4MOST mock dataset, as in Fig. 7. The black lines separate blocks of extragalactic and galactic classes. All the discussed differences between the models directly translate to the 4MOST mock dataset. Again, the MCD variants prevail as the models with highest accuracy and best uncertainties, independently of the number of used evaluation samples. The confusion matrix of the MCD model on the mock dataset (Fig. 12) shows very high accuracy for all galactic classes and AGN, probably due to the specificness of these classes within the class set. The 30% misclassification of red galaxies as bright clusters can be explained by the cross-contamination of these classes in the training data (Zhong et al. 2024). Finally, a set of one-versus-rest ROC curves is shown in Fig. B.1.

6. Conclusions We have shown that Bayesian approximate inference with Monte Carlo dropout is able to keep up with and even exceed the reference neural network in its predictive performance of the Kclassification of astronomical spectra. In addition, the model provides well-calibrated predictive posteriors on the probability simplex, while the associated increase in inference time for 100 samples, relative to a standard CNN, appears to be a negligible investment. Such a model could be suitable for the 4MOST classification pipeline, potentially applied to a more powerful deep learning architecture. In this set-up condensed class-wise uncertainties could be easily obtained by marginalizing the predictive posterior for each class and compute its standard deviation (assuming normal marginal), its two-sided percentiles, or its differential entropy. These class uncertainties could then be used to guide object selection and quality assessment in downstream tasks. Especially when a large number of objects makes manual verification unfeasible, Bayesian uncertainties allow the quantification of erArticle number, page 10

References Caelen, O. 2017, ANN MATH ARTIF INTEL, 81, 429 Chlaily, S., Ratha, D., Lozou, P., & Marinoni, A. 2023, IEEE Transactions on Signal Processing, 71, 3710 De Jong, R. S., Agertz, O., Berbel, A. A., et al. 2019, arXiv preprint arXiv:1903.02464 Depeweg, S., Hernandez-Lobato, J.-M., Doshi-Velez, F., & Udluft, S. 2018, in ICML, PMLR, 1184–1193 Gal, Y. 2016, Uncertainty in deep learning (phd thesis, University of Cambridge) Gal, Y. & Ghahramani, Z. 2016, in Proceedings of Machine Learning Research, Vol. 48, Proceedings of The 33rd ICML, ed. M. F. Balcan & K. Q. Weinberger (New York, New York, USA: PMLR), 1050–1059 Gawlikowski, J., Tassi, C. R. N., Ali, M., et al. 2023, AIR, 56, 1513 Goan, E. & Fookes, C. 2020, Case Studies in Applied Bayesian Data Science: CIRM Jean-Morlet Chair, Fall 2018, 45 Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. 2017, in ICML, PMLR, 1321– 1330 Hoffman, M. D., Blei, D. M., Wang, C., & Paisley, J. 2013, JMLR, 14, 1303 Hofman, P., Sale, Y., & Hüllermeier, E. 2024, arXiv preprint arXiv:2404.12215 Hortúa, H. J., Volpi, R., Marinelli, D., & Malagò, L. 2020, PRD, 102, 103509 Ioffe, S. & Szegedy, C. 2015, in ICML, pmlr, 448–456 Jönsson, H., Holtzman, J. A., Prieto, C. A., et al. 2020, AJ, 160, 120 Jospin, L. V., Laga, H., Boussaid, F., Buntine, W., & Bennamoun, M. 2022, IEEE Computational Intelligence Magazine, 17, 29 Kendall, A. & Gal, Y. 2017, NeurIPS, 30 Kingma, D. P. et al. 2017, Variational inference & deep learning: A new synthesis Kranzlein, M., Liu, N. F., & Schneider, N. 2021, arXiv preprint arXiv:2109.07494 Lakshminarayanan, B., Pritzel, A., & Blundell, C. 2017, NeurIPS, 30 Lin, Y.-C. & Wu, J.-H. P. 2021, PRD, 103, 063034 Loshchilov, I. & Hutter, F. 2017, arXiv preprint arXiv:1711.05101 Malinin, A. & Gales, M. 2018, NeurIPS, 31 Mancarella, M., Kennedy, J., Bose, B., & Lombriser, L. 2022, PRD, 105, 023531 Möller, A. & de Boissière, T. 2020, MNRAS, 491, 4277 Mukhoti, J., Dokania, P. K., Torr, P. H., & Gal, Y. 2020, arXiv preprint arXiv:2012.13220 Murphy, K. P. 2023, Probabilistic machine learning: Advanced topics (MIT press) Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., & Tran, D. 2019, in CVPR workshops No. 7 Ong, V. M.-H., Nott, D. J., & Smith, M. S. 2018, JCGS, 27, 465 Paszke, A., Gross, S., Massa, F., et al. 2019, NeurIPS, 32 Rezende, D. & Mohamed, S. 2015, in ICML, PMLR, 1530–1538 Silva Filho, T., Song, H., Perello-Nieto, M., et al. 2023, Mach. Learn., 112, 3211 Teimoorinia, H., Shishehchi, S., Archinuk, F., et al. 2024, arXiv preprint arXiv:2407.12151 Thummel, L., Bose, B., Pourtsidou, A., & Lombriser, L. 2024, MNRAS, 535, 3141 Tran, M.-N., Nguyen, N., Nott, D., & Kohn, R. 2020, JCGS, 29, 97 Vranken, J. F., van de Leur, R. R., Gupta, D. K., et al. 2021, EHJ-DH, 2, 401 Zhong, F., Napolitano, N. R., Heneka, C., et al. 2024, MNRAS, 532, 643 Zhou, X., Li, N., Zou, H., et al. 2025, MNRAS, 536, 2260

Simon Barton et al.: Bayesian classification of astronomical spectra with class uncertainties

1.00

Appendix A: 4MOST mock dataset

Flux Flux

0.98

0.0

True Positive Rate

Dyn (S/N=408)

0.2 GalHR (S/N=67)

0.0

Flux

Flux

Flux

0.2

ESN (S/N=17)

2

0.40 0.2 0.0 0.8

0.94

0.92

COSMO_AGN (S/N=106)

0.4 0.2 0.0 4000

WAVES (S/N=31) 5000

0.90 0.00

6000 7000 Wavelength [Å]

8000

9000

Fig. A.1: Selected example spectra of the mock dataset with the highest S/N of the respective class.

0.6

1 - Gini impurity

0.4 0.2 0.40

Dyn GalHR GalDiskLR MCsn ESN

0.35 0.30 0.25 Density

(999.6) K5 (997.4) K3 (997.7) K1 (999.0) G2 (997.7) F9 (999.0) F5 (999.4) A0 (996.5) Q (997.9) Q-BL (996.6) G (994.5) G-SF (998.9) G-SB (997.7) G-AGN

0.96

COSMO_AGN ClusB WAVES RedGAL tides_host

0.02

0.04 0.06 False Positive Rate

as negative. This one-versus-rest approach isolates the model’s ability to distinguish each class from the remainder, independent of overall classification performance. All models exhibit high area under the curve (AUC) values (> 0.99), indicating excellent discriminative ability. Although overall accuracy is not perfect, the consistently high AUCs suggest that the models rank instances correctly with high confidence, even when the final, threshold-based prediction is incorrect. In many such cases, the true class still receives a relatively high predicted probability. This implies that the models capture meaningful structure in the data and effectively encode class relationships in their output distributions. In particular, a good ranking accounts for systematic label-related misclassifications, such as those between adjacent stellar spectral types.

Appendix C: Learning curves

0.15

Appendix D: S/N confusion

0.05 0.00

0

10

20

30

40 S/N

50

60

70

80

0.10

Fig. B.1: Receiver operator curve (ROC) of the MCD model on the SDSS dataset. In brackets AUC × 1000.

0.20 0.10

0.08

Figure D.1 summarizes the highest S/N for each misclassification pair in both datasets. These plots reveals to which extend certain classes are confused due to bad signal in contrast to other reasons. Especially the few off-block cases are thus completely explained by high noise. The dominant inter-block confusions, such as RedGAL → ClusB, on the other hand, may not be seen as caused by noise, given the generally low mean S/N of 8.7 for red galaxies.

Fig. A.2: S/N distribution (bottom) and resulting expected random-guessing accuracy (top) as in Fig. 2 for the mock dataset.

Appendix E: Uncertainty distributions

Appendix B: Receiver operating characteristic

Histograms of discussed uncertainty metrics are compared in Fig. E.1. All metrics correlate with prediction correctness (green vs red).

In the multi-class setting, each ROC curve (Fig. B.1) is computed by treating one class as the positive case and all the others Article number, page 11

A&A proofs: manuscript no. aa56350-25

MCD100 (p = 0.125)

0.95

0.95

0.90

0.90 Accuracy

Accuracy

Baseline CNN

0.85 Training 1 Validation 1 Training 2 Validation 2 Training 3 Validation 3

0.80 0.75 0.70

0

10

20

Epochs

30

40

0.85 Training 1 Validation 1 Training 2 Validation 2 Training 3 Validation 3

0.80 0.75 0.70

50

0

20

0.95

0.90

0.90

0.85 Training 1 Validation 1 Training 2 Validation 2 Training 3 Validation 3

0.70

0

20

40

Epochs

60

80

100

Accuracy

Accuracy

0.95

0.75

Epochs

60

80

100

VI (f = 1, = 0.3)

Dirichlet

0.80

40

0.85 Training 1 Validation 1 Training 2 Validation 2 Training 3 Validation 3

0.80 0.75 0.70

0

20

40

Epochs

60

80

100

Fig. C.1: Training curves for all models, showing the accuracies as a function of the number of trained epochs, for the three training runs on the 4MOST mock dataset. Loss functions are omitted for clarity. A strong separation between training and validation curves would indicate overfitting. A difference between runs would indicate unstable convergence and local loss minima. The stars indicate the training state with the lowest validation loss that were used for the results above.

Article number, page 12

Simon Barton et al.: Bayesian classification of astronomical spectra with class uncertainties

Prediction

402 97%

0.6

0.8

1.0

268 3% 67 100%

GalHR

199 6%

3 1%

0 1%

5 0%

4 0% 21 2 95% 1% 1 2 9 0% 1% 92% 19 2 29% 8% 7 1%

21 4% 5 3% 98 63% 9 0%

32 5% 5 0% 204 99%

RedGAL

850 98%

GalDiskLR

WAVES

17 100%

ESN

367 94% 66 100%

COSMO_AGN ClusB 1 0%

WAVES RedGAL

MCsn

tides_host

tides_host

MCsn

0.4

ClusB

Dyn

0.2

GalDiskLR

60 24 33 39 7 4 G-AGN 93% 1% 2% 4% 0% 0% 18 60 23 6 2 G-SB 1% 96% 3% 0% 0% 41 51 51 38 6 13 22 G-SF 5% 7% 84% 3% 0% 0% 0% 45 16 24 47 9 46 39 32 8 10 G 3% 1% 4% 91% 0% 1% 0% 0% 0% 0% 20 2 56 28 Q-BL 0% 0% 92% 8% 13 33 22 5 28 7 6 4 8 Q 1%6 0%2 2% 7% 90% 0% 0% 0% 0% 0% 0% 3 13 100 65 7 A0 0% 0% 96% 4% 0% 55 112 8 76 F5 5% 94% 0% 1% 129 96 70 32 F9 92% 5% 3% 0% 4 54 97 115 6 G2 0% 1% 4% 95% 0% 6 78 102 74 K1 0% 3% 93% 5% 41 66 103 79 K3 0% 6% 91% 3% 4 6 73 110 K5 0% 0% 4% 96%

0.0

COSMO_AGN

1.0

ESN

0.8

GalHR

0.6

Dyn

0.4

Actual

0.2

G-AGN G-SB G-SF G Q-BL Q A0 F5 F9 G2 K1 K3 K5

Actual

0.0

Prediction

Fig. D.1: Confusion matrix. Top labels are the maximum S/N value for every actual–prediction pair. The bottom labels and colouring are as in Fig. 7. Left: SDSS dataset. Right: 4MOST mock dataset.

Article number, page 13

A&A proofs: manuscript no. aa56350-25

Correct Incorrect CNN

N/A

N/A

Correct Incorrect

Correct Incorrect

Correct Incorrect

Dirichlet

Correct Incorrect

N/A

Correct Incorrect

Correct Incorrect

Correct Incorrect

Correct Incorrect

Correct Incorrect

Correct Incorrect

Correct Incorrect

VI

MCD

Correct Incorrect

0.00

0.05

0.10 Mean std

0.15

0.0

0.5

1.0 1.5 2.0 Shannon entropy

2.5

0

1

2 NLL

3

4

0.00

0.05 0.10 Brier score

0.15

Fig. E.1: Comparison of different metrics of uncertainty on the SDSS dataset. Abscissa and ordinate labels apply to full columns and rows, respectively. Correct and incorrect histogram bars are independently normalized to one.

Article number, page 14

Record · ID 1006873 · SHA-256 8ae8e7cc11bb1220
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.