Astronomy & Astrophysics manuscript no. main July 7, 2026
©ESO 2026
Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification Raphaël Bonnet-Guerrini1, 2 , Bruno Sanchez2 , Dominique Fouchez2 , Benjamin Racine2 , Maya Guy3, 4 , Mariam Sabalbal5 , Manal Yassine2 , and Vincenzo Piuri1
arXiv:2607.05393v1 [astro-ph.IM] 6 Jul 2026
1
Università degli Studi di Milano, Department of Computer Science, Milan 20133, Italy e-mail: [email protected] ⋆ 2 Aix Marseille Univ, CNRS/IN2P3, CPPM, Marseille, France 3 Université Côte d’Azur, INRIA, CNRS, Laboratoire J.A.Dieudonné, Maasai team, Nice 06000, France 4 Université Côte d’Azur, Observatoire de la Côte d’Azur, CNRS, Laboratoire Lagrange, Nice 06000 France 5 Université de Liège, STAR Institute, Liège 4000, Belgium
ABSTRACT Context. Time-domain surveys generate large numbers of transient candidates, making Real-Bogus classification a critical step in
automated discovery pipelines. Obtaining reliable human-labeled training sets is costly, while community-provided labels can be noisy and survey-dependent. Aims. We aim to develop a Real-Bogus classification framework that can be trained without human-labeled data using physically motivated injected transients and bogus-dominated survey data, that remains robust under strong class contamination, and that provides well-calibrated and reliable uncertainty quantification suitable for downstream decision-making. Methods. We combine simulated transient injections with a heavily contaminated survey class and train a dual-network model using an asymmetric co-teaching strategy for classes with different label-noise levels. We evaluate performance on a labeled benchmark subset and analyze the learned representation with latent-space visualization tools. For uncertainty quantification (UQ), we compare standard approaches (MC dropout and deep ensembles) and propose a low-cost hybrid strategy that exploits the dual-network setting to improve calibration. We further extend the evaluation to the light-curve domain to assess recovery of light-curve classes. Results. The method achieves strong Real-Bogus performance on the labeled subset and remains stable under severe class contamination. It recovers transient light-curve classes with high fidelity, while single-source identification is fundamentally limited by intrinsic ambiguity in light-curve-derived labels. Our hybrid UQ approach achieves competitive calibration relative to substantially more expensive ensemble baselines. Latent-space analyses indicate that uncertainty aligns with the decision boundary and reveal that the model is able to identify subclasses within the bogus population. Conclusions. Our results show that injection-driven, weakly supervised training can enable scalable and consistent Real-Bogus classification without human-labeled training data while providing calibrated uncertainties. The methodology is well suited for transfer to forthcoming surveys by re-running the injection-based training pipeline. Key words. Astronomical instrumentation, methods and techniques - Methods: data analysis - Methods: statistical
1. Introduction The new Vera C. Rubin Legacy Survey of Space and Time (LSST) (Abell et al. 2009; Ivezić et al. 2019) features a time-domain component that will detect transients throughout its 10-year run using Difference Image Analysis (DIA) LSST Science Collaboration et al. (2009); Ivezić et al. (2019); Liu et al. (2024). In DIA, a template image is subtracted from a newly observed image to identify changes associated with new or variable sources (Alard & Lupton 1998; Alard 2000). This method is crucial for detecting transient events in crowded fields or under varying observational conditions. Typically, supernovae occupy the short-duration region of the luminosity-timescale phase space and therefore benefit from high-cadence observations that can identify them before they fade (Kulkarni & Kasliwal 2009). Because DIA depends on accurate image alignment and point-spread-function matching, it also produces false positives ("bogus") arising from noise, arti⋆
Corresponding author
facts, imperfect subtraction, cosmic rays, bad pixels, and atmospheric effects. Illustrated examples of spurious and successful DIA are shown in Fig. 1. In this work, we use the term source to denote a single-epoch DIA detection, and the term object to denote the association of sources across epochs at a consistent sky position; the magnitudes of the sources associated with a given object across time form a light curve. While light curves provide the richest information for transient typing, they are typically only informative after multiple detections have been accumulated. Relying solely on light-curve-based methods typically implies that identification occurs only after, or at best late in, the visible window of the transient. However, many downstream science goals require rapid spectroscopic follow-up (Matheson et al. 2013; Najita et al. 2016). At LSST-scale, alert brokers and downstream transient pipelines require early transient-bogus discrimination to suppress spurious detections and prioritize candidates before applying more computationally intensive light-curve classifiers Article number, page 1
A&A proofs: manuscript no. main
(Narayan et al. 2018; Möller et al. 2020; Carrasco-Davis et al. 2021; Vujčić et al. 2025). Historically, distinguishing real transients from bogus detections has relied on a combination of algorithmic quality flags and manual human inspection. With the rate of detection expected from the Vera C. Rubin Observatory, manual inspection is not feasible (Graham et al. 2024). In recent years, machine learning has demonstrated remarkable success in various domains, including image classification. Traditional image classification relied heavily on hand-crafted features and domain expertise. However, with the advent of Deep Learning (DL), convolutional neural networks (CNNs) (LeCun et al. 1989) have become the standard for image classification tasks (Hinton et al. 2012), including astronomical photometry applications (Dieleman et al. 2015; Lanusse et al. 2017). CNNs automatically learn hierarchical features from raw pixel data, significantly improving accuracy and reducing the need for manual feature engineering. Following these advances, DL methods have been developed for real-bogus classification (Sedaghat & Mahabal 2018; Reyes et al. 2018; Carrasco-Davis et al. 2021; Hosenie et al. 2021; Killestein et al. 2021; Goode et al. 2022; Makhlouf et al. 2022; Chen et al. 2023). Despite these advancements, we identify two main challenges for deploying DL-based methods for real-bogus classification. Challenge (i) DL often requires large labeled training sets,
which can be expensive and time-consuming to obtain. Realbogus classification typically suffers from lack of labeled data, most methods in the literature therefore rely on supervised learning with human-labeled datasets (Sedaghat & Mahabal 2018; Reyes et al. 2018; Carrasco-Davis et al. 2021; Hosenie et al. 2021; Goode et al. 2022; Makhlouf et al. 2022; Chen et al. 2023). While transient simulation through artificial source injection is well established (Kessler et al. 2015; Sánchez et al. 2022), bogus artifacts are heterogeneous, survey-specific, and difficult to model realistically. In addition, cross-survey domain shift has been observed for real-bogus classification, implying that methods trained for one telescope do not necessarily transfer directly to another (Cabrera-Vives et al. 2023). To meet the labeling demands of DL in the large-survey era, astronomy has increasingly relied on collaborative labeling. While successful, these efforts still incur significant overhead and produce labels with nonnegligible inter-annotator disagreement, requiring careful aggregation and bias correction (Lintott et al. 2010). The use of simulated data has been explored in Killestein et al. (2021) for point spread functions (PSFs) of minor planets superimposed on galaxy images. To date, unsupervised approaches have been tested (Mong et al. 2022), as have active-learning methods that reduce the amount of labeled data required to achieve strong performance (Liu et al. 2025). Nevertheless, real-bogus classification remains largely constrained by the need for survey-specific human-labeled training data. To overcome challenge (i), we propose a human-label-free approach based on the following intuition: the unlabeled survey detections are dominated by bogus examples, while realistic transient examples can be generated through source injection. Transient appearances can be simulated realistically at the image level through injections. By treating survey detections as a noisy bogus class and injected examples as transients during training, we formulate the task as a weakly supervised learning (WSL) problem with class-dependent label noise. Building Article number, page 2
on existing WSL methods, we introduce Asym-Co-teaching, a variant of co-teaching designed for asymmetric, class-dependent noise. To evaluate the method, we construct training sets with different contamination levels and compare the performance of competing approaches across these settings. Challenge (ii) DL models are often treated as black boxes
because their decisions arise from many nonlinear transformations over high-dimensional internal representations (Lipton 2016). Most explainable AI (XAI) methods explain a prediction by using a simplified proxy (or attribution) mechanism around a specific input (Ribeiro et al. 2016; Shrikumar et al. 2017; Binder et al. 2016; Lundberg & Lee 2017). Such techniques have been explored for real-bogus transient classification (Reyes et al. 2018), but they are inherently local: they provide insight for a single image and struggle to explain the global behavior of the model. Recently, mechanistic interpretability has emerged as an approach for identifying global structures in the computations learned by neural networks (Elhage et al. 2021; Cammarata et al. 2020; Sharkey et al. 2025). Early work often attempted to associate individual neurons, filters, or layers with interpretable behaviors (Zeiler & Fergus 2013; Olah et al. 2017), but the recognition of polysemanticity (Elhage et al. 2022) has increasingly shifted attention toward the geometry of latent representations. In this paper, we use latent space to denote the internal feature representation learned by an intermediate layer of the model. Beyond interpretability, standard deterministic deep learning models usually return point estimates and do not, by default, provide well-calibrated predictive uncertainty. Yet scientific inference requires uncertainty quantification (UQ) that can be propagated to downstream analyses. As highlighted in (DESC et al. 2026), adapting UQ methods to astronomy-specific settings is therefore essential for trustworthy scientific analyses. This is especially critical for cosmology-oriented transients targeted in this study (e.g., Type Ia supernovae and other rare events used for population-level inference), where real-bogus classification occurs early in the pipeline: missed detections and, more importantly, systematic biases in detection or classification efficiency as a function of brightness, host-galaxy properties, or redshift directly affect the survey selection function and can propagate to bias cosmological measurements. To address challenge (ii), we probe the model through latent-space analysis using dimensionality-reduction and visualization tools to seek global structure in the learned representation. Because interpretability analyses are primarily qualitative, we complement them with a systematic evaluation of uncertainty quantification (UQ) methods. This comparison is motivated by two considerations: first, epistemic uncertainty provides a natural representation of lack of knowledge in scientific inference, especially when ground truth is incomplete; second, uncertainty estimates in deep learning are often fragile, potentially miscalibrated, and sensitive to distribution shift. We therefore assess their reliability empirically in our setting and adapt existing UQ methods to the specific structure of Asym-Co-teaching. In this paper, we address these two challenges as follows. Section 2 introduces the dataset, the injection procedure, and the construction of the ground truth. Section 3 then presents the Real-Bogus classification task, the model architecture, and the training and optimization strategy. Next, we describe in Section 4 our experimental methodology for evaluating WSL approaches including the Co-teaching framework and Asym-
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
Co-teaching, our adaptation of Co-teaching to class-dependent asymmetric noise. In Section 5, we propose a methodology to assess UQ in this setting and adapt existing UQ methods to AsymCo-teaching. We present the latent space exploration and visualization in Section 6. The results are presented in Section 7, including UQ and WSL experiments as well as an extension to light-curve classification and inference. We discuss the implications and limitations of our findings in Section 8, and conclude in Section 9.
2. Data and injection process We use imaging data from the Hyper Suprime-Cam (HSC) on the 8.2 m Subaru Telescope, located on Maunakea, Hawai‘i (Miyazaki et al. 2017). HSC comprises 104 science CCDs of size 2048 × 4096 pixels with a pixel scale of 0.168 arcsec pix−1 (Miyazaki et al. 2017). For this study, we use two HSC-based datasets processed with the LSST Science Pipelines (Sánchez et al. 2022; Liu et al. 2024), whose DIA implementation follows the image-subtraction framework of Alard & Lupton (1998). To keep end-to-end processing computationally tractable during method development, we use rc2_subset as the main training dataset. This is a subset of the HSC SSP survey, used for regular testing of the LSST Data Release Production, and consists of six central science CCDs in the COSMOS field for eight visits in each of the grizy filters (40 visits in total). For inference and light-curve follow-up, we use a larger HSC-UDEEP COSMOS subset containing 103 science CCDs over 862 visits. All images are processed through the LSST Science Pipelines, including difference image analysis (DIA), source detection, and measurement. Template
Difference
(b) Bogus detection
(a) Transient candidate
Science
Fig. 1: Example of difference imaging analysis (DIA) stamps. Each triplet shows the Science image (left), the Template image (middle), and the Difference (Science−Template=Difference) (right). Top: a real transient produces a compact, PSF-like positive residual in the subtraction. Bottom: a bogus detection exhibiting a structured residual (dipole-like), characteristic of subtraction artifacts.
2.1. Difference Image Analysis
The first step of difference-image analysis (DIA) is to construct a deep template (reference) image by coadding multiple exposures of the same field; the template represents the static sky for that field. A science image (single-epoch exposure) is a new observation of the same field acquired at a later epoch. The science
image is aligned photometrically and astrometrically to the template, after which PSF matching and subtraction are performed. Ideally, the resulting difference image contains only sources that have changed between the two epochs (Fig. 1). In practice, most detections on difference images are spurious. Common sources of bogus detections include imperfect astrometric registration, PSF or photometric mismatches, variable atmospheric conditions, detector defects, cosmic rays, and subtraction artifacts. Robust Real-Bogus filtering is therefore essential to suppress these artifacts before downstream transient characterization. From the resulting difference images, candidate point-like detections are identified and represented as 30×30 pixel cutouts. Each cutout is normalized independently by clipping pixel values to the 1st-99th percentile range, applying an arcsinh transform, and then performing robust standardization using the cutout median and median absolute deviation. These cutouts form the input to our real-bogus classifier. 2.2. Supernova-Like injection
Many injection pipelines generate source catalogs by sampling positions and magnitudes from generic or ad hoc distributions. In this work, we instead adopt an injection scheme tailored to supernova-like point sources. This choice reflects the scope of our analysis: our primary downstream application is the detection of SNe Ia-like transients, and we therefore design the injection population to approximate the region of astrophysical parameter space most relevant to that use case. The resulting injection model combines host-galaxy information with simple priors on source location and brightness. Host galaxies are identified using the LSST pipeline extendedness criterion, a binary flag (0 = point-like, 1 = extended) based on a threshold on the ratio between PSF flux and model flux. For each selected galaxy we estimate the size and orientation of its projected elliptical profile, and use these quantities to define the injection-position prior. The projected distance of the injection from the galaxy center is sampled from a normal distribution centered on the host position, with variance set by the semi-major axis of the galaxy’s elliptical shape. Mathematically, the offset dinj is sampled as dinj ∼ N(µ = 0, σ2 = a2 ), where a is the semi-major axis of the galaxy. This produces a host-associated population of synthetic point sources whose projected offsets follow the size and orientation of the underlying galaxy. We then sample the injected magnitude as minj ∼ U(mhost − 1, mhost + 3), which is intended to span a broad range of source-to-host contrast ratios while remaining tied to the observed host population. In addition, we include a 10% hostless component. For these injections, positions are drawn uniformly over the image footprint, and magnitudes are sampled independently in each band from empirical distributions fitted to the UDEEP catalog. This component is intended to prevent the training set from being restricted to host-associated events only. The resulting injection catalogs are ingested by the LSST pipeline middleware and injected at the image level as part of the standard processing workflow1 . The injected exposures are then passed through the same DIA, detection, and measurement steps as the real survey data. To obtain an approximately balanced training set at the detection level, we tune the number of injected sources so that the number of recovered injected detections is comparable to the 1
lsst.source.injection Article number, page 3
A&A proofs: manuscript no. main
number of detections obtained from the corresponding real survey processing. In rc2_subset, the real survey processing contains 43 607 detections, while our injection procedure produces Ninj = 89 570 artificial point sources, of which 51 550 are recovered by the pipeline. The unrecovered injections predominantly populate the faint end of the injected distribution (Appendix A), and the recovered sample remains close to balanced relative to the survey detections. To verify that injections remain sparse at the image level, we estimate the affected pixel fraction. Averaging over the 40 × 6 = 240 CCD images gives ⟨Nsrc ⟩ ≃ 373.2 injections per CCD. Approximating the footprint of each source by a disk of radius R = 3 FWHM ≃ 12.5 pixels affects a fractional area of about 2%. This is an approximate upper bound, since source footprints may overlap, and indicates that injections remain sparse on the scale of the full detector. We combine the recovered injection-driven detections with detections from the corresponding real survey processing (i.e. detections not associated with injected truth) to form the baseline training set. The baseline training set therefore consists of single-detection cutouts with two training labels: recovered injections are labeled Transient, while all survey detections are labeled Bogus, although this latter class is expected to contain genuine transients that the model aims to recover at inference time.
Step SNR < 5 Time window 6 Minimum nights Negative flux Point source host Low flux ratio
Objects discarded 157, 112 568, 540 3, 130, 0642 352 95 34
Objects remaining 3, 699, 427 3, 130, 887 823 471 376 342
Table 1: Summary of the filtering steps used for light-curve selection for human labeling. For each step, we report the number of objects discarded and the number remaining afterward.
To further suppress contamination, we incorporate hostgalaxy information from the HSC deep coadds. For each object, we identify the epoch of maximum absolute difference flux and match its sky position to the nearest template source within 1′′ . Objects are retained if they are either hostless or associated with an extended host, and if the transient-to-host flux ratio exceeds 1.4 in the i band. These cuts intentionally favor well-measured SN-like transients over low-contrast or ambiguous cases and therefore define a high-purity evaluation subset rather than a representative sample of the full alert stream. The numbers of discarded objects are reported in Table 1.
2.3. Constructing the evaluation set
Because our method relies on a noisy class in which genuine transients are mislabeled during training, we need an independent ground truth for evaluation. We therefore construct a manually curated evaluation set from HSC-UDEEP light curves. Neither human labels nor light curve level information is used for model training and inference, filtering and manual inspection is used only at the evaluation stage. Since rc2_subset does not provide sufficiently sampled light curves for this purpose, we derive the evaluation set from the more densely sampled UDEEP dataset. We first filter the UDEEP light curves to isolate wellmeasured SN-like candidates, and then manually inspect the remaining objects to define the set of real detections used for evaluation.
Class
Objects (N = 306) Sources (N = 4,820)
SN-like OT Transient
33 (10.8%) 41 (13.4%) 74 (24.2%)
1,211 (25.1%) 1,340 (27.8%) 2,551 (52.9%)
Bogus
232 (75.8%)
2,269 (47.1%)
Total
306 (100%)
4,820 (100%)
Table 2: Detailed class composition of the manually curated evaluation set after manual labeling and removal of objects labeled as Unknown. OT denotes other transient or variable objects. We report both the number of objects and the number of source detections in each class.
Light-curve filtering. We apply a sequence of cuts to isolate a
compact set of well-sampled transient-like light curves for manual inspection. First, we remove individual source measurements with signal-to-noise ratio SNR < 5. Objects with no remaining measurements above this threshold are discarded. Next, we identify the epoch of maximum flux for each object and retain only light curves whose informative measurements lie within a window of [−30, +100] days around maximum light. We then require at least six distinct observing nights with valid detections and reject objects whose mean PSF flux across retained epochs is negative, as these are indicative of systematically spurious or over-subtracted signals. This cut may remove genuine astrophysical negative residuals or fading events, such as disappearing sources or long-term variables. These objects are outside the scope of the present high-purity SN-like curated set, which is designed to evaluate the recovery of positive transient excesses rather than to provide a complete census of variable phenomena. 2 The large number of sources discarded at this step is mostly due to single observations, which account for 2, 959, 860 sources in UDEEP.
Article number, page 4
Human labeling.
The 342 filtered light curves are then inspected manually using the coadded template image, bandaveraged science and difference images, and the multi-band light curve itself (See App. B). We assign each object to one of four categories: SN-like, Other transient or variable, Bogus, or Unknown. Objects labeled Unknown are excluded from the binary evaluation set, leaving 306 objects and 4,820 source detections for quantitative analysis. Although 232 of the 306 objects are labeled Bogus, the evaluation set is nearly balanced at the source-detection level because transient light curves typically produce detections over more epochs. Bogus-labeled objects can also contribute multiple detections, for example when subtraction residuals at a fixed location near a bright star recur across epochs owing to imperfect masking, PSF matching, or astrometric registration. These residuals often appear as dipole-like artifacts rather than isolated single-epoch events. We emphasize that this set serves as a manually curated evaluation set rather than spectroscopic truth.
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
Flatten
ŷ
FC2 1
3× 3×
x : (30, 30, 2)
30
×
30
15
×
15
7
×
7
256
MaxPool 2×2
Dropout2D
MaxPool 2×2
Dropout2D
MaxPool 2×2
Dropout2D
FC1 64 ReLU
Sigmoid
Dropout 0.25
ConvBlock3 256
ConvBlock2 128
ConvBlock1 64
Fig. 2: Schematic of the CNN architecture selected by hyperparameter optimization.
3. Deep Learning classification 3.1. Classes and Confusion Matrix
After injection, the training data consist of two labeled groups: – injected detections, labeled as Transient; – survey detections, labeled as Bogus, although this class is contaminated by genuine transients present in the survey data. The objective of training is therefore not simply to separate injected from survey detections, but to learn a decision rule that separates transient-like from bogus-like detections despite label noise in the survey class. At training time, the confusion matrix must be interpreted relative to the observed training labels, not as a ground-truth confusion matrix: – True positives (TP) are injected data classified as transient. – True negatives (TN) are survey data classified as bogus. – False positives (FP) are survey data classified as transient. – False negatives (FN) are injected data classified as bogus. In this setting, the FP category is of particular interest, because it contains the survey detections reclassified by the model as transient and may therefore include genuine transients hidden inside the noisy bogus-labeled class. By contrast, the FN category corresponds to injected detections that the model fails to recover as transients. Since the injected class is treated as the clean reference class during training, understanding and minimizing this failure mode is important. 3.2. Network Architecture
The network architecture consists of three convolutional blocks (L = 3). Each block comprises a 3 × 3 convolutional layer, followed by batch normalization, a ReLU activation, 2×2 max pooling, and spatial dropout regularization. The number of filters in the three convolutional layers is parameterized by a single base width F, and follows a geometric progression (F → 2F → 4F). This coupling enforces a monotonic increase in channel width and avoids degenerate configurations while still allowing the overall capacity of the model to be controlled through F.
Each input x ∈ R30×30×2 is a two-channel cutout, constructed by stacking the difference and template images (coadded). We do not include the science image explicitly, in order to reduce computational cost and input dimensionality while preserving both the subtraction residual and the static host-galaxy context. The convolutional feature extractor computes hL = ConvBlock3 ◦ ConvBlock2 ◦ ConvBlock1 (x), where each ConvBlockℓ consists of convolution, batch normalization, ReLU, max pooling, and dropout. After convolutional feature extraction, the feature maps are flattened and passed through two fully connected layers, where the hidden dimensionality depends on the spatial reduction factor and a configurable hidden size, as will be discussed next. The classifier computes f (x) = FC2 Dropout ReLU FC1 Flatten(hL ) , and the final prediction is obtained via ŷ = σ f (x) , with σ(·) denoting the sigmoid activation function for binary classification (transient vs. bogus detection). 3.3. Hyperparameter Optimization
The CNN hyperparameters are tuned via Bayesian optimization using the Optuna framework (Akiba et al. 2019), with the objective of minimizing the loss. To keep the search statistically efficient under a limited budget of 50 trials, we adopt a structured and low-dimensional hyperparameterization. Architecturally, only the base convolutional width F (which determines the filter configuration (F, 2F, 4F)) and the number of units in the dense layer are tuned. Similarly, a single base dropout rate (B) controls regularization throughout the network: convolutional blocks apply scaled rates ((0.5B, 0.5B, 0.5B)), while the dense layer uses B. This geometric progression enforces smooth capacity scaling and reduces redundant architectural degrees of freedom. The search is performed using Optuna’s Tree-structured Parzen Estimator (TPE) sampler, introduced in (Bergstra et al. Article number, page 5
A&A proofs: manuscript no. main
survey data (y = 0); and (ii) a training label ỹ (stored as is_injection), which is the label used for model training and may be intentionally corrupted. In the baseline dataset, y = ỹ for all samples. To construct noisy variants, we apply a one-sided, classconditional corruption mechanism: we randomly select Nmis injection-origin examples (y = 1) and flip their training label from ỹ = 1 to ỹ = 0. This contaminates the survey-data training class with transient-like samples. Because labels are flipped in only one direction, we additionally subsample survey-origin Hyperparameter Value examples so that the two training-label classes remain approxiBatch size (batch_size) 128 −4 mately balanced, N(ỹ = 1) ≃ N(ỹ = 0). Learning rate (learning_rate) 1.62 × 10 Using this procedure, we generate three datasets with η ≃ Base filters F (model_params.base_filters) 64 15%, 25%, and 35% contamination in the survey-data class. TaDense units (model_params.units) 64 ble 4 summarizes both the training-label distribution (ỹ) and the Dense dropout (model_params.base_dropout) 0.25 underlying provenance distribution (y). Table 3: Best architectural hyperparameters obtained from the Bayesian optimization. 4.1. Co-teaching 2011), which models the objective function with non-parametric density estimators in hyperparameter space. To reduce computational cost, the Bayesian optimization is carried out on a randomly selected 30% subset of the available dataset, reserved specifically for hyperparameter search. The best configuration found by the search is summarized in Table 3 and represented in Fig. 2. This configuration is used to train all the models we compare.
3.4. Early Stopping and Model selection strategy
Early stopping is commonly based on validation loss or validation accuracy. In our setting, however, these criteria are not ideal because the survey-labeled bogus class is contaminated by genuine transients. For model selection, we instead monitor the false negative rate (FNR) on the injected class, that is, the fraction of injected detections classified as bogus. This quantity is not used as the training objective itself; rather, it is used as an early-stopping and checkpoint-selection criterion. The intuition is that the injected class has the most trustworthy labels available during training, so preserving high recall on this class helps avoid selecting models that overfit the noisy survey labels. We therefore use injected-class FNR as a proxy criterion for checkpoint selection, while the network itself is still optimized using the standard training loss. This strategy biases model selection toward sensitivity to transient-like detections without explicitly rewarding memorization of the noisy bogus-labeled survey class.
4. Weakly Supervised Learning for Injection-based Transient–Bogus Classification Supervised classifiers can be sensitive to label noise. While moderate levels of corruption may be tolerated (Rolnick et al. 2017), they can bias the learned decision boundary and degrade performance (Arpit et al. 2017). As calibration pipelines improve in modern wide-field surveys (e.g., Padmanabhan et al. 2008; Burke et al. 2017; Huang et al. 2022) and DIA implementations continue to reduce subtraction artifacts and false positives (Liu et al. 2024), the bogus fraction among DIA candidates may decrease, and the survey-labeled class will contain a larger fraction of genuine transients, weakening our injection-based label-purity assumptions. We therefore study the robustness of our framework under controlled one-sided label corruption by artificially contaminating the survey-labeled training class with injection-origin samples. The baseline dataset, described in Section 2.2, contains 51 550 injection-origin detections and 43 607 survey-origin detections, for a total of 95 157 examples. In this baseline, the injection provenance and the training label coincide. For each example, we retain two binary variables: (i) a provenance label y (stored as spy_injected), indicating whether the cutout truly originates from an injection (y = 1) or from Article number, page 6
Co-teaching (Han et al. 2018) is a method used in machine learning to handle datasets with noisy labels. Two models are initialized and trained simultaneously on the same dataset. During each training iteration, each model selects a subset of training samples with the smallest loss (i.e., the samples it is most confident about). Each network then uses the small-loss samples selected by the other network to update its parameters. The intuition is that samples with smaller losses are more likely to be correctly labeled, while samples with higher losses are more likely to be noisy. This cross-teaching mechanism ensures that networks do not reinforce their own biases and are less likely to overfit to noisy labels. Over time, as the networks learn, they become better at identifying clean samples, and the training process becomes more robust to label noise3 . In our setting, standard Co-teaching has two practical limitations. First, it requires the user to specify a forget rate, i.e. the fraction of samples assumed to be corrupted; in our experiments this quantity is set using prior knowledge from the controllednoise setup (Appendix C). Second, the sample-selection step is performed globally within each mini-batch, without explicitly accounting for class-dependent noise asymmetry. In our application, however, the noise is concentrated primarily in the surveylabeled bogus class, while the injection-labeled transient class is intended to remain comparatively clean. As a result, the standard selection rule may discard informative transient examples rather than focusing rejection on the noisy class. A formal description of standard Co-teaching is given in Appendix D. 4.2. Asymmetric Co-teaching
In our application, the label noise is highly asymmetric: the fraction of corrupted labels is much higher in the bogus class than in the transient class. In addition, for our target use case, missing a true transient is typically more costly than passing some additional false detections to downstream filtering stages. Standard Co-teaching uses a single forget rate and selects small-loss samples over the whole mini-batch, independently of their class. As a result, there is no guarantee that the highloss samples discarded during training are predominantly noisy examples from the noisy majority class (bogus); the procedure may instead discard rare but informative transient examples that happen to incur higher loss. To address this issue, we propose 3 All parameter configurations for each Weakly Supervised Learning and Uncertainty Quantification method are presented in Appendix C.
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
Dataset
η [%]
Total N
Baseline Low-noise Medium-noise High-noise
0.00 15.09 24.99 35.23
95 157 68 821 44 666 30 475
Observed training labels N(ỹ = 1) N(ỹ = 0) (inj.-labeled) (survey-labeled) 51 550 43 607 34 398 34 423 22 349 22 317 15 262 15 213
True provenance N(y = 1) N(y = 0) (injection-origin) (survey-origin) 51 550 43 607 39 593 29 228 27 927 16 739 20 621 9 854
Nmis 0 5 195 5 578 5 359
Table 4: Summary of the datasets used to evaluate robustness to controlled asymmetric label noise. For each dataset, we report both the observed training labels ỹ (used for optimization) and the true provenance labels y (used only for analysis). Noise is introduced by flipping a subset of injection-origin samples from ỹ = 1 to ỹ = 0, so that the survey-labeled class contains a controlled fraction of hidden injection-origin samples. The contamination fraction is η = P(y = 1 | ỹ = 0) = Nmis /N(ỹ = 0), where Nmis = N(y = 1, ỹ = 0) is the number of injection-origin samples relabeled as survey-labeled for training. The total number of samples decreases with increasing noise because survey-origin samples are subsampled to keep the two training-label classes approximately balanced. Asymmetric Co-teaching (Asym-Co-teaching), which extends Co-teaching by introducing class-specific forget rates and classwise selection of small-loss samples.
Keep the smallest-loss fraction ρ(t) = 1−r(t) over the whole mini-batch.
Similarly to Co-teaching, the fractions of samples kept in each class at iteration t are therefore 1 − ra (t) and 1 − rb (t), and we set kt,a = (1 − ra (t)) |It,a | , kt,b = (1 − rb (t)) |It,b | .
Loss,
Co-Teaching:
We introduce class-specific forget rates ! t ra (t) = min ra · , ra , Tk ! t rb (t) = min rb · , rb , Tk
Sorted samples
Survey
Asymmetric Co-Teaching:
For network 1, we sort the samples within each class according to their loss: (1) (1) ℓi(1) (t), t (t) ≤ ℓ t (t) ≤ · · · ≤ ℓ t i i
Injected
1,a
|It,a |,a
2,b
|It,b |,b
Loss,
where (it1,a , . . . , it|It,a |,a ) is a permutation of It,a and (it1,b , . . . , it|It,b |,b ) is a permutation of It,b . The sets of reliable samples selected by network 1 at iteration t are then
Sorted samples within class
R1,t,a = {it1,a , . . . , itkt,a ,a },
Fig. 3: Illustration of sample-selection rules in Co-teaching variants. Top: standard Co-teaching applies a class-independent forget rate and selects small-loss samples globally within the minibatch. Bottom: Asym-Co-teaching performs class-wise smallloss selection with class-specific forget rates. Hollow markers denote discarded samples. Let: Bt = {(xi , ỹi )}i∈It denote the mini-batch at epoch t, where It is the index set of samples in the batch and ỹi ∈ {a, b} is the observed training label. We define the class-specific index sets It,a = {i ∈ It : ỹi = a},
2,a
(1) (1) ℓi(1) (t), t (t) ≤ ℓ t (t) ≤ · · · ≤ ℓ t i i 1,b
Keep class-specific smallest-loss fractions ρa(t), ρb(t) within each class.
R1,t,b = {it1,b , . . . , itkt,b ,b }.
Equivalently, they can be characterized as X X R1,t,a = arg min ℓi(1) (t), R1,t,b = arg min ℓi(1) (t). S ⊂It,a |S |=kt,a i∈S
S ⊂It,b |S |=kt,b i∈S
Analogously, using the losses ℓi(2) (t) of network 2, we define the sets R2,t,a , R2,t,b . For each mini-batch, both networks compute individual losses, rank the samples by loss within each class, and then train on the small-loss samples selected by the peer network. The Asymmetric Co-teaching losses are thus given by L(1) Asym-co (t) =
X 1 BCE fθ1 (xi ), ỹi |R2,t,a | + |R2,t,b | i∈R 2,t,a ! X + BCE fθ1 (xi ), ỹi , i∈R2,t,b
It,b = {i ∈ It : ỹi = b}.
(2) 1 L(2) BCE Asym-co (t) = |R 1,t,a | + |R1,t,b | i∈R 1,t,a X
The per-sample losses of network 1 and network 2 on this batch are defined by their Binary Cross Entropy (BCE) as ℓi(1) (t) = BCE fθ1 (xi ), ỹi ,
(1)
ℓi(2) (t) = BCE fθ2 (xi ), ỹi ,
i ∈ It .
+
X
fθ2 (xi ), ỹi
!
BCE fθ2 (xi ), ỹi .
i∈R1,t,b
Article number, page 7
A&A proofs: manuscript no. main
These class-specific forget rates determine the fraction of samples discarded from each class at epoch t, with remember rates ρc (t) = 1 − rc (t) controlling the number of samples retained for training.
5. Uncertainty Quantification for injection-based Real-Bogus Classification Uncertainty quantification is fundamental to scientific inference. Although machine learning can be viewed as an extension of statistical modeling, obtaining uncertainty estimates that are simultaneously well calibrated, predictive, and computationally efficient remains challenging, particularly in deep learning (Guo et al. 2017). In this work, we evaluate and compare several uncertainty quantification approaches for injection-based realbogus classification. 5.1. Evaluation of the uncertainty methods
We evaluate uncertainty methods along two complementary axes: (i) the quality of their probabilistic predictions, assessed with standard calibration and scoring metrics; and (ii) the extent to which the resulting uncertainty quantification correlates with physically meaningful indicators of classification difficulty. Calibration metrics assess whether predicted probabilities are statistically reliable estimates of true class frequencies (Niculescu-Mizil & Caruana 2005). For probabilistic evaluation, we report the negative loglikelihood (NLL), the Brier score, and the expected calibration error (ECE). NLL and Brier score are proper scoring rules, and therefore reflect both calibration and sharpness of probabilistic predictions, while ECE more directly quantifies the mismatch between predicted confidence and empirical accuracy. For binary classification, the negative log-likelihood is N
NLL = −
1 X yn log p̄n + (1 − yn ) log(1 − p̄n ) , N n=1
(3)
where p̄n denotes the predictive probability assigned to the positive class for sample n, and yn ∈ {0, 1} is the observed label. Lower NLL indicates better probabilistic predictions and strongly penalizes overconfident errors, making it particularly sensitive to miscalibration. The Brier score is N
BS =
1 X ( p̄n − yn )2 . N n=1
(4)
Lower Brier scores indicate better probabilistic predictions. The expected calibration error (ECE) measures how well a model’s estimated probabilities match the true probabilities. This is done by calculating the weighted average error of the estimated probabilities. ECE =
M X |Bm | m=1
N
|acc(Bm ) − conf(Bm )| ,
where di = rank(xi )−rank(yi ) is the difference between the ranks of corresponding values, and n is the number of samples. 5.2. Deep Ensembles
Deep Ensembles (or ensembles) provide a simple and effective approach to predictive uncertainty in deep learning (Lakshminarayanan et al. 2017). Unlike explicitly Bayesian neural-network methods (MacKay 1992), they require no modification of the base architecture and often provide strong empirical uncertainty quantification. They also differ from methods such as MC dropout because independent initialization and training encourage exploration of different regions of the loss landscape (Fort et al. 2020). For our application, we are independently training CNN models. Each model, denoted by fm (·), shares the same architecture (see Section 3.2) but is initialized with different random weights. All models are trained separately on the same dataset following the standard procedure described previously. At inference time, the ensemble prediction is obtained by averaging the individual model probabilities. Beyond improving prediction stability, the ensemble also facilitates uncertainty quantification. We quantify this uncertainty from the dispersion of the ensemble predictions. In practice, we use the standard deviation of the member probabilities for a given input x, with larger dispersion indicating greater predictive disagreement. Further details are given in Appendix F.
(5) 5.3. Repulsive Ensembles
where Bm is the set of predictions falling in bin m, acc(Bm ) is the empirical accuracy in that bin, and conf(Bm ) =
In addition to calibration, we assess the quality of our uncertainty quantification by examining its correlation with two physically meaningful and interpretable quantities: the signal-to-noise ratio (SNR) and the distance from maximum brightness for supernova samples. At low SNR, the pixel-level evidence in a single-epoch difference-image cutout may be insufficient to reliably distinguish a genuine PSF-like astrophysical transient from a chance noise fluctuation with similar morphology, implying an effective information limit for transient-bogus classification. Similarly, sources far from the peak luminosity of a supernova are expected to be more difficult for the model to classify, because supernovae observed during their rise or decline phases have less distinctive photometric signatures. Note that this metric can be computed only for sources that are part of light curves labeled as SN-like (see Section 2.3), excluding sources that are unlabeled or classified as other transient or variable or Bogus. While these two quantities are correlated, it is important to note that SNR is not evenly distributed across the training classes, as shown in Appendix E.1, and is therefore expected to follow a similar tendency. Using two metrics with different classdependent behavior makes the evaluation more robust. To determine the correlation between the uncertainty estimator method and the physical quantities, we use the Spearman rank correlation coefficient ρ computed as: P 6 ni=1 di2 ρ=1− , (6) n(n2 − 1)
1 X p̄i . |Bm | i∈B m
Article number, page 8
Repulsive Ensembles (D’Angelo & Fortuin 2021) extend standard deep ensembles by introducing an interaction term during training that encourages ensemble members to represent diverse functions while still fitting the data well. In practice, the repulsion is imposed in function space rather than directly in weight
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
space, since different parameter settings can correspond to similar predictive functions. The motivation is to increase functional diversity across ensemble members, thereby improving uncertainty estimates relative to ensembles that remain concentrated around a narrow region of the loss landscape. At inference time, predictions are aggregated in the same way as for standard deep ensembles, and uncertainty is quantified from the dispersion across ensemble members.
MC Dropout defines a local variational distribution around each configuration by keeping dropout active at test time. This motivates the approximation N
q(θ) =
1 X qn (θ). N n=1
with corresponding predictive approximation
5.4. MC Dropout
Monte Carlo Dropout (MC Dropout) (Gal & Ghahramani 2016) is a widely used uncertainty quantification method in deep learning. It can be interpreted as a particular BNN construction in which Bernoulli dropout masks induce an approximate distribution over the weights; running the model multiple times with dropout enabled at inference then amounts to Monte Carlo sampling from the resulting predictive posterior. A practical advantage of the method is that it requires no architectural change beyond the use of dropout during training. The same trained network can then be sampled multiple times at inference by applying different dropout masks. Further formal details are given in Appendix F.
(9)
p(y∗ | x∗ , D) ≈
Z
p(y∗ | x∗ , θ) q(θ) dθ N
1 X = N n=1
Z
(10)
p(y∗ | x∗ , θ) qn (θ) dθ
(11)
≈
1 X Ez∼q(z) p(y∗ | x∗ , θ̂n , z) N n=1
(12)
≈
1 XX p y∗ | x∗ , θ̂n , zn,m , N M n=1 m=1
N
N
M
(13)
where zn,m ∼ q(z) denotes a dropout mask (sampled independently for each network n and Monte Carlo pass m).
5.5. Uncertainty for Co-teaching Methods
As discussed above, deep ensembles provide a strong baseline for uncertainty quantification by capturing variability across different modes of the loss landscape, and they can be further diversified through repulsive training. However, these models are computationally expensive: an ensemble of N networks requires N full training runs and N forward passes per sample at inference time. In the Co-teaching setting, two networks are already trained from different random initializations and can therefore be viewed as a small coupled ensemble, although they are not independent in the same sense as a standard deep ensemble. Because an ensemble of two networks may provide only limited diversity, we augment the Co-teaching pair with MC Dropout at inference time. This allows multiple stochastic predictions from each of the two trained models and increases the number of predictive samples at low additional training cost. The predictive probability now becomes: p̄(x∗ ) =
N M N M 1 XX 1 XX pn,m (x∗ ) = σ fn,m (x∗ ; θ̂n ; zn,m ) , N M n=1 m=1 N M n=1 m=1
(7) where N = 2 is the number of Co-teaching networks, from different initializations, composing our pseudo-ensemble, zn,m the Dropout Mask and M is the number of Dropout masks that will be applied to both models. Following the same logic, we can estimate the uncertainty from the second centered moment: N
VarCo-Ens-Dropout (x∗ ) =
M
1 XX pn,m (x∗ ) − p̄(x∗ ) 2 . N M n=1 m=1
(8)
This construction can be viewed as a heuristic mixture approximation in which each trained Co-teaching network defines one component and MC Dropout provides local stochastic samples around that component. In a Bayesian framework as a mixture approximate posterior, N the N networks converge to different solutions (modes) {θ̂n }n=1 due to independent initializations and training dynamics, while
6. Latent Space Representation Analysis of learned internal representations can provide useful insight into model behavior. In this work, we study the latent representation learned by the classifier by extracting hidden activations and visualizing them in two dimensions. This analysis is intended as a global representation-level probe of the model, rather than a full mechanistic interpretation of its internal computations (Sharkey et al. 2025). While much recent work focuses on identifying interpretable concepts in latent representations using sparse autoencoders (SAEs) (Wu & Walmsley 2025), we adopt a simpler approach based on direct exploration of latent activations through nonlinear dimensionality reduction. 6.1. Feature Extraction
The dimensionality of the latent representation depends on the layer under consideration. We focus on the output of the final hidden layer of the network (Fig. 2), where we expect highlevel abstractions and task-relevant patterns to be strongly represented. For each input sample, we record the activation vector at this layer. For a dataset of N samples, this gives a feature matrix F ∈ RN×d , where d denotes the dimensionality of the selected layer’s output. When N exceeds a predefined maximum sample size nmax , we apply stratified random sampling to obtain a representative subset while preserving the class distribution. For ensemblebased and co-teaching models, features are extracted from the first model in the ensemble. 6.2. Dimensionality Reduction with UMAP
The high dimensionality of F makes direct interpretation infeasible. We therefore use Uniform Manifold Approximation and Projection (UMAP) (McInnes et al. 2020) to compute a twodimensional embedding suitable for visualization. Article number, page 9
A&A proofs: manuscript no. main
Fig. 4: Two-dimensional UMAP projection of the embeddings from the penultimate dense layer of the classifier. Each point corresponds to one candidate and is colored by outcome relative to the labels: TP (yellow ×), TN (red •), FP (orange ⋆), and FN (gray ■). Insets (a-e) show representative Coadd (template) and Diff (difference) cutouts sampled from the indicated regions: (a) negative-Diff region, dominated by negative residuals characteristic of subtraction artifacts or over-subtraction; (b) low-SNR region, where candidates are difficult to distinguish from background fluctuations; (c) dipole region, showing positive/negative residual astrometric misalignment; (d) low-SNR transient region, containing faint point-source-like residuals consistent with marginal transient detections; and (e) high-SNR transient region, containing compact, symmetric positive residuals characteristic of confident transient-like candidates. The cutouts are provided as illustrative examples only, they were assembled manually to make the static figure informative and are therefore approximate with respect to exact UMAP locations and local neighborhoods. The interactive versions of the UMAP visualizations (with different overlays) are available in the Zenodo supplementary material (DOI:10.5281/zenodo.18434676). UMAP constructs a weighted nearest-neighbor graph in the original feature space and optimizes a low-dimensional embedding that preserves local neighborhood structure. As with other nonlinear embedding methods, the visualization is primarily useful for exploring local organization; global distances in the 2D projection should not be over-interpreted. In contrast to linear projections such as PCA (Wold et al. 1987), UMAP can capture nonlinear structure; similarly to tSNE (van der Maaten & Hinton 2008), it primarily emphasizes preservation of local neighborhood relationships. In practice, UMAP often produces embeddings with more apparent global continuity than t-SNE, although global distances are not guaranteed to be preserved. We compute a two-dimensional UMAP embedding with n_components=2 and Euclidean distance in the original feature space. We vary n_neighbors and min_dist over a small grid: – n_neighbors: size of the local neighborhood used to approximate the manifold, Article number, page 10
– min_dist: minimum distance between embedded points, controlling cluster compactness, – n_components: target embedding dimensionality (set to 2), – metric: Euclidean distance in the original feature space. Specifically, we perform a grid search over n_neighbors ∈ {5, 10, 15, 30, 50}, min_dist ∈ {0.0, 0.1, 0.25, 0.5}. and select the configuration using a simple visualization heuristic based on the spread of the embedding coordinates. This heuristic is used only to avoid overly collapsed visualizations and should not be interpreted as an objective measure of representation quality. Using an interactive graph built with the Bokeh library, we visualize the UMAP representation together with the corresponding confusion-matrix categories, as well as image SNR, predicted class probability, and model uncertainty when available. This interactive tool facilitates qualitative exploration of
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
model behavior and helps identify structured failure modes or mislabeled examples. The resulting embeddings are used as a qualitative analysis tool to inspect class structure, misclassifications, uncertainty patterns, and potential label noise in the learned representation.
Noise B
15%
7. Results 7.1. Weakly supervised learning
Table 5 compares weakly supervised learning (WSL) approaches under increasing label-noise rates on training sets, reporting subtype-specific recall for the positive class (survey transients) and specificity for the negative class (bogus detections), together with aggregate accuracy and receiver operating characteristic (ROC). When trained on the baseline dataset (current HSC injected survey data), all methods achieve strong performance (Acc ≈ 0.91-0.92; ROC ≈ 0.97), with Co-teaching and Asym-Coteaching providing the strongest overall balance across metrics. As we artificially increase the label noise in the baseline, the Standard model degrades substantially (e.g., at 35% noise, Acc = 0.812 and ROC = 0.903), whereas the co-teaching variants remain markedly more robust. In particular, Co-teaching and Asym-Co-teaching consistently maintain high ROC AUC across noise levels (≈ 0.95) and improve accuracy relative to the Standard approach. At higher noise (25-35%), Asym-Co-teaching offers the best aggregate results (Acc ≈ 0.87-0.89; ROC ≈ 0.95), indicating that co-teaching, especially with asymmetric handling of label noise, provides the most stable performance as annotation quality deteriorates. The comparable recall obtained for the SN-like and OT subclasses supports the hypothesis that, at the single-image level, SN-like injections provide useful training examples for detecting other transient classes as well. In addition to the two methods presented in Section 4, we also evaluated Stochastic Co-teaching (de Vos et al. 2023), which we briefly describe in Appendix D. It exhibits pronounced trade-offs across the noise regimes. This instability was observed throughout training, with difficulty in tuning the α and β parameters and in selecting the delay before rejection begins. In most training cases, Stochastic Co-teaching collapsed to rejecting predominantly one class, leading to unstable training.
25%
35%
WSL method Standard Co-t Asym-Co-t Standard Co-t Asym-Co-t Standard Co-t Asym-Co-t Standard Co-t Asym-Co-t
SN-like 0.885 0.919 0.898 0.798 0.920 0.843 0.792 0.882 0.887 0.764 0.807 0.812
OT 0.914 0.928 0.905 0.815 0.925 0.838 0.753 0.897 0.875 0.761 0.779 0.833
B 0.926 0.915 0.944 0.934 0.872 0.932 0.937 0.899 0.903 0.868 0.876 0.925
Acc 0.912 0.920 0.922 0.866 0.899 0.883 0.849 0.894 0.891 0.812 0.831 0.871
ROC 0.966 0.972 0.971 0.946 0.963 0.959 0.939 0.959 0.951 0.903 0.921 0.950
Table 5: Evaluation of weak supervision learning methods under varying label noise conditions. For the positive class (survey transients), recall is reported separately for Supernova like (SNlike) and other transient or variable types (OT). For the negative class, specificity is shown for bogus detections (B). Overall accuracy (Acc) and ROC AUC summarize aggregate performance. duction in ECE, indicating both improved probabilistic accuracy and substantially better calibration. These results indicate that our Ensemble-MC-Dropout inference strategy, combined with dual-network architecture, produces probability estimates that more accurately reflect true class frequencies. The lower NLL, in particular, demonstrates that our method effectively reduces overconfident incorrect predictions, a critical property for downstream tasks. The simultaneous improvement in three metrics confirms that our uncertainty quantification approach achieves superior calibration without sacrificing robustness to extreme probabilities. Metric Standard MC Dropout Deep Ensemble Repulsive Ensemble Our method
NLL 0.453 0.334 0.314 0.308 0.292
Brier Score 0.074 0.069 0.065 0.065 0.063
ECE 0.0633 0.0405 0.0535 0.0511 0.0375
7.2. Uncertainty Quantification
Table 6: Negative Log-Likelihood (NLL), Brier Score and Expected Calibration Error (ECE) on 11 bins for the standard model, the Deep Ensemble, Repulsive Ensemble and our method, Ensemble-MC-Dropout for Co-teaching
Table 6 summarizes probabilistic performance and calibration for the different uncertainty-quantification strategies. All UQ approaches improve over the standard deterministic baseline, with consistent reductions in NLL, Brier score, and ECE. The standard CNN achieves an NLL of 0.453 and a Brier score of 0.074, illustrating the comparatively poorer probabilistic quality of deterministic predictions. This behavior is easily noticeable in Fig. 5. Our method, Ensemble-MC-Dropout for Co-teaching (Section 5.5) achieves the best overall performance across the three reported metrics, with the lowest NLL (0.292), the lowest Brier score (0.063), and the lowest ECE (0.0375). Deep Ensembles and Repulsive Ensembles remain competitive, particularly on NLL and Brier score, but are consistently outperformed by our approach, with the clearest advantage observed in ECE. Relative to the standard model, this corresponds to a 35.5% reduction in NLL, a 14.9% reduction in Brier score, and a 40.8% re-
In addition to calibration, a useful uncertainty estimator should assign larger uncertainty to samples that are intrinsically more difficult to classify, such as low-SNR detections or supernovae observed far from peak brightness. Table 7 reports Spearman correlations between uncertainty and these two physically meaningful quantities. All methods show the expected negative correlation with SNR (ρ < 0), indicating that uncertainty increases as signal quality degrades. Our method achieves the strongest correlation (ρ = −0.296), marginally outperforming Deep Ensembles (ρ = −0.290) and substantially improving upon MC Dropout (ρ = −0.239). All methods also exhibit the expected positive correlation with distance from maximum brightness. Here, Deep Ensembles achieve the highest correlation (ρ = 0.221), with our method performing nearly identically (ρ = 0.218). Article number, page 11
A&A proofs: manuscript no. main
Fraction of Positives
1.0
7.3. Latent space representation
0.8 0.6
Perfectly calibrated Standard MC Dropout Ensemble Rep Ensemble Asym Co-teaching
0.4 0.2
Residual
0.0 0.50 0.25 0.00 0.25 0.50
0.0
0.2
0.4
0.6
Mean Predicted Probability
0.8
1.0
Fig. 5: The top panel shows the fraction of positive samples as a function of the mean predicted probability, using 11 equally spaced probability bins. The dashed diagonal indicates perfect calibration. The bottom panel shows the residuals with respect to perfect calibration.
Overall, these results show that the proposed method produces uncertainty quantification that are physically meaningful and competitive with substantially more expensive ensemblebased baselines, while requiring only two jointly trained networks augmented with MC Dropout at inference time, which greatly reduces training cost compared with large ensembles. Physical value MC Dropout Deep Ensemble Repulsive Ensemble Our method
ρ SNR −0.239 −0.290 −0.271 −0.296
ρ Max. bright 0.176 0.221 0.198 0.218
Table 7: Spearman rank correlation coefficients (ρ) between predicted uncertainty and signal-to-noise ratio (SNR) and distance from maximum brightness (Max. bright, in days). Negative ρSNR values indicate that uncertainty appropriately increases for lowSNR detections, while positive ρMax. bright values indicate higher uncertainty for supernovae observed far from peak luminosity. It is also informative to examine the class-conditional values of ρSNR , reported in Appendix E.2. The overall correlation between uncertainty and SNR is driven primarily by the truepositive subset. A weaker but still visible trend remains in the true-negative subset, whereas the correlation is much less pronounced for misclassified examples. This suggests that SNR is most directly informative for uncertainty within the transient-like detections, where low signal quality more strongly affects classification confidence. This interpretation is consistent with the class-dependent SNR distributions shown in Appendix E.1, and with the latent-space visualization in App. G (interactive version available on Zenodo), where low-SNR samples appear more diffusely distributed within the transient region than in the bogus region. Article number, page 12
The projection displayed in Fig. 4 highlights a broad separation between negative and positive manifolds, with most ambiguous cases and misclassifications concentrated near their interface, thereby shaping the decision boundary. One key finding of the UMAP representation of the latent space is that it allows the user to understand which patterns are being identified by the model even within a class. In addition to the class information, UQ and SNR can be displayed (see Appendix G). The visualization is also useful for inspecting misclassified cutouts and identifying structured failure modes or potential label noise. We therefore use it as a qualitative complement to the quantitative analyses, rather than as a standalone measure of model performance. 7.4. Towards object identification from source classifications
Although the model classifies DIA sources independently, its predictions can be grouped at the object level to perform lightcurve-level transient-bogus discrimination. Using the method described in Section 5.5, we aggregate source-level predictions within each labeled light curve and assign an object-level label by majority vote. The resulting performance is summarized in Table H.1. Figure 6 shows the distribution, for each light curve in the labeled evaluation set, of the fraction of sources classified as transient; the corresponding cumulative counts are shown in App. H. A first notable result is the strong internal consistency of the source-level predictions within individual light curves: 232/306 light curves (75.8%) have at least 90% of their sources assigned to the correct class. For the transient class specifically, 79.7% of light curves have more than 80% of their sources classified as transient. On this curated evaluation set, aggregation of the sourcelevel predictions yields very strong object-level performance: 304/306 light curves are assigned the correct binary label (99.35% accuracy), with only two errors overall (one false positive and one false negative). At the source level, class separation also remains strong, with similar recall for supernovae and other transients or variables (0.887 and 0.896, respectively) and a bogus specificity of 0.951. The uncertainty quantification remains consistent with observational difficulty. For supernova detections, uncertainty is lowest near maximum light (mean uncertainty ≃ 0.074) and increases substantially for observations far from peak (> 50 days; mean uncertainty ≃ 0.180), where the corresponding accuracy falls to 0.75. This behavior is consistent with the uncertainty analysis presented above and suggests that epoch-level uncertainty could be useful for down-weighting ambiguous, lowinformation measurements in downstream light-curve analyses. Together, these results indicate that single-epoch, perdetection classification scores (and their associated uncertainties) are sufficiently stable to support reliable light-curve identification after grouping, while also providing a principled signal for down-weighting ambiguous, low-information epochs during downstream light-curve-based analyses. Encouraged by the strong performance obtained on the labeled set (Table H.1), we then deploy the same grouping rule as an inference-time filter on the full UDEEP dataset, which contains 3, 856, 539 distinct objects. Among these, 269, 968 objects have more than 50% of their sources classified as transient. However, a more detailed inspection shows that most of these objects
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ All light curves 180
Transient light curves
a)
160
b)
Bogus light curves c)
Count
140 120 100 80 60 40 20 0 0
20
40
60
80
Predicted transient (%)
100 0
20
40
60
80
Predicted transient (%)
100 100
80
60
40
20
0
Predicted bogus (%)
Fig. 6: Light-curve-level classification score distributions. Shown are the distributions of the fraction of source detections classified as transient for all light curves (a) and for true transient light curves (b), together with the distribution of the fraction of source detections classified as bogus for true bogus light curves (c). are observed on only a single night. Requiring at least five different nights of observation reduces this set to 1, 770 transient candidates, including 108 with more than 50 distinct observations.
8. Discussion and open questions The aim of this work was to address the two challenges introduced in Section 1: learning real-bogus classification without human-labeled training data, and obtaining uncertainty quantification and representations that remain interpretable and useful for downstream analysis. Our experiments show that the proposed framework recovers strong source-level performance and, after grouping detections, very strong object-level performance on the manually curated evaluation set (Section 7.4). Although the single-source metric may appear less reliable, it is important to note that this limitation is not solely algorithmic. Indeed, our labels are derived from observed light curves, and a fraction of events is intrinsically ambiguous given the available observables. This sets an effective upper bound on achievable performance for the single-source task. Inspection of misclassified cutouts with the interactive visualization supports this interpretation. Combined with in-depth latent-space visualization, these results help explain the observed behavior. A key strength of the method is its portability across surveys with limited adaptation. In the context of forthcoming datasets such as LSST, it remains an open question whether direct inference on survey data, despite the expected domain shift, is sufficient, or whether performance and calibration are best preserved by re-running the full pipeline with injection, training, and inference using survey-specific simulations. In both cases, the methodology is survey-agnostic, with the main challenges being adaptation to other data-management tools, as we have shown that our method is robust to extreme noise rates. Beyond the present scope, weakly supervised methods appear promising both for our current approach, fully human labelfree training, and as a way to mitigate the inherent label noise expected in community-labeled datasets. In parallel, asymmetric co-teaching could potentially be leveraged to estimate dataset noise levels, for example by exploiting the training dynamics to infer class-dependent noise rates and, in turn, a robust Real-Bogus ratio. While noise-rate estimation is discussed in de Vos et al. (2023) for a global noise level, extending it to class-
conditional noise in survey data is an interesting direction for future work. Regarding uncertainty quantification (UQ), it is noteworthy that our approach matches or outperforms standard baselines, including MC dropout and large ensembles. Given that ensembles with N = 50 are expected to explore a broader area of the loss landscape than methods that differ mainly at inference time (Fort et al. 2020), this result suggests that the training coupling induced by Asym-Co-teaching may play a central role. An important question for future work is whether the gain arises primarily from Asym-Co-teaching and the interaction between the two models during training, or whether the combination of MC dropout and a small ensemble could serve as a robust alternative to large ensembles. Finally, we find that UQ correlates with physically motivated quantities and improves calibration. Consistent with this, latentspace projections indicate that the UQ captures the expected decision boundary, whereas SNR is less structured in representation space, particularly in the bogus region. This supports the hypothesis that the model relies on features that are not well summarized by SNR alone, motivating further work to identify the dominant image and contextual factors driving both predictions and uncertainties. In addition, it would be valuable to explore concept-based interpretability methods applied to the learned representations (Kim et al. 2018), as well as more advanced methods such as (Wu & Walmsley 2025), to identify human-interpretable concepts associated with both astrophysical transients and bogus detections. Such concepts may guide targeted diagnostics, improve the realism of injections, and potentially lead to explicit bogus simulation.
9. Conclusion We presented a human-label-free approach to Real-Bogus classification for transient candidates. The method is trained using physically motivated injected transient classes together with a highly contaminated survey class, enabling learning without relying on human-labeled training data. This setup allows us to build on weakly supervised learning and to develop a training strategy dedicated to an asymmetric-noise regime. We show that the approach remains robust under noise levels exceeding the transient-to-bogus ratio typical of current pipelines, supporting its applicability to future surveys. Article number, page 13
A&A proofs: manuscript no. main
We also presented a comparative study of uncertainty quantification methods using both calibration and physically motivated metrics. Building on the dual-network setting of coteaching, we introduced a hybrid strategy combining elements of deep ensembles and MC dropout. This approach reduces computational cost while achieving UQ performance comparable to, or better than, the state-of-the-art baselines considered here. To support interpretation, we developed a visualization framework enabling in-depth analysis of model behavior. Finally, we extended our evaluation to light-curve space, where we obtain nearperfect classification accuracy on the labeled set. Author contributions. R.B.G. contributed to the data processing with the LSST Gen3 pipeline, and carried out the methodological development, including the physically motivated injection strategy, weakly supervised learning framework, and uncertainty quantification methods. He performed the main analysis, implemented the ML4transients library, and wrote the manuscript. B.S. provided expertise on the LSST Gen2 and Gen3 difference-imaging pipeline, including resolving compatibility issues, supporting database management, and supervising the project. D.F. conceived the project, initiated the data-processing strategy and latent-space analysis, and provided overall scientific supervision. B.R. initiated the analysis of light curves and difference-image cutouts and implemented an initial CNN model for real–bogus classification. M.Y. led the initial data-processing effort, running the LSST Gen2 pipeline for HSC and developing an early source-injection framework. M.S. further developed and systematically evaluated the CNN, including performance studies and the incorporation of additional inputs. M.G. improved the CNN through hyperparameter optimization and developed the UMAP-based analysis and visualization tools. V.P. provided guidance on interpretability. Acknowledgements. Part of this work was supported by the European Union’s Horizon Europe research and innovation programme under the Marie Sklodowska-Curie grant agreement No 101168829, Challenging AI with Challenges from Physics: How to solve fundamental problems in Physics by AI and vice versa (AIPHY).
References Abell, P. A., Allison, J., Anderson, S. F., et al. 2009, LSST Science Book, Version 2.0, Tech. rep., LSST Science Collaboration, 596 pages. Also available at full resolution at http://www.lsst.org/lsst/scibook Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. 2019, Optuna: A Nextgeneration Hyperparameter Optimization Framework Alard, C. 2000, A&AS, 144, 363 Alard, C. & Lupton, R. H. 1998, The Astrophysical Journal, 503, 325–331 Arpit, D., Jastrzebski, ˛ S., Ballas, N., et al. 2017, in Proceedings of Machine Learning Research, Vol. 70, Proceedings of the 34th International Conference on Machine Learning, ed. D. Precup & Y. W. Teh (PMLR), 233–242 Bergstra, J., Bardenet, R., Bengio, Y., & Kégl, B. 2011, in Advances in Neural Information Processing Systems, ed. J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, & K. Weinberger, Vol. 24 (Curran Associates, Inc.) Binder, A., Montavon, G., Bach, S., Müller, K.-R., & Samek, W. 2016, Layerwise Relevance Propagation for Neural Networks with Local Renormalization Layers Burke, D. L., Rykoff, E. S., Allam, S., et al. 2017, The Astronomical Journal, 155, 41 Cabrera-Vives, G., Bolivar, C., Förster, F., et al. 2023, Domain Adaptation via Minimax Entropy for Real/Bogus Classification of Astronomical Alerts Cammarata, N., Carter, S., Goh, G., et al. 2020, Distill, https://distill.pub/2020/circuits Carrasco-Davis, R., Reyes, E., Valenzuela, C., et al. 2021, The Astronomical Journal, 162, 231 Chen, Z., Zhou, W., Sun, G., et al. 2023, TransientViT: A novel CNN - Vision Transformer hybrid real/bogus transient classifier for the Kilodegree Automatic Transient Survey D’Angelo, F. & Fortuin, V. 2021, in Proceedings of the 35th International Conference on Neural Information Processing Systems, NIPS ’21 (Red Hook, NY, USA: Curran Associates Inc.)
Article number, page 14
de Vos, B. D., Jansen, G. E., & Išgum, I. 2023, Scientific Reports, 13, 16875 DESC, L., Aubourg, E., Avestruz, C., et al. 2026, Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration Dieleman, S., Willett, K. W., & Dambre, J. 2015, Monthly Notices of the Royal Astronomical Society, 450, 1441–1459 Elhage, N., Hume, T., Olsson, C., et al. 2022, Toy Models of Superposition Elhage, N., Nanda, N., Olsson, C., et al. 2021, Transformer Circuits Thread, https://transformer-circuits.pub/2021/framework/index.html Fort, S., Hu, H., & Lakshminarayanan, B. 2020, Deep Ensembles: A Loss Landscape Perspective Gal, Y. & Ghahramani, Z. 2016, Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning Goode, S., Cooke, J., Zhang, J., et al. 2022, Monthly Notices of the Royal Astronomical Society, 513, 1742 Graham, M. L., Bellm, E. C., Guy, L. P., et al. 2024, Rubin Observatory Technical Note DMTN-102 Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. 2017, On Calibration of Modern Neural Networks Han, B., Yao, Q., Yu, X., et al. 2018, Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels Hinton, G., Krizhevsky, A., Sutskever, I., & Rachmad, Y. 2012, Advances in Neural Information Processing Systems, 1097 Hosenie, Z., Bloemen, S., Groot, P., et al. 2021, Experimental Astronomy, 51, 319–344 Huang, B., Xiao, K., & Yuan, H. 2022, Photometric calibration methods for wide-field photometric surveys Ivezić, Ž., Kahn, S. M., Tyson, J. A., et al. 2019, ApJ, 873, 111 Kessler, R., Marriner, J., Childress, M., et al. 2015, The Astronomical Journal, 150, 172 Killestein, T. L., Lyman, J., Steeghs, D., et al. 2021, Monthly Notices of the Royal Astronomical Society, 503, 4838–4854 Kim, B., Wattenberg, M., Gilmer, J., et al. 2018, Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV) Kulkarni, S. R. & Kasliwal, M. M. 2009, Transients in the Local Universe Lakshminarayanan, B., Pritzel, A., & Blundell, C. 2017, Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles Lanusse, F., Ma, Q., Li, N., et al. 2017, Monthly Notices of the Royal Astronomical Society, 473, 3895 LeCun, Y., Boser, B., Denker, J. S., et al. 1989, Neural computation, 1, 541 Lintott, C., Schawinski, K., Bamford, S., et al. 2010, Monthly Notices of the Royal Astronomical Society, 410, 166–178 Lipton, Z. C. 2016, CoRR, abs/1606.03490 [1606.03490] Liu, S., Wood-Vasey, W. M., Armstrong, R., et al. 2024, The Astrophysical Journal, 967, 10 Liu, Y., Fan, L., Hu, L., et al. 2025, Astronomy & Astrophysics, 693, A105 LSST Science Collaboration, Abell, P. A., Allison, J., et al. 2009, arXiv e-prints, arXiv:0912.0201 Lundberg, S. & Lee, S.-I. 2017, A Unified Approach to Interpreting Model Predictions MacKay, D. J. 1992, Neural Computation, 4, 448 Makhlouf, K., Turpin, D., Corre, D., et al. 2022, Astronomy & Astrophysics, 664, A81 Matheson, T., Fan, X., Green, R., et al. 2013, Spectroscopy in the Era of LSST McInnes, L., Healy, J., & Melville, J. 2020, UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction Miyazaki, S., Komiyama, Y., Kawanomoto, S., et al. 2017, Publications of the Astronomical Society of Japan, 70, S1 Mong, Y.-L., Ackley, K., Killestein, T. L., et al. 2022, Monthly Notices of the Royal Astronomical Society, 518, 752–762 Möller, A., Peloton, J., Ishida, E. E. O., et al. 2020, Monthly Notices of the Royal Astronomical Society, 501, 3272–3288 Najita, J., Willman, B., Finkbeiner, D. P., et al. 2016, Maximizing Science in the Era of LSST: A Community-Based Study of Needed US Capabilities Narayan, G., Zaidi, T., Soraisam, M. D., et al. 2018, The Astrophysical Journal Supplement Series, 236, 9 Niculescu-Mizil, A. & Caruana, R. 2005, in Proceedings of the 22nd International Conference on Machine Learning, ICML ’05 (New York, NY, USA: Association for Computing Machinery), 625–632 Olah, C., Mordvintsev, A., & Schubert, L. 2017, Distill, https://distill.pub/2017/feature-visualization Padmanabhan, N., Schlegel, D. J., Finkbeiner, D. P., et al. 2008, The Astrophysical Journal, 674, 1217–1233 Reyes, E., Estevez, P. A., Reyes, I., et al. 2018, in 2018 International Joint Conference on Neural Networks (IJCNN) (IEEE), 1–8 Ribeiro, M. T., Singh, S., & Guestrin, C. 2016, "Why Should I Trust You?": Explaining the Predictions of Any Classifier Rolnick, D., Veit, A., Belongie, S., & Shavit, N. 2017, in Advances in Neural Information Processing Systems
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ Sedaghat, N. & Mahabal, A. 2018, Monthly Notices of the Royal Astronomical Society, 476, 5365–5376 Sharkey, L., Chughtai, B., Batson, J., et al. 2025, Open Problems in Mechanistic Interpretability Shrikumar, A., Greenside, P., Shcherbina, A., & Kundaje, A. 2017, Not Just a Black Box: Learning Important Features Through Propagating Activation Differences Sánchez, B. O., Kessler, R., Scolnic, D., et al. 2022, The Astrophysical Journal, 934, 96 van der Maaten, L. & Hinton, G. 2008, Journal of Machine Learning Research, 9, 2579 Vujčić, V., Srećković, V. A., Babarogić, S., & Aleksić, J. 2025, Contributions of the Astronomical Observatory Skalnate Pleso, 55, 95 Wold, S., Esbensen, K., & Geladi, P. 1987, Chemometrics and Intelligent Laboratory Systems, 2, 37, proceedings of the Multivariate Statistical Workshop for Geologists and Geochemists Wu, J. F. & Walmsley, M. 2025, Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders Zeiler, M. D. & Fergus, R. 2013, Visualizing and Understanding Convolutional Networks
Article number, page 15
A&A proofs: manuscript no. main
(a) Injection Catalog
Appendix A: Injection magnitude
g
6000
Number of Sources
We compare the magnitude distributions of injected sources, recovered injections, and real detections with positive flux. As expected, the recovered injections are systematically brighter than the full injected population, reflecting the detection efficiency of the pipeline at faint magnitudes. The positive-flux real detections have a broadly comparable magnitude distribution, and the resulting sample remains approximately balanced between real and recovered injected detections.
r
i
z
N = 86,973 = 24.76 = 1.94
y
5000 4000 3000 2000 1000 0
(b) Injections After Pipeline g
Number of Sources
5000
r
i
z
y
N = 51,520 = 23.91 = 1.44 Recov. = 59.2%
4000 3000 2000 1000 0
Number of Sources
(c) Real Sources After Pipeline g
r
i
16
18
z
N = 22,558 = 24.86 = 1.88
y
2000 1500 1000 500 0
14
20
22
24
AB Magnitude
26
28
Fig. A.1: Magnitude distributions for (a) injected sources, (b) recovered injections after the pipeline, and (c) real sources with positive flux. In each panel we report the number of sources N and the best-fitting Gaussian parameters (µ, σ) describing the distribution.
Article number, page 16
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
Appendix B: Light curve labeling tool
Appendix D: WSL formalism
Examples of the actual labeling tool are available in the Zenodo supplementary material (DOI: 10.5281/zenodo.18434676).
D.1. Co-teaching
Formally, let Bt = {(xi , ỹi )}i∈It denote the mini-batch at iteration (or epoch) t, where It is the set of indices of the samples in the batch and |Bt | = |It |. The per-sample losses of network 1 and network 2 on this batch are defined by their Binary Cross Entropy (BCE) as ℓi(1) (t) = BCE fθ1 (xi ), ỹi , ℓi(2) (t) = BCE fθ2 (xi ), ỹi , i ∈ It . The forget rate is defined as ! t r(t) = min r · , r , Tk
(D.1)
where T k is a predefined epoch after which the forget rate remains constant. Given the forget rate, the fraction of samples kept in the batch at iteration t is 1 − r(t). We set kt = (1 − r(t)) |Bt | = (1 − r(t)) |It | .
Fig. B.1: Representative examples of a bogus candidate (top) and a supernova-like transient (bottom). Left: multi-band light curves in detected bands as a function of MJD. Right: examples of associated cutouts shown. For each object from top to bottom: coadd image, mean of the science for this band and mean of the difference image for the band.
For network 1, we sort the indices in It according to the increasing order of its loss: (1) (1) ℓi(1) t (t) ≤ ℓ t (t) ≤ · · · ≤ ℓ t (t), i i 1
2
|It |
where (it1 , . . . , it|It | ) is a permutation of It . The set of reliable samples selected by network 1 at iteration t is then defined as X ℓi(1) (t). R1,t = {it1 , . . . , itkt } = arg min S ⊂It |S |=kt i∈S
Appendix C: Training configurations For reproducibility, we summarize below the training hyperparameters used for each uncertainty-estimation or robust-training method considered in this work. Unless stated otherwise, all settings follow the default values described in the main text, and we report only the parameters specific to each method (number of models or forward passes, repulsion strength and type, dropout probabilities, and the co-teaching forget-rate schedules). For the forget rates, the reported values correspond to the {Baseline, 15%, 25%, 35%} sets.
The Co-teaching losses are then given by 1 X L(1) BCE fθ1 (xi ), ỹi , co-teach (t) = |R | 2,t i∈R 2,t
1 L(2) co-teach (t) = |R | 1,t i∈R
X
(D.2) BCE fθ2 (xi ), ỹi ,
1,t
where R1,t and R2,t denote the sets of reliable samples selected at iteration t by networks 1 and 2, respectively. D.2. Stochastic Co-teaching
– Ensemble: num_models = 50. – Repulsive ensemble: num_models = 50, λrep = 0.1, repulsion_type = "prediction", T = 1.0. – MC Dropout: fwd_passes = 50, pdense = 0.25, pconv = 0.125. – Co-teaching: forget_rate ∈ {0.06, 0.16, 0.26, 0.36}. – Asym. Co-teaching: (forget_rate0 , forget_rate1 ) ∈ {(0.05, 0.01), (0.15, 0.01), (0.25, 0.01), (0.35, 0.01)}. We select the forget-rate schedules using prior knowledge estimated from the label noise inferred by the standard training method. In particular, we use the observed false-positive and false-negative fractions on the training set as approximate indicators of mislabeling. This suggests a noise level of about 5% for the survey (baseline) class and about 1% for the injected class; the reported forget rates are chosen to reflect these estimates.
Standard Co-teaching requires the user to specify a forget rate in advance. Stochastic Co-teaching (de Vos et al. 2023) replaces this fixed rate with a stochastic rejection rule. At iteration t, each network samples a threshold zt ∼ Beta(α, β), where α and β control the expected strictness of sample rejection. For a mini-batch Bt = {(xi , ỹi )}i∈It , network j ∈ {1, 2} outputs a probability p̂(i j) (t) = σ fθ j (xi ) . The probability assigned to the observed training label is ( j) if ỹi = 1, p̂i (t), p(i j) (t) = 1 − p̂( j) (t), if ỹ = 0. i i Article number, page 17
A&A proofs: manuscript no. main
Samples with p(i j) (t) ≤ zt are rejected, and the remaining samples are passed to the peer network for updating, as in standard Co-teaching. This stochastic rule removes the need to fix a single deterministic forget rate, but in our experiments it remained difficult to tune and often led to unstable class-wise rejection behavior. For this reason, we report it only as an auxiliary baseline and do not retain it among the main methods.
Appendix E: Signal-to-noise ratio by class
photometric SNR and more on heterogeneous artifact morphology. Correlations for false positives and false negatives are small and unstable in sign, indicating that misclassifications are not explained by SNR alone. ρ SNR MC Dropout Deep Ensemble Repulsive Ensemble Our method
TP −0.368 −0.434 −0.420 −0.549
FP 0.089 −0.028 0.084 −0.038
TN −0.157 −0.192 −0.162 −0.166
FN 0.063 −0.105 0.123 0.108
E.1. Signal-to-noise ratio distribution
The SNR is computed as: SNR =
|psfFlux| , psfFluxErr
where psfFlux is the PSF-fitted flux measured on the difference image and psfFluxErr is its associated uncertainty. Figure E.1 shows that the SNR distributions differ substantially between the two training classes. In particular, detections with SNR ≤ 5 are nearly absent from the injected class because such faint events are less likely to be recovered by the detection pipeline. The injected class also extends to systematically higher SNR values than the survey class. This class asymmetry is expected from the construction of the dataset, but it should be kept in mind when interpreting the relationship between SNR and uncertainty.
Table E.1: Spearman rank correlation coefficients (ρ) between predicted uncertainty and signal-to-noise ratio (SNR) for each predictive class, True Positive (TP), False Positive (FP), True Negative (TN) and False Negative (FN).
Appendix F: Uncertainty Quantification formalism F.1. Deep Ensemble
During inference, the ensemble prediction is obtained by averaging the individual model probabilities: M
p̄(x) =
1 X σ( f (x; θ̂m )), M m=1
(F.1)
where σ(·) denotes the sigmoid activation function, and θ̂m the trained weights of the m-th model. The final predicted class label is then given by: ŷensemble = I[ p̄(x) > 0.5],
(F.2)
with I[·] representing the indicator function. Beyond improving prediction stability, the ensemble also facilitates uncertainty quantification. The predictive uncertainty for a given input x is estimated as the variance of the individual model probabilities: M
Varens (x) =
1 X σ( fm (x)) − p̄(x) 2 . M m=1
(F.3)
This measure reflects the degree of agreement among ensemble members, with higher variance indicating greater predictive uncertainty. F.2. Repulsive ensemble
Fig. E.1: Signal-to-noise ratio analysis of the rc2_subset dataset by class. Panel (a) shows the SNR distribution of the injected class, panel (b) the SNR distribution of the survey class, panel (c) the corresponding cumulative distribution functions, and panel (d) the distribution split below and above SNR = 5.
E.2. Signal-Noise-Ratio correlation with Uncertainty Quantification per class
Table E.1 reports the Spearman correlation between predicted uncertainty and SNR on the manually curated evaluation set, stratified by predictive outcome. The strongest negative correlation is observed for true positives, indicating that uncertainty for correctly identified transient detections is strongly linked to signal quality. For true negatives, the correlation is weaker, suggesting that uncertainty on bogus detections depends less directly on Article number, page 18
Operationally, the interaction term of the repulsive ensemble is encoded through an update of the form ϕ( fit ) = ∇ fit log p( fit |D) − R ∇ fit k( fit , f jt ) M (F.4) j=1 , where the first term promotes agreement with the data while the second term introduces repulsion via a kernel k(·, ·) measuring similarity between ensemble functions. The operator R(·) aggregates the repulsive contributions across the other ensemble members, pushing fi away from functions that are already represented in the ensemble. This repulsivity is motivated by the desire of the uncertainty quantification method to cover a larger area of the loss landscape and avoid ensembles that are accurate but overly concentrated around a single mode. At inference time, the procedure remains the same as for Deep Ensembles: predictions are obtained by aggregating the ensemble outputs, and uncertainty can be quantified from the dispersion across members.
Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ
F.3. MC Dropout
Let θ̂ be the trained weights. Given N dropout masks, we sample the dropout masks zn ∼ q(z) (Bernoulli masks based on the dropout rate of the training), so that for each input x∗ we perform N forward passes with these different masks. For pass n, let fn (x∗ ; θ̂, zn ) denote the pre-sigmoid logit of the network. The corresponding predictive probability is pn (x∗ ) = p(y∗ = 1 | x∗ , θ̂, zn ) = σ fn (x∗ ; θ̂, zn ) ,
(F.5)
with σ the sigmoid function of the last layer. The MC dropout predictive probability is then approximated by the first moment: p̄(x∗ ) =
N N 1 X 1 X pn (x∗ ) = σ fn (x∗ ; θ̂, zn ) . N n=1 N n=1
(F.6)
The (epistemic) predictive variance can be approximated by the second centered moment: N
VarMCdropout (x∗ ) =
1 X pn (x∗ ) − p̄(x∗ ) 2 . N n=1
(F.7)
Fig. G.2: Static UMAP representation of the latent space with signal-to-noise ratio shown as a color overlay. Values are capped at SNR = 15. Symbols represent model predictions. The manually curated evaluation set is shown; accordingly, no displayed sources have SNR < 5.
Appendix G: UMAP representation The UMAP projection can be overlaid with additional quantities such as predictive uncertainty and signal-to-noise ratio. These overlays provide a qualitative view of how uncertainty and SNR are distributed across the learned representation. Figure G.1 shows that high-uncertainty samples concentrate near the interface between the main regions of the embedding. Figure G.2 shows that SNR only partially follows the same structure, with a clearer trend in the transient-like region than in the bogus region. These visual patterns are consistent with the quantitative results reported in Section 7.2.
Fig. G.1: Static UMAP representation of the latent space with predictive uncertainty shown as a color overlay. Symbols represent model predictions. The manually curated evaluation set is shown, and uncertainty is computed using the method introduced in Section 5.5. The displayed uncertainty is the per-sample standard deviation of the predicted probability.
Appendix H: Cumulative light-curve counts Article number, page 19
A&A proofs: manuscript no. main
Threshold (%)
All LCs
SN
other transients or variables
Transients (all)
Bogus
≤5 ≤ 10 ≤ 15 ≤ 20 ≤ 25 ≤ 30 ≤ 35 ≤ 40 ≤ 45 ≤ 50 ≤ 55 ≤ 60 ≤ 65 ≤ 70 ≤ 75 ≤ 80 ≤ 85 ≤ 90 ≤ 95 ≤ 100
174 190 207 215 222 223 227 231 232 232 232 233 238 243 244 248 256 265 279 306
33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 33 (100.00%) 30 ( 90.91%) 28 ( 84.85%) 28 ( 84.85%) 26 ( 78.79%) 22 ( 66.67%) 19 ( 57.58%) 11 ( 33.33%) 3 ( 9.09%)
41 (100.00%) 41 (100.00%) 41 (100.00%) 41 (100.00%) 41 (100.00%) 41 (100.00%) 41 (100.00%) 40 ( 97.56%) 40 ( 97.56%) 40 ( 97.56%) 40 ( 97.56%) 40 ( 97.56%) 39 ( 95.12%) 35 ( 85.37%) 34 ( 82.93%) 33 ( 80.49%) 28 ( 68.29%) 23 ( 56.10%) 16 ( 39.02%) 8 ( 19.51%)
74 (100.00%) 74 (100.00%) 74 (100.00%) 74 (100.00%) 74 (100.00%) 74 (100.00%) 74 (100.00%) 73 ( 98.65%) 73 ( 98.65%) 73 ( 98.65%) 73 ( 98.65%) 73 ( 98.65%) 69 ( 93.24%) 63 ( 85.14%) 62 ( 83.78%) 59 ( 79.73%) 50 ( 67.57%) 42 ( 56.76%) 27 ( 36.49%) 11 ( 14.86%)
174 ( 75.00%) 190 ( 81.90%) 207 ( 89.22%) 215 ( 92.67%) 222 ( 95.69%) 223 ( 96.12%) 227 ( 97.84%) 230 ( 99.14%) 231 ( 99.57%) 231 ( 99.57%) 231 ( 99.57%) 232 (100.00%) 232 (100.00%) 232 (100.00%) 232 (100.00%) 232 (100.00%) 232 (100.00%) 232 (100.00%) 232 (100.00%) 232 (100.00%)
Table H.1: Cumulative number of light-curves below a predicted-transient threshold (All LCs, Bogus) and above the bin lower edge for transient subclasses (SN-like, other transients or variables, Transients all). Percentages are cumulative within each class/subclass.
Article number, page 20