ConceptioArchivearXiv CS
arXiv CSopen access

A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

A welding penetration prediction model for laser welding process based on self-supervised learning using physicsinformed neural networks Sen Lia, Xiaoying Liua, Xiaojian Xua, Chendong Shaoa, Yaqi Wanga, Ling Lana,b, Xinhua Tanga, Haichao Cuia,* a

Shanghai Key Laboratory of Materials Laser Processing and Modification, School of Materials

Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, PR China. b

Shanghai Shipbuilding Technology Research Institute, Shanghai 200032, China.

Abstract: The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in achieving defect-free welded joints. Accurate prediction of the penetration state is therefore essential for ensuring weld quality. To this end, this paper introduces SimPhysNet, a novel algorithm that achieves high classification accuracy in laser welding penetration prediction using only a limited number of labelled images. This approach effectively overcomes the limitations of supervised learning classification algorithms, which are hindered in industrial applications by their dependence on extensive, high-quality labelled data. The core of SimPhysNet is a unique selfsupervised learning paradigm that embeds physical priors into a contrastive learning framework. By incorporating a physics-informed neural network (PINN), the model is guided to extract physically meaningful features of the molten pool and keyhole from a large set of unlabelled data, while three image augmentation tasks further enhance its generalization capabilities. Subsequently, a few-shot learning strategy, based on prototypical networks, enables robust classification by constructing class representations from a minimal set of labelled images. Experimental results demonstrate that SimPhysNet achieves a classification accuracy of 96.06% using only 200 labelled images (approximately 5% of the total labelled dataset), which is comparable to the performance of conventional supervised learning algorithms that utilize the entire labelled dataset. This work presents a new, efficient, and highly accurate method, providing the way for the intelligent automation of laser welding. Keywords: Laser welding; Penetration prediction; Self-supervised learning; Physics-informed neural network

*

Corresponding author: [email protected]

1. Introduction In recent years, laser welding technology has been extensively applied across various industries such as nuclear energy, shipbuilding, aerospace, defense equipment, automotive manufacturing, and railway vehicle production[1]. Compared to conventional welding techniques, laser welding offers significant advantages, including low thermal input, high weld depth-to-width ratios, rapid welding speed, and a high level of automation. Consequently, it is regarded as one of the most promising advanced joining technologies of the 21st century. However, achieving high-quality weld joints in laser welding requires precise control and optimization of process parameters. Therefore, the development of a reliable model for the entire laser deep penetration welding process would enable real-time adjustment of welding parameters to ensure full penetration and superior joint quality. During the laser welding process, a substantial amount of signals are generated, primarily including acoustic and optical signals [2]. In recent years, optical signals have attracted growing attention from researchers due to their rich information content, ease of acquisition, and strong interpretability. For instance, Kang et al. [3] developed a deep learning model using optical emission spectroscopy (OES) to enable real-time prediction of the cross-sectional geometry of welds in aluminum/copper laser dissimilar welding (LDW) processes, thereby facilitating an intuitive assessment of welding quality. Similarly, Li et al. [4] designed an online monitoring system for laser weld geometry based on a cost-effective charge-coupled device (CCD) camera and a multi-task convolutional neural network, achieving real-time prediction of weld depth and width with high accuracy. However, despite their impressive accuracy, these deep learning models are predominantly supervised, creating a significant barrier to their widespread industrial application. The deployment of such data-driven models in industrial settings is frequently hindered by a significant practical challenge: the prohibitive cost associated with acquiring large-scale, highquality labeled datasets. In the context of laser welding, establishing the ground truth for penetration state is fundamentally a labor-intensive, post-process activity. Each label necessitates careful inspection of the completed weld bead—typically by examining its backside—to confirm full penetration. When this verification process must be scaled to thousands or tens of thousands of samples to adequately cover the diverse range of process parameters and material conditions, it becomes a critical bottleneck. This reliance on extensive post-process verification impedes the rapid development and iteration of models and limits their scalability in dynamic production environments. Therefore, developing a methodology that can achieve high prediction accuracy with minimal reliance on labeled data is of paramount importance for the intelligent automation of laser welding. To address these challenges, researchers have proposed Few-Shot learning methods. Few-Shot learning aims to achieve efficient learning with a limited number of samples, thereby reducing reliance on large labeled datasets. One of the most significant approach in this field is metric-based learning [5]. This approach involves training a model to map input samples into a low-dimensional embedding space, where samples from the same class are closely clustered, while those from different classes are distinctly separated. During inference, the class of a query sample is determined by calculating its distance to samples in the support set within this embedding space. Raffin et al. [6] found that Prototypical Networks exhibit outstanding performance under conditions of limited

data availability, achieving an accuracy of 94.4% in quality monitoring tasks related to data scarcity in hook-type winding laser welding. This provides a data-efficient solution for industrial visual inspection. Addressing the challenge of limited small-sample data in industrial weld seam inspection, Zhu et al. [7] proposed a metric learning approach based on a distance metric classification mechanism, which enhances detection accuracy under extremely constrained data conditions. Liu et al. [8] introduced a novel algorithm called FS-Classifier, which employs unsupervised learning for pre-training the feature extractor and integrates it with an optimized prototypical network for small-sample classification. This method achieved high-precision classification of gas tungsten arc welding defects using small-sample datasets, further demonstrating the feasibility of combining unsupervised pre-training with prototypical network fine-tuning. While these works demonstrate the potential of combining unsupervised pre-training with few-shot learning, their pre-training objectives are purely data-driven. For complex physical processes like laser welding, such approaches may fail to capture the underlying causal relationships, instead learning superficial visual correlations. This limits their ability to generalize robustly from a feature space that is not grounded in the process physics. In recent years, integrating physical information into deep learning networks has attracted significant attention. Physics-informed neural networks (PINNs) incorporate physical laws as prior knowledge into the neural network training process. This approach ensures that while the model learns from data, it remains consistent with fundamental physical principles. Consequently, PINNs can leverage physical constraints to guide the learning of accurate data features using limited experimental datasets, aligning with the goal of achieving high classification accuracy with fewer samples. Chen et al. [9] proposed a hybrid CNN approach that integrates physical prior knowledge by converting welding-related physical laws into image features that inform the deep learning process, thereby significantly improving the accuracy of variable polarity plasma arc (VPPA) welding quality prediction. Similarly, Lu et al.[10] developed a physics-informed CNN-LSTM (PICNN-LSTM) for real-time monitoring of laser filler-wire welding. Their architecture couples spatiotemporal molten pool features with an analytical heat transfer model to infer thermal-related variables. By employing physics-aware activation functions to constrain these parameters within meaningful ranges, they achieved robust penetration depth prediction even when trained on only 10% of the dataset. Li et al. [11] introduced an improved M-PINN for predicting the fatigue life of double-sided welds by incorporating implicit physical models within activation functions and explicit physical constraints into loss functions, which significantly enhance prediction accuracy under conditions of limited data availability. Although PINNs have shown success in supervised contexts for solving forward and inverse problems, their potential as a regularizer within selfsupervised feature learning frameworks remains largely untapped. Existing studies typically use PINNs to approximate a specific solution. In contrast, our work re-purposes this paradigm to enforce physical consistency on a high-dimensional feature representation, a fundamentally different and novel application. To address the aforementioned challenges, this paper introduces SimPhysNet, a novel framework designed to achieve high-accuracy penetration prediction from a limited set of labeled images. The primary contributions of this work are summarized as follows: 1. A novel fusion of physics and self-supervision for feature learning: Unlike conventional self-

supervised methods that rely solely on data augmentation, our framework pioneers the integration of physical priors, governed by partial differential equations, are embedded directly into a contrastive learning loss function. This approach is distinct from prior works, as it utilizes fundamental physical laws as an implicit supervisory signal to guide the feature extractor in learning physically meaningful representations from unlabeled data, thereby ensuring that the learned features are not merely correlational but are grounded in the underlying process dynamics. 2. Synergistic integration of pre-training and few-shot learning: The model architecture is structured in two stages. An initial pre-training stage leverages unlabeled data to establish a robust and generalized feature space. Subsequently, a fine-tuning stage employs a prototypical networkbased few-shot learning strategy to achieve high classification accuracy with minimal labeled data. This synergistic design effectively mitigates the overfitting problem commonly encountered in datalimited scenarios. 3. Application-tailored image augmentation strategies: Three supervised pretext tasks— rotation prediction, Gaussian blur level prediction, and random cropping localization—are designed and integrated into the pre-training stage. These tasks compel the network to learn hierarchical features, from global molten pool morphology to fine-grained local details, which are critical for discriminating between penetration states and enhancing the model's overall generalization capabilities.

2. Materials and methods 2.1 Experimental setup The experimental setup primarily comprises three key components: the laser welding module, the robotic motion control system, and the coaxial image acquisition module. The laser welding module consists of a fiber laser welding machine, a water-cooling unit, and a laser welding head. During the welding process, the workpiece remains stationary, while a Yaskawa robot arm precisely controls the movement of the welding head. Prior to initiating the welding procedure, the 808 nm auxiliary laser light source and the focal length of the coaxial camera are carefully adjusted to ensure the acquisition of a clear, stable, and well-illuminated image. Throughout the welding operation, images are captured and transmitted in real-time to a connected PC for data storage. The output power of the fiber laser is controlled by the PC, whereas other critical welding parameters, such as welding speed and defocus distance, are controlled via the robotic system. The laser employed in this study is the YLS-20000 fiber laser with a spot-ring configuration, capable of delivering a maximum output power of 20 kW. The ring laser has a beam parameter product (BPP) of 15.377 𝑚𝑚𝑚𝑚 ∙ 𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚, with a focused spot diameter of 596.364 µm and a rayleigh length of 26.25 mm. The dot laser exhibits a BPP of 3.385 𝑚𝑚𝑚𝑚 ∙ 𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚, a focused spot diameter of 197.535 µm, and a rayleigh length of 2.88 mm. A schematic illustration of the laser welding setup is shown in Fig. 1. In the experimental configuration, a coaxial complementary metal oxide semiconductor (CMOS) camera was utilized to capture top-down images of the molten pool. An 808 nm auxiliary light source was integrated into the imaging system. During image acquisition,

both the CMOS camera and the auxiliary light source were synchronously triggered by the control PC, ensuring stable illumination and improved image quality. The imaging system employs a 1/4″ monochrome CMOS sensor positioned at a 0.0° angle. An 808 nm narrow bandpass filter (NBF) with a full width at half maximum (FWHM) of 30 nm is integrated into the setup. Key camera parameters were set with an exposure time of 36 μs, a gamma value of 1.2, and a black level of 1.0%. Computer

Coaxial CMOS

20kW Fiber Laser

Laser head

Auxiliary light source

Welding direction

Robot

Fig. 1 The schematic of the welding system. 2.2 Experimental method The welding material utilized is COST-E, with dimensions of 100 𝑚𝑚𝑚𝑚 × 80 𝑚𝑚𝑚𝑚 (length × width) as illustrated in Fig. 2. The specimens are provided in five different thicknesses: 3.0 mm, 4.0 mm, 6.0 mm, 8.0 mm, and 10.0 mm, with chemical compositions detailed in Tab. 1. Prior to welding, the specimens are thoroughly cleaned using anhydrous alcohol to remove any contaminants that could potentially affect the weld joint quality. During the welding process, argon gas is used as a shielding gas to protect the weld area. The argon is supplied through three parallel copper tubes positioned behind the weld seam, with a flow rate of 25 𝐿𝐿/𝑚𝑚𝑚𝑚𝑚𝑚 and a purity of 99.99%. The welding parameters are summarized in Tab. 2. By varying the plate thickness 𝑇𝑇[𝑚𝑚𝑚𝑚], welding speed 𝑉𝑉𝑤𝑤 [𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚], spot laser power 𝑃𝑃𝑑𝑑 [𝑘𝑘𝑘𝑘], and defocus amount 𝐷𝐷𝑓𝑓 [𝑚𝑚𝑚𝑚], various weld

joint conditions are obtained.

Tab. 1 Chemical composition of COST-E(wt.%) [12]. COST-E

C

Mn

Si

Ni

Cr

Mo

W

Nb

V

Fe

(wt.%)

0.12

0.45

0.1

0.75

10.5

1.0

1.0

0.05

0.2

Bal.

Tab. 2 Welding parameters. 𝐷𝐷𝑓𝑓 (mm) 𝑃𝑃𝑑𝑑 (kW) 𝑉𝑉𝑤𝑤 (m/min)

Parameter Value

[3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0]

𝑍

𝑇𝑇 (mm)

[0.0, -5.0] [1.0, 1.5, 2.0] [3.0, 4.0, 6.0, 8.0, 10.0]

𝑋𝑋

𝑌𝑌

Fig. 2 Schematics of the laser welding experiment. 2.3 Dataset acquisition Throughout the welding process, a considerable number of molten pool images, representing both fully penetrated and non-penetrated welding states, were captured in real-time using the experimental setup described in Section 2.1. Fig. 3 illustrates the front and rear surfaces as well as 𝑁𝑁

𝐿𝐿 , the cross-section of the weld after welding. The labeled dataset is denoted as 𝐷𝐷𝐿𝐿 = {𝑥𝑥𝑖𝑖,𝐿𝐿 , 𝑦𝑦𝑖𝑖,𝐿𝐿 }𝑖𝑖=1 𝑁𝑁𝑈𝑈 and the unlabeled dataset is denoted as 𝐷𝐷𝑈𝑈 = {𝑥𝑥𝑖𝑖,𝑈𝑈 }𝑖𝑖=1 , where N denotes the total number of images in each respective dataset. Here, 𝑥𝑥𝑖𝑖,𝐿𝐿 and 𝑥𝑥𝑖𝑖,𝑈𝑈 denote the molten pool images, and 𝑥𝑥𝑖𝑖,𝐿𝐿 ∈

{1, … , 𝐾𝐾} represent the corresponding annotation labels, with 𝐾𝐾 indicating the total number of categories. In the context of the laser welding penetration dataset, 𝐾𝐾 = 2. The unlabeled images were collected from routine welding experiments. The composition of the dataset is illustrated in Tab. 3, and representative samples are depicted in Fig. 4. Additionally, the following formula is used to calculate the model's accuracy:

𝑇𝑇𝑇𝑇 + 𝑇𝑇𝑇𝑇 (1) × 100% 𝑇𝑇𝑇𝑇 + 𝑇𝑇𝑇𝑇 + 𝐹𝐹𝐹𝐹 + 𝐹𝐹𝐹𝐹 where TP (true positive), FP (false positive), FN (false negative), and TN (true negative) represent the number of correctly and incorrectly classified samples. 𝐴𝐴𝐴𝐴𝐴𝐴 =

Penetrated Topside Backside

Non-penetrated 2mm

2mm

Fig. 3 Schematic diagrams illustrating the topside and backside views of the weld.

Tab. 3 Composition of the datasets. ("-" denotes the absence of the corresponding content) Datasets 𝐷𝐷𝐿𝐿

Classes

Train set

Validation set

Penetrated (Class 0)

1400

600

Non-penetrated (Class 1)

1400

600

-

8000

-

𝐷𝐷𝑈𝑈

𝑃𝑃𝑝𝑝 : 3.0 𝑘𝑘𝑘𝑘, 𝑉𝑉𝑤𝑤 : 2.0𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚 𝑃𝑃𝑝𝑝 : 7.0 𝑘𝑘𝑘𝑘,𝑉𝑉𝑤𝑤 : 2.0𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚 𝑃𝑃𝑝𝑝 : 9.0 𝑘𝑘𝑘𝑘, 𝑉𝑉𝑤𝑤 : 2.0𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚 𝐷𝐷𝑓𝑓 : 0.0 𝑚𝑚𝑚𝑚, 𝑇𝑇: 4.0 𝑚𝑚𝑚𝑚 𝐷𝐷𝑓𝑓 : −5.0 𝑚𝑚𝑚𝑚, 𝑇𝑇: 8.0 𝑚𝑚𝑚𝑚 𝐷𝐷𝑓𝑓 : −5.0 𝑚𝑚𝑚𝑚, 𝑇𝑇: 10.0 𝑚𝑚𝑚𝑚

Class 0 “Penetrated”

Class 1 “non-penetrated”

𝑃𝑃𝑝𝑝 : 5.0 𝑘𝑘𝑘𝑘, 𝑉𝑉𝑤𝑤 : 1.0𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚 𝑃𝑃𝑝𝑝 : 7.0 𝑘𝑘𝑘𝑘, 𝑉𝑉𝑤𝑤 : 2.0𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚 𝑃𝑃𝑝𝑝 : 2.0 𝑘𝑘𝑘𝑘, 𝑉𝑉𝑤𝑤 : 1.0𝑚𝑚/𝑚𝑚𝑚𝑚𝑚𝑚 𝐷𝐷𝑓𝑓 : −5.0 𝑚𝑚𝑚𝑚, 𝑇𝑇: 8.0 𝑚𝑚𝑚𝑚 𝐷𝐷𝑓𝑓 : −5.0 𝑚𝑚𝑚𝑚, 𝑇𝑇: 10.0 𝑚𝑚𝑚𝑚 𝐷𝐷𝑓𝑓 : 0.0 𝑚𝑚𝑚𝑚, 𝑇𝑇: 6.0 𝑚𝑚𝑚𝑚

Labeled dataset 𝑫𝑳

Unlabeled dataset 𝑫𝑼

Fig. 4 The samples of labeled dataset 𝐷𝐷𝐿𝐿 and unlabeled dataset 𝐷𝐷𝑈𝑈 .

3. Details of the network

As discussed in the first section, a laser molten pool classification model with strong generalization performance and high accuracy typically depends on large-scale, high-quality datasets. However, the annotation of such datasets is both labor-intensive and time-consuming. To address this issue, this paper proposes an unsupervised classification network, termed SimPhysNet, which achieves substantial classification accuracy using only a limited number of labeled samples. Fig. 5 presents the detailed architecture of the proposed network, which consists of two key components: fist, a self-supervised learning strategy is applied to train the backbone network, enabling it to efficiently extract informative features from molten pool images. Second, a few-shot classification method is utilized to train the feature classifier. Notably, to enhance the effectiveness of the self-supervised learning phase, a PINN is incorporated into the training process. The fundamental hypothesis underpinning the SimPhysNet architecture is that for complex physical systems such as laser welding, optimal feature representations must be consistent with governing physical principles. A purely data-driven model risks learning spurious correlations from image data, particularly when labeled samples are scarce. Consequently, the core methodological innovation of SimPhysNet lies in the fusion of a data-driven contrastive learning objective with physics-based constraints imposed by a PINN module. The PINN functions as a form of regularization, ensuring that the feature embedding space aligns with the laws of heat conduction and energy distribution, thereby enhancing the model's robustness and generalization performance without requiring additional annotated data.

𝒕𝟏

𝒕𝟐

Update weights 𝑳𝑷𝑰𝑵𝑵

𝑷(𝒙, 𝒕, 𝜽)

Transfer weights

Transfer weights

𝑴(𝑬(𝒙, 𝒕, 𝜽))

𝑳𝒄𝒐𝒏 + 𝑳𝒂𝒖𝒈

(a) Pre-training stage Update weights 𝑳𝑷𝑰𝑵𝑵

𝑷(𝒙, 𝒕, 𝜽)

fine-tuning stage

𝒅

𝑬(𝒙, 𝒕, 𝜽)

𝒑

𝑳𝑭𝑺𝑳

(b) Fine-tuning stage Support set

Query set

Predict output

Unlabeled dataset True label

Welding time

Feature vector

Fig. 5 The overall architecture of the proposed network in a 2-way 1-shot paradigm. The model embedding backbone is denoted by 𝑀𝑀(𝐸𝐸(𝑥𝑥, 𝑡𝑡, 𝜃𝜃)) and euclidean distance function is denoted by 𝑑𝑑. 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐 , 𝐿𝐿𝑎𝑎𝑎𝑎𝑎𝑎 and 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 denote the contrastive loss, image task loss and PINN loss in the pretraining stage, 𝐿𝐿𝐹𝐹𝐹𝐹𝐹𝐹 denote the few-shot loss in the fine-tuning stage. 𝑝𝑝 denotes the predicted logits/output from the model. 3.1 Pre-training stage 3.1.1 Contrastive learning Self-supervised learning methods typically operate on large volumes of unlabeled images. These methods do not rely on manually annotated category labels; instead, they leverage the data itself as supervisory signals to learn feature representations, which can subsequently be utilized in downstream tasks. Specifically, the contrastive learning approach adopted in this study aims to train an encoder that generates similar representations for instances within the same class, while maximizing the dissimilarity between the representations of different classes. By adhering to this principle, the model is capable of extracting meaningful features from extensive unlabeled datasets, thereby significantly enhancing its generalization performance and robustness. Fig. 6 illustrates the core implementation framework of the Simsiam algorithm[13]. Prior to being fed into the network, the molten pool image x undergoes random data augmentation, resulting in two transformed views, 𝑥𝑥1 and 𝑥𝑥2 . These augmented images are then processed through the encoder network 𝐸𝐸(∙). The encoder network comprises a backbone architecture and a projection head implemented as a MLP, where ResNet-18[14] is employed as the backbone network in this work. Following the 𝐸𝐸1 (∙) encoder, an additional MLP-based prediction head 𝑀𝑀(∙) is introduced. This prediction network maps the high-dimensional feature representation generated by 𝐸𝐸1 (∙) to align with the feature space produced by 𝐸𝐸2 (∙). Consequently, the transformed outputs for the two augmented images are expressed as 𝑟𝑟1 = 𝑀𝑀(𝐸𝐸1 (𝑥𝑥1 )) and 𝑘𝑘2 = 𝐸𝐸2 (𝑥𝑥2 ) . Evidently, the optimization objective of the network is to maximize the similarity between 𝑞𝑞1 and 𝑘𝑘2 . To achieve

this, the model minimizes their negative cosine similarity: 𝑞𝑞2 𝑘𝑘1 ∙ 𝐿𝐿𝐷𝐷 (𝑘𝑘1 , 𝑞𝑞2 ) = − 2 ‖𝑘𝑘1 ‖ ‖𝑞𝑞2 ‖2

(2)

Where ‖∙‖2 denotes the L2 norm. Furthermore, the loss function for contrastive learning can be defined as: 1 𝑞𝑞2 𝑘𝑘2 𝑞𝑞1 1 𝑘𝑘1 (3) ∙ + ∙ � 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐 = 𝐿𝐿𝐷𝐷 (𝑘𝑘1 , 𝑞𝑞2 ) + 𝐿𝐿𝐷𝐷 (𝑘𝑘2 , 𝑞𝑞1 ) = − � 2 2 2 2 ‖ ‖𝑞𝑞 ‖ ‖ ‖𝑞𝑞 2‖𝑘𝑘1 2‖𝑘𝑘2 2 2 2 1‖

An important optimization of this contrastive learning approach is the application of the stopgradient (stopgrad) operation. This enables the network to achieve effective learning outcomes without relying on negative sample pairs or momentum encoders. Accordingly, Equations (1) and (2) are reformulated as follows: (4) 𝐿𝐿𝐷𝐷 (𝑘𝑘1 , 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠(𝑞𝑞2 ))

1 1 (5) 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐 = 𝐿𝐿𝐷𝐷 (𝑘𝑘1 , 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠(𝑞𝑞2 )) + 𝐿𝐿𝐷𝐷 (𝑘𝑘2 , 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠(𝑞𝑞1 )) 2 2 Here, 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠(∙) signifies the stop-gradient operation. Equation (4) defines the loss function of the Simsiam algorithm. Contrastive Loss 𝒌𝒌𝟏

Predictor – ReLU +BN (3-layer)

𝒌𝒌𝟐

𝒒𝟏

Projector – ReLU +BN (3-layer)

𝒒𝟐

Stop-gradient

Encoder – ResNet-18 𝒙𝟏

KNN Evaluation

𝒙𝟐

Feature vector

Normalization – z-score Augmentations – ColorJitter | Grayscale | HFlip | VFlip Rotate

Blur

Crop

KNN (2) Training

Train

Evaluation

Image 𝒙

Fig. 6 Contrastive Learning process design. To effectively capture the high-dimensional features of the laser welding melt pool, this study introduces three branched prediction heads designed to perform distinct image augmentation tasks. The first task involves predicting the rotation angle, where the input image is randomly rotated by a predefined angle 𝛼𝛼. The second task consists of applying random Gaussian blur. The third task involves random cropping, wherein the input image is cropped to a resolution of 224 × 244 pixels. The prediction head responsible for estimating the rotation angle is denoted as 𝐴𝐴1 (∙), the prediction head for Gaussian blur parameters is denoted as 𝐴𝐴2 (∙), and the prediction head for random cropping is denoted as 𝐴𝐴3 (∙). Consequently, the aforementioned image augmentation tasks are achieved by optimizing the following loss function.

1 1 𝐿𝐿𝑎𝑎𝑎𝑎𝑎𝑎 (𝑥𝑥1 ) + 𝐿𝐿𝑎𝑎𝑎𝑎𝑎𝑎 (𝑥𝑥2 ) 2 2 1 = �𝛼𝛼1 𝐿𝐿𝑅𝑅𝑅𝑅𝑅𝑅 (𝑥𝑥1 , 𝜃𝜃) + 𝛽𝛽1 𝐿𝐿𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏 �𝑥𝑥1 , 𝜎𝜎𝑐𝑐 , 𝜎𝜎𝑑𝑑 , 𝜎𝜎𝑥𝑥 , 𝜎𝜎𝑦𝑦 � + 𝛾𝛾1 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 �𝑥𝑥1 , 𝐶𝐶𝑥𝑥 , 𝐶𝐶𝑦𝑦 �� 2 1 (6) + �𝛼𝛼2 𝐿𝐿𝑅𝑅𝑅𝑅𝑅𝑅 (𝑥𝑥2 , 𝜃𝜃) + 𝛽𝛽2 𝐿𝐿𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏 �𝑥𝑥2 , 𝜎𝜎𝑐𝑐 , 𝜎𝜎𝑑𝑑 , 𝜎𝜎𝑥𝑥 , 𝜎𝜎𝑦𝑦 � + 𝛾𝛾2 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 �𝑥𝑥2 , 𝐶𝐶𝑥𝑥 , 𝐶𝐶𝑦𝑦 �� 2 (7) 𝛼𝛼1 + 𝛽𝛽1 + 𝛾𝛾1 = 1, 𝛼𝛼2 + 𝛽𝛽2 + 𝛾𝛾2 = 1 Here, 𝐿𝐿𝑅𝑅𝑅𝑅𝑅𝑅 , 𝐿𝐿𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏 , and 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 denote the loss functions corresponding to the angle prediction 𝐿𝐿𝑎𝑎𝑎𝑎𝑎𝑎 =

task, Gaussian blur prediction task, and random cropping task, respectively. The coefficients 𝛼𝛼, 𝛽𝛽, and 𝛾𝛾 are used to balance the contribution of these three loss terms. 𝜃𝜃 represents the random rotation angle applied to the image, while σx and σy indicate the coordinates of the point with the

highest Gaussian blur intensity in the image, 𝜎𝜎𝑐𝑐 denotes the magnitude of the blur at that point, and 𝜎𝜎𝑑𝑑 is the decay rate of the Gaussian blur intensity. 𝐶𝐶𝑥𝑥 and 𝐶𝐶𝑦𝑦 are the coordinates of the center of

the cropped image within the original image, with the cropped region having a fixed size of 224 × 224 pixels. Furthermore, the root mean square error (RMSE) function is employed as the loss function for the aforementioned processes. Therefore, the loss function for the angle prediction task can be written as follows: 𝑛𝑛

1 𝐿𝐿𝑅𝑅𝑅𝑅𝑅𝑅 (𝑥𝑥1 , 𝜃𝜃) = � �‖𝜃𝜃𝑖𝑖 − 𝜃𝜃�𝑖𝑖 ‖2 � 𝑛𝑛

1 2

(8)

𝑖𝑖=1

Here, 𝜃𝜃𝑖𝑖 represents the actual rotation angle, and 𝜃𝜃�𝑖𝑖 denotes the predicted value of 𝐴𝐴1 (∙).

Similarly, the loss functions for the Gaussian blur prediction task and the random cropping task can be expressed as follows: 1 2

𝑛𝑛

𝑛𝑛

1 1 𝐿𝐿𝑏𝑏𝑏𝑏𝑏𝑏𝑏𝑏 �𝑥𝑥1 , 𝜎𝜎𝑐𝑐 , 𝜎𝜎𝑑𝑑 , 𝜎𝜎𝑥𝑥 , 𝜎𝜎𝑦𝑦 � = � �‖𝜎𝜎𝑐𝑐 − 𝜎𝜎�𝑐𝑐 ‖2 � + � �‖𝜎𝜎𝑑𝑑 − 𝜎𝜎�𝑑𝑑 ‖2 � 𝑛𝑛 𝑛𝑛 𝑖𝑖=1

𝑖𝑖=1

1 2

𝑛𝑛

𝑛𝑛

1 1 + � �‖𝜎𝜎𝑥𝑥 − 𝜎𝜎�𝑥𝑥 ‖2 � + � �‖𝜎𝜎𝑦𝑦 − 𝜎𝜎�𝑦𝑦 ‖2 � 𝑛𝑛 𝑛𝑛 𝑛𝑛

𝑖𝑖=1

1 2

𝑛𝑛

𝑖𝑖=1

1 1 2 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 �𝑥𝑥2 , 𝐶𝐶𝑥𝑥 , 𝐶𝐶𝑦𝑦 � = � �‖𝐶𝐶𝑥𝑥 − 𝐶𝐶̂𝑥𝑥 ‖2 � + � ��𝐶𝐶𝑦𝑦 − 𝐶𝐶̂𝑦𝑦 � � 𝑛𝑛 𝑛𝑛 𝑖𝑖=1

𝑖𝑖=1

1 2

Both take forms similar to Equation (8) and will not be elaborated further here.

1 2

1 2

(9) (10)

In the task of classifying the top surface molten pool in laser welding, the morphological characteristics of the molten pool are the primary focus of the network. Additionally, we recognize that this aspect contributes to image enhancement. Based on these two premises, we have formulated and designed three tasks. Firstly, the image rotation operation compels the model to learn orientation-invariant global features rather than relying on position-fixed local features. This suggests that the network must analyze the global structures within the image to determine the applied rotational transformation.

Given the context of laser welding penetration classification, where image content is relatively simple and the primary targets are the molten pool and small pores, the network can focus on the morphology of these elements to predict the rotation angle. Secondly, Gaussian blurring reduces image details such as textures and edges, thereby forcing the model to learn more generalized global features and avoid over-fitting to specific sharp features present in the training set. Moreover, during actual image acquisition, the prevalence of blurring is a direct consequence of real-world industrial laser welding conditions. Specifically, blurs originate from physical phenomena such as metal vapor and fume interference, plasma formation, rapid molten pool dynamics that can induce motion blur, and minor focus drift or mechanical vibrations from the robotic arm during prolonged operation. These factors introduce optical obstructions or motion artifacts. By simulating such conditions through data augmentation, the model's robustness and adaptability to real production environments are significantly improved. Thirdly, by randomly cropping local regions of the image, the model is compelled to extract discriminative features from limited pixel areas, thereby increasing its sensitivity to fine-grained details. During molten pool classification, this encourages the model to focus on key indicators such as the presence or absence of keyholes, the morphology of molten pool edges, and subtle features such as spatter, enabling it to distinguish minor differences between similar regions. In summary, the design of these three aforementioned tasks ensures that the feature vectors ultimately captured by the model contain both global and detailed features. Moreover, this approach can be regarded as a form of image enhancement, contributing to improved model generalization and robustness. 3.1.2 Physics informed neural network The problem that PINN aim to address can be formulated as: 𝑁𝑁[𝑢𝑢(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧)|𝜆𝜆] = 0, 𝑥𝑥 ∈ ℝ𝐷𝐷 , 𝑡𝑡 ∈ ℝ

(11)

where, 𝑁𝑁[𝑢𝑢(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧)|𝜆𝜆] denotes the differential operator used to describe the real physical system, and 𝑢𝑢(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧) represents the solution of the system. The primary function of PINN is to estimate the system's solution 𝑢𝑢(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧) using a deep neural network 𝐷𝐷(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧|𝜃𝜃). It is noteworthy that this network provides only an approximate solution to 𝑢𝑢(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧). Let 𝜲𝜲 = (𝑥𝑥, 𝑦𝑦, 𝑧𝑧) represent the spatial coordinates. During training, the governing equation (11) acts as a physical constraint. The optimization objective is to minimize a loss function composed of the data misfit and the PDE 𝑁𝑁

𝑑𝑑 be the set of training data points on the initial or boundary conditions, and residual. Let {𝑡𝑡𝑖𝑖 , 𝜲𝜲𝑖𝑖 }i=1

𝑁𝑁

𝑝𝑝 be the set of collocation points sampled within the physical domain. The composite loss {𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 }𝑗𝑗=1

function can be expressed as:

𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 = �

𝑁𝑁𝑑𝑑

1 2

𝑁𝑁𝑝𝑝

1 2

1 1 2 �‖𝑢𝑢(𝑡𝑡𝑖𝑖 , 𝜲𝜲𝑖𝑖 ) − 𝐷𝐷(𝑡𝑡𝑖𝑖 , 𝜲𝜲𝑖𝑖 )‖2 � + � ��𝑁𝑁�𝐷𝐷�𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 �|𝜆𝜆�� � 𝑁𝑁𝑑𝑑 𝑁𝑁𝑝𝑝 𝑖𝑖=1

𝑗𝑗=1

(12)

Here, the first term enforces the model to match the known data, while the second term enforces it to satisfy the governing physical law. It is evident that Equation (12) can also be regarded as a variant of the RMSE By explicitly embedding the physical governing equations into the construction of loss function, the PINN enables a progressive approximation of the deep learning model toward the true physical behavior. Compared to conventional numerical computation methods such as the finite element method (FEM) and the finite difference method (FDM), PINN does not rely on extensive and meticulously designed meshes. Instead, it employs automatic differentiation to update the model parameters, allowing the derivatives to be computed automatically during the training process. Essentially, this approach constructs a differentiable framework within the function space, leveraging the approximation capabilities of deep neural networks to transform the analytical solutions of physical equations into parameter optimization problems. In SimPhysNet, our application of PINN leverages a specific network architecture to guide image feature learning by enforcing physical consistency. The process begins with the highdimensional features extracted from molten pool images by the backbone encoder. These visual features are then fed into dedicated physical parameter prediction modules. These modules are designed to infer crucial physical parameters (such as the spatial boundaries of the welding zone and heat source characteristics) directly from the image features. Subsequently, a MLP approximates the temperature field 𝑢𝑢(𝑡𝑡, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧) based on these image-inferred physical parameters. The PINN loss function then evaluates how well this feature-derived physical representation (the approximated temperature field) adheres to fundamental physical laws, specifically the heat conduction PDE and boundary conditions. By backpropagating this physical consistency loss through the network, the backbone encoder is implicitly forced to learn features that correspond to physically plausible molten pool states. This mechanism serves as a powerful regularization, guiding the model to extract features that are not merely visually correlated but are inherently consistent with the underlying physics, thereby enhancing generalization and robustness. In this study, we utilize the partial differential equation (PDE) of heat conduction in a semiinfinite medium to develop a PINN capable of predicting the state of any sampling point at any given time during the welding process: 𝜕𝜕𝑇𝑇 − 𝛼𝛼∇2 𝑇𝑇 = 0 𝜕𝜕𝑡𝑡

(13)

where, 𝑇𝑇 is temperature and 𝑡𝑡 is time. 𝛼𝛼 denotes the thermal diffusivity of the material, which is

assumed to be a constant with a value of 1.17 × 10−5 𝑚𝑚2 ∙ 𝑠𝑠 −1 in our experiments. Therefore, the loss function for the thermal conduction term should be designed as follows: 𝑛𝑛

2

1 2

1 𝜕𝜕𝐷𝐷 �𝑡𝑡 , 𝜲𝜲 � 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 = � � � 𝑃𝑃𝑃𝑃𝑃𝑃 𝑗𝑗 𝑗𝑗 − 𝛼𝛼∇2 𝐷𝐷𝑃𝑃𝑃𝑃𝑃𝑃 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 �� � 𝑁𝑁𝑝𝑝 𝜕𝜕𝑡𝑡

(14)

𝑇𝑇(𝑡𝑡 = 0, 𝑥𝑥, 𝑦𝑦, 𝑧𝑧) = 26℃

(15)

𝑗𝑗=1

The following equations are satisfied at the boundaries:

The loss function for the boundary condition term should be formulated as follows: 𝑛𝑛

1 2

1 2 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼 = � ��𝑇𝑇�𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 � − 𝐷𝐷𝐼𝐼𝐼𝐼𝐼𝐼 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 �� � 𝑁𝑁𝑝𝑝

(16)

𝑗𝑗=1

In the context of welding, the heat conduction equation for the transient temperature field, known as the Fourier-Kirchhoff partial differential equation[15], is mathematically described as follows: 𝜌𝜌𝑐𝑐𝑝𝑝 ∙

𝜕𝜕𝑇𝑇 = ∇ ∙ (𝜆𝜆∇𝑇𝑇) + 𝑞𝑞𝑣𝑣 𝜕𝜕𝑡𝑡 𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕 𝜕𝜕𝜕𝜕 = � �𝜆𝜆 � + �𝜆𝜆 � + �𝜆𝜆 �� + 𝑞𝑞𝑣𝑣 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕 𝜕𝜕𝜕𝜕

(17)

where, 𝜌𝜌 denotes density, 𝑐𝑐𝑝𝑝 signifies specific heat, and 𝜆𝜆 indicates thermal conductivity. The

volumetric heat source density 𝑞𝑞𝑣𝑣 quantifies the amount of heat generated per unit time per unit volume of material, and it is utilized for modeling the heat source in the welding process.

In this study, the conical volumetric heat source model [15] with a Gaussian distribution is applied to simulate the heat source distribution during the laser welding process, as shown in Fig. 7(a). The heat source energy density, 𝑞𝑞𝑣𝑣 |𝐶𝐶𝐶𝐶𝐶𝐶 [𝑊𝑊 ∙ 𝑚𝑚−3 ] , can be mathematically represented as follows: 𝑞𝑞𝑣𝑣 (𝑥𝑥, 𝑦𝑦, 𝑧𝑧, 𝑡𝑡)|𝐶𝐶𝐶𝐶𝐶𝐶 =

1 9𝜂𝜂𝜂𝜂𝑒𝑒 3 3[(𝑥𝑥 − 𝑉𝑉𝑤𝑤 𝑡𝑡)2 + 𝑦𝑦 2 ] 𝑒𝑒𝑒𝑒𝑒𝑒 �− � 𝜋𝜋(𝑒𝑒 3 − 1) (𝑧𝑧𝑒𝑒 − 𝑧𝑧𝑖𝑖 )(𝑟𝑟𝑒𝑒2 + 𝑟𝑟𝑒𝑒 𝑟𝑟𝑖𝑖 + 𝑟𝑟𝑖𝑖2 ) 𝑟𝑟02 (𝑧𝑧) 𝑟𝑟𝑖𝑖 − 𝑟𝑟𝑒𝑒 (𝑧𝑧 − 𝑧𝑧𝑒𝑒 ) 𝑟𝑟0 (𝑧𝑧) = 𝑟𝑟𝑒𝑒 + 𝑧𝑧𝑖𝑖 − 𝑧𝑧𝑒𝑒

(18)

(19)

where, 𝑃𝑃[𝑊𝑊] is the laser power, 𝜂𝜂[−] is laser absorption rate and 𝑉𝑉𝑤𝑤 [𝑚𝑚 ∙ 𝑠𝑠 −1 ] the welding speed. 𝑟𝑟𝑒𝑒 [𝑚𝑚] and 𝑟𝑟𝑖𝑖 [𝑚𝑚] represent the radii of the upper surface 𝑧𝑧 = 𝑧𝑧𝑒𝑒 and lower surface 𝑧𝑧 = 𝑧𝑧𝑖𝑖 of the heat source, respectively. 𝑟𝑟0 is the heat source radius at the plane 𝑧𝑧 , functioning as a linearly decreasing distribution from the upper surface to the lower surface. Based on the aforementioned analysis, the loss term associated with energy conservation of the conical heat source should be formulated as follows:

𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶 = �

𝑛𝑛

2

1 𝜕𝜕𝐷𝐷 �𝑡𝑡 , 𝜲𝜲 � � �𝜌𝜌𝑐𝑐𝑝𝑝 𝐶𝐶𝐶𝐶𝐶𝐶 𝑗𝑗 𝑗𝑗 − �∇ ∙ �𝜆𝜆∇𝐷𝐷𝐶𝐶𝐶𝐶𝐶𝐶 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 �� + 𝑞𝑞𝑣𝑣 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 ��� � 𝑁𝑁𝑝𝑝 𝜕𝜕𝑡𝑡 𝑗𝑗=1

1 2

(20)

Furthermore, considering that a single heat source model may not fully characterize the thermal distribution in laser welding[16], this study employs a hybrid approach based on the Goldak’s double ellipsoidal heat source model[17], as illustrated in Fig. 7(b). The corresponding mathematical formulation is presented as follows:

𝑞𝑞𝑣𝑣1 (𝑥𝑥, 𝑦𝑦, 𝑧𝑧, 𝑡𝑡)|𝐺𝐺𝐺𝐺𝐺𝐺 =

6√3𝑓𝑓1 𝛷𝛷 −3(𝑥𝑥 − 𝑉𝑉𝑤𝑤 𝑡𝑡)2 3𝑦𝑦 2 3𝑧𝑧 2 𝑒𝑒𝑒𝑒𝑒𝑒 � � 𝑒𝑒𝑒𝑒𝑒𝑒 �− � 𝑒𝑒𝑒𝑒𝑒𝑒 �− � 2 2 𝑎𝑎 𝑏𝑏 𝑎𝑎𝑎𝑎𝑐𝑐1 𝜋𝜋 3/2 𝑐𝑐12

6√3𝑓𝑓1 𝛷𝛷 −3(𝑥𝑥 − 𝑉𝑉𝑤𝑤 𝑡𝑡)2 3𝑦𝑦 2 3𝑧𝑧 2 𝑒𝑒𝑒𝑒𝑒𝑒 � � 𝑒𝑒𝑒𝑒𝑒𝑒 �− � 𝑒𝑒𝑒𝑒𝑒𝑒 �− � 𝑎𝑎2 𝑏𝑏 2 𝑎𝑎𝑎𝑎𝑐𝑐2 𝜋𝜋 3/2 𝑐𝑐22 2𝑐𝑐1 2𝑐𝑐2 𝑓𝑓1 = , 𝑓𝑓2 = , 𝑓𝑓 + 𝑓𝑓2 = 2, 𝑞𝑞𝑣𝑣 (𝑥𝑥, 𝑦𝑦, 𝑧𝑧, 𝑡𝑡)|𝐺𝐺𝐺𝐺𝐺𝐺 = 𝑞𝑞𝑣𝑣1 + 𝑞𝑞𝑣𝑣2 𝑐𝑐1 + 𝑐𝑐2 𝑐𝑐1 + 𝑐𝑐2 1

𝑞𝑞𝑣𝑣2 (𝑥𝑥, 𝑦𝑦, 𝑧𝑧, 𝑡𝑡)|𝐺𝐺𝐺𝐺𝐺𝐺 =

(21)

(22)

(23)

where 𝛷𝛷 is the effective laser power. 𝑓𝑓1 and 𝑓𝑓2 specify the proportion of energy distributed to the front and rear sections of the double ellipsoidal heat source. The semi-axis lengths 𝑎𝑎 and 𝑏𝑏 represent the width and depth of the heat source, respectively, while the semi-axis lengths 𝑐𝑐1 and 𝑐𝑐2 denote the lengths of the front and rear sections of the heat source, respectively. The loss term for Goldak heat source energy conservation model is consistent with Equation (20).

Fig. 7 The forms of two types of welding heat sources: (a) Conical heat source model[15], (b)Goldak heat source model[17]. To prevent each term from being driven toward zero during the optimization process of the PINN, we have extended the RMSE algorithm by incorporating a penalty term, as exemplified in Equation (14): 𝑛𝑛

2

1 2

1 𝜕𝜕𝐷𝐷 �𝑡𝑡 , 𝜲𝜲 � 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 = � � � 𝑃𝑃𝑃𝑃𝑃𝑃 𝑗𝑗 𝑗𝑗 − 𝛼𝛼∇2 𝐷𝐷𝑃𝑃𝑃𝑃𝑃𝑃 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 �� � 𝑛𝑛 𝜕𝜕𝑡𝑡 𝑗𝑗=1

−1

𝜕𝜕𝐷𝐷𝑃𝑃𝑃𝑃𝑃𝑃 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 � +𝜔𝜔 ∙ ��� � + 𝜀𝜀� 𝜕𝜕𝑡𝑡

−1

+ ��𝛼𝛼∇2 𝐷𝐷𝑃𝑃𝑃𝑃𝑃𝑃 �𝑡𝑡𝑗𝑗 , 𝜲𝜲𝑗𝑗 �� + 𝜀𝜀� �

(24)

In this manner, as each component approaches zero, the penalty term increases significantly, thereby steering the optimization direction away from zero. Here, 𝜔𝜔 denotes the penalty coefficient, which is consistently set to 0.05 in our experimental configuration. Additionally, 𝜀𝜀 = 1𝑒𝑒⁻⁶ is introduced as a regularization term to avoid division by zero. A critical challenge in training PINNs is balancing the contributions of disparate loss terms (𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 , 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼 , and 𝜆𝜆𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒 ). Static weights require extensive manual tuning and may fail to adapt to

the changing dynamics of the training process. To overcome this limitation, a hybrid adaptive weighting methodology is implemented. This is a sophisticated, multi-stage approach designed to autonomously regulate the influence of each loss component through three key stages: loss signal

stabilization, normalization, and uncertainty-based dynamic weighting. First, to mitigate the effects of high-frequency noise inherent in the stochastic optimization process, each loss signal is stabilized using an exponential moving average (EMA). The smoothed , for each component 𝑘𝑘 ∈ {𝑃𝑃𝑃𝑃𝑃𝑃, 𝐼𝐼𝐼𝐼𝐼𝐼, 𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒} at training step 𝑡𝑡 is calculated as: loss, 𝐿𝐿𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠ℎ 𝑘𝑘 = 𝛼𝛼 ∙ 𝐿𝐿𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠ℎ(𝑡𝑡−1) + (1 − 𝛼𝛼) ∙ 𝐿𝐿𝑡𝑡𝑘𝑘 𝐿𝐿𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠ℎ(𝑡𝑡) 𝑘𝑘 𝑘𝑘

where 𝐿𝐿𝑡𝑡𝑘𝑘 is the raw loss at the current step and 𝛼𝛼 is a smoothing factor.

(25)

Second, to ensure all loss terms operate on a comparable numerical scale, the smoothed losses are normalized. Each smoothed loss is divided by its value at the beginning of training, 𝐿𝐿𝑘𝑘,𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖 , yielding a normalized loss 𝐿𝐿�𝑘𝑘 : 𝐿𝐿�𝑘𝑘 =

+ 𝜀𝜀 𝐿𝐿𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠ℎ(𝑡𝑡) 𝑘𝑘 𝐿𝐿𝑘𝑘,𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖 + 𝜀𝜀

(26)

This normalization prevents any single term from dominating the total loss merely due to its magnitude.

Finally, the core of this method is the dynamic weight calculation, which combines a predefined initial weight, 𝑤𝑤𝑘𝑘,𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖 , with a learnable, uncertainty-based parameter, 𝑠𝑠𝑘𝑘 . This approach allows for the incorporation of prior knowledge while enabling fine-grained, dynamic adjustments. The final weight 𝑤𝑤𝑘𝑘 for each loss is: 𝑤𝑤𝑘𝑘 = 𝑤𝑤𝑘𝑘,𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖𝑖 ∙ 𝑒𝑒𝑒𝑒𝑒𝑒(−𝑠𝑠𝑘𝑘 )

(27)

The total loss function for the PINN is then formulated as the sum of the weighted, normalized losses and a regularization term on the learnable parameters. A ReLU function is applied to ensure non-negative loss contributions. 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 =

𝑘𝑘∈{𝑃𝑃𝑃𝑃𝑃𝑃,𝐼𝐼𝐼𝐼𝐼𝐼,𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒}

1 �𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅 � 𝑤𝑤𝑘𝑘 ∙ 𝐿𝐿�𝑘𝑘 � + |𝑠𝑠𝑘𝑘 |� 2

(28)

In this formulation, the model learns to down-weigh tasks with higher uncertainty (larger 𝑠𝑠𝑘𝑘 ), while the |𝑠𝑠𝑘𝑘 | term regularizes the learning process to prevent any weight from collapsing to zero. This adaptive mechanism provides a robust and autonomous framework for balancing the multifaceted objectives in PINN training. The architecture of the PINN is illustrated in Fig. 8. A multilayer perceptron (MLP)[18] is employed to approximate the solution 𝑢𝑢(𝑡𝑡, 𝜲𝜲) . This approximated solution is then utilized to construct the heat conduction loss 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 , boundary condition loss 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼 , and heat source loss 𝐿𝐿𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒 . The parameters of the MLP are trained using a gradient descent optimization method based on the backpropagation of the loss function.

FC 7

FC

𝒙 𝒚 𝒛 𝑻

Heat source param

FC 32

𝒕𝒆𝒎𝒑𝒆𝒓𝒂𝒕𝒖𝒓𝒆_𝒑𝒓𝒆𝒅

SoftPlus

FC 64

FC

5

𝑪𝒐𝒏𝒊𝒄𝒂𝒍_𝒉𝒆𝒂𝒕_𝒔𝒐𝒖𝒓𝒄𝒆_𝒑𝒓𝒆𝒅

FC

32

SoftPlus

FC

64

FC

𝑮𝒐𝒍𝒅𝒂𝒌𝒌_𝒉𝒆𝒂𝒕_𝒔𝒐𝒖𝒓𝒄𝒆_𝒑𝒓𝒆𝒅

FC

ReLu

SE Block

FC

LHS module

6

[𝒙𝒎𝒂𝒙 , 𝒙𝒎𝒊𝒏 ] [𝒚𝒎𝒂𝒙 , 𝒚𝒎𝒊𝒏 ] [𝒛𝒎𝒂𝒙 , 𝒛𝒎𝒊𝒏 ]

SoftPlus

FC

32

SoftPlus

FC

64 64

256

FC

FC

𝒄𝒐𝒐𝒓𝒅_𝒑𝒓𝒆𝒅

Fig. 8 Overview of the PINN. A multilayer perceptron (MLP) is employed to approximate the solution 𝑢𝑢(𝑡𝑡, 𝜲𝜲), which is subsequently utilized to construct the heat conduction loss 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 boundary condition loss 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼 , and heat source loss 𝐿𝐿𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒 . The derivatives of 𝑢𝑢 are computed using automatic differentiation in PyTorch. The parameters of the MLP are trained using a gradient descent optimization method based on the backpropagation of the loss function. Latin hypercube sampling (LHS) is a highly efficient data sampling method[19]. Compared to conventional Monte Carlo sampling methods, LHS demonstrates superior performance in terms of sampling efficiency, reduced sampling variance, and lower correlation among multidimensional variables. By partitioning the parameter space into equally probable sub-intervals and selecting one sample point independently from each sub-interval, LHS effectively mitigates the clustering effect commonly observed in conventional random sampling approaches[20]. Consider a 𝑑𝑑-dimensional parameter space, where each dimension is divided into 𝑁𝑁 equiprobable sub-intervals. Let 𝑖𝑖 = 1, 2, … , 𝑁𝑁 denote the sample index and 𝑗𝑗 = 1, 2, … , 𝑑𝑑 represent the dimension index. For each dimension 𝑗𝑗 , a random permutation 𝜋𝜋𝑗𝑗 is generated, such that 𝜋𝜋𝑗𝑗 (𝑖𝑖) specifies the sub-interval

containing the 𝑖𝑖-th sample in the 𝑗𝑗-th dimension. The actual sample value in the 𝑗𝑗-th dimension is then determined according to the following equation: 𝑥𝑥𝑖𝑖𝑖𝑖 =

𝜋𝜋𝑗𝑗 (𝑖𝑖) − 𝑢𝑢𝑖𝑖𝑖𝑖 𝑁𝑁

(29)

In this context, 𝑢𝑢𝑖𝑖𝑖𝑖 represents a random number drawn from the uniform distribution within

dynamically determined spatial boundaries. This approach ensures that each sub-interval contains at least one sample while maintaining randomness within individual intervals. The spatial parameter predictor processes high-dimensional image features to infer dynamic spatial boundaries for the relevant physical domains along the 𝑥𝑥 , 𝑦𝑦 , and 𝑧𝑧 axes, represented as [𝑥𝑥𝑚𝑚𝑚𝑚𝑚𝑚 , 𝑥𝑥𝑚𝑚𝑚𝑚𝑚𝑚 ], [𝑦𝑦𝑚𝑚𝑚𝑚𝑚𝑚 , 𝑦𝑦𝑚𝑚𝑚𝑚𝑚𝑚 ], [𝑧𝑧𝑚𝑚𝑚𝑚𝑚𝑚 , 𝑧𝑧𝑚𝑚𝑚𝑚𝑚𝑚 ] . Subsequently, the LHS module utilizes these inferred threedimensional spatial boundaries to generate a set of 𝑁𝑁𝑆𝑆 pseudo-random physical coordinates (𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) . Compared with simple random sampling or uniform grid-based approaches, LHS achieves more uniform and comprehensive coverage of the three-dimensional physical space, thereby enhancing the robustness and accuracy of PINN physical loss calculations. Additionally, heat source parameters are predicted by dedicated heat source predictors (Goldak_heat_source_pred and Conical_heat_source_pred). The PINN then performs computations at each generated sampling point (𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) , effectively integrating image-derived features into the physical consistency

constraints of the model. The pre-training stage of the PINN is illustrated in Alg. 1. A three-layer MLP is utilized to estimate the heat source parameters required for the loss function computation. To ensure a continuously differentiable neural representation, the Tanh activation function is adopted across intermediate layers, while the SoftPlus activation function is used in the output layer to enforce positivity of the predicted values. Alg. 1 Algorithm of pre-training stage. Input: 𝐷𝐷𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 , 𝐷𝐷𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 , 𝐷𝐷𝑚𝑚𝑚𝑚𝑚𝑚 ∈ 𝐷𝐷𝑈𝑈 , set of random transformations 𝐹𝐹, rotation degrees 𝑅𝑅, gaussian blurs 𝛴𝛴, crop images 𝐶𝐶, contrastive network 𝐸𝐸𝜃𝜃 , 𝑀𝑀𝜃𝜃 , PINN network 𝑃𝑃𝜃𝜃 , learning rate 𝛼𝛼 01: randomly initialize 𝜃𝜃 02: for 𝑖𝑖 in epochs 03: 𝑏𝑏 ← 𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅(𝐷𝐷𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 ) //Select mini-batch samples 04: for (𝑥𝑥, 𝑡𝑡) ← 𝑏𝑏 do 05: draw random rotation degrees 𝑟𝑟1 , 𝑟𝑟2 ∼ 𝑅𝑅 06: 𝑥𝑥1 ← 𝑟𝑟1 (𝑥𝑥), 𝑥𝑥2 ← 𝑟𝑟2 (𝑥𝑥) 07: 08: 09: 10: 11: 12: 13: 14: 15: 16: 17: 18: 19: 20: 21: 22: 23: 24: 25: 26: 27: 28: 29: 30: 31:

draw random gaussian blurs 𝜎𝜎1 , 𝜎𝜎2 ∼ 𝛴𝛴 𝑥𝑥1 ← 𝜎𝜎1 (𝑥𝑥1 ), 𝑥𝑥2 ← 𝜎𝜎2 (𝑥𝑥2 ) draw random crop images 𝑐𝑐1 , 𝑐𝑐2 ∼ 𝐶𝐶 𝑥𝑥1 ← 𝑐𝑐1 (𝑥𝑥1 ), 𝑥𝑥2 ← 𝑐𝑐2 (𝑥𝑥2 )

draw random transformations 𝑓𝑓1 , 𝑓𝑓2 ∼ 𝐹𝐹 𝑥𝑥1 ← 𝑓𝑓1 (𝑥𝑥1 ), 𝑥𝑥2 ← 𝑓𝑓2 (𝑥𝑥1 ) 𝑘𝑘1 ← 𝐸𝐸𝜃𝜃 (𝑥𝑥1 ), 𝑘𝑘2 ← 𝐸𝐸𝜃𝜃 (𝑥𝑥2 ), 𝑞𝑞1 ← 𝑀𝑀𝜃𝜃 (𝑘𝑘1 ), 𝑞𝑞2 ← 𝑀𝑀𝜃𝜃 (𝑘𝑘2 ) 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐 ← 1/2 ‖𝑘𝑘1 , 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠(𝑞𝑞2 )‖2 + 1/2 ‖𝑘𝑘2 , 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠(𝑞𝑞1 )‖2 (𝐿𝐿𝑅𝑅𝑅𝑅𝑅𝑅 , 𝐿𝐿𝐵𝐵𝐵𝐵𝐵𝐵𝐵𝐵 , 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 ) ← 1/2 ‖𝑥𝑥1 , 𝑟𝑟1 ‖2 + 1/2 ‖𝑥𝑥2 , 𝑟𝑟2 ‖2

𝑝𝑝1 ← 𝑃𝑃(𝑟𝑟1 ), 𝑝𝑝2 ← 𝑃𝑃(𝑟𝑟2 ) 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 ← 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝐸𝐸𝜃𝜃 (𝑥𝑥), 𝑡𝑡) 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 ← 𝐿𝐿𝐿𝐿𝐿𝐿_𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚(𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝) // Generate 𝑁𝑁𝑆𝑆 sampling points in the physical space 𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 ← 𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺_ℎ𝑒𝑒𝑒𝑒𝑒𝑒_𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝐸𝐸𝜃𝜃 (𝑥𝑥), 𝑡𝑡) 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 ← 𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶_ℎ𝑒𝑒𝑒𝑒𝑒𝑒_𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝐸𝐸𝜃𝜃 (𝑥𝑥), 𝑡𝑡) 𝑇𝑇 ← 𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑡𝑡) 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 ← 0 for each sampled point (𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) in 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 do: �(𝑡𝑡, 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) ← 𝑃𝑃𝜃𝜃 (𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑇𝑇𝑗𝑗 , 𝑡𝑡, 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝑈𝑈 �, 𝑡𝑡, 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃_𝑗𝑗 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 (𝑈𝑈 �, 𝑡𝑡, 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼_𝑗𝑗 ← 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼 (𝑈𝑈 �, 𝑡𝑡, 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶_𝑗𝑗 ← 𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶 (𝑈𝑈 �, 𝑡𝑡, 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝐺𝐺𝐺𝐺𝐺𝐺_𝑗𝑗 ← 𝐿𝐿𝐺𝐺𝐺𝐺𝐺𝐺 (𝑈𝑈 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 + 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃_𝑗𝑗 + 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼_𝑗𝑗 + 𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶_𝑗𝑗 + 𝐿𝐿𝐺𝐺𝐺𝐺𝐺𝐺_𝑗𝑗

end for 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 /𝑁𝑁𝑆𝑆

32: 33: 34: 35: 36: 37:

𝐿𝐿 ← 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐 + 𝜆𝜆1 ∙ (𝐿𝐿𝑅𝑅𝑅𝑅𝑅𝑅 + 𝐿𝐿𝐵𝐵𝐵𝐵𝐵𝐵𝐵𝐵 + 𝐿𝐿𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 ) + 𝜆𝜆2 ∙ 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃

end for 𝜃𝜃 ← 𝜃𝜃 − 𝛼𝛼∇𝜃𝜃 𝐿𝐿 𝐴𝐴𝐴𝐴𝐴𝐴 ← 𝐾𝐾𝐾𝐾𝐾𝐾(𝐸𝐸𝜃𝜃 , 𝑀𝑀𝜃𝜃 , 𝑃𝑃𝜃𝜃 , 𝐷𝐷𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡 , 𝐷𝐷𝑚𝑚𝑚𝑚𝑚𝑚 ) end for return 𝐸𝐸𝜃𝜃

3.2 Fine-tuning stage

By leveraging the deep neural network architecture presented in Section 3.1, the proposed algorithm effectively extracts high-dimensional semantic features from laser molten pool images. Conventional approaches typically adopt parameter transfer learning strategies, fine-tuning pretrained models on specific welding process datasets for weld penetration state classification. However, under limited training sample conditions, traditional fine-tuning methods are prone to overfitting, resulting in feature space distribution shifts and significant degradation in model generalization performance. To mitigate this challenge, this study introduces a novel algorithm based on the prototype network framework. The approach preserves the feature extraction capabilities of the pre-trained network while constructing class prototype vectors and establishing a metric space mapping mechanism, thereby enabling compact feature representation and maintaining generalization of classification boundaries in few-shot training scenarios. Specifically, prototype networks learn a metric space 𝛺𝛺, in which each class occupies a distinct region, with class centers maximally separated. For a new input sample 𝑥𝑥, the Euclidean distance between 𝑥𝑥 and each class center are calculated, and the class corresponding to the minimum distance is assigned as the predicted outcome. This underlying principle is illustrated in Fig. 9.

𝑪𝟐

𝑪𝟏

𝒙

Fig. 9 Prototypical network in a few-shot learning. After completing the aforementioned two training stages, the proposed model demonstrates the

ability to adapt to few-shot learning tasks. The training process during the fine-tuning stage is outlined in Alg. 2. As illustrated in Alg. 2, 𝑁𝑁𝑠𝑠 samples are first randomly selected from 𝐷𝐷𝐿𝐿 to form the support set 𝑆𝑆𝑘𝑘 and 𝑁𝑁𝑄𝑄 samples are selected as the query set 𝑄𝑄𝑘𝑘 (line 6).

For samples belonging to the same categories, it is assumed that the high-dimensional feature vectors extracted by the same neural network exhibit a certain level of similarity, whereas feature vectors from different categories may differ significantly. Therefore, the mean of the feature vectors corresponding to each category's samples is calculated and used as the support vector 𝑐𝑐𝑘𝑘 , which can be mathematically expressed as: 𝑐𝑐𝑘𝑘 = (|𝑆𝑆𝑘𝑘 |)−1 �

(𝑥𝑥𝑖𝑖 ,𝑦𝑦𝑖𝑖 ,𝑡𝑡𝑖𝑖 )∈𝑆𝑆𝑘𝑘

𝐸𝐸𝜃𝜃 (𝑥𝑥𝑖𝑖 )

(30)

Consequently, to classify a molten pool image, its feature vector must first be calculated and compared to 𝑐𝑐𝑘𝑘 using the Euclidean distance (line 11). A smaller distance indicates a higher likelihood that the molten pool image belongs to that class. Therefore, the goal of few-shot learning is to ensure that support vectors effectively represent their respective classes by minimizing the distance between feature vectors sharing the same class label. The cross-entropy loss function is employed to update the network parameters (line 12). In the fine-tuning stage, PINN also functions as an additional supervisory signal. In classification tasks, the model must ensure that its physical interpretation of the images is precise and consistent with physical laws. This approach enhances the model's generalization capability by enforcing that the learned classification features maintain physical consistency, thereby making the model's decisions more reliable and aligned with physical intuition when encountering previously unseen samples. The configuration of PINN is analogous to that outlined in Alg. 1 and is not further elaborated here (line 20-29). The coefficient 𝛽𝛽 for the PINN network loss is set to 0.1 in our experiments. Alg. 2 Algorithm of fine-tuning stage. 𝑁𝑁𝐿𝐿 Input: labeled training dataset 𝐷𝐷𝐿𝐿 = {𝑥𝑥𝑖𝑖,𝐿𝐿 , 𝑡𝑡𝑖𝑖,𝐿𝐿 , 𝑦𝑦𝑖𝑖,𝐿𝐿 }𝑖𝑖=1 , Euclidean distance 𝑑𝑑 [⋅,⋅], learning rate 𝛼𝛼 01: receive 𝐸𝐸𝜃𝜃 from pre-training stage 02: for 𝑖𝑖 in epochs 𝑁𝑁𝑤𝑤𝑤𝑤𝑤𝑤 ← 2 // Fix the number of categories N to 2 03: 04: 05: 06:

𝐿𝐿𝐹𝐹𝐹𝐹𝐹𝐹 ← 0 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 ← 0 (𝑆𝑆𝑘𝑘 , 𝑄𝑄𝑘𝑘 ) ← 𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅𝑅(𝐷𝐷𝐿𝐿 , 𝑁𝑁𝑤𝑤𝑤𝑤𝑤𝑤 , 𝐾𝐾𝑠𝑠ℎ𝑜𝑜𝑜𝑜 , 𝑁𝑁𝑄𝑄 ) // Select support sets and query sets

07: 08: 09: 10: 11: 12: 13:

for 𝒌𝒌 in �1, … , 𝑁𝑁𝑤𝑤𝑤𝑤𝑤𝑤 � do 𝑐𝑐𝑘𝑘 ← (|𝑆𝑆𝑘𝑘 |)−1 ∑(𝑥𝑥𝑖𝑖 ,𝑦𝑦𝑖𝑖 ,𝑡𝑡𝑖𝑖 )∈𝑆𝑆𝑘𝑘 𝐸𝐸𝜃𝜃 (𝑥𝑥𝑖𝑖 )

end for for (𝑥𝑥𝑖𝑖 , 𝑦𝑦𝑖𝑖 , 𝑡𝑡𝑖𝑖 ) in 𝑄𝑄𝑘𝑘 do 𝑑𝑑 ← 𝑑𝑑[𝐸𝐸𝜃𝜃 (𝑥𝑥𝑖𝑖 ), 𝑐𝑐𝑘𝑘 ] + 𝑙𝑙𝑙𝑙𝑙𝑙 ∑𝑘𝑘 ′ 𝑒𝑒𝑒𝑒𝑒𝑒(−𝑑𝑑[𝐸𝐸𝜃𝜃 (𝑥𝑥𝑖𝑖 ), 𝑐𝑐𝑘𝑘′ ]) 𝐿𝐿𝐹𝐹𝐹𝐹𝐹𝐹 ← 𝐿𝐿𝐹𝐹𝐹𝐹𝐹𝐹 ∙ (|𝑄𝑄𝑘𝑘 𝑆𝑆𝑘𝑘 |)−1 𝑑𝑑 end for

14: 15: 16: 17: 18: 19: 20: 21: 22: 23: 24: 25: 26: 27: 28: 29: 30: 31: 32: 33: 34: 35: 36:

for (𝑋𝑋𝑖𝑖 , 𝑌𝑌𝑖𝑖 , 𝑡𝑡𝑖𝑖 ) in 𝑆𝑆𝑘𝑘 do 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 ← 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝐸𝐸𝜃𝜃 (𝑋𝑋𝑖𝑖 ), 𝑡𝑡𝑖𝑖 ) 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 ← 𝐿𝐿𝐿𝐿𝐿𝐿_𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚(𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝) 𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 ← 𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺𝐺_ℎ𝑒𝑒𝑒𝑒𝑒𝑒_𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝐸𝐸𝜃𝜃 (𝑋𝑋𝑖𝑖 ), 𝑡𝑡𝑖𝑖 ) 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝 ← 𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶_ℎ𝑒𝑒𝑒𝑒𝑒𝑒_𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝐸𝐸𝜃𝜃 (𝑋𝑋𝑖𝑖 ), 𝑡𝑡𝑖𝑖 )

𝑇𝑇 ← 𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡𝑡_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝(𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑡𝑡𝑖𝑖 ) 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃_𝑖𝑖 ← 0 for each sampled point (𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) in 𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠𝑠_𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐 do �(𝑡𝑡𝑖𝑖 , 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) ← 𝑃𝑃𝜃𝜃 (𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔𝑔_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐𝑐_𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝𝑝, 𝑇𝑇𝑗𝑗 , 𝑡𝑡𝑖𝑖 , 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝑈𝑈 �, 𝑡𝑡𝑖𝑖 , 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃_𝑗𝑗 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃 (𝑈𝑈 �, 𝑡𝑡𝑖𝑖 , 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼_𝑗𝑗 ← 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼 (𝑈𝑈 �, 𝑡𝑡𝑖𝑖 , 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶_𝑗𝑗 ← 𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶 (𝑈𝑈 �, 𝑡𝑡𝑖𝑖 , 𝑥𝑥𝑗𝑗 , 𝑦𝑦𝑗𝑗 , 𝑧𝑧𝑗𝑗 ) 𝐿𝐿𝐺𝐺𝐺𝐺𝐺𝐺_𝑗𝑗 ← 𝐿𝐿𝐺𝐺𝐺𝐺𝐺𝐺 (𝑈𝑈 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃_𝑖𝑖 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃_𝑖𝑖 + 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃_𝑗𝑗 + 𝐿𝐿𝐼𝐼𝐼𝐼𝐼𝐼_𝑗𝑗 + 𝐿𝐿𝐶𝐶𝐶𝐶𝐶𝐶_𝑗𝑗 + 𝐿𝐿𝐺𝐺𝐺𝐺𝐺𝐺_𝑗𝑗

end for 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃_𝑖𝑖 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃_𝑖𝑖 /𝑁𝑁𝑆𝑆 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 + 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃_𝑖𝑖

end for 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 ← 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑁𝑁𝑖𝑖 /(|𝑆𝑆𝑘𝑘 |)−1

𝐿𝐿 ← 𝛽𝛽 ∙ 𝐿𝐿𝑃𝑃𝑃𝑃𝑃𝑃𝑃𝑃 + 𝐿𝐿𝐹𝐹𝐹𝐹𝐹𝐹 𝜃𝜃 ← 𝜃𝜃 − 𝛼𝛼∇𝜃𝜃 𝐿𝐿 end for return 𝐸𝐸𝜃𝜃

4. Experimental verification and analysis 4.1 Experimental results for pre-training stage The experimental framework in this study is implemented using PyTorch, with the training environment configured on an Intel i7-13900KF CPU, an NVIDIA RTX 2080 Ti GPU, and 32 GB of RAM. The training configuration employs stochastic gradient descent (SGD)[21] as the optimizer, initialized with a learning rate of 0.001, and adopts the OneCycleLR scheduling strategy for dynamic learning rate adjustment. A batch size of 256 is used over 800 training epochs, and automatic mixed precision (AMP) is enabled to enhance computational efficiency. To evaluate model performance on downstream classification tasks, the K-nearest neighbors algorithm (𝐾𝐾 = 2)[22] is applied to assess classification accuracy based on high-dimensional features extracted from the trained model and test data. Specifically, the model is first trained on a laser welding penetration dataset comprising two categories, subsequently, KNN calculates Top-1 accuracy to quantify the discriminative capability of the learned features. This approach not only measures the model's effectiveness but also verifies the quality of the high-dimensional feature representations in supporting classification tasks. Additionally, ablation experiments are conducted to investigate the contribution of individual components within the proposed SimPhysNet architecture. The results of these ablation experiments are summarized in Tab. 4 and illustrated in Fig. 10. Training dynamics,

including accuracy and loss trajectories across epochs, are depicted in Fig. 11. Tab. 4 The ablation study of SimPhysNet. Best results are displayed in bold and the runner-up results are underlined. No. Method Acc(%)↑ 1

Simsiam

83.15

2

+Rot

83.77(+0.62)

3

+Blur

84.59(+0.82)

4

+Crop

85.27(+0.68)

5

+Conical heat source

87.43(+2.16)

6

+ Goldak heat source

88.17(+0.74)

7

+Muti-scales

90.15(+1.97)

8

+ PDE

90.84(+0.69)

9

+ IBC

91.50(+0.66)

10

+Variable 𝛼𝛼

92.04(+0.54)

11

No Latin

90.88(-1.16)

12

points → 200

92.84(+1.96)

13

points → 500

93.68(+0.84)

14

points → 1000

94.55(+0.88)

15

Tanh → SoftPlus

94.90(+0.34)

16

Enhanced RMSE (SimPhysNet)(ours)

95.14(+0.24)

Fig. 10 The ablation study of SimPhysNet. The bar chart represents model accuracy, the shaded bars represent unimplemented modifications, and the line graph represents the performance differences between ach modification and the preceding ones. Experimental results indicate that the three image enhancement tasks proposed in this study contribute to varying degrees of performance improvement in the model. Notably, the Gaussian blur task achieves the most significant enhancement, which may be attributed to the prevalence of blurred images in the 𝐷𝐷𝑈𝑈 dataset, thereby underscoring the effectiveness of image augmentation strategies in improving model generalization. In the ablation experiments of the PINN framework, laser welding performance is found to be primarily influenced by the Conical heat source, yielding a 2.16% improvement in accuracy. The Goldak heat source also contributes to model performance, albeit to a lesser extent, with a marginal gain of 0.54%. Given that heat source dimensions may vary under different welding conditions and inspired by multi-scale architectures in deep learning, a multi-scale heat source approach—incorporating scales {0.25, 0.5, 1.0, 1.5, 2.0} is implemented to account for diverse heat source sizes across parameter settings, leading to a 1.77% performance gain. Furthermore, the inclusion of thermal conduction loss and boundary condition loss results in improvements of 0.69% and 0.66%, respectively. Considering that the material's thermal conductivity coefficient α is temperature-dependent in practical scenarios, a fully connected neural network layer is introduced to model the nonlinear relationship between thermal conductivity 𝛼𝛼

and temperature 𝑇𝑇 . Ablation experiments reveal that the proposed modification yields a 0.54% improvement in model performance. Notably, the exclusion of Latin hypercube sampling leads to a significant degradation in accuracy by -1.16%. The experimental results further indicate that increasing the number of sampling points enhances model accuracy; however, this also substantially increases model complexity. To achieve an optimal trade-off between predictive accuracy and computational cost, the number of sampling points is set to 1000 in this study. With regard to activation function optimization, replacing Tanh with SoftPlus ensures physically predictions— particularly non-negative values for heat source parameters resulting in an additional 0.35% performance gain. Furthermore, the improved RMSE algorithm contributes to more stable loss computation, resulting in a 0.24% improvement in overall model performance. Comparing the baseline Simsiam (No. 1 in Tab. 4, 83.15% accuracy) with the full SimPhysNet incorporating all PINN components (No. 14 in Tab. 4, 94.55% accuracy), the introduction of physical constraints directly yields a substantial improvement of 11.40% in classification accuracy. This significant gain underscores PINN's direct contribution to enhancing classification performance and implicitly improves model robustness by guiding feature learning towards physically consistent representations that are less prone to overfitting to spurious correlations, thereby better generalizing to real-world complexities. Furthermore, we trained several contrastive learning models, and Tab. 5 illustrates the accuracy of different contrastive learning algorithms on the 𝐷𝐷𝑈𝑈 dataset. The results indicate that the proposed pre-training method achieves higher accuracy, which can be attributed primarily to the constraints imposed by PINN on the model, as well as the enhancement of model generalization performance through three image augmentation tasks. These factors enable the pre-training stage to more effectively extract meaningful features from molten pool images. PINN can treat physical laws as a robust, unlabeled, and prior supervisory signal. In contrast to traditional unsupervised learning approaches, which mainly focus on uncovering the intrinsic structure of data, such as clustering or dimensionality reduction, without relying on external “labels” to guide the learning process, the core mechanism of PINN in this study involves loss function that integrates multiple residual terms derived from physical laws. The calculation of these residual terms does not depend on any labeled “ground truth” data; rather, it requires only that the outputs of SimPhysNet minimize these residuals. This self-supervised learning mechanism enables SimPhysNet to learn features that are both physically plausible and robust, even in the absence of explicit output labels. Tab. 5 Model accuracies of different networks. Best results are displayed in bold and the runnerup results are underlined. Method Acc(%)↑ P(%)↑ R(%)↑ F1(%)↑ MoCo[23]

76.46

71.15

80.07

75.37

MoCov2[24]

77.94

72.73

81.25

76.41

BYOL[25]

80.07

76.59

77.91

80.02

SimCLR[26]

80.33

77.18

79.61

79.09

SimCLRv2[27]

81.44

78.92

81.03

81.36

Simsiam[13]

83.15

80.97

85.38

83.12

SimPhysNet (Ours)

95.14

95.55

90.82

94.82

Fig. 11 The losses and accuracies of SimPhysNet use KNN. (a) model No.8;(b) model No.9;(c) model No.14;(d) model No.16. During the pre-training stage, Grad-CAM[28] was employed to visualize the trained model. It was observed that, as training progressed, the model gradually focused its attention on the molten pool and keyhole regions, demonstrating increasingly higher activation values in these areas relative to others. This suggests that the model effectively acquired the semantic features of molten pool images during pre-training. The corresponding attention maps are displayed in Fig. 12. Original

100 epoch

200 epoch

400 epoch

600 epoch

800 epoch

Fig. 12 Activation map produced by Grad-CAM for SimPhysNet. 4.2 Experimental results for fine-tuning stage Due to the limited number of categories in the laser welding penetration classification dataset, which consists of only two classes, namely penetrated and non-penetrated, this study evaluates the performance of SimPhysNet under 2-way 1-shot, 5-shot, 10-shot, and 20-shot settings. In each task,

two test classes (2-way) are selected consistently, with each class comprising 1, 5, 10, and 20 support samples (K-shot), respectively, along with 1,000 query samples. To evaluated model performance, 1,000 unseen tasks were randomly sampled from the test classes. The performance of the proposed SimPhysNet was further compared with several commonly used few-shot learning (FSL) methods. Under a 95% confidence interval, Tab. 6 reports the classification accuracies across the four aforementioned settings. As shown, SimPhysNet achieves the highest accuracy in all experimental configurations. Compared to the CSS method, SimPhysNet achieves an approximately 3% improvement. These results demonstrate that an effective pre-training strategy significantly contributes to model performance in few-shot classification tasks. Moreover, the observed performance gains validate the effectiveness of PINN in generating diverse feature representations within the embedding model, thereby enhancing generalization capabilities. Tab. 6 Performance comparison of the existing FSL methods. Best results are displayed in bold and the runner-up results are underlined. Method 1-shot 5-shot 10-shot 20-shot MatchingNet[29]

36.91 ± 0.63

62.75 ± 0.94

67.32 ± 0.85

72.96 ± 0.32

ProtoNet[30]

42.42 ± 0.74

70.20 ± 0.61

73.50 ± 0.70

76.89 ± 0.37

RelationNet[31]

53.44 ± 0.72

69.32 ± 0.57

72.00 ± 0.47

78.30 ± 0.30

ProtoTransfer[32]

38.74 ± 0.71

67.47 ± 0.81

72.46 ± 0.77

74.84 ± 0.31

ProtoNet+jig[33]

44.24 ± 0.77

70.04 ± 0.55

75.13 ± 0.63

77.37 ± 0.24

ProtoNet+rot[33]

57.78 ± 0.65

68.35 ± 0.69

73.81 ± 0.49

78.46 ± 0.15

ProtoNet+rot+jig[33]

59.33 ± 0.51

71.77 ± 0.67

76.43 ± 0.57

79.09 ± 0.21

CSS[34]

62.85 ± 0.84

72.08 ± 0.73

76.49 ± 0.39

80.59 ± 0.22

SimPhysNet (Ours)

65.30 ± 0.45

75.07 ± 0.36

80.10 ± 0.26

85.64 ± 0.18

Additionally, this study investigates the impact of support set size 𝑁𝑁𝑠𝑠 on classification accuracy. The model's performance is presented in Fig. 13, which illustrates the prediction accuracy as the support set size increases from 1 to 200. As shown in the figure, the classification accuracy improves progressively with increasing 𝑁𝑁𝑠𝑠 , achieving 97.48% of the accuracy attained by the supervised learning model when 𝑁𝑁𝑠𝑠 reaches 200. When the support set is small, the model exhibits relatively lower classification accuracy. This phenomenon can be attributed to the fact that the prototype network calculates the distance between the features of the query image and the class prototypes derived from the support set. When the support set adequately covers the various states of the welding process, it more accurately represents the underlying distribution of sample images, thereby enhancing classification performance. Conversely, a limited support set leads to less reliable prototype estimation and consequently degrades accuracy. In summary, experimental results demonstrate that the proposed algorithm effectively learns discriminative features from molten pool images during the pre-training stage. Moreover, by leveraging the prototype network's few-shot fine-tuning strategy, the model achieves high classification accuracy even with a limited number of labeled samples. The training curves during the fine-tuning process are presented in Fig. 14

Fig. 13 The results of different training-key N in SimPhysNet’s fine-tuning stage.

Fig. 14 The losses and accuracies of SimPhysNet on 2-way. 4.3 Compare to supervised learning methods The algorithm proposed in this paper aims to address the issue in supervised learning algorithms that rely on a large amount of high-quality labeled data. Therefore, it is necessary to compare the performance of both approaches. The supervised model is trained on the 𝐷𝐷𝐿𝐿 dataset, with data augmentation methods consistent with those used in SimPhysNet. The dataset is split into training and validation sets with a ratio of 7:3. The supervised learning model employed is ResNet34, which achieves an accuracy of 98.84% after 100 training epochs. It is observed that the supervised model exhibits higher accuracy compared to SimPhysNet. Notably, the proposed contrastive learning pre-training combined with a few-shot fine-tuning strategy demonstrates excellent performance even when the number of labeled samples is significantly reduced: when the labeled data amounts to only 100 samples, the classification accuracy decreases by 4.97% to 93.87%; when the sample size increases to 200, the accuracy reaches 96.35% (a decrease of only 2.49%). These experimental results indicate that the feature representation space obtained through contrastive learning pre-training stage strong generalization capabilities. When combined with a few-shot learning strategy, it can approach the performance of supervised learning. This is particularly advantageous in industrial settings where labeled resources are limited, as it maintains

high classification accuracy. Consequently, the method demonstrates substantial engineering application value in practical manufacturing scenarios. The algorithm proposed in this paper addresses a critical limitation of supervised learning algorithms that rely on large volumes of high-quality labeled data. To evaluate its effectiveness, a comparative analysis of performance between the proposed approach and conventional supervised learning is conducted. The supervised model is trained on the 𝐷𝐷𝐿𝐿 dataset, with data augmentation techniques aligned with those utilized in SimPhysNet. The dataset is partitioned into training and validation subsets in a 7:3 ratio. A ResNet-34 architecture is adopted as the supervised learning model, achieving a classification accuracy of 98.84% after 100 epochs of training. The results demonstrate that the supervised model outperforms SimPhysNet in terms of accuracy. Notably, the proposed method—contrastive learning pre-training followed by few-shot fine-tuning— demonstrates robust performance even in scenarios with limited labeled data. Specifically, when only 100 labeled samples are available, the classification accuracy declines marginally to 93.87% (a reduction of 4.97%); with 200 labeled samples, accuracy reaches 96.35%, representing a decrease of merely 2.49%. These findings suggest that the feature representation space derived from contrastive learning pre-training possesses robust generalization capabilities. When integrated with a few-shot learning strategy, the method can closely approximate the performance of fully supervised learning. This characteristic is particularly beneficial in industrial environments where labeled data are scarce, enabling high classification accuracy with minimal annotation effort. Beyond accuracy, the practical utility of SimPhysNet hinges on its real-time performance. We measured the inference speed of the trained SimPhysNet model on an NVIDIA RTX 2080 Ti GPU The average inference time per molten pool image was found to be approximately 5.13ms, demonstrating its capability for rapid prediction. This high inference speed, combined with its robust performance on limited labeled data, confirms SimPhysNet's significant potential for real-time, online monitoring and control in laser welding processes. UMAP [35] visualization was performed on both the supervised learning model and SimPhysNet, as depicted in Fig. 15. The figure illustrates that, as the number of pre-training epochs increases, SimPhysNet gradually separates the two categories of images into distinct feature clusters, achieving intra-class compactness within each cluster. The pre-training stage enables the model to effectively extract discriminative features from molten pool images. Consequently, during the subsequent fine-tuning stage with limited labeled samples, only minor adjustments to the network weights are required, allowing SimPhysNet to minimize intra-class distance while maximizing inter-class distance. These results indicate that SimPhysNet can accurately classify penetration states, thereby validating the effectiveness and reliability of the proposed approach.

Fig. 15 UMAP was utilized to visualize the output features of the model, with points of varying colors representing different classes. (a)-(e) depict the results of SimPhysNet under a 2-way 200shot setting. (f) depicts the performance of the supervised learning (ResNet-34) on the training set. (For an explanation of the color references in the legend of this figure, please refer to the online version of this paper.) Fig. 16 presents the attention maps produced by the supervised learning algorithm and SimPhysNet. A comparative analysis of their attention distributions over the molten pool images reveals that both the keyhole and the molten pool regions play a critical role in the classification decisions made by the models. The high degree of similarity between the attention patterns generated by the two algorithms indicates that SimPhysNet effectively captures and learns salient visual features from the input images, demonstrating its capacity to emulate the behavior of fully supervised models.

Original

ResNet-34

SimPhysNet

Fig. 16 Activation map produced by Grad-CAM for supervised learning models, including ResNet-34 and SimPhysNet. To more clearly demonstrate the performance across the three image enhancement tasks, we simultaneously visualized the results of all three enhancement methods. Fig. 17 presents the attention maps produced by SimPhysNet under different rotation angles. It is evident that, regardless of the rotation angle, SimPhysNet consistently and accurately captures the molten pool, keyhole,

and spatter regions. This suggests that SimPhysNet enhances the informational content of the images by integrating rotational transformations, thereby enabling more robust feature representations during training process. Original

Rot 0°

Rot 45°

Rot 90°

Rot 135°

Rot 180°

Rot 0°

Rot 45°

(a)

Rot 90°

Rot 135°

Rot 180°

(b)

Fig. 17 The results of the attention maps generated by Grad-CAM across various rotation angles. (a) Visualization results of SimPhysNet and (b) is generated by supervised learning (ResNet-34). Fig. 18 illustrates the attention maps generated by SimPhysNet under various gaussian blur parameters. The first row presents original images without gaussian blur, the second row shows results with different levels of artificially applied gaussian blur, and the fourth row displays images blurred during actual welding processes. Compared to the unblurred images, SimPhysNet exhibits minimal shift in attention localization across varying blur intensities. However, the spatial extent of attention gradually expands as the intensity of blurring increases. This expansion is attributed to the suppression of fine-grained details by gaussian blur, which leads the model to rely more heavily on global structural information. Notably, when applied to real-world welding images, SimPhysNet successfully captures regions containing critical features. These observations indicate that incorporating gaussian blur during training enhances the model's robustness and generalization capability in practical welding scenarios. Original

SimPhysNet

ResNet-34

Original

SimPhysNet

ResNet-34

Fig. 18 The results of attention maps generated by Grad-CAM under various Gaussian blur parameters using supervised learning (ResNet-34) and SimPhysNet. Fig. 19 presents the attention maps generated by our method under different cropping configurations. The figure comprises six images, each corresponding to a distinct cropping position. A comparative analysis of the attention maps reveals that image cropping encourages the model to

extract discriminative features from fine-grained image details. Consequently, the model demonstrates heightened focuses on the molten pool boundaries and smaller-scale spatter regions. Original

SimPhysNet

ResNet-34

Original

SimPhysNet

ResNet-34

Fig. 19 The results of attention maps generated by Grad-CAM at Various Cropping Positions using supervised learning (ResNet-34) and SimPhysNet. Through the visualization analysis of the model attention (Fig. 17,Fig. 18 and Fig. 19), we clearly demonstrate the advantages of SimPhysNet. The Grad-CAM results reveal its exceptional robustness achieved through image enhancement tasks. Whether the images are rotated, blurred, or partially occluded, the model accurately captures critical information. This series of visualization results compellingly confirms that the pre-trained framework employed by SimPhysNet effectively enhances the quality of feature extraction and the model's generalization performance, thereby ensuring its accuracy and reliability in practical welding applications. Furthermore, to more effectively evaluate the performance of SimPhysNet, we conducted comparative experiments by training SVM[36], Random Forest[37], and XGBoost[38] models on the same 𝐷𝐷𝐿𝐿 dataset. The data augmentation and dataset partitioning strategies were kept consistent with those applied during supervi; sed learning. As shown in Tab. 7, the proposed training framework demonstrates superior feature representation capability. Moreover, compared to the baseline methods, the fine-tuning process achieves higher classification accuracy using only a limited number of labeled samples, which indicates that the pre-training stage effectively captures discriminative features. Experimental results indicate that, in comparison to conventional machine learning approaches, the proposed method exhibits enhanced generalization ability and adaptability across different settings. Tab. 7 Model accuracies of different networks. Best results are displayed in bold and the runnerup results are underlined. Pre-training Method Acc(%)↑ P(%)↑ R(%)↑ F1(%)↑ features SVM[36] 88.13 83.66 87.20 87.96 Random forest[37]

-

85.18

85.53

85.13

82.59

XGBoost[38]

-

86.22

90.85

89.54

88.88

ResNet-34[14]

-

98.84

97.87

97.62

98.53

SVM[36]

81.45

86.11

90.04

90.48

Random forest[37]

89.08

88.83

87.39

86.71

XGBoost[38]

90.32

93.10

92.74

92.45

SimPhysNet (Our proposed)

95.14

95.55

90.82

94.82

In summary, SimPhysNet demonstrates higher accuracy than other algorithms reported in the literature, particularly when trained on smaller datasets. This approach effectively mitigates the high annotation costs typically associated with deep learning methods and exhibits superior generalization performance and adaptability in data-limited scenarios. By leveraging a large corpus of unlabeled data through a physics-informed self-supervised framework, SimPhysNet achieves a classification accuracy of 96.35% using only 5% of the total labeled dataset. This performance, representing a marginal decline of just 2.49% relative to a fully supervised model, demonstrates that the vast amounts of unlabeled process data typically generated in industrial environments can be effectively utilized to construct highly accurate predictive models.

5. Conclusion Laser welding penetration classification typically relies on a large volume of labeled data. To address this limitation, this paper proposes a deep learning algorithm, termed SimPhysNet, which integrates two key components: a contrastive learning-based pre-training feature extractor and a few-shot classification-based feature classifier. The main contributions are summarized as follows: 1)

2)

3)

By embedding physical priors governed by partial differential equations into the contrastive learning loss function, a physics-informed feature extraction framework is developed. This approach incorporates physical constraint terms to guides the learned feature space toward alignment with the underlying energy transfer dynamics of the actual welding process. Consequently, the model achieves enhanced prediction accuracy and improved interpretability, thereby increasing the reliability of its outputs. Design of three image augmentation strategies for pre-training: Three supervised pretext tasks are introduced during the pre-training stage, which are rotation prediction, Gaussian blur level prediction, and random cropping localization. Rotation and blur prediction tasks encourage the network to capture global structural features, while random cropping localization promotes sensitivity to fine-grained details in weld pool images, such as melt pool boundaries. Experimental results demonstrate that these tasks effectively facilitate the contrastive learning process, enabling the network to extract discriminative features from molten pool imagery and improving performance in downstream classification tasks. Building upon the concept of prototype networks, a laser welding penetration classifier was developed. This classifier utilizes the robust feature space established during pretraining to compute class-specific prototypes. By employing metric-based inference, it attains exceptional few-shot classification performance. This is demonstrated by the model's high accuracy across multiple scenarios, achieving 65.30% in a 2-way 1-shot setting. More notably, when trained with only 5% of the labeled dataset, the model maintains an accuracy of 96.35%, which is highly comparable to its fully supervised

counterpart. his finding suggests a viable pathway to circumvent the data annotation bottleneck, substantially reducing the cost and effort required to deploy intelligent monitoring systems in practical manufacturing scenarios. The ability to approach supervised performance with minimal labeled data signifies the considerable engineering application value and scalability of the proposed methodology.

Declarations Ethics approval Not applicable. Consent to participate The authors declare their consent to participate. Consent for publication The authors declare their consent for publication. Conflict of interest The authors declare no competing interests.

Funding The authors gratefully acknowledge the financial support from the National Key R&D Program of China (No. 2023YFB3407800), National Natural Science Foundation of China (No. U2141213).

References [1] B. Wang, S.J. Hu, L. Sun, T. Freiheit, Intelligent welding system technologies: Stateof-the-art review and perspectives, J. Manuf. Syst. 56 (2020) 373–391. https://doi.org/10.1016/j.jmsy.2020.06.020. [2] W. Cai, J. Wang, P. Jiang, L. Cao, G. Mi, Q. Zhou, Application of sensing techniques and artificial intelligence-based methods to laser welding real-time monitoring: A critical review of recent literature, J. Manuf. Syst. 57 (2020) 1–18. https://doi.org/10.1016/j.jmsy.2020.07.021. [3] S. Kang, S. Jeon, K. Ryu, J. Shin, Online monitoring of weld cross-sectional shape using optical emission spectroscopy and neural network during laser dissimilar welding, Eng. Appl. Artif. Intell. 141 (2025) 109847. https://doi.org/10.1016/j.engappai.2024.109847. [4] H. Li, H. Ren, Z. Liu, F. Huang, G. Xia, Y. Long, In-situ monitoring system for weld geometry of laser welding based on multi-task convolutional neural network model, Measurement 204 (2022) 112138. https://doi.org/10.1016/j.measurement.2022.112138. [5] X. Zhang, C. Wang, Y. Tang, Z. Zhou, X. Lu, A Survey of Few-Shot Learning and Its Application in Industrial Object Detection Tasks, in: Y. Wang, K. Martinsen, T. Yu, K.

Wang (Eds.), Lect. Notes Electr. Eng., Springer, Singapore, 2022: pp. 637–647. https://doi.org/10.1007/978-981-19-0572-8_81. [6] T. Raffin, A. Mayr, M. Baader, N. Laube, A. Kühl, J. Franke, Potentials of few-shot learning for quality monitoring in laser welding of hairpin windings, Procedia CIRP 118 (2023) 901–906. https://doi.org/10.1016/j.procir.2023.06.155. [7] T. Zhu, S. Zhu, J. Zhu, W. Song, C. Li, H. Ge, J. Gu, A Deep Meta-Metric Learning Method for Few-Shot Weld Seam Visual Detection, in: 2022 IEEE Int. Conf. Robot. Biomim. ROBIO, 2022: pp. 1167–1173. https://doi.org/10.1109/ROBIO55434.2022.10012017. [8] Q. Liu, R. Xiao, Y. Xu, J. Xu, S. Chen, A defect classification algorithm for gas tungsten arc welding process based on unsupervised learning and few-shot learning strategy, J. Manuf. Process. 131 (2024) 1219–1229. https://doi.org/10.1016/j.jmapro.2024.09.084. [9] S. Chen, T. Li, F. Jiang, G. Zhang, S. Fang, Enhancing VPPA welding quality prediction: A hybrid model integrating prior physical knowledge and CNN analysis, J. Manuf. Process. 131 (2024) 1282–1295. https://doi.org/10.1016/j.jmapro.2024.09.089. [10] R. Lu, M. Lou, Y. Xia, Y. Li, Real-time penetration depth prediction via physicsinformed learning from molten pool surface morphology in laser filler wire welding, J. Intell. Manuf. (2025). https://doi.org/10.1007/s10845-025-02738-7. [11] X. Li, Z. Fu, J. Shu, B. Ji, B. Ji, A modified physics-informed neural network to fatigue life prediction of deck-rib double-side welded joints, Int. J. Fatigue 189 (2024) 108566. https://doi.org/10.1016/j.ijfatigue.2024.108566. [12] A.D. Gianfrancesco, Materials for Ultra-Supercritical and Advanced UltraSupercritical Power Plants, Woodhead Publishing, 2016. [13] X. Chen, K. He, Exploring Simple Siamese Representation Learning, (2020). https://doi.org/10.48550/arXiv.2011.10566. [14] K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: 2016: pp. 770–778. https://doi.org/10.1109/CVPR.2016.90. [15] M. Behúlová, E. Babalová, Heat source models for numerical simulation of laser welding processes – a short review, J. Phys. Conf. Ser. 2712 (2024) 012018. https://doi.org/10.1088/1742-6596/2712/1/012018. [16] P. Hartwig, L. Scheunemann, J. Schröder, On the numerical treatment of heat sources in laser beam welding processes, PAMM 23 (2023) e202200220. https://doi.org/10.1002/pamm.202200220. [17]

G. Liang, G. Qin, P. Cao, H. Wang, Evolutions of temperature field and stress

field in narrow gap oscillating laser welding process based on equivalent heat source, J. Mater. Res. Technol. 28 (2024) 154–167. https://doi.org/10.1016/j.jmrt.2023.11.262. [18] S. Wang, S. Sankaran, H. Wang, P. Perdikaris, An Expert’s Guide to Training Physics-informed Neural Networks, (2023). https://doi.org/10.48550/arXiv.2308.08468. [19] W. Guo, F. Deng, Z. Meng, L. Hua, H. Mao, J. Su, A hybrid back-propagation neural network and intelligent algorithm combined algorithm for optimizing microcellular foaming injection molding process parameters, J. Manuf. Process. 50 (2020) 528–538. https://doi.org/10.1016/j.jmapro.2019.12.020. [20] D.E. Huntington, C.S. Lyrintzis, Improvements to and limitations of Latin hypercube sampling, Probabilistic Eng. Mech. 13 (1998) 245–253. https://doi.org/10.1016/S0266-8920(97)00013-1. [21] S. Ruder, An overview of gradient descent optimization algorithms, (2017). https://doi.org/10.48550/arXiv.1609.04747. [22] T. Cover, P. Hart, Nearest neighbor pattern classification, IEEE Trans. Inf. Theory 13 (1967) 21–27. https://doi.org/10.1109/TIT.1967.1053964. [23] K. He, H. Fan, Y. Wu, S. Xie, R. Girshick, Momentum Contrast for Unsupervised Visual Representation Learning, (2020). https://doi.org/10.48550/arXiv.1911.05722. [24] X. Chen, H. Fan, R. Girshick, K. He, Improved Baselines with Momentum Contrastive Learning, (2020). https://doi.org/10.48550/ARXIV.2003.04297. [25] J.-B. Grill, F. Strub, F. Altché, C. Tallec, P.H. Richemond, E. Buchatskaya, C. Doersch, B.A. Pires, Z.D. Guo, M.G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, M. Valko, Bootstrap your own latent: A new approach to self-supervised Learning, (2020). https://doi.org/10.48550/arXiv.2006.07733. [26] T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A Simple Framework for Contrastive Learning of Visual Representations, (2020). https://doi.org/10.48550/arXiv.2002.05709. [27] T. Chen, S. Kornblith, K. Swersky, M. Norouzi, G.E. Hinton, Big SelfSupervised Models are Strong Semi-Supervised Learners, in: Adv. Neural Inf. Process. Syst., Curran Associates, Inc., 2020: pp. 22243–22255. https://proceedings.neurips.cc/paper/2020/hash/fcbc95ccdd551da181207c0c1400c655Abstract.html (accessed August 14, 2025). [28] R.R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, GradCAM: Visual Explanations from Deep Networks via Gradient-based Localization, Int. J.

Comput. Vis. 128 (2020) 336–359. https://doi.org/10.1007/s11263-019-01228-7. [29] (n.d.).

O. Vinyals, C. Blundell, T. Lillicrap, Matching Networks for One Shot Learning,

[30] J. Snell, K. Swersky, R. Zemel, Prototypical Networks for Few-shot Learning, in: Adv. Neural Inf. Process. Syst., Curran Associates, Inc., 2017. https://proceedings.neurips.cc/paper_files/paper/2017/hash/cb8da6767461f2812ae4290eac7 cbc42-Abstract.html (accessed July 2, 2025). [31] F. Sung, Y. Yang, L. Zhang, T. Xiang, P.H.S. Torr, T.M. Hospedales, Learning to Compare: Relation Network for Few-Shot Learning, in: 2018 IEEECVF Conf. Comput. Vis. Pattern Recognit., 2018: pp. 1199–1208. https://doi.org/10.1109/CVPR.2018.00131. [32] C. Medina, A. Devos, M. Grossglauser, Self-Supervised Prototypical Transfer Learning for Few-Shot Classification, (2020). https://doi.org/10.48550/arXiv.2006.11325. [33] J.-C. Su, S. Maji, B. Hariharan, When Does Self-supervision Improve Few-shot Learning?, (2020). https://doi.org/10.48550/arXiv.1910.03560. [34] Y. An, H. Xue, X. Zhao, L. Zhang, Conditional Self-Supervised Learning for Few-Shot Classification, in: Proc. Thirtieth Int. Jt. Conf. Artif. Intell., International Joint Conferences on Artificial Intelligence Organization, Montreal, Canada, 2021: pp. 2140–2146. https://doi.org/10.24963/ijcai.2021/295. [35] L. McInnes, J. Healy, J. Melville, UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, (2020). https://doi.org/10.48550/arXiv.1802.03426. [36] C. Cortes, V. Vapnik, Support-vector networks, Mach. Learn. 20 (1995) 273– 297. https://doi.org/10.1007/BF00994018. [37] L. Breiman, Random Forests, https://doi.org/10.1023/A:1010933404324.

Mach.

Learn.

45

(2001)

5–32.

[38] T. Chen, C. Guestrin, XGBoost, in: Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., 2016: pp. 785–794. https://doi.org/10.1145/2939672.2939785.

Record · ID 307046 · SHA-256 6a0dbc1c5a3d46bb
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.