Conceptio › Archive › arXiv CS
arXiv CSopen access

Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

1

Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data

arXiv:2609.21932v1 [cs.LG] 18 Sep 2026

Khoa Tran† , Ho-Si-Hung Nguyen* , Phone Wai Yan Moe† , Hung-Cuong Trinh, and Thi-Hoang-Giang Tran

Abstract—Joint remaining useful life (RUL) prediction and capacity estimation require representations of both gradual degradation and recent battery behavior. This paper presents a cross-expert framework using partial-charging measurements without measured historical full-cycle capacity as an input. The RUL Expert encodes nominal 10-min segments from ten cycles sampled within a 30-cycle history using a pretrained gated recurrent unit (GRU) encoder, a two-dimensional convolutional neural network (2D-CNN), and a temporal GRU. The Capacity Expert processes statistical descriptors of nominal 40-min segments from ten consecutive cycles using a 2D-CNN and a Transformer. A feature-wise linear modulation module uses the short-term representation to condition the long-term representation for joint prediction. Training comprises supervised autoencoder pretraining, independent expert pretraining, and fusion training with frozen experts. On two public battery-aging datasets, the reference configuration achieves mean RUL rootmean-square errors of 143.69 and 161.10 cycles and capacity errors of 12.36 and 7.28 mAh, respectively. On Dataset I, fusion reduces both mean errors relative to either standalone expert. The results demonstrate a trade-off between RUL and capacity accuracy: the proposed method attains the lowest reported RUL RMSE among the compared methods on both datasets, whereas several baselines yield lower capacity errors. Index Terms—Battery capacity estimation, feature-wise linear modulation, lithium-ion batteries, multi-task learning, partial charging, remaining useful life.

I. I NTRODUCTION Lithium-ion batteries (LIBs) are widely used in electric vehicles (EVs) [4] and stationary energy-storage systems because of their high energy density, long cycle life, and established manufacturing infrastructure. Transportation electrification and renewable-energy integration are increasing the demand for reliable battery systems. For example, [23] projects that battery * Khoa Tran and Phone Wai Yan Moe contributed equally to this work. † Corresponding author: Ho-Si-Hung Nguyen (e-mail: [email protected]). Khoa Tran is with the Data Science Laboratory, Faculty of Information Technology, Ton Duc Thang University, Ho Chi Minh City 70000, Vietnam (e-mail: [email protected]). Ho-Si-Hung Nguyen is with the Faculty of Electrical Engineering, The University of Danang—University of Science and Technology, 54 Nguyen Luong Bang, Lien Chieu, Da Nang 550000, Vietnam (e-mail: [email protected]). Phone Wai Yan Moe is with AIWARE Limited Company, 17 Huynh Man Dat Street, Hoa Cuong Bac Ward, Hai Chau District, Da Nang 550000, Vietnam (e-mail: [email protected]). Hung-Cuong Trinh is with the Natural Language Processing and Knowledge Discovery Research Group, Faculty of Information Technology, Ton Duc Thang University, Ho Chi Minh City 70000, Vietnam (e-mail: [email protected]). Thi-Hoang-Giang Tran is with the Faculty of Project and Industrial Management, The University of Danang—University of Science and Technology, 54 Nguyen Luong Bang, Lien Chieu, Da Nang 550000, Vietnam (e-mail: [email protected]).

demand for e-mobility and stationary storage in the European Union will reach approximately 1 TWh by 2030. Reliable assessment of battery degradation is therefore important for managing an expanding population of batteries throughout their service lives. Battery management systems (BMSs) [8] monitor battery condition to support operating and maintenance decisions and avoid unnecessary replacement. Two complementary indicators are state of health (SOH) [6], [7], [15] and remaining useful life (RUL) [1], [2], [12]. Capacity-based SOH quantifies the available capacity relative to the nominal capacity, whereas cycle-based RUL describes the remaining cycles before end of life (EOL): Qc × 100%, (1) SOHc = Qnom RULc = cEOL − c, (2) where Qc is the available capacity at cycle c, Qnom is the nominal capacity, and cEOL is the cycle at which the EOL criterion is reached. Joint capacity estimation and RUL prediction therefore characterize both the present condition and the anticipated service life of a battery. Model-based approaches, including equivalent circuit models (ECMs) [3], [11], relate measured signals to battery states through parameterized representations. Their performance depends on model assumptions, parameter identification, and state initialization, which can become challenging as operating conditions change. Data-driven methods provide an alternative by learning degradation-related relationships directly from observations. However, their accuracy and practical applicability depend on the representativeness of the training data and the availability of the required measurements during operation. Recent RUL prediction methods investigate different representations of battery degradation. Ge et al. [9] introduce a Koopman-inspired Degradation Model with Embedded Operating Conditions (EKIDM). Charging or relaxation voltage sequences and operating-condition variables are encoded into a latent degradation representation, while a condition-dependent Koopman operator models its evolution. Wu et al. [24] use multiple battery signals within a network employing deformable depthwise convolution. Learnable sampling offsets adapt feature extraction to the observed signals, providing a flexible approach to degradation modeling. Yu et al. [25] extract degradation features from the first 100 complete cycles, primarily using discharge capacity–voltage curves, Q(V ). Features at multiple time scales are selected and processed by a Transformer-based hybrid network. Although such observations can support early-life prediction, complete

2

cycling records and reliable discharge-capacity measurements may be unavailable during routine EV operation. Moreover, predictions based on early operating conditions may require reassessment when subsequent usage differs from the training conditions. For capacity forecasting, Qian et al. [18] combine empirical mode decomposition (EMD), a cycle-consistent adversarial network (CycleGAN), and a condition-focused Transformer. Operating-condition signals are decomposed into intrinsic mode functions and transformed into representations used to generate target-condition signals. These signals and historical capacity sequences are then used for forecasting. Peng et al. [16] propose the MIC-BOA-TiDE framework, which combines historical capacity, battery attributes, and selected health indicators. The maximal information coefficient guides feature selection, a dense encoder–decoder forecasts capacity, and kernel density estimation provides prediction intervals from the error distribution. Both approaches illustrate the usefulness of historical capacity information, but obtaining accurate fullcycle capacity measurements during routine operation remains challenging. Partial-charging measurements offer a more accessible source of degradation information. Ma et al. [14] use voltage, accumulated charge, and their differences within a charging interval from 80% state of charge (SOC) to the first occurrence of 3.6 V. An important distinction is whether accumulated charge is referenced to the beginning of the full charging cycle or to the start of the observed segment. If the former reference is retained, values within a partial segment can depend on measurements collected before that segment. This dependence can complicate deployment when charging begins at different SOC levels. Segment-local integration addresses this input requirement by resetting accumulated charge at the beginning of the available observation, without requiring the preceding charging trajectory. Song et al. [21] further demonstrate the potential of partialcharging information for SOH estimation. Their framework extracts the charging time over a specified voltage interval, voltage standard deviation, and root-mean-square voltage from constant-current charging. Kernel principal component analysis combines these indicators, and a Bayesian LSTM estimates SOH and predictive variance. A collaborative learning strategy transfers degradation information across operating conditions. This approach supports the practicality of charging-derived indicators, while highlighting the value of uncertainty assessment under changing conditions. These studies motivate three considerations for practical joint prediction. First, the observation history and charging protocol determine what information is available to a prognostic model, as illustrated by early-cycle and partialcharging formulations [14], [25]. Second, methods developed primarily for RUL or capacity/SOH emphasize different prediction objectives [9], [18], [21]. Although joint prediction has been investigated [14], it remains useful to examine whether complementary temporal representations can improve the balance between the two tasks. Third, the reference and availability of capacity-related inputs must be explicit: accumulated charge derived from measurements and historical full-

cycle capacity are distinct quantities with different acquisition requirements [14], [16], [18]. The contrast between completecycle features and partial-charging indicators also motivates reducing observation requirements without assuming that every existing method requires complete cycles [21], [25]. To address these considerations, this study proposes a crossexpert framework comprising an RUL Expert, a Capacity Expert, and a feature-wise linear modulation (FiLM) module [17]. The RUL Expert learns cycle-level embeddings from short charging segments sampled across an extended cycle history. The Capacity Expert processes statistical descriptors from longer charging segments over recent consecutive cycles. Both experts are pretrained to predict RUL and capacity; their names indicate their intended temporal roles rather than exclusive prediction targets. During fusion, the shortterm representation generates feature-wise scaling and shifting vectors that modulate the long-term representation. A shared two-output regression head then estimates RUL and capacity. Neither branch requires measured historical full-cycle capacity as an input, although capacity and RUL labels are required during supervised training. The main contributions are summarized as follows: 1) Complementary partial-charging representations: Nominal 10-min charging segments from ten cycles sampled at stride three within a 30-cycle window are paired with nominal 40-min segments from ten consecutive cycles. Both views begin at the first observed charging sample above 3.1 V. Accumulated charge and elapsed time are referenced to each segment’s start. Lag-12 differences are computed within the 30-cycle window, with unavailable differences set to zero. 2) Cross-expert feature fusion: A pretrained gated recurrent unit (GRU) encoder, a two-dimensional convolutional neural network (2D-CNN), and a temporal GRU process the long-term view. A 2D-CNN and Transformer process the short-term statistical descriptors. FiLM conditions the long-term representation on the short-term representation for joint RUL prediction and capacity estimation. 3) Three-stage optimization: Supervised GRUautoencoder pretraining combines reconstruction with RUL and capacity supervision. Independent expert pretraining then retains the transferred encoder as a frozen feature extractor. Finally, both experts are frozen while the FiLM module and shared regression head are trained. 4) Experimental assessment: Comparisons on two public battery-aging datasets and ablation studies on Dataset I examine input representations, charging-segment duration, cycle history, temporal backbones, and fusion strategies. The analysis evaluates RUL improvements alongside capacity-estimation trade-offs. II. P ROPOSED M ETHOD As illustrated in Fig. 1, the proposed framework comprises an RUL Expert, a Capacity Expert, and a feature-wise linear modulation (FiLM) module for cross-expert knowledge fusion.

3

The two experts capture complementary temporal contexts. The RUL Expert processes nominal 10-min partial charging segments sampled from a 30-cycle window to characterize gradual degradation. Lag-12 inter-cycle differences are computed within this window. The Capacity Expert processes nominal 40-min charging segments from the 10 most recent consecutive cycles to characterize short-term battery behavior. The framework is optimized in three stages: 1) supervised GRU-autoencoder pretraining, 2) independent multi-task pretraining of both experts, and 3) FiLM-based fusion training with both experts frozen. Both experts are pretrained to predict RUL and capacity; their names identify their respective temporal representations rather than exclusive prediction targets. Predictions are issued after the required charging segments have been observed. Each fusion sample pairs the two temporal views with the same RUL and capacity targets. A. RUL Expert The RUL Expert characterizes degradation over an extended cycle history. Its pipeline consists of long-term data preparation, a pretrained gated recurrent unit (GRU) encoder [5] for cycle-level representation learning, a 2D-CNN for local feature extraction, and a temporal GRU for modeling degradation across cycles. 1) Long-Term Data Processing: For cycle c, the measured current, voltage, and temperature sequences are denoted by Ic , Vc , and Tc , with corresponding timestamps tc . Bold symbols denote sequences, whereas Ic,j , Vc,j , and Tc,j denote individual samples. A nominal 10-min charging segment is extracted beginning at the first sample satisfying V > 3.1 V. Using partial charging segments supports practical batterymanagement applications in which complete charge–discharge records may be unavailable. Let Mc be the number of samples in the extracted segment. A second-order Savitzky–Golay filter [19] G(·) is independently applied to current, voltage, and temperature: Īc = G(Ic ),

V̄c = G(Vc ), T̄c = G(Tc ).

(3)

The odd filter window is selected as approximately 25% of Mc and is clipped to satisfy the segment-length and polynomialorder constraints. Elapsed time is measured relative to the beginning of the segment: τc,j = tc,j − tc,0 ,

j = 0, . . . , Mc − 1.

(4)

The charge accumulated within the observed segment is obtained by numerical integration of the filtered current: QL c,0 = 0, QL c,j =

j X

I¯c,k (τc,k − τc,k−1 ) ,

j = 1, . . . , Mc − 1.

(5)

k=1

Current and time are expressed in compatible units, such that amperes and hours yield ampere-hours. Here, QL c,j is the charge accumulated only within the observed partial segment and is distinct from the full-cycle capacity used as a prediction target.

The four long-term channels are assembled as  L  V̄c T̄c τ c ∈ RMc ×4 . ZL c = Qc

(6)

Each channel is linearly interpolated over sample positions to a fixed length of 50: L 50×4 XL , c = I50 (Zc ) ∈ R

(7)

where I50 (·) denotes channel-wise interpolation. To explicitly represent degradation changes, a 12-cycle lag difference is computed within each 30-cycle observation window. Let XL w,p denote the representation at chronological position p ∈ {0, . . . , 29} in window w. Then ( 0, p < 12, (8) ∆XL w,p = L L Xw,p − Xw,p−12 , p ≥ 12. The first 12 difference tensors are zero because their reference cycles precede the window. No observations outside the 30cycle window are used. The base and difference channels are concatenated as  e L = Concatch XL , ∆XL ∈ R50×8 . X (9) w,p w,p w,p Min–max normalization ML is applied independently to each resampled-position–channel pair. Its parameters are estimated exclusively from the training set and subsequently reused for validation and testing:   50×8 b L = ML X eL X . (10) w,p w,p ∈ R Ten cycles are selected at stride three from the 30-cycle window, i.e., pk = 3k for k = 0, . . . , 9. Thus, positions 28 and 29 are not selected for this branch. The long-term model input is therefore   L L L b b b HL = Stack X , X , . . . , X cycle w w,0 w,3 w,27 (11) 10×50×8 ∈R . 2) GRU-Autoencoder: The GRU-autoencoder learns compact cycle-level representations from partial charging measurements. Its encoder maps each 50 × 8 input sequence to a 64-dimensional embedding. The decoder projects this embedding to 128 dimensions and repeats it across 50 sequence positions. A two-layer GRU followed by a linear output layer reconstructs the eight input channels at each position. Reconstruction is non-autoregressive: previously reconstructed samples are not fed back into the decoder. The embeddings from 10 sampled cycles are also stacked and processed by an auxiliary GRU predictor for joint RUL prediction and capacity estimation. Table I summarizes the architecture. During Stage I, the encoder, decoder, and auxiliary predictor are optimized jointly using reconstruction and prediction losses. This supervision encourages the embeddings to retain input information relevant to battery degradation. After pretraining, only the encoder is transferred to the RUL Expert and frozen; the decoder and auxiliary predictor are discarded.

4

Cross-Expert Framework: Three-Stage Training Supervised GRU-Autoencoder Pretraining

STAGE I

Learn cycle embeddings using reconstruction and RUL / capacity supervision.

Long-term input

Encoder Cycle embeddings

Decoder Input reconstruction

Reconstruction loss MSE

Auxiliary Predictor Prediction head

Supervised losses RUL MSE + capacity MSE

Independent Expert Pretraining

STAGE II

Transfer the pretrained encoder to the RUL Expert; train the two experts independently.

RUL Expert

Capacity Expert

Long-term degradation history

Short-term degradation history

Long-Term Data Processing

Short-Term Data Processing

30-cycle window; 10 selected cycles; stride 3

10-cycle window; 10 consecutive cycles; stride 1

10-min segment from V > 3.1 V; measured I, V, T, t

40-min segment from V > 3.1 V; measured I, V, T, t

Q, V, T, t + difference features; 10 x 50 x 8

Q, V, T, t; 10 x 50 x 4

Long-Term Feature Extraction

Short-Term Feature Extraction 2D-CNN + Transformer

Pretrained Encoder

Supervised losses RUL MSE + capacity MSE

2D-CNN + GRU

Supervised losses RUL MSE + capacity MSE

STAGE III

Cross-Expert Knowledge Fusion Training

Freeze both pretrained experts; train FiLM fusion with supervision from both tasks.

Pretrained RUL Expert

Pretrained Capacity Expert

Cross-Expert Knowledge Fusion (FiLM)

Capacity estimation

RUL prediction

FROZEN TRAINABLE

Supervised losses RUL MSE + 0.5.capacity MSE

Fig. 1. Overview of the proposed cross-expert framework and its three-stage training procedure. Stage I learns cycle-level embeddings through supervised GRU-autoencoder pretraining. Stage II independently trains the RUL and Capacity Experts using long- and short-term degradation histories, respectively. Stage III freezes both experts and optimizes the FiLM-based fusion module and shared regression head for joint RUL prediction and capacity estimation. Tensor dimensions in the diagram omit the batch axis.

3) Long-Term Degradation Feature Extraction: The pretrained encoder independently processes the 10 cycle sequences in HL w . Their embeddings are stacked into a 10 × 64 degradation map, which is processed by two 2D convolutional layers. Mean pooling over the embedding-position axis produces a sequence of 10 cycle-level feature vectors, each with 128 dimensions.

A two-layer temporal GRU models dependencies across the 128 sampled cycles. Its final top-layer hidden state, hL , w ∈ R is supplied to the two-output regression head during Stage II and to the FiLM fusion module during Stage III. Table II summarizes the architecture.

B. Capacity Expert The Capacity Expert characterizes recent battery behavior using longer partial charging segments from 10 consecutive cycles. Channel-wise statistical descriptors summarize each segment, and a 2D-CNN followed by a Transformer encoder models the resulting cycle sequence. 1) Short-Term Data Processing: For each cycle, a nominal 40-min charging segment is extracted beginning at the first sample satisfying V > 3.1 V. Savitzky–Golay filtering is not applied in this branch. Elapsed time and accumulated charge are calculated using Eqs. (4) and (5), respectively, with the original current replacing the filtered current and the summation taken over the 40-min segment. Its locally accumulated charge is denoted by QSc,j to distinguish it from

5

TABLE I S UPERVISED GRU-AUTOENCODER P RETRAINING A RCHITECTURE

Component Operation Input Long-term window Merge batch/cycle axes Encoder GRU, 2 layers Final top-layer state Linear, 128 → 64 Decoder Linear, 64 → 128 Repeat over 50 steps GRU, 2 layers Linear, 128 → 8 Auxiliary Restore cycle axis predictor GRU, 2 layers Final top-layer state Two-output regression head

Output shape B × 10 × 50 × 8 10B × 50 × 8 10B × 50 × 128 10B × 128 10B × 64 10B × 128 10B × 50 × 128 10B × 50 × 128 10B × 50 × 8 B × 10 × 64 B × 10 × 128 B × 128 B×2

B denotes the mini-batch size. All three GRUs use 128 hidden units per layer and inter-layer dropout of 0.1. The decoder and auxiliary predictor form separate branches receiving the encoder embeddings. The decoder reconstructs each cycle independently, whereas the predictor models the sequence of 10 cycle embeddings to estimate RUL and capacity. TABLE II RUL E XPERT A RCHITECTURE

Component Input

Operation Long-term window Merge batch/cycle axes Pretrained GRU encoder, final state Encoder Linear, 128 → 64 Restore cycle axis Long-Term Add channel axis Degradation Conv2D, 1 → 32 Feature Extraction Conv2D, 32 → 128 Mean over embedding axis Transpose Temporal GRU, final state Predictor Two-output regression head

Output shape B × 10 × 50 × 8 10B × 50 × 8 10B × 128 10B × 64 B × 10 × 64 B × 1 × 10 × 64 B × 32 × 10 × 64 B × 128 × 10 × 64 B × 128 × 10 B × 10 × 128 B × 128 B×2

Both GRUs have two layers with 128 hidden units per layer and inter-layer dropout of 0.1. Both convolutions use 3 × 3 kernels, unit stride, and padding of one. Each convolution is followed by ReLU; dropout of 0.1 follows the first ReLU. The transferred encoder remains frozen during Stage II, while the expert CNN, temporal GRU, and regression head are trained. Stage III uses the 128-dimensional expert representation before the regression head.

QL c,j . Raw-signal and elapsed-time symbols in this subsection refer to the short-term segment. The resulting channels are  ZSc = QSc

Vc

Tc

 S τ c ∈ RMc ×4 ,

(12)

where McS denotes the number of samples in the extracted short-term segment. Let p ∈ {0, . . . , 9} index its chronological cycle position within the short-term window associated with sample w. Channel-wise interpolation and training-set min– max normalization produce  b S = MS I50 (ZS ) ∈ R50×4 . X w,p w,p

(13)

The normalization parameters are fitted only on the training set and are reused without modification for validation and testing.

TABLE III C APACITY E XPERT A RCHITECTURE

Component Input Short-Term Degradation Feature Extraction

Predictor

Operation Statistical descriptors Add channel axis Conv2D, 1 → 32 Conv2D, 32 → 128 Mean over feature axis Transpose to cycle sequence Linear, 128 → 128 Add sinusoidal encoding Transformer encoder, 2 layers Mean over cycle tokens Two-output regression head

Output shape B × 10 × 28 B × 1 × 10 × 28 B × 32 × 10 × 28 B × 128 × 10 × 28 B × 128 × 10 B × 10 × 128 B × 10 × 128 B × 10 × 128 B × 10 × 128 B × 128 B×2

B denotes the mini-batch size. Both convolutions use 3 × 3 kernels, unit stride, and padding of one, followed by ReLU; dropout of 0.1 follows the first ReLU. The Transformer has an embedding dimension of 128, four attention heads, a feedforward dimension of 512, and dropout of 0.1. Fixed sinusoidal positional encoding supports up to 10 cycle tokens. The pooled representation is extracted before the prediction head for Stage III.

For each channel q ∈ {1, . . . , 4}, seven descriptors are computed over the 50 normalized, resampled positions:  ⊤ dSw,p,q = µ, σ, min, max, median, σ 2 , skew ∈ R7 . (14) Here, µ, σ, and σ 2 denote the mean, standard deviation, and variance, respectively; the remaining entries denote the minimum, maximum, median, and skewness of the same channel sequence. Concatenating the descriptors of all four channels yields the cycle-level statistical vector  sSw,p = Concat dSw,p,1 , . . . , dSw,p,4 ∈ R28 . (15) The vectors from 10 consecutive cycles, sampled at stride one, are stacked to form the Capacity Expert input:  ASw = Stackcycle sSw,0 , . . . , sSw,9 ∈ R10×28 . (16) 2) Short-Term Degradation Feature Extraction: The Capacity Expert processes the statistical descriptor matrix ASw ∈ R10×28 using a 2D-CNN followed by a Transformer encoder [22]. The CNN preserves the cycle and statisticalfeature axes while expanding the channel dimension to 128. Averaging over the statistical-feature axis and transposing the result produces a sequence of 10 cycle tokens, each with 128 features. A learned linear projection maps these tokens to the Transformer embedding space, and fixed sinusoidal positional encoding represents their order within the cycle window. The resulting sequence is processed by two Transformer encoder layers. Mean pooling over the cycle axis yields the shortterm representation hSw ∈ R128 , which is passed to a twooutput regression head for joint RUL prediction and capacity estimation during Stage II. During Stage III, this representation is extracted before the regression head and supplied to the fusion module. Table III summarizes the architecture. C. Cross-Expert Knowledge Fusion The fusion module integrates the long-term representa128 tion hL and the short-term representation hSw ∈ w ∈ R

6

R128 through FiLM. The short-term representation generates feature-wise scaling and shifting vectors that modulate the long-term representation: e w = γ(hS ) ⊙ hL + β(hS ), h w w w

(17)

where ⊙ denotes element-wise multiplication, and γ, β : R128 → R128 are learned mappings. This modulation allows recent battery behavior to condition the long-term degradation representation. e w is passed to a shared regression The fused representation h head comprising a linear layer, ReLU activation, dropout, and a final linear layer with two outputs for RUL prediction and capacity estimation:    e w + b1 + b2 ∈ R2 , bw = W2 Dropout ReLU W1 h y (18) where W1 ∈ Rdh ×128 and W2 ∈ R2×dh are learned weight matrices, b1 and b2 are bias vectors with dimensions dh and 2, respectively, and dh denotes the head’s hidden dimension. The bw are the scaled RUL and capacity estimates. two entries of y Physical units are restored before evaluation as described in Section III-D. During Stage III, both pretrained experts remain frozen. The frozen experts are used in evaluation mode to disable dropout. Only the FiLM module and shared regression head are optimized, preserving the expert representations while learning their interaction for joint prediction. D. Three-Stage Training Strategy Let the RUL, capacity, and reconstruction losses be LRUL = MSE(b yRUL , yRUL ), LCap = MSE(b yCap , yCap ), b L , HL ), Lrec = MSE(H

(19)

where the prediction and target vectors collect the correspondb L denotes ing task values over a mini-batch. The tensor H the reconstructed long-term input after restoring its batch and cycle axes. Each MSE is averaged over all elements of its corresponding tensors. The prediction losses are computed on the scaled targets and corresponding outputs defined in Section III-D, using the model trained in the respective stage. 1) Stage I: Supervised GRU-Autoencoder Pretraining: The GRU autoencoder is trained using the long-term partial charging sequences. The encoder produces cycle-level embeddings, the decoder reconstructs the encoder input, and an auxiliary predictor predicts both RUL and capacity from the embedding sequence. All three components are jointly optimized using LI = LRUL + LCap + Lrec .

(20)

The resulting encoder is transferred to the RUL Expert and frozen in the subsequent stages. 2) Stage II: Independent Expert Pretraining: The RUL and Capacity Experts are trained independently. Both experts use a two-output regression head and receive supervision from the RUL and capacity targets. For expert e ∈ {L, S}, the objective is (e) (e) (e) LII = LRUL + LCap . (21)

For the RUL Expert, the transferred encoder remains frozen while the 2D-CNN, temporal GRU, and regression head are optimized. For the Capacity Expert, the 2D-CNN, Transformer, and regression head are optimized. The two experts are optimized separately, without feature exchange or shared trainable parameters during this stage. 3) Stage III: Cross-Expert Fusion Training: After Stage II, both pretrained experts are frozen. The FiLM module and shared regression head are optimized using LIII = LRUL + λCap LCap ,

(22)

where λCap controls the relative contribution of the capacity estimation loss. The reference configuration uses λCap = 1; λCap = 0.5 is evaluated separately as an RUL-prioritized alternative in the loss-weight ablation. During this stage, only the fusion module and shared regression head are updated, preserving the representations learned by both experts during pretraining. III. E XPERIMENTAL S ETUP A. Datasets Two public battery-aging datasets are considered. Dataset I, released by [20], contains 124 commercial A123 APR18650M1A lithium iron phosphate (LFP)/graphite cells, with a nominal capacity of 1.1 Ah and a nominal voltage of 3.3 V. The cells were cycled at 30 ◦ C under different fast-charging protocols, yielding cycle lives of approximately 150–2300 cycles. The dataset includes electrical and surfacetemperature measurements and provides a benchmark for degradation prediction under charging-protocol variability. Dataset II, released by Ma et al. [14], comprises 77 commercial LFP/graphite cells with a nominal capacity of 1.1 Ah. The cells were tested at 30 ◦ C using an identical charging procedure and 77 different multistage discharge protocols. The charging procedure includes 5C charging to 80% state of charge, followed by 1C charging to 3.6 V and a constantvoltage stage. The dataset contains more than 140 000 cycles, with cell lifetimes of approximately 1100–2700 cycles. It complements Dataset I by emphasizing discharge-protocol variability. Both source studies use 80% of nominal capacity as the end-of-life criterion. B. Data Partitioning and Evaluation Protocol Both datasets are partitioned by battery identity, ensuring that all windows from a given cell belong to the same partition. For Dataset I, we adopt the cell grouping reported in the supplementary information of [20], using 41 cells for training, 43 for validation, and 40 for testing. The source study’s primary test set is used here for validation, while its secondary test set serves as the test set. Training and validation cells come from the 2017-05-12 and 2017-06-30 batches; the 40 test cells come from the later 2018-04-12 batch. Thus, Dataset I evaluates unseen cells under a manufacturing-batch and charging-protocol shift. For Dataset II, we follow the predefined division of 55 development cells and 22 test cells in [14]. The 55 development cells are further divided into training and validation subsets

7

using an approximately 90:10 ratio at the cell level. The remaining 22 cells are reserved exclusively for final testing. All learned preprocessing parameters are fitted using only the training partition and remain fixed during validation and testing. Rolling windows contain only measurements available at the prediction time. Model configurations, loss weights, and checkpoints are selected using the validation partition; the test partitions are reserved for final performance evaluation. Unless otherwise indicated, predictions are averaged over five runs using fixed data partitions. This prediction-level averaging forms a five-model ensemble. Evaluation metrics are then computed separately for each test cell and summarized as the mean and standard deviation across test cells, m̄±sm . The standard deviation therefore quantifies between-cell variability in predictive performance, rather than variability across runs. C. Implementation Details The neural networks are implemented in PyTorch 2.5.1 and trained using the Adam optimizer [13] with a learning rate of 10−4 . Data preprocessing and numerical computations use NumPy 2.2.6, SciPy 1.15.3, pandas 2.3.3, and scikitlearn 1.7.2. XGBoost 3.2.0 is used for the tree-based regression experiments. All experiments are conducted on a workstation equipped with an NVIDIA GeForce RTX 5060 GPU, an AMD Ryzen 7 CPU, and 32 GB of system RAM. D. Evaluation Metrics During training, RUL and capacity targets are scaled as yCap yRUL , yeCap = , (23) yeRUL = 3000 sQ where sQ = 1.1 Ah for Dataset I and sQ = 1300 mAh for Dataset II, matching the native capacity units of each dataset. At inference, predictions are multiplied by their respective scaling factors to recover the original units before evaluation. Capacity values are converted to mAh for consistent reporting across datasets. Performance is evaluated using root-mean-square error (RMSE), coefficient of determination (R2 ), and mean absolute percentage error (MAPE): v u N u1 X RMSE = t (yi − ŷi )2 , (24) N i=1 PN (yi − ŷi )2 2 , (25) R = 1 − Pi=1 N 2 i=1 (yi − ȳ) N 100 X yi − ŷi MAPE = , (26) N i=1 yi where N is the number of evaluated samples for one test cell, yi and ŷi P are the true and predicted values, respectively, N and ȳ = N −1 i=1 yi . RMSE is reported in cycles for RUL and in mAh for capacity, whereas MAPE is reported as a percentage. Lower RMSE and MAPE and higher R2 indicate better predictive performance. The MAPE expression assumes nonzero targets; R2 requires a nonconstant target sequence. The cell-level metric values are averaged with equal weight across test cells, rather than pooling all windows across cells.

E. Baseline Methods Three groups of comparisons are considered. First, rawcurve GRU, LSTM, and Transformer baselines use nominal 40-min charging observations above 3.1 V without autoencoder embeddings or handcrafted statistics. Second, representation and architecture ablations compare XGBoost regression on statistical descriptors, raw resampled curves, and learned embeddings, together with recurrent, convolutional, and Transformer-based expert backbones. Third, fusion comparisons use the same two pretrained experts with alternative feature- or prediction-level fusion operators. The comparison additionally includes MSFEH [25], structural pruning [10], Ma et al. [14], Ge et al. [9], and Qian et al. [18]. Their numerical results are presented as reference comparisons. Differences in input availability, test partitions, and metric aggregation must be considered when interpreting rankings against external methods. IV. R ESULTS AND D ISCUSSION This section reports performance on both datasets and ablations on Dataset I. The reference FiLM configuration uses λCap = 1; the RUL-prioritized configuration uses λCap = 0.5. Unless stated otherwise, deviations quantify variation across test cells after predictions are averaged over five runs. Meanonly tables omit these deviations for compactness. A. RUL Prediction Performance Table V reports the comparison on both datasets, and Table IV separates the standalone experts and fusion configurations on Dataset I. With equal task weights, FiLM achieves an RUL RMSE of 143.69 cycles, R2 = 0.74, and MAPE of 10.91% on Dataset I. Its RMSE is 43.18%, 48.13%, and 32.36% lower than those of the raw-curve GRU, LSTM, and Transformer, respectively. These comparisons assess the complete pipeline, including its observations and representations, rather than isolating fusion. On Dataset II, the proposed method achieves an RUL RMSE of 161.10 cycles, R2 = 0.88, and MAPE of 7.15%. Relative to the raw-curve Transformer, RMSE decreases from 172.11 to 161.10 cycles (6.40%). The proposed method has the lowest reported RUL RMSE in both dataset blocks. However, MSFEH has a lower Dataset I MAPE (9.82%), so the proposed method does not lead every RUL metric. External-method rankings remain conditional on protocol compatibility. On Dataset I, reducing λCap to 0.5 lowers RMSE to 141.64 cycles, a 1.43% reduction, while increasing capacity error. This small mean difference does not establish statistical significance. Evaluation on two separately trained datasets demonstrates performance in two settings, not zero-shot transfer between datasets. B. Capacity Estimation Performance Table IV shows that equal-weight FiLM achieves a capacity RMSE of 12.36 ± 5.32 mAh, R2 = 0.90 ± 0.08, and MAPE of 0.99 ± 0.45% on Dataset I. Mean RMSE is 14.23% lower than that of the Capacity Expert and 63.91% lower than that

8

First Data normalized performance (larger = better) RUL R2

RUL MAPE

Cap RMSE

Second Data normalized performance (larger = better)

Proposed Ge2025 Qian2026 GRU LSTM Transformer

RUL MAPE

RUL RMSE

Cap R2

RUL R2

Cap RMSE

RUL RMSE

Cap R2

Cap MAPE

Proposed Ge2025 Qian2026 GRU LSTM Transformer

Cap MAPE

Fig. 2. Normalized metric profiles for (a) Dataset I and (b) Dataset II, using the six methods with complete metrics shown in the legend. Each metric is min–max normalized separately within each dataset; error metrics are reversed so that larger values indicate better performance. Normalization uses the displayed rounded values. Polygon area is not an aggregate performance metric. TABLE IV P ERFORMANCE COMPARISON ON DATASET I. VALUES ARE MEAN CELL - LEVEL METRICS ; STANDARD DEVIATIONS ARE OMITTED . T HE TWO F I LM CONFIGURATIONS USE DIFFERENT CAPACITY- LOSS WEIGHTS . RUL Prediction Method GRU LSTM Transformer RUL Expert Capacity Expert FiLM, λCap = 1 FiLM, λCap = 0.5

Capacity Estimation

RMSE (cycles)

R2

MAPE (%)

RMSE (mAh)

R2

MAPE (%)

252.90 277.03 212.44 147.08 151.04 143.69 141.64

0.30 0.13 0.44 0.72 0.70 0.74 0.74

18.48 21.46 16.93 11.44 11.21 10.91 10.91

9.39 24.13 15.13 34.25 14.41 12.36 15.50

0.93 0.64 0.85 0.27 0.86 0.90 0.84

0.59 1.85 1.11 2.98 1.10 0.99 1.26

of the RUL Expert. Together with the lower RUL error, this supports the value of combining their representations in the reference configuration. The raw-curve GRU nevertheless has a lower Dataset I capacity RMSE of 9.39 mAh and MAPE of 0.59%, at the expense of higher RUL error. With λCap = 0.5, FiLM capacity RMSE increases to 15.50 mAh, exceeding that of the standalone Capacity Expert. Improvement in both tasks relative to the experts therefore applies to equal-weight fusion, not every loss configuration. On Dataset II, the proposed method achieves a capacity RMSE of 7.28 mAh, R2 = 0.99, and MAPE of 0.50%. Several reference methods and the raw-curve GRU and LSTM have lower capacity errors. The strongest observed advantage is therefore RUL prediction; the method does not provide uniformly superior capacity estimation. The radar profiles in Fig. 2 visualize this trade-off without defining a combined ranking. C. Ablation Studies The following comparisons use Dataset I. Preliminary XGBoost experiments investigate representation choices, whereas subsequent neural-network experiments investigate expert

backbones and fusion. Their different prediction heads and preprocessing conditions preclude interpreting every comparison as a one-factor ablation of the final model. 1) Input Features and Statistical Descriptors: Figs. 3 and 4 summarize the input-feature and statistical-descriptor ablations, respectively. As shown in Fig. 3, adding time to the Q, V, T inputs reduces RUL RMSE from 250.31 to 219.81 cycles, while increasing capacity RMSE from 20.24 to 23.34 mAh. The lag-12 configuration with Savitzky–Golay filtering achieves RMSEs of 216.83 cycles and 26.03 mAh for RUL and capacity, respectively. Filtering reduces RUL error for the lag9 and lag-12 configurations but increases it for lag-6; capacity error increases for all three lag configurations. Thus, feature augmentation and filtering have task-dependent effects. These comparisons remain exploratory because several configurations were evaluated in a single run. For statistical inputs, Fig. 4 shows that extending the fourchannel descriptors from means alone to mean, standard deviation, minimum, maximum, median, variance, and skewness reduces RUL RMSE from 228.52 to 216.40 cycles. Adding kurtosis increases the error to 223.98 cycles. The sevenstatistic set produces the 28-dimensional cycle descriptor used

9

TABLE V P ERFORMANCE COMPARISON ON DATASETS I AND II. RUL RMSE IS REPORTED IN CYCLES , CAPACITY RMSE IN M A H , AND MAPE IN PERCENT. GRU, LSTM, AND T RANSFORMER USE RAW PARTIAL - CHARGING CURVES (V > 3.1 V, t:t+40 MIN ), WITHOUT AN AUTOENCODER OR HANDCRAFTED STATISTICS . A DASH DENOTES AN UNAVAILABLE RESULT. T HE PROPOSED METHOD USES λCap = 1. Dataset I Method

Dataset II

RUL Prediction

Capacity Estimation

RUL Prediction

RMSE ↓

R2 ↑

MAPE ↓

RMSE ↓

R2 ↑

MAPE ↓

RMSE ↓

R2 ↑

MAPE ↓

RMSE ↓

R2 ↑

MAPE ↓

MSFEH [25] Structural Pruning [10] Ma et al. [14] Ge2025 [9] Qian2026 [18]

161.81 274.71 290.76

0.11 0.02

9.82 21.67 23.25

13.51 32.91

0.87 0.32

1.17 2.76

192.17 186 171.82 173.57

0.858 0.804 0.86 0.84

8.72 8.15 8.01

6.18 2.57 3.13 5.13

0.995 0.999 1.00 1.00

0.176 0.21 0.34

GRU LSTM Transformer

252.90 277.03 212.44

0.30 0.13 0.44

18.48 21.46 16.93

9.39 24.13 15.13

0.93 0.64 0.85

0.59 1.85 1.11

176.98 192.88 172.11

0.86 0.83 0.83

8.41 9.31 8.30

3.07 4.62 19.73

1.00 1.00 0.95

0.21 0.33 1.64

Proposed method

143.69

0.74

10.91

12.36

0.90

0.99

161.10

0.88

7.15

7.28

0.99

0.50

Without SG

With SG

(a) RUL prediction

(b) Capacity estimation

275

237.26 219.81

225

30.04

30

250.31

223.09

219.59

223.34 214.83

26.80

25.90

235.38

25 216.83

200 175

RMSE (mAh)

20

23.34

22.61

21.39

20.24

26.03 22.34

15

w Lag it h 12 SG

ho Lag ut -1 SG 2 it w

L it ag h -6 SG w

ho L ut ag SG -6 it w

L it ag h -9 SG w

ho L ut ag SG -9 w

it

Q , + V,T SG ,t

w Lag it h 12 SG

ho Lag ut -1 SG 2 it w

L it ag h -6 SG w

L it ag h -9 SG w

ho L ut ag SG -6 it w

w

it

ho L ut ag SG -9

0 Q , + V,T SG ,t

100 Q ,V ,T ,t

5

Q ,V ,T

125

Q ,V ,T ,t

10

150

Q ,V ,T

RMSE (cycles)

250

Capacity Estimation

Fig. 3. Input-feature ablation on Dataset I: (a) RUL prediction RMSE and (b) capacity estimation RMSE. SG denotes Savitzky–Golay filtering. Bars show mean cell-level errors; some preliminary configurations use a single run rather than five-run prediction averaging. Standard deviations are omitted, and the RUL axis starts at 100 cycles. Lower values indicate better performance.

by the Capacity Expert. However, the subset comprising mean, standard deviation, minimum, and maximum achieves a lower capacity RMSE of 22.67 mAh, compared with 23.96 mAh for the selected set. Descriptor selection therefore reflects a trade-off between the two tasks, rather than simultaneous minimization of both errors. 2) Charging Segment Duration and Cycle History: Tables VI–IX report the window experiments. For partialcharging embeddings with an XGBoost head, a nominal 10min segment gives an RUL RMSE of 159.38 cycles, compared with 316.30, 286.84, and 285.96 cycles for 20-, 30-, and 40-min segments. Conversely, statistical representations favor longer segments for capacity: their capacity RMSEs are 32.41, 14.74, 10.86, and 12.77 mAh for 10, 20, 30, and 40 min, respectively. The 40-min statistical input gives the lowest RUL RMSE in that group (199.98 cycles), whereas 30 min gives the lowest capacity RMSE. The selected 40-min view should therefore not be described as universally optimal. With the 2D-CNN–GRU embedding backbone, sampling ten cycles over a 30-cycle history reduces RUL RMSE from 171.84 to 147.08 cycles relative to ten consecutive cycles. For

the statistical 2D-CNN–Transformer, ten consecutive cycles give an RUL RMSE of 151.04 cycles, versus 154.43 cycles for the extended history, with similar capacity RMSEs of 14.41 and 14.27 mAh, respectively. These observations motivate complementary history lengths, while providing limited evidence that the shorter history itself improves capacity accuracy. 3) Temporal Feature Extraction Architectures: Capacity Expert ablation: Table X evaluates statistical-input backbones with ten cycles sampled across a 30-cycle history. The 2DCNN–Transformer gives the lowest capacity RMSE in that table (14.27 mAh), compared with 22.18, 23.65, and 23.51 mAh for the 2D-CNN–LSTM, GRU, and Mamba alternatives. Its RUL RMSE is 154.43 cycles; the 2D-CNN–LSTM is slightly lower at 153.47 cycles. This table supports the Capacity Expert architecture under an extended-history ablation, rather than reporting its final consecutive-cycle setting. Table XII evaluates the same statistical family with ten consecutive cycles, matching the Capacity Expert history in Section II-B. The 2D-CNN–Transformer yields 151.04 cycles and 14.41 mAh, which are the standalone Capacity Expert values used in the fusion comparison. Thus, the extended and

10

Q,V,T

Q,V,T,t

(a) RUL prediction

(b) Capacity estimation 40 35.78

300

35

287.40

33.29

32.87

30.57 262.38

30 250.30

250

240.99 231.14

229.51

228.52

26.92 26.40

250.39 231.14

239.56

233.93 223.98

216.40

200

RMSE (mAh)

RMSE (cycles)

253.84

27.83 26.12

25

26.12 22.82

22.67

24.24

23.96

22.82

20 15 10

150

5

Cumulative statistical descriptors

si s +

K ur

to

ew +

Sk

Va r +

n ia ed M +

+ M M in ax

d St +

n ea M

si s to +

+

+

K ur

Sk

ew

Va r +

n ia ed M +

+

+ M M in ax

St +

ea M

d

0 n

100

Cumulative statistical descriptors

Fig. 4. Statistical-descriptor ablation on Dataset I using XGBoost: (a) RUL prediction RMSE and (b) capacity estimation RMSE. Descriptors are added cumulatively from left to right for the Q, V, T and Q, V, T, t inputs. Values above the bars are the reported point estimates. The RUL axis starts at 100 cycles. Lower RMSE indicates better performance. TABLE VI S TATISTICAL DESCRIPTORS : VOLTAGE - INTERVAL AND DURATION ABLATION ON DATASET I WITH XGB OOST. VOLTAGES ARE IN V; t:t+d DENOTES A NOMINAL d- MIN INTERVAL BEGINNING AT THE FIRST CHARGING SAMPLE ABOVE 3.1 V. T HESE ARE PRELIMINARY REPRESENTATION COMPARISONS ; LABELS DISTINGUISH THE RECORDED PREPROCESSING CONFIGURATIONS .

RUL prediction

Configuration RMSE (cycles) [3.4, 3.6] [3.3, 3.5] [3.2, 3.4] [3.1, 3.3] [3.3, 3.6] [3.2, 3.6] [3.1, 3.6] V > 3.1 V, (t:t+10) V > 3.1 V, (t:t+20) V > 3.1 V, (t:t+30) V > 3.1 V, (t:t+40)

216.40 229.88 ± 162.24 269.71 ± 169.62 256.43 ± 142.50 255.40 ± 158.06 226.29 ± 145.81 205.64 ± 140.99 214.55 ± 145.24 267.25 ± 156.63 220.62 ± 135.48 199.98 ± 138.06

R2

Capacity estimation MAPE (%)

0.44 16.51 0.36 ± 0.47 17.57 ± 7.64 0.03 ± 0.98 21.25 ± 10.84 0.01 ± 1.49 20.36 ± 10.68 0.23 ± 0.47 19.74 ± 6.68 0.40 ± 0.40 17.14 ± 6.12 0.51 ± 0.36 15.72 ± 5.90 0.45 ± 0.40 16.44 ± 5.69 0.14 ± 0.49 21.61 ± 7.49 0.40 ± 0.42 16.38 ± 6.50 0.51 ± 0.37 14.72 ± 6.39

consecutive histories provide similar capacity accuracy, while the consecutive history gives lower mean RUL error for this backbone. RUL Expert ablation: Table XI evaluates embedding-input backbones with the extended history. The 2D-CNN–GRU achieves the lowest RUL RMSE in that table (147.08 cycles), compared with 162.81, 173.48, and 167.28 cycles for the corresponding LSTM, Transformer, and Mamba hybrids. A standalone GRU gives 159.08 cycles. Table XIII provides the consecutive-cycle embedding comparison; the 2D-CNN–GRU error increases to 171.84 cycles. The selected RUL and Capacity backbones therefore reflect different input representations and temporal contexts. 4) Autoencoder Pretraining: Table XIV compares autoencoder architectures. In the reported 10-min embedding experiment with an XGBoost head, the GRU autoencoder achieves an RUL RMSE of 157.87 cycles. The LSTM autoencoder is slightly better at 156.31 cycles, whereas Ti-MAE, TS-MAE, H3MAE, and TimeMAE yield 195.36, 208.55, 213.76, and

RMSE (mAh)

R2

23.96 0.66 21.31 ± 9.30 0.71 ± 0.25 36.46 ± 18.28 0.08 ± 0.91 39.38 ± 18.76 -0.06 ± 0.90 26.29 ± 9.91 0.57 ± 0.32 26.26 ± 7.66 0.57 ± 0.25 24.69 ± 6.76 0.63 ± 0.20 32.41 ± 6.48 0.37 ± 0.32 14.74 ± 4.26 0.87 ± 0.09 10.86 ± 3.45 0.93 ± 0.05 12.77 ± 3.83 0.90 ± 0.07

MAPE (%) 1.70 1.58 ± 0.72 2.83 ± 1.61 3.18 ± 1.74 1.90 ± 0.82 1.90 ± 0.64 1.71 ± 0.56 2.45 ± 0.66 1.18 ± 0.35 0.81 ± 0.26 0.93 ± 0.35

230.42 cycles, respectively. GRU and LSTM therefore perform similarly for RUL, and the GRU should not be described as the unique best encoder. Capacity results also differ: H3MAE reaches 27.87 mAh, compared with 32.64 mAh for GRU. The raw-curve 10-min configuration without an autoencoder yields an RUL RMSE of 237.17 cycles. Although the embedding result is lower, the source records associate the selected embedding experiment with full charge–discharge data. Until preprocessing is matched, this comparison cannot isolate the benefit of pretraining. A definitive pretraining ablation requires the same downstream architecture and observations, comparing supervised pretraining, reconstruction-only pretraining, and random initialization under the same training budget. 5) Individual Experts and Fusion Strategies: The RUL Expert gives an RUL RMSE of 147.08 cycles but a capacity RMSE of 34.25 mAh. The Capacity Expert gives 151.04 cycles and 14.41 mAh, respectively. Equal-weight FiLM improves both mean errors to 143.69 cycles and 12.36 mAh, consistent with complementary expert information.

11

TABLE VII GRU- AUTOENCODER EMBEDDINGS FROM PARTIAL - CHARGE DATA : VOLTAGE - INTERVAL AND DURATION ABLATION ON DATASET I WITH XGB OOST. VOLTAGES ARE IN V; t:t+d DENOTES A NOMINAL d- MIN INTERVAL BEGINNING AT THE FIRST CHARGING SAMPLE ABOVE 3.1 V. T HESE ARE PRELIMINARY REPRESENTATION COMPARISONS ; LABELS DISTINGUISH THE RECORDED PREPROCESSING CONFIGURATIONS .

RUL prediction

Configuration

MAPE (%)

RMSE (mAh)

R2

MAPE (%)

216.83 ± 141.44 0.43 ± 0.38 16.96 ± 6.36 175.66 ± 144.57 0.61 ± 0.38 12.84 ± 6.93 206.67 ± 123.98 0.37 ± 0.66 16.46 ± 7.95 193.84 ± 133.89 0.28 ± 1.42 15.85 ± 11.79 232.85 ± 135.60 0.33 ± 0.40 18.28 ± 6.21 226.82 ± 139.97 0.38 ± 0.39 17.64 ± 6.00 228.23 ± 145.49 0.37 ± 0.41 17.67 ± 6.28 159.38 ± 126.86 0.68 ± 0.29 12.21 ± 6.35 316.30 ± 155.38 -0.17 ± 0.48 25.93 ± 5.75 286.84 ± 145.60 0.03 ± 0.42 23.36 ± 5.68 285.96 ± 144.86 0.05 ± 0.40 23.29 ± 5.49

26.03 ± 8.23 22.15 ± 10.24 31.84 ± 13.22 33.74 ± 17.58 27.94 ± 12.77 28.10 ± 13.75 28.69 ± 12.13 36.66 ± 7.78 21.13 ± 8.75 19.59 ± 6.12 23.77 ± 7.85

0.57 ± 0.29 0.68 ± 0.31 0.34 ± 0.48 0.20 ± 0.75 0.48 ± 0.44 0.47 ± 0.49 0.47 ± 0.39 0.19 ± 0.41 0.72 ± 0.28 0.77 ± 0.16 0.65 ± 0.26

1.93 ± 0.75 1.63 ± 0.90 2.37 ± 0.96 2.59 ± 1.49 2.05 ± 1.03 2.10 ± 1.14 2.14 ± 0.96 3.07 ± 0.77 1.67 ± 0.82 1.50 ± 0.59 1.90 ± 0.72

RMSE (cycles) [3.4, 3.6] [3.3, 3.5] [3.2, 3.4] [3.1, 3.3] [3.3, 3.6] [3.2, 3.6] [3.1, 3.6] V > 3.1 V, (t:t+10) V > 3.1 V, (t:t+20) V > 3.1 V, (t:t+30) V > 3.1 V, (t:t+40)

Capacity estimation

R2

TABLE VIII GRU- AUTOENCODER EMBEDDINGS FROM THE CONFIGURATION LABELED “ PARTIAL CHARGE ”: VOLTAGE - INTERVAL AND DURATION ABLATION ON DATASET I WITH XGB OOST. VOLTAGES ARE IN V; t:t+d DENOTES A NOMINAL d- MIN INTERVAL BEGINNING AT THE FIRST CHARGING SAMPLE ABOVE 3.1 V. T HESE ARE PRELIMINARY REPRESENTATION COMPARISONS ; LABELS DISTINGUISH THE RECORDED PREPROCESSING CONFIGURATIONS .

RUL prediction

Configuration RMSE (cycles) [3.4, 3.6] [3.3, 3.5] [3.2, 3.4] [3.1, 3.3] [3.3, 3.6] [3.2, 3.6] [3.1, 3.6] V > 3.1 V, (t:t+10) V > 3.1 V, (t:t+20) V > 3.1 V, (t:t+30) V > 3.1 V, (t:t+40)

Capacity estimation

R2

182.75 ± 135.86 0.56 ± 0.37 243.08 ± 157.18 0.30 ± 0.45 551.83 ± 181.54 -2.60 ± 0.34 357.70 ± 171.81 -0.49 ± 0.58 209.27 ± 141.63 0.47 ± 0.37 236.89 ± 149.26 0.33 ± 0.43 253.05 ± 147.59 0.24 ± 0.42 157.87 ± 130.12 0.70 ± 0.28 334.89 ± 157.53 -0.30 ± 0.47 321.28 ± 156.41 -0.20 ± 0.46 331.90 ± 165.46 -0.26 ± 0.51

MAPE (%)

RMSE (mAh)

R2

15.25 ± 7.86 18.68 ± 5.13 0.79 ± 0.12 20.15 ± 7.02 30.88 ± 7.17 0.43 ± 0.26 47.12 ± 2.54 136.07 ± 19.51 -9.73 ± 3.30 29.70 ± 5.95 39.46 ± 16.38 -0.04 ± 0.81 17.48 ± 6.77 22.31 ± 8.90 0.68 ± 0.28 20.00 ± 6.80 22.89 ± 8.69 0.67 ± 0.27 21.38 ± 6.44 23.08 ± 8.31 0.67 ± 0.21 11.91 ± 5.62 32.64 ± 7.76 0.36 ± 0.36 27.74 ± 5.62 20.24 ± 7.48 0.75 ± 0.21 26.29 ± 5.82 26.12 ± 8.18 0.59 ± 0.29 26.64 ± 5.43 33.11 ± 8.79 0.34 ± 0.36

Figure 5 compares predictions from the standalone experts and the fused model on Dataset I. FiLM reduces the mean celllevel RUL RMSE from 147.08 to 143.69 cycles and capacity RMSE from 14.41 to 12.36 mAh. However, substantial underestimation remains at high RUL values, and improvements are not uniform across samples. The figure therefore complements the aggregate metrics by illustrating the remaining prediction errors. Table XV compares the fusion operators. Attention pooling achieves the lowest mean RUL RMSE (143.55 cycles), followed closely by concatenation (143.57 cycles) and FiLM (143.69 cycles). These differences are small relative to the between-cell variability; paired cell-level comparisons are needed to assess their statistical significance. FiLM achieves the lowest mean capacity RMSE (12.36 mAh), compared with 15.12 mAh for concatenation and 18.67 mAh for attention

MAPE (%) 1.44 ± 0.44 2.10 ± 0.56 12.18 ± 1.77 3.52 ± 1.54 1.66 ± 0.88 1.73 ± 0.83 1.81 ± 0.80 2.73 ± 0.74 1.61 ± 0.69 2.25 ± 0.67 2.89 ± 0.76

pooling. It also yields the lowest mean RUL MAPE among the evaluated fusion operators (10.91%). These results support FiLM as a favorable trade-off between RUL prediction and capacity estimation, although it does not minimize RUL RMSE. Cross-attention, bilinear fusion, and the Hadamard product give capacity RMSEs of 30.37, 32.06, and 44.37 mAh, respectively. The prediction-level mixture of experts gives 35.56 mAh. More elaborate interaction mechanisms do not improve the reported results under these training conditions. The graph-attention experiments additionally change preprocessing and therefore cannot isolate the effect of the fusion operator alone. 6) Capacity Loss Weight in Fusion Training: Table XVI reports the loss-weight comparison. Reducing the capacity coefficient from 1 to 0.5 lowers mean RUL RMSE by

12

TABLE IX R AW INTERPOLATED FEATURES WITHOUT AN AUTOENCODER : VOLTAGE - INTERVAL AND DURATION ABLATION ON DATASET I WITH XGB OOST. VOLTAGES ARE IN V; t:t+d DENOTES A NOMINAL d- MIN INTERVAL BEGINNING AT THE FIRST CHARGING SAMPLE ABOVE 3.1 V. T HESE ARE PRELIMINARY REPRESENTATION COMPARISONS ; LABELS DISTINGUISH THE RECORDED PREPROCESSING CONFIGURATIONS .

RUL prediction

Configuration RMSE (cycles) [3.4, 3.6] [3.3, 3.5] [3.2, 3.4] [3.1, 3.3] [3.3, 3.6] [3.2, 3.6] [3.1, 3.6] V > 3.1 V, (t:t+10) V > 3.1 V, (t:t+20) V > 3.1 V, (t:t+30) V > 3.1 V, (t:t+40)

R2

Capacity estimation MAPE (%)

RMSE (mAh)

241.94 ± 147.19 0.32 ± 0.38 18.63 ± 5.47 28.79 ± 7.76 162.27 ± 149.71 0.69 ± 0.35 11.29 ± 6.73 21.32 ± 9.54 287.75 ± 182.36 -0.77 ± 3.93 24.14 ± 20.29 35.15 ± 13.56 256.54 ± 162.09 -0.42 ± 3.13 21.85 ± 18.32 39.04 ± 17.09 261.79 ± 134.44 0.20 ± 0.36 20.19 ± 4.75 29.14 ± 9.32 240.85 ± 127.05 0.31 ± 0.34 18.26 ± 4.94 31.12 ± 9.69 245.70 ± 133.04 0.29 ± 0.35 18.76 ± 4.73 32.55 ± 10.02 237.17 ± 153.88 0.35 ± 0.44 17.92 ± 5.89 34.49 ± 6.70 306.44 ± 157.22 -0.07 ± 0.47 24.48 ± 5.84 17.37 ± 4.69 316.40 ± 165.56 -0.15 ± 0.54 24.34 ± 6.61 27.11 ± 6.76 342.38 ± 162.34 -0.38 ± 0.62 27.28 ± 6.75 41.84 ± 5.63

R2

MAPE (%)

0.49 ± 0.31 0.71 ± 0.23 0.21 ± 0.54 -0.01 ± 0.78 0.47 ± 0.34 0.39 ± 0.39 0.33 ± 0.42 0.29 ± 0.34 0.81 ± 0.10 0.56 ± 0.24 0.01 ± 0.25

2.10 ± 0.62 1.47 ± 0.69 2.61 ± 1.13 3.08 ± 1.53 2.11 ± 0.84 2.30 ± 0.89 2.45 ± 0.88 2.57 ± 0.71 1.44 ± 0.45 2.36 ± 0.67 3.51 ± 0.53

TABLE X C APACITY E XPERT BACKBONE ABLATION WITH STATISTICAL DESCRIPTORS FROM NOMINAL 40- MIN SEGMENTS : TEN CYCLES SAMPLED AT STRIDE THREE WITHIN A 30- CYCLE WINDOW. DATASET I RESULTS ARE MEAN ± STANDARD DEVIATION ACROSS TEST CELLS AFTER AVERAGING PREDICTIONS OVER FIVE RUNS .

RUL prediction

Configuration RMSE (cycles) 1D CNN 2D CNN LSTM GRU Transformer Mamba 1D CNN +LSTM 1D CNN +GRU 1D CNN +Transformer 1D CNN +Mamba 2D CNN + LSTM 2D CNN + GRU 2D CNN + Transformer 2D CNN + Mamba

R2

229.30 ± 141.97 0.36 ± 0.40 168.62 ± 117.76 0.60 ± 0.38 254.76 ± 158.96 0.19 ± 0.54 232.09 ± 142.38 0.32 ± 0.47 264.13 ± 153.24 0.16 ± 0.48 292.75 ± 160.49 -0.02 ± 0.52 256.46 ± 157.64 0.20 ± 0.49 256.83 ± 153.31 0.20 ± 0.48 269.99 ± 156.72 0.12 ± 0.51 259.35 ± 151.85 0.19 ± 0.46 153.47 ± 113.44 0.70 ± 0.27 155.17 ± 113.80 0.69 ± 0.26 154.43 ± 115.44 0.66 ± 0.40 168.57 ± 122.63 0.64 ± 0.30

2.05 cycles, but increases capacity RMSE by 3.14 mAh (25.40%). Increasing the coefficient to 2 gives 144.63 cycles and 12.85 mAh and does not improve on equal weighting for either mean RMSE. The (1.5, 2) configuration also changes the overall loss scale, so it is not strictly equivalent to varying only λCap while holding the optimization dynamics fixed. The λCap = 0.5 configuration prioritizes mean RUL RMSE, whereas λCap = 1 provides the better observed capacity result. The equal-weight configuration is retained as the reference for joint-task comparison; the 0.5 configuration is reported separately. Loss weights act on the normalized targets defined in Section III-D. Removing one task loss produces very poor performance on the unsupervised output; this is expected because that output is no longer trained toward its target and

Capacity estimation MAPE (%)

RMSE (mAh)

R2

18.53 ± 6.70 27.78 ± 7.78 0.54 ± 0.30 12.96 ± 6.92 21.46 ± 10.40 0.68 ± 0.30 21.19 ± 8.44 42.50 ± 10.42 -0.06 ± 0.56 19.55 ± 8.19 50.73 ± 14.20 -0.55 ± 0.87 21.48 ± 6.97 25.72 ± 6.75 0.61 ± 0.20 24.38 ± 6.96 33.21 ± 7.62 0.35 ± 0.31 20.87 ± 7.57 32.70 ± 8.16 0.38 ± 0.33 20.94 ± 7.35 33.45 ± 7.17 0.35 ± 0.28 22.24 ± 7.43 35.54 ± 10.91 0.25 ± 0.50 21.02 ± 7.01 27.77 ± 8.30 0.54 ± 0.29 11.52 ± 5.67 22.18 ± 6.28 0.70 ± 0.17 11.57 ± 5.46 23.65 ± 8.03 0.65 ± 0.24 12.07 ± 7.03 14.27 ± 5.35 0.87 ± 0.11 12.87 ± 6.13 23.51 ± 8.32 0.65 ± 0.25

MAPE (%) 1.88 ± 0.73 1.81 ± 0.99 3.26 ± 0.96 3.73 ± 1.19 1.54 ± 0.66 2.32 ± 0.66 2.13 ± 0.80 2.37 ± 0.70 2.26 ± 1.07 2.09 ± 0.73 1.48 ± 0.68 1.70 ± 0.79 1.09 ± 0.51 1.62 ± 0.79

does not by itself prove beneficial cross-task transfer. D. Computational Efficiency The architecture uses ten tokens per expert and 128dimensional expert representations. During fusion training, freezing the experts eliminates their parameter updates; their forward computations are still required unless features are cached. At inference, the Stage-I reconstruction decoder and auxiliary prediction backbone are discarded, and both expert representations are processed by FiLM and a shared regression head. No measured parameter-count, memory, or timing comparison is available. Consequently, the experiments do not

13

TABLE XI RUL E XPERT BACKBONE ABLATION WITH CYCLE EMBEDDINGS FROM THE NOMINAL 10- MIN CONFIGURATION : TEN CYCLES SAMPLED AT STRIDE THREE WITHIN A 30- CYCLE WINDOW. DATASET I RESULTS ARE MEAN ± STANDARD DEVIATION ACROSS TEST CELLS AFTER AVERAGING PREDICTIONS OVER FIVE RUNS .

Configuration 1D CNN 2D CNN LSTM GRU Transformer Mamba 1D CNN +LSTM 1D CNN +GRU 1D CNN +Transformer 1D CNN +Mamba 2D CNN + LSTM 2D CNN + GRU 2D CNN + Transformer 2D CNN + Mamba

RUL prediction

Capacity estimation

RMSE (cycles)

R2

MAPE (%)

RMSE (mAh)

R2

MAPE (%)

164.82 ± 129.34 165.92 ± 134.24 164.22 ± 133.49 159.08 ± 135.97 159.03 ± 124.58 161.62 ± 124.14 173.12 ± 128.26 165.31 ± 128.12 164.45 ± 127.81 163.61 ± 121.35 162.81 ± 131.40 147.08 ± 121.88 173.48 ± 125.65 167.28 ± 132.85

0.61 ± 0.39 0.63 ± 0.36 0.59 ± 0.50 0.63 ± 0.43 0.65 ± 0.33 0.67 ± 0.27 0.61 ± 0.34 0.62 ± 0.37 0.64 ± 0.33 0.66 ± 0.28 0.64 ± 0.35 0.72 ± 0.27 0.60 ± 0.35 0.65 ± 0.31

13.44 ± 8.27 12.93 ± 7.72 13.46 ± 9.53 12.80 ± 8.68 12.82 ± 7.57 12.73 ± 6.18 14.07 ± 7.15 13.18 ± 7.85 13.07 ± 7.27 12.86 ± 5.93 13.21 ± 7.45 11.44 ± 6.20 14.19 ± 7.39 13.11 ± 6.42

23.78 ± 8.69 30.94 ± 10.28 32.96 ± 9.00 29.04 ± 7.69 49.42 ± 8.47 30.53 ± 8.53 32.61 ± 8.22 33.60 ± 8.49 44.91 ± 8.21 31.22 ± 8.27 30.73 ± 9.67 34.25 ± 9.19 43.42 ± 9.29 32.68 ± 8.32

0.64 ± 0.28 0.40 ± 0.43 0.33 ± 0.42 0.48 ± 0.30 -0.44 ± 0.60 0.43 ± 0.37 0.35 ± 0.39 0.31 ± 0.39 -0.20 ± 0.55 0.40 ± 0.36 0.41 ± 0.39 0.27 ± 0.45 -0.12 ± 0.55 0.34 ± 0.39

1.73 ± 0.89 2.69 ± 1.00 2.88 ± 0.87 2.39 ± 0.76 4.49 ± 0.79 2.50 ± 0.87 2.77 ± 0.84 2.90 ± 0.81 4.02 ± 0.79 2.64 ± 0.87 2.53 ± 0.87 2.98 ± 0.85 3.84 ± 0.83 2.81 ± 0.80

TABLE XII C APACITY E XPERT BACKBONE ABLATION WITH STATISTICAL DESCRIPTORS FROM NOMINAL 40- MIN SEGMENTS : TEN CONSECUTIVE CYCLES , MATCHING THE FINAL C APACITY E XPERT HISTORY. DATASET I RESULTS ARE MEAN ± STANDARD DEVIATION ACROSS TEST CELLS AFTER AVERAGING PREDICTIONS OVER FIVE RUNS .

Configuration 1D CNN 2D CNN LSTM GRU Transformer Mamba 1D CNN +LSTM 1D CNN +GRU 1D CNN +Transformer 1D CNN +Mamba 2D CNN + LSTM 2D CNN + GRU 2D CNN + Transformer 2D CNN + Mamba

RUL prediction

Capacity estimation

RMSE (cycles)

R2

MAPE (%)

RMSE (mAh)

R2

MAPE (%)

252.00 ± 145.98 174.16 ± 114.40 282.48 ± 165.72 224.75 ± 137.11 280.56 ± 153.59 282.61 ± 154.53 252.63 ± 153.53 270.67 ± 151.88 289.18 ± 158.28 270.74 ± 151.92 152.26 ± 110.38 154.69 ± 113.42 151.04 ± 109.64 153.35 ± 108.58

0.28 ± 0.41 0.60 ± 0.33 0.07 ± 0.56 0.39 ± 0.42 0.10 ± 0.47 0.10 ± 0.45 0.25 ± 0.47 0.15 ± 0.47 0.03 ± 0.52 0.16 ± 0.45 0.70 ± 0.29 0.71 ± 0.26 0.70 ± 0.30 0.71 ± 0.24

19.64 ± 6.08 13.10 ± 6.39 22.77 ± 7.85 18.29 ± 7.66 22.48 ± 6.47 22.60 ± 6.18 20.24 ± 7.44 21.91 ± 6.98 23.67 ± 7.21 21.57 ± 6.60 11.33 ± 5.97 11.07 ± 5.42 11.21 ± 5.94 11.05 ± 5.17

24.63 ± 6.32 21.61 ± 10.81 43.27 ± 7.94 50.92 ± 13.13 36.94 ± 8.09 28.13 ± 6.48 28.55 ± 7.82 33.99 ± 6.95 33.87 ± 9.44 31.94 ± 6.28 18.29 ± 5.14 17.41 ± 6.70 14.41 ± 6.28 20.68 ± 7.78

0.63 ± 0.19 0.67 ± 0.32 -0.08 ± 0.40 -0.55 ± 0.80 0.22 ± 0.31 0.54 ± 0.22 0.52 ± 0.27 0.33 ± 0.28 0.33 ± 0.39 0.41 ± 0.23 0.80 ± 0.13 0.80 ± 0.16 0.86 ± 0.13 0.72 ± 0.22

2.06 ± 0.56 1.80 ± 1.02 3.49 ± 0.77 3.63 ± 1.13 2.27 ± 0.81 1.83 ± 0.47 1.79 ± 0.78 2.39 ± 0.71 1.97 ± 0.96 2.28 ± 0.63 1.22 ± 0.55 1.26 ± 0.65 1.10 ± 0.62 1.41 ± 0.73

establish a computational advantage or real-time execution capability. In addition, five-run prediction averaging requires evaluating five trained instances unless an alternative deployment scheme is adopted. The nominal charging duration describes measurement availability and must be distinguished from computational inference latency. E. Limitations and Practical Implications The experiments cover two laboratory datasets of nominally similar chemistry at controlled ambient temperature. Generalization to other chemistries, variable temperatures, and field

operating conditions remains untested. RUL also depends on future usage, so forecasts learned under the recorded protocols may change under a different operating policy. Substantial between-cell variation remains despite prediction averaging. Cell-level error distributions, age-dependent evaluation, and calibrated prediction intervals would provide a fuller assessment than mean metrics. External comparison protocols and preliminary preprocessing configurations also limit conclusions about universal superiority. Segment-local charge integration removes the need for a full-cycle capacity measurement as an input, but does not

14

TABLE XIII RUL E XPERT HISTORY- CONTROL ABLATION WITH CYCLE EMBEDDINGS FROM THE NOMINAL 10- MIN CONFIGURATION : TEN CONSECUTIVE CYCLES . DATASET I RESULTS ARE MEAN ± STANDARD DEVIATION ACROSS TEST CELLS AFTER AVERAGING PREDICTIONS OVER FIVE RUNS .

RUL prediction

Configuration 1D CNN 2D CNN LSTM GRU Transformer Mamba 1D CNN +LSTM 1D CNN +GRU 1D CNN +Transformer 1D CNN +Mamba 2D CNN + LSTM 2D CNN + GRU 2D CNN + Transformer 2D CNN + Mamba

Capacity estimation

RMSE (cycles)

R2

MAPE (%)

RMSE (mAh)

R2

MAPE (%)

175.55 ± 128.72 177.55 ± 137.34 165.74 ± 126.14 169.88 ± 130.22 170.05 ± 128.06 169.05 ± 124.18 157.60 ± 133.26 162.19 ± 123.59 171.53 ± 128.20 167.47 ± 126.38 169.52 ± 137.17 171.84 ± 131.04 179.29 ± 123.41 168.27 ± 130.70

0.62 ± 0.31 0.61 ± 0.36 0.64 ± 0.33 0.64 ± 0.31 0.65 ± 0.29 0.65 ± 0.29 0.68 ± 0.33 0.66 ± 0.31 0.63 ± 0.32 0.65 ± 0.30 0.59 ± 0.46 0.65 ± 0.29 0.61 ± 0.29 0.66 ± 0.30

13.66 ± 6.57 14.03 ± 7.36 13.06 ± 7.36 13.09 ± 6.71 13.05 ± 5.95 13.03 ± 6.50 12.33 ± 7.16 12.58 ± 7.14 13.36 ± 6.81 12.80 ± 6.73 13.74 ± 9.17 13.32 ± 5.69 14.23 ± 6.36 13.01 ± 6.02

25.53 ± 7.93 39.93 ± 9.75 31.63 ± 8.70 33.58 ± 8.92 45.58 ± 8.54 29.79 ± 8.46 36.73 ± 8.23 33.79 ± 7.47 50.46 ± 8.38 34.90 ± 8.52 31.03 ± 8.41 31.17 ± 8.39 46.07 ± 8.40 27.62 ± 7.71

0.60 ± 0.25 0.03 ± 0.53 0.38 ± 0.38 0.31 ± 0.41 -0.23 ± 0.56 0.45 ± 0.35 0.19 ± 0.42 0.31 ± 0.37 -0.50 ± 0.64 0.26 ± 0.41 0.40 ± 0.37 0.40 ± 0.36 -0.26 ± 0.57 0.52 ± 0.31

1.91 ± 0.81 3.36 ± 1.03 2.70 ± 0.83 2.88 ± 0.86 4.09 ± 0.81 2.48 ± 0.84 3.10 ± 0.81 2.91 ± 0.72 4.57 ± 0.81 3.05 ± 0.84 2.49 ± 0.75 2.56 ± 0.79 4.13 ± 0.79 2.24 ± 0.76

TABLE XIV C OMPARISON OF AUTOENCODER ARCHITECTURES ON DATASET I. A LL ENCODERS ARE TRAINED WITH RECONSTRUCTION , RUL, AND CAPACITY SUPERVISION , FOLLOWED BY XGB OOST REGRESSION . R ESULTS ARE REPORTED AS THE MEAN ± STANDARD DEVIATION ACROSS TEST CELLS , CALCULATED BY AVERAGING PREDICTIONS OVER FIVE RUNS .

Configuration GRU LSTM Ti-MAE TS-MAE H3MAE TimeMAE

RUL prediction RMSE (cycles)

R2

157.87 ± 130.12 156.31 ± 120.96 195.36 ± 144.24 208.55 ± 147.32 213.76 ± 149.45 230.42 ± 145.57

0.70 ± 0.28 0.69 ± 0.27 0.55 ± 0.36 0.47 ± 0.41 0.45 ± 0.42 0.36 ± 0.46

Capacity estimation MAPE (%)

RMSE (mAh)

11.91 ± 5.62 32.64 ± 7.76 11.88 ± 5.96 34.46 ± 7.19 14.91 ± 6.09 29.02 ± 10.48 15.87 ± 6.62 31.27 ± 10.18 16.59 ± 6.74 27.87 ± 8.53 17.56 ± 6.68 31.85 ± 10.61

eliminate data-availability constraints. Prediction requires the stated cycle history, current and temperature measurements, and sufficiently long charging observations. Missing or interrupted segments, cells whose observed voltage already exceeds the threshold, and sensor errors require explicit handling before deployment. The availability of a nominal 40min charging segment is an observation requirement, not a guarantee for every operating cycle. V. C ONCLUSION This paper presents a three-stage framework for joint RUL prediction and capacity estimation from partial-charging observations. A GRU-autoencoder-based long-term view and a statistical short-term view are processed by specialized temporal backbones and combined through FiLM, with the shortterm representation conditioning the long-term representation. The reference configuration achieves mean RUL RMSEs of 143.69 and 161.10 cycles and capacity RMSEs of 12.36 and

R2

MAPE (%)

0.36 ± 0.36 0.30 ± 0.33 0.46 ± 0.40 0.39 ± 0.40 0.52 ± 0.31 0.36 ± 0.41

2.73 ± 0.74 2.77 ± 0.70 2.28 ± 0.95 2.33 ± 0.97 2.01 ± 0.80 2.32 ± 0.98

7.28 mAh on Datasets I and II, respectively. On Dataset I, equal-weight fusion improves both mean errors relative to the standalone experts. A capacity-loss weight of 0.5 further reduces RUL RMSE to 141.64 cycles but increases capacity RMSE to 15.50 mAh. The results support complementary feature fusion while showing that the proposed method does not minimize capacity error across all baselines. Future work will examine broader operating conditions, incomplete charging observations, uncertainty calibration, and measured deployment costs.

DATA AND C ODE AVAILABILITY The datasets are described in [14], [20]. Code inquiries may be directed to AIWARE Limited Company at https://aiware. website.

15

(a) RUL before vs after FiLM 1750

(b) Capacity before vs after FiLM

Before (Task 3.26) After (FiLM) Ideal (y = x)

Before (Task 3.51) After (FiLM) Ideal (y = x)

1050

Predicted Capacity (mAh)

Predicted RUL (cycles)

1500 1250 1000 750 500

1000

950

900

250 250

500

750

1000

1250

Actual RUL (cycles)

1500

1750

900

950

1000

Actual Capacity (mAh)

1050

Fig. 5. Predicted versus actual values on Dataset I before and after FiLM fusion: (a) standalone RUL Expert versus FiLM fusion; (b) standalone Capacity Expert versus FiLM fusion. The fusion configuration uses λCap = 1. All models are evaluated on the same test windows. The dashed line denotes ideal agreement, y = x. TABLE XV F USION - OPERATOR COMPARISON ON DATASET I USING THE SAME TWO FROZEN EXPERTS AND THEIR ORIGINAL PREPROCESSING . T HE F I LM REFERENCE USES EQUAL TASK - LOSS WEIGHTS . VALUES REPRODUCE THE REPORTED MEAN ± STANDARD DEVIATION ACROSS TEST CELLS AFTER AVERAGING PREDICTIONS OVER FIVE RUNS .

RUL prediction

Configuration

RMSE (cycles) Mixture of Experts (gate) 152.02 ± 113.19 Element-wise Addition 144.01 ± 106.53 Feature Concatenation 143.57 ± 103.43 Attention Pooling 143.55 ± 103.48 Cross-Attention 169.58 ± 129.38 Gated Multimodal Unit (GMU) 144.59 ± 104.30 Bilinear 153.80 ± 115.65 Max Pooling 143.77 ± 104.43 Hadamard Product 145.23 ± 107.23 FiLM (feature modulation) 143.69 ± 107.15

R2 0.71 ± 0.26 0.74 ± 0.23 0.73 ± 0.24 0.74 ± 0.23 0.64 ± 0.30 0.73 ± 0.24 0.70 ± 0.25 0.73 ± 0.24 0.74 ± 0.23 0.74 ± 0.23

R EFERENCES [1] Zaid Allal, Hassan N Noura, Flavien Vernier, Ola Salman, and Khaled Chahine. Machine learning for fuel cell remaining useful life prediction: A review. Energy Conversion and Management: X, page 101597, 2026. [2] Zhengyi Bao, Tingting Luo, Mingyu Gao, Zhiwei He, Kejie Gao, and Jiahao Nie. A lightweight and term-arbitrary memory network for remaining useful life prediction of li-ion battery. IEEE Transactions on Instrumentation and Measurement, 74:1–13, 2025. [3] Chun Chang, Shaojin Wang, Chen Tao, Jiuchun Jiang, Yan Jiang, and Lujun Wang. An improvement of equivalent circuit model for state of health estimation of lithium-ion batteries based on mid-frequency and low-frequency electrochemical impedance spectroscopy. Measurement, 202:111795, 2022. [4] Weidong Chen, Jun Liang, Zhaohua Yang, and Gen Li. A review of lithium-ion battery for electric vehicle applications and beyond. Energy Procedia, 158:4363–4368, 2019.

Capacity estimation MAPE (%)

RMSE (mAh)

11.92 ± 5.85 35.56 ± 9.87 11.06 ± 5.26 14.94 ± 6.80 11.32 ± 5.40 15.12 ± 7.50 11.31 ± 5.30 18.67 ± 8.34 13.10 ± 6.15 30.37 ± 9.95 11.37 ± 5.34 20.02 ± 8.47 12.08 ± 5.73 32.06 ± 9.70 11.28 ± 5.42 18.40 ± 8.23 11.34 ± 5.15 44.37 ± 12.22 10.91 ± 5.17 12.36 ± 5.32

R2

MAPE (%)

0.22 ± 0.43 0.84 ± 0.13 0.84 ± 0.14 0.76 ± 0.20 0.42 ± 0.37 0.73 ± 0.23 0.36 ± 0.38 0.77 ± 0.20 -0.22 ± 0.69 0.90 ± 0.08

3.13 ± 0.92 1.23 ± 0.57 1.27 ± 0.66 1.59 ± 0.75 2.65 ± 0.96 1.70 ± 0.78 2.56 ± 0.91 1.57 ± 0.73 3.74 ± 1.19 0.99 ± 0.45

[5] Kyunghyun Cho, Bart Van Merriënboer, Çağlar Gulçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder–decoder for statistical machine translation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1724–1734, 2014. [6] Guangzhong Dong, Fukang Shen, Li Sun, Mingming Zhang, and Jingwen Wei. A bayesian inferred health prognosis and state of charge estimation for power batteries. IEEE Transactions on Instrumentation and Measurement, 74:1–12, 2024. [7] Shiyi Fu, Hongtao Fan, Zhaorui Jin, Fan Ji, Yulin Tao, Yachao Dong, Xunyuan Chen, Minghao Shao, Shuyu Yuan, Yu Wang, et al. Recent progress in state of health estimation for lithium-ion batteries: From laboratory to practical application. Renewable and Sustainable Energy Reviews, 226:116323, 2026. [8] Hossam A Gabbar, Ahmed M Othman, and Muhammad R Abdussami. Review of battery management systems (bms) development and industrial standards. Technologies, 9(2):28, 2021. [9] Yang Ge, Xingxing Jiang, and Benlian Xu. Deep koopman operator-

16

TABLE XVI F USION - LOSS ABLATION ON DATASET I WITH FROZEN EXPERTS . T HE TWO COEFFICIENTS WEIGHT THE RUL AND CAPACITY LOSSES ; NO RECONSTRUCTION TERM IS USED . λCap DENOTES THE CAPACITY COEFFICIENT. R EPORTED MEAN ± STANDARD DEVIATION ACROSS TEST CELLS AFTER AVERAGING PREDICTIONS OVER FIVE RUNS .

RUL prediction

Configuration RMSE (cycles)

Capacity estimation

R2

MAPE (%)

λRUL = 1, λCap = 1 143.69 ± 107.15 0.74 ± 0.23 RUL loss only (λRUL = 1, λCap = 0) 144.09 ± 108.54 0.74 ± 0.23 Q loss only (λRUL = 0, λCap = 1) 778.07 ± 179.51 -6.55 ± 0.96 λRUL = 1, λCap = 0.5 141.64 ± 105.12 0.74 ± 0.23 λRUL = 1, λCap = 2 144.63 ± 107.08 0.74 ± 0.23 λRUL = 1.5, λCap = 2 143.62 ± 106.15 0.74 ± 0.23

based remaining useful life prediction of lithium-ion batteries under multi-condition scenarios. Journal of Energy Storage, 119:116369, 2025. [10] Yang Ge, Jiaxin Ma, and Guodong Sun. A structural pruning method for lithium-ion batteries remaining useful life prediction model with multihead attention mechanism. Journal of Energy Storage, 86:111396, 2024. [11] Arijit Guha and Amit Patra. Online estimation of the electrochemical impedance spectrum and remaining useful life of lithium-ion batteries. IEEE Transactions on Instrumentation and Measurement, 67(8):1836– 1849, 2018. [12] TaiLong Jing, Sheng Du, Cong Wang, Chao Wu, Witold Pedrycz, Mohamed Sharaf, and ZhiWu Li. A hybrid data-driven granular model for battery remaining useful life prediction. IEEE Transactions on Instrumentation and Measurement, 74:1–12, 2025. [13] Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. [14] Guijun Ma, Songpei Xu, Benben Jiang, Cheng Cheng, Xin Yang, Yue Shen, Tao Yang, Yunhui Huang, Han Ding, and Ye Yuan. Real-time personalized health status prediction of lithium-ion batteries using deep transfer learning. Energy & Environmental Science, 15(10):4083–4094, 2022. [15] Gabriele Patrizi, Fabio Canzanella, Nedka Dechkova Nikiforova, Rossella Berni, G Geoffrey Vining, and Lorenzo Ciani. A novel framework for estimating battery health indicator under different operating conditions. IEEE Transactions on Instrumentation and Measurement, 2026. [16] Tian Peng, Zhongzheng Mo, Jie Chen, Chenghao Sun, Zhi Wang, Muhammad Shahzad Nazir, and Chu Zhang. A novel mic-boa-tide fusion framework with kernel density estimation for point and probabilistic remaining useful life prediction of lithium-ion batteries. Applied Energy, 410:127570, 2026. [17] Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018. [18] Zhuoyi Qian, Zhen Chen, and Ershun Pan. A transfer condition-focused model for battery capacity forecast. Reliability Engineering & System Safety, 267:111928, 2026. [19] Abraham Savitzky and Marcel JE Golay. Smoothing and differentiation of data by simplified least squares procedures. Analytical chemistry, 36(8):1627–1639, 1964. [20] Kristen A Severson, Peter M Attia, Norman Jin, Nicholas Perkins, Benben Jiang, Zi Yang, Michael H Chen, Muratahan Aykol, Patrick K Herring, Dimitrios Fraggedakis, et al. Data-driven prediction of battery cycle life before capacity degradation. Nature energy, 4(5):383–391, 2019. [21] Yuchen Song, Xinyi Zhang, Yuhang Du, Shumei Cui, and Datong Liu. Uncertainty-oriented collaborative learning for on-orbit state-of-health estimation of satellite lithium-ion batteries considering multi-operating conditions. Applied Energy, 409:127457, 2026. [22] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. [23] Elias Vollert, Dominic Bresser, Marcel Weil, et al. Assessing the critical role of graphite in the carbon footprint of lithium-ion battery production. Journal of Energy Storage, 161:121887, 2026. [24] Xiankui Wu, Penghua Li, Zhongwei Deng, Zhitao Liu, Mekhrdod S Kurboniyon, Sheng Xiang, and Gang Yin. Ldnet-rul: lightweight de-

RMSE (mAh)

10.91 ± 5.17 12.36 ± 5.32 10.87 ± 5.17 993.27 ± 12.24 72.99 ± 5.24 18.77 ± 9.02 10.91 ± 5.36 15.50 ± 6.56 11.22 ± 5.17 12.85 ± 5.96 11.03 ± 5.18 13.08 ± 5.04

R2

MAPE (%)

0.90 ± 0.08 -559.18 ± 102.52 0.75 ± 0.20 0.84 ± 0.14 0.89 ± 0.10 0.89 ± 0.08

0.99 ± 0.45 93.07 ± 1.00 1.56 ± 0.84 1.26 ± 0.56 1.06 ± 0.52 1.05 ± 0.44

formable neural network for remaining useful life prognostics of lithiumion batteries. IEEE Transactions on Power Electronics, 40(9):13514– 13528, 2025. [25] Qiuyu Yu, Fujin Wang, Zhi Zhai, Shiyu Zheng, Bingchen Liu, Zhibin Zhao, and Xuefeng Chen. Multi-time scale feature extraction for early prediction of battery rul and knee point using a hybrid deep learning approach. Journal of Energy Storage, 117:116024, 2025.

Record · ID 1006846 · SHA-256 eb61dca2c98e5a83
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.