Dynamic monitoring of salicylic acid and vitamin B2 during the apple boiling process based on fluorescence spectroscopy and machine learning - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Food Chem X . 2026 Apr 6;35:103818. doi: 10.1016/j.fochx.2026.103818 Search in PMC Search in PubMed View in NLM Catalog Add to search Dynamic monitoring of salicylic acid and vitamin B 2 during the apple boiling process based on fluorescence spectroscopy and machine learning Haoran Xu Haoran Xu a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Haoran Xu a , Jiaqi Zheng Jiaqi Zheng a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Jiaqi Zheng a , Ze Tao Ze Tao b Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Ze Tao b , Yong Tan Yong Tan a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Yong Tan a , Xia Xiao Xia Xiao c Jilin Academy of Agricultural Sciences, Northeast Innovation Center for Agricultural Science and Technology of China, Changchun, Jilin 130033, China Find articles by Xia Xiao c , Chunyu Liu Chunyu Liu a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Chunyu Liu a, ⁎ , Xing Teng Xing Teng c Jilin Academy of Agricultural Sciences, Northeast Innovation Center for Agricultural Science and Technology of China, Changchun, Jilin 130033, China Find articles by Xing Teng c, ⁎ , Zheng Li Zheng Li a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Zheng Li a , Yi Zhang Yi Zhang a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Yi Zhang a , Ye Wang Ye Wang a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China Find articles by Ye Wang a , Xin Chen Xin Chen d China-aid Agricultural Technology Demonstration Centre (CATDC), Gwebi Agricultural College, Harare, Mashonaland West Province, Zimbabwe Find articles by Xin Chen d Author information Article notes Copyright and License information a Jilin Provincial Key Laboratory of International Science and Technology Cooperation on Spectral Application Technology and Agricultural Intelligent Development, Provincial Key Laboratory of Spectral Exploration Science and Technology, School of Physics, Changchun University of Science and Technology, Changchun, Jilin 130022, China b Changchun University of Science and Technology, Changchun, Jilin 130022, China c Jilin Academy of Agricultural Sciences, Northeast Innovation Center for Agricultural Science and Technology of China, Changchun, Jilin 130033, China d China-aid Agricultural Technology Demonstration Centre (CATDC), Gwebi Agricultural College, Harare, Mashonaland West Province, Zimbabwe ⁎ Corresponding authors. [email protected] [email protected] Received 2025 Nov 12; Revised 2026 Apr 1; Accepted 2026 Apr 2; Collection date 2026 Apr. © 2026 The Authors This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/). PMC Copyright notice PMCID: PMC13091386 PMID: 42006659 Abstract Thermal processing affects nutrient retention in fruits, yet quantitative in-process monitoring remains challenging. Excitation-emission matrix (EEM) fluorescence combined with principal component analysis (PCA) was used to track changes in nutritional components during apple boiling. Overlapped fluorescence bands were resolved using Gaussian multi-peak empirical function fitting to derive peak descriptors. re regions were screened using minimum redundancy-maximum relevance (mRMR) and χ 2 tests; the synthetic minority over-sampling technique (SMOTE) was used for class balancing. Support vector machine (SVM) achieved 90.84–97.76% accuracy for boiling-stage classification ( κ = 0.9019–0.9761). Partial least squares (PLS) provided the best concentration prediction for both vitamin B 2 ( R 2 = 0 . 9734 , root mean square error (RMSE) = 0.0239) and salicylic acid ( R 2 = 0 . 9614 , RMSE = 0.0187). Kinetics showed distinct maxima: salicylic acid (11.588–28.796 ng/mL) peaked at 30–35 min (75–88 °C), while vitamin B 2 (2.164–6.001 ng/mL) peaked at 54–58 min at 100 °C. This peak-shape-aware EEM machine-learning framework supports non-destructive, quantitative process monitoring for nutrient retention during thermal processing. Keywords: EEM, Multi-peak fitting, Classification modeling, Time-series prediction, Dynamic monitoring Highlights • EEM fluorescence combined with PCA enables dynamic monitoring during apple boiling. • Gaussian multi-peak fitting resolves overlapping spectra for key nutrient signals. • mRMR, chi-squared tests, and SMOTE enhance feature selection and model robustness. • Integrated machine-learning models (RF, PLS, CNN, SVM) achieve R 2 = 0.9734. • An EEM-ML framework enables time-resolved, non-destructive monitoring. 1. Introduction Thermal processing of fruits ( Narra et al., 2024 ) such as boiling and steaming pervades human diets and the food industry and strongly influences nutrient retention and transformation. Apples, among the most widely consumed fruits worldwide, contain vitamin C and diverse bioactive phytochemicals, including polyphenols, which contribute to antioxidant and anti-inflammatory activities and are associated with metabolic and cardiovascular health benefits ( Zhang et al., 2023 ). Thermal processing such as boiling markedly alters these compounds’ concentrations and bioavailability ( Rickman et al., 2007 , Srichamnong et al., 2016 ). Vitamin C readily undergoes oxidative degradation at elevated temperatures, whereas phenolics such as salicylic acid can experience heat-induced release that enhances bioavailability ( De Santiago et al., 2018 , Gacnik et al., 2021 ). Conventional analytical techniques including high-performance liquid chromatography HPLC, UV–Vis spectrophotometry, and gas chromatography-mass spectrometry GC-MS ( Han et al., 2025 , Ma et al., 2015 ) quantify nutrients in fruits and processed products with high accuracy and specificity. However, laborious sample pretreatment, long turnaround, and destructive operation prevent these methods from supporting real-time, continuous monitoring during processing. Consequently, these approaches do not track nutrient dynamics in fruit processing in a real-time, continuous, non-destructive manner, which constrains fine-grained process optimization and nutrient preservation. Modern analytical technologies have opened new avenues for food quality assessment. Electrochemical biosensors rapidly quantify antioxidants in fruits, including polyphenols and vitamin C, with simple operation and low cost ( Hosseinikebria et al., 2025 , Jia et al., 2025 ). Raman spectroscopy and Fourier-transform infrared spectroscopy, FTIR, provide non-destructive, highly sensitive measurements and support rapid evaluation of sugar content, acidity, and ripeness in fruits ( Grabska et al., 2023 , Wei et al., 2025 ). Biosensors track antioxidant activity in juice in real time, and Raman spectroscopy detects volatile compounds in fruits ( Wu et al., 2023 ). However, these technologies face limitations: food matrices such as pH and turbidity undermine biosensor stability. Raman and infrared spectra remain difficult to interpret in complex matrices. Most studies emphasize static quality assessment and fail to capture dynamic nutrient changes during processing with precision ( Zhang et al., 2024 ). These gaps motivate the development of efficient techniques for dynamic monitoring in food science. Fluorescence spectroscopy offers a non-destructive, highly sensitive, and rapid tool for food quality assessment ( Ma et al., 2025 , Zaroual et al., 2025 ). EEM provides multidimensional molecular information, and suits qualitative and quantitative analysis of nutrients in complex food systems ( Gu et al., 2024 ). Recent work increasingly integrates fluorescence spectroscopy with machine learning in food analysis ( Lun et al., 2025 , Yuan et al., 2025 ). Chen et al. combined excitation-emission matrix fluorescence with machine learning to identify and quantify adulteration in camellia oil, achieving accuracies above 90% ( Chen et al., 2023 ). Venturini et al. applied fluorescence spectroscopy combined with a one-dimensional convolutional neural network (1D-CNN) to olive oil quality evaluation, demonstrating that physicochemical properties could be predicted rapidly from a single fluorescence spectrum ( Venturini et al., 2023 ). Shen et al. developed a compact three-dimensional fluorescence spectroscopy platform for rapid food-safety applications and improved efficiency with PLS modeling ( Shen et al., 2024 ). Moreover, SMOTE ( Zhang et al., 2022 ) and feature selection methods such as mRMR enhance model performance and address overlapping fluorescence signals and high-dimensional data ( Alsalem, 2025 , Zhao et al., 2019 ). Despite these advances, applications of fluorescence spectroscopy to dynamic nutrient monitoring in fruit processing remain limited, and key bottlenecks persist in resolving overlapping signals in complex matrices and achieving real-time monitoring. Fluorescence spectroscopy also has inherent limitations. It primarily probes fluorophores or analytes that can be tracked via fluorescent products and complexes, and therefore cannot directly quantify many nutrients that are non-fluorescent or have extremely weak fluorescence. In addition, fluorescence signals can be influenced by matrix effects (e.g., turbidity and inner-filter effect), as well as physicochemical changes such as pH and temperature, and spectral overlap may further complicate interpretation. Accordingly, fluorescence is most reliable for time-resolved tracking and modeling of target compounds with confirmed fluorescence signatures under well-controlled conditions, rather than as a universal assay for all nutrients. This study systematically monitors the dynamic trajectories of salicylic acid and vitamin B 2 during apple boiling by coupling EEM with Gaussian multi-peak fitting. It further integrates ensemble classification models (SVM, KNN, RF) and time-series prediction models (PLS, 1D-CNN, SVM), together with feature selection via mRMR and chi-squared ( C h i 2 ) tests and data balancing using SMOTE, to elucidate how both compounds evolve under different processing conditions. The workflow applies EEM to dynamic nutrient monitoring in apple boiling, uses Gaussian multi-peak fitting to effectively separate overlapping fluorescence signals and improve the identification accuracy of salicylic acid and vitamin B 2 , combines classification and forecasting models to reveal temporal patterns of nutrient change, and employs SMOTE to mitigate class imbalance and enhance model robustness. The findings provide a scientific basis for optimizing fruit-processing parameters and maximizing nutrient retention, and they advance the integration of fluorescence spectroscopy with machine learning for dynamic monitoring in food processing, offering a new pathway toward efficient, real-time food-quality assessment. 2. Materials and methods 2.1. Materials 2.1.1. Apparatus and reagents We selected Luochuan Fuji apples marketed in December 2024, as shown in Fig. 1 . Fruits with diameters of 7–8 cm and masses of 180–200 g were chosen to ensure uniform size and color. All apples were obtained from the same batch to minimize individual variability. A total of nine apples were used in this study (three independent boiling runs, three apples per run). The samples were stored at 4 °C and equilibrated to 25 °C prior to use. Fig. 1. Open in a new tab Apple sample picture ( a ) Front view ( b ) Top view. Analytical-grade reagents were employed in all experiments, including vitamin C (99%), salicylic acid (99%), and vitamin B 2 (97.5%–99%), purchased from Aladdin Biochemical Technology Co., Ltd. (Shanghai, China). Ultrapure water with a resistivity of 18.2 M Ω cm at 25 °C was used throughout the study. 2.2. Experimental method Fluorescence spectrophotometer (Agilent Technologies, Cary Eclipse G9800A, pulsed xenon lamp 80 Hz, maximum scan speed 24 000 nm/min, wavelength accuracy ± 1 . 5 nm , wavelength range 200–900 nm), spectrometer (Optosky, ATP2400, range 191–1102 nm, slit 50 nm, resolution 1.5 nm), four-window quartz cuvette (12.5 × 12.5 × 45 mm), water bath (Bona Technology Co. Ltd, HH-1, 220 V/300 W), and electronic balance (Xiamen Qunlong Instruments, QL-2003, capacity 200 g, accuracy 0.001 g) were employed for the measurements. Samples were illuminated using LED sources at 365 nm (10 W) and 295 nm (3 W) provided by Zhongshan Yanxizao Lighting. EEM: Cary Eclipse G9800A (Agilent Technologies) was utilized with both excitation and emission slits set to 10 nm and a scan speed of 600 nm/min. Excitation was scanned from 200 to 430 nm and emission from 300 to 600 nm. During boiling, spectra were acquired at representative time points covering both the heating and isothermal stages. Two-dimensional fluorescence (time-series monitoring): Use the setup shown in Fig. 2 . Illuminate the sample with 295 nm and 365 nm LED sources, set integration times to 500 ms and 200 ms, respectively, and set the filter parameter to 0. Starting at 0 min, withdraw 5 mL of solution every 2 min into quartz cuvette for immediate measurement, yielding 31 time points (0–60 min). The temperature of the boiling medium was recorded at each sampling time point, and the temperature ranges reported for kinetic maxima were based on these measurements. Measurements were performed immediately after sampling using a fixed handling procedure to minimize cooling-related variability and to approximate near-real-time monitoring. At each time point, acquire five replicates per sample and average the results. Select apples with matched size and similar appearance, repeat the entire procedure three times, and keep the water-bath setpoint (100 °C), total duration, and sampling interval identical across all samples. After each measurement, return the aliquot to the water bath to maintain system consistency. Fig. 2. Open in a new tab Optical path schematic diagram. (1) Computer (2) Spectrometers (3) Sample cuvette (4) Light source (5) Optical fiber (6) Optical platform. Sample preparation: Peel the apples and cut them into uniform cubes of 1.0 × 1.0 × 0.5 cm. Weigh 30.0 ± 0.1 g, immerse in 300 mL ultrapure water for 1 min without stirring, then heat in a water bath with the setpoint at 100 °C. The temperature of the boiling medium was recorded at each sampling time point and increased gradually during heating, reaching 100 °C at approximately 52–54 min and remaining at 100 °C thereafter. Define the entire boiling process (including heating and the isothermal stage) as 0–60 min. Analytical-grade standards: Dissolve the analytical standards in ultrapure water to prepare a series of solutions for constructing fluorescence intensity-concentration calibration curves. Salicylic acid concentrations were 3 . 4 , 4 . 5 , 6 . 0 , 9 . 1 , 12 . 1 , 16 . 1 , 21 . 5 , 28 . 7 , 38 . 2 , 50 . 9 , and 67 . 9 ng/mL , whereas vitamin B 2 concentrations were 0 . 442 , 0 . 737 , 1 . 23 , 2 . 04 , 3 . 41 , 5 . 68 , 9 . 47 , and 15 . 8 ng/mL . For each concentration, fluorescence spectra were recorded five times, and the averaged intensity was used for analysis. 3. Data processing and model construction 3.1. Data pre-processing Each dataset comprised 155 spectra corresponding to heating times from 0 to 60 min, with samples collected every 2 min (31 time points, five spectra per time point). Three independent datasets were collected for model development, resulting in a total of 155 × 3 spectra. For the two-dimensional time-series spectra (LED-based monitoring), each spectrum was truncated to the wavelength range of 300–700 nm; the EEM measurements followed the excitation/emission ranges described in Section 2.1 . For classification, labels were assigned according to the sampling time points, consistent with the observed spectral evolution. For prediction tasks, to reduce overly optimistic estimates due to correlated replicates, data partitioning was performed in a group-wise manner based on independent boiling runs (datasets), rather than random shuffling; replicate spectra from the same time point were kept within the same split. All data were standardized to zero mean and unit variance, and PCA was applied to retain 95% of the explained variance. To improve signal quality, the two-dimensional spectra were smoothed using a moving-average filter. 3.2. Feature selection To minimize interference from irrelevant features, feature selection was performed using the mRMR ( Ahmed et al., 2022 , Peng et al., 2005 ) and chi-square ( χ 2 ) ( Ahmed et al., 2020 , Curebal and Dag, 2024 ) methods. The mRMR algorithm was evaluated using the mutual information difference (MID) criterion, while the χ 2 test was applied with default parameters. From each method, the top 50 features were retained, and their intersection (approximately 30–40 features) was used for further analysis. During model training, features were incrementally added, and the optimal number of features was determined based on macro-F1 (for classification) or R 2 (for prediction) performance. The final models utilized 10–20 features for classification and 30 features for prediction ( Xie et al., 2023 ). 3.3. Classification models development To classify the concentration zones of salicylic acid and vitamin B 2 during boiling, three models were employed: SVM,KNN, and RF ( Arslan et al., 2025 , Hu et al., 2025 ). The dataset was divided into training and testing subsets at a 4:1 ratio, and model stability was evaluated using 10-fold cross-validation. To address class imbalance, the synthetic minority oversampling technique (SMOTE, k = 5 ) was applied to generate synthetic samples for minority classes. Hyperparameters were optimized through grid search. For SVM, the radial basis function (RBF) kernel was used with C ∈ [ 0 . 1 , 100 ] and γ ∈ [ 0 . 01 , 10 ] , for KNN, k ∈ [ 3 , 9 ] was explored using Euclidean distance and inverse-distance weighting, and for RF, n estimators ∈ [ 100 , 200 ] and min-samples-leaf=5 were evaluated. The classification cost matrix was weighted according to class sample sizes to penalize misclassifications that could disproportionately affect food quality. Each model was trained for ten iterations, and the one achieving the highest macro- F 1 score was selected. Model performance was assessed using precision, recall, macro- F 1 , weighted- F 1 , ROC-AUC, and kappa ( Chicco and Jurman, 2020 , Christen et al., 2024 ). A normalized confusion matrix and learning curves were further analyzed to evaluate performance consistency and overfitting risk ( Conciatori et al., 2024 ). Feature importance was determined using the mRMR and chi-square ( χ 2 ) methods, and the intersection of their top-ranked features (10–20) was retained to optimize classification accuracy for salicylic acid and vitamin B 2 . 3.4. Prediction model development The concentration dynamics of salicylic acid and vitamin B 2 during boiling were predicted using RF, PLS, 1D-CNN, and SVM models ( Ouyang et al., 2024 , Sun et al., 2024 ). Feature selection was performed by ranking wavelength variables in descending order of variance, and the top 30 features were retained to capture spectral characteristics associated with changes in food composition. For time-series forecasting, a sliding-window approach (window size = 7) was used to generate sequential input data within each independent boiling run. To reduce information leakage from correlated replicates, model evaluation was conducted using run-wise (grouped) data partitioning and cross-validation, i.e., training and testing were separated by independent boiling runs and replicate spectra from the same time point were kept within the same split. Hyperparameters were optimized through grid search with the following configurations: for RF, 150 trees with a minimum of 5 samples per leaf; for PLS, the number of latent variables was tuned within the range [ 4 , 8 ] ; for the 1D-CNN, three convolutional layers (kernel size = 3 , ReLU activation) were used, along with the Adam optimizer (learning rate = 0 . 001 ), batch normalization, and a dropout rate of 0.3 to mitigate overfitting; and for SVM, the radial basis function (RBF) kernel was used with C ∈ [ 0 . 01 , 100 ] and ɛ ∈ [ 0 . 01 , 1 ] . Model training was parallelized to enhance computational efficiency. Model performance was quantified using the mean squared error (MSE), root-mean-square error (RMSE), mean absolute error (MAE), coefficient of determination ( R 2 ), and cross-validated mean squared error (CV-MSE) ( Karunasingha, 2022 , Lee et al., 2023 ). Between-model differences were analyzed using two-sample t -tests, and the residual standard deviation was used to assess prediction stability, ensuring reliable concentration estimation for food-processing applications (see Fig. 3 ). Fig. 3. Open in a new tab EEM of apple water solution at different time points (a) 0 min (b) 15 minutes (c) 30 min (d) 45 min (e) 60 minutes. 4. Result and discussion 4.1. Spectral analysis 4.1.1. EEM spectral measurement and analysis From the EEM maps obtained at five time points for the apple samples, substantial changes in the amount of leached fluorescent substances were clearly observed as the heating temperature and duration increased. At 0 min, strong fluorescence signals appeared at Ex/Em = 275/340 nm and 230/320 nm, primarily originating from apple flavonoids. At 15, 30, 45, and 60 min, as shown in Fig. 4 , the main fluorescence region was concentrated near 275/340 nm. However, molecular aggregation broadened the spectral envelope, indicating that EEM alone was insufficient for precise molecular identification. Fig. 4. Open in a new tab EEM pure at different concentrations (a)aqueous solution of vitamin B 2 with a concentration of 41 . 2 ng/mL (b)aqueous solution of Salicylic acid with a concentration of 66 . 7 ng/mL . Analytical-standard EEM maps ( Fig. 4 ) revealed two excitation bands for salicylic acid at 300 and 230 nm, and two for vitamin B 2 at 370 and 280 nm. Comparison with these standards confirmed that boiling facilitated the release of salicylic acid and vitamin B 2 from apples. Since most other fluorescent components could not be quantitatively analyzed, subsequent investigations focused on the dynamic behavior of salicylic acid and vitamin B 2 during boiling. The excitation band of the EEM was set to 200–430 nm, enabling successful detection of fluorescence signals from salicylic acid and vitamin B 2 , with emission ranges of 400–450 nm and 520–530 nm, respectively. Salicylic acid, a key phenolic compound in apples, exhibited strong fluorescence under 230–240 nm and 280–300 nm excitation, consistent with deep-UV imaging of the peel and outer cortex. Vitamin B 2 produced a characteristic emission peak at 520–530 nm under 270–280 nm excitation, and its high quantum yield allowed detection even at trace concentrations. The apple matrix likely contained additional fluorophores, including hydroxycinnamic acids (230–250 nm and 280–320 nm excitation, 400–475 nm emission) and flavonoids (230–260 nm and 280–350 nm excitation, 350–500 nm emission), whose signals overlapped with those of salicylic acid and vitamin B 2 , complicating spectral assignment. To resolve these overlaps, Gaussian multi-peak fitting was applied to the EEM, successfully isolating the characteristic peaks of vitamin B 2 . The excitation band of 200–430 nm was selected based on reported fluorescence properties of apple phenolics and vitamins and refined through experimental verification to ensure specificity. Multi-peak fitting markedly improved identification accuracy, particularly in the presence of background fluorescence from protein-like substances (220–290 nm excitation, 300–350 nm emission). Compared with one-dimensional fluorescence, EEM analysis provided richer molecular information and was well suited for dynamic monitoring of complex food matrices. EEM at five representative time points were used for fluorophore screening and assignment, whereas spectra at 2-min intervals were used for dynamic fitting and quantification. Fluorescence spectra were acquired at uniform time intervals under 365 nm and 295 nm excitation. Fluorescence spectra were acquired at uniform time intervals under 365 nm and 295 nm excitation. The presence of salicylic acid and vitamin B 2 was confirmed ( Fig. 5 ). Pronounced temporal variations in overall fluorophore abundance were observed during boiling. Minor signals attributable to intermediates of vitamin C oxidation were detected, but were not pursued further because their levels were low and quantitative standards were unavailable ( Fig. 6 ). In the emission window of vitamin B 2 , fluorescence from neighboring carbonyl and ketone-like species partially overlapped the vitamin B 2 band, which hindered direct peak integration and therefore precluded reliable direct quantification. Gaussian multi-peak deconvolution was therefore adopted. Here, the Gaussian components are used as an empirical and numerically stable basis to decompose overlapping bands, rather than to assert a unique physical line shape (which can be influenced by instrumental broadening, reabsorption, and matrix effects). We chose Gaussian multi-peak fitting for targeted tracking of two known bands in a dense time series; Parallel Factor Analysis (PARAFAC) is powerful for global decomposition but was not required for our targeted quantification task. Because vitamin B 2 emission shifts with concentration and solvent polarity, peak positions were anchored using gradient-concentration experiments and in situ measurements. Model complexity was determined following a minimum-parameter principle: a two-component model with centers near 400 nm and 520–530 nm yielded stable fits with structureless residuals across time points, whereas additional components did not materially improve fit quality and reduced parameter interpretability. A two-component model centered at 400 nm and 520–530 nm was selected for subsequent fitting. Fig. 5. Open in a new tab Time-series fluorescence emission spectra of the apple boiling samples acquired under two excitation wavelengths. (a) Excitation at 295 nm. (b) Excitation at 365 nm. Spectra at successive sampling times (0–60 min, 2-min interval) are overlaid to illustrate the temporal evolution of fluorescence during boiling. For clarity, only the effective emission region is displayed, excluding excitation-related scattering and wavelength ranges with negligible signal. Fig. 6. Open in a new tab Fluorescence emission spectra of analytical-grade standards measured at a series of concentration levels under two excitation conditions. (a) Salicylic acid concentration-gradient spectra (excitation = 295 nm; spectra are shown in the effective emission window). (b) vitamin B 2 concentration-gradient spectra (excitation = 365 nm; spectra are shown in the effective emission window). For each panel, curves correspond to increasing concentration (from bottom to top). The purpose of this figure is to demonstrate the concentration-dependent spectral evolution that supports mapping fluorescence intensity to concentration; therefore, an explicit concentration bar is not displayed on the axes, while the concentration settings are provided in the Methods and Supplementary information. Gaussian multi-peak fitting was applied to the EEM data, which markedly improved spectral resolution and enabled clear tracking of the temporal evolution of vitamin B 2 , as illustrated in Fig. 7 . The fitted curve cleanly separated the characteristic peak of the target compound and exhibited excellent agreement with the analytical standard, confirming the reliability of the peak assignment. As shown in Fig. 8 , the fluorescence intensities of salicylic acid and vitamin B 2 varied nonlinearly over time initially increasing and then decreasing consistent with heat-induced molecular release followed by subsequent degradation. Salicylic acid is an aromatic organic acid consisting of a benzene ring ( C 6 H 6 ) substituted with a hydroxyl group at the ortho (2-) position ( − OH ) and a carboxyl group at the 1-position ( − COOH ). The ortho configuration facilitates the formation of an intramolecular hydrogen bond, which enhances molecular stability. Salicylic acid has a molecular weight of 138 . 12 g/mol , the aromatic ring contributes to hydrophobicity, whereas the hydroxyl and carboxyl groups confer hydrophilicity, resulting in moderate overall polarity. Fig. 7. Open in a new tab Multi-peak fitting results of the fluorescence spectra of apple samples shown in the informative emission window (cropped for clarity). (a) Fitting for the component centered at 520–530 nm; (b) fitting for the component centered at 400 nm. Portions dominated by excitation-related artifacts and low-information tails were omitted for better readability. Fig. 8. Open in a new tab The fluorescence intensity of apple sample aqueous solution changes over time under excitation of different wavelength LEDs. (a) Vitamin B 2 under excitation of 365 nm LED (b) Salicylic acid under excitation of 295 nm LED. Salicylic acid was characterized as sparingly soluble in cold water, with a solubility of approximately 0 . 2 g / 100 mL at 20 ∘ C , but readily soluble in hot or boiling water. The low aqueous solubility was attributed to the hydrophobic nature of the benzene ring, whereas the enhanced solubility in polar solvents such as ethanol and hot water arose from the presence of the polar hydroxyl and carboxyl groups. In addition, the intramolecular hydrogen bond within the salicylic acid molecule weakened its intermolecular interactions with water, thereby further reducing solubility at low temperatures. During boiling, the concentration of salicylic acid was observed to increase progressively, as the compound exhibited high thermal stability and was continuously released from the apple matrix. Because the release rate exceeded the degradation rate at typical heating temperatures, the salicylic acid content increased up to approximately 36 min . Prolonged heating beyond this period resulted in a slight decline due to thermal decomposition and oxidation. Vitamin B 2 is a flavin compound consisting of an isoalloxazine nucleus-a tricyclic, nitrogen-containing heterocycle-linked to a ribitol side chain, with a molecular weight of 376 . 36 g/mol . The isoalloxazine core is responsible for its fluorescence, while the ribitol chain enhances hydrophilicity. Multiple polar functional groups (e.g., hydroxyl and amide) confer partial water solubility, although the compound remains only sparingly soluble overall. Vitamin B 2 exhibits water solubility of approximately 12 mg / 100 mL at 27 . 5 ∘ C . It is thermally stable in neutral or acidic aqueous media but dissolves readily and becomes unstable under alkaline conditions. The compound is photosensitive to both visible and ultraviolet light and undergoes irreversible photolysis to lumiflavin or related isoalloxazine derivatives. Thermal degradation is strongly pH-dependent, occurring slowly in acidic environments but accelerating in alkaline solution, where heating can oxidize or cleave the isoalloxazine ring, leading to the loss of bioactivity. Light exposure and alkalinity promote decomposition, whereas under neutral or acidic conditions at 100 ∘ C , thermal breakdown remains limited. Within the range of 60 – 100 ∘ C , vitamin B 2 is relatively stable, and the mildly acidic conditions of apple boiling further suppress oxidation and decomposition, facilitating its release into the aqueous phase. After sufficient leaching, however, prolonged heating shifts the balance such that oxidative degradation exceeds release, and the vitamin B 2 concentration begins to decline after approximately 54 min . As shown in Fig. 8 , the peak profiles of vitamin B 2 and salicylic acid exhibited pronounced time dependence, with their rates of change following distinct temporal patterns. These dynamic spectral variations provide a solid foundation for the subsequent classification modeling of concentration-distribution categories and the time-series prediction modeling of concentration-change trends. Salicylic acid and vitamin B 2 exhibited pronounced concentration variations during boiling. To enable quantitative analysis, analytical standards were measured over the same spectral bands as the samples to construct concentration gradients. The gradient datasets were fitted with high accuracy ( Fig. 9 , Fig. 10 ), yielding robust spectrum-to-concentration calibration models for quantitative estimation. In this work, the regression models were trained to predict fluorescence spectra (or their spectral representations), and concentrations were subsequently obtained by applying the calibration mapping to the predicted spectra. By comparing these calibration models with the sample spectra, the concentrations of vitamin B 2 and salicylic acid during boiling were estimated to be 2.164–6.001 ng/mL and 11.588–28.796 ng/mL, respectively. These results provide a solid foundation for subsequent classification modeling and time-series prediction analyses. Fig. 9. Open in a new tab The concentration gradient of Vitamin B 2 , with the average value taken from five measurements. Fig. 10. Open in a new tab The concentration gradient of Salicylic acid with the average value taken from five measurements. Boiling markedly altered the intensity of compounds leached from apples. The salicylic-acid-associated feature increased during the early heating stage, consistent with heat-enhanced release and mass transfer from disrupted tissues. With prolonged heating, the signal gradually declined, plausibly due to oxidative loss, thermal degradation, or fluorescence quenching as the matrix evolved. For vitamin B 2 , the fitted component near 520–530 nm showed a similar rise-then-fall trend, suggesting that extraction dominated initially, whereas degradation and quenching became more influential after extended heating. Because the protocol included a heating-up period followed by a near-boiling plateau, these dynamics likely reflect the combined effects of temperature evolution and exposure time, together with time-dependent changes in the matrix and optical properties. Quantitatively, the salicylic-acid signal increased by 146.15% and reached a maximum at approximately 32 min, then decreased by 19.35% to a minimum before 56 min. Vitamin B 2 fluorescence increased by 178.15% before 56 min and subsequently declined by 9.62% before 60 min. Cell-wall disruption during boiling likely promoted the release of other phenolic fluorophores (e.g., hydroxycinnamic acids and flavonoids), increasing spectral overlap and complexity, while concomitant changes in the apple matrix (including pH and turbidity) further modulated the fluorescence dynamics. Monitoring accuracy was improved by combining Gaussian multi-peak fitting with targeted excitation at 295 nm and 365 nm, which helped separate overlapping bands and isolate the contributions of salicylic acid and vitamin B 2 from other fluorophores (e.g., protein-like species). Nevertheless, matrix effects during boiling-including increased turbidity, light scattering, and inner-filter effects-can still compromise apparent intensity and stability. Therefore, further refinement of preprocessing and correction strategies will be pursued to mitigate these interferences. 5. Optimization of classification models with SVM, KNN, and RF as comparisons 5.1. Comparative analysis and optimization process of vitamin B 2 classification models To characterize the temporal evolution of vitamin B 2 during boiling, a classification framework was established. Because fluorescence from other species obscured the riboflavin peak, peak separation followed by Gaussian fitting was applied to isolate the target signal. Feature vectors derived from the fitted peak parameters were then used to train the classifiers. Temporal labels were initially defined by partitioning the trajectory into contiguous 2 min bins (0–2 min as the first class, followed by successive 2 min intervals), yielding 30 classes in total. To ensure consistency with the labeling strategy used for salicylic acid and to reduce overly fine temporal granularity, an alternative coarse-grained scheme was further adopted: the first 40 min were grouped into 4 min classes, while all time points from 42 to 60 min were merged into a single class. In addition, considering the actual concentration-evolution pattern of vitamin B 2 during boiling, we further introduced a 15-class formulation in which the entire 0–60 min trajectory was uniformly partitioned into 4 min intervals (one class per 4 min). This 15-class scheme better reflects the underlying physicochemical dynamics of vitamin B 2 release and degradation and therefore provides a more physically meaningful labeling strategy. Performance was assessed for SVM, KNN, and RF models (see Table 1 ). Table 1. Evaluation indicators of the original classification model of vitamin B 2 . Accuracy Macro F1 Kappa Fold SVM 45 . 33 ± 9 . 86 % 0.2256 0 . 4354 ± 0 . 1018 45.333% KNN 47 . 17 ± 11 . 17 % 0.2356 0 . 4544 ± 0 . 1152 46.567% RF 40 . 67 ± 9 . 09 % 0.2011 0 . 3873 ± 0 . 0937 40.668% Open in a new tab The baseline classifiers exhibited relatively limited performance. The SVM achieved the accuracy of 45 . 33 ± 9 . 86 % , macro-F1 score of 0.2256, κ = 0 . 4354 ± 0 . 1018 , and fold-wise cross-validation (CV) accuracy of 45.33%. The KNN model yielded slightly higher performance, with the accuracy of 47 . 17 ± 11 . 17 % , macro-F1 score of 0.2356, κ = 0 . 4544 ± 0 . 1152 , and CV accuracy of 46.57%. In contrast, the RF classifier demonstrated the lowest performance, achieving the accuracy of 40 . 67 ± 9 . 09 % , macro-F1 score of 0.2011, κ = 0 . 3873 ± 0 . 0937 , and CV accuracy of 40.67%. The relatively large standard deviations, such as 11.17% observed for the KNN model, indicate unstable discrimination of the spectral data and limited generalization capability, primarily attributed to unoptimized hyperparameters and the intrinsic spectral complexity, encompassing feature redundancy, noise interference, and class imbalance. The classification models were optimized by analyzing the temporal evolution of vitamin B 2 fluorescence intensity and redefining the labeling scheme. Specifically, from 0 to 40 min, every two consecutive time points were grouped into a single class, corresponding to 4 min per class, whereas from 42 to 60 min, all time points were merged into one class (see Table 2 ). Table 2. Evaluation metrics for the vitamin B 2 classification model post initial optimization. Accuracy Macro F1 Kappa Fold SVM 56 . 21 ± 7 . 20 % 0.3250 0 . 5374 ± 0 . 0752 56.208% KNN 64 . 62 ± 11 . 63 % 0.4932 0 . 6172 ± 0 . 1278 64.624% RF 55 . 62 ± 10 . 14 % 0.3661 0 . 5226 ± 0 . 1110 55.625% Open in a new tab A marked improvement in model performance was observed after the initial optimization. The SVM achieved an accuracy of 56 . 21 % ± 7 . 20 % , macro- F 1 score of 0.3250, κ = 0 . 5374 ± 0 . 0752 , and cross-validation accuracy of 56.21%. The KNN model performed best, reaching 64 . 62 % ± 11 . 63 % accuracy, macro- F 1 score of 0.4932, κ = 0 . 6172 ± 0 . 1278 , and cross-validation accuracy of 64.62%. The RF model achieved 55 . 62 % ± 10 . 14 % accuracy, macro- F 1 score of 0.3661, κ = 0 . 5226 ± 0 . 1110 , and cross-validation accuracy of 55.63%. Among the three models, KNN demonstrated the strongest classification capability at this stage. KNN’s higher macro- F 1 score indicated better class balance. However, the large standard deviation for KNN (11.63%) and the variability for RF (10.14%) suggest that further optimization is required to improve stability. The classifiers were further optimized by applying SMOTE oversampling and hyperparameter tuning to correct class-count imbalance. To mitigate overfitting potentially introduced by oversampling, class-weighted cost matrix and regularization terms were incorporated. Each model was trained for 30 iterations, and the configuration yielding the best overall performance was selected. Based on the 10-fold cross-validation results, the Friedman test on model accuracies yielded p = 0 . 2913 , indicating that the performance differences among SVM, KNN, and RF are not statistically significant at the 0.05 level. Among the three models, KNN achieved the best overall performance with an accuracy of 97 . 847 % ± 1 . 764 % , Macro- F 1 of 0.9780, and κ = 0 . 976 ± 0 . 019 ( Table 3 ). RF followed with an accuracy of 97 . 189 % ± 2 . 210 % , Macro- F 1 of 0.9708, and κ = 0 . 969 ± 0 . 024 , while SVM obtained an accuracy of 96 . 697 % ± 2 . 202 % , Macro- F 1 of 0.9658, and κ = 0 . 964 ± 0 . 024 . In addition, the SVM achieved a one-vs-rest Macro AUC of 0.9994 on the hold-out evaluation, further supporting the strong discriminability of the vitamin B 2 spectral data under the 11-class formulation. Table 3. Final evaluation metrics for the optimized vitamin B 2 classification model under the 11-class formulation. Accuracy Macro F1 Kappa Fold SVM 96 . 70 ± 2 . 20 % 0.9658 0 . 9637 ± 0 . 0242 96.697% KNN 97 . 85 ± 1 . 76 % 0.9780 0 . 9763 ± 0 . 0194 97.847% RF 97 . 19 ± 2 . 21 % 0.9708 0 . 9691 ± 0 . 0243 97.189% Open in a new tab Table 4 summarizes the final performance of the optimized vitamin B 2 classification models under the 15-class formulation after SMOTE. Overall, KNN achieved the best classification performance, with an accuracy of 97 . 21 ± 1 . 71 % , a macro- F 1 score of 0.9718, and a κ of 0 . 9702 ± 0 . 0184 , indicating strong discriminative capability and robust agreement across all 15 classes. RF ranked second, reaching an accuracy of 95 . 88 ± 1 . 91 % , a macro- F 1 score of 0.9580, and a κ of 0 . 9558 ± 0 . 0205 , demonstrating stable performance but slightly reduced sensitivity to fine-grained temporal boundaries. In comparison, SVM showed the lowest overall performance among the three models, with an accuracy of 94 . 30 ± 3 . 46 % , a macro- F 1 score of 0.9419, and a κ of 0 . 9390 ± 0 . 0371 , and also exhibited the largest standard deviation, suggesting higher variability across cross-validation folds. The fold accuracies (94.302%, 97.214%, and 95.877% for SVM, KNN, and RF, respectively) were consistent with the averaged results, further confirming that KNN provides the most reliable and accurate framework for vitamin B 2 recognition in the 15-class setting. Table 4. Final evaluation metrics for the optimized vitamin B 2 classification model under the 15-class formulation. Accuracy Macro F1 Kappa Fold SVM 94 . 30 ± 3 . 46 % 0.9419 0 . 9390 ± 0 . 0371 94.302% KNN 97 . 21 ± 1 . 71 % 0.9718 0 . 9702 ± 0 . 0184 97.214% RF 95 . 88 ± 1 . 91 % 0.9580 0 . 9558 ± 0 . 0205 95.877% Open in a new tab As shown in Fig. 11 , Fig. 12 , Fig. 13 , Fig. 14 and Table 4 , under the 15-class formulation (after SMOTE), KNN achieved the best overall performance, yielding an accuracy of 97 . 21 ± 1 . 71 % , a macro- F 1 score of 0.9718, and a κ statistic of 0 . 9702 ± 0 . 0184 . RF ranked second with an accuracy of 95 . 88 ± 1 . 91 % , a macro- F 1 score of 0.9580, and a κ statistic of 0 . 9558 ± 0 . 0205 , whereas SVM showed comparatively lower performance, with an accuracy of 94 . 30 ± 3 . 46 % , a macro- F 1 score of 0.9419, and a κ statistic of 0 . 9390 ± 0 . 0371 , together with the largest inter-fold variation. The confusion matrices indicate that errors were limited and primarily occurred between temporally adjacent or near-adjacent classes (e.g., Class 7–8, 10–11, and 14–15), suggesting that most misclassifications were driven by the gradual spectral evolution along the boiling timeline rather than random confusion. Fig. 11. Open in a new tab Training curve of vitamin B 2 classification model for apple samples. Fig. 12. Open in a new tab Confusion matrix and normalized confusion matrix of vitamin B 2 classification model. The upper side is the confusion matrix and the lower side is the normalized confusion matrix. Fig. 13. Open in a new tab PCA dimensionality reduction of the vitamin B 2 classification model for apple samples. Fig. 14. Open in a new tab Feature selection for vitamin B 2 classification model of apple samples. The learning curves further show that the train-test gap decreased as the training set expanded for all models. RF exhibited a relatively larger gap at small sample sizes, indicating higher sensitivity to sample scarcity, while SVM and KNN became more stable as more samples were included. Notably, KNN maintained very high training accuracy throughout; however, its test accuracy increased steadily and remained high, implying that any overfitting effect was limited in practice under the current dataset scale. Classification modeling was treated as a key step for dynamically monitoring analyte variations during apple boiling. SVM, KNN, and RF were employed to classify the temporal states of vitamin B 2 and salicylic acid based on fluorescence spectra. For vitamin B 2 , the initial fine-grained labeling (each time point as a separate class) resulted in limited performance, with accuracies ranging from 40.67% to 47.17% and macro- F 1 scores from 0.2011 to 0.2356, reflecting spectral complexity and limited separability under overly fine temporal partitioning. After redefining the class scheme and applying imbalance handling and hyperparameter optimization, the vitamin B 2 models achieved substantially improved and more stable performance. Although KNN and RF achieved slightly higher peak accuracies than SVM, they exhibited larger train–test gaps, whereas SVM showed smaller gaps and better cross-batch robustness while maintaining competitive accuracy; therefore, SVM was selected as the best-performing model. The best SVM configuration reached a macro- F 1 score of approximately 0.97 and a high κ statistic, indicating markedly enhanced discriminative power. 5.1.1. Comparison and analysis of salicylic acid classification models Classification models were constructed for salicylic acid during boiling by leveraging its prominent fluorescence peak to characterize the temporal variation in peak intensity. Classes were defined by grouping the 0–28 min range such that every two consecutive time points formed one class, whereas the 30–60 min range was merged into a single class. Three algorithms (SVM, KNN, and RF) were evaluated for classification performance (see Table 5 ). Table 5. Evaluation indicators of the salicylic acid classification model. Accuracy Kappa Fold SVM 99 . 4303 ± 0 . 9173 % 0 . 9937 ± 0 . 0101 100.0000% KNN 96 . 5820 ± 1 . 9771 % 0 . 9624 ± 0 . 0217 94.2308% RF 92 . 0464 ± 2 . 4861 % 0 . 9125 ± 0 . 0273 90.3846% Open in a new tab Consistent with the vitamin B 2 analysis, the SVM achieved the best overall performance for salicylic acid, reaching an accuracy of 99 . 43 % ± 0 . 92 % with a κ statistic of 0 . 9937 ± 0 . 0101 , indicating near-perfect agreement and strong stability across folds. The KNN model followed with an accuracy of 96 . 58 % ± 1 . 98 % and a κ statistic of 0 . 9624 ± 0 . 0217 . In contrast, RF exhibited weaker performance, achieving an accuracy of 92 . 05 % ± 2 . 49 % and a κ statistic of 0 . 9125 ± 0 . 0273 . The fold-wise accuracies were 100.00% (SVM), 94.23% (KNN), and 90.38% (RF), further confirming the reduced generalization of RF compared with SVM and KNN. As shown in Fig. S1–S4, the SVM also demonstrated the strongest recall and the cleanest confusion-matrix pattern, with misclassifications occurring only at a few isolated time points. Notably, the near-ceiling performance is consistent with the prominent salicylic-acid fluorescence signature and the coarse-grained temporal grouping; meanwhile, the low inter-fold variance and the confusion-matrix analysis suggest no obvious loss of generalization. Overall, these results support SVM as a robust and practical small-model framework for process monitoring of both salicylic acid and vitamin B 2 based on fluorescence spectra during boiling. In this context, SVM remains sensitive to feature quality, and its performance benefits from mRMR-based feature selection. Classification modeling was treated as a key step for dynamically monitoring analyte variations during apple boiling. SVM, KNN, and RF were employed to classify the temporal and concentration states of vitamin B 2 and salicylic acid based on fluorescence spectra. For vitamin B 2 , the initial fine-grained labeling (each time point as a separate class) led to limited performance, with accuracies ranging from 40.67% to 47.17% and macro- F 1 scores between 0.2011 and 0.2356, primarily due to severe class imbalance and spectral complexity. After introducing SMOTE to mitigate imbalance and performing hyperparameter optimization, the vitamin B 2 models exhibited markedly improved and stable performance (e.g., macro- F 1 increased to the 0.97–0.99 range alongside high κ ), indicating substantially enhanced discriminative power and robustness. For salicylic acid, leveraging its prominent fluorescence signature together with a physically motivated temporal grouping yielded consistently high cross-validated performance, with SVM achieving an accuracy of 99 . 43 % ± 0 . 92 % and a κ statistic of 0 . 9937 ± 0 . 0101 , while KNN and RF attained accuracies of 96 . 58 % ± 1 . 98 % and 92 . 05 % ± 2 . 49 % , together with κ statistics of 0 . 9624 ± 0 . 0217 and 0 . 9125 ± 0 . 0273 , respectively. Across both analytes, SVM provided the most reliable overall performance, showing higher accuracy and κ than KNN and RF and smaller inter-fold variability, which supports its adaptability to dynamic spectral data. KNN approached SVM on vitamin B 2 but was less stable for salicylic acid, suggesting higher sensitivity to local noise and motivating further feature refinement. RF degraded more noticeably on salicylic acid, implying sensitivity to spectral variability and distributional shifts during tree splitting. Overall, combining appropriate class definition with imbalance handling and feature selection enables robust small-model monitoring for fluorescence-based boiling processes. 5.2. A comparison of time series prediction using RF, PLS, 1D-CNN and SVM 5.2.1. Comparison and analysis of vitamin 2 prediction models Time-series prediction models were developed to estimate the concentrations of vitamin B 2 and salicylic acid. For vitamin B 2 , the spectra were preprocessed using Gaussian multi-peak fitting at 400 nm and 520–530 nm, and four forecasting models RF, PLS, 1D-CNN, and SVM were trained on the fitted peak data. The forecasting results for vitamin B 2 are summarized in Table 6 and illustrated in Fig. 15 , Fig. 16 . Table 6. Evaluation metrics for the vitamin B 2 predictive model. RF PLS 1D-CNN SVM MSE 4 . 1861 × 1 0 − 3 5 . 713 × 1 0 − 4 3 . 5147 × 1 0 − 3 4 . 8896 × 1 0 − 3 MAE 0.041463 0.011502 0.039637 0.033176 RMSE 0.0647 0.023902 0.059285 0.069922 R 2 0.80515 0.97341 0.8364 0.77242 CVMSE 5 . 7456 × 1 0 − 4 1 . 6747 × 1 0 − 5 2 . 3591 × 1 0 − 3 4 . 7516 × 1 0 − 4 Train times/s 3.8571 0.007008 25.448 33.827 CV times/s 3.0463 0.024945 21.87 49.089 Open in a new tab Fig. 15. Open in a new tab Residual values of the vitamin B 2 prediction model. Fig. 16. Open in a new tab Comparison chart of predicted and original vitamin B 2 values. The time-series forecasting performance for vitamin B 2 was evaluated across the four models. PLS achieved the best results, with MSE = 0.0005713, MAE = 0.0115, RMSE = 0.0239, CV-MSE = 1 . 6747 × 1 0 − 5 , and R 2 = 0 . 9734 , indicating excellent accuracy and goodness of fit. The SVM ranked second, yielding MSE = 0.003112, MAE = 0.0298, RMSE = 0.0558, CV-MSE = 0.0010375, and R 2 = 0 . 8551 . The 1D-CNN and RF models produced MSE values of 0.003457 and 0.004165, with corresponding R 2 values of 0.8391 and 0.8061, respectively. PLS also demonstrated the shortest runtime (training = 0.0197 s, cross-validation = 0.0484 s), whereas SVM incurred the highest computational cost (training = 36.04 s, cross-validation = 60.09 s), confirming PLS’s computational efficiency. As shown in Fig. 15 , Fig. 16 and summarized in Table 6 , both PLS and SVM produced tightly distributed residuals and closely tracked the ground truth during training, yielding accurate and practically valuable forecasts. Fig. 15 , Fig. 16 show the fitted curves and residual distributions for the spectral-intensity forecasting task. Both SVM and PLS substantially outperformed 1D-CNN and RF. The prediction curve of the SVM nearly coincided with the ground truth, and its residuals were tightly confined within the range of [ − 0 . 05 , 0 . 05 ] with minimal fluctuation, indicating that the model effectively captured the nonlinear structure of the spectral signals and maintained strong generalization capability. Similarly, the PLS model closely tracked the ground truth and exhibited nearly symmetric, highly concentrated residuals, demonstrating that its linear framework successfully represented the principal variation in spectral intensity while offering a clear computational advantage. The 1D-CNN captured the overall trend better than RF but still exhibited a systematic bias, with residuals concentrated within [ − 0 . 20 , − 0 . 05 ] . This bias is attributed to the network architecture’s limited ability to extract key discriminative spectral features, which shifts the nonlinear fit. In contrast, RF performed the worst, with prediction curves deviating markedly from the ground truth and residuals clustering in [ − 0 . 80 , − 0 . 20 ] , indicating a strong systematic bias. These results suggest that RF struggled to establish stable decision-split rules under high-dimensional spectral features and temporal dynamics, leading to underfitting. Based on the synthesized results, SVM and PLS were identified as the most suitable predictive models for this study. The SVM excelled at modeling nonlinear spectral features, whereas PLS combined high predictive accuracy with exceptional computational efficiency. Both models demonstrated strong potential for practical applications. By contrast, 1D-CNN and RF exhibited weaker predictive performance and would benefit from further optimization of network architecture or feature-extraction strategies. These findings are consistent with previous studies reporting the superiority of SVM and PLS in spectroscopic modeling, thereby reinforcing their applicability and robustness for complex spectral prediction tasks. 5.2.2. Comparison and analysis of vitamin B 2 prediction models The salicylic acid data were preprocessed by retaining the 370–580 nm spectral region, and RF, PLS, 1D-CNN, and SVM models were trained. The time-series prediction results for salicylic acid are presented in Table 7 . Table 7. Evaluation metrics for salicylic acid predictive modeling. RF PLS 1D-CNN SVM MSE 1 . 1623 × 1 0 − 2 3 . 5018 × 1 0 − 4 1 . 7461 × 1 0 − 3 2 . 5667 × 1 0 − 3 MAE 0.056078 0.01421 0.030869 0.028953 RMSE 0.10781 0.018713 0.041786 0.050663 R 2 −0.28165 0.96139 0.80746 0.71698 CVMSE 6 . 8879 × 1 0 − 4 2 . 5501 × 1 0 − 4 1 . 9616 × 1 0 − 3 1 . 3377 × 1 0 − 3 Train times/s 8.384 0.23085 15.453 112.16 CV times/s 36.602 0.79741 26.748 775.37 Open in a new tab Time-series prediction results for salicylic acid are summarized in Table 7 . PLS achieved the best performance (MSE = 3 . 5018 × 1 0 − 4 , MAE = 0.01421, RMSE = 0.01871, CVMSE = 2 . 5501 × 1 0 − 4 , R 2 = 0 . 96139 ), followed by 1D-CNN ( R 2 = 0 . 80746 ) and SVM ( R 2 = 0 . 71698 ). In contrast, RF yielded the poorest generalization on the independent test set (MSE = 1 . 1623 × 1 0 − 2 , RMSE = 0.10781, R 2 = − 0 . 28165 ). For clarity, the regression target in this section is the fluorescence spectrum (i.e., the multi-wavelength intensity vector within 370–580 nm), rather than concentration. Concentrations are obtained in a subsequent step using calibration relationships established from gradient-concentration experiments (e.g., mapping deconvolved peak intensity and area to concentration). Therefore, accurate spectral forecasting provides an indirect yet practical route to concentration estimation within the calibrated range, while preserving the full spectral structure for kinetic interpretation. The negative R 2 for RF indicates that the sum of squared prediction errors on the external test set exceeded the variance of the reference spectra (i.e., worse than a mean-spectrum baseline). This behavior is consistent with the limited extrapolation capability of tree ensembles under inter-run spectral variability and noise and with the high-dimensional and strongly collinear nature of spectroscopic data under a small-sample regime. In addition, RF was trained independently for each wavelength, and errors can accumulate when performance is aggregated at the spectrum level. By contrast, PLS explicitly exploits the covariance structure between predictors and spectra and is robust for collinear spectroscopic variables, which explains its consistently superior accuracy and efficiency in this task. Fig. S5–S6 further support these conclusions: PLS produced residuals most concentrated around zero, whereas RF exhibited the most dispersed residuals with systematic bias. Considering both analytes, PLS provides the most robust and computationally efficient model for spectral time-series modeling in this framework. Fig. S5–S6 show that the PLS model produced the most concentrated residuals near zero, indicating excellent agreement between predictions and the ground truth with minimal fitting error. The SVM model also exhibited a tight residual distribution but showed a slight systematic bias. In contrast, the 1D-CNN generated more dispersed residuals with a noticeable negative bias, revealing limitations in capturing spectral nonlinearity. The RF model displayed the most dispersed residuals and a clear systematic bias, confirming its poor adaptability to the complex and high-noise characteristics of the salicylic-acid spectra. Considering both vitamin B 2 and salicylic acid, PLS proved to be the most robust and efficient method for spectral time-series modeling. It maintained a consistently high goodness of fit and low prediction error across analytes while offering superior computational efficiency. SVM and 1D-CNN showed potential for nonlinear spectral modeling but remained limited by either accuracy or efficiency, whereas RF was unsuitable for spectral time-series prediction, particularly in modeling salicylic acid. These findings emphasize that careful model selection is essential to balance predictive accuracy and computational cost, identifying PLS as the optimal choice within this framework. The results demonstrate that SVM provides excellent discriminative performance in the classification tasks for both analytes, with high accuracy and Kappa values indicating strong adaptability to spectral characteristics. The regression analysis highlights the superiority of PLS, whose predictive accuracy and computational efficiency establish it as the preferred model for forecasting vitamin B 2 and salicylic acid concentrations. The pronounced performance decline of RF particularly in salicylic-acid classification and regression suggests a heightened sensitivity to the compound’s spectral complexity and noise. The inter-analyte performance differences primarily reflect whether the models were trained on raw spectra or peak-fitted spectra. Considering all evaluation metrics and the results of vitamin B 2 forecasting, PLS emerges as the most efficient and reliable approach for rapid and accurate on-site spectral detection. Future research could integrate advanced spectral preprocessing with chemical validation assays to further elucidate how analyte-specific physicochemical properties influence model performance and generalization. Time-series prediction models, including PLS, 1D-CNN, and SVM, were employed to quantify the concentration trajectories of salicylic acid and vitamin B 2 during the boiling process. Among these models, PLS exhibited the best performance, achieving MSE= 0.0003502–0.0005713 and R 2 =0.9614–0.9734, indicating high precision in capturing nonlinear dynamic variations. The 1D-CNN and SVM models followed, achieving R 2 =0.8075–0.8391 and 0.7170–0.8551, respectively, and both exhibited reliable predictive capability. The synergy between classification and time-series prediction establishes a comprehensive modeling framework: classification identifies concentration states such as “high” or “low,” while prediction quantifies their temporal evolution. Compared with conventional static spectral analysis, dynamic time-series prediction must accommodate temporal dependence and nonlinear spectral variation. Our multi-time-point acquisition strategy, combined with targeted spectral band selection, effectively captures these temporal dynamics. PLS also offers excellent computational efficiency (training time 0.0197–0.2309 s), making it well suited for real-time monitoring applications, whereas SVM incurs substantially higher computational cost (training time 36.04–112.16 s), thus requiring a trade-off between accuracy and efficiency in practical deployment. 6. Conclusion EEM spectroscopy combined with peak-shape-aware spectral deconvolution and machine-learning models enabled quantitative, non-destructive tracking of fluorophore-related nutrients during apple boiling. Gaussian multi-peak fitting was applied as an empirical basis-function deconvolution to separate overlapped fluorescence contributions and to derive peak descriptors for model input. For boiling-stage classification, the SVM model achieved the highest performance, with accuracy above 99% and high κ . For concentration inference using concentration-gradient calibration, PLS regression provided stable performance for both analytes (vitamin B 2 : R 2 = 0 . 9734 , RMSE = 0.0239; salicylic acid: R 2 = 0 . 9614 , RMSE = 0.0187). The random-forest model showed poor generalization for salicylic-acid spectra ( R 2 < 0 ), indicating limited suitability for this dataset. Estimated concentration ranges during boiling were 2.164–6.001 ng/mL for vitamin B 2 and 11.588–28.796 ng/mL for salicylic acid. Salicylic acid reached its maximum at 30–35 min (75–88 °C), whereas vitamin B 2 peaked at 54–58 min at 100 °C. These results define time–temperature intervals associated with maximal signal intensity under the present conditions. Limitations include sensitivity primarily to fluorophore-related compounds and potential matrix-dependent optical interferences. Further validation with complementary chemical assays and expanded analyte panels is recommended to improve specificity and generalizability. CRediT authorship contribution statement Haoran Xu: Writing – original draft, Visualization, Validation, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Jiaqi Zheng: Methodology, Data curation. Ze Tao: Writing – original draft, Conceptualization. Yong Tan: Supervision, Conceptualization. Xia Xiao: Visualization. Chunyu Liu: Supervision, Methodology. Xing Teng: Visualization. Zheng Li: Visualization. Yi Zhang: Visualization. Ye Wang: Visualization. Xin Chen: Supervision. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgments This work was supported by the Key Research and Development Program of Jilin Province under Grant No. 20240304191SF and by the Jilin Province Science and Technology Development Plan Project under Grant No. 20230101187JC In addition, this work was supported by the Provincial Undergraduate Innovation and Entrepreneurship Training Program of Changchun University of Science and Technology (no grant number available), including the projects “Research on a rice nutrient-solution concentration monitoring system based on fluorescence spectroscopy” and “Traceability study of rice from different origins based on spectral detection technology”. Footnotes Appendix A Supplementary material related to this article can be found online at https://doi.org/10.1016/j.fochx.2026.103818 . Contributor Information Chunyu Liu, Email: [email protected]. Xing Teng, Email: [email protected]. Appendix A. Supplementary data The following is the Supplementary material related to this article. MMC S1 Figures showing training, evaluation, feature selection, PCA, and prediction of salicylic acid classification in apples. mmc1.docx (1.5MB, docx) Data availability Data will be made available on request. References Ahmed Y.A., Huda S., Al-rimy B.A.S., Alharbi N., Saeed F., Ghaleb F.A., Ali I.M. A weighted minimum redundancy maximum relevance technique for ransomware early detection in industrial IoT. Sustainability. 2022;14(3):1231. doi: 10.3390/su14031231. [ DOI ] [ Google Scholar ] Ahmed Y.A., Koçer B., Huda S., Al-rimy B.A.S., Hassan M.M. A system call refinement-based enhanced Minimum Redundancy Maximum Relevance method for ransomware early detection. Journal of Network and Computer Applications. 2020;167 doi: 10.1016/j.jnca.2020.102753. [ DOI ] [ Google Scholar ] Alsalem K. A hybrid time series forecasting approach integrating fuzzy clustering and machine learning for enhanced power consumption prediction. Scientific Reports. 2025;15:6447. doi: 10.1038/s41598-025-91123-8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Arslan M., Zareef M., Afzal M., Tahir H.E., Shi J., Muhammad A., Rakha A., Zou X. A smart olfactory visualization system based on colorimetric sensor array and chemometrics for the identification of rice adulteration. Food Chemistry. 2025;492(Pt 3) doi: 10.1016/j.foodchem.2025.145581. [ DOI ] [ PubMed ] [ Google Scholar ] Chen A.-Q., Wu H.-L., Wang T., Wang X.-Z., Sun H.-B., Yu R.-Q. Intelligent analysis of excitation-emission matrix fluorescence fingerprint to identify and quantify adulteration in camellia oil based on machine learning. Talanta. 2023;251 doi: 10.1016/j.talanta.2022.123733. [ DOI ] [ PubMed ] [ Google Scholar ] Chicco D., Jurman G. The advantages of the matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics. 2020;21(1):6. doi: 10.1186/s12864-019-6413-7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Christen P., Hand D., Kirielle N. A review of the F-measure: Its history, properties, criticism, and alternatives. ACM Computing Surveys. 2024;56(3) doi: 10.1145/3606367. [ DOI ] [ Google Scholar ] Conciatori M., Valletta A., Segalini A. Improving the quality evaluation process of machine learning algorithms applied to landslide time series analysis. Computers and Geosciences. 2024;184 doi: 10.1016/j.cageo.2024.105531. [ DOI ] [ Google Scholar ] Curebal, F., & Dag, H. (2024). Enhancing Malware Classification: A Comparative Study of Feature Selection Models with Parameter Optimization. In 2024 systems and information engineering design symposium (pp. 511–516). Charlottesville, VA, USA: 10.1109/SIEDS61124.2024.10534669. [ DOI ] De Santiago E., Domínguez-Fernández M., Cid C., De Peña M.-P. Impact of cooking process on nutritional composition and antioxidants of cactus cladodes (Opuntia ficus-indica) Food Chemistry. 2018;240:1055–1062. doi: 10.1016/j.foodchem.2017.08.039. [ DOI ] [ PubMed ] [ Google Scholar ] Gacnik S., Veberič R., Hudina M., Marinovic S., Halbwirth H., Mikulič-Petkovšek M. Salicylic and methyl salicylic acid affect quality and phenolic profile of apple fruits three weeks before the harvest. Plants. 2021;10(9):1807. doi: 10.3390/plants10091807. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Grabska J., Beć K.B., Ueno N., Huck C.W. Analyzing the quality parameters of apples by spectroscopy from Vis/NIR to NIR region: A comprehensive review. Foods. 2023;12(10):1946. doi: 10.3390/foods12101946. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Gu H., Hu L., Dong Y., Chen Q., Wei Z., Lv R., zhou Q. Evolving trends in fluorescence spectroscopy techniques for food quality and safety: A review. Journal of Food Composition and Analysis. 2024;131 doi: 10.1016/j.jfca.2024.106212. [ DOI ] [ Google Scholar ] Han X., Wei Y., Yuan L., Yin X., Liu Y., Wang C., Jiang X., Li T., Liu Q. Characterization of flavor profiles of wines produced with Coniella vitis-infected grapes by GC–MS, HPLC, and sensory analysis. Food Chemistry. 2025;471 doi: 10.1016/j.foodchem.2025.142820. [ DOI ] [ PubMed ] [ Google Scholar ] Hosseinikebria S., Khazaei M., Dervisevic M., Judicpa M.A., Tian J., Razal J.M., Voelcker N.H., Nilghaz A. Electrochemical biosensors: The beacon for food safety and quality. Food Chemistry. 2025;475 doi: 10.1016/j.foodchem.2025.143284. [ DOI ] [ PubMed ] [ Google Scholar ] Hu X., Zeng J., Dai M., Li A., Liang Y., Lu W., Peng J., Tian J., Chen M., Huang D. Hyperspectral-driven PSO–SVM model and optimized CNN–LSTM–Attention fusion network for qualitative and quantitative non-destructive detection of adulteration in strong-aroma Baijiu. Food Chemistry. 2025;490 doi: 10.1016/j.foodchem.2025.145197. [ DOI ] [ PubMed ] [ Google Scholar ] Jia X., Guo H., Hao Y., Shi J., Chai R., Wang S., Wu H., Feng Y., Ji W., Wu S. Dual recognition electrochemical sensor for detection of gallic acid in teas, apples and grapes and derived products. Food Chemistry. 2025;490 doi: 10.1016/j.foodchem.2025.145063. [ DOI ] [ PubMed ] [ Google Scholar ] Karunasingha D.S.K. Root mean square error or mean absolute error? Use their ratio as well. Information Sciences. 2022;585:609–629. doi: 10.1016/j.ins.2021.11.036. [ DOI ] [ Google Scholar ] Lee D.H., Lee D., Han S., Seo S., Lee B.J., Ahn J. Deep residual neural network for predicting aerodynamic coefficient changes with ablation. Aerospace Science and Technology. 2023;136 doi: 10.1016/j.ast.2023.108207. [ DOI ] [ Google Scholar ] Lun Z., Wu X., Dong J., Wu B. Deep learning-enhanced spectroscopic technologies for food quality assessment: Convergence and emerging frontiers. Foods. 2025;14(13):2350. doi: 10.3390/foods14132350. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Ma B., Chen J., Zheng H., Fang T., Ogutu C., Li S., Han Y., Wu B. Comparative assessment of sugar and malic acid composition in cultivated and wild apples. Food Chemistry. 2015;172:86–91. doi: 10.1016/j.foodchem.2014.09.032. [ DOI ] [ PubMed ] [ Google Scholar ] Ma S., Li Y., Peng Y., Nie S., Bai X., Zhang J. Simultaneous prediction of pungency and color values of paprika via LED-induced fluorescence. Food Chemistry. 2025;493(Pt 1) doi: 10.1016/j.foodchem.2025.145745. [ DOI ] [ PubMed ] [ Google Scholar ] Narra F., Piragine E., Benedetti G., Ceccanti C., Florio M., Spezzini J., Troisi F., Giovannoni R., Martelli A., Guidi L. Impact of thermal processing on polyphenols, carotenoids, glucosinolates, and ascorbic acid in fruit and vegetables and their cardiovascular benefits. Comprehensive Reviews in Food Science and Food Safety. 2024;23(6) doi: 10.1111/1541-4337.13426. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Ouyang Q., Fan Z., Chang H., Shoaib M., Chen Q. Analyzing TVB-N in snakehead by Bayesian-optimized 1D-CNN using molecular vibrational spectroscopic techniques: Near-infrared and Raman spectroscopy. Food Chemistry. 2024;464(Pt 2) doi: 10.1016/j.foodchem.2024.141701. [ DOI ] [ PubMed ] [ Google Scholar ] Peng H., Long F., Ding C. Feature selection based on mutual information: criteria of max-dependency, max-relevance, and min-redundancy. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2005;27(8):1226–1238. doi: 10.1109/TPAMI.2005.159. [ DOI ] [ PubMed ] [ Google Scholar ] Rickman J.C., Barrett D.M., Bruhn C.M. Nutritional comparison of fresh, frozen and canned fruits and vegetables. Part 1. Vitamins C and B and phenolic compounds. Journal of the Science of Food and Agriculture. 2007;87(6):930–944. doi: 10.1002/jsfa.2825. [ DOI ] [ Google Scholar ] Shen F., Feng X., Li Y., Lin X., Cai F. Compact three-dimensional fluorescence spectroscopy and its application in food safety. LWT. 2024;202 doi: 10.1016/j.lwt.2024.116324. [ DOI ] [ Google Scholar ] Srichamnong W., Thiyajai P., Charoenkiatkul S. Conventional steaming retains tocols and γ -oryzanol better than boiling and frying in the jasmine rice variety khao dok mali 105. Food Chemistry. 2016;191:113–119. doi: 10.1016/j.foodchem.2015.05.027. [ DOI ] [ PubMed ] [ Google Scholar ] Sun P., Lin S., Li X., Li D. Different stages of flavor variations among canned antarctic krill (Euphausia superba): Based on GC-IMS and PLS-DA. Food Chemistry. 2024;459 doi: 10.1016/j.foodchem.2024.140465. [ DOI ] [ PubMed ] [ Google Scholar ] Venturini F., Sperti M., Michelucci U., Gucciardi A., Martos V.M., Deriu M.A. Extraction of physicochemical properties from the fluorescence spectrum with 1D convolutional neural networks: Application to olive oil. Journal of Food Engineering. 2023;336 doi: 10.1016/j.jfoodeng.2022.111198. [ DOI ] [ Google Scholar ] Wei C., Zhang J., Li G., Zhong Y., Ye Z., Wang H., Li K., Wu Y., Wu Y., Luo H., Sun Q., Weng Z. Rapid and non-destructive detection of formaldehyde adulteration in shrimp based on deep learning-assisted portable Raman spectroscopy. Food Chemistry. 2025;492(Pt 1) doi: 10.1016/j.foodchem.2025.145343. [ DOI ] [ PubMed ] [ Google Scholar ] Wu L., Tang X., Wu T., Zeng W., Zhu X., Hu B., Zhang S. A review on current progress of Raman-based techniques in food safety: From normal Raman spectroscopy to SESORS. Food Research International. 2023;169 doi: 10.1016/j.foodres.2023.112944. [ DOI ] [ PubMed ] [ Google Scholar ] Xie S., Zhang Y., Lv D., Chen X., Lu J., Liu J. A new improved maximal relevance and minimal redundancy method based on feature subset. Journal of Supercomputing. 2023;79(3):3157–3180. doi: 10.1007/s11227-022-04763-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Yuan Y., Ji Z., Fan Y., Xu Q., Shi C., Lyu J., Ertbjerg P. Deep learning-assisted fluorescence spectroscopy for food quality and safety analysis. Trends in Food Science & Technology. 2025;156 doi: 10.1016/j.tifs.2024.104821. [ DOI ] [ Google Scholar ] Zaroual H., El Hadrami E.M., Farah A., Ez zoubi Y., Chénè C., Karoui R. Detection and quantification of extra virgin olive oil adulteration by other grades of olive oil using front-face fluorescence spectroscopy and different multivariate analysis techniques. Food Chemistry. 2025;479 doi: 10.1016/j.foodchem.2025.143736. [ DOI ] [ PubMed ] [ Google Scholar ] Zhang Z., Li Y., Zhao S., Qie M., Bai L., Gao Z., Liang K., Zhao Y. Rapid analysis technologies with chemometrics for food authenticity field: A review. Current Research in Food Science. 2024;8 doi: 10.1016/j.crfs.2024.100676. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Zhang Z., Liu H., Chen D., Zhang J., Li H., Shen M., Pu Y., Zhang Z., Zhao J., Hu J. SMOTE-based method for balanced spectral nondestructive detection of moldy apple core. Food Control. 2022;141 doi: 10.1016/j.foodcont.2022.109100. [ DOI ] [ Google Scholar ] Zhang Y., Zeng M., Zhang X., Yu Q., Zeng W., Yu B., Gan J., Zhang S., Jiang X. Does an apple a day keep away diseases? Evidence and mechanism of action. Food Science & Nutrition. 2023;11(9):4926–4947. doi: 10.1002/fsn3.3487. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Zhao Z., Anand R., Wang M. 2019 IEEE international conference on data science and advanced analytics. 2019. Maximum relevance and minimum redundancy feature selection methods for a marketing machine learning platform; pp. 442–452. [ DOI ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials MMC S1 Figures showing training, evaluation, feature selection, PCA, and prediction of salicylic acid classification in apples. mmc1.docx (1.5MB, docx) Data Availability Statement Data will be made available on request. Articles from Food Chemistry: X are provided here courtesy of Elsevier ACTIONS View on publisher site PDF (3.3 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top