Early Detection of Water Stress by Plant Electrophysiology: Machine Learning for Irrigation Management Eduard Buss1*, Till Aust1 and Heiko Hamann1
arXiv:2604.28038v1 [cs.LG] 30 Apr 2026
1*
Department of Computer and Information Science, University of Konstanz, Universitätsstraße 10, Konstanz, 78464, Baden-Württemberg, Germany.
*Corresponding author(s). E-mail(s): [email protected]; Contributing authors: [email protected]; [email protected]; Abstract Purpose: Fast detection of plant stress is key to plant phenotyping, precision agriculture, and automated crop management. In particular, efficient irrigation management requires early identification of water stress to optimize resource use while maintaining crop performance. Direct physiological sensing offers the potential to detect stress responses before visible symptoms appear. Methods: In this study, we recorded electrophysiological signals from greenhouse-grown tomato plants subjected to water stress and developed a framework based on machine learning for online stress detection. The recorded time-series data were processed using a processing pipeline that includes statistical feature extraction and selection, automated machine learning or alternatively deep learning, and probability calibration. Results: Across multiple input time horizons, we found that a 30-minute look-back window strikes the best balance between rapid decision-making and classification performance. Using automated machine learning, the framework achieved classification accuracies of up to 92%, outperforming deep learning approaches. Sequential backward selection reduced the feature set while maintaining performance. Importantly, the framework detects transitions from healthy to stressed states in recordings that were not included in the training set. Conclusion: Overall, we provide a decision-support tool for farmers and establish a foundation for biofeedback-driven irrigation control to improve resource efficiency in (semi-)autonomous crop production systems.
1
Keywords: Precision Agriculture, Resource Management, Automated Machine Learning, Deep Learning, Decision Support System, Plant Electrophysiology
1 Introduction The European Environment Agency reports that 31% of total water abstraction in Europe is attributed to agriculture, primarily driven by irrigation (European Environment Agency 2025). As water demand continues to increase due to demographic growth, economic activities, and climate change, agriculture remains dependent on water availability to ensure yield stability, crop quality, and farmer resilience against irregular rainfall patterns (European Court of Auditors 2021). Agriculturally used water is sourced from surface water bodies such as rivers and lakes, groundwater reserves, rainwater harvesting, and reclaimed wastewater. However, its use in agriculture directly influences overall water quantity and quality through pollution from fertilizers and pesticides (Thompson et al. 2020). The European Common Agricultural Policy (CAP), as the primary framework governing agriculture and rural development, promotes more efficient water use to mitigate water scarcity (European Commission, Directorate-General for Agriculture and Rural Development 2025). Sustainable and efficient management of natural resources constitutes a key strategic objective of CAP. Traditional irrigation practices, such as commercial pre-programmed irrigation controllers, distribute the same amount of water over the whole field, ignoring local differences in soil conditions and the plants’ real hydration requirements (Abioye et al. 2020). There has been an increased focus on the development of advanced irrigation systems in recent decades due to the growing challenges posed by water scarcity, climatic variability, and the need for efficient resource use. Key advances have been in irrigation monitoring, for example, via the Internet of Things (IoT), and control systems (Abioye et al. 2020). These efforts have established precision irrigation: data-driven irrigation strategies that modulate water delivery optimized in quantity and timing, adapted to the water requirements of crops across varying field conditions and growth stages to reduce water consumption (Abioye et al. 2020). Agricultural sensor networks should not only monitor environmental conditions but also directly assess plant physiological status, as environmental variables are only indirect predictors of plant stress. This concept has been described as the Internet of Plants (IoP) (Steeneken et al. 2023), where networks of sensors continuously monitor plant health and detect both biotic and abiotic stress. Similar to human healthmonitoring technologies such as wearable sensors, dedicated devices are emerging that directly measure plant physiological signals (Ataei Kachouei et al. 2023). Typical sensing modalities include sap flow sensors, stem diameter sensors, multispectral imaging, and electrophysiological recordings. To achieve field-scale coverage with high spatial resolution, these technologies adopt concepts from the IoT. They typically operate autonomously, communicate wirelessly, rely on low-power electronics, and may incorporate energy-harvesting modules such as solar power. Although IoP systems have
2
strong potential to provide actionable insights and to enable (semi-)autonomous crop management. However, there are open questions about which physiological parameters to monitor and which sensing technologies are most suitable. Plants, as sessile organisms, must continuously adapt to complex and changing environmental conditions (Johns et al. 2021). They rely on sophisticated sensing systems to perceive environmental cues and transmit locally detected information throughout the plant to coordinate systemic responses. A key class of internal signals is ion fluxes at the cellular level, including Ca2+ , K+ , Cl– , and H+ , which generate electrical signals that propagate through the plant. One example is the action potential, a short transient change in membrane potential. In species such as Dionaea muscipula (Venus flytrap) and Mimosa pudica, these signals trigger rapid leaf movements for prey capture or surface reduction (Fromm and Lautner 2007). Despite their significance, measuring and analyzing plant electrical signals under real-world constraints remains challenging. Recording is typically performed either invasively using metal electrodes (Li et al. 2021), which cause tissue damage, or non-invasively using surface electrodes (Meder et al. 2021), which are dependent on surface moisture and are affected by plant growth and mechanical movement (e.g., wind). Signal analysis outside controlled laboratory settings is further complicated by multiple stimuli that simultaneously trigger overlapping electrical signals (Steeneken et al. 2023; Li et al. 2021). Plants also differ in physiological state, morphology (e.g., number of leaves, height), and developmental stage, leading to inter-individual variability in recorded signals (Huber and Bauerle 2016). These factors require advanced analytical approaches, including machine learning, to reliably extract meaningful patterns from complex and heterogeneous data. Recent studies have explored the use of machine learning methods to analyze plant electrophysiological signals during abiotic stress like drought (Najdenovska et al. 2021; Tran et al. 2019; Zhou et al. 2025), salinity (Bhadra et al. 2023; Zhou et al. 2025), temperature (Aust et al. 2025; Buss et al. 2023), ozone (Bhadra et al. 2023; Aust et al. 2025), nutrient deficiencies (González I Juclà et al. 2023; Najdenovska et al. 2021) as well as biotic stresses like spider mites (Najdenovska et al. 2021) or caterpillars (Reissig et al. 2021). For instance, Najdenovska et al. (2021) recorded electrophysiological signals from 36 tomato plants exposed to drought, nutrient deficiencies, and spidermite infestation. Electrical activity was measured using PhytlSigns devices (Vivent SA) with electrodes inserted into the stem. From the signals, 34 statistical features were extracted across multiple temporal windows, and an XGBoost classifier achieved accuracies of up to 85% for a 1 min window. In this work, we extend their approach in two ways. First, we extract a much larger set of statistical descriptors using the tsfresh framework (Christ et al. 2018). Second, we use automated machine learning (AutoML) to optimize the preprocessing and classification pipeline instead of manually selecting a classifier. Furthermore, rather than completely stopping irrigation, we study graded irrigation regimes with varying water amounts to better reflect realistic agricultural conditions. González I Juclà et al. (2023) studied electrical activity in 16 tomato plants under nitrogen deficit in a greenhouse using PhytlSigns sensors (Vivent SA). Plants were grown in coconut fiber and monitored for 15 days, initially under a standard nutrient
3
solution (normal state) and later with nitrogen reduced to one-third (stressed state). Four deep learning architectures were evaluated on time series windows of 1 s to 30 s, with the encoder model achieving up to 99% accuracy. Classifier certainty over the full measurement period revealed a transition from the normal to the stressed state. Similarly, we aim to detect the transition from healthy to stressed plants using classifier certainty. In contrast, we evaluate model generalizability using separate training, validation, and test sets. Because we focus on drought stress, we include control plants maintained under optimal conditions to account for environmental fluctuations in semi-controlled greenhouses. To address potential bias in classifier certainty, we apply probability calibration. While both studies use PhytlSigns sensors, we additionally validate this sensing approach using our own system, PhytoNode. Aust et al. (Aust et al. 2025) also employed the PhytoNode to investigate elevated temperature and ozone levels based on the electrophysiological response of ivy (Hedera helix ) plants. They trained a fully convolutional network (FCN) on the recorded signals and deployed the trained model directly on the microcontroller-based PhytoNode. Validation accuracies of up to 86% and 88% were achieved for temperature and ozone stimuli, respectively. The network was transferred to the microcontroller using their custom toolchain, Mbed Torch Fusion OS, which enables direct deployment of PyTorch models onto the PhytoNode. Subsequently, they successfully detected increased temperature and ozone levels in online classification experiments. Discrete predictions, such as ‘healthy’ or ‘stressed’, alone are insufficient for practical decision-making. Prediction confidence is equally important, as incorrect actions directly affect resource use and economic outcomes. For example, a stress prediction with only 52% confidence may require human review or higher thresholds for automated interventions. Depending on the classifier and training procedure, predicted probabilities can be systematically miscalibrated (Guo et al. 2017). Neural networks trained with cross-entropy often become overconfident, while ensemble methods such as random forests may yield underconfident estimates. Guo et al. (Guo et al. 2017) evaluated six calibration methods, including isotonic regression, Platt scaling, temperature scaling, histogram binning, Bayesian binning, vector scaling, and matrix scaling for neural network classifiers and identified temperature scaling as a simple and effective approach. We contribute to the concept of the Internet of Plants through the deployment of our self-sustaining sensor node, called PhytoNode (Buss et al. 2025), that communicates wirelessly and harvests solar energy. Our sensor node measures plant electrophysiological signals, and we have validated it successfully in outdoor experimental setups (Buss et al. 2026b). In this work, we extend the application from environmental monitoring to water-stress detection for precision irrigation. We include the plant physiological state in the irrigation system and form a bio-hybrid system of plant, PhytoNode, and machine-learning-based analysis. An approach that can function as a decision-support tool for farmers.
4
Fig. 1: Experimental Setup. Left: 4 plants under saturated conditions and 4 plants receiving 400 ml (June 8). Middle: All 16 plants. Right: Our PhytoNode attached to a tomato plant in a greenhouse with a close-up of the silver electrodes.
2 Materials and Methods We measured the electrical potential of tomato plants using our low-power, selfsustaining sensor node, called PhytoNode (see Fig. 1, right) (Buss et al. 2025). It is enclosed in a weather-resistant housing and includes a 4,800 mAh LiPo battery that is recharged by a solar panel that delivers up to 5 W under greenhouse lighting conditions. Each PhytoNode can simultaneously record electrical signals from two plants and transmit the data via Bluetooth Low Energy (BLE) to a data sink, a Raspberry Pi 4 in this study. Electric differential potentials (EDP) per plant were acquired using two silver-coated electrodes inserted into the stem, one near the base and the other at least 30 cm higher along the stem. The signals were sampled at 10 Hz. We recorded the electrical differential potential of 16 tomato plants continuously over a period of 18 days, from June 4th to June 22nd, 2025 (see Fig. 1, left and middle). During the initial phase (June 4th to 8th), all plants were irrigated daily with 400 mL of water, as this amount resulted in a dry or minimally moist soil surface the following day. After this period, the plants were divided into four groups, each subjected to a distinct irrigation regime for the remainder of the experiment duration: four plants were placed in a bucket with constant water access (saturated condition), four plants continued receiving 400 mL per day (ideal condition), four plants received 200 mL per day (moderate water deficit), and the remaining four plants received 100 mL per day (strong water deficit). In addition to the electrical potential, soil moisture was monitored at 0.1 Hz using standard capacitive soil moisture sensors. All recorded and processed data are available online (Buss et al. 2026a).
3 Machine Learning Analysis We employ two distinct machine learning pipelines: (1) a feature-based AutoML pipeline (Sec. 3.1) and (2) an end-to-end deep learning pipeline with hyperparameter optimization (Sec. 3.2). The former comprises (1a) preprocessing of raw measurements, time-window slicing and class label assignment, (1b) statistical feature extraction for each time window, (1c) AutoML to identify suitable classifier pipelines, and 5
(1d) sequential backward selection to determine a minimal feature subset. The second pipeline includes (2a) the same preprocessing step as in (1a), extended by robust z-score normalization, (2b) hyperparameter optimization of three neural network architectures, and (2c) extended training of the best-performing configurations. Finally, we apply class certainty calibration (Sec. 3.3) to the resulting classifiers to mitigate internal biases and analyze the temporal characteristics of the predicted confidence scores on previously unseen data (Sec. 4.4). (1a & 2a) Preprocessing, Time Window Slicing and Label Assignment. We resampled the electrical potential to 1 Hz, as this smoothed the data without significant distortion. Subsequently, all time series were sliced into fixed time windows of 1 min, 5 min, 30 min, 1 h, and 6 h. The dataset was labeled using two approaches, following a strategy similar to that described in (González I Juclà et al. 2023). In both cases, measurements from the initial days were assigned to the healthy class, whereas measurements from the final days were labeled according to the specific classification scheme adopted in each approach. In the first approach, a binary classification task (healthy vs. stressed) was defined. The final days of the treatment groups overwatered, 200 ml, and 100 ml were labeled as ‘stressed’, irrespective of the specific irrigation regime, while the initial days across these groups were labeled as ‘healthy’. In the second approach, a multiclass classification framework was employed to distinguish between overwatered, underwatered, and healthy conditions. Measurements from the final days of the overwatered group were assigned to the ‘overwatered’ class, while measurements from the final days of the four plants in the 100 ml group were labeled as ‘underwatered’. The initial days of both groups were used to define the ‘healthy’ class. For both approaches, intermediate days were not used for training but instead served to evaluate the temporal behavior of the classifier, specifically to identify the point at which predictions transition from healthy to stressed conditions. Our analysis primarily focuses on the binary classification approach, since the multiclass setting exhibited strong validation performance but limited generalization to unseen test data (see Sec. 4.2 for details).
3.1 Feature-Based Modeling and AutoML (1b) Feature Extraction. For each time window, we extracted approximately 700 time series characteristics, also called features (see Tab. 2) using the Python library tsfresh (Christ et al. 2018). Feature values resulting in numerical overflows (±inf) were replaced by the corresponding feature mean. Subsequently, features with variance below 0.01 were removed, as such features provide limited discriminative power. For numerical stability and feature comparability, features are normalized using min–max xi −min(x) scaling given by xi, norm = max(x)−min(x) where xi represents the current sample, and x the entire set of one feature. In addition to the four plants that received 400 ml over the entire period, we excluded two further plants (one each from the overwatered and 100 ml groups) from the training process and used them as test plants. The remaining data was split into training and validation sets using an 80/20 ratio. The split was stratified to preserve the class and plant data distributions across both subsets. 6
(1c) AutoML. We identified a potential classification pipeline using the AutoML framework Naive AutoML (Mohr and Wever 2023). AutoML aims to identify suitable preprocessing and learning algorithms that achieve strong generalization performance on a given dataset. This involves both selecting an appropriate pipeline and determining an optimal set of hyperparameters. First, it searches for a pipeline composed of preprocessing steps and a single predictor implemented in the Python framework sklearn. During this stage, candidate pipelines are evaluated using their default hyperparameter settings. In a second optimization phase, the best-performing pipeline is selected, and its hyperparameters are further optimized via random search. The pipeline search was done using all computed features, allowing up to 100 hyperparameter optimization iterations, and we used accuracy as evaluation metric. (1d) Feature Selection with Sequential Backward Selection. Feature selection was performed to identify a compact subset of discriminative features while avoiding unnecessary dimensionality, redundancy, and noise that may degrade classification performance. We applied a staged approach by first using the Mutual information (MI) as a univariate preselection criterion to reduce the dimensionality of the high-dimensional feature space generated by tsfresh (Vergara and Estévez 2014). MI quantifies the statistical dependency between an individual feature and the class labels and was computed for all extracted features. The features were ranked in descending order of MI, and the top 200 features were retained as candidate features. The second stage includes multivariate feature selection using sequential backward selection (SBS) as implemented in the mlxtend Python framework (Raschka 2018). SBS was applied to the candidate feature set to identify feature subsets with strong joint discriminative power. The algorithm was initialized with the full set of candidate features and iteratively removed one feature at a time. At each iteration, all possible feature subsets obtained by excluding a single feature were evaluated, and the subset yielding the highest classification performance was retained. This process was repeated until only one feature remained. Classification performance during feature selection was assessed using accuracy, which is appropriate given the balanced datasets. Furthermore, performance is assessed using five Group K-Fold splits, in which the classifier is trained five times, and in each split, two plants are excluded from training and used for validation.
3.2 Deep Learning Models and Hyperparameter Optimization (2a) Robust Z-Score Normalization. Compared to classical machine learning approaches, deep learning models are capable of learning feature representations directly from raw measurements. To provide robust normalization that is resilient to outliers, we apply a commonly used robust z-score transformation based on the median and interquartile range (IQR). The normalized value is defined as z = x−median(x) IQR(x) where z denotes the normalized time series, x the raw observation, median(x) the median of the time series, and IQR = Q3 − Q1 represents the interquartile range, i.e., the spread of the central 50% of the data. This transformation normalizes electrical potentials for comparability across plants and experiments, reducing physiological and experimental variations. Simultaneously, it provides a robust alternative to standard 7
z-score normalization by reducing sensitivity to outliers caused by transient artifacts, such as mechanical perturbations during plant handling or sensor-related noise. (2b) Hyperparameter Search with Optuna. Hyperparameter search is a labor intensive task in machine learning. This section describes the used DL architectures in combination with Optuna (Akiba et al. 2019). Optuna is an open-source framework for hyperparameter optimization that uses sampling algorithms to identify optimal hyperparameter configurations within a predefined search space. We explore this space across 100 trials using the Tree-structured Parzen Estimator (TPE) sampling strategy. The HyperbandPruner enables early termination of unpromising trials by employing multiple successive halving procedures. Trials are allocated to different brackets: early brackets evaluate a large number of hyperparameter configurations with a limited training budget (exploration), whereas later brackets focus on fewer configurations with increased training resources (exploitation). Trough recurrent comparisons in a bracket only one third survives and will not get pruned. Each model configuration is trained for up to 100 epochs with early stopping after 15 epochs without improvement of the validation loss. We define the hyperparameter search space of each model in Table 1. Next we are describing our fundamental architectures, which are the starting point for the hyperparameter optimization.
Convolutional Neural Network (CNN). Our first and simplest model is a standard CNN. It serves as our baseline deep learning classifier as it has already been used to classify electrical plant signals (González I Juclà et al. 2023). Our architecture is composed of sequential convolutional blocks, each consisting of a 1D convolutional layer, batch normalization, and a ReLU activation function. We add a max pooling and dropout layer after each second convolutional block to avoid excessive dropouts and shrinking of the feature space. The resulting feature maps are aggregated using an 1D adaptive average pooling layer to make the architecture independent of the input length followed by a flatten operation. This produces a one-dimensional feature vector that is passed to a fully connected artificial neural network (ANN) consisting of one or several (optimized by Optuna) linear layers with ReLU activations. With each successive layer, the dimensionality of the feature vector is reduced by half, except for the final layer, which outputs a vector whose length corresponds to the number of classes in the classification task. We search for general learning hyperparameters (learning rate, weight decay, batch size) as well as model specific hyperparameters such as number of convolutional blocks, kernel size, channel dimension, dropout rate and the number of linear layers (see Table 1) using Optuna as a optimization framework. InceptionTime. The second model, InceptionTime, is an ensemble of CNN models (Ismail Fawaz et al. 2020). We adapt the PyTorch implementation provided by Campos et al. (2023) (Campos et al. 2023). The receptive field of each CNN model differs and is defined by its kernel size which enables the network to extract temporal features at different scales. This property makes it well suited for classifying electrical signals that exhibit both short-term dynamics (e.g., ion fluxes) and long-term patterns (e.g.,
8
stress accumulation or circadian rhythms). Our architecture is composed of multiple sequential Inception modules organized into residual blocks. Each module shares the same structural design but uses different initial weights and applies convolutional filters of multiple lengths in parallel to capture both short-term and long-term temporal patterns. The module begins with a bottleneck layer that is a learned channel-mixing layer to reduce the input dimensionality to decrease computational cost. The bottleneck output is then passed to several parallel convolutional layers with different kernel sizes. An additional parallel branch performs max pooling on the raw input, followed by another bottleneck layer, to provide robustness to small perturbations. All parallel convolutions are first concatenated, then batch-normalized, and activated using a ReLU nonlinearity. These Inception modules are arranged sequentially within residual blocks, with skip connections linking the input and output of each block. A global average polling layer complete sthe architecture. Using Optuna, we perform hyperparameter optimization over both, general learning parameters and architectural choices. This includes the number of residual blocks, the number of Inception modules per block, the number of parallel convolutions per module, kernel sizes, bottleneck channel counts, and the number of output channels in each module (Table 1). Importantly, we only optimize the output channel size of the first block, as it is doubled in each subsequent block. Similarly, we specify the maximum kernel size, which is reduced by 20% for each further parallel convolution.
Mamba. Our next architecture is the selective State Space Model (SSM) known as Mamba (Gu and Dao 2024). This sequential model integrates principles from recurrent neural networks, convolutional neural networks, and classical continuous-time state space models. A standard state space system is expressed as h′ (t) = Ah(t) + Bx(t), y (t) = Ch(t) + Dx(t) a formulation widely used in control theory (Gu et al. 2021). Here, h(t) denotes the system’s hidden state, x(t) the current input, and y (t) the output. The transition matrix A control the change of the hidden state, while the input matrix B determines how new input signals influence the updated state h′ . Likewise, the output matrices C and D map the internal state and the input, respectively, to the output y (t), with D functioning analogously to a skip connection. We adopt this modeling principle because prior work has demonstrated that SSM architectures can capture long-range temporal dependencies spanning 10,000 time steps or more, while maintaining linear computational complexity with respect to input length (Gu et al. 2021). In contrast, self-attention based architectures typically struggle at such temporal horizons and exhibit quadratic scaling. This property makes SSMs particularly well suited for electrophysiological plant measurements, which are frequently sampled at high rates often up to 500 Hz. Our model is a compact sequence encoder built around Mamba blocks. The network begins with a linear projection layer that maps the raw input to the model’s internal dimensionality. The projected sequence is then passed through one or more identical processing layers (mamba layer). Each layer applies layer normalization, followed by a Mamba block as described by Gu et al. (2024) (Gu and Dao 2024), then a dropout layer. The output of this layer is combined with the original layer input through a residual connection.
9
After the final mamba layer, the sequence of hidden states is condensed into a single vector using mean pooling across the temporal dimension. This pooled representation provides a global summary of the entire input sequence. A final linear layer maps this representation to the model’s output dimension which is the number of classes in our classification. We use Optuna for automated hyperparameter optimization that covers again general learning parameters and structural design choices. The Mamba related parameters include the dimensionality of the input projection, the hidden size of the model, the kernel size of the initial convolution, and the expansion factor used in the block’s gating mechanism. Beyond the Mamba block itself, we search only for the number of stacked Mamba layers (see Table 1 for details).
Table 1: Hyperparameter search space for different model architectures optimized with Optuna (log-scaled where indicated). Model
Hyperparameter
Range
General
Learning Rate Weight Decay Batch Size
[10−4 , 10−2 ]log [10−6 , 10−3 ]log 2k , k ∈ {4, 5, 6, 7}
CNN
# Conv. Blocks Kernel Size Channel Dimension Dropout # Linear Layers
2k, k ∈ {1, . . . , 10} 2k + 1, k ∈ {2, . . . , 7} 2k , k ∈ {1, . . . , 10} [0, 0.5] {1, . . . , 6}
InceptionTime
# Residual Blocks # Inception Modules # Parallel Convolutions Output Channels Kernel Size
{2, . . . , 6} {1, 2, 3, 4} {2, . . . , 6} 2k , k ∈ {3, 4, 5, 6} 2k + 1, k ∈ {2, . . . , 7}
Mamba
# Mamba Layers Input Dimension Hidden Size Kernel Size Expansion
{1, . . . , 8} 2k , k ∈ {4, . . . , 9} 2k , k ∈ {3, 4, 5, 6} {2, 3, 4} {1, . . . , 5}
3.3 Confidence Calibration with Temperature Scaling Temperature scaling rescales the classifier’s logits zi by a scalar temperature T , producing zi /T , before applying a sigmoid or softmax to obtain calibrated probabilities. Temperature scaling is argmax-invariant and does not change classification accuracy as T is optimized after training and applied uniformly to all logits. The effect of calibration is visualized using reliability diagrams, which compare empirical class frequencies (based on true labels) with predicted certainties. These diagrams
10
100 electrophysiology
20
soil moisture 90
0
−20 −40
80
t1
t2
t3
70 60
soil moisture [%]
EDP [mV]
40
06-05 06-07 06-09 06-11 06-13 06-15 06-17 06-19 06-21 date [MM-DD]
Fig. 2: Exemplary EDP (blue) and soil moisture (red) of a plant irrigated with 200 mL after day 4 (yellow dashed). Intervals t1 , t2 , t3 are used for training, validation, and testing. Gray–white shading indicates 24-h cycles.
are typically combined with histograms showing the distribution of prediction confidences (Guo et al. 2017). Calibration performance is further quantified using two commonly used metrics for binary classification: the Brier score and the Adaptive Calibration Error (ACE) (Glenn et al. 1950; Nixon et al. 2019). The Brier score is defined as N 1 X 2 (yi − p̂i ) , (1) Brier = N i=1 the mean squared error between the predicted probabilities p̂i and the corresponding true labels yi across all samples N . While the Brier score measures how closely predicted probabilities match the true labels on a per-sample basis, ACE quantifies how well predicted confidences match the empirical label frequencies. To compute ACE, we first bin the predicted confidences using quantile binning with M = 20, so that each bin m contains approximately 5% of the samples. For each bin, the empirical accuracy acc(m) is computed as the fraction of the true labels of the samples assigned to that bin. In addition, the mean predicted confidence conf(m) is obtained by averaging the predicted certainties of the samples within the same bin. The ACE is then calculated as M 1 X ACE = | acc(m) − conf(m)| , (2) M m=1 for the average absolute difference between acc(m) and conf(m) across all bins.
4 Results and Discussion We illustrate the classification task using an exemplary measurement of a tomato plant that received 200 mL of water after the fourth day (see Fig. 2). As described in Sec. 2, we use the first three days (interval t1 ) and label them as ‘healthy’, and we use the last three days (interval t3 ) and label them as ‘stressed’. These data are used to train, validate, and test the classifier. The interval t2 in between is analyzed to assess classification certainty over the experiment duration on unseen data and to identify the transition from ‘healthy’ to ‘stressed’. We analyze multiple classifier input 11
intervals as shorter time windows enable faster decisions but may contain insufficient information for reliable predictions (see Tab. 2). Naturally, the length of the measured time series increases from 60 data points at a window length of 1 min to 21,600 data points at 6 h intervals. Conversely, in the training set, the number of available time series decreases from 69,112 at 1 min windows to 195 at 6 h windows. The number of extracted features per window length remains approximately constant at around 670 features, except for the 1 min interval, which retains only 421 features per time series. This reduction results from the shorter time-series length and the preprocessing step that removes features with variance below 0.01. 1 min
data set l
s
5 min f
l
s
30 min f
l
s
1h f
l
s
6h f
l
s
f
training 60 69,112 421 300 13,822 685 1,800 2,302 669 3,600 1,158 664 21,600 195 653 validation 60 17,278 421 300 3,458 685 1,800 578 669 3,600 282 664 21,600 45 653 test 60 17,278 421 300 3,456 685 1,800 576 669 3,600 288 664 21,600 48 653
Table 2: All data sets of different timeseries length l, data set size s, and number of features f based on varying look-back horizons.
4.1 Automated Machine Learning and Feature Selection Since identifying an appropriate classification pipeline, including preprocessing steps, a classifier, and a well-tuned set of hyperparameters, requires substantial effort, we employed NaiveAutoML on the datasets listed in Table 2. NaiveAutoML selected Histogram Gradient Boosting (HGB) as the best-performing classifier for all window lengths except for the 1 min interval. The Extra Trees Classifier (ETC) and Random Forest Classifier were excluded from this window because they yielded only a marginal performance gain of less than 1% and would require a different probability calibration approach, which compromises methodological consistency and comparability across time windows. HGB is an ensemble method that sequentially builds decision trees, where each tree is trained to correct the residual errors of the previous ensemble. Unlike classical gradient boosting, HGB discretizes continuous features into histograms prior to training, improving computational efficiency. Furthermore, the NaiveAutoML pipeline applied an additional variance-threshold preprocessing step for the 1 h window lengths to remove features with zero variance. Although features with variance below 0.1 had already been excluded during prior preprocessing, NaiveAutoML may re-identify constant features due to its internal K-fold cross-validation procedure (Mohr and Wever 2023). Within individual folds, feature variance can decrease to zero, leading to their removal. All investigated pipelines achieved 100% training accuracy, indicating that the models have sufficient capacity to fully fit the training data (Table 3). Validation accuracy decreases with increasing look-back horizon, whereas test accuracy increases. For short windows, samples are temporally closer and therefore more similar, which can inflate validation accuracy without reflecting true 12
horizon
pipeline
1 min 5 min 30 min 1h 6h
HGB HGB HGB VT + HGB HGB
AutoML acc. [%] train val. test 100 100 100 100 100
92.2 92.6 91.0 90.1 77.8
62.6 76.3 83.2 84.0 89.6
# features 96 193 182 200 15
SBS acc. [%] train val. test 100 100 100 100 100
89.6 92.1 91.4 90.1 82.2
61.6 75.6 82.5 82.3 87.5
Table 3: AutoML-determined optimal classification performance using all features (AutoML acc.) and classification performance after sequential backward selection (SBS acc.).
generalization. Larger windows capture more temporal variability, making validation more challenging but improving generalization on the test set. Our feature selection pipeline first ranks features by MI and then identifies an optimal subset using SBS. Initially, the top 200 features were considered. For the 1 min horizon, this was reduced to 100 due to the large training set (69,112 samples) and the combinatorial complexity of SBS, reducing the required model fits from 20,099 to 5,049. Despite this reduction, validation accuracy decreased by only 2.6% using 96 selected features (see Tab. 3). For the 5 min, 30 min, and 1 h horizons, SBS selects 193, 182, and 200 features, respectively. In all cases, training, validation, and test accuracies remain nearly identical to those obtained with the full feature set, with the largest decreases of 0.5% in validation accuracy (5 min) and 1.7% in test accuracy (1 h). Thus, up to 479 features (5 min) can be removed without meaningful performance loss. For the 6 h horizon, the highest validation accuracy (82.2%) is obtained using only 15 features, with a corresponding test accuracy of 87.5% (highest over all horizons). However, due to the smaller dataset size at this horizon, classifiers are more sensitive to individual samples and should therefore be interpreted cautiously.
4.2 Optuna and Deep Learning Similarly to NaiveAutoML, we employ Optuna to identify suitable hyperparameter configurations for the selected models (CNN, InceptionTime, and Mamba). The full optimization procedure is documented and publicly available (Buss et al. 2026a). To avoid partially trained models, the best configurations were retrained with an increased number of epochs (200) to ensure full convergence. We conducted multiple independent runs with randomly selected seeds (42, 123, 236, 679, 999) for dataset and model initialization. The resulting mean accuracies and standard deviations are reported in Table 4. No single model consistently outperforms the others across all horizons. Performance is horizon-dependent: at 1 min, CNN achieves the best result (63.56% ± 1.16%); at 5 min, InceptionTime achieves the best result (70.99% ± 1.62%); at 30 min, Mamba performs best (80.21% ± 3.91%); and at 1 h and 6 h, CNN achieves the highest accuracies (84.77% ± 2.86% and 96.96% ± 1.06%, respectively). Overall, CNN provides the best average performance across horizons (77.57%) and exhibits comparatively
13
CNN
Horizon 1 min 5 min 30 min 1h 6h
InceptionTime
train [%]
val. [%]
test [%]
train [%]
val. [%]
test [%]
79.45 ± 1.55 80.49 ± 1.05 84.31 ± 1.04 77.05 ± 1.86 83.15 ± 1.03
71.02 ± 0.81 76.07 ± 0.63 76.93 ± 0.75 75.86 ± 0.56 70.00 ± 2.54
63.56 ± 1.16 68.97 ± 1.97 73.58 ± 2.83 84.77 ± 2.86 96.96 ± 1.06
74.82 ± 2.79 81.07 ± 2.11 89.00 ± 1.31 89.21 ± 4.78 80.26 ± 2.86
70.81 ± 0.53 76.84 ± 0.60 82.58 ± 0.28 80.56 ± 2.23 70.87 ± 6.24
62.25 ± 1.25 70.99 ± 1.62 71.06 ± 2.38 77.26 ± 5.51 88.26 ± 5.07
Mamba
Horizon train [%] 1 min 5 min 30 min 1h 6h
72.66 ± 0.47 83.62 ± 3.27 81.52 ± 1.99 89.02 ± 3.66 67.07 ± 7.80
val. [%]
test [%]
70.55 ± 0.15 62.29 ± 0.79 76.19 ± 0.72 68.84 ± 2.01 78.12 ± 1.29 80.21 ± 3.91 81.33 ± 1.84 83.79 ± 2.96 63.91 ± 6.82 58.26 ± 11.94
Table 4: DL accuracies for the training, validation, and test sets across all look-back horizons, computed over five independent runs with different random seeds.
high robustness, with a mean standard deviation of 1.98% across runs. InceptionTime is competitive (73.96%) but shows higher variability (3.17%), while Mamba performs worse on average (70.68%) and demonstrates the highest variability (4.32%), particularly at longer horizons. DL models generally perform worse compared to the HGB baseline. For the training accuracy, HGB consistently achieves 100%, whereas DL models show lower performance, with gaps ranging from about -11% (e.g., InceptionTime at 30 min) to -33% (Mamba at 6 h). On the validation set, these differences decrease, ranging from -9% (Mamba at 1 h) to -19% (Mamba at 6 h). On the test set, the performance gap narrows further and becomes strongly dependent on the prediction horizon. On average, DL models achieve comparable performance to HGB, with a slight mean improvement of approximately +1% ± 1%, primarily driven by CNN (+2%). At short horizons (5 min and 30 min), DL models consistently underperform, with decreases of −6% ± 1% and −8% ± 5%, respectively. At 1 h, CNN and Mamba outperform HGB by about +2% and +1%, respectively. At 6 h, CNN shows the largest improvement (+9%), followed by InceptionTime (+1%), whereas Mamba performs substantially worse with a difference of -29%. Accordingly, despite partially competitive performance on the test set, we do not further pursue DL approaches for the binary classification task, as HGB provides more consistent and reliable performance across all horizons.
Multiclass Classification. Both the AutoML and DL approaches were evaluated on a multiclass classification task involving the classes ‘healthy’, ‘overwatered’, and ‘underwatered’ as described in Sec. 3. While the models achieved high performance on the training and validation sets, they consistently failed to generalize to the test data. For example, with this 30-min interval, the HGB model reached accuracies of 100%, 95%, and 48% on the
14
horizon
accuracy [%]
T
NLL uncal. cal.
Brier [%] uncal. cal.
ACE [%] uncal. cal.
1 min 5 min 30 min 1h 6h
89.6 92.1 91.3 90.1 82.2
1.58 3.54 2.97 1.81 1.77
0.2680 0.3991 0.3929 0.3130 0.4590
7.80 6.93 7.19 7.61 14.00
4.09 6.07 5.92 5.15 14.84
0.2416 0.1968 0.2261 0.2599 0.3958
7.47 5.88 6.51 7.43 12.91
0.60 0.35 0.86 2.93 8.19
Table 5: Certainty calibration for each interval, including validation accuracy, temperature scaling parameter T , and the uncalibrated and calibrated negative log-loss (NLL), Brier score, and adaptive calibration error (ACE)
training, validation, and test sets, respectively. Similarly, the CNN achieved 87%, 84%, and 48% across the same splits. These results were consistent across different intervals and model types. Overall, the models were not able to reliably distinguish between specific stress types, but rather only between the classes ‘healthy’ and ‘stressed’ plants.
4.3 Impact of Certainty Calibration Certainty calibration adjusts the classifier’s predicted probabilities to better match empirical outcomes, as classifier pipelines can produce under- or overconfident predictions. We therefore apply temperature scaling to the selected model and reduced feature set for all time windows. As expected, classification accuracy remains unchanged because temperature scaling rescales predicted probabilities without altering class labels (Table 5). For all horizons we obtain T > 1, indicating systematic overconfidence. Calibration improves all probabilistic metrics, including negative log-likelihood (NLL), Brier score, and adaptive calibration error (ACE), with the strongest improvement observed for ACE, which directly measures the mismatch between predicted confidence and empirical frequency. For the 1, 5, and 30 min horizons, calibration reduces ACE to below 1%, indicating improved probabilistic reliability. In contrast, the 1 h and 6 h horizons retain higher ACE values (2.93% and 8.13%), which may result from residual miscalibration not captured by a single global temperature parameter or from increased variance due to the smaller sample size. To illustrate the overconfident behavior of the classifier and the effect of certainty calibration, we show the reliability diagram and the distribution of predicted probabilities for the 1 min horizon (as a representative example) in Fig. 3. Most predictions are concentrated at the extreme probabilities of 0 and 1 (see bins of highest frequency in Fig. 3, right). Temperature scaling shifts probability mass from the extremes toward intermediate confidence levels. For example, 37% of the samples (6,377 observations) fall into the highest probability bin (100%) before calibration, whereas this proportion decreases to 30% (5,249 observations) after calibration. The reliability curve is constructed using quantile binning. The calibration shifts the curve closer to the ideal diagonal, indicating improved agreement between predicted probabilities and empirical frequencies.
15
predicted probability [%]
observed frequency [%]
1
0.5
0
ideal uncalib. calib.
1 0.8 0.6 0.4
uncalib. calib.
0.2 0 −0.4 −0.2
0 0.5 1 predicted probability [%]
0
0.2
0.4
density [%] (before ← — → after)
Fig. 3: (Un)calibrated reliability diagram and histogram data distribution.
4.4 Temporal Detection of Class Transitions After training the classifiers, the selected models were applied to the full recording period covering the three intervals t1 , t2 , and t3 (Fig. 2). Fig. 4 (top) shows the classification results for the example of a 5-min window of a tomato plant receiving 200 mL of water (same data as in Fig. 2). For now, we focus on the purple dots indicating stress certainty in percentages, that is, the classifier’s output. Using a 50% decision threshold, probabilities below 50% are assigned to the healthy class and values above 50% to the stressed class. The certainties indicate a daily pattern: values increase during daytime, shifting toward the stressed class, and decrease at night toward the healthy state. While early measurements are dominated by low certainty, the proportion accumulating near 100% increases over time, indicating a progressive transition toward the stress state. To visualize the dynamic trend, Locally Weighted Scatterplot Smoothing (LOWESS) was applied to the classification certainties. LOWESS is a non-parametric regression method that fits local weighted linear models within a defined neighborhood of data points. To avoid bias from training data, LOWESS was applied only to the central interval t2 . For the upper panel in Fig. 4, a neighborhood size of 50% of the data was chosen to capture the overall progression of stress development. The resulting smoothed curve crosses the 50% decision threshold approximately four days after irrigation reduction. For the lower panel, LOWESS was first applied to each plant individually, using a smaller neighborhood size of 4% (corresponding to approximately 12-hour time windows) to resolve daily fluctuations. The four lines shown in Fig. 4 are averages of these plant-level LOWESS curves for each treatment group. We observe 24-hour periodicity across all groups, with maxima around midday and minima at night. All reduced-irrigation treatments (100 mL, 200 mL) and the overwatered group show a gradual increase in stress certainty, whereas the control group (400 mL, green line) remains consistently below the decision threshold. The 100 mL and 200 mL groups cross the threshold on day four, whereas the overwatered group 16
100 80 60 40 20 0
100 ml
200 ml
400 ml
LOWESS
overwatered
100 80 60 40 20 0
stress certainty [%]
EDP [mV] stress certainty [%]
EDP stress certainty 40 20 0 −20 −40
06-05 06-07 06-09 06-11 06-13 06-15 06-17 06-19 06-21 date [MM-DD]
Fig. 4: Classification results for the intermediate time interval t2 (June 7 to June 19, green dashed) lines. Top: Classification outcomes at 5-minute intervals (purple dots) with a 50% LOWESS-smoothed trend of the certainty values. Bottom: Averaged LOWESS trends (4% smoothing) for each subgroup.
reaches it approximately two days later and has generally lower values. The reduced daily amplitude in the control group may reflect more stable physiological dynamics under non-stressed conditions. Given the small group size (n=4), these comparisons should be interpreted cautiously due to limited statistical power.
Precision-Recall. We analyze the precision–recall (PR) curve for two reasons: (1) to compare classification performance using the area under the PR curve (AUPRC), and (2) to adjust classifier behavior by adapting the decision threshold according to farming priorities. tp tp The precision is defined as tp+f p and the recall as tp+f n with tp, f p and f n being true positives, false positives, and false negatives, respectively. With respect to (1), unlike single-threshold metrics, AUPRC summarizes overall ranking performance across all decision thresholds, providing a threshold-independent measure of class separability. Validation performance remains high (0.91 to 0.976), while test performance increases with longer time windows (Table 6). The 1 min window yields the lowest test AUPRC (0.581), while the 6 h window achieves the highest (0.937). With respect to (2), the pipeline can function either as a decision-support tool or as a direct control signal for irrigation. Since priorities may shift (e.g., water conservation vs. growth maximization), the decision threshold can be adjusted using the precision–recall curve (Fig. 5). Treating stressed as the positive class, recall measures the proportion of stressed plants correctly identified, while precision gives the share
17
data set
1 min
5 min
30 min
1h
6h
validation test
0.962 0.581
0.976 0.728
0.971 0.871
0.948 0.871
0.910 0.937
Table 6: Area under the precision–recall curve (AUPRC), all time windows.
of true positive stress alarms. Typically, we face a tradeoff here. For example, high precision with low recall reduces false positives (false alarms), but increase false negatives (missed stressed plants), whereas high recall increases detection at the cost of more false positives. At a 50% decision threshold, precision and recall are both 91% (red cross, Fig. 5). Increasing recall to 95% (green cross) to identify a greater proportion of stressed plants reduces precision to 87% and increases false positives. This adjustment corresponds to a certainty threshold of 32.1%. In summary, the PR curve enables systematic tuning of the classifier to application-specific requirements. 1min
5min
30min
1h
0.5 recall [%]
1
0
6h
train OP
sel. OP
0.5 recall [%]
1
precision [%]
1
0.8
0.6 0
Fig. 5: Precision–recall curves of the best classifier of each window applied to the validation (left), and test (right) datasets. The red cross marks the performance of the 30 min classifier at a threshold of 50%, while the green cross represents the same classifier at a recall of 95% (threshold of 32%) to capture more stressed plants at the expense of precision.
5 Discussion and Conclusion The objective of this work is to develop a decision-support system for irrigation management based on real-time crop physiological measurements. Electrophysiological responses from 16 tomato plants were recorded under different irrigation regimes and analyzed using DL combined with hyperparameter optimization and a machine 18
learning pipeline comprising AutoML and feature selection. Lastly, post hoc certainty calibration was conducted to remove internal classifier biases. While deep learning models (especially CNNs) slightly outperform HGB on test data at short (1 min) and long horizons (1 h and 6 h), we chose HGB for further analysis because it provides more consistent performance across all horizons. In addition, its decision-making process can be traced through the individual decision trees, making the classification results more interpretable. A key challenge was selecting an input time horizon that balances performance and decision latency. Although the 5 min horizon achieved the highest validation accuracy (92.1%), it performed worse at the test data with only 76% accuracy compared to 82% for 30 min window and showed a 0.143 lower test AUPRC (0.871 for 30 min) with a larger validation–test discrepancy of the PR-curves. This indicates reduced reliability, particularly for adaptive thresholding. The 30 min horizon provides a favorable tradeoff between responsiveness and false negative/positive rates and is recommended for implementation. Classifier-internal certainty biases can hinder reliable identification of the transition from healthy to stressed states. Temperature scaling corrected the overconfidence of histogram gradient boosting, achieving near-ideal calibration with adaptive calibration errors below 1%. A transition from healthy to stressed states occurred after 4 days in the 100 mL and 200 mL groups and after 6 days in the overwatered group, whereas the 400 mL control group remained stable. This indicates that the electrophysiological changes were driven by irrigation treatments rather than environmental variation. In addition, by excluding two plants from training and testing them independently confirmed generalization to unseen individuals. This suggests that the detected patterns are stimulus-specific rather than individual-specific, despite inherent electrophysiological variability among individual plants. These findings should be interpreted in light of several limitations. The study involved 16 tomato plants under controlled conditions, limiting generalizability to other cultivars, species, seasons, and field environments. Future work will include larger populations and greater environmental variability across seasons. Moreover, only specific irrigation regimes were investigated; additional studies will address other abiotic and biotic stresses, including nutrient deficiencies, diseases, and pests, to broaden applicability. The default transition point, defined by a 50% certainty threshold, does not represent definitive physiological stress onset and requires biological validation. This work proposes a data-driven method for detecting the transition from healthy to visibly stressed states. Future studies will integrate plant science expertise and adaptive decision thresholds (see Fig. 5) to improve physiological relevance. Another possible approach is to deploy the classifier on independent plant groups using distinct class-specific decision thresholds. Irrigation would be regulated based on model outputs, and performance would be evaluated using independent measures such as multispectral imaging or biomass accumulation. In summary, we presented a machine learning–guided framework to support irrigation decision-making. Our proposed system can serve as the basis for a biofeedback loop that enables automated irrigation control directly based on plant physiological responses. Such an approach has the potential to improve resource-use efficiency and enhance sustainability in crop production systems.
19
Declarations Conflict of interest The authors declare that they have no conflict of interest.
Data availability Data is available online (Buss et al. 2026a).
References Abioye EA, Abidin MSZ, Mahmud MSA, et al (2020) A review on monitoring and advanced control strategies for precision irrigation. Computers and Electronics in Agriculture 173:105441 Akiba T, Sano S, Yanase T, et al (2019) Optuna: A next-generation hyperparameter optimization framework. In: The 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp 2623–2631 Ataei Kachouei M, Kaushik A, Ali MA (2023) Internet of things-enabled food and plant sensors to empower sustainability. Advanced Intelligent Systems 5(12):2300321 Aust T, Heck CK, Buss E, et al (2025) Embedded deep learning for bio-hybrid plant sensors to detect increased heat and ozone levels. In: 2025 IEEE SENSORS, IEEE Bhadra N, Chatterjee SK, Das S (2023) Multiclass classification of environmental chemical stimuli from unbalanced plant electrophysiological data. PLoS One 18(5):e0285321 Buss E, Aust T, Wahby M, et al (2023) Stimulus classification with electrical potential and impedance of living plants: comparing discriminant analysis and deep-learning methods. Bioinspiration & biomimetics 18(2):025003 Buss E, Aust T, Hamburger O, et al (2025) Phytonode upgraded: Energy-efficient long-term environmental monitoring using phytosensing. In: Future of Information and Communication Conference, Springer, pp 119–138 Buss E, Aust T, Hamann H (2026a) Early detection of water stress by plant electrophysiology: Machine learning for irrigation management. URL https://doi.org/10. 5281/zenodo.18873964 Buss E, Aust T, Hamann H (2026b) When plants respond: Electrophysiology and machine learning for green monitoring systems. In: Conference on Biomimetic and Biohybrid Systems, Springer, pp 249–261 Campos D, Zhang M, Yang B, et al (2023) Lightts: Lightweight time series classification with adaptive ensemble distillation. Proc ACM Manag Data 1(2):171:1–171:27. 20
https://doi.org/10.1145/3589316, URL https://doi.org/10.1145/3589316 Christ M, Braun N, Neuffer J, et al (2018) Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests (tsfresh–A Python package). Neurocomputing 307:72–77 European Commission, Directorate-General for Agriculture and Rural Development (2025) The 28 cap strategic plans underway – summary of implementation in 2023-2024: Facts and figures. Executive summary, European Commission (DG Agriculture and Rural Development), Belgium European Court of Auditors (2021) Sustainable water use in agriculture: Cap funds more likely to promote greater rather than more efficient water use. Special Report 20, European Court of Auditors, Luxembourg European Environment Agency (2025) Water abstraction by source and economic sector in europe. URL https://www.eea.europa.eu/en/analysis/indicators/ water-abstraction-by-source-and, accessed: 4 March 2026 Fromm J, Lautner S (2007) Electrical signals and their physiological significance in plants. Plant, cell & environment 30(3):249–257 Glenn WB, et al (1950) Verification of forecasts expressed in terms of probability. Monthly weather review 78(1):1–3 González I Juclà D, Najdenovska E, Dutoit F, et al (2023) Detecting stress caused by nitrogen deficit using deep learning techniques applied on plant electrophysiological data. Scientific Reports 13(1):9633 Gu A, Dao T (2024) Mamba: Linear-time sequence modeling with selective state spaces. In: First conference on language modeling Gu A, Goel K, Ré C (2021) Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:211100396 Guo C, Pleiss G, Sun Y, et al (2017) On calibration of modern neural networks. In: International Conference on Machine Learning, PMLR, pp 1321–1330 Huber AE, Bauerle TL (2016) Long-distance plant signaling pathways in response to multiple stressors: the gap in knowledge. Journal of Experimental Botany 67(7):2063–2079 Ismail Fawaz H, Lucas B, Forestier G, et al (2020) Inceptiontime: Finding alexnet for time series classification. Data Mining and Knowledge Discovery 34(6):1936–1962 Johns S, Hagihara T, Toyota M, et al (2021) The fast and the furious: rapid long-range signaling in plants. Plant Physiology 185(3):694–706
21
Li JH, Fan LF, Zhao DJ, et al (2021) Plant electrical signals: A multidisciplinary challenge. Journal of Plant Physiology 261:153418 Meder F, Saar S, Taccola S, et al (2021) Ultraconformable, self-adhering surface electrodes for measuring electrical signals in plants. Advanced Materials Technologies 6(4):2001182 Mohr F, Wever M (2023) Naive automated machine learning. Machine Learning 112(4):1131–1170 Najdenovska E, Dutoit F, Tran D, et al (2021) Identifying general stress in commercial tomatoes based on machine learning applied to plant electrophysiology. Applied Sciences 11(12):5640 Nixon J, Dusenberry MW, Zhang L, et al (2019) Measuring calibration in deep learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, IEEE, pp 38–41 Raschka S (2018) Mlxtend: Providing machine learning and data science utilities and extensions to python’s scientific computing stack. The Journal of Open Source Software 3(24) Reissig GN, Oliveira TFdC, Oliveira RPd, et al (2021) Fruit herbivory alters plant electrome: evidence for fruit-shoot long-distance electrical signaling in tomato plants. Frontiers in Sustainable Food Systems 5:657401 Steeneken PG, Kaiser E, Verbiest GJ, et al (2023) Sensors in agriculture: towards an internet of plants. Nature Reviews Methods Primers 3(1):60 Thompson RB, Incrocci L, van Ruijven J, et al (2020) Reducing contamination of water bodies from european vegetable production systems. Agricultural water management 240:106258 Tran D, Dutoit F, Najdenovska E, et al (2019) Electrophysiological assessment of plant status outside a Faraday cage using supervised machine learning. Scientific reports 9(1):17073 Vergara JR, Estévez PA (2014) A review of feature selection methods based on mutual information. Neural computing and applications 24(1):175–186 Zhou J, Fan P, Zhou S, et al (2025) Machine learning-assisted implantable plant electrophysiology microneedle sensor for plant stress monitoring. Biosensors and Bioelectronics 271:117062
22