A Scalable Pattern Mining Workflow for Interpretable Machine Log Analysis in High-Performance Computing Environments 1st Shilpika
2nd Bethany Lusch
3rd Eric Pershey
Argonne Leadership Computing Facility Argonne Leadership Computing Facility Argonne Leadership Computing Facility Argonne National Laboratory Argonne National Laboratory Argonne National Laboratory Lemont, IL, USA Lemont, IL, USA Lemont, IL, USA [email protected] [email protected] [email protected]
arXiv:2607.19143v1 [cs.DC] 21 Jul 2026
4th Carlo Graziani
5th Venkatram Vishwanath
6th Michael E. Papka
Argonne Leadership Computing Facility Argonne Leadership Computing Facility Argonne Leadership Computing Facility Argonne National Laboratory Argonne National Laboratory Argonne National Laboratory Lemont, IL, USA Lemont, IL, USA Lemont, IL, USA [email protected] [email protected] [email protected] Department of Computer Science University of Illinois Chicago Chicago, IL, USA [email protected]
Abstract—Modern supercomputers housed in High Performance Computing (HPC) environments generate massive volumes of log data daily, revealing intricate information and performance metrics about these complex systems. The sheer size and heterogeneous nature of HPC logs, especially text data, pose significant challenges for traditional analytical techniques. Consequently, more complex workflows are necessary for pattern extraction when analyzing these logs, enabling the discovery of underlying patterns and anomalies that may indicate system faults and help predict future failures and inefficiencies. Our log analysis workflow investigates a combination of advanced pattern-matching and mining techniques applied to HPC log analysis. By systematically identifying frequent log patterns and pattern sequences in log messages and storing them in a finitestate automaton, such as the Aho-Corasick automaton, our workflow enables automated detection of frequent errors and fault events. To extract these patterns and sequences, we leverage information about system hierarchy and message priority. We then correlate and cluster the identified error sequences with job logs, revealing groups of applications with similar or dissimilar error signatures. This approach yields insights that inform improvements and guide real-time monitoring efforts. Our research establishes that pattern mining is vital for unlocking the full potential of log data by enabling real-time analysis and contributing to more resilient, scalable HPC systems. We demonstrate the effectiveness of our approach through summary statistics and a case study on an exascale-class system supercomputer. Index Terms—high-performance computing, log analysis, text data analysis, pattern mining, sequence mining, exascale supercomputer
I. I NTRODUCTION Log analysis in high-performance computing (HPC) systems presents a distinctive research challenge because mod-
ern supercomputers generate massive volumes of telemetry and event data across hardware, system software, runtimes, schedulers, storage, interconnects, and user applications, making the problem simultaneously one of scale, heterogeneity, and systems interpretation. Park et al. argue that HPC log processing has become a big-data problem, not only because the data are voluminous and largely unstructured, but also because meaningful analysis requires knowledge of hardware and software behavior across multiple layers of the system stack [1]. These difficulties become even more severe at the leadership scale. Karimi et al. describe parsing more than 600 million production logs from the Frontier supercomputer over a four-week period and emphasize that such logs arise from diverse runtime, hardware, and software layers with inconsistent formats, thereby complicating structure extraction, pattern discovery, and downstream mining [2]. A further challenge is that failures and anomalies in HPC environments are rarely isolated; instead, they often exhibit temporal and inter-event dependencies, which makes root-cause analysis dependent on robust event-correlation techniques rather than simple message counting or isolated inspection [3]. In addition, many state-of-the-art anomaly-detection methods struggle in HPC settings because operational logs frequently exhibit irregular and ambiguous time-based sequences, reducing the effectiveness of methods originally developed for more regular enterprise logs [4]. Finally, even when logs are available for study, privacy constraints and anonymization may reduce their usefulness for failure diagnosis, creating a tension between data protection and analytic fidelity [5]. Overall, the central
challenge in HPC log analysis is not merely storing or parsing large datasets, but transforming noisy, heterogeneous, multilayer operational traces into timely and actionable insight for anomaly detection, failure diagnosis, and system optimization. Log data can be broadly categorized into text-based, eventbased, and numerical data being collected from multiple sources in a supercomputer. Analyzing each data type presents distinct challenges and requires tailored solutions that account for its specific structure, representation, and processing requirements. In this work, we focus our efforts on analyzing text-based log data. To mitigate and overcome some of the challenges pertaining to analyzing text-based logs, especially in high-performance computing systems, we have built an endto-end workflow that not only processes and stores raw log representations but also helps predict the anomalies. The contributions of this work are listed below: 1. Retrospective and real-time analysis of raw text-based logs. 2. Analysis of raw text-based logs by extracting patterns from raw log messages and storing them in an AhoCorasick automaton. 3. Two-level sequence mining workflow, first level to extract sequences at the intra-node process level (or thread level) and the subsequent second level to extract sequences at the job or node level. Here, a job refers to a user application running on the supercomputer. 4. Storage of sequences in a second Aho-Corasick automaton. The automaton can be queried for searches. The queries include subsequences that can be autocompleted using a trie data structure in the automaton implementation. The remainder of this paper is organized as follows. The related work section reviews the key concepts, prior work, and theoretical foundations relevant to the study. The methodology, the experiments, and the evaluation and discussion sections then describe the proposed approach, experimental setup, implementation details, and the criteria used to assess performance. Finally, the conclusion summarizes the main findings, discusses their implications, and outlines potential directions for future research. II. R ELATED W ORK Significant efforts have been made in the past to improve HPC system resilience through log analysis. A first line of work focuses on parsing and mining unstructured HPC logs with modern language models. Karimi et al. [2] present an instruction-tuned large language model (LLM) for parsing and mining heterogeneous logs from leadership-class systems, demonstrating large-scale analysis on Frontier. Closely related, Karimi et al. [6] introduce EPIC, a generative-AI platform for operational data analytics over multimodal HPC data, including text logs, images, and tabular records. Our work focuses on analyzing text-based data, and the resulting compressed message patterns extracted can serve as input to pipelines that employ large language models.
A second line of work studies failure prediction and joboutcome prediction from log data. Alharthi et al. [7] propose Clairvoyant, a transformer-decoder approach for nodefailure prediction from system logs in large-scale systems. They extend this direction with Time Machine, a generative model for real-time failure and lead-time prediction in HPC environments [8]. Park et al. [9] analyze production scheduler logs to study and predict job failures, while Brown et al. [10] provide a longitudinal analysis of failures on the Theta Cray XC40 supercomputer. Zhang et al. [11] propose SentiLog, which applies sentiment-inspired analysis to parallel file system logs for anomaly detection. Egersdoerfer et al. [12] then introduce ModuleLog, which organizes log events by sourcecode module to improve downstream anomaly detection, and subsequently propose ClusterLog, which clusters semantically and sentimentally similar log keys to reduce irregularity and ambiguity in temporal sequences before anomaly detection [4]. Our workflow incorporates a mechanism, implemented via the automaton trie data structure, that allows autocompletion of an identified chain of patterns. A third and complementary direction is the multifidelity, multiscale visual analytics path led by Shilpika et al. Rather than focusing only on text logs, this work integrates hardware logs, job logs, and environment logs collected at different temporal resolutions. The works begin with MELA, a visual analytics tool for studying multifidelity HPC system logs [13]. It is extended by a multi-level, multi-scale framework that uses multiresolution dynamic mode decomposition (mrDMD) to reveal spatial-temporal patterns across heterogeneous HPC logs [14]. The most recent installment introduces an incremental mrDMD formulation to support faster analysis of continuously collected multifidelity system data [15]. Although these works are not restricted to text-only logs, they are highly relevant because they address cross-source correlation and multiscale interpretation of operational log data in production supercomputers, which our work addresses, although for log analysis in text format. Finally, Spell [16] and Drain [17] among others are log parsing frameworks that are effective in controlled settings but rely on heuristics and manual pattern design, making them fragile in the face of dynamic, evolving templates and domainspecific variations. Our workflow can accommodate evolving patterns and templates and relies little on manual intervention. Taken together, these studies show that contemporary HPC log-analysis research is moving in three complementary directions: (i) LLM-based parsing and mining of unstructured logs, (ii) prediction and anomaly detection from system to scheduler logs, and (iii) multifidelity, multiscale analysis that correlates textual and operational logs across subsystems. A. Pattern Extraction and Sequence Mining We use Aho–Corasick automata [18], [19] for the extraction and storage of patterns and sequences, as it has shown to be an effective mechanism in the past. Past works that have employed the Aho-Corasick automaton for log analysis include: Halas (2014) studies algorithms,
Retrospective Data Analysis : Save & Extract Patterns
Save Sequence Strings 1
Unrecognized pid * exited with status *
Log Messages
Apid * shepherd * exited successfully Opening file * in mode * …….
Aho-Corasick Automaton (Pat-AHO)
Assign Unique IDs to 1218 Extracted 899 Patterns 1636 …..
Sequence Mining to Extract Sequences
Convert Sequences to Strings
“1218-899-1636-2155-817” “1218-899-1636-2155-1937” “1218-899-1636-2155-368” “1218-899-1636-2155-1647” “1218-899-2155-1904” …..
3
2
2
3
7
112
23
37
4
9
Aho-Corasick Automaton 234 (Seq-AHO)
Select the appropriate Sequence Mining algorithm
Feedback for pattern updates
1
3 22 3
379
Real-Time Data Analysis : Unrecognized pid * exited with status *
Log Messages
Extract Patterns
Apid * shepherd * exited successfully Opening file * in mode * ATOM_SOCKET not set * disabling *
Past unseen messages are saved into the automaton
Assign Previously Stored Unique IDs to Extracted Patterns
1218 899 1636 2155 ……..
Convert Sequences to Strings
“1218-899-1636-2155” Autocomplete Sequence . Strings . “1218-899-1636-2155” – {817, 1937, 368, 1647} . . . .
Returns next likely pattern sequence
Fig. 1. Log Pattern Mining Workflow Overview
including Aho–Corasick, for matching log records and builds a log-matching tool [20]. Bellekens et al. (2014) present GLoP, a GPU library for security log analysis with GPU-accelerated Aho–Corasick [21]. Potharaju et al. (2013) use Aho–Corasick inside NetSieve to compute frequencies of repeated phrases extracted from network trouble tickets [22]. Konchagin et al. (2018) apply an Aho–Corasick-based method to search sub-traces in event logs for process mining, extending it to simultaneous search across several traces [23]. Wang et al. (2017) use Aho–Corasick to locate dictionary phrases and compute phrase frequencies in historical IT service tickets, then use those phrases for knowledge-base construction and incident inference [24]. Tuck et al. (2004) modify Aho–Corasick for intrusion detection, focusing on memoryefficient matching in Snort-style IDS pipelines [25]. Tan and Sherwood (2005) describe a high-throughput Aho–Corasickbased architecture for intrusion detection and prevention [26]. Dimopoulos et al. (2007) present a memory-efficient reconfigurable Aho–Corasick for intrusion detection systems such as Snort [27]. In our work, we employ an Aho-Corasick algorithm [18], [19] in conjunction with sequence mining to efficiently store and extract log message patterns and sequences, and to predict sequences through autocompletion of these log message sequences. Our work also uses sequential pattern mining algorithms like TUP [28], [29] and CLH-Miner [30] to group co-occuring patterns. In the past, sequence mining was applied to complex event sequences [31]–[34], or console logs in conjunction with PCA for online problem detection [35]. In our work, we deal with raw text data; we first convert it into event-like patterns and then apply sequence mining algorithms to extract meaningful sequences. III. M ETHODOLOGY Figure 1 covers the workflow of our analysis pipeline. There are two main processing channels: the first is retrospec-
tive analysis, and the second is real-time analysis. Our code will be made open source upon publication. A. System Overview During the retrospective analysis, Figure 1-top, we process historical text-based logs from multiple sources, specifically syslogs, job logs, communication logs, jobtraces, sensor logs, and application logs. The non-variable part of the log messages is retained and saved in the Aho-Corasick automaton [18], [19], [36], which we will refer to as a pat-AHO, since it is used to extract patterns from raw log messages. Each of the resulting patterns is then saved in a database (RedisDB [37]) and assigned a unique identifier. These messages are henceforth referred to as patterns, and the identifiers are pattern IDs. Since a single log message can span multiple lines in the log data file, we then mine the co-occurring log patterns using sequence mining algorithms. This gives us an idea of which log messages are co-occurring. Henceforth, we will refer to sets of patterns that co-occur as sequences. Our analysis pipeline is flexible and can adapt to any user-specified sequence mining algorithm, including transformer-based ones. In this work, we employ sequential pattern mining, specifically TUP [28], [29] and CLH-Miner [30]. The results of these algorithms are a sequence of pattern IDs that the algorithm identifies as co-occurring. These sequences are then stored in a different Aho-Corasick automaton, which we will refer to as a seq-AHO, since it is used to save sequences of patterns. We also store the sequence occurrence counts in the seq-AHO. During real-time analysis, Figure 1-bottom, we process log messages in real time. The logs are first stripped of the variable part, and then checked to see whether the resulting pattern was previously saved in the automaton. The search complexity of the automaton is O(text length + number of matches). Since we only look for exact matches and not substrings, the time complexity of the pattern search in the automaton depends on the size of the input (one log line) rather than
the size of the automaton. Additionally, multiple log lines can be processed in parallel, leading to fast pattern lookup. If a new pattern is encountered, the pattern is registered in the automaton and RedisDB with a new pattern ID. We then look up the seq-AHO to determine whether the incoming stream of pattern IDs (or sequence) exists. If the sequence is present, we can perform a lookup in the seq-AHO (trie data structure) to identify the next likely pattern IDs and their past occurrences count. An example is shown in the lower right of Figure 1. An incoming stream of pattern IDs (or sequence) could be “1218899-1636-2155”, where each pattern ID is joined by a hyphen to form a sequence. If the seq-AHO has sequences with prefix “1218-899-1636-2155” stored in the trie data structure, then a lookup could autocomplete the likely next sequence of pattern IDs,“1218-899-1636-2155” – {817, 1937, 368, 1647}. B. Sequence Mining Algorithms TUP [28], [29] is an algorithm for identifying the top K patterns with high importance. Since log lines can typically contain multiple repeated lines of informational and warning messages, and fewer lines of critical error messages, we assign an importance score either using the priority field in syslogs or manually assigning user importance scores. This weighs the fewer critical messages to have higher priority while mining patterns using TUP. CLH-Miner [30] is an algorithm for cross-level highimportance patterns. Here, the cross-level refers to a taxonomy of patterns. The taxonomy helps classify the hierarchy of log messages, such as node-level, user-level, component-level, or job-level. The resulting group of patterns (or sequences) will contain either the nodes, users, components, or jobs that have “seen” the sequences. Assigning the importance scores forces the algorithm to focus more on the rare critical messages. However, setting the importance scores to an equal value will treat all log message types equally. Section IV-(A, B) covers the data summary statistics of the pattern mining (pat-AHO) and the sequence mining (pat-AHO) part of the workflow, respectively. Additionally, Section IV-B explains how the sequence mining algorithms work in succession to extract the overall log pattern sequence. IV. E XPERIMENTS A. Log Pattern Summary Statistics This subsection presents the summary statistics of logs from an exascale supercomputer grouped by project, job, user, and clustered trend identified from the pattern mining part of the workflow. 1) Project Trends: The Figure 2 shows a stacked bar plot summarizing the distribution of log message patterns (pat-AHO) of the supercomputer across projects in a day on February 2026, with the x-axis showing distinct log message patterns (pattern IDs) and the y-axis showing the corresponding counts. Each bar is stacked by project, showing how different projects contribute to the total occurrence of each pattern. The plot shows all pattern IDs whose count crosses 10, 000. The plot reveals a clear structure in the distribution.
While individual projects differ in the frequency of specific patterns, groups of projects exhibit similar compositions of recurring messages. The plot also reveals patterns that are unique to a particular project for the duration of the processed log. This suggests shared operational behaviors, common software components, or comparable execution environments. Overall, the observed similarity in pattern usage indicates that log message distributions can provide meaningful signals for characterizing project-level behavior and may support downstream tasks such as workload profiling, anomaly detection, and automated grouping of related system activities. 2) User and Message Type Trends: The top plot Figure 3 provides a bar plot of truncated log patterns (pat-AHO) within the selected group of identifiers (message types), which indicates the source component or subsystem generating a log message. The bottom plot of Figure 3 provides a bar plot of truncated log patterns grouped by user. The dark gray dotted vertical lines split the identifier groups, and the light gray ones split the user groups. Here, we present results for January, February, and March 2026 of the supercomputer where the per-user job count is greater than 10. This visualization shows that several users across different projects exhibit similar log message patterns, while others exhibit unique patterns. This indicates cross and unique commonalities in user activity or system behavior. These shared patterns suggest that similarities observed at the project level (Section IV-(A-1) may also extend to the user level, providing further evidence of recurring operational structures and offering useful insight for behavioral profiling, workload characterization, and anomaly analysis. 3) Job Timeline Trends: Until now, we have observed trends in projects, identifiers, and user groups filtered by frequency criteria. Here, we look more in-depth at patterns within a job. The temporal scatter plot, Figure 4, illustrates how log messages generated by user jobs are distributed over time, with time shown on the x-axis and individual user jobs shown on the y-axis. Each point represents a log message occurrence and is colored by username, enabling comparison of temporal activity pattern frequency across users and jobs. The points that are overlapping are spread out using a force-directed algorithm. The size of the points indicates the count of the error pattern at that timestamp. The visualization reveals that jobs associated with the same user often exhibit similar temporal structures in their log message distributions. However, these patterns vary in density: for some jobs, log messages are distributed broadly across the execution timeline, whereas in others they appear more temporally concentrated. This suggests that user-specific job behavior can manifest through recurring temporal signatures, while differences in spread may reflect variation in workload phases, runtime characteristics, or system interactions. 4) Cluster Trends: Here, we applied pattern mining (pat-AHO) to the scheduler, console, sensor, and other jobtrace information on logs that indicated communication failure messages. The PCA plot, Figure 5, summarizes node-level variation in log message behavior by projecting a node-bypattern frequency matrix into a lower-dimensional space. In
Project ID
774067 261958 705093 196481 636216 435366 464769 865170 236205 803998 147928 525967 463902 892622 575709 463948 589636 542015 557671 277383 440238 811304 570581 209461 631950 899007 442174 492580 139466 536856 952781 965375 429995 794914 922796 737940
Count
Project-100017 Project-100018 Project-100019 Project-100029 Project-100030 Project-100034 Project-100035 Project-100043 Project-100047 Project-100051
Pattern ID
Fig. 2. Most frequent log patterns grouped by projects. The x-axis shows distinct log message patterns (pattern IDs) and the y-axis shows the corresponding counts in log scale. (a) Patterns Grouped by Message Identifier
Count (Log)
Identifier
(b) Patterns Grouped by Username Count (Log)
Username User
Extracted Patterns
Fig. 3. Most frequent log patterns grouped by identifiers (top) and users (bottom). Identifiers indicate the source component or subsystem generating a log message. The dark gray dotted vertical lines split the identifier groups, and the light gray ones split the user groups. The x-axis shows prefixes of extracted patterns.
this representation, each point corresponds to a node, and the input features are the frequencies of observed log message patterns. The resulting projection shows distinct clusters corresponding to the head node, error-prone nodes, the remaining compute nodes, and the server node, indicating that different node roles and failure states produce separable log-message signatures. Notably, the separation of error nodes from the head-node cluster helps identify nodes that appear to be losing communication with the head node. These results suggest that analysis of log-pattern frequencies can support visual diagnosis of system state, role-specific behavior, and communication-
related anomalies in large-scale computing environments. B. Case Study
So far, we have seen how extracting patterns from log messages helps gain useful insights into system usage and identify trends at various levels of the system hierarchy. In this section, we will examine how extracting temporally aligned patterns can reveal more insights into the machines and their usage. We will refer to these temporally aligned patterns as
8408032
TABLE I I MPORTANCE S CORES FOR S YSLOG P RIORITY
User 1 User 2 User 3 User 4
8408299 8408906 8404938
Priority 0 1 2 3 4 5 6 7
8408175 8408225
Job IDs
8408250 8408281 8409355 8409444
Type EMERGENCY ALERT CRITICAL ERROR WARNING NOTICE INFORMATIONAL DEBUG
Importance Score 100,000 90,000 70,000 60,000 8,000 2,000 1 1
8408612
8408275 8408282
sequence that contains importance information. For this case study, we have set k to 10. The importance scores can either be user-defined or set by using the syslog message priority field. Setting the importance score to an equal number gives equal importance to all log patterns. The TUP algorithm extracts the top-k process-level sequences, weighted or prioritized by userspecific importance scores. We use the syslog message priority for syslog sequence extraction. Table I shows the values we Fig. 4. Error patterns in a force-directed scatterplot of jobs versus timestamps. have set for each priority type in this study. We found that Each point is a log patternCOMMUNICATION seen in a process on a job’sWITH node. The clusters are JOBTRACE: LOST HEADNODE setting a larger difference in importance scores between lowercolored by usernames. The size of the points indicates the count of the error pattern at that timestamp. and higher-priority log messages yields more accurate results. Log messages can be duplicated, and the same messages may appear in large bursts within a small timeframe, as seen in Failed Nodes point sizes in Figure 4 and Figure 7. We retain only the unique Job Server messages at every 10-second interval at the process level. The HeadNode importance scores are weighted by the count of occurrence PCA-1 and the weight specified in Table I. 2) Level 2 Sequence Extraction: At Level 2, we first assign Fig. 5. PCA plot of the node-level variation in log messages computed by unique IDs to process-level sequences extracted at Level 1, projecting a node-by-pattern frequency matrix into a lower-dimensional space using PCA. refer Figure 6. We then group the sequences by jobs and nodes. Following this step, we apply the CLH-Miner Algorithm. The CLH-Miner algorithm discovers cross-level high-importance sequences. A sequence is a group of patterns that co-occur sequences in a dataset containing importance scores. Here, within defined rules. The rules specify the maximum time we can define a hierarchy of nodes or jobs that have seen window length and the number of sequences to extract within the sequences previously extracted at Level 1. For example, it. if the sequence {1,2,3} is reported on multiple nodes, we can For this case study, we use syslogs of an exascale super- provide this information to the CLH-Miner, and, along with computer for a subset of jobs on the 30th March 2026. Each sequence grouping, the algorithm identifies nodes that have log line corresponds to a process within a node in a user’s reported similar sets of sequences {1,2,3}–{4,5,6}. Note that job. Throughout the runtime of a multinode job, numerous at Level 2, we extract sequences of process-level sequences. processes on a node record messages that reflect their internal These sequences are then stored in the seq-AHO automaton state. Therefore, to extract useful sequences relating to a pro- for retrieval and future analysis. cess within a node, we first group the log patterns by job, node, Figure 7 shows the jobs along the y-axis and the timeline and process. We then apply sequence mining algorithms at two along the x-axis. We have filtered 8 jobs belonging to two levels, refer Figure 6. At Level 1, we group the patterns per users running on the system on March 30th. The underlined process within a node and extract process-specific sequences. jobs are the ones that failed with a non-zero exit code, and This is an embarrassingly parallel problem to compute. Level each failed with a user termination from the shell. Each point 2 then applies the sequence-mining algorithm to the Level 1 is a log message pattern reported by a process within a node. sequences to extract common sequences across nodes or jobs. To avoid overlapping points, we have used a force-directed We apply two separate sequence mining algorithms at each algorithm. The top three jobs of user-2 two uses lesser nodes level. Level 1 uses the “TUP algorithm,” and Level 2 uses the than the bottom three jobs. We see that the bottom three jobs “CLH-Miner Algorithm”. share a unique visual signature compared to the top three jobs 1) Level 1 Sequence Extraction: The TUP algorithm dis- and the jobs belonging to user 1. At first glance, it seems that covers the top-k high-importance sequences in a complex event the pattern signature may be unique across users. However, 8408294 8408296 8408297 8408331 8408367 8408943 8408948 8408949 8409293
PCA-2
March 30, 2026
03 AM
06 AM
09 AM
Timestamp 12 PM
03 PM
06 PM
09 PM
Level 1 Sequence Mining to Extract Sequences
Extracted Patterns Jobs→ Nodes→ Processes
Level 2
Extract Extract Log Log Log Pattern Sequences Log Pattern Sequences per Process Patterns Sequences Across Jobs and Nodes Cross Level Top-K Sequences High-Priority Weighted High-Priority Weighted
Fig. 6. Overview of the sequence mining workflow.
we are interested in the sequences of these patterns over time and across jobs or nodes. When we extract the process-level sequences at Level 1 and job-level sequences at Level 2. We found that job id 8408172 has a very unique log message signature pertaining to bad network connection reported by node x4610c4s3b0n0. Job Ids 8400999 and 8405540, belonging to user-1, have a unique signature, RPC-unknownnone-:-error-JSON-decoding-error-at-byte-of-..., where a remote procedure client called a backend/helper program that was supposed to return JSON, but instead received an unexpected text leading to a runtime error. User-1 jobs also reported hardware vendor-specific messages. All eight jobs reported an application abort and clean-up messages, which were captured as a separate stream of sequences. Since we set the startup and cleanup messages as low-importance, these informational types of messages were not captured by the sequence mining algorithms. In this case study, we conclude that even though summary statistics and other visual signatures can indicate high-level trends, a deeper analysis can yield more distinct signatures. We are interested in understanding the behavior and identifying message signatures. This could eventually drive further analysis, including root-cause identification, which is not in the scope of this work. V. E VALUATION AND D ISCUSSION In this section, we evaluate the effectiveness of our workflow in extracting patterns and sequences from syslogs of an exascale supercomputer. We have analyzed syslogs for the months of January, February, and March of 2026. The total size of the dataset is roughly 900GB. In this section, we process only the syslogs pertaining to jobs on the supercomputer. We evaluated our workflow on an NVIDIA DGX A100 node with two AMD Rome CPUs and 1TB of DDR4 Memory. A. Analytical and Throughput Evaluation
Table II shows the retrieval time for patterns (pat-AHO) and sequences (seq-AHO) stored in the automaton. Note that to store information relevant to 900GB of log data, we only require a storage space of 8.8MB for patterns and 12MB for sequences in pickle format. The ID column is a unique file ID. We vary the input character length from 60 to 1000 characters and measure the time required to retrieve these patterns. The search complexity of the automaton is
TABLE II AUTOMATON PATTERN AND S EQUENCE R ETRIEVAL T IME
Automaton
pat-AHO
seq-AHO
Size
Data Size
String Length
8.8MB 8.8MB 8.8MB 8.8MB 12MB 12MB 12MB 12MB
900GB 900GB 900GB 900GB 900GB 900GB 900GB 900GB
65 131 659 989 72 145 729 1094
Time (ms) 0.35 0.33 1.12 1.47 0.42 0.63 1.89 3.05
TABLE III L EVEL 1 S EQUENCE R ETRIEVAL : T IME AND ACCURACY
ID
Jobs
Proce- Log sses Lines
High Priority Lines
0 1 2 3 4 5 6 7 8 9
5 4 8 2 2 5 4 8 65 2
1808 319 3420 2202 1982 864 2282 189 4681 3689
12956 388 13737 26742 42891 9800 281066 10052 182043 108543
15454 965 18648 27767 44254 11337 281665 10646 196624 109099
Accu- Time racy Per Process (s) 100.0 0.3371 100.0 0.3348 100.0 0.3356 96.21 0.3681 100.0 0.3366 100.0 0.3431 100.0 0.3391 100.0 0.3303 100.0 0.3590 100.0 0.3368
Total Time (s) 609.41 106.7 1147.8 810.4 667.0 296.3 773.8 62.4 1680.47 1242.3
O(text length + number of matches) [18]. Here, we only look for exact matches as we are not interested in substring matching. However, if the user requires substring matching, it is an option they can choose. Hence, retrieval time is dependent on the text length, which is in a few milliseconds as seen in Table II. Note that the seq-AHO input is a string of pattern IDs extracted by the sequence mining Level 1 and Level 2 output (Figure 6). Table III shows the time and accuracy at Level 1 of the sequence mining part of the workflow for a subset of failed jobs in the dataset. At Level 1, we retrieve sets of co-occurring patterns. As mentioned previously, a single log message can span multiple log lines. We use level 1 to extract these temporally co-occurring log lines for a process on a node of a supercomputer. Each row of Table III contains a unique file ID (ID), total number of jobs (Jobs), total number of processes (Processes), total number of log lines (Log Lines), total number of high-importance (priority less than 5.0) log lines (High Priority Lines), accuracy of retrieval (Accuracy), time for retrieval of the sequences per process in seconds
8400999
RPC JSON Decoding Error RPC Signal Error
8405540
8408133
Job IDs
8408172
Using Vendor-specific GPU Process Management Interface for Exascale (PMIx) library
8408224
8408263
Bad Connection 8408308 User 1
8408359
User 2 Non-zero Exit Status
March 30
01 AM
02 AM
03 AM
04 AM
05 AM
06 AM
Timestamp
07 AM
08 AM
09 AM
10 AM
11 AM
Fig. 7. Error patterns in a force-directed scatterplot of jobs versus timestamps. Each point is a log pattern seen in a process of a job’s node. The size of the points indicates the count of the error pattern at that timestamp. The clusters are colored by usernames. The annotations show summaries of the sequences that are identified by the sequence mining workflow in Figure 6. The underlined y-axis labels are jobs with non-zero exit status.
(Time Per Process (s)), and total time for retrieval of the sequences in seconds (Total Time (s)). Accuracy is measured by whether the algorithm retrieves all high-priority sequences. We are only looking for high-priority sequences, as they are fewer in number and more critical than warning, alert, and informational messages, which dominate the log lines. Note that our sequence mining results include informational, warning, and alert patterns in the temporal vicinity of the highpriority patterns. Table III shows that we can retrieve all highpriority messages with over 96% accuracy and, in most cases, 100% accuracy. Note that level 1 sequence mining operates on a single process at a time, making it an embarrassingly parallel problem. We can isolate each job and processes within a job’s node independently of one another, and each takes roughly 0.3 seconds of compute time. Table IV shows the Time and Accuracy at Level 2 of the Sequence Mining part of the workflow for a subset of failed jobs in the dataset. At Level 2, we retrieve sets of co-
occurring sequences. Each of these co-occurring sequences is a process-level sequence from the output of level 1. Each row of Table IV contains a unique file ID (ID), the total number of jobs (Jobs), total processes (Processes), total log lines (Log Lines), total high-importance (priority less than 5.0) log lines (High Priority Lines), minimum importance (Min Importance), accuracy of retrieval, and time for retrieval of the sequences (Time in seconds) and total time (Total time in seconds). The total time includes time for grouping by nodes or jobs and hierarchy creation. Minimum importance is a user-defined threshold that determines the minimum importance score for the output sequence. Using this field, we can retrieve only the relevant sequences and eliminate noisy informational message patterns. However, if the user prefers to view more warning alerts and informational messages, then the min support value should be reduced. Accuracy is measured by whether the algorithm retrieves all high-priority sequences. Here again, we are only looking for high-priority sequences. Table IV
TABLE IV L EVEL 2 S EQUENCE R ETRIEVAL : T IME AND ACCURACY
ID 0 0 0 0 1 1 1 1 2 2 2 2 3 3 3 3 4 4 4 4 5 5 5 5 6 6 6 6 7 7 7 7 8 8 8 8 9 9 9 9
Jobs 2 2 2 2 4 4 4 4 4 4 4 4 5 5 5 5 4 4 4 4 2 2 2 2 8 8 8 8 5 5 5 5 8 8 8 8 2 2 2 2
Processes 2202 2202 2202 2202 2282 2282 2282 2282 166 166 166 166 1808 1808 1808 1808 319 319 319 319 1982 1982 1982 1982 3420 3420 3420 3420 864 864 864 864 464 464 464 464 3689 3689 3689 3689
Log Lines 27767 27767 27767 27767 281665 281665 281665 281665 1714 1714 1714 1714 15454 15454 15454 15454 965 965 965 965 44254 44254 44254 44254 18648 18648 18648 18648 11337 11337 11337 11337 27747 27747 27747 27747 109099 109099 109099 109099
High Priority Lines 26742 26742 26742 26742 281066 281066 281066 281066 829 829 829 829 12956 12956 12956 12956 388 388 388 388 42891 42891 42891 42891 13737 13737 13737 13737 9800 9800 9800 9800 3345 3345 3345 3345 108543 108543 108543 108543
shows that we can retrieve all high-priority messages with 100% accuracy. Note that if level 1 missed a high-importance pattern, then that would be missed by level 2 as well, since the input to level 2 is the output of level 1. The accuracy only indicates whether level 2 sequence mining retrieves all the high-importance sequences identified by level 1. Note that level 2 sequence mining can be programmed to operate at the node level or the job level. Here, we have grouped sequences from all jobs that meet the specified minimum importance score. Since we are only looking for high-importance patterns, the level 2 part of the workflow roughly takes 0.5 seconds to compute. The total time includes time for grouping by nodes or jobs and hierarchy creation, and the total processing time is on average 2.5 seconds. Our results highlight the value of pattern mining for realtime monitoring, fault characterization, and enhancing resilience in modern HPC systems. The compressed data, in the form of patterns and sequences, can be used as input for other types of analysis, for example, large language model training
Min Importance 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000 10 100 1000 5000
Accuracy 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0
Time (s) 0.57 0.57 0.57 0.57 0.58 0.57 0.58 0.57 0.56 0.56 0.56 0.56 0.56 0.58 0.57 0.56 0.56 0.56 0.56 0.56 0.59 0.58 0.58 0.57 0.56 0.60 0.57 0.56 0.62 0.59 0.61 0.59 0.67 0.69 0.66 0.67 0.56 0.56 0.56 0.57
Total Time (s) 2.23 2.60 2.95 3.31 3.10 3.52 3.94 4.35 2.25 2.58 2.92 3.25 1.57 1.92 2.26 2.61 1.32 1.65 1.98 2.32 1.80 2.15 2.50 2.86 2.05 2.40 2.75 3.10 1.68 2.06 2.45 2.83 3.037 3.5 3.9 4.3 2.83 2.06 2.46 2.85
and prompting. The log size can be prohibitive for RAGlike (retrieval augmented generation) workflows. The pattern extraction also enables us to use sequence mining approaches, such as the TUP and CLH-Miner algorithms, which have previously been used for event data analysis. As a part of the future work for this workflow, we plan to automate the setting of the minimum support variable. Currently, the drawback of this approach is that the user must make an informed decision based on their domain expertise to set the minimum support, which may not be ideal in some cases. Another drawback of this approach is that we rely on priority or importance information for the sequence mining approach, which must be set by the user for some datasets. Another part of future work is to automate the detection of message importance through a preprocessing data analytics step. We plan to extend our current workflow for root-cause analysis, which is not in the scope of this study because we are processing only text data, whereas supercomputer logs are multivariate, spanning text, events, and numerical data. For
effective root-cause analysis, a holistic view with multivariate data processing pipelines is necessary. VI. C ONCLUSION This work presents a scalable log analysis workflow for extracting, detecting, and interpreting fault-related patterns in large-scale HPC environments. By combining pattern mining algorithms with the Aho-Corasick finite-state automata for efficient sequence detection, the proposed approach enables systematic identification of frequent error patterns and error sequences from heterogeneous system logs. Associating these mined sequences and patterns across logs, clustering, and visualizations, further allows supercomputer jobs to be grouped by shared error signatures, providing insight into recurring failure modes, anomalous behavior, and system–application interactions. Through use cases and summary statistics on an exascale-class supercomputer, we demonstrate that the workflow can uncover meaningful error patterns and support automated analysis of large volumes of log data. These results highlight the value of pattern mining as a foundation for real-time monitoring, fault characterization, and reliability improvement in modern HPC systems. The compressed format of data in patterns can be used as input for other types of analysis, for example, AI and LLM training, and prompting. The pattern extraction also enables us to use sequence mining approaches, such as the TUP and CLH-Miner algorithms, which have previously been used for event data analysis. Our workflow enables online anomaly detection and predictive failure analysis with near 100% accuracy for rare, highpriority messages, and integrates with operational monitoring frameworks to provide high-speed support in building more resilient, efficient supercomputing infrastructures. VII. ACKNOWLEDGMENTS This research used resources of the Argonne Leadership Computing Facility, which is a U.S. Department of Energy Office of Science User Facility operated under contract DEAC02-06CH11357. All authors were supported by the Office of Science, U.S. Department of Energy, under contract DEAC02-06CH11357. R EFERENCES [1] B. H. Park, S. Hukerikar, R. Adamson, and C. Engelmann, “Big data meets hpc log analytics: Scalable approach to understanding systems at extreme scale,” in 2017 IEEE International Conference on Cluster Computing (CLUSTER), 2017, pp. 758–765. [2] A. M. Karimi, J. Y. Choi, C. Q. Cao, and A. Khan, “Instruction-tuned llms for parsing and mining unstructured logs on leadership hpc systems,” 2026. [Online]. Available: https://arxiv.org/abs/2604.05168 [3] X. Fu, R. Ren, J. Zhan, W. Zhou, Z. Jia, and G. Lu, “Logmaster: Mining event correlations in logs of large-scale cluster systems,” in 2012 IEEE 31st Symposium on Reliable Distributed Systems, 2012, pp. 71–80. [4] C. Egersdoerfer, D. Dai, and D. Zhang, “Clusterlog: Clustering logs for effective log-based anomaly detection,” 2023. [Online]. Available: https://arxiv.org/abs/2301.07846 [5] S. Ghiasvand and F. M. Ciorba, “Assessing data usefulness for failure analysis in anonymized system logs,” in 2018 17th International Symposium on Parallel and Distributed Computing (ISPDC), 2018, pp. 164–171.
[6] A. M. Karimi, W. Shin, J. Hines, T. Ghosal, N. S. Sattar, and F. Wang, “Epic: Generative ai platform for accelerating hpc operational data analytics,” 2025. [Online]. Available: https://arxiv.org/abs/2509.16212 [7] K. A. Alharthi, A. Jhumka, S. Di, and F. Cappello, “Clairvoyant: a log-based transformer-decoder for failure prediction in largescale systems,” in Proceedings of the 36th ACM International Conference on Supercomputing, ser. ICS ’22. New York, NY, USA: Association for Computing Machinery, 2022. [Online]. Available: https://doi.org/10.1145/3524059.3532374 [8] K. A. Alharthi, A. Jhumka, S. Di, L. Gui, F. Cappello, and S. McIntosh-Smith, “Time machine: Generative real-time model for failure (and lead time) prediction in HPC systems,” in 2023 53rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN), 2023, pp. 508–521. [Online]. Available: https://doi.org/10.1109/DSN58367.2023.00054 [9] J.-W. Park, X. Huang, and C.-H. Lee, “Analyzing and predicting job failures from HPC system log,” The Journal of Supercomputing, vol. 80, pp. 435–462, 2024. [Online]. Available: https://doi.org/10.1007/ s11227-023-05482-y [10] K. A. Brown, T. Mallick, Z. Lan, R. B. Ross, and C. D. Carothers, “Analyzing a lifetime of failures on a cray xc40 supercomputer,” in Proceedings of the Cray User Group, ser. CUG ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 103–114. [Online]. Available: https://doi.org/10.1145/3757348.3757360 [11] D. Zhang, D. Dai, R. Han, and M. Zheng, “Sentilog: Anomaly detecting on parallel file systems via log-based sentiment analysis,” in Proceedings of the 13th ACM Workshop on Hot Topics in Storage and File Systems, ser. HotStorage ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 86–93. [Online]. Available: https://doi.org/10.1145/3465332.3470873 [12] C. Egersdoerfer, D. Dai, and D. Zhang, “ModuleLog: Module based approach to anomaly detection in parallel file system logs,” https: //ninercommons.charlotte.edu/record/4157?ln=en&v=pdf, 2022, [Accessed 07-05-2026]. [13] F. Shilpika, B. Lusch, M. Emani, V. Vishwanath, M. E. Papka, and K.-L. Ma, “Mela: A visual analytics tool for studying multifidelity hpc system logs,” in 2019 IEEE/ACM Industry/University Joint International Workshop on Data-center Automation, Analytics, and Control (DAAC), 2019, pp. 13–18. [14] S. Shilpika, B. Lusch, M. Emani, F. Simini, V. Vishwanath, M. E. Papka, and K.-L. Ma, “A multi-level, multi-scale visual analytics approach to assessment of multifidelity hpc systems,” in 2024 IEEE 24th International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2024, pp. 478–488. [15] Shilpika, B. Lusch, V. Vishwanath, and M. E. Papka, “An incremental multi-level, multi-scale approach to assessment of multifidelity hpc systems,” in Proceedings of the SC ’24 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, ser. SC-W ’24. IEEE Press, 2025, p. 1576–1587. [Online]. Available: https://doi.org/10.1109/SCW63240.2024.00197 [16] M. Du and F. Li, “Spell: Streaming parsing of system event logs,” in 2016 IEEE 16th International Conference on Data Mining (ICDM), 2016, pp. 859–864. [17] P. He, J. Zhu, Z. Zheng, and M. R. Lyu, “Drain: An online log parsing approach with fixed depth tree,” in 2017 IEEE International Conference on Web Services (ICWS), 2017, pp. 33–40. [18] A. V. Aho and M. J. Corasick, “Efficient string matching: an aid to bibliographic search,” Commun. ACM, vol. 18, no. 6, p. 333–340, Jun. 1975. [Online]. Available: https://doi.org/10.1145/360825.360855 [19] https://en.wikipedia.org/wiki/AhoCorasick algorithm, [Accessed 06-052026]. [20] F. Halas, “Efficient multi-pattern log matching in go language,” 2014. [Online]. Available: https://is.muni.cz/th/xbva1/thesis.pdf [21] X. J. A. Bellekens, C. Tachtatzis, R. C. Atkinson, C. Renfrew, and T. Kirkham, “Glop: Enabling massively parallel incident response through gpu log processing,” in Proceedings of the 7th International Conference on Security of Information and Networks, ser. SIN ’14. New York, NY, USA: Association for Computing Machinery, 2014, p. 295–301. [Online]. Available: https://doi.org/10.1145/2659651.2659700 [22] R. Potharaju, N. Jain, and C. Nita-Rotaru, “Juggling the jigsaw: towards automated problem inference from network trouble tickets,” in Proceedings of the 10th USENIX Conference on Networked Systems Design and Implementation, ser. nsdi’13. USA: USENIX Association, 2013, p. 127–142.
[23] A. M. Konchagin and A. A. Kalenkova, “On the efficient application of aho-corasick algorithm in process mining,” in Analysis of Images, Social Networks and Texts. Cham: Springer International Publishing, 2018, pp. 371–377. [24] Q. Wang, W. Zhou, C. Zeng, T. Li, L. Shwartz, and G. Y. Grabarnik, “Constructing the knowledge base for cognitive it service management,” in 2017 IEEE International Conference on Services Computing (SCC), 2017, pp. 410–417. [25] N. Tuck, T. Sherwood, B. Calder, and G. Varghese, “Deterministic memory-efficient string matching algorithms for intrusion detection,” in IEEE INFOCOM 2004, vol. 4, 2004, pp. 2628–2639 vol.4. [26] L. Tan and T. Sherwood, “A high throughput string matching architecture for intrusion detection and prevention,” in Proceedings of the 32nd Annual International Symposium on Computer Architecture (ISCA), 2005, pp. 112–122. [27] V. Dimopoulos, I. Papaefstathiou, and D. N. Pnevmatikatos, “A memoryefficient reconfigurable aho-corasick fsm implementation for intrusion detection systems,” in International Conference on Embedded Computer Systems: Architectures, Modeling and Simulation (IC-SAMOS), 2007, pp. 186–193. [28] S. Rathore, S. Dawar, V. Goyal, and D. Patel, “philippe-fournierviger.com,” https://www.philippe-fournier-viger.com/spmf/TUP.pdf, [Accessed 07-05-2026]. [29] S. Wan, J. Chen, W. Gan, G. Chen, and V. Goyal, “THUE: discovering top-k high utility episodes,” CoRR, vol. abs/2106.14830, 2021. [Online]. Available: https://arxiv.org/abs/2106.14830 [30] P. Fournier-Viger, Y. Wang, J. C.-W. Lin, J. M. Luna, and S. Ventura, “Mining cross-level high utility itemsets,” in Trends in Artificial Intelligence Theory and Applications. Artificial Intelligence Practices: 33rd International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, IEA/AIE 2020, Kitakyushu, Japan, September 22-25, 2020, Proceedings. Berlin, Heidelberg: Springer-Verlag, 2020, p. 858–871. [Online]. Available: https://doi.org/10.1007/978-3-030-55789-8 73 [31] P. Fournier-Viger, J. Li, J. C.-W. Lin, T. T. Chi, and R. Uday Kiran, “Mining cost-effective patterns in event logs,” Know.-Based Syst., vol. 191, no. C, Mar. 2020. [Online]. Available: https://doi.org/10.1016/j. knosys.2019.105241 [32] I. Mavroudopoulos, T. Toliopoulos, C. Bellas, A. Kosmatopoulos, and A. Gounaris, “Sequence detection in event log files,” in International Conference on Extending Database Technology, 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:232283442 [33] J. Chen, P. Wang, S. Du, and W. Wang, “Log pattern mining for distributed system maintenance,” Complexity, vol. 2020, no. 1, p. 6628165, 2020. [Online]. Available: https://onlinelibrary.wiley.com/doi/ abs/10.1155/2020/6628165 [34] M. Leemans and W. M. P. van der Aalst, “Discovery of frequent episodes in event logs,” in Data-Driven Process Discovery and Analysis, P. Ceravolo, B. Russo, and R. Accorsi, Eds. Cham: Springer International Publishing, 2015, pp. 1–31. [35] W. Xu, L. Huang, A. Fox, D. Patterson, and M. Jordan, “Online system problem detection by mining patterns of console logs,” in 2009 Ninth IEEE International Conference on Data Mining, 2009, pp. 588–597. [36] “pyahocorasick — ahocorasick documentation — pyahocorasick.readthedocs.io,” https://pyahocorasick.readthedocs.io/en/latest/ #testimonials, [Accessed 06-05-2026]. [37] “Redis reference — redis.io,” https://redis.io/docs/latest/develop/ reference/, [Accessed 06-05-2026].