arXiv:2605.02194v1 [cs.DC] 4 May 2026
Statistical Characterization of IO500 Submission Data: Performance Distributions, Correlations, and Log-Derived Insights Julian Kunkel, Aasish Kumar Sharma, Anila Ghazanfar, Sepehr Mahmoodianhamedani, and Sascha Safenreider Georg-August-Universität Göttingen / GWDG mbH, Göttingen, Germany {julian.kunkel, aasish-kumar.sharma, anila.ghazanfar, s.mahmoodian, sascha.safenreiter}@gwdg.de
Abstract. The IO500 benchmark has become the community standard for evaluating HPC storage system performance, yet the detailed data contained in its submission packages remains largely unexplored beyond aggregate leaderboard rankings. We present a statistical characterization of 61 IO500 submissions from four competition lists (ISC21 through SC22), examining score distributions, inter-phase correlations, and insights derived from detailed log files that accompany each submission. Our analysis reveals that IO500 scores span four orders of magnitude. Spearman correlation analysis shows strong within-domain clustering for both bandwidth (rs = 0.78 to 0.96) and metadata (rs = 0.89 to 0.98) phases, with the composite sub-scores exhibiting rs = 0.92 at per-node level (Pearson r = 0.53). Log-level analysis uncovers file-system-specific patterns in IOR close-time overhead, straggler behavior during the stonewall wear-down phase, and parallel-find load imbalance that are invisible in aggregate scores. These findings demonstrate that IO500 submission packages constitute a valuable research resource for understanding storage system behavior. The full submission dataset is now publicly available at https://github.com/IO500/submission-data, and our analysis scripts at https://gitlab-ce.gwdg.de/hpc-team/io500-analysis, enabling reproducible follow-up studies.
Keywords: IO500 · HPC storage · benchmark analysis · I/O performance · statistical characterization
1
Introduction
Benchmarking storage systems is central to procurement decisions, deployment validation, and performance understanding in high-performance computing. The IO500 [7,9] has become the community standard for this purpose, combining bandwidth measurements (IOR) with metadata evaluations (MDTest) into a composite score. Since its establishment, hundreds of submissions have been
2
J. Kunkel et al.
collected, each containing aggregate scores and detailed log files with per-process timing, stonewall statistics, and architectural metadata. Despite this wealth of data, published analyses have focused primarily on leaderboard rankings and top-performing systems [8]. The detailed submission packages constitute a largely untapped research resource. Systematic analysis can reveal patterns in storage system behavior, relationships between benchmark phases, and performance characteristics that aggregate scores obscure. We present a descriptive statistical characterization of 61 IO500 submissions from four competition lists spanning 2021 to 2022 (ISC21, SC21, ISC22, SC22). Our contributions are: 1. Descriptive characterization of IO500 score distributions and variability across file systems, interconnects, and deployment scales. 2. Correlation analysis between benchmark phases using Spearman rank correlations with false discovery rate correction, identifying distinct bandwidth and metadata clusters. 3. Log-derived insights from per-process timing data, revealing file-systemspecific patterns in close-time overhead, straggler behavior, and parallel-find load balancing. We emphasize that our findings describe patterns in this specific dataset and should be interpreted cautiously when generalizing. All analysis scripts and the full IO500 submission data are publicly available (see Section 4).
2
Background: The IO500 Benchmark
The IO500 benchmark suite [7,9] characterizes storage performance through two primary components and a find phase. IOR [17] measures bandwidth in two configurations. IOR-easy uses a single shared file with large sequential I/O, emulating optimized scientific applications. IOR-hard uses file-per-process access with 47,008-byte interleaved records, stressing small-file and metadata handling typical of checkpoint/restart workloads. Both report throughput in GiB/s. MDTest [10] measures metadata operation rates (create, stat, read, delete) in two modes. MDTest-easy operates within a single shared directory. MDTesthard uses a directory-per-process layout with small data payloads. Results are in kIOPS (thousands of I/O operations per second). Parallel find (pfind) measures namespace traversal rate across the file trees created by MDTest, reported in kIOPS. Scoring. Composite scores use geometric means: ScoreBW = Sior-easy-w · Sior-easy-r · Sior-hard-w · Sior-hard-r
1/4
ScoreMD = Smd-easy-w · Smd-easy-s · Smd-hard-w · Smd-hard-s · Sfind p Scoreoverall = ScoreBW × ScoreMD
(1) 1/5
(2) (3)
Title Suppressed Due to Excessive Length
3
The overall score thus combines fundamentally different units (bandwidth and metadata rate) via geometric aggregation. Submission packages include per-process CSV timing data (start/end times per operation), stonewall throughput (performance at the moment the 300-second time limit is reached), and file-closing durations. The stonewall mechanism ensures write phases run for at least 300 seconds; afterward, all processes complete in-flight operations in a “wear-down” phase, emulating bulk-synchronous behavior and exposing straggler effects.
3
Related Work
IOR [17] and MDTest [10] have long been used individually for HPC storage characterization. Kunkel et al. [7] unified them into the IO500 with standardized configurations and composite scoring. Monnier et al. [14] combined IO500 with the Mistral profiling tool for platform-level storage evaluation. I/O variability and workload characterization have received extensive study. Lofstead et al. [12] analyzed variability in petascale storage systems; Luu et al. [13] conducted cross-platform I/O studies. Carns et al. [2] and Snyder et al. [18] developed continuous characterization via Darshan, and Lockwood et al. [11] proposed holistic analysis through the UMAMI framework. Major parallel file systems including Lustre [16], GPFS [15], and DAOS [4,3] have been studied individually using various benchmarks. Despite this body of work, no prior study systematically analyzes the IO500 submission corpus itself as a statistical dataset, examining inter-phase relationships and extracting insights from the per-process log files that accompany each submission. Our work addresses this gap.
4
Methodology
Our analysis pipeline (Figure 1) proceeds from data acquisition through statistical analysis and visualization. 4.1
Data Acquisition and Scope
We analyze 61 IO500 submissions obtained from VI4IO archive [19] snapshots, spanning four competition lists: ISC21 (June 2021), SC21 (November 2021), ISC22 (June 2022), and SC22 (November 2022). The distribution is 24 submissions from 2021 lists and 37 from 2022 lists. These submissions were selected because their full packages (including log files) were available at data collection time. Since our initial analysis, the IO500 community has released the full submission dataset on GitHub [5], now containing data from 131+ sites. We use the original 61-submission subset; extending to the full repository is future work (Section 9). Our analysis scripts are at https://gitlab-ce.gwdg.de/hpc-team/ io500-analysis.
4
J. Kunkel et al. 1. Data Acquisition (VI4IO Archive, 61 submissions)
2. Cleaning & Normalization
3. Validation (completeness, consistency, outlier checks)
4. Derived Metrics (per-node, per-process, CV)
5. Statistical Analysis (Spearman, Kruskal-Wallis, FDR)
6. Visualization & Interpretation
Fig. 1: Analysis pipeline for IO500 repository characterization.
4.2
Data Cleaning and Quality
Cleaning steps include: file system name standardization (e.g., grouping GPFS and Spectrum Scale), interconnect speed normalization to Gb/s, and exclusion of records with missing values on a per-analysis basis. Data quality limitations. Self-reported submission metadata contains known quality issues. Interconnect speed is sometimes ambiguous: some submissions report throughput exceeding nominal single-NIC speed, likely because NIC count per node is not consistently reported. We treat interconnect speed as a categorical grouping variable rather than a precise measurement and document these inconsistencies rather than attempting correction. 4.3
Derived Metrics and Statistical Methods
To compare systems of different scales, we compute per-node and per-process normalized scores. Per-node normalization removes the dominant effect of deployment scale, enabling architectural efficiency comparison. We note that dividing the composite overall score (a geometric mean spanning GiB/s and kIOPS) by node count produces a mixed-unit ratio; individual phase normalizations are more directly interpretable. Correlation analysis. We use Spearman rank correlation (rs ) as our primary measure, appropriate for the skewed, non-normal distributions typical of benchmark data. Pearson correlations are reported for comparison. Multiple comparisons across the correlation matrix are controlled via Benjamini-Hochberg false discovery rate (FDR) correction [1] at α = 0.05. Group comparisons. Kruskal-Wallis H-tests [6] compare per-node performance across interconnect speed groups, with eta-squared (η 2 = H/(n−1)) effect sizes reported alongside p-values. The independence assumption may be violated when multiple submissions originate from the same organization; reported p-values are therefore approximate.
Title Suppressed Due to Excessive Length
5
Table 1: Dataset composition: 61 IO500 submissions by file system and interconnect. File System
N
Lustre 27 GPFS/SpectrumScale 12 DAOS 10 WekaFS 7 BeeGFS / Other 5
Interconnect
N
IB HDR (200 Gb/s) 22 IB EDR (100 Gb/s) 18 Omni-Path (100 Gb/s) 11 Other / Unknown 10
Table 2: Summary statistics for IO500 metrics across 61 submissions. Metric
Min Median Mean
Overall Score 3.2 Score BW (GiB/s) 0.6 Score MD (kIOPS) 10.5 IOR-easy Write 1.8 IOR-hard Write 0.0 MDTest-easy Create 12.0 MDTest-hard Create 0.1
5
Max CV
254 1,128 36,850 4.17 88 156 3,422 2.79 837 10,641 396,873 4.82 113 220 3,360 2.00 24.8 122 3,717 3.91 768 7,639 278,101 4.66 279 3,016 102,419 4.40
Dataset Characteristics
Table 1 summarizes the composition of our dataset. Lustre dominates with 27 submissions (44%), followed by GPFS/Spectrum Scale (12), DAOS (10), WekaFS (7), and BeeGFS/others (5). InfiniBand HDR (200 Gb/s) and EDR (100 Gb/s) together account for 40 of 61 submissions. Public IO500 submissions likely over-represent well-tuned configurations. The 2021-2022 window captures a period of rapid evolution, particularly the emergence of DAOS.
6
Results
6.1
Score Distributions
Figure 2 shows overall IO500 scores spanning approximately four orders of magnitude. Table 2 presents summary statistics; all metrics exhibit coefficients of variation exceeding 2.4, with “hard” configurations (IOR-hard, MDTest-hard) consistently showing higher variability than their “easy” counterparts. This indicates that small-I/O and per-directory metadata performance are more sensitive to architectural differences than large sequential I/O. The mean substantially exceeds the median in all cases, confirming right-skewed distributions driven by a few high-performance systems.
6
J. Kunkel et al. Score
ScoreBW
ScoreMD
IOR.easy.write
IOR.easy.read
IOR.hard.write
IOR.hard.read
MD.easy.write
MD.easy.stat
MD.easy.delete
MD.hard.write
MD.hard.read
MD.hard.stat
MD.hard.delete
find
1e+07 MadFS
1e+04 1e+01
1e+07 1e+04
Score
1e+01
1e+07 1e+04 1e+01
1e+07 1e+04 1e+01 Qumulo
0
20
40
60 0
20
40
60 0
20
40
60 0
20
40
60
Position (sorted by score)
Fig. 2: IO500 scores across 61 submissions, colored by file system type (log scale). Scores span approximately four orders of magnitude, reflecting substantial heterogeneity in system scale and architecture.
6.2
Correlation Analysis
Figure 3 presents Spearman rank correlation matrices for per-node and perprocess normalized scores. Two distinct clusters emerge: IOR bandwidth phases correlate strongly within-domain (rs = 0.78 to 0.96), as do MDTest metadata phases (rs = 0.89 to 0.98). Cross-domain correlations between IOR and MDTest phases, while still positive, are generally weaker in the per-process normalized view, supporting the IO500’s multi-dimensional scoring design. The composite sub-scores ScoreBW and ScoreMD exhibit a strong correlation (rs = 0.92, Pearson r = 0.53) at the per-node level, despite the weaker component-level cross-domain correlations. This is a consequence of geometric mean aggregation combined with residual scale effects: systems with more resources tend to achieve higher scores in both domains. Pearson correlations show consistent patterns with somewhat different magnitudes due to outlier influence. The rank-based Spearman results are more robust to the skewed distributions observed above and are our primary measure. Interpretation. These correlation patterns suggest that some benchmark phases capture partially overlapping performance characteristics in this dataset. However, we caution against interpreting observed correlations as evidence that any phase is redundant in general. The IO500 is designed to cover a broad space of
IOR.easy.read
7
IOR.easy.write
ScoreBW
IOR.hard.write
IOR.hard.read
Score
MD.easy.write
MD.hard.read
MD.hard.write
MD.easy.stat
MD.hard.stat
ScoreMD
MD.hard.delete
find
MD.easy.delete
IOR.easy.write
IOR.easy.read
ScoreBW
IOR.hard.write
IOR.hard.read
MD.easy.stat
MD.hard.stat
Score
MD.easy.write
find
MD.hard.write
ScoreMD
MD.hard.read
MD.hard.delete
MD.easy.delete
Title Suppressed Due to Excessive Length
1
MD.easy.delete
1
MD.easy.delete
MD.hard.delete
find
0.8
MD.hard.read
0.8
MD.hard.delete 0.6
0.6
ScoreMD
ScoreMD
MD.hard.write
0.4
find
MD.hard.stat
0.4
MD.easy.stat 0.2
MD.easy.write
0.2
MD.hard.write
Score
0
MD.hard.stat
MD.hard.read
0
MD.easy.write −0.2
−0.2
MD.easy.stat
Score
IOR.hard.read
−0.4
IOR.hard.write
IOR.hard.read
−0.4
IOR.hard.write −0.6
−0.6
ScoreBW
ScoreBW
IOR.easy.read
−0.8
IOR.easy.write
IOR.easy.write
−0.8
IOR.easy.read −1
−1
(a) Per-node normalized
(b) Per-process normalized
200 GBit
25 GBit
32 GBit
NA
10.0
1.0
0.1 0
20
40
Position (ordered by IO500 total score)
(a) Overall score
60
100 GBit
30.0
200 GBit
25 GBit
32 GBit
NA
10.0 3.0 1.0 0.3 0
20
40
60
MDTest easy write kOPS per Node
Score per Node
100 GBit 100.0
IOR Easy Write GiB/s per Node
Fig. 3: Spearman rank correlation matrices for IO500 phase scores. Color intensity and circle size encode correlation magnitude. All displayed correlations are significant after Benjamini-Hochberg FDR correction at α = 0.05. 100 GBit
1e+03
25 GBit
32 GBit
NA
1e+01 1e+00 1e−01 0
20
40
60
Position (sorted by MDTest score)
Position (sorted by IOR score)
(b) IOR-easy write
200 GBit
1e+02
(c) MDTest-easy write
Fig. 4: Per-node IO500 performance grouped by reported interconnect speed. Higher speeds are associated with higher per-node performance, with substantial within-group variability. possible storage system behaviors; correlation patterns in 61 well-tuned systems may not hold for emerging architectures or differently configured deployments. Design decisions about benchmark composition should consider the full range of intended use cases, not correlation patterns from a single sample. 6.3
Interconnect Speed and Per-Node Performance
Figure 4 shows per-node scores grouped by interconnect speed. Higher speeds are generally associated with higher per-node performance. Kruskal-Wallis tests yield H = 4.40, p = 0.111, η 2 = 0.11 for overall per-node score, indicating a moderate effect size but falling short of statistical significance at α = 0.05. For IOR-easy per-node bandwidth, the association is stronger (H = 11.16, p = 0.004, η 2 = 0.28), consistent with the expectation that large sequential I/O is more directly constrained by network bandwidth than metadata operations.
8
J. Kunkel et al. IOR.easy.write
IOR.easy.read
IOR.hard.write Lustre
10000
IOR.hard.read
1000 100 Lustre WekaFS
10
Qumulo
1 MD.easy.write
MD.easy.stat
MD.easy.delete
MD.hard.write
MD.hard.delete
find
Runtime
10000 1000 100 10
OceanStor
WekaFS
1 MD.hard.read 10000
DAOS
MD.hard.stat Lustre
OceanStor
1000 100 10 1
Vast 0
20
Qumulo 40 60 0
20
40
60 0
20
40
OPENFS OceanStor Pacific 60 0 20 40 60
Position (sorted by score)
Fig. 5: Runtime distributions across IO500 phases. Write phases are bounded below by the 300-second stonewall; some read/stat phases complete in under 10 seconds (possible residual caching effects). We use associational language deliberately. Confounding factors including storage media type, file system architecture, storage server count, and tuning quality vary simultaneously across submissions. Group sizes are unequal (as few as n = 5), limiting statistical power. Reported interconnect speeds may not reflect effective per-node bandwidth due to unreported NIC counts (Section 4.2). These limitations prevent causal conclusions about the role of interconnect speed.
7
Log-Derived Insights
Beyond aggregate scores, IO500 submission packages contain detailed log files that reveal performance characteristics invisible in summary metrics. 7.1
Runtime Distributions and Close-Time Overhead
Figure 5 shows phase runtimes. Write phases consistently meet or exceed the 300second stonewall, with some extending beyond 3,600 seconds during the weardown phase. Some read and stat phases complete in approximately 10 seconds or less, raising questions about residual caching despite IO500’s read-after-write rules. We flag these submissions as potentially cache-affected without excluding them.
Title Suppressed Due to Excessive Length
DAOS
1e+01 1e+00 1e−01 1e−02 1e−03
Lustre
ior−hard−write
1e+01 1e+00 1e−01 1e−02 DAOS 1e−03 0.00
ior−hard−read
Overhead in s
Kapok
ior−easy−write
1e+01 1e+00 1e−01 1e−02 1e−03
ior−easy−read
DAOS OceanStor
1e+01 1e+00 1e−01 1e−02 DAOS 1e−03 MadFS
9
0.25
0.50
0.75
1.00
Fig. 6: IOR close-time overhead by file system. Lustre shows up to tens of seconds (cache flush and metadata finalization); DAOS shows negligible overhead. Close time is included in the IOR timing measurement.
A particularly valuable piece of hidden information in IOR log files is the time spent in the file close operation. Figure 6 shows close-time overhead by file system. Lustre systems exhibit close times of up to tens of seconds, reflecting the cost of flushing data from client caches and completing metadata operations. DAOS shows negligible close overhead, consistent with its persistent-memorybased architecture [4,3]. This observation is practically important: close time is included in the IOR timing measurement and represents genuine application-visible I/O cost. The variation across file systems highlights that different architectures distribute the cost of data persistence differently between write and close phases.
10
J. Kunkel et al. WekaFS
Lustre Lustre
WekaFS
Kapok Qumulo
BPFS
5
Ratio stonewall/score
Vast BPFS
4 Lustre
3 2 DAOS
1 WekaFS 0 0.00
0.25
0.50
0.75
1.00
Quantiles
(a) IOR-easy write
(b) IOR-hard write
Fig. 7: Stonewall-relative Q-Q plots for IOR writes. Values near 1.0 indicate uniform process completion; larger values reveal stragglers in the wear-down phase. IOR-hard shows substantially larger deviations (up to 2× to 5×). 7.2
Stonewall-Relative Analysis
To understand per-process behavior during the stonewall phase, we compute the ratio of each process’s total runtime to the stonewall duration. Figure 7 shows these ratios as Q-Q plots for IOR write phases. IOR-easy processes finish near the stonewall time (ratio ≈ 1.0) with moderate tail deviations, as expected for the shared-file access pattern where the file system can balance load across storage targets. IOR-hard shows substantially larger deviations (up to 2× to 5×), reflecting the straggler sensitivity of the file-perprocess pattern. MDTest phases exhibit analogous patterns: MDTest-hard with larger deviations than MDTest-easy. File-system-specific patterns are visible. In particular, some Lustre submissions on HDD-based storage show the largest stonewall-relative deviations, while WekaFS submissions display characteristic step-like patterns attributable to internal storage bucket oversubscription. 7.3
Process-Level Straggler Patterns
Individual submissions reveal architecture-specific straggler patterns (Figure 8). The WekaFS submission shows clustered groups of slow processes, consistent with its distributed architecture where processes mapped to the same storage bucket experience correlated contention. The Lustre submission shows a contiguous range of slow MPI ranks, consistent with Lustre’s OST striping model where a single overloaded storage target affects all mapped processes. These contrasting patterns demonstrate that process-level log analysis provides architectural insights entirely invisible in aggregate scores. The straggler structure reveals how each file system distributes load and where contention bottlenecks arise. 7.4
Parallel Find Load Imbalance
The pfind phase reveals extreme load imbalance: a single process may check over 5 million files while others check approximately 100,000. This arises from
Title Suppressed Due to Excessive Length
330
Runtime
Runtime
700
11
500
320
310
300
300 0.00
0.25
0.50
0.75
Quantiles
(a) WekaFS: clustered stragglers
1.00
0.00
0.25
0.50
0.75
1.00
Quantiles
(b) Lustre: contiguous stragglers
Fig. 8: Per-process Q-Q plots for IOR-hard write from individual submissions, illustrating contrasting straggler architectures. (a) WekaFS: clustered groups suggest storage-bucket-level contention. (b) Lustre: contiguous slow ranks suggest OST-level hotspots. the inherent difficulty of parallelizing directory tree traversal, compounded by MDTest-hard’s per-process directory structure which cannot be efficiently distributed across find processes. Job-stealing mechanisms partially mitigate the skew but do not eliminate it. Runtime metrics (time spent in job stealing, active utilization) reveal that most processes spend the majority of their time waiting rather than actively traversing, indicating opportunities for improved scheduling in the find implementation.
8
Practical Implications
Our analysis yields actionable insights for several stakeholder groups: System procurement. The composite IO500 score obscures trade-offs between bandwidth and metadata performance. Per-node analysis reveals substantial architectural efficiency differences hidden when raw scores (dominated by node count) are compared. Evaluators should examine per-node scores alongside aggregate rankings and prioritize the phase scores most relevant to their workload profile. Benchmarking practice. Log-level analysis exposes bottlenecks (close-time overhead, straggler behavior) invisible in aggregate throughput numbers. Stonewallrelative analysis provides a standardized way to assess process-level performance uniformity. We recommend that benchmark consumers inspect timing breakdowns, not just headline throughput. IO500 community. Submission packages are a valuable research resource beyond rankings. The public GitHub dataset [5] enables longitudinal studies and crosssite modeling. We encourage submitters to provide accurate metadata (particularly NIC count and storage target count) to improve cross-submission analysis quality.
12
J. Kunkel et al.
Storage system design. The file-system-specific straggler patterns (Section 7.3) provide actionable information about load balancing behavior. Clustered patterns (WekaFS) and contiguous patterns (Lustre) point to different architectural bottlenecks and may guide optimization efforts for storage target mapping and load distribution.
9
Conclusion and Future Work
We have presented a statistical characterization of 61 IO500 submissions from the 2021-2022 competition lists. Our analysis documents substantial score variability (CV > 2 across all metrics), identifies distinct bandwidth and metadata correlation clusters (rs = 0.78 to 0.98 within, rs = 0.70 to 0.95 across domains), and demonstrates the value of log-level analysis for uncovering file-system-specific performance patterns in close-time overhead, straggler behavior, and find load imbalance. Limitations. Our 61-submission dataset is a subset of available data; observed patterns may not generalize. Self-reported metadata contains quality issues, particularly regarding interconnect specifications. Statistical tests assume independence between submissions, which is violated when multiple submissions originate from the same site. Future work should extend this analysis to the full GitHub dataset [5] (131+ sites) for temporal trend analysis and more robust statistical conclusions. Specific directions include longitudinal per-node efficiency tracking, process-level straggler modeling from CSV log files, and automated anomaly detection to flag unusual submissions. Acknowledgments. This work was conducted at the Institute of Computer Science, Georg-August-Universität Göttingen, in collaboration with GWDG (Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen). We thank the IO500 community and the VI4IO initiative for making submission data publicly available.
References 1. Benjamini, Y., Hochberg, Y.: Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological) 57(1), 289–300 (1995). https://doi.org/10.1111/j. 2517-6161.1995.tb02031.x 2. Carns, P., Harms, K., Allcock, W., Bacon, C., Lang, S., Latham, R., Ross, R.: Understanding and improving computational science storage access through continuous characterization. ACM Transactions on Storage (TOS) 7(3), 1– 26 (2011). https://doi.org/10.1145/2027066.2027068, https://doi.org/10. 1145/2027066.2027068 3. Hennecke, M.: Understanding daos storage performance scalability. In: Proceedings of the HPC Asia 2023 Workshops. pp. 1–14 (2023). https://doi.org/10.1145/ 3581576.3581577, https://doi.org/10.1145/3581576.3581577
Title Suppressed Due to Excessive Length
13
4. Intel Corporation: DAOS: Distributed asynchronous object storage (2024), https: //docs.daos.io/, accessed: 2025-10-03 5. IO500 Community: IO500 submission data repository (2025), https://github. com/IO500/submission-data, accessed: 2026-03-24 6. Kruskal, W.H., Wallis, W.A.: Use of ranks in one-criterion variance analysis. Journal of the American Statistical Association 47(260), 583–621 (1952). https: //doi.org/10.1080/01621459.1952.10483441 7. Kunkel, J., Bent, J., Lofstead, J., Markomanolis, G.S.: Establishing the io500 benchmark. White Paper (2016), https://pdsw.org/pdsw-discs17/wips/ kunkel-wip-pdsw-discs17.pdf 8. Kunkel, J., Betke, E.: Tracking user-perceived i/o slowdown via probing. In: International Conference on High Performance Computing. pp. 169–182. Springer (2019). https://doi.org/10.1007/978-3-030-34356-9_15, https://doi.org/ 10.1007/978-3-030-34356-9_15 9. Kunkel, J., Lofstead, J., Bent, J.: The IO500 – benchmarking the I/O for supercomputers. In: High Performance Computing: ISC High Performance 2019 International Workshops. pp. 580–590. Springer (2019). https://doi.org/10.1007/ 978-3-030-34356-9_44 10. LLNL: mdtest – HPC metadata benchmark (2007), https://github.com/LLNL/ mdtest, available at https://github.com/LLNL/mdtest 11. Lockwood, G.K., Yoo, W., Byna, S., Wright, N.J., Snyder, S., Harms, K., Nault, Z., Carns, P.: Umami: a recipe for generating meaningful metrics through holistic i/o performance analysis. In: Proceedings of the 2nd Joint International Workshop on Parallel Data Storage & Data Intensive Scalable Computing Systems. pp. 55– 60 (2017). https://doi.org/10.1145/3149393.3149395, https://doi.org/10. 1145/3149393.3149395 12. Lofstead, J., Zheng, F., Liu, Q., Klasky, S., Oldfield, R., Kordenbrock, T., Schwan, K., Wolf, M.: Managing variability in the io performance of petascale storage systems. In: SC’10: Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis. pp. 1– 12. IEEE (2010). https://doi.org/10.1109/SC.2010.32, https://doi.org/10. 1109/SC.2010.32 13. Luu, H., Winslett, M., Gropp, W., Ross, R., Carns, P., Harms, K., Prabhat, M., Byna, S., Yao, Y.: A multiplatform study of I/O behavior on petascale supercomputers. In: Proceedings of the 24th International Symposium on HighPerformance Parallel and Distributed Computing. pp. 33–44 (2015). https://doi. org/10.1145/2749246.2749269, https://doi.org/10.1145/2749246.2749269 14. Monnier, N., Lofstead, J., Lawson, M., Curry, M.: Profiling platform storage using io500 and mistral. In: 2019 IEEE/ACM Fourth International Parallel Data Systems Workshop (PDSW). pp. 60–73. IEEE (2019). https://doi.org/10.1109/ PDSW49588.2019.00011, https://doi.org/10.1109/PDSW49588.2019.00011 15. Schmuck, F., Haskin, R.: GPFS: A shared-disk file system for large computing clusters. In: Conference on File and Storage Technologies (FAST 02) (2002), https:// www.usenix.org/legacy/events/fast02/full_papers/schmuck/schmuck_html/ 16. Schwan, P., et al.: Lustre: Building a file system for 1000-node clusters. In: Proceedings of the 2003 Linux Symposium. vol. 2003, pp. 380–386 (2003), https: //www.landley.net/kdocs/mirror/ols2003.pdf 17. Shan, H., Shalf, J.: Using ior to analyze the i/o performance for hpc platforms (2007), https://escholarship.org/uc/item/9111c60j
14
J. Kunkel et al.
18. Snyder, S., Carns, P., Harms, K., Ross, R., Lockwood, G.K., Wright, N.J.: Modular HPC I/O characterization with Darshan. In: 2016 5th Workshop on Extreme-Scale Programming Tools (ESPT). pp. 9–17. IEEE (2016). https://doi.org/10.1109/ ESPT.2016.006, https://doi.org/10.1109/ESPT.2016.006 19. VI4IO Community: VI4IO: Virtual institute for I/O (2023), https://vi4io.org, accessed: 2025-10-03