Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets Yoonsik Hong1,∗ , Diego Klabjan1
arXiv:2606.25811v1 [q-fin.TR] 24 Jun 2026
Abstract Commodity futures can be represented hierarchically, with underlying assets at the upper level and individual futures contracts at the lower level. Entities at each level can be connected by edges reflecting inherent correlations, with cross-level edges capturing contract-to-underlying asset connections. Building on our observations of these structures, we propose a hierarchical graph learning approach for calendar spread (CS) strategies in commodity futures markets, addressing two significant gaps in the machine-learning literature: (i) the absence of learningbased methods for CS strategies in futures markets, and (ii) the lack of consideration of maturity-dependent interrelationships across commodity futures. We first establish the efficacy of CS strategies by analytically showing that CS strategies can possess higher risk-adjusted returns, measured by the information ratio, and lower risk, measured by variance and delta, than long-only strategies. We then introduce a method to convert learning-based predictions into CS positions. Next, we develop a hierarchical graph learning method that predicts futures price movements by utilizing the maturity-dependent interrelationships, thereby yielding a CS trading algorithm. Empirical results on commodity futures markets traded on the Chicago Mercantile Exchange Group demonstrate that our method outperforms benchmark models in both prediction and trading performance. We find that maturity-dependent interrelationships across commodity futures are instrumental in prediction and that CS trading based on hierarchical graph learning is effective for statistical arbitrage. Keywords: Calendar Spread Strategy, Graph Learning, Statistical Arbitrage, Commodity Futures, Deep Learning
1. Introduction We address the problem of exploiting statistical arbitrage opportunities in commodity futures markets through a calendar spread (CS) strategy based on hierarchical graph learning that captures information embedded in maturitydependent interrelationships across commodity futures. A futures contract is an agreement between two parties in which the contract price is fixed at initiation, and at maturity, one party delivers the underlying asset (or settles in cash), while the other pays the predetermined price (Hull, Treepongkaruna, Colwell, Heaney and Pitt, 2013). If the underlying asset is a commodity, the contract is referred to as a commodity futures contract. Statistical arbitrage refers to cash flows that generate profits by exploiting mispricing under acceptable risk (Lazzarino, Berrill, Šević et al., 2018), and deep-learning–based approaches for capturing such opportunities have been actively studied across various financial markets (e.g., Guijarro-Ordonez, Pelger and Zanotti (2025); Hong and Klabjan (2025a)). To exploit statistical arbitrage in commodity futures markets, we employ a CS strategy that takes long and short positions in futures contracts with the same underlying commodity but different times to maturity (TTMs). No prior learning-based studies address CS trading, despite its widespread use and extensive empirical studies in futures markets (Clarke, De Silva and Thorley, 2013; Szymanowska, De Roon, Nijman and Van Den Goorbergh, 2014; Boons and Prado, 2019; CME Group, n.d.b). Trading strategies in existing learning-based futures studies—such as single-contract directional trading and top-K selection—are highly likely to be exposed to underlying asset price risk. Such exposure to volatile underlying price movements, rather than relatively stable cost-of-carry components (e.g., interest rates and storage costs), can deteriorate risk-adjusted trading performance. Motivated by this limitation, ∗ Corresponding author
1 Department of Industrial Engineering and Management Sciences, Northwestern University, Evanston, 60208, IL, USA
[email protected] (Y. Hong); [email protected] (D. Klabjan)
Y. Hong and D. Klabjan, Preprint, June 25, 2026
1
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
we show that CS strategies can mitigate risk and yield superior risk-adjusted returns. We also introduce a method to convert learning-based predictions into CS positions. Most existing learning-based studies overlook TTM-dependent interrelationships across commodity futures. Commodity futures contracts are economically linked (Chng, 2009; Vacha, Janda, Kristoufek and Zilberman, 2013; Cai, Fang, Chang, Tian and Hamori, 2020; Furlong and Ingenito, 1996; Hammoudeh, Sari and Ewing, 2009; Boakye, Heimonen and Junttila, 2024; Liu, Lu, Zong and Xie, 2023), but this structure is rarely incorporated into learning-based models. Although some recent studies (Hu, Tan, Liu and Yin, 2025; Tan, Hu, Liu and Yin, 2024) adopt graph learning to model relationships among commodity futures, they do not distinguish these relationships by TTM. However, these relationships are inherently TTM-dependent (Brennan, 1976; Gibson and Schwartz, 1990; Schwartz, 1997; Casassus and Collin-Dufresne, 2005). For example, gasoline and diesel inventories may remain low for three months due to a war, but they are expected to recover after six months as supply chains recover. When predicting the price of diesel futures with a six-month TTM, the price information of gasoline futures with a three-month TTM should not be treated in the same way as that of gasoline futures with a six-month TTM. Ignoring TTM thus limits the ability of graph-learning models to capture TTM-dependent information. To address this gap, we propose a hierarchical graph learning approach. First, we verify the efficacy of CS strategies in terms of risk and risk-adjusted return. We theoretically establish their advantages over long-only (LO) strategies. Since any trading position can be expressed as a linear combination of CS and LO positions, we compare the two strategies. We show that CS strategies possess lower variance and risk exposure than LO strategies under some assumptions. Moreover, we show that if an LO strategy outperforms an uninformed strategy by a certain threshold, then a CS strategy derived from the LO strategy outperforms the LO strategy in terms of risk-adjusted return. We also present a method to assess whether the assumptions are realistic. We further introduce a method to map learning-based predictions into CS positions. Next, we develop a hierarchical graph learning approach that captures maturity-dependent relationships across commodity futures. We construct a bi-level graph in which underlying commodities form upper-level nodes and individual futures contracts form lower-level nodes, connected via cross-commodity, commodity-contract, and crosscontract relationships. Since commodities do not share a common contract maturity grid, we introduce a graph lifting method that constructs TTM-aligned virtual futures nodes for each commodity on a virtual TTM grid shared across commodities. On this lifted graph, we design a bi-level convolution architecture that alternates between (i) propagating information across commodities through virtual futures at the same TTM (TTM-aligned propagation) and (ii) capturing intra-commodity term structure dynamics via message passing over neighboring maturities (differentTTM propagation). By distinguishing relationships based on TTM through distinct parameterizations and iteratively integrating information across both dimensions, the proposed method captures interrelationships across commodity futures while explicitly accounting for maturity dependence. Our experiments demonstrate that the proposed method generates statistical arbitrage and that both CS strategies and maturity-dependent interrelationships are effective. We provide evidence that CS strategies can exhibit lower variance, lower exposure risk, and higher IRs than LO strategies in commodity futures markets. Our hierarchical graph learning achieves the lowest mean squared error with statistical significance, and its CS trading attains the best risk-adjusted returns by capturing TTM-dependent interrelationships among commodity futures, compared to other learning-based methods that either ignore such interrelationships or do not account for their maturity dependence. This paper makes the following main contributions. • To the best of our knowledge, this study is the first to propose a hierarchical graph learning approach for predicting commodity futures price movements while accounting for maturity-dependent interrelationships across commodity futures. • We establish the efficacy of calendar spread strategies by presenting propositions showing that they can exhibit higher information ratios, lower variance, and lower exposure risk than long-only strategies in commodity futures markets. • Empirical results show that the hierarchical graph learning method outperforms benchmark methods with statistical significance in prediction tasks. We find that maturity-dependent interrelationships across commodity futures are instrumental in predictive performance. • Our calendar spread strategy attains daily information and Sortino ratios of 0.0846 and 0.1241, corresponding to improvements of 96% and 82% over the benchmarks, and 75% and 113% over the S&P 500, respectively. These Y. Hong and D. Klabjan, Preprint, June 25, 2026
2
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
results indicate that calendar spread trading based on hierarchical graph learning in futures markets is effective for statistical arbitrage.
2. Related Work No learning-based studies address CS strategies in futures markets. Existing work applies various learning-based methods to prediction and trading, including LSTM and GRU models (Li, Li, Liu, Zhu and Wei, 2022; Zhang, Wang and Wang, 2020; Hung and Chen, 2021), econometric regression (Wang and Zhang, 2024; Angelidis, Sakkas and Tessaromatis, 2025), reinforcement learning (Gabrielsson and Johansson, 2015; Massahi and Mahootchi, 2024; Kaur, Sidhu and Bhathal, 2025), clustering-based approaches (Fengqian and Chao, 2020), and AI-agent systems incorporating vision-language models (Wang, Wu, Zhang, Tan, Chen and Lin, 2025). These studies span agricultural futures (Li et al., 2022; Liu et al., 2023; Wang et al., 2025; Fengqian and Chao, 2020), energy futures (Zhang et al., 2020; Sun, Wang, Ni, Cao and Liu, 2019; Hu et al., 2025; Baruník and Malinska, 2016), metal futures (Massahi and Mahootchi, 2024; Kaur et al., 2025; Fengqian and Chao, 2020; Sun et al., 2019; Tan et al., 2024), equity index futures (Hung and Chen, 2021; Gabrielsson and Johansson, 2015; Fengqian and Chao, 2020), and broad cross-market commodity universes (Wang and Zhang, 2024; Angelidis et al., 2025). Many learning-based studies focus solely on price or return prediction (Li et al., 2022; Zhang et al., 2020; Hung and Chen, 2021), while trading-oriented work considers single-contract directional trading (Gabrielsson and Johansson, 2015; Fengqian and Chao, 2020; Wang et al., 2025; Massahi and Mahootchi, 2024; Kaur et al., 2025; Sun et al., 2019), top-K trading (Hu et al., 2025), cross-commodity long-short trading (Wang and Zhang, 2024; Angelidis et al., 2025), or cross-commodity pairs trading (Liu et al., 2023). However, none of the prior learning-based work addresses CS strategies. We conjecture that a key reason is the lack of clear theoretical guarantees on the efficacy of CS strategies, despite their widespread empirical use in futures markets (Clarke et al., 2013; Szymanowska et al., 2014; Boons and Prado, 2019; CME Group, n.d.b). Accordingly, we provide theoretical justification and demonstrate that CS trading based on hierarchical graph learning is effective compared with benchmark methods. Existing learning-based studies largely overlook TTM-dependent interrelationships across commodity futures. Futures contracts across different commodities are economically interconnected through production chains (Chng, 2009), substitution and complementarity effects (Vacha et al., 2013), and shared macroeconomic drivers (Cai et al., 2020; Furlong and Ingenito, 1996; Hammoudeh et al., 2009; Boakye et al., 2024; Liu et al., 2023) (e.g., oil and agricultural markets (Ji and Fan, 2012; Mitchell, 2008; Nazlioglu, 2011)), yet most learning-based approaches treat contracts in isolation. Although a few recent studies (Hu et al., 2025; Tan et al., 2024) incorporate cross-futures relationships through graph learning, they treat commodity futures contracts as generic nodes and construct connections without conditioning on TTM. However, interrelationships across commodity futures are inherently TTM-dependent, as contracts with different TTMs exhibit heterogeneous sensitivities to inventory conditions, short-term supply disruptions, production constraints, and long-run demand and substitution dynamics (Brennan, 1976; Gibson and Schwartz, 1990; Schwartz, 1997; Casassus and Collin-Dufresne, 2005). Ignoring TTM thus conflates distinct economic channels and limits the ability of graph architectures to capture TTM-dependent relationships across commodity futures. Their methods (Hu et al., 2025; Tan et al., 2024) also cannot address node birth and death, as their focus on high-frequency trading assumes a fixed contract universe. In contrast, we consider daily trading, where node birth and death arise naturally from contract maturities. Through our hierarchical graph learning method, we highlight the importance of TTM dependence in commodity futures interrelationships and demonstrate the instrumental role of such interrelationships.
3. Notation and Preliminaries For a finite set and vectors 𝐱𝑖 ∈ ℝ𝑝 for 𝑖 ∈ , let [𝐱𝑖 ]𝑖∈ ∈ ℝ||×𝑝 denote the row-wise stacked matrix in lexicographical order of . For an index set , if 𝑥𝑖 ∈ ℝ is defined for 𝑖 ∈ ′ ⊂ and not defined for 𝑖 ∈ ⧵ ′ , we let sum(𝑥𝑖 ; 𝑖 ∈ ), avg(𝑥𝑖 ; 𝑖 ∈ ) and std(𝑥𝑖 ; 𝑖 ∈ ) denote the sum, average, and standard deviation over ′ , respectively. We denote by 𝕀{⋅} the indicator function. For 𝑎 ≤ 𝑏, let [𝑎 ∶ 𝑏], [𝑎 ∶ 𝑏), (𝑎 ∶ 𝑏], and (𝑎 ∶ 𝑏) denote the corresponding ∑ ̂ (𝑥; 𝐱) = 1 ∑𝑛 𝕀{𝑥 ≤𝑥} integer intervals. For 𝐱 = [𝑥𝑖 ]𝑖∈[1∶𝑛] ∈ ℝ𝑛 , we define cdf (𝑥; 𝐱) = 1𝑛 𝑛𝑖=1 𝕀{𝑥𝑖 ≤𝑥} and cdf 𝑖=1 𝑖 𝑛+𝜖 where 𝜖 > 0 is a small constant. For any random variable 𝑋, let 𝑋̃ denote a realization and 𝑋̂ denote its mean estimate. Let 𝟏 denote a vector of ones. For 𝑥 ∈ ℝ, define [𝑥]+ = max{0, 𝑥}. We let 𝑊⋅ and 𝐛⋅ denote a real-valued Y. Hong and D. Klabjan, Preprint, June 25, 2026
3
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
learnable parameter matrix and a real-valued learnable bias vector, respectively, where the subscripts indicate their indices, and let 𝜙 denote an activation function. We index trading dates by 𝑡 ∈ ℕ in chronological order, without skipped integers. Let ⊂ ℕ denote the index set of commodities, with max = ||. For each 𝑐 ∈ , let 𝑐 ⊂ ℕ denote the index set of futures contracts with underlying commodity 𝑐, ordered by increasing maturity, and satisfying max 𝑐 = |𝑐 |. We denote a futures contract by (𝑐, 𝑑) ∈ × 𝑐 with maturity 𝑇𝑐𝑑 . We denote by 𝑐 = {𝑇𝑐𝑑 ∶ 𝑑 ∈ 𝑐 } the maturity grid of commodity 𝑐 ∈ . We refer to 𝑇𝑐𝑑 − 𝑡 as the TTM of contract (𝑐, 𝑑) at time 𝑡 ∈ ℕ. For 𝑡 ∈ ℕ and 𝑐 ∈ , we let 𝑆𝑡𝑐 > 0 denote the price of commodity 𝑐 at time 𝑡, with return 𝑅𝑡𝑐 = 𝑆𝑡𝑐 ∕𝑆𝑡−1,𝑐 − 1. We assume that 𝑆𝑡𝑐 is defined for all 𝑡 ∈ ℕ and 𝑐 ∈ . For 𝑡 ∈ ℕ and (𝑐, 𝑑) ∈ × 𝑐 , let 𝑉𝑡𝑐𝑑 ≥ 0 denote the trading volume of futures contract (𝑐, 𝑑) at time 𝑡, and let 𝐹𝑡𝑐𝑑 > 0 denote its price at time 𝑡 for all 𝑡 ≤ 𝑇𝑐𝑑 . Assume that the margin ratio is one for every position (long or short) across all contracts and time, and that there are no transaction costs. The return of futures contract (𝑐, 𝑑) is then defined as 𝑅𝑡𝑐𝑑 = 𝐹𝑡𝑐𝑑 ∕𝐹𝑡−1,𝑐𝑑 − 1 when 𝑡 ≤ 𝑇𝑐𝑑 . We assume that all 𝑆𝑡𝑐 , 𝑅𝑡𝑐 , 𝑉𝑡𝑐𝑑 , 𝐹𝑡𝑐𝑑 , and 𝑅𝑡𝑐𝑑 are random variables, and that, for all (𝑐, 𝑑) ∈ × 𝑐 and 𝑡 ≤ 𝑇𝑐𝑑 , 𝔼|𝑅𝑡𝑐𝑑 | < ∞, 𝔼𝑅2𝑡𝑐𝑑 < ∞, and Var(𝑅𝑡𝑐𝑑 ) > 0.
(1)
We define the history at time 𝑡 as 𝑡 = {(𝑡′ , 𝐻̃ 𝑡′ ) ∶ 𝑡′ ∈ [1 ∶ 𝑡]}, where 𝐻̃ 𝑡′ = {(𝑐, 𝑑, 𝑉̃𝑡′ 𝑐𝑑 , 𝐹̃𝑡′ 𝑐𝑑 ) ∶ (𝑐, 𝑑) ∈ × 𝑐 , 𝑉̃𝑡′ 𝑐𝑑 > 0} is an observable sample at 𝑡′ . The price of a futures contract (𝑐, 𝑑) ∈ × 𝑐 at time 𝑡 is expressed as (2)
𝐹𝑡𝑐𝑑 = 𝑆𝑡𝑐 exp(𝑄𝑡𝑐𝑑 (𝑇𝑐𝑑 − 𝑡)) > 0 where the random variable 𝑄𝑡𝑐𝑑 ∈ ℝ is the net cost-of-carry rate. Its return is then expressed as
(3)
𝑅𝑡𝑐𝑑 = (1 + 𝑅𝑡𝑐 ) exp(𝐴𝑡𝑐𝑑 ) − 1
where 𝐴𝑡𝑐𝑑 = 𝑄𝑡𝑐𝑑 (𝑇𝑐𝑑 − 𝑡) − 𝑄𝑡−1,𝑐𝑑 (𝑇𝑐𝑑 − 𝑡 + 1) represents the change in the carry cost from 𝑡 − 1 to 𝑡. The spot delta of (𝑐, 𝑑) is Δ𝑡𝑐𝑑 =
𝜕𝐹𝑡𝑐𝑑 = exp(𝑄𝑡𝑐𝑑 (𝑇𝑐𝑑 − 𝑡)) > 0, 𝜕𝑆𝑡𝑐
(4)
which represents the risk exposure to the underlying commodity price (Hull et al., 2013). We define the set of contracts observable from 𝑡−1 to 𝑡 as 𝑈𝑡 = {(𝑐, 𝑑) ∈ ×𝑐 ∶ 𝑉𝑡𝑐𝑑 𝑉𝑡−1,𝑐𝑑 > 0} with 𝑛𝑡 = |𝑈𝑡 |. Let 𝐶𝑡 = {𝑐 ∶ (𝑐, 𝑑) ∈ 𝑈𝑡 } and 𝐷𝑡𝑐 = {𝑑 ∶ (𝑐, 𝑑) ∈ 𝑈𝑡 } denote the corresponding sets of commodities and contracts, respectively, with 𝑛𝑡𝑐 = |𝐷𝑡𝑐 |. We denote the realization of 𝑈𝑡 by 𝑈̃ 𝑡 = {(𝑐, 𝑑) ∈ × 𝑐 ∶ 𝑉̃𝑡𝑐𝑑 𝑉̃𝑡−1,𝑐𝑑 > 0}, and define 𝐶̃𝑡 , 𝐷̃ 𝑡𝑐 , 𝑛̃𝑡 , and 𝑛̃𝑡𝑐 , analogously. We denote the trading universe at time 𝑡 by 𝑈̂ 𝑡 ⊂ {(𝑐, 𝑑) ∈ × 𝑐 ∶ 𝑡 ≤ 𝑇𝑐𝑑 }, and define 𝐶̂𝑡 , 𝐷̂ 𝑡𝑐 , 𝑛̂ 𝑡 , and 𝑛̂ 𝑡𝑐 analogously. Note that 𝑈̂ 𝑡 is determined at 𝑡 − 2. In the following, we consider a single investor implementing our method.
4. Proposed Method To exploit statistical arbitrage in commodity futures markets, we propose a hierarchical graph learning method to predict futures price movements for CS trading. We first formally define a CS strategy, establish its efficacy through theoretical propositions, and introduce a projection method that maps predictions into CS positions. We then formulate the prediction problem for CS trading and develop a hierarchical graph learning architecture that captures TTMdependent interrelationships among commodity futures.
4.1. Calendar Spread Strategy For 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶̂𝑡 , we define a CS strategy for 𝑐 as a trading strategy that yields a position weight vector 𝐰𝑡𝑐 = [𝑤𝑡𝑐𝑑 ]𝑑∈𝐷̂ ∈ ℝ𝑛̂𝑡𝑐 such that 𝟏⊤ 𝐰𝑡𝑐 = 0 and ‖𝐰𝑡𝑐 ‖1 = 1. This implies symmetric long and short positions: 𝑡𝑐
∑
∑
1 . 2
(5)
Y. Hong and D. Klabjan, Preprint, June 25, 2026
4
𝑑∈𝐷̂ 𝑡𝑐
[𝑤𝑡𝑐𝑑 ]+ =
𝑑∈𝐷̂ 𝑡𝑐
[−𝑤𝑡𝑐𝑑 ]+ =
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
We let 𝑡𝑐,CS = {𝐰𝑡𝑐 ∈ ℝ𝑛̂𝑡𝑐 ∶ 𝟏⊤ 𝐰𝑡𝑐 = 0, ‖𝐰𝑡𝑐 ‖1 = 1} denote the set of CS weights for 𝑐 at 𝑡, and 𝑡𝑐,LO = {𝐰𝑡𝑐 ∈ ℝ𝑛̂𝑡𝑐 ∶ 𝑤𝑡𝑐𝑑 ≥ 0, ‖𝐰𝑡𝑐 ‖1 = 1} denote the set of LO weights for 𝑐 at 𝑡. Since any position can be expressed as a linear combination of CS and LO positions, we focus on these two classes. We evaluate both classes of strategies in terms of their risk and risk-adjusted return. For 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶̂𝑡 , a position weight vector 𝐰𝑡𝑐 = [𝑤𝑡𝑐𝑑 ]𝑑∈𝐷̂ ∈ ℝ𝑛̂𝑡𝑐 is determined by an investor at 𝑡 − 2, built in markets at 𝑡 − 1. Then, 𝑡𝑐 the resulting position is cleared at 𝑡, yielding the return given by ∑ Π(𝐰𝑡𝑐 ) = 𝑤𝑡𝑐𝑑 𝑅𝑡𝑐𝑑 (6) 𝑑∈𝐷̂ 𝑡𝑐
Note that Π(𝐰𝑡𝑐 ) depends on 𝑅𝑡𝑐𝑑 ; for simplicity, we omit this dependence. Moreover, we define the return variance √ of 𝐰𝑡𝑐 as Var(Π(𝐰𝑡𝑐 )), the IR of 𝐰𝑡𝑐 as IR(Π(𝐰𝑡𝑐 )) = 𝔼(Π(𝐰𝑡𝑐 ))∕ Var(Π(𝐰𝑡𝑐 )) if Var(Π(𝐰𝑡𝑐 )) > 0, and the delta of 𝐰𝑡𝑐 as ∑ ∑ 𝑤𝑡𝑐𝑑 𝐹𝑡−1,𝑐𝑑 ) = 𝑤𝑡𝑐𝑑 Δ𝑡−1,𝑐𝑑 . (7) Δ(𝐰𝑡𝑐 ) = 𝜕𝑆 𝜕 ( 𝑡−1,𝑐
𝑑∈𝐷̂ 𝑡𝑐
𝑑∈𝐷̂ 𝑡𝑐
Risk is commonly measured by return variance (Markowitz, 1952) or by delta (Hull et al., 2013), and investment objectives are often framed as maximizing risk-adjusted return (Sharpe, 1966, 1998). Thus, we analyze LO and CS strategies under these metrics in Propositions 1 to 3, with proofs in Appendix Section A. Proposition 1 implies that CS strategies can have lower variance than LO strategies, thereby entailing lower risk. Since the variance of 𝑅𝑡𝑐𝑑 is driven more by the underlying commodity return 𝑅𝑡𝑐 than by the cost-of-carry change 𝐴𝑡𝑐𝑑 in Eq. (3) (Gospodinov and Ng, 2013; Maslyuk and Smyth, 2009), returns of contracts with the same underlying commodity tend to be positively correlated. Based on this property, we assume that such correlations admit nonnegative lower bounds in Proposition 1. Part (i) shows that an LO position has higher variance than all CS positions when correlations exceed a threshold depending on the LO position. Part (ii) removes this dependence and shows that all CS positions have lower variance than all LO positions when correlations exceed a bound determined by the variance range ratio and the number of contracts. ′ ̂ ̂ Proposition √ 1. Fix 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶𝑡 . For 𝑑, 𝑑 ∈ 𝐷𝑡𝑐 , let 𝜎𝑡𝑐𝑑𝑑 ′ = Cov(𝑅𝑡𝑐𝑑 , 𝑅𝑡𝑐𝑑 ′ ), 𝜌𝑡𝑐𝑑𝑑 ′ = Corr(𝑅𝑡𝑐𝑑 , 𝑅𝑡𝑐𝑑 ′ ) = 𝜎𝑡𝑐𝑑𝑑 ′ ∕ 𝜎𝑡𝑐𝑑𝑑 𝜎𝑡𝑐𝑑𝑑 ′ , and 𝜅𝑡𝑐2 = (max𝑑∈𝐷̂ 𝜎𝑡𝑐𝑑𝑑 )∕(min𝑑∈𝐷̂ 𝜎𝑡𝑐𝑑𝑑 ). 𝑡𝑐 𝑡𝑐 (i) If min ′ ̂ 𝜌𝑡𝑐𝑑𝑑 ′ ≥ 0 and 𝐰 ∈ 𝑡𝑐,LO (𝐷̂ 𝑡𝑐 ) satisfies 𝑑,𝑑 ∈𝐷𝑡𝑐
min 𝜌𝑡𝑐𝑑𝑑 ′ > (𝜅𝑡𝑐2 − 2‖𝐰‖22 )(3 − 2‖𝐰‖22 )−1 ,
𝑑,𝑑 ′ ∈𝐷̂ 𝑡𝑐
(8)
then for all 𝐰′ ∈ 𝑡𝑐,CS (𝐷̂ 𝑡𝑐 ), Var(Π(𝐰)) > Var(Π(𝐰′ )). (ii) If min𝑑,𝑑 ′ ∈𝐷̂ 𝜌𝑡𝑐𝑑𝑑 ′ ≥ 0 and 𝑡𝑐
min 𝜌𝑡𝑐𝑑𝑑 ′ > (𝜅𝑡𝑐2 𝑛̂ 𝑡𝑐 − 2)(3𝑛̂ 𝑡𝑐 − 2)−1 ,
𝑑,𝑑 ′ ∈𝐷̂ 𝑡𝑐
(9)
then for all (𝐰, 𝐰′ ) ∈ 𝑡𝑐,LO (𝐷̂ 𝑡𝑐 ) × 𝑡𝑐,CS (𝐷̂ 𝑡𝑐 ), Var(Π(𝐰)) > Var(Π(𝐰′ )). Proposition 2 shows that if the expected return of an LO strategy 𝐰 exceeds that of the uninformed equal-weight strategy 𝐰̄ by a threshold, then the CS strategy 𝐰̌ derived from 𝐰 attains a positive IR that exceeds that of the LO strategy. The CS weight vector 𝐰̌ is obtained by demeaning and rescaling 𝐰, thereby preserving the relative ordering of ̌ ≥ 0 in part (ii), the first inequality in each part ensures weights while satisfying the CS constraints. Because Var(Π(𝐰)) that Π(𝐰) has nonzero variance. Since 𝐰 ≥ 0, this nonzero-variance condition typically holds when contracts with the same underlying commodity are positively correlated, as is commonly observed in the market (see Proposition 1). Moreover, the first inequality in part (ii) is satisfied under Proposition 1, for instance. The second inequality in parts (i) ̄ under and (ii) excludes the degenerate case addressed in part (iii). Each part specifies a threshold on 𝔼Π(𝐰) − 𝔼Π(𝐰) different conditions: the variance ratio in (i), the position difference in (ii), and zero in (iii). In part (iii), 𝐰̌ yields a risk-free arbitrage, which can be viewed as having an infinite IR, and in parts (i) and (ii), the CS strategy achieves a positive IR that exceeds that of the LO strategy. Therefore, if an LO strategy exceeds the uninformed benchmark beyond Y. Hong and D. Klabjan, Preprint, June 25, 2026
5
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
the corresponding threshold, investing in the LO position is inferior to the derived CS position. Since CS positions are dollar-neutral, their Sharpe and information ratios coincide, whereas the Sharpe benchmark for LO positions is the risk-free rate (Sharpe, 1966, 1998). Thus, Proposition 2 also implies that CS strategies can attain higher Sharpe ratios than LO strategies. ̄ ̄ 1 ∈ 𝑡𝑐,CS Proposition 2. Fix 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶̂𝑡 . Let 𝐰 ∈ 𝑡𝑐,LO , 𝐰̄ = 𝟏∕𝑛̂ 𝑡𝑐 ∈ 𝑡𝑐,LO , and 𝐰̌ = (𝐰 − 𝐰)∕‖𝐰 − 𝐰‖ ̄ with 𝐰 ≠ 𝐰. ̄ > 0, and (i) If Var(Π(𝐰)) > 0, Var(Π(𝐰 − 𝐰)) √ ̄ ̄ > [𝔼Π(𝐰)]+ Var(Π(𝐰 − 𝐰))∕ Var(Π(𝐰)), (10) 𝔼Π(𝐰) − 𝔼Π(𝐰) ̌ > [IR(Π(𝐰))]+ . then IR(Π(𝐰)) ̌ Var(Π(𝐰 − 𝐰)) ̄ > 0, and (ii) If Var(Π(𝐰)) > Var(Π(𝐰)), (11)
̄ > ‖𝐰 − 𝐰‖ ̄ 1 [𝔼Π(𝐰)]+ , 𝔼Π(𝐰) − 𝔼Π(𝐰) ̌ > [IR(Π(𝐰))]+ . then IR(Π(𝐰)) ̄ = 0, and (iii) If Var(Π(𝐰)) > 0, Var(Π(𝐰 − 𝐰))
(12)
̄ > 0, 𝔼Π(𝐰) − 𝔼Π(𝐰) ̌ > 0, Var(Π(𝐰)) ̌ = 0 and IR(Π(𝐰)) < ∞. then 𝔼Π(𝐰)
Proposition 3 implies that CS strategies can have lower risk exposure to the underlying commodity price. Part (i) ̌ obtained by demeaning an LO vector 𝐰 and scaling, has lower gives a condition under which the CS weight vector 𝐰, ̄ ̌ ̄ as both are delta. In particular, if ‖𝐰 − 𝐰‖1 ≥ 1, then 𝐰 always has lower risk exposure regardless of Δ(𝐰) and Δ(𝐰), max min positive by Eq. (4). Under the assumption that the delta range ratio, Δ𝑡−1,𝑐 ∕Δ𝑡−1,𝑐 , is less than three, part (ii) implies that, almost surely, every CS strategy is less exposed to the underlying commodity price than every LO strategy. ̂ Proposition 3. Fix 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶. ̄ ̄ 1 ∈ 𝑡𝑐,CS with 𝐰 ≠ 𝐰. ̄ If Δ(𝐰) ̄ > (i) Let 𝐰 ∈ 𝑡𝑐,LO , 𝐰̄ = 𝟏∕𝑛̂ 𝑡𝑐 ∈ 𝑡𝑐,LO , and 𝐰̌ = (𝐰 − 𝐰)∕‖𝐰 − 𝐰‖ ̄ 1 )Δ(𝐰), then Δ(𝐰) ̌ < Δ(𝐰). (1 − ‖𝐰 − 𝐰‖ (ii) Assume that Δ𝑡−1,𝑐𝑑 is an a.s. finite random variable for every 𝑑 ∈ 𝐷̂ 𝑡𝑐 . Let Δmin = min{ess inf Δ𝑡−1,𝑐𝑑 ∶ 𝑑 ∈ 𝑡−1,𝑐 max max min ̂ ̂ 𝐷𝑡𝑐 } ∈ ℝ and Δ = max{ess sup Δ𝑡−1,𝑐𝑑 ∶ 𝑑 ∈ 𝐷𝑡𝑐 } ∈ ℝ. If 3Δ >Δ , then 𝑡−1,𝑐
𝑡−1,𝑐
𝑡−1,𝑐
(13)
ess inf Δ(𝐰) > ess sup Δ(𝐰)
𝐰∈𝑡𝑐,LO
𝐰∈tc,CS
We assess the realism of the inequality assumptions in the second parts of Propositions 1 and 3 by estimating the relevant parameters in the trading universe 𝑈̂ 𝑡 . Since 𝑈̂ 𝑡 is determined based on 𝑡−2 , as discussed below in Eq. (25), some contracts in 𝑈̂ 𝑡 may not be observable at 𝑡 − 1 or 𝑡, i.e., 𝑉̃𝑡−1,𝑐𝑑 𝑉̃𝑡𝑐𝑑 = 0. Thus, in the propositions, we replace 𝐷̂ 𝑡𝑐 , ′ =𝐷 ̂ 𝑡𝑐 ∩ 𝐷̃ 𝑡𝑐 , 𝑛̂ ′ = |𝐷̂ ′ |, and 𝐶̂ ′ = {𝑐 ∈ 𝐶̂𝑡 ∩ 𝐶̃𝑡 ∶ 𝑛̂ ′ > 1}, respectively. For a CS strategy to be 𝑛̂ 𝑡𝑐 , and 𝐶̂𝑡 with 𝐷̂ 𝑡𝑐 𝑡𝑐 𝑡𝑐 𝑡 𝑡𝑐 implemented, at least two contracts on the same underlying asset are required; thus, we impose the constraint 𝑛̂ ′𝑡𝑐 > 1 in ′ ×𝐷 ̂ ′ , we estimate 𝜎̂ 𝑡𝑐𝑑𝑑 ′ as the sample covariance of {(𝑅̃ 𝑡′ 𝑐𝑑 , 𝑅̃ 𝑡′ 𝑐𝑑 ′ ) ∶ 𝐶̂𝑡′ . For each 𝑡 ∈ ℕ, 𝑐 ∈ 𝐶̂𝑡′ , and (𝑑, 𝑑 ′ ) ∈ 𝐷̂ 𝑡,𝑐 𝑡,𝑐 ′ min ̂ ̃ 𝑡 ∈ [𝑡 − 𝑛sam ∶ 𝑡]}, assuming 𝜎̂ 𝑡𝑐𝑑𝑑 ′ ≈ 𝜎̂ 𝑡′ 𝑐𝑑𝑑 ′ over 𝑡′ ∈ [𝑡 − 𝑛min sam ∶ 𝑡). Since 𝑈𝑡 and 𝑈𝑡 ensure price data availability on min [𝑡−𝑛sam −1 ∶ 𝑡−2]—discussed in Eq. (28)—and [𝑡−1 ∶ 𝑡], respectively, we use returns over [𝑡−𝑛min sam ∶ 𝑡] for estimation. √ We then estimate 𝜌̂𝑡𝑐𝑑𝑑 ′ = 𝜎̂ 𝑡𝑐𝑑𝑑 ′ ∕ 𝜎̂ 𝑡𝑐𝑑𝑑 𝜎̂ 𝑡𝑐𝑑𝑑 ′ , min𝑑,𝑑 ′ ∈𝐷̂ ′ 𝜌̂𝑡𝑐𝑑𝑑 ′ , and 𝜅̂ 𝑡𝑐2 = (max𝑑∈𝐷̂ ′ 𝜎̂ 𝑡𝑐𝑑𝑑 )∕(min𝑑∈𝐷̂ ′ 𝜎̂ 𝑡𝑐𝑑𝑑 ). For 𝑡𝑐 𝑡𝑐 𝑡𝑐 ̂ max ∕Δ ̂ min = max{𝐹̃𝑡−1,𝑐𝑑 ∕𝐹̃𝑡−1,𝑐𝑑 ′ ∶ 𝑑, 𝑑 ′ ∈ 𝐷̂ ′ } from Eqs. (2) and (4). We each 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶̂𝑡 , we estimate Δ 𝑡−1,𝑐
𝑡−1,𝑐
then check the following inequalities for each 𝑡 and 𝑐:
𝑡𝑐
𝜌̂𝑡𝑐,min − (𝜅̂ 2 𝑛̂ ′𝑡𝑐 − 2)(3𝑛̂ ′𝑡𝑐 − 2)−1 > 0,
(14)
𝜌̂𝑡𝑐,min ≥ 0, ̂ max ∕Δ ̂ min < 3, Δ
(15)
𝑡𝑐
𝑡𝑐
Y. Hong and D. Klabjan, Preprint, June 25, 2026
(16) 6
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
where Eqs. (14) and (15) correspond to (ii) in Proposition 1, and Eq. (16) corresponds to 3Δmin > Δmax in 𝑡−1,𝑐 𝑡−1,𝑐 Proposition 3. Since CS strategies can exhibit lower risk and higher IR yet remain largely unexplored in the machine-learning literature, we focus on CS strategies and extend the framework from a single commodity to multiple commodities. We define a CS strategy for 𝑈̂ 𝑡 as a strategy that yields a position weight vector 𝐰𝑡 = [𝑤𝑡𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ ℝ𝑛̂𝑡 satisfying 𝑡 ‖𝐰𝑡 ‖1 = 1 and (17)
𝟏⊤ 𝐰𝑡𝑐 = 0, ∀𝑐 ∈ 𝐶̂𝑡 .
Unlike the single-commodity definition, the 𝓁1 -norm constraint is imposed on 𝐰𝑡 , rather than on each 𝐰𝑡𝑐 . A CS strategy for 𝑈̂ 𝑡 is a convex combination of the weight vectors of CS strategies, one for each 𝑐 ∈ 𝐶̂𝑡 , after extending their dimensions to ℝ𝑛̂𝑡 . Given a prediction vector 𝐘̂ 𝑡 = [𝑌̂𝑡𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ ℝ𝑛̂𝑡 , we construct a CS weight vector 𝐰𝑡 for 𝑈̂ 𝑡 using a projection 𝑡 method adapted from Hong and Klabjan (2025b), where 𝑌̂𝑡𝑐𝑑 is discussed below. Specifically, we compute 𝐰𝑡 = 𝐘̂ ◦𝑡 ∕‖𝐘̂ ◦𝑡 ‖1
(18)
𝐘̂ ◦𝑡 = [𝑌̂𝑡𝑐𝑑 − 𝑌̂𝑡𝑐∙ ](𝑐,𝑑)∈𝑈̂ , 𝑡 ∑ 1 𝑌̂ ′ . 𝑌̂𝑡𝑐∙ = 𝑛̂ 𝑡𝑐 ′ ̂ 𝑡𝑐𝑑
(19)
where
(20)
𝑑 ∈𝐷𝑡𝑐
That is, we center the prediction vector by subtracting the commodity-wise average 𝑌̂𝑡𝑐∙ in Eq. (19), and then scale it in Eq. (18). This computation is equivalent to projecting 𝐘̂ 𝑡 onto {𝐰𝑡 ∶ 𝟏⊤ 𝐰𝑡𝑐 = 0, ∀𝑐 ∈ 𝐶̂𝑡 } under the Euclidean norm and then scaling; see Appendix Section B. Thus, 𝐰𝑡 satisfies 𝟏⊤ 𝐰𝑡𝑐 = 0, ∀𝑐 ∈ 𝐶̂𝑡 and ‖𝐰𝑡 ‖1 = 1. Since projection yields the closest vector satisfying the constraints, the resulting vector from Eqs. (18) to (20) is, up to scaling, the CS weight vector closest to the prediction vector. We consider a daily trading scheme with daily predictions and periodic retraining. Given 𝑛f it ∈ ℕ, we let {𝑡𝑘 ∶ 𝑘 ∈ [1 ∶ 𝑛f it ]} denote the retraining dates with 𝑡𝑘 < 𝑡𝑘+1 and 𝑡𝑛f it +1 = ∞. For each 𝑘 ∈ [1 ∶ 𝑛f it ], a prediction model 𝑓𝑘 is trained on 𝑡𝑘 when 𝑡𝑘 becomes available, i.e., after the markets close on 𝑡𝑘 and before they open on 𝑡𝑘 + 1. On each trading date 𝑡 ∈ 𝑘test = [𝑡𝑘 ∶ 𝑡𝑘+1 ), we predict price movements 𝐘𝑡+2 = [𝑌𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ ℝ𝑛̂𝑡+2 𝑡+2 from 𝑡 + 1 to 𝑡 + 2 as 𝐘̂ 𝑡+2 = [𝑌̂𝑡+2,𝑐𝑑 ] = 𝑓𝑘 (𝑡 ) ∈ ℝ𝑛̂𝑡+2 where 𝑌𝑡+2,𝑐𝑑 is defined below. Based on these ̂ (𝑐,𝑑)∈𝑈𝑡+2
predictions, CS position vector 𝐰𝑡+2 = [𝑤𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ ℝ𝑛̂𝑡+2 is computed via Eqs. (18) to (20). The prediction 𝑡+2 and CS position vectors are generated after the markets close at 𝑡 (or after training if 𝑡 = 𝑡𝑘 ) and before they open at 𝑡 + 1. In the markets, the CS position vector is then established at 𝑡 + 1 and liquidated at 𝑡 + 2, with trades executed at prices 𝐹𝑡+1,𝑐𝑑 and 𝐹𝑡+2,𝑐𝑑 , respectively. No datum in 𝐻𝑡′ with 𝑡′ ≥ 𝑡𝑘 + 1 is used to train 𝑓𝑘 , and no datum in 𝐻𝑡′ with 𝑡′ ≥ 𝑡 + 1 is used to produce (𝐘̂ 𝑡+2 , 𝐰𝑡+2 ), thereby precluding look-ahead bias.
4.2. Prediction Model For each 𝑘 ∈ [1 ∶ 𝑛f it ], given history 𝑡𝑘 , our prediction problem is to train a model 𝑓𝑘 (𝑡 ) = [𝑌̂𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ℝ𝑛̂𝑡+2 that predicts [𝑌𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ , so as to minimize
𝑡+2
∈
𝑡+2
∑ 𝑡∈𝑘test
[ ( )] 𝔼 𝑙𝑝 [𝑌̂𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ , [𝑌𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ , 𝑡+2
𝑡+2
(21)
where 𝑘test = [𝑡𝑘 ∶ 𝑡𝑘+1 ) and 𝑙𝑝 is a loss function. For 𝑡 ∈ ℕ and (𝑐, 𝑑) ∈ 𝑈̂ 𝑡 , the target 𝑌𝑡𝑐𝑑 is defined as the normalized, commodity-wise demeaned return ̂ (𝑅◦ ; [𝑅◦ ] 𝑌𝑡𝑐𝑑 = Φ−1 (cdf 𝑡𝑐𝑑 𝑡𝑐𝑑 (𝑐,𝑑)∈𝑈̂ ))
(22)
Y. Hong and D. Klabjan, Preprint, June 25, 2026
7
𝑡
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
where Φ−1 is the standard normal quantile function, and the commodity-wise demeaned return is defined as (23)
𝑅◦𝑡𝑐𝑑 = 𝑅𝑡𝑐𝑑 − 𝑅∙𝑡𝑐
̂ (𝑅◦ ; [𝑅◦ ] with 𝑅∙𝑡𝑐 = avg(𝑅𝑡𝑐𝑑 ; 𝑑 ∈ 𝐷̂ 𝑡𝑐 ). By the probability integral transform, cdf ) is approximately 𝑡𝑐𝑑 𝑡𝑐𝑑 (𝑐,𝑑)∈𝑈̂ 𝑡 −1 uniformly distributed on [0, 1], and Φ maps it to an approximately standard normal variable. Since the standardnormal target can stabilize optimization (LeCun, Bottou, Orr and Müller, 2002; Glorot and Bengio, 2010; Hong, Kim, Kim and Choi, 2023), we adopt this prediction target. The motivation for using 𝑅◦𝑡𝑐𝑑 is discussed below. When addressing a task related to Π(𝐰𝑡 ) for a CS strategy, we can predict 𝑅◦𝑡𝑐𝑑 in lieu of 𝑅𝑡𝑐𝑑 . For 𝐰𝑡 ∈ 𝑡,CS , its ∑ ∑ ∑ ∑ ∑ ∑ strategy return satisfies Π(𝐰𝑡 ) = 𝑐∈𝐶̂ 𝑑∈𝐷̂ 𝑤𝑡𝑐𝑑 𝑅𝑡𝑐𝑑 = 𝑐∈𝐶̂ ( 𝑑∈𝐷̂ 𝑤𝑡𝑐𝑑 𝑅𝑡𝑐𝑑 − 𝑅∙𝑡𝑐 𝑑∈𝐷̂ 𝑤𝑡𝑐𝑑 ) = 𝑐∈𝐶̂ 𝑡 𝑡𝑐 𝑡 𝑡𝑐 𝑡𝑐 𝑡 ∑ ∑ ◦ ∙ (𝑐,𝑑)∈𝑈̂ 𝑡 𝑅𝑡𝑐𝑑 𝑤𝑡𝑐𝑑 where each equality follows from Eq. (6), Eq. (17), algebraic 𝑑∈𝐷̂ 𝑡𝑐 𝑤𝑡𝑐𝑑 (𝑅𝑡𝑐𝑑 − 𝑅𝑡𝑐 ) = manipulation, and Eq. (23), in that order. Thus, 𝑅◦𝑡𝑐𝑑 can replace 𝑅𝑡𝑐𝑑 in Eq. (6) to obtain the return of a CS strategy. As 𝑅◦𝑡𝑐𝑑 can be less exposed to the underlying commodity, 𝑅◦𝑡𝑐𝑑 can be more predictable than 𝑅𝑡𝑐𝑑 . Fix (𝑐, 𝑑) ∈ 𝑈̂ 𝑡 . The quantity 𝑅◦𝑡𝑐𝑑 can be interpreted as the return of a CS strategy for 𝑐 with a position vector [𝕀{𝑑 ′ =𝑑} ]𝑑 ′ ∈𝐷̂ − 𝑛̂1 𝟏 ∈ 𝑡𝑐
𝑡𝑐
ℝ𝑛̂𝑡𝑐 , which becomes a weight vector after scaling. In contrast, 𝑅𝑡𝑐𝑑 corresponds to the return of a LO strategy for 𝑐 with weight [𝕀{𝑑 ′ =𝑑} ]𝑑 ′ ∈𝐷̂ . By Proposition 3 and its experimental evidence in Section 5, 𝑅◦𝑡𝑐𝑑 has lower exposure to 𝑡𝑐 the underlying commodity than 𝑅𝑡𝑐𝑑 in a large fraction of cases. In addition, the data for underlying commodities are intrinsically incomplete.1 Therefore, we predict 𝑅◦𝑡𝑐𝑑 . For each 𝑘 ∈ [1 ∶ 𝑛f it ], we train 𝑓𝑘 on {(𝑡 , [𝑌̃𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∩𝑈̃ ) ∶ 𝑡 ∈ 𝑘f it }, validate on {(𝑡 , [𝑌̃𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∩𝑈̃ ) ∶ 𝑡+2 𝑡+2 𝑡+2 𝑡+2 test }, where f it ∩ val = ∅, f it ∪ val ⊂ [1 ∶ 𝑡 − 2], 𝑡 ∈ val }, and evaluate on {(𝑡 , [𝑌̃𝑡+2,𝑐𝑑 ] ̂ ̃ )∶𝑡∈ 𝑘 𝑘
(𝑐,𝑑)∈𝑈𝑡+2 ∩𝑈𝑡+2
𝑘
𝑘
𝑘
𝑘
𝑘
and 𝑘test = [𝑡𝑘 ∶ 𝑡𝑘+1 ). Since 𝑈̂ 𝑡+2 is determined at 𝑡 based solely on 𝑡 , some contracts in 𝑈̂ 𝑡+2 may not be observable at 𝑡 + 1 or 𝑡 + 2, i.e., 𝑉̃𝑡+1,𝑐𝑑 𝑉̃𝑡+2,𝑐𝑑 = 0; thus, we use 𝑈̂ 𝑡+2 ∩ 𝑈̃ 𝑡+2 in lieu of 𝑈̂ 𝑡+2 for the model training and evaluation: ̂ (𝑅̃ ◦ ; [𝑅̃ ◦ ] 𝑌̃𝑡𝑐𝑑 = Φ−1 (cdf 𝑡𝑐𝑑 (𝑐,𝑑)∈𝑈̂ ∩𝑈̃ )), 𝑡𝑐𝑑 𝑡
𝑡
𝑅̃ ◦𝑡𝑐𝑑 = 𝑅̃ 𝑡𝑐𝑑 − 𝑅̃ ∙𝑡𝑐 ,
𝑅̃ ∙𝑡𝑐 = avg(𝑅̃ 𝑡𝑐𝑑 ; 𝑑 ∈ 𝐷̂ 𝑡𝑐 ∩ 𝐷̃ 𝑡𝑐 ).
(24)
Note that the training and validation datasets are constructed solely from 𝑡𝑘 , which is given for training 𝑓𝑘 in the problem, and do not use any future data, i.e., 𝑡′ for 𝑡′ > 𝑡𝑘 , thereby ensuring no look-ahead bias.2 Additionally, 𝑘test ⋃ for 𝑘 ∈ [1 ∶ 𝑛f it ] partition [𝑡1 ∶ 𝑡𝑛f it +1 ), i.e., 𝑘∈[1∶𝑛f it ] 𝑘test = [𝑡1 ∶ 𝑡𝑛f it +1 ) with 𝑘test ∩ 𝑘test = ∅ for 𝑘 ≠ 𝑘′ . ′ Architecture Overview. To solve the problem, we propose a bi-level hierarchical graph learning architecture 𝑓𝑘 (𝑡 ) = (𝑓𝑘H ◦𝑓𝑘B ◦𝑓𝑘L ◦𝑓𝑘P )(𝑡 ) = [𝑌̂𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ . The preprocessing function 𝑓𝑘P maps 𝑡 to a bi-level hierarchical 𝑡+2 graph, and the graph lifting function 𝑓𝑘L aligns TTM grids across commodities by extending the graph along the TTM dimension. The bi-level convolution function 𝑓𝑘B then performs TTM-conditioned propagation through the extended bi-level graph, followed by the head 𝑓𝑘H , which outputs predictions [𝑌̂𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ . Throughout these functions, our 𝑡+2 architecture captures TTM-dependent information embedded in interrelationships among commodity futures contracts. The preprocessing function 𝑓𝑘P (𝑡 ) outputs a bi-level graph 𝐺𝑡 = (𝐶̂𝑡+2 ∪ 𝑈̂ 𝑡+2 , 𝐸𝑡CC+ ∪ 𝐸𝑡CC− ∪ 𝐸𝑡CU ∪ 𝐸𝑡UC ∪ 𝐸𝑡UU , 𝑋𝑡 ), illustrated in Figure 1. The commodities in 𝐶̂𝑡+2 form the upper-level nodes, and their contracts in 𝑈̂ 𝑡+2 constitute the lower-level nodes. The upper-level nodes are connected via 𝐸𝑡CC+ ⊂ 𝐶̂𝑡+2 × 𝐶̂𝑡+2 and 𝐸𝑡CC− ⊂ 𝐶̂𝑡+2 × 𝐶̂𝑡+2 , which denote the edge sets corresponding to positive and negative correlations between commodities, respectively. The cross-level connections are given by 𝐸𝑡CU ⊂ 𝐶̂𝑡+2 × 𝑈̂ 𝑡+2 and 𝐸𝑡UC ⊂ 𝑈̂ 𝑡+2 × 𝐶̂𝑡+2 , linking each underlying commodity and its contracts. The lower-level nodes are connected via 𝐸𝑡UU ⊂ 𝑈̂ 𝑡+2 × 𝑈̂ 𝑡+2 , which links each contract to its two nearest-maturity contracts (one shorter and one longer) with the same underlying commodity3 , reflecting the economic insight (Brennan, 1976; Gibson and Schwartz, 1990; Schwartz, 1997; Casassus and Collin-Dufresne, 2005) that futures min contracts with similar TTMs share similar economic characteristics. The matrix 𝑋𝑡 ∈ ℝ𝑛̂𝑡+2 ×𝑛sam contains node features for 𝑈̂ 𝑡+2 . 1 True spot markets for some underlying commodities do not exist (Gibson and Schwartz, 1990), and thus their spot price data are not directly observable. 2 The pair ( ̃ 𝑡𝑘 −2 , [𝑌𝑡𝑘 ,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ 𝑡 ∩𝑈̃ 𝑡 ) is the latest sample that can be used for training or validation. By Eqs. (22) and (23) and the definition 𝑘 𝑘 of 𝑈𝑡 , the latest historical data involved in (𝑡 −2 , [𝑌̃𝑡 ,𝑐𝑑 ] ̂ ̃ ) are 𝑡 , thereby ensuring no look-ahead bias. (𝑐,𝑑)∈𝑈𝑡 ∩𝑈𝑡 𝑘 𝑘 𝑘 𝑘 𝑘 𝑘 3 For each commodity, contracts with the smallest or largest maturity have only one edge in 𝐸 UU 𝑡
Y. Hong and D. Klabjan, Preprint, June 25, 2026
8
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Figure 1: Bi-Level Hierarchical Graph
To ensure proper backtesting, we define the trading universe 𝑈̂ 𝑡+2 as (25)
′ 𝑈̂ 𝑡+2 = {(𝑐, 𝑑) ∈ 𝑈̂ 𝑡+2 ∶ 𝜎̂ 𝑡𝑐𝑥 > 0} ′ is defined as the set of (𝑐, 𝑑) ∈ × such that where 𝜎𝑡𝑐𝑥 is defined below, and 𝑈̂ 𝑡+2 𝑐
(26)
𝑡 + 2 ≤ 𝑇𝑐𝑑 ,
(27)
max
𝑇𝑐𝑑 − 𝑡 ≤ 𝜏 , ∏ 𝑉̃𝑡′ 𝑐𝑑 > 0.
(28)
𝑡′ ∈(𝑡−𝑛min sam ,𝑡]
Here, 𝜏 max ∈ ℕ and 𝑛min sam ∈ ℕ are hyperparameters. The constraint in Eq. (25) ensures 𝑋𝑡 is well-defined, as discussed below. Eq. (26) ensures that position vector determined at 𝑡 can be liquidated at 𝑡 + 2, and Eq. (27) excludes long-TTM ′ is traded in the market contracts, which tend to be illiquid (Xu and Zhang, 2023). Eq. (28) implies that (𝑐, 𝑑) ∈ 𝑈̂ 𝑡+2 on all of the most recent 𝑛min sam trading dates, thereby ensuring liquidity. Unlike Hu et al. (2025); Tan et al. (2024), our trading universe evolves over time, allowing node creation and deletion. To enable neural networks to capture information in 𝑡 , we define the node feature matrix 𝑋𝑡 = [𝐱𝑡𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ 𝑡+2
min
min
ℝ𝑛̂𝑡+2 ×𝑛sam where 𝐱𝑡𝑐𝑑 = [𝑥𝑡𝑐𝑑𝜏 ]𝜏∈[0∶𝑛min ) ∈ ℝ𝑛sam and sam
̂ (𝑥′ ; [𝑥′ ′ ] 𝑥𝑡𝑐𝑑𝜏 = Φ−1 (cdf 𝑡𝑐𝑑𝜏 𝑡𝑐𝑑𝜏 (𝑐,𝑑)∈𝑈̂
min ′ 𝑡+2 ,𝜏 ∈[0∶𝑛sam )
))
(29)
min min min ̂ ̂′ for 𝑡 ∈ [𝑛min sam ∶ ∞), (𝑐, 𝑑) ∈ 𝑈𝑡+2 , and 𝜏 ∈ [0 ∶ 𝑛sam ). For 𝑡 ∈ [𝑛sam ∶ ∞), (𝑐, 𝑑) ∈ 𝑈𝑡+2 , and 𝜏 ∈ [0 ∶ 𝑛sam ), we define
𝑥′𝑡𝑐𝑑𝜏 = 𝑥′′ ̂ 𝑡𝑐𝑥 𝑡𝑐𝑑𝜏 ∕𝜎
(30)
′′ ̂′ 𝜎̂ 𝑡𝑐𝑥 = std(𝑥′′ ; 𝜏 ∈ [0 ∶ 𝑛min sam ), (𝑐, 𝑑 ) ∈ 𝑈𝑡+2 ), 𝑡𝑐𝑑 ′′ 𝜏 𝑥′′ = log(𝐹̃𝑡−𝜏,𝑐𝑑 ) − 𝜇̂ 𝑥 − 𝜇̂ 𝑥 + 𝜇̂ 𝑥 ,
(31)
where
𝑡𝑐⋅𝜏 𝑡𝑐⋅⋅ 𝑡𝑐𝑑𝜏 𝑡𝑐𝑑⋅ 𝑥 ′ min ̃ 𝜇̂ 𝑡𝑐𝑑⋅ = avg(log(𝐹𝑡−𝜏 ′ ,𝑐𝑑 ); 𝜏 ∈ [0 ∶ 𝑛sam )),
Y. Hong and D. Klabjan, Preprint, June 25, 2026
(32) (33) 9
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets 𝑥 ′ 𝜇̂ 𝑡𝑐⋅𝜏 = avg(log(𝐹̃𝑡−𝜏,𝑐𝑑 ′′ ); 𝑑 ′′ ∈ {𝑑 ′ ∶ (𝑐, 𝑑 ′ ) ∈ 𝑈̂ 𝑡+2 }),
(34)
𝑥 ′′ ′ ′ ̂′ 𝜇̂ 𝑡𝑐⋅⋅ = avg(log(𝐹̃𝑡−𝜏 ′ ,𝑐𝑑 ′′ ); 𝜏 ′ ∈ [0 ∶ 𝑛min sam ), 𝑑 ∈ {𝑑 ∶ (𝑐, 𝑑 ) ∈ 𝑈𝑡+2 }).
(35)
Eqs. (32) to (35) perform two-way ANOVA-style centering of log prices. To account for differing price scales across commodities, Eq. (30) scales the centered values within each commodity. Thus, each node feature contains centered and scaled log-price information over the most recent 𝑛min sam days. Since standard-normal features can stabilize optimization (LeCun et al., 2002; Glorot and Bengio, 2010; Hong et al., 2023), we normalize 𝑥′𝑡𝑐𝑑𝜏 in Eq. (29). To capture inter-commodity relationships, we define the inter-commodity edge sets 𝐸𝑡CC+ = {(𝑐, 𝑐 ′ ) ∈ 𝐶̂𝑡+2 × ∗ > 0 is a predefined ̂ 𝐶𝑡+2 ∶ 𝜌̂𝑡𝑐𝑐 ′ ≥ 𝜌∗ } and 𝐸𝑡CC− = {(𝑐, 𝑐 ′ ) ∈ 𝐶̂𝑡+2 × 𝐶̂𝑡+2 ∶ 𝜌̂𝑡𝑐𝑐 ′ ≤ −𝜌∗ }. Here, 𝜌⋃ threshold, and 𝜌̂𝑡𝑐𝑐 ′ is the Pearson correlation computed from the accumulated sample set 𝑡′ ∈[1∶𝑡] 𝐵𝑡′ 𝑐𝑐 ′ , provided that |{𝑡′ ∈ [1 ∶ 𝑡] ∶ 𝐵𝑡′ 𝑐𝑐 ′ ≠ ∅}| ≥ 𝑛min sam . If the minimum number of observation time points is not met, we set 𝜌̂𝑡𝑐𝑐 ′ = 0. A sample batch 𝐵𝑡𝑐𝑐 ′ is defined as the set of maturity-aligned pairs of estimated-and-demeaned returns: 𝐵𝑡𝑐𝑐 ′ = {(𝑅̂ ◦𝑡𝑐 (𝑇 ), 𝑅̂ ◦𝑡𝑐 ′ (𝑇 )) ∶ 𝑇 ∈ (𝑐 ∪ 𝑐 ′ ) ∩ dom(𝑅̂ ◦𝑡𝑐 ) ∩ dom(𝑅̂ ◦𝑡𝑐 ′ ), 𝑇 − 𝑡 ≤ 𝜏 max }. Here, 𝑅̂ ◦𝑡𝑐 (𝑇 ) is the demeaned return estimation function with domain dom(𝑅̂ ◦𝑡𝑐 ). Since maturities are not aligned across commodities— i.e., a commodity-𝑐 futures contract with maturity 𝑇 , or its return, may not exist—we linearly interpolate 𝑅̂ ◦𝑡𝑐 (𝑇 ) = 𝑅̃ ∗ + −𝑅̃ ∗𝑡𝑐𝑑 −
𝑡𝑐𝑑 (𝑇 − 𝑇𝑐𝑑 − ) + 𝑅̃ ∗𝑡𝑐𝑑 − where (𝑑 − , 𝑑 + ) minimizes |𝑑 + − 𝑑 − | subject to 𝑇 ∈ [𝑇𝑐𝑑 − , 𝑇𝑐𝑑 + ] and both 𝑅̃ ∗𝑡𝑐𝑑 + and 𝑇𝑐𝑑 + −𝑇𝑐𝑑 − 𝑅̃ ∗𝑡𝑐𝑑 − are well-defined. We define
(36)
′ ) 𝑅̃ ∗𝑡𝑐𝑑 = 𝑅̃ 𝑡𝑐𝑑 − avg(𝑅̃ 𝑡𝑐𝑑 ′ ; 𝑑 ′ ∈ 𝐷𝑡𝑐
′ | > 1, where 𝐷′ = {𝑑 ′ ∈ 𝐷 ∶ 𝑇 max }. Consistent with the prediction target Eq. (22), for 𝑑 ∈ 𝐷𝑡𝑐 with |𝐷𝑡𝑐 𝑡𝑐 𝑐𝑑 ′ − 𝑡 ≤ 𝜏 𝑡𝑐 which is based on demeaned returns Eq. (23), we use demeaned returns for 𝑅̂ ◦𝑡𝑐 (𝑇 ). CC+ CC− ∪ 𝐸 CU ∪ The graph lifting function 𝑓𝑘L (𝐺𝑡 ) outputs an augmented graph 𝐺𝑡0 = (𝐶̂𝑡+2,0 ∪ 𝑈̂ 𝑡+2 , 𝐸𝑡0 ∪ 𝐸𝑡0 𝑡0 UC UU 𝐸𝑡0 ∪ 𝐸𝑡 , 𝑍𝑡0 ). It transforms 𝐶̂𝑡+2 into an extended node set 𝐶̂𝑡+2,0 = {(𝑐, 𝑗) ∶ 𝑗 ∈ [0 ∶ 𝑛bas ]} of virtual futures contracts where 𝑗 indexes a virtual basis TTM 𝜏𝑗 = 𝑗 ⋅ (𝜏 max ∕𝑛bas ). This aligns maturity grids 𝑐 across commodities 𝑐 ∈ 𝐶̂𝑡+2 onto a shared virtual TTM grid {𝜏𝑗 ∶ 𝑗 ∈ [0 ∶ 𝑛bas ]}, since commodities generally do not share identical futures maturities, i.e., 𝑐 ′ ≠ 𝑐 ′′ for 𝑐 ′ , 𝑐 ′′ ∈ . Accordingly, the edges of 𝐺𝑡 are also extended along the TTM CC+ CC− , 𝐸 CU , 𝐸 UC . Specifically, 𝐸 CC+ = {((𝑛, 𝑗), (𝑛′ , 𝑗)) ∶ (𝑛, 𝑛′ ) ∈ 𝐸 CC+ , 𝑗 ∈ [0 ∶ 𝑛 ]}, dimension, yielding 𝐸𝑡0 , 𝐸𝑡0 bas 𝑡 𝑡0 𝑡0 𝑡0 CU CC− and 𝐸 UC defined analogously. While preserving the ′ ′ 𝐸𝑡0 = {((𝑛, 𝑗), 𝑛 ) ∶ (𝑛, 𝑛 ) ∈ 𝐸𝑡CU , 𝑗 ∈ [0 ∶ 𝑛bas ]} with 𝐸𝑡0 𝑡0 CC+ CC− connects the virtual nodes with the same virtual TTMs, and 𝐸 CU ∪ 𝐸 UC connects original topology, 𝐸𝑡0 ∪ 𝐸𝑡0 𝑡0 𝑡0 virtual nodes and their original futures nodes. Next, 𝑓𝑘L produces the initial embedding 𝑍𝑡0 = [𝐳𝑡,0,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ where 𝑡+2 𝐳𝑡,0,𝑐𝑑 = Dropout(𝜙(𝑊L,𝑘 𝐱𝑡𝑐𝑑 + 𝐛L,𝑘 )) ∈ ℝ𝑛hid for (𝑐, 𝑑) ∈ 𝑈̂ 𝑡+2 . Here, 𝑛hid ∈ ℕ is a hyperparameter. B B )(𝐺 ) with 𝑙 The bi-level convolution function 𝑓𝑘B (𝐺𝑡0 ) = (𝑓𝑘,𝑙 ◦ … ◦𝑓𝑘,1 𝑡0 con layers outputs 𝐺𝑡,𝑙con , where each con
B maps 𝐺 layer function 𝑓𝑘𝑙 𝑡,𝑙−1 to 𝐺𝑡𝑙 . Each layer output graph 𝐺𝑡𝑙 shares the same node and edge sets but differs CC+ CC− ∪ in its node embedding matrix 𝑍𝑡𝑙 = [𝐳𝑡𝑙𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ ℝ𝑛̂𝑡+2 ×𝑛hid , so that 𝐺𝑡𝑙 = (𝐶̂𝑡+2,0 ∪ 𝑈̂ 𝑡+2 , 𝐸𝑡0 ∪ 𝐸𝑡0 𝑡+2
CU ∪ 𝐸 UC ∪ 𝐸 UU , 𝑍 ). Each convolution layer function 𝑓 B consists of four operations, illustrated in Figure 2: 𝐸𝑡0 𝑡𝑙 𝑡 𝑘𝑙 𝑡0 UU ◦𝑔 CU ◦𝑔 CC ◦𝑔 UC where all the operation outputs share the same node and edge sets but have distinct node B 𝑓𝑘𝑙 = 𝑔𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 UC elevates embeddings from lower-level contract nodes to TTM-aligned virtual futures nodes embeddings. First, 𝑔𝑘𝑙 CC propagates information among virtual futures contracts sharing the same TTM. Then, at the upper-level. Next, 𝑔𝑘𝑙 CU transfers the aggregated upper-level information back to lower-level contract nodes. Finally, 𝑔 UU propagates 𝑔𝑘𝑙 𝑘𝑙 information across contracts with the nearest maturities within the same commodities. Throughout these steps, we CC and 𝑔 UU , distinguish contract relationships based on maturity differences and extract information accordingly via 𝑔𝑘𝑙 𝑘𝑙 in contrast to existing approaches (Hu et al., 2025; Tan et al., 2024) that treat all relationships in the same manner regardless of maturity. Each step is discussed next. UC transforms 𝑍 UC ] The elevating operator 𝑔𝑘𝑙 ∈ ℝ𝑛̂𝑡+2 ×𝑛hid into 𝑍𝑡𝑙UC = [𝐳𝑡𝑙𝑐𝑗 𝑡,𝑙−1 = [𝐳𝑡,𝑙−1,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ (𝑐,𝑗)∈𝐶̂ ×[0∶𝑛 ] ∈
ℝ𝑛̂𝑡+2 (𝑛bas +1)×𝑛hid , via linear interpolation UC 𝐳𝑡𝑙𝑐𝑗 =
𝐳𝑡,𝑙−1,𝑐𝑑 + − 𝐳𝑡,𝑙−1,𝑐𝑑 − 𝑇𝑐𝑑 + − 𝑇𝑐𝑑 −
𝑡+2
⋅ (𝑡 + 𝜏𝑗 − 𝑇𝑐𝑑 − ) + 𝐳𝑡,𝑙−1,𝑐𝑑 − ∈ ℝ𝑛hid
Y. Hong and D. Klabjan, Preprint, June 25, 2026
𝑡+2
bas
(37) 10
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Figure 2: Bi-Level Convolution
for 𝑡 + 𝜏𝑗 ∈ [𝑇𝑐,min 𝐷̂ , 𝑇𝑐,max 𝐷̂ ], where 𝑑 − , 𝑑 + ∈ 𝐷̂ 𝑡+2,𝑐 minimize |𝑇𝑐𝑑 + − 𝑇𝑐𝑑 − | subject to 𝑡 + 𝜏𝑗 ∈ [𝑇𝑐𝑑 − , 𝑇𝑐𝑑 + ]. 𝑡+2,𝑐 𝑡+2,𝑐 Since futures contracts with similar TTMs share similar economic characteristics (Brennan, 1976; Gibson and Schwartz, 1990; Schwartz, 1997; Casassus and Collin-Dufresne, 2005), we use interpolation based on neighboring maturities. For 𝑡 + 𝜏𝑗 ∉ [𝑇𝑐,min 𝐷̂ , 𝑇𝑐,max 𝐷̂ ], we use the end point values for the extrapolation by setting 𝑡+2,𝑐 𝑡+2,𝑐 UC ∗ ⋅ min 𝐷̂ 𝑡+2,𝑐 + 𝕀{𝜏𝑗 +𝑡>𝑇 ⋅ max 𝐷̂ 𝑡+2,𝑐 . Since both interpolation 𝐳𝑡𝑙𝑐𝑗 = 𝐳𝑡,𝑙−1,𝑐𝑑 ∗ with 𝑑 = 𝕀{𝜏𝑗 +𝑡<𝑇 𝑐,min 𝐷̂ 𝑡+2,𝑐 } 𝑐,max 𝐷̂ 𝑡+2,𝑐 } ∑ UC ] ⊤ 𝐩𝑡𝑐𝑑 𝐳𝑡,𝑙−1,𝑐𝑑 ∈ ℝ(𝑛bas +1)×𝑛hid and extrapolation are linear, for each 𝑐 ∈ 𝐶̂𝑡+2 , we can write [𝐳𝑡𝑙𝑐𝑗 𝑗∈[0∶𝑛bas ] = 𝑑∈𝐷̂ 𝑡+2,𝑐
for some constant vectors 𝐩𝑡𝑐𝑑 ∈ ℝ𝑛bas +1 . Note that 𝐩𝑡𝑐𝑑 does not depend on 𝑙 and 𝑍𝑡,𝑙−1 but depends on 𝑡, 𝜏𝑗 , 𝑇𝑐𝑑 + , and 𝑇𝑐𝑑 − . Thus, 𝐩𝑡𝑐𝑑 can be precomputed before training and inference, and shares the same values for all layers. As a UC , each 𝐳UC contains the information of a virtual future contract with TTM 𝜏 for commodity 𝑐. result of 𝑔𝑘𝑙 𝑗 𝑡𝑙𝑐𝑗 CC transforms 𝑍 UC into The inter-commodity propagation 𝑔𝑘𝑙 𝑡𝑙
𝑍𝑡𝑙𝐶𝐶 = 𝜙([𝑍𝑡𝑙UC 𝑍𝑡𝑙𝐶𝐶+ 𝑍𝑡𝑙𝐶𝐶− ]𝑊CC,𝑘𝑙 + 𝟏𝐛⊤ ) ∈ ℝ𝑛̂𝑡+2 (𝑛bas +1)×𝑛hid CC,𝑘𝑙
(38)
+ 𝐶𝐶+ where [𝑍𝑡𝑙UC 𝑍𝑡𝑙𝐶𝐶+ 𝑍𝑡𝑙𝐶𝐶− ] ∈ ℝ𝑛̂𝑡+2 (𝑛bas +1)×(3𝑛hid ) , 𝑍𝑡𝑙𝐶𝐶+ = 𝜙(CONV(𝐶̂𝑡+2,0 , 𝐸𝑡0 , 𝑍𝑡𝑙𝑈 𝐶 ; 𝜃𝐶𝐶,𝑙 )) ∈ ℝ𝑛̂𝑡+2 (𝑛bas +1)×𝑛hid , and 𝑍 𝐶𝐶− = 𝜙(CONV(𝐶̂𝑡+2,0 , 𝐸 𝐶𝐶− , 𝑍 𝑈 𝐶 ; 𝜃 − )) ∈ ℝ𝑛̂𝑡+2 (𝑛bas +1)×𝑛hid . Here, CONV denotes a graph convolution, 𝑡𝑙
𝑡0
𝑡𝑙
𝐶𝐶,𝑙
+ 𝐶𝐶+ 𝐶𝐶− connect nodes with the same virtual − and 𝜃𝐶𝐶,𝑙 and 𝜃𝐶𝐶,𝑙 are distinct learnable parameters. Since 𝐸𝑡0 and 𝐸𝑡0 TTM, each node (𝑐, 𝑗 ∗ ) ∈ 𝐶̂𝑡+2,0 aggregates information 𝐳UC′ ∗ from other nodes with the same TTM 𝑗 ∗ but different 𝑡𝑙𝑐 𝑗
CC propagates TTMcommodities 𝑐 ′ , while preserving the original topology of 𝐸𝑡𝐶𝐶+ and 𝐸𝑡𝐶𝐶− . In other words, 𝑔𝑘𝑙 UC ] ̂ aligned information [𝐳𝑡𝑙𝑐𝑗 𝑗∈[0∶𝑛bas ] across commodities 𝑐 ∈ 𝐶𝑡+2 . CU converts 𝑍 𝐶𝐶 = [𝐳CC ] CU ] The lowering function 𝑔𝑘𝑙 ∈ ℝ𝑛̂𝑡+2 (𝑛bas +1)×𝑛hid into 𝑍𝑡𝑙𝐶𝑈 = [𝐳𝑡𝑙𝑐𝑑 (𝑐,𝑑)∈𝑈̂ 𝑡+2 ∈ 𝑡𝑙 𝑡𝑙𝑐𝑗 𝑐∈𝐶̂𝑡+2 ,𝑗∈[0∶𝑛bas ] 𝑛 ̂ ×𝑛 ℝ 𝑡+2 hid via linear interpolation CU 𝐳𝑡𝑙𝑐𝑑 =
CC CC 𝐳𝑡𝑙𝑐,𝑗 − +1 − 𝐳𝑡𝑙𝑐𝑗 −
𝜏𝑗 − +1 − 𝜏𝑗 −
CC 𝑛hid ⋅ (𝑇𝑐𝑑 − 𝑡 − 𝜏𝑗 − ) + 𝐳𝑡𝑙𝑐𝑗 − ∈ ℝ
(39)
where 𝑗 − ∈ [0 ∶ 𝑛bas ) satisfies 𝑇𝑐𝑑 − 𝑡 ∈ [𝜏𝑗 − , 𝜏𝑗 − +1 ]. Since the intervals [𝜏𝑗 − , 𝜏𝑗 − +1 ] partition an interval [0, 𝜏 max ], the minimization used for the indexes in Eq. (37) is not required. Since this interpolation is linear, it can be precomputed, as discussed in the elevating operator Eq. (37). Y. Hong and D. Klabjan, Preprint, June 25, 2026
11
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets UU maps 𝑍 𝐶𝑈 to 𝑍 = [𝐳 The intra-commodity propagation 𝑔𝑘𝑙 𝑡𝑙 𝑡𝑙𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ 𝑡𝑙
𝑡+2
∈ ℝ𝑛̂𝑡+2 ×𝑛hid where (40)
CU UU 𝐳𝑡𝑙𝑐𝑑 = 𝜙(𝑊UU,𝑘𝑙 [𝐳𝑡𝑙𝑐𝑑 ; 𝐳𝑡𝑙𝑐𝑑 ] + 𝐛UU,𝑘𝑙 ) ∈ ℝ𝑛hid
with
∑
UU 𝐳𝑡𝑙𝑐𝑑 =
(41)
UU∗ 𝑛hid 𝐳𝑡𝑙𝑐𝑑𝑑 , ′ ∈ ℝ
𝑑 ′ ∈{𝑑 ′′ ∶(𝑑 ′′ ,𝑑)∈𝐸𝑡UU } UU∗ CU CU 𝑛hid 𝐳𝑡𝑙𝑐𝑑𝑑 , ′ = 𝜙(𝑊UU∗,𝑘𝑙 ⋅ LayerNorm(𝐳𝑡𝑙𝑐𝑑 − 𝐳𝑡𝑙𝑐𝑑 ′ ) + 𝐛UU∗,𝑘𝑙 ) ∈ ℝ
(42)
CU with neighboring information 𝐳UU . Eq. (41) aggregates In Eq. (40), 𝐳𝑡𝑙𝑐𝑑 integrates the target contract embedding 𝐳𝑡𝑙𝑐𝑑 𝑡𝑙𝑐𝑑 UU∗ from neighboring contracts in 𝐸 UU . Each message 𝐳UU∗ is computed from the normalized difference messages 𝐳𝑡𝑙𝑐𝑑𝑑 ′ ′ 𝑡 𝑡𝑙𝑐𝑑𝑑 between the target and source embeddings as in Eq. (42) to capture relative deviations along the term structure UU within the same commodity. Consequently, 𝑔𝑘𝑙 extracts intra-commodity information from nearest-maturity contracts of the same underlying commodity, reflecting the economic insight that futures with similar TTMs exhibit similar characteristics (Brennan, 1976; Gibson and Schwartz, 1990; Schwartz, 1997; Casassus and Collin-Dufresne, 2005). Throughout the bi-level convolution 𝑓𝑘B , each futures contract node integrates its information with that of related nodes while explicitly accounting for their TTMs, in contrast to existing approaches (Hu et al., 2025; Tan et al., 2024) that ignore the maturity dependence of these relationships. In their approaches, interactions between futures contracts are treated without differentiation by maturity, so information from contracts with similar maturities is processed in the same manner as that from contracts with dissimilar maturities. In our method, however, each convolution layer B uses two distinct message passing functions: 𝑔 CC , which propagates information across different commodities at 𝑓𝑘,𝑙 𝑘𝑙 UU , which propagates information across nearby TTMs within the same commodity. By stacking the same TTM, and 𝑔𝑘𝑙 B , the model propagates information across contracts with larger maturity gaps through successive multiple layers of 𝑓𝑘,𝑙 local interactions, thereby capturing higher-order dependencies over the joint commodity-maturity space. This design enables maturity-dependent propagation by explicitly distinguishing the relationships based on TTM while capturing inter-commodity and intra-commodity structures. The head 𝑓𝑘H (𝐺𝑡𝑘𝑙con ) transforms the final node embedding matrix 𝑍𝑡,𝑙con into a prediction vector [𝑌̂𝑡+2,𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ . 𝑡+2 For each (𝑐, 𝑑) ∈ 𝑈̂ 𝑡+2 , the head 𝑓𝑘H (𝐺𝑡𝑘𝑙con ) outputs 𝑌̂𝑡+2,𝑐𝑑 = 𝑊H,𝑘 𝐳𝑡𝑙con 𝑐𝑑 + 𝐛H,𝑘 ∈ ℝ. During the training, we minimize the mean squared error between 𝑌̂𝑡+2,𝑐𝑑 and 𝑌𝑡+2,𝑐𝑑 for (𝑐, 𝑑) ∈ 𝑈̂ 𝑡+2 ∩ 𝑈̃ 𝑡+2 .
5. Experiments We conduct experiments to address the following questions. (Q1) Can CS strategies be more effective than LO strategies for statistical arbitrage in commodity futures markets in terms of risk and risk-adjusted return? (Q2) Are TTM-dependent interrelationships among futures contracts instrumental for statistical arbitrage? (Q3) Are both inter-commodity and intra-commodity relationships instrumental? Specifically, we verify Eqs. (14) to (16) in the data, together with trading examples of LO baselines, for (Q1); compare our method against benchmarks in prediction and trading for (Q2); and conduct ablation studies for (Q3). Baselines, Benchmarks, and Ablations. We consider the following LO strategy baselines for (Q1). (B1) 𝐰𝑡 = 𝐕̂ 𝑡 ∕𝟏⊤ 𝐕̂ 𝑡 where 𝐕̂ 𝑡 = [Φ(𝑌̂𝑡𝑐𝑑 )](𝑐,𝑑)∈𝑈̂
𝑡
(B2) 𝐰𝑡 = 𝐕̂ 𝑡 ∕𝟏⊤ 𝐕̂ 𝑡 where 𝐕̂ 𝑡 = cdf (𝑌̂𝑡𝑐𝑑 ; [𝑌̂𝑡𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ) 𝑡 (B3) 𝐰𝑡 = 𝐕̂ 𝑡 ∕𝟏⊤ 𝐕̂ 𝑡 where 𝐕̂ 𝑡 = [𝑌̂𝑡𝑐𝑑 − 𝑌̂𝑡,min ](𝑐,𝑑)∈𝑈̂ and 𝑌̂𝑡,min = min{𝑌̂𝑡𝑐 ′ 𝑑 ′ ∶ (𝑐 ′ , 𝑑 ′ ) ∈ 𝑈̂ 𝑡 } 𝑡
(B4) 𝐰𝑡 = 𝑛̂1 𝟏 + 𝑘𝐰CS 𝑡 where 𝑘 = min{ 𝑡
position vector for 𝑈̂ 𝑡
1 CS ∶ (𝑐, 𝑑) ∈ 𝑈̂ 𝑡 , 𝑤CS < 0} and 𝐰CS 𝑡 = [𝑤𝑡𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ 𝑡 is a CS weight 𝑡𝑐𝑑 𝑛̂ 𝑡 |𝑤CS | 𝑡𝑐𝑑
Y. Hong and D. Klabjan, Preprint, June 25, 2026
12
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
The first equations of (B1) to (B4) ensure the unit sum requirement of LO positions, by scaling for (B1) to (B3) and by using the zero sum property of CS positions for (B4). The where clauses guarantee the nonnegative weight requirement of LO positions. From Eqs. (22) and (24), the first baseline transforms 𝑌̂𝑡𝑐𝑑 into a prediction for ̂ 𝑅̃ ◦ ; [𝑅̃ ◦ ] cdf( ). The second baseline directly maps 𝑌̂𝑡𝑐𝑑 to its empirical cumulative distribution value 𝑡𝑐𝑑 𝑡𝑐𝑑 (𝑐,𝑑)∈𝑈̂ 𝑡 ∩𝑈̃ 𝑡 within the cross section [𝑌̂𝑡𝑐𝑑 ] ̂ . The third baseline subtracts the minimum value of the predictions. The last (𝑐,𝑑)∈𝑈𝑡
baseline is an inverse transformation of 𝐰̌ in Proposition 2. The four resulting baseline LO positions preserve the ordering of the prediction or weight values: Φ and cdf in the first two baselines are monotonically increasing, and the translations in the last two do not change the order. Thus, (B1) to (B4) are natural transformations from a prediction vector to an LO position vector. The benchmarks for (Q2) include ridge linear regression, LightGBM (Ke, Meng, Finley, Wang, Chen, Ma, Ye and Liu, 2017), and multi-layered perceptrons (MLP) to assess the importance of the information extracted from commodity futures interrelationships. These methods take the node features 𝐱𝑡𝑐𝑑 , whose components are defined in Eq. (29), as predictor vectors, and use 𝑌𝑡+2,𝑐𝑑 defined in Eq. (22) as the prediction target, consistent with our method. Ridge linear regression is adopted due to its strong theoretical grounding, robustness to multicollinearity, and role as a canonical linear benchmark. As gradient-boosted tree methods have been shown to outperform neural networks on numerous tabular datasets (Ke et al., 2017; Grinsztajn, Oyallon and Varoquaux, 2022; Shwartz-Ziv and Armon, 2022; McElfresh, Khandagale, Valverde, Prasad C, Ramakrishnan, Goldblum and White, 2023; Shmuel, Glickman and Lazebnik, 2024), such as 𝐱𝑡𝑐𝑑 in our study, we include LightGBM as a benchmark. We also include MLP as a fundamental neural network baseline. The MLP consists of 𝑙con hidden layers with 𝑛hid hidden units, followed by a final linear output layer. Since these benchmarks are non-graph methods, comparing them with our method shows the value of incorporating relational information in commodity futures. We also include graph neural networks (GNNs) as benchmarks for (Q2) to assess whether the TTM-dependence is instrumental in extracting information from commodity futures relationships. To predict 𝑌𝑡+2,𝑐𝑑 given 𝑡 , the GNNs use a graph (𝑈̂ 𝑡+2 , 𝐸𝑡𝑈 𝑈 + ∪ 𝐸𝑡𝑈 𝑈 − , 𝑋𝑡 ), where nodes represent contracts, edges encode positive and negative correlations, and node features follow Eq. (29). Consistent with our method, we define 𝐸𝑡𝑈 𝑈 + = {((𝑐, 𝑑), (𝑐 ′ , 𝑑 ′ )) ∈ 𝑈̂ 𝑡+2 × 𝑈̂ 𝑡+2 ∶ 𝜌̂𝑡𝑐𝑑𝑐 ′ 𝑑 ′ ≥ 𝜌∗ } and 𝐸𝑡𝑈 𝑈 − = {((𝑐, 𝑑), (𝑐 ′ , 𝑑 ′ )) ∈ 𝑈̂ 𝑡+2 × 𝑈̂ 𝑡+2 ∶ 𝜌̂𝑡𝑐𝑑𝑐 ′ 𝑑 ′ ≤ −𝜌∗ }. Here, 𝜌̂𝑡𝑐𝑑𝑐 ′ 𝑑 ′ is the Pearson correlation computed from the accumulated pairs of (𝑅̃ ∗𝑡′′ 𝑐𝑑 , 𝑅̃ ∗𝑡′′ 𝑐 ′ 𝑑 ′ ) over 𝑡′′ ∈ [1 ∶ 𝑡], using only time points where both returns are available and at least 𝑛min sam observations exist. Since GNNs capture contract relationships but ignore TTMs, they serve to evaluate the importance of TTM dependence. The GNN consists of an initial embedding layer, 𝑙con interim layers, and a final head. The initial layer embeds 𝑋𝑡 𝐺 = 𝜙(𝑋 𝑊 ⊤ 𝑛̂ 𝑡 ×𝑛hid . Each interim layer propagates information over 𝐸 𝑈 𝑈 + and 𝐸 𝑈 𝑈 − : for as 𝑍𝑡,0 𝑡 G,𝑘0 + 𝟏𝐛G,𝑘0 ) ∈ ℝ 𝑡 𝑡 𝑙 ∈ [1 ∶ 𝑙con ], the 𝑙-th interim layer output is 𝐺 𝑍𝑡,𝑙 =𝜙
([
] ) 𝐺 + 𝑛̂ 𝑡 ×𝑛hid 𝐺 − ; 𝜃𝐺,𝑙 )) 𝑊𝐺,𝑘𝑙 + 𝟏𝐛⊤ ; 𝜃𝐺,𝑙 )), 𝜙(CONV(𝑈̂ 𝑡+2 , 𝐸𝑡𝑈 𝑈 − , 𝑍𝑡,𝑙−1 𝜙(CONV(𝑈̂ 𝑡+2 , 𝐸𝑡𝑈 𝑈 + , 𝑍𝑡,𝑙−1 𝐺,𝑘𝑙 ∈ ℝ (43)
− and 𝜃 + are learnable parameters. A final head outputs a prediction where CONV denotes a graph convolution, and 𝜃𝐺,𝑙 𝐺,𝑙 vector [𝑌̂𝑡+2,𝑐𝑑 ] = 𝑍 𝐺 𝑊𝐺,𝑘,𝐻 + 𝟏𝐛⊤ ∈ ℝ𝑛̂𝑡 . ̂ (𝑐,𝑑)∈𝑈𝑡+2
𝑡,𝑙con
𝐺,𝑘,𝐻
We consider the following ablations for (Q3), keeping all other components unchanged: B = 𝑔 UU ◦𝑔 CU ◦𝑔 CC ◦𝑔 UC with 𝑓 B = 𝑔 UU , (A1) Replace 𝑓𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 B = 𝑔 UU ◦𝑔 CU ◦𝑔 CC ◦𝑔 UC with 𝑓 B = 𝑔 CU ◦𝑔 CC ◦𝑔 UC . (A2) Replace 𝑓𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙 𝑘𝑙
These variants remove inter-commodity and intra-commodity information, respectively. Comparing them with the full model verifies the necessity of the two relationships in (Q3). Experimental settings. We use daily commodity futures price data from LSEG Datastream accessed via Wharton Research Data Services.4 The data span from Aug. 1977 to Dec. 2025 and consist of , 𝑐 and 𝑐 for each 𝑐 ∈ , and 𝑡 for futures contracts listed on CME, CBOT, NYMEX, COMEX, and eCBOT and traded in U.S. dollars. Trading dates are defined as weekdays on which the number of traded contract tickers exceeds 50% of the average over the preceding 4 See Appendix Section C for details on data processing.
Y. Hong and D. Klabjan, Preprint, June 25, 2026
13
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets Table 1 Hyperparameter Grids Method Our method, GNN, (A2) Ridge MLP,(A1) LightGBM
Hyperparameter Grid 4 5 ∗ CONV ∈ {GCN, { SAGE, GAT}, #params} ∈ {10 , 10 }, 𝜌 ∈ {0.1, 0.2, 0.3}, 𝑙conv ∈ {1, 2, 3} regularizer ∈ 10−10+0.1 𝑖 ∶ 𝑖 ∈ [0 ∶ 200] #params ∈ {104 , 105 }, 𝑙conv ∈ {1, 2, 3} learning_rate ∈ {0.02, 0.05, 0.1}, num_leaves ∈ {127, 255}, min_child_weight ∈ {100, 3000}, min_child_samples ∈ {20, 1000}, num_round ∈ {100, 500, 1000}, (top_rate, other_rate) ∈ {(0.05, 0.05), (0.05, 0.10), (0.10, 0.10), (0.15, 0.10), (0.15, 0.25), (0.20, 0.10), (0.25, 0.10)}
Figure 3: 1-Year Rolling Average Rate of Satisfaction of Conditions (ii) in Propositions 1 and 3
365 days. Since future trading dates are unavailable, TTM and the difference of maturities are estimated using calendar dates. To ensure sufficient training samples, we set 𝑡1 as the first trading date in 2016, define 𝑡𝑘 as the first trading date of each subsequent year, and set 𝑛f it = 10. We conduct stratified sampling for 𝑘val to mitigate the influence of period-specific market effects, treating each month as a separate stratum. We set 𝜏 max to one year, 𝑛bas = 52 (the number of weeks in a year), and 𝑛min sam = 28 that ensures 99% of contracts to be traded in the testing period. Statistics of the trading universe are provided in Appendix Section E. Since 𝑈̂ 𝑡 can contain unobservable contracts, i.e., contracts with ̂ 𝑡𝑐 ) = ∑ ̂ ̃ 𝑤𝑡𝑐𝑑 𝑅̃ 𝑡𝑐𝑑 . 𝑉̃𝑡−1,𝑐𝑑 𝑉̃𝑡,𝑐𝑑 = 0, 𝑅̃ 𝑡,𝑐𝑑 is unavailable for such contracts; thus, we compute Eq. (6) as Π(𝐰 𝑑∈𝐷𝑡𝑐 ∩𝐷𝑡𝑐 We evaluate test performance using hyperparameters selected by validation for each period. For each 𝑘 ∈ [1 ∶ 𝑛f it ] and each method, we train models over its hyperparameter grid on 𝑘f it , select the best configuration based on validation performance on 𝑘val , and evaluate on 𝑘test . The grids are given in Table 1. We use GCN (Kipf and Welling, 2017), GAT (Veličković, Cucurull, Casanova, Romero, Liò and Bengio, 2018), and SAGE (Hamilton, Ying and Leskovec, 2017) due to their wide usage (Fey and Lenssen, 2019). Although larger models may improve performance, the ranges of #params, 𝑙conv , 𝜌∗ are appropriate for the purpose of this study based on the selected configuration counts in Appendix Section D. For each model, 𝑛hid is chosen to achieve the closest match to the target #params. We set 𝜙 to the SiLU (Ramachandran, Zoph and Le, 2017; Hendrycks and Gimpel, 2023), and use default CONV hyperparameters from the original papers. For LightGBM, we use the grid generated from the best configurations reported in its original paper appendix (Ke et al., 2017), set remaining parameters to defaults in (Microsoft, 2026a,b), and apply early stopping as in (Ke et al., 2017). In total, 8,790 models are trained across ten periods: 540 each for our method, GNN, and (A2); 2,010 for ridge; 60 each for MLP and (A1); and 5,040 for LightGBM. Experimental results are summarized in Tables 2 to 4 and Figures 3 and 4. HGL denotes our hierarchical graph learning method, with HGL-GAT/GCN/SAGE specifying CONV = GAT∕GCN∕SAGE; GNN-GAT/GCN/SAGE are defined analogously. HGL and GNN are the best convolution methods selected from GAT, GCN, and SAGE per period based on validation performance. In Table 2 and Figure 4, the daily CS position vectors are computed via Eqs. (18) to (20) where 𝑌̂𝑡𝑐𝑑 are the corresponding model predictions. In Table 2, the daily LO position vectors (except for the equal-weight (EW) strategy and S&P 500) are computed from the predictions or CS weights of HGL. The EW strategy assigns equal positive weights across contracts. In Table 2, IR and SR measure risk-adjusted return; Vol and MDD measure risk; Tvr proxies transaction costs; and Cor measures diversification benefits relative to the S&P 500 (lower Y. Hong and D. Klabjan, Preprint, June 25, 2026
14
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Figure 4: Cumulative Sum of Profit and Loss, under the trading scheme in which each day half of the portfolio weight is long and the other half is short. The jumps in 2020 and 2022 may be associated with COVID-19 and the Russo-Ukrainian war, respectively.
Table 2 Trading Performance Across Methods (IR: information ratio; SR: Sortino ratio; Ret(%): average daily return (%); Vol(%): daily return volatility (%); Hit: hit ratio; MDD: maximum drawdown (absolute); Tvr: average daily turnover; Cor: correlation with the S&P 500). Under typical margin ratios, the actual daily returns (except for the S&P 500) can exceed the Ret values reported below by at least a factor of 8.33 = 1∕0.12. For example, the actual HGL daily return can exceed 0.08634%. This is because the Ret values below are reported under the unit-margin assumption in Section 3, whereas typical margin ratios in practice are 3–12% (CME Group, n.d.a) and even lower for spreads (CME Group, n.d.b). Strategy
Method
CS
HGL Ridge MLP LGBM GNN (A1) (A2) (B1) (B2) (B3) (B4) EW S&P 500
LO
IR
SR
Ret (%)
Vol (%)
MDD
Hit
Tvr
Cor
0.08459 0.03299 0.03233 0.04324 0.01492 0.07583 0.04420 0.03266 0.03606 0.03825 0.03691 0.03224 0.04836
0.12409 0.04306 0.04963 0.06834 0.02131 0.11027 0.06526 0.04398 0.04927 0.05234 0.05018 0.04337 0.05815
0.01036 0.00517 0.00487 0.00645 0.00254 0.00827 0.00609 0.02605 0.02785 0.02994 0.02944 0.02576 0.05525
0.12249 0.15685 0.15050 0.14922 0.17031 0.10904 0.13789 0.79765 0.77239 0.78279 0.79753 0.79910 1.14240
0.02002 0.05074 0.05609 0.03893 0.05710 0.02367 0.03028 0.34652 0.31870 0.31691 0.33267 0.34957 0.38249
0.53959 0.50696 0.51134 0.51612 0.50378 0.53442 0.52129 0.53124 0.53004 0.52845 0.53084 0.53124 0.54715
0.55726 0.62866 0.51328 0.65431 0.59353 0.53363 0.53653 0.04710 0.27418 0.24609 0.21805 0.03947
-0.01850 0.04576 0.02198 0.03262 0.04165 -0.03402 0.01235 0.30209 0.29924 0.29897 0.30171 0.30227 1.00000
is better).5 HGL achieves the highest risk-adjusted returns, measured by IR and SR, indicating the strongest statistical arbitrage performance.6 We address each of (Q1) to (Q3) based on these results. (Q1) We find strong empirical support for the effectiveness of CS strategies over LO strategies for statistical arbitrage, as the assumptions in Propositions 1 and 3 hold with high probability. Specifically, on average, 81.109% of (𝑡, 𝑐) pairs in our trading universe satisfy both Eqs. (14) and (15) associated with Proposition 1, and 99.976% satisfy Eq. (16) for Proposition 3. In Figure 3, the conditions in Propositions 1 and 3 hold for at least 70% and 99%, respectively, indicating that CS strategies exhibit lower variance and delta in a large fraction of cases. Moreover, by Proposition 2 (ii), if an LO strategy achieves sufficiently high expected returns as characterized therein, then the demeaned CS strategy can ̌ holds in a large fraction of cases. Note that these percentages outperform the LO strategy, as Var(Π(𝐰)) > Var(Π(𝐰)) do not imply CS underperformance otherwise; rather, they indicate that theoretical superiority is not established in those cases. 5 Additionally, yearly metrics are provided in Appendix Section F.
6 Investment objectives in finance typically focus on maximizing risk-adjusted returns (Sharpe, 1966, 1998).
Y. Hong and D. Klabjan, Preprint, June 25, 2026
15
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets Table 3 MSE of Prediction Models. The first row indicates the model selected from GAT, GCN, and SAGE based on validation performance for each period. Model
MSE
Model
MSE
HGL HGL-GAT HGL-GCN HGL-SAGE
1.125511 1.125746 1.125513 1.125342
GNN GNN-GAT GNN-GCN GNN-SAGE
1.126562 1.126513 1.126587 1.126562
LGBM MLP
1.126221 1.126497
Ridge
1.127108
Table 4 P-values from Paired Statistical Tests of MSE Differences between HGL and Benchmark Models. The first two tests assess the significance of the differences, while the last two columns evaluate the underlying assumptions: the Shapiro–Wilk test assesses the normality of the paired differences (for the t-test), and the symmetry test (Miao, Gel and Gastwirth, 2006) examines distributional symmetry (for the Wilcoxon test).
GNN GNN-GAT GNN-GCN GNN-SAGE MLP Ridge LGBM
Paired T
Wilcoxon Signed Rank
Shapiro-Wilk
Symmetry
6.27e-05 1.68e-04 1.96e-05 6.27e-05 6.22e-05 8.01e-09 2.14e-03
8.14e-05 2.95e-05 4.88e-05 8.14e-05 2.29e-05 5.52e-08 1.36e-03
7.81e-10 1.59e-02 3.34e-05 7.81e-10 6.80e-03 1.42e-01 1.32e-01
8.74e-01 9.99e-01 9.99e-01 8.74e-01 2.13e-01 2.43e-01 9.13e-01
Furthermore, Table 2 provides an empirical example of the effectiveness of CS trading. HGL achieves more than twice the IR and SR, more than six times lower volatility, and more than fifteen times lower MDD than the LO baselines (B1) to (B4) and the EW strategy. These results are consistent with the propositions, showing that CS strategies can deliver both lower risk and higher risk-adjusted return than LO strategies. Moreover, HGL outperforms the S&P 500, with over 75% higher IR, 113% higher SR, nine times lower volatility, and nineteen times lower MDD, indicating strong investment potential. It also exhibits near-zero correlation with the S&P 500, implying substantial diversification benefits when invested alongside traditional equities. At typical margin ratios employed by CME Group (CME Group, n.d.a,n), HGL’s daily return is 188% and 56% higher than those of the LO baselines and the S&P 500, respectively. (Q2) From the outperformance of HGL in Figure 4 and tables 2 to 4, we find that TTM-dependent interrelationships between futures contracts are instrumental for both trading and prediction. The cumulative sum P&L of HGL is stable and upward-sloping in Figure 4, compared to the other benchmarks. In Table 2, HGL consistently outperforms the non-graph CS benchmarks (Ridge, MLP, LGBM, and GNN) across all the key metrics: higher IR, SR, daily return, and hit ratio; and lower volatility and MDD. These results stem from the better prediction of HGL. All HGL models achieve lower MSEs (Table 3) than the benchmarks, with differences statistically significant at the 0.01 level (Table 4) based on paired tests listed in Hong and Klabjan, 2025a using daily MSEs. Since HGL outperforms the non-graph methods, futures interrelationships are informative; since it also outperforms GNNs, TTM dependence is important for capturing these relationships. (Q3) We confirm that both inter- and intra-commodity relationships are instrumental. In Table 2, although (A1) performs better than (A2), neither achieves IR and SR comparable to HGL, indicating that both types of relationships are informative. Hence, both are important for extracting information from futures contract relationships. Future work. Although our method generates statistical arbitrage and demonstrates the effectiveness of CS strategies and the importance of TTM-dependent futures interrelationships, there remains room for future research. Due to limited and contract-specific data availability, this study does not account for margin ratios or transaction costs; incorporating these factors is an important direction. While HGL captures the TTM-dependent interrelationship information, the exact nature of the extracted information remains unclear due to the black-box nature of deep learning Y. Hong and D. Klabjan, Preprint, June 25, 2026
16
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
methods, motivating further study. Although CS strategies can outperform LO strategies, their combination may yield additional gains under certain conditions, which presents an interesting direction for further study.
6. Conclusion This paper proposes a hierarchical graph learning approach for CS strategies in commodity futures markets, bridging two key gaps in the machine-learning literature: (i) the absence of learning-based methods for CS strategies in futures markets, and (ii) the lack of consideration for maturity-dependent interrelationships among commodity futures. We establish the efficacy of CS strategies by proving that they can exhibit higher risk-adjusted returns and lower risk than long-only strategies. We also present a method to transform predictions into CS positions. Next, we develop a hierarchical graph learning method that predicts futures price movements by leveraging TTMdependent interrelationships among futures, thereby yielding a trading algorithm. The experiments demonstrate that the hierarchical graph learning and resulting CS positions outperform benchmarks in prediction and trading, respectively. We find that maturity-dependent interrelationships across commodity futures are important and that CS trading based on hierarchical graph learning is effective for statistical arbitrage.
Y. Hong and D. Klabjan, Preprint, June 25, 2026
17
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Declaration on the Use of Generative AI and AI-Assisted Technologies in Manuscript Preparation In preparing this work, the authors used ChatGPT to assist with English-language editing, such as grammar correction and translation.
References Angelidis, T., Sakkas, A., Tessaromatis, N., 2025. Predicting commodity returns: Time series vs. cross sectional prediction models. Journal of Commodity Markets 38, 100475. Baruník, J., Malinska, B., 2016. Forecasting the term structure of crude oil futures prices with neural networks. Applied Energy 164, 366–379. Boakye, E.O., Heimonen, K., Junttila, J., 2024. Commodity markets and the global macroeconomy: Evidence from machine learning and GVAR. Empirical Economics 67, 1919–1965. Boons, M., Prado, M.P., 2019. Basis-momentum. The Journal of Finance 74, 239–279. Brennan, M.J., 1976. The supply of storage, in: The Economics of Futures Trading. Springer, pp. 100–107. Cai, X.J., Fang, Z., Chang, Y., Tian, S., Hamori, S., 2020. Co-movements in commodity markets and implications in diversification benefits. Empirical Economics 58, 393–425. Casassus, J., Collin-Dufresne, P., 2005. Stochastic convenience yield implied from commodity futures and interest rates. The Journal of Finance 60, 2283–2331. Chng, M.T., 2009. Economic linkages across commodity futures: Hedging and trading implications. Journal of Banking & Finance 33, 958–970. Clarke, R.G., De Silva, H., Thorley, S., 2013. Fundamentals of futures and options. CFA Institute Research Foundation 3. CME Group, n.d.a. The benefits of futures margins. URL: https://www.cmegroup.com/education/courses/ understanding-the-benefits-of-futures/the-benefits-of-futures-margins. online; accessed 30 March 2026. CME Group, n.d.b. Understanding futures spreads: futures spread overview. URL: https://www.cmegroup.com/education/courses/ understanding-futures-spreads/futures-spread-overview. online; accessed 30 March 2026. Fengqian, D., Chao, L., 2020. An adaptive financial trading system using deep reinforcement learning with candlestick decomposing features. IEEE Access 8, 63666–63678. Fey, M., Lenssen, J.E., 2019. Fast graph representation learning with PyTorch Geometric. arXiv preprint arXiv:1903.02428. Furlong, F., Ingenito, R., 1996. Commodity prices and inflation. Economic Review-Federal Reserve Bank of San Francisco , 27–47. Gabrielsson, P., Johansson, U., 2015. High-frequency equity index futures trading using recurrent reinforcement learning with candlesticks, in: 2015 IEEE Symposium Series on Computational Intelligence, IEEE. pp. 734–741. Gibson, R., Schwartz, E.S., 1990. Stochastic convenience yield and the pricing of oil contingent claims. The Journal of Finance 45, 959–976. Glorot, X., Bengio, Y., 2010. Understanding the difficulty of training deep feedforward neural networks, in: Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, PMLR. pp. 249–256. Gospodinov, N., Ng, S., 2013. Commodity prices, convenience yields, and inflation. Review of Economics and Statistics 95, 206–219. Grinsztajn, L., Oyallon, E., Varoquaux, G., 2022. Why do tree-based models still outperform deep learning on typical tabular data?, in: Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 507–520. Guijarro-Ordonez, J., Pelger, M., Zanotti, G., 2025. Deep learning statistical arbitrage. Management Science 0. Hamilton, W., Ying, Z., Leskovec, J., 2017. Inductive representation learning on large graphs, in: Advances in Neural Information Processing Systems, Curran Associates, Inc. Hammoudeh, S., Sari, R., Ewing, B.T., 2009. Relationships among strategic commodities and with financial variables: A new look. Contemporary Economic Policy 27, 251–264. Hendrycks, D., Gimpel, K., 2023. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415. Hong, Y., Kim, Y., Kim, J., Choi, Y., 2023. Index tracking via learning to predict market sensitivities, in: Proceedings of SAI Intelligent Systems Conference, Springer. pp. 111–131. Hong, Y., Klabjan, D., 2025a. Graph learning for foreign exchange rate prediction and statistical arbitrage, in: Proceedings of the 6th ACM International Conference on AI in Finance, ACM. pp. 692–699. Hong, Y., Klabjan, D., 2025b. Statistical arbitrage in options markets by graph learning and synthetic long positions. arXiv preprint arXiv:2508.14762. Hu, M., Tan, Z., Liu, B., Yin, G., 2025. Graph portfolio: High-frequency factor predictors via heterogeneous continual GNNs. IEEE Transactions on Knowledge and Data Engineering 37, 4104–4116. Hull, J., Treepongkaruna, S., Colwell, D., Heaney, R., Pitt, D., 2013. Fundamentals of futures and options markets. 8th ed., Pearson Higher Education AU. Hung, C.C., Chen, Y.J., 2021. DPP: Deep predictor for price movement from candlestick charts. Plos One 16, e0252404. Ji, Q., Fan, Y., 2012. How does oil price volatility affect non-energy commodity markets? Applied Energy 89, 273–280. Kaur, B., Sidhu, B.K., Bhathal, G.S., 2025. A hybrid deep reinforcement learning approach for algorithmic trading in commodity futures markets. Applied Computational Intelligence and Soft Computing 2025, 5993683. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., Liu, T.Y., 2017. LightGBM: A highly efficient gradient boosting decision tree, in: Advances in Neural Information Processing Systems, Curran Associates, Inc. Kipf, T.N., Welling, M., 2017. Semi-supervised classification with graph convolutional networks, in: International Conference on Learning Representations. Lazzarino, M., Berrill, J., Šević, A., et al., 2018. What is statistical arbitrage? Theoretical Economics Letters 8, 888–908. LeCun, Y., Bottou, L., Orr, G.B., Müller, K.R., 2002. Efficient backprop, in: Neural networks: Tricks of the Trade. Springer, pp. 9–50.
Y. Hong and D. Klabjan, Preprint, June 25, 2026
18
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets Li, J., Li, G., Liu, M., Zhu, X., Wei, L., 2022. A novel text-based framework for forecasting agricultural futures using massive online news headlines. International Journal of Forecasting 38, 35–50. Liu, J., Lu, L., Zong, X., Xie, B., 2023. Nonlinear relationships in soybean commodities pairs trading-test by deep reinforcement learning. Finance Research Letters 58, 104477. Markowitz, H., 1952. Portfolio selection. The Journal of Finance 7, 77–91. Maslyuk, S., Smyth, R., 2009. Cointegration between oil spot and future prices of the same and different grades in the presence of structural change. Energy Policy 37, 1687–1693. Massahi, M., Mahootchi, M., 2024. A deep q-learning based algorithmic trading system for commodity futures markets. Expert Systems with Applications 237, 121711. McElfresh, D., Khandagale, S., Valverde, J., Prasad C, V., Ramakrishnan, G., Goldblum, M., White, C., 2023. When do neural nets outperform boosted trees on tabular data?, in: Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 76336–76369. Miao, W., Gel, Y.R., Gastwirth, J.L., 2006. A new test of symmetry about an unknown median, in: Random Walk, Sequential Analysis and Related Topics: A Festschrift in Honor of Yuan-Shih Chow. World Scientific, pp. 199–214. Microsoft, 2026a. Lightgbm documentation. URL: https://lightgbm.readthedocs.io/en/latest/Python-Intro.html. online; accessed 30 March 2026. Microsoft, 2026b. Lightgbm python package. URL: https://pypi.org/project/lightgbm/. online; accessed 30 March 2026. Mitchell, D., 2008. A note on rising food prices. World Bank Policy Research Working Paper . Nazlioglu, S., 2011. World oil and agricultural commodity prices: Evidence from nonlinear causality. Energy Policy 39, 2935–2943. Ramachandran, P., Zoph, B., Le, Q.V., 2017. Searching for activation functions. arXiv preprint arXiv:1710.05941. Schwartz, E.S., 1997. The stochastic behavior of commodity prices: Implications for valuation and hedging. The Journal of Finance 52, 923–973. Sharpe, W.F., 1966. Mutual fund performance. The Journal of Business 39, 119–138. Sharpe, W.F., 1998. The Sharpe ratio. Streetwise—the Best of the Journal of Portfolio Management 3, 169–85. Shmuel, A., Glickman, O., Lazebnik, T., 2024. A comprehensive benchmark of machine and deep learning across diverse tabular datasets. arXiv preprint arXiv:2408.14817. Shwartz-Ziv, R., Armon, A., 2022. Tabular data: Deep learning is not all you need. Information Fusion 81, 84–90. Sun, T., Wang, J., Ni, J., Cao, Y., Liu, B., 2019. Predicting futures market movement using deep neural networks, in: 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA), IEEE. pp. 118–125. Szymanowska, M., De Roon, F., Nijman, T., Van Den Goorbergh, R., 2014. An anatomy of commodity futures risk premia. The Journal of Finance 69, 453–482. Tan, Z., Hu, M., Liu, B., Yin, G., 2024. Futures quantitative investment with heterogeneous continual graph neural network, in: 2024 IEEE International Conference on Data Mining (ICDM), IEEE. pp. 851–856. Vacha, L., Janda, K., Kristoufek, L., Zilberman, D., 2013. Time-frequency dynamics of biofuel-fuel-food system. Energy Economics 40, 233–241. Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., Bengio, Y., 2018. Graph attention networks, in: International Conference on Learning Representations. Wang, J., Wu, J., Zhang, G., Tan, M., Chen, S., Lin, Z., 2025. Agricultural futures trading decision using AI agent with multiscale candlestick analysis. IEEE Transactions on Computational Social Systems . Wang, S., Zhang, T., 2024. Predictability of commodity futures returns with machine learning models. Journal of Futures Markets 44, 302–322. Xu, X., Zhang, Y., 2023. Neural network predictions of the high-frequency CSI300 first distant futures trading volume. Financial Markets and Portfolio Management 37, 191–207. Zhang, L., Wang, J., Wang, B., 2020. Energy market prediction with novel long short-term memory network: Case study of energy futures index volatility. Energy 211, 118634.
Y. Hong and D. Klabjan, Preprint, June 25, 2026
19
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
A. Proofs of Propositions A.1. Proof of Proposition 1 Proof. Fix 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶̂𝑡 . We suppress the subscripts (𝑡, 𝑐) and write 𝑤𝑑 = 𝑤𝑡𝑐𝑑 , 𝜎𝑑𝑑√′ = 𝜎𝑡𝑐𝑑𝑑 ′ , 𝜌𝑑𝑑 ′ = 𝜌𝑡𝑐𝑑𝑑 ′ , 𝑛̂ = 𝑛̂ 𝑡𝑐 , 𝜅 = 𝜅𝑡𝑐 , 𝑅𝑑 = 𝑅𝑡𝑐𝑑 , 𝐷̂ = 𝐷̂ 𝑡𝑐 , LO = 𝑡𝑐,LO , and CS = 𝑡𝑐,CS . Let 𝜎𝑑 = 𝜎𝑑𝑑 , 𝜎min = min𝑑∈𝐷̂ 𝜎𝑑 , 𝜎max = max𝑑∈𝐷̂ 𝜎𝑑 , and 𝜌∗ = min𝑑,𝑑 ′ ∈𝐷̂ 𝜌𝑑𝑑 ′ . Let 𝐰 = [𝑤𝑑 ]𝑑∈𝐷̂ ∈ LO and 𝐰′ = [𝑤′𝑑 ]𝑑∈𝐷̂ ∈ CS be arbitrary. (i) Assume that 𝜌∗ > (𝜅 2 − 2‖𝐰‖22 )(3 − 2‖𝐰‖22 )−1 and 𝜌∗ ≥ 0. We obtain Eq. (A.1) from Eq. (6) and Eq. (A.2) ∑ from the definition of 𝜌∗ . Pulling out 𝜌∗ 𝑑∈𝐷̂ 𝑤2𝑑 𝜎𝑑2 from the first term of Eq. (A.2) and adding it to the second term ∑ 2 2 for all 𝑑 ∈ 𝐷, ∗ ∑ 2 ̂ of Eq. (A.2), we have Eq. (A.3) since 𝜌∗ ′ ̂ 𝑤𝑑 𝑤𝑑 ′ 𝜎𝑑 𝜎𝑑 ′ = 𝜌 ( ̂ 𝑤𝑑 𝜎𝑑 ) . Since 𝜎 ≥ 𝜎 𝑑,𝑑 ∈𝐷
𝑑∈𝐷
and 𝟏⊤ 𝐰 = 1, we obtain Eq. (A.4). ∑ ∑ Var(Π(𝐰)) = 𝑤𝑑 𝑤𝑑 ′ 𝜎𝑑𝑑 ′ = 𝑤2𝑑 𝜎𝑑2 + 𝑑,𝑑 ′ ∈𝐷̂
∑
≥
𝑑∈𝐷̂
𝑑∈𝐷̂
= (1 − 𝜌∗ )
∑
min
(A.1)
𝑤𝑑 𝑤𝑑 ′ 𝜎𝑑 𝜎𝑑 ′ 𝜌𝑑𝑑 ′
′ ̂ 𝑑,𝑑 ′ ∈𝐷∶𝑑≠𝑑
∑
𝑤2𝑑 𝜎𝑑2 + 𝜌∗
∑
𝑑
(A.2)
𝑤𝑑 𝑤𝑑 ′ 𝜎𝑑 𝜎𝑑 ′
′ ̂ 𝑑,𝑑 ′ ∈𝐷∶𝑑≠𝑑
𝑤2𝑑 𝜎𝑑2 + 𝜌∗ (
𝑑∈𝐷̂ ∗
( 2 ≥ 𝜎min (1 − 𝜌 )‖𝐰‖22 + 𝜌
∑
(A.3)
𝑤𝑑 𝜎𝑑 )2
𝑑∈𝐷̂ ) ∗
(A.4)
We obtain Eq. (A.5) from Eq. (6). We have Eq. (A.6) since 𝑤′𝑑 = [𝑤′𝑑 ]+ − [−𝑤′𝑑 ]+ . Since 𝜎𝑑 𝜎𝑑 ′ 𝜌∗ ≤ 𝜎𝑑 𝜎𝑑 ′ 𝜌𝑑𝑑 ′ = ∑ ∑ 𝜎𝑑𝑑 ′ ≤ 𝜎𝑑 𝜎𝑑 ′ , we obtain Eq. (A.8). Since 𝐰′ is a CS position weight vector, we have 𝑑∈𝐷̂ [𝑤′𝑑 ]+ = 𝑑∈𝐷̂ [−𝑤′𝑑 ]+ = ∑ ′ + 2 since 𝜎 ≤ 𝜎 ′ + 2 2 ∕4 = 𝜎 2 (∑ 1∕2 from Eq. (5). Then, we have 𝜎max 𝑑 max , which 𝑑∈𝐷̂ [𝑤𝑑 ] ) ≥ ( 𝑑∈𝐷̂ 𝜎𝑑 [𝑤𝑑 ] ) ∑ max ′ 2 yields the upper bound of the first term of Eq. (A.9). Analogously, we obtain 𝜎max ∕4 ≥ ( 𝑑∈𝐷̂ 𝜎𝑑 [−𝑤𝑑 ]+ )2 , which ∑ yields the upper bound of the second term of Eq. (A.9). Therefore, we obtain Eq. (A.10) since 𝑑∈𝐷̂ 𝜎𝑑 [𝑤′𝑑 ]+ ≥ ∑ ∑ 𝜎min 𝑑∈𝐷̂ [𝑤′𝑑 ]+ = 𝜎min ∕2 and 𝑑∈𝐷̂ 𝜎𝑑 [−𝑤′𝑑 ]+ ≥ 𝜎min ∕2. Var(Π(𝐰′ )) =
∑ 𝑑,𝑑 ′ ∈𝐷̂
=
∑
𝑤′𝑑 𝑤′𝑑 ′ 𝜎𝑑𝑑 ′
(A.5)
) )( ( 𝜎𝑑𝑑 ′ [𝑤′𝑑 ]+ − [−𝑤′𝑑 ]+ [𝑤′𝑑 ′ ]+ − [−𝑤′𝑑 ′ ]+
(A.6)
𝑑,𝑑 ′ ∈𝐷̂
=
∑
∑
𝜎𝑑 𝜎𝑑 ′ [𝑤′𝑑 ]+ [𝑤′𝑑 ′ ]+ +
𝑑,𝑑 ′ ∈𝐷̂
( =
∑
𝜎𝑑𝑑 ′ [−𝑤′𝑑 ]+ [−𝑤′𝑑 ′ ]+ − 2
∑
∑
𝜎𝑑 𝜎𝑑 ′ [−𝑤′𝑑 ]+ [−𝑤′𝑑 ′ ]+ − 2𝜌∗
𝑑,𝑑 ′ ∈𝐷̂
)2 𝜎𝑑 [𝑤′𝑑 ]+
( +
𝑑∈𝐷̂
∑
𝑑∈𝐷̂
𝜎𝑑𝑑 ′ [𝑤′𝑑 ]+ [−𝑤′𝑑 ′ ]+
𝑑,𝑑 ′ ∈𝐷̂
𝑑,𝑑 ′ ∈𝐷̂
𝑑,𝑑 ′ ∈𝐷̂
≤
∑
𝜎𝑑𝑑 ′ [𝑤′𝑑 ]+ [𝑤′𝑑 ′ ]+ +
𝜎𝑑 𝜎𝑑 ′ [𝑤′𝑑 ]+ [−𝑤′𝑑 ′ ]+
𝑑,𝑑 ′ ∈𝐷̂
)2 𝜎𝑑 [−𝑤′𝑑 ]+
∑
(A.7)
( − 2𝜌∗
∑
)( 𝜎𝑑 [𝑤′𝑑 ]+
𝑑∈𝐷̂
𝜎 𝜎 1 1 1 2 2 2 ≤ 𝜎max ( + ) − 2𝜌∗ ⋅ min ⋅ min = (𝜎max − 𝜌∗ 𝜎min ) 4 4 2 2 2
∑
) 𝜎𝑑 [−𝑤′𝑑 ]+
(A.8) (A.9)
𝑑∈𝐷̂
(A.10)
Since 𝟏⊤ 𝐰 = 1 and 𝑤𝑑 ≥ 0, we have ‖𝐰‖22 ≤ (𝟏⊤ 𝐰)2 = 1
(A.11)
and thus 3 − 2‖𝐰‖22 > 0. From this, since 𝜌∗ > (𝜅 2 − 2‖𝐰‖22 )(3 − 2‖𝐰‖22 )−1 , we obtain 𝜌∗ (3 − 2‖𝐰‖22 ) > 𝜅 2 − 2‖𝐰‖22 , ) 2 ‖𝐰‖22 + 𝜌∗ (1 − ‖𝐰‖22 ) > 𝜅 2 − 𝜌∗ , (
Y. Hong and D. Klabjan, Preprint, June 25, 2026
(A.12) (A.13) 20
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
( ) 1 2 2 2 𝜎min (1 − 𝜌∗ )‖𝐰‖22 + 𝜌∗ > (𝜎max − 𝜌∗ 𝜎min ), 2
(A.14)
where the last inequality follows from the definition of 𝜅 2 . From Eqs. (A.4), (A.10) and (A.14), we obtain Var(Π(𝐰)) > Var(Π(𝐰′ )). (ii) Assume that min𝑑,𝑑 ′ ∈𝐷̂ 𝜌𝑡𝑐𝑑𝑑 ′ ≥ 0 and min𝑑,𝑑 ′ ∈𝐷̂ 𝜌𝑡𝑐𝑑𝑑 ′ > (𝜅 2 𝑛̂ − 2)(3𝑛̂ − 2)−1 . We have 𝜅 2 < 3 since 2
−3)𝑛̂ , 𝑛̂ is natural, and every correlation coefficient does not exceed one. Since 𝟏⊤ 𝐰 = 1, (𝜅 2 𝑛̂ − 2)(3𝑛̂ − 2)−1 = 1 + (𝜅3𝑛−2 ̂ by the Cauchy-Schwarz inequality and Eq. (A.11), we obtain 1∕𝑛̂ ≤ ‖𝐰‖22 ≤ 1. We thus have
(A.15)
3 − 2𝑛̂ −1 ≥ 3 − 2‖𝐰‖22 , (𝜅 2 𝑛̂ − 2)(3𝑛̂ − 2)−1 = 1 +
𝜅2 − 3 3 − 2𝑛̂ −1
≥
𝜅2 − 3 3 − 2‖𝐰‖22
+ 1 = (𝜅 2 − 2‖𝐰‖22 )(3 − 2‖𝐰‖22 )−1 .
(A.16)
From this and the assumption, we obtain min𝑑,𝑑 ′ ∈𝐷̂ 𝜌𝑡𝑐𝑑𝑑 ′ > (𝜅 2 𝑛̂ − 2)(3𝑛̂ − 2)−1 ≥ (𝜅 2 − 2‖𝐰‖22 )(3 − 2‖𝐰‖22 )−1 . From (i), we obtain Var(Π(𝐰)) > Var(Π(𝐰′ )).
A.2. Proof of Proposition 2 ̄ ̄ 1 ∈ 𝑡𝑐,CS with 𝐰 ≠ 𝐰. ̄ Since both Proof. Let 𝐰 ∈ 𝑡𝑐,LO , 𝐰̄ = 𝟏∕𝑛̂ 𝑡𝑐 ∈ 𝑡𝑐,LO , and 𝐰̌ = (𝐰 − 𝐰)∕‖𝐰 − 𝐰‖ expectation and Eq. (6) are linear, we have (A.17)
̄ = 𝔼Π(𝐰) − 𝔼Π(𝐰), ̄ 𝔼Π(𝐰 − 𝐰) ̄ ̄ 𝔼Π(𝐰 − 𝐰) 𝔼Π(𝐰) − 𝔼Π(𝐰) ̌ = 𝔼Π(𝐰) = . ̄ 1 ̄ 1 ‖𝐰 − 𝐰‖ ‖𝐰 − 𝐰‖
(A.18)
Due to the linearity of Eq. (6) and the property of variance, we have ) ( ̄ ̄ Var(Π(𝐰 − 𝐰)) Π(𝐰 − 𝐰) ̌ = Var = Var(Π(𝐰)) . 2 ̄ 1 ‖𝐰 − 𝐰‖ ̄ ‖𝐰 − 𝐰‖1
(A.19)
̄ > 0, we have From Eqs. (A.18) and (A.19), since Var(Π(𝐰 − 𝐰)) ̄ ̄ 1 𝔼[Π(𝐰 − 𝐰)]∕‖𝐰 − 𝐰‖ ̄ ̌ 𝔼Π(𝐰 − 𝐰) 𝔼Π(𝐰) ̄ ̌ =√ =√ =√ = IR(Π(𝐰 − 𝐰)). IR(Π(𝐰)) ̄ ̌ Var(Π(𝐰 − 𝐰)) Var(Π(𝐰)) ̄ ̄ 2 Var(Π(𝐰 − 𝐰))∕‖𝐰 − 𝐰‖
(A.20)
1
̄ > 0, and Var(Π(𝐰)) > 0. From Eqs. (10) and (A.17), we obtain (i) Assume that Eq. (10) holds, Var(Π(𝐰 − 𝐰)) + ̄ ̄ = √ 𝔼Π(𝐰−𝐰) ̌ > [IR(Π(𝐰))]+ due to Eq. (A.20). IR(Π(𝐰 − 𝐰)) > √[𝔼Π(𝐰)] = [IR(Π(𝐰))]+ , yielding IR(Π(𝐰)) ̄ Var(Π(𝐰−𝐰))
Var(Π(𝐰)
̄ > 0, and Var(Π(𝐰)) > Var(Π(𝐰)). ̌ From Eq. (A.20), we have the first (ii) Assume that Eq. (11) holds, Var(Π(𝐰−𝐰)) ̄ ̌ = Var(Π(𝐰−2𝐰)) , two equalities of Eq. (A.21). By the assumption and Eq. (A.19), we obtain Var(Π(𝐰)) > Var(Π(𝐰)) ̄ 1 ‖𝐰−𝐰‖
̄ > 0 by Eqs. (11) which yields the inequality between the third and fourth terms in Eq. (A.21) since 𝔼[Π(𝐰 − 𝐰)] ̄ > ‖𝐰 − 𝐰‖ ̄ 1 [𝔼Π(𝐰)]+ , yielding the inequality between the and (A.17). By Eqs. (11) and (A.17), we have 𝔼Π(𝐰 − 𝐰) ̌ > [IR(Π(𝐰))]+ : fourth and fifth terms in Eq. (A.21). From the following, we have IR(Π(𝐰)) ̄ 1 [𝔼Π(𝐰)]+ ‖𝐰 − 𝐰‖ ̄ ̄ 𝔼[Π(𝐰 − 𝐰)] 𝔼[Π(𝐰 − 𝐰)] ̌ = IR(Π(𝐰−𝐰)) ̄ =√ > > = [IR(Π(𝐰))]+ . IR(Π(𝐰)) √ √ ̄ ̄ 1 Var(Π(𝐰)) ‖𝐰 − 𝐰‖ ̄ 1 Var(Π(𝐰)) Var(Π(𝐰 − 𝐰)) ‖𝐰 − 𝐰‖ (A.21) ̄ = 0, and Var(Π(𝐰)) > 0. We have 𝔼Π(𝐰) ̌ > 0 by Eqs. (12) (iii) We assume that Eq. (12) holds, Var(Π(𝐰 − 𝐰)) ̌ = 0 by Eq. (A.19) and the zero-variance assumption. Let 𝐰 = [𝑤𝑐𝑑 ]𝑑∈𝐷̂ . From Eq. (6), and (A.18) and Var(Π(𝐰)) 𝑡𝑐 ∑ ∑ 𝔼Π(𝐰) = 𝔼 𝑑∈𝐷̂ 𝑤𝑐𝑑 𝑅𝑡𝑐𝑑 < 𝑑∈𝐷̂ 𝔼𝑅𝑡𝑐𝑑 < ∞, where the first inequality is because ‖𝐰‖1 = 1 and 𝑤𝑐𝑑 ≥ 0, and 𝑡𝑐 𝑡𝑐 the last inequality follows from Eq. (1). Thus, we obtain IR(Π(𝐰)) < ∞ since Var(Π(𝐰)) > 0. Y. Hong and D. Klabjan, Preprint, June 25, 2026
21
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
A.3. Proof of Proposition 3 Proof. Fix 𝑡 ∈ ℕ and 𝑐 ∈ 𝐶̂𝑡 . ̄ ̄ 1 ∈ 𝑡𝑐,CS with 𝐰 ≠ 𝐰. ̄ Assume (i) Let 𝐰 ∈ 𝑡𝑐,LO , 𝐰̄ = 𝟏∕𝑛̂ 𝑡𝑐 ∈ 𝑡𝑐,LO , and 𝐰̌ = (𝐰 − 𝐰)∕‖𝐰 − 𝐰‖ ̄ Δ(𝐰)−Δ(𝐰) ̄ > (1 − ‖𝐰 − 𝐰‖ ̄ 1 )Δ(𝐰). Then, we have Δ(𝐰) > that Δ(𝐰) . Since Δ(⋅) is linear by its definition Eq. (7), ̄ ‖𝐰−𝐰‖ ̄ Δ(𝐰)−Δ(𝐰) 𝐰−𝐰̄ ̌ Thus, we have Δ(𝐰) > Δ(𝐰). ̌ = Δ( ‖𝐰− ) = Δ(𝐰). ̄ 1 ̄ 1 ‖𝐰−𝐰‖ 𝐰‖
1
(ii) For each 𝑑 ∈ 𝐷̂ 𝑡𝑐 , we have
Δ𝑡−1,𝑐𝑑 ≥ ess inf Δ𝑡−1,𝑐𝑑 ≥ Δmin 𝑡−1,𝑐
a.s.,
(A.22)
Δ𝑡−1,𝑐𝑑 ≤ ess sup Δ𝑡−1,𝑐𝑑 ≤ Δmax 𝑡−1,𝑐
a.s.
(A.23)
∑ ∑ = Δmin a.s., where Fix any 𝐰 = [𝑤𝑑 ]𝑑∈𝐷̂ ∈ 𝑡𝑐,LO . From Eq. (7), Δ(𝐰) = 𝑑∈𝐷̂ 𝑤𝑑 Δ𝑡−1,𝑐𝑑 ≥ 𝑑∈𝐷̂ 𝑤𝑑 Δmin 𝑡−1,𝑐 𝑡−1,𝑐 𝑡𝑐 𝑡𝑐 𝑡𝑐 the inequality follows from Eq. (A.22), and the last equality is because 𝐰 ∈ 𝑡𝑐,LO . Therefore, we have (A.24)
ess inf Δ(𝐰) ≥ Δmin . 𝑡−1,𝑐
𝐰∈𝑡𝑐,LO
Redefine 𝐰 = [𝑤𝑑 ]𝑑∈𝐷̂ ∈ 𝑡𝑐,CS . From Eq. (7), we obtain 𝑡𝑐
Δ(𝐰) =
∑
[𝑤𝑑 ]+ Δ𝑡−1,𝑐𝑑 −
𝑑∈𝐷̂ 𝑡𝑐
≤
∑
∑
[−𝑤𝑑 ]+ Δ𝑡−1,𝑐𝑑
𝑑∈𝐷̂ 𝑡𝑐
[𝑤𝑑 ]+ Δmax − 𝑡−1,𝑐
𝑑∈𝐷̂ 𝑡𝑐
= 21 Δmax − 21 Δmin 𝑡−1,𝑐 𝑡−1,𝑐
∑
[−𝑤𝑑 ]+ Δmin 𝑡−1,𝑐
a.s.
𝑑∈𝐷̂ 𝑡𝑐
a.s.,
where the first equality is because 𝑤𝑑 = [𝑤𝑑 ]+ − [−𝑤𝑑 ]+ , the inequality is due to Eqs. (A.22) and (A.23), and the last equality is due to Eq. (5). Hence, we have ess sup Δ(𝐰) ≤ 21 (Δmax − Δmin ). 𝑡−1,𝑐 𝑡−1,𝑐
(A.25)
𝐰∈tc,CS
If 3Δmin > Δmax , then Δmin > 12 (Δmax − Δmin ), yielding ess inf 𝐰∈𝑡𝑐,LO Δ(𝐰) > ess sup𝐰∈𝑡𝑐,CS Δ(𝐰) due to 𝑡−1,𝑐 𝑡−1,𝑐 𝑡−1,𝑐 𝑡−1,𝑐 𝑡−1,𝑐 Eqs. (A.24) and (A.25).
B. CS Weight Projection via Group Demeaning Given a prediction vector 𝐘̂ 𝑡 = [𝑌̂𝑡𝑐𝑑 ](𝑐,𝑑)∈𝑈̂ ∈ ℝ𝑛̂𝑡 , we obtain a projected CS weight vector 𝐰𝑡 by computing: 𝑡
𝐰𝑡 = 𝐘̂ ◦𝑡 ∕‖𝐘̂ ◦𝑡 ‖1
(B.1)
𝐘̂ ◦𝑡 = [𝑌̂𝑡𝑐𝑑 − 𝑌̂𝑡𝑐∙ ](𝑐,𝑑)∈𝑈̂ , 𝑡 ∑ 1 ∙ 𝑌̂𝑡𝑐 = 𝑌̂ ′ . 𝑛̂ 𝑡𝑐 ′ ̂ 𝑡𝑐𝑑
(B.2)
where
(B.3)
𝑑 ∈𝐷𝑡𝑐
Eq. (B.1) corresponds to a scaling that enforces ‖𝐰𝑡 ‖1 = 1. From the proof below, Eqs. (B.2) and (B.3) yield the projection of 𝐘̂ 𝑡 onto {𝐰𝑡 ∈ ℝ𝑛̂𝑡 ∶ 𝟏⊤ 𝐰𝑡𝑐 = 0, ∀𝑐 ∈ 𝐶̂𝑡 }. Therefore, 𝐰𝑡 is, up to scaling, the CS weight vector closest to 𝐘̂ 𝑡 . We now show that Eqs. (B.2) and (B.3) are the Euclidean projection of 𝐘̂ 𝑡 onto {𝐰𝑡 ∈ ℝ𝑛̂𝑡 ∶ 𝐸𝑞. (17)}. The Euclidean projection is given by 𝐘̂ ′𝑡 = (𝐼𝑛̂𝑡 − 𝐺𝑡⊤ (𝐺𝑡 𝐺𝑡⊤ )−1 𝐺𝑡 )𝐘̂ 𝑡
(B.4)
Y. Hong and D. Klabjan, Preprint, June 25, 2026
22
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets Table C.1 Summary of Maturity Date Replacements (Total Contracts: 26,461). The first row reports the number and ratio of issues before the replacements. Each subsequent row reports the number of replacements made and the remaining number of issues. Replacement Type Original Replacement 1 Replacement 2 Replacement 3
Replacements
Remaining Issues
Ratio (%)
0 78 8 147
233 155 147 0
0.88 0.59 0.56 0.00
where 𝐼𝑛̂𝑡 ∈ ℝ𝑛̂𝑡 ×𝑛̂𝑡 is the identity matrix, 𝐺𝑡 = [𝐠𝑡𝑐 ]𝑐∈𝐶̂ ∈ ℝ|𝐶𝑡 |×𝑛̂𝑡 , and 𝐠𝑡𝑐 = [𝕀{𝑐 ′ =𝑐} ](𝑐 ′ ,𝑑 ′ )∈𝑈̂ ∈ ℝ𝑛̂𝑡 , because 𝑡 𝑡 Eq. (17) is equivalent to 𝐺𝑡 𝐰𝑡 = 0. Note that ̂
1 ⊤̂ 𝐠 𝐘. 𝑌̂𝑡𝑐∙ = 𝑛̂ 𝑡𝑐 𝑡𝑐 𝑡
(B.5)
The (𝑐 ′ , 𝑐 ′′ ) element of 𝐺𝑡 𝐺𝑡⊤ is 𝐠⊤ 𝐠 ′′ = 𝑛̂ 𝑡𝑐 ′ 𝕀{𝑐 ′ =𝑐 ′′ } . From this, we have 𝐺𝑡 𝐺𝑡⊤ = diag𝑐∈𝐶𝑡 (𝑛̂ 𝑡𝑐 ) and 𝑡𝑐 ′ 𝑡𝑐 𝐺𝑡⊤ (𝐺𝑡 𝐺𝑡⊤ )−1 𝐺𝑡 𝐘̂ 𝑡
(B.6)
⊤ ̂ = [𝐠𝑡𝑐1 𝐠𝑡𝑐2 … ] ⋅ diag𝑐∈𝐶𝑡 (1∕𝑛̂ 𝑡𝑐 ) ⋅ [𝐠⊤ 𝑡𝑐1 ; 𝐠𝑡𝑐2 ; … ]𝐘𝑡
(B.7)
∑ 1 ̂ = 𝐠𝑡𝑐 𝐠⊤ 𝑡𝑐 𝐘𝑡 𝑛 ̂ 𝑡𝑐 ̂ 𝑐∈𝐶𝑡 ∑ = 𝑌̂𝑡𝑐∙ 𝐠𝑡𝑐 ∈ ℝ𝑛̂𝑡
(B.8) (B.9)
𝑐∈𝐶̂𝑡
by Eq. (B.5), where 𝑐1 < 𝑐2 are the two smallest elements in 𝐶̂𝑡 . The component of Eq. (B.9) corresponding to (𝑐, 𝑑) is the commodity-wise average 𝑌̂𝑡𝑐∙ for 𝑐. From this and Eq. (B.4), the component of 𝐘̂ ′𝑡 = (𝐼𝑛̂𝑡 − 𝐺𝑡⊤ (𝐺𝑡 𝐺𝑡⊤ )−1 𝐺𝑡 )𝐘̂ 𝑡 corresponding to (𝑐, 𝑑) is 𝑌̂𝑡𝑐𝑑 − 𝑌̂𝑡𝑐∙ . Hence, we have 𝐘̂ ′𝑡 = [𝑌̂𝑡𝑐𝑑 − 𝑌̂𝑡𝑐∙ ](𝑐,𝑑)∈𝑈̂ , implying that Eqs. (B.2) and (B.4) 𝑡 coincide. Therefore, Eqs. (B.2) and (B.3) are the Euclidean projection of 𝐘̂ 𝑡 onto {𝐰𝑡 ∈ ℝ𝑛̂𝑡 ∶ 𝟏⊤ 𝐰𝑡𝑐 = 0, ∀𝑐 ∈ 𝐶̂𝑡 }.
C. Data Processing The maturity dates of futures contracts are initially estimated using their expiration dates, i.e., the LastTrdDate field in the DSFutContrInfo table of the LSEG Datastream database provided by WRDS. However, for some contracts, the estimated maturity dates based on LastTrdDate differ substantially from the expiration month recorded in the ContrDate field of the same table, which reports only the year and month rather than the exact date. To address this discrepancy, we further refine the estimated maturity dates by applying the following replacements in sequence: • Replacement 1: a valid EndDate field in the DSFutContrChg table, • Replacement 2: a valid MaxDate, or • Replacement 3: AppContrDate. Following LSEG’s recommendation, we use the EndDate field. MaxDate denotes the latest date on which the contract is traded in the price-volume table DSFutContrVal. Here, AppContrDate denotes the approximated maturity date. It is computed by adding the median difference between LastTrdDate and the 15th day of ContrDate, calculated from contracts with the same underlying commodity, to the 15th day of ContrDate. If the median is unavailable, we use the 15th day of ContrDate. For a value of EndDate and MaxDate to be valid, it should not be significantly different from AppContrDate. In addition, EndDate should not be less than MaxDate to be valid. We use the DSFutContrVal table for the price-volume data, along with the DSFutContrInfo table. Following the indication of the CurrUnitCode field in the DSFutContrInfo table, we convert the unit of the price data. Y. Hong and D. Klabjan, Preprint, June 25, 2026
23
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets Table D.1 Hyperparameter Selection Counts under Yearly Cross-Validation HGL #params 𝑙conv 𝜌∗
CONV
10000 100000 1 2 3 0.1 0.2 0.3 GAT GCN SAGE
10 3 7 10
3 2 5
(A1)
(A2)
3 7 1 1 8
10 7 3 9 1
10
1 2 7
Individual
MLP
5 5 1 4 5 4 2 4
6 4 6 4
10
Figure E.1: Time-Series Characteristics of the Trading Universe.
However, some contracts whose underlying commodity has other contracts requiring unit conversion, as indicated by CurrUnitCode, are not indicated for scaling but exhibit substantially different price scales. We therefore rescale these price data accordingly.
D. Hyperparameter Selection Statistics Each entry in Table D.1 reports the number of times a hyperparameter value is selected across the cross-validation periods.
E. Trading Universe Characteristics Figure E.1 summarizes the time-series characteristics of the trading universe used in our experiments.
F. Annual Trading Performance Metrics Tables F.1 and F.2 present annual trading performance metrics.
Y. Hong and D. Klabjan, Preprint, June 25, 2026
24
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Table F.1 Annual Trading Performance Metrics (IR: information ratio; SR: Sortino ratio; Ret(%): average daily return (%); Vol(%): daily return volatility (%); Hit: hit ratio; MDD: maximum drawdown (absolute); Tvr: average daily turnover; Cor: correlation with the S&P 500)
Metric
Year
IR
2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025
SR
Ret (%)
Vol (%)
MDD
Hit
HGL
Ridge
MLP
CS LGBM
GNN
(A1)
(A2)
(B1)
(B2)
(B3)
(B4)
EW
S&P 500
0.0611 0.1800 0.1211 -0.0174 0.1214 0.0185 0.1123 0.0478 0.0906 0.1435 0.0909 0.2980 0.1719 -0.0236 0.1591 0.0280 0.2017 0.0719 0.1340 0.2469 0.0043 0.0101 0.0120 -0.0013 0.0213 0.0023 0.0231 0.0048 0.0111 0.0160 0.0708 0.0560 0.0990 0.0755 0.1752 0.1251 0.2057 0.1014 0.1222 0.1112 0.0126 0.0046 0.0108 0.0084 0.0187 0.0142 0.0200 0.0083 0.0097 0.0077 0.5397 0.5720 0.5595 0.4841 0.5573 0.5198 0.5179 0.5280 0.5397 0.5783
0.0110 -0.0047 -0.0132 0.0036 0.0448 0.0100 0.0243 0.0656 0.0605 0.1225 0.0150 -0.0070 -0.0149 0.0066 0.0561 0.0154 0.0303 0.0963 0.0830 0.2250 0.0012 -0.0004 -0.0022 0.0004 0.0091 0.0016 0.0059 0.0089 0.0093 0.0182 0.1053 0.0852 0.1682 0.1016 0.2041 0.1583 0.2423 0.1360 0.1530 0.1486 0.0154 0.0144 0.0300 0.0104 0.0230 0.0216 0.0507 0.0169 0.0152 0.0160 0.5119 0.4680 0.4960 0.4722 0.5296 0.5278 0.5020 0.5320 0.5278 0.5020
-0.0245 0.0536 0.0567 -0.0587 0.0494 -0.0964 0.0756 0.0599 0.0146 0.1179 -0.0360 0.0969 0.0798 -0.0778 0.0715 -0.1471 0.1477 0.1092 0.0241 0.1766 -0.0019 0.0046 0.0071 -0.0050 0.0127 -0.0120 0.0188 0.0075 0.0017 0.0153 0.0796 0.0860 0.1251 0.0849 0.2567 0.1242 0.2490 0.1245 0.1182 0.1300 0.0151 0.0103 0.0177 0.0186 0.0319 0.0289 0.0205 0.0100 0.0171 0.0091 0.5119 0.5080 0.4921 0.5119 0.5336 0.4603 0.4861 0.5160 0.5119 0.5823
0.0017 0.1394 0.0746 -0.0299 0.1049 -0.0888 0.0525 0.0482 0.0225 0.0997 0.0024 0.2430 0.1149 -0.0414 0.1692 -0.1299 0.0998 0.0898 0.0330 0.1751 0.0002 0.0096 0.0091 -0.0028 0.0229 -0.0113 0.0129 0.0070 0.0033 0.0138 0.0917 0.0690 0.1216 0.0931 0.2180 0.1276 0.2451 0.1450 0.1474 0.1386 0.0144 0.0059 0.0190 0.0144 0.0286 0.0277 0.0267 0.0152 0.0145 0.0094 0.5159 0.5680 0.5040 0.5040 0.5573 0.4603 0.4861 0.5240 0.5159 0.5261
0.0503 0.0618 -0.0292 -0.0663 0.0797 -0.1063 0.0244 0.0359 -0.0023 0.0781 0.0697 0.0979 -0.0352 -0.0824 0.1123 -0.1421 0.0434 0.0605 -0.0036 0.1406 0.0055 0.0052 -0.0056 -0.0077 0.0212 -0.0159 0.0065 0.0043 -0.0003 0.0122 0.1096 0.0847 0.1901 0.1156 0.2654 0.1496 0.2651 0.1205 0.1373 0.1568 0.0154 0.0075 0.0346 0.0258 0.0237 0.0414 0.0441 0.0150 0.0204 0.0168 0.5198 0.5360 0.4802 0.5000 0.5296 0.4405 0.4861 0.5200 0.4960 0.5301
0.0547 0.1413 0.1412 -0.0111 0.1004 0.0181 0.0785 0.0456 0.0982 0.1337 0.0723 0.2433 0.2136 -0.0124 0.1412 0.0283 0.1316 0.0723 0.1383 0.2114 0.0037 0.0076 0.0134 -0.0008 0.0143 0.0019 0.0151 0.0043 0.0096 0.0137 0.0679 0.0536 0.0951 0.0748 0.1420 0.1039 0.1919 0.0947 0.0982 0.1022 0.0090 0.0066 0.0066 0.0101 0.0144 0.0152 0.0209 0.0089 0.0072 0.0088 0.5357 0.5600 0.5714 0.5198 0.5375 0.5119 0.5179 0.5280 0.5317 0.5301
0.0020 0.0769 0.0514 -0.0258 0.0979 -0.0603 0.0566 0.0763 0.0168 0.1228 0.0028 0.1216 0.0797 -0.0334 0.1396 -0.0917 0.0933 0.1279 0.0244 0.2058 0.0002 0.0056 0.0063 -0.0025 0.0212 -0.0077 0.0125 0.0093 0.0020 0.0141 0.0776 0.0723 0.1232 0.0956 0.2162 0.1273 0.2211 0.1222 0.1200 0.1148 0.0145 0.0063 0.0142 0.0126 0.0186 0.0199 0.0267 0.0095 0.0180 0.0075 0.5000 0.5320 0.5079 0.5437 0.5613 0.4762 0.5020 0.5400 0.5119 0.5382
0.0592 0.0186 -0.0394 0.0356 0.0343 0.1468 0.0618 -0.0605 -0.0218 0.0423 0.1058 0.0295 -0.0523 0.0529 0.0402 0.1908 0.0822 -0.0959 -0.0353 0.0589 0.0449 0.0093 -0.0217 0.0198 0.0343 0.1228 0.0786 -0.0463 -0.0143 0.0328 0.7580 0.4986 0.5518 0.5556 0.9990 0.8360 1.2709 0.7647 0.6548 0.7745 0.0905 0.0746 0.1186 0.0881 0.2884 0.0668 0.1862 0.1539 0.1362 0.0939 0.5198 0.5360 0.5040 0.5357 0.5573 0.5833 0.5697 0.4760 0.5159 0.5141
0.0625 0.0229 -0.0343 0.0372 0.0387 0.1461 0.0639 -0.0612 -0.0152 0.0513 0.1109 0.0370 -0.0458 0.0554 0.0462 0.1920 0.0873 -0.0963 -0.0248 0.0712 0.0459 0.0112 -0.0187 0.0201 0.0372 0.1206 0.0774 -0.0448 -0.0097 0.0389 0.7350 0.4913 0.5472 0.5404 0.9597 0.8257 1.2109 0.7316 0.6422 0.7583 0.0852 0.0673 0.1087 0.0798 0.2724 0.0682 0.1683 0.1440 0.1319 0.0886 0.5159 0.5480 0.5198 0.5357 0.5613 0.5754 0.5498 0.4720 0.5119 0.5100
0.0649 0.0260 -0.0341 0.0377 0.0399 0.1454 0.0700 -0.0575 -0.0129 0.0546 0.1161 0.0420 -0.0455 0.0563 0.0476 0.1905 0.0961 -0.0906 -0.0210 0.0764 0.0477 0.0130 -0.0188 0.0206 0.0388 0.1204 0.0860 -0.0428 -0.0085 0.0426 0.7362 0.4977 0.5503 0.5455 0.9708 0.8282 1.2292 0.7449 0.6593 0.7797 0.0825 0.0686 0.1103 0.0783 0.2737 0.0674 0.1677 0.1447 0.1291 0.0893 0.5238 0.5480 0.5159 0.5278 0.5534 0.5754 0.5578 0.4640 0.5079 0.5100
0.0607 0.0257 -0.0326 0.0343 0.0419 0.1489 0.0681 -0.0586 -0.0169 0.0490 0.1084 0.0412 -0.0436 0.0508 0.0495 0.1951 0.0921 -0.0919 -0.0274 0.0691 0.0461 0.0129 -0.0180 0.0191 0.0421 0.1238 0.0856 -0.0448 -0.0112 0.0383 0.7598 0.5004 0.5512 0.5579 1.0046 0.8320 1.2573 0.7647 0.6639 0.7803 0.0896 0.0727 0.1115 0.0870 0.2842 0.0668 0.1749 0.1557 0.1325 0.0924 0.5238 0.5440 0.5040 0.5397 0.5573 0.5754 0.5657 0.4680 0.5119 0.5181
0.0587 0.0181 -0.0398 0.0354 0.0339 0.1469 0.0612 -0.0607 -0.0225 0.0416 0.1050 0.0287 -0.0529 0.0526 0.0397 0.1908 0.0813 -0.0962 -0.0363 0.0577 0.0446 0.0090 -0.0220 0.0197 0.0339 0.1230 0.0780 -0.0465 -0.0148 0.0322 0.7600 0.4988 0.5522 0.5566 1.0016 0.8370 1.2742 0.7667 0.6549 0.7745 0.0912 0.0751 0.1194 0.0889 0.2896 0.0667 0.1876 0.1547 0.1368 0.0942 0.5198 0.5360 0.5040 0.5357 0.5573 0.5833 0.5737 0.4760 0.5119 0.5141
0.0614 0.1712 -0.0324 0.1474 0.0367 0.1261 -0.0495 0.1009 0.1202 0.0541 0.0804 0.2296 -0.0395 0.1871 0.0417 0.1787 -0.0793 0.1665 0.1585 0.0680 0.0505 0.0723 -0.0351 0.1136 0.0797 0.1033 -0.0754 0.0833 0.0959 0.0640 0.8211 0.4221 1.0831 0.7711 2.1705 0.8190 1.5232 0.8261 0.7974 1.1835 0.0820 0.0281 0.2145 0.0699 0.3825 0.0527 0.2562 0.1065 0.0872 0.2041 0.5238 0.5720 0.5198 0.5952 0.5731 0.5675 0.4303 0.5440 0.5714 0.5743
Y. Hong and D. Klabjan, Preprint, June 25, 2026
LO
25
Hierarchical Graph Learning for Calendar Spread Strategies in Commodity Futures Markets
Table F.2 Annual Trading Performance Metrics (IR: information ratio; SR: Sortino ratio; Ret(%): average daily return (%); Vol(%): daily return volatility (%); Hit: hit ratio; MDD: maximum drawdown (absolute); Tvr: average daily turnover; Cor: correlation with the S&P 500)
Metric
Year
Tvr
2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025
Cor
HGL
Ridge
MLP
CS LGBM
GNN
(A1)
(A2)
(B1)
(B2)
(B3)
(B4)
EW
S&P 500
0.5276 0.5583 0.6625 0.6274 0.5313 0.6155 0.5350 0.5039 0.5400 0.5149 0.0300 0.0426 0.0235 -0.0898 0.0267 -0.0747 -0.0512 0.0499 -0.0012 -0.1068
0.6454 0.6547 0.6705 0.4144 0.6429 0.6048 0.7176 0.8716 0.4508 0.6623 0.0181 -0.0565 0.1229 -0.0621 0.1062 -0.0416 0.0293 0.0573 0.0017 0.0305
0.4692 0.6300 0.4946 0.5628 0.5680 0.4857 0.5349 0.4988 0.4357 0.5022 0.1579 0.0582 0.0838 -0.0365 0.0783 -0.0602 0.0167 0.0062 0.0170 -0.1755
0.6413 0.7193 0.6540 0.6654 0.6852 0.6504 0.5913 0.6557 0.6471 0.6805 0.1066 0.0358 0.1115 -0.0151 0.0943 -0.1115 0.0275 0.0155 -0.0009 -0.0624
0.5683 0.6138 0.7327 0.6017 0.5764 0.5687 0.5300 0.5084 0.6230 0.6556 0.1552 0.0161 0.0506 -0.0259 0.0663 -0.0496 0.0216 0.0741 0.0515 0.0287
0.5884 0.5790 0.5235 0.5541 0.5615 0.4720 0.6061 0.4757 0.4480 0.5749 0.0461 0.0129 0.0588 0.0345 -0.0394 -0.1307 -0.0444 0.0283 0.0102 -0.1505
0.5114 0.5902 0.5890 0.5415 0.5032 0.4963 0.5649 0.5332 0.5234 0.5621 0.1137 0.0352 0.1405 -0.0353 0.0140 -0.0737 -0.0025 0.0493 0.0144 -0.0576
0.0582 0.0554 0.0500 0.0524 0.0472 0.0510 0.0483 0.0462 0.0471 0.0496 0.3874 0.0387 0.3021 0.3595 0.5252 0.2850 0.1496 0.2752 0.1219 0.3426
0.2663 0.2882 0.3197 0.3078 0.2592 0.2984 0.2646 0.2595 0.2696 0.2667 0.3834 0.0288 0.3035 0.3561 0.5189 0.2754 0.1509 0.2799 0.1272 0.3353
0.2512 0.2566 0.2444 0.2735 0.2532 0.2353 0.2492 0.2120 0.2576 0.2881 0.3816 0.0319 0.3048 0.3538 0.5224 0.2806 0.1492 0.2789 0.1261 0.3294
0.2262 0.2318 0.2257 0.2481 0.2231 0.2180 0.2262 0.1948 0.2159 0.2324 0.3892 0.0405 0.3027 0.3554 0.5257 0.2833 0.1483 0.2779 0.1196 0.3340
0.0547 0.0520 0.0431 0.0469 0.0430 0.0467 0.0466 0.0433 0.0426 0.0462 0.3878 0.0391 0.3017 0.3599 0.5252 0.2854 0.1497 0.2748 0.1218 0.3434
1.0000 1.0000 1.0000 1.0000 1.0000 1.0000 1.0000 1.0000 1.0000 1.0000
Y. Hong and D. Klabjan, Preprint, June 25, 2026
LO
26