ConceptioArchivearXiv CS
arXiv CSopen access

A Semiparametric Framework for Stochastic Fundamental Diagram Modeling

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

A Semiparametric Framework for Stochastic Fundamental Diagram Modeling Pengnan Chia , Xiaoliang Maa , Magnus Janssona , Magnus Nordenvaadb a KTH Royal Institute of Technology, Stockholm, Sweden b Swedish Transport Administration, Borlänge, Sweden

arXiv:2607.15907v1 [cs.LG] 17 Jul 2026

Abstract The stochastic fundamental diagram (SFD) provides a probabilistic description of the relationship between traffic density and flow or speed, enabling uncertainty-aware traffic modeling. However, existing stochastic models frequently struggle to accommodate rigorous physical constraints while retaining sufficient flexibility to capture complex nonlinear patterns. To address this, we propose a novel semiparametric SFD modeling framework by leveraging specially designed functional forms. These functions intrinsically satisfy physical constraints defined on the moments of the conditional flow distribution given traffic density while incorporating neural-network-based structures to capture complex empirical patterns. We derive a system of moment-matching equations to convert physical constraints into the parameterization of the conditional distribution, proving that a unique solution exists for the location–scale family of distributions, thereby guaranteeing model well-posedness. Furthermore, we demonstrate that the framework can be extended to non-location–scale distributions, including those requiring additional boundary constraints. Empirical evaluations on the real-world dataset reveal that our approach consistently outperforms representative baselines, delivering superior probabilistic accuracy and robust uncertainty quantification, particularly in congested regimes. Overall, the proposed framework provides a theoretically grounded and flexible foundation for stochastic traffic flow modeling. Keywords: Stochastic fundamental diagram; Semiparametric modeling; Traffic flow modeling; Heteroscedastic uncertainty; Neural network.

1. Introduction The fundamental diagram (FD) describes the intrinsic relationship among three macroscopic traffic variables, i.e., traffic density ρ, flow rate q, and space-mean speed v, under steady-state conditions. Since traffic flow is the product of speed and density, the fundamental diagram is typically represented by either a speed-density or flow-density relationship. Studies on the fundamental diagram provide critical insights on the vehicle movement patterns on a road segment, enabling the development of models for traffic simulation, congestion detection, and traffic management strategies. Classical FD models typically assume a deterministic relationship between traffic density and speed, where each density corresponds to a uniquely speed value. The seminal work of the Greenshields model (Greenshields et al., 1935) establishes a linear speed-density relationship, followed by several refinements such as the Greenberg model (Greenberg, 1959) and the Newell model (Newell, 1961). Although these classical models have provided valuable insights into traffic flow theory, they inherently overlook the stochastic fluctuations and heterogeneous driver behaviors that characterize traffic in real world. This limitation has motivated the development of stochastic approaches to FD modeling, which aims to capture inherent uncertainty present in real-world traffic observations. To better capture the random characteristics inherent in road traffic flow, the stochastic fundamental diagram (SFD) has been proposed. Instead of mapping density to a single deterministic flow value, the SFD characterizes a probability distribution of flow or speed for a given density. Existing SFD models, however, are often closely tied to deterministic FD formulations. One common approach treats the parameters of a deterministic model, such as free-flow speed and Email addresses: [email protected] (Pengnan Chi), [email protected] (Xiaoliang Ma), [email protected] (Magnus Jansson), [email protected] (Magnus Nordenvaad)

jam density, as random variables that follow a prescribed distribution (Jabari et al., 2014; Zhou and Zhu, 2020; Bai et al., 2021; Cheng et al., 2024). Another prevalent approach is to integrate deterministic FD structures with stochastic processes to represent random traffic dynamics (Ni et al., 2018; Lei et al., 2024; Cheng et al., 2022). While conceptually well-defined, these approaches inherit limitations from the deterministic models on which they are built. In particular, widely used single-regime deterministic models often have limited capacity to represent complex traffic states and transitions across the full range of observed densities (Ni, 2015). Consequently, relying on predetermined deterministic forms in certain traffic regimes may introduce systematic bias that undermines modeling accuracy. In addition, there is no consensus on the appropriate deterministic formulation to adopt in SFD modeling (Lei et al., 2024; Cheng et al., 2022; Wang et al., 2021). This lack of agreement stems from the different underlying assumptions adopted by existing models, making it difficult to establish clear and unified guidelines for model selection. These limitations motivate the development of a modeling framework with enhanced nonlinear modeling capabilities as well as strict adherence to traffic flow physics. In this study, we adopt a semiparametric modeling strategy that integrates physically motivated constraints with data-driven components. Rather than relying on a predefined deterministic fundamental diagram, the proposed approach enforces rigorously-defined physical properties while allowing functional relationships to be learned from data. Specifically, a semiparametric formulation is constructed in which physically constrained structures ensure consistency with traffic flow theory, while neural network (NN) based correction terms are introduced to capture complex patterns observed in empirical traffic data. In addition, traffic flow conditioned on density is modeled through a probabilistic distribution capable of representing asymmetric variability commonly observed in real traffic conditions. The distribution parameters are modeled as semiparametric functions of density, enabling the model to represent the full conditional distribution of traffic flow across different density regimes. The main contributions of this paper are summarized as follows: • A semiparametric modeling framework is proposed for estimating the stochastic flow-density relationship in the traffic fundamental diagram, which enforces physical constraints while flexibly capturing complex patterns in empirical data. • A physically constrained parameterization strategy is developed for selected probabilistic distributions; the strategy is proven to be valid for the location-scale family and demonstrates strong extensibility to other complex distributions. • The proposed approach is validated using extensive experiments on real-world traffic data, with appropriate evaluation workflow to assess modeling accuracy and uncertainty quantification. 2. Related Works 2.1. Deterministic Fundamental Diagrams Traditionally, deterministic fundamental diagrams characterize traffic flow primarily through the speed–density relationship. The Greenshields model (Greenshields et al., 1935), which assumes a linear relationship between speed and density, was the first to formalize this concept. Since then, extensive research has sought to improve the accuracy and interpretability of deterministic fundamental diagram models. Based on their functional forms, these models are commonly classified into three major categories: logarithmic, exponential, and polynomial formulations. The detailed mathematical expressions for representative models in each category are summarized in Table 1. 2.2. Stochastic Fundamental Diagrams The most common approach to modeling SFD is to extend existing deterministic models by introducing random parameters. In this framework, selected parameters of deterministic models are treated as random variables to capture variability in steady-state traffic conditions. For example, Jabari et al. (2014) and Zhou and Zhu (2020) modeled microscopic parameters such as headway and speed adaptation time as random variables and derived the macroscopic fundamental diagram relations under steady-state assumptions. Similarly, Qu et al. (2017) and Wang et al. (2021) estimated model parameters at different distribution quantiles, resulting in discrete stochastic representations of the fundamental diagram. model with random variables and obtained the stochastic speed-density relationship. Building on this approach, Ahmed et al. (2021) further analyzed traffic flow heterogeneity arising from different vehicle types. 2

Category

Model

Polynomial function-based

Greenshields Greenshields et al. (1935) Jayakrishnan Jayakrishnan et al. (1995) S3 Cheng et al. (2021)

Formulation ρ v = vfree (1 − ρ jam )   ρ v = vmin + (vfree − vmin ) 1 − ρjam

v= 

vfree  m  m2 1+ ρ ρ critical

Exponential function-based

NewellNewell (1961) Papageorgiou Papageorgiou et al. (1989) Del Castillo Del Castillo and Benitez (1995) Wang Wang et al. (2011)

Logarithmic function-based

Greenberg Greenberg (1959)

    λ 1 1 v = vfree 1 − exp − vfree − ρ ρjam   α  ρ v = vfree exp − α1 ρjam n io h vjam  ρ v = vfree 1 − exp vfree 1 − jam ρ v = vcritical + 

vfree −vcritical  θ2

1+exp

v = vcritical log

ρ−ρcritical θ1

 ρjam  ρ

Table 1: Three Categories of Deterministic Models

In addition to randomizing existing model parameters, Bai et al. (2021) introduced additional random variables to explicitly represent traffic heterogeneity. Cheng et al. (2024) investigated analytical properties of the model parameter when making other parameters as random variables. In addition to random parameter formulations, another commonly adopted approach is to treat a deterministic fundamental diagram as the expected value and introduce stochasticity through an explicit random process. In this framework, the deterministic model describes the mean traffic behavior while stochastic variability is modeled separately. For example, Ni et al. (2018) applied a complex deterministic model (Ni et al., 2016) for the expected values and characterized the stochasticity using the Maxwell–Boltzmann distribution inspired by ideal gas theory. Similarly, Liu et al. (2023) extend the deterministic fundamental diagrams using Gaussian process regression to incorporate additional traffic variables. Along the same line, Lei et al. (2024)and Cheng et al. (2022) adopted the deterministic models as the expectation function and introduced stochasticity with Gaussian process regression. Some studies depart entirely from deterministic fundamental diagram formulations. Instead, they derive stochastic traffic relationships directly from traffic flow dynamics or alternative stochastic representations. For example, several studies model traffic evolution based on the continuity equation to describe macroscopic traffic dynamics without relying on predefined fundamental diagrams (Shi et al., 2021; Storm et al., 2022; Ngoduy, 2011). In a different line of study, Zhang et al. (2025) and Zhang et al. (2025) described the platoon behavior using Markov chains and derived the macroscopic stochastic FD. Moreover, Bramich et al. (2023) employed an advanced statistical regression framework, namely generalized additive models for location, scale, and shape (GAMLSS) (Rigby and Stasinopoulos, 2005), to develop the SFD model. 3. Methodology 3.1. Problem Definition This study treats traffic flow as a random variable conditioned on density. Let ρ ∈ [0, ρjam ] and q ∈ R≥0 denote the traffic density and traffic flow, respectively, where ρjam is the traffic jam density. The objective is to model the conditional probability distribution of traffic flow given density, denoted by p(q | ρ). We assume that, for each fixed density ρ, the conditional distribution p(q | ρ) admits an analytical probability density function (PDF) belonging to a parametric family fPDF , parameterized by an n-dimensional vector η ∈ Rn :  p(q | ρ) = fPDF q | ρ; η . 3

(1)

The admissible distributions are required to satisfy fundamental physical properties of traffic flow. Under the zero-density condition (ρ = 0), the absence of vehicles implies that both traffic flow and its variability vanish. Under the jam-density condition (ρ = ρjam ), vehicles are fully congested, resulting in zero flow with negligible variability. For intermediate density, the traffic flow value must be positive, and its variability must also be positive. These physical requirements impose the following moment-based constraints on the conditional distribution p(q | ρ). Let E[·] and S[·] denote the expectation and standard deviation, respectively:    Eq∼p(q|ρ) [q] = Sq∼p(q|ρ) [q] = 0, ρ ∈ {0, ρjam },     (2) 0 < Eq∼p(q|ρ) [q] < ∞, ρ ∈ (0, ρjam ),      0 < Sq∼p(q|ρ) [q] < ∞, ρ ∈ (0, ρjam ). The modeling problem addressed in this study is therefore to construct a flexible yet physically consistent parameterization of p(q | ρ) that satisfies these constraints while accurately capturing the stochastic variability observed in empirical traffic data. 3.2. Semiparametric Framework for SFD Models A common data-driven approach to modeling stochastic traffic flow relationships is to learn distributional parameters as functions of traffic density using nonparametric models such as neural networks (Bishop, 1994; Theis and Bethge, 2015). While such approaches offer substantial modeling flexibility and expressive capability, they do not inherently guarantee physical consistency with respect to the fundamental properties of traffic flow. In particular, without explicitly enforcing the boundary and non-negativity conditions specified in Eq. (2), unconstrained nonparametric models may produce physically implausible behavior, especially in regimes near zero density or jam density. To address this limitation, this study adopts a semiparametric modeling framework in which physical constraints are explicitly embedded into the parameterization of the conditional distribution p(q | ρ). The central idea begins with separating the distributional parameters into constrained and unconstrained components. Definition 1 (Partition of Distribution Parameters). Let η ∈ Rn (n ≥ 2) denote the parameter vector governing the conditional distribution p(q | ρ). The parameter vector η is defined by a mutually exclusive partition: h i ηT = ηTcon ηTfree , (3) where: • ηcon ∈ R2 represents the constrained component determined by construction through physical constraints on the first two moments (the conditional expectation and standard deviation). • ηfree ∈ Rn−2 represents the free component that is unconstrained and capture higher-order distributional characteristics. Under this framework, the distributional parameters are derived by integrating data-driven learning with strict physical constraints. The unconstrained parameters ηfree are learned directly from data using a neural network, leveraging its universal approximation capability (Hornik et al., 1989). In contrast, the constrained parameters ηcon are obtained analytically by enforcing consistency between the theoretical moments of the distribution and the macroscopic traffic characteristics. Let µ(η) and σ(η) denote the theoretical expectation and standard deviation derived from the probability density function fPDF (q | ρ; η). Meanwhile, we introduce two semiparametric functions m(ρ) and s(ρ), representing the conditional expectation and standard deviation of traffic flow as functions of traffic density. These functions, whose detailed structure will be discussed in Section 3.3, satisfy the physical constraints defined in Eq. (2) by construction. The consistency between the distributional moments and macroscopic traffic flow characteristics is described by a moment-matching system as follows: h iT  Eq∼ fPDF [q] = µ ηTcon ηTfree = m(ρ; φ, h), h (4) iT  Sq∼ fPDF [q] = σ ηTcon ηTfree = s(ρ; φ, h), 4

where the semiparametric functions consist of a parametric component governed by domain-specific parameters φ (e.g., jam density) and a nonparametric correction term h that captures the deviations from the underling traffic flow relationship. The correction term h and the unconstrained distribution parameters ηfree are generally unknown and must be inferred from data. To provide a flexible representation of their dependence on traffic density, we introduce a neural network mapping that generates these values directly from the density input. That is, a single neural network model DNN(·), parameterized by θ, is employed to jointly model the latent correction term h and the unconstrained parameters ηfree : h iT hT ηTfree = DNN(ρ; θ). (5) The constrained parameters ηcon are subsequently determined by solving the moment-matching system in Eq. (4). For the proposed framework to be well-defined, it is necessary that this system admits a unique solution. To guarantee this property, we establish the following proposition for location–scale families of distributions. Proposition 1 (Unique Solvability for Moment Matching). Suppose the conditional distribution fPDF (q | ρ; η) belongs to a location-scale family with well-defined first and second moments. Let ηcon represent the location and scale parameters. Then, for any intermediate density ρ ∈ (0, ρjam ) and any admissible unconstrained parameter vector ηfree , the moment-matching system defined in Eq. (4) admits a unique solution for ηcon . Proof. Let ηcon = [η1 , η2 ]T , where η1 is the location parameter and η2 > 0 is the scale parameter. By the definition of a location-scale family, a random variable X ∼ fPDF (q | ρ; η) can be expressed as an affine transformation of a base random variable Z: X = η1 + η2 Z, where Z follows the standardized distribution with location η1 = 0 and scale η2 = 1. Crucially, since the shape of Z is governed solely by the unconstrained parameters ηfree , we can express its expectation and variance as explicit functions of ηfree , denoted by µ0 (·) and σ20 (·) respectively: E[Z] = µ0 (ηfree ),

Var[Z] = σ20 (ηfree ).

By the linearity of expectation and properties of variance, the moments of X are given by: E[X] = η1 + η2 µ0 (ηfree ),

Var[X] = η22 σ20 (ηfree ).

For any given density ρ ∈ (0, ρjam ), the corresponding flow expectation m(ρ) and standard deviation s(ρ) are admissible provided they satisfy the boundary constraints in Eq. (2), i.e., 0 < m(ρ) < ∞,

0 < s(ρ) < ∞.

To enforce consistency, the moment-matching system is specified as follows: η1 + η2 µ0 (ηfree ) = m(ρ),

η22 σ20 (ηfree ) = s2 (ρ).

Since η2 > 0 and s(ρ) > 0, taking the positive square root of the variance equation yields the unique solutions for the location and scale parameters: µ0 (ηfree ) s(ρ) η1 = m(ρ) − s(ρ) , η2 = . σ0 (ηfree ) σ0 (ηfree ) The above proposition guarantees that the moment-matching system admits a unique solution for location-scale families of distributions. Consequently, once the semiparametric functions m(ρ; φ, h) and s(ρ; φ, h) and the unconstrained parameters ηfree are specified, the constrained parameters ηcon are uniquely determined. Because the latent correction term h and the unconstrained parameters ηfree are generated by the neural network defined in Eq. (5), the complete parameter vector η becomes a deterministic function of density ρ, parameterized by the neural network weights θ and the physical parameters φ. We denote this mapping as h i ηT (ρ; θ, φ) = ηTcon (ρ; θ, φ) ηTfree (ρ; θ) (6) 5

Figure 1: Flowchart for Stochastic FD modeling under Semiparametric Framework

The resulting modeling procedure can therefore be summarized as the following computational chain: θ

φ

Eq. (5)

Eq. (4)

ρ −−−−−→ (h, ηfree ) −−−−−→ (ηcon , ηfree ) ≡ η −→ fPDF (q | ρ; η).

(7)

N Finally, model parameters are estimated via maximum likelihood using the observed dataset {(ρi , qi )}i=1 :

θ̂, φ̂ = arg max θ,φ

N X

 log fPDF qi | ρi ; η(ρi ; θ, φ) .

(8)

i=1

The semiparametric framework is therefore well defined for location-scale families of distributions, ensuring that the resulting conditional distribution satisfies the imposed physical constraints by model construction. Other distribution families may also be incorporated into the framework, although additional modifications may be required. The extension of the framework will be further discussed in Section 7. The overall modeling procedure is intuitively summarized in Fig. 1. 3.3. Structure Specification of Semiparametric Functions This section specifies candidate functional forms for the expectation m(ρ) and standard deviation s(ρ) of traffic flow q given density ρ, associated with the conditional distribution p(q | ρ). Under the physical constraints in Eq. (2), both m(ρ) and s(ρ) belong to the same function class that satisfies specific boundary conditions. Definition 2 (Class of Physically Admissible Functions). Let Fphy denote the class of functions  Fphy = f ∈ C([0, ρjam ]) | f (ρ) ≥ 0, f (0) = 0, f (ρjam ) = 0 ,

(9)

where C([0, ρjam ]) denotes the space of continuous real-valued functions defined on the interval [0, ρjam ]. Among functions in Fphy , the quadratic flow-density relationship implied by the classical Greenshields model represents the simplest parametric example. Building on this baseline, we introduce two semiparametric extensions that preserve physical admissibility while enabling data-driven flexibility through NN-based correction terms. 3.3.1. Quadratic Function with Neural Correction (QwNC) The classical Greenshields model assumes a linear relationship between traffic density ρ and space-mean speed v. Since traffic flow is defined as q = ρv, this assumption implies a quadratic relationship between density and flow. The Greenshields formulation is given by ! ρ vfree vG (ρ) = vfree 1 − , qG (ρ) = ρ(ρjam − ρ) , (10) ρjam ρjam where vfree denotes the free-flow speed and ρjam the jam density. The resulting quadratic function attains zero flow at both ρ = 0 and ρ = ρjam , and therefore belongs to Fphy . 6

To enhance modeling flexibility while preserving physical consistency, the constant coefficient vfree /ρjam is replaced by a density-dependent correction term c(ρ). This term corresponds to one element of the latent vector h generated by the neural network defined in Eq. (5). The resulting Quadratic Function with Neural Correction is expressed as fQwNC (ρ) = ρ(ρjam − ρ)c(ρ).

(11)

To ensure numerical stability and enforce physical non-negativity, all positive quantities are mapped through the Softplus function σ(x) = ln(1 + e x ). In particular, the jam density is parameterized as ρjam = σ(ρ j ), where ρ j is an element of the physical parameter vector φ. The density-dependent correction term c(ρ) is constrained analogously via σ(c(ρ)), yielding the final formulation   fQwNC (ρ) = ρ σ(ρ j ) − ρ σ c(ρ) . (12) Although the QwNC formulation is motivated by extending the mean flow relationship implied by the Greenshields model, it satisfies the physical constraints in Eq. (2) by construction. Accordingly, the same functional form can be leveraged to model both the conditional expectation and the conditional standard deviation of traffic flow. 3.3.2. Beta Function with Neural Correction (BwNC) To further enhance structural flexibility beyond the quadratic form, we introduce a functional specification inspired by the shape characteristics of the Beta distribution. Specifically, two shape parameters, a and b, are introduced to  transform the quadratic term ρ(σ(ρ j ) − ρ) into the more flexible form ρa σ(ρ j ) − ρ b . The parameters a and b regulate the curvature near low and high density regimes, respectively, enabling the model to capture skewed or asymmetric patterns observed in empirical data. To ensure physical admissibility and numerical stability, the Softplus function is applied to the exponents, enforcing positivity. In addition, the density difference term, (σ(ρ j ) − ρ), is rectified using a max(0, ·) operator to prevent negative values, thereby guaranteeing that the power function remains well-defined across the entire density domain. The resulting Beta Function with Neural Correction is given by   fBwNC (ρ) = ρσ(a) max 0, σ(ρ j ) − ρ σ(b) σ c(ρ) . (13) This formulation preserves the physical boundary conditions and provides additional flexibility to capture asymmetric relationships. The BwNC formulation can be used to model both the expectation and the standard deviation of conditional traffic flow distribution. 4. Specification of SFD Models Given the semiparametric framework introduced in the preceding section, a SFD model is fully specified once the conditional distribution p(q | ρ) and the semiparametric forms of the conditional expectation m(ρ) and standard deviation s(ρ) are defined. In this section, the framework is instantiated by assuming that the conditional distribution follows a Skew-Normal distribution. For completeness, the Gaussian distribution is discussed as a special case obtained by setting skewness parameter as zero. 4.1. Skew-Normal Distribution The stochastic fundamental diagram model is instantiated by assuming the conditional distribution p(q | ρ) to follow the Skew-Normal distribution with parameter as η = {ξ, ω, α}: p(q | ρ) = fSkewNormal (q | ρ; ξ, ω, α) r !  21 1  q − ξ 2 q − ξ = exp − Φ α· , πω 2 ω ω

(14)

where Φ(·) denote the standard normal cumulative distribution function. Following the Proposition 1, location parameter ξ and scale parameter ω are regarded as constrained, i.e., ηcon = [ξ, ω]T . While the skewness parameter α forms the unconstrained parameter ηfree = [α]T , which is learned directly from data. 7

Table 2: Summary of Functional Forms and Parameters for Stochastic FD Models under Skew-Normal Distribution

m(ρ; φ, h) s(ρ; φ, h) φ h DNN mapping DNN(ρ; θ) θ

SN-QwNC   ρ σ(ρ j ) − ρ σ c1   ρ σ(ρ j ) − ρ σ c2 ρj c1 , c2

SN-BwNC

  ρσ(a1 ) max 0, σ(ρ j ) − ρ σ(b1 ) σ c1   ρσ(a2 ) max 0, σ(ρ j ) − ρ σ(b2 ) σ c2 ρ j , a1 , b1 , a2 , b2 c1 , c2 ρ → [c1 c2 α]T W3 silu(W2 silu(W1 ρ + b1 ) + b2 ) + b3 W1 , W2 , W3 , b1 , b2 , b3

Two model variants based on Skew-Normal distribution, namely SN-QwNC and SN-BwNC, are established based on the functional forms defined in Section 3.3. The SN-QwNC model applies the fQwNC for functions of the expectation m(ρ; φ, h) and standard deviation s(ρ; φ, h), whereas the SN-BwNC model utilizes the fBwNC . A summary of the functional forms and parameter sets is provided in Table 2. For both the SN-QwNC and SN-BwNC specifications, the vector h contains two correction components, one associated with the conditional expectation and one with the conditional standard deviation. The domain-specific parameter vector φ differs across the two variants: in the SN-QwNC model, φ contains only density-related parameters (e.g., jam density), whereas in the SN-BwNC model it additionally includes parameters controlling the curvature of the flow–density relationship. The correction terms h and the skewness parameter α are learned jointly using a neural network. Specifically, a three-layer multilayer perceptron (MLP) with Sigmoid Linear Unit (SiLU) activation functions (Hendrycks and Gimpel, 2016) is employed. The constrained parameters {ξ, ω} are determined by matching the theoretical moments of the Skew-Normal distribution to the semiparametric functions m(ρ; φ, h) and s(ρ; φ, h). This yields the system of equations as follows: !1/2 α2 2 = m(ρ; φ, h), 1 + α2 π ! α2 2 2 ω 1− = s2 (ρ; φ, h), 1 + α2 π

ξ+ω

(15)

from which ξ and ω can be expressed explicitly as functions of α, m(ρ; φ, h), and s(ρ; φ, h), i.e., α2 2 ω = s(ρ; φ, h) 1 − 1 + α2 π

!−1/2

α2 2 ξ = m(ρ; φ, h) − ω 1 + α2 π

, (16)

!1/2 .

Consequently, all distributional parameters {ξ, ω, α} become deterministic functions of the density ρ, parameterized by the neural network weights θ and the domain-specific parameters φ. Model parameters are learned by minimizing a composite loss function consisting of the negative log-likelihood and a jam-density regularization term, i.e., ℓ = ℓNLL + λ · ℓReg .

(17)

N Given a dataset D = {(ρi , qi )}i=1 of observed density-flow pairs, the negative log-likelihood under the Skew-Normal model is  N  Y  ℓNLL = − log p(qi |ρi ) i=1 (18) !2 X ! N N 1 X qi − ξ i qi − ξi log Φ αi = N log ωi + − , 2 i=1 ωi ωi i=1

8

Table 3: Summary of Functional Forms and Parameters for Stochastic FD Models under Normal Distribution

m(ρ; φ, h) s(ρ; φ, h) φ h DNN mapping DNN(ρ; θ) θ

N-QwNC   ρ σ(ρ j ) − ρ σ c1   ρ σ(ρ j ) − ρ σ c2 ρj c1 , c2

N-BwNC

  ρσ(a1 ) max 0, σ(ρ j ) − ρ σ(b1 ) σ c1   ρσ(a2 ) max 0, σ(ρ j ) − ρ σ(b2 ) σ c2 ρ j , a1 , b1 , a2 , b2 c1 , c2 ρ → [c1 c2 ]T W3 silu(W2 silu(W1 ρ + b1 ) + b2 ) + b3 W1 , W2 , W3 , b1 , b2 , b3

where hi , αi = DNN(ρi ; θ), −1/2  α2i 2     , ωi = s(ρi ; φ, hi ) 1 − 1 + α2i π  1/2  α2i 2   . ξi = m(ρi ; φ, hi ) − ωi  1 + α2i π A regularization term is introduced to stabilize the optimization process and to enforce physical consistency of the estimated jam density: N X  ℓReg = max ρi − σ(ρ j ), 0 , (19) i=1

where σ(ρ j ) denotes the jam density reparameterized by Softplus function. The rationale of this regularization is to encourage the learned jam density to be no smaller than any observed traffic density in the dataset. Whenever the estimated jam density falls below observed densities, a penalty is incurred to prevent physically implausible solutions. No penalty is applied when the estimated jam density exceeds all observed densities. A relatively large weighting factor λ, e.g., λ = 102 , is employed to ensure effective enforcement of the regularization during the training phase. In practice, the regularization rapidly guides the jam density estimate toward a physically admissible range while maintaining stable convergence of the likelihood-based optimization. 4.2. Normal Distribution The Normal distribution can be considered as a degenerate case of the Skew-Normal distribution obtained by setting the skewness parameter α to zero. Under this assumption, the conditional distribution of traffic flow given density becomes symmetric, and only the location and scale parameters, ξ and ω, are required. Consistent with the semiparametric framework, these parameters are treated as constrained parameters ηcon , while no distributional parameter is left as free, i.e., ηfree = ∅. The conditional probability density function is given by ! 1 1  q − ξ 2 p(q | ρ) = √ exp − . (20) 2 ω 2π ω Analogous to the Skew-Normal specification, two Normal distribution based SFD models are defined: N-QwNC and N-BwNC. These models employ the same semiparametric functional forms for the conditional expectation and standard deviation as their Skew-Normal counterparts. The only structural difference lies in the neural network output, which no longer includes a skewness parameter. A summary of the functional forms and associated parameters is provided in Table 3. For the Normal distribution, the constrained parameters are obtained as a special case of Eq (16) by setting the skewness parameter α = 0. Under this condition, the location and scale parameters reduce to ω = s(ρ; φ, h),

ξ = m(ρ; φ, h). 9

(21)

Since there is no skewness parameter, the neural network only outputs the correction terms associated with the N semiparametric functions. Given a dataset D = {(ρi , qi )}i=1 , the loss function is defined as: ℓ = ℓNLL + λℓReg N X

N

1 X qi − m(ρi ; φ, h) = log s(ρi ; φ, h) + 2 i=1 s(ρi ; φ, h) i=1

!2 +λ

N X

 max ρi − σ(ρ j ), 0 .

(22)

i=1

5. Computational Experiments 5.1. MAGIC Dataset The proposed models are evaluated using the MAGIC (Multiple Conditions UAV Group-based High-fidelity Comprehensive Vehicle Trajectory) dataset (Ma et al., 2022), which provides high-resolution vehicle trajectory data collected by unmanned aerial vehicles (UAVs). The data were recorded over a three-hour period (7:40–10:40 AM) along a 4 km segment of the Shanghai Inner Ring Road. The dataset contains detailed information on vehicle type, position, velocity, and acceleration at a sampling frequency of 25 Hz. From the full dataset, a 200 m two-lane road segment was extracted to ensure coverage of a wide range of traffic conditions, spanning free-flow to congested regimes. Traffic variables were aggregated into 1 second intervals. The aggregated speed is measured in kilometers per hour (km/h), while traffic density is expressed in vehicles per kilometer per lane (veh/km/lane). Traffic flow is computed as the product of density and speed, with units of vehicles per hour per lane (veh/h/lane). 5.2. Experiment Setup A major challenge in our evaluation is the limited and imbalanced nature of the dataset. Since only three hours of tabular data were collected from a single location, the limited data volume introduces high variance in the evaluation metrics. Furthermore, the skewed distribution toward low-density samples could overemphasize model performance in the free-flow state while ignoring the performance in the congested state. Stratified k-fold cross-validation combined with a sample-weighting scheme (as illustrated in Figure 2) is used to address the aforementioned challenge. The procedure is detailed as follows: • Stratified Partitioning. The entire density range is partitioned into b = 10 equal bins, with Nm data points in m-th bin. The data within each bin is further split into k = 5 equal folds. A complete fold is constructed by taking one partition from each bin without replacement, resulting in a total of k folds. • Sample Weighting. For each experimental run, one fold T j ( j ∈ {1, . . . , k}) serves as the test set, while the remaining k − 1 folds are used for training. Each test sample i ∈ T j is weighted inversely proportional to its bin’s total population. Specifically, the weight is 1 wi = , (23) Nb(i) where b(i) is the bin index of sample i. The weight wi is explicitly integrated into the calculation of each specific evaluation metric (detailed in Section 5.3) to ensure an unbiased assessment. Let S j denote the aggregated score of a given metric evaluated on the test fold T j . After iterating through all k folds, the model’s final performance is summarized by the mean (µS ) and standard deviation (σS ) of the k aggregated scores: v u t k k X 2 1 1 X µS = S j , σS = S j − µS . (24) k j=1 k j=1

10

Figure 2: Stratified 5-fold cross-validation combined with a sample-weighting scheme

5.3. Metrics Five metrics are used for a comprehensive evaluation of the proposed models. The first two are probabilistic metrics, which evaluate the quality of the predicted output distribution. The remaining three are deterministic metrics, which assess the accuracy of the predicted conditional mean. For notation clarity across all metrics, let qi denote the ground  truth flow observation for sample i. Let fPDF · | ρi ; η(ρi ; θ, φ) denote the predicted conditional probability density function given density ρi , where the distribution parameters η are produced by a model parameterized by learnable weights θ and φ. 5.3.1. Probabilistic Metrics Weighted Continuous Ranked Probability Score (WCRPS). The CRPS is a proper scoring rule commonly used to evaluate the probabilistic forecasts. For a predictive distribution f and an actual observation y, it is defined as: !2 Z ∞ Z x CRPS( f, y) = f (t)dt − 1 x≥y dx. (25) −∞

−∞

Since integrating the cumulative distribution is often computationally intractable, CRPS can be approximated in a expectation form using Monte Carlo simulation (Gneiting and Raftery, 2007):   1   CRPS( f, y) = EX∼ f |X − y| − EX,X ′ ∼ f |X − X ′ | . (26) 2 When incorporating the sample weights wi , the aggregated Weighted CRPS for the j-th test fold T j is calculated as: P   i∈T j wi · CRPS fPDF · | ρi ; η(ρi ; θ, φ) , qi P . (27) S j,WCRPS = i∈T j wi Weighted Negative Log-Likelihood (WNLL). The Negative Log-Likelihood measures the goodness-of-fit by penalizing low probabilities assigned to the true observations. By incorporating the weight, the WNLL for the test fold is given by: P  i∈T j wi · log fPDF qi | ρi ; η(ρi ; θ, φ) P . (28) S j,WNLL = − i∈T j wi 5.3.2. Deterministic Metrics For deterministic evaluation, we collapse the predictive distribution into a single point estimate q̂i by taking its mathematical expectation: q̂i = EX∼ fPDF(·|ρi ;η(ρi ;θ,φ)) [X]. (29) 11

Weighted Mean Absolute Error (WMAE). The WMAE measures the average magnitude of the absolute errors, weighted by the density bin frequency to prevent bias toward free-flow states: P i∈T j wi |q̂i − qi | P S j,WMAE = . (30) i∈T j wi Root Weighted Mean Square Error (RWMSE). To heavily penalize larger prediction errors while accounting for the skewed dataset, we compute the Root Weighted Mean Square Error as follows: sP 2 i∈T j wi (q̂i − qi ) P S j,RWMSE = . (31) i∈T j wi Weighted Mean Absolute Percentage Error (WMAPE). To provide a relative error perspective, we define the WMAPE as the ratio of the weighted absolute errors to the weighted absolute true values: P i∈T j wi |q̂i − qi | S j,WMAPE = P . (32) i∈T j wi |qi | 5.4. Model Setup During the training phase, the model (configured with a hidden size of 16) is updated in mini-batches of 128 via the Adam optimizer (Kingma and Ba, 2014). We employ a weight decay of 10−5 for regularization, alongside β coefficients of (0.9, 0.99). The learning rate is initially set to 10−2 and explicitly decayed to 10−3 halfway through the 200-epoch training process. To ensure numerical stability, log Φ(·) is computed using the optimized function log_ndtr available in PyTorch (Paszke et al., 2019) and JAX (Frostig et al., 2019). 5.5. Baselines This study utilizes three recently proposed stochastic FD models as baselines, which are representative of two prevailing modeling strategies: 1) random parameter modeling and 2) stochastic error/variance modeling. A common principle in both strategies is the modification of deterministic FD models. To obtain accurate deterministic FD models, the weighted least squares method, adopted in (Qu et al., 2015), is applied to mitigate the influence of unbalanced data during the training phase. 5.5.1. Random parameter modeling The work of Cheng et al. (2024) extends the deterministic S3 model (Cheng et al., 2021) by introducing stochasticity into its parameters. Since the original formulation is not explicitly named in the original paper, we refer to it as S3+LN for clarity. The model is defined as follows. The S3 speed-density and flow-density relations are: vS3 (ρ) = 

vfree ρ 1 + ρcritical



m  m2 ,

ρ vfree qS3 (ρ) =   ρ m  m2 1 + ρcritical

(33)

In the stochastic formulation, the free-flow speed and critical speed are treated as random variables following LogNormal distributions: ln(vfree ) ∼ N(µ f , σ2f ), ln(vcritical ) ∼ N(µc , σ2c ), (34) and the parameter follows

   µ f − µc σ2f + σ2c  1  . ∼ N  , m 2 ln 2 (2 ln 2)2 

(35)

This framework contains five trainable parameters: {µ f , σ f , µc , σc , ρcritical }. The free-flow speed and critical speed are parameterized by {µ f , σ f } and {µc , σc }, respectively. The critical density ρcritical is estimated using the weighted least squares approach. 12

Table 4: Performance of all evaluated stochastic FD models

Flow-Density Relation Model

CRPS (↓)

NLL (↓)

MAE (↓)

RWMSE (↓)

WMAPE% (↓)

S3 + LN (Cheng et al., 2024) S3 + GP (Lei et al., 2024) GS + GP (Lei et al., 2024)

312.977±12.062 165.896±1.902 172.785±2.587

– 7.120±0.013 7.139±0.013

448.944±9.658 179.611±3.073 204.973±5.739

551.809±10.923 270.631±6.595 274.391±8.685

45.594±1.175 18.238±0.216 20.816±0.632

SN-QwNC N-QwNC SN-BwNC N-BwNC

124.936±3.629 127.917±2.564 122.643±3.864 124.444±4.020

6.663±0.037 6.749±0.040 6.633±0.053 6.683±0.046

172.636±5.740 173.303±4.080 171.736±4.616 172.594±5.425

247.023±8.740 247.880±7.560 246.799±8.124 247.220±8.621

17.529±0.531 17.597±0.338 17.438±0.413 17.525±0.488

Speed-Density Relation Model

CRPS (↓)

NLL (↓)

MAE (↓)

RWMSE (↓)

WMAPE% (↓)

S3 + LN (Cheng et al., 2024) S3 + GP (Lei et al., 2024) GS + GP (Lei et al., 2024)

4.566±0.130 2.997±0.049 3.069±0.034

– 3.166±0.014 3.185±0.013

6.637±0.142 3.700±0.061 3.929±0.072

7.795±0.180 6.894±0.161 5.493±0.149

24.756±0.573 13.801±0.254 14.654±0.291

SN-QwNC N-QwNC SN-BwNC N-BwNC

2.590±0.072 2.661±0.047 2.548±0.057 2.592±0.059

2.708±0.039 2.795±0.041 2.679±0.054 2.728±0.048

3.605±0.084 3.634±0.062 3.594±0.069 3.606±0.068

5.320±0.140 5.368±0.110 5.322±0.128 5.321±0.127

13.448±0.338 13.556±0.259 13.407±0.279 13.450±0.271

5.5.2. Stochastic Error/Variance Modeling The work of Lei et al. (2024) extends deterministic fundamental diagram (FD) models by modeling the residual error using a Gaussian Process (GP). In this formulation, the overall speed v(ρ) is represented as the sum of a deterministic baseline model fdet (ρ) and a stochastic residual term ϵ(ρ): v(ρ) = fdet (ρ) + ϵ(ρ),  ϵ(ρ) ∼ GP 0, k(ρ, ρ′ ) ,

(36)

where the error term ϵ(ρ) follows a zero-mean Gaussian Process governed by a Radial Basis Function (RBF) kernel k(ρ, ρ′ ): ! (ρ − ρ′ )2 ′ 2 k(ρ, ρ ) = σ exp − . (37) 2l2 This kernel introduces two additional trainable hyperparameters: the signal variance σ2 and the length-scale l, which control the magnitude and smoothness of the stochastic component, respectively. The Greenshields and S3 deterministic models are used as baselines for the GP framework, resulted in two variants of models denoted as GS+GP and S3+GP. 6. Results 6.1. Metric Performance Performance metrics for probabilistic estimation across all evaluated models are summarized in Table 4. Five evaluation metrics are reported for both speed-density and flow-density relations. The reported values represent statistical mean and standard deviation using 5-fold cross-validation. The NLL metric is omitted for the S3+LN model because its probability density function does not admit a close-form expression. Overall, the experimental results reveal that stochastic FD models developed under the proposed semiparametric framework consistently outperform the baseline approaches. Among all baseline models, the S3+LN model exhibits the poorest performance across most of the metrics. In contrast, the Gaussian Process (GP) regression models achieve substantially better performance relative to the S3+LN approach. 13

(a) SN-BwNC

(b) N-BwNC

(c) SN-QwNC

(d) N-QwNC

Figure 3: Modeling results from Stochastic FD Models under the Semiparametric Framework

The table also reveals performance differences among the model variants. For the baseline models, the GP regression based on the deterministic S3 model outperforms the formulation based on the Greenshields model. This advantage is consistent across cross-validation folds, as reflected by the relatively small standard deviations. 14

(a) Greenshields Model with Gaussian Process

(b) S3 Model with Gaussian Process

(c) S3 Model with Log-Normal Distribution

Figure 4: Modeling results from Baseline Stochastic FD Models

Among the models proposed in this study, semiparametric models employing the Skew-Normal (SN) distribution consistently outperform their Normal (N) distribution counterparts, indicating that modeling asymmetric uncertainty improves predictive performance. Furthermore, the BwNC formulation demonstrates slight better mean performance 15

than the QwNC formulation. However, the considerable overlap in their standard deviations suggests that the performance difference between the two formulations is relatively modest. 6.2. Modeling Analysis Figure 3 illustrates stochastic FD modeling results from semiparametric models proposed in this study, while Figure 4 shows the corresponding results from baseline models. In these visualizations, the red curve represents the conditional expected flow or speed for each density value, while the blue shading areas demonstrate the 90% and 99% confidence intervals. The semiparametric models produce physically consistent FD relationships. Specifically, both flow and speed approach zero when traffic density approaches the estimated jam density. The estimated jam density is approximately 150 veh/km/lane, which is consistent with the typical observations in the traffic environment. Furthermore, the predicted uncertainty covers the majority of the observed data points and gradually shrink as density approaches the jam density in the congestion regime. Furthermore, models based on Skew-Normal distribution are able to capture the transition between free flow and congestion regimes, which occurs around a density of 35 veh/km/lane. In addition, the SN-BwNC and N-QwNC models generate smooth functional relationships without noticeable overfitting, suggesting that the semiparametric formulation provides a good balance between model flexibility and structural regularization. The results for the S3+LN model reveal a pronounced bias toward higher values in heavy congestion regimes, as illustrated in Figure 4c. Specifically, in both the speed-density and flow-density relationships, the empirical observations in the congested region consistently fall outside the 99% confidence interval of the model estimation. A comparison of model parameters between S3+LN and the deterministic S3 model (Table A.5) indicates that this performance degradation stems from the inherent inability of the deterministic model to characterize heavy congestion regimes. Treating parameters as random variables primarily serves to introduce uncertainty without substantially improving the underlying fitting quality of the deterministic framework. Consequently, S3+LN inherits the structural drawbacks of the S3 model, leading to persistent systematic bias in higher density regions. In contrast, Gaussian Process (GP) regression demonstrates a superior capacity for stochastic modeling, though it encounters challenges in generating consistent uncertainty across the entire density range. As shown in Figure 4b, the inherent bias of the S3 model is successfully mitigated in the S3+GP configuration. Similarly, Figure 4a displays the results for S3+GP, where the Gaussian Process effectively introduces necessary nonlinearity to the standard Greenshields speed–density model. However, the performance in the heavy congestion region remains suboptimal. The scarcity of training data in high-density states leads to a significant localized increase in model uncertainty. This variance is further magnified when transformed into the flow-density relation, where the multiplicative effect of density results in a wide, undesirable predictive interval that limits the model’s practical utility for congestion forecasting. 6.3. Performance on Different Density Regions The CRPS and MAE comparison of the stochastic FD models under different density regions is presented in Figure 5. The density is categorized into four regimes: free flow (0-20 veh/km), transition (20-60 veh/km), light congestion (60-100 veh/km), and heavy congestion (100-150 veh/km). Semiparametric models consistently outperform baseline counterparts across all density regimes. This improvement is most pronounced in the heavy congestion region, where baseline models typically struggle with high variance. Regarding the flow-density relationship, a divergence in performance trends is observed as density increases. While the CRPS for baseline models degrades significantly in high-density states, the semiparametric models exhibit enhanced predictive accuracy and stability under heavy congestion. For the speed-density relationship, all models follow a non-monotonic performance curve in the speed-density relationship, with the highest error rates (worst performance) appeared in the transition region, likely due to the inherent instability of traffic flow during phase shifts. The deterministic performance of semiparametric SFD models is comparable to that of the Gaussian Process (GP) models. This suggests that the potential driver for improved stochastic modeling is the integration of physical constraints, which effectively regularizes the model’s uncertainty in high-congestion regions.

16

Figure 5: CRPS and MAE Comparison on Different Density Regions

6.4. Ablation on Regularization Loss The influence of the regularization term ℓReg on training stability is evaluated in Figure 6, which depicts the negative log-likelihood (NLL) loss across five cross-validation folds. The blue and orange curves represent the training trajectories with and without the inclusion of ℓReg , respectively. The experimental results indicate that ℓReg is indispensable for numerical stability in both the SN-QwNC and SN-BwNC architectures. In the absence of this regularization term, both models exhibit highly volatile loss curves and a significant reduction in convergence speed. Notably, the SN-BwNC model fails to achieve convergence in two out of the five folds when ℓReg is removed. These ablation results confirm that the proposed regularization loss is a critical component for ensuring both the stability and the overall effectiveness of the stochastic training process. 7. Discussion The effectiveness of the proposed framework relies on its ability to propagate first and second-moment constraints to the distribution parameters, specifically the constrained parameters ηcon . Therefore, after determining constrained parameters, the primary theoretical challenge lies in the existence of a solution for the moment-matching system in Eq. (4). For a distribution f and values 0 < m, s < ∞, the problem simplifies to:    =m  EX∼ f [X] (38)    VarX∼ f [X] = s2 17

Figure 6: The convergence patterns during model training with v.s. without regularization loss in the objective function.

As established in Proposition 1, any distribution belonging to a location-scale family guarantees a unique solution to this system. However, the applicability of our framework extends beyond this family: while some distributions naturally yield unique solutions, others become feasible once appropriate parameter constraints are imposed. 7.1. Feasible Distribution from Non-Location-scale Family 7.1.1. Gamma Distribution A primary example of a feasible non-location-scale distribution is the Gamma distribution, with probability density function as: xα−1 e−x/θ f (x; α, θ) = α , for x > 0, (39) θ Γ(α) where Γ(·) is the gamma function, and {α, θ} are positive values. The first two moments are defined as: E[X] = αθ,

Var(X) = αθ2 .

(40)

By defining the constrained parameters as ηcon = [α, θ]T , we obtain a unique solution for the system of equations in Eq. (38): s2 m2 α= 2. (41) θ= , m s 7.1.2. Log-Normal Distribution The Log-Normal distribution is another widely used non-location-scale family that fits our framework. With its probability density function defined by ! 1 (ln x − µ)2 f (x; µ, σ) = , for x > 0, (42) √ exp − 2σ2 xσ 2π where µ ∈ R and σ > 0, its first two moments are: E[X] = eµ+σ /2 , 2

Var(X) = (eσ − 1)e2µ+σ . 2

18

2

(43)

Here, the constrained parameters are defined as ηcon = [µ, σ]T . To adapt this to our framework, we substitute these theoretical moments into Eq. (38). This yields a unique, closed-form solution for the parameters: !! 1 ! 1 s2 2 s2 (44) , µ = ln m − ln 1 + 2 . σ = ln 1 + 2 2 m m 7.2. Distributions Requiring Additional Constraints 7.2.1. Truncated Distributions While many distributions inherently support the moment-matching system, certain families, such as truncated distributions, require additional mathematical constraints to remain feasible. As demonstrated by Bhatia and Davis (2000), the variance of any bounded probability distribution is strictly limited by its expectation, as well as its absolute lower bound L and upper bound U. This relationship is expressed as: Var[X] ≤ (E[X] − L)(U − E[X]).

(45)

Consequently, the condition 0 < s < ∞ does not guarantee a solution for the system in Eq. (38), particularly when s is excessively large. Translating the mathematical limitation into our physical constraints, an upper bound is needed for the standard deviation:    Eq∼p(q|ρ) [q] = Sq∼p(q|ρ) [q] = 0, ρ ∈ {0, ρjam },     (46) 0 < Eq∼p(q|ρ) [q] < ∞, ρ ∈ (0, ρjam ),      0 < Sq∼p(q|ρ) [q] ≤ S max (ρ), ρ ∈ (0, ρjam ). S max (ρ) represents its theoretical upper bound derived from Eq. (45), given by: q   S max (ρ) = Eq∼p(q|ρ) [q] − L(ρ) U(ρ) − Eq∼p(q|ρ) [q] ,

(47)

where L(ρ) and U(ρ) denote the physical lower bound (e.g., minimum flow) and upper bound (i.e., maximum capacity) of flow q at density ρ, respectively. To guarantee the existence of a valid solution for the moment-matching equations, the functional design of the standard deviation s(ρ) must explicitly depend on the expectation function m(ρ), as well as the local boundary functions L(ρ) and U(ρ). 7.2.2. Gaussian Mixture Model Another complex scenario arises when employing mixture distributions, such as a two-component Gaussian Mixture Model (GMM). The probability density function is given by: p(x) = kN(x; µ1 , σ21 ) + (1 − k)N(x; µ2 , σ22 ),

(48)

where k ∈ (0, 1) is the mixing weight. The theoretical mean and variance can be analytically expressed as: E[X] = kµ1 + (1 − k)µ2 , Var(X) = kσ21 + (1 − k)σ22 + k(1 − k)(µ1 − µ2 )2 .

(49)

The parameters are partitioned into constrained parameters ηcon = [µ1 , µ2 ]T , and unconstrained parameters ηfree = [k, σ1 , σ2 ]T . Equating the moments to m and s2 provides the following symmetric closed-form solutions: r r 1−k k µ1 = m ± ∆, µ2 = m ∓ ∆, (50) k 1−k h i where ∆ = s2 − kσ21 + (1 − k)σ22 > 0. A real solution necessitates ∆ > 0, thereby introducing a lower bound for the standard deviation. Consequently, the physical constraints must be adjusted:    Eq∼p(q|ρ) [q] = Sq∼p(q|ρ) [q] = 0, ρ ∈ {0, ρjam },     (51) 0 < Eq∼p(q|ρ) [q] < ∞, ρ ∈ (0, ρjam ),      S min (ρ) < Sq∼p(q|ρ) [q] < ∞, ρ ∈ (0, ρjam ). q where S min (ρ) = k(ρ)σ21 (ρ) + (1 − k(ρ))σ22 (ρ). Ultimately, the design of the standard deviation function s(ρ) must carefully accommodate the selected free parameters ηfree to ensure the condition ∆ > 0 is continuously satisfied. 19

8. Conclusion In this study, we proposed a novel semiparametric framework for modeling stochastic fundamental diagrams. By partitioning distribution parameters into a constrained component and a data-driven neural component, our framework reconciles physical consistency with empirical flexibility. This semiparametric design ensures that essential traffic properties, such as boundary conditions at zero and jam density, are satisfied by construction while capturing the complex stochastic patterns inherent in traffic data. From a theoretical perspective, we established that the underlying moment-matching system admits a unique solution for location–scale families, guaranteeing the well-posedness of the parameterization. We further demonstrated the framework’s extensibility by deriving feasibility conditions for non-location–scale families, including truncated and mixture-based distributions in the final discussion. Empirical evaluation using the MAGIC dataset demonstrates that our semiparametric approach consistently outperforms baseline models. Specifically, the proposed models provide superior probabilistic accuracy and robust uncertainty quantification, particularly in congested regimes where the conventional stochastic extension of deterministic models often exhibit systematic bias. Our results further highlight that incorporating asymmetric distributions significantly enhances the model’s ability to capture observed empirical variance. In summary, this framework provides a flexible, theoretically grounded foundation for stochastic FD modeling. Future research will explore the integration of diverse distribution families, the extension of the model to multi-regime and multi-lane dynamics, and the inclusion of spatiotemporal dependencies to further improve the fidelity of stochastic traffic flow representations. Acknowledgment The authors thank the Swedish Transport Administration for supporting this work. Appendix A. Parameters for the Log-Normal-Based S3 Model Table A.5 shows marginal difference between the S3 model parameters obtained using the original method (Cheng et al., 2021) (i.e., the weighted least squares method) and those derived from the method by Cheng et al. (2024). Table A.5: Parameter Value Comparison between Deterministic Model and Log-Normal-Based Stochastic Model

S3 + WLSM Expected Parameter Value vfree vcritical m ρcritical

67.773±0.259 41.008±0.238 2.760±0.039 48.554±0.170

Parameter as Random Variable ln(vfree ) ∼ N(µ = 4.194±0.003 , σ = 0.104±0.004 ) ln(vcritical ) ∼ N(µ = 3.635±0.002 , σ = 0.265±0.003 ) 1 ∼ N(µ = 0.404±0.003 , σ = 0.205±0.001 ) m Constant

66.290±0.205 37.887±0.057 2.478±0.017 48.554±0.170

20

References Ahmed, A., Ngoduy, D., Adnan, M., Baig, M.A.U., 2021. On the fundamental diagram and driving behavior modeling of heterogeneous traffic flow using UAV-based data. Transportation Research Part A: Policy and Practice 148, 100–115. Bai, L., Wong, S.C., Xu, P., Chow, A.H.F., Lam, W.H.K., 2021. Calibration of stochastic link-based fundamental diagram with explicit consideration of speed heterogeneity. Transportation Research Part B: Methodological 150, 524–539. Bhatia, R., Davis, C., 2000. A better bound on the variance. The American Mathematical Monthly 107, 353–357. Bishop, C.M., 1994. Mixture density networks. Technical Report. Aston University. Bramich, D.M., Menendez, M., Ambühl, L., 2023. FitFun: A modelling framework for successfully capturing the functional form and noise of observed traffic flow–density–speed relationships. Transportation Research Part C: Emerging Technologies 151, 104068. Cheng, Q., Lin, Y., Zhou, X.S., Liu, Z., 2024. Analytical formulation for explaining the variations in traffic states: A fundamental diagram modeling perspective with stochastic parameters. European Journal of Operational Research 312, 182–197. Cheng, Q., Liu, Z., Lin, Y., Zhou, X.S., 2021. An S-shaped three-parameter (S3) traffic stream model with consistent car following relationship. Transportation Research Part B: Methodological 153, 246–271. Cheng, Z., Wang, X., Chen, X., Trépanier, M., Sun, L., 2022. Bayesian calibration of traffic flow fundamental diagrams using Gaussian processes. IEEE Open Journal of Intelligent Transportation Systems 3, 763–771. Del Castillo, J.M., Benitez, F.G., 1995. On the functional form of the speed-density relationship—II: Empirical investigation. Transportation Research Part B: Methodological 29, 391–406. Frostig, R., Johnson, M.J., Leary, C., 2019. Compiling machine learning programs via high-level tracing, in: SysML Conference 2018. Gneiting, T., Raftery, A.E., 2007. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association 102, 359–378. doi:10.1198/016214506000001437. Greenberg, H., 1959. An analysis of traffic flow. Operations Research 7, 79–85. Greenshields, B.D., Bibbins, J.R., Channing, W.S., Miller, H.H., 1935. A study of traffic capacity, in: Highway Research Board Proceedings, Washington, D.C.. pp. 448–477. Hendrycks, D., Gimpel, K., 2016. Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415 . Hornik, K., Stinchcombe, M., White, H., 1989. Multilayer feedforward networks are universal approximators. Neural Networks 2, 359–366. Jabari, S.E., Zheng, J., Liu, H.X., 2014. A probabilistic stationary speed–density relation based on Newell’s simplified car-following model. Transportation Research Part B: Methodological 68, 205–223. Jayakrishnan, R., Tsai, W.K., Chen, A., 1995. A dynamic traffic assignment model with traffic-flow relationships. Transportation Research Part C: Emerging Technologies 3, 51–72. Kingma, D.P., Ba, J., 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 . Lei, Y.Z., Gong, Y., Yang, X.T., 2024. Unraveling stochastic fundamental diagrams with empirical knowledge: Modeling, limitations, and future directions. Transportation Research Part C: Emerging Technologies 169, 104851. 21

Liu, Z., Lyu, C., Wang, Z., Wang, S., Liu, P., Meng, Q., 2023. A Gaussian-process-based data-driven traffic flow model and its application in road capacity analysis. IEEE Transactions on Intelligent Transportation Systems 24, 1544–1563. Ma, W., Zhong, H., Wang, L., Jiang, L., Abdel-Aty, M., 2022. MAGIC dataset: Multiple conditions unmanned aerial vehicle group-based high-fidelity comprehensive vehicle trajectory dataset. Transportation Research Record 2676, 793–805. Newell, G.F., 1961. Nonlinear effects in the dynamics of car following. Operations Research 9, 209–229. Ngoduy, D., 2011. Multiclass first-order traffic model using stochastic fundamental diagrams. Transportmetrica 7, 111–125. Ni, D., 2015. Traffic flow theory: Characteristics, experimental methods, and numerical techniques. ButterworthHeinemann. Ni, D., Hsieh, H.K., Jiang, T., 2018. Modeling phase diagrams as stochastic processes with application in vehicular traffic flow. Applied Mathematical Modelling 53, 106–117. Ni, D., Leonard, J.D., Jia, C., Wang, J., 2016. Vehicle longitudinal control and traffic stream modeling. Transportation Science 50, 1016–1031. Papageorgiou, M., Blosseville, J.M., Hadj-Salem, H., 1989. Macroscopic modelling of traffic flow on the Boulevard Périphérique in Paris. Transportation Research Part B: Methodological 23, 29–47. Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al., 2019. PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems 32. Qu, X., Wang, S., Zhang, J., 2015. On the fundamental diagram for freeway traffic: A novel calibration approach for single-regime models. Transportation Research Part B: Methodological 73, 91–102. Qu, X., Zhang, J., Wang, S., 2017. On the stochastic fundamental diagram for freeway traffic: Model development, analytical properties, validation, and extensive applications. Transportation Research Part B: Methodological 104, 256–271. Rigby, R.A., Stasinopoulos, D.M., 2005. Generalized additive models for location, scale and shape. Journal of the Royal Statistical Society Series C: Applied Statistics 54, 507–554. Shi, R., Mo, Z., Huang, K., Di, X., Du, Q., 2021. A physics-informed deep learning paradigm for traffic state and fundamental diagram estimation. IEEE Transactions on Intelligent Transportation Systems 23, 11688–11698. Storm, P.J., Mandjes, M., van Arem, B., 2022. Efficient evaluation of stochastic traffic flow models using Gaussian process approximation. Transportation Research Part B: Methodological 164, 126–144. Theis, L., Bethge, M., 2015. Generative image modeling using spatial LSTMs. Advances in Neural Information Processing Systems 28. Wang, H., Li, J., Chen, Q.Y., Ni, D., 2011. Logistic modeling of the equilibrium speed–density relationship. Transportation Research Part A: Policy and Practice 45, 554–566. Wang, S., Chen, X., Qu, X., 2021. Model on empirically calibrating stochastic traffic flow fundamental diagram. Communications in Transportation Research 1, 100015. Zhang, X., Sun, J., Sun, J., 2025. On the stochastic fundamental diagram: A general micro-macroscopic traffic flow modeling framework. Communications in Transportation Research 5, 100163. Zhou, J., Zhu, F., 2020. Modeling the fundamental diagram of mixed human-driven and connected automated vehicles. Transportation Research Part C: Emerging Technologies 115, 102614. 22

Record · ID 381753 · SHA-256 e72ccec3840ed91d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.