Conceptio › Archive › arXiv CS
arXiv CSopen access

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

C AUSAL A RENA : B ENCHMARKING C AUSAL D ISCOVERY IN THE F OUNDATION M ODEL E RA

arXiv:2609.11897v1 [cs.LG] 10 Sep 2026

A P REPRINT Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye* School of Artificial Intelligence, Nanjing University E-mail: {lizr, liusy, wangtz, yehj}@lamda.nju.edu.cn

A BSTRACT Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining–evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.

1

Introduction

Causal discovery aims to recover directed causal relationships from observational or interventional data and is fundamental to scientific reasoning and intervention-based decision making [1, 2]. Because the true causal graph is rarely known in real systems, empirical evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data. SCMs make it possible to generate observational and interventional samples while retaining known ground-truth structure, and have therefore become a standard basis for comparing classical, neural, and pretrained causal discovery methods. Yet causal discovery remains difficult to evaluate consistently. Different studies use different graph families, structural mechanisms, noise distributions, dimensions, sample sizes, intervention protocols, thresholding rules, and aggregation conventions. Performance can also depend strongly on the synthetic generator itself, as illustrated by generator artifacts such as varsortability [3]. Consequently, a strong result on one benchmark does not necessarily indicate robust causal discovery across other causal environments. The emergence of causal discovery foundation models (CDFMs) makes this problem more fundamental. Methods such as AVICI [4], SEA [5], Arrow [6], CauScale [7], CDFM [8], FoundCause [9], and TabCausal [10] are pretrained over large collections of synthetic or mixed causal environments. Their benchmark performance can therefore reflect not only causal structure-learning ability, but also overlap between pretraining and evaluation—from similar graph families and mechanisms to the same or closely related SCM generators. Moreover, once a fixed SCM benchmark becomes public, it can itself become future pretraining data. Thus, in the foundation-model era, the test distribution is no longer merely an evaluation environment; it may also become part of the training distribution. ∗

Corresponding author.

Per-SCM pairwise win rate ←

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Synthetic

Semantic

Formula

Real

best

A P REPRINT

second-best

1.0

☆

0.8

★

★★ ☆

★

☆

0.6

☆

50% pairwise wins

0.4 0.2 0.0

s

s re

do

n Ra

m

g Re

S DA

A GM DA

CD

SD

TE

NO

P

L -M

S AR

e aL

Sc

Ca

M

GA

LiN

RS

EA

T NO

SP

IG

w ro Ar M) 7 (3

A SE )

6M 0. (~

PC

FM CD M) .7 (9

I IC AV M) .2 (4

IS

CD

ES

GI

al se us au Ca 4M) dC 9M) b . n 13 Ta (4 ou ( F

Figure 1: Observation-only pairwise win profile across benchmark families. For each SCM/dataset, method scores are first averaged over repeated runs. On each unit, every method pair is compared separately on F1 and SHD; higher F1 and lower SHD each count as a win, ties count as 0.5, and the reported win rate pools wins from both metrics. Each bar reports this pooled pairwise win fraction within one benchmark family. Methods are ordered by their mean family-wise win rate from left to right; stars mark the best and second-best method within each family. Values such as 139M under pretrained method names denote parameter counts in millions for the evaluated checkpoints. Existing resources address different parts of this problem. Synthetic frameworks such as Benchpress [11] and CausalProfiler [12] support controlled and scalable evaluation, while grounded resources such as Sachs [13], CausalDynamics [14], CausalBench [15], and Causal Chambers [16] provide semantic, scientific, or physical context. However, these resources typically emphasize different graph distributions, domains, tasks, or protocols, making scores difficult to compare directly across methods and especially across pretrained models. These limitations suggest four requirements for evaluating causal discovery in the foundation-model era. Breadth is needed because performance should not be determined by a single synthetic prior. Freshness is increasingly important because no fixed public SCM suite can remain unseen indefinitely; an evaluation framework should therefore support newly constructed causal environments as pretrained models evolve. Grounding complements arbitrary synthetic generators with causal systems whose variables, edges, and interventions have meaningful operational or scientific interpretations. Finally, diagnosability requires going beyond a single leaderboard score to identify which structures, mechanisms, and environments cause methods to succeed or fail. Motivated by these requirements, we introduce CausalArena2 , a unified and evolvable benchmark for causal discovery. Its three SCM families play complementary roles. Synthetic SCMs provide controlled breadth over graph structures, mechanisms, noise, dimensions, and intervention settings. Semantic operational SCMs ground causal structures in auditable real-world processes and provide a natural basis for introducing newly authored environments in future benchmark versions. Formula-grounded scientific SCMs anchor causal mechanisms in explicit scientific equations, making mechanism-level successes and failures easier to diagnose. All three families compile to a common executable SCM specification and share the same observational and interventional evaluation protocol; public real-world tables with published graphs serve as an additional external-validity check. The current arena contains 1,200 executable SCM specifications—1,000 synthetic, 100 semantic operational, and 100 formula-grounded scientific SCMs—and evaluates representative classical, neural, and pretrained causal discovery methods under common data and scoring interfaces. The initial public release exposes a balanced half of each generated SCM family, while the remaining audited SCMs are reserved for held-out leaderboard evaluation and will be released progressively through versioned updates. Beyond pooled scores, we analyze performance by SCM family, graph and mechanism properties, sample size, intervention protocol, and semantic or scientific setting. The shared interface and protocol are fixed and standardized so that later rounds can introduce new SCMs, domains, and evaluation settings without rewriting how methods are run or scored. Our experiments yield several consistent conclusions. Rankings of classical, neural, and pretrained methods shift substantially across Synthetic, Semantic, Formula, and real-data settings: no method dominates every slice, and strong performance on synthetic SCMs does not reliably transfer to semantic, scientific, or real tables. Extra samples and interventions help methods unevenly, so a single pooled score is an incomplete summary. For pretrained models especially, these patterns caution against interpreting leaderboard numbers without considering possible pretraining– evaluation overlap. Beyond ranking methods, CausalArena aims to make such failures diagnosable and, through a shared executable interface, to support successive evaluation rounds that can guide and re-test progress as new algorithms and foundation models appear. 2

Hugging Face dataset: https://huggingface.co/datasets/LAMDA-Tabular/CausalArena; leaderboard: https:// huggingface.co/spaces/LAMDA-Tabular/CausalArena-leaderboard.

2

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Figure 1 summarizes this ranking instability under a shared observation-only protocol. Methods are ordered by mean pairwise win rate across families, yet the family-wise bars show that the leader and second place often change from Synthetic to Semantic, Formula, and real data; a method that wins broadly on one family can sit mid-pack on another. The figure therefore motivates evaluating causal discovery under multiple SCM regimes, with rankings read across families as well as in pooled form. Our main contributions are: • We formulate causal discovery evaluation in the foundation-model era around four requirements—breadth, freshness, grounding, and diagnosability—highlighting the additional challenge introduced by pretraining–evaluation overlap. • We introduce CausalArena, which combines broad synthetic SCMs, semantically grounded operational SCMs, and formula-grounded scientific SCMs under a shared observational and interventional protocol whose interface can evolve by admitting new environments without changing how methods are compared. • We conduct a broad evaluation of classical, neural, and pretrained causal discovery methods and provide family-, mechanism-, sample-size-, intervention-, and real-data analyses, revealing substantial ranking shifts and distinct failure profiles across evaluation environments.

2

Related Work

2.1

Causal Discovery Methods

Causal discovery aims to recover directed causal structure from observational or interventional data, typically represented by a directed acyclic graph (DAG) associated with an underlying SCM [1, 2]. Classical approaches differ substantially in their assumptions and learning principles. Constraint- and search-based methods infer structure from conditionalindependence relations or graph scores, including PC [1, 17, 18], GIES [19], IGSP [20], and CDIS [21]. Functional approaches exploit assumptions on causal mechanisms or noise, such as LiNGAM [22] and DAS [23], which builds on additive-model structure learning [24, 25]. Continuous-optimization methods instead formulate structure learning as differentiable optimization, including NOTEARS [26], NOTEARS-MLP [27], DAGMA [28], and SDCD [29]. RandomRegress serves as a simple random-order regression control for sanity-checking scores against trivial baselines. Since these approaches rely on different structural, functional, and distributional assumptions, their relative performance can change substantially across data-generating conditions. 2.2

Causal Discovery Foundation Models

More recently, amortized and pretrained structure learners have shifted causal discovery toward foundation-model-style inference. AVICI [4] amortizes structure prediction over large collections of synthetic SCMs; SEA [5] aggregates classical discovery estimates from sampled variable subsets; Arrow [6] factorizes graphs into skeletons and topological orders with an acyclicity guarantee; CauScale [7] targets neural discovery at large graph scales; CDFM [8] treats unknown mechanisms as latents in a variational foundation-model recipe; FoundCause [9] injects pairwise causal statistics and models latent confounding; and TabCausal [10] pretrains across diverse tabular causal environments, including interventional settings. Unlike classical algorithms that are typically fitted or executed independently on each dataset, these models can encode priors learned from large synthetic or mixed causal environments. This changes the interpretation of benchmark performance: a test score may reflect not only causal structure-learning ability, but also similarity between the model’s pretraining environments and the evaluation graphs, mechanisms, or generators. Comparisons among pretrained methods are therefore particularly sensitive to differences in their training distributions and evaluation protocols. 2.3

Evaluation of Causal Discovery

Synthetic SCMs are the dominant tool for scalable causal discovery evaluation because they generate samples together with known ground-truth graphs. Standard evaluations sample a DAG, assign structural mechanisms and noise distributions, and generate observational or interventional data [1, 2]. Many method papers consequently introduce their own synthetic evaluation suites, including those accompanying NOTEARS [26], DAGMA [28], SDCD [29], DCDI [30], and NODAGS-Flow [31]. More general frameworks such as Benchpress [11], CausalProfiler [12], and DECI’s CSuite [32] systematize generation and comparison across broader configurations. However, synthetic evaluation is itself sensitive to generator design: varsortability and related artifacts [3], for example, demonstrate that methods may exploit properties of a particular data generator rather than recover causal structure in a more general sense. Complementary generator designs such as unitless unrestricted Markov-consistent (UUMC) SCM sampling [33] aim to reduce such nonphysical sortability patterns when constructing synthetic benchmarks. A complementary line of work evaluates causal discovery on systems with semantic, scientific, or physical grounding. 3

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Table 1: Comparison of representative causal discovery evaluation resources and CausalArena. Symbols summarize typical properties of the cited resources: ✓ indicates general support, ◦ indicates setting-dependent support, and – indicates that the property is typically absent. Resource / benchmark type Real / expert graphs [13, 34, 36] Method-specific synthetic [26, 28–31] Synthetic frameworks [11, 12, 32] Domain simulators [14, 15, 37, 39] Time-series CD benchmarks [37, 38, 40] Physical real-world testbeds [16] Semantic / LLM graph resources [41, 42] Foundation-model evaluation [4, 5, 8–10] CausalArena

Regenerable SCM Synthetic breadth Grounded semantics Scientific mechanisms Shared protocol Obs.+Int. Tabular DAG – ✓ ✓ ✓ ◦ – – ✓ ✓

✓ – – ✓ ✓ ✓ ✓ – ✓

– – ✓ – – – – ✓ ✓

◦ – – ✓ ✓ ✓ – – ✓

– – ✓ – – – – – ✓

◦ ◦ ◦ ◦ ◦ ✓ – ◦ ✓

✓ ✓ ✓ ◦ – ✓ – ✓ ✓

Classic resources include the Sachs protein-signaling data [13] and reference Bayesian networks in bnlearn [34], while CauseEffectPairs [35] remains a standard benchmark for bivariate cause–effect orientation, a narrower task than full multivariate DAG recovery. OCDB [36] develops real-data evaluation with graph metrics designed for more comparable structure assessment. Domain-oriented resources provide richer generative mechanisms: CausalTime [37] introduces realistically generated temporal data, CausalRivers [38] scales real-world hydrology time-series evaluation with constructed ground-truth graphs, CausalDynamics [14] targets dynamical systems derived from scientific equations, CausalBench [15] focuses on perturbational biological networks, causalAssembly [39] models industrial production processes, and Causal Chambers [16] provides physical systems with controlled interventions. TimeGraph [40] further stresses nonstationarity, irregular sampling, missingness, and latent confounding in synthetic temporal settings. These settings improve realism and mechanistic grounding, but are typically specialized to particular domains, system classes, or data modalities. Semantic information has also become increasingly relevant as language models are introduced into causal reasoning and graph construction. CausalGraphBench [41] evaluates language-model-based graph discovery against curated causal structures, while PromptBN/ReActBN [42] studies language-assisted Bayesian-network construction in low- or no-data regimes; related work surveys the use of LLMs for causal extraction, prior injection, and graph refinement [43]. Such resources evaluate a different source of causal information from purely numerical SCM benchmarks, since predictions can depend on semantic priors as well as statistical evidence. Table 1 summarizes this landscape: existing resources typically cover only a subset of regenerable SCMs, synthetic breadth, semantic or scientific grounding, a shared protocol, and joint observational–interventional tabular DAG evaluation. Synthetic frameworks emphasize breadth and protocol, but not grounding; domain simulators and physical testbeds supply meaning, but rarely a shared cross-domain protocol; foundation-model evaluations often stay within regenerable synthetic settings without operational or scientific grounding. As a result, strengths remain distributed across separate resources, and scores are hard to compare—especially for pretrained models whose training distributions may overlap with the benchmark environments. This motivates a common evaluation framework that combines these axes while retaining enough structure for detailed analysis.

3

Benchmark Formulation and Design

3.1

Task and Evaluation Formulation

We consider causal discovery over d observed variables X = (X1 , . . . , Xd ) generated by an SCM S. An SCM specifies a DAG G = (V, E) together with structural assignments  Xj := fj XPaG (j) , ϵj , j = 1, . . . , d, (1) where PaG (j) denotes the parents of Xj , fj is its causal mechanism, and ϵj is an exogenous noise variable [1, 2]. Let AG ∈ {0, 1}d×d denote the directed adjacency matrix of G. The task is to recover AG from samples generated by S. Here S supplies the controlled evaluation environment together with known ground-truth structure. Each executable SCM induces an observational distribution PS (X) and, after a do-style perturbation on intervention targets I, an interventional distribution PS (X | do(I)). We therefore construct two evaluation settings from the same SCM: an observation-only dataset Dobs and an observation-plus-intervention dataset Dobs+int . A method M receives one of these datasets and returns bG = M(D), A (2) which is scored against AG with graph-recovery metrics such as F1 and SHD.

4

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Classical and foundation-model pipelines differ in how M is obtained. A classical (or per-dataset neural) method is applied independently to each evaluation instance, bi = M(Di ), A (3) so the learner is fitted or executed only on Di . A CDFM instead first learns reusable parameters from a collection of pretraining environments Epre , bi = Mθ (Di ), θ = Pretrain(Epre ), A (4) N as in amortized and pretrained causal discovery [4, 8, 10]. For a benchmark B = {(Di , Ai )}i=1 , a conventional aggregate score is h  i bi , Ai , Score(M; B) = Aggi m A (5) where m is a graph-recovery metric and Agg aggregates over instances and replicates. 3.2

Design Requirements

Eqs. (3)–(5) make the Introduction’s difficulties precise, and they constrain what a useful Score(M; B) can mean. First, the score is defined only relative to a chosen B: different papers use different graph families, mechanisms, noise processes, sample sizes, and intervention protocols, so a high score on one B need not transfer to another. Second, for a CDFM the reported quantity is Score(Mθ ; B) with θ = Pretrain(Epre ), and similarity between Epre and B may occur from an exact SCM or graph to a shared family or generator, so a high score may mix causal discovery ability with pretraining–evaluation affinity. Third, once a fixed public B is released, its SCMs can later enter Epre , and the test distribution can overlap with later training. If B is narrow, the score tracks one generator prior; if it is fixed while Epre grows, the score can drift toward overlap; if it contains only anonymous tables, it says little about named processes or equations; if only a pooled value is reported, it hides where methods succeed or fail. CausalArena is therefore organized as a single evaluation arena whose benchmark B is built to answer these needs together. Concretely, three SCM families—synthetic, semantic operational, and formula-grounded scientific—are compiled into one executable specification and exported under the same observational and interventional interfaces. Classical, neural, and pretrained methods return directed adjacency estimates under those interfaces; scores are reported as a pooled Score(M; B) and broken down by family, mechanism, protocol, and domain factors. Public real-world tables with published graphs serve as an additional external-validity check alongside the three families. Figure 2 summarizes this organization. Let ϕ(S) denote construction factors of an SCM (graph family, mechanism and noise tags, dimension, domain, intervention protocol, and related metadata), and write slices of the arena as {Bc } indexed by factors in ϕ. The constraints above motivate the following four design criteria, stated in terms of ϕ and {Bc }. Breadth. Because Score(M; B) depends on the chosen B, the arena must support a diverse collection of slices {Bc }—graph families, mechanisms, noise, dimensions, sample sizes, interventions, and difficulty—so that performance is assessed across many causal environments. Freshness. Because a CDFM score is Score(Mθ ; B) with θ = Pretrain(Epre ), a fixed public B cannot remain outside Epre indefinitely. Freshness requires that the evaluation pool can grow: if Epre expands, the arena should admit Snew under a shared executable interface and common scoring rules. CausalArena therefore keeps the interface and protocol fixed and standardized, so that newly authored SCMs can be added as the arena evolves. bG remains meaningful outside synthetic Grounding. Because anonymous variation of ϕ alone leaves open whether A generators, some environments must attach operational or scientific meaning to variables, edges, and interventions. A bG and an interpretable claim about a named process or equation. recovered edge is then a binary entry in A Diagnosability. Because a pooled Score(M; B) leaves success and failure unexplained, evaluation should retain ϕ(S) and report slice scores h i bi , Ai ) , Scorec (M) = Aggi: c(i)=c m(A (6) so that rankings can be read as failure profiles across environments, together with a global order. These four requirements pull in different directions at once: breadth favors large regenerable grids; grounding favors named processes and equations; freshness favors newly authored environments over time; diagnosability favors explicit factors and slice reports. A single SCM family is unlikely to carry all four equally, which leads to the three complementary families next. 5

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Synthetic SCMs Topology Sample size & dimension Mechanism families

Shared interface

Interventions & difficulty Deterministic quality gates

Executable SCMs

Semantic + Formula SCMs (agentic)

Generate planning

Causal Discovery Methods

Unperturbed SCM

Graphical search

Analysis Family-wise comparison

Functional assumptions

Root / noise distributions

Seed & reference grounding

Observational split

A P REPRINT

Graph drafting & review

Interventional split

Continuous optimization

perturbations

Pretrained

Real SCMs

Intervention analysis

Evaluation Predicted Adjacency matrix

Executable translation

Validation & release Public real tables with known graphs

Human + LLM audit

Sample-size analysis

Scored against ground truth

Semantic analysis

Runtime analysis

...

Figure 2: Overview of the CausalArena framework. Synthetic, semantic operational, and formula-grounded scientific SCMs are compiled into executable specifications under a shared observational and interventional interface. Classical, neural, and pretrained causal discovery methods recover directed adjacency matrices, which are evaluated and analyzed across SCM families, mechanisms, sample sizes, intervention protocols, and domains. Public real-world tables with published graphs provide an additional external-validity check. 3.3

Three Complementary SCM Families

CausalArena assigns the requirements to three SCM families that do different jobs, then compiles them to one executable representation so that sampling, interventions, wrappers, and scoring stay comparable. The point of complementarity is coverage of roles that conflict if forced into one generator: scale and factor control on one side, operational and scientific meaning on the other, and an authoring path for new environments as models evolve. Synthetic SCMs (breadth, and factor-level diagnosis). Breadth needs a regenerable factor grid: topology, mechanism and noise tags, dimension, interventions, and difficulty can be varied systematically and at large d. Anonymous synthetic SCMs are the practical carrier of that grid. Because ϕ(S) is explicit, the same family also supports diagnosability through slice scores in Eq. (6). What they leave open is meaning: an edge is a matrix entry without an operational or scientific story. Semantic operational SCMs (operational grounding and freshness). Grounding needs variables and edges that mean something in a process—a policy setting, a queue length, a sensor reading, a counted outcome. Semantic SCMs supply that operational meaning and keep construction human-auditable. Because new scenarios can be authored through the same interface, they also provide a practical route to freshness as pretrained models evolve: new Snew can be added without changing the protocol. They do not replace the synthetic grid: operational scenarios are harder to enumerate at the same factor resolution and scale. Formula-grounded scientific SCMs (mechanistic grounding and diagnosis). Grounding also needs scientific or engineering equations with units and validity ranges. Formula SCMs keep named identities exact given their parents and place noise on instruments and readouts, so an error is an optical, chemical, or mechanical claim about a named mechanism. That structure strengthens mechanism-level diagnosability on equation-backed systems. Like semantic SCMs, they complement the synthetic family by supplying meaning that anonymous generators omit, at a cost in how densely ϕ can be swept. Taken together, the synthetic family mainly carries breadth and factor-level diagnosability; the semantic family carries operational grounding and a path to freshness; the formula family strengthens mechanistic grounding and equation-level diagnosability. The shared protocol is what makes these roles jointly readable as one arena. Public real tables remain an external wrapper check. Construction and validation details are in Section 4. The shared evaluation protocol and compared methods are specified there as well (Section 4.5). 6

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

3.4

A P REPRINT

Evolvable Evaluation for Foundation Models

The dependence of Eq. (4) on Epre makes benchmark release fundamentally different for CDFMs. Publishing executable SCMs is valuable for reproducibility, inspection, and reuse, while those same SCMs can later enter pretraining corpora. A permanently hidden test would make construction and evaluation hard to inspect and reproduce, and public in-distribution / out-of-distribution splits alone leave the tension unresolved once SCMs have been released. CausalArena therefore treats pretraining overlap as part of how a result is interpreted. Training disclosure and labeled splits clarify what a reported score uses; the shared executable interface and common protocol allow newly constructed graphs, mechanisms, and scenarios to be added as the arena evolves. In this sense, freshness is enabled by evolving the evaluation pool under a fixed interface. Release policy, contamination risks, and their limitations are discussed further in Section 7.

4

Benchmark Implementation

This section turns the design in Section 3 into an executable arena: how the three SCM families are built and validated, what the current pool covers, which method families are compared, and how they share one evaluation protocol. Concretely, each SCM specification provides a DAG G, structural mechanisms {fj }, exogenous or root distributions, intervention definitions, and the metadata required for generation and analysis. Once compiled, all families use the shared sampling and scoring interfaces in Section 4.5, with wrapper-level details in Section B. Their construction procedures reflect their different roles in CausalArena. Synthetic SCMs are generated programmatically to provide controlled breadth over causal structures and data-generating mechanisms. Semantic operational SCMs are constructed around meaningful real-world processes to provide grounding and an extensible source of newly authored environments. Formula-grounded SCMs anchor parts of the causal system in explicit scientific equations, providing mechanistic grounding and enabling more interpretable diagnosis. All SCMs pass family-specific validation and diversity checks before entering the evaluation pool. The current arena contains 1,200 executable SCM specifications: 1,000 synthetic configurations, 100 semantic operational scenarios, and 100 formula-grounded scientific scenarios. The public Hugging Face package currently releases half of each family. The following sections describe how each family is constructed and validated, then summarize the shared protocol and compared methods. 4.1

Synthetic SCMs: Controlled Breadth

The synthetic family is designed to systematically cover variations in the components of an SCM. We programmatically vary graph topology and dimension, structural mechanisms, root and noise distributions, dependencies among root variables, intervention settings, and difficulty factors. These factors yield different combinations of G, {fj }, and noise distributions, while retaining exact ground-truth structure for evaluation. The current pool contains 1,000 configurations over dimensions {10, 20, 30, 50, 100}, spanning 10 graph families, 16 mechanism families, 14 root-distribution families, 15 noise families, 4 root-dependency families, and 9 difficulty settings. The design includes high-dimensional configurations for stress testing and deliberately varies factors that can otherwise become implicit generator priors. Detailed family definitions and sampling rules are given in Section A.1. Before release, deterministic quality gates check DAG validity, mechanism executability, numerical ranges, degenerate variables, and the intended coverage factors. These checks ensure that controlled variation in the generator corresponds to valid evaluation instances rather than numerical or implementation artifacts. Programmatic audits are detailed in Section A.2. 4.2

Semantic Operational SCMs: Grounded and Extensible Environments

The semantic family moves beyond unnamed synthetic variables by constructing SCMs around operational processes with interpretable variables, causal links, measurements, and interventions. Nodes are labeled by what they mean in that process, and directed edges carry reasons for why changing a parent can affect a child. Interventions change settable fields in the same executable SCM, and descendants are recomputed accordingly. Each specification also stores variable definitions, measurement notes, intervention meanings, and edge reasons, so the causal structure is inspectable in ways that unnamed synthetic graphs are not, while retaining the same tabular input and directed adjacency target used by the other families. The current semantic pool contains 100 scenarios from 10 operational domains: cybersecurity and IT operations, education, finance and credit, government services, healthcare delivery, housing and real estate, manufacturing, public 7

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

health, urban transportation, and water and sanitation. Each domain contains 10 scenarios, with 17–25 variables and 25–67 directed edges per scenario. Because semantic SCMs are authored through a reusable construction pipeline, new operational scenarios can later be added without changing the evaluation interface. This property provides a practical route toward the evolvability and future freshness discussed in Section 3.2; the current release itself does not assume that semantic knowledge is unseen by all pretrained models.

Figure 3: Example Semantic and Formula SCMs (full released cards). Node colors follow the legend, and arrows indicate directed causal edges. Top: wildfire smoke, shelter behavior, and respiratory symptoms (public health; 20 variables, 26 edges). Bottom: barometric pressure reduction to sea level (earth systems; 16 variables, 18 edges). Example. Figure 3 (top) is the public-health scenario Wildfire smoke and shelter behavior: 20 variables and 26 edges. Think of one smoke day at one home, and read the panel left to right. Local outdoor PM2.5 is the true air near the house. The official PM2.5 monitor sees the same plume, but not perfectly, so its reading can differ from the local value. If that monitor reading is high enough, a shelter-in-place alert may be issued (this alert can be changed). The alert and the previous indoor PM2.5 reading together shape perceived smoke risk. Higher perceived risk then shortens hours outdoors and increases outdoor exposure reduction (for example through the N95 mask program). A second path is about the building. Housing leakage sets effective infiltration. Providing a portable filter raises device CADR (clean-air delivery rate). Room volume and adherence friction also matter. Local outdoor PM2.5, infiltration, filtration, and behavior then determine true indoor PM2.5. From that true indoor level, a consumer sensor gives a noisy reading, and an indoor exceedance flag is set. Daily inhaled PM2.5 dose combines hours outdoors, outdoor exposure reduction, true indoor PM2.5, and some direct outdoor contribution. Respiratory vulnerability then turns that dose into respiratory symptoms. Edges therefore have a plain reading: removing “portable filter → true indoor PM2.5” drops a real control people can change, and treating the consumer-sensor reading as the true indoor level mixes a measurement with the physical state. The executable SCM keeps the variable definitions, interventions, measurements, and edge justifications. 4.3

Formula-Grounded Scientific SCMs: Mechanistic Grounding

The formula-grounded family builds SCMs in which some mechanisms are named scientific or engineering equations, with units and valid ranges written down. Each scenario starts from one or more such equations together with the

8

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

surrounding measurement and correction steps. Those equations define selected structural mechanisms fj . Formula nodes are computed exactly from their parents. Randomness sits on inputs, instrument readings, or other non-formula quantities, not on arbitrarily jittering the equation itself. After sampling, residual checks verify that rows that were not intervened still match the scientific relation. The current pool contains 100 scenarios spanning 10 scientific and engineering domains: biology and ecology, astronomy, chemistry, mechanics and fluids, earth systems, electromagnetism, energy systems, materials and structures, optics and waves, and thermodynamics. Each domain contributes 10 scenarios, with 16–25 variables and 18–50 directed edges per scenario. Explicit equations do more than make the setting look scientific: they write down known mechanism structure that later analysis can check against. The formula family therefore supports both scientific grounding and the mechanism-level diagnosability emphasized in Section 3.2. Family-specific validation checks formula orientation, units, numerical validity ranges, and equation residuals. Example. Figure 3 (bottom) is the earth-systems scenario Barometric pressure profile: 16 variables and 18 edges. The panel has three short chains that meet in one sea-level pressure formula. First, pressure. A raw barometer pressure reading is corrected by a calibration drift bias. That drift depends on the observation window (day vs. night, changeable), because temperature cycling shifts the instrument offset. Raw pressure and drift together yield calibrated station pressure. Second, height. Surveyed elevation and GPS altitude are combined with an altitude fusion weight (also changeable) into fused station height, which then sets effective gravity. Third, temperature. The same observation window sets solar exposure. Ambient air temperature, biased by solar exposure, becomes the shielded thermometer reading. Relative humidity and that shielded temperature then determine layer-mean virtual temperature. Finally, calibrated station pressure pstn , fused height z, effective gravity geff , and layer-mean virtual temperature Tv enter the hypsometric sea-level reduction   geff z SLP = pstn exp , Rd Tv with dry-air gas constant Rd . Altimeter setting QNH is that sea-level pressure converted to hPa and rounded. The card therefore mixes exact named equations with ordinary instrument and calibration steps: changing the fusion weight or observation window recomputes later nodes in the same executable SCM, and residual checks keep non-intervened formula rows consistent with the reduction equation. 4.4

Construction Workflow, Validation, and Diversity Control

Semantic and formula-grounded SCMs share a staged agentic construction workflow consisting of seed and reference grounding, scenario planning, graph drafting, domain review, executable translation, adversarial review, deterministic validation, and release gating. Human and LLM-assisted review focuses on variable necessity, causal-edge rationale, intervention meaning, measurement realism, and potential template or alias artifacts. The formula-grounded family additionally verifies equation orientation, units, valid ranges, and residual consistency. The synthetic family does not require the same semantic authoring process; instead, deterministic programmatic checks validate graph structure, mechanisms, distributions, numerical behavior, and coverage constraints. Across all three families, diversity control removes redundant graphs, near-duplicate scenarios, repeated templates, unsupported edges, and low-information variables before inclusion in the evaluation pool. These family-specific procedures ultimately produce the same executable SCM interface. Consequently, differences across the three families arise from their causal environments, while sampling, interventions, method wrappers, adjacency targets, and scoring remain shared. Construction ablations in Section 5 examine the effects of reference grounding, planning, and graph review, while the complete staged workflow and release gates are provided in Section A.3. 4.5

Shared Protocol and Compared Methods

All families are scored under one shared protocol. For every executable SCM Si , CausalArena stores the groundtruth adjacency Ai and exports two matched tables from the same graph: an observation-only split Dobs from the unperturbed SCM, and an observation-plus-intervention split Dobs+int obtained by do-style perturbations with descendant recomputation. Intervention targets are recorded in masks shipped with the data. On the main leaderboard, the observation-only split uses nobs = 1000 rows; the mixed split uses nobs = 800 observational and nint = 200 interventional rows. Methods that do not consume intervention indicators are run only on supported splits, and unsupported cells are recorded as missing.

9

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Every method wrapper receives standardized tabular inputs and returns a directed adjacency estimate. Each SCM is evaluated with multiple independent replicates (five in the current main protocol); replicate scores are averaged before aggregation over families, domains, dimensions, or the full arena. Failed, timed-out, or invalid runs are recorded explicitly. Primary metrics are directed-edge precision, recall, F1, and SHD relative to Ai ; slice scores follow Eq. (6), so rankings can be read by family, mechanism, domain, sample size, or intervention protocol as well as in pooled form. Thresholding, CPDAG/PAG conversion rules, runtime budgets, and wrapper defaults are specified in Section B. Under this protocol the arena compares 18 methods with common wrappers: a random-order regression control (RandomRegress); graphical search (CDIS, GIES, IGSP, and PC); functional-assumption methods (DAS and LiNGAM); continuous optimization (DAGMA, NOTEARS, NOTEARS-MLP, and SDCD); and pretrained or amortized methods (Arrow, AVICI, CauScale, CDFM, FoundCause, SEA, and TabCausal). The next section reports what these methods do under the current release.

5

Results and Observations

We evaluate the methods and SCM pool from Section 4 under the shared protocol in Section B. The focus here is on observations: how rankings move across families and splits, how real tables and protocol choices change those rankings, and what that implies for foundation-model evaluation. Factor-level and Semantic/Formula analyses follow in Section 6. 5.1

Main Results

Figure 1 previews cross-family ranking instability under the shared observation-only protocol. For each SCM or real dataset, we first average F1 and SHD separately over repeated runs, yielding one score per metric per unit. The intro figure then compares every method pair on F1 and on SHD independently (higher F1 wins; lower SHD wins; ties count as 0.5) and reports the pooled win fraction over both metric-wise comparisons within each family (full matrix in Figure 19). The tables and plots below use the same per-unit, replicate-averaged scores, summarized as family-level means. Table 2 compares all evaluated methods on Synthetic, Semantic, and Formula benchmarks. Figure 4 plots directed-edge precision against recall and F1 against SHD, with obs-only and obs+int shown together. Complete method-wise metrics are in Section C.1 for Synthetic, Section C.2 for Semantic, and Section C.3 for Formula. Table 2: Main benchmark results (F1 / SHD). Cells report mean (standard deviation). Columns labeled obs+int are observationalplus-interventional, not intervention-only. Higher F1 and lower SHD are better; bold and underline mark the best and second-best F1 and SHD separately in each column. Avg. is the unweighted mean over the available family–split cells for that method, so methods without intervention support are averaged over three obs-only cells, not six. SHD is not commensurate across families of different graph size. Type

Method

Syn. obs

Syn. obs+int

Sem. obs

Sem. obs+int

Form. obs

Form. obs+int

Avg.

Control

RandomRegress

0.28 (0.07) / 218.0 (308.7)

–

0.25 (0.07) / 93.2 (69.0)

–

0.20 (0.06) / 108.3 (69.9)

–

0.24 (0.07) / 139.8 (149.2)

Graphical search

CDIS GIES IGSP PC

0.38 (0.14) / 143.0 (351.2) 0.40 (0.13) / 140.1 (352.9) 0.40 (0.17) / 31.7 (9.0) 0.40 (0.16) / 31.9 (8.9) 0.23 (0.18) / 31.0 (6.8) 0.23 (0.18) / 30.9 (6.9) 0.34 (0.16) / 68.1 (122.6) 0.48 (0.13) / 123.6 (174.1) 0.48 (0.11) / 134.0 (180.7) 0.55 (0.13) / 33.4 (15.2) 0.57 (0.13) / 32.9 (15.4) 0.47 (0.12) / 36.7 (14.1) 0.52 (0.13) / 34.7 (14.0) 0.51 (0.13) / 65.9 (68.9) 0.41 (0.12) / 135.6 (191.9) 0.39 (0.12) / 132.8 (185.7) 0.31 (0.16) / 40.8 (13.8) 0.31 (0.16) / 39.9 (12.9) 0.21 (0.15) / 43.3 (16.9) 0.16 (0.16) / 40.7 (15.9) 0.30 (0.15) / 72.2 (72.8) 0.31 (0.12) / 115.4 (148.9) – 0.33 (0.15) / 33.5 (9.4) – 0.22 (0.16) / 31.1 (6.8) – 0.29 (0.14) / 60.0 (55.0)

Functional DAS assumptions LiNGAM

0.34 (0.11) / 142.6 (179.8) 0.21 (0.12) / 109.6 (142.2)

– –

0.21 (0.06) / 59.5 (15.4) 0.20 (0.08) / 38.9 (9.3)

– –

0.16 (0.06) / 64.8 (15.5) 0.24 (0.08) / 34.7 (8.0)

– –

DAGMA 0.28 (0.12) / 107.2 (138.6) – 0.18 (0.07) / 42.0 (9.1) – 0.18 (0.07) / 39.7 (8.6) – Continuous NOTEARS 0.28 (0.11) / 104.2 (136.0) – 0.20 (0.07) / 38.1 (8.5) – 0.18 (0.08) / 35.2 (7.3) – optimization NOTEARS-MLP 0.35 (0.10) / 107.0 (136.3) – 0.22 (0.08) / 44.9 (10.7) – 0.19 (0.06) / 44.1 (8.9) – SDCD 0.38 (0.09) / 135.5 (165.2) 0.38 (0.09) / 148.1 (174.1) 0.36 (0.09) / 51.2 (16.0) 0.36 (0.09) / 51.7 (15.8) 0.28 (0.08) / 61.5 (16.0) 0.29 (0.09) / 61.6 (16.3)

Pretrained

Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

0.34 (0.13) / 106.0 (140.6) – 0.31 (0.09) / 36.6 (9.5) – 0.32 (0.15) / 104.5 (139.8) 0.35 (0.16) / 104.2 (139.8) 0.28 (0.13) / 35.5 (9.3) 0.30 (0.13) / 34.8 (9.3) 0.41 (0.11) / 85.8 (107.1) 0.46 (0.10) / 84.6 (106.6) 0.25 (0.07) / 50.2 (19.2) 0.29 (0.08) / 48.6 (18.8) 0.42 (0.14) / 136.4 (176.5) – 0.43 (0.07) / 46.0 (12.4) – 0.62 (0.12) / 80.8 (111.8) – 0.63 (0.11) / 26.8 (10.5) – 0.40 (0.13) / 99.2 (129.1) – 0.33 (0.11) / 41.8 (17.9) – 0.50 (0.16) / 93.6 (127.3) 0.50 (0.15) / 96.0 (128.4) 0.42 (0.13) / 36.1 (10.1) 0.45 (0.13) / 35.0 (10.2)

0.27 (0.08) / 35.9 (8.6) – 0.33 (0.12) / 31.8 (8.2) 0.41 (0.12) / 29.3 (8.0) 0.21 (0.07) / 65.9 (23.4) 0.28 (0.09) / 59.6 (23.1) 0.41 (0.08) / 45.8 (13.9) – 0.52 (0.12) / 33.4 (12.3) – 0.24 (0.09) / 52.8 (19.7) – 0.47 (0.13) / 29.5 (9.5) 0.53 (0.14) / 26.6 (9.2)

0.23 (0.08) / 89.0 (70.2) 0.22 (0.10) / 61.1 (53.2) 0.21 (0.09) / 63.0 (52.1) 0.22 (0.09) / 59.2 (50.6) 0.26 (0.08) / 65.3 (52.0) 0.34 (0.09) / 84.9 (67.2) 0.31 (0.10) / 59.5 (52.9) 0.33 (0.14) / 56.7 (52.4) 0.32 (0.09) / 65.8 (49.7) 0.42 (0.10) / 76.1 (67.6) 0.59 (0.12) / 47.0 (44.9) 0.32 (0.11) / 64.6 (55.6) 0.48 (0.14) / 52.8 (49.1)

The main observation is that rankings depend strongly on the evaluation environment. FoundCause is strongest among the supported observation-only slices (F1 0.62/0.63/0.52 on Synthetic/Semantic/Formula), GIES is competitive on several intervention-aware slices (best Semantic obs+int F1 0.57), and TabCausal is among the stronger pretrained methods, especially with interventional evidence (Formula obs+int F1 0.53). No method wins every slice or metric: F1 and SHD sometimes disagree, and the order changes across Synthetic, Semantic, and Formula. A single pooled score would hide these moves. 10

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Control

Graphical

Functional

Optimization

Synthetic better

FoundCause

Recall

0.5 F1 = 0.3

CDFM

GIES

0.4

0.6

PC

0.2

0.4

CDIS SEA NT-MLPCauScale Arrow AVICI DAGMA NOTEARS

0.3 0.2

LiNGAM

0.1 0.2

FoundCause GIES

0.4

0.1

0.6

F1 (higher is better)

0.6

0.8

0.4

0.4

0.3

0.2 0.1

0.6

0.6

0.8

0.6

0.2

CDFM CDIS SDCD DAS

IGSP SEA Arrow

RandReg →

0.4

0.6

better

FoundCause

0.5 TabCausal GIES

GIES

0.5

AVICI NT-MLP PC NOTEARS DAGMA

0.2

AVICI

Precision

FoundCause

0.4

GIES

Arrow

NT-MLP LiNGAM DAS PC CDIS DAGMA NOTEARS

Precision

FoundCause

TabCausal

SDCD RandReg SEA IGSP

PC Arrow DAS NT-MLP LiNGAM AVICI DAGMA NOTEARS

better

CauScale

FoundCause CDFM GIES

0.3CauScale

CauScale

0.7

TabCausal

0.6

0.4

SDCD TabCausal CDIS RandReg IGSP SEA

0.2

better

0.5

better

0.5

Precision

0.7

0.7

CDFM

TabCausal

RandReg DAS

0.3

obs+int

Formula better

0.7

0.5

SDCD IGSP

obs-only

Semantic

0.7 0.6

Pretrained

A P REPRINT

CDFM

CDIS

TabCausal CDFM SEA PC Arrow SDCD RandReg → IGSP AVICI CauScale NOTEARS LiNGAM NT-MLP DAS DAGMA

0.4 0.3 0.2

LiNGAM

AVICI

0.3

Arrow

SDCD

CDIS LiNGAM IGSP SEA RandReg → PC DAGMA NT-MLP CauScale NOTEARS DAS

0.2 0.1

80

100

120

140

SHD (lower is better)

160

20

30

40

50

SHD (lower is better)

60

20

30

40

50

60

70

SHD (lower is better)

Figure 4: Precision–recall and F1–SHD on the main benchmark. Top row: directed-edge precision versus recall. Bottom row: F1 versus SHD. Circles are obs-only, diamonds are obs+int, and a line joins the two when both exist. Shaded corners and the italic better marker indicate the preferred region (higher precision and recall; higher F1 and lower SHD). Axes are cropped to the plotted methods; RandomRegress is marked at the right edge when its SHD lies off-scale. Each method is labeled once; RandReg is RandomRegress and NT-MLP is NOTEARS-MLP. Gray curves are constant F1. SHD axes are independent across families. Figure 4 makes the same pattern geometric. Most methods occupy a mid-precision, mid-recall band; FoundCause sits farther toward the high-F1 corner on observation-only slices, while RandomRegress is the high-SHD outlier (off-scale on the bottom row). The obs+int setting is a fixed-budget mixed-evidence protocol: it replaces 200 observational rows with interventional rows instead of adding interventions on top of the same observational sample. Thus, moving from obs-only to obs+int changes both the information type and the sample composition, producing modest gains or trade-offs between F1 and SHD depending on the method. Semantic is more spread out than Formula, where several methods collapse toward similar precision–recall and F1–SHD values. Factor-level and semantic analyses in Section 6 explain these shifts. 5.2

Real-data Check

Figure 5 evaluates the same wrappers on real scored tabular datasets with public samples and published directed graphs. The observation-only sources are six CD-CSG datasets [44]: Abalone, Auto MPG, Cardiac Arrhythmia, Concrete Compressive Strength, Deutscher Wetterdienst, and Ozone. The interventional sources are Sachs flow-cytometry [13]; PetShop high traffic, low traffic, temporal traffic 1, and temporal traffic 2 [45]; and the Causal Chambers Light Tunnel and Wind Tunnel experiments [16]. The observation-only aggregate combines these six CD-CSG datasets with observation-only conversions of the seven interventional datasets, obtained by dropping intervention rows; the interventional aggregate uses the seven datasets with intervention rows. On these tables the synthetic ranking does not carry over. CDFM leads observation-only F1 (0.39), TabCausal is second (0.34) and then leads observation-plus-intervention (0.46), while FoundCause falls to 0.21. Figure 5 shows why the means are only a coarse summary: within a method the dataset markers span a wide F1 range, so a few tables dominate both the average and the standard deviation. Because obs-only averages include six additional CD-CSG datasets, intervention effects should be read on the seven paired interventional sources: TabCausal, AVICI, CauScale, and SDCD improve there, GIES and IGSP are nearly flat, and CDIS drops. This track checks whether the same wrappers still

11

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Type

Method

Obs-only

Obs-only

Obs+int

Control

RandomRegress 0.20 (0.14) / 59.9 – CDIS 0.28 (0.22) / 99.5 0.14 (0.17) / 45.4 GIES 0.26 (0.15) / 48.0 0.29 (0.13) / 118.7 Graphical IGSP 0.18 (0.22) / 31.1 0.11 (0.12) / 59.1 PC 0.18 (0.14) / 28.9 – DAS 0.19 (0.23) / 61.9 – Functional LiNGAM 0.13 (0.17) / 24.3 – DAGMA 0.16 (0.14) / 31.8 – NOTEARS 0.16 (0.15) / 29.1 – Optim. NOTEARS-MLP 0.21 (0.19) / 55.0 – SDCD 0.22 (0.17) / 44.5 0.27 (0.12) / 103.4 Arrow 0.15 (0.13) / 70.5 – AVICI 0.22 (0.18) / 26.0 0.28 (0.12) / 42.6 CauScale 0.08 (0.08) / 36.2 0.19 (0.07) / 62.1 0.39 (0.14) / 38.5 – Pretrained CDFM FoundCause 0.21 (0.23) / 98.3 – SEA 0.21 (0.14) / 31.3 – TabCausal 0.34 (0.24) / 28.4 0.46 (0.11) / 38.4

A P REPRINT

Obs+int

CauScale LiNGAM Arrow NOTEARS DAGMA IGSP PC DAS RandomRegress NOTEARS-MLP SEA FoundCause AVICI SDCD GIES CDIS TabCausal CDFM

IGSP

CDIS

CauScale

SDCD

AVICI

GIES

TabCausal

0.0

0.2

0.4

0.6

0.8

1.0

0.0

0.2

F1 CD-CSG

0.4

0.6

0.8

1.0

F1 Sachs

PetShop

Chambers

Mean

Figure 5: Real-data check. Left: compact F1 / SHD results, with cells reporting mean (std) over datasets (obs-only: 13; obs+int: 7) and SHD shown to one decimal for space. Right: per-dataset real-data F1; each marker is one scored table, and the tick is the method mean. recover published graphs on public real tables, next to the executable SCM arena. Those graphs are expert-specified or experimentally motivated in the source papers; they are not independently verified unique true DAGs. Complete method-wise and per-dataset metrics are in Section C.4. 5.3

Sample Size and Intervention Protocol

These experiments keep the synthetic SCM collection fixed and change sample size or intervention design. The sample-size suite fixes d = 30, uses 100 synthetic SCMs, and varies the observation-only sample size n ∈ {100, 200, 500, 1000, 2000, 5000, 10000}. Figure 6a plots F1 against n for all evaluated methods. The corresponding F1/SHD table is in Section C.5. Most of the F1 gain occurs between n = 100 and n = 1,000; beyond n = 2,000 the curves flatten. FoundCause remains highest throughout (0.43 at n = 100, 0.64 at n = 1,000, 0.66 at n = 10,000), and TabCausal has the next largest rise (0.33 to 0.57). These absolute levels should be read with the suite’s fixed d = 30 SCM mix in mind: the curves isolate how F1 changes with n, while a method’s height can also reflect affinity to that held-fixed graph, mechanism, and noise composition. Search and optimization methods such as GIES, IGSP, CDIS, SDCD, and NOTEARS-MLP improve through n = 1,000 and then stall. NOTEARS, DAGMA, LiNGAM, and RandomRegress barely move. Arrow and CauScale are not monotone: both peak near n = 1,000–2,000 and then drop at n = 10,000 (Arrow 0.36 → 0.33, CauScale 0.41 → 0.35), so extra observational samples do not uniformly help pretrained wrappers. The intervention-protocol suite also fixes d = 30 with 100 SCMs and compares the default resampling intervention with fixed-setpoint, parameter-shift, high-intervention-fraction, multi-target-per-sample, and dense mixed-intervention settings. Table 3 reports F1 and SHD; Figure 6b plots the F1 change relative to Default. Table 3: Intervention-protocol sensitivity. Cells report F1 / SHD; non-default protocols put the gray F1 change relative to Default on the second line. The last column is the standard deviation of each method’s F1 across intervention protocols (lower is better). Bold and underline mark the best and second-best F1, SHD, and F1 SD separately. Method

Default

CDIS

0.397 / 48.1

GIES

0.480 / 59.7

IGSP

0.400 / 53.7

SDCD

0.380 / 73.0

AVICI

0.390 / 43.4

CauScale

0.451 / 43.6

TabCausal

0.539 / 45.6

Fixed

Shift

Multi

Dense

High-frac

0.396 / 48.3

0.397 / 48.3

0.382 / 48.9

0.365 / 47.2

0.355 / 47.6

∆F1 -0.001

∆F1 +0.000

∆F1 -0.015

∆F1 -0.032

∆F1 -0.041

0.480 / 59.7

0.429 / 64.1

0.483 / 59.3

0.497 / 56.3

0.506 / 56.7

∆F1 +0.000

∆F1 -0.051

∆F1 +0.004

∆F1 +0.017

∆F1 +0.026

0.398 / 54.2

0.394 / 54.4

0.127 / 54.6

0.248 / 51.8

0.312 / 52.0

∆F1 -0.001

∆F1 -0.006

∆F1 -0.273

∆F1 -0.152

∆F1 -0.088

0.381 / 72.7

0.375 / 73.2

0.385 / 72.4

0.404 / 68.3

0.398 / 70.3

∆F1 +0.002

∆F1 -0.004

∆F1 +0.005

∆F1 +0.024

∆F1 +0.018

0.385 / 43.7

0.365 / 44.6

0.385 / 43.7

0.371 / 44.3

0.404 / 42.9

∆F1 -0.005

∆F1 -0.024

∆F1 -0.005

∆F1 -0.019

∆F1 +0.014

0.458 / 43.0

0.418 / 45.2

0.438 / 45.7

0.424 / 46.7

0.484 / 40.8

∆F1 +0.006

∆F1 -0.034

∆F1 -0.014

∆F1 -0.027

∆F1 +0.032

0.538 / 45.8

0.517 / 47.6

0.538 / 45.7

0.526 / 46.1

0.542 / 44.8

∆F1 -0.001

∆F1 -0.022

∆F1 -0.001

∆F1 -0.013

∆F1 +0.003

12

F1 SD 0.017 0.024 0.100 0.010 0.013 0.022 0.009

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Observation-only sample-size sensitivity

Mean F1

0.5 0.4 0.3 0.2 100

200

500

1k

2k

5k

F1 vs. Default

FoundCause TabCausal GIES CDFM CDIS SDCD IGSP SEA NOTEARS-MLP DAS AVICI CauScale Arrow PC DAGMA NOTEARS RandomRegress LiNGAM

0.6

10k

Observation-only sample size Control Graphical

Functional Optimization

0.05 0.00 0.05 0.10 0.15 0.20 0.25 0.30

A P REPRINT

F1 change across intervention protocols -0.041 -0.051 -0.088 -0.152

-0.273

CDIS

IGSP

GIES

Pretrained

Fixed

Shift

SDC

D

I AVIC

Multi

Dense

Cau

e Scal

ausa

TabC

l

High-frac

Figure 6: Sample-size and intervention-protocol sensitivity. (a) Mean F1 on the synthetic d=30 observation-only suite; line styles encode method class, and the right column lists full method names ordered by final F1. (b) F1 change relative to the default mixed protocol for each method that supports the intervention-protocol suite; negative values are drops. The F1-variation column further separates methods whose scores stay similar from those that move: TabCausal and SDCD vary least across these protocols (F1 SD 0.009 and 0.010), while IGSP is most sensitive (F1 SD 0.100). Figure 6b localizes that drop. Fixed-setpoint stays near Default for every method; the bars that stand out are IGSP under multi-target, dense mixed, and high-intervention-fraction interventions (∆F1 −0.27, −0.15, and −0.09). This is consistent with IGSP’s reliance on conditional-independence and invariance tests, which need sufficient per-intervention samples to be reliable; prior benchmark discussions have likewise noted that IGSP can underperform when only a small number of intervention samples are available per node. GIES, which uses the same interventional tables, rises under dense mixed and high-intervention-fraction (∆F1 +0.02 and +0.03) and drops mainly under parameter-shift. Complete protocol-wise metrics are in Section C.6. 5.4

Ablation Studies

Construction pipeline. Figure 7 compares successive stages of the agentic construction pipeline using an LLM-judge quality rubric on 20 sampled scenarios, 10 from Semantic and 10 from Formula. The full construction process is described in Section A. The reference pack adds external factual context; planning organizes variables and causal chains before graph drafting; graph review removes aliases, unsupported shortcuts, and low-information support nodes. In Figure 7 the average rises from 3.80 (seed only) to 3.87 after the reference pack, 4.37 after planning, and 4.48 after graph review. The reference pack mainly lifts semantic/mechanism diversity; planning is the largest step, raising causal plausibility, mechanism grounding, benchmark validity, and specifiability together; graph review adds a smaller further gain on the same criteria. Structural diversity is already high at the seed stage and stays flat. The score pattern supports keeping these stages. Construction-quality scores by pipeline stage LLM-judge score (1–5)

5.0

Causal plausibility Mechanism grounding Benchmark validity Specifiability Structural diversity Sem./mech. diversity Average

4.5

4.0

3.5

A0 Seed

A1 + ref.

A2 + plan

A3 + review

Figure 7: Construction-pipeline ablation. LLM-judge stage curves over 20 scenarios (10 Semantic, 10 Formula); the thicker line is the reported average. Criterion-level means with standard deviations are in Table 20.

13

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

5.5

A P REPRINT

Runtime and Resource

Pretrained and amortized wrappers pay a one-time load before they score a graph. On a short evaluation shard that load can dwarf the inference itself. We therefore time nested runs that score k graphs after one load, using a synthetic sample balanced across dimensions and graph families, and fit T (k) = Tfixed + kTgraph , where Tfixed is the load and Tgraph is the extra seconds for one more graph once the model is already in memory. The nested k grid is larger at low d, where each extra graph is cheap and a short run would make the load/inference split unstable, and smaller at high d, where each extra graph is expensive: k ∈ {100, 200, 500, 1000} at d = 10, k ∈ {50, 100, 200, 500} at d = 20, 30, and k ∈ {20, 50, 100, 200} at d = 50, 100. Table 4: Load and per-graph time for pretrained/amortized methods. Fits use a synthetic sample balanced across dimensions d ∈ {10, 20, 30, 50, 100} and graph families. Tfixed is the fitted one-time load, averaged over d. Tgraph is the extra seconds for one more graph after the model is loaded, reported by d and as an unweighted mean over those d. Nested k is {100, 200, 500, 1000} at d=10, {50, 100, 200, 500} at d=20, 30, and {20, 50, 100, 200} at d=50, 100: larger k at low d for a stable split, smaller k at high d because extra graphs cost more. Lower time is better. Bold is lowest in a time column and underline is second; R2 columns are not ranked. All method–d fits have R2 > 0.97, and all but one exceed 0.98. Per-graph seconds Tgraph Tfixed

d=10

d=20

d=30

d=50

d=100

Mean

Mean R2

Min R2

Arrow AVICI CDFM TabCausal CauScale SEA FoundCause

14.47 55.84 10.02 8.13 11.07 28.14 23.46

0.023 0.031 0.068 0.140 0.140 0.232 1.146

0.040 0.052 0.098 0.171 0.145 0.278 1.861

0.065 0.081 0.108 0.170 0.190 0.310 2.310

0.081 0.104 0.249 0.224 0.398 0.341 5.729

0.176 0.215 0.472 0.346 1.170 0.993 12.725

0.077 0.096 0.199 0.210 0.409 0.431 4.754

0.9905 0.9933 0.9947 0.9993 0.9902 0.9951 0.9995

0.9837 0.9897 0.9922 0.9979 0.9718 0.9898 0.9984

Observed total time (sec.) Observed total time (sec.)

Method

Arrow

AVICI

60 40 20 0

250 200 150 100 50 0

0

500

CauScale

1000

100 80 60 40 20 0

0

500

SEA

300

CDFM 100 80 60 40 20 0

1000

2500 2000 1500 1000 500 0

200 100 0

500

1000

Graphs in one run k

0

0

500

1000

Graphs in one run k d = 10

d = 20

TabCausal

150 100 50

0

500

1000

FoundCause

0

500

1000

Graphs in one run k

Dimension d = 30 d = 50

0 Arrow AVICI CDFM TabCausal CauScale SEA FoundCause

0

0

500

1000

Load intercept

25

50

Tfixed (sec.)

d = 100

Figure 8: Pretrained runtime fit quality. Each line panel is one pretrained/amortized wrapper on its own vertical scale. Points are observed total wrapper time on a synthetic sample balanced across dimensions and graph families; a line is the fitted T (k) = Tfixed + kTgraph within dimension d. Dots on the line mean the split is stable; a steeper line, on that panel’s scale, costs more per extra graph. The last panel shows the fitted load Tfixed . Table 4 and Figure 8 split that cost. All method–d fits have R2 > 0.97, and all but CauScale at d = 50 exceed 0.98. After load, Arrow and AVICI are the cheapest extra graphs, CDFM and TabCausal are next, CauScale and SEA are slower, and FoundCause sits in a higher band. TabCausal also has the shortest mean load (8.13s); AVICI has the longest (55.84s). The later comparison still drops Tfixed . A pretrained wrapper is usually loaded once and then applied to 14

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

many graphs—a leaderboard pool, a dimension or family shard, or a batch of related systems—so the one-time cost is Tfixed /k per graph and becomes small once k is large. The comparable quantity that remains is therefore Tgraph . Table 5 reports one comparable per-graph number for every method. At each d, the nested-k sample draws equally many graphs from each graph family, matching the synthetic-main mix, so the ranking is comparable; Tfixed is omitted because a large-k run amortizes the load. Table 5: Per-graph runtime on the synthetic mix. Non-pretrained entries are mean end-to-end wrapper time on the synthetic main benchmark, mean (standard deviation). † marks the unweighted mean of per-d Tgraph from nested-k fits. At each d, the sample draws equally many graphs from each graph family, matching the synthetic-main mix, so the times are comparable; Tfixed is omitted because it amortizes at large k. Lower is better. Rank 1 2 3 4 5 6

Method Arrow AVICI CDFM TabCausal CauScale SEA

Sec./graph

Rank

†

0.077 0.096† 0.199† 0.210† 0.409† 0.431†

Method

7 8 9 10 11 12

RandomRegress GIES FoundCause DAGMA PC IGSP

Sec./graph

Rank

0.82 (0.53) 2.78 (1.19) 4.754† 5.33 (3.01) 20.36 (146.76) 26.30 (131.27)

13 14 15 16 17 18

Method NOTEARS LiNGAM DAS CDIS NOTEARS-MLP SDCD

Sec./graph 30.46 (57.50) 36.17 (59.43) 57.60 (62.02) 75.75 (253.70) 80.27 (99.35) 192.64 (115.28)

Accuracy–cost frontier on the main synthetic pool FoundCause

Average F1

0.6

l Optima

0.4

CDFM AVICI

0.3

GIES

TabCausal

0.5

SDCD

CDIS

SEA PC

IGSP

Arrow CauScale DAGMA RandomRegress

0.2 10−1

NOTEARS

101

100

LiNGAM NOTEARS-MLP DAS

102

Seconds per graph (log scale) Warm amortized inference

End-to-end wrapper time

Pareto frontier

Figure 9: Accuracy–cost frontier. Average F1 from Table 2 against the comparable per-graph times in Table 5. Diamonds use pretrained Tgraph from a sample balanced across synthetic dimensions and graph families, omitting Tfixed ; circles use mean end-to-end time on the synthetic main benchmark. Average F1 follows the table Avg. column, which averages only the family–split cells available for each method. The dashed line is the Pareto frontier; Optimal points toward higher F1 and lower time. Figure 9 uses the same comparable per-graph times. For pretrained methods the x-coordinate is Tgraph with Tfixed omitted, because a large-scale causal-discovery run amortizes the load. The useful frontier is not simply the set of fastest methods. Arrow and AVICI are very cheap but have lower average F1; CDFM and TabCausal move the frontier upward at nearly the same cost scale; GIES remains the strongest low-cost graphical-search baseline; and FoundCause reaches the highest average F1 among its supported observation-only cells at a larger Tgraph . Several classical or optimization methods are dominated in this view: they are slower than the frontier methods while also having lower average F1. Documented pretraining priors do not cover this benchmark equally, so part of a pretrained score could sit on documented support. A later observation-only check in Section 6.2 splits categories by each method’s own audited prior and finds only modest ID-to-OOD shrinks in relative lead for that reference group; the coarse order is stable in a coverage-mix reading, with the limits of that diagnostic stated there.

6

Analysis

A released suite still needs to be checked for breadth and for whether scores can be traced to specific graph, mechanism, or domain properties. Following the diagnosability requirement in Section 3.2, this section reports synthetic factor analyses, a check of pretrained methods on categories outside their documented training support, and Semantic/Formula domain breakdowns. 15

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

6.1

A P REPRINT

Synthetic Benchmark Behavior Analysis

Interaction score. Aggregate leaderboards rank methods, but they do not say why the ranks differ. We analyze the synthetic obs-only benchmark with factor-specific interaction scores. For a method m and factor category c, let Fm,c be the mean F1 on that category. We report Im,c = Fm,c − F̄m,· − F̄·,c + F̄·,· .

(7)

This double-centering removes each method’s overall strength and each category’s overall difficulty. A positive value means that a method is stronger on that category than expected from both marginal averages. For graph-level factors, categories are assigned per graph. For edge-level factors, mechanisms and noises are assigned to an edge by the child node’s structural equation.

scale free -0.02 0.05 0.00 -0.06 -0.08 -0.06 0.00 0.03 0.04 0.05 0.03 -0.01 0.01 -0.00 -0.06 0.01 0.06 0.00 small world

0.04 -0.00 0.01 0.03 0.00 -0.01 -0.02 -0.00 0.00 -0.00 0.01 -0.01 -0.02 0.01 -0.03 -0.02 0.01 -0.00

tree -0.03 0.02 -0.01 -0.03 -0.09 -0.06 0.00 0.04 0.04 0.06 -0.01 -0.01 0.02 -0.01 -0.03 -0.01 0.06 0.03

0.15 0.10 0.05

chain 0.03 -0.02 0.02 0.04 -0.04 -0.02 -0.02 0.03 0.02 0.02 -0.03 -0.00 -0.01 -0.01 -0.03 -0.02 0.03 0.03 0.00

layered -0.01 -0.01 0.01 0.01 0.02 0.04 -0.02 -0.01 -0.03 -0.03 0.01 0.01 0.02 -0.02 0.04 0.02 -0.03 -0.01 collider rich -0.00 0.00 -0.01 -0.01 0.05 0.03 -0.02 -0.03 -0.02 -0.03 0.02 -0.00 -0.00 0.00 0.02 0.01 -0.01 -0.01 dense local modules 0.02 0.01 0.02 0.03 -0.05 -0.04 0.02 0.01 0.02 0.02 -0.03 0.01 -0.02 0.02 -0.02 -0.01 0.00 -0.02 block sparse layered -0.01 -0.01 0.01 0.01 -0.06 -0.03 0.02 0.05 0.02 0.05 -0.05 0.01 0.01 -0.01 -0.00 -0.01 -0.01 0.00

−0.05 −0.10 −0.15

double-centered F1 interaction

er -0.00 0.01 -0.01 -0.01 0.05 0.04 -0.02 -0.02 -0.02 -0.04 0.04 -0.00 -0.01 -0.00 0.01 0.02 -0.01 -0.02

bipartite layered -0.02 -0.05 -0.04 -0.03 0.19 0.12 0.05 -0.09 -0.07 -0.11 0.01 -0.00 0.00 0.03 0.10 0.01 -0.09 -0.01

s

m

o nd Ra

es gr Re

M AS D GA N i L

l I PC DIS IES GSP ARS MLP MA CD row ale FM use SEA VIC usa I G SD Ar G C A a Sc CD Ca TE RS DA u C d b O Ca N EA un Ta T Fo O N

Figure 10: Synthetic graph-family factor interactions. Each cell is a double-centered F1 interaction on the observationonly synthetic benchmark. Graph structure. Figure 10 shows that methods differ by graph topology even after category difficulty is removed. The clearest contrast is the bipartite-layered family: CDIS (+0.192), GIES (+0.120), and FoundCause (+0.098) are relatively stronger, while DAGMA (−0.108), NOTEARS (−0.089), and AVICI (−0.089) are relatively weaker. For CDIS/GIES, this is consistent with graphical-search behavior: bipartite and collider-rich structures create many separation and v-structure signals that conditional-independence or equivalence-class search can exploit. FoundCause is pretrained, but its pairwise-statistics pathway and triangular refinement module are also designed to reason over higher-order motifs such as chains and colliders, which plausibly explains why it benefits from the same structured signals. By contrast, continuous-optimization methods and AVICI are more favorable on sparse branching families such as trees or scale-free graphs, where recovering a weighted acyclic parent structure is closer to their optimization or amortized parent-selection bias. Scale effects. Dimension is ordered, so Figure 11 reports it as curves. Almost every method loses F1 as d grows, but the slopes differ. FoundCause stays highest throughout (0.67 at d = 10, 0.56 at d = 100) while declining. CauScale is the exception that rises (0.35 to 0.48), matching its design goal and training setup for large graphs: its synthetic training includes graph sizes up to hundreds of nodes, and its reduction/tied-attention design is optimized for scaling. GIES, CDIS, and SDCD are comparatively flat. Arrow, CDFM, and AVICI drop sharply by d = 100 (to 0.21, 0.27, and 0.16); this likely reflects dense pairwise decoding, skeleton/order or edge-head uncertainty accumulating over many variable pairs, and training distributions whose strongest support is not centered on our high-dimensional heterogeneous synthetic pool. TabCausal also falls (0.56 to 0.36) but remains above those collapsing pretrained methods. Both TabCausal and FoundCause combine broad synthetic task variation with explicit pairwise or motif-level signals, which helps preserve relative performance as the variable set grows.

16

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Observation-only F1 versus dimension

0.7 0.6

FoundCause

Mean F1

0.5

CauScale GIES SDCD TabCausal CDIS IGSP NOTEARS-MLP SEA CDFM DAS RandomRegress DAGMA NOTEARS Arrow PC

0.4 0.3 0.2 0.1

AVICI LiNGAM

10

20

30

50

100

Number of variables d Control

Graphical

Functional

Optimization

Pretrained

Figure 11: Synthetic dimension scaling. Mean observation-only F1 versus the number of variables d. Line styles follow method families, as in the sample-size plot. CauScale is the only method whose F1 rises with d. Mechanism and noise effects. Figure 12a shows a clear split between linear-identifiability assumptions and nonlinear mechanism coverage. On the linear-additive row, LiNGAM is the strongest positive method (+0.069), which is expected because this setting matches its linear non-Gaussian identification principle. NOTEARS is nominally a linear-SEM method, but its squared-loss objective is closer to a linear-Gaussian score and does not exploit non-Gaussian residual information for orientation; since most linear-additive cases in this pool are not Gaussian-noise linear SEMs, NOTEARS is not helped in the same way (−0.023). GIES follows LiNGAM closely (+0.053 on linear-additive, and a row-wise mechanism correlation of about 0.95 with LiNGAM), even though the algorithms are different. The common pattern is that both are most comfortable when causal effects are globally simple and separable, and both lose relative advantage when the edge signal is local, nonlinear, or non-smooth. The tree-based row creates the strongest mechanism split. AVICI (+0.100), CDFM (+0.086), and NOTEARS-MLP (+0.082) are strongly positive, while GIES (−0.091), LiNGAM (−0.074), CauScale (−0.054), and TabCausal (−0.041) are negative after double-centering. This row corresponds to piecewise and threshold-like parent effects: the relevant parent set can be recoverable when a model has enough nonlinear function capacity or has seen similar nonlinear signatures during pretraining, but the same structure is hard for linear/non-Gaussian assumptions or score-based search to express as a stable global dependence pattern. The negative values for some pretrained models mark relative specialization: their overall accuracy may remain strong, but their advantage is less concentrated on this particular piecewise mechanism class. AVICI is especially mechanism-sensitive in this panel: its interaction variance across mechanism rows is the largest among the pretrained methods, with large positive values on tree-based, polynomial-power (+0.097), and RFF-GP (+0.083) rows but a large negative value on linear-additive (−0.081). We therefore read AVICI as less stable across mechanism families than broader or more motif-oriented pretrained methods such as TabCausal and FoundCause, without attributing this pattern to a single mechanism property. The noise panel separates whether scores move more with the noise family or with the mechanism family. Overall, noise families create weaker method separation than child mechanisms: the average row spread is smaller for noise than for mechanisms, although mixture-outlier and RFF-heteroskedastic noise are strong exceptions. Graphical-search methods are comparatively flat across noise rows, whereas pretrained methods show more visible noise-dependent swings than graphical search and most continuous-optimization methods. This sensitivity is less distinctive in the graph-family and mechanism panels, where variation is shared across several method classes. Noise families therefore show how each method uses residual shape, scale variation, and support constraints. The largest contrasts come from noises that change more than marginal variance. Mixture-outlier noise injects sparse contamination, so methods that rely on smooth losses, ordering heuristics, or stable low-order associations can be pulled toward spurious supports; this is consistent with negative interactions for SDCD (−0.097) and DAS (−0.061). LiNGAM is the main exception (+0.126): in this slice, outliers amplify non-Gaussian residual cues, which its identification principle can sometimes exploit. RFF-heteroskedastic noise produces a different split because the noise scale changes with parent values, so the edge evidence is a conditional-variance signal. NOTEARS-MLP is strongly positive (+0.097), plausibly because its nonlinear regression model can absorb associated nonlinear parent-dependent patterns, while CDFM is also positive (+0.069), consistent with explicit heteroskedastic coverage in its pretraining prior. Several pretrained methods are negative on this row: even a broad pretraining prior can remain sensitive to variance-driven 17

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

activation additive 0.01 -0.00 0.00 0.00 0.02 0.01 0.01 0.01 0.02 -0.01 0.00 -0.01 -0.03 0.00 -0.01 0.00 -0.01 -0.00 discretization assignment 0.03 0.03 -0.02 -0.01 0.02 0.01 -0.00 0.04 0.05 -0.01 0.01 -0.04 -0.04 0.01 -0.03 -0.03 -0.00 -0.03

0.100

exp log -0.01 -0.01 0.01 0.01 0.00 0.01 0.01 0.00 -0.02 0.00 0.01 -0.00 0.01 -0.02 0.00 0.01 -0.02 0.01

0.075

linear additive -0.02 -0.04 0.07 0.00 -0.00 0.05 0.02 -0.02 -0.05 -0.01 0.01 0.04 0.02 -0.04 0.02 0.03 -0.08 0.01 linear additive interaction -0.02 -0.04 0.02 0.01 0.01 0.02 0.01 -0.02 -0.04 -0.00 0.00 0.02 0.03 -0.03 0.02 0.02 -0.03 0.02 mixture regime -0.02 -0.02 0.02 0.00 -0.01 0.03 0.01 -0.02 -0.04 -0.01 0.02 0.03 0.02 -0.04 0.02 0.03 -0.05 0.02 multiplicative 0.02 0.03 0.01 -0.02 0.01 0.00 0.00 0.03 0.01 -0.05 0.02 -0.03 0.02 -0.00 -0.00 -0.02 -0.03 0.00 neural random 0.01 0.03 0.03 -0.01 0.02 0.03 0.01 0.04 -0.00 -0.03 0.02 -0.01 -0.02 -0.02 -0.01 -0.02 -0.06 -0.01 piecewise threshold -0.02 -0.04 0.03 0.00 -0.01 0.02 0.00 -0.03 -0.04 -0.01 -0.00 0.00 0.03 -0.02 0.03 0.03 -0.01 0.03 polynomial power -0.01 -0.04 -0.03 -0.00 -0.02 -0.05 -0.01 -0.02 0.01 0.04 -0.03 -0.01 0.02 0.03 0.01 -0.00 0.10 0.02 rational ratio 0.00 0.02 0.01 0.00 -0.01 0.00 -0.00 -0.00 0.01 0.03 -0.00 0.04 -0.01 0.00 -0.03 0.01 -0.03 -0.05

0.050 0.025 0.000 −0.025 −0.050

rff gp 0.01 0.04 -0.04 -0.01 0.01 -0.05 -0.03 -0.00 0.03 -0.02 -0.02 -0.02 0.01 0.03 0.02 -0.04 0.08 0.02

−0.075

saturation clip -0.02 -0.01 0.02 -0.00 -0.00 0.03 0.01 -0.01 -0.03 -0.01 0.02 0.02 0.00 -0.03 0.01 0.02 -0.05 0.01

−0.100

double-centered F1 interaction

aggregation nonlinear -0.01 0.01 -0.04 0.01 -0.02 -0.04 -0.02 -0.02 0.00 0.04 -0.03 0.02 -0.00 0.03 -0.01 0.01 0.07 -0.01

tree based 0.04 0.03 -0.07 0.02 -0.02 -0.09 -0.01 0.01 0.08 0.04 -0.04 -0.02 -0.05 0.09 -0.03 -0.03 0.10 -0.04 trig periodic 0.01 0.01 -0.02 -0.01 0.01 -0.01 -0.00 0.01 0.02 -0.01 0.00 -0.02 -0.01 0.01 -0.01 -0.02 0.03 -0.00 s S M PC IS ES SP RS LP A D w le M se EA ICI al s o a F M C I es DA GA CD G IG EA S-M AG SD Arr Sc CD Cau S AV au gr N T e u i C d L a R b O AR D n C N E u Ta om T Fo O nd N a R

(a) Mechanism interactions (child-node mechanism family). beta -0.03 -0.06 -0.03 -0.01 -0.00 -0.02 -0.01 -0.03 -0.04 -0.01 0.02 0.00 0.05 -0.02 0.05 0.03 0.05 0.05 exponential -0.02 -0.03 0.01 -0.00 -0.02 -0.02 -0.00 -0.02 -0.03 0.01 0.00 0.01 0.04 -0.01 0.03 0.02 0.02 0.02

0.10

gaussian 0.01 0.02 -0.02 0.00 0.00 0.01 0.00 0.01 0.00 0.00 0.01 0.01 -0.02 -0.01 -0.01 -0.00 -0.01 0.01 gumbel 0.00 0.01 0.01 0.00 0.00 0.00 -0.01 0.00 0.00 -0.01 0.00 0.01 -0.01 -0.00 -0.00 -0.00 -0.02 -0.00

0.05

heteroskedastic -0.01 0.00 -0.04 -0.00 0.01 0.00 -0.00 -0.01 -0.02 0.00 0.02 0.01 0.01 -0.02 0.01 0.01 0.00 0.02 laplace 0.00 0.00 0.02 -0.00 0.00 0.01 -0.00 0.00 0.01 0.00 0.00 -0.00 -0.00 0.01 -0.02 -0.01 -0.01 -0.00 mixture outlier 0.00 -0.06 0.13 -0.01 -0.01 -0.01 -0.01 0.01 0.01 -0.01 -0.10 -0.04 0.05 0.03 0.04 -0.03 0.05 -0.05

0.00

multiplicative lognormal -0.03 -0.04 -0.02 -0.01 -0.01 -0.01 -0.00 -0.02 -0.03 0.01 0.01 0.00 0.05 -0.02 0.03 0.03 0.02 0.03 quantized 0.01 0.03 -0.02 0.00 -0.01 0.01 0.01 0.01 0.01 0.00 0.01 0.00 -0.02 -0.00 -0.02 -0.00 -0.03 -0.01 rff heteroskedastic

−0.05

0.06 0.04 0.00 0.02 0.03 0.01 -0.00 0.03 0.10 -0.02 -0.02 -0.02 -0.09 0.07 -0.07 -0.06 -0.02 -0.06

student t -0.00 0.00 0.01 -0.01 -0.00 -0.00 0.00 0.00 -0.00 -0.01 0.00 -0.00 0.01 0.00 -0.00 -0.00 -0.00 0.00

−0.10

double-centered F1 interaction

censored 0.01 0.03 -0.03 0.00 0.01 0.02 0.01 0.01 0.01 0.00 0.02 -0.03 -0.04 0.01 -0.02 0.00 -0.00 -0.03

target r2 gaussian 0.01 0.05 -0.02 0.01 0.00 0.00 0.02 0.03 0.00 0.02 0.00 0.01 -0.04 -0.01 -0.03 -0.01 -0.05 -0.02 uniform -0.02 0.01 -0.01 -0.01 -0.01 -0.01 -0.00 -0.02 -0.03 0.01 0.01 0.02 0.00 -0.02 0.01 0.01 0.00 0.04 l I s e e S M PC IS ES SP RS LP A D A M C row al FM us SE VIC usa I IG A es DA GA CD G A E S-M AG SD Ar uSc CD Ca a gr N T e i C R D L R b O nd Ca N EA Ta om ou T F d O n N Ra

(b) Noise interactions (child-node noise family).

Figure 12: Synthetic mechanism and noise interactions. Each cell is a double-centered F1 interaction on the observation-only synthetic benchmark. evidence. RandomRegress also has visible tendencies across noise rows, for example on RFF-heteroskedastic noise (+0.059); because it is a fixed-random-order regression control, these tendencies reflect sensitivity of nonzero regression support to scale/noise artifacts. Root dependency effects. Figure 13 isolates graph-level dependence among root variables. Independent-root markers sit near zero for most methods, as expected for the setting closest to the standard SCM assumption. Search methods such as CDIS and GIES have wide rows: uniform-ball markers (triangles) lie on the positive side and common-factor correlated-normal markers (squares) on the negative side. NOTEARS, NOTEARS-MLP, and DAGMA reverse that

18

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Root-dependency profiles RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM

Independent Corr. Gaussian Random cov. Uniform ball

FoundCause SEA TabCausal −0.075

−0.050

−0.025

0.000

0.025

0.050

0.075

0.100

Double-centered F1 interaction

Figure 13: Synthetic root-dependency interactions. Each row is one method; markers are the four root-dependency categories, using double-centered F1. Wider rows are more root-sensitive. The uniform-ball versus common-factor (correlated-normal) reversal is the left–right swap of triangles and squares. order. A plausible explanation is that smooth root covariance violates exogenous independence and can be mistaken by graphical-search procedures as root–root edges or v-structures, whereas global continuous objectives can sometimes absorb this covariance with fewer structural disruptions. Bounded uniform-ball roots favor CDIS/GIES and FoundCause, consistent with a setting where support constraints and sharper conditional changes provide useful search or motif cues but are harder to express with smooth additive objectives. Together, these views show that a synthetic score can be traced to graph structure, local mechanisms, noise processes, and root-dependency patterns. 6.2

Pretraining-Exposure OOD Score Analysis

The synthetic factor analysis above shows where methods are relatively strong or weak. For pretrained and foundationstyle methods we add an observation-only check on categories they do not list as pretraining support. For each method and factor axis, we split categories by whether they appear in that method’s documented prior, and compare the method to a fixed graphical-search reference (CDIS, GIES, IGSP, and PC) on the same held-out categories: OOD Rm,a =

OOD OOD Fm,a − F̄ref,a OOD F̄ref,a

.

(8)

The figure reports 100ROOD as a percent of the reference F1 (e.g. +60% means 60% above the reference mean), not an additive F1 increment. We use this reference group because these four methods form a coherent graphical-search family among the evaluated baselines. They still rely on standard causal-discovery assumptions, but they are comparatively less tied to a specific parametric mechanism family, neural pretraining simulator, or fixed variable ordering. Their average is a useful proxy for how hard a category is under common non-pretrained graph-search procedures. All four axes use the same scenario-level F1 rows as the main benchmark. For graph family, each scenario has one category; for mechanism, noise, and root-source axes, a scenario contributes to every factor category present in its SCM, and category scores are macro-averaged before forming ID and OOD summaries. The overall OOD score is OOD the unweighted mean of Rm,a over these four factor axes. A positive value means that the method stays above the four-method graphical-search reference on held-out categories, and larger values mean a larger relative lead. We also report an ID–OOD gap: the method’s relative gain on supported categories minus its relative gain on held-out categories. A larger positive gap means weaker generalization under this split, because the lead over the reference is more concentrated on documented support. Smaller gaps mean more stable generalization: a gap of zero, or a small negative gap, means the OOD relative gain is not below the ID relative gain. Taking the union of all pretrained methods’ documented priors would leave too few held-out categories for a stable factor-level analysis, and that union would keep shrinking as new pretrained models appear. A per-method split asks a narrower question: given what a method documents as pretraining support, does its relative advantage remain outside that support? The prior labels are audited from the corresponding papers or code. For the plot, source-level priors are mapped conservatively to CausalArena’s factor tags and counted only when they have a clear counterpart in the current synthetic 19

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

pool. Figure 14 visualizes this mapping as a documented prior-coverage profile over the four axes, with covered/total counts beside the radar. The mean column is the unweighted average of the four axis ratios, matching the four-axis averaging used for the OOD scores; we do not use radar-polygon area, which would mix axis order with the square-root radius. The full audited category mapping is released with the benchmark artifacts. Documented prior coverage Graph (10 cats)

Method

Graph Mech. Noise Root Mean

Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

2/10 4/10 2/10 4/10 4/10 2/10 5/10

2/16 2/16 3/16 5/16 13/16 3/16 8/16

3/15 4/15 1/15 7/15 6/15 1/15 8/15

3/14 1/14 1/14 5/14 5/14 1/14 4/14

18% 22% 13% 38% 49% 13% 45%

Root (14 cats)

Mechanism (16 cats)

10% 25% 50% 80%

Noise (15 cats) FoundCause TabCausal

CauScale SEA

CDFM Arrow

AVICI

Figure 14: Documented prior coverage for pretrained methods. Source-level priors are mapped to CausalArena graph, mechanism, noise, and root-family tags. The radar shows axis-wise coverage ratios on a square-root radial scale. Table rows are alphabetical. Cells are covered/total tags on that axis; Mean is the unweighted average of the four ratios, reported as a percentage. Bold is the highest value in a column and underline is second; ties share the mark. CauScale and SEA have the same mapped coverage, so their radar markers overlap. Figure 14 shows that the pretrained methods differ in documented support: FoundCause has the broadest mapped mechanism coverage (13/16), TabCausal is more balanced across graph, mechanism, and noise axes (5/10, 8/16, and 8/15), CDFM covers a relatively broad noise/root set, and AVICI has reasonable graph-family coverage but narrow mechanism and root coverage. CauScale and SEA share the same mapped counts. Those coverage differences are why the check is defined per method. Figure 15 then asks whether a method’s lead survives on categories it does not document as pretraining support. The top row is easiest to read as a location relative to the reference tick. FoundCause’s ID and OOD bars both sit well to the right of the ticks on every axis, with OOD relative gains near +60%. TabCausal remains above the ticks with smaller but still positive gains. CauScale is modestly above the reference, most clearly on mechanism and noise. CDFM is close to the ticks; SEA straddles them. Arrow and AVICI fall to the left of the ticks on several axes, so their OOD relative gains are negative. Holding out undocumented categories therefore changes how we read the scores more than it changes the coarse ranking: FoundCause and TabCausal stay first and second, but the bars show whether that order is confined to documented support. The bottom-left panel summarizes the same OOD comparison: it is the unweighted mean of ROOD over the four axes, i.e., how far each method sits above the four-baseline mean on held-out categories. FoundCause remains far above that reference (about +60%). TabCausal is next (+22%), with positive gains on all four axes. CauScale is also positive overall (+10%), especially on mechanism and noise, while CDFM is only slightly positive (+1%) and SEA is near the reference (−2%). Arrow and AVICI have negative overall OOD gains (−17% and −25%). The middle row and the bottom-right panel measure generalization as a drop from ID to OOD in relative gain. A larger positive bar means weaker generalization: the method’s lead over the reference is more concentrated on documented support. Smaller values mean more stable generalization, and a gap of zero, or a small negative gap, means OOD is not worse than ID. Green is a positive ID–OOD gap. Under this reference group, 20 of 28 method–axis gaps are positive, and every overall mean gap is nonnegative or nearly zero. The bottom-right panel is the mean of those four middle-row gaps, so smaller values indicate more stable generalization. FoundCause is the exception that stays flat (overall gap +0%), so its large lead is not concentrated on documented categories. TabCausal’s advantage shrinks on held-out categories (overall gap +4%) but remains substantial. CauScale’s gap is small (+1%). AVICI has the largest overall gap (+6%), driven by graph (+9%) and especially held-out root-family categories (+10%), consistent with its narrow mapped root support (1/14). Arrow sits below the reference on most held-out axes; CDFM’s positive OOD gains are smaller than those of FoundCause, TabCausal, and CauScale despite broader noise/root coverage. Broad 20

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Pretraining-exposure OOD diagnostic: ID/OOD performance by documented factor exposure ID F1 OOD F1

Method

Graph

Mechanism

R OOD (%)

negative ID-OOD gap

Noise

R OOD (%)

Root

R OOD (%)

R OOD (%)

FoundCause

+59%

+62%

+60%

+59%

TabCausal

+25%

+24%

+20%

+18%

CauScale

+4%

+13%

+15%

+6%

SEA

-1%

-1%

-4%

-2%

CDFM

+4%

+2%

-3%

+2%

Arrow

-13%

-17%

-20%

-17%

AVICI

-21%

0.0

0.2

0.4

-23%

0.6

0.0

0.2

0.4

0.6

-28%

0.0

0.2

raw F1

Graph ID-OOD gap

Mechanism ID-OOD gap

Noise ID-OOD gap

-1%

-1%

−10

−10

ID gain − OOD gain (%)

-1%

+5%

0

10

−10

ID gain − OOD gain (%)

0

+10%

10

−10

ID gain − OOD gain (%)

Overall OOD relative gain

ID gain − OOD gain (%)

+4%

+22% +1%

+10%

CDFM

+4%

+1%

SEA

+3%

-2%

Arrow

10

+0%

+60%

CauScale

0

Overall ID-OOD relative-advantage gap

FoundCause TabCausal

+1%

+4% +1%

10

+1% +6%

-1%

0

+5%

+1%

+9%

AVICI

+1%

+1%

-0%

+3%

Arrow

0.6

Root ID-OOD gap

-2%

-0%

+8%

0.4

+6%

-1% +10%

SEA CDFM

0.2

raw F1

+1%

+6%

0.0

+2%

+3%

CauScale

-28%

0.6

raw F1

TabCausal

AVICI

0.4

raw F1

FoundCause

Method

reference mean F1 for same split positive ID-OOD gap

+1%

-17%

+6%

-25%

−20

0

20

40

60

−1

mean relative F1 gain over reference (100R, %)

0

1

2

3

4

5

6

7

mean ID-OOD gap (%)

Figure 15: Observation-only pretraining-exposure OOD diagnostics. For each pretrained method we split synthetic obs-only categories into documented support (ID) and held-out categories (OOD), using that method’s own audited prior. Columns are the four factor axes. Top row: raw F1 on ID (teal) and OOD (pink). The dark tick on each bar is the graphical-search reference mean (CDIS, GIES, IGSP, PC) on the same ID or OOD slice; a bar to the right of its tick is better than the reference. The number at the right of each OOD bar is the relative gain ROOD from Equation (8), shown as a percent of the reference F1. Middle row: ID relative gain minus OOD relative gain, in the same percent units. A larger positive value means a larger drop from ID to OOD in that relative lead, so generalization is weaker; smaller values mean more stable generalization, and a gap of zero or a small negative value means OOD is not worse than ID. Green is a positive gap (advantage shrinks on OOD); red is a negative gap (advantage grows on OOD). Bottom row: unweighted means over the four axes. Left: overall ROOD , how far the method sits above the four-baseline mean on held-out categories, again as a percent. Right: the mean of the middle-row gaps, how much larger the ID relative gain is than the OOD relative gain; smaller values indicate more stable generalization. documented pretraining support helps, but architecture, how coverage is balanced across factors, and how the model uses nonparametric graph signals still matter. Those four-axis means also let us back out a crude full-ID / full-OOD F1 for the same average that Figure 9 reports. Let F be that average F1, c the unweighted mean documented coverage over the four axes, and δ=

1 + R̄ID −1 1 + R̄OOD 21

(9)

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

the relative gap between the mean ID and mean OOD leads in Figure 15. Treating the reported score as a coverage mix, F = c F̂ ID + (1 − c) F̂ OOD ,

F̂ ID = (1 + δ) F̂ OOD ,

(10)

gives F̂ OOD = F/(1 + cδ) and F̂ ID = F (1 + δ)/(1 + cδ). This is only an estimate: c and δ are four-axis means, and F is the family–split average already used in Figure 9.

0.6

FoundCause

Optimal TabCausal

Average F1

0.5

Control Graphical search Functional assumptions Continuous optimization

full-ID estimate full-OOD estimate pretrained ID OOD span

GIES

CDFM

0.4 AVICI

0.3

Arrow

CDIS

SEA IGSP

CauScale PC RandReg

0.2

DAGMA

0.1

SDCD

1

10 Seconds per graph (log scale)

NT-MLP

LiNGAMNOTEARS

DAS

100

Figure 16: Accuracy–cost view with full-ID / full-OOD estimates. Same axes as Figure 9: average F1 against comparable per-graph time, with pretrained x equal to Tgraph . There is no Pareto curve. Each pretrained method is a vertical pair: the teal diamond is F̂ ID and the pink diamond is F̂ OOD from Equation (10); the observed F lies on that segment. Circles are non-pretrained methods at their reported F1. Optimal points toward higher F1 and lower time. Figure 16 plots the same ranking as a vertical F1 interval. The segments are short. FoundCause stays near 0.59 at either endpoint. TabCausal moves from about 0.48 (full OOD) to 0.50 (full ID) around its reported 0.49. CDFM is 0.42–0.43; AVICI has the longest relative span (0.32–0.35), consistent with its larger ID–OOD gap and narrower mapped support. CauScale, SEA, and Arrow barely move. The coarse pretrained order is therefore not an artifact of documented coverage: replacing the observed mix with either a fully ID or a fully OOD reading does not reorder FoundCause, TabCausal, and the cheaper lower-F1 group. 6.3

Semantic and Formula Benchmark Behavior Analysis

Semantic analysis protocol. The previous subsections use anonymous synthetic factors. Semantic and Formula SCMs let us look at operational processes and scientific formula families. We report self-centered F1 shifts: a method’s F1 on a domain minus its own family average. Positive values mark domains where the method scores higher than its own family average; negative values mark domains where it scores lower. The plots show which semantic or scientific settings each method handles better or worse. Example scenarios. Figure 3 and the examples in Sections 4.2 and 4.3 introduce two released SCMs: a wildfire smoke / shelter process and a barometric pressure reduction to sea level. We reuse those scenarios here when interpreting domain-level Semantic and Formula behavior: a missing or extra edge is then a claim about that smoke day (e.g., treating a consumer indoor reading as true indoor PM2.5, or dropping filter provision as a parent of indoor concentration) or a claim about the pressure reduction (e.g., writing the sea-level formula without fused height, or treating a thermometer reading as if it were already layer-mean virtual temperature). Released Semantic and Formula SCMs, including these two, pass several review and validation passes before they leave the construction pipeline (Section A.3). We focus on observation-only behavior, restricting the comparison to methods evaluated on the same split. In Figure 17, the heatmaps show method-domain self-centered F1 shifts, and the side bars report each method’s standard deviation across domains. A larger standard deviation means that the method is more domain-sensitive after subtracting its own average performance. 22

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Domain-level patterns. After each method’s own average is removed, some domains still sit consistently above or below that baseline. On Semantic, cybersecurity and IT operations is usually weak: many methods score lower there than on their own Semantic average. Education, healthcare delivery, manufacturing, and water and sanitation are the usual strong ones, though not for every method. The side bars repeat this by method. DAS, LiNGAM, and DAGMA hardly move across Semantic domains; PC, CDIS, AVICI, and TabCausal move more. On Formula, energy systems and thermodynamics tend to be easier. Astronomy, electromagnetism, and optics and waves more often invert a law, pass through a geometric identity, or put a sensor chain around the equation, and that is where several methods lose relative F1. Observation-only semantic and formula domain specialization

+0.06

+0.00

-0.03

-0.03

+0.02

-0.03

-0.01

-0.01

CDIS +0.07

-0.06

+0.06

-0.03

+0.04

-0.01

+0.07

-0.06

+0.03

-0.11

GIES +0.06

+0.04

+0.02

-0.07

-0.05

-0.03

+0.10

-0.00

+0.01

-0.08

0.03 0.06

-0.02

+0.02

+0.01

-0.03

-0.02

-0.01

-0.02

+0.13

-0.07

-0.00

-0.02

+0.01

+0.06

-0.04

-0.01

-0.01

Std

+0.06

-0.01

+0.01

-0.02

-0.01

-0.03

-0.03

+0.11

-0.06

-0.08

+0.07

-0.02

+0.06

-0.04

0.03 0.07

0.06

-0.01

+0.01

+0.07

-0.02

-0.07

-0.03

+0.05

-0.03

+0.02

+0.00

PC +0.08

-0.06

+0.04

-0.04

-0.04

+0.01

+0.10

-0.06

+0.05

-0.09

DAS +0.04

-0.01

+0.03

-0.03

-0.01

+0.01

+0.03

-0.01

-0.03

-0.01

LiNGAM +0.05

+0.01

-0.02

-0.03

-0.01

+0.01

+0.01

+0.00

+0.01

-0.02

IGSP

Formula domains: self-centered F1

Std

+0.00

-0.03

+0.03

+0.03

-0.09

-0.03

-0.03

+0.12

-0.05

+0.10

-0.05

+0.01

+0.10

-0.02

+0.01

-0.02

-0.01

-0.05

-0.01

+0.05

-0.05

-0.02

+0.00

+0.03

+0.01

+0.01

+0.01

-0.04

+0.01

-0.02

-0.00

0.02

+0.00

+0.03

-0.03

+0.01

+0.02

-0.01

-0.01

+0.00

+0.02

-0.04

0.04 0.06 0.02

0.07 0.05 0.05 0.02 0.02

DAGMA +0.03

+0.01

+0.00

-0.01

+0.02

+0.01

+0.02

-0.02

-0.02

-0.04

0.02

+0.01

-0.01

-0.01

-0.00

+0.02

+0.00

-0.01

-0.02

-0.00

+0.01

NOTEARS +0.06

+0.01

+0.01

-0.03

-0.03

+0.02

+0.00

-0.01

+0.02

-0.04

0.03

+0.03

-0.01

+0.02

-0.01

+0.02

-0.02

+0.01

-0.03

+0.01

-0.02

0.02

NOTEARS-MLP +0.02

+0.00

+0.02

+0.05

+0.02

-0.01

+0.04

-0.02

-0.05

-0.06

0.04

-0.01

-0.02

+0.02

-0.01

+0.01

-0.01

-0.03

+0.03

-0.00

+0.01

0.02

0.04

+0.00

-0.00

+0.01

-0.03

+0.01

-0.04

+0.07

-0.02

+0.04

-0.04

-0.00

+0.02

-0.01

+0.02

-0.01

+0.00

-0.02

-0.01

+0.03

-0.03

+0.04

+0.01

-0.10

-0.02

-0.01

-0.10

+0.06

+0.01

+0.10

+0.01

SDCD +0.05

-0.01

+0.04

+0.00

-0.02

-0.01

+0.04

-0.03

+0.00

-0.07

Arrow +0.08

+0.02

-0.00

-0.04

-0.05

+0.01

+0.04

-0.02

+0.01

-0.04

AVICI +0.05

-0.02

+0.00

-0.05

-0.09

-0.01

+0.07

-0.01

+0.13

-0.08

CDFM +0.01

-0.01

+0.01

+0.01

+0.01

-0.01

-0.00

+0.02

-0.00

-0.02

FoundCause +0.05

+0.02

-0.01

-0.05

-0.01

-0.07

+0.06

-0.03

+0.06

-0.02

SEA +0.11

+0.01

+0.00

-0.06

-0.02

-0.02

+0.08

+0.01

-0.02

-0.09

TabCausal +0.01

+0.02

+0.04

-0.07

-0.04

-0.04

+0.10

-0.01

+0.10

-0.10

0.04 0.07 0.01

-0.01

-0.00

+0.02

-0.01

-0.02

+0.01

+0.02

+0.00

-0.00

-0.01

-0.05

+0.01

-0.02

-0.03

-0.03

-0.00

+0.10

-0.03

+0.09

-0.05

+0.01

+0.03

-0.04

-0.03

-0.01

-0.04

+0.03

+0.02

+0.04

-0.01

+0.01

+0.09

-0.08

-0.01

-0.01

-0.09

+0.08

-0.07

+0.10

-0.02

0.01

re

a hc

alt

y n n ic er nt ce ng ng rit bl ba t tio at an it me es usi lty uri Pu alth uca Ur por W ems ecu d IT Fin cred ern rvic Ho rea fact s st rs an he Ed d Gov se an d anu sy ybe r n n t a a M C

0.03

0.06

0.07

0.05

−0.05

0.01 0.05

0.06

0.00

0.00

0.02

0.05

He

0.10

0.05

F1 shift from method average

Semantic domains: self-centered F1 RandomRegress +0.02

oics ro try an ds rm cs ct cs is ch flui The ami Ele neti hem C n Me nd g y a d a m

s y y ls gy rth rg s tic ria s om Ea ms Op ves colo te re Ene em ron e e t t Ma uctu st wa ioys As d s r B t s an d an

0.03

−0.10 0.07

0.00

0.05

sy

Figure 17: Observation-only Semantic and Formula domain specialization. Method-level summaries. A domain heatmap only locates the gain or drop. It does not describe the processes or equations behind that domain. To read the Semantic and Formula benchmarks at that grain, we used an LLM agent to summarize each method’s characteristic strengths and weaknesses from the observation-only results together with the scenario descriptions. Table 6 reports those summaries: how the method tends to behave on Semantic scenarios, how it tends to behave on Formula scenarios, and a short overall sketch. Table 6: Method-level behavior on the Semantic and Formula benchmarks. An LLM agent summarized each method’s characteristic strengths and weaknesses from the observation-only scores and the scenario descriptions. The table reports those qualitative directions; it does not list raw scores. Method

Semantic benchmark direction

Formula benchmark direction

RandomRegress

Appears strong only when the variable order accidentally matches a simple operational sequence; collapses when the story requires distinguishing multiple paths, coordination, or behavioral response.

Shows the same order artifact on formula A sanity control, not a semantic reasoner; cases: it can match local support by chance, its large swings expose ordering and sparsebut fails on coupled equations, conservation regression artifacts. relations, and measurement chains.

CDIS

Favors locally clear signal-to-decision struc- Handles small directed formula chains bet- Works best when the causal story is locally tures, where an observed input flows into ter than coupled multiplicative systems or separable; hidden demand, shared drivers, a screening or triage step and then to an shared-parameter equations. and cascading structure blur the conditional outcome; weakens when demand, policy resignals it relies on. sponse, backlog, or containment propagates through the system.

GIES

Prefers stable, sparse, stage-like workflows More reliable on static physical con- Score/search behavior is strongest for stable in which monitoring signals lead to inter- straints than on readout chains, ex- staged mechanisms and less reliable when ventions or outcomes; degrades under accu- ponential/periodic transformations, or the observed variables are indirect traces of mulated pressure, capacity limits, and cross- observation-heavy formula contexts. a latent process. module propagation.

IGSP

Performs well on observable service work- Handles some direct multiplicative or Benefits from explicit workflows such as processing, maintenance, vac- power-law chains, but struggles with ther- flow/intervention alignment; sparse cination, or patching; weakens when rare modynamic, absorption, and exchange- event evidence and scientific readout layers positives, hidden eligibility, or surveillance process relations. make its direction tests less decisive. logic determine the graph.

23

Overall profile

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Formula benchmark direction

A P REPRINT

Method

Semantic benchmark direction

PC

Recovers local screening and inspection cas- Handles low-coupling formula systems bet- Conditional-independence search is useful cades better than adaptive, behavioral, or ter than equations with shared parameters, when the semantic story decomposes into policy-response processes. multiple outputs, or derived readouts. local tests, but brittle under adaptation, common drivers, and indirect measurements.

DAS

Shows mild advantages on anomaly- Works better on direct physical-quantity Has no sharp semantic specialization; it is monitoring and escalation chains, but loses chains than on biological, load, or pump- mainly helped when the ordering signal is advantage in outreach, engagement, and like coupled systems. simple and directly observable. priority-queue settings.

LiNGAM

Favors continuous, one-way trigger-to- More plausible on monotone scaling rela- Its linear non-Gaussian orientation prior can action stories; weakens when capacity, ad- tions and weaker on nonlinear response, help on direct driving chains, but does not herence, follow-up management, or satura- buffering, oscillation, or multi-input scien- match capacity-limited or strongly nonlintion controls the outcome. tific systems. ear semantic processes.

DAGMA

Works comparatively well when a bottle- Fits smooth, compressible formula skele- Continuous optimization favors centralized, neck or allocation lever explains a measur- tons better than load, saturation, or shared- smooth causal structure; diffuse coordiable downstream outcome; weakens on dis- mechanism systems. nation and non-smooth process logic are tributed monitoring, tracing, and coordinaharder to represent. tion chains.

NOTEARS

Favors direct risk-or-demand-to-action sto- Can recover some near-monotone physical Linear DAG optimization is most useful ries; degrades when the system involves chains, but loses reliability on complex mul- when a semantic process can be compressed propagation, containment, or multi-stage tiplicative, thermal, and interference rela- into a direct risk-to-action edge pattern. response. tions.

NOTEARS-MLP

The nonlinear SEM helps on triage and resource-allocation chains, but not enough for saturation, alert fatigue, overflow, or backlog dynamics.

SDCD

Performs better when an early warning sig- Is relatively favorable on scale-law or Most effective for monotone, observable nal leads to a resource intervention, and power-law relations, and weaker on mea- intervention chains; less robust when a hidworse when capacity constraints or long- surement readouts and geometric observa- den state or measurement layer mediates run accumulation modulate the effect. tion chains. the causal relation.

Arrow

Favors event-response stories: an incident is Handles some engineering and geometric Better aligned with event-response semandetected and then remediated or mitigated. relations, but weakens on biochemical re- tics than with strategic allocation or delayed It is weaker on planning, scheduling, and sponse curves and buffered systems. planning. resource-prioritization tasks.

AVICI

Strongest when a concrete trigger or in- Covers several common physical-law struc- Its pretrained prior helps with local triggertervention changes a downstream met- tures, but its performance varies noticeably to-outcome chains, while multi-actor orgaric; weaker in administrative coordination, across formula types and readout contexts. nization and instrumented scientific conworkload, and prioritization cascades. texts remain harder.

CauScale

Works best on short, separable signal-to- Very strong when calibrated physical inputs Especially good at compact causal chains; action stories; weakens on drift, spillover, directly determine a law output, but much system-level equilibrium, propagation, and allocation competition, and equilibrium- weaker once control, calibration, or instru- control layers are the main failure mode. like processes. ment loops wrap the formula.

CDFM

Does not show a stable semantic specializa- The formula readout is similarly flat, with The residual evidence does not support a tion in this release; strong and weak residual no reliable direction that separates favorable clear method-specific semantic profile, so cases are close after centering. and unfavorable scientific contexts. we avoid over-interpreting it.

FoundCause

Favors direct readiness, allocation, and re- Handles some clearly structured formula Structural refinement appears helpful for sponse chains; weakens on drift, spillover, chains, but struggles when shared physical direct response chains, but less sufficient cascade, and feedback-like processes. parameters, vessels, instruments, or heat- for long-range propagation and sharedexchange contexts mediate the relation. mechanism systems.

SEA

Favors short detection-to-intervention Works better on classic single-law physics Stronger on short protocol-like causal stostories and quality/adherence-to-outcome than on field-strength, wave/readout, or ries; slower accumulation and cross-system chains; weakens under retention, allocation, complex multiplicative relations. propagation remain difficult. cascade, and spillover dynamics.

TabCausal

Favors clear policy or operational levers Strong on direct formula pipelines from whose effects flow to downstream out- physical inputs to derived quantities, but comes; weakens on load propagation, back- weaker when instrumentation, regime log, and cascade-heavy systems. switches, or measurement chains surround the base law.

Captures closed-form nonlinear formulas better than linear NOTEARS, yet remains weak on control policies, threshold gates, and process feedback.

Overall profile

Extra function flexibility helps, but systemlevel semantics such as accumulation, saturation, and feedback remain the limiting factor.

Its strongest semantic profile is clear intervention point plus directed downstream chain; multi-stage propagation and instrumented readout chains remain the main weaknesses.

Operational-semantic failure modes. The following two paragraphs paraphrase Table 6; they are LLM-assisted qualitative readings, not a separate human coding of errors. On Semantic cases, the summaries look better on a short visible chain and worse once backlog, drift, or a cascade takes over. RandomRegress looks strong only when the column order happens to match a simple sequence; it fails when the case needs several paths, coordination, or a behavioral response. PC, CDIS, IGSP, and GIES do better when a visible input goes into screening, inspection, or a staged workflow, and worse when demand, backlog, capacity, or an adaptive response moves through the rest of the system. IGSP also drops when rare events or hidden eligibility set the graph. DAS has only a mild edge on anomaly monitoring and escalation, and loses it on outreach and queues. LiNGAM prefers a one-way trigger-to-action chain, and drops when capacity, adherence, or saturation sets the outcome. DAGMA, NOTEARS, NOTEARS-MLP, and 24

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

SDCD like one bottleneck, risk signal, triage step, or early warning that moves a measurable next action; they struggle with coordination, overflow, backlog, and long-run accumulation. Arrow fits detect-then-fix better than planning or scheduling. AVICI, CauScale, FoundCause, SEA, and TabCausal likewise prefer a short visible action that moves a downstream outcome, and they weaken on drift, spillover, workload, competing allocation, or cascade. CDFM is too flat in this release to assign a Semantic profile. Scientific-formula failure modes. On Formula cases, they look better on a short equation chain, and worse once extra instruments sit around that equation. RandomRegress can match a few nearby variables by chance, then fail on coupled equations, conservation relations, and measurement chains. PC, CDIS, IGSP, and GIES handle small directed formulas better than shared parameters, extra outputs, or readout chains; IGSP is also weak on thermodynamic, absorption, and exchange relations. DAS prefers a direct physical-quantity chain over a biological, load, or pump-like coupled system. LiNGAM is more plausible on monotone scaling than on nonlinear response, buffering, oscillation, or several inputs. DAGMA and NOTEARS can fit a smooth or near-monotone formula, but not load, saturation, thermal, or interference structure. NOTEARS-MLP helps on a closed-form nonlinear formula, but not on control policies, threshold gates, or process feedback. SDCD likes power-law relations and not measurement or geometric readouts. CauScale is strong when calibrated physical inputs directly set a law output, and much weaker once control, calibration, or instrument loops wrap the law. FoundCause and TabCausal like a clear input-to-derived-quantity chain, and weaken when vessels, instruments, heat exchange, regime switches, or measurement chains sit around the law. Arrow handles some engineering and geometric relations, not biochemical curves or buffered systems. SEA prefers a classic single-law physics case over field, wave, or complex product relations. AVICI covers several common physical-law forms, but the score still moves with the formula type and the readout. CDFM again has no reliable Formula direction.

7

Discussion

7.1

Open and Evolvable Evaluation

The most direct validation of a causal discovery claim is often feedback in a real system: whether intervening on a recovered parent changes the intended downstream quantity. During method development, however, such closed-loop checks are usually unavailable at scale, so the field relies on benchmarks with known graphs. CausalArena is designed for that development setting. Using executable SCMs rather than only fixed sampled tables makes graph structure, mechanisms, interventions, and sampling procedures inspectable and reproducible, and it lets later releases add new SCMs while preserving the interfaces in Section 3. For CDFMs, openness also changes how a score should be read. Once an SCM, its graph, or its generator becomes public, related environments may enter later pretraining. This does not make a public benchmark useless, but it does mean that a result should be interpreted together with the model’s training context: performance on a familiar generator distribution and performance on newly introduced environments answer different questions. We therefore treat CausalArena as an evolvable arena. Versioned public releases provide reproducible reference evaluations, while the common SCM interface supports newly authored graphs, mechanisms, and scenarios as pretrained models evolve. This is how the freshness requirement in Section 3.2 is pursued without giving up transparency. 7.2

Interpreting Foundation-Model Results

To make pretrained comparisons easier to read, CausalArena ties results to explicit benchmark versions and release states. Leaderboard entries are encouraged to disclose whether public CausalArena SCMs, generated samples, scenario metadata, or closely related generators were used during development or pretraining. The current package distinguishes public and reserved subsets: 500 synthetic, 50 semantic, and 50 formula-grounded SCMs are public, while the remaining audited SCMs are reserved for held-out leaderboard evaluation and progressive release in later benchmark versions. Results in this paper use the complete 1,200-SCM pool and are labeled by this benchmark version. These practices make pretraining–evaluation overlap easier to discuss; they are not a claim that overlap can be removed once and for all. As earlier versions become public, their SCMs may eventually enter future training corpora. Continued evaluation of new foundation models therefore benefits from periodically extending the SCM pool and reporting scores together with the benchmark version and training disclosure. The semantic construction pipeline is one practical route for adding newly authored operational environments in later releases. 7.3

Scope and Future Extensions

The current release focuses on multivariate tabular causal structure recovery with known directed graph ground truth. Synthetic SCMs emphasize controlled breadth, semantic SCMs emphasize operational grounding, and formula-grounded 25

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

SCMs emphasize explicit scientific mechanisms. Public real-world datasets serve as an external-validity check; their published graphs should be read as reference structures, with the caveat that they are not uniquely verified causal ground truth. The same executable interface can later support additional directions: new domains and SCM families, larger graphs, latent confounding, and neighboring causal tasks such as intervention-target reasoning, effect estimation, and counterfactual inference. Such extensions can reuse the current versioning and evaluation infrastructure while remaining comparable with earlier releases.

8

Conclusion

We introduced CausalArena, an evolvable benchmark for causal discovery in the foundation-model era. CausalArena combines broad synthetic SCMs, semantically grounded operational SCMs, and formula-grounded scientific SCMs under a shared observational and interventional evaluation protocol. Together, these families support broad coverage, meaningful grounding, detailed diagnosis, and future extension with newly constructed causal environments. Experiments across classical, neural, and pretrained causal discovery methods show that rankings vary substantially across SCM families, mechanisms, sample sizes, and intervention settings, and that strong performance in one evaluation regime does not necessarily transfer to others. For pretrained models, these results further highlight the importance of interpreting benchmark scores in light of possible pretraining–evaluation overlap. By coupling reproducible public reference suites with a versioned and extensible SCM interface, CausalArena provides a foundation for continued evaluation as causal discovery methods and foundation models evolve.

References [1] Peter Spirtes, Clark N. Glymour, and Richard Scheines. Causation, Prediction, and Search. MIT Press, 2 edition, 2000. [2] Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. Elements of Causal Inference: Foundations and Learning Algorithms. MIT Press, 2017. [3] Alexander Reisach, Christof Seiler, and Sebastian Weichwald. Beware of the simulated DAG! Causal discovery benchmarks may be easy to game. In NeurIPS, 2021. [4] Lars Lorch, Scott Sussex, Jonas Rothfuss, Andreas Krause, and Bernhard Schölkopf. Amortized inference for causal structure learning. In NeurIPS, 2022. [5] Menghua Wu, Yujia Bao, Regina Barzilay, and Tommi Jaakkola. Sample, estimate, aggregate: A recipe for causal discovery foundation models. Transactions on Machine Learning Research, 2025. [6] Ryan Thompson, He Zhao, Daniel M. Steinberg, and Edwin V. Bonilla. Arrow: A foundation model for causal discovery, 2026. arXiv preprint arXiv:2605.07204. [7] Bo Peng, Sirui Chen, Jiaguo Tian, Yu Qiao, and Chaochao Lu. CauScale: Neural causal discovery at scale. In ICML, 2026. [8] Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, and Peng Cui. CDFM: Towards a general-purpose causal discovery foundation model, 2026. arXiv preprint arXiv:2607.11508. [9] Patrick Blöbaum, Krishnakumar Balasubramanian, and Shiva Prasad Kasiviswanathan. FoundCause: Causal discovery with latent confounders from observational data, 2026. arXiv preprint arXiv:2606.17516. [10] Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, and Han-Jia Ye. TabCausal: Pretraining across causal environments for tabular causal discovery, 2026. arXiv preprint arXiv:2605.31156. [11] Felix L. Rios, Giusi Moffa, and Jack Kuipers. Benchpress: A versatile platform for structure learning in causal and probabilistic graphical models. Journal of Statistical Software, 114(12):1–43, 2025. [12] Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas Chesneau, Marc Schoenauer, and Özgür Şimşek. CausalProfiler: Generating synthetic benchmarks for rigorous and transparent evaluation of causal machine learning, 2025. arXiv preprint arXiv:2511.22842. [13] Karen Sachs, Omar Perez, Dana Pe’er, Douglas A. Lauffenburger, and Garry P. Nolan. Causal protein-signaling networks derived from multiparameter single-cell data. Science, 308(5721):523–529, 2005. [14] Benjamin Herdeanu, Juan Nathaniel, Carla Roesch, Jatan Buch, Gregor Ramien, Johannes Haux, and Pierre Gentine. CausalDynamics: A large-scale benchmark for structural discovery of dynamical causal models. In NeurIPS, 2025. 26

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

[15] Mathieu Chevalley, Yusuf H. Roohani, Arash Mehrjou, Jure Leskovec, and Patrick Schwab. A large-scale benchmark for network inference from single-cell perturbation data. Communications Biology, 8(1):412, 2025. [16] Juan L. Gamella, Jonas Peters, and Peter Bühlmann. Causal chambers as a real-world physical testbed for AI methodology. Nature Machine Intelligence, 7:107–118, 2025. [17] Markus Kalisch and Peter Bühlmann. Estimating high-dimensional directed acyclic graphs with the PC-algorithm. Journal of Machine Learning Research, 8(22):613–636, 2007. [18] Diego Colombo and Marloes H. Maathuis. Order-independent constraint-based causal structure learning. Journal of Machine Learning Research, 15(1):3741–3782, 2014. [19] Alain Hauser and Peter Bühlmann. Characterization and greedy learning of interventional Markov equivalence classes of directed acyclic graphs. Journal of Machine Learning Research, 13:2409–2464, 2012. [20] Yuhao Wang, Liam Solus, Karren D. Yang, and Caroline Uhler. Permutation-based causal inference algorithms with interventions. In NeurIPS, 2017. [21] Haoyue Dai, Ignavier Ng, Jianle Sun, Zeyu Tang, Gongxu Luo, Xinshuai Dong, Peter Spirtes, and Kun Zhang. When selection meets intervention: Additional complexities in causal discovery. In ICLR, 2025. [22] Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, and Antti Kerminen. A linear non-gaussian acyclic model for causal discovery. Journal of Machine Learning Research, 7(72):2003–2030, 2006. [23] Francesco Montagna, Nicoletta Noceti, Lorenzo Rosasco, Kun Zhang, and Francesco Locatello. Scalable causal discovery with score matching. In CLeaR, 2023. [24] Peter Bühlmann, Jonas Peters, and Jan Ernest. CAM: Causal additive models, high-dimensional order search and penalized regression. The Annals of Statistics, 42(6):2526–2556, 2014. [25] Patrik O. Hoyer, Dominik Janzing, Joris M. Mooij, Jonas Peters, and Bernhard Schölkopf. Nonlinear causal discovery with additive noise models. In NeurIPS, 2008. [26] Xun Zheng, Bryon Aragam, Pradeep Ravikumar, and Eric P. Xing. DAGs with NO TEARS: Continuous optimization for structure learning. In NeurIPS, 2018. [27] Xun Zheng, Chen Dan, Bryon Aragam, Pradeep Ravikumar, and Eric P. Xing. Learning sparse nonparametric DAGs. In AISTATS, 2020. [28] Kevin Bello, Bryon Aragam, and Pradeep Ravikumar. DAGMA: Learning DAGs via m-matrices and a logdeterminant acyclicity characterization. In NeurIPS, 2022. [29] Achille Nazaret, Justin Hong, Elham Azizi, and David Blei. Stable differentiable causal discovery. In ICML, pages 37413–37445, 2024. [30] Philippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien, and Alexandre Drouin. Differentiable causal discovery from interventional data. In NeurIPS, 2020. [31] Muralikrishnna G. Sethuraman, Romain Lopez, Rahul Mohan, Faramarz Fekri, Tommaso Biancalani, and JanChristian Hütter. NODAGS-Flow: Nonlinear cyclic causal structure learning. In AISTATS, pages 6371–6387, 2023. [32] Tomas Geffner, Javier Antoran, Adam Foster, Wenbo Gong, Chao Ma, Emre Kiciman, Amit Sharma, Angus Lamb, Martin Kukla, Nick Pawlowski, Miltiadis Allamanis, and Cheng Zhang. Deep end-to-end causal inference, 2022. arXiv preprint arXiv:2202.02195. [33] Rebecca J. Herman, Jonas Wahl, Urmi Ninad, and Jakob Runge. Unitless unrestricted Markov-consistent SCM generation: Better benchmark datasets for causal discovery. In CLeaR, pages 1506–1531, 2025. [34] Marco Scutari. Learning Bayesian networks with the bnlearn R package. Journal of Statistical Software, 35(3): 1–22, 2010. [35] Joris M. Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, and Bernhard Schölkopf. Distinguishing cause from effect using observational data: Methods and benchmarks. Journal of Machine Learning Research, 17 (32):1–102, 2016. [36] Wei Zhou, Hong Huang, Guowen Zhang, Ruize Shi, Kehan Yin, Yuanyuan Lin, and Bang Liu. OCDB: Revisiting causal discovery with a comprehensive benchmark and evaluation framework, 2024. arXiv preprint arXiv:2406.04598. [37] Yuxiao Cheng, Ziqian Wang, Tingxiong Xiao, Qin Zhong, and Jinliang He. CausalTime: Realistically generated time-series for benchmarking of causal discovery. In ICLR, 2024.

27

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

[38] Gideon Stein, Maha Shadaydeh, Jan Blunk, Niklas Penzel, and Joachim Denzler. CausalRivers - scaling up benchmarking of causal discovery for real-world time-series. In ICLR, 2025. [39] Konstantin Göbler, Tobias Windisch, Mathias Drton, Tim Pychynski, Martin Roth, and Steffen Sonntag. causalAssembly: Generating realistic production data for benchmarking causal discovery. In CLeaR, pages 609–642, 2024. [40] Muhammad Hasan Ferdous, Emam Hossain, and Md Osman Gani. TimeGraph: Synthetic benchmark datasets for robust time-series causal discovery. In KDD, pages 5425–5435, 2025. [41] Nikolay Babakov, Ehud Reiter, and Alberto Bugarín. CausalGraphBench: a benchmark for evaluating language models capabilities of causal graph discovery. In ACL, pages 240–258, 2025. [42] Yinghuan Zhang, Yufei Zhang, Parisa Kordjamshidi, and Zijun Cui. Bayesian network structure discovery using large language models. Transactions on Machine Learning Research, 2026. Also available as arXiv:2511.00574. [43] Guangya Wan, Yunsheng Lu, Yuqi Wu, Mengxuan Hu, and Sheng Li. Large language models for causal discovery: Current landscape and future directions. In IJCAI, pages 10687–10695, 2025. [44] Boxiang Zhao, Shuliang Wang, Lianhua Chi, Qi Li, Xiaojia Liu, and Jing Geng. Causal discovery via causal star graphs. ACM Transactions on Knowledge Discovery from Data, 17(7):98:1–98:24, 2023. [45] Michaela Hardt, William Roy Orchard, Patrick Blöbaum, Elke Kirschbaum, and Shiva Kasiviswanathan. The PetShop dataset: Finding causes of performance issues across microservices. In CLeaR, pages 957–978, 2024.

Appendix This appendix provides benchmark data details, the evaluation protocol, and complete numerical result tables.

A

Benchmark Data Details

This section records the data construction details behind the current release. We use paper-facing names, and separate generator sampling rules from realized release statistics. A.1

Synthetic SCM Pool

The synthetic main benchmark contains 1,000 SCM configurations. It crosses 10 graph families with five dimensionalities d ∈ {10, 20, 30, 50, 100} and 20 SCMs per graph-family/dimension pair. Each configuration is generated with both an observation-only split and an observation-plus-intervention split; generated files enter the release only after passing the audit rules in Section A.2. Graph structure. • Erdos–Renyi: random ordered graphs with edge density sampled from 0.05–0.25. • Scale-free: preferential-attachment-style skeletons oriented by an acyclic order, with density 0.04–0.22. • Chain: sparse path-like DAGs with density 0.03–0.16 and long directed propagation. • Tree: branching sparse DAGs with density 0.03–0.16. • Layered: layer-respecting DAGs with density 0.04–0.20. • Bipartite layered: two-stage or multi-stage bipartite-like DAGs with density 0.04–0.20. • Small-world: locally clustered DAGs with short directed paths, density 0.04–0.20. • Block-sparse layered: block-structured layered DAGs with sparse cross-block edges, density 0.04–0.20. • Dense local-module: modular DAGs with denser within-module connectivity, density 0.06–0.24. • Collider-rich: DAGs that deliberately increase v-structure frequency, density 0.06–0.24. • Validity filters: candidate graphs are resampled until they are DAGs with no isolated variables, sufficient graph depth, and at least ⌈0.8d⌉ directed edges. The optional in-degree cap is sampled from no explicit cap, 4, or 5 for d ≤ 20, from 4, 5, or 6 for 20 < d ≤ 50, and from 5, 6, or 8 for d = 100.

28

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Family mixing. • Mechanism/noise family count: the number of active local mechanism and noise families increases with dimension: 2–4 at d = 10, 3–5 at d = 20, 4–6 at d = 30, 5–8 at d = 50, and 6–10 at d = 100. • Root-family count: the number of active root families depends on the number of root variables: one family for a single root, 1–2 for at most three roots, 2–4 for at most eight roots, and 3–6 otherwise. • Local assignment: after active families are selected, node-level mechanisms, root families, and noise families are assigned across the graph so that a single SCM can contain heterogeneous local structural equations and noise channels. Mechanisms. √ • Linear additive: parent weights are scaled by 0.8/ k and biases are sampled in [−0.5, 0.5]. • Polynomial additive: powers 2–4 are used, with quadratic coefficients in [−0.5, 0.5] and cubic coefficients in [−0.15, 0.15]. √ • Multiplicative parent product: parent effects are multiplied after weight scaling by 0.8/ k and bias sampling in [−0.4, 0.4]. • Random-function product: transformed parent responses are multiplied and rescaled with output scale in [0.4, 1.6]. Together with multiplicative parent product, this form is scored as one multiplicative family, so the analyses use 16 mechanism tags. √ • Additive interaction: up to three parent pairs are sampled, with interaction weights scaled by 0.5/ #pairs. • Rational/ratio: denominator-safe ratio transforms avoid near-zero denominators through bounded offsets and clipping. • Exponential/log: exponential inputs are clipped to [−3, 3] and paired with log or signed-log style transforms where needed. • Trigonometric/periodic: periodic parent transforms use sine/cosine-style responses with randomized frequency and phase. • Piecewise/threshold: cutoffs are sampled in [−0.6, 0.6], low slopes in [0.2, 1.0], high slopes in [1.0, 2.5], and offsets in [−1.5, 1.5]. • Saturation/clipping: saturating nonlinearities use response scales in [0.8, 2.5]. • Mixture/regime switching: gates use scales in [1, 4], cutoffs in [−0.5, 0.5], and alternate response scales in [0.5, 1.8]. • Random neural: one-hidden-layer neural mechanisms use 3–8 hidden units. • RFF/GP-style: random-feature mechanisms use 64–256 features, length scales in [0.25, 3.0], and output scales in [0.4, 2.5]. • Activation additive: activation families include ReLU, leaky-ReLU, softplus, sigmoid, absolute value, sign, step, hard-tanh, ReLU6, ELU, SELU, SiLU, modulo, round, and rank transforms; modulo periods are in [0.75, 3.0] and rounding steps in [0.2, 1.0]. • Tree-based: shallow tree mechanisms use depth 2–4 with thresholds in [−0.8, 0.8]. • Discretization/assignment: discretized mechanisms use 3–7 bins or clusters. • Aggregation nonlinearities: max or log-sum-exp aggregation uses temperatures in [0.3, 1.5]. Root distributions. • Uniform: lower bounds are sampled in [−2, 0] and upper bounds in [0.5, 3.0]. • Gaussian: the mean is drawn from N (0, 1) and the standard deviation from [0.5, 2.0]. • Truncated Gaussian: the mean is drawn from N (0, 0.82 ), the standard deviation from [0.4, 1.4], and samples are clipped to mean ±2.5 standard deviations. • Lognormal: the log mean is sampled in [−0.3, 0.7] and the log standard deviation in [0.25, 0.9]. • Gamma: concentration is sampled in [1, 5] and rate in [0.5, 2]. • Scaled beta: shape parameters are sampled in [0.8, 5], then values are mapped to a sampled interval with lower endpoint in [−1, 0] and upper endpoint in [0.8, 3.0]. 29

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

• Exponential: the rate is sampled in [0.4, 2.5]. • Gumbel: the location is sampled in [−1, 1] and the scale in [0.3, 1.8]. • Bernoulli: probabilities are sampled in [0.15, 0.85]. • Poisson: rates are sampled in [1, 12]. • Ordinal categorical: the number of ordered levels is sampled from {3, . . . , 6}; category probabilities are produced by a softmax over i.i.d. normal logits, and samples take ordered values 0, . . . , k − 1. • Mixture Gaussian: the mixture probability is sampled in [0.2, 0.8], the two means in [−2, −0.2] and [0.2, 2], and component standard deviations in [0.3, 1.0]. • Multinomial discrete: the number of unordered categories is sampled from {2, . . . , 10}; logits are sampled in [−1.5, 1.5], converted to probabilities by softmax, and the sampled category index is standardized. • Zipf/power-law: support size is sampled from {8, . . . , 20}, probabilities are proportional to v −α with α ∈ [2, 4], sampled values are capped at 10, and the result is centered. • Root dependency: 800 of the 1,000 synthetic SCM configurations use independent standardized Gaussian root blocks. The remaining 200 use multi-root dependent blocks: a common-factor Gaussian with ρ ∈ [0.25, 0.75] and random signs; a random-covariance Gaussian whose latent dimension is sampled from 1, . . . , r for r root variables with covariance LL⊤ + 0.15I; or a uniform-ball sampler that draws a random direction, radius U 1/r , scale in [0.75, 2.0], and center coordinates in [−0.5, 0.5] before column standardization. Structural noise. • Gaussian: additive Gaussian noise uses scale [0.05, 0.45]. • Uniform: additive bounded noise uses scale [0.05, 0.45]. • Laplace: heavy-tailed additive noise uses scale [0.03, 0.35]. • Student-t: degrees of freedom are sampled in 2.5–8 and scale in [0.03, 0.30]. • Multiplicative lognormal: multiplicative noise uses σ ∈ [0.03, 0.25]. • Heteroskedastic: the scale is base + slope · σ(|s|) with base in [0.02, 0.15] and slope in [0.02, 0.20]. • RFF-heteroskedastic: parent values pass through 64–256 random Fourier features with length scale [0.25, 3.0] and output scale [0.3, 1.8]; a softplus transform converts this latent function into a sample-specific noise scale. • Target-R2 Gaussian: the noise variance is chosen from the empirical signal variance to target R2 ∈ [0.1, 0.9]. • Mixture-outlier: a base scale in [0.02, 0.15] is mixed with rare outliers whose probability is [0.01, 0.08] and scale is [0.8, 2.5]. • Quantized: Gaussian-noisy values are rounded to a grid with step size [0.05, 0.35]. • Censored: Gaussian-noisy values are clipped at their empirical 3rd and 97th percentiles. • Beta: centered beta noise uses shape parameters in [1, 10] and scale in [0.05, 0.45]. • Exponential: centered one-sided disturbances use rate [0.8, 4.0] and scale [0.03, 0.35]. • Gumbel: centered extreme-value disturbances use scale [0.03, 0.35]. • No-noise option: deterministic local equations are allowed, but the generator limits no-noise assignments so that at most 10% of nodes in a graph use this option. Difficulty variants. • Schedule: within each graph-family/dimension block, the generator cycles through one clean baseline case, two single-stressor cases, and one paired-stressor case. • Signal and noise stressors: weak-signal cases multiply structural mechanism strength by 0.45; high-noise cases add extra Gaussian structural noise with scale 0.35 after the selected node noise. • Observation stressors: feature warping applies signed-log, tanh, or signed-square-root transforms to about 55% of columns; outlier contamination perturbs about 2.5% of entries with heavy-tailed jumps; missingness/imputation affects about 5% of entries and fills them with column means plus small jitter; feature corruption permutes about 3% of entries within the same column. • Intervention stressors: interventional distribution shift applies shift/scale perturbations to selected interventional columns; soft interventions blend the original structural value with the intervened value using softness 0.45. 30

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Interventions. • Main split: the obs+int split uses 800 observational rows and 200 interventional rows. • Target-pool size: each SCM samples an intervention target pool whose size grows with dimension: 2–4 targets for d = 10, 3–6 for d = 20, 4–8 for d = 30, 5–12 for d = 50, and 8–20 for d = 100. • Row assignment: every interventional row selects one target from this pool using balanced assignment, applies a hard do-resampling intervention, records the binary intervention mask, and recomputes all descendants. • Intervention value: the do-resampling value is drawn uniformly from the target’s empirical baseline 10th–90th percentile range. Intervention targets may include root or structural variables unless a sensitivity suite explicitly restricts targets to structural nodes. A.2

Synthetic Audit and Protocol-sensitivity Suites

Synthetic SCMs are audited before release. Graph-level checks require an acyclic graph, no isolated variables, minimum depth 3 for chain/tree/small-world/scale-free families and 2 otherwise, and at least ⌈0.8d⌉ edges. Dataset-level checks require finite values, no constant or near-constant continuous variables, no discrete variable with a single category occupying more than 99% of rows, boundary mass at explicit clipping boundaries at most 10%, no near-duplicate variable pair above correlation 0.99, and intervention masks consistent with the target pool. For obs+int files, every recorded target must receive at least one affected row, target values must change relative to their observational mechanism, and descendants are regenerated under the intervened values. The sample-size suite uses a fixed d = 30 representative subset of the synthetic pool, with 10 SCMs per graph family and n ∈ {100, 200, 500, 1000, 2000, 5000, 10000}. In the main text we report the observation-only sample-size results to isolate observational sample count. The intervention-protocol suite also uses 100 SCMs at d = 30, but changes the intervention generator while keeping the same graph/mechanism pool: • Default resampling uses the main protocol: 800 observational rows, 200 interventional rows, one target per interventional row, all sampled targets eligible, and hard do-resampling from the target’s 10th–90th percentile range. • Fixed setpoint keeps the 800/200 split and one target per row, but sets targets to one of their 20th, 50th, or 80th percentile baseline values with jitter equal to 1% of the interquartile range. • Parameter shift keeps the 800/200 split, restricts targets to structural non-root nodes, and changes the target mechanism by scaling parent effects, shifting a bias/threshold by 0.3–0.8 interquartile ranges, or changing nonlinear strength by a factor sampled from 0.6–0.85 or 1.15–1.6. • High intervention fraction uses 200 observational rows and 800 interventional rows with single-target hard resampling. • Multi-target per sample keeps the 800/200 split but samples 2–4 targets per interventional row and applies hard resampling to each selected target. • Dense mixed intervention combines the harder settings: 200 observational rows, 800 interventional rows, 2–4 structural targets per interventional row, and a balanced mixture of hard resampling, fixed-setpoint interventions, and parameter shifts. A.3

Agentic Construction Workflow and Quality Gates

Semantic and Formula SCMs are built with staged agentic workflows. We describe the paper-facing stages and quality gates without exposing implementation filenames. Each stage leaves an auditable record, and failed checks are routed back to graph design or executable specification before release. Stage 1: seed and reference grounding. The benchmark starts from human-curated seed banks. Semantic seeds cover a fixed 10-domain plan with 10 scenarios per domain and specify the setting, causal question, measurement context, plausible actions or interventions, and the operational ambiguity to expose. Formula seeds specify a scientific task family, possible formula anchors or formula-selection policy, response-variable intent, units, validity conditions, and experimental context. The first agentic stage collects compact grounding evidence: Semantic scenarios record concrete variables, records, sensors, actions, queues, delays, capacities, and outcomes; Formula scenarios record equations, constants, units, validity domains, instruments, calibration/readout variables, and causal-orientation constraints.

31

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Stage 2: structure and mechanism planning. Before any full graph is written, the workflow diagnoses several structurally distinct options and selects a candidate. The graph-design target is 15–25 observed variables with batchaverage size close to 20; the Formula validator keeps a wider 10–30-node safety envelope for justified deviations, although the released Formula pool remains within 16–25 variables. Edge count is controlled by the selected perscenario structural plan, with no single global quota; each plan must state expected edge-count, depth, root-count, and sink-count ranges, and the later graph must report its realized node, edge, root, sink, depth, and motif alignment. A mechanism-execution plan then binds each planned node to a root sampler, mechanism family, noise family, and range policy. This plan must include at least three scenario-justified executable mechanism families and normally keeps the allowed fraction of generic bounded-sigmoid children at or below 0.25. Stage 3: graph drafting, review, and repair. The graph stage turns the selected plan into variables, directed edges, intervention handles, and local mechanisms. Every variable must define its operational meaning, measurement scale, unit, numerical range, data source or sampling story, observation process, intervenability status, and child-level structural mechanism over exactly its graph parents. Every edge must have a local mechanism explanation. Semantic reviews check causal plausibility, measurement realism, no vague latent placeholders, no hidden parents, realistic intervention semantics, and no alias, unit-conversion, or tiny-noise near-copy variables. Measurement, score, status, and readout nodes are kept only when they change the causal mechanism, observation process, intervention semantics, or diagnostic logic. Formula reviews additionally check formula correctness, unit compatibility, validity ranges, solve-for and response-variable orientation, invalid reverse orientations, and justified calibration/readout nodes. Blocking findings trigger graph revision and re-review. Stage 4: executable translation. After graph approval, a translation table maps graph variables and edges into executable root distributions, structural operations, node-level noise, range policies, intervention resamplers, saved adjacency entries, and metadata. The executable specification must use legal structured operations over graph parents. It cannot introduce hidden parents, omit graph parents from the child operation, use batch-level row statistics, or rely on prose-only mechanisms. Formula specifications also store equation identifiers, inputs, outputs, constants, validity constraints, and residual checks; formula outputs are recomputed from parents and are never clipped merely to pass a range check. Interventional rows resample the flagged targets, record the binary intervention channel, and recompute descendants. Stage 5: deterministic validation and release gate. Final validation checks schema validity, acyclicity, unique node identifiers, valid edge references, finite generated values, nonconstant columns, compatible observation-only and mixed-interventional files, target changes under intervention, descendant recomputation, and agreement among saved adjacency, generated columns, masks, graph metadata, and executable specification. The data gates require continuous variables to avoid more than 10% exact mass at a valid-range boundary unless the graph declares a real boundary mechanism; continuous/non-discrete near-copy pairs are revised when absolute correlation reaches 0.90 without scientific justification; binary, categorical, ordinal, and grid-valued nodes are revised when one category dominates without a rare-event or pass-heavy justification. Formula SCMs must additionally pass residual checks for nonintervened formula rows within numerical tolerance. Anti-template checks block duplicate support topologies, repeated coarse role topologies, slot-filled mechanism text, generic readout/flag/calibration padding, and mechanism/root/noisefamily dominance. Only scenarios that pass review, executable validation, adversarial audit, and the final release gate enter the benchmark package. A.4

Semantic SCM Pool

Semantic scenario cards target operational DAGs with observable variables, actionable intervention handles, and scenario-specific measurement channels. The released semantic pool contains 100 operational scenarios, with 10 scenarios in each of 10 domains: cybersecurity and IT operations, education, finance and credit, government services, healthcare delivery, housing and real estate, manufacturing, public health, urban transportation, and water and sanitation. The realized release contains 17–25 variables (mean 22.89), 25–67 directed edges (mean 40.28), and 1–12 intervention targets (mean 3.82). Core metadata fields include variable definitions, units/ranges, observability, intervenability, parent lists, sampling policies, edge rationales, mechanism descriptions, monotonicity/sign tags, intervention resamplers, raw-value summaries, graph artifact paths, and generator-spec artifact paths. For example, the school-start-time scenario records that delaying the bell changes planned pickup schedules and circadian alignment, while grace-window policy affects whether an arrival is counted as tardy. Figure 18 reports the released Semantic mix by mechanism and by noise. The largest mechanism shares are event/count processes (15.3%), cutoffs (8.5%), stock-flow accumulation (8.3%), and ordinal rules (8.1%). On the noise side, 43.2% of nodes have no extra noise; the rest are led by normal (26.2%), then lognormal, Bernoulli, and count-style families.

32

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Semantic: mechanisms

Semantic: noise 15.3%

Event/count Threshold/piecewise Stock-flow Ordinal/categorical Regime switch Physical law Saturating response Sigmoid index Deterministic formula Clamp/min-max Product Measurement error Exp/log Ratio/normalization Linear-affine Other

7.5

7.5%

Lognormal

6.2%

Bernoulli

10.0

12.5

15.0

Negative binomial

3.8%

Beta

2.8%

Gamma

2.6%

Mixture normal

1.8%

Binomial

1.5% 4.4%

Other

13.4%

5.0

26.2%

Normal

6.0% 5.9% 5.9% 4.8% 4.8% 4.4% 4.4% 4.0% 3.7% 3.6% 2.9%

2.5

43.2%

No extra noise

8.5% 8.3% 8.1%

0.0

17.5

0

10

Formula: mechanisms

20.4%

5

10

15

20

30

40

50

72.8%

No extra noise

9.2% 8.6% 7.1% 6.2% 5.9% 4.1% 3.1% 2.6% 2.5% 2.0% 1.9% 1.9% 1.5% 1.5% 1.3% 1.2% 0.9%

0

20

Formula: noise

18.1%

Explicit equation Measurement error Linear-affine Physical law Calibration offset Ratio/normalization Threshold Product Exp/log Saturating response Power law Ordinal/categorical Additive index Stock-flow Clamp Sigmoid index Regime switch Trig/periodic Other

A P REPRINT

25

Share in released pool (%)

Student-t

6.3%

Beta

5.0%

Normal

5.0%

Laplace

4.4%

Mixture normal

4.2%

Other

2.3%

0

20

40

60

80

Share in released pool (%)

Figure 18: Composition of the released Semantic and Formula pools. Mechanism bars are edge shares; noise bars are node shares, from one representative file per released SCM. Remaining unlabeled families are grouped as Other. Formula nodes themselves have no extra noise, so that bar is large. A.5

Formula-grounded SCM Pool

Formula scenario cards target scientific DAGs whose key dependencies are grounded in explicit equations, causalorientation contracts, units, validity ranges, and measurement/readout semantics. The released formula pool contains 100 equation-grounded scenarios, with 10 scenarios in each of 10 domains: biology and ecology, astronomy, chemistry, mechanics and fluids, earth systems, electromagnetism, energy systems, materials and structures, optics and waves, and thermodynamics. The realized release contains 16–25 variables (mean 22.68), 18–50 directed edges (mean 32.97), 1–17 intervention targets (mean 5.78), and explicit equation/residual metadata for formula-derived variables. Core metadata fields include variable definitions, units/ranges, constants, formula identifiers, equation inputs/outputs, solve-for variables, causal-orientation contracts, invalid reverse-orientation notes, measurement readouts, intervention resamplers, raw summaries, and residual checks on non-intervened rows. For example, the thin-lens scenario records the lens equation for target image distance, the Airy-disk diffraction relation, and the magnification equation, then adds actuator backlash, exposure, saturation, and quality-control readouts downstream. Figure 18 also shows the Formula mix. The largest mechanism shares are explicit equations (18.1%), measurement-error edges (9.2%), linear-affine relations (8.6%), physical laws (7.1%), and calibration offsets (6.2%). Formula nodes are computed exactly from their parents, so they carry no extra noise; that is why “No extra noise” is 72.8% of nodes. The remaining noisy nodes are mostly measurement and readout terms: Student-t, beta, normal, Laplace, or mixture-normal. Both families use do-style intervention records with descendant recomputation. The difference is what the edges write down: a running process in Semantic, a named equation plus instruments in Formula.

B

Detailed Evaluation Protocol

This section records the evaluation conventions needed to interpret the reported results. File schemas, wrapper I/O details, and runnable commands are released with the code; here we summarize the protocol choices affecting comparability.

33

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

B.1

A P REPRINT

Evaluation Settings

Each executable SCM is evaluated under two exported splits from the same underlying graph. The obs-only split contains only unperturbed samples and uses nobs = 1000 in the main leaderboard. The obs+int split contains observational rows plus do-style interventional rows with descendant recomputation and uses nobs = 800, nint = 200 in the main leaderboard. Intervention targets are recorded in the exported masks. Methods that do not consume intervention indicators are evaluated only on supported splits and are marked missing otherwise. Each generated SCM benchmark scenario or synthetic configuration is evaluated with 5 random replicates, a fixed compromise between wall-clock cost and stable aggregate estimates. For a fixed method and split, we run the method on each replicate, score against the shared ground-truth adjacency, average replicate scores to a scenario/configuration-level value, and then aggregate over families, domains, dimensions, or the full arena. Failed, timed-out, or invalid runs are recorded explicitly and never silently removed. B.2

Metrics

Given a predicted directed graph Ĝ and ground-truth adjacency G, we report directed-edge precision, recall, F1, and structural Hamming distance (SHD). Higher F1 and lower SHD are better. Compact tables report F1/SHD as mean (standard deviation); full F1 and SHD tables appear in Section C. For PC and GIES, remaining undirected CPDAG edges are randomly oriented with a fixed per-file seed before these DAG metrics are computed, so the scores are directed-graph scores, not equivalence-class scores. Appendix complete tables additionally report structural intervention distance (SID), the number of ordered pairs whose intervention distributions would be misidentified from Ĝ; normalized SHD nSHD = SHD/(d(d − 1)); and, when a method returns edge scores, AUROC and average precision (AP). Hard-graph methods without edge scores therefore have missing AUROC/AP entries. Runtime and failure status are reported separately from accuracy metrics. B.3

Runtime, Hardware, and Timeout Policy

GPU-based wrappers were evaluated on NVIDIA RTX 4090 GPUs when GPU inference was required. CPU baselines ran on AMD EPYC 7763 CPU cores with bounded process-level parallel workers. The current audited protocol uses a one-hour wall-time budget per graph row. Runs are executed in shards, typically with five graph/replicate rows per shard, so the scheduler assigns a shard-level budget proportional to the number of rows plus a small startup guard. The scheduler records completed rows, partial rows, empty failures, timeout statuses, and elapsed time in status files. The main-text runtime ranking uses two protocols over the same synthetic mix: mean end-to-end wrapper time on the synthetic main benchmark for non-pretrained methods, and nested-k load/inference fits on a sample balanced across dimensions and graph families for pretrained methods. Reduced configurations are used only when the full configuration has been verified to be impractical under the benchmark budget, and these cases are recorded. The current real-data table is complete after these recorded repairs. In the real-data obs-int-as-obs conversion, PC timed out on four larger datasets (Wind Tunnel, Light Tunnel, PetShop temporal traffic 1, and PetShop temporal traffic 2) after a long best-effort budget; the reduced repair uses α = 0.001, at most 200 sampled rows, constant-column dropping, and fixed-seed random orientation, with predictions mapped back to the original variable dimension. CDIS timed out or failed to complete on the same four datasets; its reduced repair caps the conditioning-set size at 1 and preserves the native PAG output. SEA required a preprocessing repair on the two Causal Chamber datasets because constant columns produced a smaller returned matrix; the repaired wrapper drops constant columns internally and maps predictions back to the original dimension. B.4

Method Settings

Unless noted below, methods are run once for the leaderboard with wrapper defaults and no hyperparameter search on benchmark labels. Library defaults not listed below are left unchanged by the wrapper. For pretrained or amortized methods, the checkpoint/source named below is fixed across generated-SCM and real-data runs. Methods that output equivalence-class objects preserve the native output when available, and directed-edge F1/SHD are computed after the deterministic conversion rules below. • RandomRegress. A fixed-seed random-order regression baseline. It outputs a directed graph from nonzero regression support. • CDIS. Uses observational and intervention groups under the wrapper default conditioning search. For directed DAG metrics, PAG endpoint uncertainty is converted to a directed graph using the wrapper’s liberal orientation rule. The reduced repair caps conditioning-set size at 1 only for cases verified to exceed the budget. 34

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

• GIES. Uses the interventional score with the BIC-style penalty left at the wrapper default. Undirected edges in the returned essential-graph/CPDAG-like object are randomly oriented with a fixed per-file seed for DAG metrics. • IGSP. Uses Gaussian CI-style tests with α = αinv = 10−3 , depth 4, and 5 runs by default. It reports a directed graph directly. • PC. Uses stable PC with Fisher-z CI tests and α = 0.05 by default. Undirected CPDAG edges are randomly oriented with a fixed per-file seed for DAG metrics. The reduced real-data repair uses α = 0.001 and at most 200 rows after the full configuration times out. • DAS. Uses the additive spline wrapper with pruning parameters ηG = ηH = 0.001, α = 0.05, 10 splines, degree 3, and parent range 5–20. • LiNGAM. Uses the standard DirectLiNGAM wrapper and reports the returned directed graph. • DAGMA. Uses the wrapper’s automatic linear/nonlinear selection with default optimization settings and default weight threshold 0.3. • NOTEARS. Uses linear NOTEARS with λ1 = 0.1 and fixed adjacency threshold 0.3. • NOTEARS-MLP. Uses hidden size 10, λ1 = λ2 = 0.01, max iteration 100, and fixed adjacency threshold 0.3. • SDCD. Uses the two-stage wrapper with the default stage-1/stage-2 schedule and model threshold 0.1; highdimensional budget reductions are logged when invoked. • Arrow. Uses the released ArrowFM base checkpoint arrowfm-base.pt and its official decoded binary output. • AVICI. Uses the released AVICI scm-v0 pretrained model (checkpoint checkpoint_0300000.pkl), loaded by avici.load_pretrained(download="scm-v0"), and fixed threshold 0.5 for probabilistic adjacency outputs. • CauScale. Uses the released synthetic checkpoint synthetic/auprc=0.905_migrated.ckpt, feature sample size 500 by default, and fixed threshold 0.5 for probabilistic outputs. • CDFM. Uses the released CDFM checkpoint model.safetensors and its official output logic. • FoundCause. Uses the released checkpoint checkpoint.pt. The wrapper scores saved edge probabilities with a fixed 0.5 threshold. The released official binary decision was also scored; we report the 0.5-threshold output because it had higher average F1/SHD on the main, sample-size, and semantic/formula checks. This is a documented output rule, not a per-dataset hyperparameter search. • SEA. Uses the released obs-only FCI checkpoint fci_synthetic/model_best_epoch=373_auprc=0.842.ckpt, with 500 inference batches and FCI batch size 500. We do not run SEA as an interventional method: its interventional recipe uses GIES on sampled batches and is reported to need at least 250 observations per batch [5], while the main obs+int split has only nint = 200 interventional rows in total, further split across several targets. The repaired real-data wrapper drops constant columns internally and maps predictions back to the original variable dimension. • TabCausal. Uses the CausalArena paper-snapshot checkpoint tabcausal-v2.pt through the released inference wrapper and fixed threshold 0.5 for probabilistic adjacency outputs.

C

Complete Results

This section reports the numerical result views behind the compact main-text tables and figures. Overall benchmark tables include F1, precision, recall, SHD, SID, normalized SHD (nSHD), AUROC, and AP when the corresponding method output supports the metric; hard-graph methods without edge scores therefore have missing AUROC/AP entries. nSHD is computed as SHD/(d(d − 1)). Values are reported as mean±standard deviation across heterogeneous graph instances or datasets. The reported standard deviation is a descriptive spread and can exceed the mean for nonnegative metrics such as SHD. Bold and underline mark the best and second-best values within each comparable block, separately for every metric. To keep the appendix readable, factor-level decompositions use compact F1 views; the released CSV reports retain the full precision/recall/SHD/SID columns for each slice. C.1

Synthetic Benchmark

Table 7 reports the overall synthetic benchmark metrics. Tables 8 to 11 decompose the same benchmark by graph family, dimension, root-dependency mode, and difficulty component. Tables 12 and 13 give compact edge-level F1 views by mechanism and noise labels.

35

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

0.24 0.28 0.17 0.15 0.14 0.39 0.29 0.06 0.15 0.10 0.20 0.11 0.21 0.14 0.03 0.04 0.01 RandomRegress 0.37 0.33 0.22 0.19 0.44 0.37 0.20 0.24 0.15 0.36 0.26 0.29 0.25 0.18 0.07 0.02 DAS 0.76 0.52 0.39 0.24 0.66 0.43 0.37 0.33 0.23 0.48 0.45 0.33 0.36 0.32 0.11 0.08 DAGMA 0.72 0.63 0.41 0.31 0.53 0.48 0.37 0.38 0.28 0.43 0.39 0.38 0.33 0.18 0.14 0.03 SDCD 0.83 0.67 0.48 0.30 0.65 0.56 0.45 0.44 0.32 0.58 0.51 0.43 0.42 0.36 0.13 0.07 NOTEARS-MLP 0.85 0.78 0.61 0.59 0.77 0.74 0.58 0.60 0.49 0.68 0.59 0.57 0.60 0.44 0.34 0.13 CauScale 0.86 0.81 0.76 0.69 0.70 0.24 0.32 0.18 0.19 0.32 0.42 0.22 0.29 0.28 0.10 0.07 LiNGAM 0.61 0.56 0.34 0.47 0.35 0.23 0.38 0.37 0.24 0.52 0.46 0.38 0.37 0.33 0.12 0.09 NOTEARS 0.71 0.63 0.57 0.52 0.44 0.26 0.76 0.53 0.39 0.62 0.52 0.50 0.44 0.35 0.20 0.09 IGSP 0.94 0.80 0.63 0.63 0.55 0.42 0.68 0.62 0.33 0.66 0.50 0.49 0.43 0.37 0.15 0.09 Arrow 0.85 0.76 0.67 0.62 0.56 0.40 0.82 0.63 0.47 0.74 0.60 0.60 0.54 0.42 0.24 0.09 SEA 0.90 0.85 0.77 0.72 0.68 0.51 0.81 0.76 0.61 0.67 0.45 0.40 0.32 0.29 0.13 0.06 PC 0.80 0.64 0.52 0.57 0.42 0.32 0.68 0.48 0.38 0.34 0.26 0.45 0.40 0.31 0.15 0.04 CDFM 0.89 0.74 0.55 0.61 0.49 0.41 0.58 0.54 0.48 0.50 0.40 0.55 0.47 0.39 0.16 0.12 AVICI 0.79 0.71 0.67 0.62 0.57 0.43 0.78 0.62 0.50 0.51 0.40 0.60 0.55 0.33 0.25 0.06 CDIS 0.86 0.75 0.64 0.67 0.58 0.40 0.71 0.63 0.56 0.57 0.46 0.68 0.60 0.53 0.39 0.11 GIES 0.97 0.82 0.68 0.82 0.64 0.56 0.72 0.67 0.65 0.63 0.58 0.71 0.69 0.61 0.67 0.24 TabCausal 0.96 0.93 0.89 0.86 0.87 0.66 0.90 0.88 0.80 0.85 0.76 0.87 0.85 0.84 0.75 0.61 FoundCause 0.99 0.98 0.92 0.97 0.93 0.87 0.93 0.91 0.91 0.91 0.91 0.94 0.96 0.88 0.94 0.89 0.76

Win rate of row method over column method

1.0

0.8

0.6

0.4

0.2

res s DA DA S GM NO TE SDC A AR D S Ca -MLP uS c LiN ale NO GAM TE AR S IGS Arr P ow SE A PC CD FM AV IC CD I IS Tab GIE Fo Cau S un sa dC l au se

0.0

Ra n

do

mR eg

Row method

Per-SCM Pairwise Win-Rate Matrix

Compared method

Figure 19: Complete observation-only pairwise win-rate matrix. Each cell reports how often the row method wins over the column method across per-SCM or per-dataset comparisons, after averaging repeated runs within the same unit. On each unit, F1 and SHD are compared separately (higher F1 wins, lower SHD wins, ties count as 0.5), and wins from both metrics are pooled into one rate. Methods are ordered by mean family-wise win rate, matching Figure 1. Red cells indicate that the row method wins more often; blue cells indicate the opposite. Table 7: Complete synthetic benchmark metrics. Values are mean±standard deviation over replicate-level rows. Split

Method

N F1

Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs.

RandomRegress 5k CDIS 5k GIES 5k IGSP 5k PC 5k DAS 5k LiNGAM 5k DAGMA 5k NOTEARS 5k NOTEARS-MLP 5k SDCD 5k Arrow 5k AVICI 5k CauScale 5k

0.28±0.10 0.38±0.16 0.48±0.15 0.41±0.15 0.31±0.14 0.34±0.12 0.21±0.14 0.28±0.14 0.28±0.14 0.35±0.12 0.38±0.11 0.34±0.15 0.32±0.17 0.41±0.13

Prec.

Rec.

SHD

nSHD

0.24±0.11 0.47±0.20 0.46±0.17 0.43±0.18 0.35±0.15 0.33±0.13 0.53±0.23 0.45±0.17 0.55±0.18 0.43±0.14 0.36±0.13 0.52±0.16 0.64±0.18 0.57±0.18

0.36±0.11 0.34±0.16 0.52±0.14 0.40±0.14 0.29±0.15 0.36±0.14 0.14±0.11 0.22±0.13 0.20±0.12 0.31±0.14 0.43±0.13 0.27±0.15 0.24±0.15 0.34±0.12

218.0±310.0 0.103±0.054 143.0±352.0 0.064±0.045 123.6±174.9 0.063±0.039 135.6±192.3 0.063±0.034 115.4±149.0 0.062±0.034 142.6±180.1 0.077±0.039 109.6±142.2 0.060±0.029 107.2±138.6 0.059±0.031 104.2±136.0 0.057±0.029 107.0±136.4 0.060±0.033 135.5±165.8 0.076±0.042 106.0±140.7 0.056±0.031 104.5±139.7 0.054±0.026 85.8±107.2 0.054±0.033

SID

AUROC

1003.6±1488.4 0.50±0.06 826.6±1357.4 – 798.1±1141.9 – 783.3±1239.3 – 990.2±1596.1 – 722.1±1128.5 – 664.6±1139.3 – 697.4±1136.9 – 669.1±1127.0 0.50±0.08 742.4±1145.4 0.74±0.07 787.9±1137.8 – 661.5±1132.9 0.76±0.08 652.3±1129.6 0.75±0.09 605.2±983.2 0.84±0.06

AP 0.13±0.07 – – – – – – – 0.21±0.09 0.29±0.12 – 0.35±0.15 0.37±0.16 0.44±0.12

Continued on next page

36

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Split

Method

N F1

Obs. Obs. Obs. Obs.

CDFM FoundCause SEA TabCausal

5k 5k 5k 5k

0.42±0.15 0.43±0.13 0.46±0.20 136.4±176.9 0.071±0.036 654.8±1072.2 0.83±0.08 0.43±0.18 0.62±0.13 0.69±0.16 0.59±0.16 80.8±112.0 0.042±0.025 524.4±937.8 0.90±0.06 0.66±0.14 0.40±0.15 0.58±0.18 0.33±0.15 99.2±129.2 0.053±0.029 646.9±1105.6 0.84±0.07 0.44±0.14 0.50±0.16 0.64±0.17 0.44±0.18 93.6±127.3 0.050±0.035 622.5±1067.0 0.83±0.09 0.51±0.18

5k 5k 5k 5k 5k 5k 5k

0.40±0.15 0.49±0.20 0.36±0.15 140.1±353.8 0.061±0.043 808.5±1330.8 – – – 0.48±0.14 0.44±0.15 0.55±0.14 134.0±181.4 0.067±0.038 859.7±1217.2 – 0.39±0.14 0.42±0.18 0.37±0.14 132.8±186.1 0.062±0.033 780.7±1234.3 – – 0.38±0.11 0.34±0.12 0.46±0.13 148.1±174.9 0.084±0.043 808.8±1155.8 – – 0.35±0.17 0.65±0.17 0.26±0.17 104.2±139.7 0.053±0.026 651.3±1129.4 0.75±0.10 0.38±0.17 0.46±0.12 0.59±0.16 0.39±0.12 84.6±106.7 0.053±0.032 598.2±968.8 0.84±0.06 0.48±0.12 0.50±0.15 0.61±0.15 0.46±0.19 96.0±128.4 0.052±0.038 622.9±1060.5 0.84±0.09 0.52±0.18

Obs.+int CDIS Obs.+int GIES Obs.+int IGSP Obs.+int SDCD Obs.+int AVICI Obs.+int CauScale Obs.+int TabCausal

Prec.

Rec.

SHD

nSHD

SID

A P REPRINT

AUROC

AP

Table 8: Compact synthetic graph-family breakdown. Cells report Obs./Obs+int F1. Method

Erdos– Renyi

RandomRegress 0.27/– CDIS 0.42/0.42 GIES 0.50/0.48 IGSP 0.37/0.36 PC 0.28/– DAS 0.34/– LiNGAM 0.18/– DAGMA 0.23/– NOTEARS 0.24/– NOTEARS-MLP 0.32/– SDCD 0.41/0.40 Arrow 0.33/– AVICI 0.30/0.31 CauScale 0.39/0.43 CDFM 0.41/– FoundCause 0.62/– SEA 0.40/– TabCausal 0.47/0.48

Scale-free

Chain

Tree

Layered

Colliderrich

Bipartite layered

Dense local Small-world modules

0.27/– 0.30/0.32 0.42/0.46 0.41/0.39 0.25/– 0.39/– 0.21/– 0.34/– 0.30/– 0.39/– 0.41/0.41 0.33/– 0.38/0.41 0.42/0.46 0.42/– 0.56/– 0.41/– 0.51/0.53

0.30/– 0.33/0.36 0.45/0.50 0.38/0.37 0.34/– 0.31/– 0.22/– 0.30/– 0.30/– 0.37/– 0.34/0.35 0.33/– 0.34/0.37 0.40/0.45 0.40/– 0.59/– 0.37/– 0.52/0.53

0.27/– 0.31/0.35 0.44/0.48 0.43/0.40 0.30/– 0.38/– 0.22/– 0.37/– 0.33/– 0.41/– 0.39/0.38 0.36/– 0.40/0.43 0.46/0.49 0.43/– 0.61/– 0.41/– 0.55/0.56

0.27/– 0.40/0.42 0.52/0.50 0.39/0.37 0.32/– 0.33/– 0.21/– 0.26/– 0.27/– 0.32/– 0.39/0.39 0.35/– 0.29/0.31 0.43/0.46 0.40/– 0.66/– 0.41/– 0.47/0.48

0.27/– 0.41/0.42 0.49/0.47 0.38/0.36 0.28/– 0.33/– 0.19/– 0.24/– 0.23/– 0.32/– 0.39/0.40 0.33/– 0.30/0.31 0.40/0.44 0.41/– 0.63/– 0.39/– 0.48/0.48

0.27/– 0.58/0.58 0.61/0.51 0.46/0.44 0.29/– 0.30/– 0.18/– 0.19/– 0.20/– 0.29/– 0.41/0.38 0.35/– 0.24/0.24 0.42/0.47 0.46/– 0.73/– 0.42/– 0.47/0.44

0.31/– 0.34/0.39 0.45/0.46 0.44/0.41 0.35/– 0.37/– 0.23/– 0.32/– 0.30/– 0.39/– 0.37/0.35 0.37/– 0.34/0.37 0.41/0.47 0.45/– 0.61/– 0.40/– 0.52/0.52

0.30/– 0.34/0.36 0.44/0.48 0.36/0.35 0.31/– 0.30/– 0.19/– 0.25/– 0.24/– 0.32/– 0.36/0.37 0.30/– 0.30/0.33 0.36/0.40 0.40/– 0.57/– 0.35/– 0.46/0.48

Blocksparse layered 0.30/– 0.34/0.38 0.47/0.47 0.45/0.43 0.34/– 0.35/– 0.23/– 0.36/– 0.34/– 0.40/– 0.36/0.34 0.38/– 0.33/0.36 0.45/0.49 0.43/– 0.64/– 0.41/– 0.52/0.52

Table 9: Compact synthetic dimension breakdown. Cells report Obs./Obs+int F1. Method RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

10

20

30

50

100

0.34/– 0.36/0.44 0.49/0.55 0.49/0.47 0.42/– 0.43/– 0.29/– 0.34/– 0.33/– 0.39/– 0.38/0.40 0.45/– 0.41/0.46 0.35/0.43 0.54/– 0.67/– 0.46/– 0.56/0.58

0.31/– 0.39/0.42 0.49/0.49 0.45/0.42 0.34/– 0.38/– 0.23/– 0.30/– 0.30/– 0.38/– 0.39/0.39 0.40/– 0.40/0.43 0.39/0.44 0.49/– 0.65/– 0.44/– 0.56/0.56

0.28/– 0.39/0.40 0.49/0.48 0.42/0.40 0.31/– 0.34/– 0.21/– 0.29/– 0.28/– 0.36/– 0.39/0.38 0.36/– 0.36/0.38 0.41/0.46 0.44/– 0.63/– 0.42/– 0.53/0.54

0.25/– 0.38/0.38 0.47/0.45 0.36/0.34 0.26/– 0.30/– 0.18/– 0.27/– 0.26/– 0.34/– 0.38/0.36 0.29/– 0.28/0.29 0.44/0.47 0.37/– 0.61/– 0.39/– 0.47/0.47

0.23/– 0.36/0.35 0.47/0.44 0.32/0.30 0.20/– 0.25/– 0.13/– 0.22/– 0.21/– 0.29/– 0.37/0.36 0.21/– 0.16/0.16 0.48/0.49 0.27/– 0.56/– 0.27/– 0.36/0.35

37

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Table 10: Compact synthetic root-dependency breakdown. Cells report Obs./Obs+int F1. Method

correlated normal

independent roots

random covariance normal

uniform ball

0.23/– 0.27/0.29 0.37/0.40 0.35/0.34 0.25/– 0.33/– 0.21/– 0.32/– 0.30/– 0.38/– 0.36/0.35 0.32/– 0.32/0.34 0.43/0.46 0.38/– 0.58/– 0.36/– 0.43/0.46

0.29/– 0.38/0.41 0.49/0.49 0.41/0.39 0.31/– 0.34/– 0.21/– 0.28/– 0.28/– 0.35/– 0.39/0.38 0.35/– 0.33/0.35 0.41/0.46 0.43/– 0.63/– 0.40/– 0.51/0.51

0.22/– 0.30/0.30 0.39/0.40 0.37/0.35 0.26/– 0.33/– 0.19/– 0.31/– 0.29/– 0.37/– 0.36/0.35 0.30/– 0.28/0.29 0.42/0.45 0.34/– 0.54/– 0.36/– 0.42/0.43

0.30/– 0.45/0.47 0.54/0.51 0.43/0.40 0.32/– 0.32/– 0.20/– 0.24/– 0.24/– 0.32/– 0.38/0.38 0.34/– 0.32/0.34 0.42/0.46 0.44/– 0.66/– 0.40/– 0.51/0.50

RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

Table 11: Compact synthetic difficulty-component breakdown. Cells report Obs./Obs+int F1; multi-component cases contribute to each active component. Method

Baseline

Distribution shift

Feature corruption

Feature warping

High noise

Missingness

Outliers

Soft interventions

Weak signal

RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

0.28/– 0.40/0.42 0.51/0.52 0.42/0.41 0.33/– 0.34/– 0.26/– 0.31/– 0.30/– 0.36/– 0.40/0.40 0.36/– 0.35/0.38 0.44/0.49 0.46/– 0.66/– 0.42/– 0.54/0.54

0.28/– 0.39/0.43 0.50/0.46 0.43/0.41 0.32/– 0.35/– 0.23/– 0.30/– 0.29/– 0.36/– 0.40/0.37 0.38/– 0.34/0.36 0.44/0.47 0.45/– 0.65/– 0.43/– 0.53/0.51

0.27/– 0.35/0.37 0.45/0.46 0.38/0.35 0.31/– 0.32/– 0.18/– 0.28/– 0.27/– 0.34/– 0.36/0.36 0.34/– 0.31/0.34 0.39/0.44 0.40/– 0.59/– 0.39/– 0.48/0.49

0.27/– 0.37/0.40 0.46/0.48 0.39/0.38 0.30/– 0.33/– 0.21/– 0.29/– 0.28/– 0.35/– 0.39/0.39 0.32/– 0.29/0.32 0.40/0.46 0.41/– 0.59/– 0.40/– 0.43/0.45

0.30/– 0.39/0.40 0.52/0.53 0.46/0.43 0.31/– 0.41/– 0.16/– 0.30/– 0.28/– 0.39/– 0.42/0.41 0.38/– 0.29/0.30 0.45/0.49 0.42/– 0.64/– 0.43/– 0.51/0.52

0.27/– 0.36/0.38 0.45/0.46 0.39/0.37 0.31/– 0.32/– 0.20/– 0.29/– 0.29/– 0.35/– 0.36/0.37 0.34/– 0.32/0.35 0.40/0.44 0.41/– 0.60/– 0.39/– 0.49/0.50

0.26/– 0.31/0.32 0.36/0.37 0.30/0.29 0.23/– 0.30/– 0.06/– 0.16/– 0.14/– 0.32/– 0.28/0.29 0.24/– 0.31/0.34 0.30/0.38 0.27/– 0.51/– 0.24/– 0.39/0.42

0.28/– 0.38/0.41 0.49/0.45 0.41/0.40 0.31/– 0.34/– 0.22/– 0.29/– 0.28/– 0.36/– 0.40/0.38 0.35/– 0.34/0.34 0.43/0.44 0.44/– 0.64/– 0.41/– 0.52/0.50

0.30/– 0.38/0.41 0.50/0.51 0.43/0.41 0.31/– 0.35/– 0.22/– 0.27/– 0.26/– 0.35/– 0.40/0.38 0.35/– 0.28/0.30 0.41/0.45 0.43/– 0.62/– 0.40/– 0.48/0.47

38

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Table 12: Compact synthetic edge-level mechanism breakdown. Cells report Obs./Obs+int edge-F1 from aggregated TP/FP/FN counts. Method

activation additive

aggregation nonlin.

discrete assign.

exp log

linear additive

linear interact.

mixture regime

mult.

neural random

piecewise threshold

RandomRegress 0.06/– CDIS 0.34/0.34 GIES 0.48/0.47 IGSP 0.34/0.32 PC 0.23/– DAS 0.28/– LiNGAM 0.07/– DAGMA 0.22/– NOTEARS 0.06/– NOTEARS-MLP 0.27/– SDCD 0.37/0.37 Arrow 0.22/– AVICI 0.20/0.21 CauScale 0.42/0.44 CDFM 0.21/– FoundCause 0.55/– SEA 0.33/– TabCausal 0.38/0.39

0.10/– 0.32/0.32 0.43/0.42 0.32/0.31 0.25/– 0.30/– 0.12/– 0.28/– 0.12/– 0.28/– 0.39/0.39 0.26/– 0.29/0.30 0.46/0.48 0.25/– 0.57/– 0.34/– 0.48/0.49

0.04/– 0.32/0.32 0.45/0.44 0.31/0.29 0.20/– 0.29/– 0.04/– 0.19/– 0.04/– 0.28/– 0.36/0.35 0.16/– 0.18/0.19 0.39/0.40 0.20/– 0.52/– 0.28/– 0.36/0.35

0.08/– 0.37/0.36 0.52/0.51 0.38/0.36 0.27/– 0.30/– 0.09/– 0.27/– 0.08/– 0.27/– 0.43/0.42 0.26/– 0.23/0.24 0.50/0.52 0.23/– 0.61/– 0.38/– 0.45/0.45

0.08/– 0.39/0.39 0.59/0.58 0.42/0.40 0.30/– 0.31/– 0.13/– 0.29/– 0.09/– 0.24/– 0.42/0.42 0.33/– 0.20/0.21 0.54/0.56 0.23/– 0.65/– 0.42/– 0.43/0.44

0.08/– 0.40/0.39 0.55/0.53 0.39/0.38 0.29/– 0.29/– 0.12/– 0.29/– 0.09/– 0.26/– 0.46/0.45 0.30/– 0.24/0.25 0.54/0.57 0.24/– 0.65/– 0.40/– 0.46/0.47

0.07/– 0.38/0.38 0.56/0.55 0.40/0.39 0.29/– 0.32/– 0.11/– 0.28/– 0.08/– 0.25/– 0.44/0.44 0.31/– 0.22/0.23 0.53/0.56 0.23/– 0.64/– 0.41/– 0.44/0.45

0.04/– 0.34/0.35 0.47/0.47 0.34/0.33 0.21/– 0.31/– 0.05/– 0.19/– 0.03/– 0.23/– 0.38/0.37 0.20/– 0.18/0.19 0.47/0.49 0.21/– 0.57/– 0.31/– 0.38/0.37

0.05/– 0.37/0.36 0.51/0.50 0.36/0.34 0.23/– 0.33/– 0.07/– 0.22/– 0.04/– 0.24/– 0.39/0.38 0.23/– 0.16/0.18 0.45/0.47 0.21/– 0.57/– 0.32/– 0.38/0.37

0.08/– 0.37/0.36 0.54/0.54 0.39/0.38 0.28/– 0.29/– 0.11/– 0.28/– 0.09/– 0.25/– 0.42/0.42 0.28/– 0.25/0.26 0.54/0.56 0.24/– 0.65/– 0.41/– 0.46/0.47

Method RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

polynomial power

rational ratio

rff gp

saturation clip

tree based

trig periodic

0.07/– 0.31/0.31 0.42/0.41 0.32/0.31 0.23/– 0.25/– 0.07/– 0.27/– 0.09/– 0.27/– 0.37/0.36 0.22/– 0.31/0.31 0.47/0.49 0.24/– 0.57/– 0.32/– 0.48/0.48

0.07/– 0.32/0.33 0.47/0.46 0.33/0.32 0.23/– 0.31/– 0.10/– 0.27/– 0.08/– 0.24/– 0.39/0.39 0.27/– 0.18/0.20 0.44/0.46 0.21/– 0.54/– 0.34/– 0.37/0.37

0.05/– 0.32/0.31 0.40/0.39 0.29/0.28 0.19/– 0.30/– 0.05/– 0.19/– 0.04/– 0.26/– 0.35/0.35 0.18/– 0.27/0.28 0.44/0.46 0.21/– 0.56/– 0.26/– 0.46/0.46

0.07/– 0.37/0.36 0.54/0.53 0.39/0.37 0.27/– 0.32/– 0.10/– 0.26/– 0.07/– 0.25/– 0.43/0.42 0.29/– 0.21/0.22 0.50/0.52 0.22/– 0.62/– 0.39/– 0.42/0.42

0.03/– 0.21/0.21 0.27/0.29 0.22/0.22 0.15/– 0.20/– 0.02/– 0.17/– 0.04/– 0.25/– 0.28/0.28 0.10/– 0.21/0.22 0.30/0.31 0.19/– 0.44/– 0.19/– 0.34/0.35

0.05/– 0.33/0.33 0.45/0.44 0.33/0.31 0.22/– 0.29/– 0.06/– 0.22/– 0.05/– 0.30/– 0.39/0.39 0.20/– 0.24/0.25 0.44/0.45 0.21/– 0.56/– 0.30/– 0.41/0.41

39

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

A P REPRINT

Table 13: Compact synthetic edge-level noise breakdown. Cells report Obs./Obs+int edge-F1 from aggregated TP/FP/FN counts. Method

beta

RandomRegress 0.09/– CDIS 0.39/0.40 GIES 0.51/0.50 IGSP 0.39/0.39 PC 0.28/– DAS 0.28/– LiNGAM 0.11/– DAGMA 0.29/– NOTEARS 0.09/– NOTEARS-MLP 0.31/– SDCD 0.51/0.51 Arrow 0.30/– AVICI 0.33/0.34 CauScale 0.57/0.60 CDFM 0.25/– FoundCause 0.68/– SEA 0.42/– TabCausal 0.54/0.54 Method RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

C.2

censored

exponential

gaussian

gumbel

heterosk.

laplace

mixture outlier

mult. lognormal

quantized

0.06/– 0.34/0.34 0.49/0.48 0.34/0.32 0.23/– 0.31/– 0.06/– 0.23/– 0.06/– 0.26/– 0.38/0.38 0.20/– 0.21/0.22 0.41/0.44 0.22/– 0.55/– 0.33/– 0.37/0.37

0.08/– 0.35/0.35 0.49/0.48 0.38/0.37 0.27/– 0.29/– 0.11/– 0.28/– 0.09/– 0.30/– 0.46/0.46 0.28/– 0.27/0.29 0.54/0.56 0.24/– 0.64/– 0.39/– 0.49/0.49

0.06/– 0.34/0.34 0.48/0.47 0.35/0.33 0.24/– 0.31/– 0.07/– 0.24/– 0.07/– 0.27/– 0.38/0.38 0.24/– 0.21/0.22 0.44/0.47 0.21/– 0.57/– 0.33/– 0.43/0.42

0.06/– 0.34/0.33 0.48/0.47 0.34/0.32 0.24/– 0.30/– 0.08/– 0.24/– 0.06/– 0.25/– 0.38/0.38 0.25/– 0.20/0.21 0.46/0.47 0.22/– 0.57/– 0.34/– 0.41/0.40

0.08/– 0.37/0.37 0.51/0.50 0.37/0.35 0.26/– 0.32/– 0.09/– 0.27/– 0.08/– 0.29/– 0.43/0.43 0.28/– 0.25/0.26 0.50/0.52 0.22/– 0.61/– 0.38/– 0.46/0.47

0.06/– 0.33/0.33 0.48/0.47 0.34/0.32 0.23/– 0.29/– 0.08/– 0.24/– 0.06/– 0.24/– 0.37/0.36 0.23/– 0.20/0.21 0.45/0.47 0.22/– 0.56/– 0.32/– 0.39/0.39

0.05/– 0.32/0.32 0.46/0.45 0.32/0.30 0.22/– 0.22/– 0.09/– 0.22/– 0.05/– 0.20/– 0.31/0.32 0.19/– 0.26/0.27 0.50/0.53 0.24/– 0.60/– 0.30/– 0.36/0.40

0.09/– 0.37/0.37 0.51/0.50 0.38/0.38 0.27/– 0.29/– 0.11/– 0.29/– 0.09/– 0.30/– 0.49/0.49 0.28/– 0.28/0.30 0.56/0.58 0.24/– 0.65/– 0.41/– 0.49/0.50

0.06/– 0.32/0.33 0.48/0.47 0.34/0.32 0.23/– 0.31/– 0.07/– 0.24/– 0.06/– 0.26/– 0.36/0.37 0.23/– 0.19/0.20 0.43/0.46 0.21/– 0.55/– 0.32/– 0.39/0.39

RFF heterosk.

student t

target r2 gaussian

uniform

0.02/– 0.25/0.25 0.37/0.36 0.22/0.20 0.14/– 0.21/– 0.03/– 0.11/– 0.02/– 0.13/– 0.23/0.21 0.11/– 0.08/0.08 0.25/0.26 0.17/– 0.40/– 0.16/– 0.24/0.23

0.06/– 0.34/0.35 0.48/0.48 0.35/0.33 0.24/– 0.30/– 0.08/– 0.24/– 0.07/– 0.26/– 0.39/0.39 0.24/– 0.23/0.23 0.47/0.50 0.22/– 0.58/– 0.34/– 0.42/0.42

0.06/– 0.34/0.33 0.48/0.46 0.36/0.33 0.24/– 0.34/– 0.07/– 0.26/– 0.06/– 0.28/– 0.36/0.34 0.24/– 0.17/0.17 0.42/0.43 0.20/– 0.54/– 0.32/– 0.40/0.37

0.08/– 0.35/0.36 0.50/0.49 0.37/0.36 0.26/– 0.33/– 0.10/– 0.28/– 0.08/– 0.30/– 0.44/0.44 0.29/– 0.25/0.27 0.49/0.51 0.23/– 0.62/– 0.38/– 0.51/0.50

Semantic Benchmark

Table 14 reports the full Semantic benchmark metrics. Domain-level behavior is analyzed in Section 6.3. Table 14: Complete Semantic benchmark metrics. Values are mean±standard deviation over replicate-level rows. Split

Method

Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs.

RandomRegress 500 CDIS 500 GIES 500 IGSP 500 PC 500 DAS 500 LiNGAM 500 DAGMA 500 NOTEARS 500 NOTEARS-MLP 500 SDCD 500 Arrow 500 AVICI 500 CauScale 500 CDFM 500 FoundCause 500 SEA 500 TabCausal 500

Obs.+int CDIS

N F1

Prec.

Rec.

SHD

nSHD

SID

0.25±0.09 0.20±0.09 0.38±0.10 93.2±69.3 0.183±0.124 192.7±58.2 0.40±0.17 0.55±0.23 0.32±0.15 31.7±9.2 0.063±0.016 123.5±31.2 0.55±0.14 0.54±0.16 0.57±0.14 33.4±15.7 0.066±0.028 107.9±53.4 0.31±0.18 0.34±0.22 0.30±0.16 40.8±14.4 0.081±0.025 148.0±48.5 0.33±0.16 0.43±0.21 0.27±0.14 33.5±9.7 0.067±0.017 139.5±36.6 0.21±0.07 0.19±0.07 0.24±0.08 59.5±15.9 0.119±0.026 185.4±42.9 0.20±0.09 0.40±0.18 0.14±0.06 38.9±9.5 0.078±0.017 137.2±31.7 0.18±0.07 0.25±0.10 0.15±0.06 42.0±9.2 0.084±0.014 166.8±43.1 0.20±0.08 0.36±0.14 0.14±0.06 38.1±8.5 0.076±0.014 146.1±35.4 0.22±0.09 0.25±0.11 0.20±0.08 44.9±10.9 0.089±0.017 179.3±46.8 0.36±0.10 0.32±0.10 0.43±0.12 51.2±16.4 0.102±0.028 150.2±45.2 0.31±0.10 0.46±0.14 0.24±0.09 36.6±9.6 0.073±0.017 134.0±33.5 0.28±0.14 0.54±0.20 0.19±0.11 35.5±9.3 0.071±0.016 125.4±31.9 0.25±0.08 0.28±0.13 0.24±0.08 50.2±19.5 0.100±0.034 152.3±37.0 0.43±0.08 0.42±0.10 0.48±0.11 46.0±13.1 0.092±0.024 122.4±38.2 0.63±0.12 0.68±0.12 0.61±0.14 26.8±10.8 0.053±0.020 77.6±31.1 0.33±0.11 0.41±0.17 0.30±0.10 41.8±18.4 0.083±0.032 136.6±34.7 0.42±0.13 0.50±0.15 0.38±0.13 36.1±10.1 0.073±0.021 128.0±35.8

AUROC

AP

0.55±0.05 0.14±0.05 – – – – – – – – – – – – – – 0.59±0.06 0.20±0.06 0.69±0.06 0.18±0.07 – – 0.69±0.08 0.30±0.09 0.72±0.08 0.32±0.12 0.73±0.07 0.23±0.08 0.82±0.06 0.44±0.10 0.90±0.05 0.68±0.12 0.80±0.07 0.34±0.13 0.76±0.09 0.41±0.14

500 0.40±0.18 0.55±0.24 0.31±0.15 31.9±9.2 0.064±0.016 124.2±31.7 –

– Continued on next page

40

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Split

Method

Obs.+int GIES Obs.+int IGSP Obs.+int SDCD Obs.+int AVICI Obs.+int CauScale Obs.+int TabCausal

C.3

N F1 500 500 500 500 500 500

Prec.

Rec.

SHD

0.57±0.14 0.56±0.16 0.60±0.15 32.9±15.8 0.31±0.18 0.35±0.22 0.29±0.16 39.9±13.4 0.36±0.10 0.32±0.10 0.44±0.12 51.7±16.3 0.30±0.13 0.56±0.18 0.21±0.11 34.8±9.4 0.29±0.09 0.32±0.13 0.28±0.09 48.6±19.2 0.45±0.13 0.53±0.14 0.40±0.13 35.0±10.2

AUROC

A P REPRINT

nSHD

SID

0.066±0.029 0.079±0.023 0.103±0.028 0.070±0.016 0.096±0.033 0.070±0.021

100.0±49.0 – – 147.9±47.0 – – 150.8±46.4 – – 124.4±31.1 0.72±0.08 0.34±0.12 148.2±38.2 0.74±0.07 0.27±0.09 123.9±37.7 0.77±0.09 0.45±0.14

AP

Formula Benchmark

Table 15 reports the full Formula benchmark metrics. Scientific-domain behavior is analyzed in Section 6.3. Table 15: Complete Formula benchmark metrics. Values are mean±standard deviation over replicate-level rows. Split

Method

Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs.

RandomRegress 500 CDIS 500 GIES 500 IGSP 500 PC 500 DAS 500 LiNGAM 500 DAGMA 500 NOTEARS 500 NOTEARS-MLP 500 SDCD 500 Arrow 500 AVICI 500 CauScale 500 CDFM 500 FoundCause 500 SEA 500 TabCausal 500

Obs.+int CDIS Obs.+int GIES Obs.+int IGSP Obs.+int SDCD Obs.+int AVICI Obs.+int CauScale Obs.+int TabCausal

C.4

N F1

500 500 500 500 500 500 500

Prec.

Rec.

SHD

nSHD

SID

AUROC

AP

0.20±0.08 0.14±0.07 0.38±0.12 108.3±70.1 0.221±0.134 187.2±60.8 0.56±0.05 0.12±0.05 0.23±0.19 0.35±0.28 0.18±0.16 31.0±6.9 0.064±0.015 106.8±33.1 – – 0.47±0.14 0.42±0.15 0.54±0.15 36.7±14.8 0.076±0.031 95.0±51.6 – – 0.21±0.16 0.20±0.17 0.23±0.17 43.3±17.3 0.089±0.035 137.9±49.0 – – 0.22±0.18 0.29±0.23 0.18±0.15 31.1±7.0 0.064±0.015 111.9±32.9 – – 0.16±0.07 0.13±0.06 0.21±0.08 64.8±16.5 0.133±0.034 171.5±46.6 – – 0.24±0.10 0.34±0.15 0.19±0.08 34.7±8.6 0.071±0.017 122.2±36.4 – – 0.18±0.07 0.20±0.09 0.16±0.07 39.7±8.7 0.082±0.018 163.5±50.4 – – 0.19±0.09 0.26±0.12 0.15±0.07 35.2±7.5 0.073±0.017 139.9±44.9 0.60±0.06 0.16±0.06 0.19±0.07 0.20±0.08 0.19±0.08 44.1±9.1 0.091±0.020 177.4±57.9 0.69±0.06 0.15±0.05 0.28±0.09 0.22±0.07 0.42±0.13 61.5±16.3 0.127±0.036 141.8±50.2 – – 0.27±0.09 0.33±0.12 0.24±0.09 35.9±8.9 0.074±0.020 118.4±35.4 0.73±0.08 0.25±0.08 0.33±0.12 0.46±0.16 0.26±0.11 31.8±8.4 0.066±0.018 102.4±32.2 0.78±0.07 0.34±0.12 0.21±0.08 0.17±0.08 0.29±0.11 65.9±24.1 0.134±0.043 142.1±44.2 0.70±0.08 0.17±0.07 0.41±0.08 0.35±0.09 0.53±0.11 45.8±14.5 0.094±0.031 96.0±34.6 0.83±0.05 0.39±0.10 0.52±0.12 0.49±0.14 0.59±0.14 33.4±12.7 0.069±0.024 70.4±36.3 0.88±0.06 0.56±0.15 0.24±0.10 0.23±0.11 0.27±0.11 52.8±20.6 0.108±0.038 123.8±39.6 0.75±0.06 0.21±0.09 0.47±0.13 0.51±0.14 0.45±0.14 29.5±9.5 0.061±0.021 93.1±36.2 0.82±0.08 0.46±0.16 0.23±0.19 0.35±0.28 0.18±0.16 30.9±7.1 0.52±0.14 0.47±0.15 0.60±0.15 34.7±14.7 0.16±0.16 0.16±0.17 0.17±0.17 40.7±16.3 0.29±0.09 0.22±0.08 0.43±0.13 61.6±16.8 0.41±0.13 0.55±0.16 0.34±0.12 29.3±8.2 0.28±0.10 0.24±0.11 0.37±0.11 59.6±23.9 0.53±0.14 0.57±0.14 0.51±0.14 26.6±9.2

0.064±0.015 106.1±32.7 – – 0.072±0.031 78.6±47.8 – – 0.083±0.032 130.7±47.2 – – 0.127±0.036 140.5±52.0 – – 0.060±0.017 93.1±32.9 0.81±0.07 0.42±0.13 0.121±0.042 129.3±44.9 0.74±0.07 0.26±0.11 0.055±0.020 82.9±36.8 0.83±0.08 0.54±0.16

Real-data Benchmark

Table 16 reports the complete real-data method summary. Table 17 lists every method–dataset pair, including observationonly and observation-plus-intervention splits. The six observation-only sources are the CD-CSG datasets Abalone, Auto MPG, Cardiac Arrhythmia, Concrete Compressive Strength, Deutscher Wetterdienst, and Ozone. The seven interventional sources are Sachs flow-cytometry; PetShop high traffic, low traffic, temporal traffic 1, and temporal traffic 2; and the Causal Chambers Light Tunnel and Wind Tunnel experiments. Table 16: Complete real-data method metrics. Obs.-only includes the six native observation-only datasets plus observation-only conversions of the seven interventional datasets. Split

Method

Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs.

RandomRegress 13 CDIS 13 GIES 13 IGSP 13 PC 13 DAS 13 LiNGAM 13 DAGMA 13 NOTEARS 13 NOTEARS-MLP 13 SDCD 13 Arrow 13 AVICI 13

N F1 0.20±0.14 0.28±0.22 0.26±0.15 0.18±0.22 0.18±0.14 0.19±0.23 0.13±0.17 0.16±0.14 0.16±0.15 0.21±0.19 0.22±0.17 0.15±0.13 0.22±0.18

Prec.

Rec.

SHD

0.15±0.11 0.24±0.20 0.21±0.15 0.15±0.21 0.23±0.17 0.14±0.17 0.24±0.30 0.17±0.14 0.22±0.20 0.20±0.17 0.21±0.18 0.13±0.12 0.27±0.18

0.31±0.21 0.40±0.23 0.39±0.22 0.23±0.27 0.18±0.15 0.30±0.36 0.09±0.12 0.18±0.19 0.17±0.19 0.29±0.30 0.28±0.22 0.22±0.21 0.21±0.23

59.8±55.9 0.208±0.148 99.5±133.7 0.204±0.118 48.0±43.5 0.201±0.153 31.1±23.5 0.175±0.165 28.8±21.0 0.183±0.159 61.9±55.1 0.207±0.141 24.3±17.8 0.164±0.151 31.8±27.1 0.197±0.184 29.1±24.0 0.194±0.186 55.0±106.0 0.185±0.147 44.5±38.7 0.202±0.161 70.5±102.6 0.232±0.152 26.0±20.7 0.143±0.144

nSHD

SID

AUROC

144.2±186.5 – 179.8±222.8 – 164.5±210.2 – 65.9±71.8 – 63.8±55.2 – 130.7±126.7 – 50.5±49.0 – 62.4±54.5 – 57.6±49.5 – 146.8±300.8 – 89.8±95.7 – 133.9±175.7 0.65±0.16 55.4±52.4 0.68±0.21

AP – – – – – – – – – – – 0.20±0.16 0.29±0.21

Continued on next page

41

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Split

Method

N F1

Obs. Obs. Obs. Obs. Obs.

CauScale CDFM FoundCause SEA TabCausal

13 13 13 13 13

0.08±0.08 0.17±0.28 0.08±0.09 36.2±34.0 0.134±0.098 82.9±100.1 0.64±0.19 0.23±0.16 0.39±0.14 0.33±0.20 0.64±0.27 38.5±34.5 0.189±0.172 72.9±77.6 0.78±0.15 0.40±0.25 0.21±0.23 0.21±0.28 0.31±0.22 98.3±147.0 0.181±0.114 239.8±295.2 0.76±0.13 0.37±0.22 0.21±0.14 0.28±0.26 0.22±0.16 31.3±26.0 0.158±0.136 70.9±71.6 0.68±0.20 0.28±0.16 0.34±0.24 0.33±0.24 0.40±0.30 28.4±19.3 0.182±0.180 61.8±60.3 0.70±0.17 0.38±0.24

7 7 7 7 7 7 7

0.14±0.17 0.13±0.16 0.15±0.19 45.4±10.5 0.034±0.007 94.9±50.5 – – 0.29±0.13 0.20±0.11 0.54±0.10 118.7±53.8 0.097±0.053 403.1±215.0 – – 0.11±0.12 0.09±0.10 0.15±0.17 59.1±24.3 0.043±0.013 157.6±130.0 – – 0.27±0.12 0.22±0.13 0.43±0.08 103.4±55.0 0.081±0.030 252.4±144.6 – – 0.28±0.12 0.42±0.20 0.23±0.12 42.6±12.1 0.033±0.005 88.1±49.0 0.72±0.07 0.26±0.16 0.19±0.07 0.20±0.07 0.20±0.08 62.1±17.2 0.048±0.010 135.1±81.4 0.74±0.05 0.20±0.10 0.46±0.11 0.50±0.16 0.45±0.09 38.4±9.9 0.029±0.008 87.6±46.9 0.74±0.08 0.41±0.13

Obs.+int CDIS Obs.+int GIES Obs.+int IGSP Obs.+int SDCD Obs.+int AVICI Obs.+int CauScale Obs.+int TabCausal

Prec.

Rec.

SHD

nSHD

SID

AUROC

A P REPRINT

AP

Table 17: Complete real-data dataset-level metrics. Each row is one method on one scored real dataset. Dataset

Split

Method

Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Abalone Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Auto MPG Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia Arrhythmia

Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs. Obs.

RandomRegress 0.00 CDIS 0.27 GIES 0.36 IGSP 0.23 PC 0.09 DAS 0.00 LiNGAM 0.00 DAGMA 0.00 NOTEARS 0.00 NOTEARS-MLP 0.00 SDCD 0.00 Arrow 0.00 AVICI 0.53 CauScale 0.00 CDFM 0.50 FoundCause 0.00 SEA 0.00 TabCausal 0.07 RandomRegress 0.25 CDIS 0.57 GIES 0.50 IGSP 0.29 PC 0.00 DAS 0.00 LiNGAM 0.00 DAGMA 0.00 NOTEARS 0.00 NOTEARS-MLP 0.67 SDCD 0.44 Arrow 0.44 AVICI 0.00 CauScale 0.00 CDFM 0.33 FoundCause 0.00 SEA 0.40 TabCausal 0.22 RandomRegress 0.50 CDIS 0.33 GIES 0.00 IGSP 0.00 PC 0.29 DAS 0.67 LiNGAM 0.00 DAGMA 0.00 NOTEARS 0.00 NOTEARS-MLP 0.00 SDCD 0.57 Arrow 0.00 AVICI 0.40 CauScale 0.00 CDFM 0.75 FoundCause 0.40 SEA 0.00 TabCausal 0.86

F1 Precision Recall 0.00 0.20 0.24 0.16 0.07 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.42 0.00 0.33 0.00 0.00 0.05 0.20 0.50 0.40 0.25 0.00 0.00 0.00 0.00 0.00 0.50 0.33 0.33 0.00 0.00 0.22 0.00 0.50 0.17 0.40 0.33 0.00 0.00 0.25 0.50 0.00 0.00 0.00 0.00 0.50 0.00 0.50 0.00 0.60 0.50 0.00 0.75

0.00 0.43 0.71 0.43 0.14 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.71 0.00 1.00 0.00 0.00 0.14 0.33 0.67 0.67 0.33 0.00 0.00 0.00 0.00 0.00 1.00 0.67 0.67 0.00 0.00 0.67 0.00 0.33 0.33 0.67 0.33 0.00 0.00 0.33 1.00 0.00 0.00 0.00 0.00 0.67 0.00 0.33 0.00 1.00 0.33 0.00 1.00

SHD

SID nSHD

24.0 16.0 18.0 19.0 18.0 23.0 11.0 8.0 7.0 9.0 18.0 20.0 9.0 12.0 11.0 11.0 15.0 22.0 4.0 3.0 4.0 4.0 5.0 5.0 4.0 6.0 6.0 3.0 4.0 4.0 6.0 3.0 6.0 4.0 3.0 5.0 3.0 3.0 5.0 5.0 4.0 3.0 4.0 4.0 4.0 3.0 3.0 4.0 3.0 3.0 2.0 3.0 4.0 1.0

34.0 0.429 26.0 0.286 20.0 0.321 24.0 0.339 27.0 0.321 30.0 0.411 14.0 0.196 11.0 0.143 8.0 0.125 15.0 0.161 31.0 0.321 28.0 0.357 12.0 0.161 18.0 0.214 24.0 0.196 11.0 0.196 23.0 0.268 29.0 0.393 6.0 0.333 2.0 0.250 4.0 0.333 5.0 0.333 7.0 0.417 9.0 0.417 6.0 0.333 9.0 0.500 9.0 0.500 3.0 0.250 5.0 0.333 5.0 0.333 9.0 0.500 3.0 0.250 8.0 0.500 5.0 0.333 3.0 0.250 8.0 0.417 5.0 0.250 5.0 0.250 9.0 0.417 7.0 0.417 7.0 0.333 3.0 0.250 4.0 0.333 6.0 0.333 6.0 0.333 3.0 0.250 4.0 0.250 4.0 0.333 2.0 0.250 3.0 0.250 2.0 0.167 3.0 0.250 5.0 0.333 1.0 0.083

Continued on next page

42

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Dataset

Split

Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete Concrete DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD DWD Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Ozone Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs Sachs

Obs. RandomRegress 0.22 Obs. CDIS 0.32 Obs. GIES 0.25 Obs. IGSP 0.33 Obs. PC 0.00 Obs. DAS 0.16 Obs. LiNGAM 0.00 Obs. DAGMA 0.06 Obs. NOTEARS 0.06 Obs. NOTEARS-MLP 0.06 Obs. SDCD 0.00 Obs. Arrow 0.15 Obs. AVICI 0.00 Obs. CauScale 0.00 Obs. CDFM 0.22 Obs. FoundCause 0.23 Obs. SEA 0.00 Obs. TabCausal 0.09 Obs. RandomRegress 0.36 Obs. CDIS 0.80 Obs. GIES 0.36 Obs. IGSP 0.75 Obs. PC 0.00 Obs. DAS 0.57 Obs. LiNGAM 0.00 Obs. DAGMA 0.15 Obs. NOTEARS 0.15 Obs. NOTEARS-MLP 0.29 Obs. SDCD 0.00 Obs. Arrow 0.29 Obs. AVICI 0.20 Obs. CauScale 0.00 Obs. CDFM 0.40 Obs. FoundCause 0.86 Obs. SEA 0.31 Obs. TabCausal 0.57 Obs. RandomRegress 0.00 Obs. CDIS 0.00 Obs. GIES 0.00 Obs. IGSP 0.00 Obs. PC 0.29 Obs. DAS 0.50 Obs. LiNGAM 0.00 Obs. DAGMA 0.44 Obs. NOTEARS 0.44 Obs. NOTEARS-MLP 0.44 Obs. SDCD 0.25 Obs. Arrow 0.00 Obs. AVICI 0.57 Obs. CauScale 0.00 Obs. CDFM 0.50 Obs. FoundCause 0.00 Obs. SEA 0.33 Obs. TabCausal 0.00 Obs. RandomRegress 0.28 Obs. CDIS 0.26 Obs. GIES 0.18 Obs. IGSP 0.31 Obs. PC 0.32 Obs. DAS 0.14 Obs. LiNGAM 0.30 Obs. DAGMA 0.35 Obs. NOTEARS 0.34 Obs. NOTEARS-MLP 0.29 Obs. SDCD 0.33 Obs. Arrow 0.29 Obs. AVICI 0.22 Obs. CauScale 0.13 Obs. CDFM 0.33 Obs. FoundCause 0.26 Obs. SEA 0.32 Obs. TabCausal 0.37 Obs.+int CDIS 0.26 Obs.+int GIES 0.37

Method

F1 Precision Recall 0.14 0.24 0.17 0.23 0.00 0.10 0.00 0.04 0.04 0.04 0.00 0.11 0.00 0.00 0.13 0.17 0.00 0.06 0.29 0.67 0.29 0.75 0.00 0.40 0.00 0.11 0.11 0.20 0.00 0.20 0.17 0.00 0.25 1.00 0.22 0.40 0.00 0.00 0.00 0.00 0.25 0.40 0.00 0.33 0.33 0.33 0.20 0.00 0.50 0.00 0.33 0.00 0.33 0.00 0.20 0.26 0.14 0.24 0.33 0.11 0.57 0.43 0.56 0.27 0.32 0.29 0.43 0.20 0.28 0.23 0.35 0.35 0.26 0.27

SHD

SID nSHD

0.50 26.0 0.50 17.0 0.50 22.0 0.62 20.0 0.00 21.0 0.38 26.0 0.00 14.0 0.12 26.0 0.12 26.0 0.12 27.0 0.00 33.0 0.25 22.0 0.00 14.0 0.00 8.0 0.62 26.0 0.38 20.0 0.00 8.0 0.25 33.0 0.50 5.0 1.00 2.0 0.50 5.0 0.75 2.0 0.00 6.0 1.00 6.0 0.00 8.0 0.25 9.0 0.25 9.0 0.50 8.0 0.00 8.0 0.50 8.0 0.25 5.0 0.00 5.0 1.00 7.0 1.0 0.75 0.50 7.0 1.00 6.0 0.00 5.0 0.00 5.0 0.00 5.0 0.00 5.0 0.33 4.0 0.67 3.0 0.00 4.0 0.67 4.0 4.0 0.67 0.67 4.0 0.33 4.0 0.00 5.0 2.0 0.67 0.00 3.0 1.00 5.0 0.00 3.0 0.33 4.0 0.00 4.0 0.45 39.0 0.25 26.0 0.25 37.0 0.45 34.0 0.30 24.0 0.20 38.0 0.20 18.0 0.30 19.0 0.25 18.0 0.30 24.0 0.35 24.0 0.30 26.0 0.15 19.0 0.10 21.0 0.40 25.0 0.30 28.0 0.30 22.0 0.40 22.0 0.25 26.0 0.60 35.0

35.0 0.361 27.0 0.236 28.0 0.306 22.0 0.278 35.0 0.292 32.0 0.361 20.0 0.194 32.0 0.361 33.0 0.361 33.0 0.375 39.0 0.458 32.0 0.306 22.0 0.194 8.0 0.111 56.0 0.361 24.0 0.278 8.0 0.111 43.0 0.458 8.0 0.250 2.0 0.100 8.0 0.250 3.0 0.100 10.0 0.300 6.0 0.300 9.0 0.400 11.0 0.450 11.0 0.450 11.0 0.400 11.0 0.400 10.0 0.400 11.0 0.250 6.0 0.250 12.0 0.350 1.0 0.050 9.0 0.350 7.0 0.300 7.0 0.417 6.0 0.417 7.0 0.417 8.0 0.417 5.0 0.333 5.0 0.250 4.0 0.333 5.0 0.333 5.0 0.333 5.0 0.333 7.0 0.333 8.0 0.417 4.0 0.167 3.0 0.250 9.0 0.417 4.0 0.250 4.0 0.333 6.0 0.333 40.0 – 34.0 – 78.0 – 54.0 – 77.0 – 64.0 – 40.0 – 45.0 – 44.0 – 69.0 – 62.0 – 44.0 – 44.0 – 47.0 – 48.0 – 56.0 – 59.0 – 32.0 – 34.0 – 34.0 –

Continued on next page

43

A P REPRINT

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Dataset

Split

Sachs Sachs Sachs Sachs Sachs PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop HT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop LT PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1

Obs.+int IGSP 0.31 Obs.+int SDCD 0.34 Obs.+int AVICI 0.41 Obs.+int CauScale 0.31 Obs.+int TabCausal 0.38 Obs. RandomRegress 0.29 Obs. CDIS 0.35 Obs. GIES 0.29 Obs. IGSP 0.18 Obs. PC 0.35 Obs. DAS 0.10 Obs. LiNGAM 0.30 Obs. DAGMA 0.13 Obs. NOTEARS 0.11 Obs. NOTEARS-MLP 0.26 Obs. SDCD 0.14 Obs. Arrow 0.15 Obs. AVICI 0.17 Obs. CauScale 0.20 Obs. CDFM 0.35 Obs. FoundCause 0.18 Obs. SEA 0.26 Obs. TabCausal 0.47 Obs.+int CDIS 0.34 Obs.+int GIES 0.19 Obs.+int IGSP 0.16 Obs.+int SDCD 0.17 Obs.+int AVICI 0.40 Obs.+int CauScale 0.25 Obs.+int TabCausal 0.51 Obs. RandomRegress 0.23 Obs. CDIS 0.35 Obs. GIES 0.30 Obs. IGSP 0.20 Obs. PC 0.33 Obs. DAS 0.13 Obs. LiNGAM 0.13 Obs. DAGMA 0.12 Obs. NOTEARS 0.07 Obs. NOTEARS-MLP 0.11 Obs. SDCD 0.22 Obs. Arrow 0.23 Obs. AVICI 0.23 Obs. CauScale 0.15 Obs. CDFM 0.34 Obs. FoundCause 0.13 Obs. SEA 0.25 Obs. TabCausal 0.23 Obs.+int CDIS 0.36 Obs.+int GIES 0.30 Obs.+int IGSP 0.20 Obs.+int SDCD 0.16 Obs.+int AVICI 0.21 Obs.+int CauScale 0.18 Obs.+int TabCausal 0.40 Obs. RandomRegress 0.15 Obs. CDIS 0.10 Obs. GIES 0.23 Obs. IGSP 0.00 Obs. PC 0.28 Obs. DAS 0.05 Obs. LiNGAM 0.07 Obs. DAGMA 0.13 Obs. NOTEARS 0.15 Obs. NOTEARS-MLP 0.19 Obs. SDCD 0.21 Obs. Arrow 0.10 Obs. AVICI 0.19 Obs. CauScale 0.12 Obs. CDFM 0.26 Obs. FoundCause 0.17 Obs. SEA 0.27 Obs. TabCausal 0.41 Obs.+int CDIS 0.00

Method

F1 Precision Recall 0.24 0.33 0.50 0.32 0.36 0.21 0.33 0.22 0.15 0.33 0.07 0.67 0.19 0.21 0.29 0.11 0.16 0.28 0.16 0.29 0.11 0.29 0.43 0.31 0.13 0.13 0.11 0.39 0.23 0.50 0.16 0.33 0.23 0.19 0.31 0.10 0.24 0.19 0.13 0.15 0.17 0.29 0.39 0.13 0.24 0.08 0.22 0.25 0.33 0.21 0.19 0.10 0.23 0.17 0.37 0.09 0.06 0.15 0.00 0.43 0.04 0.18 0.21 0.44 0.29 0.15 0.09 0.29 0.10 0.19 0.11 0.25 0.41 0.00

SHD

SID nSHD

0.45 35.0 0.35 24.0 0.35 16.0 0.30 24.0 0.40 21.0 0.48 89.0 0.38 51.0 0.40 74.0 0.21 70.0 0.38 51.0 0.14 94.0 0.19 36.0 0.10 43.0 0.07 41.0 0.24 49.0 0.21 90.0 0.14 58.0 0.12 44.0 0.26 76.0 0.45 62.0 0.45 163.0 0.24 49.0 0.52 45.0 0.38 53.0 0.43 142.0 0.19 70.0 0.43 162.0 0.40 47.0 0.29 66.0 0.52 40.0 0.37 97.0 0.37 54.0 0.44 78.0 0.21 60.0 0.35 54.0 0.19 96.0 0.09 44.0 0.09 46.0 0.05 45.0 0.09 49.0 0.30 80.0 0.19 47.0 0.16 42.0 0.19 77.0 0.58 82.0 0.28 146.0 0.28 66.0 0.21 51.0 0.40 54.0 0.53 102.0 0.21 60.0 0.44 183.0 0.19 53.0 0.19 65.0 0.44 50.0 0.33 144.0 0.30 210.0 0.47 122.0 0.00 43.0 0.21 42.0 0.09 125.0 0.05 44.0 0.09 44.0 0.09 40.0 0.14 43.0 0.35 98.0 0.12 74.0 0.14 44.0 0.16 82.0 0.44 94.0 0.42 160.0 0.30 57.0 0.42 44.0 0.00 43.0

62.0 – – 42.0 34.0 – 34.0 – 51.0 – 208.0 0.060 152.0 0.034 361.0 0.050 200.0 0.047 150.0 0.034 286.0 0.063 99.0 0.024 127.0 0.029 105.0 0.028 170.0 0.033 190.0 0.061 147.0 0.039 105.0 0.030 199.0 0.051 99.0 0.042 637.0 0.110 132.0 0.033 122.0 0.030 159.0 0.036 528.0 0.096 211.0 0.047 398.0 0.109 103.0 0.032 139.0 0.045 84.0 0.027 230.0 0.059 144.0 0.033 344.0 0.048 209.0 0.037 148.0 0.033 271.0 0.059 128.0 0.027 135.0 0.028 129.0 0.027 151.0 0.030 182.0 0.049 140.0 0.029 145.0 0.026 214.0 0.047 228.0 0.050 517.0 0.089 174.0 0.040 144.0 0.031 146.0 0.033 462.0 0.062 214.0 0.037 450.0 0.112 153.0 0.032 188.0 0.040 159.0 0.030 458.0 0.088 538.0 0.128 552.0 0.074 113.0 0.026 115.0 0.026 289.0 0.076 121.0 0.027 131.0 0.027 114.0 0.024 138.0 0.026 251.0 0.060 207.0 0.045 125.0 0.027 240.0 0.050 186.0 0.057 557.0 0.098 192.0 0.035 156.0 0.027 113.0 0.026

Continued on next page

44

A P REPRINT

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Dataset

Split

PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T1 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 PetShop T2 Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Light Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel

Obs.+int GIES 0.21 Obs.+int IGSP 0.00 Obs.+int SDCD 0.23 Obs.+int AVICI 0.18 Obs.+int CauScale 0.16 Obs.+int TabCausal 0.47 Obs. RandomRegress 0.09 Obs. CDIS 0.13 Obs. GIES 0.24 Obs. IGSP 0.00 Obs. PC 0.24 Obs. DAS 0.05 Obs. LiNGAM 0.07 Obs. DAGMA 0.13 Obs. NOTEARS 0.15 Obs. NOTEARS-MLP 0.28 Obs. SDCD 0.20 Obs. Arrow 0.10 Obs. AVICI 0.19 Obs. CauScale 0.12 Obs. CDFM 0.26 Obs. FoundCause 0.17 Obs. SEA 0.26 Obs. TabCausal 0.41 Obs.+int CDIS 0.00 Obs.+int GIES 0.21 Obs.+int IGSP 0.11 Obs.+int SDCD 0.21 Obs.+int AVICI 0.18 Obs.+int CauScale 0.21 Obs.+int TabCausal 0.44 Obs. RandomRegress 0.14 Obs. CDIS 0.08 Obs. GIES 0.45 Obs. IGSP 0.00 Obs. PC 0.06 Obs. DAS 0.07 Obs. LiNGAM 0.48 Obs. DAGMA 0.31 Obs. NOTEARS 0.40 Obs. NOTEARS-MLP 0.13 Obs. SDCD 0.35 Obs. Arrow 0.12 Obs. AVICI 0.00 Obs. CauScale 0.22 Obs. CDFM 0.46 Obs. FoundCause 0.08 Obs. SEA 0.07 Obs. TabCausal 0.50 Obs.+int CDIS 0.00 Obs.+int GIES 0.53 Obs.+int IGSP 0.00 Obs.+int SDCD 0.51 Obs.+int AVICI 0.39 Obs.+int CauScale 0.18 Obs.+int TabCausal 0.67 Obs. RandomRegress 0.09 Obs. CDIS 0.07 Obs. GIES 0.21 Obs. IGSP 0.00 Obs. PC 0.16 Obs. DAS 0.00 Obs. LiNGAM 0.35 Obs. DAGMA 0.27 Obs. NOTEARS 0.17 Obs. NOTEARS-MLP 0.00 Obs. SDCD 0.19 Obs. Arrow 0.06 Obs. AVICI 0.17 Obs. CauScale 0.05 Obs. CDFM 0.37 Obs. FoundCause 0.29 Obs. SEA 0.23 Obs. TabCausal 0.23

Method

F1 Precision Recall 0.13 0.00 0.16 0.26 0.15 0.48 0.05 0.08 0.16 0.00 0.35 0.04 0.18 0.21 0.44 0.41 0.14 0.09 0.29 0.10 0.19 0.10 0.26 0.41 0.00 0.14 0.08 0.15 0.25 0.17 0.43 0.14 0.04 0.55 0.00 0.18 0.06 0.86 0.28 0.47 0.07 0.54 0.07 0.00 0.47 0.86 0.05 1.00 0.78 0.00 0.42 0.00 0.46 0.79 0.24 0.84 0.07 0.04 0.17 0.00 0.50 0.00 0.46 0.22 0.14 0.00 0.21 0.03 0.28 1.00 0.42 0.43 0.23 0.29

SHD

SID nSHD

0.47 145.0 510.0 0.088 0.00 43.0 113.0 0.026 0.37 96.0 279.0 0.059 0.14 47.0 118.0 0.029 0.16 67.0 156.0 0.041 0.47 42.0 119.0 0.026 0.21 166.0 587.0 0.101 0.37 207.0 593.0 0.126 0.49 121.0 552.0 0.074 0.00 43.0 113.0 0.026 0.19 43.0 124.0 0.026 0.09 125.0 289.0 0.076 0.05 44.0 121.0 0.027 0.09 44.0 131.0 0.027 0.09 40.0 114.0 0.024 0.21 41.0 125.0 0.025 0.35 107.0 261.0 0.065 0.12 74.0 207.0 0.045 0.14 44.0 125.0 0.027 0.16 83.0 244.0 0.051 0.44 94.0 186.0 0.057 0.42 165.0 542.0 0.101 0.26 51.0 159.0 0.031 0.42 44.0 156.0 0.027 0.00 43.0 113.0 0.026 0.47 139.0 665.0 0.085 0.19 107.0 404.0 0.065 0.35 100.0 277.0 0.061 0.14 47.0 128.0 0.029 0.26 72.0 274.0 0.044 0.44 44.0 121.0 0.027 0.14 83.0 122.0 0.059 0.30 400.0 432.0 0.284 0.39 52.0 62.0 0.037 0.00 57.0 57.0 0.041 0.04 63.0 81.0 0.045 0.11 140.0 197.0 0.100 0.33 41.0 39.0 0.029 0.33 84.0 85.0 0.060 0.35 57.0 64.0 0.041 0.54 402.0 1127.0 0.286 0.26 49.0 62.0 0.035 0.44 349.0 627.0 0.248 0.00 57.0 63.0 0.041 0.14 56.0 52.0 0.040 42.0 0.029 0.32 41.0 0.44 530.0 714.0 0.377 0.04 55.0 55.0 0.039 0.37 40.0 41.0 0.028 0.00 57.0 57.0 0.041 0.72 71.0 199.0 0.050 0.00 57.0 57.0 0.041 0.56 59.0 123.0 0.042 0.26 46.0 40.0 0.033 0.14 75.0 90.0 0.053 0.56 29.0 37.0 0.021 0.12 93.0 134.0 0.094 0.31 299.0 376.0 0.301 0.26 81.0 114.0 0.082 0.00 42.0 42.0 0.042 0.10 40.0 44.0 0.040 0.00 121.0 218.0 0.122 0.29 44.0 51.0 0.044 0.33 76.0 83.0 0.077 0.21 81.0 107.0 0.082 0.00 53.0 58.0 0.053 0.17 60.0 62.0 0.060 0.17 226.0 282.0 0.228 0.12 49.0 53.0 0.049 0.02 41.0 41.0 0.041 0.33 45.0 48.0 0.045 0.21 44.0 46.0 0.044 0.24 66.0 99.0 0.067 0.19 52.0 58.0 0.052 Continued on next page

45

A P REPRINT

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

C.5

Dataset

Split

Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel Wind Tunnel

Obs.+int CDIS Obs.+int GIES Obs.+int IGSP Obs.+int SDCD Obs.+int AVICI Obs.+int CauScale Obs.+int TabCausal

Method

F1 Precision Recall 0.00 0.20 0.00 0.28 0.16 0.08 0.36

SHD

SID nSHD

0.00 0.00 42.0 0.12 0.60 197.0 0.00 0.00 42.0 0.20 0.48 100.0 0.50 0.10 42.0 0.10 0.07 66.0 0.48 0.29 43.0

42.0 0.042 424.0 0.199 42.0 0.042 198.0 0.101 41.0 0.042 65.0 0.067 42.0 0.043

A P REPRINT

Sample-size Scaling

Table 18 reports the numerical F1/SHD values for the observation-only sample-size scaling suite, which fixes d = 30 and varies the number of samples. Table 18: Complete sample-size sensitivity results. Cells report F1/SHD for the observation-only d = 30 synthetic suite. Method

n = 100

n = 200

n = 500

n = 1000

n = 2000

n = 5000

n = 10000

RandomRegress CDIS GIES IGSP PC DAS LiNGAM DAGMA NOTEARS NOTEARS-MLP SDCD Arrow AVICI CauScale CDFM FoundCause SEA TabCausal

0.25/76.6 0.23/51.1 0.35/61.8 0.25/52.3 0.25/50.6 0.22/61.5 0.20/51.3 0.27/53.0 0.26/49.2 0.26/78.2 0.22/118.1 0.32/47.5 0.22/58.0 0.26/66.5 0.34/70.5 0.43/49.7 0.26/80.8 0.33/47.6

0.28/76.8 0.29/49.4 0.41/56.2 0.31/51.8 0.27/50.2 0.28/61.5 0.20/50.5 0.29/50.1 0.28/48.2 0.32/57.9 0.26/100.0 0.34/46.5 0.31/47.5 0.33/49.7 0.38/65.5 0.52/45.2 0.34/58.6 0.42/43.9

0.28/84.1 0.36/48.4 0.46/54.2 0.38/52.2 0.30/50.8 0.32/65.4 0.21/49.7 0.29/49.6 0.28/47.7 0.35/51.5 0.35/75.3 0.35/46.3 0.35/44.9 0.40/44.6 0.42/61.2 0.61/39.0 0.41/45.7 0.49/41.4

0.29/88.8 0.39/50.8 0.49/54.6 0.41/54.3 0.31/52.5 0.35/67.5 0.22/49.5 0.29/49.1 0.28/47.4 0.36/49.9 0.38/68.7 0.36/46.6 0.35/44.5 0.41/44.6 0.44/58.0 0.64/35.6 0.42/45.0 0.53/39.2

0.28/99.2 0.40/54.2 0.50/57.9 0.43/58.9 0.32/54.3 0.36/71.0 0.21/49.4 0.29/48.8 0.28/47.5 0.36/49.9 0.40/62.8 0.36/47.5 0.35/44.3 0.40/47.0 0.46/54.5 0.65/34.7 0.42/45.1 0.55/39.0

0.27/113.0 0.42/60.0 0.49/65.3 0.43/68.8 0.31/57.9 0.35/104.7 0.21/49.5 0.29/49.0 0.28/47.3 0.37/49.4 0.41/59.2 0.34/50.0 0.36/43.7 0.37/52.6 0.48/52.4 0.65/34.1 0.42/44.7 0.57/38.1

0.26/125.7 0.43/61.7 0.49/70.6 0.43/76.8 0.31/60.1 0.36/87.1 0.21/49.5 0.29/49.0 0.27/47.5 0.36/49.6 0.43/54.2 0.33/53.5 0.36/40.4 0.35/56.3 0.48/51.2 0.66/33.7 0.42/44.8 0.57/38.6

The sample-size suite fixes d = 30 and varies only the number of observational samples. Higher F1 and lower SHD are better; bold and underline mark the best and second-best F1 and SHD separately in each column.

C.6

Intervention-protocol Sensitivity

Table 19 reports the full numerical values for the intervention-protocol sensitivity suite. Table 19: Complete intervention-protocol sensitivity metrics. All rows use the mixed observational/interventional split. Protocol

Method

Default Default Default Default Default Default Default

CDIS 500 GIES 500 IGSP 500 SDCD 500 AVICI 500 CauScale 500 TabCausal 500

0.40±0.12 0.48±0.16 0.35±0.10 48.1±19.4 0.055±0.022 237.9±117.7 – – 0.48±0.11 0.43±0.13 0.56±0.11 59.7±31.4 0.069±0.036 251.3±78.4 – – 0.40±0.11 0.44±0.15 0.38±0.11 53.7±26.9 0.062±0.031 232.6±107.2 – – 0.38±0.09 0.33±0.11 0.48±0.11 73.0±25.5 0.084±0.029 246.0±92.0 – – 0.39±0.13 0.66±0.15 0.29±0.12 43.4±17.3 0.050±0.020 189.6±109.6 0.77±0.09 0.42±0.12 0.45±0.11 0.56±0.12 0.39±0.11 43.6±18.7 0.050±0.022 201.3±108.4 0.86±0.05 0.47±0.11 0.54±0.12 0.62±0.14 0.49±0.13 40.5±17.8 0.047±0.020 187.9±96.7 0.85±0.08 0.56±0.14

Fixed setpoint CDIS 500 Fixed setpoint GIES 500 Fixed setpoint IGSP 500 Fixed setpoint SDCD 500 Fixed setpoint AVICI 500 Fixed setpoint CauScale 500 Fixed setpoint TabCausal 500

0.40±0.12 0.48±0.16 0.35±0.10 48.3±19.6 0.056±0.022 237.7±119.2 – – 0.48±0.12 0.43±0.14 0.56±0.12 59.7±31.7 0.069±0.036 253.1±79.2 – – 0.40±0.12 0.43±0.15 0.38±0.11 54.2±26.7 0.062±0.031 233.2±105.7 – – 0.38±0.09 0.33±0.11 0.48±0.11 72.7±26.5 0.084±0.030 242.8±91.2 – – 0.38±0.12 0.65±0.15 0.28±0.11 43.7±17.1 0.050±0.020 189.3±109.8 0.77±0.09 0.41±0.12 0.46±0.10 0.57±0.11 0.39±0.11 43.0±18.6 0.049±0.021 198.7±106.2 0.85±0.05 0.47±0.11 0.54±0.12 0.61±0.14 0.49±0.13 40.6±17.7 0.047±0.020 186.6±94.4 0.85±0.08 0.55±0.14

Parameter shift CDIS Parameter shift GIES Parameter shift IGSP

N F1

Prec.

Rec.

SHD

nSHD

500 0.40±0.12 0.48±0.16 0.35±0.10 48.3±19.6 0.055±0.023 500 0.43±0.11 0.39±0.12 0.50±0.10 64.1±32.1 0.074±0.037 500 0.39±0.11 0.43±0.15 0.38±0.10 54.4±27.0 0.062±0.031

SID

AUROC

240.0±120.0 – 279.8±79.1 – 235.9±107.2 –

AP

– – – Continued on next page

46

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Protocol

C.7

Method

N F1

Prec.

Rec.

SHD

nSHD

SID

AUROC

A P REPRINT

AP

Parameter shift SDCD 500 Parameter shift AVICI 500 Parameter shift CauScale 500 Parameter shift TabCausal 500

0.38±0.09 0.33±0.11 0.47±0.11 0.37±0.13 0.62±0.15 0.27±0.11 0.42±0.10 0.52±0.12 0.36±0.11 0.52±0.12 0.59±0.15 0.47±0.13

73.2±25.5 0.084±0.029 244.8±88.5 – – 44.6±17.7 0.051±0.020 192.5±109.1 0.76±0.09 0.39±0.13 45.2±18.8 0.052±0.022 209.8±106.1 0.84±0.06 0.41±0.11 41.8±18.0 0.048±0.021 191.4±94.1 0.85±0.08 0.53±0.15

Multi-target Multi-target Multi-target Multi-target Multi-target Multi-target Multi-target

CDIS 500 GIES 500 IGSP 500 SDCD 500 AVICI 500 CauScale 500 TabCausal 500

0.38±0.12 0.47±0.16 0.33±0.10 48.9±19.4 0.056±0.022 240.6±120.8 – – – 0.48±0.12 0.44±0.14 0.56±0.11 59.3±31.2 0.068±0.036 252.8±79.0 – 0.13±0.19 0.14±0.22 0.12±0.19 54.6±21.7 0.063±0.025 210.2±114.3 – – 0.38±0.09 0.34±0.11 0.48±0.11 72.4±25.9 0.083±0.030 243.8±93.4 – – 0.39±0.12 0.64±0.14 0.28±0.11 43.7±17.2 0.050±0.020 189.9±109.7 0.77±0.08 0.41±0.12 0.44±0.10 0.51±0.12 0.39±0.11 45.7±18.7 0.053±0.021 210.0±104.7 0.84±0.06 0.44±0.11 0.54±0.12 0.62±0.14 0.49±0.13 40.6±17.6 0.047±0.020 186.3±97.5 0.85±0.08 0.55±0.14

Dense mixed Dense mixed Dense mixed Dense mixed Dense mixed Dense mixed Dense mixed

CDIS 500 GIES 500 IGSP 500 SDCD 500 AVICI 500 CauScale 500 TabCausal 500

0.36±0.10 0.52±0.15 0.29±0.09 47.2±17.9 0.054±0.021 210.4±113.2 – – 0.50±0.11 0.45±0.13 0.57±0.11 56.3±28.0 0.065±0.032 251.3±76.9 – – 0.25±0.16 0.35±0.22 0.20±0.13 51.8±21.0 0.060±0.024 214.3±112.1 – – – 0.40±0.09 0.36±0.12 0.49±0.11 68.3±23.6 0.079±0.027 232.1±88.9 – 0.37±0.12 0.64±0.15 0.27±0.11 44.3±17.1 0.051±0.020 190.3±111.1 0.77±0.08 0.40±0.11 0.42±0.10 0.50±0.13 0.38±0.10 46.7±18.8 0.054±0.022 211.4±104.3 0.82±0.06 0.41±0.11 0.53±0.11 0.61±0.13 0.48±0.12 41.0±16.9 0.047±0.019 186.9±98.2 0.85±0.07 0.53±0.13

High fraction High fraction High fraction High fraction High fraction High fraction High fraction

CDIS 500 GIES 500 IGSP 500 SDCD 500 AVICI 500 CauScale 500 TabCausal 500

0.36±0.10 0.51±0.15 0.28±0.09 47.6±17.7 0.055±0.020 211.5±114.9 – – – 0.51±0.12 0.46±0.14 0.58±0.12 56.7±29.9 0.065±0.034 238.2±82.5 – 0.31±0.10 0.43±0.15 0.25±0.09 52.0±22.4 0.060±0.026 224.1±109.5 – – 0.40±0.10 0.35±0.12 0.49±0.10 70.3±25.5 0.081±0.029 237.3±93.0 – – 0.40±0.13 0.67±0.14 0.30±0.12 42.9±17.2 0.049±0.020 187.2±109.4 0.77±0.09 0.43±0.12 0.48±0.10 0.60±0.12 0.41±0.11 40.8±17.6 0.047±0.020 193.3±107.6 0.86±0.05 0.51±0.11 0.54±0.11 0.62±0.13 0.49±0.12 40.0±16.8 0.046±0.019 181.7±97.0 0.86±0.07 0.56±0.13

Construction Ablation

Table 20 repeats the criterion-level construction-quality scores with standard deviations for the same 20 sampled scenarios summarized in Figure 7. Table 20: Agentic construction ablation quality scores. Stage

Variant

Cases

A0 A1 A2 A3

Seed only + reference pack + planning + graph review

20 20 20 20

Plaus.

Mech.

Bench.

Spec.

Single avg. Struct. div. Sem./mech. div. Avg.

3.650 (0.462) 3.375 (0.626) 3.400 (0.447) 3.375 (0.535) 3.450 (0.492) 3.650 (0.328) 3.275 (0.499) 3.375 (0.358) 3.425 (0.520) 3.431 (0.398) 4.225 (0.302) 4.275 (0.302) 4.000 (0.363) 4.225 (0.302) 4.181 (0.291) 4.375 (0.275) 4.400 (0.262) 4.200 (0.377) 4.400 (0.262) 4.344 (0.262)

4.500 4.500 4.500 4.500

4.500 5.000 5.000 5.000

3.800 3.871 4.371 4.479

Scores are on a 1–5 scale judged by an external LLM evaluator. Single-graph criteria are mean (standard over 20 cases. The final average also includes two batch-level diversity criteria. Higher is better; bold and underline mark the best and second-best values in each score column. Ties share the same mark. deviation)

47

Record · ID 673494 · SHA-256 16635ab33e94e3b4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.