Mono-Z Dark Matter Search with Neural Spline Flows Using CMS Run 2015D Open Data Hitesh Rasineni
1,†
and Chebrolu Bhavishya
2,†
1
2
VIT-AP University, Amaravati, 522241, India Mohan Babu University, Tirupati, 517102, India
arXiv:2607.13771v1 [cs.LG] 15 Jul 2026
Abstract
full kinematic phase space without requiring a miss threshold. The fitted signal hard upper ET We report a search for dark matter (DM) strengths are nonzero in the nominal fits, driven produced in association with a leptonically by a high-E miss background-modelling residual √ T decaying Z boson at s = 13 TeV, us- rather than evidence for a DM signal. Combining CMS Run 2015D open data correspond- ing the µµ and ee channels in a simultaneous ing to an integrated luminosity of 2.32 fb−1 SR+VR binned profile-likelihood fit, where the and simplified-model Monte Carlo simulation. validation region (50 ≤ E miss < 100 GeV) proT Events are selected in the mono-Z → ℓ+ ℓ− vides an independent background normalisation final state in two parallel lepton-flavor chan- constraint, we set observed (expected) 95% CL nels, Z → µ+ µ− and Z → e+ e− , requir- upper limits on the signal-strength parameter ing an opposite-sign same-flavor dilepton pair of µ < 0.0177 (0.0018) for the scalar mediator, with invariant mass 60 < mℓℓ < 120 GeV. A µ < 0.0362 (0.0039) for the vector mediator, miss ≥ 50 GeV, signal region is defined by ET and µ < 0.0498 (0.0069) for the axial-vector miss |∆ϕ(ET , Z)| > 2.5, and njets ≤ 1; a con- mediator, corresponding to cross-section upmiss < 50 GeV provides the trol region with ET per limits of σ95 < 1.76 × 10−9 pb (scalar), SM density model used in the likelihood-ratio σ95 < 9.90 × 10−4 pb (vector), and σ95 < score. Forty kinematic observables are ex- 9.24 × 10−2 pb and 7.89 × 10−3 pb (axial-vector, tracted from MINIAOD and MINIAODSIM two benchmark points). The observed limits inputs, cleaned with physics-motivated bounds exceed the expected by factors of roughly 7–12 and data-driven tail clipping, and reduced to across all three mediator hypotheses and three a 37-dimensional feature vector comprising 35 independent background model variants, driven physics quantities augmented with two binary by a residual background-modelling discrepancy jet-presence indicators. Five Neural Spline in the high-E miss tail (E miss ≥ 100 GeV, 2.4– T T Flows (NSFs) [1, 2] are trained independently: 2.5% of SR events) rather than evidence of a two channel-specific SM flows learn the back- DM signal. To our knowledge this is the first ground density from control-region Drell–Yan application of Neural Spline Flow likelihoodevents (70%/30% train/validation split), and ratio scoring to a mono-Z DM search on CMS three mediator-specific DM flows learn the sig- open Run 2015D data in both the µµ and ee nal density from Monte Carlo samples with channels simultaneously. vector, axial-vector, and scalar s-channel mediators. For each mediator hypothesis h, the per-event test statistic is the log-likelihood ratio 1 Introduction Sh (x) = log p(x | DMh ) − log p(x | SMℓℓ ),
Astrophysical evidence points to a non-baryonic dark component of the Universe, but the microwhich concentrates signal sensitivity across the scopic identity of dark matter (DM) remains † † These authors contributed equally and share first unknown. Among the experimental approaches, authorship. high-energy colliders can produce DM in con1
trolled environments and probe its couplings to Standard Model (SM) particles through missing miss ) recoiling against transverse momentum (ET visible final-state objects [3]. The mono-Z topology—a leptonically decaying Z boson balmiss —offers a relatively clean siganced by ET nature: the dilepton system constrains the Z mass, electroweak backgrounds are smaller than in hadronic mono-jet searches, and the channel is sensitive to simplified models with s-channel mediators that also radiate an on-shell Z boson [4, 5, 6].
score definition used here; Section 1.2 summarizes our contributions and the paper structure. All region definitions, training splits, and numerical results reported below follow our analysis pipeline unless explicitly attributed to external work.
1.1
Analysis objective
The physics goal is to test whether the number and kinematics of events in a signal region are miss , or consistent with SM production of Z + ET Published LHC searches in this final state require a DM contribution. The working signal have progressed from early reinterpretations hypotheses are DM pair production with vector, and multivariate projections to full Run 2 re- axial-vector, and scalar mediators, simulated sults from CMS and ATLAS [7, 8, 9]. These with the same lepton selection as data. miss or related analyses typically bin events in ET Operationally, we partition events into a conkinematic variables, fit parametric or template trol region (CR) and a signal region (SR). The background models, and compare data to sig- CR (E miss < 50 GeV) is dominated by Drell– T nal hypotheses defined within the ATLAS/CMS Yan Z+jets and is used only to train SM denDark Matter Forum benchmark framework [6]. sity models; it is disjoint from the SR by conComplementary mono-Z and mono-Z ′ phe- struction. The SR applies E miss ≥ 50 GeV, T nomenology has also been studied through dilep- |∆ϕ(E miss , Z)| > 2.5, and njets ≤ 1. The T ton mass spectra, angular distributions, and SR imposes no explicit upper E miss requireT miss topologies [10, 11, 12, 13]. resonance-plus-ET ment; in practice the cleaned dataset caps miss at approximately 200 GeV (Section 3), We investigate an alternative strategy that ET miss values keeps the mono-Z event selection but replaces so events enter the fit through ET hand-crafted discriminant variables with a up to this bound. The validation interval miss < 100 GeV) additionally serves likelihood-ratio score built from learned event (50 ≤ ET densities. Neural Spline Flows (NSFs) [1, 2] as a background-normalisation sideband in the model the SM background and three DM me- simultaneous SR+VR fit: it constrains the perdiator hypotheses as continuous distributions channel normalisation nuisances using VR data over a fixed set of cleaned kinematic features; independently of the SR, providing a genuine the per-event score is the log-density difference sideband-based background prediction. between the selected DM hypothesis and the Two SM flows are trained, one on CR doublechannel-specific SM hypothesis. Normalizing muon events and one on CR double-electron flows supply exact density evaluation for corre- events (70%/30% train/validation split in each lated inputs [2, 14], and surjective or simulation- channel). Three DM flows are trained, one each aware extensions have been explored for collider on the full vector, axial-vector, and scalar MC event modeling and anomaly detection [15, 16]. samples, and are shared across lepton channels. Our implementation uses CMS Run 2015D open The branch lepton flavor (11 for electrons, data [17, 18] for SM-dominated control samples 13 for muons) selects the channel but is not and publicly released MonoZToLL MC sam- included in the NSF input vector. ples with vector, axial-vector, and scalar meFor an event x in the SR and mediator hydiators [19, 20, 21, 22] for signal studies, with pothesis h ∈ {vector, axial, scalar}, one DM flow trained per mediator model and identical physics-feature definitions in the µµ Sh (x) = log p(x | DMh ) − log p(x | SMµµ or ee ). and ee channels. (1) Section 2 reviews the mono-Z literature and Large positive values indicate kinematics more related machine-learning methods in more de- characteristic of the trained DM sample than tail. Section 1.1 states the analysis regions and of the corresponding SM CR model. 2
Before interpreting the SR, we check the SMmiss validation band only flow in a 50–100 GeV ET to verify stable channel-specific SM behavior. The primary search constructs six SR score arrays, Shµµ and Shee for the three mediator hypotheses, and fits three combined binned µµ+ee likelihoods, one per signal hypothesis [23].
1.2
and electroweak-boson couplings from ATLAS miss measurements. Z+ET Related signatures broaden the mono-Z program beyond a single on-shell Z recoiling miss in against DM alone. Ref. [10] studied ET association with a dilepton or dijet resonance from a Z ′ boson, highlighting sensitivity in mass regions where standard resonance searches lose efficiency. Ref. [11] pointed out that DM with scalar mediators can induce loop-level box corrections to Drell–Yan production, producing a “monocline” feature near mℓℓ ∼ 2mχ rather than a narrow resonance. Ref. [24] discussed pseudoscalar-portal fermion DM in two-Higgsdoublet setups, where chiral couplings suppress direct detection while preserving thermal relic targets. Ref. [25] compared mono-X searches with direct searches for the mediating particle across several simplified models, finding that mono-X is most competitive when the spectrum is compressed or when electroweak bosons participate directly in the mediator–DM coupling. Ref. [12] showed that lepton angular distributions in the Collins–Soper frame can distinguish spin-0, spin-1, and spin-2 mediator scenarios in mono-Z production and that shape information miss can strengthen coupling beyond inclusive ET limits.
Contributions and paper outline
This paper documents: 1. extraction of 40 event-level features from CMS open data and MonoZToLL MC; 2. a cleaning stage with fixed bounds, quantile-based tail rules, and referenceaware caps for MC samples; 3. CR/SR definitions, NSF training protocol, and likelihood-ratio search strategy; 4. final SR profile-likelihood fits and CLs limits for the Run 2015D dataset.
Section 3 lists data and MC inputs. Section 4 describes feature extraction and cleaning. Section 5 defines CR, validation, and SR event categories. Section 6 summarizes NSF architecture and training. Section 7 details the score, fit, and cross-checks. Section 8 presents distri2.2 butions and limits. Section 9 concludes.
Simplified models, benchmarks, and experimental searches
The ATLAS/CMS Dark Matter Forum report [6] codified simplified models —s-channel vector, axial-vector, scalar, and pseudoscalar 2.1 Collider constraints and mono-Z mediators, t-channel colored mediators, and phenomenology EFT benchmarks—together with recommended Early collider studies showed that effective-field- parameter scans and implementation guidance miss + X analyses. These models theory (EFT) descriptions of DM coupling to for Run 2 ET quarks and gluons can yield limits on interac- underpin the MonoZToLL open-data samples tion scales that complement, and in some kine- used in our signal studies; vector, axial-vector, matic regimes surpass, direct-detection bounds— and scalar mediator samples are each extracted particularly for light DM and spin-dependent and assigned a separate DM flow for hypothesisscattering [3]. The mono-Z channel was pro- specific scoring [19, 20, 21, 22, 6]. posed as an electroweak counterpart to mono-jet On the experimental side, CMS published √ and mono-photon searches: a Z boson recoils mono-Z searches at s = 13 TeV interpretagainst a pair of stable DM particles, with the ing spin-0 and spin-1 mediators, invisible Higgs miss +ℓ+ ℓ− scenarios [8, 9]. Z emitted from initial-state quarks or an in- decays, and other ET ternal line in a simplified model [4]. Ref. [5] The most recent CMS result uses the full Run 2 extended mono-X reinterpretations to events dataset (137 fb−1 ) and sets limits on mediator with a reconstructed Z → ℓ+ ℓ− system, deriv- and DM masses for vector, axial-vector, scalar, ing limits on EFT operators with quark bilinear and pseudoscalar mediators, with comparisons
2
Literature Review
3
to direct-detection cross sections [9]. Ref. [7] projected the mono-Z reach at 13 TeV with a multivariate likelihood discriminant, reporting roughly a factor-of-two improvement over cutand-count methods once background normalization systematics exceed a few percent. Highmass dilepton searches from ATLAS [26, 27] bound Z ′ resonances and four-fermion contact interactions; while not mono-Z analyses, they constrain overlapping Z-associated new-physics scenarios probed in Refs. [10, 13]. Ref. [13] applied CMS open data from Run 1 to dimuon angular distributions at high mµµ , interpreting the data in a mono-Z ′ simplified model and finding agreement with SM Drell–Yan expectations.
2.3
normalizing-flow densities trained separately on control-region data and signal MC. Our analysis sits at the intersection of these threads: we adopt the mono-Z simplified-model context and CR/SR philosophy of the experimental and Forum literature [6, 9], but estimate SM and DM event densities with NSFs and use their log-ratio as the test statistic, following asymptotic inference prescriptions [23] in the signal region. To our knowledge, a published mono-Z DM search combining NSF-based likelihood ratios with CMS open Run 2015D data in both µµ and ee channels has not been reported; the present work documents that pipeline through the final SR likelihood fit.
Density estimation and learning- 3 based search methods
Datasets and Signal Model
SM events come from CMS Run 2015D Normalizing flows map a simple base density to MINIAOD open data: the DoubleMuon pria target distribution through invertible trans- mary dataset (record 24127) [17] and the Doudataset (record 24132) [18], formations, enabling both sampling and exact bleEG primary √ likelihood evaluation [2]. Neural Spline Flows both at s = 13 TeV. These samples provide replace elementwise affine couplings with mono- opposite-sign dilepton events after analyzertonic rational-quadratic splines, increasing ex- level selection; they are used as SM-dominated pressivity while retaining analytic inversion [1]. inputs for CR training and for the SR search. Ref. [14] examined how several flow architec- The Run 2015D dataset corresponds to an intetures scale with input dimensionality on toy grated luminosity of approximately 2.32 fb−1 . HEP-like datasets, a relevant consideration for The extracted local TTrees contain 1,738,984 multivariate collider features. Ref. [15] intro- selected DoubleMuon events split across four duced surjective and stochastic flow layers to output files and 1,418,467 selected DoubleEG handle permutation symmetry, variable parti- events split across eight output files. The stored cle multiplicities, and discrete quantum num- SM data schema contains the 40 physics feabers in matrix-element-level and detector-level tures plus audit-only provenance and qualityevent generation, as well as anomaly-detection control branches used to validate trigger and bstudies. Ref. [16] proposed simulation-assisted tag behavior; the row counts are summarized in likelihood-free anomaly detection (SALAD), Appendix A.1, and the complete feature schema reweighting background simulation to data in with formulas is given in Appendix B.1. sidebands before supervised learning in a signalThe signal inputs are MonoZToLL sensitive region—a hybrid of data-driven and MINIAODSIM samples generated in the √ simulation-based approaches distinct from our same s = 13 TeV era for vector, axial-vector, explicit two-hypothesis density ratio. Closely re- and scalar mediators [6]. The vector sample lated anomaly-detection methods that use side- is record 16630 [19]; the axial-vector inputs band density estimates include ANODE [28] combine records 16575 and 16597 [20, 21]; and CATHODE [29]; both construct signal- the scalar sample is record 16601 [22]. The region classifiers from control-region density extracted selected samples contain 27,602 ratios, a strategy adjacent to our direct NSF vector-mediator events, 45,710 axial-vector log-ratio score. Ref. [7] remains the closest events, and 32,792 scalar-mediator events. The prior mono-Z reference for multivariate discrim- raw MC inputs are organized as ROOT source ination, though it used engineered kinematic files; the axial-vector training set combines inputs and a likelihood function rather than four source files (one for Mχ = 10 GeV, 4
MV = 20 GeV and three for Mχ = 50 GeV, reads the selectedPatTrigger path-name, MV = 200 GeV), totaling about 100,000 pathLastFilterAccepted, and filter-label colraw entries before the final mono-Z selec- lections. When path names are available, events tion. The same 40 physics branches listed are matched to the channel-specific HLT prein Appendix B.1 are written for data and fixes; when path names are empty, the extracMC; the MC output adds only the auxiliary tor falls back to accepted final-filter labels such gen weight branch. Although no explicit as the double-muon or double-electron filter miss bound is imposed in the selection, tokens. Trigger decisions are accumulated in upper ET the data-driven tail clipping applied during TRIGGER AUDIT and the retained event mask miss at approximately is applied before physics-feature writing. For feature cleaning caps ET 200 GeV in the extracted dataset; events the data b veto, the notebooks also decode above this value are absent from the SR score the MiniAOD pat::Jet::pairDiscriVector distribution. The axial-vector signal sample payload using uproot’s AsBinary interpretacombines two physically distinct benchmark tion. A slow decoder discovers the discriminapoints — Mχ = 10 GeV, MV = 20 GeV tor labels, a cached fixed-layout fast path ex(σ = 1.856 pb) and Mχ = 50 GeV, tracts the CSVv2 scores, and a validation step MV = 200 GeV (σ = 0.158 pb) — stored checks the fast decoder against the slow decoder across four ROOT files. Because the two points on sampled events. The medium 2015/76X differ in cross-section by a factor of ∼12, a CSVv2 working point (0.800) is then used to single effective cross-section would not corre- compute n bjets, max btag csvv2, and the bspond to any real benchmark; cross-section veto audit branches. Cleaning is applied only limits are therefore reported separately for to SM data TTrees. The SM cleaner clips fixedeach axial-vector point using the same fitted µ. domain variables, estimates high-quantile caps Generator-level cross-sections were extracted for heavy-tailed variables, and uses the opposite directly from the GenRunInfoProduct branch SM lepton channel as the reference sample when of each MINIAODSIM ROOT file. deriving channel-consistent caps. The DoubleMuon cleaner therefore references DoubleEG parts, and the DoubleEG cleaner references 4 Feature Extraction and DoubleMuon parts. The merged SM bounds are Cleaning exported to Output/sm feature bounds.json; DM MC ROOT files are not cleaned by Raw MINIAOD and MINIAODSIM files are cleaner.py. Instead, the exported SM bounds read directly with uproot, with jagged col- are applied later at DataLoader time, so genuine lections handled in awkward and converted to high-E miss or high-p DM tails are preserved T T fixed per-event NumPy arrays after object se- in the extracted MC files until the training lection. The analysis products are ROOT pipeline applies the common SM feature doEvents TTrees written with uproot.recreate, main. The fixed bounds and tail-rule families mktree, and extend; CSV files are used only used by the cleaning script are summarized in for cut-flow, file-audit, and EDA summary out- Appendix C.1. miss , dilepputs. Per-event quantities include ET ton and Z kinematics, jet and b-tag summaries, hadronic recoil variables, and angular correlations (40 physics features in the extracted The cleaned NSF feature manifest retains schema). The complete list of stored branches 35 extracted physics features and excludes and formulas for derived observables is given in lepton flavor, n bjets, gen weight, rho, Appendix B.1; Table 1 lists the selection cuts n vertices, and raw HT; lepton flavor is implemented in the extraction notebooks. retained only for channel routing and is The DoubleMuon and DoubleEG extrac- not passed to any flow. Two binary jettors use a custom trigger-object decoder presence indicators computed at DataLoader because the 2015 MiniAOD trigger infor- time (has jet1= ⊮[njets ≥ 1], has jet2= mation is not exposed as simple one- ⊮[njets ≥ 2]) are appended to form the final branch-per-HLT-path booleans. The decoder 37-dimensional NSF input. 5
Table 1: Event and object selections applied during feature extraction. Stage
DoubleMuon / DoubleEG data
Trigger
Dataset-level HLT audit; retained events pass the extractor mask Lepton multiplic- Leading same-flavor pair ity Lepton pT pT (ℓ1 ) > 25 GeV, pT (ℓ2 ) > 20 GeV Charge Opposite sign Muon acceptance |ηµ | < 2.4 Electron accep- |ηe | < 2.4, ECAL transition region tance excluded (1.4442 < |ηe | < 1.566) Isolation Relative isolation < 0.15 for both leptons Electron ID Spring15 non-triggering MVA working point Impact parame- Audited as lep1 dxy sig feature; not ter applied as an event selection cut in either data or MC Dilepton mass 60 < mℓℓ < 120 GeV Jets pT > 30 GeV, |η| < 2.4, lepton overlap removed b veto Medium CSVv2 veto, nb = 0, audited in output
MonoZToLL MC Dataset-level trigger audit; retained events pass the extractor mask Best valid same-flavor pair Same Opposite sign |ηµ | < 2.4 |ηe | < 2.4, ECAL transition region excluded Same Not applied Audited as lep1 dxy sig feature; not applied as an event selection cut in either data or MC Same Same n bjets retained for schema parity
Figure 1: Control-region distributions of log p(x | NSF SM) for the electron (left) and muon (right) channels. Overlaid Gaussian fits summarise the location and width of the dominant peak.
6
5
Control and Signal Regions
Region masks are built from miss , Z)|, and n |∆ϕ(ET jets after cleaning:
residual (Section 7.2), not a genuine 8σ discovery. Early stopping uses validation negative miss , ET log-likelihood. The five NSFs share the following hyperparameters: 8 masked autoregressive coupling transforms with rational-quadratic splines (8 bins per spline, randomised feature permutation per transform), conditioner networks with two hidden layers of 256 units each, and a standardnormal base distribution. Training uses the Adam optimiser (η = 3×10−4 , λ = 10−5 weight decay) with a 500-step linear warm-up followed by cosine annealing to ηmin = 10−6 over 200 epochs, a mini-batch size of 2048, and gradient clipping at ∥∇∥2 ≤ 1. Early stopping monitors validation NLL with a patience of 20 epochs and restores the best-epoch weights. In the present training runs, validation NLL continued to decrease at epoch 200 for all five flows; early stopping was not triggered for any flow before reaching MAX EPOCHS. Every best.pt checkpoint therefore corresponds to the final epoch rather than a converged minimum, and we flag potential residual undertraining as a limitation affecting the learned densities and reported limits. The discrete feature n jets is dequantized during training by adding uniform noise U(−0.5, 0.5); raw integer values are used at inference.
miss < 50 GeV. • CR: ET miss < 100 GeV, with • Validation: 50 ≤ ET miss the same |∆ϕ(ET , Z)| and jet cuts as the SR. miss ≥ 50 GeV, |∆ϕ(E miss , Z)| > 2.5, • SR: ET T and njets ≤ 1. miss < 50 GeV) and SR (E miss ≥ The CR (ET T 50 GeV) are disjoint by construction at the miss = 50 GeV boundary. The validation interET miss range and is reserved val lies inside the SR ET for SM closure checks before interpreting the full SR scores in data. The SR imposes no exmiss requirement; in practice the plicit upper ET miss at approximately cleaned dataset caps ET 200 GeV (Section 3).
6
Neural Spline Flow Models
Each NSF maps a feature vector x ∈ Rd to a base density through a composition of coupling layers with rational-quadratic spline transforms [1]. We implement flows in PyTorch using zuko , with hyperparameters (layer count, hidAll five flows operate in the same standardden width, bin count) fixed before training and ized feature space. SM flows fit a per-feature recorded in run manifests. z-score scaler on their respective training splits Training protocol: and save the parameters to a JSON manifest. DM flows reuse the NSF SM mumu scaler rather • NSF SM mumu: CR double-muon data, 70% than fitting new scale parameters, ensuring that train / 30% validation. DM event densities are evaluated in the SM fea• NSF SM ee: CR double-electron data, 70% ture domain. Before standardization, DM MC features are clipped to the SM bounds exported train / 30% validation. in Output/sm feature bounds.json; this premiss tails while preventing • NSF DM vector: full vector-mediator MC, serves genuine high-ET out-of-domain extrapolation of the learned SM 70% train / 30% validation, all flavors. density. • NSF DM axial: full axial-vector MC, 70% The total training campaign therefore contrain / 30% validation, all flavors. sists of five independent NSF runs: two channel• NSF DM scalar: full scalar-mediator MC, specific SM flows and three mediator-specific DM flows. Closure checks use the 50–100 GeV 70% train / 30% validation, all flavors. validation band to verify the channel-specific † The significance is capped at 8.0 due to SM models before SR score arrays are interfloating-point underflow of the chi-square sur- preted. The axial-vector µµ score distribuvival function at large q0 . The underlying q0 val- tion shows a localised bin-wobble pattern near ues (233–327) indicate a background-modelling Sh ≈ 0–50, consistent with residual NSF spline 7
Figure 2: Pre-unblinding validation-region score distributions comparing observed data with normalised dark-matter Monte Carlo for the electron (top row) and muon (bottom row) channels. miss < 100 GeV band rather than backgroundThese plots test signal sensitivity in the 50 ≤ ET modelling closure.
8
Figure 3: Post-fit SR score distributions in the electron (top row) and muon (bottom row) channels for each mediator hypothesis. Black points are observed SR data; blue histograms are best-fit SM background prediction obtained from the simultaneous SR+VR profile-likelihood fit; red curves show the best-fit S + B model.
9
Table 2: Summary of the final simultaneous SR+VR profile-likelihood results for the three dark matter mediator hypotheses. The VR constrains the background normalisation independently; the SR is used for signal search. The 68% and 95% expected bands are quoted for the expected 95% CL upper limit on µ. Cross-section limits use the generator cross-sections; the two axialvector values correspond to the two benchmark points. Mediator Scalar Vector Axial-vector
µ̂ 1.41 × 10−2 3.08 × 10−2 4.24 × 10−2
σµ̂ 0.00136 0.00245 0.00357
p0
Z
µ95 obs
µ95 exp
68% exp. band / 95% exp. band
95 [pb] σobs
≃0 ≃0 ≃0
8.0†
0.0177 0.0362 0.0498
0.0018 0.0039 0.0069
[0.00156, 0.00163] / [0.00154, 0.00166] [0.00326, 0.00364] / [0.00309, 0.00382] [0.00560, 0.00659] / [0.00517, 0.00715]
1.76 × 10−9 9.90 × 10−4 9.24 × 10−2 / 7.89 × 10−3
8.0† 8.0†
instability from non-convergence; this feature appears in all three fit variants and is therefore a property of the trained flow rather than an artifact of any single background model.
MC contributes to both SR and VR components with the same µ. Asymptotic formulae [23] are used for the discovery test statistic and CLs upper limits.
7
7.1
Likelihood-Ratio Search
The search uses Sh as the per-event test statistic for each mediator hypothesis (Eq. (1)). For µµ SR events, vector, axial-vector, and scalar DM log densities are compared with NSF SM mumu; for ee SR events the same three DM log densities are compared with NSF SM ee. This yields six score arrays in total. We perform three simultaneous SR+VR binned profile-likelihood fits for signal strength µ, one for each DM hypothesis. Each fit combines the µµ and ee score histograms from both the SR and VR simultaneously: 2 2 − 21 θee , ln L = ln LSR + ln LVR − 12 θµµ
(2)
where LSR and LVR are binned Poisson likelihoods over the score histograms in each region, the signal strength µ and the per-channel normalisation nuisances θµµ , θee are shared between regions, and the background templates in both regions are derived from a single-step VR→SR shape transfer: the SM VR score histogram is renormalised to the respective region yield and used as the nominal background prediction for that region. The observed data histograms are the raw SR and VR score counts, which differ from the renormalised VR background template by real Poisson fluctuations and by genuine shape differences between the VR and SR score distributions. The VR component constrains the background normalisation using ∼6,970 (µµ) and ∼6,644 (ee) VR events independently of the SR, providing a genuine sideband-based background prediction. Signal
Systematic uncertainties
The profile-likelihood fit includes independent Gaussian-constrained normalisation nuisance parameters for the µµ and ee channels with a prior width of 5% (σ = 0.05). In the simultaneous SR+VR fit the normalisation nuisances are constrained additionally by the VR data; the posterior uncertainty on θnorm is approximately 0.17 (compared to the prior of 1.0 in sigma units), corresponding to an effective normalisation constraint of ∼0.85% from the VR statistics alone. Integrated luminosity (propagated). The Run 2015D dataset corresponds to L = 2.32 ± 0.037 fb−1 (uncertainty 1.6%), taken from the CMS precision luminosity measurement [30]. A correlated log-normal luminosity nuisance is applied across both channels. Lepton efficiency (propagated). Combined identification, reconstruction, and trigger efficiency uncertainties of approximately 3% per channel are applied as independent per-channel log-normal nuisances, following the treatment in the CMS Run 2 mono-Z search at the same centre-of-mass energy [8]. Lepton momentumscale uncertainties of 1% (µ) and 2%/5% (electrons, barrel/endcap) are included as additional shape nuisances on the lepton pT distributions. MET resolution and pileup (not propmiss resolution and agated). Evaluating ET pileup-reweighting uncertainties requires rescoring SR, VR, and CR events with shifted
10
Figure 4: Signal-region NSF score distributions for scalar, vector, and axial-vector mediator hypotheses. Data from the electron and muon channels are shown together with the corresponding dark matter signal templates. inputs, which constitutes new analysis work not so the tail population appears as a systematic performed in this study. These sources are ac- upward fluctuation in the fit. knowledged as a gap in the systematic budget. To characterise this effect, we tested linear miss -dependent models of the and quadratic ET CR/VR score trend, validating each against held-out VR bins before extrapolating to the SR tail (Section 7.3). Neither functional form reliably predicts the observed tail behaviour: the linear model underpredicts the true SRtail mean score by 90–115 units, while the quadratic model overpredicts by 300–330 units. The limited tail statistics (160–181 events across 100 GeV) do not support more flexible extrapoHigh-ETmiss tail modelling residual lation.
Combined systematic budget. Treating luminosity as fully correlated across channels and lepton efficiency as per-channel independent, the combined propagated external system√ atic is approximately 1.62 + 32 + 32 ≈ 4.5% per channel, on top of the 0.85% data-driven normalisation constraint. The dominant uncermiss tail modtainty is the unresolved high-ET elling residual (Section 7.2).
7.2
miss ≥ 100 GeV constitute 2.4– Events with ET 2.5% of the SR (∼160–181 events across both channels) but carry mean NSF scores 140–195 units higher than the VR bulk — a shift exceeding the score distribution’s own standard deviation of 40–60 units. The single-step VR background has negligible support at these scores,
The VR-shape background is therefore used as nominal, and the resulting inflation of µ̂, q0 , and the observed-to-expected limit ratios reflects this unresolved modelling residual rather than evidence of DM production.
11
7.3
NSF CR→SR extrapolation vali- (CR), validation-region (VR), and signal-region (SR) samples. The Standard Model density dation
flows were trained only on CR events with To validate that the NSF SM density, trained miss < 50 GeV. The final statistical interpretamiss < 50 GeV), ET only on the control region (ET miss ≥ 50 GeV, tion uses the full SR selection, ET extrapolates in a characterised way into the miss |∆ϕ(ET , Z)| > 2.5, and njets ≤ 1. Although signal region, we measured the mean NSF score miss bound is imposed in miss across CR and VR bins for no explicit upper ET as a function of ET miss at the selection, the cleaned dataset caps ET all three mediator hypotheses and both lepton approximately 200 GeV, and the VR interval channels (six combinations), then tested two miss < 100 GeV) is retained as a val(50 ≤ ET extrapolation models against held-out data. idation and stress-test sample rather than as the nominal background template for the final Procedure. Events in the CR and VR were limits. miss in 10 GeV intervals. For each binned by ET For each mediator hypothesis the NSF score bin the mean NSF score ⟨Sh ⟩ was computed. miss VR bins (80–100 GeV) The two highest-ET Sh (x) = log p(x | DMh ) − log p(x | SMℓℓ ) (3) were held out as a validation set. Linear and miss were fit to the rewas constructed separately for the vector, quadratic functions of ET axial-vector, and scalar mediator hypotheses. maining bins and extrapolated to the SR tail miss Large positive values correspond to signal-like (ET ≥ 100 GeV), where the true mean score events, while negative values indicate SM-like was measured independently. kinematics. Control-region closure tests show stable Results. The linear fit undershot the held-out channel-specific behaviour of the trained SM VR bins by 0.3–0.6σ (systematic across all six flows. The CR distributions of log p(x | miss ∼ combinations). When extrapolated to ET 150 GeV it overshot the true SR-tail mean score SM) peak near ∼ 125 with similar widths by 90–115 units. The quadratic fit overshot in the µµ and ee channels, consistent with the held-out bins by 1.1–3.3σ (in the opposite a well-converged density model on Drell–Yandirection from the linear case) and overshot the dominated events. Validation-region score distributions show true SR-tail mean by 300–330 units; the positive curvature in all six fits indicates the quadratic clear separation between observed data and miss rise simulated dark matter events for all mediator form was dominated by the steep low-ET miss . hypotheses. The strongest discrimination is rather than the true behaviour at higher ET observed for the scalar mediator, whose signal distribution occupies a largely distinct region of Conclusion. Neither functional form reliably score space. Consistent behaviour is observed extrapolates the CR-trained NSF density into miss SR tail. Combined with the in both lepton channels, demonstrating that the the high-ET NSF score generalises across detector signatures limited tail statistics (160–181 events across and event topologies. 100 GeV), this demonstrates that the NSF’s CRThe SR score distributions are shown in trained density does not extrapolate smoothly miss signal Fig. 4. In all cases, the observed data are conor predictably into the sparse high-ET centrated at lower score values, whereas simregion. This validates and motivates the treatulated dark matter events populate the highment of the resulting background modelling score tail. The scalar hypothesis produces the residual as an explicit, quantified limitation largest signal-background separation, followed (Section 7.2) rather than an uncorrected predicby the vector and axial-vector models. tion. A binned profile-likelihood fit was performed using the NSF score distributions. Initial histograms were constructed using 30 score bins 8 Results and subsequently merged to ensure a minimum The normalizing-flow likelihood-ratio discrimi- nominal background yield of 20 events per bin. nant was evaluated using disjoint control-region The nominal background shapes are derived 12
Figure 5: Asymptotic CLs scans for the scalar, vector, and axial-vector mediator hypotheses from the simultaneous SR+VR fit. The VR constrains the background normalisation; the SR provides the signal search sensitivity. The horizontal dashed line marks the 95% CL exclusion threshold; vertical lines indicate the observed and expected upper limits on µ. from a single-step VR→SR transfer, while the final statistical interpretation uses a simultaneous SR+VR profile-likelihood fit with independent 5% normalization nuisance parameters in the µµ and ee channels. The validation region acts as a normalisation sideband and constrains the background prediction through nuisanceparameter profiling, providing an independent cross-check of the SR-only result shown in Appendix E.5. A pure VR-extrapolated shape alternative was also investigated but failed closure in the high-score tail; the corresponding VR-only post-fit distributions are provided in Appendix E.4 as a diagnostic study and are not used for the final statistical interpretation. The fitted score distributions for all mediator hypotheses and both lepton channels are shown in Fig. 3. In every case the observed data show a pull relative to the VR-based background prediction, most pronounced in the high-score tail (Section 7.2). The discovery test statistic q0 is
large (233–327 across the three hypotheses; Table 2), formally corresponding to a significance capped at Z = 8.0 by floating-point underflow; miss we attribute this to the characterized high-ET background-modelling residual rather than a dark matter signal (Sections 7.2–7.3). Upper limits on the signal strength parameter µ were derived using the asymptotic CLs method. The resulting limits are summarized in Table 2, while the corresponding confidencelevel scans are shown in Fig. 5. The scalar mediator provides the strongest exclusion, yielding an observed 95% confidence-level upper limit of µ < 0.0177. The vector and axial-vector hypotheses yield observed limits of µ < 0.0362 and µ < 0.0498, respectively. For all three mediator hypotheses the observed limits exceed the median expected limits by factors of roughly 7–12 across the nominal and cross-check background constructions, miss tail modreflecting the unresolved high-ET
13
elling residual described in Section 7.2. The superior performance of the scalar model is consistent with its significantly stronger separation in NSF score space. The axial-vector hypothesis exhibits the weakest sensitivity owing to the greater overlap between signal and background score distributions.
9
Summary and Outlook
We presented a mono-Z DM search that replaces hand-crafted cut flows with likelihoodratio scores from Neural Spline Flows trained on CR data and DM MC. The workflow— open-data extraction, audited cleaning, disjoint CR/SR design, and three combined µµ+ee fits from five trained flows—provides a reproducible template for density-based searches at the LHC. The final simultaneous SR+VR fits yield observed 95% CL limits of µ < 0.0177 (scalar), µ < 0.0362 (vector), and µ < 0.0498 (axialvector), corresponding to the cross-section limits reported in Table 2. The observed-tomiss expected gaps are driven by the high-ET tail modelling residual rather than by a dark matter signal. A pure VR-extrapolated shape alternative fails closure in the high-score tail (Appendix E.4) and is not used for the final limits.
14
Appendix A
Dataset and Extraction Summary
A.1
Rows read and extracted
Table 3 summarizes the row counts used after extraction. For Run 2015D data, raw rows read are aggregated from the per-file audit CSVs under Data/Extracted. For DM MC, the source files are ROOT files; the Rows read column reports the total raw entries processed before extraction, and the extracted rows are the selected event counts in the output TTrees and EDA summaries. Table 3: Dataset-level rows read and rows extracted for the analysis inputs. Dataset source
Source files
Rows read
Rows extracted
1068 3568 1 4 1
51,342,919 92,865,644 ∼50,000 ∼100,000 ∼50,000
1,738,984 1,418,467 27,602 45,710 32,792
DoubleMuon data DoubleEG data DM vector MC DM axial-vector MC DM scalar MC
A.2
HLT trigger paths and b-tag configuration
Table 4 summarizes the trigger and b-tag settings recorded by the extraction notebooks. Trigger decisions are read from selectedPatTrigger path-name and final-filter collections when available; the filter tokens are used as the fallback audit path for 2015 MiniAOD files with empty path-name payloads. The data extractors apply the medium CSVv2 b-veto, while the DM MC extractor retains n bjets only for schema compatibility. Table 4: HLT trigger prefixes, fallback filter tokens, and b-tag configuration used during event extraction. Channel
HLT prefixes
Fallback accepted-filter tokens
b-tag ment
DoubleMuon data
HLT_Mu17_TrkIsoVVL_Mu8_TrkIsoVVL _v, HLT_Mu17_TrkIsoVVL_TkMu8_Trk IsoVVL_v, HLT_IsoMu20_v, HLT_IsoT kMu20_v
hltDiMuonGlb17Glb8RelTrkIsoFiltered0 p4, hltDiMuonGlb17Trk8RelTrkIsoFilte red0p4, hltDiMuonGlb17Glb8RelTrkIsoF iltered0p4DzFiltered0p2, hltDiMuonGlb 17Trk8RelTrkIsoFiltered0p4DzFiltered 0p2
CSVv2 medium veto, WP = 0.800
DoubleEG data
HLT_Ele17_Ele12_CaloIdL_TrackIdL _IsoVL_DZ_v, HLT_Ele23_Ele12_Ca loIdL_TrackIdL_IsoVL_DZ_v, HLT_ Ele27_WPLoose_Gsf_v
hltEle17Ele12CaloIdLTrackIdLIsoVLDZF ilter, hltEle23Ele12CaloIdLTrackIdLIs oVLDZFilter, hltEle27WPLooseGsfTrackI soFilter, hltEle27noerWPLooseGsfTrac kIsoFilter
CSVv2 medium veto, WP = 0.800
MonoZToLL MC
Dataset-level trigger definition; named HLT paths are audited when TriggerResults are present.
All simulated events retained by the uproot-only trigger decoder.
n_bjets retained; no veto applied.
15
treat-
A.3
Signal MC parameter points
Table 5 lists the signal samples used for the three DM density models. The parameter values are taken from the CERN Open Data sample names in the local bibliography. Table 5: MonoZToLL signal MC parameter points and extracted event counts. Mediator sample
CERN record / DOI
Vector, V Mx-1 Mv-500 gDMgQ-1 Axial-vector, A Mx-10 Mv-20 gDMgQ-1 Axial-vector, A Mx-50 Mv-200 gDMgQ-1 Scalar, EWK Scalar Mx-100 Lambda-3000
Record 16630 / 10.7483/OPENDATA.CMS.YTLH.E0N7 Record 16575 / 10.7483/OPENDATA.CMS.YB3Q.XTXY Record 16597 / 10.7483/OPENDATA.CMS.34IE.KN6I Record 16601 / 10.7483/OPENDATA.CMS.NW7F.NFGG
16
mχ [GeV]
Mediator scale [GeV]
Extracted events
1 10 50 100
mV = 500 mV = 20 mV = 200 Λ = 3000
27,602 part of combined 45,710 part of combined 45,710 32,792
B
Feature Definitions
B.1
Physics feature branches
The extraction notebooks write the 40 physics branches listed in Tables 6 and 7 to the output Events TTrees. The DoubleMuon, DoubleEG, and MonoZToLL DM extractors use the same branch set; data outputs add audit/provenance branches, while DM MC adds gen weight. We use ∆ϕ(a, b) = |((a − b + π) mod 2π) − π|, mZ = 91.1876 GeV, and ϵ = 10−6 for protected divisions. Table 6: Physics feature branches written by the extraction notebooks, part 1 of 2. Derived features include the formula used in the analysis code. Branch
Category
Definition or formula
met pt
MET
met phi min dphi met jets
MET Jet/MET angle
dphi met jet1
Jet/MET angle
dphi met jet2
Jet/MET angle
met over sqrtHT hadronic recoil pt
MET/recoil Recoil
Missing transverse momentum magnitude read from the MiniAOD MET object. Azimuthal angle of the MiniAOD MET object. minj ∆ϕ(ϕMET , ϕj ) over selected jets; filled with π when no selected jet is present. ∆ϕ(ϕMET , ϕj1 ) for the leading selected jet; absent slot filled with the extraction sentinel, imputed at DataLoader time to 0, and masked with has jet1. ∆ϕ(ϕMET , ϕj2 ) for the subleading selected jet; absent slot filled with the extraction sentinel, imputed at DataLoader time to 0, and masked with has √ jet2. miss / H for H > 5 GeV, otherwise 0. ET T T miss cos ϕ Z With hx = −(ET MET + pT cos ϕZ ) and q miss Z hy = −(ET sin ϕMET + pT sin ϕZ ), h2x + h2y . q 2 Z2 Z2 max[(E1 + E2 )2 − (pZ x + py + pz ), 0]. p 2 2 (px1 + px2 ) + (py1 + py2 ) , with pxi = pTi cos ϕi , pyi = pTi sin ϕi . Z Z asinh(pZ z /pT ) for pT > 1 GeV, otherwise 0. Z Z atan2(py , px ). ∆ϕ(ϕMET , ϕZ ). miss / max(|pZ |, ϵ), clipped to [0, 10] at extraction. ET T miss − pZ |. |ET T |mℓℓ − mZ |. hx cos ϕZ + hy sin ϕZ . −hx sin ϕZ + hy cos ϕZ . u∥ / max(|pZ T |, ϵ), clipped to [−10, 10]. u⊥ / max(|pZ T |, ϵ), clipped to [−10, 10]. Leading selected lepton transverse momentum. Leading selected lepton pseudorapidity. Subleading selected lepton transverse momentum. Subleading selected lepton pseudorapidity.
mll
Dilepton
pt Z eta Z pseudo phi Z dphi met Z met over ptZ abs met minus ptZ abs mll minus mZ u parallel u perp u parallel over ptZ u perp over ptZ lep1 pt lep1 eta lep2 pt lep2 eta
Dilepton/Z Dilepton/Z Dilepton/Z MET/Z angle MET/Z MET/Z Dilepton Recoil Recoil Recoil/Z Recoil/Z Lepton Lepton Lepton Lepton
B.2
NSF input feature order
Table 8 gives the 37-dimensional input order used when loading the trained NSF checkpoints and scoring events. This order is read from Output/results/scoring audit.json; has jet1 and has jet2 are computed from n jets before clipping.
17
Table 7: Physics feature branches written by the extraction notebooks, part 2 of 2. Branch
Category
deltaR ll deltaphi ll lepton flavor
Dilepton angle Dilepton angle Channel
lep1 reliso lep2 reliso
Lepton ID Lepton ID
cos theta star
Dilepton angle
lep1 dxy sig dphi met lep1 n jets
Lepton ID MET/lepton angle Jets
n bjets
Jets
jet1 pt
Jets
jet1 eta
Jets
jet2 pt
Jets
jet2 eta
Jets
HT n vertices rho
Jets Event Event
DataLoader-computed features: Jet indicator has jet1
Jet indicator
has jet2
Definition or formula p (η1 − η2 )2 + ∆ϕ(ϕ1 , ϕ2 )2 . ∆ϕ(ϕ1 , ϕ2 ). Absolute PDG lepton ID: 13 for muon-pair events and 11 for electron-pair events. Leading lepton relative isolation from MiniAOD isolation components. Subleading lepton relative isolation from MiniAOD isolation components. Cosine of the negative lepton angle in the reconstructed dilepton rest frame, computed by boosting the negative lepton into the Z-candidate rest frame and projecting onto the reconstructed Z direction. Leading lepton transverse impact-parameter significance, |dxy |/σ(dxy ). ∆ϕ(ϕMET , ϕℓ1 ). Number of selected jets with pT > 30 GeV, |η| < 2.4, and lepton-overlap removal. Number of selected jets passing the medium CSVv2 working point in data; retained in DM MC for schema parity. Leading selected-jet transverse momentum; absent slot filled with 0 and masked with has jet1. Leading selected-jet pseudorapidity; absent slot filled with the extraction sentinel, imputed at DataLoader time to 0, and masked with has jet1. Subleading selected-jet transverse momentum; absent slot filled with 0 and masked with has jet2. Subleading selected-jet pseudorapidity; absent slot filled with the extraction sentinel, imputed at DataLoader time to 0, and masked with has jet2. Scalar sum of selected-jet transverse momenta. Number of reconstructed primary vertices. Event energy-density variable read from the MiniAOD fixed-grid ρ branch. ⊮[njets ≥ 1]; computed at DataLoader time, not stored in the extracted ROOT TTree. Appended to the NSF input vector to allow the flow to condition on jet presence. ⊮[njets ≥ 2]; computed at DataLoader time. Appended to the NSF input vector to allow the flow to condition on second-jet presence.
Table 8: NSF input feature order used for SM and DM scoring. Index 1 2 3 4 5 6 7 8 9 10 11 12 13
Feature met pt met phi min dphi met jets dphi met jet1 dphi met jet2 met over sqrtHT hadronic recoil pt mll pt Z eta Z pseudo phi Z dphi met Z met over ptZ
Index 14 15 16 17 18 19 20 21 22 23 24 25 26
Feature abs met minus ptZ abs mll minus mZ u parallel u perp u parallel over ptZ u perp over ptZ lep1 pt lep1 eta lep2 pt lep2 eta deltaR ll deltaphi ll lep1 reliso
18
Index 27 28 29 30 31 32 33 34 35 36 37
Feature lep2 reliso cos theta star lep1 dxy sig dphi met lep1 jet1 pt jet1 eta jet2 pt jet2 eta has jet1 has jet2 n jets
C
Data Cleaning
C.1
Fixed domain bounds
Table 9 lists the fixed-domain bounds applied by the SM cleaner. Features not shown in the fixed-bound table are either pass-through variables or receive only the tail-rule treatment in Table 11.
Table 9: Fixed bounds applied during SM cleaning. Feature(s)
Bound
met phi, phi Z min dphi met jets, dphi met jet1, dphi met jet2, dphi met Z, dphi met lep1, deltaphi ll mll eta Z pseudo lep1 eta, lep2 eta, jet1 eta, jet2 eta met over ptZ u parallel over ptZ, u perp over ptZ abs mll minus mZ, deltaR ll, n jets, n bjets, HT, n vertices, rho lep1 reliso, lep2 reliso cos theta star lepton flavor, gen weight
[−π, π] [0, π] [60, 120] GeV [−5, 5] [−2.4, 2.4] [0, 10] [−10, 10] lower bound 0 [0, 0.15] [−1, 1] pass-through
C.2
Data-derived tail caps
Table 10 lists the numerical caps stored in Output/sm feature bounds.json for variables governed by data-derived tail rules. Positive rules clip to [0, cap], while symmetric rules clip to [−cap, +cap].
Table 10: Data-derived tail caps used by the SM feature-domain cleaner. Feature met pt met over sqrtHT hadronic recoil pt pt Z abs met minus ptZ u parallel u perp lep1 pt lep2 pt lep1 dxy sig jet1 pt jet2 pt HT
Lower bound
Upper bound
0.000 0.000 0.000 0.000 0.000 -563.434 -158.807 0.000 0.000 0.000 0.000 0.000 0.000
200.000 20.516 564.334 556.310 517.988 563.434 158.807 433.151 182.855 10.000 629.442 422.450 1071.170
19
Rule family positive positive positive positive positive symmetric symmetric positive positive positive, reference-min positive positive positive
C.3
Tail-rule families
Table 11: High-quantile tail rules used by the SM cleaner. A positive rule bounds the feature below by zero and above by the estimated cap; a symmetric rule applies [−cap, +cap].
Feature(s)
Rule type
Minimum cap
met pt, abs met minus ptZ met over sqrtHT hadronic recoil pt, pt Z u parallel, u perp lep1 pt, jet1 pt lep2 pt, jet2 pt lep1 dxy sig HT
positive positive positive symmetric positive positive positive, cross-channel reference-min positive
200 10 300 80 150, 100 80, 50 10 150
20
D
Region Definitions and Event Yields
D.1
Region selection criteria
Table 12 records the CR, VR, and SR masks from Output/results/scoring audit.json. The VR is nested in the SR kinematic definition and is used as a 50–100 GeV closure and stress-test region. Table 12: Analysis region definitions used during scoring.
D.2
Region
Selection
CR VR SR
miss < 50 GeV ET miss < 100 GeV, |∆ϕ(E miss , Z)| > 2.5, n 50 ≤ ET jets ≤ 1 T miss miss , Z)| > 2.5, n ET ≥ 50 GeV, |∆ϕ(ET jets ≤ 1
Event counts per region and channel
Table 13 summarizes the region counts stored in the scoring audit. DM MC is scored in both SM channel scaler spaces, so the mediator-specific VR and SR counts are identical for the µµ and ee scoring passes. Table 13: Event counts after cleaning and region assignment. Sample SM µµ data SM ee data DM vector MC DM axial-vector MC DM scalar MC
Total
CR
VR
SR
1,738,984 1,418,467 27,602 45,710 32,792
1,709,765 1,392,023 – – –
6,970 6,644 5,608 9,983 378
7,151 6,804 16,259 18,443 25,884
21
E
Fit Configuration and Numerical Results
E.1
Fit configuration
Table 14 lists the final simultaneous SR+VR fit settings recorded in paper/json/sr+vr/fit summary.json and associated output files. The fit combines the µµ and ee score histograms for each mediator hypothesis. Table 14: Profile-likelihood fit configuration.
Setting
Value
Statistical method Initial score bins Minimum background per merged bin Signal-strength scan points Fit regions VR normalisation constraint Normalization nuisance width Random seed Fit packages
χ2 asymptotic CLs approximation 30 20.0 67 SR + VR simultaneous Per-channel θnorm , prior width 0.05 0.05 per channel 42 numpy 1.26.4, scipy, iminuit 2.32.0
E.2
Full fit results
Table 15 reports the numerical results from paper/json/sr+vr/fit summary.json. All three miss signal-plus-background fits converged; the large q0 values are attributed to the high-ET background-modelling residual discussed in Section 7.2. Table 15: Observed and expected fit results for the three mediator hypotheses from the simultaneous SR+VR profile-likelihood fit.
Mediator Vector Axial-vector Scalar
E.3
µ̂
σµ
µ95 obs
µ95 exp
68% exp. band
95% exp. band
0.0308 0.0424 0.0141
0.00245 0.00357 0.00136
0.0362 0.0498 0.0177
0.0039 0.0069 0.0018
[0.00326, 0.00364] [0.00560, 0.00659] [0.00156, 0.00163]
[0.00309, 0.00382] [0.00517, 0.00715] [0.00154, 0.00166]
Background estimation method note
The simultaneous SR+VR fit uses a single-step VR→SR shape transfer: the SM VR score histogram is renormalised to the region yield and used as the nominal background template for that region. The VR component of the likelihood constrains the per-channel normalisation nuisances θnorm using ∼6,970 (µµ) and ∼6,644 (ee) VR events; the posterior uncertainty on θnorm is ∼0.17 in sigma units, significantly tighter than the 5% Gaussian prior alone. Signal MC contributes to both SR and VR with the same signal strength µ, so the VR signal contamination is accounted for. This approach avoids a pure SR-direct circularity while keeping the SR score shape as the primary discriminant. Bins with zero nominal background are excluded from the likelihood.
22
Table 16: VR-extrapolated validation fit. The large positive fitted signal strengths and weakened observed limits indicate VR-to-SR shape bias; these numbers are not used for the final interpretation. Mediator µ̂ σµ Z µ95 µ95 68% exp. band / 95% exp. band exp obs Vector Axial-vector Scalar
E.4
0.0392 0.0587 0.0160
0.00286 0.00434 0.00132
8.0 8.0 8.0
0.0454 0.0694 0.0195
0.0043 0.0073 0.0018
[0.00359, 0.00399] / [0.00339, 0.00421] [0.00597, 0.00701] / [0.00554, 0.00761] [0.00151, 0.00163] / [0.00145, 0.00168]
VR-extrapolated validation fit
miss < 100 GeV) score As a sideband stress test, an alternative fit used the SM VR (50 ≤ ET histograms, normalized to the SR yield, as the nominal background template. This procedure does not close in the high-score tail: the VR template underpredicts the SR tail, and the signal-plus-background fit absorbs the resulting shape difference as a positive signal strength. Table 16 therefore records the VR-extrapolated fit as a diagnostic rather than as the final limit result. With a signal-strength scan extended to bracket µ̂ for all three mediators, the VR-extrapolated fit yields observed limits 9.5–10.8× weaker than expected, consistent in direction and magnitude miss tail residual, with the nominal fit’s 7.2–9.8× gap (Table 2). This corroborates that the high-ET not a background-construction artifact specific to one method, drives the observed/expected discrepancy.
E.5
SR-only validation fit
As an additional robustness check, the likelihood fit was repeated using only the SR score distributions without the VR normalisation constraint. The fitted signal strengths remain miss residual, while the resulting positive with capped Z = 8.0, as expected from the same high-ET limits are compatible with those obtained from the nominal simultaneous SR+VR fit. The corresponding CLs scans and post-fit distributions are shown in Figs. 8 and 9. Table 17: SR-only validation fit. The fitted signal strengths remain positive with capped Z = 8.0, miss residual, while the resulting limits are compatible with as expected from the same high-ET those obtained from the nominal simultaneous SR+VR fit. The Z values are capped as in Table 2. Mediator Vector Axial-vector Scalar
µ̂
σµ
Z
µ95 obs
µ95 exp
68% exp. band / 95% exp. band
3.78 × 10−2
0.00301 0.00459 0.00183
8.0 8.0 8.0
0.0443 0.0669 0.0230
0.0041 0.0075 0.0019
[0.00335, 0.00373] / [0.00322, 0.00393] [0.00606, 0.00713] / [0.00562, 0.00779] [0.00164, 0.00172] / [0.00159, 0.00176]
5.57 × 10−2 1.85 × 10−2
23
F
Additional Validation Figures
F.1
VR-extrapolated CLs scans
Figure 6: CLs scans for the VR-extrapolated validation fit. Observed CLs crosses the 95% threshold for all three mediators once the scan range brackets µ̂; see Table 16.
24
F.2
VR-extrapolated post-fit distributions
Figure 7: Post-fit score distributions for the VR-extrapolated validation fit. The signal-plusbackground component compensates for the VR-to-SR high-score tail mismatch, confirming that the fitted positive signal strength is a background-modeling bias rather than evidence for DM in SM data.
25
F.3
SR-only CLs scans
Figure 8: Asymptotic CLs scans for the scalar, vector, and axial-vector mediator hypotheses from the SR-direct fit. The horizontal dashed line marks the 95% CL exclusion threshold; vertical lines indicate the observed and expected upper limits on µ.
26
F.4
SR-only post-fit distributions
Figure 9: Post-fit SR score distributions in the electron (top row) and muon (bottom row) channels for each mediator hypothesis. Black points are observed SR data; blue histograms are the SR-direct SM background template; red curves show the best-fit S + B model.
27
√ collisions at s = 13 TeV. Eur. Phys. J. C, 81:13, 2021.
References
[1] Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. Neural [10] Marcelo Autran, Kevin Bauer, Tongyan spline flows. Advances in Neural InformaLin, and Daniel Whiteson. mono-Z′ : tion Processing Systems, 32, 2019. searches for dark matter in events with a resonance and missing transverse energy. [2] George Papamakarios, Eric Nalisnick, Phys. Rev. D, 92:115014, 2015. Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. Normaliz[11] Wolfgang Altmannshofer, Patrick J. Fox, ing flows for probabilistic modeling and Roni Harnik, Graham D. Kribs, and Nirinference. J. Mach. Learn. Res., 22:1–64, mal Raj. Dark matter signals in dilepton 2021. production at hadron colliders. Phys. Rev. D, 91:015005, 2015. [3] Jessica Goodman, Masahiro Ibe, Arvind Rajaraman, William Shepherd, Tim M. P. Tait, and Hai-Bo Yu. Constraints on [12] Daneng Yang and Qiang Li. Probing the dark sector through mono-Z boson leptonic dark matter from colliders. Phys. Rev. D, decays. Phys. Rev. D, 97:015022, 2018. 82:116010, 2010. [4] Nicole F. Bell, Ahmad J. Galea, James B. [13] S. Elgammal. Angular distribution study Dent, Thomas D. Jacques, Lawrence M. for high mass dimuon pairs in CMS open Krauss, and Thomas J. Weiler. Searching 2012 data and for mono-Z′ model. arXiv for dark matter at the LHC with a mono-Z. preprint, 2024. Phys. Rev. D, 86:096011, 2012. [14] Humberto Reyes-González and Riccardo [5] Linda M. Carpenter, Andrew Nelson, Torre. Testing the boundaries: NormalizChase Shimmin, Tim M. P. Tait, and ing flows for higher dimensional data sets. Daniel Whiteson. Collider searches for SciPost Phys., 13:047, 2022. dark matter in events with a Z boson and missing energy. Phys. Rev. D, 87:074005, [15] Rob Verheyen. Event generation and den2013. sity estimation with surjective normalizing flows. SciPost Phys., 13:047, 2022. See [6] ATLAS and CMS Collaborations. Dark arXiv:2205.01697 for publication details. matter benchmark models for early LHC run-2 searches: Report of the ATLAS/CMS dark matter forum. arXiv [16] Anders Andreassen, Benjamin Nachman, and David Shih. Simulation assisted preprint, 2015. Editors: Antonio Boveia likelihood-free anomaly detection. Phys. and Caterina Doglioni. Rev. D, 101:055004, 2020. [7] Alexandre Alves and Kuver Sinha. Searches for dark matter at the LHC: [17] CMS Collaboration. DoubleMuon priA multivariate analysis in the mono-Z mary dataset in MINIAOD format from channel. Phys. Rev. D, 92:115005, 2015. RunD of 2015 (/DoubleMuon/Run2015D16Dec2015-v1/MINIAOD). CERN Open [8] CMS Collaboration. Search for new physics Data Portal, Record 24127, 2021. DOI: in events with a leptonically decaying Z 10.7483/OPENDATA.CMS.H3TX.ZJZX. boson and a large transverse momentum imbalance in proton–proton collisions at √ s = 13 TeV. Phys. Rev. D, 97:092005, [18] CMS Collaboration. DoubleEG primary dataset in MINIAOD format from RunD of 2018. 2015 (/DoubleEG/Run2015D-08Jun2016[9] CMS Collaboration. Search for dark matv1/MINIAOD). CERN Open Data Portal, ter produced in association with a leptoniRecord 24132, 2021. DOI: 10.7483/OPENcally decaying Z boson in proton–proton DATA.CMS.6ULE.YZJW. 28
√ [19] CMS Collaboration. Simulated dataset s = 13 TeV with the ATLAS detector. DarkMatter MonoZToLL V Mx-1 MvPhys. Lett. B, 761:372–392, 2016. 500 gDMgQ-1 TuneCUETP8M1 13TeVmadgraph in MINIAODSIM format [27] ATLAS Collaboration. Search for new high-mass phenomena in the dilepton final for 2015 collision data. CERN Open state using 36 fb−1 of proton–proton colliData Portal, Record 16630, 2021. DOI: √ sion data at s = 13 TeV with the ATLAS 10.7483/OPENDATA.CMS.YTLH.E0N7. detector. Phys. Lett. B, 776:318–338, 2018. [20] CMS Collaboration. Simulated dataset DarkMatter MonoZToLL A Mx-10 Mv- [28] Benjamin Nachman and David Shih. Anomaly detection with density estima20 gDMgQ-1 TuneCUETP8M1 13TeVtion. Phys. Rev. D, 101:075042, 2020. madgraph in MINIAODSIM format for 2015 collision data. CERN [29] Anna Hallin et al. Classifying anomalies Open Data Portal, Record 16575, through outer density estimation. Phys. 2021. DOI: 10.7483/OPENRev. D, 106:055006, 2022. DATA.CMS.YB3Q.XTXY. [30] CMS Collaboration. Precision luminosity [21] CMS Collaboration. Simulated dataset measurement in proton-proton collisions √ DarkMatter MonoZToLL A Mx-50 Mvat s = 13 TeV in 2015 and 2016 at cms. 200 gDMgQ-1 TuneCUETP8M1 13TeVEur. Phys. J. C, 81:800, 2021. madgraph in MINIAODSIM format for 2015 collision data. CERN Open Data Portal, Record 16597, 2021. DOI: 10.7483/OPENDATA.CMS.34IE.KN6I. [22] CMS Collaboration. Simulated dataset DarkMatter MonoZToLL EWK Scalar Mx100 Lambda3000 TuneCUETP8M1 13TeV-madgraph in MINIAODSIM format for 2015 collision data. CERN Open Data Portal, Record 16601, 2021. DOI: 10.7483/OPENDATA.CMS.NW7F.NFGG. [23] Glen Cowan, Kyle Cranmer, Eilam Gross, and Ofer Vitells. Asymptotic formulae for likelihood-based tests of new physics. Eur. Phys. J. C, 71:1554, 2011. [24] Asher Berlin, Stefania Gori, Tongyan Lin, and Lian-Tao Wang. Pseudoscalar portal dark matter. Phys. Rev. D, 92:015005, 2015. [25] Seng Pei Liew, Michele Papucci, Alessandro Vichi, and Kathryn M. Zurek. Mono-X versus direct searches: Simplified models for dark matter at the LHC. JHEP, 03:100, 2017. [26] ATLAS Collaboration. Search for highmass new phenomena in the dilepton final state using proton–proton collisions at 29