New machine‑learning pipeline aims to find unexpected particles produced with the Higgs boson
This paper develops and tests a machine‑learning strategy to look for unexpected new particles that appear together with a Higgs boson. The method, called HAXAD (Higgs And X Anomaly Detection), uses the Higgs decay to two photons — a clean signature with good mass resolution — as an anchor to search broadly for anomalies in collider data.
The analysis is built as a three‑stage pipeline. First, many measured quantities from each collision are compressed into a lower‑dimensional representation called an embedding. The authors introduce two new embedding approaches, one unsupervised (which trains only on data) and one semi‑supervised (which uses some labeled examples). Second, a generative background estimator named CATHODE (Classifying Anomalies Through Outer Density Estimation) learns the usual background distribution from the Higgs mass sidebands and interpolates it into the Higgs signal window. Third, a weakly supervised classifier based on CWoLa (classification without labels) is trained to pick out regions where data differ from the estimated background. The paper also adds a new statistical inference framework that produces both signal‑agnostic and signal‑specific upper limits on production cross sections.
The tests use Monte Carlo simulated proton–proton collisions at a center‑of‑mass energy of 13 TeV. Detector effects are modeled with Delphes, and the study assumes an integrated luminosity of 470 fb−1 (roughly the combined Run 2 and Run 3 dataset). The dominant background is non‑resonant diphoton plus jets production, and a variety of Standard Model Higgs production modes are included. To study sensitivity, the authors generate many beyond‑Standard‑Model signal scenarios: extended Higgs sectors and scalar resonances, heavy vector triplets, several supersymmetry setups (including R‑parity violating and colored SUSY), and top‑quark flavor‑changing processes. The paper reports that the updated HAXAD steps increase signal sensitivity compared with the earlier method and that, for many tested signals, it matches or exceeds the best individual cut‑based search limits on the same final state.