MM–GSA: a way to measure input importance from fixed observational data using supervised learning
What the paper is about: The authors introduce MM–GSA, a method for Global Sensitivity Analysis (GSA) when you only have a fixed set of observations and cannot run the underlying model at chosen inputs. Instead of querying the model, MM–GSA uses supervised learning to build a metamodel that approximates the systematic relationship between inputs and the output. From that learned relationship the method reports two complementary measures of input relevance.
What the researchers did: MM–GSA combines a model-agnostic estimator of the first‑order Sobol' index and a new “trigger” or structural index. The first‑order Sobol' index measures how much of the systematic response variability can be attributed to a single input. The structural index measures how much predictive performance depends on having a given predictor available across alternative sets of predictors — in other words, whether including that variable helps prediction compared with leaving it out. Both indices are computed from the metamodel’s predictions and can be built with different supervised learning algorithms.
How it works, at a high level: The procedure first fits a metamodel to the observational dataset to learn the expected response as a function of predictors. The first‑order index is estimated by looking at how the metamodel’s expected output changes when each input varies. The structural index is estimated by comparing predictive performance across many fitted predictor subsets and tracking whether the target input is present. These two angles give different but complementary views: one asks about contribution to output variability, the other asks about usefulness for prediction.
Why it matters: Classical variance‑based GSA assumes you can evaluate the true model at chosen inputs, which is impossible for many real datasets collected observationally. MM–GSA translates GSA questions to that setting. The paper proves that both proposed estimators are statistically consistent under mild regularity conditions, and it shows a variable‑selection property for the structural index when inputs are independent. The authors also report Monte Carlo experiments and an application to NHANES (the U.S. National Health and Nutrition Examination Survey) to illustrate finite‑sample behavior.