New method helps estimate effects of several treatments when some instruments are flawed
This paper studies how to learn causal effects when researchers have more than one treatment and some instruments may be invalid. An instrumental variable (IV) is a piece of information used to tease apart cause and effect in observational data when there are unmeasured confounders. But if an IV directly affects the outcome or is correlated with hidden confounders, it is “invalid” and can spoil conclusions. The authors show that identification and inference are harder when there are multiple treatments, and they propose new rules and a robust confidence-interval method to handle that case.
The authors first analyze identification: when can the true vector of treatment effects be recovered from many instruments that may include invalid ones? In the single-treatment case each instrument points to a single candidate effect, so majority or plurality rules can pick out the true value. With multiple treatments a single instrument no longer gives a single candidate. Geometrically, each relevant instrument defines a (p_d − 1)-dimensional hyperplane in the p_d-dimensional space of treatment effects (p_d is the number of treatments). The paper introduces a generalized plurality rule that requires the true effect vector to lie on strictly more of these instrument-defined hyperplanes than any false value. They also give a generalized majority rule that provides a simpler sufficient condition when enough instruments are valid. When p_d = 1 these rules reduce to familiar single-treatment conditions.
The paper then turns to inference. Many existing methods first select a set of instruments deemed “valid” and then compute confidence intervals. But selection can fail in finite samples, especially when some invalid instruments are only weakly different from valid ones. The authors call these locally invalid IVs. To guard against selection mistakes they propose a sampling-based confidence interval. The idea is to introduce randomness when solving the linear systems that decide instrument validity, repeatedly generate candidate valid sets, and for each set compute a two-stage least squares (TSLS) confidence interval for each treatment effect. Aggregating these intervals yields a final confidence set. The authors prove that, under standard regularity conditions, this sampling confidence interval achieves the correct asymptotic coverage and has length shrinking at the usual parametric rate n^{-1/2}.