Finite-sample bounds clarify when localized conformal prediction gives reliable, local intervals
Conformal prediction is a popular way to wrap any black‑box model with a guarantee that prediction sets cover the true answer at a stated rate on average. But “on average” can hide big errors at particular covariate values. Randomly localized conformal prediction (RLCP) tries to fix that by giving more weight to calibration data near the test point. This paper gives the first finite‑sample, high‑probability guarantees for the actual localized prediction set that RLCP produces. The results control both the local coverage error (how often the set misses a true label near the point) and the set’s length compared with an ideal oracle interval.
To get these guarantees the authors study two cases. First they fix a scoring rule (a way to rank how surprising a candidate label is) and assume a mild smoothness condition on how the score distribution changes with the covariates (a Hölder condition), together with standard assumptions on the covariate density and the kernel used for localization. Under those assumptions they prove that, with high probability over the calibration sample and conditional on the realized auxiliary localization centre, the coverage error and the excess length are uniformly small for every test point inside the realized neighbourhood. The error splits into two parts: a localization bias that shrinks like a power of the bandwidth h (written O(h^β)), and a calibration term that shrinks as the effective local calibration sample grows. Balancing these two terms recovers the usual nonparametric rate (roughly n^{-β/(2β+d)} up to log factors), which makes explicit the familiar bandwidth bias–variance tradeoff in this setting.
Second, the paper treats learned scores that are fitted on an independent training sample. When the learned score targets a pivotal quantity (a score whose relevant quantile is the same everywhere, as in conformalized quantile regression), the localization bias vanishes. In that case the local guarantees decompose cleanly into a calibration term and a score‑estimation term. The analysis shows that better score learning directly sharpens the localized RLCP guarantees, and that estimation error affects the threshold comparison in a controlled, linear way.