Quantum field sampling as a steering problem: learning a finite‑time control to reach the Gibbs measure
This paper shows a new way to do stochastic quantization. Instead of waiting for a fictitious Langevin dynamics to drift to equilibrium, the authors pose quantization as a finite‑time optimal control problem. A simple, solvable “free” theory supplies a reference dynamics (an Ornstein–Uhlenbeck process). A learned, residual control force then steers trajectories so that, at a chosen finite fictitious time T, the ensemble matches the desired Gibbs (Boltzmann) measure for the interacting theory.
Mathematically the authors combine a penalty for changing the reference dynamics with a terminal cost that contains the full action. The optimal correction to the reference drift is a Doob transform, which can be written as the gradient of a log expectation of future Boltzmann weights. In practice they parametrize the residual force with a neural network and train it to minimize the control objective. Because the exact path weights are known from the theory, each trajectory can be reweighted exactly; imperfect training therefore raises the variance of observable estimates but does not produce a biased model.