New math shows how to learn discrete diffusion samplers for locally dependent categorical data
Many problems require drawing samples from high-dimensional categorical distributions where each variable only depends on a few neighbors. Examples include short-memory language models, Potts and Ising systems in physics, and some protein models. This paper studies discrete diffusion methods for that task and gives the first end-to-end statistical guarantees that tie learning error, sampling error, and data size together when the target distribution has local dependence described by a low-order Markov random field (MRF).
The authors focus on a commonly used noising scheme called uniform noising. In a discrete diffusion model one runs a forward Markov chain to corrupt data toward a simple reference distribution, then learns a reverse-time “score” that tells how to reconstruct clean samples. The main technical insight of the paper is a new “pinning decomposition” of that discrete score. It shows that, under uniform noising, the score factors into known time-dependent weights (simple binomial-type factors) multiplied by target-dependent partial marginals. In plain terms, the part that depends on how much noise was added separates cleanly from the part that depends on the unknown data law.
That separation lets the same target-specific quantities be reused across all noise levels. The authors exploit this by proposing a weight-sharing neural score learner. Instead of training separate networks for each noise level, one network stores the small number of target-dependent parameters implied by the MRF and shares them across times. They pair this learner with τ-leaping, a Poisson jump scheme used to simulate the reverse chain, and they analyze how finite data affects the learned score and the final samples. Their analysis gives explicit sample-complexity bounds that depend on the vocabulary size S, the interaction order d of the MRF (how many coordinates interact at once), and the number of observed samples n.