Telescopic Language Models let one trained model serve many compute budgets
This paper introduces Telescopic Language Models (TLMs), a way to train a single language model that works well at many different sizes or computation budgets. In plain terms, the same trained model can be run with fewer layers when you need to save compute, or with all layers when you want the best quality. The authors show a training recipe that makes the model a valid next-token predictor at every intermediate depth, rather than only at the full size or at a few hand-chosen sizes.
The core idea is to train a standard Transformer model so that partial versions of it are also correct language models. At every training step the method picks one random truncated prefix of the model’s layers and trains that short version to predict the next token. It also trains the full model in the same step as an “anchor.” That means two forward-and-backward passes per step, but no change to the model architecture and nothing extra needed at inference time. The authors call this supervision strategy stochastic prefix supervision with a full anchor.
The team tested the idea on a proxy suite with 200 million parameters and the same data stream for all methods (20 billion tokens from FineWeb-Edu). They compared TLM to “fixed-exit” suites such as Matryoshka Language Model Suites (MLMS), where only a few layer depths are trained for deployment. In the baselines that only supervise a few fixed exits, the model performs very poorly at other depths — with perplexities in the range 10^2 to 10^5, which is near chance. A single TLM run produced a valid language model at each of twenty layer prefixes. Quantitatively, TLM reduced the area under the quality-versus-compute curve by about 43–44% compared with fixed-exit suites, while matching full-capacity performance and costing about 12% less GPU time per run.
Why this matters: in practice one deployed model often needs to serve many compute budgets (for example, mobile vs. server settings). The usual approach is to train or compress a separate model for each budget point, which is expensive. TLMs offer a single artifact that can be trimmed at inference to meet different budgets, saving training and maintenance cost. The paper also shows that how often you sample each prefix during training (the “prefix sampling density”) is a tuning knob. If you concentrate training on a few depths, you can recover the fixed-exit quality at those depths, but you lose the smooth continuum of good performance across all depths.