Imagine3D-LLM: a language model that first “imagines” a compact 3D scene from multiple photos before answering
This paper asks a simple question: can a multimodal language model do better on 3D and spatial questions if it first builds a small internal 3D sketch of the scene? The authors introduce Imagine3D-LLM, a model that learns to assemble a compact 3D representation from several images and then uses that imagined scene when it answers questions. The goal is to help the model combine evidence from different viewpoints into a coherent 3D understanding.
To do this, the researchers attach a few learnable “summary” tokens to the image inputs. Those tokens are decoded into a compact 3D representation called Gaussian Splatting. During training the model is given two signals: the usual language objective of predicting the next token (what comes next in text) and a photometric reconstruction loss that supervises the summary tokens. That reconstruction loss encourages the learned 3D summary to reproduce the appearance of the input views. Although only the summary tokens receive this direct supervision, the authors find the training also strengthens how the model’s internal image features match across frames.
The motivation comes from how people think about space. Instead of relying on exact pixel-by-pixel geometry, humans tend to match objects across views, judge how viewpoints relate, and build a coarse 3D layout. Prior methods either pushed very fine-grained pixel correspondence between views or fused features from specialized 3D geometry models. Imagine3D-LLM tries a middle path: a compact, learned 3D “imagination” that the language model conditions on when reasoning.
Why this matters: the paper reports that Imagine3D-LLM consistently outperforms earlier approaches on several spatial reasoning and 3D understanding benchmarks. The key claim is that letting the model imagine a compact 3D scene can be more effective than giving it raw pixel-level geometry. The method also appears to propagate 3D-aware signals throughout the model, improving cross-frame correspondence beyond the few supervised tokens.