Generative Cinematographer lets artists edit camera and object motion in 3D from a single image
This paper introduces Generative Cinematographer (GenCine), a system that turns one photo into a simple 3D scene that artists can edit. From that single image, an artist can draw a camera path and move parts of the foreground. The system then generates a video where camera and object motions are consistent with the 3D layout.
To make objects movable, the authors lift the image into an editable 3D scaffold. Artists mark foreground regions and attach local 3D motion handles. Several handles can move different parts of the same subject. That gives a piecewise-rigid approximation to non-rigid motion. In other words, complex object motion is built from several small rigid moves. The method does not use a physics simulator or a category-specific prior (a model tuned to a particular object type).
The paper explains how these human edits are turned into signals a video model can follow. The edits are projected into guidance maps. These maps record where each controlled region appears in every frame, give each handle a fixed color across frames, and encode the current 3D positions of controlled points in the same world coordinate system as the background. Using a shared coordinate system lets the method describe object motion relative to the scene even when the camera moves.
For training, the authors recover controls from motion observed in real videos and also use synthetic videos that include ground-truth geometry and trajectories. They train a lightweight guidance branch and small adapter modules (LoRA adapters) on top of a pretrained Wan video model so the generator learns to follow the guidance maps. This training strategy mixes real recovered signals and exact synthetic data to teach the system how to follow artist controls.
In experiments, the authors report that GenCine produces consistent camera-relative motion, better geometric consistency when the viewpoint changes, and strong controllability across a range of real-world scenes. Important caveats are that the method depends on a pretrained video model and on the training data it uses (recovered controls from real videos and synthetic ground-truth). The paper presents experimental evidence for the claims, but the system’s generality beyond the tested scenes and its behavior on very different inputs are not guaranteed by the abstract.