Multimodal Flow: a fully continuous model that generates language and images together
This paper introduces Multimodal Flow, a new way to generate both text and images using one continuous model. Unlike many recent multimodal systems, it avoids turning images into discrete image tokens (a process called quantization) or mixing separate generation steps for text and pixels. Instead, the model works in a single continuous embedding space and treats text and images as the same kind of data, so they can be generated by one shared process.
To make this work the authors arrange text tokens and image regions into ordered "hyperchunks" — blocks of continuous vectors that keep text order and image spatial layout. A single backbone network, called a chunk-causal flow, learns a vector field over these hyperchunks by a technique called Flow Matching. In plain terms, Flow Matching teaches the model how to move points in embedding space so they become realistic text-and-image chunks. The model also uses joint attention to let text and image parts interact, and it keeps small, modality-specific feed-forward modules to handle the parts that need different processing.
During training the model predicts several target hyperchunks in parallel. At inference time it generates hyperchunks one after another. The authors built a version called MF-1 and pretrained it on multimodal data at three sizes: about 0.6 billion, 1.2 billion, and 1.6 billion parameters. With 150 billion pretraining tokens, MF-1 scored an average of 82.8 on the combined GenEval and DPG-Bench suites, and 75.3 across VQAv2, MMBench, and POPE — benchmarks used to test multimodal understanding and reasoning. Under matched data, optimization, and parameter budgets, the paper reports that Multimodal Flow outperforms representative hybrid and discrete models.
This approach matters because it removes a common visual bottleneck: converting images into discrete tokens can lose information and force different generation rules for each modality. A single continuous generative process makes it easier for text and images to share representations and interact. The model design also aims to be simpler to train and to sample from, since it uses one unified flow rather than multiple modality-specific procedures. The authors have released their code and models publicly at https://github.com/hustvl/Multimodal-Flow.