Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga +5 more
2/12/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 4 weaker matches are hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{zhou2026transfusion,
  title  = {Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model},
  author = {Chunting Zhou and Lili Yu and Arun Babu and Kushal Tirumala and Michihiro Yasunaga and Leonid Shamis and Jacob Kahn and Xuezhe Ma and Luke S. Zettlemoyer and Omer Levy},
  year   = {2026},
  url    = {https://api.semanticscholar.org/CorpusID:fe29071f8db5c333fcced241f70fb6f82e8452e6},
  journal = {ICLR 2025 2025}
}

Discussion