Multimodal Large Language Models Make Text-to-Image Generative Models Align Better

Xun Wu, Shaohan Huang, Guolong Wang, Jing Xiong, Furu Wei
2/3/2026

Abstract

Recent studies have demonstrated the exceptional potentials of leveraging human preference datasets to refine text-to-image generative models

DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 5 weaker matches are hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{wu2026multimodal,
  title  = {Multimodal Large Language Models Make Text-to-Image Generative Models Align Better},
  author = {Xun Wu and Shaohan Huang and Guolong Wang and Jing Xiong and Furu Wei},
  year   = {2026},
  doi    = {10.52202/079017-2584},
  url    = {https://doi.org/10.52202/079017-2584},
  journal = {NEURIPS 2024 2024}
}

Discussion