Failures to Find Transferable Image Jailbreaks Between Vision-Language Models

Rylan Schaeffer, Dan Valentine, Luke Bailey, James Chua, Cristobal Eyzaguirre +11 more
2/12/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 1 weaker match is hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{schaeffer2026failures,
  title  = {Failures to Find Transferable Image Jailbreaks Between Vision-Language Models},
  author = {Rylan Schaeffer and Dan Valentine and Luke Bailey and James Chua and Cristobal Eyzaguirre and Zane Durante and Joe Benton and Brando Miranda and Henry Sleight and Tony Tong Wang and John Hughes and Rajashree Agrawal and Mrinank Sharma and Scott Emmons and Oluwasanmi Koyejo and Ethan Perez},
  year   = {2026},
  url    = {https://api.semanticscholar.org/CorpusID:e05e9839b6544c78ccb10e86c2719328fc0d1562},
  journal = {ICLR 2025 2025}
}

Discussion