Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation Editing

Peiran Dong, Bingjie Wang, Song Guo, Junxiao Wang, Jie Zhang +1 more
2/3/2026

Abstract

Recent advancements in vision-language-to-image (VL2I) diffusion generation have made significant progress. While generating images from broad vision-language inputs holds promise, it also raises concerns about potential misuse, such as copying artistic styles without permission, which could have legal and social consequences. Therefore, it’s crucial to establish governance frameworks to ensure ethical and copyright integrity, especially with widely used diffusion models. To address these issues, researchers have explored various approaches, such as dataset filtering, adversarial perturbations, machine unlearning, and inference-time refusals. However, these methods often lack either scalability or effectiveness. In response, we propose a new framework called causal representation editing (CRE), which extends representation editing from large language models (LLMs) to diffusion-based models. CRE enhances the efficiency and flexibility of safe content generation by intervening at diffusion timesteps causally linked to unsafe concepts. This allows for precise removal of harmful content while preserving acceptable content quality, demonstrating superior effectiveness, precision and scalability compared to existing methods. CRE can handle complex scenarios, including incomplete or blurred representations of unsafe concepts, offering a promising solution to challenges in managing harmful content generation in diffusion-based models.

DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 1 weaker match is hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{dong2026towards,
  title  = {Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation Editing},
  author = {Peiran Dong and Bingjie Wang and Song Guo and Junxiao Wang and Jie Zhang and Zicong Hong},
  year   = {2026},
  doi    = {10.52202/079017-0404},
  url    = {https://doi.org/10.52202/079017-0404},
  journal = {NEURIPS 2024 2024}
}

Discussion