Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron
Yiran Zhao, Wenxuan Zhang, Yuxi Xie, Anirudh Goyal, Kenji Kawaguchi +1 more
2/12/2026
No confident code match yet
We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 2 weaker matches are hidden by default — verify before relying on them.
No code implementations found yet.
Know of an implementation? Let us know in the comments below!
@article{zhao2026understanding,
title = {Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron},
author = {Yiran Zhao and Wenxuan Zhang and Yuxi Xie and Anirudh Goyal and Kenji Kawaguchi and Michael Shieh},
year = {2026},
url = {https://api.semanticscholar.org/CorpusID:f20a924bb7e7a59fd11a36ff2ed9733bc2329d49},
journal = {ICLR 2025 2025}
}