Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron

Yiran Zhao, Wenxuan Zhang, Yuxi Xie, Anirudh Goyal, Kenji Kawaguchi +1 more
2/12/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 2 weaker matches are hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{zhao2026understanding,
  title  = {Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron},
  author = {Yiran Zhao and Wenxuan Zhang and Yuxi Xie and Anirudh Goyal and Kenji Kawaguchi and Michael Shieh},
  year   = {2026},
  url    = {https://api.semanticscholar.org/CorpusID:f20a924bb7e7a59fd11a36ff2ed9733bc2329d49},
  journal = {ICLR 2025 2025}
}

Discussion