Understanding and Minimising Outlier Features in Transformer Training

Bobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag, Thomas Hofmann
2/3/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 1 weaker match is hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{he2026understanding,
  title  = {Understanding and Minimising Outlier Features in Transformer Training},
  author = {Bobby He and Lorenzo Noci and Daniele Paliotta and Imanol Schlag and Thomas Hofmann},
  year   = {2026},
  url    = {https://www.semanticscholar.org/paper/e32803eb36116f3e218563d69b60171eddaf7024},
  journal = {NEURIPS 2024 2024}
}

Discussion