Attention with Markov: A Curious Case of Single-layer Transformers

A. Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi +2 more
2/12/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. Any repos shown below are weak matches — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{makkuva2026attention,
  title  = {Attention with Markov: A Curious Case of Single-layer Transformers},
  author = {A. Makkuva and Marco Bondaschi and Adway Girish and Alliot Nagle and Martin Jaggi and Hyeji Kim and Michael Gastpar},
  year   = {2026},
  url    = {https://api.semanticscholar.org/CorpusID:d7c0143730c3b59766697182433188570c547b60},
  journal = {ICLR 2025 2025}
}

Discussion