metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Alex Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric Schulz
2/12/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 5 weaker matches are hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{kipnis2026metabench,
  title  = {metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models},
  author = {Alex Kipnis and Konstantinos Voudouris and Luca M. Schulze Buschoff and Eric Schulz},
  year   = {2026},
  url    = {https://api.semanticscholar.org/CorpusID:d4ed201bfc47a02e001766c294ae6afd38b87079},
  journal = {ICLR 2025 2025}
}

Discussion