CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models

Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal +6 more
2/12/2026
DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. Any repos shown below are weak matches — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{jin2026cladder,
  title  = {CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models},
  author = {Zhijing Jin and Yuen Chen and Felix Leeb and Luigi Gresele and Ojasv Kamal and Zhiheng Lyu and Kevin Blin and Fernando Gonzalez Adauto and Max Kleiman-Weiner and Mrinmaya Sachan and Bernhard Schölkopf},
  year   = {2026},
  doi    = {10.48550/arXiv.2312.04350},
  url    = {https://doi.org/10.48550/arXiv.2312.04350},
  journal = {NEURIPS 2023 2023}
}

Discussion