WhodunitBench: Evaluating Large Multimodal Agents via Murder Mystery Games

Junlin Xie, Ruifei Zhang, Zhihong Chen, Xiang Wan, Guanbin Li
2/3/2026

Abstract

Recently, large language models (LLMs) have achieved superior performance, empowering the development of large multimodal agents (LMAs). An LMA expected to perform practical tasks must possess a range of capabilities, including multimodal perception, interaction, reasoning, and decision-making skills. However, existing benchmarks are limited in assessing compositional skills and actions ♡ Equal contribution ♣ Corresponding authors.

DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 1 weaker match is hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{xie2026whodunitbench,
  title  = {WhodunitBench: Evaluating Large Multimodal Agents via Murder Mystery Games},
  author = {Junlin Xie and Ruifei Zhang and Zhihong Chen and Xiang Wan and Guanbin Li},
  year   = {2026},
  doi    = {10.52202/079017-2751},
  url    = {https://doi.org/10.52202/079017-2751},
  journal = {NEURIPS 2024 2024}
}

Discussion