MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
MT-Video-Bench introduces a multi-turn dialogue benchmark that evaluates multimodal LLMs on holistic video understanding.

65%
OverallWorth a look
?
OverallWorth a lookVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
?1 reader voted. Vote to see how they split.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel1/20reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
MT-Video-Bench is a timely multi-turn video benchmark that exposes multi-turn collapse, though its "holistic" scoring lacks rigorous agreement, error analysis, or validated temporal coherence checks.
Abstract
Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Moore Wang, Yongqian Wen, Yuanxing Zhang, Haoxuan Hu, Zhiyu Pan, Yibing Huang, Zhidong Gan, Yonghong Lin, An Ping, Shihao Li, Yanghai Wang, Tianhao Peng, Jiaheng Liu. Findings of the Association for Computational Linguistics: ACL 2026. 2026.