XL-DocBench introduces a human-verified benchmark for extra-long professional document understanding spanning thousands of pages with multi-page evidence, showing current systems still struggle with long-context structured reasoning.
DocAtlas treats long-document understanding as a mutable-state interaction process via a document harness with search, memory, and review tools, reaching 71.4% on MMLongBench-Doc and boosting a 4B VLM to 63.7% via reinforcement learning.