
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
MiroEval benchmarks multimodal deep research agents via process and outcome evaluation across 100 real-world tasks, finding process quality predicts outcomes and multimodal tasks reduce scores by 3, 10 points.
Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 69 on Hugging Face · Code ★ 52
Only vote on papers you've read. Sign in with GitHub to vote.