57%Worth a look?Worth a lookVote to see the scoreNeurIPS 2026Esker / INSA Lyon / LIRISSeoul National University, SeoulMcGillOhio State University, ColumbusMcGill University, McGillAgent benchmarks & environmentsWebArena-Pro: A Heterogeneous, Multimodal, Reproducible Benchmark for Web AgentsImene Kerboua, Fatemeh Pesaran Zadeh, Xing Han Lu, Weijian Qi and 18 moreParis Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026– ReadersNo votes yet1/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 1 of 20 reviewers recommend itlenient 1/5medium 0/10strict 0/5
83%Must read?Must readVote to see the scoreNeurIPS 2026Seoul NationalSeoul National University, SeoulDatasets & benchmarksA Matched-Budget Audit Framework for Recaptioned Image-Text Supervision DistributionsA matched-budget audit framework profiles recaptioned image-text distributions via five axes and controllable basic units, showing released captions increase supported units by 3.39 to 6.36 and revealing a long-vs-dense frontier.Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae YuSydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026– ReadersNo votes yet13/20 AI panelreviewers recommend itReaders and the AI panel: vote on this paper to see what they said.Worth readingNot for meOnly vote on papers you've read. Sign in with GitHub to vote.AI panel: 13 of 20 reviewers recommend itlenient 3/5medium 9/10strict 1/5