Good Papers

Showing Agent benchmarks & environments Show all papers

84%Must read
?Must readVote to see the score

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

CheckerBench evaluates long-horizon agents on synthesizing static-analysis checkers across 300 CVE-derived tasks, finding best Pass@1 reaches 45.33%.

Hang He, Li Wang, Hao Chen, Yuchen Shao and 8 more

Published Oct 6, 2026 · ▲ 52 on Hugging Face · Code

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

MiniCorp: The Last Mile of the AI Agent Firm

MiniCorp is a simulated office environment that generates longitudinal, counterfactual enterprise data to study autonomous AI-run companies and train adaptive agents.

Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin and 5 more

Published Oct 5, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
86%Must read
?Must readVote to see the score

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

AutoSciBench autonomously generates and iteratively adapts scientific agent benchmarks via recipes and concepts, reducing solver accuracy by over 22 points versus human benchmarks while improving quality ratings.

Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards and 6 more

Published Oct 4, 2026 · ▲ 17 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
91%Must read

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

UndoBench separates tool-using agent competence from fault recovery via paired enterprise workflow trials, finding 83.54% nominal success but only 46.72% recovery success with phase-dependent vulnerabilities.

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
89%Must read
?Must readVote to see the score

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

LMBuild evaluates LLM agents on generating buildable, functional 3D structures and finds physical operability and functional affordance remain challenging despite improved soundness.

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang and 8 more

Published Oct 3, 2026 · ▲ 24 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

HyperBrowseComp introduces a multilingual, multimodal web-browsing benchmark of 423 hard questions requiring obscure evidence discovery, and current agents perform poorly against human baselines.

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han and 13 more

Published Oct 2, 2026 · ▲ 53 on Hugging Face

100% Readers1 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 1/5
88%Must read
?Must readVote to see the score

Self-Supervised Scaling of Terminal Environments for Scientific Domains

Software-in-the-loop reconstruction scales terminal-agent training by deriving verified tasks from existing scientific workflows without manual references, improving Terminal-Bench 2 performance to 53.56%.

Zhongzhi Li, Yucheng Shi, Zongxia Li, Junyao Yang and 7 more

Published Oct 2, 2026 · 0 citations · ▲ 10 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

World Editing: Intervening on Executable Worlds at Increasing Depth

World editing intervenes on executable environments at increasing depth via IGMWorld and IGMBench, where top agents achieve 78.2% task success with reliability declining by depth.

Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh and 14 more

Published Oct 1, 2026 · ▲ 14 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 3/5
86%Must read
?Must readVote to see the score

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

SimuVerity benchmarks text-to-executable Simulink generation across engineering domains, finding best agents score only 42.86 and structural similarity poorly predicts engineering performance.

Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo and 8 more

Published Oct 1, 2026 · ▲ 47 on Hugging Face · Code ★ 19

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
88%Must read
?Must readVote to see the score

Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

Human audit of 165 WebArena-Lite tasks recovers 5.45, 8.49% evaluator-missed successes, reveals trajectory errors like looping, and shows guide text and MASM improve results.

Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

OSWorld-Science benchmarks VLM agents on 146 expert scientific software tasks, showing state-of-the-art models still struggle with scientific workflows and harness design.

Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li and 27 more

Published Sep 30, 2026 · 0 citations · ▲ 63 on Hugging Face · Code ★ 5

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

WorldAuditBench benchmarks interactive 3D world auditing with multimodal agents, finding success rates of 6.6% to 42.3% versus 83.4% human performance.

Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 102 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

OSWorld-Pro introduces process-based evaluation with 2,800 subgoals across 300 tasks, revealing top models achieve only 75.7% subgoal success versus 83.4% end-state performance and identifying distinct failure modes like irrelevant actions and click errors.

Zhilin Wang, Shaokun Zhang, Yifan Zhang, Hao E. Zhang and 8 more

Published Sep 21, 2026 · 0 citations · ▲ 15 on Hugging Face

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

MTAC-IFBench benchmarks multi-turn instruction-following in agentic coding via progressive constraints, revealing rapid performance degradation in current code agents as sessions lengthen.

Bosi Wen, Cunxiang Wang, Jiayi Gui, Haoke Zhang and 5 more

Published Sep 14, 2026 · 0 citations

– ReadersNo votes yet. 1 from authors or colleagues not counted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

HarnessDev evaluates LLMs creating and evolving agent harnesses, finding generated harnesses lag human references on coding and search but match them on writing and ML tasks, with unstable, model-dependent evolution gains.

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei and 15 more

Published Sep 1, 2026 · 0 citations · ▲ 566 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

EnvHarness: Awakening Static Worlds for Agent Learning

EnvHarness wraps static environments with programmable components to reshape agent behavior without altering underlying logic, improving benchmarks by up to 9.0 points while enabling continuous policy-environment co-evolution.

Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan and 13 more

Published Aug 20, 2026 · 0 citations · ▲ 175 on Hugging Face · Code ★ 619

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Workflow-GYM benchmarks long-horizon professional GUI workflows, showing top agents achieve only ~30% success due to stage omission, error propagation, and objective drift.

Liya Zhu, Jingzhe Ding, Jian Zhang, Jianbo Xue and 36 more

Published Jun 9, 2026 · 0 citations · ▲ 221 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Agents' Last Exam

ALE introduces a benchmark evaluating AI agents on long-horizon, economically valuable real-world tasks across 13 industry clusters, finding current full pass rates below 1%.

Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang and 36 more

Published Jun 3, 2026 · 0 citations · ▲ 392 on Hugging Face · Code ★ 1,083

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench introduces 153 real-world online tasks across 144 platforms to evaluate AI agents, finding frontier models complete only about a third of them.

Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du and 24 more

Published Apr 9, 2026 · 0 citations · ▲ 377 on Hugging Face · Code ★ 958

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
90%Must read

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

TASTE reverses benchmark construction by evolving tool sequences to automatically generate harder, broader-coverage agent tasks that expose severe performance drops and saturation in existing benchmarks.

Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026 · ▲ 74 on Hugging Face · Code ★ 4

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

AudioAgentBench: Evaluating Multi-Turn Voice Agents on Real-World Tasks

Gardenia Liu, Yi-Hao Peng, Kamryn Ohly, Oliver Johansson and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AutoDataBench: How Far Are LLM Agents from Autonomously Engineering Post-Training Data Pipelines?

Qiaoyu Tang, Hao Xiang, Le Yu, Yaojie Lu and 9 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RoutingBench: Can Agentic Routing Analysis Scale to Production Datacenter Networks?

Wenlong Ding, Zhixiong Niu, Jianan Yang, Fajun Zhang and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

Junjue Wang, Weihao Xuan, Heli Qi, Pengyu Dai and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Z-AXIS: From Deterministic Ground to Agentic Depth for Enterprise Evaluation

Zifan Song, Mianzhi Chang, Ziyang Liao, haiyan xu and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

M4Bench: Evaluating Procedural Specification for Clinical EHR Derivation Agents

Hannes Ill, Rafi Al Attrach, Rajna Fani, Ahram Han and 4 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CEO-Bench: Can Agents Play the Long Game?

Haozhe Chen, Karthik Narasimhan, Zhuang Liu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · Code ★ 82

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

Kean Shi, Zihang Li, Tianyi Ma, Zengji Tu and 11 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SudoBench: A Contextual Authorization Benchmark for LLM Agents

Vincent Siu, Tianneng Shi, Shangding Gu, Zhun Wang and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

Liu Dai, Haina Wang, Weikang Wan, Hao Su

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

Zixin CHEN, Peng Liu, Rui SHENG, Haobo Li and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

OS-Omni: A Cross-Platform Benchmark for Generalist Computer-Using Agents

Hui Shen, Yunta Hsieh, Jianing Ma, Ziyuan Liu and 36 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

PCBInnoBench: Benchmarking LLM Agents on Real-World PCB Design

Fuyuan Xia, Jiaxin Hu, Kewei Yu, Zihan Zhang and 12 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Synthetic Tasks for Training AutoResearch

Ziyang Cai, Seyyedamirhossein Saeidi, Harkirat Singh Behl

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Synthetic Web: Benchmarking Language Agents under Adversarial Search Ranking

Shrey Shah, Levent Ozgur

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

SciResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Lei Xiong, Kun Luo, Ziyi Xia, Wenbo Zhang and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ScrapeBench: Evaluating Legal Compliance of AI Agents in Website Scraping

Joseph Marvin Imperial, Daniel Slate, Noam Kolt

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

MedFlowBench: Auditing Medical Agents in Full-Study Workflows

Weixiang Shen, Chengzhi Shen, Che Liu, Junde Wu and 10 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Evaluating Physical Reasoning in LLM Agents Requires Construction Benchmarks

Wenhao Deng, Tian Xia, Yuru Jiang, Joemon M Jose and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

Yun-Shiuan Chuang, Ruixuan Tu, Chengtao Dai, You Li and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CREF: Forecasting Benchmarks for the Age of Agents

Andreas Auer, Abdul Fatir Ansari, Oleksandr Shchur, Xiyuan Zhang and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

ABHBench: Evaluating Moral Decision-Making of Foundation Models from an Agentic Perspective

Chaoran Li, Yukun Li, Zeyuan Zhao, Shao Zhang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Task Success Is Not Enough: Side-effect-Aware Evaluation of Tool-Using Language Model Agents

Jiaju Huang, Shaobin Chen, Xinglong Liang, Xinyu Ma and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

See, Read, Compare: Candidate-Aware Verification for Agent Test-Time Scaling

Xinyu Ye, Yongliang Wu, Xingyu Zhu, Peng Xia and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PhysAgentGym: Free Physical-Law Verifiers for Training Small Code-Reasoning Agents

Lama Moukheiber

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ClawBenchPro: Benchmarking How Well Agent Harnesses Work

Yufei Liu, Yirong Zeng, Yuxiang He, Yutai Hou and 9 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

From Benchmark to Adoption: Coding Agent Evaluation Needs User-Level Harness

Dasol Hong, Ji Yong Cho, Bumsoo Kang, Moontae Lee

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

TrialAgentBench: Evaluating AI Agents for Clinical-Trial Analysis and Long-Horizon Drug-Development Decisions

Bradley M Segal, William J Bolton, Philip Torr, David Clifton and 1 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ContractBench: Can LLM Agents Preserve Observation Contracts?

Jicheng Wang, Yifeng He, Zili Wang, Hanwen Xing and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DDBench: A Benchmark for Agentic Debugging on Distributed Systems

Yibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Argus: A Cross-Regime Benchmark for the Transferability of Uncertainty Quantification in Computer-Use Agents

DIVAKE KUMAR, Sina Tayebati, Devashri Naik, Amanda Rios and 4 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
69%Highly rated
?Highly ratedVote to see the score

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench

Qingyun Zou, Feng Yu, Hongshi Tan, Jiahao Cui and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
3/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 3 of 20 reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

WebArena-Pro: A Heterogeneous, Multimodal, Reproducible Benchmark for Web Agents

Imene Kerboua, Fatemeh Pesaran Zadeh, Xing Han Lu, Weijian Qi and 18 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

JARVIS-Bench: Benchmarking Personal Intelligence Agents on Long-Horizon Real-User Daily Traces

Weizhi Zhang, Wei-Chieh Huang, Yueqing Liang, Liwei Jiang and 36 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

BEAKER: An Expert-Curated Benchmark for Embodied Brains in Self-Driving Chemical Laboratories

Fei Lin, Tengchao Zhang, Ziyang Gong, Xiaotong Yu and 14 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AM-Bench: A Unified Taxonomy and Evaluation Suite for Agentic Misalignment

Eric Zhang, Terry J Zhang, Chijioke Ugwuanyi, Jerick Shi and 2 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

ShopGym converts live e-commerce sites into reproducible simulated shops and synthesizes diverse benchmark tasks, showing synthetic shops preserve live structural properties and agent performance correlations.

Yuanzheng Zhu, Mingyu Zhao, Chinmay Savadikar, Han Li and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
91%Must read
?Must readVote to see the score

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents

DiagEval uses trajectory-conditioned diagnostic probes to disambiguate evaluator errors from software defects in GUI-agent evaluations, recovering over 45% of misattributed failures and improving accuracy substantially.

Sirui Hong, Liuzhijie, Tengfei Li, Wei Tao and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 2/5
80%Must read
?Must readVote to see the score

AI GAMESTORE: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games

AI GameStore proposes evaluating general intelligence via scalable synthesis of human games, finding frontier vision-language models score under 10% of human averages on most generated games.

Lance Ying, Ryan Truong, Prafull Sharma, Kaiya Zhao and 8 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

PokeAgent is a large-scale Pokémon benchmark with battling and speedrunning tracks that expose major gaps between LLMs, RL agents, and human experts.

Seth Karten, Jake Grigsby, Tersoo Upaa, Junik Bae and 27 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
Show 20 more papers