Good Papers

Showing Web & GUI agents Show all papers

80%Must read
?Must readVote to see the score

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

AutoGUIWorld uses image generators to synthesize GUI interaction trajectories without running software, improving OSWorld scores to 40.8% and ScienceBoard success to 32.2%.

Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang and 17 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face · Code ★ 9

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
88%Must read
?Must readVote to see the score

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

GUI-HARVEST optimizes executable harnesses for frozen GUI agents by aligning visual effects, comparing task runs, and consolidating failure patterns into reusable source edits, improving OSWorld-Verified by up to 12.33 points.

Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 5 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
91%Must read

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

DeskForge generates dense desktop supervision via controllable real-app environments, yielding 1.2M observations that improve GUI grounding and long-horizon computer-use task completion.

A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 10 on Hugging Face · Code ★ 4

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
67%Highly rated
?Highly ratedVote to see the score

A Subgoal-driven RL Framework for Improving Long-Horizon Web Agents

Taiyi Wang, Sian Gooding, Florian Hartmann, Oriana Riva and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

From Click Imitation to Transition Equivalence: Rethinking Supervision for GUI Agents

Zhiming Lin, Tianxiang Xu, zizhao zhang, Yixue Liu and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Yifan Zhang, Yutong Dai, Viraj Prabhu, Zhiyuan Hu and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

From Local Skills to Long-Horizon Tasks: Progressive Skill Exploration for LLM Web Agents

Xueyang Feng, Xiaohe Bo, Xu Chen, Quanyu Dai and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

Junhyuk Kwon, Seungjoon Lee, Hyejin Park, Kyle Min and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Beyond Task Success: Probing Cognitive Primitives in Web Agents

Xunjian Yin, Tianchen Guan, Jinao Wang, Weili Cao and 7 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
88%Must read
?Must readVote to see the score

PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control

PAGER closes the semantic-execution gap for point-precise geometric GUI control via dependency-structured planning and pixel-level execution, achieving 4.1x higher task success than general baselines.

Jingxuan Wei, Xi Bai, Shan Liu, caijun jia and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 15 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
86%Must read
?Must readVote to see the score

EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience

EvoCUA evolves computer-use agents via synthetic experience loops, achieving 56.7% success on OSWorld to set a new open-source state-of-the-art.

Taofeng Xue, Chong Peng, Mianqiu Huang, Linsen Guo and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 92 on Hugging Face · Code ★ 352

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
89%Must read
?Must readVote to see the score

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

OpenWebRL enables open online RL for visual web agents, with a 4B model reaching 67% Online-Mind2Web and 64% DeepShop success using minimal initialization data.

Rui Yang, Qianhui Wu, Yuxi Chen, Hao Bai and 6 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 19 on Hugging Face · Code ★ 52

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning

VSearcher uses reinforcement learning to turn static multimodal models into long-horizon web search agents that surpass proprietary models on multimodal search benchmarks.

RUIYANG ZHANG, Qianguo Sun, Chao Song, Yiyan Qi and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization

SMTL replaces sequential reasoning with parallel evidence acquisition for efficient long-horizon agentic search, achieving state-of-the-art results on multiple benchmarks with far fewer reasoning steps.

Chengjun Yu, Shu XU, Jiaqi Wu, Qianben Chen and 20 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 23 on Hugging Face · Code ★ 3

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
88%Must read
?Must readVote to see the score

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

OpenSearch-VL introduces an open-source recipe training multimodal deep search agents via curated data, diverse tools, and multi-turn fatal-aware GRPO, achieving over 10-point benchmark gains comparable to proprietary models.

Shuang Chen, Kaituo Feng, Hangting Chen, Wenxuan Huang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 57 on Hugging Face · Code ★ 289

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

ScreenSearch: Uncertainty-Aware OS Exploration

ScreenSearch combines structural screen retrieval with ambiguity-aware PUCT search to explore desktop OS states, collecting over 1M screenshots across 11 apps and showing ambiguity reduction alone is insufficient for exploration.

Michael Solodko, Justin Wagle

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

SkillMigrator learns reusable web skills via transferable interaction patterns matched by layout similarity to reduce LLM actions 8-10% across WebArena and Mind2Web.

Shiqi He, Yue Cui, Feijie Wu, Xinyu Ma and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents

ToolCUA learns optimal GUI-tool switching via scaled synthetic interleaved trajectories, tool-bootstrapped reinforcement learning, and online agentic rewards, achieving 46.85% accuracy on OSWorld-MCP.

XuHao Hu, Xi Zhang, Haiyang Xu, Kyle Qiao and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 26 on Hugging Face · Code ★ 62

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

Mecha-nudges for Machines

Mecha-nudging changes online choice environments to systematically influence AI agents without harming human usability, and Etsy listings show a 0.143-bit rise in machine-usable information after ChatGPT's release.

Giulio Frey, Kawin Ethayarajh

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
86%Must read
?Must readVote to see the score

Code2World: A GUI World Model via Renderable Code Generation

Code2World uses renderable code generation for GUI world modeling, achieving top next-UI prediction and boosting Android navigation success by up to 9.5%.

Yuhao Zheng, Li'an Zhong, Yi Wang, Rui Dai and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 186 on Hugging Face · Code ★ 312

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
Show 20 more papers