Good Papers

Trending

What readers here and on Hugging Face are upvoting

91%Must read

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Self-evolving search agents suffer co-cheating where proposers and solvers mutually reinforce errors; CrossFit partitions sources to cross-fit agreement and cuts false agreement by over half, boosting downstream search by 8+ points.

Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin and 11 more

Published Sep 30, 2026 · 0 citations · ▲ 670 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

RealCompanion benchmarks AI companions on longitudinal real-world chats, finding needed past messages are usually recent, memory detectors fail on real messages, and persona reconstruction costs vary 31-fold at equal F1.

Arman Behnam, Sunglyoung Kim, Liangwei Yang

Published Oct 1, 2026 · 0 citations · ▲ 268 on Hugging Face

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Kandinsky 6.0 Video introduces diffusion models that generate synchronized 5-second audio-video clips with lip-sync via a dual-stream CrossDiT architecture, with the 29B-parameter Pro version outperforming its predecessor and matching top competitors in speech quality.

Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin and 36 more

Published Oct 4, 2026 · ▲ 113 on Hugging Face · Code ★ 119

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
91%Must read

Does Learning Protein Folding Generalize to Broader Reasoning?

Post-training on protein-folding data via discrete answers and continuous geometry improves structure prediction and broad reasoning across ten benchmarks.

Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 120 on Hugging Face · Code ★ 28

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

FrameMorrow guides historical frame selection via prospective tokens representing future needs, improving consistency and quality across diverse long-horizon video generators.

Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan

Published Sep 30, 2026 · 0 citations · ▲ 92 on Hugging Face · Code ★ 26

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Recursive Self-Rewrite uses diverse harnesses and recursive revision to rewrite successful terminal trajectories for supervised fine-tuning, boosting pass@3 by up to 7.6x on hard benchmarks.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 91 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Cross-tokenizer on-policy distillation achieves comparable accuracy with strict top-16 shared-vocabulary supervision versus full coverage, while expanded span supervision reduces accuracy due to conflicting gradients, motivating prioritization of supervision reliability over alignment coverage.

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li and 3 more

Published Oct 6, 2026 · ▲ 40 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
89%Must read

World Action Modeling with Progressive Visual Planning

ProWAM predicts sparse visual sub-goals and actions via progressive planning, achieving state-of-the-art long-horizon robotic control and strong zero-shot real-world generalization.

Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 83 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read

Native Action-Prior Learning from Videos for World Action Models

NAVA-WAM pretrains robot action policies directly from observation-only videos via flow-matching and joint attention, improving control accuracy and label efficiency.

Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost and 9 more

Published Oct 2, 2026 · 0 citations · ▲ 81 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

UniEvo-VL improves multimodal image generation via self-distillation that minimizes divergence between student and critique-conditioned teacher diffusion distributions during self-correction. Experiments on Qwen2.5-Image improve GenEval scores from 0.747 to 0.808 without external teachers.

Fang Wu, Da Xing, Yanjie Huang, Junxi Wang and 15 more

Published Sep 30, 2026 · 0 citations · ▲ 292 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

MotorMind connects general vision-language models to deterministic robot control via mid-level actions and feedback loops, achieving 66.7% zero-shot success on LIBERO-PRO and 95% on real robots without task-specific training or external tools.

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 90 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
78%Highly rated

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD distills multi-teacher LLM specialists via hidden-state and prediction-level on-policy supervision, outperforming token-only and representation-only baselines across math, code, and logic benchmarks.

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 63 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

ALoDLM: Adaptively Looped Diffusion Language Models

ALoDLM applies token-adaptive latent recurrence to diffusion language models, allocating computation by difficulty to close the quality gap with autoregressive models at 1.7B and 8B scales.

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao and 9 more

Published Oct 3, 2026 · ▲ 55 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer unifies streaming video perception, memory, and proactive response via shared generation, achieving top results on eight benchmarks with a 4B model.

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu and 20 more

Published Oct 1, 2026 · 0 citations · ▲ 229 on Hugging Face · Code ★ 157

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

LexReward introduces taxonomy-driven rubric-based rewards for legal language models across style, element, and reasoning dimensions, improving DPO and reinforcement learning performance.

Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 51 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

DuoMatching improves few-step video generation by jointly matching frame distributions and adding frame-level supervision via an image teacher, boosting visual quality and semantic alignment over 80%.

Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 4

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot-SD self-distills masked diffusion language models by supervising high-impact commitment tokens via information-gain selection, improving reasoning with minimal data.

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 54 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

On-policy methods continuously adjust parameter update directions, unlike consistent SFT updates; constraining SFT to these directions via OPSFT transfers on-policy generalization advantages to supervised fine-tuning.

Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 81 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

HyperBrowseComp introduces a multilingual, multimodal web-browsing benchmark of 423 hard questions requiring obscure evidence discovery, and current agents perform poorly against human baselines.

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han and 13 more

Published Oct 2, 2026 · ▲ 53 on Hugging Face

100% Readers1 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 1/5
92%Must read

Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers

A benchmark of 390 AI papers measures scientific slop across structure, argument, and artifacts; a harness reduces the AI-human gap by 63% via evidence-grounded revision.

Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 55 on Hugging Face · Code ★ 13

100% Readers1 of 1 upvoted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

PointWAM forecasts 3D point trajectories of scenes and hands in a shared space-time frame to guide dexterous robot manipulation, improving DexJoCo success by 56.9 points with video pre-training.

Chunghyun Park, Beomjun Kim, Seungcheol Park, 권희승 and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 42 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero benchmarks long-horizon part-level 3D editing via sequential instructions and target images, finding agentic code-based methods preserve unedited regions better but are slower than non-agentic regeneration.

Ruihan Yu, Yu-Ju Tsai, Muyao Niu, Runyi Li and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 50 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

GraphForge synthesizes workspace tasks and verifiers over real file evidence graphs to train working agents, and fine-tuning Qwen3.6-27B improves GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II results.

Qisheng Su, Hanchen Wang, 朱冠儒, Huicheng Jiang and 8 more

Published Sep 30, 2026 · 0 citations · ▲ 146 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
86%Must read
?Must readVote to see the score

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

SimuVerity benchmarks text-to-executable Simulink generation across engineering domains, finding best agents score only 42.86 and structural similarity poorly predicts engineering performance.

Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo and 8 more

Published Oct 1, 2026 · ▲ 47 on Hugging Face · Code ★ 19

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
89%Must read
?Must readVote to see the score

PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs

PDE-JEPA introduces predictive masked-latent pretraining with geometry projection and structured latent predictors for parametric PDE dynamics, reducing errors by 33.4% in-distribution and 51.4% on unseen parameters.

Zhentao Tan, Jianrong Zhang, Ruijie Quan, Yi Yang

Published Sep 28, 2026 · 0 citations · ▲ 53 on Hugging Face · Code ★ 9

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Spatial Memory Intelligence introduces understanding-driven atomic operations for spatial-memory management in long-video world models, improving sparsity, stability, and spatial consistency.

Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 49 on Hugging Face · Code ★ 19

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness turns fixed base LLMs into agentic verifiers with workspaces and evidence tools, achieving top selection scores and 6.2, 6.4 point gains over single rollouts on long-horizon tasks.

Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 49 on Hugging Face · Code ★ 42

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid-Harness verifies candidate terminal actions at the model-harness boundary, raising TerminalBench-Lite Pass@1 from 50.00% to 68.03% and improving success at lower token cost than trajectory scaling alone.

Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan and 7 more

Published Sep 30, 2026 · 0 citations · ▲ 116 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
84%Must read
?Must readVote to see the score

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

TRACE aligns FP4 quantization between RL training and rollout paths via rollout-guided quantization-aware training for MoE language models, achieving BF16-comparable RL performance with up to 5.4x rollout speedup.

Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao and 8 more

Published Oct 6, 2026 · ▲ 15 on Hugging Face

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

EvoDuet co-evolves solutions and web queries via a retrieval gate to boost LLM discovery gains up to 82.3% across optimization tasks.

Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 109 on Hugging Face · Code ★ 4

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
89%Must read
?Must readVote to see the score

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Multilingual GSM-Symbolic introduces matched math problems across 15 languages to show model size, resource level, reasoning, and typology determine cross-lingual transfer, with size and reasoning closing low-resource gaps but not typological ones.

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart and 21 more

Published Oct 2, 2026 · 0 citations · ▲ 47 on Hugging Face · Code ★ 4

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

Sharpening Tax in Post-Training

Post-training sharpens base model behaviors at the cost of solution coverage, introducing a quantifiable "Sharpening Tax"; a posterior-tempered group sampler reduces this tax while boosting accuracy.

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov and 6 more

Published Oct 1, 2026 · 0 citations · ▲ 102 on Hugging Face · Code ★ 25

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

WorldAuditBench benchmarks interactive 3D world auditing with multimodal agents, finding success rates of 6.6% to 42.3% versus 83.4% human performance.

Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 102 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 2/5
86%Must read
?Must readVote to see the score

World Embedding Benchmark

World Embedding Benchmark evaluates physical encoding via 8,000 simulation cases, finding alignment trades off against quantitative recoverability and retrieval improves video generation fidelity.

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 44 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

PoS maintains explicit belief states for long-horizon LLM agents, detects belief trapping, and recovers to achieve top results across four benchmarks.

Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 93 on Hugging Face · Code ★ 28

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

RSIGame uses recursive self-improvement with local and global loops to autonomously refine generated games, surpassing one-shot GPT-5.5 scores while cutting generation tokens by 11x.

Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou and 9 more

Published Sep 30, 2026 · 0 citations · ▲ 91 on Hugging Face · Code ★ 125

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Hierarchical Continuous Diffusion Language Models

HC-DLM couples discrete token generation with a continuous latent trajectory via a unified variational denoising objective, outperforming diffusion baselines on Sudoku, Countdown, and language modeling.

Hui Ren, Zihan Li, Chang Liu, Huidong Liu and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 89 on Hugging Face · Code ★ 57

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

World Observer jointly generates actor and panoramic observer views to continuously model out-of-view dynamics via shared geometric warping and observer sinks.

Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 85 on Hugging Face · Code ★ 30

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler automates curriculum learning for agent harness optimization via non-stationary bandits that adapt training scenarios to evolving failure patterns, boosting Pass@1 by 4.4, 7.5 points.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han and 7 more

Published Oct 1, 2026 · ▲ 82 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems

SciUtopia is a closed-loop LLM simulation framework modeling entire academic ecosystems; it finds resubmission amplifies reviewer burden, cautious exploration balances impact and diversity, and inequality can emerge without cumulative funding advantage.

Yiqiao Jin, Yiyang Wang, Lucheng Fu, Bing He and 7 more

Published Oct 1, 2026 · 0 citations · ▲ 40 on Hugging Face · Code ★ 29

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Rhetorical robustness requires stable judgments across content-preserving rewrites and discrimination across papers; SciCore improves both via dual-branch science-core review.

Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 76 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

ROWBench: Do Video Models Render What the Program Specifies?

PROWBench evaluates video models' fidelity to program-specified world events via replayable world records and VLM-based logic-render and interaction metrics.

Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu and 5 more

Published Oct 1, 2026 · 0 citations · ▲ 70 on Hugging Face · Code ★ 31

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

LLM agents show strong source preferences across search domains that can override item quality, though supplying missing information or countering preconceptions reduces this bias.

Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 37 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Frontier models follow unreliable external guidance; training improves selective reliance, identifying it as a key agent reliability dimension.

Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

EgoTools introduces a dataset and benchmark for egocentric tool-use reasoning, showing current models struggle with visual grounding while training improves performance.

Shulin Tian, Junsu Kim, Shuai Liu, Hao Li and 16 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face · Code ★ 9

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

DARA corrects batch-level reward imbalance via inverse-square-root active-group density weights, accelerating multi-reward RL training by up to 65% with no objective change.

Tong Zheng, Skylar Zhai, Zhan Cheng, Tianming Sha and 6 more

Published Sep 30, 2026 · 0 citations · ▲ 64 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

OSWorld-Science benchmarks VLM agents on 146 expert scientific software tasks, showing state-of-the-art models still struggle with scientific workflows and harness design.

Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li and 27 more

Published Sep 30, 2026 · 0 citations · ▲ 63 on Hugging Face · Code ★ 5

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

InterEvolve evolves reward programs at test time via an LLM agent and numerical optimizer to compose a humanoid controller's skills for novel loco-manipulation tasks. The approach releases latent controller competence through object-aware forward-backward models, generating novel strategies that tra

Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

AutoGUIWorld uses image generators to synthesize GUI interaction trajectories without running software, improving OSWorld scores to 40.8% and ScienceBoard success to 32.2%.

Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang and 17 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face · Code ★ 9

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision JEV models show ordinal scale-utilization bias, compressing decisions to 26, 76% of gold support despite high accuracy, but BA-LoRA post-training improves utilization to 86%.

Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 60 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

Video Generation Models: A Survey of Post-Training and Alignment

This survey reviews post-training and alignment strategies for video generation, framing them as implicit or explicit alignment across four methodological categories to improve controllability and reliability.

Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im and 9 more

Published Sep 30, 2026 · ▲ 60 on Hugging Face · Code ★ 208

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
83%Must read
?Must readVote to see the score

Scaling and Distilling Text Embeddings for Better Diffusibility

Scaling and distilling text embeddings improves latent diffusion by yielding more connected, diffusible spaces that boost generative performance beyond autoregressive baselines.

Zekai Zhang, Yunjie Tian, Yanjin He, Xiaoyan Zhang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

CheckerBench evaluates long-horizon agents on synthesizing static-analysis checkers across 300 CVE-derived tasks, finding best Pass@1 reaches 45.33%.

Hang He, Li Wang, Hao Chen, Yuchen Shao and 8 more

Published Oct 6, 2026 · ▲ 52 on Hugging Face · Code

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Beyond the Current Scene: Event-Referential Grasping with Active View Selection

BeyondSCe enables zero-shot event-referential grasping via active view selection, achieving 76% and 77% success on visible and occluded targets versus 40% and 55% baselines.

Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 51 on Hugging Face · Code ★ 8

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

MemAdapter counters memory-induced sycophancy via counterfactual induction, context-aware reflection, and evidence-based reasoning to improve memory reliability across diverse scenarios.

Ruqing Ning, Haibo Meng, Zhishang Xiang, Zerui Chen and 3 more

Published Oct 4, 2026 · ▲ 27 on Hugging Face · Code ★ 21

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Ego2Act evaluates egocentric video generation on multi-step goal-directed manipulation, showing models skip steps and fail at fine-grained physical dynamics.

Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Language Models that Play Chess and Explain Their Moves

Queen, a 4B-parameter chess-language model, plays at grandmaster level and explains moves via cross-attention to a silent expert encoder and iterative Bellman-style explanation distillation, surpassing larger frontier models.

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Published Oct 2, 2026 · 0 citations · ▲ 31 on Hugging Face · Code ★ 18

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

In-Distribution Forcing for Long Video Generation at Test Time

In-Distribution Forcing prevents out-of-distribution key-value drift via self-caching to extend short video models to minute-scale generation.

Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 29 on Hugging Face · Code ★ 3

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ProAR: Learning Prospective Reasoning with Autoregressive Video Models

ProAR introduces goal-frame prediction and future self-alignment to enable goal-directed reasoning in autoregressive video models, surpassing baselines with 25% training steps.

Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen

Published Oct 2, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Local Support Learning

Local Support Learning pairs weight adapters with GMM gating to keep updates local, resolving catastrophic forgetting in LLMs up to 7B parameters without prior data.

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Published Oct 1, 2026 · 0 citations · ▲ 25 on Hugging Face · Code ★ 13

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

LMBuild evaluates LLM agents on generating buildable, functional 3D structures and finds physical operability and functional affordance remain challenging despite improved soundness.

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang and 8 more

Published Oct 3, 2026 · ▲ 24 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
84%Must read
?Must readVote to see the score

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

4DCodeBench benchmarks agents reconstructing dynamic scenes from video as executable graphics code, finding strong static models fail at complex dynamics.

Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang and 5 more

Published Oct 2, 2026 · 0 citations · ▲ 23 on Hugging Face · Code ★ 67

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

MetaRubric fixes vacuous rubric credit via evidence-aware optimization and counterfactual rubric adaptation, improving PubMedQA accuracy by up to 20.40 points over static-judge GRPO.

Yuxuan Fan, Jaehong Yoon

Published Oct 2, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

FrugalEvo pairs expensive LLM strategy exploration with cheap LLM implementation and caching to maximize optimization gain per cost under a budget, outperforming baselines on 10 tasks at significantly lower expense.

Hui Chen, Xuan Qi, James Zhao, Zhaopeng Feng and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 5

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

EVISKILL: Grounding Skill Evolution in Replayable Evidence

EVISKILL grounds LLM skill evolution in replayable evidence cards linking edits to supporting contexts, using targeted replay for verification and global validation for incorporation.

Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang and 1 more

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 8

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Learning to Predict Distributions over Weight Updates for Test-Time Adaptation

Query-conditioned hypernetworks predict distributions over LoRA weight updates from input queries, enabling test-time scaling via sampled adapted models that outperform deterministic and token-sampling baselines.

Azal Ahmad Khan, Keshav Ramji, Tahira Naseem, Ali Anwar and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise

POTER uses optimal transport geometry between training and reference distributions to downweight mislabeled or shortcut-aligned samples, achieving state-of-the-art worst-group accuracy with a single training stage.

Sung Ho Jo, Seonghwi Kim, Wonsang Yun, Minwoo Chae

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published Oct 1, 2026 · 0 citations

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

73%Highly rated
?Highly ratedVote to see the score

Structure-agnostic Causal Representation Learning

SaCRL jointly identifies causal structure and learns invariant representations via soft optimization over HSIC-based invariance violations without prior structural knowledge. It guarantees structure identification, invariance satisfaction, and out-of-distribution generalization while achieving state

Arman Behnam, Binghui Wang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Diptych: Scoped, AI-Interpreted Comparison for Reference Listening in Music Production

Diptych lets musicians define comparison scopes for reference listening, helping surface differences experts partially support while avoiding overreaching AI judgments.

Chongjun Zhong, Abhinaba Roy, Archishman Ghosh, Kejun Zhang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 20 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 1/5
90%Must read
?Must readVote to see the score

From Evidence to Action: How Tool-Using Agents Fail

Tool-using agents often fail by acting before establishing required evidence or leaving multi-action workflow prerequisites unresolved, despite accurate static action assessment. SafeActBench reveals failures stem from how agents use established evidence during execution, not just missing informatio

Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai and 2 more

Published Oct 6, 2026 · ▲ 20 on Hugging Face · Code ★ 3

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
90%Must read
?Must readVote to see the score

World Action Learning via Interaction-Centric Spectral Latent Guidance

WING distills interaction-centric latent actions from egocentric video and uses cross-embodiment spectral low-frequency guidance to transfer them to robot policies, achieving high success rates on LIBERO, RoboTwin, RoboCasa, and real-world tasks.

Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 20 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Self-generated feedback in long-horizon test-time training causes weight updates that improve synthetic text but degrade real-text prediction, and settlement on independent evidence prevents this failure.

Cheng Luo, Bing Li, Bernard Ghanem

Published Oct 4, 2026 · ▲ 19 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models use fixed-point convergence to enable truncated training, shared KV caches, faster prefill, and faster RL updates, while a learned depth prior and orthogonal input injection improve perplexity across scales.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen and 3 more

Published Oct 5, 2026 · ▲ 17 on Hugging Face · Code ★ 30

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

OmniReasoning introduces a benchmark, data engine, and self-distillation method to improve audio-visual joint reasoning, boosting Qwen3-Omni-30B-A3B-Thinking by up to 12.8 points.

Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu and 10 more

Published Sep 30, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT online test-time trains agents by self-distilling verified deployment trajectories into LoRA weights via a frozen hindsight model, improving long-horizon success and efficiency without external teachers or memory retrieval.

Haodong Lu, Dong Gong

Published Oct 4, 2026 · ▲ 17 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

AutoSciBench autonomously generates and iteratively adapts scientific agent benchmarks via recipes and concepts, reducing solver accuracy by over 22 points versus human benchmarks while improving quality ratings.

Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards and 6 more

Published Oct 4, 2026 · ▲ 17 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

Skill2Real learns simulation skills via a shared robot API with a Proposer-Verifier-Governor loop, transferring frozen skill hierarchies to real robots without fine-tuning to reach 78.75% real-world manipulation success.

Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 16 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

CANOPY adaptively compresses multimodal evidence via hierarchical region scoring and targeted retrieval, improving QA accuracy while reducing input tokens by up to 27.7%.

Hyojeong Yun, Jueun Kim, Wook-Shin Han

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Representation-Space MMD for Diffusion Language Models

Post-training minimizes representation-space MMD between diffusion language model outputs and references via retained token features, improving perplexity, accuracy, and parallel decoding.

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev and 6 more

Published Oct 5, 2026 · ▲ 16 on Hugging Face · Code ★ 10

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 1/5
83%Must read
?Must readVote to see the score

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

LOOM stabilizes looped MoE via bounded residual updates and per-loop routers to scale loops to 9, 12, cutting perplexity from 9.62 to 7.77.

Di He, Pengxiang Li, Da Chang, Qingyan Meng and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face · Code ★ 8

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

FailBank turns runtime shield feedback into persistent VLA policy updates via failure-bank self-evolution, raising success rates up to 25.4 points and cutting policy-induced cost up to 35.6%.

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 15 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
91%Must read

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

OnePO enables RL-only medical domain adaptation via adaptive objective evolution and teacher retirement, yielding HuatuoGPT-3 that surpasses frontier models.

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu and 6 more

Published Oct 5, 2026 · ▲ 15 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
86%Must read
?Must readVote to see the score

UNREAL: Unifying Retrieval and Long-Context with a Single Model

UNREAL unifies retrieval and long-context evidence selection via frozen LLM representations with minimal parameters, outperforming state-of-the-art retrievers and improving long-context accuracy substantially.

Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel and 3 more

Published Oct 6, 2026 · ▲ 15 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
80%Must read
?Must readVote to see the score

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

SearchJev is a fast calibrated System-1 model that scores search decisions directly without autoregressive generation, improving decision quality, speed, and calibration over same-size language models.

Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu and 5 more

Published Oct 4, 2026 · ▲ 14 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

Collective Bias Mitigation via Model Routing and Collaboration

Collective Bias Mitigation routes queries among diverse LLMs and fosters collaboration to substantially reduce bias over single-model baselines.

Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 14 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Targeted bias injection via closed-loop activation steering exploits diffusion language model denoising trajectories to steer frozen models toward adversarial demographic answers with minimal corruption.

Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed, Muhammad Umer Sheikh and 2 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
83%Must read
?Must readVote to see the score

World Editing: Intervening on Executable Worlds at Increasing Depth

World editing intervenes on executable environments at increasing depth via IGMWorld and IGMBench, where top agents achieve 78.2% task success with reliability declining by depth.

Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh and 14 more

Published Oct 1, 2026 · ▲ 14 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

RobotUse: Allocating Computation, Context, and Decisions

RobotUse organizes robot computation, context, and decisions around revisable physical actions via visual target selection and persistent playbooks, achieving 45% RoboLab success and real-world learning.

Junhoo Lee, Injun Baek, Seungyeon Kim, Suhyun Jeon and 3 more

Published Oct 4, 2026 · ▲ 14 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation integrates RL teacher gradients via loss averaging, Adam smoothing, and BF16 rounding, with averaging rules significantly altering math accuracy outcomes.

Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems

Benign multi-agent LLM planners disguise secrets to help developers evade oversight, with rare per-episode leaks compounding to high breach risk across repeated exchanges.

Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng

Published Sep 30, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read

Certification of Real Images through Calibrated Content Authentication

Deepfake detectors degrade to 76% accuracy and near-zero under attacks, so calibrated reconstruction-based authentication bounds false real-image certification to 1%.

Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi and 1 more

Published Oct 5, 2026 · ▲ 13 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

SEER uses self-evolving event reasoning and retrieval to dynamically optimize forecasting with exogenous events via reflective memory and causal knowledge, outperforming state-of-the-art baselines.

Mingtian Tan, Palash Goyal, Mihir Parmar, Sarkar Snigdha Sarathi Das and 5 more

Published Oct 2, 2026 · ▲ 13 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Efficient reasoning training differs in impact: faithfulness usually drops due to inconsistency, but monitorability remains robust.

Samuel Lewis-Lim, Xingwei Tan, Mario Sänger, Zhixue Zhao and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

PerturbBot breaks vision-language-action shortcut priors via perturbative training while GroundingFscore diagnoses shortcut reliance, enabling healthier scaling without altering inference.

Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang and 3 more

Published Oct 3, 2026 · ▲ 13 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
88%Must read
?Must readVote to see the score

When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

RACE predicts subskill transition timing to extend VLA action chunks reliably, reducing stop-and-go idle time ~5x on real robots while improving success rates.

Seonghoon Yu, Dongwon Kim, HyungRok Jung, Yoonjae Baek and 3 more

Published Oct 5, 2026 · ▲ 13 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
91%Must read

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

LoHi trades per-frame resolution for denser temporal sampling via low-resolution streams plus sparse high-resolution frames, boosting long-video accuracy up to 10.6 points and cutting front-end latency up to 7x.

Sixun Dong, Wei Li, Andong Deng, Qi Qian and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Oct 3, 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
89%Must read
?Must readVote to see the score

AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

AGO is an evidence-first quality gate for RAG that treats judge errors and missing data as explicit outcomes, using stratified beta-binomial gates and mandatory meta-evaluation to reduce unsafe promotion to 22.2%-35.1% versus 29.3%-41.8% for naive gates.

Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero, Giuseppe Santoro and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated

RealtimeWAM: One-Step Asynchronous World Action Models

RealtimeWAM uses teacher-anchored consistency distillation and cross-expert wavefront pipelining for one-step asynchronous action generation, achieving near-lossless performance with ~25x speedup.

Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong and 6 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 2,880

0% Readers0 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

Base Models Can Reason By Taking a Cue From Training Data

Fixing initial token cues in base models boosts reasoning to match RL performance, with effects traced to training data associations that can be causally edited.

Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat and 2 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

UndoBench separates tool-using agent competence from fault recovery via paired enterprise workflow trials, finding 83.54% nominal success but only 46.72% recovery success with phase-dependent vulnerabilities.

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5