Good Papers

Trending at NeurIPS 2026

Orals, spotlights and posters people are talking about

All sessions

Must read

This year's highest-rated papers

See all

Most debated

Where the reviewers can't agree

See all

All papers

How scores work
78%Highly rated
?Highly ratedVote to see the score

Agta hunter-gatherer oral microbiomes are shaped by contact network structure

Agta hunter-gatherer oral microbiomes resemble Central African foragers more than neighbors, with contact networks predicting bacterial transmission and central individuals as supersharers.

Federico Musciotto, Begoña Dobón, Michael John Greenacre, Álex Mira and 12 more

Published Dec 31, 2030 · 0 citations

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

EngramEdit updates LLM facts via conditional memory by computing target representations and penalizing shared embedding changes, achieving near-perfect edits with preserved unrelated knowledge.

Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu and 3 more

Published Oct 7, 2026 · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

PersonTTS amortizes agentic test-time scaling policy discovery across personalized multi-dimensional accuracy, latency, and cost requirements via experience reuse and distilled procedural guidance.

Xinglin Wang, Zishen Liu, Tong Zheng, Shaoxiong Feng and 8 more

Published Oct 7, 2026 · ▲ 9 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026RL for LLMs

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

SkillForge co-evolves LLM agent skills via fitness-driven lifecycles of trial, active, stable, and retired states, achieving up to 7.8% relative success improvement over baselines.

Yuyao Ge, Yiwei Wang, Yuchen He, Baolong Bi and 4 more

Published Oct 7, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

RoboQuest: Generalist Physical Agents that Search, Inspect and Test

RoboQuest benchmarks embodied agents on search, inspection, and testing tasks requiring interactive evidence gathering, finding frontier models succeed in only 23% of episodes and mostly fail by exploring too little.

Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria

Published Oct 7, 2026 · ▲ 3 on Hugging Face · Code ★ 1

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

QuadTok uses a hierarchical quadtree tokenizer to dynamically allocate tokens by visual complexity, achieving efficient autoregressive image generation with 2.08 gFID on ImageNet 256×256.

Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang and 3 more

Published Oct 7, 2026 · ▲ 8 on Hugging Face · Code

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

UltraText Bench is a bilingual benchmark of 432 prompts for evaluating dense visual text rendering in image generation across fidelity, clarity, placement, and scene quality. It reveals trade-offs across 24 model configurations and sharp performance declines from easy to hard workloads.

Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li and 12 more

Published Oct 7, 2026 · ▲ 12 on Hugging Face · Code ★ 3

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation

AdSpark introduces a 300K dataset and six-dimension benchmark for product-centric ad video generation, revealing key challenges in product preservation and multi-shot storytelling.

Zhifei Yang, Zhao Jiang, Keyang Lu, Honghe Zhu and 5 more

Published Oct 7, 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

RobotWorld benchmarks multimodal agents on 84 simulated robot tasks spanning manipulation to aerial control, revealing sophisticated but unreliable physical-world execution that varies across models.

Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu and 29 more

Published Oct 7, 2026 · ▲ 17 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

RunningTab introduces an environment-side tab tracking task requirements, file excerpts, and unread candidates for direct workspace interaction, improving agent deliverable accuracy over model-side tracking.

Jinheon Baek, Soyeong Jeong, Yumin Choi, Dongsu Han and 1 more

Published Oct 7, 2026 · ▲ 21 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Tetris3D: 3D Scene Generation With Objects That Fit Together

Tetris3D explicitly conditions each object's shape and pose on neighboring geometry and physical relations to recover physically coherent 3D scenes, achieving state-of-the-art generation and stability.

Jaeyeong Kim, Jinhyuk Jang, Jongmin Lee, Kyehong Park and 1 more

Published Oct 7, 2026 · ▲ 25 on Hugging Face · Code ★ 11

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score
arXivPrivacy

Inverting Multi-Vector Visual Document Indices

Multi-vector document indices can be inverted to recover readable pages because patch vectors retain document layout and vision-language features, achieving high source retrieval and substantial text reconstruction.

Zhuchenyang Liu, Yao Zhang, Yu Xiao

Published Oct 7, 2026 · ▲ 14 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Long-WAM: Scaling the Context of World-Action Models

Long-WAM scales causal world-action model context via autoregressive pretraining, raising robot success up to 78.7% and enabling real-time deployment.

Wei Huang, Bohan Zhang, Chenzhi Liu, Isabella Liu and 12 more

Published Oct 7, 2026 · ▲ 23 on Hugging Face · Code ★ 2,662

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

SGF+ separates context-writing and denoising parameters to fix conflicting gradients, improving autoregressive video quality and enabling continuous generation up to 24 hours without long-video training.

Zihan Su, Junhao Zhuang, Yaowei Li, Siwen Lu and 9 more

Published Oct 7, 2026 · ▲ 23 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

GRACE: Generation-aware latent compression for efficient video generation

GRACE compresses pretrained video autoencoders via frozen base latents and learned residuals aligned to frozen DiT features, cutting Wan2.1 tokens 8x and latency 11.1x without quality loss.

Jiyoung Kim, Paul Hyunbin Cho, Jisu Nam, Donghoon Lee and 4 more

Published Oct 7, 2026 · ▲ 35 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

Hybrid attention models combine full and sliding-window or linear attention, revealing a context-extension seesaw effect, positional biases causing short-context traps, and a sliding-window linear attention method achieving 16x training-free length extrapolation with perfect NIAH retrieval at 64k co

Xiaoran Liu, Ziwei He, Xipeng Qiu

Published Oct 7, 2026 · ▲ 1 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

RLHND repurposes a video diffusion backbone to track physically consistent hand poses and estimate dense tactile contact and force from monocular egocentric video, achieving state-of-the-art results and improving robot learning retargeting.

Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo and 1 more

Published Oct 7, 2026 · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents

This paper constructs an abstract oracle machine unifying agents and workflows via a stack-query automaton, and implements it as ArchNights.

Kefan Liu, Fengning Ou, Yelin Luo, Jingdi Lei

Published Oct 7, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Co-Evolving Robot Orchestrators and Policies through Deployment

Robo-COP co-evolves robot orchestrators and policies during deployment, improving held-out success to 73.8% in simulation and 50.0% in real-world tasks.

Xilun Zhang, Maggie Wang, Erik Bauer, Hong-Xing Yu and 3 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

DLoop: Looped Speculative Decoding

DLoop loops multiple speculative drafting stages before target verification, reducing target model passes and improving wall-clock speedup by 5, 41% losslessly.

Geonmo Gu, Byeongho Heo, HeeJae Jun, Yoohoon Kang and 3 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

PhysEvo: Astra Can Act, Let It

PhysEvo enables frozen-model robotic self-improvement via recursive tool and skill revision, achieving 62% success on 42 tasks and 84% on real-world manipulation.

Wenqing Tian, Zeyu Zhang, Zhaocheng Liu, Fengwei Liu and 2 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

On-Policy Distillation with Negative-Policy Rollouts

Negative-Policy OPD improves on-policy distillation by using negative-policy rollouts to supply explicit negative signals, boosting performance across scales and reasoning tasks.

Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo

Published Oct 6, 2026 · ▲ 11 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Recurrent Looped Transformer

Recurrent Looped Transformer merges a parallel encoder with a recurrent decoder to grow per-token computation depth with sequence length, achieving near-perfect length generalization on parity and permutation tasks versus chance-level Transformers.

Yifan Zhang, Jichen Feng, Shihan Qin

Published Oct 6, 2026 · ▲ 16 on Hugging Face · Code ★ 908

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness

Recursive Game Creator uses recursive Designer-Builder-Player-Reviewer loops to advance agentic games from prototypes to entertaining products, achieving 77.89 on GameCraft-Bench and 53.2% success on GameASG-Bench with higher user ratings.

Jiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu

Published Oct 6, 2026 · ▲ 21 on Hugging Face · Code ★ 6

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

DecepEval: A Benchmark for Evaluating Deception in LLM Agents

DecepEval introduces a benchmark of 1,532 instances across 28 scenarios and proposes a framework showing inducements increase LLM deception rates.

Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen and 7 more

Published Oct 6, 2026 · ▲ 25 on Hugging Face · Code

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

WorldSonus: Bringing Sound to Worlds

WorldSonus generates interactive, spatially aligned stereo audio for world models via streaming causal diffusion with real-time factor 0.41, matching state-of-the-art quality on open benchmarks.

Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang and 8 more

Published Oct 6, 2026 · ▲ 21 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

On KL-Regularized Policy Optimization

KLPO anchors KL regularization to the sampler for closed-form updates without importance weights, critics, or grouped rollouts, and subsumes SPPO, GPO, REBEL, and BPO.

Yifan Zhang

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code ★ 182

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

nanoMuse: An Open-Source Personal Agent for Every Device You Own

nanoMuse is an open-source GPL-3.0 personal agent running on every device with shared memory, local control, and a user-chosen model.

Guangyi Liu, Yong Liu, Jiangning Zhang

Published Oct 6, 2026 · ▲ 6 on Hugging Face · Code ★ 229

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

CheckerBench evaluates long-horizon agents on synthesizing static-analysis checkers across 300 CVE-derived tasks, finding best Pass@1 reaches 45.33%.

Hang He, Li Wang, Hao Chen, Yuchen Shao and 8 more

Published Oct 6, 2026 · ▲ 52 on Hugging Face · Code

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine co-evolves policies, curricula, and judges via adaptive diagnosis and coactive calibration to sustain embodied reasoning self-improvement. Experiments on driving and navigation show continuous gains in both policy and judge capability.

Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao and 9 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

Attacca trains visual goal-conditioned policies on complete search-to-interact trajectories with decoupled goal images and behavioral-phase conditioning to improve long-horizon embodied task success by up to 7x.

Gyusik Seo, Jaehong Yoon

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code ★ 3

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

UNREAL: Unifying Retrieval and Long-Context with a Single Model

UNREAL unifies retrieval and long-context evidence selection via frozen LLM representations with minimal parameters, outperforming state-of-the-art retrievers and improving long-context accuracy substantially.

Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel and 3 more

Published Oct 6, 2026 · ▲ 21 on Hugging Face

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Sensor-Language-Action Models

Sensor-Language-Action modeling unifies multimodal sensors, language, and actions via a semantic interface, and OpenSLA achieves superior hierarchical prediction and explanation with zero-shot generalization.

Yuekai Xu, Zitao Shuai, Yuzhe Yang

Published Oct 6, 2026 · ▲ 1 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

90%Must read
?Must readVote to see the score

World Models' Last Exam in Physics

World Models' Last Exam in Physics benchmarks video models via 40 measurement-based physics tasks, finding the best model scores 57.76/100 with widespread inconsistencies.

Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin and 7 more

Published Oct 6, 2026 · ▲ 2 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

CtrlCache accelerates interactive video world models via control-aware caching that detects action changes to reuse transformer residuals and apply frequency-mixed history guidance, achieving up to 1.41x speedups with improved quality.

Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh and 1 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face · Code

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Sherpa: Teaching LLMs to Teach Adaptively

Sherpa uses multi-turn reinforcement learning to train LLM teachers that adapt instructions to diverse student archetypes, improving student performance by 20.5 points and pedagogy scores to 79.2%.

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen and 1 more

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

TRACE aligns FP4 quantization between RL training and rollout paths via rollout-guided quantization-aware training for MoE language models, achieving BF16-comparable RL performance with up to 5.4x rollout speedup.

Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao and 8 more

Published Oct 6, 2026 · ▲ 80 on Hugging Face

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

90%Must read
?Must readVote to see the score

From Evidence to Action: How Tool-Using Agents Fail

Tool-using agents often fail by acting before establishing required evidence or leaving multi-action workflow prerequisites unresolved, despite accurate static action assessment. SafeActBench reveals failures stem from how agents use established evidence during execution, not just missing informatio

Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai and 2 more

Published Oct 6, 2026 · ▲ 35 on Hugging Face · Code ★ 6

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Cross-tokenizer on-policy distillation achieves comparable accuracy with strict top-16 shared-vocabulary supervision versus full coverage, while expanded span supervision reduces accuracy due to conflicting gradients, motivating prioritization of supervision reliability over alignment coverage.

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li and 3 more

Published Oct 6, 2026 · ▲ 157 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS bootstraps reusable agent memory from self-generated practice tasks without oracles, improving success rates by up to 15.9 points across benchmarks.

Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard and 2 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Building Rome from a Single Image

A redesigned object-centric generator partitions scenes into distance-adaptive chunks, captures 2D-3D correspondence, and trains on 4,000 outdoor scenes to outperform baselines in indoor and outdoor mesh generation.

Jiraphon Yenphraphai, Fang Li, Tianshuo Xu, Depu Meng and 4 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Towards In-Parameter Memory Augmentation for Large Language Models

This survey organizes in-parameter memory augmentation for LLMs by parameter placement and acquisition time to enable reusable parametric knowledge at deployment.

Haoyu Huang, Zhongwei Xie, Jiaxin Bai, Yisen Gao and 5 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Personal-Agent Mediated Recommendation with Cross-Platform User History

Personal-agent mediated recommendation balances cross-platform user history against platform rankings via the MediateRec benchmark and PAMO optimization to improve rescue-harm trade-offs.

Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley and 1 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

AdvSim2Real co-evolves tasks, adversarial injections, and a web agent in a simulator, boosting 4B agent completion by 33.6% against unseen adaptive attacks and transferring gains to real browsers.

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov and 2 more

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

NeMo-DCR bit-exactly refits trillion-parameter policies by streaming delta-compressed weight changes via affine mappings and XOR masks, cutting 1T cross-region refits from 87.5 minutes to 150 seconds.

Songlin Jiang, Zhiyu Li, Terry Kong, Yu Yao and 4 more

Published Oct 6, 2026 · ▲ 10 on Hugging Face · Code ★ 2,048

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation

EmbodiedSmith unifies asset, scene, and task generation in a recursive self-improvement loop to scale embodied simulation data, improving generation success and robot policy generalization across diverse embodiments and physics.

Yikai Qin, Yifei Deng, Mingjian Liang, Wenxuan Song and 12 more

Published Oct 6, 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

HiPLEX factorizes full-duplex speech policies into timing and content controllers to jointly optimize interaction dynamics via reinforcement learning. It lowers takeover rates, reduces interruption latency, and improves human-like turn timing versus GRPO.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Speculative tool execution predicts tool calls from partial ASR to run them during speech, cutting median voice-agent response latency from 5.79 s to 4.60 s.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026 · ▲ 6 on Hugging Face

100% Readers1 of 1 upvoted
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

ALIVE inserts objects that interact with video contents via first-frame editing and interaction guidance, outperforming baselines on interaction and insertion benchmarks.

Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou

Published Oct 6, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Dev-Primitives turn repository artifacts into active, resident-LLM agents with self-modification interfaces, and HERMES improves software engineering benchmarks by 12.4% over baselines while cutting inference costs by 26.2%.

Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang

Published Oct 6, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

D-OPCD distills agent harness improvements into diffusion weights via on-policy context distillation, raising direct text-to-image scores from 60.52 to 65.09 and enabling continual co-evolution.

Wenxuan Wang, Zekai Liu, Weinan Zhang, Yu Cheng and 1 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

VepAgent integrates causal-transition reasoning with tool-augmented reinforcement learning for video event prediction, achieving state-of-the-art results on FutureBench and NEPBench.

Qiutong Chen, Yuchan Guo, Zhenlong Yuan, Haobo Yang and 9 more

Published Oct 5, 2026 · 0 citations · ▲ 5 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Improving Proactive AI Assistance with Hierarchical Procedural Understanding

ProactiveCoach introduces hierarchical procedural data and an adaptive guidance router that improves proactive AI assistance timing and granularity adaptation by up to 57.1%.

Jin-Seop Lee, TaeYeon Won, SeongJun Jung, Junghoon Kim and 4 more

Published Oct 5, 2026 · 0 citations · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Minimal Witness Reinforcement Learning

MWRL learns minimal sufficient witnesses via union-based credit assignment to recover diverse alternatives from black-box verifiers, outperforming methods that yield single or redundant solutions.

T. Y. Tsui, Zihao Ye, Pengxiang Cai, Yanchao Li and 2 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face · Code

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

TRIAGE stabilizes native NVFP4 RL by direction-aware mismatch diagnosis and selective gradient rebalancing, achieving full-precision math reasoning with 2.3x rollout throughput.

Zhen Li, Shuai Zhang, Yanggan Gu, Yiming Zhang and 6 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Structuring MoE Expert Selection for Agentic Reinforcement Learning

Agentic RL on MoE models reveals structured expert routing by operation type; hierarchical routing control and entropy gating improve success rates over 10 points.

Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh, Sanjoy Chowdhury and 2 more

Published Oct 5, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco

Turba provides an open-source Moroccan fertilizer recommendation stack with programmatic access, versioned data, and loadable crop-specific machine learning surrogates for reproducible benchmarking.

Abdelghani Belgaid, Zakaria Mahmoud, Fahd Chibani, Oumnia Ennaji and 2 more

Published Oct 5, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

58%Worth a look
?Worth a lookVote to see the score

Empirical Variational Autoencoder

Empirical Variational Autoencoder learns autoregressive latent priors empirically via one linear layer to close the VAE prior-posterior gap, yielding high-fidelity sequential generation competitive with diffusion models at much faster inference.

Kaede Shiohara

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code ★ 6

0% Readers0 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Execution-Aligned Progressive Noise for Consistent Asynchronous Replanning in Generative Robot Policies

Execution-Aligned Progressive Noise structures noise across chunks and time to maintain consistent generative states during asynchronous replanning, achieving up to 96.7% real-robot success.

Di Wu, Ping Liu, Xuhua Chen, He Zheng and 2 more

Published Oct 5, 2026

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

Stepped MoE unifies elastic architectures and sparse gating to adapt model capacity to deployment constraints and input requirements, outperforming dense counterparts by 2-5%.

Arnav Kundu, Zhaoyang Xu, Bairu Hou, Chang Gao and 2 more

Published Oct 5, 2026 · ▲ 1 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated

RealtimeWAM: One-Step Asynchronous World Action Models

RealtimeWAM uses teacher-anchored consistency distillation and cross-expert wavefront pipelining for one-step asynchronous action generation, achieving near-lossless performance with ~25x speedup.

Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong and 6 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face · Code ★ 2,881

0% Readers0 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

OnePO enables RL-only medical domain adaptation via adaptive objective evolution and teacher retirement, yielding HuatuoGPT-3 that surpasses frontier models.

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu and 6 more

Published Oct 5, 2026 · ▲ 22 on Hugging Face · Code ★ 17

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026RL for LLMs

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

RGPO adaptively uses ground-truth rationales as temporary scaffolds to generate improved responses for on-policy RL, then transfers only higher-reward model outputs back, reducing reward sparsity and improving text and multimodal reasoning.

Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde and 2 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

WildMatch adapts pretrained image matchers to wildlife identification using only identity labels, improving retrieval accuracy and learning transferable matching priors without keypoint annotations.

Turhan Can Kargin, Piotr Kubaty, Ekaterina Rostovskaya, Izabela Wierzbowska and 2 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

S2PD switches from autoregressive to parallel diffusion during denoising to enforce physical and logical consistency with faster sampling than fully serial methods.

Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari

Published Oct 5, 2026 · 0 citations · ▲ 1 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

MEND: RL For Flow Models via Proximal Velocity Matching

MEND uses proximal velocity matching to cap rewards and accept only cost-effective sample moves, outperforming prior flow-model RL methods in far fewer updates without KL penalties or reference models.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli and 1 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Conditional Trajectory Peaks: Single-Pass Multimodal Policies over Action Chunks

Conditional Trajectory Peaks predicts multimodal action-chunk candidates in a single pass, achieving 97.25% LIBERO success and 3× faster inference while preserving diverse behaviors.

Di Wu, Rongtian Shen, Ping Liu, Xuhua Chen and 3 more

Published Oct 5, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Learning to Read the Contextual Tokens in Diffusion Transformers

A framework maps diffusion transformer contextual tokens through a frozen LLM to reveal they encode global emerging scene semantics early, inspiring contextual alignment that improves generation quality.

Omer Dahary, Etai Sella, Hadar Averbuch-Elor, Daniel Cohen-Or and 1 more

Published Oct 5, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Agents judge failing retrieval results useless but rarely stop; enforcing answers after five useless judgments improves success and fixes stopping.

Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu and 3 more

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

JLD: Perceptual Distance Through A Jacobian Lens

JLD defines a perceptual image distance via a Jacobian-derived metric tensor from frozen vision encoders, achieving state-of-the-art correlation with human judgments and resolution robustness.

Shreshth Saini, Balu Adsumilli, Alan C. Bovik

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 3

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Hybrid Linear Attention introduces query-dependent chunk-level routing for Gated DeltaNet, improving long-context benchmarks by up to 5.57 points via adaptive recurrent memory composition.

Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He and 2 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

MiniCorp: The Last Mile of the AI Agent Firm

MiniCorp is a simulated office environment that generates longitudinal, counterfactual enterprise data to study autonomous AI-run companies and train adaptive agents.

Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin and 5 more

Published Oct 5, 2026 · ▲ 20 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

66%Highly rated

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

APO learns shared LLM aligner initializations via grouped gradient coordination to enable few-shot personalization under heterogeneous preferences and scarce feedback, improving over baselines with 20 local examples.

Liyan Yang, Yige Yuan, Zhiqin Yang

Published Oct 5, 2026 · ▲ 6 on Hugging Face

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Highly rated
?Highly ratedVote to see the score

InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation

InterMimicGen retargets human motion capture to humanoids and self-evolves tracking data via iterative simulation-verified augmentation to scale dexterous loco-manipulation.

Yucheng Zhang, Sirui Xu, Jinhong Li, Liuyu Bian and 7 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

ACG-Bench evaluates arm-wise compositional generalization in dual-arm vision-language-action models via AE-VLA, which achieves 21.53% simulated and 39% real-world success versus under 6% baselines.

Zaibin Zhang, Binghao Ran, Yuhan Wu, Zhongbo Zhang and 9 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Agentic retrieval combining LLM reasoning with dense retrieval improves nDCG@10 by 8.7 points over standard retrieval but requires 107 seconds and 764K input tokens per query.

Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira and 7 more

Published Oct 5, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

On-policy power distillation trains models to generate sharpened answers directly, improving single-sample math reasoning by up to 27.3 points and outperforming multi-candidate sampling and reward-based methods.

Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa and 3 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

RACE predicts subskill transition timing to extend VLA action chunks reliably, reducing stop-and-go idle time ~5x on real robots while improving success rates.

Seonghoon Yu, Dongwon Kim, HyungRok Jung, Yoonjae Baek and 3 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
89%Must read

Certification of Real Images through Calibrated Content Authentication

Deepfake detectors degrade to 76% accuracy and near-zero under attacks, so calibrated reconstruction-based authentication bounds false real-image certification to 1%.

Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi and 1 more

Published Oct 5, 2026 · ▲ 13 on Hugging Face · Code ★ 3

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

Base Models Can Reason By Taking a Cue From Training Data

Fixing initial token cues in base models boosts reasoning to match RL performance, with effects traced to training data associations that can be causally edited.

Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat and 2 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 6

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

HLA-WM combines geometry-guided retrieval with recurrent linear attention to fix long-range forgetting in video world models, improving 60-second consistency metrics by up to 28.5% with 12× lower memory and no retraining.

Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu and 3 more

Published Oct 5, 2026 · ▲ 9 on Hugging Face · Code ★ 4

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
78%Highly rated

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

LoGRA reduces LLM reinforcement learning memory by up to 45.7% via low-rank gradient sketches and predicted-KL step control, enabling 27B-parameter training on single nodes.

Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li and 4 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Closing the Context Gap: Activation Alignment for Tabular In-Context Learning

Activation alignment trains a linear map to align partial-context student activations with full-context teacher activations, significantly improving tabular in-context learning efficiency and recovering much of the performance gap.

Yoel Zeldes

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
ICLR 2027Tabular data

Adapting prior-data fitted networks for tabular anomaly detection

Frozen and fine-tuned TabPFN representations for tabular anomaly detection yield ZEN and FOCUS, surpassing all ADBench baselines in AUROC despite unsupervised deployment and contaminated reference sets.

Maximilian Bershtman, Niv Cohen

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Learning to Learn a Language

Prior-Fitted Language Model, trained solely on synthetic non-linguistic data, learns to infer and predict real languages from context with frozen weights, achieving strong cross-lingual compression and reasoning without ever seeing real text.

Lennart Carstens-Behrens, Holger Fröhlich

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

SoK: Semantic Decision Engines in Network Control Loops

Systematizing 139 semantic decision engine families reveals most miss network deadlines and verification, with only four reporting deadline attainment; unverified decisions reverse admission verdicts under queued execution, prompting minimum reporting rules and a research agenda.

Delong Li, Chen Li, Xu Wang, Haochen Gong and 2 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Targeted bias injection via closed-loop activation steering exploits diffusion language model denoising trajectories to steer frozen models toward adversarial demographic answers with minimal corruption.

Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed, Muhammad Umer Sheikh and 2 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models

FairRSFM benchmarks remote sensing foundation models by biome to expose hidden ecological performance disparities and tests debiasing methods without backbone updates. Aggregate metrics consistently mask large biome-dependent gaps, though mitigation effectiveness varies by model and task.

Md Aminur Hossain, Omkumar Vaghasiya, Rajeev Ranjan Dwivedi, Vinod Kurmi and 1 more

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

What Matters for Latent Reasoning with Flow Matching

FLaRe uses flow matching for latent reasoning that is useful, diverse, explainable, refinable and efficient, reaching 97% of explicit chain-of-thought accuracy at 25% latency.

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

Published Oct 5, 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Capability-Driven Self-Evolution of Agent Memory

PrisMem drives agent memory self-evolution via capability-specific guidance, dependency-aware selection, and trace-guided integration, outperforming baselines by up to 10.54 points on million-token benchmarks.

Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu and 7 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Representation-Space MMD for Diffusion Language Models

Post-training minimizes representation-space MMD between diffusion language model outputs and references via retained token features, improving perplexity, accuracy, and parallel decoding.

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev and 6 more

Published Oct 5, 2026 · ▲ 24 on Hugging Face · Code ★ 12

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models use fixed-point convergence to enable truncated training, shared KV caches, faster prefill, and faster RL updates, while a learned depth prior and orthogonal input injection improve perplexity across scales.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen and 3 more

Published Oct 5, 2026 · ▲ 20 on Hugging Face · Code ★ 35

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads

Agentic RAG evaluation budgets favor broader question coverage over repeated reads or trajectories, reducing standard error by up to 33% at fixed token cost.

Jingjie Ning, Xueqi Li, Yibo Kong

Published Oct 4, 2026 · 0 citations · ▲ 14 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

77%Highly rated

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Kandinsky 6.0 Video introduces diffusion models that generate synchronized 5-second audio-video clips with lip-sync via a dual-stream CrossDiT architecture, with the 29B-parameter Pro version outperforming its predecessor and matching top competitors in speech quality.

Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin and 36 more

Published Oct 4, 2026 · ▲ 132 on Hugging Face · Code ★ 151

100% Readers1 of 1 upvoted
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting

Mobile-4DGS enables real-time static and dynamic Gaussian splatting on mobile devices via compact appearance modeling, explicit 4D motion, and depth-order reuse, substantially reducing storage and rendering overhead.

Xiaobiao Du, Beixi Hao, Zhen Fang, Tianqing Zhu and 2 more

Published Oct 4, 2026 · 0 citations · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

AutoSciBench autonomously generates and iteratively adapts scientific agent benchmarks via recipes and concepts, reducing solver accuracy by over 22 points versus human benchmarks while improving quality ratings.

Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards and 6 more

Published Oct 4, 2026 · ▲ 27 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

EVISKILL: Grounding Skill Evolution in Replayable Evidence

EVISKILL grounds LLM skill evolution in replayable evidence cards linking edits to supporting contexts, using targeted replay for verification and global validation for incorporation.

Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang and 1 more

Published Oct 4, 2026 · ▲ 36 on Hugging Face · Code ★ 26

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
88%Must read
?Must readVote to see the score

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

Feasible-future decoding reranks VLA actions by future safe-completion mass, reducing cumulative safety costs by up to 57.5% without retraining or rollouts.

Tu Nguyen, Matthieu Zimmer, Vu Anh Vu, Ziyi Wang and 3 more

Published Oct 4, 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

LiFT: Loop Flow Transformers

Loop Flow Transformers loop a shared diffusion transformer with depth-indexed regression targets, improving generation with more inference compute and fewer parameters than dense models.

Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent and 1 more

Published Oct 4, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
91%Must read

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

UndoBench separates tool-using agent competence from fault recovery via paired enterprise workflow trials, finding 83.54% nominal success but only 46.72% recovery success with phase-dependent vulnerabilities.

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published Oct 4, 2026 · ▲ 12 on Hugging Face · Code ★ 2

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

RobotUse: Allocating Computation, Context, and Decisions

RobotUse organizes robot computation, context, and decisions around revisable physical actions via visual target selection and persistent playbooks, achieving 45% RoboLab success and real-world learning.

Junhoo Lee, Injun Baek, Seungyeon Kim, Suhyun Jeon and 3 more

Published Oct 4, 2026 · ▲ 15 on Hugging Face · Code ★ 5

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT online test-time trains agents by self-distilling verified deployment trajectories into LoRA weights via a frozen hindsight model, improving long-horizon success and efficiency without external teachers or memory retrieval.

Haodong Lu, Dong Gong

Published Oct 4, 2026 · ▲ 19 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
71%Highly rated

Code2Games: Enabling Coding Agents for Gaming World Generation

Code2Games coordinates scene analysis and gameplay planning via shared representations to generate consistent gaming worlds and adapt them to Unreal Engine 5. The framework improves visual quality, interactive fidelity, and playable-game quality over direct coding-agent generation on the GameCode4D

Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao and 1 more

Published Oct 4, 2026 · ▲ 6 on Hugging Face · Code ★ 5

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
83%Must read
?Must readVote to see the score

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

DiVeR improves verifier-guided VLA test-time scaling by reweighting learning toward decision-critical states using action representation dispersion, boosting success without extra annotations or overhead.

Seongheon Park, Heecheol Kim, Shulin Tian, Lilika Makabe and 4 more

Published Oct 4, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
80%Must read
?Must readVote to see the score

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

SearchJev is a fast calibrated System-1 model that scores search decisions directly without autoregressive generation, improving decision quality, speed, and calibration over same-size language models.

Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu and 5 more

Published Oct 4, 2026 · ▲ 20 on Hugging Face · Code ★ 8

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read
?Must readVote to see the score

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Self-generated feedback in long-horizon test-time training causes weight updates that improve synthetic text but degrade real-text prediction, and settlement on independent evidence prevents this failure.

Cheng Luo, Bing Li, Bernard Ghanem

Published Oct 4, 2026 · ▲ 23 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

MemAdapter counters memory-induced sycophancy via counterfactual induction, context-aware reflection, and evidence-based reasoning to improve memory reliability across diverse scenarios.

Ruqing Ning, Haibo Meng, Zhishang Xiang, Zerui Chen and 3 more

Published Oct 4, 2026 · ▲ 40 on Hugging Face · Code ★ 27

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Prism introduces dynamic sparse attention via adaptive macro-zone block shapes guided by visual variance and cross-modal attention for 2K joint video-audio generation, yielding 2.5x training speedup and improved quality.

Shuyuan Tu, Qi Tian, Yinming Huang, Yue Wu and 7 more

Published Oct 4, 2026 · ▲ 8 on Hugging Face · Code ★ 56

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
75%Highly rated
?Highly ratedVote to see the score

Large-scale analysis of AlphaFold structures reveals organism-specific physicochemical signatures

Large-scale AlphaFold structure analysis reveals organism-specific physicochemical signatures reconstructed via DE-STRESS metrics across 48 proteomes and PDB structures.

Michael J. Stam

Published Oct 4, 2026 · 0 citations

100% Readers1 of 1 upvoted
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

Self-evolving reasoning models collapse due to invalid and repeated self-generated questions; R-Quest uses validity and novelty feedback to sustain gains across ten rounds and outperform R-Zero by 17.32 points.

Jinyuan Li, Chengsong Huang, Langlin Huang, Donghong Cai and 3 more

Published Oct 3, 2026 · 0 citations · ▲ 9 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

LoHi trades per-frame resolution for denser temporal sampling via low-resolution streams plus sparse high-resolution frames, boosting long-video accuracy up to 10.6 points and cutting front-end latency up to 7x.

Sixun Dong, Wei Li, Andong Deng, Qi Qian and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Oct 3, 2026 · ▲ 13 on Hugging Face · Code ★ 1

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution

ConEx bridges saliency maps with concept reasoning via automatic concept discovery to generate faithful, human-interpretable visual explanations.

Yehonatan Elisha, Oren Barkan, Ziv Weiss Haddad, Noam Koenigstein

Published Oct 3, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

DiffGate gates on-policy teacher guidance by trajectory failure and group difficulty to combine dense token-level updates with outcome-level GRPO rewards, improving student pass rates.

Karn Tiwari, Varnith Chordia, Prathosh A P

Published Oct 3, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Learning Discriminative Geometry for Drifting Models

Drifting models suffer from poor pixel-space performance because representation geometry controls KDE sample weighting; persistent representation learning learns discriminative geometry from raw pixels, cutting FID by 82, 95% without pretrained encoders.

Doudou Zhang, Wenwen Hou, Yilin Chen, Qi Chen

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

DistScene: Object-to-Scene Distillation for 3D Scene Generation

DistScene generates compositional 3D scenes from single images via object-to-scene distillation, improving spatial coherence by modeling environments as explicit components with shared coordinate frames.

Kunming Luo, Hongyu Yan, Ken Deng, Chengcheng Zhou and 5 more

Published Oct 3, 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

PerturbBot breaks vision-language-action shortcut priors via perturbative training while GroundingFscore diagnoses shortcut reliance, enabling healthier scaling without altering inference.

Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang and 3 more

Published Oct 3, 2026 · ▲ 13 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
88%Must read

PaLoRA: Paced Low-Rank Adaptation for Continual Learning

PaLoRA derives an optimal rank-aware pacing law for LoRA continual learning that adaptively restricts gradient scaling to prevent forgetting, improving long-horizon benchmark accuracy by 4%.

Yuxuan Li, Fanhu Zeng, Hao Tang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published Oct 3, 2026 · ▲ 9 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Agentic discovery of blood biomarker from distilled private health records

Distilling private health records into a released scoring tool enables agentic discovery of CBC biomarkers that improve diagnostic AUC over literature baselines without exposing patient data.

Seffi Cohen, Liat Antwarg Friedman, Amir Anisman, Ruth Johnson and 6 more

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
91%Must read

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

CurveCodec 2 predicts quantized skeletal curves from past values and learns residual entropy models to reduce animation storage to 0.22-0.37x of ACL with verified error bounds.

Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 5

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 4/5
88%Must read
?Must readVote to see the score

UnAct: Gradient-Free Unlearning via Targeted Activation Intervention

UnAct uses gradient-free targeted activation interventions to unlearn model classes from few forget images without gradients, labels, or retained data, matching or exceeding SSD and LFSSD accuracy across datasets and preventing network collapse with scarce data.

Saeed Abdul Muizz, Aayat Rafiq, Iqra Altaf Gillani, Janibul Bashir

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

The Numerical Linear Algebra of Large Language Models

This survey explains large language model core concepts to numerical analysts and highlights key numerical linear algebra contributions to LLM techniques.

Abdelkader Baggag, Yousef Saad

Published Oct 3, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

LMBuild evaluates LLM agents on generating buildable, functional 3D structures and finds physical operability and functional affordance remain challenging despite improved soundness.

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang and 8 more

Published Oct 3, 2026 · ▲ 32 on Hugging Face · Code ★ 5

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

ALoDLM: Adaptively Looped Diffusion Language Models

ALoDLM applies token-adaptive latent recurrence to diffusion language models, allocating computation by difficulty to close the quality gap with autoregressive models at 1.7B and 8B scales.

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao and 9 more

Published Oct 3, 2026 · ▲ 62 on Hugging Face · Code ★ 21

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
76%Highly rated
?Highly ratedVote to see the score

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

On-policy distillation improves language models via implicit teacher rewards without expanding capabilities, but unreliable preferences cause reward hacking into repetitive outputs, mitigated by masking bad responses and SFT initialization.

Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 18 on Hugging Face · Code ★ 2

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

WebFovea hardens vision-based web agent harnesses across four stages, raising hidden-set scores from 31.0 to 57.0 by fixing click mismatches, silent actions, and token contamination rather than model reasoning.

Jiangang Han

Published Oct 2, 2026 · 0 citations · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Recursive Self-Rewrite uses diverse harnesses and recursive revision to rewrite successful terminal trajectories for supervised fine-tuning, boosting pass@3 by up to 7.6x on hard benchmarks.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 100 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Periscope: Extending Frozen Language Models Beyond Their Context Window

Periscope arranges text chunks in a grid to build an evidence map via local and strided probes, letting frozen language models answer questions across multi-million-token contexts with sublinear cost and small GPU memory.

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi and 4 more

Published Oct 2, 2026 · ▲ 8 on Hugging Face · Code ★ 1

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
90%Must read
?Must readVote to see the score

World Action Learning via Interaction-Centric Spectral Latent Guidance

WING distills interaction-centric latent actions from egocentric video and uses cross-embodiment spectral low-frequency guidance to transfer them to robot policies, achieving high success rates on LIBERO, RoboTwin, RoboCasa, and real-world tasks.

Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 24 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

Rubric-privileged on-policy distillation before reinforcement learning improves open-ended task scores and reduces reward hacking versus supervised fine-tuning baselines.

Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

DuoMatching improves few-step video generation by jointly matching frame distributions and adding frame-level supervision via an image teacher, boosting visual quality and semantic alignment over 80%.

Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 69 on Hugging Face · Code ★ 33

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Fixed token codes can replace trainable input embeddings in 1.7B-scale language models, removing 100.7M parameters while preserving substantial capabilities without requiring token-specific vectors.

A. Bochkov

Published Oct 2, 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
89%Must read
?Must readVote to see the score

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

EyeRobot 2.0 uses active gaze and fixation-relative frames to enable precise bimanual manipulation with only a single stereo camera, outperforming passive stereo by 40% in real-world trials and doubling ego-plus-wrist success under occlusion.

Kush Hari, Justin H. Kerr, Nidhya Shivakumar, Samarth Mahapatra and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

Skill2Real learns simulation skills via a shared robot API with a Proposer-Verifier-Governor loop, transferring frozen skill hierarchies to real robots without fine-tuning to reach 78.75% real-world manipulation success.

Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 17 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Efficient reasoning training differs in impact: faithfulness usually drops due to inconsistency, but monitorability remains robust.

Samuel Lewis-Lim, Xingwei Tan, Mario Sänger, Zhixue Zhao and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 15 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Collective Bias Mitigation via Model Routing and Collaboration

Collective Bias Mitigation routes queries among diverse LLMs and fosters collaboration to substantially reduce bias over single-model baselines.

Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 18 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Multilingual GSM-Symbolic introduces matched math problems across 15 languages to show model size, resource level, reasoning, and typology determine cross-lingual transfer, with size and reasoning closing low-resource gaps but not typological ones.

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart and 21 more

Published Oct 2, 2026 · 0 citations · ▲ 50 on Hugging Face · Code ★ 4

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ProAR: Learning Prospective Reasoning with Autoregressive Video Models

ProAR introduces goal-frame prediction and future self-alignment to enable goal-directed reasoning in autoregressive video models, surpassing baselines with 25% training steps.

Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen

Published Oct 2, 2026 · 0 citations · ▲ 31 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

LLM agents show strong source preferences across search domains that can override item quality, though supplying missing information or countering preconceptions reduces this bias.

Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 43 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

PointWAM forecasts 3D point trajectories of scenes and hands in a shared space-time frame to guide dexterous robot manipulation, improving DexJoCo success by 56.9 points with video pre-training.

Chunghyun Park, Beomjun Kim, Seungcheol Park, 권희승 and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 55 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

4DCodeBench benchmarks agents reconstructing dynamic scenes from video as executable graphics code, finding strong static models fail at complex dynamics.

Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang and 5 more

Published Oct 2, 2026 · 0 citations · ▲ 28 on Hugging Face · Code ★ 89

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

MetaRubric fixes vacuous rubric credit via evidence-aware optimization and counterfactual rubric adaptation, improving PubMedQA accuracy by up to 20.40 points over static-judge GRPO.

Yuxuan Fan, Jaehong Yoon

Published Oct 2, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot-SD self-distills masked diffusion language models by supervising high-impact commitment tokens via information-gain selection, improving reasoning with minimal data.

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 58 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

World Embedding Benchmark

World Embedding Benchmark evaluates physical encoding via 8,000 simulation cases, finding alignment trades off against quantitative recoverability and retrieval improves video generation fidelity.

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 52 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

FrugalEvo pairs expensive LLM strategy exploration with cheap LLM implementation and caching to maximize optimization gain per cost under a budget, outperforming baselines on 10 tasks at significantly lower expense.

Hui Chen, Xuan Qi, James Zhao, Zhaopeng Feng and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

In-Distribution Forcing for Long Video Generation at Test Time

In-Distribution Forcing prevents out-of-distribution key-value drift via self-caching to extend short video models to minute-scale generation.

Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Language Models that Play Chess and Explain Their Moves

Queen, a 4B-parameter chess-language model, plays at grandmaster level and explains moves via cross-attention to a silent expert encoder and iterative Bellman-style explanation distillation, surpassing larger frontier models.

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Published Oct 2, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 40

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read

Native Action-Prior Learning from Videos for World Action Models

NAVA-WAM pretrains robot action policies directly from observation-only videos via flow-matching and joint attention, improving control accuracy and label efficiency.

Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost and 9 more

Published Oct 2, 2026 · 0 citations · ▲ 86 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation focuses agent distillation on capabilities beyond the fixed harness via action preferences and validity checks, improving long-horizon agent performance without task rewards.

Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

Published Oct 2, 2026 · 0 citations · ▲ 7 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

HyperBrowseComp introduces a multilingual, multimodal web-browsing benchmark of 423 hard questions requiring obscure evidence discovery, and current agents perform poorly against human baselines.

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han and 13 more

Published Oct 2, 2026 · ▲ 58 on Hugging Face

100% Readers1 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 1/5
86%Must read
?Must readVote to see the score

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

ADSD links numerical diagnosis to reusable solver self-improvement, reducing mean solver error by nearly 71x across four challenging numerical domains.

Peter Chen, Wotao Yin

Published Oct 2, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

SEER uses self-evolving event reasoning and retrieval to dynamically optimize forecasting with exogenous events via reflective memory and causal knowledge, outperforming state-of-the-art baselines.

Mingtian Tan, Palash Goyal, Mihir Parmar, Sarkar Snigdha Sarathi Das and 5 more

Published Oct 2, 2026 · ▲ 18 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
83%Must read
?Must readVote to see the score
arXivPrivacy

What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

Split-language-model gradients boost token recovery to 97.38% and document reconstruction to 37.77%, so split traffic requires per-token and per-document leakage reporting.

Georgios Politis, Evangelos Pappas

Published Oct 2, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
88%Must read
?Must readVote to see the score

Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction

SHIFT predicts harness utility via a local LLM to search multi-agent structures per query, achieving ~80% mean accuracy across benchmarks while reducing execution tokens by 32%.

Som Sagar, Shasha Li, Hejie Cui, Ransalu Senanayake and 1 more

Published Oct 2, 2026 · ▲ 10 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
80%Must read

Learning Latent Protein Languages for Autoregressive Generation

Learned latent protein languages improve autoregressive generation scaling, speed, and quality versus amino-acid and coordinate token models.

Mahdi Pourmirzaei, Farzaneh Esmaili, Amir Ziashahabi, Mohammadreza Pourmirzaei and 1 more

Published Oct 2, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
91%Must read

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT scores text-to-image alignment via expected agreement between image and caption answers, boosting negation accuracy to 88% and exceeding fine-tuned evaluators on human correlation benchmarks.

Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Foresight: planning future perception in streaming VLMs without retraining

FORESIGHT uses dual-stream anticipatory planning in frozen streaming VLMs to dynamically configure future perception, improving online benchmarks by up to 18.7 points without retraining.

Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Self-Supervised Scaling of Terminal Environments for Scientific Domains

Software-in-the-loop reconstruction scales terminal-agent training by deriving verified tasks from existing scientific workflows without manual references, improving Terminal-Bench 2 performance to 53.56%.

Zhongzhi Li, Yucheng Shi, Zongxia Li, Junyao Yang and 7 more

Published Oct 2, 2026 · 0 citations · ▲ 10 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

OmniConfess mitigates omni-modal hallucinations by producing token-level confessions of evidential dependence to correct unsupported commitments across text, image, audio, and video.

Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou and 3 more

Published Oct 2, 2026 · 0 citations · ▲ 6 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

COSMI: COmpositional Synthesis of Multi-object Interactions

COSMI synthesizes multi-object interactions by composing local single-object clips, yielding a 222k-sequence dataset and a diffusion model that generalizes to unseen object-interaction pairs with higher contact accuracy.

Daniel Eskandar, Ilya A. Petrov, Gerard Pons‐Moll

Published Oct 2, 2026 · 0 citations · ▲ 7 on Hugging Face · Code ★ 9

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search

StepCAD combines an LLM CAD policy with geometry-guided search to recover executable CAD programs from 3D meshes, achieving up to 87.2% relative IoU gains over baselines.

G.H. Nehme, Faez Ahmed

Published Oct 1, 2026 · 0 citations · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

ViGAR factorizes manipulation into a visual subgoal planner and executor sharing world-model representations, achieving 82% success on compositional tasks and enabling in-context behavior changes without parameter updates.

Shukai Gong, Xuanran Zhai, Yintianrun Zhang, Ruopeng Cui and 15 more

Published Oct 1, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

UniWAM: Unified World-Action Model

UniWAM unifies physical reasoning, world generation, and action prediction to achieve state-of-the-art embodied performance and log-linear co-training scaling.

Wenxuan Song, Jiayi Chen, Jingbo Wang, Shuai Zhou and 12 more

Published Oct 1, 2026 · 0 citations · ▲ 9 on Hugging Face · Code ★ 40

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

VIEScore2 unifies synthetic image evaluation via grid-based joint score and defect localization predictions, outperforming zero-shot VLMs in correlation and spatial accuracy.

Xianda Du, Max Ku, Weiming Ren, Zhi Rui Tam and 4 more

Published Oct 1, 2026 · ▲ 16 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
88%Must read
?Must readVote to see the score

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

DMAD recasts distribution matching as adversarial distillation with discriminator heads to eliminate auxiliary score models, achieving state-of-the-art few-step image, video, and audio-video generation.

Zhengming Yu, 袁俊坤, Haotian Yang, Gordon Guocheng Qian and 7 more

Published Oct 1, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise

POTER uses optimal transport geometry between training and reference distributions to downweight mislabeled or shortcut-aligned samples, achieving state-of-the-art worst-group accuracy with a single training stage.

Sung Ho Jo, Seonghwi Kim, Wonsang Yun, Minwoo Chae

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published Oct 1, 2026 · 0 citations

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero benchmarks long-horizon part-level 3D editing via sequential instructions and target images, finding agentic code-based methods preserve unedited regions better but are slower than non-agentic regeneration.

Ruihan Yu, Yu-Ju Tsai, Muyao Niu, Runyi Li and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 56 on Hugging Face · Code ★ 13

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

Kinematic MeanFlow decouples MeanFlow's time derivative via a kinematic identity to stabilize one-step robotic action generation, cutting latency by up to 74% while outperforming multi-step flow matching.

Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao

Published Oct 1, 2026 · 0 citations · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Cross-lingual MoE router alignment improves multilingual LLM performance by aligning router outputs across languages instead of hidden states.

Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Hierarchical Continuous Diffusion Language Models

HC-DLM couples discrete token generation with a continuous latent trajectory via a unified variational denoising objective, outperforming diffusion baselines on Sudoku, Countdown, and language modeling.

Hui Ren, Zihan Li, Chang Liu, Huidong Liu and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 89 on Hugging Face · Code ★ 57

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

AutoGUIWorld uses image generators to synthesize GUI interaction trajectories without running software, improving OSWorld scores to 40.8% and ScienceBoard success to 32.2%.

Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang and 17 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face · Code ★ 10

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Sharpening Tax in Post-Training

Post-training sharpens base model behaviors at the cost of solution coverage, introducing a quantifiable "Sharpening Tax"; a posterior-tempered group sampler reduces this tax while boosting accuracy.

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov and 6 more

Published Oct 1, 2026 · 0 citations · ▲ 104 on Hugging Face · Code ★ 25

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Scaling and Distilling Text Embeddings for Better Diffusibility

Scaling and distilling text embeddings improves latent diffusion by yielding more connected, diffusible spaces that boost generative performance beyond autoregressive baselines.

Zekai Zhang, Yunjie Tian, Yanjin He, Xiaoyan Zhang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

ROWBench: Do Video Models Render What the Program Specifies?

PROWBench evaluates video models' fidelity to program-specified world events via replayable world records and VLM-based logic-render and interaction metrics.

Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu and 5 more

Published Oct 1, 2026 · 0 citations · ▲ 71 on Hugging Face · Code ★ 32

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

PoS maintains explicit belief states for long-horizon LLM agents, detects belief trapping, and recovers to achieve top results across four benchmarks.

Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 93 on Hugging Face · Code ★ 28

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

InterEvolve evolves reward programs at test time via an LLM agent and numerical optimizer to compose a humanoid controller's skills for novel loco-manipulation tasks. The approach releases latent controller competence through object-aware forward-backward models, generating novel strategies that tra

Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

World Observer jointly generates actor and panoramic observer views to continuously model out-of-view dynamics via shared geometric warping and observer sinks.

Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 85 on Hugging Face · Code ★ 31

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

Biomedical retrieval encoders adapt to typed decision models via SBERT2S1, with retrieval pretraining helping residual heads but not cross-heads, and cross-entropy outperforming RLCD by 2.5, 3.0 points after fixing biased reward normalization.

Pritam Deka

Published Oct 1, 2026 · 0 citations · ▲ 11 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Octrees as an Explicit 3D Language

OctLLM represents 3D geometry via sparse octree tokens and trains separate 3D branches to achieve state-of-the-art multimodal 3D generation and understanding without degrading language ability.

Ran Dan, Si‐Tong Wei, Pengfei Xiong, Wei Zhang and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 15 on Hugging Face · Code ★ 12

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

CANOPY adaptively compresses multimodal evidence via hierarchical region scoring and targeted retrieval, improving QA accuracy while reducing input tokens by up to 27.7%.

Hyojeong Yun, Jueun Kim, Wook-Shin Han

Published Oct 1, 2026 · 0 citations · ▲ 19 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation integrates RL teacher gradients via loss averaging, Adam smoothing, and BF16 rounding, with averaging rules significantly altering math accuracy outcomes.

Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 17 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

LOOM stabilizes looped MoE via bounded residual updates and per-loop routers to scale loops to 9, 12, cutting perplexity from 9.62 to 7.77.

Di He, Pengxiang Li, Da Chang, Qingyan Meng and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 8

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer unifies streaming video perception, memory, and proactive response via shared generation, achieving top results on eight benchmarks with a 4B model.

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu and 20 more

Published Oct 1, 2026 · 0 citations · ▲ 232 on Hugging Face · Code ★ 169

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Spatial Memory Intelligence introduces understanding-driven atomic operations for spatial-memory management in long-video world models, improving sparsity, stability, and spatial consistency.

Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 54 on Hugging Face · Code ★ 20

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD distills multi-teacher LLM specialists via hidden-state and prediction-level on-policy supervision, outperforming token-only and representation-only baselines across math, code, and logic benchmarks.

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 67 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read

World Action Modeling with Progressive Visual Planning

ProWAM predicts sparse visual sub-goals and actions via progressive planning, achieving state-of-the-art long-horizon robotic control and strong zero-shot real-world generalization.

Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 92 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Local Support Learning

Local Support Learning pairs weight adapters with GMM gating to keep updates local, resolving catastrophic forgetting in LLMs up to 7B parameters without prior data.

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Published Oct 1, 2026 · 0 citations · ▲ 27 on Hugging Face · Code ★ 14

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

RealCompanion benchmarks AI companions on longitudinal real-world chats, finding needed past messages are usually recent, memory detectors fail on real messages, and persona reconstruction costs vary 31-fold at equal F1.

Arman Behnam, Sunglyoung Kim, Liangwei Yang

Published Oct 1, 2026 · 0 citations · ▲ 274 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness turns fixed base LLMs into agentic verifiers with workspaces and evidence tools, achieving top selection scores and 6.2, 6.4 point gains over single rollouts on long-horizon tasks.

Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 55 on Hugging Face · Code ★ 54

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems

SciUtopia is a closed-loop LLM simulation framework modeling entire academic ecosystems; it finds resubmission amplifies reviewer burden, cautious exploration balances impact and diversity, and inequality can emerge without cumulative funding advantage.

Yiqiao Jin, Yiyang Wang, Lucheng Fu, Bing He and 7 more

Published Oct 1, 2026 · 0 citations · ▲ 43 on Hugging Face · Code ★ 40

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Selection-based Structured Reasoning replaces open-ended reasoning with selection among reusable candidates, cutting per-turn latency over 90% while matching leading small-model search agents' success rates.

Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 14 on Hugging Face · Code ★ 7

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

AGO is an evidence-first quality gate for RAG that treats judge errors and missing data as explicit outcomes, using stratified beta-binomial gates and mandatory meta-evaluation to reduce unsafe promotion to 22.2%-35.1% versus 29.3%-41.8% for naive gates.

Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero, Giuseppe Santoro and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 17 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

GUI-HARVEST optimizes executable harnesses for frozen GUI agents by aligning visual effects, comparing task runs, and consolidating failure patterns into reusable source edits, improving OSWorld-Verified by up to 12.33 points.

Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

World Editing: Intervening on Executable Worlds at Increasing Depth

World editing intervenes on executable environments at increasing depth via IGMWorld and IGMBench, where top agents achieve 78.2% task success with reliability declining by depth.

Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh and 14 more

Published Oct 1, 2026 · ▲ 18 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 3/5
80%Must read
?Must readVote to see the score

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler automates curriculum learning for agent harness optimization via non-stationary bandits that adapt training scenarios to evolving failure patterns, boosting Pass@1 by 4.4, 7.5 points.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han and 7 more

Published Oct 1, 2026 · ▲ 82 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations

MEA is a multi-agent framework that uses reward-driven optimization to generate faithful natural-language explanations across tabular, text, and vision data, outperforming baselines by up to 34%.

Yuyang Cheng, R. Ravi, Srivarshinee Sridhar, Sriparna Saha and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 6 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

The AI Theorist reveals excitonic structure in $α$-RuCl$_3$

AI Theorist autonomously develops a first-principles model identifying distinct excitonic states with contrasting selection rules in α-RuCl3 optical spectra.

Hongjian Zhou, Xianfan Nie, Sean Wu, Tarun Patel and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 4 on Hugging Face

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies

OpenRUA gives off-the-shelf coding agents only ROS 2 terminal access to act as zero-shot visuomotor policies, achieving 99% on CaP-Bench and 87% on LIBERO-PRO without bespoke harnesses or training, and reveals emergent perception and closed-loop control behaviors.

Zhaoyang Chu, Earl T. Barr, Claire Le Goues, Peter W. O’Hearn and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 5 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

DeskForge generates dense desktop supervision via controllable real-app environments, yielding 1.2M observations that improve GUI grounding and long-horizon computer-use task completion.

A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 6

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Labels Override Definitions in Jev-Style Typed Decision Models

Typed decision models exhibit option-label bias because prompts prepend labels to definitions, letting label semantics override rules; removing labels or altering formatting fixes it.

Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram

Published Oct 1, 2026 · 0 citations · ▲ 5 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

LLM2Jev extracts calibrated Jev-style decisions from LLM token probabilities via training-free inference or tree-factorized fine-tuning, showing strong 4B models already match specialized decision models while fine-tuning mainly helps weaker backbones and specific tasks without degrading generation.

Yinheng Li, Justin Wagle

Published Oct 1, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

SimuVerity benchmarks text-to-executable Simulink generation across engineering domains, finding best agents score only 42.86 and structural similarity poorly predicts engineering performance.

Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo and 8 more

Published Oct 1, 2026 · ▲ 53 on Hugging Face · Code ★ 19

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
80%Highly rated
?Highly ratedVote to see the score
ICML 2026WorkshopOffline RL

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

Decision Titan applies test-time training to offline RL, enabling long-term dependencies 20x beyond context windows and 1.7x length generalization while revealing time embeddings and encoding as critical factors.

Jude Waide, Robert Lieck

Published Oct 1, 2026 · 0 citations

100% Readers1 of 1 upvoted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

66%Highly rated
?Highly ratedVote to see the score

Have an LLM Write Your Anomaly Detector: Autonomous Discovery of Compact, Interpretable Detectors for Time Series

An LLM autonomously searches for short NumPy anomaly detectors that lead the TSB-AD benchmark using spectral features and covariance-aware distances without neural networks or GPUs.

David Berghaus

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

64%Worth a look
?Worth a lookVote to see the score

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

OMAF proposes a one-step flow policy framework for online multi-agent reinforcement learning, achieving up to 3.4x higher returns and 10.5x sample efficiency over baselines.

Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du and 1 more

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing

SmoothOperator modulates per-sample label smoothing via embedding prominence to reduce over-alignment, boosting open-set recognition AUROC by up to 4.7%.

Thiru Thillai Nadarasar Bahavan, Yu Xia, Sachith Seneviratne, Halgamuge Saman

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Structure-agnostic Causal Representation Learning

SaCRL jointly identifies causal structure and learns invariant representations via soft optimization over HSIC-based invariance violations without prior structural knowledge. It guarantees structure identification, invariance satisfaction, and out-of-distribution generalization while achieving state

Arman Behnam, Binghui Wang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

ATPO uses adaptive Tversky reinforcement learning for controllable multi-label video safety detection, raising Jaccard Index to 75.44 on SafeWatch-Bench-Real while enabling steerable precision-recall trade-offs.

Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

GLoC-EHR: Evidence-Cited Clinical Reasoning over Global Context and Local EHR Events

GLoC-EHR uses global and local EHR memories for evidence-cited clinical reasoning, achieving top MIMIC-IV macro AUROC while reducing unsupported citations.

Chaiho Shin, Kwangsoo Kim

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Does Scaling Reinforcement Learning Really Require More Training?

SURGE extracts stronger policies from fixed RL histories via spectral fusion of checkpoints, exceeding native training-curve accuracy without extra training or inference cost.

Bangji Yang, Jiajun Fan, MA Hongba, Ruihan Guo and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

A new benchmark shows LLMs lack structural mathematical understanding with discovery as the key bottleneck, and a primitive-guided self-distillation framework repairs reasoning to boost performance.

Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin and 6 more

Published Oct 1, 2026 · 0 citations · ▲ 10 on Hugging Face · Code ★ 8

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

BDA treats multi-LLM council moves as observations of a per-agent reliability model to yield calibrated answer probabilities and robustly handle persistent adversaries without extra LLM calls.

Ionel Eduard Stan, Paolo Napoletano

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Let the Heads Talk: Beyond Diagonal Graph Attention

Top-A learns edge-conditioned off-diagonal cross-head routes in multi-head attention that preserve diagonal paths, improving interaction-dependent tasks without benefiting heterophily.

Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

AURAL uses adaptive latent reasoning with joint chunk prediction to match chain-of-thought performance while cutting first-token latency 11.8x versus explicit reasoning.

Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu and 11 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

Pretrained ASR backbones enable training-free wake-word detection, and PCA-based pruning retains performance at 50% encoder reduction.

Hwayeon Kim, Youngwon Choi, Hyeonyu Kim

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices

FedSAP uses budget-constrained tri-state channel allocation to partition models into global, private, and dropped channels for heterogeneous federated domain generalization, improving accuracy by up to 4.92 points under 80% pruning.

Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Post-training multi-token prediction heads on ~2.5B chain-of-thought tokens match pretraining speedups with 10^3-10^4x less data, while chain-aware verification and adaptive head selection boost throughput up to 16%.

Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne and 7 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

Human audit of 165 WebArena-Lite tasks recovers 5.45, 8.49% evaluator-missed successes, reveals trajectory errors like looping, and shows guide text and MASM improve results.

Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

Sentence-specificity predictors disagree across technical corpora, and ranking LLM revisions improves selection only for certain models and predictors.

Rocker D’Antonio, Thomas Benton Townsend, Dimitrios Michael Manias

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations

This survey unifies weight, prompt, and embedding adaptation methods for large language models into one taxonomy, analyzing trade-offs and cross-paradigm relationships.

Jungwon Park, Changin Choi, Jimyeong Kim, Nojun Kwak and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Ego2Act evaluates egocentric video generation on multi-step goal-directed manipulation, showing models skip steps and fail at fine-grained physical dynamics.

Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 33 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Cog-VADU reformulates video anomaly detection as sequential cognitive reasoning via recurrent chain-of-thought prompting and cross-modal re-ranking for training-free zero-shot performance.

Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

LLM novelty judges are unstable: small prompt changes alter verdicts on over half of identical idea pairs and shift accuracy by over 50 points, undermining automated ideation evaluations.

Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

FAULT turns self-diagnosed errors into step-level credit via terminal outcome anchoring and evidence-checked cost learning, recovering 95% signal coverage on ALFWorld and improving long-horizon agentic RL.

Yihua Zhu, Qianying Liu, Weixu Qiao, Xuan Ren and 9 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

FedFit reduces federated LLM fine-tuning overhead via vector-bank adapter parameterization and quantization, resolving LoRA aggregation conflicts to achieve up to 100x compression with comparable perplexity.

Hang Zou, Chao Zhang, Yuzhi Yang, Yu Tian and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Autoregressive Drillhole Modelling Under Distribution Shift

DrillBench benchmarks autoregressive drillhole modeling, finding lithology-sequence models transfer more robustly than spatial methods, and combining pretraining with retrieval improves cross-province generalization.

Yihao Ding, Daniel Yitian Su, Yiran Zhang, Christopher M. Gonzalez and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation

ReSolve reuses candidate reasoning via selective generative moderation to boost math accuracy and cut token use versus voting and self-consistency.

Bangji Yang, Jiajun Fan, MA Hongba, Xi Zhu and 5 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints

Importance-guided allocation directs limited $\ell_1$ perturbation budgets toward model-sensitive regions via fixed clean-gradient priors, boosting attack success by 2.52, 17.70 points across ten robust configurations without increasing global consumption.

Melika Shirian, Kianoosh Vadaei

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

SALD: Self-Referenced Advantage Learning for Diffusion Models

SALD self-references diffusion training via dual noise-level error differences and spectral residuals to improve generation without teachers or extra parameters.

Aryan Das, Surjo Dey, Koushik Biswas, Swalpa Kumar Roy and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Iterative Policy Refinement through Semantic Rollout Analysis

A closed-loop framework iteratively refines structured imitation-learning policies via LLM analysis of rollout tables, improving performance by up to 15% and cutting compute 75%.

Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

FSG-RL connects subproblem graphs with Python code and multi-verifier feedback to improve math reasoning, raising final-answer accuracy from 43.25% to 67.50% over supervised fine-tuning.

Zihan Liu, Xurong Xie

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

QK-Wanda couples query and key pruning scores via cross-projection deletion costs, reducing QK reconstruction error by 60% at 50% sparsity and improving downstream perplexity on some large models.

Ivan Ilin, Peter Richtárik

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

BanglaDial-Abuse introduces 1,000 synthetic abusive Bangla sentences across four regional dialects for four-class dialect identification, achieving 0.37, 0.56 lexical Jaccard similarity with distinct lexical spaces.

Hasin Almas Sifat

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score
arXivDeep RL

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

The Homomorphic Advantage Operator stabilizes FHE-based reinforcement learning by centering TD targets to eliminate Bellman drift, achieving zero approximation-bound breaches and 18-point accuracy gains without extra multiplicative depth.

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Capturing In-Context Learning Dynamics with Task Operators

Task Operator captures ICL as stable per-task affine attention transformations, enabling efficient zero-shot replay that nearly matches in-context performance and scales beyond context limits.

Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Yo-ByT5 is a byte-level Yorùbá diacritic restoration model matching mT5-base accuracy with half the parameters and superior text fidelity.

Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

MCIR-M quantifies unique predictive information beyond dependent neighbors via a normalized conditional ratio in [0,1], yielding stable dependence-aware global rankings under multicollinearity and near-duplicates.

Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers

Routing entropy in attention-residual transformers fails as a reliability signal beyond model confidence, with no robust calibration gains and low sensitivity to injected effects.

Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

SCOPE-AD: Sequential cost-aware ordinal-belief planning with energy-based models for diagnostic agents

SCOPE-AD sequentially selects diagnostic tests via ordinal-belief planning and energy-based policies, achieving 77.70% ADNI macro-F1 at $50.46 average cost versus far pricier full-modality evaluation.

Ziwen Yu, Ivan Koychev, Elizabeth Coulthard, Ting Zhou and 6 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Lingtai: What Concept Geometry Reveals—and Does Not Reveal—About LLM Inference

Lingtai introduces a training-free concept telemetry layer that reveals inference-time uncertainty-linked activity and execution-specific trajectory structures in LLMs without tracking correctness.

Jiangang Chen

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

iADD analyzes diffusion policy optimization to show only-latter-timestep updates harm diversity, then proposes incremental Feynman-Kac training that improves alignment-diversity tradeoffs across tasks.

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

From Knowledge Access to Source Learning: Developing Source-Specific Competence

SourceLearn develops reusable source-specific competence via persistent source models and dual learning mechanisms, outperforming retrieval and memory baselines by up to 22.6 points.

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 9 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Learning to Predict Distributions over Weight Updates for Test-Time Adaptation

Query-conditioned hypernetworks predict distributions over LoRA weight updates from input queries, enabling test-time scaling via sampled adapted models that outperform deterministic and token-sampling baselines.

Azal Ahmad Khan, Keshav Ramji, Tahira Naseem, Ali Anwar and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

CARM prevents opposing token-level probability changes from canceling in sequence-level masking by using absolute log-ratios, improving RL reasoning and code benchmarks over geometric-mean masking.

Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

Clock Diffusion introduces semi-autoregressive continuous diffusion language models with position-dependent noise schedules, efficient training and sampling, and Cache Grab acceleration to achieve state-of-the-art diffusion likelihoods and competitive reasoning performance.

Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky and 6 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Code Owns the Simulation, Jev Owns the Evaluation

Judgment models excel at evaluation but fail at simulation, yet pairing them with code simulation yields expert control.

Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems

Open-source LLM multi-agent systems face orchestration and execution issues mostly caused by workflow, tool integration, and memory problems, primarily solved by workflow optimization.

Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li, Peng Liang and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

XAI Evaluation Cards: A Practical Method for Designing Human-Centred XAI Evaluations

XAI Evaluation Cards provide a card-sorting method to systematically design human-centered evaluations of explainable AI systems across disciplines.

Kristýna Sirka Kacafírková, Ivania Donoso-Guzmán, Denis Parra, Katrien Verbert and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

A typed settlement contract audits concurrent LLM agent actions, showing joint policies complete 59% more six-agent doorway tasks than conservative rejection, with exact replay of 156 checkpoints and rejection of 1,332 corruptions.

Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Calibration-risk routing for controlled world-model adaptation

MC-WM partitions target data to select lower-calibration-risk world models and weights imagined policy updates via learned confidence, evaluated across 541 MuJoCo shift executions.

Yifan F. Zhang, Liang Zheng

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Semifactual Credit-Augmented Policy Optimization

SCAPO improves RL reasoning by assigning token-level credit via semifactual stability, boosting AIME accuracy over GRPO by up to 5.63 points.

Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 21 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation

ReSAIL selects PI-sensitive interaction steps and regularizes student outputs to prevent collapse in iterative agent self-distillation, boosting final-cycle success by 22.5%.

Shengjie Jin, Hengbo Xu, Zelong Sun, YuJie Guo and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 3 on Hugging Face · Code ★ 7

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers

A benchmark of 390 AI papers measures scientific slop across structure, argument, and artifacts; a harness reduces the AI-human gap by 63% via evidence-grounded revision.

Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 13

0% Readers0 of 1 upvoted
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency

DeCoPrune uses denoising-consistency to prune over 85% of KV-cache tokens in autoregressive video diffusion, preserving long-range recall and accelerating continuation generation by over 4x.

Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou and 2 more

Published Sep 30, 2026 · 0 citations · ▲ 1 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Learning Functional Subspaces for Neural Network Compression

LSP learns low-rank subspaces end-to-end via joint orthogonal projector optimization to reduce transformer memory and compute while outperforming local criteria at high compression ratios.

Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

DARA corrects batch-level reward imbalance via inverse-square-root active-group density weights, accelerating multi-reward RL training by up to 65% with no objective change.

Tong Zheng, Skylar Zhai, Zhan Cheng, Tianming Sha and 6 more

Published Sep 30, 2026 · 0 citations · ▲ 64 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

RSIGame uses recursive self-improvement with local and global loops to autonomously refine generated games, surpassing one-shot GPT-5.5 scores while cutting generation tokens by 11x.

Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou and 9 more

Published Sep 30, 2026 · 0 citations · ▲ 91 on Hugging Face · Code ★ 126

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

EvoDuet co-evolves solutions and web queries via a retrieval gate to boost LLM discovery gains up to 82.3% across optimization tasks.

Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 109 on Hugging Face · Code ★ 5

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

OSWorld-Science benchmarks VLM agents on 146 expert scientific software tasks, showing state-of-the-art models still struggle with scientific workflows and harness design.

Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li and 27 more

Published Sep 30, 2026 · 0 citations · ▲ 63 on Hugging Face · Code ★ 5

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision JEV models show ordinal scale-utilization bias, compressing decisions to 26, 76% of gold support despite high accuracy, but BA-LoRA post-training improves utilization to 86%.

Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 60 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Rhetorical robustness requires stable judgments across content-preserving rewrites and discrimination across papers; SciCore improves both via dual-branch science-core review.

Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 76 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Frontier models follow unreliable external guidance; training improves selective reliance, identifying it as a key agent reliability dimension.

Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 68 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

EgoTools introduces a dataset and benchmark for egocentric tool-use reasoning, showing current models struggle with visual grounding while training improves performance.

Shulin Tian, Junsu Kim, Shuai Liu, Hao Li and 16 more

Published Sep 30, 2026 · 0 citations · ▲ 67 on Hugging Face · Code ★ 9

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid-Harness verifies candidate terminal actions at the model-harness boundary, raising TerminalBench-Lite Pass@1 from 50.00% to 68.03% and improving success at lower token cost than trajectory scaling alone.

Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan and 7 more

Published Sep 30, 2026 · 0 citations · ▲ 117 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

WorldAuditBench benchmarks interactive 3D world auditing with multimodal agents, finding success rates of 6.6% to 42.3% versus 83.4% human performance.

Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 102 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Self-evolving search agents suffer co-cheating where proposers and solvers mutually reinforce errors; CrossFit partitions sources to cross-fit agreement and cuts false agreement by over half, boosting downstream search by 8+ points.

Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin and 11 more

Published Sep 30, 2026 · 0 citations · ▲ 672 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

UniEvo-VL improves multimodal image generation via self-distillation that minimizes divergence between student and critique-conditioned teacher diffusion distributions during self-correction. Experiments on Qwen2.5-Image improve GenEval scores from 0.747 to 0.808 without external teachers.

Fang Wu, Da Xing, Yanjie Huang, Junxi Wang and 15 more

Published Sep 30, 2026 · 0 citations · ▲ 292 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems

Benign multi-agent LLM planners disguise secrets to help developers evade oversight, with rare per-episode leaks compounding to high breach risk across repeated exchanges.

Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng

Published Sep 30, 2026 · 0 citations · ▲ 17 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Diptych: Scoped, AI-Interpreted Comparison for Reference Listening in Music Production

Diptych lets musicians define comparison scopes for reference listening, helping surface differences experts partially support while avoiding overreaching AI judgments.

Chongjun Zhong, Abhinaba Roy, Archishman Ghosh, Kejun Zhang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 21 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

GraphForge synthesizes workspace tasks and verifiers over real file evidence graphs to train working agents, and fine-tuning Qwen3.6-27B improves GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II results.

Qisheng Su, Hanchen Wang, 朱冠儒, Huicheng Jiang and 8 more

Published Sep 30, 2026 · 0 citations · ▲ 146 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

FrameMorrow guides historical frame selection via prospective tokens representing future needs, improving consistency and quality across diverse long-horizon video generators.

Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan

Published Sep 30, 2026 · 0 citations · ▲ 105 on Hugging Face · Code ★ 29

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

LexReward introduces taxonomy-driven rubric-based rewards for legal language models across style, element, and reasoning dimensions, improving DPO and reinforcement learning performance.

Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie and 4 more

Published Sep 30, 2026 · 0 citations · ▲ 62 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

OmniReasoning introduces a benchmark, data engine, and self-distillation method to improve audio-visual joint reasoning, boosting Qwen3-Omni-30B-A3B-Thinking by up to 12.8 points.

Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu and 10 more

Published Sep 30, 2026 · 0 citations · ▲ 23 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

FailBank turns runtime shield feedback into persistent VLA policy updates via failure-bank self-evolution, raising success rates up to 25.4 points and cutting policy-induced cost up to 35.6%.

Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 28 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

Does Learning Protein Folding Generalize to Broader Reasoning?

Post-training on protein-folding data via discrete answers and continuous geometry improves structure prediction and broad reasoning across ten benchmarks.

Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 129 on Hugging Face · Code ★ 30

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Magic-W0 is a structured world-action foundation model that couples 3D geometry, motion, and future semantics with action generation via layer-aligned interaction, achieving top simulated scores and strong real-robot adaptation.

Xuhua Chen, 尹貞漢, Yuan Zhang, Lingfeng Zhang and 14 more

Published Sep 30, 2026 · 0 citations · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

We introduce CIA to measure chain-of-thought alignment with internal reasoning, finding low alignment that post-training improves substantially while maintaining accuracy.

Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi

Published Sep 30, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Generator attribution falls sharply after rewriting, and provenance-based selection does not clearly outperform quality-score selection for recursive model training.

Joss Armstrong

Published Sep 30, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Two-step non-uniform flow denoising reduces VLA inference time from 61.6 ms to 22 ms by exploiting early-stage velocity stability. A distributed real-time framework and garment-folding evaluation show joint model-system optimization preserves task success with lower latency.

Di Wu, Rongtian Shen, Ping Liu, Yan Shen and 7 more

Published Sep 30, 2026 · 0 citations · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Video Generation Models: A Survey of Post-Training and Alignment

This survey reviews post-training and alignment strategies for video generation, framing them as implicit or explicit alignment across four methodological categories to improve controllability and reliability.

Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im and 9 more

Published Sep 30, 2026 · ▲ 60 on Hugging Face · Code ★ 208

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
89%Must read
?Must readVote to see the score

QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code

QuantCode specializes LLMs for executable trading code via framework pretraining and validated fine-tuning, boosting backtest success to 83.5% while revealing specialization trade-offs in tool use and repair.

Alexey Chernysh, Orkhan Ekhtibarov, Dmitry Zmitrovich

Published Sep 30, 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

Beyond the Current Scene: Event-Referential Grasping with Active View Selection

BeyondSCe enables zero-shot event-referential grasping via active view selection, achieving 76% and 77% success on visible and occluded targets versus 40% and 55% baselines.

Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 51 on Hugging Face · Code ★ 8

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Dream4ACT introduces action views to unify cross-embodiment joint actions as shared visual representations, enabling joint video-action modeling and 88.98% RoboTwin 2.0 success with training-free multiview recovery.

Xiangyu Zhu, Jin Xu, Yue Guo, Xin Wu and 5 more

Published Sep 30, 2026 · 0 citations · ▲ 8 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales

GFD-OPD fixes diffusion on-policy distillation by reducing student-teacher gaps and preventing classifier-free guidance error amplification, achieving state-of-the-art compression results.

Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang and 5 more

Published Sep 30, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

STEPQuant spatially and temporally quantizes delta-rule recurrent states to 6-bit with near-FP32 accuracy, cutting serving memory by up to 68.7%.

Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu and 5 more

Published Sep 29, 2026 · 0 citations · ▲ 21 on Hugging Face · Code ★ 83

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

RIDE extrapolates RL-induced representation residuals for stable on-policy distillation, surpassing output-space methods and matching or exceeding RL teachers.

Hao Li, Meijia Chen, Weijie Ren, Donghan Li and 3 more

Published Sep 29, 2026 · 0 citations · ▲ 577 on Hugging Face · Code ★ 6

– ReadersNo votes yet. 1 from authors or colleagues not counted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

MotorMind connects general vision-language models to deterministic robot control via mid-level actions and feedback loops, achieving 66.7% zero-shot success on LIBERO-PRO and 95% on real robots without task-specific training or external tools.

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 109 on Hugging Face

0% Readers0 of 1 upvoted
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing

Self-compensating VLA adapts online to robot execution errors via residual feedback, improving success over 30 points on physical arms and outperforming training-time robustness methods on RoboStress.

Sohyun Lee, Yoonjae Baek, Jaesang Won, Jinnyeong Kim and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 26 on Hugging Face

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts

Predictive credit for scientific explanations is measured via paired forecasts, but gains over descriptions remain unconfirmed across Tox21, OpenML, and controlled settings.

Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li

Published Sep 29, 2026 · 0 citations · ▲ 101 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

EVOKE improves LLM agent transfer by ranking actions under diverse goals at fixed states to elicit pretrained world knowledge for robust decision-making.

Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li and 7 more

Published Sep 29, 2026 · 0 citations · ▲ 78 on Hugging Face · Code ★ 7

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Pretrained transformers stop following references after 1.4, 3.6 lines, but a rank-8 LoRA at one early layer extends computation to 50, 160 lines without changing frozen weights.

Zehao Jin, Ruixuan Deng, 君然 王

Published Sep 29, 2026 · 0 citations · ▲ 75 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

A Builder learns reusable meta-skills from target feedback to construct execution harnesses that boost target performance on unseen tasks. Meta-skills improve macro-average scores by 8.95 points over no-skill construction and 12.02 over direct delivery.

Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 83 on Hugging Face · Code ★ 9

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

E-MoE improves few-step non-factorized diffusion language models via Mixture-of-Experts routing as a discrete shared latent, boosting sample quality without extra active parameters.

Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 66 on Hugging Face · Code

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

LoopVL: Recurrent Visual Intelligence

LoopVL applies recurrent loop transformers to vision-language models via iterative shared-module updates, outperforming larger non-recurrent models and exhibiting visual aha moments.

Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li and 8 more

Published Sep 29, 2026 · 0 citations · ▲ 470 on Hugging Face · Code ★ 172

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

AREX-2 synthesizes long-horizon reflective trajectories to train a Qwen3.8-27B agent that self-improves at test time, achieving strong results on MLE-bench, Frontier-CS, and deep research benchmarks while scaling with iteration budget.

Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei and 10 more

Published Sep 29, 2026 · 0 citations · ▲ 139 on Hugging Face · Code ★ 33

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

Rollout-Marginal Distillation scores autoregressive video chunks independently against a chunk teacher to prevent error accumulation, then applies video-level distillation to restore temporal coherence.

Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 5

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Tail-Influence Sampling for CVaR Policy Evaluation

Tail-Influence Sampling allocates evaluation budgets by tail influence to estimate CVaR with oracle variance and lower MSE than rollouts.

Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 27 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

Embodied progress reward models fail at long tasks due to missing context, but ProgressCompass supplies needed context to cut progress estimation errors by up to 82%.

Jianshu Zhang, Keyi Wu, Chengxuan Qian, Xiyuan Yang and 5 more

Published Sep 29, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Video2Skill: From Streaming Experience to Reusable Embodied Skills

Video2Skill benchmarks streaming embodied skill discovery, showing VLMs group manipulation events poorly and rarely expand skill libraries despite supervised fine-tuning.

Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu and 5 more

Published Sep 29, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

Adaptive Reward Routing dynamically routes updates and balances rewards during forward-process RL for joint audio-video diffusion, consistently improving quality, alignment, and synchronization over fixed baselines.

Songlin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 138 on Hugging Face

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

Proactive LLM agents need joint optimization of task capability, temporal compute allocation, and user trust, with Proactivity-Gym exposing evaluation gaps and human preference for unobtrusive assistance.

Jio Oh, Seunghyun Do, Youngjun Lee, Steven Euijong Whang and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 28 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

HelixWorld: A Real-time Interactive Audio-Visual World Model

HelixWorld is a real-time interactive audio-visual world model that synchronizes visual scenes and spatial stereo sound under user control at 24 FPS, surpassing silent models in acoustic immersion.

Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang and 12 more

Published Sep 29, 2026 · 0 citations · ▲ 34 on Hugging Face · Code ★ 623

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

WEFT evolves whole agentic interaction systems for tool-use post-training, outperforming environment-scaling baselines by up to 12.27 points across benchmarks.

Bo Mao, Hang He, Linting Wang, Lizhi Lin and 16 more

Published Sep 29, 2026 · 0 citations · ▲ 24 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

On-policy methods continuously adjust parameter update directions, unlike consistent SFT updates; constraining SFT to these directions via OPSFT transfers on-policy generalization advantages to supervised fine-tuning.

Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 86 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

Triadic linear attention uses 3D tensor states via triadic outer products to scale recurrent state size efficiently, substantially improving long-context modeling and recall.

Oliver Sieberling, Bharat Runwal, David Jin, Ryan Chin and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 10

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation

RASO retrieves and adapts external skills via cross-harness adaptation to initialize and iteratively update agent skills, outperforming non-retrieval baselines across benchmarks.

Jaewon Chu, Ji Soo Lee, Jihwan Park, Dohwan Ko and 7 more

Published Sep 29, 2026 · ▲ 63 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
78%Highly rated

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Imagine3D-LLM learns compact scene representations via learnable summary tokens decoded into 3D Gaussian splatting with photometric loss, improving multi-view spatial reasoning benchmarks.

Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han and 9 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published Sep 29, 2026 · ▲ 72 on Hugging Face · Code ★ 37

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
80%Must read
?Must readVote to see the score

AIM: Agentic Idea Management for Automated Research

AIM autonomously manages research ideas via Bayesian-inspired selection and auditing to outperform baselines by up to 4.9 points and speed up search up to 3.1x.

Hyeong Kyu Choi, Bhavana Dalvi Mishra, Jiefeng Chen, Mihir Parmar and 6 more

Published Sep 29, 2026 · 0 citations · ▲ 51 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

PixelUMM is an encoder-free unified model for image and video understanding and generation that represents images as spatial patches and videos as spatiotemporal tubelets, achieving competitive performance across tasks.

Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren and 7 more

Published Sep 29, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 161

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Context Language Models

Context language models treat context as self-modified files to learn context management, outperforming external strategies with lower compute and enabling in-context and parametric learning of management strategies.

Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li and 9 more

Published Sep 29, 2026 · 0 citations · ▲ 43 on Hugging Face · Code ★ 595

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation

Interpolated Policy Distillation mixes student and teacher token distributions to balance trajectory quality and learnability, outperforming off-policy and on-policy distillation across reasoning benchmarks.

Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang and 4 more

Published Sep 29, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Multilinguality in Hybrid Attention LLMs

Hybrid attention LLMs develop cross-lingual alignment tied to recurrent and full-attention layer ordering, with a spike at the first full-attention layer; distillation shows starting with full attention learns up to 2.5× faster.

Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz and 1 more

Published Sep 28, 2026 · 0 citations · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

PersonaDose calibrates activation-steering controllers to control language-model persona traits by requested intensity, reducing targeting errors to 4.7-6.2 points across models.

Zehao Jin, Junran Wang, Ruixuan Deng, Jiahao Chen and 3 more

Published Sep 28, 2026 · 0 citations · ▲ 54 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Agent Priors-guided Policy Learning

Agent Priors-guided Policy Learning embeds structural priors in skill interfaces to enable compositional and out-of-distribution skill generalization.

Puming (Oscar) Jiang, Tao Hu, Haozhe Du, Yibo Li and 3 more

Published Sep 28, 2026 · 0 citations · ▲ 85 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

WM-VLM adds a world-model branch to vision-language models to generate intermediate visual states for spatial reasoning, outperforming baselines by up to 39.25 points on mental rotation tasks.

Yuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen and 4 more

Published Sep 28, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

In controlled strong-to-weak distillation, rollout policy is less central than token-level KL direction and learning rate, though on-policy data can improve generalization on harder reasoning tasks.

Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar

Published Sep 28, 2026 · 0 citations · ▲ 197 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

Raw-turn selection via a typed decision model matches LLM extraction at tight budgets but falls behind at generous budgets, explaining conflicting memory results.

Rishabh Sharma, Rishika Lall

Published Sep 28, 2026 · 0 citations · ▲ 20 on Hugging Face · Code ★ 1

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

ReaLVR fixes latent reasoning's weak visual grounding by supervising latent tokens with visual evidence, improving reasoning across scales up to 235B.

Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li and 9 more

Published Sep 28, 2026 · 0 citations · ▲ 217 on Hugging Face · Code ★ 46

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

LVMT: Video Mask Transformer for Long-term Video Segmentation

LVMT introduces a GRU-based propagation module and truncated query propagation training to achieve state-of-the-art long-term video segmentation at 10x faster speeds.

Narges Norouzi, Niccolò Cavagnero, Idil Esen Zulfikar, Bastian Leibe and 2 more

Published Sep 28, 2026 · 0 citations · ▲ 18 on Hugging Face · Code ★ 9

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Data Unlearning via Inverse Distillation

Inverse Distillation Unlearning unifies multi-step model distillation and data removal via a min-max objective that recovers only retained data, reducing forgotten-class generation without retained examples or extra classifiers.

Aleksei Leonov, Nikita Kornilov, Zhenhe Zhang, Evgeny Burnaev and 2 more

Published Sep 28, 2026 · 0 citations · ▲ 20 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs

PDE-JEPA introduces predictive masked-latent pretraining with geometry projection and structured latent predictors for parametric PDE dynamics, reducing errors by 33.4% in-distribution and 51.4% on unseen parameters.

Zhentao Tan, Jianrong Zhang, Ruijie Quan, Yi Yang

Published Sep 28, 2026 · 0 citations · ▲ 58 on Hugging Face · Code ★ 9

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

SlimWise decouples MoE expert pruning across prefill and decode phases to boost serving throughput without sacrificing accuracy via direct KV cache reuse and selective distillation.

Gunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin and 2 more

Published Sep 28, 2026 · 0 citations · ▲ 12 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

MM-ABC is a mobile manipulation foundation model combining multi-level vision features, future imagination supervision, and masked joint attention to coordinate arm-base actions, achieving up to 99.1% success across benchmarks and 83% in real-world tasks.

Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao and 5 more

Published Sep 28, 2026 · 0 citations · ▲ 6 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

How to Loop MoE: Flatten the Experts, Untie the Attention

Foil improves looped MoE by flattening experts and untying attention, reducing pretraining loss by 0.012 nat and improving routing balance and confidence.

Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly and 5 more

Published Sep 28, 2026 · 0 citations · ▲ 9 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Learning Multimodal Embeddings with Evidence-Aligned Readout

EviAlign organizes multimodal evidence into semantic units and reads states at their boundaries to build retrieval embeddings, with co-designed semantic organization and boundary readout yielding 76.9 Recall@1 across 12 tasks.

Zirong Chen, Fuda Ye, Enjun Du, Junfu Pu and 7 more

Published Sep 27, 2026 · 0 citations · ▲ 6 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

SWE-Game: Can Coding Agents Build the Games We Want?

SWE-Game benchmarks coding agents on 247 Godot game tasks, finding best scores remain below 60 and executable checks outperform video judges in evaluation.

Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou and 7 more

Published Sep 27, 2026 · 0 citations · ▲ 14 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception

MinkowskiPE applies Lorentz transformations via joint spacetime positional encoding to make attention depend only on relative displacement, improving molecular dynamics and video prediction with far fewer parameters.

Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao and 3 more

Published Sep 27, 2026 · 0 citations · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features

Existing token-reweighting methods cannot reverse harmful SFT features; SCALE uses frozen SFT deltas with entropy-guided gates to suppress, reverse, or extrapolate them, improving math and code results.

Cunchun Li, Haonan He, Yifan Gao, Minglei Li and 3 more

Published Sep 27, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

YuE2 unifies symbolic and audio music generation through symbolic planning, producing readable scores and full-song audio that outperform public baselines and rival proprietary generators.

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu and 31 more

Published Sep 27, 2026 · 0 citations · ▲ 246 on Hugging Face · Code ★ 10,927

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

VisionHOPE formulates visual backbones as self-modifying learning systems with coupled co-evolving memories and proves stable non-expansive dynamics, achieving competitive results on ImageNet-1K, COCO, and ADE20K.

Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao and 7 more

Published Sep 27, 2026 · 0 citations · ▲ 323 on Hugging Face · Code ★ 880

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Raven: The Harness of Harnesses for Composable Agentic Intelligence

Raven autonomously constructs modular agent harnesses and orchestrates them across domains via a multi-agent ecosystem, significantly outperforming state-of-the-art systems on complex long-horizon tasks.

EverMind AI

Published Sep 27, 2026 · 0 citations · ▲ 566 on Hugging Face · Code ★ 5,252

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

X-Tree learns reusable hierarchical skills from agent trajectories and improves success rates up to 5.8% across web and science benchmarks.

Sitao Cheng, Xunjian Yin, Zhiyuan Sun, Yuxuan Li and 3 more

Published Sep 26, 2026 · 0 citations · ▲ 76 on Hugging Face · Code ★ 2

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

HeteroFold enables prefill-free cross-family KV cache transfer between frozen heterogeneous LLM agents, accelerating 32K context transfer up to 10.7x while matching text-based multi-agent performance.

Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo and 2 more

Published Sep 26, 2026 · 0 citations · ▲ 95 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS

DEFINE decouples speaker identity and accent in zero-shot TTS via separate audio exemplars and a single guidance weight, generalizing accent control beyond training accents with high speaker similarity.

Ambuj Mehrish, Abhinaba Roy, Alex Ivanov, T. Ahmed and 1 more

Published Sep 26, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Adaptive Latent Capacity for World Models

ALeWM learns adaptive-width JEPA world models via MixSIGReg and prefix sampling, concentrating predictive information in early latent coordinates to improve planning with lower capacity.

Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer and 1 more

Published Sep 26, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

GeoCR: Learning a Generalist Cloud Removal Prior from Heterogeneous Observations

GeoCR learns a generalist cloud removal prior from heterogeneous observations, enabling direct inference and LoRA adaptation across diverse sensors without dataset-specific fine-tuning.

Jeonghyeok Do, Munchurl Kim

Published Sep 26, 2026 · 0 citations · ▲ 5 on Hugging Face · Code ★ 2

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

KeyRec creates bounded visual memory via recent caches and structured event banks to enable efficient long-video and streaming understanding with only 10% of visual tokens, outperforming compressed baselines.

Zihan Chen, Xuejian Rong, Xiaojuan Wang, Boqing Gong and 3 more

Published Sep 26, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation

GeoSET is a generalist SAR-to-EO translation model pretrained on 3 million diverse pairs and adapted via LoRA, achieving state-of-the-art results across six benchmarks.

Jeonghyeok Do, Munchurl Kim

Published Sep 26, 2026 · 0 citations · ▲ 6 on Hugging Face · Code ★ 3

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins

PluginRSI evolves agent harnesses via reusable plugins, improving over existing methods and accelerating optimization on unseen tasks.

Yaorui Shi, Yuchun Miao, Yuxin Chen, Jiayuan Zhang and 4 more

Published Sep 26, 2026 · 0 citations · ▲ 7 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward Modeling

ScopeIF improves LLM instruction-following via graded reward modeling and scope-aware constraints, enabling small models to match frontier performance.

Bosi Wen, Yilin Niu, Xiaoying Ning, Ying Zhang and 2 more

Published Sep 26, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis

CEG-Agent introduces causal error graphs and a taxonomy separating anomalies, errors, and failures to diagnose agentic traces, achieving state-of-the-art results on the CEG-Bench benchmark.

Shu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen and 6 more

Published Sep 26, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding

DMM refines multi-agent action intents iteratively via communication to fix incompatible joint actions, achieving near-perfect success on 1,600 MovingAI tasks and scaling to over one million agents.

Valeriy Vyaltsev, Anton Andreychuk, Taisia Zlotnikova, Konstantin Yakovlev and 2 more

Published Sep 25, 2026 · 0 citations · ▲ 62 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Receiver-Conditioned Latent Communication gives 94% CacheBack

CacheBack uses receiver-conditioned filtering of sender KV caches via attention weights to cut transferred state by 75%, boosting multi-agent accuracy by 14.7 points and reducing latency 3.2x versus text.

Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang and 1 more

Published Sep 25, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

TimeBraid: Unifying Time Series and Language for Understanding and Forecasting

TimeBraid unifies time-series and language models via interleaved residual attention to jointly enable understanding and forecasting, matching larger specialized models.

Xinyue Wang, Jiacheng Pang, Kun Zhou, Kexin Zhang and 6 more

Published Sep 24, 2026 · 0 citations · Code ★ 3

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification

Automatic taxonomies reveal retrieval and verifier reasoning failures persist across scaled medical retrieve-then-verify systems, showing fundamental open-ended evaluation limits.

Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi and 4 more

Published Sep 24, 2026 · 0 citations · ▲ 19 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score
ICASSP 2027Speech synthesis

Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning

Accent analogy guidance subtracts estimated accent directions for cross-lingual voice cloning, raising speaker similarity above identity-accent trade-off curves across several open TTS models.

Yoomee Cho, Jisun Lee

Published Sep 24, 2026 · 0 citations · ▲ 5 on Hugging Face · Code ★ 5

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score
Proceedings of the Human Factors and Ergonomics Society Annual Meeting 2026Pennsylvania StateZhejiangHuman-AI interaction

When Automated Vehicles Cannot Retaliate: Driver-Initiated Takeover Under Aggression in Mixed Traffic

Simulator study finds drivers initiate takeovers under sharp braking and high aggression because automated vehicles cannot retaliate to human drivers.

Haitao Chen, Yiqi Zhang

Published Sep 24, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control

NowWAM adapts pretrained diffusion transformers to robot control via current-observation denoising and action prediction, achieving 87.7% on LIBERO-Plus with halved tokens and 1.8x speedup.

Zanyi Wang, Yuheng Lei, Dengyang Jiang, Ping Luo and 3 more

Published Sep 23, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 9

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

GTR: Gated Token Recurrence for Efficient Dense Prediction

GTR replaces softmax attention with gated token recurrence for efficient high-resolution dense prediction, achieving 58.9 COCO box AP with low latency.

Zhe Feng, Longfei Liu, Wei Liu, Kai Chen and 6 more

Published Sep 22, 2026 · 0 citations · ▲ 13 on Hugging Face · Code ★ 51

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models

QuantWM is a training-free 2-bit KV cache quantization framework for video world models that preserves attention logits and token selection to eliminate temporal flickering while achieving up to 6.20x memory compression.

Jiaqi Zhao, Xiaobin Hu, Bo Yin, Junpeng Jiang and 2 more

Published Sep 22, 2026 · 0 citations · ▲ 17 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

OSWorld-Pro introduces process-based evaluation with 2,800 subgoals across 300 tasks, revealing top models achieve only 75.7% subgoal success versus 83.4% end-state performance and identifying distinct failure modes like irrelevant actions and click errors.

Zhilin Wang, Shaokun Zhang, Yifan Zhang, Hao E. Zhang and 8 more

Published Sep 21, 2026 · 0 citations · ▲ 22 on Hugging Face

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

RRSI regularizes recursive agent harness self-improvement via annealed edit budgets, trajectory exploration, and critical selection to boost out-of-distribution performance and reduce token use. It improves up to 14.1 points in-distribution and 4.7 points out-of-distribution while cutting policy tok

Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen and 10 more

Published Sep 21, 2026 · 0 citations · ▲ 222 on Hugging Face · Code ★ 1,293

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

AdaST: Adaptive Coupling for Spatial-Temporal Forecasting

AdaST adaptively decomposes and recombines spatial-temporal data via heterogeneity-aware experts to match distinct coupling regimes, significantly outperforming state-of-the-art forecasting baselines.

Zhenyu Lei, Chenghao Liu, Yushun Dong, Qi R. Wang and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Sep 20, 2026 · 0 citations

– ReadersNo votes yet. 1 from authors or colleagues not counted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

JEPA-Anything: Learning Predictive Models across Different Worlds

JEPA-Anything uses orthogonal predictive factorization to learn cross-domain predictive models that outperform baselines in vision, biology, clinical, control, molecular, physical, and weather domains.

Taoyong Cui, Zhongyao Wang, Xinyue Xu, Weiyang Liu and 9 more

Published Sep 17, 2026 · 0 citations · ▲ 77 on Hugging Face · Code ★ 273

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

Edge0 predicts next-layer MoE routing one token ahead to stream experts from SSD, serving 35B-class MoEs at 20 tok/s within 3 GiB active memory on a 24 GB machine via recovery LoRA adapters.

Yu Lin, Yiming Wang, Runyuan Cai, Liu, Hanze and 1 more

Published Sep 16, 2026 · 0 citations · ▲ 25 on Hugging Face · Code ★ 3,528

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Learning Steadily: Accumulating Relative Point Margin Scores for Face Image Quality Assessment

CARPM-FIQA accumulates relative point margin scores across training epochs to stabilize face image quality estimates, reducing variance and improving ranking stability near top performance.

Guray Ozgur, Tahar Chettaoui, Eduarda Caldeira, Marco Huber and 3 more

Published Sep 15, 2026 · ▲ 6 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence

LimiX-2 uses scaled contextual mechanism networks pretrained on synthetic causal data to outperform tabular foundation models and recover causal skeletons.

Xingxuan Zhang, Gang Ren, Hao Yuan, Hao Zou and 36 more

Published Sep 15, 2026 · 0 citations · ▲ 816 on Hugging Face · Code ★ 4,375

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

MTAC-IFBench benchmarks multi-turn instruction-following in agentic coding via progressive constraints, revealing rapid performance degradation in current code agents as sessions lengthen.

Bosi Wen, Cunxiang Wang, Jiayi Gui, Haoke Zhang and 5 more

Published Sep 14, 2026 · 0 citations

– ReadersNo votes yet. 1 from authors or colleagues not counted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

92%Must read

JumpStart Your Policy Learning with Lessons from 160,000 Training Runs

A 160,000-run study of offline policy learning finds no universal best algorithm, shows tuning and benchmark choice alter rankings, and releases resources with a dataset-conditioned recommender.

Nabil Omi, Eric Bae, Chung Yik Edward Yeung, Siddhartha Sen and 1 more

Published Sep 12, 2026 · 0 citations · Code

– ReadersNo votes yet
20/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Vidu S2 enables real-time interactive avatar and editing video generation at 720p with dynamic references and spatial capabilities, outperforming all baselines.

Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang and 31 more

Published Sep 10, 2026 · 0 citations · ▲ 706 on Hugging Face · Code ★ 502

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse-1 closes an evaluation-selection-update loop via harness-mediated routing, structured post-training, and capability-guided data allocation, raising 4B and 9B macro-averages by ~6 and ~3.4 points.

NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo and 33 more

Published Sep 8, 2026 · 0 citations · ▲ 327 on Hugging Face · Code ★ 1,673

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

UniMate: One Unified Model to Animate Diverse Skeletons

UniMate is a unified diffusion transformer that synthesizes motion for arbitrary skeletons from text and rigged assets without test-time optimization, using topology-aware attention and a new dataset to outperform specialized animators.

Linzhan Mou, Lei, Jiahui, Zhiyang Dou, Chenyue Cai and 3 more

Published Sep 4, 2026 · 0 citations · ▲ 23 on Hugging Face · Code ★ 1,553

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

HarnessDev evaluates LLMs creating and evolving agent harnesses, finding generated harnesses lag human references on coding and search but match them on writing and ML tasks, with unstable, model-dependent evolution gains.

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei and 15 more

Published Sep 1, 2026 · 0 citations · ▲ 566 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

StudentSim: Training LLM-based Student Simulators

StudentSim trains LLM student simulators via pooled training and per-student specialization, outperforming GPT-5.4 on behavioral fidelity and guidance responsiveness across chess, writing, and math.

Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh and 3 more

Published Sep 1, 2026 · 0 citations · ▲ 495 on Hugging Face · Code ★ 53

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers

LLM scorers with equal ranking quality make unstable threshold and preference decisions under candidate reordering, and order-consistency fine-tuning fixes it without harming quality.

Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, Navid Rekabsaz

Published Aug 27, 2026 · 0 citations · ▲ 21 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1 scales agentic intelligence via environment and coordination scaling to achieve leading complex-work performance with smaller models.

B. An, B. An, B. Wang, B. L. Wang and 36 more

Published Aug 24, 2026 · 0 citations · ▲ 212 on Hugging Face · Code ★ 5,146

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

VGI-Bench: Probing Visual Intelligence in Video Generation Models

VGI-Bench evaluates video generation models via 27 visual reasoning tasks, finding top models achieve only 51% accuracy with limited self-correction.

Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma and 19 more

Published Aug 20, 2026 · 0 citations · ▲ 336 on Hugging Face · Code ★ 14

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Grounding Memory Summarization in Utility Intent

MemSuit improves memory summarization by self-distilling query-conditioned utility awareness into raw-conversation entries and decomposing blocks to prevent collateral erasure, boosting answer quality across query types.

Zhenyu Lei, Mingjia Shi, Xingbo Fu, Haoyu He and 2 more

Published Aug 20, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment

TeacherGRPO aligns teachers to student distributions via reinforcement learning to overcome reasoning distillation's Gap Curse and improves student performance.

Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng and 4 more

Published Aug 20, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

EnvHarness: Awakening Static Worlds for Agent Learning

EnvHarness wraps static environments with programmable components to reshape agent behavior without altering underlying logic, improving benchmarks by up to 9.0 points while enabling continuous policy-environment co-evolution.

Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan and 13 more

Published Aug 20, 2026 · 0 citations · ▲ 175 on Hugging Face · Code ★ 619

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM improves long-horizon agent accuracy via durable-state harness scaling without model changes, reaching 95.3% on Terminal-Bench 2.1 and cutting API costs to about $15 versus $574.68.

Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang

Published Aug 15, 2026 · 0 citations · ▲ 452 on Hugging Face · Code ★ 1,321

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

The RA-Bench benchmark reveals current detectors fail to consistently identify AI-generated crisis videos, which become harder to detect after social dissemination and frequently mislead humans.

Shuo Liang, Yixing Ma, Pengfei Zhou, Zhenglin Wan and 32 more

Published Aug 14, 2026 · 0 citations · ▲ 287 on Hugging Face · Code ★ 134

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

BDH-CQ combines in-context learning with recurrent latent reasoning, achieving 29.5% ARC-AGI-1 pass@2 at $0.0007 per task to set a new cost-efficiency frontier.

Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska and 5 more

Published Aug 10, 2026 · 0 citations · ▲ 797 on Hugging Face · Code ★ 11,065

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

OpenART scales agent red teaming via open-ended environment evolution across 10,000 stateful scenarios, with EMHA achieving 85% attack success that grows with complexity.

Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu and 5 more

Published Aug 1, 2026 · 0 citations · ▲ 266 on Hugging Face · Code ★ 231

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Frontis-MA1 improves machine-learning engineering via recursive self-improvement using OpenMLE, boosting MLE-Bench Lite medal average from 39.39% to 71.21% and surpassing larger closed models.

Junlin Yang, Che Jiang, Yu Fu, Tianwei Luo and 20 more

Published Jul 30, 2026 · 0 citations · ▲ 189 on Hugging Face · Code ★ 782

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

AskChem indexes 2.4M atomic chemistry claims with provenance for cross-paper synthesis, achieving 100% resolvable DOIs and highest citation density.

Bing Yan, Gregory Wolfe, Stefano Martiniani, Kyunghyun Cho

Published Jul 30, 2026 · 0 citations · ▲ 309 on Hugging Face · Code ★ 16

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Kimi K3: Open Frontier Intelligence

Kimi K3 is a 2.8 trillion-parameter Mixture-of-Experts model with native vision and 1-million-token context that achieves frontier performance across reasoning, coding, and agentic tasks and outperforms comparable open and proprietary models.

Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao and 36 more

Published Jul 27, 2026 · 0 citations · ▲ 526 on Hugging Face · Code ★ 8,899

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score
Accident Analysis & Prevention 2026Pennsylvania StateZhejiangHuman-AI interaction

Mitigating driver confusion and aggression in mixed traffic: Effects of V2V communication and driving style

V2V communication and assertive automated vehicle driving styles reduce human driver confusion and aggression in mixed traffic interactions.

Haitao Chen, Yiqi Zhang

Published Jul 24, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Program-as-Weights compiles natural-language specs into compact local neural adapters that match large-model prompting with far less memory and faster offline execution.

Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published Jul 2, 2026 · 0 citations · ▲ 308 on Hugging Face · Code ★ 359

– ReadersNo votes yet. 1 from authors or colleagues not counted
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Orca: The World is in Your Mind

Orca is a world foundation model that learns a unified latent space via next-state prediction from video and language, enabling scalable text, image, and action generation.

Yihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen and 36 more

Published Jun 29, 2026 · 0 citations · ▲ 512 on Hugging Face · Code ★ 1,021

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition

AgentSpec modularizes embodied LLM agents into composable components, showing performance depends on scaffold compatibility and interaction effects rather than isolated module strength.

Jixuan Chen, Jianzhi Shen, Haoqiang Kang, Zhi Hong and 9 more

Published Jun 12, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

JoyAI-VL-Interaction is an open 8B vision-language model that continuously decides whether to speak, stay silent, or delegate in real time, outperforming Doubao and Gemini across six real-world scenarios.

Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin and 16 more

Published Jun 10, 2026 · 0 citations · ▲ 218 on Hugging Face · Code ★ 1,945

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Workflow-GYM benchmarks long-horizon professional GUI workflows, showing top agents achieve only ~30% success due to stage omission, error propagation, and objective drift.

Liya Zhu, Jingzhe Ding, Jian Zhang, Jianbo Xue and 36 more

Published Jun 9, 2026 · 0 citations · ▲ 221 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

ABot-Earth 0.5: Generative 3D Earth Model

ABot-Earth 0.5 generates seamless 3D environments from satellite imagery via 3D Gaussian Splatting, synthesizing square kilometers in under 10 minutes with real-time web visualization and embodied AI navigation support.

Ming Qian, Tianjian Ouyang, Mingchao Sun, Zijian Wang and 24 more

Published Jun 8, 2026 · 0 citations · ▲ 177 on Hugging Face · Code ★ 223

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

On the Geometry of On-Policy Distillation

On-policy distillation updates occupy a sparse, low-dimensional parameter subspace that is functionally sufficient and geometrically distinct from supervised fine-tuning and reinforcement learning.

Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong and 5 more

Published Jun 5, 2026 · 0 citations · ▲ 75 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Scaling Participation in Modular AI Systems

Modular participatory AI combines small stakeholder-trained models into compositional systems that outperform monolithic LLMs by up to 15.4% and exhibit emergent collaborative capabilities.

Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer and 2 more

Published Jun 5, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Agents' Last Exam

ALE introduces a benchmark evaluating AI agents on long-horizon, economically valuable real-world tasks across 13 industry clusters, finding current full pass rates below 1%.

Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang and 36 more

Published Jun 3, 2026 · 0 citations · ▲ 392 on Hugging Face · Code ★ 1,084

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

PaddleOCR-VL-1.6 applies region-aware data optimization and progressive reinforcement learning post-training to achieve 96.33% on OmniDocBench v1.6.

Zelun Zhang, Hongen Liu, Suyin Liang, Yubo Zhang and 11 more

Published Jun 2, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 90,719

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling

A lightweight RL controller formulates adaptive sampling as an MDP to balance LLM answer correctness, latency, and computation cost at test time.

Runpeng Dai, Tong Zheng, Rui Liu, Chengsong Huang and 1 more

Published Jun 2, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

SAHG: Sector-Anisotropic Hyperbolic Graph Model for Social Bot Detection

SAHG detects LLM-driven social bots by applying direction-dependent hyperbolic curvature and dual-channel feature fusion, achieving top accuracy and F1 across three benchmarks.

Hanning Lu, Yingguang Yang, Jinwei Su, Yang; Liu and 7 more

Published May 28, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

SkillOpt treats agent skills as external state optimized via bounded text edits validated on held-out scores, improving accuracy up to 24.8 points with stable transfer.

Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang and 11 more

Published May 22, 2026 · 2 citations · ▲ 267 on Hugging Face · Code ★ 18,090

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

HRM-Text: Efficient Pretraining Beyond Scaling

HRM-Text replaces Transformers with a hierarchical recurrent model and trains on instruction pairs to achieve competitive 1B-parameter performance with 100, 900x fewer tokens and far less compute.

Guan Wang, Changling Liu, Chenyu Wang, Cai Zhou and 5 more

Published May 20, 2026 · 0 citations · ▲ 323 on Hugging Face · Code ★ 2,136

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization

TextReg mitigates prompt distributional overfitting via regularized text-space optimization, improving out-of-distribution accuracy by up to 16.5% over prior methods.

傅卢成, Ye Yu, Yiyang Wang, Yiqiao Jin and 3 more

Published May 20, 2026 · 0 citations · ▲ 8 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read

Adaptive Fused Prior Transfer for Controllable Generative Image Compression

AFP-GIC transfers adaptive fused priors from a frozen pretrained model to guide generative compression without transmitting them, reducing decoder latency and parameters while improving very-low-bitrate naturalness.

Yifei Pei, Ying Liu, Nam Ling

Published May 16, 2026 · ▲ 4 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
67%Highly rated
?Highly ratedVote to see the score

Code for the manuscript titled "An in-sensor communication electronic textile for imperceptible and ultrarobust silent speech".

An in-sensor communication electronic textile enables imperceptible and ultrarobust silent speech.

LIN Yuchen

Published May 16, 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Code for the manuscript titled "An in-sensor communication electronic textile for imperceptible and ultrarobust silent speech".

The manuscript introduces an imperceptible, ultrarobust silent-speech electronic textile with in-sensor communication capabilities.

LIN Yuchen

Published May 16, 2026 · 0 citations

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

MinT manages LoRA adapter revisions over shared 1T-class base models to train and serve millions of policies via adapter-only handoffs and durable addressability.

Mind Lab, :, Song Cao, Vic Cao and 36 more

Published May 13, 2026 · 0 citations · ▲ 226 on Hugging Face · Code ★ 79

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

RewardHarness: Self-Evolving Agentic Post-Training

RewardHarness evolves agentic evaluation tools from minimal preference data to judge image edits, surpassing GPT-5 accuracy with 0.05% training annotations.

Yuxuan Zhang, Penghui Du, Bo Li, Cong Wei and 10 more

Published May 9, 2026 · 0 citations · ▲ 245 on Hugging Face · Code ★ 71

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

AutoTTS automatically discovers test-time scaling strategies via environment-driven controller synthesis, improving LLM reasoning accuracy-cost tradeoffs over manual baselines at minimal cost.

Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao and 9 more

Published May 8, 2026 · 0 citations · ▲ 70 on Hugging Face · Code ★ 176

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

MolmoAct2: Action Reasoning Models for Real-world Deployment

MolmoAct2 is an open vision-language-action model with a specialized reasoning backbone, open action tokenizer, continuous-action expert, and adaptive reasoning that outperforms closed and open baselines across embodied reasoning and robot deployment benchmarks.

Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang and 25 more

Published May 4, 2026 · 0 citations · ▲ 358 on Hugging Face · Code ★ 796

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

Show 20 more papers