Good Papers

Trending at NeurIPS 2026

Orals, spotlights and posters people are talking about

All sessions

Must read

This year's highest-rated papers

See all

Most debated

Where the reviewers can't agree

See all

All papers

How scores work
78%Highly rated
?Highly ratedVote to see the score

Agta hunter-gatherer oral microbiomes are shaped by contact network structure

Agta hunter-gatherer oral microbiomes resemble Central African foragers more than neighbors, with contact networks predicting bacterial transmission and central individuals as supersharers.

Federico Musciotto, Begoña Dobón, Michael John Greenacre, Álex Mira and 12 more

Published Dec 31, 2030 · 0 citations

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

LEGO replaces depth lifting with a learned view synthesizer that supplies structural egocentric conditions for diffusion models, outperforming explicit reconstruction pipelines.

Suhwan Cho, Yonwoo Choi, Soongjin Kim, Jicheol Park and 1 more

Published Oct 8, 2026 · ▲ 2 on Hugging Face · Code ★ 14

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

SpatialOPSD self-distills verified spatial coding traces into tool-free multimodal models via repetition-aware distillation, outperforming SFT and GRPO on spatial and out-of-distribution benchmarks.

Rongxue Li, Meng Yang, Yiru Mao, Yongliang Tao and 5 more

Published Oct 8, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training

DreamTrue is a cross-embodiment robot world model that uses image-space action conditions and counterfactual post-training to improve action following and physical plausibility, cutting human-assessed interaction defects from 48.12% to 6.25%.

Junyan Li, Ruizhi Li, Yu Liu, Xiangshuo Liu and 5 more

Published Oct 8, 2026 · ▲ 13 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

SparseDecoding uses autoregressive generation activations to calibrate LLM pruning and an optimized sparse matrix-vector kernel, achieving up to 1.48x decoding speedup with higher accuracy.

Qitong Wang, Xinwei Niu, Mingluo Su, Shanwei Zhao and 2 more

Published Oct 8, 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

TokenRouter: Efficient Serving System for Token-Level LLM Routing

TokenRouter is a serving system for token-level LLM routing that uses request-centric programming and delayed-batching subservers to achieve up to 64x higher decoding throughput.

Tianyu Fu, Tengxuan Liu, Ruoxi Wang, Yixin Dong and 3 more

Published Oct 8, 2026 · ▲ 22 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

USDCraft: Geometrically Grounded Programmatic Modeling of Articulated 3D Assets for Simulation

USDCraft uses LLM-written executable programs grounded in partial geometry to generate simulation-ready articulated 3D assets, achieving top articulation recovery and real-to-sim-to-real robot manipulation.

Chuanrui Zhang, Zaijia Yang, Duomin Wang, Lu Shi and 3 more

Published Oct 8, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals

InfiLoop introduces loop-native residual connections with learned temporal decay to preserve correct reasoning states across tens of thousands of iterative updates, reaching 97.9% Sudoku-Extreme accuracy and continuous improvement past 20,000 steps.

Pengxiang Li, Dilxat Muhtar, Di He, Guinan Su and 2 more

Published Oct 8, 2026 · ▲ 2 on Hugging Face · Code ★ 1

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

From Prompting to Composing: A Spatial Canvas Interface for Poster Generation

Spatial Canvas Interface enables direct spatial composition for poster generation via semantic, identity, text, and pixel bindings, and Compo achieves stronger compositional controllability than prompting-based alternatives.

Yitong Wang, Fangyun Wei, Jinjing Zhao, Sirui Zhang and 4 more

Published Oct 8, 2026 · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Reasoning-Informed Visual Editing

RISEBench++ introduces reasoning-based visual editing benchmarks spanning six reasoning dimensions, and RISE-Agent outperforms existing editing approaches, though top models achieve only 56.6% accuracy.

Xue Yang, Peiyuan Zhang, Yilun Zhu, Qihao Yang and 11 more

Published Oct 8, 2026 · ▲ 2 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers

Mixed-modality retrievers suffer "chaos in the text": irrelevant text degrades retrieval more than images as modalities mix, due to text preference; proposed Trident balances views via multi-positive contrastive learning to fix it.

Yubo Sun, Chunyi Peng, Yukun Yan, Zhenghao Liu and 5 more

Published Oct 8, 2026 · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

VibeEdit: Image Editing with Canvas Instructions

VibeEdit enables precise image editing via on-image spatial marks and notes, outperforming text-only methods on disambiguating similar objects.

Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu and 8 more

Published Oct 8, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

SuperNav: An Agentic Navigation System for Any Task in Any Scene

SuperNav equips a frozen multimodal LLM with navigation tools and visual-point interfaces to handle diverse tasks across environments without navigation-specific fine-tuning, outperforming baseline methods.

Jinkai Zhang, Jingyi Xu, Yuanhong Yu, Jiarui Guo and 4 more

Published Oct 8, 2026 · ▲ 32 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

V-CoLA: Vision Token Compression with Linear Attention

V-CoLA provides training-free vision token compression for linear-attention vision-language models, preserving 99.5% performance with 50% tokens and up to 6.15x faster prefill.

Hao Jiang, Yiru Mao, Tianpeng Bu, Hao Zhou and 8 more

Published Oct 8, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

AgentGarten: Code Worlds for Evolving Agents

AgentGarten couples game engines and simulators with a neural renderer to build interactive code worlds, letting agents learn from just four rounds of experience instead of millions.

Jiawei Chi, Shangchen Miao, Zhiyuan Shi, Kailu Wu and 10 more

Published Oct 8, 2026 · ▲ 10 on Hugging Face · Code ★ 54

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute

A memory layer stores 50-million-token model states to NVMe, loading them byte-exact 2.8, 4.3x faster than recompute with 82, 98% long-range recall accuracy.

Sietse Schelpe

Published Oct 7, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

RoboJEPA: Scaling Robotic Latent World Models

RoboJEPA scales multi-embodiment robotic world models via JEPA, finding imagination error and planning improve predictably with compute according to power laws, and deploys zero-shot for long-horizon real-robot tasks.

Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan, Sarath Chandar and 8 more

Published Oct 7, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

DSReg: Provably Recovering Individual World Latents without Reconstruction

DSReg recovers individual world latents without reconstruction via structural diversity and dependency-sparsity regularization. It achieves signed-permutation identifiability post hoc on any linearly identified JEPA representation.

Yujia Zheng, David Klindt, Randall Balestriero, Bernhard Schölkopf

Published Oct 7, 2026 · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

MIMESIS is a purpose-built user simulator trained on human conversations with reasoning supervision and realistic behavioral patterns that outperforms frontier models in fidelity and yields stronger agent generalization via multi-turn reinforcement learning and coaching self-distillation.

Hoang Phan, Dat Huynh, Andrey Zhmoginov, Qi Zeng and 6 more

Published Oct 7, 2026 · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

Iris-3B scales pixel-space diffusion to 3B parameters and competitive text-to-image quality, but fine-tuning shows no significant downstream gains over latent models.

Hanqiu Li Cai, Chema Garabito

Published Oct 7, 2026 · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Δ-MOPD transfers teacher-minus-base logit shifts rather than endpoint policies to multi-teacher on-policy distillation, improving composed and phased routing results. Removing inherited base pull reduces target-student divergence, yielding gains of up to 4.11 Math and 1.95 benchmark points. Target c

Hejian Sang, Zhengze Zhou, Shayan Mohajer Hamidi, Xiaomin Li and 2 more

Published Oct 7, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

System Switch: When Should a Fast Decision Model Stop and Think?

A fast actor defers to a reasoning vision-language model via a confidence gate in real-time Doom, improving decisions proportionally to calibration but failing to reach exits in closed-loop play.

Gian Luca Bailo

Published Oct 7, 2026 · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score
arXivDeep RL

Q-Learning with Scalar Adjoint Matching

Pretrained flow policies' velocity Jacobians concentrate diagonally, enabling a scalar adjoint that eliminates per-step vector-Jacobian products, and combining this with a value penalty yields SQAM, which improves hard OGBench success by 18, 35 points and outperforms supervised fine-tuning on real b

Yonghoon Dong, Minsung Yoon, Jaehyuk Kim, Jungwoo Park and 2 more

Published Oct 7, 2026 · ▲ 9 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

UniSkill proposes actor-aligned skill edits via shared policy trajectories and contrastive action feedback, achieving 98.4% ALFWorld and 84.7% WebShop success without proposal rollouts.

Yifei Lu, Cheng Liu, Dianzhi Yu, Hui Xiang and 3 more

Published Oct 7, 2026 · ▲ 3 on Hugging Face · Code ★ 1

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

EngramEdit updates LLM facts via conditional memory by computing target representations and penalizing shared embedding changes, achieving near-perfect edits with preserved unrelated knowledge.

Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu and 3 more

Published Oct 7, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

PersonTTS amortizes agentic test-time scaling policy discovery across personalized multi-dimensional accuracy, latency, and cost requirements via experience reuse and distilled procedural guidance.

Xinglin Wang, Zishen Liu, Tong Zheng, Shaoxiong Feng and 8 more

Published Oct 7, 2026 · ▲ 16 on Hugging Face · Code ★ 2

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026RL for LLMs

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

SkillForge co-evolves LLM agent skills via fitness-driven lifecycles of trial, active, stable, and retired states, achieving up to 7.8% relative success improvement over baselines.

Yuyao Ge, Yiwei Wang, Yuchen He, Baolong Bi and 4 more

Published Oct 7, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

RoboQuest: Generalist Physical Agents that Search, Inspect and Test

RoboQuest benchmarks embodied agents on search, inspection, and testing tasks requiring interactive evidence gathering, finding frontier models succeed in only 23% of episodes and mostly fail by exploring too little.

Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria

Published Oct 7, 2026 · ▲ 3 on Hugging Face · Code ★ 2

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

QuadTok uses a hierarchical quadtree tokenizer to dynamically allocate tokens by visual complexity, achieving efficient autoregressive image generation with 2.08 gFID on ImageNet 256×256.

Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang and 3 more

Published Oct 7, 2026 · ▲ 12 on Hugging Face · Code ★ 6

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

UltraText Bench is a bilingual benchmark of 432 prompts for evaluating dense visual text rendering in image generation across fidelity, clarity, placement, and scene quality. It reveals trade-offs across 24 model configurations and sharp performance declines from easy to hard workloads.

Deyuan Liu, Yihao Hu, Jingxuan Zhang, Xingying Li and 12 more

Published Oct 7, 2026 · ▲ 64 on Hugging Face · Code ★ 15

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation

AdSpark introduces a 300K dataset and six-dimension benchmark for product-centric ad video generation, revealing key challenges in product preservation and multi-shot storytelling.

Zhifei Yang, Zhao Jiang, Keyang Lu, Honghe Zhu and 5 more

Published Oct 7, 2026 · ▲ 16 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

RobotWorld benchmarks multimodal agents on 84 simulated robot tasks spanning manipulation to aerial control, revealing sophisticated but unreliable physical-world execution that varies across models.

Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu and 29 more

Published Oct 7, 2026 · ▲ 29 on Hugging Face · Code ★ 25

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

RunningTab: Direct Workspace Interaction with Environment-Side Tabs

RunningTab introduces an environment-side tab tracking task requirements, file excerpts, and unread candidates for direct workspace interaction, improving agent deliverable accuracy over model-side tracking.

Jinheon Baek, Soyeong Jeong, Yumin Choi, Dongsu Han and 1 more

Published Oct 7, 2026 · ▲ 34 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Tetris3D: 3D Scene Generation With Objects That Fit Together

Tetris3D explicitly conditions each object's shape and pose on neighboring geometry and physical relations to recover physically coherent 3D scenes, achieving state-of-the-art generation and stability.

Jaeyeong Kim, Jinhyuk Jang, Jongmin Lee, Kyehong Park and 1 more

Published Oct 7, 2026 · ▲ 35 on Hugging Face · Code ★ 22

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
arXivPrivacy

Inverting Multi-Vector Visual Document Indices

Multi-vector document indices can be inverted to recover readable pages because patch vectors retain document layout and vision-language features, achieving high source retrieval and substantial text reconstruction.

Zhuchenyang Liu, Yao Zhang, Yu Xiao

Published Oct 7, 2026 · ▲ 18 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read

Long-WAM: Scaling the Context of World-Action Models

Long-WAM scales causal world-action model context via autoregressive pretraining, raising robot success up to 78.7% and enabling real-time deployment.

Wei Huang, Bohan Zhang, Chenzhi Liu, Isabella Liu and 12 more

Published Oct 7, 2026 · ▲ 88 on Hugging Face · Code ★ 2,687

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

SGF+ separates context-writing and denoising parameters to fix conflicting gradients, improving autoregressive video quality and enabling continuous generation up to 24 hours without long-video training.

Zihan Su, Junhao Zhuang, Yaowei Li, Siwen Lu and 9 more

Published Oct 7, 2026 · ▲ 51 on Hugging Face · Code ★ 52

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

GRACE: Generation-aware latent compression for efficient video generation

GRACE compresses pretrained video autoencoders via frozen base latents and learned residuals aligned to frozen DiT features, cutting Wan2.1 tokens 8x and latency 11.1x without quality loss.

Jiyoung Kim, Paul Hyunbin Cho, Jisu Nam, Donghoon Lee and 4 more

Published Oct 7, 2026 · ▲ 64 on Hugging Face · Code ★ 17

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

Hybrid attention models combine full and sliding-window or linear attention, revealing a context-extension seesaw effect, positional biases causing short-context traps, and a sliding-window linear attention method achieving 16x training-free length extrapolation with perfect NIAH retrieval at 64k co

Xiaoran Liu, Ziwei He, Xipeng Qiu

Published Oct 7, 2026 · ▲ 30 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

RLHND repurposes a video diffusion backbone to track physically consistent hand poses and estimate dense tactile contact and force from monocular egocentric video, achieving state-of-the-art results and improving robot learning retargeting.

Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo and 1 more

Published Oct 7, 2026 · ▲ 19 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents

This paper constructs an abstract oracle machine unifying agents and workflows via a stack-query automaton, and implements it as ArchNights.

Kefan Liu, Fengning Ou, Yelin Luo, Jingdi Lei

Published Oct 7, 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

SanSi loops language model layers to make typed decisions without text generation, reaching 72.0% accuracy across 59 sources and extending solvable reasoning depth beyond training limits.

Shuyu Gan, Young-Jun Lee, Dongyeop Kang

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Self-Retrospection Distillation turns post-hoc experience into pre-interaction foresight predictions, boosting agent success up to 60.6% versus 0.0% for RLVR when rewards are uniform.

Haoxiang Zhang, Qinglin Chen, Hiroaki Hayashi, Zhuofeng Li and 8 more

Published Oct 6, 2026 · ▲ 14 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Agent Plasticity: Measuring Self-Improvement Through Experience

Agent plasticity measures how efficiently agents turn experience into future held-out gains, showing frontier models diverge sharply in self-improvement efficiency and transfer.

Harman Singh, Anton Bakhtin, Rulin Shao, Gabriel Synnaeve and 7 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Cross-tokenizer on-policy distillation achieves comparable accuracy with strict top-16 shared-vocabulary supervision versus full coverage, while expanded span supervision reduces accuracy due to conflicting gradients, motivating prioritization of supervision reliability over alignment coverage.

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li and 3 more

Published Oct 6, 2026 · ▲ 165 on Hugging Face

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

A self-learning scientific agent for X-ray diffraction

Gan Jiang learns reusable X-ray diffraction skills by diagnosing failures and revising code, achieving up to 96.30% single-phase identification accuracy without retraining.

Bin Cao, Huichi Zhou, Runyu Yang, Jingsong Li and 6 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use

CADFather autonomously reconstructs editable CAD programs via coordinated vision-language planning, complementary proposal tools, and numerical optimization, achieving strong validity and accuracy across benchmark datasets without retraining.

Gennadiy Savrasov, Maksim Elistratov, Nikita Gavrilov, Albert Garifullin and 6 more

Published Oct 6, 2026 · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Task-Sufficient Contraction: Source Selection for Machine Information Interfaces

Task-Sufficient Contraction defines pre-selected source reductions that preserve entire downstream problem families, yielding exact rate-regret curves for fixed-action machines and affine quadratic loss via projected sources.

Joss Armstrong

Published Oct 6, 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Co-Evolving Robot Orchestrators and Policies through Deployment

Robo-COP co-evolves robot orchestrators and policies during deployment, improving held-out success to 73.8% in simulation and 50.0% in real-world tasks.

Xilun Zhang, Maggie Wang, Erik Bauer, Hong-Xing Yu and 3 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

DLoop: Looped Speculative Decoding

DLoop loops multiple speculative drafting stages before target verification, reducing target model passes and improving wall-clock speedup by 5, 41% losslessly.

Geonmo Gu, Byeongho Heo, HeeJae Jun, Yoohoon Kang and 3 more

Published Oct 6, 2026 · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

PhysEvo: Astra Can Act, Let It

PhysEvo enables frozen-model robotic self-improvement via recursive tool and skill revision, achieving 62% success on 42 tasks and 84% on real-world manipulation.

Wenqing Tian, Zeyu Zhang, Zhaocheng Liu, Fengwei Liu and 2 more

Published Oct 6, 2026 · ▲ 27 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

On-Policy Distillation with Negative-Policy Rollouts

Negative-Policy OPD improves on-policy distillation by using negative-policy rollouts to supply explicit negative signals, boosting performance across scales and reasoning tasks.

Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo

Published Oct 6, 2026 · ▲ 16 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Recurrent Looped Transformer

Recurrent Looped Transformer merges a parallel encoder with a recurrent decoder to grow per-token computation depth with sequence length, achieving near-perfect length generalization on parity and permutation tasks versus chance-level Transformers.

Yifan Zhang, Jichen Feng, Shihan Qin

Published Oct 6, 2026 · ▲ 26 on Hugging Face · Code ★ 915

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness

Recursive Game Creator uses recursive Designer-Builder-Player-Reviewer loops to advance agentic games from prototypes to entertaining products, achieving 77.89 on GameCraft-Bench and 53.2% success on GameASG-Bench with higher user ratings.

Jiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu

Published Oct 6, 2026 · ▲ 79 on Hugging Face · Code ★ 29

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

DecepEval: A Benchmark for Evaluating Deception in LLM Agents

DecepEval introduces a benchmark of 1,532 instances across 28 scenarios and proposes a framework showing inducements increase LLM deception rates.

Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen and 7 more

Published Oct 6, 2026 · ▲ 66 on Hugging Face · Code ★ 1

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

WorldSonus: Bringing Sound to Worlds

WorldSonus generates interactive, spatially aligned stereo audio for world models via streaming causal diffusion with real-time factor 0.41, matching state-of-the-art quality on open benchmarks.

Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang and 8 more

Published Oct 6, 2026 · ▲ 33 on Hugging Face · Code ★ 53

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

On KL-Regularized Policy Optimization

KLPO anchors KL regularization to the sampler for closed-form updates without importance weights, critics, or grouped rollouts, and subsumes SPPO, GPO, REBEL, and BPO.

Yifan Zhang

Published Oct 6, 2026 · ▲ 19 on Hugging Face · Code ★ 182

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

CheckerBench evaluates long-horizon agents on synthesizing static-analysis checkers across 300 CVE-derived tasks, finding best Pass@1 reaches 45.33%.

Hang He, Li Wang, Hao Chen, Yuchen Shao and 8 more

Published Oct 6, 2026 · ▲ 53 on Hugging Face · Code ★ 1

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine co-evolves policies, curricula, and judges via adaptive diagnosis and coactive calibration to sustain embodied reasoning self-improvement. Experiments on driving and navigation show continuous gains in both policy and judge capability.

Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao and 9 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

81%Must read
?Must readVote to see the score

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

Attacca trains visual goal-conditioned policies on complete search-to-interact trajectories with decoupled goal images and behavioral-phase conditioning to improve long-horizon embodied task success by up to 7x.

Gyusik Seo, Jaehong Yoon

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code ★ 3

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

UNREAL: Unifying Retrieval and Long-Context with a Single Model

UNREAL unifies retrieval and long-context evidence selection via frozen LLM representations with minimal parameters, outperforming state-of-the-art retrievers and improving long-context accuracy substantially.

Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel and 3 more

Published Oct 6, 2026 · ▲ 21 on Hugging Face

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Sensor-Language-Action Models

Sensor-Language-Action modeling unifies multimodal sensors, language, and actions via a semantic interface, and OpenSLA achieves superior hierarchical prediction and explanation with zero-shot generalization.

Yuekai Xu, Zitao Shuai, Yuzhe Yang

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code ★ 6

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

90%Must read
?Must readVote to see the score

World Models' Last Exam in Physics

World Models' Last Exam in Physics benchmarks video models via 40 measurement-based physics tasks, finding the best model scores 57.76/100 with widespread inconsistencies.

Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin and 7 more

Published Oct 6, 2026 · ▲ 7 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

CtrlCache accelerates interactive video world models via control-aware caching that detects action changes to reuse transformer residuals and apply frequency-mixed history guidance, achieving up to 1.41x speedups with improved quality.

Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh and 1 more

Published Oct 6, 2026 · ▲ 2 on Hugging Face · Code

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

81%Must read
?Must readVote to see the score

Sherpa: Teaching LLMs to Teach Adaptively

Sherpa uses multi-turn reinforcement learning to train LLM teachers that adapt instructions to diverse student archetypes, improving student performance by 20.5 points and pedagogy scores to 79.2%.

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen and 1 more

Published Oct 6, 2026 · ▲ 9 on Hugging Face · Code ★ 6

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

TRACE aligns FP4 quantization between RL training and rollout paths via rollout-guided quantization-aware training for MoE language models, achieving BF16-comparable RL performance with up to 5.4x rollout speedup.

Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao and 8 more

Published Oct 6, 2026 · ▲ 93 on Hugging Face

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

90%Must read
?Must readVote to see the score

From Evidence to Action: How Tool-Using Agents Fail

Tool-using agents often fail by acting before establishing required evidence or leaving multi-action workflow prerequisites unresolved, despite accurate static action assessment. SafeActBench reveals failures stem from how agents use established evidence during execution, not just missing informatio

Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai and 2 more

Published Oct 6, 2026 · ▲ 38 on Hugging Face · Code ★ 7

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS bootstraps reusable agent memory from self-generated practice tasks without oracles, improving success rates by up to 15.9 points across benchmarks.

Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard and 2 more

Published Oct 6, 2026 · ▲ 10 on Hugging Face · Code ★ 6

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Building Rome from a Single Image

A redesigned object-centric generator partitions scenes into distance-adaptive chunks, captures 2D-3D correspondence, and trains on 4,000 outdoor scenes to outperform baselines in indoor and outdoor mesh generation.

Jiraphon Yenphraphai, Fang Li, Tianshuo Xu, Depu Meng and 4 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Towards In-Parameter Memory Augmentation for Large Language Models

This survey organizes in-parameter memory augmentation for LLMs by parameter placement and acquisition time to enable reusable parametric knowledge at deployment.

Haoyu Huang, Zhongwei Xie, Jiaxin Bai, Yisen Gao and 5 more

Published Oct 6, 2026 · ▲ 7 on Hugging Face · Code ★ 3

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Personal-Agent Mediated Recommendation with Cross-Platform User History

Personal-agent mediated recommendation balances cross-platform user history against platform rankings via the MediateRec benchmark and PAMO optimization to improve rescue-harm trade-offs.

Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley and 1 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

AdvSim2Real co-evolves tasks, adversarial injections, and a web agent in a simulator, boosting 4B agent completion by 33.6% against unseen adaptive attacks and transferring gains to real browsers.

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov and 2 more

Published Oct 6, 2026 · ▲ 7 on Hugging Face · Code ★ 3

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

NeMo-DCR bit-exactly refits trillion-parameter policies by streaming delta-compressed weight changes via affine mappings and XOR masks, cutting 1T cross-region refits from 87.5 minutes to 150 seconds.

Songlin Jiang, Zhiyu Li, Terry Kong, Yu Yao and 4 more

Published Oct 6, 2026 · ▲ 13 on Hugging Face · Code ★ 2,052

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation

EmbodiedSmith unifies asset, scene, and task generation in a recursive self-improvement loop to scale embodied simulation data, improving generation success and robot policy generalization across diverse embodiments and physics.

Yikai Qin, Yifei Deng, Mingjian Liang, Wenxuan Song and 12 more

Published Oct 6, 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

HiPLEX factorizes full-duplex speech policies into timing and content controllers to jointly optimize interaction dynamics via reinforcement learning. It lowers takeover rates, reduces interruption latency, and improves human-like turn timing versus GRPO.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026 · ▲ 6 on Hugging Face

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Speculative tool execution predicts tool calls from partial ASR to run them during speech, cutting median voice-agent response latency from 5.79 s to 4.60 s.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026 · ▲ 8 on Hugging Face

100% Readers1 of 1 upvoted
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

ALIVE inserts objects that interact with video contents via first-frame editing and interaction guidance, outperforming baselines on interaction and insertion benchmarks.

Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code ★ 3

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Dev-Primitives turn repository artifacts into active, resident-LLM agents with self-modification interfaces, and HERMES improves software engineering benchmarks by 12.4% over baselines while cutting inference costs by 26.2%.

Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang

Published Oct 6, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation

Trace2Env builds agentic language world models from interaction traces to simulate interactive environments, improving observation fidelity and long-horizon consistency over prompt-based methods without requiring executable systems.

Quanyu Long, Xiao Chen, Jianda Chen, Haozhen Zhang and 3 more

Published Oct 5, 2026 · 0 citations · ▲ 120 on Hugging Face · Code ★ 17

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

AgentMonBench evaluates monitors identifying consequential agent decisions using evidence-grounded behavior graphs that improve verification across eight models.

Zhongxiang Sun, Jiahao Yan, Hongkang Zhao, Haojie Ding and 4 more

Published Oct 5, 2026 · 0 citations · ▲ 7 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Optimizing the Optimizer: Language Models Discover Faster Molecular Relaxation

An agent rewrites the Sella optimizer into AutoSella, cutting force calls to 40, 77% of Sella's count across benchmarks without using DFT gradients.

Artem Tsypin, Vladimir Deshchenya, Kuzma Khrabrov, Denis Potapov and 3 more

Published Oct 5, 2026 · 0 citations · ▲ 23 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score
arXivPrivacy

Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

Analytical Memory Units attach derivation graphs to cached results and use lineage-gated retrieval to block unauthorized derived access, cutting cross-department leakage by 18.8, 25.5% with 90% lineage completeness.

Venkata M Sangaraju, Sudhir Vissa

Published Oct 5, 2026 · ▲ 1 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

VepAgent integrates causal-transition reasoning with tool-augmented reinforcement learning for video event prediction, achieving state-of-the-art results on FutureBench and NEPBench.

Qiutong Chen, Yuchan Guo, Zhenlong Yuan, Haobo Yang and 9 more

Published Oct 5, 2026 · 0 citations · ▲ 53 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Improving Proactive AI Assistance with Hierarchical Procedural Understanding

ProactiveCoach introduces hierarchical procedural data and an adaptive guidance router that improves proactive AI assistance timing and granularity adaptation by up to 57.1%.

Jin-Seop Lee, TaeYeon Won, SeongJun Jung, Junghoon Kim and 4 more

Published Oct 5, 2026 · 0 citations · ▲ 9 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

D-OPCD distills agent harness improvements into diffusion weights via on-policy context distillation, raising direct text-to-image scores from 60.52 to 65.09 and enabling continual co-evolution.

Wenxuan Wang, Zekai Liu, Weinan Zhang, Yu Cheng and 1 more

Published Oct 5, 2026 · ▲ 10 on Hugging Face · Code ★ 7

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Minimal Witness Reinforcement Learning

MWRL learns minimal sufficient witnesses via union-based credit assignment to recover diverse alternatives from black-box verifiers, outperforming methods that yield single or redundant solutions.

T. Y. Tsui, Zihao Ye, Pengxiang Cai, Yanchao Li and 2 more

Published Oct 5, 2026 · ▲ 18 on Hugging Face · Code

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

TRIAGE stabilizes native NVFP4 RL by direction-aware mismatch diagnosis and selective gradient rebalancing, achieving full-precision math reasoning with 2.3x rollout throughput.

Zhen Li, Shuai Zhang, Yanggan Gu, Yiming Zhang and 6 more

Published Oct 5, 2026 · ▲ 29 on Hugging Face · Code ★ 6

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Structuring MoE Expert Selection for Agentic Reinforcement Learning

Agentic RL on MoE models reveals structured expert routing by operation type; hierarchical routing control and entropy gating improve success rates over 10 points.

Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh, Sanjoy Chowdhury and 2 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco

Turba provides an open-source Moroccan fertilizer recommendation stack with programmatic access, versioned data, and loadable crop-specific machine learning surrogates for reproducible benchmarking.

Abdelghani Belgaid, Zakaria Mahmoud, Fahd Chibani, Oumnia Ennaji and 2 more

Published Oct 5, 2026 · 0 citations · ▲ 2 on Hugging Face · Code ★ 5

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

58%Worth a look
?Worth a lookVote to see the score

Empirical Variational Autoencoder

Empirical Variational Autoencoder learns autoregressive latent priors empirically via one linear layer to close the VAE prior-posterior gap, yielding high-fidelity sequential generation competitive with diffusion models at much faster inference.

Kaede Shiohara

Published Oct 5, 2026 · ▲ 11 on Hugging Face · Code ★ 9

0% Readers0 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Execution-Aligned Progressive Noise for Consistent Asynchronous Replanning in Generative Robot Policies

Execution-Aligned Progressive Noise structures noise across chunks and time to maintain consistent generative states during asynchronous replanning, achieving up to 96.7% real-robot success.

Di Wu, Ping Liu, Xuhua Chen, He Zheng and 2 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

Stepped MoE unifies elastic architectures and sparse gating to adapt model capacity to deployment constraints and input requirements, outperforming dense counterparts by 2-5%.

Arnav Kundu, Zhaoyang Xu, Bairu Hou, Chang Gao and 2 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated

RealtimeWAM: One-Step Asynchronous World Action Models

RealtimeWAM uses teacher-anchored consistency distillation and cross-expert wavefront pipelining for one-step asynchronous action generation, achieving near-lossless performance with ~25x speedup.

Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong and 6 more

Published Oct 5, 2026 · ▲ 20 on Hugging Face · Code ★ 2,885

0% Readers0 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

OnePO enables RL-only medical domain adaptation via adaptive objective evolution and teacher retirement, yielding HuatuoGPT-3 that surpasses frontier models.

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu and 6 more

Published Oct 5, 2026 · ▲ 29 on Hugging Face · Code ★ 19

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026RL for LLMs

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

RGPO adaptively uses ground-truth rationales as temporary scaffolds to generate improved responses for on-policy RL, then transfers only higher-reward model outputs back, reducing reward sparsity and improving text and multimodal reasoning.

Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde and 2 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

WildMatch adapts pretrained image matchers to wildlife identification using only identity labels, improving retrieval accuracy and learning transferable matching priors without keypoint annotations.

Turhan Can Kargin, Piotr Kubaty, Ekaterina Rostovskaya, Izabela Wierzbowska and 2 more

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

S2PD switches from autoregressive to parallel diffusion during denoising to enforce physical and logical consistency with faster sampling than fully serial methods.

Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari

Published Oct 5, 2026 · 0 citations · ▲ 2 on Hugging Face · Code ★ 8

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

MEND: RL For Flow Models via Proximal Velocity Matching

MEND uses proximal velocity matching to cap rewards and accept only cost-effective sample moves, outperforming prior flow-model RL methods in far fewer updates without KL penalties or reference models.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli and 1 more

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Conditional Trajectory Peaks: Single-Pass Multimodal Policies over Action Chunks

Conditional Trajectory Peaks predicts multimodal action-chunk candidates in a single pass, achieving 97.25% LIBERO success and 3× faster inference while preserving diverse behaviors.

Di Wu, Rongtian Shen, Ping Liu, Xuhua Chen and 3 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Learning to Read the Contextual Tokens in Diffusion Transformers

A framework maps diffusion transformer contextual tokens through a frozen LLM to reveal they encode global emerging scene semantics early, inspiring contextual alignment that improves generation quality.

Omer Dahary, Etai Sella, Hadar Averbuch-Elor, Daniel Cohen-Or and 1 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Agents judge failing retrieval results useless but rarely stop; enforcing answers after five useless judgments improves success and fixes stopping.

Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu and 3 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

JLD: Perceptual Distance Through A Jacobian Lens

JLD defines a perceptual image distance via a Jacobian-derived metric tensor from frozen vision encoders, achieving state-of-the-art correlation with human judgments and resolution robustness.

Shreshth Saini, Balu Adsumilli, Alan C. Bovik

Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Hybrid Linear Attention introduces query-dependent chunk-level routing for Gated DeltaNet, improving long-context benchmarks by up to 5.57 points via adaptive recurrent memory composition.

Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He and 2 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

MiniCorp: The Last Mile of the AI Agent Firm

MiniCorp is a simulated office environment that generates longitudinal, counterfactual enterprise data to study autonomous AI-run companies and train adaptive agents.

Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin and 5 more

Published Oct 5, 2026 · ▲ 26 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
67%Highly rated

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

APO learns shared LLM aligner initializations via grouped gradient coordination to enable few-shot personalization under heterogeneous preferences and scarce feedback, improving over baselines with 20 local examples.

Liyan Yang, Yige Yuan, Zhiqin Yang

Published Oct 5, 2026 · ▲ 7 on Hugging Face

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
80%Highly rated
?Highly ratedVote to see the score

InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation

InterMimicGen retargets human motion capture to humanoids and self-evolves tracking data via iterative simulation-verified augmentation to scale dexterous loco-manipulation.

Yucheng Zhang, Sirui Xu, Jinhong Li, Liuyu Bian and 7 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

ACG-Bench evaluates arm-wise compositional generalization in dual-arm vision-language-action models via AE-VLA, which achieves 21.53% simulated and 39% real-world success versus under 6% baselines.

Zaibin Zhang, Binghao Ran, Yuhan Wu, Zhongbo Zhang and 9 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Agentic retrieval combining LLM reasoning with dense retrieval improves nDCG@10 by 8.7 points over standard retrieval but requires 107 seconds and 764K input tokens per query.

Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira and 7 more

Published Oct 5, 2026 · ▲ 17 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

On-policy power distillation trains models to generate sharpened answers directly, improving single-sample math reasoning by up to 27.3 points and outperforming multi-candidate sampling and reward-based methods.

Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa and 3 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

RACE predicts subskill transition timing to extend VLA action chunks reliably, reducing stop-and-go idle time ~5x on real robots while improving success rates.

Seonghoon Yu, Dongwon Kim, HyungRok Jung, Yoonjae Baek and 3 more

Published Oct 5, 2026 · ▲ 29 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
89%Must read

Certification of Real Images through Calibrated Content Authentication

Deepfake detectors degrade to 76% accuracy and near-zero under attacks, so calibrated reconstruction-based authentication bounds false real-image certification to 1%.

Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi and 1 more

Published Oct 5, 2026 · ▲ 16 on Hugging Face · Code ★ 3

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

Base Models Can Reason By Taking a Cue From Training Data

Fixing initial token cues in base models boosts reasoning to match RL performance, with effects traced to training data associations that can be causally edited.

Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat and 2 more

Published Oct 5, 2026 · ▲ 22 on Hugging Face · Code ★ 19

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

HLA-WM combines geometry-guided retrieval with recurrent linear attention to fix long-range forgetting in video world models, improving 60-second consistency metrics by up to 28.5% with 12× lower memory and no retraining.

Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu and 3 more

Published Oct 5, 2026 · ▲ 18 on Hugging Face · Code ★ 7

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
78%Highly rated

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

LoGRA reduces LLM reinforcement learning memory by up to 45.7% via low-rank gradient sketches and predicted-KL step control, enabling 27B-parameter training on single nodes.

Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li and 4 more

Published Oct 5, 2026 · ▲ 117 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Closing the Context Gap: Activation Alignment for Tabular In-Context Learning

Activation alignment trains a linear map to align partial-context student activations with full-context teacher activations, significantly improving tabular in-context learning efficiency and recovering much of the performance gap.

Yoel Zeldes

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code ★ 3

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
ICLR 2027Tabular data

Adapting prior-data fitted networks for tabular anomaly detection

Frozen and fine-tuned TabPFN representations for tabular anomaly detection yield ZEN and FOCUS, surpassing all ADBench baselines in AUROC despite unsupervised deployment and contaminated reference sets.

Maximilian Bershtman, Niv Cohen

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Learning to Learn a Language

Prior-Fitted Language Model, trained solely on synthetic non-linguistic data, learns to infer and predict real languages from context with frozen weights, achieving strong cross-lingual compression and reasoning without ever seeing real text.

Lennart Carstens-Behrens, Holger Fröhlich

Published Oct 5, 2026 · ▲ 8 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

SoK: Semantic Decision Engines in Network Control Loops

Systematizing 139 semantic decision engine families reveals most miss network deadlines and verification, with only four reporting deadline attainment; unverified decisions reverse admission verdicts under queued execution, prompting minimum reporting rules and a research agenda.

Delong Li, Chen Li, Xu Wang, Haochen Gong and 2 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Targeted bias injection via closed-loop activation steering exploits diffusion language model denoising trajectories to steer frozen models toward adversarial demographic answers with minimal corruption.

Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed, Muhammad Umer Sheikh and 2 more

Published Oct 5, 2026 · ▲ 15 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models

FairRSFM benchmarks remote sensing foundation models by biome to expose hidden ecological performance disparities and tests debiasing methods without backbone updates. Aggregate metrics consistently mask large biome-dependent gaps, though mitigation effectiveness varies by model and task.

Md Aminur Hossain, Omkumar Vaghasiya, Rajeev Ranjan Dwivedi, Vinod Kurmi and 1 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code ★ 1

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

What Matters for Latent Reasoning with Flow Matching

FLaRe uses flow matching for latent reasoning that is useful, diverse, explainable, refinable and efficient, reaching 97% of explicit chain-of-thought accuracy at 25% latency.

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

Published Oct 5, 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Capability-Driven Self-Evolution of Agent Memory

PrisMem drives agent memory self-evolution via capability-specific guidance, dependency-aware selection, and trace-guided integration, outperforming baselines by up to 10.54 points on million-token benchmarks.

Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu and 7 more

Published Oct 5, 2026 · ▲ 18 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Representation-Space MMD for Diffusion Language Models

Post-training minimizes representation-space MMD between diffusion language model outputs and references via retained token features, improving perplexity, accuracy, and parallel decoding.

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev and 6 more

Published Oct 5, 2026 · ▲ 30 on Hugging Face · Code ★ 12

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models use fixed-point convergence to enable truncated training, shared KV caches, faster prefill, and faster RL updates, while a learned depth prior and orthogonal input injection improve perplexity across scales.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen and 3 more

Published Oct 5, 2026 · ▲ 29 on Hugging Face · Code ★ 38

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

How corner is a corner case? Percentile control for highway scenario generation

The study defines scenario cornerness via history-conditioned risk percentiles and uses a guided diffusion model to generate multi-agent highway futures with precise percentile control, achieving 98.75% request fulfillment within tight tolerance on highD.

Jiaxi Liu, Hang Zhou, Hangyu Li, Yifan Wang and 4 more

Published Oct 4, 2026 · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video

CoDance learns reactive compliant human-humanoid interaction from video via multi-link force-aware training, enabling sustained partnered dancing with adaptive footsteps and compliant contact on a physical humanoid.

Zhuoqun Chen, Shucheng Jia, Boyuan Chen

Published Oct 4, 2026 · ▲ 1 on Hugging Face · Code ★ 5

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
75%Highly rated
?Highly ratedVote to see the score

Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads

Agentic RAG evaluation budgets favor broader question coverage over repeated reads or trajectories, reducing standard error by up to 33% at fixed token cost.

Jingjie Ning, Xueqi Li, Yibo Kong

Published Oct 4, 2026 · 0 citations · ▲ 19 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting

Mobile-4DGS enables real-time static and dynamic Gaussian splatting on mobile devices via compact appearance modeling, explicit 4D motion, and depth-order reuse, substantially reducing storage and rendering overhead.

Xiaobiao Du, Beixi Hao, Zhen Fang, Tianqing Zhu and 2 more

Published Oct 4, 2026 · 0 citations · ▲ 16 on Hugging Face · Code ★ 4

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score
arXivMusic

SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision

SheetSage2 unifies synthetic supervision, structured decoding, and autoregressive distillation to transcribe coherent lead sheets, outperforming prior systems on 12 of 15 benchmark metrics.

Junyan Jiang, Ruibin Yuan, Jiahao Pan, Wei Xue and 3 more

Published Oct 4, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

UndoBench separates tool-using agent competence from fault recovery via paired enterprise workflow trials, finding 83.54% nominal success but only 46.72% recovery success with phase-dependent vulnerabilities.

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published Oct 4, 2026 · ▲ 20 on Hugging Face · Code ★ 2

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
77%Highly rated

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Kandinsky 6.0 Video introduces diffusion models that generate synchronized 5-second audio-video clips with lip-sync via a dual-stream CrossDiT architecture, with the 29B-parameter Pro version outperforming its predecessor and matching top competitors in speech quality.

Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin and 36 more

Published Oct 4, 2026 · ▲ 155 on Hugging Face · Code ★ 222

100% Readers1 of 1 upvoted
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
86%Must read
?Must readVote to see the score

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

AutoSciBench autonomously generates and iteratively adapts scientific agent benchmarks via recipes and concepts, reducing solver accuracy by over 22 points versus human benchmarks while improving quality ratings.

Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards and 6 more

Published Oct 4, 2026 · ▲ 29 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

EVISKILL: Grounding Skill Evolution in Replayable Evidence

EVISKILL grounds LLM skill evolution in replayable evidence cards linking edits to supporting contexts, using targeted replay for verification and global validation for incorporation.

Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang and 1 more

Published Oct 4, 2026 · ▲ 41 on Hugging Face · Code ★ 26

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
88%Must read
?Must readVote to see the score

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

Feasible-future decoding reranks VLA actions by future safe-completion mass, reducing cumulative safety costs by up to 57.5% without retraining or rollouts.

Tu Nguyen, Matthieu Zimmer, Vu Anh Vu, Ziyi Wang and 3 more

Published Oct 4, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

LiFT: Loop Flow Transformers

Loop Flow Transformers loop a shared diffusion transformer with depth-indexed regression targets, improving generation with more inference compute and fewer parameters than dense models.

Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent and 1 more

Published Oct 4, 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
75%Highly rated
?Highly ratedVote to see the score

RobotUse: Allocating Computation, Context, and Decisions

RobotUse organizes robot computation, context, and decisions around revisable physical actions via visual target selection and persistent playbooks, achieving 45% RoboLab success and real-world learning.

Junhoo Lee, Injun Baek, Seungyeon Kim, Suhyun Jeon and 3 more

Published Oct 4, 2026 · ▲ 19 on Hugging Face · Code ★ 5

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT online test-time trains agents by self-distilling verified deployment trajectories into LoRA weights via a frozen hindsight model, improving long-horizon success and efficiency without external teachers or memory retrieval.

Haodong Lu, Dong Gong

Published Oct 4, 2026 · ▲ 23 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
71%Highly rated

Code2Games: Enabling Coding Agents for Gaming World Generation

Code2Games coordinates scene analysis and gameplay planning via shared representations to generate consistent gaming worlds and adapt them to Unreal Engine 5. The framework improves visual quality, interactive fidelity, and playable-game quality over direct coding-agent generation on the GameCode4D

Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao and 1 more

Published Oct 4, 2026 · ▲ 7 on Hugging Face · Code ★ 28

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
83%Must read
?Must readVote to see the score

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

DiVeR improves verifier-guided VLA test-time scaling by reweighting learning toward decision-critical states using action representation dispersion, boosting success without extra annotations or overhead.

Seongheon Park, Heecheol Kim, Shulin Tian, Lilika Makabe and 4 more

Published Oct 4, 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
80%Must read
?Must readVote to see the score

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

SearchJev is a fast calibrated System-1 model that scores search decisions directly without autoregressive generation, improving decision quality, speed, and calibration over same-size language models.

Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu and 5 more

Published Oct 4, 2026 · ▲ 22 on Hugging Face · Code ★ 13

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read
?Must readVote to see the score

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Self-generated feedback in long-horizon test-time training causes weight updates that improve synthetic text but degrade real-text prediction, and settlement on independent evidence prevents this failure.

Cheng Luo, Bing Li, Bernard Ghanem

Published Oct 4, 2026 · ▲ 25 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

MemAdapter counters memory-induced sycophancy via counterfactual induction, context-aware reflection, and evidence-based reasoning to improve memory reliability across diverse scenarios.

Ruqing Ning, Haibo Meng, Zhishang Xiang, Zerui Chen and 3 more

Published Oct 4, 2026 · ▲ 54 on Hugging Face · Code ★ 27

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Prism introduces dynamic sparse attention via adaptive macro-zone block shapes guided by visual variance and cross-modal attention for 2K joint video-audio generation, yielding 2.5x training speedup and improved quality.

Shuyuan Tu, Qi Tian, Yinming Huang, Yue Wu and 7 more

Published Oct 4, 2026 · ▲ 15 on Hugging Face · Code ★ 78

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
75%Highly rated
?Highly ratedVote to see the score

Large-scale analysis of AlphaFold structures reveals organism-specific physicochemical signatures

Large-scale AlphaFold structure analysis reveals organism-specific physicochemical signatures reconstructed via DE-STRESS metrics across 48 proteomes and PDB structures.

Michael J. Stam

Published Oct 4, 2026 · 0 citations

100% Readers1 of 1 upvoted
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

LoHi trades per-frame resolution for denser temporal sampling via low-resolution streams plus sparse high-resolution frames, boosting long-video accuracy up to 10.6 points and cutting front-end latency up to 7x.

Sixun Dong, Wei Li, Andong Deng, Qi Qian and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Oct 3, 2026 · ▲ 18 on Hugging Face · Code ★ 1

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
89%Must read
?Must readVote to see the score

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

Self-evolving reasoning models collapse due to invalid and repeated self-generated questions; R-Quest uses validity and novelty feedback to sustain gains across ten rounds and outperform R-Zero by 17.32 points.

Jinyuan Li, Chengsong Huang, Langlin Huang, Donghong Cai and 3 more

Published Oct 3, 2026 · 0 citations · ▲ 64 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

NAMVIS replaces diffusion with geometry-conditioned next-scale autoregression for fast, parallel multi-view synthesis that outperforms diffusion baselines in quality and speed.

Ramil Khafizov, Ilya Statsenko, Ruslan Rakhimov, Artem Komarichev and 2 more

Published Oct 3, 2026 · 0 citations · ▲ 16 on Hugging Face · Code ★ 5

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution

ConEx bridges saliency maps with concept reasoning via automatic concept discovery to generate faithful, human-interpretable visual explanations.

Yehonatan Elisha, Oren Barkan, Ziv Weiss Haddad, Noam Koenigstein

Published Oct 3, 2026 · 0 citations · ▲ 2 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

DiffGate gates on-policy teacher guidance by trajectory failure and group difficulty to combine dense token-level updates with outcome-level GRPO rewards, improving student pass rates.

Karn Tiwari, Varnith Chordia, Prathosh A P

Published Oct 3, 2026 · ▲ 15 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Learning Discriminative Geometry for Drifting Models

Drifting models suffer from poor pixel-space performance because representation geometry controls KDE sample weighting; persistent representation learning learns discriminative geometry from raw pixels, cutting FID by 82, 95% without pretrained encoders.

Doudou Zhang, Wenwen Hou, Yilin Chen, Qi Chen

Published Oct 3, 2026 · ▲ 5 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

DistScene: Object-to-Scene Distillation for 3D Scene Generation

DistScene generates compositional 3D scenes from single images via object-to-scene distillation, improving spatial coherence by modeling environments as explicit components with shared coordinate frames.

Kunming Luo, Hongyu Yan, Ken Deng, Chengcheng Zhou and 5 more

Published Oct 3, 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

PerturbBot breaks vision-language-action shortcut priors via perturbative training while GroundingFscore diagnoses shortcut reliance, enabling healthier scaling without altering inference.

Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang and 3 more

Published Oct 3, 2026 · ▲ 16 on Hugging Face · Code ★ 7

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
88%Must read

PaLoRA: Paced Low-Rank Adaptation for Continual Learning

PaLoRA derives an optimal rank-aware pacing law for LoRA continual learning that adaptively restricts gradient scaling to prevent forgetting, improving long-horizon benchmark accuracy by 4%.

Yuxuan Li, Fanhu Zeng, Hao Tang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published Oct 3, 2026 · ▲ 10 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Agentic discovery of blood biomarker from distilled private health records

Distilling private health records into a released scoring tool enables agentic discovery of CBC biomarkers that improve diagnostic AUC over literature baselines without exposing patient data.

Seffi Cohen, Liat Antwarg Friedman, Amir Anisman, Ruth Johnson and 6 more

Published Oct 3, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
91%Must read
?Must readVote to see the score

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

CurveCodec 2 predicts quantized skeletal curves from past values and learns residual entropy models to reduce animation storage to 0.22-0.37x of ACL with verified error bounds.

Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 6

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 4/5
88%Must read
?Must readVote to see the score

UnAct: Gradient-Free Unlearning via Targeted Activation Intervention

UnAct uses gradient-free targeted activation interventions to unlearn model classes from few forget images without gradients, labels, or retained data, matching or exceeding SSD and LFSSD accuracy across datasets and preventing network collapse with scarce data.

Saeed Abdul Muizz, Aayat Rafiq, Iqra Altaf Gillani, Janibul Bashir

Published Oct 3, 2026 · ▲ 5 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

The Numerical Linear Algebra of Large Language Models

This survey explains large language model core concepts to numerical analysts and highlights key numerical linear algebra contributions to LLM techniques.

Abdelkader Baggag, Yousef Saad

Published Oct 3, 2026 · ▲ 7 on Hugging Face · Code

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

LMBuild evaluates LLM agents on generating buildable, functional 3D structures and finds physical operability and functional affordance remain challenging despite improved soundness.

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang and 8 more

Published Oct 3, 2026 · ▲ 37 on Hugging Face · Code ★ 5

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

ALoDLM: Adaptively Looped Diffusion Language Models

ALoDLM applies token-adaptive latent recurrence to diffusion language models, allocating computation by difficulty to close the quality gap with autoregressive models at 1.7B and 8B scales.

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao and 9 more

Published Oct 3, 2026 · ▲ 72 on Hugging Face · Code ★ 24

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
70%Highly rated
?Highly ratedVote to see the score

Retrieval-Centric Deep Learning in Growing Nonparametric Neural Networks

Replacing linear attention with RBF and softmax kernels yields principled retrieval-centric deep learning for growing neural networks that store key-value pairs per training point, improving image classification and linking to fixed-network optimizers via MesaNet and DeltaNet.

Maximilian Schlegel, Rajai Nasser, Seijin Kobayashi, Yanick Schimpf and 5 more

Published Oct 2, 2026 · 0 citations · ▲ 2 on Hugging Face · Code ★ 3

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models

FEM-ASM separates storage, execution, and neural coordination via document states, executable skills, and residual operators, yielding incomplete 75% token reconstruction and mixed arithmetic results without competitive general capability.

А. Бочков

Published Oct 2, 2026 · 0 citations

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

92%Must read

World Action Learning via Interaction-Centric Spectral Latent Guidance

WING distills interaction-centric latent actions from egocentric video and uses cross-embodiment spectral low-frequency guidance to transfer them to robot policies, achieving high success rates on LIBERO, RoboTwin, RoboCasa, and real-world tasks.

Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 27 on Hugging Face

100% Readers1 of 1 upvoted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

FrugalEvo pairs expensive LLM strategy exploration with cheap LLM implementation and caching to maximize optimization gain per cost under a budget, outperforming baselines on 10 tasks at significantly lower expense.

Hui Chen, Xuan Qi, James Zhao, Zhaopeng Feng and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 27 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Collective Bias Mitigation via Model Routing and Collaboration

Collective Bias Mitigation routes queries among diverse LLMs and fosters collaboration to substantially reduce bias over single-model baselines.

Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 19 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Multilingual GSM-Symbolic introduces matched math problems across 15 languages to show model size, resource level, reasoning, and typology determine cross-lingual transfer, with size and reasoning closing low-resource gaps but not typological ones.

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart and 21 more

Published Oct 2, 2026 · 0 citations · ▲ 50 on Hugging Face · Code ★ 4

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ProAR: Learning Prospective Reasoning with Autoregressive Video Models

ProAR introduces goal-frame prediction and future self-alignment to enable goal-directed reasoning in autoregressive video models, surpassing baselines with 25% training steps.

Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen

Published Oct 2, 2026 · 0 citations · ▲ 31 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

PointWAM forecasts 3D point trajectories of scenes and hands in a shared space-time frame to guide dexterous robot manipulation, improving DexJoCo success by 56.9 points with video pre-training.

Chunghyun Park, Beomjun Kim, Seungcheol Park, 권희승 and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 59 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

FastOPD: On-Policy Distillation for Lightweight VLA Deployment

FastOPD distills large vision-language-action models via on-policy flow-map distillation with self-consistency, achieving 84% of teacher performance in two steps with 78.1% lower latency.

Yoojin Oh, Jeongsol Kim, Yeonwoo Seo, Jangho Park and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

LLM agents show strong source preferences across search domains that can override item quality, though supplying missing information or countering preconceptions reduces this bias.

Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 44 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Recursive Self-Rewrite uses diverse harnesses and recursive revision to rewrite successful terminal trajectories for supervised fine-tuning, boosting pass@3 by up to 7.6x on hard benchmarks.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 102 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

World Embedding Benchmark

World Embedding Benchmark evaluates physical encoding via 8,000 simulation cases, finding alignment trades off against quantitative recoverability and retrieval improves video generation fidelity.

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 52 on Hugging Face · Code ★ 6

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot-SD self-distills masked diffusion language models by supervising high-impact commitment tokens via information-gain selection, improving reasoning with minimal data.

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 60 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

MetaRubric fixes vacuous rubric credit via evidence-aware optimization and counterfactual rubric adaptation, improving PubMedQA accuracy by up to 20.40 points over static-judge GRPO.

Yuxuan Fan, Jaehong Yoon

Published Oct 2, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 2

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

4DCodeBench benchmarks agents reconstructing dynamic scenes from video as executable graphics code, finding strong static models fail at complex dynamics.

Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang and 5 more

Published Oct 2, 2026 · 0 citations · ▲ 29 on Hugging Face · Code ★ 108

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

DuoMatching improves few-step video generation by jointly matching frame distributions and adding frame-level supervision via an image teacher, boosting visual quality and semantic alignment over 80%.

Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 79 on Hugging Face · Code ★ 46

100% Readers1 of 1 upvoted
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated

Native Action-Prior Learning from Videos for World Action Models

NAVA-WAM pretrains robot action policies directly from observation-only videos via flow-matching and joint attention, improving control accuracy and label efficiency.

Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost and 9 more

Published Oct 2, 2026 · 0 citations · ▲ 87 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Language Models that Play Chess and Explain Their Moves

Queen, a 4B-parameter chess-language model, plays at grandmaster level and explains moves via cross-attention to a silent expert encoder and iterative Bellman-style explanation distillation, surpassing larger frontier models.

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Published Oct 2, 2026 · 0 citations · ▲ 36 on Hugging Face · Code ★ 46

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

In-Distribution Forcing for Long Video Generation at Test Time

In-Distribution Forcing prevents out-of-distribution key-value drift via self-caching to extend short video models to minute-scale generation.

Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 47 on Hugging Face · Code ★ 5

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

WebFovea hardens vision-based web agent harnesses across four stages, raising hidden-set scores from 31.0 to 57.0 by fixing click mismatches, silent actions, and token contamination rather than model reasoning.

Jiangang Han

Published Oct 2, 2026 · 0 citations · ▲ 17 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

On-policy distillation improves language models via implicit teacher rewards without expanding capabilities, but unreliable preferences cause reward hacking into repetitive outputs, mitigated by masking bad responses and SFT initialization.

Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 27 on Hugging Face · Code ★ 7

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Periscope: Extending Frozen Language Models Beyond Their Context Window

Periscope arranges text chunks in a grid to build an evidence map via local and strided probes, letting frozen language models answer questions across multi-million-token contexts with sublinear cost and small GPU memory.

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi and 4 more

Published Oct 2, 2026 · ▲ 9 on Hugging Face · Code ★ 1

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
80%Must read
?Must readVote to see the score

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

Rubric-privileged on-policy distillation before reinforcement learning improves open-ended task scores and reduces reward hacking versus supervised fine-tuning baselines.

Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 8 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Fixed token codes can replace trainable input embeddings in 1.7B-scale language models, removing 100.7M parameters while preserving substantial capabilities without requiring token-specific vectors.

A. Bochkov

Published Oct 2, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
89%Must read
?Must readVote to see the score

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

EyeRobot 2.0 uses active gaze and fixation-relative frames to enable precise bimanual manipulation with only a single stereo camera, outperforming passive stereo by 40% in real-world trials and doubling ego-plus-wrist success under occlusion.

Kush Hari, Justin H. Kerr, Nidhya Shivakumar, Samarth Mahapatra and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

Skill2Real learns simulation skills via a shared robot API with a Proposer-Verifier-Governor loop, transferring frozen skill hierarchies to real robots without fine-tuning to reach 78.75% real-world manipulation success.

Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 17 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Efficient reasoning training differs in impact: faithfulness usually drops due to inconsistency, but monitorability remains robust.

Samuel Lewis-Lim, Xingwei Tan, Mario Sänger, Zhixue Zhao and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 15 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation focuses agent distillation on capabilities beyond the fixed harness via action preferences and validity checks, improving long-horizon agent performance without task rewards.

Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

Published Oct 2, 2026 · 0 citations · ▲ 8 on Hugging Face · Code ★ 2

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

HyperBrowseComp introduces a multilingual, multimodal web-browsing benchmark of 423 hard questions requiring obscure evidence discovery, and current agents perform poorly against human baselines.

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han and 13 more

Published Oct 2, 2026 · ▲ 59 on Hugging Face

100% Readers1 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 1/5
86%Must read
?Must readVote to see the score

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

ADSD links numerical diagnosis to reusable solver self-improvement, reducing mean solver error by nearly 71x across four challenging numerical domains.

Peter Chen, Wotao Yin

Published Oct 2, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
75%Highly rated
?Highly ratedVote to see the score

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

SEER uses self-evolving event reasoning and retrieval to dynamically optimize forecasting with exogenous events via reflective memory and causal knowledge, outperforming state-of-the-art baselines.

Mingtian Tan, Palash Goyal, Mihir Parmar, Sarkar Snigdha Sarathi Das and 5 more

Published Oct 2, 2026 · ▲ 29 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
83%Must read
?Must readVote to see the score
arXivPrivacy

What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

Split-language-model gradients boost token recovery to 97.38% and document reconstruction to 37.77%, so split traffic requires per-token and per-document leakage reporting.

Georgios Politis, Evangelos Pappas

Published Oct 2, 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
88%Must read
?Must readVote to see the score

Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction

SHIFT predicts harness utility via a local LLM to search multi-agent structures per query, achieving ~80% mean accuracy across benchmarks while reducing execution tokens by 32%.

Som Sagar, Shasha Li, Hejie Cui, Ransalu Senanayake and 1 more

Published Oct 2, 2026 · ▲ 12 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
80%Must read

Learning Latent Protein Languages for Autoregressive Generation

Learned latent protein languages improve autoregressive generation scaling, speed, and quality versus amino-acid and coordinate token models.

Mahdi Pourmirzaei, Farzaneh Esmaili, Amir Ziashahabi, Mohammadreza Pourmirzaei and 1 more

Published Oct 2, 2026 · ▲ 6 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
91%Must read

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT scores text-to-image alignment via expected agreement between image and caption answers, boosting negation accuracy to 88% and exceeding fine-tuned evaluators on human correlation benchmarks.

Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Foresight: planning future perception in streaming VLMs without retraining

FORESIGHT uses dual-stream anticipatory planning in frozen streaming VLMs to dynamically configure future perception, improving online benchmarks by up to 18.7 points without retraining.

Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Self-Supervised Scaling of Terminal Environments for Scientific Domains

Software-in-the-loop reconstruction scales terminal-agent training by deriving verified tasks from existing scientific workflows without manual references, improving Terminal-Bench 2 performance to 53.56%.

Zhongzhi Li, Yucheng Shi, Zongxia Li, Junyao Yang and 7 more

Published Oct 2, 2026 · 0 citations · ▲ 12 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

OmniConfess mitigates omni-modal hallucinations by producing token-level confessions of evidential dependence to correct unsupported commitments across text, image, audio, and video.

Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou and 3 more

Published Oct 2, 2026 · 0 citations · ▲ 7 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

COSMI: COmpositional Synthesis of Multi-object Interactions

COSMI synthesizes multi-object interactions by composing local single-object clips, yielding a 222k-sequence dataset and a diffusion model that generalizes to unseen object-interaction pairs with higher contact accuracy.

Daniel Eskandar, Ilya A. Petrov, Gerard Pons‐Moll

Published Oct 2, 2026 · 0 citations · ▲ 8 on Hugging Face · Code ★ 18

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance

Inherit-MAS evolves multi-agent workflows via inheritance and selective execution reuse, boosting benchmark scores while cutting token usage up to 35%.

Songtao Wei, Yi Li, Zhichun Guo, Bingzhe Li

Published Oct 1, 2026 · 0 citations · ▲ 1 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

AutoGUIWorld uses image generators to synthesize GUI interaction trajectories without running software, improving OSWorld scores to 40.8% and ScienceBoard success to 32.2%.

Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang and 17 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face · Code ★ 10

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

PROWBench: Do Video Models Render What the Program Specifies?

PROWBench evaluates video models' fidelity to program-specified world events via replayable world records and VLM-based logic-render and interaction metrics.

Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu and 5 more

Published Oct 1, 2026 · 0 citations · ▲ 72 on Hugging Face · Code ★ 34

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

InterEvolve evolves reward programs at test time via an LLM agent and numerical optimizer to compose a humanoid controller's skills for novel loco-manipulation tasks. The approach releases latent controller competence through object-aware forward-backward models, generating novel strategies that tra

Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 61 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Local Support Learning

Local Support Learning pairs weight adapters with GMM gating to keep updates local, resolving catastrophic forgetting in LLMs up to 7B parameters without prior data.

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Published Oct 1, 2026 · 0 citations · ▲ 27 on Hugging Face · Code ★ 15

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

CANOPY adaptively compresses multimodal evidence via hierarchical region scoring and targeted retrieval, improving QA accuracy while reducing input tokens by up to 27.7%.

Hyojeong Yun, Jueun Kim, Wook-Shin Han

Published Oct 1, 2026 · 0 citations · ▲ 23 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation integrates RL teacher gradients via loss averaging, Adam smoothing, and BF16 rounding, with averaging rules significantly altering math accuracy outcomes.

Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 20 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Hierarchical Continuous Diffusion Language Models

HC-DLM couples discrete token generation with a continuous latent trajectory via a unified variational denoising objective, outperforming diffusion baselines on Sudoku, Countdown, and language modeling.

Hui Ren, Zihan Li, Chang Liu, Huidong Liu and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 90 on Hugging Face · Code ★ 58

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

LOOM stabilizes looped MoE via bounded residual updates and per-loop routers to scale loops to 9, 12, cutting perplexity from 9.62 to 7.77.

Di He, Pengxiang Li, Da Chang, Qingyan Meng and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 9

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer unifies streaming video perception, memory, and proactive response via shared generation, achieving top results on eight benchmarks with a 4B model.

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu and 20 more

Published Oct 1, 2026 · 0 citations · ▲ 233 on Hugging Face · Code ★ 181

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

World Observer jointly generates actor and panoramic observer views to continuously model out-of-view dynamics via shared geometric warping and observer sinks.

Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 87 on Hugging Face · Code ★ 31

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read
?Must readVote to see the score

Sharpening Tax in Post-Training

Post-training sharpens base model behaviors at the cost of solution coverage, introducing a quantifiable "Sharpening Tax"; a posterior-tempered group sampler reduces this tax while boosting accuracy.

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov and 6 more

Published Oct 1, 2026 · 0 citations · ▲ 104 on Hugging Face · Code ★ 25

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

PoS maintains explicit belief states for long-horizon LLM agents, detects belief trapping, and recovers to achieve top results across four benchmarks.

Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 93 on Hugging Face · Code ★ 28

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Spatial Memory Intelligence introduces understanding-driven atomic operations for spatial-memory management in long-video world models, improving sparsity, stability, and spatial consistency.

Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 20

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

World Action Modeling with Progressive Visual Planning

ProWAM predicts sparse visual sub-goals and actions via progressive planning, achieving state-of-the-art long-horizon robotic control and strong zero-shot real-world generalization.

Fei Zhang, Zhaochong An, Duncan Frost, Yikai Wang and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 95 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Latent-MOPD distills multi-teacher LLM specialists via hidden-state and prediction-level on-policy supervision, outperforming token-only and representation-only baselines across math, code, and logic benchmarks.

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 69 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

RealCompanion benchmarks AI companions on longitudinal real-world chats, finding needed past messages are usually recent, memory detectors fail on real messages, and persona reconstruction costs vary 31-fold at equal F1.

Arman Behnam, Sunglyoung Kim, Liangwei Yang

Published Oct 1, 2026 · 0 citations · ▲ 275 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness turns fixed base LLMs into agentic verifiers with workspaces and evidence tools, achieving top selection scores and 6.2, 6.4 point gains over single rollouts on long-horizon tasks.

Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 60

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero benchmarks long-horizon part-level 3D editing via sequential instructions and target images, finding agentic code-based methods preserve unedited regions better but are slower than non-agentic regeneration.

Ruihan Yu, Yu-Ju Tsai, Muyao Niu, Runyi Li and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 17

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems

SciUtopia is a closed-loop LLM simulation framework modeling entire academic ecosystems; it finds resubmission amplifies reviewer burden, cautious exploration balances impact and diversity, and inequality can emerge without cumulative funding advantage.

Yiqiao Jin, Yiyang Wang, Lucheng Fu, Bing He and 7 more

Published Oct 1, 2026 · 0 citations · ▲ 44 on Hugging Face · Code ★ 46

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search

StepCAD combines an LLM CAD policy with geometry-guided search to recover executable CAD programs from 3D meshes, achieving up to 87.2% relative IoU gains over baselines.

G.H. Nehme, Faez Ahmed

Published Oct 1, 2026 · 0 citations · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

ViGAR factorizes manipulation into a visual subgoal planner and executor sharing world-model representations, achieving 82% success on compositional tasks and enabling in-context behavior changes without parameter updates.

Shukai Gong, Xuanran Zhai, Yintianrun Zhang, Ruopeng Cui and 15 more

Published Oct 1, 2026 · 0 citations · ▲ 4 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

UniWAM: Unified World-Action Model

UniWAM unifies physical reasoning, world generation, and action prediction to achieve state-of-the-art embodied performance and log-linear co-training scaling.

Wenxuan Song, Jiayi Chen, Jingbo Wang, Shuai Zhou and 12 more

Published Oct 1, 2026 · 0 citations · ▲ 48 on Hugging Face · Code ★ 91

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

VIEScore2 unifies synthetic image evaluation via grid-based joint score and defect localization predictions, outperforming zero-shot VLMs in correlation and spatial accuracy.

Xianda Du, Max Ku, Weiming Ren, Zhi Rui Tam and 4 more

Published Oct 1, 2026 · ▲ 26 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
88%Must read
?Must readVote to see the score

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

DMAD recasts distribution matching as adversarial distillation with discriminator heads to eliminate auxiliary score models, achieving state-of-the-art few-step image, video, and audio-video generation.

Zhengming Yu, 袁俊坤, Haotian Yang, Gordon Guocheng Qian and 7 more

Published Oct 1, 2026 · 0 citations · ▲ 6 on Hugging Face · Code ★ 121

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

81%Must read
?Must readVote to see the score

Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise

POTER uses optimal transport geometry between training and reference distributions to downweight mislabeled or shortcut-aligned samples, achieving state-of-the-art worst-group accuracy with a single training stage.

Sung Ho Jo, Seonghwi Kim, Wonsang Yun, Minwoo Chae

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published Oct 1, 2026 · 0 citations

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

Kinematic MeanFlow decouples MeanFlow's time derivative via a kinematic identity to stabilize one-step robotic action generation, cutting latency by up to 74% while outperforming multi-step flow matching.

Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao

Published Oct 1, 2026 · 0 citations · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Cross-lingual MoE router alignment improves multilingual LLM performance by aligning router outputs across languages instead of hidden states.

Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 2 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Scaling and Distilling Text Embeddings for Better Diffusibility

Scaling and distilling text embeddings improves latent diffusion by yielding more connected, diffusible spaces that boost generative performance beyond autoregressive baselines.

Zekai Zhang, Yunjie Tian, Yanjin He, Xiaoyan Zhang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 57 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

Biomedical retrieval encoders adapt to typed decision models via SBERT2S1, with retrieval pretraining helping residual heads but not cross-heads, and cross-entropy outperforming RLCD by 2.5, 3.0 points after fixing biased reward normalization.

Pritam Deka

Published Oct 1, 2026 · 0 citations · ▲ 11 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Octrees as an Explicit 3D Language

OctLLM represents 3D geometry via sparse octree tokens and trains separate 3D branches to achieve state-of-the-art multimodal 3D generation and understanding without degrading language ability.

Ran Dan, Si‐Tong Wei, Pengfei Xiong, Wei Zhang and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 15 on Hugging Face · Code ★ 17

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Selection-based Structured Reasoning replaces open-ended reasoning with selection among reusable candidates, cutting per-turn latency over 90% while matching leading small-model search agents' success rates.

Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang and 8 more

Published Oct 1, 2026 · 0 citations · ▲ 16 on Hugging Face · Code ★ 7

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

AGO is an evidence-first quality gate for RAG that treats judge errors and missing data as explicit outcomes, using stratified beta-binomial gates and mandatory meta-evaluation to reduce unsafe promotion to 22.2%-35.1% versus 29.3%-41.8% for naive gates.

Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero, Giuseppe Santoro and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 19 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

GUI-HARVEST optimizes executable harnesses for frozen GUI agents by aligning visual effects, comparing task runs, and consolidating failure patterns into reusable source edits, improving OSWorld-Verified by up to 12.33 points.

Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 12 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

World Editing: Intervening on Executable Worlds at Increasing Depth

World editing intervenes on executable environments at increasing depth via IGMWorld and IGMBench, where top agents achieve 78.2% task success with reliability declining by depth.

Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh and 14 more

Published Oct 1, 2026 · ▲ 29 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 3/5
80%Must read
?Must readVote to see the score

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler automates curriculum learning for agent harness optimization via non-stationary bandits that adapt training scenarios to evolving failure patterns, boosting Pass@1 by 4.4, 7.5 points.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han and 7 more

Published Oct 1, 2026 · ▲ 82 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score

MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations

MEA is a multi-agent framework that uses reward-driven optimization to generate faithful natural-language explanations across tabular, text, and vision data, outperforming baselines by up to 34%.

Yuyang Cheng, R. Ravi, Srivarshinee Sridhar, Sriparna Saha and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 7 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

The AI Theorist reveals excitonic structure in $α$-RuCl$_3$

AI Theorist autonomously develops a first-principles model identifying distinct excitonic states with contrasting selection rules in α-RuCl3 optical spectra.

Hongjian Zhou, Xianfan Nie, Sean Wu, Tarun Patel and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 5 on Hugging Face

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies

OpenRUA gives off-the-shelf coding agents only ROS 2 terminal access to act as zero-shot visuomotor policies, achieving 99% on CaP-Bench and 87% on LIBERO-PRO without bespoke harnesses or training, and reveals emergent perception and closed-loop control behaviors.

Zhaoyang Chu, Earl T. Barr, Claire Le Goues, Peter W. O’Hearn and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 6 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

DeskForge generates dense desktop supervision via controllable real-app environments, yielding 1.2M observations that improve GUI grounding and long-horizon computer-use task completion.

A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 15 on Hugging Face · Code ★ 7

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Labels Override Definitions in Jev-Style Typed Decision Models

Typed decision models exhibit option-label bias because prompts prepend labels to definitions, letting label semantics override rules; removing labels or altering formatting fixes it.

Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram

Published Oct 1, 2026 · 0 citations · ▲ 7 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

LLM2Jev extracts calibrated Jev-style decisions from LLM token probabilities via training-free inference or tree-factorized fine-tuning, showing strong 4B models already match specialized decision models while fine-tuning mainly helps weaker backbones and specific tasks without degrading generation.

Yinheng Li, Justin Wagle

Published Oct 1, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

SimuVerity benchmarks text-to-executable Simulink generation across engineering domains, finding best agents score only 42.86 and structural similarity poorly predicts engineering performance.

Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo and 8 more

Published Oct 1, 2026 · ▲ 53 on Hugging Face · Code ★ 19

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
80%Highly rated
?Highly ratedVote to see the score
ICML 2026WorkshopOffline RL

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

Decision Titan applies test-time training to offline RL, enabling long-term dependencies 20x beyond context windows and 1.7x length generalization while revealing time embeddings and encoding as critical factors.

Jude Waide, Robert Lieck

Published Oct 1, 2026 · 0 citations

100% Readers1 of 1 upvoted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Have an LLM Write Your Anomaly Detector: Autonomous Discovery of Compact, Interpretable Detectors for Time Series

An LLM autonomously searches for short NumPy anomaly detectors that lead the TSB-AD benchmark using spectral features and covariance-aware distances without neural networks or GPUs.

David Berghaus

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

65%Worth a look
?Worth a lookVote to see the score

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

OMAF proposes a one-step flow policy framework for online multi-agent reinforcement learning, achieving up to 3.4x higher returns and 10.5x sample efficiency over baselines.

Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du and 1 more

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

SmoothOperator: Enhancing Representations for Fine-grained Open-set Recognition via Modulated Label Smoothing

SmoothOperator modulates per-sample label smoothing via embedding prominence to reduce over-alignment, boosting open-set recognition AUROC by up to 4.7%.

Thiru Thillai Nadarasar Bahavan, Yu Xia, Sachith Seneviratne, Halgamuge Saman

Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

73%Highly rated
?Highly ratedVote to see the score

Structure-agnostic Causal Representation Learning

SaCRL jointly identifies causal structure and learns invariant representations via soft optimization over HSIC-based invariance violations without prior structural knowledge. It guarantees structure identification, invariance satisfaction, and out-of-distribution generalization while achieving state

Arman Behnam, Binghui Wang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published Oct 1, 2026 · 0 citations

0% Readers0 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

ATPO uses adaptive Tversky reinforcement learning for controllable multi-label video safety detection, raising Jaccard Index to 75.44 on SafeWatch-Bench-Real while enabling steerable precision-recall trade-offs.

Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

GLoC-EHR: Evidence-Cited Clinical Reasoning over Global Context and Local EHR Events

GLoC-EHR uses global and local EHR memories for evidence-cited clinical reasoning, achieving top MIMIC-IV macro AUROC while reducing unsupported citations.

Chaiho Shin, Kwangsoo Kim

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Does Scaling Reinforcement Learning Really Require More Training?

SURGE extracts stronger policies from fixed RL histories via spectral fusion of checkpoints, exceeding native training-curve accuracy without extra training or inference cost.

Bangji Yang, Jiajun Fan, MA Hongba, Ruihan Guo and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

A new benchmark shows LLMs lack structural mathematical understanding with discovery as the key bottleneck, and a primitive-guided self-distillation framework repairs reasoning to boost performance.

Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin and 6 more

Published Oct 1, 2026 · 0 citations · ▲ 14 on Hugging Face · Code ★ 9

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

BDA treats multi-LLM council moves as observations of a per-agent reliability model to yield calibrated answer probabilities and robustly handle persistent adversaries without extra LLM calls.

Ionel Eduard Stan, Paolo Napoletano

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Let the Heads Talk: Beyond Diagonal Graph Attention

Top-A learns edge-conditioned off-diagonal cross-head routes in multi-head attention that preserve diagonal paths, improving interaction-dependent tasks without benefiting heterophily.

Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

AURAL uses adaptive latent reasoning with joint chunk prediction to match chain-of-thought performance while cutting first-token latency 11.8x versus explicit reasoning.

Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu and 11 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

Pretrained ASR backbones enable training-free wake-word detection, and PCA-based pruning retains performance at 50% encoder reduction.

Hwayeon Kim, Youngwon Choi, Hyeonyu Kim

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices

FedSAP uses budget-constrained tri-state channel allocation to partition models into global, private, and dropped channels for heterogeneous federated domain generalization, improving accuracy by up to 4.92 points under 80% pruning.

Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Post-training multi-token prediction heads on ~2.5B chain-of-thought tokens match pretraining speedups with 10^3-10^4x less data, while chain-aware verification and adaptive head selection boost throughput up to 16%.

Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne and 7 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

Human audit of 165 WebArena-Lite tasks recovers 5.45, 8.49% evaluator-missed successes, reveals trajectory errors like looping, and shows guide text and MASM improve results.

Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

Sentence-specificity predictors disagree across technical corpora, and ranking LLM revisions improves selection only for certain models and predictors.

Rocker D’Antonio, Thomas Benton Townsend, Dimitrios Michael Manias

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations

This survey unifies weight, prompt, and embedding adaptation methods for large language models into one taxonomy, analyzing trade-offs and cross-paradigm relationships.

Jungwon Park, Changin Choi, Jimyeong Kim, Nojun Kwak and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Ego2Act evaluates egocentric video generation on multi-step goal-directed manipulation, showing models skip steps and fail at fine-grained physical dynamics.

Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 33 on Hugging Face · Code ★ 4

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Cog-VADU reformulates video anomaly detection as sequential cognitive reasoning via recurrent chain-of-thought prompting and cross-modal re-ranking for training-free zero-shot performance.

Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

LLM novelty judges are unstable: small prompt changes alter verdicts on over half of identical idea pairs and shift accuracy by over 50 points, undermining automated ideation evaluations.

Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

FAULT turns self-diagnosed errors into step-level credit via terminal outcome anchoring and evidence-checked cost learning, recovering 95% signal coverage on ALFWorld and improving long-horizon agentic RL.

Yihua Zhu, Qianying Liu, Weixu Qiao, Xuan Ren and 9 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization

FedFit reduces federated LLM fine-tuning overhead via vector-bank adapter parameterization and quantization, resolving LoRA aggregation conflicts to achieve up to 100x compression with comparable perplexity.

Hang Zou, Chao Zhang, Yuzhi Yang, Yu Tian and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Autoregressive Drillhole Modelling Under Distribution Shift

DrillBench benchmarks autoregressive drillhole modeling, finding lithology-sequence models transfer more robustly than spatial methods, and combining pretraining with retrieval improves cross-province generalization.

Yihao Ding, Daniel Yitian Su, Yiran Zhang, Christopher M. Gonzalez and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ReSolve: Reusing Candidate Reasoning through Selective Generative Moderation

ReSolve reuses candidate reasoning via selective generative moderation to boost math accuracy and cut token use versus voting and self-consistency.

Bangji Yang, Jiajun Fan, MA Hongba, Xi Zhu and 5 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints

Importance-guided allocation directs limited $\ell_1$ perturbation budgets toward model-sensitive regions via fixed clean-gradient priors, boosting attack success by 2.52, 17.70 points across ten robust configurations without increasing global consumption.

Melika Shirian, Kianoosh Vadaei

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

SALD: Self-Referenced Advantage Learning for Diffusion Models

SALD self-references diffusion training via dual noise-level error differences and spectral residuals to improve generation without teachers or extra parameters.

Aryan Das, Surjo Dey, Koushik Biswas, Swalpa Kumar Roy and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Iterative Policy Refinement through Semantic Rollout Analysis

A closed-loop framework iteratively refines structured imitation-learning policies via LLM analysis of rollout tables, improving performance by up to 15% and cutting compute 75%.

Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning

FSG-RL connects subproblem graphs with Python code and multi-verifier feedback to improve math reasoning, raising final-answer accuracy from 43.25% to 67.50% over supervised fine-tuning.

Zihan Liu, Xurong Xie

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning

QK-Wanda couples query and key pruning scores via cross-projection deletion costs, reducing QK reconstruction error by 60% at 50% sparsity and improving downstream perplexity on some large models.

Ivan Ilin, Peter Richtárik

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

BanglaDial-Abuse introduces 1,000 synthetic abusive Bangla sentences across four regional dialects for four-class dialect identification, achieving 0.37, 0.56 lexical Jaccard similarity with distinct lexical spaces.

Hasin Almas Sifat

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score
arXivDeep RL

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

The Homomorphic Advantage Operator stabilizes FHE-based reinforcement learning by centering TD targets to eliminate Bellman drift, achieving zero approximation-bound breaches and 18-point accuracy gains without extra multiplicative depth.

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Capturing In-Context Learning Dynamics with Task Operators

Task Operator captures ICL as stable per-task affine attention transformations, enabling efficient zero-shot replay that nearly matches in-context performance and scales beyond context limits.

Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Yo-ByT5 is a byte-level Yorùbá diacritic restoration model matching mT5-base accuracy with half the parameters and superior text fidelity.

Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

MCIR-M quantifies unique predictive information beyond dependent neighbors via a normalized conditional ratio in [0,1], yielding stable dependence-aware global rankings under multicollinearity and near-duplicates.

Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Auditing Routing Entropy as an Uncertainty Signal in Attention-Residual Transformers

Routing entropy in attention-residual transformers fails as a reliability signal beyond model confidence, with no robust calibration gains and low sensitivity to injected effects.

Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

SCOPE-AD: Sequential cost-aware ordinal-belief planning with energy-based models for diagnostic agents

SCOPE-AD sequentially selects diagnostic tests via ordinal-belief planning and energy-based policies, achieving 77.70% ADNI macro-F1 at $50.46 average cost versus far pricier full-modality evaluation.

Ziwen Yu, Ivan Koychev, Elizabeth Coulthard, Ting Zhou and 6 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Lingtai: What Concept Geometry Reveals—and Does Not Reveal—About LLM Inference

Lingtai introduces a training-free concept telemetry layer that reveals inference-time uncertainty-linked activity and execution-specific trajectory structures in LLMs without tracking correctness.

Jiangang Chen

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

iADD analyzes diffusion policy optimization to show only-latter-timestep updates harm diversity, then proposes incremental Feynman-Kac training that improves alignment-diversity tradeoffs across tasks.

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

From Knowledge Access to Source Learning: Developing Source-Specific Competence

SourceLearn develops reusable source-specific competence via persistent source models and dual learning mechanisms, outperforming retrieval and memory baselines by up to 22.6 points.

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 12 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

75%Highly rated
?Highly ratedVote to see the score

Learning to Predict Distributions over Weight Updates for Test-Time Adaptation

Query-conditioned hypernetworks predict distributions over LoRA weight updates from input queries, enabling test-time scaling via sampled adapted models that outperform deterministic and token-sampling baselines.

Azal Ahmad Khan, Keshav Ramji, Tahira Naseem, Ali Anwar and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

CARM prevents opposing token-level probability changes from canceling in sequence-level masking by using absolute log-ratios, improving RL reasoning and code benchmarks over geometric-mean masking.

Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

Clock Diffusion introduces semi-autoregressive continuous diffusion language models with position-dependent noise schedules, efficient training and sampling, and Cache Grab acceleration to achieve state-of-the-art diffusion likelihoods and competitive reasoning performance.

Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky and 6 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Code Owns the Simulation, Jev Owns the Evaluation

Judgment models excel at evaluation but fail at simulation, yet pairing them with code simulation yields expert control.

Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan and 3 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Understanding Issues, Causes and Solutions in Open-Source LLM-based Multi-Agent Systems

Open-source LLM multi-agent systems face orchestration and execution issues mostly caused by workflow, tool integration, and memory problems, primarily solved by workflow optimization.

Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li, Peng Liang and 2 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

XAI Evaluation Cards: A Practical Method for Designing Human-Centred XAI Evaluations

XAI Evaluation Cards provide a card-sorting method to systematically design human-centered evaluations of explainable AI systems across disciplines.

Kristýna Sirka Kacafírková, Ivania Donoso-Guzmán, Denis Parra, Katrien Verbert and 1 more

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

A typed settlement contract audits concurrent LLM agent actions, showing joint policies complete 59% more six-agent doorway tasks than conservative rejection, with exact replay of 156 checkpoints and rejection of 1,332 corruptions.

Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Calibration-risk routing for controlled world-model adaptation

MC-WM partitions target data to select lower-calibration-risk world models and weights imagined policy updates via learned confidence, evaluated across 541 MuJoCo shift executions.

Yifan F. Zhang, Liang Zheng

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images

DreamBooth personalization on synthetic images degrades fidelity via CFG-inflated high-frequency guidance; ReGain corrects this by frequency-band scaling at sampling to recover 51-64% fidelity without real photos.

Shubhang Bhatnagar, Ishan Bhatnagar, Viraj Shah, Narendra Ahuja

Published Sep 30, 2026 · 0 citations

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

OSWorld-Science benchmarks VLM agents on 146 expert scientific software tasks, showing state-of-the-art models still struggle with scientific workflows and harness design.

Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li and 27 more

Published Sep 30, 2026 · 0 citations · ▲ 63 on Hugging Face · Code ★ 5

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision JEV models show ordinal scale-utilization bias, compressing decisions to 26, 76% of gold support despite high accuracy, but BA-LoRA post-training improves utilization to 86%.

Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 60 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

DARA corrects batch-level reward imbalance via inverse-square-root active-group density weights, accelerating multi-reward RL training by up to 65% with no objective change.

Tong Zheng, Skylar Zhai, Zhan Cheng, Tianming Sha and 6 more

Published Sep 30, 2026 · 0 citations · ▲ 64 on Hugging Face · Code ★ 2

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Frontier models follow unreliable external guidance; training improves selective reliance, identifying it as a key agent reliability dimension.

Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao and 1 more

Published Sep 30, 2026 · 0 citations · ▲ 68 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

Rhetorical robustness requires stable judgments across content-preserving rewrites and discrimination across papers; SciCore improves both via dual-branch science-core review.

Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen and 3 more

Published Sep 30, 2026 · 0 citations · ▲ 76 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

EgoTools introduces a dataset and benchmark for egocentric tool-use reasoning, showing current models struggle with visual grounding while training improves performance.

Shulin Tian, Junsu Kim, Shuai Liu, Hao Li and 16 more

Published Sep 30, 2026 · 0 citations · ▲ 70 on Hugging Face · Code ★ 10

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

RSIGame uses recursive self-improvement with local and global loops to autonomously refine generated games, surpassing one-shot GPT-5.5 scores while cutting generation tokens by 11x.

Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou and 9 more

Published Sep 30, 2026 · 0 citations · ▲ 91 on Hugging Face · Code ★ 128

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Self-evolving search agents suffer co-cheating where proposers and solvers mutually reinforce errors; CrossFit partitions sources to cross-fit agreement and cuts false agreement by over half, boosting downstream search by 8+ points.

Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin and 11 more

Published Sep 30, 2026 · 0 citations · ▲ 675 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

Show 20 more papers