Good Papers

Trending at NeurIPS 2026

Orals, spotlights and posters people are talking about

All sessions

Must read

This year's highest-rated papers

See all

Most debated

Where the reviewers can't agree

See all

All papers

How scores work
78%Highly rated
?Highly ratedVote to see the score

Agta hunter-gatherer oral microbiomes are shaped by contact network structure

Agta hunter-gatherer oral microbiomes resemble Central African foragers more than neighbors, with contact networks predicting bacterial transmission and central individuals as supersharers.

Federico Musciotto, Begoña Dobón, Michael John Greenacre, Álex Mira and 12 more

Published Dec 31, 2030 · 0 citations

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

Attacca trains visual goal-conditioned policies on complete search-to-interact trajectories with decoupled goal images and behavioral-phase conditioning to improve long-horizon embodied task success by up to 7x.

Gyusik Seo, Jaehong Yoon

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code ★ 3

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

UNREAL: Unifying Retrieval and Long-Context with a Single Model

UNREAL unifies retrieval and long-context evidence selection via frozen LLM representations with minimal parameters, outperforming state-of-the-art retrievers and improving long-context accuracy substantially.

Edan Kinderman, Elad Hoffer, Yochai Blau, Brian Chmiel and 3 more

Published Oct 6, 2026 · ▲ 15 on Hugging Face

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Sensor-Language-Action Models

Sensor-Language-Action modeling unifies multimodal sensors, language, and actions via a semantic interface, and OpenSLA achieves superior hierarchical prediction and explanation with zero-shot generalization.

Yuekai Xu, Zitao Shuai, Yuzhe Yang

Published Oct 6, 2026 · ▲ 1 on Hugging Face · Code ★ 1

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

90%Must read
?Must readVote to see the score

World Models' Last Exam in Physics

World Models' Last Exam in Physics benchmarks video models via 40 measurement-based physics tasks, finding the best model scores 57.76/100 with widespread inconsistencies.

Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin and 7 more

Published Oct 6, 2026 · ▲ 2 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine co-evolves policies, curricula, and judges via adaptive diagnosis and coactive calibration to sustain embodied reasoning self-improvement. Experiments on driving and navigation show continuous gains in both policy and judge capability.

Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao and 9 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

CtrlCache accelerates interactive video world models via control-aware caching that detects action changes to reuse transformer residuals and apply frequency-mixed history guidance, achieving up to 1.41x speedups with improved quality.

Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh and 1 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face · Code

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

CheckerBench evaluates long-horizon agents on synthesizing static-analysis checkers across 300 CVE-derived tasks, finding best Pass@1 reaches 45.33%.

Hang He, Li Wang, Hao Chen, Yuchen Shao and 8 more

Published Oct 6, 2026 · ▲ 52 on Hugging Face · Code

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

82%Must read
?Must readVote to see the score

Sherpa: Teaching LLMs to Teach Adaptively

Sherpa uses multi-turn reinforcement learning to train LLM teachers that adapt instructions to diverse student archetypes, improving student performance by 20.5 points and pedagogy scores to 79.2%.

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen and 1 more

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

100% Readers1 of 1 upvoted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

TRACE aligns FP4 quantization between RL training and rollout paths via rollout-guided quantization-aware training for MoE language models, achieving BF16-comparable RL performance with up to 5.4x rollout speedup.

Xin Wang, Hao Yu, Zhengyang Zhuge, Bochao Mao and 8 more

Published Oct 6, 2026 · ▲ 15 on Hugging Face

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

90%Must read
?Must readVote to see the score

From Evidence to Action: How Tool-Using Agents Fail

Tool-using agents often fail by acting before establishing required evidence or leaving multi-action workflow prerequisites unresolved, despite accurate static action assessment. SafeActBench reveals failures stem from how agents use established evidence during execution, not just missing informatio

Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai and 2 more

Published Oct 6, 2026 · ▲ 20 on Hugging Face · Code ★ 3

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Cross-tokenizer on-policy distillation achieves comparable accuracy with strict top-16 shared-vocabulary supervision versus full coverage, while expanded span supervision reduces accuracy due to conflicting gradients, motivating prioritization of supervision reliability over alignment coverage.

Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li and 3 more

Published Oct 6, 2026 · ▲ 40 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS bootstraps reusable agent memory from self-generated practice tasks without oracles, improving success rates by up to 15.9 points across benchmarks.

Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard and 2 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Building Rome from a Single Image

A redesigned object-centric generator partitions scenes into distance-adaptive chunks, captures 2D-3D correspondence, and trains on 4,000 outdoor scenes to outperform baselines in indoor and outdoor mesh generation.

Jiraphon Yenphraphai, Fang Li, Tianshuo Xu, Depu Meng and 4 more

Published Oct 6, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Towards In-Parameter Memory Augmentation for Large Language Models

This survey organizes in-parameter memory augmentation for LLMs by parameter placement and acquisition time to enable reusable parametric knowledge at deployment.

Haoyu Huang, Zhongwei Xie, Jiaxin Bai, Yisen Gao and 5 more

Published Oct 6, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Personal-Agent Mediated Recommendation with Cross-Platform User History

Personal-agent mediated recommendation balances cross-platform user history against platform rankings via the MediateRec benchmark and PAMO optimization to improve rescue-harm trade-offs.

Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley and 1 more

Published Oct 6, 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

AdvSim2Real co-evolves tasks, adversarial injections, and a web agent in a simulator, boosting 4B agent completion by 33.6% against unseen adaptive attacks and transferring gains to real browsers.

Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov and 2 more

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

NeMo-DCR bit-exactly refits trillion-parameter policies by streaming delta-compressed weight changes via affine mappings and XOR masks, cutting 1T cross-region refits from 87.5 minutes to 150 seconds.

Songlin Jiang, Zhiyu Li, Terry Kong, Yu Yao and 4 more

Published Oct 6, 2026 · ▲ 6 on Hugging Face · Code ★ 2,048

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation

EmbodiedSmith unifies asset, scene, and task generation in a recursive self-improvement loop to scale embodied simulation data, improving generation success and robot policy generalization across diverse embodiments and physics.

Yikai Qin, Yifei Deng, Mingjian Liang, Wenxuan Song and 12 more

Published Oct 6, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

HiPLEX factorizes full-duplex speech policies into timing and content controllers to jointly optimize interaction dynamics via reinforcement learning. It lowers takeover rates, reduces interruption latency, and improves human-like turn timing versus GRPO.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution

Speculative tool execution predicts tool calls from partial ASR to run them during speech, cutting median voice-agent response latency from 5.79 s to 4.60 s.

Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park and 4 more

Published Oct 6, 2026 · ▲ 1 on Hugging Face

100% Readers1 of 1 upvoted
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

ALIVE inserts objects that interact with video contents via first-frame editing and interaction guidance, outperforming baselines on interaction and insertion benchmarks.

Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

Dev-Primitives turn repository artifacts into active, resident-LLM agents with self-modification interfaces, and HERMES improves software engineering benchmarks by 12.4% over baselines while cutting inference costs by 26.2%.

Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang

Published Oct 6, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco

Turba provides an open-source Moroccan fertilizer recommendation stack with programmatic access, versioned data, and loadable crop-specific machine learning surrogates for reproducible benchmarking.

Abdelghani Belgaid, Zakaria Mahmoud, Fahd Chibani, Oumnia Ennaji and 2 more

Published Oct 5, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

58%Worth a look
?Worth a lookVote to see the score

Empirical Variational Autoencoder

Empirical Variational Autoencoder learns autoregressive latent priors empirically via one linear layer to close the VAE prior-posterior gap, yielding high-fidelity sequential generation competitive with diffusion models at much faster inference.

Kaede Shiohara

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code ★ 4

0% Readers0 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Execution-Aligned Progressive Noise for Consistent Asynchronous Replanning in Generative Robot Policies

Execution-Aligned Progressive Noise structures noise across chunks and time to maintain consistent generative states during asynchronous replanning, achieving up to 96.7% real-robot success.

Di Wu, Ping Liu, Xuhua Chen, He Zheng and 2 more

Published Oct 5, 2026

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

Stepped MoE: Segment-Level Routing with Configurable Inference Complexity

Stepped MoE unifies elastic architectures and sparse gating to adapt model capacity to deployment constraints and input requirements, outperforming dense counterparts by 2-5%.

Arnav Kundu, Zhaoyang Xu, Bairu Hou, Chang Gao and 2 more

Published Oct 5, 2026 · ▲ 1 on Hugging Face

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated

RealtimeWAM: One-Step Asynchronous World Action Models

RealtimeWAM uses teacher-anchored consistency distillation and cross-expert wavefront pipelining for one-step asynchronous action generation, achieving near-lossless performance with ~25x speedup.

Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong and 6 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 2,880

0% Readers0 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

OnePO enables RL-only medical domain adaptation via adaptive objective evolution and teacher retirement, yielding HuatuoGPT-3 that surpasses frontier models.

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu and 6 more

Published Oct 5, 2026 · ▲ 15 on Hugging Face · Code ★ 12

100% Readers1 of 1 upvoted
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026RL for LLMs

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

RGPO adaptively uses ground-truth rationales as temporary scaffolds to generate improved responses for on-policy RL, then transfers only higher-reward model outputs back, reducing reward sparsity and improving text and multimodal reasoning.

Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde and 2 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

WildMatch adapts pretrained image matchers to wildlife identification using only identity labels, improving retrieval accuracy and learning transferable matching priors without keypoint annotations.

Turhan Can Kargin, Piotr Kubaty, Ekaterina Rostovskaya, Izabela Wierzbowska and 2 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

S2PD switches from autoregressive to parallel diffusion during denoising to enforce physical and logical consistency with faster sampling than fully serial methods.

Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari

Published Oct 5, 2026 · 0 citations · ▲ 1 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

MEND: RL For Flow Models via Proximal Velocity Matching

MEND uses proximal velocity matching to cap rewards and accept only cost-effective sample moves, outperforming prior flow-model RL methods in far fewer updates without KL penalties or reference models.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli and 1 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Conditional Trajectory Peaks: Single-Pass Multimodal Policies over Action Chunks

Conditional Trajectory Peaks predicts multimodal action-chunk candidates in a single pass, achieving 97.25% LIBERO success and 3× faster inference while preserving diverse behaviors.

Di Wu, Rongtian Shen, Ping Liu, Xuhua Chen and 3 more

Published Oct 5, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

Learning to Read the Contextual Tokens in Diffusion Transformers

A framework maps diffusion transformer contextual tokens through a frozen LLM to reveal they encode global emerging scene semantics early, inspiring contextual alignment that improves generation quality.

Omer Dahary, Etai Sella, Hadar Averbuch-Elor, Daniel Cohen-Or and 1 more

Published Oct 5, 2026 · ▲ 5 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Agents judge failing retrieval results useless but rarely stop; enforcing answers after five useless judgments improves success and fixes stopping.

Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu and 3 more

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 3/5
89%Must read
?Must readVote to see the score

JLD: Perceptual Distance Through A Jacobian Lens

JLD defines a perceptual image distance via a Jacobian-derived metric tensor from frozen vision encoders, achieving state-of-the-art correlation with human judgments and resolution robustness.

Shreshth Saini, Balu Adsumilli, Alan C. Bovik

Published Oct 5, 2026 · ▲ 3 on Hugging Face · Code ★ 3

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
86%Must read
?Must readVote to see the score

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Hybrid Linear Attention introduces query-dependent chunk-level routing for Gated DeltaNet, improving long-context benchmarks by up to 5.57 points via adaptive recurrent memory composition.

Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He and 2 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

MiniCorp: The Last Mile of the AI Agent Firm

MiniCorp is a simulated office environment that generates longitudinal, counterfactual enterprise data to study autonomous AI-run companies and train adaptive agents.

Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin and 5 more

Published Oct 5, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
66%Highly rated

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

APO learns shared LLM aligner initializations via grouped gradient coordination to enable few-shot personalization under heterogeneous preferences and scarce feedback, improving over baselines with 20 local examples.

Liyan Yang, Yige Yuan, Zhiqin Yang

Published Oct 5, 2026 · ▲ 5 on Hugging Face

0% Readers0 of 1 upvoted
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
80%Highly rated
?Highly ratedVote to see the score

InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation

InterMimicGen retargets human motion capture to humanoids and self-evolves tracking data via iterative simulation-verified augmentation to scale dexterous loco-manipulation.

Yucheng Zhang, Sirui Xu, Jinhong Li, Liuyu Bian and 7 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

100% Readers1 of 1 upvoted
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

69%Highly rated
?Highly ratedVote to see the score

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

ACG-Bench evaluates arm-wise compositional generalization in dual-arm vision-language-action models via AE-VLA, which achieves 21.53% simulated and 39% real-world success versus under 6% baselines.

Zaibin Zhang, Binghao Ran, Yuhan Wu, Zhongbo Zhang and 9 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face

0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Agentic retrieval combining LLM reasoning with dense retrieval improves nDCG@10 by 8.7 points over standard retrieval but requires 107 seconds and 764K input tokens per query.

Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira and 7 more

Published Oct 5, 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

On-policy power distillation trains models to generate sharpened answers directly, improving single-sample math reasoning by up to 27.3 points and outperforming multi-candidate sampling and reward-based methods.

Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa and 3 more

Published Oct 5, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models

RACE predicts subskill transition timing to extend VLA action chunks reliably, reducing stop-and-go idle time ~5x on real robots while improving success rates.

Seonghoon Yu, Dongwon Kim, HyungRok Jung, Yoonjae Baek and 3 more

Published Oct 5, 2026 · ▲ 13 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
89%Must read

Certification of Real Images through Calibrated Content Authentication

Deepfake detectors degrade to 76% accuracy and near-zero under attacks, so calibrated reconstruction-based authentication bounds false real-image certification to 1%.

Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi and 1 more

Published Oct 5, 2026 · ▲ 13 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

Base Models Can Reason By Taking a Cue From Training Data

Fixing initial token cues in base models boosts reasoning to match RL performance, with effects traced to training data associations that can be causally edited.

Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat and 2 more

Published Oct 5, 2026 · ▲ 12 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

HLA-WM combines geometry-guided retrieval with recurrent linear attention to fix long-range forgetting in video world models, improving 60-second consistency metrics by up to 28.5% with 12× lower memory and no retraining.

Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu and 3 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

LoGRA reduces LLM reinforcement learning memory by up to 45.7% via low-rank gradient sketches and predicted-KL step control, enabling 27B-parameter training on single nodes.

Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li and 4 more

Published Oct 5, 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Closing the Context Gap: Activation Alignment for Tabular In-Context Learning

Activation alignment trains a linear map to align partial-context student activations with full-context teacher activations, significantly improving tabular in-context learning efficiency and recovering much of the performance gap.

Yoel Zeldes

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
ICLR 2027Tabular data

Adapting prior-data fitted networks for tabular anomaly detection

Frozen and fine-tuned TabPFN representations for tabular anomaly detection yield ZEN and FOCUS, surpassing all ADBench baselines in AUROC despite unsupervised deployment and contaminated reference sets.

Maximilian Bershtman, Niv Cohen

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Learning to Learn a Language

Prior-Fitted Language Model, trained solely on synthetic non-linguistic data, learns to infer and predict real languages from context with frozen weights, achieving strong cross-lingual compression and reasoning without ever seeing real text.

Lennart Carstens-Behrens, Holger Fröhlich

Published Oct 5, 2026 · ▲ 7 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

SoK: Semantic Decision Engines in Network Control Loops

Systematizing 139 semantic decision engine families reveals most miss network deadlines and verification, with only four reporting deadline attainment; unverified decisions reverse admission verdicts under queued execution, prompting minimum reporting rules and a research agenda.

Delong Li, Chen Li, Xu Wang, Haochen Gong and 2 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Targeted bias injection via closed-loop activation steering exploits diffusion language model denoising trajectories to steer frozen models toward adversarial demographic answers with minimal corruption.

Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed, Muhammad Umer Sheikh and 2 more

Published Oct 5, 2026 · ▲ 14 on Hugging Face · Code ★ 3

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models

FairRSFM benchmarks remote sensing foundation models by biome to expose hidden ecological performance disparities and tests debiasing methods without backbone updates. Aggregate metrics consistently mask large biome-dependent gaps, though mitigation effectiveness varies by model and task.

Md Aminur Hossain, Omkumar Vaghasiya, Rajeev Ranjan Dwivedi, Vinod Kurmi and 1 more

Published Oct 5, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

What Matters for Latent Reasoning with Flow Matching

FLaRe uses flow matching for latent reasoning that is useful, diverse, explainable, refinable and efficient, reaching 97% of explicit chain-of-thought accuracy at 25% latency.

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

Published Oct 5, 2026 · ▲ 10 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Capability-Driven Self-Evolution of Agent Memory

PrisMem drives agent memory self-evolution via capability-specific guidance, dependency-aware selection, and trace-guided integration, outperforming baselines by up to 10.54 points on million-token benchmarks.

Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu and 7 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Representation-Space MMD for Diffusion Language Models

Post-training minimizes representation-space MMD between diffusion language model outputs and references via retained token features, improving perplexity, accuracy, and parallel decoding.

Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev and 6 more

Published Oct 5, 2026 · ▲ 24 on Hugging Face · Code ★ 12

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models use fixed-point convergence to enable truncated training, shared KV caches, faster prefill, and faster RL updates, while a learned depth prior and orthogonal input injection improve perplexity across scales.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen and 3 more

Published Oct 5, 2026 · ▲ 20 on Hugging Face · Code ★ 35

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents

AutoSciBench autonomously generates and iteratively adapts scientific agent benchmarks via recipes and concepts, reducing solver accuracy by over 22 points versus human benchmarks while improving quality ratings.

Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards and 6 more

Published Oct 4, 2026 · ▲ 17 on Hugging Face

100% Readers1 of 1 upvoted
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

EVISKILL: Grounding Skill Evolution in Replayable Evidence

EVISKILL grounds LLM skill evolution in replayable evidence cards linking edits to supporting contexts, using targeted replay for verification and global validation for incorporation.

Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang and 1 more

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 8

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
88%Must read
?Must readVote to see the score

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

Feasible-future decoding reranks VLA actions by future safe-completion mass, reducing cumulative safety costs by up to 57.5% without retraining or rollouts.

Tu Nguyen, Matthieu Zimmer, Vu Anh Vu, Ziyi Wang and 3 more

Published Oct 4, 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

LiFT: Loop Flow Transformers

Loop Flow Transformers loop a shared diffusion transformer with depth-indexed regression targets, improving generation with more inference compute and fewer parameters than dense models.

Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent and 1 more

Published Oct 4, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
91%Must read
?Must readVote to see the score

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

UndoBench separates tool-using agent competence from fault recovery via paired enterprise workflow trials, finding 83.54% nominal success but only 46.72% recovery success with phase-dependent vulnerabilities.

Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah

Published Oct 4, 2026 · ▲ 11 on Hugging Face · Code ★ 1

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
76%Highly rated
?Highly ratedVote to see the score

RobotUse: Allocating Computation, Context, and Decisions

RobotUse organizes robot computation, context, and decisions around revisable physical actions via visual target selection and persistent playbooks, achieving 45% RoboLab success and real-world learning.

Junhoo Lee, Injun Baek, Seungyeon Kim, Suhyun Jeon and 3 more

Published Oct 4, 2026 · ▲ 14 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
83%Must read
?Must readVote to see the score

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

ASCENT online test-time trains agents by self-distilling verified deployment trajectories into LoRA weights via a frozen hindsight model, improving long-horizon success and efficiency without external teachers or memory retrieval.

Haodong Lu, Dong Gong

Published Oct 4, 2026 · ▲ 17 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
71%Highly rated

Code2Games: Enabling Coding Agents for Gaming World Generation

Code2Games coordinates scene analysis and gameplay planning via shared representations to generate consistent gaming worlds and adapt them to Unreal Engine 5. The framework improves visual quality, interactive fidelity, and playable-game quality over direct coding-agent generation on the GameCode4D

Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao and 1 more

Published Oct 4, 2026 · ▲ 6 on Hugging Face · Code ★ 4

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
83%Must read
?Must readVote to see the score

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

DiVeR improves verifier-guided VLA test-time scaling by reweighting learning toward decision-critical states using action representation dispersion, boosting success without extra annotations or overhead.

Seongheon Park, Heecheol Kim, Shulin Tian, Lilika Makabe and 4 more

Published Oct 4, 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
80%Must read
?Must readVote to see the score

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

SearchJev is a fast calibrated System-1 model that scores search decisions directly without autoregressive generation, improving decision quality, speed, and calibration over same-size language models.

Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu and 5 more

Published Oct 4, 2026 · ▲ 20 on Hugging Face · Code ★ 8

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
89%Must read
?Must readVote to see the score

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Self-generated feedback in long-horizon test-time training causes weight updates that improve synthetic text but degrade real-text prediction, and settlement on independent evidence prevents this failure.

Cheng Luo, Bing Li, Bernard Ghanem

Published Oct 4, 2026 · ▲ 23 on Hugging Face · Code ★ 1

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy

MemAdapter counters memory-induced sycophancy via counterfactual induction, context-aware reflection, and evidence-based reasoning to improve memory reliability across diverse scenarios.

Ruqing Ning, Haibo Meng, Zhishang Xiang, Zerui Chen and 3 more

Published Oct 4, 2026 · ▲ 40 on Hugging Face · Code ★ 27

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Prism introduces dynamic sparse attention via adaptive macro-zone block shapes guided by visual variance and cross-modal attention for 2K joint video-audio generation, yielding 2.5x training speedup and improved quality.

Shuyuan Tu, Qi Tian, Yinming Huang, Yue Wu and 7 more

Published Oct 4, 2026 · ▲ 8 on Hugging Face · Code ★ 56

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
72%Highly rated

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Kandinsky 6.0 Video introduces diffusion models that generate synchronized 5-second audio-video clips with lip-sync via a dual-stream CrossDiT architecture, with the 29B-parameter Pro version outperforming its predecessor and matching top competitors in speech quality.

Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin and 36 more

Published Oct 4, 2026 · ▲ 132 on Hugging Face · Code ★ 151

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
75%Highly rated
?Highly ratedVote to see the score

Large-scale analysis of AlphaFold structures reveals organism-specific physicochemical signatures

Large-scale AlphaFold structure analysis reveals organism-specific physicochemical signatures reconstructed via DE-STRESS metrics across 48 proteomes and PDB structures.

Michael J. Stam

Published Oct 4, 2026 · 0 citations

100% Readers1 of 1 upvoted
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

LoHi trades per-frame resolution for denser temporal sampling via low-resolution streams plus sparse high-resolution frames, boosting long-video accuracy up to 10.6 points and cutting front-end latency up to 7x.

Sixun Dong, Wei Li, Andong Deng, Qi Qian and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Oct 3, 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
72%Highly rated
?Highly ratedVote to see the score

ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution

ConEx bridges saliency maps with concept reasoning via automatic concept discovery to generate faithful, human-interpretable visual explanations.

Yehonatan Elisha, Oren Barkan, Ziv Weiss Haddad, Noam Koenigstein

Published Oct 3, 2026 · 0 citations · ▲ 1 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

DiffGate gates on-policy teacher guidance by trajectory failure and group difficulty to combine dense token-level updates with outcome-level GRPO rewards, improving student pass rates.

Karn Tiwari, Varnith Chordia, Prathosh A P

Published Oct 3, 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

Learning Discriminative Geometry for Drifting Models

Drifting models suffer from poor pixel-space performance because representation geometry controls KDE sample weighting; persistent representation learning learns discriminative geometry from raw pixels, cutting FID by 82, 95% without pretrained encoders.

Doudou Zhang, Wenwen Hou, Yilin Chen, Qi Chen

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 3

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 2/5
72%Highly rated
?Highly ratedVote to see the score

DistScene: Object-to-Scene Distillation for 3D Scene Generation

DistScene generates compositional 3D scenes from single images via object-to-scene distillation, improving spatial coherence by modeling environments as explicit components with shared coordinate frames.

Kunming Luo, Hongyu Yan, Ken Deng, Chengcheng Zhou and 5 more

Published Oct 3, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

PerturbBot breaks vision-language-action shortcut priors via perturbative training while GroundingFscore diagnoses shortcut reliance, enabling healthier scaling without altering inference.

Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang and 3 more

Published Oct 3, 2026 · ▲ 13 on Hugging Face · Code ★ 6

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
88%Must read

PaLoRA: Paced Low-Rank Adaptation for Continual Learning

PaLoRA derives an optimal rank-aware pacing law for LoRA continual learning that adaptively restricts gradient scaling to prevent forgetting, improving long-horizon benchmark accuracy by 4%.

Yuxuan Li, Fanhu Zeng, Hao Tang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published Oct 3, 2026 · ▲ 9 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Agentic discovery of blood biomarker from distilled private health records

Distilling private health records into a released scoring tool enables agentic discovery of CBC biomarkers that improve diagnostic AUC over literature baselines without exposing patient data.

Seffi Cohen, Liat Antwarg Friedman, Amir Anisman, Ruth Johnson and 6 more

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
91%Must read
?Must readVote to see the score

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

CurveCodec 2 predicts quantized skeletal curves from past values and learns residual entropy models to reduce animation storage to 0.22-0.37x of ACL with verified error bounds.

Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 5

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 4/5
88%Must read
?Must readVote to see the score

UnAct: Gradient-Free Unlearning via Targeted Activation Intervention

UnAct uses gradient-free targeted activation interventions to unlearn model classes from few forget images without gradients, labels, or retained data, matching or exceeding SSD and LFSSD accuracy across datasets and preventing network collapse with scarce data.

Saeed Abdul Muizz, Aayat Rafiq, Iqra Altaf Gillani, Janibul Bashir

Published Oct 3, 2026 · ▲ 4 on Hugging Face · Code ★ 1

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
67%Highly rated
?Highly ratedVote to see the score

The Numerical Linear Algebra of Large Language Models

This survey explains large language model core concepts to numerical analysts and highlights key numerical linear algebra contributions to LLM techniques.

Abdelkader Baggag, Yousef Saad

Published Oct 3, 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

LMBuild evaluates LLM agents on generating buildable, functional 3D structures and finds physical operability and functional affordance remain challenging despite improved soundness.

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang and 8 more

Published Oct 3, 2026 · ▲ 32 on Hugging Face · Code ★ 5

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

ALoDLM: Adaptively Looped Diffusion Language Models

ALoDLM applies token-adaptive latent recurrence to diffusion language models, allocating computation by difficulty to close the quality gap with autoregressive models at 1.7B and 8B scales.

Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao and 9 more

Published Oct 3, 2026 · ▲ 62 on Hugging Face · Code ★ 21

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
88%Must read
?Must readVote to see the score

Periscope: Extending Frozen Language Models Beyond Their Context Window

Periscope arranges text chunks in a grid to build an evidence map via local and strided probes, letting frozen language models answer questions across multi-million-token contexts with sublinear cost and small GPU memory.

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi and 4 more

Published Oct 2, 2026 · ▲ 8 on Hugging Face · Code ★ 1

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
90%Must read
?Must readVote to see the score

World Action Learning via Interaction-Centric Spectral Latent Guidance

WING distills interaction-centric latent actions from egocentric video and uses cross-embodiment spectral low-frequency guidance to transfer them to robot policies, achieving high success rates on LIBERO, RoboTwin, RoboCasa, and real-world tasks.

Zhiming Liu, Yikun Miao, Ying Chen, Hongrui Yin and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 20 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

Rubric-privileged on-policy distillation before reinforcement learning improves open-ended task scores and reduces reward hacking versus supervised fine-tuning baselines.

Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

DuoMatching improves few-step video generation by jointly matching frame distributions and adding frame-level supervision via an image teacher, boosting visual quality and semantic alignment over 80%.

Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 4

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale

Fixed token codes can replace trainable input embeddings in 1.7B-scale language models, removing 100.7M parameters while preserving substantial capabilities without requiring token-specific vectors.

A. Bochkov

Published Oct 2, 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 3/5
medium 4/10
strict 2/5
89%Must read
?Must readVote to see the score

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

EyeRobot 2.0 uses active gaze and fixation-relative frames to enable precise bimanual manipulation with only a single stereo camera, outperforming passive stereo by 40% in real-world trials and doubling ego-plus-wrist success under occlusion.

Kush Hari, Justin H. Kerr, Nidhya Shivakumar, Samarth Mahapatra and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 10 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

Skill2Real learns simulation skills via a shared robot API with a Proposer-Verifier-Governor loop, transferring frozen skill hierarchies to real robots without fine-tuning to reach 78.75% real-world manipulation success.

Xincheng He, Siyu Ma, Chang Yu, Yunuo Chen and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 16 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Efficient reasoning training differs in impact: faithfulness usually drops due to inconsistency, but monitorability remains robust.

Samuel Lewis-Lim, Xingwei Tan, Mario Sänger, Zhixue Zhao and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Collective Bias Mitigation via Model Routing and Collaboration

Collective Bias Mitigation routes queries among diverse LLMs and fosters collaboration to substantially reduce bias over single-model baselines.

Mingzhe Du, Luu Anh Tuan, Xiaobao Wu, Yichong Huang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 14 on Hugging Face

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

Multilingual GSM-Symbolic: What determines capability transfer across languages?

Multilingual GSM-Symbolic introduces matched math problems across 15 languages to show model size, resource level, reasoning, and typology determine cross-lingual transfer, with size and reasoning closing low-resource gaps but not typological ones.

Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart and 21 more

Published Oct 2, 2026 · 0 citations · ▲ 50 on Hugging Face · Code ★ 4

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ProAR: Learning Prospective Reasoning with Autoregressive Video Models

ProAR introduces goal-frame prediction and future self-alignment to enable goal-directed reasoning in autoregressive video models, surpassing baselines with 25% training steps.

Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen

Published Oct 2, 2026 · 0 citations · ▲ 26 on Hugging Face · Code ★ 2

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It

LLM agents show strong source preferences across search domains that can override item quality, though supplying missing information or countering preconceptions reduces this bias.

Jonghyun Song, Haewon Park, Jeonghoon Shim, Woojung Song and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 43 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

PointWAM: 3D World Action Modeling for Dexterous Robotic Manipulation

PointWAM forecasts 3D point trajectories of scenes and hands in a shared space-time frame to guide dexterous robot manipulation, improving DexJoCo success by 56.9 points with video pre-training.

Chunghyun Park, Beomjun Kim, Seungcheol Park, 권희승 and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 55 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Recursive Self-Rewrite uses diverse harnesses and recursive revision to rewrite successful terminal trajectories for supervised fine-tuning, boosting pass@3 by up to 7.6x on hard benchmarks.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 100 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
13/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

84%Must read
?Must readVote to see the score

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

4DCodeBench benchmarks agents reconstructing dynamic scenes from video as executable graphics code, finding strong static models fail at complex dynamics.

Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang and 5 more

Published Oct 2, 2026 · 0 citations · ▲ 28 on Hugging Face · Code ★ 89

100% Readers1 of 1 upvoted
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

MetaRubric fixes vacuous rubric credit via evidence-aware optimization and counterfactual rubric adaptation, improving PubMedQA accuracy by up to 20.40 points over static-judge GRPO.

Yuxuan Fan, Jaehong Yoon

Published Oct 2, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot-SD self-distills masked diffusion language models by supervising high-impact commitment tokens via information-gain selection, improving reasoning with minimal data.

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao and 2 more

Published Oct 2, 2026 · 0 citations · ▲ 58 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

World Embedding Benchmark

World Embedding Benchmark evaluates physical encoding via 8,000 simulation cases, finding alignment trades off against quantitative recoverability and retrieval improves video generation fidelity.

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 52 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

FrugalEvo pairs expensive LLM strategy exploration with cheap LLM implementation and caching to maximize optimization gain per cost under a budget, outperforming baselines on 10 tasks at significantly lower expense.

Hui Chen, Xuan Qi, James Zhao, Zhaopeng Feng and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 5

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

In-Distribution Forcing for Long Video Generation at Test Time

In-Distribution Forcing prevents out-of-distribution key-value drift via self-caching to extend short video models to minute-scale generation.

Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Language Models that Play Chess and Explain Their Moves

Queen, a 4B-parameter chess-language model, plays at grandmaster level and explains moves via cross-attention to a silent expert encoder and iterative Bellman-style explanation distillation, surpassing larger frontier models.

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Published Oct 2, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 40

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read

Native Action-Prior Learning from Videos for World Action Models

NAVA-WAM pretrains robot action policies directly from observation-only videos via flow-matching and joint attention, improving control accuracy and label efficiency.

Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost and 9 more

Published Oct 2, 2026 · 0 citations · ▲ 86 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation focuses agent distillation on capabilities beyond the fixed harness via action preferences and validity checks, improving long-horizon agent performance without task rewards.

Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee

Published Oct 2, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

HyperBrowseComp introduces a multilingual, multimodal web-browsing benchmark of 423 hard questions requiring obscure evidence discovery, and current agents perform poorly against human baselines.

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han and 13 more

Published Oct 2, 2026 · ▲ 58 on Hugging Face

100% Readers1 of 1 upvoted
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 4/5
medium 1/10
strict 1/5
86%Must read
?Must readVote to see the score

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

ADSD links numerical diagnosis to reusable solver self-improvement, reducing mean solver error by nearly 71x across four challenging numerical domains.

Peter Chen, Wotao Yin

Published Oct 2, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

SEER uses self-evolving event reasoning and retrieval to dynamically optimize forecasting with exogenous events via reflective memory and causal knowledge, outperforming state-of-the-art baselines.

Mingtian Tan, Palash Goyal, Mihir Parmar, Sarkar Snigdha Sarathi Das and 5 more

Published Oct 2, 2026 · ▲ 13 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 0/5
83%Must read
?Must readVote to see the score
arXivPrivacy

What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

Split-language-model gradients boost token recovery to 97.38% and document reconstruction to 37.77%, so split traffic requires per-token and per-document leakage reporting.

Georgios Politis, Evangelos Pappas

Published Oct 2, 2026 · ▲ 8 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
88%Must read
?Must readVote to see the score

Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction

SHIFT predicts harness utility via a local LLM to search multi-agent structures per query, achieving ~80% mean accuracy across benchmarks while reducing execution tokens by 32%.

Som Sagar, Shasha Li, Hejie Cui, Ransalu Senanayake and 1 more

Published Oct 2, 2026 · ▲ 10 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
80%Must read

Learning Latent Protein Languages for Autoregressive Generation

Learned latent protein languages improve autoregressive generation scaling, speed, and quality versus amino-acid and coordinate token models.

Mahdi Pourmirzaei, Farzaneh Esmaili, Amir Ziashahabi, Mohammadreza Pourmirzaei and 1 more

Published Oct 2, 2026 · ▲ 5 on Hugging Face · Code

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
91%Must read

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT scores text-to-image alignment via expected agreement between image and caption answers, boosting negation accuracy to 88% and exceeding fine-tuned evaluators on human correlation benchmarks.

Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei and 1 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Foresight: planning future perception in streaming VLMs without retraining

FORESIGHT uses dual-stream anticipatory planning in frozen streaming VLMs to dynamically configure future perception, improving online benchmarks by up to 18.7 points without retraining.

Ashok Prasad Neupane, Dipan Bartaula, Ankit Belbase, Saugat Adhikari and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Self-Supervised Scaling of Terminal Environments for Scientific Domains

Software-in-the-loop reconstruction scales terminal-agent training by deriving verified tasks from existing scientific workflows without manual references, improving Terminal-Bench 2 performance to 53.56%.

Zhongzhi Li, Yucheng Shi, Zongxia Li, Junyao Yang and 7 more

Published Oct 2, 2026 · 0 citations · ▲ 10 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

OmniConfess mitigates omni-modal hallucinations by producing token-level confessions of evidential dependence to correct unsupported commitments across text, image, audio, and video.

Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou and 3 more

Published Oct 2, 2026 · 0 citations · ▲ 6 on Hugging Face · Code ★ 1

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

Show 20 more papers