Good Papers

Showing Deep RL Show all papers

76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

MEND: RL For Flow Models via Proximal Velocity Matching

MEND uses proximal velocity matching to cap rewards and accept only cost-effective sample moves, outperforming prior flow-model RL methods in far fewer updates without KL penalties or reference models.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli and 1 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score
arXivDeep RL

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

The Homomorphic Advantage Operator stabilizes FHE-based reinforcement learning by centering TD targets to eliminate Bellman drift, achieving zero approximation-bound breaches and 18-point accuracy gains without extra multiplicative depth.

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

iADD analyzes diffusion policy optimization to show only-latter-timestep updates harm diversity, then proposes incremental Feynman-Kac training that improves alignment-diversity tradeoffs across tasks.

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli and 1 more

Published Oct 1, 2026 · 0 citations · ▲ 4 on Hugging Face · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Tail-Influence Sampling for CVaR Policy Evaluation

Tail-Influence Sampling allocates evaluation budgets by tail influence to estimate CVaR with oracle variance and lower MSE than rollouts.

Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 23 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026DeepMindDeep RL

Delightful Distributed Policy Gradient

Delightful Policy Gradient gates distributed updates with delight (advantage times surprisal) to suppress high-surprisal failures while preserving rare successes, outperforming importance-weighted methods under staleness, bugs, and corruption.

Ian Osband

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

0% Readers0 of 1 upvoted
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 3/5
medium 10/10
strict 3/5
90%Must read
?Must readVote to see the score

Modeling quantum neural network gradient with reinforcement learning

RLQ-Grad uses reinforcement learning to propose quantum neural network updates without differentiating circuits, avoiding barren plateaus and scaling with parameters rather than Hilbert space dimension to achieve orders-of-magnitude faster training and higher accuracy.

Nhan Luu, Trung D Luu, Ngoc Nam Pham, Thang C Truong

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

A Primal-dual Approach for Semi-Infinitely Constrained Reinforcement Learning

Di Wang, Liangyu Zhang, Haishan Ye, Guang Dai and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026XidianDeep RL

Reinforcement Learning for View-Adaptive Distillation in 3D Gaussian Compression

Hongji Zhao, Mingrui Zhu, Xin Wei, Nannan Wang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

Boyang Li, Matthew Kim, Sylvia Herbert

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling

Nicholas Corrado, Wenyuan Huang, Josiah Hanna

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

ROLLVERIFY: BRIDGING EFFICIENCY AND ACCURACY IN LONG-TAIL ROLLOUT REINFORCEMENT LEARNING

Yongqiang Yao, Jingru Tan, Kaihuan Liang, Zixin Yin and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

When Does Knowing the State Help? Diagnosing Process vs. Outcome Reward Design

Wenpei Shao, Ross Jacobucci

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Interactive Combinatorial Reinforcement Learning for Knowledge Graph Reasoning

Jun Nie, Yonggang Zhang, Tongliang Liu, Chengqi Zhang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026DalhousieDeep RL

Instability of Meta-Learning Intrinsic Rewards for Policy Gradient Reinforcement Learning

Dilith Jayakody, Domenic Rosati, Janarthanan Rajendran

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

LOCU: Löwdin-Orthogonalized Constraint Updates for Multi-Constraint Policy Optimization

Joonyoung Lim, Younghwan Yoo

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Exploring Lifelong Adaptation: In-Context Reinforcement Learning in Non-Stationary Environments

Ye Wang, Kaiqian Cui, Xinrun Xu, Tao Zhang and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Cross-Question Reliable Reinforcement Learning

Hector G. Rodriguez, Marcus Rohrbach

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking During RL

Aria Wong, Joshua Engels, Neel Nanda

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Gradient Routing Localizes and Removes Unintended Behaviors in RL

Jake Ward, Shawn Hu, Aria Wong, Nathan Hu and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

A Refined Sample-Complexity Analysis of Robust Policy Optimization under Decaying Actor Stepsizes

Swetha Ganesh, Vaneet Aggarwal

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Implicit Goal Conditioning via Value Disaggregation

Shashwat Saxena, Mehul Goel, Sreyas Venkataraman, Sarvesh Patil and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

NitroBox: Lightning-Fast Sandbox for Large-Scale RL Training

Yuzhou Nie, Ruilin Zhou, Zhaorun Chen, Jingyang Zhang and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Reward-Estimated Hypergradient for Bilevel Reinforcement Learning with Black-Box Follower

Shigeki Kusaka, Mikoto Kudo, Takumi Tanabe, Akifumi Wachi and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Uncertainty-Guided Reward Labeling for Reinforcement Learning under Limited Feedback

Renhao Zhang, Shreyas Chaudhari, Bruno Silva

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Bootstrapped Bipartite Actor-Critic for Diffusion RL

Tianze Zhu, yinuo Wang, Letian Tao, Tianyi Zhang and 5 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

EDEN: Emergent Dynamics in Evolutionary Neural-networks for Robust Continuous Control

Chi Zhang, Jinge Li, Yifei Wang, Lei Wang and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

CPA: Efficient and Stable FP4 RL Training via Cross-Precision Alignment

Gu Gong, Yining Wei, Yuechen Tao, Tianyuan Wu and 10 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Mitigating Compounding Errors in Online Reinforcement Learning via Optimal Transport Regularized Flow Matching

Boxiang Tao, Lei Guo, Bin Wang, Zexin Wang and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

RL-Inf: Tracking Non-local Training Data Influence for Online Reinforcement Learning

Shixuan Liu, Cheng Tang, Yuzheng Hu, Fan Wu and 2 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

BECON: Belief-Conditioned Constrained Multi-Objective Reinforcement Learning under Drifting Preferences and Budgets

Bui Trong Duc, Huynh Thi Thanh Binh

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

What Kind of Diffusion Models Do We Need in Online Reinforcement Learning?

Zihao Wu, Hongyao Tang, Yi Ma, YAN ZHENG and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Higher-Order Action Supervision Makes A Strong Policy Class

Peng Cheng, Yunxian Hou, Zhi Zhou, Qian Zhang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Towards a Unified Model for Flexible Job Shop Scheduling Problems

Inguk Choi, Woo-Jin Shin, Sang-Hyun Cho, Hyun-Jung Kim

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

A Surrogate Perspective on Convergence of Fixed-Target DQN

Zichu Liu, Nneka M Okolo, Ryan D'Orazio, Danilo Vucetic and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026TulaneDeep RL

Behavior-Discriminative Reward Shaping for Reward-Robust Reinforcement Learning

Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan Zheng

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

H-GenPO: Hierarchical Generative Policy Optimization via the Option-Critic Framework

Wonhyeok Choi, Minwoo Choi, Jaeyeul Kim, Kyumin Hwang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Foundation Pareto Flow Policy for Multi-Objective Reinforcement Learning

Zhanjiang Yang, Lijun Sun, Yueming Li, Meng Li and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Adaptive Scheduling Pipeline For Multi-Instance Asynchronous Reinforcement Learning

Salah Chikhi, Luis H Ruiz, Entong Li, Li Zeng

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

DGRAF: Observation-Quality-Aware Reinforcement Learning for Dynamic Reconfigurable Batteries

Jiasong Chen, Jingwei Hu, Zheng Fang, Zhihong Zhang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Canopy: Tree-Aware Rollout Scheduling for Agent Reinforcement Learning

Feiyuan Zhang, LI Pengbo, Ziniu Li, Yuhao Jiang and 8 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Feasible Policy Optimization for Safe Reinforcement Learning

Yujie Yang, Yuanxu Sun, Wenyu Li, Beiyan Jiang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Preference-Guided Adversarial Policy Optimization for Long-Tail Robust Driving

Tong Nie, Yihong Tang, Junlin He, Yuewen Mei and 4 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Reconciling Operational Energy Trilemma: A Heterogeneous Risk-Constrained MDP Framework with Residual Policy Learning

Yujian Ye, Siqi Qian, Yizhi Wu, Tianxiang Cui and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning Reusable Options by Decomposing Neural Policies

Parnian Behdin, Reza Abdollahzadeh, Kiarash Aghakasiri, Levi Lelis

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

S$^2$-RL: Sample-Set Dual Reinforcement Learning for Generative Semantic Segmentation Dataset Distillation

Haoyu Wang, Fei Zhou, Qingqing Qiu, Lei Zhang and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Grouped Adaptive Head Mixing for Personalized Multi-Task Federated Reinforcement Learning

Yiran Pang, Zhen Ni, Dimitris Pados, Xiangnan Zhong

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Subprocess-Constrained Markov Decision Processes

Jiarui Gan, Debmalya Mandal

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning to Undo: Transfer Reinforcement Learning under State Space Transformations

Mridul Mahajan, Aldo Pacchiano, Xuezhou Zhang

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Distribution-Adaptive Policy Optimization

Yuxiao He, Ziqi Wang, Xingzhou Lou, Xiaoqian Liu and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Iterative Gumbel Planning for Continuous Control

Shaohuai Liu, Weirui Ye, Yilun Du, Le Xie

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Decoupling is the Key: Scaling Deep Value Networks in Reinforcement Leanring

Yunsheng Xue, Ziyi Zhang, zhihao wu, Youfang Lin

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Discovering Programmatic Policies from Reinforcement Learning-Based Traffic Signal Controllers

Lindong Xie, Yang Zhang, Beiyu Song, XING Zeren and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

SAMPPO: Structure-Aware Mirror Proximal Policy Optimization

Corinna Cortes, Mehryar Mohri, Yutao Zhong

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Physics-Informed Optimal Control by Control-Only Supervision with Error Guarantees on Value and Policy

Zihua Wang, Xiaopei Jiao, Yunfeng Cai

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Solving Stochastic Control under Multiplicative and Internal Noise via Constrained Optimization

Ruben Moreno Bote, Francesco Damiani, Dmytro Grytskyy

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning Process Rewards via Visitation Matching for Efficient RL

Raymond Tsao, Andrew Wagenmaker, Sergey Levine

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Prevailing Bisimulation Metric Learning Is Biased: Implicit Regularization and Its Remedy

Junqi Lu, Ruixiang Sun, Xin Li, Gaopeng Peng and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Anatomy of Off-Policy Policy Gradient: Importance Sampling, KL Regularization, and Baselines

Haoqun Cao, Yurun Yuan, Tengyang Xie

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Cost of Mismatch: Noise Amplification in Zeroth-Order Reinforcement Learning

Lianmin Chen, Junbin Qiu, Chenxing Wei, Yao SHU and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Adaptive Robust Estimator for Policy Optimization in Reinforcement Learning

Zhongyi Li, Wan Tian, Jingyu Chen, Kangyao Huang and 7 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

$R^2E$: A Role-driven Reward Evolutionary Framework for Automated Reward Function Design

Shouhao Chang, Xuan Liu, Hongye Zhu, Xinning Chen and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A Constrained Bi-level Optimization Framework for Constrained Preference-Based Reinforcement Learning

Yue Mao, Siyuan Xu, Shicheng Liu, Minghui Zhu

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

DiversePlace: Diversity-Seeking Curriculum Reinforcement Learning for Macro Placement

Wenrui Zhou, Jiashun Liu, Wenji Fang, Zhiyao Xie and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RL-Guided Contraction of Symbolic Tensor Networks for Quantum Circuit Equivalence

Suhaib Al-Rousan, Christian Schilling, Max Tschaikowski, Kim Larsen

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

EpiPivot: Learning to Control the Simplex Method under Epistemic Uncertainty

Guantao Zhao, Mahdi Noorizadegan, Shihao Yang, Nicoleta Serban

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Learning to Search, Searching to Learn: A Closed-Loop Framework for Large-Scale Vehicle Routing

Yongji Fu, Yi Zhou, Gaojie Jin, Guanqun Cao

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Control First, Robustness Next: Decoupled Representation Learning for Visual RL Generalization

heo chanyong, Hyelyn Jeong, Jongchan Park, Seungjun Oh and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Permanent and Transient Representations for Continual Reinforcement Learning

Nishanth Anand, Doina Precup

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Skill-Adaptive Noise Scheduling for Diffusion Policies

Woo Kyung Kim, Gwangpyo Yoo, Eunyoung Park, Honguk Woo

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Effectiveness of Curriculum Learning Depends on Reward Sparsity and Competing Optima

John Vastola, Ann Huang, Satpreet Harcharan Singh, Samuel J Gershman and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

DUDS: Dual-stage Data Selection for Efficient Reinforcement Learning with Verifiable Rewards

Hongling Zheng, Li Shen, Zichuan Lin, Jiafei Lyu and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Provably Efficient Representation Learning for Low-Rank CMDPs

Kaixuan Liu, GUOJUN XIONG, Shengpu Tang, Wanyun Si and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

DSSP: Diffusion State Space Policy with Hierarchical Full-History Conditioning

Zhiyuan Guan, Jianshu Hu, Han Fang, Yunpeng Jiang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Multi-Stage Planning from Single-Stage Data: Reinforcement Learning Helps Composition but Requires Anchoring

Boyuan Zheng, Zidong Liu, Yingyu Liang

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Local Policy Manifolds for Efficient Multi-Objective Reinforcement Learning

Qiyue Xia, Tianwei Wang, J. Michael Herrmann

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Replay-buffer engineering for noise-robust quantum circuit optimization

Akash Kundu, Sebastian Feld

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

C2G-BENCH: A Cyber-Physical Evaluation Benchmark for Hierarchical Reinforcement Learning in Grid-Interactive Hyperscale Data Centers

Vineet Gundecha, Sahand Ghorbanpour, Sifat Abdullah, Mahasweta Chakraborti and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Variance-Optimal State Resampling for Reinforcement Learning with Verifiable Rewards

Yibo Wang, Yuanxin Liu, Zhaoran Wang

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Reinforcement Learning Agents Are Swimmers

Juan Rojas, Jacob Adamczyk, Abhishek Naik, Volodymyr Makarenko and 5 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Sharp Convergence and Sample Complexity of Policy Mirror Descent for Average-Reward MDPs

Enes Arda, Atilla Eryilmaz

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

JAXtari: High-Throughput and Easy-to-Modify Arcade Learning Environment

Quentin Delfosse, Raban Emunds, Paul Seitz, Sebastian Wette and 4 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PDHFormer: Progressive Dual-Head Transformer for Behavioral Choice Prediction

Hao Zhou, Jing Chen, Yaoxin Wu, Jie Gao and 1 more

Paris Poster Session 4, Thu, Dec 10, 5:30 PM–7:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Diffusion-Enhanced GFlowNet for Solving Vehicle Routing Problems

Ni Zhang, Zhiqin Zhang, Ling Pan, Hoong Chuin Lau and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decision-Focused Learning in MDPs: An Occupancy Measure Approach

Zihao Zhao, Ashwath K Karunakaram, Ali Eshragh, Yuexing Li and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

SimpleEvol: Efficient Intelligence Conversion via Less Human Prior in Automated Heuristic Design

Jianghan Zhu, Cong Zhang, Rongjie Zhu, CHI ZHANG and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

When Does Non-Uniform Replay Matter in Reinforcement Learning?

Michal Korniak, Mikołaj Czarnecki, Yarden As, Piotr Miłoś and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Differentially Private Sparse Reward Estimation with Preference Feedback

Meng Ding, Mingxi Lei, Jie Zhang, Jinyan Liu and 1 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Meta-Reinforcement Learning with Zero-Shot Reinforcement Learning

Jake Grigsby, Siddhant Agarwal, Yu Lei, Leonidas Varveropoulos and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026U AlbertaDeep RL

Deep Double Q-learning

Deep Double Q-learning trains two independent Q-functions to decouple selection and evaluation, reducing overestimation and outperforming Double DQN across 47 of 57 Atari games.

Prabhat Nagarajan, Martha White, Marlos C. Machado

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Hadamard Representation: Scaffolding Performance Across Model-free RL

Hadamard Representation replaces hidden layers with element-wise products of two layers, reducing neuron dormancy and increasing effective rank to consistently improve model-free RL performance.

Jacob Eeuwe Kooi, Zhao Yang, Mark Hoogendoorn, Vincent Francois-Lavet

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

PPO in the Fisher-Rao geometry

FR-PPO leverages Fisher-Rao geometry to provide monotonic policy improvement guarantees and sub-linear convergence without dependence on state or action space dimensions.

Razvan-Andrei Lascu, David Siska, Lukasz Szpruch

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

TriSearch: Learning to Optimize Triangulations via Bistellar Flips

TriSearch uses reinforcement learning and circuit-based flip representations to optimize triangulations across dimensions, discovering more Calabi-Yau triangulations than existing samplers.

Yiran Wang, Guido Montufar

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies

CRAFT combines dense counterfactual advantages with grounded residual corrections to reduce variance and bias in closed-loop autonomous driving fine-tuning, achieving strong Bench2Drive gains.

Keyu Chen, Nanfei Ye, Yida Wang, Wenchao Sun and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Reward Inflation: A Healthy Stimulus for Reinforcement Learning

Gradually scaling rewards during training accelerates RL adaptation by upweighting recent transitions, sustaining gradients to prevent neuron dormancy, and improving performance across diverse tasks.

Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

HyCO: A Hybrid Neural Solver for Combinatorial Optimization

HyCO hybridizes RL and diffusion solvers for combinatorial optimization, theoretically and empirically achieving lower regret via adaptive prefix construction and regime-switching.

Yuheng Li, Di Yang, Haipeng Chen, Yanhai Xiong

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

D³PO fixes multi-objective RL via decomposed per-objective updates, late preference weighting, and diversity regularization to recover dense Pareto fronts with one policy.

Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge and 1 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Factorized Spectral Representations for Reinforcement Learning

FaStR applies CP tensor decomposition to transition dynamics via contrastive learning, yielding separate state, action, and next-state encoders that shrink sample complexity and enable cross-actuator transfer.

Junyi Wu, Dan Li

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

Predictive representation learning combined with high-capacity value approximation drives scalable multitask RL, with the simple model-free MR.Q outperforming world-model methods across continuous control tasks.

Johan Obando Ceron, Lu Li, Scott Fujimoto, Pierre-Luc Bacon and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

Diagnosis-driven tension management adapts online RL to deployment-specific prior validity shifts, rejecting universal benchmarks for flexible, evidence-guided optimization.

Guozheng Ma, Lu Li, Zilin Wang, Pierre-Luc Bacon and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Video Models Can Reason with Verifiable Rewards

VideoRLVR applies reinforcement learning with verifiable rewards to video diffusion models, improving rule-consistent visual reasoning and cutting training latency 40% via early-step optimization.

Tinghui Zhu, Sheng Zhang, James Yipeng Huang, Selena Song and 4 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 9 on Hugging Face · Code ★ 28

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

TGPO: Temporal Grounded Policy Optimization for Signal Temporal Logic Tasks

TGPO decomposes signal temporal logic into timed subgoals and invariant constraints for hierarchical reinforcement learning, achieving 31.6% higher success rates than baselines on complex long-horizon robotics tasks.

Yue Meng, Fei Chen, Chuchu Fan

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 3 on Hugging Face · Code ★ 2

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance

FLAG augments states with flow latents to localize sampling and prevent importance weight collapse, enabling scalable high-dimensional expressive MaxEnt-RL policies with state-of-the-art performance.

Sungha Kim, Gawon Lee, Jusuk Lee, Jonghae Park and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

ReFPO: Reflow Regularization for Flow Matching Policy Gradients

ReFPO adds explicit reflow regularization to flow matching policy gradients, stabilizing training and enabling high-fidelity one-step inference that matches multi-step performance across control tasks.

Ge Wang, Yibo Peng, Fan Feng, Shenhao Yan and 9 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities

Long-horizon Q-learning penalizes n-step value bound violations via hinge losses to stabilize bootstrapping and outperform standard TD methods.

Armaan A Abraham, Lucy Xiaoyang Shi, Chelsea Finn

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Equivariant Reinforcement Learning for Clifford Quantum Circuit Synthesis

A qubit-relabeling-equivariant, size-agnostic reinforcement learning agent synthesizes near-optimal Clifford circuits across qubit counts, outperforming Qiskit on large instances.

Richie Yeung, Aleks Kissinger, Rob Cornish

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Embodied Neurocomputation: A Framework for Interfacing Biological Neural Cultures with Scaled Task-Driven Validation

The Embodied Neurocomputation framework optimizes biological neural network encoding for closed-loop navigation, identifying 12 configurations that outperform silicon DQN agents under identical interaction budgets.

Johnson Zhou, Daniel Tanneberg, Forough Habibollahi, Alon Loeffler and 11 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence

Bellman residual minimization for policy optimization achieves stable convergence with function approximation but lacks extensive study; foundational control results are established.

Donghwan Lee, Hyukjun Yang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

The Smart Buildings Control Suite: A Diverse Open Source Benchmark to Evaluate and Scale HVAC Control Policies for Sustainability

The Smart Buildings Control Suite is an open-source HVAC benchmark using multi-year data from 11 buildings and scalable simulators to test control policies across diverse climates and structures.

Judah Goldfeder, Victoria Dean, Zixin Jiang, Xuezheng Wang and 3 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

FASTER: Value-Guided Sampling for Fast RL

FASTER traces test-time sampling gains to earlier denoising stages via a denoising-space MDP that filters action candidates early, reducing compute while improving diffusion-based RL policy performance.

Perry Dong, Alexander Swerdlow, Dorsa Sadigh, Chelsea Finn

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Verifying Neural Networks with Reinforcement Learning

RSB uses reinforcement learning to optimize neural network verification branching heuristics, solving 11% more instances and cutting branch exploration by 50%.

Hai Duong, Thanh Le, ThanhVu Nguyen

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Learning to Cut: Reinforcement Learning for Benders Decomposition

RLBD trains a reinforcement learning policy to adaptively select Benders cuts, substantially improving computational efficiency and generalizing across problem variations.

Haochen Cai, Xian Yu

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces

DGRL enables efficient reinforcement learning in discrete action spaces up to 10^20 via distance-guided exploration and regression-based updates, improving performance by up to 66%.

Heiko Hoppe, Fabian Akkerman, Wouter van Heeswijk, Maximilian Schiffer

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction

ASCPO guarantees high-probability state-wise safety constraints in model-free RL without strong assumptions, significantly outperforming existing methods on continuous robot control tasks.

Weiye Zhao, Feihan Li

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026U VirginiaDeep RL

On the Divergence of Differential Temporal Difference Learning without Local Clocks

Differential temporal difference learning converges with local clocks but can diverge with global clocks in average-reward reinforcement learning.

David Antrobius, Shangtong Zhang

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score
NeurIPS 2026U VirginiaDeep RL

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

Latent Q-Barrier shielding filters actions via predicted future costs to improve safe in-context reinforcement learning reward-safety tradeoffs under out-of-distribution deployment shifts.

Minjae Kwon, Amir Moeini, Shangtong Zhang, Lu Feng

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Spectral Alignment in Forward–Backward Representations via Temporal Abstraction

Temporal abstraction acts as a low-pass filter reducing effective successor representation rank to fix spectral mismatch in forward-backward representations and stabilize long-horizon continuous control learning.

Seyed Mahdi Basiri Azad, Jasper Hoffmann, Iman Nematollahi, Hao Zhu and 2 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization

Unifying GCRL and MISL as control maximization reveals formulation-specific bounds linking diverse pretraining skills to downstream goal sensitivity.

Alireza Modirshanechi, Benjamin Eysenbach, Peter Dayan, Eric Schulz

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing

TOPPO balances critic gradients to fix PPO's multi-task ill-conditioning, outperforming SAC baselines with fewer parameters and steps.

Yuanpeng Li, Rui Miao, Gefei Lin, Annie Qu

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Reward-Conditioned Reinforcement Learning

RCRL conditions agents on reward parameterizations via replay counterfactual rewards, improving sample efficiency and enabling zero-shot behavioral adaptation without extra interaction.

Michal Nauman, Marek Cygan, Pieter Abbeel

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting

Flexible discounting in distributional reinforcement learning captures expressive temporal and risk preferences, fixing existing multi-horizon optimality issues.

Mehrdad Moghimi, Anthony Coache, Hyejin Ku

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026PurdueDeep RL

Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework

Natural policy gradient equals doubly smoothed policy iteration, a Bellman-operator framework achieving global geometric convergence and O((1-γ)^{-1} log ε^{-1}) iteration complexity without modified stepsizes or regularization.

Phalguni Nanda, Zaiwei Chen

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment

TMPO replaces scalar reward maximization with trajectory-level reward distribution matching via Softmax Trajectory Balance, improving diffusion alignment diversity by 9.1% while avoiding reward hacking and mode collapse.

Jiaming Li, Chenyu Zhu, Zhiyuan Ma, Nanxi Yi and 8 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

Reinforcement learning agents adapt Large Hadron Collider trigger thresholds online to maximize signal efficiency while maintaining background rates within tolerance bands, improving in-tolerance intervals by up to 56% on real CMS collision data without fine-tuning.

Zixin Ding, Shaghayegh Emami, Giovanna Salvi, Cecilia Tosciri and 6 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 3

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Behavioral Foundation Models for Quality Diversity

BFM-QD searches a behavioral foundation model's latent space for diverse high-performing policies and outperforms parameter-space quality-diversity methods, especially in sparse and deceptive tasks.

Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Learning to Solve Compositional Geometry Routing Problems

DiCon solves compositional geometry routing via differential attention suppressing weak actions and contrastive learning for robust representations, yielding broad generalization.

Mingfeng Fan, Jianan Zhou, Jiaqi Cheng, Yifeng Zhang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Zero-Shot Instruction Following in RL via Structured LTL Representations

A GNN encodes LTL instructions as Boolean formula sequences conditioning a policy, improving zero-shot multi-event RL instruction following in complex environments.

Mathias Jackermeier, Mattia Giuri, Jacques Cloete, Alessandro Abate

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Understanding Goal Generalisation in Sequential Reinforcement Learning

Salient features drive reinforcement learning goal generalization, early goals persist to affect later ones, and latent policy gradients predict out-of-distribution behavior accurately.

Jason Brown, Edward Young

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

CPPO is an on-policy contrastive RL method deriving advantages from contrastive Q-values via PPO without rewards or replay buffers, outperforming prior CRL baselines in 14 of 18 tasks and matching or exceeding hand-crafted-reward PPO in 12 of 18.

Asim Osman, Sasha Abramowitz, Mark Bergh, Ulrich Armel Mbou Sob and 12 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 3/5
88%Must read
?Must readVote to see the score

Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning

Concave scalarized multi-objective RL suffers biased gradients that cause O(ε⁻⁴) sample complexity; multi-level Monte Carlo NPG achieves optimal O(ε⁻²).

Swetha Ganesh, Jason Chia, Vaneet Aggarwal

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 4/5
80%Must read
?Must readVote to see the score

PRECISE: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models

PRECISE introduces an SDE-consistent stochastic sampler balancing exploration and stability for RL post-training of flow-matching models, enabling faster, more stable reward optimization with significantly reduced training time.

Bo Peng, Tao Huang, Weijie Kong, Junzhe Li and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Generative Actor-Critic with Soft Bridge Policies

SoftGAC proposes a soft generative actor-critic with stochastic bridge policies that expose a tractable MaxEnt objective via single-pass sampling, outperforming diffusion and flow baselines on continuous control with lower latency.

Ke He, Le He, Shunpu Tang, Yafei Wang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 2/5
medium 5/10
strict 0/5
89%Must read
?Must readVote to see the score

Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning

IBD treats an RL agent's actions as randomized interventions and uses per-dimension two-sample tests with FDR correction to identify controllable observation dimensions, matching oracle returns across 12 continuous-control tasks with up to 100 distractors.

Jiaxin Liu, Anzhe Cheng, Paul Bogdan

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026U AlbertaDeep RL

The Laplacian Keyboard: Beyond the Linear Span

Laplacian Keyboard hierarchically combines Laplacian eigenvectors into a behavior library with a meta-policy, exceeding linear span limits for better zero-shot approximation and sample efficiency.

Siddarth Chandrasekar, Marlos C. Machado

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning

NASDAQ normalizes low-dimensional observations to balance dynamics prediction losses and couples value learning with short-term value and next-observation prediction, achieving strong sample efficiency and faster training across diverse domains.

Xinwei Liu, Junyuan Liang, Zicong Hong, Jianting Zhang and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026POSTECHDeep RL

Delayed homomorphic reinforcement learning for environments with delayed feedback

DHRL uses MDP homomorphism to collapse redundant delayed states, reducing sample complexity and improving actor-critic performance on continuous control.

Jongsoo Lee, Jangwon Kim, Soohee Han

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 3/5
medium 5/10
strict 0/5
72%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026XiamenDeep RL

AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning

AlphaPareto uses LLM-guided multi-objective reinforcement learning to discover formulaic trading alphas that adapt to evolving pools and outperform competitors on real-world data.

Yingbo Zhao, Zeyu Yang, Zhoufan Zhu

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control

ParetoSlider trains one diffusion model with continuous preference weights to approximate the full Pareto front, enabling inference-time navigation of trade-offs between conflicting generative goals without retraining.

Shelly Golan, Michael Finkelson, Ariel Bereslavsky, Yotam Nitzan and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 15 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning

Low-rank adaptation regularizes critic learning by constraining updates to low-dimensional subspaces via frozen base weights, reducing loss and improving off-policy RL performance.

Yuan Zhuang, Yuexin Bian, Sihong He, Jie Feng and 6 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks

SATR constrains gradient-free RSNN updates via signal-adaptive KL trust regions, improving stability and matching PPO-LSTM returns on continuous control benchmarks.

Jinhao Li, Yuhao Sun, Zhiyuan Ma, Hao He and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces

Actor-accelerated PDA learns a policy network to approximate PDA optimization subproblems, speeding up continuous-action reinforcement learning while preserving convergence guarantees and outperforming PPO.

Ji Gao, Caleb Ju, Guanghui Lan, Zhaohui Tong

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

GARDO: Reinforcing Diffusion Models without Reward Hacking

GARDO selectively regularizes high-uncertainty diffusion samples and adaptively updates reference models to prevent reward hacking while preserving diversity and sample efficiency.

Haoran He, Yuxiao YE, Jie Liu, Jiajun Liang and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 30 on Hugging Face · Code ★ 63

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions

Gaussian trust region reshaping replaces monotonic divergence penalties with bounded non-monotonic constraints, unlocking efficient behavior transitions in non-stationary reinforcement learning.

Bingxu Liu, Jiashun Liu, Johan Obando Ceron, Hao Wang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs

Augmented Lagrangian framework achieves global last-iterate convergence for constrained MDPs with tabular, log-linear, and nonlinear policies.

Michael Lu, Max Lin, Mo Chen, Sharan Vaswani

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 2/5
medium 3/10
strict 1/5
83%Must read
?Must readVote to see the score

Utility-Constrained Policy Optimization

A practical utility-constrained MDP framework enables risk-sensitive safety constraints and flexible limit adjustments without retraining, matching or outperforming Safety Gymnasium baselines.

Mehrdad Moghimi, Bernardo Avila Pires

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
86%Must read
?Must readVote to see the score
NeurIPS 2026ColumbiaDeep RL

Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients

Hybrid Policy Optimization combines pathwise and score-function gradients via simulator backpropagation to train policies in hybrid action spaces with unbiased mixed gradients, outperforming PPO on high-dimensional control tasks.

Matias Alvo, Daniel Russo, Yashodhan Kanoria

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models

Flow-DPPO replaces PPO ratio clipping with exact KL divergence constraints for flow matching models, improving reward, stability, and multi-objective alignment.

Bowen Ping, Xiangxin Zhou, Penghui Qi, Minnan Luo and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 42 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Mirror descent actor-critic methods for entropy regularised MDPs in general spaces: stability and convergence

Policy mirror descent with inexact TD actor-critic converges for entropy-regularized MDPs in general spaces, with sublinear or linear rates under sufficient TD steps.

Denis Zorba, David Siska, Lukasz Szpruch

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 2/5
medium 2/10
strict 2/5