Good Papers

Showing Deep RL Show all papers

76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

MEND: RL For Flow Models via Proximal Velocity Matching

MEND uses proximal velocity matching to cap rewards and accept only cost-effective sample moves, outperforming prior flow-model RL methods in far fewer updates without KL penalties or reference models.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli and 1 more

Published Oct 5, 2026 · ▲ 2 on Hugging Face · Code ★ 1

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 0/5
83%Must read
?Must readVote to see the score
arXivDeep RL

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

The Homomorphic Advantage Operator stabilizes FHE-based reinforcement learning by centering TD targets to eliminate Bellman drift, achieving zero approximation-bound breaches and 18-point accuracy gains without extra multiplicative depth.

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

Published Oct 1, 2026 · 0 citations

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 9/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score
arXivDeep RL

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

iADD analyzes diffusion policy optimization to show only-latter-timestep updates harm diversity, then proposes incremental Feynman-Kac training that improves alignment-diversity tradeoffs across tasks.

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli and 1 more

Published Oct 1, 2026 · 0 citations · Code

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 2/5
medium 7/10
strict 1/5
83%Must read
?Must readVote to see the score

Tail-Influence Sampling for CVaR Policy Evaluation

Tail-Influence Sampling allocates evaluation budgets by tail influence to estimate CVaR with oracle variance and lower MSE than rollouts.

Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov and 1 more

Published Sep 29, 2026 · 0 citations · ▲ 23 on Hugging Face · Code ★ 1

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score
NeurIPS 2026DeepMindDeep RL

Delightful Distributed Policy Gradient

Delightful Policy Gradient gates distributed updates with delight (advantage times surprisal) to suppress high-surprisal failures while preserving rare successes, outperforming importance-weighted methods under staleness, bugs, and corruption.

Ian Osband

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

0% Readers0 of 1 upvoted
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 3/5
medium 10/10
strict 3/5
90%Must read
?Must readVote to see the score

Modeling quantum neural network gradient with reinforcement learning

RLQ-Grad uses reinforcement learning to propose quantum neural network updates without differentiating circuits, avoiding barren plateaus and scaling with parameters rather than Hilbert space dimension to achieve orders-of-magnitude faster training and higher accuracy.

Nhan Luu, Trung D Luu, Ngoc Nam Pham, Thang C Truong

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
45%Niche pick
?Niche pickVote to see the score

A Primal-dual Approach for Semi-Infinitely Constrained Reinforcement Learning

Di Wang, Liangyu Zhang, Haishan Ye, Guang Dai and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026XidianDeep RL

Reinforcement Learning for View-Adaptive Distillation in 3D Gaussian Compression

Hongji Zhao, Mingrui Zhu, Xin Wei, Nannan Wang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning

Boyang Li, Matthew Kim, Sylvia Herbert

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Finite-Memory Control of POMDPs: Fundamental Limits and Efficient Design

Emrecan Kutay, Atilla Eryilmaz, Ness Shroff

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling

Nicholas Corrado, Wenyuan Huang, Josiah Hanna

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ROLLVERIFY: BRIDGING EFFICIENCY AND ACCURACY IN LONG-TAIL ROLLOUT REINFORCEMENT LEARNING

Yongqiang Yao, Jingru Tan, Kaihuan Liang, Zixin Yin and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Accelerating Safe Reinforcement Learning with Massive Parallelism

Joonyoung Lim, Younghwan Yoo

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

When Does Knowing the State Help? Diagnosing Process vs. Outcome Reward Design

Wenpei Shao, Ross Jacobucci

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Interactive Combinatorial Reinforcement Learning for Knowledge Graph Reasoning

Jun Nie, Yonggang Zhang, Tongliang Liu, Chengqi Zhang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026DalhousieDeep RL

Instability of Meta-Learning Intrinsic Rewards for Policy Gradient Reinforcement Learning

Dilith Jayakody, Domenic Rosati, Janarthanan Rajendran

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

LOCU: Löwdin-Orthogonalized Constraint Updates for Multi-Constraint Policy Optimization

Joonyoung Lim, Younghwan Yoo

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Exploring Lifelong Adaptation: In-Context Reinforcement Learning in Non-Stationary Environments

Ye Wang, Kaiqian Cui, Xinrun Xu, Tao Zhang and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Cross-Question Reliable Reinforcement Learning

Hector G. Rodriguez, Marcus Rohrbach

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking During RL

Aria Wong, Joshua Engels, Neel Nanda

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
Show 20 more papers