Good Papers
NeurIPS 2026Deep RLU Alberta

The Laplacian Keyboard: Beyond the Linear Span

Laplacian Keyboard hierarchically combines Laplacian eigenvectors into a behavior library with a meta-policy, exceeding linear span limits for better zero-shot approximation and sample efficiency.

Siddarth Chandrasekar, Marlos C. Machado

Published 2026Sydney Poster Session 4 · Wed, Dec 9, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

71%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel6/20reviewers recommend it
lenient 3/5
medium 3/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
The Laplacian Keyboard earns praise for task-agnostic compositional control that outperforms linear span limits, though critics question whether the meta-policy cost and missing option-critic ablations erode its sample-efficiency gains.

Abstract

Across scientific disciplines, Laplacian eigenvectors serve as a fundamental basis for simplifying complex systems, from signal processing to quantum mechanics. In reinforcement learning (RL), they similarly form a basis over the state space, enabling reward functions to be approximated by projection onto a small set of eigenvectors. This projection makes zero-shot control possible, but it also imposes a fundamental limitation: the induced policies are only as expressive as the linear span of the chosen eigenvectors. We introduce the Laplacian Keyboard (LK), a hierarchical framework that goes beyond this linear span. LK constructs a task-agnostic library of behaviors from these eigenvectors, forming a behavior basis guaranteed to contain the optimal policy for any reward within the linear span. A meta-policy learns to stitch these behaviors dynamically, enabling efficient learning of policies outside the original linear constraints. We establish theoretical bounds on zero-shot approximation error and demonstrate empirically that LK improves over the zero-shot solution while achieving better sample efficiency compared to standard RL methods.