LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches
LoGRA reduces LLM reinforcement learning memory by up to 45.7% via low-rank gradient sketches and predicted-KL step control, enabling 27B-parameter training on single nodes.
Published Oct 5, 2026▲ 11 on Hugging FacearXiv ↗
Only vote on papers you've read. Sign in with GitHub to vote.
LoGRA delivers a practical, load-bearing memory win, cutting RL training memory by ~46% and sustaining 27B-parameter models for over 1,100 steps on a single node where dense Adam fails, though its low-rank sketches and predicted-KL guardrail…
Abstract
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. Across reasoning tasks, LoGRA reduces average training memory by up to 45.7\% without sacrificing performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra}{Molt library}.