Good Papers

Equivariant Reinforcement Learning for Clifford Quantum Circuit Synthesis

A qubit-relabeling-equivariant, size-agnostic reinforcement learning agent synthesizes near-optimal Clifford circuits across qubit counts, outperforming Qiskit on large instances.

Richie Yeung, Aleks Kissinger, Rob Cornish

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 3/5
medium 4/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
The paper delivers a striking equivariant, size-agnostic policy that finds near-optimal Clifford circuits at scale, though its reliance on all-to-all connectivity, absence of ablation, and lack of non-Clifford or topological benchmarks leave its practical synthesis value…

Abstract

We consider the problem of synthesizing Clifford quantum circuits for devices with all-to-all qubit connectivity. We approach this task as a reinforcement learning problem in which an agent learns to discover a sequence of elementary Clifford gates that reduces a given symplectic matrix representation of a Clifford circuit to the identity. This formulation permits a simple learning curriculum based on random walks from the identity. We introduce a novel neural network architecture that is equivariant to qubit relabelings of the symplectic matrix representation, and which is size-agnostic, allowing a single learned policy to be applied across different qubit counts without circuit splicing or network reparameterization. On six-qubit Clifford circuits, the largest regime for which optimal references are available, our agent finds circuits within one two-qubit gate of optimality in milliseconds per instance, and finds optimal circuits in 99.2% of instances within seconds per instance. After continued training on ten-qubit instances, the agent scales to unseen Clifford tableaus with up to thirty qubits, including targets generated from circuits with over a thousand Clifford gates, where it achieves lower average two-qubit gate counts than Qiskit's Aaronson-Gottesman and greedy Clifford synthesizers.