Good Papers

Elastic Spectral State Space Models for Train-Once Budgeted Inference

ES-SSM enables train-once deployment across budgets by truncating spectral SSM channels, yielding smooth quality-cost curves and competitive compact models.

Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang

Published 2026Atlanta Poster Session 3 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall C1arXiv ↗OpenReview ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Modern sequence models are typically trained at a fixed computational capacity, while real-world applications require deployment across platforms with different resource constraints. Existing approaches either train or distill separate compact models for selected budgets, or use elastic architectures that shrink dimensions such as width, depth, Feed-Forward Network (FFN) size, or SSM state dimension. In this paper, we propose Elastic Spectral State Space Models (ES-SSM), a train-once, export-many sequence modeling framework that gains elasticity through spectral approximation of the SSM sequence operator. ES-SSM builds on Hankel spectral filtering for state space models, where long-range token mixing is represented through fixed Hankel spectral channels that define an operator-level approximation resolution. To make this spectral elasticity reliably deployable, ES-SSM combines input-adaptive channel-wise gates with budget dropout, training the same spectral prefixes that are used by compact models at inference time. This encourages low-index channels to become predictive on their own, while higher-index channels provide refinements for larger budgets. We evaluate ES-SSM across byte-level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning benchmarks. Across these settings, a single trained ES-SSM can be truncated to competitive compact models compared with modern Transformer and SSM baselines at similar parameter scales. Furthermore, by testing under various runtime budgets, we observe smooth and stable quality-cost curves over a wide range of truncation levels.