Good Papers

Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion

Orthrus unifies autoregressive and diffusion views in transformers to enable lossless parallel token generation with up to 7.8x speedup and O(1) memory overhead.

Van Chien Nguyen, Chaitra Hegde, Van-Cuong Pham, Ryan Rossi, Franck Dernoncourt, Thien H Nguyen

Published 2026Atlanta Poster Session 1 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall C1▲ 12 on Hugging FaceCode ★ 482arXiv ↗OpenReview ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/20reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

We introduce Orthrus, a simple and efficient dual-architecture framework that unifies the exact generation fidelity of autoregressive Large Language Models (LLMs) with the high-speed parallel token generation of diffusion models. The sequential nature of standard autoregressive decoding represents a fundamental bottleneck for high-throughput inference. While diffusion language models attempt to break this barrier via parallel generation, they suffer from significant performance degradation, high training costs, and a lack of rigorous convergence guarantees. Orthrus resolves this dichotomy natively. Designed to seamlessly integrate into existing Transformers, the framework augments a frozen LLM with a lightweight, trainable module to create a parallel diffusion view alongside the standard autoregressive view. In this unified system, both views attend to the exact same high-fidelity Key-Value (KV) cache; the autoregressive head executes context pre-filling to construct accurate KV representations, while the diffusion head executes parallel generation. By employing an exact consensus mechanism between the two views, Orthrus guarantees lossless inference, delivering up to a 7.8x speedup with only an O(1) memory cache overhead and minimal parameter additions.