
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training
Prism introduces dynamic sparse attention via adaptive macro-zone block shapes guided by visual variance and cross-modal attention for 2K joint video-audio generation, yielding 2.5x training speedup and improved quality.
Published Oct 4, 2026 · ▲ 4 on Hugging Face · Code ★ 30
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.









