78%Highly rated
?Highly ratedVote to see the score

Native Audio-Visual Alignment for Generation
NAVA proposes native audio-visual alignment with an Align-then-Fuse MMDiT architecture for joint audio-video generation, achieving superior synchronization, video quality, and timbre control with 6.3B parameters.
Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 31 on Hugging Face · Code ★ 226
– ReadersNo votes yet
11/20 AI panelreviewers recommend it
Readers and the AI panel: vote on this paper to see what they said.
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5