Good Papers

Showing Unified multimodal models Show all papers

72%Highly rated
?Highly ratedVote to see the score

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Prism introduces dynamic sparse attention via adaptive macro-zone block shapes guided by visual variance and cross-modal attention for 2K joint video-audio generation, yielding 2.5x training speedup and improved quality.

Shuyuan Tu, Qi Tian, Yinming Huang, Yue Wu and 7 more

Published Oct 4, 2026 · ▲ 4 on Hugging Face · Code ★ 30

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
72%Highly rated

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Kandinsky 6.0 Video introduces diffusion models that generate synchronized 5-second audio-video clips with lip-sync via a dual-stream CrossDiT architecture, with the 29B-parameter Pro version outperforming its predecessor and matching top competitors in speech quality.

Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin and 36 more

Published Oct 4, 2026 · ▲ 113 on Hugging Face · Code ★ 119

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Octrees as an Explicit 3D Language

OctLLM represents 3D geometry via sparse octree tokens and trains separate 3D branches to achieve state-of-the-art multimodal 3D generation and understanding without degrading language ability.

Ran Dan, Si‐Tong Wei, Pengfei Xiong, Wei Zhang and 2 more

Published Oct 1, 2026 · 0 citations · ▲ 11 on Hugging Face · Code ★ 12

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
72%Highly rated

UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

UniEvo-VL improves multimodal image generation via self-distillation that minimizes divergence between student and critique-conditioned teacher diffusion distributions during self-correction. Experiments on Qwen2.5-Image improve GenEval scores from 0.747 to 0.808 without external teachers.

Fang Wu, Da Xing, Yanjie Huang, Junxi Wang and 15 more

Published Sep 30, 2026 · 0 citations · ▲ 292 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL

Adaptive Reward Routing dynamically routes updates and balances rewards during forward-process RL for joint audio-video diffusion, consistently improving quality, alignment, and synchronization over fixed baselines.

Songlin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang and 4 more

Published Sep 29, 2026 · 0 citations · ▲ 138 on Hugging Face

– ReadersNo votes yet
6/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 6 of 20 reviewers recommend it
lenient 2/5
medium 4/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

HelixWorld: A Real-time Interactive Audio-Visual World Model

HelixWorld is a real-time interactive audio-visual world model that synchronizes visual scenes and spatial stereo sound under user control at 24 FPS, surpassing silent models in acoustic immersion.

Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang and 12 more

Published Sep 29, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 623

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

PixelUMM is an encoder-free unified model for image and video understanding and generation that represents images as spatial patches and videos as spatiotemporal tubelets, achieving competitive performance across tasks.

Cong Wei, Xuanchi Ren, Bryan Chu, Weiming Ren and 7 more

Published Sep 29, 2026 · 0 citations · ▲ 32 on Hugging Face · Code ★ 161

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
71%Highly rated
?Highly ratedVote to see the score

Orca: The World is in Your Mind

Orca is a world foundation model that learns a unified latent space via next-state prediction from video and language, enabling scalable text, image, and action generation.

Yihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen and 36 more

Published Jun 29, 2026 · 0 citations · ▲ 511 on Hugging Face · Code ★ 1,021

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

ERNIE 5.0 Technical Report

ERNIE 5.0 is a trillion-parameter unified autoregressive multimodal model using sparse MoE and elastic training to support diverse understanding and generation tasks.

Haifeng Wang, Hua Wu, Tian Wu, Yu Sun and 36 more

Published Feb 4, 2026 · 2 citations · ▲ 268 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

LTX-2: Efficient Joint Audio-Visual Foundation Model

LTX-2 is an open-source 14B/5B audiovisual transformer generating synchronized high-quality video and audio with state-of-the-art open-source quality at low computational cost.

Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman and 25 more

Published Jan 6, 2026 · 0 citations · ▲ 197 on Hugging Face · Code ★ 9,606

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Tuna-2 replaces vision encoders with patch embeddings for end-to-end pixel-space multimodal understanding and generation, achieving state-of-the-art results that outperform encoder-based designs at scale.

Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen and 11 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 70 on Hugging Face · Code ★ 756

– ReadersNo votes yet. 1 from authors or colleagues not counted
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

MSCR: Jointly Balancing Modality Utilization and Discovering Synergistic Information

Xinyu Chen, Liangjian Wen, Jiang Duan, Dongkai Wang and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MEME: Lightweight Hierarchical Mixture-of-Experts for Unified Affective Computing

Yinan Zhang, Haoyu Zhang, Tianshu Yu

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DiM$^3$: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging

Zijing Wang, Mingyang Wang, Ercong Nie, Yongkang Liu and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Benchmarking and Optimizing Multimodal Structured Generation: The OracleGraph Dataset and PRISM Framework

Zhihong Sun, Qi Fan, Pan Liu, Yang Yi and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

FuseAdapt: Adaptation-Space Fusion for Multi-Modal Semantic Segmentation with Missing Modalities

Xin Zhang, Robby Tan

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

COSMIO: A Benchmark for Cross-Survey Modality Imputation

Dichang Zhang, Yixuan Shao, Jiali Cui, Haotian Yin and 7 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Process

Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou and 6 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SING-CH: Task-Scale-Agnostic Lifelong Cross-Modal Hashing on Statistical Manifolds

Haoran Yang, Junge Chen, Jun Long, Zhan Yang

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization

Dongchao Yang, Yuanyuan Wang, Songxiang Liu, Dading Chong and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
Show 20 more papers