Good Papers

Showing Video generation Show all papers

90%Must read
?Must readVote to see the score

World Models' Last Exam in Physics

World Models' Last Exam in Physics benchmarks video models via 40 measurement-based physics tasks, finding the best model scores 57.76/100 with widespread inconsistencies.

Mingju Gao, Qingle Liu, Yuzhao Peng, Xinjie Lin and 7 more

Published Oct 6, 2026 · ▲ 2 on Hugging Face

100% Readers1 of 1 upvoted
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

ALIVE: Interaction-Aligned Object Insertion for First-Frame-Guided Video Editing

ALIVE inserts objects that interact with video contents via first-frame editing and interaction guidance, outperforming baselines on interaction and insertion benchmarks.

Zhenghong Zhou, Zhe Lin, Jiebo Luo, Yuqian Zhou

Published Oct 6, 2026 · ▲ 3 on Hugging Face · Code

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

S2PD switches from autoregressive to parallel diffusion during denoising to enforce physical and logical consistency with faster sampling than fully serial methods.

Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari

Published Oct 5, 2026 · 0 citations · ▲ 1 on Hugging Face · Code ★ 4

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

89%Must read
?Must readVote to see the score

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

HLA-WM combines geometry-guided retrieval with recurrent linear attention to fix long-range forgetting in video world models, improving 60-second consistency metrics by up to 28.5% with 12× lower memory and no retraining.

Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu and 3 more

Published Oct 5, 2026 · ▲ 6 on Hugging Face · Code ★ 2

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

DuoMatching improves few-step video generation by jointly matching frame distributions and adding frame-level supervision via an image teacher, boosting visual quality and semantic alignment over 80%.

Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 17 on Hugging Face · Code ★ 4

100% Readers1 of 1 upvoted
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

In-Distribution Forcing for Long Video Generation at Test Time

In-Distribution Forcing prevents out-of-distribution key-value drift via self-caching to extend short video models to minute-scale generation.

Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 35 on Hugging Face · Code ★ 4

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

ROWBench: Do Video Models Render What the Program Specifies?

PROWBench evaluates video models' fidelity to program-specified world events via replayable world records and VLM-based logic-render and interaction metrics.

Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu and 5 more

Published Oct 1, 2026 · 0 citations · ▲ 70 on Hugging Face · Code ★ 31

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

World Observer jointly generates actor and panoramic observer views to continuously model out-of-view dynamics via shared geometric warping and observer sinks.

Hyunwook Choi, Dahyun Chung, Hyunsung Kim, Siyoon Jin and 3 more

Published Oct 1, 2026 · 0 citations · ▲ 85 on Hugging Face · Code ★ 30

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Spatial Memory Intelligence introduces understanding-driven atomic operations for spatial-memory management in long-video world models, improving sparsity, stability, and spatial consistency.

Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie and 4 more

Published Oct 1, 2026 · 0 citations · ▲ 54 on Hugging Face · Code ★ 20

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Ego2Act evaluates egocentric video generation on multi-step goal-directed manipulation, showing models skip steps and fail at fine-grained physical dynamics.

Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui and 9 more

Published Oct 1, 2026 · 0 citations · ▲ 33 on Hugging Face · Code ★ 3

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency

DeCoPrune uses denoising-consistency to prune over 85% of KV-cache tokens in autoregressive video diffusion, preserving long-range recall and accelerating continuation generation by over 4x.

Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou and 2 more

Published Sep 30, 2026 · 0 citations · ▲ 1 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

FrameMorrow guides historical frame selection via prospective tokens representing future needs, improving consistency and quality across diverse long-horizon video generators.

Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan

Published Sep 30, 2026 · 0 citations · ▲ 105 on Hugging Face · Code ★ 29

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Video Generation Models: A Survey of Post-Training and Alignment

This survey reviews post-training and alignment strategies for video generation, framing them as implicit or explicit alignment across four methodological categories to improve controllability and reliability.

Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im and 9 more

Published Sep 30, 2026 · ▲ 60 on Hugging Face · Code ★ 208

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

Rollout-Marginal Distillation scores autoregressive video chunks independently against a chunk teacher to prevent error accumulation, then applies video-level distillation to restore temporal coherence.

Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang and 2 more

Published Sep 29, 2026 · 0 citations · ▲ 18 on Hugging Face · Code ★ 2

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

91%Must read

QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models

QuantWM is a training-free 2-bit KV cache quantization framework for video world models that preserves attention logits and token selection to eliminate temporal flickering while achieving up to 6.20x memory compression.

Jiaqi Zhao, Xiaobin Hu, Bo Yin, Junpeng Jiang and 2 more

Published Sep 22, 2026 · 0 citations · ▲ 13 on Hugging Face

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Vidu S2 enables real-time interactive avatar and editing video generation at 720p with dynamic references and spatial capabilities, outperforming all baselines.

Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang and 31 more

Published Sep 10, 2026 · 0 citations · ▲ 706 on Hugging Face · Code ★ 502

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

VGI-Bench: Probing Visual Intelligence in Video Generation Models

VGI-Bench evaluates video generation models via 27 visual reasoning tasks, finding top models achieve only 51% accuracy with limited self-correction.

Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma and 19 more

Published Aug 20, 2026 · 0 citations · ▲ 336 on Hugging Face · Code ★ 14

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Kandinsky 5.0 introduces state-of-the-art 6B image and 2B/19B video generation models with optimized training and inference for high-speed, high-quality synthesis.

Arkhipkin, Vladimir, Korviakov, Vladimir, Gerasimenko, Nikolai, Parkhomenko, Denis and 21 more

Published Nov 19, 2025 · 0 citations · ▲ 236 on Hugging Face · Code ★ 834

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

LongCat-Video Technical Report

LongCat-Video is a 13.6B DiT video generation model that efficiently produces high-quality minutes-long 720p videos via coarse-to-fine generation, block sparse attention, and multi-reward RLHF.

Meituan LongCat Team, Xunliang Cai, Qilong Huang, Zhuoliang Kang and 7 more

Published Oct 25, 2025 · ▲ 44 on Hugging Face · Code ★ 9,022

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 3/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MilliVid: Adaptive Latents for Long-Range Consistency in Video Generation

Ishaan Chandratreya, David Charatan, Basile Van Hoorick, Sergey Zakharov and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Bidirectional Sparse Attention for Faster Video Diffusion Training

Chenlu Zhan, Wen Li, Jun Zhang, chuyu shen and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion

Yongji Long, Shijun Liang, Jintao Li, Yun Li

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ReGen: Agentic Video World Modeling with Synergized Reasoning and Generation

Chenguo Lin, Yu Tang, Weiqiao Zheng, Enhua Jiang and 10 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

AEON: Unifying Video and 3D World Models

Weiqi Zhang, Wenyuan Zhang, Yu-Shen Liu, Junsheng Zhou

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Spatially-Grounded Long Video Generation with Self Geometry Forcing

Chenguo Lin, Panwang Pan, Bowen Xue, Ruijie Lu and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

NIKA: Efficient Neural Video Representation via Structured Latent Diversity

Slater R Victoroff, Madison May

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

StableAvatar: Ultra-Long Audio-Driven Avatar Video Generation

Shuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Latent Motion Alignment for Video Diffusion

Nick Stracke, Kolja Bauer, Björn Ommer

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score
NeurIPS 2026Video generation

Freshness-Gated Imagination: Step-Level Trust for Latent World Models

Chunqi Guo

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

RISE-Video: Can Video Generators Decode Implicit World Rules?

Mingxin Liu, Shuran Ma, Shibei Meng, Xiangyu Zhao and 8 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

PhysEval: Quantifying the Gap Between Video Generation and World Physical Laws

Hongchu Zeng, Sijing Wu, Yanhan Zhou, Yunhao Li and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

HorizonComposer: Spatiotemporally Consistent Driving Video Editing with Enriched Traffic Semantics

Mauricio Soroco, Yuqiu Liu, Zaid Tasneem, Francesco Pittaluga and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

PhaseDance: Capturing Rhythm and Expressivity in Dance Modeling

Meongeun Kim, Taehui Lee, Soomin Park, Sung-Hee Lee

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AR-Edit: Training-Free Streaming Video Editing without Inversion

Hovhannes Margaryan, Vicky Kalogeiton, Quentin Bammey, Christian Sandor

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PAI-Actor: Cinematic Multi-Actor Character Replacement in Dynamic Scenes

Heyuan Gao, Bangxun Tang, Yiren Song, Guian Fang and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

SAX: Advancing Video Diffusion Models for Sequential Action Execution

Haoyu Wang, Baorui Ma, Donglin Di, Shiliang Zhang

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Persistent Planar Memory for Video World Models

Yuze He, Bin Tan, Zelin Gao, Yujun Shen and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Recreating Video Arenas via Automated Preference Scoring

Yue Zhao, Aniket Gupta, Juze Zhang, Tiange Xiang and 4 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Not All Layers Are Equal in Image-to-Video Transfer

Thinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina Roitberg

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PlasticMem: Adding Temporal Reasoning to Diffusions for Consistent Long Video Generation

Yufei Huang, Chao Liang, Zerong Zheng, Tianshu Hu and 5 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ApertureAttn: Native 4K Video Generation with Image-Only Supervision

Ruonan Yu, Zigeng Chen, Zhenxiong Tan, Songhua Liu and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

TKCAM: Text and Keyframe to Camera Trajectory Generation

Haozhe Yang, Zhiyang Dou, Zekai Gu, Cheng Lin and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AsymHP: Load-Balanced Sparse Attention for Video Diffusion Transformers

Xinwei Qiang, Yue Guan, Ruihan Zhu, Mihir Jagtap and 6 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

FakeParts: a New Family of AI-Generated Video Forgeries

Ziyi LIU, Firas Gabetni, Awais H SANI, Xi Wang and 4 more

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

FacEDiT: Talking Head Video Editing via Facial Motion Infilling

Sung-Bin Kim, Joohyun Chang, David Harwath, Tae-Hyun Oh

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Transforming Image Editors into Video Editors

Feng Wang, Zijie Li, Ceyuan Yang, Alan Yuille and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity

Chengfeng Han, Baole Ai, Xianlu Bian, Jie Yao and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

GeoMemory: Geometry-Indexed Memory for Long-Horizon Interactive Video Generation

Junchao Huang, Xinting Hu, Boyao Han, Shaoshuai Shi and 3 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

LogicDirector: Enforcing Temporal Composition in Text-to-Video Generation

Yujiang Pu, Zixu Cheng, Shaogang Gong, Handong Zhao and 1 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Slide P2V-Bench: A Cross-Domain Benchmark for Slide-Centric Scientific Paper-to-Presentation Video Generation

Ligang Huang, Zhou Liu, Wentao Zhang

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Physics-Constrained Generative World Model for Off-Road Terrain via Post-hoc Projection

Yuan Zhou, Hao Yu, Ruiran Cao, Haoran Yang and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

WorldForge: Forging Unified World Modeling into Video Generation

Boming Tan, Xiangdong Zhang, Ning Liao, Jingtao Zhang and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Dynamics-Aware Sparse Attention for Efficient Autoregressive Video Diffusion

Yuanyu He, Zhuokun Chen, Yefei He, Zhiwei Tang and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Beyond Copy-Paste: How Well Do Subject-Driven Video Models Understand Their Subjects?

Zun Wang, Kenan Deng, Daniel Blackburn, Linlin Lu and 4 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Learning to Correct Geometry in Generated Videos

Weijie Lyu, Tiancheng SHEN, Xiangtai Li, Yujing Wang and 4 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MedCache: Training-Free Spatially Aware Caching for Accelerated Medical Video Generation

Ufaq Khan, Umair Nawaz, Sathira Silva, Numan Saeed and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Learning-to-Memorize: Dynamic Context Management for Long-Horizon Autoregressive Video Generation

Haowei Zhu, Qijie Wang, Jia Li, XING WANG and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

TriTD: Tri-Partite Trajectory–Distribution Distillation for Real-Time Autoregressive Video Generation

Jia Li, Xurui Peng, Haowei Zhu, Tianyu Zhao and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Decomposed Graded Verifier for Generative World Modeling

Bowei Liu, Xinchen Zhang, Xuhuan Li, Kaian Jiang and 9 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

UniMoE-World: A Unified Mixture-of-Experts Architecture for Scalable Multi-Control Video Generation World Modeling

Jianjie Fang, Yongyan Xu, Ziyou Wang, Yuchao Huang and 12 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

PhysFlow: Physics-Intrinsic Velocity Regularization for Motion-Intensive Video Generation

Xianglong Guo, Chang Yu, Haobo Xu, Junhao Ma and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Position: Next-Generation Game Engines Should Be Built on Interactive Generative Video

Jiwen Yu, Yiran Qin, Haoxuan Che, Quande Liu and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Rate-Constrained Edge Metadata for Sender–Receiver Generative Video Super-Resolution

Jiaqi Guo, Mingzhen Li, Haohong Wang, Aggelos Katsaggelos

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MotionCFG: Boosting Motion Dynamics via Semantic Motion Sharpening

Byungjun Kim, Soobin Um, Jong Chul Ye

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

ENACT: Single-Image Human-Scene Interaction Motion from Language via Foundation-Model Orchestration

Sreehari Rajan, Kunal Kamalkishor Bhosikar, Charu Sharma, Nikos Athanasiou

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Meta Inverse Prompting for Video Generative Models

Yerin Jung, Jongheon Jeong

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

What Should a Streaming Video Model Remember?

Haonan Ge, Yiwei Wang, Hang Wu, Yujun Cai

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Tri-Prompting: Controllable Video Generation with Scene, Subject, and Motion Prompts

Zhenghong Zhou, Xiaohang Zhan, Zhiqin Chen, Soo Ye Kim and 7 more

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Motion Forcing: Decoupling Ego and Object Motion via Sparse Inputs for Structured Video Generation

Tianshuo Xu, ZhiFei Chen, Leyi Wu, Hao LU and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ReFree-S2V: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance

Salaheldin Youssry Abdellah Elsadek Mohamed, M. Hamza Mughal, Rishabh Dabral, Christian Theobalt

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Video Generation Research Needs a Legally Sustainable Data Infrastructure

Wenhao Wang, Biao Wu, Jiabin Luo, Xinyu Zhang

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife

Yuzhuo Li, Di Zhao, Xinyu Zhang, Daniel Wilson and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

Coarse-to-Real: Generative Rendering for Populated Dynamic Scenes

C2R generates realistic, temporally consistent urban crowd videos from coarse 3D simulations via a neural renderer guided by text and a synthetic-real domain-hedging strategy.

Gonzalo Gomez-Nogales, Yicong Hong, Chongjian GE, Peiye Zhuang and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

NEvo uses neural-guided evolutionary video synthesis to generate brain-region-optimized dynamic stimuli that surpass handcrafted localizers and reveal visual cortex temporal selectivity differences.

Yingtian Tang, Sogand Salehi, Ming Zhou, Amir Zamir and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory

IAMFlow is a training-free identity-aware memory framework that tracks persistent entities across prompts to generate consistent long narrative videos, achieving best benchmark scores and faster inference.

Jinzhuo Liu, Jiangning Zhang, Wencan Jiang, Yabiao Wang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning

Equilibrium Forcing removes noise conditioning from video diffusion to enable adaptive closed-loop inference that improves generation quality and consistency.

Hansen Lillemark, Alex Rojas, Zachary Novack, Runqian (Ray) Wang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

VideoMLA replaces per-head video diffusion KV caches with shared low-rank latents to cut memory by 92.7% and improve long-horizon streaming quality and throughput.

Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil K Akan and 3 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 24 on Hugging Face · Code ★ 18

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

PermaVid disentangles video memory into RGB appearance and depth structure with edit-aware updates to maintain long-term consistency across modifications.

Shuai Yang, Bingjie Gao, Ziwei Liu, Jiaqi Wang and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 44

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
83%Must read
?Must readVote to see the score

RefDecoder: Enhancing Visual Generation with Conditional Video Decoding

RefDecoder conditions video VAE decoders on reference images via attention to recover lost detail, boosting reconstruction PSNR by up to 2.1 dB and improving consistency across video generation tasks without retraining.

Xiang Fan, Yuheng Wang, Bohan Fang, Jason Ren and 1 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction

FreeSpec analyzes long-video diffusion via singular-spectrum analysis and proposes spectral reconstruction to reduce spectral concentration, preserving dynamics and consistency without training.

Fangda Chen, Shanshan Zhao, Longrong Yang, Chuanfu Xu and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling

AtlasVid decouples global-local video diffusion modeling to generate 4K long videos with 60.9x speedup via low-resolution proxies and hierarchical attention, training only at 720P.

Ziyang Mai, Yuyao Zhang, Yu-Wing Tai

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 1/5
88%Must read
?Must readVote to see the score

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

EntityBench introduces 140-episode multi-shot video benchmark with per-shot entity schedules and three-pillar evaluation, showing explicit per-entity memory yields highest character fidelity.

Ruozhen He, Meng Wei, Ziyan Yang, Vicente Ordonez

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 3 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

74%Highly rated
?Highly ratedVote to see the score

OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation

OmniMem enables scalable long-video generation via sparse full-range KV retrieval with adaptive window exclusion and query-shared per-head scattered access, improving dynamic degree by 52.3% while preserving consistency.

Lin Zhao, Yushu Wu, Yifan Gong, Yanzhi Wang and 1 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

Temporal Backtracking Search for Test-time Generative Video Reasoning

Temporal Backtracking Search improves video reasoning by searching over the temporal axis and restarting from verified prefixes rather than resampling from scratch, achieving 22.7% versus 0.7% best-of-N out-of-distribution.

SeJoon Jun, Zheng Ding, Huangyuan Su, Weirui Ye and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Addressable Memory for Video World Models

Interactive video world models lose visual persistence beyond training horizons because temporal RoPE offsets become out-of-distribution; WorldTrace assigns compressed memory slots virtual in-distribution positions to restore addressability, boosting temporal consistency by 15.5% and episodic recall

Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 16 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer

MotionGrounder enables multi-object motion transfer via a diffusion transformer with flow-based motion signals, object-caption alignment loss, and a new object grounding score. It outperforms baselines in multi-object controllable video generation.

Samuel Teodoro, Yun Chen, Agus Gunawan, Soo Ye Kim and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time

DynaTokens teaches dynamics to frozen camera-controlled video models via scene-specific learnable tokens, improving simultaneous dynamics and camera control over full fine-tuning.

Ziqi Ma, Hongqiao Chen, Georgia Gkioxari

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

Parameterized Stripe Attention for Efficient Video Generation

Parameterized stripe attention exploits periodic diagonal stripe structures in video DiT attention to unify sparse patterns in one hardware-efficient kernel, achieving 1.57× speedups over FlashAttention-3 with minimal quality loss.

xingyu jia, Baole Ai, Ang Wang, Kang Zhao and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

71%Highly rated
?Highly ratedVote to see the score

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

FaithfulFaces improves identity-preserving video generation via pose-shared identity alignment and achieves state-of-the-art consistency across pose changes and occlusions.

Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, bing ma and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 5/5
medium 2/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

World Models as Group Actions

World models are formalized as group actions to enforce compositional dynamics via identity, inverse, and composition consistency, improving structural metrics without harming visual quality.

Zijie Wang, Wei Zhang, Weiming Zhang, Fanqi Zhang and 3 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

ViTeX-Bench introduces a 387-video benchmark and evaluation protocol for high-fidelity video scene text editing, finding that accuracy, temporal stability, and edit locality remain hard to balance together.

Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
83%Must read
?Must readVote to see the score

Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

Grounded-Exo2Ego couples geometric anchoring with semantic grounding and camera relocalization to robustly generate egocentric video from exocentric inputs, outperforming prior methods on EgoExo4D.

Shengze Wang, Michael Stengel, Tianye Li, Wookie Park and 4 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

OpenCoF: Learning to Reason Through Video Generation

OpenCoF introduces a 17K video reasoning dataset and Wan-CoF model that improves chain-of-frame reasoning via diverse temporal supervision and reasoning tokens.

Xinyan Chen, Renrui Zhang, Ziyu Guo, Dongzhi JIANG and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 26 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

78%Highly rated
?Highly ratedVote to see the score

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

SplitMoE replaces uniform token-wise routing with split semantic and generic experts, improving video diffusion convergence, routing coherence, and generation quality over load-balanced MoEs.

Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 6 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

HumanScore: Benchmarking Human Motions in Generated Videos

HumanScore benchmarks human motion in AI videos via six metrics, revealing gaps between visual plausibility and biomechanical fidelity across 13 models.

Tiange Xiang, Yusu Fang, Tian Tan, Narayan Schütz and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

76%Highly rated
?Highly ratedVote to see the score

Retrieve What’s Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

COVRAG uses depth-based coverage maps and residual-gain frame selection to improve long-horizon geometric consistency in autoregressive video generation with low latency.

Minseok Joo, Dogyun Park, Taehoon Lee, Kyujin Lee and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 25 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

ODEWorld learns continuous latent velocity fields via ODEs to enable arbitrary-resolution world modeling, solving representation collapse and excelling at video generation and robotic control.

Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao and 4 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 14 on Hugging Face · Code ★ 62

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

MonarchRT: Efficient Attention for Real-Time Video Generation

Monarch-RT factorizes video diffusion attention via Monarch matrices to reach 95% sparsity without quality loss, enabling 16 FPS real-time generation on one GPU with 1.4-11.8x kernel speedups.

Krish Agarwal, Zhuoming Chen, cheng Luo, Yongqi Chen and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan and 7 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 54 on Hugging Face · Code ★ 177

– ReadersNo votes yet
4/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 4 of 20 reviewers recommend it
lenient 1/5
medium 3/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

CRePE encodes tokens as depth-aware distributions along curved unified-camera rays to unify camera control, lens geometry, and external geometry guidance in video generation.

Seonghyun Jin, youngmin Kim, Sunwoo Park, Jong Chul Ye

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

88%Must read
?Must readVote to see the score

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

PhyMotion evaluates human video motion via physics-simulated 3D trajectory rewards across kinematics, contact, and dynamics, improving RL post-training realism by +68 Elo.

Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim and 5 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 6 on Hugging Face · Code ★ 49

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

ONE-SHOT factorizes compositional video generation via spatial-decoupled motion injection and hybrid context integration, achieving fine-grained human-environment control and minute-level consistency without 3D alignment.

Fengyuan Yang, Luying Huang, Jiazhi Guan, Quanwei Yang and 7 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 9 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Next Forcing uses multi-chunk prediction to accelerate convergence 2.3x, boost high-frame-rate accuracy 93.1%, and double inference speed for world models.

Gangwei Xu, Qihang Zhang, Jiaming Zhou, Xing Zhu and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 136

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

LaMo extracts self-supervised latent motion priors from unlabeled videos via motion drift loss and prior guidance, improving physical consistency in video diffusion without external supervision.

Bo Jiang, Depu Meng, yihan hu, Yichen Xie and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

86%Must read
?Must readVote to see the score

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

ARGUS introduces multi-view identity mosaic injection and counterfactual training to preserve subject identity across motion, viewpoint changes, and occlusions in video generation.

Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong and 4 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

83%Must read
?Must readVote to see the score

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

ReMind trains video diffusion transformers to use cache memory for evolving hidden states across interruptions via memory-oriented curricula and PM-RoPE, achieving best STEVO-Bench scores without catastrophic forgetting.

Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng and 4 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

72%Highly rated
?Highly ratedVote to see the score

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

EverAnimate restores drifted latent flows via persistent memory and restorative matching, improving long human animation quality and identity consistency over minutes.

WUYANG LI, Yang Gao, Mariam Hassan, Lan Feng and 3 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 0/5
80%Must read
?Must readVote to see the score

Trimming the Long-Tail of Visual World Modeling Evaluation

Tailor-Bench evaluates visual world models on rare physical interactions via regular, unconventional, and impossible scenarios, revealing long-tail performance gaps and superficial visual-pattern reliance.

Bingxuan Li, Yining Hong, Cheng Qian, Hyeonjeong Ha and 5 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 40 on Hugging Face · Code ★ 1

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

80%Must read
?Must readVote to see the score

Solaris: Building a Multiplayer Video World Model in Minecraft

Solaris introduces a multiplayer video world model and automated data system for Minecraft, collecting 12.64M frames to enable consistent multi-agent simulation via staged training that outperforms single-player baselines.

Oscar Michel, Georgy Savva, Daohan Lu, Suppakit Waiwitlikhit and 6 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

70%Highly rated
?Highly ratedVote to see the score

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO

RAVEN trains real-time autoregressive video models via interleaved history-denoising rollouts and CM-GRPO, improving long-horizon quality over distilled baselines.

Yanzuo Lu, Ronglai Zuo, Jiankang Deng

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 14 on Hugging Face · Code ★ 134

– ReadersNo votes yet
5/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 5 of 20 reviewers recommend it
lenient 2/5
medium 3/10
strict 0/5
80%Must read
?Must readVote to see the score

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

SocialDirector uses training-free cross-attention modulation to control actor-action mapping and directional targets in multi-person video generation, significantly improving interaction fidelity.

LIANGYANG OUYANG, Ruicong Liu, Caixin Kang, Yifei Huang and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
89%Must read
?Must readVote to see the score

DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation

DSAQuant aligns video diffusion quantization with denoising stages to preserve visual details and improve compressed text-to-video generation quality.

Shuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Temporal Concentration from Rollout Errors: Implicit Preference Optimization For Text-to-Video Diffusion

cIPO aligns text-to-video diffusion by deriving implicit preferences from reconstruction errors and concentrating optimization on high-error temporal segments to fix sparse artifacts.

henglin liu, Fangyuan Kong, Jing Wang, Yizhou Lin and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

CoDMD adds a copula-aware relational regularizer to distribution matching distillation that improves few-step video generation, achieving 84.46/84.87 VBench scores at 4 steps with ~25× speedup.

Wenhu Zhang, Kun Cheng, Changyuan Wang, Shiyao Li and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 3/5
medium 10/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Bernini: Latent Semantic Planning for Video Diffusion

Bernini unites MLLM semantic planning with DiT video rendering via latent ViT embeddings to achieve state-of-the-art video generation and editing.

Chenchen Liu, Junyi Chen, Lei Li, Lu Chi and 7 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 20 on Hugging Face

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

DyMoS rebalances reference-frame self-attention in image-to-video models to boost motion dynamics without retraining or altering inputs.

Wooseok Jeon, Seungho Park, Seunghyun Shin, Sangeyl Lee and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

PILA injects physics-structured latent guidance into frozen video generators via mixture-of-experts alignment, achieving state-of-the-art physical plausibility and visual quality.

Cong Wang, Hanxin Zhu, Jiayi Luo, Yonglin Tian and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 8/10
strict 0/5
74%Highly rated
?Highly ratedVote to see the score

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents

SPIRAL uses sequential planning and reflective agents in a closed loop to generate long-horizon action-conditioned videos with iterative refinement and self-evolving post-training.

Yu Yang, Yue Liao, Jianbiao Mei, Baisen Wang and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
80%Must read
?Must readVote to see the score

Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

ARL2 replaces cross-frame attention with fixed-size recurrent linear states, achieving linear-time scaling, constant memory, and improved temporal consistency in autoregressive video diffusion.

Kunyang Li, Mubarak Shah, Yuzhang Shang

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5