Good Papers

Showing papers from Youtu Lab, Tencent Show all papers

80%Must read
?Must readVote to see the score

Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory

IAMFlow is a training-free identity-aware memory framework that tracks persistent entities across prompts to generate consistent long narrative videos, achieving best benchmark scores and faster inference.

Jinzhuo Liu, Jiangning Zhang, Wencan Jiang, Yabiao Wang and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

L2P: Unlocking Latent Potential for Pixel Generation

L2P transfers pre-trained latent diffusion models to pixel space via frozen intermediate layers and synthetic data, enabling efficient 4K generation with near-source performance.

Zhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang and 6 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 36 on Hugging Face · Code ★ 195

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
88%Must read
?Must readVote to see the score

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Counterfactual Evidence Disentanglement (CED) audits vision-language model grounding by comparing evidence-region and non-evidence-region support drops inside GRPO, improving visual reasoning across benchmarks without inference overhead or evidence annotations.

Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen and 5 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 16 on Hugging Face · Code ★ 4

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 3/5
74%Highly rated
?Highly ratedVote to see the score

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents

SPIRAL uses sequential planning and reflective agents in a closed loop to generate long-horizon action-conditioned videos with iterative refinement and self-evolving post-training.

Yu Yang, Yue Liao, Jianbiao Mei, Baisen Wang and 7 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 0/5
76%Highly rated
?Highly ratedVote to see the score

VicEdit: Learning to Edit Videos from Visual In-Context Examples

VicEdit enables visual in-context video editing via multi-modal guidance and achieves state-of-the-art results on instruction and visual reference tasks.

Yuji Wang, Teng Hu, Yuheng Chen, Ran Yi and 5 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 2/5