Good Papers

Showing Video-language models Show all papers

91%Must read

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

LoHi trades per-frame resolution for denser temporal sampling via low-resolution streams plus sparse high-resolution frames, boosting long-video accuracy up to 10.6 points and cutting front-end latency up to 7x.

Sixun Dong, Wei Li, Andong Deng, Qi Qian and 3 more

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published Oct 3, 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet. 1 from authors or colleagues not counted
17/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 3/5
86%Must read
?Must readVote to see the score

World Embedding Benchmark

World Embedding Benchmark evaluates physical encoding via 8,000 simulation cases, finding alignment trades off against quantitative recoverability and retrieval improves video generation fidelity.

Yiqi Liu, Ruifeng Yuan, Yang Wang, Long Li and 6 more

Published Oct 2, 2026 · 0 citations · ▲ 44 on Hugging Face · Code ★ 4

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
88%Must read
?Must readVote to see the score

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer unifies streaming video perception, memory, and proactive response via shared generation, achieving top results on eight benchmarks with a 4B model.

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu and 20 more

Published Oct 1, 2026 · 0 citations · ▲ 232 on Hugging Face · Code ★ 169

100% Readers1 of 1 upvoted
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

KeyRec creates bounded visual memory via recent caches and structured event banks to enable efficient long-video and streaming understanding with only 10% of visual tokens, outperforming compressed baselines.

Zihan Chen, Xuejian Rong, Xiaojuan Wang, Boqing Gong and 3 more

Published Sep 26, 2026 · 0 citations · ▲ 11 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5
83%Must read
?Must readVote to see the score

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

JoyAI-VL-Interaction is an open 8B vision-language model that continuously decides whether to speak, stay silent, or delegate in real time, outperforming Doubao and Gemini across six real-world scenarios.

Dingyu Yao, Junhao Zhou, Chenxu Yang, Chuanyu Qin and 16 more

Published Jun 10, 2026 · 0 citations · ▲ 218 on Hugging Face · Code ★ 1,944

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Video-MME-v2 introduces a progressive tri-level benchmark with group-based non-linear evaluation that reveals substantial gaps between top models and human experts in video understanding.

Chaoyou Fu, Haozhi Yuan, Yuhao Dong, Yi-Fan Zhang and 15 more

Published Apr 6, 2026 · 0 citations · ▲ 231 on Hugging Face · Code ★ 362

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 2/5
65%Worth a look
?Worth a lookVote to see the score

MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues

MT-Video-Bench introduces a multi-turn dialogue benchmark that evaluates multimodal LLMs on holistic video understanding.

Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Moore Wang and 12 more

Published 2026 · 0 citations · Code ★ 21

100% Readers1 of 1 upvoted
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya YANG and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

A Training-Free Video Moment Retrieval Framework via Selection from Multiple Visual Prompts

Xun Mo, Shengtao Guo, Yongwei Nie, Fei Ma and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

AdShot: Benchmarking Multimodal Large Language Models for Video Advertisement Clipping

Wen Xie, Om Rastogi, Sai S Rangoju, Gijs Overgoor and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Static-Dynamic Disentanglement for Efficient Multi-Frame Vision-Language-Action Models

Weikang Qiu, Huashuo Lei, Tinglin Huang, Aosong Feng and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

CineMME: Benchmarking Fine-Grained Perception and Plot Reasoning in Multimodal Large Language Models

Jin Liu, Dawei Du, Yexiang Liu, Sijie Zhu and 6 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video LLMs

Jongseo Lee, Hyuntak Lee, Sunghun Kim, Sooa Kim and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

TeMPO: Frame-Causal Token Compression for Efficient Video Large Language Model

Yingxin Lai, Bo Xu, Yun-ze Pan, Zhiliang Zhu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

Breaking the Second Barrier: Sub-Second Timestamped Omni-Modal Captioning

Zhihe Yang, Xin LI, Hao Tan, Zhao Zhong and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

ActO: Extracting Action Representations from MLLM Embeddings for Video World Models

Runjia Li, Minghao Chen, Junyu Xie, Philip Torr and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

The VLM as Sensor: Bayesian Active Search for Long Video Understanding

Chong Tang, Sannara EK, Dirk Koch, Robert Mullins and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

Diagnosing and Correcting Bias in MLLM for Long Video Understanding

Xusheng Liang, Jianqiao Sun, Hao Zhang, Yulei Niu and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Learning Visual Speech Representations via Cross-Modal Distillation and Joint Face-Lip Modeling

Jingxuan Zhang, Genshun Wan, Jia Pan, Jianqing Gao and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Behavior Pack Optimization for Video MLLM Post-Training

Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang and 11 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

67%Highly rated
?Highly ratedVote to see the score

FrameRouter: Frame Budget Routing for Long-Video Understanding on Video-MME-v2 and Beyond

Bingjun Luo, Jialin Guo, Siqi Li

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

45%Niche pick
?Niche pickVote to see the score

Causal-VLM: Dense Causal Captioning in Videos

Asmar Nadeem, Mahrukh Awan, Muhammad Awais, Robert Dawes and 2 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

57%Worth a look
?Worth a lookVote to see the score

From Finding to Linking: Benchmarking and Advancing Cross-Long-Video Reasoning for Multimodal LLMs

Meng Luo, Zikang Zhou, Shanqing Xu, Shize Zhang and 4 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Iris: Empowering Video MLLMs with High-Frequency Pose Priors via Spatiotemporal Binding

Jiahang Zhang, Yushuo Guan, Yuanxing Zhang, Pengfei Wan and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

Prysma: Efficient Modality Adaptation for SLO-aware LLM-based Video Question Answering

Botao Zhu, Xiaoyi Fan, Yifei Zhu

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 2/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

CondenseVLA: Learnable History Condensation for Efficient Multi-Frame VLA

Feiyang Hong, Yujie Wei, Bo Zhao, Xiu Su and 2 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

VidHalluDoctor: Learning Video Differences to Mitigate Hallucinations in Vision-Language Models

紫云 戴, Yang Fu, Henghui Ding

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

AdaAlloc: Adaptive Visual Token Allocation for Long-Video Question Answering

Joungbin An, Kristen Grauman

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

MoRe-DVC: Motion Retrieval-Augmented Generation for Detailed Video Captioning

Vipul Baghel, Ravi Hegde

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Event-Centric Perception in Weak-Signal Physical Streams with Multimodal LLMs

Chi Xu, Mengdi Jin, Jiaxing Li, William I Atlas and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Training-free Focus-Ambient Retention for Memory-Efficient Video Large Language Models

Haozheng Zeng, Qi Bi, Gui-Song Xia

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

VideoSleuth: A Narrative-Centric Agentic System with Video-Audio Native MLLM for Long-Form Video Understanding

Conglin Li, Yang Li, Qi Zhang, Guanhua Chen and 2 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
80%Must read
?Must readVote to see the score

EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding

EVIDENT routes video temporal grounding through explicit visual entity evidence to improve cross-domain robustness via parameter-efficient adapters and distillation.

Geo Ahn, Jiwook Han, Youngrae Kim, Joonseok Lee and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
88%Must read
?Must readVote to see the score

AdaCodec: A Predictive Visual Code for Video MLLMs

AdaCodec uses predictive visual codes to send full reference frames only when unpredictable, cutting video MLLM tokens by 7x while improving long-video benchmark scores and reducing latency.

Haowen Hou, Zhen Huang, Zheming Liang, Qingyi Si and 7 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 7 on Hugging Face

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

FLARE introduces a long-video audiovisual retrieval benchmark with simulated user queries, revealing caption-based performance fails to transfer and audio-language alignment remains a bottleneck.

QiJie You, Hao Liang, Mingrui Chen, Bohan Zeng and 3 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
83%Must read
?Must readVote to see the score

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

VideoOdyssey benchmarks ultra-long video understanding via continuous certificates averaging 16 minutes, revealing MLLM failures in continuous reasoning and omni-modal perception.

He Haichen, Jiayi Zhou, Sifeng SHANG, Yihan Hu and 2 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

ViDiC: Video Difference Captioning

ViDiC introduces a video difference captioning task and ViDiC-1K benchmark that reveals large multimodal models struggle with fine-grained comparative video perception.

Jiangtao Wu, Shihao Li, Zhaozhou Bian, Jialu Chen and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 29 on Hugging Face · Code ★ 15

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 3/10
strict 1/5
74%Highly rated
?Highly ratedVote to see the score

K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

K9-Bench evaluates multimodal LLMs on dog videos via 5,000 QA pairs and finds limited zero-shot reasoning over long-horizon canine interactions.

Khush Attarde, Yusuf Ali, Megha Thukral, Divye Bhutani and 2 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 5/5
medium 4/10
strict 0/5
86%Must read
?Must readVote to see the score

VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

VidVec extracts intermediate-layer MLLM embeddings for video-text retrieval via calibration and text-only alignment, achieving state-of-the-art zero-shot results without video fine-tuning.

Issar Tzachor, Dvir Samuel, Rami Ben-Ari

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026 · ▲ 24 on Hugging Face · Code

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

CaC advances video reward models via hierarchical spatiotemporal concentrating, improving fine-grained anomaly accuracy by 25.7% and reducing generated-video anomalies by 11.7%.

Jiyuan Wang, Huan Ouyang, Chunyu Lin, Dewen Fan and 14 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5
83%Must read
?Must readVote to see the score

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

DELTAVID improves video MLLM fine-grained spatiotemporal perception by training cross-video difference spotting, boosting performance across multiple video understanding benchmarks.

Yankai Yang, Yancheng Long, Bin Wen, Fan Yang and 3 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

IPIBench evaluates interactive proactive intelligence of MLLMs on continuous video streams, revealing unstable proactive triggering and weak reactive-proactive coordination, while IPI-Agent improves both via temporal gating.

Jinzhao Li, Yinuo Chen, Wenxuan Song, Yijia Lei and 4 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Motion-o: Trajectory-Grounded Video Reasoning

Motion-o introduces trajectory-grounded video reasoning via explicit motion chains of thought, improving trajectory-faithful reasoning across benchmarks without architectural changes.

Bishoy Galoaa, Shayda Moezzi, Xiangyu Bai, Sarah Ostadabbas

Atlanta Poster Session 1, Wed, Dec 9, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 2 on Hugging Face · Code ★ 8

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
86%Must read
?Must readVote to see the score

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video

SoSVQA extracts 10K pedagogically structured QA pairs from children's video to train VLMs via GRPO, yielding major reasoning gains on NExT-QA, Video-MME, and MotionBench that match proprietary systems despite far less training data.

Bishoy Galoaa, Xiangyu Bai, Sarah Ostadabbas

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 1 on Hugging Face

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
89%Must read
?Must readVote to see the score

MedHorizon: Towards Long-context Medical Video Understanding in the Wild

MedHorizon benchmarks long medical video understanding via sparse evidence retrieval and multi-hop reasoning, with top models reaching only 41.1% accuracy.

Bodong Du, Bowen Liu, Yang YU, Xinpeng Ding and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
80%Must read
?Must readVote to see the score

Harnessing Streaming Video in the Wild

Streaming-Train-248K and Streaming Harness adapt VLMs to real-time video streams with proactive interaction, 12-hour memory, and sub-second latency.

Dingyu Yao, Shuhuan Gu, Qingyi Si, Junhao Zhou and 7 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

SLVMBench: Skill Learning from Video Memory

SLVMBench evaluates video-LLMs on learning skills from long video streams and applying them in real time, revealing substantial performance degradation.

Yudong Yang, Guangzhi Sun, Yixuan Li, Wei Li and 2 more

Paris Poster Session 2, Wed, Dec 9, 5:00 PM–7:00 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
76%Highly rated
?Highly ratedVote to see the score

Not Another Text Benchmark: Putting the “Visual" Back in Visual Question Answering for Large Video Models

Three visual benchmarks for video understanding expose large video model weaknesses when reasoning through visual queries instead of text options.

Rwiddhi Chakraborty, Yinong O Wang, Cheng Zhang, Fan Bai and 5 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
10/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 10 of 20 reviewers recommend it
lenient 4/5
medium 5/10
strict 1/5
88%Must read
?Must readVote to see the score

GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs

GridProbe scores frame evidence via frozen VLM answer-space probing and adaptive selection to reduce long-video attention costs with minimal accuracy loss. It matches monolithic baselines on Video-MME-v2 at 3.36x lower compute and Pareto-dominates baselines on LongVideoBench.

Mohamed Eltahir, Ayash, Ali Habibullah, Tanveer Hussain and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 2 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 1/5
86%Must read
?Must readVote to see the score

CoRDS: Coreset-Based Representative and Diverse Selection for Streaming Video Understanding

CoRDS selects coreset subsets of key-value caches to cover accumulated visual history geometry, improving streaming video understanding with fixed memory budgets.

Ailar Mahdizadeh, Puria Azadi Moghadam, Muchen Li, Xiangteng He and 1 more

Paris Poster Session 6, Fri, Dec 11, 2:30 PM–4:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 1/5
86%Must read
?Must readVote to see the score

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

MVVBench benchmarks 4D multi-view video reasoning requiring joint cross-view temporal integration, finding vision-language models fail due to temporal mis-localization and cross-view identity breaks, with inference-time scaffolding and reinforcement learning yielding substantial gains.

Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim and 2 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
80%Must read
?Must readVote to see the score

ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

ReSCUE enables simultaneous sign language translation on continuous unsegmented video via inference-aware training, stabilized re-translation, and online sentence commitment, achieving low-latency quality near offline oracles.

Sihan Ren, Gaozheng Li, Yuanshang Quan, Yiming Qin and 4 more

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 0/5
89%Must read
?Must readVote to see the score

EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs

EchoPrune treats redundant video tokens as temporal echoes and prunes them via query relevance and reconstruction error, letting VideoLLMs process up to 20x more frames for +8.6% accuracy and 5.6x faster prefilling.

Jiameng Li, Minye Wu, Jiezhang Cao, Aleksei Tiulpin and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 10/10
strict 2/5
91%Must read
?Must readVote to see the score

Don't Pause! Every prediction matters in a streaming video

SPOT-Bench introduces multi-turn proactive queries and Timeliness-F1 to evaluate real-time streaming video perception; AsynKV improves streaming behavior by scaling compute during dead-time to match offline detection.

Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener, Yale Song and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
17/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 17 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 3/5
78%Highly rated
?Highly ratedVote to see the score

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

A diagnostic taxonomy maps AVLM failure signatures to targeted development interventions, enabling traceable industry-scale video moderation system improvements.

Shuchang Ye, Jinqiang Yu, Zhujun Xiao, Yajing Kong and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
86%Must read
?Must readVote to see the score

A Closer Look at Dynamic Scene Graph Generation in the Era of Multimodal Large Language Models

Revisiting dynamic scene graph generation via top-down reasoning, temporal relation sets, and importance-aware finetuning achieves state-of-the-art results.

Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh W Peng, Adriel Kuek and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 1/5
88%Must read
?Must readVote to see the score

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

VideoLLMs lose temporal representations across layers, so Temporal Activation Injection reinforces fading divergence at inference to improve video reasoning without training.

Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 5 on Hugging Face · Code ★ 2

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 4/5
medium 9/10
strict 2/5
80%Must read
?Must readVote to see the score

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

StreamGaze introduces a benchmark for evaluating gaze-guided temporal and proactive reasoning in streaming videos, revealing large performance gaps between state-of-the-art MLLMs and humans.

Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton, Ryan Rossi and 5 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 8 on Hugging Face · Code ★ 28

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
86%Must read
?Must readVote to see the score

BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation

BusterX introduces GenBuster-200K, GenBuster-Bench, and an MLLM baseline that detects AI-generated video via reasoning chains, outperforming leading models in accuracy and explanation quality.

Haiquan Wen, Yiwei He, Zhenglin Huang, Tianxiao Li and 6 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · Code ★ 51

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
78%Highly rated
?Highly ratedVote to see the score

COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection

COPRA adapts frozen vision-language models via reinforcement learning-generated, input-specific parameter updates for video anomaly detection, improving cross-domain performance and generalizing to video understanding tasks.

Darryl Cherian Jacob, Xinyu Liu, Kai Wang, Pan He

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
78%Highly rated
?Highly ratedVote to see the score

V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models

V-CAST prunes video tokens via curvature-guided temporal budgets and dual-anchor spatial selection, achieving 98.6% original performance with 86.4% latency.

Xinying Lin, Xuyang Liu, Yiyu Wang, Teng Ma and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
71%Highly rated
?Highly ratedVote to see the score

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

ThinkJEPA combines dense JEPA dynamics with sparse VLM reasoning via dual pathways to improve long-horizon latent world modeling and trajectory prediction.

Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan and 4 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026 · ▲ 21 on Hugging Face · Code ★ 58

– ReadersNo votes yet
7/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 7 of 20 reviewers recommend it
lenient 4/5
medium 2/10
strict 1/5
80%Must read
?Must readVote to see the score

CAM: Question Answering on Entity-Centric Videos with Continuous Extraction and Adaptive Querying

CAM uses continuous knowledge graph extraction and adaptive multi-method querying to answer entity-centric video questions, improving accuracy by up to 23 points over baselines.

Yizhou Tian, Zizhe Chen, Shiyuan Deng, Garry YANG and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

A benchmark of likely public-domain Hollywood films tests multimodal models on narrative understanding, finding vision-language models near chance and audio-visual models below human performance.

David Bamman, Kent K Chang, Allison Cooper, Juishan Hsu and 6 more

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 2/5
88%Must read
?Must readVote to see the score

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

MAGIC-Video unifies episodic, semantic, and visual content via a multimodal memory graph and narrative chain for agentic ultra-long video reasoning, outperforming prior agentic systems by up to 10.1 points.

Jiazheng Li, Chi-Hao Wu, Yunze Liu, Kaize Ding and 2 more

Atlanta Poster Session 5, Fri, Dec 11, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
15/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 15 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 2/5